Orelon logoOrelon
요금

Prompting for AI Images and Video: A Practical System

2026년 9월 15일 · Orelon Team 작성

AI 동영상 템플릿 둘러보기

영감을 위해 커뮤니티 창작물 몇 개를 둘러본 다음, 템플릿을 열어 Orelon에서 계속 만들어 보세요.

A practical prompting system for AI images and video: prompt structure, motion control, references, review criteria, and iteration habits that scale.

Most disappointing AI generations are not model failures. They are description failures. The model rendered exactly what the words allowed, and those words left so much open that the only sensible answer was an average one: a generic face, flat light, a meaningless background, motion that drifts. Sharpen the description and the same model suddenly behaves like a different product.

This is a practical system for prompting still images and moving shots. It covers how to structure a prompt, how to control motion and continuity, how to iterate without losing an afternoon, how to review results with consistent criteria, and how to store what works so a team can reuse it. The principles are tool-agnostic, and the workflow maps directly onto Orelon, where a still can be approved and then animated.

Why a clearer description beats a better model

A model cannot infer intent. Ask for 'a warrior in a forest' and it must guess the century, the season, the weather, the camera, the framing, the mood, and the lighting. With no information to work from, it returns the statistical middle of all those possibilities — competent, forgettable, and nearly impossible to art-direct.

Treat a prompt as a compressed creative brief. A good brief does not describe every pixel; it establishes intent, constraints, and a reference point so the person executing it can make sound decisions on their own. A prompt should state what matters, exclude what would ruin the shot, and leave the rest to the model.

There is an economic argument too. Every generation spends time and attention. Two extra minutes spent writing usually removes several rounds of re-rolling, and the prompt doubles as documentation: when someone asks months later for the same look but at dawn, the saved prompt is how you deliver it instead of guessing.

Prompting is also a transferable skill. Interfaces change; the vocabulary of shot size, lighting direction, motion type, and material behaviour does not. Learn the vocabulary once and you can move between tools without starting from zero.

The six blocks of a strong image prompt

Most strong prompts contain the same functional parts, and they are easiest to write in this order.

Block What it carries Short example
Subject and intent who or what, action, wardrobe, emotion, purpose 'a bicycle courier mid-delivery, soaked and determined, editorial documentary'
Composition shot size, camera height, placement, negative space 'medium shot, eye level, subject on the right third, clean wall on the left'
Lens and light focal length, depth of field, light direction and quality '35mm, shallow depth of field, hard side light through blinds'
Style and finish medium, era, grade, grain, texture 'tungsten-balanced look, teal shadows, subtle halation, light grain'
Exclusions artefacts and elements that must not appear 'no text, no logos, no extra limbs, no lens flare'
Emphasis what to prioritise if the model must compromise 'the coat is the hero detail; the background can stay soft'

Subject and intent

Vague subjects produce vague pictures. Compare 'a lighthouse keeper' with 'a lighthouse keeper in his sixties, salt-stiffened beard, oilskin coat, holding a brass lamp in a doorway at dawn, tired but resolute.' The second version delivers wardrobe, prop, setting, time of day, and emotional register in one pass.

Add an intent phrase that names the purpose: 'editorial portrait', 'product hero shot on a seamless backdrop', 'key art for a thriller'. Intent steers a model toward the visual conventions of a genre, which is often more effective than a stack of adjectives.

Composition

Framing is the fastest way to make an image look deliberate rather than accidental. Name the shot size, the camera height, and where the subject sits. Mentioning negative space is a professional habit: 'subject left, clean wall to the right' leaves room for titles, captions, or a logo later.

Lens and light

Lens language does real work. '85mm at a wide aperture' implies a compressed background and shallow focus; '24mm stopped down' implies environmental context and deep focus. Two consistent lens cues will make your images feel shot rather than assembled.

Light deserves more words than anything else. Specify direction, quality, colour, and source. One sentence such as 'warm practical lamp from the left, cold window fill from behind, soft rim light on the coat' outperforms five style adjectives.

Style and finish

Style is where prompts most often fail, because people stack incompatible references. Hyperrealism, anime, clay, and 3D render in one prompt tells the model only that you are unsure. Pick one medium, one era, one finish. If you want a photographic result, film-stock references are efficient: a stock name implies palette, contrast curve, and grain in a few syllables. If you want an illustrated result, name the technique, such as 'hand-inked line work with flat spot colour.' Naming living artists raises ethical and legal questions most teams would rather avoid.

Exclusions and emphasis

Exclusion lists remain one of the highest-leverage habits in prompting. Typical offenders: text, watermarks, duplicated limbs, warped hands, plastic skin, oversaturated colour, unrequested lens flares, background clutter. Keep the list short and specific — ten exclusions aimed at problems you actually see beat a hundred generic ones.

For emphasis, put the most important element first and avoid weighting two contradictory ideas equally. If the coat matters more than the environment, show it through order and repetition rather than aggressive weighting that distorts everything else.

Worked example: from a vague idea to an approved key frame

Start with a rough line: 'a chef in a kitchen, dramatic.'

Round one returns something plausible and instantly forgettable. Rewrite with intent: 'editorial food documentary, a chef in her forties in a small professional kitchen before service, white jacket with rolled sleeves, flour on her forearms, standing at a steel counter, quietly focused.'

Add framing and light: 'medium shot, eye level, subject left of centre, stainless steel and hanging pans behind, hard window light from camera right, deep shadows on the left wall, practical heat lamp glowing warm in the background.'

Add lens and finish: '40mm, moderate depth of field, slightly cool white balance with warm lamp highlights, fine grain, no halation.'

Add exclusions: 'no text, no logos, no extra hands, no plastic-looking skin, no lens flare.'

The second version takes about ninety seconds longer to write and returns something you can build a sequence around. That trade is the entire discipline in miniature.

The same method for product shots

Product work rewards precision over atmosphere. Replace mood words with material words: brushed aluminium, matte cardboard, condensation on glass, visible weave. Specify the background treatment explicitly — seamless sweep, textured stone, gradient falloff — and state whether the surface should be reflective or dry. If a label must be legible, say so, and generate at a size large enough that lettering survives.

The same method for portraits and characters

Characters need identity anchors more than adjectives. Age range, face structure, hair length and texture, wardrobe silhouette, and posture do more for consistency than any list of moods. Build one approved reference frame, then describe only what changes between shots. This single habit removes most of the drift people blame on the model.

The same method for landscapes and establishing shots

Wide shots are compositional problems. Decide where the horizon sits, how much sky you want, and which element leads the eye into depth — a road, a river, a row of lights. Name the time of day and the weather, because atmosphere communicates scale better than any detail. Then add one human-scale element for reference: a lone figure, a parked car, a lit window.

Iteration: running a loop that teaches you something

Prompting is an editing practice, not a lottery. A reliable loop looks like this: write the full prompt, generate a small batch, pick the closest result, then change exactly one variable. If the light is wrong, change the light and nothing else. If the framing is wrong, change the framing and nothing else.

Keep a scratchpad with each variant and a one-line note about what changed. After ten rounds you will have a personal map of what actually moves the output, calibrated to your taste and your project. That map is worth more than any list of prompts written by someone else.

Batch size matters. Three to six variants per round is enough to reveal a trend without drowning in options. Two to four per camera setup is usually right for video. When a strong result appears, save the prompt immediately; do not trust memory after the next experiment.

Decision criteria while iterating

  • Does the frame read at thumbnail size? If the subject disappears when you shrink it, composition is the problem, not detail.
  • Is the light doing one clear thing? Two competing sources usually mean an unclear brief.
  • Does the eye land where you intended?
  • Would this frame survive as one shot inside an edited sequence, alongside its neighbours and the same grade?
  • Can you reproduce it? If you cannot explain why it worked, you cannot repeat it.

Video prompting: what changes when a frame becomes a shot

Video inherits everything above and adds a harder requirement: the model must keep a world coherent while it changes. A still can hide uncertainty inside one frame; a shot has to survive time. Your prompt therefore needs to state what is in frame, what is allowed to change, what must stay identical, and how the camera behaves.

Temporal consistency

State continuity anchors explicitly: wardrobe, hairstyle, props, light colour, time of day, identity. A sentence such as 'the same character, same coat and haircut throughout, lighting unchanged' is not filler — it is a constraint that reduces flicker and drift.

Physics and weight

Realism in generated motion comes mostly from material behaviour. Describe how cloth folds, how water refracts, how smoke drifts, how dust catches light, how hair lifts in wind. Add weight cues: 'heavy gait', 'shoulders shift with each step', 'the bag swings and settles.' These give the model something to simulate rather than something to decorate.

Watch for contradictory physics. Slow motion plus fast, energetic movement pulls in two directions, as does a calm ocean plus crashing waves. Choose one physical reality and stay consistent about scale, speed, and mass.

Camera language that reads

Camera moves are among the strongest controls available, and the vocabulary comes from film grammar: dolly in, truck left, crane up, orbit, handheld follow, slow push-in, rack focus, whip pan. Pick one move per shot and give it a speed and duration: 'slow two-second push-in, no other camera movement.' Combining three moves usually produces a wobble that reads as an error rather than a style choice.

One action per shot

Walking to a window, picking up a cup, turning, and smiling is four shots compressed into one prompt, and the model will either rush them or smear them. Write the single beat you need — 'she walks slowly to the window and stops' — then generate the next beat separately. Sequences are built in the edit, not inside one prompt.

References and style locking across a sequence

Reference frames are the most reliable consistency tool available. Instead of describing a character again in every prompt, generate or upload one strong frame and let the model treat it as the anchor. The prompt then only describes what changes: 'same character and lighting, now walking through the doorway, camera follows.'

Style locking follows the same logic. Build one style sentence — lens, lighting, grade, grain — and paste it at the end of every prompt in a sequence. Because it never changes, the model treats it as a constant, and shots start to feel like they belong to the same film rather than the same folder.

This is also where stills and motion reinforce each other. Stills are quick to explore and easy to approve; once a look is signed off, animate from it. Building the key frame first in Orelon's image workspace, then moving it into video generation, is far cheaper than discovering a look is wrong after twenty animated attempts. You can see the flow in the image workspace and the video workspace.

Worked sequence: eight shots from one style sentence

Imagine a short film: a night market in the rain, one character, a lost phone, a quiet resolution. Start with the constant: 'shot on fast 35mm, tungsten and neon mix, wet asphalt reflections, cool shadows with warm practical highlights, subtle grain.' That sentence goes into every prompt below. Only the shot-specific line changes.

  1. Establishing wide: 'rain-soaked night market alley, steam from food stalls, crowds blurred, camera drifts slowly forward, no shake.'
  2. Insert: 'close-up of a phone screen glowing on wet pavement, raindrops striking the glass, shallow focus, static camera.'
  3. Character introduction: 'a young woman in a dark green jacket crouches into frame, reaches for the phone, camera holds at eye level.'
  4. Detail: 'her hand lifts the phone, water runs off the case, neon reflections ripple across her knuckles.'
  5. Reaction: 'medium close-up, she looks up past the camera, rain on her face, slow push-in, expression shifting from worry to relief.'
  6. Environmental beat: 'vendors continue around her, out of focus, steam and umbrella silhouettes, camera tracks left at walking pace.'
  7. Turn: 'she walks away from camera into the crowd, back to lens, lights streaking on wet ground, handheld follow.'
  8. Button: 'wide shot from the far end of the alley, she disappears into the crowd, rain continues, static camera.'

Notice what is missing. No prompt describes plot. Each has one camera move, one action, and the shared style sentence. Continuity comes from the repeated style line plus a reference frame, not from re-describing the character with fresh synonyms every time.

Templates, naming, and a reusable prompt library

Once a project works, turn it into a template. A serviceable structure for a reusable shot prompt:

shot type + subject and action + environment + lighting + lens and camera + style sentence + exclusions

Store templates at the project level and name individual prompts so the name explains them at a glance: market_shot03_pushin_v2. Add a one-line note about what changed from the previous version. That feels bureaucratic for a solo creator and becomes essential the moment two people touch the same sequence.

Keep a small library by use case — establishing shots, product hero frames, character close-ups, transitions — rather than one enormous file. Collections are searchable under deadline. If you would rather adapt than start from a blank page, browse Orelon's prompt library and copy the structure, not the content.

Maintain separate versions of each template for widescreen and vertical, because reframing after the fact is always worse than composing correctly the first time. Ready-made starting points help here too, and Orelon's templates give you a scaffold to adapt rather than a blank page.

For teams, treat templates as shared infrastructure: a naming convention, a one-line note per version, and one agreed place to store approved style sentences. When a new person joins mid-project, the library is the onboarding document.

Review criteria, common mistakes, and a pre-flight check

Good review is systematic, not mood-based. Before approving, ask five questions: does it match the approved style sentence, does it read at small size, is the light doing one clear thing, is the subject where the composition needs it, and would it cut cleanly against the shots either side? A frame that fails two of these is cheaper to regenerate than to repair in post.

Common mistakes that waste hours:

  • Stacking conflicting styles. One medium, one finish.
  • Describing a story instead of a frame.
  • Skipping exclusions, then fighting the same artefact forever.
  • Changing five variables at once, then learning nothing.
  • Substituting quality words for direction. Generic praise adds little next to specified light, lens, and framing.
  • Forgetting the delivery format until after generation.
  • Assuming continuity. Models do not remember your character between prompts; supply the reference or restate the anchors.
  • Not saving the winning prompt.
  • Accepting a frame that only works at full size. It will fail on a phone screen and in an edit.

A pre-flight check takes twenty seconds and saves whole sessions: subject named, intent stated, framing chosen, light described in one sentence, style sentence attached, exclusions listed, aspect ratio set, one action and one camera move for video.

FAQ

How long should a prompt be? Long enough to remove ambiguity, short enough to stay coherent. For stills, two to four sentences covering subject, framing, light, and style usually beat a paragraph of adjectives. For video, add one action and one camera move, then stop.

Do exclusions still matter? Yes. They are especially effective against recurring artefacts such as text, watermarks, extra limbs, and plastic skin. Aim them at problems you actually see rather than pasting a long generic block.

How do I keep a character consistent across shots? Anchor with a reference frame, repeat a fixed set of continuity descriptors, and hold lighting, wardrobe, and lens constant across the sequence. Approve the reference before animating anything.

What is the fastest way to improve? Iterate on one subject for twenty generations, changing one variable at a time, and write down what moved the result. Narrow deliberate practice beats browsing collections of prompts.

Should vertical video be prompted differently? Yes. Vertical framing favours tighter shots, centred or stacked composition, and fewer background elements. Keep the subject centred enough to survive cropping.

How many variants per shot? Three to six per round for stills, two to four per camera setup for video. More than that usually means the prompt is underspecified rather than that you need more options.

Can one prompt produce a whole sequence? Not reliably. Generate individual shots that share a style sentence and a reference frame, then assemble them in an edit where you control timing and rhythm.

How do I judge between two decent results? Prefer the one that matches the approved style sentence, reads clearly at small size, and needs the least repair. Then save the prompt that produced it, with a note on what made it work.

Do I need to know film theory? No, but ten minutes of vocabulary pays for itself. Shot size, lighting direction, and camera-move names give you precise words for things you already recognise, which is exactly what a prompt needs.

Put the system to work

Strong prompts are less about hidden syntax and more about clear thinking: a specified subject, deliberate light, credible physics, one action, one camera move, a constant style sentence, and a short list of exclusions. Everything else is iteration, and iteration is where taste gets built.

If you want to apply this immediately, start with one key frame, approve the look, then animate it. Orelon is built for that flow — cinematic ideas in motion, from a first still to a finished sequence, with prompt and template collections that keep your work consistent as it grows. Explore the blog for more workflow guides, or open the workspace and rewrite your weakest prompt with the six blocks above.