Learn how to write prompts for dating-app clips, viral social video, and award-worthy AI short films, with workflows, examples, and fixes.
Most AI video fails for a boring reason: the prompt described a mood instead of a shot. "Woman looking at her phone, sad, cinematic" hands a model almost nothing to build from. "A woman in her early thirties in a mustard cardigan sits on the edge of a bed at dawn, phone screen lighting her face from below, shallow depth of field, slow push-in" hands it a scene. The distance between those two prompts is the entire craft.
This guide covers three things that look unrelated and are not: prompts for relationship and dating-app-style social clips, vertical video built to survive a scroll, and short films that hold up when a jury watches them twice. All three reward the same discipline — deciding what the camera sees before asking a model to render it.
Why vague prompts produce forgettable footage
A video model is a literal collaborator. It does not infer that "sad" means a particular posture, a particular light, or a particular hour. Every vague noun is a decision handed to a random number generator, and random decisions compound until the footage looks like everyone else's.
Specificity is density, not length
A strong prompt is packed, not long. It names subject, action, camera behavior, light, and duration, then stops. Prompts that run three hundred words usually contain three competing scenes and produce a soft average of all three. If you cannot summarize your prompt in one sentence aloud, it is probably two shots.
Three viewers, three kinds of pressure
- The scroller decides in under two seconds whether the first frame is legible at thumbnail size.
- The collaborator or client needs a look that repeats across a series, not one lucky clip.
- The jury watches for intent, because technical polish stopped being rare a while ago.
Same skill, three pressure tests. Write for the strictest audience you actually have, and the others tend to come along on their own.
The anatomy of a shot prompt that holds attention
Treat prompt writing as a shot list with the camera department attached. Four layers do most of the work.
Subject, action, and one source of tension
Name a person, an age range, a wardrobe detail, and what they are doing in the first two seconds. "A man in his forties, half-tucked shirt, reading a message and then setting the phone face-down" is playable. "A man checking his phone" is wallpaper.
Then add a single tension cue. It can be physical — a paused thumb, a half-step backward — or environmental, like a second coffee cup that is not his. Tension is what stops the thumb, and it costs about six words.
Camera, lens, and a single movement
Camera language separates footage that reads as directed from footage that reads as generated:
- Framing: extreme close-up, medium close-up, wide establishing shot, over-the-shoulder
- Lens feel: 24mm for intimacy with slight distortion, 50mm neutral, 85mm compression for isolation
- Movement: slow push-in, handheld follow, locked-off tripod, lateral dolly, orbit
- Height: eye level, low angle, slight high angle looking down at hands
Pick one movement per generation. Two movements in one prompt produce drift that audiences read as instability rather than style.
Light with a nameable source
Lighting is where most prompts dissolve into the word "cinematic." Replace it with a source: a practical lamp on camera left, cold window light behind the subject, the phone screen as the only key. Then state one color intention — muted teal shadows with warm skin, or sodium-vapor orange against wet asphalt.
Time of day and weather are efficient shorthand. Golden hour, overcast midday, blue hour after rain. Models handle those conditions far better than abstract adjectives, because weather and hour imply a physical lighting setup rather than a feeling.
Duration, pacing, and aspect ratio
Decide clip length before generating. A three-second insert of hands deserves one action and no camera move. A ten-second beat can carry a movement and a small emotional change. Lock the aspect ratio early — 9:16 vertical for feeds, 2.39:1 for a wide cut — because reframing after generation always costs detail. Pacing is a writing decision, not an editing rescue.
A working prompt for dating-app and relationship clips
Relationship-adjacent content is a good test case because it runs on micro-emotions: anticipation, misread signals, the pause before a reply. Here is a structure that reliably produces usable takes.
Subject: woman, early thirties, oversized knit cardigan, hair half up
Action: reads a message, exhales through the nose, types two words, deletes them
Setting: kitchen counter, morning, cereal bowl pushed aside
Camera: medium close-up, 50mm feel, slow push-in, eye level
Light: cold window light from camera left, warm pendant lamp behind, mixed color temperature
Grade: muted greens, warm skin tones, soft contrast
Duration: 5 seconds, single action, no cuts
Aspect: 9:16 vertical
It works because it names a beat with a beginning, middle, and end: read, exhale, delete. The delete is the story. Remove it and you have a stock clip.
Two variants worth building from the same skeleton:
- The mistaken match. Same structure, action becomes a double-take at the screen and a short laugh. Warm grade, golden hour, 24mm feel for closeness.
- The aftermath. Wide shot, subject small at the end of a long table, phone face-down, single overhead practical. Eight seconds, locked-off camera.
Generate three to five variations per prompt with different seeds, then choose on expression rather than resolution. Facial timing is the least controllable variable in AI video, so plan a selection step instead of expecting one perfect take.
Vertical-first composition for feed video
A vertical frame is a portrait, not a cropped landscape. Compose for it. Keep eyes in the upper third, leave headroom for captions, and place visual interest in the vertical center, because a thumb covers the bottom third.
Rules that hold up across platforms:
- One idea per clip. A vertical viewer gets one takeaway before the swipe.
- Text-safe framing. Reserve a band through the middle for on-screen text so it never collides with a face.
- Shorter holds. A four-second shot feels long in a feed and short in a short film. Adjust per destination.
- Sound-first hooks. Design the first frame around a sound cue you can place in the edit, so audio and image reinforce each other.
If you are producing a series, build the template once — framing rules, caption position, grade, font — and reuse it. Starting from a structured video template is faster than re-deciding composition for every clip, and it keeps an audience from feeling that each post came from a different channel.
Keeping one character recognizable across a series
Consistency is the difference between a channel and a pile of clips. Three mechanisms do the heavy lifting.
Character bible. Write down age, face shape, hair, wardrobe palette, and two signature accessories. Keep it to one page. Paste the relevant lines verbatim into every prompt. Rewording a description between shots is the most common cause of a face changing mid-series.
Reference frames. Generate a clean, front-lit portrait first, approve it, and use it as the visual anchor for later shots. Some tools accept an image reference; others respond better to an extremely detailed description repeated exactly. Test both and keep what your pipeline handles best.
Scene continuity. If a series lives in one apartment, lock the layout: kitchen left, window behind the table, hallway visible in the background. Repeat that spatial description even in close-ups, so background details stay stable between shots.
When a sequence needs a still to anchor motion, generating it as an AI image first gives you a reference to match instead of a description to guess at. It also gives collaborators something concrete to approve before any motion exists, which shortens review cycles considerably.
A production pipeline from beat sheet to final cut
A repeatable order saves more time than any single prompt trick.
Write beats before prompts
List six to twelve beats in plain language. "She reads the message." "She rewrites it." "She sends it." Only after the beats work do they become shot prompts. Skipping this step is why so many AI projects end with beautiful footage and no story.
Approve a look frame
Produce one still that fixes light, palette, and wardrobe, and approve it before generating motion. Changing the look after ten clips exist means regenerating all ten, usually under deadline pressure.
Prompt in shot-sized units
One action, one camera move, one light setup per generation. Slower per prompt, much faster overall, because you rarely discard a take over a single broken element such as a warped hand or a drifting background.
Run a selection pass
Generate three to five takes per shot and judge on expression and motion quality. Resolution can be handled later; a dead-eyed performance cannot.
Assemble to the beats
Cut to the beat sheet, not to the prettiest clips. The most common failure in AI short films is a gorgeous sequence with no throughline.
Fix rhythm with sound early
Add ambience and one music bed before the first full assembly. Sound exposes pacing problems that are invisible on a silent timeline, and it makes decisions about trimming far easier.
Grade once, consistently
Apply one grade across every clip. Consistent color hides small continuity differences between generations better than any other single fix.
Archive prompts next to clips
Save the final prompt text beside the clip it produced. Your durable asset is not the clip; it is the prompt that reliably generates it. Group prompts by function — establishing shot, reaction, insert, transition, closing beat — and note the lens, light source, and duration that worked. A curated prompt library turns a creative process into a repeatable one and is the fastest way to halve generation time on the next project.
What festival juries actually look at
Selection is not a rendering contest. Juries watch hundreds of technically clean submissions, so technical polish stops differentiating almost immediately. What remains is intent: does the film know what it is about, and does every choice serve that?
Plan around a few structural realities:
- Runtime discipline. Short means short. Many competitions publish tight limits, and six focused minutes beat fifteen loose ones almost every time.
- Disclosure rules. Competitions increasingly ask whether generative tools were used and how. Read the rules for every festival you enter and follow the disclosure instructions exactly, whether that is a form field, a written statement, or an on-screen note. Policies for AI-assisted work keep changing, so check the current published criteria instead of relying on what you remember from a previous season.
- Story before spectacle. A one-room film about two people avoiding a conversation often outperforms a sweeping generated landscape, because the audience has something to hold onto.
- Sound as authorship. An original score, considered ambience, and clean dialogue handling read as craft decisions. Borrowed audio reads as filler.
A practical move: choose two or three festivals whose deadlines you can realistically hit, work backward from those dates, and treat the paperwork as part of production rather than an afterthought. Filmmakers who plan the submission like a shoot day are the ones who actually finish.
Choosing the right model for the shot
Not every generation deserves the same engine, and the honest way to choose is to test with your own material rather than a demo reel.
A workable decision framework:
- Faces and reactions. Prioritize whatever holds identity and micro-expression best. Run your reaction prompt three times and compare eyes and mouth timing.
- Motion-heavy beats. Look for coherent physics on walking, doors, water, and fabric. If limbs smear at mid-motion, the take is unusable no matter how good the still frame looks.
- Long unbroken takes. Test a single eight-second action with one camera move. Long takes expose drift faster than anything else.
- Texture and product inserts. Shallow-focus detail shots are usually the easiest win, so test them last and spend your comparison time on the hard shots.
Score each candidate on expression stability, motion coherence, prompt adherence, and iteration speed — how fast you can get three takes back. Iteration speed quietly matters most, because you will generate dozens of variations before a project is finished. If you want a starting point for comparison, browse Seedance 2.5 examples and the alternatives hub to see where different tools are strongest before you commit a week of work to one pipeline.
Common mistakes and how to fix them
| Mistake | What it looks like | Fix |
|---|---|---|
| Adjective stacking | "Epic, moody, cinematic, dreamy" | Replace each adjective with a physical source: a lamp, a window, an hour |
| Multiple actions in one prompt | The subject drifts or morphs midway | One action, one camera move, one light setup |
| No duration stated | Clips that end before the beat lands | State seconds and name the single action |
| Reworded character description | The face shifts between clips | Copy the character bible lines verbatim |
| Silent rough cuts | Pacing problems hidden until late | Add scratch audio before the first assembly |
| Skipping disclosure rules | A submission rejected on a technicality | Read each festival's rules and follow them exactly |
| Generating before deciding | Hundreds of clips with no story | Decide the story first, then generate only what the edit needs |
Worked example: a forty-second vertical piece
Imagine a forty-second vertical short about someone deciding whether to reply.
- Beat 1 (0:00–0:04). Close-up, thumb hovering over a keyboard. No camera move. Cold window light, muted palette.
- Beat 2 (0:04–0:09). Medium close-up, push-in, she types, deletes, looks up. A warm lamp enters frame and motivates the next scene.
- Beat 3 (0:09–0:15). Wide, kitchen at golden hour, subject small in frame, kettle sound cue.
- Beat 4 (0:15–0:24). Insert, phone face-down on the counter, shallow focus, faint vibration suggested in sound.
- Beat 5 (0:24–0:33). Reaction shot, half-laugh, 85mm compression, locked off.
- Beat 6 (0:33–0:40). Return to the opening frame, thumb moves, screen brightens off-camera. Hold two seconds before the cut.
Six shots, one wardrobe change maximum, one grade, one music bed. Every prompt names a single action, a single movement, and a light source. That is the entire recipe, and it scales from a feed post to a ten-minute narrative with the same shot discipline.
FAQ
How long should a prompt be? Long enough to name subject, action, camera, light, and duration — usually 40 to 90 words. Past that, prompts tend to describe competing scenes instead of one clear shot.
Do vertical and widescreen need different prompts? Yes, at least in framing language. Vertical favors close-ups, centered subjects, and shorter holds. Widescreen rewards negative space, lateral movement, and longer beats. Write the version you intend to deliver instead of cropping later.
How many takes per shot? Three to five is a practical range. Fewer and you accept weak expressions; more and you spend the session reviewing instead of editing. Choose on performance.
Can a character stay consistent across a long project? Yes, with a written character bible, a verbatim description in every prompt, and an approved reference frame. Expect small drift and plan to hide it with a consistent grade and careful shot order.
Is AI-assisted work eligible for festivals? Often, but eligibility and disclosure requirements vary by competition and change over time. Check each festival's current criteria and follow the stated process instead of assuming.
What makes a social clip feel cinematic rather than cheap? Three things: one clear subject, motivated light with a nameable source, and a single camera movement at a deliberate pace. Most footage that reads as cheap is vague about all three.
Which language should I write prompts in? Use the language your tool handles best, then keep a translated copy of your character bible so collaborators work from the same descriptions. Consistency matters more than the choice of language.
How do I stop every generation from looking the same? Change one variable per round — light source, lens, or time of day — and hold the rest fixed. Changing everything at once makes it impossible to learn what actually caused the difference.
What if I have no composer or sound designer? Start with library ambience and one licensed music bed, then replace the track under the strongest beat first. Sound is the fastest way to make shot-sized fragments feel like a film.
Should I generate at final resolution? Generate at a workable resolution, make selections, then finish the chosen takes. Spending time on high-resolution output for clips you will never use is the most common way to slow a project down.
How do I keep a series from feeling repetitive? Change the location, the time of day, or the emotional register every few posts while keeping framing, grade, and caption position fixed. The identity stays recognizable; the subject stays new.
Make your next idea cinematic
The pattern behind all of this is straightforward: decide what the camera sees, write it precisely, generate in shot-sized units, and keep what you learn. Prompts are not a shortcut around directing — they are where directing happens now.
When you are ready to test the method, open the AI video generator, describe one shot completely — subject, action, camera, light, duration — and generate three takes. Keep the prompt that worked, then browse the Orelon blog for the next step: turning a series of strong shots into a story that holds from the first frame to the last.

