Learn how multi-image fusion keeps AI characters consistent across shots, scenes, and styles, with seed sets, prompt patterns, drift fixes, and FAQs.
A character who looks like a different person in scene two will sink a series faster than any plot hole. Audiences forgive a wobbling camera move, a slightly flat line reading, or a background that does not quite match. They do not forgive a lead whose jawline, eye spacing, and hairline change every time the camera cuts. Multi-image fusion is the technique that makes that problem survivable: instead of describing a person in words and hoping the generator lands in the same place twice, you supply several images of the same face and let the model blend them into a stable identity you can reuse across shots, scenes, and even visual styles.
This guide covers how fusion actually works, how to build a reference set that survives dozens of generations, and how to run a production workflow that keeps a character recognizable across an entire series. It also covers what most quick tutorials skip: what to do when drift appears anyway, when fusion is the wrong tool for the job, and how to hold wardrobe, props, and lighting steady alongside the face.
Why Consistency Still Breaks AI Video
Every frame a video model produces is a fresh sample. The model does not carry a memory of who your character is; it re-derives a face from your prompt and whatever reference input it receives on each pass. Tiny numeric differences between passes compound. The nose gains a millimeter, the eyes drift a few pixels apart, the skin tone warms by half a stop. Individually those shifts are invisible. Across a ten-shot sequence they turn one person into three cousins.
Three failure modes show up again and again:
- Identity drift. The face stays in the same family but shifts in structure: jaw width, brow height, eye spacing, hairline, apparent age.
- Wardrobe and prop drift. A jacket changes cut, a scarf grows a pattern, a coffee cup becomes a different shape between cuts.
- Style drift. Grade, grain, and lens character change from shot to shot, which makes the identity shift feel even more severe than it actually is.
Text prompts alone cannot fix this. Descriptors such as "sharp jawline, dark wavy hair, green eyes, early thirties" describe a range of plausible people, not one individual. Two renders that both satisfy the prompt can look like siblings rather than the same person. Motion makes it worse: a model that produces a convincing static portrait may distort the face the moment the head turns, because nothing in the prompt defines what the back of that head looks like.
Fusion attacks the root cause. Instead of asking the model to invent a person from adjectives, you give it evidence and let the evidence do the work.
What Multi-Image Fusion Actually Does
Conditioning on several views instead of one
A single reference tells a model what a face looks like from exactly one angle in one lighting condition. Feed it four to six references covering different angles, expressions, and light, and the model has to reconcile them into one shared identity. That reconciliation is the fusion. The result is a representation — often described as an identity embedding or character token — that captures structure rather than surface detail, which is why it holds up better when you change the pose or the setting.
The two dials that matter: identity weight and prompt weight
Identity weight controls how strongly your references constrain the face. Prompt weight controls pose, framing, expression, action, and setting. Too little identity weight and you get drift. Too much and the character looks pasted in: the same stiff headshot under every lighting condition, refusing to emote. The useful band sits in the middle, where the face is locked but the performance is free. Finding that band takes two or three test renders per character, which is far cheaper than discovering the problem on shot forty.
Fusion versus training a dedicated character model
Fine-tuning and fusion solve the same problem from opposite directions. A fine-tune bakes identity into model weights, which can deliver very high fidelity for one specific look, but it needs a clean dataset, takes time to build, and can fight you when you want the same character in a different visual style. Fusion is a runtime conditioning step: minutes rather than hours, no dataset, and it travels more gracefully across styles. For most series work it is the practical default. A trained model is a specialist tool for a locked, high-volume look where the style never changes.
Where fusion sits in a real pipeline
The reliable pattern is stills first, motion second. Lock the character in image space, get approval, then animate approved frames. Orelon supports both halves of that loop: Create Image for building and testing the seed set, and Create Video for turning locked frames into moving shots. Skipping the stills stage and jumping straight to motion is the most common reason a promising character falls apart by shot three.
Building a Character Seed Set That Holds Up
The reference set is the largest single lever on final quality. A weak set cannot be rescued by clever prompting, and a strong set forgives a lot of sloppy language later.
The five-frame starter kit
| Frame | What it teaches the model | Practical note |
|---|---|---|
| Neutral front, even light | Core proportions | Avoid dramatic shadows; use a plain background |
| Three-quarter turn | Cheekbone and jaw depth | The most useful angle for dialogue coverage |
| Profile | Nose, chin, and hairline silhouette | Structural errors surface here immediately |
| Expression variant | How the face moves | A genuine smile or laugh prevents "mask" artifacts |
| Full body with wardrobe | Proportion and costume | Locks silhouette continuity across wide shots |
Add a sixth reference whenever the story demands a specific condition, such as dusk light, rain, or a uniform the character wears in half the episodes.
Match references to the conditions you plan to shoot in
If a series lives at golden hour, include at least one golden-hour frame. If it is mostly interiors under warm practicals, include one of those. When every reference is neutral studio light, the model has no example of how that face behaves in the conditions you actually care about, so it relights the character and the identity reads slightly off even when the geometry is correct.
Reference hygiene: what to leave out
Skip filters, heavy retouching, sunglasses, hats, extreme wide-angle distortion, motion blur, and other people in frame. Most importantly, avoid frames that were themselves generated. Recycling AI output as reference feeds drift forward — you are conditioning on an already-shifted version of the face, and each generation compounds the error. Start from real photographs whenever you can.
How many references is enough
Four to six is the sweet spot. One is too few, and ten or more can dilute the weighting until the model averages your character into a blend of features nobody intended. Five clean, well-lit, high-resolution frames will outperform twenty casual snapshots taken at a party, because the model learns structure from consistency, not from volume.
A Practical Workflow From Seed to Scene
Step 1: write a character bible
Keep it to one page: physical structure, wardrobe, props, palette, mannerisms, and the exact phrasing you intend to reuse. Include a short "do not say" list for descriptors that contradict your references. A shared document prevents collaborators from writing prompts that quietly fight each other, and it makes handoffs to editors or compositors trivial.
Step 2: generate wide, curate narrow
Produce around twenty candidate stills from your seed set, then keep five. Judge structure, not beauty. A gorgeous render with an unusual jaw width will cause problems in every subsequent shot, while a plainer frame with accurate proportions will carry the whole series.
Step 3: freeze an identity block and reuse it verbatim
Write one identity paragraph and paste it unchanged into every prompt in the production. Never paraphrase it, never reorder the words, never "improve" it mid-project. The model responds to phrasing as well as content, so a rewritten block is effectively a different character.
Step 4: test motion before committing to a sequence
Render one short clip where the character turns their head, stands up, and takes a few steps. Motion is where fusion stress shows first. If the face collapses during the turn, fix the reference set now rather than after you have produced thirty shots that all share the flaw.
Step 5: generate shot by shot and keep a ledger
Name files with a pattern that encodes the important variables, such as lead_a_seedv2_shot014_take3. Record the prompt, the reference set version, the seed, and the result in a simple sheet. When you need to regenerate a shot weeks later, the ledger is the difference between a five-minute fix and a full reshoot. Delete obvious failures immediately so nobody reuses them by accident.
Step 6: repair drift surgically
When a shot goes wrong, regenerate that shot. Do not reroll the whole sequence, and do not accept a mediocre frame because "the character is close enough." The usual fix is to re-anchor with the reference closest to the problem angle — the profile frame for a profile shot, the three-quarter frame for dialogue coverage — and to trim any contradictory descriptor from the action text.
Wardrobe, props, and palette continuity
Faces get the attention, but costume drift breaks immersion just as fast. Create a wardrobe lock sheet with one approved image per outfit and one sentence describing it. Treat recurring props the same way: the phone, the mug, the pendant, the car. Fix a three-color palette per character or per location and reuse those color words in every prompt. Backgrounds deserve a line too, since a room that rearranges itself between cuts reads as a continuity error even when the actor is perfect. Starting from a template and adapting it to your look is faster than building every shot prompt from a blank page.
Prompt Patterns That Keep a Face Stable
The most reliable structure separates identity from action. Identity stays frozen; everything else changes per shot.
IDENTITY (frozen, never edit):
[references A-E] woman, late twenties, oval face, high cheekbones,
straight nose, hazel eyes set wide, dark brown hair in a low bun,
thin scar above left brow
SHOT: medium close-up, shallow depth of field
ACTION: she sets down the cup and looks toward the door
WARDROBE: charcoal wool coat, cream turtleneck
LIGHT: warm window light from camera left, soft falloff
CAMERA: slow push in, eye level
AVOID: plastic skin, warped hands, duplicate jewelry, heavy grain
A few rules make this pattern work:
- Do not repeat physical descriptors inside the action block. Restating eye color or hair length invites contradictions with the references.
- Describe only what changes: pose, expression, action, framing, time of day.
- Use the same nouns for wardrobe every time. "Charcoal wool coat" and "dark grey jacket" are two different garments to a model.
- Keep the avoid list about artifacts and anomalies, not about identity. Putting facial traits in a negative prompt can push the model away from your references.
- Save your working patterns. A prompt library you can browse and clone removes most of the guesswork from a long production, and Orelon's prompt library is a good place to study structures other creators rely on.
Diagnosing Identity Drift: Mistakes and Fixes
When a character starts to slide, work through the causes in order rather than rerolling blindly.
- Too few references. Two images cannot describe a head. Add a profile and a three-quarter frame and re-test before changing anything else.
- Contradictory text descriptors. If the prompt says "soft round face" while the references show a narrow jaw, the model splits the difference. Delete the descriptor.
- Mixed lighting temperatures in the seed set. One daylight frame and one tungsten frame teach the model two different skin tones. Normalize the references or accept a warmer, less stable character.
- One dominant reference. If a single image carries far more visual weight than the others, the output becomes a copy of that angle. Balance the set or reduce the outlier's influence.
- Blending two characters in one pool. Ensembles need separate reference sets and separate prompts. Mixing them produces a third face that belongs to neither character.
- Ignoring aspect ratio and crop. Expanding from square to widescreen can stretch features subtly, so test at your delivery ratio before shooting the full sequence.
- Chaining generations. Generating from an already-generated frame accumulates error. Always return to the original references as your anchor.
- Changing the seed mid-sequence. A new seed can produce a new face even with identical references. Note the seed and keep it stable unless you deliberately want a variation.
A quick triage habit helps: compare the failing shot against your neutral front reference at 100 percent zoom. If the geometry is right and only the lighting feels wrong, the problem is in the shot prompt. If the geometry has moved, the problem is in the reference set or the identity weight.
Choosing the Right Approach: Decision Criteria
| Situation | Approach | Why |
|---|---|---|
| A single appearance in one scene | One strong reference image | Fusion setup is unnecessary overhead for a one-off |
| Recurring lead across many episodes | Multi-image fusion with five to six references | Best balance of setup cost and stability |
| Locked look, very high volume | A trained character model | Highest fidelity once the style is fixed |
| Same character in two art styles | Fusion with style-diverse references | Runtime conditioning adapts better than baked weights |
| Ensemble cast | One reference pool per character | Prevents feature blending between people |
| Fast client revisions | Fusion plus a versioned asset library | Regenerate a single shot without rebuilding the character |
If you are unsure, start with fusion. It is reversible, cheap to iterate, and you can always escalate to a trained model later once the design is truly locked.
Scaling a Character Across a Series
Once a character works, the goal shifts from creation to maintenance. Store each character in its own folder containing the approved reference set, the frozen identity block, the wardrobe sheet, the palette, and a handful of golden shots that represent peak quality. Version the folder, not individual files, so you never mix a reference from one revision with prompts written for another.
Batch your work by character rather than by scene. Generating all of a lead's shots in one session keeps conditions stable and makes drift easier to spot, because you are comparing like with like. Finish with a dedicated quality pass where you watch the sequence at normal speed — drift that is invisible frame by frame often becomes obvious in motion, and the reverse is also true.
Finally, plan for format variations. Vertical cuts, teaser clips, and social edits should be generated from the same locked stills rather than cropped from finished video, so the character stays sharp. Keep a short note in the Orelon blog habit loop: what changed, what broke, and which fix worked. That log becomes the most valuable document in your production.
FAQ
How many reference images do I really need? Four to six well-lit, high-resolution frames covering front, three-quarter, profile, an expression variation, and full body. More than eight rarely helps and can dilute the identity toward an average.
Can I use AI-generated images as references? Technically yes, practically it is risky. Any flaw in the generated reference becomes part of the character definition and compounds with each new generation. Use real photographs when possible.
Why does my character change when I switch visual styles? Style conditioning and identity conditioning compete. Lower the style strength, raise identity weight, or include a reference rendered in the target style so the model has an example of how that face should look in it.
What causes a character to look right in stills but wrong in motion? Motion exposes structure the model has not learned. Add profile and three-quarter references, then test with a head-turn clip before producing the full sequence.
Should I put facial traits in the negative prompt? No. Negatives that describe facial features can push the model away from your references. Reserve them for artifacts like warped hands, plastic skin, or duplicate jewelry.
How do I keep two characters from blending into each other? Keep separate reference pools and separate identity blocks, and avoid prompts that describe both people in one generation. Generate each character in their own passes and combine in the edit.
Is multi-image fusion enough for a full episodic series? For most series, yes, provided you maintain the reference set, freeze the identity block, and run a quality pass on every cut. The discipline of versioning and logging matters more than the specific tool you choose.
Bring Your Character to Life in Orelon
Consistency is not a single setting you switch on; it is a habit built from good references, a frozen identity block, disciplined versioning, and quick surgical fixes when drift appears. Get those four things right and your character will survive lighting changes, costume changes, and forty shots of coverage without losing the face your audience recognized in the first scene.
Start with a seed set and a few test stills in Create Image, then move to Create Video once the character is locked. Orelon is built for cinematic ideas in motion, so the same identity you approve in a test frame is the one you can carry through an entire sequence.



