Build a repeatable AI sound pipeline for short-form video: writing narration, directing synthetic voices, generating music beds, syncing, and mixing.
A short video gets roughly one second to earn the next one, and most of that second is sonic. Before a viewer reads a caption or registers the framing, they have already heard a voice, a beat, and the room around them. That first sonic impression decides whether the rest of the video exists for them. It is why the most underrated stage of any AI video pipeline is not the generation model but the sound wrapped around it: narration that carries the idea, music that sets the emotional temperature, and a mix that keeps both intelligible on a phone speaker at arm's length. What follows is a complete, repeatable workflow for producing AI narration and AI music for vertical short-form video, from the first line of script to the final loudness check.
Why audio decides whether a short video gets watched
Short-form feeds are scoring machines. Every swipe is a fresh judgment, and the earliest part of that judgment is audio. A confident, well-paced voice registers before the viewer consciously processes what is on screen. When a voice is flat, rushed, or buried under a music bed, the swipe happens before the story has a chance to start.
Three conditions make sound quality more decisive than it used to be:
- Volume of content. Thousands of vertical videos compete for the same few seconds, so anything that sounds average reads as background noise.
- Mobile-first playback. Most viewers watch on a small speaker or a single earbud in a noisy room. Quiet, dense, or overly subtle mixes simply disappear.
- Captions changed expectations. Because captions carry comprehension, audio no longer has to explain. It has to perform. Voice becomes tone and personality; music becomes rhythm and mood.
The practical consequence is simple: treat audio as a first-class production stage, not a cleanup step. Decide what the voice is doing and what the music is doing before you generate a single frame.
The three sound layers of a vertical video
Almost every effective short video is a stack of three layers, and each has a different job. When a video feels off and nobody can articulate why, usually one layer is doing another layer's work.
Narration
Narration carries meaning, structure, and personality. It should be intelligible at low volume and should never fight the visuals for attention. If the narration restates the caption word for word, it is not a layer — it is a duplicate.
Music bed
The bed sets emotional temperature and tempo. It tells the viewer how to feel about what they are seeing: tense, curious, warm, playful, triumphant. Its most important quality is restraint, because a bed that competes with speech damages both elements.
Effects and ambience
Transition sounds, whooshes, light room tone, and small impacts create the physical sense of space and motion. This is the layer most AI-native creators skip, and it is often what separates something that looks generated from something that feels edited.
Before generating anything, write one sentence per layer. Example: narration — calm, slightly amused explainer; music — warm low pulse under 90 BPM with no lead melody; effects — soft whoosh on each scene change and a thin room tone underneath. That sentence prevents most of the problems you would otherwise fix late at night in the timeline.
Writing narration that synthetic voices can perform
Synthetic voices are excellent at consistency and terrible at rescuing weak writing. A sentence that reads fine silently can sound clipped or robotic when a model performs it, because the model takes punctuation, spacing, and rhythm literally.
Read for breath, not for grammar
Read every line out loud at a natural pace. If you have to rush to fit a phrase into one breath, the model will rush it too, and the rush reads as mechanical. Break long clauses into two shorter sentences. Vertical narration almost never suffers from short sentences.
Punctuate for performance
Commas, em dashes, and periods are performance instructions. A period buys a beat of silence. A comma buys a small lift. An ellipsis suggests hesitation, though it is easy to overuse. If you want a pause inside a grammatically continuous sentence, restructure the sentence rather than stacking punctuation.
Write the hook as speech
The first line should be something a person could actually say to a friend. Spoken hooks are shorter, more concrete, and often start mid-thought. "This took me eleven tries" outperforms "In this video, I will explain the process of creating a short video with AI." One useful exercise: write the script, delete every adjective that does not change meaning, and read it back. What survives is usually the version worth generating.
Budget words against runtime
A comfortable speaking rate is roughly 140 to 165 words per minute. That means a 30-second video holds about 70 to 85 spoken words, and a 60-second video holds about 140 to 165. Write to that budget before you generate anything, and add a five-word cushion for breaths.
Casting and directing an AI voice
Voice libraries are large enough that selection becomes the hardest creative decision in the pipeline. Ignore novelty voices and evaluate on four practical criteria.
Timbre and perceived age
Pick a voice whose perceived age matches the point of view. Advice about a career change lands differently from a teenager explaining a game. Timbre also determines how the voice sits in a mix: brighter voices cut through music more easily, while darker voices need more space carved out for them.
Range, pace, and micro-variation
Generate the same three sentences in several candidate voices and listen back to back. Good narration has micro-variation in pitch and speed: it lifts slightly on the important word and settles at the end of a thought. A voice that delivers every sentence at identical energy will exhaust the viewer within fifteen seconds.
Pronunciation, names, and numbers
Proper nouns, brand names, acronyms, and numbers cause most pronunciation failures. Test them early, before you build a whole video around a voice. Many systems accept phonetic respelling or a pronunciation hint, which is much faster than regenerating a full take.
Direction you can actually write down
Most voice tools accept a style or delivery instruction. Useful directions are concrete: "warm, conversational, unhurried, ending sentences downward," or "energetic but not shouty, slight smile in the tone." Vague directions like "natural" or "engaging" rarely change the output because they do not describe a physical behavior.
Multilingual delivery
If your audience is bilingual, decide whether narration should be in one language, subtitled in another, or fully duplicated. Duplicating the voice track is inexpensive once the script and timing are locked, and it doubles the reach of a single visual edit. Keep sentence lengths similar across languages so the visuals stay aligned with either version.
Generating a music bed that supports the voice
Music generation models respond to description better than to genre labels alone. "Lo-fi hip hop" gives you a mood but says nothing about arrangement, so specify the role the track must play.
Describe instrumentation, energy, and space
Strong prompts combine three things: instrumentation, energy, and space. For example: warm analog synth pad, slow pulse, no drums, spacious, slightly nostalgic, leaves room for a speaking voice. Naming what should be absent is as useful as naming what should be present. Ask for "instrumental, no lead melody in the vocal range, soft transients" and you remove most of the reasons a bed fights narration.
Build the bed in four sections
You rarely need a full song. You need four functional pieces:
- Opening hit — one to three seconds that establish mood and mask the cut into the video.
- Steady bed — a loopable middle section with stable energy and nothing distracting in the midrange.
- Lift — a small build at the reveal or the punchline.
- Button — a clean ending sting, or a fade that finishes two frames before the video does.
Generating four short pieces and editing them together gives you more control than generating one long track and hoping the important moments land in the right places.
Match tempo to content, not to taste
Tempo sets how the viewer feels about pacing. Under about 80 BPM reads as calm or reflective; 90 to 110 BPM feels like a workflow or explainer; 120 BPM and up pushes urgency and energy. Pick the tempo first, then let the edit follow it.
Syncing voice, music, and generated visuals
Generated clips are rarely the exact length you want, which creates the classic problem: the narration says one thing while the picture shows another.
Cut on sentence boundaries
Build the edit around sentence boundaries, not round numbers of seconds. Generate a little extra footage for each shot, then trim so the picture changes when the narration changes thought. Viewers forgive a slightly abrupt cut far more readily than a voice finishing a sentence over unrelated imagery.
Keep handles on both ends
When you generate clips, keep one to two seconds of extra action at the start and end. Those handles are what let you slide visuals against audio without frozen frames. It is the easiest habit to build for smoother edits, and it helps to generate and assemble in the same place so timing notes survive the handoff. You can start from Orelon's video creation workspace to keep prompting, generating, and timing decisions inside one project.
Use music accents only where they mean something
Beat-matching every visual cut to a musical accent gets tiring within twenty seconds. Save the accents for the two or three moments that matter: the reveal, the turn, the final line.
Handle lip movement honestly
If a character is visible and speaking, either match the mouth to the narration or avoid framing the face during speech. The reliable techniques are cutting to hands, objects, or environment while the voice continues, or treating the voice as internal monologue with the character silent on screen. Imprecise lip sync reads as uncanny faster than almost any other flaw.
The end-to-end production workflow
Here is the full sequence in production order. Each step is quick, and skipping one usually costs more time than doing it.
- Write the audio brief. One line for narration, one for music, one for effects.
- Draft and trim the script. Run the breath test, then read it aloud against a timer.
- Generate two or three narration takes. Compare them back to back at low volume on a phone speaker, not on headphones.
- Lock the narration. Timing decisions cascade from it, so never build visuals around a scratch read.
- Generate music in sections. Opening, bed, lift, button — then assemble the bed under the locked voice.
- Generate visuals to the narration's beats. Give each shot a clear job and keep handles at both ends. If you want a proven shot pattern instead of a blank prompt, browsing Orelon templates is faster than inventing structure from scratch.
- Layer effects last and sparingly. One transition sound per cut is usually enough; two becomes a signature; five becomes noise.
- Mix, check on a phone, export. Then watch it once with the sound off to confirm the captions and visuals still carry the idea.
If you are still choosing models and formats for the visual half, Orelon's prompt library and the Orelon blog cover prompt structure and shot planning so the audio and video stages stay aligned.
Mixing and loudness for phone speakers
Mixing for vertical video is mostly about intelligibility, not perfection. You want the voice clearly on top, the music clearly underneath, and nothing popping or clipping.
A practical starting point:
| Element | Rough target |
|---|---|
| Narration peak | around -6 to -3 dB |
| Music bed under speech | roughly 12 to 18 dB below the voice |
| Effect hits | audible but never louder than narration |
| Final ceiling | below 0 dB with no clipping |
The exact numbers matter less than the relationship between them. If you cannot understand the voice on a laptop speaker at 30 percent volume, the bed is too loud no matter what the meter says.
Techniques that help almost every short video
- High-pass the voice around 80 to 100 Hz to remove rumble that eats headroom.
- Cut the bed, do not boost the voice. Lowering music is faster and sounds more natural than pushing narration up.
- Check in mono. A phone speaker is effectively mono, and wide stereo music can partially vanish.
- Leave one real silence. Two-thirds of a second of music-only space before a reveal buys a surprising amount of attention.
- Avoid hard ducking on every sentence. Constant sidechain pumping is audible. A modest, constant reduction of the bed usually sounds cleaner.
Captions are part of the sound decision
Because a large share of viewers watch muted, captions are not a separate accessibility chore — they are part of how the piece is heard. Keep caption phrases short enough to read in one glance, place them away from faces, and make sure the words on screen match the narration exactly. A mismatch between what is read and what is heard breaks trust immediately.
Mistakes, QA, and troubleshooting
Most sound problems in short-form video fall into a small set of causes. Use this as a pre-export checklist.
| Symptom | Likely cause | Fix |
|---|---|---|
| Narration sounds robotic | Long clauses, no micro-pauses, monotone delivery | Split sentences, regenerate with a delivery direction |
| Voice feels muddy | Music has a lead line in the vocal range | Regenerate an instrumental bed with no midrange melody |
| Volume jumps between shots | Voice generated in separate sessions with different settings | Regenerate all lines in one session with one voice preset |
| Video feels rushed | Script written to the word budget of a longer runtime | Cut words, not breaths |
| Ending feels abrupt | Music stops with the voice | Add a two-second button or fade two frames before the last image |
| Everything sounds loud and flat | Over-compression and clipping on export | Lower levels, re-check peaks, keep silence in the edit |
Beyond the table, four habits prevent most of these issues before they happen. Lock narration before visuals. Generate voice in a single session whenever possible. Keep two voices maximum unless the format genuinely requires a cast. And audition every mix on the worst playback device you own, because that is where much of the audience lives.
FAQ
How long should narration be in a 30-second video? Around 70 to 85 spoken words at a natural pace. Write to that budget before generating audio; trimming a synthetic take is slower than trimming a script.
Should I generate narration before or after the visuals? Narration first. Lock the voice, then generate visuals to its timing. Working in the other order forces awkward pacing compromises and makes sync problems much harder to hide.
Can one synthetic voice carry an entire channel? Yes, and consistency is usually an advantage because repetition builds recognition. If you want variety, vary pacing, script style, and shot rhythm before you vary the voice itself.
How do I stop a music bed from fighting the voice? Ask for instrumental, drum-light, spacious tracks with no lead melody in the midrange, lower the bed before raising the voice, and leave at least one clear silence in the edit.
How many narration takes should I generate? Two or three per section is the useful range. One take removes your ability to compare delivery, and ten takes turns a ten-minute task into an hour.
What should I do about loudness before exporting? Aim for a voice that stays clear at low volume, keep the final peak below the clipping point, and avoid heavy limiting. If the mix only sounds good at full volume, it will sound broken on a phone.
Is AI-generated music safe to publish? It depends on the terms attached to the generator and on the platform where you publish. Read the license for each track you use, keep a record of it, and prefer tools that state clearly how commercial use works.
Put sound first on your next project
Better AI video work rarely comes from more effects. It comes from deciding what each layer is responsible for, generating three clean takes instead of thirty messy ones, and mixing so a voice cuts through a phone speaker in a noisy room. Narration carries the idea, music carries the feeling, effects carry the physicality, and the edit holds all three together.
When you are ready to run the whole pipeline in one place, start a project on Orelon, write your three-line audio brief, and generate the first scene with the voice already in mind. That single change in order — sound before spectacle — is what makes a vertical video feel intentional rather than assembled.



