Part VIII — Voice, Narration, Music, and Sound Design
Chapter 41. Dialogue Recording, Generation, and Performance Control#
In this chapter
41.1 Do not treat each line as an independent file#
Generating line by line resets breath, pitch and emotion every time. Generate by emotional section within one scene, then split at natural pauses. Short lines needing precise lip sync can be generated separately, provided they match the surrounding context.
41.2 Three layers of script#
The text layer holds the lines. The performance layer adds objective, stress, pauses, interruptions and emotion. The recording layer adds pronunciation, file IDs and technical parameters.
dialogue_take:
line: E001_L012
speaker: VO_LINYUN_01
text: Then the acquirer's representative will attend.
objective: make the room acknowledge her formal standing
emphasis: acquirer's
pace: controlled
pause_before_ms: 420
tail_hold_ms: 300
emotion_in: restrained_2
emotion_out: dominance_3
lip_sync: true
41.3 Direct by intent#
"Angrier" is far weaker than "hold the volume down so they have to lean in to hear you." Give the performer or the model an action, a listener, and an emotion that must not be exposed. Performance beats an emotion label every time.
Adjust only one or two variables per pass: stress, pace, pause, volume or breath. Asking at once for angrier, more restrained, more powerful and more natural is self-contradictory.
41.4 Breath and pauses#
Breath is the interface between performance and editing. Keep natural inhales; removing noise must not turn a character into a machine that never breathes. Write the reason for an important pause: thinking, suppressing, waiting for a reaction, or letting evidence land.
Final pause lengths are decided together with picture and music, not locked entirely in the text.
41.5 Interruptions and overlaps#
Generate each line complete and clean, then design the overlap on the timeline. The interrupted character keeps an unfinished ending; the interrupter's entry needs attack. Do not let two voices at the same frequency mask each other for long.
Subtitles usually favor the dominant line, with the secondary voice simplified or offset.
41.6 Human, TTS and hybrid#
Human performance suits complex emotion and live interaction. TTS suits scale, revision and multiple languages. A hybrid can use a human-directed sample to control a synthetic voice, or put human performance on key climaxes and synthesis on ordinary passages. Hybrids must unify timbre, room sense and loudness.
Choose by project risk and licensing, not by ideology.
41.7 The lip-sync workflow#
Approve the final dialogue audio before generating or syncing mouths. Keep sentences short and the frontal mouth clear. Split long lines into a short lip-synced opening, off-screen continuation over an insert or reaction, and a return to the character at the end.
Lip-sync QC checks phonemes, expression, identity, teeth and head movement together. Accurate sync with a collapsed face is still a failure.
41.8 Audio cleanup#
Keep the original dry recording; denoise, de-plosive, control sibilance, equalize and lightly compress on a copy. Keep processing consistent per character, so different episodes do not sound like different microphone worlds.
Excessive denoising creates underwater artifacts; excessive compression makes every line equally forceful.
41.9 Take management#
Store takes, performance differences, state and the reason for selection per line or section. A candidate cannot
be called line_final_2.wav.
E001_L012__VO_LINYUN_01__take03__restrained.wav
Lock the chosen take; lip sync and subtitles reference that audio hash.
41.10 The scene as a recording unit#
Generate or record a full performance pass for a scene first, letting the character travel from entry state to exit. The pass need not all be used in the cut, and it establishes continuity of breath, pitch and strategy. Afterwards, do pickups only on key lines, lip-sync lines and failures — carrying the preceding and following lines as context and compared at the same monitoring level.
When several characters cannot genuinely be recorded together, produce a temporary scene partner track. The second character performs while listening to the first character's approved or near-approved version, then return to fix the first character's reactions and interruptions. Generating both sides in complete isolation produces two people who are each emphasizing and neither listening.
41.11 Dialogue editing is not removing every imperfection#
Editing first handles wrong words, plosives, clicks, over-long silences and obvious timbre jumps, then preserves meaningful inhales, swallows, hesitations and unfinished endings. The position of a breath serves intent: an inhale before a threat means something different from the exhale after finishing. Uniformly reducing every breath deletes the structure of the performance.
When splicing pickups, check three kinds of continuity: pitch and resonance, background and processing chain, and emotional momentum. Technically seamless but with a character suddenly moving from restraint to recitation is still a bad edit.
41.12 Dialogue lock and change discipline#
Once lip sync, subtitles and the final mix depend on it, dialogue audio becomes a source of truth. Any word change, take change or altered pause produces a new version and notifies lip sync, subtitles, editing and music. You cannot quietly move a line 300 milliseconds in the final mix while subtitles and mouths still reference the old hash.
An urgent wording fix assesses the least-impact route: can it be covered off screen, can it keep the original duration, does the frontal lip sync need redoing, does it change the other character's reaction. Fix semantic errors first; "another version sounds nicer" is usually not enough to trigger expensive rework after lock.
41.13 The interruption in the boardroom#
Lin Wei starts to say that Lin Yun has no standing whatsoever, and Lin Yun cuts in before the word standing: the authorization is right here. In production, Lin Wei's complete line is kept as a clean asset first, then exported with an unfinished ending. Lin Yun's inhale enters 80 milliseconds early and her line begins over Lin Wei's final consonant, while authorization stays clear. The crowd's ambience drops at the moment of the interruption, and the document landing on the table follows immediately on here.
Subtitles do not display both full lines at once: the first appears and is replaced by the second the moment the interruption lands. Sound, action and subtitles jointly show power being taken away, rather than two audio files simply layered.
41.14 SOP for dialogue audio#
First, organize the scene by emotional sections. Second, produce the performance script. Third, generate or record complete sections. Fourth, choose performances and split naturally. Fifth, design interruptions and overlaps. Sixth, clean while preserving the dry recording. Seventh, lock the final takes. Eighth, produce lip sync. Ninth, force-align subtitles. Tenth, review inside scene sound and music.
41.15 Fault tree#
Symptom: every line sounds like a fresh start. Generation was line by line. Re-record by emotional section or use overlapping context.
Symptom: the voice is natural and the scene does not advance. Emotion was pursued without an action objective. Re-direct by objective.
Symptom: the interruption sounds like two files stacked. The unfinished ending and the frequency space were never designed. Redo the entry and the mix.
Symptom: lip sync is accurate and stiff. Sentences are too long, or only phonemes were optimized. Split the lines and preserve performance shots.
41.16 Checklist, exercises and deliverables#
Check that work proceeds by section rather than by line; that every line has an objective; that stress and pauses are recorded; that overlaps are designed; that dry recordings are preserved; that final takes are locked; that lip sync uses the final audio; and that subtitles are re-aligned.
Exercise one: record the same line three ways — as attack, as probe, as status assertion. Exercise two: design a natural interruption. Exercise three: split a twelve-second sentence into lip sync plus off-screen coverage.
Deliverables for this chapter: the performance script, dialogue takes, take decisions, clean dialogue, the overlap map, lip-sync audio, and alignment data.
41.17 The scene take map: holding the performance arc#
Dialogue records more than which take number was chosen per line. A scene take map lays the whole scene out by beat, marking each character's strategy, intensity, breath, pauses, interruptions and final selection. Reviewers listen horizontally to one character from entrance to exit, which prevents brilliant individual lines from producing emotion that lurches up and down.
scene_take_map:
scene: E001_SC03
beats:
B01: {linyun: T03, intensity: 2, tactic: observe}
B02: {linwei: T02, intensity: 3, tactic: humiliate}
B03: {linyun: T05, intensity: 2, tactic: withhold}
B04: {linyun: T04, intensity: 3, tactic: prove}
bridges:
B03_to_B04: inhale_from_T04_alt2
rejected:
- take: B04_T01
reason: peaks too early, undermining the later suppression
The best individual line does not necessarily belong to the best scene. If T05's outburst is thrilling but takes the character to her peak before the evidence actually appears, choose the more restrained T04. A take decision must state the scene function; "more feeling" is not a sufficient reason on its own.
41.18 Continuity of breath, mouth sounds and physical action#
Breath is the seam between performance and editing. If a character takes a deep inhale to speak in the previous shot, the next cannot begin with a completely airless clean attack. Rapid breathing after crying cannot return to broadcast steadiness immediately after a cut. Preserve meaningful inhales, swallows and fabric movement; remove distracting noise without processing a person into a body-less voice.
Mouth sounds, plosives and sibilance are treated by perceptibility. Removing them completely makes close voices sound unnatural; too much becomes uncomfortable on headphones. Fix obvious technical problems first, then judge across the whole scene and on the target device. Periodic breathing, repeated mouth sounds and odd tails in generated audio deserve separate tagging, because they can accumulate into a machine rhythm over long passages.
Sound follows physical action: standing changes breath support and vocal distance; turning away changes the high end; moving closer lowers the reverb ratio. Even when the picture is generated as separate shots, the acoustic perspective should stay continuous along the blocking.
41.19 The dialogue master and the lip-sync change protocol#
Only the finally approved dialogue master drives lip sync, subtitles and music. Temporary takes may be used in a rough cut and must be clearly marked, so downstream does not mistake them for locked. The master stores the original dry recording, the edit session, the processed version, timecode, word-level alignment and a content hash.
Changing one word affects phonemes, duration, lip sync, subtitle breaks, picture cut points and music ducking. Change impact first determines whether a local replacement inside the same breath group is possible; if the replacement damages the delivery or the background noise, redo the complete thought unit. Never force a new word into old lip sync and rely on fast cutting to hide a key line.
After the change, regress the breaths either side, the room tone, the character state, the subtitles and the shot length. Matching audio content is not enough — performance momentum must connect too. For already-published multilingual versions, a master change triggers re-evaluation per language rather than automatically applying the same word count.
A note on sources#
Audio generation is fast, and performance continuity depends on context, splitting and the final timeline. Push-button line-by-line dubbing does not assemble itself into a scene.