中
Chapter 40. Voice Casting and the Voice Bible

Part VIII — Voice, Narration, Music, and Sound Design

Chapter 40. Voice Casting and the Voice Bible#

In this chapter
40.1 Voice is a character asset too40.2 Characters must be distinguishable first40.3 The voice card40.4 Audition scenarios40.5 Golden voice samples40.6 Parameters and versions40.7 The pronunciation dictionary40.8 Rights and alternatives40.9 Voice QC40.10 Split voice identity into a stable layer and a state layer40.11 The joint casting matrix40.12 Regression lines and drift thresholds40.13 Designing three female voices for Backlit Takeover40.14 SOP for voice casting40.15 Fault tree40.16 Checklist, exercises and deliverables40.17 The voice identity contract40.18 Auditions must test relationships, not only monologue40.19 Voice regression sets and replacement plansA note on sources

40.1 Voice is a character asset too#

A voice that changes across episodes breaks identity as badly as a changing face. A voice bible stores more than "female, mature." It records timbre, range, pace, accent, breath, articulation, an emotional ceiling, generation parameters, licensing and prohibitions.

40.2 Characters must be distinguishable first#

Principal characters have to be identifiable on a phone speaker. Do not let Lin Yun, Lin Wei and the narrator all use the same low-mid, slow "premium female voice." Differences can come from resonance placement, sentence endings, rhythm, breath and emotional control rather than exaggerated accents.

40.3 The voice card#

voice:
  id: VO_LINYUN_01
  owner: CH_LINYUN
  age_read: 27-32
  timbre: neutral-to-low, clear without weight
  resonance: moderately forward
  pace: 4.2-4.8 characters per second; slightly slower under pressure
  articulation: technical terms precise, sentence endings falling
  breath: contained, few audible sighs
  emotional_ceiling: 4_of_5
  pressure_shift: no increase in volume, pauses shorten
  forbidden: [cute or coquettish, broadcast announcer, sustained breathiness]
  voice_engine: provider_voice_id
  parameter_version: 3
  rights: commercial_project_worldwide_3y

Pace is a starting point; the actual lines still require performance.

40.4 Audition scenarios#

Test every character with the same set: neutral explanation, a short command, a low threat, a fast rebuttal, suppressed emotion, tears, shouting, personal names, amounts, English abbreviations and an interruption. Auditioning one gentle line lets timbre collapse show up later at the climax.

Place candidates into a real scene with the music bed rather than listening for "a nice voice" in isolation.

40.5 Golden voice samples#

After approval, keep dry golden samples: neutral, under pressure, intimate, angry and exhausted. Later generations are compared against them for pitch, pace, timbre and articulation. Extreme-emotion samples never become the default reference for all lines.

40.6 Parameters and versions#

A fixed voice ID does not guarantee identical output after a service update. Record model, version, stability, style strength, speed, pitch and post-processing. Run regression lines before switching versions.

You cannot re-tune by ear on a different machine each episode until it sounds "about right."

40.7 The pronunciation dictionary#

Unify names, characters with several readings, company names, amounts, dates and English terms. The dictionary records the text, its phonetic form, a spoken example, the applicable characters and a version.

A character's own name pronounced differently in episode one and episode ten is far more noticeable than a small change of timbre.

40.8 Rights and alternatives#

Confirm consent from any human voice source, the permitted scope of synthetic use, the term, the territory and derivative versions. Never clone a real performer or public figure without permission. Prepare an alternative route and the cost of re-recording for key characters.

40.9 Voice QC#

Compare identity, perceived age, pace, articulation, noise, emotion and legibility on a phone. Audio similarity scoring can flag pitch drift; it cannot judge whether the performance is right.

40.10 Split voice identity into a stable layer and a state layer#

The stable layer holds base timbre, perceived age, resonance placement, accent range, articulation habits and normal range. The state layer holds fatigue, injury, drink, post-crying, telephone filtering, distance, secretive whispering and public speaking. A project cannot swap a voice because a character is ill, and it cannot keep a character identical in every state for the sake of consistency.

Every state derives from the approved voice and states its permitted range. Post-crying can add nasal resonance and broken breathing without suddenly raising the perceived age. A telephone state changes spectrum and space without changing pace habits. A disguise can change politeness and sentence endings while retaining one subtle consonant feature for a later reveal.

40.11 The joint casting matrix#

Cast by putting principal characters in pairs inside the same conflict, not by scoring them individually. The matrix checks whether pitch region, pace, resonance, breath, sentence endings, verbal density and emotional method overlap. If Lin Yun and Lin Wei share a base range, let Lin Yun's endings fall with pauses that come from thinking, and Lin Wei's endings hold with pauses used to make others wait. Neither needs an exaggerated register to be distinguishable.

Narration enters the matrix too. If an omniscient narrator resembles the protagonist's interior voice, the audience misjudges the narrative permissions. The system must state whether the narrator is the character herself, her future self, a third-party observer or a platform-style presenter — and decide from that whether it shares identity with a character's voice.

40.12 Regression lines and drift thresholds#

Keep 8–12 regression lines per core character covering neutral, fast, quiet, proper nouns, numbers, suppressed emotion and high intensity. Regenerate them on any change of model, voice ID, parameters or post-processing and compare blind against the golden samples. Review asks at minimum whether it is still the same person, whether the character relationships have changed, and whether the extreme states remain usable.

Small pitch differences can be matched in post. Changes to the bone structure of the timbre, to consonant habits, or to perceived age cannot be repaired with EQ. Regression results are pass, conditional pass or block; a conditional pass must state which scenes it covers, so "usable for ordinary dialogue" is never extended to "usable for sobbing."

40.13 Designing three female voices for Backlit Takeover#

Lin Yun, Lin Wei and Zhou Lan all need restraint and a sense of power, and searching for "premium female voice" would make them highly homogeneous. The final design does not rely on pitch. Lin Yun uses short phrasing and precise articulation on technical terms, and her volume drops as pressure rises. Lin Wei is slightly faster and uses half-sentence pauses to observe the room, holding her endings up in public. Zhou Lan speaks the most slowly, uses almost no filler, and lets silence force other people to explain first.

Testing all three in one scene, the team could tell who controlled the room with the picture turned off. The anonymous call in the mother mystery changes spectrum and distance temporarily while retaining Zhou Lan's characteristic long pause — a sound clue the audience can go back and find.

40.14 SOP for voice casting#

First, extract vocal behavior from the character logic. Second, design the differences jointly. Third, test with one shared set of audition scenarios. Fourth, choose inside the real mix environment. Fifth, build the voice card, golden samples and pronunciation dictionary. Sixth, record parameters and licensing. Seventh, monitor versions with regression lines.

40.15 Fault tree#

Symptom: each voice sounds good alone and they blur together. Casting was not joint. Compare same-scene dialogue on a phone speaker.

Symptom: the voice changes at the climax. Candidates were never tested at their emotional ceiling. Re-audition with extreme scenes.

Symptom: timbre drifts across episodes. Parameters, model version or post-processing went unrecorded. Restore the approval chain and run regression.

Symptom: names are always mispronounced. There is no project pronunciation dictionary. Maintain it centrally and forbid line-by-line hand fixes.

40.16 Checklist, exercises and deliverables#

Check that principals are distinguishable; that auditions cover the emotional range; that parameters and versions are recorded; that golden samples cover several states; that the pronunciation dictionary is unified; that rights are explicit; and that version updates run regression.

Exercise one: design different language and voice algorithms for three female characters. Exercise two: produce ten standard audition texts. Exercise three: build a pronunciation dictionary of twenty proper nouns.

Deliverables for this chapter: the voice bible, voice cards, the audition matrix, golden voice samples, the pronunciation dictionary, the rights record, and the voice regression set.

40.17 The voice identity contract#

Upgrade the voice card into a voice identity contract, stating which features constitute this character's voice and which may change with the story. The stable layer covers approximate pitch center, resonance placement, pace habits, articulation clarity, breath ratio, perceived age and accent boundaries. The state layer covers fatigue, post-crying nasality, lowered volume, public address and private whispering. The stable layer cannot jump without a story reason; the state layer must be triggered by an event.

voice_identity_contract:
  character: CH_LINYUN
  stable:
    pitch_center: medium_low_female
    resonance: forward_but_not_bright
    pace: 4.2_to_4.8_chars_per_second
    articulation: precise_endings
    age_read: late_20s_to_early_30s
  variable:
    public_control: reduced_breath_and_longer_phrases
    private_fear: shorter_phrases_and_audible_inhale
    post_cry: light_nasal_resonance
  forbidden:
    - broadcast announcer delivery
    - untriggered juvenile softening
    - a fixed rise at the end of every line

The contract does not require identical acoustic parameters on every line. What is genuinely consistent is that the audience still recognizes one person across different emotions and believes the change came from experience. Regression review compares both the stable layer and the plausibility of the states, rather than judging every natural performance variation as drift.

40.18 Auditions must test relationships, not only monologue#

Beyond neutral statements, quiet secrets, anger and crying, auditions add two- and three-person relationship tests. The same line — say it again — carries different status, distance and attack depending on whether it is addressed to a superior, a sister or an enemy. A candidate who only sounds good alone and cannot form a contrast with the other characters is unsuited to a season.

For joint casting, listen blind with every principal at the same level and processing, checking whether pitch region, pace, sibilance, accent and performance energy sit too close together. Do not force apart voices that were similar to begin with using EQ in post; differences in base timbre and rhythm belong to the casting stage.

Auditions should also use the project's real difficulties: personal names, company names, numbers, legal terminology, fast interruptions, quiet key words and long narration. A candidate who is stunning on one emotional line but frequently misreads core nouns, or runs out of support on long passages, is very expensive to produce.

40.19 Voice regression sets and replacement plans#

Keep 8–15 regression lines per character covering neutral, high emotion, low volume, fast delivery, long sentences, proper nouns and overlaps with other characters. On any change of model, parameters, recording environment or voice performer, compare identity, clarity, performance range and how well it takes post processing using that same set.

A commercial project must define the replacement route in advance: if the original voice becomes unavailable, does the whole season get recast and re-recorded, are only unreleased episodes replaced, or does a licensed voice model continue it? Contracts must cover training, cloning, derivatives, multiple languages and withdrawal conditions. Technically possible is not the same as permitted.

After a replacement, do not compare voiceprints alone. Put the new voice into dialogue, music and phone playback, and assess whether the established character relationships have changed. A brighter voice can make a restrained Lin Yun read as younger, which affects her credibility. Replacement is recasting a character, not swapping a file format.

A note on sources#

AI comic-drama production often leaves timbre selection to the last step in an editing application. This chapter raises voice to a cross-episode identity asset, referenced by the dialogue, narration and mixing chapters that follow.