中
Chapter 46. Mixing, Loudness, and the Phone Speaker

Part VIII — Voice, Narration, Music, and Sound Design

Chapter 46. Mixing, Loudness, and the Phone Speaker#

In this chapter
46.1 Clarity always outranks loudness46.2 The dialogue chain46.3 Frequency conflict46.4 A loudness starting point46.5 The relationship between dialogue and music46.6 Dynamic range46.7 Four listening passes46.8 Mono and phase46.9 Consistency across episodes46.10 Export and measurement46.11 Build a static balance first, then automate46.12 Phone intelligibility is not simply more high frequency46.13 Read loudness reports in sections46.14 Stems and the capacity to revise46.15 Phone acceptance on Backlit Takeover46.16 SOP for mixing46.17 Fault tree46.18 Checklist, exercises and deliverables46.19 Narrative priority automation46.20 Platform transcoding and a real device matrix46.21 Revisions, international versions and M&EA note on sources

46.1 Clarity always outranks loudness#

Microdrama's first audio job is that dialogue and narration are understood. If music, low-frequency hits or Foley bury a key word, the mix has failed however cinematic it sounds.

The usual priority is key dialogue and narration, narrative sound effects, ambience, music, then decorative effects. A climax may change that temporarily without losing information.

46.2 The dialogue chain#

Edit and remove obvious noise first, then apply gentle EQ, compression, de-essing and loudness automation. Keep the timbral differences between characters; do not compress everyone into one broadcast voice.

Clean on a copy of the dry recording and keep the original. Excessive denoising and compression produce an underwater quality, breath pumping and listener fatigue.

46.3 Frequency conflict#

Clarity in a female voice frequently conflicts with piano, strings and synthesizer midrange. Use music stems, dynamic EQ or an arrangement that clears the band, rather than only lowering the overall level. Phones lack low end, so important impacts need a perceptible midrange component too.

46.4 A loudness starting point#

Online episodes can start testing around -14 LUFS integrated with true peak no higher than -1 dBTP, and platforms may normalize again. Short ads, different platforms and regions require measurement; the numbers are not absolute acceptance criteria.

What matters more is that episodes, and ads against episodes, do not change dramatically. Record the measurement tool and its version.

46.5 The relationship between dialogue and music#

When a voice is present, music around 12–18 dB below it is a starting point, adjusted by frequency and scene. A low whisper needs more space; shouting does not justify pushing music down to inaudibility.

After automatic ducking, hand-correct the key lines so the music does not pump on every syllable.

46.6 Dynamic range#

Phone environments need moderate dynamic control, and being equally loud throughout is exhausting. Preserve the contrast between drops, quiet passages and hits. Quiet sections should still sit above the audible threshold of a real viewing environment.

Do not use a peak limiter to solve a balance problem.

46.7 Four listening passes#

Headphones check detail, noise and space. A phone speaker checks dialogue and midrange. A cheap Bluetooth speaker or laptop checks poor devices. Watching muted checks subtitles and visuals. You can also test the entrance at low volume with simulated traffic noise.

Passing once on professional monitors is not finished.

46.8 Mono and phase#

Phones may be close to mono. Check that dialogue, music and effects do not cancel when the stereo image folds down. Keep critical narrative sound stable in the center and use width as support.

46.9 Consistency across episodes#

Establish reference passages for dialogue, narration, music, hits and ambience. Compare each new episode's loudness and timbre against them. Different editors must not each pursue louder.

46.10 Export and measurement#

Keep separate tracks, stems, the mix master, the limited master and the platform outputs. Re-measure loudness, peak, channels, sample rate and sync after export, in case encoding changed something.

46.11 Build a static balance first, then automate#

Choose one representative dialogue passage and establish base levels, EQ and compression with minimal automation, so dialogue, ambience and music sit in a sensible relationship. If the base chain does not hold, drawing volume line by line only creates patches nobody can maintain.

Then automate by scene intent: clear space for music before key words; raise intelligibility on a whisper without changing its intimate distance; restore room tone after a hit; adjust the bed by thought segment under long narration. Automation has an entry and a recovery; parameters must not twitch on every word.

46.12 Phone intelligibility is not simply more high frequency#

Phone speakers have limited low end and the listening environment is noisy, so blindly boosting the high end amplifies sibilance and fatigue. Solve dialogue editing, music arrangement and midrange masking first, then apply gentle EQ, dynamic control and loudness automation. Key consonants must be clear without turning every character into a bright advertising voice.

A low-frequency impact needs a midrange transient, an environmental response and brief dynamic contrast alongside it. A rumble at 40 Hz may disappear entirely on a phone; adding table contact, a metallic layer or the crowd pausing keeps the narrative impact.

46.13 Read loudness reports in sections#

Meeting integrated LUFS across an episode does not mean it is audible locally. Also check short-term loudness, dialogue range, true peak, silence or near-silence anomalies, left/right balance and mono fold-down. The opening, the paywall cut and the ending must not be hidden by an overall average.

Platforms may transcode and normalize, so before release listen back to an actual private upload of the file. When platform compression causes sibilance, distortion or loudness changes, adjust the platform master rather than the lossless mix master. Each platform output keeps its own version and measurement record.

46.14 Stems and the capacity to revise#

Deliver at minimum dialogue, narration, music, ambience, Foley/effects and full-mix stems, all starting from the same time zero. For international versions requiring M&E, confirm that dialogue-related breath, fabric and action sound is not entirely printed into the dialogue track — otherwise removing the language empties the scene.

Stems are the basis for later wording changes, ad re-cuts, other languages and censored versions. Keeping only one limited stereo master means every local change either loses quality or forces a full re-mix.

46.15 Phone acceptance on Backlit Takeover#

The boardroom scene was rich on monitors, and in the first phone test Lin Yun's quiet delivery of "the father's account" was masked by piano and HVAC midrange, while the low frequency of the document landing almost vanished. The team did not simply raise the dialogue: the piano stem was pulled for two seconds, HVAC ducked slightly under the key words, a short midrange contact was added to the paper, and 250 milliseconds of space was left after the document landed.

Overall loudness barely changed, and both the key words and the action now hold on a quiet phone. That is what device acceptance is for: it checks whether narrative information survives real playback conditions rather than chasing a bigger number.

46.16 SOP for mixing#

First, organize dry recordings and audio assets. Second, clean dialogue and unify per-character chains. Third, establish room tone and Foley. Fourth, score using stems. Fifth, handle frequency and ducking. Sixth, automate key passages and dynamics. Seventh, measure loudness and peaks. Eighth, run the four listening passes and the mono check. Ninth, re-check after platform export.

46.17 Fault tree#

Symptom: the music is already low and the dialogue is still unclear. This is frequency conflict, not overall level. Pull a stem or use dynamic EQ.

Symptom: no impact on a phone. The effect relies only on sub-bass. Add a midrange transient and an environmental response.

Symptom: the whole episode is tiring. Compression and limiting are excessive and there is no quiet tier. Restore dynamics and drops.

Symptom: distortion after upload. True peak or encoding headroom is insufficient. Lower the ceiling and re-check the platform file.

46.18 Checklist, exercises and deliverables#

Check that key words are clear; that frequencies make room; that loudness is stable; that true peak is safe; that ducking is smooth; that phone and mono pass; that episodes are consistent; and that platform files are re-measured.

Exercise one: clear space for dialogue without reducing the overall level. Exercise two: produce headphone and phone listening records. Exercise three: compare an over-compressed version against one that preserves dynamics. Exercise four: establish a cross-episode loudness reference.

Deliverables for this chapter: the dialogue chain, the mix session, the loudness report, device QC, the mono check, stems, and platform masters.

46.19 Narrative priority automation#

Mix automation serves the hierarchy of information first. Tag each beat with its first-attention sound, its second-attention sound and its background. On a key line, dialogue is first, the document landing second, music recedes. In a wordless investigation montage, narration is first, evidence Foley second, music maintains momentum. For half a second after an identity reveal, the environmental reaction can be first.

Figure 46-1 Six classes of sound yielding by narrative priority and producing a phone master

Figure 46-1 Dialogue, narration, music, ambience, Foley and effects are not set once and left. Each story beat reassigns the foreground, and automation curves hand attention over smoothly. The final mix is then verified on phones, headphones, small speakers and through platform transcoding.

The curves are not a mechanical rule that higher means more important. Foley for the document landing needs to be foreground for tens or hundreds of milliseconds, and music may return during a pause in the dialogue. The real goal is that the audience knows what to listen to at every moment without feeling other layers of the world suddenly disappear.

mix_priority:
  beat: B04_authorization_reveal
  primary: dialogue_linyun
  secondary: prop_authorization_hit
  background: [room_tone, crowd_low, music_pulse]
  automation:
    music_duck_db: -4
    room_tone_duck_db: -1.5
    recovery_ms: 480

Priority is not a fixed level. Music can lift briefly in a pause, ambience recovers after an action peak, and the whole hierarchy changes when a character enters subjective hearing. Draw the narrative automation first and let compressors and dynamic EQ execute it; relying on one uniform sidechain pushes music down identically on every sentence and flattens the performance's breathing.

46.20 Platform transcoding and a real device matrix#

A lossless master passing does not mean the released file passes. Platforms may change loudness, encoding, channels, peaks and high-frequency character. Create a test upload per platform, download or stream the actual transcode, and listen back on a low-end phone, a mid-range phone, headphones and a small speaker. Record the device, the volume setting, ambient noise and the timecodes of any problem.

The aim is not identical results on every device. It is that the core information survives: key dialogue is intelligible, principals are distinguishable, action has causality, music state remains perceptible, and there is no harsh sibilance or pumping. Sub-bass and wide stereo can shrink on a phone, provided the narrative function has a midrange or temporal backup.

Platform masters derive from the same approved mix, each storing its limiter, true peak, encoding headroom and measurement report. Do not overwrite the lossless master to raise loudness for one platform, and do not let a compressed file from a test upload become the source for the next version.

46.21 Revisions, international versions and M&E#

Group the session from day one into dialogue, narration, music, ambience, Foley and design, exporting stems from one time zero. Decide whether dialogue-adjacent fabric, breath and action sound belongs to the performance or to M&E; if it is all printed into the Chinese dialogue, characters lose their bodies when the language is removed for an international version.

When one line changes, replace the dry recording and the lip sync first, then recompute the dialogue processing, ducking, subtitles and local loudness — never paste a new line onto an already limited full mix. When music changes, check phase, dynamics and key lines. When picture changes, check every synchronized event. Each kind of change has a minimum regression pack.

An international version is not muting the Chinese track and dubbing another language over it. Duration, stress and cultural register differ, which may require re-cutting picture, rebuilding cue elastic regions and redoing subtitles. M&E supplies the shared world, and each language still needs full device and platform acceptance.

A note on sources#

Loudness numbers are only a starting point. Real acceptance happens on target devices in real environments, judged by dialogue clarity, emotional hierarchy and consistency across episodes.