中
Chapter 49. The Subtitle System

Part IX — Editing, Subtitles, Pacing, and Finishing

Chapter 49. The Subtitle System#

In this chapter
49.1 Subtitles come from the final audio49.2 Break by sense unit49.3 Timing49.4 Position and avoidance49.5 Typeface and style49.6 Speakers and narration49.7 Numbers and proper nouns49.8 Multiple languages49.9 Accessibility49.10 Subtitle QC49.11 Reading speed and exposure time49.12 Priority between subtitles and evidence49.13 Subtitle source of truth and the revision chain49.14 Line breaks are an interpretation of the performance49.15 Subtitling the amount in Backlit Takeover49.16 SOP for subtitles49.17 Fault tree49.18 Checklist, exercises and deliverables49.19 Subtitle objects bind to audio and text versions49.20 Dynamic avoidance is a timed track49.21 Multilingual subtitling is not equal-length substitutionA note on sources

49.1 Subtitles come from the final audio#

Script text is not a source of timing. Force-align once dialogue and narration are locked, then correct names, numbers, interruptions and overlaps by hand. Any audio change invalidates the related subtitles.

49.2 Break by sense unit#

Break lines by meaning and breath, not by a fixed character count. Keep verb and object, name and title, number and unit together wherever possible. Each screen carries only what the audience can read in one pass.

49.3 Timing#

A subtitle may appear a very short time before the sound to aid reading, without revealing a reversal early. The out point gives enough reading time without running into the next speaker. Interrupted lines can use a dash or simply cut, preserving the performance.

49.4 Position and avoidance#

The default is the lower safe area, with faces, hands, documents, phones and platform UI taking priority. Each shot can define a subtitle avoidance map. When a moving prop enters, subtitles need automated position or shorter text.

49.5 Typeface and style#

Choose a typeface that is legible on phones and clearly licensed. Test size, leading, outline and shadow on light and dark backgrounds. Use emphasis only for a few key words — names, amounts, results — and avoid a screen full of color.

49.6 Speakers and narration#

Most dialogue is identified by the shot and needs no name tag per line. Off-screen lines, multi-person overlaps and narration can be distinguished by a consistent style, and color must not be the only cue, for viewers with color vision differences.

49.7 Numbers and proper nouns#

Display and pronunciation may differ while the fact stays identical. The voice can say a figure in words while the subtitle uses digits. Keep the format consistent across the series and reference the fact ledger.

49.8 Multiple languages#

After translation, re-break, re-time and re-avoid. Languages differ in length, so Chinese positions and sizes cannot be copied across. Keep subtitles editable rather than burning the only version into the master.

49.9 Accessibility#

Where needed, annotate key off-screen sounds, phone calls, door sounds and music information — while a commercial episode's subtitles should not drown in unrelated effect labels. Produce a separate accessibility caption version.

49.10 Subtitle QC#

Check text against sound, speaker, spelling, timing, reading speed, breaks, platform occlusion, avoidance of faces and proof objects, font licensing, and overflow in other languages. Play on a real phone at normal speed; do not pause to admire.

subtitle_event:
  id: SUB_E001_012
  start: 00:31.200
  end: 00:33.500
  speaker: CH_LINYUN
  text: Then the acquirer's representative will attend.
  emphasis: acquirer's representative
  position: lower_center_shifted_up
  avoid: PROP_AUTHORIZATION_title

49.11 Reading speed and exposure time#

Reading speed should not rest on one fixed character-count threshold. Short but unfamiliar names, amounts and company names need longer than ordinary speech; familiar short phrases can go faster. Exposure time is also affected by picture complexity: when the audience must watch a face, a document and an action at once, reading capacity falls.

In practice, generate subtitles from speech timing first, then check each item's character density, minimum exposure, shot changes and cognitive load. When it cannot be read, first rewrite it conversationally, split the sense units, extend the shot, or let the picture carry the information. Reducing the type size comes last. Type size is not a capacity control.

49.12 Priority between subtitles and evidence#

When dialogue is explaining a screen or a document, subtitles easily cover that same evidence. Establish a visual priority per shot: faces, action, primary proof fields, subtitles, decorative graphics. If a primary proof field must be read, subtitles can move up, split into two, delay their entry, or the line can carry over to a reaction shot.

A subtitle avoidance map is not one fixed safe box. It can carry time: the lower area is available from 0 to 1.2 seconds; the authorization enters the lower center from 1.2 to 2.4, so subtitles move up; after 2.4 they return. Automated systems should read prop and face tracks rather than centering mechanically for a whole episode.

49.13 Subtitle source of truth and the revision chain#

Subtitle text comes from a transcript of the final audio, and names, amounts, dates and terminology must be verified against the fact ledger. Speech recognition output is a candidate, never an override of project truth. Each subtitle stores an audio_hash and the terminology list version; a change to the audio or to a factual field marks it stale automatically.

When someone shortens a subtitle for readability, distinguish compressing the display from changing the line. Compression must not change a promise, a negation, a number, or who is responsible. If the voice says one person did not authorize another to transfer funds, the subtitle cannot compress away the subject, because the relationship may be story evidence.

49.14 Line breaks are an interpretation of the performance#

The same sentence can be broken to emphasize the judgment, or broken elsewhere to produce a different ambiguity. Subtitle breaks follow the speaking action and the stress rather than merely balancing visual line lengths.

Preserve the form of interruptions, lies, self-corrections and talking over someone. An unfinished sentence completed by the subtitle reveals the character's intent early. Stutters and repetitions need not all be transcribed literally; keep the parts that matter to the performance.

49.15 Subtitling the amount in Backlit Takeover#

Lin Yun says that three years ago, on the same day, three transfers totalling a large sum went out. The original subtitle put the whole sentence on two lines while the picture displayed three ledger rows, and a phone could not serve both. The new version splits it into two sense units: the time phrase appears first with the date column highlighted; on the cut to the amount column, the subtitle shows the count of transfers and the total in digits.

No fact is lost, and voice, subtitle and graphics now divide the work. The amount's format is unified by the fact ledger, and the ad version, the episode and the English version all reference the same field.

49.16 SOP for subtitles#

First, lock the final audio. Second, force-align. Third, break by sense unit. Fourth, unify proper nouns and numbers. Fifth, apply platform safe areas. Sixth, build avoidance for key shots. Seventh, review on a phone. Eighth, export editable, burned-in and accessibility versions.

49.17 Fault tree#

Symptom: subtitles drift further as it goes on. Timing was estimated from the script. Re-align to final audio.

Symptom: subtitles cover the document. There is no per-shot avoidance map. Move, shorten or delay the subtitle.

Symptom: the type is too small. You are trying to fit a complete written sentence. Rewrite, split across screens, or let the picture carry the information.

Symptom: other languages overflow. Only translation happened, with no re-layout. Do breaks and type-size testing independently.

49.18 Checklist, exercises and deliverables#

Check that timing comes from audio; that breaks follow sense units; that nothing is revealed early; that key picture areas are avoided; that the font is licensed; that numbers are unified; that it is legible on a phone; and that other languages are tested independently.

Exercise one: produce line-level timecode for a minute of audio. Exercise two: re-lay five subtitles that cover proof objects. Exercise three: localize Chinese subtitles into a language that runs longer.

Deliverables for this chapter: forced alignment, the subtitle track, the style guide, the avoidance map, the terminology list, accessibility captions, and subtitle QC.

49.19 Subtitle objects bind to audio and text versions#

Every subtitle stores its text source of truth, the dialogue master hash, word-level timing, speaker, language, style and placement policy. After an audio take change, the old alignment invalidates automatically. When the script changes wording without re-recording, the system reports a content conflict. Subtitles must never become a fourth independent version of the lines.

subtitle_item:
  id: SUB_E001_042
  speaker: CH_LINYUN
  audio_ref: DLG_E001_B04_MASTER@sha256
  text_ref: LINE_E001_B04@v07
  in: 00:41.080
  out: 00:43.420
  break_after: authorization
  placement_policy: avoid_document_lower_center

Forced alignment supplies candidate timings, and humans still correct for performance and reading. A pause does not always mean a new subtitle; an interruption may require an early disappearance; a breathy line may need longer exposure. The machine synchronizes fact; the person interprets delivery.

49.20 Dynamic avoidance is a timed track#

Subtitle position cannot be fixed per shot. Faces, gestures, evidence and platform UI move, so the avoidance map changes its safe region over time. When subtitles move from the bottom to the top, do it at a natural break rather than jumping up and down every second, and keep a stable baseline within a scene.

Figure 49-1 Dynamic subtitle avoidance as faces, evidence and platform UI change

Figure 49-1 In the first phase subtitles use the lower area. When evidence lifts into frame, subtitles migrate upward at a break. Once platform UI appears, the right side stays excluded. Dynamic avoidance is not subtitles chasing objects around — it is disciplined switching between a small number of stable layouts.

The readable-area diagram on the right updates by platform and shot. The green face region, red evidence region and blue platform region are not absolute rectangles; they are attention occupied in the current frame. Subtitles take the remaining area while preserving minimum type size, line length and visual stability.

Avoidance priority is usually: critical evidence and mouths may never be covered; eyes and hand action come next; ordinary background may be covered. When no usable area remains, re-time the subtitle, split the sentence, or adjust the shot — do not shrink the type below legibility.

Accept burned-in and sidecar subtitles separately. Sidecar styling may be controlled by the platform, so supply safe position assumptions. Burned-in versions get checked for compression edges, outline, color contrast and stability against different backgrounds.

49.21 Multilingual subtitling is not equal-length substitution#

Translation preserves the line's function, relationships and information first, then handles length. English, Spanish and other languages may run longer than Chinese and need fresh breaks and exposure; Chinese timings cannot simply be copied. Key proper nouns, amounts and legal terminology are locked by the terminology list and reviewed by a native speaker.

Keep cultural rewriting separate from factual translation. Where forms of address, kinship or business structures are unfamiliar in the target market, they can be clarified in the line or the subtitle — without changing story evidence. Record the reason and the approver for each rewrite, so different episodes do not use different renderings.

Multilingual QA includes spot back-translation, comparison against picture, reading speed, punctuation, local number and date formats, sensitive terms and platform character support. Correct subtitles are not enough; they must be legible against the real picture on real devices.

A note on sources#

Subtitles are a core narrative layer on mobile, not decoration generated automatically at the end in an editing application. They must be designed together with audio, composition, graphic facts and platform UI.