中
Chapter 61. Test Design and Metric Diagnosis

Part XII — Creatives, Distribution, and Data Iteration

Chapter 61. Test Design and Metric Diagnosis#

In this chapter
61.1 Metric diagnosis, the unit of experiment, and real ROI61.2 Comment evidence, decision rules, and the operating flow61.3 Metric definitions, sampling uncertainty, and data quality61.4 Diagnostic layers, experiment conflicts, and the event dictionary61.5 Metric governance, causal estimation, and stopping disciplineA note on sources

61.1 Metric diagnosis, the unit of experiment, and real ROI#

Metrics are symptoms in a funnel.

Impression to click, three seconds, qualified view, completion, next episode, cut point, payment and payback each answer a different question. Play count cannot prove profit, and click-through cannot prove the story.

The unit of experiment.

Fix audience, platform, time, creative, landing version, cut point and budget. Randomize or otherwise control everything else. With insufficient sample, record a direction — do not declare a law.

The metrics ledger.

experiment:
  id: EXP_HOOK_001
  hypothesis: proof reversal beats bloodline secret
  variants: [AD_PROOF_A, AD_IDENTITY_B]
  controlled: [audience, bid, destination, duration]
  primary_metric: qualified_view_rate
  guardrails: [E001_completion, E002_start]
  decision_rule: defined in advance

Reading combinations.

Low click-through points at the first frame and the promise. High click-through with poor three-second retention points at whether the entrance is honored. Good three-second retention with poor completion points at the middle. High completion with low next-episode entry points at the closing question. High return with low payment points at the cut point, the price or the payment path. High payment with poor later retention points at repayment.

ROI and true cost.

Revenue minus media, platform, production, failed generation, labor and licensing is the project's return. Popularity and gross top-ups are not profit.

61.2 Comment evidence, decision rules, and the operating flow#

Coding comments.

Code comments by comprehension, taking sides, prediction, demanding the next episode, questioning the logic, complaining about the picture, complaining about payment, and asking about the schedule — linked to timecode and version. Do not chase one highly upvoted opinion.

Write the decision rule first.

Before the experiment, define success, continue-collecting, stop and anomaly handling, so nobody picks a metric to explain the result afterwards. Platform volatility and small samples must retain their uncertainty.

SOP.

First, write a falsifiable hypothesis. Second, choose primary and guardrail metrics. Third, control the variables. Fourth, set budget and stop conditions. Fifth, check the data path. Sixth, run and record versions. Seventh, diagnose using comments together with episode data. Eighth, take only the next step the evidence supports.

Fault tree.

Symptom: good clicks and no profit. Upper-funnel metrics were optimized while cut point, cost and payback were ignored.

Symptom: every conclusion differs. Sample, entrance and version were uncontrolled. Rebuild the experiment record.

Symptom: the team only reads averages. Different hooks attract different people. Segment by creative family and subsequent behavior.

Checklist, exercises and deliverables.

Check that the hypothesis is falsifiable; that primary and guardrail metrics were chosen in advance; that variables are explicable; that the data path is connected; that cost is complete; that uncertainty is retained; and that comments are coded.

Exercise one: write an experiment for two hook types. Exercise two: diagnose four funnel combinations. Exercise three: compute ROI including labor and failed generation.

Deliverables for this chapter: the experiment plan, the metrics ledger, decision rules, the funnel diagnosis, comment coding, and the ROI report.

61.3 Metric definitions, sampling uncertainty, and data quality#

Metric definitions must be fixed.

Whether three-second retention is denominated by playback starts, qualified impressions or clicking users produces different results. Each metric records its formula, event names, window, deduplication rule, time zone and data source. Platform-native data and your own payment data join on a shared experiment ID.

metric_definition:
  name: E001_completion
  numerator: unique_users_reaching_95_percent
  denominator: unique_users_starting_E001
  window: 24h_after_click
  exclusions: [internal_test, duplicate_device_events]

After a definition changes, do not plot it continuously against history; mark a break or recompute.

A worked experiment.

Suppose proof version A and identity version B run against the same audience, budget and episode. A has lower click-through and higher three-second retention, E001 completion and E002 entry. B clicks harder and sheds viewers in the episode's first ten seconds. The conclusion is not that proof is always better; it is that the identity fantasy B sells was not sufficiently honored in the landing's opening.

Two distinguishable follow-ups exist: change B's landing so the identity authority appears sooner, or keep the landing and adjust B's promise. Do not change both the ad and the episode and then announce a cause.

Attribution windows and cross-device behavior.

A user may pay later, or watch on another device. The window chosen changes ROI. Record direct click, view-assisted and remarketing attribution separately, so one conversion is not credited to several creatives.

Where platform black-box data cannot be fully verified, report the limitation rather than manufacturing precise attribution. Make the next decision from trends and controlled experiments.

Sample and uncertainty.

No single sample size suits every project. Baseline, effect size, volatility and the cost of the decision determine it together. Expensive scale-up needs stronger evidence; cheaply eliminating an obvious failure can move faster.

Report intervals and sample sizes, not only percentages to one decimal. With insufficient sample the status is inconclusive, not a forced winner. Repeatedly peeking and stopping at a high point inflates false positives, so define the checking cadence in advance.

Data quality checks.

Before spending, test the events: impression, click, playback start, three seconds, completion, next episode, cut point, payment success, refund. Check for loss, duplication, time zones, versions and bot traffic. When the event chain breaks, pause commercial conclusions.

One platform update can change autoplay or a metric definition. When a number moves sharply, check instrumentation and traffic composition before explaining it as a content change.

61.4 Diagnostic layers, experiment conflicts, and the event dictionary#

Complete ROI.

contribution profit = revenue
                    - media spend
                    - platform and payment fees
                    - refunds
                    - amortized episode and creative production
                    - failed generation and human repair
                    - licensing, support and infrastructure

Forecasting scale must also account for creative fatigue, marginal traffic becoming more expensive, and producing later episodes. Positive ROI on a small sample does not grow linearly with budget.

The diagnostic matrix.

High clicks with low completion: check promise mismatch, the opening shot and the middle. Low clicks with high completion: the episode may be good and the creative weak. High completion with low cut-point conversion: check accumulation, price and payment friction. High payment with low later retention: check repayment. Every layer healthy and profit poor: check media price, refunds and production cost.

Any metric anomaly should produce at least two competing explanations and one experiment that distinguishes them.

Segmentation and heterogeneity.

An average can hide opposite users. Segment by new versus returning, creative family, device, region, payment history and entrance — with a business reason stated in advance rather than hunting for an accidental winner across dozens of slices.

A creative that is slightly worse overall may work well in remarketing to people who have watched the first three episodes; move where it runs rather than eliminating it. Reports display sample size and uncertainty per segment.

Experiment conflicts.

One person entering several creative and pricing experiments at once contaminates the result. An experiment registry manages mutually exclusive groups, priority and cooldown. When an urgent platform change or a major sale creates an external shock, mark the data window rather than attributing environmental change to content.

An event dictionary comes before dashboards.

Before drawing any chart, define impression, click, playback start, qualified three seconds, 95 percent completion, next-episode start, payment page open, payment success, refund and day-seven retention. Each event records its trigger condition, deduplication key, anonymized user ID, creative ID, episode version, timestamp, platform and schema version.

event_definition:
  event: episode_complete
  fires_when: playhead_reaches_95_percent
  dedupe: user_id + episode_id + episode_version + 24h
  required_dimensions: [creative_id, destination_id, platform, locale]
  late_arrival_window: 48h
  owner: analytics

Without episode_version, data from before and after a re-cut mixes. Without creative_id, you cannot tell which promise brought high-quality users. Without refunds, high top-ups can conceal dissatisfaction.

61.5 Metric governance, causal estimation, and stopping discipline#

Primary, guardrail and diagnostic metrics.

The primary metric decides the experiment. Guardrails stop local optimization damaging the whole. Diagnostics help explain. Comparing proof against identity, the primary can be E002 entry rate, the guardrails post-payment retention into episode three and refunds, while click-through, three-second retention and comments serve diagnosis.

Allow one primary decision metric per experiment, or the team will pick afterwards when results are mixed. A guardrail can block a winner: if the identity version raises E002 entry while post-payment retention falls significantly, it does not scale.

The decision memo.

An experiment ends with a one-page decision memo rather than charts alone: hypothesis, versions, sample, data quality, the primary result, intervals, guardrails, competing explanations, the decision, the next experiment, and prohibited inferences. The last item prevents conclusions inflating — "the proof version won in this cold-start audience" must not become "every market prefers professional revenge."

decision:
  status: directional_win
  winner: AD_PROOF_A
  evidence: E002_start_up, refund_unchanged
  uncertainty: sample_small_in_returning_users
  next: test_first_frame_within_proof_family
  do_not_infer: identity_family_is_globally_invalid

A funnel diagnosis exercise.

Fictional results: the proof version at 1.8 percent click-through, 61 percent E001 completion, 44 percent E002 entry; the identity version at 2.6 percent click-through, 39 percent completion, 21 percent entry. The identity version is not a better creative — it is a stronger click with a weaker match. The next step holds the identity promise and tests a landing that confirms the identity faster; only if that still fails does its budget fall.

Meanwhile the proof version is stable among viewers aged 25–34 and clicks less while paying more among those over 45. Budget allocation should follow contribution profit rather than overall click-through rank. All figures here are teaching fictions; a real project must use its own baselines and uncertainty.

Use a causal diagram to state what you are estimating.

Platform, audience, time, bid, creative and landing version all move at once, so comparing two group averages easily misleads. Before an experiment, draw the minimum causal diagram: the treatment variable, the outcome, confounders, mediators and guardrails. A creative affects clicks, and click selection then affects who enters the episode — so post-click retention differences cannot all be attributed to the episode.

Prefer randomization. Where it is impossible, record the assignment mechanism and use stratification, matching or time controls cautiously. No statistical method repairs missing versions and wrong events. This book offers no universal significance threshold; a real project should have someone with experimental experience design it against risk, sample and business cost.

Holdouts, long-term effects, and cannibalization.

A short-term click win may only have attracted people who would have watched anyway, or moved conversions from another creative family. Retain a small holdout to observe incremental entry, payment, refunds and long-term retention. With sufficient budget, assess audience cannibalization between creatives.

Long-term metrics need not block every fast decision. Use tiered gates: day-one data eliminates obvious mismatches, a few days decide limited scale-up, and a longer window verifies contribution profit and retention. Each stage makes only the decision its evidence supports.

Sequential experiments and the discipline of stopping.

Watching continuously and stopping at the first lead exaggerates the winner. The plan specifies minimum sample, checking windows, and conditions for winning, futility, risk and stopping. Where sequential decisions are genuinely needed, use appropriate methods and retain a record of every look rather than quietly re-comparing in the background.

Commercially, you may stop early for P0/P1 issues, refund anomalies or platform risk. Changing the primary metric because you dislike the result is not permitted. A decision memo is still produced after an early stop, stating why it stopped and which inferences cannot be drawn.

A note on sources#

Distribution data is evidence about a specific entrance, version and moment. Fixed definitions, honest attribution and pre-committed decision rules are what turn that evidence into decisions a team can defend later.