Research
DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation
Overview Research area: AI-generated video and multimodal evaluation, specifically benchmarking short-drama (micro-drama) generation as a multi-stage production chain rather than a single video-clip t

In inglese
- arXiv
- 2609.00646
- Published
- 2026-09-01
- Authors
- Haoyuan Shi, Mingtao Chen, Shuo Jiang, Ziyan Chen, Xuyi Sheng, Yiming Liu, Ying Zhang, Miao Wang, Jianxiang Lu, Fanyang Lu, Songyuanyi Lu, Xiele Wu, Zhichao Hu, Yuhong Liu, Richeng Xuan
AI summary
Overview
Research area: AI-generated video and multimodal evaluation, specifically benchmarking short-drama (micro-drama) generation as a multi-stage production chain rather than a single video-clip task.
Technical level: Advanced. It assumes familiarity with multi-stage generative pipelines, agent frameworks, rubric-based human annotation, and correlation metrics such as PLCC/SRCC.
Scope: One sentence: DramaChain Bench is an end-to-end benchmark that scores all six stages of a commercial-style short-drama production chain — script/storyboard, keyframes, shot video and the finished drama — on items the chain itself produced, using a shared 63-dimension rubric, three-annotator human labels with localised defect attributions, and an agentic automated judge that reproduces the human model ranking.
What This Paper Is About
Existing video-generation benchmarks test only the video stage, and they do so with inputs that were authored for the test (reverse-engineered prompts, shots cut from existing footage) rather than produced by the pipeline that would feed that stage in deployment. That makes two questions unanswerable: whether each stage actually honours the original script intent rather than just its immediate prompt, and whether individual shots stay coherent once assembled into multi-episode releases. DramaChain Bench addresses this by building a benchmark on a self-built short-drama production pipeline, annotating every stage exhaustively, and validating an automated judge that can score new models without new human annotation.
Key Contributions
- A chain-level evaluation design (DramaChain Dimensions): five evaluation axes — input fidelity (F), internal consistency (C), generation plausibility (P), visual quality (Q) and cinematic expressiveness (E) — defined once and instantiated at six production granularities, resolving into 63 leaf dimensions over 5,785 items. A defect can therefore be followed along the chain (for example, a storyboard is scored on P for executability, and the shots made from it are scored on P again).
- A pipeline that generates, rather than authors, the evaluated items (DramaChain Agent): built on the ViMax framework with its chain reverse-engineered from commercial platforms, producing 20 short dramas × 3 consecutive episodes = 60 episodes, with forks occurring only at the stage under test so competing models receive identical upstream artefacts and prompt strings.
- An annotation system with mandatory spatio-temporal attribution (DramaChain Labeling System): three professional annotators independently score each item; any deduction below full marks must be localised in space or time and tagged from a fixed defect vocabulary, yielding 17,488 valid scores and 255,925 traceable attribution records from 543 annotators.
- A validated automated judge (DramaChain Agentic Judge): a multi-round agentic scorer that discards low-correlation metrics, gathers evidence via routed measurements and tool calls, and judges against a per-item checklist written in the same rubric text and tag vocabulary as the humans — reaching a mean PLCC of 0.918 against the human panel.
Main Findings
- The strongest chain barely clears the usability bar. Using gpt-5.5-xhigh for storyboarding, gpt-image-2 for keyframes and seedance-2.0 for video, the finished short drama scores 3.30 on a 5-point scale, against 3.0 as the usable line — a margin of only 0.3 — and the best single-shot result is 3.24.
- Quality decays progressively down the chain. Average scores fall from 3.78 (storyboard design) to 3.62 (single keyframes), 3.30 (episode keyframes), 3.09 (single-shot video) and 2.93 (episode video), with a modest rebound to 3.03 for the finished short drama. No pixel- or video-generation stage reaches the text stage's 3.78, and the three video granularities all sit within 0.1 of the usable line.
- Defects cascade instead of staying local. Degrading one upstream stage costs a clip 0.07–0.13 points but the finished short drama 0.25–0.83 points, so shipped quality is not determined by video generation alone.
- Leading models are invariant to granularity within a modality. gpt-image-2 tops both keyframe stages (3.99 and 3.82), and seedance-2.0 leads all three video stages (3.24, 3.11, 3.30).
- Discrimination tracks reference explicitness, not difficulty. Average top-to-bottom gaps are 1.22 for F input fidelity, 0.70 for C internal consistency, 0.64 for Q visual quality, 0.53 for P generation plausibility and 0.42 for E cinematic expressiveness, because F is validated against upstream outputs and C against sibling artefacts, while P, Q and E have no external referent. Human annotation separates models most on F too, but ranks E second widest where the automated board ranks it last.
- Composite scores hide large dimensional trade-offs. seedream-5.0-pro and nano-banana-pro differ by only 0.02 in composite on single keyframes, yet seedream-5.0-pro leads by 0.39 on F and lags by 0.60 on P. For single-shot video, kling-3.0-omni and pixverse-c1 differ by 0.01 overall while kling-3.0-omni is 0.16 higher on P and 0.22 lower on E.
- The automated judge reproduces the human board. Mean model-level PLCC is 0.918 and mean SRCC 0.881 across the six stages, with SRCC = 1.000 for short dramas. Per-stage PLCC ranges from 0.755 (single-shot video) to 0.973 (episode keyframes).
- Item-level agreement is uneven across stages. Against the leave-one-out human baseline (hb), storyboard design reaches 2.35× baseline consistency and single keyframes 1.19×, while all three video stages fall below baseline (0.79×, 0.85× and 0.85×). Item-level PLCC ranges from 0.203 (episode video) to 0.536 (single keyframe), and the mean absolute error is 0.495 with 38.2%–67.5% of scores landing within ±0.5 of consensus.
- Most off-the-shelf specialist metrics do not track human judgement. Their absolute SRCC is no greater than 0.2; fourteen photography-oriented metrics correlate negatively, and AIGC-specific perceptual models stay below 0.2, because photography-domain priors assume real photographs while AIGC models treat short-drama stylisation as anomalous. Three leaf dimensions, including V-D2, are invalidated by near-random human baselines.
- The judge can admit new models for free. Eight additional models were scored by DramaChain Agentic Judge alone on exactly the same items, so they enter the board without additional annotation cost.
Methodology in Plain English
The authors did not write test prompts and collect outputs; they built the production line first. DramaChain Agent runs the same six-stage chain commercial platforms run — write the script, break it into a storyboard, render keyframe images, animate them into shot-level video, assemble the finished drama — fully automatically, and also produces the intermediate artefacts a real line needs (character portrait sheets, scene and element reference sheets, per-shot descriptions, a camera tree). To make sure the benchmark measures the models and not the pipeline, the chain was benchmarked against three commercial short-drama platforms end to end with no human editing, frame selection or best-of-n reruns; expert review found shot division, character and scene consistency, camera language and audio landing in the same band, with differences being stylistic.
Diversity was controlled by construction: 20 dramas across channel, period and genre (20 distinct sub-genres, one per drama), crossed with visual style (live action 10 : animated 10) and dialogue language (Chinese 16 : English 4, with scene descriptions all in Chinese). Each shot slot is generated once by each participating model, and the pipeline forks only at the stage under test, so everything upstream is shared.
Annotation was treated as an engineering artefact. Three annotators score every item independently on a five-level decidable rubric, and any deduction must be localised (bounding box, time interval, or shot index) and tagged with a reason from that dimension's fixed list. Recruiters screened by stage — storyboard annotators needed directing or screenwriting backgrounds, others needed substantial image or video annotation experience — across supplier teams with their own quality-control reviewers, who sample one-fifth of outputs by default. Nine heuristic rules flag submissions for revision, including annotation speed far below the granularity-matched median, deductions with no attribution, and near-zero correlation with leave-one-out consensus.
For automation, the authors first audited the specialist-metric stack they had assembled (detection/segmentation, identity and visual similarity, scene consistency, style consistency, aesthetic quality, motion, audio analysis and cross-model corroboration) against human labels, and discarded what did not agree. The surviving measurements act as auxiliary cues, never as direct scores. DramaChain Agentic Judge then scores each leaf dimension independently over multiple agentic rounds, routing relevant measurements to each target dimension, optionally invoking perception tools such as regional zoom or timestamp-based frame extraction, and producing verdicts that carry the same defect tags and localisation as human annotations so disagreements can be cross-checked directly. Stage-specific judging configurations were used: storyboard design averages gpt-5.6-sol-xhigh and a claude-fable-5 variant, keyframes are judged primarily by doubao-seed-2.1-pro, and shot-level video by gemini-3.1-pro.
Why This Matters
Impact on research. It shifts evaluation from "is this clip good?" to "which stage of the chain caused this defect, and how far did it travel?", supplying the per-stage, per-dimension, localised supervision that prior benchmarks lack. It also delivers a cautionary result for the wider metric ecosystem: most existing specialist metrics score below |SRCC| 0.2 against expert short-drama judgement, so hybrid pipelines that feed small-model measurements straight into a large-model verdict may be building on signals that do not track human perception. Finally, it shows a validated automated judge can expand a benchmark board without re-annotating, which changes the economics of leaderboard maintenance.
Real-world applications:
- Production-line diagnosis for short-drama studios: identifying whether a delivery defect originates in the storyboard, the keyframes or the video model, rather than blaming the video stage by default.
- Model selection and procurement: choosing models per dimension (F versus P versus E) instead of per composite score, since models that tie on the composite can differ by 0.39–0.60 on a single axis.
- Automated quality control at scale: the agentic judge's defect tags and localisation can be used for triage on new episodes without a fresh three-annotator panel, which the paper states costs a median 23 minutes per item.
- Evaluating agent frameworks and generation pipelines as systems: the fork-only-at-the-stage-under-test design allows one-stage substitutions and apples-to-apples comparison of competing components.
Industry relevance. The paper frames short drama as one of the most important deployment scenarios for generative video, citing global micro-drama revenue of USD 11 billion in 2025 expected to reach USD 14 billion in 2026, with China at roughly 83% of 2025 revenue, Omdia forecasting USD 3 billion of non-China revenue in 2026, 33,000 titles drawing almost 700 million domestic viewers in 2025, and regulators noting rapid growth in AI-generated animated micro-dramas. Commercial platforms (OiiOii, Flova, XiaoYunQue, LibTV) and academic agent systems have converged on the same staged paradigm, which is what makes per-stage evaluation both feasible and necessary. The authors state they will release DramaChain Dimensions, the DramaChain Agentic Judge framework, and a curated subset of benchmark data.
Future Directions
- Extend to the unevaluated input paradigm. DramaChain Agent implements four ways to drive the video stage (fflf2v, grid2v, multiref2v and keyframe2v), but only the first three are evaluated here — keyframe2v remains untested.
- Close the item-level agreement gap on video. Automated scoring beats the human leave-one-out baseline at the text and image stages but falls below it at all three video stages (0.79×, 0.85×, 0.85×), and the agent's own bias runs negative for single-shot video (−0.35) while running positive elsewhere. Narrowing this is the obvious next step if the judge is to replace annotation outright.
- Broaden coverage beyond the current corpus. The benchmark covers Chinese 16 : English 4 with all scene descriptions in Chinese, 20 dramas and 3 consecutive episodes each; whether the rubric and judge hold for other languages, longer series and other production cultures is not established by the reported results.
- Decide what to do with weakly anchored dimensions. Three leaf dimensions, including V-D2, are invalidated by near-random human baselines, and P, Q and E show smaller top-to-bottom gaps than F and C. The paper argues such gaps reflect benchmark limitations rather than comparable model quality, leaving open how those axes should be re-specified.
Target Audience
Researchers and engineers working on generative video, multi-agent media-production pipelines, and automated evaluation — particularly those building benchmarks for multi-stage systems or training judge models that must agree with expert human panels. It is also directly useful to practitioners at short-drama and streaming platforms choosing models per production stage or setting up automated quality control, and to evaluation researchers interested in rubric design, defect attribution, and the reliability limits of off-the-shelf perceptual metrics. Readers need some grounding in benchmark methodology and correlation-based agreement analysis; the production-pipeline framing itself is explained accessibly.
Authors’ abstract
Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video, and the finished short drama. Most existing benchmarks evaluate solely the video-generation stage using pre-authored inputs instead of real upstream pipeline outputs. This leaves two critical questions unanswerable: whether each stage adheres to the original script intent (rather than only its immediate input prompt), and whether disparate shots remain coherent after assembly into multi-episode releases. We present DramaChain Bench, the first short-drama benchmark that evaluates every stage of the complete production chain. It is built upon three in-house systems sharing one dimension system, DramaChain Dimensions: five evaluation axes instantiated at every stage, resolving into 63 leaf dimensions. DramaChain Agent is calibrated against commercial short-drama platforms in both workflow and finished short-drama quality, enabling stage-wise fair comparison across models. DramaChain Labeling System has each of the 5,785 items scored independently by three professional annotators, with all defects spatio-temporally localised and selected from a predefined defect list. This process produces 17,488 valid scores and 255,925 traceable attribution records. The human annotations confirm that upstream defects cascade across the pipeline, demonstrating that final episode quality is not governed by video generation alone. DramaChain Agentic Judge then scores every leaf dimension automatically, gathering evidence over multiple agentic rounds before judging against a per-item checklist; it reproduces the model ranking at a mean PLCC of 0.918, enough to admit new models at no annotation cost.