Research
DataMagic: Authoring Data Videos through Declarative Multi-Agent Orchestration
DataMagic: Authoring Data Videos through Declarative Multi-Agent Orchestration Overview Research area: Human-Computer Interaction — data video generation, data storytelling, declarative visualization

- arXiv
- 2609.33403
- Published
- 2026-09-27
- Authors
- Yupeng Xie, Zhenyang Wang, Liangwei Wang, Jiayi Zhu, Zhouan Shen, Yuyu Luo
AI summary
DataMagic: Authoring Data Videos through Declarative Multi-Agent OrchestrationOverview
- Research area: Human-Computer Interaction — data video generation, data storytelling, declarative visualization specification, and multi-agent systems.
- Technical level: Advanced. The paper combines a formal declarative intermediate representation (DVSpec) with a multi-agent generation pipeline and a full interactive web system.
- Scope: The paper presents DataMagic, a system that generates multi-scene data videos directly from raw tabular data and a user query, using a declarative specification and a "Generate-then-Orchestrate" multi-agent strategy, evaluated on 109 real-world samples plus a within-subjects user study (N=12).
What This Paper Is About
Producing data videos — dynamic charts combined with voice narration and synchronized animation — normally requires expertise across data analysis, narrative design, and video editing. Existing tools either generate static charts with no narrative or animation, require pre-made charts as input, or use pixel-level video models that cannot guarantee numerical accuracy or trace visual elements back to data records. The paper's goal is end-to-end generation from raw tabular data to a complete, data-faithful, narratively coherent data video, while still allowing humans to refine the result.
Key Contributions
- DVSpec — a declarative specification for data videos that unifies charts, narration, and animations with their temporal relationships through data-driven semantic references and narration-indexed triggering.
- Multi-agent framework — a "Generate-then-Orchestrate" two-stage strategy for parallel candidate scene generation and global narrative orchestration.
- Interactive system — a complete web-based system supporting three complementary interaction modes (canvas manipulation, script editing, natural language commands) over a shared DVSpec state.
- Comprehensive evaluation — systematic experiments on 109 real-world samples and a within-subjects user study (N=12) confirming practical utility.
Main Findings
- Direct LLM generation is weak: Even the most advanced LLM evaluated (GPT-5) achieves only 2.13/5 average quality, with execution success rates between 48.62% and 86.24% across the four direct-generation baselines (DeepSeek-V3.2 at 48.62% and 1.91 average, Gemini-2.5-Pro at 66.06% and 2.22, GPT-5 at 86.24% and 2.13, Claude-Sonnet-4 at 84.40% and 2.18).
- DataMagic raises quality and reliability: DataMagic improves quality to 3.89 (+83%) with success rates above 95%. Its best configuration, DataMagic (Claude-Sonnet-4), reaches 98.17% execution rate and a 3.89 average score, with the largest gains in the Animation and Narrative dimensions.
- Model-agnostic behavior: DataMagic variants across DeepSeek-V3.2 (96.33%, 3.44), Gemini-2.5-Pro (97.25%, 3.47), GPT-5 (95.41%, 3.38), and Claude-Sonnet-4 (98.17%, 3.89) all substantially outperform their direct-generation counterparts, confirming the framework does not depend on one specific LLM.
- Anti-hallucination through provenance: DVSpec's semantic references bind animations to data attribute values (for example,
{"sale_date": "2025-01-30"}) rather than rendering identifiers such as DOM IDs, so references remain valid when data ordering, chart type, or dataset changes, and every visual element is traceable to an underlying data record. - Declarative synchronization removes manual timestamps: Animation triggers are expressed as narration indices (
trigger ∈ [0, k-1] ∪ {null}) and resolved at render time from actual text-to-speech audio durations, so editing narration text automatically re-aligns timing. - Ablation confirms both pipeline stages matter: Removing the Story Planner drops the average score from 3.89 to 3.44; removing Orchestration drops it to 3.54.
- Human efficiency gain: Compared with a conversational LLM workflow, DataMagic significantly improves creation efficiency — a 79.7% reduction in task time — and reduces perceived cognitive load (user study, N=12).
- Evaluator reliability: Automated scoring by Gemini-2.5-Pro on a stratified sample of 60 videos correlates strongly with averaged expert scores from 3 domain experts (Pearson r=0.91, p<0.001; MAE=0.38). The paper states that dimension-level results follow, but the provided content is truncated at that point.
- Evaluation data composition: 60 datasets in total — 36 from T2R-bench (public statistics from 19 domains) and 24 from DAComp-DA (real enterprise data such as sales and HR) — yielding 109 samples. Data scale: small (<100 rows, 6 datasets), medium (100–1000 rows, 21 datasets), large (>1000 rows, 33 datasets, up to 150,000 rows).
Methodology in Plain English
The authors begin from the observation that a good data video is a structured narrative, not a pile of visual pieces: it is a sequence of scenes, each built around one analytical insight, ordered by narrative logic such as macro-to-micro or phenomenon-before-cause.
They therefore model generation as a hierarchical orchestration problem with three decision levels — macro narrative structure, scene-level content, and micro animation detail — and address two obstacles: representing charts, narration, and animation uniformly, and searching a huge design space for coherent compositions.
DVSpec is their answer to the representation problem. A video is a metadata block plus an ordered scene sequence, where each scene is a self-contained 5-tuple (id, type, content, narration, animation). Scenes can be typed (chart, opening, stat_cards, closing). Narration is a list of segments that, after text-to-speech, get precise time ranges. Animation effects are 4-tuples (type, target_data, trigger, style). Two design choices carry the weight: target_data uses data key-value pairs instead of rendering IDs, and trigger uses a narration index instead of an absolute timestamp.
The pipeline has two phases. In candidate scene generation, a Story Planner decomposes the user query into orthogonal analytical subtasks; a Data Manager writes Python code to extract, filter, and aggregate a micro-dataset per subtask; and a Visual Designer picks chart types (bar, line, area, scatter, pie, heatmap in the current implementation) and extracts key statistics. Narration and animation are deliberately left blank here. In global narrative orchestration, a Narration Director selects a scene subset by insight value and query coverage, orders it using a narrative pattern (Freytag's Pyramid by default; otherwise inverted-pyramid, comparison-driven, time-driven, or drill-down), and generates narration using a sliding-window context so adjacent scenes connect. An Animation Coordinator then analyzes entities mentioned in each narration segment and binds them to visual elements with the right trigger index. Orchestration targets five quality dimensions (Intent, Insight, Narrative, Animation, Aesthetic) under default constraints of 60–120 seconds total duration and k ≤ 7 initial scenes.
The system is a web application built with React and Remotion, with four panels: Data & History, Preview & Script, Scene Timeline & Narration Editor, and AI Edit. It follows a "scoped refinement" principle — an edit updates explicit DVSpec fields in one scene and re-renders only that scene, rather than rerunning the whole pipeline. Three interaction modes are supported: canvas manipulation (clicking and natural-language edits like @Scene 4: change to treemap and highlight top product), data-driven Q&A that answers questions by querying the underlying table through DVSpec's data bindings rather than inferring from pixels, and scene generation that turns a Q&A insight into a new scene inserted into the sequence.
Why This Matters
Impact on research. The paper reframes end-to-end data video generation as a hierarchical content-orchestration problem rather than a pixel-synthesis problem, and shows that a declarative intermediate representation can be the shared state linking automated generation with human editing. It also provides a two-tier benchmark setup — execution rate plus five quality dimensions — across 109 samples, and evidence that multi-stage decomposition beats single-call LLM generation even when the underlying model is the same.
Real-world applications.
- Business reporting: turning sales, HR, or operational tables into narrated multi-scene videos for stakeholders.
- Journalism and public communication: converting public statistics (the T2R-bench domains, such as environmental resources and public management) into explainer videos.
- Education: producing data-driven instructional clips where narration and chart highlights stay synchronized.
- Exploratory data analysis: the data-driven Q&A and scene-generation modes turn the finished video into an interactive, queryable data interface rather than a static deliverable.
Industry relevance. The system is model-agnostic, works on raw tabular data rather than pre-made charts, and reports an 79.7% reduction in task time versus a conversational LLM workflow — properties that matter for teams producing data videos at scale without multidisciplinary staffing.
Future Directions
- Broadening chart and rendering coverage. The authors note that the available chart range is jointly shaped by agent prompts, underlying model capabilities, and implemented rendering components, and can be widened by extending those components and generation rules.
- Scaling human evaluation. The reported user study has N=12; a larger and more varied participant pool would test how the three interaction modes hold up across different authoring styles.
- Testing the declarative layer beyond tabular single-table settings. The evaluation is restricted to single-table datasets refined from DAComp-DA and T2R-bench, so multi-table and relational settings remain untested in the reported content.
- Exploring interaction-mode tradeoffs. The paper describes three complementary modes over a shared DVSpec state but does not report performance or preference comparisons among canvas manipulation, script editing, and natural language commands.
Target Audience
This paper is most valuable to HCI and visualization researchers working on data storytelling, chart animation, and declarative visualization grammars; to practitioners building AI-assisted content-authoring tools; and to engineers designing multi-agent pipelines where outputs must remain verifiable and traceable to source data. Readers should be comfortable with concepts such as intermediate representations, agent orchestration, and text-to-speech-driven synchronization.
Authors’ abstract
Data videos communicate data insights through dynamic charts, voice narration, and synchronized animations, and have become a widely adopted form of data storytelling. However, producing them requires expertise in data analysis, narrative design, and video editing. Static visualization tools lack narrative and animation capabilities; authoring tools rely on pre-prepared charts rather than raw data; and pixel-level models generate videos end-to-end but cannot guarantee data accuracy or provenance. End-to-end automatic generation faces two core challenges: how to uniformly represent charts, narration, and animations together with their temporal relationships, and how to efficiently search a vast design space for narrative-coherent compositions. We present DataMagic, which authors data videos from raw tabular data through declarative multi-agent orchestration. First, the declarative specification DVSpec unifies charts, narration, and animations with data-bound references and declarative synchronization, ensuring data provenance and automatic audio-visual alignment. Second, a "Generate-then-Orchestrate" multi-agent strategy generates candidate scenes in parallel and then optimizes narrative coherence through global orchestration. DVSpec provides a shared state for three complementary interaction modes, bridging full automation with fine-grained human control. Evaluations on 109 real-world samples show that even the most advanced LLM (e.g., GPT-5) achieves only 2.13/5 with execution success rates between 48.62% and 86.24%; DataMagic improves quality to 3.89 (+83%) with success rates above 95%, with the most significant gains in animation and narrative dimensions. A user study shows that, compared to a conversational LLM workflow, DataMagic improves creation efficiency (79.7% reduction in task time) and reduces perceived cognitive load. Project page: https://github.com/HKUSTDial/DataMagic.