Research
MaLiang-Harness: A Programmable Path to Image and Video Generation
MaLiang-Harness: A Programmable Path to Image and Video Generation Overview Research area: Computer vision, specifically multimodal large language model (MLLM) driven image and video generation throug

- arXiv
- 2609.34309
- Published
- 2026-09-28
- Authors
- Haoyu Zhao, Zihao Zhang, Xudong Wang, Jiaxi Gu, Zuxuan Wu, Yu-Gang Jiang, Shuicheng Yan
AI summary
MaLiang-Harness: A Programmable Path to Image and Video GenerationOverview
Research area: Computer vision, specifically multimodal large language model (MLLM) driven image and video generation through executable programs rather than direct pixel synthesis.
Technical level: Intermediate to Advanced. The paper assumes familiarity with MLLM agent design, rendering backends, and evaluation methodology, though the core idea (write code, render it, look at the result, fix the code) is accessible.
Scope in one sentence: The paper introduces a stateful framework, MaLiang-Harness, that organizes MLLM-driven image and video generation as a persistent loop of program construction, visual inspection, and revision, then benchmarks 11 MLLMs on two new task suites.
What This Paper Is About
Programs can run without errors yet still produce an image or video that violates what the user asked for — an object appears in the wrong place, or an animated event happens at the wrong moment. The authors name this discrepancy the Program-to-Visual (P2V) gap and argue that generating runnable code is only the beginning of visual creation. Their goal is a framework that keeps the evolving visual program, its full edit history, and its visual verification anchored to a single shared revision reference, so an MLLM can diagnose and repair visual failures rather than treating successful execution as success.
Key Contributions
-
Definition of the Program-to-Visual (P2V) gap as the discrepancy between program-level correctness and satisfaction of visual requirements, motivating a stateful formulation of visual program generation through construction, inspection, and revision.
-
MaLiang-Harness, a unified framework for programmable image and video generation built on three mechanisms: Persistent Executable Generation (PEG) state, which preserves programs and task context; Traceable Generation Process (TGP), which connects edits to rendered evidence; and Revision-aware Editing and Verification (REV), which supports restoration of historical content and checks the current revision before completion.
-
A unified programmatic visual generation interface across rendering backends (Canvas, SVG, Scene2d, and Three.js), in which each backend retains its native executable representation while sharing a common protocol for state management, rendering, and requirement review. Images and videos share the same creation framework: at a fixed program revision, an image is rendered at a specified content time, while a video samples the temporal behavior encoded by the program.
-
Two benchmarks and a multi-model evaluation: MaLiang-IBench (50 text-to-image prompts) and MaLiang-VBench (13 text-to-video prompts), used to evaluate 11 MLLMs on the image benchmark and four on the video benchmark across generation success, visual quality, and computational cost, alongside a comparison to public general-capability scores.
Main Findings
-
Generation success is not the same as visual requirement satisfaction. GPT-5.6-Luna and GPT-5.6-Terra each successfully generate 46 of 50 images, but only 22 and 24 respectively satisfy all three quality criteria (alignment, aesthetics, composition).
-
GPT-6-Astra leads on both benchmarks. It reports 100% generation success on both, with 96.0% of image tasks and 76.9% of video tasks meeting all quality thresholds, versus 86.0% and 38.5% for GPT-5.6-Sol. These percentages use the full task set of each benchmark as the denominator.
-
Quality failures concentrate in prompt adherence. All successfully generated images from the three GPT-6 models meet the aesthetics and composition thresholds; their remaining quality failures concern prompt adherence.
-
A large gap separates the GPT family from DeepSeek and Kimi on image tasks. GPT models report 92–100% generation success, versus 12–40% for DeepSeek and Kimi.
-
Token-budget exhaustion explains only part of the failures. It accounts for five of DeepSeek-V4.1-Flash's 13 video failures, one of Kimi-K2.6's 10, and three of GPT-5.6-Sol's six. The paper attributes it potentially to reasoning stagnation and unbounded agentic loops.
-
Motion is the most restrictive video criterion for Astra. All 13 of its videos meet the aesthetics threshold, but only 10 meet the motion-coherence threshold. Requiring composition and motion coherence reduces GPT-5.6-Sol's joint count from seven to five.
-
Cost and quality trade off. GPT-5.6-Sol produces 43 qualifying images at 3.28 minutes per image; Astra produces 48 at 3.70 minutes per image. On video, Astra achieves full task completion at 10.00 minutes per successful video, while GPT-5.6-Sol requires 9.27 minutes per successful video but completes fewer tasks.
-
General capability scores only partially predict visual program generation. Across 11 models, the AA Intelligence Index score and drawing quality pass rate show a positive rank correlation (Spearman ρ = 0.65). GPT-5.6-Luna and GPT-6-Luna share a displayed index score of 37, yet their drawing quality pass rates are 44% and 88% — a 44-percentage-point gap. Kimi-K3 scores 44 on the index but satisfies all visual criteria on only 18% of tasks.
-
Mean quality scores require care. DeepSeek and Kimi generate only 6–20 images, so their means describe a limited subset; Kimi-K3's alignment score of 4.56 is based on just nine images.
-
Richer backends add realism; finer brushwork alone does not. With a path-tracing backend, GPT-6-Astra constructs a breakfast still life without generated image assets, and raising the final sample count from 512 to 1024 reduces visible noise. A separate experiment that refines an existing Canvas illustration's brushwork preserves composition but remains stylized, with grass coverage becoming less visible and stone markings and mountain textures attenuated.
-
The construction record distinguishes drawn content from assembled content. In the six-panel "research figure" case, GPT-6-Astra specifies layout, architecture diagrams, labels, and plots in code from user-supplied synthetic data, while generating photographic-style assets through four recorded asset-generation attempts; the final composition uses crops from the third attempt.
-
Evaluation caveat stated by the authors: MLLM verdicts (with GPT-6-Sol as judge) remain self-assessments rather than independent measures of perceptual quality, and DeepSeek-V4-Pro is evaluated without visual feedback.
Methodology in Plain English
The harness takes a prompt and an output specification, and lets an MLLM plan the appearance, spatial composition, and temporal dynamics before selecting an executable representation and a compatible backend. The model writes drawing and animation code, optionally incorporating user-provided assets or assets obtained through image generation or search when rich appearance is hard to express in code alone. The backend renders the result into an image or video.
Three mechanisms organize the loop. The PEG state at revision k bundles the visual program with its backend identifier, associated assets (retained with content hashes), scene attributes describing spatial composition and temporal dynamics, the generation context (prompt, output specification, requirements, and current plan), and the revision index. Commits use a validated edit function, and the harness retains the preceding snapshot so earlier versions can be revisited rather than overwritten.
The TGP records each operation as a tuple of the executed operation, its inputs, its execution result including returned errors, and the source and resulting revisions. Operation indices are distinct from revision indices, since rendering and inspection do not commit state updates. Each rendered observation is linked to the revision the renderer evaluated, along with the actual output, sampled timestamps, and any crop specification.
The REV layer associates a review — the revision it assesses, requirement-specific visual evidence, and a verdict of pass, fail, or uncertain — with each visual requirement. A review applies to the current state only when its revision index matches the current one. After any commit, even a plan-only change, the current revision must be reviewed again. Restoration from a historical revision commits a new state while preserving current requirements. Temporal requirements use ordered samples spanning the specified interval at three or more distinct timestamps; the authors note such samples support temporal assessment but do not establish continuity between frames.
Delivery requires three conditions: export checks pass (source files and assets exist and match recorded hashes; format, dimensions, and video timing conform to the specification), every checkpoint passes for the current revision, and every mandatory requirement has evidence from the current revision with a passing verdict. Evaluation reports generation success, failures, token-limit failures, generation time, model calls, and token usage, plus quality scored by GPT-6-Sol on a five-point scale for prompt alignment, aesthetics, and composition, with motion coherence added for video using 12 temporally ordered frames from each successful video.
Why This Matters
The paper reframes visual generation evaluation: a model that runs code without crashing has not necessarily produced the requested image. By tying every visual assessment to a specific revision and requiring re-verification after every commit, it gives a concrete vocabulary (P2V gap, PEG, TGP, REV) for studying how MLLMs translate executable code into visual outcomes, and it shows that general benchmarks are imperfect predictors of that ability.
Real-world applications:
- Diagram and figure production. The "research figure" case shows code-driven layout, labels, and data plots combined with generated photographic assets, with provenance tracking that separates drawn elements from assembled imagery.
- Design iteration with inspectable history. Because edits, renderings, and reviews share a revision reference, a designer can compare revisions and locate which change caused a visual defect such as the biscuit surface artifacts introduced and then reduced during path-traced refinement.
- Animated and explanatory video. The MaLiang-VBench tasks cover sequences of actions, such as constructing a channel, filling it with water, and moving boats downstream.
- Asset provenance and auditability. Content hashing of programs and assets, plus export checks that files match recorded hashes, supports tracing how a delivered visual was constructed.
Industry relevance: The reported 92–100% generation success by GPT-family models versus 12–40% for DeepSeek and Kimi, combined with the 44-percentage-point quality gap between two models with identical general index scores, gives teams a reason to evaluate MLLMs on task-specific visual generation rather than relying on general capability rankings alone. The cost tables (time per qualified image, time per successful video, calls and tokens per case) provide a direct basis for budgeting agentic visual pipelines.
Future Directions
-
Closing the remaining quality gap within successful runs. Even GPT-6-Astra leaves 2 of 50 image tasks and 3 of 13 video tasks short of all quality thresholds, and its motion-coherence count (10 of 13) lags its aesthetics count (13 of 13), suggesting temporal requirement satisfaction is the harder target.
-
Controlled comparisons for richer rendering backends. The authors state that controlled comparisons are still needed to establish the contribution of more expressive renderers such as path tracing to perceptual realism.
-
Preserving detail during refinement-driven code edits. The hyperrealism-inspired experiment reduced grass coverage and attenuated stone markings and mountain textures, which the authors frame as a need to preserve meaningful visual detail when translating refinement instructions into code.
-
Independent perceptual evaluation. The paper notes its MLLM verdicts remain self-assessments, leaving open how human or non-MLLM judging would change the reported quality picture, and whether the temporal sampling approach can be extended beyond ordered samples to establish frame-to-frame continuity.
Target Audience
Researchers and engineers working on MLLM agents, code-generating visual systems, and controllable image or video generation; benchmark designers interested in separating execution success from requirement satisfaction; and practitioners who need inspectable, revision-aware pipelines for figure production, design iteration, or animation. Readers looking for diffusion or flow-matching architecture innovations will not find them here — the paper deliberately explores a non-diffusion, non-flow-matching programmable path in which renderers turn executable programs into images and videos.
Authors’ abstract
Executable programs offer explicit control over how images and videos are constructed, but generating runnable code is only the beginning of visual creation. A program can execute correctly while violating the requested composition, appearance, or motion. We define this discrepancy as the Program-to-Visual (P2V) gap and introduce MaLiang-Harness, a unified framework for organizing MLLM-driven visual generation into a persistent process of construction, inspection, and revision. Its central design is to make the evolving visual program, its construction history, and its verification share a common revision reference. We define the Persistent Executable Generation (PEG) state as preserving programs and task context. Traceable Generation Process (TGP) connects edits to rendered evidence, and Revision-aware Editing and Verification (REV) supports restoration and checks the current revision before completion. Together, these mechanisms coordinate planning, execution, and visual feedback across rendering backends. We evaluate 11 powerful closed-source MLLMs on MaLiang-IBench and four on MaLiang-VBench, measuring generation success, visual quality, and computational cost. GPT-6-Astra achieves 100% generation success on both benchmarks, with 96.0% of image tasks and 76.9% of video tasks meeting all quality thresholds. The comparison also reveals a mismatch between general capability scores and visual generation performance, with similarly scored models differing substantially in their ability to satisfy visual requirements. MaLiang-Harness provides a systematic basis for studying how MLLMs translate executable code into visual outcomes, exposing both the potential of programmable generation and the limitations of general benchmarks as predictors of this ability. The project is available at https://github.com/gulucaptain/MaLiang-Harness.