Research
CoVA-SFT: A Large-Scale Dataset for Chain of Visual Abstractions
Overview Research area: Natural language processing and multimodal reasoning — specifically chain-of-thought (CoT) prompting, supervised fine-tuning data, and benchmarks for models that reason over vi

- arXiv
- 2608.28958
- Published
- 2026-08-29
- Authors
- Tsung-Han Wu, Heekyung Lee, Anya Ji, Haoming Chen, Trevor Darrell, Joseph E. Gonzalez, David M. Chan
AI summary
Overview
Research area: Natural language processing and multimodal reasoning — specifically chain-of-thought (CoT) prompting, supervised fine-tuning data, and benchmarks for models that reason over visual representations.
Technical level: Advanced. The paper assumes familiarity with chain-of-thought reasoning, interleaved multimodal models, and supervised fine-tuning dataset construction.
Scope: The paper introduces CoVA-SFT, a large-scale supervised fine-tuning corpus for teaching multimodal language models to interleave text with visual abstractions when solving textual reasoning problems, plus a companion held-out benchmark, CoVA-Bench.
What This Paper Is About
Chain-of-thought reasoning helps large language models break problems into intermediate steps, but text-only CoT forces a model to describe visual and spatial structure in prose, which the authors describe as an awkward serialization of visual problems. Architectures can accept visual input, but the community has lacked a large, multi-step, self-corrected dataset that teaches models to build and maintain an internal visual workspace while reasoning. CoVA-SFT is presented as that dataset, together with a benchmark for evaluating it reproducibly.
Key Contributions
- CoVA-SFT dataset: A highly structured corpus of 51.9K samples containing over 222K multimodal reasoning steps, organized across 5 distinct layout families and 17 complex tasks.
- CoVA-Bench: A companion benchmark of 1,700 held-out test samples spanning the same 17 tasks, intended to make evaluation reproducible.
- Training signal design: The corpus supplies explicit rationale formulations, agentic renderings, and verification loops, which the authors say teach multimodal language models to interleave text and visual abstractions.
- Empirical validation: Evidence that models fine-tuned on CoVA-SFT outperform all interleaved CoT baselines by more than 2x on average on CoVA-Bench, while still trailing strong text-only CoT baselines.
Main Findings
- Large-scale, structured supervision: CoVA-SFT comprises 51.9K samples and over 222K multimodal reasoning steps, spread over 5 layout families and 17 tasks, indicating the dataset is designed to cover varied visual structuring rather than a single format.
- Interleaving is trained, not assumed: The dataset explicitly includes rationales, agentic renderings, and verification loops, framing visual abstraction as a learned behavior rather than a byproduct of architecture.
- Fine-tuning beats interleaved baselines: Models fine-tuned on CoVA-SFT outperform all interleaved CoT baselines by more than 2x on average on CoVA-Bench.
- Text-only CoT remains stronger: Despite the gains, the fine-tuned models still fall short of strong text-only CoT baselines, which the authors present as an open challenge rather than a solved problem.
- Reproducible evaluation is provided: CoVA-Bench offers 1,700 held-out test samples over the same task set, so results can be compared on a fixed split. The abstract does not report per-task scores, metric definitions, model sizes, training compute, or baseline identities.
Methodology in Plain English
The authors' approach is to build the missing training resource rather than propose a new architecture. They assemble a supervised fine-tuning corpus in which each example walks through a reasoning problem as a sequence of interleaved steps, some of them visual abstractions and some of them text. The steps are structured around defined layout families and task types, and are annotated with explicit rationales, rendering actions, and checks that verify intermediate results — essentially teaching the model to draw, inspect, and correct a workspace as it goes. They then hold out 1,700 samples as a benchmark and test whether models trained on the corpus do better than existing interleaved chain-of-thought approaches. The abstract does not describe how the data was collected or generated, which base models were fine-tuned, or how training was configured.
Why This Matters
Impact on research: The work shifts attention from architecture alone to training data as the bottleneck for multimodal reasoning. By releasing both a corpus and a held-out benchmark, it gives the field a shared target and a common split for measuring whether interleaved visual reasoning actually improves. The finding that interleaved models still lose to text-only CoT is a useful negative result that frames where the difficulty lies.
Real-world applications (extrapolated from the abstract's framing of textual problems that require visual workspaces):
- Diagram- and geometry-style questions posed in text, where a model must hold spatial structure rather than describe it in prose.
- Chart, table, and graph reasoning, where intermediate plots or layouts can be rendered and re-checked.
- Planning and navigation problems such as route or path reasoning, where an internal map is more natural than a verbal description.
- Multi-step troubleshooting or procedure following, where verification loops catch errors before the final answer.
Industry relevance: Companies building multimodal assistants, document understanding tools, and agentic systems have a direct interest in whether interleaved reasoning beats plain text CoT. The result reported here — a large gain over interleaved baselines but not over strong text-only CoT — suggests that shipping systems should not assume visual interleaving is automatically better, and that further data and method work is needed before it is.
Future Directions
- Close the gap with text-only CoT: The central open question is why fine-tuned interleaved models still trail strong text-only chain-of-thought baselines, and what data or training changes would fix it.
- Scale and diversify the corpus: Extending beyond 51.9K samples, 5 layout families, and 17 tasks, and testing whether layout diversity transfers to unseen families, is a natural next step.
- Isolate which ingredients matter: The abstract groups rationales, agentic renderings, and verification loops together; determining each component's separate effect would clarify the design.
- Test transfer beyond the benchmark: Whether gains on CoVA-Bench carry over to held-out task families, real multimodal inputs, or downstream agentic use is left unresolved, as the abstract reports evaluation only on CoVA-Bench.
Target Audience
Researchers and engineers working on multimodal language models, chain-of-thought reasoning, and instruction-tuning datasets. It is most useful to those who build or evaluate interleaved text-and-vision reasoning systems, and to practitioners deciding whether to invest in visual-abstraction training data for their own models. Readers looking for a new model architecture or for detailed training and evaluation methodology will need more than this abstract provides, since the full text was not available.
Authors’ abstract
Chain-of-thought (CoT) reasoning has dramatically improved large language models (LLMs) by allowing them to decompose problems into intermediate steps. While CoT is widely effective for linguistic tasks, text-only CoT forces models to serialize visual problems into awkward prose. Although architectural solutions exist to process visual inputs, the community lacks a massive, multi-step, self-corrected dataset to teach models how to build and maintain internal visual workspaces when solving purely textual reasoning problems. To address this limitation, we introduce CoVA-SFT, a highly structured corpus of 51.9K samples containing over 222K multimodal reasoning steps across 5 distinct layout families and 17 complex tasks, and CoVA-Bench, a companion benchmark of 1,700 held-out test samples spanning the same tasks for reproducible evaluation. By providing explicit rationale formulations, agentic renderings, and verification loops, CoVA-SFT teaches multimodal language models to interleave text and visual abstractions. We validate the dataset by demonstrating that models fine-tuned on CoVA-SFT outperform all interleaved CoT baselines by more than 2x on average on CoVA-Bench, though they still fall short of strong text-only CoT baselines, highlighting open challenges for future work.