Skip to content
AI.info

Research

TwiFF (Think With Future Frames): A Large-Scale Dataset for Dynamic Visual Reasoning

TwiFF (Think With Future Frames): A Large-Scale Dataset for Dynamic Visual Reasoning Overview Research area: Computer Vision / Multimodal Large Language Models — specifically Visual Chain-of-Thought (

TwiFF (Think With Future Frames): A Large-Scale Dataset for Dynamic Visual Reasoning
arXiv
2602.10675
Published
2026-02-11
Authors
Junhua Liu, Zhangcheng Wang, Zhike Han, Ningli Wang, Guotao Liang, Kun Kuang

AI summary

TwiFF (Think With Future Frames): A Large-Scale Dataset for Dynamic Visual Reasoning

Overview

Research area: Computer Vision / Multimodal Large Language Models — specifically Visual Chain-of-Thought (VCoT) reasoning for video and dynamic scenarios.

Technical level: Advanced. The paper assumes familiarity with multimodal LLMs, chain-of-thought prompting, video generation, and benchmark evaluation protocols.

Scope: The paper introduces TwiFF-2.7M (a large-scale dataset of temporally grounded visual chain-of-thought data), TwiFF-Bench (an evaluation benchmark), and TwiFF (a unified model trained on the dataset), all aimed at moving visual reasoning beyond static images into dynamic, future-state prediction.

What This Paper Is About

Existing Visual Chain-of-Thought methods let multimodal models use images as intermediate reasoning steps, but they only reason about visual content already present in the input image, so they work for static tasks such as maze navigation, geometric reasoning, visual search, and jigsaw assembly, and fail on dynamic tasks that require anticipating future states. The paper addresses this gap by building a large-scale dataset from real video clips whose reasoning steps depict actual future frames, plus a benchmark and a trained model that jointly generate future action frames and textual reasoning.

Key Contributions

  1. TwiFF-2.7M — described as the first large-scale, dynamic VCoT dataset, comprising step-by-step visual-textual rationales derived from over 2.7 million video clips (2,708,318 VCoT instances), constructed specifically so models can anticipate future states beyond static observations.
  2. TwiFF-Bench — a benchmark of 1,078 samples that evaluates both the plausibility of intermediate reasoning steps and the correctness of final predictions in open-ended, temporally evolving scenarios, addressing the limitation of benchmarks that score only end-task accuracy.
  3. The TwiFF model — a unified model trained on TwiFF-2.7M that iteratively generates future action frames and textual reasoning, reported to outperform Textual Chain-of-Thought (TCoT) and static VCoT approaches on dynamic visual question answering.
  4. Three empirical findings — that VCoT effectiveness comes from the synergy of textual and visual modalities (neither alone suffices), that physically plausible visual cues are critical for answer accuracy while misleading cues degrade it, and that visual cues in dynamic scenarios have inherent information-compression potential that suppresses action-irrelevant noise.

Main Findings

  • Large gains over the base model: Relative to the base Bagel-7B model, TwiFF improves TwiFF-Bench CoT scores by 28.8% and answer scores by 41.6%, and improves answer scores by 21.0% on the out-of-distribution Seed-Bench-R1 benchmark.
  • Per-task improvements on TwiFF-Bench (vs. Bagel): +27.7% CoT and +41.3% answer on Instructional, +32.2% CoT and +44.8% answer on Predictive, +27.2% CoT and +41.4% answer on Camera.
  • Per-split improvements on Seed-Bench-R1 (vs. Bagel): +20.6% on L1, +12.8% on L2, and +27.9% on L3.
  • Second-best overall, behind Qwen3VL-8B only: TwiFF reaches a TwiFF-Bench CoT/answer of 2.95/2.62 and a Seed-Bench-R1 average of 1.62, while Qwen3VL-8B reaches 2.84/2.44 and 1.87. The paper attributes the residual gap to architectural improvements in Qwen3VL, since Bagel is built on Qwen2.5, and notes Qwen3VL-8B is the only competing method that surpasses TwiFF.
  • Dynamic VCoT beats TCoT and static VCoT: TwiFF surpasses tool-augmented VCoT (DeepEyes) and generative VCoT baselines (Zebra-CoT, ThinkMorph), as well as TCoT-style Qwen2.5VL-7B, InternVL3.5-8B, Bagel-7B, and Janus-Pro-7B.
  • Fewer training steps still competitive: TwiFF-Lite, trained for 6,000 steps on a 300K-sample subset (versus 36,000 steps on the full dataset for TwiFF), reaches 2.90/2.55 on TwiFF-Bench and 1.62 on Seed-Bench-R1.
  • Multimodal synergy is necessary: On TwiFF-Bench, TwiFF-Lite outperforms TwiFF-Text by 3.4% and TwiFF-Image by 14.7%. On Seed-Bench-R1 the gaps widen to 11.0% and 18.2%.
  • Single-modality CoT generalizes poorly out of distribution: Relative to Bagel, TwiFF-Text and TwiFF-Image gain only 9.0% and 2.2% on Seed-Bench-R1, whereas TwiFF-Lite gains 20.9%.
  • Visual cue fidelity matters: Replacing the model's first generated image with the ground-truth future frame (TwiFF-True) raises CoT score to 3.56 (+20.7%) and answer score to 3.14 (+19.8%). Replacing it with a duplicate of the input question image (TwiFF-False) lowers CoT to 2.92 (−1.0%) and answer to 2.57 (−1.9%).
  • Information compression potential: Discarding only the input question image after the first generated frame (TwiFF-Comp) costs 2.7% CoT and 4.6% answer score, whereas discarding both the input image and the first generated image (TwiFF-Drop) costs 24.4% CoT and 14.1% answer score.
  • Reasoning quality correlates with answer quality: Higher-quality VCoTs consistently correlate with higher answer scores on TwiFF-Bench.
  • Reported data quality: Of 10,000 randomly sampled TwiFF-2.7M instances assessed by Qwen3VL for answerability and logical coherence, only 7.3% were classified as unacceptable.
  • Composition: TwiFF-2.7M is 70.6% Instructional, 17.5% Predictive, and 11.9% Camera; TwiFF-Bench is 71.2% Instructional, 15.6% Predictive, and 13.2% Camera.
  • Reasoning length and temporal span: About 64% of TwiFF-2.7M VCoT instances use one image, 29% use two, and 7% use between three and seven. Most samples cover action durations under 10 seconds, with the longest extending beyond 40 seconds.

Methodology in Plain English

The dataset is derived from Panda-70M, a YouTube-sourced video dataset, through a three-stage pipeline.

Stage 1 (coarse filter) removes low-quality or low-motion clips using four criteria: clips with an Unmasked Teacher caption-matching score below 0.43 are discarded; only clips labeled "desirable" by Panda-70M are kept (excluding static foreground images, screen-in-screen compositions, and screen recordings); clips shorter than 2 seconds are removed; and clips whose maximum inter-frame optical flow magnitude across 8 uniformly sampled frames is below 4 are dropped as insufficiently dynamic. At most four clips per original video are retained for source diversity. This yields 10,596,462 high-dynamics clips.

Stage 2 (event extraction) uses a multimodal LLM (InternVL3.5-8B) to classify each clip as Instructional, Predictive, Camera, or Undesirable (discarded), then select at least two key frames capturing the cause, process, and outcome, plus textual descriptions. Eight frames per clip are uniformly sampled for this. This yields 3,075,048 high-quality event instances.

Stage 3 (VCoT generation) orders key frames temporally, uses the earliest frame as the query image, and generates a question conditioned on that frame along with the event category and description. The answer then alternates reasoning text and subsequent frames in an interleaved format (reasoning about frame i−1, then frame i, then reasoning about frame i, and so on, ending in a final answer). This yields the 2,708,318 instances of TwiFF-2.7M.

Benchmark construction: Following the official Panda-70M split, TwiFF-Bench contains 1,078 question-answer pairs from its test subset using the same pipeline, with zero overlap with training data, and all generated samples were manually filtered to remove flawed reasoning, incorrect answers, or overly open-ended responses.

Model and training: The researchers finetune Bagel-7B with a learning rate of 2×10⁻⁵ and cosine decay for 36,000 steps on the full dataset. Ablation variants TwiFF-Lite, TwiFF-Text (textual VCoT components only), and TwiFF-Image (visual components only, with literal "frame_i" text retained so the model can decide when to stop generating) are each trained for 6,000 steps on the same 300K-sample subset.

Evaluation: GPT-5.1-2025-11-13 is used as a judge, scoring CoT reasonableness and answer correctness separately on a 0–5 scale against a reference VCoT and ground-truth answer drawn from actual future events. The judge is instructed not to penalize a CoT solely for omitting explicit image references. For out-of-distribution evaluation, Seed-Bench-R1 (4,676 samples from EPIC-Kitchens-100 and Ego4D) is reformulated as open-ended questions with only the current observation and question, and only the answer score is judged because no reference reasoning traces exist. To prevent infinite loops, DeepEyes is capped at 5 tool invocations, and ThinkMorph, Zebra-CoT, and TwiFF at 8 image generations.

Why This Matters

Impact on research: The paper reframes VCoT from reasoning about what is visible in a static input to reasoning about what happens next, and argues this requires grounding in real video rather than tool-augmented synthetic pipelines. It also contributes an evaluation paradigm that scores intermediate reasoning plausibility, not just final answers, and reports evidence that physically faithful visual cues causally improve answer accuracy while misleading cues degrade it.

Real-world applications:

  • Step-by-step instructional guidance, such as cooking, mechanical assembly, or repair, where a system must anticipate what action comes next.
  • Predictive forecasting of physical outcomes, such as whether an object will fall or a structure will collapse.
  • Camera work prescription, such as how a shot should be framed or moved to convey a specific meaning or emotion.
  • Embodied and navigation settings, where the paper cites prior findings that generating video-based predictions of physical-world actions improves model accuracy in robotic action decision-making and navigation.

Industry relevance: For companies building multimodal assistants, video-editing or cinematography tools, robotics and autonomous navigation systems, and tutorial or coaching platforms, the paper offers both training data and a benchmark for models that must reason about future visual states. The information-compression finding (TwiFF-Comp losing only 2.7% CoT and 4.6% answer score) is directly relevant to efficiency and context-management design in multi-step visual reasoning.

Future Directions

  • Reinforcement learning on CoT plausibility. The paper explicitly proposes exploring novel reinforcement learning approaches for multimodal reasoning through the lens of CoT plausibility, given its finding that visual cue authenticity drives answer accuracy.
  • Closing the gap to stronger base architectures. TwiFF trails Qwen3VL-8B, attributed to architectural improvements in Qwen3VL relative to the Qwen2.5-based Bagel, raising the question of how the dataset would perform on a stronger unified backbone.
  • Broadening dynamic scenario coverage. The paper notes Zebra-CoT's embodied data covers relatively narrow scenarios, and that TwiFF-2.7M spans three domains (instructional, predictive, camera); extending dynamic VCoT to other domains and longer time horizons (samples beyond 40 seconds exist but are rare) remains open.
  • Reducing reliance on judge models and reference traces. Evaluation depends on GPT-5.1 as a judge with reference VCoTs and ground-truth answers, while Seed-Bench-R1 lacks reference reasoning traces entirely, leaving the question of how to score reasoning quality without references.

Target Audience

Researchers and engineers working on multimodal large language models, visual chain-of-thought reasoning, video understanding, and dataset or benchmark construction. It will also interest practitioners in robotics, embodied AI, video generation, and instructional-media applications who need models to reason about future visual states rather than only static images. Readers should already be comfortable with multimodal model training and evaluation terminology.

Authors’ abstract

Visual Chain-of-Thought (VCoT) has emerged as a promising paradigm for enhancing multimodal reasoning by integrating visual perception into intermediate reasoning steps. However, existing VCoT approaches are largely confined to static scenarios and struggle to capture the temporal dynamics essential for tasks such as instruction, prediction, and camera motion. To bridge this gap, we propose TwiFF-2.7M, the first large-scale, temporally grounded VCoT dataset derived from $2.7$ million video clips, explicitly designed for dynamic visual question and answer. Accompanying this, we introduce TwiFF-Bench, a high-quality evaluation benchmark of $1,078$ samples that assesses both the plausibility of reasoning trajectories and the correctness of final answers in open-ended dynamic settings. Building on these foundations, we propose the TwiFF model, a unified modal that synergistically leverages pre-trained video generation and image comprehension capabilities to produce temporally coherent visual reasoning cues-iteratively generating future action frames and textual reasoning. Extensive experiments demonstrate that TwiFF significantly outperforms existing VCoT methods and Textual Chain-of-Thought baselines on dynamic reasoning tasks, which fully validates the effectiveness for visual question answering in dynamic scenarios. Our code and data is available at https://github.com/LiuJunhua02/TwiFF.

Read the original paper