Research
Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model
Overview Research area: Computer vision and generative video modeling — specifically physical plausibility in video world models, physics-aware video captioning, and agentic self-evolution of language

- arXiv
- 2609.40358
- Published
- 2026-09-30
- Authors
- Liming Lu, Xianzheng Ma, Wenkun He, Guanqi Zhan, Yilin Zhao, Junyu Chen, Mengyao Xu, Jiaojiao Fan, Wenhang Ge, Yuchao Gu, Yunze Liu, Boyi Li, Zhen Dong, Victor Prisacariu, Ming-Yu Liu, Song Han, Han Cai
AI summary
Overview
Research area: Computer vision and generative video modeling — specifically physical plausibility in video world models, physics-aware video captioning, and agentic self-evolution of language representations.
Technical level: Advanced. The paper assumes familiarity with video diffusion backbones, vision-language models (VLMs), LoRA fine-tuning, classifier-free guidance, and assertion-level caption evaluation.
Scope: The paper proposes Physis-Lang, a framework that treats structured, self-evolving natural language as the shared representation for curating physical training data, supervising video-model fine-tuning, and guiding inference in video world models.
What This Paper Is About
Video world models can produce videos that look convincing but violate basic physics — objects deform or vanish, collisions resolve implausibly, fluids move unnaturally, and causal events occur out of order. Most existing fixes assume language is too weak to carry physical knowledge, so they bolt on extra visual, latent, numerical, or planning-based signals. This paper revisits that assumption: it asks whether language itself, made explicit about entities, causes, interactions, governing principles, temporal evolution, and effects, and then iteratively optimized, can serve as a single shared representation of physics across the whole video-generation pipeline.
Key Contributions
- Language as a physical world representation. The authors reframe language not as a conditioning interface but as a representation that can carry, refine, and apply physical knowledge in video world models.
- A self-evolving agent system with a physics-aware critic and PhysCapBench. An agentic loop evaluates and revises the physics-captioning instruction without updating the captioner's weights. PhysCapBench, a held-out benchmark of 246 physics-rich videos and 3,794 human-verified physical assertions, measures physical coverage (recall) and claim faithfulness (precision) and provides the stopping criterion for evolution.
- A language-guided data engine. A GPT-5.5-based diagnosis agent summarizes the model's failures into a category-level deficiency profile; physics-domain tags on a large video gallery are then matched at the text level to retrieve training videos that cover the missing physical processes, which are re-captioned with the evolved guidelines.
- PhysThinker, a low-cost open-source replacement. Two 4B vision-language models — PhysThinker-C for physics-aware video captioning and PhysThinker-U for inference-time prompt upsampling — distill the physics reasoning of the proprietary captioner and upsampler.
Main Findings
- Consistent gains over the base model. Fine-tuning Cosmos3-Nano with Physis-Lang improves all four evaluated physical benchmarks: Physics-IQ Verified 40.23 to 43.41, PhyGround 65.18 to 69.90, PhyGenBench 61.67 to 71.04, and VideoPhy-2 60.41 to 68.02.
- Best among open-source, competitive with or better than the leading closed-source model. Physis-Lang-enhanced Cosmos3-Nano surpasses Veo 3.1 on Physics-IQ Verified (43.41 vs. 34.99), PhyGround (69.90 vs. 69.24), and PhyGenBench (71.04 vs. 65.63). On VideoPhy-2 it trails Veo 3.1 by 0.85 points on the full set (68.02 vs. 68.87) while outperforming it by 3.93 points on the hard split (62.36 vs. 58.43).
- Generalization across families and scales. Mean improvement is +7.05 points for Wan2.1-14B and +6.22 for Cosmos3-Nano-16B. Within the Cosmos 3 family, gains are +3.24 on Edge-4B, +6.22 on Nano-16B, and +5.02 on Super-64B.
- Evolution improves captions and downstream generation. On PhysCapBench, F1 rises from 78.64 at Iteration 1 to 87.82 at Iteration 9. Holding the pretrained Cosmos3-Nano fixed and changing only the inference captions from Iterations 1, 4, 8, and 9, PhyGenBench scores rise from 64.17 to 65.63, 65.83, and 67.29 respectively.
- Retrieved data helps. Adding language-guided retrieved videos improves benchmarks by an average of 3.01 points (PhyGenBench +2.92, Physics-IQ Verified +2.73, VideoPhy-2 +3.39). Category-level gains on VideoPhy-2 are positive across all frequently represented categories, including chemical processes (+8.00), fracture mechanics (+7.45), and cloth deformation (+7.19).
- Prompting, fine-tuning, and negative guidance each add value. Physical reasoning prompting gives a zero-shot gain of 61.67 to 63.33 without fine-tuning, and beats a manually designed physics prompt (61.86 to 63.33). Supervised fine-tuning adds 3.55 points (63.33 to 66.88) and 3.75 points (67.29 to 71.04) depending on the data. Physics negative prompts add 2.50 points (65.62 to 68.12) and 4.16 points (66.88 to 71.04).
- PhysThinker trades accuracy for cost. On Wan2.1-14B, the commercial captioner and upsampler give +7.05 mean at roughly $24.12K; commercial upsampler with PhysThinker-C gives +6.76 at roughly $0.12K; fully local PhysThinker-C plus PhysThinker-U gives +4.76 at zero cost.
- General video quality is preserved. The authors state the fine-tuning strategy does not compromise general video generation quality, with details deferred to their appendix.
Methodology in Plain English
The authors start by enriching video captions. A base caption describing visible content is augmented with a physics_reasoning field that spells out entities, dynamics, and causal relations, plus a scene-specific physics_negative_prompt describing implausible evolutions to be suppressed at inference.
To make those captions good, they build a physics-aware critic. Precision decomposes a generated caption into atomic, independently verifiable claims and checks each against the video, classifying it correct, incorrect, or uncertain; micro-averaged precision is the fraction of decided claims that are correct. Recall takes the human-curated atomic physical assertions for a video and asks whether each is stated or unambiguously entailed by the caption. F1 combines the two.
An evolution agent then runs a generate–critic–revise loop over a fixed 20-video development set (273 human-verified atomic assertions). At each iteration a fixed captioner produces captions, the critic produces scores and a diagnostic report, and the evolution agent rewrites the captioning instruction — never the captioner's weights. After each iteration the prompt is scored on PhysCapBench; the loop stops when gains saturate and the best prompt is used to re-caption the training corpus.
Separately, a GPT-5.5-based diagnosis agent watches generated videos, labels physical failures by category (for example rigid-body motion, collision, and fluid dynamics), and aggregates them into a deficiency profile. That profile is matched against physics tags on a large video gallery so that retrieved videos target physically missing domains rather than merely similar-looking content.
For training, the final captions supervise fine-tuning of existing backbones without changing their architectures or objectives, using LoRA on attention projections. For inference, an upsampler expands short prompts (and optionally an image) into the structured physics fields, and the negative prompt is used as negative conditioning. To avoid the cost of proprietary models at scale, the physics reasoning of the GPT-based captioner and upsampler is distilled into two 4B VLMs.
Why This Matters
Impact on research. The paper challenges a widely held design assumption — that physics must be injected through non-linguistic channels such as reward models, intermediate features, retrieved samples, or simulators. It argues that the bottleneck is how language is used, not language itself, and offers a reusable recipe: evolve the instruction with a physics-aware critic, benchmark it with assertion-level precision and recall, and reuse the resulting language across data curation, training, and inference. PhysCapBench also gives the community a standalone evaluation target for physics-aware captioning.
Real-world applications.
- Robotics and embodied intelligence, where a model must anticipate the physical consequences of actions rather than merely render plausible frames.
- Autonomous systems and driving simulation, where incorrect predictions of contact, deformation, or causal ordering are costly.
- Interactive simulation and content creation, where physically implausible generations break immersion or mislead.
- Cost-sensitive large-scale annotation pipelines, where the distilled PhysThinker models replace commercial APIs for captioning and prompt upsampling.
Industry relevance. The models are trained and evaluated at production scale — 183K videos, backbones from 4B to 64B parameters, and comparisons against proprietary Veo 3.1 — with an explicit accuracy-versus-API-cost accounting. That makes the approach directly relevant to teams building video generation infrastructure who need both quality and controllable inference costs.
Future Directions
- Determining how far the caption-evolution loop can be pushed: the reported trajectory saturates after nine iterations, and the authors do not report whether restarting evolution with a different development set or critic yields further gains.
- Extending the shared physical-language representation beyond T2V and I2V to other conditioned generation settings; the current training mix is 70% text-to-video, 20% image-to-video, and 10% video-to-video on Cosmos3-Nano, with no video-to-video conditioning at all for Cosmos3-Edge.
- Closing the remaining gap on VideoPhy-2, where the method still trails Veo 3.1 by 0.85 points on the full set despite a 3.93-point lead on the hard split.
- Recovering the accuracy lost by going fully local: PhysThinker-C plus PhysThinker-U costs nothing but drops the mean gain from +7.05 to +4.76 on Wan2.1-14B, leaving room to narrow that gap.
Target Audience
Researchers and engineers working on video generation, world models, and embodied AI who care about physical plausibility; practitioners building production video-generation pipelines who need to weigh captioning quality against annotation cost; and anyone interested in agentic self-evolution methods, physics-aware captioning benchmarks, or distilling reasoning from large proprietary models into small open ones. Readers without background in video diffusion models or VLM-based evaluation will find the method sections demanding.
Authors’ abstract
Video world models are expected to predict how the physical world evolves, yet they often produce visually plausible videos that violate basic physical principles. Existing approaches commonly assume that natural language is insufficient to represent the physical knowledge required for reliable generation, and therefore introduce additional visual, latent, numerical, or planning-based signals. We revisit this assumption and introduce Physis-Lang, a self-evolving framework that treats physical language as a shared and optimizable representation across data curation, model training, and video generation. Physis-Lang represents physical processes through language that describes their relevant entities, causes, interactions, governing principles, temporal evolution, and effects. To improve this representation, we construct PhysCapBench, which decomposes physical processes into atomic assertions and evaluates captions using recall and precision. An agentic loop iteratively analyzes assertion-level errors and refines the instruction used to produce physical captions. Physis-Lang further converts model deficiencies into textual descriptions and uses language-guided retrieval to identify visually diverse videos that cover missing physical processes. Experiments on four widely used physical video benchmarks with Wan and Cosmos backbones demonstrate consistent improvements in physical plausibility. Notably, starting from open-source Cosmos3-Nano backbones, our Physis-Lang-enhanced models surpass the leading proprietary Veo 3.1 model.