Video world models are expected to predict how the physical world evolves, yet they often produce visually plausible videos that violate basic physical principles. Existing approaches commonly assume that natural language is insufficient to represent the physical knowledge required for reliable generation, and therefore introduce additional visual, latent, numerical, or planning-based signals. We revisit this assumption and introduce Physis-Lang, a self-evolving framework that treats physical language as a shared and optimizable representation across data curation, model training, and video generation. Physis-Lang represents physical processes through language that describes their relevant entities, causes, interactions, governing principles, temporal evolution, and effects. To improve this representation, we construct PhysCapBench, which decomposes physical processes into atomic assertions and evaluates captions using recall and precision. An agentic loop iteratively analyzes assertion-level errors and refines the instruction used to produce physical captions. Physis-Lang further converts model deficiencies into textual descriptions and uses language-guided retrieval to identify visually diverse videos that cover missing physical processes. Experiments on four widely used physical video benchmarks with Wan and Cosmos backbones demonstrate consistent improvements in physical plausibility. Notably, starting from open-source Cosmos3-Nano backbones, our Physis-Lang-enhanced models surpass the leading proprietary Veo 3.1 model.
Figures & tables
Figure 1 : Overview of Physis-Lang . Existing approaches treat language mainly as a conditioning interface and introduce physical knowledge through separate visual, latent, numerical, or planning-based signals. Physis-Lang instead treats physical language as a shared and optimizable representation: an agentic loop evaluates and refines physical descriptions, language-guided retrieval expands the training data toward missing physical processes, and the evolved language supervises training and guides inference toward physically plausible video generation.
Figure 2 : Overview of our method . Physical processes are represented through natural language, including a base caption, explicit physics reasoning, and scene-adaptive negative physics descriptions. A PhysCapBench-driven agentic loop evaluates and refines this physical language, which is then used to expand the training data, supervise video-model training, and guide inference toward physically plausible generation.
Figure 3 : Overview of our physics-aware critic and physics caption benchmark . The critic evaluates generated captions from complementary precision and recall perspectives by verifying atomic claims against the input video and measuring the coverage of human-curated physical assertions.
Table 4
Generator
PhyGenBench
Physics-IQ Verified
VideoPhy-2
PhyGround
Mean Δ
Across model families
Wan2.1-14B
56.67 → 65.83
27.87 → 35.15
57.02 → 65.65
61.52 → 64.64
+7.05
Cosmos3-Nano-16B
61.67 → 71.04
40.23 → 43.41
60.41 → 68.02
65.18 → 69.90
+6.22
Across Cosmos scales
Cosmos3-Edge-4B
48.96 → 52.50
32.80 → 34.69
32.99 → 40.61
66.66 → 66.56
+3.24
Cosmos3-Nano-16B
61.67 → 71.04
40.23 → 43.41
60.41 → 68.02
65.18 → 69.90
+6.22
Table 5 : Effectiveness on different backbones. Each cell shows the improvement brought by our method. Mean Δ is the average improvement across the four benchmarks.
Figure 4 : From caption evolution to video generation. Left: PhysCapBench F1 at different iterations. Right: physics-aware video generation performance of the same pretrained Cosmos3 model with inference captions produced at different iterations.
Figure 7
ID
SFT data
Prompt
Score
ID
SFT data
Prompt
Score
A
None
Base
61.67
E
WISA
Base + P
65.62
B
None
M
61.86
F
WISA
Base + P + N
68.12
C
None
Base + P
63.33
G
WISA + retrieved
Base + P
66.88
D
None
Base + P + N
67.29
H
WISA + retrieved
Base + P + N
71.04
Table 7 : Prompting and SFT on PhyGenBench. Base: base caption; M: manually designed physics caption; P: physics_reasoning ; N: physics_negative_prompt .
Captioner
Upsampler
PhyGenBench
Physics-IQ Verified
VideoPhy-2
PhyGround
Mean Δ
Cost (USD)
Pretrained Wan2.1-14B
56.67
27.87
57.02 / 43.82
61.52
–
–
Commercial
Commercial
65.83
35.15
65.65 / 59.55
64.64
+7.05
∼ 24.12K
PhysThinker-C
Commercial
65.00
34.71
65.99 / 54.49
64.40
+6.76
∼ 0.12K
PhysThinker-C
PhysThinker-U
63.33
34.13
60.07 / 55.06
64.60
+4.76
0
Table 8: PhysThinker replacement on Wan2.1-14B. VideoPhy-2 reports All / Hard; Cost denotes API expenditure for training-data captioning and benchmark upsampling (USD).
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : F1-score trajectory across prompt self-evolution iterations on PhysCapBench . The dashed purple line denotes the F1 score of base caption. The five callouts summarize the representative prompt modifications introduced at Iterations 2, 3, 4, 8 and 9. The star identifies the best-performing iteration. The annotations describe the prompt-design changes, whereas all plotted F1 scores are obtained from PhysCapBench.
Figure 7 : Sliding-window mechanism for text encoding in Wan2.1.
Figure 8 : Qualitative comparisons on different physics benchmarks. Each case compares frames generated by the Cosmos3-Nano base model using base caption (Baseline) with those generated by our trained model using base caption augmented with our physics_reasoning and physics_negative_prompt (Ours) at identical timestamps. Official prompts and selected physics reasoning and negative-prompt excerpts are shown, with bold phrases highlighting physical constraints. Negative excerpts describe violations to avoid. All frames are uncropped.
Figure 9 : Qualitative results on driving and robotics. All sequences are generated by our trained model using base caption augmented with our physics_reasoning and physics_negative_prompt . Each example includes four uncropped video frames with timestamps, a caption excerpt, and selected physics reasoning and negative-prompt excerpts.
Figure 10 : Cup drop, impact, and fracture. Six chronological frames accompany the reference physical assertions in their original wording and rank order. Phase labels summarize the displayed frames; assertion IDs are not temporal alignments.
Figure 11 : Newton’s cradle and momentum transfer. Six chronological frames accompany the reference physical assertions in their original wording and rank order. Phase labels summarize the displayed frames; assertion IDs are not temporal alignments.
Figure 12 : Visualization for the attention scores of visual and text tokens. When adding physics_reasoning into the SFT and inference process, the generative model can better capture the keywords for the core physical process, and generate videos with physical realism.
Caption
Δ (%)
…Regions directly under fingertips COMPRESS first and most strongly. …
19.2
… HEAT TRANSFER lowers the average molecular kinetic energy of the juice, …
52.2
… GRAVITY still acts downward on all gas and particle mass, …
22.5
…The change in light speed at the lens surfaces BENDS the rays. …
48.7
Appendix
Table 9 : Relative increase of attention scores of physical keywords after fine-tuning.
Figure 13 : Different evaluators for videos generated by CogVideoX1.5-5B and Cosmos3-Nano. While the auto-evaluator of VideoPhy-2 tends to assign similar mid-range scores to videos with substantially different levels of physical realism, GPT-5.5 provides more discriminative assessments, better reflecting the differences in physical realism across videos.
Generative Model
Pretrained VLM
GPT-5.5
CogVideoX1.5-5B
23.01
49.41
Wan2.1-T2V-14B
23.52
57.02
Cosmos3-Nano
22.50
60.41
Veo-3.1
23.69
68.87
Appendix
Table 10 : Comparison of VideoPhy-2 scores under different evaluators.
Model
VBench-I2V
Wan2.1-14B (Pretrained)
86.86
Wan2.1-14B (Physis-Lang)
87.51
Cosmos3-Nano (Pretrained)
88.32
Cosmos3-Nano (Physis-Lang)
88.69
Appendix
Table 11 : Comparison of different models on VBench.