
NVIDIA’s Physis-Lang Adds Physics-Aware Language to Video Models
Published by AINave Editorial
Video models can produce convincing images of events that do not make physical sense. Physis-Lang, a framework from researchers at NVIDIA, MIT and the University of Oxford, tries to address that gap with language that describes not just what happens, but why and how a scene changes. The reported Cosmos3-Nano comparisons beat Google’s Veo 3.1 on three of four benchmarks, with an important exception on VideoPhy-2’s full set. The researchers describe Physis-Lang as a shared language representation for data curation, model training and inference.
Captions describe causes, not just visible events
A conventional caption might say that butter melts as it warms. Physis-Lang adds a physics_reasoning field to describe entities, causes, interactions, governing principles, changes over time and effects. It also creates a scene-specific negative prompt that names likely physical mistakes, such as a stone floating on water, for use during inference. The intent is to give the model cues about how a scene should unfold, not just a list of objects and actions. That language is used across the framework’s data, training and inference stages.
The captioning model stays frozen in the framework’s refinement loop. A critic scores captions for precision and recall, and an evolution agent revises the captioning instructions based on those scores and claim-level failures. On PhysCapBench, which the article describes as 246 videos with 3,794 human-verified assertions, caption F1 rose from 78.64 at iteration one to 87.82 at iteration nine. It dipped to 76.28 at iteration two, so refinement did not improve steadily. The reported benchmark and iteration results show both the gain and that early setback.
The benchmark lead has a clear exception
The comparison with Veo 3.1 is strong on several reported tests, but the benchmark-by-benchmark results matter more than a general claim of superiority. The article reports Physis-Lang on Cosmos3-Nano ahead on PhyGenBench, Physics-IQ Verified and PhyGround, while Veo 3.1 scores higher on VideoPhy-2’s full set. Cosmos3-Nano leads again on VideoPhy-2’s Hard split. The reported scores are shown below.
| Benchmark | Cosmos3-Nano with Physis-Lang | Veo 3.1 |
|---|---|---|
| PhyGenBench | 71.04 | 65.63 |
| Physics-IQ Verified | 43.41 | 34.99 |
| PhyGround | 69.90 | 69.24 |
| VideoPhy-2, full set | 68.02 | 68.87 |
| VideoPhy-2, Hard split | 62.36 | 58.43 |
The comparison is specific to these evaluations: the article says PhyGenBench and VideoPhy-2 use a GPT-5.5 evaluator. It also reports a separate prompting-only result: physics reasoning and negative prompts raised a frozen Cosmos3-Nano from 61.67 to 67.29 on PhyGenBench. That result is distinct from the fine-tuning comparison, which used LoRA on attention projections without changing the architecture or objective. The reported evaluation conditions and separate prompting result help define what the scores do, and do not, compare.
Language also guides which videos go into training
The framework’s data curation targets physical content linked to identified failure categories, rather than selecting clips only for visual similarity. The reported training set contains 183,000 videos: 71,000 filtered from WISA-80K and 112,000 retrieved clips. Retrieval alone added an average of 3.01 points across three benchmarks, according to the article. That makes the approach more than a prompt-writing technique: it uses physics descriptions to influence what the model learns from as well as what it is asked to generate. The training-set composition and retrieval result are reported in the article.
The practical distinction is that physics-aware language can intervene at several stages, while the benchmark results remain task- and protocol-specific. As of September 30, 2026, the article said only the paper was public, with code and weights not yet available; that status is a dated report, not confirmation of current availability. The article’s release-status note was checked on that date.



















