NVIDIA’s Physis-Lang Adds Physics-Aware Language to Video Models
marktechpost.com

NVIDIA’s Physis-Lang Adds Physics-Aware Language to Video Models

Tech News
5 min read

Published by AINave Editorial

TL;DRPhysis-Lang uses captions that describe causes and physical interactions to guide video-model data selection, training and inference. Reported Cosmos3-Nano results beat Veo 3.1 on three of four benchmark comparisons, but not on VideoPhy-2’s full set.

Video models can produce convincing images of events that do not make physical sense. Physis-Lang, a framework from researchers at NVIDIA, MIT and the University of Oxford, tries to address that gap with language that describes not just what happens, but why and how a scene changes. The reported Cosmos3-Nano comparisons beat Google’s Veo 3.1 on three of four benchmarks, with an important exception on VideoPhy-2’s full set. The researchers describe Physis-Lang as a shared language representation for data curation, model training and inference.

Captions describe causes, not just visible events

A conventional caption might say that butter melts as it warms. Physis-Lang adds a physics_reasoning field to describe entities, causes, interactions, governing principles, changes over time and effects. It also creates a scene-specific negative prompt that names likely physical mistakes, such as a stone floating on water, for use during inference. The intent is to give the model cues about how a scene should unfold, not just a list of objects and actions. That language is used across the framework’s data, training and inference stages.

The captioning model stays frozen in the framework’s refinement loop. A critic scores captions for precision and recall, and an evolution agent revises the captioning instructions based on those scores and claim-level failures. On PhysCapBench, which the article describes as 246 videos with 3,794 human-verified assertions, caption F1 rose from 78.64 at iteration one to 87.82 at iteration nine. It dipped to 76.28 at iteration two, so refinement did not improve steadily. The reported benchmark and iteration results show both the gain and that early setback.

The benchmark lead has a clear exception

The comparison with Veo 3.1 is strong on several reported tests, but the benchmark-by-benchmark results matter more than a general claim of superiority. The article reports Physis-Lang on Cosmos3-Nano ahead on PhyGenBench, Physics-IQ Verified and PhyGround, while Veo 3.1 scores higher on VideoPhy-2’s full set. Cosmos3-Nano leads again on VideoPhy-2’s Hard split. The reported scores are shown below.

Benchmark Cosmos3-Nano with Physis-Lang Veo 3.1
PhyGenBench 71.04 65.63
Physics-IQ Verified 43.41 34.99
PhyGround 69.90 69.24
VideoPhy-2, full set 68.02 68.87
VideoPhy-2, Hard split 62.36 58.43

The comparison is specific to these evaluations: the article says PhyGenBench and VideoPhy-2 use a GPT-5.5 evaluator. It also reports a separate prompting-only result: physics reasoning and negative prompts raised a frozen Cosmos3-Nano from 61.67 to 67.29 on PhyGenBench. That result is distinct from the fine-tuning comparison, which used LoRA on attention projections without changing the architecture or objective. The reported evaluation conditions and separate prompting result help define what the scores do, and do not, compare.

Language also guides which videos go into training

The framework’s data curation targets physical content linked to identified failure categories, rather than selecting clips only for visual similarity. The reported training set contains 183,000 videos: 71,000 filtered from WISA-80K and 112,000 retrieved clips. Retrieval alone added an average of 3.01 points across three benchmarks, according to the article. That makes the approach more than a prompt-writing technique: it uses physics descriptions to influence what the model learns from as well as what it is asked to generate. The training-set composition and retrieval result are reported in the article.

The practical distinction is that physics-aware language can intervene at several stages, while the benchmark results remain task- and protocol-specific. As of September 30, 2026, the article said only the paper was public, with code and weights not yet available; that status is a dated report, not confirmation of current availability. The article’s release-status note was checked on that date.

FAQs

Physis-Lang is a framework that uses language to describe physical causes, interactions, principles and changes over time in video scenes. The approach applies that representation to data curation, model training and inference. The researchers present it as a shared physical-language representation.

Sources

Latest Tech News