NVIDIA's Physis-Lang trains video models on self-evolving physics captions, beats Veo 3.1 on 3 of 4 benchmarks
Physis-Lang (NVIDIA, MIT, Oxford) adds a physics reasoning field and a scene-specific negative prompt to video captions. An agent refines the captioning instruction in a loop while the captioner itself stays frozen. PhyGenBench: Cosmos3-Nano + Physis-Lang scores 71.04, vs 65.63 for Veo 3.1 and 61.6…
Read the full story at r/machinelearningnews ↗
Timeline · 1 report
- 2026-09-30 07:37 · r/machinelearningnews
NVIDIA's Physis-Lang trains video models on self-evolving physics captions, beats Veo 3.1 on 3 of 4 benchmarks