Coding Assistants A 7.89 GB Two-Bit Quant of Qwen3.8-27B Beats the 54 GB Original at Tool Calling AI Agents AWS Open-Sources a Physical AI Toolchain That Wires SageMaker to Nvidia Isaac Sim, GR00T and Cosmos News The Best Video Model Scores Just 57.76 Out of 100 on a New Physics Benchmark Built From 40 Measured Experiments Business Biren Technology Returns to the Market a Third Time in Ten Months With a $515 Million Share Sale News Singapore's MAS Turns AI Guidance Into Binding Rules for Banks, Phased In Through 2028 Business Hone Raises $60 Million Seed at a Reported $285 Million to Sell Agents That Own an Outcome Coding Assistants AI Coding Agents Add 30% More Code but Not More Software, Harvard Study Finds Security One Prompt, an Entire Region of Agents: Inside the Patched AgentCorruption Flaw in AWS Bedrock AgentCore Coding Assistants A 7.89 GB Two-Bit Quant of Qwen3.8-27B Beats the 54 GB Original at Tool Calling AI Agents AWS Open-Sources a Physical AI Toolchain That Wires SageMaker to Nvidia Isaac Sim, GR00T and Cosmos News The Best Video Model Scores Just 57.76 Out of 100 on a New Physics Benchmark Built From 40 Measured Experiments Business Biren Technology Returns to the Market a Third Time in Ten Months With a $515 Million Share Sale News Singapore's MAS Turns AI Guidance Into Binding Rules for Banks, Phased In Through 2028 Business Hone Raises $60 Million Seed at a Reported $285 Million to Sell Agents That Own an Outcome Coding Assistants AI Coding Agents Add 30% More Code but Not More Software, Harvard Study Finds Security One Prompt, an Entire Region of Agents: Inside the Patched AgentCorruption Flaw in AWS Bedrock AgentCore

The Best Video Model Scores Just 57.76 Out of 100 on a New Physics Benchmark Built From 40 Measured Experiments

The Best Video Model Scores Just 57.76 Out of 100 on a New Physics Benchmark Built From 40 Measured Experiments

A 40-task benchmark called World Models Last Exam in Physics, built by Einsia.AI with Peking University, Tsinghua and Navers Lab, measured 1,280 generated videos from eight models against observable physical relationships. The top scorer, Seedance 2.5, reached 57.76 out of 100, and free-fall clips that looked continuous averaged just 26.61.

A ball falls, a beam of light strikes a mirror, an ice cube melts in a glass. Every major video generator renders these scenes convincingly, and most of them get the physics wrong. A new benchmark called World Models' Last Exam in Physics, described in arXiv paper 2610.08791 by a team from Einsia.AI, Navers Lab, Peking University and Tsinghua University, tries to measure the gap, and the gap is large: the best of eight video generators scored 57.76 out of 100.

The design is deliberately narrow. The authors built 40 controlled tasks spanning nine categories of physical phenomena, including mechanics, optics, fluids, phase change, heat, electromagnetism and surface tension. Each task pairs an initial image and a generation prompt with predefined physical criteria, so the model is not asked to make something plausible but to make something specific happen in a way that obeys a measurable relationship. Eight models were run across the suite, producing 1,280 videos in total.

The methodological point is that the scoring is measurement-based rather than purely aesthetic. Where a vision-language model is used to judge consistency, the paper reports that a Qwen3.6-27B model applies task-specific prompts that check whether the required subject is present, whether the specified event occurs, and whether objects and processes stay continuous. A ball can follow a perfectly smooth trajectory and still violate free-fall; two pendulums can swing steadily with the wrong period-to-length relationship.

The results separate looking right from being right. Seedance 2.5 led the field at 57.76 out of 100, and even that leaves substantial room. Free-fall videos that passed basic continuity screens averaged just 26.61 out of 100, meaning clips that looked fine often failed the underlying acceleration relationship. The authors also flag that observability failures, where the camera cannot see the relevant quantity, can suppress physical scores, so passing the benchmark does not prove full physical consistency.

Transparency is thinner than the numbers. At the time the paper surfaced, no public code, weights or dataset links were attached, and the authors state their own limitations. The scores should be read as a diagnostic of observable physics in generated video, not a universal measure of world-model intelligence.

The practical stakes are concrete. If a robot or an agent uses a generated video to rehearse pushing an object, it needs the model to represent force and friction correctly, not just render a convincing clip. Benchmarks that reward visual realism without testing whether anything obeys nature have been rewarding the wrong thing, and this paper is a corrective aimed squarely at that habit.

For an industry spending billions on world models as the training substrate for physical AI, a 57.76 ceiling is a useful piece of calibration. The visuals have been outrunning the physics for some time. This is one of the first benchmarks to put a number on how far.

Comments (0)

Log in to join the discussion

Log In

No comments yet