Simulation-Grounded · Trajectory-Level Evaluation

PhysEval.

Measuring whether the motion in generated video obeys physics — scored against an exact, simulated ground-truth trajectory, with no human and no judge in the loop.

METRIC TPA — Trajectory–Physics Alignment
SIM-GROUNDED TRAJECTORY-LEVEL NO JUDGE IN THE LOOP FULLY REPRODUCIBLE
Δ = TPA
τ* simulated reference τ̂ generated (recovered)

A clip can be pixel-perfect yet physically wrong — here the generated object floats under miscalibrated gravity. PhysEval scores the gap Δ between the recovered trajectory τ̂ and the simulator's exact reference τ*.

The problem

Visual realism is not physical realism.

Text-to-video models now synthesize multi-second clips with convincing texture, lighting, and camera movement — and the field increasingly treats them as emerging world models for planning, robotics, and simulation. That ambition hinges on a property the eye does not verify directly: the motion must obey the laws of physics. A clip can be pixel-perfect while a ball drifts upward, a collision conjures momentum from nothing, or a pendulum swings at the wrong period.

83.3% / 93.5%
of clips from leading models contain at least one human-visible physical glitch (exocentric / egocentric). The problem is far from solved.
— reported by Physion-Eval, CVPR 2026
74–90%
of those obvious glitches are missed by the strongest MLLM critic (Gemini 3.0 Pro). Building evaluation on VLM judges rests on shaky ground.
— reported by Physion-Eval, CVPR 2026

Our ability to improve physical realism is bounded by our ability to measure it — and today's measurements are indirect. Human studies are slow, costly, and subjective; a single verdict fuses three separate questions (did the right objects appear, was the render clean, was the motion correct). MLLM critics promised scale but track objects poorly and reason weakly about contact. Both share a deeper flaw: they compare a video to a textual description or a human impression — never to a numerical statement of what the physics should have been.

The core idea

Invert the protocol: author the scene in a simulator.

Physical realism is, operationally, a claim about trajectories — the path, speed, and timing of moving objects. If the correct trajectory is known, a clip's departure from it can be measured directly, objectively, and on a continuous scale. A physics engine supplies that trajectory for free.

Rather than describe a scene in words and ask whether the video looks right, we author the scene inside a physics engine, render the same specification into a prompt, and ask whether the generated motion follows the trajectory the simulator computed.
GUARANTEE 01

Certified realizability

Every prompted situation provably admits a physical solution — so a failing model cannot hide behind an ill-posed task. Even an impossible prompt gets a quantitative target, via a simulated counterfactual.

GUARANTEE 02

Exact reference trajectory τ*

The simulator exports the precise trajectory of every moving body, together with the governing law and the parameters that produced it. Evaluation becomes a well-posed registration-and-comparison problem — not a proxy, an invariant check, or another learned opinion.

The metric — TPA

Recover the trajectory, register it, score the gap.

Scoring a generated clip is a problem of trajectory recovery and comparison. The pipeline localizes the moving object, follows it through pans and zooms, registers the recovered path to the simulated reference, and reads out a single objective number.

01Promptable segmentationgated by motion saliency
→
02Camera-motion compensationremove model pans & zooms
→
03Dense point trackingrecover image-space τ̂
→
04Restricted similarity registrationalign τ̂ to τ*
→
TPA4-way comparisonone objective score
SUB-SCORE 1

Trajectory geometry

Does the shape of the path agree?

SUB-SCORE 2

Fitted law parameters

Implied gravity g, restitution e — the latent quantities, read directly from the motion.

SUB-SCORE 3

Kinematic profile

Is the velocity / acceleration profile consistent?

SUB-SCORE 4

Event timing

Apex, impact, and contact at the right moments?

Objective — no annotator, no judge Continuous — degree, not yes/no Diagnostic — where & how it breaks Reproducible — same clip, same score

The benchmark

Every test case is born in a simulator.

Each case is a parametric scene — a projectile launched at a chosen speed, a block released on a measured incline, two carts meeting at a fixed restitution — in which every initial condition is set explicitly, and each ships a certified-realizable prompt, the governing law, the exact reference trajectory, and a simulated counterfactual track.

≈10 physics families 100s parametric scenarios 1000s sim-anchored prompts 3 interaction-complexity tiers + counterfactual (anti-physics) track
ORGANIZED BY INTERACTION COMPLEXITY

Single-body · Multi-body · Continuum

From projectiles, free-fall, and inclines, to collisions and momentum transfer, to continuum media — a taxonomy grouped by how objects interact.

THREE PROMPT LEVELS

Action · Consequence · Cinematic

The same simulated scene, phrased from a bare action description up to a fully cinematic narrative — isolating how prompt style affects physical fidelity.

Code, scenarios, and simulators will be released. Scale figures reflect the benchmark's design.

Where PhysEval sits

Two paradigms for one gap.

Facing the same gap, the community has split into two paths. The perceptual / human-reasoning path (exemplified by Physion-Eval) scales expert annotation over real reference videos. The simulation-grounded path (PhysEval) builds ground truth in a physics engine and measures deviation objectively. They are complementary.

Dimension Perceptual / human-reasoning
e.g. Physion-Eval
Simulation-grounded
PhysEval
Source of truthReal reference videos + expert annotationExact trajectory τ* exported by a physics engine
Who judgesHumans, or trained VLM criticsA deterministic algorithm — no human, no judge
OutputTimestamped glitches + categories + explanations (qualitative)A continuous TPA score + 4-way decomposition (quantitative)
What it measuresWhether a person perceives a glitchWhether the motion obeys the law — and by how much
Latent quantities (g, e, momentum)Inferred only indirectlyFit directly from the trajectory
ReproducibilitySubject to annotator subjectivity / noiseFully deterministic — same clip, same number
CoverageReal-world, multi-physics, ego + exo (broad)Physics with a single object trajectory (precise; currently narrower)

Physion-Eval scales the human eye; PhysEval takes the eye out of the loop. One asks whether a person would notice; the other asks whether the physics is actually right — and by how much. Notably, it is these strongest human- and judge-based benchmarks that, with their own data, confirm the very gap an objective, law-level metric is built to close.

North star

Beyond the paper.

PhysEval's paper is a starting point. The goal is to make physical correctness a standard, measurable foundation for video generation and world models.

1

Sim-to-real

Bring the objective, law-level metric to real reference video — uniting the perceptual paradigm's breadth with the simulation paradigm's precision.

2

From evaluation to training signal

TPA is continuous, decomposable, and differentiable — a physics-alignment reward. Pair categorical failure signals (where) with TPA (how much, which way).

3

Physics-domain expansion

Fluids, articulated and soft bodies, contact-rich manipulation — toward real-world dynamical complexity.

4

A physics check-up for world models

A standard diagnostic suite — an ImageNet/GLUE-style reference point for "does this model understand physics?"

5

A living leaderboard

Tracking physical-realism progress across model generations over time.

6

Physical correctness as a first-class metric

Reportable, comparable, and trackable — alongside visual quality and prompt adherence.

The researcher

Ray (Zerui) Wang, Ph.D.

AI Research Scientist — Evaluation · Generative & Multimodal Models

Ray's work centers on a single question: how do we measure what generative and multimodal models actually get right — and where they fail? He earned his Ph.D. in Computer Engineering from Concordia University (2025) on explainability and adversarial robustness for video transformers, with 12 peer-reviewed publications including first-author work at ICSE, ACM Transactions (TOMM), IEEE Transactions on Cloud Computing, and IEEE Access, and 40+ peer reviews for IEEE Transactions and AAAI.

His research builds evaluation infrastructure, perceptual metrics, and adversarial failure harnesses for transformer-based models across video, image, text, and audio. He designed STAA — the first single-pass spatio-temporal attention-attribution method for video transformers, defining faithfulness and monotonicity at under 150 ms latency (a 97% reduction over prior art) so evaluation can run as CI inside the training loop. He built the first joint spatio-temporal adversarial attack on video transformers (84.6% success across 20,000 videos) and showed the same machinery, reversed, cuts model vulnerability by over 50%. And he shipped XAIport / XAIpipeline, a cloud-agnostic microservice framework that turns evaluation from a manual notebook into an automated MLOps stage — validated at 500+ evaluation pipelines.

Ph.D. 2025
Computer Engineering · Concordia
12
peer-reviewed publications
ICSE · TOMM · IEEE Trans
first-author venues
40+
peer reviews (IEEE Trans, AAAI)

Why PhysEval needs exactly this background.

PhysEval sits at an unusual four-way intersection — rigorous ML evaluation, video-transformer expertise, production 3D vision, and classical physics simulation. Ray's path runs through all four.

Need: an objective, reproducible metric

Evaluation & perceptual metrics

PhysEval is a new perceptual metric. Ray's Ph.D. is exactly this — designing perceptual metrics (faithfulness, monotonicity), reproducible eval-as-CI, and benchmark harnesses across heterogeneous model backends.

Need: understanding of video models

Video-transformer expertise

PhysEval evaluates video generation. Ray interpreted and red-teamed video transformers on Kinetics-400 and Something-Something V2, and built the first joint spatio-temporal attacks and defenses on them.

Need: recover trajectories from clips

3D vision — segmentation, tracking, reconstruction

TPA recovers motion via promptable segmentation + camera-compensated point tracking. Ray ships SAM3D segmentation and Gaussian Splatting / NeRF reconstruction from phone video in production at Maket.AI.

Need: ground truth born in a simulator

Physics simulation & numerical methods

PhysEval's reference trajectories come from a physics engine. Ray built multi-physics and fluid-transport simulations and implemented finite-volume, finite-element, Runge–Kutta, and spectral solvers for Navier–Stokes and advection–diffusion (Université de Montréal; M.Sc. Process Systems Engineering, TU Dortmund). He can author the scenes, run the engines, and fit the governing laws.

Experience

Selected publications

Full list of 12 publications on Google Scholar.

Education

Ph.D., Computer Engineering
Concordia University
2020 — 2025

Explainability & adversarial robustness for video transformers.

M.Sc., Process Systems Engineering
Technical University Dortmund
2014 — 2018

Modeling, PDEs, control theory, computational simulation.

B.Sc., Chemical Engineering
China University of Mining and Technology
2010 — 2014

Get in touch

Let's make physical correctness measurable.

I'm joining industry as a Research Scientist to build exactly this — objective, trajectory-level evaluation of physical realism, and the training signals that follow from it. If you're working on video generation, world models, or evaluation, I'd love to talk.

Website
xaiport.com
Open to Research Scientist roles — evaluation, physical realism, and world models.