Simulation-Grounded · Trajectory-Level Evaluation
Measuring whether the motion in generated video obeys physics — scored against an exact, simulated ground-truth trajectory, with no human and no judge in the loop.
A clip can be pixel-perfect yet physically wrong — here the generated object floats under miscalibrated gravity. PhysEval scores the gap Δ between the recovered trajectory τ̂ and the simulator's exact reference τ*.
The problem
Text-to-video models now synthesize multi-second clips with convincing texture, lighting, and camera movement — and the field increasingly treats them as emerging world models for planning, robotics, and simulation. That ambition hinges on a property the eye does not verify directly: the motion must obey the laws of physics. A clip can be pixel-perfect while a ball drifts upward, a collision conjures momentum from nothing, or a pendulum swings at the wrong period.
Our ability to improve physical realism is bounded by our ability to measure it — and today's measurements are indirect. Human studies are slow, costly, and subjective; a single verdict fuses three separate questions (did the right objects appear, was the render clean, was the motion correct). MLLM critics promised scale but track objects poorly and reason weakly about contact. Both share a deeper flaw: they compare a video to a textual description or a human impression — never to a numerical statement of what the physics should have been.
The core idea
Physical realism is, operationally, a claim about trajectories — the path, speed, and timing of moving objects. If the correct trajectory is known, a clip's departure from it can be measured directly, objectively, and on a continuous scale. A physics engine supplies that trajectory for free.
Every prompted situation provably admits a physical solution — so a failing model cannot hide behind an ill-posed task. Even an impossible prompt gets a quantitative target, via a simulated counterfactual.
The simulator exports the precise trajectory of every moving body, together with the governing law and the parameters that produced it. Evaluation becomes a well-posed registration-and-comparison problem — not a proxy, an invariant check, or another learned opinion.
The metric — TPA
Scoring a generated clip is a problem of trajectory recovery and comparison. The pipeline localizes the moving object, follows it through pans and zooms, registers the recovered path to the simulated reference, and reads out a single objective number.
Does the shape of the path agree?
Implied gravity g, restitution e — the latent quantities, read directly from the motion.
Is the velocity / acceleration profile consistent?
Apex, impact, and contact at the right moments?
The benchmark
Each case is a parametric scene — a projectile launched at a chosen speed, a block released on a measured incline, two carts meeting at a fixed restitution — in which every initial condition is set explicitly, and each ships a certified-realizable prompt, the governing law, the exact reference trajectory, and a simulated counterfactual track.
From projectiles, free-fall, and inclines, to collisions and momentum transfer, to continuum media — a taxonomy grouped by how objects interact.
The same simulated scene, phrased from a bare action description up to a fully cinematic narrative — isolating how prompt style affects physical fidelity.
Code, scenarios, and simulators will be released. Scale figures reflect the benchmark's design.
Where PhysEval sits
Facing the same gap, the community has split into two paths. The perceptual / human-reasoning path (exemplified by Physion-Eval) scales expert annotation over real reference videos. The simulation-grounded path (PhysEval) builds ground truth in a physics engine and measures deviation objectively. They are complementary.
| Dimension | Perceptual / human-reasoning e.g. Physion-Eval |
Simulation-grounded PhysEval |
|---|---|---|
| Source of truth | Real reference videos + expert annotation | Exact trajectory τ* exported by a physics engine |
| Who judges | Humans, or trained VLM critics | A deterministic algorithm — no human, no judge |
| Output | Timestamped glitches + categories + explanations (qualitative) | A continuous TPA score + 4-way decomposition (quantitative) |
| What it measures | Whether a person perceives a glitch | Whether the motion obeys the law — and by how much |
| Latent quantities (g, e, momentum) | Inferred only indirectly | Fit directly from the trajectory |
| Reproducibility | Subject to annotator subjectivity / noise | Fully deterministic — same clip, same number |
| Coverage | Real-world, multi-physics, ego + exo (broad) | Physics with a single object trajectory (precise; currently narrower) |
Physion-Eval scales the human eye; PhysEval takes the eye out of the loop. One asks whether a person would notice; the other asks whether the physics is actually right — and by how much. Notably, it is these strongest human- and judge-based benchmarks that, with their own data, confirm the very gap an objective, law-level metric is built to close.
North star
PhysEval's paper is a starting point. The goal is to make physical correctness a standard, measurable foundation for video generation and world models.
Bring the objective, law-level metric to real reference video — uniting the perceptual paradigm's breadth with the simulation paradigm's precision.
TPA is continuous, decomposable, and differentiable — a physics-alignment reward. Pair categorical failure signals (where) with TPA (how much, which way).
Fluids, articulated and soft bodies, contact-rich manipulation — toward real-world dynamical complexity.
A standard diagnostic suite — an ImageNet/GLUE-style reference point for "does this model understand physics?"
Tracking physical-realism progress across model generations over time.
Reportable, comparable, and trackable — alongside visual quality and prompt adherence.
The researcher
AI Research Scientist — Evaluation · Generative & Multimodal Models
Ray's work centers on a single question: how do we measure what generative and multimodal models actually get right — and where they fail? He earned his Ph.D. in Computer Engineering from Concordia University (2025) on explainability and adversarial robustness for video transformers, with 12 peer-reviewed publications including first-author work at ICSE, ACM Transactions (TOMM), IEEE Transactions on Cloud Computing, and IEEE Access, and 40+ peer reviews for IEEE Transactions and AAAI.
His research builds evaluation infrastructure, perceptual metrics, and adversarial failure harnesses for transformer-based models across video, image, text, and audio. He designed STAA — the first single-pass spatio-temporal attention-attribution method for video transformers, defining faithfulness and monotonicity at under 150 ms latency (a 97% reduction over prior art) so evaluation can run as CI inside the training loop. He built the first joint spatio-temporal adversarial attack on video transformers (84.6% success across 20,000 videos) and showed the same machinery, reversed, cuts model vulnerability by over 50%. And he shipped XAIport / XAIpipeline, a cloud-agnostic microservice framework that turns evaluation from a manual notebook into an automated MLOps stage — validated at 500+ evaluation pipelines.
PhysEval sits at an unusual four-way intersection — rigorous ML evaluation, video-transformer expertise, production 3D vision, and classical physics simulation. Ray's path runs through all four.
PhysEval is a new perceptual metric. Ray's Ph.D. is exactly this — designing perceptual metrics (faithfulness, monotonicity), reproducible eval-as-CI, and benchmark harnesses across heterogeneous model backends.
PhysEval evaluates video generation. Ray interpreted and red-teamed video transformers on Kinetics-400 and Something-Something V2, and built the first joint spatio-temporal attacks and defenses on them.
TPA recovers motion via promptable segmentation + camera-compensated point tracking. Ray ships SAM3D segmentation and Gaussian Splatting / NeRF reconstruction from phone video in production at Maket.AI.
PhysEval's reference trajectories come from a physics engine. Ray built multi-physics and fluid-transport simulations and implemented finite-volume, finite-element, Runge–Kutta, and spectral solvers for Navier–Stokes and advection–diffusion (Université de Montréal; M.Sc. Process Systems Engineering, TU Dortmund). He can author the scenes, run the engines, and fit the governing laws.
Production multimodal generation pipelines; SAM3D segmentation and Gaussian Splatting / NeRF reconstruction from phone video; real-time multimodal AI chat driving iterative design loops.
XAI, perceptual metrics, adversarial evaluation, and MLOps for video transformers. STAA (IEEE Access 2025); adversarial attack/defense (ACM TOMM 2025); XAIport (ICSE 2024); multi-cloud benchmarking (IEEE TCC 2024).
LLM research pipelines; reused his XAI evaluation framework to score startup defensibility along data-moat, representation, retrieval, evaluation-rigor, and monitoring axes.
LLM fine-tuning for candidate–job matching; NLP automation suite (semantic search, summarization).
Numerical simulation and dynamical-systems modeling — fluid transport, multi-physics coupling; PDE solvers (Navier–Stokes, advection–diffusion) in Python / MATLAB / COMSOL, validated against analytical benchmarks.
Full list of 12 publications on Google Scholar.
Explainability & adversarial robustness for video transformers.
Modeling, PDEs, control theory, computational simulation.
Get in touch
I'm joining industry as a Research Scientist to build exactly this — objective, trajectory-level evaluation of physical realism, and the training signals that follow from it. If you're working on video generation, world models, or evaluation, I'd love to talk.