Testing more than visual polish
An arXiv preprint has introduced a test for robot video models that exposes a problem hidden by polished footage: a rollout can appear plausible while missing the scene's actual geometry, object positions or motion, and while failing when the video is converted back into robot actions. The work asks whether a generated sequence preserves the underlying 3D scene state and can support executable behavior, addressing what the authors describe as a gap in existing evaluation.
Cosmos, the best of four evaluated models on the reported composite measures, reached a task-focused composite score called RoboPhyscore of 0.6330, equal to 92.7% of the study's Ground Truth, and an Average Full Score, or AFS, of 2.7797. Yet its object-localization score retained only 61.2% of Ground Truth and its trajectory-accuracy score 73.2%. These are benchmark comparisons, not causal evidence, and the study does not establish performance beyond its simulated setup.
A test built around the scene
RoboPhys-3D organizes 50 complementary metrics into 18 sub-dimensions and four evaluation levels. AFS is the broad summary, while RoboPhyscore is the more compact, task-relevant score. It retains eight non-task-success measures only when each measure's lowest Pearson correlation, a measure of how two sets of values move together, with Action Planner, Data Engine and VLM-2 is greater than 0.70.
For each generated rollout, the protocol compares the video with reference material processed through the same reconstruction pipeline. It also decodes the generated video into executable robot actions and replays those actions in the simulator. The shared setup is intended to show whether an apparent error came from reconstructing the 3D scene or from generating the rollout itself, while linking visual and geometric scores to simulated task behavior.
The experiments covered four open-source video world models: Wan 2.2, CogVideoX 1.5, Cosmos 3 and RoboDreamer. The models were post-trained on the benchmark, and episodes were divided into 70% for training, 15% for validation and 15% for testing. The reported ranking therefore describes the study's split-based evaluation rather than a purely zero-shot comparison.
The benchmark covers four task regimes and 50 tasks, with 100 episodes per task and five camera views per episode. In total, it includes 5,000 ground-truth episodes, 25,000 ground-truth videos, 20,000 reconstructed 3D scenes and 120,000 generated videos. The evaluation scored 145,000 videos.
Cosmos leads, but the detail tells a different story
The aggregate result becomes harder to interpret when the components are separated. Cosmos had a task-level score of 2.4139, compared with 2.3500 for Ground Truth, while its execution-grounded task success retained 85.9% of Ground Truth. The different results show that a strong aggregate or task-related score does not guarantee that every object and movement was represented accurately.
CogVideoX provides another sharp contrast. Its collision-safety score was 0.9650, yet its task-success retention was only 16.5% of Ground Truth. A high safety score was not accompanied by high task-success retention in this comparison, supporting the paper's view that perceptual plausibility, geometric consistency, state accuracy and executable success are complementary checks.
The result also changed with the reconstruction method used to turn video into a 3D scene. VGGT retained 94.7% of the ground-truth RoboPhyscore and VGGT-Ω retained 96.1%, while 4C4D retained 85.4% and 4DGS 78.2%. These are relative score comparisons, so they indicate sensitivity to the evaluation substrate rather than uncertainty intervals around the estimates.
Checks against people and executable actions
The automated measures were compared with a separate human check. Seventy participants with normal or corrected-to-normal vision, who had not seen the evaluated samples or the design of RoboPhyscore, gave holistic ratings on a 10-point scale. Each sample was rated independently by at least three people. RoboPhyscore had a Pearson correlation of 0.9761 with the human ratings and a Spearman rank correlation of 0.8962. AFS had a Pearson correlation of 0.9746, a Spearman rank correlation of 0.9715 and a mean absolute error of 0.0284.
The authors also tested how prompt detail changed the scores. With Wan, the results were 0.4962 for the instruction condition, 0.5474 for the previous-prompt condition and 0.5782 for the current-prompt condition. Cosmos moved from 0.6519 to 0.6541 and then 0.6800 across the same conditions. The reported gains over instruction-only prompting were 16.5% for Wan and 4.3% for Cosmos. Because this was a within-study ablation without randomized prompt assignment or causal estimation, it shows a score difference under these conditions, not that extra prompt detail caused the improvement.
When the generated videos were turned into actions, DreamGen ranked best among the compared inverse dynamics models, or IDMs. It had the lowest reported mean absolute error, 0.0112, and the highest Action Planner and Data Engine success scores, 0.9825 and 0.5875. Compared with MIDM, the reported differences were a 59.0% lower error and success scores 9.2% and 23.7% higher. The paper reports no uncertainty intervals for this IDM comparison.
A useful benchmark with a narrow reach
The authors say the benchmark currently uses a single simulation platform and embodiment. They identify cross-embodiment, cross-dataset and sim-to-real evaluation, as well as broader VLA-policy comparisons, as future work. The reported results are therefore a detailed comparison inside this setup. They do not establish that Cosmos will transfer to other robots, datasets or the physical world, or that a higher composite score guarantees higher task success.
Paper data and sources
Original title: RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction
Authors: Tianyi Wang, Jiazhou Chen, Yiming Xu et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text