Preprint

Physical AI benchmarks often rank models in similar ways

Preprint audit finds strong overlap among benchmark rankings and suggests that a smaller testing suite may retain much of the useful information.

Physical-AI benchmarks often rank models in similar ways, suggesting that a longer test list does not automatically provide more independent evidence. An audit of 12 physical-AI benchmarks found positive relationships across every benchmark pair. The mean Spearman rank correlation was 0.487, a measure of how similarly two rankings order the same models.

The clearest overlap was between EmbSpatial and CV-Bench. Their correlation was 0.876 across 49 models, with an interval from 0.78 to 0.93. Where2Place and RefSpatial-Bench were close too, with a correlation of 0.860 across 50 models and an interval from 0.73 to 0.93. The intervals show that the precise strength of each relationship is estimated rather than known with complete certainty.

The audit's basic question

The study was built around three linked questions: how much information the 12 benchmarks share, whether redundancy affects pooled averages, and how small a suite can remain while still separating models and covering abilities that are not repeated across tests.

The researchers began with a registry containing 152 models and 51 candidate physical-AI benchmarks. A benchmark had to have results reported for at least five candidate models to enter the screening process. The final matrix retained 51 models, each with scores for at least eight of the 12 selected benchmarks. The screen considered how densely each benchmark was represented, how recent it was and how much variety it added.

The score matrix combined published results with targeted evaluation runs. For missing evaluations, the researchers used each benchmark's official code and prompts when those were available, used greedy decoding and followed the benchmark's own answer-parsing rule. Those additional runs contributed 159 data points. The authors did not systematically reproduce the published matrix.

Some tests were easier to predict from the rest

A separate analysis used leave-one-model-out ridge regression, a prediction method that estimates a held-out model's score from its results on the other benchmarks. This tests how much a benchmark can be reconstructed from the rest of the suite. Across the highlighted benchmarks, the reported fit values ranged from 0.319 to 0.727.

Where2Place, RefSpatial-Bench and ERQA were the most predictable from the remaining suite, with fit values of 0.727, 0.721 and 0.706. RealWorldQA, RoboSpatial and BLINK were the least predictable, at 0.319, 0.351 and 0.378. A higher fit means the remaining benchmarks were better at predicting the withheld score.

That pattern is informative, but it is not a standalone verdict on benchmark quality. A score that is easy to predict from the rest may add less distinct information to the combined picture. A score that is harder to predict may capture something different, but the same unpredictability can also reflect measurement noise.

Why the choice of tests can change the ranking

The issue becomes important when benchmark scores are averaged. Giving every test equal weight can allow closely related benchmarks to give a repeated ability more influence than a distinct ability. That means adding another test can change a pooled model ranking even when it adds little independent information.

The audit compared rankings before and after substitute benchmark pairs were collapsed, alongside comparisons between the full collection and smaller benchmark combinations. The exercise examines how the composition of a suite shapes its summary ranking. It does not show that performance on one benchmark causes performance on another.

A smaller suite is presented as a practical option

The researchers also tested smaller combinations through greedy forward selection, adding benchmarks one at a time according to a utility that balanced model separation with non-redundant information. Under that utility, they present the resulting compact suite as a defensible choice rather than a uniquely correct answer.

The selection depends on the chosen utility and on how model discrimination is measured. Because the procedure is greedy, it has no guarantee of finding the globally best combination. A different objective or selection method could therefore produce a different compact suite.

What the findings do not settle

This is an audit of aggregate benchmark scores, not item-by-item responses. That limits what the researchers can say about whether particular questions are duplicated and makes direct estimates of measurement error unavailable.

The evidence also comes from a mixture of model cards, benchmark papers and targeted runs, without a systematic reproduction of the published score matrix. The comparisons should therefore be read as an audit of the assembled results, not as an independent re-evaluation of every score.

Finally, the benchmark results do not establish success on a downstream physical-system task. The audit shows how scores relate within this benchmark collection, but it does not show that the shared or residual signal transfers to manipulation or navigation.

The main message for anyone comparing physical-AI models is simple: counting benchmarks is not the same as counting independent evidence. The audit found broad agreement among model rankings, especially for two benchmark pairs, while several other tests were less predictable from the rest. A shorter suite may be useful, but its design choices remain part of the result.

Paper data and sources

Original title: A Statistical Audit of Physical AI Benchmark Redundancy
Authors: Zaruhi Navasardyan, Hrant Davtyan
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.