Skip to content
PPHYSICAL AI GUIDESTART HERE →

Field guide

Are Physical AI Benchmarks Measuring the Same Thing? A Redundancy and Ranking Audit

A practical audit of benchmark overlap, rank sensitivity, general capability, compact suites, and what current physical AI leaderboards do not establish.

By Physical AI Guide Editorial TeamPublished Updated

The short answer

Benchmark selection is part of a physical AI result. If a suite measures the same ability several times, a simple average gives that ability several votes. A model’s rank can then reflect the composition of the suite as much as a broad physical capability.

A new statistical audit makes that risk measurable. The authors assembled 51 models across 12 physical AI benchmarks, filling a dense matrix with 405 scores grouped as published results and 159 scores from their own runs under official protocols. They report two close substitute pairs, a strong general vision-language factor, and a material ranking change after duplicate abilities receive less weight (paper).

The most useful result is not a new top model. It is a warning about the instrument:

  • the average pairwise Spearman correlation among the 12 benchmarks is 0.487;
  • EmbSpatial and CV-Bench correlate at 0.876;
  • Where2Place and RefSpatial-Bench correlate at 0.860;
  • collapsing those two pairs moves 22 of 51 models by at least three places;
  • four selected benchmarks retain 78.5% of the paper’s full-suite utility;
  • the first principal component of the physical suite correlates at 0.952 with a separate general-benchmark component.

These are author-run statistical findings from one matrix. They do not show that the four-benchmark subset is uniquely correct. They also do not establish physical robot performance. None of the 12 benchmarks measures downstream task success on hardware.

Our guide to reading robot benchmark success rates explains how tasks, embodiments, denominators, execution stacks, and failure rules constrain one reported percentage. This audit adds another layer: even correctly reported benchmark scores can form a misleading aggregate when the suite repeats shared signal.

What the audit actually measured

The paper began with a registry of 51 physical AI benchmarks and 152 models. It retained 12 benchmarks based on reporting density, recency, and intended diversity. A model entered the final set if it had scores for at least eight of the 12 columns.

The final suite is dominated by spatial and visual reasoning, but it contains several formats:

Benchmark group Examples Output or evidence tested
Pointing RefSpatial-Bench, Where2Place An image coordinate inside a target mask
Single-image relations EmbSpatial, CV-Bench, OmniSpatial, RealWorldQA Relative position, depth, count, or inferred viewpoint
Multi-view and video MindCube, SAT, VSI-Bench Spatial consistency across views or video
Embodied framing ERQA, RoboSpatial Graspability, placement, feasibility, or robot-centric reasoning
Multi-image perception BLINK General visual perception across paired or multiple images

Ten benchmarks use multiple-choice answers. RefSpatial-Bench and Where2Place require a point. Every score is placed on a 0 to 100 scale.

This scope is important. The paper uses “physical AI benchmark” for tests related to spatial, visual, and embodied reasoning. It does not evaluate action policies closing a control loop on physical robots. The matrix can test whether benchmark scores overlap. It cannot establish whether any selected test predicts grasp success, collision avoidance, intervention rate, or uptime.

Why a simple average can double-count an ability

Suppose a suite has two pointing tests, one video-spatial test, and one contact-control test. An equal average does not give each underlying ability equal weight. Pointing receives half of the total because it has two columns.

That is the central aggregation problem. Equal benchmark weights are not equal capability weights when benchmarks overlap.

The paper checks pairwise overlap with Spearman rank correlation. This asks whether two benchmarks order models similarly without requiring the raw score scales to match. All pairwise correlations in the 12-benchmark matrix are positive, with a mean of 0.487.

Two pairs stand out:

Benchmark pair Spearman correlation Overlapping models Defensible reading
EmbSpatial and CV-Bench 0.876 49 The two rankings carry closely related signal in this model set
Where2Place and RefSpatial-Bench 0.860 50 Two pointing formats behave like close substitutes in this matrix
ERQA and RealWorldQA 0.758 Pairwise-complete rows Strong overlap, but below the paper’s substitute threshold
RealWorldQA and OmniSpatial 0.748 Pairwise-complete rows Substantial shared ranking signal

Correlation does not prove that two datasets contain duplicate items or identical skills. Shared model families, common training data, prompt format, measurement noise, or a broad general-capability factor can all contribute. The correct conclusion is narrower: these benchmark columns do not provide independent votes in this observed matrix.

The rank sensitivity result

To make the weighting effect concrete, the authors first rank models by an equal average across all 12 columns. They then replace each close substitute pair with one averaged column. The revised suite has ten columns, eight unchanged benchmarks and two combined abilities.

Twenty-two of 51 models move by at least three places. The paper reports examples in both directions:

  • MiMo-Embodied-7B falls nine places;
  • Gemini Robotics-ER 1.5 falls eight;
  • GPT-4o rises nine;
  • Claude Sonnet 4 rises eight.

This does not prove that the ten-column ranking is correct. Collapsing a pair is itself an editorial and statistical choice. The experiment proves something more useful: a material part of model standing came from how often the chosen suite measured abilities where that model was relatively strong.

A leaderboard should therefore publish sensitivity, not just one order. At minimum, show what happens when:

  1. correlated columns are grouped;
  2. task families receive equal weight;
  3. each benchmark receives equal weight;
  4. missing scores are excluded, imputed, or penalized;
  5. score gaps are retained or reduced to pairwise wins;
  6. model-family duplicates are handled separately.

If reasonable choices produce different top groups, the uncertainty belongs in the result.

Reconstructability: overlap beyond one pair

Pairwise correlation can miss distributed redundancy. A benchmark may not have one obvious twin, yet several other columns together may reconstruct much of its variance.

The audit fits ridge regressions that predict each benchmark from the other 11 and evaluates them with leave-one-model-out cross-validation. Reported reconstructability is highest for:

  • Where2Place, R² 0.727;
  • RefSpatial-Bench, R² 0.721;
  • ERQA, R² 0.706.

It is lowest for:

  • RealWorldQA, R² 0.319;
  • RoboSpatial, R² 0.351;
  • BLINK, R² 0.378.

A low R² can mean useful distinct information. It can also mean noise. Without repeated evaluation across prompts, decoding seeds, parser choices, and resampled items, uniqueness and unreliability are difficult to separate.

That is a critical evidence boundary. A benchmark does not become valuable merely because its rankings disagree with everything else. Disagreement deserves diagnosis.

The general-capability problem

The audit’s strongest caution appears in its external comparison. The authors add nine general language and vision benchmarks that were not designed to test physical or spatial reasoning. They estimate the first principal component of the 12 physical benchmarks, then compare it with the general anchors.

The physical component correlates at:

  • 0.952 with the general-benchmark principal component;
  • 0.950 with MMStar;
  • 0.942 with Video-MME;
  • 0.820 with MMMU.

After the authors regress out the external general axis, mean pairwise correlation among the physical benchmarks falls from 0.487 to 0.250. The number of benchmark pairs above 0.5 falls from 34 of 66 to six.

This supports a bounded statement: roughly half of the shared score structure in this matrix aligns with general vision-language capability. It does not show that physical reasoning is absent. Three relationships remain strong after residualization, including both substitute pairs and ERQA with MindCube.

The practical risk is clear. A leaderboard labeled “physical AI” can substantially rank broad visual and language competence while leaving control, contact, state estimation, and recovery untested.

For robot foundation models, pair benchmark results with evidence from the actual control stack. Our embodiment gap guide maps the adaptation, calibration, action, contact, and safety work that a vision-language score does not cover.

What the four-benchmark subset means

The paper selects a compact suite with a greedy utility. A candidate benchmark receives value when it does two things:

  1. spreads models apart, measured with the Gini coefficient of scores;
  2. adds variance not reconstructed by benchmarks already selected.

Under that formula, the first four are:

  1. RefSpatial-Bench, precise spatial localization by pointing;
  2. MindCube, spatial consistency across limited views;
  3. VSI-Bench, spatial reasoning over video;
  4. BLINK, multi-image perceptual primitives.

Together they retain 78.5% of the cumulative utility assigned to all 12. Where2Place enters fifth, bringing the total to 85.0%. The final four columns add 4.4 percentage points between them.

The subset is a defensible compression under the paper’s objective. It is not a universal benchmark standard because:

  • forward selection is greedy and has no optimality guarantee;
  • discrimination is defined with one chosen statistic;
  • marginal information depends on the available model sample;
  • some models lack one of the four selected scores;
  • utility is calculated on aggregate benchmark scores, not item responses;
  • no selected benchmark measures physical task execution.

A compact suite can reduce evaluation cost and repeated measurement. It cannot replace coverage that was never present.

Why the compact leaderboard is secondary

The authors fit a Bradley-Terry model on the selected four benchmarks. Each benchmark acts like a judge for a pair of models, voting for the higher score. The fit uses 4,231 pairwise observations across 51 models.

This avoids treating a ten-point gap on one benchmark as automatically ten times a one-point gap on another. It also introduces its own choices. Every benchmark vote has equal weight, score magnitude is discarded, and several top models have only three of the four core scores.

The resulting ranking is best read as a worked example of one compact aggregation method. It is not an ordering of complete robots, action policies, safety systems, or deployed products. We do not reproduce the top-ten table because that would foreground the least durable part of the study.

The durable contribution is the audit procedure:

  • build a score-provenance matrix;
  • measure overlap;
  • identify general versus domain-specific shared structure;
  • test rank sensitivity;
  • select a smaller suite under an explicit utility;
  • state what the selected tests still omit.

Evidence limits in the source matrix

The paper is unusually direct about several limitations.

Published scores were not systematically reproduced

Of 564 filled cells, 379 are direct transcriptions from model cards or papers, 25 are medians of conflicting published values, one is borrowed from a twin model, and 159 come from the authors’ own runs. The team validated its implementation on models with published results, but it did not conduct a systematic reproduction study.

A dense matrix can therefore be more comparable than scattered vendor tables without being perfectly homogeneous. Model versions, prompts, serving changes, parser behavior, and undocumented evaluation details may remain.

Missing data still matters

The model inclusion rule requires at least eight of 12 scores, not all 12. Pairwise correlations use available overlapping rows. The compact leaderboard can include a model with only three core scores.

Future releases should publish the complete matrix and every missing-score rule in a machine-readable artifact. Readers need to reproduce rank changes under complete-case, pairwise, and explicit imputation choices.

Aggregate scores hide item structure

The audit works at benchmark level because public reports rarely expose item-level responses. That permits useful covariance analysis, but it cannot locate duplicate items, estimate per-item reliability, or show whether overlap comes from one task family inside a broad benchmark.

Statistical uniqueness is not physical validity

A benchmark can add new variance while measuring an irrelevant shortcut. Conversely, two correlated benchmarks can differ in robustness, contamination, annotation quality, or diagnostic usefulness. Selection should combine statistical audits with construct validity and downstream tests.

The study is observational

The analysis does not intervene on training data or objectives. It can show covariance, not why models perform similarly across tests. No result establishes that the residual “physical” component causes better manipulation or navigation.

A publication checklist for physical AI leaderboards

Before trusting one aggregate rank, ask for this evidence card.

Field Minimum disclosure
Construct The ability each benchmark is intended to measure
Format Inputs, outputs, prompts, parser, and scoring rule
Matrix Every model-by-benchmark score and missing cell
Provenance Published transcription, conflicting value, or new evaluation run
Versioning Model endpoint, checkpoint, benchmark commit, dataset split, and date
Reliability Repeated prompts, decoding seeds, parser sensitivity, and item uncertainty
Overlap Pairwise correlation and multivariate reconstructability
General factor Comparison with broad vision and language anchors
Weighting Equal benchmark, equal capability group, and alternative weights
Sensitivity Rank movement across defensible aggregation choices
Physical link Separate evidence connecting the score to robot task outcomes
Limits Explicitly unmeasured control, contact, safety, intervention, and deployment fields

This complements the trial-level checklist in our dynamic robot benchmark guide. Statistical independence and operational completeness answer different questions. A serious evaluation needs both.

Evidence verdict

Claim Classification Confidence Why
The audited suite contains substantial redundant signal Author-run statistical audit Medium-high Two pairs exceed 0.86 correlation and several columns are reconstructable, but the result depends on one model matrix
Redundancy can materially change equal-average rankings Direct sensitivity result in the audit Medium-high Twenty-two of 51 models move at least three places after two pairs are collapsed
Four benchmarks are sufficient for all physical AI evaluation Not supported High The subset retains 78.5% under one utility and contains no physical control benchmark
The suite substantially reflects general capability Author-run external-factor analysis Medium-high Physical PC1 correlates at 0.952 with general PC1, with bounded model overlap and observational limits
A unique benchmark is necessarily a valid benchmark Not supported High Low predictability can come from distinct signal or noise
The compact ranking orders complete robot systems Not supported High The tests evaluate model responses, not downstream robot execution, safety, or operation
Benchmark families should receive explicit weights Practical inference from rank sensitivity Medium-high Equal column weights silently repeat abilities when correlated tests receive separate columns

What to watch next

  1. Release of the complete score matrix, provenance labels, evaluation configurations, and analysis code.
  2. Independent reproduction of the 159 newly evaluated cells.
  3. Sensitivity to alternative missing-score rules and model-family sampling.
  4. Item-level overlap analysis inside the 12 benchmarks.
  5. Repeated runs that separate measurement noise from genuinely unique signal.
  6. Alternative utility functions for compact-suite selection.
  7. Validation on later models without reselecting the suite.
  8. A bridge from spatial and embodied reasoning scores to matched physical manipulation or navigation outcomes.
  9. Benchmark groups that cover action, contact, uncertainty, latency, recovery, and intervention.
  10. Leaderboards that publish rank intervals or stability bands instead of one precise order.

Bottom line

The audit demonstrates that counting benchmarks is not the same as counting independent evidence. In its 51-model matrix, two duplicated ability pairs materially affect ranking, a compact subset preserves much of the chosen utility, and a broad general-capability axis explains a large share of common movement.

The correct response is not to discard spatial benchmarks or adopt one four-test standard. It is to make aggregation inspectable. Publish the matrix, provenance, overlap, weights, missing-data rules, and sensitivity. Then connect model scores to physical trials without pretending that a visual reasoning leaderboard already measures robot reliability.

A physical AI leaderboard is a measurement design. Its ranking is credible only when readers can see which abilities received a vote, how often they voted, and what physical behavior remains outside the test.

Frequently asked questions

Are physical AI benchmarks redundant?

A 2026 author-run audit found substantial overlap among 12 spatial and embodied reasoning benchmarks evaluated across 51 models. Two benchmark pairs had Spearman correlations above 0.86, and the median benchmark had roughly half of its variance reconstructed from the other 11. This is evidence of overlap in that matrix, not proof that every physical AI benchmark is redundant.

Does benchmark redundancy change model rankings?

It can. In the audited matrix, collapsing two highly correlated benchmark pairs into two combined columns moved 22 of 51 models by at least three ranking places under an equally weighted average. The revised order is not a uniquely correct leaderboard. It demonstrates that suite composition can materially affect rank.

Which four benchmarks did the audit select?

Under the paper's greedy utility, the selected four were RefSpatial-Bench, MindCube, VSI-Bench, and BLINK. Together they retained 78.5% of the utility assigned to all 12 benchmarks. That result depends on the chosen discrimination and marginal-information formula and is not a universal minimum suite.

Do these benchmarks measure physical robot performance?

No. The 12 audited benchmarks test visual, spatial, pointing, multi-view, video, and embodied reasoning through model responses. None measures downstream task success on a physical robot. They should not be presented as evidence of manipulation reliability, navigation reliability, safety, uptime, or deployment.

What should a benchmark report include?

Publish the model-by-benchmark matrix, score provenance, prompts, parsers, decoding settings, missing-score rules, task and item weights, uncertainty, rank sensitivity under alternative weights, artifact versions, and a clear statement of what the benchmark does not test.