SIMSimulated data. Every score, rate and price on this site is illustrative.
Instrument

Ecosystem

Benchmarks, open datasets and taxonomies: what we adopt, where the gap is.

Which vocabularies to adopt so our rows join with everyone else's, and which cell is genuinely empty.i

No public failure taxonomy blames the hardware

Each one charges every failure to the policy, so the hardware column is empty.

Taxonomy axis coverage · the empty column is the opening

All failure taxonomies compared
TaxonomyClassesPolicy axisHardware axisEnv axisNote

Open datasets record successes

They are large, and almost none label failures.

Open dataset scale · and whether failures are labelled

All datasets
DatasetTrajectoriesBodiesTasksFailures labelled?Note

Adopt four things as they are

Each one lets our results line up with work others already publish.

  1. 01RoboArena's double-blind pairwise protocol. Adopt it as is; don't invent a ranking method.i
  2. 02Table30's specialist / generalist tracks. One policy per task and one policy per body are different questions; don't report them together.
  3. 03Open X-Embodiment's skill vocabulary. Align task_family to it so our rows join its 1M trajectories.i
  4. 04SimplerEnv's perturbation axes. Visual Matching and Variant Aggregation map straight onto our A/B-tier perturbation table.

Don't build another simulation benchmark

LIBERO, CALVIN and RoboCasa already fill that cell well. Our only structural advantage is the real robot.i

Every row links to its source

The full comparison behind this page.

Benchmarks
SystemKindBodiesTasks What it gets rightWhat it lacksOur stance
Crosswalk: our classes and theirs

Our classes ↔ theirs.i

External taxonomyTheir classOur classRelationNote