Ecosystem
Benchmarks, open datasets and taxonomies: what we adopt, where the gap is.
Which vocabularies to adopt so our rows join with everyone else's, and which cell is genuinely empty.i
No public failure taxonomy blames the hardware
Each one charges every failure to the policy, so the hardware column is empty.
Taxonomy axis coverage · the empty column is the opening
All failure taxonomies compared
| Taxonomy | Classes | Policy axis | Hardware axis | Env axis | Note |
|---|
Open datasets record successes
They are large, and almost none label failures.
Open dataset scale · and whether failures are labelled
All datasets
| Dataset | Trajectories | Bodies | Tasks | Failures labelled? | Note |
|---|
Adopt four things as they are
Each one lets our results line up with work others already publish.
- 01RoboArena's double-blind pairwise protocol. Adopt it as is; don't invent a ranking method.i
- 02Table30's specialist / generalist tracks. One policy per task and one policy per body are different questions; don't report them together.
- 03Open X-Embodiment's skill vocabulary. Align
task_familyto it so our rows join its 1M trajectories.i - 04SimplerEnv's perturbation axes. Visual Matching and Variant Aggregation map straight onto our A/B-tier perturbation table.
Don't build another simulation benchmark
LIBERO, CALVIN and RoboCasa already fill that cell well. Our only structural advantage is the real robot.i
Every row links to its source
The full comparison behind this page.
Benchmarks
| System | Kind | Bodies | Tasks | What it gets right | What it lacks | Our stance |
|---|
Crosswalk: our classes and theirs
Our classes ↔ theirs.i
| External taxonomy | Their class | Our class | Relation | Note |
|---|