SIMSimulated data. Every score, rate and price on this site is illustrative.
Instrument

Judge your own runs

For labs with robots: $0.60 a run tells you how your policy compares.

See the three commands

You pay per run, you get a comparison

Role A · data producerUniversity labs · funded AI teams that can't get robots · cell renters
I pay
$0.60 per judged run i
I get
A free protocol, cheap judging and comparability: I finally know whether 57.5% is high or low
RC gets
The failure distribution of a machine RC never bought i

Three commands from robot to ranking ILLUSTRATIVE

$ pip install rc-eval && rc-eval init
  ✓ found   xArm 7 (ufactory)   @ 192.168.1.42
  ✓ found   Wuji Hand 2        @ /dev/ttyUSB0
  ✓ cameras 3 (scene / wrist / top)   sync ±1.8 ms
  body_hash  b7c1…  → new config, registered as lineage#0x41 v1

$ rc-eval run --task tabletop.pick_place --n 40 --policy my-vla:v3
  [40/40] ██████████████████████  18 min 22 s
  bundle  ./runs/2026-09-20-a41/  3.1 GB  blake3:7c1d…
  reset   avg 96 s  ← goes into your own unit economics

$ rc-eval judge --upload
  judged 40 · rc-judge-v1 · $24.00
  attempt  95.0%    success  57.5%
  top failures  grasp.slip (14)   perception.reflectance (3)

  ▸ row added to the public table (anonymized)  XE-Table 2026.3 / cf_4488
  ▸ median success, same task + body family = 61%  — you are 3.5pt below
The last line is why they stay. They can judge on their own; they can't get comparability with one body.i

Eleven fields, all required

None of them can be filled in afterwards. Missing any one, the API rejects the run.i

All eleven fields
  • Policy: vendor / name / version / checkpoint or endpoint snapshot
  • Body: arm + hand + DoF + mount height + config hash
  • Cell and last calibration date
  • Task id and tier
  • Random seed
  • Environment: color temperature / clutter count / table friction / temperature
  • Reset method and reset seconds
  • Operator minutes
  • Rights basis (one of four)
  • Bundle content hash
  • Sync error between the three cameras and joint streams

Your runs never train a rival policy

What RC takes

Submitted runs, in anonymized, aggregated form, for judge calibration and the public table.i

What RC won't do

RC may train judges on these runs, never control policies that compete with the submitter.i