SIMSimulated data. Every score, rate and price on this site is illustrative.
Instrument

How it works

Six layers from the robot cell to the signed report, and the rules the system enforces.

Who it's for

ROLE A Data producers

I have robots. I want to know if my policy is any good without building my own eval.

See their terminal →
ROLE B Evaluators and buyers

I have no robots and won't buy three to try. I need a result I can defend.

See their order form →
ROLE C Judge API

What happens inside the machine after one call comes in.

See the seven steps →
SECTION · SHADOW MODE Shadow mode

A policy runs read-only next to a human operator; we measure where it diverges and how that changes across bodies.

Open Shadow mode →

An instrument, not a leaderboard

We don't measure "who is best". We measure how this unit did, this time, under these conditions.i

rc-eval open source · free

Protocol and harness anyone runs on their own robot.i

Judge API metered

Hosted, versioned judge, billed per call.i

RC Verified not self-serve

RC operators, RC cells, private B-tier tasks: a signed result others can cite.

Six layers

L0PHYSICAL · CELLTable, arm, hand, three cameras, local GPU, e-stop, reset fixturefleet real · no cell yet
Components i
Eval cell hardwarenew~1 wk / cell
Cell manifest + capability contractextend1.5
Serial-level fleet registry (RCSN)new1.5
E-stop (API)reuse0
≈ 3 pw of listed work
Where it stands today
  • Fleet is real: 60+ humanoids, 89 arms, 41 hands.
  • No eval cell has been built yet.
  • Only xArm has a hardware e-stop input; Piper, YAM and OpenArm have none.
  • No serial-level registry in code. The ERP holds serials, but most were never scanned.
L1CAPTURE · RC-RUNNEROn-site agent, synchronized recording, immutable bundleparts exist · 0 real recordings
Components i
rc-runner (binds a cell into one run)new4
Robot drivers (9 teleop agents)reuse0.5
Sensor agents (7 RGBD + 4 LiDAR)reuse0.5
Time base ts_ns + ts_mono_nsreuse0
Bundle writer + sealnew2
Rights + attribution at t=0new0.5
Operator / reset timingnew0.5
Edge install / autostartreuse0.5
≈ 8.5 pw of listed work
Where it stands today
  • Time base, robot agents and sensor agents exist and run with --mock.
  • Shadow capture writes the obs / pred / div contract (PROTOCOL.md v1–v9).
  • 6 sessions on disk, all synthetic. Real recordings: 0.
  • rc-runner and the bundle writer are not built.
L2RECORD · RUN BUNDLEGCS, content-hashed, provably unmodifiedspec + 9 real clips
Components i
Bundle uploader (resumable)extend1.5
Object store (GCS buckets, IAM)extend0.5
Bundle verifier (hash, seal, clock health)new1.5
Ingest job runnerreuse1
Run index (Postgres)newin schema work
≈ 4.5 pw of listed work
Where it stands today
  • Bundle format is specified in run-bundle.md.
  • 9 real task clips are packaged as bundles; the dry run passes.
  • The platform pipeline still writes to /tmp, not GCS, so it is not end to end.
  • Verifier and run index are not built.
Go deeper Ledger → · Schema →
L3JUDGE · JUDGE + OPERATOR CONSOLEEvery judge-vs-operator disagreement is recordedjudge built · human ceiling unknown
Components i
judge-svc (versioned VLM scorer)new5
VLM backend router + failoverreuse0.5
Per-call cost accountingreuse0.2
Judge version registryextend1.5
Operator Console (blind until submit)extend5
Disagreement ledgernew1.5
Re-judge orchestratornew2
Recusal enforcement (DB constraint)new1
≈ 16.7 pw of listed work
Where it stands today
  • Judge runtime is built and mutation-tested (37/37).
  • Measured cost: $0.0697 per judgment (gpt-4o, 3-model ensemble, 6 calls).
  • 900 blind labels sit on 900 different runs, none double-labeled, so no κ can be read. Cause: rci_review_next() only hands out unlabeled runs.
  • The Operator Console exists as a demo page.
L4LEDGER · COEFFICIENT TABLE + RCSN PASSPORTSAppend-onlydemo ledger live · 0 real rows
Components i
Coefficient table (append-only)new2
RCSN passportnew3
Compliance fields (CN 29-digit ID, FCC, Covered List)new1 + legal review
Fault → passport writernew1
≈ 6 pw of listed work
Where it stands today
  • A demo ledger runs in Supabase with 1,979 demo runs, every one flagged is_demo.
  • Real coefficient-table rows: 0.
  • Passport tables and their 7 invariants exist, on demo data.
L5PRODUCT · REPORTS / ATTESTATIONS / QUERY / LIVEOnly A-tier runs can go livethis site · eval revenue $0
Components i
Report rendererreuse3
Attestation signer / verifierextend3
Attestation delivery (share links)reuse0.5
Public A-tier leaderboardextend2
Query APInew2
24/7 A-tier streamextend4
rc-eval open-source harnessnew4
CLI / MCPextend1.5
Booking → paymentreuse1.5
Judge API metering → chargingextend2.5
≈ 24 pw of listed work
Where it stands today
  • This site, including a Shadow report rendered from real evalsvc output on synthetic input.
  • Eval revenue: $0. Metering exists but only estimates; nothing is charged.
  • Stream, attestation signing and the public query API are not built.

Six rules

Structural constraints enforced by the system, not policies.

  1. 01B-tier task content is exposed to no one through any interface i
  2. 02Anyone who has seen a B-tier task does not score it (recusal, system-enforced)
  3. 03attempt and success are never merged into one score
  4. 04Exclusion = counts_against_policy == false, looked up in the taxonomy, never a hard-coded id i
  5. 05Buyers pay for acceptance testing; vendors may pay for VOC, compatibility or support, never for a verdict
  6. 06Public reports anonymize bodies by default i

The demo ledger, live

What the ledger records today. Every number below is aggregated live from seeded demo runs.

DEMO DATA Read live from the database; scores are seeded.

Per body · success vs attempt

success_rate attempt_rate

Cross-body σ · same policy, different body i

σ success σ attempt

Task reference footage

Human reference clips of each task, filed by setting › environment. Not recordings of these runs; the line under each clip says what it actually shows, and a click opens the task's row in the table.

Show the footage wall

Latest runs

Show the 12 most recent runs

The zeros that haven't moved

Humanoids / arms / hands60+ · 89 · 41
Eval revenue$0
Table rows0
Real runs0
Citing papers0
M0 size i480 runs
M0 effort2 eng-wk + 3 op-days