How it works
Six layers from the robot cell to the signed report, and the rules the system enforces.
Who it's for
I have robots. I want to know if my policy is any good without building my own eval.
See their terminal → ROLE B Evaluators and buyersI have no robots and won't buy three to try. I need a result I can defend.
See their order form → ROLE C Judge APIWhat happens inside the machine after one call comes in.
See the seven steps →A policy runs read-only next to a human operator; we measure where it diverges and how that changes across bodies.
Open Shadow mode →An instrument, not a leaderboard
We don't measure "who is best". We measure how this unit did, this time, under these conditions.i
rc-eval open source · free
Protocol and harness anyone runs on their own robot.i
Judge API metered
Hosted, versioned judge, billed per call.i
RC Verified not self-serve
RC operators, RC cells, private B-tier tasks: a signed result others can cite.
Six layers
L0PHYSICAL · CELLTable, arm, hand, three cameras, local GPU, e-stop, reset fixturefleet real · no cell yet
| Eval cell hardware | new | ~1 wk / cell |
| Cell manifest + capability contract | extend | 1.5 |
| Serial-level fleet registry (RCSN) | new | 1.5 |
| E-stop (API) | reuse | 0 |
- Fleet is real: 60+ humanoids, 89 arms, 41 hands.
- No eval cell has been built yet.
- Only xArm has a hardware e-stop input; Piper, YAM and OpenArm have none.
- No serial-level registry in code. The ERP holds serials, but most were never scanned.
L1CAPTURE · RC-RUNNEROn-site agent, synchronized recording, immutable bundleparts exist · 0 real recordings
| rc-runner (binds a cell into one run) | new | 4 |
| Robot drivers (9 teleop agents) | reuse | 0.5 |
| Sensor agents (7 RGBD + 4 LiDAR) | reuse | 0.5 |
| Time base ts_ns + ts_mono_ns | reuse | 0 |
| Bundle writer + seal | new | 2 |
| Rights + attribution at t=0 | new | 0.5 |
| Operator / reset timing | new | 0.5 |
| Edge install / autostart | reuse | 0.5 |
- Time base, robot agents and sensor agents exist and run with --mock.
- Shadow capture writes the obs / pred / div contract (PROTOCOL.md v1–v9).
- 6 sessions on disk, all synthetic. Real recordings: 0.
- rc-runner and the bundle writer are not built.
L2RECORD · RUN BUNDLEGCS, content-hashed, provably unmodifiedspec + 9 real clips
| Bundle uploader (resumable) | extend | 1.5 |
| Object store (GCS buckets, IAM) | extend | 0.5 |
| Bundle verifier (hash, seal, clock health) | new | 1.5 |
| Ingest job runner | reuse | 1 |
| Run index (Postgres) | new | in schema work |
- Bundle format is specified in run-bundle.md.
- 9 real task clips are packaged as bundles; the dry run passes.
- The platform pipeline still writes to /tmp, not GCS, so it is not end to end.
- Verifier and run index are not built.
L3JUDGE · JUDGE + OPERATOR CONSOLEEvery judge-vs-operator disagreement is recordedjudge built · human ceiling unknown
| judge-svc (versioned VLM scorer) | new | 5 |
| VLM backend router + failover | reuse | 0.5 |
| Per-call cost accounting | reuse | 0.2 |
| Judge version registry | extend | 1.5 |
| Operator Console (blind until submit) | extend | 5 |
| Disagreement ledger | new | 1.5 |
| Re-judge orchestrator | new | 2 |
| Recusal enforcement (DB constraint) | new | 1 |
- Judge runtime is built and mutation-tested (37/37).
- Measured cost: $0.0697 per judgment (gpt-4o, 3-model ensemble, 6 calls).
- 900 blind labels sit on 900 different runs, none double-labeled, so no κ can be read. Cause: rci_review_next() only hands out unlabeled runs.
- The Operator Console exists as a demo page.
L4LEDGER · COEFFICIENT TABLE + RCSN PASSPORTSAppend-onlydemo ledger live · 0 real rows
| Coefficient table (append-only) | new | 2 |
| RCSN passport | new | 3 |
| Compliance fields (CN 29-digit ID, FCC, Covered List) | new | 1 + legal review |
| Fault → passport writer | new | 1 |
- A demo ledger runs in Supabase with 1,979 demo runs, every one flagged is_demo.
- Real coefficient-table rows: 0.
- Passport tables and their 7 invariants exist, on demo data.
L5PRODUCT · REPORTS / ATTESTATIONS / QUERY / LIVEOnly A-tier runs can go livethis site · eval revenue $0
| Report renderer | reuse | 3 |
| Attestation signer / verifier | extend | 3 |
| Attestation delivery (share links) | reuse | 0.5 |
| Public A-tier leaderboard | extend | 2 |
| Query API | new | 2 |
| 24/7 A-tier stream | extend | 4 |
| rc-eval open-source harness | new | 4 |
| CLI / MCP | extend | 1.5 |
| Booking → payment | reuse | 1.5 |
| Judge API metering → charging | extend | 2.5 |
- This site, including a Shadow report rendered from real evalsvc output on synthetic input.
- Eval revenue: $0. Metering exists but only estimates; nothing is charged.
- Stream, attestation signing and the public query API are not built.
Six rules
Structural constraints enforced by the system, not policies.
- 01B-tier task content is exposed to no one through any interface i
- 02Anyone who has seen a B-tier task does not score it (recusal, system-enforced)
- 03
attemptandsuccessare never merged into one score - 04Exclusion =
counts_against_policy == false, looked up in the taxonomy, never a hard-coded id i - 05Buyers pay for acceptance testing; vendors may pay for VOC, compatibility or support, never for a verdict
- 06Public reports anonymize bodies by default i
The demo ledger, live
What the ledger records today. Every number below is aggregated live from seeded demo runs.
DEMO DATA Read live from the database; scores are seeded.
Per body · success vs attempt
Cross-body σ · same policy, different body i
Task reference footage
Human reference clips of each task, filed by setting › environment. Not recordings of these runs; the line under each clip says what it actually shows, and a click opens the task's row in the table.