SIMSimulated data. Every score, rate and price on this site is illustrative.
Instrument

Judge API

For developers: send one run, get a verdict and a table row for $0.60.

Read the API spec

One run in, one verdict out

Role C · Judge APIInside one $0.60 call
Input
One run bundle: 3 video streams + joint/torque streams + 11 attribution fields
Output
attempt / success / failure_class / two confidences, plus one row in the public table
Cost · price
$0.181 · $0.60 → 69.8% gross margin i
Surface
OpenAPI 3.1 · 25 endpoints · 4 webhook events · 58 fully defined schemas

Seven steps, no operator in the room

Seven steps at a glance

Cost of one $0.60 call ILLUSTRATIVE

None of the seven steps needs an operator in the room. Cell capacity is bound to people; judge capacity is bound to GPUs, so it should be priced cheap.i
The seven steps in detail TARGETS
  1. 1Intake · validateCheck the bundle hash and all 11 attribution fields. One missing = rejected. i~40 ms
  2. 2Frame samplingDense around contact events i~1.2 s
  3. 3Three-model ensembleThree VLMs judge independently. attempt and success are asked separately, never merged.$0.13 · ~6 s
  4. 4Vote + two confidencesMajority vote; attempt_confidence and success_confidence reported apart i~10 ms
  5. 5Write three ledgersjudgments = judge, operator_labels = human, runs = what happened. Mismatch → calibration set.~25 ms
  6. 6Merge into the public tableAnonymized row into coefficients. Cells with n < 40 show insufficient.async
  7. 7MeteringBilled per call; the key's rights_granted is recorded i$0.60

Every disagreement makes the judge better

↺ Feedback: where the recursion actually is
Step 5's disagreement set → trains rc-judge-v2 → all past runs can be re-judged.i
Every run makes the instrument more accurate, not the robot.

One call, two separate answers ILLUSTRATIVE

Request body
# attribution fields are required; one missing = rejected
POST /v1/judgments
Authorization: Bearer rck_live_…
Idempotency-Key: 8f2a41-01

{
  "bundle_uri": "gs://rc-runs/8f2a41/",
  "policy": { "vendor":"example-lab", "name":"policy-a",
              "version":"6.0.2", "checkpoint":"sha256:9f3c…" },
  "body":   { "arm":"xArm7", "hand":"WujiHand2", "dof":20,
              "mount_h_mm":740, "calibrated_at":"2026-09-16" },
  "task":   { "id":"tabletop.pick_place.mug_022", "tier":"A" },
  "env":    { "lighting_k":4200, "clutter":3, "friction_mu":0.42 },
  "seed":   "0x5F3A9C21",
  "reset":  { "type":"manual", "seconds":42, "operator_min":2.4 },
  "judge":  "rc-judge-v1"
}
200 OK
{
  "judgment_id": "jd_5c19",
  "attempt": true,              // language-layer property, always returned alone
  "success": false,             // body-layer property, always returned alone
  "failure_class": "grasp.slip",
  "counts_against_policy": true,  // looked up in the taxonomy, not a hard-coded list
  "attempt_confidence": 0.91,
  "success_confidence": 0.86,
  "charged_to": { "rcsn": null, "cell_id": "cell-sf-03" },
  "judge_version": "rc-judge-v1",
  "table_row_id": "cf_4471"
}

Three lines held by structure, not rules

No POST /v1/attestations

Attestations are read-only in the API and issued offline: the endpoint does not exist.i

B-tier differencing: k = 5

At least 5 tasks per cell, no task_id returned; narrow queries are widened, not refused.i

Exclusion by taxonomy, not by list

20 of 54 classes don't count against the policy, and the number grows.i

Two more rules

B-tier isolation is per connection, not per query. If isolation depends on "don't SELECT those columns", one ORM SELECT * leaks a B-tier task. The Judge API connects with a database role that has no privilege at all on the four content columns, so it cannot leak even by mistake.

Unclassified is not exempt. A failure with an empty failure_class still counts against the policy. Otherwise refusing to label becomes a free way to shrink the denominator and inflate your success rate.