Judge API
For developers: send one run, get a verdict and a table row for $0.60.
One run in, one verdict out
- Input
- One run bundle: 3 video streams + joint/torque streams + 11 attribution fields
- Output
- attempt / success / failure_class / two confidences, plus one row in the public table
- Cost · price
- $0.181 · $0.60 → 69.8% gross margin i
- Surface
- OpenAPI 3.1 · 25 endpoints · 4 webhook events · 58 fully defined schemas
Seven steps, no operator in the room
Seven steps at a glance
Cost of one $0.60 call ILLUSTRATIVE
The seven steps in detail TARGETS
- 1Intake · validateCheck the bundle hash and all 11 attribution fields. One missing = rejected. i~40 ms
- 2Frame samplingDense around contact events i~1.2 s
- 3Three-model ensembleThree VLMs judge independently. attempt and success are asked separately, never merged.$0.13 · ~6 s
- 4Vote + two confidencesMajority vote;
attempt_confidenceandsuccess_confidencereported apart i~10 ms - 5Write three ledgers
judgments= judge,operator_labels= human,runs= what happened. Mismatch → calibration set.~25 ms - 6Merge into the public tableAnonymized row into
coefficients. Cells with n < 40 showinsufficient.async - 7MeteringBilled per call; the key's
rights_grantedis recorded i$0.60
Every disagreement makes the judge better
Step 5's disagreement set → trains
rc-judge-v2 → all past runs can be re-judged.iEvery run makes the instrument more accurate, not the robot.
One call, two separate answers ILLUSTRATIVE
Request body
# attribution fields are required; one missing = rejected POST /v1/judgments Authorization: Bearer rck_live_… Idempotency-Key: 8f2a41-01 { "bundle_uri": "gs://rc-runs/8f2a41/", "policy": { "vendor":"example-lab", "name":"policy-a", "version":"6.0.2", "checkpoint":"sha256:9f3c…" }, "body": { "arm":"xArm7", "hand":"WujiHand2", "dof":20, "mount_h_mm":740, "calibrated_at":"2026-09-16" }, "task": { "id":"tabletop.pick_place.mug_022", "tier":"A" }, "env": { "lighting_k":4200, "clutter":3, "friction_mu":0.42 }, "seed": "0x5F3A9C21", "reset": { "type":"manual", "seconds":42, "operator_min":2.4 }, "judge": "rc-judge-v1" }
200 OK
{
"judgment_id": "jd_5c19",
"attempt": true, // language-layer property, always returned alone
"success": false, // body-layer property, always returned alone
"failure_class": "grasp.slip",
"counts_against_policy": true, // looked up in the taxonomy, not a hard-coded list
"attempt_confidence": 0.91,
"success_confidence": 0.86,
"charged_to": { "rcsn": null, "cell_id": "cell-sf-03" },
"judge_version": "rc-judge-v1",
"table_row_id": "cf_4471"
}
Three lines held by structure, not rules
No POST /v1/attestations
Attestations are read-only in the API and issued offline: the endpoint does not exist.i
B-tier differencing: k = 5
At least 5 tasks per cell, no task_id returned; narrow queries are widened, not refused.i
Exclusion by taxonomy, not by list
20 of 54 classes don't count against the policy, and the number grows.i
Two more rules
B-tier isolation is per connection, not per query. If isolation depends on "don't SELECT those columns", one ORM SELECT * leaks a B-tier task. The Judge API connects with a database role that has no privilege at all on the four content columns, so it cannot leak even by mistake.
Unclassified is not exempt. A failure with an empty failure_class still counts against the policy. Otherwise refusing to label becomes a free way to shrink the denominator and inflate your success rate.