SIMSimulated data. Every score, rate and price on this site is illustrative.
Instrument

Docs

Every reference document in one place, for anyone who needs the detail behind the site.

Spec28 files

Schema, API, task format, rights terms and pricing, as written.

Below ›
SchemaLive database map

Every table, field and foreign key, read from the deployed database.

Open ›
Shadow contractPROTOCOL.md

What Shadow mode records, and the nine rules it follows.

Below ›
LedgerTie-outs

Six identities that must always hold across the tables.

Open ›
EcosystemOther benchmarks

What we adopt from other benchmarks and datasets, and where the gap is.

Open ›
PlanPlan and status

What is real, what is demo, and who does what next.

Open ›

28 files, each with one job

13,956 lines across 28 files in docs/rc-instrument/.i

FileRole
architecture.mdSix-layer stack, L0 physical → L5 products.
schema.sql · data-model.mdCanonical PostgreSQL schema and its reader's guide.
failure-taxonomy.yamlControlled failure vocabulary; the only source for the failure_classes table.
judge-api.yamlOpenAPI spec for the Judge API. Not yet deployed.
task-protocol.md · tasks/Open YAML task format: what is published, what is withheld, how private tasks are stored.
run-bundle.mdOn-disk layout and hashing of one immutable episode.
data-rights-and-tos.mdDraft clause text for counsel. Not legal advice.
orders-and-pricing.md · unit-economics.pyPricing proposal; the script generates every margin figure.
reports.mdThe six customer-facing deliverables.
validate_*.py · mutation_test_*.py · derive_ambiguities.pyChecks, proof the checks can fail, and generated lists.
All files, with line counts
FileLinesType
Size by file and by type

Spec size by file · 13,956 lines

By type

One file defines what Shadow mode records

One file, PROTOCOL.md, defines what Shadow mode captures, infers and shows.i Status and owners are on Plan and status.

Ver.Rule
v1Three streams (obs / pred / div) stored apart, joined on t_ns; the policy need not be present at capture. i
v2Record once, run every model offline: a policy × body matrix. A do-nothing baseline is required; divergence is not a capability ranking.
v3Obs must carry object and target pose; decision-tick density gates task admission. i
v4n ≥ 40 makes one number printable, not two comparable. Closed-loop pred records both policy and commanded action; the envelope is part of cell identity.
v5A statistic from capped data is a bound, not a measurement. Of two caps, the tighter one wins silently.
v6Rotation deltas must declare their reference frame (body vs world).
v7A ratio can be produced entirely by normalization. i
v8Every library ships a zero-difference control. i
v9The tool point must be declared, not guessed from link order in the file.

Two decisions are still open

Both need someone to decide, not more analysis.

  1. 01Who carries the cell body's depreciation. Hardware showroom stock: a $500 week earns 48.8% margin. Eval-owned: the same week earns 3.2%. ILLUSTRATIVE
  2. 02Who physically holds B-tier task content. Lock box, manifest and access log need one named person. i

The schema has never run against a live database

It parses cleanly, and the first milestone (M0) needs none of it.

Parse-clean is not run-clean
At delivery, schema.sql parsed with pglast (176 statements) but had never run against a live database, and judge-api.yaml had no OpenAPI 3.1 conformance run.i
M0 needs none of these 13,956 lines

3 models × 4 bodies × 40 = 480 runs, paper trial sheets, human scoring, eval_real_robot.py for aggregation: about 2 engineer-weeks + 3 operator-days. The spec exists so nothing has to be redone after the first real result, not to gate it.

First-deploy acceptance test

The only proof that the B tier is isolated at the connection level. Until it passes, do not state that as verified.

1  psql -f schema.sql
2  SELECT rc_refresh_metrics(false);   ← false: CONCURRENTLY cannot fill an empty matview
3  SELECT * FROM v_schema_invariants;  ← expect 0 on every row
4  # connect as the API's own role; this must fail:
   SELECT * FROM tasks;              → permission denied for table tasks
Never connect the Judge API as the table owner. The owner bypasses column grants and RLS, so the protection is gone while every test stays green.i
Derived, not typed

No number on this page is hand-typed: spec counts come from disk, taxonomy counts from the live rci_taxonomy. A generator, derive_ambiguities.py, replaced the last hand-written list and runs --check in CI.

Why: the same error happened four times. Hard-coding 2 of the 20 excluded classes; evidence downgrades flipping a denominator; a hand-written ambiguity list disagreeing with the view's ordering; and this page once calling 17 (the leaf count) the total of 20. Each time a hand-written copy drifted from a computed one. The generator caught drift on its first run.