SIMSimulated data. Every score, rate and price on this site is illustrative.
Instrument

Label an episode标注一个片段

Two operators label the same run blind, and their agreement (κ) is the ceiling any judge is read against.两名操作员对同一个 run 双人盲标,他们之间的一致率 κ 是读任何评判模型时的上限。

No ceiling exists yet: 900 blind labels sit on 900 different runs, so no run has two.上限目前还不存在:900 条盲标落在 900 个不同的 run 上,没有一个 run 有两条标签。 i

Episode片段

STAND-IN
0.00 s

Rubric评分表

on this item本条用时 0 s

1 · Attempt: did the actor try the task?1 · 尝试:执行者有没有尝试这个任务? i
2 · Success: required state at the last frame?2 · 成功:最后一帧世界是否处于要求的状态? i
3 · Failure class (only when success is no)3 · 失败类别(仅当未成功) i

Keyboard快捷键

A/Z attempt尝试 · S/X success成功 · 1–9 failure root失败根类 · Space play播放 · ←→ scrub拖动. Keys are ignored while you type in a field.在输入框里打字时快捷键不生效。

Agreement: human A vs human B一致率 κ:人 A 对人 B

ILLUSTRATIVE Attempt and success get separate κ values, always A-vs-B, never blended.尝试和成功各算一个 κ,永远是 A 对 B,从不合并。 i

Cohen's κ, as computed on this page本页计算 Cohen's κ 的公式

po = (items where A = B) / n
pe = Σc (A's share of c) × (B's share of c)
κ = (po − pe) / (1 − pe)  ·  undefined when pe = 1

The same formula as与以下实现同一公式: cohen_kappa() in rc-instrument-app/judge/evals/judge_bench.py · v_human_ceiling in rc-instrument-app/supabase/0003_judge_loop.sql. This page recomputes it in JS from the 40 demo pairs; nothing is hard-coded.本页用 JS 从 40 对演示数据现算,没有写死任何结果。

Why a judge's κ means nothing yet为什么评判模型的 κ 目前毫无意义

A judge is read against human-vs-human agreement, and no run has two human labels yet.评判模型要对照人与人之间的一致率来读,而目前没有一个 run 有两条人工标签。

Show why, and what labelling costs展开原因与标注成本

No ceiling, no scale没有上限,就没有刻度

κ = 0.82 against operators could be near the limit or mediocre; only human-vs-human agreement tells you which.对操作员 κ = 0.82 可能接近极限,也可能很平庸;只有人与人之间的一致率能说明是哪种。

More展开

Today v_human_ceiling returns n = 0: the 900 blind labels sit on 900 distinct runs. Promotion gate G7 refuses any candidate unless human ceiling − incumbent κ ≥ 0.03, so a calibration set bought without double-labelling can't promote a judge. They have to be bought together.今天 v_human_ceiling 返回 n = 0:900 条盲标落在 900 个不同的 run 上。晋级门 G7 要求 人类上限 − 在任 κ ≥ 0.03,否则拒绝任何候选;所以不带双人盲标买来的校准集无法让评判模型晋级。两者必须一起买。

Source:来源: judge/LIMITS.md §4, judge/LABELLING-ASK.md.

The queue can't fix it by itself队列自己修不好

rci_review_next() only serves runs with no label yet, so it never hands a second operator a labelled run.rci_review_next() 只发还没有标签的 run,所以永远不会把已标过的 run 交给第二个人。

More展开

The clause is where not exists (… operator_labels o where o.run_id = r.id). By construction that is how you end up with 900 labels on 900 runs. The assignment at the top of this page (you are B, A is hidden) needs a queue change first.条件是 where not exists (… operator_labels o where o.run_id = r.id)。按这个构造,900 条标签就只能落在 900 个 run 上。本页顶部的分配方式(你是 B,A 隐藏)需要先改队列。

Source:来源: rc-instrument-app/supabase/0008_blind_review.sql.

What labelling costs标注成本

A17 h
B52 h
C89 h

Planning figure: 6 operator-minutes per item, 8 per adjudication.规划数:每条 6 操作员分钟,每次仲裁 8 分钟。 i

Judge cost, and confident-wrong cases评判成本,与高置信的错判

$0.0697 per judgment on gpt-4o; the cheap tier called a failed run a success at 0.98 confidence.gpt-4o 每次评判 $0.0697;便宜档把一个失败的 run 以 0.98 的置信度判成了成功。

Show the cost basis and the three cases展开成本依据与三个案例

Judge cost basis评判成本基准 MEASURED实测

$0.0697

per judgment on gpt-4o, pinned ensemble n = 3 (6 model calls).每次评判,gpt-4o,固定集成 n = 3(6 次模型调用)。

Basis and caveat依据与注意

6 model calls per judgment: attempt and outcome × 3. Measured from each response's usage: 36,416 input tokens (24,624 of them prompt-cache hits) and 942 output tokens per judgment.每次评判 6 次调用:尝试与结果各 × 3。按每次响应的 usage 实测:每次评判 36,416 个输入 token(其中 24,624 个命中提示缓存)、942 个输出 token。

Success agreement 7/8 on 8 convenience-sampled pilot clips whose labels were produced by an AI agent, not a human. That is not a κ and can't be read as accuracy.在 8 个方便抽样的试点片段上成功一致 7/8,这些片段的标签是 AI 智能体产出的,不是人。这不是 κ,也不能当准确率读。

Source:来源: rc-instrument-app/judge/LIMITS.md (cost table, "What this means in practice").

The cheap tier fails in the dangerous direction便宜档错在危险的方向 MEASURED实测

Episode片段Truth真值gpt-4.1-miniConf.置信度
insert straw into drinkfailed失败succeeded成功0.98
take SD card from pouchfailed失败succeeded成功0.93
take pillfailed失败succeeded成功0.62

The straw never goes in, and the model said success at 0.98: a confident wrong answer ships silently.吸管根本没插进去,模型却以 0.98 判成功:高置信的错判会悄无声息地发出去。

Why the rubric is strict为什么评分表这么严

A wrong answer at high confidence inflates every published success rate. That's why the rubric asks for success strictly against the last frame, and why every disagreement goes to a third reviewer, who sees both labels but never the judge's output.高置信的错判会抬高每一个公开的成功率。所以评分表严格按最后一帧判成功,所有分歧都交给第三位复核人,他看得到两条标签,但永远看不到评判模型的输出。

Source:来源: rc-instrument-app/judge/LIMITS.md, "The cheap tier is not good enough". "Truth" there is the AI-agent pilot label on n = 8.那里的「Truth」是 AI 智能体在 n = 8 上的试点标签。

RC Shadow · labeling demo.RC Shadow · 标注演示。 Schema: judge/evals/golden/manifest.schema.json · classes: docs/rc-instrument/failure-taxonomy.yaml (v1). Nothing on this page sends data anywhere.本页不会发出任何数据。