SIM模拟数据。本站所有分数、比率与价格均为示意。
Instrument
EN 联系

Label an episode标注一个片段

Two operators label the same run blind, and their agreement (κ) is the ceiling any judge is read against.两名操作员对同一个 run 双人盲标,他们之间的一致率 κ 是读任何评判模型时的上限。

No ceiling exists yet: 900 blind labels sit on 900 different runs, so no run has two.上限目前还不存在:900 条盲标落在 900 个不同的 run 上,没有一个 run 有两条标签。 i

Episode片段

STAND-IN
0.00 s

Rubric评分表

on this item本条用时 0 s

1 · Attempt: did the actor try the task?1 · 尝试:执行者有没有尝试这个任务? i
2 · Success: required state at the last frame?2 · 成功:最后一帧世界是否处于要求的状态? i
3 · Failure class (only when success is no)3 · 失败类别(仅当未成功) i

Keyboard快捷键

A/Z attempt尝试 · S/X success成功 · 1–9 failure root失败根类 · Space play播放 · ←→ scrub拖动. Keys are ignored while you type in a field.在输入框里打字时快捷键不生效。

Agreement: human A vs human B一致率 κ:人 A 对人 B

ILLUSTRATIVE Attempt and success get separate κ values, always A-vs-B, never blended.尝试和成功各算一个 κ,永远是 A 对 B,从不合并。 i

Cohen's κ, as computed on this page本页计算 Cohen's κ 的公式

po = (items where A = B) / n
pe = Σc (A's share of c) × (B's share of c)
κ = (po − pe) / (1 − pe)  ·  undefined when pe = 1

The same formula as与以下实现同一公式: cohen_kappa() in rc-instrument-app/judge/evals/judge_bench.py · v_human_ceiling in rc-instrument-app/supabase/0003_judge_loop.sql. This page recomputes it in JS from the 40 demo pairs; nothing is hard-coded.本页用 JS 从 40 对演示数据现算,没有写死任何结果。

Why a judge's κ means nothing yet为什么评判模型的 κ 目前毫无意义

A judge is read against human-vs-human agreement, and no run has two human labels yet.评判模型要对照人与人之间的一致率来读,而目前没有一个 run 有两条人工标签。

Show why, and what labelling costs展开原因与标注成本

No ceiling, no scale没有上限,就没有刻度

κ = 0.82 against operators could be near the limit or mediocre; only human-vs-human agreement tells you which.对操作员 κ = 0.82 可能接近极限,也可能很平庸;只有人与人之间的一致率能说明是哪种。

More展开

Today v_human_ceiling returns n = 0: the 900 blind labels sit on 900 distinct runs. Promotion gate G7 refuses any candidate unless human ceiling − incumbent κ ≥ 0.03, so a calibration set bought without double-labelling can't promote a judge. They have to be bought together.今天 v_human_ceiling 返回 n = 0:900 条盲标落在 900 个不同的 run 上。晋级门 G7 要求 人类上限 − 在任 κ ≥ 0.03,否则拒绝任何候选;所以不带双人盲标买来的校准集无法让评判模型晋级。两者必须一起买。

Source:来源: judge/LIMITS.md §4, judge/LABELLING-ASK.md.

The queue can't fix it by itself队列自己修不好

rci_review_next() only serves runs with no label yet, so it never hands a second operator a labelled run.rci_review_next() 只发还没有标签的 run,所以永远不会把已标过的 run 交给第二个人。

More展开

The clause is where not exists (… operator_labels o where o.run_id = r.id). By construction that is how you end up with 900 labels on 900 runs. The assignment at the top of this page (you are B, A is hidden) needs a queue change first.条件是 where not exists (… operator_labels o where o.run_id = r.id)。按这个构造,900 条标签就只能落在 900 个 run 上。本页顶部的分配方式(你是 B,A 隐藏)需要先改队列。

Source:来源: rc-instrument-app/supabase/0008_blind_review.sql.

What labelling costs标注成本

A17 h
B52 h
C89 h

Planning figure: 6 operator-minutes per item, 8 per adjudication.规划数:每条 6 操作员分钟,每次仲裁 8 分钟。 i

Judge cost, and confident-wrong cases评判成本,与高置信的错判

$0.0697 per judgment on gpt-4o; the cheap tier called a failed run a success at 0.98 confidence.gpt-4o 每次评判 $0.0697;便宜档把一个失败的 run 以 0.98 的置信度判成了成功。

Show the cost basis and the three cases展开成本依据与三个案例

Judge cost basis评判成本基准 MEASURED实测

$0.0697

per judgment on gpt-4o, pinned ensemble n = 3 (6 model calls).每次评判,gpt-4o,固定集成 n = 3(6 次模型调用)。

Basis and caveat依据与注意

6 model calls per judgment: attempt and outcome × 3. Measured from each response's usage: 36,416 input tokens (24,624 of them prompt-cache hits) and 942 output tokens per judgment.每次评判 6 次调用:尝试与结果各 × 3。按每次响应的 usage 实测:每次评判 36,416 个输入 token(其中 24,624 个命中提示缓存)、942 个输出 token。

Success agreement 7/8 on 8 convenience-sampled pilot clips whose labels were produced by an AI agent, not a human. That is not a κ and can't be read as accuracy.在 8 个方便抽样的试点片段上成功一致 7/8,这些片段的标签是 AI 智能体产出的,不是人。这不是 κ,也不能当准确率读。

Source:来源: rc-instrument-app/judge/LIMITS.md (cost table, "What this means in practice").

The cheap tier fails in the dangerous direction便宜档错在危险的方向 MEASURED实测

Episode片段Truth真值gpt-4.1-miniConf.置信度
insert straw into drinkfailed失败succeeded成功0.98
take SD card from pouchfailed失败succeeded成功0.93
take pillfailed失败succeeded成功0.62

The straw never goes in, and the model said success at 0.98: a confident wrong answer ships silently.吸管根本没插进去,模型却以 0.98 判成功:高置信的错判会悄无声息地发出去。

Why the rubric is strict为什么评分表这么严

A wrong answer at high confidence inflates every published success rate. That's why the rubric asks for success strictly against the last frame, and why every disagreement goes to a third reviewer, who sees both labels but never the judge's output.高置信的错判会抬高每一个公开的成功率。所以评分表严格按最后一帧判成功,所有分歧都交给第三位复核人,他看得到两条标签,但永远看不到评判模型的输出。

Source:来源: rc-instrument-app/judge/LIMITS.md, "The cheap tier is not good enough". "Truth" there is the AI-agent pilot label on n = 8.那里的「Truth」是 AI 智能体在 n = 8 上的试点标签。

RC Shadow · labeling demo.RC Shadow · 标注演示。 Schema: judge/evals/golden/manifest.schema.json · classes: docs/rc-instrument/failure-taxonomy.yaml (v1). Nothing on this page sends data anywhere.本页不会发出任何数据。