4.8 KiB
4.8 KiB
T6.2 — Judge calibration set
| Field | Value |
|---|---|
| Phase | P6 — Learning loop |
| Size | M — 1 to 3 days |
| Status | Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | T6.3. Promotion without calibration is the self-critique failure mode |
Goal
Replay a fixed set of tasks with known verifier ground truth through the judge on a schedule, and track judge-versus-verifier agreement.
Facts (inlined — no spec read needed)
- This is a precondition, not polish. Under one resident model the judge is the agent's model. Self-critique using the same model family being graded is admissible only with this bootstrap: validation against verifiable tasks.
- Without it the grader's bias is unmeasured and the optimizer will find it. And under one-state convergence (T6.1) there is no competing variant whose divergence would make that visible.
- Judge-versus-verifier agreement on the calibration set is the only
instrument that can see the grader drifting toward its own bias. It belongs
in the metrics list beside
Indeterminatecount and judge order-inconsistency. - It costs zero agent runs, since the episodes are already recorded. Re-running the episodes is the obvious wrong implementation, which is why the run count is asserted.
- Verification decides, grading explains. Agreement is measured against the verifier's ground truth, never the other way round.
Steps
- Define the calibration set: a stable list of
(TaskId, episode, known verifier outcome)drawn from recorded history. Store it explicitly — a set that is re-derived each cycle is not a fixed set. - Schedule the calibration pass on an interval, decoupled from promotion.
- For each entry, call the judge on the recorded episodes. No agent execution, no replay, no new runs.
- Compute agreement overall and per rubric dimension, so a judge biased on one dimension is caught with the dimension named rather than showing as a small overall dip.
- Emit agreement as a metric tagged by judge, model and rubric version.
- Instrument an agent-run counter across the calibration pass and assert it is zero.
- Build two judge stubs: one agreeing 95% of the time, one deliberately biased against a chosen rubric dimension.
Acceptance
- A judge stub agreeing with the verifier 95% of the time reports ~95%.
- A judge deliberately biased against a rubric dimension is caught and the dimension named.
- Calibration consumes zero agent runs — asserted by count, since re-running episodes is the obvious wrong implementation.
Verify
Harness: two judge stubs — one agreeing with the verifier 95% of the time, one biased against a single named rubric dimension. Plus an agent-run counter at the executor boundary.
Integration test — tests/it_judge_calibration.rs:
- Run the calibration pass with the 95% stub over ≥200 calibration entries. Assert reported agreement is ~95% within sampling error. State the tolerance from the sample size rather than guessing.
- Run with the dimension-biased stub. Assert the report names the dimension, not just a lower overall figure.
- Assert the agent-run counter is exactly 0 across the whole pass. The episodes are already recorded; re-running them is the obvious wrong implementation and it is expensive rather than incorrect-looking.
- Assert the calibration set is stored and stable — run the pass twice and assert the same entries were used both times. A set re-derived each cycle is not a fixed set and its trend is meaningless.
- Assert agreement is emitted as a metric tagged by judge, model and rubric version.
- Assert T6.3 refuses to promote when no calibration result exists or the latest is stale — the blocking relationship must be enforced, not documented.
Command: cargo test -p loop calibration
False pass:
- Step 2 asserting only that overall agreement dropped. A judge biased on one dimension of five moves the aggregate by a few points, which is indistinguishable from noise — the per-dimension breakdown is the instrument.
- Step 4 omitted: a set rebuilt from "recent episodes" each cycle drifts with the workload, so the metric measures the workload rather than the judge.
- Step 6 omitted, leaving calibration as a dashboard number that gates nothing — which is precisely the self-critique failure mode this exists to prevent.
Traps
- Re-executing the calibration tasks to "get fresh episodes". It costs agent runs and measures the wrong thing.
- Reporting only an aggregate agreement number, which hides single-dimension bias — the exact failure the loop will exploit.
Background (not required to do this task): rust-agentic-sys.md §1, §14.2, §15, §16 · rust-agentic-task.md