4.0 KiB
M5.2 — Labeler calibration
| Field | Value |
|---|---|
| Phase | M5 — Post-training |
| Size | M — 1–3 days |
| Status | ⬜ Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | M5.1 |
Goal
Measure how good the proxy is, before training a policy to imitate it.
Facts (inlined — no spec read needed)
M5.1's labels are distant supervision. Training on them without measuring agreement means the policy learns the 32B model's bias and reports it as improvement — and under a single-model deployment there is no competing variant whose divergence would make that visible.
Method: hand-label a stratified holdout, compare, and report more than accuracy.
- Sample size: 100 chunks is enough to distinguish 0.7 from 0.9 agreement.
- Stratify by the labeler's own answer — 50 it called evidence, 50 it did not. Random sampling from a corpus where ~5% is evidence gives ~5 positives, and the positive class is the one that matters.
- Report Cohen's κ, not raw agreement. With a 95/5 class balance, a labeler that always says "no" scores 95% agreement and is useless. κ corrects for chance.
- Also report precision and recall on the positive class separately. They fail differently: low recall silently starves memory, low precision pollutes it.
Disagreements are the artifact. Read them; they are usually either a genuinely ambiguous chunk or a question that is too vague — and the second is fixable at M1.2 and worth much more than a better labeler.
Steps
mem label sample --project P --n 100 --stratifiedwrites a blind worksheet — chunk text and question, no labeler answer visible.- Hand-label it. Record the human labels separately.
mem label calibratejoins them; reports agreement, κ, precision, recall, and the confusion matrix.- Dump all disagreements with both justifications side by side.
- Commit the holdout and the human labels; they are reusable for every future labeler change.
- Gate: κ ≥ 0.6 before the labels are used for training.
Acceptance
- Worksheet hides the labeler's answer.
- κ, precision, recall and the confusion matrix are all reported.
- Disagreements are dumped with justifications.
- The holdout is committed and reusable.
Verify
Harness: a synthetic labeled set with known agreement, so the statistics themselves are testable.
Integration test — tests/it_calibration.rs:
a1_worksheet_is_blind— assert the labeler's answer appears nowhere in the output file.a2_stratified— assert the sample is ~50/50 by labeler answer, not corpus proportional.a3_kappa_correct— feed a set with hand-computed κ; assert the reported value matches to 3 decimals.a4_kappa_vs_accuracy— a synthetic all-negative labeler on a 95/5 set: assert accuracy > 0.9 and κ ≈ 0. This is the assertion that justifies reporting κ at all.a5_confusion_matrix— all four cells match hand counts.a6_disagreements_dumped— count equals off-diagonal total; each carries both justifications.a7_holdout_stable— rerunning the sampler with the same seed reproduces the same chunks.
Command: cargo test -p mem-cli calibration
False pass:
- Reporting accuracy only. On this class balance it is nearly meaningless, and it will look excellent right up until the trained policy learns to always answer "no".
- Hand-labeling with the labeler's answer visible. Anchoring makes agreement look high and the whole exercise decorative — assertion 1 is a real safeguard, not hygiene.
Traps
- Sampling proportionally. A 5%-positive corpus yields five positives in a hundred, and precision on the positive class — the number that decides whether memory gets polluted — is estimated from five examples.
- Treating low κ as "the labeler needs a better prompt". Check the questions first: an ambiguous standing question makes evidence genuinely undecidable, and no labeler can fix that.
Background: DESIGN.md — Risks