Files

101 lines
4.0 KiB
Markdown
Raw Permalink Normal View History

2026-08-19 09:52:07 -07:00
# M5.2 — Labeler calibration
| Field | Value |
|---|---|
| Phase | M5 — Post-training |
| Size | M — 13 days |
| Status | ⬜ Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | M5.1 |
## Goal
Measure how good the proxy is, before training a policy to imitate it.
## Facts (inlined — no spec read needed)
M5.1's labels are distant supervision. Training on them without measuring
agreement means the policy learns the 32B model's bias and reports it as
improvement — and under a single-model deployment there is no competing variant
whose divergence would make that visible.
Method: hand-label a stratified holdout, compare, and report more than accuracy.
- **Sample size**: 100 chunks is enough to distinguish 0.7 from 0.9 agreement.
- **Stratify** by the labeler's own answer — 50 it called evidence, 50 it did
not. Random sampling from a corpus where ~5% is evidence gives ~5 positives,
and the positive class is the one that matters.
- **Report Cohen's κ, not raw agreement.** With a 95/5 class balance, a labeler
that always says "no" scores 95% agreement and is useless. κ corrects for
chance.
- Also report precision and recall on the positive class separately. They fail
differently: low recall silently starves memory, low precision pollutes it.
Disagreements are the artifact. Read them; they are usually either a genuinely
ambiguous chunk or a question that is too vague — and the second is fixable at
M1.2 and worth much more than a better labeler.
## Steps
1. `mem label sample --project P --n 100 --stratified` writes a blind
worksheet — chunk text and question, **no** labeler answer visible.
2. Hand-label it. Record the human labels separately.
3. `mem label calibrate` joins them; reports agreement, κ, precision, recall,
and the confusion matrix.
4. Dump all disagreements with both justifications side by side.
5. Commit the holdout and the human labels; they are reusable for every future
labeler change.
6. Gate: κ ≥ 0.6 before the labels are used for training.
## Acceptance
- Worksheet hides the labeler's answer.
- κ, precision, recall and the confusion matrix are all reported.
- Disagreements are dumped with justifications.
- The holdout is committed and reusable.
## Verify
**Harness:** a synthetic labeled set with known agreement, so the statistics
themselves are testable.
**Integration test**`tests/it_calibration.rs`:
1. `a1_worksheet_is_blind` — assert the labeler's answer appears nowhere in the
output file.
2. `a2_stratified` — assert the sample is ~50/50 by labeler answer, not corpus
proportional.
3. `a3_kappa_correct` — feed a set with hand-computed κ; assert the reported
value matches to 3 decimals.
4. `a4_kappa_vs_accuracy` — a synthetic all-negative labeler on a 95/5 set:
assert accuracy > 0.9 **and** κ ≈ 0. This is the assertion that justifies
reporting κ at all.
5. `a5_confusion_matrix` — all four cells match hand counts.
6. `a6_disagreements_dumped` — count equals off-diagonal total; each carries both
justifications.
7. `a7_holdout_stable` — rerunning the sampler with the same seed reproduces the
same chunks.
**Command:** `cargo test -p mem-cli calibration`
**False pass:**
- Reporting accuracy only. On this class balance it is nearly meaningless, and it
will look excellent right up until the trained policy learns to always answer
"no".
- Hand-labeling with the labeler's answer visible. Anchoring makes agreement look
high and the whole exercise decorative — assertion 1 is a real safeguard, not
hygiene.
## Traps
- Sampling proportionally. A 5%-positive corpus yields five positives in a
hundred, and precision on the positive class — the number that decides whether
memory gets polluted — is estimated from five examples.
- Treating low κ as "the labeler needs a better prompt". Check the *questions*
first: an ambiguous standing question makes evidence genuinely undecidable, and
no labeler can fix that.
---
Background: [DESIGN.md](../DESIGN.md) — Risks