Files
poimen-memory/tasks/M5.2-labeler-calibration.md
T

101 lines
4.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# M5.2 — Labeler calibration
| Field | Value |
|---|---|
| Phase | M5 — Post-training |
| Size | M — 13 days |
| Status | ⬜ Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | M5.1 |
## Goal
Measure how good the proxy is, before training a policy to imitate it.
## Facts (inlined — no spec read needed)
M5.1's labels are distant supervision. Training on them without measuring
agreement means the policy learns the 32B model's bias and reports it as
improvement — and under a single-model deployment there is no competing variant
whose divergence would make that visible.
Method: hand-label a stratified holdout, compare, and report more than accuracy.
- **Sample size**: 100 chunks is enough to distinguish 0.7 from 0.9 agreement.
- **Stratify** by the labeler's own answer — 50 it called evidence, 50 it did
not. Random sampling from a corpus where ~5% is evidence gives ~5 positives,
and the positive class is the one that matters.
- **Report Cohen's κ, not raw agreement.** With a 95/5 class balance, a labeler
that always says "no" scores 95% agreement and is useless. κ corrects for
chance.
- Also report precision and recall on the positive class separately. They fail
differently: low recall silently starves memory, low precision pollutes it.
Disagreements are the artifact. Read them; they are usually either a genuinely
ambiguous chunk or a question that is too vague — and the second is fixable at
M1.2 and worth much more than a better labeler.
## Steps
1. `mem label sample --project P --n 100 --stratified` writes a blind
worksheet — chunk text and question, **no** labeler answer visible.
2. Hand-label it. Record the human labels separately.
3. `mem label calibrate` joins them; reports agreement, κ, precision, recall,
and the confusion matrix.
4. Dump all disagreements with both justifications side by side.
5. Commit the holdout and the human labels; they are reusable for every future
labeler change.
6. Gate: κ ≥ 0.6 before the labels are used for training.
## Acceptance
- Worksheet hides the labeler's answer.
- κ, precision, recall and the confusion matrix are all reported.
- Disagreements are dumped with justifications.
- The holdout is committed and reusable.
## Verify
**Harness:** a synthetic labeled set with known agreement, so the statistics
themselves are testable.
**Integration test**`tests/it_calibration.rs`:
1. `a1_worksheet_is_blind` — assert the labeler's answer appears nowhere in the
output file.
2. `a2_stratified` — assert the sample is ~50/50 by labeler answer, not corpus
proportional.
3. `a3_kappa_correct` — feed a set with hand-computed κ; assert the reported
value matches to 3 decimals.
4. `a4_kappa_vs_accuracy` — a synthetic all-negative labeler on a 95/5 set:
assert accuracy > 0.9 **and** κ ≈ 0. This is the assertion that justifies
reporting κ at all.
5. `a5_confusion_matrix` — all four cells match hand counts.
6. `a6_disagreements_dumped` — count equals off-diagonal total; each carries both
justifications.
7. `a7_holdout_stable` — rerunning the sampler with the same seed reproduces the
same chunks.
**Command:** `cargo test -p mem-cli calibration`
**False pass:**
- Reporting accuracy only. On this class balance it is nearly meaningless, and it
will look excellent right up until the trained policy learns to always answer
"no".
- Hand-labeling with the labeler's answer visible. Anchoring makes agreement look
high and the whole exercise decorative — assertion 1 is a real safeguard, not
hygiene.
## Traps
- Sampling proportionally. A 5%-positive corpus yields five positives in a
hundred, and precision on the positive class — the number that decides whether
memory gets polluted — is estimated from five examples.
- Treating low κ as "the labeler needs a better prompt". Check the *questions*
first: an ambiguous standing question makes evidence genuinely undecidable, and
no labeler can fix that.
---
Background: [DESIGN.md](../DESIGN.md) — Risks