Files
poimen-memory/tasks/M5.2-labeler-calibration.md

4.0 KiB
Raw Permalink Blame History

M5.2 — Labeler calibration

Field Value
Phase M5 — Post-training
Size M — 13 days
Status Not started
Flags
Spec inlined below
Blocks M5.1

Goal

Measure how good the proxy is, before training a policy to imitate it.

Facts (inlined — no spec read needed)

M5.1's labels are distant supervision. Training on them without measuring agreement means the policy learns the 32B model's bias and reports it as improvement — and under a single-model deployment there is no competing variant whose divergence would make that visible.

Method: hand-label a stratified holdout, compare, and report more than accuracy.

  • Sample size: 100 chunks is enough to distinguish 0.7 from 0.9 agreement.
  • Stratify by the labeler's own answer — 50 it called evidence, 50 it did not. Random sampling from a corpus where ~5% is evidence gives ~5 positives, and the positive class is the one that matters.
  • Report Cohen's κ, not raw agreement. With a 95/5 class balance, a labeler that always says "no" scores 95% agreement and is useless. κ corrects for chance.
  • Also report precision and recall on the positive class separately. They fail differently: low recall silently starves memory, low precision pollutes it.

Disagreements are the artifact. Read them; they are usually either a genuinely ambiguous chunk or a question that is too vague — and the second is fixable at M1.2 and worth much more than a better labeler.

Steps

  1. mem label sample --project P --n 100 --stratified writes a blind worksheet — chunk text and question, no labeler answer visible.
  2. Hand-label it. Record the human labels separately.
  3. mem label calibrate joins them; reports agreement, κ, precision, recall, and the confusion matrix.
  4. Dump all disagreements with both justifications side by side.
  5. Commit the holdout and the human labels; they are reusable for every future labeler change.
  6. Gate: κ ≥ 0.6 before the labels are used for training.

Acceptance

  • Worksheet hides the labeler's answer.
  • κ, precision, recall and the confusion matrix are all reported.
  • Disagreements are dumped with justifications.
  • The holdout is committed and reusable.

Verify

Harness: a synthetic labeled set with known agreement, so the statistics themselves are testable.

Integration testtests/it_calibration.rs:

  1. a1_worksheet_is_blind — assert the labeler's answer appears nowhere in the output file.
  2. a2_stratified — assert the sample is ~50/50 by labeler answer, not corpus proportional.
  3. a3_kappa_correct — feed a set with hand-computed κ; assert the reported value matches to 3 decimals.
  4. a4_kappa_vs_accuracy — a synthetic all-negative labeler on a 95/5 set: assert accuracy > 0.9 and κ ≈ 0. This is the assertion that justifies reporting κ at all.
  5. a5_confusion_matrix — all four cells match hand counts.
  6. a6_disagreements_dumped — count equals off-diagonal total; each carries both justifications.
  7. a7_holdout_stable — rerunning the sampler with the same seed reproduces the same chunks.

Command: cargo test -p mem-cli calibration

False pass:

  • Reporting accuracy only. On this class balance it is nearly meaningless, and it will look excellent right up until the trained policy learns to always answer "no".
  • Hand-labeling with the labeler's answer visible. Anchoring makes agreement look high and the whole exercise decorative — assertion 1 is a real safeguard, not hygiene.

Traps

  • Sampling proportionally. A 5%-positive corpus yields five positives in a hundred, and precision on the positive class — the number that decides whether memory gets polluted — is estimated from five examples.
  • Treating low κ as "the labeler needs a better prompt". Check the questions first: an ambiguous standing question makes evidence genuinely undecidable, and no labeler can fix that.

Background: DESIGN.md — Risks