feat(M5.1-M5.2): Add evidence labeler and calibration infrastructure
M5.1 — Evidence Labeler (distant supervision):
- EvidenceLabel struct: chunk_sha, t, label, why, model, ts
- LabelerConfig: configurable model_id, max_tokens, max_context
- make_label_prompt(): question + chunk in 16K context budget
- parse_label_response(): extract yes/no + 1-sentence justification
- fits_context_budget(): verify prompt fits reasoning model limits
- Unit tests: 8/8 passing
M5.2 — Labeler Calibration (Cohen's kappa):
- CalibrationResults: tp/tn/fp/fn, accuracy, kappa, precision, recall, f1
- Cohen's kappa formula (corrects for class imbalance, unlike accuracy)
- CalibrationSample: blind worksheet (hides labeler answers from human)
- stratified_sample(): 50/50 positive/negative (not corpus-proportional)
- passes_gate(): kappa >= 0.6 threshold
- Unit tests: 6/6 passing
Integration tests:
tests/it_labeler.rs: 11 tests, all passing
- a1: One label per chunk
- a2: Keyed by sha (survives re-chunking)
- a3: Context budget respected
- a4: Justifications preserved
- a5: Label structure correct
- a6: No tools in prompt (reasoning model requirement)
- a7: Parse variations (YES/no/Yes/No)
- a8-a11: Serialization, rate reporting, edge cases
tests/it_calibration.rs: 12 tests, all passing
- a1: Worksheet blind (labeler answers hidden)
- a2: Stratified sampling (attempts 50/50)
- a3: Kappa perfect agreement = 1.0
- a4: Kappa vs accuracy (high accuracy ≠ good kappa)
- a5: Confusion matrix (all 4 cells tracked)
- a6: Precision/recall separated
- a7: Gate threshold kappa >= 0.6
- a8: F1 score computed
- a9-a12: Roundtrips, disagreement analysis, formula validation
Files created:
crates/mem-llm/src/labeler.rs (250 LOC)
crates/mem-llm/src/calibration.rs (280 LOC)
tests/it_labeler.rs (200 LOC)
tests/it_calibration.rs (300 LOC)
Architecture:
M5.1: Question + Chunk → Reasoning Model → Label + Why
M5.2: Labeler Labels + Human Labels → Kappa + Confusion Matrix → Gate
Blocks: M5.3 (corpus export)
Depends: M4.3 ✓