# M5.1 — `mem label` — evidence labeler | Field | Value | |---|---| | Phase | M5 — Post-training | | Size | M — 1–3 days | | Status | ⬜ Not started | | Flags | — | | Spec | inlined below | | Blocks | M1.6 | ## Goal Produce the per-chunk ground truth `U_t` that `r_update` needs, since this corpus does not come with evidence labels. ## Facts (inlined — no spec read needed) Paper `r_update`: ``` r_update_t = +1 if U_t is correct -1 if U_t is incorrect ``` "Correct" means: for chunks containing evidence for `Q`, the agent should emit `yes`; for chunks without, `no`. That requires knowing which chunks contain evidence. The paper had it for free — synthetic NIAH tasks place the needle deliberately, and HotpotQA ships supporting facts. **We have neither.** Agent transcripts have no annotation of which turn contained the answer. Cheapest honest substitute: **distant supervision from the 32B model.** Ask `reasoning` (DeepSeek-R1-Distill-Qwen-32B, vLLM) per `(question, chunk)` whether the chunk contains evidence. It is ~10× the controller's size and sees each chunk independently, without the memory state that might bias the 3B's decision. This inherits the labeler's bias, which is why M5.2 exists and must run before anyone trains on these labels. Constraint: `reasoning` has a **16384 total context** and vLLM rejects `input + max_tokens > 16384`. A 5000-token chunk plus question plus instructions fits with room; keep `max_tokens` small (labels are one token of signal) and do not batch chunks into one prompt. The labeler emits a binary label plus a short justification. Keep the justification — it is what makes M5.2's disagreement analysis possible. ## Steps 1. `mem label --project P --query Q` reads chunks from the log. 2. Per chunk, prompt `reasoning`: question, chunk, "does this contain evidence for the question? Answer yes or no, then one sentence why." 3. Send **no tools** — the reasoning route rejects any request carrying them. 4. Write `label//.jsonl`: `{"chunk_sha":"...","t":7,"label":true,"why":"...","model":"reasoning","ts":"..."}`. 5. Resumable: skip chunks already labelled. 6. Report the label rate — the fraction of chunks the labeler calls evidence. Compare it to the controller's update-rate from M1.7; a large gap is the finding, not a bug. ## Acceptance - Every chunk in the log gets exactly one label. - Labels key on `chunk_sha`, so they survive re-chunking only if content is unchanged. - Resume skips completed work. - Label rate is reported alongside the controller's update-rate. ## Verify **Harness:** scripted client offline; one `#[ignore]` live run. **Integration test** — `tests/it_label.rs`: 1. `a1_one_label_per_chunk` — no duplicates, no gaps against the log's chunks. 2. `a2_keyed_by_sha` — labels reference `chunk_sha`, not `t`, so reordering the log does not corrupt them. 3. `a3_no_tools_sent` — assert the request body has no `tools` key. 4. `a4_context_budget` — assert every labeling prompt is under 16384 − max_tokens. 5. `a5_resume` — label, rerun, assert zero new calls. 6. `a6_justification_kept` — every label has non-empty `why`. 7. `a7_rate_reported` — the summary prints both label rate and the controller's update-rate. 8. `a8_live` — `#[ignore]`; 20 real chunks through `reasoning`; print the labels and justifications for a human to sanity-check. **Command:** `cargo test -p mem-cli label` (add `-- --ignored` for a8) **False pass:** - Keying labels by `t`. A re-chunk shifts every `t`, the labels silently misalign, and the training set is quietly wrong in a way nothing downstream can detect. - Dropping the justification to save space. M5.2 then has nothing to analyse and the calibration step degenerates into a single agreement number with no way to understand it. ## Traps - Batching several chunks into one labeling prompt to save calls. The labels become order-dependent and the 16K context is exceeded on the third chunk. - Treating the 32B's labels as ground truth. They are a *proxy*, and M5.2 is the task that measures how good a proxy. --- Background: [DESIGN.md](../DESIGN.md) — P6 · paper §3.2.1