4.2 KiB
M5.1 — mem label — evidence labeler
| Field | Value |
|---|---|
| Phase | M5 — Post-training |
| Size | M — 1–3 days |
| Status | ⬜ Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | M1.6 |
Goal
Produce the per-chunk ground truth U_t that r_update needs, since this corpus
does not come with evidence labels.
Facts (inlined — no spec read needed)
Paper r_update:
r_update_t = +1 if U_t is correct
-1 if U_t is incorrect
"Correct" means: for chunks containing evidence for Q, the agent should emit
<check>yes</check>; for chunks without, <check>no</check>. That requires
knowing which chunks contain evidence.
The paper had it for free — synthetic NIAH tasks place the needle deliberately, and HotpotQA ships supporting facts. We have neither. Agent transcripts have no annotation of which turn contained the answer.
Cheapest honest substitute: distant supervision from the 32B model. Ask
reasoning (DeepSeek-R1-Distill-Qwen-32B, vLLM) per (question, chunk) whether
the chunk contains evidence. It is ~10× the controller's size and sees each chunk
independently, without the memory state that might bias the 3B's decision.
This inherits the labeler's bias, which is why M5.2 exists and must run before anyone trains on these labels.
Constraint: reasoning has a 16384 total context and vLLM rejects
input + max_tokens > 16384. A 5000-token chunk plus question plus instructions
fits with room; keep max_tokens small (labels are one token of signal) and do
not batch chunks into one prompt.
The labeler emits a binary label plus a short justification. Keep the justification — it is what makes M5.2's disagreement analysis possible.
Steps
mem label --project P --query Qreads chunks from the log.- Per chunk, prompt
reasoning: question, chunk, "does this contain evidence for the question? Answer yes or no, then one sentence why." - Send no tools — the reasoning route rejects any request carrying them.
- Write
label/<project>/<query-id>.jsonl:{"chunk_sha":"...","t":7,"label":true,"why":"...","model":"reasoning","ts":"..."}. - Resumable: skip chunks already labelled.
- Report the label rate — the fraction of chunks the labeler calls evidence. Compare it to the controller's update-rate from M1.7; a large gap is the finding, not a bug.
Acceptance
- Every chunk in the log gets exactly one label.
- Labels key on
chunk_sha, so they survive re-chunking only if content is unchanged. - Resume skips completed work.
- Label rate is reported alongside the controller's update-rate.
Verify
Harness: scripted client offline; one #[ignore] live run.
Integration test — tests/it_label.rs:
a1_one_label_per_chunk— no duplicates, no gaps against the log's chunks.a2_keyed_by_sha— labels referencechunk_sha, nott, so reordering the log does not corrupt them.a3_no_tools_sent— assert the request body has notoolskey.a4_context_budget— assert every labeling prompt is under 16384 − max_tokens.a5_resume— label, rerun, assert zero new calls.a6_justification_kept— every label has non-emptywhy.a7_rate_reported— the summary prints both label rate and the controller's update-rate.a8_live—#[ignore]; 20 real chunks throughreasoning; print the labels and justifications for a human to sanity-check.
Command: cargo test -p mem-cli label (add -- --ignored for a8)
False pass:
- Keying labels by
t. A re-chunk shifts everyt, the labels silently misalign, and the training set is quietly wrong in a way nothing downstream can detect. - Dropping the justification to save space. M5.2 then has nothing to analyse and the calibration step degenerates into a single agreement number with no way to understand it.
Traps
- Batching several chunks into one labeling prompt to save calls. The labels become order-dependent and the 16K context is exceeded on the third chunk.
- Treating the 32B's labels as ground truth. They are a proxy, and M5.2 is the task that measures how good a proxy.
Background: DESIGN.md — P6 · paper §3.2.1