Files
poimen-memory/tasks/M5.3-training-corpus-export.md
T

115 lines
4.7 KiB
Markdown
Raw Normal View History

2026-08-19 09:52:07 -07:00
# M5.3 — Training corpus export
| Field | Value |
|---|---|
| Phase | M5 — Post-training |
| Size | M — 13 days |
| Status | ⬜ Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | M5.1 |
## Goal
Turn the log plus labels into trajectories verl can train on — the boundary
between the Rust side and the Python side.
## Facts (inlined — no spec read needed)
**The JSONL log is the boundary.** Rust produces it; Python consumes it. Nothing
else crosses, which is what keeps the two halves independent.
A training example is a **trajectory**, not a turn: the paper's advantage mixes a
trajectory-level term with a turn-level one (`Â = α·Â_traj + (1−α)·Â_turn`,
α=0.9), so turns must stay grouped by run.
Per turn, verl needs: the exact prompt sent, the exact response, and the rewards.
```jsonl
{"trajectory_id":"01HXYZ","turns":[
{"t":1,"prompt":"<full assembled prompt>","response":"<think>...</next>",
"r_update":-1,"parsed":true},
...],
"r_exit":-0.5,"r_format":1,"r_outcome":null}
```
Reward assembly, from the paper:
- `r_update_t` = +1 if the recorded `U_t` matches M5.1's label, 1 otherwise.
- `r_exit` — one value per trajectory: `0` if `t_exit == t_last_evidence`,
`0.75` if earlier, `0.5` if later. `t_last_evidence` is the largest `t` whose
label is true. **Note L1 runs never exit** (exit gate off), so every L1
trajectory is a "late" exit at 0.5 unless the exit signal is taken from the
recorded `E_t` rather than the loop's behaviour — take it from the record.
- `r_format` = 1 only if **every** turn in the trajectory parsed, 0 otherwise.
Strict, because a malformed turn may be caused by the previous one.
- `r_outcome` is null. We have no answer-correctness signal; the paper's
`is_equiv(A, Â)` has no analogue in extraction. Say so explicitly rather than
fabricating one.
The prompt must be the **exact bytes sent**, not re-assembled. Re-assembly drifts
from what the model actually saw the moment M1.3's template changes.
## Steps
1. `mem export --project P --format verl --out corpus/`.
2. Join log turns to labels by `chunk_sha`.
3. Reconstruct each turn's prompt from the recorded request if `MEM_LLM_RECORD`
captured it; otherwise fail loudly rather than re-assembling.
4. Compute rewards as above; carry `r_outcome: null` through.
5. Group by run into trajectories, ordered by `t`.
6. Emit a summary: trajectories, turns, positive/negative `r_update` split,
`r_format` pass rate, `t_last_evidence` distribution.
7. Refuse to export if M5.2's κ is below the threshold or absent.
## Acceptance
- Turns are grouped into trajectories, ordered.
- Prompts are byte-exact recordings, never re-assembled.
- `r_format` is 0 for a trajectory with any unparsed turn.
- Export is refused without calibration.
## Verify
**Harness:** log fixture with known labels and a deliberately unparsed turn.
**Integration test**`tests/it_export.rs`:
1. `a1_trajectory_grouping` — turns grouped by run, `t` ascending, none lost.
2. `a2_r_update_signs` — matching label → +1, mismatching → 1, checked per turn.
3. `a3_r_format_strict` — a trajectory with one unparsed turn scores 0 overall,
not per turn.
4. `a4_r_exit_from_record` — assert `r_exit` derives from the recorded `E_t`, not
from whether the loop stopped. A fixture where the gate said `end` at t=5 but
the loop continued must score as an exit at 5.
5. `a5_prompt_is_recorded_bytes` — assert the exported prompt equals the recorded
request body; corrupt the recording and assert export fails rather than
silently re-assembling.
6. `a6_r_outcome_null` — assert the field is present and null, not omitted and not
zero.
7. `a7_refuses_without_calibration` — no κ file → non-zero exit naming M5.2.
8. `a8_summary_counts` — reported splits match a hand count on the fixture.
**Command:** `cargo test -p mem-cli export`
**False pass:**
- Re-assembling prompts at export time. Every assertion except 5 passes, and the
policy is trained on prompts the model never saw — which shows up as a training
run that will not converge, with no obvious cause.
- Applying `r_format` per turn. It looks more granular and it is wrong: the paper
is explicit that the strictness exists because a bad turn may be caused by the
previous one.
- Emitting `r_outcome: 0` instead of null. Zero is a real reward value and the
trainer will use it as signal.
## Traps
- Deriving `r_exit` from loop behaviour at L1. The gate is switched off there, so
every trajectory scores as a late exit and the exit signal becomes constant
noise the policy cannot learn from.
- Joining labels by `t`. Same failure as M5.1's trap — a re-chunk misaligns
everything silently.
---
Background: [DESIGN.md](../DESIGN.md) — P6 · paper §3.2