# M5.3 — Training corpus export | Field | Value | |---|---| | Phase | M5 — Post-training | | Size | M — 1–3 days | | Status | ⬜ Not started | | Flags | — | | Spec | inlined below | | Blocks | M5.1 | ## Goal Turn the log plus labels into trajectories verl can train on — the boundary between the Rust side and the Python side. ## Facts (inlined — no spec read needed) **The JSONL log is the boundary.** Rust produces it; Python consumes it. Nothing else crosses, which is what keeps the two halves independent. A training example is a **trajectory**, not a turn: the paper's advantage mixes a trajectory-level term with a turn-level one (`Â = α·Â_traj + (1−α)·Â_turn`, α=0.9), so turns must stay grouped by run. Per turn, verl needs: the exact prompt sent, the exact response, and the rewards. ```jsonl {"trajectory_id":"01HXYZ","turns":[ {"t":1,"prompt":"","response":"...", "r_update":-1,"parsed":true}, ...], "r_exit":-0.5,"r_format":1,"r_outcome":null} ``` Reward assembly, from the paper: - `r_update_t` = +1 if the recorded `U_t` matches M5.1's label, −1 otherwise. - `r_exit` — one value per trajectory: `0` if `t_exit == t_last_evidence`, `−0.75` if earlier, `−0.5` if later. `t_last_evidence` is the largest `t` whose label is true. **Note L1 runs never exit** (exit gate off), so every L1 trajectory is a "late" exit at −0.5 unless the exit signal is taken from the recorded `E_t` rather than the loop's behaviour — take it from the record. - `r_format` = 1 only if **every** turn in the trajectory parsed, 0 otherwise. Strict, because a malformed turn may be caused by the previous one. - `r_outcome` is null. We have no answer-correctness signal; the paper's `is_equiv(A, Â)` has no analogue in extraction. Say so explicitly rather than fabricating one. The prompt must be the **exact bytes sent**, not re-assembled. Re-assembly drifts from what the model actually saw the moment M1.3's template changes. ## Steps 1. `mem export --project P --format verl --out corpus/`. 2. Join log turns to labels by `chunk_sha`. 3. Reconstruct each turn's prompt from the recorded request if `MEM_LLM_RECORD` captured it; otherwise fail loudly rather than re-assembling. 4. Compute rewards as above; carry `r_outcome: null` through. 5. Group by run into trajectories, ordered by `t`. 6. Emit a summary: trajectories, turns, positive/negative `r_update` split, `r_format` pass rate, `t_last_evidence` distribution. 7. Refuse to export if M5.2's κ is below the threshold or absent. ## Acceptance - Turns are grouped into trajectories, ordered. - Prompts are byte-exact recordings, never re-assembled. - `r_format` is 0 for a trajectory with any unparsed turn. - Export is refused without calibration. ## Verify **Harness:** log fixture with known labels and a deliberately unparsed turn. **Integration test** — `tests/it_export.rs`: 1. `a1_trajectory_grouping` — turns grouped by run, `t` ascending, none lost. 2. `a2_r_update_signs` — matching label → +1, mismatching → −1, checked per turn. 3. `a3_r_format_strict` — a trajectory with one unparsed turn scores 0 overall, not per turn. 4. `a4_r_exit_from_record` — assert `r_exit` derives from the recorded `E_t`, not from whether the loop stopped. A fixture where the gate said `end` at t=5 but the loop continued must score as an exit at 5. 5. `a5_prompt_is_recorded_bytes` — assert the exported prompt equals the recorded request body; corrupt the recording and assert export fails rather than silently re-assembling. 6. `a6_r_outcome_null` — assert the field is present and null, not omitted and not zero. 7. `a7_refuses_without_calibration` — no κ file → non-zero exit naming M5.2. 8. `a8_summary_counts` — reported splits match a hand count on the fixture. **Command:** `cargo test -p mem-cli export` **False pass:** - Re-assembling prompts at export time. Every assertion except 5 passes, and the policy is trained on prompts the model never saw — which shows up as a training run that will not converge, with no obvious cause. - Applying `r_format` per turn. It looks more granular and it is wrong: the paper is explicit that the strictness exists because a bad turn may be caused by the previous one. - Emitting `r_outcome: 0` instead of null. Zero is a real reward value and the trainer will use it as signal. ## Traps - Deriving `r_exit` from loop behaviour at L1. The gate is switched off there, so every trajectory scores as a late exit and the exit signal becomes constant noise the policy cannot learn from. - Joining labels by `t`. Same failure as M5.1's trap — a re-chunk misaligns everything silently. --- Background: [DESIGN.md](../DESIGN.md) — P6 · paper §3.2