Files
poimen-memory/tasks/M5.3-training-corpus-export.md
T

115 lines
4.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# M5.3 — Training corpus export
| Field | Value |
|---|---|
| Phase | M5 — Post-training |
| Size | M — 13 days |
| Status | ⬜ Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | M5.1 |
## Goal
Turn the log plus labels into trajectories verl can train on — the boundary
between the Rust side and the Python side.
## Facts (inlined — no spec read needed)
**The JSONL log is the boundary.** Rust produces it; Python consumes it. Nothing
else crosses, which is what keeps the two halves independent.
A training example is a **trajectory**, not a turn: the paper's advantage mixes a
trajectory-level term with a turn-level one (`Â = α·Â_traj + (1−α)·Â_turn`,
α=0.9), so turns must stay grouped by run.
Per turn, verl needs: the exact prompt sent, the exact response, and the rewards.
```jsonl
{"trajectory_id":"01HXYZ","turns":[
{"t":1,"prompt":"<full assembled prompt>","response":"<think>...</next>",
"r_update":-1,"parsed":true},
...],
"r_exit":-0.5,"r_format":1,"r_outcome":null}
```
Reward assembly, from the paper:
- `r_update_t` = +1 if the recorded `U_t` matches M5.1's label, 1 otherwise.
- `r_exit` — one value per trajectory: `0` if `t_exit == t_last_evidence`,
`0.75` if earlier, `0.5` if later. `t_last_evidence` is the largest `t` whose
label is true. **Note L1 runs never exit** (exit gate off), so every L1
trajectory is a "late" exit at 0.5 unless the exit signal is taken from the
recorded `E_t` rather than the loop's behaviour — take it from the record.
- `r_format` = 1 only if **every** turn in the trajectory parsed, 0 otherwise.
Strict, because a malformed turn may be caused by the previous one.
- `r_outcome` is null. We have no answer-correctness signal; the paper's
`is_equiv(A, Â)` has no analogue in extraction. Say so explicitly rather than
fabricating one.
The prompt must be the **exact bytes sent**, not re-assembled. Re-assembly drifts
from what the model actually saw the moment M1.3's template changes.
## Steps
1. `mem export --project P --format verl --out corpus/`.
2. Join log turns to labels by `chunk_sha`.
3. Reconstruct each turn's prompt from the recorded request if `MEM_LLM_RECORD`
captured it; otherwise fail loudly rather than re-assembling.
4. Compute rewards as above; carry `r_outcome: null` through.
5. Group by run into trajectories, ordered by `t`.
6. Emit a summary: trajectories, turns, positive/negative `r_update` split,
`r_format` pass rate, `t_last_evidence` distribution.
7. Refuse to export if M5.2's κ is below the threshold or absent.
## Acceptance
- Turns are grouped into trajectories, ordered.
- Prompts are byte-exact recordings, never re-assembled.
- `r_format` is 0 for a trajectory with any unparsed turn.
- Export is refused without calibration.
## Verify
**Harness:** log fixture with known labels and a deliberately unparsed turn.
**Integration test**`tests/it_export.rs`:
1. `a1_trajectory_grouping` — turns grouped by run, `t` ascending, none lost.
2. `a2_r_update_signs` — matching label → +1, mismatching → 1, checked per turn.
3. `a3_r_format_strict` — a trajectory with one unparsed turn scores 0 overall,
not per turn.
4. `a4_r_exit_from_record` — assert `r_exit` derives from the recorded `E_t`, not
from whether the loop stopped. A fixture where the gate said `end` at t=5 but
the loop continued must score as an exit at 5.
5. `a5_prompt_is_recorded_bytes` — assert the exported prompt equals the recorded
request body; corrupt the recording and assert export fails rather than
silently re-assembling.
6. `a6_r_outcome_null` — assert the field is present and null, not omitted and not
zero.
7. `a7_refuses_without_calibration` — no κ file → non-zero exit naming M5.2.
8. `a8_summary_counts` — reported splits match a hand count on the fixture.
**Command:** `cargo test -p mem-cli export`
**False pass:**
- Re-assembling prompts at export time. Every assertion except 5 passes, and the
policy is trained on prompts the model never saw — which shows up as a training
run that will not converge, with no obvious cause.
- Applying `r_format` per turn. It looks more granular and it is wrong: the paper
is explicit that the strictness exists because a bad turn may be caused by the
previous one.
- Emitting `r_outcome: 0` instead of null. Zero is a real reward value and the
trainer will use it as signal.
## Traps
- Deriving `r_exit` from loop behaviour at L1. The gate is switched off there, so
every trajectory scores as a late exit and the exit signal becomes constant
noise the policy cannot learn from.
- Joining labels by `t`. Same failure as M5.1's trap — a re-chunk misaligns
everything silently.
---
Background: [DESIGN.md](../DESIGN.md) — P6 · paper §3.2