(plan) system review and break down plans
This commit is contained in:
@@ -0,0 +1,114 @@
|
||||
# M5.3 — Training corpus export
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| Phase | M5 — Post-training |
|
||||
| Size | M — 1–3 days |
|
||||
| Status | ⬜ Not started |
|
||||
| Flags | — |
|
||||
| Spec | inlined below |
|
||||
| Blocks | M5.1 |
|
||||
|
||||
## Goal
|
||||
|
||||
Turn the log plus labels into trajectories verl can train on — the boundary
|
||||
between the Rust side and the Python side.
|
||||
|
||||
## Facts (inlined — no spec read needed)
|
||||
|
||||
**The JSONL log is the boundary.** Rust produces it; Python consumes it. Nothing
|
||||
else crosses, which is what keeps the two halves independent.
|
||||
|
||||
A training example is a **trajectory**, not a turn: the paper's advantage mixes a
|
||||
trajectory-level term with a turn-level one (`Â = α·Â_traj + (1−α)·Â_turn`,
|
||||
α=0.9), so turns must stay grouped by run.
|
||||
|
||||
Per turn, verl needs: the exact prompt sent, the exact response, and the rewards.
|
||||
|
||||
```jsonl
|
||||
{"trajectory_id":"01HXYZ","turns":[
|
||||
{"t":1,"prompt":"<full assembled prompt>","response":"<think>...</next>",
|
||||
"r_update":-1,"parsed":true},
|
||||
...],
|
||||
"r_exit":-0.5,"r_format":1,"r_outcome":null}
|
||||
```
|
||||
|
||||
Reward assembly, from the paper:
|
||||
|
||||
- `r_update_t` = +1 if the recorded `U_t` matches M5.1's label, −1 otherwise.
|
||||
- `r_exit` — one value per trajectory: `0` if `t_exit == t_last_evidence`,
|
||||
`−0.75` if earlier, `−0.5` if later. `t_last_evidence` is the largest `t` whose
|
||||
label is true. **Note L1 runs never exit** (exit gate off), so every L1
|
||||
trajectory is a "late" exit at −0.5 unless the exit signal is taken from the
|
||||
recorded `E_t` rather than the loop's behaviour — take it from the record.
|
||||
- `r_format` = 1 only if **every** turn in the trajectory parsed, 0 otherwise.
|
||||
Strict, because a malformed turn may be caused by the previous one.
|
||||
- `r_outcome` is null. We have no answer-correctness signal; the paper's
|
||||
`is_equiv(A, Â)` has no analogue in extraction. Say so explicitly rather than
|
||||
fabricating one.
|
||||
|
||||
The prompt must be the **exact bytes sent**, not re-assembled. Re-assembly drifts
|
||||
from what the model actually saw the moment M1.3's template changes.
|
||||
|
||||
## Steps
|
||||
|
||||
1. `mem export --project P --format verl --out corpus/`.
|
||||
2. Join log turns to labels by `chunk_sha`.
|
||||
3. Reconstruct each turn's prompt from the recorded request if `MEM_LLM_RECORD`
|
||||
captured it; otherwise fail loudly rather than re-assembling.
|
||||
4. Compute rewards as above; carry `r_outcome: null` through.
|
||||
5. Group by run into trajectories, ordered by `t`.
|
||||
6. Emit a summary: trajectories, turns, positive/negative `r_update` split,
|
||||
`r_format` pass rate, `t_last_evidence` distribution.
|
||||
7. Refuse to export if M5.2's κ is below the threshold or absent.
|
||||
|
||||
## Acceptance
|
||||
|
||||
- Turns are grouped into trajectories, ordered.
|
||||
- Prompts are byte-exact recordings, never re-assembled.
|
||||
- `r_format` is 0 for a trajectory with any unparsed turn.
|
||||
- Export is refused without calibration.
|
||||
|
||||
## Verify
|
||||
|
||||
**Harness:** log fixture with known labels and a deliberately unparsed turn.
|
||||
|
||||
**Integration test** — `tests/it_export.rs`:
|
||||
1. `a1_trajectory_grouping` — turns grouped by run, `t` ascending, none lost.
|
||||
2. `a2_r_update_signs` — matching label → +1, mismatching → −1, checked per turn.
|
||||
3. `a3_r_format_strict` — a trajectory with one unparsed turn scores 0 overall,
|
||||
not per turn.
|
||||
4. `a4_r_exit_from_record` — assert `r_exit` derives from the recorded `E_t`, not
|
||||
from whether the loop stopped. A fixture where the gate said `end` at t=5 but
|
||||
the loop continued must score as an exit at 5.
|
||||
5. `a5_prompt_is_recorded_bytes` — assert the exported prompt equals the recorded
|
||||
request body; corrupt the recording and assert export fails rather than
|
||||
silently re-assembling.
|
||||
6. `a6_r_outcome_null` — assert the field is present and null, not omitted and not
|
||||
zero.
|
||||
7. `a7_refuses_without_calibration` — no κ file → non-zero exit naming M5.2.
|
||||
8. `a8_summary_counts` — reported splits match a hand count on the fixture.
|
||||
|
||||
**Command:** `cargo test -p mem-cli export`
|
||||
|
||||
**False pass:**
|
||||
- Re-assembling prompts at export time. Every assertion except 5 passes, and the
|
||||
policy is trained on prompts the model never saw — which shows up as a training
|
||||
run that will not converge, with no obvious cause.
|
||||
- Applying `r_format` per turn. It looks more granular and it is wrong: the paper
|
||||
is explicit that the strictness exists because a bad turn may be caused by the
|
||||
previous one.
|
||||
- Emitting `r_outcome: 0` instead of null. Zero is a real reward value and the
|
||||
trainer will use it as signal.
|
||||
|
||||
## Traps
|
||||
|
||||
- Deriving `r_exit` from loop behaviour at L1. The gate is switched off there, so
|
||||
every trajectory scores as a late exit and the exit signal becomes constant
|
||||
noise the policy cannot learn from.
|
||||
- Joining labels by `t`. Same failure as M5.1's trap — a re-chunk misaligns
|
||||
everything silently.
|
||||
|
||||
---
|
||||
|
||||
Background: [DESIGN.md](../DESIGN.md) — P6 · paper §3.2
|
||||
Reference in New Issue
Block a user