4.4 KiB
M5.6 — M5 composition gate
| Field | Value |
|---|---|
| Phase | M5 — Post-training |
| Size | L — 3+ days |
| Status | ⬜ Not started |
| Flags | gate |
| Spec | inlined below |
| Blocks | all of M5 |
Goal
Establish that the trained adapter is actually better than prompting — and that "better" was measured on data it never saw.
Facts (inlined — no spec read needed)
The comparison is adapter vs. the prompted baseline, on a held-out project, using the same corpus and the same prompt.
The paper's own result sets the expectation: Figure 9 shows RL helps but the prompted workflow already works, and the gains concentrate on harder tasks. So a modest improvement is the realistic success case; a dramatic one is a reason to check for leakage first.
Metrics, and why each is present:
| metric | why |
|---|---|
| update accuracy, evidence-present | recall — does it catch evidence |
| update accuracy, evidence-free | precision — does it resist noise. This is the one that collapses at α=1 |
| exact-exit ratio | did the exit gate learn anything |
| format correctness | the parser's job got easier or harder |
| memory token curve | the Figure 6 shape — flat, or climbing to the cap |
| wall clock per run | the paper claims up to 400% speedup with the exit gate |
Held-out project is non-negotiable. Same-project trajectories share file paths, error strings and phrasing; measuring on them is measuring memorisation.
Report both, always: an adapter that improves evidence-present accuracy while degrading evidence-free is worse for this system, because polluted memory degrades every subsequent turn.
Steps
- Pick a project excluded from training. Ingest it with the prompted baseline; record all metrics.
- Ingest the same project with the adapter, identical prompt and chunking.
- Compare on the M5.2 holdout labels.
- Assert the thresholds below.
- Plot the memory-token curve for both; commit it as the Figure 6 analogue.
- Commit
expected/m5.6.txt; diff. - If the adapter loses, keep it versioned and record why — a negative result with the reason is worth more than a rerun with different hyperparameters.
Acceptance
- Evidence-present accuracy ≥ baseline.
- Evidence-free accuracy ≥ baseline. Never traded away.
- Format correctness ≥ baseline.
- Memory token curve flat, not climbing to the cap.
- Measured on a project absent from training.
Verify
Harness: live gateway with both models; the held-out project; M5.2's labels.
Integration test — verify/m5.6.sh diffed against expected/m5.6.txt:
a1_holdout_is_unseen— assert the eval project appears in no training trajectory. Check the corpus, not the config.a2_evidence_present_accuracy— adapter ≥ baseline; print both.a3_evidence_free_accuracy— adapter ≥ baseline; print both.a4_format_correctness— adapter ≥ baseline.a5_memory_curve_flat— slope below the M1.8 bound for both; assert the adapter is no worse.a6_exit_ratio_reported— print exact/early/late exit ratios. Advisory: the exit gate is off at L1, so this measures the signal, not behaviour.a7_same_prompt— assert both runs used a byte-identical prompt template.a8_wall_clock— report both; no threshold, since the exit gate is off at L1.
Command: bash verify/m5.6.sh | diff - expected/m5.6.txt
False pass:
- Comparing on a project that was in training. Everything improves and none of it generalises — assertion 1 checks the corpus rather than trusting the split config.
- Reporting a single combined accuracy. It hides the precision/recall trade that matters most: a model that says "yes" more often scores better on evidence-present and pollutes memory, and the combined number can improve while the system gets worse.
- Different prompts between runs. Any measured difference then attributes to the adapter and is partly the prompt — assertion 7 is cheap and removes the doubt.
Traps
- Retraining until the gate passes without changing anything principled. That is fitting the gate, and the held-out project stops being held out the third time you look at it.
- Reading a small improvement as failure. The paper's own baseline works; the adapter's value here is as much about stability on evidence-free chunks as headline accuracy.
Background: DESIGN.md — Verification, Risks · paper Fig 9