# M3.7.6 — M3.7 composition gate | Field | Value | |---|---| | Phase | M3.7 — Tool context | | Size | M — 1–3 days | | Status | ⬜ Not started | | Flags | gate | | Spec | inlined below | | Blocks | all of M3.7 | | Depends | M3.7.3, M3.7.4, M3.7.5, M3.7.7, M3.7.8 | ## Goal Prove the lookup answers real failures from history rather than handing back documentation, and that adding it changed nothing upstream. ## Facts (inlined — no spec read needed) The phase's claim is narrow and testable: **given a failure this project has solved before, the lookup returns the fix.** Everything else is machinery in service of that. **The measurement is a replay, not an A/B.** Collect real failures with known resolutions from session history and CI. Hold out half. Ingest the first half, then replay *all* of them against `/memory/context` and score: | Signal | What it proves | |---|---| | tier-1 hit rate on ingested failures | signature normalisation actually stabilises (M3.7.7) | | tier-2 recall on held-out failures | symptom projections generalise beyond exact repeats (M3.7.8) | | tier-3 rate on ingested failures | how often the system falls back to docs when it should have known | A high tier-3 rate on failures already in the corpus is the phase failing, and it is the number to watch. It means retrieval exists and does not fire. **No dependency on the orchestrator.** The consumer is any HTTP client — `pi`, curl, an MCP call. Nothing here requires `Poimen/workflows` to execute tool calls, which is what made the earlier version of this gate unrunnable. **Upstream must be undisturbed.** This phase adds a standing query (`tool-failures`), a second vector kind, and two tables. Each can perturb things that were green: update-rate for M1.8, recall width for M3.6.6, rebuild parity for M2.8. Re-assert all three. **Cost belongs in the result.** The lookup sits in front of real work. Report p50 and p95 for each tier separately — a 900ms tier-2 is a different product than a 40ms tier-1, and the averages hide it. ## Steps 1. Assemble ≥40 real failures with known resolutions across ≥4 tools; commit the set before running anything. 2. Split 50/50 into ingested and held-out. 3. Ingest the first half through the normal path — sessions, `tool-failures` standing query, gate, symptom projections, signatures. 4. Replay all 40 against `/memory/context`; record tier, rank of the correct answer, latency. 5. Re-run M1.8, M2.8 and M3.6.6. 6. Emit `expected/m3.7-gate.txt` with per-tool tier rates and latency percentiles; commit it, same rule as M1.8. ## Acceptance - Tier-1 hit rate on ingested failures ≥ 0.80. - Tier-2 returns the correct memory in the top 3 for ≥ 0.50 of held-out failures. - Tier-3 rate on ingested failures ≤ 0.10. - M1.8, M2.8 and M3.6.6 unchanged except for the added standing query. - Tier-1 p95 under 50ms; tier-2 p95 under 500ms. ## Verify **Harness:** live gateway, real database, the committed failure set. Long-running, `#[ignore]` by default, same posture as M1.8. **Integration test** — `tests/it_m3_7_gate.rs`: 1. `a1_tier1_hit_rate` — replay the ingested half; assert ≥ 0.80 return tier 1, print per-tool. 2. `a2_tier3_rate_bounded` — on that same half, assert ≤ 0.10 fall through to tier 3. This is the "retrieval exists but never fires" detector. 3. `a3_heldout_recall` — the held-out half; assert the correct memory is in the top 3 for ≥ 0.50, proving symptom projections generalise rather than memorise. 4. `a4_symptom_ablation` — delete `kind='symptom'` vectors, re-run `a3`; assert recall drops measurably. Without this, `a3` could be satisfied by the text vector alone and M3.7.8 would be dead weight. 5. `a5_signature_stability` — for failures appearing more than once in the set, assert every occurrence produced the same `sig_sha`. 6. `a6_precedence_held` — across the whole replay, assert no response placed an R result above a tier-1 or tier-2 lesson. 7. `a7_m1_8_unchanged` — re-run M1.8; per-query numbers match the committed baseline, `tool-failures` the only addition. 8. `a8_m2_8_rebuild` — drop and rebuild with vectors, signatures and supersede rows present; assert byte-identical. 9. `a9_m3_6_6_still_green` — re-run the M3.6 gate in full. 10. `a10_latency_by_tier` — p50/p95 per tier; assert the two thresholds. 11. `a11_no_orchestrator_dependency` — run the whole gate with `Poimen/workflows` absent; assert it completes. **Command:** `cargo test --workspace m3_7_gate -- --ignored --nocapture` **False pass:** - Replaying the ingested half only. It measures memorisation; `a3` on held-out data is the one that says anything about a failure you have not seen before. - Skipping `a4`. A symptom index that is empty, or full of paraphrase, passes every other assertion here — the ablation is the only proof it contributes. - Counting a tier-1 hit without checking the returned memory is the *right* one. A signature collision produces a confident wrong answer, which is worse than tier 3. - Building the failure set from failures the system already handles well. Fix the set first, commit it, then run. ## Traps - Curating resolutions after seeing what retrieval returns. The known-good answer for each failure has to be written down before the first replay. - Reading a low tier-1 rate as a retrieval problem. It is almost always normalisation (M3.7.7); check `mem sig explain` on the misses before touching anything downstream. - Letting the ingested half leak into the held-out half through near-duplicate failures. Split by incident, not by log file. --- Background: [DESIGN.md](../DESIGN.md) — tool context · [M1.8](M1.8-m1-gate.md) · [M3.6.6](M3.6.6-m3.6-gate.md)