125 lines
5.6 KiB
Markdown
125 lines
5.6 KiB
Markdown
# M3.7.6 — M3.7 composition gate
|
||
|
||
| Field | Value |
|
||
|---|---|
|
||
| Phase | M3.7 — Tool context |
|
||
| Size | M — 1–3 days |
|
||
| Status | ⬜ Not started |
|
||
| Flags | gate |
|
||
| Spec | inlined below |
|
||
| Blocks | all of M3.7 |
|
||
| Depends | M3.7.3, M3.7.4, M3.7.5, M3.7.7, M3.7.8 |
|
||
|
||
## Goal
|
||
|
||
Prove the lookup answers real failures from history rather than handing back
|
||
documentation, and that adding it changed nothing upstream.
|
||
|
||
## Facts (inlined — no spec read needed)
|
||
|
||
The phase's claim is narrow and testable: **given a failure this project has
|
||
solved before, the lookup returns the fix.** Everything else is machinery in
|
||
service of that.
|
||
|
||
**The measurement is a replay, not an A/B.** Collect real failures with known
|
||
resolutions from session history and CI. Hold out half. Ingest the first half,
|
||
then replay *all* of them against `/memory/context` and score:
|
||
|
||
| Signal | What it proves |
|
||
|---|---|
|
||
| tier-1 hit rate on ingested failures | signature normalisation actually stabilises (M3.7.7) |
|
||
| tier-2 recall on held-out failures | symptom projections generalise beyond exact repeats (M3.7.8) |
|
||
| tier-3 rate on ingested failures | how often the system falls back to docs when it should have known |
|
||
|
||
A high tier-3 rate on failures already in the corpus is the phase failing, and it
|
||
is the number to watch. It means retrieval exists and does not fire.
|
||
|
||
**No dependency on the orchestrator.** The consumer is any HTTP client — `pi`,
|
||
curl, an MCP call. Nothing here requires `Poimen/workflows` to execute tool calls,
|
||
which is what made the earlier version of this gate unrunnable.
|
||
|
||
**Upstream must be undisturbed.** This phase adds a standing query
|
||
(`tool-failures`), a second vector kind, and two tables. Each can perturb things
|
||
that were green: update-rate for M1.8, recall width for M3.6.6, rebuild parity for
|
||
M2.8. Re-assert all three.
|
||
|
||
**Cost belongs in the result.** The lookup sits in front of real work. Report p50
|
||
and p95 for each tier separately — a 900ms tier-2 is a different product than a
|
||
40ms tier-1, and the averages hide it.
|
||
|
||
## Steps
|
||
|
||
1. Assemble ≥40 real failures with known resolutions across ≥4 tools; commit the
|
||
set before running anything.
|
||
2. Split 50/50 into ingested and held-out.
|
||
3. Ingest the first half through the normal path — sessions, `tool-failures`
|
||
standing query, gate, symptom projections, signatures.
|
||
4. Replay all 40 against `/memory/context`; record tier, rank of the correct
|
||
answer, latency.
|
||
5. Re-run M1.8, M2.8 and M3.6.6.
|
||
6. Emit `expected/m3.7-gate.txt` with per-tool tier rates and latency
|
||
percentiles; commit it, same rule as M1.8.
|
||
|
||
## Acceptance
|
||
|
||
- Tier-1 hit rate on ingested failures ≥ 0.80.
|
||
- Tier-2 returns the correct memory in the top 3 for ≥ 0.50 of held-out failures.
|
||
- Tier-3 rate on ingested failures ≤ 0.10.
|
||
- M1.8, M2.8 and M3.6.6 unchanged except for the added standing query.
|
||
- Tier-1 p95 under 50ms; tier-2 p95 under 500ms.
|
||
|
||
## Verify
|
||
|
||
**Harness:** live gateway, real database, the committed failure set. Long-running,
|
||
`#[ignore]` by default, same posture as M1.8.
|
||
|
||
**Integration test** — `tests/it_m3_7_gate.rs`:
|
||
1. `a1_tier1_hit_rate` — replay the ingested half; assert ≥ 0.80 return tier 1,
|
||
print per-tool.
|
||
2. `a2_tier3_rate_bounded` — on that same half, assert ≤ 0.10 fall through to
|
||
tier 3. This is the "retrieval exists but never fires" detector.
|
||
3. `a3_heldout_recall` — the held-out half; assert the correct memory is in the
|
||
top 3 for ≥ 0.50, proving symptom projections generalise rather than memorise.
|
||
4. `a4_symptom_ablation` — delete `kind='symptom'` vectors, re-run `a3`; assert
|
||
recall drops measurably. Without this, `a3` could be satisfied by the text
|
||
vector alone and M3.7.8 would be dead weight.
|
||
5. `a5_signature_stability` — for failures appearing more than once in the set,
|
||
assert every occurrence produced the same `sig_sha`.
|
||
6. `a6_precedence_held` — across the whole replay, assert no response placed an R
|
||
result above a tier-1 or tier-2 lesson.
|
||
7. `a7_m1_8_unchanged` — re-run M1.8; per-query numbers match the committed
|
||
baseline, `tool-failures` the only addition.
|
||
8. `a8_m2_8_rebuild` — drop and rebuild with vectors, signatures and supersede
|
||
rows present; assert byte-identical.
|
||
9. `a9_m3_6_6_still_green` — re-run the M3.6 gate in full.
|
||
10. `a10_latency_by_tier` — p50/p95 per tier; assert the two thresholds.
|
||
11. `a11_no_orchestrator_dependency` — run the whole gate with `Poimen/workflows`
|
||
absent; assert it completes.
|
||
|
||
**Command:** `cargo test --workspace m3_7_gate -- --ignored --nocapture`
|
||
|
||
**False pass:**
|
||
- Replaying the ingested half only. It measures memorisation; `a3` on held-out
|
||
data is the one that says anything about a failure you have not seen before.
|
||
- Skipping `a4`. A symptom index that is empty, or full of paraphrase, passes
|
||
every other assertion here — the ablation is the only proof it contributes.
|
||
- Counting a tier-1 hit without checking the returned memory is the *right* one.
|
||
A signature collision produces a confident wrong answer, which is worse than
|
||
tier 3.
|
||
- Building the failure set from failures the system already handles well.
|
||
Fix the set first, commit it, then run.
|
||
|
||
## Traps
|
||
|
||
- Curating resolutions after seeing what retrieval returns. The known-good answer
|
||
for each failure has to be written down before the first replay.
|
||
- Reading a low tier-1 rate as a retrieval problem. It is almost always
|
||
normalisation (M3.7.7); check `mem sig explain` on the misses before touching
|
||
anything downstream.
|
||
- Letting the ingested half leak into the held-out half through near-duplicate
|
||
failures. Split by incident, not by log file.
|
||
|
||
---
|
||
|
||
Background: [DESIGN.md](../DESIGN.md) — tool context · [M1.8](M1.8-m1-gate.md) · [M3.6.6](M3.6.6-m3.6-gate.md)
|