Files
poimen-memory/tasks/M3.7.6-m3.7-gate.md

5.6 KiB
Raw Permalink Blame History

M3.7.6 — M3.7 composition gate

Field Value
Phase M3.7 — Tool context
Size M — 13 days
Status Not started
Flags gate
Spec inlined below
Blocks all of M3.7
Depends M3.7.3, M3.7.4, M3.7.5, M3.7.7, M3.7.8

Goal

Prove the lookup answers real failures from history rather than handing back documentation, and that adding it changed nothing upstream.

Facts (inlined — no spec read needed)

The phase's claim is narrow and testable: given a failure this project has solved before, the lookup returns the fix. Everything else is machinery in service of that.

The measurement is a replay, not an A/B. Collect real failures with known resolutions from session history and CI. Hold out half. Ingest the first half, then replay all of them against /memory/context and score:

Signal What it proves
tier-1 hit rate on ingested failures signature normalisation actually stabilises (M3.7.7)
tier-2 recall on held-out failures symptom projections generalise beyond exact repeats (M3.7.8)
tier-3 rate on ingested failures how often the system falls back to docs when it should have known

A high tier-3 rate on failures already in the corpus is the phase failing, and it is the number to watch. It means retrieval exists and does not fire.

No dependency on the orchestrator. The consumer is any HTTP client — pi, curl, an MCP call. Nothing here requires Poimen/workflows to execute tool calls, which is what made the earlier version of this gate unrunnable.

Upstream must be undisturbed. This phase adds a standing query (tool-failures), a second vector kind, and two tables. Each can perturb things that were green: update-rate for M1.8, recall width for M3.6.6, rebuild parity for M2.8. Re-assert all three.

Cost belongs in the result. The lookup sits in front of real work. Report p50 and p95 for each tier separately — a 900ms tier-2 is a different product than a 40ms tier-1, and the averages hide it.

Steps

  1. Assemble ≥40 real failures with known resolutions across ≥4 tools; commit the set before running anything.
  2. Split 50/50 into ingested and held-out.
  3. Ingest the first half through the normal path — sessions, tool-failures standing query, gate, symptom projections, signatures.
  4. Replay all 40 against /memory/context; record tier, rank of the correct answer, latency.
  5. Re-run M1.8, M2.8 and M3.6.6.
  6. Emit expected/m3.7-gate.txt with per-tool tier rates and latency percentiles; commit it, same rule as M1.8.

Acceptance

  • Tier-1 hit rate on ingested failures ≥ 0.80.
  • Tier-2 returns the correct memory in the top 3 for ≥ 0.50 of held-out failures.
  • Tier-3 rate on ingested failures ≤ 0.10.
  • M1.8, M2.8 and M3.6.6 unchanged except for the added standing query.
  • Tier-1 p95 under 50ms; tier-2 p95 under 500ms.

Verify

Harness: live gateway, real database, the committed failure set. Long-running, #[ignore] by default, same posture as M1.8.

Integration testtests/it_m3_7_gate.rs:

  1. a1_tier1_hit_rate — replay the ingested half; assert ≥ 0.80 return tier 1, print per-tool.
  2. a2_tier3_rate_bounded — on that same half, assert ≤ 0.10 fall through to tier 3. This is the "retrieval exists but never fires" detector.
  3. a3_heldout_recall — the held-out half; assert the correct memory is in the top 3 for ≥ 0.50, proving symptom projections generalise rather than memorise.
  4. a4_symptom_ablation — delete kind='symptom' vectors, re-run a3; assert recall drops measurably. Without this, a3 could be satisfied by the text vector alone and M3.7.8 would be dead weight.
  5. a5_signature_stability — for failures appearing more than once in the set, assert every occurrence produced the same sig_sha.
  6. a6_precedence_held — across the whole replay, assert no response placed an R result above a tier-1 or tier-2 lesson.
  7. a7_m1_8_unchanged — re-run M1.8; per-query numbers match the committed baseline, tool-failures the only addition.
  8. a8_m2_8_rebuild — drop and rebuild with vectors, signatures and supersede rows present; assert byte-identical.
  9. a9_m3_6_6_still_green — re-run the M3.6 gate in full.
  10. a10_latency_by_tier — p50/p95 per tier; assert the two thresholds.
  11. a11_no_orchestrator_dependency — run the whole gate with Poimen/workflows absent; assert it completes.

Command: cargo test --workspace m3_7_gate -- --ignored --nocapture

False pass:

  • Replaying the ingested half only. It measures memorisation; a3 on held-out data is the one that says anything about a failure you have not seen before.
  • Skipping a4. A symptom index that is empty, or full of paraphrase, passes every other assertion here — the ablation is the only proof it contributes.
  • Counting a tier-1 hit without checking the returned memory is the right one. A signature collision produces a confident wrong answer, which is worse than tier 3.
  • Building the failure set from failures the system already handles well. Fix the set first, commit it, then run.

Traps

  • Curating resolutions after seeing what retrieval returns. The known-good answer for each failure has to be written down before the first replay.
  • Reading a low tier-1 rate as a retrieval problem. It is almost always normalisation (M3.7.7); check mem sig explain on the misses before touching anything downstream.
  • Letting the ingested half leak into the held-out half through near-duplicate failures. Split by incident, not by log file.

Background: DESIGN.md — tool context · M1.8 · M3.6.6