3.6 KiB
M3.4 — M3 composition gate
| Field | Value |
|---|---|
| Phase | M3 — L2 synthesis and retrieval |
| Size | M — 1–3 days |
| Status | ⬜ Not started |
| Flags | gate |
| Spec | inlined below |
| Blocks | all of M3 |
Goal
Prove the whole stack answers a real question with real provenance — the first point at which the system is useful rather than merely correct.
Facts (inlined — no spec read needed)
The gate is a small set of known-answer questions with hand-written expected sources, committed to the repo. Each names a fact that genuinely appears in the ingested sessions and the session it appears in.
Seed set, all drawn from real work in this corpus:
| question | expected to cite |
|---|---|
| why did requests over 10KB fail? | the Kong body-buffer / client_body_buffer_size finding |
why did Authorization: Bearer return 401? |
the Kong key-auth header finding |
| what causes the 504 on a cold ornith start? | the timeouts-on-Ingress-vs-Service finding |
This is a retrieval quality gate, so it is graded, not boolean: report hit rate at k=5 and provenance precision (fraction of cited sources that actually contain the fact). A single failing question is information, not necessarily a stop.
The gate also re-runs mem verify including the level invariant, because M3.1 is
the change most likely to break it.
Steps
verify/known-answers.yaml— question, expected node, expected source.verify/m3.4.shruns each throughmem query --format json.- Compute hit rate at 5 and provenance precision; print both.
- Assert the thresholds below.
- Run
mem verify; assert zero violations. - Assert L2 exists and its provenance resolves two hops.
- Commit
expected/m3.4.txt; diff.
Acceptance
- Hit rate at k=5 ≥ 0.8 on the known-answer set.
- Provenance precision ≥ 0.9 — a citation that does not contain the fact is worse than no citation.
mem verifyclean, including level consistency.- Every L2 hit resolves to L0 sources.
Verify
Harness: live gateway, seeded database from real logs. On-demand, not per-push.
Integration test — verify/m3.4.sh diffed against expected/m3.4.txt:
a1_hit_rate— ≥ 0.8, print actual.a2_provenance_precision— for each cited source, fetch the L0 text and assert it contains the expected fact substring; ≥ 0.9.a3_verify_clean— zero violations.a4_l2_two_hop— every L2 hit resolves through L1 to a real source.a5_rerank_contributes— hit rate with reranking is ≥ hit rate without. If reranking makes it worse, the wiring is wrong (probably index mapping, M3.2).a6_no_cross_project— a question about another project returns nothing from this one.
Command: bash verify/m3.4.sh | diff - expected/m3.4.txt
False pass:
- Writing the known-answer set after seeing what the system returns. It then measures nothing. Write the questions and expected sources from the sessions first, independently of any query output.
- Measuring hit rate without provenance precision. A system that returns the right note with fabricated citations scores 1.0 on hits and is untrustworthy — assertion 2 is the one that matters for whether anyone can act on an answer.
Traps
- Tuning
kuntil the hit rate passes. k=50 will hit almost everything and the metric stops meaning anything; the gate specifies k=5 for that reason. - Accepting a5 failing as "reranker is just not helping". It far more often means
the
indexmapping in M3.2 is wrong and scores are attached to the wrong documents.
Background: DESIGN.md — Verification