Files
poimen-memory/tasks/M3.4-m3-gate.md

3.6 KiB
Raw Permalink Blame History

M3.4 — M3 composition gate

Field Value
Phase M3 — L2 synthesis and retrieval
Size M — 13 days
Status Not started
Flags gate
Spec inlined below
Blocks all of M3

Goal

Prove the whole stack answers a real question with real provenance — the first point at which the system is useful rather than merely correct.

Facts (inlined — no spec read needed)

The gate is a small set of known-answer questions with hand-written expected sources, committed to the repo. Each names a fact that genuinely appears in the ingested sessions and the session it appears in.

Seed set, all drawn from real work in this corpus:

question expected to cite
why did requests over 10KB fail? the Kong body-buffer / client_body_buffer_size finding
why did Authorization: Bearer return 401? the Kong key-auth header finding
what causes the 504 on a cold ornith start? the timeouts-on-Ingress-vs-Service finding

This is a retrieval quality gate, so it is graded, not boolean: report hit rate at k=5 and provenance precision (fraction of cited sources that actually contain the fact). A single failing question is information, not necessarily a stop.

The gate also re-runs mem verify including the level invariant, because M3.1 is the change most likely to break it.

Steps

  1. verify/known-answers.yaml — question, expected node, expected source.
  2. verify/m3.4.sh runs each through mem query --format json.
  3. Compute hit rate at 5 and provenance precision; print both.
  4. Assert the thresholds below.
  5. Run mem verify; assert zero violations.
  6. Assert L2 exists and its provenance resolves two hops.
  7. Commit expected/m3.4.txt; diff.

Acceptance

  • Hit rate at k=5 ≥ 0.8 on the known-answer set.
  • Provenance precision ≥ 0.9 — a citation that does not contain the fact is worse than no citation.
  • mem verify clean, including level consistency.
  • Every L2 hit resolves to L0 sources.

Verify

Harness: live gateway, seeded database from real logs. On-demand, not per-push.

Integration testverify/m3.4.sh diffed against expected/m3.4.txt:

  1. a1_hit_rate — ≥ 0.8, print actual.
  2. a2_provenance_precision — for each cited source, fetch the L0 text and assert it contains the expected fact substring; ≥ 0.9.
  3. a3_verify_clean — zero violations.
  4. a4_l2_two_hop — every L2 hit resolves through L1 to a real source.
  5. a5_rerank_contributes — hit rate with reranking is ≥ hit rate without. If reranking makes it worse, the wiring is wrong (probably index mapping, M3.2).
  6. a6_no_cross_project — a question about another project returns nothing from this one.

Command: bash verify/m3.4.sh | diff - expected/m3.4.txt

False pass:

  • Writing the known-answer set after seeing what the system returns. It then measures nothing. Write the questions and expected sources from the sessions first, independently of any query output.
  • Measuring hit rate without provenance precision. A system that returns the right note with fabricated citations scores 1.0 on hits and is untrustworthy — assertion 2 is the one that matters for whether anyone can act on an answer.

Traps

  • Tuning k until the hit rate passes. k=50 will hit almost everything and the metric stops meaning anything; the gate specifies k=5 for that reason.
  • Accepting a5 failing as "reranker is just not helping". It far more often means the index mapping in M3.2 is wrong and scores are attached to the wrong documents.

Background: DESIGN.md — Verification