Files
poimen-memory/tasks/M3.6.6-m3.6-gate.md
T

6.7 KiB
Raw Blame History

M3.6.6 — M3.6 composition gate

Field Value
Phase M3.6 — Reference corpora
Size M — 13 days
Status Not started
Flags gate
Spec inlined below
Blocks all of M3.6
Depends M3.6.1, M3.6.2, M3.6.3, M3.6.4, M3.6.5

Goal

Answer the question no single task in this phase can: did adding documentation change the memory system? It must not have.

Facts (inlined — no spec read needed)

Every task here was verified alone. What none of them own is the property that makes the phase safe: a corpus is additive to retrieval and invisible to everything else. Three ways that can silently fail, and this gate exists for them.

1. The M1.8 metric can be gamed by accident. Update-rate is chunks_used / chunks_seen, and M1 fails above 30%. Documentation is evidence-free against almost any standing question, so a corpus routed through the controller would push the ratio down and make M1.8 easier to pass while the memory got worse. Any implementation that improves a quality metric by adding unrelated text has inverted it. The gate re-runs M1.8 and asserts the numbers are unchanged, not merely still-passing.

2. Retrieval quality can degrade without any test noticing. Each task asserts its own behaviour on its own fixture. The composite risk is a corpus that outcompetes real project memory in the candidate pool — invisible to M3.6.5's unit fixture, obvious on the real poimen log with a real corpus loaded.

3. The cycle guard has two producers now. M4.2 writes skills to the manifest, M3.6.4 writes reference sections. They share a normaliser. Skills exclusion regressing when a corpus is added is the failure that no test in either phase catches, because each tests only its own kind.

Swappable parts. The phase claims two seams are real: DocCorpusSource is just another RecordSource, and the corpus is just another projection input. Prove both — swap the doc tree for a differently-shaped one and re-run, and rebuild the whole store from the log with the corpus present.

Steps

  1. Establish the baseline: on a clean store, run mem ingest --project poimen for all standing queries; record the M1.8 summary table.
  2. mem ref add --project poimen --corpus homelab-knowledge <tree> against a real corpus of at least 200 chunks.
  3. Re-run the full ingest. Diff the M1.8 summary against the baseline.
  4. Run the assertions below.
  5. Emit expected/m3.6-gate.txt with the summary; commit it. Later runs diff against it and a changed expectation is a reviewable claim, same rule as M1.8.
  6. Sample 10 abstentions and 10 R hits; eyeball whether the floor is set sanely. Advisory, as M1.8's judge audit is.

Acceptance

  • M1.8's numbers are identical before and after the corpus exists.
  • Default query output is byte-identical before and after.
  • No L1 or L2 node has an R parent.
  • Skills exclusion behaviour is unchanged with a corpus loaded.
  • Drop and rebuild reproduces the mixed store byte-identically.

Verify

Harness: live gateway, real corpus, real poimen log. Long-running; nightly or on-demand, #[ignore] by default, same posture as M1.8.

Integration testtests/it_m3_6_gate.rs:

  1. a1_update_rate_identical — per standing query, assert update-rate before and after the corpus is added is equal, not merely both under 0.30. Equality is the assertion; a threshold check here would pass the exact failure described above.
  2. a2_chunks_seen_identicalchunks_seen per run is unchanged, proving no R chunk entered the recurrence.
  3. a3_no_controller_calls_during_ref_ingest — run mem ref add under a chat transport that panics on request; assert it completes. Embeddings are allowed, controller calls are not, so the fake must distinguish the two endpoints.
  4. a4_default_query_byte_identical — snapshot default mem query output for 10 fixed questions before and after; assert byte-identical.
  5. a5_no_r_parentsSELECT count(*) FROM memory_edge WHERE parent_sha IN (SELECT sha256 FROM memory_node WHERE level='R') is 0 on the live store.
  6. a6_l2_stream_excludes_r — re-run L2 synthesis; assert its input stream contained only L1 nodes and the resulting L2 memories have no R ancestor.
  7. a7_skill_exclusion_unregressed — re-run M4.2's fixtures against the store with the corpus loaded; assert identical exclusion decisions.
  8. a8_rebuild_mixed_store — drop database and vault, mem rebuild --from-log, assert byte-identical across all four levels.
  9. a9_source_seam_swappable — point mem ref add at a structurally different tree (deep nesting, no headings in one file, one non-UTF8 file) and assert it ingests or fails cleanly, never partially.
  10. a10_corpus_does_not_starve_recall — for 10 project questions, assert the top-5 default hits are the same nodes as the pre-corpus baseline, with the corpus present in the table.
  11. a11_m5_export_excludes_r — run the M5.3 training-corpus export shape; assert zero R records appear. R carries no gate decision, so its presence would poison r_update labels with rows that have no ground truth.

Command: cargo test --workspace m3_6_gate -- --ignored --nocapture

False pass:

  • Asserting update-rate is still below 30% instead of unchanged. That is the precise shape that goes green while the gate is being fed documentation — assertion 1 must be equality.
  • Running the gate with a corpus small enough not to matter. 200 chunks is a floor, not a suggestion; below it, assertions 4 and 10 pass because the corpus never reaches the candidate pool.
  • Allowing a3's fake transport to reject all HTTP. Reference ingest legitimately calls the embeddings endpoint; a blanket panic passes the assertion for the wrong reason and would also pass if ingest did nothing at all.
  • Rebuilding into a fresh database rather than dropping the live one. A rebuild that never exercises deletion has not proved the projections are droppable.

Traps

  • Comparing M1.8 summaries by eye. The numbers move in the third decimal when the gateway is under load; the committed expected/ file plus an explicit tolerance is the only version of this that stays honest over months.
  • Treating a changed baseline as a corpus problem. If update-rate shifts, first confirm the gateway model has not changed underneath — reasoning and the 3B controller are both moving targets, and misattributing that to this phase burns a day.
  • Skipping assertion 11 because M5 is not built. The export shape is checkable now, and discovering R in the training corpus during M5.3 means re-running an expensive labelling pass.

Background: DESIGN.md — reference corpora, the tier model · M1.8 · M4.2