6.7 KiB
M3.6.6 — M3.6 composition gate
| Field | Value |
|---|---|
| Phase | M3.6 — Reference corpora |
| Size | M — 1–3 days |
| Status | ⬜ Not started |
| Flags | gate |
| Spec | inlined below |
| Blocks | all of M3.6 |
| Depends | M3.6.1, M3.6.2, M3.6.4, M3.6.5 |
Goal
Answer the question no single task in this phase can: did adding documentation change the memory system? It must not have.
Facts (inlined — no spec read needed)
Every task here was verified alone. What none of them own is the property that makes the phase safe: a corpus is additive to retrieval and invisible to everything else. Three ways that can silently fail, and this gate exists for them.
1. The M1.8 metric can be gamed by accident. Update-rate is
chunks_used / chunks_seen, and M1 fails above 30%. Documentation is
evidence-free against almost any standing question, so a corpus routed through
the controller would push the ratio down and make M1.8 easier to pass while the
memory got worse. Any implementation that improves a quality metric by adding
unrelated text has inverted it. The gate re-runs M1.8 and asserts the numbers are
unchanged, not merely still-passing.
2. Retrieval quality can degrade without any test noticing. Each task asserts its own behaviour on its own fixture. The composite risk is a corpus that outcompetes real project memory in the candidate pool — invisible to M3.6.5's unit fixture, obvious on the real poimen log with a real corpus loaded.
3. The cycle guard has two producers now. M4.2 writes skills to the manifest, M3.6.4 writes reference sections. They share a normaliser. Skills exclusion regressing when a corpus is added is the failure that no test in either phase catches, because each tests only its own kind.
Swappable parts. The phase claims two seams are real: DocCorpusSource is
just another RecordSource, and the corpus is just another projection input.
Prove both — swap the doc tree for a differently-shaped one and re-run, and
rebuild the whole store from the log with the corpus present.
Steps
- Establish the baseline: on a clean store, run
mem ingest --project poimenfor all standing queries; record the M1.8 summary table. mem ref add --project poimen --corpus homelab-knowledge <tree>against a real corpus of at least 200 chunks.- Re-run the full ingest. Diff the M1.8 summary against the baseline.
- Run the assertions below.
- Emit
expected/m3.6-gate.txtwith the summary; commit it. Later runs diff against it and a changed expectation is a reviewable claim, same rule as M1.8. - Sample 10 abstentions and 10 R hits; eyeball whether the floor is set sanely. Advisory, as M1.8's judge audit is.
Acceptance
- M1.8's numbers are identical before and after the corpus exists.
- Default query output is byte-identical before and after.
- No L1 or L2 node has an R parent.
- Skills exclusion behaviour is unchanged with a corpus loaded.
- Drop and rebuild reproduces the mixed store byte-identically.
Verify
Harness: live gateway, real corpus, real poimen log. Long-running; nightly or
on-demand, #[ignore] by default, same posture as M1.8.
Integration test — tests/it_m3_6_gate.rs:
a1_update_rate_identical— per standing query, assert update-rate before and after the corpus is added is equal, not merely both under 0.30. Equality is the assertion; a threshold check here would pass the exact failure described above.a2_chunks_seen_identical—chunks_seenper run is unchanged, proving no R chunk entered the recurrence.a3_no_controller_calls_during_ref_ingest— runmem ref addunder a chat transport that panics on request; assert it completes. Embeddings are allowed, controller calls are not, so the fake must distinguish the two endpoints.a4_default_query_byte_identical— snapshot defaultmem queryoutput for 10 fixed questions before and after; assert byte-identical.a5_no_r_parents—SELECT count(*) FROM memory_edge WHERE parent_sha IN (SELECT sha256 FROM memory_node WHERE level='R')is 0 on the live store.a6_l2_stream_excludes_r— re-run L2 synthesis; assert its input stream contained only L1 nodes and the resulting L2 memories have no R ancestor.a7_skill_exclusion_unregressed— re-run M4.2's fixtures against the store with the corpus loaded; assert identical exclusion decisions.a8_rebuild_mixed_store— drop database and vault,mem rebuild --from-log, assert byte-identical across all four levels.a9_source_seam_swappable— pointmem ref addat a structurally different tree (deep nesting, no headings in one file, one non-UTF8 file) and assert it ingests or fails cleanly, never partially.a10_corpus_does_not_starve_recall— for 10 project questions, assert the top-5 default hits are the same nodes as the pre-corpus baseline, with the corpus present in the table.a11_m5_export_excludes_r— run the M5.3 training-corpus export shape; assert zero R records appear. R carries no gate decision, so its presence would poisonr_updatelabels with rows that have no ground truth.
Command: cargo test --workspace m3_6_gate -- --ignored --nocapture
False pass:
- Asserting update-rate is still below 30% instead of unchanged. That is the precise shape that goes green while the gate is being fed documentation — assertion 1 must be equality.
- Running the gate with a corpus small enough not to matter. 200 chunks is a floor, not a suggestion; below it, assertions 4 and 10 pass because the corpus never reaches the candidate pool.
- Allowing
a3's fake transport to reject all HTTP. Reference ingest legitimately calls the embeddings endpoint; a blanket panic passes the assertion for the wrong reason and would also pass if ingest did nothing at all. - Rebuilding into a fresh database rather than dropping the live one. A rebuild that never exercises deletion has not proved the projections are droppable.
Traps
- Comparing M1.8 summaries by eye. The numbers move in the third decimal when the
gateway is under load; the committed
expected/file plus an explicit tolerance is the only version of this that stays honest over months. - Treating a changed baseline as a corpus problem. If update-rate shifts, first
confirm the gateway model has not changed underneath —
reasoningand the 3B controller are both moving targets, and misattributing that to this phase burns a day. - Skipping assertion 11 because M5 is not built. The export shape is checkable now, and discovering R in the training corpus during M5.3 means re-running an expensive labelling pass.
Background: DESIGN.md — reference corpora, the tier model · M1.8 · M4.2