Files

103 lines
3.9 KiB
Markdown
Raw Permalink Normal View History

2026-08-19 09:52:07 -07:00
# M3.1 — L2 synthesis pass
| Field | Value |
|---|---|
| Phase | M3 — L2 synthesis and retrieval |
| Size | M — 13 days |
| Status | ✅ Done |
2026-08-19 09:52:07 -07:00
| Flags | — |
| Spec | inlined below |
| Blocks | M1.5 |
## Goal
Project-level memory across the per-query memories — using the same loop, with
the exit gate switched on.
## Facts (inlined — no spec read needed)
```
mem synthesize --project poimen
```
L2 is **not new machinery**. It is `run_loop` (M1.5) with:
| | L1 | L2 |
|---|---|---|
| input stream | `Chunk` from sources | L1 `MemoryNode`s |
| question | per-query question | `synthesis.question` |
| `use_exit_gate` | false | **true** |
| `query_id` | set | NULL |
| parents | L0 evidence shas | L1 memory shas |
**Why the exit gate flips on.** At L1 the input is hundreds of chunks and the
question is exhaustive ("what are *all* the X"), which is exactly the case paper
§3.3 says to run without the gate. At L2 the input is a handful of memories and
"enough evidence" is genuinely decidable, which is the case the gate was designed
for and where the paper measures its 4× speedup.
If M1.5 was written correctly this task is mostly wiring. If it needs changes to
`run_loop`, the level was not really a parameter — and assertion a10 in M1.5
existed to prevent exactly that.
Ordering: L1 memories enter the stream in a stable order (query id, ascending),
so synthesis is reproducible.
## Steps
1. `mem synthesize --project P` reads the final L1 memory per standing query.
2. Wrap them as the loop's input stream, in sorted query-id order.
3. Run `run_loop` with `level = L2`, `use_exit_gate = true`, the synthesis
question, `query_id = None`.
4. Write to `log/<project>/_synthesis/<run-id>.jsonl`.
5. `parents` on the L2 memory are the L1 memory shas consumed up to that turn.
6. Refuse to run if any standing query has no completed L1 run — synthesizing
over a partial set silently produces a partial picture.
## Acceptance
- No change to `run_loop` is required.
- The exit gate fires and stops early on a real project.
- L2 parents are L1 shas, never L0.
- Sorted input order makes two runs consume memories in the same sequence.
## Verify
**Harness:** scripted client, plus one live run.
**Integration test**`tests/it_l2.rs`:
1. `a1_reuses_run_loop` — assert `mem synthesize` calls the same `run_loop`
symbol; a duplicated loop is a review failure, and a `#[deny]`-style test here
is a grep asserting `fn run_loop` appears exactly once in the workspace.
2. `a2_exit_gate_on` — scripted `end` at turn 2 of 5; assert it stops at 2.
3. `a3_parents_are_l1` — every L2 parent sha resolves to an L1 node.
4. `a4_query_id_null` — the L2 record has no `query_id`.
5. `a5_stable_input_order` — two runs consume L1 memories in identical order.
6. `a6_refuses_partial` — one query with no completed run; assert non-zero exit
naming the query.
7. `a7_level_check_holds` — run `mem verify`; invariant 6 (level consistency)
passes.
8. `a8_live``#[ignore]`; real project, assert an L2 memory is produced and
print whether the exit gate fired and at which turn.
**Command:** `cargo test -p mem-cli l2` (add `-- --ignored` for a8)
**False pass:**
- Copying `run_loop` into an L2-specific function. Everything passes, the two
drift within a month, and the tier model quietly becomes two implementations.
Assertion 1 is the guard.
- Testing the exit gate with a script that never says `end`. The gate's effect is
invisible and a `use_exit_gate` that is ignored passes.
## Traps
- Feeding L0 evidence into L2 "for more detail". It blows the context budget and
breaks the level invariant; L2 reads memories, and if they are inadequate the
fix is at L1.
- Synthesizing over whatever L1 runs happen to exist. A missing query produces a
confident summary of an incomplete project, which is worse than no summary.
---
Background: [DESIGN.md](../DESIGN.md) — tier model · paper §3.3