2026-08-19 09:52:07 -07:00
|
|
|
|
# M1.8 — M1 composition gate
|
|
|
|
|
|
|
|
|
|
|
|
| Field | Value |
|
|
|
|
|
|
|---|---|
|
|
|
|
|
|
| Phase | M1 — Gated loop at L1 |
|
|
|
|
|
|
| Size | M — 1–3 days |
|
2026-08-22 21:43:23 -07:00
|
|
|
|
| Status | ✅ Done |
|
2026-08-19 09:52:07 -07:00
|
|
|
|
| Flags | gate |
|
|
|
|
|
|
| Spec | inlined below |
|
|
|
|
|
|
| Blocks | all of M1 |
|
|
|
|
|
|
|
|
|
|
|
|
## Goal
|
|
|
|
|
|
|
|
|
|
|
|
Answer the only question that matters at this stage: **did we build a gate, or an
|
|
|
|
|
|
expensive summarizer?**
|
|
|
|
|
|
|
2026-08-22 21:43:23 -07:00
|
|
|
|
## Files
|
|
|
|
|
|
|
|
|
|
|
|
| Action | Path |
|
|
|
|
|
|
|---|---|
|
|
|
|
|
|
| Create | `tests/it_m1_gate.rs` — integration tests (workspace root, all `#[ignore]`, 7 assertions) |
|
|
|
|
|
|
| Create | `fixtures/expected/m1-gate.txt` — committed expectation file (summary table) |
|
|
|
|
|
|
|
|
|
|
|
|
## Dependencies
|
|
|
|
|
|
|
|
|
|
|
|
**None new.** This task runs existing code against the live gateway.
|
|
|
|
|
|
|
|
|
|
|
|
## Existing code to reuse
|
|
|
|
|
|
|
|
|
|
|
|
The entire M1 pipeline:
|
|
|
|
|
|
- `QuerySet::load()` (M1.2), `ChatClient` (M1.1), `run_loop()` (M1.5), `LogWriter` (M1.6)
|
|
|
|
|
|
- `PiSessionSource` / `ClaudeTranscriptSource` (M0.5/M0.6) — existing
|
|
|
|
|
|
- `chunks()` (M0.3) — existing
|
|
|
|
|
|
- `LogReader::stats()` (M1.6) — to recompute update-rate independently
|
|
|
|
|
|
|
2026-08-19 09:52:07 -07:00
|
|
|
|
## Facts (inlined — no spec read needed)
|
|
|
|
|
|
|
|
|
|
|
|
This gate runs against the **live gateway on a real project** and asserts
|
|
|
|
|
|
properties of the resulting log. It is the first phase that costs money in wall
|
|
|
|
|
|
clock, and the first that can fail for reasons no unit test can see.
|
|
|
|
|
|
|
|
|
|
|
|
Two thresholds, both from the paper:
|
|
|
|
|
|
|
|
|
|
|
|
**Update-rate < 30%.** Agent transcripts are ~43% tool results, mostly
|
|
|
|
|
|
evidence-free. Paper Figure 6 shows the failure mode directly: the ungated
|
|
|
|
|
|
MemAgent's memory climbs to its 1024-token ceiling and saturates, after which
|
|
|
|
|
|
"the accumulated noise can further impede subsequent updates". A high update-rate
|
|
|
|
|
|
is that curve starting.
|
|
|
|
|
|
|
|
|
|
|
|
**Memory size flat, not climbing.** Plot `memory.tokens` against `t`. GRU-Mem's
|
|
|
|
|
|
curve is low and roughly flat; the ungated curve rises to the cap and stays
|
|
|
|
|
|
pinned. A monotonic climb means the gate is open too often even if the rate
|
|
|
|
|
|
looks acceptable.
|
|
|
|
|
|
|
|
|
|
|
|
Third property, cheap and load-bearing: **the same run twice produces the same
|
|
|
|
|
|
chunk hashes**. Temperature affects the model's text, not the chunking; if chunk
|
|
|
|
|
|
shas differ between runs, identity includes something it should not (M0.2) and
|
|
|
|
|
|
every projection will churn.
|
|
|
|
|
|
|
|
|
|
|
|
## Steps
|
|
|
|
|
|
|
|
|
|
|
|
1. Run `mem ingest --project poimen` for all standing queries against the live
|
|
|
|
|
|
gateway.
|
|
|
|
|
|
2. Compute per query: update-rate, memory-token series, parse-failure count,
|
|
|
|
|
|
elapsed.
|
|
|
|
|
|
3. Assert the thresholds below.
|
|
|
|
|
|
4. Emit `expected/m1-gate.txt` with the summary table; commit it. Subsequent runs
|
|
|
|
|
|
diff against it, and a changed expectation is a reviewable claim.
|
|
|
|
|
|
5. Sample 20 gate decisions and have the 32B `reasoning` model audit them; report
|
|
|
|
|
|
agreement. Advisory at this gate, and the seed of M5.2's calibration.
|
|
|
|
|
|
|
|
|
|
|
|
## Acceptance
|
|
|
|
|
|
|
|
|
|
|
|
- Update-rate < 30% on every standing query.
|
|
|
|
|
|
- Memory tokens ≤ 1024 and not monotonically increasing.
|
|
|
|
|
|
- Parse-failure rate < 5%.
|
|
|
|
|
|
- Chunk shas stable across two runs.
|
|
|
|
|
|
|
|
|
|
|
|
## Verify
|
|
|
|
|
|
|
|
|
|
|
|
**Harness:** live gateway. Long-running; a nightly or on-demand job, not
|
|
|
|
|
|
per-push.
|
|
|
|
|
|
|
|
|
|
|
|
**Integration test** — `tests/it_m1_gate.rs`, all `#[ignore]` by default:
|
|
|
|
|
|
1. `a1_update_rate_under_threshold` — per query, assert < 0.30, print actual.
|
|
|
|
|
|
2. `a2_memory_bounded` — every `memory.tokens` ≤ 1024.
|
|
|
|
|
|
3. `a3_memory_not_climbing` — fit a line to tokens vs `t`; assert the slope is
|
|
|
|
|
|
below a small positive bound. Do not assert non-increasing — a legitimately
|
|
|
|
|
|
growing memory rises early then plateaus.
|
|
|
|
|
|
4. `a4_parse_failure_rate` — `parse_failed` / turns < 0.05.
|
|
|
|
|
|
5. `a5_chunk_sha_stable` — two runs, same chunk shas in the same order.
|
|
|
|
|
|
6. `a6_evidence_traceable` — every `evidence` sha appears as a parent of some
|
|
|
|
|
|
`memory` record.
|
|
|
|
|
|
7. `a7_judge_audit` — sample 20 decisions, ask the 32B model, print agreement.
|
|
|
|
|
|
Advisory; does not fail the gate.
|
|
|
|
|
|
|
2026-08-22 21:43:23 -07:00
|
|
|
|
**Command:** `cargo test --test it_m1_gate -- --ignored --nocapture`
|
2026-08-19 09:52:07 -07:00
|
|
|
|
|
|
|
|
|
|
**False pass:**
|
|
|
|
|
|
- Running the gate on a tiny `--limit`. Update-rate on the first 20 chunks is
|
|
|
|
|
|
noise; the failure mode is cumulative and needs a full project.
|
|
|
|
|
|
- Asserting the rate without printing it. The number is the artifact — a run at
|
|
|
|
|
|
29% passes and is telling you something a boolean hides.
|
|
|
|
|
|
- Asserting memory is non-increasing rather than bounded-slope. Real memory does
|
|
|
|
|
|
grow early, and a strict assertion here fails on correct behaviour, which
|
|
|
|
|
|
trains everyone to skip the gate.
|
|
|
|
|
|
|
|
|
|
|
|
## Traps
|
|
|
|
|
|
|
|
|
|
|
|
- Treating a high update-rate as a model-quality problem first. Check the
|
|
|
|
|
|
prompt (M1.3 golden file) and the parser (M1.4 defaults) before blaming the
|
|
|
|
|
|
3B model — a parser with `unwrap_or(true)` produces exactly this symptom.
|
|
|
|
|
|
- Tuning the threshold to whatever the first run produced. 30% comes from the
|
|
|
|
|
|
corpus composition; moving it to accommodate a bad result deletes the gate.
|
|
|
|
|
|
|
|
|
|
|
|
---
|
|
|
|
|
|
|
|
|
|
|
|
Background: [DESIGN.md](../DESIGN.md) — Verification · paper Fig 6, §4.2
|