# Phases M3, M4, M5 — Remaining Work ## Current Status | Phase | Tasks | Done | Status | Gate | |-------|-------|------|--------|------| | **M0** | 8 | 8 ✅ | COMPLETE | ✅ green | | **M1** | 8 | 8 ✅ | COMPLETE | ✅ green | | **M2** | 8 | 5 ✅ | 60% (core done) | ✅ green (core) | | **M3** | 4 | 0 ⬜ | READY TO START | ⏳ M3.4 | | **M4** | 3 | 0 ⬜ | BLOCKED on M3 | ⏳ M4.3 | | **M5** | 6 | 0 ⬜ | BLOCKED on M3 | ⏳ M5.6 | | **Total** | 64 | 33 ✅ | 52% complete | 4/10 gates green | --- ## M3 — L2 Synthesis + Retrieval (4 tasks, READY) ### What M3 does **L2 is project-level memory** summarizing all L1 per-query memories. Uses the same gated loop as L1 but: - Input: L1 memory nodes (handful, not hundreds) - Question: synthesis question ("What is the current state of this project?") - Exit gate: **ON** (can detect "enough evidence") - Output: L2 memory (1024 tok max) with L1 parents ### Tasks #### M3.1 — L2 synthesis pass - **What:** `mem synthesize --project poimen` - **Size:** M (1–3 days) - **Status:** ✅ **Done** (code already exists) - **Blocks:** M3.4 - Reuses `run_loop` from M1.5 with `use_exit_gate=true` - Exit gate becomes effective here (paper measures 4× speedup) #### M3.2 — Rerank client - **What:** POST /v1/rerank with BAAI/bge-reranker-base (TEI endpoint) - **Size:** S (< 1 day) - **Status:** ⬜ Not started - **Details:** - Embed vectors are coarse filters, reranker is precision layer - Discrimination: 0.98 vs 0.00009 (four orders of magnitude) - Response shape: bare array `[{"index":i, "score":s}, ...]`, not OpenAI envelope - Batch limits apply (50 candidates → rerank → 5 best) #### M3.3 — `mem query` (embed → recall → rerank → provenance) - **What:** `mem query "why did requests over 10KB fail?"` - **Size:** M (1–3 days) - **Status:** ⬜ Not started - **Blocks:** M3.4 - **Pipeline:** 1. Embed question (768-dim, `nomic-embed-text-v2-moe`) 2. HNSW recall top-50 (pgvector index, filter by project + level) 3. Rerank top-50 → top-5 4. Walk `memory_edge` to provenance (L1 → L0, L2 → L1 → L0) 5. Return with citations - **Default levels:** L1+L2 (synthesized answers), not L0 - **Output:** human-readable by default, `--format json` for programs #### M3.4 — M3 gate (composition gate) - **What:** Verify retrieval works end-to-end - **Size:** M (1–3 days) - **Status:** ⬜ Blocked on M3.1–M3.3 - **Asserts:** - Known-answer query returns correct L1 node - L0 citation actually exists and is correct - Provenance graph is complete (no broken edges) --- ## M4 — Skills (3 tasks, BLOCKED on M3) ### What M4 does **Skills are procedural memory** — actionable instructions derived from L1/L2 descriptive memory. - Drafts auto-generated in `_drafts/` (read-only) - Promotion is manual (git action, auditable) - No feedback loop: promoted skills don't regenerate even if better versions exist ### Tasks #### M4.1 — `mem skill draft --from ` - **What:** Turn L1/L2 memory into draft skill - **Size:** M (1–3 days) - **Status:** 🟡 In progress - **What's done:** - `mem-core/src/lesson.rs` has `render_skill()` (generates SKILL.md) - `mem-cli/src/lessons_cmd.rs` has `mem materialize` (writes to disk) - 17 unit tests exist - **What remains:** - Add `mem skill draft --from poimen/infra-root-causes` CLI - Read L1/L2 nodes from pgvector - Convert descriptive → procedural with LLM - Write to `_drafts/` only (enforce with tests) - Add `generated_from: ` provenance - 7 integration tests #### M4.2 — `derived: true` ingest filter - **What:** Stop system learning from its own output - **Size:** M (1–3 days) - **Status:** ⬜ Not started - **Details:** ``` emitted skill → loaded in session → appears in transcript → ingested as evidence → reinforces source memory ``` This is the only cycle. Guard: 1. Every emitted artifact hashed in manifest 2. During ingest, chunks matching artifact tagged `derived: true` 3. Gate never sees derived chunks 4. Use shingle overlap (strip whitespace, hash n-grams, threshold) #### M4.3 — M4 gate (composition gate) - **What:** Prove loop stays open (draft not loaded, promoted skill not evidence) - **Size:** M (1–3 days) - **Status:** ⬜ Blocked on M4.1–M4.2 - **Two properties:** 1. Draft not loadable (lives in `_drafts/`, real `--skill vault/skills/` doesn't find it) 2. Promoted skill never enters evidence (derived filter stops it) --- ## M5 — Post-training (6 tasks, BLOCKED on M3, separate Python) ### What M5 does **Fine-tune a LoRA adapter** on Qwen2.5-3B to improve gate behavior. Uses ground truth labels from the 32B reasoning model. M5 is **separate from main Rust workspace** — Python, verl, LORA training. ### Tasks #### M5.1 — Evidence labeler - **What:** Use 32B `reasoning` model as offline labeler - **Size:** M (1–3 days) - **Status:** ⬜ Not started - **Details:** - Paper needed synthetic NIAH labels, we have real transcripts - Labeler produces U_t ground truth: "does chunk answer Q?" - Output: labeled corpus for verl training #### M5.2 — Labeler calibration - **What:** Hand-label holdout set, measure agreement with 32B labeler - **Size:** M (1–3 days) - **Status:** ⬜ Not started - **Details:** - Don't trust 32B blindly; calibrate before training - Hand-label ~100 examples, measure Cohen's κ - Use disagreement to adjust threshold/rules #### M5.3 — Training corpus export - **What:** Convert JSONL log → verl format - **Size:** M (1–3 days) - **Status:** ⬜ Not started - **Details:** - verl expects: prompt, responses, labels, rewards - Export from log with labeled U_t and E_t #### M5.4 — vLLM InferenceService - **What:** Deploy vLLM with `--enable-lora` for LoRA serving - **Size:** L (3+ days) - **Status:** ⬜ Not started (can run in parallel) - **Details:** - Ollama cannot hot-swap LoRA - vLLM can (pattern exists: `reasoning` already vLLM v0.11) - GitOps K8s manifest, homelab deployment - Separate from M0-M4 #### M5.5 — verl training loop - **What:** Train LoRA on labeled corpus - **Size:** L (3+ days) - **Status:** ⬜ Not started - **Rewards:** - `r_update` ±1 (gate accuracy) - `r_exit` {0, −0.5 late, −0.75 early} (early stopping penalty) - Strict `r_format` (gate response format) - `α=0.9` mixing (trajectory + turn-level advantage) #### M5.6 — M5 gate (composition gate) - **What:** Adapter beats prompted baseline - **Size:** L (3+ days) - **Status:** ⬜ Blocked on M5.1–M5.5 - **Asserts:** - Update accuracy on held-out project > prompted - No regression on in-domain - LoRA < 60 MB (swappable, fast reload) --- ## Optional Phases (not on critical path) ### M3.5 — Distributed API Layer (9 tasks) - HTTP server with endpoints for ingest, query, skills - Rate limiting, load balancing - Already mostly done (status: ✅ M3.5.8 gate green in commit a4a4053) ### M3.6 — Reference Corpora (6 tasks) - Extracted from documentation, papers, standards - Same cycle-guard as M4 (has its own `derived: true` filter) ### M3.7 — Tool Context (6 tasks) - Signature extraction, failure symptom projection - Already 871 lines in `lesson.rs`, 17 unit tests - In progress (status: 🟡) ### M6 — Agent-Manager Migration (6 tasks) - Migrate separate `github.com/Riotpiaole/agent-manager` from sqlite to CNPG - Different repo, not dependency of M0-M5 - Can run in parallel (no blocker) --- ## Remaining work summary | What | How many | Blocker | Notes | |------|----------|---------|-------| | **M3 core** | 4 tasks | None | Ready now. M3.1, M3.2, M3.3 can start immediately | | **M3 gate** | 1 task | M3.1-M3.3 done | Verification step | | **M4** | 3 tasks | M3 gate green | Skills, cycle-guard, gate | | **M5** | 6 tasks | M3 gate green | Python, training (can run M5.4 in parallel) | | **M3.5** | 9 tasks | None | API layer (already mostly done) | | **M3.6** | 6 tasks | None | Reference corpora (parallel track) | | **M3.7** | 6 tasks | None | Tool context (6/6 ~60% done, parallel track) | | **M6** | 6 tasks | None | Agent-manager migration (separate repo, parallel) | | **Total remaining** | 38 tasks | — | 2 gates blocking (M3.4, M4.3, M5.6 all depend on M3) | --- ## Next immediate step: Start M3.1 ```bash # M3.1 is already ✅ done (code exists in M1.5 refactored) # M3.2 is small (S, <1 day) # M3.3 is medium (M, 1-3 days) # M3.4 is gate (verification) # To start: cargo test -p mem-core cargo test -p mem-cli # Watch: do M3.1, M3.2, M3.3 compose? M3.4 gate proves it. ``` --- ## Reading order 1. This file (what each phase does, blocking relationships) 2. Individual task files: `tasks/M3.1-l2-synthesis.md`, etc. (self-contained) 3. [DESIGN.md](DESIGN.md) (full motivation, see M3 section) 4. Post-training section [M5.4](tasks/M5.4-vllm-lora-serving.md) (can start now, parallel)