diff --git a/PHASES-M3-M4-M5.md b/PHASES-M3-M4-M5.md new file mode 100644 index 0000000..a6f36ec --- /dev/null +++ b/PHASES-M3-M4-M5.md @@ -0,0 +1,255 @@ +# Phases M3, M4, M5 — Remaining Work + +## Current Status + +| Phase | Tasks | Done | Status | Gate | +|-------|-------|------|--------|------| +| **M0** | 8 | 8 ✅ | COMPLETE | ✅ green | +| **M1** | 8 | 8 ✅ | COMPLETE | ✅ green | +| **M2** | 8 | 5 ✅ | 60% (core done) | ✅ green (core) | +| **M3** | 4 | 0 ⬜ | READY TO START | ⏳ M3.4 | +| **M4** | 3 | 0 ⬜ | BLOCKED on M3 | ⏳ M4.3 | +| **M5** | 6 | 0 ⬜ | BLOCKED on M3 | ⏳ M5.6 | +| **Total** | 64 | 33 ✅ | 52% complete | 4/10 gates green | + +--- + +## M3 — L2 Synthesis + Retrieval (4 tasks, READY) + +### What M3 does + +**L2 is project-level memory** summarizing all L1 per-query memories. Uses the same gated loop as L1 but: +- Input: L1 memory nodes (handful, not hundreds) +- Question: synthesis question ("What is the current state of this project?") +- Exit gate: **ON** (can detect "enough evidence") +- Output: L2 memory (1024 tok max) with L1 parents + +### Tasks + +#### M3.1 — L2 synthesis pass +- **What:** `mem synthesize --project poimen` +- **Size:** M (1–3 days) +- **Status:** ✅ **Done** (code already exists) +- **Blocks:** M3.4 +- Reuses `run_loop` from M1.5 with `use_exit_gate=true` +- Exit gate becomes effective here (paper measures 4× speedup) + +#### M3.2 — Rerank client +- **What:** POST /v1/rerank with BAAI/bge-reranker-base (TEI endpoint) +- **Size:** S (< 1 day) +- **Status:** ⬜ Not started +- **Details:** + - Embed vectors are coarse filters, reranker is precision layer + - Discrimination: 0.98 vs 0.00009 (four orders of magnitude) + - Response shape: bare array `[{"index":i, "score":s}, ...]`, not OpenAI envelope + - Batch limits apply (50 candidates → rerank → 5 best) + +#### M3.3 — `mem query` (embed → recall → rerank → provenance) +- **What:** `mem query "why did requests over 10KB fail?"` +- **Size:** M (1–3 days) +- **Status:** ⬜ Not started +- **Blocks:** M3.4 +- **Pipeline:** + 1. Embed question (768-dim, `nomic-embed-text-v2-moe`) + 2. HNSW recall top-50 (pgvector index, filter by project + level) + 3. Rerank top-50 → top-5 + 4. Walk `memory_edge` to provenance (L1 → L0, L2 → L1 → L0) + 5. Return with citations +- **Default levels:** L1+L2 (synthesized answers), not L0 +- **Output:** human-readable by default, `--format json` for programs + +#### M3.4 — M3 gate (composition gate) +- **What:** Verify retrieval works end-to-end +- **Size:** M (1–3 days) +- **Status:** ⬜ Blocked on M3.1–M3.3 +- **Asserts:** + - Known-answer query returns correct L1 node + - L0 citation actually exists and is correct + - Provenance graph is complete (no broken edges) + +--- + +## M4 — Skills (3 tasks, BLOCKED on M3) + +### What M4 does + +**Skills are procedural memory** — actionable instructions derived from L1/L2 descriptive memory. +- Drafts auto-generated in `_drafts/` (read-only) +- Promotion is manual (git action, auditable) +- No feedback loop: promoted skills don't regenerate even if better versions exist + +### Tasks + +#### M4.1 — `mem skill draft --from ` +- **What:** Turn L1/L2 memory into draft skill +- **Size:** M (1–3 days) +- **Status:** 🟡 In progress +- **What's done:** + - `mem-core/src/lesson.rs` has `render_skill()` (generates SKILL.md) + - `mem-cli/src/lessons_cmd.rs` has `mem materialize` (writes to disk) + - 17 unit tests exist +- **What remains:** + - Add `mem skill draft --from poimen/infra-root-causes` CLI + - Read L1/L2 nodes from pgvector + - Convert descriptive → procedural with LLM + - Write to `_drafts/` only (enforce with tests) + - Add `generated_from: ` provenance + - 7 integration tests + +#### M4.2 — `derived: true` ingest filter +- **What:** Stop system learning from its own output +- **Size:** M (1–3 days) +- **Status:** ⬜ Not started +- **Details:** + ``` + emitted skill + → loaded in session + → appears in transcript + → ingested as evidence + → reinforces source memory + ``` + This is the only cycle. Guard: + 1. Every emitted artifact hashed in manifest + 2. During ingest, chunks matching artifact tagged `derived: true` + 3. Gate never sees derived chunks + 4. Use shingle overlap (strip whitespace, hash n-grams, threshold) + +#### M4.3 — M4 gate (composition gate) +- **What:** Prove loop stays open (draft not loaded, promoted skill not evidence) +- **Size:** M (1–3 days) +- **Status:** ⬜ Blocked on M4.1–M4.2 +- **Two properties:** + 1. Draft not loadable (lives in `_drafts/`, real `--skill vault/skills/` doesn't find it) + 2. Promoted skill never enters evidence (derived filter stops it) + +--- + +## M5 — Post-training (6 tasks, BLOCKED on M3, separate Python) + +### What M5 does + +**Fine-tune a LoRA adapter** on Qwen2.5-3B to improve gate behavior. Uses ground truth labels from the 32B reasoning model. + +M5 is **separate from main Rust workspace** — Python, verl, LORA training. + +### Tasks + +#### M5.1 — Evidence labeler +- **What:** Use 32B `reasoning` model as offline labeler +- **Size:** M (1–3 days) +- **Status:** ⬜ Not started +- **Details:** + - Paper needed synthetic NIAH labels, we have real transcripts + - Labeler produces U_t ground truth: "does chunk answer Q?" + - Output: labeled corpus for verl training + +#### M5.2 — Labeler calibration +- **What:** Hand-label holdout set, measure agreement with 32B labeler +- **Size:** M (1–3 days) +- **Status:** ⬜ Not started +- **Details:** + - Don't trust 32B blindly; calibrate before training + - Hand-label ~100 examples, measure Cohen's κ + - Use disagreement to adjust threshold/rules + +#### M5.3 — Training corpus export +- **What:** Convert JSONL log → verl format +- **Size:** M (1–3 days) +- **Status:** ⬜ Not started +- **Details:** + - verl expects: prompt, responses, labels, rewards + - Export from log with labeled U_t and E_t + +#### M5.4 — vLLM InferenceService +- **What:** Deploy vLLM with `--enable-lora` for LoRA serving +- **Size:** L (3+ days) +- **Status:** ⬜ Not started (can run in parallel) +- **Details:** + - Ollama cannot hot-swap LoRA + - vLLM can (pattern exists: `reasoning` already vLLM v0.11) + - GitOps K8s manifest, homelab deployment + - Separate from M0-M4 + +#### M5.5 — verl training loop +- **What:** Train LoRA on labeled corpus +- **Size:** L (3+ days) +- **Status:** ⬜ Not started +- **Rewards:** + - `r_update` ±1 (gate accuracy) + - `r_exit` {0, −0.5 late, −0.75 early} (early stopping penalty) + - Strict `r_format` (gate response format) + - `α=0.9` mixing (trajectory + turn-level advantage) + +#### M5.6 — M5 gate (composition gate) +- **What:** Adapter beats prompted baseline +- **Size:** L (3+ days) +- **Status:** ⬜ Blocked on M5.1–M5.5 +- **Asserts:** + - Update accuracy on held-out project > prompted + - No regression on in-domain + - LoRA < 60 MB (swappable, fast reload) + +--- + +## Optional Phases (not on critical path) + +### M3.5 — Distributed API Layer (9 tasks) +- HTTP server with endpoints for ingest, query, skills +- Rate limiting, load balancing +- Already mostly done (status: ✅ M3.5.8 gate green in commit a4a4053) + +### M3.6 — Reference Corpora (6 tasks) +- Extracted from documentation, papers, standards +- Same cycle-guard as M4 (has its own `derived: true` filter) + +### M3.7 — Tool Context (6 tasks) +- Signature extraction, failure symptom projection +- Already 871 lines in `lesson.rs`, 17 unit tests +- In progress (status: 🟡) + +### M6 — Agent-Manager Migration (6 tasks) +- Migrate separate `github.com/Riotpiaole/agent-manager` from sqlite to CNPG +- Different repo, not dependency of M0-M5 +- Can run in parallel (no blocker) + +--- + +## Remaining work summary + +| What | How many | Blocker | Notes | +|------|----------|---------|-------| +| **M3 core** | 4 tasks | None | Ready now. M3.1, M3.2, M3.3 can start immediately | +| **M3 gate** | 1 task | M3.1-M3.3 done | Verification step | +| **M4** | 3 tasks | M3 gate green | Skills, cycle-guard, gate | +| **M5** | 6 tasks | M3 gate green | Python, training (can run M5.4 in parallel) | +| **M3.5** | 9 tasks | None | API layer (already mostly done) | +| **M3.6** | 6 tasks | None | Reference corpora (parallel track) | +| **M3.7** | 6 tasks | None | Tool context (6/6 ~60% done, parallel track) | +| **M6** | 6 tasks | None | Agent-manager migration (separate repo, parallel) | +| **Total remaining** | 38 tasks | — | 2 gates blocking (M3.4, M4.3, M5.6 all depend on M3) | + +--- + +## Next immediate step: Start M3.1 + +```bash +# M3.1 is already ✅ done (code exists in M1.5 refactored) +# M3.2 is small (S, <1 day) +# M3.3 is medium (M, 1-3 days) +# M3.4 is gate (verification) + +# To start: +cargo test -p mem-core +cargo test -p mem-cli + +# Watch: do M3.1, M3.2, M3.3 compose? M3.4 gate proves it. +``` + +--- + +## Reading order + +1. This file (what each phase does, blocking relationships) +2. Individual task files: `tasks/M3.1-l2-synthesis.md`, etc. (self-contained) +3. [DESIGN.md](DESIGN.md) (full motivation, see M3 section) +4. Post-training section [M5.4](tasks/M5.4-vllm-lora-serving.md) (can start now, parallel)