Files
poimen-memory/PHASES-M3-M4-M5.md
T

256 lines
8.7 KiB
Markdown
Raw Normal View History

# Phases M3, M4, M5 — Remaining Work
## Current Status
| Phase | Tasks | Done | Status | Gate |
|-------|-------|------|--------|------|
| **M0** | 8 | 8 ✅ | COMPLETE | ✅ green |
| **M1** | 8 | 8 ✅ | COMPLETE | ✅ green |
| **M2** | 8 | 5 ✅ | 60% (core done) | ✅ green (core) |
| **M3** | 4 | 0 ⬜ | READY TO START | ⏳ M3.4 |
| **M4** | 3 | 0 ⬜ | BLOCKED on M3 | ⏳ M4.3 |
| **M5** | 6 | 0 ⬜ | BLOCKED on M3 | ⏳ M5.6 |
| **Total** | 64 | 33 ✅ | 52% complete | 4/10 gates green |
---
## M3 — L2 Synthesis + Retrieval (4 tasks, READY)
### What M3 does
**L2 is project-level memory** summarizing all L1 per-query memories. Uses the same gated loop as L1 but:
- Input: L1 memory nodes (handful, not hundreds)
- Question: synthesis question ("What is the current state of this project?")
- Exit gate: **ON** (can detect "enough evidence")
- Output: L2 memory (1024 tok max) with L1 parents
### Tasks
#### M3.1 — L2 synthesis pass
- **What:** `mem synthesize --project poimen`
- **Size:** M (13 days)
- **Status:** ✅ **Done** (code already exists)
- **Blocks:** M3.4
- Reuses `run_loop` from M1.5 with `use_exit_gate=true`
- Exit gate becomes effective here (paper measures 4× speedup)
#### M3.2 — Rerank client
- **What:** POST /v1/rerank with BAAI/bge-reranker-base (TEI endpoint)
- **Size:** S (< 1 day)
- **Status:** ⬜ Not started
- **Details:**
- Embed vectors are coarse filters, reranker is precision layer
- Discrimination: 0.98 vs 0.00009 (four orders of magnitude)
- Response shape: bare array `[{"index":i, "score":s}, ...]`, not OpenAI envelope
- Batch limits apply (50 candidates → rerank → 5 best)
#### M3.3 — `mem query` (embed → recall → rerank → provenance)
- **What:** `mem query "why did requests over 10KB fail?"`
- **Size:** M (13 days)
- **Status:** ⬜ Not started
- **Blocks:** M3.4
- **Pipeline:**
1. Embed question (768-dim, `nomic-embed-text-v2-moe`)
2. HNSW recall top-50 (pgvector index, filter by project + level)
3. Rerank top-50 → top-5
4. Walk `memory_edge` to provenance (L1 → L0, L2 → L1 → L0)
5. Return with citations
- **Default levels:** L1+L2 (synthesized answers), not L0
- **Output:** human-readable by default, `--format json` for programs
#### M3.4 — M3 gate (composition gate)
- **What:** Verify retrieval works end-to-end
- **Size:** M (13 days)
- **Status:** ⬜ Blocked on M3.1M3.3
- **Asserts:**
- Known-answer query returns correct L1 node
- L0 citation actually exists and is correct
- Provenance graph is complete (no broken edges)
---
## M4 — Skills (3 tasks, BLOCKED on M3)
### What M4 does
**Skills are procedural memory** — actionable instructions derived from L1/L2 descriptive memory.
- Drafts auto-generated in `_drafts/` (read-only)
- Promotion is manual (git action, auditable)
- No feedback loop: promoted skills don't regenerate even if better versions exist
### Tasks
#### M4.1 — `mem skill draft --from <note>`
- **What:** Turn L1/L2 memory into draft skill
- **Size:** M (13 days)
- **Status:** 🟡 In progress
- **What's done:**
- `mem-core/src/lesson.rs` has `render_skill()` (generates SKILL.md)
- `mem-cli/src/lessons_cmd.rs` has `mem materialize` (writes to disk)
- 17 unit tests exist
- **What remains:**
- Add `mem skill draft --from poimen/infra-root-causes` CLI
- Read L1/L2 nodes from pgvector
- Convert descriptive → procedural with LLM
- Write to `_drafts/` only (enforce with tests)
- Add `generated_from: <L2-sha>` provenance
- 7 integration tests
#### M4.2 — `derived: true` ingest filter
- **What:** Stop system learning from its own output
- **Size:** M (13 days)
- **Status:** ⬜ Not started
- **Details:**
```
emitted skill
→ loaded in session
→ appears in transcript
→ ingested as evidence
→ reinforces source memory
```
This is the only cycle. Guard:
1. Every emitted artifact hashed in manifest
2. During ingest, chunks matching artifact tagged `derived: true`
3. Gate never sees derived chunks
4. Use shingle overlap (strip whitespace, hash n-grams, threshold)
#### M4.3 — M4 gate (composition gate)
- **What:** Prove loop stays open (draft not loaded, promoted skill not evidence)
- **Size:** M (13 days)
- **Status:** ⬜ Blocked on M4.1M4.2
- **Two properties:**
1. Draft not loadable (lives in `_drafts/`, real `--skill vault/skills/` doesn't find it)
2. Promoted skill never enters evidence (derived filter stops it)
---
## M5 — Post-training (6 tasks, BLOCKED on M3, separate Python)
### What M5 does
**Fine-tune a LoRA adapter** on Qwen2.5-3B to improve gate behavior. Uses ground truth labels from the 32B reasoning model.
M5 is **separate from main Rust workspace** — Python, verl, LORA training.
### Tasks
#### M5.1 — Evidence labeler
- **What:** Use 32B `reasoning` model as offline labeler
- **Size:** M (13 days)
- **Status:** ⬜ Not started
- **Details:**
- Paper needed synthetic NIAH labels, we have real transcripts
- Labeler produces U_t ground truth: "does chunk answer Q?"
- Output: labeled corpus for verl training
#### M5.2 — Labeler calibration
- **What:** Hand-label holdout set, measure agreement with 32B labeler
- **Size:** M (13 days)
- **Status:** ⬜ Not started
- **Details:**
- Don't trust 32B blindly; calibrate before training
- Hand-label ~100 examples, measure Cohen's κ
- Use disagreement to adjust threshold/rules
#### M5.3 — Training corpus export
- **What:** Convert JSONL log → verl format
- **Size:** M (13 days)
- **Status:** ⬜ Not started
- **Details:**
- verl expects: prompt, responses, labels, rewards
- Export from log with labeled U_t and E_t
#### M5.4 — vLLM InferenceService
- **What:** Deploy vLLM with `--enable-lora` for LoRA serving
- **Size:** L (3+ days)
- **Status:** ⬜ Not started (can run in parallel)
- **Details:**
- Ollama cannot hot-swap LoRA
- vLLM can (pattern exists: `reasoning` already vLLM v0.11)
- GitOps K8s manifest, homelab deployment
- Separate from M0-M4
#### M5.5 — verl training loop
- **What:** Train LoRA on labeled corpus
- **Size:** L (3+ days)
- **Status:** ⬜ Not started
- **Rewards:**
- `r_update` ±1 (gate accuracy)
- `r_exit` {0, 0.5 late, 0.75 early} (early stopping penalty)
- Strict `r_format` (gate response format)
- `α=0.9` mixing (trajectory + turn-level advantage)
#### M5.6 — M5 gate (composition gate)
- **What:** Adapter beats prompted baseline
- **Size:** L (3+ days)
- **Status:** ⬜ Blocked on M5.1M5.5
- **Asserts:**
- Update accuracy on held-out project > prompted
- No regression on in-domain
- LoRA < 60 MB (swappable, fast reload)
---
## Optional Phases (not on critical path)
### M3.5 — Distributed API Layer (9 tasks)
- HTTP server with endpoints for ingest, query, skills
- Rate limiting, load balancing
- Already mostly done (status: ✅ M3.5.8 gate green in commit a4a4053)
### M3.6 — Reference Corpora (6 tasks)
- Extracted from documentation, papers, standards
- Same cycle-guard as M4 (has its own `derived: true` filter)
### M3.7 — Tool Context (6 tasks)
- Signature extraction, failure symptom projection
- Already 871 lines in `lesson.rs`, 17 unit tests
- In progress (status: 🟡)
### M6 — Agent-Manager Migration (6 tasks)
- Migrate separate `github.com/Riotpiaole/agent-manager` from sqlite to CNPG
- Different repo, not dependency of M0-M5
- Can run in parallel (no blocker)
---
## Remaining work summary
| What | How many | Blocker | Notes |
|------|----------|---------|-------|
| **M3 core** | 4 tasks | None | Ready now. M3.1, M3.2, M3.3 can start immediately |
| **M3 gate** | 1 task | M3.1-M3.3 done | Verification step |
| **M4** | 3 tasks | M3 gate green | Skills, cycle-guard, gate |
| **M5** | 6 tasks | M3 gate green | Python, training (can run M5.4 in parallel) |
| **M3.5** | 9 tasks | None | API layer (already mostly done) |
| **M3.6** | 6 tasks | None | Reference corpora (parallel track) |
| **M3.7** | 6 tasks | None | Tool context (6/6 ~60% done, parallel track) |
| **M6** | 6 tasks | None | Agent-manager migration (separate repo, parallel) |
| **Total remaining** | 38 tasks | — | 2 gates blocking (M3.4, M4.3, M5.6 all depend on M3) |
---
## Next immediate step: Start M3.1
```bash
# M3.1 is already ✅ done (code exists in M1.5 refactored)
# M3.2 is small (S, <1 day)
# M3.3 is medium (M, 1-3 days)
# M3.4 is gate (verification)
# To start:
cargo test -p mem-core
cargo test -p mem-cli
# Watch: do M3.1, M3.2, M3.3 compose? M3.4 gate proves it.
```
---
## Reading order
1. This file (what each phase does, blocking relationships)
2. Individual task files: `tasks/M3.1-l2-synthesis.md`, etc. (self-contained)
3. [DESIGN.md](DESIGN.md) (full motivation, see M3 section)
4. Post-training section [M5.4](tasks/M5.4-vllm-lora-serving.md) (can start now, parallel)