M3 (4 tasks, READY): - M3.1: L2 synthesis (✅ done, code exists) - M3.2: Rerank client (⬜ not started, S size) - M3.3: mem query (⬜ not started, M size) - M3.4: Composition gate (⏳ blocked on M3.1-3.3) M4 (3 tasks, BLOCKED on M3): - M4.1: skill draft (🟡 60% done, lesson.rs exists) - M4.2: derived filter (⬜ not started, cycle-guard) - M4.3: Gate (⏳ blocked on M4.1-4.2) M5 (6 tasks, BLOCKED on M3, separate Python): - M5.1-5.3: Labeling, corpus export - M5.4: vLLM LoRA serving (can run in parallel) - M5.5: verl training loop - M5.6: Gate (adapter beats baseline) 64 total tasks: 33 done (52%), 3 in progress, 28 remaining 4/10 gates green
8.7 KiB
8.7 KiB
Phases M3, M4, M5 — Remaining Work
Current Status
| Phase | Tasks | Done | Status | Gate |
|---|---|---|---|---|
| M0 | 8 | 8 ✅ | COMPLETE | ✅ green |
| M1 | 8 | 8 ✅ | COMPLETE | ✅ green |
| M2 | 8 | 5 ✅ | 60% (core done) | ✅ green (core) |
| M3 | 4 | 0 ⬜ | READY TO START | ⏳ M3.4 |
| M4 | 3 | 0 ⬜ | BLOCKED on M3 | ⏳ M4.3 |
| M5 | 6 | 0 ⬜ | BLOCKED on M3 | ⏳ M5.6 |
| Total | 64 | 33 ✅ | 52% complete | 4/10 gates green |
M3 — L2 Synthesis + Retrieval (4 tasks, READY)
What M3 does
L2 is project-level memory summarizing all L1 per-query memories. Uses the same gated loop as L1 but:
- Input: L1 memory nodes (handful, not hundreds)
- Question: synthesis question ("What is the current state of this project?")
- Exit gate: ON (can detect "enough evidence")
- Output: L2 memory (1024 tok max) with L1 parents
Tasks
M3.1 — L2 synthesis pass
- What:
mem synthesize --project poimen - Size: M (1–3 days)
- Status: ✅ Done (code already exists)
- Blocks: M3.4
- Reuses
run_loopfrom M1.5 withuse_exit_gate=true - Exit gate becomes effective here (paper measures 4× speedup)
M3.2 — Rerank client
- What: POST /v1/rerank with BAAI/bge-reranker-base (TEI endpoint)
- Size: S (< 1 day)
- Status: ⬜ Not started
- Details:
- Embed vectors are coarse filters, reranker is precision layer
- Discrimination: 0.98 vs 0.00009 (four orders of magnitude)
- Response shape: bare array
[{"index":i, "score":s}, ...], not OpenAI envelope - Batch limits apply (50 candidates → rerank → 5 best)
M3.3 — mem query (embed → recall → rerank → provenance)
- What:
mem query "why did requests over 10KB fail?" - Size: M (1–3 days)
- Status: ⬜ Not started
- Blocks: M3.4
- Pipeline:
- Embed question (768-dim,
nomic-embed-text-v2-moe) - HNSW recall top-50 (pgvector index, filter by project + level)
- Rerank top-50 → top-5
- Walk
memory_edgeto provenance (L1 → L0, L2 → L1 → L0) - Return with citations
- Embed question (768-dim,
- Default levels: L1+L2 (synthesized answers), not L0
- Output: human-readable by default,
--format jsonfor programs
M3.4 — M3 gate (composition gate)
- What: Verify retrieval works end-to-end
- Size: M (1–3 days)
- Status: ⬜ Blocked on M3.1–M3.3
- Asserts:
- Known-answer query returns correct L1 node
- L0 citation actually exists and is correct
- Provenance graph is complete (no broken edges)
M4 — Skills (3 tasks, BLOCKED on M3)
What M4 does
Skills are procedural memory — actionable instructions derived from L1/L2 descriptive memory.
- Drafts auto-generated in
_drafts/(read-only) - Promotion is manual (git action, auditable)
- No feedback loop: promoted skills don't regenerate even if better versions exist
Tasks
M4.1 — mem skill draft --from <note>
- What: Turn L1/L2 memory into draft skill
- Size: M (1–3 days)
- Status: 🟡 In progress
- What's done:
mem-core/src/lesson.rshasrender_skill()(generates SKILL.md)mem-cli/src/lessons_cmd.rshasmem materialize(writes to disk)- 17 unit tests exist
- What remains:
- Add
mem skill draft --from poimen/infra-root-causesCLI - Read L1/L2 nodes from pgvector
- Convert descriptive → procedural with LLM
- Write to
_drafts/only (enforce with tests) - Add
generated_from: <L2-sha>provenance - 7 integration tests
- Add
M4.2 — derived: true ingest filter
- What: Stop system learning from its own output
- Size: M (1–3 days)
- Status: ⬜ Not started
- Details:
This is the only cycle. Guard:
emitted skill → loaded in session → appears in transcript → ingested as evidence → reinforces source memory- Every emitted artifact hashed in manifest
- During ingest, chunks matching artifact tagged
derived: true - Gate never sees derived chunks
- Use shingle overlap (strip whitespace, hash n-grams, threshold)
M4.3 — M4 gate (composition gate)
- What: Prove loop stays open (draft not loaded, promoted skill not evidence)
- Size: M (1–3 days)
- Status: ⬜ Blocked on M4.1–M4.2
- Two properties:
- Draft not loadable (lives in
_drafts/, real--skill vault/skills/doesn't find it) - Promoted skill never enters evidence (derived filter stops it)
- Draft not loadable (lives in
M5 — Post-training (6 tasks, BLOCKED on M3, separate Python)
What M5 does
Fine-tune a LoRA adapter on Qwen2.5-3B to improve gate behavior. Uses ground truth labels from the 32B reasoning model.
M5 is separate from main Rust workspace — Python, verl, LORA training.
Tasks
M5.1 — Evidence labeler
- What: Use 32B
reasoningmodel as offline labeler - Size: M (1–3 days)
- Status: ⬜ Not started
- Details:
- Paper needed synthetic NIAH labels, we have real transcripts
- Labeler produces U_t ground truth: "does chunk answer Q?"
- Output: labeled corpus for verl training
M5.2 — Labeler calibration
- What: Hand-label holdout set, measure agreement with 32B labeler
- Size: M (1–3 days)
- Status: ⬜ Not started
- Details:
- Don't trust 32B blindly; calibrate before training
- Hand-label ~100 examples, measure Cohen's κ
- Use disagreement to adjust threshold/rules
M5.3 — Training corpus export
- What: Convert JSONL log → verl format
- Size: M (1–3 days)
- Status: ⬜ Not started
- Details:
- verl expects: prompt, responses, labels, rewards
- Export from log with labeled U_t and E_t
M5.4 — vLLM InferenceService
- What: Deploy vLLM with
--enable-lorafor LoRA serving - Size: L (3+ days)
- Status: ⬜ Not started (can run in parallel)
- Details:
- Ollama cannot hot-swap LoRA
- vLLM can (pattern exists:
reasoningalready vLLM v0.11) - GitOps K8s manifest, homelab deployment
- Separate from M0-M4
M5.5 — verl training loop
- What: Train LoRA on labeled corpus
- Size: L (3+ days)
- Status: ⬜ Not started
- Rewards:
r_update±1 (gate accuracy)r_exit{0, −0.5 late, −0.75 early} (early stopping penalty)- Strict
r_format(gate response format) α=0.9mixing (trajectory + turn-level advantage)
M5.6 — M5 gate (composition gate)
- What: Adapter beats prompted baseline
- Size: L (3+ days)
- Status: ⬜ Blocked on M5.1–M5.5
- Asserts:
- Update accuracy on held-out project > prompted
- No regression on in-domain
- LoRA < 60 MB (swappable, fast reload)
Optional Phases (not on critical path)
M3.5 — Distributed API Layer (9 tasks)
- HTTP server with endpoints for ingest, query, skills
- Rate limiting, load balancing
- Already mostly done (status: ✅ M3.5.8 gate green in commit
a4a4053)
M3.6 — Reference Corpora (6 tasks)
- Extracted from documentation, papers, standards
- Same cycle-guard as M4 (has its own
derived: truefilter)
M3.7 — Tool Context (6 tasks)
- Signature extraction, failure symptom projection
- Already 871 lines in
lesson.rs, 17 unit tests - In progress (status: 🟡)
M6 — Agent-Manager Migration (6 tasks)
- Migrate separate
github.com/Riotpiaole/agent-managerfrom sqlite to CNPG - Different repo, not dependency of M0-M5
- Can run in parallel (no blocker)
Remaining work summary
| What | How many | Blocker | Notes |
|---|---|---|---|
| M3 core | 4 tasks | None | Ready now. M3.1, M3.2, M3.3 can start immediately |
| M3 gate | 1 task | M3.1-M3.3 done | Verification step |
| M4 | 3 tasks | M3 gate green | Skills, cycle-guard, gate |
| M5 | 6 tasks | M3 gate green | Python, training (can run M5.4 in parallel) |
| M3.5 | 9 tasks | None | API layer (already mostly done) |
| M3.6 | 6 tasks | None | Reference corpora (parallel track) |
| M3.7 | 6 tasks | None | Tool context (6/6 ~60% done, parallel track) |
| M6 | 6 tasks | None | Agent-manager migration (separate repo, parallel) |
| Total remaining | 38 tasks | — | 2 gates blocking (M3.4, M4.3, M5.6 all depend on M3) |
Next immediate step: Start M3.1
# M3.1 is already ✅ done (code exists in M1.5 refactored)
# M3.2 is small (S, <1 day)
# M3.3 is medium (M, 1-3 days)
# M3.4 is gate (verification)
# To start:
cargo test -p mem-core
cargo test -p mem-cli
# Watch: do M3.1, M3.2, M3.3 compose? M3.4 gate proves it.