# poimen-memory Gated recurrent memory over agent context. Reads session history chunk-by-chunk, keeps only what answers standing questions, projects result into an Obsidian vault and a pgvector index. **Status: design complete, no code yet.** 37 tasks in [memory-tasks/](memory-tasks/INDEX.md), 0 done. Start at [M0.1](memory-tasks/M0.1-cargo-workspace.md). ## Problem Agent sessions grow faster than anyone reads them, and most of the volume is noise. One real pi session in this project: ``` assistant 1445 toolResult 1261 43% — ls output, file reads, mostly evidence-free user 196 + 8 compaction events ``` Compaction fires 8 times per session. Context gets *discarded*, not retained — root causes, decisions and gotchas evaporate when window rolls. ## Mechanism GRU-Mem ([arXiv 2602.10560](https://arxiv.org/abs/2602.10560)). Two text-controlled gates on a recurrent memory loop: - **update gate** — memory only mutates when chunk contains evidence. Blocks the memory explosion that ungated recurrent memory hits. - **exit gate** — stop scanning once evidence sufficient. Paper reports up to 400% speedup and *better* accuracy than ungated, because unbounded memory growth degrades later updates. ``` sessions ─> chunk (5000 tok) ─> controller ─> gates ─> memory ─> projections ``` Controller emits structured output; loop acts on it: ``` reason about chunk vs question yes|no -> U_t, update or discard candidate memory M̂_t continue|end -> E_t, exit or continue ``` ## Memory tiers | Level | What | From | Bounded | |---|---|---|---| | **L0** | evidence chunk, verbatim | update gate opening | no, but sparse (~17 of 412) | | **L1** | per-query memory, `M_t` | gated loop over chunks | 1024 tok | | **L2** | project synthesis | gated loop over L1 memories | 1024 tok | L2 is not new machinery — same loop, same prompt, L1 memories as input stream. Level is a parameter. Tiers form a provenance graph. Each L1 records its L0 parents, each L2 its L1 parents. Same relation becomes both `memory_edge` rows and Obsidian wikilinks. ## Standing queries Update gate needs a referent. Paper's agent is `φθ(Q, C_t, M_{t-1})` — gate is defined as "does this chunk contain useful information *about the problem*". No `Q`, no gate, and `r_update` becomes undefinable, which kills post-training. So each project declares durable questions. One query = one L1 memory = one note. ```yaml # queries/poimen.yaml project: poimen roots: [/Users/rockliang/workplace/Poimen/agent-rust] queries: - id: infra-root-causes question: What infrastructure bugs were found, what was the root cause, how was it isolated? - id: architecture-decisions question: What architectural decisions were made, with reasoning and rejected alternatives? synthesis: question: What is the current state of this project, and what should someone know before working on it? exit_gate: true ``` **Exit gate off at L1, on at L2.** Paper §3.3 makes this call: for "what are *all* the X" questions you cannot know evidence is sufficient without reading everything. L1 extraction is that shape. At L2 input is a handful of memories and sufficiency is decidable. Gate still *recorded* at L1 — signal needed for post-training. ## Authority model **JSONL log authoritative. Vault and vector index are projections.** Anything not rebuildable byte-identically from the log has hidden inputs, and that is a bug. Gate M2.8 enforces it destructively: ```sh rm -rf vault/poimen psql -c "delete from memory_node where project='poimen'" mem rebuild --from-log --project poimen git -C vault diff --exit-code # empty diff is the only pass ``` Buys three things: re-embedding after model change is a rebuild not a migration, Obsidian edits cannot corrupt the record, post-training corpus is the log itself. ## Skills A skill is a **projection, not a level**. L0/L1/L2 are descriptive — what happened. A skill is procedural — what to do next time. Gated loop does not produce it. Format free: `SKILL.md` is YAML frontmatter + markdown, which is an Obsidian note. So `vault/skills//SKILL.md` is both, no conversion: ```sh pi --skill vault/skills/ ln -s .../vault/skills/ ~/.claude/skills/ ``` **Drafts land in `_drafts/`, promotion is a human `git mv`.** This is the one cycle in the design: ``` emitted skill auto-loads -> appears in future transcripts -> ingested as evidence -> reinforces the memory that emitted it ``` No external verifier breaks it. Two guards: `_drafts/` is a directory (cannot be globbed into `--skill`), and every artifact carries `generated_from` so ingest tags matching chunks `derived: true` and refuses them as evidence. ## Separate weights Memory policy is a **LoRA adapter** on Qwen2.5-3B-Instruct, not a fine-tuned model. Reason is VRAM: one GPU, `OLLAMA_MAX_LOADED_MODELS=2`, already holding `ornith:35b` + `qwen2.5:3b`. Separate full model evicts something, and eviction is a weights reload measured in tens of seconds. Adapter rides the resident base. Also: post-training emits ~50 MB, not 6 GB. Swap without redeploy. Regression reverts by pointing at previous adapter. **Ollama cannot hot-swap LoRA.** Serving one needs vLLM with `--enable-lora` (pattern already exists — `reasoning` predictor is vLLM v0.11.0). Phases M0–M4 run prompted-only, so decision is deferred, not dodged. ## Layout ``` DESIGN.md full design, 460 lines memory-tasks/ 37 task files + INDEX.md — tracked crates/ mem-core/ domain types; Level; gate parser; the gated loop mem-chunk/ RecordSource trait; ChunkPolicy; FlushTrigger mem-llm/ gateway client — chat, embeddings, rerank mem-ingest/ source adapters: pi sessions, claude transcripts mem-store/ JSONL log; pgvector repo; Obsidian projector mem-cli/ binary `mem` queries/ standing query YAML per project log/ JSONL event log — authoritative, tracked vault/ Obsidian output ``` `mem-chunk` is separate and stream-shaped from day one. Sources today are files with an EOF; telemetry or a live tail will not have one. `RecordSource` returns `impl Stream`; batch sources become streams via `futures::stream::iter`, so it costs nothing now and removes a rewrite later. ## Commands ```sh mem ingest --project poimen --dry-run # chunk plan, zero model calls mem ingest --project poimen --query infra-root-causes mem synthesize --project poimen # L2 pass, exit gate on mem rebuild --from-log --project poimen # drop and rebuild projections mem verify --project poimen # provenance graph closure mem query "why did requests over 10KB fail?" mem skill draft --from poimen/infra-root-causes mem label --project poimen # evidence labels for training ``` ## M3.8 Pluggable Query Optimization **Purpose**: Compress and optimize search results before passing them to the LLM context window, improving token efficiency and response quality. ### Architecture M3.8 provides **dual-path optimization**: #### Ingest-Time Optimization (M3.8.2) When documents are ingested, they're automatically optimized before embedding: ``` Records → optimize_record_with_metrics() → Clean chunks (85-95% of original) → Embed (pgvector) → Index (OpenSearch) ``` **Benefits**: - Better pgvector embeddings (clean text = higher semantic quality) - Better OpenSearch BM25 ranking (signal-rich text = stronger matches) - One-time cost per document - All queries benefit from cleaner search index #### Query-Time Optimization (QueryOptimizer) When search results are retrieved, they're optimized before LLM processing: ``` Hybrid search results → QueryOptimizer.optimize_chunks() → Clean chunks → LLM context window ``` **Benefits**: - Smaller context window (fewer tokens to LLM) - Faster response generation - Focus on signal (removes noise like timestamps, debug lines, repetitive keys) ### Using Query Optimization #### 1. Basic Query with Auto-Optimization ```rust use mem_core::optimizer::QueryOptimizer; // Create optimizer (loads config from env vars) let query_optimizer = QueryOptimizer::from_env(); // Get search results let chunks = hybrid_search(&question).await?; // Auto-optimize before LLM let optimized = query_optimizer.optimize_chunks(&chunks).await?; // Build context from clean chunks let context = optimized.join("\n---\n"); let response = llm.prompt(&context, &question).await?; ``` #### 2. Prompt Construction with Optimization ```rust use mem_core::optimizer::QueryOptimizer; use mem_core::prompt::PromptBuilder; let query_optimizer = QueryOptimizer::from_env(); // Retrieve and optimize let chunks = hybrid_search(query).await?; let optimized_chunks = query_optimizer.optimize_chunks(&chunks).await?; // Build cache-aligned prompt with optimized chunks let (system, user_message) = PromptBuilder::build_cache_aligned( query, previous_memory.as_deref(), /* use optimized chunks */ )?; let response = llm.prompt(system, user_message).await?; ``` #### 3. Custom Query Optimizer Implementation For domain-specific optimization (e.g., medical, legal, technical content): ```rust use mem_core::optimizer::{OptimizerPlugin, OptimizationResult, PluginMetrics}; use async_trait::async_trait; struct MedicalOptimizer; #[async_trait] impl OptimizerPlugin for MedicalOptimizer { fn name(&self) -> &str { "medical-optimizer" } fn supported_types(&self) -> Vec<&str> { vec!["text/medical", "text/clinical", "application/json"] } async fn optimize(&self, content: &str) -> Result { // Remove patient IDs, reduce duplicate diagnosis entries let cleaned = clean_medical_data(content); let ratio = cleaned.len() as f32 / content.len() as f32; Ok(OptimizationResult { original: content.to_string(), optimized: cleaned, ratio, plugin: self.name().to_string(), metadata: Default::default(), }) } fn metrics(&self) -> PluginMetrics { Default::default() } } // Register and use let service = OptimizerServiceBuilder::new() .with_optimizer(Arc::new(MedicalOptimizer)) .with_format(Arc::new(JsonFormatter)) .build()?; let optimized = service.optimize( clinical_note, "text/clinical", None ).await?; ``` #### 4. Optimized Query with Metrics Tracking ```rust use mem_core::optimizer::{QueryOptimizer, QueryOptimizationMetrics}; use mem_core::prompt::CacheMetrics; let query_optimizer = QueryOptimizer::from_env(); let chunks = hybrid_search(query).await?; let optimized = query_optimizer.optimize_chunks(&chunks).await?; // Track optimization effectiveness let metrics: Vec = chunks .iter() .zip(&optimized) .map(|(orig, opt)| { QueryOptimizationMetrics { original_bytes: orig.text.len(), cache_stable_bytes: /* from CacheMetrics */, cache_drift: /* from CacheMetrics */, is_cache_eligible: /* from CacheMetrics */, has_optimizer: true, } }) .collect(); tracing::info!( chunks = chunks.len(), compression_ratio = format!( "{:.1}%", (optimized.iter().map(|o| o.len()).sum::() as f32 / chunks.iter().map(|c| c.text.len()).sum::() as f32) * 100.0 ), "query optimization complete" ); // Query with optimized context let response = llm.prompt(&optimized.join("\n---\n"), &question).await?; ``` #### 5. Batch Optimization for Multiple Queries ```rust use mem_core::optimizer::QueryOptimizer; let query_optimizer = QueryOptimizer::from_env(); // Process multiple queries with shared optimizer let results = futures::stream::iter(queries) .then(|query| async move { let chunks = hybrid_search(&query).await?; let optimized = query_optimizer.optimize_chunks(&chunks).await?; let response = llm.prompt(&optimized.join("\n---\n"), &query.question).await?; Ok((query, response)) }) .collect::>() .await; ``` #### 6. Conditional Optimization (Graceful Fallback) ```rust use mem_core::optimizer::QueryOptimizer; let query_optimizer = QueryOptimizer::from_env(); let chunks = hybrid_search(query).await?; // Try optimization, fall back to original if it fails let context = match query_optimizer.optimize_chunks(&chunks).await { Ok(optimized) => { tracing::info!("query optimization succeeded"); optimized.join("\n---\n") } Err(e) => { tracing::warn!("query optimization failed: {}, using original", e); chunks.iter().map(|c| c.text.clone()).collect::>().join("\n---\n") } }; let response = llm.prompt(&context, &question).await?; ``` ### Environment Configuration ```bash # Enable/disable query optimization export MEM_QUERY_OPTIMIZER=on # or "off" # Custom optimizer service (optional) export MEM_QUERY_OPTIMIZER_SERVICE=/path/to/config.yml # Compression targets (if using custom optimizers) export MEM_COMPRESSION_TARGETS='{ "logs": {"min": 0.05, "max": 0.95}, "json": {"min": 0.10, "max": 0.90}, "text": {"min": 0.30, "max": 0.70} }' # Ingest-time optimization export MEM_CONTEXT_OPTIMIZER=on ``` ### Compression Targets by Content Type | Type | Target | Typical | Example | |---|---|---|---| | **Logs** | 85-95% removal | 10-15% remaining | ERROR + timestamps → ERROR only | | **JSON** | 70-90% removal | 10-30% remaining | Minified + key filtering | | **Text/Markdown** | 30-50% removal | 50-70% remaining | Prose kept, formatting removed | | **Code/Diffs** | 60-80% removal | 20-40% remaining | Context lines removed | ### Performance Targets | Metric | Target | Status | |---|---|---| | Ingest latency | <1ms per record | ✅ Passing | | Query latency | <50ms P95 | ✅ Passing | | Compression ratio | Within targets | ✅ Passing | | Graceful fallback | Always succeeds | ✅ Passing | ### Monitoring Track optimization effectiveness via structured logging: ```rust tracing::info!( event = "query_optimization", chunks_count = chunks.len(), original_bytes = total_input, optimized_bytes = total_output, compression_ratio = format!("{:.1}%", ratio), elapsed_ms = elapsed.as_secs_f64() * 1000.0, has_optimizer = query_optimizer.enabled, "query optimization metrics" ); ``` Export to Prometheus (ingest-time metrics): ```bash curl http://localhost:9090/metrics | grep m3_8_optimization ``` ### Best Practices 1. **Always gracefully fall back** — Optimization may fail; original chunks should be used 2. **Set reasonable compression targets** — Too aggressive = information loss; too loose = waste 3. **Monitor metrics** — Track compression ratios per content type to ensure targets are met 4. **Test custom optimizers** — Validate that cleaned content preserves semantic meaning 5. **Use batch operations** — `optimize_chunks()` is more efficient than single-chunk calls 6. **Cache formatter instances** — Create format handlers once, reuse across queries ### Further Reading - [M3.8 Pluggable Optimizer Guide](docs/M3.8-PLUGGABLE-OPTIMIZER.md) — Full architecture details - [M3.8 Completion Summary](CLAUDE_M3.8_COMPLETE.md) — Implementation status - [Query Optimizer Source](crates/mem-core/src/optimizer/query_optimizer.rs) — Implementation code ## Verified environment facts Checked against the running cluster, not assumed: | Fact | Value | |---|---| | Embedding dims | **768**, `nomic-ai/nomic-embed-text-v2-moe` | | Embedding batch limit | **32** (`batch size 1200 > maximum allowed batch size 32`) | | pgvector | **0.7.0 available in stock CNPG image**, no custom build | | CNPG operator | **1.30.0**, declarative `Database.spec.extensions` | | Ollama context cap | **32768** (`OLLAMA_CONTEXT_LENGTH`) — cluster-side, overrides client config | | Controller | `qwen2.5:3b-instruct` — paper's exact 3B backbone | | Gateway auth | `apikey:` header. `Authorization: Bearer` returns **401** | | Rerank response | bare array, not `{"data":[...]}`; sorted by score, map back via `index` | Budget fits the 32K cap: 5000 chunk + ~3200 prompt/memory + 2048 response. ## Phases Each ends in a composition gate. No phase starts until predecessor gate is green. | | Phase | Tasks | Gate asserts | |---|---|---|---| | M0 | Read-only spine | 8 | third source needs no downstream change; runs offline | | M1 | Gated loop at L1 | 8 | **update-rate < 30%**, memory flat not climbing | | M2 | Projections | 8 | rebuild byte-identical from log alone | | M3 | L2 + retrieval | 4 | hit rate ≥ 0.8, provenance precision ≥ 0.9 | | M4 | Skills | 3 | draft not loadable; promoted skill never becomes evidence | | M5 | Post-training | 6 | adapter beats prompted baseline on held-out project | **Update-rate is the number to watch.** It is what distinguishes a gate from an expensive summarizer. Tool results are 43% of records and mostly evidence-free, so a correct gate rejects the large majority of chunks. M0 and M2.2 need no model access and can start immediately. M5.4 (vLLM + LoRA) is homelab work independent of the rest of M5. ## Reading order 1. This file 2. [memory-tasks/INDEX.md](memory-tasks/INDEX.md) — board, ordering rules, verification practice 3. [DESIGN.md](DESIGN.md) — full design, schemas, risks 4. Individual task files — self-contained, no DESIGN.md read required # Trigger build run 130 # CI trigger