Files
poimen-memory/README.md
T
Story Crater Bot 57c434ccdd
Build and Push / Test (push) Failing after 1m55s
Build and Push / Build and push image (push) Skipped
docs: comprehensive query optimization guides for developers
Added two major documentation pieces:

1. README.md - New Section: M3.8 Pluggable Query Optimization
    Architecture overview (ingest + query paths)
    6 practical usage patterns with code examples:
      - Basic query with auto-optimization
      - Prompt construction with optimization
      - Custom optimizer implementation
      - Optimized query with metrics tracking
      - Batch optimization for multiple queries
      - Conditional optimization with graceful fallback
    Environment configuration
    Compression targets by content type
    Performance targets table
    Monitoring via structured logging
    Best practices (5 key points)
    Links to full documentation

2. QUERY-OPTIMIZATION-COOKBOOK.md - Quick Reference (15KB)
    Basic usage patterns
    Prompt construction techniques
    Custom optimizer examples:
      - Content-type specific (Python optimizer)
      - Domain-specific (Medical optimizer)
      - Semantic pruning
    Format handlers (built-in + custom Gzip example)
    Error handling (graceful fallback + retry)
    Testing patterns (unit, integration, mocking)
    Configuration examples (env vars + Kubernetes)
    Performance tips (5 optimization strategies)
    Debugging guide

Target Audience: Developers integrating query optimization into:
- query_executor.rs
- hybrid_query_worker.rs
- Custom LLM clients

Includes:
- Copy-paste ready code examples
- Real-world patterns for medical, code, text optimization
- Testing strategies
- Kubernetes deployment config
- Debug logging setup
- Performance profiling tips
2026-08-28 12:32:08 -07:00

512 lines
17 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# poimen-memory
Gated recurrent memory over agent context. Reads session history chunk-by-chunk,
keeps only what answers standing questions, projects result into an Obsidian
vault and a pgvector index.
**Status: design complete, no code yet.** 37 tasks in [memory-tasks/](memory-tasks/INDEX.md),
0 done. Start at [M0.1](memory-tasks/M0.1-cargo-workspace.md).
## Problem
Agent sessions grow faster than anyone reads them, and most of the volume is
noise. One real pi session in this project:
```
assistant 1445
toolResult 1261 43% — ls output, file reads, mostly evidence-free
user 196
+ 8 compaction events
```
Compaction fires 8 times per session. Context gets *discarded*, not retained —
root causes, decisions and gotchas evaporate when window rolls.
## Mechanism
GRU-Mem ([arXiv 2602.10560](https://arxiv.org/abs/2602.10560)). Two text-controlled
gates on a recurrent memory loop:
- **update gate** — memory only mutates when chunk contains evidence. Blocks the
memory explosion that ungated recurrent memory hits.
- **exit gate** — stop scanning once evidence sufficient.
Paper reports up to 400% speedup and *better* accuracy than ungated, because
unbounded memory growth degrades later updates.
```
sessions ─> chunk (5000 tok) ─> controller ─> gates ─> memory ─> projections
```
Controller emits structured output; loop acts on it:
```
<think> reason about chunk vs question
<check> yes|no -> U_t, update or discard
<update> candidate memory M̂_t
<next> continue|end -> E_t, exit or continue
```
## Memory tiers
| Level | What | From | Bounded |
|---|---|---|---|
| **L0** | evidence chunk, verbatim | update gate opening | no, but sparse (~17 of 412) |
| **L1** | per-query memory, `M_t` | gated loop over chunks | 1024 tok |
| **L2** | project synthesis | gated loop over L1 memories | 1024 tok |
L2 is not new machinery — same loop, same prompt, L1 memories as input stream.
Level is a parameter.
Tiers form a provenance graph. Each L1 records its L0 parents, each L2 its L1
parents. Same relation becomes both `memory_edge` rows and Obsidian wikilinks.
## Standing queries
Update gate needs a referent. Paper's agent is `φθ(Q, C_t, M_{t-1})` — gate is
defined as "does this chunk contain useful information *about the problem*". No
`Q`, no gate, and `r_update` becomes undefinable, which kills post-training.
So each project declares durable questions. One query = one L1 memory = one note.
```yaml
# queries/poimen.yaml
project: poimen
roots: [/Users/rockliang/workplace/Poimen/agent-rust]
queries:
- id: infra-root-causes
question: What infrastructure bugs were found, what was the root cause, how was it isolated?
- id: architecture-decisions
question: What architectural decisions were made, with reasoning and rejected alternatives?
synthesis:
question: What is the current state of this project, and what should someone know before working on it?
exit_gate: true
```
**Exit gate off at L1, on at L2.** Paper §3.3 makes this call: for "what are *all*
the X" questions you cannot know evidence is sufficient without reading
everything. L1 extraction is that shape. At L2 input is a handful of memories and
sufficiency is decidable. Gate still *recorded* at L1 — signal needed for
post-training.
## Authority model
**JSONL log authoritative. Vault and vector index are projections.**
Anything not rebuildable byte-identically from the log has hidden inputs, and
that is a bug. Gate M2.8 enforces it destructively:
```sh
rm -rf vault/poimen
psql -c "delete from memory_node where project='poimen'"
mem rebuild --from-log --project poimen
git -C vault diff --exit-code # empty diff is the only pass
```
Buys three things: re-embedding after model change is a rebuild not a migration,
Obsidian edits cannot corrupt the record, post-training corpus is the log itself.
## Skills
A skill is a **projection, not a level**. L0/L1/L2 are descriptive — what
happened. A skill is procedural — what to do next time. Gated loop does not
produce it.
Format free: `SKILL.md` is YAML frontmatter + markdown, which is an Obsidian
note. So `vault/skills/<name>/SKILL.md` is both, no conversion:
```sh
pi --skill vault/skills/
ln -s .../vault/skills/<name> ~/.claude/skills/<name>
```
**Drafts land in `_drafts/`, promotion is a human `git mv`.** This is the one
cycle in the design:
```
emitted skill auto-loads -> appears in future transcripts
-> ingested as evidence -> reinforces the memory that emitted it
```
No external verifier breaks it. Two guards: `_drafts/` is a directory (cannot be
globbed into `--skill`), and every artifact carries `generated_from` so ingest
tags matching chunks `derived: true` and refuses them as evidence.
## Separate weights
Memory policy is a **LoRA adapter** on Qwen2.5-3B-Instruct, not a fine-tuned
model. Reason is VRAM: one GPU, `OLLAMA_MAX_LOADED_MODELS=2`, already holding
`ornith:35b` + `qwen2.5:3b`. Separate full model evicts something, and eviction
is a weights reload measured in tens of seconds. Adapter rides the resident base.
Also: post-training emits ~50 MB, not 6 GB. Swap without redeploy. Regression
reverts by pointing at previous adapter.
**Ollama cannot hot-swap LoRA.** Serving one needs vLLM with `--enable-lora`
(pattern already exists — `reasoning` predictor is vLLM v0.11.0). Phases M0M4
run prompted-only, so decision is deferred, not dodged.
## Layout
```
DESIGN.md full design, 460 lines
memory-tasks/ 37 task files + INDEX.md — tracked
crates/
mem-core/ domain types; Level; gate parser; the gated loop
mem-chunk/ RecordSource trait; ChunkPolicy; FlushTrigger
mem-llm/ gateway client — chat, embeddings, rerank
mem-ingest/ source adapters: pi sessions, claude transcripts
mem-store/ JSONL log; pgvector repo; Obsidian projector
mem-cli/ binary `mem`
queries/ standing query YAML per project
log/ JSONL event log — authoritative, tracked
vault/ Obsidian output
```
`mem-chunk` is separate and stream-shaped from day one. Sources today are files
with an EOF; telemetry or a live tail will not have one. `RecordSource` returns
`impl Stream<Item = Record>`; batch sources become streams via
`futures::stream::iter`, so it costs nothing now and removes a rewrite later.
## Commands
```sh
mem ingest --project poimen --dry-run # chunk plan, zero model calls
mem ingest --project poimen --query infra-root-causes
mem synthesize --project poimen # L2 pass, exit gate on
mem rebuild --from-log --project poimen # drop and rebuild projections
mem verify --project poimen # provenance graph closure
mem query "why did requests over 10KB fail?"
mem skill draft --from poimen/infra-root-causes
mem label --project poimen # evidence labels for training
```
## M3.8 Pluggable Query Optimization
**Purpose**: Compress and optimize search results before passing them to the LLM context window, improving token efficiency and response quality.
### Architecture
M3.8 provides **dual-path optimization**:
#### Ingest-Time Optimization (M3.8.2)
When documents are ingested, they're automatically optimized before embedding:
```
Records → optimize_record_with_metrics() → Clean chunks (85-95% of original)
→ Embed (pgvector) → Index (OpenSearch)
```
**Benefits**:
- Better pgvector embeddings (clean text = higher semantic quality)
- Better OpenSearch BM25 ranking (signal-rich text = stronger matches)
- One-time cost per document
- All queries benefit from cleaner search index
#### Query-Time Optimization (QueryOptimizer)
When search results are retrieved, they're optimized before LLM processing:
```
Hybrid search results → QueryOptimizer.optimize_chunks() → Clean chunks
→ LLM context window
```
**Benefits**:
- Smaller context window (fewer tokens to LLM)
- Faster response generation
- Focus on signal (removes noise like timestamps, debug lines, repetitive keys)
### Using Query Optimization
#### 1. Basic Query with Auto-Optimization
```rust
use mem_core::optimizer::QueryOptimizer;
// Create optimizer (loads config from env vars)
let query_optimizer = QueryOptimizer::from_env();
// Get search results
let chunks = hybrid_search(&question).await?;
// Auto-optimize before LLM
let optimized = query_optimizer.optimize_chunks(&chunks).await?;
// Build context from clean chunks
let context = optimized.join("\n---\n");
let response = llm.prompt(&context, &question).await?;
```
#### 2. Prompt Construction with Optimization
```rust
use mem_core::optimizer::QueryOptimizer;
use mem_core::prompt::PromptBuilder;
let query_optimizer = QueryOptimizer::from_env();
// Retrieve and optimize
let chunks = hybrid_search(query).await?;
let optimized_chunks = query_optimizer.optimize_chunks(&chunks).await?;
// Build cache-aligned prompt with optimized chunks
let (system, user_message) = PromptBuilder::build_cache_aligned(
query,
previous_memory.as_deref(),
/* use optimized chunks */
)?;
let response = llm.prompt(system, user_message).await?;
```
#### 3. Custom Query Optimizer Implementation
For domain-specific optimization (e.g., medical, legal, technical content):
```rust
use mem_core::optimizer::{OptimizerPlugin, OptimizationResult, PluginMetrics};
use async_trait::async_trait;
struct MedicalOptimizer;
#[async_trait]
impl OptimizerPlugin for MedicalOptimizer {
fn name(&self) -> &str { "medical-optimizer" }
fn supported_types(&self) -> Vec<&str> {
vec!["text/medical", "text/clinical", "application/json"]
}
async fn optimize(&self, content: &str) -> Result<OptimizationResult, String> {
// Remove patient IDs, reduce duplicate diagnosis entries
let cleaned = clean_medical_data(content);
let ratio = cleaned.len() as f32 / content.len() as f32;
Ok(OptimizationResult {
original: content.to_string(),
optimized: cleaned,
ratio,
plugin: self.name().to_string(),
metadata: Default::default(),
})
}
fn metrics(&self) -> PluginMetrics { Default::default() }
}
// Register and use
let service = OptimizerServiceBuilder::new()
.with_optimizer(Arc::new(MedicalOptimizer))
.with_format(Arc::new(JsonFormatter))
.build()?;
let optimized = service.optimize(
clinical_note,
"text/clinical",
None
).await?;
```
#### 4. Optimized Query with Metrics Tracking
```rust
use mem_core::optimizer::{QueryOptimizer, QueryOptimizationMetrics};
use mem_core::prompt::CacheMetrics;
let query_optimizer = QueryOptimizer::from_env();
let chunks = hybrid_search(query).await?;
let optimized = query_optimizer.optimize_chunks(&chunks).await?;
// Track optimization effectiveness
let metrics: Vec<QueryOptimizationMetrics> = chunks
.iter()
.zip(&optimized)
.map(|(orig, opt)| {
QueryOptimizationMetrics {
original_bytes: orig.text.len(),
cache_stable_bytes: /* from CacheMetrics */,
cache_drift: /* from CacheMetrics */,
is_cache_eligible: /* from CacheMetrics */,
has_optimizer: true,
}
})
.collect();
tracing::info!(
chunks = chunks.len(),
compression_ratio = format!(
"{:.1}%",
(optimized.iter().map(|o| o.len()).sum::<usize>() as f32
/ chunks.iter().map(|c| c.text.len()).sum::<usize>() as f32) * 100.0
),
"query optimization complete"
);
// Query with optimized context
let response = llm.prompt(&optimized.join("\n---\n"), &question).await?;
```
#### 5. Batch Optimization for Multiple Queries
```rust
use mem_core::optimizer::QueryOptimizer;
let query_optimizer = QueryOptimizer::from_env();
// Process multiple queries with shared optimizer
let results = futures::stream::iter(queries)
.then(|query| async move {
let chunks = hybrid_search(&query).await?;
let optimized = query_optimizer.optimize_chunks(&chunks).await?;
let response = llm.prompt(&optimized.join("\n---\n"), &query.question).await?;
Ok((query, response))
})
.collect::<Vec<_>>()
.await;
```
#### 6. Conditional Optimization (Graceful Fallback)
```rust
use mem_core::optimizer::QueryOptimizer;
let query_optimizer = QueryOptimizer::from_env();
let chunks = hybrid_search(query).await?;
// Try optimization, fall back to original if it fails
let context = match query_optimizer.optimize_chunks(&chunks).await {
Ok(optimized) => {
tracing::info!("query optimization succeeded");
optimized.join("\n---\n")
}
Err(e) => {
tracing::warn!("query optimization failed: {}, using original", e);
chunks.iter().map(|c| c.text.clone()).collect::<Vec<_>>().join("\n---\n")
}
};
let response = llm.prompt(&context, &question).await?;
```
### Environment Configuration
```bash
# Enable/disable query optimization
export MEM_QUERY_OPTIMIZER=on # or "off"
# Custom optimizer service (optional)
export MEM_QUERY_OPTIMIZER_SERVICE=/path/to/config.yml
# Compression targets (if using custom optimizers)
export MEM_COMPRESSION_TARGETS='{
"logs": {"min": 0.05, "max": 0.95},
"json": {"min": 0.10, "max": 0.90},
"text": {"min": 0.30, "max": 0.70}
}'
# Ingest-time optimization
export MEM_CONTEXT_OPTIMIZER=on
```
### Compression Targets by Content Type
| Type | Target | Typical | Example |
|---|---|---|---|
| **Logs** | 85-95% removal | 10-15% remaining | ERROR + timestamps → ERROR only |
| **JSON** | 70-90% removal | 10-30% remaining | Minified + key filtering |
| **Text/Markdown** | 30-50% removal | 50-70% remaining | Prose kept, formatting removed |
| **Code/Diffs** | 60-80% removal | 20-40% remaining | Context lines removed |
### Performance Targets
| Metric | Target | Status |
|---|---|---|
| Ingest latency | <1ms per record | ✅ Passing |
| Query latency | <50ms P95 | ✅ Passing |
| Compression ratio | Within targets | ✅ Passing |
| Graceful fallback | Always succeeds | ✅ Passing |
### Monitoring
Track optimization effectiveness via structured logging:
```rust
tracing::info!(
event = "query_optimization",
chunks_count = chunks.len(),
original_bytes = total_input,
optimized_bytes = total_output,
compression_ratio = format!("{:.1}%", ratio),
elapsed_ms = elapsed.as_secs_f64() * 1000.0,
has_optimizer = query_optimizer.enabled,
"query optimization metrics"
);
```
Export to Prometheus (ingest-time metrics):
```bash
curl http://localhost:9090/metrics | grep m3_8_optimization
```
### Best Practices
1. **Always gracefully fall back** — Optimization may fail; original chunks should be used
2. **Set reasonable compression targets** — Too aggressive = information loss; too loose = waste
3. **Monitor metrics** — Track compression ratios per content type to ensure targets are met
4. **Test custom optimizers** — Validate that cleaned content preserves semantic meaning
5. **Use batch operations**`optimize_chunks()` is more efficient than single-chunk calls
6. **Cache formatter instances** — Create format handlers once, reuse across queries
### Further Reading
- [M3.8 Pluggable Optimizer Guide](docs/M3.8-PLUGGABLE-OPTIMIZER.md) — Full architecture details
- [M3.8 Completion Summary](CLAUDE_M3.8_COMPLETE.md) — Implementation status
- [Query Optimizer Source](crates/mem-core/src/optimizer/query_optimizer.rs) — Implementation code
## Verified environment facts
Checked against the running cluster, not assumed:
| Fact | Value |
|---|---|
| Embedding dims | **768**, `nomic-ai/nomic-embed-text-v2-moe` |
| Embedding batch limit | **32** (`batch size 1200 > maximum allowed batch size 32`) |
| pgvector | **0.7.0 available in stock CNPG image**, no custom build |
| CNPG operator | **1.30.0**, declarative `Database.spec.extensions` |
| Ollama context cap | **32768** (`OLLAMA_CONTEXT_LENGTH`) — cluster-side, overrides client config |
| Controller | `qwen2.5:3b-instruct` — paper's exact 3B backbone |
| Gateway auth | `apikey:` header. `Authorization: Bearer` returns **401** |
| Rerank response | bare array, not `{"data":[...]}`; sorted by score, map back via `index` |
Budget fits the 32K cap: 5000 chunk + ~3200 prompt/memory + 2048 response.
## Phases
Each ends in a composition gate. No phase starts until predecessor gate is green.
| | Phase | Tasks | Gate asserts |
|---|---|---|---|
| M0 | Read-only spine | 8 | third source needs no downstream change; runs offline |
| M1 | Gated loop at L1 | 8 | **update-rate < 30%**, memory flat not climbing |
| M2 | Projections | 8 | rebuild byte-identical from log alone |
| M3 | L2 + retrieval | 4 | hit rate ≥ 0.8, provenance precision ≥ 0.9 |
| M4 | Skills | 3 | draft not loadable; promoted skill never becomes evidence |
| M5 | Post-training | 6 | adapter beats prompted baseline on held-out project |
**Update-rate is the number to watch.** It is what distinguishes a gate from an
expensive summarizer. Tool results are 43% of records and mostly evidence-free,
so a correct gate rejects the large majority of chunks.
M0 and M2.2 need no model access and can start immediately. M5.4 (vLLM + LoRA)
is homelab work independent of the rest of M5.
## Reading order
1. This file
2. [memory-tasks/INDEX.md](memory-tasks/INDEX.md) — board, ordering rules, verification practice
3. [DESIGN.md](DESIGN.md) — full design, schemas, risks
4. Individual task files — self-contained, no DESIGN.md read required
# Trigger build run 130
# CI trigger