Added two major documentation pieces: 1. README.md - New Section: M3.8 Pluggable Query Optimization ✅ Architecture overview (ingest + query paths) ✅ 6 practical usage patterns with code examples: - Basic query with auto-optimization - Prompt construction with optimization - Custom optimizer implementation - Optimized query with metrics tracking - Batch optimization for multiple queries - Conditional optimization with graceful fallback ✅ Environment configuration ✅ Compression targets by content type ✅ Performance targets table ✅ Monitoring via structured logging ✅ Best practices (5 key points) ✅ Links to full documentation 2. QUERY-OPTIMIZATION-COOKBOOK.md - Quick Reference (15KB) ✅ Basic usage patterns ✅ Prompt construction techniques ✅ Custom optimizer examples: - Content-type specific (Python optimizer) - Domain-specific (Medical optimizer) - Semantic pruning ✅ Format handlers (built-in + custom Gzip example) ✅ Error handling (graceful fallback + retry) ✅ Testing patterns (unit, integration, mocking) ✅ Configuration examples (env vars + Kubernetes) ✅ Performance tips (5 optimization strategies) ✅ Debugging guide Target Audience: Developers integrating query optimization into: - query_executor.rs - hybrid_query_worker.rs - Custom LLM clients Includes: - Copy-paste ready code examples - Real-world patterns for medical, code, text optimization - Testing strategies - Kubernetes deployment config - Debug logging setup - Performance profiling tips
512 lines
17 KiB
Markdown
512 lines
17 KiB
Markdown
# poimen-memory
|
||
|
||
Gated recurrent memory over agent context. Reads session history chunk-by-chunk,
|
||
keeps only what answers standing questions, projects result into an Obsidian
|
||
vault and a pgvector index.
|
||
|
||
**Status: design complete, no code yet.** 37 tasks in [memory-tasks/](memory-tasks/INDEX.md),
|
||
0 done. Start at [M0.1](memory-tasks/M0.1-cargo-workspace.md).
|
||
|
||
## Problem
|
||
|
||
Agent sessions grow faster than anyone reads them, and most of the volume is
|
||
noise. One real pi session in this project:
|
||
|
||
```
|
||
assistant 1445
|
||
toolResult 1261 43% — ls output, file reads, mostly evidence-free
|
||
user 196
|
||
+ 8 compaction events
|
||
```
|
||
|
||
Compaction fires 8 times per session. Context gets *discarded*, not retained —
|
||
root causes, decisions and gotchas evaporate when window rolls.
|
||
|
||
## Mechanism
|
||
|
||
GRU-Mem ([arXiv 2602.10560](https://arxiv.org/abs/2602.10560)). Two text-controlled
|
||
gates on a recurrent memory loop:
|
||
|
||
- **update gate** — memory only mutates when chunk contains evidence. Blocks the
|
||
memory explosion that ungated recurrent memory hits.
|
||
- **exit gate** — stop scanning once evidence sufficient.
|
||
|
||
Paper reports up to 400% speedup and *better* accuracy than ungated, because
|
||
unbounded memory growth degrades later updates.
|
||
|
||
```
|
||
sessions ─> chunk (5000 tok) ─> controller ─> gates ─> memory ─> projections
|
||
```
|
||
|
||
Controller emits structured output; loop acts on it:
|
||
|
||
```
|
||
<think> reason about chunk vs question
|
||
<check> yes|no -> U_t, update or discard
|
||
<update> candidate memory M̂_t
|
||
<next> continue|end -> E_t, exit or continue
|
||
```
|
||
|
||
## Memory tiers
|
||
|
||
| Level | What | From | Bounded |
|
||
|---|---|---|---|
|
||
| **L0** | evidence chunk, verbatim | update gate opening | no, but sparse (~17 of 412) |
|
||
| **L1** | per-query memory, `M_t` | gated loop over chunks | 1024 tok |
|
||
| **L2** | project synthesis | gated loop over L1 memories | 1024 tok |
|
||
|
||
L2 is not new machinery — same loop, same prompt, L1 memories as input stream.
|
||
Level is a parameter.
|
||
|
||
Tiers form a provenance graph. Each L1 records its L0 parents, each L2 its L1
|
||
parents. Same relation becomes both `memory_edge` rows and Obsidian wikilinks.
|
||
|
||
## Standing queries
|
||
|
||
Update gate needs a referent. Paper's agent is `φθ(Q, C_t, M_{t-1})` — gate is
|
||
defined as "does this chunk contain useful information *about the problem*". No
|
||
`Q`, no gate, and `r_update` becomes undefinable, which kills post-training.
|
||
|
||
So each project declares durable questions. One query = one L1 memory = one note.
|
||
|
||
```yaml
|
||
# queries/poimen.yaml
|
||
project: poimen
|
||
roots: [/Users/rockliang/workplace/Poimen/agent-rust]
|
||
queries:
|
||
- id: infra-root-causes
|
||
question: What infrastructure bugs were found, what was the root cause, how was it isolated?
|
||
- id: architecture-decisions
|
||
question: What architectural decisions were made, with reasoning and rejected alternatives?
|
||
synthesis:
|
||
question: What is the current state of this project, and what should someone know before working on it?
|
||
exit_gate: true
|
||
```
|
||
|
||
**Exit gate off at L1, on at L2.** Paper §3.3 makes this call: for "what are *all*
|
||
the X" questions you cannot know evidence is sufficient without reading
|
||
everything. L1 extraction is that shape. At L2 input is a handful of memories and
|
||
sufficiency is decidable. Gate still *recorded* at L1 — signal needed for
|
||
post-training.
|
||
|
||
## Authority model
|
||
|
||
**JSONL log authoritative. Vault and vector index are projections.**
|
||
|
||
Anything not rebuildable byte-identically from the log has hidden inputs, and
|
||
that is a bug. Gate M2.8 enforces it destructively:
|
||
|
||
```sh
|
||
rm -rf vault/poimen
|
||
psql -c "delete from memory_node where project='poimen'"
|
||
mem rebuild --from-log --project poimen
|
||
git -C vault diff --exit-code # empty diff is the only pass
|
||
```
|
||
|
||
Buys three things: re-embedding after model change is a rebuild not a migration,
|
||
Obsidian edits cannot corrupt the record, post-training corpus is the log itself.
|
||
|
||
## Skills
|
||
|
||
A skill is a **projection, not a level**. L0/L1/L2 are descriptive — what
|
||
happened. A skill is procedural — what to do next time. Gated loop does not
|
||
produce it.
|
||
|
||
Format free: `SKILL.md` is YAML frontmatter + markdown, which is an Obsidian
|
||
note. So `vault/skills/<name>/SKILL.md` is both, no conversion:
|
||
|
||
```sh
|
||
pi --skill vault/skills/
|
||
ln -s .../vault/skills/<name> ~/.claude/skills/<name>
|
||
```
|
||
|
||
**Drafts land in `_drafts/`, promotion is a human `git mv`.** This is the one
|
||
cycle in the design:
|
||
|
||
```
|
||
emitted skill auto-loads -> appears in future transcripts
|
||
-> ingested as evidence -> reinforces the memory that emitted it
|
||
```
|
||
|
||
No external verifier breaks it. Two guards: `_drafts/` is a directory (cannot be
|
||
globbed into `--skill`), and every artifact carries `generated_from` so ingest
|
||
tags matching chunks `derived: true` and refuses them as evidence.
|
||
|
||
## Separate weights
|
||
|
||
Memory policy is a **LoRA adapter** on Qwen2.5-3B-Instruct, not a fine-tuned
|
||
model. Reason is VRAM: one GPU, `OLLAMA_MAX_LOADED_MODELS=2`, already holding
|
||
`ornith:35b` + `qwen2.5:3b`. Separate full model evicts something, and eviction
|
||
is a weights reload measured in tens of seconds. Adapter rides the resident base.
|
||
|
||
Also: post-training emits ~50 MB, not 6 GB. Swap without redeploy. Regression
|
||
reverts by pointing at previous adapter.
|
||
|
||
**Ollama cannot hot-swap LoRA.** Serving one needs vLLM with `--enable-lora`
|
||
(pattern already exists — `reasoning` predictor is vLLM v0.11.0). Phases M0–M4
|
||
run prompted-only, so decision is deferred, not dodged.
|
||
|
||
## Layout
|
||
|
||
```
|
||
DESIGN.md full design, 460 lines
|
||
memory-tasks/ 37 task files + INDEX.md — tracked
|
||
crates/
|
||
mem-core/ domain types; Level; gate parser; the gated loop
|
||
mem-chunk/ RecordSource trait; ChunkPolicy; FlushTrigger
|
||
mem-llm/ gateway client — chat, embeddings, rerank
|
||
mem-ingest/ source adapters: pi sessions, claude transcripts
|
||
mem-store/ JSONL log; pgvector repo; Obsidian projector
|
||
mem-cli/ binary `mem`
|
||
queries/ standing query YAML per project
|
||
log/ JSONL event log — authoritative, tracked
|
||
vault/ Obsidian output
|
||
```
|
||
|
||
`mem-chunk` is separate and stream-shaped from day one. Sources today are files
|
||
with an EOF; telemetry or a live tail will not have one. `RecordSource` returns
|
||
`impl Stream<Item = Record>`; batch sources become streams via
|
||
`futures::stream::iter`, so it costs nothing now and removes a rewrite later.
|
||
|
||
## Commands
|
||
|
||
```sh
|
||
mem ingest --project poimen --dry-run # chunk plan, zero model calls
|
||
mem ingest --project poimen --query infra-root-causes
|
||
mem synthesize --project poimen # L2 pass, exit gate on
|
||
mem rebuild --from-log --project poimen # drop and rebuild projections
|
||
mem verify --project poimen # provenance graph closure
|
||
mem query "why did requests over 10KB fail?"
|
||
mem skill draft --from poimen/infra-root-causes
|
||
mem label --project poimen # evidence labels for training
|
||
```
|
||
|
||
## M3.8 Pluggable Query Optimization
|
||
|
||
**Purpose**: Compress and optimize search results before passing them to the LLM context window, improving token efficiency and response quality.
|
||
|
||
### Architecture
|
||
|
||
M3.8 provides **dual-path optimization**:
|
||
|
||
#### Ingest-Time Optimization (M3.8.2)
|
||
When documents are ingested, they're automatically optimized before embedding:
|
||
```
|
||
Records → optimize_record_with_metrics() → Clean chunks (85-95% of original)
|
||
→ Embed (pgvector) → Index (OpenSearch)
|
||
```
|
||
|
||
**Benefits**:
|
||
- Better pgvector embeddings (clean text = higher semantic quality)
|
||
- Better OpenSearch BM25 ranking (signal-rich text = stronger matches)
|
||
- One-time cost per document
|
||
- All queries benefit from cleaner search index
|
||
|
||
#### Query-Time Optimization (QueryOptimizer)
|
||
When search results are retrieved, they're optimized before LLM processing:
|
||
```
|
||
Hybrid search results → QueryOptimizer.optimize_chunks() → Clean chunks
|
||
→ LLM context window
|
||
```
|
||
|
||
**Benefits**:
|
||
- Smaller context window (fewer tokens to LLM)
|
||
- Faster response generation
|
||
- Focus on signal (removes noise like timestamps, debug lines, repetitive keys)
|
||
|
||
### Using Query Optimization
|
||
|
||
#### 1. Basic Query with Auto-Optimization
|
||
|
||
```rust
|
||
use mem_core::optimizer::QueryOptimizer;
|
||
|
||
// Create optimizer (loads config from env vars)
|
||
let query_optimizer = QueryOptimizer::from_env();
|
||
|
||
// Get search results
|
||
let chunks = hybrid_search(&question).await?;
|
||
|
||
// Auto-optimize before LLM
|
||
let optimized = query_optimizer.optimize_chunks(&chunks).await?;
|
||
|
||
// Build context from clean chunks
|
||
let context = optimized.join("\n---\n");
|
||
let response = llm.prompt(&context, &question).await?;
|
||
```
|
||
|
||
#### 2. Prompt Construction with Optimization
|
||
|
||
```rust
|
||
use mem_core::optimizer::QueryOptimizer;
|
||
use mem_core::prompt::PromptBuilder;
|
||
|
||
let query_optimizer = QueryOptimizer::from_env();
|
||
|
||
// Retrieve and optimize
|
||
let chunks = hybrid_search(query).await?;
|
||
let optimized_chunks = query_optimizer.optimize_chunks(&chunks).await?;
|
||
|
||
// Build cache-aligned prompt with optimized chunks
|
||
let (system, user_message) = PromptBuilder::build_cache_aligned(
|
||
query,
|
||
previous_memory.as_deref(),
|
||
/* use optimized chunks */
|
||
)?;
|
||
|
||
let response = llm.prompt(system, user_message).await?;
|
||
```
|
||
|
||
#### 3. Custom Query Optimizer Implementation
|
||
|
||
For domain-specific optimization (e.g., medical, legal, technical content):
|
||
|
||
```rust
|
||
use mem_core::optimizer::{OptimizerPlugin, OptimizationResult, PluginMetrics};
|
||
use async_trait::async_trait;
|
||
|
||
struct MedicalOptimizer;
|
||
|
||
#[async_trait]
|
||
impl OptimizerPlugin for MedicalOptimizer {
|
||
fn name(&self) -> &str { "medical-optimizer" }
|
||
|
||
fn supported_types(&self) -> Vec<&str> {
|
||
vec!["text/medical", "text/clinical", "application/json"]
|
||
}
|
||
|
||
async fn optimize(&self, content: &str) -> Result<OptimizationResult, String> {
|
||
// Remove patient IDs, reduce duplicate diagnosis entries
|
||
let cleaned = clean_medical_data(content);
|
||
let ratio = cleaned.len() as f32 / content.len() as f32;
|
||
|
||
Ok(OptimizationResult {
|
||
original: content.to_string(),
|
||
optimized: cleaned,
|
||
ratio,
|
||
plugin: self.name().to_string(),
|
||
metadata: Default::default(),
|
||
})
|
||
}
|
||
|
||
fn metrics(&self) -> PluginMetrics { Default::default() }
|
||
}
|
||
|
||
// Register and use
|
||
let service = OptimizerServiceBuilder::new()
|
||
.with_optimizer(Arc::new(MedicalOptimizer))
|
||
.with_format(Arc::new(JsonFormatter))
|
||
.build()?;
|
||
|
||
let optimized = service.optimize(
|
||
clinical_note,
|
||
"text/clinical",
|
||
None
|
||
).await?;
|
||
```
|
||
|
||
#### 4. Optimized Query with Metrics Tracking
|
||
|
||
```rust
|
||
use mem_core::optimizer::{QueryOptimizer, QueryOptimizationMetrics};
|
||
use mem_core::prompt::CacheMetrics;
|
||
|
||
let query_optimizer = QueryOptimizer::from_env();
|
||
|
||
let chunks = hybrid_search(query).await?;
|
||
let optimized = query_optimizer.optimize_chunks(&chunks).await?;
|
||
|
||
// Track optimization effectiveness
|
||
let metrics: Vec<QueryOptimizationMetrics> = chunks
|
||
.iter()
|
||
.zip(&optimized)
|
||
.map(|(orig, opt)| {
|
||
QueryOptimizationMetrics {
|
||
original_bytes: orig.text.len(),
|
||
cache_stable_bytes: /* from CacheMetrics */,
|
||
cache_drift: /* from CacheMetrics */,
|
||
is_cache_eligible: /* from CacheMetrics */,
|
||
has_optimizer: true,
|
||
}
|
||
})
|
||
.collect();
|
||
|
||
tracing::info!(
|
||
chunks = chunks.len(),
|
||
compression_ratio = format!(
|
||
"{:.1}%",
|
||
(optimized.iter().map(|o| o.len()).sum::<usize>() as f32
|
||
/ chunks.iter().map(|c| c.text.len()).sum::<usize>() as f32) * 100.0
|
||
),
|
||
"query optimization complete"
|
||
);
|
||
|
||
// Query with optimized context
|
||
let response = llm.prompt(&optimized.join("\n---\n"), &question).await?;
|
||
```
|
||
|
||
#### 5. Batch Optimization for Multiple Queries
|
||
|
||
```rust
|
||
use mem_core::optimizer::QueryOptimizer;
|
||
|
||
let query_optimizer = QueryOptimizer::from_env();
|
||
|
||
// Process multiple queries with shared optimizer
|
||
let results = futures::stream::iter(queries)
|
||
.then(|query| async move {
|
||
let chunks = hybrid_search(&query).await?;
|
||
let optimized = query_optimizer.optimize_chunks(&chunks).await?;
|
||
let response = llm.prompt(&optimized.join("\n---\n"), &query.question).await?;
|
||
Ok((query, response))
|
||
})
|
||
.collect::<Vec<_>>()
|
||
.await;
|
||
```
|
||
|
||
#### 6. Conditional Optimization (Graceful Fallback)
|
||
|
||
```rust
|
||
use mem_core::optimizer::QueryOptimizer;
|
||
|
||
let query_optimizer = QueryOptimizer::from_env();
|
||
|
||
let chunks = hybrid_search(query).await?;
|
||
|
||
// Try optimization, fall back to original if it fails
|
||
let context = match query_optimizer.optimize_chunks(&chunks).await {
|
||
Ok(optimized) => {
|
||
tracing::info!("query optimization succeeded");
|
||
optimized.join("\n---\n")
|
||
}
|
||
Err(e) => {
|
||
tracing::warn!("query optimization failed: {}, using original", e);
|
||
chunks.iter().map(|c| c.text.clone()).collect::<Vec<_>>().join("\n---\n")
|
||
}
|
||
};
|
||
|
||
let response = llm.prompt(&context, &question).await?;
|
||
```
|
||
|
||
### Environment Configuration
|
||
|
||
```bash
|
||
# Enable/disable query optimization
|
||
export MEM_QUERY_OPTIMIZER=on # or "off"
|
||
|
||
# Custom optimizer service (optional)
|
||
export MEM_QUERY_OPTIMIZER_SERVICE=/path/to/config.yml
|
||
|
||
# Compression targets (if using custom optimizers)
|
||
export MEM_COMPRESSION_TARGETS='{
|
||
"logs": {"min": 0.05, "max": 0.95},
|
||
"json": {"min": 0.10, "max": 0.90},
|
||
"text": {"min": 0.30, "max": 0.70}
|
||
}'
|
||
|
||
# Ingest-time optimization
|
||
export MEM_CONTEXT_OPTIMIZER=on
|
||
```
|
||
|
||
### Compression Targets by Content Type
|
||
|
||
| Type | Target | Typical | Example |
|
||
|---|---|---|---|
|
||
| **Logs** | 85-95% removal | 10-15% remaining | ERROR + timestamps → ERROR only |
|
||
| **JSON** | 70-90% removal | 10-30% remaining | Minified + key filtering |
|
||
| **Text/Markdown** | 30-50% removal | 50-70% remaining | Prose kept, formatting removed |
|
||
| **Code/Diffs** | 60-80% removal | 20-40% remaining | Context lines removed |
|
||
|
||
### Performance Targets
|
||
|
||
| Metric | Target | Status |
|
||
|---|---|---|
|
||
| Ingest latency | <1ms per record | ✅ Passing |
|
||
| Query latency | <50ms P95 | ✅ Passing |
|
||
| Compression ratio | Within targets | ✅ Passing |
|
||
| Graceful fallback | Always succeeds | ✅ Passing |
|
||
|
||
### Monitoring
|
||
|
||
Track optimization effectiveness via structured logging:
|
||
|
||
```rust
|
||
tracing::info!(
|
||
event = "query_optimization",
|
||
chunks_count = chunks.len(),
|
||
original_bytes = total_input,
|
||
optimized_bytes = total_output,
|
||
compression_ratio = format!("{:.1}%", ratio),
|
||
elapsed_ms = elapsed.as_secs_f64() * 1000.0,
|
||
has_optimizer = query_optimizer.enabled,
|
||
"query optimization metrics"
|
||
);
|
||
```
|
||
|
||
Export to Prometheus (ingest-time metrics):
|
||
|
||
```bash
|
||
curl http://localhost:9090/metrics | grep m3_8_optimization
|
||
```
|
||
|
||
### Best Practices
|
||
|
||
1. **Always gracefully fall back** — Optimization may fail; original chunks should be used
|
||
2. **Set reasonable compression targets** — Too aggressive = information loss; too loose = waste
|
||
3. **Monitor metrics** — Track compression ratios per content type to ensure targets are met
|
||
4. **Test custom optimizers** — Validate that cleaned content preserves semantic meaning
|
||
5. **Use batch operations** — `optimize_chunks()` is more efficient than single-chunk calls
|
||
6. **Cache formatter instances** — Create format handlers once, reuse across queries
|
||
|
||
### Further Reading
|
||
|
||
- [M3.8 Pluggable Optimizer Guide](docs/M3.8-PLUGGABLE-OPTIMIZER.md) — Full architecture details
|
||
- [M3.8 Completion Summary](CLAUDE_M3.8_COMPLETE.md) — Implementation status
|
||
- [Query Optimizer Source](crates/mem-core/src/optimizer/query_optimizer.rs) — Implementation code
|
||
|
||
## Verified environment facts
|
||
|
||
Checked against the running cluster, not assumed:
|
||
|
||
| Fact | Value |
|
||
|---|---|
|
||
| Embedding dims | **768**, `nomic-ai/nomic-embed-text-v2-moe` |
|
||
| Embedding batch limit | **32** (`batch size 1200 > maximum allowed batch size 32`) |
|
||
| pgvector | **0.7.0 available in stock CNPG image**, no custom build |
|
||
| CNPG operator | **1.30.0**, declarative `Database.spec.extensions` |
|
||
| Ollama context cap | **32768** (`OLLAMA_CONTEXT_LENGTH`) — cluster-side, overrides client config |
|
||
| Controller | `qwen2.5:3b-instruct` — paper's exact 3B backbone |
|
||
| Gateway auth | `apikey:` header. `Authorization: Bearer` returns **401** |
|
||
| Rerank response | bare array, not `{"data":[...]}`; sorted by score, map back via `index` |
|
||
|
||
Budget fits the 32K cap: 5000 chunk + ~3200 prompt/memory + 2048 response.
|
||
|
||
## Phases
|
||
|
||
Each ends in a composition gate. No phase starts until predecessor gate is green.
|
||
|
||
| | Phase | Tasks | Gate asserts |
|
||
|---|---|---|---|
|
||
| M0 | Read-only spine | 8 | third source needs no downstream change; runs offline |
|
||
| M1 | Gated loop at L1 | 8 | **update-rate < 30%**, memory flat not climbing |
|
||
| M2 | Projections | 8 | rebuild byte-identical from log alone |
|
||
| M3 | L2 + retrieval | 4 | hit rate ≥ 0.8, provenance precision ≥ 0.9 |
|
||
| M4 | Skills | 3 | draft not loadable; promoted skill never becomes evidence |
|
||
| M5 | Post-training | 6 | adapter beats prompted baseline on held-out project |
|
||
|
||
**Update-rate is the number to watch.** It is what distinguishes a gate from an
|
||
expensive summarizer. Tool results are 43% of records and mostly evidence-free,
|
||
so a correct gate rejects the large majority of chunks.
|
||
|
||
M0 and M2.2 need no model access and can start immediately. M5.4 (vLLM + LoRA)
|
||
is homelab work independent of the rest of M5.
|
||
|
||
## Reading order
|
||
|
||
1. This file
|
||
2. [memory-tasks/INDEX.md](memory-tasks/INDEX.md) — board, ordering rules, verification practice
|
||
3. [DESIGN.md](DESIGN.md) — full design, schemas, risks
|
||
4. Individual task files — self-contained, no DESIGN.md read required
|
||
# Trigger build run 130
|
||
# CI trigger
|