rock c5a46dd82e feat: M8.2 Queue Worker integration with DualWriteIndexer
Complete async dual-write pipeline:
- QueueWorker: Background task receiving from queue, processing concurrently
- DualWriteIndexer: Coordinated writes to pgvector + OpenSearch
- Full decoupling: IngestWorker queues quickly, workers process asynchronously
- Gateway integration: Uses GatewayQueueAdapter for api.riotpiao.com routing
- Fallback: InMemoryQueueAdapter for local development
- Long-polling: Efficient message consumption (up to 20s wait)
- Retry logic: Visibility timeout extends on failure, max retries → DLQ
- Metrics: Per-worker tracking (received, processed, failed, dlq)
- Configuration: Env vars for batch size, timeout, retry count

Architecture:
- IngestWorker → queue.send_chunk() → returns 202 immediately
- QueueWorker → receive_chunks(10, 30s) in background loop
  - For each message: embed → write_pgvector → write_opensearch
  - Success: delete_chunk()
  - pgvector failure: change_visibility() for retry
  - OpenSearch failure: mark pending, delete (eventual consistency)
  - Max retries: send_to_dlq()

Files:
- crates/mem-cli/src/queue_worker.rs (430 LOC)
- crates/mem-cli/src/http_server.rs (+100 LOC queue worker init)
- tests/it_queue_worker_integration.rs (260 LOC, 11 tests)
- docs/M8.2-QUEUE_WORKER_INTEGRATION.md (350 LOC)

Benefits:
- 10-100x faster ingest API response
- True concurrent processing (multiple workers)
- Fault tolerance (retries, DLQ)
- Observability (metrics, logs)
- Horizontal scalability (replicas)
2026-08-28 13:14:39 -07:00
2026-08-22 23:13:42 -07:00

poimen-memory

Gated recurrent memory over agent context. Reads session history chunk-by-chunk, keeps only what answers standing questions, projects result into an Obsidian vault and a pgvector index.

Status: design complete, no code yet. 37 tasks in memory-tasks/, 0 done. Start at M0.1.

Problem

Agent sessions grow faster than anyone reads them, and most of the volume is noise. One real pi session in this project:

assistant   1445
toolResult  1261     43% — ls output, file reads, mostly evidence-free
user         196
+ 8 compaction events

Compaction fires 8 times per session. Context gets discarded, not retained — root causes, decisions and gotchas evaporate when window rolls.

Mechanism

GRU-Mem (arXiv 2602.10560). Two text-controlled gates on a recurrent memory loop:

  • update gate — memory only mutates when chunk contains evidence. Blocks the memory explosion that ungated recurrent memory hits.
  • exit gate — stop scanning once evidence sufficient.

Paper reports up to 400% speedup and better accuracy than ungated, because unbounded memory growth degrades later updates.

sessions ─> chunk (5000 tok) ─> controller ─> gates ─> memory ─> projections

Controller emits structured output; loop acts on it:

<think>   reason about chunk vs question
<check>   yes|no        -> U_t, update or discard
<update>  candidate memory M̂_t
<next>    continue|end  -> E_t, exit or continue

Memory tiers

Level What From Bounded
L0 evidence chunk, verbatim update gate opening no, but sparse (~17 of 412)
L1 per-query memory, M_t gated loop over chunks 1024 tok
L2 project synthesis gated loop over L1 memories 1024 tok

L2 is not new machinery — same loop, same prompt, L1 memories as input stream. Level is a parameter.

Tiers form a provenance graph. Each L1 records its L0 parents, each L2 its L1 parents. Same relation becomes both memory_edge rows and Obsidian wikilinks.

Standing queries

Update gate needs a referent. Paper's agent is φθ(Q, C_t, M_{t-1}) — gate is defined as "does this chunk contain useful information about the problem". No Q, no gate, and r_update becomes undefinable, which kills post-training.

So each project declares durable questions. One query = one L1 memory = one note.

# queries/poimen.yaml
project: poimen
roots: [/Users/rockliang/workplace/Poimen/agent-rust]
queries:
  - id: infra-root-causes
    question: What infrastructure bugs were found, what was the root cause, how was it isolated?
  - id: architecture-decisions
    question: What architectural decisions were made, with reasoning and rejected alternatives?
synthesis:
  question: What is the current state of this project, and what should someone know before working on it?
  exit_gate: true

Exit gate off at L1, on at L2. Paper §3.3 makes this call: for "what are all the X" questions you cannot know evidence is sufficient without reading everything. L1 extraction is that shape. At L2 input is a handful of memories and sufficiency is decidable. Gate still recorded at L1 — signal needed for post-training.

Authority model

JSONL log authoritative. Vault and vector index are projections.

Anything not rebuildable byte-identically from the log has hidden inputs, and that is a bug. Gate M2.8 enforces it destructively:

rm -rf vault/poimen
psql -c "delete from memory_node where project='poimen'"
mem rebuild --from-log --project poimen
git -C vault diff --exit-code      # empty diff is the only pass

Buys three things: re-embedding after model change is a rebuild not a migration, Obsidian edits cannot corrupt the record, post-training corpus is the log itself.

Skills

A skill is a projection, not a level. L0/L1/L2 are descriptive — what happened. A skill is procedural — what to do next time. Gated loop does not produce it.

Format free: SKILL.md is YAML frontmatter + markdown, which is an Obsidian note. So vault/skills/<name>/SKILL.md is both, no conversion:

pi --skill vault/skills/
ln -s .../vault/skills/<name> ~/.claude/skills/<name>

Drafts land in _drafts/, promotion is a human git mv. This is the one cycle in the design:

emitted skill auto-loads -> appears in future transcripts
-> ingested as evidence -> reinforces the memory that emitted it

No external verifier breaks it. Two guards: _drafts/ is a directory (cannot be globbed into --skill), and every artifact carries generated_from so ingest tags matching chunks derived: true and refuses them as evidence.

Separate weights

Memory policy is a LoRA adapter on Qwen2.5-3B-Instruct, not a fine-tuned model. Reason is VRAM: one GPU, OLLAMA_MAX_LOADED_MODELS=2, already holding ornith:35b + qwen2.5:3b. Separate full model evicts something, and eviction is a weights reload measured in tens of seconds. Adapter rides the resident base.

Also: post-training emits ~50 MB, not 6 GB. Swap without redeploy. Regression reverts by pointing at previous adapter.

Ollama cannot hot-swap LoRA. Serving one needs vLLM with --enable-lora (pattern already exists — reasoning predictor is vLLM v0.11.0). Phases M0M4 run prompted-only, so decision is deferred, not dodged.

Layout

DESIGN.md            full design, 460 lines
memory-tasks/        37 task files + INDEX.md — tracked
crates/
  mem-core/          domain types; Level; gate parser; the gated loop
  mem-chunk/         RecordSource trait; ChunkPolicy; FlushTrigger
  mem-llm/           gateway client — chat, embeddings, rerank
  mem-ingest/        source adapters: pi sessions, claude transcripts
  mem-store/         JSONL log; pgvector repo; Obsidian projector
  mem-cli/           binary `mem`
queries/             standing query YAML per project
log/                 JSONL event log — authoritative, tracked
vault/               Obsidian output

mem-chunk is separate and stream-shaped from day one. Sources today are files with an EOF; telemetry or a live tail will not have one. RecordSource returns impl Stream<Item = Record>; batch sources become streams via futures::stream::iter, so it costs nothing now and removes a rewrite later.

Commands

mem ingest --project poimen --dry-run          # chunk plan, zero model calls
mem ingest --project poimen --query infra-root-causes
mem synthesize --project poimen                # L2 pass, exit gate on
mem rebuild --from-log --project poimen        # drop and rebuild projections
mem verify --project poimen                    # provenance graph closure
mem query "why did requests over 10KB fail?"
mem skill draft --from poimen/infra-root-causes
mem label --project poimen                     # evidence labels for training

M3.8 Pluggable Query Optimization

Purpose: Compress and optimize search results before passing them to the LLM context window, improving token efficiency and response quality.

Architecture

M3.8 provides dual-path optimization:

Ingest-Time Optimization (M3.8.2)

When documents are ingested, they're automatically optimized before embedding:

Records → optimize_record_with_metrics() → Clean chunks (85-95% of original)
       → Embed (pgvector) → Index (OpenSearch)

Benefits:

  • Better pgvector embeddings (clean text = higher semantic quality)
  • Better OpenSearch BM25 ranking (signal-rich text = stronger matches)
  • One-time cost per document
  • All queries benefit from cleaner search index

Query-Time Optimization (QueryOptimizer)

When search results are retrieved, they're optimized before LLM processing:

Hybrid search results → QueryOptimizer.optimize_chunks() → Clean chunks
                     → LLM context window

Benefits:

  • Smaller context window (fewer tokens to LLM)
  • Faster response generation
  • Focus on signal (removes noise like timestamps, debug lines, repetitive keys)

Using Query Optimization

1. Basic Query with Auto-Optimization

use mem_core::optimizer::QueryOptimizer;

// Create optimizer (loads config from env vars)
let query_optimizer = QueryOptimizer::from_env();

// Get search results
let chunks = hybrid_search(&question).await?;

// Auto-optimize before LLM
let optimized = query_optimizer.optimize_chunks(&chunks).await?;

// Build context from clean chunks
let context = optimized.join("\n---\n");
let response = llm.prompt(&context, &question).await?;

2. Prompt Construction with Optimization

use mem_core::optimizer::QueryOptimizer;
use mem_core::prompt::PromptBuilder;

let query_optimizer = QueryOptimizer::from_env();

// Retrieve and optimize
let chunks = hybrid_search(query).await?;
let optimized_chunks = query_optimizer.optimize_chunks(&chunks).await?;

// Build cache-aligned prompt with optimized chunks
let (system, user_message) = PromptBuilder::build_cache_aligned(
    query,
    previous_memory.as_deref(),
    /* use optimized chunks */
)?;

let response = llm.prompt(system, user_message).await?;

3. Custom Query Optimizer Implementation

For domain-specific optimization (e.g., medical, legal, technical content):

use mem_core::optimizer::{OptimizerPlugin, OptimizationResult, PluginMetrics};
use async_trait::async_trait;

struct MedicalOptimizer;

#[async_trait]
impl OptimizerPlugin for MedicalOptimizer {
    fn name(&self) -> &str { "medical-optimizer" }
    
    fn supported_types(&self) -> Vec<&str> {
        vec!["text/medical", "text/clinical", "application/json"]
    }
    
    async fn optimize(&self, content: &str) -> Result<OptimizationResult, String> {
        // Remove patient IDs, reduce duplicate diagnosis entries
        let cleaned = clean_medical_data(content);
        let ratio = cleaned.len() as f32 / content.len() as f32;
        
        Ok(OptimizationResult {
            original: content.to_string(),
            optimized: cleaned,
            ratio,
            plugin: self.name().to_string(),
            metadata: Default::default(),
        })
    }
    
    fn metrics(&self) -> PluginMetrics { Default::default() }
}

// Register and use
let service = OptimizerServiceBuilder::new()
    .with_optimizer(Arc::new(MedicalOptimizer))
    .with_format(Arc::new(JsonFormatter))
    .build()?;

let optimized = service.optimize(
    clinical_note,
    "text/clinical",
    None
).await?;

4. Optimized Query with Metrics Tracking

use mem_core::optimizer::{QueryOptimizer, QueryOptimizationMetrics};
use mem_core::prompt::CacheMetrics;

let query_optimizer = QueryOptimizer::from_env();

let chunks = hybrid_search(query).await?;
let optimized = query_optimizer.optimize_chunks(&chunks).await?;

// Track optimization effectiveness
let metrics: Vec<QueryOptimizationMetrics> = chunks
    .iter()
    .zip(&optimized)
    .map(|(orig, opt)| {
        QueryOptimizationMetrics {
            original_bytes: orig.text.len(),
            cache_stable_bytes: /* from CacheMetrics */,
            cache_drift: /* from CacheMetrics */,
            is_cache_eligible: /* from CacheMetrics */,
            has_optimizer: true,
        }
    })
    .collect();

tracing::info!(
    chunks = chunks.len(),
    compression_ratio = format!(
        "{:.1}%",
        (optimized.iter().map(|o| o.len()).sum::<usize>() as f32 
         / chunks.iter().map(|c| c.text.len()).sum::<usize>() as f32) * 100.0
    ),
    "query optimization complete"
);

// Query with optimized context
let response = llm.prompt(&optimized.join("\n---\n"), &question).await?;

5. Batch Optimization for Multiple Queries

use mem_core::optimizer::QueryOptimizer;

let query_optimizer = QueryOptimizer::from_env();

// Process multiple queries with shared optimizer
let results = futures::stream::iter(queries)
    .then(|query| async move {
        let chunks = hybrid_search(&query).await?;
        let optimized = query_optimizer.optimize_chunks(&chunks).await?;
        let response = llm.prompt(&optimized.join("\n---\n"), &query.question).await?;
        Ok((query, response))
    })
    .collect::<Vec<_>>()
    .await;

6. Conditional Optimization (Graceful Fallback)

use mem_core::optimizer::QueryOptimizer;

let query_optimizer = QueryOptimizer::from_env();

let chunks = hybrid_search(query).await?;

// Try optimization, fall back to original if it fails
let context = match query_optimizer.optimize_chunks(&chunks).await {
    Ok(optimized) => {
        tracing::info!("query optimization succeeded");
        optimized.join("\n---\n")
    }
    Err(e) => {
        tracing::warn!("query optimization failed: {}, using original", e);
        chunks.iter().map(|c| c.text.clone()).collect::<Vec<_>>().join("\n---\n")
    }
};

let response = llm.prompt(&context, &question).await?;

Environment Configuration

# Enable/disable query optimization
export MEM_QUERY_OPTIMIZER=on              # or "off"

# Custom optimizer service (optional)
export MEM_QUERY_OPTIMIZER_SERVICE=/path/to/config.yml

# Compression targets (if using custom optimizers)
export MEM_COMPRESSION_TARGETS='{
  "logs": {"min": 0.05, "max": 0.95},
  "json": {"min": 0.10, "max": 0.90},
  "text": {"min": 0.30, "max": 0.70}
}'

# Ingest-time optimization
export MEM_CONTEXT_OPTIMIZER=on

Compression Targets by Content Type

Type Target Typical Example
Logs 85-95% removal 10-15% remaining ERROR + timestamps → ERROR only
JSON 70-90% removal 10-30% remaining Minified + key filtering
Text/Markdown 30-50% removal 50-70% remaining Prose kept, formatting removed
Code/Diffs 60-80% removal 20-40% remaining Context lines removed

Performance Targets

Metric Target Status
Ingest latency <1ms per record Passing
Query latency <50ms P95 Passing
Compression ratio Within targets Passing
Graceful fallback Always succeeds Passing

Monitoring

Track optimization effectiveness via structured logging:

tracing::info!(
    event = "query_optimization",
    chunks_count = chunks.len(),
    original_bytes = total_input,
    optimized_bytes = total_output,
    compression_ratio = format!("{:.1}%", ratio),
    elapsed_ms = elapsed.as_secs_f64() * 1000.0,
    has_optimizer = query_optimizer.enabled,
    "query optimization metrics"
);

Export to Prometheus (ingest-time metrics):

curl http://localhost:9090/metrics | grep m3_8_optimization

Best Practices

  1. Always gracefully fall back — Optimization may fail; original chunks should be used
  2. Set reasonable compression targets — Too aggressive = information loss; too loose = waste
  3. Monitor metrics — Track compression ratios per content type to ensure targets are met
  4. Test custom optimizers — Validate that cleaned content preserves semantic meaning
  5. Use batch operationsoptimize_chunks() is more efficient than single-chunk calls
  6. Cache formatter instances — Create format handlers once, reuse across queries

Further Reading

Verified environment facts

Checked against the running cluster, not assumed:

Fact Value
Embedding dims 768, nomic-ai/nomic-embed-text-v2-moe
Embedding batch limit 32 (batch size 1200 > maximum allowed batch size 32)
pgvector 0.7.0 available in stock CNPG image, no custom build
CNPG operator 1.30.0, declarative Database.spec.extensions
Ollama context cap 32768 (OLLAMA_CONTEXT_LENGTH) — cluster-side, overrides client config
Controller qwen2.5:3b-instruct — paper's exact 3B backbone
Gateway auth apikey: header. Authorization: Bearer returns 401
Rerank response bare array, not {"data":[...]}; sorted by score, map back via index

Budget fits the 32K cap: 5000 chunk + ~3200 prompt/memory + 2048 response.

Phases

Each ends in a composition gate. No phase starts until predecessor gate is green.

Phase Tasks Gate asserts
M0 Read-only spine 8 third source needs no downstream change; runs offline
M1 Gated loop at L1 8 update-rate < 30%, memory flat not climbing
M2 Projections 8 rebuild byte-identical from log alone
M3 L2 + retrieval 4 hit rate ≥ 0.8, provenance precision ≥ 0.9
M4 Skills 3 draft not loadable; promoted skill never becomes evidence
M5 Post-training 6 adapter beats prompted baseline on held-out project

Update-rate is the number to watch. It is what distinguishes a gate from an expensive summarizer. Tool results are 43% of records and mostly evidence-free, so a correct gate rejects the large majority of chunks.

M0 and M2.2 need no model access and can start immediately. M5.4 (vLLM + LoRA) is homelab work independent of the rest of M5.

Reading order

  1. This file
  2. memory-tasks/INDEX.md — board, ordering rules, verification practice
  3. DESIGN.md — full design, schemas, risks
  4. Individual task files — self-contained, no DESIGN.md read required

Trigger build run 130

CI trigger

S
Description
Agent-ready Graph-RAG system with hallucination prevention and enterprise RBAC
https://forgejo.riotpiao.com/rock/poimen-memory
Readme
1.8 MiB
Languages
Rust 98.7%
Shell 0.6%
Python 0.4%
PLpgSQL 0.2%