docs: add M3.8 context optimizer to memory-flow.md
This commit is contained in:
+184
-1
@@ -7,7 +7,9 @@
|
||||
4. [Agent Context Flow](#agent-context-flow)
|
||||
5. [System Architecture](#system-architecture)
|
||||
6. [OpenSearch + JWT Authentication](#opensearch--jwt-authentication)
|
||||
7. [Pod Infrastructure](#pod-infrastructure)
|
||||
7. [M3.7.7 → M3.7.8: Failure Diagnosis Pipeline](#m377--m378-failure-diagnosis-pipeline)
|
||||
8. [M3.8: Context Optimizer](#m38-context-optimizer)
|
||||
9. [Pod Infrastructure](#pod-infrastructure)
|
||||
---
|
||||
|
||||
## Read Flow
|
||||
@@ -911,6 +913,187 @@ mem sig explain --tool=npm --query="unable to resolve dependency"
|
||||
|
||||
---
|
||||
|
||||
## M3.8: Context Optimizer
|
||||
|
||||
Headroom-inspired pre-LLM compression layer. Sits between hybrid search
|
||||
retrieval and the LLM gateway. Search indexes stay at full fidelity.
|
||||
|
||||
### Why
|
||||
|
||||
Agent transcripts are ~43% tool results. Raw evidence chunks contain timestamps,
|
||||
temp paths, ANSI codes, verbose JSON arrays, and passing test output. Feeding
|
||||
this noise to the GRU-Mem gate wastes tokens, risks hallucination on irrelevant
|
||||
details, and breaks LLM provider KV cache (dynamic content in prefix).
|
||||
|
||||
### Where It Sits
|
||||
|
||||
```
|
||||
Query → Hybrid Search (pgvector 60% + OpenSearch 40%)
|
||||
│
|
||||
│ full-fidelity chunks (untouched)
|
||||
▼
|
||||
┌───────────────────────┐
|
||||
│ CONTEXT OPTIMIZER │
|
||||
│ │
|
||||
│ 1. Magika Detect │ ML content type classification (<1ms)
|
||||
│ 2. CacheAligner │ Move timestamps/UUIDs to tail
|
||||
│ 3. Compressor │ Per-type compression
|
||||
│ 4. CCR Store │ Cache originals for retrieval
|
||||
│ │
|
||||
└───────────┬───────────┘
|
||||
│
|
||||
│ optimized chunks (30–90% smaller)
|
||||
▼
|
||||
┌───────────────────────┐
|
||||
│ Cache-Aligned Prompt │ system | query | turn
|
||||
└───────────┬───────────┘
|
||||
│
|
||||
▼
|
||||
LLM Gateway
|
||||
```
|
||||
|
||||
**Critical invariant:** Search indexes (pgvector + OpenSearch) NEVER see
|
||||
compressed content. Compression only happens in the prompt assembly path.
|
||||
|
||||
### Stage 1: Content Detection (Magika ML)
|
||||
|
||||
Google’s Magika ONNX model classifies content type in <1ms. No LLM calls,
|
||||
no network — embedded model runs locally.
|
||||
|
||||
```rust
|
||||
use magika::Session;
|
||||
|
||||
let magika = magika::Session::new()?;
|
||||
let result = magika.identify_content_sync(content.as_bytes())?;
|
||||
let label = result.info().label; // "json", "python", "diff", etc.
|
||||
```
|
||||
|
||||
Falls back to regex heuristics when Magika confidence < 0.7.
|
||||
|
||||
| Magika Label | Our Type | Compressor |
|
||||
|---|---|---|
|
||||
| `json`, `jsonl` | Json | JsonCrusher (70–90% savings) |
|
||||
| `python`, `rust`, `go`, `typescript` | Code | CodeCompressor (40–70%) |
|
||||
| `diff` | Diff | DiffCompressor (60–80%) |
|
||||
| `yaml`, `toml`, `ini` | Config | passthrough |
|
||||
| `txt` + log patterns | Log | LogCompressor (85–95%) |
|
||||
| fallback | Text | TextCompressor (30–50%) |
|
||||
|
||||
### Stage 2: CacheAligner
|
||||
|
||||
LLM providers cache based on exact prefix match. A single changing timestamp
|
||||
early in the prompt invalidates the entire KV cache.
|
||||
|
||||
CacheAligner detects dynamic patterns and moves them to the context tail:
|
||||
|
||||
```
|
||||
BEFORE (cache miss every call):
|
||||
"At 2026-08-28T09:15:00Z, run abc123 failed with..."
|
||||
↑ timestamp + run ID break prefix match
|
||||
|
||||
AFTER (cache hit on prefix):
|
||||
"Run failed with..." ← stable prefix (cached)
|
||||
"[ctx: t=2026-08-28T09:15:00Z, run=abc123]" ← dynamic tail
|
||||
```
|
||||
|
||||
Reuses normalisation patterns from M3.7.7 `lesson.rs` (timestamp, SHA,
|
||||
path, line:col, duration, temp path regexes).
|
||||
|
||||
### Stage 3: Per-Type Compression
|
||||
|
||||
**JsonCrusher** — Statistical field analysis on JSON arrays:
|
||||
- Measures per-field variance, uniqueness, distribution boundaries
|
||||
- Allocation: 30% start (schema), 15% end (recency), 55% importance
|
||||
- Keeps: all keys, structure, error/null/boolean fields, boundary items
|
||||
- Drops: homogeneous mid-array elements, long string values
|
||||
|
||||
**LogCompressor** — Reuses M3.7.7 signature extraction:
|
||||
- `markers()` for error line detection (npm, cargo, kubectl, docker)
|
||||
- `is_cascade()` for noise suppression
|
||||
- `strip_ansi()` for cleanup
|
||||
- Keeps: error lines, stack traces, exit codes, FAIL markers
|
||||
- Drops: INFO/DEBUG noise, passing tests, repeated patterns
|
||||
|
||||
**CodeCompressor** — Signature preservation (opt-in):
|
||||
- Keeps: imports, function/method signatures, type annotations
|
||||
- Drops: function bodies, inline comments, blank lines
|
||||
- Simple brace-counting heuristics (not full AST parser)
|
||||
|
||||
**DiffCompressor** — Change-only extraction:
|
||||
- Keeps: `+`/`-` lines (actual changes), hunk headers (`@@`)
|
||||
- Drops: unchanged context lines
|
||||
|
||||
**TextCompressor** — Token importance scoring:
|
||||
- Reuses M3.7.8 stop words for low-value token detection
|
||||
- Keeps: high-entropy tokens (IDs, hashes, error codes)
|
||||
- Drops: filler words, repeated phrases
|
||||
|
||||
### Stage 4: CCR Store (Compress-Cache-Retrieve)
|
||||
|
||||
Compression is aggressive but reversible. Full originals cached with SHA256
|
||||
hash. Retrieval hint injected into compressed output:
|
||||
|
||||
```
|
||||
[compressed evidence...]
|
||||
<!-- CCR:7f3a8bc... — full content available -->
|
||||
```
|
||||
|
||||
If the model needs more detail, it can request the original via hash lookup.
|
||||
LRU cache with TTL (default 1hr, matches gate run duration).
|
||||
|
||||
### Compression Targets
|
||||
|
||||
| Content Type | Ratio | Speed | Preserved |
|
||||
|---|---|---|---|
|
||||
| JSON arrays | 70–90% | ~1ms | All keys, structure, boundaries |
|
||||
| Build logs | 85–95% | ~1ms | Errors, stack traces, exit codes |
|
||||
| Source code | 40–70% | ~2ms | Signatures, imports, types |
|
||||
| Unified diffs | 60–80% | ~1ms | Change lines, hunk headers |
|
||||
| Plain text | 30–50% | ~2ms | High-entropy tokens |
|
||||
|
||||
### Code Reuse
|
||||
|
||||
| Existing Module | Reused For |
|
||||
|---|---|
|
||||
| `lesson.rs` normalise() | CacheAligner pattern detection |
|
||||
| `lesson.rs` markers() | LogCompressor error detection |
|
||||
| `lesson.rs` is_cascade() | LogCompressor noise suppression |
|
||||
| `lesson.rs` strip_ansi() | Pre-processing cleanup |
|
||||
| `symptom_projection.rs` STOP_WORDS | TextCompressor low-value tokens |
|
||||
|
||||
### Configuration
|
||||
|
||||
```bash
|
||||
# Enable/disable (default: on)
|
||||
MEM_CONTEXT_OPTIMIZER=on
|
||||
|
||||
# Per-compressor toggle
|
||||
MEM_COMPRESS_JSON=on
|
||||
MEM_COMPRESS_LOGS=on
|
||||
MEM_COMPRESS_CODE=off # opt-in
|
||||
MEM_COMPRESS_DIFF=on
|
||||
MEM_COMPRESS_TEXT=on
|
||||
|
||||
# CCR store
|
||||
MEM_CCR_ENABLED=on
|
||||
MEM_CCR_MAX_ENTRIES=1000
|
||||
MEM_CCR_TTL_SECS=3600
|
||||
```
|
||||
|
||||
### Task Breakdown (M3.8.x)
|
||||
|
||||
| Task | What | Status |
|
||||
|---|---|---|
|
||||
| M3.8.1 | ContentRouter (Magika) + all compressors + CCR | ⬜ |
|
||||
| M3.8.2 | CacheAligner integration with PromptBuilder | ⬜ |
|
||||
| M3.8.3 | Compression benchmarks + ratio tuning | ⬜ |
|
||||
| M3.8.4 | Composition gate | ⬜ |
|
||||
|
||||
See `docs/CONTEXT_OPTIMIZER.md` for full design.
|
||||
Inspired by [Headroom](https://docs.headroomlabs.ai/docs/how-compression-works).
|
||||
|
||||
---
|
||||
|
||||
## Pod Infrastructure
|
||||
|
||||
Complete pod inventory deployed in `poimen` namespace.
|
||||
|
||||
Reference in New Issue
Block a user