Files
poimen-memory/tasks/M3.8.5-compression-benchmarks.md
T
Story Crater Bot ecd8f510f3
Build and Push / Test (push) Failing after 1m54s
Build and Push / Build and push image (push) Skipped
docs: update M3.8 task specs (M3.8.3-6 detailed)
M3.8.3  COMPLETE (7 tests)
- MetricsCollector: per-project aggregation
- Structured logging (tracing)
- Prometheus export format

M3.8.4  IMPLICIT (no work needed)
- Query path already clean (no compression)
- Only cache_metrics() uses optimizer (for observability)

M3.8.5  ACTIVE (16 tests spec'd)
- Compression ratio benchmarks (5 tests: log/json/text/diff/mixed)
- Search quality validation (8 tests: pgvector/opensearch/fusion)
- Performance baseline (3 tests: latency/throughput/memory)

M3.8.6  PENDING (13 gate assertions)
- Safety (6): no data loss, deterministic, structure preservation
- Performance (4): latency p99 <3ms, throughput 1000+/sec, memory <100MB
- Quality (3): compression targets, search improvement, cache accuracy

Project progress: 64/78 complete (82%), 8/13 gates green
Total M3.8 tests: 103 (62+5+7+0+16+13)
2026-08-28 11:46:09 -07:00

5.3 KiB
Raw Blame History

M3.8.5 — Compression Benchmarks & Search Quality Validation

Field Value
Phase M3.8 — Context optimization
Size M — 12 days
Status Not started
Depends M3.8.1, M3.8.2, M3.8.3
Blocks M3.8.6

Goal

Validate that M3.8 optimization improves search quality (pgvector + OpenSearch) without sacrificing performance.

Deliverables

1. Compression Ratio Benchmarks

Test file: crates/mem-ingest/tests/it_optimizer_benchmarks.rs (200 LOC)

Benchmark each content type on real ingest sources:

#[tokio::test]
async fn benchmark_pi_session_compression() {
    // Load real Pi session transcript
    let source = PiSessionSource::new("fixtures/transcripts/pi-session-sample.json")?;
    
    let optimizer = ContextOptimizer::new()?;
    let metrics = Arc::new(Mutex::new(OptimizationMetrics::default()));
    
    let mut stream = source.records();
    while let Some(record) = stream.next().await {
        let _ = optimize_record_with_metrics(record?, &optimizer, &metrics)?;
    }
    
    let m = metrics.lock().unwrap();
    
    // Verify targets
    assert!(m.compression_ratio() >= 85.0, "log compression >= 85%");
    assert!(m.compression_ratio() <= 95.0, "log compression <= 95%");
    
    tracing::info!(
        ratio = m.compression_ratio(),
        "pi_session compression ratio"
    );
}

Tests (5):

  • benchmark_pi_session_compression (log: 85-95%)
  • benchmark_claude_transcript_compression (mixed: 60-80%)
  • benchmark_doc_corpus_compression (text: 30-50%)
  • benchmark_aggregate_compression_all_sources
  • benchmark_compression_ratio_per_compressor

2. Search Quality Metrics

Test file: tests/it_m3_8_search_quality.rs (300 LOC)

Measure pgvector + OpenSearch impact of optimization:

#[tokio::test]
async fn test_pgvector_embedding_quality() {
    // Before optimization: noisy content
    let noisy = "ERROR: failed\nINFO: debug\nTRACE: verbose\nERROR: connection";
    let noisy_embedding = embed(noisy).await?;
    
    // After optimization: clean content
    let clean = "ERROR: failed\nERROR: connection";
    let clean_embedding = embed(clean).await?;
    
    // Measure similarity
    let similarity = cosine_similarity(&noisy_embedding, &clean_embedding);
    
    // Optimized version should be nearly identical
    // (stop words/debug lines don't carry semantic info)
    assert!(similarity > 0.95, "embeddings should be similar");
}

Tests (8):

  • test_pgvector_embedding_quality (cosine similarity)
  • test_pgvector_vector_magnitude_preserved (length variance)
  • test_opensearch_bm25_score_improvement (ranking boost)
  • test_opensearch_noise_reduction (fewer false matches)
  • test_hybrid_fusion_score_stability (60% sem + 40% lex)
  • test_search_latency_with_optimization (<10ms end-to-end)
  • test_compression_does_not_break_semantic_meaning
  • test_multi_chunk_search_consistency

3. Performance Baseline

Measure optimization overhead:

#[tokio::test]
async fn test_optimization_latency_p99() {
    let optimizer = ContextOptimizer::new()?;
    let mut latencies = Vec::new();
    
    for i in 0..1000 {
        let record = make_large_record();  // 10KB+ content
        
        let start = std::time::Instant::now();
        let _optimized = optimizer.optimize(&record.text)?;
        latencies.push(start.elapsed());
    }
    
    latencies.sort();
    let p99 = latencies[990];  // 99th percentile
    
    assert!(p99 < Duration::from_millis(3), "p99 latency < 3ms");
    
    tracing::info!(
        p50_ms = latencies[500].as_secs_f64() * 1000.0,
        p99_ms = p99.as_secs_f64() * 1000.0,
        "optimization latency"
    );
}

Tests (3):

  • test_optimization_latency_p99 (<3ms)
  • test_throughput_sustained (1000+ records/sec)
  • test_memory_usage_bounded (<100MB cache)

4. Test Fixtures

Create test data in fixtures/benchmarks/:

  • pi-session-sample.json — Real Pi transcript (varies: 80-95% compression)
  • claude-transcript.json — Claude chat (varies: 60-85% compression)
  • markdown-docs.txt — Markdown content (varies: 40-60% compression)
  • json-output.json — Structured data (varies: 70-90% compression)
  • mixed-logs.txt — Mixed log output (varies: 85-95% compression)

5. Summary Report

After benchmarks run, generate docs/M3.8.5-BENCHMARKS.md:

# M3.8 Compression Benchmarks

## Compression Ratios

| Content Type | Target | Measured | Status |
|---|---|---|---|
| Logs | 85-95% | 89.2% | ✅ |
| JSON | 70-90% | 78.5% | ✅ |
| Text | 30-50% | 42.1% | ✅ |
| Diffs | 60-80% | 71.3% | ✅ |
| Mixed | 60-75% | 68.9% | ✅ |

## Search Quality Impact

- pgvector embedding similarity: 0.96 (before vs after)
- OpenSearch BM25 ranking: +18% MRR
- Hybrid search fusion: stable

## Performance

- Latency p99: 1.8ms
- Throughput: 1200 records/sec
- Cache memory: 32MB typical

Acceptance Criteria

All compression targets met (measured >= target) 16 new tests (5 compression + 8 search + 3 perf) No performance regressions (<3ms per record) Search quality improves (pgvector + OpenSearch) Benchmark report generated Fixtures checked in (reusable for future A/B testing)

Success Metrics

  • Compression ratio: log 89.2%, json 78.5%, text 42.1%
  • Embedding similarity: >0.95 (noisy vs clean)
  • Search latency: <10ms end-to-end
  • Optimization overhead: <2ms p99