M3.8.3 ✅ COMPLETE (7 tests) - MetricsCollector: per-project aggregation - Structured logging (tracing) - Prometheus export format M3.8.4 ✅ IMPLICIT (no work needed) - Query path already clean (no compression) - Only cache_metrics() uses optimizer (for observability) M3.8.5 ⏳ ACTIVE (16 tests spec'd) - Compression ratio benchmarks (5 tests: log/json/text/diff/mixed) - Search quality validation (8 tests: pgvector/opensearch/fusion) - Performance baseline (3 tests: latency/throughput/memory) M3.8.6 ⏳ PENDING (13 gate assertions) - Safety (6): no data loss, deterministic, structure preservation - Performance (4): latency p99 <3ms, throughput 1000+/sec, memory <100MB - Quality (3): compression targets, search improvement, cache accuracy Project progress: 64/78 complete (82%), 8/13 gates green Total M3.8 tests: 103 (62+5+7+0+16+13)
5.3 KiB
M3.8.5 — Compression Benchmarks & Search Quality Validation
| Field | Value |
|---|---|
| Phase | M3.8 — Context optimization |
| Size | M — 1–2 days |
| Status | ⬜ Not started |
| Depends | M3.8.1, M3.8.2, M3.8.3 |
| Blocks | M3.8.6 |
Goal
Validate that M3.8 optimization improves search quality (pgvector + OpenSearch) without sacrificing performance.
Deliverables
1. Compression Ratio Benchmarks
Test file: crates/mem-ingest/tests/it_optimizer_benchmarks.rs (200 LOC)
Benchmark each content type on real ingest sources:
#[tokio::test]
async fn benchmark_pi_session_compression() {
// Load real Pi session transcript
let source = PiSessionSource::new("fixtures/transcripts/pi-session-sample.json")?;
let optimizer = ContextOptimizer::new()?;
let metrics = Arc::new(Mutex::new(OptimizationMetrics::default()));
let mut stream = source.records();
while let Some(record) = stream.next().await {
let _ = optimize_record_with_metrics(record?, &optimizer, &metrics)?;
}
let m = metrics.lock().unwrap();
// Verify targets
assert!(m.compression_ratio() >= 85.0, "log compression >= 85%");
assert!(m.compression_ratio() <= 95.0, "log compression <= 95%");
tracing::info!(
ratio = m.compression_ratio(),
"pi_session compression ratio"
);
}
Tests (5):
benchmark_pi_session_compression(log: 85-95%)benchmark_claude_transcript_compression(mixed: 60-80%)benchmark_doc_corpus_compression(text: 30-50%)benchmark_aggregate_compression_all_sourcesbenchmark_compression_ratio_per_compressor
2. Search Quality Metrics
Test file: tests/it_m3_8_search_quality.rs (300 LOC)
Measure pgvector + OpenSearch impact of optimization:
#[tokio::test]
async fn test_pgvector_embedding_quality() {
// Before optimization: noisy content
let noisy = "ERROR: failed\nINFO: debug\nTRACE: verbose\nERROR: connection";
let noisy_embedding = embed(noisy).await?;
// After optimization: clean content
let clean = "ERROR: failed\nERROR: connection";
let clean_embedding = embed(clean).await?;
// Measure similarity
let similarity = cosine_similarity(&noisy_embedding, &clean_embedding);
// Optimized version should be nearly identical
// (stop words/debug lines don't carry semantic info)
assert!(similarity > 0.95, "embeddings should be similar");
}
Tests (8):
test_pgvector_embedding_quality(cosine similarity)test_pgvector_vector_magnitude_preserved(length variance)test_opensearch_bm25_score_improvement(ranking boost)test_opensearch_noise_reduction(fewer false matches)test_hybrid_fusion_score_stability(60% sem + 40% lex)test_search_latency_with_optimization(<10ms end-to-end)test_compression_does_not_break_semantic_meaningtest_multi_chunk_search_consistency
3. Performance Baseline
Measure optimization overhead:
#[tokio::test]
async fn test_optimization_latency_p99() {
let optimizer = ContextOptimizer::new()?;
let mut latencies = Vec::new();
for i in 0..1000 {
let record = make_large_record(); // 10KB+ content
let start = std::time::Instant::now();
let _optimized = optimizer.optimize(&record.text)?;
latencies.push(start.elapsed());
}
latencies.sort();
let p99 = latencies[990]; // 99th percentile
assert!(p99 < Duration::from_millis(3), "p99 latency < 3ms");
tracing::info!(
p50_ms = latencies[500].as_secs_f64() * 1000.0,
p99_ms = p99.as_secs_f64() * 1000.0,
"optimization latency"
);
}
Tests (3):
test_optimization_latency_p99(<3ms)test_throughput_sustained(1000+ records/sec)test_memory_usage_bounded(<100MB cache)
4. Test Fixtures
Create test data in fixtures/benchmarks/:
pi-session-sample.json— Real Pi transcript (varies: 80-95% compression)claude-transcript.json— Claude chat (varies: 60-85% compression)markdown-docs.txt— Markdown content (varies: 40-60% compression)json-output.json— Structured data (varies: 70-90% compression)mixed-logs.txt— Mixed log output (varies: 85-95% compression)
5. Summary Report
After benchmarks run, generate docs/M3.8.5-BENCHMARKS.md:
# M3.8 Compression Benchmarks
## Compression Ratios
| Content Type | Target | Measured | Status |
|---|---|---|---|
| Logs | 85-95% | 89.2% | ✅ |
| JSON | 70-90% | 78.5% | ✅ |
| Text | 30-50% | 42.1% | ✅ |
| Diffs | 60-80% | 71.3% | ✅ |
| Mixed | 60-75% | 68.9% | ✅ |
## Search Quality Impact
- pgvector embedding similarity: 0.96 (before vs after)
- OpenSearch BM25 ranking: +18% MRR
- Hybrid search fusion: stable
## Performance
- Latency p99: 1.8ms
- Throughput: 1200 records/sec
- Cache memory: 32MB typical
Acceptance Criteria
✅ All compression targets met (measured >= target) ✅ 16 new tests (5 compression + 8 search + 3 perf) ✅ No performance regressions (<3ms per record) ✅ Search quality improves (pgvector + OpenSearch) ✅ Benchmark report generated ✅ Fixtures checked in (reusable for future A/B testing)
Success Metrics
- Compression ratio: log 89.2%, json 78.5%, text 42.1%
- Embedding similarity: >0.95 (noisy vs clean)
- Search latency: <10ms end-to-end
- Optimization overhead: <2ms p99