docs: update M3.8 task specs (M3.8.3-6 detailed)
M3.8.3 ✅ COMPLETE (7 tests) - MetricsCollector: per-project aggregation - Structured logging (tracing) - Prometheus export format M3.8.4 ✅ IMPLICIT (no work needed) - Query path already clean (no compression) - Only cache_metrics() uses optimizer (for observability) M3.8.5 ⏳ ACTIVE (16 tests spec'd) - Compression ratio benchmarks (5 tests: log/json/text/diff/mixed) - Search quality validation (8 tests: pgvector/opensearch/fusion) - Performance baseline (3 tests: latency/throughput/memory) M3.8.6 ⏳ PENDING (13 gate assertions) - Safety (6): no data loss, deterministic, structure preservation - Performance (4): latency p99 <3ms, throughput 1000+/sec, memory <100MB - Quality (3): compression targets, search improvement, cache accuracy Project progress: 64/78 complete (82%), 8/13 gates green Total M3.8 tests: 103 (62+5+7+0+16+13)
This commit is contained in:
@@ -0,0 +1,184 @@
|
||||
# M3.8.5 — Compression Benchmarks & Search Quality Validation
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| Phase | M3.8 — Context optimization |
|
||||
| Size | M — 1–2 days |
|
||||
| Status | ⬜ Not started |
|
||||
| Depends | M3.8.1, M3.8.2, M3.8.3 |
|
||||
| Blocks | M3.8.6 |
|
||||
|
||||
## Goal
|
||||
|
||||
Validate that M3.8 optimization improves search quality (pgvector + OpenSearch)
|
||||
without sacrificing performance.
|
||||
|
||||
## Deliverables
|
||||
|
||||
### 1. Compression Ratio Benchmarks
|
||||
|
||||
Test file: `crates/mem-ingest/tests/it_optimizer_benchmarks.rs` (200 LOC)
|
||||
|
||||
Benchmark each content type on real ingest sources:
|
||||
|
||||
```rust
|
||||
#[tokio::test]
|
||||
async fn benchmark_pi_session_compression() {
|
||||
// Load real Pi session transcript
|
||||
let source = PiSessionSource::new("fixtures/transcripts/pi-session-sample.json")?;
|
||||
|
||||
let optimizer = ContextOptimizer::new()?;
|
||||
let metrics = Arc::new(Mutex::new(OptimizationMetrics::default()));
|
||||
|
||||
let mut stream = source.records();
|
||||
while let Some(record) = stream.next().await {
|
||||
let _ = optimize_record_with_metrics(record?, &optimizer, &metrics)?;
|
||||
}
|
||||
|
||||
let m = metrics.lock().unwrap();
|
||||
|
||||
// Verify targets
|
||||
assert!(m.compression_ratio() >= 85.0, "log compression >= 85%");
|
||||
assert!(m.compression_ratio() <= 95.0, "log compression <= 95%");
|
||||
|
||||
tracing::info!(
|
||||
ratio = m.compression_ratio(),
|
||||
"pi_session compression ratio"
|
||||
);
|
||||
}
|
||||
```
|
||||
|
||||
Tests (5):
|
||||
- `benchmark_pi_session_compression` (log: 85-95%)
|
||||
- `benchmark_claude_transcript_compression` (mixed: 60-80%)
|
||||
- `benchmark_doc_corpus_compression` (text: 30-50%)
|
||||
- `benchmark_aggregate_compression_all_sources`
|
||||
- `benchmark_compression_ratio_per_compressor`
|
||||
|
||||
### 2. Search Quality Metrics
|
||||
|
||||
Test file: `tests/it_m3_8_search_quality.rs` (300 LOC)
|
||||
|
||||
Measure pgvector + OpenSearch impact of optimization:
|
||||
|
||||
```rust
|
||||
#[tokio::test]
|
||||
async fn test_pgvector_embedding_quality() {
|
||||
// Before optimization: noisy content
|
||||
let noisy = "ERROR: failed\nINFO: debug\nTRACE: verbose\nERROR: connection";
|
||||
let noisy_embedding = embed(noisy).await?;
|
||||
|
||||
// After optimization: clean content
|
||||
let clean = "ERROR: failed\nERROR: connection";
|
||||
let clean_embedding = embed(clean).await?;
|
||||
|
||||
// Measure similarity
|
||||
let similarity = cosine_similarity(&noisy_embedding, &clean_embedding);
|
||||
|
||||
// Optimized version should be nearly identical
|
||||
// (stop words/debug lines don't carry semantic info)
|
||||
assert!(similarity > 0.95, "embeddings should be similar");
|
||||
}
|
||||
```
|
||||
|
||||
Tests (8):
|
||||
- `test_pgvector_embedding_quality` (cosine similarity)
|
||||
- `test_pgvector_vector_magnitude_preserved` (length variance)
|
||||
- `test_opensearch_bm25_score_improvement` (ranking boost)
|
||||
- `test_opensearch_noise_reduction` (fewer false matches)
|
||||
- `test_hybrid_fusion_score_stability` (60% sem + 40% lex)
|
||||
- `test_search_latency_with_optimization` (<10ms end-to-end)
|
||||
- `test_compression_does_not_break_semantic_meaning`
|
||||
- `test_multi_chunk_search_consistency`
|
||||
|
||||
### 3. Performance Baseline
|
||||
|
||||
Measure optimization overhead:
|
||||
|
||||
```rust
|
||||
#[tokio::test]
|
||||
async fn test_optimization_latency_p99() {
|
||||
let optimizer = ContextOptimizer::new()?;
|
||||
let mut latencies = Vec::new();
|
||||
|
||||
for i in 0..1000 {
|
||||
let record = make_large_record(); // 10KB+ content
|
||||
|
||||
let start = std::time::Instant::now();
|
||||
let _optimized = optimizer.optimize(&record.text)?;
|
||||
latencies.push(start.elapsed());
|
||||
}
|
||||
|
||||
latencies.sort();
|
||||
let p99 = latencies[990]; // 99th percentile
|
||||
|
||||
assert!(p99 < Duration::from_millis(3), "p99 latency < 3ms");
|
||||
|
||||
tracing::info!(
|
||||
p50_ms = latencies[500].as_secs_f64() * 1000.0,
|
||||
p99_ms = p99.as_secs_f64() * 1000.0,
|
||||
"optimization latency"
|
||||
);
|
||||
}
|
||||
```
|
||||
|
||||
Tests (3):
|
||||
- `test_optimization_latency_p99` (<3ms)
|
||||
- `test_throughput_sustained` (1000+ records/sec)
|
||||
- `test_memory_usage_bounded` (<100MB cache)
|
||||
|
||||
### 4. Test Fixtures
|
||||
|
||||
Create test data in `fixtures/benchmarks/`:
|
||||
|
||||
- `pi-session-sample.json` — Real Pi transcript (varies: 80-95% compression)
|
||||
- `claude-transcript.json` — Claude chat (varies: 60-85% compression)
|
||||
- `markdown-docs.txt` — Markdown content (varies: 40-60% compression)
|
||||
- `json-output.json` — Structured data (varies: 70-90% compression)
|
||||
- `mixed-logs.txt` — Mixed log output (varies: 85-95% compression)
|
||||
|
||||
### 5. Summary Report
|
||||
|
||||
After benchmarks run, generate `docs/M3.8.5-BENCHMARKS.md`:
|
||||
|
||||
```markdown
|
||||
# M3.8 Compression Benchmarks
|
||||
|
||||
## Compression Ratios
|
||||
|
||||
| Content Type | Target | Measured | Status |
|
||||
|---|---|---|---|
|
||||
| Logs | 85-95% | 89.2% | ✅ |
|
||||
| JSON | 70-90% | 78.5% | ✅ |
|
||||
| Text | 30-50% | 42.1% | ✅ |
|
||||
| Diffs | 60-80% | 71.3% | ✅ |
|
||||
| Mixed | 60-75% | 68.9% | ✅ |
|
||||
|
||||
## Search Quality Impact
|
||||
|
||||
- pgvector embedding similarity: 0.96 (before vs after)
|
||||
- OpenSearch BM25 ranking: +18% MRR
|
||||
- Hybrid search fusion: stable
|
||||
|
||||
## Performance
|
||||
|
||||
- Latency p99: 1.8ms
|
||||
- Throughput: 1200 records/sec
|
||||
- Cache memory: 32MB typical
|
||||
```
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
✅ All compression targets met (measured >= target)
|
||||
✅ 16 new tests (5 compression + 8 search + 3 perf)
|
||||
✅ No performance regressions (<3ms per record)
|
||||
✅ Search quality improves (pgvector + OpenSearch)
|
||||
✅ Benchmark report generated
|
||||
✅ Fixtures checked in (reusable for future A/B testing)
|
||||
|
||||
## Success Metrics
|
||||
|
||||
- Compression ratio: log 89.2%, json 78.5%, text 42.1%
|
||||
- Embedding similarity: >0.95 (noisy vs clean)
|
||||
- Search latency: <10ms end-to-end
|
||||
- Optimization overhead: <2ms p99
|
||||
Reference in New Issue
Block a user