Deleted 31 completed task files: - M0.x: 8 tasks (cargo, domain types, recordsource, tokenizer, adapters, gate) - M1.x: 8 tasks (llm-chat, standing-query, prompt template, parser, loop, log, e2e, gate) - M3.x: 4 tasks (l2-synthesis, rerank, mem-query, gate) - M3.5.x: 8 tasks (http-server, ingest, query, federation, skills, projects, rate-limiting, gate) - M3.6.1: DocCorpusSource (heading-boundary chunking) - M4.1-2: skill-draft, derived-filter Updated INDEX.md: - Removed M0 & M1 phase sections (archived in git history) - Updated progress table: 65 active tasks (42✅ + 2🟡 + 21⬜) - Updated status: M0/M1 complete, M3/M3.5 gates passing, M4.1-2 done - Noted M3.5.10 JWT auth implementation complete (awaiting image rollout) - Cleaned up broken links to deleted task files Total test count: 239 passing, 2 ignored (up from 196 at M3.4) Ready for M4.3 gate composition, M5 post-training, M7 source connectors.
4.2 KiB
4.2 KiB
M8.9 — M8 composition gate: hybrid search proves its value
| Field | Value |
|---|---|
| Phase | M8 — Hybrid Search |
| Size | M — 1 day |
| Status | ⬜ |
| Flags | gate |
| Spec | inlined below |
| Blocks | — |
| Depends | M8.1–M8.8 all ✅ |
Goal
Prove the hybrid search system works end-to-end and is measurably better than semantic-only. This gate verifies the composition — individual tasks pass their own tests, but only the gate proves they compose correctly.
Properties to verify
P1: Dual-write consistency
Every chunk in pgvector has a corresponding document in OpenSearch with the same ID, and vice versa. Zero orphans.
-- pgvector IDs not in OpenSearch
SELECT id FROM chunks WHERE opensearch_pending = true;
-- Must return 0 rows (after background retry has run)
# OpenSearch document count must equal pgvector chunk count for same project
PG_COUNT=$(psql -t -c "SELECT count(*) FROM chunks WHERE project='test'")
OS_COUNT=$(curl -sk "https://opensearch:9200/vault-test/_count" | jq .count)
# PG_COUNT == OS_COUNT
P2: Hybrid outperforms single-engine
From docs/BENCHMARK_RESULTS.md:
- Hybrid NDCG@10 > semantic-only NDCG@10.
- Hybrid NDCG@10 > lexical-only NDCG@10.
- If this fails for a specific query category, it must be documented with reasoning.
P3: Fallback works under failure
- Stop OpenSearch. Query endpoint still responds with semantic results.
- Start OpenSearch. Query endpoint returns hybrid results.
- Response
search_strategyfield accurately reports which mode was used.
P4: JWT auth enforced end-to-end
- Query without token → 401.
- Query with valid token → 200.
- OpenSearch rejects requests from non-Memory-Service pods (NetworkPolicy).
- JWT token forwarded from Memory Service to OpenSearch (not admin credentials).
P5: No regression on existing tests
All pre-existing tests still pass. cargo test green. No #[ignore] added in M8.
P6: Latency budget met
- Hybrid query: p95 < 500ms.
- Semantic-only query: p95 < 200ms (must not regress from adding hybrid path).
- Fallback to semantic: p95 < 250ms (minimal overhead from failed OpenSearch attempt).
Gate test
#!/bin/bash
set -euo pipefail
echo "=== M8 Gate: Hybrid Search ==="
# P5: All tests pass
cargo test 2>&1 | tail -1
# Expected: test result: ok. X passed; 0 failed
# P1: Dual-write consistency
PG=$(psql -t -c "SELECT count(*) FROM chunks WHERE project='test' AND opensearch_pending=false")
OS=$(curl -sk "https://opensearch:9200/vault-test/_count" | jq .count)
[ "$PG" -eq "$OS" ] && echo "P1 PASS: $PG chunks in both stores" || echo "P1 FAIL: pg=$PG os=$OS"
# P2: Hybrid > single-engine
grep -A1 "hybrid" docs/BENCHMARK_RESULTS.md | grep -oP '[\d.]+' | head -1
# Must be highest NDCG in the table
# P3: Fallback
kubectl scale statefulset/opensearch -n poimen --replicas=0
sleep 5
STRATEGY=$(curl -s -H "Authorization: Bearer $TOKEN" \
"http://localhost:8080/memory/query?query=test&project=test" | jq -r .search_strategy)
[ "$STRATEGY" = "semantic_fallback" ] && echo "P3 PASS: fallback works" || echo "P3 FAIL: $STRATEGY"
kubectl scale statefulset/opensearch -n poimen --replicas=2
sleep 30
STRATEGY=$(curl -s -H "Authorization: Bearer $TOKEN" \
"http://localhost:8080/memory/query?query=test&project=test" | jq -r .search_strategy)
[ "$STRATEGY" = "Hybrid" ] && echo "P3 PASS: hybrid restored" || echo "P3 FAIL: $STRATEGY"
# P4: Auth
STATUS=$(curl -s -o /dev/null -w '%{http_code}' "http://localhost:8080/memory/query?query=test&project=test")
[ "$STATUS" = "401" ] && echo "P4 PASS: no-auth rejected" || echo "P4 FAIL: $STATUS"
echo "=== M8 Gate Complete ==="
Acceptance
All six properties pass. If P2 fails (hybrid not better), the gate does NOT pass — go back and fix M8.7 (index tuning) or M8.4 (fusion algorithm).
False passes to check
- P1 looks green but IDs don't match. Run a JOIN, not just count comparison.
- P3 looks green but fallback latency is 30s (timeout, not fast fail). Check p95 < 250ms.
- P5 looks green but test count dropped. Compare
cargo test 2>&1 | grep 'test result'against last known count (currently 239+).
Artifacts
- Gate script (inline above)
docs/BENCHMARK_RESULTS.md(from M8.8)- All M8.1–M8.8 artifacts