Deleted 31 completed task files: - M0.x: 8 tasks (cargo, domain types, recordsource, tokenizer, adapters, gate) - M1.x: 8 tasks (llm-chat, standing-query, prompt template, parser, loop, log, e2e, gate) - M3.x: 4 tasks (l2-synthesis, rerank, mem-query, gate) - M3.5.x: 8 tasks (http-server, ingest, query, federation, skills, projects, rate-limiting, gate) - M3.6.1: DocCorpusSource (heading-boundary chunking) - M4.1-2: skill-draft, derived-filter Updated INDEX.md: - Removed M0 & M1 phase sections (archived in git history) - Updated progress table: 65 active tasks (42✅ + 2🟡 + 21⬜) - Updated status: M0/M1 complete, M3/M3.5 gates passing, M4.1-2 done - Noted M3.5.10 JWT auth implementation complete (awaiting image rollout) - Cleaned up broken links to deleted task files Total test count: 239 passing, 2 ignored (up from 196 at M3.4) Ready for M4.3 gate composition, M5 post-training, M7 source connectors.
3.7 KiB
3.7 KiB
M8.7 — Index tuning: HNSW parameters + OpenSearch analyzers
| Field | Value |
|---|---|
| Phase | M8 — Hybrid Search |
| Size | M — 1–2 days |
| Status | ⬜ |
| Flags | — |
| Spec | inlined below |
| Blocks | M8.9 |
| Depends | M8.2 (data in both stores), M8.5 (can query both stores) |
Goal
Tune pgvector index parameters and OpenSearch analyzers for retrieval accuracy. Measure baseline NDCG before and after tuning. This is engineering, not research — change one parameter, measure, keep or revert.
Design
pgvector tuning
Current: ivfflat index with default lists.
Target: Switch to hnsw index (pgvector 0.5.0+). HNSW provides better recall than IVFFlat at the cost of slower index builds and more memory.
-- Drop old index
DROP INDEX IF EXISTS chunks_embedding_idx;
-- Create HNSW index
CREATE INDEX chunks_embedding_hnsw_idx
ON chunks USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 64);
Parameters:
m = 16— max connections per node (default 16, higher = better recall, more memory).ef_construction = 64— build-time search width (default 64, higher = better recall, slower build).ef_search = 40— query-time search width (set viaSET hnsw.ef_search = 40).
Tuning approach:
- Baseline: measure recall@50 with IVFFlat.
- Switch to HNSW with defaults.
- Measure recall@50 again.
- If recall@50 ≥ 0.95, keep defaults. Otherwise increase
ef_constructionto 128.
OpenSearch tuning
Analyzer changes:
- Add
edge_ngramtokenizer for typo tolerance oncontentfield. - Add
synonymfilter for common abbreviations:k8s → kubernetes,db → database,cfg → config. - Keep
standardanalyzer as primary, addsearch_analyzerfor queries.
Field boost tuning:
content^2.0(default — most important).section_title^1.8(headings are very relevant).source^1.0(file paths are useful but shouldn't dominate).breadcrumb^0.8(context, not content).
BM25 parameters:
k1 = 1.2(term frequency saturation — default is fine).b = 0.75(length normalization — default is fine).- Don't tune these unless baseline NDCG < 0.7.
Steps
- Create test query set: 20 queries with known-relevant documents.
- Measure baseline NDCG@10 for pgvector (semantic-only) and OpenSearch (lexical-only).
- Switch pgvector from IVFFlat to HNSW. Measure NDCG@10 again.
- Update OpenSearch index template with synonym filter + edge_ngram. Reindex. Measure.
- Record all measurements in
docs/INDEX_TUNING_RESULTS.md. - Keep changes that improve NDCG. Revert changes that don't.
Acceptance
- pgvector uses HNSW index (verify with
\d+ chunksin psql). - OpenSearch index template includes synonym filter.
- NDCG@10 measurements recorded for before/after each change.
- No regression: post-tuning NDCG ≥ pre-tuning NDCG for both engines.
- pgvector query latency < 150ms (p95) after HNSW switch.
- OpenSearch query latency < 100ms (p95) after analyzer changes.
Verify
# Check pgvector index type
psql -c "\d+ chunks" | grep hnsw
# Check OpenSearch analyzer
curl -k -H "Authorization: Bearer $TOKEN" \
https://opensearch-internal:9200/vault-test/_settings | jq '.*.settings.index.analysis'
# Run NDCG measurement
cargo run -- bench-search --queries fixtures/search_queries.yaml --output docs/INDEX_TUNING_RESULTS.md
False pass: HNSW index created but ef_search set to 1, making recall worse than IVFFlat. Check recall@50 explicitly — it should be ≥ 0.95.
Artifacts
- SQL migration (drop IVFFlat, create HNSW)
- Updated OpenSearch index template
docs/INDEX_TUNING_RESULTS.md(measurements)fixtures/search_queries.yaml(test query set)