# M3.5.3 — GET /query endpoint: HNSW recall, rerank, edge-walk to L0 | Field | Value | |---|---| | Phase | M3.5 — Distributed API Layer | | Size | M — 1–3 days | | Status | ✅ Done | | Flags | — | | Spec | inlined below | | Blocks | M3.5.4, M3.5.8 | | Depends | M3.5.1, M3.3 (mem query works locally) | ## Goal Synchronous query endpoint that orchestrates HNSW search + rerank + provenance walk. Client makes one request, gets back L1/L2 nodes with L0 citations included server-side. ## Design **Request:** ``` GET /memory/query?query=why+did+requests+over+10KB+fail&project=poimen&level=L1,L2&limit=5 ``` Query params: - `query` (required, URL-encoded) — user question or search text - `project` (optional) — filter to one project; if omitted, search all projects - `level` (optional, comma-separated) — `L1,L2` (default) or `L0,L1,L2`; filters by node level - `limit` (optional, integer, default 5) — how many top results to return - `timeout_seconds` (optional, integer, default 5) — abort if search exceeds this time **Response:** ```json { "query": "why did requests over 10KB fail", "project": "poimen", "level_filter": ["L1", "L2"], "results": [ { "level": "L1", "sha256": "abc...", "text": "Kong body buffer was 8MB...", "query_score": 0.92, "rerank_score": 0.94, "parents": [ { "level": "L0", "sha256": "xyz...", "source": "pi:2026-07-21-019f857d", "text": "...Kong body buffer limit...", "timestamp": "2026-07-21T16:23:59Z" } ] }, ... ], "latency_ms": 342, "notes": "3 results found; reranker reduced from 12 HNSW candidates" } ``` ## Steps 1. `GET /memory/query` handler signature: ```rust async fn query_handler( Query(params): Query, Extension(store): Extension>, Extension(llm): Extension>, ) -> Result> ``` 2. Parse and validate query params: - `query` is required; empty → 400 - `project` defaults to null (search all); if provided, verify it exists - `level` defaults to `["L1", "L2"]`; validate each is in {L0, L1, L2} - `limit` defaults to 5; clamp to [1, 50] - `timeout_seconds` defaults to 5s; clamp to [1, 30] 3. Embed the query (calls M2.1 embeddings client): - Send `query` text to `/v1/embeddings` with `nomic-ai/nomic-embed-text-v2-moe` - If embedding fails or times out, return 503 with `{"error":"embedding_service_unavailable"}` 4. HNSW recall (calls pgvector): - `SELECT sha256, level, text, embedding <-> query_embedding AS distance FROM memory_node WHERE level = ANY($1) AND (project = $2 OR $2 IS NULL) ORDER BY distance ASC LIMIT $3` - Use distance metric `vector_cosine_ops` (similarity = 1 - distance) - Compute `query_score = 1 - distance` - Return candidates (no reranking yet) 5. Rerank (calls M3.2 rerank client): - Collect top K=3×limit candidates (e.g., 15 for limit=5) - Send to `/v1/rerank` with passages=candidates and query - Parse `bge-reranker-base` response, extract score per candidate - Compute `rerank_score = raw_score / 100` (reranker outputs [0,100]) 6. Sort by rerank_score descending, take top `limit` results 7. Edge walk (L1→L0, L2→L1): - For each result, query `memory_edge` to find parent nodes - Fetch parent node text from `memory_node` - Include in `parents` array (ordered by edge precedence if tracked, else by sha256) 8. Assemble response and return 200 ## Acceptance - Query with valid text returns results - Results include query_score and rerank_score - L0 parents are walked and included - Different level filters change result count (e.g., L0 only returns more results) - Timeout parameter is respected - Query too short (e.g., single char) handled gracefully (400 or empty result, not crash) ## Verify **Harness:** Integration tests against server + pgvector repo populated with known nodes. **Setup:** Load `tests/fixtures/memory_nodes.jsonl` into test pgvector DB before each test. Nodes include L0 (evidence), L1 (per-query memory), and L2 (synthesis) with known text and relationships. **Integration test** — `tests/it_query_endpoint.rs`: 1. `a1_basic_query_returns_results` — GET /query?query=Kong+body returns 200 with `results` array. 2. `a2_scores_are_present` — result items include `query_score` and `rerank_score`, both floats in [0,1]. 3. `a3_l0_parents_included` — L1 result has `parents` array containing L0 nodes. 4. `a4_level_filter_l0_only` — GET /query?level=L0 returns L0 nodes only (check level field). 5. `a5_level_filter_l1_l2` — GET /query?level=L1,L2 returns only L1 and L2 (no L0). 6. `a6_project_filter_works` — ingest into two projects, query with `project=poimen` — result.project matches. 7. `a7_limit_respected` — GET /query?limit=3 returns ≤3 results. 8. `a8_query_score_before_rerank` — query_score from HNSW comes before rerank; rerank_score ≤ query_score (reranker should not boost beyond HNSW recall). 9. `a9_timeout_enforced` — manually slow the embedding service (mock delay 10s), GET /query with `timeout_seconds=1` returns 503. 10. `a10_empty_query_returns_400` — GET /query (no query param) or GET /query?query= returns 400. **Command:** `cargo test -p mem-cli query_endpoint` **False pass:** - Testing only the happy path. Timeout, missing parent, embedding failure — all return different error codes. - Results sorted by query_score, not rerank_score. Reranking must reorder the results. - Parent nodes fetched but never asserted. A result with empty `parents` passes all checks. - query_score computed correctly but rerank_score always zero. Both must be present and in [0,1]. ## Traps - Timeout is wall-clock time, not per-service timeout. A 5s timeout that calls embedding (200ms) + HNSW (100ms) + rerank (500ms) should complete in <5s total, not each. Use `tokio::time::timeout()` around the entire handler. - HNSW uses `<->` operator for cosine distance (0 = opposite, 1 = same). 1 - distance is correct for similarity; do not invert again. - Reranker scores are [0,100]; dividing by 100 gives [0,1]. Not dividing is a common bug. - Embedding cache: the same query text submitted twice should reuse the embedding (save 200ms). Easy to forget. --- Background: [DESIGN.md § Distributed API Layer](../DESIGN.md#distributed-api-layer-homelab-frontend)