Files
poimen-memory/tasks/M3.6.5-query-levels-and-floor.md

137 lines
6.4 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# M3.6.5 — Query: filter-then-recall, R opt-in, relevance floor
| Field | Value |
|---|---|
| Phase | M3.6 — Reference corpora |
| Size | M — 13 days |
| Status | ⬜ Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | M3.6.6 |
| Depends | M3.6.2, M3.3, M3.2, M2.7 |
## Goal
Make R reachable on request, unreachable by default, and stop the retriever
answering questions it has no evidence for.
## Facts (inlined — no spec read needed)
```
mem query "why did requests over 10KB fail?" # L1,L2 — unchanged
mem query --levels R "kubectl describe pod" # reference only
mem query --levels L1,L2,R "..." # both, R marked in output
mem query "…" --min-score 0.4 # override the floor
```
**Filter before recall, not after.** M3.3 recalls `10×k` from HNSW and reranks
down to `k`. A corpus is typically an order of magnitude larger than the project's
own memory, so R rows compete for those 50 candidate slots even when the caller
excluded them — and post-filtering then returns three results instead of five,
quietly. The level predicate belongs in the SQL that drives the HNSW scan. The
existing `(project, level)` index already supports it.
**Abstention.** With a corpus loaded, every question has *something* moderately
close, so unconditional top-k starts returning plausible prose for questions the
memory cannot answer — worse than an empty result, because it reads as an answer.
If the best post-rerank score is below the floor, return no hits and say why:
```
no hits above relevance floor (best 0.21 < 0.35 threshold)
try --min-score to lower it, or --levels R to search reference docs
```
The floor applies to the **reranked** score, not cosine distance. M3.2's own
fixture separates a relevant from an irrelevant passage by four orders of
magnitude; cosine distance does not, which is why the floor cannot live at the
recall stage.
**R is visually distinct in output.** A reference hit prints its source URI and
heading path where a project hit prints provenance. A caller must never have to
infer from wording whether an answer came from this cluster's history or from
upstream documentation.
**R has no provenance walk.** M3.3 walks `memory_edge` one hop for L1 and two for
L2. R has no edges by construction (M3.6.2), so the walk is skipped rather than
returning empty — and `mem verify` gains the assertion that makes that safe.
## Steps
1. Push the level filter into the recall query; assert candidate width is `10×k`
*after* filtering.
2. `--levels` accepts `R`; default remains `L1,L2`.
3. Apply the relevance floor to reranked scores; `--min-score` overrides,
`--min-score 0` disables.
4. Abstention message names the best score, the threshold, and the two escapes.
5. Render R hits with source URI and heading path; suppress the provenance walk.
6. Exit code: abstention is exit 0 with no hits, not an error — it is a valid
answer. Unresolvable project stays non-zero (M3.3 assertion 7).
7. `mem verify --edges` asserts no `memory_edge` row names an R sha as parent.
## Acceptance
- Default query over a database containing a large corpus returns exactly the
same hits as before the corpus was added.
- `--levels R` returns reference sections with URI and heading path.
- A question with no good match returns nothing and explains itself.
- Lowering `--min-score` surfaces the suppressed hits.
- `mem verify` rejects a hand-inserted `L1 -> R` edge.
## Verify
**Harness:** seeded database with the poimen log *plus* a reference corpus large
enough to dominate raw recall — at least 10× the project node count. Live
reranker for scoring assertions, deterministic embedder elsewhere.
**Integration test**`tests/it_query_levels.rs`:
1. `a1_default_unchanged_by_corpus` — snapshot default query results before and
after ingesting the corpus; assert byte-identical output. This is the
headline assertion of the task.
2. `a2_filter_before_recall` — instrument the repository; assert the SQL driving
HNSW carries the level predicate and returns `10×k` rows post-filter, not
`10×k` pre-filter then fewer.
3. `a3_levels_r_returns_reference``--levels R` returns R nodes with source URI
and heading path populated.
4. `a4_floor_abstains` — a question with no relevant content returns zero hits,
exit 0, message naming best score and threshold.
5. `a5_floor_override_recovers` — same question with `--min-score 0` returns the
suppressed hits, proving abstention is a floor and not a bug upstream.
6. `a6_floor_applies_post_rerank` — construct a case where cosine is high and
rerank is low; assert it is suppressed. The reverse ordering passes every
other assertion here.
7. `a7_r_hits_visually_distinct` — assert R hits carry no provenance block and do
carry a URI, in both human and `--format json` output.
8. `a8_no_edge_to_r` — hand-insert an `L1 -> R` edge; assert `mem verify --edges`
exits non-zero and names the offending pair.
9. `a9_mixed_levels_ordering``--levels L1,L2,R` returns both kinds ranked
together with the level labelled on every row.
**Command:** `cargo test -p mem-cli query_levels`
**False pass:**
- Running assertion 1 against a small corpus. If the corpus is smaller than the
recall width, post-filtering and pre-filtering give the same answer and
assertion 2 is the only thing separating them — the fixture size is part of the
test.
- Testing abstention with a query that matches nothing at all. Zero recall
returns zero hits regardless of the floor; the fixture needs a *weak but
non-empty* match, or assertion 4 passes with the floor unimplemented.
- Asserting `--levels R` works without asserting the default excludes R. Both
directions are the contract.
## Traps
- Applying the floor to the first-stage cosine score. Cosine on `nomic` puts
unrelated text closer than intuition suggests; a floor there either suppresses
good hits or does nothing, depending on the corpus.
- Treating abstention as an error exit. Callers wrap `mem query` in scripts; a
non-zero exit for "no confident answer" turns a normal outcome into a pipeline
failure and the floor gets disabled within a week.
- Letting the reranker see 50 R candidates and 3 project candidates in a mixed
query. The reranker is not calibrated across levels, and the corpus wins on
fluency; recall per level, then merge.
---
Background: [DESIGN.md](../DESIGN.md) — reference corpora, pgvector, retrieval