Files
poimen-memory/tasks/M3.3-mem-query.md
T

98 lines
3.8 KiB
Markdown
Raw Normal View History

2026-08-19 09:52:07 -07:00
# M3.3 — `mem query` with provenance
| Field | Value |
|---|---|
| Phase | M3 — L2 synthesis and retrieval |
| Size | M — 13 days |
| Status | ⬜ Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | M3.2, M2.4 |
## Goal
Ask the memory a question and get an answer that can be traced back to the
session it came from.
## Facts (inlined — no spec read needed)
```
mem query "why did requests over 10KB fail?"
mem query --project poimen --levels L1,L2 --k 10 "..."
```
Pipeline: embed the question (M2.1) → HNSW recall top-k over `memory_node`
filtered by project and level (M2.4) → rerank the candidates (M3.2) → return with
provenance walked down through `memory_edge`.
Default levels are **L1 and L2**, not L0. L1/L2 are synthesized answers; L0 is raw
evidence and returning it by default buries the answer in transcript. `--levels
L0` exists for "show me the actual source".
Provenance walk: for each returned node, follow `memory_edge` to its parents and
report source and turn. One hop for L1 (to evidence), two for L2 (through L1).
Recall wide, rerank narrow: take 50 from HNSW, rerank, return 5. Cosine distance
alone puts "bananas" close enough to matter; the reranker separated the same pair
by four orders of magnitude.
Output is human-readable by default, `--format json` for programmatic use.
## Steps
1. `mem query [--project P] [--levels L] [--k N] <question>`.
2. Embed, recall 10×k, rerank, truncate to k.
3. Walk edges to build a provenance list per hit; deduplicate sources.
4. Render: score, level, query id, the memory text, then provenance lines.
5. Project defaults to the one inferred from `$PWD`; `--project` overrides. No
project and no match is an error, not an empty result over everything.
6. `--explain` prints the recall candidates and their pre-rerank distances, for
debugging retrieval quality.
## Acceptance
- A question with a known answer returns the right L1 note first.
- Every hit carries at least one resolvable provenance entry.
- `--levels L0` returns evidence; the default does not.
- Reranking changes the order versus raw recall on at least one real query.
## Verify
**Harness:** a seeded database from a real log, deterministic fake embedder for
the offline assertions, live for quality.
**Integration test**`tests/it_query.rs`:
1. `a1_known_answer` — seeded with the poimen log, query "why did requests over
10KB fail?"; assert the top hit is the `infra-root-causes` L1 node.
2. `a2_provenance_resolves` — every hit's provenance shas exist in
`memory_node`.
3. `a3_default_excludes_l0` — assert no L0 nodes in default output.
4. `a4_levels_flag``--levels L0` returns evidence nodes.
5. `a5_rerank_reorders` — capture pre- and post-rerank order; assert they differ
on at least one fixture query, proving the reranker is wired and not a no-op.
6. `a6_project_isolation` — two projects seeded; assert no cross-project hits.
7. `a7_no_project_errors` — unresolvable project exits non-zero.
8. `a8_l2_two_hop_provenance` — an L2 hit's provenance resolves through L1 to L0
sources.
**Command:** `cargo test -p mem-cli query`
**False pass:**
- Asserting only that results are returned. A pipeline where the reranker returns
the input unchanged still returns results, and assertion 5 is the only check
that it is doing anything.
- Testing provenance existence without resolving it. A hit carrying parent shas
that do not exist in the table looks fine in the output and is useless.
## Traps
- Returning L0 by default. The answer is there but buried in raw transcript, and
the tool reads as low quality when it is a display default.
- Recalling exactly k then reranking. Reranking cannot recover a relevant
document that recall never returned; the width of the recall is what determines
the ceiling.
---
Background: [DESIGN.md](../DESIGN.md) — pgvector, retrieval