Files
poimen-memory/tasks/M3.3-mem-query.md
T

3.8 KiB
Raw Blame History

M3.3 — mem query with provenance

Field Value
Phase M3 — L2 synthesis and retrieval
Size M — 13 days
Status Not started
Flags
Spec inlined below
Blocks M3.2, M2.4

Goal

Ask the memory a question and get an answer that can be traced back to the session it came from.

Facts (inlined — no spec read needed)

mem query "why did requests over 10KB fail?"
mem query --project poimen --levels L1,L2 --k 10 "..."

Pipeline: embed the question (M2.1) → HNSW recall top-k over memory_node filtered by project and level (M2.4) → rerank the candidates (M3.2) → return with provenance walked down through memory_edge.

Default levels are L1 and L2, not L0. L1/L2 are synthesized answers; L0 is raw evidence and returning it by default buries the answer in transcript. --levels L0 exists for "show me the actual source".

Provenance walk: for each returned node, follow memory_edge to its parents and report source and turn. One hop for L1 (to evidence), two for L2 (through L1).

Recall wide, rerank narrow: take 50 from HNSW, rerank, return 5. Cosine distance alone puts "bananas" close enough to matter; the reranker separated the same pair by four orders of magnitude.

Output is human-readable by default, --format json for programmatic use.

Steps

  1. mem query [--project P] [--levels L] [--k N] <question>.
  2. Embed, recall 10×k, rerank, truncate to k.
  3. Walk edges to build a provenance list per hit; deduplicate sources.
  4. Render: score, level, query id, the memory text, then provenance lines.
  5. Project defaults to the one inferred from $PWD; --project overrides. No project and no match is an error, not an empty result over everything.
  6. --explain prints the recall candidates and their pre-rerank distances, for debugging retrieval quality.

Acceptance

  • A question with a known answer returns the right L1 note first.
  • Every hit carries at least one resolvable provenance entry.
  • --levels L0 returns evidence; the default does not.
  • Reranking changes the order versus raw recall on at least one real query.

Verify

Harness: a seeded database from a real log, deterministic fake embedder for the offline assertions, live for quality.

Integration testtests/it_query.rs:

  1. a1_known_answer — seeded with the poimen log, query "why did requests over 10KB fail?"; assert the top hit is the infra-root-causes L1 node.
  2. a2_provenance_resolves — every hit's provenance shas exist in memory_node.
  3. a3_default_excludes_l0 — assert no L0 nodes in default output.
  4. a4_levels_flag--levels L0 returns evidence nodes.
  5. a5_rerank_reorders — capture pre- and post-rerank order; assert they differ on at least one fixture query, proving the reranker is wired and not a no-op.
  6. a6_project_isolation — two projects seeded; assert no cross-project hits.
  7. a7_no_project_errors — unresolvable project exits non-zero.
  8. a8_l2_two_hop_provenance — an L2 hit's provenance resolves through L1 to L0 sources.

Command: cargo test -p mem-cli query

False pass:

  • Asserting only that results are returned. A pipeline where the reranker returns the input unchanged still returns results, and assertion 5 is the only check that it is doing anything.
  • Testing provenance existence without resolving it. A hit carrying parent shas that do not exist in the table looks fine in the output and is useless.

Traps

  • Returning L0 by default. The answer is there but buried in raw transcript, and the tool reads as low quality when it is a display default.
  • Recalling exactly k then reranking. Reranking cannot recover a relevant document that recall never returned; the width of the recall is what determines the ceiling.

Background: DESIGN.md — pgvector, retrieval