Files
poimen-memory/tasks/M3.6.3-mem-ref-cli.md

5.6 KiB
Raw Permalink Blame History

M3.6.3 — mem ref — corpus management with replace-on-change

Field Value
Phase M3.6 — Reference corpora
Size M — 13 days
Status Not started
Flags
Spec inlined below
Blocks M3.6.6
Depends M3.6.2, M2.1

Goal

Add, list, refresh and remove reference corpora, so that re-running an ingest against changed upstream docs replaces what is there instead of stacking a second copy beside it.

Facts (inlined — no spec read needed)

mem ref add --project homelab --corpus kubectl ~/workplace/homelab/knowledge/cheatsheets
mem ref add --dry-run ...            # chunk plan only, zero model calls (M3.6.1)
mem ref list --project homelab       # corpus, docs, chunks, last ingest, drift
mem ref sync --corpus kubectl        # re-walk, replace changed docs, report
mem ref rm --corpus kubectl          # tombstone every doc in the corpus

Identity is (source_uri, doc_sha). Same URI and same sha is a no-op: no embed call, no write, exit 0 with "unchanged". Same URI and different sha is a replace: tombstone the old chunks in the log, write the new ones. A URI that has vanished from the tree on a sync is a tombstone with no successor.

Tombstone, do not delete. The log is append-only and authoritative. A {"kind":"reference_tombstone","sha256":"…","reason":"replaced"} record is what removal means; the projector drops the row and the note on replay. Deleting rows from Postgres directly makes the index un-rebuildable, which is the one thing the whole design refuses.

Embedding is the expensive step, so skip it precisely. A corpus of 400 chunks where one document changed should issue embeddings for that document's chunks only. Chunk-level sha comparison, not document-level re-embed.

list reports drift. For each corpus, re-stat the tree and compare doc shas without writing anything: 3 docs changed, 1 removed, 12 unchanged. Drift that is only discoverable by running sync means nobody runs sync.

Steps

  1. mem ref add [--project P] --corpus C [--dry-run] <path> — walk via DocCorpusSource, embed new chunks, write Reference records.
  2. Persist corpus registration (name, root path, project, last ingest) in the log as a reference_corpus record so list needs no side file.
  3. Implement chunk-level diff: existing shas for the corpus vs freshly walked shas → {new, changed, unchanged, gone}.
  4. mem ref sync — apply the diff, embedding only new and changed, emitting tombstones for gone.
  5. mem ref list — table per corpus with counts plus a dry drift check.
  6. mem ref rm — tombstone every live chunk in the corpus; leave the log intact.
  7. Exit codes: 0 on success including no-op, non-zero on unresolvable corpus or unreadable root.

Acceptance

  • add twice on an unchanged tree issues zero embedding calls the second time.
  • Editing one file and running sync re-embeds that file's chunks only.
  • Deleting a file and running sync tombstones its chunks, and it stops appearing in query results.
  • rm removes the corpus from results while leaving every record in the log.
  • list reports drift without mutating anything.

Verify

Harness: fixture tree copied to a temp dir so it can be mutated, a counting embedder that records how many texts it was asked to embed, seeded database.

Integration testtests/it_mem_ref.rs:

  1. a1_add_then_add_is_noop — run add twice; assert the embedder call count is zero on the second run and the row count is unchanged.
  2. a2_changed_doc_reembeds_only_itself — edit one file of three, sync; assert embed count equals that file's chunk count, not the corpus total.
  3. a3_replace_tombstones_predecessor — after a change, assert the old chunk sha has a tombstone record and no live row.
  4. a4_removed_doc_tombstoned — delete a file, sync; assert its chunks are gone from memory_node and present in the log.
  5. a5_rm_preserves_log — count log lines before and after rm; assert the count only grew.
  6. a6_rebuild_after_churn — after add/change/sync/rm, mem rebuild --from-log; assert the reconstructed state matches the live state exactly.
  7. a7_list_is_read_only — snapshot database and log, run list, assert both unchanged and that reported drift matches the mutations made.
  8. a8_unreadable_root_exits_nonzero — point add at a missing path; assert non-zero exit and no partial corpus registration.

Command: cargo test -p mem-cli mem_ref

False pass:

  • Asserting "no duplicate rows" instead of counting embedder calls. A sync that re-embeds everything and then upserts by sha produces a correct table and a bill; assertion 1 and 2 are the only ones that see it.
  • Verifying tombstones by querying memory_node. The row being absent is the projector working; assertion 3 has to read the log to prove the tombstone was actually written and the row was not just deleted.
  • Running the churn test without a final rebuild. Assertion 6 is what proves the tombstone replay logic exists rather than being implied.

Traps

  • Registering the corpus before the walk succeeds. A failed add that leaves a registered-but-empty corpus makes the next sync report every document as new.
  • Comparing document mtime instead of sha. Checkouts and rsync rewrite mtimes; a corpus that re-embeds on every clone costs real money on the TEI endpoint.
  • Making rm delete log records "because they are noise". That converts the log from authoritative to advisory, and nothing downstream can tell.

Background: DESIGN.md — reference corpora, JSONL event log