Files
poimen-memory/tasks/M3.6.3-mem-ref-cli.md
T

119 lines
5.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# M3.6.3 — `mem ref` — corpus management with replace-on-change
| Field | Value |
|---|---|
| Phase | M3.6 — Reference corpora |
| Size | M — 13 days |
| Status | ⬜ Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | M3.6.6 |
| Depends | M3.6.2, M2.1 |
## Goal
Add, list, refresh and remove reference corpora, so that re-running an ingest
against changed upstream docs replaces what is there instead of stacking a second
copy beside it.
## Facts (inlined — no spec read needed)
```
mem ref add --project homelab --corpus kubectl ~/workplace/homelab/knowledge/cheatsheets
mem ref add --dry-run ... # chunk plan only, zero model calls (M3.6.1)
mem ref list --project homelab # corpus, docs, chunks, last ingest, drift
mem ref sync --corpus kubectl # re-walk, replace changed docs, report
mem ref rm --corpus kubectl # tombstone every doc in the corpus
```
**Identity is `(source_uri, doc_sha)`.** Same URI and same sha is a no-op: no
embed call, no write, exit 0 with "unchanged". Same URI and different sha is a
*replace*: tombstone the old chunks in the log, write the new ones. A URI that
has vanished from the tree on a `sync` is a tombstone with no successor.
**Tombstone, do not delete.** The log is append-only and authoritative. A
`{"kind":"reference_tombstone","sha256":"…","reason":"replaced"}` record is what
removal means; the projector drops the row and the note on replay. Deleting rows
from Postgres directly makes the index un-rebuildable, which is the one thing the
whole design refuses.
**Embedding is the expensive step, so skip it precisely.** A corpus of 400 chunks
where one document changed should issue embeddings for that document's chunks
only. Chunk-level sha comparison, not document-level re-embed.
**`list` reports drift.** For each corpus, re-stat the tree and compare doc shas
without writing anything: `3 docs changed, 1 removed, 12 unchanged`. Drift that
is only discoverable by running `sync` means nobody runs `sync`.
## Steps
1. `mem ref add [--project P] --corpus C [--dry-run] <path>` — walk via
`DocCorpusSource`, embed new chunks, write `Reference` records.
2. Persist corpus registration (name, root path, project, last ingest) in the
log as a `reference_corpus` record so `list` needs no side file.
3. Implement chunk-level diff: existing shas for the corpus vs freshly walked
shas → `{new, changed, unchanged, gone}`.
4. `mem ref sync` — apply the diff, embedding only `new` and `changed`, emitting
tombstones for `gone`.
5. `mem ref list` — table per corpus with counts plus a dry drift check.
6. `mem ref rm` — tombstone every live chunk in the corpus; leave the log intact.
7. Exit codes: 0 on success including no-op, non-zero on unresolvable corpus or
unreadable root.
## Acceptance
- `add` twice on an unchanged tree issues zero embedding calls the second time.
- Editing one file and running `sync` re-embeds that file's chunks only.
- Deleting a file and running `sync` tombstones its chunks, and it stops
appearing in query results.
- `rm` removes the corpus from results while leaving every record in the log.
- `list` reports drift without mutating anything.
## Verify
**Harness:** fixture tree copied to a temp dir so it can be mutated, a counting
embedder that records how many texts it was asked to embed, seeded database.
**Integration test**`tests/it_mem_ref.rs`:
1. `a1_add_then_add_is_noop` — run `add` twice; assert the embedder call count is
zero on the second run and the row count is unchanged.
2. `a2_changed_doc_reembeds_only_itself` — edit one file of three, `sync`; assert
embed count equals that file's chunk count, not the corpus total.
3. `a3_replace_tombstones_predecessor` — after a change, assert the old chunk
sha has a tombstone record and no live row.
4. `a4_removed_doc_tombstoned` — delete a file, `sync`; assert its chunks are
gone from `memory_node` and present in the log.
5. `a5_rm_preserves_log` — count log lines before and after `rm`; assert the
count only grew.
6. `a6_rebuild_after_churn` — after add/change/sync/rm, `mem rebuild --from-log`;
assert the reconstructed state matches the live state exactly.
7. `a7_list_is_read_only` — snapshot database and log, run `list`, assert both
unchanged and that reported drift matches the mutations made.
8. `a8_unreadable_root_exits_nonzero` — point `add` at a missing path; assert
non-zero exit and no partial corpus registration.
**Command:** `cargo test -p mem-cli mem_ref`
**False pass:**
- Asserting "no duplicate rows" instead of counting embedder calls. A `sync` that
re-embeds everything and then upserts by sha produces a correct table and a
bill; assertion 1 and 2 are the only ones that see it.
- Verifying tombstones by querying `memory_node`. The row being absent is the
projector working; assertion 3 has to read the log to prove the tombstone was
actually written and the row was not just deleted.
- Running the churn test without a final rebuild. Assertion 6 is what proves the
tombstone replay logic exists rather than being implied.
## Traps
- Registering the corpus before the walk succeeds. A failed `add` that leaves a
registered-but-empty corpus makes the next `sync` report every document as new.
- Comparing document mtime instead of sha. Checkouts and rsync rewrite mtimes;
a corpus that re-embeds on every clone costs real money on the TEI endpoint.
- Making `rm` delete log records "because they are noise". That converts the log
from authoritative to advisory, and nothing downstream can tell.
---
Background: [DESIGN.md](../DESIGN.md) — reference corpora, JSONL event log