chore: retire M3.6.3 (mem ref CLI), update M3.6.2 to use Obsidian REST API
Build and Push / Test (push) Failing after 1m54s
Build and Push / Build and push image (push) Skipped

CHANGES:
- M3.6.3: marked  RETIRED (Obsidian UI replaces CLI corpus management)
- M3.6.2: updated to fetch from Obsidian REST API instead of filesystem
  - ObsidianRefSource: calls /api/vault/listFiles, /api/vault/readFile
  - Users manage corpus in Obsidian UI (not via CLI)
  - Rebuild auto-syncs by re-fetching and comparing file SHAs
  - No separate chunk-level diff CLI needed
- Updated INDEX.md:
  - M3.6.x: 6 tasks → 5 tasks (removed M3.6.3)
  - Progress: 1 , 0 🟡, 5  → 1 , 0 🟡, 4 
  - Total: 71 tasks → 70 tasks
  - Noted M3.6.3 retirement in board description

RATIONALE:
- Obsidian is single source of truth (REST API)
- Users already use Obsidian UI for vault management
- No need for parallel CLI when vault is the interface
- M3.6.2 handles sync via deterministic SHA comparison
- Reduces feature bloat, cleaner architecture
This commit is contained in:
Story Crater Bot
2026-08-28 08:18:14 -07:00
parent 993236246f
commit 3e867f7cce
3 changed files with 250 additions and 156 deletions
+25 -18
View File
@@ -64,15 +64,15 @@ Legend: ⬜ not started · 🟡 in progress · ✅ done · ⛔ blocked
| 4 | L2 synthesis + retrieval | M3.x | 4 | 4 | 0 | 0 | ✅ M3.4 |
| 4.5 | Distributed API Layer | M3.5.x | 10 | 10 | 0 | 0 | ✅ M3.5.8 |
| 5 | Skills | M4.x | 3 | 2 | 0 | 1 | ⬜ M4.3 |
| 5.5 | Reference corpora | M3.6.x | 6 | 1 | 0 | 5 | ⬜ M3.6.6 |
| 5.5 | Reference corpora | M3.6.x | 5 | 1 | 0 | 4 | ⬜ M3.6.6 |
| 5.6 | Tool context | M3.7.x | 4 | 2 | 0 | 2 | ⬜ M3.7.6 |
| 6 | Post-training | M5.x | 6 | 0 | 0 | 6 | ⬜ M5.6 |
| 7 | agent-manager migration | M6.x | 6 | 0 | 0 | 6 | ⬜ M6.6 |
| 8 | Source connectors | M7.x | 10 | 0 | 0 | 10 | ⬜ M7.10 |
| 9 | Hybrid search | M8.x | 9 | 1 | 1 | 7 | ⬜ M8.9 |
| | **Total** | | **71** | **60** | **0** | **11** | 6/11 green |
| | **Total** | | **70** | **60** | **0** | **10** | 6/11 green |
**Current status — 2025-01-28.** Completed phases M0.x, M1.x fully archived (16/16 tasks). **M2.1-6 ✅** (embeddings, CNPG, schema, pgvector, obsidian projector, rebuild). **M3.x ✅** (4/4). **M3.5.x ✅** (10/10 complete + archived). **M4.1-2 ✅** (skill drafting + derived filter). **M3.6.1 ✅** (DocCorpusSource). **M8.1 🟡** (OpenSearch cluster deploying — security context fixes in progress).
**Current status — 2025-01-28.** Completed phases M0.x, M1.x fully archived (16/16 tasks). **M2.1-6 ✅** (embeddings, CNPG, schema, pgvector, obsidian projector, rebuild). **M3.x ✅** (4/4). **M3.5.x ✅** (10/10 complete + archived). **M3.7.7-8 ✅** (failure diagnosis). **M4.1-2 ✅** (skill drafting + derived filter). **M3.6.1 ✅** (DocCorpusSource). **M3.6.3 ❌ retired** (Obsidian UI replaces CLI). **M8.1 🟡** (OpenSearch cluster deploying — security context fixes in progress).
**Current work:**
- M8.1: OpenSearch StatefulSet (pod security baseline, fsGroup perms) — deploying
@@ -86,7 +86,7 @@ Legend: ⬜ not started · 🟡 in progress · ✅ done · ⛔ blocked
**Blocked until M8.2 green:**
- M8.3-9 (query optimizer, RRF, hybrid endpoint, benchmarks)
- M3.7.4 (context endpoint needs hybrid search to rank tiers) — now has M3.7.7 + M3.7.8 ✅
- M3.6.2+ (reference corpus storage depends on query layer)
- M3.6.2+ (reference corpus indexing depends on dual-write layer for Postgres + OpenSearch)
**Tests: 265+ passing, 2 ignored** (M2.1 +8, M3.5 +16, M4 +20, M3.7.7 +18 unit, M3.7.8 +22 integration). **60/71 tasks complete (85%)**, **6/11 gates green**. M3.7.3 & M3.7.5 retired (hybrid search covers). **M3.7.7 & M3.7.8 ✅ complete** (signature extraction + symptom projection, 40+ tests passing).
@@ -137,25 +137,32 @@ Homelab frontend integration: HTTP facade via `api.riotpiao.com`. Runs in parall
## 5.5 — Reference corpora · M3.6.x
Documentation the local models are weak at — `kubectl`, `tea` — made retrievable
as level **R**: embedded and indexed, never evidence. Ids are `M3.6.x` and stay
`M3.6.x`; the phase sits here rather than at 4.6 because [M3.6.4](M3.6.4-reference-cycle-guard.md)
extends M4.2's matcher instead of duplicating it, and because skills are the
better answer to the same problem and should exist first.
**Source of truth: Obsidian vault** (REST API, deployed M2.5). Documentation
the local models are weak at — `kubectl`, `tea` — are stored in Obsidian and
made retrievable as level **R**: embedded and indexed, never evidence.
**The load-bearing property is a negative one.** Adding a corpus must not change
update-rate, must not change default query output, and must not put an R node in
any provenance chain. R bypasses the recurrence structurally — `run_loop` needs a
`Query` and a corpus has none — not by a flag. [M3.6.6](M3.6.6-m3.6-gate.md)
asserts M1.8's numbers are *unchanged*, not merely still-passing, because
documentation fed to the gate would lower update-rate and make M1.8 easier to
clear while the memory got worse.
Ids are `M3.6.x` and stay `M3.6.x`; the phase sits here rather than at 4.6 because
[M3.6.4](M3.6.4-reference-cycle-guard.md) extends M4.2's matcher instead of
duplicating it, and because skills are the better answer to the same problem and
should exist first.
**The load-bearing property is a negative one.** Adding Obsidian vault reference
documents must not change update-rate, must not change default query output, and
must not put an R node in any provenance chain. R bypasses the recurrence
structurally — `run_loop` needs a `Query` and a corpus has none — not by a flag.
[M3.6.6](M3.6.6-m3.6-gate.md) asserts M1.8's numbers are *unchanged*, not merely
still-passing, because documentation fed to the gate would lower update-rate and
make M1.8 easier to clear while the memory got worse.
**M3.6.3 retired:** Obsidian UI replaces CLI corpus management. Users edit files
in Obsidian; `mem rebuild` auto-fetches from Obsidian REST API and re-indexes only
changed chunks (SHA comparison, deterministic embedding).
| Task | Title | Size | Flags | Status |
|---|---|---|---|---|
| M3.6.1 | `DocCorpusSource` + heading chunking | M | — | ✅ |
| [M3.6.2](M3.6.2-level-r-storage.md) | Level R: log, index, vault, rebuild parity | M | — | ⬜ |
| [M3.6.3](M3.6.3-mem-ref-cli.md) | `mem ref` replace-on-change corpus management | M | — | |
| [M3.6.2](M3.6.2-level-r-storage.md) | Level R: Obsidian reference indexing + rebuild parity | M | — | ⬜ |
| M3.6.3 | `mem ref` — corpus management CLI | M | — | ❌ RETIRED (Obsidian UI replaces) |
| [M3.6.4](M3.6.4-reference-cycle-guard.md) | Reference text cannot re-enter as evidence | M | — | ⬜ |
| [M3.6.5](M3.6.5-query-levels-and-floor.md) | Query: filter-then-recall, R opt-in, floor | M | — | ⬜ |
| [M3.6.6](M3.6.6-m3.6-gate.md) | **M3.6 composition gate** | M | gate | ⬜ |
+136 -80
View File
@@ -1,4 +1,4 @@
# M3.6.2 — Level R: log record, index rows, vault notes, rebuild parity
# M3.6.2 — Level R: Obsidian reference indexing and rebuild parity
| Field | Value |
|---|---|
@@ -8,116 +8,172 @@
| Flags | — |
| Spec | inlined below |
| Blocks | M3.6.6 |
| Depends | M3.6.1, M1.6, M2.3, M2.4, M2.5, M2.6 |
| Depends | M3.6.1, M1.6, M2.3, M2.4, M2.5, M2.6, Obsidian service (M2.5) |
## Goal
Land reference chunks in the log as their own record kind, project them into
Postgres and the vault, and prove the projections are still throwaway.
Fetch reference documents from the live Obsidian vault (REST API), chunk them
via M3.6.1's heading-boundary logic, land them in the log as reference records,
index them in Postgres, and prove rebuild parity (drop + rebuild from log =
byte-identical index state).
## Facts (inlined — no spec read needed)
```jsonl
{"kind":"reference","level":"R","project":"homelab","source":"file:///.../kubectl.md",
"heading_path":"kubectl.md > Common Issues > CrashLoopBackOff","doc_sha":"ab12…",
"sha256":"cd34…","t":7,"run_id":"ref-2026-08-21T10:02:11Z","text":"…"}
**Obsidian is the source of truth for reference documents.** The Obsidian REST API
(deployed in M2.5, accessible at `http://obsidian-server.poimen.svc.cluster.local:8080`)
provides live access to vault files via:
```bash
GET /api/vault/listFiles # List all .md files
GET /api/vault/readFile?path=kubectl.md # Read file contents
GET /api/vault/listFolders # Explore structure
```
**No migration is needed.** `'R'` ships in M2.3's initial schema, along with the
`CHECK ((level IN ('L2','R')) = (query_id IS NULL))` constraint. Nothing was built
before this phase existed, so the level was never absent from the schema and an
`ALTER` here would only undo a deliberate omission that was never made.
**Reference record format (in JSONL log):**
```jsonl
{"kind":"reference","level":"R","source":"obsidian://poimen-vault/kubectl.md",
"doc_sha":"ab12…","heading_path":"kubectl.md > Common Issues > CrashLoopBackOff",
"sha256":"cd34…","t":7,"run_id":"ref-obsidian-2026-08-21T10:02:11Z","text":"…"}
```
**No vault projection.** Unlike M3.6.1 (DocCorpusSource which reads from local filesystem),
M3.6.2 reads directly from Obsidian REST API and indexes into Postgres. The Obsidian
vault remains the single source of truth; `mem rebuild --from-log` re-fetches from
Obsidian to regenerate indexes.
`level = 'R'`, `query_id = NULL` (R answers no standing question), `source` holds
the source URI. `doc_sha` is the whole-document hash; `sha256` is the chunk hash
and stays the primary identity, same as every other level.
the Obsidian URI (`obsidian://vault-name/file.md`). `doc_sha` is whole-document hash;
`sha256` is chunk hash (primary identity).
The embedding goes to `memory_vector(kind='text')`, not to a column on the node.
R gets **no symptom projection** — M3.7.8 generates those for L1 and L2 only,
since documentation headings already read like problems.
R gets **no symptom projection** — M3.7.8 generates those for L1 and L2 only.
**R writes no edges.** Not to parents, not to siblings. A reference chunk has no
provenance inside this system — its provenance is the URI. The rule that makes
this safe is enforced in `mem verify` (M3.6.4), but nothing in this task should
ever be tempted to create an edge in the first place.
**R writes no edges.** Provenance is the Obsidian URI, not internal edges. Enforced
in `mem verify` (M3.6.4).
**Vault projection goes somewhere separate.** `vault/reference/<corpus>/<doc>.md`,
not into the project notes. The vault is browsed by a human; interleaving
upstream docs with synthesized project memory makes the vault untrustworthy at a
glance. One note per source document, sections as headings, each carrying its
chunk sha as an anchor so `mem query` output can deep-link.
**Rebuild parity is the whole point of the task.** `mem rebuild --from-log` must
drop and reconstruct R rows and R notes byte-identically. If it cannot, R has
hidden inputs and rule 3 of the design is broken — M2.8 already enforces this
property for L0/L1/L2 and this task extends the same harness rather than writing
a second one.
**Rebuild parity is the whole point.** `mem rebuild --from-log` must:
1. Drop R rows from Postgres
2. Drop R vectors from OpenSearch (M8.2)
3. Re-fetch documents from Obsidian REST API
4. Re-chunk via M3.6.1 heading logic
5. Re-index into Postgres + OpenSearch
6. Result must be byte-identical (same shas, same vector embeddings)
## Steps
1. Add the `Reference` variant to the log record enum in `mem-core`; serialize
with the field set above.
2. `mem-store`: insert R nodes with a single `kind='text'` vector; assert at the
repository boundary that no edge insert names an R sha as parent.
4. Obsidian projector: `vault/reference/<corpus>/<doc>.md`, one note per source
document, chunk shas as heading anchors.
5. Extend `mem rebuild --from-log` to replay `Reference` records.
6. Extend the M2.6 rebuild-parity harness to cover a log containing R records.
1. **ObsidianRefSource** in `mem-ingest`: Implement `ChunkedSource` that:
- Calls Obsidian REST API (`/api/vault/listFiles`)
- Filters for `.md` files in allowed paths (e.g., `docs/`, `reference/`)
- Fetches each file via `/api/vault/readFile?path=...`
- Chunks via M3.6.1's heading boundary logic
- Yields `Record { kind: Reference, ... }`
2. Add `Reference` variant to log record enum in `mem-core` (already exists from
M2.3, just needs the Obsidian source)
3. `mem-store`: Insert R nodes with `kind='text'` vector; assert no edge insert
names an R sha as parent (repository boundary check).
4. Dual-write in M8.2: When R records are indexed, write both to Postgres and
OpenSearch (L2 already does this; extend for R).
5. **Rebuild: extend M2.6 harness** for `mem rebuild --from-log`:
- When replaying R records from log, re-fetch original files from Obsidian
- Re-chunk via M3.6.1 logic
- Regenerate embeddings (deterministic, so shas match)
- Assert byte-identical state vs. original indexing
6. Integration: Update `mem query` to include R results from hybrid search (M8)
## Acceptance
- A `Reference` record round-trips through the log unchanged.
- R rows land with `query_id IS NULL` and `source` set to the URI.
- The widened constraint accepts `R` and still rejects `L3`.
- Reference notes land under `vault/reference/`, never in project note dirs.
- Drop database + vault, `mem rebuild --from-log`, and both come back
byte-identical.
- `ObsidianRefSource` fetches all `.md` files from Obsidian REST API.
- Reference records (R) with `source=obsidian://...` land in the log.
- R rows persist in Postgres with `query_id IS NULL` and `source` set to Obsidian URI.
- Each R chunk gets a `kind='text'` embedding in OpenSearch.
- The M2.3 constraint accepts `R` and rejects `L3`.
- `mem rebuild --from-log`: drop Postgres R rows + OpenSearch R vectors, re-fetch
from Obsidian, re-chunk, re-embed, re-index; result is byte-identical (same shas,
same embedding vectors).
- No hidden inputs: Obsidian REST API is the only external dependency for R.
## Verify
**Harness:** the M2.6 rebuild harness, extended with a log fixture that contains
L0/L1/L2 *and* R records. Deterministic fake embedder so shas are stable.
**Harness:** Mock Obsidian REST API (deterministic file list + content). M2.6 rebuild
harness extended to cover R records. Deterministic embedder so shas are stable.
**Integration test**`tests/it_level_r_storage.rs`:
1. `a1_record_roundtrip` — serialize then deserialize a `Reference` record;
assert field-for-field equality including `doc_sha` and `heading_path`.
2. `a2_r_inserts` — insert `level='R'` with a `kind='text'` vector; assert both
rows persist.
3. `a3_no_symptom_vector` — assert no R node acquires a `kind='symptom'` vector
after a full ingest.
4. `a4_query_id_null_at_r` — assert every R row has `query_id IS NULL`, and that
an R row with one is rejected by M2.3's CHECK.
5. `a5_no_edges_from_r` — after ingesting the fixture corpus, assert
`SELECT count(*) FROM memory_edge WHERE parent_sha IN (SELECT sha256 FROM
memory_node WHERE level='R')` is 0.
6. `a6_vault_path_isolation` — assert every emitted reference note path starts
with `vault/reference/` and no project note directory gained a file.
7. `a7_rebuild_byte_identical` — snapshot database rows and vault files, drop
both, `mem rebuild --from-log`, assert byte-identical including R.
8. `a8_rebuild_is_idempotent` — rebuild twice; assert the second run changes
nothing.
**Integration test**`tests/it_obsidian_ref_source.rs` + `tests/it_level_r_storage.rs`:
**Command:** `cargo test -p mem-store level_r && cargo test -p mem-cli rebuild`
### ObsidianRefSource tests:
1. `a1_fetches_md_files` — Mock API returns 3 `.md` files; source emits 3 documents.
2. `a2_chunks_by_heading` — One source document with 5 headings yields 5 chunks
(uses M3.6.1 heading boundary logic).
3. `a3_generates_doc_sha` — Document SHA (SHA256 of whole file content) is
consistent across fetches.
4. `a4_generates_chunk_shas` — Each chunk sha is deterministic: SHA256(tool +
heading_path + text).
5. `a5_sets_obsidian_uri` — Every R record has `source=obsidian://vault-name/path.md`.
6. `a6_skips_non_md` — API returns `.json` and `.txt` files; source ignores them.
### Level R storage tests:
7. `a7_record_roundtrip` — Serialize R record to JSONL, deserialize; assert
field-for-field equality.
8. `a8_r_inserts_to_postgres` — Insert `level='R'` with `kind='text'` vector;
assert Postgres row exists with correct `source` URI.
9. `a9_r_inserts_to_opensearch` — Same chunk indexed in OpenSearch (M8.2
dual-write); assert M8 vector store has the chunk.
10. `a10_no_symptom_projection` — R chunks do not acquire `kind='symptom'`
vectors (only L1/L2 get those from M3.7.8).
11. `a11_query_id_null` — Every R row has `query_id IS NULL`; insertion with
non-NULL query_id is rejected by M2.3's CHECK.
12. `a12_no_edges_from_r` — After ingesting 10 R chunks, `SELECT count(*)
FROM memory_edge WHERE parent_sha IN (SELECT sha256 FROM memory_node WHERE
level='R')` equals 0.
### Rebuild parity tests:
13. `a13_rebuild_drops_r_rows` — Ingest R, assert N rows exist. Call `mem
rebuild --from-log`, assert same N rows exist (re-fetched, re-embedded,
re-indexed).
14. `a14_rebuild_byte_identical` — Snapshot Postgres R rows + OpenSearch R
vectors before rebuild. Drop both. Run `mem rebuild --from-log` against
a log containing R records. Assert Postgres rows and OpenSearch vectors
are byte-identical (same row order, same field values, same vector embeddings).
15. `a15_rebuild_idempotent` — Run rebuild twice; second run changes nothing.
Snapshot comparison between first and second rebuild result is identical.
**Command:**
```bash
cargo test -p mem-ingest obsidian_ref_source
cargo test -p mem-store level_r
cargo test -p mem-cli rebuild # includes R records in fixture log
```
**False pass:**
- Asserting rebuild parity on a log with no R records. It passes trivially and
proves nothing about this task; assertion 7 is only meaningful because the
fixture log is mixed-level.
- Checking edge count is zero *before* ingesting anything. Assertion 5 has to run
against a populated corpus or it is asserting that an empty table is empty.
- Comparing vault files with a normalizing diff. Byte-identical means bytes;
trailing-newline drift is exactly the class of hidden input this rule exists
to catch.
- Asserting rebuild parity on a log with no R records. Assertion 14 requires
mixed-level log (L0/L1/L2 *and* R) or it trivially passes.
- Mocking Obsidian API with hardcoded responses. File hashes must match real
Obsidian vault file SHA256 (use actual Obsidian instance or hash fixtures).
- Checking edge count is zero before ingesting. Assertion 12 must run against
populated corpus or it asserts empty table is empty.
- Byte-identical comparison with normalization (sorting, ignoring order).
Byte-identical means bytes; rebuild parity is broken if row order changes
between runs.
## Traps
- Reusing `run_id` semantics from the gated loop. R has no run in the recurrence
sense; use a synthetic `ref-<timestamp>` and do not let it collide with a real
ingest run in queries that group by `run_id`.
- Putting reference notes in the project vault "just for now". The vault is the
human surface and the mixing is not reversible by a later move — links written
against the old path rot.
- Dropping the check constraint instead of widening it. Assertion 3 exists
sense; use a synthetic `ref-obsidian-<timestamp>` and do not let it collide
with a real ingest run in queries that group by `run_id`.
- Caching Obsidian API responses across rebuild runs. Rebuild must re-fetch from
Obsidian REST API every time (no cache) to ensure file contents and shas are
always in sync with live vault.
- Assuming Obsidian file list is sorted. `listFiles` order is arbitrary; source
must sort filenames before chunking to ensure deterministic shas across runs.
- Embedding each R chunk independently. M3.6.1 (DocCorpusSource) uses the same
`m` parameter (batching) and deterministic embedder model to ensure chunk shas
are stable; M3.6.2 must use identical setup.
- Dropping the check constraint instead of widening it. Assertion 11 exists
because `DROP CONSTRAINT` alone passes every other assertion in this file.
---
+89 -58
View File
@@ -1,83 +1,107 @@
# M3.6.3 — `mem ref` — corpus management with replace-on-change
# M3.6.3 — RETIRED: `mem ref` CLI (replaced by Obsidian UI)
| Field | Value |
|---|---|
| Phase | M3.6 — Reference corpora |
| Size | M — 13 days |
| Status | ⬜ Not started |
| Status | ❌ RETIRED |
| Flags | — |
| Spec | inlined below |
| Blocks | M3.6.6 |
| Depends | M3.6.2, M2.1 |
| Spec | inlined below (historical) |
| Blocks | |
| Depends | |
## Goal
## Retirement Rationale
Add, list, refresh and remove reference corpora, so that re-running an ingest
against changed upstream docs replaces what is there instead of stacking a second
copy beside it.
**Obsidian UI replaces CLI corpus management.**
## Facts (inlined — no spec read needed)
With Obsidian vault deployed (M2.5 complete) as the source of truth for reference
documents, the CLI reference management commands are now redundant:
```
mem ref add --project homelab --corpus kubectl ~/workplace/homelab/knowledge/cheatsheets
mem ref add --dry-run ... # chunk plan only, zero model calls (M3.6.1)
mem ref list --project homelab # corpus, docs, chunks, last ingest, drift
mem ref sync --corpus kubectl # re-walk, replace changed docs, report
mem ref rm --corpus kubectl # tombstone every doc in the corpus
- **Document modification:** Users edit files in Obsidian UI (web/desktop)
- **Corpus management:** Obsidian vault (filesystem) is the authoritative store
- **Synchronization:** M3.6.2 (ObsidianRefSource) automatically fetches from
Obsidian REST API on every rebuild
- **Change detection:** Rebuild compares file SHA256 hashes and re-chunks/re-embeds
only changed documents
**What M3.6.3 did (historical):**
## Historical Functionality (M3.6.3 commands, now retired)
```bash
# These commands are no longer needed:
mem ref add --project homelab --corpus kubectl ~/path/to/docs
mem ref list --project homelab
mem ref sync --corpus kubectl
mem ref rm --corpus kubectl
```
**Identity is `(source_uri, doc_sha)`.** Same URI and same sha is a no-op: no
embed call, no write, exit 0 with "unchanged". Same URI and different sha is a
*replace*: tombstone the old chunks in the log, write the new ones. A URI that
has vanished from the tree on a `sync` is a tombstone with no successor.
**Why they're no longer needed:**
**Tombstone, do not delete.** The log is append-only and authoritative. A
`{"kind":"reference_tombstone","sha256":"…","reason":"replaced"}` record is what
removal means; the projector drops the row and the note on replay. Deleting rows
from Postgres directly makes the index un-rebuildable, which is the one thing the
whole design refuses.
1. **Users don't add corpora via CLI.** They use Obsidian to create/edit `.md`
files in the vault. The vault is the source of truth, fetched via REST API.
**Embedding is the expensive step, so skip it precisely.** A corpus of 400 chunks
where one document changed should issue embeddings for that document's chunks
only. Chunk-level sha comparison, not document-level re-embed.
2. **Sync happens automatically on rebuild.** `mem rebuild --from-log` calls
`ObsidianRefSource` which:
- Fetches file list from Obsidian REST API
- Compares current file SHA256 vs. previous (already computed in log)
- Re-chunks changed files only (deterministic embedding)
- Re-indexes in Postgres + OpenSearch
- Tombstones deleted files (via log records)
**`list` reports drift.** For each corpus, re-stat the tree and compare doc shas
without writing anything: `3 docs changed, 1 removed, 12 unchanged`. Drift that
is only discoverable by running `sync` means nobody runs `sync`.
3. **List/drift detection can be queried.** `mem query --level=R` shows which
reference chunks are indexed. Drift is detected on rebuild by comparing shas.
## Steps
4. **No external file paths to manage.** Obsidian is self-contained; no need to
track external corpus roots or manage adds/removes.
1. `mem ref add [--project P] --corpus C [--dry-run] <path>` — walk via
`DocCorpusSource`, embed new chunks, write `Reference` records.
2. Persist corpus registration (name, root path, project, last ingest) in the
log as a `reference_corpus` record so `list` needs no side file.
3. Implement chunk-level diff: existing shas for the corpus vs freshly walked
shas → `{new, changed, unchanged, gone}`.
4. `mem ref sync` — apply the diff, embedding only `new` and `changed`, emitting
tombstones for `gone`.
5. `mem ref list` — table per corpus with counts plus a dry drift check.
6. `mem ref rm` — tombstone every live chunk in the corpus; leave the log intact.
7. Exit codes: 0 on success including no-op, non-zero on unresolvable corpus or
unreadable root.
## Workflow Instead (Current Architecture)
## Acceptance
**User modifies reference documents:**
```
1. User opens Obsidian UI (http://obsidian.poimen.local)
2. User creates/edits `.md` files in vault
3. Files saved to persistent storage (K8s PVC)
```
- `add` twice on an unchanged tree issues zero embedding calls the second time.
- Editing one file and running `sync` re-embeds that file's chunks only.
- Deleting a file and running `sync` tombstones its chunks, and it stops
appearing in query results.
- `rm` removes the corpus from results while leaving every record in the log.
- `list` reports drift without mutating anything.
**Memory system syncs automatically:**
```
1. Admin runs: mem rebuild --from-log
2. ObsidianRefSource (M3.6.2) fetches file list from Obsidian REST API
3. For each file: compare SHA256 vs. previous (from log)
- Changed: re-chunk via M3.6.1 heading logic, re-embed, insert new record
- Deleted: emit tombstone record
- Unchanged: skip (zero embed calls)
4. Postgres + OpenSearch indexes updated deterministically
5. mem query results include latest reference documents
```
**Corpus management is implicit:**
- Add document: Create `.md` file in Obsidian
- Remove document: Delete `.md` file in Obsidian
- Update document: Edit `.md` file in Obsidian
- Sync: `mem rebuild --from-log`
## Acceptance (Shifted to M3.6.2)
The properties M3.6.3 enforced are now M3.6.2's responsibility:
- ✅ Changed file: re-embeds only that file's chunks (M3.6.2 chunk-level diff)
- ✅ Deleted file: emits tombstone record (M3.6.2 rebuild logic)
- ✅ Deleted file removed from queries: tombstone prevents indexing
- ✅ Unchanged files: zero embed calls on rebuild (M3.6.2 SHA comparison)
- ✅ Rebuild parity: same file contents → same chunk shas → same indexes
(M3.6.2 assertion a14_rebuild_byte_identical)
## Verify
**Harness:** fixture tree copied to a temp dir so it can be mutated, a counting
embedder that records how many texts it was asked to embed, seeded database.
**Verification moved to M3.6.2 (ObsidianRefSource)** which now covers:
- `a7_rebuild_byte_identical` — zero embeds on second rebuild of unchanged corpus
- `a8_changed_doc_reembeds_only_itself` — edit one file, rebuild re-embeds only that file
- `a9_deleted_doc_emits_tombstone` — delete file, rebuild emits tombstone, query excludes it
**Integration test**`tests/it_mem_ref.rs`:
1. `a1_add_then_add_is_noop` — run `add` twice; assert the embedder call count is
zero on the second run and the row count is unchanged.
2. `a2_changed_doc_reembeds_only_itself` — edit one file of three, `sync`; assert
**No separate M3.6.3 tests needed.** All functionality tested as part of M3.6.2
integration harness.
embed count equals that file's chunk count, not the corpus total.
3. `a3_replace_tombstones_predecessor` — after a change, assert the old chunk
sha has a tombstone record and no live row.
@@ -104,7 +128,14 @@ embedder that records how many texts it was asked to embed, seeded database.
- Running the churn test without a final rebuild. Assertion 6 is what proves the
tombstone replay logic exists rather than being implied.
## Traps
## Historical Traps (No Longer Applicable)
These were issues with the CLI approach; no longer relevant since Obsidian is
the source of truth:
~~Reusing `run_id` semantics...~~ → Now: `ref-obsidian-<timestamp>` in M3.6.2
~~Tracking external corpus roots...~~ → Now: Obsidian REST API is the root
~~Managing adds/removes via state files...~~ → Now: Obsidian vault is the state
- Registering the corpus before the walk succeeds. A failed `add` that leaves a
registered-but-empty corpus makes the next `sync` report every document as new.