# M3.6.2 — Level R: Obsidian reference indexing and rebuild parity | Field | Value | |---|---| | Phase | M3.6 — Reference corpora | | Size | M — 1–3 days | | Status | ⬜ Not started | | Flags | — | | Spec | inlined below | | Blocks | M3.6.6 | | Depends | M3.6.1, M1.6, M2.3, M2.4, M2.5, M2.6, Obsidian service (M2.5) | ## Goal Fetch reference documents from the live Obsidian vault (REST API), chunk them via M3.6.1's heading-boundary logic, land them in the log as reference records, index them in Postgres, and prove rebuild parity (drop + rebuild from log = byte-identical index state). ## Facts (inlined — no spec read needed) **Obsidian is the source of truth for reference documents.** The Obsidian REST API (deployed in M2.5, accessible at `http://obsidian-server.poimen.svc.cluster.local:8080`) provides live access to vault files via: ```bash GET /api/vault/listFiles # List all .md files GET /api/vault/readFile?path=kubectl.md # Read file contents GET /api/vault/listFolders # Explore structure ``` **Reference record format (in JSONL log):** ```jsonl {"kind":"reference","level":"R","source":"obsidian://poimen-vault/kubectl.md", "doc_sha":"ab12…","heading_path":"kubectl.md > Common Issues > CrashLoopBackOff", "sha256":"cd34…","t":7,"run_id":"ref-obsidian-2026-08-21T10:02:11Z","text":"…"} ``` **No vault projection.** Unlike M3.6.1 (DocCorpusSource which reads from local filesystem), M3.6.2 reads directly from Obsidian REST API and indexes into Postgres. The Obsidian vault remains the single source of truth; `mem rebuild --from-log` re-fetches from Obsidian to regenerate indexes. `level = 'R'`, `query_id = NULL` (R answers no standing question), `source` holds the Obsidian URI (`obsidian://vault-name/file.md`). `doc_sha` is whole-document hash; `sha256` is chunk hash (primary identity). The embedding goes to `memory_vector(kind='text')`, not to a column on the node. R gets **no symptom projection** — M3.7.8 generates those for L1 and L2 only. **R writes no edges.** Provenance is the Obsidian URI, not internal edges. Enforced in `mem verify` (M3.6.4). **Rebuild parity is the whole point.** `mem rebuild --from-log` must: 1. Drop R rows from Postgres 2. Drop R vectors from OpenSearch (M8.2) 3. Re-fetch documents from Obsidian REST API 4. Re-chunk via M3.6.1 heading logic 5. Re-index into Postgres + OpenSearch 6. Result must be byte-identical (same shas, same vector embeddings) ## Steps 1. **ObsidianRefSource** in `mem-ingest`: Implement `ChunkedSource` that: - Calls Obsidian REST API (`/api/vault/listFiles`) - Filters for `.md` files in allowed paths (e.g., `docs/`, `reference/`) - Fetches each file via `/api/vault/readFile?path=...` - Chunks via M3.6.1's heading boundary logic - Yields `Record { kind: Reference, ... }` 2. Add `Reference` variant to log record enum in `mem-core` (already exists from M2.3, just needs the Obsidian source) 3. `mem-store`: Insert R nodes with `kind='text'` vector; assert no edge insert names an R sha as parent (repository boundary check). 4. Dual-write in M8.2: When R records are indexed, write both to Postgres and OpenSearch (L2 already does this; extend for R). 5. **Rebuild: extend M2.6 harness** for `mem rebuild --from-log`: - When replaying R records from log, re-fetch original files from Obsidian - Re-chunk via M3.6.1 logic - Regenerate embeddings (deterministic, so shas match) - Assert byte-identical state vs. original indexing 6. Integration: Update `mem query` to include R results from hybrid search (M8) ## Acceptance - `ObsidianRefSource` fetches all `.md` files from Obsidian REST API. - Reference records (R) with `source=obsidian://...` land in the log. - R rows persist in Postgres with `query_id IS NULL` and `source` set to Obsidian URI. - Each R chunk gets a `kind='text'` embedding in OpenSearch. - The M2.3 constraint accepts `R` and rejects `L3`. - `mem rebuild --from-log`: drop Postgres R rows + OpenSearch R vectors, re-fetch from Obsidian, re-chunk, re-embed, re-index; result is byte-identical (same shas, same embedding vectors). - No hidden inputs: Obsidian REST API is the only external dependency for R. ## Verify **Harness:** Mock Obsidian REST API (deterministic file list + content). M2.6 rebuild harness extended to cover R records. Deterministic embedder so shas are stable. **Integration test** — `tests/it_obsidian_ref_source.rs` + `tests/it_level_r_storage.rs`: ### ObsidianRefSource tests: 1. `a1_fetches_md_files` — Mock API returns 3 `.md` files; source emits 3 documents. 2. `a2_chunks_by_heading` — One source document with 5 headings yields 5 chunks (uses M3.6.1 heading boundary logic). 3. `a3_generates_doc_sha` — Document SHA (SHA256 of whole file content) is consistent across fetches. 4. `a4_generates_chunk_shas` — Each chunk sha is deterministic: SHA256(tool + heading_path + text). 5. `a5_sets_obsidian_uri` — Every R record has `source=obsidian://vault-name/path.md`. 6. `a6_skips_non_md` — API returns `.json` and `.txt` files; source ignores them. ### Level R storage tests: 7. `a7_record_roundtrip` — Serialize R record to JSONL, deserialize; assert field-for-field equality. 8. `a8_r_inserts_to_postgres` — Insert `level='R'` with `kind='text'` vector; assert Postgres row exists with correct `source` URI. 9. `a9_r_inserts_to_opensearch` — Same chunk indexed in OpenSearch (M8.2 dual-write); assert M8 vector store has the chunk. 10. `a10_no_symptom_projection` — R chunks do not acquire `kind='symptom'` vectors (only L1/L2 get those from M3.7.8). 11. `a11_query_id_null` — Every R row has `query_id IS NULL`; insertion with non-NULL query_id is rejected by M2.3's CHECK. 12. `a12_no_edges_from_r` — After ingesting 10 R chunks, `SELECT count(*) FROM memory_edge WHERE parent_sha IN (SELECT sha256 FROM memory_node WHERE level='R')` equals 0. ### Rebuild parity tests: 13. `a13_rebuild_drops_r_rows` — Ingest R, assert N rows exist. Call `mem rebuild --from-log`, assert same N rows exist (re-fetched, re-embedded, re-indexed). 14. `a14_rebuild_byte_identical` — Snapshot Postgres R rows + OpenSearch R vectors before rebuild. Drop both. Run `mem rebuild --from-log` against a log containing R records. Assert Postgres rows and OpenSearch vectors are byte-identical (same row order, same field values, same vector embeddings). 15. `a15_rebuild_idempotent` — Run rebuild twice; second run changes nothing. Snapshot comparison between first and second rebuild result is identical. **Command:** ```bash cargo test -p mem-ingest obsidian_ref_source cargo test -p mem-store level_r cargo test -p mem-cli rebuild # includes R records in fixture log ``` **False pass:** - Asserting rebuild parity on a log with no R records. Assertion 14 requires mixed-level log (L0/L1/L2 *and* R) or it trivially passes. - Mocking Obsidian API with hardcoded responses. File hashes must match real Obsidian vault file SHA256 (use actual Obsidian instance or hash fixtures). - Checking edge count is zero before ingesting. Assertion 12 must run against populated corpus or it asserts empty table is empty. - Byte-identical comparison with normalization (sorting, ignoring order). Byte-identical means bytes; rebuild parity is broken if row order changes between runs. ## Traps - Reusing `run_id` semantics from the gated loop. R has no run in the recurrence sense; use a synthetic `ref-obsidian-` and do not let it collide with a real ingest run in queries that group by `run_id`. - Caching Obsidian API responses across rebuild runs. Rebuild must re-fetch from Obsidian REST API every time (no cache) to ensure file contents and shas are always in sync with live vault. - Assuming Obsidian file list is sorted. `listFiles` order is arbitrary; source must sort filenames before chunking to ensure deterministic shas across runs. - Embedding each R chunk independently. M3.6.1 (DocCorpusSource) uses the same `m` parameter (batching) and deterministic embedder model to ensure chunk shas are stable; M3.6.2 must use identical setup. - Dropping the check constraint instead of widening it. Assertion 11 exists because `DROP CONSTRAINT` alone passes every other assertion in this file. --- Background: [DESIGN.md](../DESIGN.md) — reference corpora, storage schemas