CHANGES: - M3.6.3: marked ❌ RETIRED (Obsidian UI replaces CLI corpus management) - M3.6.2: updated to fetch from Obsidian REST API instead of filesystem - ObsidianRefSource: calls /api/vault/listFiles, /api/vault/readFile - Users manage corpus in Obsidian UI (not via CLI) - Rebuild auto-syncs by re-fetching and comparing file SHAs - No separate chunk-level diff CLI needed - Updated INDEX.md: - M3.6.x: 6 tasks → 5 tasks (removed M3.6.3) - Progress: 1 ✅, 0 🟡, 5 ⬜ → 1 ✅, 0 🟡, 4 ⬜ - Total: 71 tasks → 70 tasks - Noted M3.6.3 retirement in board description RATIONALE: - Obsidian is single source of truth (REST API) - Users already use Obsidian UI for vault management - No need for parallel CLI when vault is the interface - M3.6.2 handles sync via deterministic SHA comparison - Reduces feature bloat, cleaner architecture
8.3 KiB
M3.6.2 — Level R: Obsidian reference indexing and rebuild parity
| Field | Value |
|---|---|
| Phase | M3.6 — Reference corpora |
| Size | M — 1–3 days |
| Status | ⬜ Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | M3.6.6 |
| Depends | M3.6.1, M1.6, M2.3, M2.4, M2.5, M2.6, Obsidian service (M2.5) |
Goal
Fetch reference documents from the live Obsidian vault (REST API), chunk them via M3.6.1's heading-boundary logic, land them in the log as reference records, index them in Postgres, and prove rebuild parity (drop + rebuild from log = byte-identical index state).
Facts (inlined — no spec read needed)
Obsidian is the source of truth for reference documents. The Obsidian REST API
(deployed in M2.5, accessible at http://obsidian-server.poimen.svc.cluster.local:8080)
provides live access to vault files via:
GET /api/vault/listFiles # List all .md files
GET /api/vault/readFile?path=kubectl.md # Read file contents
GET /api/vault/listFolders # Explore structure
Reference record format (in JSONL log):
{"kind":"reference","level":"R","source":"obsidian://poimen-vault/kubectl.md",
"doc_sha":"ab12…","heading_path":"kubectl.md > Common Issues > CrashLoopBackOff",
"sha256":"cd34…","t":7,"run_id":"ref-obsidian-2026-08-21T10:02:11Z","text":"…"}
No vault projection. Unlike M3.6.1 (DocCorpusSource which reads from local filesystem),
M3.6.2 reads directly from Obsidian REST API and indexes into Postgres. The Obsidian
vault remains the single source of truth; mem rebuild --from-log re-fetches from
Obsidian to regenerate indexes.
level = 'R', query_id = NULL (R answers no standing question), source holds
the Obsidian URI (obsidian://vault-name/file.md). doc_sha is whole-document hash;
sha256 is chunk hash (primary identity).
The embedding goes to memory_vector(kind='text'), not to a column on the node.
R gets no symptom projection — M3.7.8 generates those for L1 and L2 only.
R writes no edges. Provenance is the Obsidian URI, not internal edges. Enforced
in mem verify (M3.6.4).
Rebuild parity is the whole point. mem rebuild --from-log must:
- Drop R rows from Postgres
- Drop R vectors from OpenSearch (M8.2)
- Re-fetch documents from Obsidian REST API
- Re-chunk via M3.6.1 heading logic
- Re-index into Postgres + OpenSearch
- Result must be byte-identical (same shas, same vector embeddings)
Steps
-
ObsidianRefSource in
mem-ingest: ImplementChunkedSourcethat:- Calls Obsidian REST API (
/api/vault/listFiles) - Filters for
.mdfiles in allowed paths (e.g.,docs/,reference/) - Fetches each file via
/api/vault/readFile?path=... - Chunks via M3.6.1's heading boundary logic
- Yields
Record { kind: Reference, ... }
- Calls Obsidian REST API (
-
Add
Referencevariant to log record enum inmem-core(already exists from M2.3, just needs the Obsidian source) -
mem-store: Insert R nodes withkind='text'vector; assert no edge insert names an R sha as parent (repository boundary check). -
Dual-write in M8.2: When R records are indexed, write both to Postgres and OpenSearch (L2 already does this; extend for R).
-
Rebuild: extend M2.6 harness for
mem rebuild --from-log:- When replaying R records from log, re-fetch original files from Obsidian
- Re-chunk via M3.6.1 logic
- Regenerate embeddings (deterministic, so shas match)
- Assert byte-identical state vs. original indexing
-
Integration: Update
mem queryto include R results from hybrid search (M8)
Acceptance
ObsidianRefSourcefetches all.mdfiles from Obsidian REST API.- Reference records (R) with
source=obsidian://...land in the log. - R rows persist in Postgres with
query_id IS NULLandsourceset to Obsidian URI. - Each R chunk gets a
kind='text'embedding in OpenSearch. - The M2.3 constraint accepts
Rand rejectsL3. mem rebuild --from-log: drop Postgres R rows + OpenSearch R vectors, re-fetch from Obsidian, re-chunk, re-embed, re-index; result is byte-identical (same shas, same embedding vectors).- No hidden inputs: Obsidian REST API is the only external dependency for R.
Verify
Harness: Mock Obsidian REST API (deterministic file list + content). M2.6 rebuild harness extended to cover R records. Deterministic embedder so shas are stable.
Integration test — tests/it_obsidian_ref_source.rs + tests/it_level_r_storage.rs:
ObsidianRefSource tests:
a1_fetches_md_files— Mock API returns 3.mdfiles; source emits 3 documents.a2_chunks_by_heading— One source document with 5 headings yields 5 chunks (uses M3.6.1 heading boundary logic).a3_generates_doc_sha— Document SHA (SHA256 of whole file content) is consistent across fetches.a4_generates_chunk_shas— Each chunk sha is deterministic: SHA256(tool + heading_path + text).a5_sets_obsidian_uri— Every R record hassource=obsidian://vault-name/path.md.a6_skips_non_md— API returns.jsonand.txtfiles; source ignores them.
Level R storage tests:
a7_record_roundtrip— Serialize R record to JSONL, deserialize; assert field-for-field equality.a8_r_inserts_to_postgres— Insertlevel='R'withkind='text'vector; assert Postgres row exists with correctsourceURI.a9_r_inserts_to_opensearch— Same chunk indexed in OpenSearch (M8.2 dual-write); assert M8 vector store has the chunk.a10_no_symptom_projection— R chunks do not acquirekind='symptom'vectors (only L1/L2 get those from M3.7.8).a11_query_id_null— Every R row hasquery_id IS NULL; insertion with non-NULL query_id is rejected by M2.3's CHECK.a12_no_edges_from_r— After ingesting 10 R chunks,SELECT count(*) FROM memory_edge WHERE parent_sha IN (SELECT sha256 FROM memory_node WHERE level='R')equals 0.
Rebuild parity tests:
a13_rebuild_drops_r_rows— Ingest R, assert N rows exist. Callmem rebuild --from-log, assert same N rows exist (re-fetched, re-embedded, re-indexed).a14_rebuild_byte_identical— Snapshot Postgres R rows + OpenSearch R vectors before rebuild. Drop both. Runmem rebuild --from-logagainst a log containing R records. Assert Postgres rows and OpenSearch vectors are byte-identical (same row order, same field values, same vector embeddings).a15_rebuild_idempotent— Run rebuild twice; second run changes nothing. Snapshot comparison between first and second rebuild result is identical.
Command:
cargo test -p mem-ingest obsidian_ref_source
cargo test -p mem-store level_r
cargo test -p mem-cli rebuild # includes R records in fixture log
False pass:
- Asserting rebuild parity on a log with no R records. Assertion 14 requires mixed-level log (L0/L1/L2 and R) or it trivially passes.
- Mocking Obsidian API with hardcoded responses. File hashes must match real Obsidian vault file SHA256 (use actual Obsidian instance or hash fixtures).
- Checking edge count is zero before ingesting. Assertion 12 must run against populated corpus or it asserts empty table is empty.
- Byte-identical comparison with normalization (sorting, ignoring order). Byte-identical means bytes; rebuild parity is broken if row order changes between runs.
Traps
- Reusing
run_idsemantics from the gated loop. R has no run in the recurrence sense; use a syntheticref-obsidian-<timestamp>and do not let it collide with a real ingest run in queries that group byrun_id. - Caching Obsidian API responses across rebuild runs. Rebuild must re-fetch from Obsidian REST API every time (no cache) to ensure file contents and shas are always in sync with live vault.
- Assuming Obsidian file list is sorted.
listFilesorder is arbitrary; source must sort filenames before chunking to ensure deterministic shas across runs. - Embedding each R chunk independently. M3.6.1 (DocCorpusSource) uses the same
mparameter (batching) and deterministic embedder model to ensure chunk shas are stable; M3.6.2 must use identical setup. - Dropping the check constraint instead of widening it. Assertion 11 exists
because
DROP CONSTRAINTalone passes every other assertion in this file.
Background: DESIGN.md — reference corpora, storage schemas