Files
poimen-memory/tasks/M3.6.2-level-r-storage.md
T
Story Crater Bot 3e867f7cce
Build and Push / Test (push) Failing after 1m54s
Build and Push / Build and push image (push) Skipped
chore: retire M3.6.3 (mem ref CLI), update M3.6.2 to use Obsidian REST API
CHANGES:
- M3.6.3: marked  RETIRED (Obsidian UI replaces CLI corpus management)
- M3.6.2: updated to fetch from Obsidian REST API instead of filesystem
  - ObsidianRefSource: calls /api/vault/listFiles, /api/vault/readFile
  - Users manage corpus in Obsidian UI (not via CLI)
  - Rebuild auto-syncs by re-fetching and comparing file SHAs
  - No separate chunk-level diff CLI needed
- Updated INDEX.md:
  - M3.6.x: 6 tasks → 5 tasks (removed M3.6.3)
  - Progress: 1 , 0 🟡, 5  → 1 , 0 🟡, 4 
  - Total: 71 tasks → 70 tasks
  - Noted M3.6.3 retirement in board description

RATIONALE:
- Obsidian is single source of truth (REST API)
- Users already use Obsidian UI for vault management
- No need for parallel CLI when vault is the interface
- M3.6.2 handles sync via deterministic SHA comparison
- Reduces feature bloat, cleaner architecture
2026-08-28 08:18:14 -07:00

8.3 KiB
Raw Blame History

M3.6.2 — Level R: Obsidian reference indexing and rebuild parity

Field Value
Phase M3.6 — Reference corpora
Size M — 13 days
Status Not started
Flags
Spec inlined below
Blocks M3.6.6
Depends M3.6.1, M1.6, M2.3, M2.4, M2.5, M2.6, Obsidian service (M2.5)

Goal

Fetch reference documents from the live Obsidian vault (REST API), chunk them via M3.6.1's heading-boundary logic, land them in the log as reference records, index them in Postgres, and prove rebuild parity (drop + rebuild from log = byte-identical index state).

Facts (inlined — no spec read needed)

Obsidian is the source of truth for reference documents. The Obsidian REST API (deployed in M2.5, accessible at http://obsidian-server.poimen.svc.cluster.local:8080) provides live access to vault files via:

GET /api/vault/listFiles                          # List all .md files
GET /api/vault/readFile?path=kubectl.md           # Read file contents
GET /api/vault/listFolders                        # Explore structure

Reference record format (in JSONL log):

{"kind":"reference","level":"R","source":"obsidian://poimen-vault/kubectl.md",
 "doc_sha":"ab12…","heading_path":"kubectl.md > Common Issues > CrashLoopBackOff",
 "sha256":"cd34…","t":7,"run_id":"ref-obsidian-2026-08-21T10:02:11Z","text":"…"}

No vault projection. Unlike M3.6.1 (DocCorpusSource which reads from local filesystem), M3.6.2 reads directly from Obsidian REST API and indexes into Postgres. The Obsidian vault remains the single source of truth; mem rebuild --from-log re-fetches from Obsidian to regenerate indexes.

level = 'R', query_id = NULL (R answers no standing question), source holds the Obsidian URI (obsidian://vault-name/file.md). doc_sha is whole-document hash; sha256 is chunk hash (primary identity).

The embedding goes to memory_vector(kind='text'), not to a column on the node. R gets no symptom projection — M3.7.8 generates those for L1 and L2 only.

R writes no edges. Provenance is the Obsidian URI, not internal edges. Enforced in mem verify (M3.6.4).

Rebuild parity is the whole point. mem rebuild --from-log must:

  1. Drop R rows from Postgres
  2. Drop R vectors from OpenSearch (M8.2)
  3. Re-fetch documents from Obsidian REST API
  4. Re-chunk via M3.6.1 heading logic
  5. Re-index into Postgres + OpenSearch
  6. Result must be byte-identical (same shas, same vector embeddings)

Steps

  1. ObsidianRefSource in mem-ingest: Implement ChunkedSource that:

    • Calls Obsidian REST API (/api/vault/listFiles)
    • Filters for .md files in allowed paths (e.g., docs/, reference/)
    • Fetches each file via /api/vault/readFile?path=...
    • Chunks via M3.6.1's heading boundary logic
    • Yields Record { kind: Reference, ... }
  2. Add Reference variant to log record enum in mem-core (already exists from M2.3, just needs the Obsidian source)

  3. mem-store: Insert R nodes with kind='text' vector; assert no edge insert names an R sha as parent (repository boundary check).

  4. Dual-write in M8.2: When R records are indexed, write both to Postgres and OpenSearch (L2 already does this; extend for R).

  5. Rebuild: extend M2.6 harness for mem rebuild --from-log:

    • When replaying R records from log, re-fetch original files from Obsidian
    • Re-chunk via M3.6.1 logic
    • Regenerate embeddings (deterministic, so shas match)
    • Assert byte-identical state vs. original indexing
  6. Integration: Update mem query to include R results from hybrid search (M8)

Acceptance

  • ObsidianRefSource fetches all .md files from Obsidian REST API.
  • Reference records (R) with source=obsidian://... land in the log.
  • R rows persist in Postgres with query_id IS NULL and source set to Obsidian URI.
  • Each R chunk gets a kind='text' embedding in OpenSearch.
  • The M2.3 constraint accepts R and rejects L3.
  • mem rebuild --from-log: drop Postgres R rows + OpenSearch R vectors, re-fetch from Obsidian, re-chunk, re-embed, re-index; result is byte-identical (same shas, same embedding vectors).
  • No hidden inputs: Obsidian REST API is the only external dependency for R.

Verify

Harness: Mock Obsidian REST API (deterministic file list + content). M2.6 rebuild harness extended to cover R records. Deterministic embedder so shas are stable.

Integration testtests/it_obsidian_ref_source.rs + tests/it_level_r_storage.rs:

ObsidianRefSource tests:

  1. a1_fetches_md_files — Mock API returns 3 .md files; source emits 3 documents.
  2. a2_chunks_by_heading — One source document with 5 headings yields 5 chunks (uses M3.6.1 heading boundary logic).
  3. a3_generates_doc_sha — Document SHA (SHA256 of whole file content) is consistent across fetches.
  4. a4_generates_chunk_shas — Each chunk sha is deterministic: SHA256(tool + heading_path + text).
  5. a5_sets_obsidian_uri — Every R record has source=obsidian://vault-name/path.md.
  6. a6_skips_non_md — API returns .json and .txt files; source ignores them.

Level R storage tests:

  1. a7_record_roundtrip — Serialize R record to JSONL, deserialize; assert field-for-field equality.
  2. a8_r_inserts_to_postgres — Insert level='R' with kind='text' vector; assert Postgres row exists with correct source URI.
  3. a9_r_inserts_to_opensearch — Same chunk indexed in OpenSearch (M8.2 dual-write); assert M8 vector store has the chunk.
  4. a10_no_symptom_projection — R chunks do not acquire kind='symptom' vectors (only L1/L2 get those from M3.7.8).
  5. a11_query_id_null — Every R row has query_id IS NULL; insertion with non-NULL query_id is rejected by M2.3's CHECK.
  6. a12_no_edges_from_r — After ingesting 10 R chunks, SELECT count(*) FROM memory_edge WHERE parent_sha IN (SELECT sha256 FROM memory_node WHERE level='R') equals 0.

Rebuild parity tests:

  1. a13_rebuild_drops_r_rows — Ingest R, assert N rows exist. Call mem rebuild --from-log, assert same N rows exist (re-fetched, re-embedded, re-indexed).
  2. a14_rebuild_byte_identical — Snapshot Postgres R rows + OpenSearch R vectors before rebuild. Drop both. Run mem rebuild --from-log against a log containing R records. Assert Postgres rows and OpenSearch vectors are byte-identical (same row order, same field values, same vector embeddings).
  3. a15_rebuild_idempotent — Run rebuild twice; second run changes nothing. Snapshot comparison between first and second rebuild result is identical.

Command:

cargo test -p mem-ingest obsidian_ref_source
cargo test -p mem-store level_r
cargo test -p mem-cli rebuild  # includes R records in fixture log

False pass:

  • Asserting rebuild parity on a log with no R records. Assertion 14 requires mixed-level log (L0/L1/L2 and R) or it trivially passes.
  • Mocking Obsidian API with hardcoded responses. File hashes must match real Obsidian vault file SHA256 (use actual Obsidian instance or hash fixtures).
  • Checking edge count is zero before ingesting. Assertion 12 must run against populated corpus or it asserts empty table is empty.
  • Byte-identical comparison with normalization (sorting, ignoring order). Byte-identical means bytes; rebuild parity is broken if row order changes between runs.

Traps

  • Reusing run_id semantics from the gated loop. R has no run in the recurrence sense; use a synthetic ref-obsidian-<timestamp> and do not let it collide with a real ingest run in queries that group by run_id.
  • Caching Obsidian API responses across rebuild runs. Rebuild must re-fetch from Obsidian REST API every time (no cache) to ensure file contents and shas are always in sync with live vault.
  • Assuming Obsidian file list is sorted. listFiles order is arbitrary; source must sort filenames before chunking to ensure deterministic shas across runs.
  • Embedding each R chunk independently. M3.6.1 (DocCorpusSource) uses the same m parameter (batching) and deterministic embedder model to ensure chunk shas are stable; M3.6.2 must use identical setup.
  • Dropping the check constraint instead of widening it. Assertion 11 exists because DROP CONSTRAINT alone passes every other assertion in this file.

Background: DESIGN.md — reference corpora, storage schemas