CHANGES: - M3.6.3: marked ❌ RETIRED (Obsidian UI replaces CLI corpus management) - M3.6.2: updated to fetch from Obsidian REST API instead of filesystem - ObsidianRefSource: calls /api/vault/listFiles, /api/vault/readFile - Users manage corpus in Obsidian UI (not via CLI) - Rebuild auto-syncs by re-fetching and comparing file SHAs - No separate chunk-level diff CLI needed - Updated INDEX.md: - M3.6.x: 6 tasks → 5 tasks (removed M3.6.3) - Progress: 1 ✅, 0 🟡, 5 ⬜ → 1 ✅, 0 🟡, 4 ⬜ - Total: 71 tasks → 70 tasks - Noted M3.6.3 retirement in board description RATIONALE: - Obsidian is single source of truth (REST API) - Users already use Obsidian UI for vault management - No need for parallel CLI when vault is the interface - M3.6.2 handles sync via deterministic SHA comparison - Reduces feature bloat, cleaner architecture
182 lines
8.3 KiB
Markdown
182 lines
8.3 KiB
Markdown
# M3.6.2 — Level R: Obsidian reference indexing and rebuild parity
|
||
|
||
| Field | Value |
|
||
|---|---|
|
||
| Phase | M3.6 — Reference corpora |
|
||
| Size | M — 1–3 days |
|
||
| Status | ⬜ Not started |
|
||
| Flags | — |
|
||
| Spec | inlined below |
|
||
| Blocks | M3.6.6 |
|
||
| Depends | M3.6.1, M1.6, M2.3, M2.4, M2.5, M2.6, Obsidian service (M2.5) |
|
||
|
||
## Goal
|
||
|
||
Fetch reference documents from the live Obsidian vault (REST API), chunk them
|
||
via M3.6.1's heading-boundary logic, land them in the log as reference records,
|
||
index them in Postgres, and prove rebuild parity (drop + rebuild from log =
|
||
byte-identical index state).
|
||
|
||
## Facts (inlined — no spec read needed)
|
||
|
||
**Obsidian is the source of truth for reference documents.** The Obsidian REST API
|
||
(deployed in M2.5, accessible at `http://obsidian-server.poimen.svc.cluster.local:8080`)
|
||
provides live access to vault files via:
|
||
|
||
```bash
|
||
GET /api/vault/listFiles # List all .md files
|
||
GET /api/vault/readFile?path=kubectl.md # Read file contents
|
||
GET /api/vault/listFolders # Explore structure
|
||
```
|
||
|
||
**Reference record format (in JSONL log):**
|
||
|
||
```jsonl
|
||
{"kind":"reference","level":"R","source":"obsidian://poimen-vault/kubectl.md",
|
||
"doc_sha":"ab12…","heading_path":"kubectl.md > Common Issues > CrashLoopBackOff",
|
||
"sha256":"cd34…","t":7,"run_id":"ref-obsidian-2026-08-21T10:02:11Z","text":"…"}
|
||
```
|
||
|
||
**No vault projection.** Unlike M3.6.1 (DocCorpusSource which reads from local filesystem),
|
||
M3.6.2 reads directly from Obsidian REST API and indexes into Postgres. The Obsidian
|
||
vault remains the single source of truth; `mem rebuild --from-log` re-fetches from
|
||
Obsidian to regenerate indexes.
|
||
|
||
`level = 'R'`, `query_id = NULL` (R answers no standing question), `source` holds
|
||
the Obsidian URI (`obsidian://vault-name/file.md`). `doc_sha` is whole-document hash;
|
||
`sha256` is chunk hash (primary identity).
|
||
|
||
The embedding goes to `memory_vector(kind='text')`, not to a column on the node.
|
||
R gets **no symptom projection** — M3.7.8 generates those for L1 and L2 only.
|
||
|
||
**R writes no edges.** Provenance is the Obsidian URI, not internal edges. Enforced
|
||
in `mem verify` (M3.6.4).
|
||
|
||
**Rebuild parity is the whole point.** `mem rebuild --from-log` must:
|
||
1. Drop R rows from Postgres
|
||
2. Drop R vectors from OpenSearch (M8.2)
|
||
3. Re-fetch documents from Obsidian REST API
|
||
4. Re-chunk via M3.6.1 heading logic
|
||
5. Re-index into Postgres + OpenSearch
|
||
6. Result must be byte-identical (same shas, same vector embeddings)
|
||
|
||
## Steps
|
||
|
||
1. **ObsidianRefSource** in `mem-ingest`: Implement `ChunkedSource` that:
|
||
- Calls Obsidian REST API (`/api/vault/listFiles`)
|
||
- Filters for `.md` files in allowed paths (e.g., `docs/`, `reference/`)
|
||
- Fetches each file via `/api/vault/readFile?path=...`
|
||
- Chunks via M3.6.1's heading boundary logic
|
||
- Yields `Record { kind: Reference, ... }`
|
||
|
||
2. Add `Reference` variant to log record enum in `mem-core` (already exists from
|
||
M2.3, just needs the Obsidian source)
|
||
|
||
3. `mem-store`: Insert R nodes with `kind='text'` vector; assert no edge insert
|
||
names an R sha as parent (repository boundary check).
|
||
|
||
4. Dual-write in M8.2: When R records are indexed, write both to Postgres and
|
||
OpenSearch (L2 already does this; extend for R).
|
||
|
||
5. **Rebuild: extend M2.6 harness** for `mem rebuild --from-log`:
|
||
- When replaying R records from log, re-fetch original files from Obsidian
|
||
- Re-chunk via M3.6.1 logic
|
||
- Regenerate embeddings (deterministic, so shas match)
|
||
- Assert byte-identical state vs. original indexing
|
||
|
||
6. Integration: Update `mem query` to include R results from hybrid search (M8)
|
||
|
||
## Acceptance
|
||
|
||
- `ObsidianRefSource` fetches all `.md` files from Obsidian REST API.
|
||
- Reference records (R) with `source=obsidian://...` land in the log.
|
||
- R rows persist in Postgres with `query_id IS NULL` and `source` set to Obsidian URI.
|
||
- Each R chunk gets a `kind='text'` embedding in OpenSearch.
|
||
- The M2.3 constraint accepts `R` and rejects `L3`.
|
||
- `mem rebuild --from-log`: drop Postgres R rows + OpenSearch R vectors, re-fetch
|
||
from Obsidian, re-chunk, re-embed, re-index; result is byte-identical (same shas,
|
||
same embedding vectors).
|
||
- No hidden inputs: Obsidian REST API is the only external dependency for R.
|
||
|
||
## Verify
|
||
|
||
**Harness:** Mock Obsidian REST API (deterministic file list + content). M2.6 rebuild
|
||
harness extended to cover R records. Deterministic embedder so shas are stable.
|
||
|
||
**Integration test** — `tests/it_obsidian_ref_source.rs` + `tests/it_level_r_storage.rs`:
|
||
|
||
### ObsidianRefSource tests:
|
||
1. `a1_fetches_md_files` — Mock API returns 3 `.md` files; source emits 3 documents.
|
||
2. `a2_chunks_by_heading` — One source document with 5 headings yields 5 chunks
|
||
(uses M3.6.1 heading boundary logic).
|
||
3. `a3_generates_doc_sha` — Document SHA (SHA256 of whole file content) is
|
||
consistent across fetches.
|
||
4. `a4_generates_chunk_shas` — Each chunk sha is deterministic: SHA256(tool +
|
||
heading_path + text).
|
||
5. `a5_sets_obsidian_uri` — Every R record has `source=obsidian://vault-name/path.md`.
|
||
6. `a6_skips_non_md` — API returns `.json` and `.txt` files; source ignores them.
|
||
|
||
### Level R storage tests:
|
||
7. `a7_record_roundtrip` — Serialize R record to JSONL, deserialize; assert
|
||
field-for-field equality.
|
||
8. `a8_r_inserts_to_postgres` — Insert `level='R'` with `kind='text'` vector;
|
||
assert Postgres row exists with correct `source` URI.
|
||
9. `a9_r_inserts_to_opensearch` — Same chunk indexed in OpenSearch (M8.2
|
||
dual-write); assert M8 vector store has the chunk.
|
||
10. `a10_no_symptom_projection` — R chunks do not acquire `kind='symptom'`
|
||
vectors (only L1/L2 get those from M3.7.8).
|
||
11. `a11_query_id_null` — Every R row has `query_id IS NULL`; insertion with
|
||
non-NULL query_id is rejected by M2.3's CHECK.
|
||
12. `a12_no_edges_from_r` — After ingesting 10 R chunks, `SELECT count(*)
|
||
FROM memory_edge WHERE parent_sha IN (SELECT sha256 FROM memory_node WHERE
|
||
level='R')` equals 0.
|
||
|
||
### Rebuild parity tests:
|
||
13. `a13_rebuild_drops_r_rows` — Ingest R, assert N rows exist. Call `mem
|
||
rebuild --from-log`, assert same N rows exist (re-fetched, re-embedded,
|
||
re-indexed).
|
||
14. `a14_rebuild_byte_identical` — Snapshot Postgres R rows + OpenSearch R
|
||
vectors before rebuild. Drop both. Run `mem rebuild --from-log` against
|
||
a log containing R records. Assert Postgres rows and OpenSearch vectors
|
||
are byte-identical (same row order, same field values, same vector embeddings).
|
||
15. `a15_rebuild_idempotent` — Run rebuild twice; second run changes nothing.
|
||
Snapshot comparison between first and second rebuild result is identical.
|
||
|
||
**Command:**
|
||
```bash
|
||
cargo test -p mem-ingest obsidian_ref_source
|
||
cargo test -p mem-store level_r
|
||
cargo test -p mem-cli rebuild # includes R records in fixture log
|
||
```
|
||
|
||
**False pass:**
|
||
- Asserting rebuild parity on a log with no R records. Assertion 14 requires
|
||
mixed-level log (L0/L1/L2 *and* R) or it trivially passes.
|
||
- Mocking Obsidian API with hardcoded responses. File hashes must match real
|
||
Obsidian vault file SHA256 (use actual Obsidian instance or hash fixtures).
|
||
- Checking edge count is zero before ingesting. Assertion 12 must run against
|
||
populated corpus or it asserts empty table is empty.
|
||
- Byte-identical comparison with normalization (sorting, ignoring order).
|
||
Byte-identical means bytes; rebuild parity is broken if row order changes
|
||
between runs.
|
||
|
||
## Traps
|
||
|
||
- Reusing `run_id` semantics from the gated loop. R has no run in the recurrence
|
||
sense; use a synthetic `ref-obsidian-<timestamp>` and do not let it collide
|
||
with a real ingest run in queries that group by `run_id`.
|
||
- Caching Obsidian API responses across rebuild runs. Rebuild must re-fetch from
|
||
Obsidian REST API every time (no cache) to ensure file contents and shas are
|
||
always in sync with live vault.
|
||
- Assuming Obsidian file list is sorted. `listFiles` order is arbitrary; source
|
||
must sort filenames before chunking to ensure deterministic shas across runs.
|
||
- Embedding each R chunk independently. M3.6.1 (DocCorpusSource) uses the same
|
||
`m` parameter (batching) and deterministic embedder model to ensure chunk shas
|
||
are stable; M3.6.2 must use identical setup.
|
||
- Dropping the check constraint instead of widening it. Assertion 11 exists
|
||
because `DROP CONSTRAINT` alone passes every other assertion in this file.
|
||
|
||
---
|
||
|
||
Background: [DESIGN.md](../DESIGN.md) — reference corpora, storage schemas
|