Files
poimen-memory/tasks/M3.6.2-level-r-storage.md
T

182 lines
8.3 KiB
Markdown
Raw Normal View History

# M3.6.2 — Level R: Obsidian reference indexing and rebuild parity
| Field | Value |
|---|---|
| Phase | M3.6 — Reference corpora |
| Size | M — 13 days |
| Status | ⬜ Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | M3.6.6 |
| Depends | M3.6.1, M1.6, M2.3, M2.4, M2.5, M2.6, Obsidian service (M2.5) |
## Goal
Fetch reference documents from the live Obsidian vault (REST API), chunk them
via M3.6.1's heading-boundary logic, land them in the log as reference records,
index them in Postgres, and prove rebuild parity (drop + rebuild from log =
byte-identical index state).
## Facts (inlined — no spec read needed)
**Obsidian is the source of truth for reference documents.** The Obsidian REST API
(deployed in M2.5, accessible at `http://obsidian-server.poimen.svc.cluster.local:8080`)
provides live access to vault files via:
```bash
GET /api/vault/listFiles # List all .md files
GET /api/vault/readFile?path=kubectl.md # Read file contents
GET /api/vault/listFolders # Explore structure
```
**Reference record format (in JSONL log):**
```jsonl
{"kind":"reference","level":"R","source":"obsidian://poimen-vault/kubectl.md",
"doc_sha":"ab12…","heading_path":"kubectl.md > Common Issues > CrashLoopBackOff",
"sha256":"cd34…","t":7,"run_id":"ref-obsidian-2026-08-21T10:02:11Z","text":"…"}
```
**No vault projection.** Unlike M3.6.1 (DocCorpusSource which reads from local filesystem),
M3.6.2 reads directly from Obsidian REST API and indexes into Postgres. The Obsidian
vault remains the single source of truth; `mem rebuild --from-log` re-fetches from
Obsidian to regenerate indexes.
`level = 'R'`, `query_id = NULL` (R answers no standing question), `source` holds
the Obsidian URI (`obsidian://vault-name/file.md`). `doc_sha` is whole-document hash;
`sha256` is chunk hash (primary identity).
The embedding goes to `memory_vector(kind='text')`, not to a column on the node.
R gets **no symptom projection** — M3.7.8 generates those for L1 and L2 only.
**R writes no edges.** Provenance is the Obsidian URI, not internal edges. Enforced
in `mem verify` (M3.6.4).
**Rebuild parity is the whole point.** `mem rebuild --from-log` must:
1. Drop R rows from Postgres
2. Drop R vectors from OpenSearch (M8.2)
3. Re-fetch documents from Obsidian REST API
4. Re-chunk via M3.6.1 heading logic
5. Re-index into Postgres + OpenSearch
6. Result must be byte-identical (same shas, same vector embeddings)
## Steps
1. **ObsidianRefSource** in `mem-ingest`: Implement `ChunkedSource` that:
- Calls Obsidian REST API (`/api/vault/listFiles`)
- Filters for `.md` files in allowed paths (e.g., `docs/`, `reference/`)
- Fetches each file via `/api/vault/readFile?path=...`
- Chunks via M3.6.1's heading boundary logic
- Yields `Record { kind: Reference, ... }`
2. Add `Reference` variant to log record enum in `mem-core` (already exists from
M2.3, just needs the Obsidian source)
3. `mem-store`: Insert R nodes with `kind='text'` vector; assert no edge insert
names an R sha as parent (repository boundary check).
4. Dual-write in M8.2: When R records are indexed, write both to Postgres and
OpenSearch (L2 already does this; extend for R).
5. **Rebuild: extend M2.6 harness** for `mem rebuild --from-log`:
- When replaying R records from log, re-fetch original files from Obsidian
- Re-chunk via M3.6.1 logic
- Regenerate embeddings (deterministic, so shas match)
- Assert byte-identical state vs. original indexing
6. Integration: Update `mem query` to include R results from hybrid search (M8)
## Acceptance
- `ObsidianRefSource` fetches all `.md` files from Obsidian REST API.
- Reference records (R) with `source=obsidian://...` land in the log.
- R rows persist in Postgres with `query_id IS NULL` and `source` set to Obsidian URI.
- Each R chunk gets a `kind='text'` embedding in OpenSearch.
- The M2.3 constraint accepts `R` and rejects `L3`.
- `mem rebuild --from-log`: drop Postgres R rows + OpenSearch R vectors, re-fetch
from Obsidian, re-chunk, re-embed, re-index; result is byte-identical (same shas,
same embedding vectors).
- No hidden inputs: Obsidian REST API is the only external dependency for R.
## Verify
**Harness:** Mock Obsidian REST API (deterministic file list + content). M2.6 rebuild
harness extended to cover R records. Deterministic embedder so shas are stable.
**Integration test**`tests/it_obsidian_ref_source.rs` + `tests/it_level_r_storage.rs`:
### ObsidianRefSource tests:
1. `a1_fetches_md_files` — Mock API returns 3 `.md` files; source emits 3 documents.
2. `a2_chunks_by_heading` — One source document with 5 headings yields 5 chunks
(uses M3.6.1 heading boundary logic).
3. `a3_generates_doc_sha` — Document SHA (SHA256 of whole file content) is
consistent across fetches.
4. `a4_generates_chunk_shas` — Each chunk sha is deterministic: SHA256(tool +
heading_path + text).
5. `a5_sets_obsidian_uri` — Every R record has `source=obsidian://vault-name/path.md`.
6. `a6_skips_non_md` — API returns `.json` and `.txt` files; source ignores them.
### Level R storage tests:
7. `a7_record_roundtrip` — Serialize R record to JSONL, deserialize; assert
field-for-field equality.
8. `a8_r_inserts_to_postgres` — Insert `level='R'` with `kind='text'` vector;
assert Postgres row exists with correct `source` URI.
9. `a9_r_inserts_to_opensearch` — Same chunk indexed in OpenSearch (M8.2
dual-write); assert M8 vector store has the chunk.
10. `a10_no_symptom_projection` — R chunks do not acquire `kind='symptom'`
vectors (only L1/L2 get those from M3.7.8).
11. `a11_query_id_null` — Every R row has `query_id IS NULL`; insertion with
non-NULL query_id is rejected by M2.3's CHECK.
12. `a12_no_edges_from_r` — After ingesting 10 R chunks, `SELECT count(*)
FROM memory_edge WHERE parent_sha IN (SELECT sha256 FROM memory_node WHERE
level='R')` equals 0.
### Rebuild parity tests:
13. `a13_rebuild_drops_r_rows` — Ingest R, assert N rows exist. Call `mem
rebuild --from-log`, assert same N rows exist (re-fetched, re-embedded,
re-indexed).
14. `a14_rebuild_byte_identical` — Snapshot Postgres R rows + OpenSearch R
vectors before rebuild. Drop both. Run `mem rebuild --from-log` against
a log containing R records. Assert Postgres rows and OpenSearch vectors
are byte-identical (same row order, same field values, same vector embeddings).
15. `a15_rebuild_idempotent` — Run rebuild twice; second run changes nothing.
Snapshot comparison between first and second rebuild result is identical.
**Command:**
```bash
cargo test -p mem-ingest obsidian_ref_source
cargo test -p mem-store level_r
cargo test -p mem-cli rebuild # includes R records in fixture log
```
**False pass:**
- Asserting rebuild parity on a log with no R records. Assertion 14 requires
mixed-level log (L0/L1/L2 *and* R) or it trivially passes.
- Mocking Obsidian API with hardcoded responses. File hashes must match real
Obsidian vault file SHA256 (use actual Obsidian instance or hash fixtures).
- Checking edge count is zero before ingesting. Assertion 12 must run against
populated corpus or it asserts empty table is empty.
- Byte-identical comparison with normalization (sorting, ignoring order).
Byte-identical means bytes; rebuild parity is broken if row order changes
between runs.
## Traps
- Reusing `run_id` semantics from the gated loop. R has no run in the recurrence
sense; use a synthetic `ref-obsidian-<timestamp>` and do not let it collide
with a real ingest run in queries that group by `run_id`.
- Caching Obsidian API responses across rebuild runs. Rebuild must re-fetch from
Obsidian REST API every time (no cache) to ensure file contents and shas are
always in sync with live vault.
- Assuming Obsidian file list is sorted. `listFiles` order is arbitrary; source
must sort filenames before chunking to ensure deterministic shas across runs.
- Embedding each R chunk independently. M3.6.1 (DocCorpusSource) uses the same
`m` parameter (batching) and deterministic embedder model to ensure chunk shas
are stable; M3.6.2 must use identical setup.
- Dropping the check constraint instead of widening it. Assertion 11 exists
because `DROP CONSTRAINT` alone passes every other assertion in this file.
---
Background: [DESIGN.md](../DESIGN.md) — reference corpora, storage schemas