Files
poimen-memory/tasks/M2.3-schema-and-migrations.md
T

111 lines
4.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# M2.3 — Schema + sqlx migrations
| Field | Value |
|---|---|
| Phase | M2 — Projections |
| Size | M — 13 days |
| Status | ⬜ Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | M2.2 |
## Goal
The tables the provenance graph lives in, with the constraints that make a
malformed graph impossible rather than merely unlikely.
## Facts (inlined — no spec read needed)
```sql
CREATE TABLE memory_node (
id BIGSERIAL PRIMARY KEY,
level TEXT NOT NULL CHECK (level IN ('L0','L1','L2')),
project TEXT NOT NULL,
query_id TEXT, -- NULL at L2
run_id TEXT NOT NULL,
t INT NOT NULL,
source TEXT, -- set at L0
text TEXT NOT NULL,
sha256 TEXT NOT NULL UNIQUE, -- content identity, from M0.2
embedding vector(768) NOT NULL,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE TABLE memory_edge (
child_sha TEXT NOT NULL REFERENCES memory_node(sha256) ON DELETE CASCADE,
parent_sha TEXT NOT NULL REFERENCES memory_node(sha256) ON DELETE CASCADE,
PRIMARY KEY (child_sha, parent_sha)
);
CREATE INDEX ON memory_node USING hnsw (embedding vector_cosine_ops);
CREATE INDEX ON memory_node (project, level);
```
`sha256 UNIQUE` is what makes rebuild idempotent — re-inserting identical content
is a conflict to ignore, not a duplicate row. It is also why the hash must exclude
run ids and timestamps (M0.2).
`ON DELETE CASCADE` on both edge columns: dropping a node should not leave
dangling edges. Rebuild drops everything anyway, but a partial cleanup should not
be able to corrupt the graph.
`query_id` is NULL at L2 by design — L2 spans queries. Enforce it:
`CHECK ((level = 'L2') = (query_id IS NULL))`.
Cosine distance, not L2: these are normalised text embeddings and cosine is what
the model was trained for. `vector_cosine_ops` must match the operator the query
uses (`<=>`), or the index is silently ignored and every query is a seq scan.
## Steps
1. `migrations/0001_init.sql` with the above, plus the `query_id`/level CHECK.
2. `CREATE EXTENSION IF NOT EXISTS vector;` first — the declarative extension
(M2.2) should have run, and this makes local dev work too.
3. Wire `sqlx::migrate!()` and run at startup.
4. `sqlx prepare` for offline compile-time query checking in CI.
5. Add a `schema_version` sanity query the repo layer asserts on connect.
6. Document that changing the embedding model is a migration, because the column
width is part of the schema.
## Acceptance
- Migrations apply to a clean database and are idempotent.
- Inserting a duplicate `sha256` conflicts rather than duplicating.
- An `L2` row with a non-null `query_id` is rejected by the CHECK.
- The HNSW index is used by a cosine-distance query.
## Verify
**Harness:** a disposable database — `sqlx::test` or testcontainers with the same
image tag as production, `ghcr.io/cloudnative-pg/postgresql:16.2`.
**Integration test**`tests/it_schema.rs`:
1. `a1_migrate_clean` — apply to an empty database, assert both tables exist.
2. `a2_migrate_idempotent` — apply twice, assert no error.
3. `a3_sha_unique` — insert the same sha twice, assert a unique violation.
4. `a4_level_check``level='L3'` rejected; `level='L2'` with a `query_id`
rejected; `level='L1'` without one rejected.
5. `a5_edge_fk` — an edge referencing a missing sha is rejected.
6. `a6_cascade` — delete a node, assert its edges are gone.
7. `a7_hnsw_is_used``EXPLAIN` a `ORDER BY embedding <=> $1 LIMIT 10` query and
assert the plan contains `Index Scan` on the HNSW index, not `Seq Scan`.
**Command:** `cargo test -p mem-store schema`
**False pass:**
- Testing the schema against SQLite or plain Postgres without pgvector. It will
accept `vector(768)` as an unknown type in some configurations and every
vector assertion becomes meaningless. Use the production image.
- Omitting assertion 7. An index created with the wrong opclass exists, reports
healthy, and is never used — queries just get slower as the table grows, which
reads as a scaling problem rather than a wrong index.
## Traps
- `vector_l2_ops` with a `<=>` query, or the reverse. The index is silently
ignored. This is the single most common pgvector mistake.
- Making `query_id` NOT NULL because L1 always has one. L2 then cannot be stored,
and the workaround is a sentinel string that pollutes every group-by.
---
Background: [DESIGN.md](../DESIGN.md) — pgvector schema