7.5 KiB
M2.3 — Schema + sqlx migrations
| Field | Value |
|---|---|
| Phase | M2 — Projections |
| Size | M — 1–3 days |
| Status | ⬜ Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | M2.2 |
Goal
The tables the provenance graph lives in, with the constraints that make a malformed graph impossible rather than merely unlikely.
Facts (inlined — no spec read needed)
CREATE TABLE memory_node (
id BIGSERIAL PRIMARY KEY,
level TEXT NOT NULL CHECK (level IN ('L0','L1','L2','R')),
project TEXT NOT NULL,
query_id TEXT, -- NULL at L2 and R
run_id TEXT NOT NULL,
t INT NOT NULL,
source TEXT, -- set at L0; source URI at R
text TEXT NOT NULL,
sha256 TEXT NOT NULL UNIQUE, -- content identity, from M0.2
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE TABLE memory_edge (
child_sha TEXT NOT NULL REFERENCES memory_node(sha256) ON DELETE CASCADE,
parent_sha TEXT NOT NULL REFERENCES memory_node(sha256) ON DELETE CASCADE,
PRIMARY KEY (child_sha, parent_sha)
);
CREATE INDEX ON memory_node (project, level);
-- Vectors live outside the node: one node carries several, and a symptom
-- projection (M3.7.8) is what makes an answer findable from an error message.
CREATE TABLE memory_vector (
node_sha TEXT NOT NULL REFERENCES memory_node(sha256) ON DELETE CASCADE,
kind TEXT NOT NULL CHECK (kind IN ('text','symptom')),
embedding vector(768) NOT NULL,
PRIMARY KEY (node_sha, kind)
);
-- Partial index per kind. One index over mixed kinds forces post-filtering,
-- which starves recall exactly the way M3.6.5 describes.
CREATE INDEX ON memory_vector USING hnsw (embedding vector_cosine_ops) WHERE kind = 'text';
CREATE INDEX ON memory_vector USING hnsw (embedding vector_cosine_ops) WHERE kind = 'symptom';
-- Exact-match tier. Failures repeat verbatim; prose does not.
CREATE TABLE failure_signature (
sig_sha TEXT PRIMARY KEY, -- hash of the NORMALISED signature (M3.7.7)
node_sha TEXT NOT NULL REFERENCES memory_node(sha256) ON DELETE CASCADE,
tool TEXT NOT NULL, -- 'github-actions' | 'kubectl' | 'npm'
raw TEXT NOT NULL, -- pre-normalisation, for display
seen_count INT NOT NULL DEFAULT 1,
last_seen TIMESTAMPTZ NOT NULL
);
CREATE INDEX ON failure_signature (tool);
-- A lesson about Kong is actively harmful now that Kong is retired.
CREATE TABLE memory_supersede (
old_sha TEXT NOT NULL REFERENCES memory_node(sha256) ON DELETE CASCADE,
new_sha TEXT NOT NULL REFERENCES memory_node(sha256) ON DELETE CASCADE,
reason TEXT,
PRIMARY KEY (old_sha, new_sha)
);
embedding is deliberately not a column on memory_node. A node needs more
than one vector: the memory text as written, and a generated symptom projection
describing the errors it would explain. An L1 reads like an answer and a query
reads like a stack trace, and cosine between those two registers is mediocre —
the second vector is what closes that gap. One column cannot hold both, and
bolting on embedding_2 later is worse than a junction table now.
seen_count and last_seen are mutable and that does not break the authority
rule. Occurrences are append-only records in the JSONL log; these two fields are
a fold over them, recomputed by mem rebuild --from-log like every other
projected value.
sha256 UNIQUE is what makes rebuild idempotent — re-inserting identical content
is a conflict to ignore, not a duplicate row. It is also why the hash must exclude
run ids and timestamps (M0.2).
ON DELETE CASCADE on both edge columns: dropping a node should not leave
dangling edges. Rebuild drops everything anyway, but a partial cleanup should not
be able to corrupt the graph.
query_id is NULL at L2 and at R — L2 spans queries, R answers none. Enforce
it: CHECK ((level IN ('L2','R')) = (query_id IS NULL)).
The 'R' level ships here rather than arriving as an ALTER from M3.6. Nothing
is built yet, so widening a constraint that has never existed wrong is free, and
a migration that exists only because an earlier migration was knowingly
incomplete is debt taken on for no reason.
Cosine distance, not L2: these are normalised text embeddings and cosine is what
the model was trained for. vector_cosine_ops must match the operator the query
uses (<=>), or the index is silently ignored and every query is a seq scan.
Steps
migrations/0001_init.sqlwith the above, plus thequery_id/level CHECK.CREATE EXTENSION IF NOT EXISTS vector;first — the declarative extension (M2.2) should have run, and this makes local dev work too.- Wire
sqlx::migrate!()and run at startup. sqlx preparefor offline compile-time query checking in CI.- Add a
schema_versionsanity query the repo layer asserts on connect. - Document that changing the embedding model is a migration, because the column width is part of the schema.
Acceptance
- Migrations apply to a clean database and are idempotent.
- Inserting a duplicate
sha256conflicts rather than duplicating. - An
L2orRrow with a non-nullquery_idis rejected by the CHECK. - Both partial HNSW indexes are used by a kind-filtered cosine query.
- A node can carry a
textand asymptomvector simultaneously.
Verify
Harness: a disposable database — sqlx::test or testcontainers with the same
image tag as production, ghcr.io/cloudnative-pg/postgresql:16.2.
Integration test — tests/it_schema.rs:
a1_migrate_clean— apply to an empty database, assert both tables exist.a2_migrate_idempotent— apply twice, assert no error.a3_sha_unique— insert the same sha twice, assert a unique violation.a4_level_check—level='L3'rejected;level='L2'andlevel='R'with aquery_idrejected;level='L1'without one rejected.a5_edge_fk— an edge referencing a missing sha is rejected.a6_cascade— delete a node, assert its edges, vectors and signatures are gone.a7_hnsw_is_used—EXPLAINaWHERE kind='text' ORDER BY embedding <=> $1 LIMIT 10query; assert the plan contains anIndex Scanon the partial index, notSeq Scanand not a filter applied above a full-index scan.a8_symptom_index_separate— same forkind='symptom'; assert the plan names the other index, proving both were created and are distinguishable.a9_two_vectors_per_node— insert both kinds for one node; assert both persist and the composite primary key rejects a third of the same kind.a10_signature_unique— inserting the samesig_shatwice conflicts.a11_supersede_pair— a supersede row survives, and deleting either endpoint cascades it away.
Command: cargo test -p mem-store schema
False pass:
- Testing the schema against SQLite or plain Postgres without pgvector. It will
accept
vector(768)as an unknown type in some configurations and every vector assertion becomes meaningless. Use the production image. - Omitting assertion 7. An index created with the wrong opclass exists, reports healthy, and is never used — queries just get slower as the table grows, which reads as a scaling problem rather than a wrong index.
Traps
vector_l2_opswith a<=>query, or the reverse. The index is silently ignored. This is the single most common pgvector mistake.- Making
query_idNOT NULL because L1 always has one. L2 then cannot be stored, and the workaround is a sentinel string that pollutes every group-by.
Background: DESIGN.md — pgvector schema