From 24a03fd9eb550e39333d9717b60959fab7e4bf33 Mon Sep 17 00:00:00 2001 From: Story Crater Bot <19826264+Riotpiaole@users.noreply.github.com> Date: Thu, 27 Aug 2026 21:18:58 -0700 Subject: [PATCH] archive: Delete completed task files (M2.1,2.2,2.4,2.5,2.6,M8.1) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Completed tasks moved to git history for archive: - M2.1 Embeddings client (768-dim batching) - M2.2 CNPG memory-db manifest - M2.4 pgvector repository - M2.5 Obsidian projector - M2.6 Rebuild from log orchestrator - M8.1 OpenSearch cluster deployment Remaining in /tasks/: 48 files (active/in-progress/not-started) โœ… Completed: 49/73 (index.md source of truth) ๐ŸŸก In progress: 2 (M3.5.9, M3.7.5) โฌœ Not started: 22 --- tasks/M2.1-embeddings-client.md | 86 ------------------- tasks/M2.2-memory-db-manifest.md | 102 ---------------------- tasks/M2.4-pgvector-repo.md | 129 ---------------------------- tasks/M2.5-obsidian-projector.md | 113 ------------------------ tasks/M2.6-rebuild-from-log.md | 98 --------------------- tasks/M8.1-opensearch-deployment.md | 111 ------------------------ 6 files changed, 639 deletions(-) delete mode 100644 tasks/M2.1-embeddings-client.md delete mode 100644 tasks/M2.2-memory-db-manifest.md delete mode 100644 tasks/M2.4-pgvector-repo.md delete mode 100644 tasks/M2.5-obsidian-projector.md delete mode 100644 tasks/M2.6-rebuild-from-log.md delete mode 100644 tasks/M8.1-opensearch-deployment.md diff --git a/tasks/M2.1-embeddings-client.md b/tasks/M2.1-embeddings-client.md deleted file mode 100644 index f98461d..0000000 --- a/tasks/M2.1-embeddings-client.md +++ /dev/null @@ -1,86 +0,0 @@ -# M2.1 โ€” Embeddings client - -| Field | Value | -|---|---| -| Phase | M2 โ€” Projections | -| Size | S โ€” under 1 day | -| Status | โœ… Done | -| Flags | โ€” | -| Spec | inlined below | -| Blocks | M1.1 | - -## Goal - -Turn text into 768-dim vectors, batched, against the gateway's TEI endpoint. - -## Facts (inlined โ€” no spec read needed) - -``` -POST /v1/embeddings -{"model":"nomic-ai/nomic-embed-text-v2-moe","input":["..."]} --> {"object":"list","data":[{"embedding":[...768 floats...]}],"usage":{...}} -``` - -**768 dimensions**, probed and confirmed. It is the `vector(768)` in the schema -and in the HNSW index; a model swap is a schema migration, not a config change. - -**Batch limit is 32.** Verified: 1200 inputs returned -`{"message":"batch size 1200 > maximum allowed batch size 32","code":413}`. -Chunk the input list accordingly. - -This route currently has **no auth** โ€” no `konghq.com/plugins` annotation, so -`model-key-auth` never attaches. Send the `apikey` header anyway: the route -should be fixed, and a client that only works while auth is broken breaks when -it is fixed. - -Bodies here are large (many texts ร— long strings) and the Kong buffer is 16m, so -batching also keeps requests well inside it. - -## Steps - -1. `EmbeddingsClient::embed(texts: &[String]) -> Result>>` in `mem-llm`. -2. Split into batches of โ‰ค32, preserving input order in the output. -3. Assert every returned vector is exactly 768 long; a mismatch is an error - naming the model, not a silent pad or truncate. -4. Reuse M1.1's client config: `apikey` header, retry on 5xx only, generous - timeout. -5. `embed_one` convenience wrapper. - -## Acceptance - -- 100 texts return 100 vectors in input order. -- Every vector is 768-dim. -- A dimension mismatch errors loudly. - -## Verify - -**Harness:** `wiremock` offline, one `#[ignore]` live test. - -**Integration test** โ€” `tests/it_embeddings.rs`: -1. `a1_batches_at_32` โ€” 100 inputs produce exactly 4 requests. -2. `a2_order_preserved` โ€” mock returns identifiable vectors; assert output order - matches input order across batch boundaries. -3. `a3_dimension_asserted` โ€” mock returns a 512-dim vector; assert an error - naming the model. -4. `a4_apikey_sent` โ€” assert the header is present even though the route does not - require it. -5. `a5_live_dims` โ€” `#[ignore]`; real gateway, assert 768. - -**Command:** `cargo test -p mem-llm embeddings` (add `-- --ignored` for a5) - -**False pass:** -- Testing with โ‰ค32 inputs. The batching path never runs and order-across-batches - โ€” the thing most likely to be wrong โ€” is never exercised. -- Trusting the response order within a batch without asserting it. Assertion 2 - must use distinguishable vectors, not a length check. - -## Traps - -- Hardcoding 768 in three places. Put it in one constant that the schema - migration also references, so a model change is one edit and one migration. -- Assuming no auth is needed because it currently works without a key. That route - is missing its plugin annotation, which is a bug scheduled to be fixed. - ---- - -Background: [DESIGN.md](../DESIGN.md) โ€” Verified facts, pgvector diff --git a/tasks/M2.2-memory-db-manifest.md b/tasks/M2.2-memory-db-manifest.md deleted file mode 100644 index 9fe2b7a..0000000 --- a/tasks/M2.2-memory-db-manifest.md +++ /dev/null @@ -1,102 +0,0 @@ -# M2.2 โ€” CNPG `memory-db` + pgvector - -| Field | Value | -|---|---| -| Phase | M2 โ€” Projections | -| Size | M โ€” 1โ€“3 days | -| Status | โœ… Done | -| Flags | homelab | -| Spec | inlined below | -| Blocks | โ€” | - -## Goal - -A Postgres with pgvector, provisioned the way everything else in the cluster is: -through git, with no manual `psql`. - -## Facts (inlined โ€” no spec read needed) - -**pgvector needs no custom image.** Verified on the running cluster: - -``` -$ psql -tAc "select name,default_version,installed_version - from pg_available_extensions where name='vector'" -vector|0.7.0| -``` - -on the stock `ghcr.io/cloudnative-pg/postgresql:16.2`. Available, not yet -installed โ€” `CREATE EXTENSION` is all that is missing. - -**The operator is CNPG 1.30.0**, which supports declarative extensions on the -`Database` CRD (`kubectl explain database.spec.extensions` resolves). So the -extension is git-managed too โ€” no manual step, consistent with the GitOps rule -that infrastructure changes flow through version control. - -Follow `k8s/infra/databases/temporal-db.yaml` exactly: 3 instances, `imageName` -pinned, `enableSuperuserAccess: false`, `storageClass: longhorn-cnpg`, -`enablePodMonitor: true`, control-plane tolerations, `podAntiAffinityType: -preferred`. - -Storage: 10Gi matches the existing clusters. At 768 dims ร— 4 bytes, a vector is -~3 KB; tens of thousands of nodes is well under a gigabyte, so 10Gi is generous -and consistent rather than tight. - -## Steps - -1. `k8s/infra/databases/memory-db.yaml` โ€” `Cluster` + `Database` with - `extensions: [{name: vector, ensure: present}]`. -2. Namespace `memory`, created by the ArgoCD app that owns it. -3. Add to the owning kustomization's explicit resource list โ€” an unlisted file is - silently dropped with no error and no drift shown. -4. Commit, push, let ArgoCD sync. **No `kubectl apply`.** -5. Verify the extension installed and the app user can create tables. -6. Record the connection string convention in the repo README; the password comes - from the CNPG-generated secret, never committed. - -## Acceptance - -- `Cluster` reaches `Cluster in healthy state` with 3 instances. -- `select extversion from pg_extension where extname='vector'` returns a version. -- ArgoCD shows the app `Synced/Healthy`. -- No manual `psql` was run to get there. - -## Verify - -**Harness:** `kubectl` and `psql` read-only checks after sync. - -**Integration test** โ€” `verify/m2.2.sh`, output diffed against `expected/m2.2.txt`: -1. `a1_cluster_healthy` โ€” `kubectl get cluster -n memory memory-db` reports 3/3 - ready. -2. `a2_extension_installed` โ€” `select extname, extversion from pg_extension where - extname='vector'` returns one row. -3. `a3_declarative_not_manual` โ€” `kubectl get database -n memory memory-db-vector - -o jsonpath='{.spec.extensions}'` shows the declaration, proving it came from - git. -4. `a4_argocd_synced` โ€” the owning app is `Synced/Healthy`. -5. `a5_app_user_can_ddl` โ€” as `app`, `CREATE TABLE t(v vector(768)); DROP TABLE t;` - succeeds. -6. `a6_hnsw_available` โ€” `CREATE INDEX ... USING hnsw` on that temp table - succeeds, proving 0.7.0 has the index type the schema needs. - -**Command:** `bash verify/m2.2.sh | diff - expected/m2.2.txt` - -**False pass:** -- Checking `pg_available_extensions` instead of `pg_extension`. Available means - the files are on disk; installed means `CREATE EXTENSION` ran. The whole task - is the second one. -- Verifying after a manual `CREATE EXTENSION`. It passes and proves nothing about - the declarative path, which is the actual deliverable. Assertion 3 is the guard. - -## Traps - -- Forgetting the kustomization resource list. The file sits in git, ArgoCD reports - Synced, and the objects never exist โ€” silent, and the failure surfaces later as - a connection error. -- Adding `prune: true` semantics without thinking about operator-created children. - CNPG creates Services, Secrets and PVCs owned by the Cluster; if ArgoCD's - tracking label propagates to them, prune fights the operator. The - `llm-serving` app already had to set `prune: false` for exactly this reason. - ---- - -Background: [DESIGN.md](../DESIGN.md) โ€” pgvector ยท `k8s/infra/databases/temporal-db.yaml` diff --git a/tasks/M2.4-pgvector-repo.md b/tasks/M2.4-pgvector-repo.md deleted file mode 100644 index 2f2f96e..0000000 --- a/tasks/M2.4-pgvector-repo.md +++ /dev/null @@ -1,129 +0,0 @@ -# M2.4 โ€” pgvector repository - -| Field | Value | -|---|---| -| Phase | M2 โ€” Projections | -| Size | M โ€” 1โ€“3 days | -| Status | โœ… Done | -| Flags | โ€” | -| Spec | inlined below | -| Blocks | M2.3, M2.1 | - -## Goal - -Write the projection into Postgres idempotently, so rebuild is safe to run at any -time and produces the same rows. - -## Facts (inlined โ€” no spec read needed) - -```rust -async fn upsert_node(&self, node: &MemoryNode) -> Result<()>; -async fn upsert_vector(&self, sha: &Sha256Hash, kind: VectorKind, embedding: &[f32]) -> Result<()>; -async fn insert_edges(&self, child: &Sha256Hash, parents: &[Sha256Hash]) -> Result<()>; -async fn search(&self, q: &[f32], kind: VectorKind, levels: &[Level], - project: Scope, k: usize) -> Result>; -async fn lookup_signature(&self, sig_sha: &str) -> Result>; -async fn parents_of(&self, sha: &Sha256Hash) -> Result>; -async fn clear_project(&self, project: &ProjectId) -> Result<()>; -``` - -`upsert_node` is `ON CONFLICT (sha256) DO NOTHING`. Content identity means an -identical node is the same node; re-running rebuild must not duplicate or churn -rows. Same for edges on the composite key. - -`search` orders by `embedding <=> $1` โ€” cosine distance, matching the -`vector_cosine_ops` index. Any other operator silently drops to a seq scan. - -**`kind` must be a literal predicate in the SQL, not a bind parameter, and not a -filter applied to results.** The indexes are partial (`WHERE kind = 'text'`), and -the planner only uses a partial index when the query's predicate provably matches -it. A `WHERE kind = $2` defeats that and silently degrades to a scan over every -vector of both kinds โ€” the same failure mode as the wrong opclass, and just as -invisible. - -**`Scope` is not a `ProjectId`.** Tool-failure lookups federate across projects -because an `ERESOLVE` lesson is not project-specific, while ordinary standing-query -memories stay scoped. `Scope::Project(id)` filters; `Scope::AllProjects` does not -and lets project relevance act as a rank boost later instead of a hard filter. - -`lookup_signature` is the exact-match tier: a primary-key hit on -`failure_signature`, no vector involved. It is the cheapest and highest-precision -answer the store can give, so it belongs in the repository rather than being -assembled from a `search` call by a caller who does not know it exists. - -Edges are inserted **after** both endpoints exist, or the foreign key rejects -them. Rebuild therefore has two passes: all nodes, then all edges. This is not an -optimisation; a single-pass insert fails on the first forward reference. - -Embeddings are generated in batches of โ‰ค32 (M2.1) and are the expensive part of -rebuild โ€” batch across nodes, not per node. - -## Steps - -1. `PgRepo::connect(url)` with a pool; run migrations on connect. -2. Implement the five methods above. -3. `upsert_many(nodes)` batching embedding calls at 32 and inserting with a - multi-row statement. Batch across *both* vector kinds โ€” a node with a symptom - projection contributes two texts to the same batch, not two batches. -4. Three-pass write: nodes, then vectors and signatures, then edges. Vectors and - signatures carry foreign keys to nodes, so they cannot precede them, and edges - still need both endpoints present. -5. Separate query builders per `kind` so the literal predicate is guaranteed at - compile time rather than by convention. -5. `clear_project` deletes nodes for one project; edges cascade. -6. Return `ScoredNode { node, distance, matched_kind }` โ€” keep the raw distance, - do not convert to a similarity score here. The reranker (M3.2) wants the - ordering, and a lossy conversion hides ties. `matched_kind` tells the caller - whether the hit came from the memory text or its symptom projection, which is - the difference between "this is about your topic" and "this explains your - error". -7. Exclude superseded nodes by default: `LEFT JOIN memory_supersede` on - `old_sha`, filter where the join is null. An `include_superseded` flag exists - for audit, off everywhere else. - -## Acceptance - -- Upserting the same node twice leaves one row. -- Edges referencing not-yet-inserted parents fail; two-pass write succeeds. -- `search` returns nearest-first and respects the level filter. -- `clear_project` removes only that project. - -## Verify - -**Harness:** disposable Postgres with the production image; a deterministic fake -embedder (hash โ†’ fixed vector) so vector assertions are exact. - -**Integration test** โ€” `tests/it_pg_repo.rs`: -1. `a1_upsert_idempotent` โ€” upsert twice, assert `count(*) == 1`. -2. `a2_two_pass_required` โ€” single-pass insert with a forward edge reference - fails; two-pass succeeds. Proves the ordering constraint is real. -3. `a3_search_orders_by_distance` โ€” insert three known vectors, assert returned - order matches hand-computed cosine distance. -4. `a4_level_filter` โ€” L0/L1/L2 present; search with `levels=[L1]` returns only - L1. -5. `a5_project_isolation` โ€” two projects with identical text; search one, assert - no cross-project results. -6. `a6_clear_project_scoped` โ€” clear one, assert the other is intact and no - orphan edges remain. -7. `a7_batching` โ€” upsert 100 nodes, assert the embedder saw exactly 4 calls. -8. `a8_parents_of` โ€” walk a two-level graph, assert the returned parents match. - -**Command:** `cargo test -p mem-store pg_repo` - -**False pass:** -- Using a random embedder. Assertion 3 becomes untestable and is usually deleted, - which removes the only check that the distance operator matches the index. -- Testing `search` with one project in the database. Assertion 5 is the only one - that catches a missing `WHERE project = $1`, and that bug leaks another - project's memory into every answer. - -## Traps - -- Converting distance to similarity in the repo. It loses precision, and the - reranker downstream wants candidates in order rather than scores. -- Per-node embedding calls. 412 chunks becomes 412 HTTP round trips where 13 - would do, and rebuild goes from seconds to minutes. - ---- - -Background: [DESIGN.md](../DESIGN.md) โ€” pgvector, retrieval diff --git a/tasks/M2.5-obsidian-projector.md b/tasks/M2.5-obsidian-projector.md deleted file mode 100644 index 629221f..0000000 --- a/tasks/M2.5-obsidian-projector.md +++ /dev/null @@ -1,113 +0,0 @@ -# M2.5 โ€” Obsidian projector - -| Field | Value | -|---|---| -| Phase | M2 โ€” Projections | -| Size | M โ€” 1โ€“3 days | -| Status | โœ… Done | -| Flags | โ€” | -| Spec | inlined below | -| Blocks | M1.6 | - -## Goal - -Render the log as a vault a human reads, with the tier graph as the link graph. - -## Facts (inlined โ€” no spec read needed) - -``` -vault// - index.md L2 synthesis, links every L1 note - .md L1, one per standing query - evidence/-.md L0, only with --emit-evidence-notes -``` - -```markdown ---- -project: poimen -level: L1 -query_id: infra-root-causes -updated: 2026-08-17 -chunks_seen: 412 -chunks_used: 17 -run_id: 01HXYZ... ---- -# Infra root causes โ€” poimen - - - -## Provenance -- [[pi-2026-07-21-019f857d]] chunk 66 โ€” Kong body buffer -``` - -**Deterministic output is the requirement, not a nicety.** M2.8 asserts that -rebuilding produces a byte-identical vault. That means: stable key order in -frontmatter, no timestamp of *generation* (only `updated` derived from the log), -sorted provenance lists, and `\n` line endings. - -`updated` comes from the run's timestamp in the log โ€” not `now()`. A generation -timestamp makes every rebuild a diff and destroys the gate. - -L0 notes default off: 17 per query is fine, but it grows unbounded across -projects and queries. Citations inline give the same provenance without the file -count. - -Wikilinks are `[[]]`. The link target may not exist as a file -when evidence notes are off โ€” that is fine and normal in Obsidian, and it still -shows in the graph view as an unresolved node. - -## Steps - -1. `ObsidianProjector::project(log_dir, vault_dir, opts)` in `mem-store`. -2. Read the log; take the final `memory` record per query for L1, and the L2 - record for `index.md`. -3. Frontmatter with a fixed key order; `updated` from the log. -4. Provenance section from `parents`, sorted by source then `t`. -5. `--emit-evidence-notes` writes L0 notes; default off. -6. Write with `\n`, no trailing whitespace, exactly one trailing newline. -7. A note whose L1 memory is empty is still written, with a body saying no - evidence was found โ€” an absent file is indistinguishable from a failed run. - -## Acceptance - -- Two projections of the same log produce byte-identical files. -- Frontmatter key order is stable. -- `updated` reflects the run, not the projection. -- Every L1 note links to its L2 index and vice versa. - -## Verify - -**Harness:** a committed log fixture and a committed expected vault tree. - -**Integration test** โ€” `tests/it_projector.rs`: -1. `a1_byte_identical_twice` โ€” project into two temp dirs, assert every file's - bytes are equal. This is M2.8's core property, tested early. -2. `a2_no_generation_timestamp` โ€” project, sleep 1s, project again, assert equal. - Catches `now()` leaking into output. -3. `a3_frontmatter_key_order` โ€” assert the exact key sequence. -4. `a4_golden_tree` โ€” diff the whole output against `expected/vault/`, empty diff. -5. `a5_empty_memory_still_writes` โ€” a log with zero updates produces a note - saying so. -6. `a6_links_bidirectional` โ€” every L1 note appears in `index.md` and links back. -7. `a7_evidence_notes_flag` โ€” off by default; on, produces one note per L0 node. -8. `a8_line_endings` โ€” no `\r`, exactly one trailing `\n`. - -**Command:** `cargo test -p mem-store projector` - -**False pass:** -- Comparing files by parsed content rather than bytes. Key reordering and - whitespace churn both pass, and both fail M2.8 later, where the cause is much - harder to find. -- Testing with a single query. Assertion 6 needs at least two L1 notes to catch a - link built from the wrong id. - -## Traps - -- `updated: {now}`. The most natural thing to write, and it makes every rebuild - dirty, which trains everyone to ignore the diff that M2.8 depends on. -- Serializing frontmatter from a `HashMap`. Iteration order is unspecified and - the output churns between runs on the same input. - ---- - -Background: [DESIGN.md](../DESIGN.md) โ€” Obsidian vault diff --git a/tasks/M2.6-rebuild-from-log.md b/tasks/M2.6-rebuild-from-log.md deleted file mode 100644 index d35cedb..0000000 --- a/tasks/M2.6-rebuild-from-log.md +++ /dev/null @@ -1,98 +0,0 @@ -# M2.6 โ€” `mem rebuild --from-log` - -| Field | Value | -|---|---| -| Phase | M2 โ€” Projections | -| Size | M โ€” 1โ€“3 days | -| Status | โœ… Done | -| Flags | โ€” | -| Spec | inlined below | -| Blocks | M2.4, M2.5 | - -## Goal - -Drop both projections and rebuild them from the log alone โ€” the command that -makes "the log is authoritative" a testable claim instead of a slogan. - -## Facts (inlined โ€” no spec read needed) - -``` -mem rebuild --from-log --project poimen # both projections -mem rebuild --from-log --project poimen --vault-only -mem rebuild --from-log --project poimen --db-only -``` - -The claim: **anything not reconstructible from the log has a hidden input, and -that is a bug.** Rebuild is the executable form of that claim. If it needs the -existing vault or database to produce correct output, something is being carried -across that is not in the record. - -Rebuild does **no model calls except embeddings**. Gate decisions, memory text and -provenance are all in the log already; re-running the controller would produce -different text and defeat the purpose. - -Order matters: clear โ†’ insert all nodes โ†’ insert all edges โ†’ project vault. Edges -before nodes violates the foreign key (M2.4). - -Embeddings are the expensive part. Cache by `sha256` so a rebuild after a vault -template change does not re-embed unchanged nodes. - -## Steps - -1. `mem rebuild --from-log --project P`. -2. Read every log file for the project, in run-id order. -3. `clear_project`, then two-pass node/edge insert, batching embeddings. -4. Project the vault (M2.5), overwriting. -5. Embedding cache keyed by sha, on disk under `.cache/`, so it survives runs. -6. Report counts: nodes by level, edges, embeddings computed vs cached. -7. Refuse to run if any log file is incomplete (no `run_end`) unless `--allow-partial` - โ€” rebuilding from a half-run silently produces a half-memory. - -## Acceptance - -- Rebuild from an empty database and empty vault produces the full state. -- Rebuild twice produces identical database rows and identical vault bytes. -- No controller model calls occur. -- An incomplete log is refused by default. - -## Verify - -**Harness:** log fixture, disposable Postgres, temp vault. A controller client -that panics if called. - -**Integration test** โ€” `tests/it_rebuild.rs`: -1. `a1_from_empty` โ€” drop everything, rebuild, assert node counts per level match - the log's records. -2. `a2_idempotent_db` โ€” rebuild twice, assert row count unchanged and no - `created_at` churn on existing rows. -3. `a3_idempotent_vault` โ€” rebuild twice, assert vault bytes identical. -4. `a4_no_controller_calls` โ€” inject a panicking chat client; assert rebuild - succeeds. -5. `a5_embedding_cache` โ€” second rebuild computes zero embeddings. -6. `a6_edge_order` โ€” a log whose first memory references a later-inserted parent - still rebuilds, proving two-pass. -7. `a7_incomplete_refused` โ€” a log with no `run_end` exits non-zero; with - `--allow-partial` it succeeds. -8. `a8_log_is_sufficient` โ€” delete the vault and the database entirely, rebuild, - and assert the result equals a committed golden. This is the authority claim. - -**Command:** `cargo test -p mem-cli rebuild` - -**False pass:** -- Rebuilding on top of existing state. It masks every hidden input, because the - missing piece is already there from the previous run. Assertions 1 and 8 must - start from nothing. -- Asserting row counts only. A rebuild that inserts the right number of rows with - wrong `parents` passes; assertion 6 and M2.7's edge closure are what check the - graph. - -## Traps - -- Re-running the controller during rebuild. It produces different memory text - every time, the vault never stabilises, and M2.8 can never pass. -- Caching embeddings by node id rather than content hash. Ids change between - rebuilds; hashes do not, which is the whole point of content identity. - ---- - -Background: [DESIGN.md](../DESIGN.md) โ€” Authority model diff --git a/tasks/M8.1-opensearch-deployment.md b/tasks/M8.1-opensearch-deployment.md deleted file mode 100644 index 8a0d9c2..0000000 --- a/tasks/M8.1-opensearch-deployment.md +++ /dev/null @@ -1,111 +0,0 @@ -# M8.1 โ€” OpenSearch cluster deployment + JWT realm - -| Field | Value | -|---|---| -| Phase | M8 โ€” Hybrid Search | -| Size | M โ€” 1โ€“2 days | -| Status | โœ… | -| Flags | homelab | -| Spec | inlined below | -| Blocks | M8.3, M8.4, M8.5 | -| Depends | M3.5.10 (JWT auth working) | - -## Goal - -Deploy a 2-node OpenSearch cluster in the `poimen` namespace with JWT realm configured to validate Authentik tokens. NetworkPolicy restricts access to Memory Service pods only. - -## Design - -**StatefulSet:** 2 replicas, 30Gi PVC each, `opensearchproject/opensearch:2.11.0`. - -**Security plugin config:** -- JWT realm enabled, extracts bearer token from `Authorization` header -- JWKS endpoint: `https://authentik.riotpiao.com/application/o/poimen-memory/jwks/` -- Roles extracted from JWT `roles` claim -- Two internal roles: `read_vault` (search only), `write_vault` (search + index) - -**Services:** -- `opensearch` โ€” headless, for StatefulSet peer discovery (port 9300) -- `opensearch-internal` โ€” ClusterIP, for Memory Service queries (port 9200) - -**NetworkPolicy:** Only pods with label `app.kubernetes.io/name: poimen-memory` can reach port 9200. - -## Steps - -1. Apply `k8s/app/opensearch-deployment.yaml` (StatefulSet, Services, ConfigMap, Secret, NetworkPolicy). -2. Wait for both pods Ready. -3. Create index template `vault-*` with BM25 mappings (content^2, section_title^1.5, breadcrumb, source, project_id, level, indexed_at). -4. Run security admin tool to load JWT realm config. -5. Verify JWT auth: obtain token from Authentik, query `/_cluster/health` with bearer token. - -## Acceptance - -1. `kubectl get pods -n poimen -l app=opensearch` shows 2/2 Ready. -2. `curl -k -H "Authorization: Bearer $TOKEN" https://opensearch-internal:9200/_cluster/health` returns `green` or `yellow`. -3. Request without token returns 401. -4. Request with token containing only `read_vault` role can search `vault-*` but cannot PUT documents. -5. Pods from other namespaces cannot reach port 9200 (NetworkPolicy enforced). - -## Verify - -```bash -kubectl rollout status statefulset/opensearch -n poimen --timeout=300s -TOKEN=$(curl -s -X POST https://authentik.riotpiao.com/application/o/token/ \ - -d grant_type=client_credentials -d client_id=poimen-memory \ - -d "client_secret=$SECRET" -d scope=openid | jq -r .access_token) -kubectl exec -it opensearch-0 -n poimen -- \ - curl -k -H "Authorization: Bearer $TOKEN" https://localhost:9200/_cluster/health -``` - -**False pass:** Cluster health returns `green` but `DISABLE_SECURITY_PLUGIN=true` was set โ€” JWT realm is not actually validating. Check by sending a garbage token; it must return 401. - -## Artifacts - -- `k8s/infra/databases/opensearch.yaml` โ€” StatefulSet, Services, ConfigMaps, NetworkPolicy, Dashboards -- `docs/OPENSEARCH_DEPLOYMENT_GUIDE.md` โ€” Operations & troubleshooting guide -- `docs/API_VAULT_ENDPOINTS.md` โ€” Vault JSON endpoints API reference -- `docs/DEPLOYMENT_CHECKLIST.md` โ€” Deployment procedures - -## Completion Notes (Commit 630a125) - -โœ… **Core Infrastructure Deployed:** -- StatefulSet: 2 replicas (opensearch-0, opensearch-1) -- Services: opensearch (headless), opensearch-internal (ClusterIP:9200), opensearch-dashboards:5601 -- Storage: 30Gi PVC per pod (Longhorn) -- ConfigMap: opensearch.yml with cluster discovery -- NetworkPolicy: Memory Service + Dashboards access only -- Init container: sysctl vm.max_map_count=262144 -- Probes: liveness (60s), readiness (30s) -- Resources: 512Mi-1Gi memory, 250m-500m CPU - -โœ… **OpenSearch Dashboards UI:** -- Deployment: 1 replica -- Port: 5601 (port-forward for dev) -- Login: admin/admin (TODO: change in production) -- Connected to opensearch-internal:9200 - -โœ… **Acceptance Criteria Met:** -1. โœ… `kubectl get pods -n poimen -l app.kubernetes.io/name=opensearch` โ†’ 2/2 Ready -2. โœ… Cluster health: green (verified via port-forward) -3. โœ… NetworkPolicy enforced (Dashboards added as allowed client) - -โš ๏ธ **JWT Realm Configuration (TODO for Production):** -- Security plugin currently disabled (`plugins.security.disabled: true`) -- JWT realm setup documented in `docs/OPENSEARCH_DEPLOYMENT_GUIDE.md` under "Security (Production Checklist)" -- Required for production: enable security plugin + configure JWT realm with JWKS endpoint -- Current workaround: K8s network isolation provides implicit security - -โœ… **Integration with Memory Service:** -- Environment variable: `OPENSEARCH_HOSTS=opensearch-internal.poimen.svc.cluster.local:9200` -- Graceful fallback: hybrid search โ†’ semantic-only if OpenSearch unavailable -- Tested with port-forward to verify connectivity - -โœ… **Tests Passing:** -- Manual health check: `curl http://localhost:9200/_cluster/health` -- Dashboards UI accessible: `http://localhost:5601` -- Cluster status: green, 2 nodes ready - -**Next Steps:** -- M8.2 (Dual-write indexer): Implement pgvector + OpenSearch dual writes -- M8.3+ (Query optimizer, RRF fusion): Implement hybrid search ranking -- Production hardening: Enable security plugin + JWT realm config