(plan) system review and break down plans
This commit is contained in:
@@ -0,0 +1,122 @@
|
||||
# M6.1 — CNPG `agent-manager-db` manifest
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| Phase | M6 — agent-manager migration |
|
||||
| Size | M — 1–3 days |
|
||||
| Status | ⬜ Not started |
|
||||
| Flags | homelab |
|
||||
| Spec | inlined below |
|
||||
| Blocks | — |
|
||||
|
||||
## Goal
|
||||
|
||||
A dedicated Postgres for agent-manager's session store, provisioned the way
|
||||
everything else in the cluster is: through git, with no manual `psql`. Same
|
||||
pattern this project already used for `memory-db` (M2.2), applied to a
|
||||
different, unrelated app.
|
||||
|
||||
## Facts (inlined — no spec read needed)
|
||||
|
||||
**agent-manager is a separate repo**, not part of this Rust workspace:
|
||||
`github.com/Riotpiaole/agent-manager` (fork of `github.com/YoanWai/
|
||||
agent-manager`), local checkout `~/workplace/agent-manager`, branch
|
||||
`add-headless-spawn`. Its session store is `internal/store/store.go` — see
|
||||
M6.2/M6.3 for the schema and query port.
|
||||
|
||||
**Existing CNPG pattern, verified on the live cluster: 3 Clusters today, one
|
||||
per app, each ClusterIP-only (not LAN-reachable):**
|
||||
|
||||
| namespace/name | app |
|
||||
|---|---|
|
||||
| `cicd/forgejo-db` | Forgejo |
|
||||
| `iam/authentik-db` | Authentik |
|
||||
| `temporal/temporal-db` | Temporal |
|
||||
|
||||
**No shared/multi-tenant DB cluster** — every app gets its own dedicated
|
||||
CNPG `Cluster`. `agent-manager-db` follows the same rule; it does not join
|
||||
`memory-db` (M2.2's cluster) even though both are new Postgres instances
|
||||
touched by the same person around the same time. Different app, different
|
||||
cluster.
|
||||
|
||||
Follow `k8s/infra/databases/temporal-db.yaml` exactly, same as M2.2 did:
|
||||
`imageName` pinned, `enableSuperuserAccess: false`, `storageClass:
|
||||
longhorn-cnpg`, `enablePodMonitor: true`, control-plane tolerations,
|
||||
`podAntiAffinityType: preferred`.
|
||||
|
||||
**Instance count — open question, default to convention.** Motivation for
|
||||
this whole migration is durability-of-location, not HA (single-machine
|
||||
usage, not a multi-host shared-session requirement). The 3 existing
|
||||
clusters are all 3-instance. Default to 3 instances for consistency with
|
||||
every other app in the cluster rather than special-casing this one to 1;
|
||||
revisit only if resource pressure on the homelab nodes makes it a real
|
||||
tradeoff.
|
||||
|
||||
Storage: session rows are tiny (`sessions`, `groups`, `settings`,
|
||||
`review_*` — no blobs beyond a `snapshot TEXT` pane capture per session).
|
||||
1Gi is generous; no need for `memory-db`'s 10Gi (that one holds
|
||||
768-dim vectors).
|
||||
|
||||
## Steps
|
||||
|
||||
1. `k8s/infra/databases/agent-manager-db.yaml` — `Cluster` + `Database`,
|
||||
namespace `agent-manager`, no extensions (plain relational, no
|
||||
pgvector).
|
||||
2. Namespace `agent-manager`, created by the ArgoCD app that owns it.
|
||||
3. Add to the owning kustomization's explicit resource list — an unlisted
|
||||
file is silently dropped with no error and no drift shown (the M2.2
|
||||
task file names this exact trap).
|
||||
4. Commit, push to **both** Forgejo origin and the GitHub mirror — verify
|
||||
which `repoURL` the eventual ArgoCD `Application` for this app actually
|
||||
watches before assuming either push is the one that matters (`kong`
|
||||
app, for example, tracks the GitHub mirror specifically, not Forgejo).
|
||||
5. Let ArgoCD sync. **No `kubectl apply`.**
|
||||
6. Verify the app user can create tables (schema arrives in M6.2, but a
|
||||
throwaway `CREATE TABLE t(id text); DROP TABLE t;` proves connectivity
|
||||
here).
|
||||
|
||||
## Acceptance
|
||||
|
||||
- `Cluster` reaches `Cluster in healthy state`.
|
||||
- ArgoCD shows the app `Synced/Healthy`.
|
||||
- No manual `psql` was run to get there.
|
||||
- Service is ClusterIP-only — not reachable from the LAN directly (M6.4's
|
||||
nginx route is the only path in).
|
||||
|
||||
## Verify
|
||||
|
||||
**Harness:** `kubectl` and `psql` read-only checks after sync.
|
||||
|
||||
**Integration test** — `verify/m6.1.sh`, output diffed against
|
||||
`expected/m6.1.txt`:
|
||||
1. `a1_cluster_healthy` — `kubectl get cluster -n agent-manager
|
||||
agent-manager-db` reports all instances ready.
|
||||
2. `a2_clusterip_only` — `kubectl get svc -n agent-manager -o
|
||||
jsonpath='{.items[*].spec.type}'` contains no `LoadBalancer` or
|
||||
`NodePort`.
|
||||
3. `a3_argocd_synced` — the owning app is `Synced/Healthy`.
|
||||
4. `a4_app_user_can_ddl` — as `app`, `CREATE TABLE t(id text); DROP TABLE
|
||||
t;` succeeds.
|
||||
5. `a5_no_lan_route_yet` — connection attempt from outside the cluster
|
||||
network fails at this point in the plan (M6.4 hasn't landed).
|
||||
|
||||
**Command:** `bash verify/m6.1.sh | diff - expected/m6.1.txt`
|
||||
|
||||
**False pass:**
|
||||
- Confirming sync without checking service type. A `Cluster` can be
|
||||
healthy and `Synced` while someone fat-fingered a `LoadBalancer` type
|
||||
into the manifest, silently violating the "dedicated ingress, not raw
|
||||
LAN IP" network-path decision this whole migration made. Assertion 2 is
|
||||
the guard.
|
||||
|
||||
## Traps
|
||||
|
||||
- Forgetting the kustomization resource list (same trap M2.2 already
|
||||
named) — file sits in git, ArgoCD reports Synced, objects never exist.
|
||||
- Adding `prune: true` without accounting for CNPG-operator-created
|
||||
children (Services, Secrets, PVCs). M2.2's Traps section already hit
|
||||
this on `llm-serving`; same fix applies here (`prune: false`).
|
||||
|
||||
---
|
||||
|
||||
Background: `k8s/infra/databases/temporal-db.yaml` · [M2.2](M2.2-memory-db-manifest.md) (same pattern, different app)
|
||||
Reference in New Issue
Block a user