Files
homelab/docs/adr/0001-gitops-bootstrap-and-cd.md
T

219 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ADR 0001 — GitOps Bootstrap, CD, and Coverage
- **Status:** Proposed
- **Date:** 2026-07-23
- **Deciders:** Homelab platform owner
- **Context tags:** bootstrap, gitops, argocd, forgejo, cicd, supply-chain, disaster-recovery
One delivery system, three coupled parts:
- **Part A — Bootstrap:** bare cluster → self-hosted GitOps control plane (breaks the ArgoCD ↔ Forgejo circle).
- **Part B — Steady-state CD:** how app changes flow to the cluster once it's up.
- **Part C — Coverage & organization:** every resource captured by exactly one Application; no orphans, no duplicates.
Single 3-control-plane Talos cluster. GitOps via ArgoCD (app-of-apps under `k8s/argocd/apps/`, waves 08); CI via Forgejo Actions on a self-hosted runner.
---
## Part A — Bootstrap
### Problem — the circle
```
ArgoCD ──syncs from──▶ git repo ──hosted by──▶ Forgejo ──deployed by──▶ (the repo)
```
Forgejo hosts the repo ArgoCD reconciles. Today the circle is broken by installing Forgejo + Longhorn + CNPG **by hand** (`bootstrap-local/`, `bootstrap/phase1-5/`) — imperative, off-git, never reconciled, human-ordered rebuild.
### Decision — GitHub seed + Forgejo pull-mirror + cutover
GitHub is the seed and permanent DR source; Forgejo is steady-state truth after cutover and stays synced via a native pull-mirror.
**No adoption.** Longhorn, CNPG, and Forgejo are **pure Phase-0 manual bootstrap** — ArgoCD never manages them (this is the deliberate divergence from the "ArgoCD adopts hand-installed infra" pattern). They must exist *before* ArgoCD can sync from a git host at all, so making them Apps would re-introduce the circle. ArgoCD (seeded from GitHub) deploys only everything *after* the control plane.
```
Phase 0 (manual, once — the whole control plane):
Talos ─ Cilium CNI ─ Longhorn ─ CNPG operator ─ forgejo-db Cluster (wait Ready)
─ Forgejo + Redis ─ push repo ─ helm install argocd ─ apply root-app→GitHub
Phase 1 (ArgoCD from GitHub seed): waves 08 = monitoring · logging · iam · messaging ·
apps. Does NOT manage Longhorn / CNPG / Forgejo (those stay bootstrap-owned).
Phase 2 (cutover, once): flip root-app repoURL GitHub→Forgejo; Forgejo pull-mirrors
GitHub (live DR seed).
```
Manual order matters: CNPG operator + `forgejo-db` Cluster must be **Ready** (`.status.phase: "Cluster in healthy state"`, not just created) before `helm install forgejo`. Wait on it. `forgejo-db` = dedicated CNPG Cluster (`owner: forgejo`), preferred anti-affinity + control-plane tolerations so it can't deadlock on a partly-schedulable 3-CP.
**Options:** (A) one-time seed + manual cutover — mirror drifts. (B) **seed + pull-mirror + cutover — chosen** — safe cutover, always-current DR. (C) external-permanent — no self-hosted control plane. Rejected.
### Repository credentials (private GitHub seed)
Two read-only creds, neither needs write:
1. **ArgoCD → GitHub: SSH deploy key.** `ssh-keygen -t ed25519 -f argocd_seed -N ""`; add public key to repo → Deploy keys (read-only). Private key → ArgoCD repo-credential Secret at Phase 0 (before any secret controller):
```yaml
apiVersion: v1
kind: Secret
metadata:
name: seed-github-repo
namespace: argocd
labels: { argocd.argoproj.io/secret-type: repository }
stringData:
type: git
url: [email protected]:<org>/homelab.git # SSH; must match root-app repoURL
sshPrivateKey: |
-----BEGIN OPENSSH PRIVATE KEY-----
...
```
2. **Forgejo → GitHub: fine-grained PAT (Contents: Read).** HTTPS token for the pull-mirror; SOPS-encrypted in-repo (consumed at w1 once SOPS plugin is up).
Rotation: replace GitHub-side key/token, re-apply Secret / re-encrypt SOPS. Cutover keeps both (mirror stays live).
---
## Part B — Steady-state CD
**CI does build + push + sign only.** Dedicated controllers own promotion, reconciliation, safe rollout. Git = single source; nothing mutates cluster or CD repo by string-replace.
### Current problems
1. **`sed` image-bump** (`ci.yml:72`) — fragile, non-atomic, races, `git revert`-only rollback. "CI mutates CD repo" anti-pattern.
2. **Registry hostname split-brain** — one Forgejo addressed 4 ways; in-cluster pods pull via public ingress name (nginx hairpin, CLAUDE.md trap). `api` defaults to `:latest`.
3. **Fragmented truth** — `project: default` vs `homelab`; `api` App points at external `rock/deploy.git`.
4. **DinD runner** — privileged; RWO cache pins it to cp-1 (SPOF).
5. **No progressive delivery / supply-chain gate** — all `prune+selfHeal`, no canary/analysis/rollback, unsigned images.
### Decision
| Concern | From | To |
|---|---|---|
| Image promotion | CI `sed` → deploy repo | **ArgoCD Image Updater**, git write-back, digest-pinned |
| Build | DinD (privileged, cp-1) | **rootless BuildKit** (or Kaniko) + registry cache |
| Registry addressing | 4 hostnames | in-cluster `*.svc.cluster.local:3000`; one external push name |
| Delivery | raw auto-sync | **Argo Rollouts** canary + Prometheus `AnalysisTemplate` (app tier) |
| Supply chain | none | **cosign sign + Kyverno `verifyImages`** |
| Success signal | pod `Running` | **Playwright E2E** (video/trace) gate |
| Projects/repos | `default`+`homelab`, 2 repos | one `homelab` project, one convention |
```
push→main → runner(BuildKit): test · build·push <in-cluster-registry>/<repo>@sha256 · cosign sign
→ Image Updater watches registry, writes digest back to git
→ app-of-apps: infra/platform = plain sync ; app tier = Rollouts canary + Kyverno gate
```
**Options:** promotion — Image Updater (chosen) over `kustomize edit` (still coupled) / Flux (2nd engine). Build — BuildKit (chosen) over DinD / Kaniko. Delivery — Rollouts app-tier-only (chosen) over none / Flagger (needs mesh).
### Deployment verification (synthetic E2E)
`Running` proves scheduling, not the user path (ingress/TLS/OIDC/session/UI). Signal ladder: readiness → synthetic HTTP → **Playwright browser E2E** (gate on the top rung). Suite hits **public ingress hostnames**, records video+trace → MinIO `e2e-artifacts/` (private, lifecycle-expired). Journeys:
- **Portainer** — login page renders.
- **Authentik** — sign in (bootstrap `akadmin` from secret store), dashboard loads.
- **OAuth/OIDC federation** — each app (ArgoCD, Forgejo, Grafana…): click SSO → complete Authentik → land back authenticated. Catches client mis-registration / redirect-URI / broken provider that a green pod hides.
Runs two ways: **deploy gate** (Rollouts `AnalysisTemplate` / PostSync Job → fails sync/rollout) and **continuous smoke** (`CronJob` → Alertmanager). Creds from SOPS/Vault as env, never hardcoded. Constraint: **WebKit only** (Safari engine) per requirement. Suite lives at `tests/e2e/`; manifests at `k8s/platform/e2e/` (unwired until phase 7).
---
## Part C — GitOps coverage & organization
**Principle: every deployable directory is the source of exactly one Application, grouped by wave.** No dead aggregator kustomizations, no orphan manifests, no second app-of-apps. Layer folders stay; the tree is reconciled to the map. An App uses a git **`path:`** for manifests/kustomize bases, a **`$values` valueFile** when it only parameterizes a remote Helm chart. Secrets are the exception — one `sops-secrets` App globs all `*.enc.yaml` from repo root.
### Folder ↔ Application map (as-built)
Consolidated to two folders — `k8s/infra` (platform/infra tier) and `k8s/apps`
(workloads) — replacing the old `infrastructure/platform/security/applications`
layer folders (deleted).
```
k8s/
├── argocd/ root/ + apps/ (waves) + projects/ # engine, not an App target
├── bootstrap/ cert-manager w1 · cilium w0 · coredns w0 · ingress w1 # Phase-0 substrate
├── infra/
│ ├── longhorn/ →longhorn-config w1
│ ├── minio/ →minio-tenant (+minio-operator-values for operator) w1
│ ├── monitoring/ →monitoring-config w2 · crds/ →prometheus-crds w0
│ │ *-values →prometheus/blackbox w1
│ ├── logging/*-values →grafana/loki/promtail w2
│ ├── data-schemas/ →data-schemas w6 (CNPG operator + forgejo-db → Phase-0 manual bootstrap, NOT an App)
│ ├── iam/ →iam-jobs w3 (+authentik/vault values)
│ ├── forgejo-runner/ →forgejo-runner w3
│ └── cicd/ →cicd-config w3 (Forgejo OAuth job)
└── apps/
├── cloudflared/ duckdns/ homarr/ portainer/ temporal/ →w8
└── messaging/{kafka-cluster,queue-crd,management-service} →w57
```
### Orphans — declared in git, reconciled by NOTHING
| Path | Count | Kind | Action |
|---|---|---|---|
| `platform/monitoring/{alerts,servicemonitors,dashboards}` | 27 | PrometheusRule / ServiceMonitor / dashboard ConfigMap | wire → **`monitoring-config` App (done)** |
| `security/ci-cd/forgejo-oauth-setup-job.yaml` | 1 | Job+SA (PostSync) | wire → `cicd-config` App |
| CNPG operator + `forgejo-db` Cluster | — | CNPG install + Cluster | **Phase-0 manual bootstrap** (`bootstrap/phase2-cnpg`, `bootstrap/phase3-forgejo`), never an App — making it an App re-introduces the circle |
Root cause (monitoring): no App used `path: k8s/platform/monitoring` — only crds + valueFiles; the aggregator kustomization chain was dead. Now fixed by one `monitoring-config` App (one kustomization, **no `namespace:` transformer** — PrometheusRules live in per-app namespaces; `ServerSideApply=true` for large dashboard JSON).
### Stale / dead — delete, do NOT wire
- `applications/sqs/argocd/*` — **second competing app-of-apps** (repo `rock/kafaka-management-service.git`, project `kmsvc`).
- `applications/sqs/{queues,environments,helmfile.yaml.gotmpl,kustomization.yaml}` — helmfile-era.
- **`applications/temporal/elasticsearch.yaml`** — superseded (see Temporal datastore below).
- `applications/temporal/queues/`, root `applications/kustomization.yaml` — example / legacy overlay.
- `security/iam/authentik-migrations-job.yaml` — dead/broken (server self-migrates).
- Dead aggregator kustomizations (`infrastructure/`, `platform/`, `security/`, `ci-cd/`), `logging/minio-values.yaml`, `infrastructure/minio/test/`, scripts / `.env.example`.
**Experimental — leave unmanaged** (per `60-applications.yaml`): `applications/{dev-tools,forge,llm,shadowsocks}`.
### 🔴 Temporal datastore = PostgreSQL (hard constraint)
Temporal runs on the shared CNPG **PostgreSQL** (`ddb-cluster`), for BOTH the main store and **visibility**. **NOT Cassandra** (the Helm chart's default) and **NOT Elasticsearch.** `temporal-values.yaml` sets `elasticsearch.enabled: false` and points visibility at Postgres (`temporal-visibility-database`). Therefore `applications/temporal/elasticsearch.yaml` is dead → delete. When touching Temporal values, verify persistence stays on the pinned chart's SQL schema (chart schema drift silently reverts to Cassandra — see CLAUDE.md).
### Reconcile actions (tree → map)
1. **Restore** live-consumed paths currently staged-deleted: `infrastructure/{longhorn,minio}`, `security/iam`, `security/ci-cd/charts/forgejo-runner`, `platform/{logging,monitoring}` values + crds. Without them their Apps have no source.
2. **Add Apps:** `monitoring-config` (done), `cicd-config` (oauth job).
3. **Delete** the stale/dead list above.
4. **CNPG** (operator + `forgejo-db` Cluster) stays **Phase-0 manual bootstrap** — never an App (circular).
⚠️ The current staged deletion is **over-broad** — it wiped live GitOps sources (longhorn, minio, iam, monitoring/logging values). Committing as-is breaks those Apps' sync. Curate: restore live, delete only confirmed orphans.
ApplicationSet git-generator considered, rejected: the hand-tuned per-App settings (multi-source helm+values, `managedNamespaceMetadata`, `ServerSideApply`, waves) don't fit one template.
---
## Consequences
**Positive:** Forgejo/Longhorn/CNPG fully declarative; Phase-0 shrinks to "CNI + ArgoCD + root-app"; repeatable rebuild seeded from always-current GitHub (free off-site DR); CI shrinks to build/push/sign; digest-pinned auditable promotions; no privileged DinD; CI survives cp-1 loss; auto-rollback on SLO breach; only signed images admit; monitoring finally reconciled.
**Costs:** GitHub egress + deploy key needed at Phase 0; one cutover flip remains; two read creds (deploy key + SOPS PAT); three new controllers (Image Updater, Rollouts, Kyverno) + `Deployment`→`Rollout` migration; one-time registry-hostname normalization.
---
## Migration phases (one PR each, impact-over-risk)
0. **Consolidate bootstrap + coverage.** Single-source bootstrap; convert phase13 to ArgoCD Apps (seed repo); Forgejo pull-mirror + SOPS PAT; script = Phase-0 + cutover only. Curate the staged deletion (Part C reconcile): restore live paths, add `monitoring-config`(done)+`cicd-config`, delete stale/dead + `data/cluster`. CNPG (operator + `forgejo-db`) stays Phase-0 manual bootstrap.
1. **Kill `sed`.** `argocd-image-updater` (~w4); annotate `image-list` + git write-back; drop ci.yml step 3.
2. **Normalize registry.** In-cluster pulls via `forgejo-gitea-http.cicd.svc.cluster.local:3000`; retire `forgejo.forge.*`; pin digests.
3. **Rootless BuildKit.** Drop `az-a` nodeSelector; `--cache-to type=registry`.
4. **Argo Rollouts.** Controller (~w4); `api`/app tier → `Rollout` + canary `AnalysisTemplate`.
5. **Supply chain.** cosign (Forgejo OIDC) + Kyverno `verifyImages`.
6. **Consolidate projects.** One `homelab` project; monorepo-only vs single `deploy` repo.
7. **Synthetic E2E.** Wire `tests/e2e` suite as Rollouts/PostSync gate + `CronJob`; Portainer + Authentik + OAuth journeys.
---
## Open questions
1. ArgoCD install — `helm install` (leaning) vs Talos `extraManifests`.
2. Post-cutover primary — push GitHub (mirror pulls) vs push Forgejo (push-mirror out). Pull-from-GitHub keeps GitHub authoritative.
3. Cutover — scripted `kubectl patch` of root repoURL vs PostSync hook gated on Forgejo health.
---
## References
- `CLAUDE.md` — ArgoCD gotchas (SkipDryRunOnMissingResource, wait-for-Ready, repoURL/ingress hairpin, Helm schema drift → Temporal/Cassandra)
- `k8s/argocd/{root,apps,projects}/` · `k8s/bootstrap/` (to be superseded)
- `k8s/security/ci-cd/{example-workflows/ci.yml,charts/forgejo-runner,deploy-scaffold}`
- `tests/e2e/` (Playwright WebKit suite) · `k8s/platform/monitoring/` (monitoring-config)