14 KiB
ADR 0001 — GitOps Bootstrap, CD, and Coverage
- Status: Proposed
- Date: 2026-07-23
- Deciders: Homelab platform owner
- Context tags: bootstrap, gitops, argocd, forgejo, cicd, supply-chain, disaster-recovery
One delivery system, three coupled parts:
- Part A — Bootstrap: bare cluster → self-hosted GitOps control plane (breaks the ArgoCD ↔ Forgejo circle).
- Part B — Steady-state CD: how app changes flow to the cluster once it's up.
- Part C — Coverage & organization: every resource captured by exactly one Application; no orphans, no duplicates.
Single 3-control-plane Talos cluster. GitOps via ArgoCD (app-of-apps under k8s/argocd/apps/, waves 0–8); CI via Forgejo Actions on a self-hosted runner.
Part A — Bootstrap
Problem — the circle
ArgoCD ──syncs from──▶ git repo ──hosted by──▶ Forgejo ──deployed by──▶ (the repo)
Forgejo hosts the repo ArgoCD reconciles. Today the circle is broken by installing Forgejo + Longhorn + CNPG by hand (bootstrap-local/, bootstrap/phase1-5/) — imperative, off-git, never reconciled, human-ordered rebuild.
Decision — GitHub seed + Forgejo pull-mirror + cutover
GitHub is the seed and permanent DR source; Forgejo is steady-state truth after cutover and stays synced via a native pull-mirror.
No adoption. Longhorn, CNPG, and Forgejo are pure Phase-0 manual bootstrap — ArgoCD never manages them (this is the deliberate divergence from the "ArgoCD adopts hand-installed infra" pattern). They must exist before ArgoCD can sync from a git host at all, so making them Apps would re-introduce the circle. ArgoCD (seeded from GitHub) deploys only everything after the control plane.
Phase 0 (manual, once — the whole control plane):
Talos ─ Cilium CNI ─ Longhorn ─ CNPG operator ─ forgejo-db Cluster (wait Ready)
─ Forgejo + Redis ─ push repo ─ helm install argocd ─ apply root-app→GitHub
Phase 1 (ArgoCD from GitHub seed): waves 0–8 = monitoring · logging · iam · messaging ·
apps. Does NOT manage Longhorn / CNPG / Forgejo (those stay bootstrap-owned).
Phase 2 (cutover, once): flip root-app repoURL GitHub→Forgejo; Forgejo pull-mirrors
GitHub (live DR seed).
Manual order matters: CNPG operator + forgejo-db Cluster must be Ready (.status.phase: "Cluster in healthy state", not just created) before helm install forgejo. Wait on it. forgejo-db = dedicated CNPG Cluster (owner: forgejo), preferred anti-affinity + control-plane tolerations so it can't deadlock on a partly-schedulable 3-CP.
Options: (A) one-time seed + manual cutover — mirror drifts. (B) seed + pull-mirror + cutover — chosen — safe cutover, always-current DR. (C) external-permanent — no self-hosted control plane. Rejected.
Repository credentials (private GitHub seed)
Two read-only creds, neither needs write:
- ArgoCD → GitHub: SSH deploy key.
ssh-keygen -t ed25519 -f argocd_seed -N ""; add public key to repo → Deploy keys (read-only). Private key → ArgoCD repo-credential Secret at Phase 0 (before any secret controller):apiVersion: v1 kind: Secret metadata: name: seed-github-repo namespace: argocd labels: { argocd.argoproj.io/secret-type: repository } stringData: type: git url: [email protected]:<org>/homelab.git # SSH; must match root-app repoURL sshPrivateKey: | -----BEGIN OPENSSH PRIVATE KEY----- ... - Forgejo → GitHub: fine-grained PAT (Contents: Read). HTTPS token for the pull-mirror; SOPS-encrypted in-repo (consumed at w1 once SOPS plugin is up).
Rotation: replace GitHub-side key/token, re-apply Secret / re-encrypt SOPS. Cutover keeps both (mirror stays live).
Part B — Steady-state CD
CI does build + push + sign only. Dedicated controllers own promotion, reconciliation, safe rollout. Git = single source; nothing mutates cluster or CD repo by string-replace.
Current problems
sedimage-bump (ci.yml:72) — fragile, non-atomic, races,git revert-only rollback. "CI mutates CD repo" anti-pattern.- Registry hostname split-brain — one Forgejo addressed 4 ways; in-cluster pods pull via public ingress name (nginx hairpin, CLAUDE.md trap).
apidefaults to:latest. - Fragmented truth —
project: defaultvshomelab;apiApp points at externalrock/deploy.git. - DinD runner — privileged; RWO cache pins it to cp-1 (SPOF).
- No progressive delivery / supply-chain gate — all
prune+selfHeal, no canary/analysis/rollback, unsigned images.
Decision
| Concern | From | To |
|---|---|---|
| Image promotion | CI sed → deploy repo |
ArgoCD Image Updater, git write-back, digest-pinned |
| Build | DinD (privileged, cp-1) | rootless BuildKit (or Kaniko) + registry cache |
| Registry addressing | 4 hostnames | in-cluster *.svc.cluster.local:3000; one external push name |
| Delivery | raw auto-sync | Argo Rollouts canary + Prometheus AnalysisTemplate (app tier) |
| Supply chain | none | cosign sign + Kyverno verifyImages |
| Success signal | pod Running |
Playwright E2E (video/trace) gate |
| Projects/repos | default+homelab, 2 repos |
one homelab project, one convention |
push→main → runner(BuildKit): test · build·push <in-cluster-registry>/<repo>@sha256 · cosign sign
→ Image Updater watches registry, writes digest back to git
→ app-of-apps: infra/platform = plain sync ; app tier = Rollouts canary + Kyverno gate
Options: promotion — Image Updater (chosen) over kustomize edit (still coupled) / Flux (2nd engine). Build — BuildKit (chosen) over DinD / Kaniko. Delivery — Rollouts app-tier-only (chosen) over none / Flagger (needs mesh).
Deployment verification (synthetic E2E)
Running proves scheduling, not the user path (ingress/TLS/OIDC/session/UI). Signal ladder: readiness → synthetic HTTP → Playwright browser E2E (gate on the top rung). Suite hits public ingress hostnames, records video+trace → MinIO e2e-artifacts/ (private, lifecycle-expired). Journeys:
- Portainer — login page renders.
- Authentik — sign in (bootstrap
akadminfrom secret store), dashboard loads. - OAuth/OIDC federation — each app (ArgoCD, Forgejo, Grafana…): click SSO → complete Authentik → land back authenticated. Catches client mis-registration / redirect-URI / broken provider that a green pod hides.
Runs two ways: deploy gate (Rollouts AnalysisTemplate / PostSync Job → fails sync/rollout) and continuous smoke (CronJob → Alertmanager). Creds from SOPS/Vault as env, never hardcoded. Constraint: WebKit only (Safari engine) per requirement. Suite lives at tests/e2e/; manifests at k8s/platform/e2e/ (unwired until phase 7).
Part C — GitOps coverage & organization
Principle: every deployable directory is the source of exactly one Application, grouped by wave. No dead aggregator kustomizations, no orphan manifests, no second app-of-apps. Layer folders stay; the tree is reconciled to the map. An App uses a git path: for manifests/kustomize bases, a $values valueFile when it only parameterizes a remote Helm chart. Secrets are the exception — one sops-secrets App globs all *.enc.yaml from repo root.
Folder ↔ Application map (as-built)
Consolidated to two folders — k8s/infra (platform/infra tier) and k8s/apps
(workloads) — replacing the old infrastructure/platform/security/applications
layer folders (deleted).
k8s/
├── argocd/ root/ + apps/ (waves) + projects/ # engine, not an App target
├── bootstrap/ cert-manager w1 · cilium w0 · coredns w0 · ingress w1 # Phase-0 substrate
├── infra/
│ ├── longhorn/ →longhorn-config w1
│ ├── minio/ →minio-tenant (+minio-operator-values for operator) w1
│ ├── monitoring/ →monitoring-config w2 · crds/ →prometheus-crds w0
│ │ *-values →prometheus/blackbox w1
│ ├── logging/*-values →grafana/loki/promtail w2
│ ├── data-schemas/ →data-schemas w6 (CNPG operator + forgejo-db → Phase-0 manual bootstrap, NOT an App)
│ ├── iam/ →iam-jobs w3 (+authentik/vault values)
│ ├── forgejo-runner/ →forgejo-runner w3
│ └── cicd/ →cicd-config w3 (Forgejo OAuth job)
└── apps/
├── cloudflared/ duckdns/ homarr/ portainer/ temporal/ →w8
└── messaging/{kafka-cluster,queue-crd,management-service} →w5–7
Orphans — declared in git, reconciled by NOTHING
| Path | Count | Kind | Action |
|---|---|---|---|
platform/monitoring/{alerts,servicemonitors,dashboards} |
27 | PrometheusRule / ServiceMonitor / dashboard ConfigMap | wire → monitoring-config App (done) |
security/ci-cd/forgejo-oauth-setup-job.yaml |
1 | Job+SA (PostSync) | wire → cicd-config App |
CNPG operator + forgejo-db Cluster |
— | CNPG install + Cluster | Phase-0 manual bootstrap (bootstrap/phase2-cnpg, bootstrap/phase3-forgejo), never an App — making it an App re-introduces the circle |
Root cause (monitoring): no App used path: k8s/platform/monitoring — only crds + valueFiles; the aggregator kustomization chain was dead. Now fixed by one monitoring-config App (one kustomization, no namespace: transformer — PrometheusRules live in per-app namespaces; ServerSideApply=true for large dashboard JSON).
Stale / dead — delete, do NOT wire
applications/sqs/argocd/*— second competing app-of-apps (reporock/kafaka-management-service.git, projectkmsvc).applications/sqs/{queues,environments,helmfile.yaml.gotmpl,kustomization.yaml}— helmfile-era.applications/temporal/elasticsearch.yaml— superseded (see Temporal datastore below).applications/temporal/queues/, rootapplications/kustomization.yaml— example / legacy overlay.security/iam/authentik-migrations-job.yaml— dead/broken (server self-migrates).- Dead aggregator kustomizations (
infrastructure/,platform/,security/,ci-cd/),logging/minio-values.yaml,infrastructure/minio/test/, scripts /.env.example.
Experimental — leave unmanaged (per 60-applications.yaml): applications/{dev-tools,forge,llm,shadowsocks}.
🔴 Temporal datastore = PostgreSQL (hard constraint)
Temporal runs on the shared CNPG PostgreSQL (ddb-cluster), for BOTH the main store and visibility. NOT Cassandra (the Helm chart's default) and NOT Elasticsearch. temporal-values.yaml sets elasticsearch.enabled: false and points visibility at Postgres (temporal-visibility-database). Therefore applications/temporal/elasticsearch.yaml is dead → delete. When touching Temporal values, verify persistence stays on the pinned chart's SQL schema (chart schema drift silently reverts to Cassandra — see CLAUDE.md).
Reconcile actions (tree → map)
- Restore live-consumed paths currently staged-deleted:
infrastructure/{longhorn,minio},security/iam,security/ci-cd/charts/forgejo-runner,platform/{logging,monitoring}values + crds. Without them their Apps have no source. - Add Apps:
monitoring-config(done),cicd-config(oauth job). - Delete the stale/dead list above.
- CNPG (operator +
forgejo-dbCluster) stays Phase-0 manual bootstrap — never an App (circular).
⚠️ The current staged deletion is over-broad — it wiped live GitOps sources (longhorn, minio, iam, monitoring/logging values). Committing as-is breaks those Apps' sync. Curate: restore live, delete only confirmed orphans.
ApplicationSet git-generator considered, rejected: the hand-tuned per-App settings (multi-source helm+values, managedNamespaceMetadata, ServerSideApply, waves) don't fit one template.
Consequences
Positive: Forgejo/Longhorn/CNPG fully declarative; Phase-0 shrinks to "CNI + ArgoCD + root-app"; repeatable rebuild seeded from always-current GitHub (free off-site DR); CI shrinks to build/push/sign; digest-pinned auditable promotions; no privileged DinD; CI survives cp-1 loss; auto-rollback on SLO breach; only signed images admit; monitoring finally reconciled.
Costs: GitHub egress + deploy key needed at Phase 0; one cutover flip remains; two read creds (deploy key + SOPS PAT); three new controllers (Image Updater, Rollouts, Kyverno) + Deployment→Rollout migration; one-time registry-hostname normalization.
Migration phases (one PR each, impact-over-risk)
- Consolidate bootstrap + coverage. Single-source bootstrap; convert phase1–3 to ArgoCD Apps (seed repo); Forgejo pull-mirror + SOPS PAT; script = Phase-0 + cutover only. Curate the staged deletion (Part C reconcile): restore live paths, add
monitoring-config(done)+cicd-config, delete stale/dead +data/cluster. CNPG (operator +forgejo-db) stays Phase-0 manual bootstrap. - Kill
sed.argocd-image-updater(~w4); annotateimage-list+ git write-back; drop ci.yml step 3. - Normalize registry. In-cluster pulls via
forgejo-gitea-http.cicd.svc.cluster.local:3000; retireforgejo.forge.*; pin digests. - Rootless BuildKit. Drop
az-anodeSelector;--cache-to type=registry. - Argo Rollouts. Controller (~w4);
api/app tier →Rollout+ canaryAnalysisTemplate. - Supply chain. cosign (Forgejo OIDC) + Kyverno
verifyImages. - Consolidate projects. One
homelabproject; monorepo-only vs singledeployrepo. - Synthetic E2E. Wire
tests/e2esuite as Rollouts/PostSync gate +CronJob; Portainer + Authentik + OAuth journeys.
Open questions
- ArgoCD install —
helm install(leaning) vs TalosextraManifests. - Post-cutover primary — push GitHub (mirror pulls) vs push Forgejo (push-mirror out). Pull-from-GitHub keeps GitHub authoritative.
- Cutover — scripted
kubectl patchof root repoURL vs PostSync hook gated on Forgejo health.
References
CLAUDE.md— ArgoCD gotchas (SkipDryRunOnMissingResource, wait-for-Ready, repoURL/ingress hairpin, Helm schema drift → Temporal/Cassandra)k8s/argocd/{root,apps,projects}/·k8s/bootstrap/(to be superseded)k8s/security/ci-cd/{example-workflows/ci.yml,charts/forgejo-runner,deploy-scaffold}tests/e2e/(Playwright WebKit suite) ·k8s/platform/monitoring/(monitoring-config)