docs: rewrite CLAUDE.md for ArgoCD GitOps, track in git
This commit is contained in:
+179
-55
@@ -1,86 +1,210 @@
|
||||
# CLAUDE.md — Homelab Integration Guide
|
||||
# CLAUDE.md — Homelab Integration Guide (Example / Reusable Template)
|
||||
|
||||
**Homelab:** A bare-metal three-node Kubernetes cluster running Talos Linux with a full observability stack, SSO via Authentik, secret management via Vault, and CI/CD infrastructure (Forgejo + Argo CD, deployed).
|
||||
> **This is a sanitized template.** Copy to `CLAUDE.md`, fill in your own
|
||||
> node IPs/hostnames/secrets, and delete this notice. Nothing in this file
|
||||
> should contain real credentials, real IPs beyond illustrative examples, or
|
||||
> anything that would matter if this file became public. It's meant to be
|
||||
> shared across homelabs running similar hardware/topology (3-node bare-metal
|
||||
> Talos Kubernetes + ArgoCD GitOps), not just this one.
|
||||
|
||||
## Cluster Topology (3 control-plane HA, since 2026-07-20)
|
||||
**Homelab:** A bare-metal N-node Kubernetes cluster running Talos Linux with a
|
||||
full observability stack, SSO via Authentik, secret management via Vault, and
|
||||
CI/CD infrastructure (self-hosted git forge + Argo CD).
|
||||
|
||||
## Cluster Topology (adjust to your hardware)
|
||||
|
||||
| Node | IP | Zone | Scheduling | Storage |
|
||||
|------|----|----|-----------|---------|
|
||||
| `talos-cp-1` | .213 | az-a | schedulable (all workloads) | sole Longhorn node |
|
||||
| `talos-cp-2` | .163 | az-b | dedicated (`NoSchedule`) | none |
|
||||
| `talos-cp-3` | .166 | az-c | dedicated (`NoSchedule`) | none |
|
||||
| `<node-1>` | `<ip>` | az-a | schedulable (all workloads) | sole storage node (if single-node storage) |
|
||||
| `<node-2>` | `<ip>` | az-b | dedicated (`NoSchedule`) | none |
|
||||
| `<node-3>` | `<ip>` | az-c | dedicated (`NoSchedule`) | none |
|
||||
|
||||
3 voting etcd members peering on the LAN. Only `talos-cp-1` runs workloads and
|
||||
holds storage → stateful services are single-instance (e.g. CNPG `ddb-cluster`
|
||||
= 1 instance). API endpoint is single-homed to `.213` (no VIP yet).
|
||||
If your storage layer (Longhorn, local-path, etc.) only runs on one node,
|
||||
every stateful workload pinned to that storage class is effectively
|
||||
single-instance regardless of your control-plane HA count — document that
|
||||
explicitly here, it changes your failure-mode assumptions everywhere else.
|
||||
|
||||
## Deployment Model: ArgoCD GitOps (app-of-apps)
|
||||
|
||||
```
|
||||
git commit → git push (your git forge) → ArgoCD auto-sync → cluster
|
||||
```
|
||||
|
||||
Structure:
|
||||
- One root `Application` (`k8s/argocd/root/`) pointing at a directory of
|
||||
child `Application` manifests (`k8s/argocd/apps/*.yaml`)
|
||||
- Each child Application is either:
|
||||
- A remote Helm chart + a **second** git source (`ref: values`) supplying
|
||||
just the values file — lets you pin an upstream chart version while
|
||||
keeping your values under normal git history/review
|
||||
- A plain git directory of raw manifests (optionally with a
|
||||
`kustomization.yaml`)
|
||||
- `argocd.argoproj.io/sync-wave` annotations control ordering across
|
||||
Applications (lower number syncs first)
|
||||
|
||||
**Never `kubectl apply`/`patch`/`delete` a resource ArgoCD manages**, except:
|
||||
- Pure cleanup of stuck/dead state (e.g. deleting a permanently-failed hook
|
||||
Job so the next real sync can create a fresh one) — this is not a config
|
||||
change, just clearing wreckage that GitOps itself won't clean up
|
||||
automatically (see Gotchas below)
|
||||
- Genuine one-time bootstrap circular dependencies (e.g. Vault
|
||||
`operator init`/unseal — nothing can configure Vault's own unseal keys
|
||||
before Vault has generated them)
|
||||
|
||||
## Hard Rules (adapt freely, but keep something like these)
|
||||
|
||||
🔴 **Identify and document your storage/stateful-singleton node explicitly.**
|
||||
Whatever node holds your CSI driver's data (Longhorn, local-path, etc.),
|
||||
renaming or wiping it orphans every PVC pinned there. Name it here, in
|
||||
caps, so nobody "cleans up" it by accident.
|
||||
|
||||
🔴 **If you run multi-member etcd across a LAN + VPN/WireGuard overlay,
|
||||
pin the advertised subnet explicitly** (e.g. Talos's
|
||||
`cluster.etcd.advertisedSubnets`). Without it, etcd may advertise on the
|
||||
wrong interface and new members hang as non-promoting learners.
|
||||
|
||||
🔴 **Run your IaC formatter (terraform fmt, etc.) before every commit
|
||||
that touches infra code.** Wire this into CI as a hard gate, not a
|
||||
suggestion.
|
||||
|
||||
🔴 **Whatever your source of truth is (Terraform, ArgoCD, both) — never
|
||||
manually mutate resources it manages.** State drift is the single most
|
||||
common cause of "why did my last apply undo my manual fix" confusion.
|
||||
Fix the source, re-apply/re-sync, never bypass.
|
||||
|
||||
🔴 **Never delete a PVC without confirming replica count / backup
|
||||
freshness first.** This is always a one-way door.
|
||||
|
||||
🔴 **Decide your commit message convention up front and enforce it.**
|
||||
(This template's origin project uses: no co-authored-by footers, single-line
|
||||
commit summarizing what/why, solo-authorship assumption — adjust to your
|
||||
team's norms.)
|
||||
|
||||
🔴 **Decide your git workflow (rebase vs merge) up front and stick to it**
|
||||
cluster-wide, across every contributor/agent working in the repo.
|
||||
|
||||
🔴 **Long-running commands should not block a synchronous session** — run
|
||||
them in the background and poll, especially anything that waits on a
|
||||
Kubernetes rollout, an image pull, or a Terraform apply.
|
||||
|
||||
## GitOps / ArgoCD Gotchas (transferable to any ArgoCD-based homelab)
|
||||
|
||||
🟠 **A `kustomization.yaml` with an explicit `resources:` allowlist
|
||||
silently drops anything you forget to list.** No error, no drift shown in
|
||||
ArgoCD's UI — it just reports `Synced/Healthy` against a manifest set that
|
||||
never included your new file. Always run `kubectl kustomize <dir>/`
|
||||
locally before pushing to confirm exactly what ArgoCD will build.
|
||||
|
||||
🟠 **A top-level `namespace:` transformer in `kustomization.yaml` rewrites
|
||||
`metadata.namespace` on every resource it builds** — including RBAC
|
||||
bindings deliberately targeting a *different* namespace (e.g. granting a
|
||||
ServiceAccount in namespace A read access to Secrets in namespace B). If
|
||||
any manifest needs cross-namespace RBAC, either drop the transformer
|
||||
(safe if every resource already sets its own explicit namespace) or give
|
||||
that manifest its own Application/directory.
|
||||
|
||||
🟠 **PreSync hooks run before an Application's own normal resources are
|
||||
synced.** A PreSync Job that depends on RBAC/ServiceAccounts defined as
|
||||
plain (non-hook) resources in the *same* Application will deadlock — it
|
||||
tries to start before its own permissions exist. Use PostSync instead if
|
||||
the hook needs resources from its own Application, or move the
|
||||
prerequisite RBAC into an earlier sync-wave Application.
|
||||
|
||||
🟠 **ArgoCD hooks are not continuously reconciled by `selfHeal`.** Once a
|
||||
hook Job finishes (success, or exhausts `backoffLimit`), it's only
|
||||
deleted+recreated during an *actual new Sync operation* — not by passive
|
||||
drift detection, even with `automated.selfHeal: true` on. If you fix a
|
||||
broken hook's spec and push, the Application's `status.sync.revision` can
|
||||
show "caught up" while the live hook resource is still the old, broken
|
||||
one, because no new operation actually re-ran it. To force it: delete the
|
||||
stuck hook (clear `argocd.argoproj.io/hook-finalizer` manually if it's
|
||||
stuck `Terminating`), and if that alone doesn't trigger a fresh full sync,
|
||||
delete + re-`kubectl apply -f` the Application object itself.
|
||||
|
||||
🟠 **If you route ArgoCD's own `repoURL` through an ingress/reverse-proxy
|
||||
hostname that only listens on 80/443, don't use a non-standard port in the
|
||||
URL** — it'll silently time out trying to reach a port the proxy never
|
||||
opened, and depending on your setup this can block *every* Application's
|
||||
sync simultaneously (repo-server can't fetch git refs for anything).
|
||||
|
||||
🟠 **Don't pin exact version tags for images from registries that don't
|
||||
guarantee tag retention** (Bitnami stopped publishing versioned tags for
|
||||
free-tier images in 2025 — only `latest` + sha256 digests remain). Verify
|
||||
a tag actually exists before pinning it, or prefer minimal base images +
|
||||
a stdlib-only runtime download (e.g. Python's `urllib.request` to fetch a
|
||||
static binary) to avoid depending on any third party's tagging policy.
|
||||
|
||||
🟠 **Non-root containers can't `apk add`/`apt install` in most default
|
||||
base images** — package manager directories are root-owned. Use a
|
||||
world-writable scratch dir (`/tmp`) for anything you need to
|
||||
download/install at runtime instead.
|
||||
|
||||
🟠 **Helm does not validate unknown `values.yaml` keys.** A typo, or a
|
||||
values schema copied from the wrong chart *version's* docs/examples, is
|
||||
silently a no-op — not an error. Before concluding "this chart doesn't
|
||||
support X," clone the chart at your exact pinned version/tag and run
|
||||
`helm template` against your real values file, then diff the rendered
|
||||
output. Don't trust a chart's current `main`-branch example values file
|
||||
if you're pinned to an older release — schemas do change between major
|
||||
versions without warning in your own values file.
|
||||
|
||||
## Service Integration Routes
|
||||
|
||||
**New service? Pick your stack below:**
|
||||
**New service? Pick your stack below** (adjust doc paths to match your repo):
|
||||
|
||||
| Need | Doc | Example |
|
||||
|------|-----|---------|
|
||||
| **Authentication** | `project-usage/authentik-oidc.md` | OAuth2 login, RBAC groups, JWT tokens |
|
||||
| **Async messaging** | `project-usage/sqs-messaging.md` | Kafka topic consumers, fire-and-forget, DLQ |
|
||||
| **Async messaging** | `project-usage/sqs-messaging.md` | Queue consumers, fire-and-forget, DLQ |
|
||||
| **Object storage** | `project-usage/minio-s3.md` | File uploads, backups, log backend |
|
||||
| **CI/CD pipeline** | `project-usage/cicd-workflow.md` | GitHub Actions syntax, image push, Argo CD sync |
|
||||
| **CI/CD pipeline** | `project-usage/cicd-workflow.md` | Pipeline syntax, image push, ArgoCD sync |
|
||||
| **Workflows** | `project-usage/temporal-workflows.md` | Long-running jobs, retries, state machines |
|
||||
| **Database** | `project-usage/database-postgres.md` | CloudNativePG setup, schema migrations, replicas |
|
||||
| **Monitoring** | `project-usage/monitoring-metrics.md` | Prometheus scrape, Grafana dashboard, alerts |
|
||||
| **Secrets** | `project-usage/vault-secrets.md` | Store credentials, rotate tokens, seal/unseal |
|
||||
| **Networking** | `project-usage/networking-ingress.md` | Public HTTPS, hostname routing, TLS |
|
||||
|
||||
## Cluster Essentials
|
||||
## Cluster Essentials (fill in your own inventory)
|
||||
|
||||
**22 namespaces, 18 releases:**
|
||||
```
|
||||
Core: cert-manager, ingress-nginx, kube-system, cilium
|
||||
Storage: longhorn-system, storage (MinIO)
|
||||
Data: ddb (PostgreSQL), iam (Authentik + Vault)
|
||||
Observability: logging (Loki + Grafana), monitoring (Prometheus)
|
||||
Apps: cicd (Forgejo + Argo CD), sqs (Kafka + kmsvc), temporal, story-crater-backend
|
||||
```
|
||||
**Architecture principles (adjust to taste, but these travel well):**
|
||||
- Immutable OS (Talos, or similar — no SSH, fully declarative config)
|
||||
- Secrets in a proper secrets backend (Vault) + SOPS-encrypted manifests in
|
||||
git (`*.enc.yaml`, age-encrypted); never commit plaintext secrets or `.env`
|
||||
- ArgoCD app-of-apps as the single CD source of truth; two-phase bootstrap
|
||||
documented separately (chicken-and-egg: ArgoCD needs to exist before it
|
||||
can deploy itself declaratively — document your exact bootstrap steps)
|
||||
- Pull-based GitOps — no kubeconfig/cluster credentials ever touch your CI
|
||||
runner; the runner only needs push access to git, ArgoCD does the rest
|
||||
- Federated OIDC (one identity provider fronting every service that
|
||||
supports it)
|
||||
|
||||
**Architecture principles:**
|
||||
- Immutable OS (Talos — no SSH, declarative config)
|
||||
- Secrets in Vault + SOPS-encrypted (`*.enc.yaml`, age); never commit `.env`
|
||||
- ArgoCD app-of-apps = CD source of truth (`k8s/argocd/root` → `k8s/argocd/apps/*`); helmfile is deprecated. Two-phase bootstrap in `k8s/argocd/bootstrap/BOOTSTRAP.md`
|
||||
- Pull-based GitOps (Argo CD, no kubeconfig in CI); iterate = `git push` to Forgejo → auto-sync
|
||||
- Federated OIDC (Authentik provider for all services)
|
||||
## Deployment Checklist (per new service)
|
||||
|
||||
|
||||
## Deployment Checklist
|
||||
|
||||
- [ ] Service has Prometheus `/metrics` endpoint or ServiceMonitor
|
||||
- [ ] All credentials in Vault (never in pod env, ConfigMap, or code)
|
||||
- [ ] Ingress rule in `k8s/ingress/` with TLS cert
|
||||
- [ ] Grafana dashboard in `k8s/monitoring/dashboards/svc-<name>.yaml`
|
||||
- [ ] Alert rules in `k8s/monitoring/alerts/svc-<name>-rules.yaml` (if needed)
|
||||
- [ ] Helm release in `helmfile.yaml.gotmpl` with correct `needs:` dependencies
|
||||
|
||||
## Hard Rules
|
||||
|
||||
1. **No kubeconfig in CI** — Argo CD bridges gap (pull-based, never push secrets to runner)
|
||||
2. **Field name = variable name** — In Vault: `talos put cluster/KAFKA_BOOTSTRAP KAFKA_BOOTSTRAP="..."`
|
||||
3. **Secrets via volumes** — Never `--env` flag in pod specs (exposes in `kubectl describe`)
|
||||
4. **External services via Ingress** — All public endpoints via TLS (homelab-ca)
|
||||
5. **Never commit `.env`** — Only `.env.example` in git; real secrets in Vault
|
||||
6. **Never rename or wipe `talos-cp-1` (.213)** — sole Longhorn storage node; renaming orphans its node CR and faults every volume (permanent data loss). Rename/reprovision only the dedicated CPs.
|
||||
7. **Control-plane etcd advertises on the LAN** — keep `cluster.etcd.advertisedSubnets: ["192.168.1.0/24"]`, else Talos advertises on WireGuard and new members hang as etcd learners.
|
||||
- [ ] Prometheus `/metrics` endpoint or ServiceMonitor, if it exposes metrics
|
||||
- [ ] All credentials in your secrets backend (never in plain values.yaml,
|
||||
pod env directly, or committed anywhere in cleartext)
|
||||
- [ ] Ingress rule with TLS, if externally reachable
|
||||
- [ ] Dashboard + alert rules, if metrics are exposed
|
||||
- [ ] ArgoCD `Application` manifest added to the appropriate sync-wave file,
|
||||
**not** a standalone `helm install`/`kubectl apply` run by hand
|
||||
- [ ] Validated locally before push: `kubectl apply --dry-run=client -f`,
|
||||
`kubectl kustomize <dir>/` (if applicable), or `helm template` against
|
||||
the exact pinned chart version (if Helm-sourced)
|
||||
- [ ] After push: confirmed ArgoCD's `status.sync.revision` actually matches
|
||||
your new commit — not just that `status.sync.status` says `Synced`
|
||||
(see Gotchas — a stale hook can hide behind an otherwise-current app)
|
||||
|
||||
## Git & Release
|
||||
|
||||
**Multi-remote push:**
|
||||
```bash
|
||||
git push origin main
|
||||
```
|
||||
|
||||
**Incremental commits (service-layer grouped):**
|
||||
**Incremental commits (service-layer grouped) tend to age well:**
|
||||
- Foundation & Docs
|
||||
- Helmfile & Core Infra
|
||||
- Storage Layer
|
||||
- Core Infra (CNI, ingress, cert management, storage)
|
||||
- Observability Stack
|
||||
- IAM & Secrets
|
||||
- CI/CD & GitOps
|
||||
- Messaging Infrastructure
|
||||
- Messaging / Data Infrastructure
|
||||
- Applications & Utilities
|
||||
|
||||
Grouping by layer (rather than by day or by "misc fixes") makes it much
|
||||
easier to `git log --oneline -- <path>` your way back to *why* a given
|
||||
piece of config looks the way it does, months later.
|
||||
|
||||
Reference in New Issue
Block a user