Files
homelab/CLAUDE.md
T
Story Crater Bot d22842ca33 chore: track CLAUDE.md in git (was gitignored, now version-controlled)
CLAUDE.md was previously excluded from version control entirely (treated as
private local notes, with CLAUDE.example.md as the only git-tracked
counterpart). No longer justified - the file contains no secrets, just
architecture notes, private RFC1918 IPs, and operational lessons (same
sensitivity level as README.md, which is already tracked). Removing the
CLAUDE.md gitignore rule and committing it for the first time.
2026-07-21 20:17:14 -07:00

182 lines
11 KiB
Markdown

# CLAUDE.md — Homelab Project Reference
## Cluster Topology (3 control-plane HA)
| Node | IP | Zone | Scheduling | Storage |
|------|----|----|-----------|---------|
| `talos-cp-1` | .213 | az-a | schedulable (all workloads) | sole Longhorn node |
| `talos-cp-2` | .163 | az-b | dedicated (`NoSchedule`) | none |
| `talos-cp-3` | .166 | az-c | dedicated (`NoSchedule`) | none |
3 voting etcd members peering on the LAN. Only `talos-cp-1` runs workloads and
holds storage → stateful services are single-instance. Full detail + gotchas in
`USAGE.md` and memory `reference_talos_etcd_and_ca_gotchas`.
## Deployment Model: ArgoCD GitOps (helmfile is retired)
**helmfile.yaml.gotmpl and the `core` CLI's helmfile-era workflows described in
`USAGE.md`/`project-usage/*.md` are STALE.** Actual practice, 100% of the time:
```
git commit → git push (Forgejo) → ArgoCD auto-sync → cluster
```
App-of-apps structure: `k8s/argocd/root/homelab-root.yaml` (root Application)
`k8s/argocd/apps/*.yaml` (one file per "wave" of Applications, each
`argocd.argoproj.io/sync-wave` annotated) → each Application points at either
a remote Helm chart (+ a second `ref: values` git source for the values file)
or a plain git directory of raw manifests.
**Never `kubectl apply`/`patch`/`delete` a resource ArgoCD manages** except:
- Pure cleanup of stuck/dead state (delete a failed hook Job so the *next*
legitimate sync creates a fresh one — not a config change, just clearing
wreckage). Confirm with the user first if in doubt.
- One-time bootstrap actions with a genuine circular dependency (Vault
`operator init`/unseal — see `k8s/security/iam/VAULT-BOOTSTRAP-README.md`
equivalent scripts).
## Hard Rules (Never Violate)
🔴 **NEVER rename or wipe `talos-cp-1` (.213).** It is the sole Longhorn storage
node — all replicas are pinned to that node name. Renaming orphans its Longhorn
node CR and faults every volume (permanent data loss). Rename/reprovision only
the dedicated CPs (.163/.166), never the data node.
🔴 **Control-plane etcd must advertise on the LAN.** Keep
`cluster.etcd.advertisedSubnets: ["192.168.1.0/24"]` in the controlplane
template — without it Talos advertises on the WireGuard IP and new members hang
as non-promoting etcd learners.
🔴 **ALWAYS run `terraform fmt` after any terraform code changes.** Before commit:
```bash
terraform fmt -recursive terraform/
```
Verify no changes (clean output = formatted correctly). If files change, review diffs, commit fmt changes separately. Workflow's "Terraform Format Check" step will fail otherwise.
🔴 **NEVER manually kubectl delete/patch resources managed by Terraform.** Terraform is the source of truth for IaC-managed resources (deployments, PVCs, services in Terraform-controlled namespaces). Manual edits create state drift. If a resource is stuck:
1. Update Terraform config (variables.tf, *.tf files)
2. Run `terraform apply` (locally or via CI/CD)
3. Never bypass with manual kubectl operations
🔴 **NEVER delete a PVC unless there are replicas or backups.** A PVC deletion = permanent data loss. Verify replication status first.
🔴 **NO co-authored commit messages.** All commits are solo work. Never append `Co-Authored-By:` footer.
🔴 **Commit message format:** First line must capture what is added, what is fixed, what is removed, and why in one sentence. Example: `fix(kubernetes): use direct in-cluster auth in providers — removes kubeconfig file dependency in CI runner`. No multi-paragraph messages.
🔴 **Git workflow: rebase only, no pull/merge.** Always rebase when pulling. Use `git pull --rebase` or `git rebase main` before pushing. Keep history linear.
🔴 **Long-running commands (>10s) must run async.** Use `run_in_background: true` for Bash or spawn Agent. Don't actively wait. Prevents blocking on terraform plan, kubectl apply, downloads.
🔴 **Infrastructure changes should flow through GitOps when possible:** git commit → push → CI/CD runner (terraform apply) → ArgoCD sync. Local `terraform apply` is permitted (e.g. for local iteration, config regeneration, or when CI/CD isn't wired up for a given module) — still commit + push the resulting state/config afterward so git remains the record of truth. Manual `kubectl apply` remains disallowed for Terraform-managed resources.
## GitOps / ArgoCD Gotchas (hard-won, all confirmed live in this cluster)
🟠 **`kustomization.yaml` with an explicit `resources:` allowlist silently
drops anything not listed — no error, no drift shown.** ArgoCD reports
`Synced/Healthy` against a manifest set that never included the missing
file at all. Symptom: you commit+push a new manifest, ArgoCD says
"Synced", but the resource never appears in-cluster. Fix: check every
`kustomization.yaml` along the app's source path actually lists your new
file. `kubectl kustomize <dir>/` locally reproduces exactly what ArgoCD
will apply — always verify with it before pushing.
🟠 **A `kustomization.yaml` top-level `namespace:` transformer rewrites
`metadata.namespace` on *every* resource it builds — including RBAC
RoleBindings deliberately targeting a *different* namespace.** If any
manifest in that directory needs cross-namespace resources (e.g. a
RoleBinding granting access to Secrets in another namespace), either drop
the transformer (safe if every resource already sets its own explicit
namespace) or move that manifest to its own directory/Application.
🟠 **PreSync hooks run *before* an Application's own normal (non-hook)
resources are synced.** A PreSync-hooked Job that depends on a
ServiceAccount/RBAC defined as plain resources in the *same* Application
deadlocks: the Job tries to start before its own ServiceAccount exists.
Confirmed live — Job sat "Running" for 14+ minutes producing zero pods,
`job-controller` event log showed `serviceaccount ... not found` on every
retry. Fix: use PostSync instead (runs after that app's own resources are
applied), or split the hook into its own earlier-sync-wave Application.
🟠 **ArgoCD hooks (PreSync/PostSync) are NOT continuously reconciled by
`selfHeal` the way normal resources are.** Once a hook Job completes
(success or exhausts `backoffLimit` into Failed), it only gets
deleted+recreated (per `hook-delete-policy: BeforeHookCreation`) during an
**actual new Sync operation** — not from passive drift detection, even
with `automated.selfHeal: true`. If you fix a hook Job's spec (image,
command, RBAC) and push, `status.sync.revision` may show "caught up" while
the *live* hook resource is still running the old, broken spec — because
no new operation actually re-ran it. To force a real resync: delete the
stuck Job (clear the `argocd.argoproj.io/hook-finalizer` if it's stuck
`Terminating`), and if that alone doesn't trigger a fresh full sync, delete
+ `kubectl apply -f` the Application object itself (re-reads current git
HEAD, starts a genuinely new operation, no cascade-delete of underlying
resources since Applications don't carry a cascade finalizer by default —
confirm with `kubectl get app <name> -o jsonpath='{.metadata.finalizers}'`
first).
🟠 **ArgoCD's repo-server caches rendered manifests (~120s TTL by default,
`argocd-cm``timeout.reconciliation`).** If you change a values file and
`Application` fields both, sometimes the old rendering wins on the next
sync. Restart `argocd-repo-server` after big multi-source/values changes if
sync behavior looks stale.
🟠 **ArgoCD `repoURL` pointing at a hostname that CoreDNS rewrites to the
nginx ingress controller (for TLS termination) breaks if the URL includes
a non-standard port.** nginx only listens on 80/443 — `http://host:3000/...`
silently times out (`context deadline exceeded`) if `host` resolves to the
ingress controller, not the actual backend service. This blocked **every
single Application's sync** cluster-wide simultaneously (all showed
`Unknown` sync status) because the repo-server couldn't fetch git refs at
all. Use `https://host/...` (no port) so nginx's default TLS cert + normal
443 routing handles it.
🟠 **Bitnami Docker Hub images no longer publish versioned tags (2025
policy change) — only `latest` and sha256-pinned digests remain for their
free tier.** A pinned tag like `bitnami/kubectl:1.30` will 404/ImagePullBackOff
forever. Verify tags exist first: `curl -s "https://hub.docker.com/v2/repositories/<org>/<image>/tags?page_size=25" | jq -r '.results[].name'`.
Prefer avoiding third-party utility images entirely where possible — e.g.
`python:3.12-alpine` + stdlib `urllib.request` to fetch a static binary
(kubectl) avoids depending on any registry's tagging policy at all.
🟠 **Non-root containers (`runAsNonRoot: true`, non-zero UID) can't `apk
add` in Alpine-based images** — apk's working directories and most of
`/usr/local/bin` are root-owned. Symptom: `ERROR: Unable to open log:
Permission denied`. Use `/tmp` (world-writable) for any binary you need to
download/install at runtime, and extend `PATH` rather than writing to
`/usr/local/bin`.
🟠 **Helm does not validate unknown `values.yaml` keys — a typo'd or
wrong-schema key is silently a no-op, not an error.** Confirmed root cause
of a multi-week "Temporal doesn't support PostgreSQL" belief: the actual
chart version pinned (`temporalio/[email protected]`) uses a flat
`server.config.persistence.<store>.driver/.sql` schema, but the values file
used the *newer* chart's `datastores:`-wrapped schema (introduced in a
later major version) — silently ignored, so persistence stayed on the
chart's Cassandra default the entire time. **Before assuming "this chart
doesn't support X," clone the chart at the exact pinned tag/version and run
`helm template` with your real values — diff the rendered output, don't
trust values.yaml comments/examples from the chart's current `main`
branch, which may not match your pinned version's schema at all.**
## Workflow
- **Modify** → **Format** (`terraform fmt`) → **Apply** (locally or let CI/CD do it) → **Commit****Push**
- If fmt check fails in CI, fix locally, commit fmt changes, push again
- For k8s/ManifestsChanges: **Modify****validate** (`kubectl apply --dry-run=client -f`, or `kubectl kustomize <dir>/` if a `kustomization.yaml` is involved, or `helm template` against the exact pinned chart version for Helm-sourced Applications) → **Commit****Push** → confirm ArgoCD picked it up (check `status.sync.revision` matches your commit, not just `status.sync.status`)
## Documentation Map
- `CLAUDE.md` (this file) — private, cluster-specific, always current
- `CLAUDE.example.md` — sanitized, hardware-generic template for reuse on other 3-node Talos clusters; update alongside this file when a lesson is genuinely hardware/topology-generic (not homelab-specific secrets/IPs)
- `USAGE.md` / `project-usage/*.md`**STALE, helmfile-era.** Written for a
deprecated `helmfile apply` + `core iam`/`core secrets` CLI workflow that
does not reflect actual current practice (100% ArgoCD GitOps as of this
writing). Treat as historical reference only until rewritten; do not
follow their deployment procedures literally.
- `TROUBLESHOOTING.md` — generic Kubernetes SRE layer-before-tool methodology, still broadly applicable regardless of deployment mechanism
---
**Last updated:** 2026-07-22