Files
homelab/CLAUDE.md
T

11 KiB

CLAUDE.md — Homelab Project Reference

Cluster Topology (3 control-plane HA)

Node IP Zone Scheduling Storage
talos-cp-1 .213 az-a schedulable (all workloads) sole Longhorn node
talos-cp-2 .163 az-b dedicated (NoSchedule) none
talos-cp-3 .166 az-c dedicated (NoSchedule) none

3 voting etcd members peering on the LAN. Only talos-cp-1 runs workloads and holds storage → stateful services are single-instance. Full detail + gotchas in USAGE.md and memory reference_talos_etcd_and_ca_gotchas.

Deployment Model: ArgoCD GitOps (helmfile is retired)

helmfile.yaml.gotmpl and the core CLI's helmfile-era workflows described in USAGE.md/project-usage/*.md are STALE. Actual practice, 100% of the time:

git commit → git push (Forgejo) → ArgoCD auto-sync → cluster

App-of-apps structure: k8s/argocd/root/homelab-root.yaml (root Application) → k8s/argocd/apps/*.yaml (one file per "wave" of Applications, each argocd.argoproj.io/sync-wave annotated) → each Application points at either a remote Helm chart (+ a second ref: values git source for the values file) or a plain git directory of raw manifests.

Never kubectl apply/patch/delete a resource ArgoCD manages except:

  • Pure cleanup of stuck/dead state (delete a failed hook Job so the next legitimate sync creates a fresh one — not a config change, just clearing wreckage). Confirm with the user first if in doubt.
  • One-time bootstrap actions with a genuine circular dependency (Vault operator init/unseal — see k8s/security/iam/VAULT-BOOTSTRAP-README.md equivalent scripts).

Hard Rules (Never Violate)

🔴 NEVER rename or wipe talos-cp-1 (.213). It is the sole Longhorn storage node — all replicas are pinned to that node name. Renaming orphans its Longhorn node CR and faults every volume (permanent data loss). Rename/reprovision only the dedicated CPs (.163/.166), never the data node.

🔴 Control-plane etcd must advertise on the LAN. Keep cluster.etcd.advertisedSubnets: ["192.168.1.0/24"] in the controlplane template — without it Talos advertises on the WireGuard IP and new members hang as non-promoting etcd learners.

🔴 ALWAYS run terraform fmt after any terraform code changes. Before commit:

terraform fmt -recursive terraform/

Verify no changes (clean output = formatted correctly). If files change, review diffs, commit fmt changes separately. Workflow's "Terraform Format Check" step will fail otherwise.

🔴 NEVER manually kubectl delete/patch resources managed by Terraform. Terraform is the source of truth for IaC-managed resources (deployments, PVCs, services in Terraform-controlled namespaces). Manual edits create state drift. If a resource is stuck:

  1. Update Terraform config (variables.tf, *.tf files)
  2. Run terraform apply (locally or via CI/CD)
  3. Never bypass with manual kubectl operations

🔴 NEVER delete a PVC unless there are replicas or backups. A PVC deletion = permanent data loss. Verify replication status first.

🔴 NO co-authored commit messages. All commits are solo work. Never append Co-Authored-By: footer.

🔴 Commit message format: First line must capture what is added, what is fixed, what is removed, and why in one sentence. Example: fix(kubernetes): use direct in-cluster auth in providers — removes kubeconfig file dependency in CI runner. No multi-paragraph messages.

🔴 Git workflow: rebase only, no pull/merge. Always rebase when pulling. Use git pull --rebase or git rebase main before pushing. Keep history linear.

🔴 Long-running commands (>10s) must run async. Use run_in_background: true for Bash or spawn Agent. Don't actively wait. Prevents blocking on terraform plan, kubectl apply, downloads.

🔴 Infrastructure changes should flow through GitOps when possible: git commit → push → CI/CD runner (terraform apply) → ArgoCD sync. Local terraform apply is permitted (e.g. for local iteration, config regeneration, or when CI/CD isn't wired up for a given module) — still commit + push the resulting state/config afterward so git remains the record of truth. Manual kubectl apply remains disallowed for Terraform-managed resources.

GitOps / ArgoCD Gotchas (hard-won, all confirmed live in this cluster)

🟠 kustomization.yaml with an explicit resources: allowlist silently drops anything not listed — no error, no drift shown. ArgoCD reports Synced/Healthy against a manifest set that never included the missing file at all. Symptom: you commit+push a new manifest, ArgoCD says "Synced", but the resource never appears in-cluster. Fix: check every kustomization.yaml along the app's source path actually lists your new file. kubectl kustomize <dir>/ locally reproduces exactly what ArgoCD will apply — always verify with it before pushing.

🟠 A kustomization.yaml top-level namespace: transformer rewrites metadata.namespace on every resource it builds — including RBAC RoleBindings deliberately targeting a different namespace. If any manifest in that directory needs cross-namespace resources (e.g. a RoleBinding granting access to Secrets in another namespace), either drop the transformer (safe if every resource already sets its own explicit namespace) or move that manifest to its own directory/Application.

🟠 PreSync hooks run before an Application's own normal (non-hook) resources are synced. A PreSync-hooked Job that depends on a ServiceAccount/RBAC defined as plain resources in the same Application deadlocks: the Job tries to start before its own ServiceAccount exists. Confirmed live — Job sat "Running" for 14+ minutes producing zero pods, job-controller event log showed serviceaccount ... not found on every retry. Fix: use PostSync instead (runs after that app's own resources are applied), or split the hook into its own earlier-sync-wave Application.

🟠 ArgoCD hooks (PreSync/PostSync) are NOT continuously reconciled by selfHeal the way normal resources are. Once a hook Job completes (success or exhausts backoffLimit into Failed), it only gets deleted+recreated (per hook-delete-policy: BeforeHookCreation) during an actual new Sync operation — not from passive drift detection, even with automated.selfHeal: true. If you fix a hook Job's spec (image, command, RBAC) and push, status.sync.revision may show "caught up" while the live hook resource is still running the old, broken spec — because no new operation actually re-ran it. To force a real resync: delete the stuck Job (clear the argocd.argoproj.io/hook-finalizer if it's stuck Terminating), and if that alone doesn't trigger a fresh full sync, delete

  • kubectl apply -f the Application object itself (re-reads current git HEAD, starts a genuinely new operation, no cascade-delete of underlying resources since Applications don't carry a cascade finalizer by default — confirm with kubectl get app <name> -o jsonpath='{.metadata.finalizers}' first).

🟠 ArgoCD's repo-server caches rendered manifests (~120s TTL by default, argocd-cmtimeout.reconciliation). If you change a values file and Application fields both, sometimes the old rendering wins on the next sync. Restart argocd-repo-server after big multi-source/values changes if sync behavior looks stale.

🟠 ArgoCD repoURL pointing at a hostname that CoreDNS rewrites to the nginx ingress controller (for TLS termination) breaks if the URL includes a non-standard port. nginx only listens on 80/443 — http://host:3000/... silently times out (context deadline exceeded) if host resolves to the ingress controller, not the actual backend service. This blocked every single Application's sync cluster-wide simultaneously (all showed Unknown sync status) because the repo-server couldn't fetch git refs at all. Use https://host/... (no port) so nginx's default TLS cert + normal 443 routing handles it.

🟠 Bitnami Docker Hub images no longer publish versioned tags (2025 policy change) — only latest and sha256-pinned digests remain for their free tier. A pinned tag like bitnami/kubectl:1.30 will 404/ImagePullBackOff forever. Verify tags exist first: curl -s "https://hub.docker.com/v2/repositories/<org>/<image>/tags?page_size=25" | jq -r '.results[].name'. Prefer avoiding third-party utility images entirely where possible — e.g. python:3.12-alpine + stdlib urllib.request to fetch a static binary (kubectl) avoids depending on any registry's tagging policy at all.

🟠 Non-root containers (runAsNonRoot: true, non-zero UID) can't apk add in Alpine-based images — apk's working directories and most of /usr/local/bin are root-owned. Symptom: ERROR: Unable to open log: Permission denied. Use /tmp (world-writable) for any binary you need to download/install at runtime, and extend PATH rather than writing to /usr/local/bin.

🟠 Helm does not validate unknown values.yaml keys — a typo'd or wrong-schema key is silently a no-op, not an error. Confirmed root cause of a multi-week "Temporal doesn't support PostgreSQL" belief: the actual chart version pinned (temporalio/[email protected]) uses a flat server.config.persistence.<store>.driver/.sql schema, but the values file used the newer chart's datastores:-wrapped schema (introduced in a later major version) — silently ignored, so persistence stayed on the chart's Cassandra default the entire time. Before assuming "this chart doesn't support X," clone the chart at the exact pinned tag/version and run helm template with your real values — diff the rendered output, don't trust values.yaml comments/examples from the chart's current main branch, which may not match your pinned version's schema at all.

Workflow

  • ModifyFormat (terraform fmt) → Apply (locally or let CI/CD do it) → CommitPush
  • If fmt check fails in CI, fix locally, commit fmt changes, push again
  • For k8s/ManifestsChanges: Modifyvalidate (kubectl apply --dry-run=client -f, or kubectl kustomize <dir>/ if a kustomization.yaml is involved, or helm template against the exact pinned chart version for Helm-sourced Applications) → CommitPush → confirm ArgoCD picked it up (check status.sync.revision matches your commit, not just status.sync.status)

Documentation Map

  • CLAUDE.md (this file) — private, cluster-specific, always current
  • CLAUDE.example.md — sanitized, hardware-generic template for reuse on other 3-node Talos clusters; update alongside this file when a lesson is genuinely hardware/topology-generic (not homelab-specific secrets/IPs)
  • USAGE.md / project-usage/*.mdSTALE, helmfile-era. Written for a deprecated helmfile apply + core iam/core secrets CLI workflow that does not reflect actual current practice (100% ArgoCD GitOps as of this writing). Treat as historical reference only until rewritten; do not follow their deployment procedures literally.
  • TROUBLESHOOTING.md — generic Kubernetes SRE layer-before-tool methodology, still broadly applicable regardless of deployment mechanism

Last updated: 2026-07-22