Files
homelab/CLAUDE.md
T
Story Crater Bot d22842ca33 chore: track CLAUDE.md in git (was gitignored, now version-controlled)
CLAUDE.md was previously excluded from version control entirely (treated as
private local notes, with CLAUDE.example.md as the only git-tracked
counterpart). No longer justified - the file contains no secrets, just
architecture notes, private RFC1918 IPs, and operational lessons (same
sensitivity level as README.md, which is already tracked). Removing the
CLAUDE.md gitignore rule and committing it for the first time.
2026-07-21 20:17:14 -07:00

11 KiB

CLAUDE.md — Homelab Project Reference

Cluster Topology (3 control-plane HA)

Node IP Zone Scheduling Storage
talos-cp-1 .213 az-a schedulable (all workloads) sole Longhorn node
talos-cp-2 .163 az-b dedicated (NoSchedule) none
talos-cp-3 .166 az-c dedicated (NoSchedule) none

3 voting etcd members peering on the LAN. Only talos-cp-1 runs workloads and holds storage → stateful services are single-instance. Full detail + gotchas in USAGE.md and memory reference_talos_etcd_and_ca_gotchas.

Deployment Model: ArgoCD GitOps (helmfile is retired)

helmfile.yaml.gotmpl and the core CLI's helmfile-era workflows described in USAGE.md/project-usage/*.md are STALE. Actual practice, 100% of the time:

git commit → git push (Forgejo) → ArgoCD auto-sync → cluster

App-of-apps structure: k8s/argocd/root/homelab-root.yaml (root Application) → k8s/argocd/apps/*.yaml (one file per "wave" of Applications, each argocd.argoproj.io/sync-wave annotated) → each Application points at either a remote Helm chart (+ a second ref: values git source for the values file) or a plain git directory of raw manifests.

Never kubectl apply/patch/delete a resource ArgoCD manages except:

  • Pure cleanup of stuck/dead state (delete a failed hook Job so the next legitimate sync creates a fresh one — not a config change, just clearing wreckage). Confirm with the user first if in doubt.
  • One-time bootstrap actions with a genuine circular dependency (Vault operator init/unseal — see k8s/security/iam/VAULT-BOOTSTRAP-README.md equivalent scripts).

Hard Rules (Never Violate)

🔴 NEVER rename or wipe talos-cp-1 (.213). It is the sole Longhorn storage node — all replicas are pinned to that node name. Renaming orphans its Longhorn node CR and faults every volume (permanent data loss). Rename/reprovision only the dedicated CPs (.163/.166), never the data node.

🔴 Control-plane etcd must advertise on the LAN. Keep cluster.etcd.advertisedSubnets: ["192.168.1.0/24"] in the controlplane template — without it Talos advertises on the WireGuard IP and new members hang as non-promoting etcd learners.

🔴 ALWAYS run terraform fmt after any terraform code changes. Before commit:

terraform fmt -recursive terraform/

Verify no changes (clean output = formatted correctly). If files change, review diffs, commit fmt changes separately. Workflow's "Terraform Format Check" step will fail otherwise.

🔴 NEVER manually kubectl delete/patch resources managed by Terraform. Terraform is the source of truth for IaC-managed resources (deployments, PVCs, services in Terraform-controlled namespaces). Manual edits create state drift. If a resource is stuck:

  1. Update Terraform config (variables.tf, *.tf files)
  2. Run terraform apply (locally or via CI/CD)
  3. Never bypass with manual kubectl operations

🔴 NEVER delete a PVC unless there are replicas or backups. A PVC deletion = permanent data loss. Verify replication status first.

🔴 NO co-authored commit messages. All commits are solo work. Never append Co-Authored-By: footer.

🔴 Commit message format: First line must capture what is added, what is fixed, what is removed, and why in one sentence. Example: fix(kubernetes): use direct in-cluster auth in providers — removes kubeconfig file dependency in CI runner. No multi-paragraph messages.

🔴 Git workflow: rebase only, no pull/merge. Always rebase when pulling. Use git pull --rebase or git rebase main before pushing. Keep history linear.

🔴 Long-running commands (>10s) must run async. Use run_in_background: true for Bash or spawn Agent. Don't actively wait. Prevents blocking on terraform plan, kubectl apply, downloads.

🔴 Infrastructure changes should flow through GitOps when possible: git commit → push → CI/CD runner (terraform apply) → ArgoCD sync. Local terraform apply is permitted (e.g. for local iteration, config regeneration, or when CI/CD isn't wired up for a given module) — still commit + push the resulting state/config afterward so git remains the record of truth. Manual kubectl apply remains disallowed for Terraform-managed resources.

GitOps / ArgoCD Gotchas (hard-won, all confirmed live in this cluster)

🟠 kustomization.yaml with an explicit resources: allowlist silently drops anything not listed — no error, no drift shown. ArgoCD reports Synced/Healthy against a manifest set that never included the missing file at all. Symptom: you commit+push a new manifest, ArgoCD says "Synced", but the resource never appears in-cluster. Fix: check every kustomization.yaml along the app's source path actually lists your new file. kubectl kustomize <dir>/ locally reproduces exactly what ArgoCD will apply — always verify with it before pushing.

🟠 A kustomization.yaml top-level namespace: transformer rewrites metadata.namespace on every resource it builds — including RBAC RoleBindings deliberately targeting a different namespace. If any manifest in that directory needs cross-namespace resources (e.g. a RoleBinding granting access to Secrets in another namespace), either drop the transformer (safe if every resource already sets its own explicit namespace) or move that manifest to its own directory/Application.

🟠 PreSync hooks run before an Application's own normal (non-hook) resources are synced. A PreSync-hooked Job that depends on a ServiceAccount/RBAC defined as plain resources in the same Application deadlocks: the Job tries to start before its own ServiceAccount exists. Confirmed live — Job sat "Running" for 14+ minutes producing zero pods, job-controller event log showed serviceaccount ... not found on every retry. Fix: use PostSync instead (runs after that app's own resources are applied), or split the hook into its own earlier-sync-wave Application.

🟠 ArgoCD hooks (PreSync/PostSync) are NOT continuously reconciled by selfHeal the way normal resources are. Once a hook Job completes (success or exhausts backoffLimit into Failed), it only gets deleted+recreated (per hook-delete-policy: BeforeHookCreation) during an actual new Sync operation — not from passive drift detection, even with automated.selfHeal: true. If you fix a hook Job's spec (image, command, RBAC) and push, status.sync.revision may show "caught up" while the live hook resource is still running the old, broken spec — because no new operation actually re-ran it. To force a real resync: delete the stuck Job (clear the argocd.argoproj.io/hook-finalizer if it's stuck Terminating), and if that alone doesn't trigger a fresh full sync, delete

  • kubectl apply -f the Application object itself (re-reads current git HEAD, starts a genuinely new operation, no cascade-delete of underlying resources since Applications don't carry a cascade finalizer by default — confirm with kubectl get app <name> -o jsonpath='{.metadata.finalizers}' first).

🟠 ArgoCD's repo-server caches rendered manifests (~120s TTL by default, argocd-cmtimeout.reconciliation). If you change a values file and Application fields both, sometimes the old rendering wins on the next sync. Restart argocd-repo-server after big multi-source/values changes if sync behavior looks stale.

🟠 ArgoCD repoURL pointing at a hostname that CoreDNS rewrites to the nginx ingress controller (for TLS termination) breaks if the URL includes a non-standard port. nginx only listens on 80/443 — http://host:3000/... silently times out (context deadline exceeded) if host resolves to the ingress controller, not the actual backend service. This blocked every single Application's sync cluster-wide simultaneously (all showed Unknown sync status) because the repo-server couldn't fetch git refs at all. Use https://host/... (no port) so nginx's default TLS cert + normal 443 routing handles it.

🟠 Bitnami Docker Hub images no longer publish versioned tags (2025 policy change) — only latest and sha256-pinned digests remain for their free tier. A pinned tag like bitnami/kubectl:1.30 will 404/ImagePullBackOff forever. Verify tags exist first: curl -s "https://hub.docker.com/v2/repositories/<org>/<image>/tags?page_size=25" | jq -r '.results[].name'. Prefer avoiding third-party utility images entirely where possible — e.g. python:3.12-alpine + stdlib urllib.request to fetch a static binary (kubectl) avoids depending on any registry's tagging policy at all.

🟠 Non-root containers (runAsNonRoot: true, non-zero UID) can't apk add in Alpine-based images — apk's working directories and most of /usr/local/bin are root-owned. Symptom: ERROR: Unable to open log: Permission denied. Use /tmp (world-writable) for any binary you need to download/install at runtime, and extend PATH rather than writing to /usr/local/bin.

🟠 Helm does not validate unknown values.yaml keys — a typo'd or wrong-schema key is silently a no-op, not an error. Confirmed root cause of a multi-week "Temporal doesn't support PostgreSQL" belief: the actual chart version pinned (temporalio/[email protected]) uses a flat server.config.persistence.<store>.driver/.sql schema, but the values file used the newer chart's datastores:-wrapped schema (introduced in a later major version) — silently ignored, so persistence stayed on the chart's Cassandra default the entire time. Before assuming "this chart doesn't support X," clone the chart at the exact pinned tag/version and run helm template with your real values — diff the rendered output, don't trust values.yaml comments/examples from the chart's current main branch, which may not match your pinned version's schema at all.

Workflow

  • ModifyFormat (terraform fmt) → Apply (locally or let CI/CD do it) → CommitPush
  • If fmt check fails in CI, fix locally, commit fmt changes, push again
  • For k8s/ManifestsChanges: Modifyvalidate (kubectl apply --dry-run=client -f, or kubectl kustomize <dir>/ if a kustomization.yaml is involved, or helm template against the exact pinned chart version for Helm-sourced Applications) → CommitPush → confirm ArgoCD picked it up (check status.sync.revision matches your commit, not just status.sync.status)

Documentation Map

  • CLAUDE.md (this file) — private, cluster-specific, always current
  • CLAUDE.example.md — sanitized, hardware-generic template for reuse on other 3-node Talos clusters; update alongside this file when a lesson is genuinely hardware/topology-generic (not homelab-specific secrets/IPs)
  • USAGE.md / project-usage/*.mdSTALE, helmfile-era. Written for a deprecated helmfile apply + core iam/core secrets CLI workflow that does not reflect actual current practice (100% ArgoCD GitOps as of this writing). Treat as historical reference only until rewritten; do not follow their deployment procedures literally.
  • TROUBLESHOOTING.md — generic Kubernetes SRE layer-before-tool methodology, still broadly applicable regardless of deployment mechanism

Last updated: 2026-07-22