CLAUDE.md was previously excluded from version control entirely (treated as private local notes, with CLAUDE.example.md as the only git-tracked counterpart). No longer justified - the file contains no secrets, just architecture notes, private RFC1918 IPs, and operational lessons (same sensitivity level as README.md, which is already tracked). Removing the CLAUDE.md gitignore rule and committing it for the first time.
11 KiB
CLAUDE.md — Homelab Project Reference
Cluster Topology (3 control-plane HA)
| Node | IP | Zone | Scheduling | Storage |
|---|---|---|---|---|
talos-cp-1 |
.213 | az-a | schedulable (all workloads) | sole Longhorn node |
talos-cp-2 |
.163 | az-b | dedicated (NoSchedule) |
none |
talos-cp-3 |
.166 | az-c | dedicated (NoSchedule) |
none |
3 voting etcd members peering on the LAN. Only talos-cp-1 runs workloads and
holds storage → stateful services are single-instance. Full detail + gotchas in
USAGE.md and memory reference_talos_etcd_and_ca_gotchas.
Deployment Model: ArgoCD GitOps (helmfile is retired)
helmfile.yaml.gotmpl and the core CLI's helmfile-era workflows described in
USAGE.md/project-usage/*.md are STALE. Actual practice, 100% of the time:
git commit → git push (Forgejo) → ArgoCD auto-sync → cluster
App-of-apps structure: k8s/argocd/root/homelab-root.yaml (root Application)
→ k8s/argocd/apps/*.yaml (one file per "wave" of Applications, each
argocd.argoproj.io/sync-wave annotated) → each Application points at either
a remote Helm chart (+ a second ref: values git source for the values file)
or a plain git directory of raw manifests.
Never kubectl apply/patch/delete a resource ArgoCD manages except:
- Pure cleanup of stuck/dead state (delete a failed hook Job so the next legitimate sync creates a fresh one — not a config change, just clearing wreckage). Confirm with the user first if in doubt.
- One-time bootstrap actions with a genuine circular dependency (Vault
operator init/unseal — seek8s/security/iam/VAULT-BOOTSTRAP-README.mdequivalent scripts).
Hard Rules (Never Violate)
🔴 NEVER rename or wipe talos-cp-1 (.213). It is the sole Longhorn storage
node — all replicas are pinned to that node name. Renaming orphans its Longhorn
node CR and faults every volume (permanent data loss). Rename/reprovision only
the dedicated CPs (.163/.166), never the data node.
🔴 Control-plane etcd must advertise on the LAN. Keep
cluster.etcd.advertisedSubnets: ["192.168.1.0/24"] in the controlplane
template — without it Talos advertises on the WireGuard IP and new members hang
as non-promoting etcd learners.
🔴 ALWAYS run terraform fmt after any terraform code changes. Before commit:
terraform fmt -recursive terraform/
Verify no changes (clean output = formatted correctly). If files change, review diffs, commit fmt changes separately. Workflow's "Terraform Format Check" step will fail otherwise.
🔴 NEVER manually kubectl delete/patch resources managed by Terraform. Terraform is the source of truth for IaC-managed resources (deployments, PVCs, services in Terraform-controlled namespaces). Manual edits create state drift. If a resource is stuck:
- Update Terraform config (variables.tf, *.tf files)
- Run
terraform apply(locally or via CI/CD) - Never bypass with manual kubectl operations
🔴 NEVER delete a PVC unless there are replicas or backups. A PVC deletion = permanent data loss. Verify replication status first.
🔴 NO co-authored commit messages. All commits are solo work. Never append Co-Authored-By: footer.
🔴 Commit message format: First line must capture what is added, what is fixed, what is removed, and why in one sentence. Example: fix(kubernetes): use direct in-cluster auth in providers — removes kubeconfig file dependency in CI runner. No multi-paragraph messages.
🔴 Git workflow: rebase only, no pull/merge. Always rebase when pulling. Use git pull --rebase or git rebase main before pushing. Keep history linear.
🔴 Long-running commands (>10s) must run async. Use run_in_background: true for Bash or spawn Agent. Don't actively wait. Prevents blocking on terraform plan, kubectl apply, downloads.
🔴 Infrastructure changes should flow through GitOps when possible: git commit → push → CI/CD runner (terraform apply) → ArgoCD sync. Local terraform apply is permitted (e.g. for local iteration, config regeneration, or when CI/CD isn't wired up for a given module) — still commit + push the resulting state/config afterward so git remains the record of truth. Manual kubectl apply remains disallowed for Terraform-managed resources.
GitOps / ArgoCD Gotchas (hard-won, all confirmed live in this cluster)
🟠 kustomization.yaml with an explicit resources: allowlist silently
drops anything not listed — no error, no drift shown. ArgoCD reports
Synced/Healthy against a manifest set that never included the missing
file at all. Symptom: you commit+push a new manifest, ArgoCD says
"Synced", but the resource never appears in-cluster. Fix: check every
kustomization.yaml along the app's source path actually lists your new
file. kubectl kustomize <dir>/ locally reproduces exactly what ArgoCD
will apply — always verify with it before pushing.
🟠 A kustomization.yaml top-level namespace: transformer rewrites
metadata.namespace on every resource it builds — including RBAC
RoleBindings deliberately targeting a different namespace. If any
manifest in that directory needs cross-namespace resources (e.g. a
RoleBinding granting access to Secrets in another namespace), either drop
the transformer (safe if every resource already sets its own explicit
namespace) or move that manifest to its own directory/Application.
🟠 PreSync hooks run before an Application's own normal (non-hook)
resources are synced. A PreSync-hooked Job that depends on a
ServiceAccount/RBAC defined as plain resources in the same Application
deadlocks: the Job tries to start before its own ServiceAccount exists.
Confirmed live — Job sat "Running" for 14+ minutes producing zero pods,
job-controller event log showed serviceaccount ... not found on every
retry. Fix: use PostSync instead (runs after that app's own resources are
applied), or split the hook into its own earlier-sync-wave Application.
🟠 ArgoCD hooks (PreSync/PostSync) are NOT continuously reconciled by
selfHeal the way normal resources are. Once a hook Job completes
(success or exhausts backoffLimit into Failed), it only gets
deleted+recreated (per hook-delete-policy: BeforeHookCreation) during an
actual new Sync operation — not from passive drift detection, even
with automated.selfHeal: true. If you fix a hook Job's spec (image,
command, RBAC) and push, status.sync.revision may show "caught up" while
the live hook resource is still running the old, broken spec — because
no new operation actually re-ran it. To force a real resync: delete the
stuck Job (clear the argocd.argoproj.io/hook-finalizer if it's stuck
Terminating), and if that alone doesn't trigger a fresh full sync, delete
kubectl apply -fthe Application object itself (re-reads current git HEAD, starts a genuinely new operation, no cascade-delete of underlying resources since Applications don't carry a cascade finalizer by default — confirm withkubectl get app <name> -o jsonpath='{.metadata.finalizers}'first).
🟠 ArgoCD's repo-server caches rendered manifests (~120s TTL by default,
argocd-cm → timeout.reconciliation). If you change a values file and
Application fields both, sometimes the old rendering wins on the next
sync. Restart argocd-repo-server after big multi-source/values changes if
sync behavior looks stale.
🟠 ArgoCD repoURL pointing at a hostname that CoreDNS rewrites to the
nginx ingress controller (for TLS termination) breaks if the URL includes
a non-standard port. nginx only listens on 80/443 — http://host:3000/...
silently times out (context deadline exceeded) if host resolves to the
ingress controller, not the actual backend service. This blocked every
single Application's sync cluster-wide simultaneously (all showed
Unknown sync status) because the repo-server couldn't fetch git refs at
all. Use https://host/... (no port) so nginx's default TLS cert + normal
443 routing handles it.
🟠 Bitnami Docker Hub images no longer publish versioned tags (2025
policy change) — only latest and sha256-pinned digests remain for their
free tier. A pinned tag like bitnami/kubectl:1.30 will 404/ImagePullBackOff
forever. Verify tags exist first: curl -s "https://hub.docker.com/v2/repositories/<org>/<image>/tags?page_size=25" | jq -r '.results[].name'.
Prefer avoiding third-party utility images entirely where possible — e.g.
python:3.12-alpine + stdlib urllib.request to fetch a static binary
(kubectl) avoids depending on any registry's tagging policy at all.
🟠 Non-root containers (runAsNonRoot: true, non-zero UID) can't apk add in Alpine-based images — apk's working directories and most of
/usr/local/bin are root-owned. Symptom: ERROR: Unable to open log: Permission denied. Use /tmp (world-writable) for any binary you need to
download/install at runtime, and extend PATH rather than writing to
/usr/local/bin.
🟠 Helm does not validate unknown values.yaml keys — a typo'd or
wrong-schema key is silently a no-op, not an error. Confirmed root cause
of a multi-week "Temporal doesn't support PostgreSQL" belief: the actual
chart version pinned (temporalio/[email protected]) uses a flat
server.config.persistence.<store>.driver/.sql schema, but the values file
used the newer chart's datastores:-wrapped schema (introduced in a
later major version) — silently ignored, so persistence stayed on the
chart's Cassandra default the entire time. Before assuming "this chart
doesn't support X," clone the chart at the exact pinned tag/version and run
helm template with your real values — diff the rendered output, don't
trust values.yaml comments/examples from the chart's current main
branch, which may not match your pinned version's schema at all.
Workflow
- Modify → Format (
terraform fmt) → Apply (locally or let CI/CD do it) → Commit → Push - If fmt check fails in CI, fix locally, commit fmt changes, push again
- For k8s/ManifestsChanges: Modify → validate (
kubectl apply --dry-run=client -f, orkubectl kustomize <dir>/if akustomization.yamlis involved, orhelm templateagainst the exact pinned chart version for Helm-sourced Applications) → Commit → Push → confirm ArgoCD picked it up (checkstatus.sync.revisionmatches your commit, not juststatus.sync.status)
Documentation Map
CLAUDE.md(this file) — private, cluster-specific, always currentCLAUDE.example.md— sanitized, hardware-generic template for reuse on other 3-node Talos clusters; update alongside this file when a lesson is genuinely hardware/topology-generic (not homelab-specific secrets/IPs)USAGE.md/project-usage/*.md— STALE, helmfile-era. Written for a deprecatedhelmfile apply+core iam/core secretsCLI workflow that does not reflect actual current practice (100% ArgoCD GitOps as of this writing). Treat as historical reference only until rewritten; do not follow their deployment procedures literally.TROUBLESHOOTING.md— generic Kubernetes SRE layer-before-tool methodology, still broadly applicable regardless of deployment mechanism
Last updated: 2026-07-22