13 KiB
CLAUDE.md — Homelab Project Reference
Cluster Topology (3 control-plane HA)
| Node | IP | Zone | Scheduling | Storage |
|---|---|---|---|---|
talos-cp-1 |
.213 | az-a | schedulable (all workloads) | Longhorn replica |
talos-cp-2 |
.163 | az-b | dedicated (NoSchedule) |
Longhorn replica |
talos-cp-3 |
.166 | az-c | dedicated (NoSchedule) |
Longhorn replica |
3 voting etcd members peering on the LAN. All 3 nodes run Longhorn with 3-replica
HA storage (verified: all 17 PVCs have replicas across all nodes). Only talos-cp-1
is schedulable for workloads (cp-2/cp-3 are control-plane only). Full detail + gotchas
in USAGE.md and memory reference_talos_etcd_and_ca_gotchas.
Deployment Model: ArgoCD GitOps (helmfile is retired)
helmfile.yaml.gotmpl and the core CLI's helmfile-era workflows described in
USAGE.md/project-usage/*.md are STALE. Actual practice, 100% of the time:
git commit → git push (Forgejo) → ArgoCD auto-sync → cluster
App-of-apps structure: k8s/argocd/root/homelab-root.yaml (root Application)
→ k8s/argocd/apps/*.yaml (one file per "wave" of Applications, each
argocd.argoproj.io/sync-wave annotated) → each Application points at either
a remote Helm chart (+ a second ref: values git source for the values file)
or a plain git directory of raw manifests.
Never kubectl apply/patch/delete a resource ArgoCD manages except:
- Pure cleanup of stuck/dead state (delete a failed hook Job so the next legitimate sync creates a fresh one — not a config change, just clearing wreckage). Confirm with the user first if in doubt.
- One-time bootstrap actions with a genuine circular dependency (Vault
operator init/unseal — seek8s/security/iam/VAULT-BOOTSTRAP-README.mdequivalent scripts).
Hard Rules (Never Violate)
🔴 NEVER rename or wipe talos-cp-1 (.213). It is the sole Longhorn storage
node — all replicas are pinned to that node name. Renaming orphans its Longhorn
node CR and faults every volume (permanent data loss). Rename/reprovision only
the dedicated CPs (.163/.166), never the data node.
🔴 Control-plane etcd must advertise on the LAN. Keep
cluster.etcd.advertisedSubnets: ["192.168.1.0/24"] in the controlplane
template — without it Talos advertises on the WireGuard IP and new members hang
as non-promoting etcd learners.
🔴 ALWAYS run terraform fmt after any terraform code changes. Before commit:
terraform fmt -recursive terraform/
Verify no changes (clean output = formatted correctly). If files change, review diffs, commit fmt changes separately. Workflow's "Terraform Format Check" step will fail otherwise.
🔴 NEVER manually kubectl delete/patch resources managed by Terraform. Terraform is the source of truth for IaC-managed resources (deployments, PVCs, services in Terraform-controlled namespaces). Manual edits create state drift. If a resource is stuck:
- Update Terraform config (variables.tf, *.tf files)
- Run
terraform apply(locally or via CI/CD) - Never bypass with manual kubectl operations
🔴 NEVER delete a PVC unless there are replicas or backups. A PVC deletion = permanent data loss. Verify replication status first.
🔴 NO co-authored commit messages. All commits are solo work. Never append Co-Authored-By: footer.
🔴 Commit message format: First line must capture what is added, what is fixed, what is removed, and why in one sentence. Example: fix(kubernetes): use direct in-cluster auth in providers — removes kubeconfig file dependency in CI runner. No multi-paragraph messages.
🔴 Git workflow: rebase only, no pull/merge. Always rebase when pulling. Use git pull --rebase or git rebase main before pushing. Keep history linear.
🔴 Long-running commands (>10s) must run async. Use run_in_background: true for Bash or spawn Agent. Don't actively wait. Prevents blocking on terraform plan, kubectl apply, downloads.
🔴 Infrastructure changes should flow through GitOps when possible: git commit → push → CI/CD runner (terraform apply) → ArgoCD sync. Local terraform apply is permitted (e.g. for local iteration, config regeneration, or when CI/CD isn't wired up for a given module) — still commit + push the resulting state/config afterward so git remains the record of truth. Manual kubectl apply remains disallowed for Terraform-managed resources.
CloudNativePG (CNPG) Database Pattern
Simple ownership model — all apps use shared 'app' user:
# CNPG Cluster (bootstrap, wave 0)
bootstrap:
initdb:
database: app # Bootstrap database
owner: app # Bootstrap user (owns all databases)
# Database CR (per-app, wave 6)
spec:
name: authentik # Database name
owner: app # ← All apps use 'app' (not per-app roles)
cluster:
name: ddb-cluster
Credential Distribution (bootstrap.sh pattern):
- Source of truth: CNPG creates
ddb-cluster-appsecret inddbnamespace - Distribution:
bootstrap.shcopies secret to app namespaces:# For each app namespace (cicd, iam, etc): kubectl get secret ddb-cluster-app -n ddb -o yaml \ | sed 's/namespace: ddb/namespace: <app-namespace>/' \ | kubectl apply -f - - Apps reference local copy:
env: - name: DB_USER valueFrom: secretKeyRef: name: ddb-cluster-app # Local copy in app's namespace key: username # Always "app" - name: DB_PASSWORD valueFrom: secretKeyRef: name: ddb-cluster-app key: password
Key Points:
✅ DO:
- All apps connect as
appuser - Database CRs specify
owner: app - Isolation via separate database names (not roles)
- Copy credentials to app namespace via
bootstrap.sh - Reference local secret copy via
secretKeyRef
❌ DON'T:
- Create per-app roles in
managed.roles(CNPG doesn't transfer ownership properly) - Grant permissions via PostSync Jobs (owner already has full rights)
- Use cross-namespace
secretKeyRef(not supported) - Manually patch secrets (ArgoCD will revert)
Working Examples:
- Forgejo:
k8s/bootstrap-local/04-forgejo.yaml(referencesddb-cluster-appin cicd namespace) - Authentik:
k8s/security/iam/authentik-values.yaml(referencesddb-cluster-appin iam namespace)
For New Apps:
- Add Database CR to
k8s/data/schemas/withowner: app - Add secret copy to
bootstrap.sh(like cicd/iam examples) - Reference
ddb-cluster-appviasecretKeyRefin app's namespace - No permission grants needed — app user owns the database
GitOps / ArgoCD Gotchas (hard-won, all confirmed live in this cluster)
🟠 kustomization.yaml with an explicit resources: allowlist silently
drops anything not listed — no error, no drift shown. ArgoCD reports
Synced/Healthy against a manifest set that never included the missing
file at all. Symptom: you commit+push a new manifest, ArgoCD says
"Synced", but the resource never appears in-cluster. Fix: check every
kustomization.yaml along the app's source path actually lists your new
file. kubectl kustomize <dir>/ locally reproduces exactly what ArgoCD
will apply — always verify with it before pushing.
🟠 A kustomization.yaml top-level namespace: transformer rewrites
metadata.namespace on every resource it builds — including RBAC
RoleBindings deliberately targeting a different namespace. If any
manifest in that directory needs cross-namespace resources (e.g. a
RoleBinding granting access to Secrets in another namespace), either drop
the transformer (safe if every resource already sets its own explicit
namespace) or move that manifest to its own directory/Application.
🟠 PreSync hooks run before an Application's own normal (non-hook)
resources are synced. A PreSync-hooked Job that depends on a
ServiceAccount/RBAC defined as plain resources in the same Application
deadlocks: the Job tries to start before its own ServiceAccount exists.
Confirmed live — Job sat "Running" for 14+ minutes producing zero pods,
job-controller event log showed serviceaccount ... not found on every
retry. Fix: use PostSync instead (runs after that app's own resources are
applied), or split the hook into its own earlier-sync-wave Application.
🟠 ArgoCD hooks (PreSync/PostSync) are NOT continuously reconciled by
selfHeal the way normal resources are. Once a hook Job completes
(success or exhausts backoffLimit into Failed), it only gets
deleted+recreated (per hook-delete-policy: BeforeHookCreation) during an
actual new Sync operation — not from passive drift detection, even
with automated.selfHeal: true. If you fix a hook Job's spec (image,
command, RBAC) and push, status.sync.revision may show "caught up" while
the live hook resource is still running the old, broken spec — because
no new operation actually re-ran it. To force a real resync: delete the
stuck Job (clear the argocd.argoproj.io/hook-finalizer if it's stuck
Terminating), and if that alone doesn't trigger a fresh full sync, delete
kubectl apply -fthe Application object itself (re-reads current git HEAD, starts a genuinely new operation, no cascade-delete of underlying resources since Applications don't carry a cascade finalizer by default — confirm withkubectl get app <name> -o jsonpath='{.metadata.finalizers}'first).
🟠 ArgoCD's repo-server caches rendered manifests (~120s TTL by default,
argocd-cm → timeout.reconciliation). If you change a values file and
Application fields both, sometimes the old rendering wins on the next
sync. Restart argocd-repo-server after big multi-source/values changes if
sync behavior looks stale.
🟠 ArgoCD repoURL pointing at a hostname that CoreDNS rewrites to the
nginx ingress controller (for TLS termination) breaks if the URL includes
a non-standard port. nginx only listens on 80/443 — http://host:3000/...
silently times out (context deadline exceeded) if host resolves to the
ingress controller, not the actual backend service. This blocked every
single Application's sync cluster-wide simultaneously (all showed
Unknown sync status) because the repo-server couldn't fetch git refs at
all. Use https://host/... (no port) so nginx's default TLS cert + normal
443 routing handles it.
🟠 Bitnami Docker Hub images no longer publish versioned tags (2025
policy change) — only latest and sha256-pinned digests remain for their
free tier. A pinned tag like bitnami/kubectl:1.30 will 404/ImagePullBackOff
forever. Verify tags exist first: curl -s "https://hub.docker.com/v2/repositories/<org>/<image>/tags?page_size=25" | jq -r '.results[].name'.
Prefer avoiding third-party utility images entirely where possible — e.g.
python:3.12-alpine + stdlib urllib.request to fetch a static binary
(kubectl) avoids depending on any registry's tagging policy at all.
🟠 Non-root containers (runAsNonRoot: true, non-zero UID) can't apk add in Alpine-based images — apk's working directories and most of
/usr/local/bin are root-owned. Symptom: ERROR: Unable to open log: Permission denied. Use /tmp (world-writable) for any binary you need to
download/install at runtime, and extend PATH rather than writing to
/usr/local/bin.
🟠 Helm does not validate unknown values.yaml keys — a typo'd or
wrong-schema key is silently a no-op, not an error. Confirmed root cause
of a multi-week "Temporal doesn't support PostgreSQL" belief: the actual
chart version pinned (temporalio/[email protected]) uses a flat
server.config.persistence.<store>.driver/.sql schema, but the values file
used the newer chart's datastores:-wrapped schema (introduced in a
later major version) — silently ignored, so persistence stayed on the
chart's Cassandra default the entire time. Before assuming "this chart
doesn't support X," clone the chart at the exact pinned tag/version and run
helm template with your real values — diff the rendered output, don't
trust values.yaml comments/examples from the chart's current main
branch, which may not match your pinned version's schema at all.
Workflow
- Modify → Format (
terraform fmt) → Apply (locally or let CI/CD do it) → Commit → Push - If fmt check fails in CI, fix locally, commit fmt changes, push again
- For k8s/ManifestsChanges: Modify → validate (
kubectl apply --dry-run=client -f, orkubectl kustomize <dir>/if akustomization.yamlis involved, orhelm templateagainst the exact pinned chart version for Helm-sourced Applications) → Commit → Push → confirm ArgoCD picked it up (checkstatus.sync.revisionmatches your commit, not juststatus.sync.status)
Documentation Map
CLAUDE.md(this file) — private, cluster-specific, always currentCLAUDE.example.md— sanitized, hardware-generic template for reuse on other 3-node Talos clusters; update alongside this file when a lesson is genuinely hardware/topology-generic (not homelab-specific secrets/IPs)USAGE.md/project-usage/*.md— STALE, helmfile-era. Written for a deprecatedhelmfile apply+core iam/core secretsCLI workflow that does not reflect actual current practice (100% ArgoCD GitOps as of this writing). Treat as historical reference only until rewritten; do not follow their deployment procedures literally.TROUBLESHOOTING.md— generic Kubernetes SRE layer-before-tool methodology, still broadly applicable regardless of deployment mechanism
Last updated: 2026-07-23