Files

211 lines
11 KiB
Markdown
Raw Permalink Normal View History

# CLAUDE.md — Homelab Integration Guide (Example / Reusable Template)
> **This is a sanitized template.** Copy to `CLAUDE.md`, fill in your own
> node IPs/hostnames/secrets, and delete this notice. Nothing in this file
> should contain real credentials, real IPs beyond illustrative examples, or
> anything that would matter if this file became public. It's meant to be
> shared across homelabs running similar hardware/topology (3-node bare-metal
> Talos Kubernetes + ArgoCD GitOps), not just this one.
**Homelab:** A bare-metal N-node Kubernetes cluster running Talos Linux with a
full observability stack, SSO via Authentik, secret management via Vault, and
CI/CD infrastructure (self-hosted git forge + Argo CD).
## Cluster Topology (adjust to your hardware)
| Node | IP | Zone | Scheduling | Storage |
|------|----|----|-----------|---------|
| `<node-1>` | `<ip>` | az-a | schedulable (all workloads) | sole storage node (if single-node storage) |
| `<node-2>` | `<ip>` | az-b | dedicated (`NoSchedule`) | none |
| `<node-3>` | `<ip>` | az-c | dedicated (`NoSchedule`) | none |
If your storage layer (Longhorn, local-path, etc.) only runs on one node,
every stateful workload pinned to that storage class is effectively
single-instance regardless of your control-plane HA count — document that
explicitly here, it changes your failure-mode assumptions everywhere else.
## Deployment Model: ArgoCD GitOps (app-of-apps)
```
git commit → git push (your git forge) → ArgoCD auto-sync → cluster
```
Structure:
- One root `Application` (`k8s/argocd/root/`) pointing at a directory of
child `Application` manifests (`k8s/argocd/apps/*.yaml`)
- Each child Application is either:
- A remote Helm chart + a **second** git source (`ref: values`) supplying
just the values file — lets you pin an upstream chart version while
keeping your values under normal git history/review
- A plain git directory of raw manifests (optionally with a
`kustomization.yaml`)
- `argocd.argoproj.io/sync-wave` annotations control ordering across
Applications (lower number syncs first)
**Never `kubectl apply`/`patch`/`delete` a resource ArgoCD manages**, except:
- Pure cleanup of stuck/dead state (e.g. deleting a permanently-failed hook
Job so the next real sync can create a fresh one) — this is not a config
change, just clearing wreckage that GitOps itself won't clean up
automatically (see Gotchas below)
- Genuine one-time bootstrap circular dependencies (e.g. Vault
`operator init`/unseal — nothing can configure Vault's own unseal keys
before Vault has generated them)
## Hard Rules (adapt freely, but keep something like these)
🔴 **Identify and document your storage/stateful-singleton node explicitly.**
Whatever node holds your CSI driver's data (Longhorn, local-path, etc.),
renaming or wiping it orphans every PVC pinned there. Name it here, in
caps, so nobody "cleans up" it by accident.
🔴 **If you run multi-member etcd across a LAN + VPN/WireGuard overlay,
pin the advertised subnet explicitly** (e.g. Talos's
`cluster.etcd.advertisedSubnets`). Without it, etcd may advertise on the
wrong interface and new members hang as non-promoting learners.
🔴 **Run your IaC formatter (terraform fmt, etc.) before every commit
that touches infra code.** Wire this into CI as a hard gate, not a
suggestion.
🔴 **Whatever your source of truth is (Terraform, ArgoCD, both) — never
manually mutate resources it manages.** State drift is the single most
common cause of "why did my last apply undo my manual fix" confusion.
Fix the source, re-apply/re-sync, never bypass.
🔴 **Never delete a PVC without confirming replica count / backup
freshness first.** This is always a one-way door.
🔴 **Decide your commit message convention up front and enforce it.**
(This template's origin project uses: no co-authored-by footers, single-line
commit summarizing what/why, solo-authorship assumption — adjust to your
team's norms.)
🔴 **Decide your git workflow (rebase vs merge) up front and stick to it**
cluster-wide, across every contributor/agent working in the repo.
🔴 **Long-running commands should not block a synchronous session** — run
them in the background and poll, especially anything that waits on a
Kubernetes rollout, an image pull, or a Terraform apply.
## GitOps / ArgoCD Gotchas (transferable to any ArgoCD-based homelab)
🟠 **A `kustomization.yaml` with an explicit `resources:` allowlist
silently drops anything you forget to list.** No error, no drift shown in
ArgoCD's UI — it just reports `Synced/Healthy` against a manifest set that
never included your new file. Always run `kubectl kustomize <dir>/`
locally before pushing to confirm exactly what ArgoCD will build.
🟠 **A top-level `namespace:` transformer in `kustomization.yaml` rewrites
`metadata.namespace` on every resource it builds** — including RBAC
bindings deliberately targeting a *different* namespace (e.g. granting a
ServiceAccount in namespace A read access to Secrets in namespace B). If
any manifest needs cross-namespace RBAC, either drop the transformer
(safe if every resource already sets its own explicit namespace) or give
that manifest its own Application/directory.
🟠 **PreSync hooks run before an Application's own normal resources are
synced.** A PreSync Job that depends on RBAC/ServiceAccounts defined as
plain (non-hook) resources in the *same* Application will deadlock — it
tries to start before its own permissions exist. Use PostSync instead if
the hook needs resources from its own Application, or move the
prerequisite RBAC into an earlier sync-wave Application.
🟠 **ArgoCD hooks are not continuously reconciled by `selfHeal`.** Once a
hook Job finishes (success, or exhausts `backoffLimit`), it's only
deleted+recreated during an *actual new Sync operation* — not by passive
drift detection, even with `automated.selfHeal: true` on. If you fix a
broken hook's spec and push, the Application's `status.sync.revision` can
show "caught up" while the live hook resource is still the old, broken
one, because no new operation actually re-ran it. To force it: delete the
stuck hook (clear `argocd.argoproj.io/hook-finalizer` manually if it's
stuck `Terminating`), and if that alone doesn't trigger a fresh full sync,
delete + re-`kubectl apply -f` the Application object itself.
🟠 **If you route ArgoCD's own `repoURL` through an ingress/reverse-proxy
hostname that only listens on 80/443, don't use a non-standard port in the
URL** — it'll silently time out trying to reach a port the proxy never
opened, and depending on your setup this can block *every* Application's
sync simultaneously (repo-server can't fetch git refs for anything).
🟠 **Don't pin exact version tags for images from registries that don't
guarantee tag retention** (Bitnami stopped publishing versioned tags for
free-tier images in 2025 — only `latest` + sha256 digests remain). Verify
a tag actually exists before pinning it, or prefer minimal base images +
a stdlib-only runtime download (e.g. Python's `urllib.request` to fetch a
static binary) to avoid depending on any third party's tagging policy.
🟠 **Non-root containers can't `apk add`/`apt install` in most default
base images** — package manager directories are root-owned. Use a
world-writable scratch dir (`/tmp`) for anything you need to
download/install at runtime instead.
🟠 **Helm does not validate unknown `values.yaml` keys.** A typo, or a
values schema copied from the wrong chart *version's* docs/examples, is
silently a no-op — not an error. Before concluding "this chart doesn't
support X," clone the chart at your exact pinned version/tag and run
`helm template` against your real values file, then diff the rendered
output. Don't trust a chart's current `main`-branch example values file
if you're pinned to an older release — schemas do change between major
versions without warning in your own values file.
## Service Integration Routes
**New service? Pick your stack below** (adjust doc paths to match your repo):
| Need | Doc | Example |
|------|-----|---------|
| **Authentication** | `project-usage/authentik-oidc.md` | OAuth2 login, RBAC groups, JWT tokens |
| **Async messaging** | `project-usage/sqs-messaging.md` | Queue consumers, fire-and-forget, DLQ |
| **Object storage** | `project-usage/minio-s3.md` | File uploads, backups, log backend |
| **CI/CD pipeline** | `project-usage/cicd-workflow.md` | Pipeline syntax, image push, ArgoCD sync |
| **Workflows** | `project-usage/temporal-workflows.md` | Long-running jobs, retries, state machines |
| **Database** | `project-usage/database-postgres.md` | CloudNativePG setup, schema migrations, replicas |
| **Monitoring** | `project-usage/monitoring-metrics.md` | Prometheus scrape, Grafana dashboard, alerts |
| **Secrets** | `project-usage/vault-secrets.md` | Store credentials, rotate tokens, seal/unseal |
| **Networking** | `project-usage/networking-ingress.md` | Public HTTPS, hostname routing, TLS |
## Cluster Essentials (fill in your own inventory)
**Architecture principles (adjust to taste, but these travel well):**
- Immutable OS (Talos, or similar — no SSH, fully declarative config)
- Secrets in a proper secrets backend (Vault) + SOPS-encrypted manifests in
git (`*.enc.yaml`, age-encrypted); never commit plaintext secrets or `.env`
- ArgoCD app-of-apps as the single CD source of truth; two-phase bootstrap
documented separately (chicken-and-egg: ArgoCD needs to exist before it
can deploy itself declaratively — document your exact bootstrap steps)
- Pull-based GitOps — no kubeconfig/cluster credentials ever touch your CI
runner; the runner only needs push access to git, ArgoCD does the rest
- Federated OIDC (one identity provider fronting every service that
supports it)
## Deployment Checklist (per new service)
- [ ] Prometheus `/metrics` endpoint or ServiceMonitor, if it exposes metrics
- [ ] All credentials in your secrets backend (never in plain values.yaml,
pod env directly, or committed anywhere in cleartext)
- [ ] Ingress rule with TLS, if externally reachable
- [ ] Dashboard + alert rules, if metrics are exposed
- [ ] ArgoCD `Application` manifest added to the appropriate sync-wave file,
**not** a standalone `helm install`/`kubectl apply` run by hand
- [ ] Validated locally before push: `kubectl apply --dry-run=client -f`,
`kubectl kustomize <dir>/` (if applicable), or `helm template` against
the exact pinned chart version (if Helm-sourced)
- [ ] After push: confirmed ArgoCD's `status.sync.revision` actually matches
your new commit — not just that `status.sync.status` says `Synced`
(see Gotchas — a stale hook can hide behind an otherwise-current app)
## Git & Release
**Incremental commits (service-layer grouped) tend to age well:**
- Foundation & Docs
- Core Infra (CNI, ingress, cert management, storage)
- Observability Stack
- IAM & Secrets
- CI/CD & GitOps
- Messaging / Data Infrastructure
- Applications & Utilities
Grouping by layer (rather than by day or by "misc fixes") makes it much
easier to `git log --oneline -- <path>` your way back to *why* a given
piece of config looks the way it does, months later.