# Infrastructure Practice Playbook Standardized procedures for troubleshooting, developing, deploying, and operating the homelab platform. Each procedure explicitly calls out where the `core` CLI fits vs `kubectl`/`helmfile`/direct cluster access. See `core-cli-tools.md` for the auth/secrets domain split; see `infra-troubleshooting.md` for quick patterns and gotchas. ## Procedure A: Troubleshoot a Service or Cluster Issue 1. **Identify the domain.** - Vault/secrets: `core get`/`core put` failing, Vault unreachable. - Node/Talos: node crashes, disk full, kubelet unreachable, network issues. - Plain Kubernetes: pod CrashLoop, service 503, deployment stuck. 2. **If Vault-adjacent (secrets, authentication failing):** - Run `core secrets status` first (NOT `core auth status` — common mistake). - If `Vault UNREACHABLE`, check DNS/networking: - No wildcard DNS exists; verify manual `/etc/hosts` entries (10.6.0.1 for WireGuard, 192.168.1.160 for LAN). - Ping the Vault service: `kubectl get svc -n vault | grep vault`. - If `Vault token not cached`, run `core secrets login`, approve device code in browser. - Verify: `core secrets status` shows `✓ Authenticated`. 3. **If node/Talos-adjacent (kubelet logs, node state, services failing):** - Run `core auth status` first. - If token expired, run `core auth login-oob`, approve device code in browser. - Then run `core nodes` to list cluster nodes. - For a specific node, run `core status ` (Talos state). - Inspect Talos services: `core services ` (kubelet, etcd, controller, etc.). - Check service logs: `core logs ` (main Talos logs) or `core log-svc kubelet` (specific service). - Consult `infra-troubleshooting.md` for Pod stuck in CrashLoopBackOff and kubelet restart patterns. 4. **If plain Kubernetes (pod/deployment/service issues):** - Consult root `TROUBLESHOOTING.md` for the layer-before-tool SRE methodology (procedures 1–10). - Use `infra-troubleshooting.md` § Quick Patterns for common diagnoses: - CrashLoopBackOff: `kubectl logs -n --tail=50` + `kubectl describe pod -n | grep -A 10 Events`. - Service 503: `kubectl get endpoints -n ` (endpoints missing?) + `kubectl get pods -n -o wide` (pods not Ready?). - Helm release stuck: `helmfile status | grep -E "FAILED|UNKNOWN|PENDING"` + `helm status -n --show-resources`. - If still unclear, escalate to `kubectl get all -n ` and review resource events. 5. **For dashboards and live metrics:** - Use `core pf grafana` (port-forward to localhost:3000) rather than raw `kubectl port-forward` — keeps forwarded ports consistent. - If Prometheus unavailable, check: `kubectl get pods -n monitoring | grep prometheus`. - If ServiceMonitor not scraping, verify: `kubectl get servicemonitor -A | grep ` and inspect `.spec.selector` matches the target pod's app label. --- ## Procedure B: Launch/Develop a New Service or POC 1. **Plan the service.** - Determine namespace (e.g., `sqs`, `temporal`, `databases`, `monitoring`). - Decide if metrics exported (most should) and if OIDC-gated. - Sketch a Helm values.yaml structure (secrets, replicas, resource requests, affinity). 2. **Authenticate to Vault.** - Run `core secrets login` and approve device code in browser. - Verify: `core secrets status` shows `✓ Authenticated`. - You'll need Vault access to store service secrets in step 5. 3. **Create service Helm chart directory.** - Create `k8s//` with at minimum: - `values.yaml` (Helm values for deployment, service, replicas, resource limits). - `charts/` subdirectory for any custom local Helm charts (optional). - Follow naming conventions from `coding-standards.md`. 4. **Add Helm release to helmfile.** - Open `helmfile.yaml.gotmpl`. - Add release block under `releases:` section, following this structure: ```yaml - name: namespace: chart: / version: ~1.0 # pin major.minor, allow patch updates needs: - / # if applicable values: - k8s//values.yaml - secretsInline: DB_PASSWORD: "{{ env \"_DB_PASSWORD\" }}" ``` - Consult `coding-standards.md` for `needs:` ordering (example: sqs section shows strimzi-operator → kafka-cluster → queue-crd → management-service). - Reference real example: root helmfile's `sqs` section. 5. **If the service needs secrets (DB password, API key, OAuth secret):** - Generate value (e.g., `openssl rand -hex 32` for passwords). - Store in Vault: `core put cluster/_ _="value"`. - **Critical gotcha:** field name MUST equal variable name (e.g., `FORGEJO_ADMIN_PASSWORD=` not `value=`) per `coding-standards.md` § Vault field=variable convention. - Reference in values.yaml via `{{ env "VARIABLE_NAME" }}` (Helmfile Go template syntax, NOT shell `${VAR}`). - Do NOT hardcode secrets in values.yaml or ConfigMaps. 6. **Verify Helm syntax before deploy.** - Run `helmfile lint` (catches template errors, duplicate releases). - Run `helmfile diff -l name=` (show what will be deployed). - Review diff for correctness (verify env var substitutions, resource limits, affinity rules). 7. **Deploy the service.** - Run `helmfile apply -l name=`. - Monitor: `kubectl get pods -n -w` (watch until Running). - If pods stuck: `kubectl describe pod -n ` (check Events for SchedulingFailed, ImagePullBackOff, etc.). 8. **If the service exports `/metrics` (Prometheus format):** - Create ServiceMonitor: `k8s/monitoring/servicemonitors/svc-.yaml`. - `.spec.selector.matchLabels` must match the service's pod labels (usually `app: `). - `.spec.endpoints[0].port` must match the service port name or number exporting metrics. - Create PrometheusRule: `k8s/monitoring/alerts/svc--rules.yaml`. - Include error rate, latency, and SLO alert rules. - Use `prometheus` as the rule group. - Create Grafana dashboard: `k8s/monitoring/dashboards/svc-.yaml`. - Use 6-row template: Availability, Resources, Domain metrics, Logs, SLO, Related. - See README.md § Example Applications for a full walkthrough. - Verify scrape: `kubectl get servicemonitor -A | grep ` and check Prometheus Targets UI for green status. 9. **If OIDC/IAM-gated (admin UI, restricted API):** - Create app in Authentik: `core iam create-app "my-service" --slug my-service --redirect-uri "https://my-service.riotpiao.com/callback"`. - Bind app to group: `core iam bind-app my-service ` (e.g., `grafana-admins` for admin-only UI). - Retrieve credentials: `core iam describe-app my-service` (client ID, client secret). - Deploy secret: `kubectl create secret generic -oidc --from-literal=client-id= --from-literal=client-secret= -n `. - Reference secret in values.yaml: mount via `.spec.template.spec.containers[].env` or volumeMounts. - See `core-cli-tools.md` § Access Control Tiers for Tier A (OIDC + RBAC) vs Tier B (network perimeter only). 10. **Verify service is live.** - Pods: `kubectl get pods -n -o wide` (all Running, 1/1 Ready). - Metrics (if applicable): `kubectl get servicemonitor -A | grep ` and visit Prometheus Targets or Grafana dashboard. - Endpoint: If publicly routed via Ingress, verify `/etc/hosts` entry (10.6.0.1 for WireGuard, 192.168.1.160 for LAN) and `curl https://my-service.riotpiao.com/health` (or equivalent health endpoint). - Logs: `kubectl logs -n ` (no errors). ### Definition of Done (Per Service) - [ ] Helm chart version pinned (~1.0 format in helmfile) - [ ] All secrets in Vault (none in values.yaml or ConfigMap) - [ ] `/metrics` endpoint exported (if applicable) - [ ] ServiceMonitor resource created (if metrics exported) - [ ] PrometheusRule with error/latency/SLO alerts (if metrics exported) - [ ] Grafana dashboard (if metrics exported; 6-row template: Availability, Resources, Domain, Logs, SLO, Related) - [ ] Ingress rule (if external access needed) - [ ] OIDC integration via `core iam` (if UI component) - [ ] Verified: `helmfile diff` clean, pods Running, dashboard live or `/metrics` returning 200 --- ## Procedure C: Operate the Cluster (Node Health, Context, Cleanup) 1. **Daily health check.** - Check auth: `core auth status` (if OK, node ops will work). - List nodes: `core nodes`. - For each node, check Talos state: `core status `. - Check K8s nodes: `kubectl get nodes -o wide` (all Ready, no NotReady). - Check pod pressure: `kubectl get nodes -o json | jq '.items[] | {name: .metadata.name, memory: .status.allocatable.memory, pods: .status.allocatable.pods}'`. 2. **Troubleshoot a specific node.** - Get node IP: `core nodes` and note the IP. - Check Talos services: `core services ` (kubelet, etcd, controller should be running). - Check service logs: `core logs ` (main Talos daemon logs). - Filter to specific service: `core log-svc kubelet` (kubelet logs only). - Restart a service if needed: `core restart kubelet` (graceful kubelet restart). 3. **Pod cleanup (Failed, Evicted, Terminating pods).** - Run `core pods clean` (scans all namespaces, removes stale pods). - Verify: `kubectl get pods -A | grep -E "Failed|Evicted"` (should be empty). 4. **Switch kubectl context (when off-LAN, on WireGuard).** - List available contexts: `core config kube-list`. - Switch to WireGuard path (10.6.0.1:6443): `core config kube-use admin@homelab-cluster-1`. - **Known limitation:** `core config use ` doesn't map to WireGuard; use `kube-use` directly. - Verify: `kubectl cluster-info` shows 10.6.0.1 (not 192.168.1.213). 5. **MinIO bucket operations (if managing data/backups).** - List buckets: `core bucket list`. - Upload file: `core bucket upload `. - Download file: `core bucket download -o `. - Delete file: `core bucket delete `. 6. **Bootstrap or hardware runbooks (infrequent).** - **Fresh cluster setup:** See README.md § Bootstrap Order (14 steps). - **Adding a new Talos node:** See README.md § Adding Hardware. - Do not re-explain those long procedures here; consult README.md directly. --- ## Notes **Queue subsystem (Kafka/kmsvc/Temporal namespace auto-registration):** Already deployed and stable. If re-deploying: - Primary deploy method: `helmfile apply -l namespace=sqs` (live from root helmfile). - Alternate isolated iterate path: `k8s/sqs/helmfile.yaml.gotmpl` (not recommended for production). - Planned future: GitOps via `k8s/sqs/argocd/` (companion repo, not yet active). - **Critical rule:** Temporal namespace registration is automatic via `queue-operator`; never manually `temporal operator namespace create` for any namespace referenced by a Queue's `temporal.io/namespace` label. See `~/workplace/kmsvc-manage/CLAUDE.md` ("Temporal Namespace Registration") for the full rule and why. **Shared/reusable service repositories:** If a service's Helm chart and container image live in a separate repository, they must be: - Published as a public GitHub repository under the `Riotpiaole` organization. - Images pushed to GHCR (`ghcr.io/riotpiaole/...`) for public pullability. - Consult `coding-standards.md` § Shared/Reusable Repos for the full publishing rule. --- ## Cross-References - **core-cli-tools.md:** Auth/secrets domain split, command inventory, when to use `core` vs `kubectl`. - **coding-standards.md:** Helm naming conventions, `needs:` ordering rules, helmfile template syntax (`{{ env "VAR" }}` not `${VAR}`), Vault field=variable convention, shared-repo publishing rule. - **infra-troubleshooting.md:** Quick patterns (CrashLoopBackOff, 503, helm stuck), gotchas, hard rules. - **USAGE.md:** Exhaustive `core` command reference. - **README.md:** Bootstrap order, hardware addition, example app walkthrough, 6-row Grafana dashboard template. - **root TROUBLESHOOTING.md:** Generic Kubernetes SRE layer-before-tool methodology (10 diagnostic procedures).