Files
homelab/project-usage/infra-practice.md

199 lines
12 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Infrastructure Practice Playbook
Standardized procedures for troubleshooting, developing, deploying, and operating the homelab platform. Each procedure explicitly calls out where the `core` CLI fits vs `kubectl`/`helmfile`/direct cluster access. See `core-cli-tools.md` for the auth/secrets domain split; see `infra-troubleshooting.md` for quick patterns and gotchas.
## Procedure A: Troubleshoot a Service or Cluster Issue
1. **Identify the domain.**
- Vault/secrets: `core get`/`core put` failing, Vault unreachable.
- Node/Talos: node crashes, disk full, kubelet unreachable, network issues.
- Plain Kubernetes: pod CrashLoop, service 503, deployment stuck.
2. **If Vault-adjacent (secrets, authentication failing):**
- Run `core secrets status` first (NOT `core auth status` — common mistake).
- If `Vault UNREACHABLE`, check DNS/networking:
- No wildcard DNS exists; verify manual `/etc/hosts` entries (10.6.0.1 for WireGuard, 192.168.1.160 for LAN).
- Ping the Vault service: `kubectl get svc -n vault | grep vault`.
- If `Vault token not cached`, run `core secrets login`, approve device code in browser.
- Verify: `core secrets status` shows `✓ Authenticated`.
3. **If node/Talos-adjacent (kubelet logs, node state, services failing):**
- Run `core auth status` first.
- If token expired, run `core auth login-oob`, approve device code in browser.
- Then run `core nodes` to list cluster nodes.
- For a specific node, run `core status <ip>` (Talos state).
- Inspect Talos services: `core services <ip>` (kubelet, etcd, controller, etc.).
- Check service logs: `core logs <ip>` (main Talos logs) or `core log-svc <ip> kubelet` (specific service).
- Consult `infra-troubleshooting.md` for Pod stuck in CrashLoopBackOff and kubelet restart patterns.
4. **If plain Kubernetes (pod/deployment/service issues):**
- Consult root `TROUBLESHOOTING.md` for the layer-before-tool SRE methodology (procedures 110).
- Use `infra-troubleshooting.md` § Quick Patterns for common diagnoses:
- CrashLoopBackOff: `kubectl logs -n <ns> <pod> --tail=50` + `kubectl describe pod -n <ns> <pod> | grep -A 10 Events`.
- Service 503: `kubectl get endpoints -n <ns> <svc>` (endpoints missing?) + `kubectl get pods -n <ns> -o wide` (pods not Ready?).
- Helm release stuck: `helmfile status | grep -E "FAILED|UNKNOWN|PENDING"` + `helm status <release> -n <ns> --show-resources`.
- If still unclear, escalate to `kubectl get all -n <ns>` and review resource events.
5. **For dashboards and live metrics:**
- Use `core pf grafana` (port-forward to localhost:3000) rather than raw `kubectl port-forward` — keeps forwarded ports consistent.
- If Prometheus unavailable, check: `kubectl get pods -n monitoring | grep prometheus`.
- If ServiceMonitor not scraping, verify: `kubectl get servicemonitor -A | grep <name>` and inspect `.spec.selector` matches the target pod's app label.
---
## Procedure B: Launch/Develop a New Service or POC
1. **Plan the service.**
- Determine namespace (e.g., `sqs`, `temporal`, `databases`, `monitoring`).
- Decide if metrics exported (most should) and if OIDC-gated.
- Sketch a Helm values.yaml structure (secrets, replicas, resource requests, affinity).
2. **Authenticate to Vault.**
- Run `core secrets login` and approve device code in browser.
- Verify: `core secrets status` shows `✓ Authenticated`.
- You'll need Vault access to store service secrets in step 5.
3. **Create service Helm chart directory.**
- Create `k8s/<service>/` with at minimum:
- `values.yaml` (Helm values for deployment, service, replicas, resource limits).
- `charts/` subdirectory for any custom local Helm charts (optional).
- Follow naming conventions from `coding-standards.md`.
4. **Add Helm release to helmfile.**
- Open `helmfile.yaml.gotmpl`.
- Add release block under `releases:` section, following this structure:
```yaml
- name: <service>
namespace: <namespace>
chart: <chart-repo>/<chart-name>
version: ~1.0 # pin major.minor, allow patch updates
needs:
- <dependency-namespace>/<dependency-release> # if applicable
values:
- k8s/<service>/values.yaml
- secretsInline:
DB_PASSWORD: "{{ env \"<SERVICE>_DB_PASSWORD\" }}"
```
- Consult `coding-standards.md` for `needs:` ordering (example: sqs section shows strimzi-operator → kafka-cluster → queue-crd → management-service).
- Reference real example: root helmfile's `sqs` section.
5. **If the service needs secrets (DB password, API key, OAuth secret):**
- Generate value (e.g., `openssl rand -hex 32` for passwords).
- Store in Vault: `core put cluster/<SERVICE>_<KEY> <SERVICE>_<KEY>="value"`.
- **Critical gotcha:** field name MUST equal variable name (e.g., `FORGEJO_ADMIN_PASSWORD=` not `value=`) per `coding-standards.md` § Vault field=variable convention.
- Reference in values.yaml via `{{ env "VARIABLE_NAME" }}` (Helmfile Go template syntax, NOT shell `${VAR}`).
- Do NOT hardcode secrets in values.yaml or ConfigMaps.
6. **Verify Helm syntax before deploy.**
- Run `helmfile lint` (catches template errors, duplicate releases).
- Run `helmfile diff -l name=<service>` (show what will be deployed).
- Review diff for correctness (verify env var substitutions, resource limits, affinity rules).
7. **Deploy the service.**
- Run `helmfile apply -l name=<service>`.
- Monitor: `kubectl get pods -n <namespace> -w` (watch until Running).
- If pods stuck: `kubectl describe pod -n <namespace> <pod-name>` (check Events for SchedulingFailed, ImagePullBackOff, etc.).
8. **If the service exports `/metrics` (Prometheus format):**
- Create ServiceMonitor: `k8s/monitoring/servicemonitors/svc-<name>.yaml`.
- `.spec.selector.matchLabels` must match the service's pod labels (usually `app: <service>`).
- `.spec.endpoints[0].port` must match the service port name or number exporting metrics.
- Create PrometheusRule: `k8s/monitoring/alerts/svc-<name>-rules.yaml`.
- Include error rate, latency, and SLO alert rules.
- Use `prometheus` as the rule group.
- Create Grafana dashboard: `k8s/monitoring/dashboards/svc-<name>.yaml`.
- Use 6-row template: Availability, Resources, Domain metrics, Logs, SLO, Related.
- See README.md § Example Applications for a full walkthrough.
- Verify scrape: `kubectl get servicemonitor -A | grep <name>` and check Prometheus Targets UI for green status.
9. **If OIDC/IAM-gated (admin UI, restricted API):**
- Create app in Authentik: `core iam create-app "my-service" --slug my-service --redirect-uri "https://my-service.riotpiao.com/callback"`.
- Bind app to group: `core iam bind-app my-service <group>` (e.g., `grafana-admins` for admin-only UI).
- Retrieve credentials: `core iam describe-app my-service` (client ID, client secret).
- Deploy secret: `kubectl create secret generic <service>-oidc --from-literal=client-id=<ID> --from-literal=client-secret=<SECRET> -n <namespace>`.
- Reference secret in values.yaml: mount via `.spec.template.spec.containers[].env` or volumeMounts.
- See `core-cli-tools.md` § Access Control Tiers for Tier A (OIDC + RBAC) vs Tier B (network perimeter only).
10. **Verify service is live.**
- Pods: `kubectl get pods -n <namespace> -o wide` (all Running, 1/1 Ready).
- Metrics (if applicable): `kubectl get servicemonitor -A | grep <name>` and visit Prometheus Targets or Grafana dashboard.
- Endpoint: If publicly routed via Ingress, verify `/etc/hosts` entry (10.6.0.1 for WireGuard, 192.168.1.160 for LAN) and `curl https://my-service.riotpiao.com/health` (or equivalent health endpoint).
- Logs: `kubectl logs -n <namespace> <pod>` (no errors).
### Definition of Done (Per Service)
- [ ] Helm chart version pinned (~1.0 format in helmfile)
- [ ] All secrets in Vault (none in values.yaml or ConfigMap)
- [ ] `/metrics` endpoint exported (if applicable)
- [ ] ServiceMonitor resource created (if metrics exported)
- [ ] PrometheusRule with error/latency/SLO alerts (if metrics exported)
- [ ] Grafana dashboard (if metrics exported; 6-row template: Availability, Resources, Domain, Logs, SLO, Related)
- [ ] Ingress rule (if external access needed)
- [ ] OIDC integration via `core iam` (if UI component)
- [ ] Verified: `helmfile diff` clean, pods Running, dashboard live or `/metrics` returning 200
---
## Procedure C: Operate the Cluster (Node Health, Context, Cleanup)
1. **Daily health check.**
- Check auth: `core auth status` (if OK, node ops will work).
- List nodes: `core nodes`.
- For each node, check Talos state: `core status <ip>`.
- Check K8s nodes: `kubectl get nodes -o wide` (all Ready, no NotReady).
- Check pod pressure: `kubectl get nodes -o json | jq '.items[] | {name: .metadata.name, memory: .status.allocatable.memory, pods: .status.allocatable.pods}'`.
2. **Troubleshoot a specific node.**
- Get node IP: `core nodes` and note the IP.
- Check Talos services: `core services <ip>` (kubelet, etcd, controller should be running).
- Check service logs: `core logs <ip>` (main Talos daemon logs).
- Filter to specific service: `core log-svc <ip> kubelet` (kubelet logs only).
- Restart a service if needed: `core restart <ip> kubelet` (graceful kubelet restart).
3. **Pod cleanup (Failed, Evicted, Terminating pods).**
- Run `core pods clean` (scans all namespaces, removes stale pods).
- Verify: `kubectl get pods -A | grep -E "Failed|Evicted"` (should be empty).
4. **Switch kubectl context (when off-LAN, on WireGuard).**
- List available contexts: `core config kube-list`.
- Switch to WireGuard path (10.6.0.1:6443): `core config kube-use admin@homelab-cluster-1`.
- **Known limitation:** `core config use <talos-context>` doesn't map to WireGuard; use `kube-use` directly.
- Verify: `kubectl cluster-info` shows 10.6.0.1 (not 192.168.1.213).
5. **MinIO bucket operations (if managing data/backups).**
- List buckets: `core bucket list`.
- Upload file: `core bucket upload <bucket> <local-file>`.
- Download file: `core bucket download <bucket> <remote-file> -o <local-file>`.
- Delete file: `core bucket delete <bucket> <remote-file>`.
6. **Bootstrap or hardware runbooks (infrequent).**
- **Fresh cluster setup:** See README.md § Bootstrap Order (14 steps).
- **Adding a new Talos node:** See README.md § Adding Hardware.
- Do not re-explain those long procedures here; consult README.md directly.
---
## Notes
**Queue subsystem (Kafka/kmsvc/Temporal namespace auto-registration):** Already deployed and stable. If re-deploying:
- Primary deploy method: `helmfile apply -l namespace=sqs` (live from root helmfile).
- Alternate isolated iterate path: `k8s/sqs/helmfile.yaml.gotmpl` (not recommended for production).
- Planned future: GitOps via `k8s/sqs/argocd/` (companion repo, not yet active).
- **Critical rule:** Temporal namespace registration is automatic via `queue-operator`; never manually `temporal operator namespace create` for any namespace referenced by a Queue's `temporal.io/namespace` label. See `~/workplace/kmsvc-manage/CLAUDE.md` ("Temporal Namespace Registration") for the full rule and why.
**Shared/reusable service repositories:** If a service's Helm chart and container image live in a separate repository, they must be:
- Published as a public GitHub repository under the `Riotpiaole` organization.
- Images pushed to GHCR (`ghcr.io/riotpiaole/...`) for public pullability.
- Consult `coding-standards.md` § Shared/Reusable Repos for the full publishing rule.
---
## Cross-References
- **core-cli-tools.md:** Auth/secrets domain split, command inventory, when to use `core` vs `kubectl`.
- **coding-standards.md:** Helm naming conventions, `needs:` ordering rules, helmfile template syntax (`{{ env "VAR" }}` not `${VAR}`), Vault field=variable convention, shared-repo publishing rule.
- **infra-troubleshooting.md:** Quick patterns (CrashLoopBackOff, 503, helm stuck), gotchas, hard rules.
- **USAGE.md:** Exhaustive `core` command reference.
- **README.md:** Bootstrap order, hardware addition, example app walkthrough, 6-row Grafana dashboard template.
- **root TROUBLESHOOTING.md:** Generic Kubernetes SRE layer-before-tool methodology (10 diagnostic procedures).