Files
homelab/project-usage/infra-practice.md

12 KiB
Raw Permalink Blame History

Infrastructure Practice Playbook

Standardized procedures for troubleshooting, developing, deploying, and operating the homelab platform. Each procedure explicitly calls out where the core CLI fits vs kubectl/helmfile/direct cluster access. See core-cli-tools.md for the auth/secrets domain split; see infra-troubleshooting.md for quick patterns and gotchas.

Procedure A: Troubleshoot a Service or Cluster Issue

  1. Identify the domain.

    • Vault/secrets: core get/core put failing, Vault unreachable.
    • Node/Talos: node crashes, disk full, kubelet unreachable, network issues.
    • Plain Kubernetes: pod CrashLoop, service 503, deployment stuck.
  2. If Vault-adjacent (secrets, authentication failing):

    • Run core secrets status first (NOT core auth status — common mistake).
    • If Vault UNREACHABLE, check DNS/networking:
      • No wildcard DNS exists; verify manual /etc/hosts entries (10.6.0.1 for WireGuard, 192.168.1.160 for LAN).
      • Ping the Vault service: kubectl get svc -n vault | grep vault.
    • If Vault token not cached, run core secrets login, approve device code in browser.
    • Verify: core secrets status shows ✓ Authenticated.
  3. If node/Talos-adjacent (kubelet logs, node state, services failing):

    • Run core auth status first.
    • If token expired, run core auth login-oob, approve device code in browser.
    • Then run core nodes to list cluster nodes.
    • For a specific node, run core status <ip> (Talos state).
    • Inspect Talos services: core services <ip> (kubelet, etcd, controller, etc.).
    • Check service logs: core logs <ip> (main Talos logs) or core log-svc <ip> kubelet (specific service).
    • Consult infra-troubleshooting.md for Pod stuck in CrashLoopBackOff and kubelet restart patterns.
  4. If plain Kubernetes (pod/deployment/service issues):

    • Consult root TROUBLESHOOTING.md for the layer-before-tool SRE methodology (procedures 110).
    • Use infra-troubleshooting.md § Quick Patterns for common diagnoses:
      • CrashLoopBackOff: kubectl logs -n <ns> <pod> --tail=50 + kubectl describe pod -n <ns> <pod> | grep -A 10 Events.
      • Service 503: kubectl get endpoints -n <ns> <svc> (endpoints missing?) + kubectl get pods -n <ns> -o wide (pods not Ready?).
      • Helm release stuck: helmfile status | grep -E "FAILED|UNKNOWN|PENDING" + helm status <release> -n <ns> --show-resources.
    • If still unclear, escalate to kubectl get all -n <ns> and review resource events.
  5. For dashboards and live metrics:

    • Use core pf grafana (port-forward to localhost:3000) rather than raw kubectl port-forward — keeps forwarded ports consistent.
    • If Prometheus unavailable, check: kubectl get pods -n monitoring | grep prometheus.
    • If ServiceMonitor not scraping, verify: kubectl get servicemonitor -A | grep <name> and inspect .spec.selector matches the target pod's app label.

Procedure B: Launch/Develop a New Service or POC

  1. Plan the service.

    • Determine namespace (e.g., sqs, temporal, databases, monitoring).
    • Decide if metrics exported (most should) and if OIDC-gated.
    • Sketch a Helm values.yaml structure (secrets, replicas, resource requests, affinity).
  2. Authenticate to Vault.

    • Run core secrets login and approve device code in browser.
    • Verify: core secrets status shows ✓ Authenticated.
    • You'll need Vault access to store service secrets in step 5.
  3. Create service Helm chart directory.

    • Create k8s/<service>/ with at minimum:
      • values.yaml (Helm values for deployment, service, replicas, resource limits).
      • charts/ subdirectory for any custom local Helm charts (optional).
    • Follow naming conventions from coding-standards.md.
  4. Add Helm release to helmfile.

    • Open helmfile.yaml.gotmpl.
    • Add release block under releases: section, following this structure:
      - name: <service>
        namespace: <namespace>
        chart: <chart-repo>/<chart-name>
        version: ~1.0  # pin major.minor, allow patch updates
        needs:
          - <dependency-namespace>/<dependency-release>  # if applicable
        values:
          - k8s/<service>/values.yaml
          - secretsInline:
              DB_PASSWORD: "{{ env \"<SERVICE>_DB_PASSWORD\" }}"
      
    • Consult coding-standards.md for needs: ordering (example: sqs section shows strimzi-operator → kafka-cluster → queue-crd → management-service).
    • Reference real example: root helmfile's sqs section.
  5. If the service needs secrets (DB password, API key, OAuth secret):

    • Generate value (e.g., openssl rand -hex 32 for passwords).
    • Store in Vault: core put cluster/<SERVICE>_<KEY> <SERVICE>_<KEY>="value".
    • Critical gotcha: field name MUST equal variable name (e.g., FORGEJO_ADMIN_PASSWORD= not value=) per coding-standards.md § Vault field=variable convention.
    • Reference in values.yaml via {{ env "VARIABLE_NAME" }} (Helmfile Go template syntax, NOT shell ${VAR}).
    • Do NOT hardcode secrets in values.yaml or ConfigMaps.
  6. Verify Helm syntax before deploy.

    • Run helmfile lint (catches template errors, duplicate releases).
    • Run helmfile diff -l name=<service> (show what will be deployed).
    • Review diff for correctness (verify env var substitutions, resource limits, affinity rules).
  7. Deploy the service.

    • Run helmfile apply -l name=<service>.
    • Monitor: kubectl get pods -n <namespace> -w (watch until Running).
    • If pods stuck: kubectl describe pod -n <namespace> <pod-name> (check Events for SchedulingFailed, ImagePullBackOff, etc.).
  8. If the service exports /metrics (Prometheus format):

    • Create ServiceMonitor: k8s/monitoring/servicemonitors/svc-<name>.yaml.
      • .spec.selector.matchLabels must match the service's pod labels (usually app: <service>).
      • .spec.endpoints[0].port must match the service port name or number exporting metrics.
    • Create PrometheusRule: k8s/monitoring/alerts/svc-<name>-rules.yaml.
      • Include error rate, latency, and SLO alert rules.
      • Use prometheus as the rule group.
    • Create Grafana dashboard: k8s/monitoring/dashboards/svc-<name>.yaml.
      • Use 6-row template: Availability, Resources, Domain metrics, Logs, SLO, Related.
      • See README.md § Example Applications for a full walkthrough.
    • Verify scrape: kubectl get servicemonitor -A | grep <name> and check Prometheus Targets UI for green status.
  9. If OIDC/IAM-gated (admin UI, restricted API):

    • Create app in Authentik: core iam create-app "my-service" --slug my-service --redirect-uri "https://my-service.riotpiao.com/callback".
    • Bind app to group: core iam bind-app my-service <group> (e.g., grafana-admins for admin-only UI).
    • Retrieve credentials: core iam describe-app my-service (client ID, client secret).
    • Deploy secret: kubectl create secret generic <service>-oidc --from-literal=client-id=<ID> --from-literal=client-secret=<SECRET> -n <namespace>.
    • Reference secret in values.yaml: mount via .spec.template.spec.containers[].env or volumeMounts.
    • See core-cli-tools.md § Access Control Tiers for Tier A (OIDC + RBAC) vs Tier B (network perimeter only).
  10. Verify service is live.

    • Pods: kubectl get pods -n <namespace> -o wide (all Running, 1/1 Ready).
    • Metrics (if applicable): kubectl get servicemonitor -A | grep <name> and visit Prometheus Targets or Grafana dashboard.
    • Endpoint: If publicly routed via Ingress, verify /etc/hosts entry (10.6.0.1 for WireGuard, 192.168.1.160 for LAN) and curl https://my-service.riotpiao.com/health (or equivalent health endpoint).
    • Logs: kubectl logs -n <namespace> <pod> (no errors).

Definition of Done (Per Service)

  • Helm chart version pinned (~1.0 format in helmfile)
  • All secrets in Vault (none in values.yaml or ConfigMap)
  • /metrics endpoint exported (if applicable)
  • ServiceMonitor resource created (if metrics exported)
  • PrometheusRule with error/latency/SLO alerts (if metrics exported)
  • Grafana dashboard (if metrics exported; 6-row template: Availability, Resources, Domain, Logs, SLO, Related)
  • Ingress rule (if external access needed)
  • OIDC integration via core iam (if UI component)
  • Verified: helmfile diff clean, pods Running, dashboard live or /metrics returning 200

Procedure C: Operate the Cluster (Node Health, Context, Cleanup)

  1. Daily health check.

    • Check auth: core auth status (if OK, node ops will work).
    • List nodes: core nodes.
    • For each node, check Talos state: core status <ip>.
    • Check K8s nodes: kubectl get nodes -o wide (all Ready, no NotReady).
    • Check pod pressure: kubectl get nodes -o json | jq '.items[] | {name: .metadata.name, memory: .status.allocatable.memory, pods: .status.allocatable.pods}'.
  2. Troubleshoot a specific node.

    • Get node IP: core nodes and note the IP.
    • Check Talos services: core services <ip> (kubelet, etcd, controller should be running).
    • Check service logs: core logs <ip> (main Talos daemon logs).
    • Filter to specific service: core log-svc <ip> kubelet (kubelet logs only).
    • Restart a service if needed: core restart <ip> kubelet (graceful kubelet restart).
  3. Pod cleanup (Failed, Evicted, Terminating pods).

    • Run core pods clean (scans all namespaces, removes stale pods).
    • Verify: kubectl get pods -A | grep -E "Failed|Evicted" (should be empty).
  4. Switch kubectl context (when off-LAN, on WireGuard).

    • List available contexts: core config kube-list.
    • Switch to WireGuard path (10.6.0.1:6443): core config kube-use admin@homelab-cluster-1.
    • Known limitation: core config use <talos-context> doesn't map to WireGuard; use kube-use directly.
    • Verify: kubectl cluster-info shows 10.6.0.1 (not 192.168.1.213).
  5. MinIO bucket operations (if managing data/backups).

    • List buckets: core bucket list.
    • Upload file: core bucket upload <bucket> <local-file>.
    • Download file: core bucket download <bucket> <remote-file> -o <local-file>.
    • Delete file: core bucket delete <bucket> <remote-file>.
  6. Bootstrap or hardware runbooks (infrequent).

    • Fresh cluster setup: See README.md § Bootstrap Order (14 steps).
    • Adding a new Talos node: See README.md § Adding Hardware.
    • Do not re-explain those long procedures here; consult README.md directly.

Notes

Queue subsystem (Kafka/kmsvc/Temporal namespace auto-registration): Already deployed and stable. If re-deploying:

  • Primary deploy method: helmfile apply -l namespace=sqs (live from root helmfile).
  • Alternate isolated iterate path: k8s/sqs/helmfile.yaml.gotmpl (not recommended for production).
  • Planned future: GitOps via k8s/sqs/argocd/ (companion repo, not yet active).
  • Critical rule: Temporal namespace registration is automatic via queue-operator; never manually temporal operator namespace create for any namespace referenced by a Queue's temporal.io/namespace label. See ~/workplace/kmsvc-manage/CLAUDE.md ("Temporal Namespace Registration") for the full rule and why.

Shared/reusable service repositories: If a service's Helm chart and container image live in a separate repository, they must be:

  • Published as a public GitHub repository under the Riotpiaole organization.
  • Images pushed to GHCR (ghcr.io/riotpiaole/...) for public pullability.
  • Consult coding-standards.md § Shared/Reusable Repos for the full publishing rule.

Cross-References

  • core-cli-tools.md: Auth/secrets domain split, command inventory, when to use core vs kubectl.
  • coding-standards.md: Helm naming conventions, needs: ordering rules, helmfile template syntax ({{ env "VAR" }} not ${VAR}), Vault field=variable convention, shared-repo publishing rule.
  • infra-troubleshooting.md: Quick patterns (CrashLoopBackOff, 503, helm stuck), gotchas, hard rules.
  • USAGE.md: Exhaustive core command reference.
  • README.md: Bootstrap order, hardware addition, example app walkthrough, 6-row Grafana dashboard template.
  • root TROUBLESHOOTING.md: Generic Kubernetes SRE layer-before-tool methodology (10 diagnostic procedures).