docs: add foundation docs and cluster configuration templates
- README: cluster architecture, quick start, use cases - USAGE: stack topology, custom CLI reference - TROUBLESHOOTING: operational safety rules - .env.example: configuration template - Makefile: build shortcuts
This commit is contained in:
@@ -0,0 +1,70 @@
|
||||
# .env.example
|
||||
# Copy to .env and fill in. The real .env is gitignored — never commit it.
|
||||
# Leave any value empty and vsource will fetch it from Vault at cluster/<KEY>.
|
||||
|
||||
# ── Cluster Configuration ──────────────────────────────────────────────────────
|
||||
# Base domain for external services (Authentik, MinIO, Forgejo, etc.)
|
||||
# Example: riotpiao.homelab.com
|
||||
CLUSTER_DOMAIN=riotpiao.homelab.com
|
||||
|
||||
# Internal Kubernetes DNS names (svc.cluster.local)
|
||||
# Only change these if your cluster domain differs
|
||||
POSTGRES_HOST=ddb-cluster-rw.ddb.svc.cluster.local
|
||||
MINIO_ENDPOINT=minio.storage.svc.cluster.local:9000
|
||||
KAFKA_BOOTSTRAP=kmsvc-kafka-bootstrap.sqs.svc.cluster.local:9092
|
||||
REDIS_ADDR=kmsvc-redis-master.sqs.svc.cluster.local:6379
|
||||
|
||||
# ── Authentik ──────────────────────────────────────────────────────────────────
|
||||
# Signing/encryption key. SET ONCE — rotating invalidates all sessions and tokens.
|
||||
# openssl rand -base64 60 | tr -d '\n'
|
||||
AUTHENTIK_SECRET_KEY=
|
||||
|
||||
# Initial password for the built-in admin 'akadmin'. Change after first login.
|
||||
# openssl rand -base64 24
|
||||
AUTHENTIK_BOOTSTRAP_PASSWORD=
|
||||
|
||||
# Initial API token for 'akadmin' (automation / blueprints).
|
||||
# openssl rand -hex 32
|
||||
AUTHENTIK_BOOTSTRAP_TOKEN=
|
||||
|
||||
# Password for the bundled PostgreSQL 'authentik' user.
|
||||
# openssl rand -base64 24
|
||||
AUTHENTIK_PG_PASSWORD=
|
||||
|
||||
# ── MinIO ──────────────────────────────────────────────────────────────────────
|
||||
MINIO_ROOT_USER=
|
||||
MINIO_ROOT_PASSWORD=
|
||||
|
||||
# ── Grafana ────────────────────────────────────────────────────────────────────
|
||||
GRAFANA_ADMIN_PASSWORD=
|
||||
|
||||
# ── Forgejo ────────────────────────────────────────────────────────────────────
|
||||
FORGEJO_ADMIN_PASSWORD=
|
||||
|
||||
# ── OIDC client secrets (required) ────────────────────────────────────────────
|
||||
# These must be pre-generated and stored in Vault before running setup_talos_iam.sh.
|
||||
# talos put cluster/GRAFANA_OIDC_CLIENT_SECRET GRAFANA_OIDC_CLIENT_SECRET="$(openssl rand -hex 32)"
|
||||
GRAFANA_OIDC_CLIENT_SECRET=
|
||||
MINIO_OIDC_CLIENT_SECRET=
|
||||
AUTHENTIK_FORGEJO_CLIENT_SECRET=
|
||||
AUTHENTIK_ARGOCD_CLIENT_SECRET=
|
||||
AUTHENTIK_OLLAMA_CLIENT_SECRET=
|
||||
AUTHENTIK_TEMPORAL_CLIENT_SECRET=
|
||||
AUTHENTIK_KMSVC_CLIENT_SECRET=
|
||||
AUTHENTIK_LONGHORN_CLIENT_SECRET=
|
||||
AUTHENTIK_PORTAINER_CLIENT_SECRET=
|
||||
|
||||
# ── OIDC client IDs (optional) ────────────────────────────────────────────────
|
||||
# Leave empty to use the provider name as client_id (the safe default).
|
||||
# Only set if you need a custom client_id (e.g. after rotating a compromised credential).
|
||||
# talos put cluster/AUTHENTIK_ARGOCD_CLIENT_ID AUTHENTIK_ARGOCD_CLIENT_ID="my-custom-id"
|
||||
GRAFANA_OIDC_CLIENT_ID=
|
||||
MINIO_OIDC_CLIENT_ID=
|
||||
AUTHENTIK_FORGEJO_CLIENT_ID=
|
||||
AUTHENTIK_ARGOCD_CLIENT_ID=
|
||||
AUTHENTIK_OLLAMA_CLIENT_ID=
|
||||
AUTHENTIK_TEMPORAL_CLIENT_ID=
|
||||
|
||||
# ── CI/CD ──────────────────────────────────────────────────────────────────────
|
||||
# Forgejo Personal Access Token (from rock user) for pushing images to registry
|
||||
FORGEJO_RIOTPIAO_PAT=
|
||||
+40
@@ -0,0 +1,40 @@
|
||||
# Environment files — real values must never be committed
|
||||
.env
|
||||
|
||||
# Private CA key and generated TLS certs — ca.key must never enter the cluster or git.
|
||||
# Only ca.crt is safe to share, but we exclude the whole dir to avoid accidents.
|
||||
forge/pki/
|
||||
|
||||
# Talos machine configs — contain WireGuard private keys, bootstrap tokens, PKI
|
||||
cluster-config/controlplane.yaml
|
||||
cluster-config/worker*.yaml
|
||||
cluster-config/secrets.yaml
|
||||
cluster-config/talosconfig
|
||||
talos-forge-trust.yaml
|
||||
# Kubeconfig — contains admin client certificate + private key
|
||||
cluster-config/kubeconfig
|
||||
|
||||
*.html
|
||||
LOG.md
|
||||
project_context.md
|
||||
.claude/*
|
||||
|
||||
ca.crt
|
||||
ca.key
|
||||
ca.srl
|
||||
|
||||
forgejo.crt
|
||||
forgejo.key
|
||||
forgejo.csr
|
||||
|
||||
# Compiled test binary — Go produces a native binary named after the directory.
|
||||
# Source is k8s/storage/test/main.go; the binary has no place in version control.
|
||||
k8s/storage/test/test
|
||||
|
||||
*.key
|
||||
*.conf
|
||||
|
||||
# Allowed markdown: CLAUDE.example.md, README.md, TROUBLESHOOTING.md
|
||||
CLAUDE.md
|
||||
|
||||
skills-lock.json
|
||||
@@ -0,0 +1,72 @@
|
||||
# CLAUDE.md — Homelab Integration Guide
|
||||
|
||||
**Homelab:** A bare-metal three-node Kubernetes cluster running Talos Linux with a full observability stack, SSO via Authentik, secret management via Vault, and CI/CD infrastructure (Forgejo + Argo CD, deployed).
|
||||
|
||||
## Service Integration Routes
|
||||
|
||||
**New service? Pick your stack below:**
|
||||
|
||||
| Need | Doc | Example |
|
||||
|------|-----|---------|
|
||||
| **Authentication** | `project-usage/authentik-oidc.md` | OAuth2 login, RBAC groups, JWT tokens |
|
||||
| **Async messaging** | `project-usage/sqs-messaging.md` | Kafka topic consumers, fire-and-forget, DLQ |
|
||||
| **Object storage** | `project-usage/minio-s3.md` | File uploads, backups, log backend |
|
||||
| **CI/CD pipeline** | `project-usage/cicd-workflow.md` | GitHub Actions syntax, image push, Argo CD sync |
|
||||
| **Workflows** | `project-usage/temporal-workflows.md` | Long-running jobs, retries, state machines |
|
||||
| **Database** | `project-usage/database-postgres.md` | CloudNativePG setup, schema migrations, replicas |
|
||||
| **Monitoring** | `project-usage/monitoring-metrics.md` | Prometheus scrape, Grafana dashboard, alerts |
|
||||
| **Secrets** | `project-usage/vault-secrets.md` | Store credentials, rotate tokens, seal/unseal |
|
||||
| **Networking** | `project-usage/networking-ingress.md` | Public HTTPS, hostname routing, TLS |
|
||||
|
||||
## Cluster Essentials
|
||||
|
||||
**22 namespaces, 18 releases:**
|
||||
```
|
||||
Core: cert-manager, ingress-nginx, kube-system, cilium
|
||||
Storage: longhorn-system, storage (MinIO)
|
||||
Data: ddb (PostgreSQL), iam (Authentik + Vault)
|
||||
Observability: logging (Loki + Grafana), monitoring (Prometheus)
|
||||
Apps: cicd (Forgejo + Argo CD), sqs (Kafka + kmsvc), temporal, story-crater-backend
|
||||
```
|
||||
|
||||
**Architecture principles:**
|
||||
- Immutable OS (Talos — no SSH, declarative config)
|
||||
- Secrets in Vault (never commit `.env`, credentials in Secret volumes)
|
||||
- Helmfile = single source of truth (`helmfile.yaml.gotmpl`)
|
||||
- Pull-based GitOps (Argo CD, no kubeconfig in CI)
|
||||
- Federated OIDC (Authentik provider for all services)
|
||||
|
||||
|
||||
## Deployment Checklist
|
||||
|
||||
- [ ] Service has Prometheus `/metrics` endpoint or ServiceMonitor
|
||||
- [ ] All credentials in Vault (never in pod env, ConfigMap, or code)
|
||||
- [ ] Ingress rule in `k8s/ingress/` with TLS cert
|
||||
- [ ] Grafana dashboard in `k8s/monitoring/dashboards/svc-<name>.yaml`
|
||||
- [ ] Alert rules in `k8s/monitoring/alerts/svc-<name>-rules.yaml` (if needed)
|
||||
- [ ] Helm release in `helmfile.yaml.gotmpl` with correct `needs:` dependencies
|
||||
|
||||
## Hard Rules
|
||||
|
||||
1. **No kubeconfig in CI** — Argo CD bridges gap (pull-based, never push secrets to runner)
|
||||
2. **Field name = variable name** — In Vault: `talos put cluster/KAFKA_BOOTSTRAP KAFKA_BOOTSTRAP="..."`
|
||||
3. **Secrets via volumes** — Never `--env` flag in pod specs (exposes in `kubectl describe`)
|
||||
4. **External services via Ingress** — All public endpoints via TLS (homelab-ca)
|
||||
5. **Never commit `.env`** — Only `.env.example` in git; real secrets in Vault
|
||||
|
||||
## Git & Release
|
||||
|
||||
**Multi-remote push:**
|
||||
```bash
|
||||
git push origin main
|
||||
```
|
||||
|
||||
**Incremental commits (service-layer grouped):**
|
||||
- Foundation & Docs
|
||||
- Helmfile & Core Infra
|
||||
- Storage Layer
|
||||
- Observability Stack
|
||||
- IAM & Secrets
|
||||
- CI/CD & GitOps
|
||||
- Messaging Infrastructure
|
||||
- Applications & Utilities
|
||||
@@ -0,0 +1,328 @@
|
||||
# ── Node IPs ──────────────────────────────────────────────────────────────────
|
||||
# CP_IP has a default. All W{N}_IP variables are expected to be exported from
|
||||
# ~/.zshrc (e.g. export W1_IP=192.168.1.162). No guards — assumed always set.
|
||||
CP_IP ?= 192.168.1.213
|
||||
|
||||
export CP_IP
|
||||
|
||||
# ── Paths ─────────────────────────────────────────────────────────────────────
|
||||
TALOSCONFIG := cluster-config/coreconfig
|
||||
CP_CONFIG := cluster-config/controlplane.yaml
|
||||
SECRETS := cluster-config/secrets.yaml
|
||||
KUBECONFIG := cluster-config/kubeconfig
|
||||
|
||||
CLUSTER_NAME := homelab-cluster
|
||||
CP_ENDPOINT := https://$(CP_IP):6443
|
||||
TALOS_IMAGE := factory.core.dev/installer/613e1592b2da41ae5e265e8789429f22e121aab91cb4deb6bc3c0b6262961245:v1.13.3
|
||||
|
||||
TALOSCTL := corectl --coreconfig $(TALOSCONFIG)
|
||||
KUBECTL := kubectl --kubeconfig $(KUBECONFIG)
|
||||
|
||||
# Derive IP and config from worker number N (used by generic targets).
|
||||
# $(W$(N)_IP) expands to e.g. $(W2_IP) when N=2.
|
||||
W_IP = $(W$(N)_IP)
|
||||
W_CONFIG = cluster-config/worker-$(N).yaml
|
||||
|
||||
.DEFAULT_GOAL := help
|
||||
|
||||
# ── Help ──────────────────────────────────────────────────────────────────────
|
||||
.PHONY: help
|
||||
help:
|
||||
@echo "Homelab cluster — available targets"
|
||||
@echo ""
|
||||
@echo " Status"
|
||||
@echo " nodes kubectl get nodes"
|
||||
@echo " status-cp core node overview (control plane)"
|
||||
@echo " status-w1 core node overview (worker-1)"
|
||||
@echo " services-cp list core services (control plane)"
|
||||
@echo " services-w1 list core services (worker-1)"
|
||||
@echo ""
|
||||
@echo " Logs"
|
||||
@echo " logs-cp stream kubelet logs (control plane)"
|
||||
@echo " logs-w1 stream kubelet logs (worker-1)"
|
||||
@echo " dmesg-cp kernel dmesg (control plane)"
|
||||
@echo " dmesg-w1 kernel dmesg (worker-1)"
|
||||
@echo " log-svc-cp stream a service log (control plane) SVC=<name>"
|
||||
@echo " log-svc-w1 stream a service log (worker-1) SVC=<name>"
|
||||
@echo ""
|
||||
@echo " Config"
|
||||
@echo " gen-config regenerate controlplane.yaml + worker-N.yaml from secrets"
|
||||
@echo " apply-cp apply controlplane.yaml to CP node (live cluster)"
|
||||
@echo " apply-w1 apply cluster-config/worker-1.yaml to worker-1"
|
||||
@echo " apply-w1-insecure first-time apply to worker-1 (no certs yet)"
|
||||
@echo " apply-worker apply cluster-config/worker-N.yaml N=<num> W<N>_IP=<ip>"
|
||||
@echo " apply-worker-new first-time apply (--insecure) N=<num> W<N>_IP=<ip>"
|
||||
@echo ""
|
||||
@echo " Upgrade"
|
||||
@echo " upgrade-cp upgrade Talos on control plane"
|
||||
@echo " upgrade-w1 upgrade Talos on worker-1"
|
||||
@echo " upgrade-worker upgrade any worker N=<num> W<N>_IP=<ip>"
|
||||
@echo ""
|
||||
@echo " Shutdown / Reboot"
|
||||
@echo " shutdown-cluster graceful full shutdown (drain w1 → off w1 → off cp)"
|
||||
@echo " shutdown-cp shut down control plane only"
|
||||
@echo " shutdown-w1 shut down worker-1 only"
|
||||
@echo " shutdown-worker shut down any worker N=<num> W<N>_IP=<ip>"
|
||||
@echo " reboot-cp reboot control plane"
|
||||
@echo " reboot-w1 reboot worker-1"
|
||||
@echo " reboot-worker reboot any worker N=<num> W<N>_IP=<ip>"
|
||||
@echo ""
|
||||
@echo " Inspect (node filesystem)"
|
||||
@echo " node-ls <ip> <path> list files on a node"
|
||||
@echo " node-read <ip> <path> read a file on a node"
|
||||
@echo ""
|
||||
@echo " Maintenance"
|
||||
@echo " clean-pods delete Evicted/Failed/Terminating pods cluster-wide"
|
||||
@echo ""
|
||||
@echo " Port-forwards"
|
||||
@echo " pf-grafana localhost:3000 → Grafana"
|
||||
@echo " pf-minio localhost:9001 → MinIO console / localhost:9000 → S3 API"
|
||||
@echo " pf-loki localhost:3100 → Loki HTTP API"
|
||||
@echo " pf-portainer localhost:9000 → Portainer UI (dashboard ns)"
|
||||
@echo " pf-prometheus localhost:9090 → Prometheus UI (monitoring ns)"
|
||||
@echo " pf-longhorn localhost:8080 → Longhorn UI"
|
||||
@echo " pf-iam localhost:7000 → Authentik IAM (when deployed)"
|
||||
@echo ""
|
||||
@echo " CLI"
|
||||
@echo " cli build core-cli and install to ~/.local/bin/core"
|
||||
@echo ""
|
||||
@echo " Variables"
|
||||
@echo " CP_IP (default: 192.168.1.160)"
|
||||
@echo " W1_IP (export from ~/.zshrc — e.g. export W1_IP=192.168.1.162)"
|
||||
@echo " N (required for generic targets — worker number, e.g. N=2)"
|
||||
@echo " W<N>_IP (export from ~/.zshrc — e.g. export W2_IP=192.168.1.163)"
|
||||
@echo " SVC (required for log-svc-* targets, e.g. SVC=kubelet)"
|
||||
|
||||
# ── Status ────────────────────────────────────────────────────────────────────
|
||||
.PHONY: nodes
|
||||
nodes:
|
||||
$(KUBECTL) get nodes -o wide
|
||||
|
||||
.PHONY: status-cp
|
||||
status-cp:
|
||||
$(TALOSCTL) --nodes $(CP_IP) get members
|
||||
|
||||
.PHONY: status-w1
|
||||
status-w1:
|
||||
$(TALOSCTL) --nodes $(W1_IP) get members
|
||||
|
||||
.PHONY: services-cp
|
||||
services-cp:
|
||||
$(TALOSCTL) --nodes $(CP_IP) service
|
||||
|
||||
.PHONY: services-w1
|
||||
services-w1:
|
||||
$(TALOSCTL) --nodes $(W1_IP) service
|
||||
|
||||
# ── Logs ──────────────────────────────────────────────────────────────────────
|
||||
.PHONY: logs-cp
|
||||
logs-cp:
|
||||
$(TALOSCTL) --nodes $(CP_IP) logs kubelet -f
|
||||
|
||||
.PHONY: logs-w1
|
||||
logs-w1:
|
||||
$(TALOSCTL) --nodes $(W1_IP) logs kubelet -f
|
||||
|
||||
.PHONY: dmesg-cp
|
||||
dmesg-cp:
|
||||
$(TALOSCTL) --nodes $(CP_IP) dmesg --follow
|
||||
|
||||
.PHONY: dmesg-w1
|
||||
dmesg-w1:
|
||||
$(TALOSCTL) --nodes $(W1_IP) dmesg --follow
|
||||
|
||||
# Usage: make log-svc-cp SVC=etcd
|
||||
.PHONY: log-svc-cp
|
||||
log-svc-cp:
|
||||
ifndef SVC
|
||||
$(error SVC is not set — run: make log-svc-cp SVC=<service-name>)
|
||||
endif
|
||||
$(TALOSCTL) --nodes $(CP_IP) logs $(SVC) -f
|
||||
|
||||
.PHONY: log-svc-w1
|
||||
log-svc-w1:
|
||||
ifndef SVC
|
||||
$(error SVC is not set — run: make log-svc-w1 SVC=<service-name>)
|
||||
endif
|
||||
$(TALOSCTL) --nodes $(W1_IP) logs $(SVC) -f
|
||||
|
||||
# ── Config generation ─────────────────────────────────────────────────────────
|
||||
.PHONY: gen-config
|
||||
gen-config:
|
||||
corectl gen config $(CLUSTER_NAME) $(CP_ENDPOINT) \
|
||||
--with-secrets $(SECRETS) \
|
||||
--output-dir cluster-config/ \
|
||||
--force
|
||||
|
||||
# ── Config apply ──────────────────────────────────────────────────────────────
|
||||
.PHONY: apply-cp
|
||||
apply-cp:
|
||||
$(TALOSCTL) apply-config \
|
||||
--nodes $(CP_IP) \
|
||||
--file $(CP_CONFIG)
|
||||
|
||||
.PHONY: apply-w1
|
||||
apply-w1:
|
||||
$(TALOSCTL) apply-config \
|
||||
--nodes $(W1_IP) \
|
||||
--file cluster-config/worker-1.yaml
|
||||
|
||||
# First-time apply to worker-1 (no certs yet)
|
||||
.PHONY: apply-w1-insecure
|
||||
apply-w1-insecure:
|
||||
$(TALOSCTL) apply-config \
|
||||
--nodes $(W1_IP) \
|
||||
--file cluster-config/worker-1.yaml \
|
||||
--insecure
|
||||
|
||||
# Generic targets — derive both IP and config from N.
|
||||
# Usage: make apply-worker N=2 W2_IP=192.168.1.162
|
||||
# make apply-worker N=3 W3_IP=192.168.1.163
|
||||
.PHONY: apply-worker
|
||||
apply-worker:
|
||||
ifndef N
|
||||
$(error N is not set — run: make apply-worker N=<num> W<N>_IP=<ip>)
|
||||
endif
|
||||
$(TALOSCTL) apply-config \
|
||||
--nodes $(W_IP) \
|
||||
--file $(W_CONFIG)
|
||||
|
||||
.PHONY: apply-worker-new
|
||||
apply-worker-new:
|
||||
ifndef N
|
||||
$(error N is not set — run: make apply-worker-new N=<num> W<N>_IP=<ip>)
|
||||
endif
|
||||
$(TALOSCTL) apply-config \
|
||||
--nodes $(W_IP) \
|
||||
--file $(W_CONFIG) \
|
||||
--insecure
|
||||
|
||||
# ── Upgrade ───────────────────────────────────────────────────────────────────
|
||||
.PHONY: upgrade-cp
|
||||
upgrade-cp:
|
||||
$(TALOSCTL) upgrade \
|
||||
--nodes $(CP_IP) \
|
||||
--image $(TALOS_IMAGE) \
|
||||
--preserve
|
||||
|
||||
.PHONY: upgrade-w1
|
||||
upgrade-w1:
|
||||
$(TALOSCTL) upgrade \
|
||||
--nodes $(W1_IP) \
|
||||
--image $(TALOS_IMAGE) \
|
||||
--preserve
|
||||
|
||||
# Usage: make upgrade-worker N=2 W2_IP=192.168.1.162
|
||||
.PHONY: upgrade-worker
|
||||
upgrade-worker:
|
||||
ifndef N
|
||||
$(error N is not set — run: make upgrade-worker N=<num> W<N>_IP=<ip>)
|
||||
endif
|
||||
$(TALOSCTL) upgrade \
|
||||
--nodes $(W_IP) \
|
||||
--image $(TALOS_IMAGE) \
|
||||
--preserve
|
||||
|
||||
# ── Shutdown / Reboot ─────────────────────────────────────────────────────────
|
||||
# Full cluster: drain workers first so pods stop cleanly, then workers off,
|
||||
# then CP last (etcd must be the final process to stop).
|
||||
.PHONY: shutdown-cluster
|
||||
shutdown-cluster:
|
||||
@echo "--- draining core-worker-1 ---"
|
||||
$(KUBECTL) drain core-worker-1 --ignore-daemonsets --delete-emptydir-data
|
||||
@echo "--- shutting down worker-1 ---"
|
||||
$(TALOSCTL) --nodes $(W1_IP) shutdown
|
||||
@echo "--- shutting down control plane (last) ---"
|
||||
$(TALOSCTL) --nodes $(CP_IP) shutdown
|
||||
|
||||
.PHONY: shutdown-cp
|
||||
shutdown-cp:
|
||||
$(TALOSCTL) --nodes $(CP_IP) shutdown
|
||||
|
||||
.PHONY: shutdown-w1
|
||||
shutdown-w1:
|
||||
$(TALOSCTL) --nodes $(W1_IP) shutdown
|
||||
|
||||
# Usage: make shutdown-worker N=2 W2_IP=192.168.1.162
|
||||
.PHONY: shutdown-worker
|
||||
shutdown-worker:
|
||||
ifndef N
|
||||
$(error N is not set — run: make shutdown-worker N=<num> W<N>_IP=<ip>)
|
||||
endif
|
||||
$(TALOSCTL) --nodes $(W_IP) shutdown
|
||||
|
||||
.PHONY: reboot-cp
|
||||
reboot-cp:
|
||||
$(TALOSCTL) --nodes $(CP_IP) reboot
|
||||
|
||||
.PHONY: reboot-w1
|
||||
reboot-w1:
|
||||
$(TALOSCTL) --nodes $(W1_IP) reboot
|
||||
|
||||
# Usage: make reboot-worker N=2 W2_IP=192.168.1.162
|
||||
.PHONY: reboot-worker
|
||||
reboot-worker:
|
||||
ifndef N
|
||||
$(error N is not set — run: make reboot-worker N=<num> W<N>_IP=<ip>)
|
||||
endif
|
||||
$(TALOSCTL) --nodes $(W_IP) reboot
|
||||
|
||||
# ── Inspect ───────────────────────────────────────────────────────────────────
|
||||
# Positional args: make node-ls 192.168.1.160 /etc/kubernetes/manifests
|
||||
# $(word 2/3, $(MAKECMDGOALS)) captures the extra words; the % rule absorbs
|
||||
# them so Make doesn't error with "No rule to make target".
|
||||
.PHONY: node-ls
|
||||
node-ls:
|
||||
$(TALOSCTL) --nodes $(word 2,$(MAKECMDGOALS)) ls $(word 3,$(MAKECMDGOALS))
|
||||
|
||||
.PHONY: node-read
|
||||
node-read:
|
||||
$(TALOSCTL) --nodes $(word 2,$(MAKECMDGOALS)) read $(word 3,$(MAKECMDGOALS))
|
||||
|
||||
# Absorb positional arguments passed to node-ls / node-read
|
||||
%:
|
||||
@:
|
||||
|
||||
# ── Maintenance ───────────────────────────────────────────────────────────────
|
||||
.PHONY: clean-pods
|
||||
clean-pods:
|
||||
@echo "--- removing Failed/Evicted pods ---"
|
||||
$(KUBECTL) delete pods -A --field-selector=status.phase=Failed --ignore-not-found
|
||||
@echo "--- force-deleting stuck Terminating pods ---"
|
||||
@$(KUBECTL) get pods -A | awk '/Terminating/{print $$1, $$2}' | \
|
||||
xargs -r -n2 sh -c '$(KUBECTL) delete pod -n $$0 $$1 --force --grace-period=0' || true
|
||||
|
||||
# ── Port-forwards ─────────────────────────────────────────────────────────────
|
||||
.PHONY: pf-query
|
||||
pf-grafana:
|
||||
$(KUBECTL) port-forward -n logging svc/grafana 3000:80
|
||||
|
||||
.PHONY: pf-minio
|
||||
pf-minio:
|
||||
$(KUBECTL) port-forward -n storage svc/minio 9001:9001 &
|
||||
$(KUBECTL) port-forward -n storage svc/minio 9000:9000 &
|
||||
|
||||
.PHONY: pf-loki
|
||||
pf-loki:
|
||||
$(KUBECTL) port-forward -n logging svc/loki 3100:3100
|
||||
|
||||
.PHONY: pf-iam
|
||||
pf-iam:
|
||||
$(KUBECTL) port-forward -n iam svc/authentik-server 7000:80
|
||||
|
||||
.PHONY: pf-portainer
|
||||
pf-portainer:
|
||||
$(KUBECTL) port-forward -n dashboard svc/portainer 9000:9000
|
||||
|
||||
.PHONY: pf-prometheus
|
||||
pf-prometheus:
|
||||
$(KUBECTL) port-forward -n monitoring svc/prometheus-kube-prometheus-prometheus 9090:9090
|
||||
|
||||
.PHONY: pf-longhorn
|
||||
pf-longhorn:
|
||||
$(KUBECTL) port-forward -n longhorn-system svc/longhorn-frontend 8080:80
|
||||
|
||||
# ── CLI ───────────────────────────────────────────────────────────────────────
|
||||
.PHONY: cli
|
||||
cli:
|
||||
$(MAKE) -C core-cli install
|
||||
@@ -0,0 +1,691 @@
|
||||
# Homelab Kubernetes Cluster
|
||||
|
||||
A bare-metal three-node Kubernetes cluster running Talos Linux with **18 Helm releases** across 22 namespaces. Includes distributed storage (MinIO + Longhorn), full observability (Prometheus + Grafana + Loki), federated SSO (Authentik OIDC), secrets management (Vault), CI/CD (Forgejo + Argo CD), messaging (Kafka + kmsvc), and workflow orchestration (Temporal).
|
||||
|
||||
## Use Cases & Architecture
|
||||
|
||||
**Why this stack?**
|
||||
|
||||
This homelab replicates **production-grade cloud-native infrastructure** on bare metal, enabling:
|
||||
|
||||
1. **Learning & Prototyping** — Test distributed systems patterns (HA databases, event-driven messaging, GitOps workflows) before deploying to cloud
|
||||
2. **Self-Hosted Services** — Run applications (Story Crater, etc.) with zero cloud lock-in; full control over data, compliance, and networking
|
||||
3. **Infrastructure as Code** — Git-driven cluster state via Helmfile + Forgejo Actions + Argo CD; every change is auditable and reproducible
|
||||
4. **Observability Sandbox** — Experiment with Prometheus metrics, Loki log aggregation, and custom Grafana dashboards at scale
|
||||
|
||||
**Typical workflow:**
|
||||
|
||||
```
|
||||
Developer pushes to Forgejo (git forge)
|
||||
↓
|
||||
Forgejo Actions CI runs tests + builds OCI image
|
||||
↓
|
||||
Image pushed to Forgejo registry (private, on-cluster)
|
||||
↓
|
||||
Argo CD detects deployment repo change (pull-based GitOps)
|
||||
↓
|
||||
New pods roll out; Grafana alerts on errors/latency
|
||||
↓
|
||||
Temporal workflows coordinate long-running operations (e.g., async jobs)
|
||||
↓
|
||||
Kafka queues decouple services (fire-and-forget messaging)
|
||||
↓
|
||||
All logs + metrics centralized in Grafana for debugging
|
||||
```
|
||||
|
||||
**Architecture principles:**
|
||||
|
||||
- **Immutable OS** — Talos Linux (no SSH, declarative machine configs)
|
||||
- **No external dependencies** — All data stored locally (MinIO, CloudNativePG, Longhorn)
|
||||
- **High availability** — 3-replica databases, multi-node storage, cross-AZ readiness (on bare metal: cross-rack affinity)
|
||||
- **Federated identity** — Single Authentik OIDC provider for all services (Grafana, MinIO, Forgejo, Argo CD)
|
||||
- **Secrets at rest** — Vault + encrypted etcd; credentials never in logs or ConfigMaps
|
||||
- **Infrastructure-as-code** — Every service deployed via Helmfile; one `helmfile apply` recovers from total failure
|
||||
|
||||
## Quick Start — Deploying the Cluster
|
||||
|
||||
### 1. Bootstrap Talos Nodes
|
||||
Bootstrap each Talos node with your cluster schematic (see `CLAUDE.md` or `README` step 1–8).
|
||||
|
||||
### 2. Set Up Secrets
|
||||
All secrets are managed via environment variables sourced from `.env` (gitignored). The helmfile template expands them at deploy time.
|
||||
|
||||
**Step 1: Copy the template**
|
||||
```bash
|
||||
cp .env.example .env
|
||||
```
|
||||
|
||||
**Step 2: Populate required secrets**
|
||||
Edit `.env` and fill in cluster configuration. See `.env.example` for all options:
|
||||
|
||||
```bash
|
||||
# Cluster configuration
|
||||
CLUSTER_DOMAIN=riotpiao.homelab.com # Your cluster domain
|
||||
POSTGRES_HOST=ddb-cluster-rw.ddb.svc.cluster.local
|
||||
MINIO_ENDPOINT=minio.storage.svc.cluster.local:9000
|
||||
KAFKA_BOOTSTRAP=kmsvc-kafka-bootstrap.sqs.svc.cluster.local:9092
|
||||
REDIS_ADDR=kmsvc-redis-master.sqs.svc.cluster.local:6379
|
||||
|
||||
# Service credentials (generate with: openssl rand -hex 32)
|
||||
MINIO_ROOT_PASSWORD=<random>
|
||||
GRAFANA_ADMIN_PASSWORD=<random>
|
||||
AUTHENTIK_SECRET_KEY=<random>
|
||||
AUTHENTIK_BOOTSTRAP_PASSWORD=<random>
|
||||
AUTHENTIK_PG_PASSWORD=<random>
|
||||
```
|
||||
|
||||
**Step 3: Load and deploy**
|
||||
```bash
|
||||
# Load .env into current shell
|
||||
vsource .env
|
||||
|
||||
# Preview all changes before deployment
|
||||
helmfile diff
|
||||
|
||||
# Deploy the entire stack
|
||||
helmfile apply
|
||||
```
|
||||
|
||||
### 3. Verify Deployment
|
||||
```bash
|
||||
# Check all pods are running
|
||||
kubectl get pods -A
|
||||
|
||||
# Confirm key services are ready
|
||||
kubectl wait deploy/authentik-server -n iam --for=condition=Available --timeout=300s
|
||||
kubectl wait deploy/grafana -n logging --for=condition=Available --timeout=300s
|
||||
|
||||
# Access Grafana
|
||||
make pf-grafana # localhost:3000 (login: admin / GRAFANA_ADMIN_PASSWORD)
|
||||
```
|
||||
|
||||
### 4. First-Time Access
|
||||
|
||||
**Default Credentials:**
|
||||
- **Authentik:** https://authentik.$(CLUSTER_DOMAIN) → login: `akadmin` / `AUTHENTIK_BOOTSTRAP_PASSWORD` (change immediately)
|
||||
- **Grafana:** https://grafana.$(CLUSTER_DOMAIN) → login: `admin` / `GRAFANA_ADMIN_PASSWORD`
|
||||
- **MinIO:** https://minio.$(CLUSTER_DOMAIN) → login: `MINIO_ROOT_USER` / `MINIO_ROOT_PASSWORD`
|
||||
- **Argo CD:** https://argocd.$(CLUSTER_DOMAIN) → login via Authentik OIDC
|
||||
- **Forgejo:** https://forgejo.$(CLUSTER_DOMAIN) → login via Authentik OIDC
|
||||
|
||||
**Next Steps:**
|
||||
1. Change default passwords in each service
|
||||
2. Configure OIDC redirects (see `k8s/talos-iam/` for details)
|
||||
3. Set up GitOps: push infrastructure to Forgejo, configure Argo CD
|
||||
4. Review dashboards in Grafana (Prometheus + Loki)
|
||||
|
||||
---
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
192.168.1.0/24 (LAN)
|
||||
│
|
||||
┌──────────────────────┼──────────────────────┐
|
||||
│ │ │
|
||||
192.168.1.* 192.168.1.* 192.168.1.*
|
||||
┌────────────────┐ ┌────────────────┐ ┌────────────────┐
|
||||
│ talos-cp-1 │ │ talos-worker-1 │ │ talos-worker-2 │
|
||||
│ Control-Plane │ │ Worker │ │ Worker │
|
||||
│ + Workloads │ │ (Storage) │ │ (Storage) │
|
||||
│ (az-a) │ │ (az-b) │ │ (az-c) │
|
||||
├────────────────┤ ├────────────────┤ ├────────────────┤
|
||||
│ Pods: │ │ Pods: │ │ Pods: │
|
||||
│ • ingress-nginx│ │ • kube-system │ │ • kube-system │
|
||||
│ • authentik │ │ • storage │ │ • storage │
|
||||
│ • vault │ │ └─ minio-2 │ │ └─ minio-3 │
|
||||
│ • logging │ │ │ │ │
|
||||
│ ├─ loki │ │ │ │ │
|
||||
│ ├─ promtail │ │ │ │ │
|
||||
│ └─ grafana │ │ │ │ │
|
||||
│ • monitoring │ │ │ │ │
|
||||
│ ├─ prom │ │ │ │ │
|
||||
│ └─ blackbox │ │ │ │ │
|
||||
│ • storage │ │ │ │ │
|
||||
│ └─ minio-1 │ │ │ │ │
|
||||
│ • cicd │ │ │ │ │
|
||||
│ ├─ forgejo │ │ │ │ │
|
||||
│ ├─ argocd │ │ │ │ │
|
||||
│ └─ runner │ │ │ │ │
|
||||
│ • sqs │ │ • sqs │ │ • sqs │
|
||||
│ ├─ kafka-0 │ │ ├─ kafka-1 │ │ ├─ kafka-2 │
|
||||
│ └─ kmsvc │ │ └─ redis │ │ │
|
||||
│ • ddb │ │ • ddb │ │ • ddb │
|
||||
│ └─ postgres-0│ │ └─ postgres-1│ │ └─ postgres-2│
|
||||
│ • temporal │ │ • temporal │ │ • temporal │
|
||||
│ ├─ server │ │ ├─ cassandra │ │ ├─ cassandra │
|
||||
│ └─ cassandra │ │ │ -1 │ │ │ -2 │
|
||||
│ -0 │ │ └─ (replica) │ │ └─ (replica) │
|
||||
│ • llm │ │ │ │ │
|
||||
│ └─ ollama │ │ │ │ │
|
||||
└────────────────┘ └────────────────┘ └────────────────┘
|
||||
|
||||
Replication & Fault Tolerance:
|
||||
Data Layer:
|
||||
• MinIO: minio-1 ↔ minio-2 ↔ minio-3 (3-way active-active S3)
|
||||
• PostgreSQL: postgres-0 ↔ postgres-1 ↔ postgres-2 (primary + 2 standbys, HA streaming replication)
|
||||
• Kafka: kafka-0 ↔ kafka-1 ↔ kafka-2 (3 brokers, RF=3, min-ISR=2, cross-AZ)
|
||||
• Temporal: cassandra-0 ↔ cassandra-1 ↔ cassandra-2 (3-node distributed)
|
||||
• Loki: loki → MinIO (chunks stored in s3://loki-chunks, 10-day retention)
|
||||
|
||||
Single-Replica Services (Protected by PodDisruptionBudget minAvailable=1):
|
||||
• Observability: Grafana, Prometheus, Loki (Recreate strategy for RWO PVCs)
|
||||
• IAM: Authentik, Vault (Recreate strategy for RWO PVCs)
|
||||
• CI/CD: Argo CD server/repo-server, Forgejo (Recreate strategy for RWO PVCs)
|
||||
• LLM: Ollama
|
||||
|
||||
Networking:
|
||||
Remote access: UR_OWN.duckdns.org → home IP → cp-1
|
||||
```
|
||||
|
||||
## Stack
|
||||
|
||||
| Layer | Technology | Namespace | Purpose |
|
||||
|-------|-----------|-----------|---------|
|
||||
| OS | Talos Linux v1.13.3 | — | Immutable, Kubernetes-native OS |
|
||||
| Kubernetes | v1.36.1 | — | Container orchestration |
|
||||
| CNI | Cilium (eBPF) | kube-system | Networking, replaces kube-proxy |
|
||||
| Ingress | Nginx Ingress Controller | ingress-nginx | Reverse proxy, hostname-based routing |
|
||||
| Block Storage | Longhorn v1.7.0 | longhorn-system | Default StorageClass |
|
||||
| Object Store | MinIO (multi-AZ) | storage | S3-compatible, site-replicated across az-a/az-b |
|
||||
| IAM / SSO | Authentik | iam | OIDC provider for Grafana, MinIO, Forgejo, Argo CD |
|
||||
| Secret Store | HashiCorp Vault | iam | KV secrets backend, JWT auth via Authentik |
|
||||
| Git Forge *(planned)* | Forgejo | forge | Git server, built-in OCI registry, Actions CI |
|
||||
| CI Runner *(planned)* | Forgejo Actions + DinD | cicd | Privileged build pod; images pushed to Forgejo OCI |
|
||||
| CD *(planned)* | Argo CD | argocd | Pull-based GitOps; never holds kubeconfig in CI |
|
||||
| Log Backend | Loki (SingleBinary) | logging | 10-day retention, backed by MinIO |
|
||||
| Log Collector | Promtail (DaemonSet) | logging | Scrapes pod logs + Talos journal |
|
||||
| Metrics | kube-prometheus-stack | monitoring | Prometheus + node-exporter + kube-state-metrics |
|
||||
| Log/Metrics UI | Grafana | logging | Dashboards for Loki + Prometheus |
|
||||
| Cluster UI | Portainer CE | dashboard | Container/workload management UI |
|
||||
| LLM Inference | Ollama | llm | Local LLM model serving (open-source models) |
|
||||
|
||||
## Repository Structure
|
||||
|
||||
```
|
||||
homelab/
|
||||
├── helmfile.yaml # Single source of truth — deploys everything
|
||||
├── .env.example # Required env vars template (copy to .env, gitignored)
|
||||
│
|
||||
├── cluster-config/ # Talos + Kubernetes bootstrap
|
||||
│ ├── cilium-values.yaml
|
||||
│ ├── longhorn_bootstrap.sh
|
||||
│ ├── controlplane.yaml # gitignored — contains secrets
|
||||
│ ├── worker-1.yaml # gitignored
|
||||
│ ├── secrets.yaml # gitignored
|
||||
│ ├── kubeconfig # gitignored
|
||||
│ └── talosconfig # gitignored
|
||||
│
|
||||
├── k8s/
|
||||
│ ├── ingress/ # Nginx Ingress Controller + all Ingress rules
|
||||
│ │ ├── nginx-values.yaml
|
||||
│ │ └── ingress.yaml
|
||||
│ ├── storage/ # MinIO multi-AZ object store
|
||||
│ │ ├── minio-az-a-values.yaml
|
||||
│ │ ├── minio-az-b-values.yaml
|
||||
│ │ ├── minio-az-a-pvc.yaml
|
||||
│ │ ├── minio-service.yaml
|
||||
│ │ ├── minio-legacy-alias.yaml
|
||||
│ │ └── minio-replication-job.yaml
|
||||
│ ├── logging/ # Observability stack
|
||||
│ │ ├── loki-values.yaml
|
||||
│ │ ├── promtail-values.yaml
|
||||
│ │ └── grafana-values.yaml
|
||||
│ ├── monitoring/ # Prometheus stack + alerting + Grafana dashboards-as-code
|
||||
│ │ ├── prometheus-values.yaml
|
||||
│ │ ├── blackbox-exporter-values.yaml # Active uptime probes (feeds Service Availability dashboard)
|
||||
│ │ ├── ingress-alerts.yaml # PrometheusRule: ingress 5xx rate, p95 latency
|
||||
│ │ └── dashboards/ # ConfigMaps picked up live by Grafana's sidecar
|
||||
│ │ ├── service-availability.yaml # Uptime probes + cert expiry (operator glance)
|
||||
│ │ ├── service-golden-signals.yaml # Latency & Golden Signals (ingress RED)
|
||||
│ │ ├── service-internals.yaml # Per-service deep-dive (MinIO/Forgejo/Argo CD/Vault/Longhorn/certs)
|
||||
│ │ ├── kube-controller-health.yaml # API server RED + kube-state-metrics controller-health proxy
|
||||
│ │ ├── hardware-overview.yaml # Per-node CPU/mem/disk/network/load summary
|
||||
│ │ └── control-plane-logs.yaml # kube-system + add-on logs (Loki)
|
||||
│ ├── portainer/ # Portainer CE
|
||||
│ │ └── portainer-values.yaml
|
||||
│ ├── talos-iam/ # Authentik + Vault IAM
|
||||
│ │ ├── authentik-values.yaml
|
||||
│ │ ├── vault-values.yaml
|
||||
│ │ ├── setup_vault.sh # One-time Vault init (not replaced by Helmfile)
|
||||
│ │ └── provision_oidc.py # Authentik OIDC provisioning
|
||||
│ ├── coredns/ # CoreDNS hostname rewrites (in-cluster DNS)
|
||||
│ │ └── coredns-configmap.yaml
|
||||
│ ├── talos-ci-cd/ # CI/CD stack (planned — not yet applied)
|
||||
│ │ ├── talos_version_control.html # Implementation plan + build runbook
|
||||
│ │ ├── forgejo-values.yaml # Forgejo Helm values (gitea-charts/gitea)
|
||||
│ │ ├── argocd-values.yaml # Argo CD Helm values
|
||||
│ │ └── charts/forgejo-runner/ # Local Helm chart for the Actions runner
|
||||
│ │ ├── Chart.yaml
|
||||
│ │ ├── values.yaml
|
||||
│ │ └── templates/
|
||||
│ │ ├── deployment.yaml # Runner + DinD sidecar, Recreate strategy
|
||||
│ │ ├── pvc.yaml # runner-reg (1 Gi) + runner-dind (30 Gi)
|
||||
│ │ └── networkpolicy.yaml # Egress: forge ns + DNS + internet only
|
||||
│ └── duckdns/ # DuckDNS DDNS updater CronJob
|
||||
│
|
||||
├── talos-cli/ # Rust CLI for Vault secret access
|
||||
├── project_context.md # Authoritative live-state reference
|
||||
├── refine_cluster.md # Known-issues runbook
|
||||
└── LOG.md # Append-only change journal
|
||||
```
|
||||
|
||||
## Access
|
||||
|
||||
### Discover Ingress LoadBalancer IP
|
||||
|
||||
```bash
|
||||
# Find the external IP assigned by Cilium LB-IPAM
|
||||
kubectl get svc -n ingress-nginx ingress-nginx
|
||||
# Example output:
|
||||
# LoadBalancer IP: 192.168.1.160 (Cilium LB-IPAM assignment)
|
||||
```
|
||||
|
||||
Add to `/etc/hosts` on every client machine (Mac/Linux):
|
||||
|
||||
```
|
||||
# WireGuard access (remote — via talos-cp-1)
|
||||
10.6.0.1 grafana.riotpiao.homelab.com authentik.riotpiao.homelab.com vault.riotpiao.homelab.com minio.riotpiao.homelab.com prometheus.riotpiao.homelab.com portainer.riotpiao.homelab.com longhorn.riotpiao.homelab.com loki.riotpiao.homelab.com forgejo.riotpiao.homelab.com
|
||||
|
||||
# LAN access (on the home network — use actual LoadBalancer IP from above)
|
||||
192.168.1.160 grafana.riotpiao.homelab.com authentik.riotpiao.homelab.com vault.riotpiao.homelab.com minio.riotpiao.homelab.com prometheus.riotpiao.homelab.com portainer.riotpiao.homelab.com longhorn.riotpiao.homelab.com loki.riotpiao.homelab.com forgejo.riotpiao.homelab.com
|
||||
```
|
||||
|
||||
**Note:** `192.168.1.160` is an example Cilium LB-IPAM assignment. Verify with `kubectl get svc -n ingress-nginx ingress-nginx`.
|
||||
|
||||
Then access services at:
|
||||
|
||||
| Service | URL | Credentials |
|
||||
|---------|-----|-------------|
|
||||
| Grafana | http://grafana.riotpiao.homelab.com | admin / `GRAFANA_ADMIN_PASSWORD` or Authentik SSO |
|
||||
| Authentik | http://authentik.riotpiao.homelab.com | akadmin / see `.env` |
|
||||
| Vault | http://vault.riotpiao.homelab.com | root token / see `setup_vault.sh` output |
|
||||
| MinIO console | http://minio.riotpiao.homelab.com | `MINIO_ROOT_USER` / `MINIO_ROOT_PASSWORD` |
|
||||
| Prometheus | http://prometheus.riotpiao.homelab.com | no auth |
|
||||
| Portainer | http://portainer.riotpiao.homelab.com | set on first visit |
|
||||
| Longhorn | http://longhorn.riotpiao.homelab.com | no auth |
|
||||
| Forgejo *(planned)* | https://forgejo.forge.riotpiao.homelab.com | `rock` / `FORGEJO_ADMIN_PASSWORD`, or Authentik SSO |
|
||||
| Argo CD *(planned)* | `kubectl port-forward -n argocd svc/argocd-server 8080:443` | Authentik SSO (admins only) |
|
||||
|
||||
Grafana → "Homelab" folder has the operator dashboards (sidecar-loaded from `k8s/monitoring/dashboards/`, no restart needed on change):
|
||||
- **Service Availability & Certificate Expiration** — uptime probes + cert-manager expiry
|
||||
- **Latency & Golden Signals** — ingress request rate/error %/p50-p99 latency
|
||||
- **Kube-Controller Health** — API server RED metrics + kube-state-metrics controller-health signals
|
||||
- **Hardware Statistics** — per-node CPU/mem/disk/network/load
|
||||
- **Service Internals** — per-service deep-dive (MinIO/Forgejo/Argo CD/Vault/Longhorn)
|
||||
|
||||
## Deploy
|
||||
|
||||
```bash
|
||||
# 1. Install helmfile (once)
|
||||
brew install helmfile
|
||||
|
||||
# 2. Set credentials
|
||||
cp .env.example .env
|
||||
# edit .env with your passwords
|
||||
|
||||
# 3. Deploy everything
|
||||
helmfile apply
|
||||
|
||||
# Deploy a single stack
|
||||
helmfile apply -l namespace=logging
|
||||
helmfile apply -l name=grafana
|
||||
helmfile apply -l namespace=ingress-nginx
|
||||
|
||||
# Preview changes before applying
|
||||
helmfile diff
|
||||
```
|
||||
|
||||
## Bootstrap Order (fresh cluster)
|
||||
|
||||
```
|
||||
1. Provision nodes: talosctl apply-config (make apply-cp / apply-worker-new)
|
||||
2. Bootstrap Kubernetes: talosctl bootstrap
|
||||
3. Install Cilium: helm install cilium -f cluster-config/cilium-values.yaml
|
||||
4. Install Longhorn: bash cluster-config/longhorn_bootstrap.sh
|
||||
5. Deploy everything else: helmfile apply
|
||||
6. Vault init (one-time): bash k8s/talos-iam/setup_vault.sh
|
||||
7. Authentik OIDC: python3 k8s/talos-iam/provision_oidc.py
|
||||
8. Label worker: kubectl label node talos-worker-1 node-role.kubernetes.io/worker=
|
||||
|
||||
# ── CI/CD (planned — run after step 8) ─────────────────────────────────────
|
||||
9. Private CA + TLS: see k8s/talos-ci-cd/talos_version_control.html §12 Block 0
|
||||
10. Deploy Forgejo + runner: helmfile apply -l name=forgejo && helmfile apply -l name=forgejo-runner
|
||||
11. Authentik SSO for CI: Forgejo + Argo CD OIDC (§12 Block 1.5)
|
||||
12. Talos node CA trust: talosctl patch machineconfig (§12 Block 2)
|
||||
13. Deploy Argo CD: helmfile apply -l name=argocd (§12 Block 4)
|
||||
14. Wire deploy repo: argocd app create + push first manifests (§12 Block 4)
|
||||
```
|
||||
|
||||
## IAM & Auth Flow
|
||||
|
||||
Authentik is the central OIDC identity provider. Vault stores secrets and delegates authentication back to Authentik.
|
||||
|
||||
```
|
||||
User / core-cli
|
||||
│
|
||||
│ OAuth2 / OIDC
|
||||
▼
|
||||
Authentik (authentik.riotpiao.homelab.com)
|
||||
├── grafana app → Grafana OIDC login (group → Admin/Viewer role)
|
||||
├── minio app → MinIO OIDC login (group → readwrite/readonly policy)
|
||||
├── vault-browser → Vault UI OIDC login / `vault login -method=oidc`
|
||||
└── core-cli-shell → CLI device code flow (public client, no secret)
|
||||
│
|
||||
│ JWKS endpoint for JWT validation
|
||||
▼
|
||||
HashiCorp Vault (vault.riotpiao.homelab.com)
|
||||
├── auth/jwt — core-cli authenticates with device code JWT
|
||||
├── auth/oidc — browser/UI login via Authentik
|
||||
└── secret/ — KV v2: mcp/*, cluster/*, cloud/*
|
||||
```
|
||||
|
||||
**talos-cli device code login:**
|
||||
```bash
|
||||
core secrets login # prints URL + code → approve in browser → Vault token cached
|
||||
core put cluster/DUCKDNS_TOKEN DUCKDNS_TOKEN="abc" # field name = variable name, never `value`
|
||||
```
|
||||
|
||||
**One-time IAM setup (after `helmfile apply`):**
|
||||
```bash
|
||||
# 1. Provision OIDC apps and groups in Authentik
|
||||
GRAFANA_URL=http://grafana.riotpiao.homelab.com \
|
||||
MINIO_URL=http://minio.riotpiao.homelab.com \
|
||||
python3 k8s/talos-iam/provision_oidc.py
|
||||
|
||||
# 2. Init Vault, wire JWT + OIDC auth, seed secrets
|
||||
bash k8s/talos-iam/setup_vault.sh
|
||||
```
|
||||
|
||||
**CoreDNS hostname rewrites** (`k8s/coredns/coredns-configmap.yaml`) ensure in-cluster pods (Grafana, Vault, Forgejo runner, Argo CD) resolve internal hostnames to cluster services, avoiding hairpin NAT through LB IPs.
|
||||
|
||||
## CI/CD Pipeline *(planned)*
|
||||
|
||||
> **Implementation plan & exact build commands:** [`k8s/talos-ci-cd/talos_version_control.html`](k8s/talos-ci-cd/talos_version_control.html)
|
||||
|
||||
### Components
|
||||
|
||||
| Component | Helm chart | Namespace | Notes |
|
||||
|-----------|-----------|-----------|-------|
|
||||
| Forgejo | `gitea-charts/gitea` (Forgejo image override) | `forge` | Git + OCI registry + Actions engine; SQLite on Longhorn PVC; `strategy: Recreate` |
|
||||
| Forgejo runner | local chart `charts/forgejo-runner` | `cicd` | DinD sidecar; PodSecurity privileged; NetworkPolicy fenced |
|
||||
| Argo CD | `argo/argo-cd` | `argocd` | Pull-based CD; single replica; no external ingress (port-forward only) |
|
||||
|
||||
### Pipeline flow
|
||||
|
||||
```
|
||||
Developer
|
||||
│
|
||||
│ git push
|
||||
▼
|
||||
Forgejo (forge ns) ─── webhook ───► Runner (cicd ns)
|
||||
│ │
|
||||
│ SSO login (Authentik OIDC) │ ACTIONS_RUNTIME_TOKEN → git checkout
|
||||
▼ │ ci-registry-token (ci-bot) → docker push → Forgejo OCI
|
||||
Forgejo UI / Argo CD UI │ ci-deploy-token (ci-bot) → git commit → rock/deploy
|
||||
│
|
||||
Forgejo (rock/deploy repo) ◄─────────────┘
|
||||
│
|
||||
│ argocd-bot token (repo:read, poll every 3 min)
|
||||
▼
|
||||
Argo CD (argocd ns)
|
||||
│
|
||||
│ kubectl apply (cluster-admin ServiceAccount — never in CI)
|
||||
▼
|
||||
K8s workloads (images from Forgejo OCI)
|
||||
```
|
||||
|
||||
### Key security decisions
|
||||
|
||||
- **No kubeconfig in CI.** The runner can only push commits + OCI images. Argo CD bridges the gap autonomously.
|
||||
- **Scoped machine credentials.** `ci-bot` tokens are narrowly scoped: `package:write` for OCI, `repo:write` on `rock/deploy` only. A compromised runner cannot read other repos or call the K8s API.
|
||||
- **Private CA TLS.** Forgejo self-terminates HTTPS with a homelab CA (EC P-256, 10-year). The CA cert is distributed to Talos nodes via `machineconfig` patch and to the runner via K8s Secret. `ca.key` never enters the cluster.
|
||||
- **Authentik SSO for humans.** All interactive logins (Forgejo UI, Argo CD UI) route through Authentik. `homelab-admins` group → Forgejo admin + Argo CD `role:admin`; `homelab-devs` group → Forgejo user + no Argo CD access.
|
||||
- **Forgejo LB IP pinned.** Cilium LB-IPAM annotation `io.cilium/lb-ipam-ips: <LB_IP>` fixes the Forgejo LoadBalancer IP so the TLS SAN and DNS entries never need updating. Configure in `k8s/talos-ci-cd/forgejo-values.yaml`.
|
||||
|
||||
### Workflow file location
|
||||
|
||||
Forgejo Actions uses GitHub Actions syntax. Workflow files live in `.forgejo/workflows/` in each source repo:
|
||||
|
||||
```
|
||||
rock/source/
|
||||
└── .forgejo/
|
||||
└── workflows/
|
||||
├── ci.yml # build + test + push OCI image
|
||||
└── cd.yml # on: push to main → bump image tag in rock/deploy
|
||||
```
|
||||
|
||||
## Log Data Flow
|
||||
|
||||
```
|
||||
Pods / Talos journal (both nodes)
|
||||
│
|
||||
Promtail (DaemonSet, all nodes) reads /var/log/pods + /var/log/journal
|
||||
│
|
||||
▼
|
||||
Loki (logging ns) indexes + compacts, 10-day retention
|
||||
│ stores chunks via S3
|
||||
▼
|
||||
MinIO frontend service minio.storage.svc.cluster.local:9000
|
||||
(active-active, round-robin)
|
||||
│
|
||||
minio-az-a (cp-1) ↔ minio-az-b (worker-1) ↔ minio-az-c (worker-2)
|
||||
3-way site replication (bidirectional, automatic)
|
||||
│
|
||||
Grafana (logging ns) queries Loki + Prometheus via dashboards
|
||||
│
|
||||
Nginx Ingress → grafana.riotpiao.homelab.com browser access
|
||||
```
|
||||
|
||||
## Example Applications & Workloads
|
||||
|
||||
This cluster runs production-like applications and infrastructure services:
|
||||
|
||||
### Story Crater Backend
|
||||
**Type:** Distributed message-driven application
|
||||
**Namespace:** `story-crater-backend`
|
||||
**Architecture:**
|
||||
- gRPC server + REST gateway (Envoy)
|
||||
- PostgreSQL database (story_crater in CloudNativePG cluster)
|
||||
- Kafka topic consumers (via kmsvc message queue)
|
||||
- Temporal workflow integration for long-running operations
|
||||
- Prometheus metrics (processed messages, latency, errors)
|
||||
- Grafana dashboard (message throughput, queue depth, LLM token usage)
|
||||
|
||||
**Typical flow:**
|
||||
```
|
||||
POST /api/messages → gRPC handler
|
||||
↓
|
||||
Publish to Kafka topic (Story Crater queue)
|
||||
↓
|
||||
Message consumer processes async (may trigger LLM inference)
|
||||
↓
|
||||
Results stored in PostgreSQL + emitted as event
|
||||
↓
|
||||
Prometheus increments counters (story_crater_messages_handled_total)
|
||||
↓
|
||||
Grafana renders message throughput + duration histograms
|
||||
```
|
||||
|
||||
### Infrastructure Services (Essential)
|
||||
|
||||
| Service | Purpose | Namespace | Example Use |
|
||||
|---------|---------|-----------|-------------|
|
||||
| **Authentik** | OIDC identity provider | iam | User login, group management, SSO for Grafana/MinIO/Forgejo |
|
||||
| **Vault** | Secrets backend | iam | Database passwords, API keys, JWT token validation |
|
||||
| **CloudNativePG** | PostgreSQL 3-replica cluster | ddb | Auth database (Authentik), app database (Story Crater) |
|
||||
| **Loki** | Log aggregation | logging | Centralize pod logs, Talos kernel logs (10-day retention) |
|
||||
| **Prometheus** | Metrics collection | monitoring | Scrape kube-state-metrics, kubelet, ServiceMonitors (every 30s) |
|
||||
| **Grafana** | Observability dashboards | logging | Query Prometheus + Loki, alert on latency/error spikes |
|
||||
| **Kafka + kmsvc** | Message queue | sqs | Decouple services, async job processing, at-least-once delivery |
|
||||
| **MinIO** | S3-compatible object store | storage | Loki log chunks backend, Vault unseal keys, config backups |
|
||||
| **Longhorn** | Persistent block storage | longhorn-system | All PVCs (Postgres replicas, Kafka broker disks, MinIO) |
|
||||
|
||||
### Development Services (Optional)
|
||||
|
||||
| Service | Purpose | Namespace | Example Use |
|
||||
|---------|---------|-----------|-------------|
|
||||
| **Forgejo** | Self-hosted git + OCI registry | cicd | Version control, CI Actions runner, private Docker images |
|
||||
| **Argo CD** | Pull-based GitOps | cicd | Continuous deployment (deployment repo → Kubernetes) |
|
||||
| **Temporal** | Workflow orchestration | temporal | Schedule long-running jobs, retry logic, state machines |
|
||||
| **Portainer** | Container management UI | dashboard | Pod inspection, image management, quick debugging |
|
||||
|
||||
### Monitoring Example: Dashboard Walk-Through
|
||||
|
||||
Open **Grafana** → **Homelab** folder → **"Story Crater Backend — Service Overview"**:
|
||||
|
||||
**Row A — Availability & Golden Signals**
|
||||
- Requests/sec (blue = success, red = errors)
|
||||
- Error rate % (goal: < 0.1%)
|
||||
- p50/p95/p99 latency (goal: p99 < 500ms)
|
||||
- Alert: If error rate > 1% for 5 min, page on-call
|
||||
|
||||
**Row B — Resource Usage**
|
||||
- CPU (request/limit)
|
||||
- Memory (request/limit)
|
||||
- Restart count (goal: 0; alerts if > 2)
|
||||
|
||||
**Row C — Domain-Specific Metrics** (Story Crater only)
|
||||
- Messages processed/sec (bucketed by status: success, dlq, retry)
|
||||
- Queue depth (Kafka partitions lag)
|
||||
- Dedup window retention (FIFO redelivery tracking)
|
||||
- LLM inference tokens used/sec
|
||||
- External API call latency (e.g., OpenAI)
|
||||
|
||||
**Row D — Logs**
|
||||
- Live Loki panel: filter by pod + search for errors
|
||||
- Example: `{namespace="story-crater-backend"} | json | level="error"`
|
||||
|
||||
**Row E — Alert Status** (if SLO defined)
|
||||
- Burn rate (if consuming SLO budget)
|
||||
- Example: "30-day availability SLO = 99.5%; current burn rate = 0.2x"
|
||||
|
||||
**Row F — Related Dashboards**
|
||||
- Link to Kafka dashboard (queue depth)
|
||||
- Link to PostgreSQL dashboard (story_crater DB)
|
||||
- Link to Temporal dashboard (workflow execution times)
|
||||
|
||||
## Adding Hardware to the Cluster
|
||||
|
||||
### Step 1 — Get the Talos image
|
||||
|
||||
The image must include the same extensions as the existing nodes (`iscsi-tools` + `util-linux-tools`).
|
||||
Download from the Image Factory using the cluster's schematic ID:
|
||||
|
||||
| Format | Use case | URL |
|
||||
|--------|----------|-----|
|
||||
| ISO | USB boot (recommended for bare metal) | `https://factory.talos.dev/image/613e1592.../v1.13.3/metal-amd64.iso` |
|
||||
| RAW disk image | Write directly to drive via another machine | `https://factory.talos.dev/image/613e1592.../v1.13.3/metal-amd64.raw.xz` |
|
||||
| PXE / iPXE | Network boot — no USB needed | `https://factory.talos.dev/image/613e1592.../v1.13.3/kernel-amd64` |
|
||||
|
||||
> Full schematic ID: `613e1592b2da41ae5e265e8789429f22e121aab91cb4deb6bc3c0b6262961245`
|
||||
|
||||
---
|
||||
|
||||
### Option A — USB Bootable (recommended)
|
||||
|
||||
```bash
|
||||
curl -Lo talos-worker.iso \
|
||||
"https://factory.talos.dev/image/613e1592b2da41ae5e265e8789429f22e121aab91cb4deb6bc3c0b6262961245/v1.13.3/metal-amd64.iso"
|
||||
sudo dd if=talos-worker.iso of=/dev/sdX bs=4M status=progress && sync
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Option B — In-Memory (diskless / RAM boot via PXE)
|
||||
|
||||
Talos runs entirely from RAM. Useful for temporary nodes or hardware where you don't want to touch the existing OS.
|
||||
|
||||
```bash
|
||||
# Boot via PXE pointing to:
|
||||
# Kernel: https://factory.talos.dev/image/613e1592.../v1.13.3/kernel-amd64
|
||||
# Initrd: https://factory.talos.dev/image/613e1592.../v1.13.3/initramfs-amd64.xz
|
||||
# Cmdline: talos.platform=metal
|
||||
```
|
||||
|
||||
> Note: in-memory nodes lose state on reboot. Not suitable for Longhorn storage nodes.
|
||||
|
||||
---
|
||||
|
||||
### Option C — Direct Disk Image (headless / remote)
|
||||
|
||||
```bash
|
||||
xz -d talos-worker.raw.xz
|
||||
sudo dd if=talos-worker.raw of=/dev/sda bs=4M status=progress && sync
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Step 2 — Discover hardware in maintenance mode
|
||||
|
||||
```bash
|
||||
# Scan your LAN for the new node in maintenance mode
|
||||
nmap -sn 192.168.1.0/24
|
||||
# Note: Replace with your actual subnet (e.g., 10.0.1.0/24)
|
||||
|
||||
# Discover available disks on the node
|
||||
talosctl --nodes <maintenance-ip> --talosconfig cluster-config/talosconfig disks --insecure
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Step 3 — Prepare the worker config
|
||||
|
||||
```bash
|
||||
cp cluster-config/worker-1.yaml cluster-config/worker-N.yaml
|
||||
```
|
||||
|
||||
Edit exactly these four fields:
|
||||
|
||||
| Field | Value |
|
||||
|-------|-------|
|
||||
| `machine.network.hostname` | `talos-worker-N` |
|
||||
| `machine.network.interfaces[0].addresses` | `192.168.1.16N/24` |
|
||||
| `machine.install.disk` | disk path from Step 2 |
|
||||
| `machine.nodeLabels.topology.kubernetes.io/zone` | `az-N` |
|
||||
|
||||
---
|
||||
|
||||
### Step 4 — Apply config
|
||||
|
||||
```bash
|
||||
make apply-worker-new N=<num> WN_IP=<maintenance-ip>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Step 5 — Persist the worker IP
|
||||
|
||||
```bash
|
||||
# Substitute <WN_MAINTENANCE_IP> with the actual IP discovered in Step 2
|
||||
echo 'export W<N>_IP=192.168.1.<last-octet>' >> ~/.zshrc && source ~/.zshrc
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Step 6 — Set the node-role label
|
||||
|
||||
```bash
|
||||
kubectl label node talos-worker-N node-role.kubernetes.io/worker=
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Step 7 — Verify
|
||||
|
||||
```bash
|
||||
kubectl get nodes -w
|
||||
kubectl describe node talos-worker-N | grep -A10 Labels
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Step 8 — What automatically extends to the new node
|
||||
|
||||
| Service | Behaviour |
|
||||
|---------|-----------|
|
||||
| Cilium | New pod scheduled automatically |
|
||||
| Promtail | DaemonSet — starts immediately |
|
||||
| Nginx Ingress | DaemonSet — starts immediately, port 80/443 available on new node |
|
||||
| Longhorn | Detects new node, available for replica scheduling |
|
||||
| MinIO / Loki / Grafana | Stay on existing node (single-replica Deployments) |
|
||||
@@ -0,0 +1,630 @@
|
||||
|
||||
|
||||
# TROUBLESHOOTING.md
|
||||
|
||||
SRE Agent Troubleshooting Guide — Kubernetes / Homelab Cluster Production Incidents
|
||||
|
||||
You are an SRE agent responding to production incidents on this Kubernetes cluster. You follow a strict diagnostic methodology before taking any action. You never jump to tools before establishing the failure boundary. You never guess. You reason from evidence.
|
||||
|
||||
---
|
||||
|
||||
## RULE 0 — PRE-FLIGHT BEFORE EVERY INCIDENT
|
||||
|
||||
Before running any command, answer these three questions out loud:
|
||||
|
||||
```
|
||||
1. WHERE is the failure boundary?
|
||||
Client / Network / Pod / Controller / Infrastructure?
|
||||
|
||||
2. ALL traffic or SOME traffic?
|
||||
Complete outage = systemic. Intermittent = partial failure.
|
||||
This changes everything about where you look.
|
||||
|
||||
3. WHAT changed recently?
|
||||
Deploy / Config / Certificate renewal / Traffic spike / GitOps pipeline?
|
||||
Correlate with metrics timeline before acting.
|
||||
```
|
||||
|
||||
If you cannot answer all three — gather more information before proceeding.
|
||||
|
||||
---
|
||||
|
||||
## RULE 1 — NEVER TALK TO ETCD DIRECTLY
|
||||
|
||||
Nothing in your runbook should ever reference etcd directly.
|
||||
All state reads and writes go through the API server.
|
||||
Controllers use informer cache via Watch streams — not polling.
|
||||
|
||||
---
|
||||
|
||||
## RULE 2 — LAYER BEFORE TOOL
|
||||
|
||||
Always identify which layer is broken before selecting a tool.
|
||||
|
||||
```
|
||||
Layer 1 — Control Plane
|
||||
API Server / etcd / Controllers / CRDs
|
||||
Broken when: reconciliation loops fail, RBAC denied,
|
||||
controller crashes, Watch stream collapses
|
||||
|
||||
Layer 2 — Kubelet
|
||||
Pod lifecycle / cgroups / tmpfs mounts / probes
|
||||
Broken when: OOMKilled, CrashLoopBackOff,
|
||||
probe misconfiguration, Secret not mounted
|
||||
|
||||
Layer 3 — Networking
|
||||
CoreDNS / kube-proxy / CNI / Ingress / Load Balancer
|
||||
Broken when: pods green but traffic failing,
|
||||
DNS timeouts, empty endpoints,
|
||||
NetworkPolicy drops, IP exhaustion
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## RULE 3 — NEVER DELETE A PVC WITHOUT REPLICATION
|
||||
|
||||
🔴 **A deleted PVC = permanent data loss.** Never delete a PVC unless you have verified replicas or backups exist.
|
||||
|
||||
**Before ANY PVC deletion:**
|
||||
|
||||
```bash
|
||||
# 1. Check volume replication status
|
||||
kubectl get pvc -n <ns> <pvc>
|
||||
kubectl get pv <pv-name> -o json | jq '.spec'
|
||||
|
||||
# 2. For Longhorn volumes (storage)
|
||||
kubectl get longhorn-volume -n longhorn-system -o wide
|
||||
# Must show: State=healthy, Replicas >= 2
|
||||
|
||||
# 3. For databases (PostgreSQL)
|
||||
kubectl exec -n ddb pod/ddb-cluster-0 -- \
|
||||
psql -U postgres -c "SELECT slot_name, restart_lsn FROM pg_replication_slots;"
|
||||
# Must show: at least 1 streaming replica
|
||||
|
||||
# 4. For backup buckets (MinIO)
|
||||
# Verify backup was taken in last 24 hours
|
||||
# kubectl exec -n storage pod/minio-0 -- mc ls local/postgresql-backups/
|
||||
```
|
||||
|
||||
**If replication is not confirmed:** STOP. Do not proceed. Escalate to SRE lead.
|
||||
|
||||
---
|
||||
|
||||
## PROCEDURE 1 — CrashLoopBackOff
|
||||
|
||||
```
|
||||
NEVER start with kubectl logs.
|
||||
|
||||
Step 1 — Establish crash type
|
||||
kubectl describe pod <pod> -n <ns>
|
||||
Read: Last State → Reason → Exit Code
|
||||
|
||||
Step 2 — Exit code triage
|
||||
0 → Clean exit. Check livenessProbe config.
|
||||
1 → App error. Now run: kubectl logs <pod> -n <ns> --previous
|
||||
137 → OOMKilled. Kernel cgroup enforced memory limit.
|
||||
Run: kubectl top pod + kubectl top node
|
||||
139 → Segfault. Check binary and dependencies.
|
||||
|
||||
Step 3 — If OOMKilled (137)
|
||||
Answer before raising limit:
|
||||
A. Sawtooth memory pattern = load spike. Raise limit with headroom.
|
||||
B. Monotonic growth = memory leak. Fix the code first.
|
||||
C. kubectl describe node → MemoryPressure: True = noisy neighbor.
|
||||
Move pod, don't raise limit.
|
||||
|
||||
Step 4 — Namespace events
|
||||
kubectl get events -n <ns> --sort-by='.lastTimestamp' | tail -20
|
||||
|
||||
CRITICAL: Always use --previous for crash logs.
|
||||
Without it you get logs from current instance
|
||||
which may have lived for 3 seconds.
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## PROCEDURE 2 — TLS Handshake Failures
|
||||
|
||||
```
|
||||
Certificate Ready: True does not mean traffic is working.
|
||||
Kubernetes state layer ≠ runtime process layer.
|
||||
|
||||
Step 1 — Check Kubernetes state
|
||||
kubectl get certificate -n <ns>
|
||||
kubectl describe certificate <name> -n <ns>
|
||||
kubectl get secret <tls-secret> -n <ns> -o yaml
|
||||
|
||||
Step 2 — Check file layer (kubelet-synced tmpfs)
|
||||
kubectl exec -it <pod> -- cat /etc/certs/tls.crt
|
||||
|
||||
Step 3 — Check what LIVE PROCESS is actually serving
|
||||
openssl s_client -connect <pod-ip>:<port> </dev/null 2>/dev/null \
|
||||
| openssl x509 -noout -dates
|
||||
Old expiry = process loaded cert at startup, never reloaded.
|
||||
This bypasses Kubernetes entirely. Use this always.
|
||||
|
||||
Step 4 — Check cert-manager controller
|
||||
kubectl logs -n cert-manager deploy/cert-manager | grep -E 'ERROR|certificate'
|
||||
|
||||
Step 5 — Validate RBAC
|
||||
kubectl auth can-i create secrets \
|
||||
--as=system:serviceaccount:cert-manager:cert-manager \
|
||||
--all-namespaces
|
||||
Returns no = RBAC broken. Found root cause.
|
||||
|
||||
Trace path:
|
||||
desired resource → controller action → Secret update → workload consumption
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## PROCEDURE 3 — Pods Running and Ready But Traffic Failing
|
||||
|
||||
```
|
||||
NEVER start with ingress-nginx logs.
|
||||
Pods showing Ready does not mean traffic is flowing.
|
||||
|
||||
Step 1 — Endpoints (always first)
|
||||
kubectl get endpoints <svc> -n <ns>
|
||||
Empty = label selector mismatch, wrong port, namespace issue.
|
||||
kubectl get pods -n <ns> --show-labels
|
||||
kubectl get svc <svc> -n <ns> -o yaml | grep selector
|
||||
|
||||
Step 2 — DNS resolution inside cluster
|
||||
kubectl exec -it <pod> -- \
|
||||
nslookup <svc>.<ns>.svc.cluster.local
|
||||
Failure here = CoreDNS problem.
|
||||
|
||||
Step 3 — NetworkPolicy (silent drops)
|
||||
kubectl get networkpolicy -n <ns>
|
||||
kubectl describe networkpolicy <name> -n <ns>
|
||||
NetworkPolicy drops packets with zero error in application logs.
|
||||
GitOps can accidentally strip ingress rules.
|
||||
|
||||
Step 4 — Direct connectivity test
|
||||
kubectl exec -it <pod> -- curl -v http://<svc>:<port>/healthz
|
||||
TCP reset = port or firewall issue.
|
||||
Timeout = packet dropping, CNI or NetworkPolicy.
|
||||
|
||||
Step 5 — Ingress (only after ruling out above)
|
||||
kubectl logs -n ingress-nginx <pod> | grep -E '504|502|499|reset'
|
||||
502 = upstream pod crashed.
|
||||
503 = no healthy endpoints.
|
||||
504 = upstream too slow.
|
||||
|
||||
Step 6 — Infrastructure layer
|
||||
Check cloud load balancer health checks
|
||||
|
||||
Status code reference:
|
||||
499 → client timeout
|
||||
502 → upstream crashed
|
||||
503 → no endpoints
|
||||
504 → upstream slow
|
||||
TCP Reset → wrong port / NetworkPolicy / firewall
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## PROCEDURE 4 — Post-Rollout Latency Spike
|
||||
|
||||
```
|
||||
No crashes. No OOMKill. No failed pods. Just slow.
|
||||
Requires TWO timelines simultaneously.
|
||||
|
||||
Step 1 — Correlate timelines
|
||||
Timeline A: kubectl get events -n <ns> --sort-by='.lastTimestamp'
|
||||
Timeline B: Prometheus/metrics p99 latency graph
|
||||
If aligned = rollout caused it.
|
||||
If latency spiked before rollout completed = new code is the problem.
|
||||
|
||||
Step 2 — Decision rule
|
||||
Latency improving over time = cold start.
|
||||
Fix: readiness probe at /readyz that validates cache warmth.
|
||||
Not /healthz which only checks process is alive.
|
||||
|
||||
Latency stable and high = code regression.
|
||||
Fix: kubectl rollout undo deployment/<name> -n <ns>
|
||||
Investigate new version offline.
|
||||
|
||||
Step 3 — Five causes in order of likelihood
|
||||
1. Cold start — JVM/cache/connection pool not initialized
|
||||
2. Reduced capacity — maxUnavailable:1 during rollout
|
||||
3. Code regression — new version has performance bug
|
||||
4. Downstream saturation — database/cache throttling
|
||||
5. Connection draining race — preStop hook missing
|
||||
|
||||
Step 4 — Prevent connection drops on every rollout
|
||||
lifecycle:
|
||||
preStop:
|
||||
exec:
|
||||
command: ["sleep", "15"]
|
||||
Gives kube-proxy time to drain connections before SIGTERM.
|
||||
|
||||
Rollout safety config for live services:
|
||||
maxSurge: 1 → ceiling above desired — add first
|
||||
maxUnavailable: 0 → floor below desired — never reduce capacity
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## PROCEDURE 5 — RBAC Permission Decay
|
||||
|
||||
```
|
||||
Signal: controller returns 403 Forbidden from API server.
|
||||
Controller is ALIVE and working — it is being BLOCKED.
|
||||
This is not a connectivity issue.
|
||||
|
||||
Step 1 — Fastest confirmation
|
||||
kubectl auth can-i create secrets \
|
||||
--as=system:serviceaccount:cert-manager:cert-manager \
|
||||
--all-namespaces
|
||||
yes = RBAC fine, look elsewhere.
|
||||
no = RBAC broken, found root cause.
|
||||
|
||||
Step 2 — Diff working cluster vs broken cluster
|
||||
kubectl get clusterrolebinding <name> \
|
||||
-o yaml --context=us-cluster > /tmp/us-crb.yaml
|
||||
kubectl get clusterrolebinding <name> \
|
||||
-o yaml --context=eu-cluster > /tmp/eu-crb.yaml
|
||||
diff /tmp/us-crb.yaml /tmp/eu-crb.yaml
|
||||
|
||||
Step 3 — Three hypotheses
|
||||
A. ClusterRole/ClusterRoleBinding modified or deleted
|
||||
B. Scope changed from ClusterRoleBinding to RoleBinding
|
||||
C. ServiceAccount recreated — binding points to wrong subject
|
||||
|
||||
Step 4 — Immediate remediation (P1)
|
||||
kubectl apply -f <backup-rbac-config>
|
||||
kubectl auth can-i create secrets --as=<sa> --all-namespaces
|
||||
kubectl rollout restart deploy/<controller> -n <ns>
|
||||
kubectl get certificate -n <ns> -w
|
||||
|
||||
Step 5 — Systemic prevention
|
||||
A. Protect RBAC resources in GitOps with Prune=false
|
||||
B. CronJob every 15 minutes running kubectl auth can-i validation
|
||||
C. Pre-sync hook that blocks pipeline if RBAC check fails
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## PROCEDURE 6 — Liveness vs Readiness Probe Issues
|
||||
|
||||
```
|
||||
readinessProbe → gates traffic AND rolling update progression
|
||||
pod STAYS ALIVE if failing
|
||||
removed from endpoints
|
||||
frozen rollout = readinessProbe never passing
|
||||
|
||||
livenessProbe → gates pod SURVIVAL
|
||||
pod gets KILLED and restarted if failing
|
||||
high restart count = livenessProbe too aggressive
|
||||
|
||||
Identify from evidence alone — no commands needed:
|
||||
Frozen rollout + pods Running + 0 restarts = readinessProbe
|
||||
High restart count + exit code 1 + Running = livenessProbe
|
||||
|
||||
Step 1 — Check which container is failing
|
||||
kubectl describe pod <pod> -n <ns>
|
||||
Read Containers section — which container shows Ready: False?
|
||||
Check Events section — exact probe failure message.
|
||||
|
||||
Step 2 — Application vs platform containers
|
||||
1/2 Ready = one container passing, one failing.
|
||||
Platform-injected sidecars fail independently
|
||||
of the application container.
|
||||
Isolate the failing container before troubleshooting.
|
||||
|
||||
Step 3 — Fix probe timing
|
||||
Replace initialDelaySeconds with startupProbe:
|
||||
startupProbe:
|
||||
httpGet:
|
||||
path: /healthz/ready
|
||||
port: 8080
|
||||
failureThreshold: 30
|
||||
periodSeconds: 2
|
||||
Polls every 2s up to 60s. Passes the moment container is ready.
|
||||
No unnecessary fixed wait on every rollout.
|
||||
|
||||
Step 4 — Fix livenessProbe aggression
|
||||
livenessProbe:
|
||||
httpGet:
|
||||
path: /healthz
|
||||
port: 8080
|
||||
timeoutSeconds: 5 # was 1 — give app time to respond
|
||||
periodSeconds: 10 # was 5 — less frequent
|
||||
failureThreshold: 3
|
||||
|
||||
CRITICAL: /healthz must NEVER query external dependencies.
|
||||
Wrong: /healthz checks database connectivity
|
||||
Wrong: /healthz checks Redis connection
|
||||
Right: /healthz returns 200 if process is alive — nothing else
|
||||
External dependency checks belong in /readyz only.
|
||||
|
||||
Step 5 — Immediate remediation for frozen rollout
|
||||
kubectl rollout undo deployment/<name> -n <ns>
|
||||
kubectl rollout status deployment/<name> -n <ns>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## PROCEDURE 7 — Control Plane Component Failures
|
||||
|
||||
```
|
||||
CoreDNS down
|
||||
Signal: intermittent service-to-service failures, no pod errors
|
||||
Check: kubectl get pods -n kube-system | grep coredns
|
||||
kubectl logs -n kube-system deploy/coredns
|
||||
kubectl exec -it <pod> -- nslookup <svc>.<ns>.svc.cluster.local
|
||||
Metric: coredns_dns_request_duration_seconds p99 > 100ms
|
||||
|
||||
CNI issue (Cilium)
|
||||
Signal: pods stuck in ContainerCreating, "failed to assign IP"
|
||||
Check: kubectl describe pod <stuck-pod> | grep -A10 Events
|
||||
Check: kubectl get pods -n kube-system | grep cilium
|
||||
Metric: cilium_endpoint_creation_errors
|
||||
|
||||
kube-proxy stale rules (or Cilium networking issues)
|
||||
Signal: new services unreachable from specific nodes only
|
||||
Check: kubectl logs -n kube-system <kube-proxy-or-cilium-pod>
|
||||
Metric: kubeproxy_sync_proxy_rules_duration_seconds spike
|
||||
|
||||
Controller reconciliation loop stuck
|
||||
Signal: RBAC errors or Watch stream failures
|
||||
Check: kubectl logs -n <ns> <controller-pod>
|
||||
Look for: "forbidden", "Watch", "timeout"
|
||||
|
||||
Longhorn (storage) issues
|
||||
Signal: PVC stuck Pending, pods can't mount volumes
|
||||
Check: kubectl get pvc -A
|
||||
Check: kubectl describe pvc <name> -n <ns>
|
||||
Check: kubectl get longhorn-nodes -n longhorn-system
|
||||
Metric: longhorn_disk_capacity / longhorn_disk_reservation
|
||||
|
||||
MinIO (object storage) issues
|
||||
Signal: Loki unable to write logs, pods crash
|
||||
Check: kubectl logs -n storage deploy/minio
|
||||
Check: kubectl exec -it <minio-pod> -- mc ls storage/
|
||||
Verify: site replication status between az-a and az-b
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## PROCEDURE 8 — Certificate Expiry Incidents
|
||||
|
||||
```
|
||||
Signal: TLS handshake failures or "certificate expired" errors
|
||||
This should NEVER happen — cert-manager automates renewal 30 days early.
|
||||
|
||||
Step 1 — Check cert-manager is running
|
||||
kubectl get pods -n cert-manager
|
||||
kubectl logs -n cert-manager deploy/cert-manager | grep ERROR
|
||||
|
||||
Step 2 — Check Certificate resources
|
||||
kubectl get certificate -A
|
||||
kubectl describe certificate <name> -n <ns>
|
||||
Look for: Ready: False, "renewal" in status
|
||||
|
||||
Step 3 — Check Secret exists and contains cert
|
||||
kubectl get secret <tls-secret> -n <ns> -o yaml | grep tls.crt
|
||||
Decode and verify expiry:
|
||||
kubectl get secret <tls-secret> -n <ns> -o jsonpath='{.data.tls\.crt}' \
|
||||
| base64 -d | openssl x509 -noout -dates
|
||||
|
||||
Step 4 — Check live process cert (most important)
|
||||
openssl s_client -connect <pod-ip>:<port> </dev/null 2>/dev/null \
|
||||
| openssl x509 -noout -dates
|
||||
If old expiry here = process loaded cert at startup, never reloaded.
|
||||
Restart pod: kubectl rollout restart deploy/<name> -n <ns>
|
||||
|
||||
Step 5 — Check RBAC for cert-manager
|
||||
kubectl auth can-i create secrets \
|
||||
--as=system:serviceaccount:cert-manager:cert-manager \
|
||||
--all-namespaces
|
||||
|
||||
Step 6 — Emergency remediation (if cert truly expired)
|
||||
kubectl rollout restart deploy/cert-manager -n cert-manager
|
||||
kubectl delete certificate <name> -n <ns>
|
||||
kubectl apply -f <certificate-yaml>
|
||||
kubectl rollout restart deploy/<dependent-app> -n <ns>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## PROCEDURE 9 — Vault / Authentik IAM Issues
|
||||
|
||||
```
|
||||
Signal: Services can't authenticate, OIDC login fails, Vault sealed
|
||||
|
||||
Step 1 — Check Vault status
|
||||
kubectl get pods -n iam | grep vault
|
||||
kubectl logs -n iam deploy/vault
|
||||
kubectl exec -it <vault-pod> -n iam -- vault status
|
||||
|
||||
Step 2 — Check if Vault is sealed
|
||||
kubectl exec -it <vault-pod> -n iam -- vault status | grep Sealed
|
||||
If Sealed: true → requires unseal keys (see bootstrap docs)
|
||||
|
||||
Step 3 — Check Authentik
|
||||
kubectl get pods -n iam | grep authentik
|
||||
kubectl logs -n iam deploy/authentik-server
|
||||
kubectl describe statefulset authentik-postgresql -n iam
|
||||
|
||||
Step 4 — Check secret in Vault
|
||||
kubectl exec -it <vault-pod> -n iam -- \
|
||||
vault kv get cluster/VARIABLE_NAME
|
||||
Not found = secret never created (run setup_vault.sh)
|
||||
|
||||
Step 5 — Check OIDC provisioning
|
||||
kubectl logs -n iam job/oidc-provisioning
|
||||
If errors = run provision_oidc.py again
|
||||
|
||||
Step 6 — Verify Loki/MinIO can connect to Vault
|
||||
Check: helmfile values reference correct Vault endpoint
|
||||
Check: ServiceAccount token mounted and RBAC permitting auth
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## PROCEDURE 10 — Loki / Logging Pipeline Stalled
|
||||
|
||||
```
|
||||
Signal: No logs appearing in Grafana, Loki backend growing without limit
|
||||
|
||||
Step 1 — Check Promtail is scraping
|
||||
kubectl get pods -n logging | grep promtail
|
||||
kubectl logs -n logging ds/promtail | grep -E 'scraping|error'
|
||||
Should show: scraping every few seconds from all nodes
|
||||
|
||||
Step 2 — Check Loki can write to MinIO
|
||||
kubectl logs -n logging deploy/loki
|
||||
Look for: S3 errors, "connection refused", "write: no space"
|
||||
|
||||
Step 3 — Check MinIO is operational
|
||||
kubectl get pods -n storage | grep minio
|
||||
kubectl logs -n storage deploy/minio
|
||||
Check site replication status:
|
||||
kubectl get job -n storage | grep replication
|
||||
|
||||
Step 4 — Check disk space
|
||||
kubectl top pod -n logging
|
||||
kubectl exec -it <loki-pod> -n logging -- df -h
|
||||
If full = minio-backed storage exhausted, purge old chunks
|
||||
|
||||
Step 5 — Check PVC for Loki index
|
||||
kubectl get pvc -n logging
|
||||
kubectl describe pvc <loki-pvc> -n logging
|
||||
Bound to Longhorn PV = check node storage
|
||||
|
||||
Step 6 — Restart Loki
|
||||
kubectl rollout restart deploy/loki -n logging
|
||||
kubectl get pods -n logging -w
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Quick Command Reference
|
||||
|
||||
```bash
|
||||
# Pod state and debugging
|
||||
kubectl describe pod <pod> -n <ns>
|
||||
kubectl logs <pod> -n <ns> --previous
|
||||
kubectl get events -n <ns> --sort-by='.lastTimestamp' | tail -20
|
||||
kubectl exec -it <pod> -n <ns> -- /bin/sh
|
||||
|
||||
# Resource usage
|
||||
kubectl top pod <pod> -n <ns>
|
||||
kubectl top node
|
||||
kubectl get pvc -A
|
||||
|
||||
# Networking and service discovery
|
||||
kubectl get endpoints <svc> -n <ns>
|
||||
kubectl get svc <svc> -n <ns> -o yaml
|
||||
kubectl get pods -n <ns> --show-labels
|
||||
kubectl get networkpolicy -n <ns>
|
||||
kubectl exec -it <pod> -n <ns> -- nslookup <svc>.<ns>.svc.cluster.local
|
||||
kubectl exec -it <pod> -n <ns> -- curl -v http://<svc>:<port>/healthz
|
||||
|
||||
# TLS inspection (live process — bypasses Kubernetes)
|
||||
openssl s_client -connect <ip>:<port> </dev/null 2>/dev/null \
|
||||
| openssl x509 -noout -dates
|
||||
|
||||
# RBAC validation
|
||||
kubectl auth can-i <verb> <resource> \
|
||||
--as=system:serviceaccount:<ns>:<sa> --all-namespaces
|
||||
|
||||
# Rollout management
|
||||
kubectl rollout status deployment/<name> -n <ns>
|
||||
kubectl rollout undo deployment/<name> -n <ns>
|
||||
kubectl rollout restart deploy/<name> -n <ns>
|
||||
|
||||
# Certificates and TLS
|
||||
kubectl get certificate -n <ns>
|
||||
kubectl describe certificate <name> -n <ns>
|
||||
kubectl get secret <tls-secret> -n <ns> -o yaml
|
||||
kubectl logs -n cert-manager deploy/cert-manager | grep ERROR
|
||||
|
||||
# Control plane components
|
||||
kubectl get pods -n kube-system
|
||||
kubectl get pods -n cert-manager
|
||||
kubectl logs -n kube-system deploy/coredns
|
||||
kubectl describe node <node-name>
|
||||
|
||||
# Storage
|
||||
kubectl get pvc -A
|
||||
kubectl get longhorn-nodes -n longhorn-system
|
||||
kubectl exec -it <longhorn-pod> -n longhorn-system -- longhorn node ls
|
||||
|
||||
# IAM and secrets
|
||||
kubectl get pods -n iam
|
||||
kubectl logs -n iam deploy/vault
|
||||
kubectl logs -n iam deploy/authentik-server
|
||||
kubectl exec -it <vault-pod> -n iam -- vault status
|
||||
kubectl exec -it <vault-pod> -n iam -- vault kv get cluster/<KEY>
|
||||
|
||||
# Using the k alias (add to ~/.zshrc)
|
||||
k get pods -A
|
||||
k logs -n logging <pod>
|
||||
k describe node talos-worker-1
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Agent Behaviour Rules
|
||||
|
||||
```
|
||||
1. Never skip the three pre-flight questions
|
||||
2. Never check logs before establishing exit code
|
||||
3. Never check ingress before checking endpoints
|
||||
4. Never raise memory limit before understanding growth pattern
|
||||
5. Never assume etcd talks directly to controllers
|
||||
6. Never conflate Kubernetes state with runtime process state
|
||||
7. Always use --previous for crash logs
|
||||
8. Always correlate two timelines for post-rollout issues
|
||||
9. Always diff working cluster vs broken cluster for RBAC issues
|
||||
10. Always confirm fix with kubectl auth can-i before closing incident
|
||||
11. Always check live process cert with openssl s_client, not just Kubernetes state
|
||||
12. Always verify Vault is unsealed and accessible before troubleshooting auth issues
|
||||
13. Always check control plane components in kube-system before application logs
|
||||
14. Always rule out networking (endpoints, DNS, NetworkPolicy) before app errors
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Prevention & Observability
|
||||
|
||||
```
|
||||
Set up proactive alerts:
|
||||
|
||||
1. RBAC validation (every 15 minutes)
|
||||
for each ServiceAccount in each namespace:
|
||||
kubectl auth can-i create secrets --as=<sa> --all-namespaces
|
||||
Alert if any returns "no"
|
||||
|
||||
2. Certificate expiry monitoring
|
||||
certmanager_certificate_expiration_seconds
|
||||
Alert 15 days before expiry
|
||||
|
||||
3. Controller forbidden errors (zero tolerance)
|
||||
apiserver_request_total{code="403", user=~"system:serviceaccount:.*"}
|
||||
Alert on > 0
|
||||
|
||||
4. Watch stream collapse
|
||||
apiserver_request_total{verb="LIST"} spike
|
||||
Alert on 3x baseline within 5 minutes
|
||||
|
||||
5. Probe failures
|
||||
rate(kubelet_started_pods_total{result="failed"}[5m])
|
||||
Alert on > 0.1 per 5 minutes
|
||||
|
||||
6. Storage exhaustion
|
||||
kubelet_volume_stats_used_bytes / kubelet_volume_stats_capacity_bytes
|
||||
Alert at 80% capacity
|
||||
|
||||
7. Control plane latency
|
||||
apiserver_request_duration_seconds_sum / apiserver_request_duration_seconds_count
|
||||
Alert when p99 > 1 second
|
||||
|
||||
8. Logging pipeline lag
|
||||
loki_logql_querieslatency_seconds
|
||||
Alert when > 5 seconds
|
||||
```
|
||||
@@ -0,0 +1,218 @@
|
||||
## Cluster Architecture at a Glance
|
||||
|
||||
**Homelab** is a 2-node bare-metal Kubernetes cluster deployed with Talos Linux, designed for self-hosted services, observability, and GitOps-ready CI/CD.
|
||||
|
||||
### Deployment Stack (18 Helm releases)
|
||||
|
||||
| Layer | Component | Namespace | Purpose |
|
||||
|-------|-----------|-----------|---------|
|
||||
| **OS & Networking** | Talos Linux v1.13.3 | — | Immutable, declarative Linux |
|
||||
| **CNI** | Cilium (eBPF) | kube-system | Advanced networking, no kube-proxy |
|
||||
| **Ingress** | Nginx Ingress Controller | ingress-nginx | Reverse proxy, TLS termination |
|
||||
| **Certificates** | cert-manager + homelab-ca | cert-manager | Self-signed CA, auto-renewal |
|
||||
| **Storage (Block)** | Longhorn v1.7.0 | longhorn-system | Persistent volumes, default StorageClass |
|
||||
| **Storage (Object)** | MinIO (3-node, site-repl) | storage | S3-compatible, multi-AZ replication |
|
||||
| **Database** | CloudNativePG (3 replicas) | ddb | PostgreSQL 16 + pgvector |
|
||||
| **IAM / OIDC** | Authentik | iam | Federated OIDC provider for all services |
|
||||
| **Secrets** | HashiCorp Vault | iam | KV secrets backend, JWT auth |
|
||||
| **Logs** | Loki (SingleBinary) | logging | 10-day retention, MinIO backend |
|
||||
| **Log Collection** | Promtail (DaemonSet) | logging | Pod + kernel logs → Loki |
|
||||
| **Metrics** | Prometheus + kube-state-metrics | monitoring | Time-series metrics, service discovery |
|
||||
| **Dashboards** | Grafana | logging | Unified UI for Prometheus + Loki |
|
||||
| **Uptime Monitoring** | Blackbox Exporter | monitoring | External endpoint probes |
|
||||
| **Git Forge** | Forgejo (self-hosted) | cicd | Git server, OCI registry, Forgejo Actions |
|
||||
| **CI Runner** | Forgejo Actions Runner | cicd | Privileged build pods, image push |
|
||||
| **CD** | Argo CD | cicd | Pull-based GitOps, declarative deployments |
|
||||
| **Message Queue** | Kafka (Strimzi KRaft mode) | sqs | SQS-like message queue service |
|
||||
| **Queue Redis** | Redis | sqs | In-flight message tracking, dedup |
|
||||
| **Queue API** | kmsvc (gRPC + REST) | sqs | Message service (SendMessage, ReceiveMessage, etc.) |
|
||||
| **Workflows** | Temporal | temporal | Distributed workflow engine |
|
||||
| **Container UI** | Portainer CE | dashboard | Pod/workload management UI |
|
||||
| **Terminal** | Claude Terminal | dev-tools | Persistent dev environment (optional) |
|
||||
| **Config Reload** | Reloader | reloader | Auto-reload pods on ConfigMap/Secret changes |
|
||||
|
||||
### Dependency Chain (Release Order)
|
||||
|
||||
```
|
||||
cert-manager (root)
|
||||
↓
|
||||
├─→ ingress-nginx
|
||||
├─→ cilium (network policies)
|
||||
└─→ cloudnative-pg
|
||||
↓
|
||||
├─→ authentik (DB: authentik)
|
||||
│ ├─→ vault (uses Authentik OIDC)
|
||||
│ ├─→ forgejo (OIDC login)
|
||||
│ ├─→ argocd (OIDC login)
|
||||
│ └─→ kmsvc (JWT auth)
|
||||
│
|
||||
└─→ story-crater-backend (DB: story_crater)
|
||||
|
||||
storage (MinIO 3-node)
|
||||
├─→ loki (object backend)
|
||||
├─→ vault (unseal keys bucket)
|
||||
└─→ monitoring (Prometheus scrape)
|
||||
|
||||
monitoring (Prometheus operator)
|
||||
├─→ ingress-nginx (requires ServiceMonitor CRDs)
|
||||
├─→ authentik (requires ServiceMonitor CRDs)
|
||||
├─→ minio-az-a (requires ServiceMonitor CRDs)
|
||||
└─→ temporal (requires ServiceMonitor CRDs)
|
||||
|
||||
logging (Loki + Grafana)
|
||||
├─→ promtail (pod log collection)
|
||||
└─→ grafana (dashboards)
|
||||
|
||||
sqs (Kafka + Message Queue)
|
||||
├─→ strimzi-operator
|
||||
├─→ kafka-cluster (KRaft mode, 3 brokers)
|
||||
├─→ kmsvc-redis (in-flight tracking)
|
||||
├─→ queue-crd (operator)
|
||||
└─→ management-service (gRPC/REST API)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Custom CLI — `talos`
|
||||
|
||||
Homelab cluster control CLI (`core/`). Manages cluster nodes and Vault secrets.
|
||||
|
||||
### Secret path convention
|
||||
|
||||
All secrets live under `cluster/<VARIABLE_NAME>`. The field name is always the variable name itself (SCREAMING_SNAKE_CASE), matching the `.env` key. Example paths:
|
||||
|
||||
```
|
||||
cluster/ANTHROPIC_API_KEY
|
||||
cluster/AUTHENTIK_FORGEJO_CLIENT_ID
|
||||
cluster/AUTHENTIK_ARGOCD_CLIENT_SECRET
|
||||
```
|
||||
|
||||
### `talos put` — write a secret to Vault
|
||||
|
||||
```bash
|
||||
talos put cluster/VARIABLE_NAME VARIABLE_NAME="secret-value"
|
||||
talos put cluster/FORGEJO_ADMIN_PASSWORD FORGEJO_ADMIN_PASSWORD="$FORGEJO_ADMIN_PASSWORD"
|
||||
```
|
||||
|
||||
Field name = variable name — never `value`.
|
||||
|
||||
### `talos get` — fetch a secret from Vault
|
||||
|
||||
```bash
|
||||
talos get cluster/VARIABLE_NAME --key VARIABLE_NAME # always specify --key
|
||||
talos get cluster/FORGEJO_ADMIN_PASSWORD --key FORGEJO_ADMIN_PASSWORD
|
||||
talos get cluster/VARIABLE_NAME --json # full secret as JSON
|
||||
```
|
||||
|
||||
Note: `talos get` uses `--key` (long flag), not a positional arg — unlike `talos secrets get`.
|
||||
|
||||
### `vsource` — load a `.env` into the shell
|
||||
|
||||
zsh function (lives in `~/.zshrc`, not in the repo — can reference but cannot run directly).
|
||||
Empty `.env` values are fetched from Vault at `cluster/<KEY>`; hardcoded values pass through.
|
||||
|
||||
```zsh
|
||||
vsource # loads .env in current directory
|
||||
vsource .env.local # loads a specific file
|
||||
```
|
||||
|
||||
`.env` format — leave secrets empty, vsource resolves them from Vault:
|
||||
|
||||
```bash
|
||||
ANTHROPIC_API_KEY= # fetched from cluster/ANTHROPIC_API_KEY
|
||||
AUTHENTIK_ARGOCD_CLIENT_ID= # fetched from cluster/AUTHENTIK_ARGOCD_CLIENT_ID
|
||||
DEBUG=true # hardcoded, passed through as-is
|
||||
```
|
||||
|
||||
### Typical workflow for a generated secret
|
||||
|
||||
```bash
|
||||
# 1. Store immediately after generation (keeps secrets out of shell history)
|
||||
talos put cluster/AUTHENTIK_FORGEJO_CLIENT_SECRET AUTHENTIK_FORGEJO_CLIENT_SECRET="<paste>"
|
||||
|
||||
# 2. Use via subshell when creating K8s secrets
|
||||
kubectl create secret generic my-secret \
|
||||
--from-literal=client-secret="$(talos get cluster/AUTHENTIK_FORGEJO_CLIENT_SECRET --key AUTHENTIK_FORGEJO_CLIENT_SECRET)"
|
||||
|
||||
# 3. Or load into shell via vsource for helmfile/env-driven tools
|
||||
vsource .env && helmfile apply
|
||||
```
|
||||
|
||||
### IAM Management (Federated OIDC, Phases 1–6 Complete)
|
||||
|
||||
**Status:** ✅ Fully deployed (2026-07-02). Single federated OIDC provider (`talos-federation`) handles all service auth.
|
||||
|
||||
**Quick reference:**
|
||||
|
||||
```bash
|
||||
# View roles and capabilities
|
||||
talos iam roles list && talos iam roles describe admin
|
||||
|
||||
# Service registry (Grafana, MinIO, Forgejo, etc.)
|
||||
talos iam services list && talos iam services describe grafana
|
||||
|
||||
# Agents (admin-bot, ci-bot with auto-rotation)
|
||||
talos iam agents list && talos iam agents rotate ci-bot
|
||||
|
||||
# Role bindings (user → role with TTL)
|
||||
talos iam bindings grant [email protected] devops --expires 2026-12-31
|
||||
talos iam bindings list
|
||||
|
||||
# Audit trail (90-day retention, 12 event types)
|
||||
talos iam audit list && talos iam audit export --format json
|
||||
|
||||
# OIDC provider sync with Authentik
|
||||
talos iam providers sync-authentik
|
||||
```
|
||||
|
||||
**See `homelab/CLAUDE.md` § IAM Management for full reference** (roles, services, agents, bindings, audit, providers).
|
||||
|
||||
**Vault paths:** All IAM state stored under `cluster/iam/{federation,roles,services,agents,bindings}`.
|
||||
|
||||
### CI/CD Image Registry Authentication (Forgejo + Runner)
|
||||
|
||||
**After IAM Phase 6 changes:** All image push/pull operations via CI runner require JWT token authentication through Authentik → Vault.
|
||||
|
||||
**Push images to Forgejo registry:**
|
||||
|
||||
```bash
|
||||
# 1. Get ci-bot JWT token (runner has this injected via ServiceAccount)
|
||||
export REGISTRY_TOKEN=$(talos get cluster/iam/agents/ci-bot --key token)
|
||||
|
||||
# 2. Authenticate docker/podman to Forgejo registry
|
||||
docker login forgejo.riotpiao.homelab.com \
|
||||
--username ci-bot \
|
||||
--password "$REGISTRY_TOKEN"
|
||||
|
||||
# 3. Tag and push image
|
||||
docker tag myapp:latest forgejo.riotpiao.homelab.com/rock/myapp:latest
|
||||
docker push forgejo.riotpiao.homelab.com/rock/myapp:latest
|
||||
```
|
||||
|
||||
**Pull images in runner (automatic):**
|
||||
|
||||
```bash
|
||||
# Inside .forgejo/workflows/*.yml, runner pulls via K8s ServiceAccount
|
||||
# No explicit login needed — imagePullSecrets injected by runner pod
|
||||
image: forgejo.riotpiao.homelab.com/rock/myapp:latest
|
||||
```
|
||||
|
||||
**Runner pod setup:**
|
||||
|
||||
```bash
|
||||
# ServiceAccount (in cicd namespace) has Vault JWT auth injected
|
||||
kubectl get serviceaccount -n cicd forgejo-runner
|
||||
kubectl describe sa forgejo-runner -n cicd
|
||||
|
||||
# ImagePullSecret auto-mounted from K8s secret:
|
||||
kubectl get secret -n cicd | grep forgejo-registry
|
||||
```
|
||||
|
||||
**Vault paths for credentials:**
|
||||
|
||||
```
|
||||
cluster/iam/agents/ci-bot # JWT token for push authentication
|
||||
cluster/iam/agents/admin-bot # Alternative admin agent (if needed)
|
||||
```
|
||||
|
||||
---
|
||||
@@ -0,0 +1,324 @@
|
||||
# Flux CD Integration Planning — START HERE
|
||||
|
||||
## What Just Happened?
|
||||
|
||||
Your subagent completed **comprehensive planning documentation** for integrating Flux CD v2 with your homelab's helmfile-based infrastructure.
|
||||
|
||||
**Three complete documents created:**
|
||||
|
||||
1. **FLUX_INTEGRATION_PLAN.md** (1,810 lines)
|
||||
- Full technical specification with code examples
|
||||
- Phase-by-phase implementation roadmap
|
||||
- Conflict resolution & safety procedures
|
||||
- Testing strategy & risk assessment
|
||||
|
||||
2. **FLUX_PLANNING_SUMMARY.md** (351 lines)
|
||||
- Executive overview for stakeholders
|
||||
- Decision matrices & quick reference
|
||||
- Timeline & effort estimates
|
||||
- Success metrics
|
||||
|
||||
3. **FLUX_PLANNING_INDEX.md** (356 lines)
|
||||
- Navigation guide across all documents
|
||||
- Quick start for different audiences
|
||||
- FAQ & next steps
|
||||
|
||||
**Total:** 2,517 lines of planning documentation
|
||||
|
||||
---
|
||||
|
||||
## The Plan in 60 Seconds
|
||||
|
||||
### What Problem Are We Solving?
|
||||
|
||||
Current helmfile workflow:
|
||||
- Manual `helmfile apply` required
|
||||
- No automatic drift detection
|
||||
- No Git audit trail for changes
|
||||
- No approval gates
|
||||
- Hard to scale to multi-cluster
|
||||
|
||||
### What's the Solution?
|
||||
|
||||
Deploy **Flux CD v2** (GitOps) to:
|
||||
- Continuously reconcile cluster state from Git
|
||||
- Auto-detect & correct drift
|
||||
- Maintain full audit trail
|
||||
- Support staged rollouts with approval gates
|
||||
- Keep helmfile.yaml.gotmpl as fallback during transition
|
||||
|
||||
### How Do We Do It?
|
||||
|
||||
**3 phases, 6–8 weeks, ~99 hours:**
|
||||
|
||||
| Phase | Timeline | Work | Goal |
|
||||
|-------|----------|------|------|
|
||||
| **1** | Weeks 1–2 | Bootstrap Flux + helmfile bridge | Zero breaking changes |
|
||||
| **2** | Weeks 3–6 | Migrate 23 releases to HelmRelease CRDs | Parallel migration (4 streams) |
|
||||
| **3** | Weeks 7–8 | Enable auto-sync, metrics, runbooks | Full GitOps readiness |
|
||||
|
||||
**Key:** No downtime. Helmfile stays functional as fallback throughout.
|
||||
|
||||
---
|
||||
|
||||
## Architecture Simplified
|
||||
|
||||
```
|
||||
Git (Forgejo) ← Source of Truth
|
||||
└─→ Flux Reconciliation Loop (every 5 min)
|
||||
└─→ Kubernetes Cluster
|
||||
└─→ 23 Helm Releases (reconciled state)
|
||||
```
|
||||
|
||||
That's it. Flux watches Git. When you push changes, Flux applies them. If someone manually changes the cluster (kubectl), Flux auto-corrects on next reconciliation.
|
||||
|
||||
---
|
||||
|
||||
## Key Decisions (No Surprises)
|
||||
|
||||
| Decision | Choice | Reasoning |
|
||||
|----------|--------|-----------|
|
||||
| **Controller** | Flux v2 | Stable, battle-tested; v3 still beta |
|
||||
| **Helm** | HelmRelease CRDs | Preserves values-based workflow |
|
||||
| **Secrets** | SOPS + age | Git-stored, audited, simple |
|
||||
| **Rollout** | Phased (3×8 weeks) | Lower risk, easier debugging |
|
||||
|
||||
All decisions explained in detail in FLUX_INTEGRATION_PLAN.md §3 (Architecture Decision Matrix).
|
||||
|
||||
---
|
||||
|
||||
## What You Get
|
||||
|
||||
### By End of Phase 1 (Week 2)
|
||||
- ✅ Flux running in cluster
|
||||
- ✅ Git syncing every 60 seconds
|
||||
- ✅ Helmfile still works as fallback
|
||||
- ✅ Zero disruption to running workloads
|
||||
|
||||
### By End of Phase 2 (Week 6)
|
||||
- ✅ All 23 releases migrated to Git-based HelmRelease CRDs
|
||||
- ✅ Helmfile no longer used for deployments
|
||||
- ✅ Every release tested & verified
|
||||
- ✅ Full test suite in place
|
||||
|
||||
### By End of Phase 3 (Week 8)
|
||||
- ✅ Automatic reconciliation enabled
|
||||
- ✅ Drift detection + alerting working
|
||||
- ✅ Metrics flowing to Prometheus
|
||||
- ✅ Team trained on GitOps workflows
|
||||
- ✅ RTO < 2 hours (restore from Git if needed)
|
||||
|
||||
---
|
||||
|
||||
## How to Read the Documentation
|
||||
|
||||
### Quick Overview (10 min)
|
||||
→ **Read:** FLUX_PLANNING_SUMMARY.md
|
||||
|
||||
Start here to understand what we're doing and why. Tables, diagrams, high-level summary. Perfect for stakeholder presentations.
|
||||
|
||||
### Getting Ready to Build (1 hour)
|
||||
→ **Read:** FLUX_PLANNING_INDEX.md + FLUX_INTEGRATION_PLAN.md (Executive Summary)
|
||||
|
||||
Learn the full architecture, decision rationale, and how phases fit together.
|
||||
|
||||
### Phase 1 Implementation (Week 1–2)
|
||||
→ **Reference:** FLUX_INTEGRATION_PLAN.md §5.1 (Phase 1: Flux Bootstrap)
|
||||
|
||||
Detailed tasks:
|
||||
- 1.1: Bootstrap Flux into cluster
|
||||
- 1.2: Create Git repo structure
|
||||
- 1.3: HelmRepository CRDs (13 repos)
|
||||
- 1.4: SOPS + age setup
|
||||
- 1.5: Helmfile-bridge CronJob
|
||||
|
||||
### Phase 2 Migration (Weeks 3–6)
|
||||
→ **Reference:** FLUX_INTEGRATION_PLAN.md §5.2 (Phase 2: HelmRelease Migration)
|
||||
|
||||
Four parallel streams:
|
||||
- Stream A: Low-risk (reloader, prometheus)
|
||||
- Stream B: Medium-risk (cert-manager, ingress)
|
||||
- Stream C: High-risk secrets (authentik, vault)
|
||||
- Stream D: Complex stateful (minio, forgejo)
|
||||
|
||||
Per-release process: generate CRD → validate → deploy → test → commit
|
||||
|
||||
### Phase 3 Production Readiness (Weeks 7–8)
|
||||
→ **Reference:** FLUX_INTEGRATION_PLAN.md §5.3 (Phase 3: Continuous Reconciliation)
|
||||
|
||||
Auto-sync, metrics, runbooks, team training.
|
||||
|
||||
### Troubleshooting & Rollback
|
||||
→ **Reference:** FLUX_INTEGRATION_PLAN.md §7 (Rollback & Safety Guardrails)
|
||||
|
||||
How to recover if something breaks:
|
||||
- Suspend Flux + manual rollback
|
||||
- Git revert + auto-reconciliation
|
||||
- Disaster recovery from Git
|
||||
|
||||
### Testing Strategy
|
||||
→ **Reference:** FLUX_INTEGRATION_PLAN.md §8 (Testing Strategy)
|
||||
|
||||
Unit tests, integration tests, chaos tests, production deployment strategy.
|
||||
|
||||
---
|
||||
|
||||
## Risk Summary
|
||||
|
||||
### Main Risks & How We Handle Them
|
||||
|
||||
| Risk | Mitigation |
|
||||
|------|-----------|
|
||||
| **Flux + helmfile conflict** | Stagger reconciliation (helmfile 30min, Flux 5min) |
|
||||
| **Secret injection breaks** | Three-tier approach (SOPS + ConfigMaps + .env fallback) |
|
||||
| **Secrets leak in Git** | SOPS encryption from start + pre-commit hooks |
|
||||
| **Cluster recovery fails** | Keep helmfile as fallback; test quarterly |
|
||||
|
||||
All risks detailed with specific mitigations in FLUX_INTEGRATION_PLAN.md §9 (Risk Assessment).
|
||||
|
||||
---
|
||||
|
||||
## Timeline Reality Check
|
||||
|
||||
```
|
||||
Week 1–2: Phase 1 bootstrap (20 hrs)
|
||||
├─ 1 DevOps engineer + 1 Security engineer
|
||||
└─ 0 downtime to running workloads
|
||||
|
||||
Week 3–6: Phase 2 migration (40 hrs)
|
||||
├─ 4 parallel streams (DevOps + Ops + Security)
|
||||
└─ Release-by-release (low risk)
|
||||
|
||||
Week 7–8: Phase 3 hardening (16 hrs)
|
||||
├─ DevOps + QA
|
||||
└─ Runbooks + training
|
||||
|
||||
Total: ~99 hours (~2.5 FTE-weeks)
|
||||
6–8 calendar weeks (with parallelization)
|
||||
```
|
||||
|
||||
Actual timeline depends on:
|
||||
- Team size (4 engineers = 8 weeks; 2 engineers = 12 weeks)
|
||||
- Experience with Flux (learning curve ~40 hours)
|
||||
- Testing rigor (each phase adds 1–2 weeks)
|
||||
|
||||
---
|
||||
|
||||
## Next Actions
|
||||
|
||||
### Immediately (Today)
|
||||
|
||||
1. **Review FLUX_PLANNING_SUMMARY.md** (15 min)
|
||||
- Understand the approach
|
||||
- Check decision matrix
|
||||
- Confirm timeline is acceptable
|
||||
|
||||
2. **Share with stakeholders**
|
||||
- Security team: review SOPS approach
|
||||
- Ops team: review rollback procedures
|
||||
- Management: confirm timeline & resources
|
||||
|
||||
3. **Get approval** for:
|
||||
- Phased approach (6–8 weeks)
|
||||
- Flux v2 + HelmRelease CRDs
|
||||
- SOPS encryption for secrets
|
||||
- ~99 hours effort
|
||||
|
||||
### Week 1 (Phase 1 Kickoff)
|
||||
|
||||
1. **Assign team members**
|
||||
- DevOps lead
|
||||
- Security engineer (SOPS)
|
||||
- Ops engineer (testing)
|
||||
|
||||
2. **Bootstrap Flux**
|
||||
- `flux bootstrap git` command
|
||||
- Set up Git repo structure
|
||||
- Deploy HelmRepository CRDs
|
||||
|
||||
3. **Start helmfile-bridge development**
|
||||
- CronJob to run `helmfile apply` every 30 min
|
||||
- Test alongside Flux (staggered intervals)
|
||||
|
||||
### Weeks 3–8 (Phases 2 & 3)
|
||||
|
||||
Follow the phase roadmap in FLUX_INTEGRATION_PLAN.md with weekly syncs.
|
||||
|
||||
---
|
||||
|
||||
## Files Created
|
||||
|
||||
All in `/Users/rockliang/workplace/homelab/`:
|
||||
|
||||
1. **FLUX_INTEGRATION_PLAN.md** (55 KB)
|
||||
- Complete technical specification
|
||||
- Phase-by-phase breakdown
|
||||
- Code examples & detailed procedures
|
||||
|
||||
2. **FLUX_PLANNING_SUMMARY.md** (13 KB)
|
||||
- Executive overview
|
||||
- Decision matrices
|
||||
- Quick reference tables
|
||||
|
||||
3. **FLUX_PLANNING_INDEX.md** (13 KB)
|
||||
- Navigation guide
|
||||
- Quick start by audience
|
||||
- FAQ & related docs
|
||||
|
||||
4. **_FLUX_START_HERE.md** (this file)
|
||||
- Quick orientation
|
||||
- Next actions
|
||||
|
||||
---
|
||||
|
||||
## Questions to Ask
|
||||
|
||||
Before Phase 1 starts, clarify:
|
||||
|
||||
1. **Team capacity?** How many FTE can we dedicate?
|
||||
- 4 FTE → 8 weeks
|
||||
- 2 FTE → 12 weeks
|
||||
|
||||
2. **Timeline flexibility?** Hard deadline or can we adjust?
|
||||
- If hard: compress with more parallel streams
|
||||
- If flexible: add more testing/validation
|
||||
|
||||
3. **Flux experience on team?** Anyone used Flux before?
|
||||
- If no: add 1–2 weeks for learning curve
|
||||
- If yes: can reduce onboarding time
|
||||
|
||||
4. **Multi-cluster plans?** Will you add more clusters after homelab?
|
||||
- If yes: design for portability from start
|
||||
- If no: homelab-specific is fine
|
||||
|
||||
5. **SOPS comfort?** Any concerns about secret encryption in Git?
|
||||
- If yes: alternative is store in Vault (referenced from HelmRelease)
|
||||
- If no: SOPS is recommended
|
||||
|
||||
---
|
||||
|
||||
## Document Quality Checklist
|
||||
|
||||
The planning documentation includes:
|
||||
|
||||
- ✅ **Executive summary** — problem & solution in 1 page
|
||||
- ✅ **Current state analysis** — what we're migrating from
|
||||
- ✅ **Architecture decisions** — Flux v2, HelmRelease, SOPS (with reasoning)
|
||||
- ✅ **Detailed design** — GitRepository, Kustomization, HelmRelease CRDs
|
||||
- ✅ **3-phase roadmap** — specific tasks, timelines, deliverables, success criteria
|
||||
- ✅ **Conflict resolution** — helmfile + Flux, .env → SOPS, kubectl drift
|
||||
- ✅ **Rollback procedures** — what to do if something breaks
|
||||
- ✅ **Safety guardrails** — RBAC, audit logging, validation webhooks, approval gates
|
||||
- ✅ **Testing strategy** — unit, integration, chaos, production deployment
|
||||
- ✅ **Risk assessment** — probability, impact, mitigation for each risk
|
||||
- ✅ **Timeline & effort** — 99 hours, 6-8 weeks, team composition
|
||||
- ✅ **Useful commands** — Flux CLI cheatsheet
|
||||
- ✅ **FAQ** — downtime, rollback, recovery, cost
|
||||
|
||||
Ready for review and implementation kickoff.
|
||||
|
||||
---
|
||||
|
||||
**Status:** Planning phase complete. Ready for team discussion & approval.
|
||||
|
||||
**Next:** Review FLUX_PLANNING_SUMMARY.md, approve approach, assign Phase 1 team.
|
||||
Reference in New Issue
Block a user