# Homelab Kubernetes Cluster A bare-metal three-node Kubernetes cluster running Talos Linux with **18 Helm releases** across 22 namespaces. Includes distributed storage (MinIO + Longhorn), full observability (Prometheus + Grafana + Loki), federated SSO (Authentik OIDC), secrets management (Vault), CI/CD (Forgejo + Argo CD), messaging (Kafka + kmsvc), and workflow orchestration (Temporal). ## Use Cases & Architecture **Why this stack?** This homelab replicates **production-grade cloud-native infrastructure** on bare metal, enabling: 1. **Learning & Prototyping** — Test distributed systems patterns (HA databases, event-driven messaging, GitOps workflows) before deploying to cloud 2. **Self-Hosted Services** — Run applications (Story Crater, etc.) with zero cloud lock-in; full control over data, compliance, and networking 3. **Infrastructure as Code** — Git-driven cluster state via Helmfile + Forgejo Actions + Argo CD; every change is auditable and reproducible 4. **Observability Sandbox** — Experiment with Prometheus metrics, Loki log aggregation, and custom Grafana dashboards at scale **Typical workflow:** ``` Developer pushes to Forgejo (git forge) ↓ Forgejo Actions CI runs tests + builds OCI image ↓ Image pushed to Forgejo registry (private, on-cluster) ↓ Argo CD detects deployment repo change (pull-based GitOps) ↓ New pods roll out; Grafana alerts on errors/latency ↓ Temporal workflows coordinate long-running operations (e.g., async jobs) ↓ Kafka queues decouple services (fire-and-forget messaging) ↓ All logs + metrics centralized in Grafana for debugging ``` **Architecture principles:** - **Immutable OS** — Talos Linux (no SSH, declarative machine configs) - **No external dependencies** — All data stored locally (MinIO, CloudNativePG, Longhorn) - **High availability** — 3-replica databases, multi-node storage, cross-AZ readiness (on bare metal: cross-rack affinity) - **Federated identity** — Single Authentik OIDC provider for all services (Grafana, MinIO, Forgejo, Argo CD) - **Secrets at rest** — Vault + encrypted etcd; credentials never in logs or ConfigMaps - **Infrastructure-as-code** — Every service deployed via Helmfile; one `helmfile apply` recovers from total failure ## Quick Start — Deploying the Cluster ### 1. Bootstrap Talos Nodes Bootstrap each Talos node with your cluster schematic (see `CLAUDE.md` or `README` step 1–8). ### 2. Set Up Secrets All secrets are managed via environment variables sourced from `.env` (gitignored). The helmfile template expands them at deploy time. **Step 1: Copy the template** ```bash cp .env.example .env ``` **Step 2: Populate required secrets** Edit `.env` and fill in cluster configuration. See `.env.example` for all options: ```bash # Cluster configuration CLUSTER_DOMAIN=riotpiao.com # Your cluster domain POSTGRES_HOST=ddb-cluster-rw.ddb.svc.cluster.local MINIO_ENDPOINT=minio.storage.svc.cluster.local:9000 KAFKA_BOOTSTRAP=kmsvc-kafka-bootstrap.sqs.svc.cluster.local:9092 REDIS_ADDR=kmsvc-redis-master.sqs.svc.cluster.local:6379 # Service credentials (generate with: openssl rand -hex 32) MINIO_ROOT_PASSWORD= GRAFANA_ADMIN_PASSWORD= AUTHENTIK_SECRET_KEY= AUTHENTIK_BOOTSTRAP_PASSWORD= AUTHENTIK_PG_PASSWORD= ``` **Step 3: Load and deploy** ```bash # Load .env into current shell vsource .env # Preview all changes before deployment helmfile diff # Deploy the entire stack helmfile apply ``` ### 3. Verify Deployment ```bash # Check all pods are running kubectl get pods -A # Confirm key services are ready kubectl wait deploy/authentik-server -n iam --for=condition=Available --timeout=300s kubectl wait deploy/grafana -n logging --for=condition=Available --timeout=300s # Access Grafana make pf-grafana # localhost:3000 (login: admin / GRAFANA_ADMIN_PASSWORD) ``` ### 4. First-Time Access **Default Credentials:** - **Authentik:** https://authentik.$(CLUSTER_DOMAIN) → login: `akadmin` / `AUTHENTIK_BOOTSTRAP_PASSWORD` (change immediately) - **Grafana:** https://grafana.$(CLUSTER_DOMAIN) → login: `admin` / `GRAFANA_ADMIN_PASSWORD` - **MinIO:** https://minio.$(CLUSTER_DOMAIN) → login: `MINIO_ROOT_USER` / `MINIO_ROOT_PASSWORD` - **Argo CD:** https://argocd.$(CLUSTER_DOMAIN) → login via Authentik OIDC - **Forgejo:** https://forgejo.$(CLUSTER_DOMAIN) → login via Authentik OIDC **Next Steps:** 1. Change default passwords in each service 2. Configure OIDC redirects (see `k8s/talos-iam/` for details) 3. Set up GitOps: push infrastructure to Forgejo, configure Argo CD 4. Review dashboards in Grafana (Prometheus + Loki) --- ## Architecture ``` 192.168.1.0/24 (LAN) │ ┌──────────────────────┼──────────────────────┐ │ │ │ 192.168.1.* 192.168.1.* 192.168.1.* ┌────────────────┐ ┌────────────────┐ ┌────────────────┐ │ talos-cp-1 │ │ talos-worker-1 │ │ talos-worker-2 │ │ Control-Plane │ │ Worker │ │ Worker │ │ + Workloads │ │ (Storage) │ │ (Storage) │ │ (az-a) │ │ (az-b) │ │ (az-c) │ ├────────────────┤ ├────────────────┤ ├────────────────┤ │ Pods: │ │ Pods: │ │ Pods: │ │ • ingress-nginx│ │ • kube-system │ │ • kube-system │ │ • authentik │ │ • storage │ │ • storage │ │ • vault │ │ └─ minio-2 │ │ └─ minio-3 │ │ • logging │ │ │ │ │ │ ├─ loki │ │ │ │ │ │ ├─ promtail │ │ │ │ │ │ └─ grafana │ │ │ │ │ │ • monitoring │ │ │ │ │ │ ├─ prom │ │ │ │ │ │ └─ blackbox │ │ │ │ │ │ • storage │ │ │ │ │ │ └─ minio-1 │ │ │ │ │ │ • cicd │ │ │ │ │ │ ├─ forgejo │ │ │ │ │ │ ├─ argocd │ │ │ │ │ │ └─ runner │ │ │ │ │ │ • sqs │ │ • sqs │ │ • sqs │ │ ├─ kafka-0 │ │ ├─ kafka-1 │ │ ├─ kafka-2 │ │ └─ kmsvc │ │ └─ redis │ │ │ │ • ddb │ │ • ddb │ │ • ddb │ │ └─ postgres-0│ │ └─ postgres-1│ │ └─ postgres-2│ │ • temporal │ │ • temporal │ │ • temporal │ │ ├─ server │ │ ├─ cassandra │ │ ├─ cassandra │ │ └─ cassandra │ │ │ -1 │ │ │ -2 │ │ -0 │ │ └─ (replica) │ │ └─ (replica) │ │ • llm │ │ │ │ │ │ └─ ollama │ │ │ │ │ └────────────────┘ └────────────────┘ └────────────────┘ Replication & Fault Tolerance: Data Layer: • MinIO: minio-1 ↔ minio-2 ↔ minio-3 (3-way active-active S3) • PostgreSQL: postgres-0 ↔ postgres-1 ↔ postgres-2 (primary + 2 standbys, HA streaming replication) • Kafka: kafka-0 ↔ kafka-1 ↔ kafka-2 (3 brokers, RF=3, min-ISR=2, cross-AZ) • Temporal: cassandra-0 ↔ cassandra-1 ↔ cassandra-2 (3-node distributed) • Loki: loki → MinIO (chunks stored in s3://loki-chunks, 10-day retention) Single-Replica Services (Protected by PodDisruptionBudget minAvailable=1): • Observability: Grafana, Prometheus, Loki (Recreate strategy for RWO PVCs) • IAM: Authentik, Vault (Recreate strategy for RWO PVCs) • CI/CD: Argo CD server/repo-server, Forgejo (Recreate strategy for RWO PVCs) • LLM: Ollama Networking: Remote access: UR_OWN.duckdns.org → home IP → cp-1 ``` ## Stack | Layer | Technology | Namespace | Purpose | |-------|-----------|-----------|---------| | OS | Talos Linux v1.13.3 | — | Immutable, Kubernetes-native OS | | Kubernetes | v1.36.1 | — | Container orchestration | | CNI | Cilium (eBPF) | kube-system | Networking, replaces kube-proxy | | Ingress | Nginx Ingress Controller | ingress-nginx | Reverse proxy, hostname-based routing | | Block Storage | Longhorn v1.7.0 | longhorn-system | Default StorageClass | | Object Store | MinIO (multi-AZ) | storage | S3-compatible, site-replicated across az-a/az-b | | IAM / SSO | Authentik | iam | OIDC provider for Grafana, MinIO, Forgejo, Argo CD | | Secret Store | HashiCorp Vault | iam | KV secrets backend, JWT auth via Authentik | | Git Forge *(planned)* | Forgejo | forge | Git server, built-in OCI registry, Actions CI | | CI Runner *(planned)* | Forgejo Actions + DinD | cicd | Privileged build pod; images pushed to Forgejo OCI | | CD *(planned)* | Argo CD | argocd | Pull-based GitOps; never holds kubeconfig in CI | | Log Backend | Loki (SingleBinary) | logging | 10-day retention, backed by MinIO | | Log Collector | Promtail (DaemonSet) | logging | Scrapes pod logs + Talos journal | | Metrics | kube-prometheus-stack | monitoring | Prometheus + node-exporter + kube-state-metrics | | Log/Metrics UI | Grafana | logging | Dashboards for Loki + Prometheus | | Cluster UI | Portainer CE | dashboard | Container/workload management UI | | LLM Inference | Ollama | llm | Local LLM model serving (open-source models) | ## Repository Structure ``` homelab/ ├── helmfile.yaml # Single source of truth — deploys everything ├── .env.example # Required env vars template (copy to .env, gitignored) │ ├── cluster-config/ # Talos + Kubernetes bootstrap │ ├── cilium-values.yaml │ ├── longhorn_bootstrap.sh │ ├── controlplane.yaml # gitignored — contains secrets │ ├── worker-1.yaml # gitignored │ ├── secrets.yaml # gitignored │ ├── kubeconfig # gitignored │ └── talosconfig # gitignored │ ├── k8s/ │ ├── ingress/ # Nginx Ingress Controller + all Ingress rules │ │ ├── nginx-values.yaml │ │ └── ingress.yaml │ ├── storage/ # MinIO multi-AZ object store │ │ ├── minio-az-a-values.yaml │ │ ├── minio-az-b-values.yaml │ │ ├── minio-az-a-pvc.yaml │ │ ├── minio-service.yaml │ │ ├── minio-legacy-alias.yaml │ │ └── minio-replication-job.yaml │ ├── logging/ # Observability stack │ │ ├── loki-values.yaml │ │ ├── promtail-values.yaml │ │ └── grafana-values.yaml │ ├── monitoring/ # Prometheus stack + alerting + Grafana dashboards-as-code │ │ ├── prometheus-values.yaml │ │ ├── blackbox-exporter-values.yaml # Active uptime probes (feeds Service Availability dashboard) │ │ ├── ingress-alerts.yaml # PrometheusRule: ingress 5xx rate, p95 latency │ │ └── dashboards/ # ConfigMaps picked up live by Grafana's sidecar │ │ ├── service-availability.yaml # Uptime probes + cert expiry (operator glance) │ │ ├── service-golden-signals.yaml # Latency & Golden Signals (ingress RED) │ │ ├── service-internals.yaml # Per-service deep-dive (MinIO/Forgejo/Argo CD/Vault/Longhorn/certs) │ │ ├── kube-controller-health.yaml # API server RED + kube-state-metrics controller-health proxy │ │ ├── hardware-overview.yaml # Per-node CPU/mem/disk/network/load summary │ │ └── control-plane-logs.yaml # kube-system + add-on logs (Loki) │ ├── portainer/ # Portainer CE │ │ └── portainer-values.yaml │ ├── talos-iam/ # Authentik + Vault IAM │ │ ├── authentik-values.yaml │ │ ├── vault-values.yaml │ │ ├── setup_vault.sh # One-time Vault init (not replaced by Helmfile) │ │ └── provision_oidc.py # Authentik OIDC provisioning │ ├── coredns/ # CoreDNS hostname rewrites (in-cluster DNS) │ │ └── coredns-configmap.yaml │ ├── talos-ci-cd/ # CI/CD stack (planned — not yet applied) │ │ ├── talos_version_control.html # Implementation plan + build runbook │ │ ├── forgejo-values.yaml # Forgejo Helm values (gitea-charts/gitea) │ │ ├── argocd-values.yaml # Argo CD Helm values │ │ └── charts/forgejo-runner/ # Local Helm chart for the Actions runner │ │ ├── Chart.yaml │ │ ├── values.yaml │ │ └── templates/ │ │ ├── deployment.yaml # Runner + DinD sidecar, Recreate strategy │ │ ├── pvc.yaml # runner-reg (1 Gi) + runner-dind (30 Gi) │ │ └── networkpolicy.yaml # Egress: forge ns + DNS + internet only │ └── duckdns/ # DuckDNS DDNS updater CronJob │ ├── talos-cli/ # Rust CLI for Vault secret access ├── project_context.md # Authoritative live-state reference ├── refine_cluster.md # Known-issues runbook └── LOG.md # Append-only change journal ``` ## Access ### Discover Ingress LoadBalancer IP ```bash # Find the external IP assigned by Cilium LB-IPAM kubectl get svc -n ingress-nginx ingress-nginx # Example output: # LoadBalancer IP: 192.168.1.160 (Cilium LB-IPAM assignment) ``` Add to `/etc/hosts` on every client machine (Mac/Linux): ``` # WireGuard access (remote — via talos-cp-1) 10.6.0.1 grafana.riotpiao.com authentik.riotpiao.com vault.riotpiao.com minio.riotpiao.com prometheus.riotpiao.com portainer.riotpiao.com longhorn.riotpiao.com loki.riotpiao.com forgejo.riotpiao.com temporal.riotpiao.com temporal-grpc.riotpiao.com kmsvc.riotpiao.com # LAN access (on the home network — use actual LoadBalancer IP from above) 192.168.1.160 grafana.riotpiao.com authentik.riotpiao.com vault.riotpiao.com minio.riotpiao.com prometheus.riotpiao.com portainer.riotpiao.com longhorn.riotpiao.com loki.riotpiao.com forgejo.riotpiao.com temporal.riotpiao.com temporal-grpc.riotpiao.com kmsvc.riotpiao.com ``` **Note:** `192.168.1.160` is an example Cilium LB-IPAM assignment. Verify with `kubectl get svc -n ingress-nginx ingress-nginx`. **There is no real DNS wildcard for `*.riotpiao.com`** — every hostname must be added to `/etc/hosts` explicitly (as above) before it resolves. Adding a new Ingress host doesn't make it reachable by itself; add the line too. ### kubectl Context Two contexts exist in `cluster-config/kubeconfig`, pointed at the same cluster over different paths: | Context | Server | Use when | |---|---|---| | `admin@homelab-cluster` | `192.168.1.213:6443` (LAN) | On the home network | | `admin@homelab-cluster-1` | `10.6.0.1:6443` (WireGuard) | Remote / off-LAN | If `kubectl` commands hang or refuse the connection, switch: `kubectl config use-context admin@homelab-cluster-1`. Then access services at: | Service | URL | Credentials | |---------|-----|-------------| | Grafana | http://grafana.riotpiao.com | admin / `GRAFANA_ADMIN_PASSWORD` or Authentik SSO | | Authentik | http://authentik.riotpiao.com | akadmin / see `.env` | | Vault | http://vault.riotpiao.com | root token / see `setup_vault.sh` output | | MinIO console | http://minio.riotpiao.com | `MINIO_ROOT_USER` / `MINIO_ROOT_PASSWORD` | | Prometheus | http://prometheus.riotpiao.com | no auth | | Portainer | http://portainer.riotpiao.com | set on first visit | | Longhorn | http://longhorn.riotpiao.com | no auth | | Forgejo *(planned)* | https://forgejo.forge.riotpiao.com | `rock` / `FORGEJO_ADMIN_PASSWORD`, or Authentik SSO | | Argo CD *(planned)* | `kubectl port-forward -n argocd svc/argocd-server 8080:443` | Authentik SSO (admins only) | Grafana → "Homelab" folder has the operator dashboards (sidecar-loaded from `k8s/monitoring/dashboards/`, no restart needed on change): - **Service Availability & Certificate Expiration** — uptime probes + cert-manager expiry - **Latency & Golden Signals** — ingress request rate/error %/p50-p99 latency - **Kube-Controller Health** — API server RED metrics + kube-state-metrics controller-health signals - **Hardware Statistics** — per-node CPU/mem/disk/network/load - **Service Internals** — per-service deep-dive (MinIO/Forgejo/Argo CD/Vault/Longhorn) ## Deploy ```bash # 1. Install helmfile (once) brew install helmfile # 2. Set credentials cp .env.example .env # edit .env with your passwords # 3. Deploy everything helmfile apply # Deploy a single stack helmfile apply -l namespace=logging helmfile apply -l name=grafana helmfile apply -l namespace=ingress-nginx # Preview changes before applying helmfile diff ``` ## Bootstrap Order (fresh cluster) ``` 1. Provision nodes: talosctl apply-config (make apply-cp / apply-worker-new) 2. Bootstrap Kubernetes: talosctl bootstrap 3. Install Cilium: helm install cilium -f cluster-config/cilium-values.yaml 4. Install Longhorn: bash cluster-config/longhorn_bootstrap.sh 5. Deploy everything else: helmfile apply 6. Vault init (one-time): bash k8s/talos-iam/setup_vault.sh 7. Authentik OIDC: python3 k8s/talos-iam/provision_oidc.py 8. Label worker: kubectl label node talos-worker-1 node-role.kubernetes.io/worker= # ── CI/CD (planned — run after step 8) ───────────────────────────────────── 9. Private CA + TLS: see k8s/talos-ci-cd/talos_version_control.html §12 Block 0 10. Deploy Forgejo + runner: helmfile apply -l name=forgejo && helmfile apply -l name=forgejo-runner 11. Authentik SSO for CI: Forgejo + Argo CD OIDC (§12 Block 1.5) 12. Talos node CA trust: talosctl patch machineconfig (§12 Block 2) 13. Deploy Argo CD: helmfile apply -l name=argocd (§12 Block 4) 14. Wire deploy repo: argocd app create + push first manifests (§12 Block 4) ``` ## IAM & Auth Flow Authentik is the central OIDC identity provider. Vault stores secrets and delegates authentication back to Authentik. ``` User / core-cli │ │ OAuth2 / OIDC ▼ Authentik (authentik.riotpiao.com) ├── grafana app → Grafana OIDC login (group → Admin/Viewer role) ├── minio app → MinIO OIDC login (group → readwrite/readonly policy) ├── vault-browser → Vault UI OIDC login / `vault login -method=oidc` └── core-cli-shell → CLI device code flow (public client, no secret) │ │ JWKS endpoint for JWT validation ▼ HashiCorp Vault (vault.riotpiao.com) ├── auth/jwt — core-cli authenticates with device code JWT ├── auth/oidc — browser/UI login via Authentik └── secret/ — KV v2: mcp/*, cluster/*, cloud/* ``` **core CLI device code login:** ```bash core secrets login # prints URL + code → approve in browser → Vault token cached core put cluster/DUCKDNS_TOKEN DUCKDNS_TOKEN="abc" # field name = variable name, never `value` ``` **One-time IAM setup (after `helmfile apply`):** ```bash # 1. Provision OIDC apps and groups in Authentik GRAFANA_URL=http://grafana.riotpiao.com \ MINIO_URL=http://minio.riotpiao.com \ python3 k8s/talos-iam/provision_oidc.py # 2. Init Vault, wire JWT + OIDC auth, seed secrets bash k8s/talos-iam/setup_vault.sh ``` **CoreDNS hostname rewrites** (`k8s/coredns/coredns-configmap.yaml`) ensure in-cluster pods (Grafana, Vault, Forgejo runner, Argo CD) resolve internal hostnames to cluster services, avoiding hairpin NAT through LB IPs. ## CI/CD Pipeline *(planned)* > **Implementation plan & exact build commands:** [`k8s/talos-ci-cd/talos_version_control.html`](k8s/talos-ci-cd/talos_version_control.html) ### Components | Component | Helm chart | Namespace | Notes | |-----------|-----------|-----------|-------| | Forgejo | `gitea-charts/gitea` (Forgejo image override) | `forge` | Git + OCI registry + Actions engine; SQLite on Longhorn PVC; `strategy: Recreate` | | Forgejo runner | local chart `charts/forgejo-runner` | `cicd` | DinD sidecar; PodSecurity privileged; NetworkPolicy fenced | | Argo CD | `argo/argo-cd` | `argocd` | Pull-based CD; single replica; no external ingress (port-forward only) | ### Pipeline flow ``` Developer │ │ git push ▼ Forgejo (forge ns) ─── webhook ───► Runner (cicd ns) │ │ │ SSO login (Authentik OIDC) │ ACTIONS_RUNTIME_TOKEN → git checkout ▼ │ ci-registry-token (ci-bot) → docker push → Forgejo OCI Forgejo UI / Argo CD UI │ ci-deploy-token (ci-bot) → git commit → rock/deploy │ Forgejo (rock/deploy repo) ◄─────────────┘ │ │ argocd-bot token (repo:read, poll every 3 min) ▼ Argo CD (argocd ns) │ │ kubectl apply (cluster-admin ServiceAccount — never in CI) ▼ K8s workloads (images from Forgejo OCI) ``` ### Key security decisions - **No kubeconfig in CI.** The runner can only push commits + OCI images. Argo CD bridges the gap autonomously. - **Scoped machine credentials.** `ci-bot` tokens are narrowly scoped: `package:write` for OCI, `repo:write` on `rock/deploy` only. A compromised runner cannot read other repos or call the K8s API. - **Private CA TLS.** Forgejo self-terminates HTTPS with a homelab CA (EC P-256, 10-year). The CA cert is distributed to Talos nodes via `machineconfig` patch and to the runner via K8s Secret. `ca.key` never enters the cluster. - **Authentik SSO for humans.** All interactive logins (Forgejo UI, Argo CD UI) route through Authentik. `homelab-admins` group → Forgejo admin + Argo CD `role:admin`; `homelab-devs` group → Forgejo user + no Argo CD access. - **Forgejo LB IP pinned.** Cilium LB-IPAM annotation `io.cilium/lb-ipam-ips: ` fixes the Forgejo LoadBalancer IP so the TLS SAN and DNS entries never need updating. Configure in `k8s/talos-ci-cd/forgejo-values.yaml`. ### Workflow file location Forgejo Actions uses GitHub Actions syntax. Workflow files live in `.forgejo/workflows/` in each source repo: ``` rock/source/ └── .forgejo/ └── workflows/ ├── ci.yml # build + test + push OCI image └── cd.yml # on: push to main → bump image tag in rock/deploy ``` ## Log Data Flow ``` Pods / Talos journal (both nodes) │ Promtail (DaemonSet, all nodes) reads /var/log/pods + /var/log/journal │ ▼ Loki (logging ns) indexes + compacts, 10-day retention │ stores chunks via S3 ▼ MinIO frontend service minio.storage.svc.cluster.local:9000 (active-active, round-robin) │ minio-az-a (cp-1) ↔ minio-az-b (worker-1) ↔ minio-az-c (worker-2) 3-way site replication (bidirectional, automatic) │ Grafana (logging ns) queries Loki + Prometheus via dashboards │ Nginx Ingress → grafana.riotpiao.com browser access ``` ## Example Applications & Workloads This cluster runs production-like applications and infrastructure services: ### Story Crater Backend **Type:** Distributed message-driven application **Namespace:** `story-crater-backend` **Architecture:** - gRPC server + REST gateway (Envoy) - PostgreSQL database (story_crater in CloudNativePG cluster) - Kafka topic consumers (via kmsvc message queue) - Temporal workflow integration for long-running operations - Prometheus metrics (processed messages, latency, errors) - Grafana dashboard (message throughput, queue depth, LLM token usage) **Typical flow:** ``` POST /api/messages → gRPC handler ↓ Publish to Kafka topic (Story Crater queue) ↓ Message consumer processes async (may trigger LLM inference) ↓ Results stored in PostgreSQL + emitted as event ↓ Prometheus increments counters (story_crater_messages_handled_total) ↓ Grafana renders message throughput + duration histograms ``` ### Infrastructure Services (Essential) | Service | Purpose | Namespace | Example Use | |---------|---------|-----------|-------------| | **Authentik** | OIDC identity provider | iam | User login, group management, SSO for Grafana/MinIO/Forgejo | | **Vault** | Secrets backend | iam | Database passwords, API keys, JWT token validation | | **CloudNativePG** | PostgreSQL 3-replica cluster | ddb | Auth database (Authentik), app database (Story Crater) | | **Loki** | Log aggregation | logging | Centralize pod logs, Talos kernel logs (10-day retention) | | **Prometheus** | Metrics collection | monitoring | Scrape kube-state-metrics, kubelet, ServiceMonitors (every 30s) | | **Grafana** | Observability dashboards | logging | Query Prometheus + Loki, alert on latency/error spikes | | **Kafka + kmsvc** | Message queue | sqs | Decouple services, async job processing, at-least-once delivery | | **MinIO** | S3-compatible object store | storage | Loki log chunks backend, Vault unseal keys, config backups | | **Longhorn** | Persistent block storage | longhorn-system | All PVCs (Postgres replicas, Kafka broker disks, MinIO) | ### Development Services (Optional) | Service | Purpose | Namespace | Example Use | |---------|---------|-----------|-------------| | **Forgejo** | Self-hosted git + OCI registry | cicd | Version control, CI Actions runner, private Docker images | | **Argo CD** | Pull-based GitOps | cicd | Continuous deployment (deployment repo → Kubernetes) | | **Temporal** | Workflow orchestration | temporal | Schedule long-running jobs, retry logic, state machines | | **Portainer** | Container management UI | dashboard | Pod inspection, image management, quick debugging | ### Monitoring Example: Dashboard Walk-Through Open **Grafana** → **Homelab** folder → **"Story Crater Backend — Service Overview"**: **Row A — Availability & Golden Signals** - Requests/sec (blue = success, red = errors) - Error rate % (goal: < 0.1%) - p50/p95/p99 latency (goal: p99 < 500ms) - Alert: If error rate > 1% for 5 min, page on-call **Row B — Resource Usage** - CPU (request/limit) - Memory (request/limit) - Restart count (goal: 0; alerts if > 2) **Row C — Domain-Specific Metrics** (Story Crater only) - Messages processed/sec (bucketed by status: success, dlq, retry) - Queue depth (Kafka partitions lag) - Dedup window retention (FIFO redelivery tracking) - LLM inference tokens used/sec - External API call latency (e.g., OpenAI) **Row D — Logs** - Live Loki panel: filter by pod + search for errors - Example: `{namespace="story-crater-backend"} | json | level="error"` **Row E — Alert Status** (if SLO defined) - Burn rate (if consuming SLO budget) - Example: "30-day availability SLO = 99.5%; current burn rate = 0.2x" **Row F — Related Dashboards** - Link to Kafka dashboard (queue depth) - Link to PostgreSQL dashboard (story_crater DB) - Link to Temporal dashboard (workflow execution times) ## Adding Hardware to the Cluster ### Step 1 — Get the Talos image The image must include the same extensions as the existing nodes (`iscsi-tools` + `util-linux-tools`). Download from the Image Factory using the cluster's schematic ID: | Format | Use case | URL | |--------|----------|-----| | ISO | USB boot (recommended for bare metal) | `https://factory.talos.dev/image/613e1592.../v1.13.3/metal-amd64.iso` | | RAW disk image | Write directly to drive via another machine | `https://factory.talos.dev/image/613e1592.../v1.13.3/metal-amd64.raw.xz` | | PXE / iPXE | Network boot — no USB needed | `https://factory.talos.dev/image/613e1592.../v1.13.3/kernel-amd64` | > Full schematic ID: `613e1592b2da41ae5e265e8789429f22e121aab91cb4deb6bc3c0b6262961245` --- ### Option A — USB Bootable (recommended) ```bash curl -Lo talos-worker.iso \ "https://factory.talos.dev/image/613e1592b2da41ae5e265e8789429f22e121aab91cb4deb6bc3c0b6262961245/v1.13.3/metal-amd64.iso" sudo dd if=talos-worker.iso of=/dev/sdX bs=4M status=progress && sync ``` --- ### Option B — In-Memory (diskless / RAM boot via PXE) Talos runs entirely from RAM. Useful for temporary nodes or hardware where you don't want to touch the existing OS. ```bash # Boot via PXE pointing to: # Kernel: https://factory.talos.dev/image/613e1592.../v1.13.3/kernel-amd64 # Initrd: https://factory.talos.dev/image/613e1592.../v1.13.3/initramfs-amd64.xz # Cmdline: talos.platform=metal ``` > Note: in-memory nodes lose state on reboot. Not suitable for Longhorn storage nodes. --- ### Option C — Direct Disk Image (headless / remote) ```bash xz -d talos-worker.raw.xz sudo dd if=talos-worker.raw of=/dev/sda bs=4M status=progress && sync ``` --- ### Step 2 — Discover hardware in maintenance mode ```bash # Scan your LAN for the new node in maintenance mode nmap -sn 192.168.1.0/24 # Note: Replace with your actual subnet (e.g., 10.0.1.0/24) # Discover available disks on the node talosctl --nodes --talosconfig cluster-config/talosconfig disks --insecure ``` --- ### Step 3 — Prepare the worker config ```bash cp cluster-config/worker-1.yaml cluster-config/worker-N.yaml ``` Edit exactly these four fields: | Field | Value | |-------|-------| | `machine.network.hostname` | `talos-worker-N` | | `machine.network.interfaces[0].addresses` | `192.168.1.16N/24` | | `machine.install.disk` | disk path from Step 2 | | `machine.nodeLabels.topology.kubernetes.io/zone` | `az-N` | --- ### Step 4 — Apply config ```bash make apply-worker-new N= WN_IP= ``` --- ### Step 5 — Persist the worker IP ```bash # Substitute with the actual IP discovered in Step 2 echo 'export W_IP=192.168.1.' >> ~/.zshrc && source ~/.zshrc ``` --- ### Step 6 — Set the node-role label ```bash kubectl label node talos-worker-N node-role.kubernetes.io/worker= ``` --- ### Step 7 — Verify ```bash kubectl get nodes -w kubectl describe node talos-worker-N | grep -A10 Labels ``` --- ### Step 8 — What automatically extends to the new node | Service | Behaviour | |---------|-----------| | Cilium | New pod scheduled automatically | | Promtail | DaemonSet — starts immediately | | Nginx Ingress | DaemonSet — starts immediately, port 80/443 available on new node | | Longhorn | Detects new node, available for replica scheduling | | MinIO / Loki / Grafana | Stay on existing node (single-replica Deployments) |