Files
homelab/README.md
T

692 lines
31 KiB
Markdown
Raw Normal View History

# Homelab Kubernetes Cluster
A bare-metal three-node Kubernetes cluster running Talos Linux with **18 Helm releases** across 22 namespaces. Includes distributed storage (MinIO + Longhorn), full observability (Prometheus + Grafana + Loki), federated SSO (Authentik OIDC), secrets management (Vault), CI/CD (Forgejo + Argo CD), messaging (Kafka + kmsvc), and workflow orchestration (Temporal).
## Use Cases & Architecture
**Why this stack?**
This homelab replicates **production-grade cloud-native infrastructure** on bare metal, enabling:
1. **Learning & Prototyping** — Test distributed systems patterns (HA databases, event-driven messaging, GitOps workflows) before deploying to cloud
2. **Self-Hosted Services** — Run applications (Story Crater, etc.) with zero cloud lock-in; full control over data, compliance, and networking
3. **Infrastructure as Code** — Git-driven cluster state via Helmfile + Forgejo Actions + Argo CD; every change is auditable and reproducible
4. **Observability Sandbox** — Experiment with Prometheus metrics, Loki log aggregation, and custom Grafana dashboards at scale
**Typical workflow:**
```
Developer pushes to Forgejo (git forge)
Forgejo Actions CI runs tests + builds OCI image
Image pushed to Forgejo registry (private, on-cluster)
Argo CD detects deployment repo change (pull-based GitOps)
New pods roll out; Grafana alerts on errors/latency
Temporal workflows coordinate long-running operations (e.g., async jobs)
Kafka queues decouple services (fire-and-forget messaging)
All logs + metrics centralized in Grafana for debugging
```
**Architecture principles:**
- **Immutable OS** — Talos Linux (no SSH, declarative machine configs)
- **No external dependencies** — All data stored locally (MinIO, CloudNativePG, Longhorn)
- **High availability** — 3-replica databases, multi-node storage, cross-AZ readiness (on bare metal: cross-rack affinity)
- **Federated identity** — Single Authentik OIDC provider for all services (Grafana, MinIO, Forgejo, Argo CD)
- **Secrets at rest** — Vault + encrypted etcd; credentials never in logs or ConfigMaps
- **Infrastructure-as-code** — Every service deployed via Helmfile; one `helmfile apply` recovers from total failure
## Quick Start — Deploying the Cluster
### 1. Bootstrap Talos Nodes
Bootstrap each Talos node with your cluster schematic (see `CLAUDE.md` or `README` step 18).
### 2. Set Up Secrets
All secrets are managed via environment variables sourced from `.env` (gitignored). The helmfile template expands them at deploy time.
**Step 1: Copy the template**
```bash
cp .env.example .env
```
**Step 2: Populate required secrets**
Edit `.env` and fill in cluster configuration. See `.env.example` for all options:
```bash
# Cluster configuration
CLUSTER_DOMAIN=riotpiao.homelab.com # Your cluster domain
POSTGRES_HOST=ddb-cluster-rw.ddb.svc.cluster.local
MINIO_ENDPOINT=minio.storage.svc.cluster.local:9000
KAFKA_BOOTSTRAP=kmsvc-kafka-bootstrap.sqs.svc.cluster.local:9092
REDIS_ADDR=kmsvc-redis-master.sqs.svc.cluster.local:6379
# Service credentials (generate with: openssl rand -hex 32)
MINIO_ROOT_PASSWORD=<random>
GRAFANA_ADMIN_PASSWORD=<random>
AUTHENTIK_SECRET_KEY=<random>
AUTHENTIK_BOOTSTRAP_PASSWORD=<random>
AUTHENTIK_PG_PASSWORD=<random>
```
**Step 3: Load and deploy**
```bash
# Load .env into current shell
vsource .env
# Preview all changes before deployment
helmfile diff
# Deploy the entire stack
helmfile apply
```
### 3. Verify Deployment
```bash
# Check all pods are running
kubectl get pods -A
# Confirm key services are ready
kubectl wait deploy/authentik-server -n iam --for=condition=Available --timeout=300s
kubectl wait deploy/grafana -n logging --for=condition=Available --timeout=300s
# Access Grafana
make pf-grafana # localhost:3000 (login: admin / GRAFANA_ADMIN_PASSWORD)
```
### 4. First-Time Access
**Default Credentials:**
- **Authentik:** https://authentik.$(CLUSTER_DOMAIN) → login: `akadmin` / `AUTHENTIK_BOOTSTRAP_PASSWORD` (change immediately)
- **Grafana:** https://grafana.$(CLUSTER_DOMAIN) → login: `admin` / `GRAFANA_ADMIN_PASSWORD`
- **MinIO:** https://minio.$(CLUSTER_DOMAIN) → login: `MINIO_ROOT_USER` / `MINIO_ROOT_PASSWORD`
- **Argo CD:** https://argocd.$(CLUSTER_DOMAIN) → login via Authentik OIDC
- **Forgejo:** https://forgejo.$(CLUSTER_DOMAIN) → login via Authentik OIDC
**Next Steps:**
1. Change default passwords in each service
2. Configure OIDC redirects (see `k8s/talos-iam/` for details)
3. Set up GitOps: push infrastructure to Forgejo, configure Argo CD
4. Review dashboards in Grafana (Prometheus + Loki)
---
## Architecture
```
192.168.1.0/24 (LAN)
┌──────────────────────┼──────────────────────┐
│ │ │
192.168.1.* 192.168.1.* 192.168.1.*
┌────────────────┐ ┌────────────────┐ ┌────────────────┐
│ talos-cp-1 │ │ talos-worker-1 │ │ talos-worker-2 │
│ Control-Plane │ │ Worker │ │ Worker │
│ + Workloads │ │ (Storage) │ │ (Storage) │
│ (az-a) │ │ (az-b) │ │ (az-c) │
├────────────────┤ ├────────────────┤ ├────────────────┤
│ Pods: │ │ Pods: │ │ Pods: │
│ • ingress-nginx│ │ • kube-system │ │ • kube-system │
│ • authentik │ │ • storage │ │ • storage │
│ • vault │ │ └─ minio-2 │ │ └─ minio-3 │
│ • logging │ │ │ │ │
│ ├─ loki │ │ │ │ │
│ ├─ promtail │ │ │ │ │
│ └─ grafana │ │ │ │ │
│ • monitoring │ │ │ │ │
│ ├─ prom │ │ │ │ │
│ └─ blackbox │ │ │ │ │
│ • storage │ │ │ │ │
│ └─ minio-1 │ │ │ │ │
│ • cicd │ │ │ │ │
│ ├─ forgejo │ │ │ │ │
│ ├─ argocd │ │ │ │ │
│ └─ runner │ │ │ │ │
│ • sqs │ │ • sqs │ │ • sqs │
│ ├─ kafka-0 │ │ ├─ kafka-1 │ │ ├─ kafka-2 │
│ └─ kmsvc │ │ └─ redis │ │ │
│ • ddb │ │ • ddb │ │ • ddb │
│ └─ postgres-0│ │ └─ postgres-1│ │ └─ postgres-2│
│ • temporal │ │ • temporal │ │ • temporal │
│ ├─ server │ │ ├─ cassandra │ │ ├─ cassandra │
│ └─ cassandra │ │ │ -1 │ │ │ -2 │
│ -0 │ │ └─ (replica) │ │ └─ (replica) │
│ • llm │ │ │ │ │
│ └─ ollama │ │ │ │ │
└────────────────┘ └────────────────┘ └────────────────┘
Replication & Fault Tolerance:
Data Layer:
• MinIO: minio-1 ↔ minio-2 ↔ minio-3 (3-way active-active S3)
• PostgreSQL: postgres-0 ↔ postgres-1 ↔ postgres-2 (primary + 2 standbys, HA streaming replication)
• Kafka: kafka-0 ↔ kafka-1 ↔ kafka-2 (3 brokers, RF=3, min-ISR=2, cross-AZ)
• Temporal: cassandra-0 ↔ cassandra-1 ↔ cassandra-2 (3-node distributed)
• Loki: loki → MinIO (chunks stored in s3://loki-chunks, 10-day retention)
Single-Replica Services (Protected by PodDisruptionBudget minAvailable=1):
• Observability: Grafana, Prometheus, Loki (Recreate strategy for RWO PVCs)
• IAM: Authentik, Vault (Recreate strategy for RWO PVCs)
• CI/CD: Argo CD server/repo-server, Forgejo (Recreate strategy for RWO PVCs)
• LLM: Ollama
Networking:
Remote access: UR_OWN.duckdns.org → home IP → cp-1
```
## Stack
| Layer | Technology | Namespace | Purpose |
|-------|-----------|-----------|---------|
| OS | Talos Linux v1.13.3 | — | Immutable, Kubernetes-native OS |
| Kubernetes | v1.36.1 | — | Container orchestration |
| CNI | Cilium (eBPF) | kube-system | Networking, replaces kube-proxy |
| Ingress | Nginx Ingress Controller | ingress-nginx | Reverse proxy, hostname-based routing |
| Block Storage | Longhorn v1.7.0 | longhorn-system | Default StorageClass |
| Object Store | MinIO (multi-AZ) | storage | S3-compatible, site-replicated across az-a/az-b |
| IAM / SSO | Authentik | iam | OIDC provider for Grafana, MinIO, Forgejo, Argo CD |
| Secret Store | HashiCorp Vault | iam | KV secrets backend, JWT auth via Authentik |
| Git Forge *(planned)* | Forgejo | forge | Git server, built-in OCI registry, Actions CI |
| CI Runner *(planned)* | Forgejo Actions + DinD | cicd | Privileged build pod; images pushed to Forgejo OCI |
| CD *(planned)* | Argo CD | argocd | Pull-based GitOps; never holds kubeconfig in CI |
| Log Backend | Loki (SingleBinary) | logging | 10-day retention, backed by MinIO |
| Log Collector | Promtail (DaemonSet) | logging | Scrapes pod logs + Talos journal |
| Metrics | kube-prometheus-stack | monitoring | Prometheus + node-exporter + kube-state-metrics |
| Log/Metrics UI | Grafana | logging | Dashboards for Loki + Prometheus |
| Cluster UI | Portainer CE | dashboard | Container/workload management UI |
| LLM Inference | Ollama | llm | Local LLM model serving (open-source models) |
## Repository Structure
```
homelab/
├── helmfile.yaml # Single source of truth — deploys everything
├── .env.example # Required env vars template (copy to .env, gitignored)
├── cluster-config/ # Talos + Kubernetes bootstrap
│ ├── cilium-values.yaml
│ ├── longhorn_bootstrap.sh
│ ├── controlplane.yaml # gitignored — contains secrets
│ ├── worker-1.yaml # gitignored
│ ├── secrets.yaml # gitignored
│ ├── kubeconfig # gitignored
│ └── talosconfig # gitignored
├── k8s/
│ ├── ingress/ # Nginx Ingress Controller + all Ingress rules
│ │ ├── nginx-values.yaml
│ │ └── ingress.yaml
│ ├── storage/ # MinIO multi-AZ object store
│ │ ├── minio-az-a-values.yaml
│ │ ├── minio-az-b-values.yaml
│ │ ├── minio-az-a-pvc.yaml
│ │ ├── minio-service.yaml
│ │ ├── minio-legacy-alias.yaml
│ │ └── minio-replication-job.yaml
│ ├── logging/ # Observability stack
│ │ ├── loki-values.yaml
│ │ ├── promtail-values.yaml
│ │ └── grafana-values.yaml
│ ├── monitoring/ # Prometheus stack + alerting + Grafana dashboards-as-code
│ │ ├── prometheus-values.yaml
│ │ ├── blackbox-exporter-values.yaml # Active uptime probes (feeds Service Availability dashboard)
│ │ ├── ingress-alerts.yaml # PrometheusRule: ingress 5xx rate, p95 latency
│ │ └── dashboards/ # ConfigMaps picked up live by Grafana's sidecar
│ │ ├── service-availability.yaml # Uptime probes + cert expiry (operator glance)
│ │ ├── service-golden-signals.yaml # Latency & Golden Signals (ingress RED)
│ │ ├── service-internals.yaml # Per-service deep-dive (MinIO/Forgejo/Argo CD/Vault/Longhorn/certs)
│ │ ├── kube-controller-health.yaml # API server RED + kube-state-metrics controller-health proxy
│ │ ├── hardware-overview.yaml # Per-node CPU/mem/disk/network/load summary
│ │ └── control-plane-logs.yaml # kube-system + add-on logs (Loki)
│ ├── portainer/ # Portainer CE
│ │ └── portainer-values.yaml
│ ├── talos-iam/ # Authentik + Vault IAM
│ │ ├── authentik-values.yaml
│ │ ├── vault-values.yaml
│ │ ├── setup_vault.sh # One-time Vault init (not replaced by Helmfile)
│ │ └── provision_oidc.py # Authentik OIDC provisioning
│ ├── coredns/ # CoreDNS hostname rewrites (in-cluster DNS)
│ │ └── coredns-configmap.yaml
│ ├── talos-ci-cd/ # CI/CD stack (planned — not yet applied)
│ │ ├── talos_version_control.html # Implementation plan + build runbook
│ │ ├── forgejo-values.yaml # Forgejo Helm values (gitea-charts/gitea)
│ │ ├── argocd-values.yaml # Argo CD Helm values
│ │ └── charts/forgejo-runner/ # Local Helm chart for the Actions runner
│ │ ├── Chart.yaml
│ │ ├── values.yaml
│ │ └── templates/
│ │ ├── deployment.yaml # Runner + DinD sidecar, Recreate strategy
│ │ ├── pvc.yaml # runner-reg (1 Gi) + runner-dind (30 Gi)
│ │ └── networkpolicy.yaml # Egress: forge ns + DNS + internet only
│ └── duckdns/ # DuckDNS DDNS updater CronJob
├── talos-cli/ # Rust CLI for Vault secret access
├── project_context.md # Authoritative live-state reference
├── refine_cluster.md # Known-issues runbook
└── LOG.md # Append-only change journal
```
## Access
### Discover Ingress LoadBalancer IP
```bash
# Find the external IP assigned by Cilium LB-IPAM
kubectl get svc -n ingress-nginx ingress-nginx
# Example output:
# LoadBalancer IP: 192.168.1.160 (Cilium LB-IPAM assignment)
```
Add to `/etc/hosts` on every client machine (Mac/Linux):
```
# WireGuard access (remote — via talos-cp-1)
10.6.0.1 grafana.riotpiao.homelab.com authentik.riotpiao.homelab.com vault.riotpiao.homelab.com minio.riotpiao.homelab.com prometheus.riotpiao.homelab.com portainer.riotpiao.homelab.com longhorn.riotpiao.homelab.com loki.riotpiao.homelab.com forgejo.riotpiao.homelab.com
# LAN access (on the home network — use actual LoadBalancer IP from above)
192.168.1.160 grafana.riotpiao.homelab.com authentik.riotpiao.homelab.com vault.riotpiao.homelab.com minio.riotpiao.homelab.com prometheus.riotpiao.homelab.com portainer.riotpiao.homelab.com longhorn.riotpiao.homelab.com loki.riotpiao.homelab.com forgejo.riotpiao.homelab.com
```
**Note:** `192.168.1.160` is an example Cilium LB-IPAM assignment. Verify with `kubectl get svc -n ingress-nginx ingress-nginx`.
Then access services at:
| Service | URL | Credentials |
|---------|-----|-------------|
| Grafana | http://grafana.riotpiao.homelab.com | admin / `GRAFANA_ADMIN_PASSWORD` or Authentik SSO |
| Authentik | http://authentik.riotpiao.homelab.com | akadmin / see `.env` |
| Vault | http://vault.riotpiao.homelab.com | root token / see `setup_vault.sh` output |
| MinIO console | http://minio.riotpiao.homelab.com | `MINIO_ROOT_USER` / `MINIO_ROOT_PASSWORD` |
| Prometheus | http://prometheus.riotpiao.homelab.com | no auth |
| Portainer | http://portainer.riotpiao.homelab.com | set on first visit |
| Longhorn | http://longhorn.riotpiao.homelab.com | no auth |
| Forgejo *(planned)* | https://forgejo.forge.riotpiao.homelab.com | `rock` / `FORGEJO_ADMIN_PASSWORD`, or Authentik SSO |
| Argo CD *(planned)* | `kubectl port-forward -n argocd svc/argocd-server 8080:443` | Authentik SSO (admins only) |
Grafana → "Homelab" folder has the operator dashboards (sidecar-loaded from `k8s/monitoring/dashboards/`, no restart needed on change):
- **Service Availability & Certificate Expiration** — uptime probes + cert-manager expiry
- **Latency & Golden Signals** — ingress request rate/error %/p50-p99 latency
- **Kube-Controller Health** — API server RED metrics + kube-state-metrics controller-health signals
- **Hardware Statistics** — per-node CPU/mem/disk/network/load
- **Service Internals** — per-service deep-dive (MinIO/Forgejo/Argo CD/Vault/Longhorn)
## Deploy
```bash
# 1. Install helmfile (once)
brew install helmfile
# 2. Set credentials
cp .env.example .env
# edit .env with your passwords
# 3. Deploy everything
helmfile apply
# Deploy a single stack
helmfile apply -l namespace=logging
helmfile apply -l name=grafana
helmfile apply -l namespace=ingress-nginx
# Preview changes before applying
helmfile diff
```
## Bootstrap Order (fresh cluster)
```
1. Provision nodes: talosctl apply-config (make apply-cp / apply-worker-new)
2. Bootstrap Kubernetes: talosctl bootstrap
3. Install Cilium: helm install cilium -f cluster-config/cilium-values.yaml
4. Install Longhorn: bash cluster-config/longhorn_bootstrap.sh
5. Deploy everything else: helmfile apply
6. Vault init (one-time): bash k8s/talos-iam/setup_vault.sh
7. Authentik OIDC: python3 k8s/talos-iam/provision_oidc.py
8. Label worker: kubectl label node talos-worker-1 node-role.kubernetes.io/worker=
# ── CI/CD (planned — run after step 8) ─────────────────────────────────────
9. Private CA + TLS: see k8s/talos-ci-cd/talos_version_control.html §12 Block 0
10. Deploy Forgejo + runner: helmfile apply -l name=forgejo && helmfile apply -l name=forgejo-runner
11. Authentik SSO for CI: Forgejo + Argo CD OIDC (§12 Block 1.5)
12. Talos node CA trust: talosctl patch machineconfig (§12 Block 2)
13. Deploy Argo CD: helmfile apply -l name=argocd (§12 Block 4)
14. Wire deploy repo: argocd app create + push first manifests (§12 Block 4)
```
## IAM & Auth Flow
Authentik is the central OIDC identity provider. Vault stores secrets and delegates authentication back to Authentik.
```
User / core-cli
│ OAuth2 / OIDC
Authentik (authentik.riotpiao.homelab.com)
├── grafana app → Grafana OIDC login (group → Admin/Viewer role)
├── minio app → MinIO OIDC login (group → readwrite/readonly policy)
├── vault-browser → Vault UI OIDC login / `vault login -method=oidc`
└── core-cli-shell → CLI device code flow (public client, no secret)
│ JWKS endpoint for JWT validation
HashiCorp Vault (vault.riotpiao.homelab.com)
├── auth/jwt — core-cli authenticates with device code JWT
├── auth/oidc — browser/UI login via Authentik
└── secret/ — KV v2: mcp/*, cluster/*, cloud/*
```
**talos-cli device code login:**
```bash
core secrets login # prints URL + code → approve in browser → Vault token cached
core put cluster/DUCKDNS_TOKEN DUCKDNS_TOKEN="abc" # field name = variable name, never `value`
```
**One-time IAM setup (after `helmfile apply`):**
```bash
# 1. Provision OIDC apps and groups in Authentik
GRAFANA_URL=http://grafana.riotpiao.homelab.com \
MINIO_URL=http://minio.riotpiao.homelab.com \
python3 k8s/talos-iam/provision_oidc.py
# 2. Init Vault, wire JWT + OIDC auth, seed secrets
bash k8s/talos-iam/setup_vault.sh
```
**CoreDNS hostname rewrites** (`k8s/coredns/coredns-configmap.yaml`) ensure in-cluster pods (Grafana, Vault, Forgejo runner, Argo CD) resolve internal hostnames to cluster services, avoiding hairpin NAT through LB IPs.
## CI/CD Pipeline *(planned)*
> **Implementation plan & exact build commands:** [`k8s/talos-ci-cd/talos_version_control.html`](k8s/talos-ci-cd/talos_version_control.html)
### Components
| Component | Helm chart | Namespace | Notes |
|-----------|-----------|-----------|-------|
| Forgejo | `gitea-charts/gitea` (Forgejo image override) | `forge` | Git + OCI registry + Actions engine; SQLite on Longhorn PVC; `strategy: Recreate` |
| Forgejo runner | local chart `charts/forgejo-runner` | `cicd` | DinD sidecar; PodSecurity privileged; NetworkPolicy fenced |
| Argo CD | `argo/argo-cd` | `argocd` | Pull-based CD; single replica; no external ingress (port-forward only) |
### Pipeline flow
```
Developer
│ git push
Forgejo (forge ns) ─── webhook ───► Runner (cicd ns)
│ │
│ SSO login (Authentik OIDC) │ ACTIONS_RUNTIME_TOKEN → git checkout
▼ │ ci-registry-token (ci-bot) → docker push → Forgejo OCI
Forgejo UI / Argo CD UI │ ci-deploy-token (ci-bot) → git commit → rock/deploy
Forgejo (rock/deploy repo) ◄─────────────┘
│ argocd-bot token (repo:read, poll every 3 min)
Argo CD (argocd ns)
│ kubectl apply (cluster-admin ServiceAccount — never in CI)
K8s workloads (images from Forgejo OCI)
```
### Key security decisions
- **No kubeconfig in CI.** The runner can only push commits + OCI images. Argo CD bridges the gap autonomously.
- **Scoped machine credentials.** `ci-bot` tokens are narrowly scoped: `package:write` for OCI, `repo:write` on `rock/deploy` only. A compromised runner cannot read other repos or call the K8s API.
- **Private CA TLS.** Forgejo self-terminates HTTPS with a homelab CA (EC P-256, 10-year). The CA cert is distributed to Talos nodes via `machineconfig` patch and to the runner via K8s Secret. `ca.key` never enters the cluster.
- **Authentik SSO for humans.** All interactive logins (Forgejo UI, Argo CD UI) route through Authentik. `homelab-admins` group → Forgejo admin + Argo CD `role:admin`; `homelab-devs` group → Forgejo user + no Argo CD access.
- **Forgejo LB IP pinned.** Cilium LB-IPAM annotation `io.cilium/lb-ipam-ips: <LB_IP>` fixes the Forgejo LoadBalancer IP so the TLS SAN and DNS entries never need updating. Configure in `k8s/talos-ci-cd/forgejo-values.yaml`.
### Workflow file location
Forgejo Actions uses GitHub Actions syntax. Workflow files live in `.forgejo/workflows/` in each source repo:
```
rock/source/
└── .forgejo/
└── workflows/
├── ci.yml # build + test + push OCI image
└── cd.yml # on: push to main → bump image tag in rock/deploy
```
## Log Data Flow
```
Pods / Talos journal (both nodes)
Promtail (DaemonSet, all nodes) reads /var/log/pods + /var/log/journal
Loki (logging ns) indexes + compacts, 10-day retention
│ stores chunks via S3
MinIO frontend service minio.storage.svc.cluster.local:9000
(active-active, round-robin)
minio-az-a (cp-1) ↔ minio-az-b (worker-1) ↔ minio-az-c (worker-2)
3-way site replication (bidirectional, automatic)
Grafana (logging ns) queries Loki + Prometheus via dashboards
Nginx Ingress → grafana.riotpiao.homelab.com browser access
```
## Example Applications & Workloads
This cluster runs production-like applications and infrastructure services:
### Story Crater Backend
**Type:** Distributed message-driven application
**Namespace:** `story-crater-backend`
**Architecture:**
- gRPC server + REST gateway (Envoy)
- PostgreSQL database (story_crater in CloudNativePG cluster)
- Kafka topic consumers (via kmsvc message queue)
- Temporal workflow integration for long-running operations
- Prometheus metrics (processed messages, latency, errors)
- Grafana dashboard (message throughput, queue depth, LLM token usage)
**Typical flow:**
```
POST /api/messages → gRPC handler
Publish to Kafka topic (Story Crater queue)
Message consumer processes async (may trigger LLM inference)
Results stored in PostgreSQL + emitted as event
Prometheus increments counters (story_crater_messages_handled_total)
Grafana renders message throughput + duration histograms
```
### Infrastructure Services (Essential)
| Service | Purpose | Namespace | Example Use |
|---------|---------|-----------|-------------|
| **Authentik** | OIDC identity provider | iam | User login, group management, SSO for Grafana/MinIO/Forgejo |
| **Vault** | Secrets backend | iam | Database passwords, API keys, JWT token validation |
| **CloudNativePG** | PostgreSQL 3-replica cluster | ddb | Auth database (Authentik), app database (Story Crater) |
| **Loki** | Log aggregation | logging | Centralize pod logs, Talos kernel logs (10-day retention) |
| **Prometheus** | Metrics collection | monitoring | Scrape kube-state-metrics, kubelet, ServiceMonitors (every 30s) |
| **Grafana** | Observability dashboards | logging | Query Prometheus + Loki, alert on latency/error spikes |
| **Kafka + kmsvc** | Message queue | sqs | Decouple services, async job processing, at-least-once delivery |
| **MinIO** | S3-compatible object store | storage | Loki log chunks backend, Vault unseal keys, config backups |
| **Longhorn** | Persistent block storage | longhorn-system | All PVCs (Postgres replicas, Kafka broker disks, MinIO) |
### Development Services (Optional)
| Service | Purpose | Namespace | Example Use |
|---------|---------|-----------|-------------|
| **Forgejo** | Self-hosted git + OCI registry | cicd | Version control, CI Actions runner, private Docker images |
| **Argo CD** | Pull-based GitOps | cicd | Continuous deployment (deployment repo → Kubernetes) |
| **Temporal** | Workflow orchestration | temporal | Schedule long-running jobs, retry logic, state machines |
| **Portainer** | Container management UI | dashboard | Pod inspection, image management, quick debugging |
### Monitoring Example: Dashboard Walk-Through
Open **Grafana****Homelab** folder → **"Story Crater Backend — Service Overview"**:
**Row A — Availability & Golden Signals**
- Requests/sec (blue = success, red = errors)
- Error rate % (goal: < 0.1%)
- p50/p95/p99 latency (goal: p99 < 500ms)
- Alert: If error rate > 1% for 5 min, page on-call
**Row B — Resource Usage**
- CPU (request/limit)
- Memory (request/limit)
- Restart count (goal: 0; alerts if > 2)
**Row C — Domain-Specific Metrics** (Story Crater only)
- Messages processed/sec (bucketed by status: success, dlq, retry)
- Queue depth (Kafka partitions lag)
- Dedup window retention (FIFO redelivery tracking)
- LLM inference tokens used/sec
- External API call latency (e.g., OpenAI)
**Row D — Logs**
- Live Loki panel: filter by pod + search for errors
- Example: `{namespace="story-crater-backend"} | json | level="error"`
**Row E — Alert Status** (if SLO defined)
- Burn rate (if consuming SLO budget)
- Example: "30-day availability SLO = 99.5%; current burn rate = 0.2x"
**Row F — Related Dashboards**
- Link to Kafka dashboard (queue depth)
- Link to PostgreSQL dashboard (story_crater DB)
- Link to Temporal dashboard (workflow execution times)
## Adding Hardware to the Cluster
### Step 1 — Get the Talos image
The image must include the same extensions as the existing nodes (`iscsi-tools` + `util-linux-tools`).
Download from the Image Factory using the cluster's schematic ID:
| Format | Use case | URL |
|--------|----------|-----|
| ISO | USB boot (recommended for bare metal) | `https://factory.talos.dev/image/613e1592.../v1.13.3/metal-amd64.iso` |
| RAW disk image | Write directly to drive via another machine | `https://factory.talos.dev/image/613e1592.../v1.13.3/metal-amd64.raw.xz` |
| PXE / iPXE | Network boot — no USB needed | `https://factory.talos.dev/image/613e1592.../v1.13.3/kernel-amd64` |
> Full schematic ID: `613e1592b2da41ae5e265e8789429f22e121aab91cb4deb6bc3c0b6262961245`
---
### Option A — USB Bootable (recommended)
```bash
curl -Lo talos-worker.iso \
"https://factory.talos.dev/image/613e1592b2da41ae5e265e8789429f22e121aab91cb4deb6bc3c0b6262961245/v1.13.3/metal-amd64.iso"
sudo dd if=talos-worker.iso of=/dev/sdX bs=4M status=progress && sync
```
---
### Option B — In-Memory (diskless / RAM boot via PXE)
Talos runs entirely from RAM. Useful for temporary nodes or hardware where you don't want to touch the existing OS.
```bash
# Boot via PXE pointing to:
# Kernel: https://factory.talos.dev/image/613e1592.../v1.13.3/kernel-amd64
# Initrd: https://factory.talos.dev/image/613e1592.../v1.13.3/initramfs-amd64.xz
# Cmdline: talos.platform=metal
```
> Note: in-memory nodes lose state on reboot. Not suitable for Longhorn storage nodes.
---
### Option C — Direct Disk Image (headless / remote)
```bash
xz -d talos-worker.raw.xz
sudo dd if=talos-worker.raw of=/dev/sda bs=4M status=progress && sync
```
---
### Step 2 — Discover hardware in maintenance mode
```bash
# Scan your LAN for the new node in maintenance mode
nmap -sn 192.168.1.0/24
# Note: Replace with your actual subnet (e.g., 10.0.1.0/24)
# Discover available disks on the node
talosctl --nodes <maintenance-ip> --talosconfig cluster-config/talosconfig disks --insecure
```
---
### Step 3 — Prepare the worker config
```bash
cp cluster-config/worker-1.yaml cluster-config/worker-N.yaml
```
Edit exactly these four fields:
| Field | Value |
|-------|-------|
| `machine.network.hostname` | `talos-worker-N` |
| `machine.network.interfaces[0].addresses` | `192.168.1.16N/24` |
| `machine.install.disk` | disk path from Step 2 |
| `machine.nodeLabels.topology.kubernetes.io/zone` | `az-N` |
---
### Step 4 — Apply config
```bash
make apply-worker-new N=<num> WN_IP=<maintenance-ip>
```
---
### Step 5 — Persist the worker IP
```bash
# Substitute <WN_MAINTENANCE_IP> with the actual IP discovered in Step 2
echo 'export W<N>_IP=192.168.1.<last-octet>' >> ~/.zshrc && source ~/.zshrc
```
---
### Step 6 — Set the node-role label
```bash
kubectl label node talos-worker-N node-role.kubernetes.io/worker=
```
---
### Step 7 — Verify
```bash
kubectl get nodes -w
kubectl describe node talos-worker-N | grep -A10 Labels
```
---
### Step 8 — What automatically extends to the new node
| Service | Behaviour |
|---------|-----------|
| Cilium | New pod scheduled automatically |
| Promtail | DaemonSet — starts immediately |
| Nginx Ingress | DaemonSet — starts immediately, port 80/443 available on new node |
| Longhorn | Detects new node, available for replica scheduling |
| MinIO / Loki / Grafana | Stay on existing node (single-replica Deployments) |