Story Crater Bot e6ada95b39 fix(api): retire Kong key-auth on model routes; agent-pod builds agent-manager fork + ships coordinator.js
Kong key-auth rejected the Authorization: Bearer header every OpenAI-SDK-compatible client sends (verified: raw apikey header works, Bearer doesn't), so it's commented out and stripped from every llm-routes.yaml annotation until there's a Bearer-compatible fix. agent-pod now clones and builds the agent-manager fork from source at container start (no prebuilt binary shipped -- wrong arch and over ConfigMap's size cap) and ships coordinator.js alongside hub.js, so multiple repos can run the pipeline concurrently in one pod via kubectl exec. hub.js keeps its existing role as the container's foreground process, unchanged.
2026-08-18 17:50:52 -07:00

Homelab Kubernetes Cluster

A bare-metal three-node Kubernetes cluster running Talos Linux with 18 Helm releases across 22 namespaces. Includes distributed storage (MinIO + Longhorn), full observability (Prometheus + Grafana + Loki), federated SSO (Authentik OIDC), secrets management (Vault), CI/CD (Forgejo + Argo CD), messaging (Kafka + kmsvc), and workflow orchestration (Temporal).

Use Cases & Architecture

Why this stack?

This homelab replicates production-grade cloud-native infrastructure on bare metal, enabling:

  1. Learning & Prototyping — Test distributed systems patterns (HA databases, event-driven messaging, GitOps workflows) before deploying to cloud
  2. Self-Hosted Services — Run applications (Story Crater, etc.) with zero cloud lock-in; full control over data, compliance, and networking
  3. Infrastructure as Code — Git-driven cluster state via Helmfile + Forgejo Actions + Argo CD; every change is auditable and reproducible
  4. Observability Sandbox — Experiment with Prometheus metrics, Loki log aggregation, and custom Grafana dashboards at scale

Typical workflow:

Developer pushes to Forgejo (git forge)
    ↓
Forgejo Actions CI runs tests + builds OCI image
    ↓
Image pushed to Forgejo registry (private, on-cluster)
    ↓
Argo CD detects deployment repo change (pull-based GitOps)
    ↓
New pods roll out; Grafana alerts on errors/latency
    ↓
Temporal workflows coordinate long-running operations (e.g., async jobs)
    ↓
Kafka queues decouple services (fire-and-forget messaging)
    ↓
All logs + metrics centralized in Grafana for debugging

Architecture principles:

  • Immutable OS — Talos Linux (no SSH, declarative machine configs)
  • No external dependencies — All data stored locally (MinIO, CloudNativePG, Longhorn)
  • High availability — 3-replica databases, multi-node storage, cross-AZ readiness (on bare metal: cross-rack affinity)
  • Federated identity — Single Authentik OIDC provider for all services (Grafana, MinIO, Forgejo, Argo CD)
  • Secrets at rest — Vault + encrypted etcd; credentials never in logs or ConfigMaps
  • Infrastructure-as-code — Every service deployed via Helmfile; one helmfile apply recovers from total failure

Quick Start — Deploying the Cluster

1. Bootstrap Talos Nodes

Bootstrap each Talos node with your cluster schematic (see CLAUDE.md or README step 18).

2. Set Up Secrets

All secrets are managed via environment variables sourced from .env (gitignored). The helmfile template expands them at deploy time.

Step 1: Copy the template

cp .env.example .env

Step 2: Populate required secrets Edit .env and fill in cluster configuration. See .env.example for all options:

# Cluster configuration
CLUSTER_DOMAIN=riotpiao.com           # Your cluster domain
POSTGRES_HOST=ddb-cluster-rw.ddb.svc.cluster.local
MINIO_ENDPOINT=minio.storage.svc.cluster.local:9000
KAFKA_BOOTSTRAP=kmsvc-kafka-bootstrap.sqs.svc.cluster.local:9092
REDIS_ADDR=kmsvc-redis-master.sqs.svc.cluster.local:6379

# Service credentials (generate with: openssl rand -hex 32)
MINIO_ROOT_PASSWORD=<random>
GRAFANA_ADMIN_PASSWORD=<random>
AUTHENTIK_SECRET_KEY=<random>
AUTHENTIK_BOOTSTRAP_PASSWORD=<random>
AUTHENTIK_PG_PASSWORD=<random>

Step 3: Load and deploy

# Load .env into current shell
vsource .env

# Preview all changes before deployment
helmfile diff

# Deploy the entire stack
helmfile apply

3. Verify Deployment

# Check all pods are running
kubectl get pods -A

# Confirm key services are ready
kubectl wait deploy/authentik-server -n iam --for=condition=Available --timeout=300s
kubectl wait deploy/grafana -n logging --for=condition=Available --timeout=300s

# Access Grafana
make pf-grafana    # localhost:3000 (login: admin / GRAFANA_ADMIN_PASSWORD)

4. First-Time Access

Default Credentials:

Next Steps:

  1. Change default passwords in each service
  2. Configure OIDC redirects (see k8s/talos-iam/ for details)
  3. Set up GitOps: push infrastructure to Forgejo, configure Argo CD
  4. Review dashboards in Grafana (Prometheus + Loki)

Architecture

                         192.168.1.0/24 (LAN)
                                │
         ┌──────────────────────┼──────────────────────┐
         │                      │                      │
    192.168.1.*          192.168.1.*          192.168.1.*
  ┌────────────────┐      ┌────────────────┐    ┌────────────────┐
  │  talos-cp-1    │      │ talos-worker-1 │    │ talos-worker-2 │
  │ Control-Plane  │      │  Worker        │    │  Worker        │
  │   + Workloads  │      │  (Storage)     │    │  (Storage)     │
  │   (az-a)       │      │   (az-b)       │    │   (az-c)       │
  ├────────────────┤      ├────────────────┤    ├────────────────┤
  │ Pods:          │      │ Pods:          │    │ Pods:          │
  │ • ingress-nginx│      │ • kube-system  │    │ • kube-system  │
  │ • authentik    │      │ • storage      │    │ • storage      │
  │ • vault        │      │   └─ minio-2   │    │   └─ minio-3   │
  │ • logging      │      │                │    │                │
  │   ├─ loki      │      │                │    │                │
  │   ├─ promtail  │      │                │    │                │
  │   └─ grafana   │      │                │    │                │
  │ • monitoring   │      │                │    │                │
  │   ├─ prom      │      │                │    │                │
  │   └─ blackbox  │      │                │    │                │
  │ • storage      │      │                │    │                │
  │   └─ minio-1   │      │                │    │                │
  │ • cicd         │      │                │    │                │
  │   ├─ forgejo   │      │                │    │                │
  │   ├─ argocd    │      │                │    │                │
  │   └─ runner    │      │                │    │                │
  │ • sqs          │      │ • sqs          │    │ • sqs          │
  │   ├─ kafka-0   │      │   ├─ kafka-1   │    │   ├─ kafka-2   │
  │   └─ kmsvc     │      │   └─ redis     │    │                │
  │ • ddb          │      │ • ddb          │    │ • ddb          │
  │   └─ postgres-0│      │   └─ postgres-1│    │   └─ postgres-2│
  │ • temporal     │      │ • temporal     │    │ • temporal     │
  │   ├─ server    │      │   ├─ cassandra │    │   ├─ cassandra │
  │   └─ cassandra │      │   │   -1       │    │   │   -2       │
  │       -0       │      │   └─ (replica) │    │   └─ (replica) │
  │ • llm          │      │                │    │                │
  │   └─ ollama    │      │                │    │                │
  └────────────────┘      └────────────────┘    └────────────────┘

Replication & Fault Tolerance:
  Data Layer:
    • MinIO:      minio-1 ↔ minio-2 ↔ minio-3  (3-way active-active S3)
    • PostgreSQL: postgres-0 ↔ postgres-1 ↔ postgres-2  (primary + 2 standbys, HA streaming replication)
    • Kafka:      kafka-0 ↔ kafka-1 ↔ kafka-2  (3 brokers, RF=3, min-ISR=2, cross-AZ)
    • Temporal:   cassandra-0 ↔ cassandra-1 ↔ cassandra-2  (3-node distributed)
    • Loki:       loki → MinIO (chunks stored in s3://loki-chunks, 10-day retention)
  
  Single-Replica Services (Protected by PodDisruptionBudget minAvailable=1):
    • Observability: Grafana, Prometheus, Loki (Recreate strategy for RWO PVCs)
    • IAM: Authentik, Vault (Recreate strategy for RWO PVCs)
    • CI/CD: Argo CD server/repo-server, Forgejo (Recreate strategy for RWO PVCs)
    • LLM: Ollama

Networking:
  Remote access: UR_OWN.duckdns.org → home IP → cp-1

Stack

Layer Technology Namespace Purpose
OS Talos Linux v1.13.3 Immutable, Kubernetes-native OS
Kubernetes v1.36.1 Container orchestration
CNI Cilium (eBPF) kube-system Networking, replaces kube-proxy
Ingress Nginx Ingress Controller ingress-nginx Reverse proxy, hostname-based routing
Block Storage Longhorn v1.7.0 longhorn-system Default StorageClass
Object Store MinIO (multi-AZ) storage S3-compatible, site-replicated across az-a/az-b
IAM / SSO Authentik iam OIDC provider for Grafana, MinIO, Forgejo, Argo CD
Secret Store HashiCorp Vault iam KV secrets backend, JWT auth via Authentik
Git Forge (planned) Forgejo forge Git server, built-in OCI registry, Actions CI
CI Runner (planned) Forgejo Actions + DinD cicd Privileged build pod; images pushed to Forgejo OCI
CD (planned) Argo CD argocd Pull-based GitOps; never holds kubeconfig in CI
Log Backend Loki (SingleBinary) logging 10-day retention, backed by MinIO
Log Collector Promtail (DaemonSet) logging Scrapes pod logs + Talos journal
Metrics kube-prometheus-stack monitoring Prometheus + node-exporter + kube-state-metrics
Log/Metrics UI Grafana logging Dashboards for Loki + Prometheus
Cluster UI Portainer CE dashboard Container/workload management UI
LLM Inference Ollama llm Local LLM model serving (open-source models)

Repository Structure

homelab/
├── helmfile.yaml             # Single source of truth — deploys everything
├── .env.example              # Required env vars template (copy to .env, gitignored)
│
├── cluster-config/           # Talos + Kubernetes bootstrap
│   ├── cilium-values.yaml
│   ├── longhorn_bootstrap.sh
│   ├── controlplane.yaml     # gitignored — contains secrets
│   ├── worker-1.yaml         # gitignored
│   ├── secrets.yaml          # gitignored
│   ├── kubeconfig            # gitignored
│   └── talosconfig           # gitignored
│
├── k8s/
│   ├── ingress/              # Nginx Ingress Controller + all Ingress rules
│   │   ├── nginx-values.yaml
│   │   └── ingress.yaml
│   ├── storage/              # MinIO multi-AZ object store
│   │   ├── minio-az-a-values.yaml
│   │   ├── minio-az-b-values.yaml
│   │   ├── minio-az-a-pvc.yaml
│   │   ├── minio-service.yaml
│   │   ├── minio-legacy-alias.yaml
│   │   └── minio-replication-job.yaml
│   ├── logging/              # Observability stack
│   │   ├── loki-values.yaml
│   │   ├── promtail-values.yaml
│   │   └── grafana-values.yaml
│   ├── monitoring/           # Prometheus stack + alerting + Grafana dashboards-as-code
│   │   ├── prometheus-values.yaml
│   │   ├── blackbox-exporter-values.yaml   # Active uptime probes (feeds Service Availability dashboard)
│   │   ├── ingress-alerts.yaml             # PrometheusRule: ingress 5xx rate, p95 latency
│   │   └── dashboards/                     # ConfigMaps picked up live by Grafana's sidecar
│   │       ├── service-availability.yaml      # Uptime probes + cert expiry (operator glance)
│   │       ├── service-golden-signals.yaml     # Latency & Golden Signals (ingress RED)
│   │       ├── service-internals.yaml          # Per-service deep-dive (MinIO/Forgejo/Argo CD/Vault/Longhorn/certs)
│   │       ├── kube-controller-health.yaml      # API server RED + kube-state-metrics controller-health proxy
│   │       ├── hardware-overview.yaml           # Per-node CPU/mem/disk/network/load summary
│   │       └── control-plane-logs.yaml          # kube-system + add-on logs (Loki)
│   ├── portainer/            # Portainer CE
│   │   └── portainer-values.yaml
│   ├── talos-iam/            # Authentik + Vault IAM
│   │   ├── authentik-values.yaml
│   │   ├── vault-values.yaml
│   │   ├── setup_vault.sh    # One-time Vault init (not replaced by Helmfile)
│   │   └── provision_oidc.py # Authentik OIDC provisioning
│   ├── coredns/              # CoreDNS hostname rewrites (in-cluster DNS)
│   │   └── coredns-configmap.yaml
│   ├── talos-ci-cd/          # CI/CD stack (planned — not yet applied)
│   │   ├── talos_version_control.html  # Implementation plan + build runbook
│   │   ├── forgejo-values.yaml         # Forgejo Helm values (gitea-charts/gitea)
│   │   ├── argocd-values.yaml          # Argo CD Helm values
│   │   └── charts/forgejo-runner/      # Local Helm chart for the Actions runner
│   │       ├── Chart.yaml
│   │       ├── values.yaml
│   │       └── templates/
│   │           ├── deployment.yaml     # Runner + DinD sidecar, Recreate strategy
│   │           ├── pvc.yaml            # runner-reg (1 Gi) + runner-dind (30 Gi)
│   │           └── networkpolicy.yaml  # Egress: forge ns + DNS + internet only
│   └── duckdns/              # DuckDNS DDNS updater CronJob
│
├── talos-cli/                # Rust CLI for Vault secret access
├── project_context.md        # Authoritative live-state reference
├── refine_cluster.md         # Known-issues runbook
└── LOG.md                    # Append-only change journal

Access

Discover Ingress LoadBalancer IP

# Find the external IP assigned by Cilium LB-IPAM
kubectl get svc -n ingress-nginx ingress-nginx
# Example output:
# LoadBalancer IP: 192.168.1.160 (Cilium LB-IPAM assignment)

Add to /etc/hosts on every client machine (Mac/Linux):

# WireGuard access (remote — via talos-cp-1)
10.6.0.1  grafana.riotpiao.com authentik.riotpiao.com vault.riotpiao.com minio.riotpiao.com prometheus.riotpiao.com portainer.riotpiao.com longhorn.riotpiao.com loki.riotpiao.com forgejo.riotpiao.com temporal.riotpiao.com temporal-grpc.riotpiao.com kmsvc.riotpiao.com

# LAN access (on the home network — use actual LoadBalancer IP from above)
192.168.1.160  grafana.riotpiao.com authentik.riotpiao.com vault.riotpiao.com minio.riotpiao.com prometheus.riotpiao.com portainer.riotpiao.com longhorn.riotpiao.com loki.riotpiao.com forgejo.riotpiao.com temporal.riotpiao.com temporal-grpc.riotpiao.com kmsvc.riotpiao.com

Note: 192.168.1.160 is an example Cilium LB-IPAM assignment. Verify with kubectl get svc -n ingress-nginx ingress-nginx.

There is no real DNS wildcard for *.riotpiao.com — every hostname must be added to /etc/hosts explicitly (as above) before it resolves. Adding a new Ingress host doesn't make it reachable by itself; add the line too.

kubectl Context

Two contexts exist in cluster-config/kubeconfig, pointed at the same cluster over different paths:

Context Server Use when
admin@homelab-cluster 192.168.1.213:6443 (LAN) On the home network
admin@homelab-cluster-1 10.6.0.1:6443 (WireGuard) Remote / off-LAN

If kubectl commands hang or refuse the connection, switch: kubectl config use-context admin@homelab-cluster-1.

Then access services at:

Service URL Credentials
Grafana http://grafana.riotpiao.com admin / GRAFANA_ADMIN_PASSWORD or Authentik SSO
Authentik http://authentik.riotpiao.com akadmin / see .env
Vault http://vault.riotpiao.com root token / see setup_vault.sh output
MinIO console http://minio.riotpiao.com MINIO_ROOT_USER / MINIO_ROOT_PASSWORD
Prometheus http://prometheus.riotpiao.com no auth
Portainer http://portainer.riotpiao.com set on first visit
Longhorn http://longhorn.riotpiao.com no auth
Forgejo (planned) https://forgejo.forge.riotpiao.com rock / FORGEJO_ADMIN_PASSWORD, or Authentik SSO
Argo CD (planned) kubectl port-forward -n argocd svc/argocd-server 8080:443 Authentik SSO (admins only)

Grafana → "Homelab" folder has the operator dashboards (sidecar-loaded from k8s/monitoring/dashboards/, no restart needed on change):

  • Service Availability & Certificate Expiration — uptime probes + cert-manager expiry
  • Latency & Golden Signals — ingress request rate/error %/p50-p99 latency
  • Kube-Controller Health — API server RED metrics + kube-state-metrics controller-health signals
  • Hardware Statistics — per-node CPU/mem/disk/network/load
  • Service Internals — per-service deep-dive (MinIO/Forgejo/Argo CD/Vault/Longhorn)

Deploy

# 1. Install helmfile (once)
brew install helmfile

# 2. Set credentials
cp .env.example .env
# edit .env with your passwords

# 3. Deploy everything
helmfile apply

# Deploy a single stack
helmfile apply -l namespace=logging
helmfile apply -l name=grafana
helmfile apply -l namespace=ingress-nginx

# Preview changes before applying
helmfile diff

Bootstrap Order (fresh cluster)

1.  Provision nodes:        talosctl apply-config (make apply-cp / apply-worker-new)
2.  Bootstrap Kubernetes:   talosctl bootstrap
3.  Install Cilium:         helm install cilium -f cluster-config/cilium-values.yaml
4.  Install Longhorn:       bash cluster-config/longhorn_bootstrap.sh
5.  Deploy everything else: helmfile apply
6.  Vault init (one-time):  bash k8s/talos-iam/setup_vault.sh
7.  Authentik OIDC:         python3 k8s/talos-iam/provision_oidc.py
8.  Label worker:           kubectl label node talos-worker-1 node-role.kubernetes.io/worker=

# ── CI/CD (planned — run after step 8) ─────────────────────────────────────
9.  Private CA + TLS:       see k8s/talos-ci-cd/talos_version_control.html §12 Block 0
10. Deploy Forgejo + runner: helmfile apply -l name=forgejo && helmfile apply -l name=forgejo-runner
11. Authentik SSO for CI:   Forgejo + Argo CD OIDC (§12 Block 1.5)
12. Talos node CA trust:    talosctl patch machineconfig (§12 Block 2)
13. Deploy Argo CD:         helmfile apply -l name=argocd (§12 Block 4)
14. Wire deploy repo:       argocd app create + push first manifests (§12 Block 4)

IAM & Auth Flow

Authentik is the central OIDC identity provider. Vault stores secrets and delegates authentication back to Authentik.

  User / core-cli
       │
       │  OAuth2 / OIDC
       ▼
  Authentik  (authentik.riotpiao.com)
  ├── grafana app      → Grafana OIDC login (group → Admin/Viewer role)
  ├── minio app        → MinIO OIDC login (group → readwrite/readonly policy)
  ├── vault-browser    → Vault UI OIDC login / `vault login -method=oidc`
  └── core-cli-shell  → CLI device code flow (public client, no secret)
       │
       │  JWKS endpoint for JWT validation
       ▼
  HashiCorp Vault  (vault.riotpiao.com)
  ├── auth/jwt   — core-cli authenticates with device code JWT
  ├── auth/oidc  — browser/UI login via Authentik
  └── secret/    — KV v2: mcp/*, cluster/*, cloud/*

core CLI device code login:

core secrets login        # prints URL + code → approve in browser → Vault token cached
core put cluster/DUCKDNS_TOKEN DUCKDNS_TOKEN="abc"            # field name = variable name, never `value`

One-time IAM setup (after helmfile apply):

# 1. Provision OIDC apps and groups in Authentik
GRAFANA_URL=http://grafana.riotpiao.com \
MINIO_URL=http://minio.riotpiao.com \
python3 k8s/talos-iam/provision_oidc.py

# 2. Init Vault, wire JWT + OIDC auth, seed secrets
bash k8s/talos-iam/setup_vault.sh

CoreDNS hostname rewrites (k8s/coredns/coredns-configmap.yaml) ensure in-cluster pods (Grafana, Vault, Forgejo runner, Argo CD) resolve internal hostnames to cluster services, avoiding hairpin NAT through LB IPs.

CI/CD Pipeline (planned)

Implementation plan & exact build commands: k8s/talos-ci-cd/talos_version_control.html

Components

Component Helm chart Namespace Notes
Forgejo gitea-charts/gitea (Forgejo image override) forge Git + OCI registry + Actions engine; SQLite on Longhorn PVC; strategy: Recreate
Forgejo runner local chart charts/forgejo-runner cicd DinD sidecar; PodSecurity privileged; NetworkPolicy fenced
Argo CD argo/argo-cd argocd Pull-based CD; single replica; no external ingress (port-forward only)

Pipeline flow

Developer
  │
  │  git push
  ▼
Forgejo (forge ns)  ─── webhook ───►  Runner (cicd ns)
  │                                        │
  │  SSO login (Authentik OIDC)            │  ACTIONS_RUNTIME_TOKEN → git checkout
  ▼                                        │  ci-registry-token (ci-bot) → docker push → Forgejo OCI
Forgejo UI / Argo CD UI                    │  ci-deploy-token (ci-bot)   → git commit → rock/deploy
                                           │
Forgejo (rock/deploy repo)  ◄─────────────┘
  │
  │  argocd-bot token (repo:read, poll every 3 min)
  ▼
Argo CD (argocd ns)
  │
  │  kubectl apply (cluster-admin ServiceAccount — never in CI)
  ▼
K8s workloads (images from Forgejo OCI)

Key security decisions

  • No kubeconfig in CI. The runner can only push commits + OCI images. Argo CD bridges the gap autonomously.
  • Scoped machine credentials. ci-bot tokens are narrowly scoped: package:write for OCI, repo:write on rock/deploy only. A compromised runner cannot read other repos or call the K8s API.
  • Private CA TLS. Forgejo self-terminates HTTPS with a homelab CA (EC P-256, 10-year). The CA cert is distributed to Talos nodes via machineconfig patch and to the runner via K8s Secret. ca.key never enters the cluster.
  • Authentik SSO for humans. All interactive logins (Forgejo UI, Argo CD UI) route through Authentik. homelab-admins group → Forgejo admin + Argo CD role:admin; homelab-devs group → Forgejo user + no Argo CD access.
  • Forgejo LB IP pinned. Cilium LB-IPAM annotation io.cilium/lb-ipam-ips: <LB_IP> fixes the Forgejo LoadBalancer IP so the TLS SAN and DNS entries never need updating. Configure in k8s/talos-ci-cd/forgejo-values.yaml.

Workflow file location

Forgejo Actions uses GitHub Actions syntax. Workflow files live in .forgejo/workflows/ in each source repo:

rock/source/
└── .forgejo/
    └── workflows/
        ├── ci.yml    # build + test + push OCI image
        └── cd.yml    # on: push to main → bump image tag in rock/deploy

Log Data Flow

Pods / Talos journal (both nodes)
       │
  Promtail (DaemonSet, all nodes)   reads /var/log/pods + /var/log/journal
       │
       ▼
    Loki (logging ns)               indexes + compacts, 10-day retention
       │  stores chunks via S3
       ▼
    MinIO frontend service          minio.storage.svc.cluster.local:9000
    (active-active, round-robin)
       │
  minio-az-a (cp-1) ↔ minio-az-b (worker-1) ↔ minio-az-c (worker-2)
  3-way site replication (bidirectional, automatic)
       │
  Grafana (logging ns)              queries Loki + Prometheus via dashboards
       │
  Nginx Ingress → grafana.riotpiao.com   browser access

Example Applications & Workloads

This cluster runs production-like applications and infrastructure services:

Story Crater Backend

Type: Distributed message-driven application
Namespace: story-crater-backend
Architecture:

  • gRPC server + REST gateway (Envoy)
  • PostgreSQL database (story_crater in CloudNativePG cluster)
  • Kafka topic consumers (via kmsvc message queue)
  • Temporal workflow integration for long-running operations
  • Prometheus metrics (processed messages, latency, errors)
  • Grafana dashboard (message throughput, queue depth, LLM token usage)

Typical flow:

POST /api/messages → gRPC handler
  ↓
Publish to Kafka topic (Story Crater queue)
  ↓
Message consumer processes async (may trigger LLM inference)
  ↓
Results stored in PostgreSQL + emitted as event
  ↓
Prometheus increments counters (story_crater_messages_handled_total)
  ↓
Grafana renders message throughput + duration histograms

Infrastructure Services (Essential)

Service Purpose Namespace Example Use
Authentik OIDC identity provider iam User login, group management, SSO for Grafana/MinIO/Forgejo
Vault Secrets backend iam Database passwords, API keys, JWT token validation
CloudNativePG PostgreSQL 3-replica cluster ddb Auth database (Authentik), app database (Story Crater)
Loki Log aggregation logging Centralize pod logs, Talos kernel logs (10-day retention)
Prometheus Metrics collection monitoring Scrape kube-state-metrics, kubelet, ServiceMonitors (every 30s)
Grafana Observability dashboards logging Query Prometheus + Loki, alert on latency/error spikes
Kafka + kmsvc Message queue sqs Decouple services, async job processing, at-least-once delivery
MinIO S3-compatible object store storage Loki log chunks backend, Vault unseal keys, config backups
Longhorn Persistent block storage longhorn-system All PVCs (Postgres replicas, Kafka broker disks, MinIO)

Development Services (Optional)

Service Purpose Namespace Example Use
Forgejo Self-hosted git + OCI registry cicd Version control, CI Actions runner, private Docker images
Argo CD Pull-based GitOps cicd Continuous deployment (deployment repo → Kubernetes)
Temporal Workflow orchestration temporal Schedule long-running jobs, retry logic, state machines
Portainer Container management UI dashboard Pod inspection, image management, quick debugging

Monitoring Example: Dashboard Walk-Through

Open GrafanaHomelab folder → "Story Crater Backend — Service Overview":

Row A — Availability & Golden Signals

  • Requests/sec (blue = success, red = errors)
  • Error rate % (goal: < 0.1%)
  • p50/p95/p99 latency (goal: p99 < 500ms)
  • Alert: If error rate > 1% for 5 min, page on-call

Row B — Resource Usage

  • CPU (request/limit)
  • Memory (request/limit)
  • Restart count (goal: 0; alerts if > 2)

Row C — Domain-Specific Metrics (Story Crater only)

  • Messages processed/sec (bucketed by status: success, dlq, retry)
  • Queue depth (Kafka partitions lag)
  • Dedup window retention (FIFO redelivery tracking)
  • LLM inference tokens used/sec
  • External API call latency (e.g., OpenAI)

Row D — Logs

  • Live Loki panel: filter by pod + search for errors
  • Example: {namespace="story-crater-backend"} | json | level="error"

Row E — Alert Status (if SLO defined)

  • Burn rate (if consuming SLO budget)
  • Example: "30-day availability SLO = 99.5%; current burn rate = 0.2x"

Row F — Related Dashboards

  • Link to Kafka dashboard (queue depth)
  • Link to PostgreSQL dashboard (story_crater DB)
  • Link to Temporal dashboard (workflow execution times)

Adding Hardware to the Cluster

Step 1 — Get the Talos image

The image must include the same extensions as the existing nodes (iscsi-tools + util-linux-tools). Download from the Image Factory using the cluster's schematic ID:

Format Use case URL
ISO USB boot (recommended for bare metal) https://factory.talos.dev/image/613e1592.../v1.13.3/metal-amd64.iso
RAW disk image Write directly to drive via another machine https://factory.talos.dev/image/613e1592.../v1.13.3/metal-amd64.raw.xz
PXE / iPXE Network boot — no USB needed https://factory.talos.dev/image/613e1592.../v1.13.3/kernel-amd64

Full schematic ID: 613e1592b2da41ae5e265e8789429f22e121aab91cb4deb6bc3c0b6262961245


curl -Lo talos-worker.iso \
  "https://factory.talos.dev/image/613e1592b2da41ae5e265e8789429f22e121aab91cb4deb6bc3c0b6262961245/v1.13.3/metal-amd64.iso"
sudo dd if=talos-worker.iso of=/dev/sdX bs=4M status=progress && sync

Option B — In-Memory (diskless / RAM boot via PXE)

Talos runs entirely from RAM. Useful for temporary nodes or hardware where you don't want to touch the existing OS.

# Boot via PXE pointing to:
# Kernel:  https://factory.talos.dev/image/613e1592.../v1.13.3/kernel-amd64
# Initrd:  https://factory.talos.dev/image/613e1592.../v1.13.3/initramfs-amd64.xz
# Cmdline: talos.platform=metal

Note: in-memory nodes lose state on reboot. Not suitable for Longhorn storage nodes.


Option C — Direct Disk Image (headless / remote)

xz -d talos-worker.raw.xz
sudo dd if=talos-worker.raw of=/dev/sda bs=4M status=progress && sync

Step 2 — Discover hardware in maintenance mode

# Scan your LAN for the new node in maintenance mode
nmap -sn 192.168.1.0/24
# Note: Replace with your actual subnet (e.g., 10.0.1.0/24)

# Discover available disks on the node
talosctl --nodes <maintenance-ip> --talosconfig cluster-config/talosconfig disks --insecure

Step 3 — Prepare the worker config

cp cluster-config/worker-1.yaml cluster-config/worker-N.yaml

Edit exactly these four fields:

Field Value
machine.network.hostname talos-worker-N
machine.network.interfaces[0].addresses 192.168.1.16N/24
machine.install.disk disk path from Step 2
machine.nodeLabels.topology.kubernetes.io/zone az-N

Step 4 — Apply config

make apply-worker-new N=<num> WN_IP=<maintenance-ip>

Step 5 — Persist the worker IP

# Substitute <WN_MAINTENANCE_IP> with the actual IP discovered in Step 2
echo 'export W<N>_IP=192.168.1.<last-octet>' >> ~/.zshrc && source ~/.zshrc

Step 6 — Set the node-role label

kubectl label node talos-worker-N node-role.kubernetes.io/worker=

Step 7 — Verify

kubectl get nodes -w
kubectl describe node talos-worker-N | grep -A10 Labels

Step 8 — What automatically extends to the new node

Service Behaviour
Cilium New pod scheduled automatically
Promtail DaemonSet — starts immediately
Nginx Ingress DaemonSet — starts immediately, port 80/443 available on new node
Longhorn Detects new node, available for replica scheduling
MinIO / Loki / Grafana Stay on existing node (single-replica Deployments)
S
Description
riotpiao.com homelab GitOps repo
Readme
60 MiB
Languages
Python 48.4%
HCL 19.8%
Shell 13.7%
TypeScript 9.1%
Makefile 7.1%
Other 1.9%