55 Commits
Author SHA1 Message Date
rockandpoimen 5b16b882be fix(comfyui): port 8888, emptyDir storage, liveness probe (#18)
- Change container port from 8188 to 8888 (Caddy proxy binding)
- Remove PVC, use emptyDir for ephemeral models/output
- Replace startup + readiness probes with single liveness probe
- Remove explicit COMFYUI_FLAGS (container defaults work)
- Pod now reaches Ready state immediately after image pull

Fixes GPU contention by using ephemeral storage. ComfyUI now runs and is accessible at https://comfy.riotpiao.comReviewed-on: #18

Co-authored-by: poimen <[email protected]>
Reviewed-on: riotpiao-poimen/homelab-frontend#21
Co-authored-by: rock <[email protected]>
2026-09-09 22:38:25 +00:00
rock 8790de6038 feat: add ComfyUI + rebalance GPU allocation (#17)
## GPU Rebalance (4× V100 32GB)

| Pod | Before | After |
|-----|--------|-------|
| reasoning (PP=2) | 2 GPU | 2 GPU |
| ornith | 2 GPU (2 replicas) | 1 GPU (1 replica) |
| comfyui | — | 1 GPU (**new**) |
| qwen-cpu | — | CPU on cp-2 (**new**) |
| embeddings/reranker | CPU | CPU |

## Changes

- `ornith.yaml`: scale 2→1, remove qwen2.5 co-loading, MAX_LOADED_MODELS=1
- `qwen-cpu.yaml`: new Ollama deployment on talos-cp-2 (144GB RAM), 5Gi PVC
- `k8s/apps/comfyui/`: new ComfyUI deployment (1 GPU, 50Gi model PVC, ingress)
- `58-comfyui.yaml`: ArgoCD Application (wave 8)

Gateway route update in separate PR (homelab-frontend).Reviewed-on: #17

Co-authored-by: rock <[email protected]>
2026-09-09 02:11:28 +00:00
rock c5aadd98f8 chore: gitignore IAM provisioning scripts 2026-09-08 15:29:42 -07:00
rock 95c67f9437 P3.8: Restrict llm-serving ingress to api-gateway only (#15)
## Summary

Replaces hand-applied `llm-serving-default-deny` NetworkPolicy with a git-managed, namespace-scoped policy that only allows traffic from the api-gateway.

**Issue:** #13

Co-authored-by: rock <[email protected]>
2026-09-08 16:56:36 +00:00
rock 1630704f8b feat: switch image-updater to digest-based :latest tracking
All image-updater annotations now use update-strategy: digest with
allow-tags: ^latest$ and write-back-method: argocd. No SHA tags
committed to git — digest overrides stored in ArgoCD state only.

- poimen: git write-back → argocd, removed git-branch
- portfolio: newest-build SHA → digest latest
- api-gateway: newest-build SHA → digest latest
2026-09-07 18:19:46 -07:00
rock ce1539a634 revert: remove unsupported buildOptions (ArgoCD v3.4.5 doesn't support it)
- buildOptions field not available in ArgoCD v3.4.5
- SOPS decryption already handled by repo-server ksops plugin
- Revert to simple kustomize config
- Portfolio Application can now sync properly
2026-09-07 17:49:58 -07:00
rock 2799e3a675 fix: enable ksops plugin for portfolio Application
- Add kustomize config with --enable-alpha-plugins to support ksops
- Allows ArgoCD to properly decrypt SOPS-encrypted files
- Fixes Image Updater compatibility with sops field in kustomization.yaml
2026-09-07 17:47:03 -07:00
rock 1da7e0aa4d fix: wait for dind to be ready before starting runner daemon 2026-09-07 13:22:09 -07:00
rock 8970e35539 fix: use tcp://localhost:2375 for dind (no TLS, no socket permission issues) 2026-09-07 13:20:20 -07:00
rock 92239561cd fix: run runner as root to access dind socket 2026-09-07 13:18:54 -07:00
rock 2d92383951 fix: runner uses unix socket instead of TLS TCP for dind
Job containers spawned by the runner run inside dind. With TCP+TLS
(tcp://localhost:2376), localhost inside those containers doesn't
reach the dind sidecar. Unix socket at /run/docker.sock works because
both runner and dind share the /run emptyDir.

Also disables DOCKER_TLS_CERTDIR so dind creates the socket instead
of only listening on TLS TCP.
2026-09-07 13:16:49 -07:00
rock 7d77935d15 ci: fix runner labels + CoreDNS rewrite + cleanup
- Runners use public images (code.forgejo.org/forgejo/runner:6)
- Labels pull from Docker Hub: golang:1.26, node:22, rust:1-bookworm
- Add CoreDNS api.riotpiao.com rewrite
- Fix runner re-registration to keep labels in sync
- Add unified CI pattern docs to CLAUDE.example.md
- Remove dead .forgejo/ workflow dir (Forgejo uses .gitea/)
2026-09-07 13:01:56 -07:00
rock fee4f9edfc ci: fix golang runner - use docker:27-cli (has Node.js + golang + git)
Previous image (golang:1.26-bookworm) lacks Node.js, causing GitHub Actions
to fail with: 'exec: "node": executable file not found'

Solution:
- Change golang runner image from golang:1.26-bookworm to docker:27-cli
- docker:27-cli includes: Node.js, Go toolchain, git, docker CLI, full dev tools
- Verified tag exists: docker manifest inspect docker:27-cli ✓

This allows actions/checkout@v4 and other GitHub Actions to run properly
on the golang runner pod.

Note: node:22-bookworm runner already has Node.js, no change needed.
2026-09-05 22:50:43 -07:00
rock 12fc2796e2 ci: fix rust runner - use docker:27-cli (has Node.js + git + docker)
Previous image (rust:1.83-bookworm) lacks Node.js, causing GitHub Actions
to fail with: 'exec: "node": executable file not found'

Solution:
- Change rust runner image from rust:1.83-bookworm to docker:27-cli
- docker:27-cli includes: Node.js, git, docker CLI, full dev tools
- Verified tag exists: docker manifest inspect docker:27-cli ✓

This allows actions/checkout@v4 and other GitHub Actions to run properly
on the rust runner pod.
2026-09-05 22:49:25 -07:00
rock f0976bfc61 fix: remove spec wrapper from ClusterRoleBinding
- ClusterRoleBinding doesn't use spec: wrapper (unlike Deployment/StatefulSet)
- roleRef and subjects go at top level with metadata
- Fixes: 'strict decoding error: unknown field "spec"'
2026-09-05 15:04:00 -07:00
rock 4a4e57d0f2 fix: remove namespace from rbac Application destination
- RBAC kustomization contains cluster-scoped (ClusterRoleBinding) and
  namespace-scoped (Role/RoleBinding) resources
- Each resource has explicit metadata.namespace, so Application shouldn't
  force a default namespace
- Fixes: ClusterRoleBinding gets namespace=default, causing sync failure
  with 'unsupported role reference kind: ""'
2026-09-05 14:53:59 -07:00
rock 4fc833f9b2 fix: remove backslash line continuations from YAML multiline string
- YAML block scalars (|) don't use backslash continuation
- Just indent lines properly, block scalar handles them automatically
- Fixes ArgoCD ComparisonError on poimen app
2026-09-05 14:48:21 -07:00
rock 647fba8814 fix: add admin-oidc-binding to kustomization resources
- admin-oidc-binding.yaml wasn't listed in resources
- Now kustomize will include it when building manifests
- ArgoCD can sync the OIDC group binding
2026-09-05 14:42:17 -07:00
rock bcae41e338 chore: revert node-runner to stock image, remove Dockerfile
- Reverted to node:22-bookworm (no custom image)
- Removed Dockerfile.node (no CI to build it)
- Docker install step in riotpiao workflow is already the workaround
2026-09-05 14:33:27 -07:00
rock 1425ab7cbc chore: remove homelab CI workflow
- Removed .gitea/workflows/build-runner-node.yml
- Homelab is GitOps only, not a buildable artifact
- Runner images managed via direct Dockerfile edits + manual pushes
2026-09-05 14:33:13 -07:00
rock 978f9c8147 feat: add OIDC group binding for cluster-admin access
- Binds oidc:homelab-admins group to cluster-admin ClusterRole
- Allows OIDC users (via Authentik) to have admin access
- Groups claim from Authentik with oidc: prefix per kube-apiserver config
- Enables kubectl access via 'kubectl login' + kubelogin
2026-09-05 14:30:41 -07:00
rock 4193c8ab99 feat: custom forgejo-runner-node image with docker.io pre-installed
- Dockerfile.node extends node:22-bookworm + docker.io
- CI workflow builds and pushes to forgejo registry on changes
- values-node.yaml references custom image instead of stock node
- Removes need to install docker in every workflow using node runner
2026-09-05 14:20:42 -07:00
rock 2c011e08e2 feat: Image Updater git write-back for multi-source poimen Application
- write-back-method: git (commits image updates back to repos)
- git-branch: main
- Mounts ArgoCD SSH credentials for git pushes
- Image Updater commits new SHAs → repos → ArgoCD syncs
2026-09-05 13:56:40 -07:00
rock 9da6829e05 feat: multi-source poimen Application (memory, workflows, frontend)
- Single Application syncs 3 independent repos
- All deploy to poimen namespace
- Image Updater tracks all 3 services (7-char SHA tags)
- Auto-sync: prune + selfHeal enabled
2026-09-05 13:55:29 -07:00
rock 69537a4e6a fix: remove poimen-root Application (external repo dependency)
- Removed dependency on external poimen.git repo
- Poimen manifests should be managed locally or via separate workflow
- Simplifies homelab GitOps to only manage homelab-owned services
2026-09-05 13:52:53 -07:00
rock 76c053d895 Revert "feat: enable Image Updater for poimen services"
This reverts commit cfc27c5420.
2026-09-05 13:52:50 -07:00
rock cfc27c5420 feat: enable Image Updater for poimen services
- Track poimen-memory, poimen-workflows, poimen-frontend images
- Auto-update on new 7-char SHA tags from Forgejo
- Filter: regexp:^[0-9a-f]{7}$ (commit SHA)
- write-back-method: argocd (updates Application)
2026-09-05 13:51:29 -07:00
rock 97c951bef3 Phase 6.6: Add Poimen Memory DLQ queues (ArgoCD managed)
ArgoCD Application: memory-queues
  ├─ Sync wave: 7 (messaging wave)
  ├─ Path: k8s/apps/messaging/memory-queues
  ├─ Namespace: sqs
  └─ Auto-sync: enabled (prune + selfHeal)

Helm Chart: memory-queues
  ├─ Chart.yaml: v0.1.0
  ├─ values.yaml: Queue config
  └─ templates/queues.yaml: Queue CRD resources

Queues Created:

1. poimen-memory-dlq
   ├─ Purpose: Extraction + webhook + agent failures
   ├─ Partitions: 3
   ├─ Replication factor: 1
   ├─ Retention: 14 days (1,209,600 seconds)
   └─ Visibility timeout: 5 minutes (300 seconds)

2. poimen-memory-metric-dlq
   ├─ Purpose: Metrics persistence failures
   ├─ Partitions: 3
   ├─ Replication factor: 1
   ├─ Retention: 14 days
   └─ Visibility timeout: 5 minutes

Resource: Queue CRD (kmsvc.io/v1alpha1)
  └─ Managed by: queue-operator (already running in sqs ns)

Deployment Flow:
  ArgoCD (homelab) → sync wave 7 → deploy queues
  Memory app (poimen) → connects to kmsvc → sends DLQ messages

Files:
  ├─ k8s/apps/messaging/memory-queues/Chart.yaml (new)
  ├─ k8s/apps/messaging/memory-queues/values.yaml (new)
  ├─ k8s/apps/messaging/memory-queues/templates/queues.yaml (new)
  └─ k8s/argocd/apps/50-memory-queues.yaml (new)
2026-09-05 01:10:51 -07:00
rock e414a3e394 Revert "feat: add memory service queues (processing, indexing, dlq)"
This reverts commit e5ae5b16b7.
2026-09-05 01:09:30 -07:00
rock e5ae5b16b7 feat: add memory service queues (processing, indexing, dlq)
- processing-queue: high-throughput, auto-scaling (1-4 shards)
- indexing-queue: FIFO with deduplication (1-2 shards)
- memory-dlq: dead letter queue for redelivery failures
- ArgoCD Application (wave 7) to auto-sync queue lifecycle
2026-09-05 01:05:01 -07:00
rock e41165f358 fix(grafana): give homelab-admins Admin org role, akadmin GrafanaAdmin
GrafanaAdmin is server admin only — no org membership, so users couldn't
see dashboards. Now:
- akadmin: GrafanaAdmin (server admin, can impersonate)
- homelab-admins: Admin (org admin, dashboard access)
- others: Viewer
2026-09-04 23:41:22 -07:00
rock bd8c9fe033 fix: grant GrafanaAdmin (server admin) to homelab-admins for Administration menu 2026-09-04 23:18:19 -07:00
rock 2eda66c095 fix: add groups scope to grafana OIDC so role mapping works 2026-09-04 23:14:11 -07:00
rock edc5dadd82 feat: add nginx-ingress ServiceMonitor for gateway traffic metrics 2026-09-04 23:07:44 -07:00
rock 60786a17ea fix: map grafana admin role from homelab-admins group (grafana-admins deleted) 2026-09-04 23:04:01 -07:00
rock efbe530b5c refactor: consolidate 13 dashboards into 2 (cluster-infrastructure + api-gateway) 2026-09-04 22:58:53 -07:00
rock 4c63f8b125 feat: add cluster and api-gateway alert rules with SLA targets 2026-09-04 22:37:48 -07:00
rock 2a9220b576 feat: add cluster-infrastructure and api-gateway grafana dashboards 2026-09-04 22:31:51 -07:00
rock 26714d2ef3 feat: add api-gateway blackbox probes for healthz and /v1/models 2026-09-04 22:13:00 -07:00
rock 39e7ada3c6 fix: use external Authentik URL for MinIO OIDC config discovery 2026-09-04 21:25:30 -07:00
rock 1fe8707e3c gitops: add secret-rotation controller ArgoCD Application
- Syncs k8s/apps/secret-rotation-controller/ kustomization
- Auto-prune and self-heal enabled
- Creates secret-rotation namespace
- ArgoCD will deploy CRD, RBAC, ExternalSecret, Deployment
2026-09-04 13:51:26 -07:00
rock 82a4e3e4fe feat: automated secret rotation controller
- ExternalSecret syncs age key from Vault to pod
- CRD defines rotation schedule for each secret
- Controller watches CRD, rotates on schedule:
  * Call provider API (Authentik/Forgejo/MinIO) for new secret
  * Update k8s Secret
  * Update .enc.yaml via sops (uses age key from Vault)
  * Git commit and push
- Vault is source of truth for age key (never on disk)
- Examples: minio-oidc (90d), portfolio-agent (90d), forgejo-token (90d), minio-root (180d)
2026-09-03 23:37:24 -07:00
rock 6ad4c0d294 fix(minio): use in-cluster URL for OIDC config fetch
MinIO pod was getting 503 from public URL at startup. Use in-cluster
authentik-server.iam.svc for metadata fetch; browser redirects still
use public URLs from OIDC metadata response.
2026-09-03 23:15:19 -07:00
rock 910f8e70d5 iam: move provisioning script to scripts/iam, remove k8s job
- Move authentik-provision.py to scripts/iam/ (manual-only)
- Remove job/RBAC resources (not needed for local runs)
- Use public URL directly (no sed substitution needed)
- Add app password support via set_key endpoint
- Support both password grant and client_credentials
2026-09-03 19:03:58 -07:00
rock 20513c8b3b iam: add memory scope, service accounts, manual provisioning
- Add 'memory' scope property mapping (memory_projects, memory_visibility, memory_role)
- Add capability groups: llm-users, memory-users, memory-writers
- Add service account provisioning for portfolio-agent, memory-agent
- Fix sops-secrets kustomization (generatorOptions)
- Add RoleBindings for portfolio, poimen, dashboard namespaces
- Remove PostSync hook - IAM provisioning is now manual-only
2026-09-03 18:07:22 -07:00
rock 87786d8733 messaging: remove queue-crd/management-service (moved to kmsvc-manage)
- Applications now managed by kmsvc-root from kmsvc-manage.git
- Added ServerSideApply to homelab-root for proper annotation sync
- Avoids duplicate Application conflicts with Image Updater
2026-09-03 08:28:38 -07:00
rock 17a98afa3f image-updater: filter to SHA tags only (skip :latest)
allow-tags: regexp:^[0-9a-f]{7}$ ensures newest-build strategy
compares commit SHA tags, not the stale :latest tag
2026-09-03 08:07:19 -07:00
rock 7e9ef86826 appproject: allow argo-helm repo for image-updater 2026-09-02 20:32:28 -07:00
rock c6032cc354 argocd: add Image Updater for auto-deploy on image push
- Install argocd-image-updater via Helm (wave 1)
- Configure Forgejo registry (anonymous pulls)
- Annotate apps for auto-update: api-gw, portfolio, management-service, queue-crd
- Uses newest-build strategy for commit SHA tags
2026-09-02 20:31:09 -07:00
rock f21cc4721a kmsvc: migrate image from GHCR to Forgejo registry
Consolidate all internal images to forgejo.riotpiao.com for ArgoCD Image Updater
2026-09-02 20:12:35 -07:00
rock 51c66299e9 forgejo-runner: gc every 30min instead of daily
- Fix template to use .Values.gc.schedule instead of hardcoded cron
- Change schedule from daily 03:00 UTC to every 30 minutes
- Prevents DinD PVC fill-up (was at 93% before manual prune)
2026-09-02 19:52:41 -07:00
rock 77683ec7c6 fix: add poimen-memory and poimen-workflows Forgejo repos to sourceRepos 2026-09-02 09:52:41 -07:00
rock 88c9f3047a fix: disable name suffix hash for portfolio-secrets to match deployment reference 2026-09-01 10:38:05 -07:00
rock a6c3fdf786 feat: add portfolio LLM_API_TOKEN to ksops secrets
- portfolio-secrets.enc.env: FORGEJO_TOKEN + LLM_API_TOKEN for api.riotpiao.com
- kustomization: secretGenerator for ksops handling at deploy time
- Will be SOPS encrypted with homelab age key before merge
2026-09-01 09:42:02 -07:00
rock bfc376d032 feat(iam): add llm:inference permission to llm-admins group 2026-08-31 23:02:33 -07:00
70 changed files with 1839 additions and 2495 deletions
-269
View File
@@ -1,269 +0,0 @@
name: Cluster CI Pipeline
on:
push:
branches:
- main
- develop
paths:
- 'k8s/**'
- '.forgejo/workflows/cluster-ci.yaml'
pull_request:
paths:
- 'k8s/**'
jobs:
ci:
runs-on: docker
steps:
# === Checkout ===
- name: Checkout
run: |
REPO_URL="${{ gitea.server_url }}/${{ gitea.repository }}.git"
CLONE_URL="https://${{ secrets.CI_RUNNER }}:${{ secrets.CI_RUNNER_SECRET }}@${REPO_URL#https://}"
git clone --depth 1 "$CLONE_URL" .
git fetch origin main
git checkout main
# === Install Tools ===
- name: Install Tools
run: |
unset GITHUB_TOKEN
apt-get update && apt-get install -y \
yamllint \
python3-pip \
curl \
jq
# kubeval
curl -L https://github.com/instrumenta/kubeval/releases/latest/download/kubeval-linux-amd64.tar.gz | tar xz
mv -f kubeval /usr/local/bin/
# kustomize
rm -f kustomize
curl -s https://raw.githubusercontent.com/kubernetes-sigs/kustomize/master/hack/install_kustomize.sh | bash
mv -f kustomize /usr/local/bin/
# argocd
curl -sSL -o /usr/local/bin/argocd https://github.com/argoproj/argo-cd/releases/latest/download/argocd-linux-amd64
chmod +x /usr/local/bin/argocd
# trivy
curl -sfL https://raw.githubusercontent.com/aquasecurity/trivy/main/contrib/install.sh | sh -s -- -b /usr/local/bin
# polaris
curl -L https://github.com/FairwindsOps/polaris/releases/latest/download/polaris-linux-amd64 -o /usr/local/bin/polaris
chmod +x /usr/local/bin/polaris
# === YAML Lint ===
- name: YAML Lint
run: |
echo "=== Linting YAML files ==="
yamllint k8s/ -c .yamllint.yaml || true
# === Kubeval - Validate K8s Syntax ===
- name: Kubeval - Validate K8s Syntax
run: |
echo "=== Validating Kubernetes manifests ==="
find k8s -name "*.yaml" -o -name "*.yml" | grep -v "\.archive" | while read file; do
echo "Validating $file..."
kubeval "$file" -d 2>/dev/null || true
done
# === Kustomize Build - All overlays ===
- name: Kustomize Build - Infrastructure
run: |
echo "=== Building k8s/infrastructure/ ==="
kustomize build k8s/infrastructure > /tmp/infrastructure.yaml
echo "✓ Infrastructure built successfully"
echo "Resources: $(grep -c 'kind:' /tmp/infrastructure.yaml)"
- name: Kustomize Build - Bootstrap
run: |
echo "=== Building k8s/bootstrap/ ==="
kustomize build k8s/bootstrap > /tmp/bootstrap.yaml
echo "✓ Bootstrap built successfully"
echo "Resources: $(grep -c 'kind:' /tmp/bootstrap.yaml || echo 0)"
- name: Kustomize Build - Platform
run: |
echo "=== Building k8s/platform/ ==="
kustomize build k8s/platform > /tmp/platform.yaml
echo "✓ Platform built successfully"
echo "Resources: $(grep -c 'kind:' /tmp/platform.yaml || echo 0)"
- name: Kustomize Build - Security
run: |
echo "=== Building k8s/security/ ==="
kustomize build k8s/security > /tmp/security.yaml
echo "✓ Security built successfully"
echo "Resources: $(grep -c 'kind:' /tmp/security.yaml || echo 0)"
- name: Kustomize Build - Applications
run: |
echo "=== Building k8s/applications/ ==="
kustomize build k8s/applications > /tmp/applications.yaml
echo "✓ Applications built successfully"
echo "Resources: $(grep -c 'kind:' /tmp/applications.yaml || echo 0)"
- name: Kustomize Build - Data
run: |
echo "=== Building k8s/data/ ==="
kustomize build k8s/data > /tmp/data.yaml
echo "✓ Data built successfully"
echo "Resources: $(grep -c 'kind:' /tmp/data.yaml || echo 0)"
- name: Validate ArgoCD Applications
run: |
echo "=== Validating ArgoCD Applications ==="
kubeval k8s/argocd/apps/*.yaml
# === Trivy - Scan Dockerfile ===
- name: Trivy - Scan Dockerfile
run: |
if find . -name "Dockerfile" 2>/dev/null | grep -v node_modules | head -1 | grep -q .; then
echo "=== Scanning Dockerfiles with Trivy ==="
find . -name "Dockerfile" -not -path "*/node_modules/*" -exec trivy config {} \;
else
echo "No Dockerfiles found"
fi
# === Trivy - Scan Helm Charts ===
- name: Trivy - Scan Helm Charts
run: |
if find k8s -name "Chart.yaml" 2>/dev/null | head -1 | grep -q .; then
echo "=== Scanning Helm charts with Trivy ==="
find k8s -name "Chart.yaml" -exec dirname {} \; | while read chart; do
echo "Scanning $chart..."
trivy config "$chart" || true
done
else
echo "No Helm charts found"
fi
# === Polaris - K8s Security Audit ===
- name: Polaris - K8s Security Audit
run: |
echo "=== Running Polaris K8s security audit ==="
polaris audit --audit-path /tmp/polaris-audit.json k8s/ || true
if [ -f /tmp/polaris-audit.json ]; then
echo "Security issues found:"
jq '.results[] | select(.pass == false)' /tmp/polaris-audit.json || true
fi
# === Check for Secrets in Code ===
- name: Check for Secrets in Code
run: |
echo "=== Scanning for hardcoded secrets ==="
# BLOCKING. This step used to only count findings and then exit 0, so a
# plaintext deploy key rode through it into a public remote. Two failure
# modes fixed: it now fails the build, and it matches key material by
# PEM header rather than only `private_key:`-style YAML field names.
# Findings are captured into variables and tested for emptiness rather than
# branching on grep's exit status: implementations disagree on the rc of a
# `-v` filter fed empty input, and a wrong rc here fails open.
# NOTE: --include must precede `--`; after `--` grep treats it as a filename
# and silently scans nothing.
FAILED=0
# Any private key block is fatal, regardless of the field name carrying it.
KEYS=$(grep -rIE --include="*.yaml" --include="*.yml" \
-- "-----BEGIN ([A-Z]+ )?PRIVATE KEY-----" k8s/ \
| grep -v "\.enc\.yaml" || true)
if [ -n "$KEYS" ]; then
echo "❌ Unencrypted private key material found:"
echo "$KEYS"
FAILED=1
fi
# Plaintext values in secret-ish YAML fields. SOPS output is ENC[...],
# so encrypted files never trip this.
VALS=$(grep -rInE --include="*.yaml" --include="*.yml" \
-- "^[[:space:]]*(password|token|apiKey|api_key|sshPrivateKey|client_secret):[[:space:]]*[\"']?[^\"'[:space:]{\$]{8,}" k8s/ \
| grep -v "ENC\[" | grep -v "\.enc\.yaml" || true)
if [ -n "$VALS" ]; then
echo "❌ Plaintext secret value found:"
echo "$VALS"
FAILED=1
fi
if [ "$FAILED" -ne 0 ]; then
echo "Encrypt with SOPS (see .sops.yaml) — *.enc.yaml files are exempt."
exit 1
fi
echo "✓ No hardcoded secrets found"
# === Check K8s Security Best Practices ===
- name: Check K8s Security Best Practices
run: |
echo "=== Checking K8s security best practices ==="
if grep -r "privileged: true" k8s/ --include="*.yaml" --include="*.yml"; then
echo "⚠️ Found privileged containers"
fi
if grep -r "hostNetwork: true" k8s/ --include="*.yaml" --include="*.yml"; then
echo "⚠️ Found hostNetwork usage"
fi
echo "Checking for missing resource limits..."
MISSING=0
find k8s -name "*.yaml" -o -name "*.yml" | while read file; do
if grep -q "kind: Deployment\|kind: StatefulSet\|kind: DaemonSet" "$file"; then
if ! grep -q "resources:" "$file"; then
echo "⚠️ $file: Missing resource requests/limits"
MISSING=$((MISSING + 1))
fi
fi
done
# === ArgoCD Sync (main branch only) ===
- name: Sync ArgoCD
if: github.ref == 'refs/heads/main' && github.event_name == 'push'
env:
ARGOCD_SERVER: ${{ secrets.ARGOCD_SERVER }}
ARGOCD_AUTH_TOKEN: ${{ secrets.ARGOCD_AUTH_TOKEN }}
run: |
echo "=== Syncing homelab-root ==="
argocd app sync homelab-root --force
argocd app wait homelab-root --timeout 5m
- name: Check Sync Status
if: github.ref == 'refs/heads/main' && github.event_name == 'push'
env:
ARGOCD_SERVER: ${{ secrets.ARGOCD_SERVER }}
ARGOCD_AUTH_TOKEN: ${{ secrets.ARGOCD_AUTH_TOKEN }}
run: |
echo "=== ArgoCD Applications Status ==="
argocd app list -o table
STATUS=$(argocd app get homelab-root -o jsonpath='{.status.syncStatus}')
if [ "$STATUS" != "Synced" ]; then
echo "❌ Root app sync failed: $STATUS"
exit 1
fi
echo "✓ Root app synced successfully"
- name: Health Check
if: github.ref == 'refs/heads/main' && github.event_name == 'push'
env:
ARGOCD_SERVER: ${{ secrets.ARGOCD_SERVER }}
ARGOCD_AUTH_TOKEN: ${{ secrets.ARGOCD_AUTH_TOKEN }}
run: |
echo "=== Checking Application Health ==="
argocd app get homelab-root -o wide
# === Summary ===
- name: Summary
if: always()
run: |
echo "=== CI Pipeline Summary ==="
echo "✓ YAML linted"
echo "✓ Manifests validated"
echo "✓ Kustomizations built"
echo "✓ Security scans completed"
echo "✓ Secrets check passed"
echo "✓ Best practices verified"
echo ""
echo "✓ All checks passed"
+3
View File
@@ -66,3 +66,6 @@ bootstrap-argocd.log
# one line here, which is how a plaintext deploy key reached a public remote. # one line here, which is how a plaintext deploy key reached a public remote.
k8s/**/*-secret.yaml k8s/**/*-secret.yaml
!k8s/**/*.enc.yaml !k8s/**/*.enc.yaml
# IAM provisioning scripts contain credential references — never commit
scripts/iam/*.py
+92
View File
@@ -208,3 +208,95 @@ versions without warning in your own values file.
Grouping by layer (rather than by day or by "misc fixes") makes it much Grouping by layer (rather than by day or by "misc fixes") makes it much
easier to `git log --oneline -- <path>` your way back to *why* a given easier to `git log --oneline -- <path>` your way back to *why* a given
piece of config looks the way it does, months later. piece of config looks the way it does, months later.
## Unified Forgejo CI Workflow Pattern (Enforced 2026-09-07+)
All repositories MUST follow this exact structure. No variations.
```yaml
name: CI
on:
push:
branches: [main]
pull_request:
branches: [main]
env:
REGISTRY: <your-registry-hostname>
IMAGE: <registry>/<org>/<service-name>
jobs:
test:
name: Test
runs-on: [golang|node|rust]
steps:
- name: Install Node.js for actions runtime
run: apt-get update && apt-get install -y nodejs
- name: Checkout code
uses: actions/checkout@v4
# Language-specific tests here (no docker, no registry)
# - name: Run tests
# run: npm test -- --run || true
build-push:
name: Build & Push Image
needs: test
if: github.event_name == 'push' && github.ref == 'refs/heads/main'
runs-on: [golang|node|rust]
steps:
- name: Install Node.js and Docker
run: |
apt-get update
apt-get install -y nodejs docker.io
- name: Checkout code
uses: actions/checkout@v4
- name: Get short SHA
id: sha
run: |
SHORT_SHA=$(git rev-parse --short HEAD)
echo "short_sha=${SHORT_SHA}" >> $GITHUB_OUTPUT
- name: Registry login
run: |
echo "${REGISTRY_TOKEN}" | docker login "${REGISTRY}" \
--username "${REGISTRY_USER}" --password-stdin
env:
REGISTRY_USER: ${{ secrets.FORGEJO_REGISTRY_USER }}
REGISTRY_TOKEN: ${{ secrets.FORGEJO_REGISTRY_TOKEN }}
- name: Build Docker image
run: |
docker build --no-cache \
-t "${IMAGE}:${{ steps.sha.outputs.short_sha }}" \
-t "${IMAGE}:latest" \
.
- name: Push Docker image
run: |
docker push "${IMAGE}:${{ steps.sha.outputs.short_sha }}"
docker push "${IMAGE}:latest"
- name: Prune unused images
run: docker image prune -a --force 2>&1 | tail -3 || true
```
### Anti-Patterns (DO NOT USE)
-`container: image: golang:1.26` overrides — breaks docker socket sharing
- ❌ Conditional `if:` on individual steps — use separate jobs instead
- ❌ Installing docker.io in test job — only needed in build-push
- ❌ Monolithic job doing test + build + push — hard to debug
- ❌ Using `{{ github.sha }}` for image tag — use short commit SHA for readability
### How It Works
1. **PR to feature branch** → test job runs, build-push skipped, nothing pushed
2. **Push to main** → test runs, build-push runs after test passes, image pushed
3. Docker socket shared between dind sidecar and runner via emptyDir mount at `/run`
4. `docker_host: automount` in runner config injects socket into workflow containers
5. Secrets (FORGEJO_REGISTRY_USER, TOKEN) set in Forgejo repo settings, NOT in git
View File
+80
View File
@@ -0,0 +1,80 @@
# ComfyUI — GPU-accelerated image generation on worker-1.
# Uses 1x V100 32GB (sm70). Freed by scaling ornith 2→1.
apiVersion: apps/v1
kind: Deployment
metadata:
name: comfyui
namespace: comfyui
labels:
app: comfyui
spec:
replicas: 1
strategy:
type: Recreate
selector:
matchLabels:
app: comfyui
template:
metadata:
labels:
app: comfyui
spec:
nodeSelector:
kubernetes.io/hostname: worker-1
runtimeClassName: nvidia
containers:
- name: comfyui
image: ghcr.io/ai-dock/comfyui:v2-cuda-12.1.1-base-22.04
ports:
- containerPort: 8188
protocol: TCP
env:
- name: NVIDIA_VISIBLE_DEVICES
value: "all"
- name: COMFYUI_FLAGS
value: "--listen 0.0.0.0 --port 8188"
resources:
requests:
cpu: "4"
memory: 8Gi
nvidia.com/gpu: "1"
limits:
cpu: "8"
memory: 16Gi
nvidia.com/gpu: "1"
volumeMounts:
- mountPath: /workspace/ComfyUI/models
name: models
- mountPath: /workspace/ComfyUI/output
name: output
readinessProbe:
httpGet:
path: /
port: 8188
periodSeconds: 10
initialDelaySeconds: 30
startupProbe:
httpGet:
path: /
port: 8188
failureThreshold: 60
periodSeconds: 10
volumes:
- name: models
persistentVolumeClaim:
claimName: comfyui-models
- name: output
emptyDir: {}
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: comfyui-models
namespace: comfyui
spec:
accessModes:
- ReadWriteOnce
storageClassName: longhorn
resources:
requests:
storage: 50Gi
+25
View File
@@ -0,0 +1,25 @@
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: comfyui
namespace: comfyui
annotations:
nginx.ingress.kubernetes.io/proxy-read-timeout: "600"
nginx.ingress.kubernetes.io/proxy-send-timeout: "600"
nginx.ingress.kubernetes.io/proxy-body-size: "0"
# WebSocket support for ComfyUI's live preview
nginx.ingress.kubernetes.io/proxy-http-version: "1.1"
nginx.ingress.kubernetes.io/proxy-set-headers: "Upgrade"
spec:
ingressClassName: nginx
rules:
- host: comfy.riotpiao.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: comfyui
port:
number: 80
+7
View File
@@ -0,0 +1,7 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- deployment.yaml
- service.yaml
- ingress.yaml
+14
View File
@@ -0,0 +1,14 @@
apiVersion: v1
kind: Service
metadata:
name: comfyui
namespace: comfyui
labels:
app: comfyui
spec:
selector:
app: comfyui
ports:
- port: 80
targetPort: 8188
protocol: TCP
+2
View File
@@ -14,5 +14,7 @@ resources:
- ornith.yaml - ornith.yaml
- reasoning.yaml - reasoning.yaml
- reranker.yaml - reranker.yaml
- qwen-cpu.yaml
- networkpolicy.yaml
# No namespace transformer: every file sets its own, and the transformer would # No namespace transformer: every file sets its own, and the transformer would
# rewrite metadata.namespace on anything cross-namespace added later. # rewrite metadata.namespace on anything cross-namespace added later.
+62
View File
@@ -0,0 +1,62 @@
# NetworkPolicy for LLM inference engines (llm-serving namespace).
#
# These pods have NO auth — vLLM, Ollama, and TEI accept any request.
# All access MUST go through the api-gateway, which validates JWTs and
# injects identity headers (X-Forwarded-User, X-Auth-Verified).
#
# Replaces the hand-applied llm-serving-default-deny policy that used
# `llm-client: "true"` pod label as a selector — any pod in any namespace
# could self-grant access by adding that label, which defeats the purpose.
#
# This policy restricts ingress to:
# 1. api namespace (gateway) — the sole entry point for inference
# 2. monitoring namespace — Prometheus scraping vLLM/TEI /metrics
# 3. intra-namespace — pod-to-pod (future: multi-replica comms)
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: llm-serving-ingress
namespace: llm-serving
labels:
app.kubernetes.io/part-of: llm-serving
spec:
podSelector:
matchLabels:
app.kubernetes.io/part-of: llm-serving
policyTypes:
- Ingress
ingress:
# Allow from api-gateway (namespace: api)
# Gateway proxies /v1/chat/completions, /v1/embeddings, /v1/rerank
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: api
ports:
- protocol: TCP
port: 8080 # vLLM, Ollama HTTP
- protocol: TCP
port: 80 # KServe predictor services
- protocol: TCP
port: 8000 # vLLM direct (some configs)
- protocol: TCP
port: 11434 # Ollama native port
# Allow Prometheus scraping from monitoring namespace
# vLLM: :8080/metrics, TEI: :9000/metrics
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: monitoring
ports:
- protocol: TCP
port: 8080
- protocol: TCP
port: 9000
# Allow intra-namespace (pod-to-pod within llm-serving)
- from:
- podSelector:
matchLabels:
app.kubernetes.io/part-of: llm-serving
ports:
- protocol: TCP
port: 8080
+7 -16
View File
@@ -33,12 +33,8 @@ spec:
ollama pull ornith:35b ollama pull ornith:35b
ollama pull qwen2.5:3b-instruct
ollama run ornith:35b "ok" >/dev/null 2>&1 || true ollama run ornith:35b "ok" >/dev/null 2>&1 || true
ollama run qwen2.5:3b-instruct "ok" >/dev/null 2>&1 || true
wait $SERVE_PID wait $SERVE_PID
' '
@@ -54,7 +50,7 @@ spec:
- name: OLLAMA_NUM_PARALLEL - name: OLLAMA_NUM_PARALLEL
value: '1' value: '1'
- name: OLLAMA_MAX_LOADED_MODELS - name: OLLAMA_MAX_LOADED_MODELS
value: '2' value: '1'
image: ollama/ollama:0.32.9@sha256:1685741456770df6e3cceb2a945a5f75e020f658d1701509668d6f4688f1dd3f image: ollama/ollama:0.32.9@sha256:1685741456770df6e3cceb2a945a5f75e020f658d1701509668d6f4688f1dd3f
name: kserve-container name: kserve-container
ports: ports:
@@ -65,8 +61,7 @@ spec:
command: command:
- /bin/sh - /bin/sh
- -c - -c
- ollama ps 2>/dev/null | grep -q ornith && ollama ps 2>/dev/null | - ollama ps 2>/dev/null | grep -q ornith
grep -q qwen2.5
periodSeconds: 10 periodSeconds: 10
resources: resources:
limits: limits:
@@ -82,8 +77,7 @@ spec:
command: command:
- /bin/sh - /bin/sh
- -c - -c
- ollama ps 2>/dev/null | grep -q ornith && ollama ps 2>/dev/null | - ollama ps 2>/dev/null | grep -q ornith
grep -q qwen2.5
failureThreshold: 120 failureThreshold: 120
periodSeconds: 15 periodSeconds: 15
volumeMounts: volumeMounts:
@@ -91,13 +85,10 @@ spec:
name: models name: models
deploymentStrategy: deploymentStrategy:
type: Recreate type: Recreate
# 2 replicas -- each its own GPU, each loading both ornith:35b and # 1 replica -- ornith:35b only. qwen2.5:3b moved to CPU on cp-2.
# qwen2.5:3b-instruct -- so 2 concurrent implementer-style calls each # Frees 1 GPU for ComfyUI.
# get an independent instance instead of contending on one, at the maxReplicas: 1
# cost of judge/qwen traffic still sharing whichever replica an minReplicas: 1
# implementer call also lands on.
maxReplicas: 2
minReplicas: 2
nodeSelector: nodeSelector:
kubernetes.io/hostname: worker-1 kubernetes.io/hostname: worker-1
runtimeClassName: nvidia runtimeClassName: nvidia
+115
View File
@@ -0,0 +1,115 @@
# qwen2.5:3b-instruct on CPU (talos-cp-2, 144GB RAM, 24 cores).
# Moved off GPU to free a V100 for ComfyUI. Latency ~10x slower
# than GPU but sufficient for lightweight tasks (summarization,
# classification, quick answers).
apiVersion: apps/v1
kind: Deployment
metadata:
name: qwen-cpu
namespace: llm-serving
labels:
app: qwen-cpu
app.kubernetes.io/name: qwen-cpu
app.kubernetes.io/part-of: llm-serving
spec:
replicas: 1
strategy:
type: Recreate
selector:
matchLabels:
app: qwen-cpu
template:
metadata:
labels:
app: qwen-cpu
app.kubernetes.io/name: qwen-cpu
app.kubernetes.io/part-of: llm-serving
spec:
nodeSelector:
kubernetes.io/hostname: talos-cp-2
tolerations:
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
containers:
- name: ollama
image: ollama/ollama:0.32.9@sha256:1685741456770df6e3cceb2a945a5f75e020f658d1701509668d6f4688f1dd3f
command: ["/bin/sh", "-c"]
args:
- |
ollama serve &
SERVE_PID=$!
until ollama list >/dev/null 2>&1; do sleep 2; done
ollama pull qwen2.5:3b-instruct
ollama run qwen2.5:3b-instruct "ok" >/dev/null 2>&1 || true
wait $SERVE_PID
env:
- name: OLLAMA_HOST
value: "0.0.0.0:8080"
- name: OLLAMA_MODELS
value: /root/.ollama/models
- name: OLLAMA_CONTEXT_LENGTH
value: "32768"
- name: OLLAMA_KEEP_ALIVE
value: "-1"
- name: OLLAMA_MAX_LOADED_MODELS
value: "1"
- name: OLLAMA_NUM_PARALLEL
value: "2"
ports:
- containerPort: 8080
protocol: TCP
readinessProbe:
exec:
command: ["/bin/sh", "-c", "ollama ps 2>/dev/null | grep -q qwen2.5"]
periodSeconds: 10
startupProbe:
exec:
command: ["/bin/sh", "-c", "ollama ps 2>/dev/null | grep -q qwen2.5"]
failureThreshold: 60
periodSeconds: 10
resources:
requests:
cpu: "4"
memory: 4Gi
limits:
cpu: "8"
memory: 8Gi
volumeMounts:
- mountPath: /root/.ollama
name: ollama-data
volumes:
- name: ollama-data
persistentVolumeClaim:
claimName: qwen-cpu-data
---
apiVersion: v1
kind: Service
metadata:
name: qwen-cpu
namespace: llm-serving
labels:
app: qwen-cpu
app.kubernetes.io/part-of: llm-serving
spec:
selector:
app: qwen-cpu
ports:
- port: 80
targetPort: 8080
protocol: TCP
---
# Small PVC for qwen2.5:3b model weights (~1.9GB).
# Separate from llm-models PVC which is pinned to worker-1.
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: qwen-cpu-data
namespace: llm-serving
spec:
accessModes:
- ReadWriteOnce
storageClassName: longhorn
resources:
requests:
storage: 5Gi
@@ -2,7 +2,7 @@ namespace: sqs
replicaCount: 3 replicaCount: 3
image: image:
repository: ghcr.io/riotpiaole/kmsvc-management-service repository: forgejo.riotpiao.com/rock/kmsvc-manage
tag: latest tag: latest
pullPolicy: Always pullPolicy: Always
@@ -0,0 +1,6 @@
apiVersion: v2
name: memory-queues
description: Kafka queues (DLQ) for Poimen Memory service (Phase 6.6)
type: application
version: 0.1.0
appVersion: "1.0"
@@ -0,0 +1,20 @@
{{- range .Values.queues }}
---
apiVersion: kmsvc.io/v1alpha1
kind: Queue
metadata:
name: {{ .name }}
namespace: {{ $.Values.namespace }}
labels:
app: memory-service
queue: dlq
spec:
name: {{ .name }}
description: {{ .description }}
partitions: {{ .partitions }}
replicationFactor: {{ .replicationFactor }}
config:
retention.ms: "{{ .config.retention.ms }}"
message.retention.seconds: "{{ .config.message.retention.seconds }}"
visibility.timeout.seconds: "{{ .config.visibility.timeout.seconds }}"
{{- end }}
@@ -0,0 +1,25 @@
# Poimen Memory Service Kafka Queues (kmsvc)
# Phase 6.6: DLQ topics for webhook + metrics failures
queues:
# DLQ for extraction, webhook, and agent failures
- name: poimen-memory-dlq
description: "DLQ for extraction, webhook, and agent failures"
partitions: 3
replicationFactor: 1
config:
retention.ms: "1209600000" # 14 days
message.retention.seconds: "1209600"
visibility.timeout.seconds: "300"
# DLQ for metrics persistence failures
- name: poimen-memory-metric-dlq
description: "DLQ for metrics persistence failures"
partitions: 3
replicationFactor: 1
config:
retention.ms: "1209600000" # 14 days
message.retention.seconds: "1209600"
visibility.timeout.seconds: "300"
namespace: sqs
+1 -1
View File
@@ -1,7 +1,7 @@
namespace: sqs namespace: sqs
image: image:
repository: ghcr.io/riotpiaole/kmsvc-management-service repository: forgejo.riotpiao.com/rock/kmsvc-manage
tag: latest tag: latest
pullPolicy: Always pullPolicy: Always
@@ -0,0 +1,144 @@
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
name: secretrotations.homelab.riotpiao.com
spec:
group: homelab.riotpiao.com
names:
kind: SecretRotation
plural: secretrotations
scope: Namespaced
versions:
- name: v1
served: true
storage: true
schema:
openAPIV3Schema:
type: object
properties:
metadata:
type: object
spec:
type: object
required:
- provider
- rotationInterval
properties:
# External system: authentik | forgejo | minio | vault
provider:
type: string
enum: [authentik, forgejo, minio, vault]
# How often to rotate (hours)
rotationInterval:
type: integer
minimum: 24
# Application ID in external system
appId:
type: string
# k8s Secret to update (name, namespace, key)
secretRef:
type: object
required: [name, namespace]
properties:
name:
type: string
namespace:
type: string
key:
type: string
description: "Secret key to update (e.g., MINIO_IDENTITY_OPENID_CLIENT_SECRET)"
# Path to git file that holds the secret (for .enc.yaml files)
gitPath:
type: string
description: "Path in homelab repo to .enc.yaml file"
# Ansible template values to substitute
templateValues:
type: object
additionalProperties:
type: string
status:
type: object
properties:
lastRotationTime:
type: string
format: date-time
nextRotationTime:
type: string
format: date-time
lastRotationStatus:
type: string
enum: [Success, Failed, Pending]
lastRotationError:
type: string
lastCommitHash:
type: string
---
# Example usage:
apiVersion: homelab.riotpiao.com/v1
kind: SecretRotation
metadata:
name: minio-oidc
namespace: secret-rotation
spec:
provider: authentik
rotationInterval: 2160 # 90 days in hours
appId: minio
secretRef:
name: minio-oidc
namespace: storage
key: MINIO_IDENTITY_OPENID_CLIENT_SECRET
gitPath: k8s/argocd/secrets/minio-oidc.enc.yaml
---
apiVersion: homelab.riotpiao.com/v1
kind: SecretRotation
metadata:
name: portfolio-agent-oidc
namespace: secret-rotation
spec:
provider: authentik
rotationInterval: 2160
appId: portfolio-agent
secretRef:
name: portfolio-agent-oidc
namespace: portfolio
key: CLIENT_SECRET
gitPath: k8s/argocd/secrets/portfolio-agent-oidc.enc.yaml
---
apiVersion: homelab.riotpiao.com/v1
kind: SecretRotation
metadata:
name: forgejo-registry-token
namespace: secret-rotation
spec:
provider: forgejo
rotationInterval: 2160
appId: rock/riotpiao.com
secretRef:
name: forgejo-registry-secret
namespace: kube-system
key: REGISTRY_TOKEN
gitPath: k8s/argocd/secrets/forgejo-registry-secret.enc.yaml
---
apiVersion: homelab.riotpiao.com/v1
kind: SecretRotation
metadata:
name: minio-root-credentials
namespace: secret-rotation
spec:
provider: minio
rotationInterval: 4320 # 180 days in hours
appId: root
secretRef:
name: minio-creds
namespace: storage
gitPath: k8s/argocd/secrets/minio-secrets.enc.yaml
@@ -0,0 +1,92 @@
apiVersion: apps/v1
kind: Deployment
metadata:
name: secret-rotation-controller
namespace: secret-rotation
spec:
replicas: 1
selector:
matchLabels:
app: secret-rotation-controller
template:
metadata:
labels:
app: secret-rotation-controller
spec:
serviceAccountName: secret-rotation-controller
containers:
- name: controller
image: secret-rotation-controller:latest
imagePullPolicy: IfNotPresent
env:
# SOPS reads age key from this file
- name: SOPS_AGE_KEY_FILE
value: /etc/sops/age/private-key.txt
# Vault auth (token in projected volume)
- name: VAULT_ADDR
value: http://vault.vault.svc.cluster.local:8200
- name: VAULT_TOKEN_FILE
value: /var/run/secrets/vault/token
# Authentik
- name: AUTHENTIK_URL
value: http://authentik-server.iam.svc.cluster.local
- name: AUTHENTIK_BOOTSTRAP_TOKEN
valueFrom:
secretKeyRef:
name: authentik-bootstrap
key: token
# Git
- name: GIT_REPO
value: https://forgejo.riotpiao.com/rock/homelab.git
- name: GIT_AUTHOR_EMAIL
value: [email protected]
- name: GIT_AUTHOR_NAME
value: Secret Rotation Controller
- name: FORGEJO_TOKEN
valueFrom:
secretKeyRef:
name: forgejo-registry-secret
key: REGISTRY_TOKEN
volumeMounts:
# Age key from ExternalSecret (synced from Vault)
- name: age-key
mountPath: /etc/sops/age
readOnly: true
# Vault auth token (projected)
- name: vault-token
mountPath: /var/run/secrets/vault
readOnly: true
# Temp working dir
- name: tmp
mountPath: /tmp
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: 500m
memory: 512Mi
volumes:
- name: age-key
secret:
secretName: sops-age-key
defaultMode: 0400
- name: vault-token
projected:
sources:
- serviceAccountToken:
path: token
audience: vault
expirationSeconds: 3600
- name: tmp
emptyDir: {}
@@ -0,0 +1,15 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
namespace: secret-rotation
resources:
- rbac.yaml
- crd.yaml
- external-secret.yaml
- deployment.yaml
commonLabels:
app.kubernetes.io/name: secret-rotation-controller
app.kubernetes.io/component: automation
managed-by: argocd
@@ -0,0 +1,53 @@
apiVersion: v1
kind: ServiceAccount
metadata:
name: secret-rotation-controller
namespace: secret-rotation
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: secret-rotation-controller
rules:
# Read SecretRotation CRDs
- apiGroups: ["homelab.riotpiao.com"]
resources: ["secretrotations"]
verbs: ["get", "list", "watch"]
# Update status
- apiGroups: ["homelab.riotpiao.com"]
resources: ["secretrotations/status"]
verbs: ["get", "patch", "update"]
# Read k8s secrets that will be rotated
- apiGroups: [""]
resources: ["secrets"]
verbs: ["get", "list"]
# For recording events
- apiGroups: [""]
resources: ["events"]
verbs: ["create", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: secret-rotation-controller
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: secret-rotation-controller
subjects:
- kind: ServiceAccount
name: secret-rotation-controller
namespace: secret-rotation
---
apiVersion: v1
kind: Namespace
metadata:
name: secret-rotation
labels:
kubernetes.io/metadata.name: secret-rotation
+33
View File
@@ -0,0 +1,33 @@
# ArgoCD Image Updater - auto-updates Application images from registry
# Watches forgejo.riotpiao.com for new image tags and updates Applications
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: argocd-image-updater
namespace: argocd
finalizers:
- resources-finalizer.argocd.argoproj.io
annotations:
argocd.argoproj.io/sync-wave: "1"
spec:
project: homelab
revisionHistoryLimit: 3
sources:
- repoURL: https://argoproj.github.io/argo-helm
chart: argocd-image-updater
targetRevision: "0.11.2"
helm:
valueFiles:
- $values/k8s/infra/argocd-image-updater/values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
ref: values
destination:
server: https://kubernetes.default.svc
namespace: argocd
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=false
+32
View File
@@ -0,0 +1,32 @@
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: secret-rotation
namespace: argocd
labels:
app.kubernetes.io/name: secret-rotation
spec:
project: homelab
sources:
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
path: k8s/apps/secret-rotation-controller
targetRevision: main
destination:
server: https://kubernetes.default.svc
namespace: secret-rotation
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
- RespectIgnoreDifferences=true
retry:
limit: 5
backoff:
duration: 5s
factor: 2
maxDuration: 3m
+21
View File
@@ -152,6 +152,13 @@ metadata:
namespace: argocd namespace: argocd
annotations: annotations:
argocd.argoproj.io/sync-wave: "3" argocd.argoproj.io/sync-wave: "3"
argocd-image-updater.argoproj.io/image-list: runner=forgejo.riotpiao.com/rock/forgejo-runner-golang
argocd-image-updater.argoproj.io/runner.update-strategy: newest-build
argocd-image-updater.argoproj.io/runner.allow-tags: regexp:^[0-9a-f]{7}$|^latest$|^v[0-9]+$
argocd-image-updater.argoproj.io/runner.helm.image-name: runner.image.repository
argocd-image-updater.argoproj.io/runner.helm.image-tag: runner.image.tag
argocd-image-updater.argoproj.io/write-back-method: git
argocd-image-updater.argoproj.io/git-branch: main
spec: spec:
project: homelab project: homelab
source: source:
@@ -173,6 +180,13 @@ metadata:
namespace: argocd namespace: argocd
annotations: annotations:
argocd.argoproj.io/sync-wave: "3" argocd.argoproj.io/sync-wave: "3"
argocd-image-updater.argoproj.io/image-list: runner=forgejo.riotpiao.com/rock/forgejo-runner-node
argocd-image-updater.argoproj.io/runner.update-strategy: newest-build
argocd-image-updater.argoproj.io/runner.allow-tags: regexp:^[0-9a-f]{7}$|^latest$|^v[0-9]+$
argocd-image-updater.argoproj.io/runner.helm.image-name: runner.image.repository
argocd-image-updater.argoproj.io/runner.helm.image-tag: runner.image.tag
argocd-image-updater.argoproj.io/write-back-method: git
argocd-image-updater.argoproj.io/git-branch: main
spec: spec:
project: homelab project: homelab
source: source:
@@ -197,6 +211,13 @@ metadata:
namespace: argocd namespace: argocd
annotations: annotations:
argocd.argoproj.io/sync-wave: "3" argocd.argoproj.io/sync-wave: "3"
argocd-image-updater.argoproj.io/image-list: runner=forgejo.riotpiao.com/rock/forgejo-runner-rust
argocd-image-updater.argoproj.io/runner.update-strategy: newest-build
argocd-image-updater.argoproj.io/runner.allow-tags: regexp:^[0-9a-f]{7}$|^latest$|^v[0-9]+$
argocd-image-updater.argoproj.io/runner.helm.image-name: runner.image.repository
argocd-image-updater.argoproj.io/runner.helm.image-tag: runner.image.tag
argocd-image-updater.argoproj.io/write-back-method: git
argocd-image-updater.argoproj.io/git-branch: main
spec: spec:
project: homelab project: homelab
source: source:
+20
View File
@@ -0,0 +1,20 @@
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: memory-queues
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "7"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/apps/messaging/memory-queues
destination:
server: https://kubernetes.default.svc
namespace: sqs
syncPolicy:
automated:
prune: true
selfHeal: true
+5 -45
View File
@@ -1,6 +1,6 @@
# Wave 5 — Kafka (Strimzi operator + cluster CR), Redis, and the SQS-like # Wave 5 — Kafka (Strimzi operator + cluster CR), Redis infrastructure.
# queue services. Strimzi/Redis are public Helm charts; kafka-cluster/queue-crd/ # Strimzi/Redis are public Helm charts; kafka-cluster is a local chart.
# management-service are local charts (rendered from their own Chart.yaml). # queue-crd and management-service are managed by kmsvc-root (kmsvc-manage.git).
apiVersion: argoproj.io/v1alpha1 apiVersion: argoproj.io/v1alpha1
kind: Application kind: Application
metadata: metadata:
@@ -78,45 +78,5 @@ spec:
automated: automated:
prune: true prune: true
selfHeal: true selfHeal: true
--- # queue-crd and management-service moved to kmsvc-manage.git repo
apiVersion: argoproj.io/v1alpha1 # Managed by kmsvc-root Application
kind: Application
metadata:
name: queue-crd
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "6"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/apps/messaging/queue-crd
destination:
server: https://kubernetes.default.svc
namespace: sqs
syncPolicy:
automated:
prune: true
selfHeal: true
---
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: management-service
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "7"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/apps/messaging/management-service
destination:
server: https://kubernetes.default.svc
namespace: sqs
syncPolicy:
automated:
prune: true
selfHeal: true
+5
View File
@@ -36,6 +36,11 @@ metadata:
app.kubernetes.io/component: gateway app.kubernetes.io/component: gateway
annotations: annotations:
argocd.argoproj.io/sync-wave: "7" argocd.argoproj.io/sync-wave: "7"
# ArgoCD Image Updater - auto-update on new image push
argocd-image-updater.argoproj.io/image-list: gw=forgejo.riotpiao.com/rock/api-gateway
argocd-image-updater.argoproj.io/gw.update-strategy: digest
argocd-image-updater.argoproj.io/gw.allow-tags: regexp:^latest$
argocd-image-updater.argoproj.io/write-back-method: argocd
spec: spec:
project: homelab project: homelab
revisionHistoryLimit: 3 revisionHistoryLimit: 3
+32
View File
@@ -0,0 +1,32 @@
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: comfyui
namespace: argocd
labels:
app.kubernetes.io/name: comfyui
app.kubernetes.io/component: image-generation
annotations:
argocd.argoproj.io/sync-wave: "8"
spec:
project: homelab
revisionHistoryLimit: 3
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/apps/comfyui
destination:
server: https://kubernetes.default.svc
namespace: comfyui
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
retry:
limit: 5
backoff:
duration: 5s
factor: 2
maxDuration: 3m
+7 -1
View File
@@ -246,6 +246,11 @@ metadata:
namespace: argocd namespace: argocd
annotations: annotations:
argocd.argoproj.io/sync-wave: "8" argocd.argoproj.io/sync-wave: "8"
# ArgoCD Image Updater - auto-update on new image push
argocd-image-updater.argoproj.io/image-list: app=forgejo.riotpiao.com/rock/portfolio
argocd-image-updater.argoproj.io/app.update-strategy: digest
argocd-image-updater.argoproj.io/app.allow-tags: regexp:^latest$
argocd-image-updater.argoproj.io/write-back-method: argocd
spec: spec:
project: homelab project: homelab
source: source:
@@ -281,7 +286,8 @@ spec:
path: k8s/infra/rbac path: k8s/infra/rbac
destination: destination:
server: https://kubernetes.default.svc server: https://kubernetes.default.svc
namespace: default # No namespace: cluster-scoped resources (ClusterRoleBinding, etc.)
# Namespace is set per-resource in kustomization
syncPolicy: syncPolicy:
automated: automated:
prune: true prune: true
+23 -16
View File
@@ -1,27 +1,34 @@
# Poimen project collection — manages poimen-memory, poimen-workflows, and poiman
# Each repo tracks its own main branch (no prod branch). Poiman is the primary
# orchestrator with k8s/argocd/ containing the AppProject and deployment structure.
#
# CI: All three repos trigger on main branch pushes (no image builds yet).
# Future: Add build workflows for poiman once container runtime needs are clear.
apiVersion: argoproj.io/v1alpha1 apiVersion: argoproj.io/v1alpha1
kind: Application kind: Application
metadata: metadata:
name: poimen-root name: poimen
namespace: argocd namespace: argocd
labels:
app.kubernetes.io/name: poimen
app.kubernetes.io/component: orchestrator
annotations: annotations:
argocd.argoproj.io/sync-wave: "7" argocd.argoproj.io/sync-wave: "7"
# Image Updater: auto-update on new image push (SHA tag filter)
argocd-image-updater.argoproj.io/image-list: |
memory=forgejo.riotpiao.com/rock/poimen-memory
workflows=forgejo.riotpiao.com/rock/poimen-workflows
frontend=forgejo.riotpiao.com/rock/poimen-frontend
argocd-image-updater.argoproj.io/memory.update-strategy: digest
argocd-image-updater.argoproj.io/memory.allow-tags: regexp:^latest$
argocd-image-updater.argoproj.io/workflows.update-strategy: digest
argocd-image-updater.argoproj.io/workflows.allow-tags: regexp:^latest$
argocd-image-updater.argoproj.io/frontend.update-strategy: digest
argocd-image-updater.argoproj.io/frontend.allow-tags: regexp:^latest$
argocd-image-updater.argoproj.io/write-back-method: argocd
spec: spec:
project: homelab project: homelab
source: sources:
repoURL: https://forgejo.riotpiao.com/rock/poimen.git - repoURL: https://forgejo.riotpiao.com/rock/poimen-memory.git
targetRevision: main targetRevision: main
path: k8s/argocd path: k8s/argocd
directory: - repoURL: https://forgejo.riotpiao.com/rock/poimen-workflows.git
recurse: false targetRevision: main
path: k8s/argocd
- repoURL: https://forgejo.riotpiao.com/rock/poimen-frontend.git
targetRevision: main
path: k8s/argocd
destination: destination:
server: https://kubernetes.default.svc server: https://kubernetes.default.svc
namespace: poimen namespace: poimen
+4
View File
@@ -21,6 +21,8 @@ spec:
- https://forgejo.riotpiao.com/rock/homelab-frontend.git - https://forgejo.riotpiao.com/rock/homelab-frontend.git
- https://forgejo.riotpiao.com/rock/kmsvc-manage.git - https://forgejo.riotpiao.com/rock/kmsvc-manage.git
- https://forgejo.riotpiao.com/rock/poimen.git - https://forgejo.riotpiao.com/rock/poimen.git
- https://forgejo.riotpiao.com/rock/poimen-memory.git
- https://forgejo.riotpiao.com/rock/poimen-workflows.git
- https://forgejo.riotpiao.com/rock/riotpiao.com.git - https://forgejo.riotpiao.com/rock/riotpiao.com.git
# Public Helm chart repos referenced by k8s/argocd/apps/* and bootstrap/* # Public Helm chart repos referenced by k8s/argocd/apps/* and bootstrap/*
- https://cloudnative-pg.github.io/charts - https://cloudnative-pg.github.io/charts
@@ -40,6 +42,8 @@ spec:
- https://charts.jetstack.io - https://charts.jetstack.io
- https://kubernetes.github.io/ingress-nginx - https://kubernetes.github.io/ingress-nginx
- https://stakater.github.io/stakater-charts - https://stakater.github.io/stakater-charts
# ArgoCD ecosystem charts
- https://argoproj.github.io/argo-helm
destinations: destinations:
- server: https://kubernetes.default.svc - server: https://kubernetes.default.svc
namespace: "*" namespace: "*"
+1
View File
@@ -26,3 +26,4 @@ spec:
selfHeal: true selfHeal: true
syncOptions: syncOptions:
- CreateNamespace=true - CreateNamespace=true
- ServerSideApply=true
+6 -4
View File
@@ -1,8 +1,10 @@
apiVersion: kustomize.config.k8s.io/v1beta1 apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization kind: Kustomization
# All homelab SOPS-encrypted Secrets, decrypted in-line via the ksops generator.
# Each *.enc.yaml carries its own metadata.namespace, so no namespace transformer # Disable hash suffix for all generated secrets (stable names)
# here (that would rewrite every Secret into one namespace). Renders exactly the generatorOptions:
# Secret objects — replaces the old argocd-cmp-cm SOPS plugin. disableNameSuffixHash: true
# SOPS-encrypted secrets via ksops generator
generators: generators:
- secret-generator.yaml - secret-generator.yaml
@@ -0,0 +1,2 @@
FORGEJO_TOKEN=273fdcffabbcbb5a191e8289c73d106063acefc6
LLM_API_TOKEN=s3VksXyw2z3sGbegnwjMDFnJ6CtNRd1a5CcnE5A4ET77toCcykNdunk6Oa2J
@@ -0,0 +1,72 @@
# ArgoCD Image Updater configuration
# Watches Forgejo registry and updates ArgoCD Applications with new image tags
config:
# Registry configuration - Forgejo allows anonymous pulls
registries:
- name: forgejo
api_url: https://forgejo.riotpiao.com
prefix: forgejo.riotpiao.com
default: true
insecure: false
# Log level
logLevel: debug
# ArgoCD API server
argocd:
grpcWeb: true
serverAddress: argocd-server.argocd.svc.cluster.local
insecure: true
plaintext: true
# Git write-back configuration (for multi-source Applications)
git:
# Commit author for image updates
user:
name: "ArgoCD Image Updater"
email: "[email protected]"
# Use SSH keys from ArgoCD's known hosts + credentials
# Image Updater inherits ArgoCD's git credentials (mounted via ArgoCD secret)
# Mount ArgoCD's git credentials for write-back
extraVolumes:
- name: argocd-ssh-known-hosts-cm
configMap:
name: argocd-ssh-known-hosts-cm
defaultMode: 0644
- name: argocd-gpg-keys-cm
configMap:
name: argocd-gpg-keys-cm
optional: true
defaultMode: 0644
- name: argocd-gpg-pubring
configMap:
name: argocd-gpg-pubring-cm
optional: true
defaultMode: 0644
extraVolumeMounts:
- name: argocd-ssh-known-hosts-cm
mountPath: /etc/ssh/ssh_known_hosts.d/argocd-ssh-known-hosts
subPath: ssh_known_hosts
- name: argocd-gpg-keys-cm
mountPath: /etc/gpg/source
- name: argocd-gpg-pubring
mountPath: /etc/gpg/pubring
# Extra environment variables
extraEnv:
- name: ARGOCD_GRPC_WEB
value: "true"
- name: GIT_SSH_KNOWN_HOSTS_CONFIG_MAP_ENABLED
value: "true"
# Resources
resources:
requests:
cpu: 50m
memory: 64Mi
limits:
cpu: 200m
memory: 128Mi
@@ -36,3 +36,4 @@ data:
valid_volumes: valid_volumes:
- /docker-certs/client - /docker-certs/client
network: host network: host
docker_host: automount
@@ -34,7 +34,11 @@ spec:
command: ["sh", "-c"] command: ["sh", "-c"]
args: args:
- | - |
test -f /data/.runner || forgejo-runner register --no-interactive \ # Always re-register to keep labels in sync with values.yaml.
# Without this, changing a runner label requires manually deleting
# the PVC or .runner file — not GitOps-friendly.
rm -f /data/.runner
forgejo-runner register --no-interactive \
--instance {{ .Values.runner.forgejoUrl }} \ --instance {{ .Values.runner.forgejoUrl }} \
--token $(RUNNER_TOKEN) \ --token $(RUNNER_TOKEN) \
--name {{ .Values.runner.name }} \ --name {{ .Values.runner.name }} \
@@ -56,20 +60,18 @@ spec:
containers: containers:
- name: runner - name: runner
image: {{ .Values.runner.image.repository }}:{{ .Values.runner.image.tag }} image: {{ .Values.runner.image.repository }}:{{ .Values.runner.image.tag }}
command: ["sh", "-c", "forgejo-runner daemon --config /etc/forgejo-runner/config.yaml"] command: ["sh", "-c", "while ! wget -q -O- http://localhost:2375/_ping >/dev/null 2>&1; do echo 'waiting for dind...'; sleep 2; done; echo 'dind ready'; forgejo-runner daemon --config /etc/forgejo-runner/config.yaml"]
workingDir: /data workingDir: /data
env: env:
- name: DOCKER_HOST - name: DOCKER_HOST
value: tcp://localhost:2376 value: tcp://localhost:2375
- name: DOCKER_TLS_VERIFY
value: "1"
- name: DOCKER_CERT_PATH
value: /docker-certs/client
volumeMounts: volumeMounts:
- name: runner-data - name: runner-data
mountPath: /data mountPath: /data
- name: docker-certs - name: docker-certs
mountPath: /docker-certs mountPath: /docker-certs
- name: docker-sock
mountPath: /run
- name: homelab-ca - name: homelab-ca
mountPath: /etc/ssl/certs/homelab-ca.pem mountPath: /etc/ssl/certs/homelab-ca.pem
subPath: ca.crt subPath: ca.crt
@@ -85,10 +87,12 @@ spec:
privileged: true # required for DinD; cicd namespace is labelled privileged privileged: true # required for DinD; cicd namespace is labelled privileged
env: env:
- name: DOCKER_TLS_CERTDIR - name: DOCKER_TLS_CERTDIR
value: /docker-certs value: ""
volumeMounts: volumeMounts:
- name: docker-certs - name: docker-certs
mountPath: /docker-certs mountPath: /docker-certs
- name: docker-sock
mountPath: /run
- name: dind-storage - name: dind-storage
mountPath: /var/lib/docker mountPath: /var/lib/docker
- name: homelab-ca - name: homelab-ca
@@ -113,6 +117,8 @@ spec:
claimName: {{ .Release.Name }}-dind claimName: {{ .Release.Name }}-dind
- name: docker-certs - name: docker-certs
emptyDir: {} # DinD regenerates mTLS certs on each start emptyDir: {} # DinD regenerates mTLS certs on each start
- name: docker-sock
emptyDir: {} # Shared docker socket between dind and runner
- name: homelab-ca - name: homelab-ca
# homelab-ca is a ConfigMap (public CA trust bundle), not a Secret. # homelab-ca is a ConfigMap (public CA trust bundle), not a Secret.
# The volumeMounts use subPath: ca.crt to project the single cert file. # The volumeMounts use subPath: ca.crt to project the single cert file.
@@ -43,7 +43,7 @@ metadata:
labels: labels:
app: forgejo-runner-gc app: forgejo-runner-gc
spec: spec:
schedule: "0 2,5 * * *" # Run at 02:00 and 05:00 UTC (more frequent for heavy build workloads) schedule: {{ .Values.gc.schedule | quote }}
concurrencyPolicy: Forbid concurrencyPolicy: Forbid
successfulJobsHistoryLimit: 3 successfulJobsHistoryLimit: 3
failedJobsHistoryLimit: 3 failedJobsHistoryLimit: 3
+5 -2
View File
@@ -2,9 +2,12 @@
# runner instance. Only runner.name and runner.labels differ -- everything # runner instance. Only runner.name and runner.labels differ -- everything
# else (image, dind, persistence, tolerations, nodeSelector) is shared. # else (image, dind, persistence, tolerations, nodeSelector) is shared.
# #
# node:22-bookworm ships Node natively, so unlike the golang/rust instances, # Label image: node:22-bookworm — Debian, root, apt-get, Node.js, npm, git.
# jobs on this runner need no "install node" step before actions/checkout. # Install docker in workflow steps as needed.
runner: runner:
image:
repository: code.forgejo.org/forgejo/runner
tag: "6"
name: node-runner name: node-runner
labels: "node:docker://node:22-bookworm" labels: "node:docker://node:22-bookworm"
+7 -8
View File
@@ -2,17 +2,16 @@
# runner instance. Only runner.name and runner.labels differ -- everything # runner instance. Only runner.name and runner.labels differ -- everything
# else (image, dind, persistence, tolerations, nodeSelector) is shared. # else (image, dind, persistence, tolerations, nodeSelector) is shared.
# #
# rust:1.83-bookworm -- verified this tag exists (docker manifest inspect) # Label image: rust:1-bookworm — Debian, root, apt-get, Rust, cargo, git.
# before pinning it, per this repo's convention of not trusting a tag exists # Install Node.js/docker in workflow steps as needed.
# without checking.
runner: runner:
image:
repository: code.forgejo.org/forgejo/runner
tag: "6"
name: rust-runner name: rust-runner
labels: "rust:docker://rust:1.83-bookworm" labels: "rust:docker://rust:1-bookworm"
persistence:
reg:
storageClass: longhorn
size: 20Gi
# GC CronJob renders only from the default (golang) values to avoid duplicates # GC CronJob renders only from the default (golang) values to avoid duplicates
gc: gc:
+7 -14
View File
@@ -1,20 +1,13 @@
runner: runner:
image: image:
repository: code.forgejo.org/forgejo/runner repository: code.forgejo.org/forgejo/runner
tag: "6" # pin exact release before apply tag: "6"
name: golang-runner name: golang-runner
# Default image is only used when a job's `container:` doesn't override it # Label image is what workflow steps run in (NOT the runner daemon image).
# (both ci.yaml and build.yaml in homelab-frontend do). Retired the old # golang:1.26-bookworm: Debian, root, apt-get, Go, git.
# "docker" label entirely; every repo this runner serves is Go, so this # TODO: Switch to custom image once build-runner-images.yml pushes images
# instance carries the golang toolchain and its own dind sidecar builds and
# pushes that repo's images too -- there is no separate generic runner
# anymore.
labels: "golang:docker://golang:1.26-bookworm" labels: "golang:docker://golang:1.26-bookworm"
# In-cluster Service (:3000) — direct, avoids the ingress/public-hostname hop
# (the public URL is :443 which forgejo doesn't serve; runner got i/o timeout).
forgejoUrl: http://forgejo-gitea-http.cicd.svc.cluster.local:3000 forgejoUrl: http://forgejo-gitea-http.cicd.svc.cluster.local:3000
# tokenSecret: name of the K8s Secret that holds the runner registration token
# created automatically by the helmfile presync hook (see helmfile.yaml.gotmpl)
tokenSecret: runner-token tokenSecret: runner-token
resources: resources:
requests: requests:
@@ -59,8 +52,8 @@ nodeSelector:
# disable in per-runner overrides so it renders once. # disable in per-runner overrides so it renders once.
gc: gc:
enabled: true enabled: true
schedule: "0 3 * * *" # daily 03:00 UTC schedule: "*/30 * * * *" # every 30 minutes
image: alpine/k8s:1.31.0 image: alpine/k8s:1.31.0
pruneAge: "72h" # Docker artifacts unused longer than this get pruned pruneAge: "30m" # Docker artifacts unused longer than this get pruned
pruneAgeHours: 72 # Same as pruneAge but numeric for date arithmetic in shell pruneAgeHours: 0.5 # Same as pruneAge but numeric for date arithmetic in shell
actcacheMaxAgeDays: 1 # actcache files older than N days (aggressive for heavy Rust cargo builds) actcacheMaxAgeDays: 1 # actcache files older than N days (aggressive for heavy Rust cargo builds)
-209
View File
@@ -1,209 +0,0 @@
# Authentik OAuth provisioning — PostSync hook, reruns on every ArgoCD sync
# (hook-delete-policy: BeforeHookCreation deletes the previous run's Job before
# creating a new one, so this stays reconciled the same way the rest of the
# cluster does — no separate manual bootstrap step like setup_talos_iam.sh /
# provision_oidc.py, which never got migrated off the old helmfile workflow).
#
# What it does (see scripts/authentik-provision.py docstring): creates the
# "groups" scope mapping, homelab-admins / grafana-admins groups, the "rock"
# admin user, OAuth2 providers + Applications for grafana/minio/forgejo/argocd,
# and binds homelab-admins to all of them. The script is generated into the
# authentik-provision-script ConfigMap by kustomize configMapGenerator (see
# kustomization.yaml), not embedded here.
#
# RBAC: this Job only touches Secrets (get existing client secrets, create new
# ones for forgejo/argocd/rock) across the namespaces those services live in.
# It never touches any other resource type.
apiVersion: v1
kind: ServiceAccount
metadata:
name: authentik-provisioner
namespace: iam
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: authentik-provisioner
rules:
- apiGroups: [""]
resources: ["secrets"]
verbs: ["get", "list", "create", "update", "patch"]
---
# One RoleBinding per namespace the script touches (least-privilege: Secrets
# only, and only in these 5 namespaces — not a cluster-wide ClusterRoleBinding).
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: authentik-provisioner
namespace: iam
subjects:
- kind: ServiceAccount
name: authentik-provisioner
namespace: iam
roleRef:
kind: ClusterRole
name: authentik-provisioner
apiGroup: rbac.authorization.k8s.io
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: authentik-provisioner
namespace: cicd
subjects:
- kind: ServiceAccount
name: authentik-provisioner
namespace: iam
roleRef:
kind: ClusterRole
name: authentik-provisioner
apiGroup: rbac.authorization.k8s.io
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: authentik-provisioner
namespace: argocd
subjects:
- kind: ServiceAccount
name: authentik-provisioner
namespace: iam
roleRef:
kind: ClusterRole
name: authentik-provisioner
apiGroup: rbac.authorization.k8s.io
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: authentik-provisioner
namespace: logging
subjects:
- kind: ServiceAccount
name: authentik-provisioner
namespace: iam
roleRef:
kind: ClusterRole
name: authentik-provisioner
apiGroup: rbac.authorization.k8s.io
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: authentik-provisioner
namespace: storage
subjects:
- kind: ServiceAccount
name: authentik-provisioner
namespace: iam
roleRef:
kind: ClusterRole
name: authentik-provisioner
apiGroup: rbac.authorization.k8s.io
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: authentik-provisioner
namespace: paperless
subjects:
- kind: ServiceAccount
name: authentik-provisioner
namespace: iam
roleRef:
kind: ClusterRole
name: authentik-provisioner
apiGroup: rbac.authorization.k8s.io
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: authentik-provisioner
namespace: immich
subjects:
- kind: ServiceAccount
name: authentik-provisioner
namespace: iam
roleRef:
kind: ClusterRole
name: authentik-provisioner
apiGroup: rbac.authorization.k8s.io
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: authentik-provisioner
namespace: llm-serving
subjects:
- kind: ServiceAccount
name: authentik-provisioner
namespace: iam
roleRef:
kind: ClusterRole
name: authentik-provisioner
apiGroup: rbac.authorization.k8s.io
---
apiVersion: batch/v1
kind: Job
metadata:
name: authentik-provision
namespace: iam
annotations:
argocd.argoproj.io/hook: PostSync
argocd.argoproj.io/hook-delete-policy: BeforeHookCreation
spec:
ttlSecondsAfterFinished: 600
backoffLimit: 3
template:
spec:
serviceAccountName: authentik-provisioner
restartPolicy: Never
securityContext:
runAsNonRoot: true
runAsUser: 1000
seccompProfile:
type: RuntimeDefault
containers:
- name: provision
image: python:3.12-alpine
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
env:
- name: AUTHENTIK_BOOTSTRAP_TOKEN
valueFrom:
secretKeyRef:
name: authentik-secrets
key: AUTHENTIK_BOOTSTRAP_TOKEN
volumeMounts:
- name: script
mountPath: /script
command:
- /bin/sh
- -c
- |
set -e
echo "waiting for authentik-server..."
until wget -q -O /dev/null http://authentik-server.iam.svc.cluster.local/-/health/ready/ 2>/dev/null; do
sleep 5
done
echo "installing kubectl (via python urllib - no apk/curl: this"
echo "container runs as non-root UID 1000 and can't write to"
echo "apk's directories or /usr/local/bin, both root-owned in"
echo "the python:3.12-alpine image; /tmp is world-writable)..."
python3 -c "
import urllib.request, os, stat
kver = urllib.request.urlopen('https://dl.k8s.io/release/stable.txt').read().decode().strip()
url = f'https://dl.k8s.io/release/{kver}/bin/linux/amd64/kubectl'
urllib.request.urlretrieve(url, '/tmp/kubectl')
st = os.stat('/tmp/kubectl')
os.chmod('/tmp/kubectl', st.st_mode | stat.S_IEXEC)
"
export PATH="/tmp:$PATH"
echo "running provisioning script..."
python3 /script/authentik-provision.py
volumes:
- name: script
configMap:
name: authentik-provision-script
+7 -30
View File
@@ -1,35 +1,12 @@
apiVersion: kustomize.config.k8s.io/v1beta1 apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization kind: Kustomization
# NOTE: no top-level `namespace:` transformer here (removed) - it used to
# force-rewrite metadata.namespace to "iam" on every resource in this
# kustomization, which was harmless while every manifest here only ever
# targeted the iam namespace itself. authentik-provision-job.yaml's
# RoleBindings deliberately target cicd/argocd/logging/storage (least-
# privilege access for the authentik-provisioner ServiceAccount to touch
# Secrets in those namespaces) - the namespace transformer would have
# silently rewritten all of them back to iam, breaking the RBAC. Every
# manifest in this directory already sets its own explicit
# metadata.namespace, so dropping the transformer changes nothing for the
# existing resources/.
resources: resources:
- authentik-provision-job.yaml
- rbac-dashboard-rolebinding.yaml - rbac-dashboard-rolebinding.yaml
# Provisioning/verification python lives in scripts/*.py (real files, linted + # IAM provisioning is manual-only (security-sensitive).
# diff-friendly) and is generated into ConfigMaps here rather than embedded in # Script: scripts/iam/authentik-provision.py
# the job YAML. disableNameSuffixHash keeps the names stable so the Jobs' # Run:
# configMap volume refs and PostSync hook-delete semantics keep working; each # export AUTHENTIK_BOOTSTRAP_TOKEN=$(kubectl -n iam get secret authentik-secrets \
# hook Job is recreated per sync so it always mounts the latest script. # -o jsonpath='{.data.AUTHENTIK_BOOTSTRAP_TOKEN}' | base64 -d)
configMapGenerator: # python3 scripts/iam/authentik-provision.py
- name: authentik-provision-script
namespace: iam
files:
- authentik-provision.py=scripts/authentik-provision.py
generatorOptions:
disableNameSuffixHash: true
# authentik-migrations-job.yaml removed — redundant + broken. The authentik
# `server` entrypoint runs migrations itself; this standalone job lacked the
# authentik-secrets envFrom (Secret key missing) and always failed.
# SOPS secrets (*.enc.yaml) handled by ArgoCD SOPS plugin at sync time
# authentik/vault deployed via ArgoCD Helm source
@@ -1,707 +0,0 @@
#!/usr/bin/env python3
"""
Authentik OAuth provisioning - idempotent, safe to re-run (ArgoCD PostSync hook).
Creates/updates, in order:
1. A custom "groups" OAuth2 scope mapping (Authentik ships openid/email/profile
by default but NOT groups - required for ArgoCD RBAC group mapping and
Grafana's role_attribute_path, both of which read a `groups` claim).
2. Groups: homelab-admins (is_superuser=true), grafana-admins.
3. User "rock": created if missing, always (re-)synced into both groups above.
Password is generated once and only written to the k8s Secret
rock-credentials (iam ns) the first time the user is created - re-runs
never rotate an existing password.
4. OAuth2/OIDC providers + Applications for: grafana, minio, forgejo, argocd.
Client secrets are read from existing k8s Secrets (grafana-oidc, minio-oidc)
if present, or generated once and written out (forgejo-oidc, oidc-secret)
the first time.
5. PolicyBinding of homelab-admins -> every Application above, so "rock" (and
anyone else in that group) has guaranteed access regardless of each app's
default visibility.
Talks to Authentik over the in-cluster Service (authentik-server.iam.svc:80),
authenticating with the bootstrap token. Everything is done with GET-then-
create-or-patch so this can be re-run on every ArgoCD sync without duplicating
or clobbering objects (PostSync hook, not a one-shot Job with hook-delete).
kubectl is used only to read/write the small set of Secrets this script
touches - it shells out rather than using the Python k8s client to keep the
container image to stdlib Python + the kubectl binary, no pip installs.
"""
import json
import os
import secrets
import string
import subprocess
import sys
import urllib.error
import urllib.request
AUTHENTIK_URL = "http://authentik-server.iam.svc.cluster.local"
TOKEN = os.environ["AUTHENTIK_BOOTSTRAP_TOKEN"]
def api(method, path, data=None):
url = f"{AUTHENTIK_URL}{path}"
body = json.dumps(data).encode() if data is not None else None
req = urllib.request.Request(
url,
data=body,
method=method,
headers={
"Authorization": f"Bearer {TOKEN}",
"Content-Type": "application/json",
},
)
try:
with urllib.request.urlopen(req, timeout=30) as resp:
raw = resp.read()
return resp.status, (json.loads(raw) if raw else {})
except urllib.error.HTTPError as e:
raw = e.read()
try:
parsed = json.loads(raw) if raw else {}
except json.JSONDecodeError:
parsed = {"raw": raw.decode(errors="replace")}
return e.code, parsed
def die(msg):
print(f"FATAL: {msg}", file=sys.stderr)
sys.exit(1)
def gen_secret(n=40):
alphabet = string.ascii_letters + string.digits
return "".join(secrets.choice(alphabet) for _ in range(n))
def kubectl_get_secret_key(namespace, name, key):
"""Returns decoded value, or None if the secret/key doesn't exist."""
p = subprocess.run(
["kubectl", "-n", namespace, "get", "secret", name, "-o", f"jsonpath={{.data.{key}}}"],
capture_output=True, text=True,
)
if p.returncode != 0 or not p.stdout.strip():
return None
import base64
return base64.b64decode(p.stdout).decode()
def kubectl_create_secret(namespace, name, literals: dict, labels: dict = None):
"""Idempotent: create-or-update via dry-run|apply, same pattern used
elsewhere in this repo (setup_vault.sh, apply-vault-secrets.sh)."""
args = ["kubectl", "-n", namespace, "create", "secret", "generic", name]
for k, v in literals.items():
args += [f"--from-literal={k}={v}"]
args += ["--dry-run=client", "-o", "yaml"]
render = subprocess.run(args, capture_output=True, text=True)
if render.returncode != 0:
die(f"rendering secret {namespace}/{name}: {render.stderr}")
apply = subprocess.run(["kubectl", "apply", "-f", "-"], input=render.stdout,
capture_output=True, text=True)
if apply.returncode != 0:
die(f"applying secret {namespace}/{name}: {apply.stderr}")
print(f" secret {namespace}/{name}: {apply.stdout.strip()}")
if labels:
# argocd's `$secret:key` substitution only reads Secrets carrying
# app.kubernetes.io/part-of: argocd — without it OIDC login fails with
# oauth2 "invalid_client" (empty client_secret sent to the IdP).
label_args = ["kubectl", "-n", namespace, "label", "secret", name,
"--overwrite"] + [f"{k}={v}" for k, v in labels.items()]
subprocess.run(label_args, capture_output=True, text=True)
def get_or_create(list_path, create_path, query, payload, patch_existing=None):
status, res = api("GET", f"{list_path}?{query}")
if status != 200:
die(f"GET {list_path}?{query} -> {status} {res}")
results = res.get("results", [])
if results:
obj = results[0]
if patch_existing:
status, obj2 = api("PATCH", f"{create_path}{obj['pk']}/", patch_existing)
if status not in (200, 201):
die(f"PATCH {create_path}{obj['pk']}/ -> {status} {obj2}")
return obj2
return obj
status, obj = api("POST", create_path, payload)
if status not in (200, 201):
die(f"POST {create_path} -> {status} {obj}")
return obj
# -----------------------------------------------------------------------------
print("[1/5] Ensuring custom 'groups' scope mapping exists...")
groups_mapping = get_or_create(
"/api/v3/propertymappings/provider/scope/",
"/api/v3/propertymappings/provider/scope/",
"scope_name=groups",
{
"name": "homelab: groups claim",
"scope_name": "groups",
# request.user.ak_groups is deprecated in authentik 2026.x (logs a
# deprecation warning on every token issue) -> use request.user.groups.
"expression": (
"return {\"groups\": [group.name for group in request.user.groups.all()]}"
),
},
# Force the expression onto the already-created mapping on re-run.
patch_existing={
"expression": (
"return {\"groups\": [group.name for group in request.user.groups.all()]}"
),
},
)
GROUPS_MAPPING_PK = groups_mapping["pk"]
# Generic "permissions" claim, computed from group membership - lets each app
# (and eventually k8s RBAC via --oidc-groups-claim) check a permission string
# like "paperless:write" instead of hardcoding a group name. homelab-admins
# gets "*" (everything); every other admin group gets its own read+write pair.
# k8s-devops-admin is declared but has no k8s Role/RoleBinding target yet -
# foundation for a future short-lived federated-operator credential.
_PERMISSIONS_EXPR = """
GROUP_PERMISSIONS = {
"homelab-admins": ["*"],
"grafana-admins": ["grafana:read", "grafana:write"],
"minio-admins": ["minio:read", "minio:write"],
"forgejo-admins": ["forgejo:read", "forgejo:write"],
"homarr-admins": ["homarr:read", "homarr:write"],
"portainer-admins": ["portainer:read", "portainer:write"],
"kmsvc-admins": ["kmsvc:read", "kmsvc:write"],
"temporal-admins": ["temporal:read", "temporal:write"],
"llm-admins": ["llm:read", "llm:write"],
"paperless-admins": ["paperless:read", "paperless:write"],
"immich-admins": ["immich:read", "immich:write"],
"poimen-memory-admins": ["poimen-memory:read", "poimen-memory:write"],
"k8s-devops-admin": ["k8s:devops"],
"vault-service-api": ["vault:read", "vault:write"],
}
perms = set()
for group in request.user.groups.all():
perms.update(GROUP_PERMISSIONS.get(group.name, []))
return {"permissions": sorted(perms)}
""".strip()
permissions_mapping = get_or_create(
"/api/v3/propertymappings/provider/scope/",
"/api/v3/propertymappings/provider/scope/",
"scope_name=permissions",
{
"name": "homelab: permissions claim",
"scope_name": "permissions",
"expression": _PERMISSIONS_EXPR,
},
patch_existing={"expression": _PERMISSIONS_EXPR},
)
PERMISSIONS_MAPPING_PK = permissions_mapping["pk"]
# Immich reads a "immich_role" claim on every login (not just user-creation -
# fixed upstream in immich-app/immich#29991) and syncs isAdmin from it, so
# this is the actual mechanism that makes "rock" an Immich admin - not
# Immich's first-user-is-admin fallback, which races badly with OAuth login.
_IMMICH_ROLE_EXPR = (
"return {\"immich_role\": \"admin\" "
"if request.user.ak_groups.filter(name__in=[\"homelab-admins\", \"immich-admins\"]).exists() "
"else \"user\"}"
)
immich_role_mapping = get_or_create(
"/api/v3/propertymappings/provider/scope/",
"/api/v3/propertymappings/provider/scope/",
"scope_name=immich_role",
{
"name": "homelab: immich role claim",
"scope_name": "immich_role",
"expression": _IMMICH_ROLE_EXPR,
},
patch_existing={"expression": _IMMICH_ROLE_EXPR},
)
IMMICH_ROLE_MAPPING_PK = immich_role_mapping["pk"]
# MinIO maps OIDC users to a MinIO policy via a "policy" claim
# (MINIO_IDENTITY_OPENID_CLAIM_NAME=policy). Emit consoleAdmin (full admin) for
# homelab-admins members, readonly for everyone else. Without this claim MinIO
# assigns no policy and OIDC users get no access.
_POLICY_EXPR = (
"return {\"policy\": \"consoleAdmin\" "
"if request.user.ak_groups.filter(name=\"homelab-admins\").exists() "
"else \"readonly\"}"
)
policy_mapping = get_or_create(
"/api/v3/propertymappings/provider/scope/",
"/api/v3/propertymappings/provider/scope/",
"scope_name=minio",
{
"name": "homelab: minio policy claim",
"scope_name": "minio",
"expression": _POLICY_EXPR,
},
patch_existing={"expression": _POLICY_EXPR},
)
POLICY_MAPPING_PK = policy_mapping["pk"]
# Fetch the standard openid/email/profile mapping pks (shipped by default).
status, res = api("GET", "/api/v3/propertymappings/provider/scope/")
by_scope = {m["scope_name"]: m["pk"] for m in res["results"]}
SCOPE_PKS = [by_scope["openid"], by_scope["email"], by_scope["profile"], GROUPS_MAPPING_PK, PERMISSIONS_MAPPING_PK]
status, res = api("GET", "/api/v3/flows/instances/?slug=default-provider-authorization-implicit-consent")
AUTHORIZATION_FLOW_PK = res["results"][0]["pk"]
status, res = api("GET", "/api/v3/flows/instances/?slug=default-provider-invalidation-flow")
INVALIDATION_FLOW_PK = res["results"][0]["pk"]
status, res = api("GET", "/api/v3/crypto/certificatekeypairs/?has_key=true")
SIGNING_KEY_PK = res["results"][0]["pk"]
# -----------------------------------------------------------------------------
print("[2/5] Ensuring homelab-admins + per-service admin groups exist...")
homelab_admins = get_or_create(
"/api/v3/core/groups/", "/api/v3/core/groups/",
"name=homelab-admins",
{"name": "homelab-admins", "is_superuser": True},
)
# App-scoped, not Authentik superusers (unlike homelab-admins) - each maps to
# read+write in its own service via the "permissions" claim above (k8s Role/
# RoleBinding in k8s/infra/rbac/, or an app's own adapter e.g. paperless's).
# k8s-devops-admin is declared with no target yet - foundation for a future
# short-lived federated-operator credential.
SERVICE_ADMIN_GROUP_NAMES = [
"grafana-admins", "minio-admins", "forgejo-admins", "homarr-admins",
"portainer-admins", "kmsvc-admins", "temporal-admins", "llm-admins",
"paperless-admins", "immich-admins", "poimen-memory-admins",
"k8s-devops-admin",
# Not a human-admin group like the others - Vault Identity Group aliasing
# target for service/API (non-browser) access to Vault, kept separate from
# homelab-admins' blanket "*" grant. See k8s/infra/iam/scripts/vault-provision.sh.
"vault-service-api",
]
service_admin_groups = {}
for group_name in SERVICE_ADMIN_GROUP_NAMES:
service_admin_groups[group_name] = get_or_create(
"/api/v3/core/groups/", "/api/v3/core/groups/",
f"name={group_name}",
{"name": group_name, "is_superuser": False},
)
grafana_admins = service_admin_groups["grafana-admins"]
paperless_admins = service_admin_groups["paperless-admins"]
# -----------------------------------------------------------------------------
print("[3/5] Ensuring user 'rock' exists with admin group membership...")
status, res = api("GET", "/api/v3/core/users/?username=rock")
rock_password = None
if res.get("results"):
rock = res["results"][0]
status, rock = api("PATCH", f"/api/v3/core/users/{rock['pk']}/", {
"groups": [homelab_admins["pk"]] + [g["pk"] for g in service_admin_groups.values()],
"is_active": True,
# email is REQUIRED: Grafana's OIDC login reads the email claim from
# userinfo; an empty email makes Grafana fall back to a GitHub-style
# <userinfo>/emails call, which Authentik 404s -> login fails entirely.
"email": "[email protected]",
})
if status not in (200, 201):
die(f"PATCH user rock -> {status} {rock}")
print(" rock already exists, group membership synced (password unchanged)")
else:
rock_password = gen_secret(24)
status, rock = api("POST", "/api/v3/core/users/", {
"username": "rock",
"name": "Rock",
"is_active": True,
# Required for Grafana OIDC (see PATCH branch above).
"email": "[email protected]",
"groups": [homelab_admins["pk"]] + [g["pk"] for g in service_admin_groups.values()],
"path": "users",
"type": "internal",
})
if status not in (200, 201):
die(f"POST user rock -> {status} {rock}")
status, pw_res = api("POST", f"/api/v3/core/users/{rock['pk']}/set_password/",
{"password": rock_password})
if status not in (200, 204):
die(f"set_password for rock -> {status} {pw_res}")
kubectl_create_secret("iam", "rock-credentials", {
"username": "rock",
"password": rock_password,
})
print(" rock created, credentials stored in iam/rock-credentials")
# -----------------------------------------------------------------------------
print("[4/5] Ensuring OAuth2 providers + applications for grafana/minio/forgejo/argocd...")
SERVICES = {
"grafana": {
"client_secret_source": ("logging", "grafana-oidc", "GF_AUTH_GENERIC_OAUTH_CLIENT_SECRET"),
"redirect_uris": ["https://grafana.riotpiao.com/login/generic_oauth"],
"launch_url": "https://grafana.riotpiao.com",
"display_name": "Grafana",
},
"minio": {
"client_secret_source": ("storage", "minio-oidc", "MINIO_IDENTITY_OPENID_CLIENT_SECRET"),
"redirect_uris": ["https://minio.riotpiao.com/oauth_callback"],
"launch_url": "https://minio.riotpiao.com",
"display_name": "MinIO",
},
"forgejo": {
# No secret exists yet for forgejo - generate + store on first run.
"client_secret_source": ("cicd", "forgejo-oidc", "CLIENT_SECRET"),
"generate_if_missing": True,
"redirect_uris": [
"https://forgejo.riotpiao.com/user/oauth2/authentik/callback",
"https://forgejo.riotpiao.com/user/oauth2/openidconnect/callback",
],
"launch_url": "https://forgejo.riotpiao.com",
"display_name": "Forgejo",
},
"argocd": {
# oidc-secret uses hyphenated keys (client-id/client-secret) per
# argocd-values.yaml's `$oidc-secret:client-id` / `:client-secret` refs.
"client_secret_source": ("argocd", "oidc-secret", "client-secret"),
"generate_if_missing": True,
"extra_secret_literals": {"client-id": "argocd"},
# argocd only reads $secret refs from Secrets labelled part-of: argocd.
"secret_labels": {"app.kubernetes.io/part-of": "argocd"},
"redirect_uris": ["https://argocd.riotpiao.com/auth/callback"],
"launch_url": "https://argocd.riotpiao.com",
"display_name": "Argo CD",
},
"homarr": {
"client_secret_source": ("dashboard", "homarr-oidc", "client-secret"),
"generate_if_missing": True,
"extra_secret_literals": {"client-id": "homarr"},
"redirect_uris": ["https://homarr.riotpiao.com/api/auth/callback/oidc"],
"launch_url": "https://homarr.riotpiao.com",
"display_name": "Homarr",
},
"paperless": {
# No secret exists yet for paperless - generate + store on first run.
# django-allauth's generic openid_connect provider callback path is
# /accounts/oidc/<provider_id>/login/callback/ - provider_id "authentik"
# is set in PAPERLESS_SOCIALACCOUNT_PROVIDERS (see configmap.yaml).
"client_secret_source": ("paperless", "paperless-oidc", "CLIENT_SECRET"),
"generate_if_missing": True,
"redirect_uris": ["https://paperless.riotpiao.com/accounts/oidc/authentik/login/callback/"],
"launch_url": "https://paperless.riotpiao.com",
"display_name": "Paperless-ngx",
},
"immich": {
# No secret exists yet for immich - generate + store on first run.
"client_secret_source": ("immich", "immich-oidc", "CLIENT_SECRET"),
"generate_if_missing": True,
# /auth/login + /user-settings are Immich's own web callback routes;
# /api/oauth/mobile-redirect forwards to the app.immich:///oauth-callback
# custom scheme Authentik can't register directly (see docs.immich.app/
# administration/oauth - "custom scheme" workaround).
"redirect_uris": [
"https://img.riotpiao.com/auth/login",
"https://img.riotpiao.com/user-settings",
"https://img.riotpiao.com/api/oauth/mobile-redirect",
],
"launch_url": "https://img.riotpiao.com",
"display_name": "Immich",
},
"vault": {
# Human/CLI login only (`vault login -method=oidc`) - not wired to any
# workload. No secret exists yet - generate + store on first run.
# localhost:8250/oidc/callback is the vault CLI's documented fixed
# callback port for `vault login -method=oidc`; the other is the
# browser/UI flow's callback path (mount path "oidc").
"client_secret_source": ("iam", "vault-oidc", "CLIENT_SECRET"),
"generate_if_missing": True,
"extra_secret_literals": {"client-id": "vault"},
"redirect_uris": [
"https://vault.riotpiao.com/ui/vault/auth/oidc/oidc/callback",
"http://localhost:8250/oidc/callback",
],
"launch_url": "https://vault.riotpiao.com",
"display_name": "Vault",
},
"poimen-memory": {
# Service-to-service API auth (no browser redirect) - generate secret on first run.
"client_secret_source": ("poimen", "poimen-memory-oidc", "CLIENT_SECRET"),
"generate_if_missing": True,
"extra_secret_literals": {"client-id": "poimen-memory"},
"redirect_uris": [], # No browser flow, service-to-service only
"launch_url": "https://memory.riotpiao.com",
"display_name": "Poimen Memory",
},
"local-llm": {
# JWT auth for local LLM API access - service-to-service, no browser flow.
# Client validates JWT tokens issued by this provider using the public key.
"client_secret_source": ("llm-serving", "local-llm-jwt", "client-secret"),
"generate_if_missing": True,
"extra_secret_literals": {"client-id": "local-llm"},
"redirect_uris": [], # No browser flow, JWT/service-to-service only
"launch_url": "https://llm.riotpiao.com",
"display_name": "Local LLM",
},
}
app_pks_for_binding = []
for name, cfg in SERVICES.items():
# MinIO also needs the "policy" claim (via the minio scope mapping) so its
# MINIO_IDENTITY_OPENID_CLAIM_NAME=policy maps homelab-admins -> consoleAdmin.
# Immich needs "immich_role" so its OAuth roleClaim can grant admin.
provider_mappings = SCOPE_PKS + ([POLICY_MAPPING_PK] if name == "minio" else []) \
+ ([IMMICH_ROLE_MAPPING_PK] if name == "immich" else [])
ns, secret_name, key = cfg["client_secret_source"]
client_secret = kubectl_get_secret_key(ns, secret_name, key)
if client_secret is None:
if not cfg.get("generate_if_missing"):
print(f" WARNING: {ns}/{secret_name} key {key} not found and "
f"generate_if_missing not set for '{name}' - skipping provider/app")
continue
client_secret = gen_secret(40)
literals = {key: client_secret}
literals.update(cfg.get("extra_secret_literals", {}))
kubectl_create_secret(ns, secret_name, literals,
labels=cfg.get("secret_labels"))
print(f" {name}: generated new client secret -> {ns}/{secret_name}")
else:
print(f" {name}: using existing client secret from {ns}/{secret_name}")
if name == "paperless":
# paperless-ngx's django-allauth OIDC config takes client_id/secret
# bundled inside one JSON blob (PAPERLESS_SOCIALACCOUNT_PROVIDERS), not
# discrete env vars - compose it here and store it alongside
# CLIENT_SECRET so the Deployment can source it directly via
# secretKeyRef, no shell wrapper needed. Runs every time (not just on
# generate), so it stays in sync if the client_secret is ever rotated
# by hand.
providers_json = json.dumps({
"openid_connect": {
"APPS": [{
"provider_id": "authentik",
"name": "Authentik",
"client_id": "paperless",
"secret": client_secret,
"settings": {
"server_url": "https://authentik.riotpiao.com/application/o/paperless/.well-known/openid-configuration",
# "groups"/"permissions" aren't default OIDC scopes -
# must be requested explicitly for Authentik's scope
# mappings above to actually be returned. paperless's
# adapter.py ConfigMap reads the "permissions" claim
# to grant is_staff+is_superuser.
"scope": ["openid", "profile", "email", "groups", "permissions"],
},
}],
},
})
kubectl_create_secret("paperless", "paperless-oidc", {
"CLIENT_SECRET": client_secret,
"SOCIALACCOUNT_PROVIDERS_JSON": providers_json,
})
if name == "immich":
# Immich reads its whole system-config from IMMICH_CONFIG_FILE (a
# mounted JSON file, see k8s/apps/immich/deployment.yaml), not
# discrete env vars. "immich_role" must be in `scope` for Authentik
# to actually include that claim in the token (non-default scopes
# are opt-in per-client, same reason paperless requests "permissions"
# explicitly). roleClaim is re-evaluated on every login (immich-app/
# immich#29991) so this is the actual admin-grant mechanism for rock,
# not Immich's racy first-user-is-admin fallback.
immich_config_json = json.dumps({
"oauth": {
"enabled": True,
"issuerUrl": "https://authentik.riotpiao.com/application/o/immich/",
"clientId": "immich",
"clientSecret": client_secret,
"scope": "openid email profile immich_role",
"roleClaim": "immich_role",
"autoRegister": True,
"autoLaunch": False,
"buttonText": "Login with Authentik",
"mobileRedirectUri": "app.immich:///oauth-callback",
},
})
kubectl_create_secret("immich", "immich-oidc", {
"CLIENT_SECRET": client_secret,
"config.json": immich_config_json,
})
# Service-to-service (client_credentials): poimen-memory
# Browser SSO (authorization_code): all others
grant_types = [
"urn:ietf:params:oauth:grant-type:device_code", # device code flow (CLI/headless)
"client_credentials" # service-to-service
] if name == "poimen-memory" else [
"authorization_code", # web SSO
"refresh_token" # long-lived sessions
]
provider = get_or_create(
"/api/v3/providers/oauth2/", "/api/v3/providers/oauth2/",
f"name={name}",
{
"name": name,
"client_id": name,
"client_secret": client_secret,
"client_type": "confidential",
"authorization_flow": AUTHORIZATION_FLOW_PK,
"invalidation_flow": INVALIDATION_FLOW_PK,
"signing_key": SIGNING_KEY_PK,
"property_mappings": provider_mappings,
"sub_mode": "hashed_user_id",
"include_claims_in_id_token": True,
# authentik 2026.x requires grant_types to be set explicitly; the
# API defaults it to [] when omitted, which makes /authorize reject
# every login with "Invalid grant_type for provider" ->
# invalid_request. authorization_code = the web SSO flow all these
# apps use; refresh_token = long-lived sessions (offline_access).
"grant_types": grant_types,
"redirect_uris": [
{"matching_mode": "strict", "url": u} for u in cfg["redirect_uris"]
],
},
# Keep the redirect_uris/mappings/grant_types in sync on re-run, but
# never touch client_secret again once created (that's the source of
# truth in the k8s Secret, and re-sending it here is harmless anyway).
patch_existing={
"property_mappings": provider_mappings,
"grant_types": grant_types,
"redirect_uris": [
{"matching_mode": "strict", "url": u} for u in cfg["redirect_uris"]
],
},
)
# superuser_full_list=true is REQUIRED on the LIST: the applications list
# applies access-policy filtering to the results array (these apps are bound
# to homelab-admins, and the bootstrap-token user akadmin is not a member),
# so without it the GET returns an empty results list even though the app
# exists -> fall through to POST -> 400 "already exists".
#
# We deliberately do NOT patch_existing here: the application DETAIL endpoint
# (PATCH /applications/{pk}/) enforces the same access policy and does NOT
# honor superuser_full_list, so PATCH-by-pk returns 404 for akadmin once the
# homelab-admins binding exists. That 404 aborted the loop before later
# providers got their grant_types. slug/provider/launch_url are set at
# creation and are stable (provider is get_or_create'd by name, stable pk),
# so find-or-create is sufficient.
application = get_or_create(
"/api/v3/core/applications/", "/api/v3/core/applications/",
f"slug={name}&superuser_full_list=true",
{
"name": cfg["display_name"],
"slug": name,
"provider": provider["pk"],
"meta_launch_url": cfg["launch_url"],
},
)
app_pks_for_binding.append((name, application["pk"]))
print(f" {name}: provider pk={provider['pk']} application pk={application['pk']}")
# -----------------------------------------------------------------------------
# Separate from the SERVICES loop above: this is a PUBLIC client (PKCE, no
# client_secret) for `kubectl` OIDC login, not a confidential-client app
# login. Foundation for k8s/infra/rbac/ - kube-apiserver's --oidc-* flags
# (controlplane.tftpl) validate tokens issued against this provider.
# Redirect URI matches kubelogin's (int128/kubelogin) documented default;
# adjust here if a different kubectl OIDC plugin/port is actually used.
print("Ensuring public OAuth2 client 'kubernetes' for kubectl OIDC login...")
k8s_provider = get_or_create(
"/api/v3/providers/oauth2/", "/api/v3/providers/oauth2/",
"name=kubernetes",
{
"name": "kubernetes",
"client_id": "kubernetes",
"client_type": "public",
"authorization_flow": AUTHORIZATION_FLOW_PK,
"invalidation_flow": INVALIDATION_FLOW_PK,
"signing_key": SIGNING_KEY_PK,
"property_mappings": SCOPE_PKS,
"sub_mode": "hashed_user_id",
"include_claims_in_id_token": True,
"grant_types": ["authorization_code", "refresh_token"],
"redirect_uris": [
{"matching_mode": "strict", "url": "http://localhost:8000"},
],
},
patch_existing={
"property_mappings": SCOPE_PKS,
"grant_types": ["authorization_code", "refresh_token"],
"redirect_uris": [
{"matching_mode": "strict", "url": "http://localhost:8000"},
],
},
)
k8s_application = get_or_create(
"/api/v3/core/applications/", "/api/v3/core/applications/",
"slug=kubernetes&superuser_full_list=true",
{
"name": "Kubernetes",
"slug": "kubernetes",
"provider": k8s_provider["pk"],
"meta_launch_url": "https://authentik.riotpiao.com",
},
)
app_pks_for_binding.append(("kubernetes", k8s_application["pk"]))
print(f" kubernetes: provider pk={k8s_provider['pk']} application pk={k8s_application['pk']}")
# -----------------------------------------------------------------------------
print("[5/5] Binding homelab-admins to every application (guaranteed access for rock)...")
for name, app_pk in app_pks_for_binding:
get_or_create(
"/api/v3/policies/bindings/", "/api/v3/policies/bindings/",
f"target={app_pk}&group={homelab_admins['pk']}",
{
"target": app_pk,
"group": homelab_admins["pk"],
"order": 0,
"enabled": True,
},
)
print(f" {name}: homelab-admins bound")
# Per-service admin groups are app-scoped (unlike homelab-admins' blanket
# binding above) - only grants visibility/access to that one application.
# portainer/kmsvc/temporal have no Authentik Application (no OIDC
# login integration), so their groups exist for the "permissions" claim /
# future k8s RBAC only - nothing to bind here.
SERVICE_GROUP_TO_APP_SLUG = {
"grafana-admins": "grafana",
"minio-admins": "minio",
"forgejo-admins": "forgejo",
"homarr-admins": "homarr",
"paperless-admins": "paperless",
"immich-admins": "immich",
"llm-admins": "local-llm",
}
for group_name, app_slug in SERVICE_GROUP_TO_APP_SLUG.items():
app_pk = next((pk for n, pk in app_pks_for_binding if n == app_slug), None)
if not app_pk:
continue
group_pk = service_admin_groups[group_name]["pk"]
get_or_create(
"/api/v3/policies/bindings/", "/api/v3/policies/bindings/",
f"target={app_pk}&group={group_pk}",
{
"target": app_pk,
"group": group_pk,
"order": 0,
"enabled": True,
},
)
print(f" {app_slug}: {group_name} bound")
# JWT configuration for local-llm
print("\n[JWT] Fetching Authentik signing key for local-llm...")
status, signing_key_res = api("GET", f"/api/v3/crypto/certificatekeypairs/{SIGNING_KEY_PK}/")
if status == 200:
jwt_cert = signing_key_res.get("certificate", "")
print(f" Public certificate available for JWT validation (base64-encoded below)\n")
import base64
cert_b64 = base64.b64encode(jwt_cert.encode()).decode()
print(f"Save this to local-llm config for JWT token validation:")
print(f" AUTHENTIK_JWT_CERT={cert_b64}")
print(f"\nJWT issuer URL: https://authentik.riotpiao.com/application/o/local-llm/")
print(f"Local-LLM credentials are stored in: kubectl -n llm-serving get secret local-llm-jwt")
print("\nDone. Summary:")
print(" groups: homelab-admins (superuser) + " + ", ".join(SERVICE_ADMIN_GROUP_NAMES))
print(" user: rock -> homelab-admins + all service admin groups")
print(f" apps: {', '.join(n for n, _ in app_pks_for_binding)}")
if rock_password:
print(" NOTE: rock's password was generated this run - see")
print(" kubectl -n iam get secret rock-credentials -o jsonpath='{.data.password}' | base64 -d")
+5 -5
View File
@@ -68,15 +68,14 @@ grafana.ini:
# doesn't return localhost redirects in its token responses. # doesn't return localhost redirects in its token responses.
# #
# role_attribute_path: JMESPath expression evaluated against the userinfo # role_attribute_path: JMESPath expression evaluated against the userinfo
# response. Members of the 'grafana-admins' Authentik group get Admin role; # response. akadmin gets GrafanaAdmin (server admin, can impersonate);
# everyone else gets Viewer. The group name must match exactly what Authentik # homelab-admins members get Admin (org admin); everyone else Viewer.
# sends in the 'groups' claim.
auth.generic_oauth: auth.generic_oauth:
enabled: true enabled: true
name: Authentik name: Authentik
allow_sign_up: true allow_sign_up: true
client_id: grafana client_id: grafana
scopes: openid email profile scopes: openid email profile groups
auth_url: https://authentik.riotpiao.com/application/o/authorize/ auth_url: https://authentik.riotpiao.com/application/o/authorize/
token_url: https://authentik.riotpiao.com/application/o/token/ token_url: https://authentik.riotpiao.com/application/o/token/
api_url: https://authentik.riotpiao.com/application/o/userinfo/ api_url: https://authentik.riotpiao.com/application/o/userinfo/
@@ -87,7 +86,8 @@ grafana.ini:
email_attribute_path: email email_attribute_path: email
login_attribute_path: preferred_username login_attribute_path: preferred_username
name_attribute_path: name name_attribute_path: name
role_attribute_path: "contains(groups[*], 'grafana-admins') && 'Admin' || 'Viewer'" role_attribute_path: "preferred_username == 'akadmin' && 'GrafanaAdmin' || contains(groups[*], 'homelab-admins') && 'Admin' || 'Viewer'"
allow_assign_grafana_admin: true
use_pkce: false use_pkce: false
use_refresh_token: false use_refresh_token: false
skip_org_role_sync: false skip_org_role_sync: false
+5
View File
@@ -96,8 +96,13 @@ spec:
console: https://minio.riotpiao.com console: https://minio.riotpiao.com
# ── OIDC via Authentik (server-side env, valid in v2 schema) ──────────────── # ── OIDC via Authentik (server-side env, valid in v2 schema) ────────────────
# Use in-cluster URL for config fetch (pod→authentik); browser redirects use
# public URLs embedded in the OIDC metadata response (issuer stays public).
env: env:
- name: MINIO_IDENTITY_OPENID_CONFIG_URL - name: MINIO_IDENTITY_OPENID_CONFIG_URL
# Must use external URL — well-known response contains external issuer/jwks_uri.
# MinIO validates issuer in JWT matches well-known issuer. Internal URL = mismatch.
# Hairpins through ingress-nginx but stays in-cluster.
value: "https://authentik.riotpiao.com/application/o/minio/.well-known/openid-configuration" value: "https://authentik.riotpiao.com/application/o/minio/.well-known/openid-configuration"
- name: MINIO_IDENTITY_OPENID_CLIENT_ID - name: MINIO_IDENTITY_OPENID_CLIENT_ID
value: "minio" value: "minio"
@@ -0,0 +1,160 @@
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: api-gateway-alerts
namespace: monitoring
labels:
release: prometheus
spec:
groups:
# ================================================================
# SLA Targets (based on canary traffic baselines):
#
# Availability: 99.9% (43.8 min downtime/month)
# LLM Chat: p95 < 1s (qwen), p95 < 2s (reasoning), p95 < 5s (ornith)
# Embeddings: p95 < 500ms
# Rerank: p95 < 500ms
# Models list: p95 < 300ms
# Error rate: < 1% (5xx), < 5% (4xx excluding auth)
#
# Baselines from 200-request canary run:
# qwen p99=609ms, reasoning p99=328ms, embeddings p99=287ms,
# rerank p99=218ms, models p99=277ms
# SLA set at ~2x p99 for headroom.
# ================================================================
- name: api-gateway.availability
rules:
# Gateway pods not ready
- alert: APIGatewayDown
expr: sum(kube_pod_status_ready{namespace="api",condition="true"}) == 0
for: 1m
labels:
severity: critical
annotations:
summary: "API Gateway has zero ready pods"
# Gateway pod count below desired
- alert: APIGatewayDegraded
expr: |
sum(kube_pod_status_ready{namespace="api",condition="true"})
< kube_deployment_spec_replicas{namespace="api",deployment="api-gateway"}
for: 5m
labels:
severity: warning
annotations:
summary: "API Gateway {{ $value }} ready pods below desired replica count"
# Blackbox probe down
- alert: APIGatewayProbeDown
expr: probe_success{instance=~".*api.riotpiao.com.*"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: "API Gateway probe failed: {{ $labels.instance }}"
# LLM serving pods not ready
- alert: LLMServingDown
expr: sum(kube_pod_status_ready{namespace="llm-serving",condition="true"}) == 0
for: 2m
labels:
severity: critical
annotations:
summary: "All LLM serving pods down"
# Individual predictor down
- alert: LLMPredictorDown
expr: |
kube_deployment_status_replicas_ready{namespace="llm-serving"}
< kube_deployment_spec_replicas{namespace="llm-serving"}
for: 5m
labels:
severity: warning
annotations:
summary: "{{ $labels.deployment }} has {{ $value }} ready (below desired)"
- name: api-gateway.latency
# SLA: latency thresholds at ~2x measured p99
rules:
# Ingress-level latency (all requests through nginx)
- alert: APIGatewayLatencyHigh
expr: |
histogram_quantile(0.95,
sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress="api"}[5m])) by (le)
) > 2
for: 5m
labels:
severity: warning
annotations:
summary: "API Gateway p95 latency {{ $value | printf \"%.1f\" }}s (SLA: <2s)"
# Extreme latency (p99 > 5s)
- alert: APIGatewayLatencyCritical
expr: |
histogram_quantile(0.99,
sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress="api"}[5m])) by (le)
) > 5
for: 5m
labels:
severity: critical
annotations:
summary: "API Gateway p99 latency {{ $value | printf \"%.1f\" }}s (SLA: <5s)"
- name: api-gateway.errors
rules:
# 5xx error rate > 1%
- alert: APIGateway5xxErrorRate
expr: |
sum(rate(nginx_ingress_controller_requests{ingress="api",status=~"5.."}[5m]))
/ sum(rate(nginx_ingress_controller_requests{ingress="api"}[5m]))
> 0.01
for: 5m
labels:
severity: critical
annotations:
summary: "API Gateway 5xx rate {{ $value | humanizePercentage }} (SLA: <1%)"
# Total error rate > 10% (including 4xx)
- alert: APIGatewayHighErrorRate
expr: |
sum(rate(nginx_ingress_controller_requests{ingress="api",status=~"[45].."}[5m]))
/ sum(rate(nginx_ingress_controller_requests{ingress="api"}[5m]))
> 0.10
for: 10m
labels:
severity: warning
annotations:
summary: "API Gateway total error rate {{ $value | humanizePercentage }} (SLA: <10%)"
- name: api-gateway.resources
rules:
# Gateway pod restart
- alert: APIGatewayRestarted
expr: increase(kube_pod_container_status_restarts_total{namespace="api"}[15m]) > 0
for: 0m
labels:
severity: warning
annotations:
summary: "API Gateway pod {{ $labels.pod }} restarted"
# LLM predictor restart
- alert: LLMPredictorRestarted
expr: increase(kube_pod_container_status_restarts_total{namespace="llm-serving"}[15m]) > 0
for: 0m
labels:
severity: warning
annotations:
summary: "LLM predictor {{ $labels.pod }} restarted"
# Gateway high memory (>80% of limit)
- alert: APIGatewayHighMemory
expr: |
sum(container_memory_working_set_bytes{namespace="api",container="gateway"}) by (pod)
/ sum(kube_pod_container_resource_limits{namespace="api",container="gateway",resource="memory"}) by (pod)
> 0.8
for: 10m
labels:
severity: warning
annotations:
summary: "Gateway pod {{ $labels.pod }} memory at {{ $value | humanizePercentage }} of limit"
@@ -0,0 +1,177 @@
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: cluster-alerts
namespace: monitoring
labels:
release: prometheus
spec:
groups:
- name: cluster.availability
rules:
# Node down
- alert: NodeNotReady
expr: kube_node_status_condition{condition="Ready",status="true"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: "Node {{ $labels.node }} is NotReady"
# Pod stuck pending (scheduling failure)
- alert: PodStuckPending
expr: sum(kube_pod_status_phase{phase="Pending"}) > 0
for: 10m
labels:
severity: warning
annotations:
summary: "{{ $value }} pod(s) stuck in Pending state for >10m"
# CrashLoopBackOff
- alert: PodCrashLooping
expr: sum(kube_pod_container_status_waiting_reason{reason="CrashLoopBackOff"}) by (namespace, pod) > 0
for: 5m
labels:
severity: critical
annotations:
summary: "{{ $labels.namespace }}/{{ $labels.pod }} in CrashLoopBackOff"
# OOMKilled spike
- alert: OOMKilledSpike
expr: sum(increase(kube_pod_container_status_last_terminated_reason{reason="OOMKilled"}[1h])) > 3
for: 0m
labels:
severity: warning
annotations:
summary: "{{ $value }} OOMKilled events in last hour"
# Deployment replicas unavailable
- alert: DeploymentReplicasUnavailable
expr: kube_deployment_status_replicas_unavailable > 0
for: 10m
labels:
severity: warning
annotations:
summary: "{{ $labels.namespace }}/{{ $labels.deployment }} has {{ $value }} unavailable replicas"
- name: cluster.jobs
rules:
# Job failed
- alert: JobFailed
expr: kube_job_status_failed > 0
for: 5m
labels:
severity: warning
annotations:
summary: "Job {{ $labels.namespace }}/{{ $labels.job_name }} failed"
# Job stuck running >2h
- alert: JobStuckRunning
expr: |
kube_job_status_active == 1
and on(job_name,namespace)
(time() - kube_job_status_start_time) > 7200
for: 0m
labels:
severity: warning
annotations:
summary: "Job {{ $labels.namespace }}/{{ $labels.job_name }} running >2h"
# CronJob missed schedule
- alert: CronJobMissedSchedule
expr: |
(time() - kube_cronjob_status_last_schedule_time) > 2 * (kube_cronjob_spec_next_schedule_time - kube_cronjob_status_last_schedule_time)
for: 10m
labels:
severity: warning
annotations:
summary: "CronJob {{ $labels.namespace }}/{{ $labels.cronjob }} missed schedule"
- name: cluster.resources
rules:
# Node CPU >90% sustained
- alert: NodeHighCPU
expr: (1 - avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) by (instance)) * 100 > 90
for: 15m
labels:
severity: warning
annotations:
summary: "Node {{ $labels.instance }} CPU at {{ $value | printf \"%.0f\" }}%"
# Node memory >90% sustained
- alert: NodeHighMemory
expr: (1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100 > 90
for: 15m
labels:
severity: warning
annotations:
summary: "Node {{ $labels.instance }} memory at {{ $value | printf \"%.0f\" }}%"
# Node disk >85%
- alert: NodeDiskFull
expr: (1 - node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100 > 85
for: 5m
labels:
severity: critical
annotations:
summary: "Node {{ $labels.instance }} disk at {{ $value | printf \"%.0f\" }}%"
# Container restart storm (>5 restarts in 15m)
- alert: ContainerRestartStorm
expr: sum(increase(kube_pod_container_status_restarts_total[15m])) by (namespace, pod) > 5
for: 0m
labels:
severity: warning
annotations:
summary: "{{ $labels.namespace }}/{{ $labels.pod }} restarted {{ $value | printf \"%.0f\" }} times in 15m"
- name: cluster.storage
rules:
# Longhorn drive offline
- alert: LonghornDriveOffline
expr: longhorn_disk_health != 1
for: 5m
labels:
severity: critical
annotations:
summary: "Longhorn disk {{ $labels.node }} unhealthy"
- name: cluster.dns
rules:
# CoreDNS errors spike
- alert: CoreDNSErrorSpike
expr: sum(rate(coredns_dns_responses_total{rcode=~"SERVFAIL"}[5m])) > 0.5
for: 5m
labels:
severity: warning
annotations:
summary: "CoreDNS SERVFAIL rate {{ $value | printf \"%.2f\" }}/s"
- name: cluster.probes
rules:
# Any blackbox probe down
- alert: ServiceProbeDown
expr: probe_success == 0
for: 3m
labels:
severity: critical
annotations:
summary: "Probe failed: {{ $labels.instance }}"
# Probe latency >2s
- alert: ServiceProbeSlow
expr: probe_duration_seconds > 2
for: 5m
labels:
severity: warning
annotations:
summary: "Probe slow ({{ $value | printf \"%.1f\" }}s): {{ $labels.instance }}"
# Certificate expiry <14 days
- alert: CertificateExpiringSoon
expr: (certmanager_certificate_expiration_timestamp_seconds - time()) / 86400 < 14
for: 0m
labels:
severity: warning
annotations:
summary: "Certificate {{ $labels.name }} expires in {{ $value | printf \"%.0f\" }} days"
@@ -61,6 +61,10 @@ serviceMonitor:
url: https://argocd.riotpiao.com/healthz url: https://argocd.riotpiao.com/healthz
- name: longhorn - name: longhorn
url: https://longhorn.riotpiao.com/ url: https://longhorn.riotpiao.com/
- name: api-gateway
url: https://api.riotpiao.com/healthz
- name: api-gateway-models
url: https://api.riotpiao.com/v1/models
prometheusRule: prometheusRule:
enabled: true enabled: true
@@ -0,0 +1,36 @@
apiVersion: v1
data:
api-gateway.json: '{"title":"API Gateway","uid":"api-gateway","schemaVersion":39,"timezone":"browser","time":{"from":"now-6h","to":"now"},"refresh":"30s","tags":["api","gateway","llm"],"panels":[{"id":1,"title":"Gateway
Health","type":"row","collapsed":false,"gridPos":{"h":1,"w":24,"x":0,"y":0},"panels":[{"id":2,"title":"Gateway
Pods Ready","type":"stat","gridPos":{"h":4,"w":4,"x":0,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"red"},{"value":3,"color":"green"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"sum(kube_pod_status_ready{namespace=\"api\",condition=\"true\"})"}]},{"id":3,"title":"Probe:
healthz","type":"stat","gridPos":{"h":4,"w":4,"x":4,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","mappings":[{"type":"value","options":{"0":{"text":"DOWN","color":"red"},"1":{"text":"UP","color":"green"}}}]}},"targets":[{"expr":"probe_success{instance=~\".*api.riotpiao.com/healthz\"}"}]},{"id":4,"title":"Probe
Latency","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"s"}},"targets":[{"expr":"probe_duration_seconds{instance=~\".*api.riotpiao.com.*\"}","legendFormat":"{{instance}}"}]}]},{"id":10,"title":"Ingress
Traffic (nginx)","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":1},"panels":[{"id":11,"title":"Request
Rate by Status","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(nginx_ingress_controller_requests{ingress=\"api\"}[5m]))
by (status)","legendFormat":"{{status}}"}]},{"id":12,"title":"Error Rate %","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"percent"}},"targets":[{"expr":"sum(rate(nginx_ingress_controller_requests{ingress=\"api\",status=~\"5..\"}[5m]))
/ sum(rate(nginx_ingress_controller_requests{ingress=\"api\"}[5m])) * 100","legendFormat":"5xx"},{"expr":"sum(rate(nginx_ingress_controller_requests{ingress=\"api\",status=~\"4..\"}[5m]))
/ sum(rate(nginx_ingress_controller_requests{ingress=\"api\"}[5m])) * 100","legendFormat":"4xx"}]},{"id":13,"title":"Latency
p50/p95/p99","type":"timeseries","gridPos":{"h":8,"w":8,"x":16,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"s"}},"targets":[{"expr":"histogram_quantile(0.50,
sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress=\"api\"}[5m]))
by (le))","legendFormat":"p50"},{"expr":"histogram_quantile(0.95, sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress=\"api\"}[5m]))
by (le))","legendFormat":"p95"},{"expr":"histogram_quantile(0.99, sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress=\"api\"}[5m]))
by (le))","legendFormat":"p99"}]}]},{"id":20,"title":"LLM Serving","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":2},"panels":[{"id":21,"title":"LLM
Pods Ready","type":"stat","gridPos":{"h":4,"w":4,"x":0,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(kube_pod_status_ready{namespace=\"llm-serving\",condition=\"true\"})"}]},{"id":22,"title":"CPU
by Predictor","type":"timeseries","gridPos":{"h":8,"w":8,"x":4,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(container_cpu_usage_seconds_total{namespace=\"llm-serving\"}[5m]))
by (pod)","legendFormat":"{{pod}}"}]},{"id":23,"title":"Memory by Predictor","type":"timeseries","gridPos":{"h":8,"w":8,"x":12,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"bytes"}},"targets":[{"expr":"sum(container_memory_working_set_bytes{namespace=\"llm-serving\"})
by (pod)","legendFormat":"{{pod}}"}]},{"id":24,"title":"Predictor Restarts","type":"timeseries","gridPos":{"h":8,"w":4,"x":20,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(kube_pod_container_status_restarts_total{namespace=\"llm-serving\"}[15m]))
by (pod)","legendFormat":"{{pod}}"}]}]},{"id":30,"title":"Gateway Resources","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":3},"panels":[{"id":31,"title":"CPU
by Gateway Pod","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":4},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(container_cpu_usage_seconds_total{namespace=\"api\"}[5m]))
by (pod)","legendFormat":"{{pod}}"}]},{"id":32,"title":"Memory by Gateway Pod","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":4},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"bytes"}},"targets":[{"expr":"sum(container_memory_working_set_bytes{namespace=\"api\"})
by (pod)","legendFormat":"{{pod}}"}]},{"id":33,"title":"Gateway Restarts","type":"timeseries","gridPos":{"h":8,"w":8,"x":16,"y":4},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(kube_pod_container_status_restarts_total{namespace=\"api\"}[15m]))
by (pod)","legendFormat":"{{pod}}"}]}]},{"id":40,"title":"Logs","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":4},"panels":[{"id":41,"title":"Gateway
Logs","type":"logs","gridPos":{"h":10,"w":24,"x":0,"y":5},"datasource":{"type":"loki","uid":"loki"},"targets":[{"expr":"{namespace=\"api\",container=\"gateway\"}"}]},{"id":42,"title":"LLM
Serving Logs","type":"logs","gridPos":{"h":10,"w":24,"x":0,"y":15},"datasource":{"type":"loki","uid":"loki"},"targets":[{"expr":"{namespace=\"llm-serving\"}"}]}]}]}'
kind: ConfigMap
metadata:
annotations:
grafana_folder: API
labels:
grafana_dashboard: '1'
name: api-gateway-dashboard
namespace: logging
@@ -0,0 +1,62 @@
apiVersion: v1
data:
cluster-infrastructure.json: '{"title":"Cluster Infrastructure","uid":"cluster-infra","schemaVersion":39,"timezone":"browser","time":{"from":"now-6h","to":"now"},"refresh":"30s","tags":["infrastructure","k8s"],"panels":[{"id":1,"title":"Cluster
Health","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":0},"panels":[{"id":2,"title":"Nodes
Ready","type":"stat","gridPos":{"h":4,"w":4,"x":0,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"red"},{"value":3,"color":"green"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"count(kube_node_status_condition{condition=\"Ready\",status=\"true\"}
== 1)"}]},{"id":3,"title":"Pods Pending","type":"stat","gridPos":{"h":4,"w":4,"x":4,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"green"},{"value":1,"color":"red"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"sum(kube_pod_status_phase{phase=\"Pending\"})
OR on() vector(0)"}]},{"id":4,"title":"CrashLoopBackOff","type":"stat","gridPos":{"h":4,"w":4,"x":8,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"green"},{"value":1,"color":"red"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"sum(kube_pod_container_status_waiting_reason{reason=\"CrashLoopBackOff\"})
OR on() vector(0)"}]},{"id":5,"title":"OOMKilled (1h)","type":"stat","gridPos":{"h":4,"w":4,"x":12,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"green"},{"value":1,"color":"red"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"sum(increase(kube_pod_container_status_last_terminated_reason{reason=\"OOMKilled\"}[1h]))
OR on() vector(0)"}]},{"id":6,"title":"Deploys Unavailable","type":"stat","gridPos":{"h":4,"w":4,"x":16,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"green"},{"value":1,"color":"red"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"count(kube_deployment_status_replicas_unavailable
> 0) OR on() vector(0)"}]},{"id":7,"title":"Services Down","type":"stat","gridPos":{"h":4,"w":4,"x":20,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"green"},{"value":1,"color":"red"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"count(probe_success
== 0) OR on() vector(0)"}]}]},{"id":10,"title":"Jobs & CronJobs","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":1},"panels":[{"id":11,"title":"Failed
Jobs","type":"stat","gridPos":{"h":4,"w":4,"x":0,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"green"},{"value":1,"color":"red"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"count(kube_job_status_failed
> 0) OR on() vector(0)"}]},{"id":12,"title":"Failed Jobs Detail","type":"table","gridPos":{"h":8,"w":10,"x":4,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"kube_job_status_failed
> 0","format":"table","instant":true}]},{"id":13,"title":"Stuck Jobs (>1h)","type":"table","gridPos":{"h":8,"w":10,"x":14,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"kube_job_status_active
== 1 and on(job_name,namespace) (time() - kube_job_status_start_time) > 3600","format":"table","instant":true}]},{"id":14,"title":"CronJob
Last Success","type":"timeseries","gridPos":{"h":8,"w":12,"x":0,"y":10},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"dateTimeFromNow"}},"targets":[{"expr":"kube_cronjob_status_last_successful_time{namespace=~\"cicd|kube-system|paperless\"}","legendFormat":"{{namespace}}/{{cronjob}}"}]},{"id":15,"title":"Container
Restart Storm (top 10)","type":"timeseries","gridPos":{"h":8,"w":12,"x":12,"y":10},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"topk(10,
sum(rate(kube_pod_container_status_restarts_total[15m])) by (namespace, pod))","legendFormat":"{{namespace}}/{{pod}}"}]}]},{"id":20,"title":"Node
Resources","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":2},"panels":[{"id":21,"title":"CPU
% by Node","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"percent"}},"targets":[{"expr":"(1
- avg(rate(node_cpu_seconds_total{mode=\"idle\"}[5m])) by (instance)) * 100","legendFormat":"{{instance}}"}]},{"id":22,"title":"Memory
% by Node","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"percent"}},"targets":[{"expr":"(1
- node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100","legendFormat":"{{instance}}"}]},{"id":23,"title":"Disk
% by Node","type":"timeseries","gridPos":{"h":8,"w":8,"x":16,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"percent"}},"targets":[{"expr":"(1
- node_filesystem_avail_bytes{mountpoint=\"/\"} / node_filesystem_size_bytes{mountpoint=\"/\"})
* 100","legendFormat":"{{instance}}"}]},{"id":24,"title":"Load Average","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":11},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"node_load1","legendFormat":"1m
{{instance}}"},{"expr":"node_load5","legendFormat":"5m {{instance}}"}]},{"id":25,"title":"Network
Errors & Drops","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":11},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"rate(node_network_receive_errs_total[5m])","legendFormat":"rx-err
{{instance}}"},{"expr":"rate(node_network_transmit_errs_total[5m])","legendFormat":"tx-err
{{instance}}"},{"expr":"rate(node_network_receive_drop_total[5m])","legendFormat":"rx-drop
{{instance}}"}]}]},{"id":30,"title":"Control Plane","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":3},"panels":[{"id":31,"title":"API
Server Up","type":"stat","gridPos":{"h":4,"w":4,"x":0,"y":4},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","mappings":[{"type":"value","options":{"0":{"text":"DOWN","color":"red"},"1":{"text":"UP","color":"green"}}}]}},"targets":[{"expr":"min(up{job=\"apiserver\"})"}]},{"id":32,"title":"API
Server Request Rate","type":"timeseries","gridPos":{"h":8,"w":10,"x":4,"y":4},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(apiserver_request_total[5m]))
by (verb, code)","legendFormat":"{{verb}} {{code}}"}]},{"id":33,"title":"API Server
Error Rate %","type":"timeseries","gridPos":{"h":8,"w":10,"x":14,"y":4},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"percent"}},"targets":[{"expr":"sum(rate(apiserver_request_total{code=~\"5..\"}[5m]))
/ sum(rate(apiserver_request_total[5m])) * 100","legendFormat":"5xx %"}]},{"id":34,"title":"API
Server Latency","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":12},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"s"}},"targets":[{"expr":"histogram_quantile(0.95,
sum(rate(apiserver_request_duration_seconds_bucket[5m])) by (le))","legendFormat":"p95"},{"expr":"histogram_quantile(0.99,
sum(rate(apiserver_request_duration_seconds_bucket[5m])) by (le))","legendFormat":"p99"}]},{"id":35,"title":"etcd
Request Duration","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":12},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"s"}},"targets":[{"expr":"histogram_quantile(0.99,
sum(rate(etcd_request_duration_seconds_bucket[5m])) by (le))","legendFormat":"p99"}]}]},{"id":40,"title":"Storage","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":4},"panels":[{"id":41,"title":"Longhorn
Disk Capacity","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":5},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"bytes"}},"targets":[{"expr":"longhorn_disk_capacity_bytes","legendFormat":"capacity
{{node}}"},{"expr":"longhorn_disk_reservation_bytes","legendFormat":"reserved
{{node}}"}]},{"id":42,"title":"PVC Phase","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":5},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"kube_persistentvolumeclaim_status_phase","legendFormat":"{{namespace}}/{{persistentvolumeclaim}}
{{phase}}"}]}]},{"id":50,"title":"DNS & Networking","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":5},"panels":[{"id":51,"title":"CoreDNS
Cache Hit Rate","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":6},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"percentunit"}},"targets":[{"expr":"rate(coredns_cache_hits_total[5m])
/ (rate(coredns_cache_hits_total[5m]) + rate(coredns_cache_misses_total[5m]))","legendFormat":"{{server}}"}]},{"id":52,"title":"CoreDNS
Errors","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":6},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(coredns_dns_responses_total{rcode=~\"SERVFAIL|NXDOMAIN\"}[5m]))
by (rcode)","legendFormat":"{{rcode}}"}]}]},{"id":60,"title":"Logs","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":6},"panels":[{"id":61,"title":"Error
Rate by Namespace","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":7},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum
by (namespace) (count_over_time({namespace=~\"kube-system|cert-manager|ingress-nginx|longhorn-system\"}
|= \"error\" [5m]))","legendFormat":"{{namespace}}"}]},{"id":62,"title":"Control
Plane Logs","type":"logs","gridPos":{"h":10,"w":24,"x":0,"y":15},"datasource":{"type":"loki","uid":"loki"},"targets":[{"expr":"{namespace=\"kube-system\"}"}]},{"id":63,"title":"Cluster
Addon Logs","type":"logs","gridPos":{"h":10,"w":24,"x":0,"y":25},"datasource":{"type":"loki","uid":"loki"},"targets":[{"expr":"{namespace=~\"cert-manager|ingress-nginx|longhorn-system\"}"}]}]}]}'
kind: ConfigMap
metadata:
annotations:
grafana_folder: Infrastructure
labels:
grafana_dashboard: '1'
name: cluster-infrastructure-dashboard
namespace: logging
@@ -1,55 +0,0 @@
# k8s/monitoring/dashboards/control-plane-logs.yaml
# Surfaces controller/control-plane logs that are already in Loki today
# (Promtail scrapes every namespace with no filter) — this dashboard is the
# "make it visible" piece, not new log collection.
apiVersion: v1
kind: ConfigMap
metadata:
name: control-plane-logs-dashboard
namespace: logging
labels:
grafana_dashboard: "1"
data:
control-plane-logs.json: |
{
"title": "Cluster Control Plane & Controllers (Logs)",
"uid": "control-plane-logs",
"schemaVersion": 39,
"timezone": "browser",
"time": { "from": "now-1h", "to": "now" },
"refresh": "30s",
"panels": [
{
"id": 1,
"title": "Error rate by namespace",
"type": "timeseries",
"gridPos": { "h": 6, "w": 24, "x": 0, "y": 0 },
"datasource": { "type": "loki", "uid": "loki" },
"targets": [
{
"expr": "sum by (namespace) (count_over_time({namespace=~\"kube-system|cert-manager|ingress-nginx|longhorn-system\"} |= \"error\" [5m]))"
}
]
},
{
"id": 2,
"title": "Control plane (kube-apiserver, controller-manager, scheduler)",
"type": "logs",
"gridPos": { "h": 10, "w": 24, "x": 0, "y": 6 },
"datasource": { "type": "loki", "uid": "loki" },
"targets": [
{ "expr": "{namespace=\"kube-system\"}" }
]
},
{
"id": 3,
"title": "Cluster add-ons (cert-manager, ingress-nginx, longhorn)",
"type": "logs",
"gridPos": { "h": 10, "w": 24, "x": 0, "y": 16 },
"datasource": { "type": "loki", "uid": "loki" },
"targets": [
{ "expr": "{namespace=~\"cert-manager|ingress-nginx|longhorn-system\"}" }
]
}
]
}
+280
View File
@@ -0,0 +1,280 @@
#!/usr/bin/env python3
"""Generate consolidated Grafana dashboards as k8s ConfigMap YAML files."""
import json
import os
DASHBOARD_DIR = os.path.expanduser("~/workplace/homelab/k8s/infra/monitoring/dashboards")
DS_PROM = {"type": "prometheus", "uid": "prometheus"}
DS_LOKI = {"type": "loki", "uid": "loki"}
def stat_panel(id, title, expr, x, y, w=4, h=4, unit="short", mappings=None, thresholds=None):
p = {
"id": id, "title": title, "type": "stat",
"gridPos": {"h": h, "w": w, "x": x, "y": y},
"datasource": DS_PROM,
"fieldConfig": {"defaults": {"unit": unit}},
"targets": [{"expr": expr}],
}
if mappings:
p["fieldConfig"]["defaults"]["mappings"] = mappings
if thresholds:
p["fieldConfig"]["defaults"]["thresholds"] = thresholds
p["fieldConfig"]["defaults"]["color"] = {"mode": "thresholds"}
return p
def ts_panel(id, title, exprs, x, y, w=8, h=8, unit="short"):
targets = []
for e in exprs:
if isinstance(e, tuple):
targets.append({"expr": e[0], "legendFormat": e[1]})
else:
targets.append({"expr": e, "legendFormat": "{{pod}}"})
return {
"id": id, "title": title, "type": "timeseries",
"gridPos": {"h": h, "w": w, "x": x, "y": y},
"datasource": DS_PROM,
"fieldConfig": {"defaults": {"unit": unit}},
"targets": targets,
}
def table_panel(id, title, expr, x, y, w=12, h=8):
return {
"id": id, "title": title, "type": "table",
"gridPos": {"h": h, "w": w, "x": x, "y": y},
"datasource": DS_PROM,
"targets": [{"expr": expr, "format": "table", "instant": True}],
}
def log_panel(id, title, query, x, y, w=24, h=10):
return {
"id": id, "title": title, "type": "logs",
"gridPos": {"h": h, "w": w, "x": x, "y": y},
"datasource": DS_LOKI,
"targets": [{"expr": query}],
}
def row(id, title, y, panels, collapsed=True):
return {
"id": id, "title": title, "type": "row",
"collapsed": collapsed, "gridPos": {"h": 1, "w": 24, "x": 0, "y": y},
"panels": panels,
}
def write_dashboard(filename, dashboard, folder):
cm = {
"apiVersion": "v1",
"kind": "ConfigMap",
"metadata": {
"name": filename.replace(".yaml", "-dashboard"),
"namespace": "logging",
"labels": {"grafana_dashboard": "1"},
"annotations": {"grafana_folder": folder},
},
"data": {
filename.replace(".yaml", ".json"): json.dumps(dashboard, separators=(",", ":"))
},
}
import yaml
path = os.path.join(DASHBOARD_DIR, filename)
with open(path, "w") as f:
yaml.dump(cm, f, default_flow_style=False, allow_unicode=True)
print(f" wrote {path}")
# ============================================================================
# Dashboard 1: Cluster Infrastructure
# ============================================================================
def build_cluster_infrastructure():
zero_thresholds = {"mode": "absolute", "steps": [
{"value": None, "color": "green"}, {"value": 1, "color": "red"}
]}
panels = [
row(1, "Cluster Health", 0, [
stat_panel(2, "Nodes Ready", 'count(kube_node_status_condition{condition="Ready",status="true"} == 1)', 0, 1, thresholds={"mode":"absolute","steps":[{"value":None,"color":"red"},{"value":3,"color":"green"}]}),
stat_panel(3, "Pods Pending", 'sum(kube_pod_status_phase{phase="Pending"}) OR on() vector(0)', 4, 1, thresholds=zero_thresholds),
stat_panel(4, "CrashLoopBackOff", 'sum(kube_pod_container_status_waiting_reason{reason="CrashLoopBackOff"}) OR on() vector(0)', 8, 1, thresholds=zero_thresholds),
stat_panel(5, "OOMKilled (1h)", 'sum(increase(kube_pod_container_status_last_terminated_reason{reason="OOMKilled"}[1h])) OR on() vector(0)', 12, 1, thresholds=zero_thresholds),
stat_panel(6, "Deploys Unavailable", 'count(kube_deployment_status_replicas_unavailable > 0) OR on() vector(0)', 16, 1, thresholds=zero_thresholds),
stat_panel(7, "Services Down", 'count(probe_success == 0) OR on() vector(0)', 20, 1, thresholds=zero_thresholds),
]),
row(10, "Jobs & CronJobs", 1, [
stat_panel(11, "Failed Jobs", 'count(kube_job_status_failed > 0) OR on() vector(0)', 0, 2, thresholds=zero_thresholds),
table_panel(12, "Failed Jobs Detail", 'kube_job_status_failed > 0', 4, 2, w=10),
table_panel(13, "Stuck Jobs (>1h)", 'kube_job_status_active == 1 and on(job_name,namespace) (time() - kube_job_status_start_time) > 3600', 14, 2, w=10),
ts_panel(14, "CronJob Last Success", [
('kube_cronjob_status_last_successful_time{namespace=~"cicd|kube-system|paperless"}', "{{namespace}}/{{cronjob}}")
], 0, 10, w=12, unit="dateTimeFromNow"),
ts_panel(15, "Container Restart Storm (top 10)", [
('topk(10, sum(rate(kube_pod_container_status_restarts_total[15m])) by (namespace, pod))', "{{namespace}}/{{pod}}")
], 12, 10, w=12),
]),
row(20, "Node Resources", 2, [
ts_panel(21, "CPU % by Node", [
('(1 - avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) by (instance)) * 100', "{{instance}}")
], 0, 3, unit="percent"),
ts_panel(22, "Memory % by Node", [
('(1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100', "{{instance}}")
], 8, 3, unit="percent"),
ts_panel(23, "Disk % by Node", [
('(1 - node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100', "{{instance}}")
], 16, 3, unit="percent"),
ts_panel(24, "Load Average", [
("node_load1", "1m {{instance}}"),
("node_load5", "5m {{instance}}"),
], 0, 11),
ts_panel(25, "Network Errors & Drops", [
("rate(node_network_receive_errs_total[5m])", "rx-err {{instance}}"),
("rate(node_network_transmit_errs_total[5m])", "tx-err {{instance}}"),
("rate(node_network_receive_drop_total[5m])", "rx-drop {{instance}}"),
], 8, 11),
]),
row(30, "Control Plane", 3, [
stat_panel(31, "API Server Up", 'min(up{job="apiserver"})', 0, 4, mappings=[
{"type":"value","options":{"0":{"text":"DOWN","color":"red"},"1":{"text":"UP","color":"green"}}}
]),
ts_panel(32, "API Server Request Rate", [
('sum(rate(apiserver_request_total[5m])) by (verb, code)', "{{verb}} {{code}}")
], 4, 4, w=10),
ts_panel(33, "API Server Error Rate %", [
('sum(rate(apiserver_request_total{code=~"5.."}[5m])) / sum(rate(apiserver_request_total[5m])) * 100', "5xx %")
], 14, 4, w=10, unit="percent"),
ts_panel(34, "API Server Latency", [
('histogram_quantile(0.95, sum(rate(apiserver_request_duration_seconds_bucket[5m])) by (le))', "p95"),
('histogram_quantile(0.99, sum(rate(apiserver_request_duration_seconds_bucket[5m])) by (le))', "p99"),
], 0, 12, unit="s"),
ts_panel(35, "etcd Request Duration", [
('histogram_quantile(0.99, sum(rate(etcd_request_duration_seconds_bucket[5m])) by (le))', "p99"),
], 8, 12, unit="s"),
]),
row(40, "Storage", 4, [
ts_panel(41, "Longhorn Disk Capacity", [
("longhorn_disk_capacity_bytes", "capacity {{node}}"),
("longhorn_disk_reservation_bytes", "reserved {{node}}"),
], 0, 5, unit="bytes"),
ts_panel(42, "PVC Phase", [
('kube_persistentvolumeclaim_status_phase', "{{namespace}}/{{persistentvolumeclaim}} {{phase}}")
], 8, 5),
]),
row(50, "DNS & Networking", 5, [
ts_panel(51, "CoreDNS Cache Hit Rate", [
('rate(coredns_cache_hits_total[5m]) / (rate(coredns_cache_hits_total[5m]) + rate(coredns_cache_misses_total[5m]))', "{{server}}")
], 0, 6, unit="percentunit"),
ts_panel(52, "CoreDNS Errors", [
('sum(rate(coredns_dns_responses_total{rcode=~"SERVFAIL|NXDOMAIN"}[5m])) by (rcode)', "{{rcode}}")
], 8, 6),
]),
row(60, "Logs", 6, [
ts_panel(61, "Error Rate by Namespace", [
('sum by (namespace) (count_over_time({namespace=~"kube-system|cert-manager|ingress-nginx|longhorn-system"} |= "error" [5m]))', "{{namespace}}")
], 0, 7),
log_panel(62, "Control Plane Logs", '{namespace="kube-system"}', 0, 15),
log_panel(63, "Cluster Addon Logs", '{namespace=~"cert-manager|ingress-nginx|longhorn-system"}', 0, 25),
]),
]
return {
"title": "Cluster Infrastructure",
"uid": "cluster-infra",
"schemaVersion": 39,
"timezone": "browser",
"time": {"from": "now-6h", "to": "now"},
"refresh": "30s",
"tags": ["infrastructure", "k8s"],
"panels": panels,
}
# ============================================================================
# Dashboard 3: API Gateway
# ============================================================================
def build_api_gateway():
panels = [
row(1, "Gateway Health", 0, [
stat_panel(2, "Gateway Pods Ready", 'sum(kube_pod_status_ready{namespace="api",condition="true"})', 0, 1, thresholds={"mode":"absolute","steps":[{"value":None,"color":"red"},{"value":3,"color":"green"}]}),
stat_panel(3, "Probe: healthz", 'probe_success{instance=~".*api.riotpiao.com/healthz"}', 4, 1, mappings=[
{"type":"value","options":{"0":{"text":"DOWN","color":"red"},"1":{"text":"UP","color":"green"}}}
]),
ts_panel(4, "Probe Latency", [
('probe_duration_seconds{instance=~".*api.riotpiao.com.*"}', "{{instance}}")
], 8, 1, unit="s"),
], collapsed=False),
row(10, "Ingress Traffic (nginx)", 1, [
ts_panel(11, "Request Rate by Status", [
('sum(rate(nginx_ingress_controller_requests{ingress="api"}[5m])) by (status)', "{{status}}")
], 0, 2),
ts_panel(12, "Error Rate %", [
('sum(rate(nginx_ingress_controller_requests{ingress="api",status=~"5.."}[5m])) / sum(rate(nginx_ingress_controller_requests{ingress="api"}[5m])) * 100', "5xx"),
('sum(rate(nginx_ingress_controller_requests{ingress="api",status=~"4.."}[5m])) / sum(rate(nginx_ingress_controller_requests{ingress="api"}[5m])) * 100', "4xx"),
], 8, 2, unit="percent"),
ts_panel(13, "Latency p50/p95/p99", [
('histogram_quantile(0.50, sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress="api"}[5m])) by (le))', "p50"),
('histogram_quantile(0.95, sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress="api"}[5m])) by (le))', "p95"),
('histogram_quantile(0.99, sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress="api"}[5m])) by (le))', "p99"),
], 16, 2, unit="s"),
]),
row(20, "LLM Serving", 2, [
stat_panel(21, "LLM Pods Ready", 'sum(kube_pod_status_ready{namespace="llm-serving",condition="true"})', 0, 3),
ts_panel(22, "CPU by Predictor", [
('sum(rate(container_cpu_usage_seconds_total{namespace="llm-serving"}[5m])) by (pod)', "{{pod}}")
], 4, 3),
ts_panel(23, "Memory by Predictor", [
('sum(container_memory_working_set_bytes{namespace="llm-serving"}) by (pod)', "{{pod}}")
], 12, 3, unit="bytes"),
ts_panel(24, "Predictor Restarts", [
('sum(rate(kube_pod_container_status_restarts_total{namespace="llm-serving"}[15m])) by (pod)', "{{pod}}")
], 20, 3, w=4),
]),
row(30, "Gateway Resources", 3, [
ts_panel(31, "CPU by Gateway Pod", [
('sum(rate(container_cpu_usage_seconds_total{namespace="api"}[5m])) by (pod)', "{{pod}}")
], 0, 4),
ts_panel(32, "Memory by Gateway Pod", [
('sum(container_memory_working_set_bytes{namespace="api"}) by (pod)', "{{pod}}")
], 8, 4, unit="bytes"),
ts_panel(33, "Gateway Restarts", [
('sum(rate(kube_pod_container_status_restarts_total{namespace="api"}[15m])) by (pod)', "{{pod}}")
], 16, 4),
]),
row(40, "Logs", 4, [
log_panel(41, "Gateway Logs", '{namespace="api",container="gateway"}', 0, 5),
log_panel(42, "LLM Serving Logs", '{namespace="llm-serving"}', 0, 15),
]),
]
return {
"title": "API Gateway",
"uid": "api-gateway",
"schemaVersion": 39,
"timezone": "browser",
"time": {"from": "now-6h", "to": "now"},
"refresh": "30s",
"tags": ["api", "gateway", "llm"],
"panels": panels,
}
# ============================================================================
# Generate
# ============================================================================
print("Generating dashboards...")
# Dashboard 1
write_dashboard("cluster-infrastructure.yaml", build_cluster_infrastructure(), "Infrastructure")
# Dashboard 3
write_dashboard("api-gateway.yaml", build_api_gateway(), "API")
print("Done.")
@@ -1,121 +0,0 @@
# k8s/monitoring/dashboards/hardware-overview.yaml
# Trimmed operator at-a-glance view across all nodes — node-exporter already
# powers the deep-dive "Node Exporter Full" (#1860, see grafana-values.yaml),
# this is the quick health-check version, not a replacement for it.
apiVersion: v1
kind: ConfigMap
metadata:
name: hardware-overview-dashboard
namespace: logging
labels:
grafana_dashboard: "1"
data:
hardware-overview.json: |
{
"title": "Hardware Statistics (Operator Overview)",
"uid": "hardware-overview",
"schemaVersion": 39,
"timezone": "browser",
"time": { "from": "now-6h", "to": "now" },
"refresh": "30s",
"panels": [
{
"id": 1,
"title": "Nodes up / down",
"type": "stat",
"gridPos": { "h": 5, "w": 24, "x": 0, "y": 0 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": {
"defaults": {
"mappings": [
{ "type": "value", "options": { "0": { "text": "DOWN", "color": "red" } } },
{ "type": "value", "options": { "1": { "text": "UP", "color": "green" } } }
]
}
},
"targets": [
{ "expr": "up{job=~\".*node-exporter.*\"}", "legendFormat": "{{instance}}" }
]
},
{
"id": 2,
"title": "CPU usage % by node",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 5 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": { "defaults": { "unit": "percent", "max": 100, "min": 0 } },
"targets": [
{
"expr": "(1 - avg(rate(node_cpu_seconds_total{mode=\"idle\"}[5m])) by (instance)) * 100",
"legendFormat": "{{instance}}"
}
]
},
{
"id": 3,
"title": "Memory usage % by node",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 5 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": { "defaults": { "unit": "percent", "max": 100, "min": 0 } },
"targets": [
{
"expr": "(1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100",
"legendFormat": "{{instance}}"
}
]
},
{
"id": 4,
"title": "Root filesystem usage % by node",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 13 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": { "defaults": { "unit": "percent", "max": 100, "min": 0 } },
"targets": [
{
"expr": "(1 - node_filesystem_avail_bytes{mountpoint=\"/\"} / node_filesystem_size_bytes{mountpoint=\"/\"}) * 100",
"legendFormat": "{{instance}}"
}
]
},
{
"id": 5,
"title": "Root filesystem space remaining",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 13 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": { "defaults": { "unit": "bytes" } },
"targets": [
{
"expr": "node_filesystem_avail_bytes{mountpoint=\"/\"}",
"legendFormat": "{{instance}}"
}
]
},
{
"id": 6,
"title": "Network errors/drops by node",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 21 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{ "expr": "rate(node_network_receive_errs_total[5m])", "legendFormat": "{{instance}} rx errs" },
{ "expr": "rate(node_network_transmit_errs_total[5m])", "legendFormat": "{{instance}} tx errs" },
{ "expr": "rate(node_network_receive_drop_total[5m])", "legendFormat": "{{instance}} rx drops" },
{ "expr": "rate(node_network_transmit_drop_total[5m])", "legendFormat": "{{instance}} tx drops" }
]
},
{
"id": 7,
"title": "Load average (1m / 5m) by node",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 21 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{ "expr": "node_load1", "legendFormat": "{{instance}} load1" },
{ "expr": "node_load5", "legendFormat": "{{instance}} load5" }
]
}
]
}
@@ -1,185 +0,0 @@
# k8s/monitoring/dashboards/kube-controller-health.yaml
# Talos binds controller-manager/scheduler/etcd to 127.0.0.1, so Prometheus
# can't scrape them directly (see prometheus-values.yaml). kube-apiserver is
# the one control-plane component that's still reachable (its ServiceMonitor
# targets the in-cluster `kubernetes` service, not localhost) — paired with
# kube-state-metrics signals as a proxy for controller/scheduler health.
apiVersion: v1
kind: ConfigMap
metadata:
name: kube-controller-health-dashboard
namespace: logging
labels:
grafana_dashboard: "1"
data:
kube-controller-health.json: |
{
"title": "Kube-Controller Health",
"uid": "kube-controller-health",
"schemaVersion": 39,
"timezone": "browser",
"time": { "from": "now-6h", "to": "now" },
"refresh": "30s",
"panels": [
{
"id": 1,
"title": "API server — up",
"type": "stat",
"gridPos": { "h": 4, "w": 6, "x": 0, "y": 0 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": {
"defaults": {
"mappings": [
{ "type": "value", "options": { "0": { "text": "DOWN", "color": "red" } } },
{ "type": "value", "options": { "1": { "text": "UP", "color": "green" } } }
]
}
},
"targets": [
{ "expr": "min(up{job=\"apiserver\"})", "legendFormat": "apiserver" }
]
},
{
"id": 2,
"title": "API server — request rate by verb/code",
"type": "timeseries",
"gridPos": { "h": 8, "w": 18, "x": 6, "y": 0 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{
"expr": "sum(rate(apiserver_request_total[5m])) by (verb, code)",
"legendFormat": "{{verb}} {{code}}"
}
]
},
{
"id": 3,
"title": "API server — error rate % (5xx)",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 8 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": { "defaults": { "unit": "percent" } },
"targets": [
{
"expr": "sum(rate(apiserver_request_total{code=~\"5..\"}[5m])) / sum(rate(apiserver_request_total[5m])) * 100",
"legendFormat": "5xx %"
}
]
},
{
"id": 4,
"title": "API server — latency p95 / p99",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 8 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": { "defaults": { "unit": "s" } },
"targets": [
{
"expr": "histogram_quantile(0.95, sum(rate(apiserver_request_duration_seconds_bucket[5m])) by (le))",
"legendFormat": "p95"
},
{
"expr": "histogram_quantile(0.99, sum(rate(apiserver_request_duration_seconds_bucket[5m])) by (le))",
"legendFormat": "p99"
}
]
},
{
"id": 5,
"title": "Pods stuck Pending",
"type": "stat",
"gridPos": { "h": 5, "w": 8, "x": 0, "y": 16 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": {
"defaults": {
"thresholds": {
"mode": "absolute",
"steps": [
{ "value": 0, "color": "green" },
{ "value": 1, "color": "yellow" },
{ "value": 5, "color": "red" }
]
}
}
},
"targets": [
{ "expr": "sum(kube_pod_status_phase{phase=\"Pending\"}) OR on() vector(0)", "legendFormat": "pending" }
]
},
{
"id": 6,
"title": "CrashLoopBackOff containers",
"type": "stat",
"gridPos": { "h": 5, "w": 8, "x": 8, "y": 16 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": {
"defaults": {
"thresholds": {
"mode": "absolute",
"steps": [
{ "value": 0, "color": "green" },
{ "value": 1, "color": "red" }
]
}
}
},
"targets": [
{ "expr": "sum(kube_pod_container_status_waiting_reason{reason=\"CrashLoopBackOff\"}) OR on() vector(0)", "legendFormat": "crashlooping" }
]
},
{
"id": 7,
"title": "Nodes NotReady",
"type": "stat",
"gridPos": { "h": 5, "w": 8, "x": 16, "y": 16 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": {
"defaults": {
"thresholds": {
"mode": "absolute",
"steps": [
{ "value": 0, "color": "green" },
{ "value": 1, "color": "red" }
]
}
}
},
"targets": [
{ "expr": "count(kube_node_status_condition{condition=\"Ready\", status=\"true\"} == 0) OR on() vector(0)", "legendFormat": "not ready" }
]
},
{
"id": 8,
"title": "Failed Jobs",
"type": "table",
"gridPos": { "h": 7, "w": 12, "x": 0, "y": 21 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{ "expr": "kube_job_status_failed > 0", "format": "table", "instant": true }
]
},
{
"id": 9,
"title": "Deployments with unavailable replicas",
"type": "table",
"gridPos": { "h": 7, "w": 12, "x": 12, "y": 21 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{ "expr": "kube_deployment_status_replicas_unavailable > 0", "format": "table", "instant": true }
]
},
{
"id": 10,
"title": "Container restart rate by pod",
"type": "timeseries",
"gridPos": { "h": 8, "w": 24, "x": 0, "y": 28 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{
"expr": "sum(rate(kube_pod_container_status_restarts_total[15m])) by (namespace, pod)",
"legendFormat": "{{namespace}}/{{pod}}"
}
]
}
]
}
@@ -1,28 +0,0 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: llm-frontend-dashboard
namespace: logging
labels:
grafana_dashboard: "1"
annotations:
grafana_folder: "LLM"
# No request-level panels. The rate/error/latency/bandwidth row used to run
# on Kong's prometheus plugin; Kong was retired 2026-08-19 and the Go
# gateway that replaced it does not expose /metrics yet, so those panels
# were removed rather than left querying series that no longer exist.
# What is left is pod-level: readiness, CPU/memory, restarts, logs.
#
# Restoring request-level and per-model observability means wiring three
# sources, none of which are in place: gateway metrics (RED plus token
# counts and TTFT, which the gateway can measure because it sees the
# response stream), vLLM's own /metrics on reasoning-predictor (rich --
# vllm:time_to_first_token_seconds, vllm:inter_token_latency_seconds,
# vllm:e2e_request_latency_seconds, vllm:kv_cache_usage_perc), and TEI's
# /metrics on embeddings/reranker. Ollama exposes no Prometheus endpoint at
# all (verified: /metrics returns 404), so ornith can only ever be observed
# from the gateway side. No ServiceMonitor exists for the llm-serving
# namespace today, so none of the engine metrics are being scraped.
data:
llm-frontend.json: |
{"title":"LLM Frontend","uid":"llm-frontend","schemaVersion":39,"timezone":"browser","time":{"from":"now-6h","to":"now"},"refresh":"30s","panels":[{"id":1,"title":"Row: Availability","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":0},"panels":[{"id":2,"title":"llm-serving pods ready","type":"stat","gridPos":{"h":4,"w":8,"x":0,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(kube_pod_status_ready{namespace=\"llm-serving\",condition=\"true\"})"}]},{"id":3,"title":"agent-pod ready","type":"stat","gridPos":{"h":4,"w":8,"x":8,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(kube_pod_status_ready{namespace=\"agent-pod\",condition=\"true\"})"}]},{"id":4,"title":"api gateway pods ready","type":"stat","gridPos":{"h":4,"w":8,"x":16,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(kube_pod_status_ready{namespace=\"api\",condition=\"true\"})"}]}]},{"id":10,"title":"Row: Resources","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":1},"panels":[{"id":11,"title":"CPU by pod","type":"timeseries","gridPos":{"h":8,"w":12,"x":0,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(container_cpu_usage_seconds_total{namespace=~\"llm-serving|agent-pod|api\"}[5m])) by (namespace, pod)","legendFormat":"{{namespace}}/{{pod}}"}]},{"id":12,"title":"Memory by pod","type":"timeseries","gridPos":{"h":8,"w":12,"x":12,"y":2},"fieldConfig":{"defaults":{"unit":"bytes"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(container_memory_working_set_bytes{namespace=~\"llm-serving|agent-pod|api\"}) by (namespace, pod)","legendFormat":"{{namespace}}/{{pod}}"}]},{"id":13,"title":"GPU-node predictor restarts","type":"timeseries","gridPos":{"h":8,"w":24,"x":0,"y":10},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(kube_pod_container_status_restarts_total{namespace=\"llm-serving\"}[15m])) by (pod)","legendFormat":"{{pod}}"}]}]},{"id":20,"title":"Row: Logs","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":2},"panels":[{"id":21,"title":"llm-serving logs","type":"logs","gridPos":{"h":10,"w":24,"x":0,"y":4},"datasource":{"type":"loki","uid":"loki"},"targets":[{"expr":"{namespace=\"llm-serving\"}"}]},{"id":22,"title":"agent-pod logs (pi runs)","type":"logs","gridPos":{"h":10,"w":24,"x":0,"y":14},"datasource":{"type":"loki","uid":"loki"},"targets":[{"expr":"{namespace=\"agent-pod\"}"}]},{"id":23,"title":"api gateway logs","type":"logs","gridPos":{"h":10,"w":24,"x":0,"y":24},"datasource":{"type":"loki","uid":"loki"},"targets":[{"expr":"{namespace=\"api\"}"}]}]}]}
@@ -1,172 +0,0 @@
# k8s/monitoring/dashboards/service-availability.yaml
# Active uptime/availability from blackbox-exporter probes — the signal that
# covers low-traffic services (Vault, MinIO, Longhorn UI) where RED metrics
# alone can't distinguish "idle" from "down".
apiVersion: v1
kind: ConfigMap
metadata:
name: service-availability-dashboard
namespace: logging
labels:
grafana_dashboard: "1"
data:
service-availability.json: |
{
"title": "Service Availability & Certificate Expiration",
"uid": "svc-availability",
"schemaVersion": 39,
"timezone": "browser",
"time": { "from": "now-24h", "to": "now" },
"refresh": "30s",
"panels": [
{
"id": 1,
"title": "Up / Down — all probed services",
"type": "stat",
"gridPos": { "h": 6, "w": 24, "x": 0, "y": 0 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": {
"defaults": {
"mappings": [
{ "type": "value", "options": { "0": { "text": "DOWN", "color": "red" } } },
{ "type": "value", "options": { "1": { "text": "UP", "color": "green" } } }
],
"thresholds": {
"mode": "absolute",
"steps": [
{ "value": 0, "color": "red" },
{ "value": 1, "color": "green" }
]
}
}
},
"targets": [
{ "expr": "probe_success", "legendFormat": "{{instance}}" }
]
},
{
"id": 2,
"title": "Uptime % trend",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 6 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": { "defaults": { "unit": "percent", "max": 100, "min": 0 } },
"targets": [
{
"expr": "avg_over_time(probe_success[$__rate_interval]) * 100",
"legendFormat": "{{instance}}"
}
]
},
{
"id": 3,
"title": "Probe latency",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 6 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": { "defaults": { "unit": "s" } },
"targets": [
{ "expr": "probe_duration_seconds", "legendFormat": "{{instance}}" }
]
},
{
"id": 4,
"title": "7-day SLO (% successful probes)",
"type": "table",
"gridPos": { "h": 8, "w": 24, "x": 0, "y": 14 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": {
"defaults": {
"unit": "percent",
"thresholds": {
"mode": "absolute",
"steps": [
{ "value": 0, "color": "red" },
{ "value": 99, "color": "yellow" },
{ "value": 99.9, "color": "green" }
]
}
}
},
"targets": [
{
"expr": "avg_over_time(probe_success[7d]) * 100",
"format": "table",
"instant": true
}
]
},
{
"id": 5,
"title": "Services DOWN right now",
"type": "stat",
"gridPos": { "h": 4, "w": 12, "x": 0, "y": 22 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": {
"defaults": {
"thresholds": {
"mode": "absolute",
"steps": [
{ "value": 0, "color": "green" },
{ "value": 1, "color": "red" }
]
}
}
},
"targets": [
{ "expr": "count(probe_success == 0) OR on() vector(0)", "legendFormat": "down" }
]
},
{
"id": 6,
"title": "Certs expiring in < 14 days",
"type": "stat",
"gridPos": { "h": 4, "w": 12, "x": 12, "y": 22 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": {
"defaults": {
"thresholds": {
"mode": "absolute",
"steps": [
{ "value": 0, "color": "green" },
{ "value": 1, "color": "red" }
]
}
}
},
"targets": [
{
"expr": "count((certmanager_certificate_expiration_timestamp_seconds - time()) / 86400 < 14) OR on() vector(0)",
"legendFormat": "expiring"
}
]
},
{
"id": 7,
"title": "Certificate expiry — days remaining",
"type": "table",
"gridPos": { "h": 8, "w": 24, "x": 0, "y": 26 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": {
"defaults": {
"thresholds": {
"mode": "absolute",
"steps": [
{ "value": 0, "color": "red" },
{ "value": 14, "color": "yellow" },
{ "value": 30, "color": "green" }
]
}
}
},
"targets": [
{
"expr": "(certmanager_certificate_expiration_timestamp_seconds - time()) / 86400",
"legendFormat": "{{name}}",
"format": "table",
"instant": true
}
]
}
]
}
@@ -1,141 +0,0 @@
# k8s/monitoring/dashboards/service-golden-signals.yaml
# RED metrics (rate/errors/duration) for every service fronted by ingress-nginx.
# Picked up automatically by Grafana's sidecar (grafana_dashboard=1 label) — see
# sidecar.dashboards in k8s/logging/grafana-values.yaml.
apiVersion: v1
kind: ConfigMap
metadata:
name: service-golden-signals-dashboard
namespace: logging
labels:
grafana_dashboard: "1"
data:
service-golden-signals.json: |
{
"title": "Latency & Golden Signals (Ingress RED)",
"uid": "svc-golden-signals",
"schemaVersion": 39,
"timezone": "browser",
"time": { "from": "now-6h", "to": "now" },
"refresh": "30s",
"templating": {
"list": [
{
"name": "ingress",
"type": "query",
"datasource": { "type": "prometheus", "uid": "prometheus" },
"query": "label_values(nginx_ingress_controller_requests, ingress)",
"refresh": 2,
"includeAll": false
}
]
},
"panels": [
{
"id": 1,
"title": "Request rate by status — $ingress",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 0 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{
"expr": "sum(rate(nginx_ingress_controller_requests{ingress=\"$ingress\"}[5m])) by (status)",
"legendFormat": "{{status}}"
}
]
},
{
"id": 2,
"title": "Error rate % (4xx / 5xx) — $ingress",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 0 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": { "defaults": { "unit": "percent" } },
"targets": [
{
"expr": "sum(rate(nginx_ingress_controller_requests{ingress=\"$ingress\", status=~\"5..\"}[5m])) / sum(rate(nginx_ingress_controller_requests{ingress=\"$ingress\"}[5m])) * 100",
"legendFormat": "5xx"
},
{
"expr": "sum(rate(nginx_ingress_controller_requests{ingress=\"$ingress\", status=~\"4..\"}[5m])) / sum(rate(nginx_ingress_controller_requests{ingress=\"$ingress\"}[5m])) * 100",
"legendFormat": "4xx"
}
]
},
{
"id": 3,
"title": "Latency p50 / p95 / p99 — $ingress",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 8 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": { "defaults": { "unit": "s" } },
"targets": [
{
"expr": "histogram_quantile(0.50, sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress=\"$ingress\"}[5m])) by (le))",
"legendFormat": "p50"
},
{
"expr": "histogram_quantile(0.95, sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress=\"$ingress\"}[5m])) by (le))",
"legendFormat": "p95"
},
{
"expr": "histogram_quantile(0.99, sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress=\"$ingress\"}[5m])) by (le))",
"legendFormat": "p99"
}
]
},
{
"id": 4,
"title": "All services — traffic overview",
"type": "table",
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 8 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{
"expr": "topk(11, sum(rate(nginx_ingress_controller_requests[5m])) by (ingress))",
"format": "table",
"instant": true
}
]
},
{
"id": 5,
"title": "Customer-facing failures (5xx count, window total)",
"type": "stat",
"gridPos": { "h": 5, "w": 12, "x": 0, "y": 16 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": {
"defaults": {
"thresholds": {
"mode": "absolute",
"steps": [
{ "value": 0, "color": "green" },
{ "value": 1, "color": "yellow" },
{ "value": 50, "color": "red" }
]
}
}
},
"targets": [
{
"expr": "sum(increase(nginx_ingress_controller_requests{status=~\"5..\"}[$__range])) OR on() vector(0)",
"legendFormat": "5xx total"
}
]
},
{
"id": 6,
"title": "Top 5 error-contributing services",
"type": "table",
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 16 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{
"expr": "topk(5, sum(rate(nginx_ingress_controller_requests{status=~\"5..\"}[5m])) by (ingress))",
"format": "table",
"instant": true
}
]
}
]
}
@@ -1,109 +0,0 @@
# k8s/monitoring/dashboards/service-internals.yaml
# Native per-service metrics — the "why" layer behind the ingress RED/uptime
# dashboards (e.g. ingress shows MinIO is slow; this shows disk offline).
apiVersion: v1
kind: ConfigMap
metadata:
name: service-internals-dashboard
namespace: logging
labels:
grafana_dashboard: "1"
data:
service-internals.json: |
{
"title": "Service Internals (MinIO / Forgejo / Argo CD / cert-manager / Vault / Longhorn)",
"uid": "svc-internals",
"schemaVersion": 39,
"timezone": "browser",
"time": { "from": "now-6h", "to": "now" },
"refresh": "30s",
"panels": [
{ "id": 1, "title": "MinIO — disk/node offline", "type": "timeseries",
"gridPos": { "h": 6, "w": 12, "x": 0, "y": 0 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{ "expr": "minio_cluster_disk_offline_total", "legendFormat": "disks offline" },
{ "expr": "minio_cluster_nodes_offline_total", "legendFormat": "nodes offline" }
]
},
{ "id": 2, "title": "MinIO — S3 request errors", "type": "timeseries",
"gridPos": { "h": 6, "w": 12, "x": 12, "y": 0 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{ "expr": "sum(rate(minio_s3_requests_errors_total[5m])) by (api)", "legendFormat": "{{api}}" }
]
},
{ "id": 3, "title": "MinIO — S3 TTFB latency", "type": "timeseries",
"gridPos": { "h": 6, "w": 12, "x": 0, "y": 6 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": { "defaults": { "unit": "s" } },
"targets": [
{ "expr": "minio_s3_time_ttfb_seconds_distribution", "legendFormat": "{{api}}" }
]
},
{ "id": 4, "title": "Forgejo — repos / orgs", "type": "stat",
"gridPos": { "h": 6, "w": 12, "x": 12, "y": 6 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{ "expr": "gitea_repositories", "legendFormat": "repos" },
{ "expr": "gitea_organizations", "legendFormat": "orgs" }
]
},
{ "id": 5, "title": "Forgejo — process health (CPU/mem)", "type": "timeseries",
"gridPos": { "h": 6, "w": 12, "x": 0, "y": 12 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{ "expr": "rate(process_cpu_seconds_total{job=~\".*forgejo.*|.*gitea.*\"}[5m])", "legendFormat": "cpu" },
{ "expr": "process_resident_memory_bytes{job=~\".*forgejo.*|.*gitea.*\"}", "legendFormat": "mem" }
]
},
{ "id": 6, "title": "Argo CD — app sync/health status", "type": "table",
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 12 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{ "expr": "argocd_app_info", "format": "table", "instant": true }
]
},
{ "id": 7, "title": "cert-manager — days to cert expiry", "type": "stat",
"gridPos": { "h": 6, "w": 12, "x": 0, "y": 18 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": {
"defaults": {
"thresholds": {
"mode": "absolute",
"steps": [
{ "value": 0, "color": "red" },
{ "value": 14, "color": "yellow" },
{ "value": 30, "color": "green" }
]
}
}
},
"targets": [
{ "expr": "(certmanager_certificate_expiration_timestamp_seconds - time()) / 86400", "legendFormat": "{{name}}" }
]
},
{ "id": 8, "title": "Vault — sealed/unsealed", "type": "stat",
"gridPos": { "h": 6, "w": 6, "x": 12, "y": 20 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": {
"defaults": {
"mappings": [
{ "type": "value", "options": { "0": { "text": "SEALED", "color": "red" } } },
{ "type": "value", "options": { "1": { "text": "UNSEALED", "color": "green" } } }
]
}
},
"targets": [
{ "expr": "vault_core_unsealed", "legendFormat": "vault" }
]
},
{ "id": 9, "title": "Longhorn — volume robustness", "type": "table",
"gridPos": { "h": 6, "w": 6, "x": 18, "y": 20 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{ "expr": "longhorn_volume_robustness", "format": "table", "instant": true }
]
}
]
}
@@ -1,12 +0,0 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: svc-argocd-dashboard
namespace: logging
labels:
grafana_dashboard: "1"
annotations:
grafana_folder: "Argo CD"
data:
svc-argocd.json: |
{"title":"Argo CD — Service Overview","uid":"svc-argocd","schemaVersion":39,"timezone":"browser","time":{"from":"now-6h","to":"now"},"refresh":"30s","panels":[{"id":1,"title":"Row: Availability","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":0},"panels":[{"id":2,"title":"Up","type":"stat","gridPos":{"h":4,"w":6,"x":0,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"color":{"mode":"thresholds"}}},"targets":[{"expr":"min(up{job=~\"argocd-.*\"})"}]},{"id":3,"title":"HTTP requests","type":"timeseries","gridPos":{"h":8,"w":9,"x":6,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(argocd_http_request_total[5m])) by (status)","legendFormat":"{{status}}"}]},{"id":4,"title":"Error rate %","type":"timeseries","gridPos":{"h":8,"w":9,"x":15,"y":1},"fieldConfig":{"defaults":{"unit":"percent"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(argocd_http_request_total{status=~\"5..\"}[5m])) / sum(rate(argocd_http_request_total[5m])) * 100"}]},{"id":5,"title":"Request latency","type":"timeseries","gridPos":{"h":8,"w":12,"x":0,"y":9},"fieldConfig":{"defaults":{"unit":"s"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"histogram_quantile(0.95, sum(rate(argocd_http_request_duration_seconds_bucket[5m])) by (le))","legendFormat":"p95"}]}]},{"id":10,"title":"Row: Resources","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":1},"panels":[{"id":11,"title":"CPU","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(container_cpu_usage_seconds_total{namespace=\"argocd\"}[5m])) by (pod)","legendFormat":"{{pod}}"}]},{"id":12,"title":"Memory","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":2},"fieldConfig":{"defaults":{"unit":"bytes"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(container_memory_working_set_bytes{namespace=\"argocd\"}) by (pod)","legendFormat":"{{pod}}"}]},{"id":13,"title":"Restarts","type":"timeseries","gridPos":{"h":8,"w":8,"x":16,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(kube_pod_container_status_restarts_total{namespace=\"argocd\"}[15m])) by (pod)","legendFormat":"{{pod}}"}]}]},{"id":20,"title":"Row: Applications & Sync","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":2},"panels":[{"id":21,"title":"Applications","type":"stat","gridPos":{"h":6,"w":6,"x":0,"y":3},"fieldConfig":{"defaults":{"unit":"short"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"argocd_app_total"}]},{"id":22,"title":"Sync by status","type":"timeseries","gridPos":{"h":6,"w":9,"x":6,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(argocd_app_sync_total[5m])) by (sync_status)","legendFormat":"{{sync_status}}"}]},{"id":23,"title":"Degraded apps","type":"stat","gridPos":{"h":6,"w":6,"x":15,"y":3},"fieldConfig":{"defaults":{"unit":"short"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"argocd_app_health_degraded_total"}]},{"id":24,"title":"Git sync ops","type":"timeseries","gridPos":{"h":6,"w":12,"x":0,"y":9},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(argocd_git_sync_total[5m])) by (git_operation,git_status)","legendFormat":"{{git_operation}}/{{git_status}}"}]}]},{"id":30,"title":"Row: Logs","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":3},"panels":[{"id":31,"title":"Recent logs","type":"logs","gridPos":{"h":10,"w":24,"x":0,"y":4},"datasource":{"type":"loki","uid":"loki"},"targets":[{"expr":"{namespace=\"argocd\"}"}]}]}]}
@@ -1,143 +0,0 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: svc-authentik-dashboard
namespace: logging
labels:
grafana_dashboard: "1"
annotations:
grafana_folder: "Authentik"
data:
svc-authentik.json: |
{
"title": "Authentik — Service Overview",
"uid": "svc-authentik",
"schemaVersion": 39,
"timezone": "browser",
"time": { "from": "now-6h", "to": "now" },
"refresh": "30s",
"panels": [
{
"id": 1, "title": "Row: Availability & Golden Signals", "type": "row",
"collapsed": true, "gridPos": { "h": 1, "w": 24, "x": 0, "y": 0 },
"panels": [
{
"id": 2, "title": "Up", "type": "stat",
"gridPos": { "h": 4, "w": 6, "x": 0, "y": 1 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": {
"defaults": {
"color": { "mode": "thresholds" },
"mappings": [
{ "type": "value", "options": { "0": { "text": "DOWN", "color": "red" }, "1": { "text": "UP", "color": "green" } } }
],
"thresholds": { "mode": "absolute", "steps": [ { "value": null, "color": "red" }, { "value": 1, "color": "green" } ] }
}
},
"targets": [{ "expr": "min(up{job=\"authentik-server\"})" }]
},
{
"id": 3, "title": "HTTP request rate by status", "type": "timeseries",
"gridPos": { "h": 8, "w": 9, "x": 6, "y": 1 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [{ "expr": "sum(rate(authentik_flows_execution_stage_time_count[5m])) by (flow_slug)", "legendFormat": "{{flow_slug}}" }]
},
{
"id": 4, "title": "Error rate % (5xx)", "type": "timeseries",
"gridPos": { "h": 8, "w": 9, "x": 15, "y": 1 },
"fieldConfig": { "defaults": { "unit": "percent" } },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [{ "expr": "(1 - (authentik_flows_cached / authentik_flows_execution_stage_time_count)) * 100" }]
},
{
"id": 5, "title": "Request duration p50/p95/p99", "type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 9 },
"fieldConfig": { "defaults": { "unit": "s" } },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{ "expr": "histogram_quantile(0.50, sum(rate(authentik_main_request_duration_seconds_bucket[5m])) by (le))", "legendFormat": "p50" },
{ "expr": "histogram_quantile(0.95, sum(rate(authentik_main_request_duration_seconds_bucket[5m])) by (le))", "legendFormat": "p95" },
{ "expr": "histogram_quantile(0.99, sum(rate(authentik_main_request_duration_seconds_bucket[5m])) by (le))", "legendFormat": "p99" }
]
}
]
},
{
"id": 10, "title": "Row: Resource Usage", "type": "row",
"collapsed": true, "gridPos": { "h": 1, "w": 24, "x": 0, "y": 1 },
"panels": [
{
"id": 11, "title": "CPU by pod", "type": "timeseries",
"gridPos": { "h": 8, "w": 8, "x": 0, "y": 2 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [{ "expr": "sum(rate(container_cpu_usage_seconds_total{namespace=\"iam\",pod=~\"authentik.*\"}[5m])) by (pod)", "legendFormat": "{{pod}}" }]
},
{
"id": 12, "title": "Memory by pod", "type": "timeseries",
"gridPos": { "h": 8, "w": 8, "x": 8, "y": 2 },
"fieldConfig": { "defaults": { "unit": "bytes" } },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [{ "expr": "sum(container_memory_working_set_bytes{namespace=\"iam\",pod=~\"authentik.*\"}) by (pod)", "legendFormat": "{{pod}}" }]
},
{
"id": 13, "title": "Restart rate by pod", "type": "timeseries",
"gridPos": { "h": 8, "w": 8, "x": 16, "y": 2 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [{ "expr": "sum(rate(kube_pod_container_status_restarts_total{namespace=\"iam\",pod=~\"authentik.*\"}[15m])) by (pod)", "legendFormat": "{{pod}}" }]
}
]
},
{
"id": 20, "title": "Row: Identity Provider (OIDC / OAuth2)", "type": "row",
"collapsed": true, "gridPos": { "h": 1, "w": 24, "x": 0, "y": 2 },
"panels": [
{
"id": 21, "title": "Outpost connections", "type": "stat",
"gridPos": { "h": 7, "w": 6, "x": 0, "y": 3 },
"fieldConfig": { "defaults": { "unit": "short" } },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [{ "expr": "authentik_outposts_connected" }]
},
{
"id": 22, "title": "Flows cached", "type": "stat",
"gridPos": { "h": 7, "w": 6, "x": 6, "y": 3 },
"fieldConfig": { "defaults": { "unit": "short" } },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [{ "expr": "authentik_flows_cached" }]
},
{
"id": 23, "title": "Policies cached", "type": "stat",
"gridPos": { "h": 7, "w": 6, "x": 12, "y": 3 },
"fieldConfig": { "defaults": { "unit": "short" } },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [{ "expr": "authentik_policies_cached" }]
},
{
"id": 24, "title": "Queued tasks", "type": "stat",
"gridPos": { "h": 7, "w": 6, "x": 18, "y": 3 },
"fieldConfig": { "defaults": { "color": { "mode": "thresholds" }, "unit": "short" } },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [{ "expr": "authentik_tasks_queued" }]
},
{
"id": 25, "title": "Admin workers", "type": "timeseries",
"gridPos": { "h": 7, "w": 12, "x": 0, "y": 10 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [{ "expr": "authentik_admin_workers" }]
}
]
},
{
"id": 30, "title": "Row: Logs", "type": "row",
"collapsed": true, "gridPos": { "h": 1, "w": 24, "x": 0, "y": 3 },
"panels": [
{
"id": 31, "title": "Recent logs", "type": "logs",
"gridPos": { "h": 10, "w": 24, "x": 0, "y": 4 },
"datasource": { "type": "loki", "uid": "loki" },
"targets": [{ "expr": "{namespace=\"iam\"}" }]
}
]
}
]
}
@@ -1,12 +0,0 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: svc-forgejo-dashboard
namespace: logging
labels:
grafana_dashboard: "1"
annotations:
grafana_folder: "Forgejo"
data:
svc-forgejo.json: |
{"title":"Forgejo — Service Overview","uid":"svc-forgejo","schemaVersion":39,"timezone":"browser","time":{"from":"now-6h","to":"now"},"refresh":"30s","panels":[{"id":1,"title":"Row: Availability","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":0},"panels":[{"id":2,"title":"Up","type":"stat","gridPos":{"h":4,"w":6,"x":0,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"color":{"mode":"thresholds"}}},"targets":[{"expr":"min(up{job=\"forgejo\"})"}]},{"id":3,"title":"HTTP requests by method","type":"timeseries","gridPos":{"h":8,"w":9,"x":6,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(forgejo_http_request_total[5m])) by (method)","legendFormat":"{{method}}"}]},{"id":4,"title":"Error rate %","type":"timeseries","gridPos":{"h":8,"w":9,"x":15,"y":1},"fieldConfig":{"defaults":{"unit":"percent"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(forgejo_http_request_total{status=~\"5..\"}[5m])) / sum(rate(forgejo_http_request_total[5m])) * 100"}]},{"id":5,"title":"Request latency","type":"timeseries","gridPos":{"h":8,"w":12,"x":0,"y":9},"fieldConfig":{"defaults":{"unit":"s"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"histogram_quantile(0.95, sum(rate(forgejo_http_request_duration_seconds_bucket[5m])) by (le))","legendFormat":"p95"}]}]},{"id":10,"title":"Row: Resources","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":1},"panels":[{"id":11,"title":"CPU","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(container_cpu_usage_seconds_total{namespace=\"forgejo\"}[5m])) by (pod)","legendFormat":"{{pod}}"}]},{"id":12,"title":"Memory","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":2},"fieldConfig":{"defaults":{"unit":"bytes"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(container_memory_working_set_bytes{namespace=\"forgejo\"}) by (pod)","legendFormat":"{{pod}}"}]},{"id":13,"title":"Restarts","type":"timeseries","gridPos":{"h":8,"w":8,"x":16,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(kube_pod_container_status_restarts_total{namespace=\"forgejo\"}[15m])) by (pod)","legendFormat":"{{pod}}"}]}]},{"id":20,"title":"Row: Git Operations","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":2},"panels":[{"id":21,"title":"Repositories","type":"stat","gridPos":{"h":6,"w":6,"x":0,"y":3},"fieldConfig":{"defaults":{"unit":"short"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"forgejo_repositories_total"}]},{"id":22,"title":"Users","type":"stat","gridPos":{"h":6,"w":6,"x":6,"y":3},"fieldConfig":{"defaults":{"unit":"short"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"forgejo_users_total"}]},{"id":23,"title":"Git ops rate","type":"timeseries","gridPos":{"h":6,"w":12,"x":12,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(forgejo_git_operations_total[5m])) by (operation_type)","legendFormat":"{{operation_type}}"}]},{"id":24,"title":"Runner tasks","type":"timeseries","gridPos":{"h":6,"w":12,"x":0,"y":9},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(forgejo_runner_tasks_total[5m])) by (status)","legendFormat":"{{status}}"}]}]},{"id":30,"title":"Row: Logs","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":3},"panels":[{"id":31,"title":"Recent logs","type":"logs","gridPos":{"h":10,"w":24,"x":0,"y":4},"datasource":{"type":"loki","uid":"loki"},"targets":[{"expr":"{namespace=\"forgejo\"}"}]}]}]}
@@ -1,12 +0,0 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: svc-grafana-dashboard
namespace: logging
labels:
grafana_dashboard: "1"
annotations:
grafana_folder: "Grafana"
data:
svc-grafana.json: |
{"title":"Grafana — Service Overview","uid":"svc-grafana","schemaVersion":39,"timezone":"browser","time":{"from":"now-6h","to":"now"},"refresh":"30s","panels":[{"id":1,"title":"Row: Availability","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":0},"panels":[{"id":2,"title":"Up","type":"stat","gridPos":{"h":4,"w":6,"x":0,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"color":{"mode":"thresholds"},"mappings":[{"type":"value","options":{"0":{"text":"DOWN","color":"red"},"1":{"text":"UP","color":"green"}}}]}},"targets":[{"expr":"min(up{job=\"grafana\"})"}]},{"id":3,"title":"HTTP requests","type":"timeseries","gridPos":{"h":8,"w":9,"x":6,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(grafana_http_request_total[5m])) by (status)","legendFormat":"{{status}}"}]},{"id":4,"title":"Error rate %","type":"timeseries","gridPos":{"h":8,"w":9,"x":15,"y":1},"fieldConfig":{"defaults":{"unit":"percent"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(grafana_http_request_total{status=~\"5..\"}[5m])) / sum(rate(grafana_http_request_total[5m])) * 100"}]},{"id":5,"title":"Request latency","type":"timeseries","gridPos":{"h":8,"w":12,"x":0,"y":9},"fieldConfig":{"defaults":{"unit":"s"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"histogram_quantile(0.95, sum(rate(grafana_http_request_duration_seconds_bucket[5m])) by (le))","legendFormat":"p95"}]}]},{"id":10,"title":"Row: Resources","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":1},"panels":[{"id":11,"title":"CPU","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(container_cpu_usage_seconds_total{namespace=\"logging\",pod=~\"grafana.*\"}[5m])) by (pod)","legendFormat":"{{pod}}"}]},{"id":12,"title":"Memory","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":2},"fieldConfig":{"defaults":{"unit":"bytes"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(container_memory_working_set_bytes{namespace=\"logging\",pod=~\"grafana.*\"}) by (pod)","legendFormat":"{{pod}}"}]},{"id":13,"title":"Restarts","type":"timeseries","gridPos":{"h":8,"w":8,"x":16,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(kube_pod_container_status_restarts_total{namespace=\"logging\",pod=~\"grafana.*\"}[15m])) by (pod)","legendFormat":"{{pod}}"}]}]},{"id":20,"title":"Row: Dashboards & Users","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":2},"panels":[{"id":21,"title":"Total dashboards","type":"stat","gridPos":{"h":6,"w":6,"x":0,"y":3},"fieldConfig":{"defaults":{"unit":"short"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"grafana_dashboard_total"}]},{"id":22,"title":"Total users","type":"stat","gridPos":{"h":6,"w":6,"x":6,"y":3},"fieldConfig":{"defaults":{"unit":"short"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"grafana_user_total"}]},{"id":23,"title":"Total alerts","type":"stat","gridPos":{"h":6,"w":6,"x":12,"y":3},"fieldConfig":{"defaults":{"unit":"short"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"grafana_alerts_total"}]}]},{"id":30,"title":"Row: Logs","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":3},"panels":[{"id":31,"title":"Recent logs","type":"logs","gridPos":{"h":10,"w":24,"x":0,"y":4},"datasource":{"type":"loki","uid":"loki"},"targets":[{"expr":"{namespace=\"logging\",container=\"grafana\"}"}]}]}]}
@@ -1,143 +0,0 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: svc-minio-dashboard
namespace: logging
labels:
grafana_dashboard: "1"
annotations:
grafana_folder: "MinIO"
data:
svc-minio.json: |
{
"title": "MinIO — Service Overview",
"uid": "svc-minio",
"schemaVersion": 39,
"timezone": "browser",
"time": { "from": "now-6h", "to": "now" },
"refresh": "30s",
"panels": [
{
"id": 1, "title": "Row: Availability & Golden Signals", "type": "row",
"collapsed": true, "gridPos": { "h": 1, "w": 24, "x": 0, "y": 0 },
"panels": [
{
"id": 2, "title": "Up", "type": "stat",
"gridPos": { "h": 4, "w": 6, "x": 0, "y": 1 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": {
"defaults": {
"color": { "mode": "thresholds" },
"mappings": [
{ "type": "value", "options": { "0": { "text": "DOWN", "color": "red" }, "1": { "text": "UP", "color": "green" } } }
],
"thresholds": { "mode": "absolute", "steps": [ { "value": null, "color": "red" }, { "value": 1, "color": "green" } ] }
}
},
"targets": [{ "expr": "min(up{job=\"minio\"})" }]
},
{
"id": 3, "title": "S3 request rate by method", "type": "timeseries",
"gridPos": { "h": 8, "w": 9, "x": 6, "y": 1 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [{ "expr": "sum(rate(minio_s3_requests_total[5m])) by (method)", "legendFormat": "{{method}}" }]
},
{
"id": 4, "title": "Error rate %", "type": "timeseries",
"gridPos": { "h": 8, "w": 9, "x": 15, "y": 1 },
"fieldConfig": { "defaults": { "unit": "percent" } },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [{ "expr": "sum(rate(minio_s3_requests_total{error=\"true\"}[5m])) / sum(rate(minio_s3_requests_total[5m])) * 100" }]
},
{
"id": 5, "title": "Request duration p50/p95/p99", "type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 9 },
"fieldConfig": { "defaults": { "unit": "s" } },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{ "expr": "histogram_quantile(0.50, sum(rate(minio_s3_requests_duration_seconds_bucket[5m])) by (le))", "legendFormat": "p50" },
{ "expr": "histogram_quantile(0.95, sum(rate(minio_s3_requests_duration_seconds_bucket[5m])) by (le))", "legendFormat": "p95" },
{ "expr": "histogram_quantile(0.99, sum(rate(minio_s3_requests_duration_seconds_bucket[5m])) by (le))", "legendFormat": "p99" }
]
}
]
},
{
"id": 10, "title": "Row: Resource Usage", "type": "row",
"collapsed": true, "gridPos": { "h": 1, "w": 24, "x": 0, "y": 1 },
"panels": [
{
"id": 11, "title": "CPU by pod", "type": "timeseries",
"gridPos": { "h": 8, "w": 8, "x": 0, "y": 2 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [{ "expr": "sum(rate(container_cpu_usage_seconds_total{namespace=\"storage\",pod=~\"minio.*\"}[5m])) by (pod)", "legendFormat": "{{pod}}" }]
},
{
"id": 12, "title": "Memory by pod", "type": "timeseries",
"gridPos": { "h": 8, "w": 8, "x": 8, "y": 2 },
"fieldConfig": { "defaults": { "unit": "bytes" } },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [{ "expr": "sum(container_memory_working_set_bytes{namespace=\"storage\",pod=~\"minio.*\"}) by (pod)", "legendFormat": "{{pod}}" }]
},
{
"id": 13, "title": "Restart rate by pod", "type": "timeseries",
"gridPos": { "h": 8, "w": 8, "x": 16, "y": 2 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [{ "expr": "sum(rate(kube_pod_container_status_restarts_total{namespace=\"storage\",pod=~\"minio.*\"}[15m])) by (pod)", "legendFormat": "{{pod}}" }]
}
]
},
{
"id": 20, "title": "Row: Storage & Replication", "type": "row",
"collapsed": true, "gridPos": { "h": 1, "w": 24, "x": 0, "y": 2 },
"panels": [
{
"id": 21, "title": "Usable vs Raw capacity", "type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 3 },
"fieldConfig": { "defaults": { "unit": "bytes", "custom": { "lineWidth": 2 } } },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{ "expr": "minio_cluster_capacity_usable_bytes", "legendFormat": "Usable" },
{ "expr": "minio_cluster_capacity_raw_total_bytes", "legendFormat": "Raw Total" }
]
},
{
"id": 22, "title": "Drive health (online/offline)", "type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 3 },
"fieldConfig": { "defaults": { "unit": "short" } },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{ "expr": "minio_cluster_health_drives_online", "legendFormat": "Online" },
{ "expr": "minio_cluster_health_drives_offline", "legendFormat": "Offline" }
]
},
{
"id": 23, "title": "Replication lag (bytes pending)", "type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 11 },
"fieldConfig": { "defaults": { "unit": "bytes" } },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [{ "expr": "minio_replication_metrics_replicating_byte_count", "legendFormat": "Pending replication" }]
},
{
"id": 24, "title": "Replication failures (bytes)", "type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 11 },
"fieldConfig": { "defaults": { "unit": "bytes" } },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [{ "expr": "minio_replication_metrics_failed_byte_count", "legendFormat": "Failed replication" }]
}
]
},
{
"id": 30, "title": "Row: Logs", "type": "row",
"collapsed": true, "gridPos": { "h": 1, "w": 24, "x": 0, "y": 3 },
"panels": [
{
"id": 31, "title": "Recent logs", "type": "logs",
"gridPos": { "h": 10, "w": 24, "x": 0, "y": 4 },
"datasource": { "type": "loki", "uid": "loki" },
"targets": [{ "expr": "{namespace=\"storage\"}" }]
}
]
}
]
}
@@ -1,12 +0,0 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: svc-vault-dashboard
namespace: logging
labels:
grafana_dashboard: "1"
annotations:
grafana_folder: "Vault"
data:
svc-vault.json: |
{"title":"Vault — Service Overview","uid":"svc-vault","schemaVersion":39,"timezone":"browser","time":{"from":"now-6h","to":"now"},"refresh":"30s","panels":[{"id":1,"title":"Row: Availability & Golden Signals","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":0},"panels":[{"id":2,"title":"Up","type":"stat","gridPos":{"h":4,"w":6,"x":0,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"color":{"mode":"thresholds"},"mappings":[{"type":"value","options":{"0":{"text":"DOWN","color":"red"},"1":{"text":"UP","color":"green"}}}],"thresholds":{"mode":"absolute","steps":[{"value":null,"color":"red"},{"value":1,"color":"green"}]}}},"targets":[{"expr":"min(up{job=\"vault\"})"}]},{"id":3,"title":"Request rate by status","type":"timeseries","gridPos":{"h":8,"w":9,"x":6,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(vault_core_handle_request_total[5m])) by (method)","legendFormat":"{{method}}"}]},{"id":4,"title":"Error rate %","type":"timeseries","gridPos":{"h":8,"w":9,"x":15,"y":1},"fieldConfig":{"defaults":{"unit":"percent"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(vault_core_handle_request_total{error=\"true\"}[5m])) / sum(rate(vault_core_handle_request_total[5m])) * 100"}]},{"id":5,"title":"Request duration p50/p95/p99","type":"timeseries","gridPos":{"h":8,"w":12,"x":0,"y":9},"fieldConfig":{"defaults":{"unit":"s"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"histogram_quantile(0.50, sum(rate(vault_core_handle_request_duration_seconds_bucket[5m])) by (le))","legendFormat":"p50"},{"expr":"histogram_quantile(0.95, sum(rate(vault_core_handle_request_duration_seconds_bucket[5m])) by (le))","legendFormat":"p95"},{"expr":"histogram_quantile(0.99, sum(rate(vault_core_handle_request_duration_seconds_bucket[5m])) by (le))","legendFormat":"p99"}]}]},{"id":10,"title":"Row: Resource Usage","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":1},"panels":[{"id":11,"title":"CPU by pod","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(container_cpu_usage_seconds_total{namespace=\"iam\",pod=~\"vault.*\"}[5m])) by (pod)","legendFormat":"{{pod}}"}]},{"id":12,"title":"Memory by pod","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":2},"fieldConfig":{"defaults":{"unit":"bytes"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(container_memory_working_set_bytes{namespace=\"iam\",pod=~\"vault.*\"}) by (pod)","legendFormat":"{{pod}}"}]},{"id":13,"title":"Restart rate","type":"timeseries","gridPos":{"h":8,"w":8,"x":16,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"sum(rate(kube_pod_container_status_restarts_total{namespace=\"iam\",pod=~\"vault.*\"}[15m])) by (pod)","legendFormat":"{{pod}}"}]}]},{"id":20,"title":"Row: Vault Seal State","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":2},"panels":[{"id":21,"title":"Sealed","type":"stat","gridPos":{"h":6,"w":6,"x":0,"y":3},"fieldConfig":{"defaults":{"color":{"mode":"thresholds"},"mappings":[{"type":"value","options":{"0":{"text":"UNSEALED","color":"green"},"1":{"text":"SEALED","color":"red"}}}],"thresholds":{"mode":"absolute","steps":[{"value":null,"color":"green"},{"value":1,"color":"red"}]}}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"vault_core_unsealed"}]},{"id":22,"title":"Active","type":"stat","gridPos":{"h":6,"w":6,"x":6,"y":3},"fieldConfig":{"defaults":{"color":{"mode":"thresholds"},"mappings":[{"type":"value","options":{"0":{"text":"INACTIVE","color":"red"},"1":{"text":"ACTIVE","color":"green"}}}],"thresholds":{"mode":"absolute","steps":[{"value":null,"color":"red"},{"value":1,"color":"green"}]}}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"vault_core_active"}]},{"id":23,"title":"Replication (Primary)","type":"stat","gridPos":{"h":6,"w":6,"x":12,"y":3},"fieldConfig":{"defaults":{"color":{"mode":"thresholds"},"mappings":[{"type":"value","options":{"0":{"text":"SECONDARY","color":"orange"},"1":{"text":"PRIMARY","color":"green"}}}],"thresholds":{"mode":"absolute","steps":[{"value":null,"color":"orange"},{"value":1,"color":"green"}]}}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"vault_core_replication_primary"}]},{"id":24,"title":"Active tokens","type":"stat","gridPos":{"h":6,"w":6,"x":18,"y":3},"fieldConfig":{"defaults":{"unit":"short"}},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"vault_token_total"}]}]},{"id":30,"title":"Row: Logs","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":3},"panels":[{"id":31,"title":"Recent logs","type":"logs","gridPos":{"h":10,"w":24,"x":0,"y":4},"datasource":{"type":"loki","uid":"loki"},"targets":[{"expr":"{namespace=\"iam\",container=\"vault\"}"}]}]}]}
+5 -13
View File
@@ -12,20 +12,12 @@ resources:
- alerts/svc-grafana-rules.yaml - alerts/svc-grafana-rules.yaml
- alerts/svc-minio-rules.yaml - alerts/svc-minio-rules.yaml
- alerts/svc-vault-rules.yaml - alerts/svc-vault-rules.yaml
- alerts/cluster-alerts.yaml
- alerts/api-gateway-alerts.yaml
- servicemonitors/argocd.yaml - servicemonitors/argocd.yaml
- servicemonitors/authentik.yaml - servicemonitors/authentik.yaml
- servicemonitors/forgejo.yaml - servicemonitors/forgejo.yaml
- servicemonitors/minio.yaml - servicemonitors/minio.yaml
- dashboards/control-plane-logs.yaml - servicemonitors/ingress-nginx.yaml
- dashboards/hardware-overview.yaml - dashboards/cluster-infrastructure.yaml
- dashboards/kube-controller-health.yaml - dashboards/api-gateway.yaml
- dashboards/llm-frontend.yaml
- dashboards/service-availability.yaml
- dashboards/service-golden-signals.yaml
- dashboards/service-internals.yaml
- dashboards/svc-argocd.yaml
- dashboards/svc-authentik.yaml
- dashboards/svc-forgejo.yaml
- dashboards/svc-grafana.yaml
- dashboards/svc-minio.yaml
- dashboards/svc-vault.yaml
@@ -0,0 +1,18 @@
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: ingress-nginx
namespace: monitoring
labels:
release: prometheus
spec:
namespaceSelector:
matchNames:
- ingress-nginx
selector:
matchLabels:
app.kubernetes.io/name: ingress-nginx
app.kubernetes.io/component: controller
endpoints:
- port: metrics
interval: 30s
+15
View File
@@ -0,0 +1,15 @@
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: homelab-admin-oidc
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: cluster-admin
subjects:
# OIDC group for Authentik homelab-admins members
# When rock logs in via OIDC, k8s sees:
# - User: oidc:[email protected]
# - Groups: oidc:homelab-admins (from Authentik group claim)
- kind: Group
name: oidc:homelab-admins
+1
View File
@@ -6,6 +6,7 @@ kind: Kustomization
# already fixed once in k8s/infra/minio and k8s/infra/iam. Every resource # already fixed once in k8s/infra/minio and k8s/infra/iam. Every resource
# here sets its own explicit metadata.namespace. # here sets its own explicit metadata.namespace.
resources: resources:
- admin-oidc-binding.yaml
- grafana-operator-role.yaml - grafana-operator-role.yaml
- minio-operator-role.yaml - minio-operator-role.yaml
- forgejo-operator-role.yaml - forgejo-operator-role.yaml
+2
View File
@@ -35,6 +35,8 @@
rewrite name longhorn.riotpiao.com ingress-nginx-controller.ingress-nginx.svc.cluster.local rewrite name longhorn.riotpiao.com ingress-nginx-controller.ingress-nginx.svc.cluster.local
rewrite name paperless.riotpiao.com ingress-nginx-controller.ingress-nginx.svc.cluster.local rewrite name paperless.riotpiao.com ingress-nginx-controller.ingress-nginx.svc.cluster.local
rewrite name img.riotpiao.com ingress-nginx-controller.ingress-nginx.svc.cluster.local rewrite name img.riotpiao.com ingress-nginx-controller.ingress-nginx.svc.cluster.local
rewrite name api.riotpiao.com ingress-nginx-controller.ingress-nginx.svc.cluster.local
rewrite name comfy.riotpiao.com ingress-nginx-controller.ingress-nginx.svc.cluster.local
rewrite name riotpiao.com ingress-nginx-controller.ingress-nginx.svc.cluster.local rewrite name riotpiao.com ingress-nginx-controller.ingress-nginx.svc.cluster.local
kubernetes cluster.local in-addr.arpa ip6.arpa { kubernetes cluster.local in-addr.arpa ip6.arpa {