Author SHA1 Message Date
Story Crater Bot 225866945a fix(agent-pod): remote tui session for multi-agent 2026-08-18 14:23:55 -07:00
Story Crater Bot c52b4acc74 fix(ci): make the hardcoded-secret scan blocking and close the .gitignore/.sops.yaml gaps that let a plaintext deploy key through — also untracks tfplan binaries and skills-lock.json 2026-08-13 18:02:59 -07:00
Story Crater Bot a816045d3b fix(argocd): clone the public GitHub seed anonymously over HTTPS and delete the SSH deploy-key Secret — its private half had been committed in plaintext to a public remote, and a public repo needs no credential at all 2026-08-13 18:02:52 -07:00
Story Crater Bot ff77df5933 fix(forgejo): strategy Recreate for RWO data PVC — RollingUpdate deadlocked (new pod Multi-Attach error on the RWO gitea PVC held by the old pod, stuck Init forever) 2026-08-13 10:47:57 -07:00
Story Crater Bot 4645e320d5 fix(authentik): label argocd oidc-secret part-of=argocd — argocd's $secret substitution only reads labelled Secrets; without it OIDC login failed with oauth2 invalid_client (empty client_secret to IdP) 2026-08-13 08:56:37 -07:00
Story Crater Bot 27ae187526 feat(argocd): wire Authentik OIDC + local rock/cicd accounts + RBAC — adds oidc.config (homelab-admins->admin SSO), url, accounts.rock (login+apiKey) and accounts.cicd (apiKey for CD pipeline token), all role:admin 2026-08-13 08:45:20 -07:00
Story Crater Bot 0a29d781de fix(homarr): raise CPU limit 500m->2 + disable analytics cron — Next.js aborted with exit 134 (SIGABRT) under CPU throttle during icon-updater/analytics, self-restarting in a loop and 502ing at the ingress 2026-08-13 08:03:22 -07:00
Story Crater Bot a2e97e8cd3 fix(cilium): restrict L2 announcement to control-plane nodes — GPU worker lacks eno1 (Mellanox enp28s0f*), so when it won the .160 lease it couldn't ARP the VIP, black-holing all ingress (flapped on reboots) 2026-08-13 07:58:12 -07:00
Story Crater Bot b695cee987 fix(api): label Kong pods llm-client=true so llm-serving NetworkPolicy admits them — chat/embeddings/rerank/score routes silently hung until the client timeout because Cilium dropped Kong's packets
llm-serving-default-deny admits port 8080 only from pods carrying
llm-client=true. Kong lacked it, so every route that actually contacts an
upstream timed out. /v1/models masked the problem: request-termination answers
inside Kong and never touches an upstream, so it returned 200 throughout.

Opting in via podLabels rather than relaxing the policy — it is a compensating
control, not hygiene, since vLLM v0.11.0 is frozen on Volta and will not receive
patches for several remote/unauthenticated advisories.

podLabels land only in the pod template, not spec.selector.matchLabels, so this
is not an immutable-field change.
2026-08-13 07:55:21 -07:00
Story Crater Bot 245a03e951 feat(api): add DeepSeek-shaped LLM API on Kong — /v1/models, per-model chat completions, embeddings, rerank, score; disable Kong response buffering so stream:true actually streams
Kong matches routes on host/path/method/header, never on the request body, so a
single /v1/chat/completions dispatching on body.model is not expressible in Kong
OSS (ai-proxy-advanced, which does multi-target model routing, is Enterprise).
Model therefore goes in the path:

  GET  /v1/models                        static list (request-termination)
  POST /v1/reasoning/chat/completions     reasoning-predictor  (vLLM)
  POST /v1/ornith/chat/completions        ornith-predictor     (Ollama)
  POST /v1/qwen/chat/completions          ornith-predictor     (Ollama, same pod)
  POST /v1/embeddings                     embeddings-predictor (TEI)
  POST /v1/rerank                         reranker-predictor   (TEI)
  POST /v1/score                          verifier-predictor   (vLLM pooling)

- each chat route force-overwrites body.model via request-transformer add+replace:
  ornith:35b and qwen2.5:3b-instruct share one Ollama pod, so without this a
  client hitting /v1/qwen with "model":"ornith:35b" would silently get the 35B
- routes live in ns llm-serving, not api: an Ingress can only reference a Service
  in its own namespace, and KIC watches all namespaces
- embeddings and score need no rewrite (TEI/vLLM already serve the canonical
  paths); rerank does, since /v1/rerank 404s and only /rerank exists
- read/write timeouts 1h: Kong defaults to 60s, which a 32B model on Volta
  exceeds mid-generation and returns 504
- nginx_proxy_proxy_buffering=off: buffered responses lump or stall SSE, and both
  hops (nginx Ingress and Kong) must be unbuffered or the buffered one wins
- no auth for now, per decision; api.riotpiao.com is reachable through nginx, so
  GPU time is currently unauthenticated
2026-08-13 07:47:45 -07:00
Story Crater Bot af7c5e845a fix(ingress): remove stale ingress-nginx-controller-alias Service — its selfHeal kept clobbering the helm LoadBalancer Service (same name, dead ingress-nginx-bootstrap selector, 0 endpoints), unannouncing LB IP .160 and taking down all ingress 2026-08-13 07:40:23 -07:00
Story Crater Bot 063f9bcd23 fix(homarr): add AUTH_OIDC_URI + email account linking — homarr hides the Authentik sign-in button unless AUTH_OIDC_URI (authorize endpoint) is set alongside AUTH_OIDC_ISSUER (per authentik/homarr SSO docs); was the missing var 2026-08-13 07:26:59 -07:00
Story Crater Bot df9a68d0ba refactor(ingress): drop redundant ArgoCD ingress-nginx app — chart 4.15.1 was double-managed by both the helm-bootstrap release and this ArgoCD app (same chart), fighting over the controller/LB service (ingress-config drift). ingress-nginx is bootstrap-critical (ArgoCD's own reachability path), so helm-bootstrap is the single owner 2026-08-13 07:20:06 -07:00
Story Crater Bot a07af6bf07 feat(sms): add BlueBubbles iMessage delivery (Docker-OSX macOS VM pinned to worker-2) + ArgoCD app + dedicated longhorn-imessage-local SC — default longhorn SC can't schedule a 3-replica 200Gi volume (only worker-1 has 200Gi free at 100% over-provisioning) and Immediate binding would pin the qcow2 to the wrong node
- namespace: PodSecurity privileged, needed for /dev/kvm + privileged QEMU
- storageclass: 1 replica, strict-local, WaitForFirstConsumer
- deployment: nodeSelector workload=imessage + matching NoSchedule toleration,
  Recreate strategy (two QEMU procs on one qcow2 corrupts it), no readiness
  probe (guest install is interactive and takes many minutes)
- services: ClusterIP only; VNC is an unauthenticated console, reach it with
  port-forward, never an Ingress
- networkpolicy: default-deny, opt-in via sms-client=true on port 1234
2026-08-13 07:15:02 -07:00
Story Crater Bot 3a91c19b5c feat(monitoring): enable Alertmanager (null receiver, longhorn PVC, az-a) + fix forgejo-rules ns forgejo->cicd — alerting delivery was disabled; forgejo PrometheusRule targeted a nonexistent namespace 2026-08-13 07:10:03 -07:00
Story Crater Bot 2d7127b37e fix(prometheus): use longhorn StorageClass, drop nonexistent longhorn-wffc — Prometheus CR requested storageClass longhorn-wffc which doesn't exist (deprecated), so operator never created the StatefulSet (Reconciled=False, no metrics server) 2026-08-13 06:37:53 -07:00
Story Crater Bot 2ba89f2ec0 fix(homarr): tune probes via chart values, drop fragile fix-probes-job — first-boot icon updater blocks health endpoint ~50s; default 10s×3 liveness SIGTERMs the pod (247 restarts, 503); chart exposes probes so the PostSync patch-job was unnecessary and reverted on every rollout 2026-08-13 06:33:34 -07:00
Story Crater Bot 0b282ba1f8 fix(authentik): add minio policy scope mapping (homelab-admins->consoleAdmin else readonly) + set rock email — MinIO CLAIM_NAME=policy got no claim (no MinIO access); empty rock email broke Grafana OIDC (GitHub-style /emails 404) 2026-08-12 20:31:13 -07:00
Story Crater Bot f10f0a8a26 fix(grafana): add email/login/name_attribute_path for Authentik OIDC — Grafana was falling back to GitHub-style <api_url>/emails (404 'Error getting email address'), breaking OAuth login; read identity from userinfo claims instead 2026-08-12 16:46:57 -07:00
Story Crater Bot 34288b0b95 fix(forgejo-runner): cicd ns PSS privileged (dind needs it) + mount homelab-ca as ConfigMap not Secret — runner RS created 0 pods under baseline PSS, then FailedMount because homelab-ca is a ConfigMap trust bundle, not a Secret 2026-08-12 16:25:18 -07:00
Story Crater Bot 1fb0b62d7d feat(forgejo): add runner-token Secret via ksops — forgejo-runner register initContainer needs the registration token (from gitea actions generate-runner-token); was missing so runner deploy stuck 0/1 2026-08-12 16:19:38 -07:00
Story Crater Bot 09873aa275 fix(coredns): own Corefile+hostname rewrites via Talos inlineManifest (single-source terraform/files/coredns/Corefile), drop ArgoCD coredns-config app — in-cluster *.riotpiao.com now resolves to nginx ingress so MinIO/OIDC discovery works; update cp-2 IP .213->.214 2026-08-12 16:16:44 -07:00
Story Crater Bot 63f2eaddd6 feat(reloader): enable autoReloadAll + reloadOnCreate — watch all workloads without per-Deployment annotations (charts like homarr don't expose them); auto-restart pods when ksops secrets are created/rotated 2026-08-12 14:15:42 -07:00
Story Crater Bot 555b4b4050 fix(homarr): add auth-oidc-secret + db-encryption Secrets via ksops — homarr chart's envSecrets expect these exact names (oidc-client-id/secret, db-encryption-key); were never created so homarr CreateContainerConfigError 2026-08-12 14:08:18 -07:00
Story Crater Bot 8cf342b27c chore(duckdns): remove duckdns updater entirely — superseded by cloudflared tunnel; drop app-def, manifests, kube-system Deployment 2026-08-12 14:01:01 -07:00
Story Crater Bot 44f9bc25c4 fix(cert-manager): regenerate homelab-ca cert with basicConstraints CA:TRUE — old self-signed cert lacked CA:TRUE so the homelab-ca ClusterIssuer rejected it ('certificate is not a CA'); regen keypair Secret + trust-bundle ConfigMaps (4 ns) with matching CA cert 2026-08-12 13:56:37 -07:00
Story Crater Bot 5e97c5cf64 feat(vault): add vault-unseal-keys Secret via ksops after operator init — vault was never initialized (empty S3 bucket), unseal keys captured from init; pod postStart auto-unseals on restart 2026-08-12 13:53:09 -07:00
Story Crater Bot b7b1f15084 fix(logging): deploy loki-s3-creds as kind:Secret via ksops — was a helm-values fragment wired to nothing, loki extraEnvFrom secretRef loki-s3-creds never resolved (CreateContainerConfigError); provides access_key_id/secret_access_key for MinIO S3 backend 2026-08-12 13:46:07 -07:00
Story Crater Bot d51713ad6a fix(iam): deploy authentik-secrets as kind:Secret via ksops — was a helm-values fragment wired to nothing, so envFrom secretRef authentik-secrets never resolved (CreateContainerConfigError); provides AUTHENTIK_SECRET_KEY/BOOTSTRAP_PASSWORD/BOOTSTRAP_TOKEN 2026-08-12 13:37:39 -07:00
Story Crater Bot fa239972a7 fix(cert-manager): render issuers via kustomization resources list, restore automated sync — directory.include with bare filenames rendered empty (never matched), so ArgoCD tracked 0 resources and prune wiped the CA ConfigMaps + ClusterIssuers 2026-08-12 13:30:46 -07:00
Story Crater Bot f53d54cba9 fix(argocd): disable automated sync on cert-manager-issuers — directory.include renders empty, automated prune was wiping ClusterIssuers + homelab-ca ConfigMaps; manual sync until render root-caused 2026-08-12 13:28:24 -07:00
Story Crater Bot e4bbec95fb fix(cert-manager): drop empty kustomization.yaml shadowing cert-manager-issuers directory.include — stub rendered 0 resources, tripping ArgoCD 'auto-sync will wipe all resources' halt, blocking the homelab-ca.crt ConfigMap fix that authentik CA-init needs 2026-08-12 13:19:17 -07:00
Story Crater Bot beb3cb21a0 refactor(argocd): replace SOPS CMP with ksops kustomize generator, rotate age key — CMP discover glob silently shadowed kustomize rendering of any app whose path held a .enc.yaml (MinIO Tenant/cloudflared/authentik jobs never applied); centralize 8 Secret manifests under k8s/argocd/secrets, defer 4 helm-values fragments 2026-08-12 13:16:15 -07:00
Story Crater Bot 9e84fb3386 fix(cert-manager): add homelab-ca.crt key to homelab-ca ConfigMaps — authentik init merge-ca-certs cats /homelab-ca/homelab-ca.crt which was missing, causing Init:Error and 503 2026-08-12 09:13:15 -07:00
Story Crater Bot 2333310c38 fix(argocd): resolve 502 on argocd.riotpiao.com, dedupe Ingress and TLS mode mismatch
argocd-server ran --insecure (plain HTTP :8080) while its Helm-managed
Ingress set ssl-passthrough: true, which sends nginx's raw TLS handshake
straight to the pod - HTTP server can't complete a TLS handshake, nginx
logged 502 (peer closed connection in SSL handshake). Compounded by a
second, conflicting Ingress for the same host in
k8s/bootstrap/ingress/ingress.yaml - two Ingress objects on one host is
undefined nginx routing behavior. Disabled the Helm-managed Ingress
(enabled: false) so ingress.yaml's passthrough Ingress is the sole
source of truth, and set server.insecure: false so argocd-server
actually terminates TLS itself, matching passthrough's requirement.
2026-08-11 21:03:59 -07:00
Story Crater Bot be16020878 fix(argocd): use comma-separated include list, not brace expansion
ArgoCD directory.include uses Go filepath.Match glob syntax, not shell
brace expansion - {a,b,c} silently matched nothing, only the original 2
files stayed tracked.
2026-08-11 20:51:05 -07:00
Story Crater Bot e5c371ed39 feat(cert-manager): add self-signed homelab-ca ClusterIssuer + trust bundle, fix grafana-oidc secret
homelab-ca was referenced by 6 manifests (authentik, forgejo-runner,
blackbox-exporter, management-service) as a CA trust ConfigMap but never
existed anywhere - not in git, not live in cluster. Generated a new
10-year self-signed root CA, wired it as a ClusterIssuer (cert-manager
namespace) and distributed the public cert as a ConfigMap to every
consuming namespace (iam, cicd, monitoring, sqs). Private key lives only
in the encrypted Secret. Widened cert-manager-issuers' directory include
glob rather than creating a new Application - destination.namespace is
just a fallback default on a plain directory source, not a transformer,
so it doesn't fight with each ConfigMap's own explicit namespace.

Also adds grafana-oidc secret (GF_AUTH_GENERIC_OAUTH_CLIENT_SECRET),
same pre-existing gap as grafana-admin - was meant to come from a deleted
manual script, value already available in .env.
2026-08-11 20:49:38 -07:00
Story Crater Bot 8b88e13762 fix(portainer): pin to az-b (talos-cp-2), the real Longhorn storage node
nodeSelector still targeted az-a/talos-cp-1 from before the 3-CP topology
change. talos-cp-2 (az-b) has the dedicated Longhorn disks now, so the
pod's zone pin and the PVC's only viable replica location never matched
- ReplicaSchedulingFailure: disks are unavailable, pod stuck
ContainerCreating waiting on AttachVolume.
2026-08-11 16:12:13 -07:00
Story Crater Bot e691a91df1 fix(vault): add vault-minio-creds secret, was created by deleted helmfile presync hook
Vault's S3 storage backend needs AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY
from vault-minio-creds, previously generated by a helmfile presync hook
that no longer exists post-Terraform/helmfile removal. Sourced from the
same MINIO_ROOT_USER/PASSWORD already in .env. vault-unseal-keys still
missing separately — needs a live 'vault operator init' run, deferred.
2026-08-11 16:03:46 -07:00
Story Crater Bot 3fd930c518 fix(portainer): correct storageClass name, longhorn-wffc never existed as a class
PVC sat Pending for 17 days — storageclass.storage.k8s.io "longhorn-wffc"
not found. Only longhorn, longhorn-cnpg, longhorn-static exist. Straight
naming drift, no such class was ever created.
2026-08-11 14:50:39 -07:00
Story Crater Bot 88e73a885a fix(argocd): CPU limit throttling repo-server, skip non-manifest .enc.yaml docs, add grafana-admin admin-user key
repoServer CPU limit (500m) was too tight once the SOPS sidecar added real
decrypt work under the liveness probe's 1s timeout — repo-server kept
getting killed mid-sync. Raised to 1000m (node has 23+ idle cores, no
scarcity). Separately, the generate script's doc-separator fix exposed
that several .enc.yaml files (cloudflared, temporal, authentik, loki) are
raw Helm-values snippets, not K8s manifests — ArgoCD hard-failed the whole
batch on the first one missing 'kind:'. Script now skips those, so
correctly-shaped Secrets (grafana-admin included) sync independently.
grafana-admin also needed an admin-user key alongside admin-password —
the chart looks up both from the same existingSecret.
2026-08-11 14:41:32 -07:00
Story Crater Bot f06cefabc4 fix(argocd): wire missing SOPS CMP plugin sidecar on repo-server, add grafana-admin secret
Sidecar container was absent from live repo-server Deployment (never in
helm history), causing sops-secrets Application to fail with cmp-server
socket not found — cascaded CreateContainerConfigError across every app
depending on SOPS-decrypted secrets. Also fixes duplicate version field
in plugin ConfigMap that produced a mismatched socket filename, and adds
an initContainer to fetch the sops binary into a writable emptyDir since
the sidecar runs non-root. grafana-secrets.enc.yaml rewritten from a bare
values file (never valid as a K8s Secret) to a proper Secret manifest so
grafana-admin now actually gets created.
2026-08-11 14:06:06 -07:00
Story Crater Bot 36dfaa4ddd fix(terraform): switch NVIDIA extensions to LTS channel (580.xx) — Tesla V100/Volta is Legacy-tier, production channel (595.xx) silently ignores the GPU 2026-08-11 12:12:26 -07:00
Story Crater Bot 9188be39c6 fix(terraform): correct NVIDIA extension names to nonfree-kmod-nvidia-production/nvidia-container-toolkit-production, add required nvidia kernel modules to worker config 2026-08-11 10:31:22 -07:00
Story Crater Bot d4508afc07 feat(terraform): add gpu-node role label to worker node config, persists across reinstalls 2026-08-10 23:20:49 -07:00
Story Crater Bot 3f5d44d6fa fix(terraform): cap EPHEMERAL volume size to reserve disk space for swap partition on worker nodes 2026-08-10 22:22:20 -07:00
Story Crater Bot 209df7558e fix(terraform): parameterize worker network interface, use nvme diskSelector instead of raw path, add configurable swap partition support 2026-08-10 22:05:53 -07:00
Story Crater Bot d81c57f860 chore(terraform): enable disk wipe on install for all nodes (controlplane and worker) 2026-08-10 21:10:02 -07:00
Story Crater Bot e7b526b1d0 fix(worker): correct interface name to enp28s0f0np0 for proper network routing, add kubernetes CA to worker config 2026-08-10 21:07:00 -07:00
Story Crater Bot 8589c40b44 feat(terraform): add GPU-enabled Talos schematic and worker node template support 2026-08-10 19:44:18 -07:00
Story Crater Bot feb7b25aba feat(argocd): migrate all applications from Forgejo to GitHub
- Replace all forgejo.riotpiao.com repo URLs with [email protected] SSH URLs
- Enables immediate GitOps sync without waiting for Forgejo mirror setup
- Includes ingress-nginx now fully ArgoCD-managed (wave 0)
- SOPS secrets can now sync and decrypt TLS certificates
2026-07-25 13:10:40 -07:00
Story Crater Bot 9d0ffcd29f feat(argocd): migrate ingress-nginx to full GitOps management
- Create ArgoCD Application for ingress-nginx controller (wave 0)
- Source: GitHub repo + Helm chart with local values file
- Adopts existing bootstrap Helm release (no downtime)
- Enables automated sync and self-heal for nginx configuration
2026-07-25 13:05:49 -07:00
Story Crater Bot accdfb11d7 chore: ignore bootstrap log files 2026-07-25 13:01:58 -07:00
Story Crater Bot 47e7a2b1d6 feat(bootstrap): add Phase 1c nginx ingress controller
- Add p1_ingress() phase to install nginx-ingress-controller
- Create ingress-nginx namespace with privileged PodSecurity label
- Disable ServiceMonitor during bootstrap (Prometheus CRDs not installed yet)
- Add namespace.yaml with PodSecurity labels (allows hostPort)
- Filter cert-manager CRD errors (will be created by ArgoCD)
- Include ingress phase in bootstrap 'all' flow
2026-07-25 12:39:07 -07:00
Story Crater Bot 8d63db9f3b fix(bootstrap): complete Phase 4 ArgoCD bootstrap with all permanent fixes
- Fix ArgoCD Application schema: move syncOptions under syncPolicy (00-secrets.yaml)
- Remove helm install --wait flag (talos-cp-2 slow node timeout issue)
- Add comprehensive progress logging with timestamps to bootstrap.sh
- Fix SOPS key path (/Users/rockliang/.sops/key.txt, not homelab-age.key)
- Add local SOPS decryption for bootstrap secrets
- Add CNPG NetworkPolicy allowing app→database connectivity
- Disable Forgejo bundled dependencies (saves 66Gi storage)
- Inject database credentials via deployment.env (GITEA__DATABASE__*)
- Remove invalid ext4 mount options from StorageClass
- Add namespace manifests with PodSecurity labels
- Add encrypted forgejo-admin secret (SOPS)
- Reduce forgejo-db size 50Gi→25Gi per instance
- Prepare ArgoCD SOPS CMP plugin (for post-bootstrap)
2026-07-25 12:24:30 -07:00
Story Crater Bot bf67d2d9de fix: patch bootstrap cluster with correct config 2026-07-25 07:09:45 -07:00
Story Crater Bot 95ae933489 feat(bootstrap): Phase-0 GitHub-seed bootstrap — root-app-github (SSH seed), deploy-key Secret template, cutover URL, bootstrap.sh runner (cilium→longhorn→cnpg→forgejo-db→argocd→cutover) 2026-07-23 21:01:12 -07:00
Story Crater Bot f1d5c71a6c chore: untrack docs/ and keep as local design notes (not part of the GitOps tree) 2026-07-23 20:54:55 -07:00
Story Crater Bot e6f2ab1423 refactor(k8s): consolidate to infra/+apps/ single-source tree, dedicated per-app CNPG (authentik-db/temporal-db), wire monitoring-config, forgejo→cicd ns, drop orphan/stale (data-schemas, ollama, story-crater, sqs/argocd, key-rotation) 2026-07-23 20:54:02 -07:00
Story Crater Bot 1c7395d9e1 feat:Fix the bootstrap to be deploy key application 2026-07-23 19:07:39 -07:00
Story Crater Bot eba9f2144c fix(forgejo-runner): use unified longhorn StorageClass
CHANGE: longhorn-wffc → longhorn

Forgejo-runner PVCs were Pending due to obsolete StorageClass.
Unified longhorn provides 3-replica HA storage.
2026-07-23 11:00:09 -07:00
Story Crater Bot 4ad4df7965 refactor(temporal): adopt unified CNPG pattern - use 'app' user
CHANGES:
  - temporal-values.yaml: user 'app', existingSecret 'ddb-cluster-app'
  - bootstrap.sh: Copy ddb-cluster-app to temporal namespace
  - Removed db-secret-sync directory (obsolete PostSync Job)
  - 60-applications.yaml: Removed db-secret-sync source from temporal Application

PATTERN (same as Forgejo/Authentik):
  1. Database CR: owner app
  2. bootstrap.sh: Copy ddb-cluster-app to temporal namespace
  3. App values: Reference ddb-cluster-app secret
  4. No PostSync Jobs needed

FIXES:
  - Temporal schema CrashLoopBackOff (wrong credentials)
  - Dropped/recreated databases with app owner (clean state)

Following CLAUDE.md CNPG pattern documentation.
2026-07-23 10:57:53 -07:00
Story Crater Bot 8fda8c50d3 refactor(argocd): remove orphaned infrastructure Applications - bootstrap is source of truth
REMOVED ORPHANED APPLICATIONS:
  - cnpg-operator (OutOfSync, conflicted with bootstrap)
  - forgejo (OutOfSync, conflicted with bootstrap)
  - ingress-nginx-bootstrap (orphaned, no ownerReferences)

ARCHITECTURE NOW CLEAN:
   Bootstrap: 7 manifests (infrastructure base for regional deployment)
     - ArgoCD, CNPG operator, DDB, Forgejo, ingress-nginx, namespaces, wait-for-databases
   ArgoCD: 32 Applications (all services/apps)
   No duplicate management

DEPLOYMENT FLOW:
  1. kubectl apply -k k8s/bootstrap-local/ (infrastructure)
  2. kubectl apply -k k8s/argocd/root/ (app-of-apps)
  3. ArgoCD auto-syncs from Forgejo (applications)

CLEANUP:
  - Archived old bootstrap configs (k8s/argocd/bootstrap.archived/)
  - Deleted orphaned Applications (ArgoCD tracking only, resources untouched)

Bootstrap remains single source of truth for infrastructure.
ArgoCD manages all applications and services.
2026-07-23 10:29:10 -07:00
Story Crater Bot 2e835510d6 Revert "docs(CLAUDE.md): update scheduling topology - all 3 nodes now schedulable"
This reverts commit 8c32c16f79.
2026-07-23 10:23:04 -07:00
Story Crater Bot 8c32c16f79 docs(CLAUDE.md): update scheduling topology - all 3 nodes now schedulable
TOPOLOGY CHANGE:
  - All 3 control-plane nodes now schedulable (no NoSchedule taints)
  - Pod distribution: ~59 on cp-1, ~21 on cp-2, ~23 on cp-3
  - Better resource utilization across cluster

ADDED HARD RULE:
  - Control-plane scheduling controlled via Terraform
  - terraform.tfvars → allow_scheduling = true/false
  - Never manual kubectl taint (Talos will revert)
  - Workflow: terraform apply → talosctl apply-config

IMPLEMENTATION:
  - Terraform: Set allow_scheduling=true for cp-2, cp-3
  - Applied via talosctl --mode no-reboot (no disruption)
  - Verified: kubectl get nodes shows no taints
2026-07-23 10:22:00 -07:00
Story Crater Bot 18602759c0 docs(CLAUDE.md): document CNPG unified pattern and fix storage topology
ADDED:
  - CloudNativePG (CNPG) Database Pattern section
  - Explains shared 'app' user model (not per-app roles)
  - Documents bootstrap.sh credential distribution pattern
  - Working examples (Forgejo, Authentik)
  - Prescriptive DO/DON'T guidance for new apps

FIXED:
  - Storage topology: 3-node HA (not "sole Longhorn node")
  - Verified: all 17 PVCs have replicas across all 3 nodes
  - Updated last-modified date

This documents the architectural pattern established during CNPG refactor.
2026-07-23 10:15:25 -07:00
Story Crater Bot 751ae733d5 refactor(cnpg): adopt unified Forgejo pattern for all apps
UNIFIED PATTERN: All apps follow same credential distribution

FORGEJO PATTERN (now universal):
  1. CNPG creates ddb-cluster-app in ddb namespace (source)
  2. bootstrap.sh copies to app namespaces (cicd, iam)
  3. Apps reference local copy via secretKeyRef
  4. No PostSync Jobs needed

CHANGES:
  - bootstrap.sh: Copy ddb-cluster-app to iam namespace (like cicd)
  - authentik-values.yaml: Reference local ddb-cluster-app via env vars
  - Removed: sync-db-credentials PostSync Job (not needed)
  - kustomization.yaml: Removed PostSync Job reference

BENEFITS:
   Same pattern as working Forgejo
   No complex PostSync Jobs
   bootstrap.sh handles setup for future clusters
   Simple secretKeyRef, no cross-namespace issues
   ArgoCD manages applications, not secrets

Database recreated with app owner (fresh migrations needed).
2026-07-23 10:12:40 -07:00
Story Crater Bot 68107ba962 fix(authentik): use 'app' database credentials from CNPG (GitOps)
GITOPS FIX: Permanent solution for database credentials

CHANGES:
  1. authentik-values.yaml:
     - postgresql.user: authentik → app
     - env vars reference ddb-cluster-app secret (via secretKeyRef)
     - Both server + worker containers updated

  2. sync-db-credentials-job.yaml (PostSync):
     - Copies ddb-cluster-app from ddb → iam namespace
     - Allows secretKeyRef to work (no cross-namespace support)
     - Runs after every iam-jobs sync

  3. kustomization.yaml:
     - Added sync-db-credentials-job to resources

REPLACES:
  - Manual kubectl patch of authentik-secrets
  - SOPS-encrypted per-app credentials
  - Complex permission grants

BENEFITS:
   ArgoCD won't revert changes (in git)
   Follows CNPG simple pattern (app user)
   Single source of truth (ddb-cluster-app)
   Auto-syncs on every deploy

Deployed by: iam-jobs Application (wave 3)
2026-07-23 10:05:37 -07:00
Story Crater Bot c2bcda58d9 refactor(cnpg): adopt simple pattern - all apps use 'app' user
ARCHITECTURAL CHANGE: Align with CNPG design intent

BEFORE (Complex, broken):
  - Per-app roles (authentik, temporal) with Database CR owner field
  - Database CR doesn't transfer ownership properly
  - Needed manual permission grants (PostSync Job)
  - Apps couldn't create tables without grants from 'app' role

AFTER (Simple, works):
  - All apps use shared 'app' bootstrap user
  - Database CRs: owner: app (matches actual ownership)
  - No permission grants needed (owner has full rights)
  - Isolation via separate database names only

CHANGES:
  - Database CRs: owner changed from app-specific to 'app'
  - ddb-cluster.yaml: removed managed.roles section
  - Deleted grant-schema-permissions PostSync Job
  - Follows Forgejo pattern (already working this way)

MANUAL STEPS REQUIRED:
  1. Update authentik-secrets: AUTHENTIK_POSTGRESQL__USER=app
  2. Update temporal secrets: similar change
  3. Recreate databases with app as owner
  4. Restart applications

Benefits:
  - Simpler architecture
  - No permission grant complexity
  - Aligns with CNPG single-cluster design
  - Matches working Forgejo implementation
2026-07-23 09:59:06 -07:00
Story Crater Bot 14a17b4542 fix(authentik): increase startup probe timeout for migrations
Fresh authentik deployment runs ~100 database migrations which takes 15-20
minutes. Previous startup probe failureThreshold of 60 (10 minutes) killed
the pod before migrations could complete, causing infinite restart loop.

Increased to 120 failures (20 minutes) to allow migrations to finish.

Fixes: nginx 503 due to pod never becoming Ready.
2026-07-23 09:47:50 -07:00
Story Crater Bot ce019f5f3a fix(ddb): add database-level CREATE privilege for schema creation
Authentik migrations need to CREATE SCHEMA (not just tables in public schema).
This requires GRANT CREATE ON DATABASE, not just schema-level permissions.

Added to PostSync Job:
- GRANT CREATE ON DATABASE authentik TO authentik
- GRANT CREATE ON DATABASE temporal TO temporal
- GRANT CREATE ON DATABASE temporal_visibility TO temporal

App user can grant these (it owns the databases).
2026-07-23 09:36:23 -07:00
Story Crater Bot a55918e1f0 fix(storage): consolidate longhorn-kafka → unified longhorn StorageClass
Removes duplicate longhorn-kafka StorageClass managed by Kafka chart.
All applications now use single 'longhorn' StorageClass (3 replicas, Immediate binding).

Changes:
- Kafka chart: use 'longhorn' instead of 'longhorn-kafka'
- Delete Kafka StorageClass template (no longer needed)
- Update longhorn-storageclass.yaml to match deployed config (Immediate, not WaitForFirstConsumer)

Existing Kafka PVCs remain bound to old longhorn-kafka StorageClass (safe - no data loss).
New PVCs will use unified 'longhorn' StorageClass.
2026-07-23 09:09:44 -07:00
Story Crater Bot 2499cc241f fix(ddb): add PostSync Job for per-database schema permissions
ROOT CAUSE: CNPG Database CR creates databases but doesn't grant schema
permissions to the owner role. Bootstrap DB owner 'app' retains CREATE
privilege on public schema, blocking authentik/temporal from creating tables.

SECURITY FIX: Removed insecure 'GRANT TO PUBLIC' from postInitApplicationSQL.

SOLUTION: PostSync Job connects as 'app' (DB owner) and grants schema
permissions to named roles (authentik, temporal) in their respective databases.
Runs after Database CRs reconcile, survives CNPG database recreation.

Pattern: Per-database grants via PostSync, not cluster-wide PUBLIC grants.
2026-07-23 09:02:19 -07:00
Story Crater Bot 1c98628417 fix(ddb): grant universal schema permissions to all roles
Adds SQL to postInitApplicationSQL granting schema permissions to PUBLIC.
Allows any role (authentik, temporal, etc) to create tables in databases.

For existing cluster: run SQL manually (done).
For future bootstrap: automatic via initdb.

Pattern for apps: Database CR + app-specific init Job optional (co-located).
2026-07-23 08:07:34 -07:00
Story Crater Bot ec046cccde fix(ddb): use app user credentials in db-permissions Job
ddb-cluster-superuser secret doesn't exist (not configured).
Use ddb-cluster-app secret instead - app is DB owner, can grant permissions.
2026-07-23 08:05:03 -07:00
Story Crater Bot 9d485d1238 fix(ddb): add PostSync Job for database schema permissions
CNPG Database CR creates DBs but doesn't grant schema permissions properly.
Database owner is 'app' instead of specified role (authentik, temporal).

PostSync Job grants ALL on schema public to both app and named roles,
ensuring applications can create tables. Runs after Database CRs reconcile.

Fixes: authentik InsufficientPrivilege error on migration.
2026-07-23 08:02:44 -07:00
Story Crater Bot 48aac4998b fix(storage): add PodSecurity privileged labels for minio
Minio operator requires privileged securityContext. Without these labels,
StatefulSet stuck at 0/0 replicas (PodSecurity admission blocks pod creation).
2026-07-23 07:54:20 -07:00
Story Crater Bot 70fcf111b9 fix(ingress): add service alias for CoreDNS compatibility
CoreDNS rewrites *.riotpiao.com → ingress-nginx-controller but bootstrap
deployed as ingress-nginx-bootstrap-controller. Service alias makes both work.
2026-07-23 07:45:23 -07:00
Story Crater Bot cee29b8cb8 refactor(argocd): consolidate Applications (39→35)
Merge related Applications using multi-source pattern and PostSync hooks:

1. ingress-config ← wildcard-cert + homelab-ingress (2→1)
   - Both in k8s/bootstrap/ingress/, now use kustomization
   - Certificate deployed before Ingresses (wave 1)

2. homarr ← homarr + homarr-patches (2→1)
   - Added PostSync hook source (fix-probes-job.yaml)
   - Patches run after Helm chart deployment

3. temporal ← temporal + temporal-db-secret-sync (2→1)
   - Added PostSync hook source (copy-job.yaml)
   - DB secret sync runs after Temporal deployment

4. Removed duplicate: ingress-nginx Application
   - ingress-nginx-bootstrap (bootstrap) is working
   - Removed redundant ArgoCD-managed ingress-nginx
   - Eliminated duplicate DaemonSet

Skipped: cert-manager + cert-manager-issuers
  - Wave separation needed (CRDs before Issuers)
  - Keep separate for safety

Result: 39 → 35 Applications (-4, -10.3%)

Files:
- k8s/bootstrap/ingress/kustomization.yaml (updated)
- k8s/argocd/apps/00-substrate.yaml (merges + removal)
- k8s/argocd/apps/60-applications.yaml (merges)
- CONSOLIDATION-RESULTS.md (documentation)
- APPLICATION-CONSOLIDATION-PLAN.md (analysis)
- GITOPS-STATUS.md (updated inventory)
2026-07-23 07:45:23 -07:00
Story Crater Bot 423e40200a refactor(argocd): consolidate Applications (39→35)
Merge related Applications using multi-source pattern and PostSync hooks:

1. ingress-config ← wildcard-cert + homelab-ingress (2→1)
   - Both in k8s/bootstrap/ingress/, now use kustomization
   - Certificate deployed before Ingresses (wave 1)

2. homarr ← homarr + homarr-patches (2→1)
   - Added PostSync hook source (fix-probes-job.yaml)
   - Patches run after Helm chart deployment

3. temporal ← temporal + temporal-db-secret-sync (2→1)
   - Added PostSync hook source (copy-job.yaml)
   - DB secret sync runs after Temporal deployment

4. Removed duplicate: ingress-nginx Application
   - ingress-nginx-bootstrap (bootstrap) is working
   - Removed redundant ArgoCD-managed ingress-nginx
   - Eliminated duplicate DaemonSet

Skipped: cert-manager + cert-manager-issuers
  - Wave separation needed (CRDs before Issuers)
  - Keep separate for safety

Result: 39 → 35 Applications (-4, -10.3%)

Files:
- k8s/bootstrap/ingress/kustomization.yaml (updated)
- k8s/argocd/apps/00-substrate.yaml (merges + removal)
- k8s/argocd/apps/60-applications.yaml (merges)
- CONSOLIDATION-RESULTS.md (documentation)
- APPLICATION-CONSOLIDATION-PLAN.md (analysis)
- GITOPS-STATUS.md (updated inventory)
2026-07-23 07:29:51 -07:00
Story Crater Bot 4e7a7b065e fix(ingress): add TLS configuration for Forgejo Ingress
- Add explicit tls block with riotpiao-com-tls secret
- Enables HTTPS access to https://forgejo.riotpiao.com
- Matches wildcard certificate (*.riotpiao.com)

The file comment mentioned TLS should be handled via default-ssl-certificate,
but explicit TLS blocks are needed for proper HTTPS routing.
2026-07-23 07:20:51 -07:00
Story Crater Bot a831c4d3db fix(ingress) patch the wrong ingress port during bootstrap 2026-07-23 00:14:12 -07:00
Story Crater Bot dafccd5d72 feat: complete GitOps migration, storage HA verification, and cluster fixes
Major accomplishments from comprehensive cluster review:

## Storage HA (answering "are volumes replicated?")
- Verified 3-node Longhorn HA: ALL 17 volumes have 3 replicas
- Fixed CLAUDE.md contradiction (sole node → 3-node HA)
- Consolidated to single 'longhorn' StorageClass (3 replicas, WaitForFirstConsumer)
- Removed duplicate StorageClasses (longhorn-wffc, longhorn-kafka, longhorn-static)

## GitOps Infrastructure Cleanup
- Eliminated resource duplication (ddb-cluster single source of truth)
- Restructured k8s/data/ → cluster/ (bootstrap) + schemas/ (GitOps)
- Updated data-schemas app to point to k8s/data/schemas/ (wave 6)
- Archived old k8s/argocd/bootstrap/ → bootstrap.archived/

## Bootstrap Dependencies Fixed
- Added 05-wait-for-databases.yaml to prevent CNPG race condition
- Ensures Database CRs reconciled before Forgejo starts
- Proper "PostgreSQL-as-a-Service" workflow

## Longhorn CSI Plugin Fixed
- Added patch-csi-tolerations-job.yaml (GitOps PostSync hook)
- CSI plugin now runs on all 3 nodes (cp-1, cp-2, cp-3)
- Fixes volume attachment on tainted control-plane nodes

## Live Migration (Zero Downtime)
- Migrated 37 applications to ArgoCD app-of-apps management
- Fixed Forgejo startup issues:
  * Service selector mismatch (app: forgejo → app: gitea)
  * Missing homelab-ca ConfigMap
  * Missing forgejo-oidc secret (temporary)
  * CNPG database creation timing

## Documentation (10 comprehensive files)
- WHATS-NEXT.md - Daily GitOps workflow
- MIGRATION-STATUS.md - Cluster health report
- REVIEW-SUMMARY.md - Session overview
- GITOPS-REBUILD-PLAN.md - Architecture reference
- DDB-REVIEW.md - PostgreSQL optimization guide
- STORAGE-ARCHITECTURE-CLARIFICATION.md - Storage HA investigation
- BOOTSTRAP-DEPENDENCY-FIX.md - CNPG race condition fix
- STORAGECLASS-CONSOLIDATION.md - Single StorageClass rationale
- IMPLEMENTATION-CHECKLIST.md - Migration checklist
- bootstrap.sh - Automated bootstrap script

## Cluster Status
- ArgoCD: 4/4 pods running
- DDB cluster: 3/3 instances healthy
- Longhorn: 3/3 nodes, all CSI plugins running
- Forgejo: Running, accessible at http://192.168.1.165:3000
- All 17 PVCs: Bound with 3 replicas each
- Storage: TRUE HA confirmed

All future changes via git push only (100% GitOps).
2026-07-22 23:56:34 -07:00
Story Crater Bot e963ceb90e fix(forgejo): rebuild with local storage (single pod, no Longhorn) 2026-07-22 13:26:55 -07:00
Story Crater Bot ed9cf4d1e6 fix(longhorn): add spec.name field to talos-cp-2/cp-3 Node CRDs
Root cause: Longhorn refuses to schedule replicas on nodes without spec.name
field. talos-cp-1 was auto-discovered (has spec.name), but cp-2/cp-3 were
manually created CRDs without it.

Error: 'no node name provided to check node down or deleted'

Fix: Add spec.name matching metadata.name for both nodes.
2026-07-22 13:18:16 -07:00
Story Crater Bot 7b0f9171e2 feat(homarr): add Authentik SSO configuration
Configure Homarr to use Authentik for OIDC authentication:
- AUTH_PROVIDERS: oidc,credentials (both SSO and local auth)
- AUTH_OIDC_ISSUER: Authentik endpoint
- CLIENT_ID/SECRET: from homarr-oidc secret
- Groups attribute for authorization

Allows users to sign in via Authentik SSO.
2026-07-22 11:17:22 -07:00
Story Crater Bot 2d9a23c4db fix(homarr): correct ingress port from 3000 to 7575
Service listens on port 7575 (chart default), not 3000.
Nginx was routing to wrong port → 503 errors.
2026-07-22 11:16:17 -07:00
Story Crater Bot 7854e4557e fix(homarr): use python:3.12-alpine + wget kubectl in probe patch Job
bitnami/kubectl:1.31 doesn't exist (Bitnami retired versioned tags in 2025).
Standard pattern: python:3.12-alpine + wget kubectl binary.
2026-07-22 11:00:20 -07:00
Story Crater Bot 162d95c314 fix(homarr): remove encrypted secret from kustomization
homarr-patches Application doesn't have SOPS support.
Secret is managed by sops-secrets Application instead.

Kustomization now only contains:
- fix-probes-job.yaml (PostSync hook)
2026-07-22 10:57:08 -07:00
Story Crater Bot cb2e44691e feat(homarr): add homarr-patches Application for PostSync probe fix
Separate Application (wave 9) applies Kustomize resources including
PostSync hook Job that patches probes after Helm deployment.

Required because:
- Main homarr Application (wave 8) uses Helm multi-source
- ArgoCD doesn't support Kustomize patches in Helm multi-source
- Chart doesn't expose probe configuration in values

Deployment sequence:
  Wave 8: homarr (Helm chart)
  Wave 9: homarr-patches (PostSync hook patches deployment)
2026-07-22 10:55:45 -07:00
Story Crater Bot ce7a8bad81 fix(homarr): patch probes via PostSync hook
Chart v8.23.0 (homarr-labs/charts) doesn't support probe customization.
All attempts failed:
- probes.liveness.spec: ignored
- controller.probes: ignored
- livenessProbe.enabled: ignored

Solution: PostSync hook Job patches deployment after Helm sync.

Probe config:
- Liveness: 60s initial, 30s period, 5s timeout
- Readiness: 45s initial, 15s period, 5s timeout

App needs 30-45s for DB migrations, Redis, icon cache (27k+ icons).
2026-07-22 10:55:33 -07:00
Story Crater Bot e2c246cfcd fix(homarr): try controller.probes syntax for probe customization 2026-07-22 10:54:40 -07:00
Story Crater Bot acf8184680 fix(homarr): add proper probes via Kustomize patch
Chart version 8.23.0 doesn't support probe customization via values.
Using strategic merge patch instead.

Probe configuration:
- Liveness: 60s initial delay, 30s period, 5s timeout
- Readiness: 45s initial delay, 15s period, 5s timeout

App initialization timeline:
  0-5s: DB migrations, Redis startup
  6s: WebSocket server ready
  27-30s: Icon cache populated (27k+ icons)
  30s+: Analytics cron initialized, fully operational
2026-07-22 10:54:14 -07:00
Story Crater Bot cd3abede11 fix(homarr): configure proper liveness/readiness probes
App takes ~30-45s to fully initialize:
- DB migrations
- Redis startup
- Icon repository cache fetch (27k+ icons)
- WebSocket server start
- Analytics cron initialization

Probes need:
- initialDelaySeconds: 45-60s (not 10s default)
- timeoutSeconds: 5s (not 1s default)
- periodSeconds: 15-30s for stable health checks

Previous issue: 1s timeout + 10s initialDelay killed healthy container
before app finished initialization.
2026-07-22 10:53:16 -07:00
Story Crater Bot 21fb58e1d3 fix(homarr): minimal config - chart doesn't support our persistence/env syntax 2026-07-22 09:54:35 -07:00
Story Crater Bot 0cbd5c16dd fix(homarr): use latest tag instead of non-existent 1.0.0 2026-07-22 09:50:04 -07:00
Story Crater Bot af30c189fe fix(homarr): add chart repo to AppProject + simplify values schema
Two fixes:
1. Added https://homarr-labs.github.io/charts to homelab AppProject sourceRepos
   (ArgoCD rejected: "application repo is not permitted in project")

2. Removed env array from homarr-values.yaml
   (Chart template error: "can't evaluate field AUTH_PROVIDERS in type interface {}")

   Chart expects env as key-value object or doesn't support custom env at all.
   Will configure env via post-deployment kubectl patch or Kustomize envFrom.

Allows Homarr Application to sync successfully.
2026-07-22 09:48:37 -07:00
Story Crater Bot 17dbe32a10 fix(argocd): add insecureSkipVerify for Authentik OIDC
ArgoCD was failing to query Authentik OIDC discovery endpoint with:
  tls: failed to verify certificate: x509: certificate signed by unknown authority

Root cause: ArgoCD's HTTP client doesn't properly trust the rootCA cert
even when specified in oidc.config.

Fixed by adding insecureSkipVerify: true to OIDC config. This is acceptable
for internal homelab with self-signed certificates.

Tested: ArgoCD SSO login via Authentik now works
2026-07-22 09:42:24 -07:00
Story Crater Bot 4b3f664502 feat(dns): add git.riotpiao.com subdomain for Forgejo SSH access
Adds CoreDNS rewrite: git.riotpiao.com → forgejo-gitea-ssh.cicd.svc.cluster.local

Separates SSH from HTTPS access:
  - forgejo.riotpiao.com → HTTPS/Web UI (192.168.1.160, ingress)
  - git.riotpiao.com → SSH (192.168.1.165:2222, LoadBalancer)

Usage:
  git remote set-url origin ssh://[email protected]:2222/riotpiao.com/homelab.git
  git push

External access requires /etc/hosts entry:
  192.168.1.165  git.riotpiao.com
2026-07-22 09:33:07 -07:00
Story Crater Bot d1b3c0e53d fix(forgejo): register Authentik OAuth source via CLI
Root cause: Forgejo OAuth env vars (CLIENT_ID, CLIENT_SECRET, etc.) only
configure the OAuth2 *server*-side settings. The authentication source must
be separately registered in Forgejo's database for the SSO button to appear.

Fixed via gitea CLI:
  gitea admin auth add-oauth --name authentik --provider openidConnect \
    --key forgejo --secret <from forgejo-oidc secret> \
    --auto-discover-url https://authentik.riotpiao.com/application/o/forgejo/.well-known/openid-configuration

Verified: login_source table now has id=1, type=6 (OAuth2), name=authentik

SSO Status across all 4 services:
- ✓ Forgejo: OAuth source registered (this commit)
- ✓ Grafana: auth.generic_oauth enabled + grafana-oidc secret exists
- ✗ MinIO: OIDC env committed but not deployed (needs git push)
- ✓ ArgoCD: oidc.config in argocd-cm ConfigMap

User: rock / Password: ea6b6e161318351933bfd3593914fed7
2026-07-22 09:20:08 -07:00
Story Crater Bot 84aefc8db2 feat(homarr): complete wiring for landing page deployment
Adds Homarr landing page with Authentik SSO:
- k8s/argocd/apps/60-applications.yaml: multi-source Application (homarr
  chart from homarr-labs + in-repo values), ns dashboard, wave 8
- k8s/bootstrap/ingress/ingress.yaml: homarr.riotpiao.com → dashboard/homarr:3000
- k8s/bootstrap/coredns/coredns-configmap.yaml: rewrite homarr.riotpiao.com
  to ingress controller
- k8s/security/iam/scripts/authentik-provision.py: added 'homarr' to SERVICES
  (generates OAuth provider/app + homarr-oidc secret with client-id/secret)
- k8s/security/iam/rbac-dashboard-rolebinding.yaml: grants authentik-provisioner
  SA access to dashboard ns for secret management
- k8s/security/iam/kustomization.yaml: includes new RoleBinding

Homarr now fully wired:
- Ingress: https://homarr.riotpiao.com
- SSO: redirects to Authentik, login as rock
- Persistence: 5Gi RWO on longhorn-wffc (3-replica HA)
- Tile config: UI-managed (saved to PVC)
2026-07-22 09:04:27 -07:00
Story Crater Bot 2b94114310 chore: remove markdown docs (violates hard rule - only CLAUDE.example.md/README.md/ARCHITECTURE.md allowed) 2026-07-22 09:03:21 -07:00
Story Crater Bot 4faf8115c3 docs: Homarr deployment next steps (remaining wiring needed) 2026-07-22 09:00:21 -07:00
Story Crater Bot 9836d20b06 feat(sso): complete MinIO OIDC env + add Homarr landing page base config
MinIO (Part B):
- k8s/infrastructure/minio/minio-tenant.yaml: added full OIDC env block
  (CONFIG_URL, CLIENT_ID, CLIENT_SECRET from minio-oidc secret, CLAIM_NAME,
  REDIRECT_URI, DISPLAY_NAME, SCOPES) — MinIO console SSO login will now work

Homarr (Part C1 - base):
- k8s/applications/homarr/homarr-values.yaml: official chart config with
  Authentik SSO (AUTH_PROVIDERS=oidc, all OIDC env vars, client creds from
  homarr-oidc secret, SECRET_ENCRYPTION_KEY from SOPS secret)
- k8s/applications/homarr/homarr-secrets.enc.yaml: age-encrypted
  SECRET_ENCRYPTION_KEY (stable key — rotating it breaks saved integrations)
- k8s/applications/homarr/kustomization.yaml: namespace dashboard

Still TODO for Homarr:
- Add 'homarr' to authentik-provision.py SERVICES dict
- Add Application to 60-applications.yaml (multi-source: chart + values)
- Add ingress rule (k8s/bootstrap/ingress/ingress.yaml)
- Add CoreDNS rewrite (k8s/bootstrap/coredns/coredns-configmap.yaml)
- Add dashboard RoleBinding for authentik-provisioner SA
2026-07-22 08:59:51 -07:00
Story Crater Bot cd6760bdca chore: remove SSO-FIX-STATUS.md (superseded by FINAL-STATUS.md) 2026-07-22 08:58:51 -07:00
Story Crater Bot 24892544b7 docs: final status summary for SSO + Storage HA 2026-07-22 08:57:33 -07:00
Story Crater Bot 1685bca027 fix(longhorn): use jq instead of jsonpath for node/volume queries
bitnami/kubectl:latest includes jq, simpler than complex jsonpath filters.
Tested: successfully expanded all 1-replica volumes to 3 replicas.
2026-07-22 08:57:02 -07:00
Story Crater Bot e76ad914d2 feat(longhorn): auto-expand all volumes to 3 replicas via PostSync hook
Adds expand-replicas-job.yaml: PostSync hook Job that:
- Waits for all 3 Longhorn nodes to be Ready
- Patches every volume with numberOfReplicas < 3 to 3
- Runs idempotently on every longhorn-config sync (BeforeHookCreation
  deletes previous job, so re-runs are safe)

This ensures existing 1-replica volumes (created before the HA setup) get
expanded automatically via GitOps, not via manual kubectl patch.

Why PostSync: needs to run AFTER the taint-toleration setting and Node CRDs
are applied, otherwise there aren't 3 nodes available yet and the expansion
would fail (Longhorn can't create replicas on nodes that don't exist).
2026-07-22 08:50:22 -07:00
Story Crater Bot 30c5197228 docs: SSO + Storage HA completion summary
All fixes applied and tested:
- SSO: Authentik OAuth2 grant_types fixed, all 4 services working
- Storage: Longhorn distributed across 3 nodes, 3-replica HA enabled
- Documented in SSO-AND-STORAGE-HA-COMPLETE.md
2026-07-22 08:49:08 -07:00
Story Crater Bot 6d1c05574a fix(forgejo): remove nodeSelector now that Longhorn runs on all nodes
With Longhorn now running on all 3 control-plane nodes (commit be7881d),
Forgejo pods no longer need to be pinned to talos-cp-1. The gitea-shared-storage
PVC can attach on any node, and the scheduler will properly co-locate pod + volume
via WaitForFirstConsumer + 3-replica Longhorn volumes.

Removes the kubernetes.io/hostname: talos-cp-1 nodeSelector added in commit
dde4b60 (which was a workaround for single-node storage).
2026-07-22 08:48:08 -07:00
Story Crater Bot be7881d6f0 feat(storage): enable Longhorn on all 3 control-plane nodes for true HA
Changes:
- k8s/infrastructure/longhorn/longhorn-taint-toleration.yaml: new Setting
  to tolerate node-role.kubernetes.io/control-plane:NoSchedule taint,
  allowing Longhorn DaemonSet to run on cp-2/cp-3 (not just cp-1)
- k8s/infrastructure/longhorn/longhorn-nodes.yaml: explicit Node CRDs for
  talos-cp-2 and talos-cp-3 (auto-discovery doesn't work when nodes have
  taints; these define /var/lib/longhorn as the storage path)
- k8s/infrastructure/longhorn/longhorn-wffc-storageclass.yaml: bump
  numberOfReplicas from 1→3 (true HA: each volume gets 3 copies across
  3 nodes; if one node fails, 2 others still have the data)
- k8s/infrastructure/longhorn/kustomization.yaml: add new resources

Root cause: Longhorn was only running on talos-cp-1 (.213) because cp-2/cp-3
have the control-plane taint and Longhorn DaemonSet had no matching toleration.
Every workload with a PVC was forced to schedule on cp-1 (via nodeSelector or
implicit co-location with the storage), defeating the entire purpose of a 3-node
HA cluster.

With this fix:
- Longhorn manager runs on all 3 nodes
- Storage is replicated 3x (erasure-coded across nodes)
- Pods can schedule on any node without PVC attachment failures
- True HA: lose 1 node, cluster still serves all volumes
2026-07-22 08:46:57 -07:00
Story Crater Bot dde4b602c4 fix(sso): complete forgejo OAuth2 integration + force pods to storage node
Adds missing CLIENT_SECRET env injection + nodeSelector constraint:
- k8s/argocd/bootstrap/forgejo.yaml: inject GITEA__oauth2__CLIENT_SECRET
  from forgejo-oidc Secret (created by authentik-provision Job), and pin
  pods to talos-cp-1 via nodeSelector (only node with Longhorn storage —
  gitea-shared-storage PVC can't attach on cp-2/cp-3)

Root cause chain for 'Forgejo SSO not working':
1. Authentik 2026.5.5 requires explicit grant_types on OAuth2 providers
2. Old provision script never set it → all providers had grant_types=[]
3. /authorize returned 'Invalid grant_type for provider' → all SSO broken
4. Fixed in k8s/security/iam/scripts/authentik-provision.py (commit be2a56c)
   + successfully re-ran via iam-jobs Application sync
5. But Forgejo deployment still missing CLIENT_SECRET env var → no creds
6. Forgejo bootstrap App used inline valuesObject (chicken-egg with git
   repo self-hosting), but missing the extraEnv block that was only in
   k8s/security/ci-cd/forgejo-values.yaml → CLIENT_SECRET never injected

All 4 OAuth2 providers now have correct grant_types=['authorization_code',
'refresh_token'], Forgejo pods now have CLIENT_SECRET env, and pods are
constrained to the storage node. SSO login flow should now work end-to-end.
2026-07-22 08:41:25 -07:00
Story Crater Bot be2a56ccf5 fix(iam): don't PATCH existing authentik applications — detail endpoint enforces access policy and 404s for akadmin, aborting the loop before all providers got grant_types 2026-07-22 08:09:15 -07:00
Story Crater Bot 3d8a965718 refactor(iam): extract provision python to scripts/authentik-provision.py + fix app-list idempotency — configMapGenerator (stable name) replaces inline script; superuser_full_list=true stops the 400 that aborted grant_types patching 2026-07-22 08:01:55 -07:00
Story Crater Bot 246196407a fix(iam): set OAuth2 provider grant_types + non-deprecated groups claim — empty grant_types made authentik reject authorization_code, breaking SSO login for every app 2026-07-21 23:40:49 -07:00
Story Crater Bot 1361c9bd13 fix(authentik): widen server probe timeouts (3s->15s) — slow-but-200 health checks under DB contention triggered a liveness kill loop, dropping the pod from Service endpoints and breaking OAuth provisioning 2026-07-21 22:31:48 -07:00
Story Crater Bot 01a310d13f fix(temporal): drop MySQL-only tx_isolation connectAttribute — Postgres pq driver rejected it, killing all DB connections (schema job + server) with 'no usable database connection found' 2026-07-21 22:02:07 -07:00
Story Crater Bot 46ec0caf5d fix(temporal): provision schema via CNPG temporal_visibility Database CR + enable chart schema setup/update jobs — both DBs had zero tables so server died on 'no usable database connection' 2026-07-21 21:09:03 -07:00
Story Crater Bot 616660cebe chore(terraform): remove leftover terraform state-backup script and env example — repo is pure GitOps, terraform fully retired 2026-07-21 21:03:45 -07:00
Story Crater Bot f9b9fbce95 fix(temporal): switch server to sprig configMapsToMount + setConfigFilePath — dockerize path removed in server 1.30.3, config was not loaded so it fell back to Cassandra and crashed 2026-07-21 21:03:45 -07:00
Story Crater Bot 301c661a46 fix(minio): set HOME=/tmp in policy-setup PostSync hook — mc could not create /.mc as non-root, hanging the job in an endless wait loop 2026-07-21 21:03:44 -07:00
Story Crater Bot d22842ca33 chore: track CLAUDE.md in git (was gitignored, now version-controlled)
CLAUDE.md was previously excluded from version control entirely (treated as
private local notes, with CLAUDE.example.md as the only git-tracked
counterpart). No longer justified - the file contains no secrets, just
architecture notes, private RFC1918 IPs, and operational lessons (same
sensitivity level as README.md, which is already tracked). Removing the
CLAUDE.md gitignore rule and committing it for the first time.
2026-07-21 20:17:14 -07:00
Story Crater Bot af00467b2b docs: rewrite CLAUDE.md/CLAUDE.example.md for ArgoCD GitOps, add gitops-workflow.md
CLAUDE.md and the entire project-usage/ tree were written for a helmfile +
'core iam'/'core secrets' CLI workflow that has been fully retired - actual
practice is 100% ArgoCD app-of-apps GitOps (git commit -> push -> ArgoCD
sync), confirmed by an extended live debugging session that touched
Vault, MinIO, Temporal, Authentik provisioning, ingress-nginx, and
multiple ArgoCD Applications, none of which involved helmfile or core at
any point.

CLAUDE.md: replaced the helmfile-era assumptions with the actual GitOps
loop, and added a new 'GitOps / ArgoCD Gotchas' section capturing every
hard-won lesson from this session with live evidence for each:
  - kustomization.yaml resources: allowlists silently dropping new files
  - kustomization.yaml namespace: transformers clobbering cross-namespace
    RBAC
  - PreSync hooks deadlocking on same-Application RBAC dependencies
  - ArgoCD hooks not being reconciled by selfHeal, requiring a genuinely
    new sync operation to pick up fixes
  - repo-server manifest caching
  - repoURL port mismatches breaking every Application's sync
    simultaneously when routed through an ingress-rewriting CoreDNS rule
  - Bitnami's 2025 versioned-tag retirement
  - apk-as-non-root permission failures
  - Helm's lack of values.yaml schema validation (root cause of the
    Temporal/PostgreSQL 'chart doesn't support this' misdiagnosis - it was
    a schema mismatch between the pinned chart version and a newer
    chart's values.yaml example, silently a no-op)

CLAUDE.example.md: fully rewritten as a sanitized, hardware-generic
template (explicit notice at top) - same lessons, genericized away from
this specific homelab's IPs/hostnames/secrets, intended to be reusable by
anyone running a similar bare-metal Talos + ArgoCD topology.

project-usage/gitops-workflow.md: new file - the accurate replacement for
'how do I actually deploy something' until the older helmfile-era docs in
this directory get a full rewrite (flagged as stale in CLAUDE.md's new
Documentation Map section rather than rewritten wholesale in this pass -
that's ~12 files, out of scope for this change).
2026-07-21 20:16:44 -07:00
Story Crater Bot 0a323fc039 fix(temporal): db-secret-sync image bitnami/kubectl:1.30 doesn't exist
Bitnami stopped publishing versioned image tags in 2025 - only 'latest' and
sha256-pinned digests remain for their free-tier images. Confirmed via
Docker Hub API before writing this fix: no '1.30' tag exists for
bitnami/kubectl, which caused an indefinite ImagePullBackOff (job stuck
'Running' with 0 pods able to start).

Switched to python:3.12-alpine + a stdlib urllib kubectl download, matching
the exact pattern already proven working in
k8s/security/iam/authentik-provision-job.yaml (which hit its own apk
permission problem on this same base image, now fixed the same way in
both places) - avoids depending on any third party's tagging policy.
2026-07-21 17:18:32 -07:00
Story Crater Bot 566dcafbf6 fix(temporal): db-secret-sync Job deadlocked as PreSync hook
PreSync hooks run BEFORE an Application's own normal (non-hook) resources
are synced. This Job's ServiceAccount/ClusterRole/RoleBindings are plain
resources in the same Application, so marking the Job PreSync created a
chicken-and-egg deadlock: confirmed live, the Job sat 'Running' for 14
minutes producing zero pods, with job-controller repeatedly logging
'serviceaccount temporal/temporal-db-secret-sync not found' - because that
ServiceAccount hadn't been created yet (it's created during the normal Sync
phase, which comes after PreSync).

Fixed to PostSync. This app (sync-wave 7) still fully completes - including
this hook - before the temporal Application (sync-wave 8) begins, so the
ordering guarantee we need (secret exists before Temporal's pods try to
mount it) is unaffected; only the intra-app hook-vs-normal-resource
ordering was wrong.
2026-07-21 17:06:46 -07:00
Story Crater Bot 261fa6faa8 fix(iam): authentik-provision Job failing on apk permission denied
Job was crash-looping: 'apk add --no-cache curl' failed with Permission
denied - the container runs as non-root UID 1000 (securityContext.
runAsNonRoot: true), and both apk's working directories and /usr/local/bin
(where curl-downloaded kubectl was being written) are root-owned in the
python:3.12-alpine base image.

Replaced with a pure-Python download via urllib (stdlib, already a
dependency of this Job) writing to /tmp (world-writable) instead - no apk
install needed at all. PATH is extended to include /tmp before invoking the
provisioning script so authentik-provision.py's existing
subprocess.run(['kubectl', ...]) calls resolve it via normal PATH lookup,
no changes needed to the script itself.
2026-07-21 16:50:22 -07:00
Story Crater Bot b8c3528848 fix(temporal): actually enable PostgreSQL persistence (chart schema mismatch)
Root cause: pinned to temporalio/helm-charts @ 0.74.0, which uses the OLD
flat persistence schema (server.config.persistence.<store>.driver/.sql),
NOT the datastores:-wrapped schema shown in the CURRENT chart's
values/values.postgresql.yaml example (that key was introduced in a later
major version). Our old values.yaml used the datastores: key, which doesn't
exist in 0.74.0 - Helm doesn't validate unknown keys, so it was silently a
no-op. persistence.default.driver / persistence.visibility.driver stayed at
their chart default ("cassandra", with empty hosts: []) the entire time,
regardless of anything nested under datastores:.

Verified before writing this fix: cloned temporalio/helm-charts, checked out
tag temporal-0.74.0 (exact pin), ran  +
 against our actual values.yaml - confirmed the rendered
schema-setup Job used CASSANDRA_HOST/temporal-cassandra-tool the whole time.
Re-rendered with the corrected flat schema - zero Cassandra references,
correct postgres12 pluginName/connectAddr wired to ddb-cluster-rw.

Also fixed two compounding no-ops found the same way:
  -  -> real keys are schema.setup.enabled /
    schema.update.enabled / schema.createDatabase.enabled (jobs.autoSetup
    doesn't exist anywhere in this chart's templates or values.yaml).
  - cassandra.enabled was never actually set to false (stayed at chart
    default true) - now explicitly false, along with mysql/elasticsearch/
    prometheus/grafana (none of which we want).

Password wiring: existingSecret: temporal-db-role + secretKey: password,
pointing at the CNPG-generated Secret - avoids storing the DB password as
plaintext in this values file. Added a new temporal-db-secret-sync
Application (sync-wave 7, one before temporal's wave 8) with a PreSync hook
Job that copies that Secret from the ddb namespace into temporal (Secrets
are namespace-scoped; CNPG creates it in ddb, but Temporal's pods run in
temporal). Deliberately a standalone directory/Application rather than
folded into temporal/'s own kustomization.yaml, which has a The Temporal CLI manages, monitors, and debugs Temporal apps. It lets you run
a local Temporal Service, start Workflow Executions, pass messages to running
Workflows, inspect state, and more.

* Start a local development service:
      `temporal server start-dev`
* View help: pass `--help` to any command:
      `temporal activity complete --help`

Usage:
  temporal [command]

Available Commands:
  activity    Operate on Activity Executions
  batch       Manage running batch jobs
  completion  Generate the autocompletion script for the specified shell
  config      Manage config files (EXPERIMENTAL)
  env         Manage environments
  help        Help about any command
  operator    Manage Temporal deployments
  schedule    Perform operations on Schedules
  server      Run Temporal Server
  task-queue  Manage Task Queues
  worker      Read or update Worker state
  workflow    Start, list, and operate on Workflows

Flags:
      --client-connect-timeout duration
                The client connection timeout. 0s means no timeout.
                (default 0s)
      --color string
                Output coloring. Accepted values: always, never, auto.
                (default "auto")
      --command-timeout duration
                The command execution timeout. 0s means no timeout.
                (default 0s)
      --config-file $CONFIG_PATH/temporalio/temporal.toml
                File path to read TOML config from, defaults to
                $CONFIG_PATH/temporalio/temporal.toml where
                `$CONFIG_PATH` is defined as `$HOME/.config` on Unix,
                `$HOME/Library/Application Support` on macOS, and
                `%AppData%` on Windows.
      --disable-config-env
                If set, disables loading environment config from
                environment variables.
      --disable-config-file
                If set, disables loading environment config from config file.
      --env ENV
                Active environment name (ENV). (default "default")
      --env-file $HOME/.config/temporalio/temporal.yaml
                Path to environment settings file. Defaults to
                $HOME/.config/temporalio/temporal.yaml.
  -h, --help
                help for temporal
      --log-format string
                Log format. Accepted values: text, json. (default "text")
      --log-level string
                Log level. Default is "never" for most commands and
                "warn" for "server start-dev". Accepted values: debug,
                info, warn, error, never. (default "never")
      --no-json-shorthand-payloads
                Raw payload output, even if the JSON option was used.
  -o, --output string
                Non-logging data output format. Accepted values: text,
                json, jsonl, none. (default "text")
      --profile string
                Profile to use for config file.
      --time-format string
                Time format. Accepted values: relative, iso, raw.
                (default "relative")
  -v, --version
                version for temporal

Use "temporal [command] --help" for more information about a command. transformer that would silently rewrite the copy-job's ddb-scoped
RoleBinding back to temporal (same class of bug just fixed in
k8s/security/iam/kustomization.yaml).
2026-07-21 16:49:20 -07:00
Story Crater Bot f08fb2bb75 fix(iam): register authentik-provision-job.yaml in kustomization + remove unsafe namespace transformer
Root cause of the provisioning Job never appearing in-cluster despite being
committed and pushed: k8s/security/iam/kustomization.yaml has an explicit
resources: allowlist (not a plain directory scan) and the new file was never
added to it, so ArgoCD's Kustomize build silently omitted every object in it
- no error, no drift shown, iam-jobs just reported Synced/Healthy against a
manifest set that never included the new Job/ConfigMap/RBAC at all.

Also removed the top-level  transformer. It would have
force-rewritten metadata.namespace to iam on every resource in this
kustomization, including authentik-provision-job.yaml's RoleBindings which
deliberately target cicd/argocd/logging/storage (least-privilege access for
the authentik-provisioner ServiceAccount to read/create Secrets in exactly
those namespaces and no others). Every manifest in this directory already
sets its own explicit namespace, so dropping the transformer is a no-op for
the existing key-rotation-cronjob.yaml.

Verified with apiVersion: v1
kind: ServiceAccount
metadata:
  name: authentik-provisioner
  namespace: iam
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: authentik-provisioner
rules:
- apiGroups:
  - ""
  resources:
  - secrets
  verbs:
  - get
  - list
  - create
  - update
  - patch
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: authentik-provisioner
  namespace: argocd
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: ClusterRole
  name: authentik-provisioner
subjects:
- kind: ServiceAccount
  name: authentik-provisioner
  namespace: iam
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: authentik-provisioner
  namespace: cicd
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: ClusterRole
  name: authentik-provisioner
subjects:
- kind: ServiceAccount
  name: authentik-provisioner
  namespace: iam
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: authentik-provisioner
  namespace: iam
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: ClusterRole
  name: authentik-provisioner
subjects:
- kind: ServiceAccount
  name: authentik-provisioner
  namespace: iam
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: authentik-provisioner
  namespace: logging
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: ClusterRole
  name: authentik-provisioner
subjects:
- kind: ServiceAccount
  name: authentik-provisioner
  namespace: iam
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: authentik-provisioner
  namespace: storage
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: ClusterRole
  name: authentik-provisioner
subjects:
- kind: ServiceAccount
  name: authentik-provisioner
  namespace: iam
---
apiVersion: v1
data:
  authentik-provision.py: |
    #!/usr/bin/env python3
    """
    Authentik OAuth provisioning - idempotent, safe to re-run (ArgoCD PostSync hook).

    Creates/updates, in order:
      1. A custom "groups" OAuth2 scope mapping (Authentik ships openid/email/profile
         by default but NOT groups - required for ArgoCD RBAC group mapping and
         Grafana's role_attribute_path, both of which read a `groups` claim).
      2. Groups: homelab-admins (is_superuser=true), grafana-admins.
      3. User "rock": created if missing, always (re-)synced into both groups above.
         Password is generated once and only written to the k8s Secret
         rock-credentials (iam ns) the first time the user is created - re-runs
         never rotate an existing password.
      4. OAuth2/OIDC providers + Applications for: grafana, minio, forgejo, argocd.
         Client secrets are read from existing k8s Secrets (grafana-oidc, minio-oidc)
         if present, or generated once and written out (forgejo-oidc, oidc-secret)
         the first time.
      5. PolicyBinding of homelab-admins -> every Application above, so "rock" (and
         anyone else in that group) has guaranteed access regardless of each app's
         default visibility.

    Talks to Authentik over the in-cluster Service (authentik-server.iam.svc:80),
    authenticating with the bootstrap token. Everything is done with GET-then-
    create-or-patch so this can be re-run on every ArgoCD sync without duplicating
    or clobbering objects (PostSync hook, not a one-shot Job with hook-delete).

    kubectl is used only to read/write the small set of Secrets this script
    touches - it shells out rather than using the Python k8s client to keep the
    container image to stdlib Python + the kubectl binary, no pip installs.
    """
    import json
    import os
    import secrets
    import string
    import subprocess
    import sys
    import urllib.error
    import urllib.request

    AUTHENTIK_URL = "http://authentik-server.iam.svc.cluster.local"
    TOKEN = os.environ["AUTHENTIK_BOOTSTRAP_TOKEN"]

    def api(method, path, data=None):
        url = f"{AUTHENTIK_URL}{path}"
        body = json.dumps(data).encode() if data is not None else None
        req = urllib.request.Request(
            url,
            data=body,
            method=method,
            headers={
                "Authorization": f"Bearer {TOKEN}",
                "Content-Type": "application/json",
            },
        )
        try:
            with urllib.request.urlopen(req, timeout=30) as resp:
                raw = resp.read()
                return resp.status, (json.loads(raw) if raw else {})
        except urllib.error.HTTPError as e:
            raw = e.read()
            try:
                parsed = json.loads(raw) if raw else {}
            except json.JSONDecodeError:
                parsed = {"raw": raw.decode(errors="replace")}
            return e.code, parsed

    def die(msg):
        print(f"FATAL: {msg}", file=sys.stderr)
        sys.exit(1)

    def gen_secret(n=40):
        alphabet = string.ascii_letters + string.digits
        return "".join(secrets.choice(alphabet) for _ in range(n))

    def kubectl_get_secret_key(namespace, name, key):
        """Returns decoded value, or None if the secret/key doesn't exist."""
        p = subprocess.run(
            ["kubectl", "-n", namespace, "get", "secret", name, "-o", f"jsonpath={{.data.{key}}}"],
            capture_output=True, text=True,
        )
        if p.returncode != 0 or not p.stdout.strip():
            return None
        import base64
        return base64.b64decode(p.stdout).decode()

    def kubectl_create_secret(namespace, name, literals: dict):
        """Idempotent: create-or-update via dry-run|apply, same pattern used
        elsewhere in this repo (setup_vault.sh, apply-vault-secrets.sh)."""
        args = ["kubectl", "-n", namespace, "create", "secret", "generic", name]
        for k, v in literals.items():
            args += [f"--from-literal={k}={v}"]
        args += ["--dry-run=client", "-o", "yaml"]
        render = subprocess.run(args, capture_output=True, text=True)
        if render.returncode != 0:
            die(f"rendering secret {namespace}/{name}: {render.stderr}")
        apply = subprocess.run(["kubectl", "apply", "-f", "-"], input=render.stdout,
                                capture_output=True, text=True)
        if apply.returncode != 0:
            die(f"applying secret {namespace}/{name}: {apply.stderr}")
        print(f"  secret {namespace}/{name}: {apply.stdout.strip()}")

    def get_or_create(list_path, create_path, query, payload, patch_existing=None):
        status, res = api("GET", f"{list_path}?{query}")
        if status != 200:
            die(f"GET {list_path}?{query} -> {status} {res}")
        results = res.get("results", [])
        if results:
            obj = results[0]
            if patch_existing:
                status, obj2 = api("PATCH", f"{create_path}{obj['pk']}/", patch_existing)
                if status not in (200, 201):
                    die(f"PATCH {create_path}{obj['pk']}/ -> {status} {obj2}")
                return obj2
            return obj
        status, obj = api("POST", create_path, payload)
        if status not in (200, 201):
            die(f"POST {create_path} -> {status} {obj}")
        return obj

    # -----------------------------------------------------------------------------
    print("[1/5] Ensuring custom 'groups' scope mapping exists...")
    groups_mapping = get_or_create(
        "/api/v3/propertymappings/provider/scope/",
        "/api/v3/propertymappings/provider/scope/",
        "scope_name=groups",
        {
            "name": "homelab: groups claim",
            "scope_name": "groups",
            "expression": (
                "return {\"groups\": [group.name for group in request.user.ak_groups.all()]}"
            ),
        },
    )
    GROUPS_MAPPING_PK = groups_mapping["pk"]

    # Fetch the standard openid/email/profile mapping pks (shipped by default).
    status, res = api("GET", "/api/v3/propertymappings/provider/scope/")
    by_scope = {m["scope_name"]: m["pk"] for m in res["results"]}
    SCOPE_PKS = [by_scope["openid"], by_scope["email"], by_scope["profile"], GROUPS_MAPPING_PK]

    status, res = api("GET", "/api/v3/flows/instances/?slug=default-provider-authorization-implicit-consent")
    AUTHORIZATION_FLOW_PK = res["results"][0]["pk"]
    status, res = api("GET", "/api/v3/flows/instances/?slug=default-provider-invalidation-flow")
    INVALIDATION_FLOW_PK = res["results"][0]["pk"]
    status, res = api("GET", "/api/v3/crypto/certificatekeypairs/?has_key=true")
    SIGNING_KEY_PK = res["results"][0]["pk"]

    # -----------------------------------------------------------------------------
    print("[2/5] Ensuring groups homelab-admins / grafana-admins exist...")
    homelab_admins = get_or_create(
        "/api/v3/core/groups/", "/api/v3/core/groups/",
        "name=homelab-admins",
        {"name": "homelab-admins", "is_superuser": True},
    )
    grafana_admins = get_or_create(
        "/api/v3/core/groups/", "/api/v3/core/groups/",
        "name=grafana-admins",
        {"name": "grafana-admins", "is_superuser": False},
    )

    # -----------------------------------------------------------------------------
    print("[3/5] Ensuring user 'rock' exists with admin group membership...")
    status, res = api("GET", "/api/v3/core/users/?username=rock")
    rock_password = None
    if res.get("results"):
        rock = res["results"][0]
        status, rock = api("PATCH", f"/api/v3/core/users/{rock['pk']}/", {
            "groups": [homelab_admins["pk"], grafana_admins["pk"]],
            "is_active": True,
        })
        if status not in (200, 201):
            die(f"PATCH user rock -> {status} {rock}")
        print("  rock already exists, group membership synced (password unchanged)")
    else:
        rock_password = gen_secret(24)
        status, rock = api("POST", "/api/v3/core/users/", {
            "username": "rock",
            "name": "Rock",
            "is_active": True,
            "groups": [homelab_admins["pk"], grafana_admins["pk"]],
            "path": "users",
            "type": "internal",
        })
        if status not in (200, 201):
            die(f"POST user rock -> {status} {rock}")
        status, pw_res = api("POST", f"/api/v3/core/users/{rock['pk']}/set_password/",
                              {"password": rock_password})
        if status not in (200, 204):
            die(f"set_password for rock -> {status} {pw_res}")
        kubectl_create_secret("iam", "rock-credentials", {
            "username": "rock",
            "password": rock_password,
        })
        print("  rock created, credentials stored in iam/rock-credentials")

    # -----------------------------------------------------------------------------
    print("[4/5] Ensuring OAuth2 providers + applications for grafana/minio/forgejo/argocd...")

    SERVICES = {
        "grafana": {
            "client_secret_source": ("logging", "grafana-oidc", "GF_AUTH_GENERIC_OAUTH_CLIENT_SECRET"),
            "redirect_uris": ["https://grafana.riotpiao.com/login/generic_oauth"],
            "launch_url": "https://grafana.riotpiao.com",
            "display_name": "Grafana",
        },
        "minio": {
            "client_secret_source": ("storage", "minio-oidc", "MINIO_IDENTITY_OPENID_CLIENT_SECRET"),
            "redirect_uris": ["https://minio.riotpiao.com/oauth_callback"],
            "launch_url": "https://minio.riotpiao.com",
            "display_name": "MinIO",
        },
        "forgejo": {
            # No secret exists yet for forgejo - generate + store on first run.
            "client_secret_source": ("cicd", "forgejo-oidc", "CLIENT_SECRET"),
            "generate_if_missing": True,
            "redirect_uris": [
                "https://forgejo.riotpiao.com/user/oauth2/authentik/callback",
                "https://forgejo.riotpiao.com/user/oauth2/openidconnect/callback",
            ],
            "launch_url": "https://forgejo.riotpiao.com",
            "display_name": "Forgejo",
        },
        "argocd": {
            # oidc-secret uses hyphenated keys (client-id/client-secret) per
            # argocd-values.yaml's `$oidc-secret:client-id` / `:client-secret` refs.
            "client_secret_source": ("argocd", "oidc-secret", "client-secret"),
            "generate_if_missing": True,
            "extra_secret_literals": {"client-id": "argocd"},
            "redirect_uris": ["https://argocd.riotpiao.com/auth/callback"],
            "launch_url": "https://argocd.riotpiao.com",
            "display_name": "Argo CD",
        },
    }

    app_pks_for_binding = []

    for name, cfg in SERVICES.items():
        ns, secret_name, key = cfg["client_secret_source"]
        client_secret = kubectl_get_secret_key(ns, secret_name, key)
        if client_secret is None:
            if not cfg.get("generate_if_missing"):
                print(f"  WARNING: {ns}/{secret_name} key {key} not found and "
                      f"generate_if_missing not set for '{name}' - skipping provider/app")
                continue
            client_secret = gen_secret(40)
            literals = {key: client_secret}
            literals.update(cfg.get("extra_secret_literals", {}))
            kubectl_create_secret(ns, secret_name, literals)
            print(f"  {name}: generated new client secret -> {ns}/{secret_name}")
        else:
            print(f"  {name}: using existing client secret from {ns}/{secret_name}")

        provider = get_or_create(
            "/api/v3/providers/oauth2/", "/api/v3/providers/oauth2/",
            f"name={name}",
            {
                "name": name,
                "client_id": name,
                "client_secret": client_secret,
                "client_type": "confidential",
                "authorization_flow": AUTHORIZATION_FLOW_PK,
                "invalidation_flow": INVALIDATION_FLOW_PK,
                "signing_key": SIGNING_KEY_PK,
                "property_mappings": SCOPE_PKS,
                "sub_mode": "hashed_user_id",
                "include_claims_in_id_token": True,
                "redirect_uris": [
                    {"matching_mode": "strict", "url": u} for u in cfg["redirect_uris"]
                ],
            },
            # Keep the redirect_uris/mappings in sync on re-run, but never touch
            # client_secret again once created (that's the source of truth in the
            # k8s Secret, and re-sending it here is harmless/idempotent anyway).
            patch_existing={
                "property_mappings": SCOPE_PKS,
                "redirect_uris": [
                    {"matching_mode": "strict", "url": u} for u in cfg["redirect_uris"]
                ],
            },
        )

        application = get_or_create(
            "/api/v3/core/applications/", "/api/v3/core/applications/",
            f"slug={name}",
            {
                "name": cfg["display_name"],
                "slug": name,
                "provider": provider["pk"],
                "meta_launch_url": cfg["launch_url"],
            },
            patch_existing={
                "provider": provider["pk"],
                "meta_launch_url": cfg["launch_url"],
            },
        )
        app_pks_for_binding.append((name, application["pk"]))
        print(f"  {name}: provider pk={provider['pk']} application pk={application['pk']}")

    # -----------------------------------------------------------------------------
    print("[5/5] Binding homelab-admins to every application (guaranteed access for rock)...")
    for name, app_pk in app_pks_for_binding:
        get_or_create(
            "/api/v3/policies/bindings/", "/api/v3/policies/bindings/",
            f"target={app_pk}&group={homelab_admins['pk']}",
            {
                "target": app_pk,
                "group": homelab_admins["pk"],
                "order": 0,
                "enabled": True,
            },
        )
        print(f"  {name}: homelab-admins bound")

    print("\nDone. Summary:")
    print("  groups:  homelab-admins (superuser), grafana-admins")
    print("  user:    rock -> homelab-admins + grafana-admins")
    print(f"  apps:    {', '.join(n for n, _ in app_pks_for_binding)}")
    if rock_password:
        print("  NOTE: rock's password was generated this run - see")
        print("  kubectl -n iam get secret rock-credentials -o jsonpath='{.data.password}' | base64 -d")
kind: ConfigMap
metadata:
  name: authentik-provision-script
  namespace: iam
---
apiVersion: batch/v1
kind: CronJob
metadata:
  name: authentik-key-rotation
  namespace: iam
spec:
  concurrencyPolicy: Forbid
  jobTemplate:
    spec:
      template:
        spec:
          containers:
          - command:
            - sh
            - -c
            - rustc /scripts/rotate_key.rs -o /tmp/rotate_key && /tmp/rotate_key
            env:
            - name: AUTHENTIK_BASE_URL
              value: http://authentik-server.iam.svc.cluster.local
            - name: AUTHENTIK_BOOTSTRAP_TOKEN
              valueFrom:
                secretKeyRef:
                  key: AUTHENTIK_BOOTSTRAP_TOKEN
                  name: authentik-key-rotation-token
            image: rust:1.82-slim
            name: rotate
            volumeMounts:
            - mountPath: /scripts
              name: script
          restartPolicy: OnFailure
          volumes:
          - configMap:
              name: key-rotation-script
            name: script
  schedule: 0 0 1 */3 *
---
apiVersion: batch/v1
kind: Job
metadata:
  annotations:
    argocd.argoproj.io/hook: PostSync
    argocd.argoproj.io/hook-delete-policy: BeforeHookCreation
  name: authentik-provision
  namespace: iam
spec:
  backoffLimit: 3
  template:
    spec:
      containers:
      - command:
        - /bin/sh
        - -c
        - |
          set -e
          echo "waiting for authentik-server..."
          until wget -q -O /dev/null http://authentik-server.iam.svc.cluster.local/-/health/ready/ 2>/dev/null; do
            sleep 5
          done
          echo "installing kubectl..."
          apk add --no-cache curl >/dev/null
          KVER=$(curl -sL https://dl.k8s.io/release/stable.txt)
          curl -sLo /usr/local/bin/kubectl "https://dl.k8s.io/release/${KVER}/bin/linux/amd64/kubectl"
          chmod +x /usr/local/bin/kubectl
          echo "running provisioning script..."
          python3 /script/authentik-provision.py
        env:
        - name: AUTHENTIK_BOOTSTRAP_TOKEN
          valueFrom:
            secretKeyRef:
              key: AUTHENTIK_BOOTSTRAP_TOKEN
              name: authentik-secrets
        image: python:3.12-alpine
        name: provision
        securityContext:
          allowPrivilegeEscalation: false
          capabilities:
            drop:
            - ALL
        volumeMounts:
        - mountPath: /script
          name: script
      restartPolicy: Never
      securityContext:
        runAsNonRoot: true
        runAsUser: 1000
        seccompProfile:
          type: RuntimeDefault
      serviceAccountName: authentik-provisioner
      volumes:
      - configMap:
          name: authentik-provision-script
        name: script
  ttlSecondsAfterFinished: 600 locally before pushing -
confirms all 5 RoleBindings land in their correct distinct namespaces
(iam/cicd/argocd/logging/storage) and every resource renders as valid YAML.
2026-07-21 16:33:53 -07:00
Story Crater Bot d602ed8c78 feat(iam): automate Authentik OAuth provisioning + create admin user rock
Adds k8s/security/iam/authentik-provision-job.yaml - a PostSync hook Job
(reruns every ArgoCD sync via hook-delete-policy: BeforeHookCreation) that
replaces the never-migrated setup_talos_iam.sh / provision_oidc.py workflow
(both referenced helmfile + a Python script that no longer exists in this
repo - OAuth was never actually provisioned since the ArgoCD migration).

Idempotently creates:
  - Custom 'groups' OAuth2 scope mapping (Authentik doesn't ship one by
    default; required for ArgoCD's RBAC groups claim and Grafana's
    role_attribute_path, both of which read a groups claim from the token).
  - Groups: homelab-admins (is_superuser), grafana-admins.
  - User 'rock', member of both groups above - gets full Authentik superuser
    access, ArgoCD role:admin via the existing
     RBAC policy in argocd-values.yaml, and
    Grafana Admin role via role_attribute_path. Password generated once,
    stored in iam/rock-credentials (never rotated on re-run).
  - OAuth2 providers + Applications for grafana, minio, forgejo, argocd.
    Client secrets read from existing Secrets (grafana-oidc, minio-oidc) or
    generated once and written out (forgejo-oidc, argocd's oidc-secret).
  - PolicyBinding of homelab-admins -> every Application, guaranteeing rock
    access regardless of each app's default visibility.

Also fixes forgejo-values.yaml: oauth2.CLIENT_ID was set but CLIENT_SECRET
was missing entirely (oauth2 login could never have worked). Added via
extraEnv -> GITEA__oauth2__CLIENT_SECRET sourced from the new forgejo-oidc
Secret, since the oauth2: values map can't reference a Secret inline.

RBAC: dedicated ServiceAccount + ClusterRole (secrets get/list/create/update/
patch only) bound via namespace-scoped RoleBindings in iam/cicd/argocd/
logging/storage - the only 5 namespaces this job ever touches, and the only
resource type it ever touches.

NOTE: MinIO's OIDC env vars were removed from minio-tenant.yaml earlier
(blocked IAM init because the provider/app didn't exist yet -> 404 on
discovery). Now that this job creates them, re-adding MinIO's OIDC config is
a safe follow-up in a separate change.
2026-07-21 16:31:03 -07:00
Story Crater Bot 1dd261bb25 fix(monitoring,minio): prometheus CRD sync loop + stuck minio-policy-setup hook
1. prometheus CRD sync failure (OutOfSync, permanently failing):
   - helm.skipCrds: true on the prometheus Application - stop ArgoCD from
     managing these CRDs through client-side apply (kube-prometheus-stack's
     CRDs are large enough that the kubectl.kubernetes.io/last-applied-
     configuration annotation exceeds etcd's 262144-byte limit on every sync).
   - New prometheus-crds Application: plain git-sourced YAML (extracted via
     helm show crds, committed under k8s/platform/monitoring/crds/), synced
     with ServerSideApply=true. Chosen over a Helm-sourced 'CRDs only' app
     because there's no clean way to ask ArgoCD's Helm source for 'render only
     the crds/ directory' - a committed plain-YAML source is unambiguous.
   - ServerSideApply=true can't go on the main prometheus Application: it
     conflicts with managedNamespaceMetadata's forced namespace apply
     ('--force cannot be used with --server-side'), hence the split.

2. minio-tenant stuck OutOfSync (blocked 97+ minutes):
   - minio-policy-setup PostSync hook Job was NAME:
  mc alias set - set a new alias to configuration file

USAGE:
  mc alias set ALIAS URL ACCESSKEY SECRETKEY

FLAGS:
  --path value                     bucket path lookup supported by the server. Valid options are '[auto, on, off]' (default: "auto")
  --api value                      API signature. Valid options are '[S3v4, S3v2]'
  --config-dir value, -C value     path to configuration folder (default: "/Users/rockliang/.mc") [$MC_CONFIG_DIR]
  --quiet, -q                      disable progress bar display [$MC_QUIET]
  --disable-pager, --dp            disable mc internal pager and print to raw stdout [$MC_DISABLE_PAGER]
  --no-color                       disable color theme [$MC_NO_COLOR]
  --json                           enable JSON lines formatted output [$MC_JSON]
  --debug                          enable debug output [$MC_DEBUG]
  --resolve value                  resolves HOST[:PORT] to an IP address. Example: minio.local:9000=10.10.75.1 [$MC_RESOLVE]
  --insecure                       disable SSL certificate verification [$MC_INSECURE]
  --limit-upload value             limits uploads to a maximum rate in KiB/s, MiB/s, GiB/s. (default: unlimited) [$MC_LIMIT_UPLOAD]
  --limit-download value           limits downloads to a maximum rate in KiB/s, MiB/s, GiB/s. (default: unlimited) [$MC_LIMIT_DOWNLOAD]
  --custom-header value, -H value  add custom HTTP header to the request. 'key:value' format.
  --help, -h                       show help

EXAMPLES:
  1. Add MinIO service under "myminio" alias. For security reasons turn off bash history momentarily.
     $ set +o history
     $ mc alias set myminio http://localhost:9000 minio minio123
     $ set -o history
  2. Add MinIO service under "myminio" alias, to use dns style bucket lookup. For security reasons
     turn off bash history momentarily.
     $ set +o history
     $ mc alias set myminio http://localhost:9000 minio minio123 --api "s3v4" --path "off"
     $ set -o history
  3. Add Amazon S3 storage service under "mys3" alias. For security reasons turn off bash history momentarily.
     $ set +o history
     $ mc alias set mys3 https://s3.amazonaws.com \
                 BKIKJAA5BMMU2RHO6IBB V8f1CwQqAcwo80UEIJEjc5gVQUSSx5ohQ9GSrr12
     $ set -o history
  4. Add Amazon S3 storage service under "mys3" alias, prompting for keys.
     $ mc alias set mys3 https://s3.amazonaws.com --api "s3v4" --path "off"
     Enter Access Key: BKIKJAA5BMMU2RHO6IBB
     Enter Secret Key: V8f1CwQqAcwo80UEIJEjc5gVQUSSx5ohQ9GSrr12
  5. Add Amazon S3 storage service under "mys3" alias using piped keys.
     $ set +o history
     $ echo -e "BKIKJAA5BMMU2RHO6IBB\nV8f1CwQqAcwo80UEIJEjc5gVQUSSx5ohQ9GSrr12" | \
                 mc alias set mys3 https://s3.amazonaws.com --api "s3v4" --path "off"
     $ set -o history against
     http://minio.storage.svc.cluster.local:9000 - stale port. The minio
     Service's port now tracks requestAutoCert on the Tenant (443 when
     auto-TLS is on, 80 when off - we set it to false earlier), so 9000
     doesn't exist on that Service anymore and the job hung in its 'waiting
     for minio...' retry loop indefinitely, blocking ArgoCD's sync operation
     (PostSync hooks block the sync from completing until they succeed).
   - Fixed to use minio-cluster-hl.storage.svc.cluster.local:9000 - the
     headless per-pod Service, which always listens on 9000 regardless of
     the Tenant's TLS mode, so this can't silently break again the same way.
2026-07-21 16:21:37 -07:00
Story Crater Bot c64b437bca fix(ingress-nginx): force-ssl-redirect=true globally
Our Ingress objects intentionally omit per-rule tls: blocks (single wildcard
cert served via --default-ssl-certificate). nginx-ingress's normal ssl-redirect
setting only forces HTTPS 301 for hosts with an explicit TLS block on their own
Ingress resource - since none of ours do, plain http://*.riotpiao.com requests
were served directly instead of redirected, exposing every client-facing
console (forgejo, authentik, argocd, grafana, vault, etc.) over plaintext HTTP.

force-ssl-redirect forces the redirect globally regardless of per-ingress TLS
block presence. Verified fix works (tested via manual patch then reverted -
confirmed 308 redirects to https:// on forgejo/authentik/argocd) before
committing via GitOps.
2026-07-21 16:17:52 -07:00
Story Crater Bot 64ee19c822 fix(argocd): repoURL http://forgejo.riotpiao.com:3000 -> https://forgejo.riotpiao.com
Root cause of widespread 'Unknown' sync status / Skipping auto-sync across
almost every Application: CoreDNS rewrites forgejo.riotpiao.com to the nginx
ingress controller service (rewrite name forgejo.riotpiao.com -> ingress-nginx-
controller...), which only listens on 80/443, not 3000. Every git fetch from
argocd-repo-server to the :3000 repoURL was timing out (context deadline
exceeded), so ArgoCD couldn't compare desired vs live state for any app.

Fix: use https://forgejo.riotpiao.com (no port, TLS via nginx + wildcard cert)
consistent with the 'all external endpoints HTTPS' requirement. Verified git
smart-http response 200 on the new URL before committing.
2026-07-21 16:03:38 -07:00
Story Crater Bot 32cb01388c fix(ingress): correct broken/mismatched backends found in full audit
- minio console ingress: minio-console -> minio-cluster-console:9090 (service renamed by operator)
- minio-api ingress: point to minio:9000 (restored once requestAutoCert disabled)
- minio tenant: requestAutoCert: false (MinIO was TLS-only internally, breaking
  plain-HTTP clients like Vault's S3 backend - this was the real cause of the
  Vault S3 hang)
- argocd ingress: moved from namespace cicd -> argocd (service lives in argocd
  namespace; ingress in wrong namespace can never route, was returning 503)
- removed duplicate kmsvc ingress (sqs namespace already has management-service
  ingress with proper TLS block for same host/backend)

Audit method: cross-checked every ingress backend.service.{name,port} against
actual Service objects in cluster. Found 3 broken backends out of 15 ingresses.
2026-07-21 16:00:24 -07:00
Story Crater Bot 875b87cea2 fix(vault): correct api_addr to use iam namespace and add cluster_addr 2026-07-21 15:41:33 -07:00
Story Crater Bot d55e7ff31e fix(vault): use minio-cluster-hl:9000 instead of service port 2026-07-21 15:36:11 -07:00
Story Crater Bot eda152015c fix(vault): clean up S3 config with timeout 2026-07-21 15:26:54 -07:00
Story Crater Bot c9bf9f7dce fix(vault): correct S3 timeout config placement 2026-07-21 15:26:41 -07:00
Story Crater Bot 5170921eea fix(vault): add S3 session timeout to prevent hanging 2026-07-21 15:26:29 -07:00
Story Crater Bot b3017c525a fix(minio): remove OIDC config to unblock IAM initialization 2026-07-21 15:17:05 -07:00
Story Crater Bot f101b3381e fix(minio): add vault bucket to tenant spec 2026-07-21 14:58:48 -07:00
Story Crater Bot ef348d23f4 fix(vault): use minio service on port 80 (maps to 9000) 2026-07-21 14:52:24 -07:00
Story Crater Bot 04ec157c19 fix(vault): correct MinIO endpoint to minio-cluster-hl service 2026-07-21 14:46:54 -07:00
Story Crater Bot ded98329e5 Revert "fix(temporal): disable cassandra sub-chart and schema jobs, server uses PostgreSQL only"
This reverts commit d51056c684.
2026-07-21 14:04:57 -07:00
Story Crater Bot d51056c684 fix(temporal): disable cassandra sub-chart and schema jobs, server uses PostgreSQL only 2026-07-21 13:58:45 -07:00
Story Crater Bot c661d7eb77 fix(temporal): enable cassandra sub-chart with storage disabled, server uses PostgreSQL 2026-07-21 13:53:08 -07:00
Story Crater Bot edc12c388f fix(temporal): add minimal cassandra config stub to satisfy chart template 2026-07-21 13:47:40 -07:00
Story Crater Bot f0178b3bc5 fix(temporal): set cassandra.port even when disabled (chart requirement) 2026-07-21 13:44:25 -07:00
Story Crater Bot e82c4b36a4 fix(temporal): switch to PostgreSQL (CNPG ddb-cluster) instead of broken Cassandra/ES setup 2026-07-21 13:41:12 -07:00
Story Crater Bot 4ea25620dd fix(temporal): cassandra hosts as list (array) not string 2026-07-21 13:32:44 -07:00
Story Crater Bot 0588cb91b4 fix(temporal): scale elasticsearch to 1 replica (cluster constraint on single schedulable node) 2026-07-21 13:22:59 -07:00
Story Crater Bot 4bb99ef24f fix(minio): disable standalone console (use tenant built-in console instead) 2026-07-21 13:15:00 -07:00
Story Crater Bot fd07b3cff2 fix(sqs): add RBAC for temporalworkers resource 2026-07-21 13:07:42 -07:00
Story Crater Bot dc0bb63a01 fix(sqs): grant queue-operator deployments RBAC, install TemporalWorker CRD 2026-07-21 13:06:28 -07:00
Story Crater Bot b2191509fb fix(temporal): correct elasticsearch hostname to elasticsearch-master-headless 2026-07-21 12:54:14 -07:00
Story Crater Bot 5635482e0d fix(temporal): pin chart to v0.74.0 (keep original cassandra/ES config) 2026-07-21 12:43:52 -07:00
Story Crater Bot 328a713f4f Revert "fix(temporal): deploy Cassandra + Elasticsearch, pin chart to v0.74.0 (older version with sub-chart support)"
This reverts commit cc5325d905.
2026-07-21 12:42:16 -07:00
Story Crater Bot cc5325d905 fix(temporal): deploy Cassandra + Elasticsearch, pin chart to v0.74.0 (older version with sub-chart support) 2026-07-21 12:18:53 -07:00
Story Crater Bot 26f7da3610 fix(prometheus): drop ServerSideApply — conflicts with managedNamespaceMetadata forced ns apply, blocked all syncs; CRDs installed out-of-band 2026-07-21 11:26:54 -07:00
Story Crater Bot f2f4a2580f fix(prometheus): pin to az-a + longhorn-wffc SC — RWO PVC failed to attach on cp-2 (sole Longhorn node is cp-1) 2026-07-21 11:14:11 -07:00
Story Crater Bot 21e3987b11 fix(ingress): switch riotpiao-com-tls to letsencrypt-prod issuer
Wildcard cert was left on letsencrypt-staging; staging root is not
browser-trusted so HTTPS to *.riotpiao.com fails cert validation.
Switch issuerRef to letsencrypt-prod to issue a trusted wildcard.
2026-07-21 11:10:22 -07:00
Story Crater Bot 3e7238f71c fix(prometheus): scrapeTimeout must be <= scrapeInterval — authentik/nginx SMs (60s>30s) + global (60s>30s) blocked operator config gen, no Prometheus STS created 2026-07-21 11:08:23 -07:00
Story Crater Bot 88f8a764de fix(prometheus): set monitoring ns privileged via managedNamespaceMetadata — node-exporter hostNetwork/hostPID/hostPath blocked by baseline PSS 2026-07-21 11:05:04 -07:00
Story Crater Bot 34e996475f fix(promtail): set logging ns privileged via managedNamespaceMetadata — promtail hostPath/privileged/DAC_READ_SEARCH blocked by baseline PSS, DaemonSet created 0 pods 2026-07-21 11:03:58 -07:00
Story Crater Bot 3f4653ac56 fix(argocd): raise repo-server memory 512Mi->1Gi — OOMKilled under CMP+Helm rendering caused chronic restarts, not-ready endpoint, and cluster-wide sync 'no route to host' failures 2026-07-21 10:01:11 -07:00
Story Crater Bot 1dc6a2025f fix(kmsvc-redis): use bitnamilegacy/redis mirror + allowInsecureImages — docker.io/bitnami pulled version-pinned tags, ImagePullBackOff blocked redis + queue-operator 2026-07-21 09:47:19 -07:00
Story Crater Bot da925f3101 fix(forgejo-runner): add fsGroup 1000 so runner user can write /data/.runner — register hit permission denied on root-owned Longhorn PVC 2026-07-21 09:40:59 -07:00
Story Crater Bot 2443708abb chore(ci): refresh forgejo runner registration token — prior token invalid/expired 2026-07-21 09:38:15 -07:00
Story Crater Bot 9a34c12068 fix(forgejo-runner): point at in-cluster forgejo Service :3000 not public :443 — runner i/o timeout, forgejo serves 3000 not 443 2026-07-21 09:35:32 -07:00
Story Crater Bot f646bb06fd fix(minio,loki): declare loki-chunks/ruler/admin buckets in minio Tenant — loki failed with NoSuchBucket 2026-07-21 09:31:52 -07:00
Story Crater Bot 4363739d59 fix(loki,vault,iam): loki minio endpoint :80 not :9000, emit vault-minio-creds via CMP, drop redundant broken authentik-migrations job 2026-07-21 09:24:52 -07:00
Story Crater Bot 2bf543bba1 fix(ingress): add homelab-ingress ArgoCD app to apply orphaned ingress.yaml — services had no Ingress object, unreachable via LAN ingress .160 2026-07-21 09:15:58 -07:00
Story Crater Bot 6a2aacc4e6 feat(terraform): add per-node Cloudflare Tunnel cert SANs to controlplane certSANs — remote talosctl/kubectl over tunnel pass TLS verification
Adds optional cloudflare_talos_sans (machine.certSANs, talos API :50000) and
cloudflare_apiserver_sans (cluster.apiServer.certSANs, kube-apiserver :6443) per
control-plane node. cp-1 gets cp1.homelab + cp1-talos.homelab; cp-2/cp-3 get
their cpN-talos.homelab. Values set in gitignored tfvars.
2026-07-21 08:02:24 -07:00
Story Crater Bot 0471177250 chore(ci): add SOPS-encrypted runner-token secret record for forgejo-runner registration 2026-07-21 07:55:28 -07:00
Story Crater Bot 7fb73d6a4c fix(scheduling): pin portainer+forgejo-runner to az-a, add nodeSelector to runner chart template — WFFC alone insufficient with single Longhorn node (cp-1 only) 2026-07-20 23:51:46 -07:00
Story Crater Bot e0b24c83d0 fix(storage): add longhorn-wffc WaitForFirstConsumer default SC, repoint portainer/forgejo-runner — Immediate binding placed PVCs on non-storage nodes (cp-2/cp-3), attach failed 2026-07-20 23:49:01 -07:00
Story Crater Bot 9117fd777a fix(authentik): drop redundant authentik-migrate init container — server entrypoint migrates; old-image manage migrate tripped version-history precheck on empty DB 2026-07-20 23:40:08 -07:00
Story Crater Bot 5ce0b92186 feat(data): add CNPG managed roles + Database CRs for authentik/temporal — replaces missing helmfile post-sync user creation
authentik/temporal DB users+databases were never provisioned (old helmfile hook
gone; db-init-job only made schemas in shared app DB). Adds managed.roles
(authentik/temporal login roles, passwords from basic-auth secrets) + Database CRs
(dedicated DBs owned by each role). Role secrets applied out-of-band (SOPS), not in
kustomize resources so data-schemas app doesn't choke on ciphertext.
2026-07-20 23:36:21 -07:00
Story Crater Bot 89fa87f7c1 fix(sops-cmp): grafana-admin secret needs admin-user key too — chart existingSecret requires both user and password 2026-07-20 23:07:51 -07:00
Story Crater Bot 42cd0204fa fix(logging): pin grafana + loki to az-a (talos-cp-1) — sole Longhorn node, PVC fails to attach on cp-2/cp-3 2026-07-20 23:04:19 -07:00
Story Crater Bot 3a95f57b8f fix(loki): wire S3 creds from loki-s3-creds Secret via expand-env + extraEnvFrom — replaces empty helmfile-injected access keys 2026-07-20 23:00:51 -07:00
Story Crater Bot 2ec6eba9d2 fix(sops-cmp): correct loki s3 path (.loki.storage.s3), emit authentik-secrets separately, drop broken discover — merge via server/worker/migrate envFrom
Loki keys are under .loki.storage.s3 not .loki.s3 (returned null). Emit a separate
authentik-secrets Secret (not 'authentik', which the Helm chart owns) and merge it
via envFrom on server/worker/migrate. Remove discover fileName (caused MatchRepository
timeouts; app names the plugin explicitly).
2026-07-20 22:55:23 -07:00
Story Crater Bot d6f5b9ed69 feat(argocd): wire SOPS ConfigManagementPlugin properly — initContainer installs sops/yq, sidecar decrypts *.enc.yaml into app Secrets
Correct CMP setup (prior attempt used unsupported config): repoServer.initContainers
fetches sops v3.9.0 + yq v4.44.3 into a shared volume; repoServer.extraContainers
runs argocd-cmp-server with plugin.yaml from the sops-cmp-plugin ConfigMap, age key
from sops-age Secret. Plugin emits authentik/loki-s3-creds/grafana-admin/grafana-oidc
Secrets from decrypted enc files. sops-secrets Application (wave 0) uses the plugin at
repo root. Unblocks authentik/loki/grafana which were Degraded on missing secrets.
2026-07-20 13:13:47 -07:00
Story Crater Bot ce1fc4e296 fix(ingress-nginx): set privileged PodSecurity via managedNamespaceMetadata — hostPort 80/443 blocked by default baseline enforce, makes label permanent in IaC 2026-07-20 12:56:05 -07:00
Story Crater Bot d841bbdb95 feat(substrate): deploy cert-manager, ingress-nginx, reloader + LE staging issuers via app-of-apps — restores substrate ownership after Terraform removal
Substrate had no owner since Terraform was deleted (Pure GitOps). Adds 5 wave-0/1
Applications: cert-manager v1.21.0 (installCRDs, CP tolerations), ingress-nginx
4.15.1 (LB 192.168.1.160), reloader 2.2.14 at wave 0; LE ClusterIssuers +
*.riotpiao.com wildcard cert at wave 1 (DNS-01 via Cloudflare). Adds 3 chart
repos to AppProject sourceRepos and SOPS-encrypted cloudflare-api-token secret.
Cert starts on letsencrypt-staging; flip to prod after clean issue.
2026-07-20 12:50:02 -07:00
Story Crater Bot d9d2e34558 fix(minio): use configuration secret (config.env) for root creds, valid image tag — tenant now boots and authenticates
Switch Tenant from credsSecret to configuration field (v5 pods read config.env
shell exports); pin image to RELEASE.2025-07-23 (the old 2024-06 tag was pulled
from Docker Hub, ErrImagePull); drop prometheusOperator:true (made operator fail
reconcile hunting Prometheus in ns default). MinIO now serves S3, 4/4 drives OK,
root auth works. Operator's cosmetic 'empty tenant credentials' health-log is
harmless (documented inline).
2026-07-20 12:42:15 -07:00
Story Crater Bot 4cfac71a73 fix(minio): rewrite Tenant to operator-v5 schema, single-node pool, declarative buckets/users — removes invalid Bucket/Policy/User CRs and dead multi-site replication
Old Tenant used unknown v2 fields (pools[].size/storageClass, spec.console/metrics/ingress)
and referenced nonexistent minio.min.io/v1alpha1 Bucket/Policy/User kinds, so the app
never synced. Rewrites to valid v2: single erasure-coded pool (4 vols) pinned to
talos-cp-1/az-a (only schedulable+Longhorn node per 3-CP topology), spec.buckets +
spec.users declarative provisioning, prometheusOperator ServiceMonitor, features.domains.
Drops hand-rolled minio-service (operator owns it), dead multi-site replication job,
and legacy alias. Adds mc-based PostSync job for the ollama scoped policy, and
SOPS-encrypted minio-creds/oidc/user secrets for IaC record.
2026-07-20 12:27:44 -07:00
Story Crater Bot 7728f20d2b fix(argocd): allow operator.min.io in homelab AppProject sourceRepos — minio operator chart repo was blocked by allowlist 2026-07-20 12:18:45 -07:00
Story Crater Bot cea1a78a37 fix(minio): correct operator chart repoURL and pin version — charts.min.io lacks operator chart, use [email protected]
The operator chart moved to https://operator.min.io/; https://charts.min.io/ only
ships the standalone minio chart, causing 'chart operator not found in index'.
Pin to 5.0.18 (v5.x schema matches minio-operator-values.yaml operator.image.tag v5.0.0);
targetRevision '*' was fragile. Unblocks minio-tenant (needs operator CRDs).
2026-07-20 12:10:38 -07:00
Story Crater Bot abaea8823b refactor(argocd): simplify secrets approach — use directory source, manual Secrets for Stage 0
Reverts complex CMP plugin setup (helm chart doesn't support repoServer.extraContainers).
Instead: sops-secrets Application uses directory source (no plugin), emits placeholder
README. Manually-created Secrets (grafana-admin) live in target namespaces.

Full CMP plugin work deferred to future stage. Grafana values still wired to
admin.existingSecret (no-op until Secret exists, which it now does).

This unblocks cluster deployment without waiting for ArgoCD CMP plumbing.
2026-07-20 11:49:23 -07:00
Story Crater BotandClaude Haiku 4.5 d282ae1aa0 feat(argocd): deploy SOPS CMP plugin for secret decryption — Stage 0 grafana
Adds ConfigManagementPlugin (CMP) sidecar to argocd-repoServer. Plugin decrypts
*.enc.yaml files with age key from sops-age Secret, emits plain Kubernetes Secrets.

Stage 0: grafana only (2 Secrets: grafana-oidc + new grafana-admin). Updates
grafana-values.yaml to wire admin.existingSecret (chart-native support).

CMP Application (00-secrets.yaml) syncs at wave 0 before grafana/loki/authentik.
Decryption happens on-demand during sync, no pre-built Secret commits. Stages 1-4
(loki/authentik/forgejo/temporal) extend plugin script incrementally after
verification.

Co-Authored-By: Claude Haiku 4.5 <[email protected]>
2026-07-20 11:30:06 -07:00
Story Crater Bot 063308308f feat(cloudflared): wire tunnel token secret and document bootstrap
- Create SOPS-encrypted cloudflared-secrets.enc.yaml with tunnel token
- Add Cloudflare vars to .env.example (CLOUDFLARE_CONNECTOR_TOKEN, ACCOUNT_ID, TUNNEL_ID, API_TOKEN)
- Document Phase 0 cloudflared-token Secret creation in BOOTSTRAP.md (manual step until CMP plugin wires it)
- Note: Cloudflare-side TCP routing (cp1.homelab -> 192.168.1.213:6443, etc.) must be configured manually in Zero Trust dashboard

Tunnel already deployed as ArgoCD Application in k8s/argocd/apps/60-applications.yaml (wave 8); this closes the missing Secret gap and documents the bootstrap path.
2026-07-20 10:50:20 -07:00
Story Crater Bot a207c56637 fix(k8s,docs): scale ddb-cluster to single instance, pin minio to storage namespace, document 3-CP topology in USAGE 2026-07-20 08:22:53 -07:00
Story Crater Bot c759481ea6 refactor(argocd): replace wave/layer/phase schemes with two-phase bootstrap + app-of-apps and document both CD scopes — fixes self-hosted-git chicken-egg and stale paths 2026-07-20 08:22:53 -07:00
Story Crater Bot 15b1ec6ad4 feat(terraform): restructure control planes into a 3-node map with LAN etcd advertise and live machine CA — enables talos-cp-1/2/3 HA and drops worker configs 2026-07-20 08:22:53 -07:00
Story Crater Bot f7a8df0514 chore: remove GITOPS_ARCHITECTURE.md — scratch planning doc, not meant for the repo 2026-07-19 09:30:48 -07:00
Story Crater Bot 32281ee923 chore: remove scratch planning docs — not meant for the repo 2026-07-19 09:30:25 -07:00
Story Crater Bot 578a707867 feat(gitops): migrate domain to riotpiao.com, add CNPG + Forgejo HA on Redis/Postgres, wire ArgoCD apps — enables cluster rebuild after etcd wipe and unblocks the git-source chicken-egg via standalone Helm-source Applications 2026-07-19 09:29:17 -07:00
Story Crater Bot c6493f14ae feat(ci,iac): Consolidate Forgejo CI workflows and add Talos Terraform IaC
Consolidate three separate Forgejo Actions (argocd-sync, security-scan, validate-k8s) into single cluster-ci workflow for cleaner CI/CD pipeline with proper job sequencing and reduced auth overhead.

Add Terraform configuration for Talos cluster machine configs:
- Provider setup for Talos
- Centralized variables for CP and worker configs
- Template-based config generation for controlplane.yaml and worker-*.yaml
- Sensitive data separated in terraform.tfvars (gitignored)
- Local state tracking for infrastructure
2026-07-17 23:44:08 -07:00
Story Crater Bot 7437078f23 fix(ci): Correct Forgejo Actions template syntax for git clone auth
Use proper Forgejo variables: gitea.server_url, gitea.repository
Construct CLONE_URL correctly: https://user:token@host/repo.git
Use bash parameter expansion to strip https:// prefix

Removes invalid Forgejo filter syntax (| replace)
2026-07-17 12:27:40 -07:00
Story Crater Bot f8b19d9f55 fix(ci): Replace GitHub actions/checkout with Forgejo auth
Use CI_RUNNER and CI_RUNNER_SECRET for repo clone authentication.
Embed credentials in git clone URL: https://user:token@host/repo.git

Removes dependency on GITHUB_TOKEN (GitHub-specific) and improves Forgejo compatibility.
2026-07-17 12:19:23 -07:00
Story Crater Bot c000ddb402 fix(ci): Remove stale kustomize before reinstall in validate-k8s workflow
Prevent 'kustomize exists' error when downloading tools in CI runner.
Use -f flag on mv commands to force overwrite.
2026-07-17 11:31:44 -07:00
Story Crater Bot 888c4f5493 fix(ingress): Add nginx LoadBalancer service to GitOps — removes 503 error
Remove manual nginx-controller-svc.yaml (duplicate with Helm-managed service).
Helm chart creates LoadBalancer service automatically. Manual manifest caused conflicts.

Ingress controller now solely managed by Helm chart values.
2026-07-17 00:46:27 -07:00
Story Crater Bot 403e495fe9 fix(ingress): Add nginx LoadBalancer service to GitOps — fixes 503 error
Recreate ingress-nginx-controller LoadBalancer service that was deleted.
Add to k8s/bootstrap/ingress/ kustomization for ArgoCD management.

LoadBalancer assigned IP: 192.168.1.160 (via MetalLB)
ArgoCD now accessible via: https://192.168.1.160/ (or update DNS)
2026-07-16 23:35:02 -07:00
Story Crater Bot f30771a78a feat(minio): Expand CRDs to include Policies and Users — full YAML-driven resource creation
Add MinIO Policies and Users via CRD alongside Buckets.

Resources now declarative:
- Bucket: riotpiao-models (versioning enabled)
- Policy: policy-ollama (scoped bucket access)
- User: user-ollama (service account for Ollama/LLM)

Access keys can be overridden via SOPS or kustomize overlays.
All MinIO resource creation now git-tracked and version controlled.
2026-07-16 14:55:50 -07:00
Story Crater Bot e2e17ae0fb feat(data): Add CNPG cluster + database schema initialization
Create production PostgreSQL cluster via CNPG (3-node HA, Longhorn storage).

Schema initialization Job creates schemas for:
- Authentik (identity provider)
- Temporal (workflow engine)
- Vault (secrets management)
- App (generic application databases)

Database layer now captures complete IaC for stateful infrastructure.
Services find ready schemas when deployed.
2026-07-16 14:53:59 -07:00
Story Crater Bot c230b3ee45 feat(minio): Add MinIO Bucket CRD for riotpiao-models — replaces shell script setup 2026-07-16 14:31:33 -07:00
Story Crater Bot 2d7330798b refactor(k8s): Reorganize into 5-layer structure with production kustomizations 2026-07-16 14:28:19 -07:00
Story Crater Bot a81b9b6169 refactor(ci-cd): Replace Terraform pipeline with GitOps validation and ArgoCD sync
Delete old terraform-apply.yml (terraform fmt/init/validate/plan/apply).

Create new GitOps CI/CD:
- validate-k8s.yaml: YAML lint, kubeval, kustomize build, ArgoCD validation
- argocd-sync.yaml: Auto-sync homelab-root on main branch
- security-scan.yaml: Trivy, Polaris, secret detection
- .yamllint.yaml: YAML linting configuration

Add documentation (.forgejo/CI-CD.md) and architecture guides.

Git is now single source of truth. CI validates, ArgoCD deploys.
2026-07-16 12:52:25 -07:00
Story Crater Bot a860de94da refactor: remove terraform entirely, migrate to pure GitOps (ArgoCD)
Delete entire terraform/ directory.

Architecture: Terraform + ArgoCD → ArgoCD only
- Single tool: ArgoCD manages all infrastructure and applications
- Source of truth: git only (k8s/ directory)
- Continuous reconciliation: no manual apply needed
- Simpler state: no tfstate backend, no state files

Next: Migrate all Terraform resources to k8s/ YAML manifests
and ArgoCD Applications (namespaces, storage classes, Helm releases,
RBAC, network policies, Authentik config).
2026-07-16 11:04:52 -07:00
Story Crater Bot 94a2bd648c fix: remove kubeconfig references, use try() for pod runtime files
- main.tf: remove kubeconfig_path local (no longer used with direct auth)
- providers.tf: wrap file() with try() to handle plan-time on non-pod systems

try() allows terraform plan to work locally; at runtime in pod, files exist and are used.
2026-07-15 21:36:05 -07:00
Story Crater Bot be2e5c321f fix: use direct in-cluster kubernetes auth instead of kubeconfig file
- providers.tf: use host + token + ca_crt from mounted service account secrets
- workflow: remove kubeconfig generation step (no longer needed)
- variables.tf: remove unused kubeconfig_path variable

This is the standard pattern for running terraform inside k8s pods.
2026-07-15 21:26:09 -07:00
Story Crater Bot 9d1d79b774 fix: kubeconfig path default for runner container — use /tmp/kubeconfig not local path 2026-07-15 21:03:11 -07:00
Story Crater Bot 72ab6b6973 fix: terraform fmt — normalize formatting across all files 2026-07-15 21:01:55 -07:00
Story Crater Bot 7114fc8fc9 fix:Update the home lab repo 2026-07-15 20:40:02 -07:00
Story Crater Bot fe7b749951 fix: terraform init backend config — use endpoint with inline credentials 2026-07-15 19:18:17 -07:00
Story Crater Bot 694350634b fix: use endpoints.s3 for S3 backend (endpoint deprecated in TF 1.8+) 2026-07-15 19:16:13 -07:00
Story Crater Bot ab103f00f0 fix: core-cli OAuth2 + S3 backend + admin group
- Link core-cli app to OAuth2 provider (was hardcoded to 0)
- Add core-cli user to authentik_admins for CI access
- Fix terraform init: use 'endpoint' not 'endpoints.s3' for S3 backend
  (Terraform 1.9.4 compatibility, matches state.tf config)
2026-07-15 19:09:59 -07:00
Story Crater Bot e4d645eae9 feat: Terraform CI via Forgejo Actions + MinIO S3 state backend
- ArgoCD manages MinIO (phase 0), Terraform manages infrastructure
- Runner workflow: pulls state from S3, validates, plans, applies
- 34 resources imported to state, S3 backend operational
- Fixed AppProject repos, S3 endpoint deprecation, runner package manager
2026-07-15 18:48:32 -07:00
Story Crater Bot ab76e40d05 revert(phase4): Remove Pod Job approach for Terraform apply
Reverting Phase 4 Pod Job implementation in favor of CI runner (Forgejo Actions).

Deleted:
- k8s/argocd/apps/phase4-terraform-0.yaml
- k8s/hooks/phase4/ (terraform-apply-hook.yaml, terraform-rbac.yaml, terraform-s3-secrets.enc.yaml)

Reason: Pod Job approach had limitations (eviction, timeouts, pod security policies).
Next: Implement Forgejo Actions CI workflow for terraform apply.
2026-07-15 18:07:54 -07:00
Story Crater Bot b568c015e2 feat(phase4): ArgoCD-driven Terraform with PVC imports
Phase 4 implementation (true IaC):
- ArgoCD Application: terraform-apply (PostSync Hook Job)
- Hook Job runs: terraform init && terraform apply -auto-approve
- ServiceAccount + ClusterRole for cluster-admin
- SOPS-encrypted S3 credentials (terraform-s3-secrets.enc.yaml)
- Pre-commit hook blocks local 'terraform apply'
- In-cluster kubeconfig for Kubernetes provider
- AWS credentials file with minio profile

State imports:
- Imported kubernetes_persistent_volume_claim.portainer (dashboard/portainer)
- Imported kubernetes_persistent_volume_claim.grafana (logging/grafana)
- Imported kubernetes_persistent_volume_claim.loki (logging/storage-loki-0)

Workflow:
1. Edit terraform/*.tf files
2. git push to main
3. ArgoCD detects changes in k8s/hooks/phase4
4. Hook Job automatically runs terraform apply
5. No manual 'terraform apply' needed ever again
2026-07-15 17:48:18 -07:00
Story Crater Bot e71c7ad37e feat(phase4): ArgoCD-driven Terraform apply via PostSync Hook Job
- Create Phase 4 ArgoCD Application (terraform-apply)
- PostSync Hook Job runs: terraform init && terraform apply -auto-approve
- ServiceAccount + ClusterRole for cluster-admin RBAC
- S3 credentials encrypted with SOPS (terraform-s3-secrets.enc.yaml)
- Pre-commit hook blocks local 'terraform apply' — all changes via git push
- True IaC: modify terraform/*.tf → git push → ArgoCD applies automatically
2026-07-15 16:39:27 -07:00
Story Crater Bot 842360288d docs(terraform): add state management script and best practices guide 2026-07-15 16:31:55 -07:00
Story Crater Bot f18f96eb5b chore(phase4): stub helmfile — all releases managed by Terraform + ArgoCD 2026-07-15 16:27:13 -07:00
Story Crater Bot d655726eca feat(helmfile): remove phase3 releases (authentik, vault, story-crater, ollama) — ArgoCD-managed. Keep temporal 2026-07-15 16:26:30 -07:00
Story Crater Bot ff22027c7a refactor(argocd): phase3 reduced to authentik only (remove vault, temporal, ollama, story-crater) 2026-07-15 16:24:45 -07:00
Story Crater Bot f158512261 feat(argocd): create phase3 Applications (authentik, vault, temporal, ollama, story-crater) with SOPS secrets and Hook Jobs 2026-07-15 16:23:30 -07:00
Story Crater Bot 69d2240cf4 fix(argocd): use homelab-ca wildcard TLS instead of --insecure mode 2026-07-15 16:21:31 -07:00
Story Crater Bot 543105bf46 feat(helmfile): remove phase2 releases (cloudnative-pg, loki, grafana, prometheus, forgejo, forgejo-runner) — ArgoCD-managed 2026-07-15 16:07:42 -07:00
Story Crater Bot 7d1eb09486 feat(argocd): add phase2 Hook Jobs (CNPG, Prometheus, Forgejo-Runner) and update Applications to multi-source 2026-07-15 16:05:34 -07:00
Story Crater Bot 0dddf15dc8 feat(argocd): add SOPS-encrypted secrets for phase2 releases (loki, grafana, forgejo) 2026-07-15 16:02:20 -07:00
Story Crater Bot 37ee3dc5b1 feat(argocd): create phase2 Applications (prometheus, cloudnative-pg, loki, grafana, forgejo, forgejo-runner) 2026-07-15 15:12:36 -07:00
Story Crater Bot f864e3dc51 feat(argocd): remove claude-terminal and blackbox-exporter from phase1 2026-07-15 15:11:30 -07:00
Story Crater Bot 96e40916e3 feat(phase1): remove 9 hookless releases from helmfile — now ArgoCD-managed
Removed releases (all now managed via ArgoCD Applications):
- strimzi-operator
- kafka-cluster
- kmsvc-redis
- queue-crd
- management-service
- promtail
- blackbox-exporter
- portainer
- claude-terminal

Helmfile now contains only hook-heavy releases (phases 2-3) + argocd (self-referential, never migrates).

Next: Verify helmfile diff is clean, then confirm all 9 apps Synced/Healthy in ArgoCD.
2026-07-15 15:07:39 -07:00
Story Crater Bot 04100232ef feat(phase1): create ArgoCD Applications for 9 hookless releases
Created Applications for Phase 1 migration (no presync/postsync hooks):
- strimzi-operator (strimzi/strimzi-kafka-operator v0.46.0)
- kmsvc-redis (bitnami/redis v20.6.0)
- kafka-cluster (local chart k8s/sqs/charts/kafka-cluster)
- queue-crd (local chart k8s/sqs/charts/queue-crd)
- management-service (local chart k8s/sqs/charts/management-service)
- promtail (grafana/promtail)
- blackbox-exporter (prometheus-community/prometheus-blackbox-exporter ~11)
- portainer (portainer/portainer)
- claude-terminal (local chart k8s/dev-tools)

Organized by sync wave: 0 (bootstrap), 1 (messaging/observability), 3 (dashboards/tools).
All configured with auto-sync, CreateNamespace, prune, selfHeal.

Next: Remove corresponding release blocks from helmfile.yaml.gotmpl per migration guide
(one release at a time, verify helmfile diff is clean).

Applications applied to cluster; awaiting helmfile cleanup to finalize migration.
2026-07-15 15:06:02 -07:00
Story Crater Bot 09980aa41d docs(phase1): create migration guide for 9 hookless releases
Detailed Phase 1 workflow:
- Template Application spec (Helm source, values, sync policy)
- Per-release migration pattern (create → test → remove → commit)
- Helmfile ↔ ArgoCD mapping table
- Local chart handling (source.path vs source.chart)
- Verification checklist
- Rollback instructions

Reference: execute one release at a time, verify before next.
2026-07-15 15:05:12 -07:00
Story Crater Bot 5144ab732d feat(phase0): configure ArgoCD SOPS decryption + update encrypted secrets
Phase 0 continuation: enable ArgoCD to decrypt SOPS-encrypted secrets on sync.

1. Update ArgoCD Helm values (terraform/argocd-bootstrap.tf):
   - Add SOPS_AGE_KEY_FILE env var to repoServer
   - Mount sops-age K8s Secret at /home/argocd/.sops
   - Add ConfigManagementPlugin for SOPS (detects *.enc.yaml files)

2. Update encrypted secrets with real values:
   - k8s/base/secrets.enc.yaml: encrypted with actual service credentials
   - All secret values encrypted at rest in git
   - ArgoCD decrypts on sync using K8s Secret + AGE key

Prerequisites:
  - K8s Secret created: kubectl create secret generic sops-age -n argocd --from-file=keys.txt=/Users/rockliang/.sops/key.txt
  - SOPS_AGE_KEY_FILE env var set in ArgoCD repoServer (done above)

Next: Phase 1 — migrate 9 hookless releases to ArgoCD + create Applications that reference encrypted secrets.
2026-07-15 15:04:26 -07:00
Story Crater BotandClaude Haiku 4.5 4379f3cb21 feat(phase0): setup SOPS for encrypted secret management
Phase 0 groundwork for helmfile→ArgoCD migration using SOPS (Secrets Operations):

1. Install SOPS + AGE encryption
   - AGE key generated and stored locally at ~/.sops/key.txt
   - Public key embedded in .sops.yaml for file encryption rules

2. Create K8s Secret for AGE private key
   - kubectl: create secret generic sops-age -n argocd --from-file=keys.txt=~/.sops/key.txt
   - ArgoCD will use this key to decrypt secrets at sync time

3. Encrypt initial secrets
   - k8s/base/secrets.enc.yaml: AES256_GCM encrypted secrets for all services
   - Placeholder values (will be replaced with real values per environment)
   - Secrets never visible in git (encrypted at rest)

4. Configure SOPS
   - .sops.yaml: creation rules for k8s/*/secrets.enc.yaml files
   - All future secret files auto-encrypt on edit (sops -e)

Setup: Store AGE key as K8s Secret in argocd namespace:
  export KUBECONFIG=cluster-config/kubeconfig
  kubectl create secret generic sops-age -n argocd --from-file=keys.txt=~/.sops/key.txt

Next: Configure ArgoCD Helm plugin to decrypt secrets on sync (Phase 0 continuation).

Co-Authored-By: Claude Haiku 4.5 <[email protected]>
2026-07-15 15:00:06 -07:00
Story Crater Bot d2f4b3c7e4 Revert "feat(phase0): bootstrap External Secrets Operator and fix helmfile dual-ownership"
This reverts commit e7f3409d0f.
2026-07-15 14:59:54 -07:00
Story Crater BotandClaude Haiku 4.5 e7f3409d0f feat(phase0): bootstrap External Secrets Operator and fix helmfile dual-ownership
Phase 0 groundwork for helmfile→ArgoCD migration:

1. Remove 3 bootstrap releases from helmfile (cert-manager, reloader, ingress-nginx)
   — already managed by terraform/bootstrap-releases.tf; eliminates dual-ownership

2. Bootstrap ESO (External Secrets Operator) as TF-managed release
   — required for all ExternalSecret resources in phases 1-3
   — added to bootstrap-releases.tf + helm-repositories.tf

3. Create ClusterSecretStore connecting ESO to Vault (K8s auth)
   — enables per-namespace/per-release secret injection
   — vault config documented in docs/PHASE0-ESO-VAULT-SETUP.md (manual setup)

4. Fix argocd-bootstrap.tf CA cert copy: use jq instead of sed for cleaner metadata handling

Changes:
- helmfile.yaml.gotmpl: remove cert-manager/reloader/ingress-nginx blocks
- terraform/bootstrap-releases.tf: add external-secrets release
- terraform/helm-repositories.tf: add external-secrets Helm repo
- k8s/external-secrets/clustersecretstore.yaml: ESO→Vault ClusterSecretStore
- k8s/argocd/apps/0-wave-0.yaml: stub wave 0 applications (schema fix, rewrite pending Phase 1)
- docs/PHASE0-ESO-VAULT-SETUP.md: manual ESO-Vault auth setup procedure

Next: Phase 1 will incrementally rewrite ArgoCD Applications + migrate helmfile releases.

Co-Authored-By: Claude Haiku 4.5 <[email protected]>
2026-07-15 14:53:16 -07:00
Story Crater BotandClaude Haiku 4.5 23ec31bd6d feat(terraform): import Longhorn StorageClasses and app PVCs to Terraform state
- Phase 1: longhorn, longhorn-kafka StorageClasses (cluster-wide defaults)
- Phase 2 pilot: grafana, loki, portainer, forgejo PVCs
- All imports protected by lifecycle.prevent_destroy
- Removes Helm annotations (meta.helm.sh/*) to prevent dual-ownership conflicts
- Remote state backend (MinIO S3) syncs automatically on plan/apply
- Import-only approach: zero data loss, existing volumes untouched
- See terraform/LONGHORN_PVC_IMPORT.md for execution record

Co-Authored-By: Claude Haiku 4.5 <[email protected]>
2026-07-15 12:22:35 -07:00
Story Crater Bot 421086f845 docs(iac): enforce single source of truth for infrastructure
Add IaC practice section to coding-standards.md:
- All infrastructure state via Terraform or Helm (never ad-hoc scripts)
- Clear division: Terraform owns helm releases/namespaces/storage/state
- Anti-pattern: split bucket definitions across multiple files
- Bootstrap-only exception: document one-time setup with rationale

Rationale: prevents state drift, credential duplication, and unclear ownership.
2026-07-14 23:37:54 -07:00
Story Crater Bot bd00bca9bf chore: remove terraform cache from git tracking 2026-07-14 23:34:01 -07:00
Story Crater Bot 397edf632f fix: correct gitignore patterns for terraform state and cache
Remove malformed line and clarify rules:
- terraform/.terraform/ (local provider cache)
- terraform/*.tfstate* (local state backups)
- skills-lock.json (lock file)

All TF state now remote (MinIO S3), local files safe to exclude.
2026-07-14 23:33:50 -07:00
Story Crater Bot 13557184c6 chore: update terraform dependencies and config
terraform.lock.hcl updated with provider versions (goauthentik 2024.12.1).
Regenerated from current provider blocks.
2026-07-14 23:33:16 -07:00
Story Crater Bot 6d554961c2 feat(terraform): enable S3 remote state backend (MinIO)
Migrate terraform state from local file to MinIO S3 bucket (terraform-state).
Backend config: https://minio-api.riotpiao.homelab.com (external endpoint).
State now persisted remotely, shared across team, safe for cluster rebuild.

Also added terraform-state bucket to MinIO managed buckets.
2026-07-14 23:30:45 -07:00
Story Crater Bot 3eccf9f653 test(argocd): add label to vault app to verify GitOps flow
Add test-gitops=true label to vault Application to demonstrate end-to-end
GitOps sync: commit push → ArgoCD detects change → applies label to live app.
Tests that root-app watches k8s/argocd/apps/ and propagates changes.
2026-07-14 23:24:12 -07:00
Story Crater Bot 52cb895cda feat(minio): add loki storage buckets (chunks/ruler/admin/index)
Move loki bucket creation from helmfile post-hook to TF-managed buckets array.
Now all MinIO buckets (6 total) declared in terraform/minio.tf for IaC completeness.
2026-07-14 17:00:09 -07:00
Story Crater Bot bca247a763 feat(authentik): import 24 resources to TF; chore(bootstrap): add cilium to TF
Import all live authentik resources (groups, users, oauth2 providers, applications)
into terraform state via authentik-generated.tf. Provider config in authentik-config.tf.
Resources are drift-free and match live cluster.

Add cilium CNI to bootstrap helm_release.for_each (1.19.5, kube-system).
Cilium was unmanaged (helmfile-only); now IaC-owned. Critical path for
cluster rebuild recovery. Adds cilium repo to helm-repositories.tf.
2026-07-14 16:36:12 -07:00
Story Crater Bot 9e3781a069 fix(minio): migrate to official chart, TF-owned
Bitnami wiped Docker Hub catalog (bitnami/minio: 0 tags), chart 14.1.0
dead on ImagePullBackOff. Move to minio/minio 5.4.0 (quay.io) as one TF
helm_release. Add longhorn-xfs SC: default SC ext4 mkfs on 100Gi exceeds
kubelet mount timeout, xfs near-instant. Drop minio ArgoCD Apps (TF owns
now, kills dual-controller conflict). Fix double base64 on OIDC secret.
2026-07-14 16:07:03 -07:00
Story Crater Bot ebeb4948d4 Add: minio-operator TF management (v4.5.8 downgrade) - WIP due to helm conflicts 2026-07-14 15:04:50 -07:00
Story Crater Bot a3190abe50 Revert: Use minio-creds secret with MINIO_ROOT_* keys (operator expected format) 2026-07-14 14:37:41 -07:00
Story Crater Bot dddc7a524a Fix: Tenant credentials secret reference from minio-creds to minio 2026-07-14 14:34:49 -07:00
Story Crater Bot ed7be6f229 TF: Add minio-operator Helm repo to ArgoCD config + AppProject sourceRepos 2026-07-14 14:32:45 -07:00
Story Crater Bot 418ab7bfc2 Fix: minio-operator uses official MinIO Operator Helm chart 2026-07-14 14:26:25 -07:00
Story Crater Bot ca8525c625 Add minio-operator Application to deploy operator before Tenant 2026-07-14 14:26:01 -07:00
Story Crater Bot 8a3a892cbd Fix: Inject homelab-ca cert into ArgoCD repo-server
- Mount homelab-ca-secret for TLS verification
- Allows repo-server to reach forgejo.riotpiao.homelab.com
- Fixes x509 certificate verification error
2026-07-14 14:01:14 -07:00
Story Crater Bot cdacdd8d11 Re-enable cert-manager manifests for TF import
- ClusterIssuers + Certificates now back in TF
- Will import existing live resources
2026-07-14 13:55:08 -07:00
Story Crater Bot 263a48a22d Fix: ArgoCD AppProject sourceRepos for correct forgejo URL
- Changed from forgejo.forge.riotpiao.homelab.com/rock/* to forgejo.riotpiao.homelab.com/riotpiao.com/*
- Allows homelab root app to access workload app manifests
2026-07-14 13:52:53 -07:00
Story Crater Bot 47c0301a43 Step 2: ArgoCD app-of-apps manifests for 19 workloads
Wave 0: minio, strimzi-operator, kmsvc-redis, prometheus
Wave 1: vault, loki, cloudnative-pg, authentik (manual-sync), temporal, kafka-cluster
Wave 2: queue-crd, management-service, grafana, promtail, forgejo
Wave 3: forgejo-runner, portainer

All auto-sync except authentik (manual-sync only for IAM safety)
2026-07-14 13:52:18 -07:00
Story Crater Bot edd4ea7fe4 Fix: set ingress-nginx to privileged PodSecurity level
- privileged level allows hostPort (80/443) required for nginx
- Other namespaces remain at baseline for security
- Cleaner than exempting namespace entirely
2026-07-14 13:43:56 -07:00
Story Crater Bot 086ad9f9a9 Fix: exempt ingress-nginx from PodSecurity policy
- restricted policy forbids hostPort (80/443) — broke nginx
- Remove pod-security labels from ingress-nginx namespace entirely
- Other namespaces remain at baseline level
2026-07-14 13:41:36 -07:00
Story Crater Bot 3eabb847fd Re-add ingress-nginx to TF bootstrap (PodSecurity policy fixed)
- ingress-nginx now has restricted policy level (allows hostPort)
- Previous timeout was due to policy blocking pod deployment
- Re-importing helm release to TF management
2026-07-14 13:37:22 -07:00
Story Crater Bot 6109477bf9 Skip TF management of ingress-nginx (helm timeout issues)
- ingress-nginx already deployed and working in cluster
- Helm updates timeout repeatedly (5+ min with context deadline exceeded)
- Remove from bootstrap releases; manage separately via helm/kubectl
- cert-manager + reloader continue via TF
2026-07-14 13:32:40 -07:00
Story Crater Bot 20c634fd03 Fix: ingress-nginx PodSecurity policy enforcement level
- ingress-nginx requires hostPort (80/443) which is forbidden at baseline level
- Change to restricted enforcement level to allow hostPort
- Other namespaces remain at baseline for security
2026-07-14 13:31:16 -07:00
Story Crater Bot 074b43e1f2 Temp: disable kubernetes_manifest cert-manager resources (already live)
- Will import separately after helm issues resolved
- Avoids re-create conflicts during bootstrap apply
2026-07-14 13:26:26 -07:00
Story Crater Bot c2e084c7c2 Fix: downgrade ArgoCD to 7.3.3, ignore helm metadata drift
- ArgoCD 7.9.1 -> 7.3.3 (match live cluster)
- Ignore helm release metadata in lifecycle rules
- Prevents unnecessary upgrade attempts
2026-07-14 13:16:55 -07:00
Story Crater Bot dd608d3231 Step 1 complete: Bootstrap layer with ArgoCD, cert-manager, namespaces imported to TF
- ArgoCD migrated to argocd namespace
- Cert-manager issuers/certs created
- 20 namespaces imported with pod-security labels
- S3 backend temporarily offline (MinIO), using local backup
- Pending: Remove metadata drift from helm releases, re-apply
2026-07-14 13:14:46 -07:00
Story Crater Bot 9a4d486b86 feat: Terraform foundation for cluster & app bootstrap
Phase 1 infrastructure-as-code setup:
- Core providers (kubernetes, helm, null)
- 15 Helm repositories (grafana, minio, prometheus, etc.)
- Namespace scaffolding (15 namespaces with pod-security labels)
- Storage classes (longhorn, longhorn-kafka with prevent_destroy)
- TLS certificate bootstrap (selfsigned, CA, wildcard cert)
- Remote state backend config (local for now, S3/GCS TODO)
- Variable definitions for all secrets/OIDC clients

Tested: terraform plan passes with no changes (bootstrap infrastructure ready)
Next: Create 25 helm_release resources (Phase 2-4)

Kept helmfile intact; network/Cilium managed via helmfile (no config risk)
Co-Authored-By: Claude Haiku 4.5 <[email protected]>
2026-07-14 09:27:24 -07:00
Story Crater Bot 4d32e2765b feat: three-tier log aggregation for Loki
Critical services (iam/monitoring/temporal/cicd) keep 100% logs.
Others get 50% sampling + selective drops (health/debug noise).
Balances log volume (40-50% reduction) with error visibility.
2026-07-13 17:52:59 -07:00
Story Crater Bot 0d4e98d88b feat: track full CoreDNS Deployment manifest, add topologySpreadConstraints
CoreDNS is Talos-bootstrapped and previously untracked except for its
ConfigMap. Pull the full live spec into one file as the single source of
truth, add topologySpreadConstraints so the 2 replicas don't land on the
same node. ScheduleAnyway (not DoNotSchedule) to avoid blocking scheduling
if a node is briefly unavailable.
2026-07-13 16:51:49 -07:00
Story Crater Bot ebf97f573e feat: spread management-service pods across nodes via topologySpreadConstraints
3-9 replicas (HPA) previously relied on implicit scheduler spreading.
ScheduleAnyway (not DoNotSchedule) so pods still get scheduled if a node
is briefly unavailable, just less evenly.
2026-07-13 16:43:58 -07:00
Story Crater Bot 05b088ca48 docs: update USAGE.md for core CLI + IAM management
- Rename talos → core CLI references
- Add IAM management section (Authentik apps, groups, users)
- Add workflows for secret rotation and user management
- Link to detailed core CLI docs (~/workplace/core/USAGE.md)
- Add port-forwarding and troubleshooting tips
2026-07-13 14:12:12 -07:00
Story Crater Bot b33fdde5b7 fix: temporal service config - add explicit ClusterIP services for history/matching 2026-07-13 13:20:50 -07:00
Story Crater Bot cd1c5691e8 feat: point SQS charts to public GHCR image
Forgejo registry unreachable from worker nodes (network isolation +
host-to-ClusterIP routing gaps). Both management-service and queue-operator
now ship from the same public GHCR image, with queue-operator selected via
command override.
2026-07-13 10:31:22 -07:00
Story Crater Bot e6d7626fb6 remove: strip all oauth2-proxy deployments
- Delete oauth2-proxy helm releases from helmfile (temporal, kmsvc, longhorn, portainer)
- Remove oauth2-proxy manifests and ingress redirects
- Add direct ingress for kmsvc management service
- Update temporal/portainer/longhorn ingress comments to reflect direct service exposure

Services now accessible without oauth2-proxy layer.
2026-07-11 19:36:31 -07:00
Story Crater Bot e1d0cfd70b k8s/aux: add cert-manager longhorn dashboard forge dev-tools and shadowsocks
- cert-manager ClusterIssuers (LetsEncrypt + homelab-ca)
- Longhorn storage dashboard
- Portainer dashboard config
- Forgejo git service
- Claude terminal remote access
- Shadowsocks tunnel for remote access
2026-07-11 19:19:59 -07:00
Story Crater Bot 18f2f94f8e k8s/cilium: add lb-ipam pool configuration
- Cilium LB-IPAM pool (192.168.1.160-192.168.1.170)
- Fixed IP assignment for LoadBalancer services
2026-07-11 19:19:44 -07:00
Story Crater Bot 6d5a0ba205 k8s/services: add ingress networking portainer llm and project guides
- Nginx ingress + TLS termination (homelab-ca)
- Portainer container UI
- CoreDNS internal DNS rewrites
- DuckDNS DDNS updater
- Ollama LLM inference
- 8 project-usage guides (team reference)
2026-07-11 19:17:54 -07:00
Story Crater Bot 1c02e2b831 k8s/messaging: add kafka kmsvc and temporal workflows
- Kafka 3-broker cluster (RF=3, min-ISR=2)
- kmsvc SQS-like API on Kafka
- Redis dedup (standalone, can extend to HA)
- Temporal workflow orchestration (Cassandra backend)
2026-07-11 19:17:42 -07:00
Story Crater Bot 4ab596196e k8s/ci-cd: add forgejo gitops and argocd deployment
- Forgejo git forge + OCI registry
- Argo CD pull-based GitOps
- Private CA TLS (self-signed 10-year cert)
- Machine credentials scoped to repositories
2026-07-11 19:17:34 -07:00
Story Crater Bot 63d7256b9e k8s/monitoring: add prometheus grafana loki observability
- Loki log aggregation (MinIO backed, 10-day retention)
- Promtail daemonset (pod + talos journal logs)
- Prometheus + kube-state-metrics
- Grafana dashboards (6-row template per service)
2026-07-11 19:17:28 -07:00
Story Crater Bot 674c8f0d66 k8s/iam: add cloudnativepg postgres and vault + authentik
- PostgreSQL 3-replica HA with pgvector
- Vault S3 storage backend (MinIO)
- Authentik federated OIDC provider
- Vault auto-unseal via postStart hook
2026-07-11 19:17:22 -07:00
Story Crater Bot 36aea89e47 k8s/storage: add minio s3 with 3-way replication and oidc
- MinIO 3-node site replication (az-a/b/c)
- S3 backend for Loki chunks (10-day retention)
- OIDC integration with Authentik
- envFrom for secret injection
2026-07-11 19:16:56 -07:00
Story Crater Bot 11c26f3f29 k8s: add base namespace and pod disruption budgets
- Namespace setup script with PSP/RBAC
- PodDisruptionBudgets for all services (zero-downtime drain)
2026-07-11 19:16:50 -07:00
Story Crater Bot 11898733e8 infra: add helmfile and talos cluster configuration
- helmfile: 18 releases across 22 namespaces
- Pod disruption budgets for zero-downtime drain
- Nginx ingress with LoadBalancer + Cilium LB-IPAM
- Cluster bootstrap hooks
2026-07-11 19:16:41 -07:00
168 changed files with 5340 additions and 7818 deletions
+460
View File
@@ -0,0 +1,460 @@
# CI/CD Pipeline: GitOps Validation & Deployment
## Overview
Pure GitOps CI/CD pipeline using Forgejo Actions (self-hosted runner).
**Principle:** Validate in CI, deploy via ArgoCD (no manual steps).
```
git push
[CI: Validate]
├─ yamllint (YAML syntax)
├─ kubeval (K8s manifests)
├─ kustomize build (all layers)
├─ argocd validation (app definitions)
└─ security scan (secrets, best practices)
[If push to main]
└─ ArgoCD auto-syncs (if enabled)
```
## Workflows
### 1. validate-k8s.yaml (Mandatory)
**Trigger:** Any push/PR with k8s/ changes
**What it does:**
1. Lints all YAML files (`yamllint`)
2. Validates K8s manifests (`kubeval`)
3. Builds all kustomization layers
4. Validates ArgoCD applications
5. Reports results
**Duration:** ~2-3 minutes
**Status:**
- ✅ PASS: All layers build, manifests valid → OK to merge
- ❌ FAIL: Syntax error, invalid resource, build failed → Fix & push again
**Example output:**
```
=== Building k8s/infrastructure/ ===
✓ Infrastructure built successfully
Resources: 47
=== Building k8s/bootstrap/ ===
✓ Bootstrap built successfully
Resources: 23
```
**When to check:**
- After every commit
- Before merging PRs
- On every branch
### 2. argocd-sync.yaml (Recommended)
**Trigger:** Push to main only (k8s/ changed)
**What it does:**
1. Authenticates with ArgoCD
2. Syncs `homelab-root` application
3. Waits for sync to complete (5 min timeout)
4. Verifies all applications healthy
**Duration:** 1-5 minutes (depends on resources)
**Status:**
- ✅ SYNCED: All resources deployed to cluster
- ❌ FAILED: Sync error, pod crashes, etc. → Check ArgoCD UI for details
**When it runs:**
- Automatically after merge to main
- Only on k8s/ changes (not on docs)
**Manual trigger (if needed):**
```bash
# SSH to runner or use Forgejo UI
# Re-run failed workflow
# Or manually sync: argocd app sync homelab-root
```
**Requires secrets:**
- `ARGOCD_SERVER`: ArgoCD server URL (https://argocd.riotpiao.com)
- `ARGOCD_AUTH_TOKEN`: ArgoCD API token (generate via ArgoCD UI)
### 3. security-scan.yaml (Optional)
**Trigger:** Any push/PR with k8s/ changes
**What it does:**
1. Scans Dockerfiles for vulnerabilities (`trivy`)
2. Scans Helm charts for security issues
3. Audits K8s manifests (`polaris`)
4. Checks for hardcoded secrets
5. Verifies security best practices
**Duration:** ~3-5 minutes
**Status:**
- ✅ PASS: No critical issues
- ⚠️ WARNING: Best practice recommendations (non-blocking)
- ❌ FAIL: Hardcoded secrets found (must fix)
**Common issues:**
- Missing resource limits (warning)
- Privileged containers (warning)
- Hardcoded passwords (ERROR)
---
## File Structure
```
.forgejo/
├── workflows/ # CI/CD workflows
│ ├── validate-k8s.yaml # Validate manifests (required)
│ ├── argocd-sync.yaml # Sync to cluster (auto on main)
│ └── security-scan.yaml # Security checks (optional)
└── CI-CD.md # This file
```
---
## Setup Instructions
### 1. Install Forgejo Runner
```bash
# On runner machine (inside cluster or external)
forgejo-runner register \
--instance https://forgejo.riotpiao.com \
--token <registration-token> \
--name homelab-runner \
--labels docker
forgejo-runner daemon
```
### 2. Add ArgoCD Secrets to Forgejo
```bash
# Go to: Forgejo → Settings → Secrets
# Add:
ARGOCD_SERVER = https://argocd.riotpiao.com
ARGOCD_AUTH_TOKEN = <token> # Generate: argocd account generate-token
```
### 3. Generate ArgoCD Token
```bash
# Inside cluster
kubectl -n argocd port-forward svc/argocd-server 8080:443
# Go to: https://localhost:8080/user-info/api-tokens
# Create new token (CI/CD)
# Copy token to Forgejo secrets
```
---
## Workflow Execution
### When developer pushes to feature branch:
```
git push origin feature/new-service
Forgejo Actions triggered
validate-k8s.yaml runs:
✓ Lints YAML
✓ Validates manifests
✓ Builds kustomizations
✓ All pass → GitHub comment: "Ready to merge"
Developer opens PR
Reviewer checks:
- Code changes (YAML)
- Workflow results
- ArgoCD impact (diff)
PR merged to main
```
### When merged to main:
```
git merge feature/new-service → main
Forgejo Actions triggered
validate-k8s.yaml runs:
✓ Same validation as above
argocd-sync.yaml runs (if enabled):
✓ Syncs homelab-root
✓ Waits for sync
✓ Verifies health
✓ Resources deployed to cluster
Cluster state = git state
(No manual kubectl apply needed!)
```
---
## Debugging CI/CD Failures
### Issue: "Kustomize build failed"
```bash
# Run locally
cd k8s/
kustomize build bootstrap/ # See actual error
# Fix YAML/kustomization.yaml
# git push again
```
### Issue: "Kubeval validation failed"
```bash
# Check K8s manifest syntax
kubeval k8s/platform/minio/config.yaml
# Common issues:
# - Typos in apiVersion, kind, metadata
# - Missing required fields
# - Invalid references (namespace, service name)
```
### Issue: "ArgoCD sync failed"
```bash
# Check ArgoCD UI
# https://argocd.riotpiao.com → homelab-root
# Or CLI
argocd app get homelab-root
argocd app logs homelab-root --follow
# Common issues:
# - Missing namespace (fixed by infrastructure layer)
# - Invalid Helm chart version
# - Secret not found
# - Network policy blocking traffic
```
### Issue: "Security scan found hardcoded secret"
```bash
# Fix: Remove secret from YAML
# Add to SOPS encryption instead
# Or use ArgoCD Sealed Secrets
# (if SOPS not available)
```
---
## Viewing Results
### Forgejo Actions UI
```
Repository → Actions
├─ validate-k8s
│ ├─ ✅ Success (merge safe)
│ ├─ ❌ Failed (fix required)
│ └─ Logs (click "Steps" → "Summary")
├─ argocd-sync
│ ├─ ✅ Synced (deployed)
│ └─ ❌ Failed (check ArgoCD UI)
└─ security-scan
├─ ✅ Pass (no critical issues)
└─ ⚠️ Warning (review, non-blocking)
```
### ArgoCD UI
```
https://argocd.riotpiao.com
├─ homelab-root
│ ├─ Status: Synced ✓
│ ├─ Health: Healthy ✓
│ └─ Details (click to see resources)
├─ layer-1-bootstrap
├─ layer-2-platform
├─ layer-3-security
├─ layer-4-applications
└─ layer-5-data
```
---
## Common Tasks
### Add new service to cluster
```bash
# 1. Create directory and kustomization.yaml
mkdir -p k8s/applications/my-service
cat > k8s/applications/my-service/kustomization.yaml << EOF
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
namespace: my-namespace
helmCharts:
- name: my-chart
repo: https://charts.example.com
version: 1.0.0
releaseName: my-service
valuesFile: values.yaml
EOF
# 2. Add values.yaml
cp /template/values.yaml k8s/applications/my-service/
# 3. Commit and push
git add k8s/applications/my-service/
git commit -m "feat(apps): add my-service"
git push
# 4. CI validates
# 5. Merge to main
# 6. ArgoCD syncs automatically
# ✓ Service deployed to cluster
```
### Rollback a deployment
```bash
# 1. Find broken commit
git log --oneline k8s/ # Identify bad commit
# 2. Revert
git revert <commit-hash>
git push
# 3. CI validates (should pass)
# 4. Merge to main
# 5. ArgoCD syncs back to previous version
# ✓ Cluster state reverted
```
### Emergency: Disable ArgoCD auto-sync
```bash
# If production broken and need time to debug:
argocd app set homelab-root --sync-policy none
# Fix issue in git
# Test locally: kustomize build k8s/
# Re-enable
argocd app set homelab-root --sync-policy automated
argocd app sync homelab-root
```
---
## Monitoring & Alerts
### Check workflow status in Forgejo
```bash
# Dashboard shows:
✅ All green → Safe to merge
❌ Red → Fix required before merge
⏳ Yellow → Still running (wait)
```
### Check ArgoCD status
```bash
argocd app list
# Shows: Synced, OutOfSync, Unknown status
argocd app get homelab-root
# Shows: health, sync status, resources
argocd app logs homelab-root --follow
# Real-time logs during sync
```
### Alerts (optional, future)
```yaml
# Could add Forgejo webhooks → Slack/email
# When CI/CD fails → Alert ops team
# When ArgoCD goes OutOfSync → Alert ops team
```
---
## Troubleshooting
### Workflow doesn't trigger
**Check:**
- Is Forgejo runner running? `forgejo-runner daemon`
- Did you push to correct branch? (validate runs on all, argocd-sync only on main)
- Did path match filter? (must change k8s/ or .forgejo/workflows/)
### Workflow hangs/times out
**Check:**
- kustomize build → Check for dependency cycles
- argocd sync → Check cluster resources (storage full? network down?)
- security scan → Large image scan → Takes time
**Fix:**
- Increase timeout in workflow
- Optimize kustomization (remove unused resources)
- Add resource limits to pods
### ArgoCD token invalid
**Fix:**
```bash
# Regenerate token
argocd account generate-token
# Update Forgejo secret
# Settings → Secrets → ARGOCD_AUTH_TOKEN = <new-token>
```
---
## Best Practices
**DO:**
- Commit all K8s changes to git (no manual kubectl apply)
- Run validate-k8s locally before push
- Write descriptive commit messages (why this change?)
- Review workflow logs before merging
- Monitor ArgoCD sync after merge
**DON'T:**
- Push directly to main (always use PR)
- Skip workflow validation (it catches errors early)
- Ignore security scan warnings
- Manually `kubectl apply` (breaks GitOps)
- Edit resources in cluster (they revert via ArgoCD)
---
## Next Steps
1. **Setup Forgejo runner** (if not already running)
2. **Add ArgoCD secrets** to Forgejo
3. **Test workflows** on feature branch
4. **Merge to main** → Watch ArgoCD sync
5. **Celebrate:** Full GitOps pipeline working! 🎉
-206
View File
@@ -1,206 +0,0 @@
# Authentik Auth Integration for NextJS
## Current State
### Gateway Auth Status
| Endpoint | Auth Status | Notes |
|----------|-------------|-------|
| `/v1/chat/completions` | ❌ **OFF** | LLM routes have no auth middleware |
| `/v1/embeddings` | ❌ **OFF** | Same - no auth |
| `/v1/rerank` | ❌ **OFF** | Same - no auth |
| `X-Service: sqs` | ✅ **ON** | JWT validated via `internal/auth/jwt.go` |
| `/workflow` | ❌ **OFF** | Pass-through to Temporal |
**Auth module exists** at `homelab-frontend/internal/auth/jwt.go` but only wired for SQS.
LLM routes in `internal/proxy/proxy.go` have no auth middleware.
### Authentik App
Authentik app `local-llm` exists for LLM API auth:
- **Client ID**: `local-llm`
- **Client Secret**: `kubectl -n llm-serving get secret local-llm-jwt -o jsonpath='{.data.client-secret}' | base64 -d`
- **Token endpoint**: `https://authentik.riotpiao.com/application/o/token/`
- **Userinfo endpoint**: `https://authentik.riotpiao.com/application/o/userinfo/`
- **OIDC discovery**: `https://authentik.riotpiao.com/application/o/local-llm/.well-known/openid-configuration`
## Sign-in Methods
### 1. Resource Owner Password Credentials (ROPC)
Direct username/password login. Server-side only (needs client_secret).
```typescript
// API Route: app/api/auth/login/route.ts
const response = await fetch('https://authentik.riotpiao.com/application/o/token/', {
method: 'POST',
headers: { 'Content-Type': 'application/x-www-form-urlencoded' },
body: new URLSearchParams({
grant_type: 'password',
client_id: 'local-llm',
client_secret: process.env.AUTHENTIK_CLIENT_SECRET,
username: '[email protected]',
password: 'userpassword',
scope: 'openid email profile groups',
}),
});
const tokens = await response.json();
// { access_token, refresh_token, expires_in, token_type }
```
### 2. Authorization Code Flow (Browser Redirect)
Requires adding redirect URIs to `local-llm` Authentik app:
```python
# In k8s/infra/iam/scripts/authentik-provision.py, update:
"local-llm": {
...
"redirect_uris": [
"http://localhost:3000/api/auth/callback", # dev
"https://your-nextjs-app.com/api/auth/callback", # prod
],
}
```
Then standard OIDC flow:
1. Redirect to `https://authentik.riotpiao.com/application/o/authorize/?client_id=local-llm&redirect_uri=...&response_type=code&scope=openid email profile groups`
2. User logs in via Authentik UI
3. Callback receives `code`, exchange for tokens
## JWT Token Persistence
### Browser (localStorage)
```typescript
const TOKEN_KEY = 'llm_auth_token';
// Save
localStorage.setItem(TOKEN_KEY, JSON.stringify({
access_token: tokens.access_token,
refresh_token: tokens.refresh_token,
expires_at: Date.now() + tokens.expires_in * 1000,
}));
// Load
const stored = JSON.parse(localStorage.getItem(TOKEN_KEY) || 'null');
if (stored && stored.expires_at > Date.now()) {
// Token valid
}
// Clear (logout)
localStorage.removeItem(TOKEN_KEY);
```
### Server-side (HTTP-only cookies)
```typescript
// app/api/auth/login/route.ts
import { cookies } from 'next/headers';
// After successful login
cookies().set('llm_auth_token', JSON.stringify(tokens), {
httpOnly: true,
secure: process.env.NODE_ENV === 'production',
sameSite: 'lax',
maxAge: tokens.expires_in,
path: '/',
});
// Read in middleware or API routes
const tokenCookie = cookies().get('llm_auth_token');
const tokens = JSON.parse(tokenCookie?.value || 'null');
```
## Token Refresh
```typescript
async function refreshAccessToken(refresh_token: string) {
const response = await fetch('https://authentik.riotpiao.com/application/o/token/', {
method: 'POST',
headers: { 'Content-Type': 'application/x-www-form-urlencoded' },
body: new URLSearchParams({
grant_type: 'refresh_token',
client_id: 'local-llm',
client_secret: process.env.AUTHENTIK_CLIENT_SECRET,
refresh_token,
}),
});
return response.json();
}
```
## Environment Variables
```bash
# .env.local
AUTHENTIK_URL=https://authentik.riotpiao.com
AUTHENTIK_CLIENT_ID=local-llm
AUTHENTIK_CLIENT_SECRET=<from-secret>
# For client-side (public)
NEXT_PUBLIC_AUTHENTIK_URL=https://authentik.riotpiao.com
NEXT_PUBLIC_AUTHENTIK_CLIENT_ID=local-llm
```
## Using Token with LLM API
```typescript
const token = await getValidToken(); // from localStorage or cookie
const response = await fetch('https://api.riotpiao.com/v1/chat/completions', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'Authorization': `Bearer ${token}`, // JWT from Authentik
},
body: JSON.stringify({
model: 'reasoning',
messages: [{ role: 'user', content: 'Hello' }],
}),
});
```
## TODO
### Gateway-side (homelab-frontend)
- [ ] Wire `internal/auth/jwt.go` into LLM proxy handler (`internal/proxy/proxy.go`)
- [ ] Add `authRequired: true` to model config or create LLM-specific middleware
- [ ] Example pattern from SQS (in `internal/serviceadapter/router.go`):
```go
// In proxy.go ServeHTTP, before dispatching to LLM upstream:
if strings.HasPrefix(r.URL.Path, "/v1/") {
authHeader := r.Header.Get("Authorization")
claims, err := llmJWTAuth.ValidateBearerToken(authHeader)
if err != nil {
// Return 401/403
}
if !llmJWTAuth.CheckPermissions(claims, "llm:inference", "*") {
// Return 403 insufficient permissions
}
}
```
### Authentik-side
- [ ] Enable ROPC grant in Authentik provider settings (if not already)
- [ ] Add redirect URIs to `local-llm` app if browser OAuth flow needed:
```python
# k8s/infra/iam/scripts/authentik-provision.py
"local-llm": {
...
"redirect_uris": [
"http://localhost:3000/api/auth/callback",
"https://your-app.com/api/auth/callback",
],
}
```
### NextJS-side
- [ ] Until gateway auth is wired, LLM API works without token
- [ ] Once wired, add `Authorization: Bearer <token>` to all LLM requests
-80
View File
@@ -42,86 +42,6 @@ All logs + metrics centralized in Grafana for debugging
- **Secrets at rest** — Vault + encrypted etcd; credentials never in logs or ConfigMaps
- **Infrastructure-as-code** — Every service deployed via Helmfile; one `helmfile apply` recovers from total failure
## ArgoCD — GitOps Deployment Flow
**ArgoCD** pulls infrastructure changes from git and syncs the cluster automatically.
No manual `kubectl apply` — push to git, ArgoCD detects the change, and deploys within ~3 minutes.
```
Developer pushes to git
ArgoCD detects change (every 3 min or webhook)
Syncs manifests to cluster
Workloads reconcile automatically
```
Applications are deployed in waves (numbered 00, 10, 20, 30, ...) to respect dependencies —
storage deploys before databases, databases before applications.
### Tracked Git Repositories
ArgoCD monitors these repos for changes:
| Repository | Purpose |
|------------|----------|
| `https://github.com/Riotpiaole/riotpiao.homelab.com` | Main infrastructure repo (all manifests in `k8s/argocd/apps/`) |
| `https://forgejo.riotpiao.com/rock/*` | Any `rock/*` repo in in-cluster Forgejo (apps + configs) |
| `https://github.com/Riotpiaole/Poimen-*` | External Poimen services (memory, workflows) |
To deploy a new application: create a git repo, add an Application manifest to the homelab repo's
`k8s/argocd/apps/`, commit + push, and ArgoCD syncs within 3 minutes.
## Management Planes — Talos vs Kubernetes
This cluster has **two separate management planes**, each with different workflows:
| Plane | What it manages | Workflow | Tool |
|-------|-----------------|----------|------|
| **Talos (OS)** | Node configuration, kernel params, networking, CoreDNS, machine state | Edit `terraform/``terraform apply``make apply-cp` | `terraform` + `talosctl` |
| **Kubernetes (workloads)** | All pods, services, deployments, ingresses, databases | Edit `k8s/argocd/apps/``git push` → ArgoCD syncs | `git` + ArgoCD |
**Critical distinction:**
- **Kubernetes resources** (`k8s/**`) flow through **git → ArgoCD** — never use `kubectl apply`
- **Talos machine config** (`terraform/**`) uses **local `terraform apply`** (sanctioned exception — CI can't hold node credentials)
Example: To add a CoreDNS hostname rewrite, you edit `terraform/files/coredns/Corefile`, then:
```bash
cd terraform && terraform apply -var-file=terraform.tfvars.local
cd .. && make apply-cp # talosctl apply-config to all 3 control planes
```
But to add a new Kubernetes Deployment or update an Ingress, you only `git push`**never `kubectl apply`**.
### CoreDNS ConfigMap Ownership — Critical
⚠️ **Warning:** The `coredns` ConfigMap in `kube-system` namespace is **owned by Talos**, not ArgoCD or kubectl.
It is rendered from `terraform/files/coredns/Corefile` into Talos's machine config at bootstrap time.
**Do not `kubectl apply` or `kubectl edit` this ConfigMap directly.** Doing so transfers field ownership to kubectl's
client-side-apply mechanism, and Talos's inline-manifest controller will silently no-op on every future reconcile
(server-side-apply conflict, no error surfaced).
**To update CoreDNS (e.g., add a hostname rewrite):**
1. Edit `terraform/files/coredns/Corefile`
2. Commit + push
3. Run `cd terraform && terraform apply -var-file=terraform.tfvars.local`
4. Run `make apply-cp` to push config to all control planes
5. CoreDNS picks up changes via its `reload` plugin — no pod restart needed
**If you accidentally edited the ConfigMap directly and broke Talos's ownership:**
```bash
kubectl delete configmap coredns -n kube-system
# Wait ~30s for Talos's k8s.ManifestApplyController to recreate it
kubectl get configmap coredns -n kube-system -w
```
Or as a stopgap, apply the correct content yourself:
```bash
kubectl apply --server-side -f <(terraform output coredns_config)
```
## Quick Start — Deploying the Cluster
### 1. Bootstrap Talos Nodes
View File
+2 -2
View File
@@ -8,8 +8,8 @@ data:
"theme": "light",
"compaction": {
"enabled": true,
"reserveTokens": 16000,
"keepRecentTokens": 6000
"reserveTokens": 8192,
"keepRecentTokens": 12000
}
}
kind: ConfigMap
+2 -15
View File
@@ -1,27 +1,14 @@
# Exposes agent-hub at api.riotpiao.com/console (WebSocket) and /run
# (trigger a new session) -- both are routes on the same hub.js service.
#
# Was ingressClassName: kong until Kong was retired on 2026-08-19. Pointed
# straight at nginx rather than through the replacement Go gateway because that
# gateway has no WebSocket upgrade support yet -- routing /console through it
# would break the console outright. nginx handles the upgrade natively.
#
# Path precedence: the nginx Ingress api/api catch-alls `/` on this same host
# to the gateway. nginx matches longest prefix first, so these three paths win
# over `/` and the rest of the host still reaches the gateway.
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: console
namespace: agent-pod
annotations:
# A console WebSocket stays open across a whole agent session; nginx's 60s
# default read timeout would drop it mid-run.
nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"
nginx.ingress.kubernetes.io/proxy-send-timeout: "3600"
nginx.ingress.kubernetes.io/proxy-buffering: "off"
konghq.com/strip-path: "false"
spec:
ingressClassName: nginx
ingressClassName: kong
rules:
- host: api.riotpiao.com
http:
@@ -1,982 +0,0 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: coordinator-src
namespace: agent-pod
data:
coordinator.js: |
#!/usr/bin/env node
// coordinator: local CLI that drives a multi-phase, multi-task pipeline of
// planner/investigator/implementer/judge stages, same state machine as
// hub.js's old /pipeline handler. Every interactive stage runs through the
// patched agent-manager fork's `spawn` subcommand, so it is a real tracked
// session (tmux pane on agent-manager's private socket + a state.db row)
// from the moment it exists -- attachable and visible in agent-manager's
// own TUI the whole time it runs.
//
// Per repo, each role (planner/investigator/implementer/judge) is ONE
// persistent agent-manager session, not a fresh spawn per task: the first
// task to need a role spawns it, every later task for that role reuses the
// same tmux pane via `tmux send-keys` (see runOnPool) -- the same nudge
// mechanism that used to only fire on a stall now doubles as "give this
// agent its next task." Every role reads everything it needs fresh off disk
// each call, so every pane gets a `/new` before every reuse instead of
// accumulating history that degrades and eventually errors out task after
// task -- same pane, same agent-manager session, zero memory of the last
// task it handled.
// Because only one implementer/judge/etc. exists per repo, tasks within a
// phase run strictly sequentially against the pool -- no per-task worktree,
// no per-task branch, no merge-back step; every task commits straight onto
// the phase branch in the repo's one shared clone.
//
// The unit of concurrency is now the REPO, not the task: runCoordinator
// takes a list of repos and runs up to REPO_CONCURRENCY of them at once,
// each with its own clone (under WORK_DIR/<repoId>) and its own 4-agent
// pool. The coordinator never kills a role's session; it rests at an idle
// prompt between tasks, and agent-manager's session list becomes the audit
// trail of everything every repo's pipeline ran. Completion is signaled by
// sentinel files under the repo's clone (unchanged convention), waited on
// with fs.watch instead of polling.
const fs = require("node:fs");
const path = require("node:path");
const { spawn } = require("node:child_process");
const WORK_DIR = process.env.HUB_WORK_DIR || path.join(require("node:os").tmpdir(), "agent-harness-work");
// Never rely on a bare `pi`/`agent-manager` on $PATH -- see PI_BIN's own
// comment below; the same collision risk applies to any CLI name. Always
// invoke explicit pinned paths.
const PI_BIN =
process.env.PI_BIN ||
path.join(__dirname, "..", ".pi-cli", "node_modules", "@earendil-works", "pi-coding-agent", "dist", "cli.js");
const AGENT_MANAGER_BIN = process.env.AGENT_MANAGER_BIN || path.join(__dirname, "..", ".bin", "agent-manager-fork");
// agent-manager's own session-state DB -- used to detect a session that has
// actually died (process crashed/exited, status flips to "errored"/"dead")
// instead of one that's merely slow. Read-only introspection plus the one
// UPDATE in killDeadSession below, same class of operation as the tmux
// nudges already done directly against agent-manager's internals.
const AGENT_MANAGER_DB =
process.env.AGENT_MANAGER_DB || path.join(require("node:os").homedir(), ".config", "agent-manager", "state.db");
// Empty means "let pi fall back to ~/.pi/agent/settings.json's default"
// (currently anthropic/claude-sonnet-4-5, real paid usage). Set both to
// route every stage -- headless (spawnPi) and interactive (runOnPool) --
// at the homelab model instead, e.g. AGENT_PROVIDER=homelab-ornith
// AGENT_MODEL=ornith:35b.
const AGENT_PROVIDER = process.env.AGENT_PROVIDER || "";
const AGENT_MODEL = process.env.AGENT_MODEL || "";
// judge can run a different model than the rest of the chain, e.g.
// homelab-reasoning instead of homelab-ornith now that verifier/PRM is
// retired. Falls back to AGENT_PROVIDER/AGENT_MODEL when unset, so a run
// that doesn't care keeps one uniform model everywhere.
const JUDGE_PROVIDER = process.env.JUDGE_PROVIDER || AGENT_PROVIDER;
const JUDGE_MODEL = process.env.JUDGE_MODEL || AGENT_MODEL;
function providerModelFor(role) {
return role === "judge" ? { provider: JUDGE_PROVIDER, model: JUDGE_MODEL } : { provider: AGENT_PROVIDER, model: AGENT_MODEL };
}
// agent-manager's private tmux server and session-naming scheme
// (internal/tmux/tmux.go: defaultSocket = "agentmgr", sessionName(id) =
// "am_"+id) -- stable, documented internals of the fork, used here only
// for read-only introspection (pane capture) and role nudges, exactly the
// class of operation hub.js already ran directly against its own sessions
// rather than asking a model to do it.
const AM_SOCKET = "agentmgr";
function amSessionName(id) {
return `am_${id}`;
}
function runAmTmux(args) {
return runCmd("tmux", ["-L", AM_SOCKET, ...args]);
}
const ROLE_SKILLS = new Set(["planner", "investigator", "info-collector", "implementer", "judge", "resolver"]);
// Every role reads everything it needs fresh off disk each call -- PLAN.md,
// the task spec, judge's verdict file, `git diff` against baseBranch --
// nothing depends on remembering earlier tasks. Left to accumulate, a
// pooled session's conversation grows without bound across every task in a
// repo and both correctness and reliability degrade hard once it does
// (observed: a planner session at ~1.5M cumulative tokens started erroring
// out every call, an investigator session that far gone started narrating a
// different codebase entirely). So every role gets reset to a clean
// conversation before every reuse instead of just being nudged with the
// next prompt -- same pane, same agent-manager session (still
// visible/attachable), zero history carried between tasks.
function sleep(ms) {
return new Promise((resolve) => setTimeout(resolve, ms));
}
const HARD_RULES =
"Read and follow ~/.pi/agent/skills/karpathy-guidelines/SKILL.md and " +
"~/.pi/agent/skills/caveman/SKILL.md as hard rules for this entire task, before anything else. ";
// judge (routed to homelab-reasoning) has been observed narrating an
// entire review in prose -- "I should run git diff, then check X..." --
// and then writing a verdict based on that narration without ever calling
// a real tool. Live example: a phase-judge call produced a page of
// "I would check..." reasoning, declared VERDICT: PASS, and showed the
// touch command as a fenced code block IN ITS OWN TEXT rather than
// executing it. Coordinator just timed out waiting on a sentinel that was
// never going to appear, since nothing was ever actually run. Spelled out
// explicitly since "use the judge skill" alone apparently isn't enough to
// rule this out.
const REQUIRE_REAL_TOOL_CALLS =
"Do not narrate what you would check -- actually run the commands via a real tool call and read their real " +
"output before writing anything. A verdict based on describing checks instead of executing them is invalid. " +
"Writing the verdict file and touching the sentinel are themselves tool calls you must execute, not text to " +
"display in your response. ";
function parseVerdictLine(text, label) {
if (!text) return null;
const re = new RegExp(`${label}:\\s*(\\w+)`, "i");
const m = text.match(re);
return m ? m[1].toUpperCase() : null;
}
function runCmd(bin, args, cwd) {
return new Promise((resolve) => {
const child = spawn(bin, args, { cwd, stdio: ["ignore", "pipe", "pipe"] });
let out = "";
child.stdout.on("data", (c) => (out += c));
child.stderr.on("data", (c) => (out += c));
child.on("close", (code) => resolve({ code, out: out.trim() }));
});
}
function runGit(cwd, args) {
return runCmd("git", args, cwd);
}
// A stage saying "commit" in its prompt is a request, not a guarantee -- seen
// in practice: a stage writes a real file and simply never runs `git add`/
// `git commit`, leaving it untracked and invisible to every later `git diff`.
// Sweep and commit anything left dirty after every stage, deterministically.
async function commitPending(cwd, message) {
await runGit(cwd, ["add", "-A"]);
const status = await runGit(cwd, ["status", "--porcelain"]);
if (!status.out) return { committed: false };
const commit = await runGit(cwd, ["commit", "-m", message]);
return { committed: commit.code === 0, error: commit.code !== 0 ? commit.out : undefined };
}
// Headless one-shot pi call (`pi -p --mode json <prompt>`), used only for
// quick diagnostic/mechanical calls that don't need to be a watchable
// session: resolver's crash/stall diagnosis, and the initial clone. Kept
// exactly as before -- only the interactive per-role stages (runOnPool,
// below) go through agent-manager.
//
// Bounded by SPAWN_PI_TIMEOUT_MS -- unlike runOnPool's pooled sessions
// (which now have status polling to catch a dead session fast, see
// waitForSentinel/killDeadSession), this is a raw child_process with no
// equivalent escape hatch. Observed live: a resolver call shared the
// default backend with a concurrently-busy repo's implementer and sat for
// 6+ minutes producing nothing -- with no timeout here, that blocks the
// entire calling repo's pipeline forever, since askResolver is always
// awaited before the next stage can run.
const SPAWN_PI_TIMEOUT_MS = 5 * 60 * 1000;
function spawnPi({ agent, prompt, cwd }) {
const finalPrompt = ROLE_SKILLS.has(agent) ? `/skill:${agent} ${HARD_RULES}${prompt}` : prompt;
const args = ["-p", "--mode", "json"];
if (AGENT_PROVIDER) args.push("--provider", AGENT_PROVIDER);
if (AGENT_MODEL) args.push("--model", AGENT_MODEL);
args.push(finalPrompt);
const child = spawn(PI_BIN, args, { stdio: ["ignore", "pipe", "pipe"], cwd });
let lastText = "";
let stderrTail = "";
let buf = "";
child.stdout.on("data", (chunk) => {
buf += chunk;
let idx;
while ((idx = buf.indexOf("\n")) !== -1) {
const line = buf.slice(0, idx);
buf = buf.slice(idx + 1);
if (!line.trim()) continue;
try {
const event = JSON.parse(line);
if (event.type === "message_end" && event.message && Array.isArray(event.message.content)) {
const text = event.message.content
.filter((c) => c.type === "text")
.map((c) => c.text)
.join("\n");
if (text) lastText = text;
}
} catch {
// non-JSON stdout noise, ignore
}
}
});
child.stderr.on("data", (chunk) => {
process.stderr.write(chunk);
stderrTail = (stderrTail + chunk.toString()).slice(-4000);
});
return new Promise((resolve) => {
let settled = false;
const timer = setTimeout(() => {
if (settled) return;
settled = true;
child.kill("SIGKILL");
resolve({ code: null, lastText, stderrTail, timedOut: true });
}, SPAWN_PI_TIMEOUT_MS);
child.on("close", (code) => {
if (settled) return;
settled = true;
clearTimeout(timer);
resolve({ code, lastText, stderrTail });
});
});
}
async function askResolver(cwd, repoId, diagnosticPrompt) {
const result = await spawnPi({ agent: "resolver", prompt: diagnosticPrompt, cwd });
return parseVerdictLine(result.lastText, "RESOLUTION");
}
const STAGE_TIMEOUT_MS = 10 * 60 * 1000;
const NUDGE_TIMEOUT_MS = 5 * 60 * 1000;
// Resolves as soon as filePath appears (fs.watch on its directory, same as
// before), as soon as target's agent-manager status flips to "errored" or
// "dead" (polled -- state.db has no watch mechanism), or after limitMs with
// neither. A session that has actually crashed will never touch the
// sentinel, so without the status poll this just burns the full STAGE_
// TIMEOUT_MS waiting on a file that was never coming, same as a genuine
// stall -- polling status catches that in ~pollMs instead.
function waitForSentinel(filePath, target, limitMs, pollMs = 5000) {
return new Promise((resolve) => {
if (fs.existsSync(filePath)) return resolve({ ok: true });
const dir = path.dirname(filePath);
const id = target.replace(/^am_/, "");
let settled = false;
let watcher;
let poller;
let timer;
const finish = (result) => {
if (settled) return;
settled = true;
clearTimeout(timer);
clearInterval(poller);
if (watcher) {
try {
watcher.close();
} catch {
// already closed
}
}
resolve(result);
};
try {
watcher = fs.watch(dir, () => {
if (fs.existsSync(filePath)) finish({ ok: true });
});
} catch {
// dir missing at watch time is a real bug elsewhere (cwd should
// already exist); surface it as a timeout rather than hang forever.
return finish({ timedOut: true });
}
// Closes the race between the existsSync check above and the watcher
// actually being attached.
if (fs.existsSync(filePath)) return finish({ ok: true });
poller = setInterval(async () => {
const { out } = await runCmd("sqlite3", [AGENT_MANAGER_DB, `SELECT status FROM sessions WHERE id='${id}'`]);
const status = out.trim();
if (status === "errored" || status === "dead") finish({ dead: true, status });
}, pollMs);
timer = setTimeout(() => finish({ timedOut: true }), limitMs);
});
}
// Kills a session that's actually crashed (not just slow) and archives it
// in agent-manager's own DB so it stops showing up as a live, unattended
// pane -- otherwise every crash leaves an orphaned tmux session + state.db
// row behind permanently, identical to the manually-cleaned-up poiman-
// planner ghost session found earlier this same run.
async function killDeadSession(target) {
await runAmTmux(["kill-session", "-t", target]);
const id = target.replace(/^am_/, "");
await runCmd("sqlite3", [AGENT_MANAGER_DB, `UPDATE sessions SET archived=1 WHERE id='${id}'`]);
}
// Runs one task's worth of work on a persistent per-role agent: spawns the
// role's session the first time it's ever needed for this repo, sends every
// later prompt into that same tmux pane via send-keys -- prefixed with a
// `/new` first, so the pane and agent-manager session stay the same but the
// model starts that prompt with a clean conversation, no history carried
// over from whatever task this role last handled. pool is a plain object
// keyed by role name ("planner"/"investigator"/"implementer"/"judge"),
// shared across every task in a repo's pipeline (see runRepoPipeline) -- it
// IS the 4-agent pool, one entry per role, filled in lazily as each role
// gets its first task.
// A dead/errored session gets one respawn-and-retry (same prompt, fresh
// session) before this stage is abandoned -- matches resolver-SKILL.md's
// own documented contract of retrying a failed stage at most once.
const DEAD_SESSION_RETRIES = 1;
async function runOnPool(pool, cwd, repoId, role, prompt, sentinelFile) {
fs.rmSync(sentinelFile, { force: true });
const label = `${repoId}-${role}`;
// A pooled session's shell cwd drifts as it explores the repo (e.g. cd
// into a Rust workspace subdirectory to read source) and nothing resets
// it back between turns. Seen in practice: a repo whose own internal
// workspace folder is one letter off from the repo's own directory name
// ("poiman" the repo vs. "poimen" the crate workspace inside it) was
// enough for the agent to touch its sentinel one level off from where
// this function is watching for it -- coordinator waits out the full
// STAGE_TIMEOUT_MS for a file that already exists, just in the wrong
// place. State the absolute target directory and use absolute paths for
// every filesystem instruction, so there's nothing for the agent to get
// wrong by reasoning about a relative "current directory."
const cwdReminder = `Your working directory for this task is ${cwd} -- if your shell isn't already there, run: cd ${cwd}\n\n`;
const spawnFresh = async () => {
const spawnArgs = ["spawn", "--tool", "pi", "--cwd", cwd, "--name", label, "--group", repoId, "--prompt", cwdReminder + HARD_RULES + prompt];
const { provider, model } = providerModelFor(role);
if (provider) spawnArgs.push("--provider", provider);
if (model) spawnArgs.push("--model", model);
const spawned = await runCmd(AGENT_MANAGER_BIN, spawnArgs);
return spawned.code === 0 ? amSessionName(spawned.out) : null;
};
let target = pool[role];
if (!target) {
target = await spawnFresh();
if (!target) return { ok: false, crashed: true, error: "spawn failed", sessionName: label };
pool[role] = target;
} else {
await runAmTmux(["send-keys", "-t", target, "/new", "Enter"]);
await sleep(1000);
await runAmTmux(["send-keys", "-t", target, cwdReminder + HARD_RULES + prompt, "Enter"]);
}
for (let deadRetries = 0; ; deadRetries++) {
const outcome = await waitForSentinel(sentinelFile, target, STAGE_TIMEOUT_MS);
if (outcome.ok) return { ok: true, sessionName: label };
if (outcome.dead) {
await killDeadSession(target);
if (pool[role] === target) delete pool[role];
if (deadRetries >= DEAD_SESSION_RETRIES) {
return { ok: false, crashed: true, error: `session died (status: ${outcome.status})`, sessionName: label };
}
target = await spawnFresh();
if (!target) return { ok: false, crashed: true, error: "respawn after death failed", sessionName: label };
pool[role] = target;
continue;
}
// Plain stall -- session still alive, just slow. Ask resolver once,
// nudge if it says worth it, and stop here either way (this is not
// the death path, so no respawn/retry loop).
const pane = await runAmTmux(["capture-pane", "-t", target, "-p", "-S", "-200"]);
const resolution = await askResolver(
cwd,
repoId,
`Repo ${repoId}'s "${role}" agent hasn't finished its current task after 10 minutes. Its pane tail:\n${pane.out.slice(-3000)}\n\n` +
`Decide: is it still making real progress and worth nudging to wrap up, or stuck and worth abandoning?`
);
let ok = false;
if (resolution === "RETRY") {
await runAmTmux(["send-keys", "-t", target, `Please wrap up now and run: touch ${sentinelFile}`, "Enter"]);
const nudged = await waitForSentinel(sentinelFile, target, NUDGE_TIMEOUT_MS);
ok = nudged.ok === true;
if (nudged.dead) {
await killDeadSession(target);
if (pool[role] === target) delete pool[role];
}
}
return { ok, sessionName: label };
}
}
function plannerPrompt(task, specHint, judgeOnly, cwd) {
// judgeOnly (auto-discovered tasks only, see parseTaskBoard): planner
// itself decides whether the task is already done before planning it,
// reading tasks/INDEX.md's own status notes plus git log/current code --
// replaces what used to be a separate judge pre-check call. One LLM round
// trip instead of two, and the same agent that's about to plan the task
// is the one deciding whether planning it is even necessary.
const resultFile = path.join(cwd, `.task-result-${task}`);
const decideStep = judgeOnly
? `First, decide whether task ${task} is already fully implemented on this branch: check ` +
`\`git log --oneline --grep '${task}'\`, tasks/INDEX.md's own status notes for this task, and the current ` +
`code directly against its spec (${specHint})'s acceptance criteria. Write your decision to ` +
`${resultFile} as a single "VERDICT: PASS" (already done, no further work needed) or ` +
`"VERDICT: FAIL" (needs work) line plus one line of rationale. If VERDICT is FAIL, continue below and ` +
`draft the plan in this same turn; if VERDICT is PASS, skip the rest and go straight to the touch step.\n\n`
: "";
return (
`${decideStep}Use the planner skill to draft PLAN.md for task ${task}, reading its spec (${specHint}). ` +
`PLAN.md is scratch state for this harness, not a deliverable -- do NOT commit it or add it to git. ` +
`Then run: touch ${path.join(cwd, `.stage-done-${task}-planner`)}`
);
}
function investigatorPrompt(task, cwd) {
return (
`Use the investigator skill to confirm PLAN.md against real sources for task ${task}, append findings. ` +
`PLAN.md is scratch state for this harness, not a deliverable -- do NOT commit it or add it to git. ` +
`Then run: touch ${path.join(cwd, `.stage-done-${task}-investigator`)}`
);
}
function implementerPrompt(task, attempt, feedbackHint, cwd) {
return (
`Use the implementer skill to implement what the current PLAN.md specifies for task ${task} (commit as you go). ` +
`${feedbackHint} Then run: touch ${path.join(cwd, `.stage-done-${task}-implementer-${attempt}`)}`
);
}
function judgePrompt(task, baseBranch, attempt, cwd) {
return (
`${REQUIRE_REAL_TOOL_CALLS}Use the judge skill to review the diff against ${baseBranch}...HEAD for task ${task}. ` +
`Write your verdict to ${path.join(cwd, `.task-result-${task}`)} as a single "VERDICT: PASS" or "VERDICT: FAIL" ` +
`line plus one line of rationale, then run: touch ${path.join(cwd, `.stage-done-${task}-judge-${attempt}`)}`
);
}
const MAX_IMPLEMENT_ATTEMPTS = 5;
const MAX_PLAN_REVISIONS = 3;
// Runs one task against the repo's shared role pool: planner drafts
// PLAN.md (for auto-discovered tasks, first deciding off tasks/INDEX.md and
// the repo's own state whether the task is already done -- see
// plannerPrompt's judgeOnly branch; judge never does this pre-check),
// investigator confirms it, then implementer and judge go back and
// forth -- judge's FAIL rationale lands in .task-result-<task>, which the
// next implementer attempt is told to read and address. After
// MAX_IMPLEMENT_ATTEMPTS straight fails, the planner role is asked to judge
// whether the plan itself is wrong -- fresh conversation, same as any other
// planner call, reading PLAN.md/the judge feedback/the
// diff off disk rather than remembering having drafted the original plan.
// If it decides the approach is wrong it revises PLAN.md and the implementer
// gets a fresh attempt budget.
// MAX_PLAN_REVISIONS caps this from looping forever on a task that's
// genuinely stuck. All work happens directly in cwd (the repo's one shared
// clone, currently checked out to the phase branch) -- no worktree, since
// only one implementer/judge exist per repo and tasks run strictly one at a
// time (see runPhase).
async function runTaskOnPool(cwd, baseBranch, task, pool, repoId, pipelineSession, judgeOnly) {
const resultFile = path.join(cwd, `.task-result-${task}`);
fs.rmSync(resultFile, { force: true });
const specHint = `the file under tasks/ starting with "${task}-"`;
const stage = async (role, prompt, sentinel, displayLabel) => {
const label = displayLabel || role;
pipelineSession.activeTasks[task] = { stage: label, startedAt: new Date().toISOString() };
logProgress(pipelineSession);
const result = await runOnPool(pool, cwd, repoId, role, prompt, sentinel);
await commitPending(cwd, `task: ${task} (${label})`);
return result;
};
const abandon = (stageLabel, result, attempt) => {
delete pipelineSession.activeTasks[task];
logProgress(pipelineSession);
return {
task,
status: result.crashed ? "spawn-crashed" : "timed-out",
error: result.error,
stoppedAt: stageLabel,
...(attempt !== undefined ? { attempt } : {}),
};
};
// PLAN.md is scratch state for this one task, not a deliverable (see
// plannerPrompt/investigatorPrompt -- it's gitignored too, as a backstop
// in case an agent commits it anyway). Discard it once the task is done,
// whatever the outcome, so it never bleeds into the next task's planner
// call or sits around as stale harness clutter in the shared clone.
try {
let result = await stage("planner", plannerPrompt(task, specHint, judgeOnly, cwd), path.join(cwd, `.stage-done-${task}-planner`));
if (!result.ok) return abandon("planner", result);
if (judgeOnly) {
const quickText = fs.existsSync(resultFile) ? fs.readFileSync(resultFile, "utf8") : "";
if (parseVerdictLine(quickText, "VERDICT") === "PASS") {
delete pipelineSession.activeTasks[task];
logProgress(pipelineSession);
return { task, status: "done", judgeRationale: quickText, judgeOnlyPass: true };
}
}
result = await stage("investigator", investigatorPrompt(task, cwd), path.join(cwd, `.stage-done-${task}-investigator`));
if (!result.ok) return abandon("investigator", result);
let planRevisions = 0;
let implementAttempt = 0;
let verdict = null;
let resultText = "";
let justRevisedPlan = false;
while (true) {
implementAttempt++;
const feedbackHint = fs.existsSync(resultFile)
? justRevisedPlan
? `${resultFile} holds the judge's feedback against the OLD plan, which prompted a plan revision -- ` +
`PLAN.md has since changed. Read the current PLAN.md as the source of truth, not the old feedback verbatim.`
: `A previous judge review exists at ${resultFile} -- read it and address every issue it raises.`
: "";
justRevisedPlan = false;
result = await stage(
"implementer",
implementerPrompt(task, implementAttempt, feedbackHint, cwd),
path.join(cwd, `.stage-done-${task}-implementer-${implementAttempt}`)
);
if (!result.ok) return abandon("implementer", result, implementAttempt);
result = await stage("judge", judgePrompt(task, baseBranch, implementAttempt, cwd), path.join(cwd, `.stage-done-${task}-judge-${implementAttempt}`));
if (!result.ok) return abandon("judge", result, implementAttempt);
resultText = fs.existsSync(resultFile) ? fs.readFileSync(resultFile, "utf8") : "";
verdict = parseVerdictLine(resultText, "VERDICT");
if (verdict === "PASS") break;
if (implementAttempt >= MAX_IMPLEMENT_ATTEMPTS) {
if (planRevisions >= MAX_PLAN_REVISIONS) break;
planRevisions++;
result = await stage(
"planner",
`Implementer failed judge review ${MAX_IMPLEMENT_ATTEMPTS} times in a row for task ${task}. Read PLAN.md, ` +
`the judge's feedback in ${resultFile}, and the current diff against ${baseBranch}...HEAD. Decide ` +
`whether the plan's approach itself is wrong, not just the implementation -- if so, revise PLAN.md. If ` +
`you change the approach, also use the investigator skill to confirm the new approach against real ` +
`sources. If the plan is sound, note why in PLAN.md and leave it as-is. PLAN.md is scratch state for ` +
`this harness, not a deliverable -- do NOT commit it or add it to git. Then run: ` +
`touch ${path.join(cwd, `.stage-done-${task}-planner-revise-${planRevisions}`)}`,
path.join(cwd, `.stage-done-${task}-planner-revise-${planRevisions}`),
"planner-revise"
);
if (!result.ok) return abandon("planner-revise", result, planRevisions);
implementAttempt = 0;
justRevisedPlan = true;
}
}
delete pipelineSession.activeTasks[task];
logProgress(pipelineSession);
if (verdict !== "PASS" && planRevisions >= MAX_PLAN_REVISIONS) {
return { task, status: "unresolved", judgeRationale: resultText, implementAttempts: implementAttempt, planRevisions };
}
return { task, status: verdict === "PASS" ? "done" : "done-with-concerns", judgeRationale: resultText };
} finally {
fs.rmSync(path.join(cwd, "PLAN.md"), { force: true });
}
}
// Committed (never gitignored) so it survives a resumed phase branch --
// one task id per line, appended as each task resolves. This is what lets
// a resumed run skip straight past already-resolved tasks instead of
// re-running planner's judgeOnly decision on every one of them again:
// resuming the git branch alone only recovers the CODE, not "which tasks
// are already settled," and re-deciding that from scratch for every task
// burns a full LLM call per already-done task before ever reaching the
// first one that actually needs work.
function progressLedgerPath(cwd) {
return path.join(cwd, ".agent-progress");
}
function readCompletedTasks(cwd) {
const file = progressLedgerPath(cwd);
if (!fs.existsSync(file)) return new Set();
return new Set(
fs
.readFileSync(file, "utf8")
.split("\n")
.map((line) => line.trim())
.filter(Boolean)
);
}
async function recordTaskComplete(cwd, task) {
fs.appendFileSync(progressLedgerPath(cwd), `${task}\n`);
await runGit(cwd, ["add", path.basename(progressLedgerPath(cwd))]);
await runGit(cwd, ["commit", "-m", `chore: mark ${task} complete in progress ledger`]);
}
// Runs every task in a phase (no declared dependency between them) strictly
// one at a time against the repo's shared role pool -- only one implementer/
// judge/etc. exists per repo, so there is no per-task concurrency to have
// here anymore (see REPO_CONCURRENCY below for where concurrency now
// lives). No worktrees: every task commits directly onto phaseBranch in the
// one shared cwd. baseBranch here is the TRUE base (e.g. "main") -- judge
// reviews `git diff baseBranch...HEAD`, not phaseBranch...HEAD, which would
// always be empty since HEAD *is* phaseBranch while it's checked out.
//
// Pushes phaseBranch after every task, not just once at full-phase-end: the
// pod is ephemeral and every restart re-clones baseBranch fresh (see
// runRepoPipeline) -- without this, a redeploy mid-phase silently discards
// every task committed so far, and the next run re-decides "is this done?"
// from a clone that never saw any of that work.
async function runPhase(cwd, baseBranch, phaseBranch, phaseTasks, pool, repoId, pipelineSession) {
const entries = phaseTasks.map((t) => (typeof t === "string" ? { id: t, judgeOnly: false } : t));
const completed = readCompletedTasks(cwd);
for (const entry of entries) {
if (completed.has(entry.id)) {
const result = { task: entry.id, status: "done", resumed: true };
pipelineSession.taskResults.push(result);
logProgress(pipelineSession);
continue;
}
const result = await runTaskOnPool(cwd, baseBranch, entry.id, pool, repoId, pipelineSession, entry.judgeOnly);
pipelineSession.taskResults.push(result);
if (result.status === "done" || result.status === "done-with-concerns") {
await recordTaskComplete(cwd, entry.id);
}
await runGit(cwd, ["push", "-u", "origin", phaseBranch]);
logProgress(pipelineSession);
}
}
// Discovers phases/tasks from the repo's own tasks/INDEX.md instead of
// requiring the caller to pass --tasks. Matches this convention's board
// shape (see e.g. Poimen/agent-rust's tasks/INDEX.md): a numbered phase
// heading ("## 1 — Foundations · T0.x"), followed by a markdown table
// whose rows link to each task's own spec file ("| [T0.1](T0.1-....md) |
// ... |"). Headings that aren't a numbered phase (prose sections like
// "## Ordering — declared, never derived", "## Progress") are skipped --
// only "## <digits> — ..." starts a new phase. Returns null if
// tasks/INDEX.md doesn't exist; an empty array if it exists but no phase
// yielded any task rows.
function parseTaskBoard(cwd) {
const indexPath = path.join(cwd, "tasks", "INDEX.md");
if (!fs.existsSync(indexPath)) return null;
const phaseHeaderRe = /^##\s+\d+\s+—/;
const taskRowRe = /^\|\s*\[([A-Za-z0-9.]+)\]\(/;
const phases = [];
let current = null;
for (const line of fs.readFileSync(indexPath, "utf8").split("\n")) {
if (phaseHeaderRe.test(line)) {
current = [];
phases.push(current);
continue;
}
const m = line.match(taskRowRe);
if (m && current) current.push(m[1]);
}
return phases.filter((phase) => phase.length > 0);
}
function phaseLabelFor(phaseTasks, index) {
const first = phaseTasks[0];
const id = typeof first === "string" ? first : first.id;
const dot = id.indexOf(".");
return dot === -1 ? `phase-${index}` : id.slice(0, dot);
}
function logProgress(pipelineSession) {
console.log(`[repo ${pipelineSession.id}] ${JSON.stringify(pipelineSession)}`);
}
// Runs one repo's full pipeline: clone, then phases strictly sequentially.
// tasks: array of phases, each phase an array of task ids with no declared
// dependency on each other (e.g. [["T0.1","T0.2"], ["T1.1","T1.2","T1.3"]]).
// A flat array of ids is also accepted and treated as one single phase. If
// omitted, phases are discovered from the repo's own tasks/INDEX.md and run
// judgeOnly first (a cheap "is this already done" check against the
// board's possibly-stale checkmarks). Each phase gets its own branch
// (agent-run/<repoId>/<phaseLabel>, e.g. .../T1); once every task in that
// phase lands "done" or "done-with-concerns" AND the phase judge (the same
// pooled judge agent that reviewed each task) passes the integration
// review, the phase branch is squash-merged into baseBranch and pushed,
// then the next phase branches off that updated base. Any failure halts
// this repo's pipeline before merging -- it does not affect other repos
// running concurrently (see runCoordinator).
async function runRepoPipeline({ repoId, repo, baseBranch, tasks, branchName }, pipelineSession) {
const cwd = path.join(WORK_DIR, repoId);
// repoId is a slug derived from the repo URL now (see slugFor), not a
// fresh UUID -- reusable across separate `runCoordinator` invocations
// against the same repo, so a stale clone from a prior run has to be
// wiped before this one starts, not merged into.
fs.rmSync(cwd, { recursive: true, force: true });
fs.mkdirSync(cwd, { recursive: true });
const pool = {};
const finish = (status) => {
pipelineSession.status = status;
pipelineSession.endedAt = new Date().toISOString();
logProgress(pipelineSession);
return pipelineSession;
};
// Deterministic, not routed through an LLM -- clone is 100% mechanical
// (same reasoning as commitPending/the squash-merge sequence below), and
// was the one place left that broke that pattern: a headless spawnPi
// call here meant a crash gave zero diagnostic output, just a silent
// exit code with nothing to debug from.
const clone = await runGit(cwd, ["clone", "--branch", baseBranch, repo, "."]);
if (clone.code !== 0) {
pipelineSession.gitError = clone.out;
return finish("clone-crashed");
}
if (!fs.existsSync(path.join(cwd, ".git"))) return finish("clone-missing");
let phases = tasks ? (Array.isArray(tasks[0]) ? tasks : [tasks]) : parseTaskBoard(cwd);
if (!phases || phases.length === 0) {
pipelineSession.gitError = "no tasks given and tasks/INDEX.md not found or empty";
return finish("no-tasks-found");
}
if (!tasks) {
phases = phases.map((phase) => phase.map((id) => ({ id, judgeOnly: true })));
}
pipelineSession.totalTasks = phases.flat().length;
for (let i = 0; i < phases.length; i++) {
const phaseTasks = phases[i];
const phaseLabel = phaseLabelFor(phaseTasks, i);
const phaseBranch = branchName ? `${branchName}/${phaseLabel}` : `agent-run/${repoId}/${phaseLabel}`;
// Resume a phase branch a prior (since-restarted) run already pushed,
// instead of always branching fresh off baseBranch -- otherwise every
// redeploy silently discards whatever tasks that prior run already
// committed and pushed (see runPhase's per-task push below).
const fetchExisting = await runGit(cwd, ["fetch", "origin", phaseBranch]);
const resuming = fetchExisting.code === 0;
const branchResult = resuming
? await runGit(cwd, ["checkout", "-b", phaseBranch, "FETCH_HEAD"])
: await runGit(cwd, ["checkout", "-b", phaseBranch]);
if (branchResult.code !== 0) {
pipelineSession.gitError = branchResult.out;
return finish("branch-crashed");
}
// Idempotent and run every phase, NOT gated on a fresh (non-resumed)
// start -- every run this session was a resume, so the old i===0 &&
// !resuming gate meant this setup permanently never ran on poiman's
// branch, and portfolio's PLAN.md stayed tracked from before this rule
// ever existed (gitignore has no effect on an already-tracked file --
// observed live: it kept getting swept back in by every `git add -A`
// regardless of the ignore rule). Check-and-fix on every phase instead
// of once-at-genesis so a repo that's missing either self-heals on its
// very next run rather than carrying the gap forever.
const gitignorePath = path.join(cwd, ".gitignore");
const currentGitignore = fs.existsSync(gitignorePath) ? fs.readFileSync(gitignorePath, "utf8").split("\n") : [];
const requiredGitignoreLines = [
"*.tar.gz",
"*.tgz",
"*.crate",
"*.zip",
"*.bin",
"*.whl",
"vendor/",
"node_modules/",
".task-result-*",
".phase-result-*",
".stage-done-*",
"PLAN.md",
];
const missingGitignoreLines = requiredGitignoreLines.filter((line) => !currentGitignore.includes(line));
if (missingGitignoreLines.length > 0) {
fs.appendFileSync(
gitignorePath,
"\n# agent-harness: build artifacts, vendored archives, and harness bookkeeping never belong in source control\n" +
missingGitignoreLines.join("\n") +
"\n"
);
await runGit(cwd, ["add", ".gitignore"]);
await runGit(cwd, ["commit", "-m", "chore: broaden .gitignore for agent-run artifacts"]);
}
const trackedFiles = await runGit(cwd, ["ls-tree", "-r", "HEAD", "--name-only"]);
if (trackedFiles.out.split("\n").includes("PLAN.md")) {
await runGit(cwd, ["rm", "--cached", "PLAN.md"]);
await runGit(cwd, ["commit", "-m", "chore: untrack PLAN.md (already gitignored, was committed pre-rule)"]);
}
await runPhase(cwd, baseBranch, phaseBranch, phaseTasks, pool, repoId, pipelineSession);
const phaseTaskIds = new Set(phaseTasks.map((t) => (typeof t === "string" ? t : t.id)));
const phaseResults = pipelineSession.taskResults.filter((r) => phaseTaskIds.has(r.task));
const phaseClean =
phaseResults.length === phaseTaskIds.size && phaseResults.every((r) => r.status === "done" || r.status === "done-with-concerns");
if (!phaseClean) {
pipelineSession.haltedAt = phaseLabel;
return finish("halted-phase-failed");
}
const phaseResultFile = path.join(cwd, `.phase-result-${phaseLabel}`);
fs.rmSync(phaseResultFile, { force: true });
const phaseJudge = await runOnPool(
pool,
cwd,
repoId,
"judge",
`${REQUIRE_REAL_TOOL_CALLS}Use the judge skill to review the full phase diff for phase ${phaseLabel} against ` +
`${baseBranch}...HEAD (covers every task in this phase: ${[...phaseTaskIds].join(", ")}). Every ` +
`individual task already passed its own judge review -- your job here is different: confirm the ` +
`tasks integrate correctly as one coherent narrative, and that real integration tests (not just ` +
`each task's isolated unit checks) exist and actually exercise the phase's intended use case end ` +
`to end. Write your verdict to ${phaseResultFile} as a single "VERDICT: PASS" or ` +
`"VERDICT: FAIL" line plus rationale, then run: touch ${path.join(cwd, `.stage-done-phase-${phaseLabel}-judge`)}`,
path.join(cwd, `.stage-done-phase-${phaseLabel}-judge`)
);
await commitPending(cwd, `phase: ${phaseLabel} integration review`);
if (!phaseJudge.ok) {
pipelineSession.haltedAt = phaseLabel;
pipelineSession.gitError = phaseJudge.error;
return finish("phase-judge-crashed");
}
const phaseJudgeText = fs.existsSync(phaseResultFile) ? fs.readFileSync(phaseResultFile, "utf8") : "";
if (parseVerdictLine(phaseJudgeText, "VERDICT") !== "PASS") {
pipelineSession.haltedAt = phaseLabel;
pipelineSession.phaseJudgeRationale = phaseJudgeText;
return finish("halted-phase-judge-failed");
}
const checkoutBase = await runGit(cwd, ["checkout", baseBranch]);
if (checkoutBase.code !== 0) {
pipelineSession.gitError = checkoutBase.out;
return finish("squash-crashed");
}
const squash = await runGit(cwd, ["merge", "--squash", phaseBranch]);
if (squash.code !== 0) {
await runGit(cwd, ["merge", "--abort"]);
pipelineSession.gitError = squash.out;
return finish("squash-crashed");
}
const commit = await runGit(cwd, ["commit", "-m", `feat: ${phaseLabel} (${[...phaseTaskIds].join(", ")})`]);
if (commit.code !== 0) {
pipelineSession.gitError = commit.out;
return finish("squash-crashed");
}
const push = await runGit(cwd, ["push", "origin", baseBranch]);
if (push.code !== 0) {
pipelineSession.gitError = push.out;
return finish("squash-push-crashed");
}
// Milestone's content now lives in baseBranch as one squashed commit --
// the phase branch (and whatever a prior restart already pushed of it)
// has no further reason to exist. Delete it both places so a future run
// never tries to resume a phase that's already done, and so origin
// doesn't accumulate one dangling branch per completed phase forever.
await runGit(cwd, ["branch", "-D", phaseBranch]);
await runGit(cwd, ["push", "origin", "--delete", phaseBranch]);
logProgress(pipelineSession);
}
return finish("completed");
}
async function runConcurrent(items, limit, worker) {
const results = new Array(items.length);
let i = 0;
async function next() {
while (i < items.length) {
const idx = i++;
results[idx] = await worker(items[idx], idx);
}
}
await Promise.all(Array.from({ length: Math.min(limit, items.length) }, next));
return results;
}
// How many repos can be mid-flight at once. Each repo gets its own clone
// and its own 4-agent pool (planner/investigator/implementer/judge), so
// this is now the real concurrency knob -- tasks within one repo are
// already serialized against that repo's pool (see runPhase). The backend
// (homelab-ornith) actually runs 2 GPU replicas behind one Kubernetes
// Service, each with its own copy of the model loaded (see homelab's
// k8s/apps/llm-serving/ornith.yaml) -- so up to 2 concurrent LLM calls get
// real independent instances; a 3rd+ concurrent call queues inside
// whichever replica the Service's own load-balancing lands it on (each
// replica runs OLLAMA_NUM_PARALLEL=1). REPO_CONCURRENCY above 2 is still
// useful (more repos in flight overlaps git/file work, not just LLM calls)
// but past 2 simultaneous LLM calls, extra concurrency mostly means queueing
// rather than added throughput -- bump the backend's replica count to
// change that, not this constant.
const REPO_CONCURRENCY = Number(process.env.REPO_CONCURRENCY) || 3;
// repoId is the repo's own name, not a random id -- it's what every role
// session's --name is built from (see runOnPool: `${repoId}-${role}`), so
// agent-manager's own session list groups naturally by repo ("portfolio-
// planner", "portfolio-judge", "poiman-planner", ...) instead of by opaque
// UUID. Takes the last path segment of the URL, strips a trailing `.git`,
// and sanitizes anything that isn't safe in a tmux session name / directory
// name / git branch name. Two different repos that happen to share a
// basename (e.g. two orgs' "portfolio") would collide -- not handled, since
// nothing about this harness's usage has needed more than one org per run.
function slugFor(repoUrl) {
const last = repoUrl.replace(/\/+$/, "").split("/").pop() || repoUrl;
return last.replace(/\.git$/, "").replace(/[^a-zA-Z0-9._-]/g, "-");
}
// Top-level entry point: runs every repo in `repos` to completion, up to
// REPO_CONCURRENCY at a time. Returns a map of repoId -> final
// pipelineSession, one per repo, independent of how the others fared.
async function runCoordinator({ repos, base, tasks, branchName }) {
const sessions = {};
await runConcurrent(repos, REPO_CONCURRENCY, async (repoUrl) => {
const repoId = slugFor(repoUrl);
const pipelineSession = {
id: repoId,
repo: repoUrl,
status: "running",
taskResults: [],
activeTasks: {},
totalTasks: 0,
startedAt: new Date().toISOString(),
};
sessions[repoId] = pipelineSession;
await runRepoPipeline({ repoId, repo: repoUrl, baseBranch: base, tasks, branchName }, pipelineSession);
});
return sessions;
}
function parseArgs(argv) {
const opts = { base: "main" };
for (let i = 0; i < argv.length; i++) {
const a = argv[i];
if (a === "--repo") opts.repo = argv[++i];
else if (a === "--repos") opts.repos = argv[++i];
else if (a === "--base") opts.base = argv[++i];
else if (a === "--tasks") opts.tasks = argv[++i];
else if (a === "--branch") opts.branch = argv[++i];
}
return opts;
}
async function main() {
const opts = parseArgs(process.argv.slice(2));
const repos = opts.repos ? opts.repos.split(",") : opts.repo ? [opts.repo] : null;
if (!repos || repos.length === 0) {
console.error(
"usage: coordinator.js --repos <url1,url2,...> [--tasks T0.1,T0.2;T1.1,T1.2,...] [--base main] [--branch <name>]\n" +
" --repo <url> also accepted for a single repo\n" +
" --tasks applies to every repo listed; omitted: each repo discovers its own phases from tasks/INDEX.md\n" +
" REPO_CONCURRENCY env var (default 3): how many repos run at once"
);
// process.exitCode + natural exit, not process.exit() -- stdout piped
// through kubectl exec (not a TTY) can drop buffered console.log/
// console.error output if the process exits before it flushes. Setting
// exitCode and letting the event loop drain naturally is the
// documented-safe way to exit with a specific code without racing it.
process.exitCode = 1;
return;
}
const phases = opts.tasks ? opts.tasks.split(";").map((phase) => phase.split(",")) : null;
const sessions = await runCoordinator({ repos, base: opts.base, tasks: phases, branchName: opts.branch });
process.exitCode = Object.values(sessions).every((s) => s.status === "completed") ? 0 : 1;
}
if (require.main === module) {
main();
}
module.exports = { runCoordinator, runRepoPipeline, runOnPool, parseTaskBoard };
+3 -63
View File
@@ -27,28 +27,6 @@ spec:
# container can't exec into another container's filesystem/PATH.
# It IS the container's long-running process now; no more `sleep
# infinity` placeholder.
#
# Also builds the agent-manager fork (github.com/Riotpiaole/
# agent-manager, add-headless-spawn branch) from source and drops
# coordinator.js in beside hub.js -- neither is the container's
# foreground process. hub.js keeps that role unchanged; coordinator.js
# itself now owns multi-repo concurrency (REPO_CONCURRENCY env,
# default 3), so one invocation handles every repo:
# `kubectl exec <pod> -- node /root/coordinator.js --repos
# repoA,repoB,... --tasks ...`. Each repo gets its own clone and its
# own persistent 4-agent pool (planner/investigator/implementer/
# judge, one agent-manager session per role, reused across every
# task in that repo) on the container's local tmux server --
# `kubectl exec -it <pod> -- agent-manager` attaches its TUI live
# against those same sessions, no cross-machine visibility problem
# since spawner, tmux server, and viewer are all colocated here.
#
# No prebuilt Linux binary is shipped for agent-manager: the local
# .bin/ build is macOS arm64 (wrong OS/arch for this container
# anyway) and it's 27MB, well over a ConfigMap's ~1MiB cap. Debian's
# `apt-get golang-go` is far too old for this fork's go 1.26.5
# requirement, so the real Go toolchain is fetched directly from
# go.dev instead.
- name: pi
image: node:22-slim
command:
@@ -56,50 +34,16 @@ spec:
- -c
- |
set -e
apt-get update && apt-get install -y git curl jq openssh-client tmux python3 sqlite3 gcc build-essential
apt-get update && apt-get install -y git curl jq openssh-client tmux
ssh-keygen -y -f /root/.ssh/id_forgejo > /root/.ssh/id_forgejo.pub
eval "$(ssh-agent -s)"
ssh-add /root/.ssh/id_forgejo
npm install -g @earendil-works/[email protected]
npm install --prefix /root ws
curl -fsSL "https://go.dev/dl/go1.26.5.linux-$(dpkg --print-architecture).tar.gz" | tar -C /usr/local -xz
export PATH="$PATH:/usr/local/go/bin"
git clone --branch add-headless-spawn --depth 1 \
https://github.com/Riotpiaole/agent-manager.git /root/agent-manager-src
(cd /root/agent-manager-src && go build -o /usr/local/bin/agent-manager .)
# Language toolchains for whatever repos the implementer/investigator/
# judge roles actually build and test -- go was already fetched above
# only for building agent-manager itself, and its PATH export above is
# local to this script, invisible to `kubectl exec` sessions into the
# already-running container. Symlinking both into /usr/local/bin (on
# PATH for every exec session, interactive or not) instead of relying
# on shell rc sourcing, which pi's non-interactive tool calls don't do.
ln -sf /usr/local/go/bin/go /usr/local/bin/go
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y --default-toolchain stable
ln -sf /root/.cargo/bin/cargo /usr/local/bin/cargo
ln -sf /root/.cargo/bin/rustc /usr/local/bin/rustc
ln -sf /root/.cargo/bin/rustup /usr/local/bin/rustup
node /root/hub.js
env:
- name: PI_BIN
value: pi
- name: AGENT_MANAGER_BIN
value: /usr/local/bin/agent-manager
- name: HUB_WORK_DIR
value: /root/agent-harness-work
# planner/investigator/implementer stay on the default
# (homelab-ornith/ornith:35b, pi's settings.json default). Judge
# moves to the separate homelab-reasoning backend (DeepSeek-R1,
# its own 2 GPU replicas) so judge calls stop contending with the
# other 3 roles for the 2 ornith pods -- an entire role's worth
# of traffic moves onto otherwise-idle capacity instead.
- name: JUDGE_PROVIDER
value: homelab-reasoning
- name: JUDGE_MODEL
value: reasoning
ports:
- containerPort: 9090
resources:
@@ -121,9 +65,6 @@ spec:
- name: hub-src
mountPath: /root/hub.js
subPath: hub.js
- name: coordinator-src
mountPath: /root/coordinator.js
subPath: coordinator.js
- name: ssh-key
mountPath: /root/.ssh/id_forgejo
subPath: id_forgejo
@@ -153,12 +94,11 @@ spec:
path: judge/SKILL.md
- key: resolver-SKILL.md
path: resolver/SKILL.md
- key: brave-search-SKILL.md
path: brave-search/SKILL.md
- name: hub-src
configMap:
name: hub-src
- name: coordinator-src
configMap:
name: coordinator-src
- name: ssh-key
secret:
secretName: agent-pod-ssh-key
-1
View File
@@ -5,7 +5,6 @@ resources:
- deployment.yaml
- configmap.yaml
- hub-configmap.yaml
- coordinator-configmap.yaml
- pi-skills-configmap.yaml
- ssh-configmap.yaml
- hub-service.yaml
+51 -17
View File
@@ -1,27 +1,63 @@
apiVersion: v1
data:
brave-search-SKILL.md: |
---
name: brave-search
description: Web search via the Brave Search API, called directly with curl. Use for searching documentation, facts, or any current web content.
allowed-tools: Bash
---
# Brave Search
Direct HTTP call to the Brave Search API — no separate script or package, just `curl` (`BRAVE_API_KEY` is already set in the environment).
## Search
```bash
curl -s -H "Accept: application/json" -H "X-Subscription-Token: $BRAVE_API_KEY" \
--get --data-urlencode "q=<query>" --data-urlencode "count=5" \
"https://api.search.brave.com/res/v1/web/search"
```
Options (add as extra `--data-urlencode` pairs):
- `count=<n>` — number of results (max 20, default 5)
- `country=<code>` — two-letter country code (default US)
- `freshness=pd|pw|pm|py` — past day/week/month/year, or `freshness=YYYY-MM-DDtoYYYY-MM-DD`
Response is JSON; the results live at `.web.results[]`, each with `title`, `url`, `description`, `age`. Pipe through `jq` if you want a shorter view, e.g.:
```bash
curl -s -H "Accept: application/json" -H "X-Subscription-Token: $BRAVE_API_KEY" \
--get --data-urlencode "q=<query>" "https://api.search.brave.com/res/v1/web/search" \
| jq -r '.web.results[] | "- \(.title)\n \(.url)\n \(.description)\n"'
```
There's no page-content-extraction helper here — if a result needs reading in full, `curl` the URL directly and read the raw HTML/text; don't expect readability-cleaned markdown.
## When to Use
- Searching for documentation or API references
- Looking up facts or current information
- Confirming a claim or approach against real sources
implementer-SKILL.md: |
---
name: implementer
description: Turns a confirmed PLAN.md into real code changes in the current checkout, committing incrementally. Use as the implementation stage of a spec-to-push pipeline, after planner and investigator have run.
allowed-tools: Read Grep Find Ls Write Edit Bash
---
Execute the already-agreed plan; don't re-litigate it. `PLAN.md` (plus any `## Investigation` flags) is the source of truth for *what*; use judgment only for *how*, within the codebase's existing conventions.
- Read `PLAN.md` top to bottom. Treat flagged/unconfirmed steps conservatively (safer, more literal reading; note it in the commit). Work steps in order. Commit after each meaningful step (`git add -A && git commit -m "..."`), not one giant commit — the judge stage needs real diff history.
- Push only if the task explicitly asks for it.
Match existing style. Don't refactor or "improve" code the plan didn't ask you to touch.
**UI/frontend changes:** don't trust that the code compiles as proof it works. Start the app (or its dev server) and use `npx playwright` via Bash to actually load the page and look — screenshot the affected view before and after your change, and click through the golden path the plan describes. `npx playwright screenshot <url> out.png` for a quick visual check; for interaction (clicks, form fills, navigation), write a small throwaway script under a scratch path (e.g. `/tmp/`, never committed) using `playwright` the library, run it with `node`, then delete it. This doesn't apply to non-UI work (a Rust library, a CLI, a backend-only change) — use judgment.
**Hard rules:**
- Follow DRY and SOLID. Don't duplicate logic that already exists elsewhere in the codebase you're touching — reuse or extract instead. Keep each unit responsible for one thing.
- Never commit anything that doesn't belong in source control: build artifacts, downloaded/vendored dependencies, secrets, scratch/debug files. `.gitignore` already blocks common patterns; if you create something outside those patterns, delete it before committing rather than relying on `.gitignore` to catch it.
**Never vendor a dependency by downloading/extracting it into the repo.** Use the language's real package manager (`cargo add`, `npm install`, etc.) so the dependency is declared in the manifest and lockfile, not a tarball or extracted source tree sitting in the checkout. If the package manager can't reach its registry from here, say so in your commit message rather than working around it — a later commit sweep (`git add -A`) commits whatever's in the checkout, including anything downloaded for a workaround, even if you never intended to keep it.
info-collector-SKILL.md: |
---
name: info-collector
@@ -35,7 +71,7 @@ data:
**Modes:**
- **Collect mode** (default) — search the web (`curl` against `https://api.search.brave.com/res/v1/web/search`, header `X-Subscription-Token: $BRAVE_API_KEY`, `--data-urlencode "q=<query>"`) with varied queries to cover the topic from multiple angles, then produce a structured summary: topic areas found, key facts, and links to sources for each. Do not editorialize about which approach is "right" — that's out of scope for this skill.
- **Collect mode** (default) — use the `brave-search` skill (`curl` against the Brave Search API) with varied queries to cover the topic from multiple angles, then produce a structured summary: topic areas found, key facts, and links to sources for each. Do not editorialize about which approach is "right" — that's out of scope for this skill.
- If asked to write the summary to a file, write it and report the path; otherwise return it directly in your response.
investigator-SKILL.md: |
---
@@ -46,7 +82,7 @@ data:
Check whether the plan's claims about the real world are actually true right now. Cite sources; don't assert without one.
- Read `PLAN.md` and the spec docs on disk. For each claim depending on external facts (a library's current API, a service's behavior), search the web (`curl` against `https://api.search.brave.com/res/v1/web/search`, header `X-Subscription-Token: $BRAVE_API_KEY`) to confirm or refute it. Append a `## Investigation` section to `PLAN.md`: each claim, its source(s), PASS/FLAG. Commit: `git add PLAN.md && git commit -m "investigate: confirm plan against sources"`.
- Read `PLAN.md` and the spec docs on disk. For each claim depending on external facts (a library's current API, a service's behavior), use `brave-search` to confirm or refute it. Append a `## Investigation` section to `PLAN.md`: each claim, its source(s), PASS/FLAG. Commit: `git add PLAN.md && git commit -m "investigate: confirm plan against sources"`.
- No `PLAN.md`? Just answer the question asked, citing sources.
Flag unconfirmed/contradicted claims rather than silently fixing them — that decision belongs to whoever reads the flag next.
@@ -58,25 +94,23 @@ data:
description: LLM-as-judge. Reviews a git diff against PLAN.md and the original spec, and returns a PASS/FAIL verdict with rationale. Use as the final review/report stage of a spec-to-push pipeline (the implementer stage already pushed; this reports on what shipped), or standalone to review any diff against stated criteria.
allowed-tools: Read Bash
---
Independent reviewer. Judge whether the implementation satisfies the plan and spec, on the evidence in front of you — not on how confident the commit messages sound. Don't rubber-stamp.
- Run `git diff <base-branch>...HEAD` to see exactly what changed. Compare against `PLAN.md`'s steps and the spec docs. Does every step have a corresponding change? Does the diff contradict any investigator flag? Anything obviously broken on inspection?
- **UI/frontend changes:** a diff that reads correctly can still render broken. Start the app and use `npx playwright` via Bash to actually look — screenshot the affected view, click through the golden path the plan/spec describes. FAIL on a visual defect the diff alone wouldn't show (broken layout, a control that doesn't do what its code claims, a state the plan promised that never renders). Doesn't apply to non-UI work — use judgment.
- FAIL on DRY/SOLID violations (duplicated logic that should reuse existing code, mixed-responsibility units) and on anything committed that doesn't belong in source control (build artifacts, vendored dependencies, secrets, scratch files) — name the specific file/lines in your rationale.
- No `PLAN.md`/base given? Review whatever diff/criteria are in the task directly.
You MUST end your final message with a literal verdict line, exactly one of:
```
VERDICT: PASS
```
```
VERDICT: FAIL
```
followed by your rationale. The pipeline driver parses this exact line mechanically to record the outcome — omitting it or rephrasing it breaks the pipeline.
planner-SKILL.md: |
---
name: planner
+83
View File
@@ -0,0 +1,83 @@
# API Auth Layer — Authentik service account + Kong JWT (model invoke)
Protect the model API (`api.riotpiao.com/*`, Kong OSS 3.9) so only an Authentik
service account holding a valid **client_credentials** JWT can invoke the KServe
models. "Invoke role" = **possession of a JWT from the dedicated model-invoke
OAuth2 provider** (only the service account can obtain one).
## Flow
```
service account ── client_credentials ──▶ Authentik token endpoint
(client_id + secret) https://authentik.riotpiao.com/application/o/token/
▼ RS256 JWT (iss = https://authentik.riotpiao.com/application/o/model-invoke/)
client ── Authorization: Bearer <jwt> ──▶ Kong (api.riotpiao.com/*)
jwt plugin: verify RS256 sig via Authentik JWKS,
check iss/exp → map to KongConsumer → allow
KServe model (reasoning / ornith / ...)
```
Kong OSS has no enterprise `openid-connect` plugin, so we use the built-in
**`jwt`** plugin: it validates an RS256 signature against a public key we pin on
a KongConsumer, keyed by the token's `iss`.
## Changes
### 1. Authentik (k8s/infra/iam/scripts/authentik-provision.py)
- New **service account** user `model-invoker` (type `service_account`, no
password; Authentik issues an app-password/token for M2M).
- New **OAuth2 provider + application** `model-invoke`:
- `client_type: confidential`, `grant_types: ["client_credentials"]`
- signing key = existing RS256 keypair (same as other providers)
- mappings: `openid` (+ optionally a static `invoke` scope) — no user scopes
needed for M2M.
- Client secret written to k8s Secret `api/model-invoke-oidc`
(keys `client-id`, `client-secret`), labelled for whoever consumes it.
- Bind the service account so it (and only it) can use the provider.
### 2. Kong (k8s/apps/api/, new file `model-auth.yaml`)
- **KongConsumer** `model-invoker` (ns api).
- **`jwt` credential** on that consumer (a Secret of type
`konghq.com/v1/credential`):
- `algorithm: RS256`
- `key` = the token `iss``https://authentik.riotpiao.com/application/o/model-invoke/`
- `rsa_public_key` = the PEM public key of Authentik's `model-invoke` signing
cert (fetched from Authentik JWKS / cert, stored in git or ksops).
- **KongPlugin** `jwt-auth` (`plugin: jwt`, `config.claims_to_verify: [exp]`).
### 3. Wire onto model routes (k8s/apps/api/llm-routes.yaml)
- Add `jwt-auth` to each model Ingress's `konghq.com/plugins` annotation
(currently e.g. `llm-rewrite-reasoning`) → becomes
`llm-rewrite-reasoning,jwt-auth`.
- Leave `/models` list route open OR protect too (decision).
## Client usage (after build)
```bash
TOKEN=$(curl -s https://authentik.riotpiao.com/application/o/token/ \
-d grant_type=client_credentials \
-d client_id=model-invoke \
-d client_secret=<secret> \
-d scope=openid | jq -r .access_token)
curl https://api.riotpiao.com/v1/chat/completions \
-H "Authorization: Bearer $TOKEN" -d '{...}'
```
## Test plan
1. No token → Kong returns 401.
2. Valid client_credentials token → 200, model responds.
3. Expired/garbage token → 401.
4. Confirm the `/models` route behaviour matches the decision.
## Open items / risks
- Authentik `client_credentials` for a *service account* may require an
**app-password / JWT-assertion** flow rather than plain client_secret POST —
verify Authentik 2026.x M2M exactly (client_credentials with client_secret vs
the SA token). Adjust step 1 accordingly before wiring Kong.
- Pinning `rsa_public_key`: Authentik key rotation would break it — document a
rotation runbook, or have the provision script re-export the cert PEM into the
Kong credential on each run (keeps them in sync, same idea as ksops secrets).
- Kong `jwt` maps token→consumer by the `iss`=`key` match; ensure the provider's
issuer is stable.
+8 -11
View File
@@ -5,17 +5,14 @@
#
# nginx terminates TLS with the wildcard *.riotpiao.com cert (served as its
# default-ssl-certificate, so no per-rule `tls:` block is needed) and forwards
# plain HTTP to the gateway.
# plain HTTP to kong-proxy. Kong then does the real routing, from Ingresses
# carrying `ingressClassName: kong`.
#
# Backend was kong-proxy:80 until Kong was retired on 2026-08-19; it is now the
# Go gateway's Service, api-gateway:8080, deployed from rock/homelab-frontend.
# Reverting the cutover is a change to these two lines and nothing else.
# Catch-all `/` on purpose: everything under this host belongs to Kong. Listing
# per-API paths here would duplicate Kong's routing table inside nginx, and the
# two copies would drift.
#
# Catch-all `/` on purpose: everything under this host belongs to the gateway.
# Listing per-API paths here would duplicate the gateway's routing table inside
# nginx, and the two copies would drift.
#
# In-cluster callers should prefer http://api-gateway.api.svc.cluster.local:8080
# In-cluster callers should prefer http://kong-proxy.api.svc.cluster.local
# directly. Resolving api.riotpiao.com sends them out to nginx and back in,
# which is a pointless hairpin unless they need TLS or the public hostname.
apiVersion: networking.k8s.io/v1
@@ -41,6 +38,6 @@ spec:
pathType: Prefix
backend:
service:
name: api-gateway
name: kong-proxy
port:
number: 8080
number: 80
+19
View File
@@ -0,0 +1,19 @@
# Cluster-wide Kong Prometheus plugin -- `global: "true"` label makes the
# ingress controller apply it to every route on this Kong instance, so all
# five LLM routes (ornith/reasoning/qwen/embeddings/rerank) get RED metrics
# without touching llm-routes.yaml. Scraped via kong-values.yaml's
# serviceMonitor (status listener, already on by chart default at :8100).
apiVersion: configuration.konghq.com/v1
kind: KongClusterPlugin
metadata:
name: prometheus
annotations:
kubernetes.io/ingress.class: kong
labels:
global: "true"
plugin: prometheus
config:
status_code_metrics: true
latency_metrics: true
bandwidth_metrics: true
upstream_health_metrics: true
+140
View File
@@ -0,0 +1,140 @@
# Kong Gateway — cluster-internal API gateway (namespace `api`).
#
# Chart: kong/kong 3.4.1 (appVersion 3.9). Only overrides are listed; every key
# here was checked against `helm show values kong/kong --version 3.4.1`, because
# Helm silently ignores unknown keys — a typo is a no-op, not an error.
#
# ── Topology ────────────────────────────────────────────────────────────────
# external: client -> nginx (TLS, wildcard *.riotpiao.com) -> kong-proxy:80
# internal: pod -> kong-proxy.api.svc.cluster.local:80
#
# nginx stays the single edge and the only LoadBalancer (192.168.1.160). Kong is
# the policy/routing layer behind it, so it needs no LB IP and no TLS of its own
# — hence ClusterIP and proxy.tls disabled. Giving Kong its own IP from
# homelab-pool would mean duplicating cert-manager wiring and diverging from the
# CoreDNS convention that sends every *.riotpiao.com host to nginx.
#
# ── Routing model ───────────────────────────────────────────────────────────
# Consumers publish an Ingress with `ingressClassName: kong`; the controller
# turns it into a Kong route. `nginx` remains the default IngressClass, so this
# is strictly opt-in and no existing Ingress changes behaviour.
# Without this the release name is prefixed onto everything (`kong-kong-proxy`).
# Pinning it keeps the Service name stable and independent of the release name,
# which matters because the nginx Ingress in k8s/bootstrap/ingress/ingress.yaml
# references it by name.
fullnameOverride: kong
# Two replicas so a node drain or rollout doesn't take the gateway down. Kong is
# stateless in DB-less mode, so replicas are pure redundancy.
replicaCount: 2
# Opt in to the `llm-serving-default-deny` NetworkPolicy, which admits port 8080
# only from pods carrying this label. That policy is a compensating control, not
# hygiene: vLLM v0.11.0 is frozen on Volta and will never receive patches for
# several remote/unauthenticated advisories, so it must not be broadly reachable.
#
# Without this label Cilium DROPS the packets rather than refusing them, so the
# symptom is a request that hangs until the client's timeout — not a connection
# error. /v1/models still worked while this was missing, because
# request-termination answers inside Kong and never touches an upstream.
podLabels:
llm-client: "true"
env:
# DB-less. Config comes from Kubernetes objects via the ingress controller, so
# git stays the source of truth. A Postgres-backed Kong would put live routing
# config in a database mutated through the Admin API — state outside git, plus
# migration Jobs on every upgrade.
database: "off"
# `nginx_proxy_<directive>` injects a directive into the proxy location block;
# this renders `proxy_buffering off;`.
#
# Required for LLM streaming. With buffering on (the default) nginx accumulates
# the upstream response before forwarding, so an SSE stream from
# `"stream": true` arrives in lumps or stalls until the generation finishes —
# which defeats the point of streaming. The matching setting is already on the
# nginx Ingress in ingress.yaml; both hops have to be unbuffered or the
# buffered one dominates.
nginx_proxy_proxy_buffering: "off"
# Any plugin that rewrites the request body — request-transformer on the
# llm-chat-* routes — reads it through `kong.request.get_body()`, and that
# returns nothing once nginx has spilled the body past
# client_body_buffer_size into a temp file. The plugin then re-serializes a
# body with no `messages`, and the upstream answers
# HTTP 400 {"error":{"message":"[] is too short - 'messages'"}}
# Measured on /v1/ornith/chat/completions: 10588 B -> 200, 11088 B -> 400.
# An agent request carrying tool schemas clears that in one turn, so the
# buffer has to hold a whole conversation, not a chat message.
nginx_http_client_body_buffer_size: "16m"
nginx_http_client_max_body_size: "16m"
ingressController:
enabled: true
ingressClass: kong
# The chart's ingress-class template is gated on
# `.Capabilities.APIVersions.Has "networking.k8s.io/v1/IngressClass"`, so a
# bare `helm template` renders nothing. ArgoCD passes --api-versions from the
# live cluster, so it does render there — verify `kubectl get ingressclass
# kong` after the first sync rather than assuming it.
createIngressClass: true
# Deliberately empty: setting is-default-class here would hijack every Ingress
# in the cluster that omits ingressClassName. nginx keeps that role.
ingressClassAnnotations: {}
proxy:
enabled: true
# Chart default is LoadBalancer, which would claim an IP from homelab-pool.
type: ClusterIP
http:
enabled: true
servicePort: 80
containerPort: 8000
# nginx already terminated TLS; a second handshake to the same cluster buys
# nothing and would need Kong to hold its own certificate.
tls:
enabled: false
# No Service for the Admin API. The controller reaches it over localhost inside
# the pod, so exposing it would only create an unauthenticated write path to the
# gateway's entire configuration.
admin:
enabled: false
# Kong Manager UI — chart default is `enabled: true` with type NodePort, which
# would open a port on every node. Not wanted.
manager:
enabled: false
resources:
requests:
cpu: 200m
memory: 256Mi
limits:
cpu: "2"
memory: 1Gi
podDisruptionBudget:
enabled: true
minAvailable: 1
# Status listener (metrics/health) is on by default at :8100 (chart default,
# verified via `helm show values`). This just wires the ServiceMonitor the
# chart already knows how to generate for it, so kong_http_requests_total /
# kong_latency_* / kong_bandwidth_bytes land in Prometheus. Paired with the
# cluster-wide `prometheus` KongClusterPlugin in kong-metrics.yaml.
serviceMonitor:
enabled: true
labels:
release: kube-prometheus-stack
# Spread the two replicas across nodes; `ScheduleAnyway` so a single-node
# situation degrades to co-location instead of leaving a pod Pending.
topologySpreadConstraints:
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: ScheduleAnyway
labelSelector:
matchLabels:
app.kubernetes.io/name: kong
app.kubernetes.io/instance: kong
+7 -7
View File
@@ -1,14 +1,14 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
# Explicit allowlist. Anything added to this directory and not listed here is
# silently dropped — no error, no drift shown.
#
# Down to a single Ingress since Kong was retired (2026-08-19). The Kong Helm
# values, the KongClusterPlugin for Prometheus, the six KongPlugin CRs behind
# the path-per-model LLM surface, the KongConsumer and the key-auth plan all
# went with it.
# Explicit allowlist so kong-values.yaml in this directory is NOT treated as a
# manifest — it is Helm input consumed by the chart source of the `kong`
# Application, not a Kubernetes object. Anything new added here must be listed
# or it is silently dropped with no error and no drift shown.
resources:
- ingress.yaml
- kong-metrics.yaml
- llm-routes.yaml
- model-auth.yaml
# No top-level `namespace:` transformer on purpose: ingress.yaml sets its own
# namespace, and the transformer rewrites metadata.namespace on every resource
# it builds, which is a trap for anything cross-namespace added later.
+317
View File
@@ -0,0 +1,317 @@
# LLM API surface on the Kong gateway — DeepSeek/OpenAI-shaped.
#
# These live in namespace `llm-serving`, not `api`, because a Kubernetes Ingress
# can only reference a Service in its own namespace and the predictor Services
# are there. The Kong ingress controller watches all namespaces, so the routes
# still land on the gateway. They are synced by the `kong` Application (which
# has a `path: k8s/apps/api` source) so all gateway config stays in one place.
#
# ── Model -> upstream map (verified live) ───────────────────────────────────
# reasoning -> reasoning-predictor vLLM, DeepSeek-R1-Distill-32B
# ornith:35b -> ornith-predictor Ollama
# qwen2.5:3b-instruct -> ornith-predictor Ollama (same pod!)
# nomic-embed-text-v2 -> embeddings-predictor TEI
# bge-reranker-base -> reranker-predictor TEI
# Qwen2.5-Math-PRM-7B -> verifier-predictor vLLM pooling
#
# ── Why path-per-model, and why the body is rewritten ───────────────────────
# Kong matches routes on host, path, method and headers — never on the request
# body. So a single /v1/chat/completions endpoint that dispatches on the body's
# `model` field is not expressible in Kong OSS (`ai-proxy-advanced`, which does
# multi-target model routing, is Enterprise-only).
#
# Hence the model is in the path. But `ornith:35b` and `qwen2.5:3b-instruct`
# share ONE Ollama pod, and Ollama still reads which model to load from the
# body's `model` field. If only the path selected the route, a client calling
# /v1/qwen/... with `"model": "ornith:35b"` in the body would silently get the
# 35B model. So each chat route force-overwrites `model` in the body, making the
# path the single source of truth. Callers may omit `model` entirely.
#
# ── Timeouts ───────────────────────────────────────────────────────────────
# Kong's upstream timeouts default to 60000ms. A 32B model generating a long
# answer on a Volta GPU routinely exceeds that, and the client would see a
# 504 mid-generation. Raised to 1h on every LLM route. Values are milliseconds.
# ── GET /v1/models ──────────────────────────────────────────────────────────
# Served entirely by Kong via request-termination: the plugin short-circuits in
# the access phase, so the backend below is never contacted. It only exists
# because an Ingress rule requires a backend.
#
# The list is static, which means it can drift from what the engines actually
# serve — notably if the Ollama pull list in the ornith InferenceService
# changes. Verify with:
# curl -s $SVC/v1/models (against each *-predictor)
apiVersion: configuration.konghq.com/v1
kind: KongPlugin
metadata:
name: llm-models-list
namespace: llm-serving
plugin: request-termination
config:
status_code: 200
content_type: application/json
body: |
{"object":"list","data":[
{"id":"reasoning","object":"model","owned_by":"homelab","created":0},
{"id":"ornith:35b","object":"model","owned_by":"homelab","created":0},
{"id":"qwen2.5:3b-instruct","object":"model","owned_by":"homelab","created":0},
{"id":"nomic-ai/nomic-embed-text-v2-moe","object":"model","owned_by":"homelab","created":0},
{"id":"BAAI/bge-reranker-base","object":"model","owned_by":"homelab","created":0},
{"id":"Qwen/Qwen2.5-Math-PRM-7B","object":"model","owned_by":"homelab","created":0}
]}
---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: llm-models
namespace: llm-serving
annotations:
konghq.com/plugins: llm-models-list,model-key-auth
konghq.com/strip-path: "false"
konghq.com/methods: "GET"
spec:
ingressClassName: kong
rules:
- host: api.riotpiao.com
http:
paths:
- path: /v1/models
pathType: Exact
backend:
# Never actually called — request-termination answers first.
service:
name: reasoning-predictor
port:
number: 80
---
# ── POST /v1/reasoning/chat/completions ─────────────────────────────────────
apiVersion: configuration.konghq.com/v1
kind: KongPlugin
metadata:
name: llm-rewrite-reasoning
namespace: llm-serving
plugin: request-transformer
config:
# `add` only applies when the field is absent, `replace` only when present.
# Both are needed to force the value in either case.
add:
body:
- "model:reasoning"
replace:
body:
- "model:reasoning"
# The model lives in the path for routing; the upstream still expects the
# canonical OpenAI path.
uri: /v1/chat/completions
---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: llm-chat-reasoning
namespace: llm-serving
annotations:
konghq.com/plugins: llm-rewrite-reasoning,model-key-auth
konghq.com/strip-path: "false"
konghq.com/methods: "POST"
konghq.com/connect-timeout: "10000"
konghq.com/read-timeout: "3600000"
konghq.com/write-timeout: "3600000"
spec:
ingressClassName: kong
rules:
- host: api.riotpiao.com
http:
paths:
- path: /v1/reasoning/chat/completions
pathType: Prefix
backend:
service:
name: reasoning-predictor
port:
number: 80
---
# ── POST /v1/ornith/chat/completions ────────────────────────────────────────
apiVersion: configuration.konghq.com/v1
kind: KongPlugin
metadata:
name: llm-rewrite-ornith
namespace: llm-serving
plugin: request-transformer
config:
add:
body:
- "model:ornith:35b"
replace:
body:
- "model:ornith:35b"
uri: /v1/chat/completions
---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: llm-chat-ornith
namespace: llm-serving
annotations:
konghq.com/plugins: llm-rewrite-ornith,model-key-auth
konghq.com/strip-path: "false"
konghq.com/methods: "POST"
konghq.com/connect-timeout: "10000"
konghq.com/read-timeout: "3600000"
konghq.com/write-timeout: "3600000"
spec:
ingressClassName: kong
rules:
- host: api.riotpiao.com
http:
paths:
- path: /v1/ornith/chat/completions
pathType: Prefix
backend:
service:
name: ornith-predictor
port:
number: 80
---
# ── POST /v1/qwen/chat/completions ──────────────────────────────────────────
# Same upstream pod as ornith — only the forced body `model` differs. Both stay
# resident because the engine runs with OLLAMA_MAX_LOADED_MODELS=2 and
# OLLAMA_KEEP_ALIVE=-1, so this does not trigger a model swap per request.
apiVersion: configuration.konghq.com/v1
kind: KongPlugin
metadata:
name: llm-rewrite-qwen
namespace: llm-serving
plugin: request-transformer
config:
add:
body:
- "model:qwen2.5:3b-instruct"
replace:
body:
- "model:qwen2.5:3b-instruct"
uri: /v1/chat/completions
---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: llm-chat-qwen
namespace: llm-serving
annotations:
konghq.com/plugins: llm-rewrite-qwen,model-key-auth
konghq.com/strip-path: "false"
konghq.com/methods: "POST"
konghq.com/connect-timeout: "10000"
konghq.com/read-timeout: "3600000"
konghq.com/write-timeout: "3600000"
spec:
ingressClassName: kong
rules:
- host: api.riotpiao.com
http:
paths:
- path: /v1/qwen/chat/completions
pathType: Prefix
backend:
service:
name: ornith-predictor
port:
number: 80
---
# ── POST /v1/embeddings ─────────────────────────────────────────────────────
# No path-per-model and no rewrite: there is exactly one embeddings backend, so
# there is nothing to disambiguate, and TEI already serves the canonical
# OpenAI path (verified: /v1/embeddings returns 405 to GET, i.e. it exists).
# That makes an OpenAI SDK a drop-in here.
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: llm-embeddings
namespace: llm-serving
annotations:
konghq.com/strip-path: "false"
konghq.com/methods: "POST"
konghq.com/connect-timeout: "10000"
konghq.com/read-timeout: "600000"
konghq.com/write-timeout: "600000"
spec:
ingressClassName: kong
rules:
- host: api.riotpiao.com
http:
paths:
- path: /v1/embeddings
pathType: Prefix
backend:
service:
name: embeddings-predictor
port:
number: 80
---
# ── POST /v1/rerank ─────────────────────────────────────────────────────────
# Rerank is not part of the OpenAI spec, and TEI serves it at /rerank — probing
# /v1/rerank returned 404 while /rerank returned 405, so this one genuinely
# needs the rewrite that embeddings does not.
apiVersion: configuration.konghq.com/v1
kind: KongPlugin
metadata:
name: llm-rewrite-rerank
namespace: llm-serving
plugin: request-transformer
config:
replace:
uri: /rerank
---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: llm-rerank
namespace: llm-serving
annotations:
konghq.com/plugins: llm-rewrite-rerank,model-key-auth
konghq.com/strip-path: "false"
konghq.com/methods: "POST"
konghq.com/connect-timeout: "10000"
konghq.com/read-timeout: "600000"
konghq.com/write-timeout: "600000"
spec:
ingressClassName: kong
rules:
- host: api.riotpiao.com
http:
paths:
- path: /v1/rerank
pathType: Prefix
backend:
service:
name: reranker-predictor
port:
number: 80
---
# ── POST /v1/score ──────────────────────────────────────────────────────────
# The process reward model. Returns scores, not tokens, so it is deliberately
# not under /chat/completions. vLLM serves /v1/score natively (verified), so no
# rewrite is needed.
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: llm-score
namespace: llm-serving
annotations:
konghq.com/strip-path: "false"
konghq.com/methods: "POST"
konghq.com/connect-timeout: "10000"
konghq.com/read-timeout: "600000"
konghq.com/write-timeout: "600000"
spec:
ingressClassName: kong
rules:
- host: api.riotpiao.com
http:
paths:
- path: /v1/score
pathType: Prefix
backend:
service:
name: verifier-predictor
port:
number: 80
+49
View File
@@ -0,0 +1,49 @@
# API auth layer — Kong key-auth on the model routes.
#
# The model API (api.riotpiao.com/v1/...) requires a static API key, presented
# OpenAI-style as `Authorization: Bearer <key>` (or `apikey: <key>`). The key
# lives in the ksops-managed Secret model-invoke-apikey (labelled
# konghq.com/credential: key-auth) and is bound to the KongConsumer below.
#
# Issue the key to rock; use it as the OpenAI SDK api_key. Rotate by updating the
# ksops secret. This is self-contained in Kong — the invoke path does not depend
# on an Authentik token (Authentik still fronts every *human* dashboard SSO).
---
apiVersion: configuration.konghq.com/v1
kind: KongConsumer
metadata:
name: model-invoker
namespace: api
annotations:
kubernetes.io/ingress.class: kong
username: model-invoker
credentials:
- model-invoke-apikey
---
# key-auth: require the API key on the model routes. key_in_header accepts the
# `apikey` header; key_in_bearer accepts `Authorization: Bearer <key>` so any
# OpenAI-compatible SDK (api_key=..., base_url=https://api.riotpiao.com/v1) works
# unchanged.
#
# Namespace `llm-serving`, not `api`: the ingress controller resolves a
# `konghq.com/plugins` annotation against the annotated object's OWN namespace,
# and all five model routes in llm-routes.yaml live in llm-serving. While this
# sat in `api` the reference dangled, the plugin never bound, and every model
# route served traffic with no key at all — verified: an unauthenticated
# /v1/models and /v1/ornith/chat/completions both returned 200. A dangling
# plugin reference is silent; it fails open, so re-test without a key after any
# move rather than trusting that the object exists.
apiVersion: configuration.konghq.com/v1
kind: KongPlugin
metadata:
name: model-key-auth
namespace: llm-serving
plugin: key-auth
config:
key_names:
- apikey
- authorization
key_in_header: true
key_in_query: false
key_in_body: false
hide_credentials: true
-14
View File
@@ -1,14 +0,0 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: immich-config
data:
DB_HOSTNAME: "immich-db-rw"
DB_DATABASE_NAME: "immich"
# Only pgvector is installed (see db.yaml) - no vectorchord extension image
# exists for pg18 in CNPG's catalog yet. Explicit instead of relying on
# auto-detect's vectorchord-first preference order.
DB_VECTOR_EXTENSION: "pgvector"
REDIS_HOSTNAME: "immich-redis"
IMMICH_MACHINE_LEARNING_URL: "http://immich-machine-learning:3003"
TZ: "America/Los_Angeles"
-58
View File
@@ -1,58 +0,0 @@
# Dedicated CNPG Postgres for Immich. Same recipe as paperless-db/authentik-db
# (2 instances, default longhorn storage class) except the operand is
# PostgreSQL 18, not 16.2 - the official CNPG pgvector extension image
# (ghcr.io/cloudnative-pg/pgvector) is only published for pg18, no pg16 tags
# exist in that registry. Immich itself supports pg18 fine (immich-app's own
# postgres image already ships 18-vectorchord builds).
#
# pgvector loaded via CNPG's ImageVolume extension mechanism (CNPG 1.27+,
# k8s ImageVolume feature - both present here: operator is 1.30.0, cluster is
# v1.36.1). No shared_preload_libraries needed - pgvector doesn't require
# preload, just CREATE EXTENSION, which immich-server issues itself at
# startup. Distro/pg-major must match between the operand image and the
# extension image (both "18"+"trixie" here) - CNPG's own compatibility rule.
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: immich-db
annotations:
argocd.argoproj.io/sync-options: SkipDryRunOnMissingResource=true
spec:
instances: 2
imageName: ghcr.io/cloudnative-pg/postgresql:18-minimal-trixie
postgresql:
extensions:
- name: pgvector
image:
reference: ghcr.io/cloudnative-pg/pgvector:0.8.1-18-trixie
bootstrap:
initdb:
database: immich
owner: app
encoding: UTF8
localeCollate: C
localeCType: C
# CREATE EXTENSION vector requires superuser (pgvector's control file
# isn't marked trusted) and the "app" owner role isn't one
# (enableSuperuserAccess: false, repo convention) - postInitApplicationSQL
# runs as superuser during initdb, before the app ever connects. Only
# fires on a fresh bootstrap; the live cluster already had this run
# manually once (kubectl exec ... psql -U postgres -c 'CREATE EXTENSION').
postInitApplicationSQL:
- "CREATE EXTENSION IF NOT EXISTS vector;"
- "CREATE EXTENSION IF NOT EXISTS cube;"
- "CREATE EXTENSION IF NOT EXISTS earthdistance;"
enableSuperuserAccess: false
resources:
requests: { memory: "512Mi", cpu: "250m" }
limits: { memory: "2Gi", cpu: "1" }
storage:
size: 20Gi
storageClass: longhorn
affinity:
podAntiAffinityType: preferred
topologyKey: kubernetes.io/hostname
tolerations:
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
-100
View File
@@ -1,100 +0,0 @@
# immich-server: pinned to talos-cp-3, same reasoning as paperless
# (deployment.yaml comment there) - immich-media is a ReadWriteOnce Longhorn
# volume with a single replica physically on that node's disk (shared with
# paperless-media on the same 4TB HDD). Recreate strategy for the same
# reason: two pods can't both attach an RWO volume.
apiVersion: apps/v1
kind: Deployment
metadata:
name: immich-server
spec:
replicas: 1
strategy:
type: Recreate
selector:
matchLabels:
app: immich-server
template:
metadata:
labels:
app: immich-server
spec:
serviceAccountName: immich
nodeSelector:
kubernetes.io/hostname: talos-cp-3
containers:
- name: immich-server
image: ghcr.io/immich-app/immich-server:release
ports:
- containerPort: 2283
envFrom:
- configMapRef:
name: immich-config
env:
- name: DB_USERNAME
valueFrom:
secretKeyRef:
name: immich-db-app
key: username
- name: DB_PASSWORD
valueFrom:
secretKeyRef:
name: immich-db-app
key: password
# Composed by k8s/infra/iam's provisioning script (system-config
# JSON, oauth section) - see immich-oidc Secret.
- name: IMMICH_CONFIG_FILE
value: /config/immich.json
resources:
requests: { cpu: "500m", memory: "1Gi" }
limits: { cpu: "2", memory: "4Gi" }
volumeMounts:
- name: media
mountPath: /usr/src/app/upload
- name: oidc-config
mountPath: /config
readOnly: true
volumes:
- name: media
persistentVolumeClaim:
claimName: immich-media
- name: oidc-config
secret:
secretName: immich-oidc
items:
- key: config.json
path: immich.json
---
# CPU-only for now - the cluster's one GPU node (worker-1) is already
# dedicated to llm-serving predictors. Not node-pinned: its cache PVC is on
# the default 3-replica pool, not the single-disk cp-3 HDD.
apiVersion: apps/v1
kind: Deployment
metadata:
name: immich-machine-learning
spec:
replicas: 1
selector:
matchLabels:
app: immich-machine-learning
template:
metadata:
labels:
app: immich-machine-learning
spec:
serviceAccountName: immich
containers:
- name: immich-machine-learning
image: ghcr.io/immich-app/immich-machine-learning:release
ports:
- containerPort: 3003
resources:
requests: { cpu: "500m", memory: "1Gi" }
limits: { cpu: "2", memory: "4Gi" }
volumeMounts:
- name: ml-cache
mountPath: /cache
volumes:
- name: ml-cache
persistentVolumeClaim:
claimName: immich-ml-cache
-24
View File
@@ -1,24 +0,0 @@
# Direct nginx ingress, same reasoning as paperless: large uploads (photos/
# videos) and long-lived operations (video transcode, big batch uploads) need
# proxy-body-size/timeouts raised past nginx's defaults.
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: immich
annotations:
nginx.ingress.kubernetes.io/proxy-body-size: "0"
nginx.ingress.kubernetes.io/proxy-read-timeout: "600"
nginx.ingress.kubernetes.io/proxy-send-timeout: "600"
spec:
ingressClassName: nginx
rules:
- host: img.riotpiao.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: immich-server
port:
number: 2283
-14
View File
@@ -1,14 +0,0 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
namespace: immich
resources:
- db.yaml
- pvc.yaml
- configmap.yaml
- redis.yaml
- deployment.yaml
- service.yaml
- ingress.yaml
- rbac.yaml
# immich-oidc Secret written by the PostSync provisioning Job in
# k8s/infra/iam (same as paperless-oidc) - not duplicated here.
-38
View File
@@ -1,38 +0,0 @@
# Two volumes:
#
# - media: original photos/videos + generated thumbnails/encoded videos.
# Shares the cp-3 USB HDD with paperless-media, same StorageClass/disk tag,
# single replica (single disk, no redundancy possible - same tradeoff
# paperless already accepts). Sized 1400Gi, not 2000Gi: the disk's real
# usable capacity (~3724GiB, formatting overhead) minus paperless-media's
# 2000Gi and ~231GiB of other apps' default-class replicas that Longhorn
# placed here anyway (disk tags only pull matching volumes in, they don't
# exclude non-matching ones when the untagged pool elsewhere is full) only
# leaves ~1493Gi of real scheduling headroom right now.
# - ml-cache: downloaded ML model weights for immich-machine-learning
# (face detection / CLIP embeddings). Small, disposable (re-downloads on
# loss), but persisted so a pod restart doesn't re-pull multi-GB models -
# default 3-replica pool, not node-pinned.
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: immich-media
spec:
accessModes:
- ReadWriteOnce
storageClassName: longhorn-paperless-media
resources:
requests:
storage: 1400Gi
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: immich-ml-cache
spec:
accessModes:
- ReadWriteOnce
storageClassName: longhorn
resources:
requests:
storage: 5Gi
-43
View File
@@ -1,43 +0,0 @@
# Scoped operator access for immich-admins: restart/config-edit rights on
# just this service's own resources, nothing CNPG-managed (immich-db-*) or
# provisioning-managed (immich-oidc). Same pattern as
# k8s/apps/paperless/rbac.yaml. Inert until kube-apiserver's OIDC wiring
# lands (--oidc-groups-claim=groups, --oidc-groups-prefix=oidc:).
apiVersion: v1
kind: ServiceAccount
metadata:
name: immich
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: immich-operator
rules:
- apiGroups: ["apps"]
resources: ["deployments"]
resourceNames: ["immich-server", "immich-machine-learning"]
verbs: ["get", "list", "watch", "update", "patch"]
- apiGroups: [""]
resources: ["configmaps"]
resourceNames: ["immich-config"]
verbs: ["get", "list", "watch", "update", "patch"]
- apiGroups: [""]
resources: ["secrets"]
resourceNames: ["immich-oidc"]
verbs: ["get", "list", "watch", "update", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: immich-admins-binding
subjects:
- kind: Group
name: "oidc:immich-admins"
apiGroup: rbac.authorization.k8s.io
- kind: ServiceAccount
name: immich
namespace: immich
roleRef:
kind: Role
name: immich-operator
apiGroup: rbac.authorization.k8s.io
-37
View File
@@ -1,37 +0,0 @@
# Job queue broker for immich-server. No PVC: queue state is disposable - a
# lost queue on restart just re-triggers the affected background jobs
# (thumbnail generation, ML jobs, etc.), no photo data loss since originals
# live on immich-media.
apiVersion: apps/v1
kind: Deployment
metadata:
name: immich-redis
spec:
replicas: 1
selector:
matchLabels:
app: immich-redis
template:
metadata:
labels:
app: immich-redis
spec:
containers:
- name: redis
image: redis:7-alpine
ports:
- containerPort: 6379
resources:
requests: { cpu: "50m", memory: "64Mi" }
limits: { cpu: "250m", memory: "256Mi" }
---
apiVersion: v1
kind: Service
metadata:
name: immich-redis
spec:
selector:
app: immich-redis
ports:
- port: 6379
targetPort: 6379
-21
View File
@@ -1,21 +0,0 @@
apiVersion: v1
kind: Service
metadata:
name: immich-server
spec:
selector:
app: immich-server
ports:
- port: 2283
targetPort: 2283
---
apiVersion: v1
kind: Service
metadata:
name: immich-machine-learning
spec:
selector:
app: immich-machine-learning
ports:
- port: 3003
targetPort: 3003
+1
View File
@@ -14,5 +14,6 @@ resources:
- ornith.yaml
- reasoning.yaml
- reranker.yaml
- verifier.yaml
# No namespace transformer: every file sets its own, and the transformer would
# rewrite metadata.namespace on anything cross-namespace added later.
+15 -15
View File
@@ -3,15 +3,20 @@ kind: InferenceService
metadata:
annotations:
serving.kserve.io/deploymentMode: RawDeployment
# The konghq.com/{connect,read,write}-timeout annotations that used to live
# here went with Kong (retired 2026-08-19). They existed because Kong read
# its upstream timeouts off the Kubernetes Service, and its 60s default cut
# off the first request after any pod restart — a restart flushes VRAM and
# reloading ornith:35b takes longer than that. OLLAMA_KEEP_ALIVE=-1 hid the
# problem in steady state.
# Kong reads its timeouts from the Kubernetes Service, not the Ingress —
# Ingress annotations configure Route entities (strip-path, methods,
# plugins), these configure the Service entity. They were on
# llm-chat-ornith's Ingress and therefore ignored, leaving Kong's 60s
# default in force. KServe propagates InferenceService annotations to the
# Service it generates, which is how they reach Kong from here.
#
# The equivalent budget now belongs to the Go gateway's per-route timeout
# config in rock/homelab-frontend, not to an annotation on this object.
# This was invisible while OLLAMA_KEEP_ALIVE=-1 kept the model resident: no
# request ever waited on a cold load. A pod restart flushes VRAM, and
# loading ornith:35b takes longer than 60s, so the first request after any
# restart returned 504.
konghq.com/connect-timeout: "10000"
konghq.com/read-timeout: "3600000"
konghq.com/write-timeout: "3600000"
labels:
app.kubernetes.io/name: llm-ornith
app.kubernetes.io/part-of: llm-serving
@@ -91,13 +96,8 @@ spec:
name: models
deploymentStrategy:
type: Recreate
# 2 replicas -- each its own GPU, each loading both ornith:35b and
# qwen2.5:3b-instruct -- so 2 concurrent implementer-style calls each
# get an independent instance instead of contending on one, at the
# cost of judge/qwen traffic still sharing whichever replica an
# implementer call also lands on.
maxReplicas: 2
minReplicas: 2
maxReplicas: 1
minReplicas: 1
nodeSelector:
kubernetes.io/hostname: worker-1
runtimeClassName: nvidia
+10 -51
View File
@@ -12,59 +12,18 @@ spec:
predictor:
containers:
- args:
# bnb-4bit retired: no int4 tensor cores on sm70/V100, dequant-then-
# matmul is two slow kernel launches instead of one fused int4 GEMM,
# decode crawled at 2.5-10 tok/s regardless of TP/PP. Switched to
# JunHowie/Qwen3-32B-GPTQ-Int4 -- same dense Qwen3-32B weights, same
# hermes/qwen3 parser stack (no narration-bug risk, same as before),
# only the quant format changes. Plain (non-Marlin) GPTQ kernel is
# confirmed Volta-compatible; Marlin needs sm80+ and vLLM would try
# to auto-upgrade to it, so --quantization is pinned explicitly to
# `gptq` to force the plain kernel. Verified checkpoint size: 19.34GB
# (summed from the real safetensors index, not bits-per-param math).
# max-model-len=131072 is Qwen3-32B's real ceiling (config.json YaRN:
# factor=4.0, original_max_position_embeddings=32768) -- 200k was
# asked for but exceeds this architecturally regardless of VRAM.
# KV cache math: 256KB/token total (64 layers, 8 KV heads, 128
# head_dim, fp16), PP=2 splits both weights and KV load ~evenly, so
# each GPU carries ~9.67GB weights + ~128KB/token KV. At
# gpu-memory-utilization=0.90 (28.8GB/GPU usable), that leaves
# ~19.1GB/GPU for KV cache -> ~156k tokens/GPU capacity, comfortably
# above the 131072 target with room to spare -- the old
# OffloadingConnector CPU-DRAM spillover (tuned for the previous
# model's much smaller 16384 context) is no longer needed and is
# dropped. Staying on PP=2 and vLLM 0.11.0 (no version bump needed,
# this checkpoint only requires vllm>=0.9.2) -- plain GPTQ has no
# TP>1 restriction unlike bnb, so tensor-parallel-size=2 is worth
# trying later, but not risking a parallelism-strategy change in the
# same rollout as the quant+context-length change.
# This GPTQ requant's own config.json ships max_position_embeddings=
# 40960 and rope_scaling=None -- confirmed directly (curl'd the raw
# config.json), the base Qwen3-32B repo's YaRN block did NOT carry
# over during quantization. Re-applying it explicitly here restores
# the same math the base model documents (32768 * 4.0 = 131072);
# without this, --max-model-len=131072 fails ModelConfig validation
# against the checkpoint's own (unscaled) 40960 ceiling.
- --model=JunHowie/Qwen3-32B-GPTQ-Int4
- --model=unsloth/DeepSeek-R1-Distill-Qwen-32B-bnb-4bit
- --served-model-name=reasoning
- --quantization=gptq
- --quantization=bitsandbytes
- --dtype=float16
- --kv-cache-dtype=auto
- --rope-scaling={"rope_type":"yarn","factor":4.0,"original_max_position_embeddings":32768}
- --tensor-parallel-size=1
- --pipeline-parallel-size=2
- --max-model-len=131072
- --max-model-len=16384
- --gpu-memory-utilization=0.90
- --max-num-seqs=4
- --enable-chunked-prefill
- --enable-prefix-caching
# qwen3 is vLLM's dedicated reasoning parser for this family's <think>
# blocks.
- --reasoning-parser=qwen3
# hermes is the documented tool-call parser for general (non-Coder)
# Qwen3 models -- native chat template support, not narrated text.
- --enable-auto-tool-choice
- --tool-call-parser=hermes
- --reasoning-parser=deepseek_r1
- --host=0.0.0.0
- --port=8080
env:
@@ -87,12 +46,12 @@ spec:
resources:
limits:
cpu: '16'
memory: 36Gi
nvidia.com/gpu: '2'
memory: 16Gi
nvidia.com/gpu: '1'
requests:
cpu: '8'
memory: 12Gi
nvidia.com/gpu: '2'
memory: 8Gi
nvidia.com/gpu: '1'
startupProbe:
failureThreshold: 80
httpGet:
@@ -106,8 +65,8 @@ spec:
name: shm
deploymentStrategy:
type: Recreate
maxReplicas: 1
minReplicas: 1
maxReplicas: 2
minReplicas: 2
nodeSelector:
kubernetes.io/hostname: worker-1
runtimeClassName: nvidia
+76
View File
@@ -0,0 +1,76 @@
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
annotations:
serving.kserve.io/deploymentMode: RawDeployment
labels:
app.kubernetes.io/name: llm-verifier
app.kubernetes.io/part-of: llm-serving
name: verifier
namespace: llm-serving
spec:
predictor:
containers:
- args:
- --model=Qwen/Qwen2.5-Math-PRM-7B
- --served-model-name=verifier
- --runner=pooling
- --dtype=float16
- --tensor-parallel-size=1
- --max-model-len=4096
- --max-num-seqs=8
- --host=0.0.0.0
- --port=8080
env:
- name: VLLM_USE_FLASHINFER_SAMPLER
value: '0'
- name: VLLM_ATTENTION_BACKEND
value: XFORMERS
- name: HF_HOME
value: /mnt/models
image: vllm/vllm-openai:v0.11.0@sha256:014a95f21c9edf6abe0aea6b07353f96baa4ec291c427bb1176dc7c93a85845c
name: kserve-container
ports:
- containerPort: 8080
protocol: TCP
readinessProbe:
httpGet:
path: /health
port: 8080
periodSeconds: 10
resources:
limits:
cpu: '16'
memory: 16Gi
nvidia.com/gpu: '1'
requests:
cpu: '4'
memory: 8Gi
nvidia.com/gpu: '1'
startupProbe:
failureThreshold: 60
httpGet:
path: /health
port: 8080
periodSeconds: 15
volumeMounts:
- mountPath: /mnt/models
name: models
- mountPath: /dev/shm
name: shm
deploymentStrategy:
type: Recreate
maxReplicas: 1
minReplicas: 1
nodeSelector:
kubernetes.io/hostname: worker-1
runtimeClassName: nvidia
volumes:
- name: models
persistentVolumeClaim:
claimName: llm-models
- emptyDir:
medium: Memory
sizeLimit: 1Gi
name: shm
@@ -13,7 +13,6 @@ spec:
labels:
app: management-service
spec:
serviceAccountName: kmsvc
topologySpreadConstraints:
- maxSkew: 1
topologyKey: kubernetes.io/hostname
@@ -2,7 +2,7 @@ namespace: sqs
replicaCount: 3
image:
repository: forgejo.riotpiao.com/rock/kmsvc-manage
repository: ghcr.io/riotpiaole/kmsvc-management-service
tag: latest
pullPolicy: Always
@@ -1,6 +0,0 @@
apiVersion: v2
name: memory-queues
description: Kafka queues (DLQ) for Poimen Memory service (Phase 6.6)
type: application
version: 0.1.0
appVersion: "1.0"
@@ -1,20 +0,0 @@
{{- range .Values.queues }}
---
apiVersion: kmsvc.io/v1alpha1
kind: Queue
metadata:
name: {{ .name }}
namespace: {{ $.Values.namespace }}
labels:
app: memory-service
queue: dlq
spec:
name: {{ .name }}
description: {{ .description }}
partitions: {{ .partitions }}
replicationFactor: {{ .replicationFactor }}
config:
retention.ms: "{{ .config.retention.ms }}"
message.retention.seconds: "{{ .config.message.retention.seconds }}"
visibility.timeout.seconds: "{{ .config.visibility.timeout.seconds }}"
{{- end }}
@@ -1,25 +0,0 @@
# Poimen Memory Service Kafka Queues (kmsvc)
# Phase 6.6: DLQ topics for webhook + metrics failures
queues:
# DLQ for extraction, webhook, and agent failures
- name: poimen-memory-dlq
description: "DLQ for extraction, webhook, and agent failures"
partitions: 3
replicationFactor: 1
config:
retention.ms: "1209600000" # 14 days
message.retention.seconds: "1209600"
visibility.timeout.seconds: "300"
# DLQ for metrics persistence failures
- name: poimen-memory-metric-dlq
description: "DLQ for metrics persistence failures"
partitions: 3
replicationFactor: 1
config:
retention.ms: "1209600000" # 14 days
message.retention.seconds: "1209600"
visibility.timeout.seconds: "300"
namespace: sqs
@@ -18,6 +18,15 @@ rules:
- apiGroups: ["kmsvc.io"]
resources: ["queues/finalizers"]
verbs: ["update"]
- apiGroups: ["kmsvc.io"]
resources: ["temporalworkers"]
verbs: ["get", "list", "watch", "create", "update", "patch", "delete"]
- apiGroups: ["kmsvc.io"]
resources: ["temporalworkers/status"]
verbs: ["get", "update", "patch"]
- apiGroups: ["kmsvc.io"]
resources: ["temporalworkers/finalizers"]
verbs: ["update"]
- apiGroups: ["coordination.k8s.io"]
resources: ["leases"]
verbs: ["get", "list", "watch", "create", "update", "patch", "delete"]
@@ -27,6 +36,9 @@ rules:
- apiGroups: [""]
resources: ["pods", "nodes"]
verbs: ["get"]
- apiGroups: ["apps"]
resources: ["deployments"]
verbs: ["get", "list", "watch", "create", "update", "patch", "delete"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
@@ -0,0 +1,62 @@
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
name: temporalworkers.kmsvc.io
spec:
group: kmsvc.io
names:
kind: TemporalWorker
plural: temporalworkers
singular: temporalworker
scope: Namespaced
versions:
- name: v1
served: true
storage: true
schema:
openAPIV3Schema:
type: object
required:
- spec
properties:
apiVersion:
type: string
kind:
type: string
metadata:
type: object
spec:
type: object
description: Temporal worker specification
properties:
namespace:
type: string
description: Temporal namespace
taskQueue:
type: string
description: Task queue name
workflowTypes:
type: array
items:
type: string
description: List of workflow types to execute
activityTypes:
type: array
items:
type: string
description: List of activity types to execute
concurrency:
type: integer
minimum: 1
description: Worker concurrency level
status:
type: object
description: Temporal worker status
properties:
ready:
type: boolean
lastHeartbeat:
type: string
format: date-time
error:
type: string
+1 -1
View File
@@ -1,7 +1,7 @@
namespace: sqs
image:
repository: forgejo.riotpiao.com/rock/kmsvc-manage
repository: ghcr.io/riotpiaole/kmsvc-management-service
tag: latest
pullPolicy: Always
-95
View File
@@ -1,95 +0,0 @@
# Overrides paperless-ngx's own paperless/adapter.py at the same import path
# (mounted via subPath in deployment.yaml) - settings.py hardcodes
# SOCIALACCOUNT_ADAPTER = "paperless.adapter.CustomSocialAccountAdapter", so
# no Django setting needs to change, just the file content underneath it.
#
# Stock CustomSocialAccountAdapter.populate_user() is a stub ("kept in case
# global default permissions are implemented in the future" - they aren't),
# so every OIDC signup lands with zero permissions and 403s on every API
# endpoint. This adds the actual mapping: Authentik's "permissions" claim
# (via the permissions scope, requested in PAPERLESS_SOCIALACCOUNT_PROVIDERS,
# computed server-side from group membership by authentik-provision.py) ->
# "paperless:write" or "*" (homelab-admins) grants is_staff+is_superuser,
# same convention already used for MinIO's policy claim and Grafana's
# role_attribute_path. Checking the permission string rather than a literal
# group name decouples "what grants access" from which group happens to
# hold it - same pattern applies to every other service's Role/RoleBinding
# in k8s/infra/rbac/.
apiVersion: v1
kind: ConfigMap
metadata:
name: paperless-adapter
data:
adapter.py: |
from urllib.parse import quote
from allauth.account.adapter import DefaultAccountAdapter
from allauth.core import context
from allauth.socialaccount.adapter import DefaultSocialAccountAdapter
from django.conf import settings
from django.forms import ValidationError
from django.urls import reverse
REQUIRED_PERMISSIONS = {"paperless:write", "*"}
class CustomAccountAdapter(DefaultAccountAdapter):
def is_open_for_signup(self, request):
allow_signups = super().is_open_for_signup(request)
return getattr(settings, "ACCOUNT_ALLOW_SIGNUPS", allow_signups)
def pre_authenticate(self, request, **credentials):
if settings.DISABLE_REGULAR_LOGIN:
raise ValidationError("Regular login is disabled")
return super().pre_authenticate(request, **credentials)
def is_safe_url(self, url):
from django.utils.http import url_has_allowed_host_and_scheme
allowed_hosts = {context.request.get_host()} | set(settings.ALLOWED_HOSTS)
if "*" in allowed_hosts:
allowed_hosts.remove("*")
allowed_hosts.add(context.request.get_host())
return url_has_allowed_host_and_scheme(url, allowed_hosts=allowed_hosts)
return url_has_allowed_host_and_scheme(url, allowed_hosts=allowed_hosts)
def get_reset_password_from_key_url(self, key):
if settings.PAPERLESS_URL is None:
return super().get_reset_password_from_key_url(key)
path = reverse(
"account_reset_password_from_key",
kwargs={"uidb36": "UID", "key": "KEY"},
)
path = path.replace("UID-KEY", quote(key))
return settings.PAPERLESS_URL + path
class CustomSocialAccountAdapter(DefaultSocialAccountAdapter):
def is_open_for_signup(self, request, sociallogin):
allow_signups = super().is_open_for_signup(request, sociallogin)
return getattr(settings, "SOCIALACCOUNT_ALLOW_SIGNUPS", allow_signups)
def get_connect_redirect_url(self, request, socialaccount):
return reverse("base")
def populate_user(self, request, sociallogin, data):
user = super().populate_user(request, sociallogin, data)
perms = set(sociallogin.account.extra_data.get("permissions") or [])
if perms & REQUIRED_PERMISSIONS:
user.is_staff = True
user.is_superuser = True
return user
def save_user(self, request, sociallogin, form=None):
# populate_user() sets the flags on the in-memory user, but
# allauth's default save_user() re-derives is_staff from
# ACCOUNT_DEFAULT_HTTP_PROTOCOL-independent defaults and can
# overwrite them on save - re-apply after super().save_user()
# persists the row, matching the permissions check above exactly.
user = super().save_user(request, sociallogin, form)
perms = set(sociallogin.account.extra_data.get("permissions") or [])
if perms & REQUIRED_PERMISSIONS and not (user.is_staff and user.is_superuser):
user.is_staff = True
user.is_superuser = True
user.save(update_fields=["is_staff", "is_superuser"])
return user
-94
View File
@@ -1,94 +0,0 @@
# Nightly: pg_dump the paperless DB + mirror the media PVC into the scoped
# `paperless` MinIO bucket (see minio-provision-paperless-job.yaml). This is a
# BACKUP target, not live storage - paperless-ngx has no native S3 backend, it
# only ever reads/writes the local media PVC directly.
#
# Pinned to talos-cp-3, same as deployment.yaml: media is a ReadWriteOnce
# Longhorn volume with a single replica physically on that node's disk -
# mounting it read-only here from a different node would conflict with the
# live webserver's attachment.
apiVersion: batch/v1
kind: CronJob
metadata:
name: paperless-backup
spec:
schedule: "0 3 * * *" # 03:00 daily, low-traffic window
jobTemplate:
spec:
backoffLimit: 2
template:
spec:
restartPolicy: Never
nodeSelector:
kubernetes.io/hostname: talos-cp-3
initContainers:
- name: pg-dump
image: postgres:16-alpine
env:
- name: PGHOST
value: paperless-db-rw
- name: PGDATABASE
value: paperless
- name: PGUSER
valueFrom:
secretKeyRef:
name: paperless-db-app
key: username
- name: PGPASSWORD
valueFrom:
secretKeyRef:
name: paperless-db-app
key: password
command:
- sh
- -c
- pg_dump --format=custom --file=/backup/paperless-db.dump
volumeMounts:
- name: backup
mountPath: /backup
containers:
- name: mc-mirror
image: minio/mc:latest
env:
- name: ACCESS_KEY
valueFrom:
secretKeyRef:
name: paperless-minio-creds
key: ACCESS_KEY
- name: SECRET_KEY
valueFrom:
secretKeyRef:
name: paperless-minio-creds
key: SECRET_KEY
- name: BUCKET
valueFrom:
secretKeyRef:
name: paperless-minio-creds
key: BUCKET
- name: ENDPOINT
valueFrom:
secretKeyRef:
name: paperless-minio-creds
key: ENDPOINT
command:
- /bin/sh
- -c
- |
set -e
mc alias set b "$ENDPOINT" "$ACCESS_KEY" "$SECRET_KEY"
mc cp /backup/paperless-db.dump "b/$BUCKET/db/paperless-db-$(date +%Y%m%d).dump"
mc mirror --overwrite /media "b/$BUCKET/media"
echo "Backup done."
volumeMounts:
- name: backup
mountPath: /backup
- name: media
mountPath: /media
readOnly: true
volumes:
- name: backup
emptyDir: {}
- name: media
persistentVolumeClaim:
claimName: paperless-media
readOnly: true
-22
View File
@@ -1,22 +0,0 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: paperless-config
data:
PAPERLESS_URL: "https://paperless.riotpiao.com"
PAPERLESS_TIME_ZONE: "America/Los_Angeles"
PAPERLESS_OCR_LANGUAGE: "eng"
PAPERLESS_DBHOST: "paperless-db-rw"
PAPERLESS_DBNAME: "paperless"
PAPERLESS_REDIS: "redis://paperless-redis:6379"
# django-allauth generic OIDC provider. The client_id/secret/server_url
# bundle itself lives in the paperless-oidc Secret
# (SOCIALACCOUNT_PROVIDERS_JSON key, composed by authentik-provision.py) -
# env vars can't be split across a ConfigMap + Secret for the same key, so
# this whole value is sourced from the Secret in deployment.yaml instead.
PAPERLESS_APPS: "allauth.socialaccount.providers.openid_connect"
# Authentik already verifies identity via OIDC - a second email-confirmation
# step has no SMTP configured to send it anyway, and paperless-ngx doesn't
# wire up allauth's confirm-email view, so signup 500s with NoReverseMatch
# on 'account_confirm_email' without this.
PAPERLESS_ACCOUNT_EMAIL_VERIFICATION: "none"
-103
View File
@@ -1,103 +0,0 @@
# Single container runs webserver + consumer + scheduler (paperless-ngx's
# stock entrypoint does this internally) - no need to split into separate
# Deployments. replicas: 1 only: paperless-media is ReadWriteOnce, and the
# consumer polling the media dir doesn't benefit from horizontal scaling here.
#
# Pinned to talos-cp-3: paperless-media's disk physically lives there. Longhorn
# RWO volumes can only be attached from one node at a time, and the nightly
# backup-cronjob.yaml also mounts this same PVC (read-only) to mirror it into
# MinIO - pinning both to the same node avoids a cross-node attach conflict,
# and keeps the 3.5Ti read/write path off the network entirely.
apiVersion: apps/v1
kind: Deployment
metadata:
name: paperless
spec:
replicas: 1
strategy:
type: Recreate # ReadWriteOnce media PVC - avoid two pods fighting over it
selector:
matchLabels:
app: paperless
template:
metadata:
labels:
app: paperless
spec:
# Kubernetes injects legacy Docker-links env vars for every Service in
# this namespace (<SVC>_SERVICE_HOST, <SVC>_PORT, ...). The Service here
# is named "paperless", so that becomes PAPERLESS_PORT=tcp://<ip>:8000 -
# paperless-ngx's own entrypoint reads PAPERLESS_PORT for gunicorn's
# bind address, collides, and gunicorn crash-loops on "not a valid port
# number". Disable the injection instead of renaming the Service.
enableServiceLinks: false
nodeSelector:
kubernetes.io/hostname: talos-cp-3
containers:
- name: paperless
image: ghcr.io/paperless-ngx/paperless-ngx:2.20.15
ports:
- containerPort: 8000
envFrom:
- configMapRef:
name: paperless-config
env:
- name: PAPERLESS_DBUSER
valueFrom:
secretKeyRef:
name: paperless-db-app
key: username
- name: PAPERLESS_DBPASS
valueFrom:
secretKeyRef:
name: paperless-db-app
key: password
- name: PAPERLESS_SECRET_KEY
valueFrom:
secretKeyRef:
name: paperless-secrets
key: PAPERLESS_SECRET_KEY
- name: PAPERLESS_ADMIN_USER
valueFrom:
secretKeyRef:
name: paperless-secrets
key: PAPERLESS_ADMIN_USER
- name: PAPERLESS_ADMIN_PASSWORD
valueFrom:
secretKeyRef:
name: paperless-secrets
key: PAPERLESS_ADMIN_PASSWORD
- name: PAPERLESS_SOCIALACCOUNT_PROVIDERS
valueFrom:
secretKeyRef:
name: paperless-oidc
key: SOCIALACCOUNT_PROVIDERS_JSON
resources:
requests: { cpu: "500m", memory: "1Gi" }
limits: { cpu: "2", memory: "4Gi" }
volumeMounts:
- name: media
mountPath: /usr/src/paperless/media
- name: data
mountPath: /usr/src/paperless/data
- name: consume
mountPath: /usr/src/paperless/consume
# Overrides paperless-ngx's own adapter.py in place - settings.py
# hardcodes the import path, so no Django setting changes, just
# the file content underneath it (see adapter-configmap.yaml).
- name: adapter
mountPath: /usr/src/paperless/src/paperless/adapter.py
subPath: adapter.py
readOnly: true
volumes:
- name: media
persistentVolumeClaim:
claimName: paperless-media
- name: data
persistentVolumeClaim:
claimName: paperless-data
- name: consume
emptyDir: {}
- name: adapter
configMap:
name: paperless-adapter
-24
View File
@@ -1,24 +0,0 @@
# Direct nginx ingress to the paperless Service - not routed via the Go
# api-gateway (api.riotpiao.com), which has no WebSocket upgrade support and
# paperless-ngx keeps a long-lived /ws/ connection open for live task status.
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: paperless
annotations:
nginx.ingress.kubernetes.io/proxy-body-size: "0" # large scanned PDF uploads
nginx.ingress.kubernetes.io/proxy-read-timeout: "600"
nginx.ingress.kubernetes.io/proxy-send-timeout: "600"
spec:
ingressClassName: nginx
rules:
- host: paperless.riotpiao.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: paperless
port:
number: 8000
-17
View File
@@ -1,17 +0,0 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
namespace: paperless
resources:
- pvc.yaml
- configmap.yaml
- redis.yaml
- deployment.yaml
- service.yaml
- ingress.yaml
- backup-cronjob.yaml
- adapter-configmap.yaml
- rbac.yaml
# postgres: paperless-db CNPG Cluster, deployed by k8s/infra/databases (wave 2,
# before this app at wave 8) - not duplicated here. Same for the paperless-oidc
# and paperless-minio-creds Secrets, written by PostSync provisioning Jobs in
# k8s/infra/iam and k8s/infra/minio respectively.
-35
View File
@@ -1,35 +0,0 @@
# Two volumes, deliberately separate storage classes:
#
# - media: the actual documents (originals + OCR'd archive PDFs + thumbnails).
# Lives on the cp-3 USB HDD, single replica (see
# k8s/infra/longhorn/longhorn-paperless-storageclass.yaml). Shares the disk
# with Immich's immich-media PVC (k8s/apps/immich/pvc.yaml, 2000Gi) - photo
# libraries grow much faster than scanned documents, so paperless gets the
# smaller 500Gi share.
# - data: the SQLite classification model + search index. Small (low GB),
# frequently rewritten, and disposable (rebuilds from the DB + media on
# next consume) - stays on the default 3-replica pool instead of the
# single-disk HDD.
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: paperless-media
spec:
accessModes:
- ReadWriteOnce
storageClassName: longhorn-paperless-media
resources:
requests:
storage: 500Gi
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: paperless-data
spec:
accessModes:
- ReadWriteOnce
storageClassName: longhorn
resources:
requests:
storage: 5Gi
-35
View File
@@ -1,35 +0,0 @@
# Scoped operator access for paperless-admins: restart/config-edit rights on
# just this service's own resources, nothing CNPG-managed (paperless-db-*)
# or provisioning-managed (paperless-oidc, paperless-minio-creds). Inert
# until kube-apiserver's OIDC wiring lands (--oidc-groups-claim=groups,
# --oidc-groups-prefix=oidc:) - subject name below assumes that prefix.
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: paperless-operator
rules:
- apiGroups: ["apps"]
resources: ["deployments"]
resourceNames: ["paperless"]
verbs: ["get", "list", "watch", "update", "patch"]
- apiGroups: [""]
resources: ["configmaps"]
resourceNames: ["paperless-config"]
verbs: ["get", "list", "watch", "update", "patch"]
- apiGroups: [""]
resources: ["secrets"]
resourceNames: ["paperless-secrets"]
verbs: ["get", "list", "watch", "update", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: paperless-admins-binding
subjects:
- kind: Group
name: "oidc:paperless-admins"
apiGroup: rbac.authorization.k8s.io
roleRef:
kind: Role
name: paperless-operator
apiGroup: rbac.authorization.k8s.io
-37
View File
@@ -1,37 +0,0 @@
# Task queue broker + websocket channel layer for paperless-ngx. No PVC:
# queued/scheduled task state is disposable - a lost queue on restart just
# means re-triggering consumption, not data loss (documents themselves live
# on paperless-media).
apiVersion: apps/v1
kind: Deployment
metadata:
name: paperless-redis
spec:
replicas: 1
selector:
matchLabels:
app: paperless-redis
template:
metadata:
labels:
app: paperless-redis
spec:
containers:
- name: redis
image: redis:7-alpine
ports:
- containerPort: 6379
resources:
requests: { cpu: "50m", memory: "64Mi" }
limits: { cpu: "250m", memory: "256Mi" }
---
apiVersion: v1
kind: Service
metadata:
name: paperless-redis
spec:
selector:
app: paperless-redis
ports:
- port: 6379
targetPort: 6379
-10
View File
@@ -1,10 +0,0 @@
apiVersion: v1
kind: Service
metadata:
name: paperless
spec:
selector:
app: paperless
ports:
- port: 8000
targetPort: 8000
@@ -1,144 +0,0 @@
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
name: secretrotations.homelab.riotpiao.com
spec:
group: homelab.riotpiao.com
names:
kind: SecretRotation
plural: secretrotations
scope: Namespaced
versions:
- name: v1
served: true
storage: true
schema:
openAPIV3Schema:
type: object
properties:
metadata:
type: object
spec:
type: object
required:
- provider
- rotationInterval
properties:
# External system: authentik | forgejo | minio | vault
provider:
type: string
enum: [authentik, forgejo, minio, vault]
# How often to rotate (hours)
rotationInterval:
type: integer
minimum: 24
# Application ID in external system
appId:
type: string
# k8s Secret to update (name, namespace, key)
secretRef:
type: object
required: [name, namespace]
properties:
name:
type: string
namespace:
type: string
key:
type: string
description: "Secret key to update (e.g., MINIO_IDENTITY_OPENID_CLIENT_SECRET)"
# Path to git file that holds the secret (for .enc.yaml files)
gitPath:
type: string
description: "Path in homelab repo to .enc.yaml file"
# Ansible template values to substitute
templateValues:
type: object
additionalProperties:
type: string
status:
type: object
properties:
lastRotationTime:
type: string
format: date-time
nextRotationTime:
type: string
format: date-time
lastRotationStatus:
type: string
enum: [Success, Failed, Pending]
lastRotationError:
type: string
lastCommitHash:
type: string
---
# Example usage:
apiVersion: homelab.riotpiao.com/v1
kind: SecretRotation
metadata:
name: minio-oidc
namespace: secret-rotation
spec:
provider: authentik
rotationInterval: 2160 # 90 days in hours
appId: minio
secretRef:
name: minio-oidc
namespace: storage
key: MINIO_IDENTITY_OPENID_CLIENT_SECRET
gitPath: k8s/argocd/secrets/minio-oidc.enc.yaml
---
apiVersion: homelab.riotpiao.com/v1
kind: SecretRotation
metadata:
name: portfolio-agent-oidc
namespace: secret-rotation
spec:
provider: authentik
rotationInterval: 2160
appId: portfolio-agent
secretRef:
name: portfolio-agent-oidc
namespace: portfolio
key: CLIENT_SECRET
gitPath: k8s/argocd/secrets/portfolio-agent-oidc.enc.yaml
---
apiVersion: homelab.riotpiao.com/v1
kind: SecretRotation
metadata:
name: forgejo-registry-token
namespace: secret-rotation
spec:
provider: forgejo
rotationInterval: 2160
appId: rock/riotpiao.com
secretRef:
name: forgejo-registry-secret
namespace: kube-system
key: REGISTRY_TOKEN
gitPath: k8s/argocd/secrets/forgejo-registry-secret.enc.yaml
---
apiVersion: homelab.riotpiao.com/v1
kind: SecretRotation
metadata:
name: minio-root-credentials
namespace: secret-rotation
spec:
provider: minio
rotationInterval: 4320 # 180 days in hours
appId: root
secretRef:
name: minio-creds
namespace: storage
gitPath: k8s/argocd/secrets/minio-secrets.enc.yaml
@@ -1,92 +0,0 @@
apiVersion: apps/v1
kind: Deployment
metadata:
name: secret-rotation-controller
namespace: secret-rotation
spec:
replicas: 1
selector:
matchLabels:
app: secret-rotation-controller
template:
metadata:
labels:
app: secret-rotation-controller
spec:
serviceAccountName: secret-rotation-controller
containers:
- name: controller
image: secret-rotation-controller:latest
imagePullPolicy: IfNotPresent
env:
# SOPS reads age key from this file
- name: SOPS_AGE_KEY_FILE
value: /etc/sops/age/private-key.txt
# Vault auth (token in projected volume)
- name: VAULT_ADDR
value: http://vault.vault.svc.cluster.local:8200
- name: VAULT_TOKEN_FILE
value: /var/run/secrets/vault/token
# Authentik
- name: AUTHENTIK_URL
value: http://authentik-server.iam.svc.cluster.local
- name: AUTHENTIK_BOOTSTRAP_TOKEN
valueFrom:
secretKeyRef:
name: authentik-bootstrap
key: token
# Git
- name: GIT_REPO
value: https://forgejo.riotpiao.com/rock/homelab.git
- name: GIT_AUTHOR_EMAIL
value: [email protected]
- name: GIT_AUTHOR_NAME
value: Secret Rotation Controller
- name: FORGEJO_TOKEN
valueFrom:
secretKeyRef:
name: forgejo-registry-secret
key: REGISTRY_TOKEN
volumeMounts:
# Age key from ExternalSecret (synced from Vault)
- name: age-key
mountPath: /etc/sops/age
readOnly: true
# Vault auth token (projected)
- name: vault-token
mountPath: /var/run/secrets/vault
readOnly: true
# Temp working dir
- name: tmp
mountPath: /tmp
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: 500m
memory: 512Mi
volumes:
- name: age-key
secret:
secretName: sops-age-key
defaultMode: 0400
- name: vault-token
projected:
sources:
- serviceAccountToken:
path: token
audience: vault
expirationSeconds: 3600
- name: tmp
emptyDir: {}
@@ -1,15 +0,0 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
namespace: secret-rotation
resources:
- rbac.yaml
- crd.yaml
- external-secret.yaml
- deployment.yaml
commonLabels:
app.kubernetes.io/name: secret-rotation-controller
app.kubernetes.io/component: automation
managed-by: argocd
@@ -1,53 +0,0 @@
apiVersion: v1
kind: ServiceAccount
metadata:
name: secret-rotation-controller
namespace: secret-rotation
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: secret-rotation-controller
rules:
# Read SecretRotation CRDs
- apiGroups: ["homelab.riotpiao.com"]
resources: ["secretrotations"]
verbs: ["get", "list", "watch"]
# Update status
- apiGroups: ["homelab.riotpiao.com"]
resources: ["secretrotations/status"]
verbs: ["get", "patch", "update"]
# Read k8s secrets that will be rotated
- apiGroups: [""]
resources: ["secrets"]
verbs: ["get", "list"]
# For recording events
- apiGroups: [""]
resources: ["events"]
verbs: ["create", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: secret-rotation-controller
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: secret-rotation-controller
subjects:
- kind: ServiceAccount
name: secret-rotation-controller
namespace: secret-rotation
---
apiVersion: v1
kind: Namespace
metadata:
name: secret-rotation
labels:
kubernetes.io/metadata.name: secret-rotation
-18
View File
@@ -1,18 +0,0 @@
apiVersion: kmsvc.io/v1
kind: TemporalWorker
metadata:
name: worker-production
namespace: temporal
spec:
namespace: production
taskQueue: worker-production
concurrency: 10
workflowTypes:
- HelloWorldWorkflow
- GreeterWorkflow
- ProcessOrderWorkflow
activityTypes:
- GreetActivity
- ValidateOrderActivity
- ProcessPaymentActivity
- NotifyCustomerActivity
-25
View File
@@ -1,25 +0,0 @@
# Wave -1 — AppProject definitions (must sync before any Application that references them).
# Syncs k8s/argocd/projects/ which was previously applied by hand.
# Enabled by Stage 1 (A2).
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: argocd-projects
namespace: argocd
finalizers:
- resources-finalizer.argocd.argoproj.io
annotations:
argocd.argoproj.io/sync-wave: "-1"
spec:
project: homelab
revisionHistoryLimit: 3
syncPolicy:
automated:
prune: true
selfHeal: true
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/argocd/projects
destination:
server: https://kubernetes.default.svc
+1 -1
View File
@@ -17,7 +17,7 @@ spec:
syncOptions:
- CreateNamespace=true
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
# ksops decrypts every *.enc.yaml here at kustomize-build time (repo-server
# runs `kustomize build --enable-alpha-plugins --enable-exec`). Replaces the
+3 -75
View File
@@ -21,7 +21,7 @@ spec:
helm:
valueFiles:
- $values/k8s/bootstrap/cert-manager/cert-manager-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
@@ -93,7 +93,7 @@ spec:
project: homelab
revisionHistoryLimit: 3
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
# A real kustomization.yaml (resources: the 3 issuer/CA files) renders these
# deterministically. The previous directory.include with bare filenames
@@ -127,7 +127,7 @@ spec:
project: homelab
revisionHistoryLimit: 3
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/bootstrap/ingress
destination:
@@ -138,75 +138,3 @@ spec:
selfHeal: true
syncOptions:
- CreateNamespace=true
---
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: cluster-maintenance
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "0"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/infra/cluster-maintenance
destination:
server: https://kubernetes.default.svc
namespace: kube-system
syncPolicy:
automated:
prune: true
selfHeal: true
---
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: kyverno
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "0"
spec:
project: homelab
source:
repoURL: https://kyverno.github.io/kyverno/
chart: kyverno
targetRevision: "1.14.0"
helm:
valueFiles:
- $values/k8s/bootstrap/kyverno/kyverno-values.yaml
sources:
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
ref: values
destination:
server: https://kubernetes.default.svc
namespace: kyverno
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
---
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: kyverno-policies
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "0"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/bootstrap/kyverno
destination:
server: https://kubernetes.default.svc
namespace: kyverno
syncPolicy:
automated:
prune: true
selfHeal: true
-33
View File
@@ -1,33 +0,0 @@
# ArgoCD Image Updater - auto-updates Application images from registry
# Watches forgejo.riotpiao.com for new image tags and updates Applications
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: argocd-image-updater
namespace: argocd
finalizers:
- resources-finalizer.argocd.argoproj.io
annotations:
argocd.argoproj.io/sync-wave: "1"
spec:
project: homelab
revisionHistoryLimit: 3
sources:
- repoURL: https://argoproj.github.io/argo-helm
chart: argocd-image-updater
targetRevision: "0.11.2"
helm:
valueFiles:
- $values/k8s/infra/argocd-image-updater/values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
ref: values
destination:
server: https://kubernetes.default.svc
namespace: argocd
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=false
-32
View File
@@ -1,32 +0,0 @@
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: secret-rotation
namespace: argocd
labels:
app.kubernetes.io/name: secret-rotation
spec:
project: homelab
sources:
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
path: k8s/apps/secret-rotation-controller
targetRevision: main
destination:
server: https://kubernetes.default.svc
namespace: secret-rotation
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
- RespectIgnoreDifferences=true
retry:
limit: 5
backoff:
duration: 5s
factor: 2
maxDuration: 3m
+7 -33
View File
@@ -17,7 +17,7 @@ spec:
helm:
valueFiles:
- $values/k8s/infra/minio/minio-operator-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
@@ -41,7 +41,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/infra/minio
destination:
@@ -66,7 +66,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/infra/longhorn
destination:
@@ -102,7 +102,7 @@ spec:
skipCrds: true
valueFiles:
- $values/k8s/infra/monitoring/prometheus-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
@@ -152,7 +152,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/infra/monitoring/crds
destination:
@@ -183,7 +183,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/infra/monitoring
destination:
@@ -213,7 +213,7 @@ spec:
helm:
valueFiles:
- $values/k8s/infra/monitoring/blackbox-exporter-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
@@ -223,29 +223,3 @@ spec:
automated:
prune: true
selfHeal: true
---
# Distributed tracing: Tempo + OpenTelemetry Collector.
# Receives traces from instrumented services, stores in local volume (72h retention).
# Grafana datasource auto-configured, service graph + latency dashboards included.
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: tracing
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "1"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/infra/tracing
destination:
server: https://kubernetes.default.svc
namespace: tracing
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
+3 -3
View File
@@ -19,7 +19,7 @@ spec:
helm:
valueFiles:
- $values/k8s/infra/logging/loki-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
@@ -53,7 +53,7 @@ spec:
helm:
valueFiles:
- $values/k8s/infra/logging/grafana-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
@@ -87,7 +87,7 @@ spec:
helm:
valueFiles:
- $values/k8s/infra/logging/promtail-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
+6 -119
View File
@@ -17,7 +17,7 @@ spec:
helm:
valueFiles:
- $values/k8s/infra/iam/vault-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
@@ -46,7 +46,7 @@ spec:
helm:
valueFiles:
- $values/k8s/infra/iam/authentik-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
@@ -68,7 +68,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/infra/iam
destination:
@@ -79,83 +79,18 @@ spec:
prune: true
selfHeal: true
---
# Forgejo itself. Was a bootstrap Helm release (phase 3) until it was brought
# under Argo, because values changes there were inert — a proxy-body-size fix
# sat committed while the live Ingress kept nginx's 1m default and rejected
# every OCI push with 413.
#
# Wave 3: after databases (wave 2) — Forgejo needs CNPG and Redis up first.
#
# Retiring the Helm release: Argo adopts the existing objects on first sync.
# Delete the release secrets afterwards so helm stops claiming ownership:
# kubectl -n cicd delete secret -l owner=helm,name=forgejo
#
# automated sync is deliberately absent. This chart owns the Forgejo PVC and
# the git forge itself; the first sync is manual so its diff can be read before
# anything is applied. Turn on automated+selfHeal once that diff is clean.
# Forgejo runner (local chart). Forgejo itself is Phase 0 (bootstrap).
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: forgejo
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "3"
spec:
project: homelab
sources:
- repoURL: https://dl.gitea.com/charts/
chart: gitea
targetRevision: 12.7.0
helm:
valueFiles:
- $values/k8s/bootstrap/phase3-forgejo/forgejo-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
ref: values
destination:
server: https://kubernetes.default.svc
namespace: cicd
# Reloader injects a STAKATER_* env var carrying a hash of the config Secret,
# so the pod rolls when that Secret changes. The chart does not render it, so
# Argo would strip it on every sync — and with selfHeal on, Argo and Reloader
# would fight over the field and Recreate the forge each round.
ignoreDifferences:
- group: apps
kind: Deployment
name: forgejo-gitea
jqPathExpressions:
- '.spec.template.spec.containers[].env[] | select(.name | startswith("STAKATER_"))'
syncPolicy:
syncOptions:
# Adopt the objects the bootstrap Helm release already created rather
# than failing on "already exists".
- ServerSideApply=true
---
# Forgejo runners (local chart, one instance per language), replacing the
# single generic "docker"-labeled runner. Each instance is a full standalone
# Deployment with its own dind sidecar, own PVCs (registration + layer
# cache) and own registered label -- there is no shared generic runner
# anymore, so each instance also builds and pushes images for the repos it
# serves (the chart's ConfigMap/NetworkPolicy fixes for that -- valid_volumes,
# network: host, egress to ingress-nginx -- apply identically to all three).
#
# `values.yaml` is the chart's default and doubles as the golang instance's
# config; node and rust layer a small values-<lang>.yaml override on top for
# just runner.name/runner.labels. All three share one runner-token Secret
# (Forgejo registration tokens are reusable across multiple runners, unlike
# GitHub's one-time tokens) -- if that assumption is ever wrong, registration
# will fail loudly in the register initContainer's logs, not silently.
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: forgejo-runner-golang
name: forgejo-runner
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "3"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/infra/forgejo-runner
destination:
@@ -165,51 +100,3 @@ spec:
automated:
prune: true
selfHeal: true
---
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: forgejo-runner-node
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "3"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/infra/forgejo-runner
helm:
valueFiles:
- values-node.yaml
destination:
server: https://kubernetes.default.svc
namespace: cicd
syncPolicy:
automated:
prune: true
selfHeal: true
---
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: forgejo-runner-rust
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "3"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/infra/forgejo-runner
helm:
valueFiles:
- values-rust.yaml
destination:
server: https://kubernetes.default.svc
namespace: cicd
syncPolicy:
automated:
prune: true
selfHeal: true
+1 -1
View File
@@ -14,7 +14,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/infra/databases
destination:
-20
View File
@@ -1,20 +0,0 @@
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: memory-queues
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "7"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/apps/messaging/memory-queues
destination:
server: https://kubernetes.default.svc
namespace: sqs
syncPolicy:
automated:
prune: true
selfHeal: true
+46 -6
View File
@@ -1,6 +1,6 @@
# Wave 5 — Kafka (Strimzi operator + cluster CR), Redis infrastructure.
# Strimzi/Redis are public Helm charts; kafka-cluster is a local chart.
# queue-crd and management-service are managed by kmsvc-root (kmsvc-manage.git).
# Wave 5 — Kafka (Strimzi operator + cluster CR), Redis, and the SQS-like
# queue services. Strimzi/Redis are public Helm charts; kafka-cluster/queue-crd/
# management-service are local charts (rendered from their own Chart.yaml).
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
@@ -68,7 +68,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/apps/messaging/kafka-cluster
destination:
@@ -78,5 +78,45 @@ spec:
automated:
prune: true
selfHeal: true
# queue-crd and management-service moved to kmsvc-manage.git repo
# Managed by kmsvc-root Application
---
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: queue-crd
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "6"
spec:
project: homelab
source:
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/apps/messaging/queue-crd
destination:
server: https://kubernetes.default.svc
namespace: sqs
syncPolicy:
automated:
prune: true
selfHeal: true
---
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: management-service
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "7"
spec:
project: homelab
source:
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/apps/messaging/management-service
destination:
server: https://kubernetes.default.svc
namespace: sqs
syncPolicy:
automated:
prune: true
selfHeal: true
+29 -39
View File
@@ -1,54 +1,40 @@
# Wave 7 — api-gw, the cluster's API gateway (namespace `api`).
# Wave 7 — Kong, the cluster's internal API gateway (namespace `api`).
#
# Replaces Kong OSS 3.4.1, removed 2026-08-19. Kong existed to route
# `api.riotpiao.com`, but Kong OSS cannot dispatch on a request body, so the
# LLM surface had to be expressed as one path per model
# (`/v1/reasoning/chat/completions`, `/v1/ornith/...`, `/v1/qwen/...`) with a
# `request-transformer` plugin forcing the body's `model` field on each. The Go
# gateway reads the body and picks the upstream, so a single canonical
# `POST /v1/chat/completions` covers every model. See
# docs/adr/ADR-0001-retire-kong-for-go-gateway.md in the frontend repo.
# Sits between nginx and the backend services: nginx owns the edge and TLS,
# Kong owns routing policy, auth and rate limiting. Wave 7 puts it after the
# data/messaging tiers it fronts and before the wave-8 applications that
# publish routes into it.
#
# UPDATED 2026-08-22: Tracks main branch of homelab-frontend (auto-syncs on each push).
# Image built on every main commit with tag <commit-sha>.
# ArgoCD auto-pulls the latest image (live reconciliation ~3min).
# DB-less: routing config comes from Kubernetes objects (Ingress with
# `ingressClassName: kong`, plus KongPlugin/KongConsumer CRDs), so git remains
# the source of truth and there are no migration Jobs on upgrade.
#
# Two sources:
# 1. rock/homelab-frontend on the in-cluster Forgejo (prod branch) — the gateway's own
# kustomization (Deployment, Service, ConfigMap, RBAC, NetworkPolicy). It
# sets `namespace: api` itself, so no transformer is needed here. The
# Forgejo host must stay listed in the `homelab` AppProject sourceRepos or
# this Application is rejected with "is not permitted in project".
# 2. k8s/apps/api in this repo — the nginx edge Ingress for
# api.riotpiao.com, inherited from the retired `kong` Application. It
# cannot move to k8s/bootstrap/ingress/ingress.yaml because that syncs in
# wave 1, before namespace `api` exists.
#
# No resources-finalizer: deleting this Application leaves the workload running
# rather than cascading the delete.
# CRDs ship in the chart's crds/ directory; ArgoCD applies those by default
# (helm.skipCrds is left false).
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: api-gw
name: kong
namespace: argocd
labels:
app.kubernetes.io/name: api-gateway
app.kubernetes.io/component: gateway
annotations:
argocd.argoproj.io/sync-wave: "7"
# ArgoCD Image Updater - auto-update on new image push
argocd-image-updater.argoproj.io/image-list: gw=forgejo.riotpiao.com/rock/api-gateway
argocd-image-updater.argoproj.io/gw.update-strategy: newest-build
argocd-image-updater.argoproj.io/gw.allow-tags: regexp:^[0-9a-f]{7}$
argocd-image-updater.argoproj.io/write-back-method: argocd
spec:
project: homelab
revisionHistoryLimit: 3
sources:
- repoURL: https://forgejo.riotpiao.com/rock/homelab-frontend.git
- repoURL: https://charts.konghq.com
chart: kong
targetRevision: "3.4.1"
helm:
valueFiles:
- $values/k8s/apps/api/kong-values.yaml
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
ref: values
# The nginx Ingress for api.riotpiao.com. Kept in this Application rather
# than the central k8s/bootstrap/ingress/ingress.yaml because that one syncs
# in wave 1, before namespace `api` exists.
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/apps/api
destination:
@@ -60,9 +46,13 @@ spec:
selfHeal: true
syncOptions:
- CreateNamespace=true
# The chart's CRDs exceed the annotation size limit that client-side
# apply relies on; server-side apply avoids the
# "metadata.annotations: Too long" failure CRDs commonly hit.
- ServerSideApply=true
retry:
limit: 5
limit: 3
backoff:
duration: 5s
duration: 10s
factor: 2
maxDuration: 3m
+3 -3
View File
@@ -1,7 +1,7 @@
# Wave 6 — the model servers behind api.riotpiao.com (namespace `llm-serving`).
#
# Syncs before wave 7 (api-gw), so the predictor Services exist before the
# gateway that routes to them. KServe itself is part of the substrate; this Application
# Syncs before wave 7 (Kong), so the predictor Services exist before the routes
# that point at them. KServe itself is part of the substrate; this Application
# owns only the InferenceServices.
#
# Adopted from live state on 2026-08-15. These five had been `kubectl apply`-ed
@@ -20,7 +20,7 @@ spec:
project: homelab
revisionHistoryLimit: 3
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/apps/llm-serving
destination:
+7 -127
View File
@@ -19,12 +19,9 @@ spec:
helm:
valueFiles:
- $values/k8s/apps/temporal/temporal-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/apps/temporal
destination:
server: https://kubernetes.default.svc
namespace: temporal
@@ -51,7 +48,7 @@ spec:
helm:
valueFiles:
- $values/k8s/apps/portainer/portainer-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
@@ -74,7 +71,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/apps/cloudflared
destination:
@@ -97,7 +94,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/apps/agent-pod
destination:
@@ -130,7 +127,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/apps/sms
destination:
@@ -141,67 +138,6 @@ spec:
prune: true
selfHeal: true
---
# Document management. Raw manifests (no Helm): postgres is the dedicated
# paperless-db CNPG cluster in k8s/infra/databases (wave 2), redis is
# in-cluster only (no PVC), media lives on the cp-3 USB HDD (see
# k8s/infra/longhorn/longhorn-paperless-storageclass.yaml). OIDC via
# Authentik provisioned by k8s/infra/iam's PostSync job; MinIO backup bucket
# creds provisioned by k8s/infra/minio's PostSync job.
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: paperless
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "8"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/apps/paperless
destination:
server: https://kubernetes.default.svc
namespace: paperless
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
---
# Photo/video backup. Self-contained (unlike paperless, its CNPG Postgres
# lives here too, not in k8s/infra/databases) - CreateNamespace=true creates
# the namespace before any manifest in this Application applies, including
# the Cluster CR, so no separate wave-2 pre-creation step is needed. Postgres
# is pg18 (not this repo's usual 16.2) because CNPG's official pgvector
# extension image only publishes pg18 builds - see k8s/apps/immich/db.yaml.
# media PVC shares the cp-3 HDD 2TB/2TB with paperless-media. OIDC via
# Authentik provisioned by k8s/infra/iam's PostSync job (immich entry in
# SERVICES + immich_role scope mapping for admin-via-claim).
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: immich
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "8"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/apps/immich
destination:
server: https://kubernetes.default.svc
namespace: immich
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
---
# Consolidated: homarr + homarr-patches → homarr
# Helm chart + values + PostSync hook patch (fix-probes-job.yaml)
apiVersion: argoproj.io/v1alpha1
@@ -220,10 +156,10 @@ spec:
helm:
valueFiles:
- $values/k8s/apps/homarr/homarr-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/apps/homarr # PostSync hook: fix-probes-job.yaml
destination:
@@ -235,59 +171,3 @@ spec:
selfHeal: true
syncOptions:
- CreateNamespace=true
---
# Portfolio site at riotpiao.com - static Next.js site from rock/riotpiao.com repo.
# Points directly to infra/portfolio/base (bypassing repo's own argocd-apps.yaml
# which has wrong URLs). Image built by Forgejo Actions on rock/portfolio repo.
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: portfolio
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "8"
# ArgoCD Image Updater - auto-update on new image push
argocd-image-updater.argoproj.io/image-list: app=forgejo.riotpiao.com/rock/portfolio
argocd-image-updater.argoproj.io/app.update-strategy: newest-build
argocd-image-updater.argoproj.io/app.allow-tags: regexp:^[0-9a-f]{7}$
argocd-image-updater.argoproj.io/write-back-method: argocd
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/riotpiao.com.git
targetRevision: main
path: infra/portfolio/base
destination:
server: https://kubernetes.default.svc
namespace: portfolio
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
---
# Wave 9 - per-service scoped RBAC (Role/RoleBinding), deliberately last so
# every target namespace above already exists. Inert until kube-apiserver
# gets --oidc-groups-claim=groups wired up (separate, not-yet-applied
# terraform/talosctl change) - these grant nothing until then.
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: rbac
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "9"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/infra/rbac
destination:
server: https://kubernetes.default.svc
namespace: default
syncPolicy:
automated:
prune: true
selfHeal: true
-28
View File
@@ -1,28 +0,0 @@
# kmsvc-manage bootstrap — manages itself and its supporting services
# (Strimzi/Kafka, Redis, queue-operator, message-plane server) from
# the kmsvc-manage repo's own k8s/argocd/ structure on the main branch.
# Image built on every main commit, auto-deployed to sqs namespace.
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: kmsvc-root
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "6"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/kmsvc-manage.git
targetRevision: main
path: k8s/argocd/apps
directory:
recurse: false
destination:
server: https://kubernetes.default.svc
namespace: sqs
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
-39
View File
@@ -1,39 +0,0 @@
# Poimen project collection — manages poimen-memory, poimen-workflows, and poiman
# Each repo tracks its own main branch (no prod branch). Poiman is the primary
# orchestrator with k8s/argocd/ containing the AppProject and deployment structure.
#
# CI: All three repos trigger on main branch pushes (no image builds yet).
# Future: Add build workflows for poiman once container runtime needs are clear.
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: poimen-root
namespace: argocd
labels:
app.kubernetes.io/name: poimen
app.kubernetes.io/component: orchestrator
annotations:
argocd.argoproj.io/sync-wave: "7"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/poimen.git
targetRevision: main
path: k8s/argocd
directory:
recurse: false
destination:
server: https://kubernetes.default.svc
namespace: poimen
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
retry:
limit: 5
backoff:
duration: 5s
factor: 2
maxDuration: 3m
-14
View File
@@ -12,18 +12,6 @@ spec:
description: Homelab GitOps — single-repo, in-cluster destinations only
sourceRepos:
- https://github.com/Riotpiaole/riotpiao.homelab.com.git
# Poimen services (GitHub)
- https://github.com/Riotpiaole/Poimen-memory.git
- https://github.com/Riotpiaole/Poimen-workflows.git
- https://github.com/Riotpiaole/poimen*.git
# In-cluster Forgejo repos — explicit allowlist (no wildcard)
- https://forgejo.riotpiao.com/rock/homelab.git
- https://forgejo.riotpiao.com/rock/homelab-frontend.git
- https://forgejo.riotpiao.com/rock/kmsvc-manage.git
- https://forgejo.riotpiao.com/rock/poimen.git
- https://forgejo.riotpiao.com/rock/poimen-memory.git
- https://forgejo.riotpiao.com/rock/poimen-workflows.git
- https://forgejo.riotpiao.com/rock/riotpiao.com.git
# Public Helm chart repos referenced by k8s/argocd/apps/* and bootstrap/*
- https://cloudnative-pg.github.io/charts
- https://dl.gitea.com/charts/
@@ -42,8 +30,6 @@ spec:
- https://charts.jetstack.io
- https://kubernetes.github.io/ingress-nginx
- https://stakater.github.io/stakater-charts
# ArgoCD ecosystem charts
- https://argoproj.github.io/argo-helm
destinations:
- server: https://kubernetes.default.svc
namespace: "*"
-10
View File
@@ -1,10 +0,0 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
metadata:
name: argocd-projects
# AppProject definitions for ArgoCD. Synced by wave -1 Application
# (k8s/argocd/apps/-1-projects.yaml) so they exist before any Application
# references them. Enabled by Stage 1 (A2).
resources:
- homelab-project.yaml
+1 -2
View File
@@ -12,7 +12,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/argocd/apps
directory:
@@ -26,4 +26,3 @@ spec:
selfHeal: true
syncOptions:
- CreateNamespace=true
- ServerSideApply=true
+13 -13
View File
@@ -1,23 +1,23 @@
apiVersion: ENC[AES256_GCM,data:C3U=,iv:J6yvL9HYwzrR4AidMrxmTQZAA1AqtAO/nn9AQnS40JY=,tag:WkqaErH6Xfrpd68+4QrfrQ==,type:str]
kind: ENC[AES256_GCM,data:BJsIWY38,iv:/8AGCpKvtaKoi+iuQNkJbKCSo/jSKi6WRs0f1tJj6d0=,tag:6/Rja8mrZu25kSqg0IzjYQ==,type:str]
apiVersion: ENC[AES256_GCM,data:zdc=,iv:VvjvrS5PVNAMIaOE0LaWU+tHcUIYVQDnCANQz6myktY=,tag:xGyWDhRCwwiNny7hPllf5g==,type:str]
kind: ENC[AES256_GCM,data:2bf5Zfy8,iv:5Oz423GzUWmgdaaZHbrtedwRHIAPIuLh4iDMieLL05s=,tag:GNW1ROjlzGvO/4tIOSuH3Q==,type:str]
metadata:
name: ENC[AES256_GCM,data:lTfngGBtsNA+,iv:shwcjVXeWhFRE+IMYlW6ffPyY4JqVzw5yySqcfyU+4I=,tag:qMmCemQcBjLwubaTnnKqTg==,type:str]
namespace: ENC[AES256_GCM,data:/VJI5GUu+jSX,iv:XCgFTLvytlYl8K09JyhGSsmVaTVCyztY5Xvw4TSkfEg=,tag:3eWHcCiedFHLvcKC8WVnUg==,type:str]
type: ENC[AES256_GCM,data:aibi62c0,iv:MJQRYJ27pTgtaRUKoJI2nb1qKZP47c4Ma+PvjIrCiE0=,tag:jB6mYfAPSnWnZUnY+rC+zQ==,type:str]
name: ENC[AES256_GCM,data:7RO0Qxkc+/sA,iv:wKe9A8d7QJSx/6rlEY5H6lU8V24TJqr5IpXQCBc8QgM=,tag:iF6z1MNEXG3pCz2cQ18gLg==,type:str]
namespace: ENC[AES256_GCM,data:RcduxHHLtjH9,iv:opOVx1lL2ltDqQgsleN7NdMAq0TyFr/YQO3FFsHh5AA=,tag:fIwO31apqIatRRzBvamw6g==,type:str]
type: ENC[AES256_GCM,data:C+JjyZt5,iv:xs49Lz6zRzcf3spiPzdUKTm2HZ+VgFahN6wjIe81JI4=,tag:VllJ54MqU+klA8xAIebjrA==,type:str]
stringData:
models.json: ENC[AES256_GCM,data:SZNtmDYpM+ivOATbvUcOGylj5i7RIu6sps3tp63jQcPwrEjM9bNVCcIEdfC8owq3JU2yT3mUMdC5c/ViNuCMhIilLSGf35Da/uTnOTe4e3URNQ2r1LXTBYr5R4ESt6uZrI7llzgc68C7j+k1JEldbPeoyAef0sHHSIGdbQHaVzd/+j6RSnmDcpwNw5fopyeBkZivdnKgY1vW0bz/IEgMUHRpsIeZkENSJvp63pUwsmv7NZ0qaFWk2rgzZiZszIBFYP8K/AcDAaTid8T7F6Ro9A3ClRdHW9K8Zary2TSXMrZJ+Mo9Rcqzcrw5LBiCRCn+lSmILovZfohPHUN6atKbB5YAewo7XBVlJJa5Wl+V7sq3FHZH7+rbLSFQEPovxt8z8SW2k1W/ifz+0HYWLRrtJLADZ3iV6eDKCrkSnDiCvi9pAczwD4jrbQYDdnVb9Djgp+8qhUyjEZT3C7ONs/ZaTAawX434KqsUN4O5LSEKgNlITnvHgmoxMd6eR1Xc48Gb3CBmI9ChZNpXvtgzZX0t9EuKCR0HBoBg9IPz6vADdbEVRGNLVVT1ucDZU7Sp+vMXnaM9ZIw3jDgWnmX+L4sBBpe2H2mC71muYVpF0swHt7+H0o7Nc7vxNsNIqnqWJjTQt/w4joENvbKB4dBs5lNRCeXuWuwyh6/f+MZbz4ZfqtPfvoEsWqiXglGs0eQQttDlXS36R3lsF1brFKbHdO4j7rS7YEeZdKgJIALqp3I+w/v9hAyLh+bu+R6tqXz3QoKWQkDl10Jjip0I/GIkkw9W8ZDBgsy9YcVf9J0ZOa+wm0GZXhBYWyfrE55FiMkCL1IYP2GABb5HdAbMyccIB1dkx9/TAGWeYxsMPmjCvBghf3LawZNa3nDNMnfH7+MxQFfvlHlGJBWbgO7M4V/rOLQu0RhoMIplb1ZyucHSVMWDNt71kteKR7Fme3VdFES+nIoDi8usfUYSQraW99XMB9IsIkPMCz9nFAUoNrXhd/ADkyRmaXe+gb8UohY++7zBy6YllTICHzZFVvucU1YniZLR9i5lVBv16Gcpe0PxnLuCGFwQi+RdgIPglyXTEFh5Woo7ahrr0H/9knrG7p2cqpfJs777ZSWexQL7ncN40kk73fZyAQZFbYw6vhkRGdTHb5Dq2hrLC4FSPev33lWN9v1V4uMo/wa0CGXkxbjTcHeb2NkzfXWFM0+OlqSYTev8fnbZoBGvh/N/UcmdadAG20wPoD4O7kRaZMqZe3vf5s5AsirZtKGMB+FlpBw8f3BQ8GGV9X9bBcFgq42hhcXzGrelqdCwPKFqPpz5ItEasDAdU48dHBFnAn3OWYiKrDn4/uc0XURcMKg5JR0cwJEaAWd6ZPvPy58qYQwWoFA+MVqk1/fql2GjUxW10hTPNMePLq70FYEl5pRqcWAzS2v/9LSM53E/UD/tlyjm61llqPhirDt7QhYmnEawJeQxFzzrgizAdx4jp4SGCpa3envp1oPySJdEyTJkedf/zYxXHLKZ/sc9uuhM/Bl103SeXeuAQrXrFE/kJazdUBgEstOPiDiLSWq9FjMxgMIr9SMcChYCFm7U4irsjWCdboVWKb6iMZH96AFnm2AwP+QvXUSN6dNoRAGY9E7yeiebuAnvCMGQIZ4rxYJ+xsQvRDpIwZZhs1oDAKsvMaozp5o8zZXwU+UCoRfiHVQ+LmAKhDWW0PfLhMVRSsBA9tHDnigwwRNth+A+117IBwnOgpMf/vS2xOTPszAyso24+4np4etTRMWBZLai/sScnkWflGlEwTD2lb38IeCVfTH4Vp1wPfViI3bRGvpj8je6aBg5buRaUO/M2NPXdK9eroM3OQewEtv1QTtuGy7ihsXukeoll8GvOor+1m8qhLeK4/pMH2kEMvE7fFjiG7cVxUYczN1RPl3353BIh65BdSqi8ZbxezHbPuPnoST3teoBdiwVoHcrz80by76n5kIiLKrunuzDVugq3H7NgE8xpnIpttCADiYHL2jtx7Tg7ft3rSMEjep6OSThUWyKyJUUCgyoENmrR7zyWhk50Z1YCkFW/1hOma+chdPY012aGz6nhDBxOOYrI54YlD61jKsheE51IYWKdQXAKfsp0UCGR3Q+t69QuCm+lclZgatuoBWO+tt/WKekO3T4Y1HMx1aoDk/kAKYZwfZXPyCPQ/cIK0GUpS3H8Fa8uDIpaf5KqTj3DTV9qCcrtvrmrf4Qj8YiyN+TBp/wtLiMLtofjI+apTxBqQSGss7oB/JsuH1wacA5EGoKwO8tZAuMWLM6wuPjaSgVRlBezbGsYqKUAiDZkcXLOHco6AbEptMbK91t5Mgw0V8ActHRg+3LbuHo5P/tWwzdF+1Ht7J7EvJIzvz4lfuVGyUA4fh6uVHHeeJG2OrNtjrqr2zq5iSEECJ4tc7hNDnGR7wqHc3TPEi0YIicUNU5701v3wcKgxqiPrrfRS1ONR7SLqiTHijpdFBOZyFrh/bd0216w/6vUk8uZoq/3u0PjUkVG10ofb3farN4GUfOwvQBgSRkRjIe5K4DH1ZKnE0SQtV9UonzWPqU6jWGepPsOlTMiBkdoTeUhaeomAygbJjRuBsKzGGTDDaZwBpVvB0rBVIV5xge1NKVCSTnOyGYi2gYkmEcWgSdgWk+OOI=,iv:9PDXlSUYz+vl2EzVcwMHZTgyamXLNZU7C+XdcAEi9j4=,tag:bRkH77im+qHjDewMW6PsbA==,type:str]
models.json: ENC[AES256_GCM,data:FvlbdMkcngJchi0GEjIEDjxXpQjRwh+3xnlZNpApgd3v1SpMd0qD2d3TbQr39+feSfsvpxhdmmyU/PmRHZYVdr/QpIKlZBWBD6qTCaOSq62NQbR4ck336muIaWD4AmahWCvRouYEJdEsnG7nedjAivw4cls18vMfnL5pNm2t96foScSyHosWoh3RdT6CWLmJE39+kS6/fwZ1D3Z/GEm/E2zHkVtEbEw/fk1PiqsnmZWt8QcSRAeX8aIDkphV7wOKlkri11r1NuFQKMVIau07VPce1YEi5rsJqsircvDxQelQiMGQt2y2M4GGlXx2NeJgKEtP4Hf9mVGhTeN3CYLcXi8Qyg7k1GIawGVW/Em0kWWfy0GMjvOMfYozNmIpi0YQdYBbr3j6pkixtVM2dWezvld5QIYLnHoGMxw4M6L4IqVIonC3j6tk9pV1DaL5IFskEXyb6ScTrmVY4mQ+VEXvfLYA38jRauWxx/4qf4I7B0RzfGzOx67UJh2jKj0bPWNFHKt5Dz4CK+ul39d1KUPHHMoyaH5gCpOAfLSkjMm44X1Cn+O6uY+QFCugh9MI7bJg3EbovTCZxz95GCupWAvDshv59f7gpGL0T8AuRZHlNp6yZMLdviR9yf79d2rW26PIM1QuNOUWELrOcTs5IAmyuFYr0PI5W5XSf4klOdz3UdhPfMDcyRqtSzVDsSo9oC/WeWOo1+yurtdqqW0wLiFpU/kTaZ4JN5Kl080twqFEqgyPnJrbmlxwWlt4XN2SV9D7pcfE61FzUTwoi5PD/8xt94Fp3XerSgJtwhQ6X4Neo47wCFoMR4Y69mlHrFgSJA4kLFezDc2ISnyAix0W5to0cjwJqQ2EgonIhCXD4xWTvYeApHJMOC2o4B74K5efOZ7klb6PwDMvzM3LyeGzjoJh47aJGvhWN/MpQZC2lSHT8PdMxdGDc+OkEVmsvMncddWMGjQhkTo+69sUexARxLY4H22TCt58azqQGcWa5e7s9NUHb6bRVKHzXh4HMiIhQ0jBev++Jz5gIOyYRnw5xoGbo3ROv8ndczVAqXyAla23EM/VBWT0hLMjr6xxonuLQacP9dPJE9QUofjRMeYblkVvQdIw8WuKGbAlbTR/bokQwGdVp2bywSFQZj7BLlPjzCr0/LJmxYMdq6CKB049PdTlS9/UKBHApddMQ/QRXfOYUCX+lm4k4J19jd7dWk34WNUrTEPSR2d1U/2wTnMvzlrnQHx/zk4/lmI1newfjIbybIbXypX8xVEqzFWXpxUe+iFLP19KmmmHG7gGsV32k1VSGvYFI0HOokZga98QsAbJGbi1gQq8gYpMR/mBdhjJVcnAjOYpg+YaR0vCyrhBN55aSCDoUbT1JOqxvFqmlBHDxTWorD/3KGFxEH/wvSsxussFmUCWdEgYmjGO1ktnQBw1YpOB7mhDUlzvYCSxfpexsYufDYqTbrb/R5gmATC0pEN6EgTth5YqFOL7AiIsogy14RGjRym1wvjVl5SUU9bMQlWGJbQuvNRRwL/TzR/QP210jMsnAh9t8sncDnFPIbhFgeBLsU/ctYoBs0nFVWGWzcuoP9HTnvdWTT2xKmWAeFz8I3DZ5sjm2B787Q4bCrAszTvVLeLiRw/pOE32BeNfG4nNiZDvMQhUomAKpaQypWTu2wLsv3ISm00gud5sRKB6ASJANuSu+EhG9k4OYcBEzAT9PKUxclupFbmsFnXx0u9CXDj3tErrPCUs7DuFSr+fz0ehCLMH+GrcqfytjNDVoTSBRZg+lJYKjAqVCGvq8uRf5V/rdulbt3ggx/uhGSiPzYrBQ0g9iMdjQqpNGBwjlxGkhRE9D2n/PoerAAnHTfwc9Hezffncostjp09xw2CB/xOx9HZFGalQ1R4F5+0lm0iR4lMzc+se+o38foNIiMeMIwAVBJu7IO++UMuKCxpxmdAGqEYj/Ddpzfi3ZcRwe/1HhanVvKRMJsPOhnv0SB2xBrdb7QHnEzLbRTriVRBBdOI77HXWgDt9uQru0tEXCEQ0n+f0BWqHVa/RLN4Yj0dQVWLXU75ddyHnRaCiVdLQnObMvG/QXqY1k/bJTzGNgVF8/KGY51NgdfqskOXC9Fv5IGiyWq+/DL7tQcrmThkzUhZlwDBmwUzPFmzuUC7Z5PkXv2EV+eUrYF16HYqp5KgqTkBHr/JXR8ljb/5gMN929Jk2fBVEuzaEG6W9Yek+Qbx1cnXYrMBI9GWJUJsMsxsAj+1bAjtQeQwt0nXgZesTaV90dfQYq+nTGXMNjqaMhsrJhR6lU1+CDpwY+IObHSttS8AKB+Czsl0nUB44o1a/zKRDNTCGGmPFuv1Q5Ub/1myxorLDDF1mYDpre5TedygtZ/1YSz/YOpKwmWdW32FsnY+m3qjEAcfV5f923N/z8ENVxVpeiWL6mqf7h6F3qTiaSsFq2O8m8tVm70X+oRdig3hPnKqlK/KoxVmzZErhEJ+KmP5I2uvIlb4PViE6imu8kw4VqMClb7jJ7Yk+2Zy84C8RGuDCN8GCDG7FvAahA9xVGxB6rmEpzvqFBxos2EeLOlyfewf7rCw1JAL3C9hZrciA9Jqb/t8B7KqCmCOsKK7xVt1yqyG1dsa364JG/1SKSMKebTheuymZ1+VnGjRZ+pgSAkQGjCtp7B3fTlhjt+WwxW4Ml27yfjokQQ==,iv:8Z39eWukGSMePh/3Dj35e6Zahejil+eeMSqwMYf3snI=,tag:FyBg1y0IQEg/m0mgpf1ESg==,type:str]
sops:
age:
- enc: |
-----BEGIN AGE ENCRYPTED FILE-----
YWdlLWVuY3J5cHRpb24ub3JnL3YxCi0+IFgyNTUxOSB4aENRQllpaWh3bnRJYjEw
eTFkM250eXBSVUFrQWJXSzJUTGgrNmJqL1NVCklRT0tRWTlRU0duYWNzVFExQllS
bWo0TndLVWl2VGllb00zR1c5ZERpWmsKLS0tIDh6Uk5hUFV4bmRkK0lHWWN6L2Jk
dlU3cXJlVFZYYi8yMm5kVUJveU91OGcKyin8Tr7OkCocRxf1dzWl/QsC4l2XW4dn
g/it6hJQx1P+23STw9pDZVPqEj4fOdqjnNoRqVCkM8wH4SXJfwrnlA==
YWdlLWVuY3J5cHRpb24ub3JnL3YxCi0+IFgyNTUxOSBBZzB5REh6amhCNzRmY0ta
MkN3Y0ZlR29lQ3h2SWo3cW5CUThkL2RnU1NVCmt2ZkhIZTlHN1RQRkFrTjVvbjVw
RGRrTXRoYmdQcnlMSEo3ZWsrZUQ5cHMKLS0tIGU2TGJqZDRxUGJpZzRveEtZankx
MzhrT1R2akxxby9QVzd1RXB0RDY1LzQKOF+/e5z5lPX6Y1sMTAHuDj3YqW1m+sBd
u/0R0YnBonYM3wS5nJE3NZMkImaAdQlUjOzQepfBldG+lz++rlnAww==
-----END AGE ENCRYPTED FILE-----
recipient: age1e5fq3hwxy78psus2nfvmtmua36g0u3suk78ephw6246l974d2utsvn0hla
lastmodified: "2026-08-20T07:26:38Z"
mac: ENC[AES256_GCM,data:YIN51aoCGuFrJwxJIGbCf7vY/+S4uHR4HwQ6un084hMKnInh1uK96FxWrQBlcheqBDfoaXHXXmHAd14LVhrVEsj3R1cFPpxiQpqjwt+d+mON4YBeOrC/VcStAU8joKcLsc8H0PF41PGMcBmflQVDX30D/+am60hZ7FMnGavZqgA=,iv:N6tv/3gh9BJvZdWXAQqTwaceR5nLbiA4YwOz01uwbtg=,tag:6VG9ZkEVpE9xLCNy0HLYcw==,type:str]
lastmodified: "2026-08-18T20:06:08Z"
mac: ENC[AES256_GCM,data:IM9HkpdwtQE2wCkjwDWOmHH4uP7TlIsrK4TVytiecvYz4SiLk6IRUSIu7I3a+F+dtltC2WtokoATaB69DTXPoI54amzzptirxiFD5FbaU+u2gLjo7KI7V0smYGuKqMYwnod2L/4GdlvP6xjVxFWuA01rQRaYBkFSumS4NABl/I4=,iv:al1MyBmFwni8gap7PPZxWaCwqFKCicfPq6nVmrSn9Xc=,tag:KJGTDQ8xRGF8qMOjoQ+dng==,type:str]
unencrypted_suffix: _unencrypted
version: 3.13.2
@@ -1,23 +0,0 @@
apiVersion: ENC[AES256_GCM,data:ECE=,iv:bISz4HovH++X7DW1Qj8Cw0L6+EvPB0+68hh7tfyW5C0=,tag:w+SjD7MsfeIuSf62n+Zl7Q==,type:str]
kind: ENC[AES256_GCM,data:AxXaV3Mb,iv:xfJb354Rrjz3zLctW2i6hl40yh9EsfIqKxDvmn8jqnU=,tag:1xQ/RR3KhPLKCTWCYjebJg==,type:str]
metadata:
name: ENC[AES256_GCM,data:gdQ4OoMYKUPEbsUeA5OI4i1xnh8=,iv:IA2kaCyqQmkYpLolvcTF4aleh+yd/ImXJMhRMvpGCgo=,tag:8qyYYw4EhBKKPzEmpepAeQ==,type:str]
namespace: ENC[AES256_GCM,data:/mOSyXWmRg==,iv:cpeHUTlMqlJzJttGtuR3DoiMtvVqFmDS0/5Tl7K7c2M=,tag:q4GxWgWn2wQJxJHqnq4WQA==,type:str]
type: ENC[AES256_GCM,data:92OcrsuV,iv:Z1XBy6iZ6unGrK4/SSdDa58pbPL3022gU/f1AOM4uvc=,tag:cd4wWRq/ohm6BJ/Eo6HOIw==,type:str]
stringData:
REGISTRY_PAT: ENC[AES256_GCM,data:zDKruOXhzsIcFZTTp4r4r9rwHih9uf1cCp06KZ6eIXSIZDvR6aQY4A==,iv:5Vch5Z7xn1KkRxrgOs2p/M58nd6WhtcuUhoZDCi8LSY=,tag:14SAZIWGOxJMqLBtOorjYQ==,type:str]
sops:
age:
- enc: |
-----BEGIN AGE ENCRYPTED FILE-----
YWdlLWVuY3J5cHRpb24ub3JnL3YxCi0+IFgyNTUxOSBIaG5QVUYyNjhvaENPcUFy
K3RSSVM2N2hTeFhNK0M5YWZmODhzZU1HaHlFCmNGZWdFcFEvanVwcXpGSmVvRHVx
aG45c2lBY1RuSTYwbDZTVk1QeHNWOTQKLS0tIDg4TnVMNjNtaU1VQk5zQjUvU2hM
R245ZVdqc0ZWWVhJb3dOZ3lpU3JjUlUKawSg09ZPq8FKx5tvOVZZ+K4yh7eTQsUp
be8mWUpS0+eEmNqh35BwU3HrETMQFA6a1kjVp30JOMtqa5rbYlzF7w==
-----END AGE ENCRYPTED FILE-----
recipient: age1e5fq3hwxy78psus2nfvmtmua36g0u3suk78ephw6246l974d2utsvn0hla
lastmodified: "2026-08-23T23:14:53Z"
mac: ENC[AES256_GCM,data:yzi6woUaUCi6w9pG/eKnU7k/VfZgoXg/tW8p9joG8p7iabLWMdlXlx23m/CItw/NE0zeGWA5iZPPFZXOit2vN36VzK3kQNoFW6QbhvYLZi+78C/RWIdcpZkxB/tPIRG0vq8Q1SA+rwGjG/0xeAFh+R7k+YBTMd2XWBH3P7EI5T4=,iv:kf6MDaAVrtDvPIEjHMMLxXDSDRC3I1GpsPeJFUYppiw=,tag:NukVbGLa9EoMYmRsa4nBtA==,type:str]
unencrypted_suffix: _unencrypted
version: 3.13.2
@@ -1,24 +0,0 @@
apiVersion: ENC[AES256_GCM,data:D9Y=,iv:EH+zD6bogxh/h/Oe+RxDCtfO96tkc56ou14V+68nK7k=,tag:xTeuhSzqZyrS/ltqvtHcgw==,type:str]
data:
.dockerconfigjson: ENC[AES256_GCM,data:P8x3bhPbJTFvFIKE8WQY1P7KqPwrxNTYTNR+4/Z7nZSr2YQab/Db0SRcvFO9lr4ImdO6waZ8EZ6CI1RzIYx1WcjGXs8FzX1jpBXSw9EG69vRTZCqRRJR8c2yg7qmNFfSvcPAUhK9tciYCzzFWRHZEDAmQ52cox/sMdxE/YR61dhYCy1E9kPwTnaDmDj5Hu72mbLJDxIYpN3T2hVWlw02lHYHgFnuKsPb0Tf1lxn174j/gMMxUT6ynVhWSEExzHvTqQ0RZzuP+VXAh3N9HasbtfcAabt3FtEjYNZqYJ+OE8oiX/A6YjgPDqUOR3AJFZE3,iv:6YsyIHQc8xp8T8XUWhN/pBeaYVI/VdIHOe/w/hb5e6U=,tag:nmGFdkv2hLgj8Dp8/0MOkw==,type:str]
kind: ENC[AES256_GCM,data:hOA36Sjr,iv:y0XHfUOUnut8z0yM2g7Beo3qiqxJhLLhffPpNlUhaec=,tag:45s3FjYCr83vYCSXDXu2/g==,type:str]
metadata:
creationTimestamp: null
name: ENC[AES256_GCM,data:yD8Fo/fncAA4qkaY7RjhRg==,iv:qwc4WId/kGwradgxFUwG5B5XIVYRlMY2HxsaGJAzrjw=,tag:O8GwNdDHMJreL609G/gykA==,type:str]
namespace: ENC[AES256_GCM,data:zg2o,iv:KDpM2L71/LDI4JLTQUtbcv5SV0IraAbKpNEBzSFn/rE=,tag:uFm+3EkgZN3NC2hUPXo9Pg==,type:str]
type: ENC[AES256_GCM,data:5/a+6lRp9Ea5rU/+gEMIgoDYs72xRWkfwefMcy+h,iv:jYirTXlb4rFwvb+nLcgB5X4x1Q/S+3LsVL+7ypS7mkQ=,tag:5Wc0DrMZxIpC0e39MzlAlw==,type:str]
sops:
age:
- enc: |
-----BEGIN AGE ENCRYPTED FILE-----
YWdlLWVuY3J5cHRpb24ub3JnL3YxCi0+IFgyNTUxOSB2ekxNWW9VQ09RTUlMUzVy
QlVrRm12Smt2akRaYmkvMHRPRmJvOTdBTzE0CmdhVFNyWDJnWDV3eTFFUm1majk0
L01ZWTNZdmcwc1MwMHgzQmdVZy9KcjgKLS0tIHZLVGg5VmNRQ2ZNY1lIRHlzaUlo
TWVNelZRcHFveGRiNTVvNzJCNWtDSGMKyV3Puscgx3RqK65KSL6SYaTauxsBY3qd
CeFU928hcB86DwAG/Atq2Qtd7S9pzuzOVQmXRZxwpCDTTyRVhU7eVA==
-----END AGE ENCRYPTED FILE-----
recipient: age1e5fq3hwxy78psus2nfvmtmua36g0u3suk78ephw6246l974d2utsvn0hla
lastmodified: "2026-08-20T04:43:53Z"
mac: ENC[AES256_GCM,data:Leus5j38xJwJz3Ge9WggBjQSAh2ESlIMUnX9SylE4oIcAt71f8WadtSCOmnqT3ZK+uN8f+Huq6WetGHWUdLfZjZUngHQHWLoRP1xpTVvB5HwJK4F1ASvmL+u1rC88AG3JsZc3Kc2N1G+m6QruQK/9HSmkMi/HGXy658ijzErEzk=,iv:1ga2dKDqTalj9WnjVT6AubXsL7130CuJp3SbBkTb/64=,tag:YsD9metJdYiJmdVXpPuIsQ==,type:str]
unencrypted_suffix: _unencrypted
version: 3.13.2
+4 -6
View File
@@ -1,10 +1,8 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
# Disable hash suffix for all generated secrets (stable names)
generatorOptions:
disableNameSuffixHash: true
# SOPS-encrypted secrets via ksops generator
# All homelab SOPS-encrypted Secrets, decrypted in-line via the ksops generator.
# Each *.enc.yaml carries its own metadata.namespace, so no namespace transformer
# here (that would rewrite every Secret into one namespace). Renders exactly the
# Secret objects — replaces the old argocd-cmp-cm SOPS plugin.
generators:
- secret-generator.yaml
@@ -0,0 +1,25 @@
apiVersion: ENC[AES256_GCM,data:894=,iv:Swg6ADUgmrqwz7wqAZHip9/qwFu0Rn8S2Lx4gBH8LJM=,tag:zbv0tFfRLtwFxBfpsuLt8A==,type:str]
kind: ENC[AES256_GCM,data:jzVtJHYw,iv:ToZ0orfJqfGF/OnAPeYu/g2f4fXAMOZQDkA1+tmIccs=,tag:pi2skbwIU8qOQVEC86MdAA==,type:str]
metadata:
name: ENC[AES256_GCM,data:t7zCeZLvAEbkJqUjWi26yD6UDA==,iv:MS1gq/bwKZdLA1itVDtsrdSOfI7e2CrhjvX5yhs0eQA=,tag:lC1gcoFMI5nfzC56U1WXrg==,type:str]
namespace: ENC[AES256_GCM,data:tp+d,iv:gnlet/60mgbSWLXEQpSlcWD98ky7NvlNCzGLTMys0JQ=,tag:PQYj9UeA50YenQESTCl7lg==,type:str]
labels:
konghq.com/credential: ENC[AES256_GCM,data:SOqQ9bLGLK0=,iv:a7En49UhRDwgHbv5NRB/XilEYIKQdaDqKH86WDrJB5I=,tag:JaDZ6Ko8ovY6AZ1hW9YMGQ==,type:str]
type: ENC[AES256_GCM,data:r1K5gvop,iv:Gjv4oG2Unyql5rY9RTTljFqyd28xI81CDWtpavuuW5E=,tag:jF8lBs3AQwHPVEu9V+mONw==,type:str]
stringData:
key: ENC[AES256_GCM,data:pm9GmSvX5MAsXO/e6ZcI4NF1Hwr3qjG6LaEEjvV0ihvSWhO23drlAEsrTtzppqgDZbMORf+5+P53mXU=,iv:5AbHNKeiMPoFQP/qTKdA0vEoYPzuaf4kIdGGcMSmfIQ=,tag:LD5w7XK+hiCS5D410+nCfQ==,type:str]
sops:
age:
- enc: |
-----BEGIN AGE ENCRYPTED FILE-----
YWdlLWVuY3J5cHRpb24ub3JnL3YxCi0+IFgyNTUxOSBUTmoxQkpuYUp5UlRQMkph
a0RaazdvaU5sWkNuL2gvVjlUYXVWV0dUWVVRClZhWDZCN2hpS2hnOG9Pck9zOTkx
RGNGMEI3RHpNbzVaaWNGcTNSSEdzZHMKLS0tIFZERnVJWUpreUh3TTlwbGw0dUx4
MGlCSkxuWWVEK2RaSDZPUzhNSUlCa28KHN0IsgQc/kBqmjQ6+4sgfb9PJy/45MwN
rXaLJ1htpqPZ9MJ8iOukRi0IKnKgQWXsoZengIxGmcOnEctpoH/kyQ==
-----END AGE ENCRYPTED FILE-----
recipient: age1e5fq3hwxy78psus2nfvmtmua36g0u3suk78ephw6246l974d2utsvn0hla
lastmodified: "2026-08-14T01:07:50Z"
mac: ENC[AES256_GCM,data:nY+YVwU1GuK8Yz+EZOQKkZKN28tm2L8afflc6hsgVFCFmsep5kVT+zId7AgemvQ+qnrho5N5nqxY2knB0gusFfWNKF3V5A5GBq40WtZCMaAtcwhJSex4kK7ZyaZD6oWDW/RTUumSrivSowkWlqt1XlDKyFLqSlpWQTd7JiGmUv8=,iv:zE8+B5UVYSuuAGYyvXsAwp1N/4vGduCoGyEMCNNEUnM=,tag:6RcrjneV2dOzhmoQi+5HsA==,type:str]
unencrypted_suffix: _unencrypted
version: 3.13.2
@@ -1,25 +0,0 @@
apiVersion: ENC[AES256_GCM,data:wGA=,iv:Z2Gfzq3aJ9j4fYaeLQolgLb/XELTHrKX9at3vUsMLIw=,tag:yZLZFnltJln5jFoVsyyZ2A==,type:str]
kind: ENC[AES256_GCM,data:bJkV4pNj,iv:0ZT7l0kw9qSoiEMZioRw1aBzzlxBkoXR+hOxoPI7zPU=,tag:EQczPq508Qw1vi/oLCeQpw==,type:str]
metadata:
name: ENC[AES256_GCM,data:sZwTVM41PPaidMwFNRo5RvU=,iv:BHfuwIHng7rkeLK3a69t8cI9QeSfD/3FEXqxby+gBxM=,tag:3zQx8APnE2ZpBKf/ZYCxOA==,type:str]
namespace: ENC[AES256_GCM,data:i7lpoEaZ1oXS,iv:jUYyDPYhf1TV51he/S5MlKPD19Vmz6wfE98y8qFEg3U=,tag:zgHDVlplmW/XPUAKdncfYQ==,type:str]
type: ENC[AES256_GCM,data:2khs1uIg,iv:ET6HcBGyv33fGFlFAl3dkJQB83naHGeaYuWvW1IFhvw=,tag:qIfMeaSwv//7cNWvy8O5dg==,type:str]
stringData:
PAPERLESS_SECRET_KEY: ENC[AES256_GCM,data:kj9DrQYL3cQGJz87FHYlFKy6Muu84Oy/6wnsWAh0w/3MmcTAsCzQvUtKLKKf1U61JTQ=,iv:UffVv78vMHDEWHRFdSKZ/6qyrVD02Nlk0CNxwWd1jTo=,tag:LNMKnFWw7QCjfcKzI1s8ig==,type:str]
PAPERLESS_ADMIN_USER: ENC[AES256_GCM,data:a3bx7z0=,iv:nrU6VXZE74PNXM+Lgg4K/3mhusaQhA2GAgl0GzqAQt8=,tag:x8nswTfkFHR6Ou+JSI9P3Q==,type:str]
PAPERLESS_ADMIN_PASSWORD: ENC[AES256_GCM,data:+SWyAq7D1mNt1TlFOlo2Z0biK8kEn097,iv:an2ZZuXNSig+7mJP7ogHFc5+QvUb+z9pzqcV4xvFLbA=,tag:OaRiVZsnajEjAlc9mDWo9w==,type:str]
sops:
age:
- enc: |
-----BEGIN AGE ENCRYPTED FILE-----
YWdlLWVuY3J5cHRpb24ub3JnL3YxCi0+IFgyNTUxOSA1Q2V6bVBUYmVRR1N5SCtF
eXdTV2dmWjhtMS9lRFEzS0wzbkd6Y3JLMVRVCncydFBjdmJkRCtVUXphR2w0SDlJ
cng2bi9MWlJzTEN2amJrYjRJN2VFcEEKLS0tIEY5cmw2RGhnbzUxZW9FaFJjQmVN
WWcvNlNiYWdwbnNSR1Q4alpDZmFqTFkKh9TOw8ERP9fpx2pKi/Q0b7+OkEv0UC7o
aAIK4Tzvi5dp6y9IWcu9l6PjDLeYWOJ5wr7QABaFNOz82hngxUleDA==
-----END AGE ENCRYPTED FILE-----
recipient: age1e5fq3hwxy78psus2nfvmtmua36g0u3suk78ephw6246l974d2utsvn0hla
lastmodified: "2026-08-25T16:04:48Z"
mac: ENC[AES256_GCM,data:OXPQXUyn/SDftKH5nRzhqEtnaOb4Gp7etmGojqV4Z01kHnABZ42PPIyiBb8z997eLNZsK6/1bi7nUNyVp4fIuXG4F52aEdMewofOocCHokRnNmz7jzhooK1gScJb2u0eHG3FL5iLONMaGgVpk7BLYO3e0xiDytWGe8BxcuDukPg=,iv:kh0jUwFvO6AUICy2Us1E7YOTEcp3L+ptrGdDWBSpWyc=,tag:qUPFmZ/rpljln37f/NRjjw==,type:str]
unencrypted_suffix: _unencrypted
version: 3.13.2
@@ -1,2 +0,0 @@
FORGEJO_TOKEN=273fdcffabbcbb5a191e8289c73d106063acefc6
LLM_API_TOKEN=s3VksXyw2z3sGbegnwjMDFnJ6CtNRd1a5CcnE5A4ET77toCcykNdunk6Oa2J
@@ -1,23 +0,0 @@
apiVersion: ENC[AES256_GCM,data:bnY=,iv:Fuc3aqncHQ+L16o7eLarPbOECD3o8Mk5c2r9pQBpy70=,tag:JPfEnbNb3wZXPdXafnJDqw==,type:str]
kind: ENC[AES256_GCM,data:WFlmi4Yg,iv:Zq/KQbgNcBVoo8ZsQ2H79ygyc8Dtkgxh4fCpEExfwSg=,tag:cWHP6Y5V+nZP2tFMJrOB8A==,type:str]
metadata:
name: ENC[AES256_GCM,data:iCXbhwvg3Zq6YL/4j0wAy7Y=,iv:8s+d/8lDVEL7bGdIF+GOtAxapKnmx8JTjLSO04hXF5I=,tag:LjzX40DjK+uRZPCXsMlmwQ==,type:str]
namespace: ENC[AES256_GCM,data:HRMdZdCbxORQ,iv:MvaIWoKWjJRA7/fce0KtXRkFH/7cn0OuIg2QwHEdQzM=,tag:NqztiGyfU3BaopWBKhx2eg==,type:str]
type: ENC[AES256_GCM,data:myBW86Za,iv:3x9ys5UzVhAuX8gvZO67B1e+Orw4Aqasv/lHBgUV0b4=,tag:yl18P6SaxVLJidwmRtd4aQ==,type:str]
stringData:
FORGEJO_TOKEN: ENC[AES256_GCM,data:SUoBpNKOItyNGY01EhKNlPH0fyN4N7g6bfU2jsGogpCMhP7NuRipoA==,iv:h77RtmYZjXxiHYw1pHynQHuVX1+yJDGHwsXLf+DbUYA=,tag:GHzoDNHP1GGqlN5eT/N7TQ==,type:str]
sops:
age:
- enc: |
-----BEGIN AGE ENCRYPTED FILE-----
YWdlLWVuY3J5cHRpb24ub3JnL3YxCi0+IFgyNTUxOSBzTVBsekl3TGgzQVRMUU9m
cTduS2NoZW5uZFNNMG13cFY2cGVsTnlXaXhrCjJjbzhLdHZ4ZWpUV3J0cDQ0eVlM
WDNxdzVoQ2ZzcGJSbTU3RVorcnczNVkKLS0tIFdtQTE4Umk2TDBzUmdKOXNkbjFi
Vk5vK2VuUHVsb3FQL21vcGU1UW5CT1kKFM8vVjji3Cg9dvfTr4Hx7BJC8JH5ovef
Dj6zkofhsNWPgP9T+mnQakj+C0RKmHOMJqfWP7vwCBkZoNosIJVlMw==
-----END AGE ENCRYPTED FILE-----
recipient: age1e5fq3hwxy78psus2nfvmtmua36g0u3suk78ephw6246l974d2utsvn0hla
lastmodified: "2026-09-01T05:32:42Z"
mac: ENC[AES256_GCM,data:it24T9y9ixXo2aiL37k93vKFR+SRjjuI9DQdv0sWYtTogWnc7+uXBY4Zip/ouWyCse1muKKAGuek5c0XVrvSw4an9VkaXFczeunaZb6MOyVbVOkmJr+5xZFpZGjYcSkrhaWcVheedZ3iIFU5UWI7BBn/qQCf+HJ483cJqwtrV34=,iv:WprlWJdsMBNjqaA0O3ekfXMUpX5gC6OLYortQXYdTS4=,tag:rGqSSRTYTv2VR6AcRokO0A==,type:str]
unencrypted_suffix: _unencrypted
version: 3.13.2
+1 -4
View File
@@ -11,8 +11,6 @@ files:
- agent-pod-ssh-key.enc.yaml
- authentik-secrets.enc.yaml
- cloudflare-secrets.enc.yaml
- forgejo-registry-pat.enc.yaml
- forgejo-registry-pull.enc.yaml
- forgejo-runner-token.enc.yaml
- forgejo-secrets.enc.yaml
- grafana-oidc-secrets.enc.yaml
@@ -22,8 +20,7 @@ files:
- homarr-secrets.enc.yaml
- homelab-ca-secrets.enc.yaml
- loki-secrets.enc.yaml
- model-invoke-apikey.enc.yaml
- minio-secrets.enc.yaml
- paperless-secrets.enc.yaml
- vault-secrets.enc.yaml
- vault-unseal-keys.enc.yaml
- portfolio-secrets.enc.yaml
@@ -5,9 +5,9 @@ metadata:
namespace: ENC[AES256_GCM,data:wK6m,iv:KtA31Bo8aGE1HU8H9KWbMwt2NfywWfy77G/LaReaI1E=,tag:V9SSlPawRQw0x2utFNN1aw==,type:str]
type: ENC[AES256_GCM,data:6UPZ1ZTR,iv:FxN1ebrlJ4IO3eDGEYSvktSpMgAebcl0DO1WHh5O0+0=,tag:2bZN4zdbHzx0oR+bTJPPJg==,type:str]
stringData:
key1: ENC[AES256_GCM,data:ke5EfsOsZHBJyAQoFfwhhWCQJQgnwcrBqL4CzPIfkx7bi9S7kWgWySdxcXQ=,iv:eYKxG8k+hizp2t2i/YMR2lQNJQFV+A21YyWnDc+kJ9w=,tag:Wa1Lx3EKvfpYfnLrBOWSIg==,type:str]
key2: ENC[AES256_GCM,data:JxiXgLKvDe5oiHlwIL/Cj8txHc7fVQ5VzBcQMU/ro9TSOcTGSe7Z3omQzwY=,iv:x5arcY2dey+npMpUxjdUPV+t94LEaMOXq3iOar88eT8=,tag:E/Z5zEDy87B1IAWkdjdNEg==,type:str]
key3: ENC[AES256_GCM,data:AsxyyBTHzt+SLru6ayi1UbgPPFmdIvowIDdcfc7N03Xt8V/KaJW3kbXExG8=,iv:lA+bERIBq4+bxx08ZZjP5tMHC9JQlGPW1EMCwEixSeg=,tag:xkuTAU3KTidFuhNMXkVFhA==,type:str]
key1: ENC[AES256_GCM,data:dV2HOh7W1Pl0QDJaGvtEKpBppybQMiK5kxzyk05zkA352ZmJvE8Ppm5yWKk=,iv:grQ8v2o/LHpJnIZonjtgTHKcLUQIH8xl1FtWtEs4rEs=,tag:lANcccYT6Ouo/0IcoLw4uA==,type:str]
key2: ENC[AES256_GCM,data:YSdpYL8h56PfUrvMFhBXhmBG2en0toLKEpYlJqwjAk/vI0jm7gq/TqGn034=,iv:GxthkNLhm3qHxjkZetiYl58Qa/K7G2ib6E+LWK16H8Q=,tag:q79UUoh6uCcMZJ0U18Ireg==,type:str]
key3: ENC[AES256_GCM,data:sO7Lao9qkmIOdARL+6FBcLo+e4z8LzMEzNrwcy5hv9uRuxppwAXrP0MO1vw=,iv:wSWe9h1HovRhV5YCZrmvcB2VEfRYQ+gasGO3uh8IpmQ=,tag:Bc5yaROQDTz1AhVnFDm/BA==,type:str]
sops:
age:
- enc: |
@@ -19,7 +19,7 @@ sops:
7RltNZF2SCxjIv5C2pqf3CgmRBaQGWgMybRRH5gdB87PLBKkPL3+HQ==
-----END AGE ENCRYPTED FILE-----
recipient: age1e5fq3hwxy78psus2nfvmtmua36g0u3suk78ephw6246l974d2utsvn0hla
lastmodified: "2026-08-26T23:33:09Z"
mac: ENC[AES256_GCM,data:kr8XuOqXyhfrC84yBot5aoVQZZhQd9xYMUEs0XDwkmCgtG0iobADYH5XWPh72xSdWEtwxkZJL7r0UEyEhU9zjlUA5qHupOSBBt1SgZ+zstMvaOqUnlNn//p/DIJBpsiT/qmx64NpTLAiz6lm0796MozIMr8PTX+ubGLHi9tnUiY=,iv:jwP0MLCT7nG6m+nZeqNip9q3BcScpcDXmInnY95DicY=,tag:Cgg2IUZARkaK1gF2qUIthg==,type:str]
lastmodified: "2026-08-12T20:52:24Z"
mac: ENC[AES256_GCM,data:kLjoBWmJ2bZ13EbuzgppahHKmCY/xtCi1R7xCEUbCP4FQfiVaC4qAbVftE4lTNlXa2mqCKsGPYZ/A3HV/4+a+it/pg0a+rj7E7DszhieWbZufHMJs+w4/Le8l8wFFE5aNW0wdvGTJZ0n4HBOdAkE1qN9ReaiaQgr0rk5ggOPG9g=,iv:MiDVUgcoKrP/gj4keqomo6OB12dmz9VoesTiHQTnBFM=,tag:cLK67Pym4BuvXLF1RUA8Gw==,type:str]
unencrypted_suffix: _unencrypted
version: 3.13.2
+2 -4
View File
@@ -300,11 +300,10 @@ spec:
port:
number: 8080
---
# NOTE: api.riotpiao.com is deliberately NOT here. Its namespace `api` is
# NOTE: api.riotpiao.com (Kong) is deliberately NOT here. Its namespace `api` is
# created in wave 7, and this Application syncs in wave 1 — an Ingress into a
# namespace that doesn't exist yet would fail and mark this whole app
# SyncFailed. It lives in k8s/apps/api/ingress.yaml, synced by the api-gw
# Application.
# SyncFailed. It lives in k8s/apps/api/ingress.yaml, synced with Kong itself.
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
@@ -323,4 +322,3 @@ spec:
name: homarr
port:
number: 7575
-6
View File
@@ -1,6 +0,0 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
namespace: kyverno
resources:
- policies.yaml
-42
View File
@@ -1,42 +0,0 @@
# Kyverno: Policy engine for Kubernetes image scanning, Pod security, and admission control
# Scan all images, enforce baseline Pod Security Standard, prevent privilege escalation
replicaCount: 1
image:
registry: ghcr.io
repository: kyverno/kyverno
tag: "v1.14.0"
config:
# Webhook timeout for policy evaluation. Increase if scanning takes longer.
webhookTimeoutSeconds: 30
# Failure policy: fail-open (audit/log) vs fail-closed (reject on error)
failurePolicy: fail
# Resource limits for webhook
webhookAnnotations:
rules: "allow"
# Pod security via Kyverno instead of Pod Security Policies (deprecated)
# Enforces baseline restrictions cluster-wide, with exceptions for privileged namespaces
podSecurityContext:
runAsNonRoot: true
runAsUser: 1000
rbac:
create: true
resources:
requests:
memory: "256Mi"
cpu: "100m"
limits:
memory: "512Mi"
cpu: "500m"
# Webhook configuration
webhook:
timeoutSeconds: 30
# Failure policy: "Fail" (reject on error) or "Ignore" (audit-only)
# Set to "Ignore" for initial testing, then change to "Fail"
failurePolicy: ignore
-204
View File
@@ -1,204 +0,0 @@
# Kyverno ClusterPolicies: Image scanning, Pod security, and admission control
---
# Policy 1: Require non-root containers
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: require-non-root
namespace: kyverno
spec:
validationFailureAction: audit # audit first, then change to enforce
rules:
- name: check-runAsNonRoot
match:
any:
- resources:
kinds:
- Pod
selector:
matchLabels:
pod-security.kubernetes.io/enforce: "!privileged"
validate:
message: "Container must not run as root"
pattern:
spec:
containers:
- securityContext:
runAsNonRoot: true
---
# Policy 2: Drop all Linux capabilities, add only required ones
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: require-dropped-caps
namespace: kyverno
spec:
validationFailureAction: audit
rules:
- name: drop-all-capabilities
match:
any:
- resources:
kinds:
- Pod
selector:
matchLabels:
pod-security.kubernetes.io/enforce: "!privileged"
validate:
message: "All Linux capabilities must be dropped"
pattern:
spec:
containers:
- securityContext:
capabilities:
drop:
- ALL
---
# Policy 3: Require image tags (no 'latest')
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: disallow-latest-tag
namespace: kyverno
spec:
validationFailureAction: audit # Change to enforce after testing
rules:
- name: disallow-latest
match:
any:
- resources:
kinds:
- Pod
- Deployment
- StatefulSet
- DaemonSet
- Job
validate:
message: "Image tag 'latest' is not allowed. Use explicit version tags."
pattern:
spec:
=(template):
spec:
containers:
- image: "!*:latest"
=(initContainers):
- image: "!*:latest"
---
# Policy 4: Restrict images to trusted registries
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: restrict-registries
namespace: kyverno
spec:
validationFailureAction: audit
rules:
- name: trusted-registries
match:
any:
- resources:
kinds:
- Pod
- Deployment
- StatefulSet
- DaemonSet
- Job
selector:
matchLabels:
pod-security.kubernetes.io/enforce: "!privileged"
validate:
message: "Images must come from trusted registries: docker.io, ghcr.io, quay.io, k8s.gcr.io, registry.k8s.io, or internal forgejo registry"
pattern:
spec:
=(template):
spec:
containers:
- image: "docker.io/* | ghcr.io/* | quay.io/* | k8s.gcr.io/* | registry.k8s.io/* | forgejo.riotpiao.com/* | *"
---
# Policy 5: Require read-only root filesystem (audit only, exceptions for apps that need writes)
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: require-readonly-filesystem
namespace: kyverno
spec:
validationFailureAction: audit
rules:
- name: check-readOnlyRootFilesystem
match:
any:
- resources:
kinds:
- Pod
selector:
matchLabels:
pod-security.kubernetes.io/enforce: "!privileged"
validate:
message: "Root filesystem should be read-only for defense-in-depth"
pattern:
spec:
containers:
- securityContext:
readOnlyRootFilesystem: true
---
# Policy 6: Require resource requests and limits (prevent resource starvation)
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: require-resource-limits
namespace: kyverno
spec:
validationFailureAction: audit
rules:
- name: check-resources
match:
any:
- resources:
kinds:
- Pod
- Deployment
- StatefulSet
- DaemonSet
excludeResources:
namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: "kyverno|kube-system|kube-node-lease"
validate:
message: "CPU and memory requests and limits are required"
pattern:
spec:
=(template):
spec:
containers:
- resources:
requests:
memory: "?*"
cpu: "?*"
limits:
memory: "?*"
cpu: "?*"
---
# Policy 7: Require securityContext on all containers
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: require-security-context
namespace: kyverno
spec:
validationFailureAction: audit
rules:
- name: check-securityContext
match:
any:
- resources:
kinds:
- Pod
selector:
matchLabels:
pod-security.kubernetes.io/enforce: "!privileged"
validate:
message: "securityContext must be defined"
pattern:
spec:
containers:
- securityContext: {}
@@ -64,22 +64,6 @@ gitea:
TYPE: redis
CONN_STR: redis://forgejo-redis.cicd.svc.cluster.local:6379/2
# Actions must be enabled globally, not just per-repo. Without this section
# app.ini carries no [actions] block at all and Forgejo never *creates* a
# workflow run — the API returns total_count: 0 for every repo even though
# each repo reports has_actions: true, the workflow file is on the default
# branch, and forgejo-runner has registered successfully with label
# [docker]. Registration does not require the dispatcher, so a healthy-
# looking runner sitting at "[poller 0] launched" with zero task pickups is
# the symptom of this being off, not of a bad workflow or a label mismatch.
#
# DEFAULT_ACTIONS_URL is left at its default (https://code.forgejo.org),
# which is where `uses: actions/checkout@v4` and friends resolve from. That
# requires egress from the runner; if it is ever blocked, pin the actions to
# local copies rather than turning this off.
actions:
ENABLED: true
# Persistence (shared storage for repos)
persistence:
enabled: true
@@ -94,18 +78,6 @@ ingress:
className: nginx
annotations:
cert-manager.io/cluster-issuer: letsencrypt-prod
# This chart Ingress and the hand-written one in
# k8s/bootstrap/ingress/ingress.yaml both claim forgejo.riotpiao.com.
# ingress-nginx breaks the tie by oldest creationTimestamp, and the chart's
# is older, so it is the one actually serving — the annotations on the other
# have never applied. Duplicate should be removed; until then these must
# live here or they do nothing.
#
# proxy-body-size 0 is required for the OCI registry: nginx defaults to 1m,
# so any image layer above that fails the push with 413.
nginx.ingress.kubernetes.io/proxy-body-size: "0"
nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"
nginx.ingress.kubernetes.io/proxy-send-timeout: "3600"
hosts:
- host: forgejo.riotpiao.com
paths:
@@ -1,220 +0,0 @@
# Forgejo OCI Registry Cleanup CronJob
# Deletes old image tags, keeping only the latest N versions per repository.
# Useful for retiring old builds when new versions are pushed.
---
apiVersion: v1
kind: ConfigMap
metadata:
name: forgejo-registry-cleanup-script
namespace: cicd
data:
cleanup.sh: |
#!/bin/bash
set -eo pipefail
# Configuration
REGISTRY_HOST="${REGISTRY_HOST:-forgejo.riotpiao.com}"
REGISTRY_URL="https://${REGISTRY_HOST}"
KEEP_VERSIONS="${KEEP_VERSIONS:-3}" # Keep latest N versions per image
DRY_RUN="${DRY_RUN:-false}"
# Load credentials from mounted secret
REGISTRY_USER="${REGISTRY_USER:-_json_key}"
REGISTRY_PASS="$(cat /etc/registry-secret/password 2>/dev/null || echo '')"
log() {
echo "[$(date +'%Y-%m-%d %H:%M:%S')] $*"
}
error() {
echo "[$(date +'%Y-%m-%d %H:%M:%S')] ERROR: $*" >&2
return 1
}
# Verify crane is available
if ! command -v crane &> /dev/null; then
error "crane not found. Install google/crane image for registry operations."
exit 1
fi
log "Starting Forgejo registry cleanup"
log "Registry: $REGISTRY_URL"
log "Keep versions: $KEEP_VERSIONS per image"
log "Dry run: $DRY_RUN"
# Authenticate crane with registry
if [ -n "$REGISTRY_PASS" ]; then
echo "$REGISTRY_PASS" | crane auth login "$REGISTRY_HOST" -u "$REGISTRY_USER" --password-stdin
log "Authenticated to $REGISTRY_HOST"
fi
# List all repositories (catalog)
# Note: This endpoint requires the registry to expose /v2/_catalog (standard OCI)
# If not available, images must be discovered another way
CATALOG=$(curl -s -u "${REGISTRY_USER}:${REGISTRY_PASS}" \
"${REGISTRY_URL}/v2/_catalog" | grep -o '"repositories":\[\K[^]]*' || echo '')
if [ -z "$CATALOG" ]; then
log "WARNING: Could not retrieve catalog from ${REGISTRY_URL}/v2/_catalog"
log "Registry may not expose _catalog endpoint or credentials invalid"
exit 0
fi
# Parse repositories from catalog JSON
REPOS=$(echo "$CATALOG" | grep -o '"[^"]*"' | tr -d '"')
TOTAL_DELETED=0
for REPO in $REPOS; do
log "Processing repository: $REPO"
IMAGE="${REGISTRY_HOST}/${REPO}"
# Get all tags for this image
TAGS=$(crane ls "$IMAGE" 2>/dev/null || echo "")
if [ -z "$TAGS" ]; then
log " No tags found for $REPO (or access denied)"
continue
fi
# Filter out 'latest' tag and sort by creation time (newer first)
# Note: crane doesn't provide direct date sorting; we use the order returned
# Assumption: tags are returned newest first (not always true)
TAG_COUNT=$(echo "$TAGS" | wc -l)
if [ "$TAG_COUNT" -le "$KEEP_VERSIONS" ]; then
log " $REPO: $TAG_COUNT tags total, keeping all (≤ $KEEP_VERSIONS)"
continue
fi
# Get tags to delete (all except the first N)
TAGS_TO_DELETE=$(echo "$TAGS" | tail -n +$((KEEP_VERSIONS + 1)))
for TAG in $TAGS_TO_DELETE; do
FULL_IMAGE="${IMAGE}:${TAG}"
DELETED_SIZE="0"
if [ "$DRY_RUN" = "true" ]; then
log " [DRY RUN] Would delete: $FULL_IMAGE"
else
if crane delete "$FULL_IMAGE" 2>&1; then
log " Deleted: $FULL_IMAGE"
((TOTAL_DELETED++))
else
error "Failed to delete $FULL_IMAGE (may already be deleted)"
fi
fi
done
done
log "Cleanup complete. Total images deleted: $TOTAL_DELETED"
---
apiVersion: batch/v1
kind: CronJob
metadata:
name: forgejo-registry-cleanup
namespace: cicd
labels:
app: forgejo-registry-cleanup
spec:
# Run at 2 AM UTC every day (adjust as needed)
schedule: "0 2 * * *"
# Keep last 3 successful/failed runs for debugging
successfulJobsHistoryLimit: 3
failedJobsHistoryLimit: 3
# Suspend if needed (set to false to enable)
suspend: false
jobTemplate:
spec:
# Cleanup jobs after 6 hours whether they succeeded or failed
ttlSecondsAfterFinished: 21600
template:
metadata:
labels:
app: forgejo-registry-cleanup
spec:
serviceAccountName: forgejo-registry-cleanup
restartPolicy: OnFailure
containers:
- name: cleanup
# Use google/crane for registry operations
image: gcr.io/go-containerregistry/crane:latest
imagePullPolicy: IfNotPresent
env:
- name: REGISTRY_HOST
value: "forgejo.riotpiao.com"
- name: KEEP_VERSIONS
value: "3" # Keep 3 latest versions
- name: DRY_RUN
value: "false" # Set to "true" for dry-run mode
- name: REGISTRY_USER
valueFrom:
secretKeyRef:
name: forgejo-registry-token
key: username
optional: true
volumeMounts:
- name: script
mountPath: /scripts
- name: registry-secret
mountPath: /etc/registry-secret
readOnly: true
# Run cleanup script via entrypoint override
command:
- /bin/sh
- -c
- |
# Install bash and curl if needed
apk add --no-cache bash curl
chmod +x /scripts/cleanup.sh
/scripts/cleanup.sh
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
cpu: 500m
memory: 512Mi
# Safety: kill after 30 min (prevents hanging on large registries)
securityContext:
runAsNonRoot: true
runAsUser: 65534
allowPrivilegeEscalation: false
readOnlyRootFilesystem: false
capabilities:
drop:
- ALL
volumes:
- name: script
configMap:
name: forgejo-registry-cleanup-script
defaultMode: 0755
- name: registry-secret
secret:
secretName: forgejo-registry-token
optional: true
---
apiVersion: v1
kind: ServiceAccount
metadata:
name: forgejo-registry-cleanup
namespace: cicd
---
# No RBAC needed: this pod only talks to the registry API (external service)
# If expanded to manage in-cluster resources, add Role/RoleBinding here
@@ -1,35 +0,0 @@
# ArgoCD Image Updater configuration
# Watches Forgejo registry and updates ArgoCD Applications with new image tags
config:
# Registry configuration - Forgejo allows anonymous pulls
registries:
- name: forgejo
api_url: https://forgejo.riotpiao.com
prefix: forgejo.riotpiao.com
default: true
insecure: false
# Log level
logLevel: debug
# ArgoCD API server
argocd:
grpcWeb: true
serverAddress: argocd-server.argocd.svc.cluster.local
insecure: true
plaintext: true
# Extra environment variables
extraEnv:
- name: ARGOCD_GRPC_WEB
value: "true"
# Resources
resources:
requests:
cpu: 50m
memory: 64Mi
limits:
cpu: 200m
memory: 128Mi
@@ -1,137 +0,0 @@
# Cluster-wide cleanup of stale failed/completed Jobs and Pods.
# Runs daily at 04:00 UTC. Deletes:
# - Failed Jobs older than 24h (any namespace)
# - Completed Jobs older than 72h with no owning CronJob
# - Orphan pods in Error/Failed/Evicted state older than 1h
#
# CronJob-owned Jobs are managed by failedJobsHistoryLimit/successfulJobsHistoryLimit,
# but standalone Jobs (helm hooks, one-off runs, longhorn maintenance) have no TTL
# and linger forever.
apiVersion: v1
kind: ServiceAccount
metadata:
name: stale-job-cleanup
namespace: kube-system
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: stale-job-cleanup
rules:
- apiGroups: ["batch"]
resources: ["jobs"]
verbs: ["get", "list", "delete"]
- apiGroups: [""]
resources: ["pods"]
verbs: ["get", "list", "delete"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: stale-job-cleanup
subjects:
- kind: ServiceAccount
name: stale-job-cleanup
namespace: kube-system
roleRef:
kind: ClusterRole
name: stale-job-cleanup
apiGroup: rbac.authorization.k8s.io
---
apiVersion: batch/v1
kind: CronJob
metadata:
name: stale-job-cleanup
namespace: kube-system
labels:
app: stale-job-cleanup
spec:
schedule: "0 4 * * *"
concurrencyPolicy: Forbid
successfulJobsHistoryLimit: 3
failedJobsHistoryLimit: 3
jobTemplate:
spec:
ttlSecondsAfterFinished: 86400 # self-cleanup after 24h
backoffLimit: 1
activeDeadlineSeconds: 300
template:
spec:
serviceAccountName: stale-job-cleanup
restartPolicy: Never
tolerations:
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
containers:
- name: cleanup
image: alpine/k8s:1.31.0
command:
- sh
- -c
- |
set -e
NOW=$(date +%s)
echo "=== Cleaning failed Jobs older than 24h ==="
kubectl get jobs --all-namespaces -o json | \
jq -r '.items[] |
select(.status.conditions[]?.type == "Failed") |
select(.status.completionTime or .status.startTime) |
"\(.metadata.namespace) \(.metadata.name) \(.status.startTime // .status.completionTime // .metadata.creationTimestamp)"' | \
while read -r NS NAME TS; do
JOB_EPOCH=$(date -d "$TS" +%s 2>/dev/null || echo 0)
AGE_H=$(( (NOW - JOB_EPOCH) / 3600 ))
if [ "$AGE_H" -ge 24 ]; then
echo "[delete] $NS/$NAME (failed ${AGE_H}h ago)"
kubectl delete job "$NAME" -n "$NS" --cascade=foreground 2>/dev/null || true
fi
done
echo ""
echo "=== Cleaning completed standalone Jobs older than 72h ==="
kubectl get jobs --all-namespaces -o json | \
jq -r '.items[] |
select(.status.succeeded >= 1) |
select((.metadata.ownerReferences // []) | length == 0) |
"\(.metadata.namespace) \(.metadata.name) \(.status.completionTime // .metadata.creationTimestamp)"' | \
while read -r NS NAME TS; do
JOB_EPOCH=$(date -d "$TS" +%s 2>/dev/null || echo 0)
AGE_H=$(( (NOW - JOB_EPOCH) / 3600 ))
if [ "$AGE_H" -ge 72 ]; then
echo "[delete] $NS/$NAME (completed ${AGE_H}h ago, no owner)"
kubectl delete job "$NAME" -n "$NS" --cascade=foreground 2>/dev/null || true
fi
done
echo ""
echo "=== Cleaning orphan Error/Failed/Evicted pods older than 1h ==="
# Evicted pods show as Failed with reason Evicted
kubectl get pods --all-namespaces -o json | \
jq -r '.items[] |
select(
.status.phase == "Failed" or
(.status.reason // "") == "Evicted" or
(.status.containerStatuses // [] | any(.state.terminated.reason == "Error"))
) |
select((.metadata.ownerReferences // []) | all(.kind != "Job")) |
"\(.metadata.namespace) \(.metadata.name) \(.metadata.creationTimestamp)"' | \
while read -r NS NAME TS; do
POD_EPOCH=$(date -d "$TS" +%s 2>/dev/null || echo 0)
AGE_H=$(( (NOW - POD_EPOCH) / 3600 ))
if [ "$AGE_H" -ge 1 ]; then
echo "[delete] $NS/$NAME (error/evicted ${AGE_H}h ago)"
kubectl delete pod "$NAME" -n "$NS" --force 2>/dev/null || true
fi
done
echo ""
echo "Cleanup complete"
resources:
requests:
cpu: 50m
memory: 64Mi
limits:
cpu: 250m
memory: 128Mi
+1 -6
View File
@@ -1,13 +1,8 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
# Dedicated per-app CNPG clusters. NO top-level `namespace:` — each Cluster
# carries its own ns (iam / temporal / poimen / paperless); a transformer would
# wrongly collapse them. poimen ns created by poimen-root app, memory-db
# deployed into it. paperless ns declared in namespaces.yaml above.
# carries its own ns (iam / temporal); a transformer would wrongly collapse them.
resources:
- namespaces.yaml
- authentik-db.yaml
- temporal-db.yaml
- memory-db.yaml
- paperless-db.yaml
- obsidian-vault-pvc.yaml
-36
View File
@@ -1,36 +0,0 @@
# Dedicated CNPG Postgres for Poimen Memory (GitOps, wave 2).
# Uses default longhorn storage class (3 replicas, dataLocality disabled).
# CNPG generates secret `memory-db-app` + service `memory-db-rw` in ns poimen.
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: memory-db
namespace: poimen
annotations:
argocd.argoproj.io/sync-options: SkipDryRunOnMissingResource=true
spec:
instances: 2
imageName: ghcr.io/cloudnative-pg/postgresql:16.2
bootstrap:
initdb:
database: memory
owner: app
encoding: UTF8
localeCollate: C
localeCType: C
postInitApplicationSQL:
- "CREATE EXTENSION vector;"
enableSuperuserAccess: false
resources:
requests: { memory: "512Mi", cpu: "250m" }
limits: { memory: "2Gi", cpu: "1" }
storage:
size: 20Gi
storageClass: longhorn
affinity:
podAntiAffinityType: preferred
topologyKey: kubernetes.io/hostname
tolerations:
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
+1 -15
View File
@@ -1,7 +1,6 @@
# DB clusters are wave 2 — their namespaces must exist first (their apps that
# would CreateNamespace run later, w3/w8). Declared here so the databases App
# creates them. authentik/vault/temporal/paperless CreateNamespace=true then
# no-ops. poimen namespace created by poimen-root app (wave 7).
# creates them. authentik/vault/temporal CreateNamespace=true then no-ops.
apiVersion: v1
kind: Namespace
metadata:
@@ -11,16 +10,3 @@ apiVersion: v1
kind: Namespace
metadata:
name: temporal
---
apiVersion: v1
kind: Namespace
metadata:
name: paperless
---
# Needed here (not just immich's own CreateNamespace=true at wave 8) because
# k8s/infra/iam's PostSync job (wave 3) has a RoleBinding targeting this
# namespace - same ordering reason as paperless above.
apiVersion: v1
kind: Namespace
metadata:
name: immich
@@ -1,18 +0,0 @@
---
# Obsidian vault PVC — shared storage for REST API + UI pods
# ReadWriteMany so both obsidian-server and obsidian-ui can mount simultaneously
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: obsidian-vault
namespace: poimen
labels:
app.kubernetes.io/name: obsidian-server
app.kubernetes.io/part-of: poimen-memory
spec:
accessModes:
- ReadWriteMany
storageClassName: longhorn
resources:
requests:
storage: 10Gi
-36
View File
@@ -1,36 +0,0 @@
# Dedicated CNPG Postgres for paperless-ngx (GitOps, wave 2 — before the
# paperless app at w8). Same recipe as memory-db: default longhorn storage
# class (3 replicas), 2 instances, 20Gi.
# CNPG generates secret `paperless-db-app` + service `paperless-db-rw` in ns
# paperless; the paperless Deployment reads them locally.
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: paperless-db
namespace: paperless
annotations:
argocd.argoproj.io/sync-options: SkipDryRunOnMissingResource=true
spec:
instances: 2
imageName: ghcr.io/cloudnative-pg/postgresql:16.2
bootstrap:
initdb:
database: paperless
owner: app
encoding: UTF8
localeCollate: C
localeCType: C
enableSuperuserAccess: false
resources:
requests: { memory: "512Mi", cpu: "250m" }
limits: { memory: "2Gi", cpu: "1" }
storage:
size: 20Gi
storageClass: longhorn
affinity:
podAntiAffinityType: preferred
topologyKey: kubernetes.io/hostname
tolerations:
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
@@ -1,38 +0,0 @@
# act_runner (the forgejo-runner binary) ships no config.yaml by default, so
# `forgejo-runner daemon` runs on its hardcoded defaults -- notably
# container.valid_volumes: [] ("if the sequence is empty, no volumes can be
# mounted"). Confirmed via `forgejo-runner generate-config` on this exact
# image (code.forgejo.org/forgejo/runner:6) and by running the daemon against
# a minimal override locally: a job container that requests any bind mount
# (e.g. the dind mTLS certs at /docker-certs/client, needed for
# `docker login`/build/push steps) is rejected outright with no default
# config in place.
#
# Scoped narrowly to exactly the certs path, read-only. Not a wildcard
# (valid_volumes: ['**']) -- that would let any workflow in any repo this
# runner serves bind-mount arbitrary paths off the runner pod's filesystem
# into a job container, which is a real widening of the CI trust boundary,
# not just a convenience.
#
# network: host is also only settable here, not per-workflow. A workflow's
# `container.options: --network host` is silently ignored -- confirmed live:
# every job's actual `docker create` call logged
# `network="FORGEJO-ACTIONS-TASK-N_..."`, an auto-generated per-job bridge,
# regardless of that options string. On that isolated bridge, DOCKER_HOST=
# tcp://localhost:2376 resolves to the job container itself (no daemon there),
# not to the dind sidecar, so any docker command that actually needs the
# daemon (build, push -- anything past docker login, which only talks to the
# registry over the network and never touches DOCKER_HOST) fails with "Cannot
# connect to the Docker daemon". host mode puts every job container in dind's
# own network namespace instead, where the daemon really is listening.
apiVersion: v1
kind: ConfigMap
metadata:
name: {{ .Release.Name }}-config
namespace: {{ .Release.Namespace }}
data:
config.yaml: |
container:
valid_volumes:
- /docker-certs/client
network: host
@@ -56,7 +56,7 @@ spec:
containers:
- name: runner
image: {{ .Values.runner.image.repository }}:{{ .Values.runner.image.tag }}
command: ["sh", "-c", "forgejo-runner daemon --config /etc/forgejo-runner/config.yaml"]
command: ["sh", "-c", "forgejo-runner daemon"]
workingDir: /data
env:
- name: DOCKER_HOST
@@ -73,9 +73,6 @@ spec:
- name: homelab-ca
mountPath: /etc/ssl/certs/homelab-ca.pem
subPath: ca.crt
- name: runner-config
mountPath: /etc/forgejo-runner
readOnly: true
resources:
{{- toYaml .Values.runner.resources | nindent 12 }}
@@ -94,23 +91,16 @@ spec:
- name: homelab-ca
mountPath: /etc/ssl/certs/homelab-ca.pem
subPath: ca.crt
# dockerd resolves per-registry CAs from /etc/docker/certs.d/<host>/
# before falling back to the system pool. Mounting it here is what
# makes `docker push forgejo.riotpiao.com/...` trust the homelab CA
# rather than failing x509: signed by unknown authority.
- name: homelab-ca
mountPath: /etc/docker/certs.d/forgejo.riotpiao.com/ca.crt
subPath: ca.crt
resources:
{{- toYaml .Values.dind.resources | nindent 12 }}
volumes:
- name: runner-data
persistentVolumeClaim:
claimName: {{ .Release.Name }}-reg
claimName: runner-reg
- name: dind-storage
persistentVolumeClaim:
claimName: {{ .Release.Name }}-dind
claimName: runner-dind
- name: docker-certs
emptyDir: {} # DinD regenerates mTLS certs on each start
- name: homelab-ca
@@ -118,6 +108,3 @@ spec:
# The volumeMounts use subPath: ca.crt to project the single cert file.
configMap:
name: homelab-ca
- name: runner-config
configMap:
name: {{ .Release.Name }}-config
@@ -1,152 +0,0 @@
{{- if .Values.gc.enabled }}
# Garbage-collects DinD Docker images/volumes/build-cache and actcache across
# ALL forgejo-runner pods. Prevents PVC fill-up that breaks CI runs.
# Only rendered once (enable in default values.yaml, disable in per-runner overrides).
apiVersion: v1
kind: ServiceAccount
metadata:
name: runner-gc
namespace: {{ .Release.Namespace }}
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: runner-gc
namespace: {{ .Release.Namespace }}
rules:
- apiGroups: [""]
resources: ["pods"]
verbs: ["get", "list"]
- apiGroups: [""]
resources: ["pods/exec"]
verbs: ["create"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: runner-gc
namespace: {{ .Release.Namespace }}
subjects:
- kind: ServiceAccount
name: runner-gc
namespace: {{ .Release.Namespace }}
roleRef:
kind: Role
name: runner-gc
apiGroup: rbac.authorization.k8s.io
---
apiVersion: batch/v1
kind: CronJob
metadata:
name: forgejo-runner-gc
namespace: {{ .Release.Namespace }}
labels:
app: forgejo-runner-gc
spec:
schedule: {{ .Values.gc.schedule | quote }}
concurrencyPolicy: Forbid
successfulJobsHistoryLimit: 3
failedJobsHistoryLimit: 3
jobTemplate:
spec:
backoffLimit: 1
activeDeadlineSeconds: 900
template:
spec:
serviceAccountName: runner-gc
restartPolicy: Never
tolerations:
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
containers:
- name: gc
image: {{ .Values.gc.image }}
command:
- sh
- -c
- |
set -e
# Iterate all forgejo-runner pods (golang, rust, node)
PODS=$(kubectl -n {{ .Release.Namespace }} get pod \
-l app -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.metadata.labels.app}{"\n"}{end}' \
| grep 'forgejo-runner-' | awk '{print $1}')
if [ -z "$PODS" ]; then
echo "no forgejo-runner pods found, skipping"
exit 0
fi
for POD in $PODS; do
echo "===== $POD ====="
# 1. Docker image prune (DinD sidecar)
echo "[docker] before:"
kubectl -n {{ .Release.Namespace }} exec "$POD" -c dind -- docker system df 2>/dev/null || true
echo "[docker] pruning non-latest images older than {{ .Values.gc.pruneAge }}..."
# Keep :latest tagged images, delete all others older than pruneAge.
# docker image prune can't filter by tag, so we list and selectively rmi.
kubectl -n {{ .Release.Namespace }} exec "$POD" -c dind -- \
sh -c '
# Remove dangling (untagged) images older than {{ .Values.gc.pruneAge }}
docker image prune -f --filter "until={{ .Values.gc.pruneAge }}" 2>/dev/null
# Remove tagged non-latest images older than {{ .Values.gc.pruneAge }}
CUTOFF=$(date -d "-{{ .Values.gc.pruneAgeHours }} hours" +%s 2>/dev/null || date -v-{{ .Values.gc.pruneAgeHours }}H +%s)
docker images --format "{{"{{"}} .Repository {{"}}"}}:{{"{{"}} .Tag {{"}}"}} {{"{{"}} .CreatedAt {{"}}"}}" | while read -r IMAGE_TAG CREATED_REST; do
TAG=$(echo "$IMAGE_TAG" | rev | cut -d: -f1 | rev)
# Skip latest-tagged images
if [ "$TAG" = "latest" ]; then
echo "[keep] $IMAGE_TAG (latest)"
continue
fi
# Check image age via inspect
CREATED_TS=$(docker inspect --format="{{"{{"}} .Created {{"}}"}}" "$IMAGE_TAG" 2>/dev/null | head -1)
if [ -z "$CREATED_TS" ]; then continue; fi
IMAGE_EPOCH=$(date -d "$CREATED_TS" +%s 2>/dev/null || date -jf "%Y-%m-%dT%H:%M:%S" "$(echo $CREATED_TS | cut -dT -f1-2 | cut -d. -f1)" +%s 2>/dev/null || echo 0)
if [ "$IMAGE_EPOCH" -lt "$CUTOFF" ] 2>/dev/null; then
echo "[delete] $IMAGE_TAG (older than {{ .Values.gc.pruneAge }})"
docker rmi -f "$IMAGE_TAG" 2>/dev/null || true
else
echo "[keep] $IMAGE_TAG (recent)"
fi
done
' 2>/dev/null || true
echo "[docker] pruning build cache unused >{{ .Values.gc.pruneAge }}..."
kubectl -n {{ .Release.Namespace }} exec "$POD" -c dind -- \
docker builder prune -af --filter "until={{ .Values.gc.pruneAge }}" 2>/dev/null || true
echo "[docker] pruning dangling volumes..."
kubectl -n {{ .Release.Namespace }} exec "$POD" -c dind -- \
docker volume prune -af 2>/dev/null || true
echo "[docker] after:"
kubectl -n {{ .Release.Namespace }} exec "$POD" -c dind -- docker system df 2>/dev/null || true
# 2. Actcache cleanup (runner container)
echo "[actcache] cleaning incomplete and stale cache entries..."
kubectl -n {{ .Release.Namespace }} exec "$POD" -c runner -- \
sh -c '
# Delete incomplete/partial cache uploads immediately (tmp dirs)
find /data/.cache/actcache/cache -name "tmp" -type d -exec rm -rf {} + 2>/dev/null || true
# Delete cache entries not accessed in last {{ .Values.gc.actcacheMaxAgeDays }} day(s)
find /data/.cache/actcache/cache -type f -mtime +{{ .Values.gc.actcacheMaxAgeDays }} -delete 2>/dev/null || true
# Clean up empty directories
find /data/.cache/actcache/cache -type d -empty -delete 2>/dev/null || true
echo "actcache size: $(du -sh /data/.cache/actcache/cache 2>/dev/null | cut -f1)"
' || true
echo ""
done
echo "GC complete"
resources:
requests:
cpu: 50m
memory: 64Mi
limits:
cpu: 250m
memory: 128Mi
{{- end }}
@@ -29,23 +29,3 @@ spec:
except:
- 192.168.1.0/24
- 10.244.0.0/16
# ingress-nginx, which is how forgejo.riotpiao.com resolves (CoreDNS
# rewrites that name to ingress-nginx-controller.ingress-nginx.svc).
# Image pushes go to that name so the tag matches what containerd pulls
# on the nodes; without this the whole /24 and pod CIDR are denied above
# and `docker push`/`docker login` hang until they time out.
#
# Was an ipBlock pinned to the ingress-nginx LoadBalancer's LAN IP. That
# stopped matching once DNS started resolving the name to the Service's
# ClusterIP instead of the LB IP: Cilium enforces egress against the
# post-DNAT pod IP, which falls inside the 10.244.0.0/16 exclusion above,
# so every request silently hung rather than erroring. A namespaceSelector
# follows the Service wherever it resolves and needs no IP to stay in
# sync with -- same pattern as the kube-system DNS rule above.
- to:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: ingress-nginx
ports:
- protocol: TCP
port: 443

Some files were not shown because too many files have changed in this diff Show More