- Increase PVC from 20Gi to 100Gi (talos-cp-1 disk full)
- Move from az-a (talos-cp-1) to worker-1 (648GB available)
- Increase retention from 15d to 30d (95GB limit)
- Fixes Prometheus pod restarts due to disk pressure
Add ServiceMonitor for llm-serving namespace. Wire prometheus.io annotations to all vLLM pods (reasoning, ornith, embeddings, reranker). Scrape /metrics@8080 every 30s with proper relabeling.
Ingress api/api now backs onto api-gateway:8080; the kong Application, its
Helm values, plugins and llm-routes are removed. Gateway image v0.0.0 is in
the Forgejo registry and the pull secret is in the api namespace.