Fast GPU overwhelms VRAM during high concurrency. reasoning.yaml uses gpu-memory-utilization=0.90 but no runtime KV cache monitoring. Expose vLLM /metrics, add Prometheus scrape + Grafana panel for KV cache usage. Files: k8s/apps/llm-serving/reasoning.yaml, k8s/infra/monitoring/
No dependencies set.
The note is not visible to the blocked user.
Fast GPU overwhelms VRAM during high concurrency. reasoning.yaml uses gpu-memory-utilization=0.90 but no runtime KV cache monitoring. Expose vLLM /metrics, add Prometheus scrape + Grafana panel for KV cache usage. Files: k8s/apps/llm-serving/reasoning.yaml, k8s/infra/monitoring/