[GPU] KV Cache VRAM monitoring + PagedAttention #34

Open
opened 2026-09-11 01:27:11 +00:00 by poimen · 0 comments
Member

Fast GPU overwhelms VRAM during high concurrency. reasoning.yaml uses gpu-memory-utilization=0.90 but no runtime KV cache monitoring. Expose vLLM /metrics, add Prometheus scrape + Grafana panel for KV cache usage. Files: k8s/apps/llm-serving/reasoning.yaml, k8s/infra/monitoring/

Fast GPU overwhelms VRAM during high concurrency. reasoning.yaml uses gpu-memory-utilization=0.90 but no runtime KV cache monitoring. Expose vLLM /metrics, add Prometheus scrape + Grafana panel for KV cache usage. Files: k8s/apps/llm-serving/reasoning.yaml, k8s/infra/monitoring/
poimen added this to the LLM Production Hardening milestone 2026-09-11 01:27:11 +00:00
Sign in to join this conversation.