[Observability] HPA driven by GPU KV Cache usage #39

Open
opened 2026-09-11 01:27:33 +00:00 by poimen · 0 comments
Member

Standard K8s autoscalers scale on CPU/RAM not AI load. Hook KServe InferenceService into HPA driven by KV cache utilization metric from vLLM Prometheus. Scale ornith pods when cache >80%. Files: k8s/apps/llm-serving/ornith.yaml, new HPA manifest

Standard K8s autoscalers scale on CPU/RAM not AI load. Hook KServe InferenceService into HPA driven by KV cache utilization metric from vLLM Prometheus. Scale ornith pods when cache >80%. Files: k8s/apps/llm-serving/ornith.yaml, new HPA manifest
poimen added this to the LLM Production Hardening milestone 2026-09-11 01:27:33 +00:00
Sign in to join this conversation.