Standard K8s autoscalers scale on CPU/RAM not AI load. Hook KServe InferenceService into HPA driven by KV cache utilization metric from vLLM Prometheus. Scale ornith pods when cache >80%. Files: k8s/apps/llm-serving/ornith.yaml, new HPA manifest
Standard K8s autoscalers scale on CPU/RAM not AI load. Hook KServe InferenceService into HPA driven by KV cache utilization metric from vLLM Prometheus. Scale ornith pods when cache >80%. Files: k8s/apps/llm-serving/ornith.yaml, new HPA manifest
poimen
added this to the LLM Production Hardening milestone 2026-09-11 01:27:33 +00:00
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Standard K8s autoscalers scale on CPU/RAM not AI load. Hook KServe InferenceService into HPA driven by KV cache utilization metric from vLLM Prometheus. Scale ornith pods when cache >80%. Files: k8s/apps/llm-serving/ornith.yaml, new HPA manifest