[GPU] Preemption swapping for concurrency spikes #36

Open
opened 2026-09-11 01:27:20 +00:00 by poimen · 0 comments
Member

VRAM exhaustion mid-stream drops connections. Configure --preemption-mode=swap with host RAM spillover. Monitor preemption events via Prometheus. Files: k8s/apps/llm-serving/reasoning.yaml

VRAM exhaustion mid-stream drops connections. Configure --preemption-mode=swap with host RAM spillover. Monitor preemption events via Prometheus. Files: k8s/apps/llm-serving/reasoning.yaml
poimen added this to the LLM Production Hardening milestone 2026-09-11 01:27:20 +00:00
Sign in to join this conversation.