## GPU Rebalance (4× V100 32GB) | Pod | Before | After | |-----|--------|-------| | reasoning (PP=2) | 2 GPU | 2 GPU | | ornith | 2 GPU (2 replicas) | 1 GPU (1 replica) | | comfyui | — | 1 GPU (**new**) | | qwen-cpu | — | CPU on cp-2 (**new**) | | embeddings/reranker | CPU | CPU | ## Changes - `ornith.yaml`: scale 2→1, remove qwen2.5 co-loading, MAX_LOADED_MODELS=1 - `qwen-cpu.yaml`: new Ollama deployment on talos-cp-2 (144GB RAM), 5Gi PVC - `k8s/apps/comfyui/`: new ComfyUI deployment (1 GPU, 50Gi model PVC, ingress) - `58-comfyui.yaml`: ArgoCD Application (wave 8) Gateway route update in separate PR (homelab-frontend).Reviewed-on: #17 Co-authored-by: rock <[email protected]>
26 lines
721 B
YAML
26 lines
721 B
YAML
apiVersion: networking.k8s.io/v1
|
|
kind: Ingress
|
|
metadata:
|
|
name: comfyui
|
|
namespace: comfyui
|
|
annotations:
|
|
nginx.ingress.kubernetes.io/proxy-read-timeout: "600"
|
|
nginx.ingress.kubernetes.io/proxy-send-timeout: "600"
|
|
nginx.ingress.kubernetes.io/proxy-body-size: "0"
|
|
# WebSocket support for ComfyUI's live preview
|
|
nginx.ingress.kubernetes.io/proxy-http-version: "1.1"
|
|
nginx.ingress.kubernetes.io/proxy-set-headers: "Upgrade"
|
|
spec:
|
|
ingressClassName: nginx
|
|
rules:
|
|
- host: comfy.riotpiao.com
|
|
http:
|
|
paths:
|
|
- path: /
|
|
pathType: Prefix
|
|
backend:
|
|
service:
|
|
name: comfyui
|
|
port:
|
|
number: 80
|