
rockandpoimen
39f5fe3683
feat: add ComfyUI, rebalance GPU allocation
GPU rebalance (4× V100 32GB on worker-1):
- reasoning: 2 GPU (unchanged, PP=2 for Qwen3-32B)
- ornith: 2→1 GPU (scale to 1 replica, ornith:35b only)
- comfyui: 0→1 GPU (new)
- embeddings/reranker: 0 GPU (CPU, unchanged)
qwen2.5:3b-instruct moved to CPU on talos-cp-2 (144GB RAM).
Separate Ollama deployment + 5Gi PVC, pulls model on first start.
Gateway config updated in homelab-frontend (separate commit).
Co-authored-by: poimen <[email protected]>
2026-09-08 18:26:07 -07:00
..
2026-08-25 11:20:26 -07:00
2026-08-25 11:20:26 -07:00
2026-08-31 11:43:04 -07:00
2026-09-02 20:31:09 -07:00
2026-08-18 15:08:04 -07:00
2026-09-04 13:51:26 -07:00
2026-08-31 15:02:15 -07:00
2026-08-25 11:20:26 -07:00
2026-09-07 13:01:56 -07:00
2026-08-25 11:20:26 -07:00
2026-09-05 01:10:51 -07:00
2026-09-03 08:28:38 -07:00
2026-09-07 18:19:46 -07:00
2026-08-25 11:20:26 -07:00
2026-09-08 18:26:07 -07:00
2026-09-07 18:19:46 -07:00
2026-08-21 20:55:20 -07:00
2026-09-07 18:19:46 -07:00