fix(longhorn): isolate GPU worker from general storage scheduling
- Remove expand-replicas-job (blindly forced all volumes to 3 replicas, ignoring StorageClass settings) - Add diskSelector: 'storage' to longhorn and longhorn-cnpg StorageClasses so replicas only land on CP nodes (cp-1, cp-2, cp-3) - Tag all CP node disks with 'storage' via PostSync job (disk names are runtime-discovered, can't hardcode in Node CRs) - Disable scheduling on worker-1 Node CR — only longhorn-llm-local (diskSelector: 'llm') can use it - worker-1 is GPU-only: llm-models and comfyui use dedicated SCs
This commit is contained in:
@@ -8,11 +8,11 @@ resources:
|
||||
- longhorn-servicemonitor.yaml
|
||||
- longhorn-taint-toleration.yaml
|
||||
- longhorn-nodes.yaml
|
||||
- expand-replicas-job.yaml
|
||||
- longhorn-tag-disks-job.yaml
|
||||
- patch-csi-tolerations-job.yaml
|
||||
- longhorn-add-disks-job.yaml # Add extra disks to talos-cp-2
|
||||
# Longhorn deployed via bootstrap script or Helm.
|
||||
# These manifests configure it: unified StorageClass (default, 3 replicas),
|
||||
# Prometheus ServiceMonitor, taint toleration for control-plane nodes, explicit
|
||||
# Node CRDs for cp-2/cp-3, CSI plugin tolerations, and a PostSync hook Job
|
||||
# that ensures all existing volumes have 3 replicas.
|
||||
# that configures storage for the cluster.
|
||||
|
||||
Reference in New Issue
Block a user