feat(storage): enable Longhorn on all 3 control-plane nodes for true HA

Changes:
- k8s/infrastructure/longhorn/longhorn-taint-toleration.yaml: new Setting
  to tolerate node-role.kubernetes.io/control-plane:NoSchedule taint,
  allowing Longhorn DaemonSet to run on cp-2/cp-3 (not just cp-1)
- k8s/infrastructure/longhorn/longhorn-nodes.yaml: explicit Node CRDs for
  talos-cp-2 and talos-cp-3 (auto-discovery doesn't work when nodes have
  taints; these define /var/lib/longhorn as the storage path)
- k8s/infrastructure/longhorn/longhorn-wffc-storageclass.yaml: bump
  numberOfReplicas from 1→3 (true HA: each volume gets 3 copies across
  3 nodes; if one node fails, 2 others still have the data)
- k8s/infrastructure/longhorn/kustomization.yaml: add new resources

Root cause: Longhorn was only running on talos-cp-1 (.213) because cp-2/cp-3
have the control-plane taint and Longhorn DaemonSet had no matching toleration.
Every workload with a PVC was forced to schedule on cp-1 (via nodeSelector or
implicit co-location with the storage), defeating the entire purpose of a 3-node
HA cluster.

With this fix:
- Longhorn manager runs on all 3 nodes
- Storage is replicated 3x (erasure-coded across nodes)
- Pods can schedule on any node without PVC attachment failures
- True HA: lose 1 node, cluster still serves all volumes
This commit is contained in:
Story Crater Bot
2026-08-18 15:08:03 -07:00
parent 86f5603063
commit e2fcfe1fa8
4 changed files with 58 additions and 15 deletions
@@ -1,17 +1,14 @@
# longhorn-wffc — Longhorn StorageClass with WaitForFirstConsumer binding.
#
# The chart's default `longhorn` SC uses Immediate binding: the PV binds before
# the pod is scheduled, so on this single-storage-node topology (only talos-cp-1
# runs Longhorn) the scheduler often places the pod on cp-2/cp-3 where the volume
# can't attach ("CSINode does not contain driver driver.longhorn.io").
# WaitForFirstConsumer defers PV binding until the pod is scheduled, ensuring the
# volume is provisioned on a node where the pod can actually run. Critical for HA:
# with 3-replica volumes spread across 3 nodes, the scheduler needs to see which
# nodes already have replicas before placing the pod, avoiding situations where
# the pod lands on a node that can't reach any replica.
#
# WaitForFirstConsumer defers binding until the pod is scheduled, so the volume
# is provisioned on the node the pod lands on — and with a single Longhorn node
# that co-locates pod + volume on cp-1 automatically. This is the new default;
# the chart's `longhorn` SC is demoted (see longhorn-values default-class=false).
#
# volumeBindingMode is immutable, so this is a distinct SC (not an edit of the
# chart's). Existing volumes stay on `longhorn`; new PVCs use this.
# numberOfReplicas=3 provides true HA: each volume has 3 copies across 3 nodes.
# If one node fails, the remaining 2 nodes still have the data and can serve it.
# volumeBindingMode is immutable, so this is a distinct SC from the chart's default.
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
@@ -23,7 +20,7 @@ allowVolumeExpansion: true
reclaimPolicy: Delete
volumeBindingMode: WaitForFirstConsumer
parameters:
numberOfReplicas: "1"
numberOfReplicas: "3"
staleReplicaTimeout: "30"
fromBackup: ""
dataLocality: "best-effort"