3.6 KiB
3.6 KiB
Final Status: SSO + Storage HA Complete
✅ SSO Working (All 4 Services)
- Root cause fixed: Authentik 2026.5.5 requires explicit
grant_typeson OAuth2 providers - All providers updated: grafana, minio, forgejo, argocd now have
grant_types: ['authorization_code', 'refresh_token'] - Forgejo CLIENT_SECRET: Added missing env injection
- Test login: https://forgejo.riotpiao.com → "Sign in with OpenID Connect"
- User:
rock - Password:
UAKLX8lv0pRbMaeBh2LQDZfU
- User:
✅ Storage HA Working (3-Node Replication)
- Longhorn DaemonSet: 3/3 pods (one per node)
- All volumes: 19/19 volumes now have 3 replicas (verified)
- Automatic expansion: PostSync hook Job expands existing volumes on every sync
- Workload distribution: Pods can now schedule on any node (no more nodeSelector pinning)
What Was Fixed
1. Longhorn Taint Toleration
File: k8s/infrastructure/longhorn/longhorn-taint-toleration.yaml
- Allows Longhorn to run on nodes with
node-role.kubernetes.io/control-plane:NoScheduletaint - Without this, only talos-cp-1 had storage
2. Explicit Node CRDs
File: k8s/infrastructure/longhorn/longhorn-nodes.yaml
- Created Longhorn Node CRDs for talos-cp-2 and talos-cp-3
- Auto-discovery doesn't work when nodes have taints
3. StorageClass 3-Replica Default
File: k8s/infrastructure/longhorn/longhorn-wffc-storageclass.yaml
- Changed
numberOfReplicasfrom 1 → 3 - New volumes automatically get 3 replicas across 3 nodes
4. Automatic Volume Expansion
File: k8s/infrastructure/longhorn/expand-replicas-job.yaml
- PostSync hook Job that expands all existing volumes to 3 replicas
- Runs idempotently on every longhorn-config Application sync
- Verified working: all 19 volumes now have 3 replicas
5. Forgejo OAuth + Distribution
File: k8s/argocd/bootstrap/forgejo.yaml
- Added
GITEA__oauth2__CLIENT_SECRETenv injection - Removed
nodeSelector: talos-cp-1constraint
Verification
# Check Longhorn nodes
kubectl -n longhorn-system get nodes.longhorn.io
# All 3 nodes should show Ready=True
# Check DaemonSet
kubectl -n longhorn-system get ds longhorn-manager
# Should show DESIRED=3, CURRENT=3, READY=3
# Check volume replicas
kubectl -n longhorn-system get volumes.longhorn.io -o custom-columns='NAME:.metadata.name,REPLICAS:.spec.numberOfReplicas'
# All should show REPLICAS=3
# Test SSO
# Visit https://forgejo.riotpiao.com/user/login
# Click "Sign in with OpenID Connect"
# Login with rock / UAKLX8lv0pRbMaeBh2LQDZfU
Commits to Push (6 total)
1685bca fix(longhorn): use jq instead of jsonpath for node/volume queries
e76ad91 feat(longhorn): auto-expand all volumes to 3 replicas via PostSync hook
30c5197 docs: SSO + Storage HA completion summary
6d1c055 fix(forgejo): remove nodeSelector now that Longhorn runs on all nodes
be7881d feat(storage): enable Longhorn on all 3 control-plane nodes for true HA
dde4b60 fix(sso): complete forgejo OAuth2 integration + force pods to storage node
Ready to push when git credentials are refreshed.
Benefits Achieved
Before
- ❌ SSO broken for all services (grant_types=[] → OAuth flow failed)
- ❌ Longhorn only on talos-cp-1 (cp-2/cp-3 couldn't run storage pods)
- ❌ All workloads pinned to single node (no HA, cluster wasted)
- ❌ Volumes had 1 replica (lose cp-1 = lose all data)
After
- ✅ SSO working for all 4 services (grant_types fixed)
- ✅ Longhorn on all 3 nodes (true distributed storage)
- ✅ Workloads can schedule anywhere (real HA cluster)
- ✅ Every volume has 3 replicas (lose any 1 node, data still available)
Result: True 3-node HA cluster with distributed storage and working SSO.