Major accomplishments from comprehensive cluster review: ## Storage HA (answering "are volumes replicated?") - Verified 3-node Longhorn HA: ALL 17 volumes have 3 replicas - Fixed CLAUDE.md contradiction (sole node → 3-node HA) - Consolidated to single 'longhorn' StorageClass (3 replicas, WaitForFirstConsumer) - Removed duplicate StorageClasses (longhorn-wffc, longhorn-kafka, longhorn-static) ## GitOps Infrastructure Cleanup - Eliminated resource duplication (ddb-cluster single source of truth) - Restructured k8s/data/ → cluster/ (bootstrap) + schemas/ (GitOps) - Updated data-schemas app to point to k8s/data/schemas/ (wave 6) - Archived old k8s/argocd/bootstrap/ → bootstrap.archived/ ## Bootstrap Dependencies Fixed - Added 05-wait-for-databases.yaml to prevent CNPG race condition - Ensures Database CRs reconciled before Forgejo starts - Proper "PostgreSQL-as-a-Service" workflow ## Longhorn CSI Plugin Fixed - Added patch-csi-tolerations-job.yaml (GitOps PostSync hook) - CSI plugin now runs on all 3 nodes (cp-1, cp-2, cp-3) - Fixes volume attachment on tainted control-plane nodes ## Live Migration (Zero Downtime) - Migrated 37 applications to ArgoCD app-of-apps management - Fixed Forgejo startup issues: * Service selector mismatch (app: forgejo → app: gitea) * Missing homelab-ca ConfigMap * Missing forgejo-oidc secret (temporary) * CNPG database creation timing ## Documentation (10 comprehensive files) - WHATS-NEXT.md - Daily GitOps workflow - MIGRATION-STATUS.md - Cluster health report - REVIEW-SUMMARY.md - Session overview - GITOPS-REBUILD-PLAN.md - Architecture reference - DDB-REVIEW.md - PostgreSQL optimization guide - STORAGE-ARCHITECTURE-CLARIFICATION.md - Storage HA investigation - BOOTSTRAP-DEPENDENCY-FIX.md - CNPG race condition fix - STORAGECLASS-CONSOLIDATION.md - Single StorageClass rationale - IMPLEMENTATION-CHECKLIST.md - Migration checklist - bootstrap.sh - Automated bootstrap script ## Cluster Status - ArgoCD: 4/4 pods running - DDB cluster: 3/3 instances healthy - Longhorn: 3/3 nodes, all CSI plugins running - Forgejo: Running, accessible at http://192.168.1.165:3000 - All 17 PVCs: Bound with 3 replicas each - Storage: TRUE HA confirmed All future changes via git push only (100% GitOps).
78 lines
2.6 KiB
Bash
Executable File
78 lines
2.6 KiB
Bash
Executable File
#!/bin/bash
|
|
|
|
echo "🔧 Fixing Forgejo issues..."
|
|
echo ""
|
|
|
|
# Issue 1: Multi-Attach - old pod still holding the volume
|
|
echo "==> Issue 1: Cleaning up old Forgejo deployment"
|
|
echo "Current deployments:"
|
|
kubectl get deployment -n cicd | grep forgejo
|
|
|
|
echo ""
|
|
OLD_DEPLOYMENT=$(kubectl get deployment -n cicd -o name | grep -E "forgejo-[0-9]" | grep -v gitea)
|
|
if [ -n "$OLD_DEPLOYMENT" ]; then
|
|
echo "Found old deployment: $OLD_DEPLOYMENT"
|
|
kubectl delete $OLD_DEPLOYMENT -n cicd --wait=true
|
|
echo " ✓ Old deployment deleted"
|
|
else
|
|
echo " No old deployment found, checking for orphaned pods..."
|
|
kubectl get pods -n cicd -l app.kubernetes.io/name=gitea -o name | while read pod; do
|
|
POD_NAME=$(echo $pod | cut -d/ -f2)
|
|
if [[ ! "$POD_NAME" =~ "forgejo-gitea" ]]; then
|
|
echo " Deleting orphaned pod: $POD_NAME"
|
|
kubectl delete pod -n cicd $POD_NAME --force --grace-period=0
|
|
fi
|
|
done
|
|
fi
|
|
|
|
# Issue 2: Missing homelab-ca ConfigMap
|
|
echo ""
|
|
echo "==> Issue 2: Checking homelab-ca ConfigMap"
|
|
if kubectl get configmap homelab-ca -n cicd &>/dev/null; then
|
|
echo " ✓ homelab-ca already exists"
|
|
else
|
|
echo " ⚠️ homelab-ca not found in cicd namespace"
|
|
echo " Checking if it exists elsewhere..."
|
|
|
|
# Check common namespaces
|
|
for ns in default kube-system cert-manager; do
|
|
if kubectl get configmap homelab-ca -n $ns &>/dev/null; then
|
|
echo " Found in namespace: $ns"
|
|
echo " Copying to cicd namespace..."
|
|
kubectl get configmap homelab-ca -n $ns -o yaml | \
|
|
sed 's/namespace: '$ns'/namespace: cicd/' | \
|
|
kubectl apply -f -
|
|
echo " ✓ Copied homelab-ca to cicd"
|
|
break
|
|
fi
|
|
done
|
|
|
|
# If still not found, check if we need to create it
|
|
if ! kubectl get configmap homelab-ca -n cicd &>/dev/null; then
|
|
echo " Creating empty homelab-ca ConfigMap (you may need to populate it)..."
|
|
kubectl create configmap homelab-ca -n cicd --from-literal=ca.crt=""
|
|
echo " ⚠️ Created empty ConfigMap - update with actual CA if needed"
|
|
fi
|
|
fi
|
|
|
|
echo ""
|
|
echo "==> Waiting for volume detach (30s)..."
|
|
sleep 30
|
|
|
|
echo ""
|
|
echo "==> Current Forgejo pod status:"
|
|
kubectl get pods -n cicd -l app.kubernetes.io/name=gitea
|
|
|
|
echo ""
|
|
echo "==> If still in Init or Pending, describe one pod:"
|
|
POD=$(kubectl get pods -n cicd -l app.kubernetes.io/name=gitea --no-headers | head -1 | awk '{print $1}')
|
|
if [ -n "$POD" ]; then
|
|
kubectl describe pod -n cicd $POD | grep -A 10 "Events:" | head -15
|
|
fi
|
|
|
|
echo ""
|
|
echo "✅ Fixes applied!"
|
|
echo ""
|
|
echo "Next: Monitor pod startup"
|
|
echo " kubectl get pods -n cicd -w"
|