Major accomplishments from comprehensive cluster review: ## Storage HA (answering "are volumes replicated?") - Verified 3-node Longhorn HA: ALL 17 volumes have 3 replicas - Fixed CLAUDE.md contradiction (sole node → 3-node HA) - Consolidated to single 'longhorn' StorageClass (3 replicas, WaitForFirstConsumer) - Removed duplicate StorageClasses (longhorn-wffc, longhorn-kafka, longhorn-static) ## GitOps Infrastructure Cleanup - Eliminated resource duplication (ddb-cluster single source of truth) - Restructured k8s/data/ → cluster/ (bootstrap) + schemas/ (GitOps) - Updated data-schemas app to point to k8s/data/schemas/ (wave 6) - Archived old k8s/argocd/bootstrap/ → bootstrap.archived/ ## Bootstrap Dependencies Fixed - Added 05-wait-for-databases.yaml to prevent CNPG race condition - Ensures Database CRs reconciled before Forgejo starts - Proper "PostgreSQL-as-a-Service" workflow ## Longhorn CSI Plugin Fixed - Added patch-csi-tolerations-job.yaml (GitOps PostSync hook) - CSI plugin now runs on all 3 nodes (cp-1, cp-2, cp-3) - Fixes volume attachment on tainted control-plane nodes ## Live Migration (Zero Downtime) - Migrated 37 applications to ArgoCD app-of-apps management - Fixed Forgejo startup issues: * Service selector mismatch (app: forgejo → app: gitea) * Missing homelab-ca ConfigMap * Missing forgejo-oidc secret (temporary) * CNPG database creation timing ## Documentation (10 comprehensive files) - WHATS-NEXT.md - Daily GitOps workflow - MIGRATION-STATUS.md - Cluster health report - REVIEW-SUMMARY.md - Session overview - GITOPS-REBUILD-PLAN.md - Architecture reference - DDB-REVIEW.md - PostgreSQL optimization guide - STORAGE-ARCHITECTURE-CLARIFICATION.md - Storage HA investigation - BOOTSTRAP-DEPENDENCY-FIX.md - CNPG race condition fix - STORAGECLASS-CONSOLIDATION.md - Single StorageClass rationale - IMPLEMENTATION-CHECKLIST.md - Migration checklist - bootstrap.sh - Automated bootstrap script ## Cluster Status - ArgoCD: 4/4 pods running - DDB cluster: 3/3 instances healthy - Longhorn: 3/3 nodes, all CSI plugins running - Forgejo: Running, accessible at http://192.168.1.165:3000 - All 17 PVCs: Bound with 3 replicas each - Storage: TRUE HA confirmed All future changes via git push only (100% GitOps).
9.3 KiB
GitOps Rebuild Implementation Checklist
Use this checklist to track the rebuild implementation step-by-step.
📋 Pre-Implementation
-
Backup current state
kubectl get applications -n argocd -o yaml > backup-argocd-apps.yaml kubectl get cluster ddb-cluster -n ddb -o yaml > backup-ddb-cluster.yaml kubectl get all -n cicd -o yaml > backup-forgejo.yaml kubectl get all -n ddb -o yaml > backup-ddb.yaml -
Verify prerequisites
- kubectl configured (
kubectl cluster-info) - SOPS age key exists (
~/.sops/homelab-age.key) - ArgoCD CLI installed (
argocd version) - Git configured with Forgejo credentials
- kubectl configured (
-
⚠️ CRITICAL: Verify storage replication status
# Check Longhorn nodes (should show 3 if HA is active) kubectl get nodes.longhorn.io -n longhorn-system # Check replica counts kubectl get volumes.longhorn.io -n longhorn-system \ -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.spec.numberOfReplicas}{"\n"}{end}'- If single-node (1 replica): Deploy
k8s/infrastructure/longhorn/first - If 3-node (3 replicas): Update CLAUDE.md to reflect HA status
- See
STORAGE-ARCHITECTURE-CLARIFICATION.mdfor details
- If single-node (1 replica): Deploy
-
Review documentation
- Read
GITOPS-REBUILD-PLAN.mdcompletely - Review
DDB-REVIEW.mdfor database configuration - Understand wave structure (0-8)
- Read
🏗️ Implementation (Choose One Path)
Option A: Fresh Cluster (Recommended)
-
Provision Talos cluster
cd terraform/ terraform apply # Wait for cluster to be ready -
Run bootstrap script
./bootstrap.sh # Verify all steps complete successfully -
Push to Forgejo
git remote add forgejo https://forgejo.riotpiao.com/riotpiao.com/homelab.git git push forgejo main -
Deploy app-of-apps
kubectl apply -f k8s/argocd/projects/homelab-project.yaml kubectl apply -k k8s/argocd/root argocd app sync homelab-root --prune
Option B: Incremental Migration (Existing Cluster)
-
Apply bootstrap-local resources alongside existing
# Dry-run first kubectl apply -k k8s/bootstrap-local/ --dry-run=client # Actually apply (creates bootstrap Applications) kubectl apply -k k8s/bootstrap-local/ -
Update ArgoCD apps one wave at a time
# Wave 0 kubectl apply -f k8s/argocd/apps/00-substrate.yaml argocd app sync cert-manager ingress-nginx reloader # Verify healthy, then proceed to wave 1, 2, 3, etc. -
Update data-schemas path
# Edit 40-data.yaml (already done in this plan) kubectl apply -f k8s/argocd/apps/40-data.yaml argocd app sync data-schemas -
Delete old bootstrap Applications
# These are now in bootstrap-local/ kubectl delete application cnpg-operator -n argocd # Note: Keep forgejo as manual-sync-only
✅ Post-Implementation Verification
Wave 0-2: Infrastructure
-
cert-manager
argocd app get cert-manager kubectl get clusterissuers -A # Verify: letsencrypt-staging, letsencrypt-prod -
ingress-nginx
kubectl get pods -n ingress-nginx kubectl get svc ingress-nginx-controller -n ingress-nginx # Verify: LoadBalancer IP assigned (192.168.1.160) -
Prometheus
kubectl get pods -n monitoring kubectl get servicemonitors -A # Verify: prometheus, grafana, alertmanager running
Wave 3-5: Logging, Secrets, IAM
-
Loki/Grafana
kubectl get pods -n logging # Access: https://grafana.riotpiao.com -
SOPS secrets
argocd app get sops-secrets kubectl get secrets -n ddb | grep db-role # Verify: authentik-db-role, temporal-db-role exist -
Vault
kubectl get pods -n iam kubectl exec -n iam vault-0 -- vault status -
Authentik
kubectl get pods -n iam # Access: https://authentik.riotpiao.com
Wave 6-8: Data, Messaging, Applications
-
Data schemas
kubectl get databases -n ddb # Verify: forgejo, authentik, temporal, temporal_visibility kubectl get jobs -n ddb # Verify: db-init-job Completed -
Kafka/SQS
kubectl get pods -n sqs kubectl get kafkas -n sqs -
Temporal
kubectl get pods -n temporal kubectl logs -n temporal deployment/temporal-frontend -f # Verify: Connected to PostgreSQL # Access: https://temporal.riotpiao.com -
Cloudflared
kubectl get pods -n cloudflared kubectl logs -n cloudflared deployment/cloudflared # Verify: Tunnel connected
Overall Health
-
All Applications Synced
argocd app list # Verify: All STATUS=Synced, HEALTH=Healthy -
No stuck pods
kubectl get pods --all-namespaces | grep -vE 'Running|Completed' # (Should be empty) -
All Ingresses accessible
kubectl get ingress -A # Test each URL in browser -
PostgreSQL connections
# Forgejo kubectl exec -n cicd deployment/forgejo -- psql -h ddb-cluster-rw.ddb.svc -U app -d forgejo -c '\conninfo' # Authentik kubectl exec -n iam deployment/authentik-server -- python manage.py check --database default # Temporal kubectl exec -n temporal deployment/temporal-frontend -- tctl --db_engine postgres cluster health
🧹 Cleanup (After Successful Migration)
-
Remove old bootstrap files (optional)
# Move to archive/ mkdir -p archive/ mv k8s/argocd/bootstrap/ archive/old-bootstrap/ mv USAGE.md project-usage/ archive/old-helmfile-docs/ -
Update CLAUDE.md
# Remove references to Phase 0 manual steps # Update to point to GITOPS-REBUILD-PLAN.md -
Commit cleanup
git add -A git commit -m "chore: migrate to bootstrap-local + GitOps structure" git push forgejo main
🔄 Day-2 Validation
-
Test GitOps workflow
# Make a simple change echo "# Test comment" >> k8s/applications/temporal/temporal-values.yaml git commit -am "test: validate GitOps workflow" git push # Watch ArgoCD auto-sync watch -n 2 'argocd app get temporal | grep -A 3 "Sync Status"' -
Test rollback
git revert HEAD git push # Verify ArgoCD auto-syncs the rollback -
Test adding new application
# Create minimal app mkdir k8s/applications/test-app cat > k8s/applications/test-app/deployment.yaml << 'EOF' apiVersion: apps/v1 kind: Deployment metadata: name: nginx-test namespace: default spec: replicas: 1 selector: matchLabels: app: nginx-test template: metadata: labels: app: nginx-test spec: containers: - name: nginx image: nginx:alpine ports: - containerPort: 80 EOF # Add to ArgoCD cat >> k8s/argocd/apps/08-applications.yaml << 'EOF' --- apiVersion: argoproj.io/v1alpha1 kind: Application metadata: name: test-app namespace: argocd annotations: argocd.argoproj.io/sync-wave: "8" spec: project: homelab source: repoURL: https://forgejo.riotpiao.com/riotpiao.com/homelab.git targetRevision: main path: k8s/applications/test-app destination: server: https://kubernetes.default.svc namespace: default syncPolicy: automated: prune: true selfHeal: true EOF # Push and verify git add -A && git commit -m "test: add test application" && git push argocd app get test-app kubectl get deployment nginx-test # Cleanup kubectl delete application test-app -n argocd git revert HEAD && git push
📊 Monitoring & Alerts
-
Configure Prometheus alerts
- ArgoCD sync failures
- PostgreSQL replication lag
- PVC usage >80%
-
Set up Grafana dashboards
- ArgoCD overview
- PostgreSQL performance
- Ingress traffic
-
Document runbooks
- DDB cluster recovery (CLAUDE.md)
- Forgejo outage (breaks GitOps)
- ArgoCD degradation
🎯 Success Metrics
- Bootstrap time: <10 minutes
- Zero manual kubectl apply (except bootstrap.sh)
- No duplicate resources
- All apps Synced + Healthy
- Git is single source of truth
- Rollbacks via git revert only
🆘 Rollback Plan (If Things Go Wrong)
Rollback to Old State
# 1. Restore ArgoCD Applications
kubectl apply -f backup-argocd-apps.yaml
# 2. Restore DDB cluster (if modified)
kubectl apply -f backup-ddb-cluster.yaml
# 3. Restore Forgejo
kubectl apply -f backup-forgejo.yaml
# 4. Sync all apps to last known good commit
argocd app sync --all --revision <last-good-commit-sha>
Nuclear Option (Full Cluster Rebuild)
# 1. Export all PVCs data (Forgejo git repos, PostgreSQL data)
# (Manual backup via Longhorn UI or velero)
# 2. Destroy cluster
cd terraform/
terraform destroy
# 3. Re-provision from scratch
terraform apply
./bootstrap.sh
# Restore PVC data
Notes:
- Check off items as you complete them
- Add timestamps/notes for each major step
- Keep this checklist updated as you encounter issues