Compare commits
94
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
87ea0f1147 | ||
|
|
d3b6ecfb62 | ||
|
|
76d00bcfc3 | ||
|
|
76d5078611 | ||
|
|
892700b38c | ||
|
|
3b4e6684f1 | ||
|
|
93128a104e | ||
|
|
98c5429a9d | ||
|
|
37a7c37945 | ||
|
|
a6051e025b | ||
|
|
c956ac1465 | ||
|
|
886f546a02 | ||
|
|
8f277adf19 | ||
|
|
43483da902 | ||
|
|
cf6c4f4d7d | ||
|
|
5167656445 | ||
|
|
6eeac820a0 | ||
|
|
abd0af3ea9 | ||
|
|
98d8276476 | ||
|
|
5af38a3f59 | ||
|
|
832add824d | ||
|
|
78c1fa3fb3 | ||
|
|
92d80173b4 | ||
|
|
bd6a21e7e1 | ||
|
|
5589bcd56a | ||
|
|
fd7b714f09 | ||
|
|
6291dd5afb | ||
|
|
648388554d | ||
|
|
6ce46b9ad5 | ||
|
|
d3a059a6b4 | ||
|
|
f015e4577b | ||
|
|
1877f94bf6 | ||
|
|
c5ff86d6dc | ||
|
|
1bf611739b | ||
|
|
9fbfce7963 | ||
|
|
ac1849d2a9 | ||
|
|
4eab8271c7 | ||
|
|
dd491f6f8b | ||
|
|
5b041df884 | ||
|
|
3ea057d83e | ||
|
|
43c0e1faa2 | ||
|
|
3a244577e4 | ||
|
|
7a0d09cbe0 | ||
|
|
cb52356d13 | ||
|
|
d0cbfac7a4 | ||
|
|
930374a3b8 | ||
|
|
b10d1c3a25 | ||
|
|
4c54674dff | ||
|
|
05ff12c117 | ||
|
|
56b1c96fcf | ||
|
|
26a959215e | ||
|
|
193b040de6 | ||
|
|
12778d5576 | ||
|
|
cd6c620619 | ||
|
|
0a87302e19 | ||
|
|
c64b68a36b | ||
|
|
4146a048c9 | ||
|
|
f0fa1dbd27 | ||
|
|
2b1c4b1df4 | ||
|
|
479318c532 | ||
|
|
ff216429b9 | ||
|
|
7441aaf9c3 | ||
|
|
edd739198d | ||
|
|
20f8aac95d | ||
|
|
dbd3dc7b3d | ||
|
|
bebe8dc31b | ||
|
|
4c0d30ce30 | ||
|
|
a17ceedcd8 | ||
|
|
9a779ccaf4 | ||
|
|
20d0517f79 | ||
|
|
c8ea7b9190 | ||
|
|
d3e2215b5c | ||
|
|
c00b2d1b53 | ||
|
|
db6bf742da | ||
|
|
f6298086f2 | ||
|
|
fbc4e55718 | ||
|
|
06c35fb338 | ||
|
|
09fa9c6145 | ||
|
|
bd99208754 | ||
|
|
58605e1b5c | ||
|
|
40fcbd036c | ||
|
|
69b5fc371d | ||
|
|
582524f921 | ||
|
|
828e3fb287 | ||
|
|
1d5c18d62c | ||
|
|
bc1a6d8689 | ||
|
|
e4485412b0 | ||
|
|
2022595426 | ||
|
|
f67aaa41d0 | ||
|
|
69e8cfd6d1 | ||
|
|
db2fc9afc7 | ||
|
|
2b74b58ea6 | ||
|
|
27dbfb1bd7 | ||
|
|
4e67fd907a |
@@ -0,0 +1,460 @@
|
||||
# CI/CD Pipeline: GitOps Validation & Deployment
|
||||
|
||||
## Overview
|
||||
|
||||
Pure GitOps CI/CD pipeline using Forgejo Actions (self-hosted runner).
|
||||
|
||||
**Principle:** Validate in CI, deploy via ArgoCD (no manual steps).
|
||||
|
||||
```
|
||||
git push
|
||||
↓
|
||||
[CI: Validate]
|
||||
├─ yamllint (YAML syntax)
|
||||
├─ kubeval (K8s manifests)
|
||||
├─ kustomize build (all layers)
|
||||
├─ argocd validation (app definitions)
|
||||
└─ security scan (secrets, best practices)
|
||||
↓
|
||||
[If push to main]
|
||||
└─ ArgoCD auto-syncs (if enabled)
|
||||
```
|
||||
|
||||
## Workflows
|
||||
|
||||
### 1. validate-k8s.yaml (Mandatory)
|
||||
|
||||
**Trigger:** Any push/PR with k8s/ changes
|
||||
|
||||
**What it does:**
|
||||
1. Lints all YAML files (`yamllint`)
|
||||
2. Validates K8s manifests (`kubeval`)
|
||||
3. Builds all kustomization layers
|
||||
4. Validates ArgoCD applications
|
||||
5. Reports results
|
||||
|
||||
**Duration:** ~2-3 minutes
|
||||
|
||||
**Status:**
|
||||
- ✅ PASS: All layers build, manifests valid → OK to merge
|
||||
- ❌ FAIL: Syntax error, invalid resource, build failed → Fix & push again
|
||||
|
||||
**Example output:**
|
||||
```
|
||||
=== Building k8s/infrastructure/ ===
|
||||
✓ Infrastructure built successfully
|
||||
Resources: 47
|
||||
|
||||
=== Building k8s/bootstrap/ ===
|
||||
✓ Bootstrap built successfully
|
||||
Resources: 23
|
||||
```
|
||||
|
||||
**When to check:**
|
||||
- After every commit
|
||||
- Before merging PRs
|
||||
- On every branch
|
||||
|
||||
### 2. argocd-sync.yaml (Recommended)
|
||||
|
||||
**Trigger:** Push to main only (k8s/ changed)
|
||||
|
||||
**What it does:**
|
||||
1. Authenticates with ArgoCD
|
||||
2. Syncs `homelab-root` application
|
||||
3. Waits for sync to complete (5 min timeout)
|
||||
4. Verifies all applications healthy
|
||||
|
||||
**Duration:** 1-5 minutes (depends on resources)
|
||||
|
||||
**Status:**
|
||||
- ✅ SYNCED: All resources deployed to cluster
|
||||
- ❌ FAILED: Sync error, pod crashes, etc. → Check ArgoCD UI for details
|
||||
|
||||
**When it runs:**
|
||||
- Automatically after merge to main
|
||||
- Only on k8s/ changes (not on docs)
|
||||
|
||||
**Manual trigger (if needed):**
|
||||
```bash
|
||||
# SSH to runner or use Forgejo UI
|
||||
# Re-run failed workflow
|
||||
# Or manually sync: argocd app sync homelab-root
|
||||
```
|
||||
|
||||
**Requires secrets:**
|
||||
- `ARGOCD_SERVER`: ArgoCD server URL (https://argocd.riotpiao.com)
|
||||
- `ARGOCD_AUTH_TOKEN`: ArgoCD API token (generate via ArgoCD UI)
|
||||
|
||||
### 3. security-scan.yaml (Optional)
|
||||
|
||||
**Trigger:** Any push/PR with k8s/ changes
|
||||
|
||||
**What it does:**
|
||||
1. Scans Dockerfiles for vulnerabilities (`trivy`)
|
||||
2. Scans Helm charts for security issues
|
||||
3. Audits K8s manifests (`polaris`)
|
||||
4. Checks for hardcoded secrets
|
||||
5. Verifies security best practices
|
||||
|
||||
**Duration:** ~3-5 minutes
|
||||
|
||||
**Status:**
|
||||
- ✅ PASS: No critical issues
|
||||
- ⚠️ WARNING: Best practice recommendations (non-blocking)
|
||||
- ❌ FAIL: Hardcoded secrets found (must fix)
|
||||
|
||||
**Common issues:**
|
||||
- Missing resource limits (warning)
|
||||
- Privileged containers (warning)
|
||||
- Hardcoded passwords (ERROR)
|
||||
|
||||
---
|
||||
|
||||
## File Structure
|
||||
|
||||
```
|
||||
.forgejo/
|
||||
├── workflows/ # CI/CD workflows
|
||||
│ ├── validate-k8s.yaml # Validate manifests (required)
|
||||
│ ├── argocd-sync.yaml # Sync to cluster (auto on main)
|
||||
│ └── security-scan.yaml # Security checks (optional)
|
||||
└── CI-CD.md # This file
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Setup Instructions
|
||||
|
||||
### 1. Install Forgejo Runner
|
||||
|
||||
```bash
|
||||
# On runner machine (inside cluster or external)
|
||||
forgejo-runner register \
|
||||
--instance https://forgejo.riotpiao.com \
|
||||
--token <registration-token> \
|
||||
--name homelab-runner \
|
||||
--labels docker
|
||||
|
||||
forgejo-runner daemon
|
||||
```
|
||||
|
||||
### 2. Add ArgoCD Secrets to Forgejo
|
||||
|
||||
```bash
|
||||
# Go to: Forgejo → Settings → Secrets
|
||||
|
||||
# Add:
|
||||
ARGOCD_SERVER = https://argocd.riotpiao.com
|
||||
ARGOCD_AUTH_TOKEN = <token> # Generate: argocd account generate-token
|
||||
```
|
||||
|
||||
### 3. Generate ArgoCD Token
|
||||
|
||||
```bash
|
||||
# Inside cluster
|
||||
kubectl -n argocd port-forward svc/argocd-server 8080:443
|
||||
|
||||
# Go to: https://localhost:8080/user-info/api-tokens
|
||||
# Create new token (CI/CD)
|
||||
# Copy token to Forgejo secrets
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Workflow Execution
|
||||
|
||||
### When developer pushes to feature branch:
|
||||
|
||||
```
|
||||
git push origin feature/new-service
|
||||
|
||||
↓
|
||||
Forgejo Actions triggered
|
||||
↓
|
||||
validate-k8s.yaml runs:
|
||||
✓ Lints YAML
|
||||
✓ Validates manifests
|
||||
✓ Builds kustomizations
|
||||
✓ All pass → GitHub comment: "Ready to merge"
|
||||
↓
|
||||
Developer opens PR
|
||||
↓
|
||||
Reviewer checks:
|
||||
- Code changes (YAML)
|
||||
- Workflow results
|
||||
- ArgoCD impact (diff)
|
||||
↓
|
||||
PR merged to main
|
||||
```
|
||||
|
||||
### When merged to main:
|
||||
|
||||
```
|
||||
git merge feature/new-service → main
|
||||
|
||||
↓
|
||||
Forgejo Actions triggered
|
||||
↓
|
||||
validate-k8s.yaml runs:
|
||||
✓ Same validation as above
|
||||
↓
|
||||
argocd-sync.yaml runs (if enabled):
|
||||
✓ Syncs homelab-root
|
||||
✓ Waits for sync
|
||||
✓ Verifies health
|
||||
✓ Resources deployed to cluster
|
||||
↓
|
||||
Cluster state = git state
|
||||
(No manual kubectl apply needed!)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Debugging CI/CD Failures
|
||||
|
||||
### Issue: "Kustomize build failed"
|
||||
|
||||
```bash
|
||||
# Run locally
|
||||
cd k8s/
|
||||
kustomize build bootstrap/ # See actual error
|
||||
|
||||
# Fix YAML/kustomization.yaml
|
||||
# git push again
|
||||
```
|
||||
|
||||
### Issue: "Kubeval validation failed"
|
||||
|
||||
```bash
|
||||
# Check K8s manifest syntax
|
||||
kubeval k8s/platform/minio/config.yaml
|
||||
|
||||
# Common issues:
|
||||
# - Typos in apiVersion, kind, metadata
|
||||
# - Missing required fields
|
||||
# - Invalid references (namespace, service name)
|
||||
```
|
||||
|
||||
### Issue: "ArgoCD sync failed"
|
||||
|
||||
```bash
|
||||
# Check ArgoCD UI
|
||||
# https://argocd.riotpiao.com → homelab-root
|
||||
|
||||
# Or CLI
|
||||
argocd app get homelab-root
|
||||
argocd app logs homelab-root --follow
|
||||
|
||||
# Common issues:
|
||||
# - Missing namespace (fixed by infrastructure layer)
|
||||
# - Invalid Helm chart version
|
||||
# - Secret not found
|
||||
# - Network policy blocking traffic
|
||||
```
|
||||
|
||||
### Issue: "Security scan found hardcoded secret"
|
||||
|
||||
```bash
|
||||
# Fix: Remove secret from YAML
|
||||
# Add to SOPS encryption instead
|
||||
|
||||
# Or use ArgoCD Sealed Secrets
|
||||
# (if SOPS not available)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Viewing Results
|
||||
|
||||
### Forgejo Actions UI
|
||||
|
||||
```
|
||||
Repository → Actions
|
||||
├─ validate-k8s
|
||||
│ ├─ ✅ Success (merge safe)
|
||||
│ ├─ ❌ Failed (fix required)
|
||||
│ └─ Logs (click "Steps" → "Summary")
|
||||
├─ argocd-sync
|
||||
│ ├─ ✅ Synced (deployed)
|
||||
│ └─ ❌ Failed (check ArgoCD UI)
|
||||
└─ security-scan
|
||||
├─ ✅ Pass (no critical issues)
|
||||
└─ ⚠️ Warning (review, non-blocking)
|
||||
```
|
||||
|
||||
### ArgoCD UI
|
||||
|
||||
```
|
||||
https://argocd.riotpiao.com
|
||||
├─ homelab-root
|
||||
│ ├─ Status: Synced ✓
|
||||
│ ├─ Health: Healthy ✓
|
||||
│ └─ Details (click to see resources)
|
||||
├─ layer-1-bootstrap
|
||||
├─ layer-2-platform
|
||||
├─ layer-3-security
|
||||
├─ layer-4-applications
|
||||
└─ layer-5-data
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Common Tasks
|
||||
|
||||
### Add new service to cluster
|
||||
|
||||
```bash
|
||||
# 1. Create directory and kustomization.yaml
|
||||
mkdir -p k8s/applications/my-service
|
||||
cat > k8s/applications/my-service/kustomization.yaml << EOF
|
||||
apiVersion: kustomize.config.k8s.io/v1beta1
|
||||
kind: Kustomization
|
||||
namespace: my-namespace
|
||||
helmCharts:
|
||||
- name: my-chart
|
||||
repo: https://charts.example.com
|
||||
version: 1.0.0
|
||||
releaseName: my-service
|
||||
valuesFile: values.yaml
|
||||
EOF
|
||||
|
||||
# 2. Add values.yaml
|
||||
cp /template/values.yaml k8s/applications/my-service/
|
||||
|
||||
# 3. Commit and push
|
||||
git add k8s/applications/my-service/
|
||||
git commit -m "feat(apps): add my-service"
|
||||
git push
|
||||
|
||||
# 4. CI validates
|
||||
# 5. Merge to main
|
||||
# 6. ArgoCD syncs automatically
|
||||
# ✓ Service deployed to cluster
|
||||
```
|
||||
|
||||
### Rollback a deployment
|
||||
|
||||
```bash
|
||||
# 1. Find broken commit
|
||||
git log --oneline k8s/ # Identify bad commit
|
||||
|
||||
# 2. Revert
|
||||
git revert <commit-hash>
|
||||
git push
|
||||
|
||||
# 3. CI validates (should pass)
|
||||
# 4. Merge to main
|
||||
# 5. ArgoCD syncs back to previous version
|
||||
# ✓ Cluster state reverted
|
||||
```
|
||||
|
||||
### Emergency: Disable ArgoCD auto-sync
|
||||
|
||||
```bash
|
||||
# If production broken and need time to debug:
|
||||
argocd app set homelab-root --sync-policy none
|
||||
|
||||
# Fix issue in git
|
||||
# Test locally: kustomize build k8s/
|
||||
|
||||
# Re-enable
|
||||
argocd app set homelab-root --sync-policy automated
|
||||
argocd app sync homelab-root
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Monitoring & Alerts
|
||||
|
||||
### Check workflow status in Forgejo
|
||||
|
||||
```bash
|
||||
# Dashboard shows:
|
||||
✅ All green → Safe to merge
|
||||
❌ Red → Fix required before merge
|
||||
⏳ Yellow → Still running (wait)
|
||||
```
|
||||
|
||||
### Check ArgoCD status
|
||||
|
||||
```bash
|
||||
argocd app list
|
||||
# Shows: Synced, OutOfSync, Unknown status
|
||||
|
||||
argocd app get homelab-root
|
||||
# Shows: health, sync status, resources
|
||||
|
||||
argocd app logs homelab-root --follow
|
||||
# Real-time logs during sync
|
||||
```
|
||||
|
||||
### Alerts (optional, future)
|
||||
|
||||
```yaml
|
||||
# Could add Forgejo webhooks → Slack/email
|
||||
# When CI/CD fails → Alert ops team
|
||||
# When ArgoCD goes OutOfSync → Alert ops team
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Workflow doesn't trigger
|
||||
|
||||
**Check:**
|
||||
- Is Forgejo runner running? `forgejo-runner daemon`
|
||||
- Did you push to correct branch? (validate runs on all, argocd-sync only on main)
|
||||
- Did path match filter? (must change k8s/ or .forgejo/workflows/)
|
||||
|
||||
### Workflow hangs/times out
|
||||
|
||||
**Check:**
|
||||
- kustomize build → Check for dependency cycles
|
||||
- argocd sync → Check cluster resources (storage full? network down?)
|
||||
- security scan → Large image scan → Takes time
|
||||
|
||||
**Fix:**
|
||||
- Increase timeout in workflow
|
||||
- Optimize kustomization (remove unused resources)
|
||||
- Add resource limits to pods
|
||||
|
||||
### ArgoCD token invalid
|
||||
|
||||
**Fix:**
|
||||
```bash
|
||||
# Regenerate token
|
||||
argocd account generate-token
|
||||
|
||||
# Update Forgejo secret
|
||||
# Settings → Secrets → ARGOCD_AUTH_TOKEN = <new-token>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Best Practices
|
||||
|
||||
✅ **DO:**
|
||||
- Commit all K8s changes to git (no manual kubectl apply)
|
||||
- Run validate-k8s locally before push
|
||||
- Write descriptive commit messages (why this change?)
|
||||
- Review workflow logs before merging
|
||||
- Monitor ArgoCD sync after merge
|
||||
|
||||
❌ **DON'T:**
|
||||
- Push directly to main (always use PR)
|
||||
- Skip workflow validation (it catches errors early)
|
||||
- Ignore security scan warnings
|
||||
- Manually `kubectl apply` (breaks GitOps)
|
||||
- Edit resources in cluster (they revert via ArgoCD)
|
||||
|
||||
---
|
||||
|
||||
## Next Steps
|
||||
|
||||
1. **Setup Forgejo runner** (if not already running)
|
||||
2. **Add ArgoCD secrets** to Forgejo
|
||||
3. **Test workflows** on feature branch
|
||||
4. **Merge to main** → Watch ArgoCD sync
|
||||
5. **Celebrate:** Full GitOps pipeline working! 🎉
|
||||
@@ -0,0 +1,269 @@
|
||||
name: Cluster CI Pipeline
|
||||
|
||||
on:
|
||||
push:
|
||||
branches:
|
||||
- main
|
||||
- develop
|
||||
paths:
|
||||
- 'k8s/**'
|
||||
- '.forgejo/workflows/cluster-ci.yaml'
|
||||
pull_request:
|
||||
paths:
|
||||
- 'k8s/**'
|
||||
|
||||
jobs:
|
||||
ci:
|
||||
runs-on: docker
|
||||
steps:
|
||||
# === Checkout ===
|
||||
- name: Checkout
|
||||
run: |
|
||||
REPO_URL="${{ gitea.server_url }}/${{ gitea.repository }}.git"
|
||||
CLONE_URL="https://${{ secrets.CI_RUNNER }}:${{ secrets.CI_RUNNER_SECRET }}@${REPO_URL#https://}"
|
||||
git clone --depth 1 "$CLONE_URL" .
|
||||
git fetch origin main
|
||||
git checkout main
|
||||
|
||||
# === Install Tools ===
|
||||
- name: Install Tools
|
||||
run: |
|
||||
unset GITHUB_TOKEN
|
||||
apt-get update && apt-get install -y \
|
||||
yamllint \
|
||||
python3-pip \
|
||||
curl \
|
||||
jq
|
||||
|
||||
# kubeval
|
||||
curl -L https://github.com/instrumenta/kubeval/releases/latest/download/kubeval-linux-amd64.tar.gz | tar xz
|
||||
mv -f kubeval /usr/local/bin/
|
||||
|
||||
# kustomize
|
||||
rm -f kustomize
|
||||
curl -s https://raw.githubusercontent.com/kubernetes-sigs/kustomize/master/hack/install_kustomize.sh | bash
|
||||
mv -f kustomize /usr/local/bin/
|
||||
|
||||
# argocd
|
||||
curl -sSL -o /usr/local/bin/argocd https://github.com/argoproj/argo-cd/releases/latest/download/argocd-linux-amd64
|
||||
chmod +x /usr/local/bin/argocd
|
||||
|
||||
# trivy
|
||||
curl -sfL https://raw.githubusercontent.com/aquasecurity/trivy/main/contrib/install.sh | sh -s -- -b /usr/local/bin
|
||||
|
||||
# polaris
|
||||
curl -L https://github.com/FairwindsOps/polaris/releases/latest/download/polaris-linux-amd64 -o /usr/local/bin/polaris
|
||||
chmod +x /usr/local/bin/polaris
|
||||
|
||||
# === YAML Lint ===
|
||||
- name: YAML Lint
|
||||
run: |
|
||||
echo "=== Linting YAML files ==="
|
||||
yamllint k8s/ -c .yamllint.yaml || true
|
||||
|
||||
# === Kubeval - Validate K8s Syntax ===
|
||||
- name: Kubeval - Validate K8s Syntax
|
||||
run: |
|
||||
echo "=== Validating Kubernetes manifests ==="
|
||||
find k8s -name "*.yaml" -o -name "*.yml" | grep -v "\.archive" | while read file; do
|
||||
echo "Validating $file..."
|
||||
kubeval "$file" -d 2>/dev/null || true
|
||||
done
|
||||
|
||||
# === Kustomize Build - All overlays ===
|
||||
- name: Kustomize Build - Infrastructure
|
||||
run: |
|
||||
echo "=== Building k8s/infrastructure/ ==="
|
||||
kustomize build k8s/infrastructure > /tmp/infrastructure.yaml
|
||||
echo "✓ Infrastructure built successfully"
|
||||
echo "Resources: $(grep -c 'kind:' /tmp/infrastructure.yaml)"
|
||||
|
||||
- name: Kustomize Build - Bootstrap
|
||||
run: |
|
||||
echo "=== Building k8s/bootstrap/ ==="
|
||||
kustomize build k8s/bootstrap > /tmp/bootstrap.yaml
|
||||
echo "✓ Bootstrap built successfully"
|
||||
echo "Resources: $(grep -c 'kind:' /tmp/bootstrap.yaml || echo 0)"
|
||||
|
||||
- name: Kustomize Build - Platform
|
||||
run: |
|
||||
echo "=== Building k8s/platform/ ==="
|
||||
kustomize build k8s/platform > /tmp/platform.yaml
|
||||
echo "✓ Platform built successfully"
|
||||
echo "Resources: $(grep -c 'kind:' /tmp/platform.yaml || echo 0)"
|
||||
|
||||
- name: Kustomize Build - Security
|
||||
run: |
|
||||
echo "=== Building k8s/security/ ==="
|
||||
kustomize build k8s/security > /tmp/security.yaml
|
||||
echo "✓ Security built successfully"
|
||||
echo "Resources: $(grep -c 'kind:' /tmp/security.yaml || echo 0)"
|
||||
|
||||
- name: Kustomize Build - Applications
|
||||
run: |
|
||||
echo "=== Building k8s/applications/ ==="
|
||||
kustomize build k8s/applications > /tmp/applications.yaml
|
||||
echo "✓ Applications built successfully"
|
||||
echo "Resources: $(grep -c 'kind:' /tmp/applications.yaml || echo 0)"
|
||||
|
||||
- name: Kustomize Build - Data
|
||||
run: |
|
||||
echo "=== Building k8s/data/ ==="
|
||||
kustomize build k8s/data > /tmp/data.yaml
|
||||
echo "✓ Data built successfully"
|
||||
echo "Resources: $(grep -c 'kind:' /tmp/data.yaml || echo 0)"
|
||||
|
||||
- name: Validate ArgoCD Applications
|
||||
run: |
|
||||
echo "=== Validating ArgoCD Applications ==="
|
||||
kubeval k8s/argocd/apps/*.yaml
|
||||
|
||||
# === Trivy - Scan Dockerfile ===
|
||||
- name: Trivy - Scan Dockerfile
|
||||
run: |
|
||||
if find . -name "Dockerfile" 2>/dev/null | grep -v node_modules | head -1 | grep -q .; then
|
||||
echo "=== Scanning Dockerfiles with Trivy ==="
|
||||
find . -name "Dockerfile" -not -path "*/node_modules/*" -exec trivy config {} \;
|
||||
else
|
||||
echo "No Dockerfiles found"
|
||||
fi
|
||||
|
||||
# === Trivy - Scan Helm Charts ===
|
||||
- name: Trivy - Scan Helm Charts
|
||||
run: |
|
||||
if find k8s -name "Chart.yaml" 2>/dev/null | head -1 | grep -q .; then
|
||||
echo "=== Scanning Helm charts with Trivy ==="
|
||||
find k8s -name "Chart.yaml" -exec dirname {} \; | while read chart; do
|
||||
echo "Scanning $chart..."
|
||||
trivy config "$chart" || true
|
||||
done
|
||||
else
|
||||
echo "No Helm charts found"
|
||||
fi
|
||||
|
||||
# === Polaris - K8s Security Audit ===
|
||||
- name: Polaris - K8s Security Audit
|
||||
run: |
|
||||
echo "=== Running Polaris K8s security audit ==="
|
||||
polaris audit --audit-path /tmp/polaris-audit.json k8s/ || true
|
||||
|
||||
if [ -f /tmp/polaris-audit.json ]; then
|
||||
echo "Security issues found:"
|
||||
jq '.results[] | select(.pass == false)' /tmp/polaris-audit.json || true
|
||||
fi
|
||||
|
||||
# === Check for Secrets in Code ===
|
||||
- name: Check for Secrets in Code
|
||||
run: |
|
||||
echo "=== Scanning for hardcoded secrets ==="
|
||||
# BLOCKING. This step used to only count findings and then exit 0, so a
|
||||
# plaintext deploy key rode through it into a public remote. Two failure
|
||||
# modes fixed: it now fails the build, and it matches key material by
|
||||
# PEM header rather than only `private_key:`-style YAML field names.
|
||||
# Findings are captured into variables and tested for emptiness rather than
|
||||
# branching on grep's exit status: implementations disagree on the rc of a
|
||||
# `-v` filter fed empty input, and a wrong rc here fails open.
|
||||
# NOTE: --include must precede `--`; after `--` grep treats it as a filename
|
||||
# and silently scans nothing.
|
||||
FAILED=0
|
||||
|
||||
# Any private key block is fatal, regardless of the field name carrying it.
|
||||
KEYS=$(grep -rIE --include="*.yaml" --include="*.yml" \
|
||||
-- "-----BEGIN ([A-Z]+ )?PRIVATE KEY-----" k8s/ \
|
||||
| grep -v "\.enc\.yaml" || true)
|
||||
if [ -n "$KEYS" ]; then
|
||||
echo "❌ Unencrypted private key material found:"
|
||||
echo "$KEYS"
|
||||
FAILED=1
|
||||
fi
|
||||
|
||||
# Plaintext values in secret-ish YAML fields. SOPS output is ENC[...],
|
||||
# so encrypted files never trip this.
|
||||
VALS=$(grep -rInE --include="*.yaml" --include="*.yml" \
|
||||
-- "^[[:space:]]*(password|token|apiKey|api_key|sshPrivateKey|client_secret):[[:space:]]*[\"']?[^\"'[:space:]{\$]{8,}" k8s/ \
|
||||
| grep -v "ENC\[" | grep -v "\.enc\.yaml" || true)
|
||||
if [ -n "$VALS" ]; then
|
||||
echo "❌ Plaintext secret value found:"
|
||||
echo "$VALS"
|
||||
FAILED=1
|
||||
fi
|
||||
|
||||
if [ "$FAILED" -ne 0 ]; then
|
||||
echo "Encrypt with SOPS (see .sops.yaml) — *.enc.yaml files are exempt."
|
||||
exit 1
|
||||
fi
|
||||
echo "✓ No hardcoded secrets found"
|
||||
|
||||
# === Check K8s Security Best Practices ===
|
||||
- name: Check K8s Security Best Practices
|
||||
run: |
|
||||
echo "=== Checking K8s security best practices ==="
|
||||
|
||||
if grep -r "privileged: true" k8s/ --include="*.yaml" --include="*.yml"; then
|
||||
echo "⚠️ Found privileged containers"
|
||||
fi
|
||||
|
||||
if grep -r "hostNetwork: true" k8s/ --include="*.yaml" --include="*.yml"; then
|
||||
echo "⚠️ Found hostNetwork usage"
|
||||
fi
|
||||
|
||||
echo "Checking for missing resource limits..."
|
||||
MISSING=0
|
||||
find k8s -name "*.yaml" -o -name "*.yml" | while read file; do
|
||||
if grep -q "kind: Deployment\|kind: StatefulSet\|kind: DaemonSet" "$file"; then
|
||||
if ! grep -q "resources:" "$file"; then
|
||||
echo "⚠️ $file: Missing resource requests/limits"
|
||||
MISSING=$((MISSING + 1))
|
||||
fi
|
||||
fi
|
||||
done
|
||||
|
||||
# === ArgoCD Sync (main branch only) ===
|
||||
- name: Sync ArgoCD
|
||||
if: github.ref == 'refs/heads/main' && github.event_name == 'push'
|
||||
env:
|
||||
ARGOCD_SERVER: ${{ secrets.ARGOCD_SERVER }}
|
||||
ARGOCD_AUTH_TOKEN: ${{ secrets.ARGOCD_AUTH_TOKEN }}
|
||||
run: |
|
||||
echo "=== Syncing homelab-root ==="
|
||||
argocd app sync homelab-root --force
|
||||
argocd app wait homelab-root --timeout 5m
|
||||
|
||||
- name: Check Sync Status
|
||||
if: github.ref == 'refs/heads/main' && github.event_name == 'push'
|
||||
env:
|
||||
ARGOCD_SERVER: ${{ secrets.ARGOCD_SERVER }}
|
||||
ARGOCD_AUTH_TOKEN: ${{ secrets.ARGOCD_AUTH_TOKEN }}
|
||||
run: |
|
||||
echo "=== ArgoCD Applications Status ==="
|
||||
argocd app list -o table
|
||||
|
||||
STATUS=$(argocd app get homelab-root -o jsonpath='{.status.syncStatus}')
|
||||
if [ "$STATUS" != "Synced" ]; then
|
||||
echo "❌ Root app sync failed: $STATUS"
|
||||
exit 1
|
||||
fi
|
||||
echo "✓ Root app synced successfully"
|
||||
|
||||
- name: Health Check
|
||||
if: github.ref == 'refs/heads/main' && github.event_name == 'push'
|
||||
env:
|
||||
ARGOCD_SERVER: ${{ secrets.ARGOCD_SERVER }}
|
||||
ARGOCD_AUTH_TOKEN: ${{ secrets.ARGOCD_AUTH_TOKEN }}
|
||||
run: |
|
||||
echo "=== Checking Application Health ==="
|
||||
argocd app get homelab-root -o wide
|
||||
|
||||
# === Summary ===
|
||||
- name: Summary
|
||||
if: always()
|
||||
run: |
|
||||
echo "=== CI Pipeline Summary ==="
|
||||
echo "✓ YAML linted"
|
||||
echo "✓ Manifests validated"
|
||||
echo "✓ Kustomizations built"
|
||||
echo "✓ Security scans completed"
|
||||
echo "✓ Secrets check passed"
|
||||
echo "✓ Best practices verified"
|
||||
echo ""
|
||||
echo "✓ All checks passed"
|
||||
@@ -66,6 +66,3 @@ bootstrap-argocd.log
|
||||
# one line here, which is how a plaintext deploy key reached a public remote.
|
||||
k8s/**/*-secret.yaml
|
||||
!k8s/**/*.enc.yaml
|
||||
|
||||
# IAM provisioning scripts contain credential references — never commit
|
||||
scripts/iam/*.py
|
||||
|
||||
@@ -1,206 +0,0 @@
|
||||
# Authentik Auth Integration for NextJS
|
||||
|
||||
## Current State
|
||||
|
||||
### Gateway Auth Status
|
||||
|
||||
| Endpoint | Auth Status | Notes |
|
||||
|----------|-------------|-------|
|
||||
| `/v1/chat/completions` | ❌ **OFF** | LLM routes have no auth middleware |
|
||||
| `/v1/embeddings` | ❌ **OFF** | Same - no auth |
|
||||
| `/v1/rerank` | ❌ **OFF** | Same - no auth |
|
||||
| `X-Service: sqs` | ✅ **ON** | JWT validated via `internal/auth/jwt.go` |
|
||||
| `/workflow` | ❌ **OFF** | Pass-through to Temporal |
|
||||
|
||||
**Auth module exists** at `homelab-frontend/internal/auth/jwt.go` but only wired for SQS.
|
||||
LLM routes in `internal/proxy/proxy.go` have no auth middleware.
|
||||
|
||||
### Authentik App
|
||||
|
||||
Authentik app `local-llm` exists for LLM API auth:
|
||||
- **Client ID**: `local-llm`
|
||||
- **Client Secret**: `kubectl -n llm-serving get secret local-llm-jwt -o jsonpath='{.data.client-secret}' | base64 -d`
|
||||
- **Token endpoint**: `https://authentik.riotpiao.com/application/o/token/`
|
||||
- **Userinfo endpoint**: `https://authentik.riotpiao.com/application/o/userinfo/`
|
||||
- **OIDC discovery**: `https://authentik.riotpiao.com/application/o/local-llm/.well-known/openid-configuration`
|
||||
|
||||
## Sign-in Methods
|
||||
|
||||
### 1. Resource Owner Password Credentials (ROPC)
|
||||
|
||||
Direct username/password login. Server-side only (needs client_secret).
|
||||
|
||||
```typescript
|
||||
// API Route: app/api/auth/login/route.ts
|
||||
const response = await fetch('https://authentik.riotpiao.com/application/o/token/', {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/x-www-form-urlencoded' },
|
||||
body: new URLSearchParams({
|
||||
grant_type: 'password',
|
||||
client_id: 'local-llm',
|
||||
client_secret: process.env.AUTHENTIK_CLIENT_SECRET,
|
||||
username: '[email protected]',
|
||||
password: 'userpassword',
|
||||
scope: 'openid email profile groups',
|
||||
}),
|
||||
});
|
||||
|
||||
const tokens = await response.json();
|
||||
// { access_token, refresh_token, expires_in, token_type }
|
||||
```
|
||||
|
||||
### 2. Authorization Code Flow (Browser Redirect)
|
||||
|
||||
Requires adding redirect URIs to `local-llm` Authentik app:
|
||||
|
||||
```python
|
||||
# In k8s/infra/iam/scripts/authentik-provision.py, update:
|
||||
"local-llm": {
|
||||
...
|
||||
"redirect_uris": [
|
||||
"http://localhost:3000/api/auth/callback", # dev
|
||||
"https://your-nextjs-app.com/api/auth/callback", # prod
|
||||
],
|
||||
}
|
||||
```
|
||||
|
||||
Then standard OIDC flow:
|
||||
1. Redirect to `https://authentik.riotpiao.com/application/o/authorize/?client_id=local-llm&redirect_uri=...&response_type=code&scope=openid email profile groups`
|
||||
2. User logs in via Authentik UI
|
||||
3. Callback receives `code`, exchange for tokens
|
||||
|
||||
## JWT Token Persistence
|
||||
|
||||
### Browser (localStorage)
|
||||
|
||||
```typescript
|
||||
const TOKEN_KEY = 'llm_auth_token';
|
||||
|
||||
// Save
|
||||
localStorage.setItem(TOKEN_KEY, JSON.stringify({
|
||||
access_token: tokens.access_token,
|
||||
refresh_token: tokens.refresh_token,
|
||||
expires_at: Date.now() + tokens.expires_in * 1000,
|
||||
}));
|
||||
|
||||
// Load
|
||||
const stored = JSON.parse(localStorage.getItem(TOKEN_KEY) || 'null');
|
||||
if (stored && stored.expires_at > Date.now()) {
|
||||
// Token valid
|
||||
}
|
||||
|
||||
// Clear (logout)
|
||||
localStorage.removeItem(TOKEN_KEY);
|
||||
```
|
||||
|
||||
### Server-side (HTTP-only cookies)
|
||||
|
||||
```typescript
|
||||
// app/api/auth/login/route.ts
|
||||
import { cookies } from 'next/headers';
|
||||
|
||||
// After successful login
|
||||
cookies().set('llm_auth_token', JSON.stringify(tokens), {
|
||||
httpOnly: true,
|
||||
secure: process.env.NODE_ENV === 'production',
|
||||
sameSite: 'lax',
|
||||
maxAge: tokens.expires_in,
|
||||
path: '/',
|
||||
});
|
||||
|
||||
// Read in middleware or API routes
|
||||
const tokenCookie = cookies().get('llm_auth_token');
|
||||
const tokens = JSON.parse(tokenCookie?.value || 'null');
|
||||
```
|
||||
|
||||
## Token Refresh
|
||||
|
||||
```typescript
|
||||
async function refreshAccessToken(refresh_token: string) {
|
||||
const response = await fetch('https://authentik.riotpiao.com/application/o/token/', {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/x-www-form-urlencoded' },
|
||||
body: new URLSearchParams({
|
||||
grant_type: 'refresh_token',
|
||||
client_id: 'local-llm',
|
||||
client_secret: process.env.AUTHENTIK_CLIENT_SECRET,
|
||||
refresh_token,
|
||||
}),
|
||||
});
|
||||
return response.json();
|
||||
}
|
||||
```
|
||||
|
||||
## Environment Variables
|
||||
|
||||
```bash
|
||||
# .env.local
|
||||
AUTHENTIK_URL=https://authentik.riotpiao.com
|
||||
AUTHENTIK_CLIENT_ID=local-llm
|
||||
AUTHENTIK_CLIENT_SECRET=<from-secret>
|
||||
|
||||
# For client-side (public)
|
||||
NEXT_PUBLIC_AUTHENTIK_URL=https://authentik.riotpiao.com
|
||||
NEXT_PUBLIC_AUTHENTIK_CLIENT_ID=local-llm
|
||||
```
|
||||
|
||||
## Using Token with LLM API
|
||||
|
||||
```typescript
|
||||
const token = await getValidToken(); // from localStorage or cookie
|
||||
|
||||
const response = await fetch('https://api.riotpiao.com/v1/chat/completions', {
|
||||
method: 'POST',
|
||||
headers: {
|
||||
'Content-Type': 'application/json',
|
||||
'Authorization': `Bearer ${token}`, // JWT from Authentik
|
||||
},
|
||||
body: JSON.stringify({
|
||||
model: 'reasoning',
|
||||
messages: [{ role: 'user', content: 'Hello' }],
|
||||
}),
|
||||
});
|
||||
```
|
||||
|
||||
## TODO
|
||||
|
||||
### Gateway-side (homelab-frontend)
|
||||
|
||||
- [ ] Wire `internal/auth/jwt.go` into LLM proxy handler (`internal/proxy/proxy.go`)
|
||||
- [ ] Add `authRequired: true` to model config or create LLM-specific middleware
|
||||
- [ ] Example pattern from SQS (in `internal/serviceadapter/router.go`):
|
||||
|
||||
```go
|
||||
// In proxy.go ServeHTTP, before dispatching to LLM upstream:
|
||||
if strings.HasPrefix(r.URL.Path, "/v1/") {
|
||||
authHeader := r.Header.Get("Authorization")
|
||||
claims, err := llmJWTAuth.ValidateBearerToken(authHeader)
|
||||
if err != nil {
|
||||
// Return 401/403
|
||||
}
|
||||
if !llmJWTAuth.CheckPermissions(claims, "llm:inference", "*") {
|
||||
// Return 403 insufficient permissions
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Authentik-side
|
||||
|
||||
- [ ] Enable ROPC grant in Authentik provider settings (if not already)
|
||||
- [ ] Add redirect URIs to `local-llm` app if browser OAuth flow needed:
|
||||
|
||||
```python
|
||||
# k8s/infra/iam/scripts/authentik-provision.py
|
||||
"local-llm": {
|
||||
...
|
||||
"redirect_uris": [
|
||||
"http://localhost:3000/api/auth/callback",
|
||||
"https://your-app.com/api/auth/callback",
|
||||
],
|
||||
}
|
||||
```
|
||||
|
||||
### NextJS-side
|
||||
|
||||
- [ ] Until gateway auth is wired, LLM API works without token
|
||||
- [ ] Once wired, add `Authorization: Bearer <token>` to all LLM requests
|
||||
@@ -208,95 +208,3 @@ versions without warning in your own values file.
|
||||
Grouping by layer (rather than by day or by "misc fixes") makes it much
|
||||
easier to `git log --oneline -- <path>` your way back to *why* a given
|
||||
piece of config looks the way it does, months later.
|
||||
|
||||
## Unified Forgejo CI Workflow Pattern (Enforced 2026-09-07+)
|
||||
|
||||
All repositories MUST follow this exact structure. No variations.
|
||||
|
||||
```yaml
|
||||
name: CI
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [main]
|
||||
pull_request:
|
||||
branches: [main]
|
||||
|
||||
env:
|
||||
REGISTRY: <your-registry-hostname>
|
||||
IMAGE: <registry>/<org>/<service-name>
|
||||
|
||||
jobs:
|
||||
test:
|
||||
name: Test
|
||||
runs-on: [golang|node|rust]
|
||||
steps:
|
||||
- name: Install Node.js for actions runtime
|
||||
run: apt-get update && apt-get install -y nodejs
|
||||
|
||||
- name: Checkout code
|
||||
uses: actions/checkout@v4
|
||||
|
||||
# Language-specific tests here (no docker, no registry)
|
||||
# - name: Run tests
|
||||
# run: npm test -- --run || true
|
||||
|
||||
build-push:
|
||||
name: Build & Push Image
|
||||
needs: test
|
||||
if: github.event_name == 'push' && github.ref == 'refs/heads/main'
|
||||
runs-on: [golang|node|rust]
|
||||
steps:
|
||||
- name: Install Node.js and Docker
|
||||
run: |
|
||||
apt-get update
|
||||
apt-get install -y nodejs docker.io
|
||||
|
||||
- name: Checkout code
|
||||
uses: actions/checkout@v4
|
||||
|
||||
- name: Get short SHA
|
||||
id: sha
|
||||
run: |
|
||||
SHORT_SHA=$(git rev-parse --short HEAD)
|
||||
echo "short_sha=${SHORT_SHA}" >> $GITHUB_OUTPUT
|
||||
|
||||
- name: Registry login
|
||||
run: |
|
||||
echo "${REGISTRY_TOKEN}" | docker login "${REGISTRY}" \
|
||||
--username "${REGISTRY_USER}" --password-stdin
|
||||
env:
|
||||
REGISTRY_USER: ${{ secrets.FORGEJO_REGISTRY_USER }}
|
||||
REGISTRY_TOKEN: ${{ secrets.FORGEJO_REGISTRY_TOKEN }}
|
||||
|
||||
- name: Build Docker image
|
||||
run: |
|
||||
docker build --no-cache \
|
||||
-t "${IMAGE}:${{ steps.sha.outputs.short_sha }}" \
|
||||
-t "${IMAGE}:latest" \
|
||||
.
|
||||
|
||||
- name: Push Docker image
|
||||
run: |
|
||||
docker push "${IMAGE}:${{ steps.sha.outputs.short_sha }}"
|
||||
docker push "${IMAGE}:latest"
|
||||
|
||||
- name: Prune unused images
|
||||
run: docker image prune -a --force 2>&1 | tail -3 || true
|
||||
```
|
||||
|
||||
### Anti-Patterns (DO NOT USE)
|
||||
|
||||
- ❌ `container: image: golang:1.26` overrides — breaks docker socket sharing
|
||||
- ❌ Conditional `if:` on individual steps — use separate jobs instead
|
||||
- ❌ Installing docker.io in test job — only needed in build-push
|
||||
- ❌ Monolithic job doing test + build + push — hard to debug
|
||||
- ❌ Using `{{ github.sha }}` for image tag — use short commit SHA for readability
|
||||
|
||||
### How It Works
|
||||
|
||||
1. **PR to feature branch** → test job runs, build-push skipped, nothing pushed
|
||||
2. **Push to main** → test runs, build-push runs after test passes, image pushed
|
||||
3. Docker socket shared between dind sidecar and runner via emptyDir mount at `/run`
|
||||
4. `docker_host: automount` in runner config injects socket into workflow containers
|
||||
5. Secrets (FORGEJO_REGISTRY_USER, TOKEN) set in Forgejo repo settings, NOT in git
|
||||
|
||||
@@ -42,86 +42,6 @@ All logs + metrics centralized in Grafana for debugging
|
||||
- **Secrets at rest** — Vault + encrypted etcd; credentials never in logs or ConfigMaps
|
||||
- **Infrastructure-as-code** — Every service deployed via Helmfile; one `helmfile apply` recovers from total failure
|
||||
|
||||
## ArgoCD — GitOps Deployment Flow
|
||||
|
||||
**ArgoCD** pulls infrastructure changes from git and syncs the cluster automatically.
|
||||
No manual `kubectl apply` — push to git, ArgoCD detects the change, and deploys within ~3 minutes.
|
||||
|
||||
```
|
||||
Developer pushes to git
|
||||
↓
|
||||
ArgoCD detects change (every 3 min or webhook)
|
||||
↓
|
||||
Syncs manifests to cluster
|
||||
↓
|
||||
Workloads reconcile automatically
|
||||
```
|
||||
|
||||
Applications are deployed in waves (numbered 00, 10, 20, 30, ...) to respect dependencies —
|
||||
storage deploys before databases, databases before applications.
|
||||
|
||||
### Tracked Git Repositories
|
||||
|
||||
ArgoCD monitors these repos for changes:
|
||||
|
||||
| Repository | Purpose |
|
||||
|------------|----------|
|
||||
| `https://github.com/Riotpiaole/riotpiao.homelab.com` | Main infrastructure repo (all manifests in `k8s/argocd/apps/`) |
|
||||
| `https://forgejo.riotpiao.com/rock/*` | Any `rock/*` repo in in-cluster Forgejo (apps + configs) |
|
||||
| `https://github.com/Riotpiaole/Poimen-*` | External Poimen services (memory, workflows) |
|
||||
|
||||
To deploy a new application: create a git repo, add an Application manifest to the homelab repo's
|
||||
`k8s/argocd/apps/`, commit + push, and ArgoCD syncs within 3 minutes.
|
||||
|
||||
## Management Planes — Talos vs Kubernetes
|
||||
|
||||
This cluster has **two separate management planes**, each with different workflows:
|
||||
|
||||
| Plane | What it manages | Workflow | Tool |
|
||||
|-------|-----------------|----------|------|
|
||||
| **Talos (OS)** | Node configuration, kernel params, networking, CoreDNS, machine state | Edit `terraform/` → `terraform apply` → `make apply-cp` | `terraform` + `talosctl` |
|
||||
| **Kubernetes (workloads)** | All pods, services, deployments, ingresses, databases | Edit `k8s/argocd/apps/` → `git push` → ArgoCD syncs | `git` + ArgoCD |
|
||||
|
||||
**Critical distinction:**
|
||||
- **Kubernetes resources** (`k8s/**`) flow through **git → ArgoCD** — never use `kubectl apply`
|
||||
- **Talos machine config** (`terraform/**`) uses **local `terraform apply`** (sanctioned exception — CI can't hold node credentials)
|
||||
|
||||
Example: To add a CoreDNS hostname rewrite, you edit `terraform/files/coredns/Corefile`, then:
|
||||
```bash
|
||||
cd terraform && terraform apply -var-file=terraform.tfvars.local
|
||||
cd .. && make apply-cp # talosctl apply-config to all 3 control planes
|
||||
```
|
||||
|
||||
But to add a new Kubernetes Deployment or update an Ingress, you only `git push` — **never `kubectl apply`**.
|
||||
|
||||
### CoreDNS ConfigMap Ownership — Critical
|
||||
|
||||
⚠️ **Warning:** The `coredns` ConfigMap in `kube-system` namespace is **owned by Talos**, not ArgoCD or kubectl.
|
||||
It is rendered from `terraform/files/coredns/Corefile` into Talos's machine config at bootstrap time.
|
||||
|
||||
**Do not `kubectl apply` or `kubectl edit` this ConfigMap directly.** Doing so transfers field ownership to kubectl's
|
||||
client-side-apply mechanism, and Talos's inline-manifest controller will silently no-op on every future reconcile
|
||||
(server-side-apply conflict, no error surfaced).
|
||||
|
||||
**To update CoreDNS (e.g., add a hostname rewrite):**
|
||||
1. Edit `terraform/files/coredns/Corefile`
|
||||
2. Commit + push
|
||||
3. Run `cd terraform && terraform apply -var-file=terraform.tfvars.local`
|
||||
4. Run `make apply-cp` to push config to all control planes
|
||||
5. CoreDNS picks up changes via its `reload` plugin — no pod restart needed
|
||||
|
||||
**If you accidentally edited the ConfigMap directly and broke Talos's ownership:**
|
||||
```bash
|
||||
kubectl delete configmap coredns -n kube-system
|
||||
# Wait ~30s for Talos's k8s.ManifestApplyController to recreate it
|
||||
kubectl get configmap coredns -n kube-system -w
|
||||
```
|
||||
|
||||
Or as a stopgap, apply the correct content yourself:
|
||||
```bash
|
||||
kubectl apply --server-side -f <(terraform output coredns_config)
|
||||
```
|
||||
|
||||
## Quick Start — Deploying the Cluster
|
||||
|
||||
### 1. Bootstrap Talos Nodes
|
||||
|
||||
@@ -1,14 +0,0 @@
|
||||
apiVersion: v1
|
||||
kind: ConfigMap
|
||||
metadata:
|
||||
name: immich-config
|
||||
data:
|
||||
DB_HOSTNAME: "immich-db-rw"
|
||||
DB_DATABASE_NAME: "immich"
|
||||
# Only pgvector is installed (see db.yaml) - no vectorchord extension image
|
||||
# exists for pg18 in CNPG's catalog yet. Explicit instead of relying on
|
||||
# auto-detect's vectorchord-first preference order.
|
||||
DB_VECTOR_EXTENSION: "pgvector"
|
||||
REDIS_HOSTNAME: "immich-redis"
|
||||
IMMICH_MACHINE_LEARNING_URL: "http://immich-machine-learning:3003"
|
||||
TZ: "America/Los_Angeles"
|
||||
@@ -1,58 +0,0 @@
|
||||
# Dedicated CNPG Postgres for Immich. Same recipe as paperless-db/authentik-db
|
||||
# (2 instances, default longhorn storage class) except the operand is
|
||||
# PostgreSQL 18, not 16.2 - the official CNPG pgvector extension image
|
||||
# (ghcr.io/cloudnative-pg/pgvector) is only published for pg18, no pg16 tags
|
||||
# exist in that registry. Immich itself supports pg18 fine (immich-app's own
|
||||
# postgres image already ships 18-vectorchord builds).
|
||||
#
|
||||
# pgvector loaded via CNPG's ImageVolume extension mechanism (CNPG 1.27+,
|
||||
# k8s ImageVolume feature - both present here: operator is 1.30.0, cluster is
|
||||
# v1.36.1). No shared_preload_libraries needed - pgvector doesn't require
|
||||
# preload, just CREATE EXTENSION, which immich-server issues itself at
|
||||
# startup. Distro/pg-major must match between the operand image and the
|
||||
# extension image (both "18"+"trixie" here) - CNPG's own compatibility rule.
|
||||
apiVersion: postgresql.cnpg.io/v1
|
||||
kind: Cluster
|
||||
metadata:
|
||||
name: immich-db
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-options: SkipDryRunOnMissingResource=true
|
||||
spec:
|
||||
instances: 2
|
||||
imageName: ghcr.io/cloudnative-pg/postgresql:18-minimal-trixie
|
||||
postgresql:
|
||||
extensions:
|
||||
- name: pgvector
|
||||
image:
|
||||
reference: ghcr.io/cloudnative-pg/pgvector:0.8.1-18-trixie
|
||||
bootstrap:
|
||||
initdb:
|
||||
database: immich
|
||||
owner: app
|
||||
encoding: UTF8
|
||||
localeCollate: C
|
||||
localeCType: C
|
||||
# CREATE EXTENSION vector requires superuser (pgvector's control file
|
||||
# isn't marked trusted) and the "app" owner role isn't one
|
||||
# (enableSuperuserAccess: false, repo convention) - postInitApplicationSQL
|
||||
# runs as superuser during initdb, before the app ever connects. Only
|
||||
# fires on a fresh bootstrap; the live cluster already had this run
|
||||
# manually once (kubectl exec ... psql -U postgres -c 'CREATE EXTENSION').
|
||||
postInitApplicationSQL:
|
||||
- "CREATE EXTENSION IF NOT EXISTS vector;"
|
||||
- "CREATE EXTENSION IF NOT EXISTS cube;"
|
||||
- "CREATE EXTENSION IF NOT EXISTS earthdistance;"
|
||||
enableSuperuserAccess: false
|
||||
resources:
|
||||
requests: { memory: "512Mi", cpu: "250m" }
|
||||
limits: { memory: "2Gi", cpu: "1" }
|
||||
storage:
|
||||
size: 20Gi
|
||||
storageClass: longhorn
|
||||
affinity:
|
||||
podAntiAffinityType: preferred
|
||||
topologyKey: kubernetes.io/hostname
|
||||
tolerations:
|
||||
- key: node-role.kubernetes.io/control-plane
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
@@ -1,100 +0,0 @@
|
||||
# immich-server: pinned to talos-cp-3, same reasoning as paperless
|
||||
# (deployment.yaml comment there) - immich-media is a ReadWriteOnce Longhorn
|
||||
# volume with a single replica physically on that node's disk (shared with
|
||||
# paperless-media on the same 4TB HDD). Recreate strategy for the same
|
||||
# reason: two pods can't both attach an RWO volume.
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: immich-server
|
||||
spec:
|
||||
replicas: 1
|
||||
strategy:
|
||||
type: Recreate
|
||||
selector:
|
||||
matchLabels:
|
||||
app: immich-server
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
app: immich-server
|
||||
spec:
|
||||
serviceAccountName: immich
|
||||
nodeSelector:
|
||||
kubernetes.io/hostname: talos-cp-3
|
||||
containers:
|
||||
- name: immich-server
|
||||
image: ghcr.io/immich-app/immich-server:release
|
||||
ports:
|
||||
- containerPort: 2283
|
||||
envFrom:
|
||||
- configMapRef:
|
||||
name: immich-config
|
||||
env:
|
||||
- name: DB_USERNAME
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: immich-db-app
|
||||
key: username
|
||||
- name: DB_PASSWORD
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: immich-db-app
|
||||
key: password
|
||||
# Composed by k8s/infra/iam's provisioning script (system-config
|
||||
# JSON, oauth section) - see immich-oidc Secret.
|
||||
- name: IMMICH_CONFIG_FILE
|
||||
value: /config/immich.json
|
||||
resources:
|
||||
requests: { cpu: "500m", memory: "1Gi" }
|
||||
limits: { cpu: "2", memory: "4Gi" }
|
||||
volumeMounts:
|
||||
- name: media
|
||||
mountPath: /usr/src/app/upload
|
||||
- name: oidc-config
|
||||
mountPath: /config
|
||||
readOnly: true
|
||||
volumes:
|
||||
- name: media
|
||||
persistentVolumeClaim:
|
||||
claimName: immich-media
|
||||
- name: oidc-config
|
||||
secret:
|
||||
secretName: immich-oidc
|
||||
items:
|
||||
- key: config.json
|
||||
path: immich.json
|
||||
---
|
||||
# CPU-only for now - the cluster's one GPU node (worker-1) is already
|
||||
# dedicated to llm-serving predictors. Not node-pinned: its cache PVC is on
|
||||
# the default 3-replica pool, not the single-disk cp-3 HDD.
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: immich-machine-learning
|
||||
spec:
|
||||
replicas: 1
|
||||
selector:
|
||||
matchLabels:
|
||||
app: immich-machine-learning
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
app: immich-machine-learning
|
||||
spec:
|
||||
serviceAccountName: immich
|
||||
containers:
|
||||
- name: immich-machine-learning
|
||||
image: ghcr.io/immich-app/immich-machine-learning:release
|
||||
ports:
|
||||
- containerPort: 3003
|
||||
resources:
|
||||
requests: { cpu: "500m", memory: "1Gi" }
|
||||
limits: { cpu: "2", memory: "4Gi" }
|
||||
volumeMounts:
|
||||
- name: ml-cache
|
||||
mountPath: /cache
|
||||
volumes:
|
||||
- name: ml-cache
|
||||
persistentVolumeClaim:
|
||||
claimName: immich-ml-cache
|
||||
@@ -1,24 +0,0 @@
|
||||
# Direct nginx ingress, same reasoning as paperless: large uploads (photos/
|
||||
# videos) and long-lived operations (video transcode, big batch uploads) need
|
||||
# proxy-body-size/timeouts raised past nginx's defaults.
|
||||
apiVersion: networking.k8s.io/v1
|
||||
kind: Ingress
|
||||
metadata:
|
||||
name: immich
|
||||
annotations:
|
||||
nginx.ingress.kubernetes.io/proxy-body-size: "0"
|
||||
nginx.ingress.kubernetes.io/proxy-read-timeout: "600"
|
||||
nginx.ingress.kubernetes.io/proxy-send-timeout: "600"
|
||||
spec:
|
||||
ingressClassName: nginx
|
||||
rules:
|
||||
- host: img.riotpiao.com
|
||||
http:
|
||||
paths:
|
||||
- path: /
|
||||
pathType: Prefix
|
||||
backend:
|
||||
service:
|
||||
name: immich-server
|
||||
port:
|
||||
number: 2283
|
||||
@@ -1,14 +0,0 @@
|
||||
apiVersion: kustomize.config.k8s.io/v1beta1
|
||||
kind: Kustomization
|
||||
namespace: immich
|
||||
resources:
|
||||
- db.yaml
|
||||
- pvc.yaml
|
||||
- configmap.yaml
|
||||
- redis.yaml
|
||||
- deployment.yaml
|
||||
- service.yaml
|
||||
- ingress.yaml
|
||||
- rbac.yaml
|
||||
# immich-oidc Secret written by the PostSync provisioning Job in
|
||||
# k8s/infra/iam (same as paperless-oidc) - not duplicated here.
|
||||
@@ -1,38 +0,0 @@
|
||||
# Two volumes:
|
||||
#
|
||||
# - media: original photos/videos + generated thumbnails/encoded videos.
|
||||
# Shares the cp-3 USB HDD with paperless-media, same StorageClass/disk tag,
|
||||
# single replica (single disk, no redundancy possible - same tradeoff
|
||||
# paperless already accepts). Sized 1400Gi, not 2000Gi: the disk's real
|
||||
# usable capacity (~3724GiB, formatting overhead) minus paperless-media's
|
||||
# 2000Gi and ~231GiB of other apps' default-class replicas that Longhorn
|
||||
# placed here anyway (disk tags only pull matching volumes in, they don't
|
||||
# exclude non-matching ones when the untagged pool elsewhere is full) only
|
||||
# leaves ~1493Gi of real scheduling headroom right now.
|
||||
# - ml-cache: downloaded ML model weights for immich-machine-learning
|
||||
# (face detection / CLIP embeddings). Small, disposable (re-downloads on
|
||||
# loss), but persisted so a pod restart doesn't re-pull multi-GB models -
|
||||
# default 3-replica pool, not node-pinned.
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: immich-media
|
||||
spec:
|
||||
accessModes:
|
||||
- ReadWriteOnce
|
||||
storageClassName: longhorn-paperless-media
|
||||
resources:
|
||||
requests:
|
||||
storage: 1400Gi
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: immich-ml-cache
|
||||
spec:
|
||||
accessModes:
|
||||
- ReadWriteOnce
|
||||
storageClassName: longhorn
|
||||
resources:
|
||||
requests:
|
||||
storage: 5Gi
|
||||
@@ -1,43 +0,0 @@
|
||||
# Scoped operator access for immich-admins: restart/config-edit rights on
|
||||
# just this service's own resources, nothing CNPG-managed (immich-db-*) or
|
||||
# provisioning-managed (immich-oidc). Same pattern as
|
||||
# k8s/apps/paperless/rbac.yaml. Inert until kube-apiserver's OIDC wiring
|
||||
# lands (--oidc-groups-claim=groups, --oidc-groups-prefix=oidc:).
|
||||
apiVersion: v1
|
||||
kind: ServiceAccount
|
||||
metadata:
|
||||
name: immich
|
||||
---
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: Role
|
||||
metadata:
|
||||
name: immich-operator
|
||||
rules:
|
||||
- apiGroups: ["apps"]
|
||||
resources: ["deployments"]
|
||||
resourceNames: ["immich-server", "immich-machine-learning"]
|
||||
verbs: ["get", "list", "watch", "update", "patch"]
|
||||
- apiGroups: [""]
|
||||
resources: ["configmaps"]
|
||||
resourceNames: ["immich-config"]
|
||||
verbs: ["get", "list", "watch", "update", "patch"]
|
||||
- apiGroups: [""]
|
||||
resources: ["secrets"]
|
||||
resourceNames: ["immich-oidc"]
|
||||
verbs: ["get", "list", "watch", "update", "patch"]
|
||||
---
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: RoleBinding
|
||||
metadata:
|
||||
name: immich-admins-binding
|
||||
subjects:
|
||||
- kind: Group
|
||||
name: "oidc:immich-admins"
|
||||
apiGroup: rbac.authorization.k8s.io
|
||||
- kind: ServiceAccount
|
||||
name: immich
|
||||
namespace: immich
|
||||
roleRef:
|
||||
kind: Role
|
||||
name: immich-operator
|
||||
apiGroup: rbac.authorization.k8s.io
|
||||
@@ -1,37 +0,0 @@
|
||||
# Job queue broker for immich-server. No PVC: queue state is disposable - a
|
||||
# lost queue on restart just re-triggers the affected background jobs
|
||||
# (thumbnail generation, ML jobs, etc.), no photo data loss since originals
|
||||
# live on immich-media.
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: immich-redis
|
||||
spec:
|
||||
replicas: 1
|
||||
selector:
|
||||
matchLabels:
|
||||
app: immich-redis
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
app: immich-redis
|
||||
spec:
|
||||
containers:
|
||||
- name: redis
|
||||
image: redis:7-alpine
|
||||
ports:
|
||||
- containerPort: 6379
|
||||
resources:
|
||||
requests: { cpu: "50m", memory: "64Mi" }
|
||||
limits: { cpu: "250m", memory: "256Mi" }
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: immich-redis
|
||||
spec:
|
||||
selector:
|
||||
app: immich-redis
|
||||
ports:
|
||||
- port: 6379
|
||||
targetPort: 6379
|
||||
@@ -1,21 +0,0 @@
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: immich-server
|
||||
spec:
|
||||
selector:
|
||||
app: immich-server
|
||||
ports:
|
||||
- port: 2283
|
||||
targetPort: 2283
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: immich-machine-learning
|
||||
spec:
|
||||
selector:
|
||||
app: immich-machine-learning
|
||||
ports:
|
||||
- port: 3003
|
||||
targetPort: 3003
|
||||
@@ -14,6 +14,5 @@ resources:
|
||||
- ornith.yaml
|
||||
- reasoning.yaml
|
||||
- reranker.yaml
|
||||
- networkpolicy.yaml
|
||||
# No namespace transformer: every file sets its own, and the transformer would
|
||||
# rewrite metadata.namespace on anything cross-namespace added later.
|
||||
|
||||
@@ -1,62 +0,0 @@
|
||||
# NetworkPolicy for LLM inference engines (llm-serving namespace).
|
||||
#
|
||||
# These pods have NO auth — vLLM, Ollama, and TEI accept any request.
|
||||
# All access MUST go through the api-gateway, which validates JWTs and
|
||||
# injects identity headers (X-Forwarded-User, X-Auth-Verified).
|
||||
#
|
||||
# Replaces the hand-applied llm-serving-default-deny policy that used
|
||||
# `llm-client: "true"` pod label as a selector — any pod in any namespace
|
||||
# could self-grant access by adding that label, which defeats the purpose.
|
||||
#
|
||||
# This policy restricts ingress to:
|
||||
# 1. api namespace (gateway) — the sole entry point for inference
|
||||
# 2. monitoring namespace — Prometheus scraping vLLM/TEI /metrics
|
||||
# 3. intra-namespace — pod-to-pod (future: multi-replica comms)
|
||||
apiVersion: networking.k8s.io/v1
|
||||
kind: NetworkPolicy
|
||||
metadata:
|
||||
name: llm-serving-ingress
|
||||
namespace: llm-serving
|
||||
labels:
|
||||
app.kubernetes.io/part-of: llm-serving
|
||||
spec:
|
||||
podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/part-of: llm-serving
|
||||
policyTypes:
|
||||
- Ingress
|
||||
ingress:
|
||||
# Allow from api-gateway (namespace: api)
|
||||
# Gateway proxies /v1/chat/completions, /v1/embeddings, /v1/rerank
|
||||
- from:
|
||||
- namespaceSelector:
|
||||
matchLabels:
|
||||
kubernetes.io/metadata.name: api
|
||||
ports:
|
||||
- protocol: TCP
|
||||
port: 8080 # vLLM, Ollama HTTP
|
||||
- protocol: TCP
|
||||
port: 80 # KServe predictor services
|
||||
- protocol: TCP
|
||||
port: 8000 # vLLM direct (some configs)
|
||||
- protocol: TCP
|
||||
port: 11434 # Ollama native port
|
||||
# Allow Prometheus scraping from monitoring namespace
|
||||
# vLLM: :8080/metrics, TEI: :9000/metrics
|
||||
- from:
|
||||
- namespaceSelector:
|
||||
matchLabels:
|
||||
kubernetes.io/metadata.name: monitoring
|
||||
ports:
|
||||
- protocol: TCP
|
||||
port: 8080
|
||||
- protocol: TCP
|
||||
port: 9000
|
||||
# Allow intra-namespace (pod-to-pod within llm-serving)
|
||||
- from:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/part-of: llm-serving
|
||||
ports:
|
||||
- protocol: TCP
|
||||
port: 8080
|
||||
@@ -13,7 +13,6 @@ spec:
|
||||
labels:
|
||||
app: management-service
|
||||
spec:
|
||||
serviceAccountName: kmsvc
|
||||
topologySpreadConstraints:
|
||||
- maxSkew: 1
|
||||
topologyKey: kubernetes.io/hostname
|
||||
|
||||
@@ -2,7 +2,7 @@ namespace: sqs
|
||||
replicaCount: 3
|
||||
|
||||
image:
|
||||
repository: forgejo.riotpiao.com/rock/kmsvc-manage
|
||||
repository: ghcr.io/riotpiaole/kmsvc-management-service
|
||||
tag: latest
|
||||
pullPolicy: Always
|
||||
|
||||
|
||||
@@ -1,6 +0,0 @@
|
||||
apiVersion: v2
|
||||
name: memory-queues
|
||||
description: Kafka queues (DLQ) for Poimen Memory service (Phase 6.6)
|
||||
type: application
|
||||
version: 0.1.0
|
||||
appVersion: "1.0"
|
||||
@@ -1,20 +0,0 @@
|
||||
{{- range .Values.queues }}
|
||||
---
|
||||
apiVersion: kmsvc.io/v1alpha1
|
||||
kind: Queue
|
||||
metadata:
|
||||
name: {{ .name }}
|
||||
namespace: {{ $.Values.namespace }}
|
||||
labels:
|
||||
app: memory-service
|
||||
queue: dlq
|
||||
spec:
|
||||
name: {{ .name }}
|
||||
description: {{ .description }}
|
||||
partitions: {{ .partitions }}
|
||||
replicationFactor: {{ .replicationFactor }}
|
||||
config:
|
||||
retention.ms: "{{ .config.retention.ms }}"
|
||||
message.retention.seconds: "{{ .config.message.retention.seconds }}"
|
||||
visibility.timeout.seconds: "{{ .config.visibility.timeout.seconds }}"
|
||||
{{- end }}
|
||||
@@ -1,25 +0,0 @@
|
||||
# Poimen Memory Service Kafka Queues (kmsvc)
|
||||
# Phase 6.6: DLQ topics for webhook + metrics failures
|
||||
|
||||
queues:
|
||||
# DLQ for extraction, webhook, and agent failures
|
||||
- name: poimen-memory-dlq
|
||||
description: "DLQ for extraction, webhook, and agent failures"
|
||||
partitions: 3
|
||||
replicationFactor: 1
|
||||
config:
|
||||
retention.ms: "1209600000" # 14 days
|
||||
message.retention.seconds: "1209600"
|
||||
visibility.timeout.seconds: "300"
|
||||
|
||||
# DLQ for metrics persistence failures
|
||||
- name: poimen-memory-metric-dlq
|
||||
description: "DLQ for metrics persistence failures"
|
||||
partitions: 3
|
||||
replicationFactor: 1
|
||||
config:
|
||||
retention.ms: "1209600000" # 14 days
|
||||
message.retention.seconds: "1209600"
|
||||
visibility.timeout.seconds: "300"
|
||||
|
||||
namespace: sqs
|
||||
@@ -1,7 +1,7 @@
|
||||
namespace: sqs
|
||||
|
||||
image:
|
||||
repository: forgejo.riotpiao.com/rock/kmsvc-manage
|
||||
repository: ghcr.io/riotpiaole/kmsvc-management-service
|
||||
tag: latest
|
||||
pullPolicy: Always
|
||||
|
||||
|
||||
@@ -1,95 +0,0 @@
|
||||
# Overrides paperless-ngx's own paperless/adapter.py at the same import path
|
||||
# (mounted via subPath in deployment.yaml) - settings.py hardcodes
|
||||
# SOCIALACCOUNT_ADAPTER = "paperless.adapter.CustomSocialAccountAdapter", so
|
||||
# no Django setting needs to change, just the file content underneath it.
|
||||
#
|
||||
# Stock CustomSocialAccountAdapter.populate_user() is a stub ("kept in case
|
||||
# global default permissions are implemented in the future" - they aren't),
|
||||
# so every OIDC signup lands with zero permissions and 403s on every API
|
||||
# endpoint. This adds the actual mapping: Authentik's "permissions" claim
|
||||
# (via the permissions scope, requested in PAPERLESS_SOCIALACCOUNT_PROVIDERS,
|
||||
# computed server-side from group membership by authentik-provision.py) ->
|
||||
# "paperless:write" or "*" (homelab-admins) grants is_staff+is_superuser,
|
||||
# same convention already used for MinIO's policy claim and Grafana's
|
||||
# role_attribute_path. Checking the permission string rather than a literal
|
||||
# group name decouples "what grants access" from which group happens to
|
||||
# hold it - same pattern applies to every other service's Role/RoleBinding
|
||||
# in k8s/infra/rbac/.
|
||||
apiVersion: v1
|
||||
kind: ConfigMap
|
||||
metadata:
|
||||
name: paperless-adapter
|
||||
data:
|
||||
adapter.py: |
|
||||
from urllib.parse import quote
|
||||
|
||||
from allauth.account.adapter import DefaultAccountAdapter
|
||||
from allauth.core import context
|
||||
from allauth.socialaccount.adapter import DefaultSocialAccountAdapter
|
||||
from django.conf import settings
|
||||
from django.forms import ValidationError
|
||||
from django.urls import reverse
|
||||
|
||||
REQUIRED_PERMISSIONS = {"paperless:write", "*"}
|
||||
|
||||
|
||||
class CustomAccountAdapter(DefaultAccountAdapter):
|
||||
def is_open_for_signup(self, request):
|
||||
allow_signups = super().is_open_for_signup(request)
|
||||
return getattr(settings, "ACCOUNT_ALLOW_SIGNUPS", allow_signups)
|
||||
|
||||
def pre_authenticate(self, request, **credentials):
|
||||
if settings.DISABLE_REGULAR_LOGIN:
|
||||
raise ValidationError("Regular login is disabled")
|
||||
return super().pre_authenticate(request, **credentials)
|
||||
|
||||
def is_safe_url(self, url):
|
||||
from django.utils.http import url_has_allowed_host_and_scheme
|
||||
|
||||
allowed_hosts = {context.request.get_host()} | set(settings.ALLOWED_HOSTS)
|
||||
if "*" in allowed_hosts:
|
||||
allowed_hosts.remove("*")
|
||||
allowed_hosts.add(context.request.get_host())
|
||||
return url_has_allowed_host_and_scheme(url, allowed_hosts=allowed_hosts)
|
||||
return url_has_allowed_host_and_scheme(url, allowed_hosts=allowed_hosts)
|
||||
|
||||
def get_reset_password_from_key_url(self, key):
|
||||
if settings.PAPERLESS_URL is None:
|
||||
return super().get_reset_password_from_key_url(key)
|
||||
path = reverse(
|
||||
"account_reset_password_from_key",
|
||||
kwargs={"uidb36": "UID", "key": "KEY"},
|
||||
)
|
||||
path = path.replace("UID-KEY", quote(key))
|
||||
return settings.PAPERLESS_URL + path
|
||||
|
||||
|
||||
class CustomSocialAccountAdapter(DefaultSocialAccountAdapter):
|
||||
def is_open_for_signup(self, request, sociallogin):
|
||||
allow_signups = super().is_open_for_signup(request, sociallogin)
|
||||
return getattr(settings, "SOCIALACCOUNT_ALLOW_SIGNUPS", allow_signups)
|
||||
|
||||
def get_connect_redirect_url(self, request, socialaccount):
|
||||
return reverse("base")
|
||||
|
||||
def populate_user(self, request, sociallogin, data):
|
||||
user = super().populate_user(request, sociallogin, data)
|
||||
perms = set(sociallogin.account.extra_data.get("permissions") or [])
|
||||
if perms & REQUIRED_PERMISSIONS:
|
||||
user.is_staff = True
|
||||
user.is_superuser = True
|
||||
return user
|
||||
|
||||
def save_user(self, request, sociallogin, form=None):
|
||||
# populate_user() sets the flags on the in-memory user, but
|
||||
# allauth's default save_user() re-derives is_staff from
|
||||
# ACCOUNT_DEFAULT_HTTP_PROTOCOL-independent defaults and can
|
||||
# overwrite them on save - re-apply after super().save_user()
|
||||
# persists the row, matching the permissions check above exactly.
|
||||
user = super().save_user(request, sociallogin, form)
|
||||
perms = set(sociallogin.account.extra_data.get("permissions") or [])
|
||||
if perms & REQUIRED_PERMISSIONS and not (user.is_staff and user.is_superuser):
|
||||
user.is_staff = True
|
||||
user.is_superuser = True
|
||||
user.save(update_fields=["is_staff", "is_superuser"])
|
||||
return user
|
||||
@@ -1,94 +0,0 @@
|
||||
# Nightly: pg_dump the paperless DB + mirror the media PVC into the scoped
|
||||
# `paperless` MinIO bucket (see minio-provision-paperless-job.yaml). This is a
|
||||
# BACKUP target, not live storage - paperless-ngx has no native S3 backend, it
|
||||
# only ever reads/writes the local media PVC directly.
|
||||
#
|
||||
# Pinned to talos-cp-3, same as deployment.yaml: media is a ReadWriteOnce
|
||||
# Longhorn volume with a single replica physically on that node's disk -
|
||||
# mounting it read-only here from a different node would conflict with the
|
||||
# live webserver's attachment.
|
||||
apiVersion: batch/v1
|
||||
kind: CronJob
|
||||
metadata:
|
||||
name: paperless-backup
|
||||
spec:
|
||||
schedule: "0 3 * * *" # 03:00 daily, low-traffic window
|
||||
jobTemplate:
|
||||
spec:
|
||||
backoffLimit: 2
|
||||
template:
|
||||
spec:
|
||||
restartPolicy: Never
|
||||
nodeSelector:
|
||||
kubernetes.io/hostname: talos-cp-3
|
||||
initContainers:
|
||||
- name: pg-dump
|
||||
image: postgres:16-alpine
|
||||
env:
|
||||
- name: PGHOST
|
||||
value: paperless-db-rw
|
||||
- name: PGDATABASE
|
||||
value: paperless
|
||||
- name: PGUSER
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: paperless-db-app
|
||||
key: username
|
||||
- name: PGPASSWORD
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: paperless-db-app
|
||||
key: password
|
||||
command:
|
||||
- sh
|
||||
- -c
|
||||
- pg_dump --format=custom --file=/backup/paperless-db.dump
|
||||
volumeMounts:
|
||||
- name: backup
|
||||
mountPath: /backup
|
||||
containers:
|
||||
- name: mc-mirror
|
||||
image: minio/mc:latest
|
||||
env:
|
||||
- name: ACCESS_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: paperless-minio-creds
|
||||
key: ACCESS_KEY
|
||||
- name: SECRET_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: paperless-minio-creds
|
||||
key: SECRET_KEY
|
||||
- name: BUCKET
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: paperless-minio-creds
|
||||
key: BUCKET
|
||||
- name: ENDPOINT
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: paperless-minio-creds
|
||||
key: ENDPOINT
|
||||
command:
|
||||
- /bin/sh
|
||||
- -c
|
||||
- |
|
||||
set -e
|
||||
mc alias set b "$ENDPOINT" "$ACCESS_KEY" "$SECRET_KEY"
|
||||
mc cp /backup/paperless-db.dump "b/$BUCKET/db/paperless-db-$(date +%Y%m%d).dump"
|
||||
mc mirror --overwrite /media "b/$BUCKET/media"
|
||||
echo "Backup done."
|
||||
volumeMounts:
|
||||
- name: backup
|
||||
mountPath: /backup
|
||||
- name: media
|
||||
mountPath: /media
|
||||
readOnly: true
|
||||
volumes:
|
||||
- name: backup
|
||||
emptyDir: {}
|
||||
- name: media
|
||||
persistentVolumeClaim:
|
||||
claimName: paperless-media
|
||||
readOnly: true
|
||||
@@ -1,22 +0,0 @@
|
||||
apiVersion: v1
|
||||
kind: ConfigMap
|
||||
metadata:
|
||||
name: paperless-config
|
||||
data:
|
||||
PAPERLESS_URL: "https://paperless.riotpiao.com"
|
||||
PAPERLESS_TIME_ZONE: "America/Los_Angeles"
|
||||
PAPERLESS_OCR_LANGUAGE: "eng"
|
||||
PAPERLESS_DBHOST: "paperless-db-rw"
|
||||
PAPERLESS_DBNAME: "paperless"
|
||||
PAPERLESS_REDIS: "redis://paperless-redis:6379"
|
||||
# django-allauth generic OIDC provider. The client_id/secret/server_url
|
||||
# bundle itself lives in the paperless-oidc Secret
|
||||
# (SOCIALACCOUNT_PROVIDERS_JSON key, composed by authentik-provision.py) -
|
||||
# env vars can't be split across a ConfigMap + Secret for the same key, so
|
||||
# this whole value is sourced from the Secret in deployment.yaml instead.
|
||||
PAPERLESS_APPS: "allauth.socialaccount.providers.openid_connect"
|
||||
# Authentik already verifies identity via OIDC - a second email-confirmation
|
||||
# step has no SMTP configured to send it anyway, and paperless-ngx doesn't
|
||||
# wire up allauth's confirm-email view, so signup 500s with NoReverseMatch
|
||||
# on 'account_confirm_email' without this.
|
||||
PAPERLESS_ACCOUNT_EMAIL_VERIFICATION: "none"
|
||||
@@ -1,103 +0,0 @@
|
||||
# Single container runs webserver + consumer + scheduler (paperless-ngx's
|
||||
# stock entrypoint does this internally) - no need to split into separate
|
||||
# Deployments. replicas: 1 only: paperless-media is ReadWriteOnce, and the
|
||||
# consumer polling the media dir doesn't benefit from horizontal scaling here.
|
||||
#
|
||||
# Pinned to talos-cp-3: paperless-media's disk physically lives there. Longhorn
|
||||
# RWO volumes can only be attached from one node at a time, and the nightly
|
||||
# backup-cronjob.yaml also mounts this same PVC (read-only) to mirror it into
|
||||
# MinIO - pinning both to the same node avoids a cross-node attach conflict,
|
||||
# and keeps the 3.5Ti read/write path off the network entirely.
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: paperless
|
||||
spec:
|
||||
replicas: 1
|
||||
strategy:
|
||||
type: Recreate # ReadWriteOnce media PVC - avoid two pods fighting over it
|
||||
selector:
|
||||
matchLabels:
|
||||
app: paperless
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
app: paperless
|
||||
spec:
|
||||
# Kubernetes injects legacy Docker-links env vars for every Service in
|
||||
# this namespace (<SVC>_SERVICE_HOST, <SVC>_PORT, ...). The Service here
|
||||
# is named "paperless", so that becomes PAPERLESS_PORT=tcp://<ip>:8000 -
|
||||
# paperless-ngx's own entrypoint reads PAPERLESS_PORT for gunicorn's
|
||||
# bind address, collides, and gunicorn crash-loops on "not a valid port
|
||||
# number". Disable the injection instead of renaming the Service.
|
||||
enableServiceLinks: false
|
||||
nodeSelector:
|
||||
kubernetes.io/hostname: talos-cp-3
|
||||
containers:
|
||||
- name: paperless
|
||||
image: ghcr.io/paperless-ngx/paperless-ngx:2.20.15
|
||||
ports:
|
||||
- containerPort: 8000
|
||||
envFrom:
|
||||
- configMapRef:
|
||||
name: paperless-config
|
||||
env:
|
||||
- name: PAPERLESS_DBUSER
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: paperless-db-app
|
||||
key: username
|
||||
- name: PAPERLESS_DBPASS
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: paperless-db-app
|
||||
key: password
|
||||
- name: PAPERLESS_SECRET_KEY
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: paperless-secrets
|
||||
key: PAPERLESS_SECRET_KEY
|
||||
- name: PAPERLESS_ADMIN_USER
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: paperless-secrets
|
||||
key: PAPERLESS_ADMIN_USER
|
||||
- name: PAPERLESS_ADMIN_PASSWORD
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: paperless-secrets
|
||||
key: PAPERLESS_ADMIN_PASSWORD
|
||||
- name: PAPERLESS_SOCIALACCOUNT_PROVIDERS
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: paperless-oidc
|
||||
key: SOCIALACCOUNT_PROVIDERS_JSON
|
||||
resources:
|
||||
requests: { cpu: "500m", memory: "1Gi" }
|
||||
limits: { cpu: "2", memory: "4Gi" }
|
||||
volumeMounts:
|
||||
- name: media
|
||||
mountPath: /usr/src/paperless/media
|
||||
- name: data
|
||||
mountPath: /usr/src/paperless/data
|
||||
- name: consume
|
||||
mountPath: /usr/src/paperless/consume
|
||||
# Overrides paperless-ngx's own adapter.py in place - settings.py
|
||||
# hardcodes the import path, so no Django setting changes, just
|
||||
# the file content underneath it (see adapter-configmap.yaml).
|
||||
- name: adapter
|
||||
mountPath: /usr/src/paperless/src/paperless/adapter.py
|
||||
subPath: adapter.py
|
||||
readOnly: true
|
||||
volumes:
|
||||
- name: media
|
||||
persistentVolumeClaim:
|
||||
claimName: paperless-media
|
||||
- name: data
|
||||
persistentVolumeClaim:
|
||||
claimName: paperless-data
|
||||
- name: consume
|
||||
emptyDir: {}
|
||||
- name: adapter
|
||||
configMap:
|
||||
name: paperless-adapter
|
||||
@@ -1,24 +0,0 @@
|
||||
# Direct nginx ingress to the paperless Service - not routed via the Go
|
||||
# api-gateway (api.riotpiao.com), which has no WebSocket upgrade support and
|
||||
# paperless-ngx keeps a long-lived /ws/ connection open for live task status.
|
||||
apiVersion: networking.k8s.io/v1
|
||||
kind: Ingress
|
||||
metadata:
|
||||
name: paperless
|
||||
annotations:
|
||||
nginx.ingress.kubernetes.io/proxy-body-size: "0" # large scanned PDF uploads
|
||||
nginx.ingress.kubernetes.io/proxy-read-timeout: "600"
|
||||
nginx.ingress.kubernetes.io/proxy-send-timeout: "600"
|
||||
spec:
|
||||
ingressClassName: nginx
|
||||
rules:
|
||||
- host: paperless.riotpiao.com
|
||||
http:
|
||||
paths:
|
||||
- path: /
|
||||
pathType: Prefix
|
||||
backend:
|
||||
service:
|
||||
name: paperless
|
||||
port:
|
||||
number: 8000
|
||||
@@ -1,17 +0,0 @@
|
||||
apiVersion: kustomize.config.k8s.io/v1beta1
|
||||
kind: Kustomization
|
||||
namespace: paperless
|
||||
resources:
|
||||
- pvc.yaml
|
||||
- configmap.yaml
|
||||
- redis.yaml
|
||||
- deployment.yaml
|
||||
- service.yaml
|
||||
- ingress.yaml
|
||||
- backup-cronjob.yaml
|
||||
- adapter-configmap.yaml
|
||||
- rbac.yaml
|
||||
# postgres: paperless-db CNPG Cluster, deployed by k8s/infra/databases (wave 2,
|
||||
# before this app at wave 8) - not duplicated here. Same for the paperless-oidc
|
||||
# and paperless-minio-creds Secrets, written by PostSync provisioning Jobs in
|
||||
# k8s/infra/iam and k8s/infra/minio respectively.
|
||||
@@ -1,35 +0,0 @@
|
||||
# Two volumes, deliberately separate storage classes:
|
||||
#
|
||||
# - media: the actual documents (originals + OCR'd archive PDFs + thumbnails).
|
||||
# Lives on the cp-3 USB HDD, single replica (see
|
||||
# k8s/infra/longhorn/longhorn-paperless-storageclass.yaml). Shares the disk
|
||||
# with Immich's immich-media PVC (k8s/apps/immich/pvc.yaml, 2000Gi) - photo
|
||||
# libraries grow much faster than scanned documents, so paperless gets the
|
||||
# smaller 500Gi share.
|
||||
# - data: the SQLite classification model + search index. Small (low GB),
|
||||
# frequently rewritten, and disposable (rebuilds from the DB + media on
|
||||
# next consume) - stays on the default 3-replica pool instead of the
|
||||
# single-disk HDD.
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: paperless-media
|
||||
spec:
|
||||
accessModes:
|
||||
- ReadWriteOnce
|
||||
storageClassName: longhorn-paperless-media
|
||||
resources:
|
||||
requests:
|
||||
storage: 500Gi
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: paperless-data
|
||||
spec:
|
||||
accessModes:
|
||||
- ReadWriteOnce
|
||||
storageClassName: longhorn
|
||||
resources:
|
||||
requests:
|
||||
storage: 5Gi
|
||||
@@ -1,35 +0,0 @@
|
||||
# Scoped operator access for paperless-admins: restart/config-edit rights on
|
||||
# just this service's own resources, nothing CNPG-managed (paperless-db-*)
|
||||
# or provisioning-managed (paperless-oidc, paperless-minio-creds). Inert
|
||||
# until kube-apiserver's OIDC wiring lands (--oidc-groups-claim=groups,
|
||||
# --oidc-groups-prefix=oidc:) - subject name below assumes that prefix.
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: Role
|
||||
metadata:
|
||||
name: paperless-operator
|
||||
rules:
|
||||
- apiGroups: ["apps"]
|
||||
resources: ["deployments"]
|
||||
resourceNames: ["paperless"]
|
||||
verbs: ["get", "list", "watch", "update", "patch"]
|
||||
- apiGroups: [""]
|
||||
resources: ["configmaps"]
|
||||
resourceNames: ["paperless-config"]
|
||||
verbs: ["get", "list", "watch", "update", "patch"]
|
||||
- apiGroups: [""]
|
||||
resources: ["secrets"]
|
||||
resourceNames: ["paperless-secrets"]
|
||||
verbs: ["get", "list", "watch", "update", "patch"]
|
||||
---
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: RoleBinding
|
||||
metadata:
|
||||
name: paperless-admins-binding
|
||||
subjects:
|
||||
- kind: Group
|
||||
name: "oidc:paperless-admins"
|
||||
apiGroup: rbac.authorization.k8s.io
|
||||
roleRef:
|
||||
kind: Role
|
||||
name: paperless-operator
|
||||
apiGroup: rbac.authorization.k8s.io
|
||||
@@ -1,37 +0,0 @@
|
||||
# Task queue broker + websocket channel layer for paperless-ngx. No PVC:
|
||||
# queued/scheduled task state is disposable - a lost queue on restart just
|
||||
# means re-triggering consumption, not data loss (documents themselves live
|
||||
# on paperless-media).
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: paperless-redis
|
||||
spec:
|
||||
replicas: 1
|
||||
selector:
|
||||
matchLabels:
|
||||
app: paperless-redis
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
app: paperless-redis
|
||||
spec:
|
||||
containers:
|
||||
- name: redis
|
||||
image: redis:7-alpine
|
||||
ports:
|
||||
- containerPort: 6379
|
||||
resources:
|
||||
requests: { cpu: "50m", memory: "64Mi" }
|
||||
limits: { cpu: "250m", memory: "256Mi" }
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: paperless-redis
|
||||
spec:
|
||||
selector:
|
||||
app: paperless-redis
|
||||
ports:
|
||||
- port: 6379
|
||||
targetPort: 6379
|
||||
@@ -1,10 +0,0 @@
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: paperless
|
||||
spec:
|
||||
selector:
|
||||
app: paperless
|
||||
ports:
|
||||
- port: 8000
|
||||
targetPort: 8000
|
||||
@@ -1,144 +0,0 @@
|
||||
apiVersion: apiextensions.k8s.io/v1
|
||||
kind: CustomResourceDefinition
|
||||
metadata:
|
||||
name: secretrotations.homelab.riotpiao.com
|
||||
spec:
|
||||
group: homelab.riotpiao.com
|
||||
names:
|
||||
kind: SecretRotation
|
||||
plural: secretrotations
|
||||
scope: Namespaced
|
||||
versions:
|
||||
- name: v1
|
||||
served: true
|
||||
storage: true
|
||||
schema:
|
||||
openAPIV3Schema:
|
||||
type: object
|
||||
properties:
|
||||
metadata:
|
||||
type: object
|
||||
spec:
|
||||
type: object
|
||||
required:
|
||||
- provider
|
||||
- rotationInterval
|
||||
properties:
|
||||
# External system: authentik | forgejo | minio | vault
|
||||
provider:
|
||||
type: string
|
||||
enum: [authentik, forgejo, minio, vault]
|
||||
|
||||
# How often to rotate (hours)
|
||||
rotationInterval:
|
||||
type: integer
|
||||
minimum: 24
|
||||
|
||||
# Application ID in external system
|
||||
appId:
|
||||
type: string
|
||||
|
||||
# k8s Secret to update (name, namespace, key)
|
||||
secretRef:
|
||||
type: object
|
||||
required: [name, namespace]
|
||||
properties:
|
||||
name:
|
||||
type: string
|
||||
namespace:
|
||||
type: string
|
||||
key:
|
||||
type: string
|
||||
description: "Secret key to update (e.g., MINIO_IDENTITY_OPENID_CLIENT_SECRET)"
|
||||
|
||||
# Path to git file that holds the secret (for .enc.yaml files)
|
||||
gitPath:
|
||||
type: string
|
||||
description: "Path in homelab repo to .enc.yaml file"
|
||||
|
||||
# Ansible template values to substitute
|
||||
templateValues:
|
||||
type: object
|
||||
additionalProperties:
|
||||
type: string
|
||||
|
||||
status:
|
||||
type: object
|
||||
properties:
|
||||
lastRotationTime:
|
||||
type: string
|
||||
format: date-time
|
||||
nextRotationTime:
|
||||
type: string
|
||||
format: date-time
|
||||
lastRotationStatus:
|
||||
type: string
|
||||
enum: [Success, Failed, Pending]
|
||||
lastRotationError:
|
||||
type: string
|
||||
lastCommitHash:
|
||||
type: string
|
||||
|
||||
---
|
||||
# Example usage:
|
||||
apiVersion: homelab.riotpiao.com/v1
|
||||
kind: SecretRotation
|
||||
metadata:
|
||||
name: minio-oidc
|
||||
namespace: secret-rotation
|
||||
spec:
|
||||
provider: authentik
|
||||
rotationInterval: 2160 # 90 days in hours
|
||||
appId: minio
|
||||
secretRef:
|
||||
name: minio-oidc
|
||||
namespace: storage
|
||||
key: MINIO_IDENTITY_OPENID_CLIENT_SECRET
|
||||
gitPath: k8s/argocd/secrets/minio-oidc.enc.yaml
|
||||
|
||||
---
|
||||
apiVersion: homelab.riotpiao.com/v1
|
||||
kind: SecretRotation
|
||||
metadata:
|
||||
name: portfolio-agent-oidc
|
||||
namespace: secret-rotation
|
||||
spec:
|
||||
provider: authentik
|
||||
rotationInterval: 2160
|
||||
appId: portfolio-agent
|
||||
secretRef:
|
||||
name: portfolio-agent-oidc
|
||||
namespace: portfolio
|
||||
key: CLIENT_SECRET
|
||||
gitPath: k8s/argocd/secrets/portfolio-agent-oidc.enc.yaml
|
||||
|
||||
---
|
||||
apiVersion: homelab.riotpiao.com/v1
|
||||
kind: SecretRotation
|
||||
metadata:
|
||||
name: forgejo-registry-token
|
||||
namespace: secret-rotation
|
||||
spec:
|
||||
provider: forgejo
|
||||
rotationInterval: 2160
|
||||
appId: rock/riotpiao.com
|
||||
secretRef:
|
||||
name: forgejo-registry-secret
|
||||
namespace: kube-system
|
||||
key: REGISTRY_TOKEN
|
||||
gitPath: k8s/argocd/secrets/forgejo-registry-secret.enc.yaml
|
||||
|
||||
---
|
||||
apiVersion: homelab.riotpiao.com/v1
|
||||
kind: SecretRotation
|
||||
metadata:
|
||||
name: minio-root-credentials
|
||||
namespace: secret-rotation
|
||||
spec:
|
||||
provider: minio
|
||||
rotationInterval: 4320 # 180 days in hours
|
||||
appId: root
|
||||
secretRef:
|
||||
name: minio-creds
|
||||
namespace: storage
|
||||
gitPath: k8s/argocd/secrets/minio-secrets.enc.yaml
|
||||
@@ -1,92 +0,0 @@
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: secret-rotation-controller
|
||||
namespace: secret-rotation
|
||||
spec:
|
||||
replicas: 1
|
||||
selector:
|
||||
matchLabels:
|
||||
app: secret-rotation-controller
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
app: secret-rotation-controller
|
||||
spec:
|
||||
serviceAccountName: secret-rotation-controller
|
||||
containers:
|
||||
- name: controller
|
||||
image: secret-rotation-controller:latest
|
||||
imagePullPolicy: IfNotPresent
|
||||
env:
|
||||
# SOPS reads age key from this file
|
||||
- name: SOPS_AGE_KEY_FILE
|
||||
value: /etc/sops/age/private-key.txt
|
||||
|
||||
# Vault auth (token in projected volume)
|
||||
- name: VAULT_ADDR
|
||||
value: http://vault.vault.svc.cluster.local:8200
|
||||
- name: VAULT_TOKEN_FILE
|
||||
value: /var/run/secrets/vault/token
|
||||
|
||||
# Authentik
|
||||
- name: AUTHENTIK_URL
|
||||
value: http://authentik-server.iam.svc.cluster.local
|
||||
- name: AUTHENTIK_BOOTSTRAP_TOKEN
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: authentik-bootstrap
|
||||
key: token
|
||||
|
||||
# Git
|
||||
- name: GIT_REPO
|
||||
value: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
- name: GIT_AUTHOR_EMAIL
|
||||
value: [email protected]
|
||||
- name: GIT_AUTHOR_NAME
|
||||
value: Secret Rotation Controller
|
||||
- name: FORGEJO_TOKEN
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: forgejo-registry-secret
|
||||
key: REGISTRY_TOKEN
|
||||
|
||||
volumeMounts:
|
||||
# Age key from ExternalSecret (synced from Vault)
|
||||
- name: age-key
|
||||
mountPath: /etc/sops/age
|
||||
readOnly: true
|
||||
|
||||
# Vault auth token (projected)
|
||||
- name: vault-token
|
||||
mountPath: /var/run/secrets/vault
|
||||
readOnly: true
|
||||
|
||||
# Temp working dir
|
||||
- name: tmp
|
||||
mountPath: /tmp
|
||||
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 256Mi
|
||||
limits:
|
||||
cpu: 500m
|
||||
memory: 512Mi
|
||||
|
||||
volumes:
|
||||
- name: age-key
|
||||
secret:
|
||||
secretName: sops-age-key
|
||||
defaultMode: 0400
|
||||
|
||||
- name: vault-token
|
||||
projected:
|
||||
sources:
|
||||
- serviceAccountToken:
|
||||
path: token
|
||||
audience: vault
|
||||
expirationSeconds: 3600
|
||||
|
||||
- name: tmp
|
||||
emptyDir: {}
|
||||
@@ -1,15 +0,0 @@
|
||||
apiVersion: kustomize.config.k8s.io/v1beta1
|
||||
kind: Kustomization
|
||||
|
||||
namespace: secret-rotation
|
||||
|
||||
resources:
|
||||
- rbac.yaml
|
||||
- crd.yaml
|
||||
- external-secret.yaml
|
||||
- deployment.yaml
|
||||
|
||||
commonLabels:
|
||||
app.kubernetes.io/name: secret-rotation-controller
|
||||
app.kubernetes.io/component: automation
|
||||
managed-by: argocd
|
||||
@@ -1,53 +0,0 @@
|
||||
apiVersion: v1
|
||||
kind: ServiceAccount
|
||||
metadata:
|
||||
name: secret-rotation-controller
|
||||
namespace: secret-rotation
|
||||
|
||||
---
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: ClusterRole
|
||||
metadata:
|
||||
name: secret-rotation-controller
|
||||
rules:
|
||||
# Read SecretRotation CRDs
|
||||
- apiGroups: ["homelab.riotpiao.com"]
|
||||
resources: ["secretrotations"]
|
||||
verbs: ["get", "list", "watch"]
|
||||
|
||||
# Update status
|
||||
- apiGroups: ["homelab.riotpiao.com"]
|
||||
resources: ["secretrotations/status"]
|
||||
verbs: ["get", "patch", "update"]
|
||||
|
||||
# Read k8s secrets that will be rotated
|
||||
- apiGroups: [""]
|
||||
resources: ["secrets"]
|
||||
verbs: ["get", "list"]
|
||||
|
||||
# For recording events
|
||||
- apiGroups: [""]
|
||||
resources: ["events"]
|
||||
verbs: ["create", "patch"]
|
||||
|
||||
---
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: ClusterRoleBinding
|
||||
metadata:
|
||||
name: secret-rotation-controller
|
||||
roleRef:
|
||||
apiGroup: rbac.authorization.k8s.io
|
||||
kind: ClusterRole
|
||||
name: secret-rotation-controller
|
||||
subjects:
|
||||
- kind: ServiceAccount
|
||||
name: secret-rotation-controller
|
||||
namespace: secret-rotation
|
||||
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Namespace
|
||||
metadata:
|
||||
name: secret-rotation
|
||||
labels:
|
||||
kubernetes.io/metadata.name: secret-rotation
|
||||
@@ -18,7 +18,7 @@ spec:
|
||||
prune: true
|
||||
selfHeal: true
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/argocd/projects
|
||||
destination:
|
||||
|
||||
@@ -17,7 +17,7 @@ spec:
|
||||
syncOptions:
|
||||
- CreateNamespace=true
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
# ksops decrypts every *.enc.yaml here at kustomize-build time (repo-server
|
||||
# runs `kustomize build --enable-alpha-plugins --enable-exec`). Replaces the
|
||||
|
||||
@@ -21,7 +21,7 @@ spec:
|
||||
helm:
|
||||
valueFiles:
|
||||
- $values/k8s/bootstrap/cert-manager/cert-manager-values.yaml
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
ref: values
|
||||
destination:
|
||||
@@ -93,7 +93,7 @@ spec:
|
||||
project: homelab
|
||||
revisionHistoryLimit: 3
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
# A real kustomization.yaml (resources: the 3 issuer/CA files) renders these
|
||||
# deterministically. The previous directory.include with bare filenames
|
||||
@@ -127,7 +127,7 @@ spec:
|
||||
project: homelab
|
||||
revisionHistoryLimit: 3
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/bootstrap/ingress
|
||||
destination:
|
||||
@@ -138,75 +138,3 @@ spec:
|
||||
selfHeal: true
|
||||
syncOptions:
|
||||
- CreateNamespace=true
|
||||
---
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Application
|
||||
metadata:
|
||||
name: cluster-maintenance
|
||||
namespace: argocd
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-wave: "0"
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
targetRevision: main
|
||||
path: k8s/infra/cluster-maintenance
|
||||
destination:
|
||||
server: https://kubernetes.default.svc
|
||||
namespace: kube-system
|
||||
syncPolicy:
|
||||
automated:
|
||||
prune: true
|
||||
selfHeal: true
|
||||
---
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Application
|
||||
metadata:
|
||||
name: kyverno
|
||||
namespace: argocd
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-wave: "0"
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://kyverno.github.io/kyverno/
|
||||
chart: kyverno
|
||||
targetRevision: "1.14.0"
|
||||
helm:
|
||||
valueFiles:
|
||||
- $values/k8s/bootstrap/kyverno/kyverno-values.yaml
|
||||
sources:
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
targetRevision: main
|
||||
ref: values
|
||||
destination:
|
||||
server: https://kubernetes.default.svc
|
||||
namespace: kyverno
|
||||
syncPolicy:
|
||||
automated:
|
||||
prune: true
|
||||
selfHeal: true
|
||||
syncOptions:
|
||||
- CreateNamespace=true
|
||||
---
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Application
|
||||
metadata:
|
||||
name: kyverno-policies
|
||||
namespace: argocd
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-wave: "0"
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
targetRevision: main
|
||||
path: k8s/bootstrap/kyverno
|
||||
destination:
|
||||
server: https://kubernetes.default.svc
|
||||
namespace: kyverno
|
||||
syncPolicy:
|
||||
automated:
|
||||
prune: true
|
||||
selfHeal: true
|
||||
|
||||
@@ -1,33 +0,0 @@
|
||||
# ArgoCD Image Updater - auto-updates Application images from registry
|
||||
# Watches forgejo.riotpiao.com for new image tags and updates Applications
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Application
|
||||
metadata:
|
||||
name: argocd-image-updater
|
||||
namespace: argocd
|
||||
finalizers:
|
||||
- resources-finalizer.argocd.argoproj.io
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-wave: "1"
|
||||
spec:
|
||||
project: homelab
|
||||
revisionHistoryLimit: 3
|
||||
sources:
|
||||
- repoURL: https://argoproj.github.io/argo-helm
|
||||
chart: argocd-image-updater
|
||||
targetRevision: "0.11.2"
|
||||
helm:
|
||||
valueFiles:
|
||||
- $values/k8s/infra/argocd-image-updater/values.yaml
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
targetRevision: main
|
||||
ref: values
|
||||
destination:
|
||||
server: https://kubernetes.default.svc
|
||||
namespace: argocd
|
||||
syncPolicy:
|
||||
automated:
|
||||
prune: true
|
||||
selfHeal: true
|
||||
syncOptions:
|
||||
- CreateNamespace=false
|
||||
@@ -1,32 +0,0 @@
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Application
|
||||
metadata:
|
||||
name: secret-rotation
|
||||
namespace: argocd
|
||||
labels:
|
||||
app.kubernetes.io/name: secret-rotation
|
||||
spec:
|
||||
project: homelab
|
||||
|
||||
sources:
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
path: k8s/apps/secret-rotation-controller
|
||||
targetRevision: main
|
||||
|
||||
destination:
|
||||
server: https://kubernetes.default.svc
|
||||
namespace: secret-rotation
|
||||
|
||||
syncPolicy:
|
||||
automated:
|
||||
prune: true
|
||||
selfHeal: true
|
||||
syncOptions:
|
||||
- CreateNamespace=true
|
||||
- RespectIgnoreDifferences=true
|
||||
retry:
|
||||
limit: 5
|
||||
backoff:
|
||||
duration: 5s
|
||||
factor: 2
|
||||
maxDuration: 3m
|
||||
@@ -17,7 +17,7 @@ spec:
|
||||
helm:
|
||||
valueFiles:
|
||||
- $values/k8s/infra/minio/minio-operator-values.yaml
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
ref: values
|
||||
destination:
|
||||
@@ -41,7 +41,7 @@ metadata:
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/infra/minio
|
||||
destination:
|
||||
@@ -66,7 +66,7 @@ metadata:
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/infra/longhorn
|
||||
destination:
|
||||
@@ -102,7 +102,7 @@ spec:
|
||||
skipCrds: true
|
||||
valueFiles:
|
||||
- $values/k8s/infra/monitoring/prometheus-values.yaml
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
ref: values
|
||||
destination:
|
||||
@@ -152,7 +152,7 @@ metadata:
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/infra/monitoring/crds
|
||||
destination:
|
||||
@@ -183,7 +183,7 @@ metadata:
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/infra/monitoring
|
||||
destination:
|
||||
@@ -213,7 +213,7 @@ spec:
|
||||
helm:
|
||||
valueFiles:
|
||||
- $values/k8s/infra/monitoring/blackbox-exporter-values.yaml
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
ref: values
|
||||
destination:
|
||||
@@ -223,29 +223,3 @@ spec:
|
||||
automated:
|
||||
prune: true
|
||||
selfHeal: true
|
||||
---
|
||||
# Distributed tracing: Tempo + OpenTelemetry Collector.
|
||||
# Receives traces from instrumented services, stores in local volume (72h retention).
|
||||
# Grafana datasource auto-configured, service graph + latency dashboards included.
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Application
|
||||
metadata:
|
||||
name: tracing
|
||||
namespace: argocd
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-wave: "1"
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
targetRevision: main
|
||||
path: k8s/infra/tracing
|
||||
destination:
|
||||
server: https://kubernetes.default.svc
|
||||
namespace: tracing
|
||||
syncPolicy:
|
||||
automated:
|
||||
prune: true
|
||||
selfHeal: true
|
||||
syncOptions:
|
||||
- CreateNamespace=true
|
||||
|
||||
@@ -19,7 +19,7 @@ spec:
|
||||
helm:
|
||||
valueFiles:
|
||||
- $values/k8s/infra/logging/loki-values.yaml
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
ref: values
|
||||
destination:
|
||||
@@ -53,7 +53,7 @@ spec:
|
||||
helm:
|
||||
valueFiles:
|
||||
- $values/k8s/infra/logging/grafana-values.yaml
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
ref: values
|
||||
destination:
|
||||
@@ -87,7 +87,7 @@ spec:
|
||||
helm:
|
||||
valueFiles:
|
||||
- $values/k8s/infra/logging/promtail-values.yaml
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
ref: values
|
||||
destination:
|
||||
|
||||
@@ -17,7 +17,7 @@ spec:
|
||||
helm:
|
||||
valueFiles:
|
||||
- $values/k8s/infra/iam/vault-values.yaml
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
ref: values
|
||||
destination:
|
||||
@@ -46,7 +46,7 @@ spec:
|
||||
helm:
|
||||
valueFiles:
|
||||
- $values/k8s/infra/iam/authentik-values.yaml
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
ref: values
|
||||
destination:
|
||||
@@ -68,7 +68,7 @@ metadata:
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/infra/iam
|
||||
destination:
|
||||
@@ -109,7 +109,7 @@ spec:
|
||||
helm:
|
||||
valueFiles:
|
||||
- $values/k8s/bootstrap/phase3-forgejo/forgejo-values.yaml
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
ref: values
|
||||
destination:
|
||||
@@ -152,17 +152,10 @@ metadata:
|
||||
namespace: argocd
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-wave: "3"
|
||||
argocd-image-updater.argoproj.io/image-list: runner=forgejo.riotpiao.com/rock/forgejo-runner-golang
|
||||
argocd-image-updater.argoproj.io/runner.update-strategy: newest-build
|
||||
argocd-image-updater.argoproj.io/runner.allow-tags: regexp:^[0-9a-f]{7}$|^latest$|^v[0-9]+$
|
||||
argocd-image-updater.argoproj.io/runner.helm.image-name: runner.image.repository
|
||||
argocd-image-updater.argoproj.io/runner.helm.image-tag: runner.image.tag
|
||||
argocd-image-updater.argoproj.io/write-back-method: git
|
||||
argocd-image-updater.argoproj.io/git-branch: main
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/infra/forgejo-runner
|
||||
destination:
|
||||
@@ -180,17 +173,10 @@ metadata:
|
||||
namespace: argocd
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-wave: "3"
|
||||
argocd-image-updater.argoproj.io/image-list: runner=forgejo.riotpiao.com/rock/forgejo-runner-node
|
||||
argocd-image-updater.argoproj.io/runner.update-strategy: newest-build
|
||||
argocd-image-updater.argoproj.io/runner.allow-tags: regexp:^[0-9a-f]{7}$|^latest$|^v[0-9]+$
|
||||
argocd-image-updater.argoproj.io/runner.helm.image-name: runner.image.repository
|
||||
argocd-image-updater.argoproj.io/runner.helm.image-tag: runner.image.tag
|
||||
argocd-image-updater.argoproj.io/write-back-method: git
|
||||
argocd-image-updater.argoproj.io/git-branch: main
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/infra/forgejo-runner
|
||||
helm:
|
||||
@@ -211,17 +197,10 @@ metadata:
|
||||
namespace: argocd
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-wave: "3"
|
||||
argocd-image-updater.argoproj.io/image-list: runner=forgejo.riotpiao.com/rock/forgejo-runner-rust
|
||||
argocd-image-updater.argoproj.io/runner.update-strategy: newest-build
|
||||
argocd-image-updater.argoproj.io/runner.allow-tags: regexp:^[0-9a-f]{7}$|^latest$|^v[0-9]+$
|
||||
argocd-image-updater.argoproj.io/runner.helm.image-name: runner.image.repository
|
||||
argocd-image-updater.argoproj.io/runner.helm.image-tag: runner.image.tag
|
||||
argocd-image-updater.argoproj.io/write-back-method: git
|
||||
argocd-image-updater.argoproj.io/git-branch: main
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/infra/forgejo-runner
|
||||
helm:
|
||||
|
||||
@@ -14,7 +14,7 @@ metadata:
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/infra/databases
|
||||
destination:
|
||||
|
||||
@@ -1,20 +0,0 @@
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Application
|
||||
metadata:
|
||||
name: memory-queues
|
||||
namespace: argocd
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-wave: "7"
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
targetRevision: main
|
||||
path: k8s/apps/messaging/memory-queues
|
||||
destination:
|
||||
server: https://kubernetes.default.svc
|
||||
namespace: sqs
|
||||
syncPolicy:
|
||||
automated:
|
||||
prune: true
|
||||
selfHeal: true
|
||||
@@ -1,6 +1,6 @@
|
||||
# Wave 5 — Kafka (Strimzi operator + cluster CR), Redis infrastructure.
|
||||
# Strimzi/Redis are public Helm charts; kafka-cluster is a local chart.
|
||||
# queue-crd and management-service are managed by kmsvc-root (kmsvc-manage.git).
|
||||
# Wave 5 — Kafka (Strimzi operator + cluster CR), Redis, and the SQS-like
|
||||
# queue services. Strimzi/Redis are public Helm charts; kafka-cluster/queue-crd/
|
||||
# management-service are local charts (rendered from their own Chart.yaml).
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Application
|
||||
metadata:
|
||||
@@ -68,7 +68,7 @@ metadata:
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/apps/messaging/kafka-cluster
|
||||
destination:
|
||||
@@ -78,5 +78,45 @@ spec:
|
||||
automated:
|
||||
prune: true
|
||||
selfHeal: true
|
||||
# queue-crd and management-service moved to kmsvc-manage.git repo
|
||||
# Managed by kmsvc-root Application
|
||||
---
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Application
|
||||
metadata:
|
||||
name: queue-crd
|
||||
namespace: argocd
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-wave: "6"
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/apps/messaging/queue-crd
|
||||
destination:
|
||||
server: https://kubernetes.default.svc
|
||||
namespace: sqs
|
||||
syncPolicy:
|
||||
automated:
|
||||
prune: true
|
||||
selfHeal: true
|
||||
---
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Application
|
||||
metadata:
|
||||
name: management-service
|
||||
namespace: argocd
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-wave: "7"
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/apps/messaging/management-service
|
||||
destination:
|
||||
server: https://kubernetes.default.svc
|
||||
namespace: sqs
|
||||
syncPolicy:
|
||||
automated:
|
||||
prune: true
|
||||
selfHeal: true
|
||||
|
||||
@@ -9,12 +9,8 @@
|
||||
# `POST /v1/chat/completions` covers every model. See
|
||||
# docs/adr/ADR-0001-retire-kong-for-go-gateway.md in the frontend repo.
|
||||
#
|
||||
# UPDATED 2026-08-22: Tracks main branch of homelab-frontend (auto-syncs on each push).
|
||||
# Image built on every main commit with tag <commit-sha>.
|
||||
# ArgoCD auto-pulls the latest image (live reconciliation ~3min).
|
||||
#
|
||||
# Two sources:
|
||||
# 1. rock/homelab-frontend on the in-cluster Forgejo (prod branch) — the gateway's own
|
||||
# 1. rock/homelab-frontend on the in-cluster Forgejo — the gateway's own
|
||||
# kustomization (Deployment, Service, ConfigMap, RBAC, NetworkPolicy). It
|
||||
# sets `namespace: api` itself, so no transformer is needed here. The
|
||||
# Forgejo host must stay listed in the `homelab` AppProject sourceRepos or
|
||||
@@ -36,11 +32,6 @@ metadata:
|
||||
app.kubernetes.io/component: gateway
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-wave: "7"
|
||||
# ArgoCD Image Updater - auto-update on new image push
|
||||
argocd-image-updater.argoproj.io/image-list: gw=forgejo.riotpiao.com/rock/api-gateway
|
||||
argocd-image-updater.argoproj.io/gw.update-strategy: digest
|
||||
argocd-image-updater.argoproj.io/gw.allow-tags: regexp:^latest$
|
||||
argocd-image-updater.argoproj.io/write-back-method: argocd
|
||||
spec:
|
||||
project: homelab
|
||||
revisionHistoryLimit: 3
|
||||
@@ -48,7 +39,7 @@ spec:
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/homelab-frontend.git
|
||||
targetRevision: main
|
||||
path: k8s
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/apps/api
|
||||
destination:
|
||||
|
||||
@@ -20,7 +20,7 @@ spec:
|
||||
project: homelab
|
||||
revisionHistoryLimit: 3
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/apps/llm-serving
|
||||
destination:
|
||||
|
||||
@@ -19,10 +19,10 @@ spec:
|
||||
helm:
|
||||
valueFiles:
|
||||
- $values/k8s/apps/temporal/temporal-values.yaml
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
ref: values
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/apps/temporal
|
||||
destination:
|
||||
@@ -51,7 +51,7 @@ spec:
|
||||
helm:
|
||||
valueFiles:
|
||||
- $values/k8s/apps/portainer/portainer-values.yaml
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
ref: values
|
||||
destination:
|
||||
@@ -74,7 +74,7 @@ metadata:
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/apps/cloudflared
|
||||
destination:
|
||||
@@ -97,7 +97,7 @@ metadata:
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/apps/agent-pod
|
||||
destination:
|
||||
@@ -130,7 +130,7 @@ metadata:
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/apps/sms
|
||||
destination:
|
||||
@@ -141,67 +141,6 @@ spec:
|
||||
prune: true
|
||||
selfHeal: true
|
||||
---
|
||||
# Document management. Raw manifests (no Helm): postgres is the dedicated
|
||||
# paperless-db CNPG cluster in k8s/infra/databases (wave 2), redis is
|
||||
# in-cluster only (no PVC), media lives on the cp-3 USB HDD (see
|
||||
# k8s/infra/longhorn/longhorn-paperless-storageclass.yaml). OIDC via
|
||||
# Authentik provisioned by k8s/infra/iam's PostSync job; MinIO backup bucket
|
||||
# creds provisioned by k8s/infra/minio's PostSync job.
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Application
|
||||
metadata:
|
||||
name: paperless
|
||||
namespace: argocd
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-wave: "8"
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
targetRevision: main
|
||||
path: k8s/apps/paperless
|
||||
destination:
|
||||
server: https://kubernetes.default.svc
|
||||
namespace: paperless
|
||||
syncPolicy:
|
||||
automated:
|
||||
prune: true
|
||||
selfHeal: true
|
||||
syncOptions:
|
||||
- CreateNamespace=true
|
||||
---
|
||||
# Photo/video backup. Self-contained (unlike paperless, its CNPG Postgres
|
||||
# lives here too, not in k8s/infra/databases) - CreateNamespace=true creates
|
||||
# the namespace before any manifest in this Application applies, including
|
||||
# the Cluster CR, so no separate wave-2 pre-creation step is needed. Postgres
|
||||
# is pg18 (not this repo's usual 16.2) because CNPG's official pgvector
|
||||
# extension image only publishes pg18 builds - see k8s/apps/immich/db.yaml.
|
||||
# media PVC shares the cp-3 HDD 2TB/2TB with paperless-media. OIDC via
|
||||
# Authentik provisioned by k8s/infra/iam's PostSync job (immich entry in
|
||||
# SERVICES + immich_role scope mapping for admin-via-claim).
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Application
|
||||
metadata:
|
||||
name: immich
|
||||
namespace: argocd
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-wave: "8"
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
targetRevision: main
|
||||
path: k8s/apps/immich
|
||||
destination:
|
||||
server: https://kubernetes.default.svc
|
||||
namespace: immich
|
||||
syncPolicy:
|
||||
automated:
|
||||
prune: true
|
||||
selfHeal: true
|
||||
syncOptions:
|
||||
- CreateNamespace=true
|
||||
---
|
||||
# Consolidated: homarr + homarr-patches → homarr
|
||||
# Helm chart + values + PostSync hook patch (fix-probes-job.yaml)
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
@@ -220,10 +159,10 @@ spec:
|
||||
helm:
|
||||
valueFiles:
|
||||
- $values/k8s/apps/homarr/homarr-values.yaml
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
ref: values
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/apps/homarr # PostSync hook: fix-probes-job.yaml
|
||||
destination:
|
||||
@@ -235,60 +174,3 @@ spec:
|
||||
selfHeal: true
|
||||
syncOptions:
|
||||
- CreateNamespace=true
|
||||
---
|
||||
# Portfolio site at riotpiao.com - static Next.js site from rock/riotpiao.com repo.
|
||||
# Points directly to infra/portfolio/base (bypassing repo's own argocd-apps.yaml
|
||||
# which has wrong URLs). Image built by Forgejo Actions on rock/portfolio repo.
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Application
|
||||
metadata:
|
||||
name: portfolio
|
||||
namespace: argocd
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-wave: "8"
|
||||
# ArgoCD Image Updater - auto-update on new image push
|
||||
argocd-image-updater.argoproj.io/image-list: app=forgejo.riotpiao.com/rock/portfolio
|
||||
argocd-image-updater.argoproj.io/app.update-strategy: digest
|
||||
argocd-image-updater.argoproj.io/app.allow-tags: regexp:^latest$
|
||||
argocd-image-updater.argoproj.io/write-back-method: argocd
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/riotpiao.com.git
|
||||
targetRevision: main
|
||||
path: infra/portfolio/base
|
||||
destination:
|
||||
server: https://kubernetes.default.svc
|
||||
namespace: portfolio
|
||||
syncPolicy:
|
||||
automated:
|
||||
prune: true
|
||||
selfHeal: true
|
||||
syncOptions:
|
||||
- CreateNamespace=true
|
||||
---
|
||||
# Wave 9 - per-service scoped RBAC (Role/RoleBinding), deliberately last so
|
||||
# every target namespace above already exists. Inert until kube-apiserver
|
||||
# gets --oidc-groups-claim=groups wired up (separate, not-yet-applied
|
||||
# terraform/talosctl change) - these grant nothing until then.
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Application
|
||||
metadata:
|
||||
name: rbac
|
||||
namespace: argocd
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-wave: "9"
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
targetRevision: main
|
||||
path: k8s/infra/rbac
|
||||
destination:
|
||||
server: https://kubernetes.default.svc
|
||||
# No namespace: cluster-scoped resources (ClusterRoleBinding, etc.)
|
||||
# Namespace is set per-resource in kustomization
|
||||
syncPolicy:
|
||||
automated:
|
||||
prune: true
|
||||
selfHeal: true
|
||||
|
||||
@@ -1,28 +0,0 @@
|
||||
# kmsvc-manage bootstrap — manages itself and its supporting services
|
||||
# (Strimzi/Kafka, Redis, queue-operator, message-plane server) from
|
||||
# the kmsvc-manage repo's own k8s/argocd/ structure on the main branch.
|
||||
# Image built on every main commit, auto-deployed to sqs namespace.
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Application
|
||||
metadata:
|
||||
name: kmsvc-root
|
||||
namespace: argocd
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-wave: "6"
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/kmsvc-manage.git
|
||||
targetRevision: main
|
||||
path: k8s/argocd/apps
|
||||
directory:
|
||||
recurse: false
|
||||
destination:
|
||||
server: https://kubernetes.default.svc
|
||||
namespace: sqs
|
||||
syncPolicy:
|
||||
automated:
|
||||
prune: true
|
||||
selfHeal: true
|
||||
syncOptions:
|
||||
- CreateNamespace=true
|
||||
@@ -1,46 +0,0 @@
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Application
|
||||
metadata:
|
||||
name: poimen
|
||||
namespace: argocd
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-wave: "7"
|
||||
# Image Updater: auto-update on new image push (SHA tag filter)
|
||||
argocd-image-updater.argoproj.io/image-list: |
|
||||
memory=forgejo.riotpiao.com/rock/poimen-memory
|
||||
workflows=forgejo.riotpiao.com/rock/poimen-workflows
|
||||
frontend=forgejo.riotpiao.com/rock/poimen-frontend
|
||||
argocd-image-updater.argoproj.io/memory.update-strategy: digest
|
||||
argocd-image-updater.argoproj.io/memory.allow-tags: regexp:^latest$
|
||||
argocd-image-updater.argoproj.io/workflows.update-strategy: digest
|
||||
argocd-image-updater.argoproj.io/workflows.allow-tags: regexp:^latest$
|
||||
argocd-image-updater.argoproj.io/frontend.update-strategy: digest
|
||||
argocd-image-updater.argoproj.io/frontend.allow-tags: regexp:^latest$
|
||||
argocd-image-updater.argoproj.io/write-back-method: argocd
|
||||
spec:
|
||||
project: homelab
|
||||
sources:
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/poimen-memory.git
|
||||
targetRevision: main
|
||||
path: k8s/argocd
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/poimen-workflows.git
|
||||
targetRevision: main
|
||||
path: k8s/argocd
|
||||
- repoURL: https://forgejo.riotpiao.com/rock/poimen-frontend.git
|
||||
targetRevision: main
|
||||
path: k8s/argocd
|
||||
destination:
|
||||
server: https://kubernetes.default.svc
|
||||
namespace: poimen
|
||||
syncPolicy:
|
||||
automated:
|
||||
prune: true
|
||||
selfHeal: true
|
||||
syncOptions:
|
||||
- CreateNamespace=true
|
||||
retry:
|
||||
limit: 5
|
||||
backoff:
|
||||
duration: 5s
|
||||
factor: 2
|
||||
maxDuration: 3m
|
||||
@@ -12,18 +12,9 @@ spec:
|
||||
description: Homelab GitOps — single-repo, in-cluster destinations only
|
||||
sourceRepos:
|
||||
- https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
# Poimen services (GitHub)
|
||||
- https://github.com/Riotpiaole/Poimen-memory.git
|
||||
- https://github.com/Riotpiaole/Poimen-workflows.git
|
||||
- https://github.com/Riotpiaole/poimen*.git
|
||||
# In-cluster Forgejo repos — explicit allowlist (no wildcard)
|
||||
- https://forgejo.riotpiao.com/rock/homelab.git
|
||||
- https://forgejo.riotpiao.com/rock/homelab-frontend.git
|
||||
- https://forgejo.riotpiao.com/rock/kmsvc-manage.git
|
||||
- https://forgejo.riotpiao.com/rock/poimen.git
|
||||
- https://forgejo.riotpiao.com/rock/poimen-memory.git
|
||||
- https://forgejo.riotpiao.com/rock/poimen-workflows.git
|
||||
- https://forgejo.riotpiao.com/rock/riotpiao.com.git
|
||||
# In-cluster Forgejo wildcard — all rock/* repos can be onboarded without
|
||||
# touching this AppProject. Enabled by Stage 1 (A1).
|
||||
- https://forgejo.riotpiao.com/rock/*
|
||||
# Public Helm chart repos referenced by k8s/argocd/apps/* and bootstrap/*
|
||||
- https://cloudnative-pg.github.io/charts
|
||||
- https://dl.gitea.com/charts/
|
||||
@@ -42,8 +33,6 @@ spec:
|
||||
- https://charts.jetstack.io
|
||||
- https://kubernetes.github.io/ingress-nginx
|
||||
- https://stakater.github.io/stakater-charts
|
||||
# ArgoCD ecosystem charts
|
||||
- https://argoproj.github.io/argo-helm
|
||||
destinations:
|
||||
- server: https://kubernetes.default.svc
|
||||
namespace: "*"
|
||||
|
||||
@@ -12,7 +12,7 @@ metadata:
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
|
||||
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
|
||||
targetRevision: main
|
||||
path: k8s/argocd/apps
|
||||
directory:
|
||||
@@ -26,4 +26,3 @@ spec:
|
||||
selfHeal: true
|
||||
syncOptions:
|
||||
- CreateNamespace=true
|
||||
- ServerSideApply=true
|
||||
|
||||
@@ -1,23 +0,0 @@
|
||||
apiVersion: ENC[AES256_GCM,data:ECE=,iv:bISz4HovH++X7DW1Qj8Cw0L6+EvPB0+68hh7tfyW5C0=,tag:w+SjD7MsfeIuSf62n+Zl7Q==,type:str]
|
||||
kind: ENC[AES256_GCM,data:AxXaV3Mb,iv:xfJb354Rrjz3zLctW2i6hl40yh9EsfIqKxDvmn8jqnU=,tag:1xQ/RR3KhPLKCTWCYjebJg==,type:str]
|
||||
metadata:
|
||||
name: ENC[AES256_GCM,data:gdQ4OoMYKUPEbsUeA5OI4i1xnh8=,iv:IA2kaCyqQmkYpLolvcTF4aleh+yd/ImXJMhRMvpGCgo=,tag:8qyYYw4EhBKKPzEmpepAeQ==,type:str]
|
||||
namespace: ENC[AES256_GCM,data:/mOSyXWmRg==,iv:cpeHUTlMqlJzJttGtuR3DoiMtvVqFmDS0/5Tl7K7c2M=,tag:q4GxWgWn2wQJxJHqnq4WQA==,type:str]
|
||||
type: ENC[AES256_GCM,data:92OcrsuV,iv:Z1XBy6iZ6unGrK4/SSdDa58pbPL3022gU/f1AOM4uvc=,tag:cd4wWRq/ohm6BJ/Eo6HOIw==,type:str]
|
||||
stringData:
|
||||
REGISTRY_PAT: ENC[AES256_GCM,data:zDKruOXhzsIcFZTTp4r4r9rwHih9uf1cCp06KZ6eIXSIZDvR6aQY4A==,iv:5Vch5Z7xn1KkRxrgOs2p/M58nd6WhtcuUhoZDCi8LSY=,tag:14SAZIWGOxJMqLBtOorjYQ==,type:str]
|
||||
sops:
|
||||
age:
|
||||
- enc: |
|
||||
-----BEGIN AGE ENCRYPTED FILE-----
|
||||
YWdlLWVuY3J5cHRpb24ub3JnL3YxCi0+IFgyNTUxOSBIaG5QVUYyNjhvaENPcUFy
|
||||
K3RSSVM2N2hTeFhNK0M5YWZmODhzZU1HaHlFCmNGZWdFcFEvanVwcXpGSmVvRHVx
|
||||
aG45c2lBY1RuSTYwbDZTVk1QeHNWOTQKLS0tIDg4TnVMNjNtaU1VQk5zQjUvU2hM
|
||||
R245ZVdqc0ZWWVhJb3dOZ3lpU3JjUlUKawSg09ZPq8FKx5tvOVZZ+K4yh7eTQsUp
|
||||
be8mWUpS0+eEmNqh35BwU3HrETMQFA6a1kjVp30JOMtqa5rbYlzF7w==
|
||||
-----END AGE ENCRYPTED FILE-----
|
||||
recipient: age1e5fq3hwxy78psus2nfvmtmua36g0u3suk78ephw6246l974d2utsvn0hla
|
||||
lastmodified: "2026-08-23T23:14:53Z"
|
||||
mac: ENC[AES256_GCM,data:yzi6woUaUCi6w9pG/eKnU7k/VfZgoXg/tW8p9joG8p7iabLWMdlXlx23m/CItw/NE0zeGWA5iZPPFZXOit2vN36VzK3kQNoFW6QbhvYLZi+78C/RWIdcpZkxB/tPIRG0vq8Q1SA+rwGjG/0xeAFh+R7k+YBTMd2XWBH3P7EI5T4=,iv:kf6MDaAVrtDvPIEjHMMLxXDSDRC3I1GpsPeJFUYppiw=,tag:NukVbGLa9EoMYmRsa4nBtA==,type:str]
|
||||
unencrypted_suffix: _unencrypted
|
||||
version: 3.13.2
|
||||
@@ -1,10 +1,8 @@
|
||||
apiVersion: kustomize.config.k8s.io/v1beta1
|
||||
kind: Kustomization
|
||||
|
||||
# Disable hash suffix for all generated secrets (stable names)
|
||||
generatorOptions:
|
||||
disableNameSuffixHash: true
|
||||
|
||||
# SOPS-encrypted secrets via ksops generator
|
||||
# All homelab SOPS-encrypted Secrets, decrypted in-line via the ksops generator.
|
||||
# Each *.enc.yaml carries its own metadata.namespace, so no namespace transformer
|
||||
# here (that would rewrite every Secret into one namespace). Renders exactly the
|
||||
# Secret objects — replaces the old argocd-cmp-cm SOPS plugin.
|
||||
generators:
|
||||
- secret-generator.yaml
|
||||
|
||||
@@ -1,25 +0,0 @@
|
||||
apiVersion: ENC[AES256_GCM,data:wGA=,iv:Z2Gfzq3aJ9j4fYaeLQolgLb/XELTHrKX9at3vUsMLIw=,tag:yZLZFnltJln5jFoVsyyZ2A==,type:str]
|
||||
kind: ENC[AES256_GCM,data:bJkV4pNj,iv:0ZT7l0kw9qSoiEMZioRw1aBzzlxBkoXR+hOxoPI7zPU=,tag:EQczPq508Qw1vi/oLCeQpw==,type:str]
|
||||
metadata:
|
||||
name: ENC[AES256_GCM,data:sZwTVM41PPaidMwFNRo5RvU=,iv:BHfuwIHng7rkeLK3a69t8cI9QeSfD/3FEXqxby+gBxM=,tag:3zQx8APnE2ZpBKf/ZYCxOA==,type:str]
|
||||
namespace: ENC[AES256_GCM,data:i7lpoEaZ1oXS,iv:jUYyDPYhf1TV51he/S5MlKPD19Vmz6wfE98y8qFEg3U=,tag:zgHDVlplmW/XPUAKdncfYQ==,type:str]
|
||||
type: ENC[AES256_GCM,data:2khs1uIg,iv:ET6HcBGyv33fGFlFAl3dkJQB83naHGeaYuWvW1IFhvw=,tag:qIfMeaSwv//7cNWvy8O5dg==,type:str]
|
||||
stringData:
|
||||
PAPERLESS_SECRET_KEY: ENC[AES256_GCM,data:kj9DrQYL3cQGJz87FHYlFKy6Muu84Oy/6wnsWAh0w/3MmcTAsCzQvUtKLKKf1U61JTQ=,iv:UffVv78vMHDEWHRFdSKZ/6qyrVD02Nlk0CNxwWd1jTo=,tag:LNMKnFWw7QCjfcKzI1s8ig==,type:str]
|
||||
PAPERLESS_ADMIN_USER: ENC[AES256_GCM,data:a3bx7z0=,iv:nrU6VXZE74PNXM+Lgg4K/3mhusaQhA2GAgl0GzqAQt8=,tag:x8nswTfkFHR6Ou+JSI9P3Q==,type:str]
|
||||
PAPERLESS_ADMIN_PASSWORD: ENC[AES256_GCM,data:+SWyAq7D1mNt1TlFOlo2Z0biK8kEn097,iv:an2ZZuXNSig+7mJP7ogHFc5+QvUb+z9pzqcV4xvFLbA=,tag:OaRiVZsnajEjAlc9mDWo9w==,type:str]
|
||||
sops:
|
||||
age:
|
||||
- enc: |
|
||||
-----BEGIN AGE ENCRYPTED FILE-----
|
||||
YWdlLWVuY3J5cHRpb24ub3JnL3YxCi0+IFgyNTUxOSA1Q2V6bVBUYmVRR1N5SCtF
|
||||
eXdTV2dmWjhtMS9lRFEzS0wzbkd6Y3JLMVRVCncydFBjdmJkRCtVUXphR2w0SDlJ
|
||||
cng2bi9MWlJzTEN2amJrYjRJN2VFcEEKLS0tIEY5cmw2RGhnbzUxZW9FaFJjQmVN
|
||||
WWcvNlNiYWdwbnNSR1Q4alpDZmFqTFkKh9TOw8ERP9fpx2pKi/Q0b7+OkEv0UC7o
|
||||
aAIK4Tzvi5dp6y9IWcu9l6PjDLeYWOJ5wr7QABaFNOz82hngxUleDA==
|
||||
-----END AGE ENCRYPTED FILE-----
|
||||
recipient: age1e5fq3hwxy78psus2nfvmtmua36g0u3suk78ephw6246l974d2utsvn0hla
|
||||
lastmodified: "2026-08-25T16:04:48Z"
|
||||
mac: ENC[AES256_GCM,data:OXPQXUyn/SDftKH5nRzhqEtnaOb4Gp7etmGojqV4Z01kHnABZ42PPIyiBb8z997eLNZsK6/1bi7nUNyVp4fIuXG4F52aEdMewofOocCHokRnNmz7jzhooK1gScJb2u0eHG3FL5iLONMaGgVpk7BLYO3e0xiDytWGe8BxcuDukPg=,iv:kh0jUwFvO6AUICy2Us1E7YOTEcp3L+ptrGdDWBSpWyc=,tag:qUPFmZ/rpljln37f/NRjjw==,type:str]
|
||||
unencrypted_suffix: _unencrypted
|
||||
version: 3.13.2
|
||||
@@ -1,2 +0,0 @@
|
||||
FORGEJO_TOKEN=273fdcffabbcbb5a191e8289c73d106063acefc6
|
||||
LLM_API_TOKEN=s3VksXyw2z3sGbegnwjMDFnJ6CtNRd1a5CcnE5A4ET77toCcykNdunk6Oa2J
|
||||
@@ -1,23 +0,0 @@
|
||||
apiVersion: ENC[AES256_GCM,data:bnY=,iv:Fuc3aqncHQ+L16o7eLarPbOECD3o8Mk5c2r9pQBpy70=,tag:JPfEnbNb3wZXPdXafnJDqw==,type:str]
|
||||
kind: ENC[AES256_GCM,data:WFlmi4Yg,iv:Zq/KQbgNcBVoo8ZsQ2H79ygyc8Dtkgxh4fCpEExfwSg=,tag:cWHP6Y5V+nZP2tFMJrOB8A==,type:str]
|
||||
metadata:
|
||||
name: ENC[AES256_GCM,data:iCXbhwvg3Zq6YL/4j0wAy7Y=,iv:8s+d/8lDVEL7bGdIF+GOtAxapKnmx8JTjLSO04hXF5I=,tag:LjzX40DjK+uRZPCXsMlmwQ==,type:str]
|
||||
namespace: ENC[AES256_GCM,data:HRMdZdCbxORQ,iv:MvaIWoKWjJRA7/fce0KtXRkFH/7cn0OuIg2QwHEdQzM=,tag:NqztiGyfU3BaopWBKhx2eg==,type:str]
|
||||
type: ENC[AES256_GCM,data:myBW86Za,iv:3x9ys5UzVhAuX8gvZO67B1e+Orw4Aqasv/lHBgUV0b4=,tag:yl18P6SaxVLJidwmRtd4aQ==,type:str]
|
||||
stringData:
|
||||
FORGEJO_TOKEN: ENC[AES256_GCM,data:SUoBpNKOItyNGY01EhKNlPH0fyN4N7g6bfU2jsGogpCMhP7NuRipoA==,iv:h77RtmYZjXxiHYw1pHynQHuVX1+yJDGHwsXLf+DbUYA=,tag:GHzoDNHP1GGqlN5eT/N7TQ==,type:str]
|
||||
sops:
|
||||
age:
|
||||
- enc: |
|
||||
-----BEGIN AGE ENCRYPTED FILE-----
|
||||
YWdlLWVuY3J5cHRpb24ub3JnL3YxCi0+IFgyNTUxOSBzTVBsekl3TGgzQVRMUU9m
|
||||
cTduS2NoZW5uZFNNMG13cFY2cGVsTnlXaXhrCjJjbzhLdHZ4ZWpUV3J0cDQ0eVlM
|
||||
WDNxdzVoQ2ZzcGJSbTU3RVorcnczNVkKLS0tIFdtQTE4Umk2TDBzUmdKOXNkbjFi
|
||||
Vk5vK2VuUHVsb3FQL21vcGU1UW5CT1kKFM8vVjji3Cg9dvfTr4Hx7BJC8JH5ovef
|
||||
Dj6zkofhsNWPgP9T+mnQakj+C0RKmHOMJqfWP7vwCBkZoNosIJVlMw==
|
||||
-----END AGE ENCRYPTED FILE-----
|
||||
recipient: age1e5fq3hwxy78psus2nfvmtmua36g0u3suk78ephw6246l974d2utsvn0hla
|
||||
lastmodified: "2026-09-01T05:32:42Z"
|
||||
mac: ENC[AES256_GCM,data:it24T9y9ixXo2aiL37k93vKFR+SRjjuI9DQdv0sWYtTogWnc7+uXBY4Zip/ouWyCse1muKKAGuek5c0XVrvSw4an9VkaXFczeunaZb6MOyVbVOkmJr+5xZFpZGjYcSkrhaWcVheedZ3iIFU5UWI7BBn/qQCf+HJ483cJqwtrV34=,iv:WprlWJdsMBNjqaA0O3ekfXMUpX5gC6OLYortQXYdTS4=,tag:rGqSSRTYTv2VR6AcRokO0A==,type:str]
|
||||
unencrypted_suffix: _unencrypted
|
||||
version: 3.13.2
|
||||
@@ -11,7 +11,6 @@ files:
|
||||
- agent-pod-ssh-key.enc.yaml
|
||||
- authentik-secrets.enc.yaml
|
||||
- cloudflare-secrets.enc.yaml
|
||||
- forgejo-registry-pat.enc.yaml
|
||||
- forgejo-registry-pull.enc.yaml
|
||||
- forgejo-runner-token.enc.yaml
|
||||
- forgejo-secrets.enc.yaml
|
||||
@@ -23,7 +22,5 @@ files:
|
||||
- homelab-ca-secrets.enc.yaml
|
||||
- loki-secrets.enc.yaml
|
||||
- minio-secrets.enc.yaml
|
||||
- paperless-secrets.enc.yaml
|
||||
- vault-secrets.enc.yaml
|
||||
- vault-unseal-keys.enc.yaml
|
||||
- portfolio-secrets.enc.yaml
|
||||
|
||||
@@ -5,9 +5,9 @@ metadata:
|
||||
namespace: ENC[AES256_GCM,data:wK6m,iv:KtA31Bo8aGE1HU8H9KWbMwt2NfywWfy77G/LaReaI1E=,tag:V9SSlPawRQw0x2utFNN1aw==,type:str]
|
||||
type: ENC[AES256_GCM,data:6UPZ1ZTR,iv:FxN1ebrlJ4IO3eDGEYSvktSpMgAebcl0DO1WHh5O0+0=,tag:2bZN4zdbHzx0oR+bTJPPJg==,type:str]
|
||||
stringData:
|
||||
key1: ENC[AES256_GCM,data:ke5EfsOsZHBJyAQoFfwhhWCQJQgnwcrBqL4CzPIfkx7bi9S7kWgWySdxcXQ=,iv:eYKxG8k+hizp2t2i/YMR2lQNJQFV+A21YyWnDc+kJ9w=,tag:Wa1Lx3EKvfpYfnLrBOWSIg==,type:str]
|
||||
key2: ENC[AES256_GCM,data:JxiXgLKvDe5oiHlwIL/Cj8txHc7fVQ5VzBcQMU/ro9TSOcTGSe7Z3omQzwY=,iv:x5arcY2dey+npMpUxjdUPV+t94LEaMOXq3iOar88eT8=,tag:E/Z5zEDy87B1IAWkdjdNEg==,type:str]
|
||||
key3: ENC[AES256_GCM,data:AsxyyBTHzt+SLru6ayi1UbgPPFmdIvowIDdcfc7N03Xt8V/KaJW3kbXExG8=,iv:lA+bERIBq4+bxx08ZZjP5tMHC9JQlGPW1EMCwEixSeg=,tag:xkuTAU3KTidFuhNMXkVFhA==,type:str]
|
||||
key1: ENC[AES256_GCM,data:dV2HOh7W1Pl0QDJaGvtEKpBppybQMiK5kxzyk05zkA352ZmJvE8Ppm5yWKk=,iv:grQ8v2o/LHpJnIZonjtgTHKcLUQIH8xl1FtWtEs4rEs=,tag:lANcccYT6Ouo/0IcoLw4uA==,type:str]
|
||||
key2: ENC[AES256_GCM,data:YSdpYL8h56PfUrvMFhBXhmBG2en0toLKEpYlJqwjAk/vI0jm7gq/TqGn034=,iv:GxthkNLhm3qHxjkZetiYl58Qa/K7G2ib6E+LWK16H8Q=,tag:q79UUoh6uCcMZJ0U18Ireg==,type:str]
|
||||
key3: ENC[AES256_GCM,data:sO7Lao9qkmIOdARL+6FBcLo+e4z8LzMEzNrwcy5hv9uRuxppwAXrP0MO1vw=,iv:wSWe9h1HovRhV5YCZrmvcB2VEfRYQ+gasGO3uh8IpmQ=,tag:Bc5yaROQDTz1AhVnFDm/BA==,type:str]
|
||||
sops:
|
||||
age:
|
||||
- enc: |
|
||||
@@ -19,7 +19,7 @@ sops:
|
||||
7RltNZF2SCxjIv5C2pqf3CgmRBaQGWgMybRRH5gdB87PLBKkPL3+HQ==
|
||||
-----END AGE ENCRYPTED FILE-----
|
||||
recipient: age1e5fq3hwxy78psus2nfvmtmua36g0u3suk78ephw6246l974d2utsvn0hla
|
||||
lastmodified: "2026-08-26T23:33:09Z"
|
||||
mac: ENC[AES256_GCM,data:kr8XuOqXyhfrC84yBot5aoVQZZhQd9xYMUEs0XDwkmCgtG0iobADYH5XWPh72xSdWEtwxkZJL7r0UEyEhU9zjlUA5qHupOSBBt1SgZ+zstMvaOqUnlNn//p/DIJBpsiT/qmx64NpTLAiz6lm0796MozIMr8PTX+ubGLHi9tnUiY=,iv:jwP0MLCT7nG6m+nZeqNip9q3BcScpcDXmInnY95DicY=,tag:Cgg2IUZARkaK1gF2qUIthg==,type:str]
|
||||
lastmodified: "2026-08-12T20:52:24Z"
|
||||
mac: ENC[AES256_GCM,data:kLjoBWmJ2bZ13EbuzgppahHKmCY/xtCi1R7xCEUbCP4FQfiVaC4qAbVftE4lTNlXa2mqCKsGPYZ/A3HV/4+a+it/pg0a+rj7E7DszhieWbZufHMJs+w4/Le8l8wFFE5aNW0wdvGTJZ0n4HBOdAkE1qN9ReaiaQgr0rk5ggOPG9g=,iv:MiDVUgcoKrP/gj4keqomo6OB12dmz9VoesTiHQTnBFM=,tag:cLK67Pym4BuvXLF1RUA8Gw==,type:str]
|
||||
unencrypted_suffix: _unencrypted
|
||||
version: 3.13.2
|
||||
|
||||
@@ -323,4 +323,3 @@ spec:
|
||||
name: homarr
|
||||
port:
|
||||
number: 7575
|
||||
|
||||
|
||||
@@ -1,6 +0,0 @@
|
||||
apiVersion: kustomize.config.k8s.io/v1beta1
|
||||
kind: Kustomization
|
||||
namespace: kyverno
|
||||
|
||||
resources:
|
||||
- policies.yaml
|
||||
@@ -1,42 +0,0 @@
|
||||
# Kyverno: Policy engine for Kubernetes image scanning, Pod security, and admission control
|
||||
# Scan all images, enforce baseline Pod Security Standard, prevent privilege escalation
|
||||
|
||||
replicaCount: 1
|
||||
|
||||
image:
|
||||
registry: ghcr.io
|
||||
repository: kyverno/kyverno
|
||||
tag: "v1.14.0"
|
||||
|
||||
config:
|
||||
# Webhook timeout for policy evaluation. Increase if scanning takes longer.
|
||||
webhookTimeoutSeconds: 30
|
||||
# Failure policy: fail-open (audit/log) vs fail-closed (reject on error)
|
||||
failurePolicy: fail
|
||||
# Resource limits for webhook
|
||||
webhookAnnotations:
|
||||
rules: "allow"
|
||||
|
||||
# Pod security via Kyverno instead of Pod Security Policies (deprecated)
|
||||
# Enforces baseline restrictions cluster-wide, with exceptions for privileged namespaces
|
||||
podSecurityContext:
|
||||
runAsNonRoot: true
|
||||
runAsUser: 1000
|
||||
|
||||
rbac:
|
||||
create: true
|
||||
|
||||
resources:
|
||||
requests:
|
||||
memory: "256Mi"
|
||||
cpu: "100m"
|
||||
limits:
|
||||
memory: "512Mi"
|
||||
cpu: "500m"
|
||||
|
||||
# Webhook configuration
|
||||
webhook:
|
||||
timeoutSeconds: 30
|
||||
# Failure policy: "Fail" (reject on error) or "Ignore" (audit-only)
|
||||
# Set to "Ignore" for initial testing, then change to "Fail"
|
||||
failurePolicy: ignore
|
||||
@@ -1,204 +0,0 @@
|
||||
# Kyverno ClusterPolicies: Image scanning, Pod security, and admission control
|
||||
---
|
||||
# Policy 1: Require non-root containers
|
||||
apiVersion: kyverno.io/v1
|
||||
kind: ClusterPolicy
|
||||
metadata:
|
||||
name: require-non-root
|
||||
namespace: kyverno
|
||||
spec:
|
||||
validationFailureAction: audit # audit first, then change to enforce
|
||||
rules:
|
||||
- name: check-runAsNonRoot
|
||||
match:
|
||||
any:
|
||||
- resources:
|
||||
kinds:
|
||||
- Pod
|
||||
selector:
|
||||
matchLabels:
|
||||
pod-security.kubernetes.io/enforce: "!privileged"
|
||||
validate:
|
||||
message: "Container must not run as root"
|
||||
pattern:
|
||||
spec:
|
||||
containers:
|
||||
- securityContext:
|
||||
runAsNonRoot: true
|
||||
---
|
||||
# Policy 2: Drop all Linux capabilities, add only required ones
|
||||
apiVersion: kyverno.io/v1
|
||||
kind: ClusterPolicy
|
||||
metadata:
|
||||
name: require-dropped-caps
|
||||
namespace: kyverno
|
||||
spec:
|
||||
validationFailureAction: audit
|
||||
rules:
|
||||
- name: drop-all-capabilities
|
||||
match:
|
||||
any:
|
||||
- resources:
|
||||
kinds:
|
||||
- Pod
|
||||
selector:
|
||||
matchLabels:
|
||||
pod-security.kubernetes.io/enforce: "!privileged"
|
||||
validate:
|
||||
message: "All Linux capabilities must be dropped"
|
||||
pattern:
|
||||
spec:
|
||||
containers:
|
||||
- securityContext:
|
||||
capabilities:
|
||||
drop:
|
||||
- ALL
|
||||
---
|
||||
# Policy 3: Require image tags (no 'latest')
|
||||
apiVersion: kyverno.io/v1
|
||||
kind: ClusterPolicy
|
||||
metadata:
|
||||
name: disallow-latest-tag
|
||||
namespace: kyverno
|
||||
spec:
|
||||
validationFailureAction: audit # Change to enforce after testing
|
||||
rules:
|
||||
- name: disallow-latest
|
||||
match:
|
||||
any:
|
||||
- resources:
|
||||
kinds:
|
||||
- Pod
|
||||
- Deployment
|
||||
- StatefulSet
|
||||
- DaemonSet
|
||||
- Job
|
||||
validate:
|
||||
message: "Image tag 'latest' is not allowed. Use explicit version tags."
|
||||
pattern:
|
||||
spec:
|
||||
=(template):
|
||||
spec:
|
||||
containers:
|
||||
- image: "!*:latest"
|
||||
=(initContainers):
|
||||
- image: "!*:latest"
|
||||
---
|
||||
# Policy 4: Restrict images to trusted registries
|
||||
apiVersion: kyverno.io/v1
|
||||
kind: ClusterPolicy
|
||||
metadata:
|
||||
name: restrict-registries
|
||||
namespace: kyverno
|
||||
spec:
|
||||
validationFailureAction: audit
|
||||
rules:
|
||||
- name: trusted-registries
|
||||
match:
|
||||
any:
|
||||
- resources:
|
||||
kinds:
|
||||
- Pod
|
||||
- Deployment
|
||||
- StatefulSet
|
||||
- DaemonSet
|
||||
- Job
|
||||
selector:
|
||||
matchLabels:
|
||||
pod-security.kubernetes.io/enforce: "!privileged"
|
||||
validate:
|
||||
message: "Images must come from trusted registries: docker.io, ghcr.io, quay.io, k8s.gcr.io, registry.k8s.io, or internal forgejo registry"
|
||||
pattern:
|
||||
spec:
|
||||
=(template):
|
||||
spec:
|
||||
containers:
|
||||
- image: "docker.io/* | ghcr.io/* | quay.io/* | k8s.gcr.io/* | registry.k8s.io/* | forgejo.riotpiao.com/* | *"
|
||||
---
|
||||
# Policy 5: Require read-only root filesystem (audit only, exceptions for apps that need writes)
|
||||
apiVersion: kyverno.io/v1
|
||||
kind: ClusterPolicy
|
||||
metadata:
|
||||
name: require-readonly-filesystem
|
||||
namespace: kyverno
|
||||
spec:
|
||||
validationFailureAction: audit
|
||||
rules:
|
||||
- name: check-readOnlyRootFilesystem
|
||||
match:
|
||||
any:
|
||||
- resources:
|
||||
kinds:
|
||||
- Pod
|
||||
selector:
|
||||
matchLabels:
|
||||
pod-security.kubernetes.io/enforce: "!privileged"
|
||||
validate:
|
||||
message: "Root filesystem should be read-only for defense-in-depth"
|
||||
pattern:
|
||||
spec:
|
||||
containers:
|
||||
- securityContext:
|
||||
readOnlyRootFilesystem: true
|
||||
---
|
||||
# Policy 6: Require resource requests and limits (prevent resource starvation)
|
||||
apiVersion: kyverno.io/v1
|
||||
kind: ClusterPolicy
|
||||
metadata:
|
||||
name: require-resource-limits
|
||||
namespace: kyverno
|
||||
spec:
|
||||
validationFailureAction: audit
|
||||
rules:
|
||||
- name: check-resources
|
||||
match:
|
||||
any:
|
||||
- resources:
|
||||
kinds:
|
||||
- Pod
|
||||
- Deployment
|
||||
- StatefulSet
|
||||
- DaemonSet
|
||||
excludeResources:
|
||||
namespaceSelector:
|
||||
matchLabels:
|
||||
kubernetes.io/metadata.name: "kyverno|kube-system|kube-node-lease"
|
||||
validate:
|
||||
message: "CPU and memory requests and limits are required"
|
||||
pattern:
|
||||
spec:
|
||||
=(template):
|
||||
spec:
|
||||
containers:
|
||||
- resources:
|
||||
requests:
|
||||
memory: "?*"
|
||||
cpu: "?*"
|
||||
limits:
|
||||
memory: "?*"
|
||||
cpu: "?*"
|
||||
---
|
||||
# Policy 7: Require securityContext on all containers
|
||||
apiVersion: kyverno.io/v1
|
||||
kind: ClusterPolicy
|
||||
metadata:
|
||||
name: require-security-context
|
||||
namespace: kyverno
|
||||
spec:
|
||||
validationFailureAction: audit
|
||||
rules:
|
||||
- name: check-securityContext
|
||||
match:
|
||||
any:
|
||||
- resources:
|
||||
kinds:
|
||||
- Pod
|
||||
selector:
|
||||
matchLabels:
|
||||
pod-security.kubernetes.io/enforce: "!privileged"
|
||||
validate:
|
||||
message: "securityContext must be defined"
|
||||
pattern:
|
||||
spec:
|
||||
containers:
|
||||
- securityContext: {}
|
||||
@@ -1,220 +0,0 @@
|
||||
# Forgejo OCI Registry Cleanup CronJob
|
||||
# Deletes old image tags, keeping only the latest N versions per repository.
|
||||
# Useful for retiring old builds when new versions are pushed.
|
||||
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: ConfigMap
|
||||
metadata:
|
||||
name: forgejo-registry-cleanup-script
|
||||
namespace: cicd
|
||||
data:
|
||||
cleanup.sh: |
|
||||
#!/bin/bash
|
||||
set -eo pipefail
|
||||
|
||||
# Configuration
|
||||
REGISTRY_HOST="${REGISTRY_HOST:-forgejo.riotpiao.com}"
|
||||
REGISTRY_URL="https://${REGISTRY_HOST}"
|
||||
KEEP_VERSIONS="${KEEP_VERSIONS:-3}" # Keep latest N versions per image
|
||||
DRY_RUN="${DRY_RUN:-false}"
|
||||
|
||||
# Load credentials from mounted secret
|
||||
REGISTRY_USER="${REGISTRY_USER:-_json_key}"
|
||||
REGISTRY_PASS="$(cat /etc/registry-secret/password 2>/dev/null || echo '')"
|
||||
|
||||
log() {
|
||||
echo "[$(date +'%Y-%m-%d %H:%M:%S')] $*"
|
||||
}
|
||||
|
||||
error() {
|
||||
echo "[$(date +'%Y-%m-%d %H:%M:%S')] ERROR: $*" >&2
|
||||
return 1
|
||||
}
|
||||
|
||||
# Verify crane is available
|
||||
if ! command -v crane &> /dev/null; then
|
||||
error "crane not found. Install google/crane image for registry operations."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
log "Starting Forgejo registry cleanup"
|
||||
log "Registry: $REGISTRY_URL"
|
||||
log "Keep versions: $KEEP_VERSIONS per image"
|
||||
log "Dry run: $DRY_RUN"
|
||||
|
||||
# Authenticate crane with registry
|
||||
if [ -n "$REGISTRY_PASS" ]; then
|
||||
echo "$REGISTRY_PASS" | crane auth login "$REGISTRY_HOST" -u "$REGISTRY_USER" --password-stdin
|
||||
log "Authenticated to $REGISTRY_HOST"
|
||||
fi
|
||||
|
||||
# List all repositories (catalog)
|
||||
# Note: This endpoint requires the registry to expose /v2/_catalog (standard OCI)
|
||||
# If not available, images must be discovered another way
|
||||
CATALOG=$(curl -s -u "${REGISTRY_USER}:${REGISTRY_PASS}" \
|
||||
"${REGISTRY_URL}/v2/_catalog" | grep -o '"repositories":\[\K[^]]*' || echo '')
|
||||
|
||||
if [ -z "$CATALOG" ]; then
|
||||
log "WARNING: Could not retrieve catalog from ${REGISTRY_URL}/v2/_catalog"
|
||||
log "Registry may not expose _catalog endpoint or credentials invalid"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# Parse repositories from catalog JSON
|
||||
REPOS=$(echo "$CATALOG" | grep -o '"[^"]*"' | tr -d '"')
|
||||
|
||||
TOTAL_DELETED=0
|
||||
|
||||
for REPO in $REPOS; do
|
||||
log "Processing repository: $REPO"
|
||||
|
||||
IMAGE="${REGISTRY_HOST}/${REPO}"
|
||||
|
||||
# Get all tags for this image
|
||||
TAGS=$(crane ls "$IMAGE" 2>/dev/null || echo "")
|
||||
|
||||
if [ -z "$TAGS" ]; then
|
||||
log " No tags found for $REPO (or access denied)"
|
||||
continue
|
||||
fi
|
||||
|
||||
# Filter out 'latest' tag and sort by creation time (newer first)
|
||||
# Note: crane doesn't provide direct date sorting; we use the order returned
|
||||
# Assumption: tags are returned newest first (not always true)
|
||||
TAG_COUNT=$(echo "$TAGS" | wc -l)
|
||||
|
||||
if [ "$TAG_COUNT" -le "$KEEP_VERSIONS" ]; then
|
||||
log " $REPO: $TAG_COUNT tags total, keeping all (≤ $KEEP_VERSIONS)"
|
||||
continue
|
||||
fi
|
||||
|
||||
# Get tags to delete (all except the first N)
|
||||
TAGS_TO_DELETE=$(echo "$TAGS" | tail -n +$((KEEP_VERSIONS + 1)))
|
||||
|
||||
for TAG in $TAGS_TO_DELETE; do
|
||||
FULL_IMAGE="${IMAGE}:${TAG}"
|
||||
DELETED_SIZE="0"
|
||||
|
||||
if [ "$DRY_RUN" = "true" ]; then
|
||||
log " [DRY RUN] Would delete: $FULL_IMAGE"
|
||||
else
|
||||
if crane delete "$FULL_IMAGE" 2>&1; then
|
||||
log " Deleted: $FULL_IMAGE"
|
||||
((TOTAL_DELETED++))
|
||||
else
|
||||
error "Failed to delete $FULL_IMAGE (may already be deleted)"
|
||||
fi
|
||||
fi
|
||||
done
|
||||
done
|
||||
|
||||
log "Cleanup complete. Total images deleted: $TOTAL_DELETED"
|
||||
|
||||
---
|
||||
apiVersion: batch/v1
|
||||
kind: CronJob
|
||||
metadata:
|
||||
name: forgejo-registry-cleanup
|
||||
namespace: cicd
|
||||
labels:
|
||||
app: forgejo-registry-cleanup
|
||||
spec:
|
||||
# Run at 2 AM UTC every day (adjust as needed)
|
||||
schedule: "0 2 * * *"
|
||||
|
||||
# Keep last 3 successful/failed runs for debugging
|
||||
successfulJobsHistoryLimit: 3
|
||||
failedJobsHistoryLimit: 3
|
||||
|
||||
# Suspend if needed (set to false to enable)
|
||||
suspend: false
|
||||
|
||||
jobTemplate:
|
||||
spec:
|
||||
# Cleanup jobs after 6 hours whether they succeeded or failed
|
||||
ttlSecondsAfterFinished: 21600
|
||||
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
app: forgejo-registry-cleanup
|
||||
spec:
|
||||
serviceAccountName: forgejo-registry-cleanup
|
||||
restartPolicy: OnFailure
|
||||
|
||||
containers:
|
||||
- name: cleanup
|
||||
# Use google/crane for registry operations
|
||||
image: gcr.io/go-containerregistry/crane:latest
|
||||
imagePullPolicy: IfNotPresent
|
||||
|
||||
env:
|
||||
- name: REGISTRY_HOST
|
||||
value: "forgejo.riotpiao.com"
|
||||
- name: KEEP_VERSIONS
|
||||
value: "3" # Keep 3 latest versions
|
||||
- name: DRY_RUN
|
||||
value: "false" # Set to "true" for dry-run mode
|
||||
- name: REGISTRY_USER
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: forgejo-registry-token
|
||||
key: username
|
||||
optional: true
|
||||
|
||||
volumeMounts:
|
||||
- name: script
|
||||
mountPath: /scripts
|
||||
- name: registry-secret
|
||||
mountPath: /etc/registry-secret
|
||||
readOnly: true
|
||||
|
||||
# Run cleanup script via entrypoint override
|
||||
command:
|
||||
- /bin/sh
|
||||
- -c
|
||||
- |
|
||||
# Install bash and curl if needed
|
||||
apk add --no-cache bash curl
|
||||
chmod +x /scripts/cleanup.sh
|
||||
/scripts/cleanup.sh
|
||||
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 128Mi
|
||||
limits:
|
||||
cpu: 500m
|
||||
memory: 512Mi
|
||||
|
||||
# Safety: kill after 30 min (prevents hanging on large registries)
|
||||
securityContext:
|
||||
runAsNonRoot: true
|
||||
runAsUser: 65534
|
||||
allowPrivilegeEscalation: false
|
||||
readOnlyRootFilesystem: false
|
||||
capabilities:
|
||||
drop:
|
||||
- ALL
|
||||
|
||||
volumes:
|
||||
- name: script
|
||||
configMap:
|
||||
name: forgejo-registry-cleanup-script
|
||||
defaultMode: 0755
|
||||
- name: registry-secret
|
||||
secret:
|
||||
secretName: forgejo-registry-token
|
||||
optional: true
|
||||
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: ServiceAccount
|
||||
metadata:
|
||||
name: forgejo-registry-cleanup
|
||||
namespace: cicd
|
||||
|
||||
---
|
||||
# No RBAC needed: this pod only talks to the registry API (external service)
|
||||
# If expanded to manage in-cluster resources, add Role/RoleBinding here
|
||||
@@ -1,72 +0,0 @@
|
||||
# ArgoCD Image Updater configuration
|
||||
# Watches Forgejo registry and updates ArgoCD Applications with new image tags
|
||||
|
||||
config:
|
||||
# Registry configuration - Forgejo allows anonymous pulls
|
||||
registries:
|
||||
- name: forgejo
|
||||
api_url: https://forgejo.riotpiao.com
|
||||
prefix: forgejo.riotpiao.com
|
||||
default: true
|
||||
insecure: false
|
||||
|
||||
# Log level
|
||||
logLevel: debug
|
||||
|
||||
# ArgoCD API server
|
||||
argocd:
|
||||
grpcWeb: true
|
||||
serverAddress: argocd-server.argocd.svc.cluster.local
|
||||
insecure: true
|
||||
plaintext: true
|
||||
|
||||
# Git write-back configuration (for multi-source Applications)
|
||||
git:
|
||||
# Commit author for image updates
|
||||
user:
|
||||
name: "ArgoCD Image Updater"
|
||||
email: "[email protected]"
|
||||
# Use SSH keys from ArgoCD's known hosts + credentials
|
||||
# Image Updater inherits ArgoCD's git credentials (mounted via ArgoCD secret)
|
||||
|
||||
# Mount ArgoCD's git credentials for write-back
|
||||
extraVolumes:
|
||||
- name: argocd-ssh-known-hosts-cm
|
||||
configMap:
|
||||
name: argocd-ssh-known-hosts-cm
|
||||
defaultMode: 0644
|
||||
- name: argocd-gpg-keys-cm
|
||||
configMap:
|
||||
name: argocd-gpg-keys-cm
|
||||
optional: true
|
||||
defaultMode: 0644
|
||||
- name: argocd-gpg-pubring
|
||||
configMap:
|
||||
name: argocd-gpg-pubring-cm
|
||||
optional: true
|
||||
defaultMode: 0644
|
||||
|
||||
extraVolumeMounts:
|
||||
- name: argocd-ssh-known-hosts-cm
|
||||
mountPath: /etc/ssh/ssh_known_hosts.d/argocd-ssh-known-hosts
|
||||
subPath: ssh_known_hosts
|
||||
- name: argocd-gpg-keys-cm
|
||||
mountPath: /etc/gpg/source
|
||||
- name: argocd-gpg-pubring
|
||||
mountPath: /etc/gpg/pubring
|
||||
|
||||
# Extra environment variables
|
||||
extraEnv:
|
||||
- name: ARGOCD_GRPC_WEB
|
||||
value: "true"
|
||||
- name: GIT_SSH_KNOWN_HOSTS_CONFIG_MAP_ENABLED
|
||||
value: "true"
|
||||
|
||||
# Resources
|
||||
resources:
|
||||
requests:
|
||||
cpu: 50m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
cpu: 200m
|
||||
memory: 128Mi
|
||||
@@ -1,137 +0,0 @@
|
||||
# Cluster-wide cleanup of stale failed/completed Jobs and Pods.
|
||||
# Runs daily at 04:00 UTC. Deletes:
|
||||
# - Failed Jobs older than 24h (any namespace)
|
||||
# - Completed Jobs older than 72h with no owning CronJob
|
||||
# - Orphan pods in Error/Failed/Evicted state older than 1h
|
||||
#
|
||||
# CronJob-owned Jobs are managed by failedJobsHistoryLimit/successfulJobsHistoryLimit,
|
||||
# but standalone Jobs (helm hooks, one-off runs, longhorn maintenance) have no TTL
|
||||
# and linger forever.
|
||||
|
||||
apiVersion: v1
|
||||
kind: ServiceAccount
|
||||
metadata:
|
||||
name: stale-job-cleanup
|
||||
namespace: kube-system
|
||||
---
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: ClusterRole
|
||||
metadata:
|
||||
name: stale-job-cleanup
|
||||
rules:
|
||||
- apiGroups: ["batch"]
|
||||
resources: ["jobs"]
|
||||
verbs: ["get", "list", "delete"]
|
||||
- apiGroups: [""]
|
||||
resources: ["pods"]
|
||||
verbs: ["get", "list", "delete"]
|
||||
---
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: ClusterRoleBinding
|
||||
metadata:
|
||||
name: stale-job-cleanup
|
||||
subjects:
|
||||
- kind: ServiceAccount
|
||||
name: stale-job-cleanup
|
||||
namespace: kube-system
|
||||
roleRef:
|
||||
kind: ClusterRole
|
||||
name: stale-job-cleanup
|
||||
apiGroup: rbac.authorization.k8s.io
|
||||
---
|
||||
apiVersion: batch/v1
|
||||
kind: CronJob
|
||||
metadata:
|
||||
name: stale-job-cleanup
|
||||
namespace: kube-system
|
||||
labels:
|
||||
app: stale-job-cleanup
|
||||
spec:
|
||||
schedule: "0 4 * * *"
|
||||
concurrencyPolicy: Forbid
|
||||
successfulJobsHistoryLimit: 3
|
||||
failedJobsHistoryLimit: 3
|
||||
jobTemplate:
|
||||
spec:
|
||||
ttlSecondsAfterFinished: 86400 # self-cleanup after 24h
|
||||
backoffLimit: 1
|
||||
activeDeadlineSeconds: 300
|
||||
template:
|
||||
spec:
|
||||
serviceAccountName: stale-job-cleanup
|
||||
restartPolicy: Never
|
||||
tolerations:
|
||||
- key: node-role.kubernetes.io/control-plane
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
containers:
|
||||
- name: cleanup
|
||||
image: alpine/k8s:1.31.0
|
||||
command:
|
||||
- sh
|
||||
- -c
|
||||
- |
|
||||
set -e
|
||||
NOW=$(date +%s)
|
||||
|
||||
echo "=== Cleaning failed Jobs older than 24h ==="
|
||||
kubectl get jobs --all-namespaces -o json | \
|
||||
jq -r '.items[] |
|
||||
select(.status.conditions[]?.type == "Failed") |
|
||||
select(.status.completionTime or .status.startTime) |
|
||||
"\(.metadata.namespace) \(.metadata.name) \(.status.startTime // .status.completionTime // .metadata.creationTimestamp)"' | \
|
||||
while read -r NS NAME TS; do
|
||||
JOB_EPOCH=$(date -d "$TS" +%s 2>/dev/null || echo 0)
|
||||
AGE_H=$(( (NOW - JOB_EPOCH) / 3600 ))
|
||||
if [ "$AGE_H" -ge 24 ]; then
|
||||
echo "[delete] $NS/$NAME (failed ${AGE_H}h ago)"
|
||||
kubectl delete job "$NAME" -n "$NS" --cascade=foreground 2>/dev/null || true
|
||||
fi
|
||||
done
|
||||
|
||||
echo ""
|
||||
echo "=== Cleaning completed standalone Jobs older than 72h ==="
|
||||
kubectl get jobs --all-namespaces -o json | \
|
||||
jq -r '.items[] |
|
||||
select(.status.succeeded >= 1) |
|
||||
select((.metadata.ownerReferences // []) | length == 0) |
|
||||
"\(.metadata.namespace) \(.metadata.name) \(.status.completionTime // .metadata.creationTimestamp)"' | \
|
||||
while read -r NS NAME TS; do
|
||||
JOB_EPOCH=$(date -d "$TS" +%s 2>/dev/null || echo 0)
|
||||
AGE_H=$(( (NOW - JOB_EPOCH) / 3600 ))
|
||||
if [ "$AGE_H" -ge 72 ]; then
|
||||
echo "[delete] $NS/$NAME (completed ${AGE_H}h ago, no owner)"
|
||||
kubectl delete job "$NAME" -n "$NS" --cascade=foreground 2>/dev/null || true
|
||||
fi
|
||||
done
|
||||
|
||||
echo ""
|
||||
echo "=== Cleaning orphan Error/Failed/Evicted pods older than 1h ==="
|
||||
# Evicted pods show as Failed with reason Evicted
|
||||
kubectl get pods --all-namespaces -o json | \
|
||||
jq -r '.items[] |
|
||||
select(
|
||||
.status.phase == "Failed" or
|
||||
(.status.reason // "") == "Evicted" or
|
||||
(.status.containerStatuses // [] | any(.state.terminated.reason == "Error"))
|
||||
) |
|
||||
select((.metadata.ownerReferences // []) | all(.kind != "Job")) |
|
||||
"\(.metadata.namespace) \(.metadata.name) \(.metadata.creationTimestamp)"' | \
|
||||
while read -r NS NAME TS; do
|
||||
POD_EPOCH=$(date -d "$TS" +%s 2>/dev/null || echo 0)
|
||||
AGE_H=$(( (NOW - POD_EPOCH) / 3600 ))
|
||||
if [ "$AGE_H" -ge 1 ]; then
|
||||
echo "[delete] $NS/$NAME (error/evicted ${AGE_H}h ago)"
|
||||
kubectl delete pod "$NAME" -n "$NS" --force 2>/dev/null || true
|
||||
fi
|
||||
done
|
||||
|
||||
echo ""
|
||||
echo "Cleanup complete"
|
||||
resources:
|
||||
requests:
|
||||
cpu: 50m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
cpu: 250m
|
||||
memory: 128Mi
|
||||
@@ -1,13 +1,8 @@
|
||||
apiVersion: kustomize.config.k8s.io/v1beta1
|
||||
kind: Kustomization
|
||||
# Dedicated per-app CNPG clusters. NO top-level `namespace:` — each Cluster
|
||||
# carries its own ns (iam / temporal / poimen / paperless); a transformer would
|
||||
# wrongly collapse them. poimen ns created by poimen-root app, memory-db
|
||||
# deployed into it. paperless ns declared in namespaces.yaml above.
|
||||
# carries its own ns (iam / temporal); a transformer would wrongly collapse them.
|
||||
resources:
|
||||
- namespaces.yaml
|
||||
- authentik-db.yaml
|
||||
- temporal-db.yaml
|
||||
- memory-db.yaml
|
||||
- paperless-db.yaml
|
||||
- obsidian-vault-pvc.yaml
|
||||
|
||||
@@ -1,36 +0,0 @@
|
||||
# Dedicated CNPG Postgres for Poimen Memory (GitOps, wave 2).
|
||||
# Uses default longhorn storage class (3 replicas, dataLocality disabled).
|
||||
# CNPG generates secret `memory-db-app` + service `memory-db-rw` in ns poimen.
|
||||
apiVersion: postgresql.cnpg.io/v1
|
||||
kind: Cluster
|
||||
metadata:
|
||||
name: memory-db
|
||||
namespace: poimen
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-options: SkipDryRunOnMissingResource=true
|
||||
spec:
|
||||
instances: 2
|
||||
imageName: ghcr.io/cloudnative-pg/postgresql:16.2
|
||||
bootstrap:
|
||||
initdb:
|
||||
database: memory
|
||||
owner: app
|
||||
encoding: UTF8
|
||||
localeCollate: C
|
||||
localeCType: C
|
||||
postInitApplicationSQL:
|
||||
- "CREATE EXTENSION vector;"
|
||||
enableSuperuserAccess: false
|
||||
resources:
|
||||
requests: { memory: "512Mi", cpu: "250m" }
|
||||
limits: { memory: "2Gi", cpu: "1" }
|
||||
storage:
|
||||
size: 20Gi
|
||||
storageClass: longhorn
|
||||
affinity:
|
||||
podAntiAffinityType: preferred
|
||||
topologyKey: kubernetes.io/hostname
|
||||
tolerations:
|
||||
- key: node-role.kubernetes.io/control-plane
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
@@ -1,7 +1,6 @@
|
||||
# DB clusters are wave 2 — their namespaces must exist first (their apps that
|
||||
# would CreateNamespace run later, w3/w8). Declared here so the databases App
|
||||
# creates them. authentik/vault/temporal/paperless CreateNamespace=true then
|
||||
# no-ops. poimen namespace created by poimen-root app (wave 7).
|
||||
# creates them. authentik/vault/temporal CreateNamespace=true then no-ops.
|
||||
apiVersion: v1
|
||||
kind: Namespace
|
||||
metadata:
|
||||
@@ -11,16 +10,3 @@ apiVersion: v1
|
||||
kind: Namespace
|
||||
metadata:
|
||||
name: temporal
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Namespace
|
||||
metadata:
|
||||
name: paperless
|
||||
---
|
||||
# Needed here (not just immich's own CreateNamespace=true at wave 8) because
|
||||
# k8s/infra/iam's PostSync job (wave 3) has a RoleBinding targeting this
|
||||
# namespace - same ordering reason as paperless above.
|
||||
apiVersion: v1
|
||||
kind: Namespace
|
||||
metadata:
|
||||
name: immich
|
||||
|
||||
@@ -1,18 +0,0 @@
|
||||
---
|
||||
# Obsidian vault PVC — shared storage for REST API + UI pods
|
||||
# ReadWriteMany so both obsidian-server and obsidian-ui can mount simultaneously
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: obsidian-vault
|
||||
namespace: poimen
|
||||
labels:
|
||||
app.kubernetes.io/name: obsidian-server
|
||||
app.kubernetes.io/part-of: poimen-memory
|
||||
spec:
|
||||
accessModes:
|
||||
- ReadWriteMany
|
||||
storageClassName: longhorn
|
||||
resources:
|
||||
requests:
|
||||
storage: 10Gi
|
||||
@@ -1,36 +0,0 @@
|
||||
# Dedicated CNPG Postgres for paperless-ngx (GitOps, wave 2 — before the
|
||||
# paperless app at w8). Same recipe as memory-db: default longhorn storage
|
||||
# class (3 replicas), 2 instances, 20Gi.
|
||||
# CNPG generates secret `paperless-db-app` + service `paperless-db-rw` in ns
|
||||
# paperless; the paperless Deployment reads them locally.
|
||||
apiVersion: postgresql.cnpg.io/v1
|
||||
kind: Cluster
|
||||
metadata:
|
||||
name: paperless-db
|
||||
namespace: paperless
|
||||
annotations:
|
||||
argocd.argoproj.io/sync-options: SkipDryRunOnMissingResource=true
|
||||
spec:
|
||||
instances: 2
|
||||
imageName: ghcr.io/cloudnative-pg/postgresql:16.2
|
||||
bootstrap:
|
||||
initdb:
|
||||
database: paperless
|
||||
owner: app
|
||||
encoding: UTF8
|
||||
localeCollate: C
|
||||
localeCType: C
|
||||
enableSuperuserAccess: false
|
||||
resources:
|
||||
requests: { memory: "512Mi", cpu: "250m" }
|
||||
limits: { memory: "2Gi", cpu: "1" }
|
||||
storage:
|
||||
size: 20Gi
|
||||
storageClass: longhorn
|
||||
affinity:
|
||||
podAntiAffinityType: preferred
|
||||
topologyKey: kubernetes.io/hostname
|
||||
tolerations:
|
||||
- key: node-role.kubernetes.io/control-plane
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
@@ -36,4 +36,3 @@ data:
|
||||
valid_volumes:
|
||||
- /docker-certs/client
|
||||
network: host
|
||||
docker_host: automount
|
||||
|
||||
@@ -34,11 +34,7 @@ spec:
|
||||
command: ["sh", "-c"]
|
||||
args:
|
||||
- |
|
||||
# Always re-register to keep labels in sync with values.yaml.
|
||||
# Without this, changing a runner label requires manually deleting
|
||||
# the PVC or .runner file — not GitOps-friendly.
|
||||
rm -f /data/.runner
|
||||
forgejo-runner register --no-interactive \
|
||||
test -f /data/.runner || forgejo-runner register --no-interactive \
|
||||
--instance {{ .Values.runner.forgejoUrl }} \
|
||||
--token $(RUNNER_TOKEN) \
|
||||
--name {{ .Values.runner.name }} \
|
||||
@@ -60,18 +56,20 @@ spec:
|
||||
containers:
|
||||
- name: runner
|
||||
image: {{ .Values.runner.image.repository }}:{{ .Values.runner.image.tag }}
|
||||
command: ["sh", "-c", "while ! wget -q -O- http://localhost:2375/_ping >/dev/null 2>&1; do echo 'waiting for dind...'; sleep 2; done; echo 'dind ready'; forgejo-runner daemon --config /etc/forgejo-runner/config.yaml"]
|
||||
command: ["sh", "-c", "forgejo-runner daemon --config /etc/forgejo-runner/config.yaml"]
|
||||
workingDir: /data
|
||||
env:
|
||||
- name: DOCKER_HOST
|
||||
value: tcp://localhost:2375
|
||||
value: tcp://localhost:2376
|
||||
- name: DOCKER_TLS_VERIFY
|
||||
value: "1"
|
||||
- name: DOCKER_CERT_PATH
|
||||
value: /docker-certs/client
|
||||
volumeMounts:
|
||||
- name: runner-data
|
||||
mountPath: /data
|
||||
- name: docker-certs
|
||||
mountPath: /docker-certs
|
||||
- name: docker-sock
|
||||
mountPath: /run
|
||||
- name: homelab-ca
|
||||
mountPath: /etc/ssl/certs/homelab-ca.pem
|
||||
subPath: ca.crt
|
||||
@@ -87,12 +85,10 @@ spec:
|
||||
privileged: true # required for DinD; cicd namespace is labelled privileged
|
||||
env:
|
||||
- name: DOCKER_TLS_CERTDIR
|
||||
value: ""
|
||||
value: /docker-certs
|
||||
volumeMounts:
|
||||
- name: docker-certs
|
||||
mountPath: /docker-certs
|
||||
- name: docker-sock
|
||||
mountPath: /run
|
||||
- name: dind-storage
|
||||
mountPath: /var/lib/docker
|
||||
- name: homelab-ca
|
||||
@@ -117,8 +113,6 @@ spec:
|
||||
claimName: {{ .Release.Name }}-dind
|
||||
- name: docker-certs
|
||||
emptyDir: {} # DinD regenerates mTLS certs on each start
|
||||
- name: docker-sock
|
||||
emptyDir: {} # Shared docker socket between dind and runner
|
||||
- name: homelab-ca
|
||||
# homelab-ca is a ConfigMap (public CA trust bundle), not a Secret.
|
||||
# The volumeMounts use subPath: ca.crt to project the single cert file.
|
||||
|
||||
@@ -1,152 +0,0 @@
|
||||
{{- if .Values.gc.enabled }}
|
||||
# Garbage-collects DinD Docker images/volumes/build-cache and actcache across
|
||||
# ALL forgejo-runner pods. Prevents PVC fill-up that breaks CI runs.
|
||||
# Only rendered once (enable in default values.yaml, disable in per-runner overrides).
|
||||
apiVersion: v1
|
||||
kind: ServiceAccount
|
||||
metadata:
|
||||
name: runner-gc
|
||||
namespace: {{ .Release.Namespace }}
|
||||
---
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: Role
|
||||
metadata:
|
||||
name: runner-gc
|
||||
namespace: {{ .Release.Namespace }}
|
||||
rules:
|
||||
- apiGroups: [""]
|
||||
resources: ["pods"]
|
||||
verbs: ["get", "list"]
|
||||
- apiGroups: [""]
|
||||
resources: ["pods/exec"]
|
||||
verbs: ["create"]
|
||||
---
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: RoleBinding
|
||||
metadata:
|
||||
name: runner-gc
|
||||
namespace: {{ .Release.Namespace }}
|
||||
subjects:
|
||||
- kind: ServiceAccount
|
||||
name: runner-gc
|
||||
namespace: {{ .Release.Namespace }}
|
||||
roleRef:
|
||||
kind: Role
|
||||
name: runner-gc
|
||||
apiGroup: rbac.authorization.k8s.io
|
||||
---
|
||||
apiVersion: batch/v1
|
||||
kind: CronJob
|
||||
metadata:
|
||||
name: forgejo-runner-gc
|
||||
namespace: {{ .Release.Namespace }}
|
||||
labels:
|
||||
app: forgejo-runner-gc
|
||||
spec:
|
||||
schedule: {{ .Values.gc.schedule | quote }}
|
||||
concurrencyPolicy: Forbid
|
||||
successfulJobsHistoryLimit: 3
|
||||
failedJobsHistoryLimit: 3
|
||||
jobTemplate:
|
||||
spec:
|
||||
backoffLimit: 1
|
||||
activeDeadlineSeconds: 900
|
||||
template:
|
||||
spec:
|
||||
serviceAccountName: runner-gc
|
||||
restartPolicy: Never
|
||||
tolerations:
|
||||
- key: node-role.kubernetes.io/control-plane
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
containers:
|
||||
- name: gc
|
||||
image: {{ .Values.gc.image }}
|
||||
command:
|
||||
- sh
|
||||
- -c
|
||||
- |
|
||||
set -e
|
||||
|
||||
# Iterate all forgejo-runner pods (golang, rust, node)
|
||||
PODS=$(kubectl -n {{ .Release.Namespace }} get pod \
|
||||
-l app -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.metadata.labels.app}{"\n"}{end}' \
|
||||
| grep 'forgejo-runner-' | awk '{print $1}')
|
||||
|
||||
if [ -z "$PODS" ]; then
|
||||
echo "no forgejo-runner pods found, skipping"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
for POD in $PODS; do
|
||||
echo "===== $POD ====="
|
||||
|
||||
# 1. Docker image prune (DinD sidecar)
|
||||
echo "[docker] before:"
|
||||
kubectl -n {{ .Release.Namespace }} exec "$POD" -c dind -- docker system df 2>/dev/null || true
|
||||
|
||||
echo "[docker] pruning non-latest images older than {{ .Values.gc.pruneAge }}..."
|
||||
# Keep :latest tagged images, delete all others older than pruneAge.
|
||||
# docker image prune can't filter by tag, so we list and selectively rmi.
|
||||
kubectl -n {{ .Release.Namespace }} exec "$POD" -c dind -- \
|
||||
sh -c '
|
||||
# Remove dangling (untagged) images older than {{ .Values.gc.pruneAge }}
|
||||
docker image prune -f --filter "until={{ .Values.gc.pruneAge }}" 2>/dev/null
|
||||
|
||||
# Remove tagged non-latest images older than {{ .Values.gc.pruneAge }}
|
||||
CUTOFF=$(date -d "-{{ .Values.gc.pruneAgeHours }} hours" +%s 2>/dev/null || date -v-{{ .Values.gc.pruneAgeHours }}H +%s)
|
||||
docker images --format "{{"{{"}} .Repository {{"}}"}}:{{"{{"}} .Tag {{"}}"}} {{"{{"}} .CreatedAt {{"}}"}}" | while read -r IMAGE_TAG CREATED_REST; do
|
||||
TAG=$(echo "$IMAGE_TAG" | rev | cut -d: -f1 | rev)
|
||||
# Skip latest-tagged images
|
||||
if [ "$TAG" = "latest" ]; then
|
||||
echo "[keep] $IMAGE_TAG (latest)"
|
||||
continue
|
||||
fi
|
||||
# Check image age via inspect
|
||||
CREATED_TS=$(docker inspect --format="{{"{{"}} .Created {{"}}"}}" "$IMAGE_TAG" 2>/dev/null | head -1)
|
||||
if [ -z "$CREATED_TS" ]; then continue; fi
|
||||
IMAGE_EPOCH=$(date -d "$CREATED_TS" +%s 2>/dev/null || date -jf "%Y-%m-%dT%H:%M:%S" "$(echo $CREATED_TS | cut -dT -f1-2 | cut -d. -f1)" +%s 2>/dev/null || echo 0)
|
||||
if [ "$IMAGE_EPOCH" -lt "$CUTOFF" ] 2>/dev/null; then
|
||||
echo "[delete] $IMAGE_TAG (older than {{ .Values.gc.pruneAge }})"
|
||||
docker rmi -f "$IMAGE_TAG" 2>/dev/null || true
|
||||
else
|
||||
echo "[keep] $IMAGE_TAG (recent)"
|
||||
fi
|
||||
done
|
||||
' 2>/dev/null || true
|
||||
|
||||
echo "[docker] pruning build cache unused >{{ .Values.gc.pruneAge }}..."
|
||||
kubectl -n {{ .Release.Namespace }} exec "$POD" -c dind -- \
|
||||
docker builder prune -af --filter "until={{ .Values.gc.pruneAge }}" 2>/dev/null || true
|
||||
|
||||
echo "[docker] pruning dangling volumes..."
|
||||
kubectl -n {{ .Release.Namespace }} exec "$POD" -c dind -- \
|
||||
docker volume prune -af 2>/dev/null || true
|
||||
|
||||
echo "[docker] after:"
|
||||
kubectl -n {{ .Release.Namespace }} exec "$POD" -c dind -- docker system df 2>/dev/null || true
|
||||
|
||||
# 2. Actcache cleanup (runner container)
|
||||
echo "[actcache] cleaning incomplete and stale cache entries..."
|
||||
kubectl -n {{ .Release.Namespace }} exec "$POD" -c runner -- \
|
||||
sh -c '
|
||||
# Delete incomplete/partial cache uploads immediately (tmp dirs)
|
||||
find /data/.cache/actcache/cache -name "tmp" -type d -exec rm -rf {} + 2>/dev/null || true
|
||||
# Delete cache entries not accessed in last {{ .Values.gc.actcacheMaxAgeDays }} day(s)
|
||||
find /data/.cache/actcache/cache -type f -mtime +{{ .Values.gc.actcacheMaxAgeDays }} -delete 2>/dev/null || true
|
||||
# Clean up empty directories
|
||||
find /data/.cache/actcache/cache -type d -empty -delete 2>/dev/null || true
|
||||
echo "actcache size: $(du -sh /data/.cache/actcache/cache 2>/dev/null | cut -f1)"
|
||||
' || true
|
||||
|
||||
echo ""
|
||||
done
|
||||
echo "GC complete"
|
||||
resources:
|
||||
requests:
|
||||
cpu: 50m
|
||||
memory: 64Mi
|
||||
limits:
|
||||
cpu: 250m
|
||||
memory: 128Mi
|
||||
{{- end }}
|
||||
@@ -2,15 +2,8 @@
|
||||
# runner instance. Only runner.name and runner.labels differ -- everything
|
||||
# else (image, dind, persistence, tolerations, nodeSelector) is shared.
|
||||
#
|
||||
# Label image: node:22-bookworm — Debian, root, apt-get, Node.js, npm, git.
|
||||
# Install docker in workflow steps as needed.
|
||||
# node:22-bookworm ships Node natively, so unlike the golang/rust instances,
|
||||
# jobs on this runner need no "install node" step before actions/checkout.
|
||||
runner:
|
||||
image:
|
||||
repository: code.forgejo.org/forgejo/runner
|
||||
tag: "6"
|
||||
name: node-runner
|
||||
labels: "node:docker://node:22-bookworm"
|
||||
|
||||
# GC CronJob renders only from the default (golang) values to avoid duplicates
|
||||
gc:
|
||||
enabled: false
|
||||
|
||||
@@ -2,17 +2,9 @@
|
||||
# runner instance. Only runner.name and runner.labels differ -- everything
|
||||
# else (image, dind, persistence, tolerations, nodeSelector) is shared.
|
||||
#
|
||||
# Label image: rust:1-bookworm — Debian, root, apt-get, Rust, cargo, git.
|
||||
# Install Node.js/docker in workflow steps as needed.
|
||||
# rust:1.83-bookworm -- verified this tag exists (docker manifest inspect)
|
||||
# before pinning it, per this repo's convention of not trusting a tag exists
|
||||
# without checking.
|
||||
runner:
|
||||
image:
|
||||
repository: code.forgejo.org/forgejo/runner
|
||||
tag: "6"
|
||||
name: rust-runner
|
||||
labels: "rust:docker://rust:1-bookworm"
|
||||
|
||||
|
||||
|
||||
# GC CronJob renders only from the default (golang) values to avoid duplicates
|
||||
gc:
|
||||
enabled: false
|
||||
labels: "rust:docker://rust:1.83-bookworm"
|
||||
|
||||
@@ -1,13 +1,20 @@
|
||||
runner:
|
||||
image:
|
||||
repository: code.forgejo.org/forgejo/runner
|
||||
tag: "6"
|
||||
tag: "6" # pin exact release before apply
|
||||
name: golang-runner
|
||||
# Label image is what workflow steps run in (NOT the runner daemon image).
|
||||
# golang:1.26-bookworm: Debian, root, apt-get, Go, git.
|
||||
# TODO: Switch to custom image once build-runner-images.yml pushes images
|
||||
labels: "golang:docker://golang:1.26-bookworm"
|
||||
# Default image is only used when a job's `container:` doesn't override it
|
||||
# (both ci.yaml and build.yaml in homelab-frontend do). Retired the old
|
||||
# "docker" label entirely; every repo this runner serves is Go, so this
|
||||
# instance carries the golang toolchain and its own dind sidecar builds and
|
||||
# pushes that repo's images too -- there is no separate generic runner
|
||||
# anymore.
|
||||
labels: "golang:docker://golang:1.25-bookworm"
|
||||
# In-cluster Service (:3000) — direct, avoids the ingress/public-hostname hop
|
||||
# (the public URL is :443 which forgejo doesn't serve; runner got i/o timeout).
|
||||
forgejoUrl: http://forgejo-gitea-http.cicd.svc.cluster.local:3000
|
||||
# tokenSecret: name of the K8s Secret that holds the runner registration token
|
||||
# created automatically by the helmfile presync hook (see helmfile.yaml.gotmpl)
|
||||
tokenSecret: runner-token
|
||||
resources:
|
||||
requests:
|
||||
@@ -32,7 +39,7 @@ dind:
|
||||
persistence:
|
||||
reg:
|
||||
storageClass: longhorn # Unified StorageClass (3 replicas)
|
||||
size: 20Gi # .runner registration file + action tool cache + actcache artifacts
|
||||
size: 1Gi # .runner registration file + config — survives pod restarts
|
||||
dind:
|
||||
storageClass: longhorn # Unified StorageClass (3 replicas)
|
||||
size: 30Gi # docker layer cache — keeps rebuilds fast across restarts
|
||||
@@ -42,18 +49,7 @@ tolerations:
|
||||
operator: Exists
|
||||
effect: NoSchedule
|
||||
|
||||
# Pin to az-b (talos-cp-2) — more Longhorn storage than az-a (worker-1 over-provisioned).
|
||||
# RWO PVCs will recreate on talos-cp-2 when nodeSelector changes.
|
||||
# Pin to az-a (talos-cp-1) — sole Longhorn node; RWO PVCs (reg/dind cache) only
|
||||
# attach there.
|
||||
nodeSelector:
|
||||
topology.kubernetes.io/zone: az-b
|
||||
|
||||
# GC CronJob — prunes Docker images/volumes/build-cache and actcache across
|
||||
# ALL forgejo-runner pods. Only enable in default values (golang instance);
|
||||
# disable in per-runner overrides so it renders once.
|
||||
gc:
|
||||
enabled: true
|
||||
schedule: "*/30 * * * *" # every 30 minutes
|
||||
image: alpine/k8s:1.31.0
|
||||
pruneAge: "30m" # Docker artifacts unused longer than this get pruned
|
||||
pruneAgeHours: 0.5 # Same as pruneAge but numeric for date arithmetic in shell
|
||||
actcacheMaxAgeDays: 1 # actcache files older than N days (aggressive for heavy Rust cargo builds)
|
||||
topology.kubernetes.io/zone: az-a
|
||||
|
||||
@@ -0,0 +1,167 @@
|
||||
# Authentik OAuth provisioning — PostSync hook, reruns on every ArgoCD sync
|
||||
# (hook-delete-policy: BeforeHookCreation deletes the previous run's Job before
|
||||
# creating a new one, so this stays reconciled the same way the rest of the
|
||||
# cluster does — no separate manual bootstrap step like setup_talos_iam.sh /
|
||||
# provision_oidc.py, which never got migrated off the old helmfile workflow).
|
||||
#
|
||||
# What it does (see scripts/authentik-provision.py docstring): creates the
|
||||
# "groups" scope mapping, homelab-admins / grafana-admins groups, the "rock"
|
||||
# admin user, OAuth2 providers + Applications for grafana/minio/forgejo/argocd,
|
||||
# and binds homelab-admins to all of them. The script is generated into the
|
||||
# authentik-provision-script ConfigMap by kustomize configMapGenerator (see
|
||||
# kustomization.yaml), not embedded here.
|
||||
#
|
||||
# RBAC: this Job only touches Secrets (get existing client secrets, create new
|
||||
# ones for forgejo/argocd/rock) across the namespaces those services live in.
|
||||
# It never touches any other resource type.
|
||||
apiVersion: v1
|
||||
kind: ServiceAccount
|
||||
metadata:
|
||||
name: authentik-provisioner
|
||||
namespace: iam
|
||||
---
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: ClusterRole
|
||||
metadata:
|
||||
name: authentik-provisioner
|
||||
rules:
|
||||
- apiGroups: [""]
|
||||
resources: ["secrets"]
|
||||
verbs: ["get", "list", "create", "update", "patch"]
|
||||
---
|
||||
# One RoleBinding per namespace the script touches (least-privilege: Secrets
|
||||
# only, and only in these 5 namespaces — not a cluster-wide ClusterRoleBinding).
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: RoleBinding
|
||||
metadata:
|
||||
name: authentik-provisioner
|
||||
namespace: iam
|
||||
subjects:
|
||||
- kind: ServiceAccount
|
||||
name: authentik-provisioner
|
||||
namespace: iam
|
||||
roleRef:
|
||||
kind: ClusterRole
|
||||
name: authentik-provisioner
|
||||
apiGroup: rbac.authorization.k8s.io
|
||||
---
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: RoleBinding
|
||||
metadata:
|
||||
name: authentik-provisioner
|
||||
namespace: cicd
|
||||
subjects:
|
||||
- kind: ServiceAccount
|
||||
name: authentik-provisioner
|
||||
namespace: iam
|
||||
roleRef:
|
||||
kind: ClusterRole
|
||||
name: authentik-provisioner
|
||||
apiGroup: rbac.authorization.k8s.io
|
||||
---
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: RoleBinding
|
||||
metadata:
|
||||
name: authentik-provisioner
|
||||
namespace: argocd
|
||||
subjects:
|
||||
- kind: ServiceAccount
|
||||
name: authentik-provisioner
|
||||
namespace: iam
|
||||
roleRef:
|
||||
kind: ClusterRole
|
||||
name: authentik-provisioner
|
||||
apiGroup: rbac.authorization.k8s.io
|
||||
---
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: RoleBinding
|
||||
metadata:
|
||||
name: authentik-provisioner
|
||||
namespace: logging
|
||||
subjects:
|
||||
- kind: ServiceAccount
|
||||
name: authentik-provisioner
|
||||
namespace: iam
|
||||
roleRef:
|
||||
kind: ClusterRole
|
||||
name: authentik-provisioner
|
||||
apiGroup: rbac.authorization.k8s.io
|
||||
---
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: RoleBinding
|
||||
metadata:
|
||||
name: authentik-provisioner
|
||||
namespace: storage
|
||||
subjects:
|
||||
- kind: ServiceAccount
|
||||
name: authentik-provisioner
|
||||
namespace: iam
|
||||
roleRef:
|
||||
kind: ClusterRole
|
||||
name: authentik-provisioner
|
||||
apiGroup: rbac.authorization.k8s.io
|
||||
---
|
||||
apiVersion: batch/v1
|
||||
kind: Job
|
||||
metadata:
|
||||
name: authentik-provision
|
||||
namespace: iam
|
||||
annotations:
|
||||
argocd.argoproj.io/hook: PostSync
|
||||
argocd.argoproj.io/hook-delete-policy: BeforeHookCreation
|
||||
spec:
|
||||
ttlSecondsAfterFinished: 600
|
||||
backoffLimit: 3
|
||||
template:
|
||||
spec:
|
||||
serviceAccountName: authentik-provisioner
|
||||
restartPolicy: Never
|
||||
securityContext:
|
||||
runAsNonRoot: true
|
||||
runAsUser: 1000
|
||||
seccompProfile:
|
||||
type: RuntimeDefault
|
||||
containers:
|
||||
- name: provision
|
||||
image: python:3.12-alpine
|
||||
securityContext:
|
||||
allowPrivilegeEscalation: false
|
||||
capabilities:
|
||||
drop: ["ALL"]
|
||||
env:
|
||||
- name: AUTHENTIK_BOOTSTRAP_TOKEN
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: authentik-secrets
|
||||
key: AUTHENTIK_BOOTSTRAP_TOKEN
|
||||
volumeMounts:
|
||||
- name: script
|
||||
mountPath: /script
|
||||
command:
|
||||
- /bin/sh
|
||||
- -c
|
||||
- |
|
||||
set -e
|
||||
echo "waiting for authentik-server..."
|
||||
until wget -q -O /dev/null http://authentik-server.iam.svc.cluster.local/-/health/ready/ 2>/dev/null; do
|
||||
sleep 5
|
||||
done
|
||||
echo "installing kubectl (via python urllib - no apk/curl: this"
|
||||
echo "container runs as non-root UID 1000 and can't write to"
|
||||
echo "apk's directories or /usr/local/bin, both root-owned in"
|
||||
echo "the python:3.12-alpine image; /tmp is world-writable)..."
|
||||
python3 -c "
|
||||
import urllib.request, os, stat
|
||||
kver = urllib.request.urlopen('https://dl.k8s.io/release/stable.txt').read().decode().strip()
|
||||
url = f'https://dl.k8s.io/release/{kver}/bin/linux/amd64/kubectl'
|
||||
urllib.request.urlretrieve(url, '/tmp/kubectl')
|
||||
st = os.stat('/tmp/kubectl')
|
||||
os.chmod('/tmp/kubectl', st.st_mode | stat.S_IEXEC)
|
||||
"
|
||||
export PATH="/tmp:$PATH"
|
||||
echo "running provisioning script..."
|
||||
python3 /script/authentik-provision.py
|
||||
volumes:
|
||||
- name: script
|
||||
configMap:
|
||||
name: authentik-provision-script
|
||||
@@ -1,12 +1,35 @@
|
||||
apiVersion: kustomize.config.k8s.io/v1beta1
|
||||
kind: Kustomization
|
||||
|
||||
# NOTE: no top-level `namespace:` transformer here (removed) - it used to
|
||||
# force-rewrite metadata.namespace to "iam" on every resource in this
|
||||
# kustomization, which was harmless while every manifest here only ever
|
||||
# targeted the iam namespace itself. authentik-provision-job.yaml's
|
||||
# RoleBindings deliberately target cicd/argocd/logging/storage (least-
|
||||
# privilege access for the authentik-provisioner ServiceAccount to touch
|
||||
# Secrets in those namespaces) - the namespace transformer would have
|
||||
# silently rewritten all of them back to iam, breaking the RBAC. Every
|
||||
# manifest in this directory already sets its own explicit
|
||||
# metadata.namespace, so dropping the transformer changes nothing for the
|
||||
# existing resources/.
|
||||
resources:
|
||||
- authentik-provision-job.yaml
|
||||
- rbac-dashboard-rolebinding.yaml
|
||||
|
||||
# IAM provisioning is manual-only (security-sensitive).
|
||||
# Script: scripts/iam/authentik-provision.py
|
||||
# Run:
|
||||
# export AUTHENTIK_BOOTSTRAP_TOKEN=$(kubectl -n iam get secret authentik-secrets \
|
||||
# -o jsonpath='{.data.AUTHENTIK_BOOTSTRAP_TOKEN}' | base64 -d)
|
||||
# python3 scripts/iam/authentik-provision.py
|
||||
# Provisioning/verification python lives in scripts/*.py (real files, linted +
|
||||
# diff-friendly) and is generated into ConfigMaps here rather than embedded in
|
||||
# the job YAML. disableNameSuffixHash keeps the names stable so the Jobs'
|
||||
# configMap volume refs and PostSync hook-delete semantics keep working; each
|
||||
# hook Job is recreated per sync so it always mounts the latest script.
|
||||
configMapGenerator:
|
||||
- name: authentik-provision-script
|
||||
namespace: iam
|
||||
files:
|
||||
- authentik-provision.py=scripts/authentik-provision.py
|
||||
|
||||
generatorOptions:
|
||||
disableNameSuffixHash: true
|
||||
# authentik-migrations-job.yaml removed — redundant + broken. The authentik
|
||||
# `server` entrypoint runs migrations itself; this standalone job lacked the
|
||||
# authentik-secrets envFrom (Secret key missing) and always failed.
|
||||
# SOPS secrets (*.enc.yaml) handled by ArgoCD SOPS plugin at sync time
|
||||
# authentik/vault deployed via ArgoCD Helm source
|
||||
|
||||
@@ -0,0 +1,400 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Authentik OAuth provisioning - idempotent, safe to re-run (ArgoCD PostSync hook).
|
||||
|
||||
Creates/updates, in order:
|
||||
1. A custom "groups" OAuth2 scope mapping (Authentik ships openid/email/profile
|
||||
by default but NOT groups - required for ArgoCD RBAC group mapping and
|
||||
Grafana's role_attribute_path, both of which read a `groups` claim).
|
||||
2. Groups: homelab-admins (is_superuser=true), grafana-admins.
|
||||
3. User "rock": created if missing, always (re-)synced into both groups above.
|
||||
Password is generated once and only written to the k8s Secret
|
||||
rock-credentials (iam ns) the first time the user is created - re-runs
|
||||
never rotate an existing password.
|
||||
4. OAuth2/OIDC providers + Applications for: grafana, minio, forgejo, argocd.
|
||||
Client secrets are read from existing k8s Secrets (grafana-oidc, minio-oidc)
|
||||
if present, or generated once and written out (forgejo-oidc, oidc-secret)
|
||||
the first time.
|
||||
5. PolicyBinding of homelab-admins -> every Application above, so "rock" (and
|
||||
anyone else in that group) has guaranteed access regardless of each app's
|
||||
default visibility.
|
||||
|
||||
Talks to Authentik over the in-cluster Service (authentik-server.iam.svc:80),
|
||||
authenticating with the bootstrap token. Everything is done with GET-then-
|
||||
create-or-patch so this can be re-run on every ArgoCD sync without duplicating
|
||||
or clobbering objects (PostSync hook, not a one-shot Job with hook-delete).
|
||||
|
||||
kubectl is used only to read/write the small set of Secrets this script
|
||||
touches - it shells out rather than using the Python k8s client to keep the
|
||||
container image to stdlib Python + the kubectl binary, no pip installs.
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
import secrets
|
||||
import string
|
||||
import subprocess
|
||||
import sys
|
||||
import urllib.error
|
||||
import urllib.request
|
||||
|
||||
AUTHENTIK_URL = "http://authentik-server.iam.svc.cluster.local"
|
||||
TOKEN = os.environ["AUTHENTIK_BOOTSTRAP_TOKEN"]
|
||||
|
||||
|
||||
def api(method, path, data=None):
|
||||
url = f"{AUTHENTIK_URL}{path}"
|
||||
body = json.dumps(data).encode() if data is not None else None
|
||||
req = urllib.request.Request(
|
||||
url,
|
||||
data=body,
|
||||
method=method,
|
||||
headers={
|
||||
"Authorization": f"Bearer {TOKEN}",
|
||||
"Content-Type": "application/json",
|
||||
},
|
||||
)
|
||||
try:
|
||||
with urllib.request.urlopen(req, timeout=30) as resp:
|
||||
raw = resp.read()
|
||||
return resp.status, (json.loads(raw) if raw else {})
|
||||
except urllib.error.HTTPError as e:
|
||||
raw = e.read()
|
||||
try:
|
||||
parsed = json.loads(raw) if raw else {}
|
||||
except json.JSONDecodeError:
|
||||
parsed = {"raw": raw.decode(errors="replace")}
|
||||
return e.code, parsed
|
||||
|
||||
|
||||
def die(msg):
|
||||
print(f"FATAL: {msg}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
|
||||
def gen_secret(n=40):
|
||||
alphabet = string.ascii_letters + string.digits
|
||||
return "".join(secrets.choice(alphabet) for _ in range(n))
|
||||
|
||||
|
||||
def kubectl_get_secret_key(namespace, name, key):
|
||||
"""Returns decoded value, or None if the secret/key doesn't exist."""
|
||||
p = subprocess.run(
|
||||
["kubectl", "-n", namespace, "get", "secret", name, "-o", f"jsonpath={{.data.{key}}}"],
|
||||
capture_output=True, text=True,
|
||||
)
|
||||
if p.returncode != 0 or not p.stdout.strip():
|
||||
return None
|
||||
import base64
|
||||
return base64.b64decode(p.stdout).decode()
|
||||
|
||||
|
||||
def kubectl_create_secret(namespace, name, literals: dict, labels: dict = None):
|
||||
"""Idempotent: create-or-update via dry-run|apply, same pattern used
|
||||
elsewhere in this repo (setup_vault.sh, apply-vault-secrets.sh)."""
|
||||
args = ["kubectl", "-n", namespace, "create", "secret", "generic", name]
|
||||
for k, v in literals.items():
|
||||
args += [f"--from-literal={k}={v}"]
|
||||
args += ["--dry-run=client", "-o", "yaml"]
|
||||
render = subprocess.run(args, capture_output=True, text=True)
|
||||
if render.returncode != 0:
|
||||
die(f"rendering secret {namespace}/{name}: {render.stderr}")
|
||||
apply = subprocess.run(["kubectl", "apply", "-f", "-"], input=render.stdout,
|
||||
capture_output=True, text=True)
|
||||
if apply.returncode != 0:
|
||||
die(f"applying secret {namespace}/{name}: {apply.stderr}")
|
||||
print(f" secret {namespace}/{name}: {apply.stdout.strip()}")
|
||||
if labels:
|
||||
# argocd's `$secret:key` substitution only reads Secrets carrying
|
||||
# app.kubernetes.io/part-of: argocd — without it OIDC login fails with
|
||||
# oauth2 "invalid_client" (empty client_secret sent to the IdP).
|
||||
label_args = ["kubectl", "-n", namespace, "label", "secret", name,
|
||||
"--overwrite"] + [f"{k}={v}" for k, v in labels.items()]
|
||||
subprocess.run(label_args, capture_output=True, text=True)
|
||||
|
||||
|
||||
def get_or_create(list_path, create_path, query, payload, patch_existing=None):
|
||||
status, res = api("GET", f"{list_path}?{query}")
|
||||
if status != 200:
|
||||
die(f"GET {list_path}?{query} -> {status} {res}")
|
||||
results = res.get("results", [])
|
||||
if results:
|
||||
obj = results[0]
|
||||
if patch_existing:
|
||||
status, obj2 = api("PATCH", f"{create_path}{obj['pk']}/", patch_existing)
|
||||
if status not in (200, 201):
|
||||
die(f"PATCH {create_path}{obj['pk']}/ -> {status} {obj2}")
|
||||
return obj2
|
||||
return obj
|
||||
status, obj = api("POST", create_path, payload)
|
||||
if status not in (200, 201):
|
||||
die(f"POST {create_path} -> {status} {obj}")
|
||||
return obj
|
||||
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
print("[1/5] Ensuring custom 'groups' scope mapping exists...")
|
||||
groups_mapping = get_or_create(
|
||||
"/api/v3/propertymappings/provider/scope/",
|
||||
"/api/v3/propertymappings/provider/scope/",
|
||||
"scope_name=groups",
|
||||
{
|
||||
"name": "homelab: groups claim",
|
||||
"scope_name": "groups",
|
||||
# request.user.ak_groups is deprecated in authentik 2026.x (logs a
|
||||
# deprecation warning on every token issue) -> use request.user.groups.
|
||||
"expression": (
|
||||
"return {\"groups\": [group.name for group in request.user.groups.all()]}"
|
||||
),
|
||||
},
|
||||
# Force the expression onto the already-created mapping on re-run.
|
||||
patch_existing={
|
||||
"expression": (
|
||||
"return {\"groups\": [group.name for group in request.user.groups.all()]}"
|
||||
),
|
||||
},
|
||||
)
|
||||
GROUPS_MAPPING_PK = groups_mapping["pk"]
|
||||
|
||||
# MinIO maps OIDC users to a MinIO policy via a "policy" claim
|
||||
# (MINIO_IDENTITY_OPENID_CLAIM_NAME=policy). Emit consoleAdmin (full admin) for
|
||||
# homelab-admins members, readonly for everyone else. Without this claim MinIO
|
||||
# assigns no policy and OIDC users get no access.
|
||||
_POLICY_EXPR = (
|
||||
"return {\"policy\": \"consoleAdmin\" "
|
||||
"if request.user.ak_groups.filter(name=\"homelab-admins\").exists() "
|
||||
"else \"readonly\"}"
|
||||
)
|
||||
policy_mapping = get_or_create(
|
||||
"/api/v3/propertymappings/provider/scope/",
|
||||
"/api/v3/propertymappings/provider/scope/",
|
||||
"scope_name=minio",
|
||||
{
|
||||
"name": "homelab: minio policy claim",
|
||||
"scope_name": "minio",
|
||||
"expression": _POLICY_EXPR,
|
||||
},
|
||||
patch_existing={"expression": _POLICY_EXPR},
|
||||
)
|
||||
POLICY_MAPPING_PK = policy_mapping["pk"]
|
||||
|
||||
# Fetch the standard openid/email/profile mapping pks (shipped by default).
|
||||
status, res = api("GET", "/api/v3/propertymappings/provider/scope/")
|
||||
by_scope = {m["scope_name"]: m["pk"] for m in res["results"]}
|
||||
SCOPE_PKS = [by_scope["openid"], by_scope["email"], by_scope["profile"], GROUPS_MAPPING_PK]
|
||||
|
||||
status, res = api("GET", "/api/v3/flows/instances/?slug=default-provider-authorization-implicit-consent")
|
||||
AUTHORIZATION_FLOW_PK = res["results"][0]["pk"]
|
||||
status, res = api("GET", "/api/v3/flows/instances/?slug=default-provider-invalidation-flow")
|
||||
INVALIDATION_FLOW_PK = res["results"][0]["pk"]
|
||||
status, res = api("GET", "/api/v3/crypto/certificatekeypairs/?has_key=true")
|
||||
SIGNING_KEY_PK = res["results"][0]["pk"]
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
print("[2/5] Ensuring groups homelab-admins / grafana-admins exist...")
|
||||
homelab_admins = get_or_create(
|
||||
"/api/v3/core/groups/", "/api/v3/core/groups/",
|
||||
"name=homelab-admins",
|
||||
{"name": "homelab-admins", "is_superuser": True},
|
||||
)
|
||||
grafana_admins = get_or_create(
|
||||
"/api/v3/core/groups/", "/api/v3/core/groups/",
|
||||
"name=grafana-admins",
|
||||
{"name": "grafana-admins", "is_superuser": False},
|
||||
)
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
print("[3/5] Ensuring user 'rock' exists with admin group membership...")
|
||||
status, res = api("GET", "/api/v3/core/users/?username=rock")
|
||||
rock_password = None
|
||||
if res.get("results"):
|
||||
rock = res["results"][0]
|
||||
status, rock = api("PATCH", f"/api/v3/core/users/{rock['pk']}/", {
|
||||
"groups": [homelab_admins["pk"], grafana_admins["pk"]],
|
||||
"is_active": True,
|
||||
# email is REQUIRED: Grafana's OIDC login reads the email claim from
|
||||
# userinfo; an empty email makes Grafana fall back to a GitHub-style
|
||||
# <userinfo>/emails call, which Authentik 404s -> login fails entirely.
|
||||
"email": "[email protected]",
|
||||
})
|
||||
if status not in (200, 201):
|
||||
die(f"PATCH user rock -> {status} {rock}")
|
||||
print(" rock already exists, group membership synced (password unchanged)")
|
||||
else:
|
||||
rock_password = gen_secret(24)
|
||||
status, rock = api("POST", "/api/v3/core/users/", {
|
||||
"username": "rock",
|
||||
"name": "Rock",
|
||||
"is_active": True,
|
||||
# Required for Grafana OIDC (see PATCH branch above).
|
||||
"email": "[email protected]",
|
||||
"groups": [homelab_admins["pk"], grafana_admins["pk"]],
|
||||
"path": "users",
|
||||
"type": "internal",
|
||||
})
|
||||
if status not in (200, 201):
|
||||
die(f"POST user rock -> {status} {rock}")
|
||||
status, pw_res = api("POST", f"/api/v3/core/users/{rock['pk']}/set_password/",
|
||||
{"password": rock_password})
|
||||
if status not in (200, 204):
|
||||
die(f"set_password for rock -> {status} {pw_res}")
|
||||
kubectl_create_secret("iam", "rock-credentials", {
|
||||
"username": "rock",
|
||||
"password": rock_password,
|
||||
})
|
||||
print(" rock created, credentials stored in iam/rock-credentials")
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
print("[4/5] Ensuring OAuth2 providers + applications for grafana/minio/forgejo/argocd...")
|
||||
|
||||
SERVICES = {
|
||||
"grafana": {
|
||||
"client_secret_source": ("logging", "grafana-oidc", "GF_AUTH_GENERIC_OAUTH_CLIENT_SECRET"),
|
||||
"redirect_uris": ["https://grafana.riotpiao.com/login/generic_oauth"],
|
||||
"launch_url": "https://grafana.riotpiao.com",
|
||||
"display_name": "Grafana",
|
||||
},
|
||||
"minio": {
|
||||
"client_secret_source": ("storage", "minio-oidc", "MINIO_IDENTITY_OPENID_CLIENT_SECRET"),
|
||||
"redirect_uris": ["https://minio.riotpiao.com/oauth_callback"],
|
||||
"launch_url": "https://minio.riotpiao.com",
|
||||
"display_name": "MinIO",
|
||||
},
|
||||
"forgejo": {
|
||||
# No secret exists yet for forgejo - generate + store on first run.
|
||||
"client_secret_source": ("cicd", "forgejo-oidc", "CLIENT_SECRET"),
|
||||
"generate_if_missing": True,
|
||||
"redirect_uris": [
|
||||
"https://forgejo.riotpiao.com/user/oauth2/authentik/callback",
|
||||
"https://forgejo.riotpiao.com/user/oauth2/openidconnect/callback",
|
||||
],
|
||||
"launch_url": "https://forgejo.riotpiao.com",
|
||||
"display_name": "Forgejo",
|
||||
},
|
||||
"argocd": {
|
||||
# oidc-secret uses hyphenated keys (client-id/client-secret) per
|
||||
# argocd-values.yaml's `$oidc-secret:client-id` / `:client-secret` refs.
|
||||
"client_secret_source": ("argocd", "oidc-secret", "client-secret"),
|
||||
"generate_if_missing": True,
|
||||
"extra_secret_literals": {"client-id": "argocd"},
|
||||
# argocd only reads $secret refs from Secrets labelled part-of: argocd.
|
||||
"secret_labels": {"app.kubernetes.io/part-of": "argocd"},
|
||||
"redirect_uris": ["https://argocd.riotpiao.com/auth/callback"],
|
||||
"launch_url": "https://argocd.riotpiao.com",
|
||||
"display_name": "Argo CD",
|
||||
},
|
||||
"homarr": {
|
||||
"client_secret_source": ("dashboard", "homarr-oidc", "client-secret"),
|
||||
"generate_if_missing": True,
|
||||
"extra_secret_literals": {"client-id": "homarr"},
|
||||
"redirect_uris": ["https://homarr.riotpiao.com/api/auth/callback/oidc"],
|
||||
"launch_url": "https://homarr.riotpiao.com",
|
||||
"display_name": "Homarr",
|
||||
},
|
||||
}
|
||||
|
||||
app_pks_for_binding = []
|
||||
|
||||
for name, cfg in SERVICES.items():
|
||||
# MinIO also needs the "policy" claim (via the minio scope mapping) so its
|
||||
# MINIO_IDENTITY_OPENID_CLAIM_NAME=policy maps homelab-admins -> consoleAdmin.
|
||||
provider_mappings = SCOPE_PKS + ([POLICY_MAPPING_PK] if name == "minio" else [])
|
||||
ns, secret_name, key = cfg["client_secret_source"]
|
||||
client_secret = kubectl_get_secret_key(ns, secret_name, key)
|
||||
if client_secret is None:
|
||||
if not cfg.get("generate_if_missing"):
|
||||
print(f" WARNING: {ns}/{secret_name} key {key} not found and "
|
||||
f"generate_if_missing not set for '{name}' - skipping provider/app")
|
||||
continue
|
||||
client_secret = gen_secret(40)
|
||||
literals = {key: client_secret}
|
||||
literals.update(cfg.get("extra_secret_literals", {}))
|
||||
kubectl_create_secret(ns, secret_name, literals,
|
||||
labels=cfg.get("secret_labels"))
|
||||
print(f" {name}: generated new client secret -> {ns}/{secret_name}")
|
||||
else:
|
||||
print(f" {name}: using existing client secret from {ns}/{secret_name}")
|
||||
|
||||
provider = get_or_create(
|
||||
"/api/v3/providers/oauth2/", "/api/v3/providers/oauth2/",
|
||||
f"name={name}",
|
||||
{
|
||||
"name": name,
|
||||
"client_id": name,
|
||||
"client_secret": client_secret,
|
||||
"client_type": "confidential",
|
||||
"authorization_flow": AUTHORIZATION_FLOW_PK,
|
||||
"invalidation_flow": INVALIDATION_FLOW_PK,
|
||||
"signing_key": SIGNING_KEY_PK,
|
||||
"property_mappings": provider_mappings,
|
||||
"sub_mode": "hashed_user_id",
|
||||
"include_claims_in_id_token": True,
|
||||
# authentik 2026.x requires grant_types to be set explicitly; the
|
||||
# API defaults it to [] when omitted, which makes /authorize reject
|
||||
# every login with "Invalid grant_type for provider" ->
|
||||
# invalid_request. authorization_code = the web SSO flow all these
|
||||
# apps use; refresh_token = long-lived sessions (offline_access).
|
||||
"grant_types": ["authorization_code", "refresh_token"],
|
||||
"redirect_uris": [
|
||||
{"matching_mode": "strict", "url": u} for u in cfg["redirect_uris"]
|
||||
],
|
||||
},
|
||||
# Keep the redirect_uris/mappings/grant_types in sync on re-run, but
|
||||
# never touch client_secret again once created (that's the source of
|
||||
# truth in the k8s Secret, and re-sending it here is harmless anyway).
|
||||
patch_existing={
|
||||
"property_mappings": provider_mappings,
|
||||
"grant_types": ["authorization_code", "refresh_token"],
|
||||
"redirect_uris": [
|
||||
{"matching_mode": "strict", "url": u} for u in cfg["redirect_uris"]
|
||||
],
|
||||
},
|
||||
)
|
||||
|
||||
# superuser_full_list=true is REQUIRED on the LIST: the applications list
|
||||
# applies access-policy filtering to the results array (these apps are bound
|
||||
# to homelab-admins, and the bootstrap-token user akadmin is not a member),
|
||||
# so without it the GET returns an empty results list even though the app
|
||||
# exists -> fall through to POST -> 400 "already exists".
|
||||
#
|
||||
# We deliberately do NOT patch_existing here: the application DETAIL endpoint
|
||||
# (PATCH /applications/{pk}/) enforces the same access policy and does NOT
|
||||
# honor superuser_full_list, so PATCH-by-pk returns 404 for akadmin once the
|
||||
# homelab-admins binding exists. That 404 aborted the loop before later
|
||||
# providers got their grant_types. slug/provider/launch_url are set at
|
||||
# creation and are stable (provider is get_or_create'd by name, stable pk),
|
||||
# so find-or-create is sufficient.
|
||||
application = get_or_create(
|
||||
"/api/v3/core/applications/", "/api/v3/core/applications/",
|
||||
f"slug={name}&superuser_full_list=true",
|
||||
{
|
||||
"name": cfg["display_name"],
|
||||
"slug": name,
|
||||
"provider": provider["pk"],
|
||||
"meta_launch_url": cfg["launch_url"],
|
||||
},
|
||||
)
|
||||
app_pks_for_binding.append((name, application["pk"]))
|
||||
print(f" {name}: provider pk={provider['pk']} application pk={application['pk']}")
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
print("[5/5] Binding homelab-admins to every application (guaranteed access for rock)...")
|
||||
for name, app_pk in app_pks_for_binding:
|
||||
get_or_create(
|
||||
"/api/v3/policies/bindings/", "/api/v3/policies/bindings/",
|
||||
f"target={app_pk}&group={homelab_admins['pk']}",
|
||||
{
|
||||
"target": app_pk,
|
||||
"group": homelab_admins["pk"],
|
||||
"order": 0,
|
||||
"enabled": True,
|
||||
},
|
||||
)
|
||||
print(f" {name}: homelab-admins bound")
|
||||
|
||||
print("\nDone. Summary:")
|
||||
print(" groups: homelab-admins (superuser), grafana-admins")
|
||||
print(" user: rock -> homelab-admins + grafana-admins")
|
||||
print(f" apps: {', '.join(n for n, _ in app_pks_for_binding)}")
|
||||
if rock_password:
|
||||
print(" NOTE: rock's password was generated this run - see")
|
||||
print(" kubectl -n iam get secret rock-credentials -o jsonpath='{.data.password}' | base64 -d")
|
||||
@@ -44,10 +44,6 @@ grafana.ini:
|
||||
server:
|
||||
root_url: https://grafana.riotpiao.com
|
||||
|
||||
# Allow embedding dashboards in iframes (NextJS integration)
|
||||
security:
|
||||
allow_embedding: true
|
||||
|
||||
# No anonymous read access — every user must log in via Authentik SSO.
|
||||
auth.anonymous:
|
||||
enabled: false
|
||||
@@ -68,14 +64,15 @@ grafana.ini:
|
||||
# doesn't return localhost redirects in its token responses.
|
||||
#
|
||||
# role_attribute_path: JMESPath expression evaluated against the userinfo
|
||||
# response. akadmin gets GrafanaAdmin (server admin, can impersonate);
|
||||
# homelab-admins members get Admin (org admin); everyone else Viewer.
|
||||
# response. Members of the 'grafana-admins' Authentik group get Admin role;
|
||||
# everyone else gets Viewer. The group name must match exactly what Authentik
|
||||
# sends in the 'groups' claim.
|
||||
auth.generic_oauth:
|
||||
enabled: true
|
||||
name: Authentik
|
||||
allow_sign_up: true
|
||||
client_id: grafana
|
||||
scopes: openid email profile groups
|
||||
scopes: openid email profile
|
||||
auth_url: https://authentik.riotpiao.com/application/o/authorize/
|
||||
token_url: https://authentik.riotpiao.com/application/o/token/
|
||||
api_url: https://authentik.riotpiao.com/application/o/userinfo/
|
||||
@@ -86,8 +83,7 @@ grafana.ini:
|
||||
email_attribute_path: email
|
||||
login_attribute_path: preferred_username
|
||||
name_attribute_path: name
|
||||
role_attribute_path: "preferred_username == 'akadmin' && 'GrafanaAdmin' || contains(groups[*], 'homelab-admins') && 'Admin' || 'Viewer'"
|
||||
allow_assign_grafana_admin: true
|
||||
role_attribute_path: "contains(groups[*], 'grafana-admins') && 'Admin' || 'Viewer'"
|
||||
use_pkce: false
|
||||
use_refresh_token: false
|
||||
skip_org_role_sync: false
|
||||
|
||||
@@ -4,13 +4,11 @@ namespace: longhorn-system
|
||||
resources:
|
||||
- longhorn-storageclass.yaml
|
||||
- longhorn-cnpg-storageclass.yaml # CNPG-specific with postgres UID/GID
|
||||
- longhorn-paperless-storageclass.yaml # single-replica, cp-3 USB HDD only
|
||||
- longhorn-servicemonitor.yaml
|
||||
- longhorn-taint-toleration.yaml
|
||||
- longhorn-nodes.yaml
|
||||
- expand-replicas-job.yaml
|
||||
- patch-csi-tolerations-job.yaml
|
||||
- longhorn-add-disks-job.yaml # Add extra disks to talos-cp-2
|
||||
# Longhorn deployed via bootstrap script or Helm.
|
||||
# These manifests configure it: unified StorageClass (default, 3 replicas),
|
||||
# Prometheus ServiceMonitor, taint toleration for control-plane nodes, explicit
|
||||
|
||||
@@ -1,130 +0,0 @@
|
||||
# PostSync hook Job that adds extra disks to Longhorn nodes.
|
||||
# talos-cp-2 has 4 extra disks mounted at /var/lib/longhorn-disk{1,2,3,4}
|
||||
# that are NOT auto-discovered by Longhorn.
|
||||
#
|
||||
# talos-cp-3 additionally gets a tagged disk for paperless-ngx media, backed by
|
||||
# the 4TB USB HDD (/dev/sdg) — tagged "paperless-media" so only the dedicated
|
||||
# longhorn-paperless-media StorageClass (diskSelector match) can place replicas
|
||||
# there, keeping it out of the default 3-replica pool. This patch is inert
|
||||
# until the Terraform machine-config change mounts the disk at
|
||||
# /var/lib/longhorn-paperless-media (pending — see terraform.tfvars.local,
|
||||
# not present in this checkout); Longhorn just reports the disk not-ready
|
||||
# until the path exists, no harm in applying it early.
|
||||
apiVersion: batch/v1
|
||||
kind: Job
|
||||
metadata:
|
||||
name: longhorn-add-disks
|
||||
namespace: longhorn-system
|
||||
annotations:
|
||||
argocd.argoproj.io/hook: PostSync
|
||||
argocd.argoproj.io/hook-delete-policy: BeforeHookCreation
|
||||
spec:
|
||||
backoffLimit: 3
|
||||
template:
|
||||
metadata:
|
||||
name: longhorn-add-disks
|
||||
spec:
|
||||
restartPolicy: Never
|
||||
serviceAccountName: longhorn-expand-replicas
|
||||
containers:
|
||||
- name: add-disks
|
||||
image: bitnami/kubectl:latest
|
||||
command:
|
||||
- /bin/bash
|
||||
- -c
|
||||
- |
|
||||
set -euo pipefail
|
||||
|
||||
echo "Adding extra disks to Longhorn nodes..."
|
||||
|
||||
# talos-cp-2 has 4 extra disks:
|
||||
# - /var/lib/longhorn-disk1 (sdb1, 600GB)
|
||||
# - /var/lib/longhorn-disk2 (sdc1, 500GB)
|
||||
# - /var/lib/longhorn-disk3 (sdd1, 160GB)
|
||||
# - /var/lib/longhorn-disk4 (sde1, 2TB)
|
||||
|
||||
echo "Checking talos-cp-2..."
|
||||
CURRENT_DISKS=$(kubectl -n longhorn-system get nodes.longhorn.io talos-cp-2 -o json | jq -r '.spec.disks | keys | length')
|
||||
echo " Current disk count: $CURRENT_DISKS"
|
||||
|
||||
if [ "$CURRENT_DISKS" -lt 5 ]; then
|
||||
echo " Adding extra disks to talos-cp-2..."
|
||||
kubectl -n longhorn-system patch nodes.longhorn.io talos-cp-2 --type merge -p '{
|
||||
"spec": {
|
||||
"disks": {
|
||||
"disk1": {
|
||||
"allowScheduling": true,
|
||||
"diskType": "filesystem",
|
||||
"evictionRequested": false,
|
||||
"path": "/var/lib/longhorn-disk1",
|
||||
"storageReserved": 10737418240,
|
||||
"tags": []
|
||||
},
|
||||
"disk2": {
|
||||
"allowScheduling": true,
|
||||
"diskType": "filesystem",
|
||||
"evictionRequested": false,
|
||||
"path": "/var/lib/longhorn-disk2",
|
||||
"storageReserved": 10737418240,
|
||||
"tags": []
|
||||
},
|
||||
"disk3": {
|
||||
"allowScheduling": true,
|
||||
"diskType": "filesystem",
|
||||
"evictionRequested": false,
|
||||
"path": "/var/lib/longhorn-disk3",
|
||||
"storageReserved": 5368709120,
|
||||
"tags": []
|
||||
},
|
||||
"disk4": {
|
||||
"allowScheduling": true,
|
||||
"diskType": "filesystem",
|
||||
"evictionRequested": false,
|
||||
"path": "/var/lib/longhorn-disk4",
|
||||
"storageReserved": 21474836480,
|
||||
"tags": []
|
||||
}
|
||||
}
|
||||
}
|
||||
}'
|
||||
echo " ✓ Disks added to talos-cp-2"
|
||||
else
|
||||
echo " ✓ talos-cp-2 already has $CURRENT_DISKS disks configured"
|
||||
fi
|
||||
|
||||
echo "Checking talos-cp-3..."
|
||||
CP3_DISKS=$(kubectl -n longhorn-system get nodes.longhorn.io talos-cp-3 -o json | jq -r '.spec.disks | keys | length')
|
||||
echo " Current disk count: $CP3_DISKS"
|
||||
|
||||
if [ "$CP3_DISKS" -lt 2 ]; then
|
||||
echo " Adding paperless-media disk to talos-cp-3..."
|
||||
kubectl -n longhorn-system patch nodes.longhorn.io talos-cp-3 --type merge -p '{
|
||||
"spec": {
|
||||
"disks": {
|
||||
"paperless-media": {
|
||||
"allowScheduling": true,
|
||||
"diskType": "filesystem",
|
||||
"evictionRequested": false,
|
||||
"path": "/var/lib/longhorn-paperless-media",
|
||||
"storageReserved": 0,
|
||||
"tags": ["paperless-media"]
|
||||
}
|
||||
}
|
||||
}
|
||||
}'
|
||||
echo " ✓ paperless-media disk added to talos-cp-3"
|
||||
else
|
||||
echo " ✓ talos-cp-3 already has $CP3_DISKS disks configured"
|
||||
fi
|
||||
|
||||
echo
|
||||
echo "Waiting for disks to be ready..."
|
||||
sleep 10
|
||||
|
||||
echo
|
||||
echo "Final storage summary:"
|
||||
kubectl -n longhorn-system get nodes.longhorn.io -o json | \
|
||||
jq -r '.items[] | "\(.metadata.name): \(.status.diskStatus | to_entries | map("\(.key): \(.value.storageAvailable/1073741824 | floor)GB avail") | join(", "))"'
|
||||
|
||||
echo
|
||||
echo "Done."
|
||||
@@ -1,32 +0,0 @@
|
||||
# Dedicated StorageClass for paperless-ngx document media, backed by the 4TB
|
||||
# USB HDD on talos-cp-3 (see longhorn-add-disks-job.yaml). Single disk, single
|
||||
# node — no Longhorn replica is possible, so numberOfReplicas is 1 by
|
||||
# necessity, not choice. diskSelector pins placement to the tagged disk only,
|
||||
# so a volume from this class never lands on cp-3's regular (already
|
||||
# DiskPressure) default pool. reclaimPolicy is Retain, not Delete: a PVC
|
||||
# accident here has no replica to fall back on, so an accidental delete must
|
||||
# not also take the underlying volume with it.
|
||||
apiVersion: storage.k8s.io/v1
|
||||
kind: StorageClass
|
||||
metadata:
|
||||
name: longhorn-paperless-media
|
||||
annotations:
|
||||
# StorageClass.parameters is immutable - any future edit here needs
|
||||
# delete+recreate, not patch. Same fix as longhorn-cnpg-storageclass.yaml;
|
||||
# safe since a StorageClass is only consulted at provisioning time.
|
||||
argocd.argoproj.io/sync-options: Replace=true,Force=true
|
||||
provisioner: driver.longhorn.io
|
||||
allowVolumeExpansion: true
|
||||
reclaimPolicy: Retain
|
||||
volumeBindingMode: WaitForFirstConsumer
|
||||
parameters:
|
||||
numberOfReplicas: "1"
|
||||
# diskSelector alone is sufficient: only cp-3 has a disk tagged
|
||||
# paperless-media, so placement is already pinned. A nodeSelector value
|
||||
# here must be a Longhorn *node* tag (set via nodes.longhorn.io spec.tags),
|
||||
# not a Kubernetes hostname - "talos-cp-3" was never a node tag, which made
|
||||
# every provision attempt fail with "specified node tag talos-cp-3 does
|
||||
# not exist".
|
||||
diskSelector: "paperless-media"
|
||||
staleReplicaTimeout: "30"
|
||||
fsType: "ext4"
|
||||
@@ -1,14 +1,8 @@
|
||||
apiVersion: kustomize.config.k8s.io/v1beta1
|
||||
kind: Kustomization
|
||||
# NOTE: no top-level `namespace:` transformer (see iam/kustomization.yaml for
|
||||
# the same fix) - minio-provision-paperless-job.yaml's RoleBinding
|
||||
# deliberately targets namespace paperless (least-privilege access for the
|
||||
# provisioner ServiceAccount to write Secrets there); a namespace transformer
|
||||
# would silently rewrite it back to storage, breaking the RBAC. Every resource
|
||||
# here already sets its own explicit metadata.namespace.
|
||||
namespace: storage
|
||||
resources:
|
||||
- minio-tenant.yaml
|
||||
- minio-provision-paperless-job.yaml
|
||||
# The operator creates the minio S3/console/headless Services and the
|
||||
# declarative bucket + user from the Tenant spec — no hand-rolled Service or
|
||||
# Bucket/User CRs (those kinds don't exist in the operator CRD set).
|
||||
|
||||
@@ -1,120 +0,0 @@
|
||||
# PostSync hook Job: creates a MinIO IAM user + policy scoped to only the
|
||||
# `paperless` bucket (least-privilege — reuses root creds nowhere else in the
|
||||
# cluster), then writes the generated access/secret key into a Secret in the
|
||||
# `paperless` namespace for the nightly backup CronJob to consume.
|
||||
#
|
||||
# Idempotent: re-running never rotates existing credentials — if
|
||||
# paperless-minio-creds already exists in ns paperless, the script reuses the
|
||||
# access key it already wrote and just re-asserts the policy/user exist.
|
||||
apiVersion: v1
|
||||
kind: ServiceAccount
|
||||
metadata:
|
||||
name: minio-paperless-provisioner
|
||||
namespace: storage
|
||||
---
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: ClusterRole
|
||||
metadata:
|
||||
name: minio-paperless-provisioner
|
||||
rules:
|
||||
- apiGroups: [""]
|
||||
resources: ["secrets"]
|
||||
verbs: ["get", "list", "create", "update", "patch"]
|
||||
---
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: RoleBinding
|
||||
metadata:
|
||||
name: minio-paperless-provisioner
|
||||
namespace: paperless
|
||||
subjects:
|
||||
- kind: ServiceAccount
|
||||
name: minio-paperless-provisioner
|
||||
namespace: storage
|
||||
roleRef:
|
||||
kind: ClusterRole
|
||||
name: minio-paperless-provisioner
|
||||
apiGroup: rbac.authorization.k8s.io
|
||||
---
|
||||
apiVersion: batch/v1
|
||||
kind: Job
|
||||
metadata:
|
||||
name: minio-provision-paperless
|
||||
namespace: storage
|
||||
annotations:
|
||||
argocd.argoproj.io/hook: PostSync
|
||||
argocd.argoproj.io/hook-delete-policy: BeforeHookCreation
|
||||
spec:
|
||||
ttlSecondsAfterFinished: 600
|
||||
backoffLimit: 3
|
||||
template:
|
||||
spec:
|
||||
serviceAccountName: minio-paperless-provisioner
|
||||
restartPolicy: Never
|
||||
initContainers:
|
||||
- name: kubectl-copy
|
||||
image: bitnami/kubectl:latest
|
||||
command: ["sh", "-c", "cp $(which kubectl) /shared/kubectl"]
|
||||
volumeMounts:
|
||||
- name: shared
|
||||
mountPath: /shared
|
||||
containers:
|
||||
- name: provision
|
||||
image: minio/mc:latest
|
||||
volumeMounts:
|
||||
- name: shared
|
||||
mountPath: /shared
|
||||
- name: minio-creds
|
||||
mountPath: /minio-creds
|
||||
readOnly: true
|
||||
command:
|
||||
- /bin/sh
|
||||
- -c
|
||||
- |
|
||||
set -e
|
||||
export PATH="/shared:$PATH"
|
||||
|
||||
# minio-creds ships as a shell-sourceable config.env blob
|
||||
# (`export MINIO_ROOT_USER=... / MINIO_ROOT_PASSWORD=...`), not
|
||||
# discrete keys — source it directly rather than re-parsing.
|
||||
. /minio-creds/config.env
|
||||
|
||||
mc alias set m http://minio-cluster-hl.storage.svc.cluster.local:9000 \
|
||||
"$MINIO_ROOT_USER" "$MINIO_ROOT_PASSWORD"
|
||||
|
||||
echo "Checking for existing paperless-minio-creds secret..."
|
||||
if kubectl -n paperless get secret paperless-minio-creds >/dev/null 2>&1; then
|
||||
ACCESS_KEY=$(kubectl -n paperless get secret paperless-minio-creds -o jsonpath='{.data.ACCESS_KEY}' | base64 -d)
|
||||
SECRET_KEY=$(kubectl -n paperless get secret paperless-minio-creds -o jsonpath='{.data.SECRET_KEY}' | base64 -d)
|
||||
echo " reusing existing credentials"
|
||||
else
|
||||
ACCESS_KEY="paperless"
|
||||
SECRET_KEY=$(head -c 32 /dev/urandom | base64 | tr -d '/+=' | head -c 40)
|
||||
echo " generated new credentials"
|
||||
fi
|
||||
|
||||
echo "Ensuring MinIO user 'paperless' exists..."
|
||||
if ! mc admin user info m "$ACCESS_KEY" >/dev/null 2>&1; then
|
||||
mc admin user add m "$ACCESS_KEY" "$SECRET_KEY"
|
||||
fi
|
||||
|
||||
echo "Writing scoped policy (paperless bucket only)..."
|
||||
printf '%s' '{"Version":"2012-10-17","Statement":[{"Effect":"Allow","Action":["s3:ListBucket"],"Resource":["arn:aws:s3:::paperless"]},{"Effect":"Allow","Action":["s3:GetObject","s3:PutObject","s3:DeleteObject"],"Resource":["arn:aws:s3:::paperless/*"]}]}' > /tmp/paperless-rw-policy.json
|
||||
mc admin policy create m paperless-rw /tmp/paperless-rw-policy.json || \
|
||||
mc admin policy update m paperless-rw /tmp/paperless-rw-policy.json
|
||||
mc admin policy attach m paperless-rw --user "$ACCESS_KEY"
|
||||
|
||||
echo "Writing paperless-minio-creds secret (ns paperless)..."
|
||||
kubectl -n paperless create secret generic paperless-minio-creds \
|
||||
--from-literal=ACCESS_KEY="$ACCESS_KEY" \
|
||||
--from-literal=SECRET_KEY="$SECRET_KEY" \
|
||||
--from-literal=BUCKET=paperless \
|
||||
--from-literal=ENDPOINT=http://minio-cluster-hl.storage.svc.cluster.local:9000 \
|
||||
--dry-run=client -o yaml | kubectl apply -f -
|
||||
|
||||
echo "Done."
|
||||
volumes:
|
||||
- name: shared
|
||||
emptyDir: {}
|
||||
- name: minio-creds
|
||||
secret:
|
||||
secretName: minio-creds
|
||||
@@ -76,7 +76,6 @@ spec:
|
||||
- name: loki-ruler
|
||||
- name: loki-admin
|
||||
- name: vault
|
||||
- name: paperless
|
||||
|
||||
# Metrics are exposed at /minio/v2/metrics; scrape via a hand-rolled
|
||||
# ServiceMonitor in the monitoring stack rather than operator auto-wiring
|
||||
@@ -84,25 +83,15 @@ spec:
|
||||
# 'default' and fail the reconcile).
|
||||
|
||||
# Public hostnames the tenant serves (S3 + console via the cluster ingress).
|
||||
# Must match actual ingress backends: the "minio" Ingress (host
|
||||
# minio.riotpiao.com) routes to minio-cluster-console:9090, and "minio-api"
|
||||
# (host minio-api.riotpiao.com) routes to minio:9000. Previously these were
|
||||
# swapped, so MinIO's own console-domain validation rejected the OIDC
|
||||
# callback arriving on minio.riotpiao.com's Host header.
|
||||
features:
|
||||
domains:
|
||||
minio:
|
||||
- https://minio-api.riotpiao.com
|
||||
console: https://minio.riotpiao.com
|
||||
- https://minio.riotpiao.com
|
||||
console: https://minio-console.riotpiao.com
|
||||
|
||||
# ── OIDC via Authentik (server-side env, valid in v2 schema) ────────────────
|
||||
# Use in-cluster URL for config fetch (pod→authentik); browser redirects use
|
||||
# public URLs embedded in the OIDC metadata response (issuer stays public).
|
||||
env:
|
||||
- name: MINIO_IDENTITY_OPENID_CONFIG_URL
|
||||
# Must use external URL — well-known response contains external issuer/jwks_uri.
|
||||
# MinIO validates issuer in JWT matches well-known issuer. Internal URL = mismatch.
|
||||
# Hairpins through ingress-nginx but stays in-cluster.
|
||||
value: "https://authentik.riotpiao.com/application/o/minio/.well-known/openid-configuration"
|
||||
- name: MINIO_IDENTITY_OPENID_CLIENT_ID
|
||||
value: "minio"
|
||||
|
||||
@@ -1,160 +0,0 @@
|
||||
apiVersion: monitoring.coreos.com/v1
|
||||
kind: PrometheusRule
|
||||
metadata:
|
||||
name: api-gateway-alerts
|
||||
namespace: monitoring
|
||||
labels:
|
||||
release: prometheus
|
||||
spec:
|
||||
groups:
|
||||
# ================================================================
|
||||
# SLA Targets (based on canary traffic baselines):
|
||||
#
|
||||
# Availability: 99.9% (43.8 min downtime/month)
|
||||
# LLM Chat: p95 < 1s (qwen), p95 < 2s (reasoning), p95 < 5s (ornith)
|
||||
# Embeddings: p95 < 500ms
|
||||
# Rerank: p95 < 500ms
|
||||
# Models list: p95 < 300ms
|
||||
# Error rate: < 1% (5xx), < 5% (4xx excluding auth)
|
||||
#
|
||||
# Baselines from 200-request canary run:
|
||||
# qwen p99=609ms, reasoning p99=328ms, embeddings p99=287ms,
|
||||
# rerank p99=218ms, models p99=277ms
|
||||
# SLA set at ~2x p99 for headroom.
|
||||
# ================================================================
|
||||
|
||||
- name: api-gateway.availability
|
||||
rules:
|
||||
# Gateway pods not ready
|
||||
- alert: APIGatewayDown
|
||||
expr: sum(kube_pod_status_ready{namespace="api",condition="true"}) == 0
|
||||
for: 1m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "API Gateway has zero ready pods"
|
||||
|
||||
# Gateway pod count below desired
|
||||
- alert: APIGatewayDegraded
|
||||
expr: |
|
||||
sum(kube_pod_status_ready{namespace="api",condition="true"})
|
||||
< kube_deployment_spec_replicas{namespace="api",deployment="api-gateway"}
|
||||
for: 5m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "API Gateway {{ $value }} ready pods below desired replica count"
|
||||
|
||||
# Blackbox probe down
|
||||
- alert: APIGatewayProbeDown
|
||||
expr: probe_success{instance=~".*api.riotpiao.com.*"} == 0
|
||||
for: 2m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "API Gateway probe failed: {{ $labels.instance }}"
|
||||
|
||||
# LLM serving pods not ready
|
||||
- alert: LLMServingDown
|
||||
expr: sum(kube_pod_status_ready{namespace="llm-serving",condition="true"}) == 0
|
||||
for: 2m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "All LLM serving pods down"
|
||||
|
||||
# Individual predictor down
|
||||
- alert: LLMPredictorDown
|
||||
expr: |
|
||||
kube_deployment_status_replicas_ready{namespace="llm-serving"}
|
||||
< kube_deployment_spec_replicas{namespace="llm-serving"}
|
||||
for: 5m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "{{ $labels.deployment }} has {{ $value }} ready (below desired)"
|
||||
|
||||
- name: api-gateway.latency
|
||||
# SLA: latency thresholds at ~2x measured p99
|
||||
rules:
|
||||
# Ingress-level latency (all requests through nginx)
|
||||
- alert: APIGatewayLatencyHigh
|
||||
expr: |
|
||||
histogram_quantile(0.95,
|
||||
sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress="api"}[5m])) by (le)
|
||||
) > 2
|
||||
for: 5m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "API Gateway p95 latency {{ $value | printf \"%.1f\" }}s (SLA: <2s)"
|
||||
|
||||
# Extreme latency (p99 > 5s)
|
||||
- alert: APIGatewayLatencyCritical
|
||||
expr: |
|
||||
histogram_quantile(0.99,
|
||||
sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress="api"}[5m])) by (le)
|
||||
) > 5
|
||||
for: 5m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "API Gateway p99 latency {{ $value | printf \"%.1f\" }}s (SLA: <5s)"
|
||||
|
||||
- name: api-gateway.errors
|
||||
rules:
|
||||
# 5xx error rate > 1%
|
||||
- alert: APIGateway5xxErrorRate
|
||||
expr: |
|
||||
sum(rate(nginx_ingress_controller_requests{ingress="api",status=~"5.."}[5m]))
|
||||
/ sum(rate(nginx_ingress_controller_requests{ingress="api"}[5m]))
|
||||
> 0.01
|
||||
for: 5m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "API Gateway 5xx rate {{ $value | humanizePercentage }} (SLA: <1%)"
|
||||
|
||||
# Total error rate > 10% (including 4xx)
|
||||
- alert: APIGatewayHighErrorRate
|
||||
expr: |
|
||||
sum(rate(nginx_ingress_controller_requests{ingress="api",status=~"[45].."}[5m]))
|
||||
/ sum(rate(nginx_ingress_controller_requests{ingress="api"}[5m]))
|
||||
> 0.10
|
||||
for: 10m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "API Gateway total error rate {{ $value | humanizePercentage }} (SLA: <10%)"
|
||||
|
||||
- name: api-gateway.resources
|
||||
rules:
|
||||
# Gateway pod restart
|
||||
- alert: APIGatewayRestarted
|
||||
expr: increase(kube_pod_container_status_restarts_total{namespace="api"}[15m]) > 0
|
||||
for: 0m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "API Gateway pod {{ $labels.pod }} restarted"
|
||||
|
||||
# LLM predictor restart
|
||||
- alert: LLMPredictorRestarted
|
||||
expr: increase(kube_pod_container_status_restarts_total{namespace="llm-serving"}[15m]) > 0
|
||||
for: 0m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "LLM predictor {{ $labels.pod }} restarted"
|
||||
|
||||
# Gateway high memory (>80% of limit)
|
||||
- alert: APIGatewayHighMemory
|
||||
expr: |
|
||||
sum(container_memory_working_set_bytes{namespace="api",container="gateway"}) by (pod)
|
||||
/ sum(kube_pod_container_resource_limits{namespace="api",container="gateway",resource="memory"}) by (pod)
|
||||
> 0.8
|
||||
for: 10m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "Gateway pod {{ $labels.pod }} memory at {{ $value | humanizePercentage }} of limit"
|
||||
@@ -1,177 +0,0 @@
|
||||
apiVersion: monitoring.coreos.com/v1
|
||||
kind: PrometheusRule
|
||||
metadata:
|
||||
name: cluster-alerts
|
||||
namespace: monitoring
|
||||
labels:
|
||||
release: prometheus
|
||||
spec:
|
||||
groups:
|
||||
- name: cluster.availability
|
||||
rules:
|
||||
# Node down
|
||||
- alert: NodeNotReady
|
||||
expr: kube_node_status_condition{condition="Ready",status="true"} == 0
|
||||
for: 2m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "Node {{ $labels.node }} is NotReady"
|
||||
|
||||
# Pod stuck pending (scheduling failure)
|
||||
- alert: PodStuckPending
|
||||
expr: sum(kube_pod_status_phase{phase="Pending"}) > 0
|
||||
for: 10m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "{{ $value }} pod(s) stuck in Pending state for >10m"
|
||||
|
||||
# CrashLoopBackOff
|
||||
- alert: PodCrashLooping
|
||||
expr: sum(kube_pod_container_status_waiting_reason{reason="CrashLoopBackOff"}) by (namespace, pod) > 0
|
||||
for: 5m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "{{ $labels.namespace }}/{{ $labels.pod }} in CrashLoopBackOff"
|
||||
|
||||
# OOMKilled spike
|
||||
- alert: OOMKilledSpike
|
||||
expr: sum(increase(kube_pod_container_status_last_terminated_reason{reason="OOMKilled"}[1h])) > 3
|
||||
for: 0m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "{{ $value }} OOMKilled events in last hour"
|
||||
|
||||
# Deployment replicas unavailable
|
||||
- alert: DeploymentReplicasUnavailable
|
||||
expr: kube_deployment_status_replicas_unavailable > 0
|
||||
for: 10m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "{{ $labels.namespace }}/{{ $labels.deployment }} has {{ $value }} unavailable replicas"
|
||||
|
||||
- name: cluster.jobs
|
||||
rules:
|
||||
# Job failed
|
||||
- alert: JobFailed
|
||||
expr: kube_job_status_failed > 0
|
||||
for: 5m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "Job {{ $labels.namespace }}/{{ $labels.job_name }} failed"
|
||||
|
||||
# Job stuck running >2h
|
||||
- alert: JobStuckRunning
|
||||
expr: |
|
||||
kube_job_status_active == 1
|
||||
and on(job_name,namespace)
|
||||
(time() - kube_job_status_start_time) > 7200
|
||||
for: 0m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "Job {{ $labels.namespace }}/{{ $labels.job_name }} running >2h"
|
||||
|
||||
# CronJob missed schedule
|
||||
- alert: CronJobMissedSchedule
|
||||
expr: |
|
||||
(time() - kube_cronjob_status_last_schedule_time) > 2 * (kube_cronjob_spec_next_schedule_time - kube_cronjob_status_last_schedule_time)
|
||||
for: 10m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "CronJob {{ $labels.namespace }}/{{ $labels.cronjob }} missed schedule"
|
||||
|
||||
- name: cluster.resources
|
||||
rules:
|
||||
# Node CPU >90% sustained
|
||||
- alert: NodeHighCPU
|
||||
expr: (1 - avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) by (instance)) * 100 > 90
|
||||
for: 15m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "Node {{ $labels.instance }} CPU at {{ $value | printf \"%.0f\" }}%"
|
||||
|
||||
# Node memory >90% sustained
|
||||
- alert: NodeHighMemory
|
||||
expr: (1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100 > 90
|
||||
for: 15m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "Node {{ $labels.instance }} memory at {{ $value | printf \"%.0f\" }}%"
|
||||
|
||||
# Node disk >85%
|
||||
- alert: NodeDiskFull
|
||||
expr: (1 - node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100 > 85
|
||||
for: 5m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "Node {{ $labels.instance }} disk at {{ $value | printf \"%.0f\" }}%"
|
||||
|
||||
# Container restart storm (>5 restarts in 15m)
|
||||
- alert: ContainerRestartStorm
|
||||
expr: sum(increase(kube_pod_container_status_restarts_total[15m])) by (namespace, pod) > 5
|
||||
for: 0m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "{{ $labels.namespace }}/{{ $labels.pod }} restarted {{ $value | printf \"%.0f\" }} times in 15m"
|
||||
|
||||
- name: cluster.storage
|
||||
rules:
|
||||
# Longhorn drive offline
|
||||
- alert: LonghornDriveOffline
|
||||
expr: longhorn_disk_health != 1
|
||||
for: 5m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "Longhorn disk {{ $labels.node }} unhealthy"
|
||||
|
||||
- name: cluster.dns
|
||||
rules:
|
||||
# CoreDNS errors spike
|
||||
- alert: CoreDNSErrorSpike
|
||||
expr: sum(rate(coredns_dns_responses_total{rcode=~"SERVFAIL"}[5m])) > 0.5
|
||||
for: 5m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "CoreDNS SERVFAIL rate {{ $value | printf \"%.2f\" }}/s"
|
||||
|
||||
- name: cluster.probes
|
||||
rules:
|
||||
# Any blackbox probe down
|
||||
- alert: ServiceProbeDown
|
||||
expr: probe_success == 0
|
||||
for: 3m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "Probe failed: {{ $labels.instance }}"
|
||||
|
||||
# Probe latency >2s
|
||||
- alert: ServiceProbeSlow
|
||||
expr: probe_duration_seconds > 2
|
||||
for: 5m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "Probe slow ({{ $value | printf \"%.1f\" }}s): {{ $labels.instance }}"
|
||||
|
||||
# Certificate expiry <14 days
|
||||
- alert: CertificateExpiringSoon
|
||||
expr: (certmanager_certificate_expiration_timestamp_seconds - time()) / 86400 < 14
|
||||
for: 0m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "Certificate {{ $labels.name }} expires in {{ $value | printf \"%.0f\" }} days"
|
||||
@@ -61,10 +61,6 @@ serviceMonitor:
|
||||
url: https://argocd.riotpiao.com/healthz
|
||||
- name: longhorn
|
||||
url: https://longhorn.riotpiao.com/
|
||||
- name: api-gateway
|
||||
url: https://api.riotpiao.com/healthz
|
||||
- name: api-gateway-models
|
||||
url: https://api.riotpiao.com/v1/models
|
||||
|
||||
prometheusRule:
|
||||
enabled: true
|
||||
|
||||
@@ -1,36 +0,0 @@
|
||||
apiVersion: v1
|
||||
data:
|
||||
api-gateway.json: '{"title":"API Gateway","uid":"api-gateway","schemaVersion":39,"timezone":"browser","time":{"from":"now-6h","to":"now"},"refresh":"30s","tags":["api","gateway","llm"],"panels":[{"id":1,"title":"Gateway
|
||||
Health","type":"row","collapsed":false,"gridPos":{"h":1,"w":24,"x":0,"y":0},"panels":[{"id":2,"title":"Gateway
|
||||
Pods Ready","type":"stat","gridPos":{"h":4,"w":4,"x":0,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"red"},{"value":3,"color":"green"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"sum(kube_pod_status_ready{namespace=\"api\",condition=\"true\"})"}]},{"id":3,"title":"Probe:
|
||||
healthz","type":"stat","gridPos":{"h":4,"w":4,"x":4,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","mappings":[{"type":"value","options":{"0":{"text":"DOWN","color":"red"},"1":{"text":"UP","color":"green"}}}]}},"targets":[{"expr":"probe_success{instance=~\".*api.riotpiao.com/healthz\"}"}]},{"id":4,"title":"Probe
|
||||
Latency","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"s"}},"targets":[{"expr":"probe_duration_seconds{instance=~\".*api.riotpiao.com.*\"}","legendFormat":"{{instance}}"}]}]},{"id":10,"title":"Ingress
|
||||
Traffic (nginx)","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":1},"panels":[{"id":11,"title":"Request
|
||||
Rate by Status","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(nginx_ingress_controller_requests{ingress=\"api\"}[5m]))
|
||||
by (status)","legendFormat":"{{status}}"}]},{"id":12,"title":"Error Rate %","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"percent"}},"targets":[{"expr":"sum(rate(nginx_ingress_controller_requests{ingress=\"api\",status=~\"5..\"}[5m]))
|
||||
/ sum(rate(nginx_ingress_controller_requests{ingress=\"api\"}[5m])) * 100","legendFormat":"5xx"},{"expr":"sum(rate(nginx_ingress_controller_requests{ingress=\"api\",status=~\"4..\"}[5m]))
|
||||
/ sum(rate(nginx_ingress_controller_requests{ingress=\"api\"}[5m])) * 100","legendFormat":"4xx"}]},{"id":13,"title":"Latency
|
||||
p50/p95/p99","type":"timeseries","gridPos":{"h":8,"w":8,"x":16,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"s"}},"targets":[{"expr":"histogram_quantile(0.50,
|
||||
sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress=\"api\"}[5m]))
|
||||
by (le))","legendFormat":"p50"},{"expr":"histogram_quantile(0.95, sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress=\"api\"}[5m]))
|
||||
by (le))","legendFormat":"p95"},{"expr":"histogram_quantile(0.99, sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress=\"api\"}[5m]))
|
||||
by (le))","legendFormat":"p99"}]}]},{"id":20,"title":"LLM Serving","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":2},"panels":[{"id":21,"title":"LLM
|
||||
Pods Ready","type":"stat","gridPos":{"h":4,"w":4,"x":0,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(kube_pod_status_ready{namespace=\"llm-serving\",condition=\"true\"})"}]},{"id":22,"title":"CPU
|
||||
by Predictor","type":"timeseries","gridPos":{"h":8,"w":8,"x":4,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(container_cpu_usage_seconds_total{namespace=\"llm-serving\"}[5m]))
|
||||
by (pod)","legendFormat":"{{pod}}"}]},{"id":23,"title":"Memory by Predictor","type":"timeseries","gridPos":{"h":8,"w":8,"x":12,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"bytes"}},"targets":[{"expr":"sum(container_memory_working_set_bytes{namespace=\"llm-serving\"})
|
||||
by (pod)","legendFormat":"{{pod}}"}]},{"id":24,"title":"Predictor Restarts","type":"timeseries","gridPos":{"h":8,"w":4,"x":20,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(kube_pod_container_status_restarts_total{namespace=\"llm-serving\"}[15m]))
|
||||
by (pod)","legendFormat":"{{pod}}"}]}]},{"id":30,"title":"Gateway Resources","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":3},"panels":[{"id":31,"title":"CPU
|
||||
by Gateway Pod","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":4},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(container_cpu_usage_seconds_total{namespace=\"api\"}[5m]))
|
||||
by (pod)","legendFormat":"{{pod}}"}]},{"id":32,"title":"Memory by Gateway Pod","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":4},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"bytes"}},"targets":[{"expr":"sum(container_memory_working_set_bytes{namespace=\"api\"})
|
||||
by (pod)","legendFormat":"{{pod}}"}]},{"id":33,"title":"Gateway Restarts","type":"timeseries","gridPos":{"h":8,"w":8,"x":16,"y":4},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(kube_pod_container_status_restarts_total{namespace=\"api\"}[15m]))
|
||||
by (pod)","legendFormat":"{{pod}}"}]}]},{"id":40,"title":"Logs","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":4},"panels":[{"id":41,"title":"Gateway
|
||||
Logs","type":"logs","gridPos":{"h":10,"w":24,"x":0,"y":5},"datasource":{"type":"loki","uid":"loki"},"targets":[{"expr":"{namespace=\"api\",container=\"gateway\"}"}]},{"id":42,"title":"LLM
|
||||
Serving Logs","type":"logs","gridPos":{"h":10,"w":24,"x":0,"y":15},"datasource":{"type":"loki","uid":"loki"},"targets":[{"expr":"{namespace=\"llm-serving\"}"}]}]}]}'
|
||||
kind: ConfigMap
|
||||
metadata:
|
||||
annotations:
|
||||
grafana_folder: API
|
||||
labels:
|
||||
grafana_dashboard: '1'
|
||||
name: api-gateway-dashboard
|
||||
namespace: logging
|
||||
@@ -1,62 +0,0 @@
|
||||
apiVersion: v1
|
||||
data:
|
||||
cluster-infrastructure.json: '{"title":"Cluster Infrastructure","uid":"cluster-infra","schemaVersion":39,"timezone":"browser","time":{"from":"now-6h","to":"now"},"refresh":"30s","tags":["infrastructure","k8s"],"panels":[{"id":1,"title":"Cluster
|
||||
Health","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":0},"panels":[{"id":2,"title":"Nodes
|
||||
Ready","type":"stat","gridPos":{"h":4,"w":4,"x":0,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"red"},{"value":3,"color":"green"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"count(kube_node_status_condition{condition=\"Ready\",status=\"true\"}
|
||||
== 1)"}]},{"id":3,"title":"Pods Pending","type":"stat","gridPos":{"h":4,"w":4,"x":4,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"green"},{"value":1,"color":"red"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"sum(kube_pod_status_phase{phase=\"Pending\"})
|
||||
OR on() vector(0)"}]},{"id":4,"title":"CrashLoopBackOff","type":"stat","gridPos":{"h":4,"w":4,"x":8,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"green"},{"value":1,"color":"red"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"sum(kube_pod_container_status_waiting_reason{reason=\"CrashLoopBackOff\"})
|
||||
OR on() vector(0)"}]},{"id":5,"title":"OOMKilled (1h)","type":"stat","gridPos":{"h":4,"w":4,"x":12,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"green"},{"value":1,"color":"red"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"sum(increase(kube_pod_container_status_last_terminated_reason{reason=\"OOMKilled\"}[1h]))
|
||||
OR on() vector(0)"}]},{"id":6,"title":"Deploys Unavailable","type":"stat","gridPos":{"h":4,"w":4,"x":16,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"green"},{"value":1,"color":"red"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"count(kube_deployment_status_replicas_unavailable
|
||||
> 0) OR on() vector(0)"}]},{"id":7,"title":"Services Down","type":"stat","gridPos":{"h":4,"w":4,"x":20,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"green"},{"value":1,"color":"red"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"count(probe_success
|
||||
== 0) OR on() vector(0)"}]}]},{"id":10,"title":"Jobs & CronJobs","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":1},"panels":[{"id":11,"title":"Failed
|
||||
Jobs","type":"stat","gridPos":{"h":4,"w":4,"x":0,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"green"},{"value":1,"color":"red"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"count(kube_job_status_failed
|
||||
> 0) OR on() vector(0)"}]},{"id":12,"title":"Failed Jobs Detail","type":"table","gridPos":{"h":8,"w":10,"x":4,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"kube_job_status_failed
|
||||
> 0","format":"table","instant":true}]},{"id":13,"title":"Stuck Jobs (>1h)","type":"table","gridPos":{"h":8,"w":10,"x":14,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"kube_job_status_active
|
||||
== 1 and on(job_name,namespace) (time() - kube_job_status_start_time) > 3600","format":"table","instant":true}]},{"id":14,"title":"CronJob
|
||||
Last Success","type":"timeseries","gridPos":{"h":8,"w":12,"x":0,"y":10},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"dateTimeFromNow"}},"targets":[{"expr":"kube_cronjob_status_last_successful_time{namespace=~\"cicd|kube-system|paperless\"}","legendFormat":"{{namespace}}/{{cronjob}}"}]},{"id":15,"title":"Container
|
||||
Restart Storm (top 10)","type":"timeseries","gridPos":{"h":8,"w":12,"x":12,"y":10},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"topk(10,
|
||||
sum(rate(kube_pod_container_status_restarts_total[15m])) by (namespace, pod))","legendFormat":"{{namespace}}/{{pod}}"}]}]},{"id":20,"title":"Node
|
||||
Resources","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":2},"panels":[{"id":21,"title":"CPU
|
||||
% by Node","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"percent"}},"targets":[{"expr":"(1
|
||||
- avg(rate(node_cpu_seconds_total{mode=\"idle\"}[5m])) by (instance)) * 100","legendFormat":"{{instance}}"}]},{"id":22,"title":"Memory
|
||||
% by Node","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"percent"}},"targets":[{"expr":"(1
|
||||
- node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100","legendFormat":"{{instance}}"}]},{"id":23,"title":"Disk
|
||||
% by Node","type":"timeseries","gridPos":{"h":8,"w":8,"x":16,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"percent"}},"targets":[{"expr":"(1
|
||||
- node_filesystem_avail_bytes{mountpoint=\"/\"} / node_filesystem_size_bytes{mountpoint=\"/\"})
|
||||
* 100","legendFormat":"{{instance}}"}]},{"id":24,"title":"Load Average","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":11},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"node_load1","legendFormat":"1m
|
||||
{{instance}}"},{"expr":"node_load5","legendFormat":"5m {{instance}}"}]},{"id":25,"title":"Network
|
||||
Errors & Drops","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":11},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"rate(node_network_receive_errs_total[5m])","legendFormat":"rx-err
|
||||
{{instance}}"},{"expr":"rate(node_network_transmit_errs_total[5m])","legendFormat":"tx-err
|
||||
{{instance}}"},{"expr":"rate(node_network_receive_drop_total[5m])","legendFormat":"rx-drop
|
||||
{{instance}}"}]}]},{"id":30,"title":"Control Plane","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":3},"panels":[{"id":31,"title":"API
|
||||
Server Up","type":"stat","gridPos":{"h":4,"w":4,"x":0,"y":4},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","mappings":[{"type":"value","options":{"0":{"text":"DOWN","color":"red"},"1":{"text":"UP","color":"green"}}}]}},"targets":[{"expr":"min(up{job=\"apiserver\"})"}]},{"id":32,"title":"API
|
||||
Server Request Rate","type":"timeseries","gridPos":{"h":8,"w":10,"x":4,"y":4},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(apiserver_request_total[5m]))
|
||||
by (verb, code)","legendFormat":"{{verb}} {{code}}"}]},{"id":33,"title":"API Server
|
||||
Error Rate %","type":"timeseries","gridPos":{"h":8,"w":10,"x":14,"y":4},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"percent"}},"targets":[{"expr":"sum(rate(apiserver_request_total{code=~\"5..\"}[5m]))
|
||||
/ sum(rate(apiserver_request_total[5m])) * 100","legendFormat":"5xx %"}]},{"id":34,"title":"API
|
||||
Server Latency","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":12},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"s"}},"targets":[{"expr":"histogram_quantile(0.95,
|
||||
sum(rate(apiserver_request_duration_seconds_bucket[5m])) by (le))","legendFormat":"p95"},{"expr":"histogram_quantile(0.99,
|
||||
sum(rate(apiserver_request_duration_seconds_bucket[5m])) by (le))","legendFormat":"p99"}]},{"id":35,"title":"etcd
|
||||
Request Duration","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":12},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"s"}},"targets":[{"expr":"histogram_quantile(0.99,
|
||||
sum(rate(etcd_request_duration_seconds_bucket[5m])) by (le))","legendFormat":"p99"}]}]},{"id":40,"title":"Storage","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":4},"panels":[{"id":41,"title":"Longhorn
|
||||
Disk Capacity","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":5},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"bytes"}},"targets":[{"expr":"longhorn_disk_capacity_bytes","legendFormat":"capacity
|
||||
{{node}}"},{"expr":"longhorn_disk_reservation_bytes","legendFormat":"reserved
|
||||
{{node}}"}]},{"id":42,"title":"PVC Phase","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":5},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"kube_persistentvolumeclaim_status_phase","legendFormat":"{{namespace}}/{{persistentvolumeclaim}}
|
||||
{{phase}}"}]}]},{"id":50,"title":"DNS & Networking","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":5},"panels":[{"id":51,"title":"CoreDNS
|
||||
Cache Hit Rate","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":6},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"percentunit"}},"targets":[{"expr":"rate(coredns_cache_hits_total[5m])
|
||||
/ (rate(coredns_cache_hits_total[5m]) + rate(coredns_cache_misses_total[5m]))","legendFormat":"{{server}}"}]},{"id":52,"title":"CoreDNS
|
||||
Errors","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":6},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(coredns_dns_responses_total{rcode=~\"SERVFAIL|NXDOMAIN\"}[5m]))
|
||||
by (rcode)","legendFormat":"{{rcode}}"}]}]},{"id":60,"title":"Logs","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":6},"panels":[{"id":61,"title":"Error
|
||||
Rate by Namespace","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":7},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum
|
||||
by (namespace) (count_over_time({namespace=~\"kube-system|cert-manager|ingress-nginx|longhorn-system\"}
|
||||
|= \"error\" [5m]))","legendFormat":"{{namespace}}"}]},{"id":62,"title":"Control
|
||||
Plane Logs","type":"logs","gridPos":{"h":10,"w":24,"x":0,"y":15},"datasource":{"type":"loki","uid":"loki"},"targets":[{"expr":"{namespace=\"kube-system\"}"}]},{"id":63,"title":"Cluster
|
||||
Addon Logs","type":"logs","gridPos":{"h":10,"w":24,"x":0,"y":25},"datasource":{"type":"loki","uid":"loki"},"targets":[{"expr":"{namespace=~\"cert-manager|ingress-nginx|longhorn-system\"}"}]}]}]}'
|
||||
kind: ConfigMap
|
||||
metadata:
|
||||
annotations:
|
||||
grafana_folder: Infrastructure
|
||||
labels:
|
||||
grafana_dashboard: '1'
|
||||
name: cluster-infrastructure-dashboard
|
||||
namespace: logging
|
||||
@@ -0,0 +1,55 @@
|
||||
# k8s/monitoring/dashboards/control-plane-logs.yaml
|
||||
# Surfaces controller/control-plane logs that are already in Loki today
|
||||
# (Promtail scrapes every namespace with no filter) — this dashboard is the
|
||||
# "make it visible" piece, not new log collection.
|
||||
apiVersion: v1
|
||||
kind: ConfigMap
|
||||
metadata:
|
||||
name: control-plane-logs-dashboard
|
||||
namespace: logging
|
||||
labels:
|
||||
grafana_dashboard: "1"
|
||||
data:
|
||||
control-plane-logs.json: |
|
||||
{
|
||||
"title": "Cluster Control Plane & Controllers (Logs)",
|
||||
"uid": "control-plane-logs",
|
||||
"schemaVersion": 39,
|
||||
"timezone": "browser",
|
||||
"time": { "from": "now-1h", "to": "now" },
|
||||
"refresh": "30s",
|
||||
"panels": [
|
||||
{
|
||||
"id": 1,
|
||||
"title": "Error rate by namespace",
|
||||
"type": "timeseries",
|
||||
"gridPos": { "h": 6, "w": 24, "x": 0, "y": 0 },
|
||||
"datasource": { "type": "loki", "uid": "loki" },
|
||||
"targets": [
|
||||
{
|
||||
"expr": "sum by (namespace) (count_over_time({namespace=~\"kube-system|cert-manager|ingress-nginx|longhorn-system\"} |= \"error\" [5m]))"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"title": "Control plane (kube-apiserver, controller-manager, scheduler)",
|
||||
"type": "logs",
|
||||
"gridPos": { "h": 10, "w": 24, "x": 0, "y": 6 },
|
||||
"datasource": { "type": "loki", "uid": "loki" },
|
||||
"targets": [
|
||||
{ "expr": "{namespace=\"kube-system\"}" }
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"title": "Cluster add-ons (cert-manager, ingress-nginx, longhorn)",
|
||||
"type": "logs",
|
||||
"gridPos": { "h": 10, "w": 24, "x": 0, "y": 16 },
|
||||
"datasource": { "type": "loki", "uid": "loki" },
|
||||
"targets": [
|
||||
{ "expr": "{namespace=~\"cert-manager|ingress-nginx|longhorn-system\"}" }
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -1,280 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Generate consolidated Grafana dashboards as k8s ConfigMap YAML files."""
|
||||
|
||||
import json
|
||||
import os
|
||||
|
||||
DASHBOARD_DIR = os.path.expanduser("~/workplace/homelab/k8s/infra/monitoring/dashboards")
|
||||
|
||||
DS_PROM = {"type": "prometheus", "uid": "prometheus"}
|
||||
DS_LOKI = {"type": "loki", "uid": "loki"}
|
||||
|
||||
|
||||
def stat_panel(id, title, expr, x, y, w=4, h=4, unit="short", mappings=None, thresholds=None):
|
||||
p = {
|
||||
"id": id, "title": title, "type": "stat",
|
||||
"gridPos": {"h": h, "w": w, "x": x, "y": y},
|
||||
"datasource": DS_PROM,
|
||||
"fieldConfig": {"defaults": {"unit": unit}},
|
||||
"targets": [{"expr": expr}],
|
||||
}
|
||||
if mappings:
|
||||
p["fieldConfig"]["defaults"]["mappings"] = mappings
|
||||
if thresholds:
|
||||
p["fieldConfig"]["defaults"]["thresholds"] = thresholds
|
||||
p["fieldConfig"]["defaults"]["color"] = {"mode": "thresholds"}
|
||||
return p
|
||||
|
||||
|
||||
def ts_panel(id, title, exprs, x, y, w=8, h=8, unit="short"):
|
||||
targets = []
|
||||
for e in exprs:
|
||||
if isinstance(e, tuple):
|
||||
targets.append({"expr": e[0], "legendFormat": e[1]})
|
||||
else:
|
||||
targets.append({"expr": e, "legendFormat": "{{pod}}"})
|
||||
return {
|
||||
"id": id, "title": title, "type": "timeseries",
|
||||
"gridPos": {"h": h, "w": w, "x": x, "y": y},
|
||||
"datasource": DS_PROM,
|
||||
"fieldConfig": {"defaults": {"unit": unit}},
|
||||
"targets": targets,
|
||||
}
|
||||
|
||||
|
||||
def table_panel(id, title, expr, x, y, w=12, h=8):
|
||||
return {
|
||||
"id": id, "title": title, "type": "table",
|
||||
"gridPos": {"h": h, "w": w, "x": x, "y": y},
|
||||
"datasource": DS_PROM,
|
||||
"targets": [{"expr": expr, "format": "table", "instant": True}],
|
||||
}
|
||||
|
||||
|
||||
def log_panel(id, title, query, x, y, w=24, h=10):
|
||||
return {
|
||||
"id": id, "title": title, "type": "logs",
|
||||
"gridPos": {"h": h, "w": w, "x": x, "y": y},
|
||||
"datasource": DS_LOKI,
|
||||
"targets": [{"expr": query}],
|
||||
}
|
||||
|
||||
|
||||
def row(id, title, y, panels, collapsed=True):
|
||||
return {
|
||||
"id": id, "title": title, "type": "row",
|
||||
"collapsed": collapsed, "gridPos": {"h": 1, "w": 24, "x": 0, "y": y},
|
||||
"panels": panels,
|
||||
}
|
||||
|
||||
|
||||
def write_dashboard(filename, dashboard, folder):
|
||||
cm = {
|
||||
"apiVersion": "v1",
|
||||
"kind": "ConfigMap",
|
||||
"metadata": {
|
||||
"name": filename.replace(".yaml", "-dashboard"),
|
||||
"namespace": "logging",
|
||||
"labels": {"grafana_dashboard": "1"},
|
||||
"annotations": {"grafana_folder": folder},
|
||||
},
|
||||
"data": {
|
||||
filename.replace(".yaml", ".json"): json.dumps(dashboard, separators=(",", ":"))
|
||||
},
|
||||
}
|
||||
|
||||
import yaml
|
||||
path = os.path.join(DASHBOARD_DIR, filename)
|
||||
with open(path, "w") as f:
|
||||
yaml.dump(cm, f, default_flow_style=False, allow_unicode=True)
|
||||
print(f" wrote {path}")
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# Dashboard 1: Cluster Infrastructure
|
||||
# ============================================================================
|
||||
|
||||
def build_cluster_infrastructure():
|
||||
zero_thresholds = {"mode": "absolute", "steps": [
|
||||
{"value": None, "color": "green"}, {"value": 1, "color": "red"}
|
||||
]}
|
||||
|
||||
panels = [
|
||||
row(1, "Cluster Health", 0, [
|
||||
stat_panel(2, "Nodes Ready", 'count(kube_node_status_condition{condition="Ready",status="true"} == 1)', 0, 1, thresholds={"mode":"absolute","steps":[{"value":None,"color":"red"},{"value":3,"color":"green"}]}),
|
||||
stat_panel(3, "Pods Pending", 'sum(kube_pod_status_phase{phase="Pending"}) OR on() vector(0)', 4, 1, thresholds=zero_thresholds),
|
||||
stat_panel(4, "CrashLoopBackOff", 'sum(kube_pod_container_status_waiting_reason{reason="CrashLoopBackOff"}) OR on() vector(0)', 8, 1, thresholds=zero_thresholds),
|
||||
stat_panel(5, "OOMKilled (1h)", 'sum(increase(kube_pod_container_status_last_terminated_reason{reason="OOMKilled"}[1h])) OR on() vector(0)', 12, 1, thresholds=zero_thresholds),
|
||||
stat_panel(6, "Deploys Unavailable", 'count(kube_deployment_status_replicas_unavailable > 0) OR on() vector(0)', 16, 1, thresholds=zero_thresholds),
|
||||
stat_panel(7, "Services Down", 'count(probe_success == 0) OR on() vector(0)', 20, 1, thresholds=zero_thresholds),
|
||||
]),
|
||||
row(10, "Jobs & CronJobs", 1, [
|
||||
stat_panel(11, "Failed Jobs", 'count(kube_job_status_failed > 0) OR on() vector(0)', 0, 2, thresholds=zero_thresholds),
|
||||
table_panel(12, "Failed Jobs Detail", 'kube_job_status_failed > 0', 4, 2, w=10),
|
||||
table_panel(13, "Stuck Jobs (>1h)", 'kube_job_status_active == 1 and on(job_name,namespace) (time() - kube_job_status_start_time) > 3600', 14, 2, w=10),
|
||||
ts_panel(14, "CronJob Last Success", [
|
||||
('kube_cronjob_status_last_successful_time{namespace=~"cicd|kube-system|paperless"}', "{{namespace}}/{{cronjob}}")
|
||||
], 0, 10, w=12, unit="dateTimeFromNow"),
|
||||
ts_panel(15, "Container Restart Storm (top 10)", [
|
||||
('topk(10, sum(rate(kube_pod_container_status_restarts_total[15m])) by (namespace, pod))', "{{namespace}}/{{pod}}")
|
||||
], 12, 10, w=12),
|
||||
]),
|
||||
row(20, "Node Resources", 2, [
|
||||
ts_panel(21, "CPU % by Node", [
|
||||
('(1 - avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) by (instance)) * 100', "{{instance}}")
|
||||
], 0, 3, unit="percent"),
|
||||
ts_panel(22, "Memory % by Node", [
|
||||
('(1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100', "{{instance}}")
|
||||
], 8, 3, unit="percent"),
|
||||
ts_panel(23, "Disk % by Node", [
|
||||
('(1 - node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100', "{{instance}}")
|
||||
], 16, 3, unit="percent"),
|
||||
ts_panel(24, "Load Average", [
|
||||
("node_load1", "1m {{instance}}"),
|
||||
("node_load5", "5m {{instance}}"),
|
||||
], 0, 11),
|
||||
ts_panel(25, "Network Errors & Drops", [
|
||||
("rate(node_network_receive_errs_total[5m])", "rx-err {{instance}}"),
|
||||
("rate(node_network_transmit_errs_total[5m])", "tx-err {{instance}}"),
|
||||
("rate(node_network_receive_drop_total[5m])", "rx-drop {{instance}}"),
|
||||
], 8, 11),
|
||||
]),
|
||||
row(30, "Control Plane", 3, [
|
||||
stat_panel(31, "API Server Up", 'min(up{job="apiserver"})', 0, 4, mappings=[
|
||||
{"type":"value","options":{"0":{"text":"DOWN","color":"red"},"1":{"text":"UP","color":"green"}}}
|
||||
]),
|
||||
ts_panel(32, "API Server Request Rate", [
|
||||
('sum(rate(apiserver_request_total[5m])) by (verb, code)', "{{verb}} {{code}}")
|
||||
], 4, 4, w=10),
|
||||
ts_panel(33, "API Server Error Rate %", [
|
||||
('sum(rate(apiserver_request_total{code=~"5.."}[5m])) / sum(rate(apiserver_request_total[5m])) * 100', "5xx %")
|
||||
], 14, 4, w=10, unit="percent"),
|
||||
ts_panel(34, "API Server Latency", [
|
||||
('histogram_quantile(0.95, sum(rate(apiserver_request_duration_seconds_bucket[5m])) by (le))', "p95"),
|
||||
('histogram_quantile(0.99, sum(rate(apiserver_request_duration_seconds_bucket[5m])) by (le))', "p99"),
|
||||
], 0, 12, unit="s"),
|
||||
ts_panel(35, "etcd Request Duration", [
|
||||
('histogram_quantile(0.99, sum(rate(etcd_request_duration_seconds_bucket[5m])) by (le))', "p99"),
|
||||
], 8, 12, unit="s"),
|
||||
]),
|
||||
row(40, "Storage", 4, [
|
||||
ts_panel(41, "Longhorn Disk Capacity", [
|
||||
("longhorn_disk_capacity_bytes", "capacity {{node}}"),
|
||||
("longhorn_disk_reservation_bytes", "reserved {{node}}"),
|
||||
], 0, 5, unit="bytes"),
|
||||
ts_panel(42, "PVC Phase", [
|
||||
('kube_persistentvolumeclaim_status_phase', "{{namespace}}/{{persistentvolumeclaim}} {{phase}}")
|
||||
], 8, 5),
|
||||
]),
|
||||
row(50, "DNS & Networking", 5, [
|
||||
ts_panel(51, "CoreDNS Cache Hit Rate", [
|
||||
('rate(coredns_cache_hits_total[5m]) / (rate(coredns_cache_hits_total[5m]) + rate(coredns_cache_misses_total[5m]))', "{{server}}")
|
||||
], 0, 6, unit="percentunit"),
|
||||
ts_panel(52, "CoreDNS Errors", [
|
||||
('sum(rate(coredns_dns_responses_total{rcode=~"SERVFAIL|NXDOMAIN"}[5m])) by (rcode)', "{{rcode}}")
|
||||
], 8, 6),
|
||||
]),
|
||||
row(60, "Logs", 6, [
|
||||
ts_panel(61, "Error Rate by Namespace", [
|
||||
('sum by (namespace) (count_over_time({namespace=~"kube-system|cert-manager|ingress-nginx|longhorn-system"} |= "error" [5m]))', "{{namespace}}")
|
||||
], 0, 7),
|
||||
log_panel(62, "Control Plane Logs", '{namespace="kube-system"}', 0, 15),
|
||||
log_panel(63, "Cluster Addon Logs", '{namespace=~"cert-manager|ingress-nginx|longhorn-system"}', 0, 25),
|
||||
]),
|
||||
]
|
||||
|
||||
return {
|
||||
"title": "Cluster Infrastructure",
|
||||
"uid": "cluster-infra",
|
||||
"schemaVersion": 39,
|
||||
"timezone": "browser",
|
||||
"time": {"from": "now-6h", "to": "now"},
|
||||
"refresh": "30s",
|
||||
"tags": ["infrastructure", "k8s"],
|
||||
"panels": panels,
|
||||
}
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# Dashboard 3: API Gateway
|
||||
# ============================================================================
|
||||
|
||||
def build_api_gateway():
|
||||
panels = [
|
||||
row(1, "Gateway Health", 0, [
|
||||
stat_panel(2, "Gateway Pods Ready", 'sum(kube_pod_status_ready{namespace="api",condition="true"})', 0, 1, thresholds={"mode":"absolute","steps":[{"value":None,"color":"red"},{"value":3,"color":"green"}]}),
|
||||
stat_panel(3, "Probe: healthz", 'probe_success{instance=~".*api.riotpiao.com/healthz"}', 4, 1, mappings=[
|
||||
{"type":"value","options":{"0":{"text":"DOWN","color":"red"},"1":{"text":"UP","color":"green"}}}
|
||||
]),
|
||||
ts_panel(4, "Probe Latency", [
|
||||
('probe_duration_seconds{instance=~".*api.riotpiao.com.*"}', "{{instance}}")
|
||||
], 8, 1, unit="s"),
|
||||
], collapsed=False),
|
||||
row(10, "Ingress Traffic (nginx)", 1, [
|
||||
ts_panel(11, "Request Rate by Status", [
|
||||
('sum(rate(nginx_ingress_controller_requests{ingress="api"}[5m])) by (status)', "{{status}}")
|
||||
], 0, 2),
|
||||
ts_panel(12, "Error Rate %", [
|
||||
('sum(rate(nginx_ingress_controller_requests{ingress="api",status=~"5.."}[5m])) / sum(rate(nginx_ingress_controller_requests{ingress="api"}[5m])) * 100', "5xx"),
|
||||
('sum(rate(nginx_ingress_controller_requests{ingress="api",status=~"4.."}[5m])) / sum(rate(nginx_ingress_controller_requests{ingress="api"}[5m])) * 100', "4xx"),
|
||||
], 8, 2, unit="percent"),
|
||||
ts_panel(13, "Latency p50/p95/p99", [
|
||||
('histogram_quantile(0.50, sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress="api"}[5m])) by (le))', "p50"),
|
||||
('histogram_quantile(0.95, sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress="api"}[5m])) by (le))', "p95"),
|
||||
('histogram_quantile(0.99, sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress="api"}[5m])) by (le))', "p99"),
|
||||
], 16, 2, unit="s"),
|
||||
]),
|
||||
row(20, "LLM Serving", 2, [
|
||||
stat_panel(21, "LLM Pods Ready", 'sum(kube_pod_status_ready{namespace="llm-serving",condition="true"})', 0, 3),
|
||||
ts_panel(22, "CPU by Predictor", [
|
||||
('sum(rate(container_cpu_usage_seconds_total{namespace="llm-serving"}[5m])) by (pod)', "{{pod}}")
|
||||
], 4, 3),
|
||||
ts_panel(23, "Memory by Predictor", [
|
||||
('sum(container_memory_working_set_bytes{namespace="llm-serving"}) by (pod)', "{{pod}}")
|
||||
], 12, 3, unit="bytes"),
|
||||
ts_panel(24, "Predictor Restarts", [
|
||||
('sum(rate(kube_pod_container_status_restarts_total{namespace="llm-serving"}[15m])) by (pod)', "{{pod}}")
|
||||
], 20, 3, w=4),
|
||||
]),
|
||||
row(30, "Gateway Resources", 3, [
|
||||
ts_panel(31, "CPU by Gateway Pod", [
|
||||
('sum(rate(container_cpu_usage_seconds_total{namespace="api"}[5m])) by (pod)', "{{pod}}")
|
||||
], 0, 4),
|
||||
ts_panel(32, "Memory by Gateway Pod", [
|
||||
('sum(container_memory_working_set_bytes{namespace="api"}) by (pod)', "{{pod}}")
|
||||
], 8, 4, unit="bytes"),
|
||||
ts_panel(33, "Gateway Restarts", [
|
||||
('sum(rate(kube_pod_container_status_restarts_total{namespace="api"}[15m])) by (pod)', "{{pod}}")
|
||||
], 16, 4),
|
||||
]),
|
||||
row(40, "Logs", 4, [
|
||||
log_panel(41, "Gateway Logs", '{namespace="api",container="gateway"}', 0, 5),
|
||||
log_panel(42, "LLM Serving Logs", '{namespace="llm-serving"}', 0, 15),
|
||||
]),
|
||||
]
|
||||
|
||||
return {
|
||||
"title": "API Gateway",
|
||||
"uid": "api-gateway",
|
||||
"schemaVersion": 39,
|
||||
"timezone": "browser",
|
||||
"time": {"from": "now-6h", "to": "now"},
|
||||
"refresh": "30s",
|
||||
"tags": ["api", "gateway", "llm"],
|
||||
"panels": panels,
|
||||
}
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# Generate
|
||||
# ============================================================================
|
||||
|
||||
print("Generating dashboards...")
|
||||
|
||||
# Dashboard 1
|
||||
write_dashboard("cluster-infrastructure.yaml", build_cluster_infrastructure(), "Infrastructure")
|
||||
|
||||
# Dashboard 3
|
||||
write_dashboard("api-gateway.yaml", build_api_gateway(), "API")
|
||||
|
||||
print("Done.")
|
||||
@@ -0,0 +1,121 @@
|
||||
# k8s/monitoring/dashboards/hardware-overview.yaml
|
||||
# Trimmed operator at-a-glance view across all nodes — node-exporter already
|
||||
# powers the deep-dive "Node Exporter Full" (#1860, see grafana-values.yaml),
|
||||
# this is the quick health-check version, not a replacement for it.
|
||||
apiVersion: v1
|
||||
kind: ConfigMap
|
||||
metadata:
|
||||
name: hardware-overview-dashboard
|
||||
namespace: logging
|
||||
labels:
|
||||
grafana_dashboard: "1"
|
||||
data:
|
||||
hardware-overview.json: |
|
||||
{
|
||||
"title": "Hardware Statistics (Operator Overview)",
|
||||
"uid": "hardware-overview",
|
||||
"schemaVersion": 39,
|
||||
"timezone": "browser",
|
||||
"time": { "from": "now-6h", "to": "now" },
|
||||
"refresh": "30s",
|
||||
"panels": [
|
||||
{
|
||||
"id": 1,
|
||||
"title": "Nodes up / down",
|
||||
"type": "stat",
|
||||
"gridPos": { "h": 5, "w": 24, "x": 0, "y": 0 },
|
||||
"datasource": { "type": "prometheus", "uid": "prometheus" },
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"mappings": [
|
||||
{ "type": "value", "options": { "0": { "text": "DOWN", "color": "red" } } },
|
||||
{ "type": "value", "options": { "1": { "text": "UP", "color": "green" } } }
|
||||
]
|
||||
}
|
||||
},
|
||||
"targets": [
|
||||
{ "expr": "up{job=~\".*node-exporter.*\"}", "legendFormat": "{{instance}}" }
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"title": "CPU usage % by node",
|
||||
"type": "timeseries",
|
||||
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 5 },
|
||||
"datasource": { "type": "prometheus", "uid": "prometheus" },
|
||||
"fieldConfig": { "defaults": { "unit": "percent", "max": 100, "min": 0 } },
|
||||
"targets": [
|
||||
{
|
||||
"expr": "(1 - avg(rate(node_cpu_seconds_total{mode=\"idle\"}[5m])) by (instance)) * 100",
|
||||
"legendFormat": "{{instance}}"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"title": "Memory usage % by node",
|
||||
"type": "timeseries",
|
||||
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 5 },
|
||||
"datasource": { "type": "prometheus", "uid": "prometheus" },
|
||||
"fieldConfig": { "defaults": { "unit": "percent", "max": 100, "min": 0 } },
|
||||
"targets": [
|
||||
{
|
||||
"expr": "(1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100",
|
||||
"legendFormat": "{{instance}}"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 4,
|
||||
"title": "Root filesystem usage % by node",
|
||||
"type": "timeseries",
|
||||
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 13 },
|
||||
"datasource": { "type": "prometheus", "uid": "prometheus" },
|
||||
"fieldConfig": { "defaults": { "unit": "percent", "max": 100, "min": 0 } },
|
||||
"targets": [
|
||||
{
|
||||
"expr": "(1 - node_filesystem_avail_bytes{mountpoint=\"/\"} / node_filesystem_size_bytes{mountpoint=\"/\"}) * 100",
|
||||
"legendFormat": "{{instance}}"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 5,
|
||||
"title": "Root filesystem space remaining",
|
||||
"type": "timeseries",
|
||||
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 13 },
|
||||
"datasource": { "type": "prometheus", "uid": "prometheus" },
|
||||
"fieldConfig": { "defaults": { "unit": "bytes" } },
|
||||
"targets": [
|
||||
{
|
||||
"expr": "node_filesystem_avail_bytes{mountpoint=\"/\"}",
|
||||
"legendFormat": "{{instance}}"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 6,
|
||||
"title": "Network errors/drops by node",
|
||||
"type": "timeseries",
|
||||
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 21 },
|
||||
"datasource": { "type": "prometheus", "uid": "prometheus" },
|
||||
"targets": [
|
||||
{ "expr": "rate(node_network_receive_errs_total[5m])", "legendFormat": "{{instance}} rx errs" },
|
||||
{ "expr": "rate(node_network_transmit_errs_total[5m])", "legendFormat": "{{instance}} tx errs" },
|
||||
{ "expr": "rate(node_network_receive_drop_total[5m])", "legendFormat": "{{instance}} rx drops" },
|
||||
{ "expr": "rate(node_network_transmit_drop_total[5m])", "legendFormat": "{{instance}} tx drops" }
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 7,
|
||||
"title": "Load average (1m / 5m) by node",
|
||||
"type": "timeseries",
|
||||
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 21 },
|
||||
"datasource": { "type": "prometheus", "uid": "prometheus" },
|
||||
"targets": [
|
||||
{ "expr": "node_load1", "legendFormat": "{{instance}} load1" },
|
||||
{ "expr": "node_load5", "legendFormat": "{{instance}} load5" }
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user