5.5 KiB
Forgejo OCI Registry Cleanup
Automatic garbage collection for the Forgejo container registry. Deletes old image tags when newer versions are pushed, keeping only the latest N versions per repository.
Why
The Forgejo OCI registry stores all pushed images indefinitely. Without cleanup:
- Old/retired image versions accumulate
- Storage fills up (
longhornPVC) - Old versions clutter the UI
What it does
CronJob (forgejo-registry-cleanup):
- Runs daily at 2 AM UTC (configurable)
- Lists all images in the registry
- For each image, keeps only the latest 3 versions (configurable)
- Deletes tags for older versions
- Skips images with ≤ 3 tags (nothing to delete)
How to enable
The manifest is in k8s/bootstrap/phase3-forgejo/registry-cleanup-cronjob.yaml. It's currently disabled (suspended) because:
-
Forgejo registry auth needs to be configured
forgejo-registry-tokensecret must exist incicdnamespace- Should contain
usernameandpasswordkeys - User needs permission to delete images in the registry
-
Registry must expose
/v2/_catalogendpoint- Standard for OCI registries
- Forgejo includes this, but may be behind auth
Step 1: Create registry token
If forgejo-registry-token doesn't exist or is empty:
# As a Forgejo admin, create an API token with full scope
# https://forgejo.riotpiao.com/user/settings/tokens
# Copy the token
kubectl create secret generic forgejo-registry-token \
-n cicd \
--from-literal=username=<your-username> \
--from-literal=password=<the-api-token> \
--dry-run=client -o yaml | sops -e -i -
Or edit via k8s/argocd/secrets/forgejo-registry-token.enc.yaml:
apiVersion: v1
kind: Secret
metadata:
name: forgejo-registry-token
namespace: cicd
type: Opaque
stringData:
username: ci-bot # or any user with admin rights
password: <api-token>
Step 2: Test in dry-run mode
Before enabling for real, verify it works:
# Edit the CronJob to set DRY_RUN=true
kubectl set env cronjob/forgejo-registry-cleanup -n cicd DRY_RUN=true
# Trigger a test run
kubectl create job --from=cronjob/forgejo-registry-cleanup \
-n cicd forgejo-registry-cleanup-test
# Check logs
kubectl logs -n cicd -l job-name=forgejo-registry-cleanup-test -f
Dry-run output shows which images would be deleted without deleting them.
Step 3: Enable for real
# Set DRY_RUN=false and unsuspend
kubectl patch cronjob forgejo-registry-cleanup -n cicd \
-p '{"spec":{"suspend":false}}'
kubectl set env cronjob/forgejo-registry-cleanup -n cicd DRY_RUN=false
Configuration
Edit registry-cleanup-cronjob.yaml or patch the CronJob:
| Env var | Default | Purpose |
|---|---|---|
REGISTRY_HOST |
forgejo.riotpiao.com |
Registry hostname |
KEEP_VERSIONS |
3 |
How many recent versions to keep per image |
DRY_RUN |
false |
If true, log what would be deleted without deleting |
Schedule: Edit .spec.schedule (cron format). Current: 0 2 * * * (2 AM UTC daily).
Examples:
0 2 * * 0→ Weekly on Sunday at 2 AM0 0 1 * *→ Monthly on the 1st at midnight0 */6 * * *→ Every 6 hours
Monitoring
Check if running
# See all runs
kubectl get jobs -n cicd -l app=forgejo-registry-cleanup
# Latest run logs
kubectl logs -n cicd -l app=forgejo-registry-cleanup --tail=100 -f
Failed runs
If a job fails:
- Check logs:
kubectl logs -n cicd <job-pod> - Common issues:
- 403 Unauthorized: Registry token invalid or expired
- 404 _catalog: Registry doesn't expose catalog endpoint
- Connection refused: Registry unreachable (DNS, network policy)
Metrics
The job doesn't currently emit Prometheus metrics, but you can:
- Check pod exit codes in K8s events
- Parse logs for "Total images deleted: N"
- Set up log aggregation to alert on failures
Limitations
-
No version sorting: Tags are deleted in the order returned by the registry
- Assumption: registries return newest first (not always true)
- Fix: Parse semantic versions explicitly if needed
-
No protection for
latesttag: Iflatestis old, it will be kept but others deleted- Desired behavior: prioritize newest build + never delete
latest - Could add logic to always keep
latest+ latest N-1 tagged versions
- Desired behavior: prioritize newest build + never delete
-
No size-aware deletion: Deletes by tag count, not storage size
- Desired: keep until storage threshold is reached
- Would need registry V2 API extensions (
HEAD /v2/<image>/blobs/<digest>for size)
Customizing the script
Edit the cleanup.sh script in the ConfigMap to:
- Change sorting/selection logic
- Integrate with external systems (Slack alerts, Prometheus metrics)
- Add per-image exceptions (e.g., never delete
production-*tags) - Use
--delete-by-digestto reclaim actual disk space (not just catalog entries)
Example: Keep all tags matching v*.*.*. plus latest 2:
# In cleanup.sh, replace the tag filtering logic:
SEMVER_TAGS=$(echo "$TAGS" | grep -E '^v[0-9]+\.[0-9]+\.[0-9]+$' | sort -rV)
KEEP_TAGS="$SEMVER_TAGS $(echo "$TAGS" | head -2 | tr '\n' ' ')"
TAGS_TO_DELETE=$(echo "$TAGS" | grep -v -F "$KEEP_TAGS")
Future improvements
- Semantic version sorting (v1.0.0 > v0.9.9)
- Storage size-aware retention (keep until >80% full)
- Slack/email notifications on deletion
- Prometheus metrics export
- Per-image exception rules (YAML config)
- Integration with CI/CD pipeline (delete old PR images automatically)
Related: k8s/bootstrap/phase3-forgejo/ — Forgejo deployment manifests