Files
homelab/k8s/infra/forgejo-runner/templates/configmap.yaml
T
rock ad1ce5e4aa fix: Forgejo CI runners — proper images, shared docker socket, auto-registration
All CI workflows across all repos were broken. Three root causes:

1. Runner labels pointed to bare Alpine image (forgejo/runner:6) which has
   no Node.js, no docker, no Go/Rust, no root, no apt-get. Every workflow
   step failed immediately.

   Fix: Change labels to official Debian language images:
     golang → docker://golang:1.26-bookworm
     node   → docker://node:22-bookworm
     rust   → docker://rust:1-bookworm

2. Docker socket not shared between dind sidecar and runner container.
   dind creates /run/docker.sock in its own filesystem. Runner container
   has a separate filesystem and can't see it. Workflow containers that
   run 'docker build' get 'Cannot connect to Docker daemon'.

   Fix: Share /run between dind and runner via emptyDir volume.
   Add docker_host: automount to runner config so the socket is
   mounted into workflow containers automatically.

   Note: /var/run is a symlink to /run in Alpine — must mount at /run.

3. Init container skipped re-registration if .runner file existed on PVC.
   Changing labels in values.yaml had no effect until PVC was manually
   deleted — not GitOps-friendly.

   Fix: Always rm .runner and re-register on pod start. Labels stay
   in sync with values.yaml automatically.

Also removes custom Dockerfiles and build-runner-images workflow — no longer
needed since official images provide the language tools directly.
2026-09-06 22:23:40 -07:00

40 lines
1.9 KiB
YAML

# act_runner (the forgejo-runner binary) ships no config.yaml by default, so
# `forgejo-runner daemon` runs on its hardcoded defaults -- notably
# container.valid_volumes: [] ("if the sequence is empty, no volumes can be
# mounted"). Confirmed via `forgejo-runner generate-config` on this exact
# image (code.forgejo.org/forgejo/runner:6) and by running the daemon against
# a minimal override locally: a job container that requests any bind mount
# (e.g. the dind mTLS certs at /docker-certs/client, needed for
# `docker login`/build/push steps) is rejected outright with no default
# config in place.
#
# Scoped narrowly to exactly the certs path, read-only. Not a wildcard
# (valid_volumes: ['**']) -- that would let any workflow in any repo this
# runner serves bind-mount arbitrary paths off the runner pod's filesystem
# into a job container, which is a real widening of the CI trust boundary,
# not just a convenience.
#
# network: host is also only settable here, not per-workflow. A workflow's
# `container.options: --network host` is silently ignored -- confirmed live:
# every job's actual `docker create` call logged
# `network="FORGEJO-ACTIONS-TASK-N_..."`, an auto-generated per-job bridge,
# regardless of that options string. On that isolated bridge, DOCKER_HOST=
# tcp://localhost:2376 resolves to the job container itself (no daemon there),
# not to the dind sidecar, so any docker command that actually needs the
# daemon (build, push -- anything past docker login, which only talks to the
# registry over the network and never touches DOCKER_HOST) fails with "Cannot
# connect to the Docker daemon". host mode puts every job container in dind's
# own network namespace instead, where the daemon really is listening.
apiVersion: v1
kind: ConfigMap
metadata:
name: {{ .Release.Name }}-config
namespace: {{ .Release.Namespace }}
data:
config.yaml: |
container:
valid_volumes:
- /docker-certs/client
network: host
docker_host: automount