Daily 04:00 UTC sweeper in kube-system:
- Delete failed Jobs older than 24h (any namespace)
- Delete completed standalone Jobs older than 72h (no CronJob owner)
- Delete orphan Error/Evicted pods older than 1h
- Self-cleans via ttlSecondsAfterFinished
Self-hosted photo backup (Google Photos replacement) - raw manifests,
no Helm chart, self-contained under k8s/apps/immich including its own
CNPG Postgres. Media PVC shares the cp-3 HDD 2TB/2TB with
paperless-media.
Postgres is pg18, not this repo's usual 16.2: CNPG's official pgvector
extension image (ghcr.io/cloudnative-pg/pgvector) is only published
for pg18, loaded via CNPG's ImageVolume extension mechanism (operator
1.30.0 / k8s 1.36.1 both support it). Immich auto-manages CREATE
EXTENSION itself at startup.
OIDC via a new "immich_role" Authentik scope mapping (homelab-admins/
immich-admins -> "admin" claim, else "user"), consumed by Immich's
OAuth roleClaim setting which re-syncs isAdmin on every login - more
reliable than Immich's racy first-user-is-admin fallback. Config
composed into an immich-oidc Secret and mounted as IMMICH_CONFIG_FILE,
matching the paperless-oidc pattern. k8s RBAC (immich-operator Role +
oidc:immich-admins binding) mirrors paperless/rbac.yaml.
immich namespace pre-created in k8s/infra/databases/namespaces.yaml
(not just immich's own CreateNamespace=true) since the iam PostSync
job's RoleBinding needs it to exist before wave 8.
Adds permissions claim + per-service admin groups in Authentik, scoped
Role/RoleBinding per service, public PKCE kubernetes OAuth2 client, and
kube-apiserver OIDC extraArgs. Also fixes paperless OIDC signup permissions
via adapter override and adds CoreDNS rewrite for authentik.riotpiao.com.
homelab-root and every child Application still tracked github.com/Riotpiaole/riotpiao.homelab.com, which had diverged from origin (Forgejo) for a while - pushes to Forgejo were never picked up by ArgoCD. Repointed to forgejo.riotpiao.com/rock/homelab.git, already covered by the AppProject's rock/* wildcard.
Fixes controlplane.tftpl's install.wipe:true (should be false, live CPs already run false) and syncs coredns Corefile back to what's actually deployed (drops an unrolled-out, stale Kong-era rewrite).
The poiman repo was renamed to poimen on Forgejo; the stale repoURL made
poimen-root fail with a 301 redirect ComparisonError (ArgoCD's git
client doesn't follow redirects on smart-HTTP fetch), blocking sync for
poimen-root and everything under it.
A1: Replace per-repo Forgejo entries with https://forgejo.riotpiao.com/rock/*
wildcard so onboarding never requires touching AppProject.
A2: Add wave -1 Application for k8s/argocd/projects/ so it syncs before
any Application references the AppProject.
Also add kustomization.yaml to k8s/argocd/projects/ to make it renderable.
Enabled by Stage 1 (A1, A2).
Ingress api/api now backs onto api-gateway:8080; the kong Application, its
Helm values, plugins and llm-routes are removed. Gateway image v0.0.0 is in
the Forgejo registry and the pull secret is in the api namespace.
Values changes were inert as a bootstrap release, so the proxy-body-size fix
never reached the live Ingress. First sync is manual — the chart owns the
Forgejo PVC.
- namespace: PodSecurity privileged, needed for /dev/kvm + privileged QEMU
- storageclass: 1 replica, strict-local, WaitForFirstConsumer
- deployment: nodeSelector workload=imessage + matching NoSchedule toleration,
Recreate strategy (two QEMU procs on one qcow2 corrupts it), no readiness
probe (guest install is interactive and takes many minutes)
- services: ClusterIP only; VNC is an unauthenticated console, reach it with
port-forward, never an Ingress
- networkpolicy: default-deny, opt-in via sms-client=true on port 1234
ArgoCD directory.include uses Go filepath.Match glob syntax, not shell
brace expansion - {a,b,c} silently matched nothing, only the original 2
files stayed tracked.
homelab-ca was referenced by 6 manifests (authentik, forgejo-runner,
blackbox-exporter, management-service) as a CA trust ConfigMap but never
existed anywhere - not in git, not live in cluster. Generated a new
10-year self-signed root CA, wired it as a ClusterIssuer (cert-manager
namespace) and distributed the public cert as a ConfigMap to every
consuming namespace (iam, cicd, monitoring, sqs). Private key lives only
in the encrypted Secret. Widened cert-manager-issuers' directory include
glob rather than creating a new Application - destination.namespace is
just a fallback default on a plain directory source, not a transformer,
so it doesn't fight with each ConfigMap's own explicit namespace.
Also adds grafana-oidc secret (GF_AUTH_GENERIC_OAUTH_CLIENT_SECRET),
same pre-existing gap as grafana-admin - was meant to come from a deleted
manual script, value already available in .env.
- Replace all forgejo.riotpiao.com repo URLs with [email protected] SSH URLs
- Enables immediate GitOps sync without waiting for Forgejo mirror setup
- Includes ingress-nginx now fully ArgoCD-managed (wave 0)
- SOPS secrets can now sync and decrypt TLS certificates
Root cause: pinned to temporalio/helm-charts @ 0.74.0, which uses the OLD
flat persistence schema (server.config.persistence.<store>.driver/.sql),
NOT the datastores:-wrapped schema shown in the CURRENT chart's
values/values.postgresql.yaml example (that key was introduced in a later
major version). Our old values.yaml used the datastores: key, which doesn't
exist in 0.74.0 - Helm doesn't validate unknown keys, so it was silently a
no-op. persistence.default.driver / persistence.visibility.driver stayed at
their chart default ("cassandra", with empty hosts: []) the entire time,
regardless of anything nested under datastores:.
Verified before writing this fix: cloned temporalio/helm-charts, checked out
tag temporal-0.74.0 (exact pin), ran +
against our actual values.yaml - confirmed the rendered
schema-setup Job used CASSANDRA_HOST/temporal-cassandra-tool the whole time.
Re-rendered with the corrected flat schema - zero Cassandra references,
correct postgres12 pluginName/connectAddr wired to ddb-cluster-rw.
Also fixed two compounding no-ops found the same way:
- -> real keys are schema.setup.enabled /
schema.update.enabled / schema.createDatabase.enabled (jobs.autoSetup
doesn't exist anywhere in this chart's templates or values.yaml).
- cassandra.enabled was never actually set to false (stayed at chart
default true) - now explicitly false, along with mysql/elasticsearch/
prometheus/grafana (none of which we want).
Password wiring: existingSecret: temporal-db-role + secretKey: password,
pointing at the CNPG-generated Secret - avoids storing the DB password as
plaintext in this values file. Added a new temporal-db-secret-sync
Application (sync-wave 7, one before temporal's wave 8) with a PreSync hook
Job that copies that Secret from the ddb namespace into temporal (Secrets
are namespace-scoped; CNPG creates it in ddb, but Temporal's pods run in
temporal). Deliberately a standalone directory/Application rather than
folded into temporal/'s own kustomization.yaml, which has a The Temporal CLI manages, monitors, and debugs Temporal apps. It lets you run
a local Temporal Service, start Workflow Executions, pass messages to running
Workflows, inspect state, and more.
* Start a local development service:
`temporal server start-dev`
* View help: pass `--help` to any command:
`temporal activity complete --help`
Usage:
temporal [command]
Available Commands:
activity Operate on Activity Executions
batch Manage running batch jobs
completion Generate the autocompletion script for the specified shell
config Manage config files (EXPERIMENTAL)
env Manage environments
help Help about any command
operator Manage Temporal deployments
schedule Perform operations on Schedules
server Run Temporal Server
task-queue Manage Task Queues
worker Read or update Worker state
workflow Start, list, and operate on Workflows
Flags:
--client-connect-timeout duration
The client connection timeout. 0s means no timeout.
(default 0s)
--color string
Output coloring. Accepted values: always, never, auto.
(default "auto")
--command-timeout duration
The command execution timeout. 0s means no timeout.
(default 0s)
--config-file $CONFIG_PATH/temporalio/temporal.toml
File path to read TOML config from, defaults to
$CONFIG_PATH/temporalio/temporal.toml where
`$CONFIG_PATH` is defined as `$HOME/.config` on Unix,
`$HOME/Library/Application Support` on macOS, and
`%AppData%` on Windows.
--disable-config-env
If set, disables loading environment config from
environment variables.
--disable-config-file
If set, disables loading environment config from config file.
--env ENV
Active environment name (ENV). (default "default")
--env-file $HOME/.config/temporalio/temporal.yaml
Path to environment settings file. Defaults to
$HOME/.config/temporalio/temporal.yaml.
-h, --help
help for temporal
--log-format string
Log format. Accepted values: text, json. (default "text")
--log-level string
Log level. Default is "never" for most commands and
"warn" for "server start-dev". Accepted values: debug,
info, warn, error, never. (default "never")
--no-json-shorthand-payloads
Raw payload output, even if the JSON option was used.
-o, --output string
Non-logging data output format. Accepted values: text,
json, jsonl, none. (default "text")
--profile string
Profile to use for config file.
--time-format string
Time format. Accepted values: relative, iso, raw.
(default "relative")
-v, --version
version for temporal
Use "temporal [command] --help" for more information about a command. transformer that would silently rewrite the copy-job's ddb-scoped
RoleBinding back to temporal (same class of bug just fixed in
k8s/security/iam/kustomization.yaml).
1. prometheus CRD sync failure (OutOfSync, permanently failing):
- helm.skipCrds: true on the prometheus Application - stop ArgoCD from
managing these CRDs through client-side apply (kube-prometheus-stack's
CRDs are large enough that the kubectl.kubernetes.io/last-applied-
configuration annotation exceeds etcd's 262144-byte limit on every sync).
- New prometheus-crds Application: plain git-sourced YAML (extracted via
helm show crds, committed under k8s/platform/monitoring/crds/), synced
with ServerSideApply=true. Chosen over a Helm-sourced 'CRDs only' app
because there's no clean way to ask ArgoCD's Helm source for 'render only
the crds/ directory' - a committed plain-YAML source is unambiguous.
- ServerSideApply=true can't go on the main prometheus Application: it
conflicts with managedNamespaceMetadata's forced namespace apply
('--force cannot be used with --server-side'), hence the split.
2. minio-tenant stuck OutOfSync (blocked 97+ minutes):
- minio-policy-setup PostSync hook Job was NAME:
mc alias set - set a new alias to configuration file
USAGE:
mc alias set ALIAS URL ACCESSKEY SECRETKEY
FLAGS:
--path value bucket path lookup supported by the server. Valid options are '[auto, on, off]' (default: "auto")
--api value API signature. Valid options are '[S3v4, S3v2]'
--config-dir value, -C value path to configuration folder (default: "/Users/rockliang/.mc") [$MC_CONFIG_DIR]
--quiet, -q disable progress bar display [$MC_QUIET]
--disable-pager, --dp disable mc internal pager and print to raw stdout [$MC_DISABLE_PAGER]
--no-color disable color theme [$MC_NO_COLOR]
--json enable JSON lines formatted output [$MC_JSON]
--debug enable debug output [$MC_DEBUG]
--resolve value resolves HOST[:PORT] to an IP address. Example: minio.local:9000=10.10.75.1 [$MC_RESOLVE]
--insecure disable SSL certificate verification [$MC_INSECURE]
--limit-upload value limits uploads to a maximum rate in KiB/s, MiB/s, GiB/s. (default: unlimited) [$MC_LIMIT_UPLOAD]
--limit-download value limits downloads to a maximum rate in KiB/s, MiB/s, GiB/s. (default: unlimited) [$MC_LIMIT_DOWNLOAD]
--custom-header value, -H value add custom HTTP header to the request. 'key:value' format.
--help, -h show help
EXAMPLES:
1. Add MinIO service under "myminio" alias. For security reasons turn off bash history momentarily.
$ set +o history
$ mc alias set myminio http://localhost:9000 minio minio123
$ set -o history
2. Add MinIO service under "myminio" alias, to use dns style bucket lookup. For security reasons
turn off bash history momentarily.
$ set +o history
$ mc alias set myminio http://localhost:9000 minio minio123 --api "s3v4" --path "off"
$ set -o history
3. Add Amazon S3 storage service under "mys3" alias. For security reasons turn off bash history momentarily.
$ set +o history
$ mc alias set mys3 https://s3.amazonaws.com \
BKIKJAA5BMMU2RHO6IBB V8f1CwQqAcwo80UEIJEjc5gVQUSSx5ohQ9GSrr12
$ set -o history
4. Add Amazon S3 storage service under "mys3" alias, prompting for keys.
$ mc alias set mys3 https://s3.amazonaws.com --api "s3v4" --path "off"
Enter Access Key: BKIKJAA5BMMU2RHO6IBB
Enter Secret Key: V8f1CwQqAcwo80UEIJEjc5gVQUSSx5ohQ9GSrr12
5. Add Amazon S3 storage service under "mys3" alias using piped keys.
$ set +o history
$ echo -e "BKIKJAA5BMMU2RHO6IBB\nV8f1CwQqAcwo80UEIJEjc5gVQUSSx5ohQ9GSrr12" | \
mc alias set mys3 https://s3.amazonaws.com --api "s3v4" --path "off"
$ set -o history against
http://minio.storage.svc.cluster.local:9000 - stale port. The minio
Service's port now tracks requestAutoCert on the Tenant (443 when
auto-TLS is on, 80 when off - we set it to false earlier), so 9000
doesn't exist on that Service anymore and the job hung in its 'waiting
for minio...' retry loop indefinitely, blocking ArgoCD's sync operation
(PostSync hooks block the sync from completing until they succeed).
- Fixed to use minio-cluster-hl.storage.svc.cluster.local:9000 - the
headless per-pod Service, which always listens on 9000 regardless of
the Tenant's TLS mode, so this can't silently break again the same way.