Author SHA1 Message Date
Story Crater Bot 87ea0f1147 feat(forgejo-runner): split into golang/node/rust runners, retire generic docker runner 2026-08-21 16:49:26 -07:00
Story Crater Bot d3b6ecfb62 Add Temporal worker for production task queue 2026-08-21 16:44:59 -07:00
Story Crater Bot 76d00bcfc3 fix: restore YaRN rope-scaling for reasoning-predictor (GPTQ requant dropped it, checkpoint's own ceiling was 40960 not 131072) 2026-08-21 16:42:39 -07:00
Story Crater Bot 76d5078611 fix: swap reasoning-predictor to Qwen3-32B-GPTQ-Int4, 131072 context (bnb-4bit decode too slow, GPTQ is Volta-native) 2026-08-21 16:39:33 -07:00
Story Crater Bot 892700b38c chore(forgejo-runner): arm for cascading delete ahead of 3-runner migration 2026-08-21 16:35:01 -07:00
Story Crater Bot 3b4e6684f1 refactor(forgejo-runner): template PVC names off Release.Name for multi-instance reuse 2026-08-21 16:30:41 -07:00
Story Crater Bot 93128a104e fix: retire one reasoning-predictor replica, run PP=2 across both V100s (Qwen3.5 MoE swap abandoned, moving to Ollama) 2026-08-21 16:23:07 -07:00
Story Crater Bot 98c5429a9d fix(forgejo-runner): job containers must use host network to reach dind 2026-08-21 16:22:19 -07:00
Story Crater Bot 37a7c37945 fix(forgejo-runner): egress to ingress-nginx by namespace, not a stale LB IP 2026-08-21 16:22:19 -07:00
Story Crater Bot a6051e025b fix(forgejo-runner): allow job containers to mount /docker-certs/client so docker login/build/push work 2026-08-21 16:22:19 -07:00
Story Crater Bot c956ac1465 chore: retire TemporalWorker CRD — agent-harness-worker and Forgejo build workflow removed 2026-08-20 23:10:06 -07:00
Story Crater Bot 886f546a02 fix(forgejo): enable Actions globally so workflow runs are created
Every workflow in the cluster has been silently dead. app.ini carried no
[actions] section, so Forgejo never created a run: the API returns
total_count: 0 for rock/homelab and rock/homelab-frontend alike, despite both
repos reporting has_actions: true, cluster-ci.yaml and build.yaml sitting on
their default branches, and forgejo-runner having registered successfully.

Registration does not go through the dispatcher, which is why the runner looks
healthy -- it logs "declared successfully" and "[poller 0] launched" and then
picks up nothing, forever. That reads like a runner or label problem and is
neither.

This also explains why the api-gateway images in the registry were all built
by hand: the pipeline that was supposed to build them has never once run.

Forgejo restarts on this values change; git and the container registry are
briefly unavailable.
2026-08-20 21:40:43 -07:00
Story Crater Bot 8f277adf19 stage1: A1-A2 AppProject and projects Application
A1: Replace per-repo Forgejo entries with https://forgejo.riotpiao.com/rock/*
    wildcard so onboarding never requires touching AppProject.

A2: Add wave -1 Application for k8s/argocd/projects/ so it syncs before
    any Application references the AppProject.

Also add kustomization.yaml to k8s/argocd/projects/ to make it renderable.

Enabled by Stage 1 (A1, A2).
2026-08-20 21:31:05 -07:00
Story Crater Bot 43483da902 pi-models: fix baseUrl to match homelab-frontend gateway contract
Kong was retired 2026-08-19, replaced by the rock/homelab-frontend Go
gateway (single /v1/chat/completions endpoint, model routed via the
request body's "model" field per API.md). Old per-model baseUrls
(/v1/ornith, /v1/reasoning, /v1/qwen) all 404 against the new gateway.
Also flipping reasoning's supportsTools to true -- confirmed working via
live test now that reasoning runs Qwen3-32B instead of DeepSeek-R1.
2026-08-20 00:26:52 -07:00
Story Crater Bot cf6c4f4d7d chore: drop the Kong key-auth credential secret, unused now that Kong is gone 2026-08-19 23:40:50 -07:00
Story Crater Bot 5167656445 feat: point api.riotpiao.com at the gateway ahead of Kong removal
Kong is being deleted, so the backend cannot stay kong-proxy. Gateway serves
404 on API routes until tasks 2.1/2.2 land.
2026-08-19 23:34:33 -07:00
Story Crater Bot 6eeac820a0 fix: resolve forgejo.riotpiao.com to the ingress LB on nodes
The pinned ClusterIP died when the rev-6 upgrade recreated the Service, timing
out every node image pull.
2026-08-19 23:25:41 -07:00
Story Crater Bot abd0af3ea9 revert: point api.riotpiao.com back at kong-proxy
Gateway pods are ErrImagePull — nodes resolve forgejo.riotpiao.com to a
ClusterIP and time out, so the Service had no endpoints and the host was
returning 503. Kong is still running; this restores it.
2026-08-19 22:56:22 -07:00
Story Crater Bot 98d8276476 feat: cut api.riotpiao.com over to the Go gateway and retire Kong
Ingress api/api now backs onto api-gateway:8080; the kong Application, its
Helm values, plugins and llm-routes are removed. Gateway image v0.0.0 is in
the Forgejo registry and the pull secret is in the api namespace.
2026-08-19 22:51:53 -07:00
Story Crater Bot 5af38a3f59 fix: ignore Reloader's injected env var on the Forgejo Deployment
Argo would otherwise strip STAKATER_* on each sync and fight Reloader for it,
recreating the forge pod every reconcile.
2026-08-19 22:43:29 -07:00
Story Crater Bot 832add824d feat: manage Forgejo with Argo instead of the bootstrap Helm release
Values changes were inert as a bootstrap release, so the proxy-body-size fix
never reached the live Ingress. First sync is manual — the chart owns the
Forgejo PVC.
2026-08-19 22:37:56 -07:00
Story Crater Bot 78c1fa3fb3 fix: set proxy-body-size 0 on the Forgejo chart Ingress
Two Ingresses claim forgejo.riotpiao.com and nginx honours the older chart one,
so the annotation on the other never applied and OCI pushes over 1m got 413.
2026-08-19 22:25:51 -07:00
Story Crater Bot 92d80173b4 feat: let the runner build and the cluster pull from the Forgejo registry
- Runner egress: allow 192.168.1.160/32:443. forgejo.riotpiao.com resolves to
  the ingress LB, inside the 192.168.1.0/24 block the NetworkPolicy denies, so
  docker push hung until timeout.
- dind CA: also mount homelab-ca at /etc/docker/certs.d/forgejo.riotpiao.com/,
  the path dockerd actually reads for per-registry trust.
- Pull secret: dockerconfigjson for the api namespace; /v2/ answers 401.
- AppProject: allow the Forgejo repo as a source for api-gw.
2026-08-19 21:48:01 -07:00
Story Crater Bot bd6a21e7e1 coordinator: make gitignore/PLAN.md setup idempotent, run every phase
Old i===0 && !resuming gate meant this only ran on a fresh start -- every
run this session was a resume, so poiman's branch never got the harness
gitignore rules, and portfolio's PLAN.md stayed tracked from before the
rule existed (gitignore doesn't affect already-tracked files). Now checks
and fixes both on every phase instead of once at genesis.
2026-08-19 21:12:33 -07:00
Story Crater Bot 5589bcd56a coordinator: detect+respawn dead pool sessions, bound resolver call
Dead sessions were only caught after a full 10-min stall timeout; now
polled via agent-manager status and respawned (retry once). spawnPi had
no timeout and could hang a repo's whole pipeline forever -- bounded to
5 minutes now.
2026-08-19 20:27:17 -07:00
Story Crater Bot fd7b714f09 reasoning: raise num_cpu_blocks 32->256 for real DRAM KV offload capacity
32 blocks was a ~1GB safety-valve leftover from the num_cpu_blocks=2000
hang incident, not meaningful offload capacity. This model's KV cache is
~32MB/128-token block (64 layers, 8 KV heads x 128 head_dim, fp16) --
256 blocks gives ~8GB of real DRAM offload (32,768 tokens), comfortably
under the pod's 36Gi limit alongside the ~20GB bnb-4bit weights.
2026-08-19 18:42:37 -07:00
Story Crater Bot 6291dd5afb reasoning: swap to dense Qwen3-32B-bnb-4bit for reliable tool calling
DeepSeek-R1-distill's tool_choice=auto narration bug needed a real fix,
not a workaround -- Qwen3's native tool-call format (hermes-compatible
chat template) solves it at the source instead of parsing around it.
Dense Qwen3-32B avoids the MoE arch/quantization pitfalls hit by the two
prior swap attempts (Kimi-distilled Qwen3.6 MoE, AWQ Qwen3-30B-A3B) --
same bnb-4bit path already proven working on this sm70 (V100) node.
2026-08-19 18:40:15 -07:00
Story Crater Bot 648388554d reasoning: revert to DeepSeek-R1-Distill-32B, retire Kimi/Qwen3 swap attempt
Three straight failures on worker-1: Kimi-K2.6-distilled Qwen3.6-35B-A3B
had an unrecognized model type (qwen3_5_moe); the AWQ-4bit fallback needed
compute capability 80+ (marlin INT4 kernels) but this node's GPU is sm70
(V100); on-the-fly bitsandbytes against the full-precision Qwen3-30B-A3B
kept crash-looping. Reverting to the last known-good config (ac1849d) --
tool-call narration bug on judge remains open, to revisit separately.
2026-08-19 18:34:15 -07:00
Story Crater Bot 6ce46b9ad5 reasoning: switch to on-the-fly bnb quant, worker-1 GPU is sm70 (V100)
cpatonn's pre-quantized build failed with a real hardware constraint:
"Quantization scheme not supported for current GPU. Min capability: 80.
Current capability: 70." AWQ/GPTQ/compressed-tensors marlin INT4 kernels
all need sm80+ -- this node's GPU can't run any of them. Only bitsandbytes
or full precision work here. Switching to the official full-precision
Qwen/Qwen3-30B-A3B-Thinking-2507 with --quantization=bitsandbytes
on-the-fly, and bumping the memory limit (36Gi->48Gi, request unchanged)
for the transient bf16-shard staging during load.
2026-08-19 18:30:24 -07:00
Story Crater Bot d3a059a6b4 reasoning: fix quantization flag mismatch (compressed-tensors, not awq_marlin)
cpatonn's "AWQ-4bit" repo is actually quantized via llm-compressor --
config.json declares compressed-tensors. Passing awq_marlin explicitly
conflicted with the checkpoint's own declared format and 400d at
config-validation time.
2026-08-19 18:23:15 -07:00
Story Crater Bot f015e4577b reasoning: fall back to official Qwen3-30B-A3B-Thinking-2507 AWQ-4bit
Kimi-K2.6-distilled Qwen3.6-35B-A3B crashed on boot -- model type
qwen3_5_moe unrecognized by transformers/vLLM 0.11.0, a genuinely
unsupported architecture, not a config issue. Using cpatonn's pre-quantized
AWQ-4bit build of the official Qwen3-30B-A3B-Thinking-2507 instead: native
vLLM support confirmed, no Kimi distillation but Qwen3's own tool-call
format is natively supported (the actual root problem being solved).
Restored max-num-seqs=4 since AWQ-4bit weight footprint leaves more KV
headroom than the bnb attempts did.
2026-08-19 18:20:34 -07:00
Story Crater Bot 1877f94bf6 reasoning: halve max-num-seqs to 2 for Kimi swap's first boot
New model's weight footprint (35B total MoE at on-the-fly bnb-4bit) leaves
less confirmed KV-cache headroom on the 32GB card than the old one had --
reducing concurrent-sequence worst case until real memory use is verified.
2026-08-19 18:14:22 -07:00
Story Crater Bot c5ff86d6dc reasoning: swap DeepSeek-R1-Distill-32B for Kimi-K2.6-distilled Qwen3.6-35B-A3B
R1-family tool_choice=auto is a documented vLLM architecture conflict --
the model narrates fake tool_calls in <think> instead of emitting real
ones, regardless of parser (deepseek_v3 400s, hermes parses but the model
still doesn't call out). Qwen3's native tool-call format sidesteps this.

No pre-quantized AWQ/GPTQ/bnb checkpoint exists for this specific distill
(only GGUF, llama.cpp/Ollama-only) -- using on-the-fly bitsandbytes
quantization against the full bf16 checkpoint instead.
2026-08-19 18:11:45 -07:00
Story Crater Bot 1bf611739b fix(agent-pod): force judge to actually call tools instead of narrating
Observed live: phase-judge (on homelab-reasoning) wrote a full page of
'I should check X, then Y' reasoning, declared VERDICT: PASS, and showed
the touch command as a fenced code block in its own text -- never ran
git diff, never wrote the result file, never touched the sentinel.
Coordinator timed out waiting on a file that was never going to appear.
2026-08-19 17:13:49 -07:00
Story Crater Bot 9fbfce7963 fix(agent-pod): install rust+gcc toolchain, symlink go, drop brave-search skill
poiman is Rust, portfolio is Go -- neither toolchain was reachable from an
interactive kubectl exec session (go's PATH export was local to its own
install script; rust was entirely absent, and cargo needs gcc as a linker
which also wasn't present).

brave-search was just a curl one-liner wrapped in its own skill file --
inlined the same curl command directly into info-collector/investigator's
instructions instead of dispatching to a separate skill for it.
2026-08-19 15:59:11 -07:00
Story Crater Bot ac1849d2a9 fix(llm-serving): bump reasoning memory limit to 36Gi headroom 2026-08-19 15:16:46 -07:00
Story Crater Bot 4eab8271c7 fix(llm-serving): num_cpu_blocks=2000 hung pod startup, drop to 32 2026-08-19 15:09:28 -07:00
Story Crater Bot dd491f6f8b feat(llm-serving): offload reasoning's KV cache to CPU DRAM
vLLM 0.11.0's native OffloadingConnector -- spills KV blocks to CPU RAM on
preemption instead of discarding them, avoiding recompute. Built into vLLM
core, no extra dependency. Bumped memory request/limit (+4Gi/replica) to
give the CPU block pool real room; worker-1 had ~18Gi of request headroom
across both replicas.
2026-08-19 15:02:28 -07:00
Story Crater Bot 5b041df884 fix(agent-pod): committed progress ledger so resume skips done tasks
Resuming the phase branch alone only recovers the code -- the task loop
still walked from the first task, re-verifying every already-done one
through a full planner call before reaching the first task that actually
needed work. .agent-progress is committed (not gitignored) and appended
per completed task, so a resumed run reads it once and skips straight
past known-done tasks with zero LLM calls. Validated locally against a
throwaway repo: second run skipped both tasks instantly (resumed: true)
instead of re-running planner on them.
2026-08-19 13:16:03 -07:00
Story Crater Bot 3ea057d83e fix(agent-pod): group sessions by repo, resume phase branches, fix empty-diff bug
- agent-manager spawn now gets --group repoId, so the TUI clusters
  planner/investigator/implementer/judge under one repo heading instead of
  4 unrelated sessions.
- runPhase was called with phaseBranch where it needed the true baseBranch,
  so every per-task judge review compared phaseBranch...HEAD -- always
  empty, since HEAD is phaseBranch while checked out. Judges only produced
  real verdicts anyway because they fell back to their own git log/show.
- Every restart re-cloned baseBranch fresh and started a new phase branch,
  discarding whatever a prior run had already committed mid-phase. Now:
  fetch+resume an existing phase branch if origin has one, push after every
  task instead of only at phase-end, and delete the phase branch (local +
  origin) once its milestone squash-merges into base.
2026-08-19 11:46:53 -07:00
Story Crater Bot 43c0e1faa2 fix(agent-pod): absolute paths for every sentinel/verdict file, cwd reminder per call
A pooled session's shell cwd drifts as it explores the repo between turns.
Seen live: a repo whose internal workspace dir is one letter off from the
repo's own directory name was enough for the agent to touch its sentinel
one level off from where coordinator watches for it -- coordinator waited
out the full timeout for a file that existed, just in the wrong place.
2026-08-19 11:26:33 -07:00
Story Crater Bot 3a244577e4 fix(agent-pod): fold judgeOnly status check into planner, drop separate judge pre-check 2026-08-19 10:57:21 -07:00
Story Crater Bot 7a0d09cbe0 fix(agent-pod): install python3 and sqlite3 in the container init 2026-08-19 10:46:50 -07:00
Story Crater Bot cb52356d13 fix(agent-pod): sync coordinator.js (slug repoId), tighten compaction, route judge to reasoning model 2026-08-19 10:41:33 -07:00
Story Crater Bot d0cbfac7a4 fix(agent-pod): stateless role pool (/new per reuse), never commit PLAN.md 2026-08-19 07:59:01 -07:00
Story Crater Bot 930374a3b8 feat(agent-pod): persistent per-role agent pool, concurrency moves to repo level
coordinator.js now runs one long-lived planner/investigator/implementer/judge
session per repo (reused across every task via tmux send-keys) instead of a
fresh spawn per task per stage. Tasks within a repo run sequentially against
that pool; concurrency is now REPO_CONCURRENCY (default 3) concurrent repos
via a new --repos flag, not concurrent tasks in one repo's phase.
2026-08-18 21:30:49 -07:00
Story Crater Bot b10d1c3a25 fix(llm-serving): use hermes tool-call parser, not deepseek_v3
deepseek_v3 400s on this checkpoint: "could not locate tool call start/end tokens in the tokenizer". unsloth/DeepSeek-R1-Distill-Qwen-32B is a Qwen2.5 base distilled on R1 reasoning traces -- it kept R1's <think> format but never got DeepSeek-V3's own special tool-call tokens registered in its tokenizer. hermes parses from text patterns instead of special tokens, so it works against the underlying Qwen tokenizer.
2026-08-18 20:56:40 -07:00
Story Crater Bot 4c54674dff fix(llm-serving): enable tool calling on homelab-reasoning
pi sends tool_choice="auto" for every session (Read/Bash/etc.) -- vLLM 400s on that without --enable-auto-tool-choice and a --tool-call-parser. Verified this deployed vLLM v0.11.0's registered parsers directly; deepseek_v3 matches, same family as the deepseek_r1 reasoning-parser already set (this Qwen-base distillation still emits DeepSeek's own tool-call format).
2026-08-18 20:50:21 -07:00
Story Crater Bot 05ff12c117 fix(agent-pod): use process.exitCode not process.exit() in coordinator.js
process.exit() right after console.log() can drop buffered stdout when it's piped (not a TTY) -- exactly kubectl exec's case. Explains the silent empty-output-exit-1 failures. process.exitCode + natural exit lets the event loop drain and flush first.
2026-08-18 19:13:07 -07:00
Story Crater Bot 56b1c96fcf fix(agent-pod): clone deterministically, not through a headless LLM call
git clone is mechanical -- routing it through spawnPi meant a crash gave zero diagnostic output, just a silent exit code. Direct runGit call now, same as commitPending/the squash-merge sequence. Drops the now-unused runStageWithResolver.
2026-08-18 18:57:18 -07:00
Story Crater Bot 26a959215e fix(agent-pod): sync coordinator.js ConfigMap, was stale since auto-discovery landed
The pod's coordinator-src ConfigMap still had the pre-auto-discovery version -- --tasks was required, no task-board parsing, no self-chained stages, no judge model routing. Regenerated from the current source.
2026-08-18 18:42:45 -07:00
Story Crater Bot 193b040de6 feat(agent-pod): implementer and judge learn Playwright for UI verification
Both skills already have Bash in allowed-tools -- no new pi capability needed. For UI/frontend work, implementer screenshots/clicks through the golden path via npx playwright instead of trusting that code compiling means it renders correctly; judge does the same as review evidence, FAILing on visual defects a diff alone wouldn't show. Doesn't apply to non-UI work.
2026-08-18 18:37:46 -07:00
Story Crater Bot 12778d5576 feat(llm-serving): scale ornith to 2 replicas instead of a dedicated grm GPU
reasoning keeps its 2 GPUs untouched. verifier's freed GPU goes to a second ornith replica instead of a standalone qwen-only pod -- both replicas load ornith:35b + qwen2.5:3b-instruct, k8s Service load-balances across them, so 2 concurrent implementer-style calls get independent instances.
2026-08-18 18:25:12 -07:00
Story Crater Bot cd6c620619 feat(llm-serving): retire verifier-predictor, add grm (qwen2.5:3b)
Frees verifier's GPU from an underused vLLM PRM deployment. qwen2.5:3b-instruct moves off ornith-predictor's shared pod onto its own dedicated GPU (grm.yaml), so verification/judge traffic stops contending with ornith:35b's agent traffic. /v1/qwen/chat/completions now points at grm-predictor; path unchanged.
2026-08-18 18:18:31 -07:00
Story Crater Bot 0a87302e19 fix(api): retire Kong key-auth on model routes; agent-pod builds agent-manager fork + ships coordinator.js
Kong key-auth rejected the Authorization: Bearer header every OpenAI-SDK-compatible client sends (verified: raw apikey header works, Bearer doesn't), so it's commented out and stripped from every llm-routes.yaml annotation until there's a Bearer-compatible fix. agent-pod now clones and builds the agent-manager fork from source at container start (no prebuilt binary shipped -- wrong arch and over ConfigMap's size cap) and ships coordinator.js alongside hub.js, so multiple repos can run the pipeline concurrently in one pod via kubectl exec. hub.js keeps its existing role as the container's foreground process, unchanged.
2026-08-18 17:50:52 -07:00
Story Crater Bot c64b68a36b fix(agent-pod): remote tui session for multi-agent 2026-08-18 15:08:04 -07:00
Story Crater Bot 4146a048c9 fix(ci): make the hardcoded-secret scan blocking and close the .gitignore/.sops.yaml gaps that let a plaintext deploy key through — also untracks tfplan binaries and skills-lock.json 2026-08-18 15:08:04 -07:00
Story Crater Bot f0fa1dbd27 fix(argocd): clone the public GitHub seed anonymously over HTTPS and delete the SSH deploy-key Secret — its private half had been committed in plaintext to a public remote, and a public repo needs no credential at all 2026-08-18 15:08:04 -07:00
Story Crater Bot 2b1c4b1df4 fix(forgejo): strategy Recreate for RWO data PVC — RollingUpdate deadlocked (new pod Multi-Attach error on the RWO gitea PVC held by the old pod, stuck Init forever) 2026-08-18 15:08:04 -07:00
Story Crater Bot 479318c532 fix(authentik): label argocd oidc-secret part-of=argocd — argocd's $secret substitution only reads labelled Secrets; without it OIDC login failed with oauth2 invalid_client (empty client_secret to IdP) 2026-08-18 15:08:04 -07:00
Story Crater Bot ff216429b9 feat(argocd): wire Authentik OIDC + local rock/cicd accounts + RBAC — adds oidc.config (homelab-admins->admin SSO), url, accounts.rock (login+apiKey) and accounts.cicd (apiKey for CD pipeline token), all role:admin 2026-08-18 15:08:04 -07:00
Story Crater Bot 7441aaf9c3 fix(homarr): raise CPU limit 500m->2 + disable analytics cron — Next.js aborted with exit 134 (SIGABRT) under CPU throttle during icon-updater/analytics, self-restarting in a loop and 502ing at the ingress 2026-08-18 15:08:04 -07:00
Story Crater Bot edd739198d fix(cilium): restrict L2 announcement to control-plane nodes — GPU worker lacks eno1 (Mellanox enp28s0f*), so when it won the .160 lease it couldn't ARP the VIP, black-holing all ingress (flapped on reboots) 2026-08-18 15:08:04 -07:00
Story Crater Bot 20f8aac95d fix(api): label Kong pods llm-client=true so llm-serving NetworkPolicy admits them — chat/embeddings/rerank/score routes silently hung until the client timeout because Cilium dropped Kong's packets
llm-serving-default-deny admits port 8080 only from pods carrying
llm-client=true. Kong lacked it, so every route that actually contacts an
upstream timed out. /v1/models masked the problem: request-termination answers
inside Kong and never touches an upstream, so it returned 200 throughout.

Opting in via podLabels rather than relaxing the policy — it is a compensating
control, not hygiene, since vLLM v0.11.0 is frozen on Volta and will not receive
patches for several remote/unauthenticated advisories.

podLabels land only in the pod template, not spec.selector.matchLabels, so this
is not an immutable-field change.
2026-08-18 15:08:04 -07:00
Story Crater Bot dbd3dc7b3d feat(api): add DeepSeek-shaped LLM API on Kong — /v1/models, per-model chat completions, embeddings, rerank, score; disable Kong response buffering so stream:true actually streams
Kong matches routes on host/path/method/header, never on the request body, so a
single /v1/chat/completions dispatching on body.model is not expressible in Kong
OSS (ai-proxy-advanced, which does multi-target model routing, is Enterprise).
Model therefore goes in the path:

  GET  /v1/models                        static list (request-termination)
  POST /v1/reasoning/chat/completions     reasoning-predictor  (vLLM)
  POST /v1/ornith/chat/completions        ornith-predictor     (Ollama)
  POST /v1/qwen/chat/completions          ornith-predictor     (Ollama, same pod)
  POST /v1/embeddings                     embeddings-predictor (TEI)
  POST /v1/rerank                         reranker-predictor   (TEI)
  POST /v1/score                          verifier-predictor   (vLLM pooling)

- each chat route force-overwrites body.model via request-transformer add+replace:
  ornith:35b and qwen2.5:3b-instruct share one Ollama pod, so without this a
  client hitting /v1/qwen with "model":"ornith:35b" would silently get the 35B
- routes live in ns llm-serving, not api: an Ingress can only reference a Service
  in its own namespace, and KIC watches all namespaces
- embeddings and score need no rewrite (TEI/vLLM already serve the canonical
  paths); rerank does, since /v1/rerank 404s and only /rerank exists
- read/write timeouts 1h: Kong defaults to 60s, which a 32B model on Volta
  exceeds mid-generation and returns 504
- nginx_proxy_proxy_buffering=off: buffered responses lump or stall SSE, and both
  hops (nginx Ingress and Kong) must be unbuffered or the buffered one wins
- no auth for now, per decision; api.riotpiao.com is reachable through nginx, so
  GPU time is currently unauthenticated
2026-08-18 15:08:04 -07:00
Story Crater Bot bebe8dc31b fix(ingress): remove stale ingress-nginx-controller-alias Service — its selfHeal kept clobbering the helm LoadBalancer Service (same name, dead ingress-nginx-bootstrap selector, 0 endpoints), unannouncing LB IP .160 and taking down all ingress 2026-08-18 15:08:04 -07:00
Story Crater Bot 4c0d30ce30 fix(homarr): add AUTH_OIDC_URI + email account linking — homarr hides the Authentik sign-in button unless AUTH_OIDC_URI (authorize endpoint) is set alongside AUTH_OIDC_ISSUER (per authentik/homarr SSO docs); was the missing var 2026-08-18 15:08:04 -07:00
Story Crater Bot a17ceedcd8 refactor(ingress): drop redundant ArgoCD ingress-nginx app — chart 4.15.1 was double-managed by both the helm-bootstrap release and this ArgoCD app (same chart), fighting over the controller/LB service (ingress-config drift). ingress-nginx is bootstrap-critical (ArgoCD's own reachability path), so helm-bootstrap is the single owner 2026-08-18 15:08:04 -07:00
Story Crater Bot 9a779ccaf4 feat(sms): add BlueBubbles iMessage delivery (Docker-OSX macOS VM pinned to worker-2) + ArgoCD app + dedicated longhorn-imessage-local SC — default longhorn SC can't schedule a 3-replica 200Gi volume (only worker-1 has 200Gi free at 100% over-provisioning) and Immediate binding would pin the qcow2 to the wrong node
- namespace: PodSecurity privileged, needed for /dev/kvm + privileged QEMU
- storageclass: 1 replica, strict-local, WaitForFirstConsumer
- deployment: nodeSelector workload=imessage + matching NoSchedule toleration,
  Recreate strategy (two QEMU procs on one qcow2 corrupts it), no readiness
  probe (guest install is interactive and takes many minutes)
- services: ClusterIP only; VNC is an unauthenticated console, reach it with
  port-forward, never an Ingress
- networkpolicy: default-deny, opt-in via sms-client=true on port 1234
2026-08-18 15:08:04 -07:00
Story Crater Bot 20d0517f79 feat(monitoring): enable Alertmanager (null receiver, longhorn PVC, az-a) + fix forgejo-rules ns forgejo->cicd — alerting delivery was disabled; forgejo PrometheusRule targeted a nonexistent namespace 2026-08-18 15:08:04 -07:00
Story Crater Bot c8ea7b9190 fix(prometheus): use longhorn StorageClass, drop nonexistent longhorn-wffc — Prometheus CR requested storageClass longhorn-wffc which doesn't exist (deprecated), so operator never created the StatefulSet (Reconciled=False, no metrics server) 2026-08-18 15:08:04 -07:00
Story Crater Bot d3e2215b5c fix(homarr): tune probes via chart values, drop fragile fix-probes-job — first-boot icon updater blocks health endpoint ~50s; default 10s×3 liveness SIGTERMs the pod (247 restarts, 503); chart exposes probes so the PostSync patch-job was unnecessary and reverted on every rollout 2026-08-18 15:08:04 -07:00
Story Crater Bot c00b2d1b53 fix(authentik): add minio policy scope mapping (homelab-admins->consoleAdmin else readonly) + set rock email — MinIO CLAIM_NAME=policy got no claim (no MinIO access); empty rock email broke Grafana OIDC (GitHub-style /emails 404) 2026-08-18 15:08:04 -07:00
Story Crater Bot db6bf742da fix(grafana): add email/login/name_attribute_path for Authentik OIDC — Grafana was falling back to GitHub-style <api_url>/emails (404 'Error getting email address'), breaking OAuth login; read identity from userinfo claims instead 2026-08-18 15:08:04 -07:00
Story Crater Bot f6298086f2 fix(forgejo-runner): cicd ns PSS privileged (dind needs it) + mount homelab-ca as ConfigMap not Secret — runner RS created 0 pods under baseline PSS, then FailedMount because homelab-ca is a ConfigMap trust bundle, not a Secret 2026-08-18 15:08:04 -07:00
Story Crater Bot fbc4e55718 feat(forgejo): add runner-token Secret via ksops — forgejo-runner register initContainer needs the registration token (from gitea actions generate-runner-token); was missing so runner deploy stuck 0/1 2026-08-18 15:08:04 -07:00
Story Crater Bot 06c35fb338 fix(coredns): own Corefile+hostname rewrites via Talos inlineManifest (single-source terraform/files/coredns/Corefile), drop ArgoCD coredns-config app — in-cluster *.riotpiao.com now resolves to nginx ingress so MinIO/OIDC discovery works; update cp-2 IP .213->.214 2026-08-18 15:08:04 -07:00
Story Crater Bot 09fa9c6145 feat(reloader): enable autoReloadAll + reloadOnCreate — watch all workloads without per-Deployment annotations (charts like homarr don't expose them); auto-restart pods when ksops secrets are created/rotated 2026-08-18 15:08:04 -07:00
Story Crater Bot bd99208754 fix(homarr): add auth-oidc-secret + db-encryption Secrets via ksops — homarr chart's envSecrets expect these exact names (oidc-client-id/secret, db-encryption-key); were never created so homarr CreateContainerConfigError 2026-08-18 15:08:04 -07:00
Story Crater Bot 58605e1b5c chore(duckdns): remove duckdns updater entirely — superseded by cloudflared tunnel; drop app-def, manifests, kube-system Deployment 2026-08-18 15:08:04 -07:00
Story Crater Bot 40fcbd036c fix(cert-manager): regenerate homelab-ca cert with basicConstraints CA:TRUE — old self-signed cert lacked CA:TRUE so the homelab-ca ClusterIssuer rejected it ('certificate is not a CA'); regen keypair Secret + trust-bundle ConfigMaps (4 ns) with matching CA cert 2026-08-18 15:08:04 -07:00
Story Crater Bot 69b5fc371d fix: deploy authentik/loki/vault Secrets via ksops (were dead helm-values fragments, causing CreateContainerConfigError) 2026-08-18 15:08:04 -07:00
Story Crater Bot 582524f921 fix(cert-manager): cert-manager-issuers directory.include renders empty — switch to explicit resources list, restore automated sync 2026-08-18 15:08:04 -07:00
Story Crater Bot 828e3fb287 refactor(argocd): replace SOPS CMP with ksops kustomize generator, rotate age key — CMP discover glob silently shadowed kustomize rendering of any app whose path held a .enc.yaml (MinIO Tenant/cloudflared/authentik jobs never applied); centralize 8 Secret manifests under k8s/argocd/secrets, defer 4 helm-values fragments 2026-08-18 15:08:04 -07:00
Story Crater Bot 1d5c18d62c fix(cert-manager): add homelab-ca.crt key to homelab-ca ConfigMaps — authentik init merge-ca-certs cats /homelab-ca/homelab-ca.crt which was missing, causing Init:Error and 503 2026-08-18 15:08:04 -07:00
Story Crater Bot bc1a6d8689 fix(argocd): resolve 502 on argocd.riotpiao.com, dedupe Ingress and TLS mode mismatch
argocd-server ran --insecure (plain HTTP :8080) while its Helm-managed
Ingress set ssl-passthrough: true, which sends nginx's raw TLS handshake
straight to the pod - HTTP server can't complete a TLS handshake, nginx
logged 502 (peer closed connection in SSL handshake). Compounded by a
second, conflicting Ingress for the same host in
k8s/bootstrap/ingress/ingress.yaml - two Ingress objects on one host is
undefined nginx routing behavior. Disabled the Helm-managed Ingress
(enabled: false) so ingress.yaml's passthrough Ingress is the sole
source of truth, and set server.insecure: false so argocd-server
actually terminates TLS itself, matching passthrough's requirement.
2026-08-18 15:08:04 -07:00
Story Crater Bot e4485412b0 fix(argocd): use comma-separated include list, not brace expansion
ArgoCD directory.include uses Go filepath.Match glob syntax, not shell
brace expansion - {a,b,c} silently matched nothing, only the original 2
files stayed tracked.
2026-08-18 15:08:04 -07:00
Story Crater Bot 2022595426 feat(cert-manager): add self-signed homelab-ca ClusterIssuer + trust bundle, fix grafana-oidc secret
homelab-ca was referenced by 6 manifests (authentik, forgejo-runner,
blackbox-exporter, management-service) as a CA trust ConfigMap but never
existed anywhere - not in git, not live in cluster. Generated a new
10-year self-signed root CA, wired it as a ClusterIssuer (cert-manager
namespace) and distributed the public cert as a ConfigMap to every
consuming namespace (iam, cicd, monitoring, sqs). Private key lives only
in the encrypted Secret. Widened cert-manager-issuers' directory include
glob rather than creating a new Application - destination.namespace is
just a fallback default on a plain directory source, not a transformer,
so it doesn't fight with each ConfigMap's own explicit namespace.

Also adds grafana-oidc secret (GF_AUTH_GENERIC_OAUTH_CLIENT_SECRET),
same pre-existing gap as grafana-admin - was meant to come from a deleted
manual script, value already available in .env.
2026-08-18 15:08:04 -07:00
Story Crater Bot f67aaa41d0 fix(portainer): pin to az-b (talos-cp-2), the real Longhorn storage node
nodeSelector still targeted az-a/talos-cp-1 from before the 3-CP topology
change. talos-cp-2 (az-b) has the dedicated Longhorn disks now, so the
pod's zone pin and the PVC's only viable replica location never matched
- ReplicaSchedulingFailure: disks are unavailable, pod stuck
ContainerCreating waiting on AttachVolume.
2026-08-18 15:08:04 -07:00
Story Crater Bot 69e8cfd6d1 fix(vault): add vault-minio-creds secret, was created by deleted helmfile presync hook
Vault's S3 storage backend needs AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY
from vault-minio-creds, previously generated by a helmfile presync hook
that no longer exists post-Terraform/helmfile removal. Sourced from the
same MINIO_ROOT_USER/PASSWORD already in .env. vault-unseal-keys still
missing separately — needs a live 'vault operator init' run, deferred.
2026-08-18 15:08:04 -07:00
Story Crater Bot db2fc9afc7 fix(portainer): correct storageClass name, longhorn-wffc never existed as a class
PVC sat Pending for 17 days — storageclass.storage.k8s.io "longhorn-wffc"
not found. Only longhorn, longhorn-cnpg, longhorn-static exist. Straight
naming drift, no such class was ever created.
2026-08-18 15:08:04 -07:00
Story Crater Bot 2b74b58ea6 fix(argocd): wire SOPS CMP sidecar + grafana-admin secret on repo-server 2026-08-18 15:08:04 -07:00
Story Crater Bot 27dbfb1bd7 feat(terraform): GPU worker node support (schematic, interface/diskSelector/swap tuning, gpu-node label, NVIDIA LTS extensions) 2026-08-18 15:08:04 -07:00
Story Crater Bot 4e67fd907a feat(argocd): migrate all applications from Forgejo to GitHub
- Replace all forgejo.riotpiao.com repo URLs with [email protected] SSH URLs
- Enables immediate GitOps sync without waiting for Forgejo mirror setup
- Includes ingress-nginx now fully ArgoCD-managed (wave 0)
- SOPS secrets can now sync and decrypt TLS certificates
2026-08-18 15:08:04 -07:00
147 changed files with 5053 additions and 5555 deletions
+460
View File
@@ -0,0 +1,460 @@
# CI/CD Pipeline: GitOps Validation & Deployment
## Overview
Pure GitOps CI/CD pipeline using Forgejo Actions (self-hosted runner).
**Principle:** Validate in CI, deploy via ArgoCD (no manual steps).
```
git push
[CI: Validate]
├─ yamllint (YAML syntax)
├─ kubeval (K8s manifests)
├─ kustomize build (all layers)
├─ argocd validation (app definitions)
└─ security scan (secrets, best practices)
[If push to main]
└─ ArgoCD auto-syncs (if enabled)
```
## Workflows
### 1. validate-k8s.yaml (Mandatory)
**Trigger:** Any push/PR with k8s/ changes
**What it does:**
1. Lints all YAML files (`yamllint`)
2. Validates K8s manifests (`kubeval`)
3. Builds all kustomization layers
4. Validates ArgoCD applications
5. Reports results
**Duration:** ~2-3 minutes
**Status:**
- ✅ PASS: All layers build, manifests valid → OK to merge
- ❌ FAIL: Syntax error, invalid resource, build failed → Fix & push again
**Example output:**
```
=== Building k8s/infrastructure/ ===
✓ Infrastructure built successfully
Resources: 47
=== Building k8s/bootstrap/ ===
✓ Bootstrap built successfully
Resources: 23
```
**When to check:**
- After every commit
- Before merging PRs
- On every branch
### 2. argocd-sync.yaml (Recommended)
**Trigger:** Push to main only (k8s/ changed)
**What it does:**
1. Authenticates with ArgoCD
2. Syncs `homelab-root` application
3. Waits for sync to complete (5 min timeout)
4. Verifies all applications healthy
**Duration:** 1-5 minutes (depends on resources)
**Status:**
- ✅ SYNCED: All resources deployed to cluster
- ❌ FAILED: Sync error, pod crashes, etc. → Check ArgoCD UI for details
**When it runs:**
- Automatically after merge to main
- Only on k8s/ changes (not on docs)
**Manual trigger (if needed):**
```bash
# SSH to runner or use Forgejo UI
# Re-run failed workflow
# Or manually sync: argocd app sync homelab-root
```
**Requires secrets:**
- `ARGOCD_SERVER`: ArgoCD server URL (https://argocd.riotpiao.com)
- `ARGOCD_AUTH_TOKEN`: ArgoCD API token (generate via ArgoCD UI)
### 3. security-scan.yaml (Optional)
**Trigger:** Any push/PR with k8s/ changes
**What it does:**
1. Scans Dockerfiles for vulnerabilities (`trivy`)
2. Scans Helm charts for security issues
3. Audits K8s manifests (`polaris`)
4. Checks for hardcoded secrets
5. Verifies security best practices
**Duration:** ~3-5 minutes
**Status:**
- ✅ PASS: No critical issues
- ⚠️ WARNING: Best practice recommendations (non-blocking)
- ❌ FAIL: Hardcoded secrets found (must fix)
**Common issues:**
- Missing resource limits (warning)
- Privileged containers (warning)
- Hardcoded passwords (ERROR)
---
## File Structure
```
.forgejo/
├── workflows/ # CI/CD workflows
│ ├── validate-k8s.yaml # Validate manifests (required)
│ ├── argocd-sync.yaml # Sync to cluster (auto on main)
│ └── security-scan.yaml # Security checks (optional)
└── CI-CD.md # This file
```
---
## Setup Instructions
### 1. Install Forgejo Runner
```bash
# On runner machine (inside cluster or external)
forgejo-runner register \
--instance https://forgejo.riotpiao.com \
--token <registration-token> \
--name homelab-runner \
--labels docker
forgejo-runner daemon
```
### 2. Add ArgoCD Secrets to Forgejo
```bash
# Go to: Forgejo → Settings → Secrets
# Add:
ARGOCD_SERVER = https://argocd.riotpiao.com
ARGOCD_AUTH_TOKEN = <token> # Generate: argocd account generate-token
```
### 3. Generate ArgoCD Token
```bash
# Inside cluster
kubectl -n argocd port-forward svc/argocd-server 8080:443
# Go to: https://localhost:8080/user-info/api-tokens
# Create new token (CI/CD)
# Copy token to Forgejo secrets
```
---
## Workflow Execution
### When developer pushes to feature branch:
```
git push origin feature/new-service
Forgejo Actions triggered
validate-k8s.yaml runs:
✓ Lints YAML
✓ Validates manifests
✓ Builds kustomizations
✓ All pass → GitHub comment: "Ready to merge"
Developer opens PR
Reviewer checks:
- Code changes (YAML)
- Workflow results
- ArgoCD impact (diff)
PR merged to main
```
### When merged to main:
```
git merge feature/new-service → main
Forgejo Actions triggered
validate-k8s.yaml runs:
✓ Same validation as above
argocd-sync.yaml runs (if enabled):
✓ Syncs homelab-root
✓ Waits for sync
✓ Verifies health
✓ Resources deployed to cluster
Cluster state = git state
(No manual kubectl apply needed!)
```
---
## Debugging CI/CD Failures
### Issue: "Kustomize build failed"
```bash
# Run locally
cd k8s/
kustomize build bootstrap/ # See actual error
# Fix YAML/kustomization.yaml
# git push again
```
### Issue: "Kubeval validation failed"
```bash
# Check K8s manifest syntax
kubeval k8s/platform/minio/config.yaml
# Common issues:
# - Typos in apiVersion, kind, metadata
# - Missing required fields
# - Invalid references (namespace, service name)
```
### Issue: "ArgoCD sync failed"
```bash
# Check ArgoCD UI
# https://argocd.riotpiao.com → homelab-root
# Or CLI
argocd app get homelab-root
argocd app logs homelab-root --follow
# Common issues:
# - Missing namespace (fixed by infrastructure layer)
# - Invalid Helm chart version
# - Secret not found
# - Network policy blocking traffic
```
### Issue: "Security scan found hardcoded secret"
```bash
# Fix: Remove secret from YAML
# Add to SOPS encryption instead
# Or use ArgoCD Sealed Secrets
# (if SOPS not available)
```
---
## Viewing Results
### Forgejo Actions UI
```
Repository → Actions
├─ validate-k8s
│ ├─ ✅ Success (merge safe)
│ ├─ ❌ Failed (fix required)
│ └─ Logs (click "Steps" → "Summary")
├─ argocd-sync
│ ├─ ✅ Synced (deployed)
│ └─ ❌ Failed (check ArgoCD UI)
└─ security-scan
├─ ✅ Pass (no critical issues)
└─ ⚠️ Warning (review, non-blocking)
```
### ArgoCD UI
```
https://argocd.riotpiao.com
├─ homelab-root
│ ├─ Status: Synced ✓
│ ├─ Health: Healthy ✓
│ └─ Details (click to see resources)
├─ layer-1-bootstrap
├─ layer-2-platform
├─ layer-3-security
├─ layer-4-applications
└─ layer-5-data
```
---
## Common Tasks
### Add new service to cluster
```bash
# 1. Create directory and kustomization.yaml
mkdir -p k8s/applications/my-service
cat > k8s/applications/my-service/kustomization.yaml << EOF
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
namespace: my-namespace
helmCharts:
- name: my-chart
repo: https://charts.example.com
version: 1.0.0
releaseName: my-service
valuesFile: values.yaml
EOF
# 2. Add values.yaml
cp /template/values.yaml k8s/applications/my-service/
# 3. Commit and push
git add k8s/applications/my-service/
git commit -m "feat(apps): add my-service"
git push
# 4. CI validates
# 5. Merge to main
# 6. ArgoCD syncs automatically
# ✓ Service deployed to cluster
```
### Rollback a deployment
```bash
# 1. Find broken commit
git log --oneline k8s/ # Identify bad commit
# 2. Revert
git revert <commit-hash>
git push
# 3. CI validates (should pass)
# 4. Merge to main
# 5. ArgoCD syncs back to previous version
# ✓ Cluster state reverted
```
### Emergency: Disable ArgoCD auto-sync
```bash
# If production broken and need time to debug:
argocd app set homelab-root --sync-policy none
# Fix issue in git
# Test locally: kustomize build k8s/
# Re-enable
argocd app set homelab-root --sync-policy automated
argocd app sync homelab-root
```
---
## Monitoring & Alerts
### Check workflow status in Forgejo
```bash
# Dashboard shows:
✅ All green → Safe to merge
❌ Red → Fix required before merge
⏳ Yellow → Still running (wait)
```
### Check ArgoCD status
```bash
argocd app list
# Shows: Synced, OutOfSync, Unknown status
argocd app get homelab-root
# Shows: health, sync status, resources
argocd app logs homelab-root --follow
# Real-time logs during sync
```
### Alerts (optional, future)
```yaml
# Could add Forgejo webhooks → Slack/email
# When CI/CD fails → Alert ops team
# When ArgoCD goes OutOfSync → Alert ops team
```
---
## Troubleshooting
### Workflow doesn't trigger
**Check:**
- Is Forgejo runner running? `forgejo-runner daemon`
- Did you push to correct branch? (validate runs on all, argocd-sync only on main)
- Did path match filter? (must change k8s/ or .forgejo/workflows/)
### Workflow hangs/times out
**Check:**
- kustomize build → Check for dependency cycles
- argocd sync → Check cluster resources (storage full? network down?)
- security scan → Large image scan → Takes time
**Fix:**
- Increase timeout in workflow
- Optimize kustomization (remove unused resources)
- Add resource limits to pods
### ArgoCD token invalid
**Fix:**
```bash
# Regenerate token
argocd account generate-token
# Update Forgejo secret
# Settings → Secrets → ARGOCD_AUTH_TOKEN = <new-token>
```
---
## Best Practices
**DO:**
- Commit all K8s changes to git (no manual kubectl apply)
- Run validate-k8s locally before push
- Write descriptive commit messages (why this change?)
- Review workflow logs before merging
- Monitor ArgoCD sync after merge
**DON'T:**
- Push directly to main (always use PR)
- Skip workflow validation (it catches errors early)
- Ignore security scan warnings
- Manually `kubectl apply` (breaks GitOps)
- Edit resources in cluster (they revert via ArgoCD)
---
## Next Steps
1. **Setup Forgejo runner** (if not already running)
2. **Add ArgoCD secrets** to Forgejo
3. **Test workflows** on feature branch
4. **Merge to main** → Watch ArgoCD sync
5. **Celebrate:** Full GitOps pipeline working! 🎉
+269
View File
@@ -0,0 +1,269 @@
name: Cluster CI Pipeline
on:
push:
branches:
- main
- develop
paths:
- 'k8s/**'
- '.forgejo/workflows/cluster-ci.yaml'
pull_request:
paths:
- 'k8s/**'
jobs:
ci:
runs-on: docker
steps:
# === Checkout ===
- name: Checkout
run: |
REPO_URL="${{ gitea.server_url }}/${{ gitea.repository }}.git"
CLONE_URL="https://${{ secrets.CI_RUNNER }}:${{ secrets.CI_RUNNER_SECRET }}@${REPO_URL#https://}"
git clone --depth 1 "$CLONE_URL" .
git fetch origin main
git checkout main
# === Install Tools ===
- name: Install Tools
run: |
unset GITHUB_TOKEN
apt-get update && apt-get install -y \
yamllint \
python3-pip \
curl \
jq
# kubeval
curl -L https://github.com/instrumenta/kubeval/releases/latest/download/kubeval-linux-amd64.tar.gz | tar xz
mv -f kubeval /usr/local/bin/
# kustomize
rm -f kustomize
curl -s https://raw.githubusercontent.com/kubernetes-sigs/kustomize/master/hack/install_kustomize.sh | bash
mv -f kustomize /usr/local/bin/
# argocd
curl -sSL -o /usr/local/bin/argocd https://github.com/argoproj/argo-cd/releases/latest/download/argocd-linux-amd64
chmod +x /usr/local/bin/argocd
# trivy
curl -sfL https://raw.githubusercontent.com/aquasecurity/trivy/main/contrib/install.sh | sh -s -- -b /usr/local/bin
# polaris
curl -L https://github.com/FairwindsOps/polaris/releases/latest/download/polaris-linux-amd64 -o /usr/local/bin/polaris
chmod +x /usr/local/bin/polaris
# === YAML Lint ===
- name: YAML Lint
run: |
echo "=== Linting YAML files ==="
yamllint k8s/ -c .yamllint.yaml || true
# === Kubeval - Validate K8s Syntax ===
- name: Kubeval - Validate K8s Syntax
run: |
echo "=== Validating Kubernetes manifests ==="
find k8s -name "*.yaml" -o -name "*.yml" | grep -v "\.archive" | while read file; do
echo "Validating $file..."
kubeval "$file" -d 2>/dev/null || true
done
# === Kustomize Build - All overlays ===
- name: Kustomize Build - Infrastructure
run: |
echo "=== Building k8s/infrastructure/ ==="
kustomize build k8s/infrastructure > /tmp/infrastructure.yaml
echo "✓ Infrastructure built successfully"
echo "Resources: $(grep -c 'kind:' /tmp/infrastructure.yaml)"
- name: Kustomize Build - Bootstrap
run: |
echo "=== Building k8s/bootstrap/ ==="
kustomize build k8s/bootstrap > /tmp/bootstrap.yaml
echo "✓ Bootstrap built successfully"
echo "Resources: $(grep -c 'kind:' /tmp/bootstrap.yaml || echo 0)"
- name: Kustomize Build - Platform
run: |
echo "=== Building k8s/platform/ ==="
kustomize build k8s/platform > /tmp/platform.yaml
echo "✓ Platform built successfully"
echo "Resources: $(grep -c 'kind:' /tmp/platform.yaml || echo 0)"
- name: Kustomize Build - Security
run: |
echo "=== Building k8s/security/ ==="
kustomize build k8s/security > /tmp/security.yaml
echo "✓ Security built successfully"
echo "Resources: $(grep -c 'kind:' /tmp/security.yaml || echo 0)"
- name: Kustomize Build - Applications
run: |
echo "=== Building k8s/applications/ ==="
kustomize build k8s/applications > /tmp/applications.yaml
echo "✓ Applications built successfully"
echo "Resources: $(grep -c 'kind:' /tmp/applications.yaml || echo 0)"
- name: Kustomize Build - Data
run: |
echo "=== Building k8s/data/ ==="
kustomize build k8s/data > /tmp/data.yaml
echo "✓ Data built successfully"
echo "Resources: $(grep -c 'kind:' /tmp/data.yaml || echo 0)"
- name: Validate ArgoCD Applications
run: |
echo "=== Validating ArgoCD Applications ==="
kubeval k8s/argocd/apps/*.yaml
# === Trivy - Scan Dockerfile ===
- name: Trivy - Scan Dockerfile
run: |
if find . -name "Dockerfile" 2>/dev/null | grep -v node_modules | head -1 | grep -q .; then
echo "=== Scanning Dockerfiles with Trivy ==="
find . -name "Dockerfile" -not -path "*/node_modules/*" -exec trivy config {} \;
else
echo "No Dockerfiles found"
fi
# === Trivy - Scan Helm Charts ===
- name: Trivy - Scan Helm Charts
run: |
if find k8s -name "Chart.yaml" 2>/dev/null | head -1 | grep -q .; then
echo "=== Scanning Helm charts with Trivy ==="
find k8s -name "Chart.yaml" -exec dirname {} \; | while read chart; do
echo "Scanning $chart..."
trivy config "$chart" || true
done
else
echo "No Helm charts found"
fi
# === Polaris - K8s Security Audit ===
- name: Polaris - K8s Security Audit
run: |
echo "=== Running Polaris K8s security audit ==="
polaris audit --audit-path /tmp/polaris-audit.json k8s/ || true
if [ -f /tmp/polaris-audit.json ]; then
echo "Security issues found:"
jq '.results[] | select(.pass == false)' /tmp/polaris-audit.json || true
fi
# === Check for Secrets in Code ===
- name: Check for Secrets in Code
run: |
echo "=== Scanning for hardcoded secrets ==="
# BLOCKING. This step used to only count findings and then exit 0, so a
# plaintext deploy key rode through it into a public remote. Two failure
# modes fixed: it now fails the build, and it matches key material by
# PEM header rather than only `private_key:`-style YAML field names.
# Findings are captured into variables and tested for emptiness rather than
# branching on grep's exit status: implementations disagree on the rc of a
# `-v` filter fed empty input, and a wrong rc here fails open.
# NOTE: --include must precede `--`; after `--` grep treats it as a filename
# and silently scans nothing.
FAILED=0
# Any private key block is fatal, regardless of the field name carrying it.
KEYS=$(grep -rIE --include="*.yaml" --include="*.yml" \
-- "-----BEGIN ([A-Z]+ )?PRIVATE KEY-----" k8s/ \
| grep -v "\.enc\.yaml" || true)
if [ -n "$KEYS" ]; then
echo "❌ Unencrypted private key material found:"
echo "$KEYS"
FAILED=1
fi
# Plaintext values in secret-ish YAML fields. SOPS output is ENC[...],
# so encrypted files never trip this.
VALS=$(grep -rInE --include="*.yaml" --include="*.yml" \
-- "^[[:space:]]*(password|token|apiKey|api_key|sshPrivateKey|client_secret):[[:space:]]*[\"']?[^\"'[:space:]{\$]{8,}" k8s/ \
| grep -v "ENC\[" | grep -v "\.enc\.yaml" || true)
if [ -n "$VALS" ]; then
echo "❌ Plaintext secret value found:"
echo "$VALS"
FAILED=1
fi
if [ "$FAILED" -ne 0 ]; then
echo "Encrypt with SOPS (see .sops.yaml) — *.enc.yaml files are exempt."
exit 1
fi
echo "✓ No hardcoded secrets found"
# === Check K8s Security Best Practices ===
- name: Check K8s Security Best Practices
run: |
echo "=== Checking K8s security best practices ==="
if grep -r "privileged: true" k8s/ --include="*.yaml" --include="*.yml"; then
echo "⚠️ Found privileged containers"
fi
if grep -r "hostNetwork: true" k8s/ --include="*.yaml" --include="*.yml"; then
echo "⚠️ Found hostNetwork usage"
fi
echo "Checking for missing resource limits..."
MISSING=0
find k8s -name "*.yaml" -o -name "*.yml" | while read file; do
if grep -q "kind: Deployment\|kind: StatefulSet\|kind: DaemonSet" "$file"; then
if ! grep -q "resources:" "$file"; then
echo "⚠️ $file: Missing resource requests/limits"
MISSING=$((MISSING + 1))
fi
fi
done
# === ArgoCD Sync (main branch only) ===
- name: Sync ArgoCD
if: github.ref == 'refs/heads/main' && github.event_name == 'push'
env:
ARGOCD_SERVER: ${{ secrets.ARGOCD_SERVER }}
ARGOCD_AUTH_TOKEN: ${{ secrets.ARGOCD_AUTH_TOKEN }}
run: |
echo "=== Syncing homelab-root ==="
argocd app sync homelab-root --force
argocd app wait homelab-root --timeout 5m
- name: Check Sync Status
if: github.ref == 'refs/heads/main' && github.event_name == 'push'
env:
ARGOCD_SERVER: ${{ secrets.ARGOCD_SERVER }}
ARGOCD_AUTH_TOKEN: ${{ secrets.ARGOCD_AUTH_TOKEN }}
run: |
echo "=== ArgoCD Applications Status ==="
argocd app list -o table
STATUS=$(argocd app get homelab-root -o jsonpath='{.status.syncStatus}')
if [ "$STATUS" != "Synced" ]; then
echo "❌ Root app sync failed: $STATUS"
exit 1
fi
echo "✓ Root app synced successfully"
- name: Health Check
if: github.ref == 'refs/heads/main' && github.event_name == 'push'
env:
ARGOCD_SERVER: ${{ secrets.ARGOCD_SERVER }}
ARGOCD_AUTH_TOKEN: ${{ secrets.ARGOCD_AUTH_TOKEN }}
run: |
echo "=== Checking Application Health ==="
argocd app get homelab-root -o wide
# === Summary ===
- name: Summary
if: always()
run: |
echo "=== CI Pipeline Summary ==="
echo "✓ YAML linted"
echo "✓ Manifests validated"
echo "✓ Kustomizations built"
echo "✓ Security scans completed"
echo "✓ Secrets check passed"
echo "✓ Best practices verified"
echo ""
echo "✓ All checks passed"
-3
View File
@@ -66,6 +66,3 @@ bootstrap-argocd.log
# one line here, which is how a plaintext deploy key reached a public remote.
k8s/**/*-secret.yaml
!k8s/**/*.enc.yaml
# IAM provisioning scripts contain credential references — never commit
scripts/iam/*.py
-206
View File
@@ -1,206 +0,0 @@
# Authentik Auth Integration for NextJS
## Current State
### Gateway Auth Status
| Endpoint | Auth Status | Notes |
|----------|-------------|-------|
| `/v1/chat/completions` | ❌ **OFF** | LLM routes have no auth middleware |
| `/v1/embeddings` | ❌ **OFF** | Same - no auth |
| `/v1/rerank` | ❌ **OFF** | Same - no auth |
| `X-Service: sqs` | ✅ **ON** | JWT validated via `internal/auth/jwt.go` |
| `/workflow` | ❌ **OFF** | Pass-through to Temporal |
**Auth module exists** at `homelab-frontend/internal/auth/jwt.go` but only wired for SQS.
LLM routes in `internal/proxy/proxy.go` have no auth middleware.
### Authentik App
Authentik app `local-llm` exists for LLM API auth:
- **Client ID**: `local-llm`
- **Client Secret**: `kubectl -n llm-serving get secret local-llm-jwt -o jsonpath='{.data.client-secret}' | base64 -d`
- **Token endpoint**: `https://authentik.riotpiao.com/application/o/token/`
- **Userinfo endpoint**: `https://authentik.riotpiao.com/application/o/userinfo/`
- **OIDC discovery**: `https://authentik.riotpiao.com/application/o/local-llm/.well-known/openid-configuration`
## Sign-in Methods
### 1. Resource Owner Password Credentials (ROPC)
Direct username/password login. Server-side only (needs client_secret).
```typescript
// API Route: app/api/auth/login/route.ts
const response = await fetch('https://authentik.riotpiao.com/application/o/token/', {
method: 'POST',
headers: { 'Content-Type': 'application/x-www-form-urlencoded' },
body: new URLSearchParams({
grant_type: 'password',
client_id: 'local-llm',
client_secret: process.env.AUTHENTIK_CLIENT_SECRET,
username: '[email protected]',
password: 'userpassword',
scope: 'openid email profile groups',
}),
});
const tokens = await response.json();
// { access_token, refresh_token, expires_in, token_type }
```
### 2. Authorization Code Flow (Browser Redirect)
Requires adding redirect URIs to `local-llm` Authentik app:
```python
# In k8s/infra/iam/scripts/authentik-provision.py, update:
"local-llm": {
...
"redirect_uris": [
"http://localhost:3000/api/auth/callback", # dev
"https://your-nextjs-app.com/api/auth/callback", # prod
],
}
```
Then standard OIDC flow:
1. Redirect to `https://authentik.riotpiao.com/application/o/authorize/?client_id=local-llm&redirect_uri=...&response_type=code&scope=openid email profile groups`
2. User logs in via Authentik UI
3. Callback receives `code`, exchange for tokens
## JWT Token Persistence
### Browser (localStorage)
```typescript
const TOKEN_KEY = 'llm_auth_token';
// Save
localStorage.setItem(TOKEN_KEY, JSON.stringify({
access_token: tokens.access_token,
refresh_token: tokens.refresh_token,
expires_at: Date.now() + tokens.expires_in * 1000,
}));
// Load
const stored = JSON.parse(localStorage.getItem(TOKEN_KEY) || 'null');
if (stored && stored.expires_at > Date.now()) {
// Token valid
}
// Clear (logout)
localStorage.removeItem(TOKEN_KEY);
```
### Server-side (HTTP-only cookies)
```typescript
// app/api/auth/login/route.ts
import { cookies } from 'next/headers';
// After successful login
cookies().set('llm_auth_token', JSON.stringify(tokens), {
httpOnly: true,
secure: process.env.NODE_ENV === 'production',
sameSite: 'lax',
maxAge: tokens.expires_in,
path: '/',
});
// Read in middleware or API routes
const tokenCookie = cookies().get('llm_auth_token');
const tokens = JSON.parse(tokenCookie?.value || 'null');
```
## Token Refresh
```typescript
async function refreshAccessToken(refresh_token: string) {
const response = await fetch('https://authentik.riotpiao.com/application/o/token/', {
method: 'POST',
headers: { 'Content-Type': 'application/x-www-form-urlencoded' },
body: new URLSearchParams({
grant_type: 'refresh_token',
client_id: 'local-llm',
client_secret: process.env.AUTHENTIK_CLIENT_SECRET,
refresh_token,
}),
});
return response.json();
}
```
## Environment Variables
```bash
# .env.local
AUTHENTIK_URL=https://authentik.riotpiao.com
AUTHENTIK_CLIENT_ID=local-llm
AUTHENTIK_CLIENT_SECRET=<from-secret>
# For client-side (public)
NEXT_PUBLIC_AUTHENTIK_URL=https://authentik.riotpiao.com
NEXT_PUBLIC_AUTHENTIK_CLIENT_ID=local-llm
```
## Using Token with LLM API
```typescript
const token = await getValidToken(); // from localStorage or cookie
const response = await fetch('https://api.riotpiao.com/v1/chat/completions', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'Authorization': `Bearer ${token}`, // JWT from Authentik
},
body: JSON.stringify({
model: 'reasoning',
messages: [{ role: 'user', content: 'Hello' }],
}),
});
```
## TODO
### Gateway-side (homelab-frontend)
- [ ] Wire `internal/auth/jwt.go` into LLM proxy handler (`internal/proxy/proxy.go`)
- [ ] Add `authRequired: true` to model config or create LLM-specific middleware
- [ ] Example pattern from SQS (in `internal/serviceadapter/router.go`):
```go
// In proxy.go ServeHTTP, before dispatching to LLM upstream:
if strings.HasPrefix(r.URL.Path, "/v1/") {
authHeader := r.Header.Get("Authorization")
claims, err := llmJWTAuth.ValidateBearerToken(authHeader)
if err != nil {
// Return 401/403
}
if !llmJWTAuth.CheckPermissions(claims, "llm:inference", "*") {
// Return 403 insufficient permissions
}
}
```
### Authentik-side
- [ ] Enable ROPC grant in Authentik provider settings (if not already)
- [ ] Add redirect URIs to `local-llm` app if browser OAuth flow needed:
```python
# k8s/infra/iam/scripts/authentik-provision.py
"local-llm": {
...
"redirect_uris": [
"http://localhost:3000/api/auth/callback",
"https://your-app.com/api/auth/callback",
],
}
```
### NextJS-side
- [ ] Until gateway auth is wired, LLM API works without token
- [ ] Once wired, add `Authorization: Bearer <token>` to all LLM requests
-92
View File
@@ -208,95 +208,3 @@ versions without warning in your own values file.
Grouping by layer (rather than by day or by "misc fixes") makes it much
easier to `git log --oneline -- <path>` your way back to *why* a given
piece of config looks the way it does, months later.
## Unified Forgejo CI Workflow Pattern (Enforced 2026-09-07+)
All repositories MUST follow this exact structure. No variations.
```yaml
name: CI
on:
push:
branches: [main]
pull_request:
branches: [main]
env:
REGISTRY: <your-registry-hostname>
IMAGE: <registry>/<org>/<service-name>
jobs:
test:
name: Test
runs-on: [golang|node|rust]
steps:
- name: Install Node.js for actions runtime
run: apt-get update && apt-get install -y nodejs
- name: Checkout code
uses: actions/checkout@v4
# Language-specific tests here (no docker, no registry)
# - name: Run tests
# run: npm test -- --run || true
build-push:
name: Build & Push Image
needs: test
if: github.event_name == 'push' && github.ref == 'refs/heads/main'
runs-on: [golang|node|rust]
steps:
- name: Install Node.js and Docker
run: |
apt-get update
apt-get install -y nodejs docker.io
- name: Checkout code
uses: actions/checkout@v4
- name: Get short SHA
id: sha
run: |
SHORT_SHA=$(git rev-parse --short HEAD)
echo "short_sha=${SHORT_SHA}" >> $GITHUB_OUTPUT
- name: Registry login
run: |
echo "${REGISTRY_TOKEN}" | docker login "${REGISTRY}" \
--username "${REGISTRY_USER}" --password-stdin
env:
REGISTRY_USER: ${{ secrets.FORGEJO_REGISTRY_USER }}
REGISTRY_TOKEN: ${{ secrets.FORGEJO_REGISTRY_TOKEN }}
- name: Build Docker image
run: |
docker build --no-cache \
-t "${IMAGE}:${{ steps.sha.outputs.short_sha }}" \
-t "${IMAGE}:latest" \
.
- name: Push Docker image
run: |
docker push "${IMAGE}:${{ steps.sha.outputs.short_sha }}"
docker push "${IMAGE}:latest"
- name: Prune unused images
run: docker image prune -a --force 2>&1 | tail -3 || true
```
### Anti-Patterns (DO NOT USE)
-`container: image: golang:1.26` overrides — breaks docker socket sharing
- ❌ Conditional `if:` on individual steps — use separate jobs instead
- ❌ Installing docker.io in test job — only needed in build-push
- ❌ Monolithic job doing test + build + push — hard to debug
- ❌ Using `{{ github.sha }}` for image tag — use short commit SHA for readability
### How It Works
1. **PR to feature branch** → test job runs, build-push skipped, nothing pushed
2. **Push to main** → test runs, build-push runs after test passes, image pushed
3. Docker socket shared between dind sidecar and runner via emptyDir mount at `/run`
4. `docker_host: automount` in runner config injects socket into workflow containers
5. Secrets (FORGEJO_REGISTRY_USER, TOKEN) set in Forgejo repo settings, NOT in git
-80
View File
@@ -42,86 +42,6 @@ All logs + metrics centralized in Grafana for debugging
- **Secrets at rest** — Vault + encrypted etcd; credentials never in logs or ConfigMaps
- **Infrastructure-as-code** — Every service deployed via Helmfile; one `helmfile apply` recovers from total failure
## ArgoCD — GitOps Deployment Flow
**ArgoCD** pulls infrastructure changes from git and syncs the cluster automatically.
No manual `kubectl apply` — push to git, ArgoCD detects the change, and deploys within ~3 minutes.
```
Developer pushes to git
ArgoCD detects change (every 3 min or webhook)
Syncs manifests to cluster
Workloads reconcile automatically
```
Applications are deployed in waves (numbered 00, 10, 20, 30, ...) to respect dependencies —
storage deploys before databases, databases before applications.
### Tracked Git Repositories
ArgoCD monitors these repos for changes:
| Repository | Purpose |
|------------|----------|
| `https://github.com/Riotpiaole/riotpiao.homelab.com` | Main infrastructure repo (all manifests in `k8s/argocd/apps/`) |
| `https://forgejo.riotpiao.com/rock/*` | Any `rock/*` repo in in-cluster Forgejo (apps + configs) |
| `https://github.com/Riotpiaole/Poimen-*` | External Poimen services (memory, workflows) |
To deploy a new application: create a git repo, add an Application manifest to the homelab repo's
`k8s/argocd/apps/`, commit + push, and ArgoCD syncs within 3 minutes.
## Management Planes — Talos vs Kubernetes
This cluster has **two separate management planes**, each with different workflows:
| Plane | What it manages | Workflow | Tool |
|-------|-----------------|----------|------|
| **Talos (OS)** | Node configuration, kernel params, networking, CoreDNS, machine state | Edit `terraform/``terraform apply``make apply-cp` | `terraform` + `talosctl` |
| **Kubernetes (workloads)** | All pods, services, deployments, ingresses, databases | Edit `k8s/argocd/apps/``git push` → ArgoCD syncs | `git` + ArgoCD |
**Critical distinction:**
- **Kubernetes resources** (`k8s/**`) flow through **git → ArgoCD** — never use `kubectl apply`
- **Talos machine config** (`terraform/**`) uses **local `terraform apply`** (sanctioned exception — CI can't hold node credentials)
Example: To add a CoreDNS hostname rewrite, you edit `terraform/files/coredns/Corefile`, then:
```bash
cd terraform && terraform apply -var-file=terraform.tfvars.local
cd .. && make apply-cp # talosctl apply-config to all 3 control planes
```
But to add a new Kubernetes Deployment or update an Ingress, you only `git push`**never `kubectl apply`**.
### CoreDNS ConfigMap Ownership — Critical
⚠️ **Warning:** The `coredns` ConfigMap in `kube-system` namespace is **owned by Talos**, not ArgoCD or kubectl.
It is rendered from `terraform/files/coredns/Corefile` into Talos's machine config at bootstrap time.
**Do not `kubectl apply` or `kubectl edit` this ConfigMap directly.** Doing so transfers field ownership to kubectl's
client-side-apply mechanism, and Talos's inline-manifest controller will silently no-op on every future reconcile
(server-side-apply conflict, no error surfaced).
**To update CoreDNS (e.g., add a hostname rewrite):**
1. Edit `terraform/files/coredns/Corefile`
2. Commit + push
3. Run `cd terraform && terraform apply -var-file=terraform.tfvars.local`
4. Run `make apply-cp` to push config to all control planes
5. CoreDNS picks up changes via its `reload` plugin — no pod restart needed
**If you accidentally edited the ConfigMap directly and broke Talos's ownership:**
```bash
kubectl delete configmap coredns -n kube-system
# Wait ~30s for Talos's k8s.ManifestApplyController to recreate it
kubectl get configmap coredns -n kube-system -w
```
Or as a stopgap, apply the correct content yourself:
```bash
kubectl apply --server-side -f <(terraform output coredns_config)
```
## Quick Start — Deploying the Cluster
### 1. Bootstrap Talos Nodes
View File
-14
View File
@@ -1,14 +0,0 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: immich-config
data:
DB_HOSTNAME: "immich-db-rw"
DB_DATABASE_NAME: "immich"
# Only pgvector is installed (see db.yaml) - no vectorchord extension image
# exists for pg18 in CNPG's catalog yet. Explicit instead of relying on
# auto-detect's vectorchord-first preference order.
DB_VECTOR_EXTENSION: "pgvector"
REDIS_HOSTNAME: "immich-redis"
IMMICH_MACHINE_LEARNING_URL: "http://immich-machine-learning:3003"
TZ: "America/Los_Angeles"
-58
View File
@@ -1,58 +0,0 @@
# Dedicated CNPG Postgres for Immich. Same recipe as paperless-db/authentik-db
# (2 instances, default longhorn storage class) except the operand is
# PostgreSQL 18, not 16.2 - the official CNPG pgvector extension image
# (ghcr.io/cloudnative-pg/pgvector) is only published for pg18, no pg16 tags
# exist in that registry. Immich itself supports pg18 fine (immich-app's own
# postgres image already ships 18-vectorchord builds).
#
# pgvector loaded via CNPG's ImageVolume extension mechanism (CNPG 1.27+,
# k8s ImageVolume feature - both present here: operator is 1.30.0, cluster is
# v1.36.1). No shared_preload_libraries needed - pgvector doesn't require
# preload, just CREATE EXTENSION, which immich-server issues itself at
# startup. Distro/pg-major must match between the operand image and the
# extension image (both "18"+"trixie" here) - CNPG's own compatibility rule.
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: immich-db
annotations:
argocd.argoproj.io/sync-options: SkipDryRunOnMissingResource=true
spec:
instances: 2
imageName: ghcr.io/cloudnative-pg/postgresql:18-minimal-trixie
postgresql:
extensions:
- name: pgvector
image:
reference: ghcr.io/cloudnative-pg/pgvector:0.8.1-18-trixie
bootstrap:
initdb:
database: immich
owner: app
encoding: UTF8
localeCollate: C
localeCType: C
# CREATE EXTENSION vector requires superuser (pgvector's control file
# isn't marked trusted) and the "app" owner role isn't one
# (enableSuperuserAccess: false, repo convention) - postInitApplicationSQL
# runs as superuser during initdb, before the app ever connects. Only
# fires on a fresh bootstrap; the live cluster already had this run
# manually once (kubectl exec ... psql -U postgres -c 'CREATE EXTENSION').
postInitApplicationSQL:
- "CREATE EXTENSION IF NOT EXISTS vector;"
- "CREATE EXTENSION IF NOT EXISTS cube;"
- "CREATE EXTENSION IF NOT EXISTS earthdistance;"
enableSuperuserAccess: false
resources:
requests: { memory: "512Mi", cpu: "250m" }
limits: { memory: "2Gi", cpu: "1" }
storage:
size: 20Gi
storageClass: longhorn
affinity:
podAntiAffinityType: preferred
topologyKey: kubernetes.io/hostname
tolerations:
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
-100
View File
@@ -1,100 +0,0 @@
# immich-server: pinned to talos-cp-3, same reasoning as paperless
# (deployment.yaml comment there) - immich-media is a ReadWriteOnce Longhorn
# volume with a single replica physically on that node's disk (shared with
# paperless-media on the same 4TB HDD). Recreate strategy for the same
# reason: two pods can't both attach an RWO volume.
apiVersion: apps/v1
kind: Deployment
metadata:
name: immich-server
spec:
replicas: 1
strategy:
type: Recreate
selector:
matchLabels:
app: immich-server
template:
metadata:
labels:
app: immich-server
spec:
serviceAccountName: immich
nodeSelector:
kubernetes.io/hostname: talos-cp-3
containers:
- name: immich-server
image: ghcr.io/immich-app/immich-server:release
ports:
- containerPort: 2283
envFrom:
- configMapRef:
name: immich-config
env:
- name: DB_USERNAME
valueFrom:
secretKeyRef:
name: immich-db-app
key: username
- name: DB_PASSWORD
valueFrom:
secretKeyRef:
name: immich-db-app
key: password
# Composed by k8s/infra/iam's provisioning script (system-config
# JSON, oauth section) - see immich-oidc Secret.
- name: IMMICH_CONFIG_FILE
value: /config/immich.json
resources:
requests: { cpu: "500m", memory: "1Gi" }
limits: { cpu: "2", memory: "4Gi" }
volumeMounts:
- name: media
mountPath: /usr/src/app/upload
- name: oidc-config
mountPath: /config
readOnly: true
volumes:
- name: media
persistentVolumeClaim:
claimName: immich-media
- name: oidc-config
secret:
secretName: immich-oidc
items:
- key: config.json
path: immich.json
---
# CPU-only for now - the cluster's one GPU node (worker-1) is already
# dedicated to llm-serving predictors. Not node-pinned: its cache PVC is on
# the default 3-replica pool, not the single-disk cp-3 HDD.
apiVersion: apps/v1
kind: Deployment
metadata:
name: immich-machine-learning
spec:
replicas: 1
selector:
matchLabels:
app: immich-machine-learning
template:
metadata:
labels:
app: immich-machine-learning
spec:
serviceAccountName: immich
containers:
- name: immich-machine-learning
image: ghcr.io/immich-app/immich-machine-learning:release
ports:
- containerPort: 3003
resources:
requests: { cpu: "500m", memory: "1Gi" }
limits: { cpu: "2", memory: "4Gi" }
volumeMounts:
- name: ml-cache
mountPath: /cache
volumes:
- name: ml-cache
persistentVolumeClaim:
claimName: immich-ml-cache
-24
View File
@@ -1,24 +0,0 @@
# Direct nginx ingress, same reasoning as paperless: large uploads (photos/
# videos) and long-lived operations (video transcode, big batch uploads) need
# proxy-body-size/timeouts raised past nginx's defaults.
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: immich
annotations:
nginx.ingress.kubernetes.io/proxy-body-size: "0"
nginx.ingress.kubernetes.io/proxy-read-timeout: "600"
nginx.ingress.kubernetes.io/proxy-send-timeout: "600"
spec:
ingressClassName: nginx
rules:
- host: img.riotpiao.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: immich-server
port:
number: 2283
-14
View File
@@ -1,14 +0,0 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
namespace: immich
resources:
- db.yaml
- pvc.yaml
- configmap.yaml
- redis.yaml
- deployment.yaml
- service.yaml
- ingress.yaml
- rbac.yaml
# immich-oidc Secret written by the PostSync provisioning Job in
# k8s/infra/iam (same as paperless-oidc) - not duplicated here.
-38
View File
@@ -1,38 +0,0 @@
# Two volumes:
#
# - media: original photos/videos + generated thumbnails/encoded videos.
# Shares the cp-3 USB HDD with paperless-media, same StorageClass/disk tag,
# single replica (single disk, no redundancy possible - same tradeoff
# paperless already accepts). Sized 1400Gi, not 2000Gi: the disk's real
# usable capacity (~3724GiB, formatting overhead) minus paperless-media's
# 2000Gi and ~231GiB of other apps' default-class replicas that Longhorn
# placed here anyway (disk tags only pull matching volumes in, they don't
# exclude non-matching ones when the untagged pool elsewhere is full) only
# leaves ~1493Gi of real scheduling headroom right now.
# - ml-cache: downloaded ML model weights for immich-machine-learning
# (face detection / CLIP embeddings). Small, disposable (re-downloads on
# loss), but persisted so a pod restart doesn't re-pull multi-GB models -
# default 3-replica pool, not node-pinned.
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: immich-media
spec:
accessModes:
- ReadWriteOnce
storageClassName: longhorn-paperless-media
resources:
requests:
storage: 1400Gi
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: immich-ml-cache
spec:
accessModes:
- ReadWriteOnce
storageClassName: longhorn
resources:
requests:
storage: 5Gi
-43
View File
@@ -1,43 +0,0 @@
# Scoped operator access for immich-admins: restart/config-edit rights on
# just this service's own resources, nothing CNPG-managed (immich-db-*) or
# provisioning-managed (immich-oidc). Same pattern as
# k8s/apps/paperless/rbac.yaml. Inert until kube-apiserver's OIDC wiring
# lands (--oidc-groups-claim=groups, --oidc-groups-prefix=oidc:).
apiVersion: v1
kind: ServiceAccount
metadata:
name: immich
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: immich-operator
rules:
- apiGroups: ["apps"]
resources: ["deployments"]
resourceNames: ["immich-server", "immich-machine-learning"]
verbs: ["get", "list", "watch", "update", "patch"]
- apiGroups: [""]
resources: ["configmaps"]
resourceNames: ["immich-config"]
verbs: ["get", "list", "watch", "update", "patch"]
- apiGroups: [""]
resources: ["secrets"]
resourceNames: ["immich-oidc"]
verbs: ["get", "list", "watch", "update", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: immich-admins-binding
subjects:
- kind: Group
name: "oidc:immich-admins"
apiGroup: rbac.authorization.k8s.io
- kind: ServiceAccount
name: immich
namespace: immich
roleRef:
kind: Role
name: immich-operator
apiGroup: rbac.authorization.k8s.io
-37
View File
@@ -1,37 +0,0 @@
# Job queue broker for immich-server. No PVC: queue state is disposable - a
# lost queue on restart just re-triggers the affected background jobs
# (thumbnail generation, ML jobs, etc.), no photo data loss since originals
# live on immich-media.
apiVersion: apps/v1
kind: Deployment
metadata:
name: immich-redis
spec:
replicas: 1
selector:
matchLabels:
app: immich-redis
template:
metadata:
labels:
app: immich-redis
spec:
containers:
- name: redis
image: redis:7-alpine
ports:
- containerPort: 6379
resources:
requests: { cpu: "50m", memory: "64Mi" }
limits: { cpu: "250m", memory: "256Mi" }
---
apiVersion: v1
kind: Service
metadata:
name: immich-redis
spec:
selector:
app: immich-redis
ports:
- port: 6379
targetPort: 6379
-21
View File
@@ -1,21 +0,0 @@
apiVersion: v1
kind: Service
metadata:
name: immich-server
spec:
selector:
app: immich-server
ports:
- port: 2283
targetPort: 2283
---
apiVersion: v1
kind: Service
metadata:
name: immich-machine-learning
spec:
selector:
app: immich-machine-learning
ports:
- port: 3003
targetPort: 3003
-1
View File
@@ -14,6 +14,5 @@ resources:
- ornith.yaml
- reasoning.yaml
- reranker.yaml
- networkpolicy.yaml
# No namespace transformer: every file sets its own, and the transformer would
# rewrite metadata.namespace on anything cross-namespace added later.
-62
View File
@@ -1,62 +0,0 @@
# NetworkPolicy for LLM inference engines (llm-serving namespace).
#
# These pods have NO auth — vLLM, Ollama, and TEI accept any request.
# All access MUST go through the api-gateway, which validates JWTs and
# injects identity headers (X-Forwarded-User, X-Auth-Verified).
#
# Replaces the hand-applied llm-serving-default-deny policy that used
# `llm-client: "true"` pod label as a selector — any pod in any namespace
# could self-grant access by adding that label, which defeats the purpose.
#
# This policy restricts ingress to:
# 1. api namespace (gateway) — the sole entry point for inference
# 2. monitoring namespace — Prometheus scraping vLLM/TEI /metrics
# 3. intra-namespace — pod-to-pod (future: multi-replica comms)
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: llm-serving-ingress
namespace: llm-serving
labels:
app.kubernetes.io/part-of: llm-serving
spec:
podSelector:
matchLabels:
app.kubernetes.io/part-of: llm-serving
policyTypes:
- Ingress
ingress:
# Allow from api-gateway (namespace: api)
# Gateway proxies /v1/chat/completions, /v1/embeddings, /v1/rerank
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: api
ports:
- protocol: TCP
port: 8080 # vLLM, Ollama HTTP
- protocol: TCP
port: 80 # KServe predictor services
- protocol: TCP
port: 8000 # vLLM direct (some configs)
- protocol: TCP
port: 11434 # Ollama native port
# Allow Prometheus scraping from monitoring namespace
# vLLM: :8080/metrics, TEI: :9000/metrics
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: monitoring
ports:
- protocol: TCP
port: 8080
- protocol: TCP
port: 9000
# Allow intra-namespace (pod-to-pod within llm-serving)
- from:
- podSelector:
matchLabels:
app.kubernetes.io/part-of: llm-serving
ports:
- protocol: TCP
port: 8080
@@ -13,7 +13,6 @@ spec:
labels:
app: management-service
spec:
serviceAccountName: kmsvc
topologySpreadConstraints:
- maxSkew: 1
topologyKey: kubernetes.io/hostname
@@ -2,7 +2,7 @@ namespace: sqs
replicaCount: 3
image:
repository: forgejo.riotpiao.com/rock/kmsvc-manage
repository: ghcr.io/riotpiaole/kmsvc-management-service
tag: latest
pullPolicy: Always
@@ -1,6 +0,0 @@
apiVersion: v2
name: memory-queues
description: Kafka queues (DLQ) for Poimen Memory service (Phase 6.6)
type: application
version: 0.1.0
appVersion: "1.0"
@@ -1,20 +0,0 @@
{{- range .Values.queues }}
---
apiVersion: kmsvc.io/v1alpha1
kind: Queue
metadata:
name: {{ .name }}
namespace: {{ $.Values.namespace }}
labels:
app: memory-service
queue: dlq
spec:
name: {{ .name }}
description: {{ .description }}
partitions: {{ .partitions }}
replicationFactor: {{ .replicationFactor }}
config:
retention.ms: "{{ .config.retention.ms }}"
message.retention.seconds: "{{ .config.message.retention.seconds }}"
visibility.timeout.seconds: "{{ .config.visibility.timeout.seconds }}"
{{- end }}
@@ -1,25 +0,0 @@
# Poimen Memory Service Kafka Queues (kmsvc)
# Phase 6.6: DLQ topics for webhook + metrics failures
queues:
# DLQ for extraction, webhook, and agent failures
- name: poimen-memory-dlq
description: "DLQ for extraction, webhook, and agent failures"
partitions: 3
replicationFactor: 1
config:
retention.ms: "1209600000" # 14 days
message.retention.seconds: "1209600"
visibility.timeout.seconds: "300"
# DLQ for metrics persistence failures
- name: poimen-memory-metric-dlq
description: "DLQ for metrics persistence failures"
partitions: 3
replicationFactor: 1
config:
retention.ms: "1209600000" # 14 days
message.retention.seconds: "1209600"
visibility.timeout.seconds: "300"
namespace: sqs
+1 -1
View File
@@ -1,7 +1,7 @@
namespace: sqs
image:
repository: forgejo.riotpiao.com/rock/kmsvc-manage
repository: ghcr.io/riotpiaole/kmsvc-management-service
tag: latest
pullPolicy: Always
-95
View File
@@ -1,95 +0,0 @@
# Overrides paperless-ngx's own paperless/adapter.py at the same import path
# (mounted via subPath in deployment.yaml) - settings.py hardcodes
# SOCIALACCOUNT_ADAPTER = "paperless.adapter.CustomSocialAccountAdapter", so
# no Django setting needs to change, just the file content underneath it.
#
# Stock CustomSocialAccountAdapter.populate_user() is a stub ("kept in case
# global default permissions are implemented in the future" - they aren't),
# so every OIDC signup lands with zero permissions and 403s on every API
# endpoint. This adds the actual mapping: Authentik's "permissions" claim
# (via the permissions scope, requested in PAPERLESS_SOCIALACCOUNT_PROVIDERS,
# computed server-side from group membership by authentik-provision.py) ->
# "paperless:write" or "*" (homelab-admins) grants is_staff+is_superuser,
# same convention already used for MinIO's policy claim and Grafana's
# role_attribute_path. Checking the permission string rather than a literal
# group name decouples "what grants access" from which group happens to
# hold it - same pattern applies to every other service's Role/RoleBinding
# in k8s/infra/rbac/.
apiVersion: v1
kind: ConfigMap
metadata:
name: paperless-adapter
data:
adapter.py: |
from urllib.parse import quote
from allauth.account.adapter import DefaultAccountAdapter
from allauth.core import context
from allauth.socialaccount.adapter import DefaultSocialAccountAdapter
from django.conf import settings
from django.forms import ValidationError
from django.urls import reverse
REQUIRED_PERMISSIONS = {"paperless:write", "*"}
class CustomAccountAdapter(DefaultAccountAdapter):
def is_open_for_signup(self, request):
allow_signups = super().is_open_for_signup(request)
return getattr(settings, "ACCOUNT_ALLOW_SIGNUPS", allow_signups)
def pre_authenticate(self, request, **credentials):
if settings.DISABLE_REGULAR_LOGIN:
raise ValidationError("Regular login is disabled")
return super().pre_authenticate(request, **credentials)
def is_safe_url(self, url):
from django.utils.http import url_has_allowed_host_and_scheme
allowed_hosts = {context.request.get_host()} | set(settings.ALLOWED_HOSTS)
if "*" in allowed_hosts:
allowed_hosts.remove("*")
allowed_hosts.add(context.request.get_host())
return url_has_allowed_host_and_scheme(url, allowed_hosts=allowed_hosts)
return url_has_allowed_host_and_scheme(url, allowed_hosts=allowed_hosts)
def get_reset_password_from_key_url(self, key):
if settings.PAPERLESS_URL is None:
return super().get_reset_password_from_key_url(key)
path = reverse(
"account_reset_password_from_key",
kwargs={"uidb36": "UID", "key": "KEY"},
)
path = path.replace("UID-KEY", quote(key))
return settings.PAPERLESS_URL + path
class CustomSocialAccountAdapter(DefaultSocialAccountAdapter):
def is_open_for_signup(self, request, sociallogin):
allow_signups = super().is_open_for_signup(request, sociallogin)
return getattr(settings, "SOCIALACCOUNT_ALLOW_SIGNUPS", allow_signups)
def get_connect_redirect_url(self, request, socialaccount):
return reverse("base")
def populate_user(self, request, sociallogin, data):
user = super().populate_user(request, sociallogin, data)
perms = set(sociallogin.account.extra_data.get("permissions") or [])
if perms & REQUIRED_PERMISSIONS:
user.is_staff = True
user.is_superuser = True
return user
def save_user(self, request, sociallogin, form=None):
# populate_user() sets the flags on the in-memory user, but
# allauth's default save_user() re-derives is_staff from
# ACCOUNT_DEFAULT_HTTP_PROTOCOL-independent defaults and can
# overwrite them on save - re-apply after super().save_user()
# persists the row, matching the permissions check above exactly.
user = super().save_user(request, sociallogin, form)
perms = set(sociallogin.account.extra_data.get("permissions") or [])
if perms & REQUIRED_PERMISSIONS and not (user.is_staff and user.is_superuser):
user.is_staff = True
user.is_superuser = True
user.save(update_fields=["is_staff", "is_superuser"])
return user
-94
View File
@@ -1,94 +0,0 @@
# Nightly: pg_dump the paperless DB + mirror the media PVC into the scoped
# `paperless` MinIO bucket (see minio-provision-paperless-job.yaml). This is a
# BACKUP target, not live storage - paperless-ngx has no native S3 backend, it
# only ever reads/writes the local media PVC directly.
#
# Pinned to talos-cp-3, same as deployment.yaml: media is a ReadWriteOnce
# Longhorn volume with a single replica physically on that node's disk -
# mounting it read-only here from a different node would conflict with the
# live webserver's attachment.
apiVersion: batch/v1
kind: CronJob
metadata:
name: paperless-backup
spec:
schedule: "0 3 * * *" # 03:00 daily, low-traffic window
jobTemplate:
spec:
backoffLimit: 2
template:
spec:
restartPolicy: Never
nodeSelector:
kubernetes.io/hostname: talos-cp-3
initContainers:
- name: pg-dump
image: postgres:16-alpine
env:
- name: PGHOST
value: paperless-db-rw
- name: PGDATABASE
value: paperless
- name: PGUSER
valueFrom:
secretKeyRef:
name: paperless-db-app
key: username
- name: PGPASSWORD
valueFrom:
secretKeyRef:
name: paperless-db-app
key: password
command:
- sh
- -c
- pg_dump --format=custom --file=/backup/paperless-db.dump
volumeMounts:
- name: backup
mountPath: /backup
containers:
- name: mc-mirror
image: minio/mc:latest
env:
- name: ACCESS_KEY
valueFrom:
secretKeyRef:
name: paperless-minio-creds
key: ACCESS_KEY
- name: SECRET_KEY
valueFrom:
secretKeyRef:
name: paperless-minio-creds
key: SECRET_KEY
- name: BUCKET
valueFrom:
secretKeyRef:
name: paperless-minio-creds
key: BUCKET
- name: ENDPOINT
valueFrom:
secretKeyRef:
name: paperless-minio-creds
key: ENDPOINT
command:
- /bin/sh
- -c
- |
set -e
mc alias set b "$ENDPOINT" "$ACCESS_KEY" "$SECRET_KEY"
mc cp /backup/paperless-db.dump "b/$BUCKET/db/paperless-db-$(date +%Y%m%d).dump"
mc mirror --overwrite /media "b/$BUCKET/media"
echo "Backup done."
volumeMounts:
- name: backup
mountPath: /backup
- name: media
mountPath: /media
readOnly: true
volumes:
- name: backup
emptyDir: {}
- name: media
persistentVolumeClaim:
claimName: paperless-media
readOnly: true
-22
View File
@@ -1,22 +0,0 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: paperless-config
data:
PAPERLESS_URL: "https://paperless.riotpiao.com"
PAPERLESS_TIME_ZONE: "America/Los_Angeles"
PAPERLESS_OCR_LANGUAGE: "eng"
PAPERLESS_DBHOST: "paperless-db-rw"
PAPERLESS_DBNAME: "paperless"
PAPERLESS_REDIS: "redis://paperless-redis:6379"
# django-allauth generic OIDC provider. The client_id/secret/server_url
# bundle itself lives in the paperless-oidc Secret
# (SOCIALACCOUNT_PROVIDERS_JSON key, composed by authentik-provision.py) -
# env vars can't be split across a ConfigMap + Secret for the same key, so
# this whole value is sourced from the Secret in deployment.yaml instead.
PAPERLESS_APPS: "allauth.socialaccount.providers.openid_connect"
# Authentik already verifies identity via OIDC - a second email-confirmation
# step has no SMTP configured to send it anyway, and paperless-ngx doesn't
# wire up allauth's confirm-email view, so signup 500s with NoReverseMatch
# on 'account_confirm_email' without this.
PAPERLESS_ACCOUNT_EMAIL_VERIFICATION: "none"
-103
View File
@@ -1,103 +0,0 @@
# Single container runs webserver + consumer + scheduler (paperless-ngx's
# stock entrypoint does this internally) - no need to split into separate
# Deployments. replicas: 1 only: paperless-media is ReadWriteOnce, and the
# consumer polling the media dir doesn't benefit from horizontal scaling here.
#
# Pinned to talos-cp-3: paperless-media's disk physically lives there. Longhorn
# RWO volumes can only be attached from one node at a time, and the nightly
# backup-cronjob.yaml also mounts this same PVC (read-only) to mirror it into
# MinIO - pinning both to the same node avoids a cross-node attach conflict,
# and keeps the 3.5Ti read/write path off the network entirely.
apiVersion: apps/v1
kind: Deployment
metadata:
name: paperless
spec:
replicas: 1
strategy:
type: Recreate # ReadWriteOnce media PVC - avoid two pods fighting over it
selector:
matchLabels:
app: paperless
template:
metadata:
labels:
app: paperless
spec:
# Kubernetes injects legacy Docker-links env vars for every Service in
# this namespace (<SVC>_SERVICE_HOST, <SVC>_PORT, ...). The Service here
# is named "paperless", so that becomes PAPERLESS_PORT=tcp://<ip>:8000 -
# paperless-ngx's own entrypoint reads PAPERLESS_PORT for gunicorn's
# bind address, collides, and gunicorn crash-loops on "not a valid port
# number". Disable the injection instead of renaming the Service.
enableServiceLinks: false
nodeSelector:
kubernetes.io/hostname: talos-cp-3
containers:
- name: paperless
image: ghcr.io/paperless-ngx/paperless-ngx:2.20.15
ports:
- containerPort: 8000
envFrom:
- configMapRef:
name: paperless-config
env:
- name: PAPERLESS_DBUSER
valueFrom:
secretKeyRef:
name: paperless-db-app
key: username
- name: PAPERLESS_DBPASS
valueFrom:
secretKeyRef:
name: paperless-db-app
key: password
- name: PAPERLESS_SECRET_KEY
valueFrom:
secretKeyRef:
name: paperless-secrets
key: PAPERLESS_SECRET_KEY
- name: PAPERLESS_ADMIN_USER
valueFrom:
secretKeyRef:
name: paperless-secrets
key: PAPERLESS_ADMIN_USER
- name: PAPERLESS_ADMIN_PASSWORD
valueFrom:
secretKeyRef:
name: paperless-secrets
key: PAPERLESS_ADMIN_PASSWORD
- name: PAPERLESS_SOCIALACCOUNT_PROVIDERS
valueFrom:
secretKeyRef:
name: paperless-oidc
key: SOCIALACCOUNT_PROVIDERS_JSON
resources:
requests: { cpu: "500m", memory: "1Gi" }
limits: { cpu: "2", memory: "4Gi" }
volumeMounts:
- name: media
mountPath: /usr/src/paperless/media
- name: data
mountPath: /usr/src/paperless/data
- name: consume
mountPath: /usr/src/paperless/consume
# Overrides paperless-ngx's own adapter.py in place - settings.py
# hardcodes the import path, so no Django setting changes, just
# the file content underneath it (see adapter-configmap.yaml).
- name: adapter
mountPath: /usr/src/paperless/src/paperless/adapter.py
subPath: adapter.py
readOnly: true
volumes:
- name: media
persistentVolumeClaim:
claimName: paperless-media
- name: data
persistentVolumeClaim:
claimName: paperless-data
- name: consume
emptyDir: {}
- name: adapter
configMap:
name: paperless-adapter
-24
View File
@@ -1,24 +0,0 @@
# Direct nginx ingress to the paperless Service - not routed via the Go
# api-gateway (api.riotpiao.com), which has no WebSocket upgrade support and
# paperless-ngx keeps a long-lived /ws/ connection open for live task status.
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: paperless
annotations:
nginx.ingress.kubernetes.io/proxy-body-size: "0" # large scanned PDF uploads
nginx.ingress.kubernetes.io/proxy-read-timeout: "600"
nginx.ingress.kubernetes.io/proxy-send-timeout: "600"
spec:
ingressClassName: nginx
rules:
- host: paperless.riotpiao.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: paperless
port:
number: 8000
-17
View File
@@ -1,17 +0,0 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
namespace: paperless
resources:
- pvc.yaml
- configmap.yaml
- redis.yaml
- deployment.yaml
- service.yaml
- ingress.yaml
- backup-cronjob.yaml
- adapter-configmap.yaml
- rbac.yaml
# postgres: paperless-db CNPG Cluster, deployed by k8s/infra/databases (wave 2,
# before this app at wave 8) - not duplicated here. Same for the paperless-oidc
# and paperless-minio-creds Secrets, written by PostSync provisioning Jobs in
# k8s/infra/iam and k8s/infra/minio respectively.
-35
View File
@@ -1,35 +0,0 @@
# Two volumes, deliberately separate storage classes:
#
# - media: the actual documents (originals + OCR'd archive PDFs + thumbnails).
# Lives on the cp-3 USB HDD, single replica (see
# k8s/infra/longhorn/longhorn-paperless-storageclass.yaml). Shares the disk
# with Immich's immich-media PVC (k8s/apps/immich/pvc.yaml, 2000Gi) - photo
# libraries grow much faster than scanned documents, so paperless gets the
# smaller 500Gi share.
# - data: the SQLite classification model + search index. Small (low GB),
# frequently rewritten, and disposable (rebuilds from the DB + media on
# next consume) - stays on the default 3-replica pool instead of the
# single-disk HDD.
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: paperless-media
spec:
accessModes:
- ReadWriteOnce
storageClassName: longhorn-paperless-media
resources:
requests:
storage: 500Gi
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: paperless-data
spec:
accessModes:
- ReadWriteOnce
storageClassName: longhorn
resources:
requests:
storage: 5Gi
-35
View File
@@ -1,35 +0,0 @@
# Scoped operator access for paperless-admins: restart/config-edit rights on
# just this service's own resources, nothing CNPG-managed (paperless-db-*)
# or provisioning-managed (paperless-oidc, paperless-minio-creds). Inert
# until kube-apiserver's OIDC wiring lands (--oidc-groups-claim=groups,
# --oidc-groups-prefix=oidc:) - subject name below assumes that prefix.
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: paperless-operator
rules:
- apiGroups: ["apps"]
resources: ["deployments"]
resourceNames: ["paperless"]
verbs: ["get", "list", "watch", "update", "patch"]
- apiGroups: [""]
resources: ["configmaps"]
resourceNames: ["paperless-config"]
verbs: ["get", "list", "watch", "update", "patch"]
- apiGroups: [""]
resources: ["secrets"]
resourceNames: ["paperless-secrets"]
verbs: ["get", "list", "watch", "update", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: paperless-admins-binding
subjects:
- kind: Group
name: "oidc:paperless-admins"
apiGroup: rbac.authorization.k8s.io
roleRef:
kind: Role
name: paperless-operator
apiGroup: rbac.authorization.k8s.io
-37
View File
@@ -1,37 +0,0 @@
# Task queue broker + websocket channel layer for paperless-ngx. No PVC:
# queued/scheduled task state is disposable - a lost queue on restart just
# means re-triggering consumption, not data loss (documents themselves live
# on paperless-media).
apiVersion: apps/v1
kind: Deployment
metadata:
name: paperless-redis
spec:
replicas: 1
selector:
matchLabels:
app: paperless-redis
template:
metadata:
labels:
app: paperless-redis
spec:
containers:
- name: redis
image: redis:7-alpine
ports:
- containerPort: 6379
resources:
requests: { cpu: "50m", memory: "64Mi" }
limits: { cpu: "250m", memory: "256Mi" }
---
apiVersion: v1
kind: Service
metadata:
name: paperless-redis
spec:
selector:
app: paperless-redis
ports:
- port: 6379
targetPort: 6379
-10
View File
@@ -1,10 +0,0 @@
apiVersion: v1
kind: Service
metadata:
name: paperless
spec:
selector:
app: paperless
ports:
- port: 8000
targetPort: 8000
@@ -1,144 +0,0 @@
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
name: secretrotations.homelab.riotpiao.com
spec:
group: homelab.riotpiao.com
names:
kind: SecretRotation
plural: secretrotations
scope: Namespaced
versions:
- name: v1
served: true
storage: true
schema:
openAPIV3Schema:
type: object
properties:
metadata:
type: object
spec:
type: object
required:
- provider
- rotationInterval
properties:
# External system: authentik | forgejo | minio | vault
provider:
type: string
enum: [authentik, forgejo, minio, vault]
# How often to rotate (hours)
rotationInterval:
type: integer
minimum: 24
# Application ID in external system
appId:
type: string
# k8s Secret to update (name, namespace, key)
secretRef:
type: object
required: [name, namespace]
properties:
name:
type: string
namespace:
type: string
key:
type: string
description: "Secret key to update (e.g., MINIO_IDENTITY_OPENID_CLIENT_SECRET)"
# Path to git file that holds the secret (for .enc.yaml files)
gitPath:
type: string
description: "Path in homelab repo to .enc.yaml file"
# Ansible template values to substitute
templateValues:
type: object
additionalProperties:
type: string
status:
type: object
properties:
lastRotationTime:
type: string
format: date-time
nextRotationTime:
type: string
format: date-time
lastRotationStatus:
type: string
enum: [Success, Failed, Pending]
lastRotationError:
type: string
lastCommitHash:
type: string
---
# Example usage:
apiVersion: homelab.riotpiao.com/v1
kind: SecretRotation
metadata:
name: minio-oidc
namespace: secret-rotation
spec:
provider: authentik
rotationInterval: 2160 # 90 days in hours
appId: minio
secretRef:
name: minio-oidc
namespace: storage
key: MINIO_IDENTITY_OPENID_CLIENT_SECRET
gitPath: k8s/argocd/secrets/minio-oidc.enc.yaml
---
apiVersion: homelab.riotpiao.com/v1
kind: SecretRotation
metadata:
name: portfolio-agent-oidc
namespace: secret-rotation
spec:
provider: authentik
rotationInterval: 2160
appId: portfolio-agent
secretRef:
name: portfolio-agent-oidc
namespace: portfolio
key: CLIENT_SECRET
gitPath: k8s/argocd/secrets/portfolio-agent-oidc.enc.yaml
---
apiVersion: homelab.riotpiao.com/v1
kind: SecretRotation
metadata:
name: forgejo-registry-token
namespace: secret-rotation
spec:
provider: forgejo
rotationInterval: 2160
appId: rock/riotpiao.com
secretRef:
name: forgejo-registry-secret
namespace: kube-system
key: REGISTRY_TOKEN
gitPath: k8s/argocd/secrets/forgejo-registry-secret.enc.yaml
---
apiVersion: homelab.riotpiao.com/v1
kind: SecretRotation
metadata:
name: minio-root-credentials
namespace: secret-rotation
spec:
provider: minio
rotationInterval: 4320 # 180 days in hours
appId: root
secretRef:
name: minio-creds
namespace: storage
gitPath: k8s/argocd/secrets/minio-secrets.enc.yaml
@@ -1,92 +0,0 @@
apiVersion: apps/v1
kind: Deployment
metadata:
name: secret-rotation-controller
namespace: secret-rotation
spec:
replicas: 1
selector:
matchLabels:
app: secret-rotation-controller
template:
metadata:
labels:
app: secret-rotation-controller
spec:
serviceAccountName: secret-rotation-controller
containers:
- name: controller
image: secret-rotation-controller:latest
imagePullPolicy: IfNotPresent
env:
# SOPS reads age key from this file
- name: SOPS_AGE_KEY_FILE
value: /etc/sops/age/private-key.txt
# Vault auth (token in projected volume)
- name: VAULT_ADDR
value: http://vault.vault.svc.cluster.local:8200
- name: VAULT_TOKEN_FILE
value: /var/run/secrets/vault/token
# Authentik
- name: AUTHENTIK_URL
value: http://authentik-server.iam.svc.cluster.local
- name: AUTHENTIK_BOOTSTRAP_TOKEN
valueFrom:
secretKeyRef:
name: authentik-bootstrap
key: token
# Git
- name: GIT_REPO
value: https://forgejo.riotpiao.com/rock/homelab.git
- name: GIT_AUTHOR_EMAIL
value: [email protected]
- name: GIT_AUTHOR_NAME
value: Secret Rotation Controller
- name: FORGEJO_TOKEN
valueFrom:
secretKeyRef:
name: forgejo-registry-secret
key: REGISTRY_TOKEN
volumeMounts:
# Age key from ExternalSecret (synced from Vault)
- name: age-key
mountPath: /etc/sops/age
readOnly: true
# Vault auth token (projected)
- name: vault-token
mountPath: /var/run/secrets/vault
readOnly: true
# Temp working dir
- name: tmp
mountPath: /tmp
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: 500m
memory: 512Mi
volumes:
- name: age-key
secret:
secretName: sops-age-key
defaultMode: 0400
- name: vault-token
projected:
sources:
- serviceAccountToken:
path: token
audience: vault
expirationSeconds: 3600
- name: tmp
emptyDir: {}
@@ -1,15 +0,0 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
namespace: secret-rotation
resources:
- rbac.yaml
- crd.yaml
- external-secret.yaml
- deployment.yaml
commonLabels:
app.kubernetes.io/name: secret-rotation-controller
app.kubernetes.io/component: automation
managed-by: argocd
@@ -1,53 +0,0 @@
apiVersion: v1
kind: ServiceAccount
metadata:
name: secret-rotation-controller
namespace: secret-rotation
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: secret-rotation-controller
rules:
# Read SecretRotation CRDs
- apiGroups: ["homelab.riotpiao.com"]
resources: ["secretrotations"]
verbs: ["get", "list", "watch"]
# Update status
- apiGroups: ["homelab.riotpiao.com"]
resources: ["secretrotations/status"]
verbs: ["get", "patch", "update"]
# Read k8s secrets that will be rotated
- apiGroups: [""]
resources: ["secrets"]
verbs: ["get", "list"]
# For recording events
- apiGroups: [""]
resources: ["events"]
verbs: ["create", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: secret-rotation-controller
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: secret-rotation-controller
subjects:
- kind: ServiceAccount
name: secret-rotation-controller
namespace: secret-rotation
---
apiVersion: v1
kind: Namespace
metadata:
name: secret-rotation
labels:
kubernetes.io/metadata.name: secret-rotation
+1 -1
View File
@@ -18,7 +18,7 @@ spec:
prune: true
selfHeal: true
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/argocd/projects
destination:
+1 -1
View File
@@ -17,7 +17,7 @@ spec:
syncOptions:
- CreateNamespace=true
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
# ksops decrypts every *.enc.yaml here at kustomize-build time (repo-server
# runs `kustomize build --enable-alpha-plugins --enable-exec`). Replaces the
+3 -75
View File
@@ -21,7 +21,7 @@ spec:
helm:
valueFiles:
- $values/k8s/bootstrap/cert-manager/cert-manager-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
@@ -93,7 +93,7 @@ spec:
project: homelab
revisionHistoryLimit: 3
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
# A real kustomization.yaml (resources: the 3 issuer/CA files) renders these
# deterministically. The previous directory.include with bare filenames
@@ -127,7 +127,7 @@ spec:
project: homelab
revisionHistoryLimit: 3
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/bootstrap/ingress
destination:
@@ -138,75 +138,3 @@ spec:
selfHeal: true
syncOptions:
- CreateNamespace=true
---
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: cluster-maintenance
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "0"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/infra/cluster-maintenance
destination:
server: https://kubernetes.default.svc
namespace: kube-system
syncPolicy:
automated:
prune: true
selfHeal: true
---
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: kyverno
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "0"
spec:
project: homelab
source:
repoURL: https://kyverno.github.io/kyverno/
chart: kyverno
targetRevision: "1.14.0"
helm:
valueFiles:
- $values/k8s/bootstrap/kyverno/kyverno-values.yaml
sources:
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
ref: values
destination:
server: https://kubernetes.default.svc
namespace: kyverno
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
---
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: kyverno-policies
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "0"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/bootstrap/kyverno
destination:
server: https://kubernetes.default.svc
namespace: kyverno
syncPolicy:
automated:
prune: true
selfHeal: true
-33
View File
@@ -1,33 +0,0 @@
# ArgoCD Image Updater - auto-updates Application images from registry
# Watches forgejo.riotpiao.com for new image tags and updates Applications
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: argocd-image-updater
namespace: argocd
finalizers:
- resources-finalizer.argocd.argoproj.io
annotations:
argocd.argoproj.io/sync-wave: "1"
spec:
project: homelab
revisionHistoryLimit: 3
sources:
- repoURL: https://argoproj.github.io/argo-helm
chart: argocd-image-updater
targetRevision: "0.11.2"
helm:
valueFiles:
- $values/k8s/infra/argocd-image-updater/values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
ref: values
destination:
server: https://kubernetes.default.svc
namespace: argocd
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=false
-32
View File
@@ -1,32 +0,0 @@
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: secret-rotation
namespace: argocd
labels:
app.kubernetes.io/name: secret-rotation
spec:
project: homelab
sources:
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
path: k8s/apps/secret-rotation-controller
targetRevision: main
destination:
server: https://kubernetes.default.svc
namespace: secret-rotation
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
- RespectIgnoreDifferences=true
retry:
limit: 5
backoff:
duration: 5s
factor: 2
maxDuration: 3m
+7 -33
View File
@@ -17,7 +17,7 @@ spec:
helm:
valueFiles:
- $values/k8s/infra/minio/minio-operator-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
@@ -41,7 +41,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/infra/minio
destination:
@@ -66,7 +66,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/infra/longhorn
destination:
@@ -102,7 +102,7 @@ spec:
skipCrds: true
valueFiles:
- $values/k8s/infra/monitoring/prometheus-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
@@ -152,7 +152,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/infra/monitoring/crds
destination:
@@ -183,7 +183,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/infra/monitoring
destination:
@@ -213,7 +213,7 @@ spec:
helm:
valueFiles:
- $values/k8s/infra/monitoring/blackbox-exporter-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
@@ -223,29 +223,3 @@ spec:
automated:
prune: true
selfHeal: true
---
# Distributed tracing: Tempo + OpenTelemetry Collector.
# Receives traces from instrumented services, stores in local volume (72h retention).
# Grafana datasource auto-configured, service graph + latency dashboards included.
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: tracing
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "1"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/infra/tracing
destination:
server: https://kubernetes.default.svc
namespace: tracing
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
+3 -3
View File
@@ -19,7 +19,7 @@ spec:
helm:
valueFiles:
- $values/k8s/infra/logging/loki-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
@@ -53,7 +53,7 @@ spec:
helm:
valueFiles:
- $values/k8s/infra/logging/grafana-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
@@ -87,7 +87,7 @@ spec:
helm:
valueFiles:
- $values/k8s/infra/logging/promtail-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
+7 -28
View File
@@ -17,7 +17,7 @@ spec:
helm:
valueFiles:
- $values/k8s/infra/iam/vault-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
@@ -46,7 +46,7 @@ spec:
helm:
valueFiles:
- $values/k8s/infra/iam/authentik-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
@@ -68,7 +68,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/infra/iam
destination:
@@ -109,7 +109,7 @@ spec:
helm:
valueFiles:
- $values/k8s/bootstrap/phase3-forgejo/forgejo-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
@@ -152,17 +152,10 @@ metadata:
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "3"
argocd-image-updater.argoproj.io/image-list: runner=forgejo.riotpiao.com/rock/forgejo-runner-golang
argocd-image-updater.argoproj.io/runner.update-strategy: newest-build
argocd-image-updater.argoproj.io/runner.allow-tags: regexp:^[0-9a-f]{7}$|^latest$|^v[0-9]+$
argocd-image-updater.argoproj.io/runner.helm.image-name: runner.image.repository
argocd-image-updater.argoproj.io/runner.helm.image-tag: runner.image.tag
argocd-image-updater.argoproj.io/write-back-method: git
argocd-image-updater.argoproj.io/git-branch: main
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/infra/forgejo-runner
destination:
@@ -180,17 +173,10 @@ metadata:
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "3"
argocd-image-updater.argoproj.io/image-list: runner=forgejo.riotpiao.com/rock/forgejo-runner-node
argocd-image-updater.argoproj.io/runner.update-strategy: newest-build
argocd-image-updater.argoproj.io/runner.allow-tags: regexp:^[0-9a-f]{7}$|^latest$|^v[0-9]+$
argocd-image-updater.argoproj.io/runner.helm.image-name: runner.image.repository
argocd-image-updater.argoproj.io/runner.helm.image-tag: runner.image.tag
argocd-image-updater.argoproj.io/write-back-method: git
argocd-image-updater.argoproj.io/git-branch: main
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/infra/forgejo-runner
helm:
@@ -211,17 +197,10 @@ metadata:
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "3"
argocd-image-updater.argoproj.io/image-list: runner=forgejo.riotpiao.com/rock/forgejo-runner-rust
argocd-image-updater.argoproj.io/runner.update-strategy: newest-build
argocd-image-updater.argoproj.io/runner.allow-tags: regexp:^[0-9a-f]{7}$|^latest$|^v[0-9]+$
argocd-image-updater.argoproj.io/runner.helm.image-name: runner.image.repository
argocd-image-updater.argoproj.io/runner.helm.image-tag: runner.image.tag
argocd-image-updater.argoproj.io/write-back-method: git
argocd-image-updater.argoproj.io/git-branch: main
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/infra/forgejo-runner
helm:
+1 -1
View File
@@ -14,7 +14,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/infra/databases
destination:
-20
View File
@@ -1,20 +0,0 @@
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: memory-queues
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "7"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/apps/messaging/memory-queues
destination:
server: https://kubernetes.default.svc
namespace: sqs
syncPolicy:
automated:
prune: true
selfHeal: true
+46 -6
View File
@@ -1,6 +1,6 @@
# Wave 5 — Kafka (Strimzi operator + cluster CR), Redis infrastructure.
# Strimzi/Redis are public Helm charts; kafka-cluster is a local chart.
# queue-crd and management-service are managed by kmsvc-root (kmsvc-manage.git).
# Wave 5 — Kafka (Strimzi operator + cluster CR), Redis, and the SQS-like
# queue services. Strimzi/Redis are public Helm charts; kafka-cluster/queue-crd/
# management-service are local charts (rendered from their own Chart.yaml).
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
@@ -68,7 +68,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/apps/messaging/kafka-cluster
destination:
@@ -78,5 +78,45 @@ spec:
automated:
prune: true
selfHeal: true
# queue-crd and management-service moved to kmsvc-manage.git repo
# Managed by kmsvc-root Application
---
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: queue-crd
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "6"
spec:
project: homelab
source:
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/apps/messaging/queue-crd
destination:
server: https://kubernetes.default.svc
namespace: sqs
syncPolicy:
automated:
prune: true
selfHeal: true
---
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: management-service
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "7"
spec:
project: homelab
source:
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/apps/messaging/management-service
destination:
server: https://kubernetes.default.svc
namespace: sqs
syncPolicy:
automated:
prune: true
selfHeal: true
+2 -11
View File
@@ -9,12 +9,8 @@
# `POST /v1/chat/completions` covers every model. See
# docs/adr/ADR-0001-retire-kong-for-go-gateway.md in the frontend repo.
#
# UPDATED 2026-08-22: Tracks main branch of homelab-frontend (auto-syncs on each push).
# Image built on every main commit with tag <commit-sha>.
# ArgoCD auto-pulls the latest image (live reconciliation ~3min).
#
# Two sources:
# 1. rock/homelab-frontend on the in-cluster Forgejo (prod branch) — the gateway's own
# 1. rock/homelab-frontend on the in-cluster Forgejo — the gateway's own
# kustomization (Deployment, Service, ConfigMap, RBAC, NetworkPolicy). It
# sets `namespace: api` itself, so no transformer is needed here. The
# Forgejo host must stay listed in the `homelab` AppProject sourceRepos or
@@ -36,11 +32,6 @@ metadata:
app.kubernetes.io/component: gateway
annotations:
argocd.argoproj.io/sync-wave: "7"
# ArgoCD Image Updater - auto-update on new image push
argocd-image-updater.argoproj.io/image-list: gw=forgejo.riotpiao.com/rock/api-gateway
argocd-image-updater.argoproj.io/gw.update-strategy: digest
argocd-image-updater.argoproj.io/gw.allow-tags: regexp:^latest$
argocd-image-updater.argoproj.io/write-back-method: argocd
spec:
project: homelab
revisionHistoryLimit: 3
@@ -48,7 +39,7 @@ spec:
- repoURL: https://forgejo.riotpiao.com/rock/homelab-frontend.git
targetRevision: main
path: k8s
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/apps/api
destination:
+1 -1
View File
@@ -20,7 +20,7 @@ spec:
project: homelab
revisionHistoryLimit: 3
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/apps/llm-serving
destination:
+8 -126
View File
@@ -19,10 +19,10 @@ spec:
helm:
valueFiles:
- $values/k8s/apps/temporal/temporal-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/apps/temporal
destination:
@@ -51,7 +51,7 @@ spec:
helm:
valueFiles:
- $values/k8s/apps/portainer/portainer-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
destination:
@@ -74,7 +74,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/apps/cloudflared
destination:
@@ -97,7 +97,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/apps/agent-pod
destination:
@@ -130,7 +130,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/apps/sms
destination:
@@ -141,67 +141,6 @@ spec:
prune: true
selfHeal: true
---
# Document management. Raw manifests (no Helm): postgres is the dedicated
# paperless-db CNPG cluster in k8s/infra/databases (wave 2), redis is
# in-cluster only (no PVC), media lives on the cp-3 USB HDD (see
# k8s/infra/longhorn/longhorn-paperless-storageclass.yaml). OIDC via
# Authentik provisioned by k8s/infra/iam's PostSync job; MinIO backup bucket
# creds provisioned by k8s/infra/minio's PostSync job.
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: paperless
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "8"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/apps/paperless
destination:
server: https://kubernetes.default.svc
namespace: paperless
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
---
# Photo/video backup. Self-contained (unlike paperless, its CNPG Postgres
# lives here too, not in k8s/infra/databases) - CreateNamespace=true creates
# the namespace before any manifest in this Application applies, including
# the Cluster CR, so no separate wave-2 pre-creation step is needed. Postgres
# is pg18 (not this repo's usual 16.2) because CNPG's official pgvector
# extension image only publishes pg18 builds - see k8s/apps/immich/db.yaml.
# media PVC shares the cp-3 HDD 2TB/2TB with paperless-media. OIDC via
# Authentik provisioned by k8s/infra/iam's PostSync job (immich entry in
# SERVICES + immich_role scope mapping for admin-via-claim).
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: immich
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "8"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/apps/immich
destination:
server: https://kubernetes.default.svc
namespace: immich
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
---
# Consolidated: homarr + homarr-patches → homarr
# Helm chart + values + PostSync hook patch (fix-probes-job.yaml)
apiVersion: argoproj.io/v1alpha1
@@ -220,10 +159,10 @@ spec:
helm:
valueFiles:
- $values/k8s/apps/homarr/homarr-values.yaml
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
ref: values
- repoURL: https://forgejo.riotpiao.com/rock/homelab.git
- repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/apps/homarr # PostSync hook: fix-probes-job.yaml
destination:
@@ -235,60 +174,3 @@ spec:
selfHeal: true
syncOptions:
- CreateNamespace=true
---
# Portfolio site at riotpiao.com - static Next.js site from rock/riotpiao.com repo.
# Points directly to infra/portfolio/base (bypassing repo's own argocd-apps.yaml
# which has wrong URLs). Image built by Forgejo Actions on rock/portfolio repo.
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: portfolio
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "8"
# ArgoCD Image Updater - auto-update on new image push
argocd-image-updater.argoproj.io/image-list: app=forgejo.riotpiao.com/rock/portfolio
argocd-image-updater.argoproj.io/app.update-strategy: digest
argocd-image-updater.argoproj.io/app.allow-tags: regexp:^latest$
argocd-image-updater.argoproj.io/write-back-method: argocd
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/riotpiao.com.git
targetRevision: main
path: infra/portfolio/base
destination:
server: https://kubernetes.default.svc
namespace: portfolio
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
---
# Wave 9 - per-service scoped RBAC (Role/RoleBinding), deliberately last so
# every target namespace above already exists. Inert until kube-apiserver
# gets --oidc-groups-claim=groups wired up (separate, not-yet-applied
# terraform/talosctl change) - these grant nothing until then.
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: rbac
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "9"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
targetRevision: main
path: k8s/infra/rbac
destination:
server: https://kubernetes.default.svc
# No namespace: cluster-scoped resources (ClusterRoleBinding, etc.)
# Namespace is set per-resource in kustomization
syncPolicy:
automated:
prune: true
selfHeal: true
-28
View File
@@ -1,28 +0,0 @@
# kmsvc-manage bootstrap — manages itself and its supporting services
# (Strimzi/Kafka, Redis, queue-operator, message-plane server) from
# the kmsvc-manage repo's own k8s/argocd/ structure on the main branch.
# Image built on every main commit, auto-deployed to sqs namespace.
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: kmsvc-root
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "6"
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/kmsvc-manage.git
targetRevision: main
path: k8s/argocd/apps
directory:
recurse: false
destination:
server: https://kubernetes.default.svc
namespace: sqs
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
-46
View File
@@ -1,46 +0,0 @@
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: poimen
namespace: argocd
annotations:
argocd.argoproj.io/sync-wave: "7"
# Image Updater: auto-update on new image push (SHA tag filter)
argocd-image-updater.argoproj.io/image-list: |
memory=forgejo.riotpiao.com/rock/poimen-memory
workflows=forgejo.riotpiao.com/rock/poimen-workflows
frontend=forgejo.riotpiao.com/rock/poimen-frontend
argocd-image-updater.argoproj.io/memory.update-strategy: digest
argocd-image-updater.argoproj.io/memory.allow-tags: regexp:^latest$
argocd-image-updater.argoproj.io/workflows.update-strategy: digest
argocd-image-updater.argoproj.io/workflows.allow-tags: regexp:^latest$
argocd-image-updater.argoproj.io/frontend.update-strategy: digest
argocd-image-updater.argoproj.io/frontend.allow-tags: regexp:^latest$
argocd-image-updater.argoproj.io/write-back-method: argocd
spec:
project: homelab
sources:
- repoURL: https://forgejo.riotpiao.com/rock/poimen-memory.git
targetRevision: main
path: k8s/argocd
- repoURL: https://forgejo.riotpiao.com/rock/poimen-workflows.git
targetRevision: main
path: k8s/argocd
- repoURL: https://forgejo.riotpiao.com/rock/poimen-frontend.git
targetRevision: main
path: k8s/argocd
destination:
server: https://kubernetes.default.svc
namespace: poimen
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
retry:
limit: 5
backoff:
duration: 5s
factor: 2
maxDuration: 3m
+3 -14
View File
@@ -12,18 +12,9 @@ spec:
description: Homelab GitOps — single-repo, in-cluster destinations only
sourceRepos:
- https://github.com/Riotpiaole/riotpiao.homelab.com.git
# Poimen services (GitHub)
- https://github.com/Riotpiaole/Poimen-memory.git
- https://github.com/Riotpiaole/Poimen-workflows.git
- https://github.com/Riotpiaole/poimen*.git
# In-cluster Forgejo repos — explicit allowlist (no wildcard)
- https://forgejo.riotpiao.com/rock/homelab.git
- https://forgejo.riotpiao.com/rock/homelab-frontend.git
- https://forgejo.riotpiao.com/rock/kmsvc-manage.git
- https://forgejo.riotpiao.com/rock/poimen.git
- https://forgejo.riotpiao.com/rock/poimen-memory.git
- https://forgejo.riotpiao.com/rock/poimen-workflows.git
- https://forgejo.riotpiao.com/rock/riotpiao.com.git
# In-cluster Forgejo wildcard — all rock/* repos can be onboarded without
# touching this AppProject. Enabled by Stage 1 (A1).
- https://forgejo.riotpiao.com/rock/*
# Public Helm chart repos referenced by k8s/argocd/apps/* and bootstrap/*
- https://cloudnative-pg.github.io/charts
- https://dl.gitea.com/charts/
@@ -42,8 +33,6 @@ spec:
- https://charts.jetstack.io
- https://kubernetes.github.io/ingress-nginx
- https://stakater.github.io/stakater-charts
# ArgoCD ecosystem charts
- https://argoproj.github.io/argo-helm
destinations:
- server: https://kubernetes.default.svc
namespace: "*"
+1 -2
View File
@@ -12,7 +12,7 @@ metadata:
spec:
project: homelab
source:
repoURL: https://forgejo.riotpiao.com/rock/homelab.git
repoURL: https://github.com/Riotpiaole/riotpiao.homelab.com.git
targetRevision: main
path: k8s/argocd/apps
directory:
@@ -26,4 +26,3 @@ spec:
selfHeal: true
syncOptions:
- CreateNamespace=true
- ServerSideApply=true
@@ -1,23 +0,0 @@
apiVersion: ENC[AES256_GCM,data:ECE=,iv:bISz4HovH++X7DW1Qj8Cw0L6+EvPB0+68hh7tfyW5C0=,tag:w+SjD7MsfeIuSf62n+Zl7Q==,type:str]
kind: ENC[AES256_GCM,data:AxXaV3Mb,iv:xfJb354Rrjz3zLctW2i6hl40yh9EsfIqKxDvmn8jqnU=,tag:1xQ/RR3KhPLKCTWCYjebJg==,type:str]
metadata:
name: ENC[AES256_GCM,data:gdQ4OoMYKUPEbsUeA5OI4i1xnh8=,iv:IA2kaCyqQmkYpLolvcTF4aleh+yd/ImXJMhRMvpGCgo=,tag:8qyYYw4EhBKKPzEmpepAeQ==,type:str]
namespace: ENC[AES256_GCM,data:/mOSyXWmRg==,iv:cpeHUTlMqlJzJttGtuR3DoiMtvVqFmDS0/5Tl7K7c2M=,tag:q4GxWgWn2wQJxJHqnq4WQA==,type:str]
type: ENC[AES256_GCM,data:92OcrsuV,iv:Z1XBy6iZ6unGrK4/SSdDa58pbPL3022gU/f1AOM4uvc=,tag:cd4wWRq/ohm6BJ/Eo6HOIw==,type:str]
stringData:
REGISTRY_PAT: ENC[AES256_GCM,data:zDKruOXhzsIcFZTTp4r4r9rwHih9uf1cCp06KZ6eIXSIZDvR6aQY4A==,iv:5Vch5Z7xn1KkRxrgOs2p/M58nd6WhtcuUhoZDCi8LSY=,tag:14SAZIWGOxJMqLBtOorjYQ==,type:str]
sops:
age:
- enc: |
-----BEGIN AGE ENCRYPTED FILE-----
YWdlLWVuY3J5cHRpb24ub3JnL3YxCi0+IFgyNTUxOSBIaG5QVUYyNjhvaENPcUFy
K3RSSVM2N2hTeFhNK0M5YWZmODhzZU1HaHlFCmNGZWdFcFEvanVwcXpGSmVvRHVx
aG45c2lBY1RuSTYwbDZTVk1QeHNWOTQKLS0tIDg4TnVMNjNtaU1VQk5zQjUvU2hM
R245ZVdqc0ZWWVhJb3dOZ3lpU3JjUlUKawSg09ZPq8FKx5tvOVZZ+K4yh7eTQsUp
be8mWUpS0+eEmNqh35BwU3HrETMQFA6a1kjVp30JOMtqa5rbYlzF7w==
-----END AGE ENCRYPTED FILE-----
recipient: age1e5fq3hwxy78psus2nfvmtmua36g0u3suk78ephw6246l974d2utsvn0hla
lastmodified: "2026-08-23T23:14:53Z"
mac: ENC[AES256_GCM,data:yzi6woUaUCi6w9pG/eKnU7k/VfZgoXg/tW8p9joG8p7iabLWMdlXlx23m/CItw/NE0zeGWA5iZPPFZXOit2vN36VzK3kQNoFW6QbhvYLZi+78C/RWIdcpZkxB/tPIRG0vq8Q1SA+rwGjG/0xeAFh+R7k+YBTMd2XWBH3P7EI5T4=,iv:kf6MDaAVrtDvPIEjHMMLxXDSDRC3I1GpsPeJFUYppiw=,tag:NukVbGLa9EoMYmRsa4nBtA==,type:str]
unencrypted_suffix: _unencrypted
version: 3.13.2
+4 -6
View File
@@ -1,10 +1,8 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
# Disable hash suffix for all generated secrets (stable names)
generatorOptions:
disableNameSuffixHash: true
# SOPS-encrypted secrets via ksops generator
# All homelab SOPS-encrypted Secrets, decrypted in-line via the ksops generator.
# Each *.enc.yaml carries its own metadata.namespace, so no namespace transformer
# here (that would rewrite every Secret into one namespace). Renders exactly the
# Secret objects — replaces the old argocd-cmp-cm SOPS plugin.
generators:
- secret-generator.yaml
@@ -1,25 +0,0 @@
apiVersion: ENC[AES256_GCM,data:wGA=,iv:Z2Gfzq3aJ9j4fYaeLQolgLb/XELTHrKX9at3vUsMLIw=,tag:yZLZFnltJln5jFoVsyyZ2A==,type:str]
kind: ENC[AES256_GCM,data:bJkV4pNj,iv:0ZT7l0kw9qSoiEMZioRw1aBzzlxBkoXR+hOxoPI7zPU=,tag:EQczPq508Qw1vi/oLCeQpw==,type:str]
metadata:
name: ENC[AES256_GCM,data:sZwTVM41PPaidMwFNRo5RvU=,iv:BHfuwIHng7rkeLK3a69t8cI9QeSfD/3FEXqxby+gBxM=,tag:3zQx8APnE2ZpBKf/ZYCxOA==,type:str]
namespace: ENC[AES256_GCM,data:i7lpoEaZ1oXS,iv:jUYyDPYhf1TV51he/S5MlKPD19Vmz6wfE98y8qFEg3U=,tag:zgHDVlplmW/XPUAKdncfYQ==,type:str]
type: ENC[AES256_GCM,data:2khs1uIg,iv:ET6HcBGyv33fGFlFAl3dkJQB83naHGeaYuWvW1IFhvw=,tag:qIfMeaSwv//7cNWvy8O5dg==,type:str]
stringData:
PAPERLESS_SECRET_KEY: ENC[AES256_GCM,data:kj9DrQYL3cQGJz87FHYlFKy6Muu84Oy/6wnsWAh0w/3MmcTAsCzQvUtKLKKf1U61JTQ=,iv:UffVv78vMHDEWHRFdSKZ/6qyrVD02Nlk0CNxwWd1jTo=,tag:LNMKnFWw7QCjfcKzI1s8ig==,type:str]
PAPERLESS_ADMIN_USER: ENC[AES256_GCM,data:a3bx7z0=,iv:nrU6VXZE74PNXM+Lgg4K/3mhusaQhA2GAgl0GzqAQt8=,tag:x8nswTfkFHR6Ou+JSI9P3Q==,type:str]
PAPERLESS_ADMIN_PASSWORD: ENC[AES256_GCM,data:+SWyAq7D1mNt1TlFOlo2Z0biK8kEn097,iv:an2ZZuXNSig+7mJP7ogHFc5+QvUb+z9pzqcV4xvFLbA=,tag:OaRiVZsnajEjAlc9mDWo9w==,type:str]
sops:
age:
- enc: |
-----BEGIN AGE ENCRYPTED FILE-----
YWdlLWVuY3J5cHRpb24ub3JnL3YxCi0+IFgyNTUxOSA1Q2V6bVBUYmVRR1N5SCtF
eXdTV2dmWjhtMS9lRFEzS0wzbkd6Y3JLMVRVCncydFBjdmJkRCtVUXphR2w0SDlJ
cng2bi9MWlJzTEN2amJrYjRJN2VFcEEKLS0tIEY5cmw2RGhnbzUxZW9FaFJjQmVN
WWcvNlNiYWdwbnNSR1Q4alpDZmFqTFkKh9TOw8ERP9fpx2pKi/Q0b7+OkEv0UC7o
aAIK4Tzvi5dp6y9IWcu9l6PjDLeYWOJ5wr7QABaFNOz82hngxUleDA==
-----END AGE ENCRYPTED FILE-----
recipient: age1e5fq3hwxy78psus2nfvmtmua36g0u3suk78ephw6246l974d2utsvn0hla
lastmodified: "2026-08-25T16:04:48Z"
mac: ENC[AES256_GCM,data:OXPQXUyn/SDftKH5nRzhqEtnaOb4Gp7etmGojqV4Z01kHnABZ42PPIyiBb8z997eLNZsK6/1bi7nUNyVp4fIuXG4F52aEdMewofOocCHokRnNmz7jzhooK1gScJb2u0eHG3FL5iLONMaGgVpk7BLYO3e0xiDytWGe8BxcuDukPg=,iv:kh0jUwFvO6AUICy2Us1E7YOTEcp3L+ptrGdDWBSpWyc=,tag:qUPFmZ/rpljln37f/NRjjw==,type:str]
unencrypted_suffix: _unencrypted
version: 3.13.2
@@ -1,2 +0,0 @@
FORGEJO_TOKEN=273fdcffabbcbb5a191e8289c73d106063acefc6
LLM_API_TOKEN=s3VksXyw2z3sGbegnwjMDFnJ6CtNRd1a5CcnE5A4ET77toCcykNdunk6Oa2J
@@ -1,23 +0,0 @@
apiVersion: ENC[AES256_GCM,data:bnY=,iv:Fuc3aqncHQ+L16o7eLarPbOECD3o8Mk5c2r9pQBpy70=,tag:JPfEnbNb3wZXPdXafnJDqw==,type:str]
kind: ENC[AES256_GCM,data:WFlmi4Yg,iv:Zq/KQbgNcBVoo8ZsQ2H79ygyc8Dtkgxh4fCpEExfwSg=,tag:cWHP6Y5V+nZP2tFMJrOB8A==,type:str]
metadata:
name: ENC[AES256_GCM,data:iCXbhwvg3Zq6YL/4j0wAy7Y=,iv:8s+d/8lDVEL7bGdIF+GOtAxapKnmx8JTjLSO04hXF5I=,tag:LjzX40DjK+uRZPCXsMlmwQ==,type:str]
namespace: ENC[AES256_GCM,data:HRMdZdCbxORQ,iv:MvaIWoKWjJRA7/fce0KtXRkFH/7cn0OuIg2QwHEdQzM=,tag:NqztiGyfU3BaopWBKhx2eg==,type:str]
type: ENC[AES256_GCM,data:myBW86Za,iv:3x9ys5UzVhAuX8gvZO67B1e+Orw4Aqasv/lHBgUV0b4=,tag:yl18P6SaxVLJidwmRtd4aQ==,type:str]
stringData:
FORGEJO_TOKEN: ENC[AES256_GCM,data:SUoBpNKOItyNGY01EhKNlPH0fyN4N7g6bfU2jsGogpCMhP7NuRipoA==,iv:h77RtmYZjXxiHYw1pHynQHuVX1+yJDGHwsXLf+DbUYA=,tag:GHzoDNHP1GGqlN5eT/N7TQ==,type:str]
sops:
age:
- enc: |
-----BEGIN AGE ENCRYPTED FILE-----
YWdlLWVuY3J5cHRpb24ub3JnL3YxCi0+IFgyNTUxOSBzTVBsekl3TGgzQVRMUU9m
cTduS2NoZW5uZFNNMG13cFY2cGVsTnlXaXhrCjJjbzhLdHZ4ZWpUV3J0cDQ0eVlM
WDNxdzVoQ2ZzcGJSbTU3RVorcnczNVkKLS0tIFdtQTE4Umk2TDBzUmdKOXNkbjFi
Vk5vK2VuUHVsb3FQL21vcGU1UW5CT1kKFM8vVjji3Cg9dvfTr4Hx7BJC8JH5ovef
Dj6zkofhsNWPgP9T+mnQakj+C0RKmHOMJqfWP7vwCBkZoNosIJVlMw==
-----END AGE ENCRYPTED FILE-----
recipient: age1e5fq3hwxy78psus2nfvmtmua36g0u3suk78ephw6246l974d2utsvn0hla
lastmodified: "2026-09-01T05:32:42Z"
mac: ENC[AES256_GCM,data:it24T9y9ixXo2aiL37k93vKFR+SRjjuI9DQdv0sWYtTogWnc7+uXBY4Zip/ouWyCse1muKKAGuek5c0XVrvSw4an9VkaXFczeunaZb6MOyVbVOkmJr+5xZFpZGjYcSkrhaWcVheedZ3iIFU5UWI7BBn/qQCf+HJ483cJqwtrV34=,iv:WprlWJdsMBNjqaA0O3ekfXMUpX5gC6OLYortQXYdTS4=,tag:rGqSSRTYTv2VR6AcRokO0A==,type:str]
unencrypted_suffix: _unencrypted
version: 3.13.2
-3
View File
@@ -11,7 +11,6 @@ files:
- agent-pod-ssh-key.enc.yaml
- authentik-secrets.enc.yaml
- cloudflare-secrets.enc.yaml
- forgejo-registry-pat.enc.yaml
- forgejo-registry-pull.enc.yaml
- forgejo-runner-token.enc.yaml
- forgejo-secrets.enc.yaml
@@ -23,7 +22,5 @@ files:
- homelab-ca-secrets.enc.yaml
- loki-secrets.enc.yaml
- minio-secrets.enc.yaml
- paperless-secrets.enc.yaml
- vault-secrets.enc.yaml
- vault-unseal-keys.enc.yaml
- portfolio-secrets.enc.yaml
@@ -5,9 +5,9 @@ metadata:
namespace: ENC[AES256_GCM,data:wK6m,iv:KtA31Bo8aGE1HU8H9KWbMwt2NfywWfy77G/LaReaI1E=,tag:V9SSlPawRQw0x2utFNN1aw==,type:str]
type: ENC[AES256_GCM,data:6UPZ1ZTR,iv:FxN1ebrlJ4IO3eDGEYSvktSpMgAebcl0DO1WHh5O0+0=,tag:2bZN4zdbHzx0oR+bTJPPJg==,type:str]
stringData:
key1: ENC[AES256_GCM,data:ke5EfsOsZHBJyAQoFfwhhWCQJQgnwcrBqL4CzPIfkx7bi9S7kWgWySdxcXQ=,iv:eYKxG8k+hizp2t2i/YMR2lQNJQFV+A21YyWnDc+kJ9w=,tag:Wa1Lx3EKvfpYfnLrBOWSIg==,type:str]
key2: ENC[AES256_GCM,data:JxiXgLKvDe5oiHlwIL/Cj8txHc7fVQ5VzBcQMU/ro9TSOcTGSe7Z3omQzwY=,iv:x5arcY2dey+npMpUxjdUPV+t94LEaMOXq3iOar88eT8=,tag:E/Z5zEDy87B1IAWkdjdNEg==,type:str]
key3: ENC[AES256_GCM,data:AsxyyBTHzt+SLru6ayi1UbgPPFmdIvowIDdcfc7N03Xt8V/KaJW3kbXExG8=,iv:lA+bERIBq4+bxx08ZZjP5tMHC9JQlGPW1EMCwEixSeg=,tag:xkuTAU3KTidFuhNMXkVFhA==,type:str]
key1: ENC[AES256_GCM,data:dV2HOh7W1Pl0QDJaGvtEKpBppybQMiK5kxzyk05zkA352ZmJvE8Ppm5yWKk=,iv:grQ8v2o/LHpJnIZonjtgTHKcLUQIH8xl1FtWtEs4rEs=,tag:lANcccYT6Ouo/0IcoLw4uA==,type:str]
key2: ENC[AES256_GCM,data:YSdpYL8h56PfUrvMFhBXhmBG2en0toLKEpYlJqwjAk/vI0jm7gq/TqGn034=,iv:GxthkNLhm3qHxjkZetiYl58Qa/K7G2ib6E+LWK16H8Q=,tag:q79UUoh6uCcMZJ0U18Ireg==,type:str]
key3: ENC[AES256_GCM,data:sO7Lao9qkmIOdARL+6FBcLo+e4z8LzMEzNrwcy5hv9uRuxppwAXrP0MO1vw=,iv:wSWe9h1HovRhV5YCZrmvcB2VEfRYQ+gasGO3uh8IpmQ=,tag:Bc5yaROQDTz1AhVnFDm/BA==,type:str]
sops:
age:
- enc: |
@@ -19,7 +19,7 @@ sops:
7RltNZF2SCxjIv5C2pqf3CgmRBaQGWgMybRRH5gdB87PLBKkPL3+HQ==
-----END AGE ENCRYPTED FILE-----
recipient: age1e5fq3hwxy78psus2nfvmtmua36g0u3suk78ephw6246l974d2utsvn0hla
lastmodified: "2026-08-26T23:33:09Z"
mac: ENC[AES256_GCM,data:kr8XuOqXyhfrC84yBot5aoVQZZhQd9xYMUEs0XDwkmCgtG0iobADYH5XWPh72xSdWEtwxkZJL7r0UEyEhU9zjlUA5qHupOSBBt1SgZ+zstMvaOqUnlNn//p/DIJBpsiT/qmx64NpTLAiz6lm0796MozIMr8PTX+ubGLHi9tnUiY=,iv:jwP0MLCT7nG6m+nZeqNip9q3BcScpcDXmInnY95DicY=,tag:Cgg2IUZARkaK1gF2qUIthg==,type:str]
lastmodified: "2026-08-12T20:52:24Z"
mac: ENC[AES256_GCM,data:kLjoBWmJ2bZ13EbuzgppahHKmCY/xtCi1R7xCEUbCP4FQfiVaC4qAbVftE4lTNlXa2mqCKsGPYZ/A3HV/4+a+it/pg0a+rj7E7DszhieWbZufHMJs+w4/Le8l8wFFE5aNW0wdvGTJZ0n4HBOdAkE1qN9ReaiaQgr0rk5ggOPG9g=,iv:MiDVUgcoKrP/gj4keqomo6OB12dmz9VoesTiHQTnBFM=,tag:cLK67Pym4BuvXLF1RUA8Gw==,type:str]
unencrypted_suffix: _unencrypted
version: 3.13.2
-1
View File
@@ -323,4 +323,3 @@ spec:
name: homarr
port:
number: 7575
-6
View File
@@ -1,6 +0,0 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
namespace: kyverno
resources:
- policies.yaml
-42
View File
@@ -1,42 +0,0 @@
# Kyverno: Policy engine for Kubernetes image scanning, Pod security, and admission control
# Scan all images, enforce baseline Pod Security Standard, prevent privilege escalation
replicaCount: 1
image:
registry: ghcr.io
repository: kyverno/kyverno
tag: "v1.14.0"
config:
# Webhook timeout for policy evaluation. Increase if scanning takes longer.
webhookTimeoutSeconds: 30
# Failure policy: fail-open (audit/log) vs fail-closed (reject on error)
failurePolicy: fail
# Resource limits for webhook
webhookAnnotations:
rules: "allow"
# Pod security via Kyverno instead of Pod Security Policies (deprecated)
# Enforces baseline restrictions cluster-wide, with exceptions for privileged namespaces
podSecurityContext:
runAsNonRoot: true
runAsUser: 1000
rbac:
create: true
resources:
requests:
memory: "256Mi"
cpu: "100m"
limits:
memory: "512Mi"
cpu: "500m"
# Webhook configuration
webhook:
timeoutSeconds: 30
# Failure policy: "Fail" (reject on error) or "Ignore" (audit-only)
# Set to "Ignore" for initial testing, then change to "Fail"
failurePolicy: ignore
-204
View File
@@ -1,204 +0,0 @@
# Kyverno ClusterPolicies: Image scanning, Pod security, and admission control
---
# Policy 1: Require non-root containers
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: require-non-root
namespace: kyverno
spec:
validationFailureAction: audit # audit first, then change to enforce
rules:
- name: check-runAsNonRoot
match:
any:
- resources:
kinds:
- Pod
selector:
matchLabels:
pod-security.kubernetes.io/enforce: "!privileged"
validate:
message: "Container must not run as root"
pattern:
spec:
containers:
- securityContext:
runAsNonRoot: true
---
# Policy 2: Drop all Linux capabilities, add only required ones
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: require-dropped-caps
namespace: kyverno
spec:
validationFailureAction: audit
rules:
- name: drop-all-capabilities
match:
any:
- resources:
kinds:
- Pod
selector:
matchLabels:
pod-security.kubernetes.io/enforce: "!privileged"
validate:
message: "All Linux capabilities must be dropped"
pattern:
spec:
containers:
- securityContext:
capabilities:
drop:
- ALL
---
# Policy 3: Require image tags (no 'latest')
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: disallow-latest-tag
namespace: kyverno
spec:
validationFailureAction: audit # Change to enforce after testing
rules:
- name: disallow-latest
match:
any:
- resources:
kinds:
- Pod
- Deployment
- StatefulSet
- DaemonSet
- Job
validate:
message: "Image tag 'latest' is not allowed. Use explicit version tags."
pattern:
spec:
=(template):
spec:
containers:
- image: "!*:latest"
=(initContainers):
- image: "!*:latest"
---
# Policy 4: Restrict images to trusted registries
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: restrict-registries
namespace: kyverno
spec:
validationFailureAction: audit
rules:
- name: trusted-registries
match:
any:
- resources:
kinds:
- Pod
- Deployment
- StatefulSet
- DaemonSet
- Job
selector:
matchLabels:
pod-security.kubernetes.io/enforce: "!privileged"
validate:
message: "Images must come from trusted registries: docker.io, ghcr.io, quay.io, k8s.gcr.io, registry.k8s.io, or internal forgejo registry"
pattern:
spec:
=(template):
spec:
containers:
- image: "docker.io/* | ghcr.io/* | quay.io/* | k8s.gcr.io/* | registry.k8s.io/* | forgejo.riotpiao.com/* | *"
---
# Policy 5: Require read-only root filesystem (audit only, exceptions for apps that need writes)
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: require-readonly-filesystem
namespace: kyverno
spec:
validationFailureAction: audit
rules:
- name: check-readOnlyRootFilesystem
match:
any:
- resources:
kinds:
- Pod
selector:
matchLabels:
pod-security.kubernetes.io/enforce: "!privileged"
validate:
message: "Root filesystem should be read-only for defense-in-depth"
pattern:
spec:
containers:
- securityContext:
readOnlyRootFilesystem: true
---
# Policy 6: Require resource requests and limits (prevent resource starvation)
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: require-resource-limits
namespace: kyverno
spec:
validationFailureAction: audit
rules:
- name: check-resources
match:
any:
- resources:
kinds:
- Pod
- Deployment
- StatefulSet
- DaemonSet
excludeResources:
namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: "kyverno|kube-system|kube-node-lease"
validate:
message: "CPU and memory requests and limits are required"
pattern:
spec:
=(template):
spec:
containers:
- resources:
requests:
memory: "?*"
cpu: "?*"
limits:
memory: "?*"
cpu: "?*"
---
# Policy 7: Require securityContext on all containers
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: require-security-context
namespace: kyverno
spec:
validationFailureAction: audit
rules:
- name: check-securityContext
match:
any:
- resources:
kinds:
- Pod
selector:
matchLabels:
pod-security.kubernetes.io/enforce: "!privileged"
validate:
message: "securityContext must be defined"
pattern:
spec:
containers:
- securityContext: {}
@@ -1,220 +0,0 @@
# Forgejo OCI Registry Cleanup CronJob
# Deletes old image tags, keeping only the latest N versions per repository.
# Useful for retiring old builds when new versions are pushed.
---
apiVersion: v1
kind: ConfigMap
metadata:
name: forgejo-registry-cleanup-script
namespace: cicd
data:
cleanup.sh: |
#!/bin/bash
set -eo pipefail
# Configuration
REGISTRY_HOST="${REGISTRY_HOST:-forgejo.riotpiao.com}"
REGISTRY_URL="https://${REGISTRY_HOST}"
KEEP_VERSIONS="${KEEP_VERSIONS:-3}" # Keep latest N versions per image
DRY_RUN="${DRY_RUN:-false}"
# Load credentials from mounted secret
REGISTRY_USER="${REGISTRY_USER:-_json_key}"
REGISTRY_PASS="$(cat /etc/registry-secret/password 2>/dev/null || echo '')"
log() {
echo "[$(date +'%Y-%m-%d %H:%M:%S')] $*"
}
error() {
echo "[$(date +'%Y-%m-%d %H:%M:%S')] ERROR: $*" >&2
return 1
}
# Verify crane is available
if ! command -v crane &> /dev/null; then
error "crane not found. Install google/crane image for registry operations."
exit 1
fi
log "Starting Forgejo registry cleanup"
log "Registry: $REGISTRY_URL"
log "Keep versions: $KEEP_VERSIONS per image"
log "Dry run: $DRY_RUN"
# Authenticate crane with registry
if [ -n "$REGISTRY_PASS" ]; then
echo "$REGISTRY_PASS" | crane auth login "$REGISTRY_HOST" -u "$REGISTRY_USER" --password-stdin
log "Authenticated to $REGISTRY_HOST"
fi
# List all repositories (catalog)
# Note: This endpoint requires the registry to expose /v2/_catalog (standard OCI)
# If not available, images must be discovered another way
CATALOG=$(curl -s -u "${REGISTRY_USER}:${REGISTRY_PASS}" \
"${REGISTRY_URL}/v2/_catalog" | grep -o '"repositories":\[\K[^]]*' || echo '')
if [ -z "$CATALOG" ]; then
log "WARNING: Could not retrieve catalog from ${REGISTRY_URL}/v2/_catalog"
log "Registry may not expose _catalog endpoint or credentials invalid"
exit 0
fi
# Parse repositories from catalog JSON
REPOS=$(echo "$CATALOG" | grep -o '"[^"]*"' | tr -d '"')
TOTAL_DELETED=0
for REPO in $REPOS; do
log "Processing repository: $REPO"
IMAGE="${REGISTRY_HOST}/${REPO}"
# Get all tags for this image
TAGS=$(crane ls "$IMAGE" 2>/dev/null || echo "")
if [ -z "$TAGS" ]; then
log " No tags found for $REPO (or access denied)"
continue
fi
# Filter out 'latest' tag and sort by creation time (newer first)
# Note: crane doesn't provide direct date sorting; we use the order returned
# Assumption: tags are returned newest first (not always true)
TAG_COUNT=$(echo "$TAGS" | wc -l)
if [ "$TAG_COUNT" -le "$KEEP_VERSIONS" ]; then
log " $REPO: $TAG_COUNT tags total, keeping all (≤ $KEEP_VERSIONS)"
continue
fi
# Get tags to delete (all except the first N)
TAGS_TO_DELETE=$(echo "$TAGS" | tail -n +$((KEEP_VERSIONS + 1)))
for TAG in $TAGS_TO_DELETE; do
FULL_IMAGE="${IMAGE}:${TAG}"
DELETED_SIZE="0"
if [ "$DRY_RUN" = "true" ]; then
log " [DRY RUN] Would delete: $FULL_IMAGE"
else
if crane delete "$FULL_IMAGE" 2>&1; then
log " Deleted: $FULL_IMAGE"
((TOTAL_DELETED++))
else
error "Failed to delete $FULL_IMAGE (may already be deleted)"
fi
fi
done
done
log "Cleanup complete. Total images deleted: $TOTAL_DELETED"
---
apiVersion: batch/v1
kind: CronJob
metadata:
name: forgejo-registry-cleanup
namespace: cicd
labels:
app: forgejo-registry-cleanup
spec:
# Run at 2 AM UTC every day (adjust as needed)
schedule: "0 2 * * *"
# Keep last 3 successful/failed runs for debugging
successfulJobsHistoryLimit: 3
failedJobsHistoryLimit: 3
# Suspend if needed (set to false to enable)
suspend: false
jobTemplate:
spec:
# Cleanup jobs after 6 hours whether they succeeded or failed
ttlSecondsAfterFinished: 21600
template:
metadata:
labels:
app: forgejo-registry-cleanup
spec:
serviceAccountName: forgejo-registry-cleanup
restartPolicy: OnFailure
containers:
- name: cleanup
# Use google/crane for registry operations
image: gcr.io/go-containerregistry/crane:latest
imagePullPolicy: IfNotPresent
env:
- name: REGISTRY_HOST
value: "forgejo.riotpiao.com"
- name: KEEP_VERSIONS
value: "3" # Keep 3 latest versions
- name: DRY_RUN
value: "false" # Set to "true" for dry-run mode
- name: REGISTRY_USER
valueFrom:
secretKeyRef:
name: forgejo-registry-token
key: username
optional: true
volumeMounts:
- name: script
mountPath: /scripts
- name: registry-secret
mountPath: /etc/registry-secret
readOnly: true
# Run cleanup script via entrypoint override
command:
- /bin/sh
- -c
- |
# Install bash and curl if needed
apk add --no-cache bash curl
chmod +x /scripts/cleanup.sh
/scripts/cleanup.sh
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
cpu: 500m
memory: 512Mi
# Safety: kill after 30 min (prevents hanging on large registries)
securityContext:
runAsNonRoot: true
runAsUser: 65534
allowPrivilegeEscalation: false
readOnlyRootFilesystem: false
capabilities:
drop:
- ALL
volumes:
- name: script
configMap:
name: forgejo-registry-cleanup-script
defaultMode: 0755
- name: registry-secret
secret:
secretName: forgejo-registry-token
optional: true
---
apiVersion: v1
kind: ServiceAccount
metadata:
name: forgejo-registry-cleanup
namespace: cicd
---
# No RBAC needed: this pod only talks to the registry API (external service)
# If expanded to manage in-cluster resources, add Role/RoleBinding here
@@ -1,72 +0,0 @@
# ArgoCD Image Updater configuration
# Watches Forgejo registry and updates ArgoCD Applications with new image tags
config:
# Registry configuration - Forgejo allows anonymous pulls
registries:
- name: forgejo
api_url: https://forgejo.riotpiao.com
prefix: forgejo.riotpiao.com
default: true
insecure: false
# Log level
logLevel: debug
# ArgoCD API server
argocd:
grpcWeb: true
serverAddress: argocd-server.argocd.svc.cluster.local
insecure: true
plaintext: true
# Git write-back configuration (for multi-source Applications)
git:
# Commit author for image updates
user:
name: "ArgoCD Image Updater"
email: "[email protected]"
# Use SSH keys from ArgoCD's known hosts + credentials
# Image Updater inherits ArgoCD's git credentials (mounted via ArgoCD secret)
# Mount ArgoCD's git credentials for write-back
extraVolumes:
- name: argocd-ssh-known-hosts-cm
configMap:
name: argocd-ssh-known-hosts-cm
defaultMode: 0644
- name: argocd-gpg-keys-cm
configMap:
name: argocd-gpg-keys-cm
optional: true
defaultMode: 0644
- name: argocd-gpg-pubring
configMap:
name: argocd-gpg-pubring-cm
optional: true
defaultMode: 0644
extraVolumeMounts:
- name: argocd-ssh-known-hosts-cm
mountPath: /etc/ssh/ssh_known_hosts.d/argocd-ssh-known-hosts
subPath: ssh_known_hosts
- name: argocd-gpg-keys-cm
mountPath: /etc/gpg/source
- name: argocd-gpg-pubring
mountPath: /etc/gpg/pubring
# Extra environment variables
extraEnv:
- name: ARGOCD_GRPC_WEB
value: "true"
- name: GIT_SSH_KNOWN_HOSTS_CONFIG_MAP_ENABLED
value: "true"
# Resources
resources:
requests:
cpu: 50m
memory: 64Mi
limits:
cpu: 200m
memory: 128Mi
@@ -1,137 +0,0 @@
# Cluster-wide cleanup of stale failed/completed Jobs and Pods.
# Runs daily at 04:00 UTC. Deletes:
# - Failed Jobs older than 24h (any namespace)
# - Completed Jobs older than 72h with no owning CronJob
# - Orphan pods in Error/Failed/Evicted state older than 1h
#
# CronJob-owned Jobs are managed by failedJobsHistoryLimit/successfulJobsHistoryLimit,
# but standalone Jobs (helm hooks, one-off runs, longhorn maintenance) have no TTL
# and linger forever.
apiVersion: v1
kind: ServiceAccount
metadata:
name: stale-job-cleanup
namespace: kube-system
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: stale-job-cleanup
rules:
- apiGroups: ["batch"]
resources: ["jobs"]
verbs: ["get", "list", "delete"]
- apiGroups: [""]
resources: ["pods"]
verbs: ["get", "list", "delete"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: stale-job-cleanup
subjects:
- kind: ServiceAccount
name: stale-job-cleanup
namespace: kube-system
roleRef:
kind: ClusterRole
name: stale-job-cleanup
apiGroup: rbac.authorization.k8s.io
---
apiVersion: batch/v1
kind: CronJob
metadata:
name: stale-job-cleanup
namespace: kube-system
labels:
app: stale-job-cleanup
spec:
schedule: "0 4 * * *"
concurrencyPolicy: Forbid
successfulJobsHistoryLimit: 3
failedJobsHistoryLimit: 3
jobTemplate:
spec:
ttlSecondsAfterFinished: 86400 # self-cleanup after 24h
backoffLimit: 1
activeDeadlineSeconds: 300
template:
spec:
serviceAccountName: stale-job-cleanup
restartPolicy: Never
tolerations:
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
containers:
- name: cleanup
image: alpine/k8s:1.31.0
command:
- sh
- -c
- |
set -e
NOW=$(date +%s)
echo "=== Cleaning failed Jobs older than 24h ==="
kubectl get jobs --all-namespaces -o json | \
jq -r '.items[] |
select(.status.conditions[]?.type == "Failed") |
select(.status.completionTime or .status.startTime) |
"\(.metadata.namespace) \(.metadata.name) \(.status.startTime // .status.completionTime // .metadata.creationTimestamp)"' | \
while read -r NS NAME TS; do
JOB_EPOCH=$(date -d "$TS" +%s 2>/dev/null || echo 0)
AGE_H=$(( (NOW - JOB_EPOCH) / 3600 ))
if [ "$AGE_H" -ge 24 ]; then
echo "[delete] $NS/$NAME (failed ${AGE_H}h ago)"
kubectl delete job "$NAME" -n "$NS" --cascade=foreground 2>/dev/null || true
fi
done
echo ""
echo "=== Cleaning completed standalone Jobs older than 72h ==="
kubectl get jobs --all-namespaces -o json | \
jq -r '.items[] |
select(.status.succeeded >= 1) |
select((.metadata.ownerReferences // []) | length == 0) |
"\(.metadata.namespace) \(.metadata.name) \(.status.completionTime // .metadata.creationTimestamp)"' | \
while read -r NS NAME TS; do
JOB_EPOCH=$(date -d "$TS" +%s 2>/dev/null || echo 0)
AGE_H=$(( (NOW - JOB_EPOCH) / 3600 ))
if [ "$AGE_H" -ge 72 ]; then
echo "[delete] $NS/$NAME (completed ${AGE_H}h ago, no owner)"
kubectl delete job "$NAME" -n "$NS" --cascade=foreground 2>/dev/null || true
fi
done
echo ""
echo "=== Cleaning orphan Error/Failed/Evicted pods older than 1h ==="
# Evicted pods show as Failed with reason Evicted
kubectl get pods --all-namespaces -o json | \
jq -r '.items[] |
select(
.status.phase == "Failed" or
(.status.reason // "") == "Evicted" or
(.status.containerStatuses // [] | any(.state.terminated.reason == "Error"))
) |
select((.metadata.ownerReferences // []) | all(.kind != "Job")) |
"\(.metadata.namespace) \(.metadata.name) \(.metadata.creationTimestamp)"' | \
while read -r NS NAME TS; do
POD_EPOCH=$(date -d "$TS" +%s 2>/dev/null || echo 0)
AGE_H=$(( (NOW - POD_EPOCH) / 3600 ))
if [ "$AGE_H" -ge 1 ]; then
echo "[delete] $NS/$NAME (error/evicted ${AGE_H}h ago)"
kubectl delete pod "$NAME" -n "$NS" --force 2>/dev/null || true
fi
done
echo ""
echo "Cleanup complete"
resources:
requests:
cpu: 50m
memory: 64Mi
limits:
cpu: 250m
memory: 128Mi
+1 -6
View File
@@ -1,13 +1,8 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
# Dedicated per-app CNPG clusters. NO top-level `namespace:` — each Cluster
# carries its own ns (iam / temporal / poimen / paperless); a transformer would
# wrongly collapse them. poimen ns created by poimen-root app, memory-db
# deployed into it. paperless ns declared in namespaces.yaml above.
# carries its own ns (iam / temporal); a transformer would wrongly collapse them.
resources:
- namespaces.yaml
- authentik-db.yaml
- temporal-db.yaml
- memory-db.yaml
- paperless-db.yaml
- obsidian-vault-pvc.yaml
-36
View File
@@ -1,36 +0,0 @@
# Dedicated CNPG Postgres for Poimen Memory (GitOps, wave 2).
# Uses default longhorn storage class (3 replicas, dataLocality disabled).
# CNPG generates secret `memory-db-app` + service `memory-db-rw` in ns poimen.
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: memory-db
namespace: poimen
annotations:
argocd.argoproj.io/sync-options: SkipDryRunOnMissingResource=true
spec:
instances: 2
imageName: ghcr.io/cloudnative-pg/postgresql:16.2
bootstrap:
initdb:
database: memory
owner: app
encoding: UTF8
localeCollate: C
localeCType: C
postInitApplicationSQL:
- "CREATE EXTENSION vector;"
enableSuperuserAccess: false
resources:
requests: { memory: "512Mi", cpu: "250m" }
limits: { memory: "2Gi", cpu: "1" }
storage:
size: 20Gi
storageClass: longhorn
affinity:
podAntiAffinityType: preferred
topologyKey: kubernetes.io/hostname
tolerations:
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
+1 -15
View File
@@ -1,7 +1,6 @@
# DB clusters are wave 2 — their namespaces must exist first (their apps that
# would CreateNamespace run later, w3/w8). Declared here so the databases App
# creates them. authentik/vault/temporal/paperless CreateNamespace=true then
# no-ops. poimen namespace created by poimen-root app (wave 7).
# creates them. authentik/vault/temporal CreateNamespace=true then no-ops.
apiVersion: v1
kind: Namespace
metadata:
@@ -11,16 +10,3 @@ apiVersion: v1
kind: Namespace
metadata:
name: temporal
---
apiVersion: v1
kind: Namespace
metadata:
name: paperless
---
# Needed here (not just immich's own CreateNamespace=true at wave 8) because
# k8s/infra/iam's PostSync job (wave 3) has a RoleBinding targeting this
# namespace - same ordering reason as paperless above.
apiVersion: v1
kind: Namespace
metadata:
name: immich
@@ -1,18 +0,0 @@
---
# Obsidian vault PVC — shared storage for REST API + UI pods
# ReadWriteMany so both obsidian-server and obsidian-ui can mount simultaneously
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: obsidian-vault
namespace: poimen
labels:
app.kubernetes.io/name: obsidian-server
app.kubernetes.io/part-of: poimen-memory
spec:
accessModes:
- ReadWriteMany
storageClassName: longhorn
resources:
requests:
storage: 10Gi
-36
View File
@@ -1,36 +0,0 @@
# Dedicated CNPG Postgres for paperless-ngx (GitOps, wave 2 — before the
# paperless app at w8). Same recipe as memory-db: default longhorn storage
# class (3 replicas), 2 instances, 20Gi.
# CNPG generates secret `paperless-db-app` + service `paperless-db-rw` in ns
# paperless; the paperless Deployment reads them locally.
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: paperless-db
namespace: paperless
annotations:
argocd.argoproj.io/sync-options: SkipDryRunOnMissingResource=true
spec:
instances: 2
imageName: ghcr.io/cloudnative-pg/postgresql:16.2
bootstrap:
initdb:
database: paperless
owner: app
encoding: UTF8
localeCollate: C
localeCType: C
enableSuperuserAccess: false
resources:
requests: { memory: "512Mi", cpu: "250m" }
limits: { memory: "2Gi", cpu: "1" }
storage:
size: 20Gi
storageClass: longhorn
affinity:
podAntiAffinityType: preferred
topologyKey: kubernetes.io/hostname
tolerations:
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
@@ -36,4 +36,3 @@ data:
valid_volumes:
- /docker-certs/client
network: host
docker_host: automount
@@ -34,11 +34,7 @@ spec:
command: ["sh", "-c"]
args:
- |
# Always re-register to keep labels in sync with values.yaml.
# Without this, changing a runner label requires manually deleting
# the PVC or .runner file — not GitOps-friendly.
rm -f /data/.runner
forgejo-runner register --no-interactive \
test -f /data/.runner || forgejo-runner register --no-interactive \
--instance {{ .Values.runner.forgejoUrl }} \
--token $(RUNNER_TOKEN) \
--name {{ .Values.runner.name }} \
@@ -60,18 +56,20 @@ spec:
containers:
- name: runner
image: {{ .Values.runner.image.repository }}:{{ .Values.runner.image.tag }}
command: ["sh", "-c", "while ! wget -q -O- http://localhost:2375/_ping >/dev/null 2>&1; do echo 'waiting for dind...'; sleep 2; done; echo 'dind ready'; forgejo-runner daemon --config /etc/forgejo-runner/config.yaml"]
command: ["sh", "-c", "forgejo-runner daemon --config /etc/forgejo-runner/config.yaml"]
workingDir: /data
env:
- name: DOCKER_HOST
value: tcp://localhost:2375
value: tcp://localhost:2376
- name: DOCKER_TLS_VERIFY
value: "1"
- name: DOCKER_CERT_PATH
value: /docker-certs/client
volumeMounts:
- name: runner-data
mountPath: /data
- name: docker-certs
mountPath: /docker-certs
- name: docker-sock
mountPath: /run
- name: homelab-ca
mountPath: /etc/ssl/certs/homelab-ca.pem
subPath: ca.crt
@@ -87,12 +85,10 @@ spec:
privileged: true # required for DinD; cicd namespace is labelled privileged
env:
- name: DOCKER_TLS_CERTDIR
value: ""
value: /docker-certs
volumeMounts:
- name: docker-certs
mountPath: /docker-certs
- name: docker-sock
mountPath: /run
- name: dind-storage
mountPath: /var/lib/docker
- name: homelab-ca
@@ -117,8 +113,6 @@ spec:
claimName: {{ .Release.Name }}-dind
- name: docker-certs
emptyDir: {} # DinD regenerates mTLS certs on each start
- name: docker-sock
emptyDir: {} # Shared docker socket between dind and runner
- name: homelab-ca
# homelab-ca is a ConfigMap (public CA trust bundle), not a Secret.
# The volumeMounts use subPath: ca.crt to project the single cert file.
@@ -1,152 +0,0 @@
{{- if .Values.gc.enabled }}
# Garbage-collects DinD Docker images/volumes/build-cache and actcache across
# ALL forgejo-runner pods. Prevents PVC fill-up that breaks CI runs.
# Only rendered once (enable in default values.yaml, disable in per-runner overrides).
apiVersion: v1
kind: ServiceAccount
metadata:
name: runner-gc
namespace: {{ .Release.Namespace }}
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: runner-gc
namespace: {{ .Release.Namespace }}
rules:
- apiGroups: [""]
resources: ["pods"]
verbs: ["get", "list"]
- apiGroups: [""]
resources: ["pods/exec"]
verbs: ["create"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: runner-gc
namespace: {{ .Release.Namespace }}
subjects:
- kind: ServiceAccount
name: runner-gc
namespace: {{ .Release.Namespace }}
roleRef:
kind: Role
name: runner-gc
apiGroup: rbac.authorization.k8s.io
---
apiVersion: batch/v1
kind: CronJob
metadata:
name: forgejo-runner-gc
namespace: {{ .Release.Namespace }}
labels:
app: forgejo-runner-gc
spec:
schedule: {{ .Values.gc.schedule | quote }}
concurrencyPolicy: Forbid
successfulJobsHistoryLimit: 3
failedJobsHistoryLimit: 3
jobTemplate:
spec:
backoffLimit: 1
activeDeadlineSeconds: 900
template:
spec:
serviceAccountName: runner-gc
restartPolicy: Never
tolerations:
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
containers:
- name: gc
image: {{ .Values.gc.image }}
command:
- sh
- -c
- |
set -e
# Iterate all forgejo-runner pods (golang, rust, node)
PODS=$(kubectl -n {{ .Release.Namespace }} get pod \
-l app -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.metadata.labels.app}{"\n"}{end}' \
| grep 'forgejo-runner-' | awk '{print $1}')
if [ -z "$PODS" ]; then
echo "no forgejo-runner pods found, skipping"
exit 0
fi
for POD in $PODS; do
echo "===== $POD ====="
# 1. Docker image prune (DinD sidecar)
echo "[docker] before:"
kubectl -n {{ .Release.Namespace }} exec "$POD" -c dind -- docker system df 2>/dev/null || true
echo "[docker] pruning non-latest images older than {{ .Values.gc.pruneAge }}..."
# Keep :latest tagged images, delete all others older than pruneAge.
# docker image prune can't filter by tag, so we list and selectively rmi.
kubectl -n {{ .Release.Namespace }} exec "$POD" -c dind -- \
sh -c '
# Remove dangling (untagged) images older than {{ .Values.gc.pruneAge }}
docker image prune -f --filter "until={{ .Values.gc.pruneAge }}" 2>/dev/null
# Remove tagged non-latest images older than {{ .Values.gc.pruneAge }}
CUTOFF=$(date -d "-{{ .Values.gc.pruneAgeHours }} hours" +%s 2>/dev/null || date -v-{{ .Values.gc.pruneAgeHours }}H +%s)
docker images --format "{{"{{"}} .Repository {{"}}"}}:{{"{{"}} .Tag {{"}}"}} {{"{{"}} .CreatedAt {{"}}"}}" | while read -r IMAGE_TAG CREATED_REST; do
TAG=$(echo "$IMAGE_TAG" | rev | cut -d: -f1 | rev)
# Skip latest-tagged images
if [ "$TAG" = "latest" ]; then
echo "[keep] $IMAGE_TAG (latest)"
continue
fi
# Check image age via inspect
CREATED_TS=$(docker inspect --format="{{"{{"}} .Created {{"}}"}}" "$IMAGE_TAG" 2>/dev/null | head -1)
if [ -z "$CREATED_TS" ]; then continue; fi
IMAGE_EPOCH=$(date -d "$CREATED_TS" +%s 2>/dev/null || date -jf "%Y-%m-%dT%H:%M:%S" "$(echo $CREATED_TS | cut -dT -f1-2 | cut -d. -f1)" +%s 2>/dev/null || echo 0)
if [ "$IMAGE_EPOCH" -lt "$CUTOFF" ] 2>/dev/null; then
echo "[delete] $IMAGE_TAG (older than {{ .Values.gc.pruneAge }})"
docker rmi -f "$IMAGE_TAG" 2>/dev/null || true
else
echo "[keep] $IMAGE_TAG (recent)"
fi
done
' 2>/dev/null || true
echo "[docker] pruning build cache unused >{{ .Values.gc.pruneAge }}..."
kubectl -n {{ .Release.Namespace }} exec "$POD" -c dind -- \
docker builder prune -af --filter "until={{ .Values.gc.pruneAge }}" 2>/dev/null || true
echo "[docker] pruning dangling volumes..."
kubectl -n {{ .Release.Namespace }} exec "$POD" -c dind -- \
docker volume prune -af 2>/dev/null || true
echo "[docker] after:"
kubectl -n {{ .Release.Namespace }} exec "$POD" -c dind -- docker system df 2>/dev/null || true
# 2. Actcache cleanup (runner container)
echo "[actcache] cleaning incomplete and stale cache entries..."
kubectl -n {{ .Release.Namespace }} exec "$POD" -c runner -- \
sh -c '
# Delete incomplete/partial cache uploads immediately (tmp dirs)
find /data/.cache/actcache/cache -name "tmp" -type d -exec rm -rf {} + 2>/dev/null || true
# Delete cache entries not accessed in last {{ .Values.gc.actcacheMaxAgeDays }} day(s)
find /data/.cache/actcache/cache -type f -mtime +{{ .Values.gc.actcacheMaxAgeDays }} -delete 2>/dev/null || true
# Clean up empty directories
find /data/.cache/actcache/cache -type d -empty -delete 2>/dev/null || true
echo "actcache size: $(du -sh /data/.cache/actcache/cache 2>/dev/null | cut -f1)"
' || true
echo ""
done
echo "GC complete"
resources:
requests:
cpu: 50m
memory: 64Mi
limits:
cpu: 250m
memory: 128Mi
{{- end }}
+2 -9
View File
@@ -2,15 +2,8 @@
# runner instance. Only runner.name and runner.labels differ -- everything
# else (image, dind, persistence, tolerations, nodeSelector) is shared.
#
# Label image: node:22-bookworm — Debian, root, apt-get, Node.js, npm, git.
# Install docker in workflow steps as needed.
# node:22-bookworm ships Node natively, so unlike the golang/rust instances,
# jobs on this runner need no "install node" step before actions/checkout.
runner:
image:
repository: code.forgejo.org/forgejo/runner
tag: "6"
name: node-runner
labels: "node:docker://node:22-bookworm"
# GC CronJob renders only from the default (golang) values to avoid duplicates
gc:
enabled: false
+4 -12
View File
@@ -2,17 +2,9 @@
# runner instance. Only runner.name and runner.labels differ -- everything
# else (image, dind, persistence, tolerations, nodeSelector) is shared.
#
# Label image: rust:1-bookworm — Debian, root, apt-get, Rust, cargo, git.
# Install Node.js/docker in workflow steps as needed.
# rust:1.83-bookworm -- verified this tag exists (docker manifest inspect)
# before pinning it, per this repo's convention of not trusting a tag exists
# without checking.
runner:
image:
repository: code.forgejo.org/forgejo/runner
tag: "6"
name: rust-runner
labels: "rust:docker://rust:1-bookworm"
# GC CronJob renders only from the default (golang) values to avoid duplicates
gc:
enabled: false
labels: "rust:docker://rust:1.83-bookworm"
+16 -20
View File
@@ -1,13 +1,20 @@
runner:
image:
repository: code.forgejo.org/forgejo/runner
tag: "6"
tag: "6" # pin exact release before apply
name: golang-runner
# Label image is what workflow steps run in (NOT the runner daemon image).
# golang:1.26-bookworm: Debian, root, apt-get, Go, git.
# TODO: Switch to custom image once build-runner-images.yml pushes images
labels: "golang:docker://golang:1.26-bookworm"
# Default image is only used when a job's `container:` doesn't override it
# (both ci.yaml and build.yaml in homelab-frontend do). Retired the old
# "docker" label entirely; every repo this runner serves is Go, so this
# instance carries the golang toolchain and its own dind sidecar builds and
# pushes that repo's images too -- there is no separate generic runner
# anymore.
labels: "golang:docker://golang:1.25-bookworm"
# In-cluster Service (:3000) — direct, avoids the ingress/public-hostname hop
# (the public URL is :443 which forgejo doesn't serve; runner got i/o timeout).
forgejoUrl: http://forgejo-gitea-http.cicd.svc.cluster.local:3000
# tokenSecret: name of the K8s Secret that holds the runner registration token
# created automatically by the helmfile presync hook (see helmfile.yaml.gotmpl)
tokenSecret: runner-token
resources:
requests:
@@ -32,7 +39,7 @@ dind:
persistence:
reg:
storageClass: longhorn # Unified StorageClass (3 replicas)
size: 20Gi # .runner registration file + action tool cache + actcache artifacts
size: 1Gi # .runner registration file + config — survives pod restarts
dind:
storageClass: longhorn # Unified StorageClass (3 replicas)
size: 30Gi # docker layer cache — keeps rebuilds fast across restarts
@@ -42,18 +49,7 @@ tolerations:
operator: Exists
effect: NoSchedule
# Pin to az-b (talos-cp-2) — more Longhorn storage than az-a (worker-1 over-provisioned).
# RWO PVCs will recreate on talos-cp-2 when nodeSelector changes.
# Pin to az-a (talos-cp-1) — sole Longhorn node; RWO PVCs (reg/dind cache) only
# attach there.
nodeSelector:
topology.kubernetes.io/zone: az-b
# GC CronJob — prunes Docker images/volumes/build-cache and actcache across
# ALL forgejo-runner pods. Only enable in default values (golang instance);
# disable in per-runner overrides so it renders once.
gc:
enabled: true
schedule: "*/30 * * * *" # every 30 minutes
image: alpine/k8s:1.31.0
pruneAge: "30m" # Docker artifacts unused longer than this get pruned
pruneAgeHours: 0.5 # Same as pruneAge but numeric for date arithmetic in shell
actcacheMaxAgeDays: 1 # actcache files older than N days (aggressive for heavy Rust cargo builds)
topology.kubernetes.io/zone: az-a
+167
View File
@@ -0,0 +1,167 @@
# Authentik OAuth provisioning — PostSync hook, reruns on every ArgoCD sync
# (hook-delete-policy: BeforeHookCreation deletes the previous run's Job before
# creating a new one, so this stays reconciled the same way the rest of the
# cluster does — no separate manual bootstrap step like setup_talos_iam.sh /
# provision_oidc.py, which never got migrated off the old helmfile workflow).
#
# What it does (see scripts/authentik-provision.py docstring): creates the
# "groups" scope mapping, homelab-admins / grafana-admins groups, the "rock"
# admin user, OAuth2 providers + Applications for grafana/minio/forgejo/argocd,
# and binds homelab-admins to all of them. The script is generated into the
# authentik-provision-script ConfigMap by kustomize configMapGenerator (see
# kustomization.yaml), not embedded here.
#
# RBAC: this Job only touches Secrets (get existing client secrets, create new
# ones for forgejo/argocd/rock) across the namespaces those services live in.
# It never touches any other resource type.
apiVersion: v1
kind: ServiceAccount
metadata:
name: authentik-provisioner
namespace: iam
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: authentik-provisioner
rules:
- apiGroups: [""]
resources: ["secrets"]
verbs: ["get", "list", "create", "update", "patch"]
---
# One RoleBinding per namespace the script touches (least-privilege: Secrets
# only, and only in these 5 namespaces — not a cluster-wide ClusterRoleBinding).
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: authentik-provisioner
namespace: iam
subjects:
- kind: ServiceAccount
name: authentik-provisioner
namespace: iam
roleRef:
kind: ClusterRole
name: authentik-provisioner
apiGroup: rbac.authorization.k8s.io
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: authentik-provisioner
namespace: cicd
subjects:
- kind: ServiceAccount
name: authentik-provisioner
namespace: iam
roleRef:
kind: ClusterRole
name: authentik-provisioner
apiGroup: rbac.authorization.k8s.io
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: authentik-provisioner
namespace: argocd
subjects:
- kind: ServiceAccount
name: authentik-provisioner
namespace: iam
roleRef:
kind: ClusterRole
name: authentik-provisioner
apiGroup: rbac.authorization.k8s.io
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: authentik-provisioner
namespace: logging
subjects:
- kind: ServiceAccount
name: authentik-provisioner
namespace: iam
roleRef:
kind: ClusterRole
name: authentik-provisioner
apiGroup: rbac.authorization.k8s.io
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: authentik-provisioner
namespace: storage
subjects:
- kind: ServiceAccount
name: authentik-provisioner
namespace: iam
roleRef:
kind: ClusterRole
name: authentik-provisioner
apiGroup: rbac.authorization.k8s.io
---
apiVersion: batch/v1
kind: Job
metadata:
name: authentik-provision
namespace: iam
annotations:
argocd.argoproj.io/hook: PostSync
argocd.argoproj.io/hook-delete-policy: BeforeHookCreation
spec:
ttlSecondsAfterFinished: 600
backoffLimit: 3
template:
spec:
serviceAccountName: authentik-provisioner
restartPolicy: Never
securityContext:
runAsNonRoot: true
runAsUser: 1000
seccompProfile:
type: RuntimeDefault
containers:
- name: provision
image: python:3.12-alpine
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
env:
- name: AUTHENTIK_BOOTSTRAP_TOKEN
valueFrom:
secretKeyRef:
name: authentik-secrets
key: AUTHENTIK_BOOTSTRAP_TOKEN
volumeMounts:
- name: script
mountPath: /script
command:
- /bin/sh
- -c
- |
set -e
echo "waiting for authentik-server..."
until wget -q -O /dev/null http://authentik-server.iam.svc.cluster.local/-/health/ready/ 2>/dev/null; do
sleep 5
done
echo "installing kubectl (via python urllib - no apk/curl: this"
echo "container runs as non-root UID 1000 and can't write to"
echo "apk's directories or /usr/local/bin, both root-owned in"
echo "the python:3.12-alpine image; /tmp is world-writable)..."
python3 -c "
import urllib.request, os, stat
kver = urllib.request.urlopen('https://dl.k8s.io/release/stable.txt').read().decode().strip()
url = f'https://dl.k8s.io/release/{kver}/bin/linux/amd64/kubectl'
urllib.request.urlretrieve(url, '/tmp/kubectl')
st = os.stat('/tmp/kubectl')
os.chmod('/tmp/kubectl', st.st_mode | stat.S_IEXEC)
"
export PATH="/tmp:$PATH"
echo "running provisioning script..."
python3 /script/authentik-provision.py
volumes:
- name: script
configMap:
name: authentik-provision-script
+30 -7
View File
@@ -1,12 +1,35 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
# NOTE: no top-level `namespace:` transformer here (removed) - it used to
# force-rewrite metadata.namespace to "iam" on every resource in this
# kustomization, which was harmless while every manifest here only ever
# targeted the iam namespace itself. authentik-provision-job.yaml's
# RoleBindings deliberately target cicd/argocd/logging/storage (least-
# privilege access for the authentik-provisioner ServiceAccount to touch
# Secrets in those namespaces) - the namespace transformer would have
# silently rewritten all of them back to iam, breaking the RBAC. Every
# manifest in this directory already sets its own explicit
# metadata.namespace, so dropping the transformer changes nothing for the
# existing resources/.
resources:
- authentik-provision-job.yaml
- rbac-dashboard-rolebinding.yaml
# IAM provisioning is manual-only (security-sensitive).
# Script: scripts/iam/authentik-provision.py
# Run:
# export AUTHENTIK_BOOTSTRAP_TOKEN=$(kubectl -n iam get secret authentik-secrets \
# -o jsonpath='{.data.AUTHENTIK_BOOTSTRAP_TOKEN}' | base64 -d)
# python3 scripts/iam/authentik-provision.py
# Provisioning/verification python lives in scripts/*.py (real files, linted +
# diff-friendly) and is generated into ConfigMaps here rather than embedded in
# the job YAML. disableNameSuffixHash keeps the names stable so the Jobs'
# configMap volume refs and PostSync hook-delete semantics keep working; each
# hook Job is recreated per sync so it always mounts the latest script.
configMapGenerator:
- name: authentik-provision-script
namespace: iam
files:
- authentik-provision.py=scripts/authentik-provision.py
generatorOptions:
disableNameSuffixHash: true
# authentik-migrations-job.yaml removed — redundant + broken. The authentik
# `server` entrypoint runs migrations itself; this standalone job lacked the
# authentik-secrets envFrom (Secret key missing) and always failed.
# SOPS secrets (*.enc.yaml) handled by ArgoCD SOPS plugin at sync time
# authentik/vault deployed via ArgoCD Helm source
@@ -0,0 +1,400 @@
#!/usr/bin/env python3
"""
Authentik OAuth provisioning - idempotent, safe to re-run (ArgoCD PostSync hook).
Creates/updates, in order:
1. A custom "groups" OAuth2 scope mapping (Authentik ships openid/email/profile
by default but NOT groups - required for ArgoCD RBAC group mapping and
Grafana's role_attribute_path, both of which read a `groups` claim).
2. Groups: homelab-admins (is_superuser=true), grafana-admins.
3. User "rock": created if missing, always (re-)synced into both groups above.
Password is generated once and only written to the k8s Secret
rock-credentials (iam ns) the first time the user is created - re-runs
never rotate an existing password.
4. OAuth2/OIDC providers + Applications for: grafana, minio, forgejo, argocd.
Client secrets are read from existing k8s Secrets (grafana-oidc, minio-oidc)
if present, or generated once and written out (forgejo-oidc, oidc-secret)
the first time.
5. PolicyBinding of homelab-admins -> every Application above, so "rock" (and
anyone else in that group) has guaranteed access regardless of each app's
default visibility.
Talks to Authentik over the in-cluster Service (authentik-server.iam.svc:80),
authenticating with the bootstrap token. Everything is done with GET-then-
create-or-patch so this can be re-run on every ArgoCD sync without duplicating
or clobbering objects (PostSync hook, not a one-shot Job with hook-delete).
kubectl is used only to read/write the small set of Secrets this script
touches - it shells out rather than using the Python k8s client to keep the
container image to stdlib Python + the kubectl binary, no pip installs.
"""
import json
import os
import secrets
import string
import subprocess
import sys
import urllib.error
import urllib.request
AUTHENTIK_URL = "http://authentik-server.iam.svc.cluster.local"
TOKEN = os.environ["AUTHENTIK_BOOTSTRAP_TOKEN"]
def api(method, path, data=None):
url = f"{AUTHENTIK_URL}{path}"
body = json.dumps(data).encode() if data is not None else None
req = urllib.request.Request(
url,
data=body,
method=method,
headers={
"Authorization": f"Bearer {TOKEN}",
"Content-Type": "application/json",
},
)
try:
with urllib.request.urlopen(req, timeout=30) as resp:
raw = resp.read()
return resp.status, (json.loads(raw) if raw else {})
except urllib.error.HTTPError as e:
raw = e.read()
try:
parsed = json.loads(raw) if raw else {}
except json.JSONDecodeError:
parsed = {"raw": raw.decode(errors="replace")}
return e.code, parsed
def die(msg):
print(f"FATAL: {msg}", file=sys.stderr)
sys.exit(1)
def gen_secret(n=40):
alphabet = string.ascii_letters + string.digits
return "".join(secrets.choice(alphabet) for _ in range(n))
def kubectl_get_secret_key(namespace, name, key):
"""Returns decoded value, or None if the secret/key doesn't exist."""
p = subprocess.run(
["kubectl", "-n", namespace, "get", "secret", name, "-o", f"jsonpath={{.data.{key}}}"],
capture_output=True, text=True,
)
if p.returncode != 0 or not p.stdout.strip():
return None
import base64
return base64.b64decode(p.stdout).decode()
def kubectl_create_secret(namespace, name, literals: dict, labels: dict = None):
"""Idempotent: create-or-update via dry-run|apply, same pattern used
elsewhere in this repo (setup_vault.sh, apply-vault-secrets.sh)."""
args = ["kubectl", "-n", namespace, "create", "secret", "generic", name]
for k, v in literals.items():
args += [f"--from-literal={k}={v}"]
args += ["--dry-run=client", "-o", "yaml"]
render = subprocess.run(args, capture_output=True, text=True)
if render.returncode != 0:
die(f"rendering secret {namespace}/{name}: {render.stderr}")
apply = subprocess.run(["kubectl", "apply", "-f", "-"], input=render.stdout,
capture_output=True, text=True)
if apply.returncode != 0:
die(f"applying secret {namespace}/{name}: {apply.stderr}")
print(f" secret {namespace}/{name}: {apply.stdout.strip()}")
if labels:
# argocd's `$secret:key` substitution only reads Secrets carrying
# app.kubernetes.io/part-of: argocd — without it OIDC login fails with
# oauth2 "invalid_client" (empty client_secret sent to the IdP).
label_args = ["kubectl", "-n", namespace, "label", "secret", name,
"--overwrite"] + [f"{k}={v}" for k, v in labels.items()]
subprocess.run(label_args, capture_output=True, text=True)
def get_or_create(list_path, create_path, query, payload, patch_existing=None):
status, res = api("GET", f"{list_path}?{query}")
if status != 200:
die(f"GET {list_path}?{query} -> {status} {res}")
results = res.get("results", [])
if results:
obj = results[0]
if patch_existing:
status, obj2 = api("PATCH", f"{create_path}{obj['pk']}/", patch_existing)
if status not in (200, 201):
die(f"PATCH {create_path}{obj['pk']}/ -> {status} {obj2}")
return obj2
return obj
status, obj = api("POST", create_path, payload)
if status not in (200, 201):
die(f"POST {create_path} -> {status} {obj}")
return obj
# -----------------------------------------------------------------------------
print("[1/5] Ensuring custom 'groups' scope mapping exists...")
groups_mapping = get_or_create(
"/api/v3/propertymappings/provider/scope/",
"/api/v3/propertymappings/provider/scope/",
"scope_name=groups",
{
"name": "homelab: groups claim",
"scope_name": "groups",
# request.user.ak_groups is deprecated in authentik 2026.x (logs a
# deprecation warning on every token issue) -> use request.user.groups.
"expression": (
"return {\"groups\": [group.name for group in request.user.groups.all()]}"
),
},
# Force the expression onto the already-created mapping on re-run.
patch_existing={
"expression": (
"return {\"groups\": [group.name for group in request.user.groups.all()]}"
),
},
)
GROUPS_MAPPING_PK = groups_mapping["pk"]
# MinIO maps OIDC users to a MinIO policy via a "policy" claim
# (MINIO_IDENTITY_OPENID_CLAIM_NAME=policy). Emit consoleAdmin (full admin) for
# homelab-admins members, readonly for everyone else. Without this claim MinIO
# assigns no policy and OIDC users get no access.
_POLICY_EXPR = (
"return {\"policy\": \"consoleAdmin\" "
"if request.user.ak_groups.filter(name=\"homelab-admins\").exists() "
"else \"readonly\"}"
)
policy_mapping = get_or_create(
"/api/v3/propertymappings/provider/scope/",
"/api/v3/propertymappings/provider/scope/",
"scope_name=minio",
{
"name": "homelab: minio policy claim",
"scope_name": "minio",
"expression": _POLICY_EXPR,
},
patch_existing={"expression": _POLICY_EXPR},
)
POLICY_MAPPING_PK = policy_mapping["pk"]
# Fetch the standard openid/email/profile mapping pks (shipped by default).
status, res = api("GET", "/api/v3/propertymappings/provider/scope/")
by_scope = {m["scope_name"]: m["pk"] for m in res["results"]}
SCOPE_PKS = [by_scope["openid"], by_scope["email"], by_scope["profile"], GROUPS_MAPPING_PK]
status, res = api("GET", "/api/v3/flows/instances/?slug=default-provider-authorization-implicit-consent")
AUTHORIZATION_FLOW_PK = res["results"][0]["pk"]
status, res = api("GET", "/api/v3/flows/instances/?slug=default-provider-invalidation-flow")
INVALIDATION_FLOW_PK = res["results"][0]["pk"]
status, res = api("GET", "/api/v3/crypto/certificatekeypairs/?has_key=true")
SIGNING_KEY_PK = res["results"][0]["pk"]
# -----------------------------------------------------------------------------
print("[2/5] Ensuring groups homelab-admins / grafana-admins exist...")
homelab_admins = get_or_create(
"/api/v3/core/groups/", "/api/v3/core/groups/",
"name=homelab-admins",
{"name": "homelab-admins", "is_superuser": True},
)
grafana_admins = get_or_create(
"/api/v3/core/groups/", "/api/v3/core/groups/",
"name=grafana-admins",
{"name": "grafana-admins", "is_superuser": False},
)
# -----------------------------------------------------------------------------
print("[3/5] Ensuring user 'rock' exists with admin group membership...")
status, res = api("GET", "/api/v3/core/users/?username=rock")
rock_password = None
if res.get("results"):
rock = res["results"][0]
status, rock = api("PATCH", f"/api/v3/core/users/{rock['pk']}/", {
"groups": [homelab_admins["pk"], grafana_admins["pk"]],
"is_active": True,
# email is REQUIRED: Grafana's OIDC login reads the email claim from
# userinfo; an empty email makes Grafana fall back to a GitHub-style
# <userinfo>/emails call, which Authentik 404s -> login fails entirely.
"email": "[email protected]",
})
if status not in (200, 201):
die(f"PATCH user rock -> {status} {rock}")
print(" rock already exists, group membership synced (password unchanged)")
else:
rock_password = gen_secret(24)
status, rock = api("POST", "/api/v3/core/users/", {
"username": "rock",
"name": "Rock",
"is_active": True,
# Required for Grafana OIDC (see PATCH branch above).
"email": "[email protected]",
"groups": [homelab_admins["pk"], grafana_admins["pk"]],
"path": "users",
"type": "internal",
})
if status not in (200, 201):
die(f"POST user rock -> {status} {rock}")
status, pw_res = api("POST", f"/api/v3/core/users/{rock['pk']}/set_password/",
{"password": rock_password})
if status not in (200, 204):
die(f"set_password for rock -> {status} {pw_res}")
kubectl_create_secret("iam", "rock-credentials", {
"username": "rock",
"password": rock_password,
})
print(" rock created, credentials stored in iam/rock-credentials")
# -----------------------------------------------------------------------------
print("[4/5] Ensuring OAuth2 providers + applications for grafana/minio/forgejo/argocd...")
SERVICES = {
"grafana": {
"client_secret_source": ("logging", "grafana-oidc", "GF_AUTH_GENERIC_OAUTH_CLIENT_SECRET"),
"redirect_uris": ["https://grafana.riotpiao.com/login/generic_oauth"],
"launch_url": "https://grafana.riotpiao.com",
"display_name": "Grafana",
},
"minio": {
"client_secret_source": ("storage", "minio-oidc", "MINIO_IDENTITY_OPENID_CLIENT_SECRET"),
"redirect_uris": ["https://minio.riotpiao.com/oauth_callback"],
"launch_url": "https://minio.riotpiao.com",
"display_name": "MinIO",
},
"forgejo": {
# No secret exists yet for forgejo - generate + store on first run.
"client_secret_source": ("cicd", "forgejo-oidc", "CLIENT_SECRET"),
"generate_if_missing": True,
"redirect_uris": [
"https://forgejo.riotpiao.com/user/oauth2/authentik/callback",
"https://forgejo.riotpiao.com/user/oauth2/openidconnect/callback",
],
"launch_url": "https://forgejo.riotpiao.com",
"display_name": "Forgejo",
},
"argocd": {
# oidc-secret uses hyphenated keys (client-id/client-secret) per
# argocd-values.yaml's `$oidc-secret:client-id` / `:client-secret` refs.
"client_secret_source": ("argocd", "oidc-secret", "client-secret"),
"generate_if_missing": True,
"extra_secret_literals": {"client-id": "argocd"},
# argocd only reads $secret refs from Secrets labelled part-of: argocd.
"secret_labels": {"app.kubernetes.io/part-of": "argocd"},
"redirect_uris": ["https://argocd.riotpiao.com/auth/callback"],
"launch_url": "https://argocd.riotpiao.com",
"display_name": "Argo CD",
},
"homarr": {
"client_secret_source": ("dashboard", "homarr-oidc", "client-secret"),
"generate_if_missing": True,
"extra_secret_literals": {"client-id": "homarr"},
"redirect_uris": ["https://homarr.riotpiao.com/api/auth/callback/oidc"],
"launch_url": "https://homarr.riotpiao.com",
"display_name": "Homarr",
},
}
app_pks_for_binding = []
for name, cfg in SERVICES.items():
# MinIO also needs the "policy" claim (via the minio scope mapping) so its
# MINIO_IDENTITY_OPENID_CLAIM_NAME=policy maps homelab-admins -> consoleAdmin.
provider_mappings = SCOPE_PKS + ([POLICY_MAPPING_PK] if name == "minio" else [])
ns, secret_name, key = cfg["client_secret_source"]
client_secret = kubectl_get_secret_key(ns, secret_name, key)
if client_secret is None:
if not cfg.get("generate_if_missing"):
print(f" WARNING: {ns}/{secret_name} key {key} not found and "
f"generate_if_missing not set for '{name}' - skipping provider/app")
continue
client_secret = gen_secret(40)
literals = {key: client_secret}
literals.update(cfg.get("extra_secret_literals", {}))
kubectl_create_secret(ns, secret_name, literals,
labels=cfg.get("secret_labels"))
print(f" {name}: generated new client secret -> {ns}/{secret_name}")
else:
print(f" {name}: using existing client secret from {ns}/{secret_name}")
provider = get_or_create(
"/api/v3/providers/oauth2/", "/api/v3/providers/oauth2/",
f"name={name}",
{
"name": name,
"client_id": name,
"client_secret": client_secret,
"client_type": "confidential",
"authorization_flow": AUTHORIZATION_FLOW_PK,
"invalidation_flow": INVALIDATION_FLOW_PK,
"signing_key": SIGNING_KEY_PK,
"property_mappings": provider_mappings,
"sub_mode": "hashed_user_id",
"include_claims_in_id_token": True,
# authentik 2026.x requires grant_types to be set explicitly; the
# API defaults it to [] when omitted, which makes /authorize reject
# every login with "Invalid grant_type for provider" ->
# invalid_request. authorization_code = the web SSO flow all these
# apps use; refresh_token = long-lived sessions (offline_access).
"grant_types": ["authorization_code", "refresh_token"],
"redirect_uris": [
{"matching_mode": "strict", "url": u} for u in cfg["redirect_uris"]
],
},
# Keep the redirect_uris/mappings/grant_types in sync on re-run, but
# never touch client_secret again once created (that's the source of
# truth in the k8s Secret, and re-sending it here is harmless anyway).
patch_existing={
"property_mappings": provider_mappings,
"grant_types": ["authorization_code", "refresh_token"],
"redirect_uris": [
{"matching_mode": "strict", "url": u} for u in cfg["redirect_uris"]
],
},
)
# superuser_full_list=true is REQUIRED on the LIST: the applications list
# applies access-policy filtering to the results array (these apps are bound
# to homelab-admins, and the bootstrap-token user akadmin is not a member),
# so without it the GET returns an empty results list even though the app
# exists -> fall through to POST -> 400 "already exists".
#
# We deliberately do NOT patch_existing here: the application DETAIL endpoint
# (PATCH /applications/{pk}/) enforces the same access policy and does NOT
# honor superuser_full_list, so PATCH-by-pk returns 404 for akadmin once the
# homelab-admins binding exists. That 404 aborted the loop before later
# providers got their grant_types. slug/provider/launch_url are set at
# creation and are stable (provider is get_or_create'd by name, stable pk),
# so find-or-create is sufficient.
application = get_or_create(
"/api/v3/core/applications/", "/api/v3/core/applications/",
f"slug={name}&superuser_full_list=true",
{
"name": cfg["display_name"],
"slug": name,
"provider": provider["pk"],
"meta_launch_url": cfg["launch_url"],
},
)
app_pks_for_binding.append((name, application["pk"]))
print(f" {name}: provider pk={provider['pk']} application pk={application['pk']}")
# -----------------------------------------------------------------------------
print("[5/5] Binding homelab-admins to every application (guaranteed access for rock)...")
for name, app_pk in app_pks_for_binding:
get_or_create(
"/api/v3/policies/bindings/", "/api/v3/policies/bindings/",
f"target={app_pk}&group={homelab_admins['pk']}",
{
"target": app_pk,
"group": homelab_admins["pk"],
"order": 0,
"enabled": True,
},
)
print(f" {name}: homelab-admins bound")
print("\nDone. Summary:")
print(" groups: homelab-admins (superuser), grafana-admins")
print(" user: rock -> homelab-admins + grafana-admins")
print(f" apps: {', '.join(n for n, _ in app_pks_for_binding)}")
if rock_password:
print(" NOTE: rock's password was generated this run - see")
print(" kubectl -n iam get secret rock-credentials -o jsonpath='{.data.password}' | base64 -d")
+5 -9
View File
@@ -44,10 +44,6 @@ grafana.ini:
server:
root_url: https://grafana.riotpiao.com
# Allow embedding dashboards in iframes (NextJS integration)
security:
allow_embedding: true
# No anonymous read access — every user must log in via Authentik SSO.
auth.anonymous:
enabled: false
@@ -68,14 +64,15 @@ grafana.ini:
# doesn't return localhost redirects in its token responses.
#
# role_attribute_path: JMESPath expression evaluated against the userinfo
# response. akadmin gets GrafanaAdmin (server admin, can impersonate);
# homelab-admins members get Admin (org admin); everyone else Viewer.
# response. Members of the 'grafana-admins' Authentik group get Admin role;
# everyone else gets Viewer. The group name must match exactly what Authentik
# sends in the 'groups' claim.
auth.generic_oauth:
enabled: true
name: Authentik
allow_sign_up: true
client_id: grafana
scopes: openid email profile groups
scopes: openid email profile
auth_url: https://authentik.riotpiao.com/application/o/authorize/
token_url: https://authentik.riotpiao.com/application/o/token/
api_url: https://authentik.riotpiao.com/application/o/userinfo/
@@ -86,8 +83,7 @@ grafana.ini:
email_attribute_path: email
login_attribute_path: preferred_username
name_attribute_path: name
role_attribute_path: "preferred_username == 'akadmin' && 'GrafanaAdmin' || contains(groups[*], 'homelab-admins') && 'Admin' || 'Viewer'"
allow_assign_grafana_admin: true
role_attribute_path: "contains(groups[*], 'grafana-admins') && 'Admin' || 'Viewer'"
use_pkce: false
use_refresh_token: false
skip_org_role_sync: false
-2
View File
@@ -4,13 +4,11 @@ namespace: longhorn-system
resources:
- longhorn-storageclass.yaml
- longhorn-cnpg-storageclass.yaml # CNPG-specific with postgres UID/GID
- longhorn-paperless-storageclass.yaml # single-replica, cp-3 USB HDD only
- longhorn-servicemonitor.yaml
- longhorn-taint-toleration.yaml
- longhorn-nodes.yaml
- expand-replicas-job.yaml
- patch-csi-tolerations-job.yaml
- longhorn-add-disks-job.yaml # Add extra disks to talos-cp-2
# Longhorn deployed via bootstrap script or Helm.
# These manifests configure it: unified StorageClass (default, 3 replicas),
# Prometheus ServiceMonitor, taint toleration for control-plane nodes, explicit
@@ -1,130 +0,0 @@
# PostSync hook Job that adds extra disks to Longhorn nodes.
# talos-cp-2 has 4 extra disks mounted at /var/lib/longhorn-disk{1,2,3,4}
# that are NOT auto-discovered by Longhorn.
#
# talos-cp-3 additionally gets a tagged disk for paperless-ngx media, backed by
# the 4TB USB HDD (/dev/sdg) — tagged "paperless-media" so only the dedicated
# longhorn-paperless-media StorageClass (diskSelector match) can place replicas
# there, keeping it out of the default 3-replica pool. This patch is inert
# until the Terraform machine-config change mounts the disk at
# /var/lib/longhorn-paperless-media (pending — see terraform.tfvars.local,
# not present in this checkout); Longhorn just reports the disk not-ready
# until the path exists, no harm in applying it early.
apiVersion: batch/v1
kind: Job
metadata:
name: longhorn-add-disks
namespace: longhorn-system
annotations:
argocd.argoproj.io/hook: PostSync
argocd.argoproj.io/hook-delete-policy: BeforeHookCreation
spec:
backoffLimit: 3
template:
metadata:
name: longhorn-add-disks
spec:
restartPolicy: Never
serviceAccountName: longhorn-expand-replicas
containers:
- name: add-disks
image: bitnami/kubectl:latest
command:
- /bin/bash
- -c
- |
set -euo pipefail
echo "Adding extra disks to Longhorn nodes..."
# talos-cp-2 has 4 extra disks:
# - /var/lib/longhorn-disk1 (sdb1, 600GB)
# - /var/lib/longhorn-disk2 (sdc1, 500GB)
# - /var/lib/longhorn-disk3 (sdd1, 160GB)
# - /var/lib/longhorn-disk4 (sde1, 2TB)
echo "Checking talos-cp-2..."
CURRENT_DISKS=$(kubectl -n longhorn-system get nodes.longhorn.io talos-cp-2 -o json | jq -r '.spec.disks | keys | length')
echo " Current disk count: $CURRENT_DISKS"
if [ "$CURRENT_DISKS" -lt 5 ]; then
echo " Adding extra disks to talos-cp-2..."
kubectl -n longhorn-system patch nodes.longhorn.io talos-cp-2 --type merge -p '{
"spec": {
"disks": {
"disk1": {
"allowScheduling": true,
"diskType": "filesystem",
"evictionRequested": false,
"path": "/var/lib/longhorn-disk1",
"storageReserved": 10737418240,
"tags": []
},
"disk2": {
"allowScheduling": true,
"diskType": "filesystem",
"evictionRequested": false,
"path": "/var/lib/longhorn-disk2",
"storageReserved": 10737418240,
"tags": []
},
"disk3": {
"allowScheduling": true,
"diskType": "filesystem",
"evictionRequested": false,
"path": "/var/lib/longhorn-disk3",
"storageReserved": 5368709120,
"tags": []
},
"disk4": {
"allowScheduling": true,
"diskType": "filesystem",
"evictionRequested": false,
"path": "/var/lib/longhorn-disk4",
"storageReserved": 21474836480,
"tags": []
}
}
}
}'
echo " ✓ Disks added to talos-cp-2"
else
echo " ✓ talos-cp-2 already has $CURRENT_DISKS disks configured"
fi
echo "Checking talos-cp-3..."
CP3_DISKS=$(kubectl -n longhorn-system get nodes.longhorn.io talos-cp-3 -o json | jq -r '.spec.disks | keys | length')
echo " Current disk count: $CP3_DISKS"
if [ "$CP3_DISKS" -lt 2 ]; then
echo " Adding paperless-media disk to talos-cp-3..."
kubectl -n longhorn-system patch nodes.longhorn.io talos-cp-3 --type merge -p '{
"spec": {
"disks": {
"paperless-media": {
"allowScheduling": true,
"diskType": "filesystem",
"evictionRequested": false,
"path": "/var/lib/longhorn-paperless-media",
"storageReserved": 0,
"tags": ["paperless-media"]
}
}
}
}'
echo " ✓ paperless-media disk added to talos-cp-3"
else
echo " ✓ talos-cp-3 already has $CP3_DISKS disks configured"
fi
echo
echo "Waiting for disks to be ready..."
sleep 10
echo
echo "Final storage summary:"
kubectl -n longhorn-system get nodes.longhorn.io -o json | \
jq -r '.items[] | "\(.metadata.name): \(.status.diskStatus | to_entries | map("\(.key): \(.value.storageAvailable/1073741824 | floor)GB avail") | join(", "))"'
echo
echo "Done."
@@ -1,32 +0,0 @@
# Dedicated StorageClass for paperless-ngx document media, backed by the 4TB
# USB HDD on talos-cp-3 (see longhorn-add-disks-job.yaml). Single disk, single
# node — no Longhorn replica is possible, so numberOfReplicas is 1 by
# necessity, not choice. diskSelector pins placement to the tagged disk only,
# so a volume from this class never lands on cp-3's regular (already
# DiskPressure) default pool. reclaimPolicy is Retain, not Delete: a PVC
# accident here has no replica to fall back on, so an accidental delete must
# not also take the underlying volume with it.
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: longhorn-paperless-media
annotations:
# StorageClass.parameters is immutable - any future edit here needs
# delete+recreate, not patch. Same fix as longhorn-cnpg-storageclass.yaml;
# safe since a StorageClass is only consulted at provisioning time.
argocd.argoproj.io/sync-options: Replace=true,Force=true
provisioner: driver.longhorn.io
allowVolumeExpansion: true
reclaimPolicy: Retain
volumeBindingMode: WaitForFirstConsumer
parameters:
numberOfReplicas: "1"
# diskSelector alone is sufficient: only cp-3 has a disk tagged
# paperless-media, so placement is already pinned. A nodeSelector value
# here must be a Longhorn *node* tag (set via nodes.longhorn.io spec.tags),
# not a Kubernetes hostname - "talos-cp-3" was never a node tag, which made
# every provision attempt fail with "specified node tag talos-cp-3 does
# not exist".
diskSelector: "paperless-media"
staleReplicaTimeout: "30"
fsType: "ext4"
+1 -7
View File
@@ -1,14 +1,8 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
# NOTE: no top-level `namespace:` transformer (see iam/kustomization.yaml for
# the same fix) - minio-provision-paperless-job.yaml's RoleBinding
# deliberately targets namespace paperless (least-privilege access for the
# provisioner ServiceAccount to write Secrets there); a namespace transformer
# would silently rewrite it back to storage, breaking the RBAC. Every resource
# here already sets its own explicit metadata.namespace.
namespace: storage
resources:
- minio-tenant.yaml
- minio-provision-paperless-job.yaml
# The operator creates the minio S3/console/headless Services and the
# declarative bucket + user from the Tenant spec — no hand-rolled Service or
# Bucket/User CRs (those kinds don't exist in the operator CRD set).
@@ -1,120 +0,0 @@
# PostSync hook Job: creates a MinIO IAM user + policy scoped to only the
# `paperless` bucket (least-privilege — reuses root creds nowhere else in the
# cluster), then writes the generated access/secret key into a Secret in the
# `paperless` namespace for the nightly backup CronJob to consume.
#
# Idempotent: re-running never rotates existing credentials — if
# paperless-minio-creds already exists in ns paperless, the script reuses the
# access key it already wrote and just re-asserts the policy/user exist.
apiVersion: v1
kind: ServiceAccount
metadata:
name: minio-paperless-provisioner
namespace: storage
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: minio-paperless-provisioner
rules:
- apiGroups: [""]
resources: ["secrets"]
verbs: ["get", "list", "create", "update", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: minio-paperless-provisioner
namespace: paperless
subjects:
- kind: ServiceAccount
name: minio-paperless-provisioner
namespace: storage
roleRef:
kind: ClusterRole
name: minio-paperless-provisioner
apiGroup: rbac.authorization.k8s.io
---
apiVersion: batch/v1
kind: Job
metadata:
name: minio-provision-paperless
namespace: storage
annotations:
argocd.argoproj.io/hook: PostSync
argocd.argoproj.io/hook-delete-policy: BeforeHookCreation
spec:
ttlSecondsAfterFinished: 600
backoffLimit: 3
template:
spec:
serviceAccountName: minio-paperless-provisioner
restartPolicy: Never
initContainers:
- name: kubectl-copy
image: bitnami/kubectl:latest
command: ["sh", "-c", "cp $(which kubectl) /shared/kubectl"]
volumeMounts:
- name: shared
mountPath: /shared
containers:
- name: provision
image: minio/mc:latest
volumeMounts:
- name: shared
mountPath: /shared
- name: minio-creds
mountPath: /minio-creds
readOnly: true
command:
- /bin/sh
- -c
- |
set -e
export PATH="/shared:$PATH"
# minio-creds ships as a shell-sourceable config.env blob
# (`export MINIO_ROOT_USER=... / MINIO_ROOT_PASSWORD=...`), not
# discrete keys — source it directly rather than re-parsing.
. /minio-creds/config.env
mc alias set m http://minio-cluster-hl.storage.svc.cluster.local:9000 \
"$MINIO_ROOT_USER" "$MINIO_ROOT_PASSWORD"
echo "Checking for existing paperless-minio-creds secret..."
if kubectl -n paperless get secret paperless-minio-creds >/dev/null 2>&1; then
ACCESS_KEY=$(kubectl -n paperless get secret paperless-minio-creds -o jsonpath='{.data.ACCESS_KEY}' | base64 -d)
SECRET_KEY=$(kubectl -n paperless get secret paperless-minio-creds -o jsonpath='{.data.SECRET_KEY}' | base64 -d)
echo " reusing existing credentials"
else
ACCESS_KEY="paperless"
SECRET_KEY=$(head -c 32 /dev/urandom | base64 | tr -d '/+=' | head -c 40)
echo " generated new credentials"
fi
echo "Ensuring MinIO user 'paperless' exists..."
if ! mc admin user info m "$ACCESS_KEY" >/dev/null 2>&1; then
mc admin user add m "$ACCESS_KEY" "$SECRET_KEY"
fi
echo "Writing scoped policy (paperless bucket only)..."
printf '%s' '{"Version":"2012-10-17","Statement":[{"Effect":"Allow","Action":["s3:ListBucket"],"Resource":["arn:aws:s3:::paperless"]},{"Effect":"Allow","Action":["s3:GetObject","s3:PutObject","s3:DeleteObject"],"Resource":["arn:aws:s3:::paperless/*"]}]}' > /tmp/paperless-rw-policy.json
mc admin policy create m paperless-rw /tmp/paperless-rw-policy.json || \
mc admin policy update m paperless-rw /tmp/paperless-rw-policy.json
mc admin policy attach m paperless-rw --user "$ACCESS_KEY"
echo "Writing paperless-minio-creds secret (ns paperless)..."
kubectl -n paperless create secret generic paperless-minio-creds \
--from-literal=ACCESS_KEY="$ACCESS_KEY" \
--from-literal=SECRET_KEY="$SECRET_KEY" \
--from-literal=BUCKET=paperless \
--from-literal=ENDPOINT=http://minio-cluster-hl.storage.svc.cluster.local:9000 \
--dry-run=client -o yaml | kubectl apply -f -
echo "Done."
volumes:
- name: shared
emptyDir: {}
- name: minio-creds
secret:
secretName: minio-creds
+2 -13
View File
@@ -76,7 +76,6 @@ spec:
- name: loki-ruler
- name: loki-admin
- name: vault
- name: paperless
# Metrics are exposed at /minio/v2/metrics; scrape via a hand-rolled
# ServiceMonitor in the monitoring stack rather than operator auto-wiring
@@ -84,25 +83,15 @@ spec:
# 'default' and fail the reconcile).
# Public hostnames the tenant serves (S3 + console via the cluster ingress).
# Must match actual ingress backends: the "minio" Ingress (host
# minio.riotpiao.com) routes to minio-cluster-console:9090, and "minio-api"
# (host minio-api.riotpiao.com) routes to minio:9000. Previously these were
# swapped, so MinIO's own console-domain validation rejected the OIDC
# callback arriving on minio.riotpiao.com's Host header.
features:
domains:
minio:
- https://minio-api.riotpiao.com
console: https://minio.riotpiao.com
- https://minio.riotpiao.com
console: https://minio-console.riotpiao.com
# ── OIDC via Authentik (server-side env, valid in v2 schema) ────────────────
# Use in-cluster URL for config fetch (pod→authentik); browser redirects use
# public URLs embedded in the OIDC metadata response (issuer stays public).
env:
- name: MINIO_IDENTITY_OPENID_CONFIG_URL
# Must use external URL — well-known response contains external issuer/jwks_uri.
# MinIO validates issuer in JWT matches well-known issuer. Internal URL = mismatch.
# Hairpins through ingress-nginx but stays in-cluster.
value: "https://authentik.riotpiao.com/application/o/minio/.well-known/openid-configuration"
- name: MINIO_IDENTITY_OPENID_CLIENT_ID
value: "minio"
@@ -1,160 +0,0 @@
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: api-gateway-alerts
namespace: monitoring
labels:
release: prometheus
spec:
groups:
# ================================================================
# SLA Targets (based on canary traffic baselines):
#
# Availability: 99.9% (43.8 min downtime/month)
# LLM Chat: p95 < 1s (qwen), p95 < 2s (reasoning), p95 < 5s (ornith)
# Embeddings: p95 < 500ms
# Rerank: p95 < 500ms
# Models list: p95 < 300ms
# Error rate: < 1% (5xx), < 5% (4xx excluding auth)
#
# Baselines from 200-request canary run:
# qwen p99=609ms, reasoning p99=328ms, embeddings p99=287ms,
# rerank p99=218ms, models p99=277ms
# SLA set at ~2x p99 for headroom.
# ================================================================
- name: api-gateway.availability
rules:
# Gateway pods not ready
- alert: APIGatewayDown
expr: sum(kube_pod_status_ready{namespace="api",condition="true"}) == 0
for: 1m
labels:
severity: critical
annotations:
summary: "API Gateway has zero ready pods"
# Gateway pod count below desired
- alert: APIGatewayDegraded
expr: |
sum(kube_pod_status_ready{namespace="api",condition="true"})
< kube_deployment_spec_replicas{namespace="api",deployment="api-gateway"}
for: 5m
labels:
severity: warning
annotations:
summary: "API Gateway {{ $value }} ready pods below desired replica count"
# Blackbox probe down
- alert: APIGatewayProbeDown
expr: probe_success{instance=~".*api.riotpiao.com.*"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: "API Gateway probe failed: {{ $labels.instance }}"
# LLM serving pods not ready
- alert: LLMServingDown
expr: sum(kube_pod_status_ready{namespace="llm-serving",condition="true"}) == 0
for: 2m
labels:
severity: critical
annotations:
summary: "All LLM serving pods down"
# Individual predictor down
- alert: LLMPredictorDown
expr: |
kube_deployment_status_replicas_ready{namespace="llm-serving"}
< kube_deployment_spec_replicas{namespace="llm-serving"}
for: 5m
labels:
severity: warning
annotations:
summary: "{{ $labels.deployment }} has {{ $value }} ready (below desired)"
- name: api-gateway.latency
# SLA: latency thresholds at ~2x measured p99
rules:
# Ingress-level latency (all requests through nginx)
- alert: APIGatewayLatencyHigh
expr: |
histogram_quantile(0.95,
sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress="api"}[5m])) by (le)
) > 2
for: 5m
labels:
severity: warning
annotations:
summary: "API Gateway p95 latency {{ $value | printf \"%.1f\" }}s (SLA: <2s)"
# Extreme latency (p99 > 5s)
- alert: APIGatewayLatencyCritical
expr: |
histogram_quantile(0.99,
sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress="api"}[5m])) by (le)
) > 5
for: 5m
labels:
severity: critical
annotations:
summary: "API Gateway p99 latency {{ $value | printf \"%.1f\" }}s (SLA: <5s)"
- name: api-gateway.errors
rules:
# 5xx error rate > 1%
- alert: APIGateway5xxErrorRate
expr: |
sum(rate(nginx_ingress_controller_requests{ingress="api",status=~"5.."}[5m]))
/ sum(rate(nginx_ingress_controller_requests{ingress="api"}[5m]))
> 0.01
for: 5m
labels:
severity: critical
annotations:
summary: "API Gateway 5xx rate {{ $value | humanizePercentage }} (SLA: <1%)"
# Total error rate > 10% (including 4xx)
- alert: APIGatewayHighErrorRate
expr: |
sum(rate(nginx_ingress_controller_requests{ingress="api",status=~"[45].."}[5m]))
/ sum(rate(nginx_ingress_controller_requests{ingress="api"}[5m]))
> 0.10
for: 10m
labels:
severity: warning
annotations:
summary: "API Gateway total error rate {{ $value | humanizePercentage }} (SLA: <10%)"
- name: api-gateway.resources
rules:
# Gateway pod restart
- alert: APIGatewayRestarted
expr: increase(kube_pod_container_status_restarts_total{namespace="api"}[15m]) > 0
for: 0m
labels:
severity: warning
annotations:
summary: "API Gateway pod {{ $labels.pod }} restarted"
# LLM predictor restart
- alert: LLMPredictorRestarted
expr: increase(kube_pod_container_status_restarts_total{namespace="llm-serving"}[15m]) > 0
for: 0m
labels:
severity: warning
annotations:
summary: "LLM predictor {{ $labels.pod }} restarted"
# Gateway high memory (>80% of limit)
- alert: APIGatewayHighMemory
expr: |
sum(container_memory_working_set_bytes{namespace="api",container="gateway"}) by (pod)
/ sum(kube_pod_container_resource_limits{namespace="api",container="gateway",resource="memory"}) by (pod)
> 0.8
for: 10m
labels:
severity: warning
annotations:
summary: "Gateway pod {{ $labels.pod }} memory at {{ $value | humanizePercentage }} of limit"
@@ -1,177 +0,0 @@
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: cluster-alerts
namespace: monitoring
labels:
release: prometheus
spec:
groups:
- name: cluster.availability
rules:
# Node down
- alert: NodeNotReady
expr: kube_node_status_condition{condition="Ready",status="true"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: "Node {{ $labels.node }} is NotReady"
# Pod stuck pending (scheduling failure)
- alert: PodStuckPending
expr: sum(kube_pod_status_phase{phase="Pending"}) > 0
for: 10m
labels:
severity: warning
annotations:
summary: "{{ $value }} pod(s) stuck in Pending state for >10m"
# CrashLoopBackOff
- alert: PodCrashLooping
expr: sum(kube_pod_container_status_waiting_reason{reason="CrashLoopBackOff"}) by (namespace, pod) > 0
for: 5m
labels:
severity: critical
annotations:
summary: "{{ $labels.namespace }}/{{ $labels.pod }} in CrashLoopBackOff"
# OOMKilled spike
- alert: OOMKilledSpike
expr: sum(increase(kube_pod_container_status_last_terminated_reason{reason="OOMKilled"}[1h])) > 3
for: 0m
labels:
severity: warning
annotations:
summary: "{{ $value }} OOMKilled events in last hour"
# Deployment replicas unavailable
- alert: DeploymentReplicasUnavailable
expr: kube_deployment_status_replicas_unavailable > 0
for: 10m
labels:
severity: warning
annotations:
summary: "{{ $labels.namespace }}/{{ $labels.deployment }} has {{ $value }} unavailable replicas"
- name: cluster.jobs
rules:
# Job failed
- alert: JobFailed
expr: kube_job_status_failed > 0
for: 5m
labels:
severity: warning
annotations:
summary: "Job {{ $labels.namespace }}/{{ $labels.job_name }} failed"
# Job stuck running >2h
- alert: JobStuckRunning
expr: |
kube_job_status_active == 1
and on(job_name,namespace)
(time() - kube_job_status_start_time) > 7200
for: 0m
labels:
severity: warning
annotations:
summary: "Job {{ $labels.namespace }}/{{ $labels.job_name }} running >2h"
# CronJob missed schedule
- alert: CronJobMissedSchedule
expr: |
(time() - kube_cronjob_status_last_schedule_time) > 2 * (kube_cronjob_spec_next_schedule_time - kube_cronjob_status_last_schedule_time)
for: 10m
labels:
severity: warning
annotations:
summary: "CronJob {{ $labels.namespace }}/{{ $labels.cronjob }} missed schedule"
- name: cluster.resources
rules:
# Node CPU >90% sustained
- alert: NodeHighCPU
expr: (1 - avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) by (instance)) * 100 > 90
for: 15m
labels:
severity: warning
annotations:
summary: "Node {{ $labels.instance }} CPU at {{ $value | printf \"%.0f\" }}%"
# Node memory >90% sustained
- alert: NodeHighMemory
expr: (1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100 > 90
for: 15m
labels:
severity: warning
annotations:
summary: "Node {{ $labels.instance }} memory at {{ $value | printf \"%.0f\" }}%"
# Node disk >85%
- alert: NodeDiskFull
expr: (1 - node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100 > 85
for: 5m
labels:
severity: critical
annotations:
summary: "Node {{ $labels.instance }} disk at {{ $value | printf \"%.0f\" }}%"
# Container restart storm (>5 restarts in 15m)
- alert: ContainerRestartStorm
expr: sum(increase(kube_pod_container_status_restarts_total[15m])) by (namespace, pod) > 5
for: 0m
labels:
severity: warning
annotations:
summary: "{{ $labels.namespace }}/{{ $labels.pod }} restarted {{ $value | printf \"%.0f\" }} times in 15m"
- name: cluster.storage
rules:
# Longhorn drive offline
- alert: LonghornDriveOffline
expr: longhorn_disk_health != 1
for: 5m
labels:
severity: critical
annotations:
summary: "Longhorn disk {{ $labels.node }} unhealthy"
- name: cluster.dns
rules:
# CoreDNS errors spike
- alert: CoreDNSErrorSpike
expr: sum(rate(coredns_dns_responses_total{rcode=~"SERVFAIL"}[5m])) > 0.5
for: 5m
labels:
severity: warning
annotations:
summary: "CoreDNS SERVFAIL rate {{ $value | printf \"%.2f\" }}/s"
- name: cluster.probes
rules:
# Any blackbox probe down
- alert: ServiceProbeDown
expr: probe_success == 0
for: 3m
labels:
severity: critical
annotations:
summary: "Probe failed: {{ $labels.instance }}"
# Probe latency >2s
- alert: ServiceProbeSlow
expr: probe_duration_seconds > 2
for: 5m
labels:
severity: warning
annotations:
summary: "Probe slow ({{ $value | printf \"%.1f\" }}s): {{ $labels.instance }}"
# Certificate expiry <14 days
- alert: CertificateExpiringSoon
expr: (certmanager_certificate_expiration_timestamp_seconds - time()) / 86400 < 14
for: 0m
labels:
severity: warning
annotations:
summary: "Certificate {{ $labels.name }} expires in {{ $value | printf \"%.0f\" }} days"
@@ -61,10 +61,6 @@ serviceMonitor:
url: https://argocd.riotpiao.com/healthz
- name: longhorn
url: https://longhorn.riotpiao.com/
- name: api-gateway
url: https://api.riotpiao.com/healthz
- name: api-gateway-models
url: https://api.riotpiao.com/v1/models
prometheusRule:
enabled: true
@@ -1,36 +0,0 @@
apiVersion: v1
data:
api-gateway.json: '{"title":"API Gateway","uid":"api-gateway","schemaVersion":39,"timezone":"browser","time":{"from":"now-6h","to":"now"},"refresh":"30s","tags":["api","gateway","llm"],"panels":[{"id":1,"title":"Gateway
Health","type":"row","collapsed":false,"gridPos":{"h":1,"w":24,"x":0,"y":0},"panels":[{"id":2,"title":"Gateway
Pods Ready","type":"stat","gridPos":{"h":4,"w":4,"x":0,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"red"},{"value":3,"color":"green"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"sum(kube_pod_status_ready{namespace=\"api\",condition=\"true\"})"}]},{"id":3,"title":"Probe:
healthz","type":"stat","gridPos":{"h":4,"w":4,"x":4,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","mappings":[{"type":"value","options":{"0":{"text":"DOWN","color":"red"},"1":{"text":"UP","color":"green"}}}]}},"targets":[{"expr":"probe_success{instance=~\".*api.riotpiao.com/healthz\"}"}]},{"id":4,"title":"Probe
Latency","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"s"}},"targets":[{"expr":"probe_duration_seconds{instance=~\".*api.riotpiao.com.*\"}","legendFormat":"{{instance}}"}]}]},{"id":10,"title":"Ingress
Traffic (nginx)","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":1},"panels":[{"id":11,"title":"Request
Rate by Status","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(nginx_ingress_controller_requests{ingress=\"api\"}[5m]))
by (status)","legendFormat":"{{status}}"}]},{"id":12,"title":"Error Rate %","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"percent"}},"targets":[{"expr":"sum(rate(nginx_ingress_controller_requests{ingress=\"api\",status=~\"5..\"}[5m]))
/ sum(rate(nginx_ingress_controller_requests{ingress=\"api\"}[5m])) * 100","legendFormat":"5xx"},{"expr":"sum(rate(nginx_ingress_controller_requests{ingress=\"api\",status=~\"4..\"}[5m]))
/ sum(rate(nginx_ingress_controller_requests{ingress=\"api\"}[5m])) * 100","legendFormat":"4xx"}]},{"id":13,"title":"Latency
p50/p95/p99","type":"timeseries","gridPos":{"h":8,"w":8,"x":16,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"s"}},"targets":[{"expr":"histogram_quantile(0.50,
sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress=\"api\"}[5m]))
by (le))","legendFormat":"p50"},{"expr":"histogram_quantile(0.95, sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress=\"api\"}[5m]))
by (le))","legendFormat":"p95"},{"expr":"histogram_quantile(0.99, sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress=\"api\"}[5m]))
by (le))","legendFormat":"p99"}]}]},{"id":20,"title":"LLM Serving","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":2},"panels":[{"id":21,"title":"LLM
Pods Ready","type":"stat","gridPos":{"h":4,"w":4,"x":0,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(kube_pod_status_ready{namespace=\"llm-serving\",condition=\"true\"})"}]},{"id":22,"title":"CPU
by Predictor","type":"timeseries","gridPos":{"h":8,"w":8,"x":4,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(container_cpu_usage_seconds_total{namespace=\"llm-serving\"}[5m]))
by (pod)","legendFormat":"{{pod}}"}]},{"id":23,"title":"Memory by Predictor","type":"timeseries","gridPos":{"h":8,"w":8,"x":12,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"bytes"}},"targets":[{"expr":"sum(container_memory_working_set_bytes{namespace=\"llm-serving\"})
by (pod)","legendFormat":"{{pod}}"}]},{"id":24,"title":"Predictor Restarts","type":"timeseries","gridPos":{"h":8,"w":4,"x":20,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(kube_pod_container_status_restarts_total{namespace=\"llm-serving\"}[15m]))
by (pod)","legendFormat":"{{pod}}"}]}]},{"id":30,"title":"Gateway Resources","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":3},"panels":[{"id":31,"title":"CPU
by Gateway Pod","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":4},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(container_cpu_usage_seconds_total{namespace=\"api\"}[5m]))
by (pod)","legendFormat":"{{pod}}"}]},{"id":32,"title":"Memory by Gateway Pod","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":4},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"bytes"}},"targets":[{"expr":"sum(container_memory_working_set_bytes{namespace=\"api\"})
by (pod)","legendFormat":"{{pod}}"}]},{"id":33,"title":"Gateway Restarts","type":"timeseries","gridPos":{"h":8,"w":8,"x":16,"y":4},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(kube_pod_container_status_restarts_total{namespace=\"api\"}[15m]))
by (pod)","legendFormat":"{{pod}}"}]}]},{"id":40,"title":"Logs","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":4},"panels":[{"id":41,"title":"Gateway
Logs","type":"logs","gridPos":{"h":10,"w":24,"x":0,"y":5},"datasource":{"type":"loki","uid":"loki"},"targets":[{"expr":"{namespace=\"api\",container=\"gateway\"}"}]},{"id":42,"title":"LLM
Serving Logs","type":"logs","gridPos":{"h":10,"w":24,"x":0,"y":15},"datasource":{"type":"loki","uid":"loki"},"targets":[{"expr":"{namespace=\"llm-serving\"}"}]}]}]}'
kind: ConfigMap
metadata:
annotations:
grafana_folder: API
labels:
grafana_dashboard: '1'
name: api-gateway-dashboard
namespace: logging
@@ -1,62 +0,0 @@
apiVersion: v1
data:
cluster-infrastructure.json: '{"title":"Cluster Infrastructure","uid":"cluster-infra","schemaVersion":39,"timezone":"browser","time":{"from":"now-6h","to":"now"},"refresh":"30s","tags":["infrastructure","k8s"],"panels":[{"id":1,"title":"Cluster
Health","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":0},"panels":[{"id":2,"title":"Nodes
Ready","type":"stat","gridPos":{"h":4,"w":4,"x":0,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"red"},{"value":3,"color":"green"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"count(kube_node_status_condition{condition=\"Ready\",status=\"true\"}
== 1)"}]},{"id":3,"title":"Pods Pending","type":"stat","gridPos":{"h":4,"w":4,"x":4,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"green"},{"value":1,"color":"red"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"sum(kube_pod_status_phase{phase=\"Pending\"})
OR on() vector(0)"}]},{"id":4,"title":"CrashLoopBackOff","type":"stat","gridPos":{"h":4,"w":4,"x":8,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"green"},{"value":1,"color":"red"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"sum(kube_pod_container_status_waiting_reason{reason=\"CrashLoopBackOff\"})
OR on() vector(0)"}]},{"id":5,"title":"OOMKilled (1h)","type":"stat","gridPos":{"h":4,"w":4,"x":12,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"green"},{"value":1,"color":"red"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"sum(increase(kube_pod_container_status_last_terminated_reason{reason=\"OOMKilled\"}[1h]))
OR on() vector(0)"}]},{"id":6,"title":"Deploys Unavailable","type":"stat","gridPos":{"h":4,"w":4,"x":16,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"green"},{"value":1,"color":"red"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"count(kube_deployment_status_replicas_unavailable
> 0) OR on() vector(0)"}]},{"id":7,"title":"Services Down","type":"stat","gridPos":{"h":4,"w":4,"x":20,"y":1},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"green"},{"value":1,"color":"red"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"count(probe_success
== 0) OR on() vector(0)"}]}]},{"id":10,"title":"Jobs & CronJobs","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":1},"panels":[{"id":11,"title":"Failed
Jobs","type":"stat","gridPos":{"h":4,"w":4,"x":0,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","thresholds":{"mode":"absolute","steps":[{"value":null,"color":"green"},{"value":1,"color":"red"}]},"color":{"mode":"thresholds"}}},"targets":[{"expr":"count(kube_job_status_failed
> 0) OR on() vector(0)"}]},{"id":12,"title":"Failed Jobs Detail","type":"table","gridPos":{"h":8,"w":10,"x":4,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"kube_job_status_failed
> 0","format":"table","instant":true}]},{"id":13,"title":"Stuck Jobs (>1h)","type":"table","gridPos":{"h":8,"w":10,"x":14,"y":2},"datasource":{"type":"prometheus","uid":"prometheus"},"targets":[{"expr":"kube_job_status_active
== 1 and on(job_name,namespace) (time() - kube_job_status_start_time) > 3600","format":"table","instant":true}]},{"id":14,"title":"CronJob
Last Success","type":"timeseries","gridPos":{"h":8,"w":12,"x":0,"y":10},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"dateTimeFromNow"}},"targets":[{"expr":"kube_cronjob_status_last_successful_time{namespace=~\"cicd|kube-system|paperless\"}","legendFormat":"{{namespace}}/{{cronjob}}"}]},{"id":15,"title":"Container
Restart Storm (top 10)","type":"timeseries","gridPos":{"h":8,"w":12,"x":12,"y":10},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"topk(10,
sum(rate(kube_pod_container_status_restarts_total[15m])) by (namespace, pod))","legendFormat":"{{namespace}}/{{pod}}"}]}]},{"id":20,"title":"Node
Resources","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":2},"panels":[{"id":21,"title":"CPU
% by Node","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"percent"}},"targets":[{"expr":"(1
- avg(rate(node_cpu_seconds_total{mode=\"idle\"}[5m])) by (instance)) * 100","legendFormat":"{{instance}}"}]},{"id":22,"title":"Memory
% by Node","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"percent"}},"targets":[{"expr":"(1
- node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100","legendFormat":"{{instance}}"}]},{"id":23,"title":"Disk
% by Node","type":"timeseries","gridPos":{"h":8,"w":8,"x":16,"y":3},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"percent"}},"targets":[{"expr":"(1
- node_filesystem_avail_bytes{mountpoint=\"/\"} / node_filesystem_size_bytes{mountpoint=\"/\"})
* 100","legendFormat":"{{instance}}"}]},{"id":24,"title":"Load Average","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":11},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"node_load1","legendFormat":"1m
{{instance}}"},{"expr":"node_load5","legendFormat":"5m {{instance}}"}]},{"id":25,"title":"Network
Errors & Drops","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":11},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"rate(node_network_receive_errs_total[5m])","legendFormat":"rx-err
{{instance}}"},{"expr":"rate(node_network_transmit_errs_total[5m])","legendFormat":"tx-err
{{instance}}"},{"expr":"rate(node_network_receive_drop_total[5m])","legendFormat":"rx-drop
{{instance}}"}]}]},{"id":30,"title":"Control Plane","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":3},"panels":[{"id":31,"title":"API
Server Up","type":"stat","gridPos":{"h":4,"w":4,"x":0,"y":4},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short","mappings":[{"type":"value","options":{"0":{"text":"DOWN","color":"red"},"1":{"text":"UP","color":"green"}}}]}},"targets":[{"expr":"min(up{job=\"apiserver\"})"}]},{"id":32,"title":"API
Server Request Rate","type":"timeseries","gridPos":{"h":8,"w":10,"x":4,"y":4},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(apiserver_request_total[5m]))
by (verb, code)","legendFormat":"{{verb}} {{code}}"}]},{"id":33,"title":"API Server
Error Rate %","type":"timeseries","gridPos":{"h":8,"w":10,"x":14,"y":4},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"percent"}},"targets":[{"expr":"sum(rate(apiserver_request_total{code=~\"5..\"}[5m]))
/ sum(rate(apiserver_request_total[5m])) * 100","legendFormat":"5xx %"}]},{"id":34,"title":"API
Server Latency","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":12},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"s"}},"targets":[{"expr":"histogram_quantile(0.95,
sum(rate(apiserver_request_duration_seconds_bucket[5m])) by (le))","legendFormat":"p95"},{"expr":"histogram_quantile(0.99,
sum(rate(apiserver_request_duration_seconds_bucket[5m])) by (le))","legendFormat":"p99"}]},{"id":35,"title":"etcd
Request Duration","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":12},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"s"}},"targets":[{"expr":"histogram_quantile(0.99,
sum(rate(etcd_request_duration_seconds_bucket[5m])) by (le))","legendFormat":"p99"}]}]},{"id":40,"title":"Storage","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":4},"panels":[{"id":41,"title":"Longhorn
Disk Capacity","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":5},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"bytes"}},"targets":[{"expr":"longhorn_disk_capacity_bytes","legendFormat":"capacity
{{node}}"},{"expr":"longhorn_disk_reservation_bytes","legendFormat":"reserved
{{node}}"}]},{"id":42,"title":"PVC Phase","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":5},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"kube_persistentvolumeclaim_status_phase","legendFormat":"{{namespace}}/{{persistentvolumeclaim}}
{{phase}}"}]}]},{"id":50,"title":"DNS & Networking","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":5},"panels":[{"id":51,"title":"CoreDNS
Cache Hit Rate","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":6},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"percentunit"}},"targets":[{"expr":"rate(coredns_cache_hits_total[5m])
/ (rate(coredns_cache_hits_total[5m]) + rate(coredns_cache_misses_total[5m]))","legendFormat":"{{server}}"}]},{"id":52,"title":"CoreDNS
Errors","type":"timeseries","gridPos":{"h":8,"w":8,"x":8,"y":6},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum(rate(coredns_dns_responses_total{rcode=~\"SERVFAIL|NXDOMAIN\"}[5m]))
by (rcode)","legendFormat":"{{rcode}}"}]}]},{"id":60,"title":"Logs","type":"row","collapsed":true,"gridPos":{"h":1,"w":24,"x":0,"y":6},"panels":[{"id":61,"title":"Error
Rate by Namespace","type":"timeseries","gridPos":{"h":8,"w":8,"x":0,"y":7},"datasource":{"type":"prometheus","uid":"prometheus"},"fieldConfig":{"defaults":{"unit":"short"}},"targets":[{"expr":"sum
by (namespace) (count_over_time({namespace=~\"kube-system|cert-manager|ingress-nginx|longhorn-system\"}
|= \"error\" [5m]))","legendFormat":"{{namespace}}"}]},{"id":62,"title":"Control
Plane Logs","type":"logs","gridPos":{"h":10,"w":24,"x":0,"y":15},"datasource":{"type":"loki","uid":"loki"},"targets":[{"expr":"{namespace=\"kube-system\"}"}]},{"id":63,"title":"Cluster
Addon Logs","type":"logs","gridPos":{"h":10,"w":24,"x":0,"y":25},"datasource":{"type":"loki","uid":"loki"},"targets":[{"expr":"{namespace=~\"cert-manager|ingress-nginx|longhorn-system\"}"}]}]}]}'
kind: ConfigMap
metadata:
annotations:
grafana_folder: Infrastructure
labels:
grafana_dashboard: '1'
name: cluster-infrastructure-dashboard
namespace: logging
@@ -0,0 +1,55 @@
# k8s/monitoring/dashboards/control-plane-logs.yaml
# Surfaces controller/control-plane logs that are already in Loki today
# (Promtail scrapes every namespace with no filter) — this dashboard is the
# "make it visible" piece, not new log collection.
apiVersion: v1
kind: ConfigMap
metadata:
name: control-plane-logs-dashboard
namespace: logging
labels:
grafana_dashboard: "1"
data:
control-plane-logs.json: |
{
"title": "Cluster Control Plane & Controllers (Logs)",
"uid": "control-plane-logs",
"schemaVersion": 39,
"timezone": "browser",
"time": { "from": "now-1h", "to": "now" },
"refresh": "30s",
"panels": [
{
"id": 1,
"title": "Error rate by namespace",
"type": "timeseries",
"gridPos": { "h": 6, "w": 24, "x": 0, "y": 0 },
"datasource": { "type": "loki", "uid": "loki" },
"targets": [
{
"expr": "sum by (namespace) (count_over_time({namespace=~\"kube-system|cert-manager|ingress-nginx|longhorn-system\"} |= \"error\" [5m]))"
}
]
},
{
"id": 2,
"title": "Control plane (kube-apiserver, controller-manager, scheduler)",
"type": "logs",
"gridPos": { "h": 10, "w": 24, "x": 0, "y": 6 },
"datasource": { "type": "loki", "uid": "loki" },
"targets": [
{ "expr": "{namespace=\"kube-system\"}" }
]
},
{
"id": 3,
"title": "Cluster add-ons (cert-manager, ingress-nginx, longhorn)",
"type": "logs",
"gridPos": { "h": 10, "w": 24, "x": 0, "y": 16 },
"datasource": { "type": "loki", "uid": "loki" },
"targets": [
{ "expr": "{namespace=~\"cert-manager|ingress-nginx|longhorn-system\"}" }
]
}
]
}
@@ -1,280 +0,0 @@
#!/usr/bin/env python3
"""Generate consolidated Grafana dashboards as k8s ConfigMap YAML files."""
import json
import os
DASHBOARD_DIR = os.path.expanduser("~/workplace/homelab/k8s/infra/monitoring/dashboards")
DS_PROM = {"type": "prometheus", "uid": "prometheus"}
DS_LOKI = {"type": "loki", "uid": "loki"}
def stat_panel(id, title, expr, x, y, w=4, h=4, unit="short", mappings=None, thresholds=None):
p = {
"id": id, "title": title, "type": "stat",
"gridPos": {"h": h, "w": w, "x": x, "y": y},
"datasource": DS_PROM,
"fieldConfig": {"defaults": {"unit": unit}},
"targets": [{"expr": expr}],
}
if mappings:
p["fieldConfig"]["defaults"]["mappings"] = mappings
if thresholds:
p["fieldConfig"]["defaults"]["thresholds"] = thresholds
p["fieldConfig"]["defaults"]["color"] = {"mode": "thresholds"}
return p
def ts_panel(id, title, exprs, x, y, w=8, h=8, unit="short"):
targets = []
for e in exprs:
if isinstance(e, tuple):
targets.append({"expr": e[0], "legendFormat": e[1]})
else:
targets.append({"expr": e, "legendFormat": "{{pod}}"})
return {
"id": id, "title": title, "type": "timeseries",
"gridPos": {"h": h, "w": w, "x": x, "y": y},
"datasource": DS_PROM,
"fieldConfig": {"defaults": {"unit": unit}},
"targets": targets,
}
def table_panel(id, title, expr, x, y, w=12, h=8):
return {
"id": id, "title": title, "type": "table",
"gridPos": {"h": h, "w": w, "x": x, "y": y},
"datasource": DS_PROM,
"targets": [{"expr": expr, "format": "table", "instant": True}],
}
def log_panel(id, title, query, x, y, w=24, h=10):
return {
"id": id, "title": title, "type": "logs",
"gridPos": {"h": h, "w": w, "x": x, "y": y},
"datasource": DS_LOKI,
"targets": [{"expr": query}],
}
def row(id, title, y, panels, collapsed=True):
return {
"id": id, "title": title, "type": "row",
"collapsed": collapsed, "gridPos": {"h": 1, "w": 24, "x": 0, "y": y},
"panels": panels,
}
def write_dashboard(filename, dashboard, folder):
cm = {
"apiVersion": "v1",
"kind": "ConfigMap",
"metadata": {
"name": filename.replace(".yaml", "-dashboard"),
"namespace": "logging",
"labels": {"grafana_dashboard": "1"},
"annotations": {"grafana_folder": folder},
},
"data": {
filename.replace(".yaml", ".json"): json.dumps(dashboard, separators=(",", ":"))
},
}
import yaml
path = os.path.join(DASHBOARD_DIR, filename)
with open(path, "w") as f:
yaml.dump(cm, f, default_flow_style=False, allow_unicode=True)
print(f" wrote {path}")
# ============================================================================
# Dashboard 1: Cluster Infrastructure
# ============================================================================
def build_cluster_infrastructure():
zero_thresholds = {"mode": "absolute", "steps": [
{"value": None, "color": "green"}, {"value": 1, "color": "red"}
]}
panels = [
row(1, "Cluster Health", 0, [
stat_panel(2, "Nodes Ready", 'count(kube_node_status_condition{condition="Ready",status="true"} == 1)', 0, 1, thresholds={"mode":"absolute","steps":[{"value":None,"color":"red"},{"value":3,"color":"green"}]}),
stat_panel(3, "Pods Pending", 'sum(kube_pod_status_phase{phase="Pending"}) OR on() vector(0)', 4, 1, thresholds=zero_thresholds),
stat_panel(4, "CrashLoopBackOff", 'sum(kube_pod_container_status_waiting_reason{reason="CrashLoopBackOff"}) OR on() vector(0)', 8, 1, thresholds=zero_thresholds),
stat_panel(5, "OOMKilled (1h)", 'sum(increase(kube_pod_container_status_last_terminated_reason{reason="OOMKilled"}[1h])) OR on() vector(0)', 12, 1, thresholds=zero_thresholds),
stat_panel(6, "Deploys Unavailable", 'count(kube_deployment_status_replicas_unavailable > 0) OR on() vector(0)', 16, 1, thresholds=zero_thresholds),
stat_panel(7, "Services Down", 'count(probe_success == 0) OR on() vector(0)', 20, 1, thresholds=zero_thresholds),
]),
row(10, "Jobs & CronJobs", 1, [
stat_panel(11, "Failed Jobs", 'count(kube_job_status_failed > 0) OR on() vector(0)', 0, 2, thresholds=zero_thresholds),
table_panel(12, "Failed Jobs Detail", 'kube_job_status_failed > 0', 4, 2, w=10),
table_panel(13, "Stuck Jobs (>1h)", 'kube_job_status_active == 1 and on(job_name,namespace) (time() - kube_job_status_start_time) > 3600', 14, 2, w=10),
ts_panel(14, "CronJob Last Success", [
('kube_cronjob_status_last_successful_time{namespace=~"cicd|kube-system|paperless"}', "{{namespace}}/{{cronjob}}")
], 0, 10, w=12, unit="dateTimeFromNow"),
ts_panel(15, "Container Restart Storm (top 10)", [
('topk(10, sum(rate(kube_pod_container_status_restarts_total[15m])) by (namespace, pod))', "{{namespace}}/{{pod}}")
], 12, 10, w=12),
]),
row(20, "Node Resources", 2, [
ts_panel(21, "CPU % by Node", [
('(1 - avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) by (instance)) * 100', "{{instance}}")
], 0, 3, unit="percent"),
ts_panel(22, "Memory % by Node", [
('(1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100', "{{instance}}")
], 8, 3, unit="percent"),
ts_panel(23, "Disk % by Node", [
('(1 - node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100', "{{instance}}")
], 16, 3, unit="percent"),
ts_panel(24, "Load Average", [
("node_load1", "1m {{instance}}"),
("node_load5", "5m {{instance}}"),
], 0, 11),
ts_panel(25, "Network Errors & Drops", [
("rate(node_network_receive_errs_total[5m])", "rx-err {{instance}}"),
("rate(node_network_transmit_errs_total[5m])", "tx-err {{instance}}"),
("rate(node_network_receive_drop_total[5m])", "rx-drop {{instance}}"),
], 8, 11),
]),
row(30, "Control Plane", 3, [
stat_panel(31, "API Server Up", 'min(up{job="apiserver"})', 0, 4, mappings=[
{"type":"value","options":{"0":{"text":"DOWN","color":"red"},"1":{"text":"UP","color":"green"}}}
]),
ts_panel(32, "API Server Request Rate", [
('sum(rate(apiserver_request_total[5m])) by (verb, code)', "{{verb}} {{code}}")
], 4, 4, w=10),
ts_panel(33, "API Server Error Rate %", [
('sum(rate(apiserver_request_total{code=~"5.."}[5m])) / sum(rate(apiserver_request_total[5m])) * 100', "5xx %")
], 14, 4, w=10, unit="percent"),
ts_panel(34, "API Server Latency", [
('histogram_quantile(0.95, sum(rate(apiserver_request_duration_seconds_bucket[5m])) by (le))', "p95"),
('histogram_quantile(0.99, sum(rate(apiserver_request_duration_seconds_bucket[5m])) by (le))', "p99"),
], 0, 12, unit="s"),
ts_panel(35, "etcd Request Duration", [
('histogram_quantile(0.99, sum(rate(etcd_request_duration_seconds_bucket[5m])) by (le))', "p99"),
], 8, 12, unit="s"),
]),
row(40, "Storage", 4, [
ts_panel(41, "Longhorn Disk Capacity", [
("longhorn_disk_capacity_bytes", "capacity {{node}}"),
("longhorn_disk_reservation_bytes", "reserved {{node}}"),
], 0, 5, unit="bytes"),
ts_panel(42, "PVC Phase", [
('kube_persistentvolumeclaim_status_phase', "{{namespace}}/{{persistentvolumeclaim}} {{phase}}")
], 8, 5),
]),
row(50, "DNS & Networking", 5, [
ts_panel(51, "CoreDNS Cache Hit Rate", [
('rate(coredns_cache_hits_total[5m]) / (rate(coredns_cache_hits_total[5m]) + rate(coredns_cache_misses_total[5m]))', "{{server}}")
], 0, 6, unit="percentunit"),
ts_panel(52, "CoreDNS Errors", [
('sum(rate(coredns_dns_responses_total{rcode=~"SERVFAIL|NXDOMAIN"}[5m])) by (rcode)', "{{rcode}}")
], 8, 6),
]),
row(60, "Logs", 6, [
ts_panel(61, "Error Rate by Namespace", [
('sum by (namespace) (count_over_time({namespace=~"kube-system|cert-manager|ingress-nginx|longhorn-system"} |= "error" [5m]))', "{{namespace}}")
], 0, 7),
log_panel(62, "Control Plane Logs", '{namespace="kube-system"}', 0, 15),
log_panel(63, "Cluster Addon Logs", '{namespace=~"cert-manager|ingress-nginx|longhorn-system"}', 0, 25),
]),
]
return {
"title": "Cluster Infrastructure",
"uid": "cluster-infra",
"schemaVersion": 39,
"timezone": "browser",
"time": {"from": "now-6h", "to": "now"},
"refresh": "30s",
"tags": ["infrastructure", "k8s"],
"panels": panels,
}
# ============================================================================
# Dashboard 3: API Gateway
# ============================================================================
def build_api_gateway():
panels = [
row(1, "Gateway Health", 0, [
stat_panel(2, "Gateway Pods Ready", 'sum(kube_pod_status_ready{namespace="api",condition="true"})', 0, 1, thresholds={"mode":"absolute","steps":[{"value":None,"color":"red"},{"value":3,"color":"green"}]}),
stat_panel(3, "Probe: healthz", 'probe_success{instance=~".*api.riotpiao.com/healthz"}', 4, 1, mappings=[
{"type":"value","options":{"0":{"text":"DOWN","color":"red"},"1":{"text":"UP","color":"green"}}}
]),
ts_panel(4, "Probe Latency", [
('probe_duration_seconds{instance=~".*api.riotpiao.com.*"}', "{{instance}}")
], 8, 1, unit="s"),
], collapsed=False),
row(10, "Ingress Traffic (nginx)", 1, [
ts_panel(11, "Request Rate by Status", [
('sum(rate(nginx_ingress_controller_requests{ingress="api"}[5m])) by (status)', "{{status}}")
], 0, 2),
ts_panel(12, "Error Rate %", [
('sum(rate(nginx_ingress_controller_requests{ingress="api",status=~"5.."}[5m])) / sum(rate(nginx_ingress_controller_requests{ingress="api"}[5m])) * 100', "5xx"),
('sum(rate(nginx_ingress_controller_requests{ingress="api",status=~"4.."}[5m])) / sum(rate(nginx_ingress_controller_requests{ingress="api"}[5m])) * 100', "4xx"),
], 8, 2, unit="percent"),
ts_panel(13, "Latency p50/p95/p99", [
('histogram_quantile(0.50, sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress="api"}[5m])) by (le))', "p50"),
('histogram_quantile(0.95, sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress="api"}[5m])) by (le))', "p95"),
('histogram_quantile(0.99, sum(rate(nginx_ingress_controller_request_duration_seconds_bucket{ingress="api"}[5m])) by (le))', "p99"),
], 16, 2, unit="s"),
]),
row(20, "LLM Serving", 2, [
stat_panel(21, "LLM Pods Ready", 'sum(kube_pod_status_ready{namespace="llm-serving",condition="true"})', 0, 3),
ts_panel(22, "CPU by Predictor", [
('sum(rate(container_cpu_usage_seconds_total{namespace="llm-serving"}[5m])) by (pod)', "{{pod}}")
], 4, 3),
ts_panel(23, "Memory by Predictor", [
('sum(container_memory_working_set_bytes{namespace="llm-serving"}) by (pod)', "{{pod}}")
], 12, 3, unit="bytes"),
ts_panel(24, "Predictor Restarts", [
('sum(rate(kube_pod_container_status_restarts_total{namespace="llm-serving"}[15m])) by (pod)', "{{pod}}")
], 20, 3, w=4),
]),
row(30, "Gateway Resources", 3, [
ts_panel(31, "CPU by Gateway Pod", [
('sum(rate(container_cpu_usage_seconds_total{namespace="api"}[5m])) by (pod)', "{{pod}}")
], 0, 4),
ts_panel(32, "Memory by Gateway Pod", [
('sum(container_memory_working_set_bytes{namespace="api"}) by (pod)', "{{pod}}")
], 8, 4, unit="bytes"),
ts_panel(33, "Gateway Restarts", [
('sum(rate(kube_pod_container_status_restarts_total{namespace="api"}[15m])) by (pod)', "{{pod}}")
], 16, 4),
]),
row(40, "Logs", 4, [
log_panel(41, "Gateway Logs", '{namespace="api",container="gateway"}', 0, 5),
log_panel(42, "LLM Serving Logs", '{namespace="llm-serving"}', 0, 15),
]),
]
return {
"title": "API Gateway",
"uid": "api-gateway",
"schemaVersion": 39,
"timezone": "browser",
"time": {"from": "now-6h", "to": "now"},
"refresh": "30s",
"tags": ["api", "gateway", "llm"],
"panels": panels,
}
# ============================================================================
# Generate
# ============================================================================
print("Generating dashboards...")
# Dashboard 1
write_dashboard("cluster-infrastructure.yaml", build_cluster_infrastructure(), "Infrastructure")
# Dashboard 3
write_dashboard("api-gateway.yaml", build_api_gateway(), "API")
print("Done.")
@@ -0,0 +1,121 @@
# k8s/monitoring/dashboards/hardware-overview.yaml
# Trimmed operator at-a-glance view across all nodes — node-exporter already
# powers the deep-dive "Node Exporter Full" (#1860, see grafana-values.yaml),
# this is the quick health-check version, not a replacement for it.
apiVersion: v1
kind: ConfigMap
metadata:
name: hardware-overview-dashboard
namespace: logging
labels:
grafana_dashboard: "1"
data:
hardware-overview.json: |
{
"title": "Hardware Statistics (Operator Overview)",
"uid": "hardware-overview",
"schemaVersion": 39,
"timezone": "browser",
"time": { "from": "now-6h", "to": "now" },
"refresh": "30s",
"panels": [
{
"id": 1,
"title": "Nodes up / down",
"type": "stat",
"gridPos": { "h": 5, "w": 24, "x": 0, "y": 0 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": {
"defaults": {
"mappings": [
{ "type": "value", "options": { "0": { "text": "DOWN", "color": "red" } } },
{ "type": "value", "options": { "1": { "text": "UP", "color": "green" } } }
]
}
},
"targets": [
{ "expr": "up{job=~\".*node-exporter.*\"}", "legendFormat": "{{instance}}" }
]
},
{
"id": 2,
"title": "CPU usage % by node",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 5 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": { "defaults": { "unit": "percent", "max": 100, "min": 0 } },
"targets": [
{
"expr": "(1 - avg(rate(node_cpu_seconds_total{mode=\"idle\"}[5m])) by (instance)) * 100",
"legendFormat": "{{instance}}"
}
]
},
{
"id": 3,
"title": "Memory usage % by node",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 5 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": { "defaults": { "unit": "percent", "max": 100, "min": 0 } },
"targets": [
{
"expr": "(1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100",
"legendFormat": "{{instance}}"
}
]
},
{
"id": 4,
"title": "Root filesystem usage % by node",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 13 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": { "defaults": { "unit": "percent", "max": 100, "min": 0 } },
"targets": [
{
"expr": "(1 - node_filesystem_avail_bytes{mountpoint=\"/\"} / node_filesystem_size_bytes{mountpoint=\"/\"}) * 100",
"legendFormat": "{{instance}}"
}
]
},
{
"id": 5,
"title": "Root filesystem space remaining",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 13 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"fieldConfig": { "defaults": { "unit": "bytes" } },
"targets": [
{
"expr": "node_filesystem_avail_bytes{mountpoint=\"/\"}",
"legendFormat": "{{instance}}"
}
]
},
{
"id": 6,
"title": "Network errors/drops by node",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 21 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{ "expr": "rate(node_network_receive_errs_total[5m])", "legendFormat": "{{instance}} rx errs" },
{ "expr": "rate(node_network_transmit_errs_total[5m])", "legendFormat": "{{instance}} tx errs" },
{ "expr": "rate(node_network_receive_drop_total[5m])", "legendFormat": "{{instance}} rx drops" },
{ "expr": "rate(node_network_transmit_drop_total[5m])", "legendFormat": "{{instance}} tx drops" }
]
},
{
"id": 7,
"title": "Load average (1m / 5m) by node",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 21 },
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{ "expr": "node_load1", "legendFormat": "{{instance}} load1" },
{ "expr": "node_load5", "legendFormat": "{{instance}} load5" }
]
}
]
}

Some files were not shown because too many files have changed in this diff Show More