chore: initial commit of Go API gateway
CI / Test (push) Canceled after 0s
CI / Vet (push) Canceled after 0s
CI / Build (push) Canceled after 0s
CI / Security (govulncheck) (push) Canceled after 0s

Baseline for the Kong replacement on api.riotpiao.com. Brings the working
tree under version control for the first time: gateway source, the task
board that drives the agent runs, test fixtures, and K8s manifests.

Anchor the gateway ignore rule to the repo root. Unanchored, "gateway"
also matched the cmd/gateway/ source directory, so the program entrypoint
was excluded from every commit.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
This commit is contained in:
Story Crater Bot
2026-08-19 20:54:34 -07:00
co-authored by Claude Opus 5
commit 058f11cf2b
109 changed files with 8992 additions and 0 deletions
+367
View File
@@ -0,0 +1,367 @@
# homelab-frontend — Requirements
The contract for the Go API gateway that replaces Kong OSS on `*.riotpiao.com`.
Companion documents:
- [docs/adr/ADR-0001-retire-kong-for-go-gateway.md](docs/adr/ADR-0001-retire-kong-for-go-gateway.md) — why Kong is being retired
- [docs/MIGRATION-kong.md](docs/MIGRATION-kong.md) — exact inventory of what Kong does today and the cutover order
- [tasks/INDEX.md](tasks/INDEX.md) — the task board
All cluster facts below were verified live against context `admin@homelab-cluster`
on 2026-08-19. Re-verify before relying on any number.
---
## 0. Invariants
These hold for every surface. A change that breaks one of these is a design change,
not an implementation detail.
- **G1** — ingress-nginx owns TLS and the edge. The gateway never terminates TLS.
- **G2** — The gateway holds no Kubernetes credentials. It proxies to services that
do. Cluster-read permissions stay in atlas, out of the public edge process.
- **G3** — Public surfaces use standard protocol shapes. If an OpenAI SDK cannot
call it unmodified, the design is wrong.
- **G4** — Streaming is unbuffered end to end, and a client disconnect cancels the
upstream request rather than orphaning it.
- **G5** — Authentication is Bearer-token, validated against Authentik via JWKS
fetched at runtime. No pinned public keys, no rotation runbook.
- **G6** — Every route's timeout, body cap and concurrency limit is explicit in
configuration. No silent defaults.
- **G7** — All deployment flows through git and Argo. No `kubectl apply`, no
`helm upgrade`, no local `terraform apply`.
---
## 1. Runtime and configuration
### 1.1 Process
Single static Go binary. Reads configuration at startup, serves HTTP, exits cleanly
on SIGTERM after draining in-flight requests.
Must run with **no cluster, no kubeconfig and no credentials** so that behaviour can
be verified in a closed loop before touching live traffic. Upstreams are
configuration; pointing them at local stubs is the entire mechanism. This is a hard
requirement, not a convenience — see §7.
### 1.2 Configuration
Route and upstream configuration is declarative and loaded at startup. It must
express, per upstream: address, path rewrite, connect/read/write timeouts, maximum
body size, and whether the route requires authentication.
Configuration errors fail startup loudly. A gateway that starts with a silently
dropped route is worse than one that refuses to start.
**Configuration lives in git**, mounted as a ConfigMap and synced by Argo. Not a
CRD. A CRD would require the gateway to watch the API server, which needs RBAC and
contradicts G2 — and CRD-driven routing is precisely the indirection being retired
with Kong, where the routing table was split across six `KongPlugin` CRs, seven
Ingresses and a Helm values file.
A CRD earns its keep when someone other than the repo owner must register routes.
That is not true here. If it becomes true, the additive answer is a controller that
renders this same ConfigMap — the gateway stays credential-free either way.
### 1.3 Health
- `GET /healthz` — liveness, no upstream checks, always cheap.
- `GET /readyz` — readiness; may fail while configuration is invalid or JWKS has
never been successfully fetched.
Neither requires authentication.
---
## 2. Proxy core
### 2.1 Reverse proxying
Standard reverse proxy to configured upstreams. Connection reuse across requests.
Hop-by-hop headers stripped correctly. `X-Forwarded-*` set from the nginx-supplied
values, not fabricated.
### 2.2 Streaming
SSE and chunked responses pass through without buffering. Tokens must reach the
client as the upstream emits them, not on completion.
WebSocket upgrade must work — `agent-pod/console` depends on it.
### 2.3 Disconnect propagation
When a client disconnects, the upstream request is cancelled immediately. This is
load-bearing: an orphaned generation holds a vLLM sequence slot, and there are only
eight in the cluster.
### 2.4 Timeouts
Per-route, explicit. Current Kong values, which are deliberate and must be preserved
unless changed knowingly:
| Route class | connect | read | write |
|---|---|---|---|
| chat | 10s | 1h | 1h |
| embeddings, rerank | 10s | 10m | 10m |
The 1-hour read timeout exists because a 32B model on a Volta GPU routinely exceeds
60s. Any shorter application-level cap must be enforced *by the gateway's own
logic*, not by shortening the proxy timeout — otherwise long legitimate generations
truncate mid-stream.
---
## 3. LLM surfaces — `api.riotpiao.com`
Two protocol dialects, permanently. Both translate into one dialect-neutral canonical
request, and both pass through **one shared slot controller** before reaching a
predictor.
| Prefix | Dialect | Primary client |
|---|---|---|
| `/v1/*` | OpenAI-compatible | pi, generic OpenAI SDKs |
| `/llm/*` | Anthropic Messages | the riotpiao frontend (first-party only) |
```
/v1/* (OpenAI) /llm/* (Anthropic)
| |
+-----------+------------+
v
canonical request dialect-neutral
v
slot controller keyed by UPSTREAM, not by route
v
reasoning-predictor / ornith-predictor
```
**The slot controller is keyed by upstream and shared across dialects.** Per-dialect
semaphores are wrong: the 8 sequence slots are physical, so two independent gates
would each believe they were within budget while together exceeding it. Requests from
both surfaces contend for the same slots and the same queue, in arrival order.
Dispatch, budgets, logging and metrics all operate on the canonical request. Adding a
third dialect later must not require touching the controller.
### 3.1 Body-based model dispatch
`POST /v1/chat/completions` selects its upstream from the request body's `model`
field. This is the single most important requirement in this document: it is the
capability Kong OSS lacked, and the reason the gateway exists.
Unknown or missing `model` is a client error with a useful message listing valid
values — not a 500, and not a silent fallback to a default model.
### 3.2 Upstream map
Verified live. `served-model-name` values are what clients send.
| `model` in body | Upstream Service | Engine |
|---|---|---|
| `reasoning` | `reasoning-predictor.llm-serving:80` | vLLM, DeepSeek-R1-Distill-Qwen-32B |
| `ornith:35b` | `ornith-predictor.llm-serving:80` | Ollama |
| `qwen2.5:3b-instruct` | `ornith-predictor.llm-serving:80` | Ollama, same pods |
| `nomic-ai/nomic-embed-text-v2-moe` | `embeddings-predictor.llm-serving:80` | TEI |
| `BAAI/bge-reranker-base` | `reranker-predictor.llm-serving:80` | TEI |
`reasoning` runs 2 replicas × `--max-num-seqs=4` = **8 concurrent sequence slots
total**, `--max-model-len=16384`, `--reasoning-parser=deepseek_r1`,
`--enable-auto-tool-choice --tool-call-parser=hermes`.
Note `ornith:35b` and `qwen2.5:3b-instruct` share pods; both stay resident via
`OLLAMA_MAX_LOADED_MODELS=2` and `OLLAMA_KEEP_ALIVE=-1`, so dispatching between them
does not trigger a model swap.
### 3.3 Path rewriting
Upstreams expect canonical paths. `/v1/chat/completions` and `/v1/embeddings` pass
through unchanged. Rerank is the exception: TEI serves `/rerank`, not `/v1/rerank`,
so that route rewrites.
### 3.4 Legacy path aliases
`/v1/{reasoning,ornith,qwen}/chat/completions` must keep working during cutover —
pi is a live caller. They behave exactly as the canonical endpoint with `model`
forced to the corresponding value, overriding whatever the body says.
These are temporary. They exist to make the cutover reversible, and are removed once
callers have migrated.
### 3.5 `GET /v1/models`
Derived from the configured upstream map, never hardcoded. Kong served a static
list, and its own manifest flags that the list can drift from what the engines
actually serve. The gateway's list must be incapable of disagreeing with what
routing will accept.
OpenAI list shape: `{"object":"list","data":[{"id","object":"model","owned_by","created"}]}`.
### 3.6 Behaviour to preserve
Verified against the live endpoint:
- The upstream returns `reasoning_content` separately from `content` for the
`reasoning` model. Pass both through untouched.
- Tool calling works with explicit `tool_choice`, and is unreliable with
`tool_choice: auto` on the R1-distill model. The gateway does not compensate for
this — it is a model property, not a gateway concern. Do not add retries or
rewriting to work around it.
---
## 4. Authentication — Authentik
### 4.1 Current state
**The model API is unauthenticated today.** Confirmed live: `/v1/reasoning/chat/completions`
answers with no credentials.
Kong's `key-auth` was retired because it accepts a raw `apikey:` header but rejects
`Authorization: Bearer`, which hard-blocks every OpenAI-compatible client. See
`~/workplace/homelab/k8s/apps/api/model-auth.yaml`.
### 4.2 Requirement
Bearer tokens in `Authorization`, validated against Authentik
(`https://authentik.riotpiao.com`) by fetching and caching JWKS at runtime.
Key rotation must be handled by refetching JWKS, not by pinned PEMs. The
pinned-`rsa_public_key` approach in `AUTH-PLAN.md` and its rotation runbook exist
only to route around a Kong OSS limitation and must not be carried forward.
Service accounts obtain tokens via `client_credentials` against Authentik's token
endpoint.
### 4.3 Rollout
Auth ships behind a flag, defaulting off, and is enabled deliberately.
Enabling it breaks every current caller until they hold a token — pi included, whose
`models.json` currently sends a `customHeaders: {apikey: ...}` block that will need
replacing with a Bearer token.
### 4.4 Authorization
Beyond authentication, a token must be checked for the right to invoke the
capability it is calling. A token minted for queue access should not invoke a GPU.
---
## 5. Rate limiting and budgets
No `rate-limiting` plugin exists anywhere in the cluster today — this is net new
work, not a migration. Verified: six Kong plugins exist, none is `rate-limiting`.
Requirements, in priority order:
1. **GPU slot protection.** `reasoning` has 8 total sequence slots. Concurrent
in-flight requests to it must be capped below that, leaving operator headroom.
Excess requests queue up to a bounded depth, then are rejected with a retryable
status.
2. **Per-caller budgets.** Identified callers get a request budget over a window.
3. **Body size caps**, per route.
Rejections use RFC 9457 `application/problem+json` and set `Retry-After` where a
retry time is knowable.
---
## 6. Observability
Kong's cluster-wide `prometheus` plugin is being retired. The gateway must expose at
least equivalent signal or observability regresses at cutover: request rate,
latency, status codes, bandwidth, and upstream health, labelled by route and
upstream.
Gateway-specific signals that Kong could not provide, and which are the reason for
several requirements above: in-flight requests per upstream, queue depth, GPU slot
occupancy, and rejections by reason.
Structured logging. Every rejected request is logged with the reason. No secrets, no
tokens, no request bodies in logs.
---
## 7. Local development and verification
An agent must be able to close a change/verify loop with no cluster, no kubeconfig
and no credentials. This is a hard requirement because it determines whether work can
proceed unattended.
Concretely: it must be possible to start the gateway locally, point it at stub
upstreams, issue requests, and assert on the responses — including streaming
responses and client disconnects.
Verification of any API-shaped task means asserting on the **actual HTTP response**:
status, headers, and body. "It compiles" and "it starts" are not verification.
Parity with Kong is verified by comparing gateway and Kong responses for the same
request, for every route in the migration inventory, before cutover.
---
## 8. Deployment
Container: distroless or scratch, `runAsNonRoot`, read-only root filesystem, all
capabilities dropped, `seccompProfile: RuntimeDefault`, no shell.
Image tags are commit SHAs, never `:latest` — Argo's `selfHeal` cannot roll out a
mutable tag reliably.
NetworkPolicy: egress only to the upstreams it proxies plus Authentik; ingress from
`ingress-nginx` only.
Deployed as an Argo Application in the `homelab-root` GitOps repo. Verified live:
zero Argo Applications anywhere in the cluster source from any Forgejo URL, so
`github.com/Riotpiaole/riotpiao.homelab.com` is authoritative.
---
## 9. Capability surface — path-based
Every capability is a path prefix on the single host `api.riotpiao.com`. One DNS
record, one Cloudflare tunnel hostname, one nginx Ingress, one Service.
| Prefix | Backs onto | Status |
|---|---|---|
| `/v1/*` | `llm-serving` predictors | v1 — **reserved**, see below |
| `/sqs/*` | kmsvc management-service, Kafka/Strimzi (`sqs` ns) | future |
| `/workflow/*` | Temporal (`temporal` ns) | future |
| `/cluster/*` | atlas, separate repo `riotpiao-backend` | future |
| `/db/*` | CloudNativePG, MinIO, monitoring reads | future |
**`/v1/*` is reserved for the OpenAI-compatible surface and nothing else.** G3 pins
it: an SDK expects `/v1/chat/completions` at the base URL, so that prefix can never
be repurposed or nested. Every other capability gets its own prefix that cannot
collide with a current or future OpenAI path.
Subdomains are deliberately *not* used. Paths keep hostname configuration to a
single entry — and hostname configuration is the demonstrated failure mode here, as
the unresolved apex 403 shows. Promoting a prefix to its own subdomain later is an
additive host rule that can run alongside the path; the reverse is not, because
clients hardcode hostnames.
Notes carried from the cluster:
- Temporal namespace registration is automatic via queue-operator, never manual.
- `management-service` already exposes gRPC at `kmsvc.riotpiao.com`; the `/sqs`
prefix is a new surface, not a replacement for it.
- atlas keeps its own informers and RBAC. The gateway proxies to it and holds no
cluster credentials of its own (G2).
- `/db/*` read surfaces need particular care — see G2 before designing them.
---
## 10. Known cluster facts worth not rediscovering
- `kmsvc-redis-master.sqs:6379` has **no authentication**`ALLOW_EMPTY_PASSWORD=yes`,
TLS off. Any workload with network reach has full unauthenticated read/write. A
NetworkPolicy is the only control.
- `reasoning-predictor` listens on port **80**, not 8080.
- `prometheus-operated.monitoring` is **headless** (ClusterIP None) — egress policies
need pod selectors, not ClusterIPs.
- `agent-pod/console` is publicly routed, unauthenticated, accepts free-form prompts
into a shell-capable container, and serves a WebSocket. Putting it behind gateway
auth is a security fix, not merely a port.
- Eight `*.example.com` hosts exist on istio-class Ingresses in `llm-serving`. KServe
defaults, not public, out of scope — do not mistake them for gateway routes.