Files
homelab-frontend/REQUIREMENTS.md
T
Story Crater BotandClaude Opus 5 058f11cf2b
CI / Test (push) Canceled after 0s
CI / Vet (push) Canceled after 0s
CI / Build (push) Canceled after 0s
CI / Security (govulncheck) (push) Canceled after 0s
chore: initial commit of Go API gateway
Baseline for the Kong replacement on api.riotpiao.com. Brings the working
tree under version control for the first time: gateway source, the task
board that drives the agent runs, test fixtures, and K8s manifests.

Anchor the gateway ignore rule to the repo root. Unanchored, "gateway"
also matched the cmd/gateway/ source directory, so the program entrypoint
was excluded from every commit.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-19 20:54:34 -07:00

15 KiB
Raw Blame History

homelab-frontend — Requirements

The contract for the Go API gateway that replaces Kong OSS on *.riotpiao.com.

Companion documents:

All cluster facts below were verified live against context admin@homelab-cluster on 2026-08-19. Re-verify before relying on any number.


0. Invariants

These hold for every surface. A change that breaks one of these is a design change, not an implementation detail.

  • G1 — ingress-nginx owns TLS and the edge. The gateway never terminates TLS.
  • G2 — The gateway holds no Kubernetes credentials. It proxies to services that do. Cluster-read permissions stay in atlas, out of the public edge process.
  • G3 — Public surfaces use standard protocol shapes. If an OpenAI SDK cannot call it unmodified, the design is wrong.
  • G4 — Streaming is unbuffered end to end, and a client disconnect cancels the upstream request rather than orphaning it.
  • G5 — Authentication is Bearer-token, validated against Authentik via JWKS fetched at runtime. No pinned public keys, no rotation runbook.
  • G6 — Every route's timeout, body cap and concurrency limit is explicit in configuration. No silent defaults.
  • G7 — All deployment flows through git and Argo. No kubectl apply, no helm upgrade, no local terraform apply.

1. Runtime and configuration

1.1 Process

Single static Go binary. Reads configuration at startup, serves HTTP, exits cleanly on SIGTERM after draining in-flight requests.

Must run with no cluster, no kubeconfig and no credentials so that behaviour can be verified in a closed loop before touching live traffic. Upstreams are configuration; pointing them at local stubs is the entire mechanism. This is a hard requirement, not a convenience — see §7.

1.2 Configuration

Route and upstream configuration is declarative and loaded at startup. It must express, per upstream: address, path rewrite, connect/read/write timeouts, maximum body size, and whether the route requires authentication.

Configuration errors fail startup loudly. A gateway that starts with a silently dropped route is worse than one that refuses to start.

Configuration lives in git, mounted as a ConfigMap and synced by Argo. Not a CRD. A CRD would require the gateway to watch the API server, which needs RBAC and contradicts G2 — and CRD-driven routing is precisely the indirection being retired with Kong, where the routing table was split across six KongPlugin CRs, seven Ingresses and a Helm values file.

A CRD earns its keep when someone other than the repo owner must register routes. That is not true here. If it becomes true, the additive answer is a controller that renders this same ConfigMap — the gateway stays credential-free either way.

1.3 Health

  • GET /healthz — liveness, no upstream checks, always cheap.
  • GET /readyz — readiness; may fail while configuration is invalid or JWKS has never been successfully fetched.

Neither requires authentication.


2. Proxy core

2.1 Reverse proxying

Standard reverse proxy to configured upstreams. Connection reuse across requests. Hop-by-hop headers stripped correctly. X-Forwarded-* set from the nginx-supplied values, not fabricated.

2.2 Streaming

SSE and chunked responses pass through without buffering. Tokens must reach the client as the upstream emits them, not on completion.

WebSocket upgrade must work — agent-pod/console depends on it.

2.3 Disconnect propagation

When a client disconnects, the upstream request is cancelled immediately. This is load-bearing: an orphaned generation holds a vLLM sequence slot, and there are only eight in the cluster.

2.4 Timeouts

Per-route, explicit. Current Kong values, which are deliberate and must be preserved unless changed knowingly:

Route class connect read write
chat 10s 1h 1h
embeddings, rerank 10s 10m 10m

The 1-hour read timeout exists because a 32B model on a Volta GPU routinely exceeds 60s. Any shorter application-level cap must be enforced by the gateway's own logic, not by shortening the proxy timeout — otherwise long legitimate generations truncate mid-stream.


3. LLM surfaces — api.riotpiao.com

Two protocol dialects, permanently. Both translate into one dialect-neutral canonical request, and both pass through one shared slot controller before reaching a predictor.

Prefix Dialect Primary client
/v1/* OpenAI-compatible pi, generic OpenAI SDKs
/llm/* Anthropic Messages the riotpiao frontend (first-party only)
  /v1/*  (OpenAI)        /llm/*  (Anthropic)
        |                        |
        +-----------+------------+
                    v
            canonical request        dialect-neutral
                    v
             slot controller         keyed by UPSTREAM, not by route
                    v
        reasoning-predictor / ornith-predictor

The slot controller is keyed by upstream and shared across dialects. Per-dialect semaphores are wrong: the 8 sequence slots are physical, so two independent gates would each believe they were within budget while together exceeding it. Requests from both surfaces contend for the same slots and the same queue, in arrival order.

Dispatch, budgets, logging and metrics all operate on the canonical request. Adding a third dialect later must not require touching the controller.

3.1 Body-based model dispatch

POST /v1/chat/completions selects its upstream from the request body's model field. This is the single most important requirement in this document: it is the capability Kong OSS lacked, and the reason the gateway exists.

Unknown or missing model is a client error with a useful message listing valid values — not a 500, and not a silent fallback to a default model.

3.2 Upstream map

Verified live. served-model-name values are what clients send.

model in body Upstream Service Engine
reasoning reasoning-predictor.llm-serving:80 vLLM, DeepSeek-R1-Distill-Qwen-32B
ornith:35b ornith-predictor.llm-serving:80 Ollama
qwen2.5:3b-instruct ornith-predictor.llm-serving:80 Ollama, same pods
nomic-ai/nomic-embed-text-v2-moe embeddings-predictor.llm-serving:80 TEI
BAAI/bge-reranker-base reranker-predictor.llm-serving:80 TEI

reasoning runs 2 replicas × --max-num-seqs=4 = 8 concurrent sequence slots total, --max-model-len=16384, --reasoning-parser=deepseek_r1, --enable-auto-tool-choice --tool-call-parser=hermes.

Note ornith:35b and qwen2.5:3b-instruct share pods; both stay resident via OLLAMA_MAX_LOADED_MODELS=2 and OLLAMA_KEEP_ALIVE=-1, so dispatching between them does not trigger a model swap.

3.3 Path rewriting

Upstreams expect canonical paths. /v1/chat/completions and /v1/embeddings pass through unchanged. Rerank is the exception: TEI serves /rerank, not /v1/rerank, so that route rewrites.

3.4 Legacy path aliases

/v1/{reasoning,ornith,qwen}/chat/completions must keep working during cutover — pi is a live caller. They behave exactly as the canonical endpoint with model forced to the corresponding value, overriding whatever the body says.

These are temporary. They exist to make the cutover reversible, and are removed once callers have migrated.

3.5 GET /v1/models

Derived from the configured upstream map, never hardcoded. Kong served a static list, and its own manifest flags that the list can drift from what the engines actually serve. The gateway's list must be incapable of disagreeing with what routing will accept.

OpenAI list shape: {"object":"list","data":[{"id","object":"model","owned_by","created"}]}.

3.6 Behaviour to preserve

Verified against the live endpoint:

  • The upstream returns reasoning_content separately from content for the reasoning model. Pass both through untouched.
  • Tool calling works with explicit tool_choice, and is unreliable with tool_choice: auto on the R1-distill model. The gateway does not compensate for this — it is a model property, not a gateway concern. Do not add retries or rewriting to work around it.

4. Authentication — Authentik

4.1 Current state

The model API is unauthenticated today. Confirmed live: /v1/reasoning/chat/completions answers with no credentials.

Kong's key-auth was retired because it accepts a raw apikey: header but rejects Authorization: Bearer, which hard-blocks every OpenAI-compatible client. See ~/workplace/homelab/k8s/apps/api/model-auth.yaml.

4.2 Requirement

Bearer tokens in Authorization, validated against Authentik (https://authentik.riotpiao.com) by fetching and caching JWKS at runtime.

Key rotation must be handled by refetching JWKS, not by pinned PEMs. The pinned-rsa_public_key approach in AUTH-PLAN.md and its rotation runbook exist only to route around a Kong OSS limitation and must not be carried forward.

Service accounts obtain tokens via client_credentials against Authentik's token endpoint.

4.3 Rollout

Auth ships behind a flag, defaulting off, and is enabled deliberately.

Enabling it breaks every current caller until they hold a token — pi included, whose models.json currently sends a customHeaders: {apikey: ...} block that will need replacing with a Bearer token.

4.4 Authorization

Beyond authentication, a token must be checked for the right to invoke the capability it is calling. A token minted for queue access should not invoke a GPU.


5. Rate limiting and budgets

No rate-limiting plugin exists anywhere in the cluster today — this is net new work, not a migration. Verified: six Kong plugins exist, none is rate-limiting.

Requirements, in priority order:

  1. GPU slot protection. reasoning has 8 total sequence slots. Concurrent in-flight requests to it must be capped below that, leaving operator headroom. Excess requests queue up to a bounded depth, then are rejected with a retryable status.
  2. Per-caller budgets. Identified callers get a request budget over a window.
  3. Body size caps, per route.

Rejections use RFC 9457 application/problem+json and set Retry-After where a retry time is knowable.


6. Observability

Kong's cluster-wide prometheus plugin is being retired. The gateway must expose at least equivalent signal or observability regresses at cutover: request rate, latency, status codes, bandwidth, and upstream health, labelled by route and upstream.

Gateway-specific signals that Kong could not provide, and which are the reason for several requirements above: in-flight requests per upstream, queue depth, GPU slot occupancy, and rejections by reason.

Structured logging. Every rejected request is logged with the reason. No secrets, no tokens, no request bodies in logs.


7. Local development and verification

An agent must be able to close a change/verify loop with no cluster, no kubeconfig and no credentials. This is a hard requirement because it determines whether work can proceed unattended.

Concretely: it must be possible to start the gateway locally, point it at stub upstreams, issue requests, and assert on the responses — including streaming responses and client disconnects.

Verification of any API-shaped task means asserting on the actual HTTP response: status, headers, and body. "It compiles" and "it starts" are not verification.

Parity with Kong is verified by comparing gateway and Kong responses for the same request, for every route in the migration inventory, before cutover.


8. Deployment

Container: distroless or scratch, runAsNonRoot, read-only root filesystem, all capabilities dropped, seccompProfile: RuntimeDefault, no shell.

Image tags are commit SHAs, never :latest — Argo's selfHeal cannot roll out a mutable tag reliably.

NetworkPolicy: egress only to the upstreams it proxies plus Authentik; ingress from ingress-nginx only.

Deployed as an Argo Application in the homelab-root GitOps repo. Verified live: zero Argo Applications anywhere in the cluster source from any Forgejo URL, so github.com/Riotpiaole/riotpiao.homelab.com is authoritative.


9. Capability surface — path-based

Every capability is a path prefix on the single host api.riotpiao.com. One DNS record, one Cloudflare tunnel hostname, one nginx Ingress, one Service.

Prefix Backs onto Status
/v1/* llm-serving predictors v1 — reserved, see below
/sqs/* kmsvc management-service, Kafka/Strimzi (sqs ns) future
/workflow/* Temporal (temporal ns) future
/cluster/* atlas, separate repo riotpiao-backend future
/db/* CloudNativePG, MinIO, monitoring reads future

/v1/* is reserved for the OpenAI-compatible surface and nothing else. G3 pins it: an SDK expects /v1/chat/completions at the base URL, so that prefix can never be repurposed or nested. Every other capability gets its own prefix that cannot collide with a current or future OpenAI path.

Subdomains are deliberately not used. Paths keep hostname configuration to a single entry — and hostname configuration is the demonstrated failure mode here, as the unresolved apex 403 shows. Promoting a prefix to its own subdomain later is an additive host rule that can run alongside the path; the reverse is not, because clients hardcode hostnames.

Notes carried from the cluster:

  • Temporal namespace registration is automatic via queue-operator, never manual.
  • management-service already exposes gRPC at kmsvc.riotpiao.com; the /sqs prefix is a new surface, not a replacement for it.
  • atlas keeps its own informers and RBAC. The gateway proxies to it and holds no cluster credentials of its own (G2).
  • /db/* read surfaces need particular care — see G2 before designing them.

10. Known cluster facts worth not rediscovering

  • kmsvc-redis-master.sqs:6379 has no authenticationALLOW_EMPTY_PASSWORD=yes, TLS off. Any workload with network reach has full unauthenticated read/write. A NetworkPolicy is the only control.
  • reasoning-predictor listens on port 80, not 8080.
  • prometheus-operated.monitoring is headless (ClusterIP None) — egress policies need pod selectors, not ClusterIPs.
  • agent-pod/console is publicly routed, unauthenticated, accepts free-form prompts into a shell-capable container, and serves a WebSocket. Putting it behind gateway auth is a security fix, not merely a port.
  • Eight *.example.com hosts exist on istio-class Ingresses in llm-serving. KServe defaults, not public, out of scope — do not mistake them for gateway routes.