# Kong retirement — inventory and cutover Everything Kong does on `api.riotpiao.com` today, and where it goes. Inventory verified live against context `admin@homelab-cluster` on 2026-08-19. Source of the objects being retired: `~/workplace/homelab/k8s/apps/api/` and `k8s/argocd/apps/55-api-gateway.yaml`. ## What is running now Kong OSS 3.4.1, Helm chart from `https://charts.konghq.com`, DB-less, namespace `api`, Argo Application `kong` at sync wave 7. Two replicas. Fronted by `ingress-nginx` via Ingress `api/api`, which catch-alls `/` on `api.riotpiao.com` to `kong-proxy:80`. Eleven ReplicaSets exist on the Kong Deployment, the newest minutes old — this config is being actively iterated, so re-verify the inventory immediately before cutover. ## Routing table to port Seven `ingressClassName: kong` Ingresses. Six in `llm-serving`, one in `agent-pod`. | Method | Path | Upstream | Transform applied by Kong | |---|---|---|---| | GET | `/v1/models` | — | `request-termination`: static 200 JSON, upstream never contacted | | POST | `/v1/reasoning/chat/completions` | `reasoning-predictor:80` | force body `model=reasoning`, rewrite URI to `/v1/chat/completions` | | POST | `/v1/ornith/chat/completions` | `ornith-predictor:80` | force body `model=ornith:35b`, rewrite URI | | POST | `/v1/qwen/chat/completions` | `ornith-predictor:80` | force body `model=qwen2.5:3b-instruct`, rewrite URI | | POST | `/v1/embeddings` | `embeddings-predictor:80` | none — TEI already serves the canonical path | | POST | `/v1/rerank` | `reranker-predictor:80` | rewrite URI to `/rerank` (TEI does not serve `/v1/rerank`) | | GET/WS | `/console`, `/run`, `/sessions` | `agent-hub:9090` (`agent-pod` ns) | none, `strip-path: false` | Upstream model map, from the manifest comments and confirmed live: - `reasoning` → `reasoning-predictor` — vLLM, DeepSeek-R1-Distill-Qwen-32B, 2 replicas, `--max-num-seqs=4`, `--max-model-len=16384`, `--reasoning-parser=deepseek_r1`, `--enable-auto-tool-choice --tool-call-parser=hermes` - `ornith:35b` → `ornith-predictor` — Ollama, 2 replicas - `qwen2.5:3b-instruct` → `ornith-predictor` — same pods; both models stay resident via `OLLAMA_MAX_LOADED_MODELS=2`, `OLLAMA_KEEP_ALIVE=-1` - `nomic-ai/nomic-embed-text-v2-moe` → `embeddings-predictor` — TEI - `BAAI/bge-reranker-base` → `reranker-predictor` — TEI ### The path-per-model surface goes away The three chat paths exist only because Kong OSS cannot dispatch on the request body. The gateway serves a single `POST /v1/chat/completions` and selects the upstream from the body's `model` field. Keep the old paths as aliases during cutover so live clients do not break, then remove them once callers have migrated. pi is a live caller today. ### `/v1/models` should not be ported verbatim Kong serves a hardcoded list via `request-termination`. The manifest already flags that it can drift from what the engines actually serve. Derive the response from the gateway's configured upstream map instead, so the list cannot disagree with what routing will accept. ## Plugins being retired | Plugin | Scope | Replacement | |---|---|---| | `llm-rewrite-reasoning` / `-ornith` / `-qwen` | llm-serving | body-based dispatch in `internal/llm` | | `llm-rewrite-rerank` | llm-serving | per-upstream path rewrite in the route table | | `llm-models-list` | llm-serving | derived from the upstream map | | `prometheus` | **cluster-wide** | `internal/observability` — must expose bandwidth, latency, status codes, upstream health or observability regresses | No `rate-limiting` plugin exists anywhere in the cluster. REQUIREMENTS.md §4 Tier 2 describes it as an existing layer; it is not built. Nothing to migrate — it is net new work, and it now belongs in the gateway rather than in Kong. ## Auth: currently off, must land on `KongConsumer model-invoker` exists in namespace `api` and stays defined, but the `key-auth` plugin is commented out and every route has `model-key-auth` stripped from its `konghq.com/plugins` annotation. **The model API is unauthenticated right now.** Confirmed live 2026-08-19: a request to `/v1/reasoning/chat/completions` with no credentials returns 200. The reason is recorded in `model-auth.yaml` — Kong's `key-auth` accepts a raw `apikey:` header but rejects `Authorization: Bearer`, which blocks every OpenAI-compatible client. That is why `~/.pi/agent/models.json` carries a `customHeaders: {apikey: ...}` block. The gateway reads Bearer tokens directly and validates them against Authentik via JWKS. `AUTH-PLAN.md`'s pinned-RSA-key approach and its rotation runbook are not needed and should not be carried over. Ship auth behind a flag. Turning it on breaks every current caller until they hold a token — pi included. ## Timeouts Kong today: | Route class | connect | read | write | |---|---|---|---| | chat | 10s | **1h** | 1h | | embeddings / rerank | 10s | 10m | 10m | nginx in front sets `proxy-read-timeout: 3600`, `proxy-send-timeout: 3600`, `proxy-buffering: off`, `proxy-body-size: 0`. Those stay — they are what makes token streaming work, and the gateway needs the same treatment from nginx. The 1-hour read timeout is deliberate: a 32B model on a Volta GPU routinely exceeds 60s. Any shorter server-side cap must be enforced *in the gateway*, not by shortening the proxy timeout, or long legitimate generations get truncated mid-stream. ## Cutover Reversible at every step. Kong keeps serving until the last step. 1. Deploy the gateway alongside Kong, unexposed. Verify in-cluster against `http://homelab-frontend.api.svc.cluster.local`. 2. Compare gateway and Kong responses for every route in the table above, including a streaming chat request and a client disconnect mid-stream. 3. Repoint Ingress `api/api` from `kong-proxy:80` to the gateway Service. **This is the cutover.** Reverting is a one-line change to the same Ingress. 4. Soak. Watch gateway metrics and pi traffic. 5. Delete the seven kong-class Ingresses and the six KongPlugin CRs. 6. Remove the `kong` Application from `k8s/argocd/apps/55-api-gateway.yaml`; let Argo prune the Helm release, the CRDs and namespace leftovers. Steps 1–4 are reversible in seconds. Step 5 onward is not — do not start it until the soak is clean. All of this flows through git and Argo. No `kubectl apply`, no `helm upgrade`. ## Loose ends - `agent-pod/console` is publicly routed, unauthenticated, accepts free-form prompts into a shell-capable container, and exposes a WebSocket. Migrating it behind the gateway's auth is a security fix, not merely a port. Treat WebSocket upgrade as an explicit requirement of the proxy layer. - Eight `*.example.com` hosts exist on istio-class Ingresses in `llm-serving` (`{embeddings,ornith,reasoning,reranker}[-predictor]-llm-serving.example.com`). KServe defaults, not public, not Kong's — out of scope here, but they exist and should not be mistaken for gateway routes. - Ingress class split across the cluster is 7 kong / 17 nginx / 4 istio. Only the 7 kong ones are in scope.