feat(api): add DeepSeek-shaped LLM API on Kong — /v1/models, per-model chat completions, embeddings, rerank, score; disable Kong response buffering so stream:true actually streams
Kong matches routes on host/path/method/header, never on the request body, so a single /v1/chat/completions dispatching on body.model is not expressible in Kong OSS (ai-proxy-advanced, which does multi-target model routing, is Enterprise). Model therefore goes in the path: GET /v1/models static list (request-termination) POST /v1/reasoning/chat/completions reasoning-predictor (vLLM) POST /v1/ornith/chat/completions ornith-predictor (Ollama) POST /v1/qwen/chat/completions ornith-predictor (Ollama, same pod) POST /v1/embeddings embeddings-predictor (TEI) POST /v1/rerank reranker-predictor (TEI) POST /v1/score verifier-predictor (vLLM pooling) - each chat route force-overwrites body.model via request-transformer add+replace: ornith:35b and qwen2.5:3b-instruct share one Ollama pod, so without this a client hitting /v1/qwen with "model":"ornith:35b" would silently get the 35B - routes live in ns llm-serving, not api: an Ingress can only reference a Service in its own namespace, and KIC watches all namespaces - embeddings and score need no rewrite (TEI/vLLM already serve the canonical paths); rerank does, since /v1/rerank 404s and only /rerank exists - read/write timeouts 1h: Kong defaults to 60s, which a 32B model on Volta exceeds mid-generation and returns 504 - nginx_proxy_proxy_buffering=off: buffered responses lump or stall SSE, and both hops (nginx Ingress and Kong) must be unbuffered or the buffered one wins - no auth for now, per decision; api.riotpiao.com is reachable through nginx, so GPU time is currently unauthenticated
This commit is contained in:
@@ -35,6 +35,16 @@ env:
|
||||
# config in a database mutated through the Admin API — state outside git, plus
|
||||
# migration Jobs on every upgrade.
|
||||
database: "off"
|
||||
# `nginx_proxy_<directive>` injects a directive into the proxy location block;
|
||||
# this renders `proxy_buffering off;`.
|
||||
#
|
||||
# Required for LLM streaming. With buffering on (the default) nginx accumulates
|
||||
# the upstream response before forwarding, so an SSE stream from
|
||||
# `"stream": true` arrives in lumps or stalls until the generation finishes —
|
||||
# which defeats the point of streaming. The matching setting is already on the
|
||||
# nginx Ingress in ingress.yaml; both hops have to be unbuffered or the
|
||||
# buffered one dominates.
|
||||
nginx_proxy_proxy_buffering: "off"
|
||||
|
||||
ingressController:
|
||||
enabled: true
|
||||
|
||||
Reference in New Issue
Block a user