Files
homelab-frontend/docs/API-llm.md
T
Story Crater BotandClaude Opus 5 058f11cf2b
CI / Test (push) Canceled after 0s
CI / Vet (push) Canceled after 0s
CI / Build (push) Canceled after 0s
CI / Security (govulncheck) (push) Canceled after 0s
chore: initial commit of Go API gateway
Baseline for the Kong replacement on api.riotpiao.com. Brings the working
tree under version control for the first time: gateway source, the task
board that drives the agent runs, test fixtures, and K8s manifests.

Anchor the gateway ignore rule to the repo root. Unanchored, "gateway"
also matched the cmd/gateway/ source directory, so the program entrypoint
was excluded from every commit.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-19 20:54:34 -07:00

12 KiB

API — LLM surfaces

Two protocol dialects over the same models and the same slot controller.

Prefix Dialect Endpoint Client
/v1 OpenAI-compatible POST /v1/chat/completions pi, OpenAI SDKs
/llm Anthropic Messages POST /llm/v1/messages riotpiao frontend (first-party)

Status marks below: [LIVE] verified against the running cluster on 2026-08-19. [SPEC] the contract this gateway must implement; not built yet.


Models

model value Upstream Engine Context Notes
reasoning reasoning-predictor.llm-serving:80 vLLM, DeepSeek-R1-Distill-Qwen-32B 16384 emits reasoning_content; 8 sequence slots total
ornith:35b ornith-predictor.llm-serving:80 Ollama 131072 reliable tool calling
qwen2.5:3b-instruct ornith-predictor.llm-serving:80 Ollama 32768 same pods as ornith
nomic-ai/nomic-embed-text-v2-moe embeddings-predictor.llm-serving:80 TEI embeddings only
BAAI/bge-reranker-base reranker-predictor.llm-serving:80 TEI rerank only

reasoning runs 2 replicas x --max-num-seqs=4. Those 8 slots are the scarcest resource in the cluster and are shared across both dialects.


Authentication [SPEC]

Ships behind a flag, default off. The model API is unauthenticated today.

Authorization: Bearer <authentik-jwt>

Decided — Bearer on both surfaces

Authorization: Bearer <jwt> is the only accepted credential, on /v1 and /llm alike. One auth path, consistent with G5, validated against Authentik via JWKS.

Known divergence from Anthropic: the real Anthropic API authenticates with x-api-key and requires anthropic-version: 2023-06-01. A stock Anthropic SDK pointed at /llm will send x-api-key and get a 401.

This is accepted, not overlooked. The /llm client is the first-party riotpiao frontend, which sends whatever we tell it to. If a real Anthropic SDK ever needs to reach this gateway, accepting x-api-key as a second credential source is an additive change — a small branch in one middleware, not a redesign.

anthropic-version is accepted and ignored if present, and never required.

The 401 for an x-api-key-only request must name the problem — say that Bearer is required — rather than returning a bare 401. The Kong retirement was caused by exactly this failure mode: a gateway that rejected the header clients actually send, without saying why.


OpenAI dialect — POST /v1/chat/completions

Request [SPEC]

{
  "model": "reasoning",
  "messages": [{"role": "user", "content": "Why is wave 4 empty?"}],
  "max_tokens": 2000,
  "temperature": 0.7,
  "stream": false
}

model is required and selects the upstream. The body is forwarded byte-identical — the gateway reads model, it does not rewrite it.

Response, non-streaming [LIVE]

Captured verbatim from reasoning on 2026-08-19, abridged:

{
  "id": "chatcmpl-f17bd2fe22e4276d24e9438e40e89cea",
  "object": "chat.completion",
  "created": 1787172340,
  "model": "reasoning",
  "choices": [{
    "index": 0,
    "message": {
      "role": "assistant",
      "content": "\n\nTo find the current weather in Toronto...",
      "reasoning_content": "Okay, so I need to figure out...",
      "tool_calls": []
    },
    "finish_reason": "length"
  }],
  "usage": {"prompt_tokens": 16, "completion_tokens": 300, "total_tokens": 316}
}

reasoning_content is a sibling of content, not nested in it. This is a vLLM extension produced by --reasoning-parser=deepseek_r1; it is not part of the OpenAI spec. Pass it through untouched.

The two engines disagree on the field name [LIVE]

Verified 2026-08-19 by calling both:

Upstream Engine Reasoning field
reasoning-predictor vLLM reasoning_content
ornith-predictor Ollama reasoning

Neither is in the OpenAI spec, so neither is wrong — they are two vendor extensions that happen to mean the same thing. The gateway must recognise both when mapping to the Anthropic thinking block, or ornith:35b responses will silently lose their reasoning on the /llm surface.

Do not normalise them on the /v1 surface. That surface passes bodies through untouched, and a client asking for ornith:35b should get exactly what Ollama sent. Normalisation belongs in the canonical request model (task 2.9), which is the layer that exists to absorb precisely this kind of upstream difference.

Response, streaming [SPEC]

data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"reasoning_content":"Okay"},"finish_reason":null}]}

data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Wave"},"finish_reason":null}]}

data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]

Data-only frames, no event: lines. Terminated by the literal data: [DONE].

Legacy aliases [LIVE, being retired]

POST /v1/{reasoning,ornith,qwen}/chat/completions force model to the corresponding value regardless of the body. They exist only because Kong could not dispatch on the body. Removed once callers migrate.

GET /v1/models [SPEC]

{"object":"list","data":[{"id":"reasoning","object":"model","owned_by":"homelab","created":0}]}

Derived from the registry, never hardcoded.

Errors [SPEC]

RFC 9457 application/problem+json:

{
  "type": "https://riotpiao.com/errors/unknown-model",
  "title": "Unknown model",
  "status": 400,
  "detail": "\"gpt-4\" is not available",
  "validModels": ["reasoning", "ornith:35b", "qwen2.5:3b-instruct"]
}

Anthropic dialect — POST /llm/v1/messages [SPEC]

Path note: the Anthropic SDK appends /v1/messages to its base URL, so a base URL of https://api.riotpiao.com/llm produces exactly this path.

Request

{
  "model": "reasoning",
  "max_tokens": 2000,
  "system": "You are a cluster assistant.",
  "messages": [
    {"role": "user", "content": "Why is wave 4 empty?"}
  ],
  "stream": true
}

Differences from the OpenAI dialect that the translator must handle:

Concern OpenAI Anthropic
system prompt messages[0].role = "system" top-level system field
max_tokens optional required
content string string or block array
roles system/user/assistant/tool user/assistant only
stop stop stop_sequences

max_tokens being required is a real divergence — the gateway must either reject its absence with a clear error or apply a documented default. Pick one and state it; do not silently default.

Response, non-streaming

{
  "id": "msg_01ABC",
  "type": "message",
  "role": "assistant",
  "model": "reasoning",
  "content": [
    {"type": "thinking", "thinking": "Waves are sort keys, not a sequence..."},
    {"type": "text", "text": "Wave 4 is empty. Waves are sort keys..."}
  ],
  "stop_reason": "end_turn",
  "stop_sequence": null,
  "usage": {"input_tokens": 16, "output_tokens": 300}
}

Field mapping from the upstream OpenAI response:

Upstream Anthropic
choices[0].message.reasoning_content content[] block {"type":"thinking","thinking":...}
choices[0].message.content content[] block {"type":"text","text":...}
finish_reason: "stop" stop_reason: "end_turn"
finish_reason: "length" stop_reason: "max_tokens"
usage.prompt_tokens usage.input_tokens
usage.completion_tokens usage.output_tokens

The thinking block precedes the text block.

Response, streaming

Anthropic SSE uses named events with content-block indices, unlike OpenAI's flat frames. Verified event sequence:

event: message_start
data: {"type":"message_start","message":{"id":"msg_01ABC","type":"message","role":"assistant","model":"reasoning","content":[],"usage":{"input_tokens":16,"output_tokens":0}}}

event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"thinking","thinking":""}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"thinking_delta","thinking":"Waves are sort keys"}}

event: content_block_stop
data: {"type":"content_block_stop","index":0}

event: content_block_start
data: {"type":"content_block_start","index":1,"content_block":{"type":"text","text":""}}

event: content_block_delta
data: {"type":"content_block_delta","index":1,"delta":{"type":"text_delta","text":"Wave 4 is empty."}}

event: content_block_stop
data: {"type":"content_block_stop","index":1}

event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn","stop_sequence":null},"usage":{"output_tokens":300}}

event: message_stop
data: {"type":"message_stop"}

Three details that are easy to get wrong:

  • In message_delta, usage is a sibling of delta, not inside it.
  • The delta field name matches the delta type: thinking_delta carries .thinking, text_delta carries .text.
  • Block index 0 is thinking, index 1 is text. You only learn reasoning has ended when content first appears in an upstream chunk, so the thinking block must be closed before the text block opens. If a response has no reasoning_content at all, the text block is index 0 and no thinking block is emitted.

Queue position — non-standard extension

Anthropic's event set has no way to say "you are queued", because the stream implicitly begins after a slot is acquired. With only 8 slots, queueing is normal here.

Emitted before message_start:

event: queue
data: {"type":"queue","position":3}

This is deliberately outside the Anthropic spec. It is safe only because the client is first-party; a strict Anthropic client would ignore the unknown event and show nothing while queued.

Errors

Anthropic error shape, not RFC 9457 — the same rejection renders differently depending on which surface received it:

{"type":"error","error":{"type":"invalid_request_error","message":"Unknown model \"gpt-4\". Available: reasoning, ornith:35b, qwen2.5:3b-instruct"}}
Condition HTTP error.type
unknown or missing model 400 invalid_request_error
max_tokens absent (if required) 400 invalid_request_error
malformed JSON 400 invalid_request_error
unsupported feature requested 400 invalid_request_error
not authenticated 401 authentication_error
budget exhausted or queue full 429 rate_limit_error
upstream failure 502 api_error

Deliberately not implemented

Each returns 400 naming the unsupported feature — never a silent partial implementation:

tool use and tool_result turns, image content blocks, prompt-caching headers, the batch API, multi-block user content, thinking.budget_tokens configuration.

The target client is the riotpiao frontend. Widening scope is a code change with a test, not an accident.


Shared behaviour, both dialects

One slot controller, keyed by upstream. A /v1 request and a /llm request contend for the same 8 reasoning slots and the same queue, in arrival order. Per-dialect semaphores would each believe they were within budget while together exceeding the physical limit.

Streaming is unbuffered and a client disconnect cancels the upstream immediately. An orphaned generation holds a slot until it completes on its own, which for a 32B model on a Volta GPU can run to minutes.

Timeouts [LIVE]: chat routes are connect 10s / read 1h / write 1h. The hour is deliberate — a 32B model on this hardware routinely exceeds 60s. Any shorter application cap is enforced in gateway logic, never by shortening the proxy timeout.

Tool calling [LIVE]: reasoning honours an explicit tool_choice but returns tool_calls: [] under tool_choice: "auto" — it reasons about the tool in prose instead. ornith:35b returns finish_reason: "tool_calls" correctly under auto. This is a model property; the gateway does not compensate for it.


Examples

# OpenAI dialect
curl -s https://api.riotpiao.com/v1/chat/completions \
  -H 'content-type: application/json' \
  -d '{"model":"reasoning","messages":[{"role":"user","content":"Why is wave 4 empty?"}],"max_tokens":500}'

# Anthropic dialect, streaming
curl -N -s https://api.riotpiao.com/llm/v1/messages \
  -H 'content-type: application/json' \
  -d '{"model":"reasoning","max_tokens":500,"stream":true,
       "messages":[{"role":"user","content":"Why is wave 4 empty?"}]}'

# model list
curl -s https://api.riotpiao.com/v1/models