Files
homelab-frontend/tasks/2.12-anthropic-sse-state-machine.md
T
Story Crater BotandClaude Opus 5 058f11cf2b
CI / Test (push) Canceled after 0s
CI / Vet (push) Canceled after 0s
CI / Build (push) Canceled after 0s
CI / Security (govulncheck) (push) Canceled after 0s
chore: initial commit of Go API gateway
Baseline for the Kong replacement on api.riotpiao.com. Brings the working
tree under version control for the first time: gateway source, the task
board that drives the agent runs, test fixtures, and K8s manifests.

Anchor the gateway ignore rule to the repo root. Unanchored, "gateway"
also matched the cmd/gateway/ source directory, so the program entrypoint
was excluded from every commit.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-19 20:54:34 -07:00

3.8 KiB

2.12 — Anthropic SSE state machine (RED)

Phase: 2 — LLM surfaces Stage: RED Depends on: 2.11, 1.2, 1.3

POST /llm/v1/messages with "stream":true must emit Anthropic SSE. Anthropic uses NAMED events carrying content-block indices; the upstream emits flat OpenAI data-only chunks. Write the failing tests against the event sequence before writing a translator.

  • Every frame has both an event: line and a data: line; a bare data: frame is a failure on this surface
  • Content-Type is text/event-stream
  • Event order for a full response is exactly: message_start, then for each block content_block_start, one or more content_block_delta, content_block_stop, then message_delta, then message_stop
  • message_start carries the message envelope with the client-sent model, role assistant, empty content, and usage.input_tokens
  • The block carrying reasoning_content is index 0 with block type thinking, and its deltas are thinking_delta
  • The block carrying content is index 1 with block type text, and its deltas are text_delta
  • The end of reasoning is only knowable when content first arrives, so the arrival of the first content token MUST emit content_block_stop for index 0 before content_block_start for index 1 — the two blocks never overlap
  • If a response has no reasoning_content at all, the text block is index 0 and no thinking block is started; indices are assigned in emission order with no gaps
  • If a response has reasoning_content and never any content, the thinking block is still closed before message_delta
  • message_delta carries stop_reason and usage.output_tokens; upstream finish_reason stop becomes end_turn and length becomes max_tokens
  • message_stop is the final frame and is emitted exactly once per response
  • Translation is streaming and unbuffered: each upstream chunk is converted and flushed as it arrives, and the response is never accumulated to be inspected
  • A client disconnect mid-stream cancels the upstream request immediately and releases the slot, rather than orphaning the generation
  • An upstream failure after message_start terminates the stream with an error frame rather than a truncated but apparently successful sequence
  • The OpenAI data: [DONE] sentinel is consumed by the translator and never forwarded to a /llm client

An orphaned generation holds one of only eight vLLM sequence slots in the cluster, which is why disconnect cancellation is an acceptance criterion here and not only in the proxy layer.

Verify

curl -N -s localhost:8080/llm/v1/messages -H 'content-type: application/json' \
  -d '{"model":"reasoning","max_tokens":64,"stream":true,"messages":[{"role":"user","content":"hi"}]}' \
  | grep '^event:'
# expected: message_start, content_block_start, content_block_delta..., content_block_stop,
#           content_block_start, content_block_delta..., content_block_stop, message_delta, message_stop

curl -N -s localhost:8080/llm/v1/messages -H 'content-type: application/json' \
  -d '{"model":"reasoning","max_tokens":64,"stream":true,"messages":[{"role":"user","content":"hi"}]}' \
  | grep -n -E 'content_block_stop|"index":1' | head -3
# expected: the index 0 content_block_stop line precedes the first line mentioning index 1

timeout 1 curl -N -s localhost:8080/llm/v1/messages -H 'content-type: application/json' \
  -d '{"model":"reasoning","max_tokens":512,"stream":true,"messages":[{"role":"user","content":"long"}]}' >/dev/null
grep -c 'cancelled' /tmp/stub-reasoning.log
# expected: 1 within a second of the client going away