Files
homelab-frontend/tasks/2.12-anthropic-sse-state-machine.md
T
Story Crater BotandClaude Opus 5 058f11cf2b
CI / Test (push) Canceled after 0s
CI / Vet (push) Canceled after 0s
CI / Build (push) Canceled after 0s
CI / Security (govulncheck) (push) Canceled after 0s
chore: initial commit of Go API gateway
Baseline for the Kong replacement on api.riotpiao.com. Brings the working
tree under version control for the first time: gateway source, the task
board that drives the agent runs, test fixtures, and K8s manifests.

Anchor the gateway ignore rule to the repo root. Unanchored, "gateway"
also matched the cmd/gateway/ source directory, so the program entrypoint
was excluded from every commit.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-19 20:54:34 -07:00

65 lines
3.8 KiB
Markdown

# 2.12 — Anthropic SSE state machine (RED)
Phase: 2 — LLM surfaces
Stage: RED
Depends on: [2.11](2.11-anthropic-response-translation.md), [1.2](1.2-streaming-passthrough.md), [1.3](1.3-disconnect-propagation.md)
`POST /llm/v1/messages` with `"stream":true` must emit Anthropic SSE. Anthropic uses
NAMED events carrying content-block indices; the upstream emits flat OpenAI data-only
chunks. Write the failing tests against the event sequence before writing a translator.
- [ ] Every frame has both an `event:` line and a `data:` line; a bare `data:` frame is
a failure on this surface
- [ ] `Content-Type` is `text/event-stream`
- [ ] Event order for a full response is exactly: `message_start`, then for each block
`content_block_start`, one or more `content_block_delta`, `content_block_stop`,
then `message_delta`, then `message_stop`
- [ ] `message_start` carries the message envelope with the client-sent model, `role`
`assistant`, empty `content`, and `usage.input_tokens`
- [ ] The block carrying `reasoning_content` is index 0 with block type `thinking`, and
its deltas are `thinking_delta`
- [ ] The block carrying `content` is index 1 with block type `text`, and its deltas
are `text_delta`
- [ ] The end of reasoning is only knowable when `content` first arrives, so the
arrival of the first `content` token MUST emit `content_block_stop` for index 0
before `content_block_start` for index 1 — the two blocks never overlap
- [ ] If a response has no `reasoning_content` at all, the text block is index 0 and no
thinking block is started; indices are assigned in emission order with no gaps
- [ ] If a response has `reasoning_content` and never any `content`, the thinking block
is still closed before `message_delta`
- [ ] `message_delta` carries `stop_reason` and `usage.output_tokens`; upstream
`finish_reason` `stop` becomes `end_turn` and `length` becomes `max_tokens`
- [ ] `message_stop` is the final frame and is emitted exactly once per response
- [ ] Translation is streaming and unbuffered: each upstream chunk is converted and
flushed as it arrives, and the response is never accumulated to be inspected
- [ ] A client disconnect mid-stream cancels the upstream request immediately and
releases the slot, rather than orphaning the generation
- [ ] An upstream failure after `message_start` terminates the stream with an error
frame rather than a truncated but apparently successful sequence
- [ ] The OpenAI `data: [DONE]` sentinel is consumed by the translator and never
forwarded to a `/llm` client
An orphaned generation holds one of only eight vLLM sequence slots in the cluster,
which is why disconnect cancellation is an acceptance criterion here and not only in
the proxy layer.
## Verify
```bash
curl -N -s localhost:8080/llm/v1/messages -H 'content-type: application/json' \
-d '{"model":"reasoning","max_tokens":64,"stream":true,"messages":[{"role":"user","content":"hi"}]}' \
| grep '^event:'
# expected: message_start, content_block_start, content_block_delta..., content_block_stop,
# content_block_start, content_block_delta..., content_block_stop, message_delta, message_stop
curl -N -s localhost:8080/llm/v1/messages -H 'content-type: application/json' \
-d '{"model":"reasoning","max_tokens":64,"stream":true,"messages":[{"role":"user","content":"hi"}]}' \
| grep -n -E 'content_block_stop|"index":1' | head -3
# expected: the index 0 content_block_stop line precedes the first line mentioning index 1
timeout 1 curl -N -s localhost:8080/llm/v1/messages -H 'content-type: application/json' \
-d '{"model":"reasoning","max_tokens":512,"stream":true,"messages":[{"role":"user","content":"long"}]}' >/dev/null
grep -c 'cancelled' /tmp/stub-reasoning.log
# expected: 1 within a second of the client going away
```