Files
riotpiao.com/tasks/06-phase5-chat.md
T

2.7 KiB

Phase 5 — Surface D: chat + rate limiter

~5 days, one PR. Highest risk, highest cost — ships last, ships with its limiter, never after.

Depends on 00-decisions.md #4 (is chat in scope for v1).

RED

  • ratelimit_test.go — 7th concurrent chat queues rather than reaching vLLM
  • disconnect_test.go — client abort cancels the upstream request
  • budget_test.go — 13th message in 24h returns 429 with Retry-After
  • injection_test.go — snapshot content cannot alter system-prompt behaviour
  • context_test.go — history truncation keeps total tokens under 16384

GREEN

  • POST /api/chat — session + Turnstile, 12/day/session, 6 global concurrent, SSE token stream
  • Rate limiting Tier 3 (atlas): global chat semaphore = 6 (2 of 8 GPU slots kept as operator headroom); queue depth 20 then reject 429; per-session budget 12 msg/24h; per-request timeout 120s hard server-side; disconnect cancels upstream immediately (req.Context() threaded to vLLM request)
  • Queue position streamed as SSE {"type":"queue","position":N}
  • Prompt injection defenses: system prompt is compile-time constant, unreachable by user input; cluster snapshot digest injected in a delimited block explicitly labelled untrusted data; user message always last; no tool-calling — model reads pre-built digest, cannot query anything; max_tokens: 1500 cap
  • Context budget (16384 total): system prompt ~300, snapshot digest capped at 2000, max_tokens 1500, ~12500 for history, truncated oldest-first
  • Chat SSE events: {"type":"reasoning"|"content"|"queue"|"done"|"error"} — render reasoning_content in collapsible block (this is the demo)
  • Frontend: new chat component — collapsible reasoning block, queue position, streaming tokens
  • Prometheus metrics: atlas_chat_concurrent, atlas_chat_queue_depth, atlas_ratelimit_rejections_total, atlas_snapshot_age_seconds
  • Alertmanager rule: chat queue saturated > 5 min, snapshot age > 60s
  • pnpm audit + govulncheck in CI, fail build on high severity

REFACTOR

  • simplify pass on rate limiter + SSE chat handler

Verify

# load test at 20 concurrent clients
# GPU sequence usage never exceeds 6, no upstream 5xx, queue drains

Done — atlas v1 shipped

Success criteria (from PLAN-atlas.md):

  • Protected services redirect unauthenticated to Authentik (if still applicable post decision #3)
  • Fabricated stats gone, real numbers or removed
  • riotpiao.com is the only public hostname (I1 held)
  • Golden redaction test passes against live-cluster fixture
  • Load test: 20 concurrent clients, GPU usage ≤ 6, no 5xx, queue drains