feat(network): SSE optimization for local LLM streaming (#31 #32 #33)
CI / CI (pull_request) Successful in 3m11s

Addresses three critical network issues for LLM streaming performance:

**#33 Disable proxy buffering for SSE**
- Add X-Accel-Buffering: no header to response
- Tells nginx/Ingress to stream events immediately instead of buffering
- Paired with ResponseController.Flush() for unbuffered token delivery

**#32 HTTP/2 multiplexing for concurrent streams**
- Enable HTTP/2 in server config via http2.ConfigureServer()
- Increase MaxConnsPerHost from default (2) to 10
- ForceAttemptHTTP2 on outbound Transport for upstream connections
- Allows multiple concurrent LLM requests without blocking

**#31 TCP backpressure for streaming LLM responses**
- Set TCP_NODELAY on dialer to disable Nagle's algorithm
- Reduces latency by sending small packets immediately
- Critical for low TTFT (time-to-first-token) under load
- Upstream Transport respects backpressure when clients read slowly

**Tests added:**
- TestTCPBackpressure: Verifies TCP backpressure handling with slow client
- TestConcurrentSSEStreams: Confirms HTTP/2 multiplexing works correctly
- Both pass at 0.11s and 0.06s respectively

Fixes all three streaming performance issues in one coherent change.
This commit is contained in:
Admin Bot
2026-09-14 08:14:04 +09:00
parent 7de71180b3
commit cc9a32f53a
14 changed files with 957 additions and 13 deletions
+5
View File
@@ -6,6 +6,8 @@ namespace: api
resources:
- ci-rbac.yaml
- task-integration-test.yaml
- task-load-test.yaml
- pipeline-sse-optimization.yaml
generatorOptions:
disableNameSuffixHash: true
@@ -14,3 +16,6 @@ configMapGenerator:
- name: integration-test-script
files:
- scripts/integration-test.sh
- name: load-test-script
files:
- scripts/load-test.sh