docs: comprehensive api and testing documentation

- API.md: Full REST API documentation with examples
  * All endpoints (health, models, chat, embeddings, rerank)
  * Request/response schemas
  * Error handling (RFC 9457 problem+json)
  * Examples in bash, Python, TypeScript

- TESTING_GUIDE.md: Quick reference testing guide
  * 15 copy-paste test commands
  * Complete testing checklist
  * Troubleshooting guide
  * Performance testing examples
  * Integration test scripts

Ready for deployment verification and integration testing.
This commit is contained in:
Story Crater Bot
2026-08-19 23:55:22 -07:00
parent a8dfd5b2f0
commit c8c656046a
2 changed files with 1494 additions and 0 deletions
+560
View File
@@ -0,0 +1,560 @@
# API Testing Guide
Quick reference for testing the homelab-frontend gateway API.
## Setup
```bash
# Set base URL
export GATEWAY="https://api.riotpiao.com"
# Or for local testing
export GATEWAY="http://localhost:8080"
```
---
## Quick Tests (Copy & Paste)
### 1. Health Checks ✅
```bash
# Liveness
curl $GATEWAY/healthz | jq .
# Readiness
curl $GATEWAY/readyz | jq .
```
**Expected**: Both return `{"status":"..."}` with HTTP 200
---
### 2. List Models ✅
```bash
curl $GATEWAY/v1/models | jq '.data[] | .id'
```
**Expected Output**:
```
"reasoning"
"ornith:35b"
"qwen2.5:3b-instruct"
"nomic-ai/nomic-embed-text-v2-moe"
"BAAI/bge-reranker-base"
```
---
### 3. Chat - Basic ✅
```bash
curl -X POST $GATEWAY/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "reasoning",
"messages": [
{"role": "user", "content": "What is 2+2?"}
]
}' | jq '.choices[0].message.content'
```
**Expected**: Model responds with an answer
---
### 4. Chat - Ornith Model ✅
```bash
curl -X POST $GATEWAY/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "ornith:35b",
"messages": [
{"role": "user", "content": "Hello"}
]
}' | jq '.choices[0].message.content'
```
**Expected**: Routes to ornith model, returns response
---
### 5. Chat - Qwen Model ✅
```bash
curl -X POST $GATEWAY/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "qwen2.5:3b-instruct",
"messages": [
{"role": "user", "content": "Hi"}
]
}' | jq '.choices[0].message.content'
```
**Expected**: Routes to qwen model, returns response
---
### 6. Chat - Unknown Model (Should Error) ❌→✅
```bash
curl -X POST $GATEWAY/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "gpt-4-turbo",
"messages": []
}' | jq '.'
```
**Expected**: HTTP 400 with problem+json:
```json
{
"type": "https://api.example.com/problems/unknown-model",
"title": "Unknown Model",
"status": 400,
"detail": "Model \"gpt-4-turbo\" is not available. See valid_models for available options.",
"valid_models": ["reasoning", "ornith:35b", ...]
}
```
---
### 7. Chat - Missing Model (Should Error) ❌→✅
```bash
curl -X POST $GATEWAY/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"messages": [{"role": "user", "content": "test"}]
}' | jq '.'
```
**Expected**: HTTP 400 with problem+json (missing model)
---
### 8. Chat - Streaming ✅
```bash
curl -N -X POST $GATEWAY/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "reasoning",
"messages": [{"role": "user", "content": "count to 3"}],
"stream": true
}' | head -20
```
**Expected**:
- Multiple `data: {...}` lines (SSE chunks)
- Final `data: [DONE]`
- Chunks arrive incrementally (observable with `-N` flag)
---
### 9. Chat - Tool Calling ✅
```bash
curl -X POST $GATEWAY/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "reasoning",
"messages": [
{"role": "user", "content": "What is the weather in SF?"}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get weather for a location",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string"}
},
"required": ["location"]
}
}
}
]
}' | jq '.choices[0].message.tool_calls'
```
**Expected**: Array of tool calls (if model decides to call them), or null (if not)
---
### 10. Embeddings ✅
```bash
curl -X POST $GATEWAY/v1/embeddings \
-H 'Content-Type: application/json' \
-d '{
"model": "nomic-ai/nomic-embed-text-v2-moe",
"input": "hello world"
}' | jq '.data | length'
```
**Expected**: `1` (one embedding vector)
---
### 11. Embeddings - Multiple ✅
```bash
curl -X POST $GATEWAY/v1/embeddings \
-H 'Content-Type: application/json' \
-d '{
"model": "nomic-ai/nomic-embed-text-v2-moe",
"input": ["text 1", "text 2", "text 3"]
}' | jq '.data | length'
```
**Expected**: `3` (three embedding vectors)
---
### 12. Embeddings - Unknown Model (Should Error) ❌→✅
```bash
curl -X POST $GATEWAY/v1/embeddings \
-H 'Content-Type: application/json' \
-d '{
"model": "unknown-embed",
"input": "test"
}' | jq '.status'
```
**Expected**: `400` (client error)
---
### 13. Rerank ✅
```bash
curl -X POST $GATEWAY/v1/rerank \
-H 'Content-Type: application/json' \
-d '{
"model": "BAAI/bge-reranker-base",
"query": "machine learning",
"texts": [
"Machine learning is AI",
"Python is a language",
"Deep learning is ML"
]
}' | jq '.results'
```
**Expected**: Array of ranked results with scores:
```json
[
{"index": 0, "score": 0.95},
{"index": 2, "score": 0.85},
{"index": 1, "score": 0.15}
]
```
---
### 14. Rerank - Unknown Model (Should Error) ❌→✅
```bash
curl -X POST $GATEWAY/v1/rerank \
-H 'Content-Type: application/json' \
-d '{
"model": "unknown-rerank",
"query": "test",
"texts": ["a"]
}' | jq '.status'
```
**Expected**: `400` (client error)
---
### 15. Invalid JSON (Should Error) ❌→✅
```bash
curl -X POST $GATEWAY/v1/chat/completions \
-H 'Content-Type: application/json' \
-d 'not json' | jq '.title'
```
**Expected**: `"Invalid Request Body"` (HTTP 400)
---
## Testing Checklist
Complete this checklist to verify all endpoints:
### Health Endpoints
- [ ] GET /healthz → 200, `{"status":"alive"}`
- [ ] GET /readyz → 200, `{"status":"ready"}`
### Model Discovery
- [ ] GET /v1/models → 200, returns all 5 models
- [ ] All advertised models can be called (none 400)
### Chat Completions
- [ ] POST /v1/chat/completions (reasoning) → 200, response
- [ ] POST /v1/chat/completions (ornith:35b) → 200, response
- [ ] POST /v1/chat/completions (qwen2.5:3b-instruct) → 200, response
- [ ] POST /v1/chat/completions (unknown model) → 400, problem+json
- [ ] POST /v1/chat/completions (missing model) → 400, problem+json
- [ ] POST /v1/chat/completions (invalid JSON) → 400, problem+json
- [ ] POST /v1/chat/completions (streaming) → 200, SSE chunks
- [ ] POST /v1/chat/completions (with tools) → 200, tool_calls present/absent
### Embeddings
- [ ] POST /v1/embeddings (single input) → 200, embedding
- [ ] POST /v1/embeddings (multiple inputs) → 200, embeddings array
- [ ] POST /v1/embeddings (unknown model) → 400, problem+json
### Reranking
- [ ] POST /v1/rerank → 200, ranked results
- [ ] POST /v1/rerank (unknown model) → 400, problem+json
- [ ] Verify path is rewritten to /rerank on upstream
### Error Handling
- [ ] Unknown model lists valid_models
- [ ] Error responses are problem+json
- [ ] No 5xx for client errors (validation errors)
- [ ] Upstream errors pass through
### Streaming
- [ ] Chunks arrive incrementally
- [ ] Final `[DONE]` sentinel present
- [ ] Works for chat completions
### Tool Calling
- [ ] Tool definitions forward to upstream
- [ ] Tool calls in response
- [ ] Multi-turn with tool results
- [ ] Parallel tool calls
- [ ] Complex nested arguments preserved
---
## Troubleshooting
### 404 Responses
**Symptom**: All endpoints return `"not found"`
**Cause**: ConfigMap with models/routes not deployed
**Solution**:
```bash
kubectl -n api create configmap homelab-frontend-config \
--from-file=config.yaml=k8s/configmap.yaml
kubectl -n api rollout restart deployment/homelab-frontend
```
---
### 503 (Not Ready)
**Symptom**: `/readyz` returns 503
**Cause**: Configuration not loaded or JWKS fetch failed
**Solution**:
```bash
# Check logs
kubectl -n api logs deployment/homelab-frontend
# Check config
kubectl -n api get configmap homelab-frontend-config
```
---
### Connection Refused
**Symptom**: `Connection refused` or `Temporary failure in name resolution`
**Cause**:
- Gateway not running
- Wrong URL/hostname
- Network issue
**Solution**:
```bash
# Verify gateway is running
kubectl -n api get pods -l app=homelab-frontend
# Check service
kubectl -n api get svc homelab-frontend
# Verify ingress
kubectl -n api get ingress api
```
---
### Upstream Connection Errors
**Symptom**: `502 Bad Gateway` or `connection refused to upstream`
**Cause**: Model upstream service not reachable
**Solution**:
```bash
# Check upstreams are running
kubectl -n llm-serving get pods
# Verify addresses in ConfigMap
kubectl -n api get configmap homelab-frontend-config -o yaml
# Test connectivity from gateway pod
kubectl -n api exec deployment/homelab-frontend -- \
curl -s reasoning-predictor.llm-serving:80/healthz
```
---
### Streaming Doesn't Work
**Symptom**: Chunks arrive all at once (buffered) instead of incrementally
**Cause**: nginx buffering or client not using `-N` flag
**Solution**:
```bash
# Use -N flag
curl -N https://api.riotpiao.com/v1/chat/completions ...
# Verify nginx has buffering disabled
# Should have: proxy-buffering: off in Ingress annotations
```
---
## Performance Testing
### Load Test (Simple)
```bash
# Send 10 requests in parallel
for i in {1..10}; do
curl -X POST $GATEWAY/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"reasoning","messages":[{"role":"user","content":"Hi"}]}' &
done
wait
echo "Completed 10 requests"
```
### Concurrency Test
```bash
# Use Apache Bench (if installed)
ab -n 100 -c 10 \
-p request.json \
-T application/json \
$GATEWAY/v1/chat/completions
# Create request.json:
# {"model":"reasoning","messages":[{"role":"user","content":"test"}]}
```
### Latency Test
```bash
# Measure response time
curl -w "\nTotal time: %{time_total}s\n" \
-X POST $GATEWAY/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "reasoning",
"messages": [{"role": "user", "content": "What is AI?"}]
}' > /dev/null
```
---
## Integration Testing
### Test with Python
```bash
pip install requests
cat > test_api.py << 'EOF'
import requests
import json
gateway = "https://api.riotpiao.com"
# Test health
r = requests.get(f"{gateway}/healthz")
assert r.status_code == 200
print("✓ Health check passed")
# Test models
r = requests.get(f"{gateway}/v1/models")
assert r.status_code == 200
models = [m['id'] for m in r.json()['data']]
print(f"✓ Models: {models}")
# Test chat
r = requests.post(
f"{gateway}/v1/chat/completions",
json={"model": "reasoning", "messages": [{"role": "user", "content": "Hi"}]}
)
assert r.status_code == 200
print("✓ Chat works")
# Test unknown model error
r = requests.post(
f"{gateway}/v1/chat/completions",
json={"model": "gpt-4", "messages": []}
)
assert r.status_code == 400
assert "unknown" in r.json()['detail'].lower()
print("✓ Unknown model error correct")
# Test embeddings
r = requests.post(
f"{gateway}/v1/embeddings",
json={"model": "nomic-ai/nomic-embed-text-v2-moe", "input": "test"}
)
assert r.status_code == 200
print("✓ Embeddings work")
# Test rerank
r = requests.post(
f"{gateway}/v1/rerank",
json={"model": "BAAI/bge-reranker-base", "query": "test", "texts": ["a", "b"]}
)
assert r.status_code == 200
print("✓ Reranking works")
print("\n✅ All tests passed!")
EOF
python test_api.py
```
---
## Summary
| Category | Tests | Expected |
|----------|-------|----------|
| Health | 2 | ✅ Both 200 |
| Models | 1 | ✅ 5 models listed |
| Chat | 8 | ✅ 6 success + 2 error |
| Embeddings | 3 | ✅ 2 success + 1 error |
| Rerank | 2 | ✅ 1 success + 1 error |
| Streaming | 1 | ✅ Incremental chunks |
| Tools | 1 | ✅ Tool calls present |
| **TOTAL** | **18+** | **✅ ALL PASS** |
Once all tests pass, the gateway is production-ready! 🚀