- API.md: Full REST API documentation with examples * All endpoints (health, models, chat, embeddings, rerank) * Request/response schemas * Error handling (RFC 9457 problem+json) * Examples in bash, Python, TypeScript - TESTING_GUIDE.md: Quick reference testing guide * 15 copy-paste test commands * Complete testing checklist * Troubleshooting guide * Performance testing examples * Integration test scripts Ready for deployment verification and integration testing.
12 KiB
API Testing Guide
Quick reference for testing the homelab-frontend gateway API.
Setup
# Set base URL
export GATEWAY="https://api.riotpiao.com"
# Or for local testing
export GATEWAY="http://localhost:8080"
Quick Tests (Copy & Paste)
1. Health Checks ✅
# Liveness
curl $GATEWAY/healthz | jq .
# Readiness
curl $GATEWAY/readyz | jq .
Expected: Both return {"status":"..."} with HTTP 200
2. List Models ✅
curl $GATEWAY/v1/models | jq '.data[] | .id'
Expected Output:
"reasoning"
"ornith:35b"
"qwen2.5:3b-instruct"
"nomic-ai/nomic-embed-text-v2-moe"
"BAAI/bge-reranker-base"
3. Chat - Basic ✅
curl -X POST $GATEWAY/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "reasoning",
"messages": [
{"role": "user", "content": "What is 2+2?"}
]
}' | jq '.choices[0].message.content'
Expected: Model responds with an answer
4. Chat - Ornith Model ✅
curl -X POST $GATEWAY/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "ornith:35b",
"messages": [
{"role": "user", "content": "Hello"}
]
}' | jq '.choices[0].message.content'
Expected: Routes to ornith model, returns response
5. Chat - Qwen Model ✅
curl -X POST $GATEWAY/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "qwen2.5:3b-instruct",
"messages": [
{"role": "user", "content": "Hi"}
]
}' | jq '.choices[0].message.content'
Expected: Routes to qwen model, returns response
6. Chat - Unknown Model (Should Error) ❌→✅
curl -X POST $GATEWAY/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "gpt-4-turbo",
"messages": []
}' | jq '.'
Expected: HTTP 400 with problem+json:
{
"type": "https://api.example.com/problems/unknown-model",
"title": "Unknown Model",
"status": 400,
"detail": "Model \"gpt-4-turbo\" is not available. See valid_models for available options.",
"valid_models": ["reasoning", "ornith:35b", ...]
}
7. Chat - Missing Model (Should Error) ❌→✅
curl -X POST $GATEWAY/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"messages": [{"role": "user", "content": "test"}]
}' | jq '.'
Expected: HTTP 400 with problem+json (missing model)
8. Chat - Streaming ✅
curl -N -X POST $GATEWAY/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "reasoning",
"messages": [{"role": "user", "content": "count to 3"}],
"stream": true
}' | head -20
Expected:
- Multiple
data: {...}lines (SSE chunks) - Final
data: [DONE] - Chunks arrive incrementally (observable with
-Nflag)
9. Chat - Tool Calling ✅
curl -X POST $GATEWAY/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "reasoning",
"messages": [
{"role": "user", "content": "What is the weather in SF?"}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get weather for a location",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string"}
},
"required": ["location"]
}
}
}
]
}' | jq '.choices[0].message.tool_calls'
Expected: Array of tool calls (if model decides to call them), or null (if not)
10. Embeddings ✅
curl -X POST $GATEWAY/v1/embeddings \
-H 'Content-Type: application/json' \
-d '{
"model": "nomic-ai/nomic-embed-text-v2-moe",
"input": "hello world"
}' | jq '.data | length'
Expected: 1 (one embedding vector)
11. Embeddings - Multiple ✅
curl -X POST $GATEWAY/v1/embeddings \
-H 'Content-Type: application/json' \
-d '{
"model": "nomic-ai/nomic-embed-text-v2-moe",
"input": ["text 1", "text 2", "text 3"]
}' | jq '.data | length'
Expected: 3 (three embedding vectors)
12. Embeddings - Unknown Model (Should Error) ❌→✅
curl -X POST $GATEWAY/v1/embeddings \
-H 'Content-Type: application/json' \
-d '{
"model": "unknown-embed",
"input": "test"
}' | jq '.status'
Expected: 400 (client error)
13. Rerank ✅
curl -X POST $GATEWAY/v1/rerank \
-H 'Content-Type: application/json' \
-d '{
"model": "BAAI/bge-reranker-base",
"query": "machine learning",
"texts": [
"Machine learning is AI",
"Python is a language",
"Deep learning is ML"
]
}' | jq '.results'
Expected: Array of ranked results with scores:
[
{"index": 0, "score": 0.95},
{"index": 2, "score": 0.85},
{"index": 1, "score": 0.15}
]
14. Rerank - Unknown Model (Should Error) ❌→✅
curl -X POST $GATEWAY/v1/rerank \
-H 'Content-Type: application/json' \
-d '{
"model": "unknown-rerank",
"query": "test",
"texts": ["a"]
}' | jq '.status'
Expected: 400 (client error)
15. Invalid JSON (Should Error) ❌→✅
curl -X POST $GATEWAY/v1/chat/completions \
-H 'Content-Type: application/json' \
-d 'not json' | jq '.title'
Expected: "Invalid Request Body" (HTTP 400)
Testing Checklist
Complete this checklist to verify all endpoints:
Health Endpoints
- GET /healthz → 200,
{"status":"alive"} - GET /readyz → 200,
{"status":"ready"}
Model Discovery
- GET /v1/models → 200, returns all 5 models
- All advertised models can be called (none 400)
Chat Completions
- POST /v1/chat/completions (reasoning) → 200, response
- POST /v1/chat/completions (ornith:35b) → 200, response
- POST /v1/chat/completions (qwen2.5:3b-instruct) → 200, response
- POST /v1/chat/completions (unknown model) → 400, problem+json
- POST /v1/chat/completions (missing model) → 400, problem+json
- POST /v1/chat/completions (invalid JSON) → 400, problem+json
- POST /v1/chat/completions (streaming) → 200, SSE chunks
- POST /v1/chat/completions (with tools) → 200, tool_calls present/absent
Embeddings
- POST /v1/embeddings (single input) → 200, embedding
- POST /v1/embeddings (multiple inputs) → 200, embeddings array
- POST /v1/embeddings (unknown model) → 400, problem+json
Reranking
- POST /v1/rerank → 200, ranked results
- POST /v1/rerank (unknown model) → 400, problem+json
- Verify path is rewritten to /rerank on upstream
Error Handling
- Unknown model lists valid_models
- Error responses are problem+json
- No 5xx for client errors (validation errors)
- Upstream errors pass through
Streaming
- Chunks arrive incrementally
- Final
[DONE]sentinel present - Works for chat completions
Tool Calling
- Tool definitions forward to upstream
- Tool calls in response
- Multi-turn with tool results
- Parallel tool calls
- Complex nested arguments preserved
Troubleshooting
404 Responses
Symptom: All endpoints return "not found"
Cause: ConfigMap with models/routes not deployed
Solution:
kubectl -n api create configmap homelab-frontend-config \
--from-file=config.yaml=k8s/configmap.yaml
kubectl -n api rollout restart deployment/homelab-frontend
503 (Not Ready)
Symptom: /readyz returns 503
Cause: Configuration not loaded or JWKS fetch failed
Solution:
# Check logs
kubectl -n api logs deployment/homelab-frontend
# Check config
kubectl -n api get configmap homelab-frontend-config
Connection Refused
Symptom: Connection refused or Temporary failure in name resolution
Cause:
- Gateway not running
- Wrong URL/hostname
- Network issue
Solution:
# Verify gateway is running
kubectl -n api get pods -l app=homelab-frontend
# Check service
kubectl -n api get svc homelab-frontend
# Verify ingress
kubectl -n api get ingress api
Upstream Connection Errors
Symptom: 502 Bad Gateway or connection refused to upstream
Cause: Model upstream service not reachable
Solution:
# Check upstreams are running
kubectl -n llm-serving get pods
# Verify addresses in ConfigMap
kubectl -n api get configmap homelab-frontend-config -o yaml
# Test connectivity from gateway pod
kubectl -n api exec deployment/homelab-frontend -- \
curl -s reasoning-predictor.llm-serving:80/healthz
Streaming Doesn't Work
Symptom: Chunks arrive all at once (buffered) instead of incrementally
Cause: nginx buffering or client not using -N flag
Solution:
# Use -N flag
curl -N https://api.riotpiao.com/v1/chat/completions ...
# Verify nginx has buffering disabled
# Should have: proxy-buffering: off in Ingress annotations
Performance Testing
Load Test (Simple)
# Send 10 requests in parallel
for i in {1..10}; do
curl -X POST $GATEWAY/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"reasoning","messages":[{"role":"user","content":"Hi"}]}' &
done
wait
echo "Completed 10 requests"
Concurrency Test
# Use Apache Bench (if installed)
ab -n 100 -c 10 \
-p request.json \
-T application/json \
$GATEWAY/v1/chat/completions
# Create request.json:
# {"model":"reasoning","messages":[{"role":"user","content":"test"}]}
Latency Test
# Measure response time
curl -w "\nTotal time: %{time_total}s\n" \
-X POST $GATEWAY/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "reasoning",
"messages": [{"role": "user", "content": "What is AI?"}]
}' > /dev/null
Integration Testing
Test with Python
pip install requests
cat > test_api.py << 'EOF'
import requests
import json
gateway = "https://api.riotpiao.com"
# Test health
r = requests.get(f"{gateway}/healthz")
assert r.status_code == 200
print("✓ Health check passed")
# Test models
r = requests.get(f"{gateway}/v1/models")
assert r.status_code == 200
models = [m['id'] for m in r.json()['data']]
print(f"✓ Models: {models}")
# Test chat
r = requests.post(
f"{gateway}/v1/chat/completions",
json={"model": "reasoning", "messages": [{"role": "user", "content": "Hi"}]}
)
assert r.status_code == 200
print("✓ Chat works")
# Test unknown model error
r = requests.post(
f"{gateway}/v1/chat/completions",
json={"model": "gpt-4", "messages": []}
)
assert r.status_code == 400
assert "unknown" in r.json()['detail'].lower()
print("✓ Unknown model error correct")
# Test embeddings
r = requests.post(
f"{gateway}/v1/embeddings",
json={"model": "nomic-ai/nomic-embed-text-v2-moe", "input": "test"}
)
assert r.status_code == 200
print("✓ Embeddings work")
# Test rerank
r = requests.post(
f"{gateway}/v1/rerank",
json={"model": "BAAI/bge-reranker-base", "query": "test", "texts": ["a", "b"]}
)
assert r.status_code == 200
print("✓ Reranking works")
print("\n✅ All tests passed!")
EOF
python test_api.py
Summary
| Category | Tests | Expected |
|---|---|---|
| Health | 2 | ✅ Both 200 |
| Models | 1 | ✅ 5 models listed |
| Chat | 8 | ✅ 6 success + 2 error |
| Embeddings | 3 | ✅ 2 success + 1 error |
| Rerank | 2 | ✅ 1 success + 1 error |
| Streaming | 1 | ✅ Incremental chunks |
| Tools | 1 | ✅ Tool calls present |
| TOTAL | 18+ | ✅ ALL PASS |
Once all tests pass, the gateway is production-ready! 🚀