# API Testing Guide Quick reference for testing the homelab-frontend gateway API. ## Setup ```bash # Set base URL export GATEWAY="https://api.riotpiao.com" # Or for local testing export GATEWAY="http://localhost:8080" ``` --- ## Quick Tests (Copy & Paste) ### 1. Health Checks ✅ ```bash # Liveness curl $GATEWAY/healthz | jq . # Readiness curl $GATEWAY/readyz | jq . ``` **Expected**: Both return `{"status":"..."}` with HTTP 200 --- ### 2. List Models ✅ ```bash curl $GATEWAY/v1/models | jq '.data[] | .id' ``` **Expected Output**: ``` "reasoning" "ornith:35b" "qwen2.5:3b-instruct" "nomic-ai/nomic-embed-text-v2-moe" "BAAI/bge-reranker-base" ``` --- ### 3. Chat - Basic ✅ ```bash curl -X POST $GATEWAY/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{ "model": "reasoning", "messages": [ {"role": "user", "content": "What is 2+2?"} ] }' | jq '.choices[0].message.content' ``` **Expected**: Model responds with an answer --- ### 4. Chat - Ornith Model ✅ ```bash curl -X POST $GATEWAY/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{ "model": "ornith:35b", "messages": [ {"role": "user", "content": "Hello"} ] }' | jq '.choices[0].message.content' ``` **Expected**: Routes to ornith model, returns response --- ### 5. Chat - Qwen Model ✅ ```bash curl -X POST $GATEWAY/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{ "model": "qwen2.5:3b-instruct", "messages": [ {"role": "user", "content": "Hi"} ] }' | jq '.choices[0].message.content' ``` **Expected**: Routes to qwen model, returns response --- ### 6. Chat - Unknown Model (Should Error) ❌→✅ ```bash curl -X POST $GATEWAY/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{ "model": "gpt-4-turbo", "messages": [] }' | jq '.' ``` **Expected**: HTTP 400 with problem+json: ```json { "type": "https://api.example.com/problems/unknown-model", "title": "Unknown Model", "status": 400, "detail": "Model \"gpt-4-turbo\" is not available. See valid_models for available options.", "valid_models": ["reasoning", "ornith:35b", ...] } ``` --- ### 7. Chat - Missing Model (Should Error) ❌→✅ ```bash curl -X POST $GATEWAY/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{ "messages": [{"role": "user", "content": "test"}] }' | jq '.' ``` **Expected**: HTTP 400 with problem+json (missing model) --- ### 8. Chat - Streaming ✅ ```bash curl -N -X POST $GATEWAY/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{ "model": "reasoning", "messages": [{"role": "user", "content": "count to 3"}], "stream": true }' | head -20 ``` **Expected**: - Multiple `data: {...}` lines (SSE chunks) - Final `data: [DONE]` - Chunks arrive incrementally (observable with `-N` flag) --- ### 9. Chat - Tool Calling ✅ ```bash curl -X POST $GATEWAY/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{ "model": "reasoning", "messages": [ {"role": "user", "content": "What is the weather in SF?"} ], "tools": [ { "type": "function", "function": { "name": "get_weather", "description": "Get weather for a location", "parameters": { "type": "object", "properties": { "location": {"type": "string"} }, "required": ["location"] } } } ] }' | jq '.choices[0].message.tool_calls' ``` **Expected**: Array of tool calls (if model decides to call them), or null (if not) --- ### 10. Embeddings ✅ ```bash curl -X POST $GATEWAY/v1/embeddings \ -H 'Content-Type: application/json' \ -d '{ "model": "nomic-ai/nomic-embed-text-v2-moe", "input": "hello world" }' | jq '.data | length' ``` **Expected**: `1` (one embedding vector) --- ### 11. Embeddings - Multiple ✅ ```bash curl -X POST $GATEWAY/v1/embeddings \ -H 'Content-Type: application/json' \ -d '{ "model": "nomic-ai/nomic-embed-text-v2-moe", "input": ["text 1", "text 2", "text 3"] }' | jq '.data | length' ``` **Expected**: `3` (three embedding vectors) --- ### 12. Embeddings - Unknown Model (Should Error) ❌→✅ ```bash curl -X POST $GATEWAY/v1/embeddings \ -H 'Content-Type: application/json' \ -d '{ "model": "unknown-embed", "input": "test" }' | jq '.status' ``` **Expected**: `400` (client error) --- ### 13. Rerank ✅ ```bash curl -X POST $GATEWAY/v1/rerank \ -H 'Content-Type: application/json' \ -d '{ "model": "BAAI/bge-reranker-base", "query": "machine learning", "texts": [ "Machine learning is AI", "Python is a language", "Deep learning is ML" ] }' | jq '.results' ``` **Expected**: Array of ranked results with scores: ```json [ {"index": 0, "score": 0.95}, {"index": 2, "score": 0.85}, {"index": 1, "score": 0.15} ] ``` --- ### 14. Rerank - Unknown Model (Should Error) ❌→✅ ```bash curl -X POST $GATEWAY/v1/rerank \ -H 'Content-Type: application/json' \ -d '{ "model": "unknown-rerank", "query": "test", "texts": ["a"] }' | jq '.status' ``` **Expected**: `400` (client error) --- ### 15. Invalid JSON (Should Error) ❌→✅ ```bash curl -X POST $GATEWAY/v1/chat/completions \ -H 'Content-Type: application/json' \ -d 'not json' | jq '.title' ``` **Expected**: `"Invalid Request Body"` (HTTP 400) --- ## Testing Checklist Complete this checklist to verify all endpoints: ### Health Endpoints - [ ] GET /healthz → 200, `{"status":"alive"}` - [ ] GET /readyz → 200, `{"status":"ready"}` ### Model Discovery - [ ] GET /v1/models → 200, returns all 5 models - [ ] All advertised models can be called (none 400) ### Chat Completions - [ ] POST /v1/chat/completions (reasoning) → 200, response - [ ] POST /v1/chat/completions (ornith:35b) → 200, response - [ ] POST /v1/chat/completions (qwen2.5:3b-instruct) → 200, response - [ ] POST /v1/chat/completions (unknown model) → 400, problem+json - [ ] POST /v1/chat/completions (missing model) → 400, problem+json - [ ] POST /v1/chat/completions (invalid JSON) → 400, problem+json - [ ] POST /v1/chat/completions (streaming) → 200, SSE chunks - [ ] POST /v1/chat/completions (with tools) → 200, tool_calls present/absent ### Embeddings - [ ] POST /v1/embeddings (single input) → 200, embedding - [ ] POST /v1/embeddings (multiple inputs) → 200, embeddings array - [ ] POST /v1/embeddings (unknown model) → 400, problem+json ### Reranking - [ ] POST /v1/rerank → 200, ranked results - [ ] POST /v1/rerank (unknown model) → 400, problem+json - [ ] Verify path is rewritten to /rerank on upstream ### Error Handling - [ ] Unknown model lists valid_models - [ ] Error responses are problem+json - [ ] No 5xx for client errors (validation errors) - [ ] Upstream errors pass through ### Streaming - [ ] Chunks arrive incrementally - [ ] Final `[DONE]` sentinel present - [ ] Works for chat completions ### Tool Calling - [ ] Tool definitions forward to upstream - [ ] Tool calls in response - [ ] Multi-turn with tool results - [ ] Parallel tool calls - [ ] Complex nested arguments preserved --- ## Troubleshooting ### 404 Responses **Symptom**: All endpoints return `"not found"` **Cause**: ConfigMap with models/routes not deployed **Solution**: ```bash kubectl -n api create configmap homelab-frontend-config \ --from-file=config.yaml=k8s/configmap.yaml kubectl -n api rollout restart deployment/homelab-frontend ``` --- ### 503 (Not Ready) **Symptom**: `/readyz` returns 503 **Cause**: Configuration not loaded or JWKS fetch failed **Solution**: ```bash # Check logs kubectl -n api logs deployment/homelab-frontend # Check config kubectl -n api get configmap homelab-frontend-config ``` --- ### Connection Refused **Symptom**: `Connection refused` or `Temporary failure in name resolution` **Cause**: - Gateway not running - Wrong URL/hostname - Network issue **Solution**: ```bash # Verify gateway is running kubectl -n api get pods -l app=homelab-frontend # Check service kubectl -n api get svc homelab-frontend # Verify ingress kubectl -n api get ingress api ``` --- ### Upstream Connection Errors **Symptom**: `502 Bad Gateway` or `connection refused to upstream` **Cause**: Model upstream service not reachable **Solution**: ```bash # Check upstreams are running kubectl -n llm-serving get pods # Verify addresses in ConfigMap kubectl -n api get configmap homelab-frontend-config -o yaml # Test connectivity from gateway pod kubectl -n api exec deployment/homelab-frontend -- \ curl -s reasoning-predictor.llm-serving:80/healthz ``` --- ### Streaming Doesn't Work **Symptom**: Chunks arrive all at once (buffered) instead of incrementally **Cause**: nginx buffering or client not using `-N` flag **Solution**: ```bash # Use -N flag curl -N https://api.riotpiao.com/v1/chat/completions ... # Verify nginx has buffering disabled # Should have: proxy-buffering: off in Ingress annotations ``` --- ## Performance Testing ### Load Test (Simple) ```bash # Send 10 requests in parallel for i in {1..10}; do curl -X POST $GATEWAY/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{"model":"reasoning","messages":[{"role":"user","content":"Hi"}]}' & done wait echo "Completed 10 requests" ``` ### Concurrency Test ```bash # Use Apache Bench (if installed) ab -n 100 -c 10 \ -p request.json \ -T application/json \ $GATEWAY/v1/chat/completions # Create request.json: # {"model":"reasoning","messages":[{"role":"user","content":"test"}]} ``` ### Latency Test ```bash # Measure response time curl -w "\nTotal time: %{time_total}s\n" \ -X POST $GATEWAY/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{ "model": "reasoning", "messages": [{"role": "user", "content": "What is AI?"}] }' > /dev/null ``` --- ## Integration Testing ### Test with Python ```bash pip install requests cat > test_api.py << 'EOF' import requests import json gateway = "https://api.riotpiao.com" # Test health r = requests.get(f"{gateway}/healthz") assert r.status_code == 200 print("✓ Health check passed") # Test models r = requests.get(f"{gateway}/v1/models") assert r.status_code == 200 models = [m['id'] for m in r.json()['data']] print(f"✓ Models: {models}") # Test chat r = requests.post( f"{gateway}/v1/chat/completions", json={"model": "reasoning", "messages": [{"role": "user", "content": "Hi"}]} ) assert r.status_code == 200 print("✓ Chat works") # Test unknown model error r = requests.post( f"{gateway}/v1/chat/completions", json={"model": "gpt-4", "messages": []} ) assert r.status_code == 400 assert "unknown" in r.json()['detail'].lower() print("✓ Unknown model error correct") # Test embeddings r = requests.post( f"{gateway}/v1/embeddings", json={"model": "nomic-ai/nomic-embed-text-v2-moe", "input": "test"} ) assert r.status_code == 200 print("✓ Embeddings work") # Test rerank r = requests.post( f"{gateway}/v1/rerank", json={"model": "BAAI/bge-reranker-base", "query": "test", "texts": ["a", "b"]} ) assert r.status_code == 200 print("✓ Reranking works") print("\n✅ All tests passed!") EOF python test_api.py ``` --- ## Summary | Category | Tests | Expected | |----------|-------|----------| | Health | 2 | ✅ Both 200 | | Models | 1 | ✅ 5 models listed | | Chat | 8 | ✅ 6 success + 2 error | | Embeddings | 3 | ✅ 2 success + 1 error | | Rerank | 2 | ✅ 1 success + 1 error | | Streaming | 1 | ✅ Incremental chunks | | Tools | 1 | ✅ Tool calls present | | **TOTAL** | **18+** | **✅ ALL PASS** | Once all tests pass, the gateway is production-ready! 🚀