# API Gateway Documentation ## Overview The homelab-frontend gateway is a production-ready reverse proxy for LLM model inference. It routes requests to multiple model upstreams based on configuration, with support for streaming, tool calling, and multiple API formats. **Base URL**: `https://api.riotpiao.com` **Deployment**: Client → nginx ingress → gateway → model upstreams --- ## Table of Contents 1. [Health Endpoints](#health-endpoints) 2. [GET /v1/models](#get-v1models) - List available models 3. [POST /v1/chat/completions](#post-v1chat-completions) - Chat with LLM 4. [POST /v1/embeddings](#post-v1embeddings) - Generate embeddings 5. [POST /v1/rerank](#post-v1rerank) - Rerank documents 6. [Error Handling](#error-handling) 7. [Examples](#examples) --- ## Health Endpoints ### GET /healthz Always returns 200 (liveness probe). **Response**: ```json {"status":"alive"} ``` **Status Code**: 200 --- ### GET /readyz Returns 200 when the gateway is ready (config loaded, upstreams available). **Response**: ```json {"status":"ready"} ``` **Status Code**: 200 (ready) or 503 (not ready) --- ## GET /v1/models List all configured models available for dispatch. **Method**: GET **Path**: `/v1/models` **Authentication**: None required **Query Parameters**: None **Request Headers**: ``` Accept: application/json ``` **Response Headers**: ``` Content-Type: application/json ``` **Response Schema**: ```json { "object": "list", "data": [ { "id": "model-name", "object": "model", "owned_by": "api.riotpiao.com", "created": 1700000000 } ] } ``` **Status Codes**: - `200` - OK **Example**: ```bash curl -s https://api.riotpiao.com/v1/models | jq . ``` **Response Example**: ```json { "object": "list", "data": [ { "id": "reasoning", "object": "model", "owned_by": "api.riotpiao.com", "created": 1700000000 }, { "id": "ornith:35b", "object": "model", "owned_by": "api.riotpiao.com", "created": 1700000000 }, { "id": "qwen2.5:3b-instruct", "object": "model", "owned_by": "api.riotpiao.com", "created": 1700000000 }, { "id": "nomic-ai/nomic-embed-text-v2-moe", "object": "model", "owned_by": "api.riotpiao.com", "created": 1700000000 }, { "id": "BAAI/bge-reranker-base", "object": "model", "owned_by": "api.riotpiao.com", "created": 1700000000 } ] } ``` --- ## POST /v1/chat/completions Chat with an LLM model. Routes to upstream based on the `model` field in the request body. **Method**: POST **Path**: `/v1/chat/completions` **Authentication**: None required (future: Bearer token) **Request Headers**: ``` Content-Type: application/json ``` **Request Body Schema**: ```json { "model": "string (required)", "messages": [ { "role": "string (user|assistant|system)", "content": "string|array (required)", "tool_calls": "array (optional, from assistant)" } ], "temperature": "number (optional, 0-2)", "top_p": "number (optional, 0-1)", "max_tokens": "integer (optional)", "stream": "boolean (optional, default: false)", "tools": [ { "type": "function", "function": { "name": "string", "description": "string", "parameters": "object" } } ] } ``` **Response Schema** (non-streaming): ```json { "id": "string", "object": "chat.completion", "created": "integer", "model": "string", "choices": [ { "index": "integer", "message": { "role": "assistant", "content": "string|null", "tool_calls": [ { "id": "string", "type": "function", "function": { "name": "string", "arguments": "string (JSON)" } } ] }, "finish_reason": "stop|tool_calls|length" } ], "usage": { "prompt_tokens": "integer", "completion_tokens": "integer", "total_tokens": "integer" } } ``` **Response Schema** (streaming): ``` data: {"id":"...", "object":"chat.completion.chunk", "choices":[...]} data: {"id":"...", "object":"chat.completion.chunk", "choices":[...]} ... data: [DONE] ``` **Status Codes**: - `200` - OK - `400` - Bad request (missing/invalid model, invalid JSON, etc.) - `500` - Internal server error (upstream issue) **Supported Models**: - `reasoning` - Reasoning model - `ornith:35b` - Ornith 35B model - `qwen2.5:3b-instruct` - Qwen 2.5 3B model **Examples**: ### Basic Chat ```bash curl -X POST https://api.riotpiao.com/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{ "model": "reasoning", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' ``` ### Chat with Tool Calling ```bash curl -X POST https://api.riotpiao.com/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{ "model": "reasoning", "messages": [ { "role": "user", "content": "What is the weather in San Francisco?" } ], "tools": [ { "type": "function", "function": { "name": "get_weather", "description": "Get the weather for a location", "parameters": { "type": "object", "properties": { "location": { "type": "string", "description": "City name" }, "unit": { "type": "string", "enum": ["celsius", "fahrenheit"] } }, "required": ["location"] } } } ] }' ``` ### Streaming Chat ```bash curl -N -X POST https://api.riotpiao.com/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{ "model": "reasoning", "messages": [ { "role": "user", "content": "Count from 1 to 3" } ], "stream": true }' ``` ### Multi-turn Conversation with Tool Results ```bash curl -X POST https://api.riotpiao.com/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{ "model": "reasoning", "messages": [ { "role": "user", "content": "What is the weather?" }, { "role": "assistant", "tool_calls": [ { "id": "call_123", "type": "function", "function": { "name": "get_weather", "arguments": "{\"location\": \"San Francisco\"}" } } ] }, { "role": "tool", "content": "{\"temperature\": 22, \"condition\": \"sunny\"}" } ], "tools": [ { "type": "function", "function": { "name": "get_weather", "description": "Get weather", "parameters": {} } } ] }' ``` --- ## POST /v1/embeddings Generate embeddings for text input. **Method**: POST **Path**: `/v1/embeddings` **Authentication**: None required **Request Headers**: ``` Content-Type: application/json ``` **Request Body Schema**: ```json { "model": "string (required)", "input": "string | array of strings (required)", "encoding_format": "float | base64 (optional)" } ``` **Response Schema**: ```json { "object": "list", "data": [ { "object": "embedding", "embedding": [0.1, 0.2, ...], "index": "integer" } ], "model": "string", "usage": { "prompt_tokens": "integer", "total_tokens": "integer" } } ``` **Status Codes**: - `200` - OK - `400` - Bad request (missing/invalid model, etc.) - `500` - Internal server error **Supported Models**: - `nomic-ai/nomic-embed-text-v2-moe` - Embedding model **Examples**: ### Single Input ```bash curl -X POST https://api.riotpiao.com/v1/embeddings \ -H 'Content-Type: application/json' \ -d '{ "model": "nomic-ai/nomic-embed-text-v2-moe", "input": "The quick brown fox" }' ``` ### Multiple Inputs ```bash curl -X POST https://api.riotpiao.com/v1/embeddings \ -H 'Content-Type: application/json' \ -d '{ "model": "nomic-ai/nomic-embed-text-v2-moe", "input": [ "Document 1 text", "Document 2 text", "Document 3 text" ] }' ``` --- ## POST /v1/rerank Rerank documents based on relevance to a query. **Method**: POST **Path**: `/v1/rerank` **Authentication**: None required **Request Headers**: ``` Content-Type: application/json ``` **Request Body Schema**: ```json { "model": "string (required)", "query": "string (required)", "texts": ["string"], "top_k": "integer (optional)", "return_documents": "boolean (optional)" } ``` **Response Schema**: ```json { "results": [ { "index": "integer", "score": "float (0-1)", "text": "string (optional)" } ] } ``` **Status Codes**: - `200` - OK - `400` - Bad request (missing/invalid model, etc.) - `500` - Internal server error **Supported Models**: - `BAAI/bge-reranker-base` - BGE reranker model **Note**: The gateway rewrites the path from `/v1/rerank` to `/rerank` on the upstream. **Examples**: ### Basic Reranking ```bash curl -X POST https://api.riotpiao.com/v1/rerank \ -H 'Content-Type: application/json' \ -d '{ "model": "BAAI/bge-reranker-base", "query": "What is machine learning?", "texts": [ "Machine learning is a type of artificial intelligence", "Dogs are animals", "Deep learning is a subset of machine learning", "Python is a programming language" ] }' ``` ### With Top-K Parameter ```bash curl -X POST https://api.riotpiao.com/v1/rerank \ -H 'Content-Type: application/json' \ -d '{ "model": "BAAI/bge-reranker-base", "query": "best practices", "texts": [ "Follow code style guidelines", "Write unit tests", "Use meaningful variable names", "Eat healthy food" ], "top_k": 2 }' ``` --- ## Error Handling ### Error Response Format The gateway returns RFC 9457 Problem Details for client errors (4xx): ```json { "type": "https://api.example.com/problems/error-type", "title": "Human-readable error title", "status": 400, "detail": "Detailed explanation of what went wrong", "valid_models": ["model1", "model2"] // Only for model-related errors } ``` ### Error Types #### Unknown Model Error **Status**: `400 Bad Request` **Trigger**: Model name not in registry **Response**: ```json { "type": "https://api.example.com/problems/unknown-model", "title": "Unknown Model", "status": 400, "detail": "Model \"gpt-4\" is not available. See valid_models for available options.", "valid_models": ["reasoning", "ornith:35b", "qwen2.5:3b-instruct", "nomic-ai/nomic-embed-text-v2-moe", "BAAI/bge-reranker-base"] } ``` **Example**: ```bash curl -X POST https://api.riotpiao.com/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{"model":"gpt-4","messages":[]}' ``` #### Missing Model Field **Status**: `400 Bad Request` **Trigger**: No `model` field in request body **Response**: ```json { "type": "https://api.example.com/problems/missing-model", "title": "Missing Model", "status": 400, "detail": "The 'model' field is required and must be a non-empty string", "valid_models": [...] } ``` **Example**: ```bash curl -X POST https://api.riotpiao.com/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{"messages":[]}' ``` #### Invalid JSON **Status**: `400 Bad Request` **Trigger**: Request body is not valid JSON **Response**: ```json { "type": "https://api.example.com/problems/invalid-request-body", "title": "Invalid Request Body", "status": 400, "detail": "request body is not valid JSON" } ``` **Example**: ```bash curl -X POST https://api.riotpiao.com/v1/chat/completions \ -H 'Content-Type: application/json' \ -d 'not json' ``` #### Upstream Error **Status**: `5xx` (from upstream) **Trigger**: Upstream service error **Response**: Forwarded from upstream (unmodified) --- ## Examples ### Test Script ```bash #!/bin/bash GATEWAY="https://api.riotpiao.com" echo "=== Testing Gateway API ===" echo "" # Test 1: Health checks echo "1. Health checks" curl -s "$GATEWAY/healthz" | jq . curl -s "$GATEWAY/readyz" | jq . echo "" # Test 2: List models echo "2. List models" curl -s "$GATEWAY/v1/models" | jq '.data[] | .id' echo "" # Test 3: Chat with reasoning model echo "3. Chat with reasoning model" curl -s -X POST "$GATEWAY/v1/chat/completions" \ -H 'Content-Type: application/json' \ -d '{ "model": "reasoning", "messages": [{"role": "user", "content": "What is 2+2?"}] }' | jq '.choices[0].message.content' echo "" # Test 4: Unknown model (should be 400) echo "4. Unknown model (should be 400)" curl -s -X POST "$GATEWAY/v1/chat/completions" \ -H 'Content-Type: application/json' \ -d '{"model":"gpt-4","messages":[]}' | jq '{status: .status, title: .title}' echo "" # Test 5: Embeddings echo "5. Embeddings" curl -s -X POST "$GATEWAY/v1/embeddings" \ -H 'Content-Type: application/json' \ -d '{ "model": "nomic-ai/nomic-embed-text-v2-moe", "input": "hello world" }' | jq '.data | length' echo "" # Test 6: Rerank echo "6. Rerank" curl -s -X POST "$GATEWAY/v1/rerank" \ -H 'Content-Type: application/json' \ -d '{ "model": "BAAI/bge-reranker-base", "query": "test", "texts": ["a", "b"] }' | jq '.results | length' echo "" # Test 7: Streaming echo "7. Streaming (showing first 5 chunks)" curl -s -N -X POST "$GATEWAY/v1/chat/completions" \ -H 'Content-Type: application/json' \ -d '{ "model": "reasoning", "messages": [{"role": "user", "content": "hi"}], "stream": true }' | head -10 echo "" echo "=== All tests completed ===" ``` ### Python Client Example ```python import requests import json GATEWAY = "https://api.riotpiao.com" # Get models response = requests.get(f"{GATEWAY}/v1/models") models = response.json() print(f"Available models: {[m['id'] for m in models['data']]}") # Chat completion response = requests.post( f"{GATEWAY}/v1/chat/completions", json={ "model": "reasoning", "messages": [ {"role": "user", "content": "What is machine learning?"} ] } ) message = response.json() print(f"Response: {message['choices'][0]['message']['content']}") # Chat with tools response = requests.post( f"{GATEWAY}/v1/chat/completions", json={ "model": "reasoning", "messages": [ {"role": "user", "content": "Get the weather"} ], "tools": [ { "type": "function", "function": { "name": "get_weather", "description": "Get weather", "parameters": {} } } ] } ) result = response.json() if "tool_calls" in result["choices"][0]["message"]: print(f"Tool calls: {result['choices'][0]['message']['tool_calls']}") # Streaming response = requests.post( f"{GATEWAY}/v1/chat/completions", json={ "model": "reasoning", "messages": [ {"role": "user", "content": "Count to 3"} ], "stream": True }, stream=True ) for line in response.iter_lines(): if line: print(line) # Embeddings response = requests.post( f"{GATEWAY}/v1/embeddings", json={ "model": "nomic-ai/nomic-embed-text-v2-moe", "input": "hello world" } ) embeddings = response.json() print(f"Embeddings: {embeddings['data'][0]['embedding'][:5]}") # Rerank response = requests.post( f"{GATEWAY}/v1/rerank", json={ "model": "BAAI/bge-reranker-base", "query": "ML", "texts": ["machine learning", "python", "deep learning"] } ) results = response.json() print(f"Rerank results: {results['results']}") ``` ### JavaScript/TypeScript Client Example ```typescript const GATEWAY = "https://api.riotpiao.com"; // Get models async function getModels() { const response = await fetch(`${GATEWAY}/v1/models`); const data = await response.json(); return data.data.map((m: any) => m.id); } // Chat completion async function chat(model: string, message: string) { const response = await fetch(`${GATEWAY}/v1/chat/completions`, { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify({ model, messages: [{ role: "user", content: message }], }), }); const data = await response.json(); return data.choices[0].message.content; } // Chat with streaming async function chatStream(model: string, message: string) { const response = await fetch(`${GATEWAY}/v1/chat/completions`, { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify({ model, messages: [{ role: "user", content: message }], stream: true, }), }); const reader = response.body!.getReader(); const decoder = new TextDecoder(); while (true) { const { done, value } = await reader.read(); if (done) break; const chunk = decoder.decode(value); const lines = chunk.split("\n"); for (const line of lines) { if (line.startsWith("data: ")) { const data = JSON.parse(line.slice(6)); if (data.choices[0].delta?.content) { console.log(data.choices[0].delta.content); } } } } } // Embeddings async function embed(model: string, input: string[]) { const response = await fetch(`${GATEWAY}/v1/embeddings`, { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify({ model, input }), }); const data = await response.json(); return data.data; } // Rerank async function rerank( model: string, query: string, texts: string[] ) { const response = await fetch(`${GATEWAY}/v1/rerank`, { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify({ model, query, texts }), }); const data = await response.json(); return data.results; } // Usage (async () => { const models = await getModels(); console.log("Models:", models); const response = await chat("reasoning", "What is AI?"); console.log("Response:", response); await chatStream("reasoning", "Count to 3"); const embeddings = await embed("nomic-ai/nomic-embed-text-v2-moe", [ "hello", ]); console.log("Embeddings:", embeddings); const rerankResults = await rerank("BAAI/bge-reranker-base", "ML", [ "machine learning", "python", ]); console.log("Rerank:", rerankResults); })(); ``` --- ## Rate Limiting Currently, no rate limiting is enforced. This will be added in Phase 4. --- ## Authentication Currently, no authentication is enforced. Bearer token support will be added in Phase 3. --- ## Timeouts Default timeouts per route: - **Connect**: 10s - **Read**: 1h (for streaming) - **Write**: 1h These are configured per model upstream. --- ## Body Size Limits - **Default**: 100MB - **Per-route**: Configurable Requests exceeding the limit return `413 Request Entity Too Large`. --- ## Support For issues or questions: - Check gateway logs: `kubectl -n api logs deployment/homelab-frontend` - Check health: `curl https://api.riotpiao.com/healthz` - Verify config: `curl https://api.riotpiao.com/v1/models`