docs: comprehensive api and testing documentation
- API.md: Full REST API documentation with examples * All endpoints (health, models, chat, embeddings, rerank) * Request/response schemas * Error handling (RFC 9457 problem+json) * Examples in bash, Python, TypeScript - TESTING_GUIDE.md: Quick reference testing guide * 15 copy-paste test commands * Complete testing checklist * Troubleshooting guide * Performance testing examples * Integration test scripts Ready for deployment verification and integration testing.
This commit is contained in:
@@ -0,0 +1,934 @@
|
||||
# API Gateway Documentation
|
||||
|
||||
## Overview
|
||||
|
||||
The homelab-frontend gateway is a production-ready reverse proxy for LLM model inference. It routes requests to multiple model upstreams based on configuration, with support for streaming, tool calling, and multiple API formats.
|
||||
|
||||
**Base URL**: `https://api.riotpiao.com`
|
||||
|
||||
**Deployment**: Client → nginx ingress → gateway → model upstreams
|
||||
|
||||
---
|
||||
|
||||
## Table of Contents
|
||||
|
||||
1. [Health Endpoints](#health-endpoints)
|
||||
2. [GET /v1/models](#get-v1models) - List available models
|
||||
3. [POST /v1/chat/completions](#post-v1chat-completions) - Chat with LLM
|
||||
4. [POST /v1/embeddings](#post-v1embeddings) - Generate embeddings
|
||||
5. [POST /v1/rerank](#post-v1rerank) - Rerank documents
|
||||
6. [Error Handling](#error-handling)
|
||||
7. [Examples](#examples)
|
||||
|
||||
---
|
||||
|
||||
## Health Endpoints
|
||||
|
||||
### GET /healthz
|
||||
|
||||
Always returns 200 (liveness probe).
|
||||
|
||||
**Response**:
|
||||
```json
|
||||
{"status":"alive"}
|
||||
```
|
||||
|
||||
**Status Code**: 200
|
||||
|
||||
---
|
||||
|
||||
### GET /readyz
|
||||
|
||||
Returns 200 when the gateway is ready (config loaded, upstreams available).
|
||||
|
||||
**Response**:
|
||||
```json
|
||||
{"status":"ready"}
|
||||
```
|
||||
|
||||
**Status Code**: 200 (ready) or 503 (not ready)
|
||||
|
||||
---
|
||||
|
||||
## GET /v1/models
|
||||
|
||||
List all configured models available for dispatch.
|
||||
|
||||
**Method**: GET
|
||||
|
||||
**Path**: `/v1/models`
|
||||
|
||||
**Authentication**: None required
|
||||
|
||||
**Query Parameters**: None
|
||||
|
||||
**Request Headers**:
|
||||
```
|
||||
Accept: application/json
|
||||
```
|
||||
|
||||
**Response Headers**:
|
||||
```
|
||||
Content-Type: application/json
|
||||
```
|
||||
|
||||
**Response Schema**:
|
||||
```json
|
||||
{
|
||||
"object": "list",
|
||||
"data": [
|
||||
{
|
||||
"id": "model-name",
|
||||
"object": "model",
|
||||
"owned_by": "api.riotpiao.com",
|
||||
"created": 1700000000
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Status Codes**:
|
||||
- `200` - OK
|
||||
|
||||
**Example**:
|
||||
|
||||
```bash
|
||||
curl -s https://api.riotpiao.com/v1/models | jq .
|
||||
```
|
||||
|
||||
**Response Example**:
|
||||
```json
|
||||
{
|
||||
"object": "list",
|
||||
"data": [
|
||||
{
|
||||
"id": "reasoning",
|
||||
"object": "model",
|
||||
"owned_by": "api.riotpiao.com",
|
||||
"created": 1700000000
|
||||
},
|
||||
{
|
||||
"id": "ornith:35b",
|
||||
"object": "model",
|
||||
"owned_by": "api.riotpiao.com",
|
||||
"created": 1700000000
|
||||
},
|
||||
{
|
||||
"id": "qwen2.5:3b-instruct",
|
||||
"object": "model",
|
||||
"owned_by": "api.riotpiao.com",
|
||||
"created": 1700000000
|
||||
},
|
||||
{
|
||||
"id": "nomic-ai/nomic-embed-text-v2-moe",
|
||||
"object": "model",
|
||||
"owned_by": "api.riotpiao.com",
|
||||
"created": 1700000000
|
||||
},
|
||||
{
|
||||
"id": "BAAI/bge-reranker-base",
|
||||
"object": "model",
|
||||
"owned_by": "api.riotpiao.com",
|
||||
"created": 1700000000
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## POST /v1/chat/completions
|
||||
|
||||
Chat with an LLM model. Routes to upstream based on the `model` field in the request body.
|
||||
|
||||
**Method**: POST
|
||||
|
||||
**Path**: `/v1/chat/completions`
|
||||
|
||||
**Authentication**: None required (future: Bearer token)
|
||||
|
||||
**Request Headers**:
|
||||
```
|
||||
Content-Type: application/json
|
||||
```
|
||||
|
||||
**Request Body Schema**:
|
||||
```json
|
||||
{
|
||||
"model": "string (required)",
|
||||
"messages": [
|
||||
{
|
||||
"role": "string (user|assistant|system)",
|
||||
"content": "string|array (required)",
|
||||
"tool_calls": "array (optional, from assistant)"
|
||||
}
|
||||
],
|
||||
"temperature": "number (optional, 0-2)",
|
||||
"top_p": "number (optional, 0-1)",
|
||||
"max_tokens": "integer (optional)",
|
||||
"stream": "boolean (optional, default: false)",
|
||||
"tools": [
|
||||
{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "string",
|
||||
"description": "string",
|
||||
"parameters": "object"
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Response Schema** (non-streaming):
|
||||
```json
|
||||
{
|
||||
"id": "string",
|
||||
"object": "chat.completion",
|
||||
"created": "integer",
|
||||
"model": "string",
|
||||
"choices": [
|
||||
{
|
||||
"index": "integer",
|
||||
"message": {
|
||||
"role": "assistant",
|
||||
"content": "string|null",
|
||||
"tool_calls": [
|
||||
{
|
||||
"id": "string",
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "string",
|
||||
"arguments": "string (JSON)"
|
||||
}
|
||||
}
|
||||
]
|
||||
},
|
||||
"finish_reason": "stop|tool_calls|length"
|
||||
}
|
||||
],
|
||||
"usage": {
|
||||
"prompt_tokens": "integer",
|
||||
"completion_tokens": "integer",
|
||||
"total_tokens": "integer"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Response Schema** (streaming):
|
||||
```
|
||||
data: {"id":"...", "object":"chat.completion.chunk", "choices":[...]}
|
||||
data: {"id":"...", "object":"chat.completion.chunk", "choices":[...]}
|
||||
...
|
||||
data: [DONE]
|
||||
```
|
||||
|
||||
**Status Codes**:
|
||||
- `200` - OK
|
||||
- `400` - Bad request (missing/invalid model, invalid JSON, etc.)
|
||||
- `500` - Internal server error (upstream issue)
|
||||
|
||||
**Supported Models**:
|
||||
- `reasoning` - Reasoning model
|
||||
- `ornith:35b` - Ornith 35B model
|
||||
- `qwen2.5:3b-instruct` - Qwen 2.5 3B model
|
||||
|
||||
**Examples**:
|
||||
|
||||
### Basic Chat
|
||||
|
||||
```bash
|
||||
curl -X POST https://api.riotpiao.com/v1/chat/completions \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{
|
||||
"model": "reasoning",
|
||||
"messages": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": "What is the capital of France?"
|
||||
}
|
||||
]
|
||||
}'
|
||||
```
|
||||
|
||||
### Chat with Tool Calling
|
||||
|
||||
```bash
|
||||
curl -X POST https://api.riotpiao.com/v1/chat/completions \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{
|
||||
"model": "reasoning",
|
||||
"messages": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": "What is the weather in San Francisco?"
|
||||
}
|
||||
],
|
||||
"tools": [
|
||||
{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "get_weather",
|
||||
"description": "Get the weather for a location",
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"location": {
|
||||
"type": "string",
|
||||
"description": "City name"
|
||||
},
|
||||
"unit": {
|
||||
"type": "string",
|
||||
"enum": ["celsius", "fahrenheit"]
|
||||
}
|
||||
},
|
||||
"required": ["location"]
|
||||
}
|
||||
}
|
||||
}
|
||||
]
|
||||
}'
|
||||
```
|
||||
|
||||
### Streaming Chat
|
||||
|
||||
```bash
|
||||
curl -N -X POST https://api.riotpiao.com/v1/chat/completions \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{
|
||||
"model": "reasoning",
|
||||
"messages": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": "Count from 1 to 3"
|
||||
}
|
||||
],
|
||||
"stream": true
|
||||
}'
|
||||
```
|
||||
|
||||
### Multi-turn Conversation with Tool Results
|
||||
|
||||
```bash
|
||||
curl -X POST https://api.riotpiao.com/v1/chat/completions \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{
|
||||
"model": "reasoning",
|
||||
"messages": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": "What is the weather?"
|
||||
},
|
||||
{
|
||||
"role": "assistant",
|
||||
"tool_calls": [
|
||||
{
|
||||
"id": "call_123",
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "get_weather",
|
||||
"arguments": "{\"location\": \"San Francisco\"}"
|
||||
}
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"role": "tool",
|
||||
"content": "{\"temperature\": 22, \"condition\": \"sunny\"}"
|
||||
}
|
||||
],
|
||||
"tools": [
|
||||
{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "get_weather",
|
||||
"description": "Get weather",
|
||||
"parameters": {}
|
||||
}
|
||||
}
|
||||
]
|
||||
}'
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## POST /v1/embeddings
|
||||
|
||||
Generate embeddings for text input.
|
||||
|
||||
**Method**: POST
|
||||
|
||||
**Path**: `/v1/embeddings`
|
||||
|
||||
**Authentication**: None required
|
||||
|
||||
**Request Headers**:
|
||||
```
|
||||
Content-Type: application/json
|
||||
```
|
||||
|
||||
**Request Body Schema**:
|
||||
```json
|
||||
{
|
||||
"model": "string (required)",
|
||||
"input": "string | array of strings (required)",
|
||||
"encoding_format": "float | base64 (optional)"
|
||||
}
|
||||
```
|
||||
|
||||
**Response Schema**:
|
||||
```json
|
||||
{
|
||||
"object": "list",
|
||||
"data": [
|
||||
{
|
||||
"object": "embedding",
|
||||
"embedding": [0.1, 0.2, ...],
|
||||
"index": "integer"
|
||||
}
|
||||
],
|
||||
"model": "string",
|
||||
"usage": {
|
||||
"prompt_tokens": "integer",
|
||||
"total_tokens": "integer"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Status Codes**:
|
||||
- `200` - OK
|
||||
- `400` - Bad request (missing/invalid model, etc.)
|
||||
- `500` - Internal server error
|
||||
|
||||
**Supported Models**:
|
||||
- `nomic-ai/nomic-embed-text-v2-moe` - Embedding model
|
||||
|
||||
**Examples**:
|
||||
|
||||
### Single Input
|
||||
|
||||
```bash
|
||||
curl -X POST https://api.riotpiao.com/v1/embeddings \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{
|
||||
"model": "nomic-ai/nomic-embed-text-v2-moe",
|
||||
"input": "The quick brown fox"
|
||||
}'
|
||||
```
|
||||
|
||||
### Multiple Inputs
|
||||
|
||||
```bash
|
||||
curl -X POST https://api.riotpiao.com/v1/embeddings \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{
|
||||
"model": "nomic-ai/nomic-embed-text-v2-moe",
|
||||
"input": [
|
||||
"Document 1 text",
|
||||
"Document 2 text",
|
||||
"Document 3 text"
|
||||
]
|
||||
}'
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## POST /v1/rerank
|
||||
|
||||
Rerank documents based on relevance to a query.
|
||||
|
||||
**Method**: POST
|
||||
|
||||
**Path**: `/v1/rerank`
|
||||
|
||||
**Authentication**: None required
|
||||
|
||||
**Request Headers**:
|
||||
```
|
||||
Content-Type: application/json
|
||||
```
|
||||
|
||||
**Request Body Schema**:
|
||||
```json
|
||||
{
|
||||
"model": "string (required)",
|
||||
"query": "string (required)",
|
||||
"texts": ["string"],
|
||||
"top_k": "integer (optional)",
|
||||
"return_documents": "boolean (optional)"
|
||||
}
|
||||
```
|
||||
|
||||
**Response Schema**:
|
||||
```json
|
||||
{
|
||||
"results": [
|
||||
{
|
||||
"index": "integer",
|
||||
"score": "float (0-1)",
|
||||
"text": "string (optional)"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Status Codes**:
|
||||
- `200` - OK
|
||||
- `400` - Bad request (missing/invalid model, etc.)
|
||||
- `500` - Internal server error
|
||||
|
||||
**Supported Models**:
|
||||
- `BAAI/bge-reranker-base` - BGE reranker model
|
||||
|
||||
**Note**: The gateway rewrites the path from `/v1/rerank` to `/rerank` on the upstream.
|
||||
|
||||
**Examples**:
|
||||
|
||||
### Basic Reranking
|
||||
|
||||
```bash
|
||||
curl -X POST https://api.riotpiao.com/v1/rerank \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{
|
||||
"model": "BAAI/bge-reranker-base",
|
||||
"query": "What is machine learning?",
|
||||
"texts": [
|
||||
"Machine learning is a type of artificial intelligence",
|
||||
"Dogs are animals",
|
||||
"Deep learning is a subset of machine learning",
|
||||
"Python is a programming language"
|
||||
]
|
||||
}'
|
||||
```
|
||||
|
||||
### With Top-K Parameter
|
||||
|
||||
```bash
|
||||
curl -X POST https://api.riotpiao.com/v1/rerank \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{
|
||||
"model": "BAAI/bge-reranker-base",
|
||||
"query": "best practices",
|
||||
"texts": [
|
||||
"Follow code style guidelines",
|
||||
"Write unit tests",
|
||||
"Use meaningful variable names",
|
||||
"Eat healthy food"
|
||||
],
|
||||
"top_k": 2
|
||||
}'
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Error Handling
|
||||
|
||||
### Error Response Format
|
||||
|
||||
The gateway returns RFC 9457 Problem Details for client errors (4xx):
|
||||
|
||||
```json
|
||||
{
|
||||
"type": "https://api.example.com/problems/error-type",
|
||||
"title": "Human-readable error title",
|
||||
"status": 400,
|
||||
"detail": "Detailed explanation of what went wrong",
|
||||
"valid_models": ["model1", "model2"] // Only for model-related errors
|
||||
}
|
||||
```
|
||||
|
||||
### Error Types
|
||||
|
||||
#### Unknown Model Error
|
||||
|
||||
**Status**: `400 Bad Request`
|
||||
|
||||
**Trigger**: Model name not in registry
|
||||
|
||||
**Response**:
|
||||
```json
|
||||
{
|
||||
"type": "https://api.example.com/problems/unknown-model",
|
||||
"title": "Unknown Model",
|
||||
"status": 400,
|
||||
"detail": "Model \"gpt-4\" is not available. See valid_models for available options.",
|
||||
"valid_models": ["reasoning", "ornith:35b", "qwen2.5:3b-instruct", "nomic-ai/nomic-embed-text-v2-moe", "BAAI/bge-reranker-base"]
|
||||
}
|
||||
```
|
||||
|
||||
**Example**:
|
||||
```bash
|
||||
curl -X POST https://api.riotpiao.com/v1/chat/completions \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"model":"gpt-4","messages":[]}'
|
||||
```
|
||||
|
||||
#### Missing Model Field
|
||||
|
||||
**Status**: `400 Bad Request`
|
||||
|
||||
**Trigger**: No `model` field in request body
|
||||
|
||||
**Response**:
|
||||
```json
|
||||
{
|
||||
"type": "https://api.example.com/problems/missing-model",
|
||||
"title": "Missing Model",
|
||||
"status": 400,
|
||||
"detail": "The 'model' field is required and must be a non-empty string",
|
||||
"valid_models": [...]
|
||||
}
|
||||
```
|
||||
|
||||
**Example**:
|
||||
```bash
|
||||
curl -X POST https://api.riotpiao.com/v1/chat/completions \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"messages":[]}'
|
||||
```
|
||||
|
||||
#### Invalid JSON
|
||||
|
||||
**Status**: `400 Bad Request`
|
||||
|
||||
**Trigger**: Request body is not valid JSON
|
||||
|
||||
**Response**:
|
||||
```json
|
||||
{
|
||||
"type": "https://api.example.com/problems/invalid-request-body",
|
||||
"title": "Invalid Request Body",
|
||||
"status": 400,
|
||||
"detail": "request body is not valid JSON"
|
||||
}
|
||||
```
|
||||
|
||||
**Example**:
|
||||
```bash
|
||||
curl -X POST https://api.riotpiao.com/v1/chat/completions \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d 'not json'
|
||||
```
|
||||
|
||||
#### Upstream Error
|
||||
|
||||
**Status**: `5xx` (from upstream)
|
||||
|
||||
**Trigger**: Upstream service error
|
||||
|
||||
**Response**: Forwarded from upstream (unmodified)
|
||||
|
||||
---
|
||||
|
||||
## Examples
|
||||
|
||||
### Test Script
|
||||
|
||||
```bash
|
||||
#!/bin/bash
|
||||
|
||||
GATEWAY="https://api.riotpiao.com"
|
||||
|
||||
echo "=== Testing Gateway API ==="
|
||||
echo ""
|
||||
|
||||
# Test 1: Health checks
|
||||
echo "1. Health checks"
|
||||
curl -s "$GATEWAY/healthz" | jq .
|
||||
curl -s "$GATEWAY/readyz" | jq .
|
||||
echo ""
|
||||
|
||||
# Test 2: List models
|
||||
echo "2. List models"
|
||||
curl -s "$GATEWAY/v1/models" | jq '.data[] | .id'
|
||||
echo ""
|
||||
|
||||
# Test 3: Chat with reasoning model
|
||||
echo "3. Chat with reasoning model"
|
||||
curl -s -X POST "$GATEWAY/v1/chat/completions" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{
|
||||
"model": "reasoning",
|
||||
"messages": [{"role": "user", "content": "What is 2+2?"}]
|
||||
}' | jq '.choices[0].message.content'
|
||||
echo ""
|
||||
|
||||
# Test 4: Unknown model (should be 400)
|
||||
echo "4. Unknown model (should be 400)"
|
||||
curl -s -X POST "$GATEWAY/v1/chat/completions" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"model":"gpt-4","messages":[]}' | jq '{status: .status, title: .title}'
|
||||
echo ""
|
||||
|
||||
# Test 5: Embeddings
|
||||
echo "5. Embeddings"
|
||||
curl -s -X POST "$GATEWAY/v1/embeddings" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{
|
||||
"model": "nomic-ai/nomic-embed-text-v2-moe",
|
||||
"input": "hello world"
|
||||
}' | jq '.data | length'
|
||||
echo ""
|
||||
|
||||
# Test 6: Rerank
|
||||
echo "6. Rerank"
|
||||
curl -s -X POST "$GATEWAY/v1/rerank" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{
|
||||
"model": "BAAI/bge-reranker-base",
|
||||
"query": "test",
|
||||
"texts": ["a", "b"]
|
||||
}' | jq '.results | length'
|
||||
echo ""
|
||||
|
||||
# Test 7: Streaming
|
||||
echo "7. Streaming (showing first 5 chunks)"
|
||||
curl -s -N -X POST "$GATEWAY/v1/chat/completions" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{
|
||||
"model": "reasoning",
|
||||
"messages": [{"role": "user", "content": "hi"}],
|
||||
"stream": true
|
||||
}' | head -10
|
||||
echo ""
|
||||
|
||||
echo "=== All tests completed ==="
|
||||
```
|
||||
|
||||
### Python Client Example
|
||||
|
||||
```python
|
||||
import requests
|
||||
import json
|
||||
|
||||
GATEWAY = "https://api.riotpiao.com"
|
||||
|
||||
# Get models
|
||||
response = requests.get(f"{GATEWAY}/v1/models")
|
||||
models = response.json()
|
||||
print(f"Available models: {[m['id'] for m in models['data']]}")
|
||||
|
||||
# Chat completion
|
||||
response = requests.post(
|
||||
f"{GATEWAY}/v1/chat/completions",
|
||||
json={
|
||||
"model": "reasoning",
|
||||
"messages": [
|
||||
{"role": "user", "content": "What is machine learning?"}
|
||||
]
|
||||
}
|
||||
)
|
||||
message = response.json()
|
||||
print(f"Response: {message['choices'][0]['message']['content']}")
|
||||
|
||||
# Chat with tools
|
||||
response = requests.post(
|
||||
f"{GATEWAY}/v1/chat/completions",
|
||||
json={
|
||||
"model": "reasoning",
|
||||
"messages": [
|
||||
{"role": "user", "content": "Get the weather"}
|
||||
],
|
||||
"tools": [
|
||||
{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "get_weather",
|
||||
"description": "Get weather",
|
||||
"parameters": {}
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
)
|
||||
result = response.json()
|
||||
if "tool_calls" in result["choices"][0]["message"]:
|
||||
print(f"Tool calls: {result['choices'][0]['message']['tool_calls']}")
|
||||
|
||||
# Streaming
|
||||
response = requests.post(
|
||||
f"{GATEWAY}/v1/chat/completions",
|
||||
json={
|
||||
"model": "reasoning",
|
||||
"messages": [
|
||||
{"role": "user", "content": "Count to 3"}
|
||||
],
|
||||
"stream": True
|
||||
},
|
||||
stream=True
|
||||
)
|
||||
for line in response.iter_lines():
|
||||
if line:
|
||||
print(line)
|
||||
|
||||
# Embeddings
|
||||
response = requests.post(
|
||||
f"{GATEWAY}/v1/embeddings",
|
||||
json={
|
||||
"model": "nomic-ai/nomic-embed-text-v2-moe",
|
||||
"input": "hello world"
|
||||
}
|
||||
)
|
||||
embeddings = response.json()
|
||||
print(f"Embeddings: {embeddings['data'][0]['embedding'][:5]}")
|
||||
|
||||
# Rerank
|
||||
response = requests.post(
|
||||
f"{GATEWAY}/v1/rerank",
|
||||
json={
|
||||
"model": "BAAI/bge-reranker-base",
|
||||
"query": "ML",
|
||||
"texts": ["machine learning", "python", "deep learning"]
|
||||
}
|
||||
)
|
||||
results = response.json()
|
||||
print(f"Rerank results: {results['results']}")
|
||||
```
|
||||
|
||||
### JavaScript/TypeScript Client Example
|
||||
|
||||
```typescript
|
||||
const GATEWAY = "https://api.riotpiao.com";
|
||||
|
||||
// Get models
|
||||
async function getModels() {
|
||||
const response = await fetch(`${GATEWAY}/v1/models`);
|
||||
const data = await response.json();
|
||||
return data.data.map((m: any) => m.id);
|
||||
}
|
||||
|
||||
// Chat completion
|
||||
async function chat(model: string, message: string) {
|
||||
const response = await fetch(`${GATEWAY}/v1/chat/completions`, {
|
||||
method: "POST",
|
||||
headers: { "Content-Type": "application/json" },
|
||||
body: JSON.stringify({
|
||||
model,
|
||||
messages: [{ role: "user", content: message }],
|
||||
}),
|
||||
});
|
||||
const data = await response.json();
|
||||
return data.choices[0].message.content;
|
||||
}
|
||||
|
||||
// Chat with streaming
|
||||
async function chatStream(model: string, message: string) {
|
||||
const response = await fetch(`${GATEWAY}/v1/chat/completions`, {
|
||||
method: "POST",
|
||||
headers: { "Content-Type": "application/json" },
|
||||
body: JSON.stringify({
|
||||
model,
|
||||
messages: [{ role: "user", content: message }],
|
||||
stream: true,
|
||||
}),
|
||||
});
|
||||
|
||||
const reader = response.body!.getReader();
|
||||
const decoder = new TextDecoder();
|
||||
|
||||
while (true) {
|
||||
const { done, value } = await reader.read();
|
||||
if (done) break;
|
||||
|
||||
const chunk = decoder.decode(value);
|
||||
const lines = chunk.split("\n");
|
||||
for (const line of lines) {
|
||||
if (line.startsWith("data: ")) {
|
||||
const data = JSON.parse(line.slice(6));
|
||||
if (data.choices[0].delta?.content) {
|
||||
console.log(data.choices[0].delta.content);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Embeddings
|
||||
async function embed(model: string, input: string[]) {
|
||||
const response = await fetch(`${GATEWAY}/v1/embeddings`, {
|
||||
method: "POST",
|
||||
headers: { "Content-Type": "application/json" },
|
||||
body: JSON.stringify({ model, input }),
|
||||
});
|
||||
const data = await response.json();
|
||||
return data.data;
|
||||
}
|
||||
|
||||
// Rerank
|
||||
async function rerank(
|
||||
model: string,
|
||||
query: string,
|
||||
texts: string[]
|
||||
) {
|
||||
const response = await fetch(`${GATEWAY}/v1/rerank`, {
|
||||
method: "POST",
|
||||
headers: { "Content-Type": "application/json" },
|
||||
body: JSON.stringify({ model, query, texts }),
|
||||
});
|
||||
const data = await response.json();
|
||||
return data.results;
|
||||
}
|
||||
|
||||
// Usage
|
||||
(async () => {
|
||||
const models = await getModels();
|
||||
console.log("Models:", models);
|
||||
|
||||
const response = await chat("reasoning", "What is AI?");
|
||||
console.log("Response:", response);
|
||||
|
||||
await chatStream("reasoning", "Count to 3");
|
||||
|
||||
const embeddings = await embed("nomic-ai/nomic-embed-text-v2-moe", [
|
||||
"hello",
|
||||
]);
|
||||
console.log("Embeddings:", embeddings);
|
||||
|
||||
const rerankResults = await rerank("BAAI/bge-reranker-base", "ML", [
|
||||
"machine learning",
|
||||
"python",
|
||||
]);
|
||||
console.log("Rerank:", rerankResults);
|
||||
})();
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Rate Limiting
|
||||
|
||||
Currently, no rate limiting is enforced. This will be added in Phase 4.
|
||||
|
||||
---
|
||||
|
||||
## Authentication
|
||||
|
||||
Currently, no authentication is enforced. Bearer token support will be added in Phase 3.
|
||||
|
||||
---
|
||||
|
||||
## Timeouts
|
||||
|
||||
Default timeouts per route:
|
||||
- **Connect**: 10s
|
||||
- **Read**: 1h (for streaming)
|
||||
- **Write**: 1h
|
||||
|
||||
These are configured per model upstream.
|
||||
|
||||
---
|
||||
|
||||
## Body Size Limits
|
||||
|
||||
- **Default**: 100MB
|
||||
- **Per-route**: Configurable
|
||||
|
||||
Requests exceeding the limit return `413 Request Entity Too Large`.
|
||||
|
||||
---
|
||||
|
||||
## Support
|
||||
|
||||
For issues or questions:
|
||||
- Check gateway logs: `kubectl -n api logs deployment/homelab-frontend`
|
||||
- Check health: `curl https://api.riotpiao.com/healthz`
|
||||
- Verify config: `curl https://api.riotpiao.com/v1/models`
|
||||
Reference in New Issue
Block a user