Files
homelab-frontend/API.md
T
Story Crater Bot c8c656046a docs: comprehensive api and testing documentation
- API.md: Full REST API documentation with examples
  * All endpoints (health, models, chat, embeddings, rerank)
  * Request/response schemas
  * Error handling (RFC 9457 problem+json)
  * Examples in bash, Python, TypeScript

- TESTING_GUIDE.md: Quick reference testing guide
  * 15 copy-paste test commands
  * Complete testing checklist
  * Troubleshooting guide
  * Performance testing examples
  * Integration test scripts

Ready for deployment verification and integration testing.
2026-08-19 23:55:22 -07:00

19 KiB

API Gateway Documentation

Overview

The homelab-frontend gateway is a production-ready reverse proxy for LLM model inference. It routes requests to multiple model upstreams based on configuration, with support for streaming, tool calling, and multiple API formats.

Base URL: https://api.riotpiao.com

Deployment: Client → nginx ingress → gateway → model upstreams


Table of Contents

  1. Health Endpoints
  2. GET /v1/models - List available models
  3. POST /v1/chat/completions - Chat with LLM
  4. POST /v1/embeddings - Generate embeddings
  5. POST /v1/rerank - Rerank documents
  6. Error Handling
  7. Examples

Health Endpoints

GET /healthz

Always returns 200 (liveness probe).

Response:

{"status":"alive"}

Status Code: 200


GET /readyz

Returns 200 when the gateway is ready (config loaded, upstreams available).

Response:

{"status":"ready"}

Status Code: 200 (ready) or 503 (not ready)


GET /v1/models

List all configured models available for dispatch.

Method: GET

Path: /v1/models

Authentication: None required

Query Parameters: None

Request Headers:

Accept: application/json

Response Headers:

Content-Type: application/json

Response Schema:

{
  "object": "list",
  "data": [
    {
      "id": "model-name",
      "object": "model",
      "owned_by": "api.riotpiao.com",
      "created": 1700000000
    }
  ]
}

Status Codes:

  • 200 - OK

Example:

curl -s https://api.riotpiao.com/v1/models | jq .

Response Example:

{
  "object": "list",
  "data": [
    {
      "id": "reasoning",
      "object": "model",
      "owned_by": "api.riotpiao.com",
      "created": 1700000000
    },
    {
      "id": "ornith:35b",
      "object": "model",
      "owned_by": "api.riotpiao.com",
      "created": 1700000000
    },
    {
      "id": "qwen2.5:3b-instruct",
      "object": "model",
      "owned_by": "api.riotpiao.com",
      "created": 1700000000
    },
    {
      "id": "nomic-ai/nomic-embed-text-v2-moe",
      "object": "model",
      "owned_by": "api.riotpiao.com",
      "created": 1700000000
    },
    {
      "id": "BAAI/bge-reranker-base",
      "object": "model",
      "owned_by": "api.riotpiao.com",
      "created": 1700000000
    }
  ]
}

POST /v1/chat/completions

Chat with an LLM model. Routes to upstream based on the model field in the request body.

Method: POST

Path: /v1/chat/completions

Authentication: None required (future: Bearer token)

Request Headers:

Content-Type: application/json

Request Body Schema:

{
  "model": "string (required)",
  "messages": [
    {
      "role": "string (user|assistant|system)",
      "content": "string|array (required)",
      "tool_calls": "array (optional, from assistant)"
    }
  ],
  "temperature": "number (optional, 0-2)",
  "top_p": "number (optional, 0-1)",
  "max_tokens": "integer (optional)",
  "stream": "boolean (optional, default: false)",
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "string",
        "description": "string",
        "parameters": "object"
      }
    }
  ]
}

Response Schema (non-streaming):

{
  "id": "string",
  "object": "chat.completion",
  "created": "integer",
  "model": "string",
  "choices": [
    {
      "index": "integer",
      "message": {
        "role": "assistant",
        "content": "string|null",
        "tool_calls": [
          {
            "id": "string",
            "type": "function",
            "function": {
              "name": "string",
              "arguments": "string (JSON)"
            }
          }
        ]
      },
      "finish_reason": "stop|tool_calls|length"
    }
  ],
  "usage": {
    "prompt_tokens": "integer",
    "completion_tokens": "integer",
    "total_tokens": "integer"
  }
}

Response Schema (streaming):

data: {"id":"...", "object":"chat.completion.chunk", "choices":[...]}
data: {"id":"...", "object":"chat.completion.chunk", "choices":[...]}
...
data: [DONE]

Status Codes:

  • 200 - OK
  • 400 - Bad request (missing/invalid model, invalid JSON, etc.)
  • 500 - Internal server error (upstream issue)

Supported Models:

  • reasoning - Reasoning model
  • ornith:35b - Ornith 35B model
  • qwen2.5:3b-instruct - Qwen 2.5 3B model

Examples:

Basic Chat

curl -X POST https://api.riotpiao.com/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "reasoning",
    "messages": [
      {
        "role": "user",
        "content": "What is the capital of France?"
      }
    ]
  }'

Chat with Tool Calling

curl -X POST https://api.riotpiao.com/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "reasoning",
    "messages": [
      {
        "role": "user",
        "content": "What is the weather in San Francisco?"
      }
    ],
    "tools": [
      {
        "type": "function",
        "function": {
          "name": "get_weather",
          "description": "Get the weather for a location",
          "parameters": {
            "type": "object",
            "properties": {
              "location": {
                "type": "string",
                "description": "City name"
              },
              "unit": {
                "type": "string",
                "enum": ["celsius", "fahrenheit"]
              }
            },
            "required": ["location"]
          }
        }
      }
    ]
  }'

Streaming Chat

curl -N -X POST https://api.riotpiao.com/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "reasoning",
    "messages": [
      {
        "role": "user",
        "content": "Count from 1 to 3"
      }
    ],
    "stream": true
  }'

Multi-turn Conversation with Tool Results

curl -X POST https://api.riotpiao.com/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "reasoning",
    "messages": [
      {
        "role": "user",
        "content": "What is the weather?"
      },
      {
        "role": "assistant",
        "tool_calls": [
          {
            "id": "call_123",
            "type": "function",
            "function": {
              "name": "get_weather",
              "arguments": "{\"location\": \"San Francisco\"}"
            }
          }
        ]
      },
      {
        "role": "tool",
        "content": "{\"temperature\": 22, \"condition\": \"sunny\"}"
      }
    ],
    "tools": [
      {
        "type": "function",
        "function": {
          "name": "get_weather",
          "description": "Get weather",
          "parameters": {}
        }
      }
    ]
  }'

POST /v1/embeddings

Generate embeddings for text input.

Method: POST

Path: /v1/embeddings

Authentication: None required

Request Headers:

Content-Type: application/json

Request Body Schema:

{
  "model": "string (required)",
  "input": "string | array of strings (required)",
  "encoding_format": "float | base64 (optional)"
}

Response Schema:

{
  "object": "list",
  "data": [
    {
      "object": "embedding",
      "embedding": [0.1, 0.2, ...],
      "index": "integer"
    }
  ],
  "model": "string",
  "usage": {
    "prompt_tokens": "integer",
    "total_tokens": "integer"
  }
}

Status Codes:

  • 200 - OK
  • 400 - Bad request (missing/invalid model, etc.)
  • 500 - Internal server error

Supported Models:

  • nomic-ai/nomic-embed-text-v2-moe - Embedding model

Examples:

Single Input

curl -X POST https://api.riotpiao.com/v1/embeddings \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "nomic-ai/nomic-embed-text-v2-moe",
    "input": "The quick brown fox"
  }'

Multiple Inputs

curl -X POST https://api.riotpiao.com/v1/embeddings \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "nomic-ai/nomic-embed-text-v2-moe",
    "input": [
      "Document 1 text",
      "Document 2 text",
      "Document 3 text"
    ]
  }'

POST /v1/rerank

Rerank documents based on relevance to a query.

Method: POST

Path: /v1/rerank

Authentication: None required

Request Headers:

Content-Type: application/json

Request Body Schema:

{
  "model": "string (required)",
  "query": "string (required)",
  "texts": ["string"],
  "top_k": "integer (optional)",
  "return_documents": "boolean (optional)"
}

Response Schema:

{
  "results": [
    {
      "index": "integer",
      "score": "float (0-1)",
      "text": "string (optional)"
    }
  ]
}

Status Codes:

  • 200 - OK
  • 400 - Bad request (missing/invalid model, etc.)
  • 500 - Internal server error

Supported Models:

  • BAAI/bge-reranker-base - BGE reranker model

Note: The gateway rewrites the path from /v1/rerank to /rerank on the upstream.

Examples:

Basic Reranking

curl -X POST https://api.riotpiao.com/v1/rerank \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "BAAI/bge-reranker-base",
    "query": "What is machine learning?",
    "texts": [
      "Machine learning is a type of artificial intelligence",
      "Dogs are animals",
      "Deep learning is a subset of machine learning",
      "Python is a programming language"
    ]
  }'

With Top-K Parameter

curl -X POST https://api.riotpiao.com/v1/rerank \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "BAAI/bge-reranker-base",
    "query": "best practices",
    "texts": [
      "Follow code style guidelines",
      "Write unit tests",
      "Use meaningful variable names",
      "Eat healthy food"
    ],
    "top_k": 2
  }'

Error Handling

Error Response Format

The gateway returns RFC 9457 Problem Details for client errors (4xx):

{
  "type": "https://api.example.com/problems/error-type",
  "title": "Human-readable error title",
  "status": 400,
  "detail": "Detailed explanation of what went wrong",
  "valid_models": ["model1", "model2"] // Only for model-related errors
}

Error Types

Unknown Model Error

Status: 400 Bad Request

Trigger: Model name not in registry

Response:

{
  "type": "https://api.example.com/problems/unknown-model",
  "title": "Unknown Model",
  "status": 400,
  "detail": "Model \"gpt-4\" is not available. See valid_models for available options.",
  "valid_models": ["reasoning", "ornith:35b", "qwen2.5:3b-instruct", "nomic-ai/nomic-embed-text-v2-moe", "BAAI/bge-reranker-base"]
}

Example:

curl -X POST https://api.riotpiao.com/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"gpt-4","messages":[]}'

Missing Model Field

Status: 400 Bad Request

Trigger: No model field in request body

Response:

{
  "type": "https://api.example.com/problems/missing-model",
  "title": "Missing Model",
  "status": 400,
  "detail": "The 'model' field is required and must be a non-empty string",
  "valid_models": [...]
}

Example:

curl -X POST https://api.riotpiao.com/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"messages":[]}'

Invalid JSON

Status: 400 Bad Request

Trigger: Request body is not valid JSON

Response:

{
  "type": "https://api.example.com/problems/invalid-request-body",
  "title": "Invalid Request Body",
  "status": 400,
  "detail": "request body is not valid JSON"
}

Example:

curl -X POST https://api.riotpiao.com/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d 'not json'

Upstream Error

Status: 5xx (from upstream)

Trigger: Upstream service error

Response: Forwarded from upstream (unmodified)


Examples

Test Script

#!/bin/bash

GATEWAY="https://api.riotpiao.com"

echo "=== Testing Gateway API ==="
echo ""

# Test 1: Health checks
echo "1. Health checks"
curl -s "$GATEWAY/healthz" | jq .
curl -s "$GATEWAY/readyz" | jq .
echo ""

# Test 2: List models
echo "2. List models"
curl -s "$GATEWAY/v1/models" | jq '.data[] | .id'
echo ""

# Test 3: Chat with reasoning model
echo "3. Chat with reasoning model"
curl -s -X POST "$GATEWAY/v1/chat/completions" \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "reasoning",
    "messages": [{"role": "user", "content": "What is 2+2?"}]
  }' | jq '.choices[0].message.content'
echo ""

# Test 4: Unknown model (should be 400)
echo "4. Unknown model (should be 400)"
curl -s -X POST "$GATEWAY/v1/chat/completions" \
  -H 'Content-Type: application/json' \
  -d '{"model":"gpt-4","messages":[]}' | jq '{status: .status, title: .title}'
echo ""

# Test 5: Embeddings
echo "5. Embeddings"
curl -s -X POST "$GATEWAY/v1/embeddings" \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "nomic-ai/nomic-embed-text-v2-moe",
    "input": "hello world"
  }' | jq '.data | length'
echo ""

# Test 6: Rerank
echo "6. Rerank"
curl -s -X POST "$GATEWAY/v1/rerank" \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "BAAI/bge-reranker-base",
    "query": "test",
    "texts": ["a", "b"]
  }' | jq '.results | length'
echo ""

# Test 7: Streaming
echo "7. Streaming (showing first 5 chunks)"
curl -s -N -X POST "$GATEWAY/v1/chat/completions" \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "reasoning",
    "messages": [{"role": "user", "content": "hi"}],
    "stream": true
  }' | head -10
echo ""

echo "=== All tests completed ==="

Python Client Example

import requests
import json

GATEWAY = "https://api.riotpiao.com"

# Get models
response = requests.get(f"{GATEWAY}/v1/models")
models = response.json()
print(f"Available models: {[m['id'] for m in models['data']]}")

# Chat completion
response = requests.post(
    f"{GATEWAY}/v1/chat/completions",
    json={
        "model": "reasoning",
        "messages": [
            {"role": "user", "content": "What is machine learning?"}
        ]
    }
)
message = response.json()
print(f"Response: {message['choices'][0]['message']['content']}")

# Chat with tools
response = requests.post(
    f"{GATEWAY}/v1/chat/completions",
    json={
        "model": "reasoning",
        "messages": [
            {"role": "user", "content": "Get the weather"}
        ],
        "tools": [
            {
                "type": "function",
                "function": {
                    "name": "get_weather",
                    "description": "Get weather",
                    "parameters": {}
                }
            }
        ]
    }
)
result = response.json()
if "tool_calls" in result["choices"][0]["message"]:
    print(f"Tool calls: {result['choices'][0]['message']['tool_calls']}")

# Streaming
response = requests.post(
    f"{GATEWAY}/v1/chat/completions",
    json={
        "model": "reasoning",
        "messages": [
            {"role": "user", "content": "Count to 3"}
        ],
        "stream": True
    },
    stream=True
)
for line in response.iter_lines():
    if line:
        print(line)

# Embeddings
response = requests.post(
    f"{GATEWAY}/v1/embeddings",
    json={
        "model": "nomic-ai/nomic-embed-text-v2-moe",
        "input": "hello world"
    }
)
embeddings = response.json()
print(f"Embeddings: {embeddings['data'][0]['embedding'][:5]}")

# Rerank
response = requests.post(
    f"{GATEWAY}/v1/rerank",
    json={
        "model": "BAAI/bge-reranker-base",
        "query": "ML",
        "texts": ["machine learning", "python", "deep learning"]
    }
)
results = response.json()
print(f"Rerank results: {results['results']}")

JavaScript/TypeScript Client Example

const GATEWAY = "https://api.riotpiao.com";

// Get models
async function getModels() {
  const response = await fetch(`${GATEWAY}/v1/models`);
  const data = await response.json();
  return data.data.map((m: any) => m.id);
}

// Chat completion
async function chat(model: string, message: string) {
  const response = await fetch(`${GATEWAY}/v1/chat/completions`, {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({
      model,
      messages: [{ role: "user", content: message }],
    }),
  });
  const data = await response.json();
  return data.choices[0].message.content;
}

// Chat with streaming
async function chatStream(model: string, message: string) {
  const response = await fetch(`${GATEWAY}/v1/chat/completions`, {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({
      model,
      messages: [{ role: "user", content: message }],
      stream: true,
    }),
  });

  const reader = response.body!.getReader();
  const decoder = new TextDecoder();

  while (true) {
    const { done, value } = await reader.read();
    if (done) break;

    const chunk = decoder.decode(value);
    const lines = chunk.split("\n");
    for (const line of lines) {
      if (line.startsWith("data: ")) {
        const data = JSON.parse(line.slice(6));
        if (data.choices[0].delta?.content) {
          console.log(data.choices[0].delta.content);
        }
      }
    }
  }
}

// Embeddings
async function embed(model: string, input: string[]) {
  const response = await fetch(`${GATEWAY}/v1/embeddings`, {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({ model, input }),
  });
  const data = await response.json();
  return data.data;
}

// Rerank
async function rerank(
  model: string,
  query: string,
  texts: string[]
) {
  const response = await fetch(`${GATEWAY}/v1/rerank`, {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({ model, query, texts }),
  });
  const data = await response.json();
  return data.results;
}

// Usage
(async () => {
  const models = await getModels();
  console.log("Models:", models);

  const response = await chat("reasoning", "What is AI?");
  console.log("Response:", response);

  await chatStream("reasoning", "Count to 3");

  const embeddings = await embed("nomic-ai/nomic-embed-text-v2-moe", [
    "hello",
  ]);
  console.log("Embeddings:", embeddings);

  const rerankResults = await rerank("BAAI/bge-reranker-base", "ML", [
    "machine learning",
    "python",
  ]);
  console.log("Rerank:", rerankResults);
})();

Rate Limiting

Currently, no rate limiting is enforced. This will be added in Phase 4.


Authentication

Currently, no authentication is enforced. Bearer token support will be added in Phase 3.


Timeouts

Default timeouts per route:

  • Connect: 10s
  • Read: 1h (for streaming)
  • Write: 1h

These are configured per model upstream.


Body Size Limits

  • Default: 100MB
  • Per-route: Configurable

Requests exceeding the limit return 413 Request Entity Too Large.


Support

For issues or questions:

  • Check gateway logs: kubectl -n api logs deployment/homelab-frontend
  • Check health: curl https://api.riotpiao.com/healthz
  • Verify config: curl https://api.riotpiao.com/v1/models