23 KiB
API Reference: api.riotpiao.com
Gateway to all homelab services. Single entry point for LLM inference, workflows, queues, memory, and cluster operations.
Base URL: https://api.riotpiao.com
Table of Contents
- Health & Status
- LLM Services
- Workflow Services (Temporal)
- Queue Services (SQS)
- Memory Services
- IAM Services
- S3 Services
- Authentication
- Error Handling
Health & Status
GET /healthz
Liveness probe. Returns immediately without checks.
curl https://api.riotpiao.com/healthz
Response:
{
"status": "alive"
}
Status Code: 200
GET /readyz
Readiness probe. Returns 200 when gateway is ready (config loaded, connections available).
curl https://api.riotpiao.com/readyz
Response:
{
"status": "ready"
}
Status Code: 200 (ready) or 503 (not ready)
LLM Services
GET /v1/models
List all available models for inference.
curl https://api.riotpiao.com/v1/models
Response:
{
"object": "list",
"data": [
{
"id": "reasoning",
"object": "model",
"owned_by": "api.riotpiao.com",
"created": 1700000000
},
{
"id": "ornith:35b",
"object": "model",
"owned_by": "api.riotpiao.com",
"created": 1700000000
},
{
"id": "qwen2.5:3b-instruct",
"object": "model",
"owned_by": "api.riotpiao.com",
"created": 1700000000
},
{
"id": "nomic-ai/nomic-embed-text-v2-moe",
"object": "model",
"owned_by": "api.riotpiao.com",
"created": 1700000000
},
{
"id": "BAAI/bge-reranker-base",
"object": "model",
"owned_by": "api.riotpiao.com",
"created": 1700000000
}
]
}
Status Code: 200
POST /v1/chat/completions
Chat with an LLM model. Supports streaming, tool calling, and multi-turn conversations.
curl -X POST https://api.riotpiao.com/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "reasoning",
"messages": [
{"role": "user", "content": "What is 2+2?"}
]
}'
Request Body:
{
"model": "reasoning",
"messages": [
{
"role": "user",
"content": "string or array"
}
],
"temperature": 0.7,
"top_p": 1.0,
"max_tokens": 2048,
"stream": false,
"tools": [
{
"type": "function",
"function": {
"name": "function_name",
"description": "What it does",
"parameters": {
"type": "object",
"properties": {},
"required": []
}
}
}
]
}
Response (non-streaming):
{
"id": "chatcmpl-123",
"object": "chat.completion",
"created": 1704000000,
"model": "reasoning",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "2+2=4"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 10,
"completion_tokens": 5,
"total_tokens": 15
}
}
Response (streaming, SSE format):
data: {"id":"...","object":"chat.completion.chunk","choices":[{"delta":{"content":"2"}}]}
data: {"id":"...","object":"chat.completion.chunk","choices":[{"delta":{"content":"+"}}]}
data: {"id":"...","object":"chat.completion.chunk","choices":[{"delta":{"content":"2"}}]}
data: {"id":"...","object":"chat.completion.chunk","choices":[{"delta":{"content":"="}}]}
data: {"id":"...","object":"chat.completion.chunk","choices":[{"delta":{"content":"4"}}]}
data: [DONE]
Status Codes:
- 200 OK
- 400 Bad Request (invalid model, missing fields, invalid JSON)
- 500 Internal Server Error (upstream issue)
Available Models:
reasoning— DeepSeek-R1-Distill reasoning model (8 concurrent slots)ornith:35b— Ollama 35B modelqwen2.5:3b-instruct— Qwen 2.5 3B model
POST /v1/embeddings
Generate text embeddings.
curl -X POST https://api.riotpiao.com/v1/embeddings \
-H 'Content-Type: application/json' \
-d '{
"model": "nomic-ai/nomic-embed-text-v2-moe",
"input": ["text to embed", "another text"]
}'
Request Body:
{
"model": "nomic-ai/nomic-embed-text-v2-moe",
"input": "string or array of strings",
"encoding_format": "float"
}
Response:
{
"object": "list",
"data": [
{
"object": "embedding",
"embedding": [0.1, 0.2, 0.3, ...],
"index": 0
},
{
"object": "embedding",
"embedding": [0.4, 0.5, 0.6, ...],
"index": 1
}
],
"model": "nomic-ai/nomic-embed-text-v2-moe",
"usage": {
"prompt_tokens": 20,
"total_tokens": 20
}
}
Status Codes:
- 200 OK
- 400 Bad Request (invalid model or input)
- 500 Internal Server Error
POST /v1/rerank
Rerank documents by relevance to a query.
curl -X POST https://api.riotpiao.com/v1/rerank \
-H 'Content-Type: application/json' \
-d '{
"model": "BAAI/bge-reranker-base",
"query": "machine learning",
"texts": [
"Machine learning is AI",
"Python is a language",
"Deep learning is ML"
],
"top_k": 2
}'
Request Body:
{
"model": "BAAI/bge-reranker-base",
"query": "search query",
"texts": ["text1", "text2", "text3"],
"top_k": 2,
"return_documents": true
}
Response:
{
"results": [
{
"index": 0,
"score": 0.95,
"text": "Machine learning is AI"
},
{
"index": 2,
"score": 0.85,
"text": "Deep learning is ML"
}
]
}
Status Codes:
- 200 OK
- 400 Bad Request
- 500 Internal Server Error
Workflow Services (Temporal)
Base path: /workflow
All operations use REST with JSON body. Internally translated to gRPC (port 7233).
POST /workflow
Execute a workflow operation via action dispatch.
curl -X POST https://api.riotpiao.com/workflow \
-H 'Content-Type: application/json' \
-d '{
"action": "START_WORKFLOW",
"namespace": "default",
"payload": {
"workflow_id": "order-123",
"workflow_type": "ProcessOrder",
"task_queue": "orders_queue",
"input": {
"order_id": "123",
"amount": 99.99
}
}
}'
Supported Actions:
START_WORKFLOW
Start a new workflow execution.
Payload:
{
"workflow_id": "unique-id",
"workflow_type": "WorkflowName",
"task_queue": "queue_name",
"input": {"field": "value"},
"options": {
"workflow_execution_timeout": 3600,
"workflow_run_timeout": 1800,
"workflow_task_timeout": 300
}
}
Response:
{
"success": true,
"action": "START_WORKFLOW",
"data": {
"workflow_id": "order-123",
"run_id": "abc123def",
"start_time": "2026-08-27T..."
}
}
DESCRIBE_WORKFLOW
Get workflow execution details.
Payload:
{
"workflow_id": "order-123",
"run_id": "abc123def"
}
Response:
{
"success": true,
"action": "DESCRIBE_WORKFLOW",
"data": {
"workflow_id": "order-123",
"run_id": "abc123def",
"status": "RUNNING",
"start_time": "2026-08-27T..."
}
}
LIST_WORKFLOWS
List workflow executions with filtering.
Payload:
{
"status": "RUNNING",
"page_size": 50,
"next_page_token": ""
}
GET_WORKFLOW_HISTORY
Retrieve workflow execution history (events).
Payload:
{
"workflow_id": "order-123",
"run_id": "abc123def",
"max_events": 100
}
SIGNAL_WORKFLOW
Send a signal to a running workflow (trigger handler).
Payload:
{
"workflow_id": "order-123",
"run_id": "abc123def",
"signal_name": "payment_received",
"input": {"amount": 99.99}
}
QUERY_WORKFLOW
Query workflow state without modifying it (read-only).
Payload:
{
"workflow_id": "order-123",
"run_id": "abc123def",
"query_type": "get_status"
}
CANCEL_WORKFLOW
Request graceful cancellation.
Payload:
{
"workflow_id": "order-123",
"run_id": "abc123def"
}
TERMINATE_WORKFLOW
Stop workflow immediately.
Payload:
{
"workflow_id": "order-123",
"run_id": "abc123def",
"reason": "User cancelled"
}
Other Actions
- RESET_WORKFLOW
- UPDATE_WORKFLOW
- HEARTBEAT_ACTIVITY
- COMPLETE_ACTIVITY
- FAIL_ACTIVITY
Status Codes:
- 200 OK
- 400 Bad Request (invalid action or fields)
- 404 Not Found (workflow not found)
- 503 Service Unavailable (Temporal unreachable)
Queue Services (SQS)
Base path: /sqs (via X-Service header routing)
Operations via ServiceAdapter CRD. Use header-based dispatch:
curl -X GET https://api.riotpiao.com/ \
-H 'X-Service: sqs' \
-H 'X-Resource: queue/my-queue/messages'
Supported Resources:
queue— List queuesqueue/{name}/message— Send/receive messagesqueue/{name}/messages— Batch operations
Memory Services
Base path: / (via X-Service header)
Query and manage semantic memory for knowledge retrieval.
curl -X GET https://api.riotpiao.com/ \
-H 'X-Service: memory' \
-H 'X-Resource: query' \
-G --data-urlencode 'query=explain machine learning' \
--data-urlencode 'level=L1,L2'
Supported Resources:
query— Search memory (GET with params)skill— List/retrieve skillsproject— List/describe projectsingest— Add new documents (POST, requires memory:write capability)
IAM Services
Base path: / (via X-Service header)
Manage Authentik users, groups, and permissions.
curl -X GET https://api.riotpiao.com/ \
-H 'Authorization: Bearer <token>' \
-H 'X-Service: iam' \
-H 'X-Resource: user'
Supported Resources:
user— List/create usersuser/{id}— Get/update/delete specific userrole(group) — Manage groups/rolespermission— List permissionsflow— List authentication flows
S3 Services
Base path: / (via X-Service header)
Object storage operations via MinIO.
curl -X GET https://api.riotpiao.com/ \
-H 'X-Service: s3' \
-H 'X-Resource: bucket/my-bucket/objects'
Supported Resources:
bucket— List bucketsbucket/{name}/objects— List objects in bucketbucket/{name}/object/{key}— Get/put/delete object
Authentication
Bearer Token (JWT)
All operations except /healthz and /readyz require authentication.
curl -H 'Authorization: Bearer <jwt-token>' \
https://api.riotpiao.com/v1/models
Obtaining Tokens
Via Authentik OIDC (human login):
core auth login --username [email protected]
Via service account (programmatic):
core mwinit login --username service-account --password secret
export RIOTPIAO_TOKEN=$(cat ~/.talos/.riotpiao-auth)
curl -H "Authorization: Bearer $RIOTPIAO_TOKEN" \
https://api.riotpiao.com/v1/models
Capabilities (RBAC)
Tokens embed capabilities in claims. Required capabilities:
llm:inference—/v1/*chat/embeddings/rerankworkflow:execute—/workflowoperationsmemory:read— Memory queriesmemory:write— Memory ingestsqs:access— Queue operationss3:access— S3 operationsiam:admin— IAM management
Error Handling
Problem Details (RFC 9457)
All error responses use standardized JSON format:
{
"type": "https://api.example.com/problems/unknown-model",
"title": "Unknown Model",
"status": 400,
"detail": "Model 'gpt-4' is not available. See valid_models for available options.",
"valid_models": ["reasoning", "ornith:35b", ...]
}
HTTP Status Codes
| Status | Meaning |
|---|---|
| 200 | Success |
| 400 | Bad Request (validation error, invalid model, missing fields) |
| 401 | Unauthorized (invalid/missing token) |
| 403 | Forbidden (insufficient permissions/capability) |
| 404 | Not Found (workflow, resource) |
| 409 | Conflict (duplicate workflow_id) |
| 429 | Too Many Requests (rate limit exceeded) |
| 503 | Service Unavailable (backend unreachable) |
| 504 | Gateway Timeout (request exceeded timeout) |
Error Examples
Unknown Model:
curl -X POST https://api.riotpiao.com/v1/chat/completions \
-d '{"model":"gpt-4","messages":[]}'
# Returns 400
{
"type": "https://api.example.com/problems/unknown-model",
"title": "Unknown Model",
"status": 400,
"detail": "Model 'gpt-4' is not available",
"valid_models": ["reasoning", "ornith:35b", "qwen2.5:3b-instruct", ...]
}
Temporal Unavailable:
curl -X POST https://api.riotpiao.com/workflow \
-d '{"action":"START_WORKFLOW",...}'
# Returns 503 if Temporal server unreachable
{
"success": false,
"error": "TEMPORAL_UNAVAILABLE",
"message": "Temporal server connection not available"
}
Examples
Complete RAG Pipeline
- Query embedding:
curl -X POST https://api.riotpiao.com/v1/embeddings \
-H 'Authorization: Bearer $TOKEN' \
-d '{
"model": "nomic-ai/nomic-embed-text-v2-moe",
"input": "how do neural networks work"
}' | jq '.data[0].embedding'
- Rerank documents:
curl -X POST https://api.riotpiao.com/v1/rerank \
-H 'Authorization: Bearer $TOKEN' \
-d '{
"model": "BAAI/bge-reranker-base",
"query": "how do neural networks work",
"texts": ["doc1", "doc2", "doc3"],
"top_k": 2
}'
- Chat with context:
curl -X POST https://api.riotpiao.com/v1/chat/completions \
-H 'Authorization: Bearer $TOKEN' \
-d '{
"model": "reasoning",
"messages": [
{
"role": "user",
"content": "Based on these documents: [top-2 from rerank]\nQuestion: how do neural networks work?"
}
]
}'
Workflow Orchestration
# Start workflow
WORKFLOW_ID="order-$(date +%s)"
curl -X POST https://api.riotpiao.com/workflow \
-H 'Authorization: Bearer $TOKEN' \
-d "{
\"action\": \"START_WORKFLOW\",
\"namespace\": \"default\",
\"payload\": {
\"workflow_id\": \"$WORKFLOW_ID\",
\"workflow_type\": \"ProcessOrder\",
\"task_queue\": \"orders\",
\"input\": {\"order_id\": \"123\"}
}
}"
# Monitor execution
curl -X POST https://api.riotpiao.com/workflow \
-H 'Authorization: Bearer $TOKEN' \
-d "{
\"action\": \"DESCRIBE_WORKFLOW\",
\"namespace\": \"default\",
\"payload\": {\"workflow_id\": \"$WORKFLOW_ID\"}
}"
# Send signal when payment received
curl -X POST https://api.riotpiao.com/workflow \
-H 'Authorization: Bearer $TOKEN' \
-d "{
\"action\": \"SIGNAL_WORKFLOW\",
\"namespace\": \"default\",
\"payload\": {
\"workflow_id\": \"$WORKFLOW_ID\",
\"signal_name\": \"payment_received\",
\"input\": {\"amount\": 99.99}
}
}"
Rate Limits & Quotas
/v1/chat/completions— 8 concurrent slots (reasoning model)/v1/embeddings— 10 concurrent/v1/rerank— 10 concurrent/workflow— 100 concurrent operations- Retry-After header set on 429 responses
Timeouts
| Endpoint | Connect | Read | Write |
|---|---|---|---|
/v1/chat/completions |
10s | 1h | 1h |
/v1/embeddings |
10s | 10m | 10m |
/v1/rerank |
10s | 10m | 10m |
/workflow |
10s | 30s | 10s |
Support
- Health:
curl https://api.riotpiao.com/healthz - Readiness:
curl https://api.riotpiao.com/readyz - Models:
curl https://api.riotpiao.com/v1/models - Logs:
kubectl -n api logs deployment/homelab-frontend
LLM Inference in Workflows
The Poimen workflows system includes built-in LLM inference activities that call /v1/chat/completions via the gateway.
LLMInferenceActivity
Single-prompt LLM inference within a workflow.
Workflow Definition (Canvas Node):
{
"id": "llm-node-1",
"type": "llm-inference",
"label": "Analyze Code with LLM",
"data": {
"model": "reasoning",
"system_prompt": "You are a code analysis expert. Provide detailed feedback.",
"user_prompt": "Analyze this code for security issues: {{ previous_output.code }}",
"temperature": 0.7,
"max_tokens": 2048,
"auth_token": "{{ user.jwt_token }}"
}
}
Fields:
model(required): Model ID (reasoning, ornith:35b, ornith:13b, qwen2.5:3b)system_prompt: System instruction for the modeluser_prompt(required): User message to sendtemperature: Sampling temperature (0.0-1.0, default 0.7)max_tokens: Maximum output tokensauth_token(optional): JWT token for authenticated endpoints (propagates as Authorization: Bearer header)
Backend Implementation: The LLMInferenceActivity in the workflows backend automatically:
- Substitutes template variables (e.g.,
{{ previous_output.code }}) - Calls
/v1/chat/completionswith the resolved prompt - Returns the LLM response as activity output
- Retries on transient failures (up to 3 attempts)
- Timeouts after 120 seconds
Output:
{
"response": "The code has several security vulnerabilities...",
"model": "reasoning",
"stop_reason": "stop_sequence",
"tokens_used": 450
}
Supported Models:
reasoning— DeepSeek-R1-Distill (best for complex analysis)ornith:35b— Ollama 35Bornith:13b— Ollama 13Bqwen2.5:3b— Qwen 2.5 3B
LLMBatchInferenceActivity
Multiple-prompt LLM inference (sequential processing).
Workflow Definition:
{
"id": "llm-batch-1",
"type": "llm-batch-inference",
"label": "Batch Code Review",
"data": {
"model": "reasoning",
"system_prompt": "Review each code snippet and provide feedback.",
"prompts": [
"Review snippet 1: {{ files[0].content }}",
"Review snippet 2: {{ files[1].content }}",
"Review snippet 3: {{ files[2].content }}"
],
"auth_token": "{{ user.jwt_token }}"
}
}
Fields:
model(required): Model IDsystem_prompt: System instruction (same for all prompts)prompts(required): List of user prompts to processtemperature: Sampling temperature (0.0-1.0)auth_token(optional): JWT token for authenticated endpoints (propagates as Authorization: Bearer header)
Output:
{
"responses": [
"Snippet 1 review...",
"Snippet 2 review...",
"Snippet 3 review..."
],
"model": "reasoning",
"errors": []
}
Typical Use Cases:
- Batch code review across multiple files
- Parallel document summarization
- Comparative analysis of alternatives
- Policy compliance checking
Workflow Integration Examples
1. Code Analysis Workflow
Clone Repo → Analyze Code → LLM Security Review → Generate Report → Notify
2. Document Processing
Retrieve Documents → Embed + Index → LLM Summarize (batch) → Archive
3. Multi-Stage Review
Retrieve Memory → LLM Context Extraction → Route to Activity A/B/C → Notify
Authentication & Authorization
JWT tokens can be passed to LLM inference activities and are automatically propagated to the LLM API endpoint.
Token Flow:
Workflow Canvas
↓ (auth_token field)
Poimen Workflow Executor
↓ (passed to LLMInferenceActivity)
Activity calls LLM client
↓ (adds "Authorization: Bearer {token}" header)
homelab-frontend proxy
↓ (preserves Authorization header)
LLM Backend (reasoning/ollama/etc)
↓ (validates token)
Response returned
Example: Passing User Token from RetrieveMemory Activity
{
"id": "flow-1",
"type": "retrieve-memory",
"label": "Get User Context",
"data": {...}
}
→
{
"id": "llm-1",
"type": "llm-inference",
"label": "Analyze with User's Token",
"data": {
"model": "reasoning",
"user_prompt": "...",
"auth_token": "{{ previous_output.user_token }}"
}
}
Token Validation:
- Tokens are validated by homelab-frontend proxy (checks signature, expiration)
- Only valid tokens are propagated to LLM backend
- Invalid tokens result in 401 Unauthorized error
- Missing token (if required) results in 401 Unauthorized
Note: The auth_token field is optional. If omitted, the LLM API is called without authentication (public endpoints only).
CanvasReasonerActivity
Auto-suggest workflow connections using LLM reasoning. When you drop new activities onto the canvas, this activity analyzes them and suggests logical connections based on input/output compatibility and workflow patterns.
Workflow Definition:
{
"id": "canvas-reason-1",
"type": "canvas-reasoner",
"label": "Auto-Connect Activities",
"data": {
"nodes": "{{ workflow.nodes }}",
"edges": "{{ workflow.edges }}",
"preserve_existing": true,
"auth_token": "{{ user.jwt_token }}"
}
}
Use Cases:
- New nodes added to canvas → automatically suggest connections
- Validate workflow design → LLM reasoning explains connections
- Redesign workflow → suggest optimal activity sequence
- Data flow analysis → ensure proper input/output matching
How It Works:
- Analyzes all node types and their configurations
- Reviews existing edges (if preserving)
- Uses reasoning model to infer logical connections
- Returns suggested edges with confidence score
- Includes reasoning explanation
Output Example:
{
"suggested_edges": [
{"source": "clone-1", "target": "analyze-1"},
{"source": "analyze-1", "target": "security-scan-1"},
{"source": "security-scan-1", "target": "report-1"}
],
"reasoning": "Clone repository first, analyze code, perform security scan, generate report. Standard code review workflow.",
"confidence": 0.92
}
Fields:
nodes(required): Canvas nodes to analyzeedges(required): Current edgespreserve_existing(optional, default true): Keep existing edges and only suggest new onesauth_token(optional): JWT for LLM reasoning calls
Confidence Scores:
- 0.9-1.0: High confidence (common patterns)
- 0.7-0.9: Medium confidence (reasonable connections)
- 0.5-0.7: Low confidence (multiple valid approaches)
- <0.5: Unsure (manual review recommended)
Integration Example:
User drops 3 new nodes on canvas
↓
Workflow calls CanvasReasonerActivity
↓
LLM analyzes: Clone → Analyze → Report
↓
Returns edges + reasoning
↓
Frontend updates canvas with suggested connections
↓
User approves/rejects suggestions
Error Handling
If LLM inference fails:
- First activity retry (2-second backoff)
- Second activity retry (4-second backoff)
- Third activity retry (8-second backoff)
- If all retries fail, workflow records error and proceeds to next activity (or fails if terminal)
Common Failure Scenarios:
- Network timeout:
connection refused(retry automatically) - Model not found:
unknown model: xyz(terminal error) - Rate limited: HTTP 429 (retry with exponential backoff)
- Prompt too long:
context length exceeded(terminal error)
Performance & Cost
- Single prompt inference: ~100-500ms (model-dependent)
- Batch processing: Serial (not parallel), ~100-500ms per prompt
- Model inference costs: Free (on-premise Ollama/Reasoning models)
- Token counting: Provided in response for quota tracking
Optimization Tips:
- Use
ornith:13borqwen2.5:3bfor faster inference - Use
reasoningonly for complex analysis that needs reasoning - Cache frequently-used prompts at workflow level
- Use batch activity for multiple similar prompts (better throughput)