1185 lines
27 KiB
Markdown
1185 lines
27 KiB
Markdown
# API Reference: api.riotpiao.com
|
|
|
|
Gateway to all homelab services. Single entry point for LLM inference, workflows, queues, memory, and cluster operations.
|
|
|
|
**Base URL:** `https://api.riotpiao.com`
|
|
|
|
---
|
|
|
|
## Table of Contents
|
|
|
|
1. [Health & Status](#health--status)
|
|
2. [LLM Services](#llm-services)
|
|
3. [Workflow Services (Temporal)](#workflow-services-temporal)
|
|
4. [Queue Services (SQS)](#queue-services-sqs)
|
|
5. [Memory Services](#memory-services)
|
|
6. [IAM Services](#iam-services)
|
|
7. [S3 Services](#s3-services)
|
|
8. [Authentication](#authentication)
|
|
9. [Error Handling](#error-handling)
|
|
|
|
---
|
|
|
|
## Health & Status
|
|
|
|
### GET /healthz
|
|
|
|
Liveness probe. Returns immediately without checks.
|
|
|
|
```bash
|
|
curl https://api.riotpiao.com/healthz
|
|
```
|
|
|
|
**Response:**
|
|
```json
|
|
{
|
|
"status": "alive"
|
|
}
|
|
```
|
|
|
|
**Status Code:** 200
|
|
|
|
---
|
|
|
|
### GET /readyz
|
|
|
|
Readiness probe. Returns 200 when gateway is ready (config loaded, connections available).
|
|
|
|
```bash
|
|
curl https://api.riotpiao.com/readyz
|
|
```
|
|
|
|
**Response:**
|
|
```json
|
|
{
|
|
"status": "ready"
|
|
}
|
|
```
|
|
|
|
**Status Code:** 200 (ready) or 503 (not ready)
|
|
|
|
---
|
|
|
|
## LLM Services
|
|
|
|
### GET /v1/models
|
|
|
|
List all available models for inference.
|
|
|
|
```bash
|
|
curl https://api.riotpiao.com/v1/models
|
|
```
|
|
|
|
**Response:**
|
|
```json
|
|
{
|
|
"object": "list",
|
|
"data": [
|
|
{
|
|
"id": "reasoning",
|
|
"object": "model",
|
|
"owned_by": "api.riotpiao.com",
|
|
"created": 1700000000
|
|
},
|
|
{
|
|
"id": "ornith:35b",
|
|
"object": "model",
|
|
"owned_by": "api.riotpiao.com",
|
|
"created": 1700000000
|
|
},
|
|
{
|
|
"id": "qwen2.5:3b-instruct",
|
|
"object": "model",
|
|
"owned_by": "api.riotpiao.com",
|
|
"created": 1700000000
|
|
},
|
|
{
|
|
"id": "nomic-ai/nomic-embed-text-v2-moe",
|
|
"object": "model",
|
|
"owned_by": "api.riotpiao.com",
|
|
"created": 1700000000
|
|
},
|
|
{
|
|
"id": "BAAI/bge-reranker-base",
|
|
"object": "model",
|
|
"owned_by": "api.riotpiao.com",
|
|
"created": 1700000000
|
|
}
|
|
]
|
|
}
|
|
```
|
|
|
|
**Status Code:** 200
|
|
|
|
---
|
|
|
|
### POST /v1/chat/completions
|
|
|
|
Chat with an LLM model. Supports streaming, tool calling, and multi-turn conversations.
|
|
|
|
```bash
|
|
curl -X POST https://api.riotpiao.com/v1/chat/completions \
|
|
-H 'Content-Type: application/json' \
|
|
-d '{
|
|
"model": "reasoning",
|
|
"messages": [
|
|
{"role": "user", "content": "What is 2+2?"}
|
|
]
|
|
}'
|
|
```
|
|
|
|
**Request Body:**
|
|
```json
|
|
{
|
|
"model": "reasoning",
|
|
"messages": [
|
|
{
|
|
"role": "user",
|
|
"content": "string or array"
|
|
}
|
|
],
|
|
"temperature": 0.7,
|
|
"top_p": 1.0,
|
|
"max_tokens": 2048,
|
|
"stream": false,
|
|
"tools": [
|
|
{
|
|
"type": "function",
|
|
"function": {
|
|
"name": "function_name",
|
|
"description": "What it does",
|
|
"parameters": {
|
|
"type": "object",
|
|
"properties": {},
|
|
"required": []
|
|
}
|
|
}
|
|
}
|
|
]
|
|
}
|
|
```
|
|
|
|
**Response (non-streaming):**
|
|
```json
|
|
{
|
|
"id": "chatcmpl-123",
|
|
"object": "chat.completion",
|
|
"created": 1704000000,
|
|
"model": "reasoning",
|
|
"choices": [
|
|
{
|
|
"index": 0,
|
|
"message": {
|
|
"role": "assistant",
|
|
"content": "2+2=4"
|
|
},
|
|
"finish_reason": "stop"
|
|
}
|
|
],
|
|
"usage": {
|
|
"prompt_tokens": 10,
|
|
"completion_tokens": 5,
|
|
"total_tokens": 15
|
|
}
|
|
}
|
|
```
|
|
|
|
**Response (streaming, SSE format):**
|
|
```
|
|
data: {"id":"...","object":"chat.completion.chunk","choices":[{"delta":{"content":"2"}}]}
|
|
data: {"id":"...","object":"chat.completion.chunk","choices":[{"delta":{"content":"+"}}]}
|
|
data: {"id":"...","object":"chat.completion.chunk","choices":[{"delta":{"content":"2"}}]}
|
|
data: {"id":"...","object":"chat.completion.chunk","choices":[{"delta":{"content":"="}}]}
|
|
data: {"id":"...","object":"chat.completion.chunk","choices":[{"delta":{"content":"4"}}]}
|
|
data: [DONE]
|
|
```
|
|
|
|
**Status Codes:**
|
|
- 200 OK
|
|
- 400 Bad Request (invalid model, missing fields, invalid JSON)
|
|
- 500 Internal Server Error (upstream issue)
|
|
|
|
**Available Models:**
|
|
- `reasoning` — DeepSeek-R1-Distill reasoning model (8 concurrent slots)
|
|
- `ornith:35b` — Ollama 35B model
|
|
- `qwen2.5:3b-instruct` — Qwen 2.5 3B model
|
|
|
|
---
|
|
|
|
### POST /v1/embeddings
|
|
|
|
Generate text embeddings.
|
|
|
|
```bash
|
|
curl -X POST https://api.riotpiao.com/v1/embeddings \
|
|
-H 'Content-Type: application/json' \
|
|
-d '{
|
|
"model": "nomic-ai/nomic-embed-text-v2-moe",
|
|
"input": ["text to embed", "another text"]
|
|
}'
|
|
```
|
|
|
|
**Request Body:**
|
|
```json
|
|
{
|
|
"model": "nomic-ai/nomic-embed-text-v2-moe",
|
|
"input": "string or array of strings",
|
|
"encoding_format": "float"
|
|
}
|
|
```
|
|
|
|
**Response:**
|
|
```json
|
|
{
|
|
"object": "list",
|
|
"data": [
|
|
{
|
|
"object": "embedding",
|
|
"embedding": [0.1, 0.2, 0.3, ...],
|
|
"index": 0
|
|
},
|
|
{
|
|
"object": "embedding",
|
|
"embedding": [0.4, 0.5, 0.6, ...],
|
|
"index": 1
|
|
}
|
|
],
|
|
"model": "nomic-ai/nomic-embed-text-v2-moe",
|
|
"usage": {
|
|
"prompt_tokens": 20,
|
|
"total_tokens": 20
|
|
}
|
|
}
|
|
```
|
|
|
|
**Status Codes:**
|
|
- 200 OK
|
|
- 400 Bad Request (invalid model or input)
|
|
- 500 Internal Server Error
|
|
|
|
---
|
|
|
|
### POST /v1/rerank
|
|
|
|
Rerank documents by relevance to a query.
|
|
|
|
```bash
|
|
curl -X POST https://api.riotpiao.com/v1/rerank \
|
|
-H 'Content-Type: application/json' \
|
|
-d '{
|
|
"model": "BAAI/bge-reranker-base",
|
|
"query": "machine learning",
|
|
"texts": [
|
|
"Machine learning is AI",
|
|
"Python is a language",
|
|
"Deep learning is ML"
|
|
],
|
|
"top_k": 2
|
|
}'
|
|
```
|
|
|
|
**Request Body:**
|
|
```json
|
|
{
|
|
"model": "BAAI/bge-reranker-base",
|
|
"query": "search query",
|
|
"texts": ["text1", "text2", "text3"],
|
|
"top_k": 2,
|
|
"return_documents": true
|
|
}
|
|
```
|
|
|
|
**Response:**
|
|
```json
|
|
{
|
|
"results": [
|
|
{
|
|
"index": 0,
|
|
"score": 0.95,
|
|
"text": "Machine learning is AI"
|
|
},
|
|
{
|
|
"index": 2,
|
|
"score": 0.85,
|
|
"text": "Deep learning is ML"
|
|
}
|
|
]
|
|
}
|
|
```
|
|
|
|
**Status Codes:**
|
|
- 200 OK
|
|
- 400 Bad Request
|
|
- 500 Internal Server Error
|
|
|
|
---
|
|
|
|
## Workflow Services (Temporal)
|
|
|
|
Base path: `/workflow`
|
|
|
|
All operations use REST with JSON body. Internally translated to gRPC (port 7233).
|
|
|
|
### POST /workflow
|
|
|
|
Execute a workflow operation via action dispatch.
|
|
|
|
```bash
|
|
curl -X POST https://api.riotpiao.com/workflow \
|
|
-H 'Content-Type: application/json' \
|
|
-d '{
|
|
"action": "START_WORKFLOW",
|
|
"namespace": "default",
|
|
"payload": {
|
|
"workflow_id": "order-123",
|
|
"workflow_type": "ProcessOrder",
|
|
"task_queue": "orders_queue",
|
|
"input": {
|
|
"order_id": "123",
|
|
"amount": 99.99
|
|
}
|
|
}
|
|
}'
|
|
```
|
|
|
|
**Supported Actions:**
|
|
|
|
#### START_WORKFLOW
|
|
Start a new workflow execution.
|
|
|
|
**Payload:**
|
|
```json
|
|
{
|
|
"workflow_id": "unique-id",
|
|
"workflow_type": "WorkflowName",
|
|
"task_queue": "queue_name",
|
|
"input": {"field": "value"},
|
|
"options": {
|
|
"workflow_execution_timeout": 3600,
|
|
"workflow_run_timeout": 1800,
|
|
"workflow_task_timeout": 300
|
|
}
|
|
}
|
|
```
|
|
|
|
**Response:**
|
|
```json
|
|
{
|
|
"success": true,
|
|
"action": "START_WORKFLOW",
|
|
"data": {
|
|
"workflow_id": "order-123",
|
|
"run_id": "abc123def",
|
|
"start_time": "2026-08-27T..."
|
|
}
|
|
}
|
|
```
|
|
|
|
#### DESCRIBE_WORKFLOW
|
|
Get workflow execution details.
|
|
|
|
**Payload:**
|
|
```json
|
|
{
|
|
"workflow_id": "order-123",
|
|
"run_id": "abc123def"
|
|
}
|
|
```
|
|
|
|
**Response:**
|
|
```json
|
|
{
|
|
"success": true,
|
|
"action": "DESCRIBE_WORKFLOW",
|
|
"data": {
|
|
"workflow_id": "order-123",
|
|
"run_id": "abc123def",
|
|
"status": "RUNNING",
|
|
"start_time": "2026-08-27T..."
|
|
}
|
|
}
|
|
```
|
|
|
|
#### LIST_WORKFLOWS
|
|
List workflow executions with filtering.
|
|
|
|
**Payload:**
|
|
```json
|
|
{
|
|
"status": "RUNNING",
|
|
"page_size": 50,
|
|
"next_page_token": ""
|
|
}
|
|
```
|
|
|
|
#### GET_WORKFLOW_HISTORY
|
|
Retrieve workflow execution history (events).
|
|
|
|
**Payload:**
|
|
```json
|
|
{
|
|
"workflow_id": "order-123",
|
|
"run_id": "abc123def",
|
|
"max_events": 100
|
|
}
|
|
```
|
|
|
|
#### SIGNAL_WORKFLOW
|
|
Send a signal to a running workflow (trigger handler).
|
|
|
|
**Payload:**
|
|
```json
|
|
{
|
|
"workflow_id": "order-123",
|
|
"run_id": "abc123def",
|
|
"signal_name": "payment_received",
|
|
"input": {"amount": 99.99}
|
|
}
|
|
```
|
|
|
|
#### QUERY_WORKFLOW
|
|
Query workflow state without modifying it (read-only).
|
|
|
|
**Payload:**
|
|
```json
|
|
{
|
|
"workflow_id": "order-123",
|
|
"run_id": "abc123def",
|
|
"query_type": "get_status"
|
|
}
|
|
```
|
|
|
|
#### CANCEL_WORKFLOW
|
|
Request graceful cancellation.
|
|
|
|
**Payload:**
|
|
```json
|
|
{
|
|
"workflow_id": "order-123",
|
|
"run_id": "abc123def"
|
|
}
|
|
```
|
|
|
|
#### TERMINATE_WORKFLOW
|
|
Stop workflow immediately.
|
|
|
|
**Payload:**
|
|
```json
|
|
{
|
|
"workflow_id": "order-123",
|
|
"run_id": "abc123def",
|
|
"reason": "User cancelled"
|
|
}
|
|
```
|
|
|
|
#### Other Actions
|
|
- RESET_WORKFLOW
|
|
- UPDATE_WORKFLOW
|
|
- HEARTBEAT_ACTIVITY
|
|
- COMPLETE_ACTIVITY
|
|
- FAIL_ACTIVITY
|
|
|
|
**Status Codes:**
|
|
- 200 OK
|
|
- 400 Bad Request (invalid action or fields)
|
|
- 404 Not Found (workflow not found)
|
|
- 503 Service Unavailable (Temporal unreachable)
|
|
|
|
---
|
|
|
|
## Queue Services (SQS)
|
|
|
|
Base path: `/sqs` (via X-Service header routing)
|
|
|
|
Operations via ServiceAdapter CRD. Use header-based dispatch:
|
|
|
|
```bash
|
|
curl -X GET https://api.riotpiao.com/ \
|
|
-H 'X-Service: sqs' \
|
|
-H 'X-Resource: queue/my-queue/messages'
|
|
```
|
|
|
|
**Supported Resources:**
|
|
- `queue` — List queues
|
|
- `queue/{name}/message` — Send/receive messages
|
|
- `queue/{name}/messages` — Batch operations
|
|
|
|
---
|
|
|
|
## Memory Services
|
|
|
|
Base path: `/` (via X-Service header)
|
|
|
|
Query and manage semantic memory for knowledge retrieval.
|
|
|
|
```bash
|
|
curl -X GET https://api.riotpiao.com/ \
|
|
-H 'X-Service: memory' \
|
|
-H 'X-Resource: query' \
|
|
-G --data-urlencode 'query=explain machine learning' \
|
|
--data-urlencode 'level=L1,L2'
|
|
```
|
|
|
|
**Supported Resources:**
|
|
- `query` — Search memory (GET with params)
|
|
- `skill` — List/retrieve skills
|
|
- `project` — List/describe projects
|
|
- `ingest` — Add new documents (POST, requires memory:write capability)
|
|
|
|
---
|
|
|
|
## IAM Services
|
|
|
|
Base path: `/` (via X-Service header)
|
|
|
|
Manage Authentik users, groups, and permissions.
|
|
|
|
```bash
|
|
curl -X GET https://api.riotpiao.com/ \
|
|
-H 'Authorization: Bearer <token>' \
|
|
-H 'X-Service: iam' \
|
|
-H 'X-Resource: user'
|
|
```
|
|
|
|
**Supported Resources:**
|
|
- `user` — List/create users
|
|
- `user/{id}` — Get/update/delete specific user
|
|
- `role` (group) — Manage groups/roles
|
|
- `permission` — List permissions
|
|
- `flow` — List authentication flows
|
|
|
|
---
|
|
|
|
## S3 Services
|
|
|
|
Base path: `/` (via X-Service header)
|
|
|
|
Object storage operations via MinIO.
|
|
|
|
```bash
|
|
curl -X GET https://api.riotpiao.com/ \
|
|
-H 'X-Service: s3' \
|
|
-H 'X-Resource: bucket/my-bucket/objects'
|
|
```
|
|
|
|
**Supported Resources:**
|
|
- `bucket` — List buckets
|
|
- `bucket/{name}/objects` — List objects in bucket
|
|
- `bucket/{name}/object/{key}` — Get/put/delete object
|
|
|
|
---
|
|
|
|
## Authentication
|
|
|
|
### Bearer Token (JWT)
|
|
|
|
All operations except `/healthz` and `/readyz` require authentication.
|
|
|
|
```bash
|
|
curl -H 'Authorization: Bearer <jwt-token>' \
|
|
https://api.riotpiao.com/v1/models
|
|
```
|
|
|
|
### Obtaining Tokens
|
|
|
|
**Via Authentik OIDC (human login):**
|
|
```bash
|
|
core auth login --username [email protected]
|
|
```
|
|
|
|
**Via service account (programmatic):**
|
|
```bash
|
|
core mwinit login --username service-account --password secret
|
|
export RIOTPIAO_TOKEN=$(cat ~/.talos/.riotpiao-auth)
|
|
|
|
curl -H "Authorization: Bearer $RIOTPIAO_TOKEN" \
|
|
https://api.riotpiao.com/v1/models
|
|
```
|
|
|
|
### Capabilities (RBAC)
|
|
|
|
Tokens embed capabilities in claims. Required capabilities:
|
|
|
|
- `llm:inference` — `/v1/*` chat/embeddings/rerank
|
|
- `workflow:execute` — `/workflow` operations
|
|
- `memory:read` — Memory queries
|
|
- `memory:write` — Memory ingest
|
|
- `sqs:access` — Queue operations
|
|
- `s3:access` — S3 operations
|
|
- `iam:admin` — IAM management
|
|
|
|
---
|
|
|
|
## Error Handling
|
|
|
|
### Problem Details (RFC 9457)
|
|
|
|
All error responses use standardized JSON format:
|
|
|
|
```json
|
|
{
|
|
"type": "https://api.example.com/problems/unknown-model",
|
|
"title": "Unknown Model",
|
|
"status": 400,
|
|
"detail": "Model 'gpt-4' is not available. See valid_models for available options.",
|
|
"valid_models": ["reasoning", "ornith:35b", ...]
|
|
}
|
|
```
|
|
|
|
### HTTP Status Codes
|
|
|
|
| Status | Meaning |
|
|
|--------|---------|
|
|
| 200 | Success |
|
|
| 400 | Bad Request (validation error, invalid model, missing fields) |
|
|
| 401 | Unauthorized (invalid/missing token) |
|
|
| 403 | Forbidden (insufficient permissions/capability) |
|
|
| 404 | Not Found (workflow, resource) |
|
|
| 409 | Conflict (duplicate workflow_id) |
|
|
| 429 | Too Many Requests (rate limit exceeded) |
|
|
| 503 | Service Unavailable (backend unreachable) |
|
|
| 504 | Gateway Timeout (request exceeded timeout) |
|
|
|
|
### Error Examples
|
|
|
|
**Unknown Model:**
|
|
```bash
|
|
curl -X POST https://api.riotpiao.com/v1/chat/completions \
|
|
-d '{"model":"gpt-4","messages":[]}'
|
|
|
|
# Returns 400
|
|
{
|
|
"type": "https://api.example.com/problems/unknown-model",
|
|
"title": "Unknown Model",
|
|
"status": 400,
|
|
"detail": "Model 'gpt-4' is not available",
|
|
"valid_models": ["reasoning", "ornith:35b", "qwen2.5:3b-instruct", ...]
|
|
}
|
|
```
|
|
|
|
**Temporal Unavailable:**
|
|
```bash
|
|
curl -X POST https://api.riotpiao.com/workflow \
|
|
-d '{"action":"START_WORKFLOW",...}'
|
|
|
|
# Returns 503 if Temporal server unreachable
|
|
{
|
|
"success": false,
|
|
"error": "TEMPORAL_UNAVAILABLE",
|
|
"message": "Temporal server connection not available"
|
|
}
|
|
```
|
|
|
|
---
|
|
|
|
## Examples
|
|
|
|
### Complete RAG Pipeline
|
|
|
|
1. **Query embedding:**
|
|
```bash
|
|
curl -X POST https://api.riotpiao.com/v1/embeddings \
|
|
-H 'Authorization: Bearer $TOKEN' \
|
|
-d '{
|
|
"model": "nomic-ai/nomic-embed-text-v2-moe",
|
|
"input": "how do neural networks work"
|
|
}' | jq '.data[0].embedding'
|
|
```
|
|
|
|
2. **Rerank documents:**
|
|
```bash
|
|
curl -X POST https://api.riotpiao.com/v1/rerank \
|
|
-H 'Authorization: Bearer $TOKEN' \
|
|
-d '{
|
|
"model": "BAAI/bge-reranker-base",
|
|
"query": "how do neural networks work",
|
|
"texts": ["doc1", "doc2", "doc3"],
|
|
"top_k": 2
|
|
}'
|
|
```
|
|
|
|
3. **Chat with context:**
|
|
```bash
|
|
curl -X POST https://api.riotpiao.com/v1/chat/completions \
|
|
-H 'Authorization: Bearer $TOKEN' \
|
|
-d '{
|
|
"model": "reasoning",
|
|
"messages": [
|
|
{
|
|
"role": "user",
|
|
"content": "Based on these documents: [top-2 from rerank]\nQuestion: how do neural networks work?"
|
|
}
|
|
]
|
|
}'
|
|
```
|
|
|
|
### Workflow Orchestration
|
|
|
|
```bash
|
|
# Start workflow
|
|
WORKFLOW_ID="order-$(date +%s)"
|
|
curl -X POST https://api.riotpiao.com/workflow \
|
|
-H 'Authorization: Bearer $TOKEN' \
|
|
-d "{
|
|
\"action\": \"START_WORKFLOW\",
|
|
\"namespace\": \"default\",
|
|
\"payload\": {
|
|
\"workflow_id\": \"$WORKFLOW_ID\",
|
|
\"workflow_type\": \"ProcessOrder\",
|
|
\"task_queue\": \"orders\",
|
|
\"input\": {\"order_id\": \"123\"}
|
|
}
|
|
}"
|
|
|
|
# Monitor execution
|
|
curl -X POST https://api.riotpiao.com/workflow \
|
|
-H 'Authorization: Bearer $TOKEN' \
|
|
-d "{
|
|
\"action\": \"DESCRIBE_WORKFLOW\",
|
|
\"namespace\": \"default\",
|
|
\"payload\": {\"workflow_id\": \"$WORKFLOW_ID\"}
|
|
}"
|
|
|
|
# Send signal when payment received
|
|
curl -X POST https://api.riotpiao.com/workflow \
|
|
-H 'Authorization: Bearer $TOKEN' \
|
|
-d "{
|
|
\"action\": \"SIGNAL_WORKFLOW\",
|
|
\"namespace\": \"default\",
|
|
\"payload\": {
|
|
\"workflow_id\": \"$WORKFLOW_ID\",
|
|
\"signal_name\": \"payment_received\",
|
|
\"input\": {\"amount\": 99.99}
|
|
}
|
|
}"
|
|
```
|
|
|
|
---
|
|
|
|
## Rate Limits & Quotas
|
|
|
|
- `/v1/chat/completions` — 8 concurrent slots (reasoning model)
|
|
- `/v1/embeddings` — 10 concurrent
|
|
- `/v1/rerank` — 10 concurrent
|
|
- `/workflow` — 100 concurrent operations
|
|
- Retry-After header set on 429 responses
|
|
|
|
---
|
|
|
|
## Timeouts
|
|
|
|
| Endpoint | Connect | Read | Write |
|
|
|----------|---------|------|-------|
|
|
| `/v1/chat/completions` | 10s | 1h | 1h |
|
|
| `/v1/embeddings` | 10s | 10m | 10m |
|
|
| `/v1/rerank` | 10s | 10m | 10m |
|
|
| `/workflow` | 10s | 30s | 10s |
|
|
|
|
---
|
|
|
|
## Support
|
|
|
|
- **Health:** `curl https://api.riotpiao.com/healthz`
|
|
- **Readiness:** `curl https://api.riotpiao.com/readyz`
|
|
- **Models:** `curl https://api.riotpiao.com/v1/models`
|
|
- **Logs:** `kubectl -n api logs deployment/homelab-frontend`
|
|
|
|
---
|
|
|
|
## LLM Inference in Workflows
|
|
|
|
The Poimen workflows system includes built-in LLM inference activities that call `/v1/chat/completions` via the gateway.
|
|
|
|
### LLMInferenceActivity
|
|
|
|
Single-prompt LLM inference within a workflow.
|
|
|
|
**Workflow Definition (Canvas Node):**
|
|
```json
|
|
{
|
|
"id": "llm-node-1",
|
|
"type": "llm-inference",
|
|
"label": "Analyze Code with LLM",
|
|
"data": {
|
|
"model": "reasoning",
|
|
"system_prompt": "You are a code analysis expert. Provide detailed feedback.",
|
|
"user_prompt": "Analyze this code for security issues: {{ previous_output.code }}",
|
|
"temperature": 0.7,
|
|
"max_tokens": 2048,
|
|
"auth_token": "{{ user.jwt_token }}"
|
|
}
|
|
}
|
|
```
|
|
|
|
**Fields:**
|
|
- `model` (required): Model ID (reasoning, ornith:35b, ornith:13b, qwen2.5:3b)
|
|
- `system_prompt`: System instruction for the model
|
|
- `user_prompt` (required): User message to send
|
|
- `temperature`: Sampling temperature (0.0-1.0, default 0.7)
|
|
- `max_tokens`: Maximum output tokens
|
|
- `auth_token` (optional): JWT token for authenticated endpoints (propagates as Authorization: Bearer header)
|
|
|
|
**Backend Implementation:**
|
|
The LLMInferenceActivity in the workflows backend automatically:
|
|
1. Substitutes template variables (e.g., `{{ previous_output.code }}`)
|
|
2. Calls `/v1/chat/completions` with the resolved prompt
|
|
3. Returns the LLM response as activity output
|
|
4. Retries on transient failures (up to 3 attempts)
|
|
5. Timeouts after 120 seconds
|
|
|
|
**Output:**
|
|
```json
|
|
{
|
|
"response": "The code has several security vulnerabilities...",
|
|
"model": "reasoning",
|
|
"stop_reason": "stop_sequence",
|
|
"tokens_used": 450
|
|
}
|
|
```
|
|
|
|
**Supported Models:**
|
|
- `reasoning` — DeepSeek-R1-Distill (best for complex analysis)
|
|
- `ornith:35b` — Ollama 35B
|
|
- `ornith:13b` — Ollama 13B
|
|
- `qwen2.5:3b` — Qwen 2.5 3B
|
|
|
|
---
|
|
|
|
### LLMBatchInferenceActivity
|
|
|
|
Multiple-prompt LLM inference (sequential processing).
|
|
|
|
**Workflow Definition:**
|
|
```json
|
|
{
|
|
"id": "llm-batch-1",
|
|
"type": "llm-batch-inference",
|
|
"label": "Batch Code Review",
|
|
"data": {
|
|
"model": "reasoning",
|
|
"system_prompt": "Review each code snippet and provide feedback.",
|
|
"prompts": [
|
|
"Review snippet 1: {{ files[0].content }}",
|
|
"Review snippet 2: {{ files[1].content }}",
|
|
"Review snippet 3: {{ files[2].content }}"
|
|
],
|
|
"auth_token": "{{ user.jwt_token }}"
|
|
}
|
|
}
|
|
```
|
|
|
|
**Fields:**
|
|
- `model` (required): Model ID
|
|
- `system_prompt`: System instruction (same for all prompts)
|
|
- `prompts` (required): List of user prompts to process
|
|
- `temperature`: Sampling temperature (0.0-1.0)
|
|
- `auth_token` (optional): JWT token for authenticated endpoints (propagates as Authorization: Bearer header)
|
|
|
|
**Output:**
|
|
```json
|
|
{
|
|
"responses": [
|
|
"Snippet 1 review...",
|
|
"Snippet 2 review...",
|
|
"Snippet 3 review..."
|
|
],
|
|
"model": "reasoning",
|
|
"errors": []
|
|
}
|
|
```
|
|
|
|
**Typical Use Cases:**
|
|
- Batch code review across multiple files
|
|
- Parallel document summarization
|
|
- Comparative analysis of alternatives
|
|
- Policy compliance checking
|
|
|
|
---
|
|
|
|
### Workflow Integration Examples
|
|
|
|
**1. Code Analysis Workflow**
|
|
```
|
|
Clone Repo → Analyze Code → LLM Security Review → Generate Report → Notify
|
|
```
|
|
|
|
**2. Document Processing**
|
|
```
|
|
Retrieve Documents → Embed + Index → LLM Summarize (batch) → Archive
|
|
```
|
|
|
|
**3. Multi-Stage Review**
|
|
```
|
|
Retrieve Memory → LLM Context Extraction → Route to Activity A/B/C → Notify
|
|
```
|
|
|
|
---
|
|
|
|
### Authentication & Authorization
|
|
|
|
JWT tokens can be passed to LLM inference activities and are automatically propagated to the LLM API endpoint.
|
|
|
|
**Token Flow:**
|
|
```
|
|
Workflow Canvas
|
|
↓ (auth_token field)
|
|
Poimen Workflow Executor
|
|
↓ (passed to LLMInferenceActivity)
|
|
Activity calls LLM client
|
|
↓ (adds "Authorization: Bearer {token}" header)
|
|
homelab-frontend proxy
|
|
↓ (preserves Authorization header)
|
|
LLM Backend (reasoning/ollama/etc)
|
|
↓ (validates token)
|
|
Response returned
|
|
```
|
|
|
|
**Example: Passing User Token from RetrieveMemory Activity**
|
|
```json
|
|
{
|
|
"id": "flow-1",
|
|
"type": "retrieve-memory",
|
|
"label": "Get User Context",
|
|
"data": {...}
|
|
}
|
|
→
|
|
{
|
|
"id": "llm-1",
|
|
"type": "llm-inference",
|
|
"label": "Analyze with User's Token",
|
|
"data": {
|
|
"model": "reasoning",
|
|
"user_prompt": "...",
|
|
"auth_token": "{{ previous_output.user_token }}"
|
|
}
|
|
}
|
|
```
|
|
|
|
**Token Validation:**
|
|
- Tokens are validated by homelab-frontend proxy (checks signature, expiration)
|
|
- Only valid tokens are propagated to LLM backend
|
|
- Invalid tokens result in 401 Unauthorized error
|
|
- Missing token (if required) results in 401 Unauthorized
|
|
|
|
**Note:** The `auth_token` field is optional. If omitted, the LLM API is called without authentication (public endpoints only).
|
|
|
|
---
|
|
|
|
### CanvasReasonerActivity
|
|
|
|
Auto-suggest workflow connections using LLM reasoning. When you drop new activities onto the canvas, this activity analyzes them and suggests logical connections based on input/output compatibility and workflow patterns.
|
|
|
|
**Workflow Definition:**
|
|
```json
|
|
{
|
|
"id": "canvas-reason-1",
|
|
"type": "canvas-reasoner",
|
|
"label": "Auto-Connect Activities",
|
|
"data": {
|
|
"nodes": "{{ workflow.nodes }}",
|
|
"edges": "{{ workflow.edges }}",
|
|
"preserve_existing": true,
|
|
"auth_token": "{{ user.jwt_token }}"
|
|
}
|
|
}
|
|
```
|
|
|
|
**Use Cases:**
|
|
- New nodes added to canvas → automatically suggest connections
|
|
- Validate workflow design → LLM reasoning explains connections
|
|
- Redesign workflow → suggest optimal activity sequence
|
|
- Data flow analysis → ensure proper input/output matching
|
|
|
|
**How It Works:**
|
|
1. Analyzes all node types and their configurations
|
|
2. Reviews existing edges (if preserving)
|
|
3. Uses reasoning model to infer logical connections
|
|
4. Returns suggested edges with confidence score
|
|
5. Includes reasoning explanation
|
|
|
|
**Output Example:**
|
|
```json
|
|
{
|
|
"suggested_edges": [
|
|
{"source": "clone-1", "target": "analyze-1"},
|
|
{"source": "analyze-1", "target": "security-scan-1"},
|
|
{"source": "security-scan-1", "target": "report-1"}
|
|
],
|
|
"reasoning": "Clone repository first, analyze code, perform security scan, generate report. Standard code review workflow.",
|
|
"confidence": 0.92
|
|
}
|
|
```
|
|
|
|
**Fields:**
|
|
- `nodes` (required): Canvas nodes to analyze
|
|
- `edges` (required): Current edges
|
|
- `preserve_existing` (optional, default true): Keep existing edges and only suggest new ones
|
|
- `auth_token` (optional): JWT for LLM reasoning calls
|
|
|
|
**Confidence Scores:**
|
|
- 0.9-1.0: High confidence (common patterns)
|
|
- 0.7-0.9: Medium confidence (reasonable connections)
|
|
- 0.5-0.7: Low confidence (multiple valid approaches)
|
|
- <0.5: Unsure (manual review recommended)
|
|
|
|
**Integration Example:**
|
|
```
|
|
User drops 3 new nodes on canvas
|
|
↓
|
|
Workflow calls CanvasReasonerActivity
|
|
↓
|
|
LLM analyzes schemas: Clone (out: path,commit) → Analyze (in: path) → Report (in: metrics)
|
|
↓
|
|
Compatibility checker validates edges
|
|
↓
|
|
Returns:
|
|
- Suggested edges (Clone → Analyze → Report)
|
|
- Incompatible edges (Report → Approve [terminal sink])
|
|
- Disconnected nodes (if any isolated nodes)
|
|
- User alerts explaining issues
|
|
↓
|
|
Frontend shows:
|
|
✅ Green edges (compatible)
|
|
❌ Red warnings (incompatible)
|
|
🔌 Yellow badges (disconnected)
|
|
↓
|
|
User approves compatible edges, fixes/removes incompatible ones
|
|
```
|
|
|
|
---
|
|
|
|
### Handling Incompatible Connections
|
|
|
|
When two activities **cannot** be connected, the response includes detailed incompatibility data:
|
|
|
|
**Example: Terminal Activity Blocking Connection**
|
|
```json
|
|
{
|
|
"suggested_edges": [
|
|
{"source": "clone-1", "target": "analyze-1"}
|
|
],
|
|
"incompatible_edges": [
|
|
{
|
|
"source": "security-scan-1",
|
|
"target": "approve-1",
|
|
"reason": "ApproveWorkflowActivity accepts no inputs (terminal sink activity)",
|
|
"source_needs": "to output: issues, metrics, severity",
|
|
"target_needs": "none (approval only blocks workflow)",
|
|
"suggestion": "ApproveWorkflowActivity must be the final step. Place it after Report generates summary."
|
|
}
|
|
],
|
|
"user_alerts": [
|
|
"⚠️ security-scan-1 → approve-1: ApproveWorkflowActivity is terminal (no inputs). Place it at the end of the workflow."
|
|
]
|
|
}
|
|
```
|
|
|
|
**Example: Type Mismatch**
|
|
```json
|
|
{
|
|
"incompatible_edges": [
|
|
{
|
|
"source": "llm-inference-1",
|
|
"target": "deployment-check-1",
|
|
"reason": "Output type mismatch: string ≠ object",
|
|
"source_needs": "outputs: response (string)",
|
|
"target_needs": "inputs: deployment_plan (object)",
|
|
"suggestion": "Insert LLM transformer node to convert string response → deployment_plan object"
|
|
}
|
|
],
|
|
"user_alerts": [
|
|
"⚠️ llm-inference-1 → deployment-check-1: Type mismatch (string ≠ object). Use LLM transformation node to map outputs."
|
|
]
|
|
}
|
|
```
|
|
|
|
**Example: Disconnected Nodes**
|
|
```json
|
|
{
|
|
"disconnected_nodes": ["security-scan-1", "notify-1"],
|
|
"user_alerts": [
|
|
"🔌 Node 'SecurityScan' has no connections. Connect it or remove from canvas.",
|
|
"🔌 Node 'Notify' has no incoming edges. Check if it should receive data."
|
|
]
|
|
}
|
|
```
|
|
|
|
### Incompatibility Warning Schema
|
|
|
|
```json
|
|
{
|
|
"source": "string - source node ID",
|
|
"target": "string - target node ID",
|
|
"reason": "string - why connection fails",
|
|
"source_needs": "string - what source outputs",
|
|
"target_needs": "string - what target requires",
|
|
"suggestion": "string - how to fix it"
|
|
}
|
|
```
|
|
|
|
### Common Incompatibility Reasons
|
|
|
|
| Reason | Example | Solution |
|
|
|--------|---------|----------|
|
|
| **Terminal Activity** | Notify → CloneRepo | Can't output from sink (terminal) |
|
|
| **Type Mismatch** | string → object | Use LLM transformer node |
|
|
| **No Outputs** | Notification has no outputs | Terminal activities can't be sources |
|
|
| **No Inputs** | Approval has no inputs | Terminal activities can't accept data |
|
|
| **Semantic Mismatch** | Approval → Analysis | Doesn't make logical sense |
|
|
|
|
### Frontend Alert Display
|
|
|
|
**Sidebar UI:**
|
|
```
|
|
🚨 Connection Issues (3)
|
|
|
|
⚠️ CloneRepo-1 → Approve-1
|
|
Terminal sink can't receive inputs
|
|
[Fix] [Ignore] [Remove Node]
|
|
|
|
⚠️ LLMInference-1 → DeploymentCheck-1
|
|
Type mismatch: string → object
|
|
[Add Transformer] [Manual Map]
|
|
|
|
🔌 SecurityScan-1 (isolated)
|
|
No connections detected
|
|
[Connect] [Remove]
|
|
```
|
|
|
|
**Canvas Visual Feedback:**
|
|
- ❌ Incompatible suggested edges appear as **red dashed lines** (don't auto-add)
|
|
- ⚠️ Disconnected nodes show **yellow border** with icon
|
|
- ✅ Compatible edges appear as **green solid lines** (safe to accept)
|
|
|
|
---
|
|
|
|
### Error Handling
|
|
|
|
If LLM inference fails:
|
|
- First activity retry (2-second backoff)
|
|
- Second activity retry (4-second backoff)
|
|
- Third activity retry (8-second backoff)
|
|
- If all retries fail, workflow records error and proceeds to next activity (or fails if terminal)
|
|
|
|
**Common Failure Scenarios:**
|
|
- Network timeout: `connection refused` (retry automatically)
|
|
- Model not found: `unknown model: xyz` (terminal error)
|
|
- Rate limited: HTTP 429 (retry with exponential backoff)
|
|
- Prompt too long: `context length exceeded` (terminal error)
|
|
|
|
---
|
|
|
|
### Performance & Cost
|
|
|
|
- Single prompt inference: ~100-500ms (model-dependent)
|
|
- Batch processing: Serial (not parallel), ~100-500ms per prompt
|
|
- Model inference costs: Free (on-premise Ollama/Reasoning models)
|
|
- Token counting: Provided in response for quota tracking
|
|
|
|
**Optimization Tips:**
|
|
- Use `ornith:13b` or `qwen2.5:3b` for faster inference
|
|
- Use `reasoning` only for complex analysis that needs reasoning
|
|
- Cache frequently-used prompts at workflow level
|
|
- Use batch activity for multiple similar prompts (better throughput)
|
|
|