2026-08-19 20:52:13 -07:00
# homelab-frontend
2026-08-30 07:29:24 -07:00
Production API gateway for the homelab cluster. Single entry point (`api.riotpiao.com` ) for all services: LLM inference, workflows, queues, memory, and cluster operations.
2026-08-19 20:52:13 -07:00
2026-08-30 07:29:24 -07:00
**Status:** Live in production. Replaced Kong OSS entirely.
2026-08-19 20:52:13 -07:00
2026-08-30 07:29:24 -07:00
---
## Quick Links
- **API Reference:** See [API.md ](API.md ) — how to call every service
- **Base URL:** `https://api.riotpiao.com`
- **Source:** `ssh://git.riotpiao.com:2222/rock/homelab-frontend.git`
---
## Architecture
2026-08-19 20:52:13 -07:00
```
2026-08-30 07:29:24 -07:00
┌─────────────┬──────────────┬────────────┐
│ Browser │ SDK │ CLI │
└──────┬──────┴──────┬───────┴────┬───────┘
│ │ │
└─────────────┼────────────┘
│
HTTPS/TLS
│
┌─────────────┼────────────┐
│ Cloudflare Edge │
│ (DDoS, caching) │
└──────────┬──────────────┘
│
ingress-nginx
(SSL termination)
│
┌─────────────────────────────┐
│ homelab-frontend Gateway │
│ (routing, auth, limits) │
└──────┬───────────────────────┘
│
┌──────┴──────────────────────────────────┐
│ │
/v1/* /workflow /sqs /
(LLM) (Temporal gRPC) (Queues) (X-Service)
│ │ │ │
▼ ▼ ▼ ▼
llm-serving temporal:7233 kmsvc/Kafka IAM, S3
(vLLM, Ollama) (WorkflowService) Memory
(TEI) (gRPC bridge) (poimen)
2026-08-19 20:52:13 -07:00
```
2026-08-30 07:29:24 -07:00
**Design principles:**
- ✅ Single hostname, multiple path prefixes
- ✅ HTTP REST gateway → gRPC Temporal bridge
- ✅ Bearer token auth via Authentik (JWT + RBAC)
- ✅ Streaming unbuffered (SSE, WebSocket)
- ✅ Per-route timeouts & rate limits
- ✅ No cluster credentials held by gateway
2026-08-19 20:52:13 -07:00
2026-08-30 07:29:24 -07:00
---
2026-08-19 20:52:13 -07:00
2026-08-30 07:29:24 -07:00
## Services & Capabilities
2026-08-19 20:52:13 -07:00
2026-08-30 07:29:24 -07:00
| Service | Prefix | Upstream | Status |
|---------|--------|----------|--------|
| **LLM Chat** | `/v1/chat/completions` | llm-serving (vLLM) | ✅ Live |
| **Embeddings** | `/v1/embeddings` | llm-serving (TEI) | ✅ Live |
| **Reranking** | `/v1/rerank` | llm-serving (TEI) | ✅ Live |
| **Workflows** | `/workflow` | Temporal gRPC (7233) | ✅ Live (START, DESCRIBE, SIGNAL, QUERY, etc) |
| **Queues** | `/` + `X-Service: sqs` | kmsvc/Kafka | ⏳ Ready (ServiceAdapter) |
| **Memory** | `/` + `X-Service: memory` | poimen-memory | ✅ Live |
| **IAM** | `/` + `X-Service: iam` | Authentik API | ✅ Live |
| **S3** | `/` + `X-Service: s3` | MinIO | ✅ Live |
2026-08-19 20:52:13 -07:00
2026-08-30 07:29:24 -07:00
---
2026-08-19 20:52:13 -07:00
2026-08-30 07:29:24 -07:00
## How to Use
2026-08-19 20:52:13 -07:00
2026-08-30 07:29:24 -07:00
### 1. Get a token
2026-08-19 20:52:13 -07:00
2026-08-30 07:29:24 -07:00
**Human (OIDC device code):**
```bash
core auth login
export TOKEN = $( cat ~/.cache/talos/authentik_id_token)
```
2026-08-19 20:52:13 -07:00
2026-08-30 07:29:24 -07:00
**Service account (client credentials):**
```bash
core mwinit login --username sa-name --password secret
export TOKEN = $( cat ~/.talos/.riotpiao-auth)
```
2026-08-19 20:52:13 -07:00
2026-08-30 07:29:24 -07:00
### 2. Call any service
2026-08-19 20:52:13 -07:00
2026-08-30 07:29:24 -07:00
**Chat:**
```bash
curl -X POST https://api.riotpiao.com/v1/chat/completions \
-H "Authorization: Bearer $TOKEN " \
-d '{
"model": "reasoning",
"messages": [{"role": "user", "content": "What is 2+2?"}]
}'
```
**Workflow:**
```bash
curl -X POST https://api.riotpiao.com/workflow \
-H "Authorization: Bearer $TOKEN " \
-d '{
"action": "START_WORKFLOW",
"namespace": "default",
"payload": {
"workflow_id": "my-workflow",
"workflow_type": "MyWorkflow",
"task_queue": "default"
}
}'
```
**Memory:**
```bash
curl -X GET https://api.riotpiao.com/ \
-H "Authorization: Bearer $TOKEN " \
-H 'X-Service: memory' \
-H 'X-Resource: query' \
-G --data-urlencode 'query=explain machine learning'
```
**Full examples:** See [API.md ](API.md )
---
## Available Models
### LLM (Chat & Reasoning)
- `reasoning` — DeepSeek-R1-Distill-Qwen-32B (8 concurrent slots)
- `ornith:35b` — Ollama 35B
- `qwen2.5:3b-instruct` — Qwen 2.5 3B
### Embeddings
- `nomic-ai/nomic-embed-text-v2-moe` — Fast, multilingual
### Reranking
- `BAAI/bge-reranker-base` — Document relevance scoring
---
## Authentication
All endpoints (except `/healthz` , `/readyz` ) require:
2026-08-19 20:52:13 -07:00
```
2026-08-30 07:29:24 -07:00
Authorization: Bearer <jwt-token>
```
Tokens validated via Authentik JWKS (runtime fetched, cached, auto-rotated).
**Capabilities** (RBAC):
- `llm:inference` — `/v1/*` chat/embeddings/rerank
- `workflow:execute` — `/workflow` operations
- `memory:read` / `memory:write` — Memory operations
- `sqs:access` — Queue operations
- `s3:access` — S3 operations
- `iam:admin` — User/group management
---
## Error Handling
All errors return RFC 9457 `application/problem+json` :
```json
{
"type" : "https://api.example.com/problems/unknown-model" ,
"title" : "Unknown Model" ,
"status" : 400 ,
"detail" : "Model 'gpt-4' is not available" ,
"valid_models" : [ "reasoning" , "ornith:35b" , ... ]
}
```
**Common status codes:**
- 200 OK
- 400 Bad Request (validation, unknown model)
- 401 Unauthorized (missing/invalid token)
- 403 Forbidden (insufficient capability)
- 404 Not Found (workflow, resource)
- 429 Too Many Requests (rate limit)
- 503 Service Unavailable (backend down)
---
## Rate Limits
| Endpoint | Limit | Retry-After |
|----------|-------|-------------|
| `/v1/chat/completions` | 8 concurrent | Yes |
| `/v1/embeddings` | 10 concurrent | Yes |
| `/v1/rerank` | 10 concurrent | Yes |
| `/workflow` | 100 concurrent | Yes |
Hitting limit returns 429 with `Retry-After` header.
---
## Timeouts
| Endpoint | Connect | Read | Write |
|----------|---------|------|-------|
| `/v1/chat` | 10s | 1h | 1h |
| `/v1/embeddings` | 10s | 10m | 10m |
| `/v1/rerank` | 10s | 10m | 10m |
| `/workflow` | 10s | 30s | 10s |
Client disconnects cancel upstream request immediately (no orphaned slots).
---
## Local Development
Run without cluster, no credentials needed:
```bash
# Build
go build ./cmd/gateway
# Run locally
./gateway
# Test in another terminal
curl http://localhost:8080/healthz
```
Points upstreams at local stubs if not connected to cluster (see `internal/config` ).
---
## Code Layout
```
cmd/gateway/ Server entrypoint
2026-08-19 20:52:13 -07:00
internal/
2026-08-30 07:29:24 -07:00
server/ Router, health checks
proxy/ Reverse proxy, streaming, timeouts
temporal/ Workflow handler + gRPC bridge
serviceadapter/ X-Service dispatcher (CRD-driven)
config/ Route + upstream configuration
auth/ Authentik JWT validation
observability/ Metrics, structured logging
k8s/
configmap.yaml Route definitions
rbac.yaml Service account, roles
deployment.yaml Pod spec
networkpolicy.yaml Ingress/egress rules
testdata/ Fixtures for offline tests
2026-08-19 20:52:13 -07:00
```
2026-08-30 07:29:24 -07:00
---
2026-08-19 20:52:13 -07:00
2026-08-30 07:29:24 -07:00
## Deployment
2026-08-19 20:52:13 -07:00
2026-08-30 07:29:24 -07:00
Deployed to Kubernetes via ArgoCD:
2026-08-19 20:52:13 -07:00
2026-08-30 07:29:24 -07:00
```bash
# Check deployment
kubectl -n api get deployment homelab-frontend
# View logs
kubectl -n api logs -l app = homelab-frontend -f
# Restart
kubectl -n api rollout restart deployment/homelab-frontend
```
Configuration mounted as ConfigMap (`k8s/configmap.yaml` ).
---
## Health Checks
```bash
# Liveness (always succeeds)
curl https://api.riotpiao.com/healthz
# Readiness (waits for config + JWKS)
curl https://api.riotpiao.com/readyz
# List available models
curl https://api.riotpiao.com/v1/models
```
---
## Support
**Issues:** Check pod logs
```bash
kubectl -n api logs deployment/homelab-frontend
```
**Debug config:**
```bash
kubectl -n api get configmap homelab-frontend-config -o yaml
```
**Restart pod:**
```bash
kubectl -n api rollout restart deployment/homelab-frontend
```
**Test endpoint directly:**
```bash
kubectl -n api port-forward svc/homelab-frontend 8080:8080
curl http://localhost:8080/healthz
```
---
## See Also
- [API.md ](API.md ) — Complete API reference with examples
- `internal/` — Source code (handlers, routing, auth)
- `k8s/` — Kubernetes manifests