388 lines
12 KiB
Markdown
388 lines
12 KiB
Markdown
# Transaction-Based Observability Plan
|
|
|
|
## Overview
|
|
|
|
Implement transaction-driven observability with correlation ID propagation across all gateway services (LLM, SQS, Temporal, Memory, S3, IAM). Pattern inspired by Amazon's SQS-based transaction tracing.
|
|
|
|
## Architecture
|
|
|
|
```
|
|
Client Request
|
|
│
|
|
▼
|
|
┌─────────────────────────────────────────────────────────┐
|
|
│ Gateway │
|
|
│ ┌─────────────────────────────────────────────────┐ │
|
|
│ │ 1. Generate/extract X-Request-ID │ │
|
|
│ │ 2. Start timer │ │
|
|
│ │ 3. Validate JWT │ │
|
|
│ │ 4. Route to upstream │ │
|
|
│ │ 5. Stream response │ │
|
|
│ │ 6. Emit metrics (request_id, model, status) │ │
|
|
│ │ 7. Log with correlation ID │ │
|
|
│ └─────────────────────────────────────────────────┘ │
|
|
│ │ │
|
|
│ ▼ │
|
|
│ ┌──────────────────┐ │
|
|
│ │ /metrics │ ◄── Prometheus scrape │
|
|
│ │ (RED + tokens) │ │
|
|
│ └──────────────────┘ │
|
|
└─────────────────────────────────────────────────────────┘
|
|
│
|
|
▼
|
|
┌─────────────────────┐
|
|
│ LLM Upstream │
|
|
│ (vLLM/Ollama/TEI) │
|
|
│ X-Request-ID passed │
|
|
└─────────────────────┘
|
|
```
|
|
|
|
## Metrics
|
|
|
|
### Request Metrics (RED)
|
|
|
|
```prometheus
|
|
# Rate - requests per second
|
|
llm_requests_total{model="reasoning", status="success|error", error_type="none|timeout|upstream|auth"}
|
|
|
|
# Errors - error rate by type
|
|
llm_request_errors_total{model="reasoning", error_type="timeout|upstream_5xx|auth_failed|invalid_request"}
|
|
|
|
# Duration - latency histograms
|
|
llm_request_duration_seconds{model="reasoning", phase="total|ttft|generation"}
|
|
```
|
|
|
|
### Token Metrics
|
|
|
|
```prometheus
|
|
# Token throughput
|
|
llm_tokens_total{model="reasoning", direction="prompt|completion"}
|
|
|
|
# Tokens per request (histogram)
|
|
llm_tokens_per_request{model="reasoning", direction="prompt|completion"}
|
|
```
|
|
|
|
### Connection Metrics
|
|
|
|
```prometheus
|
|
# Active requests (gauge)
|
|
llm_active_requests{model="reasoning"}
|
|
|
|
# Upstream health
|
|
llm_upstream_health{upstream="reasoning-predictor.llm-serving:80", status="healthy|unhealthy"}
|
|
```
|
|
|
|
### Service-Specific Metrics
|
|
|
|
#### SQS (Queue Operations)
|
|
|
|
```prometheus
|
|
# Message latency
|
|
sqs_message_latency_seconds{queue="orders", operation="send|receive|delete", status="success|error"}
|
|
|
|
# Message throughput
|
|
sqs_messages_total{queue="orders", operation="send|receive|delete"}
|
|
|
|
# DLQ errors
|
|
sqs_dlq_messages_total{queue="orders", reason="timeout|invalid_format|permission_denied"}
|
|
|
|
# Queue depth (gauge)
|
|
sqs_queue_depth{queue="orders"}
|
|
|
|
# Error rate by type
|
|
sqs_errors_total{queue="orders", error_type="timeout|network|auth|rate_limit"}
|
|
```
|
|
|
|
#### Temporal (Workflow Orchestration)
|
|
|
|
```prometheus
|
|
# Workflow execution time
|
|
temporal_workflow_duration_seconds{workflow_type="ProcessOrder", status="completed|failed|timeout"}
|
|
|
|
# Workflow state transitions
|
|
temporal_workflow_state_transitions_total{workflow_type="ProcessOrder", from="PENDING", to="RUNNING|COMPLETED|FAILED"}
|
|
|
|
# Activity success rate
|
|
temporal_activity_success_rate{activity="CloneRepo|AnalyzeCode|Deploy", status="success|failure|retry"}
|
|
|
|
# Workflow failures by reason
|
|
temporal_workflow_failures_total{workflow_type="ProcessOrder", reason="timeout|activity_failed|network|permission_denied"}
|
|
|
|
# Active workflows (gauge)
|
|
temporal_active_workflows{workflow_type="ProcessOrder"}
|
|
```
|
|
|
|
#### Memory / Poimen (Semantic Search)
|
|
|
|
```prometheus
|
|
# Query latency
|
|
memory_query_latency_seconds{level="L1|L2", status="success|error"}
|
|
|
|
# Ingestion throughput
|
|
memory_documents_ingested_total{project="homelab|portfolio", level="L1|L2"}
|
|
|
|
# Rerank scoring
|
|
memory_rerank_score{query_type="similarity|relevance", percentile="p50|p95|p99"}
|
|
|
|
# Vector DB errors
|
|
memory_errors_total{operation="query|ingest", error_type="timeout|invalid_query|capacity_exceeded"}
|
|
|
|
# Cache hits
|
|
memory_cache_hits_total{level="L1|L2"}
|
|
```
|
|
|
|
#### S3 (Object Storage)
|
|
|
|
```prometheus
|
|
# Upload/download bandwidth
|
|
s3_bytes_transferred{operation="put|get", bucket="data|backups", status="success|error"}
|
|
|
|
# Object operations latency
|
|
s3_operation_duration_seconds{operation="put|get|delete|list", bucket="data"}
|
|
|
|
# Multipart upload failures
|
|
s3_multipart_failures_total{bucket="data", reason="timeout|part_mismatch|abort"}
|
|
|
|
# Bucket space usage (gauge)
|
|
s3_bucket_size_bytes{bucket="data|backups"}
|
|
|
|
# Error rate
|
|
s3_errors_total{bucket="data", error_type="network|auth|timeout|quota_exceeded"}
|
|
```
|
|
|
|
#### IAM (Identity & Access Management)
|
|
|
|
```prometheus
|
|
# User/role mutations
|
|
iam_mutations_total{operation="create_user|update_user|delete_user|add_role|remove_role", status="success|failed"}
|
|
|
|
# Permission cache
|
|
iam_permission_cache_hits_total{resource="user|role|group"}
|
|
|
|
# Auth flow latency
|
|
iam_auth_latency_seconds{flow="client_credentials|authorization_code|refresh", status="success|failed"}
|
|
|
|
# JWT validation
|
|
iam_jwt_validations_total{result="valid|expired|invalid_signature|insufficient_permissions"}
|
|
|
|
# Error rate by type
|
|
iam_errors_total{operation="list_users|get_user|create_user", error_type="auth_failed|not_found|permission_denied|network"}
|
|
```
|
|
|
|
#### Gateway (All Services)
|
|
|
|
```prometheus
|
|
# Request rate by service
|
|
gateway_requests_total{service="llm|sqs|workflow|memory|s3|iam", status="2xx|4xx|5xx"}
|
|
|
|
# Latency by service
|
|
gateway_request_duration_seconds{service="llm|sqs|workflow|memory|s3|iam", phase="total|auth|upstream"}
|
|
|
|
# Correlation ID tracking
|
|
gateway_correlation_id_errors_total{reason="missing|invalid_format|timeout"}
|
|
```
|
|
|
|
## Correlation ID Flow
|
|
|
|
### Request ID Generation
|
|
|
|
```go
|
|
func RequestIDMiddleware(next http.Handler) http.Handler {
|
|
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
|
requestID := r.Header.Get("X-Request-ID")
|
|
if requestID == "" {
|
|
requestID = uuid.New().String()
|
|
}
|
|
|
|
// Set in context for logging
|
|
ctx := context.WithValue(r.Context(), RequestIDKey, requestID)
|
|
|
|
// Echo back to client
|
|
w.Header().Set("X-Request-ID", requestID)
|
|
|
|
// Propagate to upstream
|
|
r.Header.Set("X-Request-ID", requestID)
|
|
|
|
next.ServeHTTP(w, r.WithContext(ctx))
|
|
})
|
|
}
|
|
```
|
|
|
|
### Structured Logging
|
|
|
|
```json
|
|
{
|
|
"timestamp": "2026-09-01T12:00:00Z",
|
|
"level": "info",
|
|
"message": "request completed",
|
|
"request_id": "abc-123-def",
|
|
"model": "reasoning",
|
|
"status": "success",
|
|
"duration_ms": 1523,
|
|
"tokens_prompt": 150,
|
|
"tokens_completion": 89,
|
|
"user_id": "rock",
|
|
"client_ip": "192.168.1.100"
|
|
}
|
|
```
|
|
|
|
Error case:
|
|
```json
|
|
{
|
|
"timestamp": "2026-09-01T12:00:00Z",
|
|
"level": "error",
|
|
"message": "request failed",
|
|
"request_id": "abc-123-def",
|
|
"model": "reasoning",
|
|
"status": "error",
|
|
"error_type": "upstream_timeout",
|
|
"error_detail": "context deadline exceeded",
|
|
"duration_ms": 30000,
|
|
"action": "LLM_INFERENCE",
|
|
"exception": "UPSTREAM_TIMEOUT"
|
|
}
|
|
```
|
|
|
|
## Drill-Down Query Flow
|
|
|
|
### 1. Dashboard Alert
|
|
Grafana shows spike in error rate:
|
|
```promql
|
|
rate(llm_request_errors_total{error_type="upstream_timeout"}[5m]) > 0.1
|
|
```
|
|
|
|
### 2. Error Breakdown
|
|
Click through to see error distribution:
|
|
```promql
|
|
sum by (error_type, model) (rate(llm_request_errors_total[5m]))
|
|
```
|
|
|
|
### 3. Find Affected Requests
|
|
Loki query for specific error type:
|
|
```logql
|
|
{namespace="api", app="api-gateway"}
|
|
| json
|
|
| error_type="upstream_timeout"
|
|
| line_format "{{.request_id}} {{.model}} {{.duration_ms}}ms"
|
|
```
|
|
|
|
### 4. Trace Single Request
|
|
Drill into specific request_id:
|
|
```logql
|
|
{namespace=~"api|llm-serving"} |= "abc-123-def"
|
|
```
|
|
|
|
### 5. Correlate with Upstream
|
|
If vLLM/Ollama logs also include X-Request-ID:
|
|
```logql
|
|
{namespace="llm-serving"} |= "abc-123-def"
|
|
```
|
|
|
|
## Implementation Tasks
|
|
|
|
### Phase 1: Metrics Foundation
|
|
- [ ] Create `internal/observability/metrics.go`
|
|
- Define Prometheus metrics (counters, histograms, gauges)
|
|
- Register with default registry
|
|
- [ ] Add `/metrics` endpoint to server
|
|
- [ ] Create ServiceMonitor for gateway
|
|
|
|
### Phase 2: Request ID Middleware
|
|
- [ ] Create `internal/observability/requestid.go`
|
|
- Generate/extract X-Request-ID
|
|
- Store in context
|
|
- Propagate to upstream
|
|
- [ ] Wire into server router
|
|
|
|
### Phase 3: Instrumentation
|
|
- [ ] Instrument proxy handler
|
|
- Record request start/end times
|
|
- Count tokens from response (parse SSE stream)
|
|
- Emit metrics on completion
|
|
- [ ] Update structured logging
|
|
- Add request_id to all log entries
|
|
- Add action/exception pattern for errors
|
|
|
|
### Phase 4: Dashboard
|
|
- [ ] Update `llm-frontend` dashboard
|
|
- Request rate panel
|
|
- Error rate by type panel
|
|
- Latency percentiles (p50, p95, p99)
|
|
- Token throughput panel
|
|
- Active requests gauge
|
|
- [ ] Add drill-down links to Loki
|
|
|
|
### Phase 5: Alerting
|
|
- [ ] Create PrometheusRule for LLM alerts
|
|
- High error rate
|
|
- High latency (p99 > threshold)
|
|
- Upstream unhealthy
|
|
- Token rate anomaly
|
|
|
|
## File Structure
|
|
|
|
```
|
|
internal/observability/
|
|
├── metrics.go # Prometheus metric definitions
|
|
├── requestid.go # Request ID middleware
|
|
├── instrumentation.go # Metric recording helpers
|
|
└── logging.go # Structured logging with correlation
|
|
|
|
k8s/
|
|
├── servicemonitor.yaml # Prometheus scrape config
|
|
└── prometheusrule.yaml # Alert rules
|
|
|
|
k8s/infra/monitoring/dashboards/
|
|
└── llm-frontend.yaml # Updated dashboard
|
|
```
|
|
|
|
## Metric Labels Convention
|
|
|
|
| Label | Values | Description |
|
|
|-------|--------|-------------|
|
|
| `model` | reasoning, ornith:35b, qwen2.5:3b-instruct, etc. | LLM model name |
|
|
| `status` | success, error | Request outcome |
|
|
| `error_type` | none, timeout, upstream_5xx, auth_failed, invalid_request, rate_limited | Error category |
|
|
| `phase` | total, ttft, generation | Latency measurement phase |
|
|
| `direction` | prompt, completion | Token direction |
|
|
| `upstream` | host:port | Upstream address |
|
|
|
|
## Action-Exception Pattern
|
|
|
|
For structured error tracking (Amazon-style):
|
|
|
|
| Action | Exception | Metric |
|
|
|--------|-----------|--------|
|
|
| `LLM_INFERENCE` | `UPSTREAM_TIMEOUT` | `llm_action_exception_total{action="LLM_INFERENCE", exception="UPSTREAM_TIMEOUT"}` |
|
|
| `LLM_INFERENCE` | `UPSTREAM_5XX` | `llm_action_exception_total{action="LLM_INFERENCE", exception="UPSTREAM_5XX"}` |
|
|
| `JWT_VALIDATION` | `TOKEN_EXPIRED` | `llm_action_exception_total{action="JWT_VALIDATION", exception="TOKEN_EXPIRED"}` |
|
|
| `JWT_VALIDATION` | `INVALID_SIGNATURE` | `llm_action_exception_total{action="JWT_VALIDATION", exception="INVALID_SIGNATURE"}` |
|
|
| `BODY_PARSE` | `INVALID_JSON` | `llm_action_exception_total{action="BODY_PARSE", exception="INVALID_JSON"}` |
|
|
| `MODEL_DISPATCH` | `UNKNOWN_MODEL` | `llm_action_exception_total{action="MODEL_DISPATCH", exception="UNKNOWN_MODEL"}` |
|
|
|
|
Query pattern:
|
|
```promql
|
|
# All exceptions for an action
|
|
sum by (exception) (rate(llm_action_exception_total{action="LLM_INFERENCE"}[5m]))
|
|
|
|
# Specific exception rate
|
|
rate(llm_action_exception_total{action="LLM_INFERENCE", exception="UPSTREAM_TIMEOUT"}[5m])
|
|
```
|
|
|
|
## Dependencies
|
|
|
|
```go
|
|
// go.mod additions
|
|
require (
|
|
github.com/prometheus/client_golang v1.19.0
|
|
github.com/google/uuid v1.6.0
|
|
)
|
|
```
|
|
|
|
## References
|
|
|
|
- [Prometheus Go client](https://github.com/prometheus/client_golang)
|
|
- [OpenTelemetry metrics](https://opentelemetry.io/docs/instrumentation/go/manual/#metrics)
|
|
- [Grafana Loki LogQL](https://grafana.com/docs/loki/latest/query/)
|
|
- [RED method](https://www.weave.works/blog/the-red-method-key-metrics-for-microservices-architecture/)
|