docs: Extend observability to cover all gateway services (SQS, Temporal, Memory, S3, IAM)
CI / Vet, test, build (push) Successful in 3m41s
CI / Build and push image (push) Successful in 1m18s

This commit is contained in:
Admin Bot
2026-09-03 21:32:52 -07:00
parent e1ab49dcdc
commit e4bbad5ad8
+387
View File
@@ -0,0 +1,387 @@
# Transaction-Based Observability Plan
## Overview
Implement transaction-driven observability with correlation ID propagation across all gateway services (LLM, SQS, Temporal, Memory, S3, IAM). Pattern inspired by Amazon's SQS-based transaction tracing.
## Architecture
```
Client Request
┌─────────────────────────────────────────────────────────┐
│ Gateway │
│ ┌─────────────────────────────────────────────────┐ │
│ │ 1. Generate/extract X-Request-ID │ │
│ │ 2. Start timer │ │
│ │ 3. Validate JWT │ │
│ │ 4. Route to upstream │ │
│ │ 5. Stream response │ │
│ │ 6. Emit metrics (request_id, model, status) │ │
│ │ 7. Log with correlation ID │ │
│ └─────────────────────────────────────────────────┘ │
│ │ │
│ ▼ │
│ ┌──────────────────┐ │
│ │ /metrics │ ◄── Prometheus scrape │
│ │ (RED + tokens) │ │
│ └──────────────────┘ │
└─────────────────────────────────────────────────────────┘
┌─────────────────────┐
│ LLM Upstream │
│ (vLLM/Ollama/TEI) │
│ X-Request-ID passed │
└─────────────────────┘
```
## Metrics
### Request Metrics (RED)
```prometheus
# Rate - requests per second
llm_requests_total{model="reasoning", status="success|error", error_type="none|timeout|upstream|auth"}
# Errors - error rate by type
llm_request_errors_total{model="reasoning", error_type="timeout|upstream_5xx|auth_failed|invalid_request"}
# Duration - latency histograms
llm_request_duration_seconds{model="reasoning", phase="total|ttft|generation"}
```
### Token Metrics
```prometheus
# Token throughput
llm_tokens_total{model="reasoning", direction="prompt|completion"}
# Tokens per request (histogram)
llm_tokens_per_request{model="reasoning", direction="prompt|completion"}
```
### Connection Metrics
```prometheus
# Active requests (gauge)
llm_active_requests{model="reasoning"}
# Upstream health
llm_upstream_health{upstream="reasoning-predictor.llm-serving:80", status="healthy|unhealthy"}
```
### Service-Specific Metrics
#### SQS (Queue Operations)
```prometheus
# Message latency
sqs_message_latency_seconds{queue="orders", operation="send|receive|delete", status="success|error"}
# Message throughput
sqs_messages_total{queue="orders", operation="send|receive|delete"}
# DLQ errors
sqs_dlq_messages_total{queue="orders", reason="timeout|invalid_format|permission_denied"}
# Queue depth (gauge)
sqs_queue_depth{queue="orders"}
# Error rate by type
sqs_errors_total{queue="orders", error_type="timeout|network|auth|rate_limit"}
```
#### Temporal (Workflow Orchestration)
```prometheus
# Workflow execution time
temporal_workflow_duration_seconds{workflow_type="ProcessOrder", status="completed|failed|timeout"}
# Workflow state transitions
temporal_workflow_state_transitions_total{workflow_type="ProcessOrder", from="PENDING", to="RUNNING|COMPLETED|FAILED"}
# Activity success rate
temporal_activity_success_rate{activity="CloneRepo|AnalyzeCode|Deploy", status="success|failure|retry"}
# Workflow failures by reason
temporal_workflow_failures_total{workflow_type="ProcessOrder", reason="timeout|activity_failed|network|permission_denied"}
# Active workflows (gauge)
temporal_active_workflows{workflow_type="ProcessOrder"}
```
#### Memory / Poimen (Semantic Search)
```prometheus
# Query latency
memory_query_latency_seconds{level="L1|L2", status="success|error"}
# Ingestion throughput
memory_documents_ingested_total{project="homelab|portfolio", level="L1|L2"}
# Rerank scoring
memory_rerank_score{query_type="similarity|relevance", percentile="p50|p95|p99"}
# Vector DB errors
memory_errors_total{operation="query|ingest", error_type="timeout|invalid_query|capacity_exceeded"}
# Cache hits
memory_cache_hits_total{level="L1|L2"}
```
#### S3 (Object Storage)
```prometheus
# Upload/download bandwidth
s3_bytes_transferred{operation="put|get", bucket="data|backups", status="success|error"}
# Object operations latency
s3_operation_duration_seconds{operation="put|get|delete|list", bucket="data"}
# Multipart upload failures
s3_multipart_failures_total{bucket="data", reason="timeout|part_mismatch|abort"}
# Bucket space usage (gauge)
s3_bucket_size_bytes{bucket="data|backups"}
# Error rate
s3_errors_total{bucket="data", error_type="network|auth|timeout|quota_exceeded"}
```
#### IAM (Identity & Access Management)
```prometheus
# User/role mutations
iam_mutations_total{operation="create_user|update_user|delete_user|add_role|remove_role", status="success|failed"}
# Permission cache
iam_permission_cache_hits_total{resource="user|role|group"}
# Auth flow latency
iam_auth_latency_seconds{flow="client_credentials|authorization_code|refresh", status="success|failed"}
# JWT validation
iam_jwt_validations_total{result="valid|expired|invalid_signature|insufficient_permissions"}
# Error rate by type
iam_errors_total{operation="list_users|get_user|create_user", error_type="auth_failed|not_found|permission_denied|network"}
```
#### Gateway (All Services)
```prometheus
# Request rate by service
gateway_requests_total{service="llm|sqs|workflow|memory|s3|iam", status="2xx|4xx|5xx"}
# Latency by service
gateway_request_duration_seconds{service="llm|sqs|workflow|memory|s3|iam", phase="total|auth|upstream"}
# Correlation ID tracking
gateway_correlation_id_errors_total{reason="missing|invalid_format|timeout"}
```
## Correlation ID Flow
### Request ID Generation
```go
func RequestIDMiddleware(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
requestID := r.Header.Get("X-Request-ID")
if requestID == "" {
requestID = uuid.New().String()
}
// Set in context for logging
ctx := context.WithValue(r.Context(), RequestIDKey, requestID)
// Echo back to client
w.Header().Set("X-Request-ID", requestID)
// Propagate to upstream
r.Header.Set("X-Request-ID", requestID)
next.ServeHTTP(w, r.WithContext(ctx))
})
}
```
### Structured Logging
```json
{
"timestamp": "2026-09-01T12:00:00Z",
"level": "info",
"message": "request completed",
"request_id": "abc-123-def",
"model": "reasoning",
"status": "success",
"duration_ms": 1523,
"tokens_prompt": 150,
"tokens_completion": 89,
"user_id": "rock",
"client_ip": "192.168.1.100"
}
```
Error case:
```json
{
"timestamp": "2026-09-01T12:00:00Z",
"level": "error",
"message": "request failed",
"request_id": "abc-123-def",
"model": "reasoning",
"status": "error",
"error_type": "upstream_timeout",
"error_detail": "context deadline exceeded",
"duration_ms": 30000,
"action": "LLM_INFERENCE",
"exception": "UPSTREAM_TIMEOUT"
}
```
## Drill-Down Query Flow
### 1. Dashboard Alert
Grafana shows spike in error rate:
```promql
rate(llm_request_errors_total{error_type="upstream_timeout"}[5m]) > 0.1
```
### 2. Error Breakdown
Click through to see error distribution:
```promql
sum by (error_type, model) (rate(llm_request_errors_total[5m]))
```
### 3. Find Affected Requests
Loki query for specific error type:
```logql
{namespace="api", app="api-gateway"}
| json
| error_type="upstream_timeout"
| line_format "{{.request_id}} {{.model}} {{.duration_ms}}ms"
```
### 4. Trace Single Request
Drill into specific request_id:
```logql
{namespace=~"api|llm-serving"} |= "abc-123-def"
```
### 5. Correlate with Upstream
If vLLM/Ollama logs also include X-Request-ID:
```logql
{namespace="llm-serving"} |= "abc-123-def"
```
## Implementation Tasks
### Phase 1: Metrics Foundation
- [ ] Create `internal/observability/metrics.go`
- Define Prometheus metrics (counters, histograms, gauges)
- Register with default registry
- [ ] Add `/metrics` endpoint to server
- [ ] Create ServiceMonitor for gateway
### Phase 2: Request ID Middleware
- [ ] Create `internal/observability/requestid.go`
- Generate/extract X-Request-ID
- Store in context
- Propagate to upstream
- [ ] Wire into server router
### Phase 3: Instrumentation
- [ ] Instrument proxy handler
- Record request start/end times
- Count tokens from response (parse SSE stream)
- Emit metrics on completion
- [ ] Update structured logging
- Add request_id to all log entries
- Add action/exception pattern for errors
### Phase 4: Dashboard
- [ ] Update `llm-frontend` dashboard
- Request rate panel
- Error rate by type panel
- Latency percentiles (p50, p95, p99)
- Token throughput panel
- Active requests gauge
- [ ] Add drill-down links to Loki
### Phase 5: Alerting
- [ ] Create PrometheusRule for LLM alerts
- High error rate
- High latency (p99 > threshold)
- Upstream unhealthy
- Token rate anomaly
## File Structure
```
internal/observability/
├── metrics.go # Prometheus metric definitions
├── requestid.go # Request ID middleware
├── instrumentation.go # Metric recording helpers
└── logging.go # Structured logging with correlation
k8s/
├── servicemonitor.yaml # Prometheus scrape config
└── prometheusrule.yaml # Alert rules
k8s/infra/monitoring/dashboards/
└── llm-frontend.yaml # Updated dashboard
```
## Metric Labels Convention
| Label | Values | Description |
|-------|--------|-------------|
| `model` | reasoning, ornith:35b, qwen2.5:3b-instruct, etc. | LLM model name |
| `status` | success, error | Request outcome |
| `error_type` | none, timeout, upstream_5xx, auth_failed, invalid_request, rate_limited | Error category |
| `phase` | total, ttft, generation | Latency measurement phase |
| `direction` | prompt, completion | Token direction |
| `upstream` | host:port | Upstream address |
## Action-Exception Pattern
For structured error tracking (Amazon-style):
| Action | Exception | Metric |
|--------|-----------|--------|
| `LLM_INFERENCE` | `UPSTREAM_TIMEOUT` | `llm_action_exception_total{action="LLM_INFERENCE", exception="UPSTREAM_TIMEOUT"}` |
| `LLM_INFERENCE` | `UPSTREAM_5XX` | `llm_action_exception_total{action="LLM_INFERENCE", exception="UPSTREAM_5XX"}` |
| `JWT_VALIDATION` | `TOKEN_EXPIRED` | `llm_action_exception_total{action="JWT_VALIDATION", exception="TOKEN_EXPIRED"}` |
| `JWT_VALIDATION` | `INVALID_SIGNATURE` | `llm_action_exception_total{action="JWT_VALIDATION", exception="INVALID_SIGNATURE"}` |
| `BODY_PARSE` | `INVALID_JSON` | `llm_action_exception_total{action="BODY_PARSE", exception="INVALID_JSON"}` |
| `MODEL_DISPATCH` | `UNKNOWN_MODEL` | `llm_action_exception_total{action="MODEL_DISPATCH", exception="UNKNOWN_MODEL"}` |
Query pattern:
```promql
# All exceptions for an action
sum by (exception) (rate(llm_action_exception_total{action="LLM_INFERENCE"}[5m]))
# Specific exception rate
rate(llm_action_exception_total{action="LLM_INFERENCE", exception="UPSTREAM_TIMEOUT"}[5m])
```
## Dependencies
```go
// go.mod additions
require (
github.com/prometheus/client_golang v1.19.0
github.com/google/uuid v1.6.0
)
```
## References
- [Prometheus Go client](https://github.com/prometheus/client_golang)
- [OpenTelemetry metrics](https://opentelemetry.io/docs/instrumentation/go/manual/#metrics)
- [Grafana Loki LogQL](https://grafana.com/docs/loki/latest/query/)
- [RED method](https://www.weave.works/blog/the-red-method-key-metrics-for-microservices-architecture/)