16 KiB
User Group & Client Identification Logging
Problem
Standard observability only tracks requests globally. Root cause analysis fails when:
- "Chat API fails for 10% of users" → which users? which client type?
- "SQS queue slow for portfolio team" → portfolio admins or agents?
- "Memory ingestion fails in batches" → browser JS or CLI tool?
We need: User group + client type tagged in every log, trace, and metric for surgical debugging.
Enterprise Approaches
Approach 1: OpenTelemetry Baggage + W3C Trace Context
Pattern: Inject user metadata into distributed trace from gateway, propagate via HTTP headers to all upstream services.
Implementation:
// Gateway middleware extracts JWT claims → baggage
func UserBaggageMiddleware(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
// Parse JWT (already done in auth middleware)
claims := r.Context().Value("claims").(jwt.MapClaims)
// Extract user info
userID := claims["sub"].(string)
groups := claims["groups"].([]interface{})
// Create baggage (OTel standard)
ctx := baggage.ContextWithValues(r.Context(),
baggage.NewKeyValueProperty("user_id", userID),
baggage.NewKeyValueProperty("user_group", groups[0].(string)),
baggage.NewKeyValueProperty("client_id", r.Header.Get("X-Client-ID")),
baggage.NewKeyValueProperty("sdk_version", r.Header.Get("User-Agent")),
)
// Baggage auto-propagates in outbound headers (traceparent, baggage)
next.ServeHTTP(w, r.WithContext(ctx))
})
}
Logs:
{
"trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
"span_id": "00f067aa0ba902b7",
"baggage": {
"user_id": "rock",
"user_group": "homelab-admins",
"client_id": "portfolio-web-v1.2.3",
"sdk_version": "LLM-SDK/2.0.1"
},
"message": "chat request routed to llm-serving",
"status": "success"
}
Metrics tags:
llm_requests_total{
user_group="homelab-admins",
client_type="web",
sdk_version="2.0.1",
status="success"
}
Pros:
- ✅ CNCF standard (OTel baggage spec)
- ✅ Propagates automatically via W3C headers
- ✅ Works across Kubernetes, Temporal, Loki
- ✅ SDK integration (Python, JS, Go all support it)
- ✅ No extra plumbing needed
Cons:
- ❌ Baggage size limited (~8KB per W3C spec)
- ❌ Requires OTel SDK in all services (SDK cost)
- ❌ Not all upstream services parse baggage (vLLM, Ollama don't natively)
Approach 2: Structured Logging with User Context Injection
Pattern: Gateway logs include user/client metadata; downstream services receive request ID + log to same correlator (Loki).
Implementation:
// Gateway request logger
func StructuredRequestLogger(req *http.Request) {
claims := req.Context().Value("claims").(jwt.MapClaims)
log := map[string]interface{}{
"request_id": req.Header.Get("X-Request-ID"),
"timestamp": time.Now().UTC(),
"user_id": claims["sub"],
"user_group": claims["groups"].([]interface{})[0],
"user_email": claims["email"],
"org_id": claims["org_id"], // custom claim from Authentik
// Client info
"client_type": parseClientType(req.Header.Get("User-Agent")),
"client_id": req.Header.Get("X-Client-ID"),
"sdk_version": parseSDKVersion(req.Header.Get("User-Agent")),
"client_ip": ClientIP(req),
"client_country": GeoIP(req),
// Request context
"service": "gateway",
"path": req.URL.Path,
"method": req.Method,
"model": req.Header.Get("X-Model"),
"api_version": "v1",
// Performance
"start_time": time.Now(),
}
// Log as JSON to stdout → Loki
jsonLog, _ := json.Marshal(log)
fmt.Println(string(jsonLog))
}
func parseClientType(userAgent string) string {
if strings.Contains(userAgent, "Chrome") { return "browser-chrome" }
if strings.Contains(userAgent, "Safari") { return "browser-safari" }
if strings.Contains(userAgent, "LLM-SDK/") { return "sdk-python" }
if strings.Contains(userAgent, "node-llm/") { return "sdk-node" }
if strings.Contains(userAgent, "curl") { return "cli-curl" }
if strings.Contains(userAgent, "PostmanRuntime") { return "postman" }
return "unknown"
}
Loki query to find impacted group:
{namespace="portfolio"}
| json
| user_group="portfolio-agents"
| client_type="sdk-python"
| status="error"
| stats count() by error_reason
Pros:
- ✅ Simple to implement (just structured logging)
- ✅ Works with existing Loki + Grafana
- ✅ No SDK overhead
- ✅ Full flexibility (any custom claim)
- ✅ Easy to debug (plain JSON)
Cons:
- ❌ Logs only; doesn't tag metrics automatically
- ❌ Needs Loki for log aggregation (additional infra)
- ❌ Can't correlate metrics and logs without additional work
- ❌ Log volume explodes (every request logged)
Approach 3: HTTP Headers Passthrough + Service-Level Tagging
Pattern: Gateway adds custom headers (X-User-Group, X-Client-Type, X-Org-ID); every upstream service logs them.
Implementation:
// Gateway outbound handler
func EnrichUpstreamRequest(upstream string, req *http.Request) *http.Request {
claims := req.Context().Value("claims").(jwt.MapClaims)
// Inject user metadata as headers
req.Header.Set("X-User-ID", claims["sub"].(string))
req.Header.Set("X-User-Group", claims["groups"].([]interface{})[0].(string))
req.Header.Set("X-Org-ID", claims["org_id"].(string))
req.Header.Set("X-Client-Type", parseClientType(req.Header.Get("User-Agent")))
req.Header.Set("X-Request-ID", uuid.New().String()) // existing
req.Header.Set("X-Trace-ID", req.Header.Get("X-Trace-ID")) // OpenTelemetry
return req
}
// Upstream service (vLLM, Temporal, S3) middleware
func LogUserContextMiddleware(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
log := map[string]interface{}{
"request_id": r.Header.Get("X-Request-ID"),
"trace_id": r.Header.Get("X-Trace-ID"),
"user_id": r.Header.Get("X-User-ID"),
"user_group": r.Header.Get("X-User-Group"),
"org_id": r.Header.Get("X-Org-ID"),
"client_type": r.Header.Get("X-Client-Type"),
"service": "llm-serving",
"endpoint": r.URL.Path,
}
// Log at start and end of request
jsonLog, _ := json.Marshal(log)
fmt.Println(string(jsonLog))
next.ServeHTTP(w, r)
})
}
Prometheus labels (emit from upstream service):
llm_requests_total{
user_group="homelab-admins",
org_id="portfolio",
client_type="browser-chrome",
status="success"
}
Pros:
- ✅ Works everywhere (simple HTTP headers)
- ✅ No SDK/baggage overhead
- ✅ Metrics can be tagged per service
- ✅ Compatible with all upstream services (vLLM, Ollama, Temporal, etc.)
- ✅ Easy to add/debug headers
Cons:
- ❌ Manual propagation needed in each service
- ❌ Header limits (~8KB total)
- ❌ Not standard (custom X-* headers)
- ❌ Hard to enforce consistency across teams
Approach 4: Synthetic User Journeys + Client Fingerprinting
Pattern: Portfolio + Poimen + CLI clients → identify as "portfolio-web-users", "poimen-cli-users"; track their entire journey through system.
Implementation:
// Client registration (SDK or browser)
type ClientMetadata struct {
ClientID string // generated once, persisted locally
ClientType string // "portfolio-web", "poimen-cli", "memory-api-client"
ClientVersion string
UserGroup string // from JWT
Org string
}
// Gateway: extract or generate client fingerprint
func GetClientFingerprint(req *http.Request) string {
clientID := req.Header.Get("X-Client-ID") // SDK provides
if clientID == "" {
// Browser: fingerprint from User-Agent + IP + TLS
clientID = Hash(
req.Header.Get("User-Agent"),
ClientIP(req),
req.TLS.PeerCertificates[0].SerialNumber,
)
}
return clientID
}
// Tag all downstream operations with client fingerprint
func RecordClientJourney(clientID, userGroup, clientType string, operation string, status string) {
// Prometheus
clientJourneyDuration.WithLabelValues(clientID, userGroup, clientType, operation, status).Observe(duration)
// Loki (full request chain)
log := map[string]interface{}{
"client_fingerprint": clientID,
"user_group": userGroup,
"client_type": clientType,
"operation": operation,
"status": status,
"timestamp": time.Now(),
}
LogToLoki(log)
}
Query to find impacted users:
# Alert: "Chat API failures for portfolio-web-users"
(
rate(llm_requests_total{client_type="portfolio-web", status="error"}[5m]) > 0.1
)
Grafana dashboard:
- Heatmap: client_type × user_group vs latency
- Time series: error rate by client_fingerprint
- Table: top 10 impacted client IPs/fingerprints
Pros:
- ✅ Surgical root cause ("portfolio-cli-v2.1 broken for homelab-admins")
- ✅ Can replay exact user journey from logs
- ✅ Cohort-based alerting ("if 5% of portfolio-web fail, alert")
- ✅ User behavior analysis (which clients call what)
Cons:
- ❌ Privacy concerns (fingerprinting users)
- ❌ Complex state management (client IDs, persistent storage)
- ❌ High cardinality (can explode metrics cardinality if not capped)
- ❌ Requires SDK integration
Approach 5: Event Sourcing + User Action Audit Log
Pattern: Every operation → immutable audit log entry with user, client, action, outcome. Query for "all portfolio-web users who failed chat between 2-3pm".
Implementation:
type AuditEvent struct {
EventID string // unique per operation
Timestamp time.Time
RequestID string // correlation ID
// Who
UserID string
UserGroup string
Email string
// What
Action string // "chat_completion", "sqs_send", "memory_ingest"
Service string // "llm-serving", "sqs", "memory"
ResourceID string // model name, queue name, etc.
// How
ClientType string // "browser-chrome", "sdk-python", "cli-curl"
ClientID string
ClientVersion string
ClientIP string
// Result
Status string // "success", "error", "timeout"
ErrorCode string
ErrorMessage string
Duration int64 // milliseconds
// Context
OrgID string
ProjectID string
Metadata map[string]interface{}
}
// Emit audit event (to Postgres/Clickhouse)
func EmitAuditEvent(ctx context.Context, event AuditEvent) {
db.InsertAuditEvent(event)
// Also send to Loki for real-time analysis
LogToLoki(event)
}
// SQL query: "Find all users in portfolio-agents impacted by chat failure"
// SELECT DISTINCT user_id, email FROM audit_events
// WHERE user_group='portfolio-agents'
// AND action='chat_completion'
// AND status='error'
// AND timestamp > NOW() - INTERVAL '1 hour'
// AND client_type IN ('browser-chrome', 'sdk-python')
Grafana query (using Loki or Clickhouse):
SELECT
user_group,
client_type,
COUNT(*) as total_requests,
SUM(CASE WHEN status='error' THEN 1 ELSE 0 END) as failures,
SUM(CASE WHEN status='error' THEN 1 ELSE 0 END) * 100.0 / COUNT(*) as error_rate
FROM audit_events
WHERE timestamp > NOW() - INTERVAL '1 hour'
GROUP BY user_group, client_type
ORDER BY error_rate DESC
Pros:
- ✅ Immutable audit trail (compliance, forensics)
- ✅ Can answer "Who was affected and how?" precisely
- ✅ Excellent for root cause (filter by user_group + timestamp + client_type)
- ✅ Supports RBAC/compliance reporting
- ✅ Works with both logs and metrics
Cons:
- ❌ Storage heavy (every operation stored)
- ❌ Requires additional data store (Postgres/Clickhouse)
- ❌ Latency for complex queries (especially on large datasets)
- ❌ Schema evolution complexity
Recommended Hybrid Approach for Your Setup
Combine Approaches 1 + 3 + 5:
┌─────────────────────────────────────────────────────────────┐
│ Gateway (homelab-frontend) │
│ │
│ 1. Parse JWT → extract user_id, user_group, org_id │
│ 2. Add OTel baggage (Approach 1) │
│ 3. Add X-* headers (Approach 3) │
│ 4. Emit audit event to Loki + Postgres (Approach 5) │
│ │
│ Headers sent upstream: │
│ X-Request-ID (existing) │
│ X-Trace-ID (OTel) │
│ traceparent, baggage (OTel) │
│ X-User-ID, X-User-Group, X-Org-ID │
│ X-Client-Type, X-Client-ID │
└─────────────────────────────────────────────────────────────┘
│
├──→ llm-serving (vLLM)
│ - Log headers
│ - Emit metrics {user_group, client_type}
│ - Propagate X-Request-ID
│
├──→ sqs (Temporal)
│ - Log headers
│ - Emit metrics {user_group, org_id}
│
└──→ memory (Poimen)
- Log headers
- Emit metrics {user_group, client_type}
└──→ Loki (structured logs)
└──→ Prometheus (metrics with user/client tags)
└──→ Postgres/Clickhouse (audit events)
Workflow:
- User calls gateway with Bearer token + User-Agent
- Gateway parses JWT → creates OTel baggage + X-* headers
- Gateway logs audit event (user_id, user_group, client_type, action, status)
- Headers propagated to upstream services
- Root cause query:
- Metrics:
llm_requests_total{user_group="portfolio-agents", client_type="sdk-python", status="error"} - Logs (Loki):
{namespace="portfolio"} | json | user_group="portfolio-agents" | client_type="sdk-python" | status="error" - Audit trail (Postgres):
SELECT * FROM audit_events WHERE user_group='portfolio-agents' AND client_type='sdk-python' AND timestamp>'2026-09-04'
- Metrics:
Implementation Roadmap
| Phase | What | Where | Effort |
|---|---|---|---|
| 1 | JWT claims extraction | internal/auth/jwt.go |
1 day |
| 2 | OTel baggage middleware | internal/middleware/baggage.go |
1 day |
| 3 | X-* header propagation | internal/proxy/proxy.go |
0.5 day |
| 4 | Structured logging (Loki) | internal/logging/structured.go |
1 day |
| 5 | Audit events (Postgres) | internal/audit/events.go + schema |
2 days |
| 6 | Prometheus labels per service | internal/observability/*.go |
1 day |
| 7 | Grafana dashboards | k8s/grafana-dashboards/ |
1 day |
Which Approach for Your Use Case?
User Group Identification:
- Use OTel Baggage (Approach 1) — auto-propagates
- Fallback to X- headers* (Approach 3) if OTel not available
Client Type Detection:
- Parse User-Agent in gateway
- Tag as:
browser-chrome,browser-safari,sdk-python,sdk-node,cli-curl,postman,unknown
Root Cause (Surgical Debugging):
- Use Audit Events (Approach 5) with Postgres
- Query: "Who from portfolio-agents using sdk-python failed in last hour?"
High-Cardinality Metrics:
- Avoid tagging by individual user_id (cardinality explosion)
- Use: user_group, org_id, client_type, sdk_version (bounded)
Real-Time Alerting:
- Prometheus metrics with
{user_group, client_type, status}labels - Alert: "error_rate > 5% for user_group='portfolio-agents' for 5min"