Files
homelab-frontend/USER_GROUP_TRACING.md
T
Admin Bot d3d5279942
CI / Vet, test, build (push) Failing after 1m33s
CI / Build and push image (push) Skipped
docs: Add user group & client identification logging patterns
2026-09-04 13:50:50 -07:00

16 KiB
Raw Blame History

User Group & Client Identification Logging

Problem

Standard observability only tracks requests globally. Root cause analysis fails when:

  • "Chat API fails for 10% of users" → which users? which client type?
  • "SQS queue slow for portfolio team" → portfolio admins or agents?
  • "Memory ingestion fails in batches" → browser JS or CLI tool?

We need: User group + client type tagged in every log, trace, and metric for surgical debugging.


Enterprise Approaches

Approach 1: OpenTelemetry Baggage + W3C Trace Context

Pattern: Inject user metadata into distributed trace from gateway, propagate via HTTP headers to all upstream services.

Implementation:

// Gateway middleware extracts JWT claims → baggage
func UserBaggageMiddleware(next http.Handler) http.Handler {
    return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
        // Parse JWT (already done in auth middleware)
        claims := r.Context().Value("claims").(jwt.MapClaims)
        
        // Extract user info
        userID := claims["sub"].(string)
        groups := claims["groups"].([]interface{})
        
        // Create baggage (OTel standard)
        ctx := baggage.ContextWithValues(r.Context(),
            baggage.NewKeyValueProperty("user_id", userID),
            baggage.NewKeyValueProperty("user_group", groups[0].(string)),
            baggage.NewKeyValueProperty("client_id", r.Header.Get("X-Client-ID")),
            baggage.NewKeyValueProperty("sdk_version", r.Header.Get("User-Agent")),
        )
        
        // Baggage auto-propagates in outbound headers (traceparent, baggage)
        next.ServeHTTP(w, r.WithContext(ctx))
    })
}

Logs:

{
  "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
  "span_id": "00f067aa0ba902b7",
  "baggage": {
    "user_id": "rock",
    "user_group": "homelab-admins",
    "client_id": "portfolio-web-v1.2.3",
    "sdk_version": "LLM-SDK/2.0.1"
  },
  "message": "chat request routed to llm-serving",
  "status": "success"
}

Metrics tags:

llm_requests_total{
  user_group="homelab-admins",
  client_type="web",
  sdk_version="2.0.1",
  status="success"
}

Pros:

  • CNCF standard (OTel baggage spec)
  • Propagates automatically via W3C headers
  • Works across Kubernetes, Temporal, Loki
  • SDK integration (Python, JS, Go all support it)
  • No extra plumbing needed

Cons:

  • Baggage size limited (~8KB per W3C spec)
  • Requires OTel SDK in all services (SDK cost)
  • Not all upstream services parse baggage (vLLM, Ollama don't natively)

Approach 2: Structured Logging with User Context Injection

Pattern: Gateway logs include user/client metadata; downstream services receive request ID + log to same correlator (Loki).

Implementation:

// Gateway request logger
func StructuredRequestLogger(req *http.Request) {
    claims := req.Context().Value("claims").(jwt.MapClaims)
    
    log := map[string]interface{}{
        "request_id": req.Header.Get("X-Request-ID"),
        "timestamp": time.Now().UTC(),
        "user_id": claims["sub"],
        "user_group": claims["groups"].([]interface{})[0],
        "user_email": claims["email"],
        "org_id": claims["org_id"], // custom claim from Authentik
        
        // Client info
        "client_type": parseClientType(req.Header.Get("User-Agent")),
        "client_id": req.Header.Get("X-Client-ID"),
        "sdk_version": parseSDKVersion(req.Header.Get("User-Agent")),
        "client_ip": ClientIP(req),
        "client_country": GeoIP(req),
        
        // Request context
        "service": "gateway",
        "path": req.URL.Path,
        "method": req.Method,
        "model": req.Header.Get("X-Model"),
        "api_version": "v1",
        
        // Performance
        "start_time": time.Now(),
    }
    
    // Log as JSON to stdout → Loki
    jsonLog, _ := json.Marshal(log)
    fmt.Println(string(jsonLog))
}

func parseClientType(userAgent string) string {
    if strings.Contains(userAgent, "Chrome") { return "browser-chrome" }
    if strings.Contains(userAgent, "Safari") { return "browser-safari" }
    if strings.Contains(userAgent, "LLM-SDK/") { return "sdk-python" }
    if strings.Contains(userAgent, "node-llm/") { return "sdk-node" }
    if strings.Contains(userAgent, "curl") { return "cli-curl" }
    if strings.Contains(userAgent, "PostmanRuntime") { return "postman" }
    return "unknown"
}

Loki query to find impacted group:

{namespace="portfolio"} 
| json 
| user_group="portfolio-agents" 
| client_type="sdk-python" 
| status="error"
| stats count() by error_reason

Pros:

  • Simple to implement (just structured logging)
  • Works with existing Loki + Grafana
  • No SDK overhead
  • Full flexibility (any custom claim)
  • Easy to debug (plain JSON)

Cons:

  • Logs only; doesn't tag metrics automatically
  • Needs Loki for log aggregation (additional infra)
  • Can't correlate metrics and logs without additional work
  • Log volume explodes (every request logged)

Approach 3: HTTP Headers Passthrough + Service-Level Tagging

Pattern: Gateway adds custom headers (X-User-Group, X-Client-Type, X-Org-ID); every upstream service logs them.

Implementation:

// Gateway outbound handler
func EnrichUpstreamRequest(upstream string, req *http.Request) *http.Request {
    claims := req.Context().Value("claims").(jwt.MapClaims)
    
    // Inject user metadata as headers
    req.Header.Set("X-User-ID", claims["sub"].(string))
    req.Header.Set("X-User-Group", claims["groups"].([]interface{})[0].(string))
    req.Header.Set("X-Org-ID", claims["org_id"].(string))
    req.Header.Set("X-Client-Type", parseClientType(req.Header.Get("User-Agent")))
    req.Header.Set("X-Request-ID", uuid.New().String()) // existing
    req.Header.Set("X-Trace-ID", req.Header.Get("X-Trace-ID")) // OpenTelemetry
    
    return req
}

// Upstream service (vLLM, Temporal, S3) middleware
func LogUserContextMiddleware(next http.Handler) http.Handler {
    return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
        log := map[string]interface{}{
            "request_id": r.Header.Get("X-Request-ID"),
            "trace_id": r.Header.Get("X-Trace-ID"),
            "user_id": r.Header.Get("X-User-ID"),
            "user_group": r.Header.Get("X-User-Group"),
            "org_id": r.Header.Get("X-Org-ID"),
            "client_type": r.Header.Get("X-Client-Type"),
            "service": "llm-serving",
            "endpoint": r.URL.Path,
        }
        
        // Log at start and end of request
        jsonLog, _ := json.Marshal(log)
        fmt.Println(string(jsonLog))
        
        next.ServeHTTP(w, r)
    })
}

Prometheus labels (emit from upstream service):

llm_requests_total{
  user_group="homelab-admins",
  org_id="portfolio",
  client_type="browser-chrome",
  status="success"
}

Pros:

  • Works everywhere (simple HTTP headers)
  • No SDK/baggage overhead
  • Metrics can be tagged per service
  • Compatible with all upstream services (vLLM, Ollama, Temporal, etc.)
  • Easy to add/debug headers

Cons:

  • Manual propagation needed in each service
  • Header limits (~8KB total)
  • Not standard (custom X-* headers)
  • Hard to enforce consistency across teams

Approach 4: Synthetic User Journeys + Client Fingerprinting

Pattern: Portfolio + Poimen + CLI clients → identify as "portfolio-web-users", "poimen-cli-users"; track their entire journey through system.

Implementation:

// Client registration (SDK or browser)
type ClientMetadata struct {
    ClientID      string // generated once, persisted locally
    ClientType    string // "portfolio-web", "poimen-cli", "memory-api-client"
    ClientVersion string
    UserGroup     string // from JWT
    Org           string
}

// Gateway: extract or generate client fingerprint
func GetClientFingerprint(req *http.Request) string {
    clientID := req.Header.Get("X-Client-ID") // SDK provides
    if clientID == "" {
        // Browser: fingerprint from User-Agent + IP + TLS
        clientID = Hash(
            req.Header.Get("User-Agent"),
            ClientIP(req),
            req.TLS.PeerCertificates[0].SerialNumber,
        )
    }
    return clientID
}

// Tag all downstream operations with client fingerprint
func RecordClientJourney(clientID, userGroup, clientType string, operation string, status string) {
    // Prometheus
    clientJourneyDuration.WithLabelValues(clientID, userGroup, clientType, operation, status).Observe(duration)
    
    // Loki (full request chain)
    log := map[string]interface{}{
        "client_fingerprint": clientID,
        "user_group": userGroup,
        "client_type": clientType,
        "operation": operation,
        "status": status,
        "timestamp": time.Now(),
    }
    LogToLoki(log)
}

Query to find impacted users:

# Alert: "Chat API failures for portfolio-web-users"
(
  rate(llm_requests_total{client_type="portfolio-web", status="error"}[5m]) > 0.1
)

Grafana dashboard:

  • Heatmap: client_type × user_group vs latency
  • Time series: error rate by client_fingerprint
  • Table: top 10 impacted client IPs/fingerprints

Pros:

  • Surgical root cause ("portfolio-cli-v2.1 broken for homelab-admins")
  • Can replay exact user journey from logs
  • Cohort-based alerting ("if 5% of portfolio-web fail, alert")
  • User behavior analysis (which clients call what)

Cons:

  • Privacy concerns (fingerprinting users)
  • Complex state management (client IDs, persistent storage)
  • High cardinality (can explode metrics cardinality if not capped)
  • Requires SDK integration

Approach 5: Event Sourcing + User Action Audit Log

Pattern: Every operation → immutable audit log entry with user, client, action, outcome. Query for "all portfolio-web users who failed chat between 2-3pm".

Implementation:

type AuditEvent struct {
    EventID       string    // unique per operation
    Timestamp     time.Time
    RequestID     string    // correlation ID
    
    // Who
    UserID        string
    UserGroup     string
    Email         string
    
    // What
    Action        string // "chat_completion", "sqs_send", "memory_ingest"
    Service       string // "llm-serving", "sqs", "memory"
    ResourceID    string // model name, queue name, etc.
    
    // How
    ClientType    string // "browser-chrome", "sdk-python", "cli-curl"
    ClientID      string
    ClientVersion string
    ClientIP      string
    
    // Result
    Status        string // "success", "error", "timeout"
    ErrorCode     string
    ErrorMessage  string
    Duration      int64 // milliseconds
    
    // Context
    OrgID         string
    ProjectID     string
    Metadata      map[string]interface{}
}

// Emit audit event (to Postgres/Clickhouse)
func EmitAuditEvent(ctx context.Context, event AuditEvent) {
    db.InsertAuditEvent(event)
    // Also send to Loki for real-time analysis
    LogToLoki(event)
}

// SQL query: "Find all users in portfolio-agents impacted by chat failure"
// SELECT DISTINCT user_id, email FROM audit_events 
// WHERE user_group='portfolio-agents' 
//   AND action='chat_completion' 
//   AND status='error' 
//   AND timestamp > NOW() - INTERVAL '1 hour'
//   AND client_type IN ('browser-chrome', 'sdk-python')

Grafana query (using Loki or Clickhouse):

SELECT 
  user_group,
  client_type,
  COUNT(*) as total_requests,
  SUM(CASE WHEN status='error' THEN 1 ELSE 0 END) as failures,
  SUM(CASE WHEN status='error' THEN 1 ELSE 0 END) * 100.0 / COUNT(*) as error_rate
FROM audit_events
WHERE timestamp > NOW() - INTERVAL '1 hour'
GROUP BY user_group, client_type
ORDER BY error_rate DESC

Pros:

  • Immutable audit trail (compliance, forensics)
  • Can answer "Who was affected and how?" precisely
  • Excellent for root cause (filter by user_group + timestamp + client_type)
  • Supports RBAC/compliance reporting
  • Works with both logs and metrics

Cons:

  • Storage heavy (every operation stored)
  • Requires additional data store (Postgres/Clickhouse)
  • Latency for complex queries (especially on large datasets)
  • Schema evolution complexity

Combine Approaches 1 + 3 + 5:

┌─────────────────────────────────────────────────────────────┐
│ Gateway (homelab-frontend)                                  │
│                                                             │
│ 1. Parse JWT → extract user_id, user_group, org_id        │
│ 2. Add OTel baggage (Approach 1)                           │
│ 3. Add X-* headers (Approach 3)                            │
│ 4. Emit audit event to Loki + Postgres (Approach 5)       │
│                                                             │
│ Headers sent upstream:                                      │
│   X-Request-ID (existing)                                  │
│   X-Trace-ID (OTel)                                        │
│   traceparent, baggage (OTel)                              │
│   X-User-ID, X-User-Group, X-Org-ID                        │
│   X-Client-Type, X-Client-ID                               │
└─────────────────────────────────────────────────────────────┘
         │
         ├──→ llm-serving (vLLM)
         │    - Log headers
         │    - Emit metrics {user_group, client_type}
         │    - Propagate X-Request-ID
         │
         ├──→ sqs (Temporal)
         │    - Log headers
         │    - Emit metrics {user_group, org_id}
         │
         └──→ memory (Poimen)
              - Log headers
              - Emit metrics {user_group, client_type}
              
         └──→ Loki (structured logs)
         └──→ Prometheus (metrics with user/client tags)
         └──→ Postgres/Clickhouse (audit events)

Workflow:

  1. User calls gateway with Bearer token + User-Agent
  2. Gateway parses JWT → creates OTel baggage + X-* headers
  3. Gateway logs audit event (user_id, user_group, client_type, action, status)
  4. Headers propagated to upstream services
  5. Root cause query:
    • Metrics: llm_requests_total{user_group="portfolio-agents", client_type="sdk-python", status="error"}
    • Logs (Loki): {namespace="portfolio"} | json | user_group="portfolio-agents" | client_type="sdk-python" | status="error"
    • Audit trail (Postgres): SELECT * FROM audit_events WHERE user_group='portfolio-agents' AND client_type='sdk-python' AND timestamp>'2026-09-04'

Implementation Roadmap

Phase What Where Effort
1 JWT claims extraction internal/auth/jwt.go 1 day
2 OTel baggage middleware internal/middleware/baggage.go 1 day
3 X-* header propagation internal/proxy/proxy.go 0.5 day
4 Structured logging (Loki) internal/logging/structured.go 1 day
5 Audit events (Postgres) internal/audit/events.go + schema 2 days
6 Prometheus labels per service internal/observability/*.go 1 day
7 Grafana dashboards k8s/grafana-dashboards/ 1 day

Which Approach for Your Use Case?

User Group Identification:

  • Use OTel Baggage (Approach 1) — auto-propagates
  • Fallback to X- headers* (Approach 3) if OTel not available

Client Type Detection:

  • Parse User-Agent in gateway
  • Tag as: browser-chrome, browser-safari, sdk-python, sdk-node, cli-curl, postman, unknown

Root Cause (Surgical Debugging):

  • Use Audit Events (Approach 5) with Postgres
  • Query: "Who from portfolio-agents using sdk-python failed in last hour?"

High-Cardinality Metrics:

  • Avoid tagging by individual user_id (cardinality explosion)
  • Use: user_group, org_id, client_type, sdk_version (bounded)

Real-Time Alerting:

  • Prometheus metrics with {user_group, client_type, status} labels
  • Alert: "error_rate > 5% for user_group='portfolio-agents' for 5min"