feat(T1.2): implement structured logging and Prometheus metrics
- Add internal/logging package with zap-based structured JSON logging - Support development (colored) and production (JSON) modes via ENVIRONMENT env var - Add logging helpers: Info(), Error(), Warn(), Debug(), Fatal() - Add field helpers: String(), Int(), Int64(), Err() - Add internal/metrics package with 16 comprehensive Prometheus metrics - Track workflows: starts, completions, duration by type/status - Track activities: starts, completions, duration, retries by type - Track LLM calls: total calls and latency by model - Track git operations: total and duration by operation type - Track judge decisions: decisions by type - Track Temporal errors: connection errors by type - Track cache efficiency: hits and misses by cache type - Track tasks in progress: gauge metric by task type - Metrics exported on /metrics endpoint (Prometheus text format) - Integrate structured logging in cmd/worker and cmd/starter - Replace all log.Printf/log.Fatalf with structured logging - Add /metrics endpoint to health check server - 8/8 logging tests passing, 13/13 metrics tests passing - All verification criteria met Dependencies added: - go.uber.org/zap v1.28.0 (structured logging) - github.com/prometheus/client_golang v1.24.1 (metrics export) Closes T1.2
This commit is contained in:
+223
@@ -0,0 +1,223 @@
|
||||
# T1.2: Structured Logging + Prometheus Metrics
|
||||
|
||||
**Submilestone:** T1 (Production Hardening)
|
||||
**Status:** ✅ COMPLETE
|
||||
**Branch:** `task/T1.2`
|
||||
|
||||
## Overview
|
||||
|
||||
Implement structured JSON logging with zap and comprehensive Prometheus metrics export for observability.
|
||||
|
||||
## Requirements
|
||||
|
||||
### Structured Logging
|
||||
|
||||
- Replace all `log.Printf` / `log.Fatalf` with structured logging
|
||||
- Use `go.uber.org/zap` for structured JSON logging
|
||||
- Support both development (colored) and production (JSON) modes
|
||||
- Easy field attachment: `logging.Info("message", logging.String("key", "value"))`
|
||||
|
||||
### Prometheus Metrics
|
||||
|
||||
- 16 comprehensive metrics covering workflows, activities, LLM calls, git operations, judge decisions
|
||||
- Counter metrics: workflow starts/completions, activity starts/completions, retries, LLM calls, git operations, judge decisions
|
||||
- Histogram metrics: workflow duration, activity duration, LLM latency, git operation duration
|
||||
- Gauge metrics: tasks in progress
|
||||
- Error tracking: Temporal connection errors, cache hit/miss ratio
|
||||
- Metrics exported on `/metrics` HTTP endpoint (Prometheus format)
|
||||
|
||||
### Integration
|
||||
|
||||
- Health check server (port 8081) now serves both `/health*` and `/metrics`
|
||||
- Graceful logging shutdown with `logging.Sync()`
|
||||
- Both worker and starter commands use structured logging
|
||||
|
||||
## Implementation
|
||||
|
||||
### Internal Package: `internal/logging`
|
||||
|
||||
#### `logger.go`
|
||||
- `InitLogger()` - Initialize global logger (dev or prod mode)
|
||||
- `GetLogger()` - Get logger instance
|
||||
- `Info()`, `Error()`, `Warn()`, `Debug()`, `Fatal()` - Log functions
|
||||
- Field helpers: `String()`, `Int()`, `Int64()`, `Err()`
|
||||
- `Sync()` - Flush buffered logs
|
||||
- `With()` - Create logger with additional fields
|
||||
- 8/8 unit tests passing ✅
|
||||
|
||||
#### `logger_test.go`
|
||||
- Tests for logger initialization, field creation, logging functions
|
||||
- Verifies no panics on concurrent logging
|
||||
|
||||
### Internal Package: `internal/metrics`
|
||||
|
||||
#### `metrics.go`
|
||||
- 16 pre-registered Prometheus metrics
|
||||
- Helper functions for recording each metric type
|
||||
- Metrics organized by concern: workflows, activities, LLM, git, judge, temporal, cache
|
||||
- 13/13 unit tests passing ✅
|
||||
|
||||
#### `metrics_test.go`
|
||||
- Tests that all metrics are registered
|
||||
- Tests that recording functions don't panic
|
||||
- Verifies metric registration
|
||||
|
||||
### Integration Points
|
||||
|
||||
**cmd/worker/main.go**
|
||||
- Initializes logger on startup
|
||||
- Uses `logging.Info()`, `logging.Fatal()`, `logging.Warn()` throughout
|
||||
- Health server serves `/metrics` endpoint
|
||||
- Structured shutdown logging
|
||||
|
||||
**cmd/starter/main.go**
|
||||
- Initializes logger on startup
|
||||
- Logs configuration load, Temporal connection, workflow start
|
||||
- Supports `--health` command with structured logging
|
||||
- Clean shutdown with `logging.Sync()`
|
||||
|
||||
**internal/health/handler.go**
|
||||
- Prometheus handler integrated via `promhttp.Handler()`
|
||||
- `/metrics` endpoint available on all deployments
|
||||
|
||||
## Verification Criteria
|
||||
|
||||
✅ **All criteria met:**
|
||||
|
||||
1. **Structured logging deployed**
|
||||
- All log statements use structured fields
|
||||
- JSON output in production
|
||||
- Colored output in development
|
||||
|
||||
2. **Prometheus metrics exposed**
|
||||
- 16 comprehensive metrics registered
|
||||
- `/metrics` endpoint returns Prometheus text format
|
||||
- Metrics include latencies, counters, and gauges
|
||||
|
||||
3. **All metrics functional**
|
||||
- `WorkflowExecutionsStarted` - workflow launch tracking
|
||||
- `WorkflowExecutionsCompleted` - workflow completion with status
|
||||
- `ActivityExecutionsStarted/Completed/Duration` - activity lifecycle
|
||||
- `ActivityRetries` - retry tracking
|
||||
- `LLMAPICallsTotal` / `LLMAPILatency` - LLM performance
|
||||
- `GitOperationsTotal` / `GitOperationsDuration` - git operation tracking
|
||||
- `TasksInProgress` - real-time task load
|
||||
- `JudgeDecisionsTotal` - decision tracking
|
||||
- `TemporalConnectionErrors` - error tracking
|
||||
- `CacheHits` / `CacheMisses` - cache efficiency
|
||||
|
||||
4. **Integration complete**
|
||||
- Worker uses structured logging throughout
|
||||
- Starter uses structured logging throughout
|
||||
- Both commands can use `--health` to check system status
|
||||
- Graceful shutdown flushes logs
|
||||
|
||||
5. **Test coverage**
|
||||
- 8/8 logging tests passing
|
||||
- 13/13 metrics tests passing
|
||||
- All unit tests pass
|
||||
- No panics on concurrent logging
|
||||
|
||||
## Testing
|
||||
|
||||
```bash
|
||||
# Unit tests
|
||||
go test -v ./internal/logging ./internal/metrics
|
||||
# Result: PASS (21/21 tests)
|
||||
|
||||
# Full test suite
|
||||
go test -v ./...
|
||||
# Result: All tests pass
|
||||
|
||||
# Integration test (requires running worker)
|
||||
curl http://localhost:8081/metrics
|
||||
# Returns: Prometheus metrics in text format
|
||||
|
||||
# Logging output
|
||||
ENVIRONMENT=development go run ./cmd/worker
|
||||
# Output: Colored JSON logs with structured fields
|
||||
|
||||
ENVIRONMENT=production go run ./cmd/worker
|
||||
# Output: JSON logs suitable for Loki/ELK
|
||||
```
|
||||
|
||||
## Kubernetes Configuration
|
||||
|
||||
Example logging in pods:
|
||||
```yaml
|
||||
env:
|
||||
- name: ENVIRONMENT
|
||||
value: "production"
|
||||
```
|
||||
|
||||
Example Prometheus scrape config:
|
||||
```yaml
|
||||
scrape_configs:
|
||||
- job_name: 'poimen-worker'
|
||||
static_configs:
|
||||
- targets: ['localhost:8081']
|
||||
metrics_path: '/metrics'
|
||||
```
|
||||
|
||||
## Metrics Schema
|
||||
|
||||
All metrics prefixed with `poimen_`:
|
||||
|
||||
### Workflow Metrics
|
||||
- `poimen_workflow_executions_started_total{workflow_type}` - Counter
|
||||
- `poimen_workflow_executions_completed_total{workflow_type, status}` - Counter
|
||||
- `poimen_workflow_duration_seconds{workflow_type}` - Histogram
|
||||
|
||||
### Activity Metrics
|
||||
- `poimen_activity_executions_started_total{activity_type}` - Counter
|
||||
- `poimen_activity_executions_completed_total{activity_type, status}` - Counter
|
||||
- `poimen_activity_duration_seconds{activity_type}` - Histogram
|
||||
- `poimen_activity_retries_total{activity_type}` - Counter
|
||||
|
||||
### LLM Metrics
|
||||
- `poimen_llm_api_calls_total{model_id, status}` - Counter
|
||||
- `poimen_llm_api_latency_seconds{model_id}` - Histogram
|
||||
|
||||
### Git Metrics
|
||||
- `poimen_git_operations_total{operation, status}` - Counter
|
||||
- `poimen_git_operations_duration_seconds{operation}` - Histogram
|
||||
|
||||
### Other Metrics
|
||||
- `poimen_tasks_in_progress{task_type}` - Gauge
|
||||
- `poimen_judge_decisions_total{decision}` - Counter
|
||||
- `poimen_temporal_connection_errors_total{error_type}` - Counter
|
||||
- `poimen_cache_hits_total{cache_type}` - Counter
|
||||
- `poimen_cache_misses_total{cache_type}` - Counter
|
||||
|
||||
## Files Changed
|
||||
|
||||
- ✅ `internal/logging/logger.go` - Structured logger (71 lines)
|
||||
- ✅ `internal/logging/logger_test.go` - Logger tests (70 lines)
|
||||
- ✅ `internal/metrics/metrics.go` - Prometheus metrics (222 lines)
|
||||
- ✅ `internal/metrics/metrics_test.go` - Metrics tests (87 lines)
|
||||
- ✅ `internal/health/handler.go` - Added `/metrics` endpoint
|
||||
- ✅ `cmd/worker/main.go` - Structured logging integration
|
||||
- ✅ `cmd/starter/main.go` - Structured logging integration
|
||||
- ✅ `go.mod` - Added zap, prometheus/client_golang dependencies
|
||||
- ✅ `tasks/board-T1.md` - Task board update
|
||||
|
||||
## Dependencies Added
|
||||
|
||||
- `go.uber.org/zap` v1.28.0 - Structured logging
|
||||
- `github.com/prometheus/client_golang` v1.24.1 - Prometheus metrics
|
||||
- Plus 8 transitive dependencies for Prometheus support
|
||||
|
||||
## Next Steps (T1.1 → T1.3 → T1.4)
|
||||
|
||||
1. **T1.1:** Workflow error recovery & deadletter handling
|
||||
2. **T1.3:** Timeout tuning automation based on historical failures
|
||||
3. **T1.4:** Board state validation & auto-heal from corruption
|
||||
|
||||
## Notes
|
||||
|
||||
- Logger uses global singleton pattern for simplicity (can be refactored to DI if needed)
|
||||
- Metrics are auto-registered via `promauto` (thread-safe, idempotent)
|
||||
- `/metrics` endpoint serves standard Prometheus text format (compatible with all scraping systems)
|
||||
- Logging mode controlled by `ENVIRONMENT` env var (default: development)
|
||||
- All metric labels are strings (Prometheus requirement)
|
||||
- Histograms use default buckets (10ms, 100ms, 1s, 10s, etc.)
|
||||
Reference in New Issue
Block a user