175 lines
4.8 KiB
Markdown
175 lines
4.8 KiB
Markdown
# T1.8: Health Checks for Kubernetes
|
|||
|
|
|
||
|
|
**Submilestone:** T1 (Production Hardening)
|
||
|
|
**Status:** ✅ COMPLETE
|
||
|
|
**Branch:** `task/T1.8`
|
||
|
|
|
||
|
|
## Overview
|
||
|
|
|
||
|
|
Implement comprehensive health checks for Kubernetes deployments with liveness and readiness probes.
|
||
|
|
|
||
|
|
## Requirements
|
||
|
|
|
||
|
|
### Endpoints
|
||
|
|
|
||
|
|
- **GET /health** - Full health report (JSON)
|
||
|
|
- Returns 200 if healthy, 503 if unhealthy
|
||
|
|
- Includes all component statuses, latencies, timestamps
|
||
|
|
|
||
|
|
- **GET /health/live** - Kubernetes liveness probe
|
||
|
|
- Returns 200 if service is running
|
||
|
|
- Returns 503 if not initialized
|
||
|
|
|
||
|
|
- **GET /health/ready** - Kubernetes readiness probe
|
||
|
|
- Returns 200 if service is ready to accept traffic
|
||
|
|
- Returns 503 if any component unhealthy
|
||
|
|
|
||
|
|
### Components
|
||
|
|
|
||
|
|
1. **Temporal** - Cluster connectivity check
|
||
|
|
- Attempts to get a workflow execution
|
||
|
|
- Returns healthy if Temporal responds (even with NotFound)
|
||
|
|
- Returns unhealthy if unreachable
|
||
|
|
|
||
|
|
### Features
|
||
|
|
|
||
|
|
- Periodic health check caching (30s interval) to avoid excessive checks
|
||
|
|
- JSON health reports with component status, latency, timestamp
|
||
|
|
- Separate liveness and readiness checks for K8s probes
|
||
|
|
- Graceful shutdown with health server cleanup
|
||
|
|
|
||
|
|
## Implementation
|
||
|
|
|
||
|
|
### Internal Package: `internal/health`
|
||
|
|
|
||
|
|
#### `health.go`
|
||
|
|
- `Status` type with constants: `StatusHealthy`, `StatusUnhealthy`, `StatusUnknown`
|
||
|
|
- `ComponentHealth` struct for individual component status
|
||
|
|
- `HealthReport` struct for complete health status
|
||
|
|
- `Checker` interface for health checking
|
||
|
|
- `Check()` method that performs comprehensive health check
|
||
|
|
- `IsHealthy()` for quick boolean check
|
||
|
|
- Caching mechanism to avoid repeated checks within interval
|
||
|
|
|
||
|
|
#### `handler.go`
|
||
|
|
- HTTP handler implementation
|
||
|
|
- `RegisterRoutes()` to set up endpoints on a mux
|
||
|
|
- Handlers for `/health`, `/health/live`, `/health/ready`
|
||
|
|
- Proper HTTP status codes (200 for healthy, 503 for unhealthy)
|
||
|
|
|
||
|
|
#### `health_test.go`
|
||
|
|
- Unit tests for health checker
|
||
|
|
- Tests for nil client, caching, JSON serialization
|
||
|
|
- Tests for timestamp validation
|
||
|
|
- 10/10 tests passing ✅
|
||
|
|
|
||
|
|
### Integration
|
||
|
|
|
||
|
|
**cmd/worker/main.go**
|
||
|
|
- Health check server runs on port 8081
|
||
|
|
- Runs in separate goroutine alongside worker
|
||
|
|
- Graceful shutdown on SIGINT/SIGTERM
|
||
|
|
- Waits for health server to shutdown before exiting
|
||
|
|
|
||
|
|
**cmd/starter/main.go**
|
||
|
|
- `--health` flag to run health check and exit
|
||
|
|
- Outputs JSON health report
|
||
|
|
- Returns non-zero exit code if unhealthy
|
||
|
|
|
||
|
|
## Verification Criteria
|
||
|
|
|
||
|
|
✅ **All criteria met:**
|
||
|
|
|
||
|
|
1. **Health endpoints responsive**
|
||
|
|
- GET /health returns 200 with JSON report
|
||
|
|
- GET /health/live returns 200 if running
|
||
|
|
- GET /health/ready returns 503 if Temporal unavailable
|
||
|
|
|
||
|
|
2. **Kubernetes integration**
|
||
|
|
- Can be used as livenessProbe target
|
||
|
|
- Can be used as readinessProbe target
|
||
|
|
- Port 8081 exposed for probes
|
||
|
|
|
||
|
|
3. **Component checks**
|
||
|
|
- Temporal connectivity verified via GetWorkflow call
|
||
|
|
- Caching prevents excessive health checks
|
||
|
|
- Latency measured and reported
|
||
|
|
|
||
|
|
4. **Graceful shutdown**
|
||
|
|
- Health server stops on SIGINT/SIGTERM
|
||
|
|
- Worker stops cleanly
|
||
|
|
- No hanging goroutines
|
||
|
|
|
||
|
|
5. **CLI integration**
|
||
|
|
- `starter --health` command works
|
||
|
|
- Outputs JSON report
|
||
|
|
- Exits with appropriate code
|
||
|
|
|
||
|
|
## Testing
|
||
|
|
|
||
|
|
```bash
|
||
|
|
# Unit tests
|
||
|
|
go test -v ./internal/health
|
||
|
|
# Result: PASS (10/10 tests)
|
||
|
|
|
||
|
|
# Integration test (requires Temporal)
|
||
|
|
# When Temporal unavailable:
|
||
|
|
curl http://localhost:8081/health
|
||
|
|
# Returns: 503 with status="unhealthy", components.temporal.error set
|
||
|
|
|
||
|
|
# When Temporal available:
|
||
|
|
curl http://localhost:8081/health
|
||
|
|
# Returns: 200 with status="healthy"
|
||
|
|
```
|
||
|
|
|
||
|
|
## Kubernetes Configuration
|
||
|
|
|
||
|
|
Example liveness probe:
|
||
|
|
```yaml
|
||
|
|
livenessProbe:
|
||
|
|
httpGet:
|
||
|
|
path: /health/live
|
||
|
|
port: 8081
|
||
|
|
initialDelaySeconds: 10
|
||
|
|
periodSeconds: 10
|
||
|
|
```
|
||
|
|
|
||
|
|
Example readiness probe:
|
||
|
|
```yaml
|
||
|
|
readinessProbe:
|
||
|
|
httpGet:
|
||
|
|
path: /health/ready
|
||
|
|
port: 8081
|
||
|
|
initialDelaySeconds: 5
|
||
|
|
periodSeconds: 5
|
||
|
|
```
|
||
|
|
|
||
|
|
## Files Changed
|
||
|
|
|
||
|
|
- ✅ `internal/health/health.go` - Core health checker (106 lines)
|
||
|
|
- ✅ `internal/health/handler.go` - HTTP endpoints (68 lines)
|
||
|
|
- ✅ `internal/health/health_test.go` - Unit tests (119 lines)
|
||
|
|
- ✅ `cmd/worker/main.go` - Worker integration
|
||
|
|
- ✅ `cmd/starter/main.go` - Starter health check command
|
||
|
|
- ✅ `tasks/board-T1.md` - Task board update
|
||
|
|
|
||
|
|
## Dependencies
|
||
|
|
|
||
|
|
- `go.temporal.io/sdk/client` - Already in go.mod
|
||
|
|
- `net/http` - Standard library
|
||
|
|
- `encoding/json` - Standard library
|
||
|
|
- `github.com/stretchr/testify/assert` - Already in go.mod
|
||
|
|
|
||
|
|
## Notes
|
||
|
|
|
||
|
|
- Health check server runs on `:8081` (separate from main application)
|
||
|
|
- Caching interval set to 30 seconds (configurable)
|
||
|
|
- Temporal check uses GetWorkflow with timeout for quick response
|
||
|
|
- Handler is reusable across different services
|
||
|
|
|
||
|
|
## Next Steps (T1.7 → T1.1 → T1.2)
|
||
|
|
|
||
|
|
1. **T1.7:** Immutable audit logging (track all decisions)
|
||
|
|
2. **T1.2:** Structured logging + Prometheus metrics
|
||
|
|
3. **T1.1:** Workflow error recovery & deadletter handling
|