Critical services (iam/monitoring/temporal/cicd) keep 100% logs. Others get 50% sampling + selective drops (health/debug noise). Balances log volume (40-50% reduction) with error visibility.
- Loki log aggregation (MinIO backed, 10-day retention) - Promtail daemonset (pod + talos journal logs) - Prometheus + kube-state-metrics - Grafana dashboards (6-row template per service)