- Nginx ingress + TLS termination (homelab-ca) - Portainer container UI - CoreDNS internal DNS rewrites - DuckDNS DDNS updater - Ollama LLM inference - 8 project-usage guides (team reference)
6.4 KiB
6.4 KiB
Monitoring: Prometheus, Grafana & Loki
Prometheus: prometheus-kube-prom-prometheus.monitoring.svc.cluster.local:9090
Grafana: https://grafana.riotpiao.homelab.com
Loki: loki.logging.svc.cluster.local:3100
Namespaces: monitoring, logging
When to Use
- Metrics — CPU, memory, latency, request rate, custom business metrics (Prometheus)
- Logs — Pod logs, kernel logs, trace events (Loki)
- Dashboards — Real-time visualization (Grafana)
- Alerts — Page on high error rate, latency spike, or resource exhaustion
Quick Start
1. Access Grafana:
# Browser: https://grafana.riotpiao.homelab.com
# Login: admin / GRAFANA_ADMIN_PASSWORD (from .env)
# Or via Authentik SSO
# Port-forward (if no external access)
make pf-grafana # localhost:3000
2. View dashboards:
Grafana → Dashboards → Homelab folder
- Service Availability (uptime probes + cert expiry)
- Latency & Golden Signals (request rate/error %/duration)
- Hardware Statistics (per-node CPU/mem/disk)
- Kube-Controller Health (API server metrics)
- Service Internals (per-service deep-dive)
3. Add Prometheus data source:
Grafana → Settings → Data Sources → Add
- Type: Prometheus
- URL: http://prometheus-kube-prom-prometheus.monitoring.svc.cluster.local:9090
- Save & test
4. Create dashboard:
{
"dashboard": {
"title": "My Service",
"panels": [
{
"targets": [
{
"expr": "rate(http_requests_total{service='myapp'}[5m])"
}
],
"title": "Request Rate"
}
]
}
}
Instrumentation Patterns
Go (Prometheus):
import "github.com/prometheus/client_golang/prometheus"
// Counter (monotonic increase)
requestsTotal := prometheus.NewCounterVec(
prometheus.CounterOpts{
Name: "http_requests_total",
Help: "Total requests",
},
[]string{"method", "status"},
)
// Gauge (can go up or down)
activeConnections := prometheus.NewGauge(prometheus.GaugeOpts{
Name: "active_connections",
})
// Histogram (latency buckets)
requestDuration := prometheus.NewHistogramVec(
prometheus.HistogramOpts{
Name: "http_request_duration_seconds",
Buckets: []float64{.001, .01, .1, 1},
},
[]string{"method", "path"},
)
// Record metric
requestsTotal.WithLabelValues("POST", "200").Inc()
requestDuration.WithLabelValues("POST", "/api/users").Observe(0.042)
Node.js (prom-client):
const prometheus = require('prom-client');
const httpRequestDuration = new prometheus.Histogram({
name: 'http_request_duration_seconds',
help: 'Duration of HTTP requests in seconds',
labelNames: ['method', 'path', 'status'],
buckets: [0.001, 0.01, 0.1, 1],
});
// Middleware
app.use((req, res, next) => {
const start = Date.now();
res.on('finish', () => {
const duration = (Date.now() - start) / 1000;
httpRequestDuration.labels(req.method, req.path, res.statusCode).observe(duration);
});
next();
});
// Expose metrics
app.get('/metrics', (req, res) => {
res.set('Content-Type', prometheus.register.contentType);
res.end(prometheus.register.metrics());
});
ServiceMonitor (Auto-Scrape)
Prometheus automatically discovers ServiceMonitor resources:
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: myapp
namespace: myapp-ns
spec:
selector:
matchLabels:
app: myapp
endpoints:
- port: metrics
interval: 30s
path: /metrics
Deploy & verify:
kubectl apply -f myapp-servicemonitor.yaml
# Check Prometheus targets
kubectl port-forward -n monitoring svc/prometheus-kube-prom-prometheus 9090:9090
# http://localhost:9090/targets → should show myapp
PrometheusRule (Alerting)
Define alert rules:
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: myapp-alerts
namespace: myapp-ns
spec:
groups:
- name: myapp
rules:
- alert: HighErrorRate
expr: rate(http_requests_total{status=~"5.."}[5m]) > 0.05 # > 5% error rate
for: 5m
annotations:
summary: "High error rate for {{ $labels.service }}"
- alert: HighLatency
expr: histogram_quantile(0.99, http_request_duration_seconds_bucket) > 1 # p99 > 1s
for: 10m
annotations:
summary: "High latency for {{ $labels.path }}"
Logging (Loki)
View pod logs in Grafana:
Grafana → Explore → Loki
Query: {namespace="story-crater-backend"} | json | level="error"
Label filters:
- namespace
- pod_name
- container
- level (error, warn, info)
Log retention:
- 10-day default (configurable)
- Older logs deleted automatically
- Chunks stored in MinIO (s3://loki-chunks)
Monitoring Checklist
For every service:
- Export
/metrics(Prometheus format) - Add ServiceMonitor resource
- Add PrometheusRule (errors, latency)
- Create Grafana dashboard (6 rows: Availability, Resources, Domain, Logs, SLO, Related)
- Set up Slack/PagerDuty integration (if on-call)
Key Queries
Request rate by status:
sum(rate(http_requests_total[5m])) by (status)
Error rate %:
sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m])) * 100
P99 latency:
histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))
Memory usage %:
(1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes)) * 100
Troubleshooting
Prometheus can't scrape ServiceMonitor:
# Check ServiceMonitor exists
k get servicemonitor -A | grep myapp
# Verify Service has correct labels
k get svc -A -o json | jq '.items[] | select(.metadata.name=="myapp") | .metadata.labels'
# Check Prometheus targets
kubectl port-forward -n monitoring svc/prometheus-kube-prom-prometheus 9090:9090
# http://localhost:9090/targets → look for "Down" targets with error
Metrics missing from Grafana:
# Verify pod is exporting metrics
k port-forward -n myapp-ns pod/myapp-0 8080:8080
curl http://localhost:8080/metrics | grep my_metric_name
# Check query syntax in Grafana
# Hover over query → "Query Inspector" to see response
Alert not firing:
# Check PrometheusRule is loaded
k get prometheusrule -n myapp-ns
# Verify query in Prometheus
kubectl port-forward -n monitoring svc/prometheus-kube-prom-prometheus 9090:9090
# http://localhost:9090 → Graph tab → paste query
See /TROUBLESHOOTING.md for full incident guide.