k8s/services: add ingress networking portainer llm and project guides
- Nginx ingress + TLS termination (homelab-ca) - Portainer container UI - CoreDNS internal DNS rewrites - DuckDNS DDNS updater - Ollama LLM inference - 8 project-usage guides (team reference)
This commit is contained in:
@@ -0,0 +1,265 @@
|
||||
# Monitoring: Prometheus, Grafana & Loki
|
||||
|
||||
**Prometheus:** `prometheus-kube-prom-prometheus.monitoring.svc.cluster.local:9090`
|
||||
**Grafana:** `https://grafana.riotpiao.homelab.com`
|
||||
**Loki:** `loki.logging.svc.cluster.local:3100`
|
||||
**Namespaces:** `monitoring`, `logging`
|
||||
|
||||
## When to Use
|
||||
|
||||
- **Metrics** — CPU, memory, latency, request rate, custom business metrics (Prometheus)
|
||||
- **Logs** — Pod logs, kernel logs, trace events (Loki)
|
||||
- **Dashboards** — Real-time visualization (Grafana)
|
||||
- **Alerts** — Page on high error rate, latency spike, or resource exhaustion
|
||||
|
||||
## Quick Start
|
||||
|
||||
**1. Access Grafana:**
|
||||
```bash
|
||||
# Browser: https://grafana.riotpiao.homelab.com
|
||||
# Login: admin / GRAFANA_ADMIN_PASSWORD (from .env)
|
||||
# Or via Authentik SSO
|
||||
|
||||
# Port-forward (if no external access)
|
||||
make pf-grafana # localhost:3000
|
||||
```
|
||||
|
||||
**2. View dashboards:**
|
||||
```
|
||||
Grafana → Dashboards → Homelab folder
|
||||
- Service Availability (uptime probes + cert expiry)
|
||||
- Latency & Golden Signals (request rate/error %/duration)
|
||||
- Hardware Statistics (per-node CPU/mem/disk)
|
||||
- Kube-Controller Health (API server metrics)
|
||||
- Service Internals (per-service deep-dive)
|
||||
```
|
||||
|
||||
**3. Add Prometheus data source:**
|
||||
```
|
||||
Grafana → Settings → Data Sources → Add
|
||||
- Type: Prometheus
|
||||
- URL: http://prometheus-kube-prom-prometheus.monitoring.svc.cluster.local:9090
|
||||
- Save & test
|
||||
```
|
||||
|
||||
**4. Create dashboard:**
|
||||
```json
|
||||
{
|
||||
"dashboard": {
|
||||
"title": "My Service",
|
||||
"panels": [
|
||||
{
|
||||
"targets": [
|
||||
{
|
||||
"expr": "rate(http_requests_total{service='myapp'}[5m])"
|
||||
}
|
||||
],
|
||||
"title": "Request Rate"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Instrumentation Patterns
|
||||
|
||||
**Go (Prometheus):**
|
||||
```go
|
||||
import "github.com/prometheus/client_golang/prometheus"
|
||||
|
||||
// Counter (monotonic increase)
|
||||
requestsTotal := prometheus.NewCounterVec(
|
||||
prometheus.CounterOpts{
|
||||
Name: "http_requests_total",
|
||||
Help: "Total requests",
|
||||
},
|
||||
[]string{"method", "status"},
|
||||
)
|
||||
|
||||
// Gauge (can go up or down)
|
||||
activeConnections := prometheus.NewGauge(prometheus.GaugeOpts{
|
||||
Name: "active_connections",
|
||||
})
|
||||
|
||||
// Histogram (latency buckets)
|
||||
requestDuration := prometheus.NewHistogramVec(
|
||||
prometheus.HistogramOpts{
|
||||
Name: "http_request_duration_seconds",
|
||||
Buckets: []float64{.001, .01, .1, 1},
|
||||
},
|
||||
[]string{"method", "path"},
|
||||
)
|
||||
|
||||
// Record metric
|
||||
requestsTotal.WithLabelValues("POST", "200").Inc()
|
||||
requestDuration.WithLabelValues("POST", "/api/users").Observe(0.042)
|
||||
```
|
||||
|
||||
**Node.js (prom-client):**
|
||||
```javascript
|
||||
const prometheus = require('prom-client');
|
||||
|
||||
const httpRequestDuration = new prometheus.Histogram({
|
||||
name: 'http_request_duration_seconds',
|
||||
help: 'Duration of HTTP requests in seconds',
|
||||
labelNames: ['method', 'path', 'status'],
|
||||
buckets: [0.001, 0.01, 0.1, 1],
|
||||
});
|
||||
|
||||
// Middleware
|
||||
app.use((req, res, next) => {
|
||||
const start = Date.now();
|
||||
res.on('finish', () => {
|
||||
const duration = (Date.now() - start) / 1000;
|
||||
httpRequestDuration.labels(req.method, req.path, res.statusCode).observe(duration);
|
||||
});
|
||||
next();
|
||||
});
|
||||
|
||||
// Expose metrics
|
||||
app.get('/metrics', (req, res) => {
|
||||
res.set('Content-Type', prometheus.register.contentType);
|
||||
res.end(prometheus.register.metrics());
|
||||
});
|
||||
```
|
||||
|
||||
## ServiceMonitor (Auto-Scrape)
|
||||
|
||||
**Prometheus automatically discovers ServiceMonitor resources:**
|
||||
```yaml
|
||||
apiVersion: monitoring.coreos.com/v1
|
||||
kind: ServiceMonitor
|
||||
metadata:
|
||||
name: myapp
|
||||
namespace: myapp-ns
|
||||
spec:
|
||||
selector:
|
||||
matchLabels:
|
||||
app: myapp
|
||||
endpoints:
|
||||
- port: metrics
|
||||
interval: 30s
|
||||
path: /metrics
|
||||
```
|
||||
|
||||
**Deploy & verify:**
|
||||
```bash
|
||||
kubectl apply -f myapp-servicemonitor.yaml
|
||||
|
||||
# Check Prometheus targets
|
||||
kubectl port-forward -n monitoring svc/prometheus-kube-prom-prometheus 9090:9090
|
||||
# http://localhost:9090/targets → should show myapp
|
||||
```
|
||||
|
||||
## PrometheusRule (Alerting)
|
||||
|
||||
**Define alert rules:**
|
||||
```yaml
|
||||
apiVersion: monitoring.coreos.com/v1
|
||||
kind: PrometheusRule
|
||||
metadata:
|
||||
name: myapp-alerts
|
||||
namespace: myapp-ns
|
||||
spec:
|
||||
groups:
|
||||
- name: myapp
|
||||
rules:
|
||||
- alert: HighErrorRate
|
||||
expr: rate(http_requests_total{status=~"5.."}[5m]) > 0.05 # > 5% error rate
|
||||
for: 5m
|
||||
annotations:
|
||||
summary: "High error rate for {{ $labels.service }}"
|
||||
|
||||
- alert: HighLatency
|
||||
expr: histogram_quantile(0.99, http_request_duration_seconds_bucket) > 1 # p99 > 1s
|
||||
for: 10m
|
||||
annotations:
|
||||
summary: "High latency for {{ $labels.path }}"
|
||||
```
|
||||
|
||||
## Logging (Loki)
|
||||
|
||||
**View pod logs in Grafana:**
|
||||
```
|
||||
Grafana → Explore → Loki
|
||||
Query: {namespace="story-crater-backend"} | json | level="error"
|
||||
|
||||
Label filters:
|
||||
- namespace
|
||||
- pod_name
|
||||
- container
|
||||
- level (error, warn, info)
|
||||
```
|
||||
|
||||
**Log retention:**
|
||||
- 10-day default (configurable)
|
||||
- Older logs deleted automatically
|
||||
- Chunks stored in MinIO (s3://loki-chunks)
|
||||
|
||||
## Monitoring Checklist
|
||||
|
||||
**For every service:**
|
||||
- [ ] Export `/metrics` (Prometheus format)
|
||||
- [ ] Add ServiceMonitor resource
|
||||
- [ ] Add PrometheusRule (errors, latency)
|
||||
- [ ] Create Grafana dashboard (6 rows: Availability, Resources, Domain, Logs, SLO, Related)
|
||||
- [ ] Set up Slack/PagerDuty integration (if on-call)
|
||||
|
||||
## Key Queries
|
||||
|
||||
**Request rate by status:**
|
||||
```promql
|
||||
sum(rate(http_requests_total[5m])) by (status)
|
||||
```
|
||||
|
||||
**Error rate %:**
|
||||
```promql
|
||||
sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m])) * 100
|
||||
```
|
||||
|
||||
**P99 latency:**
|
||||
```promql
|
||||
histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))
|
||||
```
|
||||
|
||||
**Memory usage %:**
|
||||
```promql
|
||||
(1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes)) * 100
|
||||
```
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
**Prometheus can't scrape ServiceMonitor:**
|
||||
```bash
|
||||
# Check ServiceMonitor exists
|
||||
k get servicemonitor -A | grep myapp
|
||||
|
||||
# Verify Service has correct labels
|
||||
k get svc -A -o json | jq '.items[] | select(.metadata.name=="myapp") | .metadata.labels'
|
||||
|
||||
# Check Prometheus targets
|
||||
kubectl port-forward -n monitoring svc/prometheus-kube-prom-prometheus 9090:9090
|
||||
# http://localhost:9090/targets → look for "Down" targets with error
|
||||
```
|
||||
|
||||
**Metrics missing from Grafana:**
|
||||
```bash
|
||||
# Verify pod is exporting metrics
|
||||
k port-forward -n myapp-ns pod/myapp-0 8080:8080
|
||||
curl http://localhost:8080/metrics | grep my_metric_name
|
||||
|
||||
# Check query syntax in Grafana
|
||||
# Hover over query → "Query Inspector" to see response
|
||||
```
|
||||
|
||||
**Alert not firing:**
|
||||
```bash
|
||||
# Check PrometheusRule is loaded
|
||||
k get prometheusrule -n myapp-ns
|
||||
|
||||
# Verify query in Prometheus
|
||||
kubectl port-forward -n monitoring svc/prometheus-kube-prom-prometheus 9090:9090
|
||||
# http://localhost:9090 → Graph tab → paste query
|
||||
```
|
||||
|
||||
See `/TROUBLESHOOTING.md` for full incident guide.
|
||||
Reference in New Issue
Block a user