# SLA: API Gateway & Platform Services ## API Gateway (api.riotpiao.com) ### Availability | Target | Measurement | Alert | |--------|------------|-------| | 99.9% uptime | `probe_success{instance=~".*api.riotpiao.com.*"}` | `APIGatewayProbeDown` fires after 2m down | | Monthly budget: 43.8 min downtime | 7-day SLO: `avg_over_time(probe_success[7d]) * 100` | | | Zero ready pods = critical | `sum(kube_pod_status_ready{namespace="api"}) == 0` | `APIGatewayDown` fires after 1m | ### Latency Baselines measured from 200-request canary run against live cluster. SLA set at ~2x measured p99 for headroom. | Endpoint | Measured p50 | Measured p99 | SLA (p95) | SLA (p99) | Alert | |----------|-------------|-------------|-----------|-----------|-------| | LLM Chat (qwen) | 514ms | 609ms | <1s | <2s | `APIGatewayLatencyHigh` | | LLM Chat (reasoning) | 300ms | 328ms | <1s | <2s | `APIGatewayLatencyHigh` | | LLM Chat (ornith:35b) | 1.2s | 1.2s | <3s | <5s | `APIGatewayLatencyCritical` | | LLM Chat (streaming) | 569ms | 628ms | <1s | <2s | `APIGatewayLatencyHigh` | | Embeddings | 189ms | 287ms | <500ms | <1s | `APIGatewayLatencyHigh` | | Rerank | 106ms | 218ms | <500ms | <1s | `APIGatewayLatencyHigh` | | Models list | 68ms | 277ms | <300ms | <500ms | `APIGatewayLatencyHigh` | | Auth rejection | 69ms | 87ms | <200ms | <500ms | (no alert, expected fast) | ### Error Rate | Target | Measurement | Alert | |--------|------------|-------| | 5xx < 1% | `nginx_ingress_controller_requests{status=~"5.."}` / total | `APIGateway5xxErrorRate` fires after 5m >1% | | Total errors < 10% | 4xx + 5xx / total | `APIGatewayHighErrorRate` fires after 10m >10% | --- ## LLM Serving (llm-serving namespace) | Target | Measurement | Alert | |--------|------------|-------| | All predictors running | replicas ready == desired per deployment | `LLMPredictorDown` fires after 5m | | Zero LLM pods = critical | `sum(ready{namespace="llm-serving"}) == 0` | `LLMServingDown` fires after 2m | | No restart storms | restart count in 15m | `LLMPredictorRestarted` on any restart | --- ## Cluster Infrastructure ### Node Health | Target | Measurement | Alert | |--------|------------|-------| | All nodes Ready | `kube_node_status_condition` | `NodeNotReady` fires after 2m | | CPU < 90% sustained | `node_cpu_seconds_total` | `NodeHighCPU` fires after 15m | | Memory < 90% sustained | `node_memory_MemAvailable_bytes` | `NodeHighMemory` fires after 15m | | Disk < 85% | `node_filesystem_avail_bytes` | `NodeDiskFull` fires after 5m (critical) | ### Pod Health | Target | Measurement | Alert | |--------|------------|-------| | No pods pending > 10m | `kube_pod_status_phase{phase="Pending"}` | `PodStuckPending` | | No CrashLoopBackOff > 5m | `kube_pod_container_status_waiting_reason` | `PodCrashLooping` (critical) | | OOMKilled < 3/hour | `kube_pod_container_status_last_terminated_reason` | `OOMKilledSpike` | | No restart storms | >5 restarts in 15m | `ContainerRestartStorm` | ### Jobs | Target | Measurement | Alert | |--------|------------|-------| | No failed jobs | `kube_job_status_failed > 0` | `JobFailed` fires after 5m | | No stuck jobs > 2h | `kube_job_status_active` + age | `JobStuckRunning` | | CronJobs on schedule | last_schedule vs next_schedule | `CronJobMissedSchedule` fires after 10m | ### Storage | Target | Measurement | Alert | |--------|------------|-------| | Longhorn drives healthy | `longhorn_disk_health` | `LonghornDriveOffline` fires after 5m (critical) | ### DNS | Target | Measurement | Alert | |--------|------------|-------| | CoreDNS SERVFAIL < 0.5/s | `coredns_dns_responses_total{rcode="SERVFAIL"}` | `CoreDNSErrorSpike` fires after 5m | ### Probes | Target | Measurement | Alert | |--------|------------|-------| | All service probes passing | `probe_success` | `ServiceProbeDown` fires after 3m (critical) | | Probe latency < 2s | `probe_duration_seconds` | `ServiceProbeSlow` fires after 5m | | Certs valid > 14 days | `certmanager_certificate_expiration_timestamp_seconds` | `CertificateExpiringSoon` | --- ## Alert Severity Levels | Severity | Meaning | Response Time | |----------|---------|--------------| | **critical** | Service down or data loss risk. Immediate impact on users. | Investigate within 15 min | | **warning** | Degraded performance or resource pressure. No immediate outage. | Investigate within 4 hours | ### Critical Alerts (require immediate action) - `APIGatewayDown` — zero gateway pods - `LLMServingDown` — zero LLM pods - `NodeNotReady` — node lost - `PodCrashLooping` — service crashing repeatedly - `NodeDiskFull` — disk > 85% - `LonghornDriveOffline` — storage unhealthy - `ServiceProbeDown` — external service unreachable - `APIGateway5xxErrorRate` — 5xx > 1% - `APIGatewayLatencyCritical` — p99 > 5s --- ## Current Alert Status Alerts firing after deployment: | Alert | State | Root Cause | |-------|-------|-----------| | `APIGatewayProbeDown` | pending | Blackbox probe for api-gateway not yet active (pod restart needed) | | `PodStuckPending` | pending | `sms/macos-bluebubbles` pending 22d (scheduling constraint) | | `PodCrashLooping` | pending | `iam/authentik-provision` job in Error state | | `DeploymentReplicasUnavailable` | pending | Same root causes above | | `ServiceProbeDown` | pending | api-gateway probe target not in blackbox yet | None are false positives. All reflect real cluster state.