# Phase 7 Operations Runbook Incident response and troubleshooting for versioning, audit, ranking, and rebuild systems. --- ## Incident: Rebuild Parity Failure **Severity**: CRITICAL | **Impact**: Data integrity at risk | **SLO**: Consistency 100% ### Detection Alert: `RebuildParityFailed` Symptoms: - Daily CI rebuild check fails - Checksums don't match before/after - Diff summary shows unexpected changes ### Immediate Actions (0-5 minutes) 1. **Acknowledge and declare incident** ```bash # In war room: incident declare rebuild-parity-$(date +%s) ``` 2. **Assess scope** ```bash # Get last successful rebuild curl -H "Authorization: Bearer $TOKEN" \ http://localhost:8080/memory/rebuild/status \ | jq '.last_rebuild | {rebuild_id, checksum, records_processed}' ``` 3. **Block writes** (if corruption suspected) ```bash # Stop ingest pipeline kubectl scale deployment memory-ingest --replicas=0 -n poimen ``` ### Investigation (5-30 minutes) 1. **Check logs for non-determinism** ```bash # Look for embedding model version changes kubectl logs -l app=memory -c worker -n poimen --since=24h | grep -i "embedding\|model\|version" # Check for floating point precision issues kubectl logs -l app=memory -c worker -n poimen --since=24h | grep -i "float\|precision\|nan" ``` 2. **Verify event log integrity** ```bash # Count records in event log vs database EVENTS=$(psql $DB -c "SELECT COUNT(*) FROM event_log WHERE project_id='poimen'" --csv | tail -1) ENTITIES=$(psql $DB -c "SELECT COUNT(*) FROM memory_entity WHERE project_id='poimen'" --csv | tail -1) echo "Event log: $EVENTS, DB entities: $ENTITIES" [ "$EVENTS" -eq "$ENTITIES" ] || echo "MISMATCH: corruption likely" ``` 3. **Check for dependency changes** ```bash # Get current container image versions kubectl get pods -l app=memory -o jsonpath='{.items[*].spec.containers[*].image}' -n poimen # Compare to expected (git tag) git show HEAD:deploy/memory-deployment.yaml | grep image: ``` ### Resolution **If embedding model changed** (likely cause): ```bash # Option 1: Revert to previous model version kubectl set env deployment memory EMBEDDINGS_MODEL=sentence-transformers/all-MiniLM-L6-v2:v0.1 -n poimen kubectl rollout restart deployment memory -n poimen # Wait for rebuild to complete sleep 300 # Re-run verification curl -X POST http://localhost:8080/memory/rebuild \ -H "Authorization: Bearer $TOKEN" \ -d '{"project":"poimen","verify":true}' ``` **If event log corrupted** (rare): ```bash # Restore from backup kubectl exec -it memory-backup-pod /bin/bash << 'EOF' pg_restore -d memory /backups/memory-$(date -d '1 day ago' +%Y-%m-%d).sql EOF # Verify curl -X POST http://localhost:8080/memory/rebuild \ -H "Authorization: Bearer $TOKEN" \ -d '{"project":"poimen","verify":true}' ``` **If timestamp drift** (clock skew): ```bash # Verify NTP sync on all nodes timedatectl status chronyc tracking # If unsync, force resync chronyc -a makestep # Re-run rebuild curl -X POST http://localhost:8080/memory/rebuild \ -H "Authorization: Bearer $TOKEN" \ -d '{"project":"poimen","verify":true}' ``` ### Prevention - Pin embedding model version in deployment - Automated tests for embedding determinism - Daily rebuild CI (catch early) - Event log backups (hourly) ### Escalation If root cause unknown after 30 min: ```bash # Escalate to engineering pagerduty trigger --title "Rebuild parity: unknown cause" \ --description "$(curl -s http://localhost:8080/memory/rebuild/status | jq -c .)" ``` --- ## Incident: Audit Chain Broken **Severity**: CRITICAL | **Impact**: Compliance violation | **SLO**: Consistency 100% ### Detection Alert: `AuditChainBroken` Symptoms: - Audit integrity check fails - Specific entity version has broken hash chain - Compliance audit fails ### Immediate Actions (0-5 minutes) 1. **Identify affected entity** ```bash # Get entity with broken chain ENTITY=$(curl -s http://localhost:8080/memory/audit/verify \ -H "Authorization: Bearer $TOKEN" \ | jq -r '.first_broken_at.entity_id') echo "Affected: $ENTITY" ``` 2. **Document for compliance** ```bash # Create incident record cat > /tmp/audit-incident.json << EOF { "timestamp": "$(date -Iseconds)", "entity_id": "$ENTITY", "broken_at_version": $(curl -s http://localhost:8080/memory/entities/$ENTITY/versions \ -H "Authorization: Bearer $TOKEN" | jq '.total'), "severity": "critical", "action": "see runbook" } EOF # Save for audit trail cp /tmp/audit-incident.json /var/log/poimen/audit-incident-$(date +%s).json ``` ### Investigation (5-30 minutes) 1. **Check for unauthorized writes** ```bash # Get who made the change curl -s http://localhost:8080/memory/entities/$ENTITY/versions \ -H "Authorization: Bearer $TOKEN" \ | jq '.versions[] | {version_num, changed_by, changed_at}' ``` 2. **Check for DB trigger bypass** ```bash # Verify immutability trigger exists psql $DB -c "SELECT trigger_name, event_object_table FROM information_schema.triggers WHERE trigger_name LIKE '%immutable%';" # If missing, recreate psql $DB -f crates/mem-store/migrations/007_versioning_schema.sql ``` 3. **Check Authentik logs for unauthorized access** ```bash # Get Authentik audit kubectl logs -l app=authentik -n iam --since=24h | grep -i "$ENTITY\|unauthorized\|denied" ``` ### Resolution **Audit logs are immutable** — cannot repair. Options: 1. **Document incident for compliance** ```bash # Create compliance report cat > /tmp/compliance-report.md << EOF # Audit Chain Integrity Incident **Date**: $(date) **Severity**: Critical **Entity Affected**: $ENTITY **Root Cause**: [Investigation finding] ## Actions Taken 1. Incident documented 2. Authentik access logs reviewed 3. Immutability trigger verified 4. [Preventive action] ## Compliance Impact - Audit trail for $ENTITY versions is compromised - Recommend manual review of $ENTITY history - All future versions protected by restored trigger EOF # Store in compliance folder cp /tmp/compliance-report.md /var/log/poimen/compliance-incidents/ ``` 2. **Restore trigger and lock down** ```bash # Re-create immutability trigger psql $DB -f crates/mem-store/migrations/007_versioning_schema.sql # Verify it worked psql $DB -c "UPDATE memory_entity_version SET operation='test' LIMIT 1" || echo "Trigger working" ``` 3. **Investigate root cause** - Was trigger accidentally dropped? - Was there an emergency maintenance window? - Was there an accidental SQL injection? - Was there a permission escalation? ### Prevention - Immutability trigger on all audit tables - Regular trigger verification (weekly) - Authentik audit log retention (1 year) - Database role separation (no direct table updates) ### Escalation After confirmation of compromise: ```bash # Notify compliance/legal team notify compliance-team << EOF Audit chain integrity compromised for entity: $ENTITY See: /var/log/poimen/compliance-incidents/audit-chain-$(date +%Y-%m-%d).md EOF ``` --- ## Incident: Version Query Timeout **Severity**: MEDIUM | **Impact**: Slow audits | **SLO**: Latency p99 < 500ms ### Detection Alert: `MemoryVersionLatencyHigh` Symptoms: - `GET /memory/entities/{id}/versions` takes > 500ms - Audit reports slow - Dashboard unresponsive ### Investigation 1. **Check query performance** ```bash # Explain the query psql $DB << EOF EXPLAIN ANALYZE SELECT * FROM memory_entity_version WHERE entity_id = 'e_kubernetes' ORDER BY version_num DESC; EOF ``` 2. **Check index status** ```bash # Verify indexes exist and are healthy psql $DB -c "SELECT schemaname, tablename, indexname FROM pg_indexes WHERE tablename LIKE 'memory_entity_version';" # Check if bloated psql $DB -c "SELECT * FROM pgstattuple('memory_entity_version');" ``` ### Resolution 1. **If indexes missing or bloated** ```bash # Rebuild indexes psql $DB << EOF REINDEX TABLE memory_entity_version; REINDEX TABLE memory_edge_version; ANALYZE memory_entity_version; ANALYZE memory_edge_version; EOF ``` 2. **If table too large, partition by entity_id** ```bash # Add range partition (future improvement) # For now, truncate old versions psql $DB << EOF DELETE FROM memory_entity_version WHERE changed_at < NOW() - INTERVAL '1 year'; EOF ``` --- ## Incident: Ranking Profile Not Working **Severity**: LOW | **Impact**: Search quality | **SLO**: None (feature)** ### Detection Symptoms: - `GET /memory/ranking/profiles` returns empty - Query with `ranking_profile: "recency_focused"` returns 400 - Ranking signals are all 0 ### Investigation 1. **Check profile endpoint** ```bash curl -v http://localhost:8080/memory/ranking/profiles \ -H "Authorization: Bearer $TOKEN" ``` 2. **Check signal computation in logs** ```bash kubectl logs -l app=memory -c worker -n poimen --tail=100 | grep -i signal ``` ### Resolution 1. **Restart ranking service** ```bash kubectl rollout restart deployment memory -n poimen ``` 2. **Verify profiles loaded** ```bash curl http://localhost:8080/memory/ranking/profiles \ -H "Authorization: Bearer $TOKEN" | jq '.profiles | length' # Should return 3 ``` --- ## Routine: Daily Health Check Run every day at 1 AM UTC: ```bash #!/bin/bash set -e TOKEN=$(get_jwt_token) API=http://localhost:8080 echo "=== Daily Memory Health Check ===" # 1. Rebuild verification echo "1. Testing rebuild parity..." REBUILD=$(curl -s -X POST $API/memory/rebuild \ -H "Authorization: Bearer $TOKEN" \ -d '{"project":"poimen","verify":true}') MATCH=$(echo $REBUILD | jq -r '.checksum.match') if [ "$MATCH" != "true" ]; then echo "❌ REBUILD PARITY FAILED" echo $REBUILD | jq '.diff_summary' exit 1 fi echo "✅ Rebuild parity: OK" # 2. Audit chain echo "2. Testing audit chain..." AUDIT=$(curl -s http://localhost:8080/memory/audit/verify \ -H "Authorization: Bearer $TOKEN") VALID=$(echo $AUDIT | jq -r '.chain_valid') if [ "$VALID" != "true" ]; then echo "❌ AUDIT CHAIN BROKEN" exit 1 fi echo "✅ Audit chain: OK" # 3. Latency check echo "3. Testing latency..." START=$(date +%s%N) curl -s $API/memory/entities/e_kubernetes/versions \ -H "Authorization: Bearer $TOKEN" > /dev/null END=$(date +%s%N) ELAPSED_MS=$(( (END - START) / 1000000 )) if [ $ELAPSED_MS -gt 500 ]; then echo "⚠️ Version query slow: ${ELAPSED_MS}ms" else echo "✅ Latency: OK (${ELAPSED_MS}ms)" fi # 4. Disk space echo "4. Checking disk..." USAGE=$(du -sh /var/lib/postgresql | cut -f1) echo " Database size: $USAGE" echo "" echo "=== Health check complete ===" ``` --- ## SLO Compliance Monitor these daily: ```promql # Availability memory:availability:slo >= 0.999 # Latency memory:query:latency:p99 < 0.2 memory:version:latency:p99 < 0.5 # Consistency memory:rebuild:success_rate == 1.0 memory:audit:chain_valid_rate == 1.0 ``` If any SLO breached, escalate to on-call engineer. --- ## Contact - **On-call**: See PagerDuty schedule - **Slack**: #poimen-alerts - **Incident**: incident declare phase7-\*