Files
homelab/k8s/llm/README.md
T
Story Crater Bot 61be24e27f k8s/services: add ingress networking portainer llm and project guides
- Nginx ingress + TLS termination (homelab-ca)
- Portainer container UI
- CoreDNS internal DNS rewrites
- DuckDNS DDNS updater
- Ollama LLM inference
- 8 project-usage guides (team reference)
2026-08-18 15:08:00 -07:00

7.0 KiB

Ollama LLM Inference Service

CPU-only LLM inference server on talos-cp-1. Single model hot-loaded (DeepSeek-R1:70b), 42GB, 70Gi memory limit.

Quick Start

Access via port-forward

kubectl -n llm port-forward svc/ollama 11434:11434
curl http://localhost:11434/api/tags

Debug pod (in-cluster)

kubectl run debug --rm -it -n llm --image=curlimages/curl \
  --labels="app.kubernetes.io/role=llm-debug" \
  --serviceaccount=llm-worker -- sh

# Inside pod
TOKEN=$(cat /var/run/secrets/kubernetes.io/serviceaccount/token)
curl -H "Authorization: Bearer $TOKEN" \
  http://ollama.llm.svc.cluster.local:11434/api/tags

Architecture

Component Value
Service ClusterIP ollama.llm.svc.cluster.local:11434
Namespace llm
Node talos-cp-1 (pinned via nodeAffinity)
Memory request 50Gi
Memory limit 70Gi
Storage 115Gi PVC (Longhorn)
Model deepseek-r1:70b (~42GB)
Max loaded 1 model
Parallelism 1 request at a time

API Endpoints

List models

curl http://ollama.llm.svc.cluster.local:11434/api/tags

Response:

{
  "models": [
    {"name": "deepseek-r1:70b", "size": 42000000000, ...}
  ]
}

Generate (non-streaming)

curl -X POST http://ollama.llm.svc.cluster.local:11434/api/generate \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-r1:70b",
    "prompt": "Why is the sky blue?",
    "stream": false
  }'

Pull model

curl -X POST http://ollama.llm.svc.cluster.local:11434/api/pull \
  -H "Content-Type: application/json" \
  -d '{"name": "deepseek-r1:70b", "stream": false}'

Operations

Check pod status

kubectl -n llm get pod -l app.kubernetes.io/name=ollama
kubectl -n llm describe pod -l app.kubernetes.io/name=ollama

View logs

kubectl -n llm logs deployment/ollama -f

Monitor download progress (bootstrap)

kubectl -n llm logs -f job/bootstrap-models -c model-download

Restart deployment

kubectl -n llm rollout restart deployment/ollama

Storage

  • PVC: ollama-models-cache, 115Gi, Longhorn StorageClass
  • Mount: /root/.ollama/models (Ollama model cache)
  • Lifecycle: RWO (Read-Write-Once), tied to talos-cp-1

Resize PVC

⚠️ PVCs can only expand, not shrink. Edit values.yaml and redeploy:

pvc:
  size: 120Gi  # increase only
vsource .env && helmfile -f helmfile.yaml.gotmpl -l name=ollama apply

Networking

NetworkPolicy

  • Default-deny ingress on Ollama pods
  • Allow from pods labeled app.kubernetes.io/name: llm-worker (port 11434)
  • Allow from pods labeled app.kubernetes.io/role: llm-debug (port 11434)

View policy:

kubectl -n llm get networkpolicy ollama

Test access from external pod (should fail):

kubectl run test --rm -it --image=curlimages/curl -- \
  curl http://ollama.llm.svc.cluster.local:11434/
# Connection timeout (correct)

Test access from debug pod (should succeed):

kubectl -n llm logs job/bootstrap-models  # verify bootstrap completed
# Then run debug pod as shown above

Configuration

Helm values (k8s/llm/charts/ollama/values.yaml)

resources:
  requests:
    cpu: 8
    memory: 50Gi
  limits:
    cpu: 16
    memory: 70Gi

env:
  OLLAMA_MAX_LOADED_MODELS: "1"
  OLLAMA_NUM_PARALLEL: "1"
  OLLAMA_MAX_QUEUE: "32"
  OLLAMA_KEEP_ALIVE: "-1"
  OLLAMA_HOST: "0.0.0.0:11434"

preloadJob:
  enabled: true
  hotModels:
    - deepseek-r1:70b

Environment variables

Variable Value Purpose
OLLAMA_MODELS /root/.ollama/models Model cache dir
OLLAMA_MAX_LOADED_MODELS 1 Max concurrent models in RAM
OLLAMA_NUM_PARALLEL 1 Parallel request threads
OLLAMA_MAX_QUEUE 32 Request queue depth
OLLAMA_KEEP_ALIVE -1 Keep model resident (never unload)
OLLAMA_HOST 0.0.0.0:11434 Bind address

Tune OLLAMA_NUM_PARALLEL based on CPU cores. Current: 1 (conservative, CPU bottleneck).

Model Management

Current model

  • Name: deepseek-r1:70b
  • Size: ~42GB
  • Quantization: Default Ollama quant
  • Status: Downloaded during pod init via bootstrap job

Change model

  1. Edit values.yaml:
preloadJob:
  hotModels:
    - deepseek-r1:32b  # or any available model
  1. Redeploy:
kubectl -n llm delete job bootstrap-models --ignore-not-found
vsource .env && helmfile -f helmfile.yaml.gotmpl -l name=ollama apply
  1. Monitor:
kubectl -n llm logs -f job/bootstrap-models -c model-download

Available models

Ollama registry: https://ollama.com/library

Examples:

  • deepseek-r1:70b (reasoning, 42GB)
  • deepseek-r1:32b (faster, 20GB)
  • llama3.1:70b (general, 41GB)
  • mistral:large (26GB)

Troubleshooting

Pod stuck in ContainerCreating

kubectl -n llm describe pod -l app.kubernetes.io/name=ollama
# Check Events section for PVC/image pull issues

Bootstrap job failing

kubectl -n llm logs job/bootstrap-models -c model-download --tail=50
# Common: model not found in registry, disk full, network timeout

Model pull timeout

# Increase pod timeout (edit deployment directly)
kubectl -n llm edit deployment ollama
# Change readinessProbe.initialDelaySeconds, livenessProbe.periodSeconds

Out of memory

Model size exceeds limit. Reduce memory.limits or choose smaller model.

kubectl top pod -n llm  # check actual usage

Cannot connect from other pods

Verify NetworkPolicy:

kubectl -n llm get networkpolicy
kubectl -n llm describe networkpolicy ollama
# Add pod label: app.kubernetes.io/name: llm-worker or app.kubernetes.io/role: llm-debug

Secrets

Ollama pod receives MinIO credentials via Secret ollama-minio (created by helmfile presync):

kubectl -n llm get secret ollama-minio -o jsonpath='{.data}' | jq

Keys: endpoint, bucket, access_key, secret_key

Used by bootstrap job to upload model blobs to MinIO (future: auto-backup).

Metrics & Observability

Prometheus scrape (if enabled)

ServiceMonitor: Not yet configured (see k8s/monitoring/dashboards/services/)

Metrics to add:

  • ollama_requests_total (counter)
  • ollama_request_duration_seconds (histogram)
  • ollama_loaded_models (gauge)

Logs

Pod logs via kubectl. No log aggregation to Loki yet.

kubectl -n llm logs deployment/ollama -f --timestamps

Cleanup

Delete Ollama completely

vsource .env && helmfile -f helmfile.yaml.gotmpl -l name=ollama destroy
# Keeps PVC (data safety). To delete: kubectl -n llm delete pvc ollama-models-cache

Delete just the model cache (keep deployment)

kubectl -n llm delete pvc ollama-models-cache
# Recreate: kubectl -n llm patch deployment ollama -p '{"spec":{"template":{"metadata":{"annotations":{"restart":"now"}}}}}'

See Also

  • Helmfile: helmfile.yaml.gotmpl (llm release block)
  • Chart: k8s/llm/charts/ollama/
  • Namespace: llm
  • Bootstrap: k8s/llm/bootstrap-models-job.yaml (manual preload fallback)