This website requires JavaScript.
Explore
Help
Register
Sign In
rock
/
homelab
Watch
1
Star
0
Fork
0
Code
Issues
Pull Requests
1
Actions
Packages
Projects
Releases
Wiki
Activity
Files
feat/gpu-inference-stack
homelab
/
k8s
/
apps
/
llm-serving
/
namespace.yaml
T
Add File
New File
Upload File
Apply Patch
Copy Permalink
Download directory as ZIP
Download directory as TAR.GZ
Delete Directory
5 lines
61 B
YAML
Raw
Permalink
Normal View
History
Unescape
Escape
feat(gpu): serve 6 models on worker-1 via KServe — vLLM v0.11.0 (bitsandbytes) + Ollama + TEI, plus RuntimeClass/privileged-PSA prereqs and a local-NVMe StorageClass, working around Volta sm_70 limits
2026-08-13 07:02:53 -07:00
apiVersion
:
v1
kind
:
Namespace
metadata
:
name
:
llm-serving
Reference in New Issue
Copy Permalink