
Model Serving Kubernetes
- 88 installs
- 44 repo stars
- Updated May 22, 2026
- bagelhole/devops-security-agent-skills
model-serving-kubernetes is a Claude Code skill for deploying ML models on Kubernetes with KServe and NVIDIA Triton, including canary deployments, autoscaling, versioning, A/B testing, and GPU resource management.
About
This skill deploys ML models on Kubernetes using KServe and NVIDIA Triton Inference Server. It covers InferenceServices, canary deployments with traffic splitting, KEDA autoscaling, and GPU resource management for production serving. A developer uses it to serve scikit-learn, PyTorch, ONNX, or LLM models at scale with versioning and A/B testing. It matters for running inference reliably across multiple model versions on GPU clusters.
- Deploys ML models on Kubernetes with KServe and NVIDIA Triton, including GPU-aware scheduling
- Covers canary deployments, traffic splitting, A/B testing, and KEDA autoscaling on Prometheus metrics
- Serves LLMs via vLLM InferenceServices with GPU resource requests and readiness probes
Model Serving Kubernetes by the numbers
- 88 all-time installs (skills.sh)
- Ranked #594 of 1,039 Cloud & Infrastructure skills by installs in the Skillselion catalog
- Data as of Jul 28, 2026 (Skillselion catalog sync)
model-serving-kubernetes capabilities & compatibility
- Capabilities
- model serving · canary deployment · autoscaling · gpu scheduling
- Works with
- kubernetes · docker · grafana
- Use cases
- devops · ci cd
- Pricing
- Free
What model-serving-kubernetes says it does
Production ML model serving with KServe and Triton — canary deployments, autoscaling, and GPU-aware scheduling.
canaryTrafficPercent: 20 # 20% to new version, 80% to stable
npx skills add https://github.com/bagelhole/devops-security-agent-skills --skill model-serving-kubernetesAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 88 |
|---|---|
| repo stars | ★ 44 |
| Last updated | May 22, 2026 |
| Repository | bagelhole/devops-security-agent-skills ↗ |
What it does
Serve ML models and LLMs at scale on Kubernetes with KServe/Triton, canary rollouts, autoscaling, and GPU scheduling.
Who is it for?
Teams serving scikit-learn, PyTorch, TensorFlow, ONNX, or LLM models at scale on Kubernetes with GPU nodes.
Skip if: Single-model local inference or environments without a Kubernetes cluster and GPU nodes.
When should I use this skill?
Serving ML models at scale, implementing canary/A-B testing, or autoscaling inference pods on GPU metrics.
What you get
Production model serving on Kubernetes with KServe/Triton, traffic-split canaries, KEDA autoscaling, and GPU-aware scheduling.
By the numbers
- supports 4 model frameworks (scikit-learn, PyTorch, TensorFlow, ONNX)
- requires Kubernetes 1.28+
Files
Model Serving on Kubernetes
Production ML model serving with KServe and Triton — canary deployments, autoscaling, and GPU-aware scheduling.
When to Use This Skill
Use this skill when:
- Serving scikit-learn, PyTorch, TensorFlow, or ONNX models at scale
- Implementing canary deployments and A/B testing for ML models
- Autoscaling inference pods based on request rate or GPU metrics
- Deploying LLMs with Triton or KServe on Kubernetes
- Managing multiple model versions with traffic splitting
Prerequisites
- Kubernetes 1.28+ with GPU nodes
- KServe installed (or Triton standalone)
kubectlandhelmconfigured- NVIDIA GPU Operator installed on cluster
KServe Installation
# Install KServe with Helm
helm repo add kserve https://kserve.github.io/helm-charts
helm repo update
helm install kserve kserve/kserve \
--namespace kserve \
--create-namespace \
--set kserve.controller.gateway.ingressGateway.className=nginx
# Verify
kubectl get pods -n kserve
kubectl get crd | grep kserveBasic InferenceService (KServe)
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
name: sklearn-iris
namespace: models
spec:
predictor:
sklearn:
storageUri: gs://kfserving-examples/models/sklearn/1.0/model
resources:
requests:
cpu: "1"
memory: 2Gi
limits:
cpu: "2"
memory: 4Gikubectl apply -f inference-service.yaml
# Get inference service URL
kubectl get inferenceservice sklearn-iris -n models
# NAME URL READY ...
# sklearn-iris http://sklearn-iris.models.example.com True
# Test prediction
curl -X POST http://sklearn-iris.models.example.com/v1/models/sklearn-iris:predict \
-H "Content-Type: application/json" \
-d '{"instances": [[6.8, 2.8, 4.8, 1.4]]}'GPU-Enabled LLM InferenceService
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
name: llama-3-8b
namespace: models
annotations:
serving.kserve.io/enable-prometheus-scraping: "true"
spec:
predictor:
containers:
- name: vllm-container
image: vllm/vllm-openai:latest
args:
- "--model"
- "meta-llama/Llama-3.1-8B-Instruct"
- "--tensor-parallel-size"
- "1"
- "--gpu-memory-utilization"
- "0.90"
ports:
- containerPort: 8080
protocol: TCP
resources:
requests:
nvidia.com/gpu: "1"
memory: "20Gi"
cpu: "4"
limits:
nvidia.com/gpu: "1"
memory: "24Gi"
cpu: "8"
readinessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 60
periodSeconds: 10
env:
- name: HUGGING_FACE_HUB_TOKEN
valueFrom:
secretKeyRef:
name: hf-token
key: token
nodeSelector:
nvidia.com/gpu.present: "true"
transformer:
containers:
- name: kserve-container
image: kserve/kserve-transformer:latestCanary Deployment (Traffic Splitting)
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
name: llama-3-8b
namespace: models
spec:
predictor:
canaryTrafficPercent: 20 # 20% to new version, 80% to stable
containers:
- name: vllm-container
image: vllm/vllm-openai:latest
args:
- "--model"
- "meta-llama/Llama-3.1-8B-Instruct-v2" # new model version
resources:
limits:
nvidia.com/gpu: "1"# Gradually increase canary traffic
kubectl patch inferenceservice llama-3-8b -n models \
--type='json' \
-p='[{"op":"replace","path":"/spec/predictor/canaryTrafficPercent","value":50}]'
# Promote canary to stable
kubectl patch inferenceservice llama-3-8b -n models \
--type='json' \
-p='[{"op":"remove","path":"/spec/predictor/canaryTrafficPercent"}]'Autoscaling with KEDA
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: llama-scaler
namespace: models
spec:
scaleTargetRef:
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
name: llama-3-8b
minReplicaCount: 1
maxReplicaCount: 5
triggers:
- type: prometheus
metadata:
serverAddress: http://prometheus-server.monitoring:9090
metricName: kserve_request_count
threshold: "10"
query: |
sum(rate(kserve_request_count_total{namespace="models",
service="llama-3-8b"}[1m]))NVIDIA Triton Inference Server
apiVersion: apps/v1
kind: Deployment
metadata:
name: triton-server
namespace: models
spec:
replicas: 2
selector:
matchLabels:
app: triton
template:
metadata:
labels:
app: triton
spec:
containers:
- name: triton
image: nvcr.io/nvidia/tritonserver:24.05-py3
args:
- "tritonserver"
- "--model-store=s3://my-model-store/models"
- "--model-control-mode=poll" # auto-load new model versions
- "--repository-poll-secs=30"
- "--metrics-port=8002"
ports:
- containerPort: 8000 # HTTP
- containerPort: 8001 # gRPC
- containerPort: 8002 # Metrics
resources:
limits:
nvidia.com/gpu: "1"
readinessProbe:
httpGet:
path: /v2/health/ready
port: 8000
initialDelaySeconds: 30Triton Model Repository Structure
s3://my-model-store/models/
├── text-classifier/
│ ├── config.pbtxt
│ ├── 1/
│ │ └── model.onnx
│ └── 2/
│ └── model.onnx # new version; auto-loaded
├── embedding-model/
│ ├── config.pbtxt
│ └── 1/
│ └── model.onnx# config.pbtxt for ONNX model
name: "text-classifier"
backend: "onnxruntime"
max_batch_size: 64
dynamic_batching {
preferred_batch_size: [16, 32]
max_queue_delay_microseconds: 1000
}
input [
{ name: "input_ids" data_type: TYPE_INT64 dims: [-1] }
{ name: "attention_mask" data_type: TYPE_INT64 dims: [-1] }
]
output [
{ name: "logits" data_type: TYPE_FP32 dims: [-1] }
]
instance_group [
{ kind: KIND_GPU count: 2 } # 2 model instances on GPU
]Model Management Commands
# List loaded models (Triton)
curl http://triton:8000/v2/models
# Load a new model version
curl -X POST http://triton:8000/v2/repository/models/text-classifier/load
# Unload a model
curl -X POST http://triton:8000/v2/repository/models/text-classifier/unload
# KServe — watch rollout status
kubectl rollout status deployment/llama-3-8b-predictor -n models
kubectl get inferenceservice llama-3-8b -n models -wCommon Issues
| Issue | Cause | Fix |
|---|---|---|
InferenceService not ready | Model loading or OOM | Check predictor pod logs; increase memory limits |
| Canary stuck at 0% | KNative routing issue | Check kubectl get ksvc -n models |
| Triton missing model | S3 permissions or path | Verify IAM role; check --model-store path |
| Low GPU utilization | Dynamic batching off | Enable dynamic_batching in Triton config |
| Autoscaler not triggering | Prometheus query wrong | Test query in Prometheus UI |
Best Practices
- Use canary deployments for all model updates — roll back in seconds if metrics degrade.
- Enable Triton dynamic batching — it can increase GPU throughput 5–10× for small models.
- Store models in S3/GCS with versioned paths (
s3://bucket/model/v1/,v2/). - Pin GPU node selectors to prevent model pods landing on CPU-only nodes.
- Monitor p99 latency and error rates per model version during canary rollouts.
Related Skills
- vllm-server - vLLM for LLM serving
- llm-inference-scaling - KEDA autoscaling
- kubernetes-ops - Core Kubernetes operations
- gpu-server-management - GPU nodes
Related skills
FAQ
What frameworks can this serve?
The skill covers serving scikit-learn, PyTorch, TensorFlow, or ONNX models at scale, plus deploying LLMs with Triton or KServe on Kubernetes.
How does canary deployment work in KServe?
You set canaryTrafficPercent on the predictor (for example 20% to the new version, 80% to stable), gradually patch it higher, then remove it to promote the canary to stable.