Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
bagelhole avatar

Llm Inference Scaling

  • 90 installs
  • 44 repo stars
  • Updated May 22, 2026
  • bagelhole/devops-security-agent-skills

llm-inference-scaling is a Claude skill that auto-scales LLM inference clusters on Kubernetes using KEDA, GPU metrics, horizontal pod autoscaling and spot instances.

About

llm-inference-scaling is a skill that auto-scales LLM inference clusters on Kubernetes using KEDA, GPU metrics and horizontal pod autoscaling. It shows how to deploy vLLM on GPU nodes, scale on queue depth and KV-cache usage via Prometheus, run queue-based batch scaling with Redis, and use spot instances for cost efficiency. A developer uses it to handle unpredictable inference traffic on a GPU fleet.

  • vLLM deployment on GPU nodes with the NVIDIA GPU Operator
  • KEDA autoscaling on vLLM queue depth and KV cache metrics
  • Spot-instance node affinity and cluster autoscaler for cost savings

Llm Inference Scaling by the numbers

  • 90 all-time installs (skills.sh)
  • Ranked #588 of 1,039 Cloud & Infrastructure skills by installs in the Skillselion catalog
  • Data as of Jul 28, 2026 (Skillselion catalog sync)
At a glance

llm-inference-scaling capabilities & compatibility

Capabilities
kubernetes ops · llm gateway · llmops platform engineering · load balancing
Works with
kubernetes · grafana · aws
Use cases
devops
Runs
Runs locally
Pricing
Free
From the docs

What llm-inference-scaling says it does

Scale LLM inference horizontally on Kubernetes with GPU-aware autoscaling, request queuing, and cost-efficient spot instance strategies.
SKILL.md
scale up if >10 requests waiting
SKILL.md
npx skills add https://github.com/bagelhole/devops-security-agent-skills --skill llm-inference-scaling

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs90
repo stars44
Last updatedMay 22, 2026
Repositorybagelhole/devops-security-agent-skills

What it does

Auto-scale vLLM/TGI inference pods on Kubernetes with KEDA, GPU metrics and spot instances to handle traffic spikes cheaply.

Who is it for?

Platform teams running a vLLM/TGI GPU fleet on Kubernetes with unpredictable inference traffic

Skip if: Apps that only call hosted inference APIs and manage no GPU pods

When should I use this skill?

LLM traffic is unpredictable, managing a vLLM/TGI fleet, or reducing inference cost with spot GPUs

What you get

GPU pods that scale on queue depth and KV-cache usage, with spot instances cutting cost

  • vLLM GPU Deployment manifests
  • KEDA ScaledObject/ScaledJob config
  • Spot-instance node affinity and cluster autoscaler setup

By the numbers

  • KEDA scales up when >10 requests are waiting
  • Scale-up trigger at KV cache >80%
  • Example maxReplicaCount of 8 for inference pods

Files

SKILL.mdMarkdownGitHub ↗

LLM Inference Scaling

Scale LLM inference horizontally on Kubernetes with GPU-aware autoscaling, request queuing, and cost-efficient spot instance strategies.

When to Use This Skill

Use this skill when:

  • LLM API traffic is unpredictable and you need to scale up/down automatically
  • Managing a fleet of vLLM or TGI inference pods on Kubernetes
  • Reducing inference costs with spot/preemptible GPU instances
  • Implementing queue-based autoscaling for batch inference jobs
  • Building a multi-model serving platform that shares GPU resources

Prerequisites

  • Kubernetes cluster with GPU nodes (NVIDIA operator installed)
  • KEDA (Kubernetes Event-Driven Autoscaler) installed
  • Prometheus with GPU metrics (dcgm-exporter or gpu-operator)
  • Helm 3+ for chart deployments

GPU Node Setup

# Install NVIDIA GPU Operator (handles drivers, container toolkit, DCGM)
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm repo update

helm install gpu-operator nvidia/gpu-operator \
  --namespace gpu-operator \
  --create-namespace \
  --set driver.enabled=true \
  --set dcgm.enabled=true \
  --set devicePlugin.enabled=true

# Verify GPU nodes are recognized
kubectl get nodes -l nvidia.com/gpu.present=true
kubectl describe node <gpu-node> | grep nvidia

vLLM Deployment with GPU Resources

apiVersion: apps/v1
kind: Deployment
metadata:
  name: vllm-llama-8b
  labels:
    app: vllm
    model: llama-3.1-8b
spec:
  replicas: 1
  selector:
    matchLabels:
      app: vllm
      model: llama-3.1-8b
  template:
    metadata:
      labels:
        app: vllm
        model: llama-3.1-8b
    spec:
      nodeSelector:
        nvidia.com/gpu.present: "true"
      tolerations:
      - key: nvidia.com/gpu
        operator: Exists
        effect: NoSchedule
      containers:
      - name: vllm
        image: vllm/vllm-openai:latest
        args:
        - "--model"
        - "meta-llama/Llama-3.1-8B-Instruct"
        - "--tensor-parallel-size"
        - "1"
        - "--gpu-memory-utilization"
        - "0.90"
        - "--max-num-seqs"
        - "128"
        resources:
          requests:
            nvidia.com/gpu: "1"
            memory: "20Gi"
            cpu: "4"
          limits:
            nvidia.com/gpu: "1"
            memory: "24Gi"
            cpu: "8"
        ports:
        - containerPort: 8000
        readinessProbe:
          httpGet:
            path: /health
            port: 8000
          initialDelaySeconds: 60
          periodSeconds: 10
        env:
        - name: HUGGING_FACE_HUB_TOKEN
          valueFrom:
            secretKeyRef:
              name: hf-token
              key: token

KEDA Autoscaling on Prometheus Metrics

apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: vllm-scaledobject
spec:
  scaleTargetRef:
    name: vllm-llama-8b
  minReplicaCount: 1
  maxReplicaCount: 8
  cooldownPeriod: 300          # 5 min before scale-down
  pollingInterval: 15
  triggers:
  - type: prometheus
    metadata:
      serverAddress: http://prometheus-server.monitoring:9090
      metricName: vllm_num_requests_waiting
      threshold: "10"           # scale up if >10 requests waiting
      query: |
        sum(vllm:num_requests_waiting{deployment="vllm-llama-8b"})
  - type: prometheus
    metadata:
      serverAddress: http://prometheus-server.monitoring:9090
      metricName: vllm_gpu_cache_usage
      threshold: "0.8"          # scale up if KV cache >80% full
      query: |
        avg(vllm:gpu_cache_usage_perc{deployment="vllm-llama-8b"})

Queue-Based Scaling (Redis + KEDA)

# ScaledJob for async batch inference
apiVersion: keda.sh/v1alpha1
kind: ScaledJob
metadata:
  name: llm-batch-inference
spec:
  jobTargetRef:
    template:
      spec:
        containers:
        - name: inference-worker
          image: myapp/inference-worker:latest
          env:
          - name: REDIS_URL
            value: redis://redis:6379
          - name: QUEUE_NAME
            value: inference-jobs
        restartPolicy: OnFailure
  minReplicaCount: 0
  maxReplicaCount: 20
  pollingInterval: 5
  successfulJobsHistoryLimit: 3
  triggers:
  - type: redis
    metadata:
      address: redis:6379
      listName: inference-jobs
      listLength: "5"           # 1 worker per 5 queued jobs

Spot Instance Strategy

# Mixed node pool: on-demand + spot GPUs
apiVersion: v1
kind: ConfigMap
metadata:
  name: cluster-autoscaler-priority-config
data:
  priorities: |
    10:  # low priority = prefer
    - .*spot.*
    50:
    - .*on-demand.*
---
# Node affinity for spot with on-demand fallback
spec:
  affinity:
    nodeAffinity:
      preferredDuringSchedulingIgnoredDuringExecution:
      - weight: 80
        preference:
          matchExpressions:
          - key: node.kubernetes.io/lifecycle
            operator: In
            values: [spot]
      - weight: 20
        preference:
          matchExpressions:
          - key: node.kubernetes.io/lifecycle
            operator: In
            values: [on-demand]

Cluster Autoscaler for GPU Nodes

# AWS EKS — enable cluster autoscaler for GPU node group
helm install cluster-autoscaler autoscaler/cluster-autoscaler \
  --namespace kube-system \
  --set autoDiscovery.clusterName=my-cluster \
  --set awsRegion=us-east-1 \
  --set rbac.serviceAccount.annotations."eks\.amazonaws\.com/role-arn"=arn:aws:iam::ACCOUNT:role/ClusterAutoscalerRole \
  --set extraArgs.skip-nodes-with-local-storage=false \
  --set extraArgs.expander=least-waste

# Annotate GPU node group for autoscaler
kubectl annotate node <node> \
  cluster-autoscaler.kubernetes.io/safe-to-evict="false"

Scaling Metrics to Monitor

# Prometheus queries for scaling decisions
# Requests waiting in vLLM queue
sum(vllm:num_requests_waiting) by (model)

# GPU KV cache utilization (>80% = bottleneck)
avg(vllm:gpu_cache_usage_perc) by (pod)

# Tokens per second throughput
sum(rate(vllm:generation_tokens_total[5m])) by (model)

# P99 time-to-first-token
histogram_quantile(0.99, rate(vllm:time_to_first_token_seconds_bucket[5m]))

Common Issues

IssueCauseFix
Pods stuck in PendingNo GPU nodes availableCheck cluster autoscaler logs; verify node group limits
Scale-up too slowCluster autoscaler delay + model load timePre-warm replicas; increase minReplicaCount
GPU fragmentationMultiple small models on large GPUsUse MIG partitioning or consolidate model sizes
Spot eviction causes errorsSpot instance reclamationAdd PodDisruptionBudget; use graceful shutdown
KEDA not scalingPrometheus query returns no dataTest query in Prometheus UI first

Best Practices

  • Set minReplicaCount: 1 to avoid cold starts; scale to 0 only for batch jobs.
  • Use PodDisruptionBudget with minAvailable: 1 to survive spot evictions.
  • Pre-pull model weights into a shared PVC to speed up pod startup by 5–10×.
  • Separate model families across node pools (A10G for 7B, A100 for 70B).
  • Use Kubernetes VPA for CPU/memory right-sizing alongside KEDA for replica count.

Related Skills

  • vllm-server - vLLM configuration and tuning
  • gpu-server-management - GPU node setup
  • model-serving-kubernetes - KServe
  • kubernetes-ops - Core Kubernetes
  • llm-cost-optimization - Cost strategies

Related skills

FAQ

What triggers scale-up?

KEDA scales on Prometheus metrics such as vLLM requests waiting (>10) and GPU KV cache usage (>80%).

How does it reduce cost?

It uses spot/preemptible GPU node affinity with on-demand fallback and a cluster autoscaler tuned for least waste.

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.