Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
bagelhole avatar

Gpu Kubernetes Operations

  • 86 installs
  • 44 repo stars
  • Updated May 22, 2026
  • bagelhole/devops-security-agent-skills

gpu-kubernetes-operations is a Claude Code skill for operating GPU-backed Kubernetes clusters, covering the NVIDIA GPU Operator, MIG partitioning, autoscaling, and DCGM monitoring for AI workloads.

About

A Claude skill for operating GPU-backed Kubernetes clusters that run AI inference and training. A developer uses it to install the NVIDIA GPU Operator, partition GPUs with MIG or time-slicing, build GPU-aware autoscaling, and monitor GPU health with DCGM and Prometheus. It targets Kubernetes 1.28+ clusters with NVIDIA GPUs like A100 and H100.

  • Installs the NVIDIA GPU Operator and device plugin on Kubernetes clusters
  • Configures MIG partitioning and GPU time-slicing to share GPUs across workloads
  • Sets up DCGM/Prometheus GPU health monitoring and alert rules

Gpu Kubernetes Operations by the numbers

  • 86 all-time installs (skills.sh)
  • Ranked #600 of 1,039 Cloud & Infrastructure skills by installs in the Skillselion catalog
  • Data as of Jul 28, 2026 (Skillselion catalog sync)
At a glance

gpu-kubernetes-operations capabilities & compatibility

Capabilities
devops · orchestration
Works with
kubernetes · grafana
Use cases
devops · orchestration
Pricing
Free
From the docs

What gpu-kubernetes-operations says it does

Operate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls.
SKILL.md
The GPU Operator automates driver, toolkit, device plugin, and DCGM deployment.
SKILL.md
npx skills add https://github.com/bagelhole/devops-security-agent-skills --skill gpu-kubernetes-operations

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs86
repo stars44
Last updatedMay 22, 2026
Repositorybagelhole/devops-security-agent-skills

What it does

Operate GPU node pools in Kubernetes for AI inference and training with scheduling, partitioning, and health monitoring.

Who is it for?

Platform teams running AI inference or training on Kubernetes GPU clusters.

Skip if: Single-server GPU setups without Kubernetes.

When should I use this skill?

Setting up GPU node pools, MIG partitioning, or troubleshooting GPU scheduling in Kubernetes.

What you get

A resilient, cost-efficient GPU Kubernetes cluster with the GPU Operator, MIG or time-slicing, and DCGM monitoring.

  • GPU Operator installation
  • MIG or time-slicing configuration
  • DCGM ServiceMonitor and GPU alert rules

By the numbers

  • Ships example MIG configs from 1g.10gb (7 instances) up to full-GPU profiles
  • Time-slicing example advertises 4 virtual GPUs per physical GPU

Files

SKILL.mdMarkdownGitHub ↗

GPU Kubernetes Operations

Run resilient and cost-efficient GPU clusters for production AI workloads.

When to Use This Skill

  • Setting up GPU node pools in Kubernetes for AI inference or training
  • Configuring NVIDIA device plugin and GPU operator
  • Implementing MIG partitioning to share GPUs across workloads
  • Building GPU-aware autoscaling policies
  • Monitoring GPU health with DCGM and Prometheus
  • Troubleshooting GPU scheduling, driver, or OOM issues

Prerequisites

  • Kubernetes 1.28+ cluster with GPU-capable nodes
  • NVIDIA GPUs (A10, L4, A100, H100, or similar)
  • NVIDIA drivers installed on nodes (535+ recommended)
  • Helm 3 for operator and plugin installation
  • Prometheus stack for metrics collection

NVIDIA GPU Operator Installation

The GPU Operator automates driver, toolkit, device plugin, and DCGM deployment.

# Add NVIDIA Helm repo
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm repo update

# Install GPU Operator
helm install gpu-operator nvidia/gpu-operator \
  --namespace gpu-operator \
  --create-namespace \
  --set driver.enabled=true \
  --set toolkit.enabled=true \
  --set devicePlugin.enabled=true \
  --set dcgmExporter.enabled=true \
  --set migManager.enabled=true \
  --set nodeStatusExporter.enabled=true \
  --version v24.3.0

# Verify installation
kubectl get pods -n gpu-operator
kubectl get nodes -o json | jq '.items[].status.allocatable["nvidia.com/gpu"]'

NVIDIA Device Plugin (Standalone)

If not using the GPU Operator, deploy the device plugin directly.

# nvidia-device-plugin.yaml
apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: nvidia-device-plugin
  namespace: kube-system
spec:
  selector:
    matchLabels:
      name: nvidia-device-plugin
  template:
    metadata:
      labels:
        name: nvidia-device-plugin
    spec:
      tolerations:
        - key: nvidia.com/gpu
          operator: Exists
          effect: NoSchedule
      priorityClassName: system-node-critical
      containers:
        - name: nvidia-device-plugin
          image: nvcr.io/nvidia/k8s-device-plugin:v0.15.0
          securityContext:
            privileged: true
          env:
            - name: FAIL_ON_INIT_ERROR
              value: "false"
            - name: DEVICE_SPLIT_COUNT
              value: "1"
            - name: DEVICE_LIST_STRATEGY
              value: "envvar"
          volumeMounts:
            - name: device-plugin
              mountPath: /var/lib/kubelet/device-plugins
      volumes:
        - name: device-plugin
          hostPath:
            path: /var/lib/kubelet/device-plugins

MIG (Multi-Instance GPU) Partitioning

MIG allows a single A100 or H100 to be split into isolated GPU instances.

# mig-config.yaml - ConfigMap for MIG Manager
apiVersion: v1
kind: ConfigMap
metadata:
  name: mig-parted-config
  namespace: gpu-operator
data:
  config.yaml: |
    version: v1
    mig-configs:
      # 7 small instances for inference microservices
      all-1g.10gb:
        - devices: all
          mig-enabled: true
          mig-devices:
            "1g.10gb": 7

      # 3 medium instances for mid-size models
      all-2g.20gb:
        - devices: all
          mig-enabled: true
          mig-devices:
            "2g.20gb": 3

      # Mixed: 1 large + 2 small
      mixed-inference:
        - devices: all
          mig-enabled: true
          mig-devices:
            "3g.40gb": 1
            "1g.10gb": 4

      # Full GPU for training (no partitioning)
      all-disabled:
        - devices: all
          mig-enabled: false
# Apply MIG profile to a node
kubectl label nodes gpu-node-01 nvidia.com/mig.config=all-1g.10gb --overwrite

# Verify MIG instances
kubectl exec -it nvidia-device-plugin-xxxxx -n kube-system -- nvidia-smi mig -lgi

# Check available MIG resources
kubectl get nodes gpu-node-01 -o json | jq '.status.allocatable | with_entries(select(.key | startswith("nvidia.com")))'

Requesting MIG Slices in Pods

# pod-with-mig.yaml
apiVersion: v1
kind: Pod
metadata:
  name: inference-small
spec:
  containers:
    - name: model
      image: registry.internal/vllm-server:latest
      resources:
        limits:
          nvidia.com/mig-1g.10gb: 1
      # For medium slice:
      # nvidia.com/mig-2g.20gb: 1
      # For large slice:
      # nvidia.com/mig-3g.40gb: 1

GPU Time-Slicing

For GPUs that do not support MIG (A10, L4), use time-slicing to share a GPU.

# time-slicing-config.yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: time-slicing-config
  namespace: gpu-operator
data:
  any: |-
    version: v1
    flags:
      migStrategy: none
    sharing:
      timeSlicing:
        renameByDefault: false
        failRequestsGreaterThanOne: false
        resources:
          - name: nvidia.com/gpu
            replicas: 4
# Apply time-slicing config
kubectl patch clusterpolicy/cluster-policy \
  --type merge \
  -p '{"spec":{"devicePlugin":{"config":{"name":"time-slicing-config","default":"any"}}}}'

# After applying, each physical GPU appears as 4 virtual GPUs
kubectl get nodes -o json | jq '.items[].status.allocatable["nvidia.com/gpu"]'
# Output: "4" per physical GPU

DCGM Monitoring

# dcgm-servicemonitor.yaml
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: dcgm-exporter
  namespace: gpu-operator
  labels:
    release: prometheus
spec:
  selector:
    matchLabels:
      app: nvidia-dcgm-exporter
  endpoints:
    - port: gpu-metrics
      interval: 15s
      path: /metrics

Key DCGM Metrics and Alert Rules

# gpu-alerts.yaml
groups:
  - name: gpu-health
    rules:
      - alert: GPUHighTemperature
        expr: DCGM_FI_DEV_GPU_TEMP > 85
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "GPU {{ $labels.gpu }} temperature above 85C on {{ $labels.node }}"

      - alert: GPUMemoryPressure
        expr: (DCGM_FI_DEV_FB_USED / DCGM_FI_DEV_FB_FREE) > 0.90
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "GPU memory above 90% on {{ $labels.node }} GPU {{ $labels.gpu }}"

      - alert: GPUECCErrors
        expr: increase(DCGM_FI_DEV_ECC_DBE_VOL_TOTAL[1h]) > 0
        labels:
          severity: critical
        annotations:
          summary: "Double-bit ECC errors detected on {{ $labels.node }} GPU {{ $labels.gpu }}"

      - alert: GPUXidErrors
        expr: increase(DCGM_FI_DEV_XID_ERRORS[5m]) > 0
        labels:
          severity: warning
        annotations:
          summary: "Xid error on {{ $labels.node }} GPU {{ $labels.gpu }}: {{ $labels.xid }}"

      - alert: GPULowUtilization
        expr: DCGM_FI_DEV_GPU_UTIL < 10 and on(pod) kube_pod_status_phase{phase="Running"} == 1
        for: 30m
        labels:
          severity: info
        annotations:
          summary: "GPU underutilized on {{ $labels.node }} - consider rightsizing"

      - alert: GPUDriverMismatch
        expr: count(count by (driver_version)(DCGM_FI_DRIVER_VERSION)) > 1
        labels:
          severity: warning
        annotations:
          summary: "Multiple GPU driver versions detected across cluster"

GPU Node Pool Configuration

# gpu-nodepool.yaml
apiVersion: v1
kind: Node
metadata:
  labels:
    gpu-type: a100
    gpu-memory: "80gb"
    gpu-mig-capable: "true"
    node-role: gpu-inference
spec:
  taints:
    - key: nvidia.com/gpu
      value: "true"
      effect: NoSchedule
---
# Inference deployment with GPU scheduling
apiVersion: apps/v1
kind: Deployment
metadata:
  name: llm-inference
  namespace: ai-serving
spec:
  replicas: 3
  selector:
    matchLabels:
      app: llm-inference
  template:
    metadata:
      labels:
        app: llm-inference
    spec:
      tolerations:
        - key: nvidia.com/gpu
          operator: Exists
          effect: NoSchedule
      nodeSelector:
        gpu-type: a100
      affinity:
        podAntiAffinity:
          preferredDuringSchedulingIgnoredDuringExecution:
            - weight: 100
              podAffinityTerm:
                labelSelector:
                  matchLabels:
                    app: llm-inference
                topologyKey: kubernetes.io/hostname
      containers:
        - name: vllm
          image: registry.internal/vllm-server:0.4.1
          resources:
            requests:
              nvidia.com/gpu: 1
              cpu: "4"
              memory: "32Gi"
            limits:
              nvidia.com/gpu: 1
              cpu: "8"
              memory: "64Gi"
          env:
            - name: CUDA_VISIBLE_DEVICES
              value: "all"

GPU Autoscaling

# gpu-hpa.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: llm-inference-hpa
  namespace: ai-serving
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: llm-inference
  minReplicas: 2
  maxReplicas: 8
  metrics:
    - type: Pods
      pods:
        metric:
          name: DCGM_FI_DEV_GPU_UTIL
        target:
          type: AverageValue
          averageValue: "75"
    - type: Pods
      pods:
        metric:
          name: inference_queue_depth
        target:
          type: AverageValue
          averageValue: "10"
  behavior:
    scaleUp:
      stabilizationWindowSeconds: 60
      policies:
        - type: Pods
          value: 2
          periodSeconds: 120
    scaleDown:
      stabilizationWindowSeconds: 300
      policies:
        - type: Pods
          value: 1
          periodSeconds: 300
---
# Cluster Autoscaler config for GPU node pools
apiVersion: v1
kind: ConfigMap
metadata:
  name: cluster-autoscaler-config
  namespace: kube-system
data:
  config: |
    expander: priority
    scale-down-delay-after-add: 10m
    scale-down-unneeded-time: 10m
    skip-nodes-with-local-storage: false
    balance-similar-node-groups: true
    expendable-pods-priority-cutoff: -10
    gpu-total:
      - min: 2
        max: 16
        gpu: nvidia.com/gpu

Scheduling Patterns

  • Use node affinity by GPU type (A10/L4/A100/H100).
  • Separate latency-critical inference from batch training.
  • Pin model replicas with anti-affinity for availability.
  • Reserve headroom for failover and rolling updates.

Cost Optimization

  • Prefer MIG slices for smaller inference services.
  • Schedule batch jobs in off-peak windows.
  • Route low-priority traffic to cheaper model tiers.
  • Use spot/preemptible instances for training workloads.
  • Monitor GPU utilization and rightsize deployments.

Troubleshooting

SymptomCheckFix
Pod stuck in Pendingkubectl describe pod for GPU resource eventsVerify node has allocatable GPUs, check taints/tolerations
CUDA OOM during inferenceModel too large for GPU memoryReduce batch size, use quantization, or use MIG slice
DCGM metrics missingServiceMonitor labels matchingVerify DCGM exporter pod is running and scrape config
Driver mismatch after upgradenvidia-smi on each nodeCordon node, drain, upgrade driver, uncordon
GPU not detectedDevice plugin pod logsRestart device plugin, check NVIDIA container toolkit
Time-slicing not workingConfigMap applied but no extra GPUsRestart device plugin pods after config change
ECC errors increasingnvidia-smi -q -d ECCSchedule node drain and hardware replacement

Related Skills

  • llm-inference-scaling - Autoscale inference workloads
  • model-serving-kubernetes - Production model serving patterns
  • gpu-server-management - Host-level GPU management fundamentals
  • multi-tenant-llm-hosting - Multi-tenant GPU sharing
  • llm-cost-optimization - Cost optimization strategies

Related skills

FAQ

What does the NVIDIA GPU Operator automate?

It automates driver, toolkit, device plugin, and DCGM deployment on GPU nodes.

How do you share a GPU across workloads?

Use MIG partitioning on A100/H100, or time-slicing on GPUs like A10 and L4.

Cloud & Infrastructureinframonitoringdeploy

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.