Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
pixel-process-ug avatar

Senior Devops

  • 64 installs
  • 1 repo stars
  • Updated March 16, 2026
  • pixel-process-ug/superkit-agents

Helps with devops & ci/cd tasks.

About

senior-devops is a Claude Code skill for devops & ci/cd. It helps solo builders move faster with AI-assisted development.

  • senior-devops
  • DevOps & CI/CD
  • AI-coding skill

Senior Devops by the numbers

  • 64 all-time installs (skills.sh)
  • +1 installs in the week ending Aug 4, 2026 (Skillselion tracking)
  • Ranked #641 of 1,435 DevOps & CI/CD skills by installs in the Skillselion catalog
  • Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/pixel-process-ug/superkit-agents --skill senior-devops

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs64
repo stars1
Last updatedMarch 16, 2026
Repositorypixel-process-ug/superkit-agents

What it does

Helps with devops & ci/cd tasks.

Files

SKILL.mdMarkdownGitHub ↗

Senior DevOps Engineer

Overview

Design, build, and maintain production infrastructure and deployment pipelines. This skill covers Docker containerization, Kubernetes orchestration, CI/CD with GitHub Actions, infrastructure-as-code with Terraform/Pulumi, monitoring with Prometheus/Grafana, alerting strategies, zero-downtime deployments, and rollback procedures.

Phase 1: Infrastructure Design

1. Define deployment topology (single server, cluster, multi-region) 2. Choose containerization strategy (Docker, Buildpacks) 3. Select orchestration platform (Kubernetes, ECS, Cloud Run) 4. Plan networking (load balancers, DNS, TLS) 5. Design secret management approach

STOP — Present infrastructure design to user for approval before implementation.

Infrastructure Decision Table

ScaleTopologyOrchestrationRecommended
Hobby / MVPSingle serverDocker ComposeRailway, Fly.io
Startup (< 100k users)Small clusterECS, Cloud RunAWS ECS, GCP Cloud Run
Growth (100k - 1M users)Multi-AZ clusterKubernetesEKS, GKE
Enterprise (1M+ users)Multi-regionKubernetes + service meshEKS/GKE + Istio
Compliance-heavyDedicated/private cloudKubernetesSelf-managed K8s

Phase 2: Pipeline Implementation

1. Build CI pipeline (lint, test, build, security scan) 2. Build CD pipeline (deploy to staging, production) 3. Configure environment-specific settings 4. Set up artifact registry (container images, packages) 5. Implement deployment strategy (blue-green, canary, rolling)

STOP — Validate pipeline config syntax and present for review.

Phase 3: Observability

1. Deploy monitoring stack (Prometheus, Grafana) 2. Configure alerting rules and escalation 3. Set up log aggregation 4. Implement distributed tracing 5. Create runbooks for common incidents

STOP — Verify monitoring covers all critical services before declaring complete.

Dockerfile Best Practices

# 1. Use specific version tags (not :latest)
FROM node:20-alpine AS base

# 2. Set working directory
WORKDIR /app

# 3. Install dependencies in separate layer (cache optimization)
FROM base AS deps
COPY package.json pnpm-lock.yaml ./
RUN corepack enable && pnpm install --frozen-lockfile --prod

FROM base AS build-deps
COPY package.json pnpm-lock.yaml ./
RUN corepack enable && pnpm install --frozen-lockfile

# 4. Build in separate stage
FROM build-deps AS builder
COPY . .
RUN pnpm build

# 5. Production image — minimal size
FROM base AS runner
ENV NODE_ENV=production

# 6. Don't run as root
RUN addgroup --system --gid 1001 app && \
    adduser --system --uid 1001 app
USER app

# 7. Copy only what's needed
COPY --from=deps /app/node_modules ./node_modules
COPY --from=builder /app/dist ./dist

# 8. Health check
HEALTHCHECK --interval=30s --timeout=3s --start-period=5s \
  CMD wget -qO- http://localhost:3000/health || exit 1

# 9. Expose port and set entrypoint
EXPOSE 3000
CMD ["node", "dist/server.js"]

Key Dockerfile Rules

RuleWhy
Multi-stage buildsMinimize image size
.dockerignore fileExclude node_modules, .git, tests
Non-root userSecurity hardening
Specific base image versionsReproducible builds
Layer ordering (deps before src)Cache efficiency
HEALTHCHECK instructionContainer health monitoring
No secrets in build args/layersPrevent credential leaks

Docker Compose Patterns

services:
  app:
    build:
      context: .
      dockerfile: Dockerfile
      target: runner
    ports:
      - "3000:3000"
    environment:
      - DATABASE_URL=postgresql://postgres:postgres@db:5432/app
      - REDIS_URL=redis://cache:6379
    depends_on:
      db:
        condition: service_healthy
      cache:
        condition: service_started
    healthcheck:
      test: ["CMD", "wget", "-qO-", "http://localhost:3000/health"]
      interval: 10s
      timeout: 5s
      retries: 3

  db:
    image: postgres:16-alpine
    volumes:
      - postgres_data:/var/lib/postgresql/data
    environment:
      POSTGRES_DB: app
      POSTGRES_USER: postgres
      POSTGRES_PASSWORD: postgres
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U postgres"]
      interval: 5s
      timeout: 3s
      retries: 5

  cache:
    image: redis:7-alpine
    volumes:
      - redis_data:/data

volumes:
  postgres_data:
  redis_data:

GitHub Actions Workflow

name: CI/CD
on:
  push:
    branches: [main]
  pull_request:
    branches: [main]

concurrency:
  group: ${{ github.workflow }}-${{ github.ref }}
  cancel-in-progress: true

jobs:
  lint-and-test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: pnpm/action-setup@v3
      - uses: actions/setup-node@v4
        with:
          node-version: 20
          cache: pnpm
      - run: pnpm install --frozen-lockfile
      - run: pnpm lint
      - run: pnpm typecheck
      - run: pnpm test -- --coverage

  security-scan:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: npx audit-ci --moderate
      - uses: aquasecurity/trivy-action@master
        with:
          scan-type: fs
          severity: HIGH,CRITICAL

  build-and-push:
    needs: [lint-and-test, security-scan]
    if: github.ref == 'refs/heads/main'
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: docker/setup-buildx-action@v3
      - uses: docker/login-action@v3
        with:
          registry: ghcr.io
          username: ${{ github.actor }}
          password: ${{ secrets.GITHUB_TOKEN }}
      - uses: docker/build-push-action@v5
        with:
          push: true
          tags: ghcr.io/${{ github.repository }}:${{ github.sha }}
          cache-from: type=gha
          cache-to: type=gha,mode=max

  deploy:
    needs: build-and-push
    runs-on: ubuntu-latest
    environment: production
    steps:
      - name: Deploy to production
        run: echo "Deploying ${{ github.sha }}"

Terraform / Pulumi Patterns

Terraform Structure

modules/
  vpc/
    main.tf, variables.tf, outputs.tf
  ecs/
    main.tf, variables.tf, outputs.tf
environments/
  staging/
    main.tf, terraform.tfvars
  production/
    main.tf, terraform.tfvars

Key IaC Rules

RuleWhy
Remote state backend (S3 + DynamoDB)Shared state, locking
State lockingPrevent concurrent modifications
Environment-specific variable filesSeparation of concerns
Module versioningReproducible shared infra
terraform plan in CICatch issues before apply
Drift detection on scheduleDetect manual changes
Tag all resourcesOwnership, cost allocation

Monitoring (Prometheus + Grafana)

USE Method (Resources)

ResourceUtilizationSaturationErrors
CPUcpu_usage_percentcpu_throttled
Memorymemory_usage_bytesoom_kills
Diskdisk_usage_percentio_waitdisk_errors
Networkbytes_totalqueue_lengtherrors_total

RED Method (Services)

  • Rate: requests per second
  • Errors: error rate per second
  • Duration: latency distribution (p50, p95, p99)

Alerting Rules

groups:
  - name: app-alerts
    rules:
      - alert: HighErrorRate
        expr: rate(http_requests_total{status=~"5.."}[5m]) / rate(http_requests_total[5m]) > 0.05
        for: 5m
        labels:
          severity: critical
      - alert: HighLatency
        expr: histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m])) > 1
        for: 5m
        labels:
          severity: warning

Alerting Best Practices

PracticeWhy
Alert on symptoms, not causesReduces noise, focuses on impact
Every alert has a runbook linkEnables fast response
Tiered severitycritical=page, warning=ticket, info=log
Aggregate before alertingAvoid flapping
Review and prune quarterlyPrevent alert fatigue

Zero-Downtime Deployment Strategies

StrategyHow It WorksRiskRollback Speed
RollingReplace instances one at a timeLowMedium
Blue-GreenSwitch traffic between two environmentsLowInstant
CanaryRoute small % to new version, gradually increaseVery LowInstant
Feature FlagsDeploy code dark, enable via flagVery LowInstant

Rollback Procedures

1. Automated: health check fails -> automatic rollback 2. Manual: kubectl rollout undo deployment/app 3. Database: forward-only migrations with backward compatibility 4. Config: revert via secret manager version

Database Migration Safety

RuleRationale
Migrations must be backward compatibleOld code + new schema must work
Never rename/drop columns in same deployTwo-phase change required
Two-phase: add column -> deploy -> remove oldZero-downtime schema evolution
Always test rollback of each migrationEnsure reversibility

Anti-Patterns / Common Mistakes

Anti-PatternWhy It Is WrongWhat to Do Instead
Manual production deploymentsNo audit trail, error-proneAutomate via CI/CD
Shared or hardcoded secretsSecurity breach riskUse secrets manager
No rollback plan before deployingStuck if deploy failsDocument rollback before every deploy
latest tag for production imagesNon-reproduciblePin specific version tags
Running containers as rootSecurity vulnerabilityUse non-root user in Dockerfile
Alert fatigue from non-actionable alertsReal issues get missedAlert on symptoms, tune thresholds
Skipping staging environmentBugs found in productionAlways deploy to staging first
Snowflake servers with manual configCannot reproduce, cannot scaleInfrastructure as code
Monitoring without alertingNobody notices problemsWire alerts to monitoring

Key Principles

  • Infrastructure as code — no manual changes to production
  • Immutable infrastructure — replace, do not patch
  • Cattle, not pets — servers are disposable
  • Shift left security — scan early in pipeline
  • Least privilege — minimal permissions everywhere
  • Automate everything that runs more than twice
  • Test the disaster recovery plan regularly

Documentation Lookup (Context7)

Use mcp__context7__resolve-library-id then mcp__context7__query-docs for up-to-date docs. Returned docs override memorized knowledge.

  • docker — for Dockerfile syntax, compose configuration, or multi-stage builds
  • kubernetes — for resource manifests, kubectl commands, or Helm charts
  • terraform — for provider configuration, resource blocks, or state management

---

Integration Points

SkillIntegration
deploymentProvides higher-level deploy pipeline orchestration
security-reviewSecurity scan stage in CI pipeline
planningInfrastructure changes are planned like features
verification-before-completionPost-deploy verification gate
finishing-a-development-branchMerge triggers deployment pipeline
mcp-builderMCP servers need containerization and deployment

Skill Type

FLEXIBLE — Adapt tooling and patterns to the project's cloud provider, team size, and operational maturity. The principles (IaC, immutability, observability) are constant; the specific tools are interchangeable.

Related skills

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.