Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
daemon-blockint-tech avatar

Incident Management Engineer

  • 26 installs
  • 7 repo stars
  • Updated May 20, 2026
  • daemon-blockint-tech/agentic-enteprises-skill

Guides incident management program design: severity models, escalation policies, on-call rotations, paging/comms tooling, postmortems, and reliability metrics.

About

Guides incident management engineering across severity and escalation models, on-call design, paging/comms tooling, incident lifecycle, and blameless postmortems. A team uses it when building incident response programs, on-call rotations, or SEV definitions and metrics.

  • Severity aligned to customer impact rather than alert noise
  • Alerting-to-paging-to-timeline integration with MTTR/MTTD metrics

Incident Management Engineer by the numbers

  • 26 all-time installs (skills.sh)
  • Ranked #885 of 1,435 DevOps & CI/CD skills by installs in the Skillselion catalog
  • Data as of Jul 29, 2026 (Skillselion catalog sync)
npx skills add https://github.com/daemon-blockint-tech/agentic-enteprises-skill --skill incident-management-engineer

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs26
repo stars7
Last updatedMay 20, 2026
Repositorydaemon-blockint-tech/agentic-enteprises-skill

What it does

Guides incident management program design: severity models, escalation policies, on-call rotations, paging/comms tooling, postmortems, and reliability metrics.

Files

SKILL.mdMarkdownGitHub ↗

Incident Management Engineer

When to Use

  • Define or revise severity levels and escalation policies
  • Design on-call rotations, schedules, and handoffs
  • Integrate alerting → paging → incident channel → ticket timeline
  • Run blameless postmortem process and action-item tracking
  • Report incident metrics and improve MTTR/MTTD
  • Configure status page and customer comms workflows for outages

When NOT to Use

  • Fix pipelines, deploys, or service code during outage → devops, fullstack-software-engineer
  • Investigate malware, phishing, or SOC alerts → soc-analyst (deep hunts → defensive-security-analyst)
  • Enterprise security IR and legal/compliance program → cybersecurity
  • Cross-team launch programs and RAID → technical-program-manager
  • Data platform-specific ops → data-system-ops-lead
  • Write customer-facing runbooks only → tech-writer-researcher
  • Single-account repro and support escalations → support-engineer

Related skills

NeedSkill
SLOs, error budgets, reliability metricssite-reliability-engineer
Pipelines, alerts stack implementationdevops
Rollback and cutover during outagedeployment-strategist
Security incident playbookscybersecurity
SOC alert triage and playbookssoc-analyst
Active CSIRT response, timelines, evidenceincident-responder
BCP/DRP, cyber recovery playbooks, restore tests, tabletopsbcm-disaster-recovery-specialist
Deep investigation, hunts, detectionsdefensive-security-analyst
Major cross-team incident coordinationtechnical-program-manager
Runbook documentationtech-writer-researcher
Customer ticket repro and engineering escalationsupport-engineer
Incident and crisis message packscommunication-lead
Exec/community customer escalation programcommunity-executive-escalations-program-manager

Core Workflows

1. Severity and escalation

1. Align severity to customer impact, not alert noise 2. Map each level: response time, who pages, comms required 3. Document escalation ladder (primary → secondary → manager → exec) 4. Review quarterly with recent incident data

See `references/severity_escalation.md` for matrix template.

2. On-call program

  • Primary + secondary coverage; no single point of failure
  • Rotation length: prefer weekly over daily for sustainability
  • Fairness: track pages per person; cap repeat pages
  • Handoff ritual with open incidents and deploy context

See `references/on_call_design.md` for rotation and handoff patterns.

3. Incident lifecycle tooling

Standard flow:

Alert → page → incident declared → comms channel → roles assigned → mitigate → resolve → postmortem
  • Auto-create incident record with timeline (who/when)
  • Integrate chat, tickets, and paging in one timeline
  • Reserve manual steps for role assignment and customer comms approval

See `references/incident_tooling.md` for integration checklist.

4. Active incident (commander-lite)

During SEV1–2:

RoleResponsibility
Incident commanderCoordinates; does not debug alone
CommunicationsInternal + external updates on cadence
Technical lead(s)Mitigation per service
  • Time-box updates (e.g., every 30 min until stable)
  • Log decisions in incident timeline
  • Defer root-cause deep dive until mitigated

See `references/incident_lifecycle.md` for phases.

5. Postmortem program

  • Blameless; focus on systems and process
  • Within 48h for SEV1–2; required before closing incident
  • Action items: owner, due date, tracked to completion
  • Share learnings broadly; link detection gaps to monitoring (devops)

See `references/postmortem_process.md` for template and metrics.

6. Metrics and improvement

Track monthly:

  • Incident count by severity
  • MTTD, MTTR (mitigation and full resolution)
  • Repeat incidents (same root cause class)
  • Postmortem action item closure rate
  • On-call load (pages per engineer)

When to load references

  • SEV matrix and escalationreferences/severity_escalation.md
  • Rotations and handoffsreferences/on_call_design.md
  • Lifecycle phasesreferences/incident_lifecycle.md
  • PagerDuty/Slack/ticket wiringreferences/incident_tooling.md
  • Postmortems and metricsreferences/postmortem_process.md

Related skills

DevOps & CI/CDmonitoring

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.