Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
affaan-m avatar

Enterprise Agent Ops

  • 1.4k installs
  • 238k repo stars
  • Updated August 5, 2026
  • affaan-m/ecc

This is a copy of enterprise-agent-ops by affaan-m - installs and ranking accrue to the original listing.

enterprise-agent-ops is an agent skill that adds runtime lifecycle management, observability, safety boundaries, and controlled rollouts to long-running AI agents in cloud-hosted production environments.

About

enterprise-agent-ops is an ECC agent skill for operating cloud-hosted or continuously running agent systems beyond single CLI sessions. It structures four operational domains: runtime lifecycle with start, pause, stop, and restart controls; observability through logs, metrics, and traces; safety controls including scopes, permissions, and kill switches; and change management with rollout, rollback, and audit trails. Baseline controls include immutable deployment artifacts, least-privilege credentials, environment-level secret injection, hard timeout and retry budgets, and audit logs for high-risk actions. Developers reach for enterprise-agent-ops when production agents need operational guardrails, incident response hooks, and governed deployment practices rather than ephemeral local agent runs.

  • Covers four operational domains: runtime lifecycle, observability, safety controls, change management
  • Enforces immutable artifacts, least-privilege credentials, hard timeouts, retry budgets and audit logs
  • Tracks five key metrics including success rate, mean retries, time to recovery, cost per task and failure class distribu
  • Provides a six-step incident response pattern that ends with regression, security checks and gradual resumption
  • Integrates with PM2, systemd, container orchestrators and CI/CD gates

Enterprise Agent Ops by the numbers

  • 1,370 all-time installs (skills.sh)
  • +84 installs in the week ending Aug 5, 2026 (Skillselion tracking)
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/affaan-m/ecc --skill enterprise-agent-ops

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs1.4k
repo stars238k
Last updatedAugust 5, 2026
Repositoryaffaan-m/ecc

How do you operate long-running AI agents in production?

Add runtime lifecycle, observability, safety boundaries and controlled rollouts to long-running AI agents.

Who is it for?

Platform engineers running cloud-hosted or continuously running AI agents who need lifecycle management, observability, safety boundaries, and governed rollouts.

Skip if: Developers running one-off local CLI agent sessions without production uptime, monitoring, or rollout governance requirements.

When should I use this skill?

The user mentions agent observability, kill switches, agent rollout, runtime lifecycle, production agent monitoring, or long-running agent operations.

What you get

Lifecycle controls, observability dashboards, safety boundaries, rollout procedures, and audit logs for agent workloads.

  • observability configuration
  • safety boundary policies
  • rollout and rollback runbooks

By the numbers

  • Structures agent operations across 4 domains: runtime lifecycle, observability, safety controls, change management

Files

SKILL.mdMarkdownGitHub ↗

Enterprise Agent Ops

Use this skill for cloud-hosted or continuously running agent systems that need operational controls beyond single CLI sessions.

Operational Domains

1. runtime lifecycle (start, pause, stop, restart) 2. observability (logs, metrics, traces) 3. safety controls (scopes, permissions, kill switches) 4. change management (rollout, rollback, audit)

Baseline Controls

  • immutable deployment artifacts
  • least-privilege credentials
  • environment-level secret injection
  • hard timeout and retry budgets
  • audit log for high-risk actions

Metrics to Track

  • success rate
  • mean retries per task
  • time to recovery
  • cost per successful task
  • failure class distribution

Incident Pattern

When failure spikes: 1. freeze new rollout 2. capture representative traces 3. isolate failing route 4. patch with smallest safe change 5. run regression + security checks 6. resume gradually

Deployment Integrations

This skill pairs with:

  • PM2 workflows
  • systemd services
  • container orchestrators
  • CI/CD gates

Related skills

How it compares

Choose enterprise-agent-ops over agent-building skills when the task is production operations and governance rather than authoring new agent prompts or tools.

FAQ

What operational domains does enterprise-agent-ops cover?

enterprise-agent-ops covers runtime lifecycle, observability with logs metrics and traces, safety controls with scopes and kill switches, and change management including rollout, rollback, and audit logging for production agents.

What baseline controls does enterprise-agent-ops require?

enterprise-agent-ops requires immutable deployment artifacts, least-privilege credentials, environment-level secret injection, hard timeout and retry budgets, and audit logs for high-risk agent actions in cloud-hosted workloads.

AI & Agent Buildingagentsautomation

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.