Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
daemon-blockint-tech avatar

Ai Lead Ops

  • 28 installs
  • 7 repo stars
  • Updated May 20, 2026
  • daemon-blockint-tech/agentic-enteprises-skill

Run AI production operations: SLOs, release governance, incident reviews, cost/capacity tracking, and vendor eval for LLM-powered features.

About

Guides AI operations leadership covering LLM reliability, model/prompt release governance, incidents, cost/capacity, and vendor management. A developer uses it when standing up AI platform ops, defining LLM SLAs, or governing model rollouts.

  • Release governance checklist with eval, red-team, canary, and rollback gates
  • SLOs, AI incident types, and unit-economics cost tracking

Ai Lead Ops by the numbers

  • 28 all-time installs (skills.sh)
  • Ranked #9,462 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
  • Data as of Jul 29, 2026 (Skillselion catalog sync)
npx skills add https://github.com/daemon-blockint-tech/agentic-enteprises-skill --skill ai-lead-ops

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs28
repo stars7
Last updatedMay 20, 2026
Repositorydaemon-blockint-tech/agentic-enteprises-skill

What it does

Run AI production operations: SLOs, release governance, incident reviews, cost/capacity tracking, and vendor eval for LLM-powered features.

Files

SKILL.mdMarkdownGitHub ↗

AI Lead Ops

When to Use

  • Standing up AI platform operations and production service reliability
  • Defining SLAs/SLOs for LLM-powered features
  • Running AI incident reviews and post-mortems
  • Governing model, prompt, and index rollouts with tiered gates
  • Tracking AI unit economics (cost per session, tokens per feature)
  • Coordinating red-team and evaluation gates before releases
  • Building team rituals and cadence across engineering, research, risk, and product
  • Managing AI vendor relationships, contracts, and bake-offs

When NOT to Use

  • Implementing memory stores or context packing code → ai-memory-developer / ai-context-engineer
  • Building RAG pipelines or agent tools → ai-engineer
  • Designing corporate AI policy or regulatory mapping → ai-risk-governance
  • General network penetration testing or enterprise security programs → cybersecurity
  • Structured token/cost improvement roadmaps with backlog → ai-token-improvement-plan-engineer
  • Commercial/enterprise AI solution architecture → applied-ai-architect-commercial-enterprise
  • Vertical AI product engineering managers and squad roadmaps → engineering-manager-vertical-ai-products

Related skills

NeedSkill
Build RAG, agents, eval harnessesai-engineer
Memory and context implementationai-memory-developer, ai-context-engineer
Risk tiering and policiesai-risk-governance
Adversarial testing executionai-redteam
CI/CD and platform incidentsdevops
Pipeline securitydevsecops
Token optimization roadmap and initiative backlogai-token-improvement-plan-engineer
Commercial/enterprise AI architectureapplied-ai-architect-commercial-enterprise
Skills portfolio governanceai-skill-manager
Safeguard inference platformml-infrastructure-engineer-safeguards
Safety classifier researchml-research-engineer-safeguards

Core Workflows

1. Operating model and cadence

RitualFrequencyOutcomes
AI ops standupDailyBlockers, incidents, deploys
Model/prompt change reviewPer releaseApprovers, eval delta
Cost reviewWeeklySpend vs budget, top features
Risk & safety syncBi-weeklyIncidents, policy gaps
Quarterly capacityQuarterlyModel roadmap, vendor contracts

Define RACI: who owns model, prompt, index, eval suite, on-call.

See `references/operating_model.md` for roles and escalation.

2. Release governance

Production promotion checklist:

  • [ ] Eval regression passed on golden + safety set
  • [ ] Red-team sign-off for tier-2+ use cases
  • [ ] Model card / change log updated
  • [ ] Canary with error and cost monitors
  • [ ] Rollback procedure tested (previous prompt + model version pinned)
  • [ ] Comms plan for customer-visible behavior change

See `references/release_governance.md` for tiered gates and canary metrics.

3. SLOs, incidents, and observability

Example SLIs:

SLINotes
AvailabilitySuccessful completion / total requests
Latencyp95 end-to-end
Quality proxyThumbs-down rate, escalation rate
SafetyPolicy violation rate post-deploy
CostUSD per successful session

AI incident types: toxic output, PII leak in logs, retrieval cross-tenant leak, runaway agent loop, vendor outage.

See `references/incidents_slos.md` for severity matrix and post-incident template.

4. Cost and capacity

  • Track tokens by model, feature, tenant
  • Set budgets and alerts at 80/100/110%
  • Optimize via routing, caching, context engineering (partner with ai-context-engineer)
  • Forecast from usage growth + model price changes

See `references/cost_capacity.md` for unit economics worksheet.

5. Vendor and eval program

  • Maintain scorecard: quality, latency, safety, price, data terms
  • Run structured bake-offs before annual renewals
  • Own central eval harness ownership and dataset hygiene

See `references/vendor_eval_program.md` for RFP topics and eval program maturity.

When to load references

  • Team cadence and RACIreferences/operating_model.md
  • Releases and canariesreferences/release_governance.md
  • SLOs and incidentsreferences/incidents_slos.md
  • Cost and capacityreferences/cost_capacity.md
  • Vendors and eval opsreferences/vendor_eval_program.md

Related skills

AI & Agent Buildingmonitoringdeploy

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.