
Dora Metrics
- 46 installs
- 325 repo stars
- Updated August 2, 2026
- athola/claude-night-market
dora-metrics is an agent skill that maps DORA signals to AI-assisted workflow failure modes and remediation levers.
About
dora-metrics is an agent skill from Claude Night Market that reframes classic DORA measurements for AI-assisted delivery. Solo builders and small teams already track deployments and incidents; this module explains what to watch when agents join the pipeline—splitting change failure rate by label, comparing lead time across adoption windows, sanity-checking time to restore after agent hotfixes, and capping deployment-frequency enthusiasm with failure rate. It references concrete CLI usage via python3 -m minister.dora_metrics with JSON output, and recommends friction responses such as hookify rules or imbue gates rather than turning off AI help. Use it when you are growing a product with heavy agent throughput and need honest signals on whether review and restore practices kept pace with speed.
- Run minister.dora_metrics with --window 30 and distinct --failure-label values (e.g. bug vs ai-bug) and compare JSON exp
- Treat AI CFR more than five percentage points above human CFR as lenient review signal, not a ban on AI assistance
- Compare lead time for 30 days before versus after agent adoption to spot velocity-for-stability tradeoffs
- Watch time-to-restore when agents ship hotfixes—incomplete RCA can inflate TRS once truth surfaces
- Pair rising deployment frequency with CFR so arbitrarily high DF from agents does not look healthy alone
Dora Metrics by the numbers
- 46 all-time installs (skills.sh)
- Ranked #760 of 1,435 DevOps & CI/CD skills by installs in the Skillselion catalog
- Security screen: MEDIUM risk (skills.sh audit)
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/athola/claude-night-market --skill dora-metricsAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 46 |
|---|---|
| repo stars | ★ 325 |
| Security audit | 1 / 2 scanners passed |
| Last updated | August 2, 2026 |
| Repository | athola/claude-night-market ↗ |
What it does
Compare DORA-style change failure rate and lead time for AI-authored versus human work so agent velocity does not hide quality regressions.
Who is it for?
Best when you're shipping with agents and already log deploys and failure labels and want minister-style JSON metrics over rolling windows.
Skip if: Skip if you have no deployment or incident labeling history and only need a single pre-launch checklist.
When should I use this skill?
Measuring delivery health after enabling or scaling agentic coding workflows and you have labeled failures or deploy history to query.
What you get
You run labeled DORA windows, interpret AI versus human CFR and lead-time tradeoffs, and add gates at friction points instead of guessing.
- JSON metric snapshots for all versus AI-labeled failure windows
- Interpretation notes on CFR, lead time, TRS, and DF tradeoffs with suggested friction gates
By the numbers
- Example commands use a 30-day --window with JSON output via python3 -m minister.dora_metrics
- AI versus human CFR gap of more than five percentage points is called out as a review-leniency signal
Files
DORA Metrics
Purpose
Compute the four DORA delivery-performance metrics (Deployment Frequency, Lead Time for Changes, Change Failure Rate, and Time to Restore Service) from local git history and the GitHub API. Classify each metric into Elite, High, Medium, or Low using thresholds from DORA's State of DevOps research, and surface the single weakest dimension as the next improvement target.
When to Use
- Engineering management retrospectives and quarterly reviews.
- Auditing whether agentic workflows (AI-assisted PRs, automated
deploys) improve velocity and stability or quietly regress them.
- Feeding a tier signal into
minister:release-health-gates.
When Not to Use
- Single-team velocity tracking that needs story-point burndowns
rather than delivery-performance evidence.
- Repositories without a clear production branch or release cadence;
DORA assumes one.
Workflow
1. Run the helper script with the desired window:
python3 -m minister.dora_metrics --window 30 --branch main2. Read the output: per-metric value, tier classification, and the bottleneck pointer.
3. For agentic-workflow audits, run the same window twice. Once filtering to AI-authored PRs (e.g., --failure-label ai-bug), once across all PRs. Compare the CFR delta. See modules/agentic-workflow-signals.md.
4. Optionally pipe --json into the tracker so trend data persists alongside release-health-gates snapshots.
5. Optionally render trend charts with kuva when reviewing multiple windows or comparing before/after an agentic-workflow change:
# Collect weekly snapshots into a TSV, then plot all four metrics
# week<TAB>metric<TAB>value
kuva line trends.tsv --x week --y value --color-by metric \
--title "DORA trends (30-day windows)" -o dora-trends.svg
# Quick terminal preview without writing a file
kuva line trends.tsv --x week --y value --color-by metric --terminalkuva reads TSV/CSV from stdin or a file path. Install once: cargo install kuva --features cli. No project source changes required. See kuva for the full plot-type reference.
Inputs
| Flag | Default | Meaning |
|---|---|---|
--window | 30 | Measurement window in days |
--branch | HEAD | Production branch |
--failure-label | bug | GitHub label marking prod failures |
--json | off | Emit JSON instead of human-readable |
--repo-path | cwd | Repository directory |
Outputs
A short text report or JSON payload with:
- Per-metric numeric value (e.g.,
4.2/day,2.1 hours,8%). - Per-metric tier (Elite, High, Medium, Low).
- Overall tier (the weakest of the four).
- Bottleneck key, identifying which metric to focus improvement on.
Tier Thresholds
See modules/thresholds.md for the complete table. Brief summary:
| Metric | Elite | High | Medium | Low |
|---|---|---|---|---|
| DF | >= 1/day | >= 1/week | >= 1/month | < 1/month |
| LT | <= 1 day | <= 1 week | <= 1 month | > 1 month |
| CFR | <= 15% | <= 30% | <= 45% | > 45% |
| TRS | < 1 hour | < 1 day | < 1 week | >= 1 week |
Verification
Confirm a DORA report is real by re-running the script over a narrower window and checking that DF and LT scale predictably. For CFR and TRS, sample two or three of the contributing GitHub issues and verify the bug (or chosen) label is correct on each.
Testing
Unit tests live in plugins/minister/tests/unit/test_dora_metrics.py. Each tier boundary is exercised at the threshold, so future contributors who adjust an inequality (> vs >=) trigger a failure rather than a silent regression. Add new tests at the threshold when extending classification logic.
Exit Criteria
- [ ] DORA report generated for the requested window.
- [ ] All four metrics classified into a tier.
- [ ] Bottleneck dimension surfaced.
- [ ] Output is readable in a terminal or as a PR comment.
Agentic Workflow Signals from DORA
DORA metrics were designed for human-driven engineering teams, but the same four numbers expose specific failure modes in AI-assisted pipelines.
What to Watch
Change Failure Rate, AI vs human
Run the metric twice with different --failure-label values:
python3 -m minister.dora_metrics --window 30 --failure-label bug --json > all.json
python3 -m minister.dora_metrics --window 30 --failure-label ai-bug --json > ai.jsonIf AI-authored CFR exceeds human-authored CFR by more than five percentage points, treat it as a signal that review is too lenient on AI output, not that AI is unsafe in general. The right response is usually adding a hookify rule or imbue gate at the friction point, not banning AI assistance.
Lead Time, before vs after agent adoption
Compute lead time for the 30 days before and after enabling an agentic workflow. If LT improved but CFR or TRS regressed, the team is trading stability for velocity. The bottleneck dimension surfaced by the skill points at which trade was made.
Time to Restore, agent-driven hotfixes
If TRS got worse after agents started shipping hotfixes, suspect incomplete root-cause analysis. The Replit incident is a case study: fast restore claims that turn out to be fabricated extend TRS once the truth surfaces.
Deployment Frequency, ceiling check
Agents can push DF arbitrarily high. Pair DF with CFR; if DF rose and CFR rose proportionally, the agent is generating noise rather than signal. A high-DF, high-CFR team produces churn.
Producing a Comparison Report
Combine two windows side-by-side:
from minister.dora_metrics import compute_metrics
# ... collect events for each cohort ...
human = compute_metrics(human_deploys, human_failures, window_days=30)
agent = compute_metrics(agent_deploys, agent_failures, window_days=30)
print("Human:", human.tier())
print("Agent:", agent.tier())
print("Human bottleneck:", human.bottleneck())
print("Agent bottleneck:", agent.bottleneck())If the bottleneck differs across cohorts, that is the most useful single output: it tells the engineering manager which guardrail is missing for which population.
Anti-Patterns
- Reporting only DF as proof of agent ROI without CFR.
- Excluding agent-authored failures from the failure label.
- Comparing against last quarter when agent adoption mid-window
invalidates the comparison.
DORA Tier Thresholds
Source: DORA's State of DevOps research. The thresholds below match the published bands; minor adjustments per release year are common but the band shape is stable.
Deployment Frequency (DF)
How often code is deployed to production. Higher is better.
| Tier | Threshold |
|---|---|
| Elite | At least once per day |
| High | Between once per week and once per day |
| Medium | Between once per month and once per week |
| Low | Less often than once per month |
Lead Time for Changes (LT)
Median time from commit to production. Lower is better.
| Tier | Threshold |
|---|---|
| Elite | Less than one day |
| High | One day to one week |
| Medium | One week to one month |
| Low | More than one month |
Change Failure Rate (CFR)
Percentage of deployments that cause a production failure. Lower is better.
| Tier | Threshold |
|---|---|
| Elite | At most 15% |
| High | 16-30% |
| Medium | 31-45% |
| Low | More than 45% |
Time to Restore Service (TRS)
Median time to recover from a production failure. Lower is better.
| Tier | Threshold |
|---|---|
| Elite | Less than one hour |
| High | Less than one day |
| Medium | Less than one week |
| Low | One week or more |
Boundary Behavior
The implementation places the boundary value in the better tier:
- DF exactly 1.0/day classifies as Elite, not High.
- LT exactly 24 hours classifies as Elite, not High.
- CFR exactly 15% classifies as Elite, not High.
- TRS exactly 1 hour classifies as High, not Elite (TRS uses strict
< for Elite to keep the "less than one hour" wording honest).
Boundary tests in plugins/minister/tests/unit/test_dora_metrics.py pin these choices.
Related skills
How it compares
Interpretation layer for DORA-style metrics in agent pipelines—not a dashboard product or generic analytics MCP by itself.
FAQ
Who is dora-metrics for?
Developers and small teams measuring delivery health while Claude Code or similar agents author a growing share of changes and hotfixes.
When should I use dora-metrics?
Use it in Grow analytics when comparing AI-labeled bugs to human CFR, in Operate monitoring after restore incidents, or in Ship launch prep when deployment frequency spikes with agents.
Is dora-metrics safe to install?
The skill describes running local minister metrics commands; review the Security Audits panel on this Prism page before installing skills from the Night Market repo.