
Zero Tolerance For Failure
- 27 installs
- 7 repo stars
- Updated May 20, 2026
- daemon-blockint-tech/agentic-enteprises-skill
Establish failure-prevention culture for mission-critical systems: HRO principles, verification gates, fail-safe design, pre-mortems, FMEA, and stop-the-line policy.
About
Guides failure-prevention culture and operational excellence for mission-critical engineering, covering HRO principles, defense-in-depth, verification gates, redundancy, pre-mortems/FMEA, and prevention metrics. Used when building failure-prevention programs and countering normalization of deviance.
- Design defense-in-depth, fail-safe, and fail-closed controls
- Facilitate pre-mortems and FMEA; track defect-escape and near-miss metrics
Zero Tolerance For Failure by the numbers
- 27 all-time installs (skills.sh)
- Ranked #686 of 1,352 Code Review & Quality skills by installs in the Skillselion catalog
- Data as of Jul 29, 2026 (Skillselion catalog sync)
npx skills add https://github.com/daemon-blockint-tech/agentic-enteprises-skill --skill zero-tolerance-for-failureAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 27 |
|---|---|
| repo stars | ★ 7 |
| Last updated | May 20, 2026 |
| Repository | daemon-blockint-tech/agentic-enteprises-skill ↗ |
What it does
Establish failure-prevention culture for mission-critical systems: HRO principles, verification gates, fail-safe design, pre-mortems, FMEA, and stop-the-line policy.
Files
Zero-Tolerance for Failure
When to Use
- Establish failure-prevention culture and operating norms for mission-critical systems
- Apply HRO principles (preoccupation with failure, reluctance to simplify, sensitivity to operations, commitment to resilience, deference to expertise)
- Design defense-in-depth, fail-safe, and fail-closed controls for software, infra, and OT
- Define verification gates, independent checks, and release hold criteria
- Architect redundancy, isolation, and graceful degradation with explicit failure modes
- Facilitate pre-mortems, FMEA, and risk registers before high-stakes change
- Author stop-the-line policy and escalation when quality or safety signals are ambiguous
- Select and track prevention metrics (defect escape, near-miss, repeat incidents, gate bypass)
- Coach leadership behaviors that counter normalization of deviance and blame theater
- Brief engineering and ops on zero-defect aspiration vs error budgets for the right domain
When NOT to Use
- Own SLI/SLO definitions, error-budget policy, and burn-rate alerting →
site-reliability-engineer - Run live incident war room, SEV classification, and status communications →
incident-management-engineer - Design backup/immutability, RTO/RPO recovery, and ransomware restore architecture →
cyber-resilience-engineer - Enforce CI compile/lint/test gates without broader prevention program →
build-validator - Facilitate sprint ceremonies, backlog grooming, or team agile transformation → agile coaching skills
- Issue HR warnings, performance plans, or legal disciplinary guidance → escalate to HR/legal
- Own classified ATO/accreditation packages without operational excellence deliverables →
classified-cyber-security-senior-manager(pair for cleared context) - Produce ADRs and integration patterns without failure-prevention lens →
senior-system-architecture(pair for architecture)
Related skills
| Need | Skill |
|---|---|
| SLOs, error budgets, reliability toil, capacity | site-reliability-engineer |
| Incident program, SEV, on-call, postmortems | incident-management-engineer |
| Recovery tiers, backup/immutability, resilience tests | cyber-resilience-engineer |
| Build/CI quality gates and merge validation | build-validator |
| Cleared program governance, inspection, escalation | classified-cyber-security-senior-manager |
| NFRs, architecture review, ADRs | senior-system-architecture |
| Enterprise BCM/DR program and tabletops | bcm-disaster-recovery-specialist |
| Active CSIRT containment and forensics | incident-responder |
Core Workflows
1. Scope, limits, and charter
Clarify what “zero tolerance” means in context—aspiration, gates, and metrics—without perfectionism traps.
See `references/zero_tolerance_scope_and_limits.md`.
2. HRO principles and operating mindset
Embed high-reliability organization behaviors in teams that face rare, catastrophic failure.
See `references/high_reliability_organization_principles.md`.
3. Prevention, verification, and gates
Layer independent checks, hold points, and evidence before irreversible change.
See `references/prevention_verification_and_gates.md`.
4. Redundancy, degradation, and fail-safe design
Specify failure modes, safe defaults, and degraded operation—not only happy path.
See `references/redundancy_degradation_and_fail_safe.md`.
5. Pre-mortem, FMEA, and risk registers
Surface latent failures before launch; maintain living risk and mitigation traceability.
See `references/pre_mortem_fmea_and_risk_registers.md`.
6. Leadership, culture, and metrics
Measure prevention; reinforce stop-the-line; reduce normalization of deviance.
See `references/leadership_culture_and_metrics.md`.
Outputs
- Failure-prevention charter — scope, principles, RACI, interfaces with SRE/IR/QA
- Gate catalog — hold points, owners, evidence required, bypass rules and audit trail
- Design review pack — fail-safe/fail-closed decisions, degradation modes, verification plan
- FMEA / pre-mortem record — failure modes, causes, controls, residual risk, owners
- Stop-the-line policy — triggers, authority, duration, comms, and restart criteria
- Metrics dashboard brief — defect escape, near-miss, repeat incidents, gate effectiveness
- Leadership playbook — behaviors, rituals, and anti-patterns (learning vs blame)
Principles
- Prevent over punish — optimize systems and norms; avoid blame theater and hidden workarounds
- Fail closed by default — ambiguous safety or auth states deny; document explicit fail-open exceptions
- Independent verification — separation between build, check, and approve for high-criticality change
- Deference to expertise — elevate domain experts at the boundary; leaders ask, not override silently
- Measure escapes and near-misses — lagging severity alone rewards luck; track what almost failed
- Stop-the-line is a gift — halting bad change is success; normalize escalation without career penalty
- Pair with peers — reliability math (SRE), incidents (IR program), recovery (resilience), builds (CI)
High-reliability organization principles
Table of contents
1. HRO overview 2. Five principles in practice 3. Behaviors by role 4. Normalization of deviance 5. Classified and safety-critical overlays 6. Assessment checklist
HRO overview
High-reliability organizations operate in domains where failures are rare but catastrophic—aviation, nuclear, healthcare, military systems, large-scale cloud control planes.
HRO success depends on collective mindfulness, not heroic individuals. This reference translates Weick/Sutcliffe-style HRO ideas into engineering and operations rituals.
Five principles in practice
| Principle | Meaning | Engineering rituals |
|---|---|---|
| Preoccupation with failure | Treat small signals as data | Near-miss reviews; weak-signal dashboards; “almost” postmortems |
| Reluctance to simplify | Resist single-story explanations | Multi-cause diagrams; required dissent in design review |
| Sensitivity to operations | Frontline context shapes decisions | Ops present in planning; runbook walkthroughs before launch |
| Commitment to resilience | Absorb surprise and recover | Graceful degradation specs; rehearsed rollback; pair cyber-resilience-engineer |
| Deference to expertise | Expertise trumps rank at the boundary | Named technical authority for stop-the-line; no silent override |
Behaviors by role
| Role | Do | Avoid |
|---|---|---|
| Executive | Fund verification; protect reporters; ask “what almost failed?” | Velocity KPIs that punish escalation |
| Engineering lead | Staff reviews; gate ownership; pre-mortems on tier-0 | Rubber-stamp “LGTM” on critical paths |
| Individual contributor | Report near-miss; halt ambiguous deploy | Quiet fixes that bypass change control |
| SRE | Pair prevention metrics with SLOs (site-reliability-engineer) | Treat all risk as budgetable without tiering |
| Incident manager | Separate learning from command (incident-management-engineer) | Blame-focused postmortems |
Normalization of deviance
Normalization of deviance occurs when repeated workaround becomes “how we do it” until an accident exposes the gap.
| Signal | Example | Response |
|---|---|---|
| Drift from standard | Manual prod hotfix weekly | Stop-the-line; fix pipeline |
| Weak alarms | Alert fatigue; muted pages | Tune + fix root cause; don’t delete signal |
| Gate bypass culture | “Emergency” every Friday | Audit bypasses; tighten criteria |
| Hero dependency | One person always saves deploy | Cross-train; automate checks |
| Documentation theater | Runbooks stale since launch | Tie releases to runbook verification |
Run a norms audit quarterly: interview operators and reviewers; compare written gates to observed practice.
Classified and safety-critical overlays
| Overlay | Additional expectation |
|---|---|
| Classified | Change boards, inspection windows, two-person integrity—classified-cyber-security-senior-manager |
| Safety-critical OT | Physical interlocks; provenance; restricted remote access |
| Regulated finance/health | Segregation of duties; evidence retention for gate decisions |
Do not substitute HRO language for legal or accreditation obligations—coordinate with compliance owners.
Assessment checklist
- [ ] Near-miss mechanism exists and is used without retaliation
- [ ] Design reviews record dissent and unresolved risks
- [ ] Stop-the-line authority is named and exercised in last 12 months
- [ ] Tier-0 changes have independent verification (not author-only)
- [ ] Repeat incidents trigger system fixes, not repeat training alone
- [ ] Executives receive escape/near-miss trends, not only outage counts
Score maturity 1–4: ad hoc → defined gates → measured escapes → continuous norm reinforcement.
Leadership, culture, and metrics
Table of contents
1. Leadership behaviors 2. Stop-the-line authority 3. Metrics that matter 4. Dashboards and reviews 5. Anti-metrics and gaming 6. Communication templates
Leadership behaviors
| Behavior | Example |
|---|---|
| Ask about near-miss | Standing question in ops review: “What almost failed?” |
| Reward escalation | Recognize stop-the-line that prevented harm |
| Fund prevention | Budget verification, test envs, review capacity |
| Model curiosity | Leaders admit unknowns; invite dissent |
| Fix systems | Repeat incident → engineering change, not slogans |
| Separate blame | Learning review distinct from HR process |
| Anti-pattern | Why it fails |
|---|---|
| “Zero incidents” without near-miss data | Hides weak signals |
| Velocity at all costs | Normalizes gate bypass |
| Shooting messenger | Stops reporting |
| Review theater | Checkboxes without independent verification |
Stop-the-line authority
Define in policy:
| Element | Guidance |
|---|---|
| Who | Any trained role on tier-0/1; named backup |
| Triggers | Ambiguous safety, failed verification, unexplained metric shift, repeat near-miss |
| Actions | Halt deploy/change; freeze config; convene rapid review |
| Duration | Until exit criteria met or explicit risk acceptance |
| Communication | Notify stakeholders; no silent halt |
| Restart | Document evidence and approver |
Success metric: stop-the-line events that prevented customer or safety impact—celebrate, do not penalize.
Metrics that matter
| Metric | Definition | Why |
|---|---|---|
| Defect escape rate | Defects found in prod / total defects (by tier) | Measures gate effectiveness |
| Critical escape count | Tier-0/1 defects in prod | Lagging severity with prevention lens |
| Near-miss rate | Reported near-misses per period (normalized) | Leading indicator if culture healthy |
| Repeat incident rate | Incidents with same root category within 90d | System fix effectiveness |
| Gate bypass rate | Emergency changes / total prod changes | Culture stress signal |
| Verification coverage | % tier-0 changes with independent check | Process adherence |
| Time-to-detect | For escapes, how long until noticed | Observability quality (pair SRE) |
| Pre-mortem yield | Mitigations implemented / identified | Risk work ROI |
Not sufficient alone: raw incident count (luck), deployment frequency without escapes, individual “error counts” for blame.
Dashboards and reviews
| Cadence | Audience | Focus |
|---|---|---|
| Weekly | Engineering + ops leads | Escapes, near-miss, open mitigations |
| Monthly | Directors | Trends, bypass audit, repeat incidents |
| Quarterly | Exec + risk | Tier coverage, norms audit, investment asks |
Pair with SRE reliability review (site-reliability-engineer)—prevention metrics explain why SLOs broke, not only that they broke.
Anti-metrics and gaming
| Gaming | Countermeasure |
|---|---|
| Under-report near-miss | Anonymous channel; leader modeling |
| Reclassify severity | Independent incident review |
| Bypass without ticket | Automated change detection |
| Close risks without mitigation | Expiry and audit of accepted risks |
Communication templates
Near-miss report (short):
What almost happened:
What prevented impact:
What could fail next time:
Proposed system fix / owner / date:Stop-the-line notice:
Halted: [change/deploy]
Reason: [signal]
Owner: [name]
Next step: [review time / criteria to resume]Executive brief (monthly):
Defect escapes (tier-0/1): [n] trend
Near-miss reports: [n] trend
Repeat incidents: [n]
Gate bypasses: [n] with top reasons
Stop-the-line events: [n] outcomes
Top systemic actions: [list]Keep narratives learning-oriented—escalate individual conduct to HR outside this skill.
Pre-mortem, FMEA, and risk registers
Table of contents
1. When to run structured risk work 2. Pre-mortem workflow 3. FMEA basics 4. Risk register pattern 5. Linking to gates and metrics 6. Facilitation tips
When to run structured risk work
| Trigger | Recommended method |
|---|---|
| New tier-0/1 service or major architecture pivot | Pre-mortem + FMEA lite |
| Large migration, cutover, or flag flip at scale | Pre-mortem + rollback drill |
| Repeat incident or serious near-miss | FMEA on affected subsystem |
| Classified or safety-critical change | Formal risk register + change board |
| “We’ve always done it this way” without recent test | Norms audit + pre-mortem |
Skip heavyweight FMEA for low-tier experiments—still document residual risk and owner.
Pre-mortem workflow
Goal: Imagine future failure, then prevent it before launch.
1. Frame — objective, scope, success criteria, irreversible steps 2. Assume failure — “It is 6 months later; we failed badly. Why?” 3. Silent brainstorm — causes, blind spots, org failures (5–10 min) 4. Share and cluster — themes: technical, process, people, external 5. Mitigate — controls, gates, tests, owners, dates 6. Residual risk — accept, defer with date, or stop launch 7. Record — store in risk register; link to change ticket
Outputs: mitigation backlog, new gates, test additions, stop-the-line triggers.
FMEA basics
Failure Mode and Effects Analysis scores:
| Field | Description |
|---|---|
| Function | What the component must do |
| Failure mode | How it can fail |
| Effect | Local and system impact |
| Cause | Root mechanisms |
| Current controls | Detection/prevention already in place |
| Severity (S) | 1–10 impact if unmitigated |
| Occurrence (O) | 1–10 likelihood |
| Detection (D) | 1–10 chance current controls catch it |
| RPN | S × O × D (or use S×O with separate detection plan) |
| Action | Design/test/process change |
| Re-score | After action implemented |
Prioritize rows with high S even if RPN moderate—catastrophic rare events matter in HRO contexts.
Risk register pattern
Maintain a living register (spreadsheet or ticket epic):
| Column | Purpose |
|---|---|
| Risk ID | Stable reference |
| Description | Clear failure scenario |
| Tier / system | Scope |
| Owner | Accountable mitigator |
| Controls | Existing + planned |
| Residual level | Low/Med/High or qualitative |
| Review date | Expiry for accepted risk |
| Linked changes | Tickets, ADRs, gates |
Accepted high residual risk requires named executive or delegated authority and expiry—no permanent “accepted” without review.
Linking to gates and metrics
| Risk work output | Downstream hook |
|---|---|
| New test requirement | CI gate or release checklist (build-validator) |
| Architecture change | ADR + design review (senior-system-architecture) |
| Ops readiness gap | Ops gate evidence |
| Monitoring gap | SLI/alert work (site-reliability-engineer) |
| Recovery gap | Resilience test (cyber-resilience-engineer) |
After launch, track did any pre-mortem scenario occur? — calibrate facilitation quality.
Facilitation tips
- Invite skeptics and operators, not only builders
- Use premortem.google style anonymity if hierarchy suppresses speech
- Separate learning from go/no-go decision in the meeting record
- Time-box; assign owners before adjourn
- Re-run when scope changes materially—do not shelf the doc
Do not use pre-mortem output for performance punishment—undermines HRO reporting culture.
Prevention, verification, and gates
Table of contents
1. Defense in depth 2. Verification layers 3. Gate catalog pattern 4. Independent checks 5. Bypass and emergency change 6. Evidence and audit trail
Defense in depth
Layer controls so no single failure defeats safety or integrity:
| Layer | Examples |
|---|---|
| Requirements | Threat model, misuse cases, safety cases |
| Design | Fail-closed authZ, least privilege, blast-radius limits |
| Implementation | Static analysis, invariant tests, fuzzing on parsers |
| Build | Reproducible builds, signed artifacts (build-validator) |
| Deploy | Canary, progressive delivery, automated rollback |
| Runtime | Rate limits, circuit breakers, anomaly detection |
| Operations | Runbook verification, game days, access reviews |
Map layers to tier (crown jewels vs standard services)—do not apply maximum depth everywhere without cost trade-off documentation.
Verification layers
| Layer | Intent | Typical artifacts |
|---|---|---|
| Self-check | Author confidence | Unit tests, local integration |
| Peer review | Catch logic and design gaps | PR review, architecture review |
| Independent verification | Separation of duties | QA sign-off, security review, ops readiness |
| Environment proof | Behavior in realistic conditions | Staging soak, shadow traffic, chaos in non-prod |
| Operational proof | Humans can run and roll back | Runbook drill, on-call shadow |
Rule: Tier-0/1 requires at least one independent layer beyond author self-check before production.
Gate catalog pattern
Document each hold point in a gate catalog:
| Field | Description |
|---|---|
| Gate ID | Stable identifier |
| Trigger | Change types that require this gate |
| Owner | Role accountable for pass/fail |
| Entry criteria | What must be true to start review |
| Evidence | Tests, scans, sign-offs, diagrams |
| Exit criteria | Objective pass conditions |
| SLA | Max wait time; escalation path |
| Bypass | Who can authorize; required post-facto review |
Example gates: architecture review, threat model, ops readiness, DR fit check (cyber-resilience-engineer), classified change board (classified-cyber-security-senior-manager).
Independent checks
| Practice | Purpose |
|---|---|
| Four-eyes on prod change | Reduce single-actor risk |
| Reviewer ≠ author | On critical merges and config pushes |
| Automated policy gates | OPA, admission control, protected branches |
| Sampling audit | Random deep review of “routine” changes |
| Red team / tabletop | Validate assumptions before launch |
Pair automated CI gates (build-validator) with human judgment for context CI cannot see.
Bypass and emergency change
| Requirement | Detail |
|---|---|
| Narrow criteria | Safety, active incident, regulatory deadline—document which |
| Time box | Maximum bypass duration; mandatory follow-up gate |
| Approver | Named role; cannot be sole implementer |
| Telemetry | Log bypass ID in change ticket and monitoring |
| After-action | Retro within 5 business days; convert to permanent fix |
Track bypass rate and repeat bypass reasons—chronic bypass indicates broken gates, not heroic ops.
Evidence and audit trail
Store for each gate passage:
- Ticket/link to change record
- Test results, scan reports, review notes
- Approver identity and timestamp
- Residual risks accepted with owner and date
Evidence must be durable and searchable for inspections—not only chat threads.
Redundancy, degradation, and fail-safe design
Table of contents
1. Failure mode thinking 2. Fail-safe vs fail-closed vs fail-open 3. Redundancy patterns 4. Graceful degradation 5. OT and classified considerations 6. Design review prompts
Failure mode thinking
For each component, document:
| Question | Output |
|---|---|
| What can fail? | Failure mode list |
| How does it fail? | Safe vs unsafe failure states |
| What is detected? | Monitoring, health checks, watchdogs |
| What is the response? | Failover, degrade, halt, manual procedure |
| What is the blast radius? | Users, data, safety, compliance |
Use FMEA (references/pre_mortem_fmea_and_risk_registers.md) for structured scoring on high-criticality systems.
Fail-safe vs fail-closed vs fail-open
| Posture | When to use | Examples |
|---|---|---|
| Fail-closed | Default deny when uncertain | AuthZ, firewall on auth failure, payment holds |
| Fail-safe | Move to physically or logically safe state | OT interlock engages; drain connections |
| Fail-open | Availability over immediate security—requires explicit approval | Rare; document compensating controls |
Default: Prefer fail-closed for security and integrity; fail-safe for safety; document any fail-open with risk owner and monitoring.
Redundancy patterns
| Pattern | Benefit | Caveat |
|---|---|---|
| N+1 capacity | Survive single node loss | Correlated failures (AZ, firmware) |
| Active/active | Low failover time | Split-brain needs fencing |
| Active/passive | Simpler consistency | Failover untested = unknown RTO |
| Geographic redundancy | Regional disaster | Data residency, replication lag |
| Diverse implementations | Common-mode bug resistance | Cost and complexity |
Redundancy without tested failover is wishful thinking—schedule exercises; pair cyber-resilience-engineer for recovery proof.
Graceful degradation
Define degraded modes explicitly:
| Mode | Capability retained | Disabled or limited | User/comms expectation |
|---|---|---|---|
| Full | All features | — | Normal |
| Degraded | Core path only | Non-critical features | Banner / status |
| Read-only | Query, no mutation | Writes | Clear messaging |
| Maintenance | None or admin only | Public API | Planned window |
Requirements:
- Deterministic defaults — feature flags and config default to safer mode
- Backpressure — shed load before unbounded queue growth
- Cascading failure containment — timeouts, bulkheads, circuit breakers
- Observability — metric per degradation mode; alert on unintended entry
OT and classified considerations
| Domain | Extra design rules |
|---|---|
| OT | Manual override visible and logged; no silent remote setpoint change |
| Classified | Redundant paths meet accreditation; no cross-domain leakage on failover |
| Dual-use cloud | Sovereignty and key custody on failover—classified-cyber-security-senior-manager |
Design review prompts
- [ ] Unsafe failure states identified and eliminated where possible
- [ ] Fail-closed/default-deny documented for auth and data paths
- [ ] Degraded modes defined with monitoring and runbooks
- [ ] Redundancy tested in last 12 months (or justified exception)
- [ ] Single points of failure listed with accepted risk owner
- [ ] Rollback/automation path exists for deploy-induced failure
Capture decisions in ADR or architecture pack—pair senior-system-architecture for NFR traceability.
Zero-tolerance scope and limits
Table of contents
1. Purpose and boundaries 2. Zero-defect aspiration vs error budgets 3. Perfectionism and blame traps 4. In scope vs out of scope 5. Domains of application 6. Charter elements
Purpose and boundaries
Zero-tolerance for failure (operational excellence lens) means building systems and culture where preventable failures are designed out, caught early, and learned from—without claiming literal zero defects everywhere.
This skill sits between:
- Reliability operations (
site-reliability-engineer) — SLOs, error budgets, capacity - Incident program (
incident-management-engineer) — SEV, on-call, postmortem process - Recovery engineering (
cyber-resilience-engineer) — backup, RTO/RPO, restore tests - Build validation (
build-validator) — automated CI gates
Own prevention philosophy, verification depth, HRO behaviors, and prevention metrics—not paging policy, restore architecture, or lint rules alone.
Zero-defect aspiration vs error budgets
| Concept | Appropriate use | Misuse |
|---|---|---|
| Zero-defect aspiration | Safety-critical, auth, payments, classified control planes, OT interlocks | Demanding 100% uptime on experimental features |
| Error budget | Customer-facing SLO trade-offs, release velocity vs reliability (site-reliability-engineer) | Excusing preventable escapes in tier-0 controls |
| Zero critical escapes | Defects that reach production in tier-0/1 without gate bypass | Counting all minor UI bugs as culture failures |
| Near-miss reporting | Learning system; mandatory in HRO contexts | Punitive scorecards tied to individual blame |
Decision rule: If failure implies harm, compromise, or regulatory breach, bias to prevention, gates, and fail-closed design. If failure implies degraded UX or revenue within agreed SLO, pair with error budgets and mitigation—still track escapes.
Perfectionism and blame traps
| Healthy zero-tolerance | Toxic pseudo zero-tolerance |
|---|---|
| Fix systems that allow repeat incidents | Fire individuals for first honest mistake |
| Celebrate stop-the-line and near-miss reports | Hide problems to meet velocity metrics |
| Invest in verification and independent review | Add paperwork without changing risk |
| Transparent residual risk with owners | “Zero incidents” slogans with no measurement |
| Learning reviews with action owners | Public shaming or “name and blame” postmortems |
Escalate HR or legal disciplinary questions out of this skill—stay on engineering and operational norms.
In scope vs out of scope
| In scope | Out of scope (route to peer skill) |
|---|---|
| Failure-prevention charter, principles, RACI | SLO/error-budget policy → site-reliability-engineer |
| Verification gates, independent checks, hold points | War-room command → incident-management-engineer |
| Fail-safe/fail-closed and degradation design | Backup/immutability architecture → cyber-resilience-engineer |
| Pre-mortem, FMEA, risk registers | CI job definitions only → build-validator |
| Stop-the-line triggers and authority | ATO/accreditation packages → classified-cyber-security-senior-manager |
| Defect escape, near-miss, repeat-incident metrics | ADR templates without prevention lens → senior-system-architecture |
Domains of application
| Domain | Prevention emphasis |
|---|---|
| Software | AuthZ fail-closed, idempotency, invariant tests, canary + rollback, feature flags with safe defaults |
| Infrastructure | Blast-radius limits, config drift detection, staged rollouts, break-glass with audit |
| Safety-critical OT | Interlocks, manual overrides logged, dual-channel commands, provenance of setpoints |
| Classified programs | Two-person rules, change boards, inspection readiness—pair classified-cyber-security-senior-manager |
Charter elements
1. Scope — tiers/systems covered; explicit exclusions 2. Principles — fail-closed, independent verification, stop-the-line 3. Governance — who can bypass gates; audit requirements 4. Interfaces — SRE, IR, QA, security, BCM 5. Metrics — defect escape, near-miss, repeat incidents, gate bypass rate 6. Review cadence — quarterly norm check; after material escape or near-catastrophe
Store charter and gate definitions in a durable, version-controlled repository accessible to builders and reviewers.