
Resilience Analysis
- 1 installs
- 404 repo stars
- Updated August 5, 2026
- aiskillstore/marketplace
resilience-analysis is a skill that assesses error handling, isolation boundaries, and recovery mechanisms in agent frameworks.
About
This skill assesses error handling, isolation boundaries, and recovery mechanisms in agent frameworks. It traces error propagation, evaluates sandboxing for code execution, and catalogs retry, fallback, and circuit-breaker patterns. A developer uses it to judge an agent framework's production readiness and identify failure modes.
- Assesses error handling and isolation in agent frameworks
- Maps error propagation, sandboxing, and recovery patterns
- Compares sandbox mechanisms from RestrictedPython to Firecracker
Resilience Analysis by the numbers
- 1 all-time installs (skills.sh)
- Ranked #14,102 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
resilience-analysis capabilities & compatibility
- Capabilities
- code review · security audit
- Use cases
- code review · security audit
- Pricing
- Free
What resilience-analysis says it does
Assess error handling, isolation boundaries, and recovery mechanisms in agent frameworks.
Assesses error handling and isolation boundaries.
npx skills add https://github.com/aiskillstore/marketplace --skill resilience-analysisAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 1 |
|---|---|
| repo stars | ★ 404 |
| Last updated | August 5, 2026 |
| Repository | aiskillstore/marketplace ↗ |
What it does
Assess an agent framework's error propagation, sandboxing, and recovery for production readiness.
Who is it for?
Engineers evaluating an agent framework's resilience and production readiness.
Skip if: General app error handling unrelated to agent frameworks.
When should I use this skill?
Tracing error propagation, evaluating sandboxing, or assessing an agent framework's failure modes.
What you get
Produces an error-propagation map, sandboxing assessment, and recovery-pattern catalog for a framework.
- error-propagation map
- sandboxing assessment
- recovery-pattern catalog
By the numbers
- 6-row sandboxing mechanisms comparison table
- 4-step analysis process (trace, isolate, catalog, assess)
Files
Resilience Analysis
Assesses error handling and isolation boundaries.
Process
1. Trace error propagation — Map exception flow from tools to agent 2. Identify isolation — Sandbox mechanisms for dangerous operations 3. Catalog recovery — Retry logic, fallbacks, circuit breakers 4. Assess boundaries — What crashes propagate vs. are contained
Error Propagation Analysis
Questions to Answer
1. Does a tool exception terminate the agent? 2. Are LLM API errors retried automatically? 3. Is parsing failure (malformed output) recoverable? 4. What happens when state updates fail?
Propagation Patterns
Crash Propagation (Dangerous)
def run_tool(self, tool, args):
return tool.execute(args) # Exception bubbles upException Wrapping
def run_tool(self, tool, args):
try:
return tool.execute(args)
except Exception as e:
raise ToolExecutionError(tool.name, e) from eError Containment
def run_tool(self, tool, args):
try:
return ToolResult(success=True, output=tool.execute(args))
except Exception as e:
return ToolResult(success=False, error=str(e))Propagation Map Template
User Input
↓
┌─────────────────────────────────────────┐
│ Agent Loop │
│ ↓ │
│ ┌─────────────────────────────────────┐ │
│ │ LLM Call │ │
│ │ • APIError → [Retry 3x / Propagate] │ │
│ │ • RateLimit → [Backoff / Propagate] │ │
│ │ • Timeout → [Retry / Propagate] │ │
│ └─────────────────────────────────────┘ │
│ ↓ │
│ ┌─────────────────────────────────────┐ │
│ │ Output Parsing │ │
│ │ • ParseError → [Retry / Contained] │ │
│ │ • ValidationError → [Contained] │ │
│ └─────────────────────────────────────┘ │
│ ↓ │
│ ┌─────────────────────────────────────┐ │
│ │ Tool Execution │ │
│ │ • ToolError → [Feedback to LLM] │ │
│ │ • Timeout → [Kill / Continue] │ │
│ │ • SecurityError → [Propagate] │ │
│ └─────────────────────────────────────┘ │
└─────────────────────────────────────────┘Sandboxing Mechanisms
Code Execution Isolation
| Mechanism | Safety Level | Performance | Complexity |
|---|---|---|---|
| None | ⚠️ Dangerous | Fast | None |
| RestrictedPython | Medium | Fast | Low |
| AST Validation | Low | Fast | Medium |
| Subprocess | Medium | Overhead | Low |
| Docker/Container | High | High overhead | Medium |
| gVisor/Firecracker | Very High | Medium overhead | High |
Detection Patterns
No Sandboxing
exec(user_code) # Direct execution
eval(expression) # Direct eval
subprocess.run(cmd, shell=True) # Shell injection riskBasic Sandboxing
# RestrictedPython
from RestrictedPython import compile_restricted
code = compile_restricted(user_code, '<string>', 'exec')
# AST validation
tree = ast.parse(user_code)
if has_dangerous_nodes(tree):
raise SecurityError()Process Isolation
# Subprocess with limits
result = subprocess.run(
['python', '-c', user_code],
timeout=30,
capture_output=True,
user='nobody' # Drop privileges
)Container Isolation
import docker
client = docker.from_env()
container = client.containers.run(
'python:3.11-slim',
command=['python', '-c', user_code],
mem_limit='256m',
network_disabled=True,
remove=True
)Recovery Patterns
Retry Logic
# Simple retry
@retry(max_attempts=3, backoff=exponential)
def call_llm(self, prompt):
return self.client.generate(prompt)
# Retry with error feedback
def call_with_retry(self, prompt, max_retries=3):
errors = []
for i in range(max_retries):
try:
return self.llm.generate(prompt)
except ParseError as e:
errors.append(str(e))
prompt = f"{prompt}\n\nPrevious errors: {errors}"
raise MaxRetriesExceeded(errors)Fallback Mechanisms
def generate(self, prompt):
try:
return self.primary_llm.generate(prompt)
except APIError:
return self.fallback_llm.generate(prompt)Circuit Breaker
class CircuitBreaker:
def __init__(self, failure_threshold=5, reset_timeout=60):
self.failures = 0
self.state = 'closed'
self.last_failure = None
def call(self, func, *args):
if self.state == 'open':
if time.time() - self.last_failure > self.reset_timeout:
self.state = 'half-open'
else:
raise CircuitOpen()
try:
result = func(*args)
self.failures = 0
self.state = 'closed'
return result
except Exception as e:
self.failures += 1
self.last_failure = time.time()
if self.failures >= self.failure_threshold:
self.state = 'open'
raiseOutput Template
## Resilience Analysis: [Framework Name]
### Error Propagation Map
| Error Source | Error Type | Handling | Propagates? |
|--------------|-----------|----------|-------------|
| LLM API | RateLimitError | Retry 3x with backoff | No |
| LLM API | APIError | Retry 1x | Yes |
| Parser | ParseError | Feed back to LLM | No |
| Tool | Exception | Wrap and feed to LLM | No |
| Tool | Timeout | Kill process | No |
| State | ValidationError | Propagate | Yes |
### Sandboxing Assessment
- **Code Execution**: [Mechanism or None]
- **File System**: [Isolated/Restricted/Open]
- **Network**: [Blocked/Filtered/Open]
- **Resource Limits**: [Memory/CPU/Time limits]
### Recovery Mechanisms
| Pattern | Implementation | Location |
|---------|---------------|----------|
| Retry | Exponential backoff, 3 attempts | llm.py:L45 |
| Fallback | Secondary model | agent.py:L120 |
| Circuit Breaker | None | - |
### Risk Assessment
- **Critical Gaps**: [List any missing protections]
- **Production Ready**: [Yes/No/Needs work]Integration
- Prerequisite:
codebase-mappingto identify execution code - Feeds into:
antipattern-catalogfor error handling issues - Related:
execution-engine-analysisfor async error handling
{
"schema_version": "2.0",
"meta": {
"generated_at": "2026-01-17T04:40:11.603Z",
"slug": "dowwie-resilience-analysis",
"source_url": "https://github.com/Dowwie/agent_framework_study/tree/main/.claude/skills/resilience-analysis",
"source_ref": "main",
"model": "claude",
"analysis_version": "3.0.0",
"source_type": "community",
"content_hash": "aada6391d7d22272631c2cea796140b7fb020336e07f12eca8e19ea2a398504f",
"tree_hash": "8361f3ecc6934334c47be222ebf88bb48b46e8086e8d4708abe0de11073b4c7e"
},
"skill": {
"name": "resilience-analysis",
"description": "Assess error handling, isolation boundaries, and recovery mechanisms in agent frameworks. Use when (1) tracing error propagation paths, (2) evaluating sandboxing for code execution, (3) understanding retry and fallback mechanisms, (4) assessing production readiness, or (5) identifying failure modes and recovery patterns.",
"summary": "Assess error handling, isolation boundaries, and recovery mechanisms in agent frameworks. Use when (...",
"icon": "🛡️",
"version": "1.0.0",
"author": "Dowwie",
"license": "MIT",
"category": "research",
"tags": [
"error-handling",
"security-analysis",
"agent-frameworks",
"resilience",
"production-readiness"
],
"supported_tools": [
"claude",
"codex",
"claude-code"
],
"risk_factors": [
"network",
"filesystem",
"scripts",
"external_commands"
]
},
"security_audit": {
"risk_level": "safe",
"is_blocked": false,
"safe_to_publish": true,
"summary": "Pure documentation skill with no executable code. Contains only markdown guidance for analyzing error handling and recovery patterns. All 45 static findings are false positives - the code snippets in SKILL.md are educational examples showing dangerous patterns to AVOID, not actual implementation code. No network, filesystem, or command execution capabilities exist in this skill.",
"risk_factor_evidence": [
{
"factor": "network",
"evidence": [
{
"file": "skill-report.json",
"line_start": 6,
"line_end": 6
}
]
},
{
"factor": "filesystem",
"evidence": [
{
"file": "skill-report.json",
"line_start": 6,
"line_end": 6
}
]
},
{
"factor": "scripts",
"evidence": [
{
"file": "SKILL.md",
"line_start": 100,
"line_end": 100
},
{
"file": "SKILL.md",
"line_start": 99,
"line_end": 99
}
]
},
{
"factor": "external_commands",
"evidence": [
{
"file": "SKILL.md",
"line_start": 99,
"line_end": 99
},
{
"file": "SKILL.md",
"line_start": 101,
"line_end": 101
},
{
"file": "SKILL.md",
"line_start": 119,
"line_end": 119
},
{
"file": "SKILL.md",
"line_start": 29,
"line_end": 32
},
{
"file": "SKILL.md",
"line_start": 32,
"line_end": 35
},
{
"file": "SKILL.md",
"line_start": 35,
"line_end": 41
},
{
"file": "SKILL.md",
"line_start": 41,
"line_end": 44
},
{
"file": "SKILL.md",
"line_start": 44,
"line_end": 50
},
{
"file": "SKILL.md",
"line_start": 50,
"line_end": 54
},
{
"file": "SKILL.md",
"line_start": 54,
"line_end": 80
},
{
"file": "SKILL.md",
"line_start": 80,
"line_end": 98
},
{
"file": "SKILL.md",
"line_start": 98,
"line_end": 102
},
{
"file": "SKILL.md",
"line_start": 102,
"line_end": 105
},
{
"file": "SKILL.md",
"line_start": 105,
"line_end": 114
},
{
"file": "SKILL.md",
"line_start": 114,
"line_end": 117
},
{
"file": "SKILL.md",
"line_start": 117,
"line_end": 125
},
{
"file": "SKILL.md",
"line_start": 125,
"line_end": 128
},
{
"file": "SKILL.md",
"line_start": 128,
"line_end": 138
},
{
"file": "SKILL.md",
"line_start": 138,
"line_end": 144
},
{
"file": "SKILL.md",
"line_start": 144,
"line_end": 160
},
{
"file": "SKILL.md",
"line_start": 160,
"line_end": 164
},
{
"file": "SKILL.md",
"line_start": 164,
"line_end": 170
},
{
"file": "SKILL.md",
"line_start": 170,
"line_end": 174
},
{
"file": "SKILL.md",
"line_start": 174,
"line_end": 199
},
{
"file": "SKILL.md",
"line_start": 199,
"line_end": 203
},
{
"file": "SKILL.md",
"line_start": 203,
"line_end": 234
},
{
"file": "SKILL.md",
"line_start": 234,
"line_end": 238
},
{
"file": "SKILL.md",
"line_start": 238,
"line_end": 239
},
{
"file": "SKILL.md",
"line_start": 239,
"line_end": 240
}
]
}
],
"critical_findings": [],
"high_findings": [],
"medium_findings": [],
"low_findings": [],
"dangerous_patterns": [],
"files_scanned": 2,
"total_lines": 426,
"audit_model": "claude",
"audited_at": "2026-01-17T04:40:11.603Z"
},
"content": {
"user_title": "Analyze agent framework resilience",
"value_statement": "Agent frameworks vary widely in error handling and recovery capabilities. This skill provides systematic methods to trace error propagation, evaluate sandboxing effectiveness, and assess production readiness before deployment.",
"seo_keywords": [
"resilience analysis",
"error handling",
"agent framework security",
"sandboxing evaluation",
"retry mechanisms",
"circuit breaker pattern",
"Claude Code",
"Claude",
"Codex",
"production readiness"
],
"actual_capabilities": [
"Trace error propagation paths through agent loops",
"Evaluate sandboxing mechanisms for code execution",
"Catalog retry logic, fallbacks, and circuit breakers",
"Assess boundary isolation between components",
"Identify failure modes and recovery patterns",
"Generate resilience assessment reports"
],
"limitations": [
"Does not execute code or modify frameworks",
"Requires codebase-mapping skill for source access",
"Cannot assess runtime behavior without deployment",
"Analysis based on static code review only"
],
"use_cases": [
{
"target_user": "Framework Developers",
"title": "Validate error handling design",
"description": "Systematically review error propagation and recovery patterns before releasing agent frameworks to production."
},
{
"target_user": "Security Auditors",
"title": "Assess isolation boundaries",
"description": "Evaluate sandboxing mechanisms and identify potential attack vectors in untrusted code execution scenarios."
},
{
"target_user": "DevOps Engineers",
"title": "Verify production readiness",
"description": "Confirm retry strategies, circuit breakers, and failure containment meet reliability requirements."
}
],
"prompt_templates": [
{
"title": "Basic Error Analysis",
"scenario": "Analyze error handling in a framework",
"prompt": "Use resilience-analysis to assess how [framework_name] handles errors. Trace the propagation paths from tool exceptions to the agent loop."
},
{
"title": "Sandboxing Review",
"scenario": "Evaluate code execution isolation",
"prompt": "Evaluate the sandboxing mechanisms in [framework_name]. What isolation level is used for code execution? Are there resource limits?"
},
{
"title": "Recovery Assessment",
"scenario": "Catalog recovery patterns",
"prompt": "Catalog all retry, fallback, and circuit breaker patterns in [framework_name]. What is the maximum retry count?"
},
{
"title": "Comprehensive Audit",
"scenario": "Full resilience review",
"prompt": "Perform a complete resilience analysis on [framework_name]. Include error propagation map, sandboxing assessment, and risk assessment."
}
],
"output_examples": [
{
"input": "Analyze error handling in AutoGPT framework",
"output": [
"## Resilience Analysis: AutoGPT",
"### Error Propagation Map",
"- LLM API errors: Retried 3x with exponential backoff, then propagated",
"- Tool exceptions: Wrapped and fed back to LLM for recovery",
"### Sandboxing Assessment",
"- Code Execution: Subprocess isolation with timeout",
"- File System: Restricted access via allowlist",
"### Risk Assessment",
"- Critical Gaps: No circuit breaker implementation",
"- Production Ready: Yes, with circuit breaker enhancement"
]
},
{
"input": "Evaluate sandboxing in LangChain",
"output": [
"## Resilience Analysis: LangChain",
"### Sandboxing Assessment",
"- Code Execution: AST validation only, no process isolation",
"- Network: Open access",
"- Resource Limits: None implemented",
"### Risk Assessment",
"- Critical Gaps: Unrestricted code execution",
"- Recommendation: Add container isolation for production"
]
},
{
"input": "Catalog recovery patterns in custom agent",
"output": [
"## Resilience Analysis: Custom Agent",
"### Recovery Mechanisms",
"- Retry: Exponential backoff, 3 attempts",
"- Fallback: Secondary LLM provider",
"- Circuit Breaker: 5 failures triggers open state",
"### Assessment",
"- Well-implemented recovery patterns",
"- Production Ready: Yes"
]
}
],
"best_practices": [
"Always trace error propagation from outermost to innermost layer to identify containment gaps",
"Use container isolation for untrusted code execution instead of relying on AST validation alone",
"Implement circuit breakers for external API calls to prevent cascade failures"
],
"anti_patterns": [
"Direct exception bubbling without wrapping or containment",
"Shell=True in subprocess calls with user-controlled input",
"No resource limits on code execution or tool operations"
],
"faq": [
{
"question": "Which agent frameworks does this skill support?",
"answer": "This skill analyzes any agent framework. It focuses on general patterns applicable across frameworks like LangChain, AutoGPT, and custom implementations."
},
{
"question": "What limits exist for analysis depth?",
"answer": "Analysis is limited to static code review. Runtime behavior and race conditions require separate testing tools."
},
{
"question": "How does this integrate with other skills?",
"answer": "Use codebase-mapping first to identify execution code paths. Results feed into antipattern-catalog for comprehensive assessment."
},
{
"question": "Is my data safe during analysis?",
"answer": "Yes. This skill only reads code files locally. No data is transmitted externally. All analysis occurs within your secure environment."
},
{
"question": "Why does analysis fail on some frameworks?",
"answer": "Frameworks using metaprogramming or code generation may hide execution paths. Consider using reflection-based analysis for such cases."
},
{
"question": "How does this compare to execution-engine-analysis?",
"answer": "Execution-engine-analysis focuses on control flow. Resilience-analysis specifically targets error handling, isolation boundaries, and recovery mechanisms."
}
]
},
"file_structure": [
{
"name": "SKILL.md",
"type": "file",
"path": "SKILL.md",
"lines": 241
}
]
}
Related skills
FAQ
What does it assess?
Error propagation, isolation boundaries, and recovery mechanisms in agent frameworks.
Which sandboxing options does it compare?
None, RestrictedPython, AST validation, subprocess, Docker containers, and gVisor/Firecracker.