Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
aiskillstore avatar

Resilience Analysis

  • 1 installs
  • 404 repo stars
  • Updated August 5, 2026
  • aiskillstore/marketplace

resilience-analysis is a skill that assesses error handling, isolation boundaries, and recovery mechanisms in agent frameworks.

About

This skill assesses error handling, isolation boundaries, and recovery mechanisms in agent frameworks. It traces error propagation, evaluates sandboxing for code execution, and catalogs retry, fallback, and circuit-breaker patterns. A developer uses it to judge an agent framework's production readiness and identify failure modes.

  • Assesses error handling and isolation in agent frameworks
  • Maps error propagation, sandboxing, and recovery patterns
  • Compares sandbox mechanisms from RestrictedPython to Firecracker

Resilience Analysis by the numbers

  • 1 all-time installs (skills.sh)
  • Ranked #14,102 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
At a glance

resilience-analysis capabilities & compatibility

Capabilities
code review · security audit
Use cases
code review · security audit
Pricing
Free
From the docs

What resilience-analysis says it does

Assess error handling, isolation boundaries, and recovery mechanisms in agent frameworks.
SKILL.md
Assesses error handling and isolation boundaries.
SKILL.md
npx skills add https://github.com/aiskillstore/marketplace --skill resilience-analysis

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs1
repo stars404
Last updatedAugust 5, 2026
Repositoryaiskillstore/marketplace

What it does

Assess an agent framework's error propagation, sandboxing, and recovery for production readiness.

Who is it for?

Engineers evaluating an agent framework's resilience and production readiness.

Skip if: General app error handling unrelated to agent frameworks.

When should I use this skill?

Tracing error propagation, evaluating sandboxing, or assessing an agent framework's failure modes.

What you get

Produces an error-propagation map, sandboxing assessment, and recovery-pattern catalog for a framework.

  • error-propagation map
  • sandboxing assessment
  • recovery-pattern catalog

By the numbers

  • 6-row sandboxing mechanisms comparison table
  • 4-step analysis process (trace, isolate, catalog, assess)

Files

SKILL.mdMarkdownGitHub ↗

Resilience Analysis

Assesses error handling and isolation boundaries.

Process

1. Trace error propagation — Map exception flow from tools to agent 2. Identify isolation — Sandbox mechanisms for dangerous operations 3. Catalog recovery — Retry logic, fallbacks, circuit breakers 4. Assess boundaries — What crashes propagate vs. are contained

Error Propagation Analysis

Questions to Answer

1. Does a tool exception terminate the agent? 2. Are LLM API errors retried automatically? 3. Is parsing failure (malformed output) recoverable? 4. What happens when state updates fail?

Propagation Patterns

Crash Propagation (Dangerous)

def run_tool(self, tool, args):
    return tool.execute(args)  # Exception bubbles up

Exception Wrapping

def run_tool(self, tool, args):
    try:
        return tool.execute(args)
    except Exception as e:
        raise ToolExecutionError(tool.name, e) from e

Error Containment

def run_tool(self, tool, args):
    try:
        return ToolResult(success=True, output=tool.execute(args))
    except Exception as e:
        return ToolResult(success=False, error=str(e))

Propagation Map Template

User Input
    ↓
┌─────────────────────────────────────────┐
│ Agent Loop                              │
│   ↓                                     │
│ ┌─────────────────────────────────────┐ │
│ │ LLM Call                            │ │
│ │ • APIError → [Retry 3x / Propagate] │ │
│ │ • RateLimit → [Backoff / Propagate] │ │
│ │ • Timeout → [Retry / Propagate]     │ │
│ └─────────────────────────────────────┘ │
│   ↓                                     │
│ ┌─────────────────────────────────────┐ │
│ │ Output Parsing                      │ │
│ │ • ParseError → [Retry / Contained]  │ │
│ │ • ValidationError → [Contained]     │ │
│ └─────────────────────────────────────┘ │
│   ↓                                     │
│ ┌─────────────────────────────────────┐ │
│ │ Tool Execution                      │ │
│ │ • ToolError → [Feedback to LLM]     │ │
│ │ • Timeout → [Kill / Continue]       │ │
│ │ • SecurityError → [Propagate]       │ │
│ └─────────────────────────────────────┘ │
└─────────────────────────────────────────┘

Sandboxing Mechanisms

Code Execution Isolation

MechanismSafety LevelPerformanceComplexity
None⚠️ DangerousFastNone
RestrictedPythonMediumFastLow
AST ValidationLowFastMedium
SubprocessMediumOverheadLow
Docker/ContainerHighHigh overheadMedium
gVisor/FirecrackerVery HighMedium overheadHigh

Detection Patterns

No Sandboxing

exec(user_code)  # Direct execution
eval(expression)  # Direct eval
subprocess.run(cmd, shell=True)  # Shell injection risk

Basic Sandboxing

# RestrictedPython
from RestrictedPython import compile_restricted
code = compile_restricted(user_code, '<string>', 'exec')

# AST validation
tree = ast.parse(user_code)
if has_dangerous_nodes(tree):
    raise SecurityError()

Process Isolation

# Subprocess with limits
result = subprocess.run(
    ['python', '-c', user_code],
    timeout=30,
    capture_output=True,
    user='nobody'  # Drop privileges
)

Container Isolation

import docker
client = docker.from_env()
container = client.containers.run(
    'python:3.11-slim',
    command=['python', '-c', user_code],
    mem_limit='256m',
    network_disabled=True,
    remove=True
)

Recovery Patterns

Retry Logic

# Simple retry
@retry(max_attempts=3, backoff=exponential)
def call_llm(self, prompt):
    return self.client.generate(prompt)

# Retry with error feedback
def call_with_retry(self, prompt, max_retries=3):
    errors = []
    for i in range(max_retries):
        try:
            return self.llm.generate(prompt)
        except ParseError as e:
            errors.append(str(e))
            prompt = f"{prompt}\n\nPrevious errors: {errors}"
    raise MaxRetriesExceeded(errors)

Fallback Mechanisms

def generate(self, prompt):
    try:
        return self.primary_llm.generate(prompt)
    except APIError:
        return self.fallback_llm.generate(prompt)

Circuit Breaker

class CircuitBreaker:
    def __init__(self, failure_threshold=5, reset_timeout=60):
        self.failures = 0
        self.state = 'closed'
        self.last_failure = None
    
    def call(self, func, *args):
        if self.state == 'open':
            if time.time() - self.last_failure > self.reset_timeout:
                self.state = 'half-open'
            else:
                raise CircuitOpen()
        
        try:
            result = func(*args)
            self.failures = 0
            self.state = 'closed'
            return result
        except Exception as e:
            self.failures += 1
            self.last_failure = time.time()
            if self.failures >= self.failure_threshold:
                self.state = 'open'
            raise

Output Template

## Resilience Analysis: [Framework Name]

### Error Propagation Map

| Error Source | Error Type | Handling | Propagates? |
|--------------|-----------|----------|-------------|
| LLM API | RateLimitError | Retry 3x with backoff | No |
| LLM API | APIError | Retry 1x | Yes |
| Parser | ParseError | Feed back to LLM | No |
| Tool | Exception | Wrap and feed to LLM | No |
| Tool | Timeout | Kill process | No |
| State | ValidationError | Propagate | Yes |

### Sandboxing Assessment
- **Code Execution**: [Mechanism or None]
- **File System**: [Isolated/Restricted/Open]
- **Network**: [Blocked/Filtered/Open]
- **Resource Limits**: [Memory/CPU/Time limits]

### Recovery Mechanisms

| Pattern | Implementation | Location |
|---------|---------------|----------|
| Retry | Exponential backoff, 3 attempts | llm.py:L45 |
| Fallback | Secondary model | agent.py:L120 |
| Circuit Breaker | None | - |

### Risk Assessment
- **Critical Gaps**: [List any missing protections]
- **Production Ready**: [Yes/No/Needs work]

Integration

  • Prerequisite: codebase-mapping to identify execution code
  • Feeds into: antipattern-catalog for error handling issues
  • Related: execution-engine-analysis for async error handling

Related skills

FAQ

What does it assess?

Error propagation, isolation boundaries, and recovery mechanisms in agent frameworks.

Which sandboxing options does it compare?

None, RestrictedPython, AST validation, subprocess, Docker containers, and gVisor/Firecracker.

AI & Agent Buildingagentsautomation

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.