
Systematic Debugging
- 29 installs
- 17 repo stars
- Updated May 14, 2026
- delphine-l/claude_global
Helps with debugging tasks.
About
systematic-debugging is a Claude Code skill for debugging. It helps solo builders move faster with AI-assisted development.
- systematic-debugging
- Debugging
- AI-coding skill
Systematic Debugging by the numbers
- 29 all-time installs (skills.sh)
- Ranked #365 of 596 Debugging skills by installs in the Skillselion catalog
- Data as of Jul 29, 2026 (Skillselion catalog sync)
npx skills add https://github.com/delphine-l/claude_global --skill systematic-debuggingAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 29 |
|---|---|
| repo stars | ★ 17 |
| Last updated | May 14, 2026 |
| Repository | delphine-l/claude_global ↗ |
What it does
Helps with debugging tasks.
Files
Systematic Debugging
Overview
Random fixes waste time and create new bugs. Quick patches mask underlying issues.
Core principle: ALWAYS find root cause before attempting fixes.
When to Use
Use for ANY technical issue: test failures, unexpected behavior, pipeline errors, build failures, Galaxy workflow errors, notebook exceptions, environment issues.
Use ESPECIALLY when:
- "Just one quick fix" seems obvious
- You've already tried multiple fixes
- Previous fix didn't work
- You don't fully understand the issue
Supporting Files
- [root-cause-tracing.md](root-cause-tracing.md) - Trace bugs backward through call chain to find the original trigger. Instrumentation techniques, stack trace analysis.
- [defense-in-depth.md](defense-in-depth.md) - Add validation at multiple layers after finding root cause. Entry point, business logic, environment guards, debug logging.
The Four Phases
Complete each phase before proceeding to the next.
Phase 1: Root Cause Investigation
BEFORE attempting ANY fix:
1. Read error messages carefully
- Don't skip past errors or warnings — they often contain the solution
- Read stack traces completely
- Note line numbers, file paths, error codes
2. Reproduce consistently
- Can you trigger it reliably? What are the exact steps?
- If not reproducible, gather more data — don't guess
3. Check recent changes
- Git diff, recent commits, new dependencies
- Config changes, environmental differences
4. Gather evidence in multi-component systems
- For pipelines (Galaxy workflow → tool → data), log what enters and exits each component
- Run once with diagnostics to see WHERE it breaks
- Then investigate that specific component
5. Trace data flow
- Where does the bad value originate? (See root-cause-tracing.md)
- Keep tracing up the call chain until you find the source
- Fix at source, not at symptom
Phase 2: Pattern Analysis
1. Find working examples — similar working code in same codebase 2. Compare against references — read reference implementation completely, don't skim 3. Identify differences — list every difference, however small 4. Understand dependencies — settings, config, environment, assumptions
Phase 3: Hypothesis and Testing
1. Form single hypothesis — "I think X is the root cause because Y" 2. Test minimally — smallest possible change, one variable at a time 3. Verify — did it work? If not, form NEW hypothesis. Don't pile fixes on top. 4. When you don't know — say so. Don't pretend. Research more.
Phase 4: Implementation
1. Create failing test/reproduction — simplest possible, automated if possible 2. Implement single fix — address root cause, ONE change, no "while I'm here" improvements 3. Verify fix — test passes? No other tests broken? Issue resolved? 4. If fix doesn't work:
- Count fixes attempted
- If < 3: return to Phase 1, re-analyze with new information
- If >= 3: STOP — question the architecture (see below)
When 3+ Fixes Fail
Pattern indicating architectural problem:
- Each fix reveals new issues in different places
- Fixes require "massive refactoring"
- Each fix creates new symptoms elsewhere
STOP and discuss with the user before attempting more fixes. This is not a failed hypothesis — it's a wrong approach.
Red Flags — STOP and Return to Phase 1
If you catch yourself thinking:
- "Quick fix for now, investigate later"
- "Just try changing X and see if it works"
- "It's probably X, let me fix that"
- "I don't fully understand but this might work"
- Proposing solutions before tracing data flow
- "One more fix attempt" when already tried 2+
Common Rationalizations
| Excuse | Reality |
|---|---|
| "Issue is simple, don't need process" | Simple issues have root causes too |
| "Emergency, no time for process" | Systematic is FASTER than guess-and-check |
| "Just try this first, then investigate" | First fix sets the pattern. Do it right. |
| "Multiple fixes at once saves time" | Can't isolate what worked. Causes new bugs. |
| "I see the problem, let me fix it" | Seeing symptoms != understanding root cause |
Quick Reference
| Phase | Key Activities | Done when |
|---|---|---|
| 1. Root Cause | Read errors, reproduce, check changes, trace data | Understand WHAT and WHY |
| 2. Pattern | Find working examples, compare | Differences identified |
| 3. Hypothesis | Form theory, test minimally | Confirmed or new hypothesis |
| 4. Implementation | Create test, fix, verify | Bug resolved, tests pass |
Attribution
Adapted from obra/superpowers systematic-debugging skill.
Defense-in-Depth Validation
When you fix a bug caused by invalid data, adding validation at one place feels sufficient. But that single check can be bypassed by different code paths, refactoring, or edge cases.
Core principle: Validate at EVERY layer data passes through. Make the bug structurally impossible.
Why Multiple Layers
- Single validation: "We fixed the bug"
- Multiple layers: "We made the bug impossible"
Different layers catch different cases:
- Entry validation catches most bugs
- Business logic catches edge cases
- Environment guards prevent context-specific dangers
- Debug logging helps when other layers fail
The Four Layers
Layer 1: Entry Point Validation
Reject obviously invalid input at the API/function boundary.
def run_analysis(input_file, output_dir):
if not input_file or not os.path.exists(input_file):
raise ValueError(f"Input file not found: {input_file}")
if not output_dir or not os.path.isdir(output_dir):
raise ValueError(f"Output directory invalid: {output_dir}")Layer 2: Business Logic Validation
Ensure data makes sense for this specific operation.
def align_sequences(reference, reads):
if os.path.getsize(reference) == 0:
raise ValueError("Reference file is empty")
# Verify format matches expectationLayer 3: Environment Guards
Prevent dangerous operations in specific contexts.
def delete_intermediate_files(directory):
# Never delete outside the project working directory
if not os.path.abspath(directory).startswith(PROJECT_ROOT):
raise ValueError(f"Refusing to delete outside project: {directory}")Layer 4: Debug Instrumentation
Capture context for forensics when issues do occur.
import logging
logger = logging.getLogger(__name__)
def process_sample(sample_id, data_path):
logger.debug(f"Processing {sample_id} from {data_path}, "
f"exists={os.path.exists(data_path)}, "
f"size={os.path.getsize(data_path) if os.path.exists(data_path) else 'N/A'}")Applying the Pattern
When you find a bug:
1. Trace the data flow — where does bad value originate? Where is it used? 2. Map all checkpoints — list every point data passes through 3. Add validation at each layer — entry, business, environment, debug 4. Test each layer — try to bypass layer 1, verify layer 2 catches it
Key Insight
Don't stop at one validation point. Each layer catches bugs the others miss: different code paths bypass entry validation, edge cases bypass business logic, and debug logging identifies structural misuse.
Adapted from obra/superpowers.
Root Cause Tracing
Bugs often manifest deep in the call stack. Your instinct is to fix where the error appears, but that's treating a symptom.
Core principle: Trace backward through the call chain until you find the original trigger, then fix at the source.
When to Use
- Error happens deep in execution (not at entry point)
- Stack trace shows long call chain
- Unclear where invalid data originated
- Multi-component systems (pipeline → tool → data)
The Tracing Process
1. Observe the Symptom
Error: File not found: /path/to/expected/output.bam2. Find Immediate Cause
What code directly causes this? What function, what line?
3. Ask: What Called This?
Trace the call chain upward:
process_alignment(input_file)
→ called by run_pipeline(sample)
→ called by batch_process(samples)
→ called by main()4. Keep Tracing Up
What value was passed? Where did it come from?
- Was the path constructed wrong?
- Was a variable empty/None?
- Was configuration missing?
5. Find Original Trigger
The root cause is often far from where the error appears:
- A config file missing a key
- An environment variable not set
- A previous pipeline step that silently produced no output
- A data format assumption that doesn't hold
Adding Diagnostic Instrumentation
When you can't trace manually, add temporary logging:
# Before the problematic operation
import traceback
print(f"DEBUG: input_file={input_file}", file=sys.stderr)
print(f"DEBUG: cwd={os.getcwd()}", file=sys.stderr)
print(f"DEBUG: exists={os.path.exists(input_file)}", file=sys.stderr)
traceback.print_stack(file=sys.stderr)Tips:
- Use
stderr(not logger — may be suppressed) - Log BEFORE the dangerous operation, not after it fails
- Include: paths, working directory, environment variables
traceback.print_stack()shows complete call chain
For Multi-Component Systems
For pipelines (Galaxy workflow → tool execution → data processing):
For EACH component boundary:
- Log what data enters
- Log what data exits
- Verify configuration propagation
- Check state at each layer
Run once → analyze evidence → identify failing component → investigateKey Principle
NEVER fix just where the error appears. Trace back to find the original trigger. Then also consider adding validation at intermediate layers (see defense-in-depth.md).
Adapted from obra/superpowers.