Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
trailofbits avatar

Graph Evolution

  • 2.6k installs
  • 6.4k repo stars
  • Updated August 4, 2026
  • trailofbits/skills

graph-evolution builds Trailmark graphs at two snapshots and reports security-focused structural diffs beyond text patches.

About

The graph-evolution skill compares Trailmark code graphs at two source snapshots to find security-relevant structural changes text diffs miss. Use cases include new attack paths, complexity shifts, blast radius growth, taint propagation changes, and privilege boundary modifications across commits, tags, or directories. Prerequisites require trailmark installed via uv pip install; manual source comparison is forbidden when the tool fails. Phase 1 creates git worktrees for before and after refs. Phase 2 builds graphs with QueryEngine.from_directory, runs engine.preanalysis for blast radius and taint data, and exports JSON summaries. Phase 3 runs both trailmark diff --json and the plugin graph_diff.py helper for subgraph membership changes; stop if either writes empty JSON. Phase 4 interprets nodes added removed modified, edges, entrypoints, and subgraph deltas into a security-focused markdown report. Phase 5 cleans worktrees. Reject rationalizations to skip preanalysis or rely on text diff alone. Related skills differential-review handles line diffs, trailmark handles single snapshots, diagramming-code handles diagrams, and genotoxic handles mutation triage.

  • Compare two git refs or directories with Trailmark structural diff.
  • engine.preanalysis required on both snapshots before diffing.
  • Run trailmark diff --json and graph_diff.py subgraph helper.
  • Surfaces attack paths, blast radius, taint, and privilege boundary shifts.
  • Five-phase workflow: worktrees, build graphs, diff, report, cleanup.

Graph Evolution by the numbers

  • 2,630 all-time installs (skills.sh)
  • +117 installs in the week ending Aug 5, 2026 (Skillselion tracking)
  • Ranked #198 of 2,203 Security skills by installs in the Skillselion catalog
  • Security screen: LOW risk (skills.sh audit)
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
At a glance

graph-evolution capabilities & compatibility

Capabilities
git worktree snapshot isolation · trailmark preanalysis and json export · native and subgraph structural diff merging · security focused markdown report generation
Use cases
security audit · code review · research
Platforms
macOS · Linux
Runs
Runs locally
Pricing
Free
From the docs

What graph-evolution says it does

Surfaces security-relevant changes that text-level diffs miss
SKILL.md
Without pre-analysis, you miss taint changes, blast radius growth, and privilege boundary shifts
SKILL.md
npx skills add https://github.com/trailofbits/skills --skill graph-evolution

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs2.6k
repo stars6.4k
Security audit3 / 3 scanners passed
Last updatedAugust 4, 2026
Repositorytrailofbits/skills

What security-relevant structural changes occurred between these two commits or tags?

Compare Trailmark code graphs between git refs to surface security-relevant structural diffs beyond text patches.

Who is it for?

Security reviews comparing audit snapshots, release tags, or branch heads in codebases with trailmark installed.

Skip if: Skip for line-level review only; use differential-review or single-snapshot trailmark instead.

When should I use this skill?

User compares commits for attack surface growth, structural evolution, or security changes text diffs miss.

What you get

A markdown report from native and subgraph diffs highlighting attack paths, taint, and boundary shifts.

  • structural security diff report
  • classified graph evolution metrics

By the numbers

  • Includes graph_diff.py and 2 reference docs for metrics and reports
  • Requires two snapshot graphs with preanalysis on each

Files

SKILL.mdMarkdownGitHub ↗

Graph Evolution

Builds Trailmark code graphs at two source snapshots and computes a structural diff. Surfaces security-relevant changes that text-level diffs miss: new attack paths, complexity shifts, blast radius growth, taint propagation changes, and privilege boundary modifications.

When to Use

  • Comparing two git refs to understand what structurally changed
  • Auditing a range of commits for security-relevant evolution
  • Detecting new attack paths created by code changes
  • Finding functions whose blast radius or complexity grew silently
  • Identifying taint propagation changes across refactors
  • Pre-release structural comparison (tag-to-tag or branch-to-branch)

When NOT to Use

  • Line-level code review (use differential-review for text-diff analysis)
  • Single-snapshot analysis (use the trailmark skill directly)
  • Diagram generation from a single snapshot (use the diagramming-code skill)
  • Mutation testing triage (use the genotoxic skill)

Rationalizations to Reject

RationalizationWhy It's WrongRequired Action
"We just need the structural diff, skip pre-analysis"Without pre-analysis, you miss taint changes, blast radius growth, and privilege boundary shiftsRun engine.preanalysis() on both snapshots
"Text diff covers what changed"Text diffs miss new attack paths, transitive complexity shifts, and subgraph membership changesUse structural diff to complement text diff
"Only added nodes matter"Removed security functions and shifted privilege boundaries are equally dangerousReview removals and modifications, not just additions
"Low-severity structural changes can be ignored"INFO-level changes (dead code removal) can mask removed security checksClassify every change, review removals for replaced functionality
"One snapshot's graph is enough for comparison"Single-snapshot analysis can't detect evolution — you need both before and afterAlways build and export both graphs
"Tool isn't installed, I'll compare manually"Manual comparison misses what graph analysis catchesInstall trailmark first

---

Prerequisites

trailmark must be installed. If uv run trailmark fails, run:

uv pip install trailmark

DO NOT fall back to "manual comparison" or reading source files as a substitute for running trailmark. The tool must be installed and used programmatically. If installation fails, report the error.

---

Quick Start

# Compare two git refs (e.g., tags, branches, commits)
# 1. Build graphs at each snapshot
# 2. Run pre-analysis on both
# 3. Compute structural diff
# 4. Generate report

# Step-by-step: see Workflow below

---

Decision Tree

├─ Need to understand what each metric means?
│  └─ Read: references/evolution-metrics.md
│
├─ Need the report output format?
│  └─ Read: references/report-format.md
│
├─ Already have two graph JSON exports?
│  └─ Jump to Phase 3 (run native diff + graph_diff.py)
│
└─ Starting from two git refs?
   └─ Start at Phase 1

---

Workflow

Graph Evolution Progress:
- [ ] Phase 1: Create snapshots (git worktrees)
- [ ] Phase 2: Build graphs + pre-analysis on both snapshots
- [ ] Phase 3: Compute structural diff
- [ ] Phase 4: Interpret diff and generate report
- [ ] Phase 5: Clean up worktrees

Phase 1: Create Snapshots

Use git worktrees to get clean copies of each ref without disturbing the working tree.

# Create temp directories for worktrees
BEFORE_DIR=$(mktemp -d)
AFTER_DIR=$(mktemp -d)

# Create worktrees (run from repo root)
git worktree add "$BEFORE_DIR" {before_ref}
git worktree add "$AFTER_DIR" {after_ref}

If comparing two directories instead of git refs, skip this phase and use the directory paths directly in Phase 2.

Phase 2: Build Graphs and Run Pre-Analysis

Build Trailmark graphs for both snapshots and run pre-analysis on each. Pre-analysis computes blast radius, taint propagation, privilege boundaries, and entrypoint enumeration.

from trailmark.query.api import QueryEngine

def build_and_export(target_dir, output_path, language="auto"):
    """Build graph, run pre-analysis, export JSON."""
    engine = QueryEngine.from_directory(target_dir, language=language)
    engine.preanalysis()
    json_str = engine.to_json()
    with open(output_path, "w") as f:
        f.write(json_str)
    return engine.summary()

import tempfile, os
work_dir = tempfile.mkdtemp(prefix="trailmark_evolution_")
before_json = os.path.join(work_dir, "before_graph.json")
after_json = os.path.join(work_dir, "after_graph.json")

before_summary = build_and_export(
    "{before_dir}", before_json
)
after_summary = build_and_export(
    "{after_dir}", after_json
)

Verify both graphs built successfully by checking the summary output. If either fails, rerun with an explicit language or comma-separated list instead of auto.

Phase 3: Compute Structural Diff

Run both:

1. Trailmark's native structural diff for nodes, edges, and entrypoints 2. The plugin's graph_diff.py helper for subgraph membership changes

Using the same work_dir from Phase 2:

trailmark diff --json "{before_dir}" "{after_dir}" > "{work_dir}/trailmark_diff.json" || \
  uv run trailmark diff --json "{before_dir}" "{after_dir}" > "{work_dir}/trailmark_diff.json"

uv run {baseDir}/scripts/graph_diff.py \
    --before "{before_json}" \
    --after "{after_json}" > "{work_dir}/subgraph_diff.json"

If either diff command fails or writes an empty JSON file, stop and report the error instead of continuing to Phase 4.

The native Trailmark diff contains:

KeyContents
summary_deltaChanges in node/edge/entrypoint counts
nodes.addedNew functions, classes, methods
nodes.removedDeleted functions, classes, methods
nodes.modifiedFunctions with changed CC, params, line span
edges.addedNew call/inheritance/import relationships
edges.removedDeleted relationships
entrypointsAdded, removed, and modified entrypoints

The subgraph diff contains:

KeyContents
subgraphsPer-subgraph membership changes (tainted, high_blast_radius, etc.)

Phase 4: Interpret Diff and Generate Report

Read both diff JSON files and generate a security-focused markdown report. See references/report-format.md for the full template.

Interpretation priorities (highest to lowest):

1. New tainted paths — nodes entering the tainted subgraph, especially if they also appear in added edges targeting sensitive functions 2. Privilege boundary changes — new or removed trust transitions from the native entrypoint/edge diff plus the subgraph diff 3. Attack surface growth — new entrypoints, especially untrusted_external, from trailmark_diff.json 4. Blast radius increases — nodes entering high_blast_radius 5. Complexity spikes — CC increases > 3 on tainted or entrypoint-reachable nodes 6. Structural additions — new nodes and edges (review needed) 7. Structural removals — verify removed security functions were replaced

Cross-reference structural changes with git diff {before_ref}..{after_ref} to add source-level context to findings.

Severity classification:

SeverityStructural Signal
CRITICALNew tainted path to sensitive function, removed auth boundary
HIGHNew entrypoint + high blast radius, large CC increase on tainted node
MEDIUMNew trust-boundary-crossing edges, moderate CC increase
LOWAdded nodes without entrypoint reachability
INFODead code removal, complexity reductions

For detailed metric definitions, see references/evolution-metrics.md.

Phase 5: Clean Up

Remove git worktrees after the report is written:

git worktree remove "{before_dir}"
git worktree remove "{after_dir}"

---

Diff Reference

trailmark diff --json BEFORE AFTER
uv run {baseDir}/scripts/graph_diff.py [OPTIONS]

Use trailmark diff for:

  • Node/edge changes
  • Added/removed/modified entrypoints
  • Human-readable structural diff reports

Use graph_diff.py for:

  • Subgraph membership changes derived from engine.preanalysis()
  • tainted, high_blast_radius, privilege_boundary, and related sets
ArgumentDefaultDescription
--beforerequiredPath to the "before" graph JSON
--afterrequiredPath to the "after" graph JSON
--indent2JSON output indentation

graph_diff.py input format: Trailmark JSON exports from engine.to_json(). graph_diff.py output: JSON structural diff for nodes, edges, and subgraphs.

---

Quality Checklist

Before delivering the report:

  • [ ] Both graphs built successfully (check summaries)
  • [ ] Pre-analysis ran on both snapshots
  • [ ] Native Trailmark diff computed and non-empty (trailmark_diff.json)
  • [ ] Subgraph diff computed and non-empty (subgraph_diff.json)
  • [ ] All subgraph changes interpreted (tainted, blast radius, etc.)
  • [ ] Critical findings include evidence (node IDs, edge diffs)
  • [ ] Severity levels assigned to all findings
  • [ ] Source-level context added via git diff cross-reference
  • [ ] Worktrees cleaned up (or temp dirs removed)
  • [ ] Report written to GRAPH_EVOLUTION_*.md

---

Integration

trailmark skill: Phase 2 uses the trailmark API for graph building and pre-analysis. All trailmark query patterns work on either snapshot's engine.

differential-review skill: Use graph-evolution for structural analysis, differential-review for line-level code review. The two are complementary — graph-evolution finds attack paths that text diffs miss, while differential-review provides git blame context and micro-adversarial analysis.

genotoxic skill: If graph-evolution reveals new high-CC tainted nodes, feed them to genotoxic for mutation testing triage.

diagramming-code skill: Generate before/after diagrams to visualize structural changes. Use call-graph or data-flow diagrams focused on changed nodes.

---

Supporting Documentation

  • [references/evolution-metrics.md](references/evolution-metrics.md)

What each structural metric means and why it matters for security

  • [references/report-format.md](references/report-format.md)

Report template, severity classification, and example findings

Related skills

How it compares

Use graph-evolution alongside text diffs when security-relevant structure—not just lines—may have shifted between releases.

FAQ

Can I skip preanalysis on one snapshot?

No. Run engine.preanalysis on both snapshots or you miss taint and blast radius changes.

What if trailmark is not installed?

Install with uv pip install trailmark; do not fall back to manual source comparison.

Which diff outputs must I read?

Both trailmark diff JSON and graph_diff.py subgraph JSON before generating the report.

Is Graph Evolution safe to install?

skills.sh reports 3 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.

Securityauditappsec

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.