
Agent Eval
- 1.4k installs
- 64.6k repo stars
- Updated August 1, 2026
- colbymchenry/codegraph
agent-eval provides documented workflows for Benchmark CodeGraph retrieval quality on a real codebase by comparing agent behavior with vs without CodeGraph. Use when the user runs /agent-eval or asks to te
About
The agent-eval skill benchmark CodeGraph retrieval quality on a real codebase by comparing agent behavior with vs without CodeGraph. Use when the user runs /agent-eval or asks to test, benchmark, audit, or validate a codegraph version (the local dev build or a published npm version) against a language's repo. # CodeGraph Quality Audit Measures how much CodeGraph helps an agent versus plain grep/read, for a chosen codegraph version on a chosen real-world repo. Drives the harness in `scripts/agent-eval/`. ## Prerequisites - `tmux` 3+, a logged-in `claude` CLI, `node`, `git` (macOS/Linux). - Run from the codegraph repo root. ## Workflow Copy this checklist: ``` - [ ] 1. Pick version (local or npm) - [ ] 2. Pick repo by size - [ ] 4. Pick harness (headless / tmux / both) - [ ] 5. Run audit.sh in the background - [ ] 6. Report results ``` **Step 1 - version.** Ask with `AskUserQuestion`: which codegraph version to test. Offer "Local dev build" and "Latest published"; the free-text "Other" lets the user type a specific version (e.g.
- `tmux` 3+, a logged-in `claude` CLI, `node`, `git` (macOS/Linux).
- Run from the codegraph repo root.
- [ ] 1. Pick version (local or npm)
- [ ] 2. Pick language
- [ ] 3. Pick repo by size
Agent Eval by the numbers
- 1,390 all-time installs (skills.sh)
- +38 installs in the week ending Aug 5, 2026 (Skillselion tracking)
- Ranked #226 of 2,715 Automation & Workflows skills by installs in the Skillselion catalog
- Security screen: MEDIUM risk (skills.sh audit)
- Data as of Aug 5, 2026 (Skillselion catalog sync)
agent-eval capabilities & compatibility
- Capabilities
- `tmux` 3+, a logged in `claude` cli, `node`, `gi · run from the codegraph repo root. · [ ] 1. pick version (local or npm) · [ ] 2. pick language · [ ] 3. pick repo by size
- Use cases
- documentation
What agent-eval says it does
# CodeGraph Quality Audit Measures how much CodeGraph helps an agent versus plain grep/read, for a chosen codegraph version on a chosen real-world repo.
Drives the harness in `scripts/agent-eval/`.
npx skills add https://github.com/colbymchenry/codegraph --skill agent-evalAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 1.4k |
|---|---|
| repo stars | ★ 64.6k |
| Security audit | 2 / 3 scanners passed |
| Last updated | August 1, 2026 |
| Repository | colbymchenry/codegraph ↗ |
How do I use agent-eval for the task described in its SKILL.md triggers?
Benchmark CodeGraph retrieval quality on a real codebase by comparing agent behavior with vs without CodeGraph. Use when the user runs /agent-eval or asks to test, benchmark, audit, or validate a cod.
Who is it for?
Teams invoking agent-eval when the user request matches documented triggers and prerequisites.
Skip if: Skip when cached docs are missing, the request is a negative trigger, or another sibling skill owns the workflow.
When should I use this skill?
Benchmark CodeGraph retrieval quality on a real codebase by comparing agent behavior with vs without CodeGraph. Use when the user runs /agent-eval or asks to test, benchmark, audit, or validate a codegraph version (the l
What you get
Step-by-step guidance grounded in agent-eval documentation and reference files.
- With-vs-without CodeGraph benchmark report
- Per-version retrieval quality audit results
By the numbers
- Requires tmux version 3 or newer alongside claude CLI, node, and git
- Uses harness located at scripts/agent-eval/ inside the codegraph repository
Files
CodeGraph Quality Audit
Measures how much CodeGraph helps an agent versus plain grep/read, for a chosen codegraph version on a chosen real-world repo. Drives the harness in scripts/agent-eval/.
Prerequisites
tmux3+, a logged-inclaudeCLI,node,git(macOS/Linux).- Run from the codegraph repo root.
Workflow
Copy this checklist:
- [ ] 1. Pick version (local or npm)
- [ ] 2. Pick language
- [ ] 3. Pick repo by size
- [ ] 4. Pick harness (headless / tmux / both)
- [ ] 5. Run audit.sh in the background
- [ ] 6. Report resultsStep 1 — version. Ask with AskUserQuestion: which codegraph version to test. Offer "Local dev build" and "Latest published"; the free-text "Other" lets the user type a specific version (e.g. 0.7.10). Map the answer to a VERSION token:
- "Local dev build" →
local - "Latest published" →
latest - a typed version → that string (e.g.
0.7.10)
Step 2 — language. Read .claude/skills/agent-eval/corpus.json. Ask with AskUserQuestion which language to test, listing the languages that have entries.
Step 3 — repo. From the chosen language's entries, ask which repo. Label each option with its size and file count, e.g. excalidraw — Medium (~600 files). Each entry carries the repo URL and a representative question.
Step 4 — harness. Ask with AskUserQuestion which harness to run, and map the answer to a MODE token:
- "Headless" →
headless—claude -pwith stream-json: exact tokens/cost and a
clean tool sequence (2 runs, fast, no TTY).
- "Interactive (tmux)" →
tmux— drives the real Claude TUI in tmux: faithful
Explore-subagent behavior, metrics from session logs (2 runs, slower).
- "Both" →
all— headless + interactive (4 runs).
Step 5 — run. Launch in the background (sets the version, clones if missing, wipes + re-indexes, runs the chosen arms — several minutes):
scripts/agent-eval/audit.sh <VERSION> <repo-name> <repo-url> "<question>" <MODE>Step 6 — report. When the job finishes, read the log and report per arm:
- Headless (
parse-run.mjs): total tool calls, fileReads, Grep/Bash,
codegraph-tool calls, duration, total cost.
- Interactive (
parse-session.mjs): the `VERDICT: codegraph_explore used Nx |
Read N | Grep/Bash N and TOKENS:` lines.
Lead with cost + tool/Read counts — they are the reliable signals; raw token in/out are confounded by subagent delegation and prompt caching. State whether codegraph reduced effort and whether both arms reached a correct answer.
Notes
- The index is rebuilt every run (
audit.shwipes.codegraph) — different
versions extract differently, so an index must be served by the same binary that built it.
audit.shtemporarily mutates the globalcodegraphinstall for the test,
then restores your dev link via local-install.sh.
- Corpus repos are cloned to
/tmp/codegraph-corpus(reused if already present). - Add or edit repos in
corpus.json(fields:name,repo,size,files,
question).
{
"_comment": "Test corpus for /agent-eval. Add entries freely. size: Small (<~150 files), Medium (~150-1500), Large (>~1500). 'question' is a representative architectural question that exercises cross-file understanding.",
"TypeScript": [
{
"name": "ky",
"repo": "https://github.com/sindresorhus/ky",
"size": "Small",
"files": "~25",
"question": "How does ky implement request retries and timeouts?"
},
{
"name": "excalidraw",
"repo": "https://github.com/excalidraw/excalidraw",
"size": "Medium",
"files": "~600",
"question": "How does Excalidraw render and update canvas elements?"
},
{
"name": "vscode",
"repo": "https://github.com/microsoft/vscode",
"size": "Large",
"files": "~10000",
"question": "How does the extension host communicate with the main process?"
}
],
"JavaScript": [
{
"name": "express",
"repo": "https://github.com/expressjs/express",
"size": "Small",
"files": "~50",
"question": "How does Express route a request through its middleware stack?"
}
],
"Go": [
{
"name": "cobra",
"repo": "https://github.com/spf13/cobra",
"size": "Small",
"files": "~50",
"question": "How does cobra parse commands and flags?"
},
{
"name": "gin",
"repo": "https://github.com/gin-gonic/gin",
"size": "Medium",
"files": "~150",
"question": "How does gin route requests through its middleware chain?"
},
{
"name": "terraform",
"repo": "https://github.com/hashicorp/terraform",
"size": "Large",
"files": "~4000",
"question": "How does Terraform build and walk the resource dependency graph?"
},
{
"name": "cosmos-sdk",
"repo": "https://github.com/cosmos/cosmos-sdk",
"size": "Large",
"files": "~5000",
"question": "How does a bank module MsgSend message reach the account balance update? Trace the cross-module call path from the bank keeper's Send handler through to the account/balance store update."
}
],
"Python": [
{
"name": "click",
"repo": "https://github.com/pallets/click",
"size": "Small",
"files": "~60",
"question": "How does click parse command-line arguments into commands?"
},
{
"name": "flask",
"repo": "https://github.com/pallets/flask",
"size": "Medium",
"files": "~90",
"question": "How does Flask dispatch a request to a view function?"
},
{
"name": "django",
"repo": "https://github.com/django/django",
"size": "Large",
"files": "~2700",
"question": "How does Django's ORM build and execute a query from a QuerySet?"
}
],
"Rust": [
{
"name": "clap",
"repo": "https://github.com/clap-rs/clap",
"size": "Medium",
"files": "~200",
"question": "How does clap parse arguments against a derived command definition?"
},
{
"name": "tokio",
"repo": "https://github.com/tokio-rs/tokio",
"size": "Large",
"files": "~700",
"question": "How does tokio schedule and run async tasks on its runtime?"
},
{
"name": "deno",
"repo": "https://github.com/denoland/deno",
"size": "Large",
"files": "~1500",
"question": "How does Deno load and execute a TypeScript module?"
}
],
"Java": [
{
"name": "gson",
"repo": "https://github.com/google/gson",
"size": "Medium",
"files": "~200",
"question": "How does Gson serialize an object to JSON?"
},
{
"name": "okhttp",
"repo": "https://github.com/square/okhttp",
"size": "Medium",
"files": "~640",
"question": "How does OkHttp process a request through its interceptor chain?"
},
{
"name": "guava",
"repo": "https://github.com/google/guava",
"size": "Large",
"files": "~3000",
"question": "How does Guava's CacheBuilder build and configure a cache?"
}
],
"Kotlin": [
{
"name": "koin",
"repo": "https://github.com/InsertKoinIO/koin",
"size": "Medium",
"files": "~300",
"question": "How does Koin resolve and inject dependencies?"
},
{
"name": "leakcanary",
"repo": "https://github.com/square/leakcanary",
"size": "Medium",
"files": "~250",
"question": "How does LeakCanary detect and analyze a memory leak?"
}
],
"Swift": [
{
"name": "alamofire",
"repo": "https://github.com/Alamofire/Alamofire",
"size": "Small",
"files": "~100",
"question": "How does Alamofire build, send, and validate a request?"
}
],
"C#": [
{
"name": "serilog",
"repo": "https://github.com/serilog/serilog",
"size": "Medium",
"files": "~250",
"question": "How does Serilog route a log event to its sinks?"
},
{
"name": "jellyfin",
"repo": "https://github.com/jellyfin/jellyfin",
"size": "Large",
"files": "~2500",
"question": "How does Jellyfin scan and identify items in a media library?"
}
],
"Ruby": [
{
"name": "sinatra",
"repo": "https://github.com/sinatra/sinatra",
"size": "Small",
"files": "~60",
"question": "How does Sinatra match a request to a route handler?"
},
{
"name": "discourse",
"repo": "https://github.com/discourse/discourse",
"size": "Large",
"files": "~3000",
"question": "How does Discourse create and render a new post?"
}
],
"PHP": [
{
"name": "slim",
"repo": "https://github.com/slimphp/Slim",
"size": "Small",
"files": "~80",
"question": "How does Slim handle a request through its middleware?"
},
{
"name": "laravel",
"repo": "https://github.com/laravel/framework",
"size": "Large",
"files": "~3000",
"question": "How does Laravel resolve and dispatch a route to a controller?"
}
],
"C": [
{
"name": "redis",
"repo": "https://github.com/redis/redis",
"size": "Large",
"files": "~600",
"question": "How does Redis parse and dispatch a client command?"
}
],
"C++": [
{
"name": "json",
"repo": "https://github.com/nlohmann/json",
"size": "Small",
"files": "~100",
"question": "How does nlohmann::json parse a JSON string into a value?"
},
{
"name": "grpc",
"repo": "https://github.com/grpc/grpc",
"size": "Large",
"files": "~3000",
"question": "How does gRPC dispatch an incoming RPC to its handler?"
}
],
"Dart": [
{
"name": "flutter",
"repo": "https://github.com/flutter/flutter",
"size": "Large",
"files": "~6000",
"question": "How does Flutter build and lay out a widget tree?"
}
],
"Svelte": [
{
"name": "shadcn-svelte",
"repo": "https://github.com/huntabyte/shadcn-svelte",
"size": "Medium",
"files": "~600",
"question": "How do shadcn-svelte components compose and apply their styling?"
}
],
"Lua": [
{
"name": "lualine.nvim",
"repo": "https://github.com/nvim-lualine/lualine.nvim",
"size": "Small",
"files": "~120",
"question": "How does lualine assemble and render its statusline sections and components?"
},
{
"name": "telescope.nvim",
"repo": "https://github.com/nvim-telescope/telescope.nvim",
"size": "Medium",
"files": "~80",
"question": "How does Telescope wire a picker to its finder, sorter, and previewer?"
},
{
"name": "kong",
"repo": "https://github.com/Kong/kong",
"size": "Large",
"files": "~1330",
"question": "How does Kong execute plugins across a request's lifecycle phases?"
}
],
"Luau": [
{
"name": "Knit",
"repo": "https://github.com/Sleitnick/Knit",
"size": "Small",
"files": "~10",
"question": "How does Knit register services and expose them to clients?"
},
{
"name": "vide",
"repo": "https://github.com/centau/vide",
"size": "Small",
"files": "~40",
"question": "How does vide track reactive sources and re-run effects when state changes?"
},
{
"name": "Fusion",
"repo": "https://github.com/dphfox/Fusion",
"size": "Medium",
"files": "~115",
"question": "How does Fusion build and update its reactive UI graph from state objects?"
}
],
"Objective-C": [
{
"name": "Masonry",
"repo": "https://github.com/SnapKit/Masonry",
"size": "Small",
"files": "~50",
"question": "How does Masonry build and activate Auto Layout constraints from its block DSL?"
},
{
"name": "FMDB",
"repo": "https://github.com/ccgus/fmdb",
"size": "Medium",
"files": "~80",
"question": "How does FMDB execute a prepared SQL statement and bind parameters?"
},
{
"name": "SDWebImage",
"repo": "https://github.com/SDWebImage/SDWebImage",
"size": "Large",
"files": "~400",
"question": "How does SDWebImage download, cache, and decode an image for a UIImageView?"
}
],
"Mixed iOS (Swift+ObjC)": [
{
"name": "Charts",
"repo": "https://github.com/danielgindi/Charts",
"size": "Small",
"files": "~270",
"question": "How does the ChartsDemo ObjC demo controller drive the Swift Charts library to animate and notify a data update?"
},
{
"name": "realm-swift",
"repo": "https://github.com/realm/realm-swift",
"size": "Medium",
"files": "~370",
"question": "How does a Swift `Realm.write { realm.add(obj) }` reach the Objective-C persistence layer?"
},
{
"name": "wikipedia-ios",
"repo": "https://github.com/wikimedia/wikipedia-ios",
"size": "Large",
"files": "~1700",
"question": "How does tapping a search result reach the article-fetch network call across the Swift / ObjC boundary?"
}
],
"React Native (legacy bridge + TurboModule)": [
{
"name": "@react-native-async-storage",
"repo": "https://github.com/react-native-async-storage/async-storage",
"size": "Small",
"files": "~60",
"question": "How does `setItem` in JS reach the native `legacy_multiSet` implementation?"
},
{
"name": "react-native-svg",
"repo": "https://github.com/software-mansion/react-native-svg",
"size": "Medium",
"files": "~700",
"question": "How does a JS `Svg.getTotalLength(...)` reach the iOS / Android native implementation via TurboModule?"
},
{
"name": "react-native-firebase",
"repo": "https://github.com/invertase/react-native-firebase",
"size": "Large",
"files": "~1100",
"question": "How does a native iOS push notification reach the JS `messaging().onMessage(...)` listener?"
}
],
"Expo Modules": [
{
"name": "expo-haptics",
"repo": "https://github.com/expo/expo/tree/main/packages/expo-haptics",
"size": "Small",
"files": "~15",
"question": "How does `Haptics.notificationAsync(...)` in JS reach `UINotificationFeedbackGenerator` in the Swift Module?"
},
{
"name": "expo-camera",
"repo": "https://github.com/expo/expo/tree/main/packages/expo-camera",
"size": "Medium",
"files": "~70",
"question": "How does a JS `CameraView.takePictureAsync(options)` reach the native AVCaptureSession / CameraDevice call?"
}
],
"React Native Fabric (view components)": [
{
"name": "react-native-segmented-control",
"repo": "https://github.com/react-native-segmented-control/segmented-control",
"size": "Small",
"files": "~25",
"question": "How does JSX `<SegmentedControl onChange={cb}/>` reach the native onChange handler on iOS/Android?"
},
{
"name": "react-native-screens",
"repo": "https://github.com/software-mansion/react-native-screens",
"size": "Medium",
"files": "~1200",
"question": "How does JSX `<ScreenStack>` reach the native RNSScreenStackView component?"
},
{
"name": "react-native-skia",
"repo": "https://github.com/Shopify/react-native-skia",
"size": "Large",
"files": "~1000",
"question": "How does a `<SkiaPictureView/>` JSX usage reach the iOS / Android native renderer?"
}
],
"R": [
{
"name": "AnomalyDetection",
"repo": "https://github.com/twitter/AnomalyDetection",
"size": "Small",
"files": "~24",
"question": "How does AnomalyDetectionTs go from the exported entry function to the underlying S-H-ESD statistical test? Name the functions on the path in order."
},
{
"name": "dplyr",
"repo": "https://github.com/tidyverse/dplyr",
"size": "Medium",
"files": "~450",
"question": "When mutate() is called on a grouped data frame, which functions handle the grouping and expression evaluation, in order, from mutate() down?"
},
{
"name": "ggplot2",
"repo": "https://github.com/tidyverse/ggplot2",
"size": "Large",
"files": "~1150",
"question": "When a ggplot object is printed, how does the plot actually get built and drawn \u2014 trace the path from print/plot to where geoms render. Name the key functions in order."
}
]
}Related skills
How it compares
Pick agent-eval when you need side-by-side CodeGraph versus baseline agent measurements on a real repo instead of synthetic unit tests alone.
FAQ
What does agent-eval do?
Benchmark CodeGraph retrieval quality on a real codebase by comparing agent behavior with vs without CodeGraph. Use when the user runs /agent-eval or asks to test, benchmark, audit, or validate a codegraph version (the l
When should I use agent-eval?
Benchmark CodeGraph retrieval quality on a real codebase by comparing agent behavior with vs without CodeGraph. Use when the user runs /agent-eval or asks to test, benchmark, audit, or validate a codegraph version (the l
What are common prerequisites?
--- name: agent-eval description: Benchmark CodeGraph retrieval quality on a real codebase by comparing agent behavior with vs without CodeGraph.
Is Agent Eval safe to install?
skills.sh reports 2 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.