Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
affaan-m avatar

Agentic Engineering

  • 6k installs
  • 238k repo stars
  • Updated August 5, 2026
  • affaan-m/everything-claude-code

agentic-engineering is an agent skill for >

About

name agentic-engineering description Operate as an agentic engineer using eval-first execution decomposition and cost-aware model routing Use when AI agents perform most implementation work and humans enforce quality and risk controls metadata origin ECC Agentic Engineering Use this skill for engineering workflows where AI agents perform most implementation work and humans enforce quality and risk controls Define completion criteria before execution Decompose work into agent-sized units Route model tiers by task complexity Measure with evals and regression checks Define capability eval and regression eval Run baseline and capture failure signatures Re-run evals and compare deltas Write test that captures desired behavior eval 2 Run test capture baseline failures 3 name agentic-engineering description Operate as an agentic engineer using eval-first execution decomposition and cost-aware model routing Use when AI agents perform most implementation work and humans enforce quality and risk controls metadata origin ECC Agentic Engineering Use this skill for engineering workflows where AI agents perform most implementation work and humans enforce quality and risk controls Define complet.

  • Define completion criteria before execution.
  • Decompose work into agent-sized units.
  • Route model tiers by task complexity.
  • Measure with evals and regression checks.
  • Define capability eval and regression eval.

Agentic Engineering by the numbers

  • 6,002 all-time installs (skills.sh)
  • +227 installs in the week ending Aug 5, 2026 (Skillselion tracking)
  • Ranked #259 of 2,153 Testing & QA skills by installs in the Skillselion catalog
  • Security screen: MEDIUM risk (skills.sh audit)
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
At a glance

agentic-engineering capabilities & compatibility

Capabilities
define completion criteria before execution. · decompose work into agent sized units. · route model tiers by task complexity. · measure with evals and regression checks. · define capability eval and regression eval.
Use cases
documentation
From the docs

What agentic-engineering says it does

--- name: agentic-engineering description: > Operate as an agentic engineer using eval-first execution, decomposition, and cost-aware model routing.
SKILL.md
Use when AI agents perform most implementation work and humans enforce quality and risk controls.
SKILL.md
metadata: origin: ECC --- # Agentic Engineering Use this skill for engineering workflows where AI agents perform most implementation work and humans enforce quality and risk controls.
SKILL.md
npx skills add https://github.com/affaan-m/everything-claude-code --skill agentic-engineering

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs6k
repo stars238k
Security audit3 / 3 scanners passed
Last updatedAugust 5, 2026
Repositoryaffaan-m/everything-claude-code

When should developers use agentic-engineering and what problem does it solve?

>

Who is it for?

Developers working with agentic-engineering patterns described in the skill documentation.

Skip if: Skip when cached docs are empty or the task is outside the skill's documented scope.

When should I use this skill?

>

What you get

Grounded guidance and workflows from SKILL.md for agentic-engineering.

  • Eval definitions
  • Task decomposition plan
  • Regression comparison report

By the numbers

  • Follows a four-step eval-first loop: define evals, baseline, implement, re-evaluate
  • Enforces four operating principles for agentic engineering discipline

Files

SKILL.mdMarkdownGitHub ↗

Agentic Engineering

Use this skill for engineering workflows where AI agents perform most implementation work and humans enforce quality and risk controls.

Operating Principles

1. Define completion criteria before execution. 2. Decompose work into agent-sized units. 3. Route model tiers by task complexity. 4. Measure with evals and regression checks.

Eval-First Loop

1. Define capability eval and regression eval. 2. Run baseline and capture failure signatures. 3. Execute implementation. 4. Re-run evals and compare deltas.

Example workflow:

1. Write test that captures desired behavior (eval)
2. Run test → capture baseline failures
3. Implement feature
4. Re-run test → verify improvements
5. Check for regressions in other tests

Task Decomposition

Apply the 15-minute unit rule:

  • Each unit should be independently verifiable
  • Each unit should have a single dominant risk
  • Each unit should expose a clear done condition

Good decomposition:

Task: Add user authentication
├─ Unit 1: Add password hashing (15 min, security risk)
├─ Unit 2: Create login endpoint (15 min, API contract risk)
├─ Unit 3: Add session management (15 min, state risk)
└─ Unit 4: Protect routes with middleware (15 min, auth logic risk)

Bad decomposition:

Task: Add user authentication (2 hours, multiple risks)

Model Routing

Choose model tier based on task complexity:

  • Haiku: Classification, boilerplate transforms, narrow edits
  • Example: Rename variable, add type annotation, format code
  • Sonnet: Implementation and refactors
  • Example: Implement feature, refactor module, write tests
  • Opus: Architecture, root-cause analysis, multi-file invariants
  • Example: Design system, debug complex issue, review architecture

Cost discipline: Escalate model tier only when lower tier fails with a clear reasoning gap.

Session Strategy

  • Continue session for closely-coupled units
  • Example: Implementing related functions in same module
  • Start fresh session after major phase transitions
  • Example: Moving from implementation to testing
  • Compact after milestone completion, not during active debugging
  • Example: After feature complete, before starting next feature

Review Focus for AI-Generated Code

Prioritize:

  • Invariants and edge cases
  • Error boundaries
  • Security and auth assumptions
  • Hidden coupling and rollout risk

Do not waste review cycles on style-only disagreements when automated format/lint already enforce style.

Review checklist:

  • [ ] Edge cases handled (null, empty, boundary values)
  • [ ] Error handling comprehensive
  • [ ] Security assumptions validated
  • [ ] No hidden coupling between modules
  • [ ] Rollout risk assessed (breaking changes, migrations)

Cost Discipline

Track per task:

  • Model tier used
  • Token estimate
  • Retries needed
  • Wall-clock time
  • Success/failure outcome

Example tracking:

Task: Implement user login
Model: Sonnet
Tokens: ~5k input, ~2k output
Retries: 1 (initial implementation had auth bug)
Time: 8 minutes
Outcome: Success

When to Use This Skill

  • Managing AI-driven development workflows
  • Planning agent task decomposition
  • Optimizing model tier selection
  • Implementing eval-first development
  • Reviewing AI-generated code
  • Tracking development costs

Integration with Other Skills

  • tdd-workflow: Combine with eval-first loop for test-driven development
  • verification-loop: Use for continuous validation during implementation
  • search-first: Apply before implementation to find existing solutions
  • coding-standards: Reference during code review phase

Related skills

Forks & variants (1)

Agentic Engineering has 1 known copy in the catalog totaling 1.4k installs. They canonicalize to this original listing.

How it compares

Pick agentic-engineering over ad-hoc agent prompting when shipped code must pass defined evals and regression checks, not just look correct.

FAQ

What does agentic-engineering do?

>

When should I invoke agentic-engineering?

>

Where is the source documentation?

Ground claims in SKILL.md excerpts and linked reference files from the cached docs.

Is Agentic Engineering safe to install?

skills.sh reports 3 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.