
Agentic Engineering
- 1.4k installs
- 238k repo stars
- Updated August 5, 2026
- affaan-m/ecc
This is a copy of agentic-engineering by affaan-m - installs and ranking accrue to the original listing.
agentic-engineering is an ECC skill that runs engineering workflows where AI agents perform most implementation work while humans maintain strict quality and risk controls through eval-first execution, decomposition, and
About
agentic-engineering is a skill from affaan-m/ecc for operating as an agentic engineer when AI agents handle most coding and humans enforce quality gates. Four operating principles require defining completion criteria before execution, decomposing work into agent-sized units, routing model tiers by task complexity, and measuring outcomes with evals and regression checks. The eval-first loop defines capability evals and regression suites before agents execute changes. Developers reach for agentic-engineering when shifting from manual implementation to supervised agent execution without sacrificing reliability, cost discipline, or verifiable completion criteria across multi-step software tasks.
- 4 operating principles that keep humans in control of quality and risk
- Eval-first loop with baseline capture and regression delta measurement
- 15-minute unit rule for task decomposition into independently verifiable pieces
- Model routing across Haiku, Sonnet, and Opus based on task complexity
- Session strategy rules for continuing vs starting fresh sessions
Agentic Engineering by the numbers
- 1,437 all-time installs (skills.sh)
- +84 installs in the week ending Aug 5, 2026 (Skillselion tracking)
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/affaan-m/ecc --skill agentic-engineeringAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 1.4k |
|---|---|
| repo stars | ★ 238k |
| Last updated | August 5, 2026 |
| Repository | affaan-m/ecc ↗ |
How do you run eval-first agentic engineering workflows?
Run engineering workflows where AI agents perform most implementation work while the human maintains strict quality and risk controls.
Who is it for?
Engineering leads delegating implementation to AI agents who need eval-first quality controls and cost-aware model routing.
Skip if: Solo manual coding sessions without agent delegation, or trivial one-line fixes that do not need decomposition or regression evals.
When should I use this skill?
A team assigns most implementation to AI agents and needs eval-first execution, decomposition, and human quality gates.
What you get
Completion criteria, agent-sized task breakdowns, model routing plans, and eval regression check results
- task decomposition plans
- eval and regression suites
By the numbers
- Documents 4 operating principles for agentic engineering
Files
Agentic Engineering
Use this skill for engineering workflows where AI agents perform most implementation work and humans enforce quality and risk controls.
Operating Principles
1. Define completion criteria before execution. 2. Decompose work into agent-sized units. 3. Route model tiers by task complexity. 4. Measure with evals and regression checks.
Eval-First Loop
1. Define capability eval and regression eval. 2. Run baseline and capture failure signatures. 3. Execute implementation. 4. Re-run evals and compare deltas.
Example workflow:
1. Write test that captures desired behavior (eval)
2. Run test → capture baseline failures
3. Implement feature
4. Re-run test → verify improvements
5. Check for regressions in other testsTask Decomposition
Apply the 15-minute unit rule:
- Each unit should be independently verifiable
- Each unit should have a single dominant risk
- Each unit should expose a clear done condition
Good decomposition:
Task: Add user authentication
├─ Unit 1: Add password hashing (15 min, security risk)
├─ Unit 2: Create login endpoint (15 min, API contract risk)
├─ Unit 3: Add session management (15 min, state risk)
└─ Unit 4: Protect routes with middleware (15 min, auth logic risk)Bad decomposition:
Task: Add user authentication (2 hours, multiple risks)Model Routing
Choose model tier based on task complexity:
- Haiku: Classification, boilerplate transforms, narrow edits
- Example: Rename variable, add type annotation, format code
- Sonnet: Implementation and refactors
- Example: Implement feature, refactor module, write tests
- Opus: Architecture, root-cause analysis, multi-file invariants
- Example: Design system, debug complex issue, review architecture
Cost discipline: Escalate model tier only when lower tier fails with a clear reasoning gap.
Session Strategy
- Continue session for closely-coupled units
- Example: Implementing related functions in same module
- Start fresh session after major phase transitions
- Example: Moving from implementation to testing
- Compact after milestone completion, not during active debugging
- Example: After feature complete, before starting next feature
Review Focus for AI-Generated Code
Prioritize:
- Invariants and edge cases
- Error boundaries
- Security and auth assumptions
- Hidden coupling and rollout risk
Do not waste review cycles on style-only disagreements when automated format/lint already enforce style.
Review checklist:
- [ ] Edge cases handled (null, empty, boundary values)
- [ ] Error handling comprehensive
- [ ] Security assumptions validated
- [ ] No hidden coupling between modules
- [ ] Rollout risk assessed (breaking changes, migrations)
Cost Discipline
Track per task:
- Model tier used
- Token estimate
- Retries needed
- Wall-clock time
- Success/failure outcome
Example tracking:
Task: Implement user login
Model: Sonnet
Tokens: ~5k input, ~2k output
Retries: 1 (initial implementation had auth bug)
Time: 8 minutes
Outcome: SuccessWhen to Use This Skill
- Managing AI-driven development workflows
- Planning agent task decomposition
- Optimizing model tier selection
- Implementing eval-first development
- Reviewing AI-generated code
- Tracking development costs
Integration with Other Skills
- tdd-workflow: Combine with eval-first loop for test-driven development
- verification-loop: Use for continuous validation during implementation
- search-first: Apply before implementation to find existing solutions
- coding-standards: Reference during code review phase
Related skills
How it compares
Pick agentic-engineering over generic pair-programming skills when agents own most implementation and you need eval-first gates plus model routing discipline.
FAQ
What is the eval-first loop in agentic-engineering?
agentic-engineering requires defining capability evals and regression checks before agents execute work, then measuring outputs against those criteria while humans enforce quality and risk controls.
How does agentic-engineering handle model selection?
agentic-engineering routes model tiers by task complexity as one of four operating principles, pairing cheaper models with simpler agent-sized units and stronger models with harder decomposition steps.