
Engineering Manager Agent Prompts Evals
- 27 installs
- 7 repo stars
- Updated May 20, 2026
- daemon-blockint-tech/agentic-enteprises-skill
Guides engineering managers leading prompt and eval teams: org design, roadmap prioritization, prompt release governance, hiring, and eval-health metrics.
About
Guides engineering managers who lead teams owning agent prompts, golden eval suites, judge programs, and prompt regression CI, covering org design, release governance, and team KPIs. A manager uses it when staffing eval work, prioritizing eval debt, or governing prompt releases.
- Prompt semver gates with waivers and risk sign-off for customer-facing agents
- Eval-health KPIs: pass rate, slice regression, judge drift
Engineering Manager Agent Prompts Evals by the numbers
- 27 all-time installs (skills.sh)
- Ranked #9,601 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
- Data as of Jul 29, 2026 (Skillselion catalog sync)
npx skills add https://github.com/daemon-blockint-tech/agentic-enteprises-skill --skill engineering-manager-agent-prompts-evalsAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 27 |
|---|---|
| repo stars | ★ 7 |
| Last updated | May 20, 2026 |
| Repository | daemon-blockint-tech/agentic-enteprises-skill ↗ |
What it does
Guides engineering managers leading prompt and eval teams: org design, roadmap prioritization, prompt release governance, hiring, and eval-health metrics.
Files
Engineering Manager, Agent Prompts & Evals
When to Use
- Build or scale a prompt + eval engineering function (central or embedded)
- Prioritize golden set, harness, and judge work with PM and risk
- Define release policy for prompt/tool changes (gates, waivers, rollback)
- Hire and level prompt engineers, eval engineers, and tech leads
- Resolve capacity conflicts between new agents and eval debt
- Report eval health (pass rate, slice regressions, judge drift) to leadership
- Partner with
ai-lead-opson production incidents tied to prompts
When NOT to Use
- Author system prompts, tool schemas, or eval cases →
prompt-engineer-agent-prompts-evals - General prompt patterns (CoT, few-shot) →
prompt-engineer - Full RAG/agent application code →
ai-engineer - Vertical AI product roadmap and GTM launches →
engineering-manager-vertical-ai-products - Adversarial red-team execution →
ai-redteam - Org-wide data or analytics management →
data-manager
Related skills
| Need | Skill |
|---|---|
| IC prompt/eval implementation | prompt-engineer-agent-prompts-evals |
| Vertical AI EM (broader) | engineering-manager-vertical-ai-products |
| AI ops and rollouts | ai-lead-ops |
| Risk tier and policy | ai-risk-governance |
| Token cost programs | ai-token-improvement-plan-engineer |
Core Workflows
1. Org design
Central eval platform vs embedded prompt owners; ratios and interfaces.
See `references/team_org_prompt_eval.md`.
2. Roadmap and prioritization
Eval debt, new agent coverage, judge calibration, CI investment.
See `references/roadmap_prioritization.md`.
3. Release governance
Prompt semver gates, waivers, coordination with risk and ops.
See `references/release_governance.md`.
4. Stakeholder partnerships
PM, applied architect, risk, red-team, platform eng.
See `references/stakeholder_partnerships.md`.
5. Hiring and development
Levels for prompt/eval specialists.
See `references/hiring_development.md`.
6. Team metrics
Pass rate, coverage, time-to-golden-case, incident linkage.
See `references/team_metrics_accountability.md`.
Output standards
- Roadmap items name eval slice, prompt surface, and risk tier
- No prompt prod change without documented baseline vs candidate
- Waivers require risk sign-off for Tier 1–2 customer-facing agents
- Escalations include trade-offs (scope, date, headcount)
When to load references
- Org →
references/team_org_prompt_eval.md - Roadmap →
references/roadmap_prioritization.md - Release →
references/release_governance.md - Stakeholders →
references/stakeholder_partnerships.md - People →
references/hiring_development.md - KPIs →
references/team_metrics_accountability.md
Hiring and Development
Profiles
| Role | Focus |
|---|---|
| Agent prompt engineer | System prompts, tools, handoffs |
| Eval engineer | Harness, datasets, CI |
| Generalist | Both on smaller teams |
Interview depth: prompt-engineer-agent-prompts-evals references.
Interview loop
| Stage | Assesses |
|---|---|
| Prompt + tool design | Given agent scenario |
| Eval design | Write 5 cases + assertions |
| Debug trace | Wrong-tool root cause |
| Behavioral | Pushback on ship without eval |
| Manager | Scope, stakeholder alignment |
Levels (summary)
| Level | Bar |
|---|---|
| IC3 | Contributes prompts/cases with review |
| IC4 | Owns one agent eval slice |
| IC5 | Owns harness standards + cross-agent patterns |
| Lead | Technical bar, review, incident commander |
| EM | Roadmap, hiring, governance, predictability |
Onboarding
| Milestone | Goal |
|---|---|
| 30d | Merge prompt or eval PR; run harness locally |
| 60d | Own a tag slice |
| 90d | Lead one release gate cycle |
Growth paths
- Deep specialist (judges, safety goldens)
- Platform eval infra
- EM or vertical AI EM pivot
Retention
- Chronic ship-without-eval culture
- Flaky CI ignored
- No path to impact beyond firefighting
Release Governance (Prompt & Eval)
Gate stack (customer-facing agents)
| Gate | Owner | Block? |
|---|---|---|
| Golden CI pass | Eval eng | Yes |
| Tag slice thresholds (safety, tools) | EM + risk | Yes |
| Judge calibration current | Judge owner | Yes if stale |
| Red-team sign-off (Tier 1–2) | ai-redteam | Yes |
| Ops kill switch tested | ai-lead-ops | Yes |
| Prompt semver + changelog | Prompt eng | Yes |
IC checklist detail: prompt-engineer-agent-prompts-evals → references/prompt_versioning_regression.md.
Waiver process
1. Document failing cases and business justification 2. Compensating control (human review, feature flag, limited audience) 3. Risk approver for tier 4. Expiry date and remediation ticket
Log all waivers; review monthly in leadership sync.
Rollback
- N-1 prompt in config; flag per version
- Eval must pass on N-1 before rollback deploy
- Post-rollback: add regression cases for failure mode
Internal / low-tier agents
Reduced gates — still require smoke golden set and owner sign-off.
Manager responsibilities
- Enforce no hotfix prompt without trace + case addition
- Align with vertical EM (
engineering-manager-vertical-ai-products) on launch calendar - Communicate eval status in ship/no-ship meetings
Roadmap and Prioritization
Backlog item template
| Field | Example |
|---|---|
| Agent / surface | Support copilot v2 |
| Outcome | Reduce wrong-tool rate 50% |
| Deliverable | 40 new goldens + tool schema rewrite |
| Risk tier | Tier 2 |
| Deps | Platform trace API v2 |
Priority buckets
| Bucket | Examples |
|---|---|
| P0 — production regression | Pass rate drop on safety/refusal slice |
| P1 — launch blocker | GA requires green golden set |
| P2 — coverage | New tool, new locale, new intent |
| P3 — efficiency | Harness speed, judge cost reduction |
Reserve 25–30% capacity for P0/P1 buffer and harness maintenance.
Trade-offs (manager decisions)
| Tension | Resolution framework |
|---|---|
| New agent vs eval debt | No new agent without minimum golden coverage bar |
| More tools vs eval combinatorics | Cap tools per agent; platform review |
| Judge quality vs cost | Calibrate quarterly; sample in prod |
| Speed vs human labels | SME budget upfront for Tier 1 |
Metrics for roadmap reviews
- Pass rate trend by tag (not aggregate only)
- Time-to-add-golden after production failure
- Open waiver count
- Judge–human agreement drift
Escalation
Escalate when: launch date conflicts with eval gate; risk refuses waiver without staffing; platform blocks harness access.
Bring options: reduce scope, delay GA, temporary human review layer.
Stakeholder Partnerships
Product management
| You provide | You need |
|---|---|
| Eval-based ship recommendations | Prioritized agents and dates |
| Clear limitations for GTM | Early prompt/tool scope lock |
| Pass-rate dashboards by slice | Acceptance of eval-driven delays |
Applied AI architect
- Escalate new tool surfaces and multi-agent handoffs
- Do not bypass ADRs for tenant/data boundaries
Risk and compliance
- Tier assignment before roadmap commit
- Waiver path for time pressure
- Policy updates → new refusal goldens within SLA
Red team
- Schedule before Tier 1–2 GA
- Convert findings to regression tags in golden set
- Not a substitute for continuous CI
AI lead ops
- Incidents caused by prompts: joint post-mortem
- Production judge sampling alignment
- Model version changes → re-baseline plan
IC team (prompt-engineer-agent-prompts-evals)
EM protects eval debt time; ICs own technical depth.
Status rhythm
| Forum | Frequency |
|---|---|
| Eval health review | Weekly |
| Launch readiness | Per milestone |
| Judge calibration | Monthly |
Anti-pattern: EM rewriting prompts in prod without IC review.
Team Metrics and Accountability
Scorecard
| Metric | Notes |
|---|---|
| Overall golden pass rate | Trend; not sole metric |
| Pass rate by tag | safety, tools, refusal, domain |
| Coverage | % prod failure modes with goldens |
| Time-to-golden | Hours/days from incident → case |
| Flaky case rate | Quarantined / total |
| Judge–human agreement | Calibration health |
| Prompt releases with full gate | % without waiver |
| P1 prompt incidents | Count and recurrence |
Targets (set per org)
- No drop on safety/refusal slice without waiver
- Time-to-golden < 5 business days for Tier 1–2
- Eval debt burn: N cases/quarter per agent
Accountability
| Event | EM action |
|---|---|
| Pass rate drop on merge | Block release; root cause |
| Repeat wrong-tool class | Platform or prompt initiative |
| Waiver spike | Review with risk; staffing plan |
| Judge drift | Pause judge-gated deploys; recalibrate |
Reporting
Weekly: slice heatmap, blockers, waivers open
Quarterly: coverage map per agent, harness investment ROI (fewer incidents, faster launches)
vs ai-lead-ops
Ops owns SLOs and incidents; this function owns preventive eval quality and prompt change discipline.
Team Org — Prompt & Eval
Functions to separate
| Function | Owns |
|---|---|
| Prompt engineering | System/developer prompts, tool schemas |
| Eval engineering | Harness, golden sets, CI gates |
| Judge program | Rubrics, human calibration |
| Agent platform | Runtime, tracing — usually ai-engineer / platform |
EM may own first three; partner for platform.
Models
| Model | When |
|---|---|
| Central prompt+eval guild | Many agents; shared standards |
| Embedded in product squads | Few agents; fast iteration |
| Hybrid | Guild sets harness + rubrics; squads own domain goldens |
Avoid: every squad forks harness; no shared pass-rate dashboard.
Interfaces
| Partner | Cadence | Topics |
|---|---|---|
| Product | Weekly | Agent backlog, launch dates |
ai-lead-ops | Bi-weekly | Incidents, rollout tiers |
ai-risk-governance | Per release | Waivers, tier |
applied-ai-architect-commercial-enterprise | As needed | Tool/prompt architecture |
ai-redteam | Per major agent | Findings → regression cases |
Staffing
- 1 eval engineer per 2–3 active agents in flight (rule of thumb)
- Dedicated judge calibration owner when >1 LLM judge in prod
- Tech lead when golden set >500 cases or multi-repo harness
Anti-patterns
- PM owns golden set without engineering versioning
- No on-call rotation for prompt-related prod regressions
- Team only reacts to fires — no eval debt budget