
Ai Engineer
- 30 installs
- 7 repo stars
- Updated May 20, 2026
- daemon-blockint-tech/agentic-enteprises-skill
Design and ship production LLM features like RAG pipelines, agent workflows, and eval harnesses with cost, latency, and safe-deploy guardrails.
About
Guides production AI engineering across LLM apps, RAG pipelines, agents, eval harnesses, and safe deployment. A developer uses it when building chatbots, copilots, retrieval systems, or integrating OpenAI/Anthropic/local models into a product.
- RAG pipeline checklist: chunk, embed, index, retrieve, rerank, generate, cite
- Pre-launch eval layers for retrieval, generation, safety, and ops
Ai Engineer by the numbers
- 30 all-time installs (skills.sh)
- Ranked #9,276 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
- Data as of Jul 29, 2026 (Skillselion catalog sync)
npx skills add https://github.com/daemon-blockint-tech/agentic-enteprises-skill --skill ai-engineerAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 30 |
|---|---|
| repo stars | ★ 7 |
| Last updated | May 20, 2026 |
| Repository | daemon-blockint-tech/agentic-enteprises-skill ↗ |
What it does
Design and ship production LLM features like RAG pipelines, agent workflows, and eval harnesses with cost, latency, and safe-deploy guardrails.
Files
AI Engineer
When to Use
- Building chatbots, copilots, or retrieval-augmented generation systems
- Designing multi-step agent workflows with tool use
- Integrating OpenAI, Anthropic, or local models into products
- Setting up RAG pipelines (chunk, embed, index, retrieve, rerank, generate)
- Building evaluation harnesses and regression suites for LLM features
- Optimizing cost/latency through model routing, caching, or context strategy
- Planning safe deployment of generative features (canary, kill switch, monitoring)
When NOT to Use
- Academic literature synthesis or research methodology →
ai-researcher - Organizational AI policy, regulation, or risk tiering →
ai-risk-governance - Adversarial safety testing and jailbreak campaigns →
ai-redteam - Prompt-only tuning without system architecture changes →
prompt-engineer - Enterprise-wide non-AI system integration ADRs →
senior-system-architecture - Token/cost improvement program planning and roadmap →
ai-token-improvement-plan-engineer - Commercial/enterprise AI solution architecture →
applied-ai-architect-commercial-enterprise - Skills portfolio governance and batch validation →
ai-skill-manager - Agent prompts, golden evals, judge rubrics →
prompt-engineer-agent-prompts-evals
Related skills
| Need | Skill |
|---|---|
| Prompt templates and agent message design | prompt-engineer |
| Offline experiments, statistics, classical ML | data-scientist |
| Papers, benchmarks, research methodology | ai-researcher |
| Policies, model cards, governance | ai-risk-governance |
| Red-team and jailbreak campaigns | ai-redteam |
| Persistent memory design | ai-memory-developer |
| Context window and token budgeting | ai-context-engineer |
| Token cost improvement plan and roadmap | ai-token-improvement-plan-engineer |
| AI production ops and release governance | ai-lead-ops |
| Cross-system boundaries and platform ADRs | senior-system-architecture |
| Commercial/enterprise AI architecture | applied-ai-architect-commercial-enterprise |
| Agent skills catalog and validation | ai-skill-manager |
| Safeguard serving stack and policy runtime | ml-infrastructure-engineer-safeguards |
| Safety model R&D and benchmark design | ml-research-engineer-safeguards |
Core Workflows
1. Solution shaping
1. Define user job, success metric, and failure modes 2. Decide: single LLM call vs RAG vs multi-step agent 3. Choose model tier (quality vs cost vs latency) 4. Identify data sources, PII boundaries, and retention 5. Plan human-in-the-loop for high-risk actions
See `references/solution_patterns.md` for RAG vs fine-tune vs agent decision tree.
2. RAG pipeline
ingest → chunk → embed → index → retrieve → rerank → generate → citeChecklist:
- [ ] Chunk size tuned on eval set
- [ ] Metadata filters for tenancy/ACL
- [ ] Hybrid search if keyword matters
- [ ] Ground answers with citations; refuse when context insufficient
- [ ] Refresh index on source updates
See `references/rag_pipeline.md` for chunking, eval metrics, and freshness.
3. Agents and tools
- Tools: narrow schemas, idempotent where possible, timeouts
- Loop: plan → act → observe → stop condition
- Cap iterations and token budget
- Log tool calls for audit; redact secrets in traces
See `references/agents_tools.md` for ReAct patterns and failure handling.
4. Evaluation before launch
| Layer | Measure |
|---|---|
| Retrieval | Recall@k, MRR on golden questions |
| Generation | Faithfulness, answer relevance (LLM-judge + human sample) |
| Safety | Refusal rate on policy violations |
| Ops | p95 latency, cost per session |
Ship only when regression suite passes on CI for golden set.
See `references/evaluation_ops.md` for datasets, CI eval, and monitoring.
5. Production operations
- Version prompts and models; canary new versions
- Monitor drift, error rate, tool failures, spend
- Kill switch for model or feature flag
- Incident runbook for toxic output or data leak
See `references/evaluation_ops.md` for production monitoring.
When to load references
- Architecture choices →
references/solution_patterns.md - RAG implementation →
references/rag_pipeline.md - Agents and tools →
references/agents_tools.md - Eval and production →
references/evaluation_ops.md
Agents and tools
Table of contents
1. Tool schema 2. Loop controls 3. Failure handling
Tool schema
- Clear name and description for when to call
- Required fields explicit; enums for fixed choices
- Validate arguments before execution
Loop controls
- Max iterations (e.g., 10)
- Max wall-clock time
- Stop when final answer tool invoked
Failure handling
- Retry transient errors with backoff
- Surface tool error to model once; do not infinite retry
- Human escalation path for repeated failures
Evaluation and operations
Table of contents
1. Golden datasets 2. CI eval 3. Production monitoring
Golden datasets
50–200 examples covering:
- Happy path
- Edge cases
- Refusal/policy cases
- Known historical failures
Version dataset with git tag.
CI eval
Run on PR when prompts, retrieval, or models change; block on regression of primary metric.
Production monitoring
- Token usage and cost per feature
- Latency p50/p95
- Refusal and escalation rates
- User thumbs down sampling for review
RAG pipeline
Table of contents
1. Chunking 2. Retrieval 3. Eval metrics
Chunking
- Start 512–1024 tokens with overlap 10–20%
- Preserve headings in chunk metadata
- Split code and prose differently
Retrieval
- Metadata filters for tenant_id mandatory in multi-tenant
- Rerank top 20 → pass top 5 to LLM
- Return "insufficient context" when scores below threshold
Eval metrics
| Metric | Meaning |
|---|---|
| Recall@k | Gold doc in top k |
| Faithfulness | Answer supported by chunks |
| Answer relevance | Addresses question |
Solution patterns
Table of contents
1. Decision tree 2. Model routing
Decision tree
| Need | Pattern |
|---|---|
| Fixed Q&A on static docs | RAG |
| Structured extraction | Single call + JSON schema |
| Multi-step tools/APIs | Agent with tools |
| Domain tone/style | System prompt + few-shot; fine-tune if insufficient |
Model routing
Route by task complexity:
- Small/fast model for classification and routing
- Large model for synthesis and hard reasoning
- Fallback on timeout with degraded response message