
Ai Redteam
- 33 installs
- 7 repo stars
- Updated May 20, 2026
- daemon-blockint-tech/agentic-enteprises-skill
Red-team LLM applications for prompt injection, jailbreaks, tool abuse, and data exfiltration, with harnesses, reporting, and mitigation retests.
About
Guides adversarial testing of AI systems including prompt injection, jailbreaks, tool abuse, data exfiltration, and multi-turn attacks. A developer uses it when red-teaming chatbots, agents, or RAG systems before launch or validating mitigations.
- LLM threat model: prompt injection, jailbreak, tool abuse, exfiltration
- Test phases from baseline to automated sweep, manual, and regression
Ai Redteam by the numbers
- 33 all-time installs (skills.sh)
- Ranked #1,467 of 2,203 Security skills by installs in the Skillselion catalog
- Data as of Jul 29, 2026 (Skillselion catalog sync)
npx skills add https://github.com/daemon-blockint-tech/agentic-enteprises-skill --skill ai-redteamAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 33 |
|---|---|
| repo stars | ★ 7 |
| Last updated | May 20, 2026 |
| Repository | daemon-blockint-tech/agentic-enteprises-skill ↗ |
What it does
Red-team LLM applications for prompt injection, jailbreaks, tool abuse, and data exfiltration, with harnesses, reporting, and mitigation retests.
Files
AI Red Team
When to Use
- Red-teaming chatbots, agents, RAG systems, or copilots before launch
- Designing safety evaluation suites and adversarial test harnesses
- Reproducing reported prompt injection or jailbreak vulnerabilities
- Validating mitigations after incidents (retesting filters, hardening)
- Running multi-turn coercion, encoding, or indirect injection campaigns
- Assessing bias, harmful output, or data exfiltration risks in LLM applications
- Scoping rules of engagement and severity rubrics for AI security testing
When NOT to Use
- Writing corporate AI policy or risk governance frameworks →
ai-risk-governance - Building production LLM features or RAG pipelines →
ai-engineer - General network/AD/infra penetration testing →
network-pentester - Authorized web/API OWASP testing (non-LLM) →
web-pentester - Enterprise adversary simulation, MITRE ATT&CK campaigns, purple team →
red-team-specialist - Binary, firmware, or protocol reverse engineering →
reverse-engineer - CI/CD pipeline security →
devsecops
Related skills
| Need | Skill |
|---|---|
| Production architecture and mitigations | ai-engineer |
| Governance sign-off and risk tiers | ai-risk-governance |
| Prompt design baselines | prompt-engineer |
| CI pipeline security | devsecops |
| Web/API OWASP pentest (non-LLM) | web-pentester |
| Network/AD/infra pentest (non-LLM) | network-pentester |
| Multi-domain pentest (non-LLM) | penetration-tester |
| Enterprise red team / adversary simulation (non-LLM) | red-team-specialist |
| Security program and pentest governance | cybersecurity |
| Deploy/monitor safeguard inference path | ml-infrastructure-engineer-safeguards |
| Safety benchmarks and classifier training | ml-research-engineer-safeguards |
| Post-incident disk/memory/log forensics and chain of custody | digital-forensics-analyst |
| Binary/protocol RE on non-LLM malware or implants | reverse-engineer |
| Security incident coordination after AI abuse | incident-responder |
Core Workflows
1. Scope and rules of engagement
1. Define target: model, app surface, tools, data stores 2. Obtain written authorization and time window 3. Agree out-of-scope (e.g., no social engineering of employees unless approved) 4. Define success criteria: critical findings, reproduction steps, severity rubric 5. Plan safe test environment (no prod customer data)
See `references/engagement_scope.md` for ROE template and severity definitions.
2. Threat model for LLM applications
| Class | Examples |
|---|---|
| Prompt injection | Instructions in user/doc content override system policy |
| Jailbreak | Role-play, encoding, multi-turn coercion |
| Tool abuse | Unauthorized API calls, parameter injection |
| Data exfiltration | RAG leaks other tenants' chunks, PII in logs |
| Supply chain | Malicious tool definitions, compromised plugins |
| Denial of service | Token burn, recursive agent loops |
See `references/attack_catalog.md` for technique families and test prompts (use ethically).
3. Test execution
Phases:
1. Baseline — document intended refusals and allowed behaviors 2. Automated sweep — harness with curated attack set + fuzz mutations 3. Manual creativity — domain-specific abuse scenarios 4. Tool/RAG focus — indirect injection via retrieved documents 5. Regression — re-run after mitigations
Log: input, output, tool calls, latency, whether guardrail fired.
See `references/testing_harness.md` for harness design and datasets.
4. Reporting
Each finding includes:
- Title and severity (impact × likelihood)
- Steps to reproduce (minimal)
- Evidence (redacted transcripts)
- Affected component
- Recommended mitigation
- Retest criteria
See `references/reporting.md` for report template and remediation tracking.
5. Mitigation validation
| Mitigation | Retest |
|---|---|
| Input/output filters | Bypass attempts with paraphrases |
| System prompt hardening | Injection via RAG context |
| Tool allowlists | Confused deputy and scope creep |
| Human approval gate | Automated agent bypass paths |
See `references/mitigations.md` for defense depth and known weak controls.
When to load references
- ROE and scope →
references/engagement_scope.md - Attack types →
references/attack_catalog.md - Harness and automation →
references/testing_harness.md - Reports →
references/reporting.md - Defenses →
references/mitigations.md
Attack catalog
Table of contents
1. Prompt injection 2. Jailbreak families 3. RAG attacks 4. Ethics
Prompt injection
- Direct: user overrides system in same turn
- Indirect: malicious content in retrieved doc or email parsed by agent
- Payload in metadata fields (titles, alt text)
Jailbreak families
- Role-play and fictional framing
- Encoding (base64, other languages)
- Multi-turn gradual compliance
- Refusal suppression ("ignore previous")
Test only on authorized systems; do not publish working exploits against third parties without permission.
RAG attacks
- Plant document in index with hidden instructions
- Poison chunk boundaries to split defenses
Ethics
Authorized testing only; minimize harm; redact PII in reports.
Engagement scope
Table of contents
1. ROE template 2. Severity rubric
ROE template
# AI red team — rules of engagement
- Target systems:
- Authorized testers:
- Window:
- Prohibited:
- Data handling:
- Emergency contact:Severity rubric
| Level | Criteria |
|---|---|
| Critical | Cross-tenant data, arbitrary tool execution, policy bypass at scale |
| High | Reliable injection affecting many users |
| Medium | Limited bypass or single-tenant leak |
| Low | Theoretical, heavy user cooperation |
Mitigations
Table of contents
1. Defense layers 2. Weak controls
Defense layers
| Layer | Control |
|---|---|
| Input | Length limits, blocklists, classifiers |
| System | Strong policy, delimiter isolation |
| RAG | Source trust tiers, chunk sanitization |
| Tools | Allowlist, authz per tool, human approval |
| Output | Policy classifier, PII redaction |
Weak controls
- Security through obscurity in system prompt alone
- Single regex filter
- Assuming fine-tuning removes jailbreaks
- Trusting retrieved HTML without sanitization
Always validate with ai-redteam regression set after changes.
Reporting
Table of contents
1. Finding template 2. Remediation tracking
Finding template
### [SEV] Title
**Component:**
**Steps:**
1.
**Impact:**
**Evidence:** (redacted transcript)
**Recommendation:**
**Retest:**Remediation tracking
Link findings to tickets; retest within 2 weeks of fix; close only on passing regression suite.
Testing harness
Table of contents
1. Harness design 2. Dataset sources
Harness design
for attack in dataset:
response = app.send(attack)
score = judge(policy_violation, leakage, tool_abuse)
log(transcript, score)- Pin model and prompt versions per run
- Parallelize with rate limits
- Compare to baseline after mitigations
Dataset sources
- Internal policy violation probes
- Public harm benchmarks (use responsibly)
- Custom cases from support tickets (sanitized)
- Mutations of failed cases from prior runs