Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
proffesor-for-testing avatar

Pentest Validation

  • 120 installs
  • 433 repo stars
  • Updated August 4, 2026
  • proffesor-for-testing/agentic-qe

Verify penetration-test findings, reproduce exploits safely, prioritize remediations, and confirm fixes with regression security checks before external audit sign-off.

About

Supports pentest validation workflows in agentic-qe by helping teams reproduce findings, prioritize fixes, document remediations, and re-verify closures so security assessments translate into shippable, audit-ready hardening rather than stale ticket backlogs.

  • Finding reproduction steps
  • Risk prioritization rubrics
  • Remediation verification
  • Safe exploit sandboxing
  • Audit-ready evidence capture

Pentest Validation by the numbers

  • 120 all-time installs (skills.sh)
  • +3 installs in the week ending Aug 4, 2026 (Skillselion tracking)
  • Ranked #953 of 2,203 Security skills by installs in the Skillselion catalog
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/proffesor-for-testing/agentic-qe --skill pentest-validation

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs120
repo stars433
Last updatedAugust 4, 2026
Repositoryproffesor-for-testing/agentic-qe

What it does

Verify penetration-test findings, reproduce exploits safely, prioritize remediations, and confirm fixes with regression security checks before external audit sign-off.

Files

SKILL.mdMarkdownGitHub ↗

Pentest Validation

<default_to_action> When validating security findings: 1. REQUIRE explicit authorization for target URL 2. SCAN with qe-security-scanner (SAST + dependency + secrets) 3. ANALYZE with qe-security-reviewer + qe-security-auditor (parallel) 4. VALIDATE with qe-pentest-validator (graduated exploitation, parallel per vuln type) 5. REPORT only confirmed findings with PoC evidence ("No Exploit, No Report") 6. UPDATE exploit playbook with new patterns

Quality Gates:

  • Authorization confirmed before ANY exploitation
  • Target URL is staging/dev (NOT production)
  • Budget cap enforced ($15 default)
  • Time cap enforced (30 min default)
  • All exploitation attempts logged

</default_to_action>

Quick Reference Card

The 4-Phase Pipeline

PhaseAgent(s)PurposeParallelism
1. Reconqe-security-scannerSAST, DAST, dependency scan, secretsInternal parallel
2. Analysisqe-security-reviewer + qe-security-auditorCode review + compliance checkBoth in parallel
3. Validationqe-pentest-validatorGraduated exploit validationPer-vuln-type parallel
4. Reportqe-quality-gate"No Exploit, No Report" filterSequential

Graduated Exploitation Tiers

TierHandlerCostLatencyUse When
1Agent Booster (WASM)$0<1msCode pattern is conclusive (eval, innerHTML, hardcoded creds)
2Haiku$0.0002~500msNeed payload test against live target
3Sonnet/Opus$0.003-$0.0152-5sFull exploit chain with data proof

When to Use This Skill

ScenarioTierEstimated Cost
PR security review (source only)1$0
Pre-release validation (staging)1-2$1-5
Full pentest validation1-3$5-15
Compliance audit evidence1-3$5-15

---

Configuration

pentest:
  target_url: https://staging.app.com    # REQUIRED for Tier 2-3
  source_repo: ./src                      # REQUIRED for Tier 1+
  exploitation_tier: 2                    # 1=pattern-only, 2=payload-test, 3=full-exploit
  vuln_types:                             # Which pipelines to run
    - injection                           # SQL, NoSQL, command injection
    - xss                                 # Reflected, stored, DOM XSS
    - auth                                # Auth bypass, session, JWT
    - ssrf                                # URL scheme abuse, metadata
  max_cost_usd: 15                        # Budget cap per run
  timeout_minutes: 30                     # Time cap per run
  require_authorization: true             # MUST confirm target ownership
  no_production: true                     # Block production URLs
  production_patterns:                    # URL patterns to block
    - "*.prod.*"
    - "api.*"
    - "www.*"

---

Safeguards (Mandatory)

Authorization Gate

Every pentest validation run MUST: 1. Display target URL and exploitation tier to user 2. Require explicit confirmation: "I own/authorized testing of this target" 3. Log authorization with timestamp 4. Block if target URL matches production patterns

What This Skill Does NOT Do

  • Full autonomous reconnaissance (Nmap, Subfinder)
  • Zero-day exploit development
  • Attack targets without explicit authorization
  • Test production systems
  • Store actual exfiltrated data (only proof of access)
  • Social engineering or phishing simulation
  • Port scanning or service discovery

---

Validation Pipelines

Injection Pipeline

AttackTier 1 (Pattern)Tier 2 (Payload)Tier 3 (Full)
SQL injectionString concat in query' OR '1'='1 response diffUNION SELECT data extraction
NoSQL injection$where, $gt in queryOperator injection testCollection enumeration
Command injectionexec(), system() callsCommand delimiter testReverse shell proof
LDAP injectionString concat in filterWildcard injectionDirectory enumeration

XSS Pipeline

AttackTier 1 (Pattern)Tier 2 (Payload)Tier 3 (Full)
Reflected XSSNo output encoding<img onerror> reflectionBrowser JS execution via qe-browser (Vibium)
Stored XSSinnerHTML assignmentPayload stored + retrievedCookie theft PoC
DOM XSSdocument.write(location)Fragment injectionDOM manipulation proof

Auth Pipeline

AttackTier 1 (Pattern)Tier 2 (Payload)Tier 3 (Full)
JWT noneNo algorithm validationModified JWT acceptedAdmin access with forged token
Session fixationNo session rotationPre-set session reusedCross-user session hijack
Credential stuffingNo rate limiting100 attempts unblockedValid credential discovery
IDORNo authorization checkAccess other user dataFull CRUD on foreign resources

SSRF Pipeline

AttackTier 1 (Pattern)Tier 2 (Payload)Tier 3 (Full)
Internal URLUser-controlled URL fetchhttp://169.254.169.254Cloud metadata extraction
DNS rebindingURL validation bypassRebind to internal IPInternal service access
Protocol smugglingURL scheme not restrictedfile:///etc/passwdFile content in response

---

Agent Coordination

Orchestration Pattern

// Phase 1: Recon (parallel scans)
await Task("Security Scan", {
  target: "./src",
  layers: { sast: true, dast: true, dependencies: true, secrets: true }
}, "qe-security-scanner");

// Phase 2: Analysis (parallel review)
await Promise.all([
  Task("Code Security Review", {
    findings: phase1Results,
    depth: "comprehensive"
  }, "qe-security-reviewer"),

  Task("Compliance Audit", {
    findings: phase1Results,
    frameworks: ["owasp-top-10"]
  }, "qe-security-auditor")
]);

// Phase 3: Validation (graduated exploitation)
await Task("Exploit Validation", {
  findings: [...phase1Results, ...phase2Results],
  target_url: "https://staging.app.com",
  exploitation_tier: 2,
  vuln_types: ["injection", "xss", "auth", "ssrf"],
  max_cost_usd: 15,
  timeout_minutes: 30
}, "qe-pentest-validator");

// Phase 4: Report ("No Exploit, No Report" gate)
await Task("Security Quality Gate", {
  findings: phase3Results.confirmedFindings,
  gate: "no-exploit-no-report",
  require_poc: true
}, "qe-quality-gate");

Finding Classification

StatusMeaningAction
confirmed-exploitableExploitation succeeded with PoCReport with evidence
likely-exploitablePartial exploitation, defenses detectedReport with caveats
not-exploitableAll exploitation attempts failedFilter from report
inconclusiveWAF/defense blocked, unclear if vulnerableReport for manual review

---

Exploit Playbook Memory

Namespace Structure

aqe/pentest/
 playbook/
  exploit/{vuln_type}/{tech_stack}/{technique}
  bypass/{defense_type}/{technique}
  payload/{vuln_type}/{variant}
 results/
  validation-{timestamp}
 poc/
  {finding_id}-poc

Learning Loop

1. Before validation: Query playbook for known patterns matching findings 2. During validation: Try known payloads first (higher success rate) 3. After validation: Store new successful patterns with confidence scores 4. Over time: Agent converges on most effective payloads per tech stack

---

Cost Optimization

Estimated Cost by Scenario

ScenarioTier MixFindingsEst. CostEst. Time
PR check (source only)100% Tier 15$0<5s
Sprint validation70% T1, 30% T215$2-55-10 min
Release validation40% T1, 40% T2, 20% T325$8-1515-30 min
Full pentest20% T1, 30% T2, 50% T340$15-3030-60 min

Cost vs Shannon Comparison

MetricShannonAQE Pentest Validation
Cost per run~$50$5-15 (graduated tiers)
Runtime60-90 min15-30 min (parallel pipelines)
False positive rateLow (exploit-proven)Low (same principle)
LearningNone (static prompts)ReasoningBank playbook

---

Success Metrics

MetricTargetMeasurement
False positive reduction>60% of findings eliminatedPre/post validator comparison
Exploit confirmation rate>80% of confirmed findings truly exploitableManual PoC verification
Cost per run<$15 USDToken tracking per pipeline
Time per run<30 minutesExecution time metrics
Playbook growth100+ patterns after 6 monthsMemory namespace count

---

Related Skills

  • security-testing - OWASP vulnerability scanning, SAST/DAST automation
  • compliance-testing - Regulatory compliance
  • api-testing-patterns - API security testing
  • chaos-engineering-resilience - Security under chaos

---

Remember

"No Exploit, No Report." A vulnerability scanner that can't prove exploitation delivers uncertain value. This skill transforms security findings from theoretical risks into proven vulnerabilities with evidence. Every confirmed finding comes with a reproducible proof-of-concept. Every false positive is eliminated before it reaches the report.

Think proof, not prediction. Don't report what MIGHT be vulnerable. Prove what IS vulnerable.

Related skills

Securityauditappseccompliance

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.