
Team Testing
- 61 installs
- 2.1k repo stars
- Updated June 18, 2026
- catlog22/claude-code-workflow
Generate and run automated tests
About
Automates testing for team-testing. Teams use this to ensure code quality before deployment.
- Test automation
- Quality assurance
Team Testing by the numbers
- 61 all-time installs (skills.sh)
- Ranked #1,159 of 2,153 Testing & QA skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/catlog22/claude-code-workflow --skill team-testingAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 61 |
|---|---|
| repo stars | ★ 2.1k |
| Last updated | June 18, 2026 |
| Repository | catlog22/claude-code-workflow ↗ |
What it does
Generate and run automated tests
Files
Team Testing
Orchestrate multi-agent test pipeline: strategist -> generator -> executor -> analyst. Progressive layer coverage (L1/L2/L3) with Generator-Critic loops for coverage convergence.
Architecture
Skill(skill="team-testing", args="task description")
|
SKILL.md (this file) = Router
|
+--------------+--------------+
| |
no --role flag --role <name>
| |
Coordinator Worker
roles/coordinator/role.md roles/<name>/role.md
|
+-- analyze -> dispatch -> spawn workers -> STOP
|
+-------+-------+-------+-------+
v v v v
[strat] [gen] [exec] [analyst]
team-worker agents, each loads roles/<role>/role.mdRole Registry
| Role | Path | Prefix | Inner Loop |
|---|---|---|---|
| coordinator | roles/coordinator/role.md | — | — |
| strategist | roles/strategist/role.md | STRATEGY-* | false |
| generator | roles/generator/role.md | TESTGEN-* | true |
| executor | roles/executor/role.md | TESTRUN-* | dynamic |
| analyst | roles/analyst/role.md | TESTANA-* | false |
Role Router
Parse $ARGUMENTS:
- Has
--role <name>-> Readroles/<name>/role.md, execute Phase 2-4 - No
--role->@roles/coordinator/role.md, execute entry router
Shared Constants
- Session prefix:
TST - Session path:
.workflow/.team/TST-<slug>-<date>/ - Team name:
testing - CLI tools:
ccw cli --mode analysis(read-only),ccw cli --mode write(modifications) - Message bus:
mcp__ccw-tools__team_msg(session_id=<session-id>, ...)
Worker Spawn Template
Coordinator spawns workers using this template:
Agent({
subagent_type: "team-worker",
description: "Spawn <role> worker",
team_name: "testing",
name: "<role>",
run_in_background: true,
prompt: `## Role Assignment
role: <role>
role_spec: <skill_root>/roles/<role>/role.md
session: <session-folder>
session_id: <session-id>
team_name: testing
requirement: <task-description>
inner_loop: <true|false>
## Progress Milestones
session_id: <session-id>
Report progress via team_msg at natural phase boundaries (context loaded -> core work done -> verification).
Report blockers immediately via team_msg type="blocker".
Report completion via team_msg type="task_complete" after final SendMessage.
Read role_spec file (@<skill_root>/roles/<role>/role.md) to load Phase 2-4 domain instructions.
Execute built-in Phase 1 (task discovery) -> role Phase 2-4 -> built-in Phase 5 (report).`
})User Commands
| Command | Action |
|---|---|
check / status | View pipeline status graph |
resume / continue | Advance to next step |
revise <TASK-ID> | Revise specific task |
feedback <text> | Inject feedback for revision |
Completion Action
When pipeline completes, coordinator presents:
AskUserQuestion({
questions: [{
question: "Testing pipeline complete. What would you like to do?",
header: "Completion",
multiSelect: false,
options: [
{ label: "Archive & Clean (Recommended)", description: "Archive session, clean up team" },
{ label: "Keep Active", description: "Keep session for follow-up work" },
{ label: "Deepen Coverage", description: "Add more test layers or increase coverage targets" }
]
}]
})Session Directory
.workflow/.team/TST-<slug>-<date>/
├── .msg/messages.jsonl # Team message bus
├── .msg/meta.json # Session metadata
├── wisdom/ # Cross-task knowledge
├── strategy/ # Strategist output
├── tests/ # Generator output (L1-unit/, L2-integration/, L3-e2e/)
├── results/ # Executor output
└── analysis/ # Analyst outputSpecs Reference
- specs/pipelines.md — Pipeline definitions and task registry
- specs/team-config.json — Team configuration
Error Handling
| Scenario | Resolution |
|---|---|
| Unknown --role value | Error with available role list |
| Role not found | Error with expected path (roles/<name>/role.md) |
| CLI tool fails | Worker fallback to direct implementation |
| GC loop exceeded | Accept current coverage with warning |
| Fast-advance conflict | Coordinator reconciles on next callback |
| Completion action fails | Default to Keep Active |
Test Quality Analyst
Analyze defect patterns, identify coverage gaps, assess GC loop effectiveness, and generate a quality report with actionable recommendations.
Phase 2: Context Loading
| Input | Source | Required |
|---|---|---|
| Task description | From task subject/description | Yes |
| Session path | Extracted from task description | Yes |
| Execution results | <session>/results/run-*.json | Yes |
| Test strategy | <session>/strategy/test-strategy.md | Yes |
| .msg/meta.json | <session>/wisdom/.msg/meta.json | Yes |
1. Extract session path from task description 2. Read .msg/meta.json for execution context (executor, generator namespaces) 3. Read all execution results:
Glob("<session>/results/run-*.json")
Read("<session>/results/run-001.json")4. Read test strategy:
Read("<session>/strategy/test-strategy.md")5. Read test files for pattern analysis:
Glob("<session>/tests/**/*")Phase 3: Quality Analysis
Analysis dimensions:
1. Coverage Analysis -- Aggregate coverage by layer:
| Layer | Coverage | Target | Status |
|---|---|---|---|
| L1 | X% | Y% | Met/Below |
2. Defect Pattern Analysis -- Frequency and severity:
| Pattern | Frequency | Severity |
|---|---|---|
| pattern | count | HIGH (>=3) / MEDIUM (>=2) / LOW (<2) |
3. GC Loop Effectiveness:
| Metric | Value | Assessment |
|---|---|---|
| Rounds | N | - |
| Coverage Improvement | +/-X% | HIGH (>10%) / MEDIUM (>5%) / LOW (<=5%) |
4. Coverage Gaps -- per module/feature:
- Area, Current %, Gap %, Reason, Recommendation
5. Quality Score:
| Dimension | Score (1-10) | Weight |
|---|---|---|
| Coverage Achievement | score | 30% |
| Test Effectiveness | score | 25% |
| Defect Detection | score | 25% |
| GC Loop Efficiency | score | 20% |
Write report to <session>/analysis/quality-report.md
Tech Profile Scan
After test analysis, emit context-aware trigger signals (based on detected codebase characteristics):
1. Check test findings → signals (test_gap, perf_sensitive) 2. Check tested code → risk signals (sql_detected, auth_detected, injection_risk) 3. Include tech_profile in Phase 5 state_update data
Phase 4: Trend Analysis & State Update
Historical comparison (if multiple sessions exist):
Glob(".workflow/.team/TST-*/.msg/meta.json")- Track coverage trends over time
- Identify defect pattern evolution
- Compare GC loop effectiveness across sessions
Update <session>/wisdom/.msg/meta.json under analyst namespace:
- Merge
{ "analyst": { quality_score, coverage_gaps, top_defect_patterns, gc_effectiveness, recommendations } }
Analyze Task
Parse user task -> detect testing capabilities -> select pipeline -> design roles.
CONSTRAINT: Text-level analysis only. NO source code reading, NO codebase exploration.
Signal Detection
| Keywords | Capability | Prefix |
|---|---|---|
| strategy, plan, layers, scope | strategist | STRATEGY |
| generate tests, write tests, create tests | generator | TESTGEN |
| run tests, execute, coverage | executor | TESTRUN |
| analyze, report, quality, defects | analyst | TESTANA |
Pipeline Mode Detection
| Condition | Pipeline |
|---|---|
| fileCount <= 3 AND moduleCount <= 1 | targeted |
| fileCount <= 10 AND moduleCount <= 3 | standard |
| Otherwise | comprehensive |
Dependency Graph
Natural ordering for testing pipeline:
- Tier 0: strategist (change analysis, no upstream dependency)
- Tier 1: generator (requires strategy)
- Tier 2: executor (requires generated tests; GC loop with generator)
- Tier 3: analyst (requires execution results)
Pipeline Definitions
Targeted: STRATEGY -> TESTGEN(L1) -> TESTRUN(L1)
Standard: STRATEGY -> TESTGEN(L1) -> TESTRUN(L1) -> TESTGEN(L2) -> TESTRUN(L2) -> TESTANA
Comprehensive: STRATEGY -> [TESTGEN(L1) || TESTGEN(L2)] -> [TESTRUN(L1) || TESTRUN(L2)] -> TESTGEN(L3) -> TESTRUN(L3) -> TESTANAComplexity Scoring
| Factor | Points |
|---|---|
| Per test layer | +1 |
| Parallel tracks | +1 per track |
| GC loop enabled | +1 |
| Serial depth > 3 | +1 |
Results: 1-2 Low, 3-5 Medium, 6+ High
Role Minimization
- Cap at 5 roles (coordinator + 4 workers)
- GC loop: generator <-> executor iterate up to 3 rounds per layer
Output
Write <session>/task-analysis.json:
{
"task_description": "<original>",
"pipeline_mode": "<targeted|standard|comprehensive>",
"capabilities": [{ "name": "<cap>", "prefix": "<PREFIX>", "keywords": ["..."] }],
"dependency_graph": { "<TASK-ID>": { "role": "<role>", "blockedBy": ["..."], "layer": "L1|L2|L3" } },
"roles": [{ "name": "<role>", "prefix": "<PREFIX>", "inner_loop": true }],
"complexity": { "score": 0, "level": "Low|Medium|High" },
"coverage_targets": { "L1": 80, "L2": 60, "L3": 40 },
"gc_loop_enabled": true
}Dispatch Tasks
Create testing task chains with correct dependencies. Supports targeted, standard, and comprehensive pipelines.
Workflow
1. Read task-analysis.json -> extract pipeline_mode and dependency_graph 2. Read specs/pipelines.md -> get task registry for selected pipeline 3. Topological sort tasks (respect blockedBy) 4. Validate all owners exist in role registry (SKILL.md) 5. For each task (in order):
- TaskCreate with structured description (see template below)
- TaskUpdate with blockedBy + owner assignment
6. Update session.json with pipeline.tasks_total 7. Validate chain (no orphans, no cycles, all refs valid)
Task Description Template
PURPOSE: <goal> | Success: <criteria>
TASK:
- <step 1>
- <step 2>
CONTEXT:
- Session: <session-folder>
- Scope: <scope>
- Layer: <L1-unit|L2-integration|L3-e2e>
- Upstream artifacts: <artifact-1>, <artifact-2>
- Shared memory: <session>/wisdom/.msg/meta.json
EXPECTED: <deliverable path> + <quality criteria>
CONSTRAINTS: <scope limits, focus areas>
---
InnerLoop: <true|false>
RoleSpec: ~ or <project>/.claude/skills/team-testing/roles/<role>/role.mdPipeline Task Registry
Targeted Pipeline
STRATEGY-001 (strategist): Analyze change scope, define test strategy
blockedBy: []
TESTGEN-001 (generator): Generate L1 unit tests
blockedBy: [STRATEGY-001], meta: layer=L1-unit
TESTRUN-001 (executor): Execute L1 tests, collect coverage
blockedBy: [TESTGEN-001], inner_loop: true, meta: layer=L1-unit, coverage_target=80%Standard Pipeline
STRATEGY-001 (strategist): Analyze change scope, define test strategy
blockedBy: []
TESTGEN-001 (generator): Generate L1 unit tests
blockedBy: [STRATEGY-001], meta: layer=L1-unit
TESTRUN-001 (executor): Execute L1 tests, collect coverage
blockedBy: [TESTGEN-001], inner_loop: true, meta: layer=L1-unit, coverage_target=80%
TESTGEN-002 (generator): Generate L2 integration tests
blockedBy: [TESTRUN-001], meta: layer=L2-integration
TESTRUN-002 (executor): Execute L2 tests, collect coverage
blockedBy: [TESTGEN-002], inner_loop: true, meta: layer=L2-integration, coverage_target=60%
TESTANA-001 (analyst): Defect pattern analysis, quality report
blockedBy: [TESTRUN-002]Comprehensive Pipeline
STRATEGY-001 (strategist): Analyze change scope, define test strategy
blockedBy: []
TESTGEN-001 (generator-1): Generate L1 unit tests
blockedBy: [STRATEGY-001], meta: layer=L1-unit
TESTGEN-002 (generator-2): Generate L2 integration tests
blockedBy: [STRATEGY-001], meta: layer=L2-integration
TESTRUN-001 (executor-1): Execute L1 tests, collect coverage
blockedBy: [TESTGEN-001], inner_loop: true, meta: layer=L1-unit, coverage_target=80%
TESTRUN-002 (executor-2): Execute L2 tests, collect coverage
blockedBy: [TESTGEN-002], inner_loop: true, meta: layer=L2-integration, coverage_target=60%
TESTGEN-003 (generator): Generate L3 E2E tests
blockedBy: [TESTRUN-001, TESTRUN-002], meta: layer=L3-e2e
TESTRUN-003 (executor): Execute L3 tests, collect coverage
blockedBy: [TESTGEN-003], inner_loop: true, meta: layer=L3-e2e, coverage_target=40%
TESTANA-001 (analyst): Defect pattern analysis, quality report
blockedBy: [TESTRUN-003]InnerLoop Flag Rules
- generator: always true (GC loop iterations within a single layer)
- executor: dynamic per pipeline mode:
- Targeted/Standard:
true(serial layer chain, single executor handles GC loops) - Comprehensive:
falsefor parallel TESTRUN tasks (TESTRUN-001 and TESTRUN-002 have independent blockedBy, each gets own worker) - strategist, analyst: always false
Dependency Validation
- No orphan tasks (all tasks have valid owner)
- No circular dependencies
- All blockedBy references exist
- Session reference in every task description
- RoleSpec reference in every task description
Log After Creation
mcp__ccw-tools__team_msg({
operation: "log",
session_id: <session-id>,
from: "coordinator",
type: "pipeline_selected",
data: { pipeline: "<mode>", task_count: <N> }
})Monitor Pipeline
Event-driven pipeline coordination. Beat model: coordinator wake -> process -> spawn -> STOP.
Constants
- SPAWN_MODE: background
- ONE_STEP_PER_INVOCATION: true
- FAST_ADVANCE_AWARE: true
- WORKER_AGENT: team-worker
- MAX_GC_ROUNDS: 3
Handler Router
| Source | Handler |
|---|---|
| Message contains [strategist], [generator], [executor], [analyst] | handleCallback |
| "capability_gap" | handleAdapt |
| "check" or "status" | handleCheck |
| "resume" or "continue" | handleResume |
| All tasks completed | handleComplete |
| Default | handleSpawnNext |
Role-Worker Map
| Prefix | Role | Role Spec | inner_loop |
|---|---|---|---|
| STRATEGY-* | strategist | ~ or <project>/.claude/skills/team-testing/roles/strategist/role.md | false |
| TESTGEN-* | generator | ~ or <project>/.claude/skills/team-testing/roles/generator/role.md | true |
| TESTRUN-* | executor | ~ or <project>/.claude/skills/team-testing/roles/executor/role.md | dynamic |
| TESTANA-* | analyst | ~ or <project>/.claude/skills/team-testing/roles/analyst/role.md | false |
handleCallback
Worker completed. Process and advance.
1. Parse message to identify role and task ID:
| Message Pattern | Role Detection |
|---|---|
[strategist] or task ID STRATEGY-* | strategist |
[generator] or task ID TESTGEN-* | generator |
[executor] or task ID TESTRUN-* | executor |
[analyst] or task ID TESTANA-* | analyst |
2. Check if progress update (inner loop) or final completion 3. Progress -> update session state, STOP 4. Completion -> mark task done via TaskUpdate(status="completed"), remove from active_workers 5. Check for checkpoints:
- TESTRUN-* completes -> read meta.json for executor.pass_rate and executor.coverage:
- (pass_rate >= 0.95 AND coverage >= target) OR gc_rounds[layer] >= MAX_GC_ROUNDS -> proceed to handleSpawnNext
- (pass_rate < 0.95 OR coverage < target) AND gc_rounds[layer] < MAX_GC_ROUNDS -> create GC fix tasks, increment gc_rounds[layer]
GC Fix Task Creation (when coverage below target):
TaskCreate({
subject: "TESTGEN-<layer>-fix-<round>: Revise <layer> tests (GC #<round>)",
description: "PURPOSE: Revise tests to fix failures and improve coverage | Success: pass_rate >= 0.95 AND coverage >= target
TASK:
- Read previous test results and failure details
- Revise tests to address failures
- Improve coverage for uncovered areas
CONTEXT:
- Session: <session-folder>
- Layer: <layer>
- Previous results: <session>/results/run-<N>.json
EXPECTED: Revised test files in <session>/tests/<layer>/
CONSTRAINTS: Only modify test files
---
InnerLoop: true
RoleSpec: ~ or <project>/.claude/skills/team-testing/roles/generator/role.md"
})
TaskCreate({
subject: "TESTRUN-<layer>-fix-<round>: Re-execute <layer> (GC #<round>)",
description: "PURPOSE: Re-execute tests after revision | Success: pass_rate >= 0.95
CONTEXT:
- Session: <session-folder>
- Layer: <layer>
- Input: tests/<layer>
EXPECTED: <session>/results/run-<N>-gc.json
---
InnerLoop: true
RoleSpec: ~ or <project>/.claude/skills/team-testing/roles/executor/role.md",
blockedBy: ["TESTGEN-<layer>-fix-<round>"]
})Update session.gc_rounds[layer]++
6. -> handleSpawnNext
handleCheck
Read-only status report, then STOP.
Worker Progress (from message bus):
Before generating status output, read worker milestones:
const progressMsgs = mcp__ccw-tools__team_msg({
operation: "list", session_id: sessionId, type: "progress", last: 50
})
const blockerMsgs = mcp__ccw-tools__team_msg({
operation: "list", session_id: sessionId, type: "blocker", last: 10
})
// Aggregate latest milestone per task
const taskProgress = {}
for (const msg of (progressMsgs.result?.messages || [])) {
const tid = msg.data?.task_id
if (tid && (!taskProgress[tid] || msg.ts > taskProgress[tid].ts)) {
taskProgress[tid] = { phase: msg.data.phase, pct: msg.data.progress_pct, ts: msg.ts }
}
}Include in status output:
- Per-worker latest milestone (phase + progress_pct) next to task status
- Active blockers section (if any blockerMsgs found)
Output:
[coordinator] Testing Pipeline Status
[coordinator] Mode: <pipeline_mode>
[coordinator] Progress: <done>/<total> (<pct>%)
[coordinator] GC Rounds: L1: <n>/3, L2: <n>/3
[coordinator] Pipeline Graph:
STRATEGY-001: <done|run|wait> test-strategy.md
TESTGEN-001: <done|run|wait> generating L1...
TESTRUN-001: <done|run|wait> blocked by TESTGEN-001
TESTGEN-002: <done|run|wait> blocked by TESTRUN-001
TESTRUN-002: <done|run|wait> blocked by TESTGEN-002
TESTANA-001: <done|run|wait> blocked by TESTRUN-*
[coordinator] Active Workers: <list with elapsed time>
[coordinator] Ready: <pending tasks with resolved deps>
[coordinator] Commands: 'resume' to advance | 'check' to refreshThen STOP.
handleResume
1. No active workers -> handleSpawnNext 2. Has active -> check each status
- completed -> mark done via TaskUpdate
- in_progress -> still running
3. Some completed -> handleSpawnNext 4. All running -> report status, STOP
handleSpawnNext
Find ready tasks, spawn workers, STOP.
1. Collect from TaskList():
- completedSubjects: status = completed
- inProgressSubjects: status = in_progress
- readySubjects: status = pending AND all blockedBy in completedSubjects
2. No ready + work in progress -> report waiting, STOP 3. No ready + nothing in progress -> handleComplete 4. Has ready -> for each ready task: a. Determine role from prefix (use Role-Worker Map) b. Check inner_loop: parse task description InnerLoop: field (NOT role.md default)
- InnerLoop: true AND same-role worker already active -> skip (worker picks up next task)
- InnerLoop: false OR no active same-role worker -> spawn new worker
c. TaskUpdate -> in_progress d. team_msg log -> task_unblocked e. Spawn team-worker:
Agent({
subagent_type: "team-worker",
description: "Spawn <role> worker for <subject>",
team_name: "testing",
name: "<role>",
run_in_background: true,
prompt: `## Role Assignment
role: <role>
role_spec: ~ or <project>/.claude/skills/team-testing/roles/<role>/role.md
session: <session-folder>
session_id: <session-id>
team_name: testing
requirement: <task-description>
inner_loop: <true|false>
## Current Task
- Task ID: <task-id>
- Task: <subject>
## Progress Milestones
session_id: <session-id>
Report progress via team_msg at natural phase boundaries (context loaded -> core work done -> verification).
Report blockers immediately via team_msg type="blocker".
Report completion via team_msg type="task_complete" after final SendMessage.
Read role_spec file to load Phase 2-4 domain instructions.
Execute built-in Phase 1 (task discovery) -> role Phase 2-4 -> built-in Phase 5 (report).`
})f. Add to active_workers
5. Parallel spawn (comprehensive pipeline):
- TESTGEN-001 + TESTGEN-002 both unblocked -> spawn both in parallel (name: "generator-1", "generator-2")
- TESTRUN-001 + TESTRUN-002 both unblocked -> spawn both in parallel (name: "executor-1", "executor-2")
6. Update session.json, output summary, STOP
handleComplete
Pipeline done. Generate report and completion action.
1. Verify all tasks (including any GC fix tasks) have status "completed" or "deleted" 2. If any tasks incomplete -> return to handleSpawnNext 3. If all complete:
- Read final state from meta.json (analyst.quality_score, executor.coverage, gc_rounds)
- Generate summary (deliverables, task count, GC rounds, coverage metrics)
4. Read session.completion_action:
- interactive -> AskUserQuestion (Archive/Keep/Deepen Coverage)
- auto_archive -> Archive & Clean (status=completed, TeamDelete)
- auto_keep -> Keep Active (status=paused)
handleAdapt
Capability gap reported mid-pipeline.
1. Parse gap description 2. Check if existing role covers it -> redirect 3. Role count < 5 -> generate dynamic role-spec in <session>/role-specs/ 4. Create new task, spawn worker 5. Role count >= 5 -> merge or pause
Fast-Advance Reconciliation
On every coordinator wake: 1. Read team_msg entries with type="fast_advance" 2. Sync active_workers with spawned successors 3. No duplicate spawns
Phase 4: State Persistence
After every handler execution: 1. Reconcile active_workers with actual TaskList states 2. Remove entries for completed/deleted tasks 3. Write updated session.json 4. STOP (wait for next callback)
Error Handling
| Scenario | Resolution |
|---|---|
| Session file not found | Error, suggest re-initialization |
| Worker callback from unknown role | Log info, scan for other completions |
| GC loop exceeded (3 rounds) | Accept current coverage with warning, proceed |
| Pipeline stall | Check blockedBy chains, report to user |
| Coverage tool unavailable | Degrade to pass rate judgment |
| Worker crash | Reset task to pending, respawn |
Coordinator Role
Orchestrate team-testing: analyze -> dispatch -> spawn -> monitor -> report.
Identity
- Name: coordinator | Tag: [coordinator]
- Responsibility: Change scope analysis -> Create team -> Dispatch tasks -> Monitor progress -> Report results
Boundaries
MUST
- Use
team-workeragent type for all worker spawns (NOTgeneral-purpose) - Follow Command Execution Protocol for dispatch and monitor commands
- Respect pipeline stage dependencies (blockedBy)
- Stop after spawning workers -- wait for callbacks
- Handle Generator-Critic cycles with max 3 iterations per layer
- Execute completion action in Phase 5
MUST NOT
- Implement domain logic (test generation, execution, analysis) -- workers handle this
- Spawn workers without creating tasks first
- Skip quality gates when coverage is below target
- Modify test files or source code directly -- delegate to workers
- Force-advance pipeline past failed GC loops
Command Execution Protocol
When coordinator needs to execute a specific phase: 1. Read commands/<command>.md 2. Follow the workflow defined in the command 3. Commands are inline execution guides, NOT separate agents 4. Execute synchronously, complete before proceeding
Entry Router
| Detection | Condition | Handler |
|---|---|---|
| Worker callback | Message contains [strategist], [generator], [executor], [analyst] | -> handleCallback (monitor.md) |
| Status check | Args contain "check" or "status" | -> handleCheck (monitor.md) |
| Manual resume | Args contain "resume" or "continue" | -> handleResume (monitor.md) |
| Capability gap | Message contains "capability_gap" | -> handleAdapt (monitor.md) |
| Pipeline complete | All tasks completed | -> handleComplete (monitor.md) |
| Interrupted session | Active session in .workflow/.team/TST-* | -> Phase 0 |
| New session | None of above | -> Phase 1 |
For callback/check/resume/adapt/complete: load @commands/monitor.md, execute handler, STOP.
Phase 0: Session Resume Check
1. Scan .workflow/.team/TST-*/session.json for active/paused sessions 2. No sessions -> Phase 1 3. Single session -> reconcile (audit TaskList, reset in_progress->pending, rebuild team, kick first ready task) 4. Multiple -> AskUserQuestion for selection
Phase 1: Requirement Clarification
TEXT-LEVEL ONLY. No source code reading.
1. Parse task description from $ARGUMENTS 2. Analyze change scope:
Bash("git diff --name-only HEAD~1 2>/dev/null || git diff --name-only --cached")3. Select pipeline:
| Condition | Pipeline |
|---|---|
| fileCount <= 3 AND moduleCount <= 1 | targeted |
| fileCount <= 10 AND moduleCount <= 3 | standard |
| Otherwise | comprehensive |
4. Clarify if ambiguous (AskUserQuestion for scope) 5. Delegate to @commands/analyze.md 6. Output: task-analysis.json 7. CRITICAL: Always proceed to Phase 2, never skip team workflow
Phase 2: Create Team + Initialize Session
1. Resolve workspace paths (MUST do first):
project_root= result ofBash({ command: "pwd" })skill_root=<project_root>/.claude/skills/team-testing
2. Generate session ID: TST-<slug>-<date> 3. Create session folder structure (strategy/, tests/L1-unit/, tests/L2-integration/, tests/L3-e2e/, results/, analysis/, wisdom/) 4. TeamCreate with team name "testing" 5. Read specs/pipelines.md -> select pipeline based on mode 6. Initialize pipeline via team_msg state_update:
mcp__ccw-tools__team_msg({
operation: "log", session_id: "<id>", from: "coordinator",
type: "state_update", summary: "Session initialized",
data: {
pipeline_mode: "<targeted|standard|comprehensive>",
pipeline_stages: ["strategist", "generator", "executor", "analyst"],
team_name: "testing",
coverage_targets: { "L1": 80, "L2": 60, "L3": 40 },
gc_rounds: {}
}
})7. Write session.json
Phase 3: Create Task Chain
Delegate to @commands/dispatch.md: 1. Read specs/pipelines.md for selected pipeline's task registry 2. Topological sort tasks 3. Create tasks via TaskCreate with blockedBy 4. Update session.json
Phase 4: Spawn-and-Stop
Delegate to @commands/monitor.md#handleSpawnNext: 1. Find ready tasks (pending + blockedBy resolved) 2. Spawn team-worker agents (see SKILL.md Spawn Template) 3. Output status summary 4. STOP
Phase 5: Report + Completion Action
1. Generate summary (deliverables, pipeline stats, GC rounds, coverage metrics) 2. Execute completion action per session.completion_action:
- interactive -> AskUserQuestion (Archive/Keep/Deepen Coverage)
- auto_archive -> Archive & Clean
- auto_keep -> Keep Active
Error Handling
| Error | Resolution |
|---|---|
| Task too vague | AskUserQuestion for clarification |
| Session corruption | Attempt recovery, fallback to manual |
| Worker crash | Reset task to pending, respawn |
| Dependency cycle | Detect in analysis, halt |
| GC loop exceeded (3 rounds) | Accept current coverage, log to wisdom, proceed |
| Coverage tool unavailable | Degrade to pass rate judgment |
Test Executor
inner_loop: dynamic — Dispatch sets per-task:truefor serial GC loops (single layer pipeline),falsefor parallel layer execution (comprehensive pipeline where TESTRUN-001 and TESTRUN-002 run independently). When false, each TESTRUN task gets its own worker.
Execute tests, collect coverage, attempt auto-fix for failures. Acts as the Critic in the Generator-Critic loop. Reports pass rate and coverage for coordinator GC decisions.
Phase 2: Context Loading
| Input | Source | Required |
|---|---|---|
| Task description | From task subject/description | Yes |
| Session path | Extracted from task description | Yes |
| Test directory | Task description (Input: <path>) | Yes |
| Coverage target | Task description (default: 80%) | Yes |
| .msg/meta.json | <session>/wisdom/.msg/meta.json | No |
1. Extract session path and test directory from task description 2. Load test specs: Run ccw spec load --category test for test framework conventions and coverage targets 3. Extract coverage target (default: 80%) 3. Read .msg/meta.json for framework info (from strategist namespace) 4. Determine test framework:
| Framework | Run Command |
|---|---|
| Jest | npx jest --coverage --json --outputFile=<session>/results/jest-output.json |
| Pytest | python -m pytest --cov --cov-report=json:<session>/results/coverage.json -v |
| Vitest | npx vitest run --coverage --reporter=json |
5. Find test files to execute:
Glob("<session>/<test-dir>/**/*")Phase 3: Test Execution + Fix Cycle
Iterative test-fix cycle (max 3 iterations):
| Step | Action |
|---|---|
| 1 | Run test command |
| 2 | Parse results: pass rate + coverage |
| 3 | pass_rate >= 0.95 AND coverage >= target -> success, exit |
| 4 | Extract failing test details |
| 5 | Delegate fix to CLI tool (gemini write mode) |
| 6 | Increment iteration; >= 3 -> exit with failures |
Bash("<test-command> 2>&1 || true")Auto-fix delegation (on failure):
Bash({
command: `ccw cli -p "PURPOSE: Fix test failures to achieve pass rate >= 0.95; success = all tests pass
TASK: • Analyze test failure output • Identify root causes • Fix test code only (not source) • Preserve test intent
MODE: write
CONTEXT: @<session>/<test-dir>/**/* | Memory: Test framework: <framework>, iteration <N>/3
EXPECTED: Fixed test files with: corrected assertions, proper async handling, fixed imports, maintained coverage
CONSTRAINTS: Only modify test files | Preserve test structure | No source code changes
Test failures:
<test-output>" --tool gemini --mode write --cd <session>`,
run_in_background: false
})Save results: <session>/results/run-<N>.json
Phase 4: Defect Pattern Extraction & State Update
Extract defect patterns from failures:
| Pattern Type | Detection Keywords |
|---|---|
| Null reference | "null", "undefined", "Cannot read property" |
| Async timing | "timeout", "async", "await", "promise" |
| Import errors | "Cannot find module", "import" |
| Type mismatches | "type", "expected", "received" |
Record effective test patterns (if pass_rate > 0.8):
| Pattern | Detection |
|---|---|
| Happy path | "should succeed", "valid input" |
| Edge cases | "edge", "boundary", "limit" |
| Error handling | "should fail", "error", "throw" |
Update <session>/wisdom/.msg/meta.json under executor namespace:
- Merge
{ "executor": { pass_rate, coverage, defect_patterns, effective_patterns, coverage_history_entry } }
Test Generator
Generate test code by layer (L1 unit / L2 integration / L3 E2E). Acts as the Generator in the Generator-Critic loop. Supports revision mode for GC loop iterations.
Phase 2: Context Loading
| Input | Source | Required |
|---|---|---|
| Task description | From task subject/description | Yes |
| Session path | Extracted from task description | Yes |
| Test strategy | <session>/strategy/test-strategy.md | Yes |
| .msg/meta.json | <session>/wisdom/.msg/meta.json | No |
1. Extract session path and layer from task description 2. Load test specs: Run ccw spec load --category test for test framework conventions and coverage targets 3. Read test strategy:
Read("<session>/strategy/test-strategy.md")3. Read source files to test (from strategy priority_files, limit 20) 4. Read .msg/meta.json for framework and scope context
5. Detect revision mode:
| Condition | Mode |
|---|---|
| Task subject contains "fix" or "revised" | Revision -- load previous failures |
| Otherwise | Fresh generation |
For revision mode:
- Read latest result file for failure details
- Read effective test patterns from .msg/meta.json
6. Read wisdom files if available
Phase 3: Test Generation
Strategy selection by complexity:
| File Count | Strategy |
|---|---|
| <= 3 files | Direct: inline Write/Edit |
| 3-5 files | Single code-developer agent |
| > 5 files | Batch: group by module, one agent per batch |
Direct generation (per source file): 1. Generate test path: <session>/tests/<layer>/<test-file> 2. Generate test code: happy path, edge cases, error handling 3. Write test file
CLI delegation (medium/high complexity):
Bash({
command: `ccw cli -p "PURPOSE: Generate <layer> tests using <framework> to achieve coverage target; success = all priority files covered with quality tests
TASK: • Analyze source files • Generate test cases (happy path, edge cases, errors) • Write test files with proper structure • Ensure import resolution
MODE: write
CONTEXT: @<source-files> @<session>/strategy/test-strategy.md | Memory: Framework: <framework>, Layer: <layer>, Round: <round>
<if-revision: Previous failures: <failure-details>
Effective patterns: <patterns-from-meta>>
EXPECTED: Test files in <session>/tests/<layer>/ with: proper test structure, comprehensive coverage, correct imports, framework conventions
CONSTRAINTS: Follow test strategy priorities | Use framework best practices | <layer>-appropriate assertions
Source files to test:
<file-list-with-content>" --tool gemini --mode write --cd <session>`,
run_in_background: false
})Output verification:
Glob("<session>/tests/<layer>/**/*")Phase 4: Self-Validation & State Update
Validation checks:
| Check | Method | Action on Fail |
|---|---|---|
| Syntax | tsc --noEmit or equivalent | Auto-fix imports/types |
| File count | Count generated files | Report issue |
| Import resolution | Check broken imports | Fix import paths |
Update <session>/wisdom/.msg/meta.json under generator namespace:
- Merge
{ "generator": { test_files, layer, round, is_revision } }
Test Strategist
Analyze git diff, determine test layers, define coverage targets, and formulate test strategy with prioritized execution order.
Phase 2: Context & Environment Detection
| Input | Source | Required |
|---|---|---|
| Task description | From task subject/description | Yes |
| Session path | Extracted from task description | Yes |
| .msg/meta.json | <session>/wisdom/.msg/meta.json | No |
1. Extract session path and scope from task description 2. Get git diff for change analysis:
Bash("git diff HEAD~1 --name-only 2>/dev/null || git diff --cached --name-only")
Bash("git diff HEAD~1 -- <changed-files> 2>/dev/null || git diff --cached -- <changed-files>")3. Detect test framework from project files:
| Signal File | Framework | Test Pattern |
|---|---|---|
| jest.config.js/ts | Jest | **/*.test.{ts,tsx,js} |
| vitest.config.ts/js | Vitest | **/*.test.{ts,tsx} |
| pytest.ini / pyproject.toml | Pytest | **/test_*.py |
| No detection | Default | Jest patterns |
4. Scan existing test patterns:
Glob("**/*.test.*")
Glob("**/*.spec.*")5. Read .msg/meta.json if exists for session context
Phase 3: Strategy Formulation
Change analysis dimensions:
| Change Type | Analysis | Priority |
|---|---|---|
| New files | Need new tests | High |
| Modified functions | Need updated tests | Medium |
| Deleted files | Need test cleanup | Low |
| Config changes | May need integration tests | Variable |
Strategy output structure:
1. Change Analysis Table: File, Change Type, Impact, Priority 2. Test Layer Recommendations:
- L1 Unit: Scope, Coverage Target, Priority Files, Patterns
- L2 Integration: Scope, Coverage Target, Integration Points
- L3 E2E: Scope, Coverage Target, User Scenarios
3. Risk Assessment: Risk, Probability, Impact, Mitigation 4. Test Execution Order: Prioritized sequence
Write strategy to <session>/strategy/test-strategy.md
Self-validation:
| Check | Criteria | Fallback |
|---|---|---|
| Has L1 scope | L1 scope not empty | Default to all changed files |
| Has coverage targets | L1 target > 0 | Use defaults (80/60/40) |
| Has priority files | List not empty | Use all changed files |
Phase 4: Wisdom & State Update
1. Write discoveries to <session>/wisdom/conventions.md (detected framework, patterns) 2. Update <session>/wisdom/.msg/meta.json under strategist namespace:
- Read existing -> merge
{ "strategist": { framework, layers, coverage_targets, priority_files, risks } }-> write back
Testing Pipelines
Pipeline definitions and task registry for team-testing.
Pipeline Selection
| Condition | Pipeline |
|---|---|
| fileCount <= 3 AND moduleCount <= 1 | targeted |
| fileCount <= 10 AND moduleCount <= 3 | standard |
| Otherwise | comprehensive |
Pipeline Definitions
Targeted Pipeline (3 tasks, serial)
STRATEGY-001 -> TESTGEN-001 -> TESTRUN-001| Task ID | Role | Dependencies | Layer | Description |
|---|---|---|---|---|
| STRATEGY-001 | strategist | (none) | — | Analyze changes, define test strategy |
| TESTGEN-001 | generator | STRATEGY-001 | L1 | Generate L1 unit tests |
| TESTRUN-001 | executor | TESTGEN-001 | L1 | Execute L1 tests, collect coverage |
Standard Pipeline (6 tasks, progressive layers)
STRATEGY-001 -> TESTGEN-001 -> TESTRUN-001 -> TESTGEN-002 -> TESTRUN-002 -> TESTANA-001| Task ID | Role | Dependencies | Layer | Description |
|---|---|---|---|---|
| STRATEGY-001 | strategist | (none) | — | Analyze changes, define test strategy |
| TESTGEN-001 | generator | STRATEGY-001 | L1 | Generate L1 unit tests |
| TESTRUN-001 | executor | TESTGEN-001 | L1 | Execute L1 tests, collect coverage |
| TESTGEN-002 | generator | TESTRUN-001 | L2 | Generate L2 integration tests |
| TESTRUN-002 | executor | TESTGEN-002 | L2 | Execute L2 tests, collect coverage |
| TESTANA-001 | analyst | TESTRUN-002 | — | Defect pattern analysis, quality report |
Comprehensive Pipeline (8 tasks, parallel windows)
STRATEGY-001 -> [TESTGEN-001 || TESTGEN-002] -> [TESTRUN-001 || TESTRUN-002] -> TESTGEN-003 -> TESTRUN-003 -> TESTANA-001| Task ID | Role | Dependencies | Layer | Description |
|---|---|---|---|---|
| STRATEGY-001 | strategist | (none) | — | Analyze changes, define test strategy |
| TESTGEN-001 | generator-1 | STRATEGY-001 | L1 | Generate L1 unit tests (parallel) |
| TESTGEN-002 | generator-2 | STRATEGY-001 | L2 | Generate L2 integration tests (parallel) |
| TESTRUN-001 | executor-1 | TESTGEN-001 | L1 | Execute L1 tests (parallel) |
| TESTRUN-002 | executor-2 | TESTGEN-002 | L2 | Execute L2 tests (parallel) |
| TESTGEN-003 | generator | TESTRUN-001, TESTRUN-002 | L3 | Generate L3 E2E tests |
| TESTRUN-003 | executor | TESTGEN-003 | L3 | Execute L3 tests, collect coverage |
| TESTANA-001 | analyst | TESTRUN-003 | — | Defect pattern analysis, quality report |
GC Loop (Generator-Critic)
Generator and executor iterate per test layer:
TESTGEN -> TESTRUN -> (if pass_rate < 0.95 OR coverage < target) -> TESTGEN-fix -> TESTRUN-fix
(if pass_rate >= 0.95 AND coverage >= target) -> next layer or TESTANA- Max iterations: 3 per layer
- After 3 iterations: accept current state with warning
Coverage Targets
| Layer | Name | Default Target |
|---|---|---|
| L1 | Unit Tests | 80% |
| L2 | Integration Tests | 60% |
| L3 | E2E Tests | 40% |
Session Directory
.workflow/.team/TST-<slug>-<YYYY-MM-DD>/
├── .msg/messages.jsonl # Message bus log
├── .msg/meta.json # Session metadata
├── wisdom/ # Cross-task knowledge
│ ├── learnings.md
│ ├── decisions.md
│ ├── conventions.md
│ └── issues.md
├── strategy/ # Strategist output
│ └── test-strategy.md
├── tests/ # Generator output
│ ├── L1-unit/
│ ├── L2-integration/
│ └── L3-e2e/
├── results/ # Executor output
│ ├── run-001.json
│ └── coverage-001.json
└── analysis/ # Analyst output
└── quality-report.md{
"team_name": "team-testing",
"team_display_name": "Team Testing",
"description": "Testing team with Generator-Critic loop, shared defect memory, and progressive test layers",
"version": "1.0.0",
"roles": {
"coordinator": {
"task_prefix": null,
"responsibility": "Change scope analysis, layer selection, quality gating",
"message_types": ["pipeline_selected", "gc_loop_trigger", "quality_gate", "task_unblocked", "error", "shutdown"]
},
"strategist": {
"task_prefix": "STRATEGY",
"responsibility": "Analyze git diff, determine test layers, define coverage targets",
"message_types": ["strategy_ready", "error"]
},
"generator": {
"task_prefix": "TESTGEN",
"responsibility": "Generate test cases by layer (unit/integration/E2E)",
"message_types": ["tests_generated", "tests_revised", "error"]
},
"executor": {
"task_prefix": "TESTRUN",
"responsibility": "Execute tests, collect coverage, auto-fix failures",
"message_types": ["tests_passed", "tests_failed", "coverage_report", "error"]
},
"analyst": {
"task_prefix": "TESTANA",
"responsibility": "Defect pattern analysis, coverage gap analysis, quality report",
"message_types": ["analysis_ready", "error"]
}
},
"pipelines": {
"targeted": {
"description": "Small scope: strategy → generate L1 → run",
"task_chain": ["STRATEGY-001", "TESTGEN-001", "TESTRUN-001"],
"gc_loops": 0
},
"standard": {
"description": "Progressive: L1 → L2 with analysis",
"task_chain": ["STRATEGY-001", "TESTGEN-001", "TESTRUN-001", "TESTGEN-002", "TESTRUN-002", "TESTANA-001"],
"gc_loops": 1
},
"comprehensive": {
"description": "Full coverage: parallel L1+L2, then L3 with analysis",
"task_chain": ["STRATEGY-001", "TESTGEN-001", "TESTGEN-002", "TESTRUN-001", "TESTRUN-002", "TESTGEN-003", "TESTRUN-003", "TESTANA-001"],
"gc_loops": 2,
"parallel_groups": [["TESTGEN-001", "TESTGEN-002"], ["TESTRUN-001", "TESTRUN-002"]]
}
},
"innovation_patterns": {
"generator_critic": {
"generator": "generator",
"critic": "executor",
"max_rounds": 3,
"convergence_trigger": "coverage >= target && pass_rate >= 0.95"
},
"shared_memory": {
"file": "shared-memory.json",
"fields": {
"strategist": "test_strategy",
"generator": "generated_tests",
"executor": "execution_results",
"analyst": "analysis_report"
},
"persistent_fields": ["defect_patterns", "effective_test_patterns", "coverage_history"]
},
"dynamic_pipeline": {
"selector": "coordinator",
"criteria": "changed_file_count + module_count + change_type"
}
},
"test_layers": {
"L1": { "name": "Unit Tests", "coverage_target": 80, "description": "Function-level isolation tests" },
"L2": { "name": "Integration Tests", "coverage_target": 60, "description": "Module interaction tests" },
"L3": { "name": "E2E Tests", "coverage_target": 40, "description": "User scenario end-to-end tests" }
},
"collaboration_patterns": ["CP-1", "CP-3", "CP-5"],
"session_dirs": {
"base": ".workflow/.team/TST-{slug}-{YYYY-MM-DD}/",
"strategy": "strategy/",
"tests": "tests/",
"results": "results/",
"analysis": "analysis/",
"messages": ".workflow/.team-msg/{team-name}/"
}
}