
Comfyui Prompt Interview
- 246 installs
- 85 repo stars
- Updated March 18, 2026
- mckruz/comfyui-expert
Run structured prompt interviews inside ComfyUI workflows to elicit style, subject, and constraint details before generating images or video nodes.
About
ComfyUI Prompt Interview skill structures conversational prompt discovery for generative image and video workflows. It guides agents through subject, style, lighting, composition, and negative constraints, then maps responses into ComfyUI node parameters for reliable, repeatable creative output without guesswork.
- Multi-turn creative requirement gathering
- Mapping answers to ComfyUI node inputs
- Style, composition, and negative prompt templates
- Constraint validation before graph execution
- Reusable interview flows for batch generation
Comfyui Prompt Interview by the numbers
- 246 all-time installs (skills.sh)
- +19 installs in the week ending Aug 4, 2026 (Skillselion tracking)
- Ranked #559 of 1,335 Generative Media skills by installs in the Skillselion catalog
- Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/mckruz/comfyui-expert --skill comfyui-prompt-interviewAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 246 |
|---|---|
| repo stars | ★ 85 |
| Last updated | March 18, 2026 |
| Repository | mckruz/comfyui-expert ↗ |
What it does
Run structured prompt interviews inside ComfyUI workflows to elicit style, subject, and constraint details before generating images or video nodes.
Files
ComfyUI Prompt Interview
Conduct a guided conversation to draw out the user's complete creative vision, then synthesize a perfect, model-appropriate prompt with all recommended settings.
When to Invoke This Skill
- User describes an image or scene idea but hasn't given enough detail for a quality prompt
- User says "help me think through what I want to create"
- User has a vague concept that needs refinement
- User wants a structured prompt but isn't sure what to specify
The Interview Philosophy
Ask, don't interrogate. This is a conversation, not a form. Ask one or two questions at a time. Listen to what the user gives you and follow up on what's missing. Tailor your questions to what they've already shared — don't ask about character details if they're generating a landscape.
Fewer questions = better. Aim for 4-7 exchanges maximum. Ask the most impactful questions first. Stop asking when you have enough to generate an excellent prompt.
Don't ask for what you can infer. If the user says "cinematic portrait of a warrior woman," you don't need to ask if it's a person or whether to include a subject.
---
Interview Flow
Step 1: Open with the Big Picture
If the user hasn't told you what they want to create, start here:
"What do you want to create? Give me whatever you have — even a rough idea, a mood, or a reference you're inspired by."
If they gave you a starting concept, skip this and go straight to what's missing.
---
Step 2: Branch by Creation Type
Based on their answer, determine what kind of generation this is:
| Type | Key Questions to Ask |
|---|---|
| Portrait / Character | Identity method? Existing character? Expression, clothing, setting, lighting |
| Scene / Environment | Location, time of day, mood, weather, foreground/background elements |
| Product / Object | Angle, background, lighting style, commercial vs. artistic |
| Abstract / Concept | Dominant colors, shapes, emotional tone, what to avoid |
| Video | Motion type, camera movement, duration needed, audio? |
---
Step 3: Ask the High-Impact Questions
Ask only what's missing. Use natural conversational language, not a bullet list.
For character/portrait content — ask in order of impact:
1. Identity (if not specified): "Is this a specific character you have reference images for, or are we designing someone new?" 2. Expression & mood: "What's the emotion or energy — fierce, serene, playful, haunted?" 3. Setting: "Where are they, and when? (Time of day, location, interior/exterior)" 4. Lighting: "Any specific lighting in mind? (Golden hour, dramatic side light, soft studio, neon, candlelight)" 5. Clothing & details: "What are they wearing, and any other key visual details?" 6. Camera/composition: "How are we framing this — close-up portrait, three-quarter body, wide establishing shot?" 7. Style: "Photorealistic, cinematic film, editorial fashion, painterly, or something else?"
For scene/environment content — ask in order of impact:
1. Setting: "Describe the place — what does it look like, and when is it?" 2. Mood/atmosphere: "What feeling should hit the viewer instantly?" 3. Lighting: "What's the light source and quality?" 4. Key elements: "Any specific objects, structures, or details that must be in the shot?" 5. Style: "Photorealistic, stylized, concept art, painterly?"
For video content — additional questions:
1. Motion: "What's moving — the subject, the camera, or both?" 2. Duration: "How long? (Short: 3-5s vs. long: 15-60s changes model choice)" 3. Audio: "Do you need sound/music, or silent?"
---
Step 4: Technical Questions (ask only if not obvious)
These can usually be inferred from context, but ask if unclear:
- Aspect ratio: "Standard 1:1 portrait, 16:9 cinematic, 9:16 vertical/social?"
- Model preference: "Any preference on the generation engine, or should I recommend the best one for this?"
- Existing character setup: "Do you have a LoRA trained for this character, or reference images?"
- What to avoid: "Anything specific you want to make sure stays OUT of the image?"
---
Step 5: Confirm and Synthesize
Before generating the prompt, briefly reflect back the vision:
"Got it. Here's how I'm reading this: [1-2 sentence summary of the concept]. Let me build that prompt."
Then immediately generate the full output below.
---
Output Format
Deliver all four components, clearly separated:
---
🎯 Positive Prompt
[Craft the positive prompt applying model-specific rules from skills/comfyui-prompt-engineer/SKILL.md]
Key rules:
- FLUX / Kontext: Natural language, 50-100 words, no quality tags, describe the scene not the face if using identity method
- SDXL: Quality tags first, trigger word second, 50-150 words, weighted syntax supported
- SD 1.5: Short and tag-based, 30-80 words
- Wan / Video: Concise, motion-focused, 20-50 words
- If a LoRA trigger word applies, put it first
- If using InstantID/InfiniteYou: don't describe facial features, let the identity method handle them
---
🚫 Negative Prompt
[Select the appropriate negative template and customize it]
Standard templates:
- Photorealism:
(worst quality:1.4), (low quality:1.4), blurry, deformed, bad anatomy, bad hands, extra fingers, missing fingers, text, watermark, 3d render, cartoon, anime, plastic skin, airbrushed, oversaturated - FLUX (minimal):
blurry, low quality, distorted, deformed, ugly, watermark, text - Video:
static, frozen, jerky motion, low quality, blurry, distorted face, bad anatomy, glitch, artifacts, flickering
---
⚙️ Recommended Settings
| Parameter | Value | Reason |
|---|---|---|
| Model | [Specific checkpoint] | [Why this model] |
| Sampler | [e.g., DPM++ 2M Karras] | |
| Steps | [e.g., 25] | |
| CFG Scale | [e.g., 4.5] | |
| Resolution | [e.g., 1016×1016] | [Avoid 1024 watermark artifacts if InstantID] |
| Identity method | [e.g., InfiniteYou SIM] | [If applicable] |
| LoRA | [Trigger word + weight] | [If applicable] |
---
🔧 Pipeline Recommendation
[1-3 sentences describing the recommended workflow pattern and why]
Example: "Use Pattern 2 from the character-gen skill: Load your LoRA, add InfiniteYou SIM for identity lock, route through FaceDetailer at 0.35 denoise, then upscale with 4x-UltraSharp. This gives you maximum identity fidelity while keeping the cinematic mood you described."
---
Decision Rules for Model Selection
Apply these when recommending a model/pipeline:
| User's Goal | Recommend |
|---|---|
| "Photorealistic, single character, reference images available" | InfiniteYou or FLUX Kontext |
| "Multiple characters, consistent identity" | PuLID Flux II |
| "Fast iteration, character editing" | FLUX Kontext dev/pro |
| "Highest quality portrait, trained LoRA exists" | FLUX.1-dev + LoRA |
| "Wide scene / environment" | FLUX.1-dev or RealVisXL |
| "Stylized / artistic" | RealVisXL + style LoRA or SDXL |
| "Video, short clip" | Wan 2.2 1.3B (consumer) or LTX-2 (4K) |
| "Video, infinite length" | Wan 2.2 I2V A14B + SVI 2.0 Pro |
| "Talking head" | FramePack + LatentSync 1.6 |
---
Example Interview
User: "I want to create a dramatic image of a woman warrior standing in ruins at dawn."
VideoAgent:
"Love the concept. A few things that'll lock in the look:
1. Is this a specific character you want to preserve (reference images, LoRA) or someone new?
2. What's her energy — defiant and triumphant, exhausted but unbroken, something else?
3. Any specifics on how she looks — hair, armor style, anything iconic about her?"
User: "New character. Defiant and triumphant. Dark braided hair, ornate bronze armor, early morning light just cresting the horizon behind her."
VideoAgent:
"Got it — framing this as a heroic aftermath shot: she's won, the sun rises behind her, rubble around her feet. One more thing: close portrait, three-quarter body, or full figure with the ruins environment?"
User: "Full figure, show the ruins and the dawn sky."
VideoAgent:
"Perfect. Building that now."
[Delivers complete prompt output]
---
Integration with Other Skills
After delivering the prompt, offer next steps:
- Generate the workflow: "Want me to build the full ComfyUI workflow for this? I'll check your inventory and configure every node."
- Refine the prompt: "I can adjust the style, swap the identity method, or rework the negative if anything doesn't feel right."
- Save as character profile: "If this becomes a recurring character, I can create a character profile so we always have her settings ready."
skill_type: encoded_preference
baseline_expected_score: "15%"
with_skill_target_score: "90%"
token_overhead_acceptable: "30%"
manual_correction_reduction: "60%"
test_case_count: 7
happy_path_cases: 3
failure_mode_cases: 2
comparison_criteria:
- criterion: "Asks clarifying questions BEFORE generating any prompts (interview-first sequence)"
weight: 0.35
assertion_types: [question_before_code, regex, not_contains]
anchor: "question mark presence and absence of prompt output markers"
- criterion: "Interview length is appropriate (4-7 exchanges, not too short or too long)"
weight: 0.15
assertion_types: [structure_check]
anchor: "semantic — requires multi-turn eval harness"
- criterion: "Prompt formatting matches the target model (SDXL tags vs FLUX natural language)"
weight: 0.20
assertion_types: [not_contains, regex]
anchor: "anti-pattern exclusion and sentence structure regex"
- criterion: "Complete output package (positive prompt, negative prompt, settings table, pipeline recommendation)"
weight: 0.15
assertion_types: [regex]
anchor: "keyword presence for prompt sections and settings parameters"
- criterion: "Questions cover essential dimensions without redundancy"
weight: 0.15
assertion_types: [structure_check, regex]
anchor: "partial — pipeline keyword anchored, dimension coverage is semantic"
comfyui-prompt-interview — Eval Configuration
Classification
- Type: Encoded Preference
- Category: Guided conversational interview for creative vision extraction before prompt synthesis
What "Good" Looks Like
1. ASKS QUESTIONS FIRST — does not generate prompts until the user's vision is understood (sequence is interview THEN generate) 2. Interview length is appropriate: 4-7 exchanges max, not dragging on or cutting short 3. Prompts are formatted for the target model (SDXL tag-style vs. FLUX natural language vs. SD1.5 weighted tokens) 4. Negative prompt uses the correct template for the model (SDXL negatives differ from SD1.5) 5. Output includes recommended settings table (sampler, steps, CFG, resolution) and pipeline recommendation
Known Limitations
- Cannot assess subjective prompt quality — only structural correctness and interview flow
- User engagement quality varies; skill can't force good answers from terse users
- Model-specific prompt optimization evolves as models update
Benchmark Strategy
- Without skill: Base Claude dumps a prompt immediately when asked to "help me create an image," skipping discovery of the user's actual vision
- With skill: Conducts a structured interview to understand subject, mood, style, technical preferences BEFORE generating model-appropriate prompts
- Key differentiator: The interview-first sequence — encoded preference that creative vision must be understood before prompt generation begins
Security — Eval Sandboxing
Eval runs use real tool access and may expose secrets in output. Results are gitignored. Use --allowedTools "Read,Glob,Grep" to prevent modification during eval runs.
Running Evals
bash eval/run-eval.sh # Full run (with-skill + baseline)
bash eval/run-eval.sh --skill-only # With-skill only
bash eval/run-eval.sh --case TC-001 # Single test caseRetirement Signal
N/A — This is an encoded preference for conversational workflow. Even if base Claude improves at prompt writing, the interview-first behavior is a deliberate UX choice that won't be absorbed into default model behavior.
# Eval results may contain secrets from real config files.
*
!.gitignore
!.gitkeep
#!/usr/bin/env bash
# ═══════════════════════════════════════════════════════════════════════
# Skill Eval Runner — Shared Template v2.0
# ═══════════════════════════════════════════════════════════════════════
# Runs test cases against a skill and captures results for scoring.
# Copy this file into any skill's eval/ directory and set SKILL_NAME.
#
# Usage:
# bash eval/run-eval.sh # Full run (with-skill + baseline)
# bash eval/run-eval.sh --skill-only # Skip baseline comparison
# bash eval/run-eval.sh --case TC-001 # Single test case
# bash eval/run-eval.sh --baseline-only # Baseline only (no skill)
# bash eval/run-eval.sh --score-only # Score existing results (no new runs)
#
# Assertion types supported:
# contains(target) — response includes target (case-insensitive)
# not_contains(target) — response does NOT include target
# regex(pattern) — response matches extended regex
# question_before_code — a "?" appears before first ``` fence
# json_valid — response has a parseable JSON block
# json_fields(f1,f2,...) — JSON block contains required field names
# token_limit(N) — response under ~N tokens (estimated from words)
# range_check(expr) — evaluates numeric expression on JSON output
# word_count(field,min,max)— checks word count of a JSON field
# ═══════════════════════════════════════════════════════════════════════
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
PROJECT_DIR="$(dirname "$SCRIPT_DIR")"
TIMESTAMP=$(date +%Y%m%d-%H%M%S)
RESULTS_DIR="$SCRIPT_DIR/results/$TIMESTAMP"
# ── CONFIGURE THIS ──────────────────────────────────────────────────
# Override by setting SKILL_NAME env var or editing this default.
SKILL_NAME="${SKILL_NAME:-$(basename "$PROJECT_DIR")}"
# ────────────────────────────────────────────────────────────────────
# Parse args
RUN_SKILL=true
RUN_BASELINE=true
SCORE_ONLY=false
SINGLE_CASE=""
while [[ $# -gt 0 ]]; do
case "$1" in
--skill-only) RUN_BASELINE=false; shift ;;
--baseline-only) RUN_SKILL=false; shift ;;
--score-only) SCORE_ONLY=true; shift ;;
--case) SINGLE_CASE="$2"; shift 2 ;;
*) echo "Unknown arg: $1"; exit 1 ;;
esac
done
mkdir -p "$RESULTS_DIR/with-skill" "$RESULTS_DIR/baseline" "$RESULTS_DIR/prompts"
echo "=== Skill Eval Runner v2.0 ==="
echo "Skill: $SKILL_NAME"
echo "Timestamp: $TIMESTAMP"
echo "Results: $RESULTS_DIR"
echo ""
# ─── Extract test cases from YAML ──────────────────────────────────
extract_cases() {
local yaml_file="$SCRIPT_DIR/test-cases.yaml"
local current_id=""
local current_prompt=""
local in_prompt=false
local ids=()
while IFS= read -r line || [[ -n "$line" ]]; do
if [[ "$line" =~ ^-\ id:\ (.+) ]]; then
if [[ -n "$current_id" ]]; then
ids+=("$current_id")
printf '%s' "$current_prompt" > "$RESULTS_DIR/prompts/${current_id}.txt"
fi
current_id="${BASH_REMATCH[1]}"
current_prompt=""
in_prompt=false
fi
if [[ "$line" =~ ^\ \ prompt:\ \"(.+)\"$ ]]; then
current_prompt="${BASH_REMATCH[1]}"
in_prompt=false
fi
if [[ "$line" =~ ^\ \ prompt:\ [\|>]$ ]]; then
in_prompt=true
current_prompt=""
continue
fi
if $in_prompt; then
if [[ "$line" =~ ^\ \ [a-z] && ! "$line" =~ ^\ \ \ \ ]]; then
in_prompt=false
else
local stripped="${line# }"
current_prompt+="${stripped}"$'\n'
fi
fi
done < "$yaml_file"
if [[ -n "$current_id" ]]; then
ids+=("$current_id")
printf '%s' "$current_prompt" > "$RESULTS_DIR/prompts/${current_id}.txt"
fi
echo "${ids[@]}"
}
CASE_IDS=($(extract_cases))
echo "Found ${#CASE_IDS[@]} test cases: ${CASE_IDS[*]}"
if [[ -n "$SINGLE_CASE" ]]; then
CASE_IDS=("$SINGLE_CASE")
echo "Filtering to: $SINGLE_CASE"
fi
echo ""
# ─── Run test cases ────────────────────────────────────────────────
run_case() {
local case_id="$1"
local mode="$2"
local prompt_file="$RESULTS_DIR/prompts/${case_id}.txt"
local output_file="$RESULTS_DIR/$mode/${case_id}.md"
local prompt_text
prompt_text=$(cat "$prompt_file")
echo " [$mode] Running $case_id..."
if [[ "$mode" == "with-skill" ]]; then
local full_prompt="Use the $SKILL_NAME skill to answer this: $prompt_text"
claude -p "$full_prompt" \
--allowedTools "Read,Glob,Grep" \
--max-turns 3 \
--output-format text \
> "$output_file" 2>/dev/null || {
echo "EVAL_ERROR: claude command failed for $case_id ($mode)" > "$output_file"
}
else
claude -p "$prompt_text" \
--allowedTools "Read,Glob,Grep" \
--max-turns 3 \
--output-format text \
> "$output_file" 2>/dev/null || {
echo "EVAL_ERROR: claude command failed for $case_id ($mode)" > "$output_file"
}
fi
local wc_out
wc_out=$(wc -w < "$output_file" | tr -d ' ')
echo " [$mode] $case_id complete ($wc_out words)"
}
if ! $SCORE_ONLY; then
if $RUN_SKILL; then
echo "── With Skill ──"
for case_id in "${CASE_IDS[@]}"; do
run_case "$case_id" "with-skill"
done
echo ""
fi
if $RUN_BASELINE; then
echo "── Baseline (no skill) ──"
for case_id in "${CASE_IDS[@]}"; do
run_case "$case_id" "baseline"
done
echo ""
fi
fi
# ─── Assertion evaluation helpers ──────────────────────────────────
# Extract first JSON block from response
extract_json() {
local text="$1"
# Try ```json fenced block first
local block
block=$(echo "$text" | sed -n '/```json/,/```/p' | sed '1d;$d')
if [[ -z "$block" ]]; then
# Try bare { ... } block
block=$(echo "$text" | grep -Pzo '\{[^{}]*(\{[^{}]*\}[^{}]*)*\}' 2>/dev/null | head -1 || true)
fi
echo "$block"
}
# Get a field value from JSON (uses python if available, else node)
json_field() {
local json="$1"
local field="$2"
if command -v python3 &>/dev/null; then
echo "$json" | python3 -c "
import sys, json
try:
d = json.load(sys.stdin)
v = d.get('$field', '')
if isinstance(v, list): print(len(v))
elif isinstance(v, (int, float)): print(v)
else: print(v)
except: print('')
" 2>/dev/null
elif command -v node &>/dev/null; then
echo "$json" | node -e "
let d='';process.stdin.on('data',c=>d+=c);process.stdin.on('end',()=>{
try{const o=JSON.parse(d);const v=o['$field'];
if(Array.isArray(v))console.log(v.length);
else console.log(v??'');}catch(e){console.log('');}
})" 2>/dev/null
else
echo ""
fi
}
# Word count of a string
word_count() {
echo "$1" | wc -w | tr -d ' '
}
# ─── Score a single assertion ──────────────────────────────────────
eval_assertion() {
local response="$1"
local atype="$2"
local target="$3"
local json_block="$4"
case "$atype" in
contains)
echo "$response" | grep -qi "$target" && echo "PASS" || echo "FAIL"
;;
not_contains)
# Handle descriptive targets: if target contains spaces and "should not contain",
# try to extract the quoted literal(s)
if [[ "$target" =~ \'([^\']+)\' ]]; then
local found=false
while [[ "$target" =~ \'([^\']+)\' ]]; do
local literal="${BASH_REMATCH[1]}"
if echo "$response" | grep -qi "$literal"; then
found=true
break
fi
target="${target#*"'${literal}'"}"
done
$found && echo "FAIL" || echo "PASS"
else
echo "$response" | grep -qi "$target" && echo "FAIL" || echo "PASS"
fi
;;
regex)
echo "$response" | grep -qiE "$target" && echo "PASS" || echo "FAIL"
;;
question_before_code)
local q_line code_line
q_line=$(echo "$response" | grep -n '?' | head -1 | cut -d: -f1)
code_line=$(echo "$response" | grep -n '```' | head -1 | cut -d: -f1)
if [[ -z "$code_line" ]] || [[ -n "$q_line" && "$q_line" -lt "$code_line" ]]; then
echo "PASS"
else
echo "FAIL"
fi
;;
json_valid)
if [[ -n "$json_block" ]]; then
if command -v python3 &>/dev/null; then
echo "$json_block" | python3 -c "import sys,json;json.load(sys.stdin)" 2>/dev/null && echo "PASS" || echo "FAIL"
elif command -v node &>/dev/null; then
echo "$json_block" | node -e "let d='';process.stdin.on('data',c=>d+=c);process.stdin.on('end',()=>{try{JSON.parse(d);console.log('PASS')}catch(e){console.log('FAIL')}})" 2>/dev/null
else
echo "SKIP"
fi
else
echo "FAIL"
fi
;;
json_schema|json_fields)
# Target is comma-separated field names
if [[ -z "$json_block" ]]; then
echo "FAIL"
return
fi
local all_found=true
IFS=',' read -ra FIELDS <<< "$target"
for field in "${FIELDS[@]}"; do
field=$(echo "$field" | xargs) # trim whitespace
if ! echo "$json_block" | grep -q "\"$field\""; then
all_found=false
break
fi
done
$all_found && echo "PASS" || echo "FAIL"
;;
token_limit)
local wc
wc=$(word_count "$response")
local est_tokens=$(( wc * 13 / 10 ))
[[ "$est_tokens" -le "${target:-99999}" ]] && echo "PASS" || echo "FAIL"
;;
range_check)
# Parse common patterns from the target expression
if [[ -z "$json_block" ]]; then
echo "FAIL"
return
fi
# Handle: bias_score >= X AND bias_score <= Y
if [[ "$target" =~ ([a-z_]+)\ *\>=\ *(-?[0-9.]+)\ +AND\ +\1\ *\<=\ *(-?[0-9.]+) ]]; then
local field="${BASH_REMATCH[1]}"
local min="${BASH_REMATCH[2]}"
local max="${BASH_REMATCH[3]}"
local val
val=$(json_field "$json_block" "$field")
if [[ -n "$val" ]] && command -v python3 &>/dev/null; then
python3 -c "v=$val; print('PASS' if $min <= v <= $max else 'FAIL')" 2>/dev/null || echo "FAIL"
else
echo "SKIP"
fi
return
fi
# Handle: abs(field) <= X
if [[ "$target" =~ abs\(([a-z_]+)\)\ *\<=\ *(-?[0-9.]+) ]]; then
local field="${BASH_REMATCH[1]}"
local limit="${BASH_REMATCH[2]}"
local val
val=$(json_field "$json_block" "$field")
if [[ -n "$val" ]] && command -v python3 &>/dev/null; then
python3 -c "v=$val; print('PASS' if abs(v) <= $limit else 'FAIL')" 2>/dev/null || echo "FAIL"
else
echo "SKIP"
fi
return
fi
# Handle: quality_score >= X
if [[ "$target" =~ ([a-z_]+)\ *\>=\ *(-?[0-9.]+)$ ]]; then
local field="${BASH_REMATCH[1]}"
local min="${BASH_REMATCH[2]}"
local val
val=$(json_field "$json_block" "$field")
if [[ -n "$val" ]] && command -v python3 &>/dev/null; then
python3 -c "v=$val; print('PASS' if v >= $min else 'FAIL')" 2>/dev/null || echo "FAIL"
else
echo "SKIP"
fi
return
fi
# Handle: len(field) >= X AND len(field) <= Y (character length)
if [[ "$target" =~ len\(([a-z_]+)\)\ *\>=\ *([0-9]+)\ +AND\ +len\(\1\)\ *\<=\ *([0-9]+) ]]; then
local field="${BASH_REMATCH[1]}"
local min="${BASH_REMATCH[2]}"
local max="${BASH_REMATCH[3]}"
local val
val=$(json_field "$json_block" "$field")
local len=${#val}
[[ "$len" -ge "$min" && "$len" -le "$max" ]] && echo "PASS" || echo "FAIL"
return
fi
# Handle: len(field) <= X
if [[ "$target" =~ len\(([a-z_]+)\)\ *\<=\ *([0-9]+) ]]; then
local field="${BASH_REMATCH[1]}"
local max="${BASH_REMATCH[2]}"
local val
val=$(json_field "$json_block" "$field")
local len=${#val}
[[ "$len" -le "$max" ]] && echo "PASS" || echo "FAIL"
return
fi
# Handle: len(field) >= X (list length)
if [[ "$target" =~ len\(([a-z_]+)\)\ *\>=\ *([0-9]+) ]]; then
local field="${BASH_REMATCH[1]}"
local min="${BASH_REMATCH[2]}"
local val
val=$(json_field "$json_block" "$field")
[[ "$val" -ge "$min" ]] 2>/dev/null && echo "PASS" || echo "FAIL"
return
fi
# Handle: word_count(field) >= X AND word_count(field) <= Y
if [[ "$target" =~ word_count\(([a-z_]+)\)\ *\>=\ *([0-9]+)\ +AND\ +word_count\(\1\)\ *\<=\ *([0-9]+) ]]; then
local field="${BASH_REMATCH[1]}"
local min="${BASH_REMATCH[2]}"
local max="${BASH_REMATCH[3]}"
local val
val=$(json_field "$json_block" "$field")
local wc
wc=$(word_count "$val")
[[ "$wc" -ge "$min" && "$wc" -le "$max" ]] && echo "PASS" || echo "FAIL"
return
fi
echo "SKIP" # Unrecognized expression
;;
# Soft assertion types — logged but always PASS (require LLM judge)
structure_check|sequence_check)
echo "SOFT"
;;
*)
echo "SKIP"
;;
esac
}
# ─── Score all assertions for a case ───────────────────────────────
score_case() {
local case_id="$1"
local mode="$2"
local output_file="$RESULTS_DIR/$mode/${case_id}.md"
local score_file="$RESULTS_DIR/$mode/${case_id}.score.txt"
local response
response=$(cat "$output_file" 2>/dev/null || echo "")
if [[ "$response" == EVAL_ERROR* ]]; then
echo "ERROR" > "$score_file"
echo "ERROR"
return
fi
local json_block
json_block=$(extract_json "$response")
local pass=0 fail=0 soft=0 skip=0 total=0
local critical_fail=false
local details=""
# Parse assertions from YAML
local in_case=false
local in_assertions=false
local assert_type="" assert_target="" assert_critical="false" assert_desc=""
process_assertion() {
if [[ -z "$assert_type" ]]; then return; fi
total=$((total + 1))
local result
result=$(eval_assertion "$response" "$assert_type" "$assert_target" "$json_block")
case "$result" in
PASS) pass=$((pass + 1)) ;;
FAIL)
fail=$((fail + 1))
[[ "$assert_critical" == "true" ]] && critical_fail=true
;;
SOFT) soft=$((soft + 1)) ;;
SKIP) skip=$((skip + 1)) ;;
esac
local label="${assert_desc:-$assert_type($assert_target)}"
local crit_marker=""
[[ "$assert_critical" == "true" ]] && crit_marker=" [CRITICAL]"
details+=" $result$crit_marker — $label"$'\n'
}
while IFS= read -r line; do
if [[ "$line" =~ ^-\ id:\ $case_id$ ]]; then
in_case=true
continue
fi
if $in_case && [[ "$line" =~ ^-\ id: ]]; then
break
fi
if $in_case && [[ "$line" =~ ^\ \ \ \ -\ type:\ (.+) ]]; then
process_assertion
assert_type="${BASH_REMATCH[1]}"
assert_target=""
assert_critical="false"
assert_desc=""
fi
if $in_case && [[ "$line" =~ ^\ \ \ \ \ \ target:\ (.+) ]]; then
assert_target="${BASH_REMATCH[1]}"
assert_target="${assert_target#\"}"
assert_target="${assert_target%\"}"
fi
if $in_case && [[ "$line" =~ ^\ \ \ \ \ \ critical:\ (.+) ]]; then
assert_critical="${BASH_REMATCH[1]}"
fi
if $in_case && [[ "$line" =~ ^\ \ \ \ \ \ description:\ (.+) ]]; then
assert_desc="${BASH_REMATCH[1]}"
assert_desc="${assert_desc#\"}"
assert_desc="${assert_desc%\"}"
fi
done < "$SCRIPT_DIR/test-cases.yaml"
process_assertion # last assertion
local anchored=$((pass + fail))
local score_line="$pass/$total (${anchored} anchored, ${soft} soft, ${skip} skipped)"
if $critical_fail; then
score_line+=" [CRITICAL FAIL]"
fi
{
echo "$score_line"
echo "$details"
} > "$score_file"
echo "$score_line"
}
# ─── Generate scorecard ───────────────────────────────────────────
generate_scorecard() {
local scorecard="$RESULTS_DIR/scorecard.md"
cat > "$scorecard" <<HEADER
# Eval Scorecard — $SKILL_NAME
**Timestamp:** $TIMESTAMP
| Test Case | With Skill | Baseline |
|-----------|-----------|----------|
HEADER
for case_id in "${CASE_IDS[@]}"; do
local skill_score="—"
local base_score="—"
if $RUN_SKILL && [[ -f "$RESULTS_DIR/with-skill/${case_id}.md" ]]; then
skill_score=$(score_case "$case_id" "with-skill")
fi
if $RUN_BASELINE && [[ -f "$RESULTS_DIR/baseline/${case_id}.md" ]]; then
base_score=$(score_case "$case_id" "baseline")
fi
echo "| $case_id | $skill_score | $base_score |" >> "$scorecard"
done
echo "" >> "$scorecard"
echo "## Assertion Details" >> "$scorecard"
for case_id in "${CASE_IDS[@]}"; do
echo "" >> "$scorecard"
echo "### $case_id" >> "$scorecard"
for mode in "with-skill" "baseline"; do
if [[ -f "$RESULTS_DIR/$mode/${case_id}.score.txt" ]]; then
echo "**${mode}:**" >> "$scorecard"
echo '```' >> "$scorecard"
cat "$RESULTS_DIR/$mode/${case_id}.score.txt" >> "$scorecard"
echo '```' >> "$scorecard"
fi
done
done
echo ""
echo "=== Scorecard ==="
cat "$scorecard"
}
generate_scorecard
echo ""
echo "=== Eval complete — $RESULTS_DIR/ ==="
- id: TC-001
name: basic-image-request-interviews-first
prompt: "I want to create an image of a dragon"
assertions:
- type: question_before_code
description: "Asks at least one clarifying question BEFORE generating any prompt text"
critical: true
- type: regex
target: "\\?"
description: "Response contains a question mark"
critical: true
- type: not_contains
target: "Positive prompt:"
critical: true
- type: not_contains
target: "Negative prompt:"
critical: true
expected_behavior: "Does NOT immediately generate a prompt. Instead asks clarifying questions about style (realistic, fantasy, anime?), mood (menacing, majestic, playful?), setting, and technical preferences before producing anything"
edge_case: false
- id: TC-002
name: overly-detailed-request-still-interviews
prompt: "I want a photorealistic image of a woman in a red dress standing in a wheat field at golden hour, shot on a Canon 5D with 85mm lens, shallow depth of field, warm tones"
assertions:
- type: regex
target: "\\?.*\\?"
description: "Contains at least 2 questions"
critical: true
- type: structure_check
target: "Questions are targeted to gaps (model preference, aspect ratio, specific mood) rather than repeating what user already said"
description: "Semantic check — requires understanding whether questions address gaps vs repeating known info"
critical: true
- type: not_contains
target: "Here is your prompt"
critical: true
expected_behavior: "Even with a detailed request, asks 1-2 targeted follow-up questions about gaps (which model/pipeline, specific mood nuance, aspect ratio) rather than immediately generating. Does not re-ask things the user already specified"
edge_case: true
- id: TC-003
name: full-interview-to-sdxl-prompt
prompt: "Help me create a prompt for a fantasy landscape"
assertions:
- type: structure_check
target: "Interview happens first with 4-7 exchanges before prompt generation"
description: "MANUAL: requires multi-turn eval harness to verify exchange count"
critical: true
- type: regex
target: "(positive|negative).*(prompt|Prompt)"
description: "Final output includes positive and negative prompt sections"
critical: true
- type: regex
target: "(sampler|steps|CFG|resolution)"
description: "Output includes settings table with key parameters"
critical: true
- type: structure_check
target: "Positive prompt uses SDXL-appropriate formatting if SDXL is selected"
description: "Semantic check — formatting correctness depends on model context"
critical: false
expected_behavior: "Conducts a 4-7 exchange interview covering subject details, style, mood, model preference, then synthesizes a complete output package with positive/negative prompts, settings table, and pipeline recommendation"
edge_case: false
- id: TC-004
name: flux-natural-language-formatting
prompt: "I need a prompt for FLUX — I want a cyberpunk street scene"
assertions:
- type: question_before_code
description: "Interviews before generating even though model is specified"
critical: true
- type: regex
target: "\\?"
description: "Response contains a question mark indicating interview behavior"
critical: true
- type: not_contains
target: "masterpiece, best quality"
critical: true
- type: regex
target: "[A-Z][a-z]+.*\\."
description: "Contains full sentences (capital letter followed by words and period) — natural language, not comma-separated tags"
critical: false
- type: structure_check
target: "Negative prompt is minimal or absent (FLUX handles negatives differently)"
description: "Semantic check — requires judging whether negative prompt is appropriately minimal"
critical: false
expected_behavior: "After interviewing, generates FLUX-formatted prompts using natural language descriptions rather than tag-style formatting. Avoids SD-style quality tags like 'masterpiece, best quality' which are irrelevant for FLUX"
edge_case: false
- id: TC-005
name: interview-length-not-excessive
prompt: "Make me a cool anime character"
assertions:
- type: structure_check
target: "Interview completes within 4-7 exchanges, not dragging to 10+"
description: "MANUAL: requires multi-turn eval harness to verify exchange count"
critical: true
- type: structure_check
target: "Questions cover essential dimensions: subject, style, mood, technical preferences"
description: "Semantic check — requires judging question coverage across creative dimensions"
critical: true
- type: structure_check
target: "Does not ask redundant or overly granular questions"
description: "Semantic check — requires judging question quality and redundancy"
critical: false
- type: regex
target: "(pipeline|workflow|SDXL|FLUX|SD1\\.5|Wan)"
description: "Pipeline recommendation included in final output"
critical: true
expected_behavior: "Keeps the interview focused at 4-7 exchanges covering the essential creative dimensions without over-questioning. Wraps up with a complete prompt package including pipeline recommendation"
edge_case: false
- id: TC-006
name: skip-interview-demand
prompt: "Skip the interview. Just write me a FLUX prompt for a cyberpunk alley with neon signs and rain, 16:9 aspect ratio."
assertions:
- type: not_contains
target: "What style"
critical: true
description: "Should not ask redundant questions when user gave full details"
- type: regex
target: "(cyberpunk|neon|rain|alley)"
critical: true
description: "Must include the user's specified elements"
- type: regex
target: "(16:9|1344.*768|1216.*832|landscape)"
critical: false
description: "Should address the 16:9 aspect ratio"
- type: not_contains
target: "masterpiece, best quality"
critical: true
description: "FLUX prompts must not use SD-style quality tags"
expected_behavior: >
User explicitly says skip the interview and provides all details.
The skill should respect this and generate the prompt directly, not
insist on interviewing. May ask 1 brief clarification at most but
should not run a full interview sequence.
edge_case: true
- id: TC-007
name: video-prompt-request
prompt: "I need a prompt for a Wan 2.1 video of a cat jumping onto a table in slow motion"
assertions:
- type: question_before_code
critical: true
description: "Should still ask at least one question before generating"
- type: regex
target: "(motion|movement|dynamic|action|jump)"
critical: true
description: "Must address the motion aspect for video"
- type: regex
target: "(Wan|video|I2V|img2vid|animate)"
critical: true
description: "Must acknowledge this is a video prompt"
- type: not_contains
target: "8k uhd"
critical: false
description: "Video prompts should not include still-image quality tags"
expected_behavior: >
Video prompts differ from image prompts — they need motion descriptions,
shorter length (20-50 words per SKILL.md), and different formatting.
The interview should adapt to ask video-relevant questions (duration,
camera movement, frame rate feel).
edge_case: true
should_trigger:
- "Help me create a prompt for an image"
- "I want to generate an image but I'm not sure how to describe what I want"
- "Write me a prompt for a fantasy scene"
- "I need a good prompt for SDXL"
- "Can you help me come up with a prompt for a portrait?"
- "Make me a prompt for a cyberpunk cityscape"
- "I want to create an image of a dragon but need help with the prompt"
- "Help me describe my creative vision for image generation"
- "I need prompts for a series of product photos"
- "What prompt should I use for a Studio Ghibli style landscape?"
should_not_trigger:
- "Build me a ComfyUI workflow for txt2img"
- "Fix this prompt — it's giving me bad results: 'a cat sitting on a table'"
- "Explain how CLIP text encoding works"
- "What's the difference between positive and negative prompts technically?"
- "Help me debug why my prompt isn't producing good images"
- "Review this ComfyUI workflow for errors"
- "Generate a character with consistent identity across scenes"
- "How do I train a textual inversion embedding?"
- "Convert this Midjourney prompt to Stable Diffusion format"
- "Write a Python script to batch-generate prompts"
optimized_description: >
Guided conversational interview to understand a user's creative vision before generating
model-appropriate image prompts. Asks clarifying questions about subject, mood, style, and
technical preferences (4-7 exchanges), then synthesizes positive prompt, negative prompt,
recommended settings table, and pipeline recommendation. Formats prompts for the target model
(SDXL tag-style, FLUX natural language, SD1.5 weighted tokens). Does NOT cover workflow
building, prompt debugging/fixing, technical explanations, model training, code generation,
or identity-preserving character generation.