
Squad Refine
- 29 installs
- Updated July 20, 2026
- steloit/squad-skills
Run a structured interview on a todo backlog card so rough titles become requirements with goals, scope, acceptance criteria, and edge cases before implementation.
About
squad-refine is an agent skill for solo builders using the Squad task pipeline who need disciplined backlog grooming without skipping safety context. Invoked as `/squad-refine <ID>`, it loads the card from the project API, insists on todo status (with a user warning otherwise), then runs a structured interview that converts a rough title and description into actionable requirements: explicit goal, bounded scope, acceptance criteria, and edge cases. Before interviewing, it hunts dependency hints in description and tags, pulls prior card implementation notes and plans, and inspects the codebase when a predecessor exists; if none is named, it asks one upfront question about prior work. The procedure is wired to shared pipeline levels, status transitions, and API endpoints documented in squad/shared.md, and it treats squad/principles.md as mandatory safety reading—not optional flavor text. Use it when backlog items are too fuzzy for an implementation agent and you want repeatable refinement inside the same repo workflow.
- Reads a backlog task via GET /api/task/$ID and enforces todo-status targeting
- Structured user interview produces goal, scope, acceptance criteria, and edge cases
- Scans dependencies, prior implementation_notes, and codebase before questioning
- Mandatory reads: ../squad/shared.md and ../squad/principles.md for pipeline safety
- Warns and confirms if the card is not in todo status
Squad Refine by the numbers
- 29 all-time installs (skills.sh)
- Ranked #1,873 of 3,282 Productivity & Planning skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/steloit/squad-skills --skill squad-refineAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 29 |
|---|---|
| Last updated | July 20, 2026 |
| Repository | steloit/squad-skills ↗ |
What it does
Run a structured interview on a todo backlog card so rough titles become requirements with goals, scope, acceptance criteria, and edge cases before implementation.
Files
Shared context: read ../squad/shared.md for pipeline levels, status transitions, API endpoints, error handling, and agent context flow.Safety principles: read ../squad/principles.md — mandatory, not optional./squad-refine <ID> — Refine Backlog Requirements
Reads a rough backlog item and refines it into concrete, actionable requirements through structured user interview.
Target: tasks in todo status (backlog). If the task is not todo, warn the user and confirm before proceeding.
Procedure
⓪ Resolve the observation gate ONCE (the consent-gate seam — see ../squad/shared.md → Abstraction Rubric)
python3 ../squad/scripts/observe.py gate >/dev/null 2>&1; OBSERVE_OK=$?
# 0 = emit corrections, non-zero = skip. Cache it; every emit below reuses it (best-effort, || true).
# Mint ONE correlation_id for this refine run for any steering emits: CID=$(python3 -c 'import uuid;print(uuid.uuid4())')
① Read the task
TASK = api GET /task/$ID
Extract: title, description, priority, level, tags, card_type
**Epic targets are containers** (`card_type:'epic'`): they hold child tasks, they are not runnable
and have no acceptance-criteria/plan of their own. If the target is an epic, do NOT run the refine
interview — point the user at its children (`GET /api/task/$ID/relationships` → `.children`) and stop.
**Terminal targets are not runnable** (`status` is `cancelled` or `done`): a cancelled (or done)
task is a non-runnable terminal — there is nothing to refine until it re-enters the pipeline. Do NOT
run the interview; warn the user (e.g. `"Task #$ID is cancelled (terminal) — reopen it before refining"`)
and point them at `POST /api/task/$ID/reopen` (which restores `cancelled`/`done` → `todo`), then stop.
(A `done` target may have been reached via the gated pipeline OR an administrative
`POST /api/task/$ID/complete`; both land on the same `done` terminal, so this one branch covers both.)
An **epic** used as a blocker **auto-completes** — when all its children reach a terminal status its
stored status rolls up to `done`/`cancelled`, satisfying the status-based readiness gate automatically
(no manual `/complete` needed). The derived epic `complete` rollup stays display-only; the stored
status (kept in sync by the rollup) is what satisfies the dep.
① ½. Look for prior implementation context (always run this before the interview)
a. Detect dependencies via the relationships API (NOT description text — the `Depends on:` convention
is retired; see `../squad/shared.md` → **Task Relationships & Epics**):
REL = api GET /task/$ID/relationships
Dependency ids = `.blocked_by[].id`. Also check `.parent` for the containing epic.
b. If a dependency found → fetch that card's implementation output:
PRIOR = api GET /task/$NNN?fields=title,implementation_notes,plan
Also inspect the actual codebase: read files, interfaces, schemas confirmed in that card.
c. If no explicit dependency → ask ONE question before the main interview:
"Is there a prior task whose implementation this builds on? (task ID or 'none')"
If the user gives an ID, fetch it as in (b).
If "none" or new work → skip, proceed with regular interview.
d. Summarize what was confirmed from prior implementation:
PRIOR_CONTEXT = {
confirmed interfaces, schemas, file paths, component names, API routes, etc.
}
This context is injected into ③ (gap analysis) and ⑤ (description synthesis).
② Display current state
Show the user their raw title + description as-is.
If PRIOR_CONTEXT exists, also show: "Prior implementation context: [summary]"
③ Analyze for gaps
Identify what's missing or vague across these dimensions:
- WHAT: What exactly should be built/changed?
- WHY: What problem does this solve? What's the motivation?
- SCOPE: What's included vs excluded?
- ACCEPTANCE: How do we know it's done?
- CONSTRAINTS: Technical limitations, compatibility, performance?
- EDGE CASES: Error states, boundary conditions?
- DEPENDENCIES: Does it depend on other tasks or external systems?
④ Interview the user — the GAP-LEDGER LOOP (MANDATORY)
Depth comes from chasing follow-ups, not from one big up-front round. The LOOP
MECHANICS below are low-freedom (re-emit the ledger → ask → probe → call the
script → obey it); the CONTENT of each question stays high-freedom (your
judgement). The stop is OWNED by `../squad/scripts/refine_ledger.py` — you do
NOT self-judge "looks done" (a known gameable failure).
1. Seed + RE-EMIT the ledger artifact. From the ③ gaps, build a fixed-schema
JSON list — and RE-EMIT IT IN FULL at the top of EVERY round (state lives in
tokens, not memory). Do NOT just "keep a ledger" in your head:
[ {"id":"g1","dimension":"WHAT","status":"OPEN","source":"original"},
{"id":"g2","dimension":"SCOPE","status":"OPEN","source":"original"}, … ]
dimension ∈ WHAT/WHY/SCOPE/ACCEPTANCE/CONSTRAINTS/EDGE/DEPS (③); status ∈
OPEN/RESOLVED; source = original | raised-by-answer-R#.
2. SELECT the highest-value OPEN gaps (those whose answers most change
scope / acceptance / level — EVPI-style), recency-first (chase the LAST
answer). Ask 1–4 as ONE AskUserQuestion round. Apply the **Clarification &
Research** rules below to decide menu-vs-research-recommend for each.
3. RECORD answers → mark those ledger entries RESOLVED (re-emit reflects it).
4. PROBE-SCAN (mandatory, non-skippable output slot). Over EACH new answer,
write to a visible `Probe-scan (R#):` slot either:
- ≥1 NEW OPEN ledger entry the answer introduced — an underspecified
concept / a vague term / an in-scope omission / an open choice —
`{"id":…,"dimension":…,"status":"OPEN","source":"raised-by-answer-R#"}`, OR
- an explicit `No new gaps: <reason>` (the diminishing-returns signal).
Grice filter every candidate question: Clear + Relevant + Informative (its
answer increases what's known); drop generic filler; NEVER re-ask a RESOLVED
entry. A genuinely CLEAR card produces `No new gaps` in round 1 — do NOT
manufacture filler gaps to keep the loop alive.
5. CALL THE STOP-GATE and OBEY its exit code. Pipe the re-emitted ledger to the
script with the round number + this round's probe-scan outcome:
printf '%s' "$LEDGER_JSON" | python3 ../squad/scripts/refine_ledger.py \
verdict --round "$R" --last-probe <new_gaps|no_new_gaps>; V=$?
# 0 STOP-CLEAN → go to verify-before-done, then ⑤
# 1 CONTINUE → loop to step 1 (R+1); the script says gaps remain
# 2 STOP-DEGRADED → cap hit with WHAT/SCOPE/ACCEPTANCE still OPEN (backstop)
# 3 STOP-ENOUGH → user escape (backstop)
The verdict is AUTHORITY. Exit 1 means MORE rounds are owed — do not synthesize.
6. BACKSTOPS (not the primary control — the probe-scan + OPEN count are):
- The user says "enough" at ANY point → run the gate with `--user-enough`
(STOP-ENOUGH, exit 3): synthesize now; any residual OPEN row → an
`## Open Questions` block in ⑤.
- The script's hard cap (exit 2, STOP-DEGRADED) → write the residual OPEN rows
into `## Open Questions` AND, since WHAT/SCOPE/ACCEPTANCE is still OPEN,
recommend `/squad-explore` or a card split rather than shipping a falsely-
"refined" card. (`--json` gives `core_unresolved` + `residual_open`.)
- A clean stop (exit 0) at the cap MAY still carry non-core residual OPEN rows
(WHAT/SCOPE/ACCEPTANCE resolved, only non-core gaps left) → write those rows
into `## Open Questions` too. (`--json` `residual_open` lists them; the script's
human output flags `N non-core row(s) → ## Open Questions`.) No explore/split
recommendation here — the core is covered, so the card is genuinely refined.
7. VERIFY-BEFORE-DONE (immediately before ⑤). Re-derive the OPEN count by
re-running the gate on the final ledger; if ANY row is still OPEN under a
synthesize verdict it can only be a cap/enough stop (exit 2/3) → carry the
residual into `## Open Questions`. A self-asserted "done" with an OPEN row on
a non-backstop path is impossible — the gate (exit 1) catches it.
Steering emit (best-effort): if an interview answer REDIRECTS the task's direction
(not a routine fill-in), emit one abstracted user_steering event (enums per ../squad/shared.md →
Abstraction Rubric: interview-redirect row). Skip for ordinary answers.
[ "$OBSERVE_OK" = 0 ] && python3 ../squad/scripts/observe.py emit "$ID" --modality corrective \
--valence na --target scope --severity trivial --attributability latent_preference \
--comment "redirected during the interview" --correlation-id "$CID" || true
── Clarification & Research (value-of-information) ──
Wired into step 2's question selection:
- VoI RULE: ask a question ONLY when its answer would MATERIALLY change what
gets built. Otherwise proceed and NOTE THE ASSUMPTION (don't ask filler).
- CLASSIFY each question you would ask:
· PREFERENCE / OWNERSHIP / IRREVERSIBLE-SCOPE fork → AskUserQuestion MENU,
each option carrying a one-line rationale (the user owns this call).
· ANALYSIS-RESOLVABLE (best practice / architecture / perf / library) →
RESEARCH it and present ONE recommendation WITH REASONING — not a menu.
- MODE-DETECTION → switch to research-then-recommend-ONE for the rest of the
thread (remember it for the session) when the user: re-sends a directive
verbatim, asks "what do you suggest", says "do web research / best practices",
grants "full freedom to re-architect", or rejects an option-set.
- PROPOSAL-NOT-COMMIT: every output is a recommendation + reasoning + a cheap
override (reject / edit / redirect), mirroring Plan Mode. Never read as locked.
- VALUE-GATED RESEARCH: research fires ONLY when value is high — the card is
design / architecture / re-architecture / high-uncertainty (best practices
materially shape the outcome) OR the user explicitly signals it. A clear /
trivial card does NO research (fast refine, near-zero extra tokens). DEPTH
scales with stakes: ONE targeted best-practices check for a moderate card; a
deeper sweep ONLY for a genuine architecture decision — never an always-on
fan-out. EDGE: a research keyword on a trivial card does NOT trigger a sweep —
the gate is VoI, not keyword-match.
- CONFIGURABLE aggressiveness `SQUAD_REFINE_RESEARCH = off | auto-by-value |
always`, default `auto-by-value`. Resolve env `SQUAD_REFINE_RESEARCH` >
committed `.squadrc` `SQUAD_REFINE_RESEARCH=` (the standard `SQUAD_*` ladder,
`../squad/shared.md` → Per-key resolution); `off` disables research entirely,
`always` researches every design-shaped question.
⑤ Synthesize the refined SPEC
Build a structured spec OBJECT — NOT a description rewrite. The human's original
request stays in `description` and is NEVER overwritten; the refined spec is a
separate first-class artifact (the `tasks.spec` shape). If PRIOR_CONTEXT exists,
ground requirements in confirmed interfaces/file paths — not assumptions.
The spec has three authored fields (the server assigns `version`):
- goal: 1–2 sentences — what this task achieves and why.
- requirements: string[] — the COMPLETE, testable set, each item a discrete
string. Describe WHAT, not HOW (no implementation hints/pseudo-code).
Use soft prefixes so intent is explicit (author convention, not
enforced by the API):
"REQ: …" core requirement
"AC: WHEN … THE SYSTEM SHALL …" acceptance criterion (EARS)
"SCOPE(IN): …" / "SCOPE(OUT): …" the in / "Not Included" boundary
"CONSTRAINT: …" technical constraint
"EDGE: …" edge case
"SOURCE: <url>" a research source (guarded; see below)
GUARDED `## Sources` (the Sources convention; reuse note in
`../squad/shared.md` → Sources Convention). EMIT `SOURCE:`
entries ONLY when external research MATERIALLY informed the card
(a non-research card carries NONE — the omit-empty rule). Two
hard guards: (1) cite ONLY sources actually consulted THIS run
as real, verifiable URLs — NEVER fabricate a citation or a
plausible-looking arXiv id; (2) a codebase fact cites `file:line`,
NOT a URL. The entries live IN `requirements[]` as `SOURCE: …`
rows (the spec is a structured object — there is no free-markdown
home for a `## Sources` section; the card view renders the rows
as the Sources block).
- qa: {question, answer}[] — one entry per interview question asked in ④
(answer = the user's chosen value; null if a question went unanswered).
Example:{ "goal": "Let admins invite members so teams can self-serve onboarding.", "requirements": [ "REQ: An admin can send an invite by email from the members page.", "AC: WHEN an admin submits a valid email THE SYSTEM SHALL create an invite and email a signed link.", "SCOPE(OUT): bulk CSV invites are not included.", "EDGE: re-inviting an existing member returns 'already a member' without creating a duplicate." ], "qa": [{ "question": "Email or OAuth invites?", "answer": "Email only for v1" }] }
⑥ Present the refined SPEC + RE-ASSESS THE LEVEL
Paraphrase the resolved scope back, then show the spec in a readable form
(goal, the requirements list, the Q&A).
RE-ASSESS THE LEVEL from the REFINED scope (scope can grow materially during the
interview and the level must follow it — e.g. an L2 that grew to
RLS proofs + e2e + multi-job CI + deploy config stayed L2, skipping plan_review
AND test). Score the refined requirements against `../squad/shared.md` →
Pipeline Levels (L1 trivial / L2 single-layer / L3 new-feature · architecture ·
multi-layer — adds test/CI surface) + `../squad/principles.md` → Card-Split
Criteria. The L1/L2/L3 rubric itself is UNCHANGED; you only re-score against it.
Ask the user to confirm with AskUserQuestion:
- "Approve & save" (write the spec)
- "Edit more" (go back to interview)
- "Cancel" (discard changes)
If the re-assessed level DIFFERS from the current level, surface the change
INSIDE this approval as an explicit choice (the level is NEVER auto-applied;
re-leveling happens ONLY here in refine, never in squad-run):
- "Apply" (save with the re-assessed level, e.g. L2 → L3)
- "Keep" (save, keep the current level)
- "Adjust" (pick a different level)
The chosen level is written in ⑦ (the "Apply the re-assessed level" line).
Steering emit (best-effort): "Approve & save" emits nothing (routine approval).
On "Edit more" OR "Cancel" emit one abstracted user_steering event (enums per
../squad/shared.md → Abstraction Rubric: the Edit-more / Cancel rows). Use the same $CID.
# "Edit more":
[ "$OBSERVE_OK" = 0 ] && python3 ../squad/scripts/observe.py emit "$ID" --modality corrective \
--valence negative --target scope --severity moderate --attributability latent_preference \
--comment "sent the spec back for edits" --correlation-id "$CID" || true
# "Cancel":
[ "$OBSERVE_OK" = 0 ] && python3 ../squad/scripts/observe.py emit "$ID" --modality corrective \
--valence negative --target scope --severity moderate --attributability ambiguous \
--comment "cancelled the refine" --correlation-id "$CID" || true
⑦ Save
If approved:
- **Mint ONE `correlation_id` for THIS save occasion**, before the spec write. The SAME
value tags both the `/spec` write AND the Refiner `/activity` note below, so the board
groups the spec snapshot + the Refiner note into one timeline stage. A re-refine is a
new save → mint a fresh id (never cache/reuse it across saves; a re-refine = new id = a
distinct grouped entry):CORRELATION_ID=$(python3 -c 'import uuid;print(uuid.uuid4())')
- **Write the SPEC via the dedicated endpoint** — the human `description` is NEVER touched.
The CAS token is the TASK `version` (same token every write uses); read it immediately
before writing. The endpoint writes `spec`, bumps `spec_version`, and emits the
`kind='spec'` provenance row. `spec.version` is server-assigned, so omit it from the body.Build the spec object from ⑤ (no version — the server stamps it).
SPEC_JSON=$(jq -n --arg goal "$GOAL" \ --argjson reqs "$REQUIREMENTS_JSON_ARRAY" --argjson qa "$QA_JSON_ARRAY" \ '{goal:$goal, requirements:$reqs, qa:$qa}') VER=$(api GET /task/$ID?fields=version -q version) ERR=$(mktemp) RESP=$(api POST /task/$ID/spec \ --json "$(jq -n --argjson spec "$SPEC_JSON" --argjson ev "$VER" --arg model "$MODEL_REFINER" --arg cid "$CORRELATION_ID" \ '{spec:$spec, expected_version:$ev, actor:"Refiner", model:$model, correlation_id:$cid}')" 2>"$ERR") RC=$?
RC 0 → RESP = { success, version, spec_version }. RC 4 → board rejected: the stderr body
($ERR) carries the board's 4xx — a 412 "Precondition failed" on a concurrent edit (re-read
version and retry ONCE; if it still 412s, surface to the user — don't loop) or a 400
malformed spec.
(On a 412 retry, KEEP the same $CORRELATION_ID — it's still the same save occasion.)
rm -f "$ERR"
- Apply the re-assessed level from ⑥ (and priority/tags if discussed) — PATCH the
`level` to the user's ⑥ choice (Apply / Keep / Adjust), and title/priority/tags
only if the interview changed them (PATCH — never `description`).
- **Declare dependencies structurally**: if the interview surfaced that this task is blocked by
another (#DEP), declare it via a `blocks` edge — NOT a `Depends on:` text line:DEP blocks ID (ID is blocked_by DEP). to is an opaque <KEY>-<seq> id string — use --arg.
Server returns 409 on a cycle (surfaced, no pre-check).
api POST /task/$DEP/relationships --json "$(jq -n --arg to "$ID" '{to:$to, type:"blocks"}')"
- Append the short Refiner activity note (the `kind='spec'` row above carries the snapshot;
this records the round count). Carry the SAME `$CORRELATION_ID` minted at the top of this
step so the board threads this note with the spec snapshot into one timeline stage.
POST /api/task/$ID/activity:
{ "actor": "Refiner", "model": "<MODEL_REFINER>", "message": "Requirements refined. N questions across M rounds.", "correlation_id": "$CORRELATION_ID" }Model Routing
Resolve MODEL_PROVIDER + the read_model helper per ../squad/shared.md → Model Resolution, then:
MODEL_REFINER=$(read_model refiner)Coach (friction review of this run)
After step ⑦ Save completes (an approved refine), dispatch the Coach per ../squad/shared.md → Coach Dispatch (reuse the Model Routing resolution above — MODEL_PROVIDER + helpers). Pass:
skill_name=squad-refinesource_task=$IDrun_summary="squad-refine refined the requirements for task $ID."trajectory= the interview Q/A rounds + the refined specfriction_signals= any board-API friction during the spec write (POST /task/:id/spec);noneif clean
Interview Tips
- If the user wrote "add login" → ask: OAuth/email? Session/JWT? Which pages need auth guards?
- If the user wrote "improve performance" → ask: Which page/API? Current latency? Target latency? Measurement method?
- If the user wrote "fix the UI" → ask: Which component? What's wrong now? Mockup/reference? Responsive?
- Prefer showing concrete options over open-ended "what do you want?"