
Ia Reflect
- 3 installs
- 28 repo stars
- Updated August 5, 2026
- iliaal/whetstone
Review a coding session for mistakes, friction, and wins, then audit skills and persist chosen lessons to memory.
About
Runs a structured session retrospective that scans the full conversation for mistakes, friction, wasted effort, and wins, then audits the skills used. A developer uses it at the end of a session to capture lessons learned and decide what to persist to memory.
- Cites the specific exchange and impact for each mistake, friction point, or win
- Proposes measurable skill-audit changes and asks which lessons to persist to memory
Ia Reflect by the numbers
- 3 all-time installs (skills.sh)
- Ranked #2,390 of 3,282 Productivity & Planning skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/iliaal/whetstone --skill ia-reflectAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 3 |
|---|---|
| repo stars | ★ 28 |
| Last updated | August 5, 2026 |
| Repository | iliaal/whetstone ↗ |
What it does
Review a coding session for mistakes, friction, and wins, then audit skills and persist chosen lessons to memory.
Files
Reflect
Success Criteria
- Every mistake/friction point cites the specific moment and its impact
- Improvements are actionable and prioritized (cap defined in step 4)
- Each skill audit proposes measurable changes (not vague suggestions)
- User is asked which items to persist to memory
- If review activity occurred, review-trap patterns are captured to persistent memory, or explicitly marked as "none"
Process
1. Session Review
Scan the full conversation. For each finding, cite the specific exchange (quote or paraphrase) and its impact.
| Category | Signal |
|---|---|
| Mistakes | Wrong outputs, incorrect assumptions, hallucinated facts |
| Friction | Repeated clarifications, verbose responses, misread intent |
| Wasted effort | Work discarded, wrong approaches tried first |
| Wins | Approaches worth repeating, smooth interactions |
Skip one-time typos, external tool failures, and issues outside agent control.
2. Review Activity Scan (if applicable)
If the session included PR or MR review activity in either direction, run this scan before moving on. Skip only if no reviews happened.
Inbound (my code was reviewed): For each review comment received:
- Did I accept it? If yes, what pattern did the reviewer catch that I missed? Is it a recurring blind spot? Capture the one-liner to persistent memory.
- Did I push back? If I was right and the reviewer was wrong, nothing to capture. If I was wrong and had to retract mid-thread, capture what I learned.
Outbound (I reviewed someone else's code): For each comment I authored:
- Was it accepted? Nothing to capture -- good call.
- Was it rejected with a valid counter? That's a review trap. Capture the pattern: what heuristic did I apply that produced a wrong comment?
"No harvestable items" is a valid outcome -- say so explicitly. Don't let the step quietly drop off.
3. Operational Learnings
Before listing improvements, scan the session for operational insights worth preserving. Apply the 5-minute filter: would knowing this save 5+ minutes in a future session? If yes, include it. Examples: a project-specific quirk, a command that failed unexpectedly, an approach that worked better than expected.
4. Improvements
Numbered list of concrete improvements, ranked by impact. Each item: one sentence, imperative, actionable. Cap at 10 items: if more surface, the bottom items are noise -- drop them rather than batching or splitting.
Ask: "Which of these should I remember for future chats?"
Save approved items to memory files at ~/.claude/projects/<project-slug>/memory/ (replace <project-slug> with the slug matching the current working directory, e.g., -home-ilia-ai-whetstone) using the Write tool with proper frontmatter (see MEMORY.md index).
5. Skill Audit (if skills were used)
For each skill invoked during the session:
A. Self-check gate -- If the skill lacks success criteria + verification loop:
- Add
## Success Criteriaat top (3-5 measurable checks) - Add
## Self-Checkat bottom: "Verify all success criteria are met before presenting output. If not, iterate (max 5 times)."
B. Token efficiency -- Flag: redundant phrasing, mergeable sections, oversized examples, "Claude already knows this" content, inert frontmatter metadata.
C. Other -- Missing edge cases, vague directives (rewrite as measurable criteria or remove), naked negations (add "do Y instead" or remove).
Present proposed changes as diffs. Ask: "Apply these? (all / pick / skip)"
6. Capture Markers
The `remember:` prefix is the highest-confidence capture signal. When the user writes a message beginning with remember:, treat everything after the colon as a memory candidate — no interpretation required. Save directly to the appropriate memory file with a one-line summary and the user's exact phrasing. Example: remember: we never use Pest, always PHPUnit → save to feedback_phpunit_over_pest.md.
Correction patterns to watch for (lower-confidence, batch these for review at /ia-reflect time):
- "no, use X" / "actually, X" / "don't use Y, use X"
- "stop doing X" / "never X"
- "that's wrong — the right way is..."
- repeated clarifications of the same thing within a session
Optional capture hook: a UserPromptSubmit hook can pattern-match the markers above into ~/.claude/learnings-queue.json as the user types, so /ia-reflect processes the queue deterministically instead of re-scanning the full transcript. Not shipped with this skill; document the convention and leave implementation to users who need it.
7. Pattern Detection
If 2+ similar tasks appear that no existing skill covers, suggest a new skill (1-2 sentence description). Create only after confirmation.
Proactive trigger: When the user corrects you, clarifies the same thing twice, or shows frustration, append: "Tip: Type /ia-reflect when you're ready -- I'll review what we can improve."
Self-Check
Before presenting output, verify all success criteria are met. If any fail, revise (max 5 iterations).
ia-reflect Specification
Intent
ia-reflect is a tool-class skill (a narrow utility scoped to a single capability). Session retrospective and skill audit. Use when asked to reflect, do a retrospective, review lessons learned, audit what went well or wrong, or review session effectiveness.
Scope
In scope:
- Behaviors described in
SKILL.mdand routed via the should_trigger phrasings indistillery/tests/fixtures/triggers/ia-reflect.jsonl. - Updates to runtime behavior, structure, trigger precision, references, and validation.
Out of scope:
- Acting as the runtime instructions themselves (those live in
SKILL.md). - Trigger phrasings already covered by adjacent
ia-*skills (validate-pluginflags >70% description overlap as DUPLICATE_TRIGGER). - <!-- to fill in: domain-specific exclusions when the skill drifts -->
Trigger Context
- Class:
tool - Hook regex:
plugins/whetstone/hooks/skill-patterns.sh->SKILL_PATTERNS[ia-reflect] - Common requests (from fixture should_trigger):
- "let's do a retrospective on this session"
- "what went wrong with the last deployment"
- "retrospective on this debugging session"
- Should not trigger for (from fixture should_not_trigger):
- "implement the webhook handler for Stripe events"
- "update the Docker compose file for local dev"
- "plan the next feature"
Source And Evidence Model
Authoritative sources:
SKILL.md-- runtime instructions and reference routing.references/*.md-- bundled supplementary content (0 file(s)).distillery/tests/fixtures/triggers/ia-reflect.jsonl-- positive and negative trigger phrasings under regression test.plugins/whetstone/hooks/skill-patterns.sh-- regex pattern that fires this skill.distillery/.eval-data/ia-reflect/-- harvested session examples (when present).
Data that must not be stored in this skill or its references:
- Secrets, credentials, tokens.
- Machine-specific filesystem paths (
/home/...,/Users/...,~/ai/...). The validator (MACHINE_PATH_LEAK) flags these as HIGH. - Private URLs, customer data, or unredacted personal information.
Coverage matrix
| Dimension | Status | Evidence |
|---|---|---|
| Trigger fixtures | complete | distillery/tests/fixtures/triggers/ia-reflect.jsonl (>=5 should_trigger, >=5 should_not_trigger) |
| Hook regex pattern | complete | plugins/whetstone/hooks/skill-patterns.sh (SKILL_PATTERNS[ia-reflect]) |
| Reference architecture | n/a | no references; SKILL.md is self-contained |
| Real-usage signal | <!-- populated by harvest-sessions when sessions exist --> | distillery/.eval-data/ia-reflect/ (created by harvest-sessions) |
Evaluation
Lightweight (run on every change):
python3 distillery/scripts/distiller.py validate-plugin --component ia-reflect
python3 distillery/scripts/distiller.py test-triggers --skill ia-reflectDeeper (when behavior risk warrants):
python3 distillery/scripts/distiller.py dspy-eval ia-reflect
python3 distillery/scripts/distiller.py diagnose-negatives ia-reflectAcceptance gates:
validate-plugin --component ia-reflectreturns 0 HIGH findings.test-triggers --skill ia-reflectreturns F1 = 1.0 with floors of 5 should_trigger and 5 should_not_trigger.- For dspy-eval, the composite score does not regress against the most recent saved baseline (see
distillery/.eval-data/ia-reflect/history.json).
Known Limitations
<!-- to fill in over time as drift surfaces. Default rule: any time diagnose-negatives surfaces a recurring failure pattern, document it here so future maintainers understand the trade-off the current implementation accepts. -->
Maintenance Notes
- Update
SKILL.mdwhen the runtime workflow, branch conditions, or output contract changes. - Update this
SPEC.mdwhen intent, scope, evidence model, evaluation gates, or maintenance expectations change. - Update the trigger fixture when adding new positive phrasings, removing stale ones, or expanding scope (the 5/5 floor is a hard validator gate).
- Update the hook regex in
skill-patterns.shwhenever fixture positives expose a missed phrasing; verify F1 = 1.0 witheval-triggersbefore committing. - Run the full release pipeline via
/release-- never bump versions or update CHANGELOG.md from a per-skill edit.