Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
kiki avatar

Scrutiny

  • Updated July 17, 2026
  • kiki-c0/skills

scrutiny is a Claude Code skill in the Testing & QA category. Discipline for pressure-testing load-bearing claims in analytical writing. Human-facing; see disciplines/scrutiny/README.md.

Key points

  • scrutiny
  • Testing & QA
  • AI-coding skill

Scrutiny by the numbers

  • Data as of Jul 18, 2026 (Skillselion catalog sync)
/plugin marketplace add kiki-c0/skills
/plugin install scrutiny@kiki-c0-skills

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Last updatedJuly 17, 2026
Repositorykiki-c0/skills

What it does

Discipline for pressure-testing load-bearing claims in analytical writing. Human-facing; see disciplines/scrutiny/README.md.

README.md

Scrutiny

A discipline for humans working with AI on analytical writing.

This is a README, not a skill Claude wields. AI systems rationalize past their own tests and produce confident weak claims as a default mode. You are the scrutineer.

The three tests

Apply to any load-bearing claim — the kind the piece depends on, not throwaway observations.

Example-survival. Ask for a specific worked example a skeptical reader could not dismantle. If the AI has to stretch, or reaches for a second that fits worse, the claim is the problem. Narrow, soften, or cut.

Uniqueness. Does the claim do work the other claims don't? If it reduces to something argued elsewhere, cut. Fewer pillars each doing real work beat more pillars with redundancy.

Falsifiability. Ask what would defeat the claim. For "X cannot Y", ask whether Y is achievable with moderate engineering. If yes, the honest claim is a tendency, not an impossibility.

When a test fails to fail

When a test gets a weak-but-not-empty answer — a serviceable example, a half-credible falsifier, a uniqueness defense that gestures at distinctness — switch tests, do not push harder. A claim that scrapes through example-survival but cannot survive uniqueness is the same outcome as a claim that fails example- survival outright. Cycling on one test rewards the AI for ratifying weak output. Switching forces independent attack surfaces.

Warning signs in AI output

  • Forced-example cycling. Second example fits worse than the first. Stop asking for more; attack the claim.
  • Illustration mismatch. The example illustrates a nearby claim, not the one under review. Test: if the opposing case fits this example too, the example is wrong.
  • Architectural overreach. Cannot, always, fundamentally where a tendency claim would be more honest. Read with suspicion.

AI pushback reflexes — both wrong

When you challenge a claim, the AI will often capitulate ("you're right") or re-up ("stronger version"). Both are position-switches driven by conversational pressure, not re-audit. Demand fresh audit: apply the three tests to the challenged claim from scratch. The audit may land on the same position, on capitulation, or on a stronger version. The discipline is the same regardless of direction.

See also

The three tests and warning signs are consistent with question- quality criteria Chris Sanders documented in his 2021 dissertation on expert investigative cognition: relevancy (the scrutineer's question must be grounded in the actual claim), specificity (a narrow range of valid answers — like a worked example), and answerability (resolvable from the text and its sources). Sanders also named refinement: when a question fails to land, reformulate rather than abandon. That is the move behind When a test fails to fail. Different domain (digital forensics), same cognitive primitives.

One sentence

Every load-bearing claim pays rent with a worked example, a unique contribution, and a named way it could be wrong. In AI-assisted writing, the person enforcing the rent is you.

Related skills

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.