Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
aaaaqwq avatar

Model Hierarchy

  • 10 installs
  • 82 repo stars
  • Updated August 2, 2026
  • aaaaqwq/claude-code-skills

model-hierarchy is a Claude Code skill that routes tasks to the cheapest capable model tier based on task complexity to cut AI agent cost.

About

model-hierarchy is a Claude Code skill that cost-optimizes agent operations by routing each task to the cheapest capable model. It classifies tasks as routine, moderate, or complex and maps them to cheap, mid, or premium model tiers, with escalation when a cheaper model already failed and a vision override for image tasks. A developer uses it when deciding which model to use, spawning sub-agents, or cutting cost. It documents a decision algorithm and per-tier model tables.

  • Routes tasks to the cheapest model that can handle them across three cost tiers
  • Classifies each task as routine, moderate, or complex before choosing a model
  • Includes a vision override so image tasks never go to text-only models

Model Hierarchy by the numbers

  • 10 all-time installs (skills.sh)
  • Ranked #11,882 of 16,556 AI & Agent Building skills by installs in the Skillselion catalog
  • Data as of Aug 3, 2026 (Skillselion catalog sync)
At a glance

model-hierarchy capabilities & compatibility

Capabilities
model routing · cost optimization · task classification · agent tooling
Use cases
orchestration · token optimization
Pricing
Free
From the docs

What model-hierarchy says it does

Route tasks to the cheapest model that can handle them. Most agent work is routine.
SKILL.md
**80% of agent tasks are janitorial.** File reads, status checks, formatting, simple Q&A. These don't need expensive models.
SKILL.md
npx skills add https://github.com/aaaaqwq/claude-code-skills --skill model-hierarchy

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs10
repo stars82
Last updatedAugust 2, 2026
Repositoryaaaaqwq/claude-code-skills

What it does

Route each agent task to the cheapest model tier that can handle it, based on task complexity.

Who is it for?

Deciding which model tier a task needs and cutting agent cost

Skip if: Handling model outages or rate limits (that is failover, not routing)

When should I use this skill?

You are deciding which model to use, spawning sub-agents, or the current model feels like overkill

What you get

Each task routed to the cheapest model tier that can handle it, reserving premium models for genuinely complex work.

  • A model-tier routing decision per task
  • A documented selectModel decision algorithm
  • Per-tier model reference tables

By the numbers

  • 3 model tiers (cheap, mid, premium)
  • Claims 80% of agent tasks are routine

Files

SKILL.mdMarkdownGitHub ↗

Model Hierarchy

Route tasks to the cheapest model that can handle them. Most agent work is routine.

Core Principle

80% of agent tasks are janitorial. File reads, status checks, formatting, simple Q&A. These don't need expensive models. Reserve premium models for problems that actually require deep reasoning.

Model Tiers

Tier 1: Cheap ($0.10-0.50/M tokens)

ModelInputOutputBest For
DeepSeek V3$0.14$0.28General routine work
GPT-4o-mini$0.15$0.60Quick responses
Claude Haiku$0.25$1.25Fast tool use
Gemini Flash$0.075$0.30High volume
GLM 5 (Zhipu)(OpenRouter Z.AI)(OpenRouter Z.AI)Routine + moderate text; 200K context; text-only — do not use for image/vision
Kimi K2.5 (Moonshot)$0.45$2.25Routine + moderate; 262K context; multimodal (text + image + video)

Text-only models (e.g. GLM 5): Do not use for any task that requires image input or vision — no photo analysis, screenshots, image-generation tools, or document/chart vision. Route to a vision-capable model (e.g. Kimi K2.5, GPT-4o, Gemini, Claude with vision, GLM-4.5V/4.6V).

Vision-capable Tier 1/2 (e.g. Kimi K2.5): Use for routine or moderate tasks that may involve images — screenshots, photo analysis, docs, image-generation orchestration — without moving to premium vision models.

Tier 2: Mid ($1-5/M tokens)

ModelInputOutputBest For
Claude Sonnet$3.00$15.00Balanced performance
GPT-4o$2.50$10.00Multimodal tasks
Gemini Pro$1.25$5.00Long context

Tier 3: Premium ($10-75/M tokens)

ModelInputOutputBest For
Claude Opus$15.00$75.00Complex reasoning
GPT-4.5$75.00$150.00Frontier tasks
o1$15.00$60.00Multi-step reasoning
o3-mini$1.10$4.40Reasoning on budget

Prices as of Feb 2026. Check provider docs for current rates.

Task Classification

Before executing any task, classify it:

ROUTINE → Use Tier 1

Requires image/vision → Do not assign to text-only models (GLM 5, etc.). Use a vision-capable model from Tier 1/2 or 3 (e.g. Kimi K2.5, GPT-4o, Gemini, Claude, GLM-4.5V).

Characteristics:

  • Single-step operations
  • Clear, unambiguous instructions
  • No judgment required
  • Deterministic output expected

Examples:

  • File read/write operations
  • Status checks and health monitoring
  • Simple lookups (time, weather, definitions)
  • Formatting and restructuring text
  • List operations (filter, sort, transform)
  • API calls with known parameters
  • Heartbeat and cron tasks
  • URL fetching and basic parsing

MODERATE → Use Tier 2

Characteristics:

  • Multi-step but well-defined
  • Some synthesis required
  • Standard patterns apply
  • Quality matters but isn't critical

Examples:

  • Code generation (standard patterns)
  • Summarization and synthesis
  • Draft writing (emails, docs, messages)
  • Data analysis and transformation
  • Multi-file operations
  • Tool orchestration
  • Code review (non-security)
  • Search and research tasks

COMPLEX → Use Tier 3

Characteristics:

  • Novel problem solving required
  • Multiple valid approaches
  • Nuanced judgment calls
  • High stakes or irreversible
  • Previous attempts failed

Examples:

  • Multi-step debugging
  • Architecture and design decisions
  • Security-sensitive code review
  • Tasks where cheaper model already failed
  • Ambiguous requirements needing interpretation
  • Long-context reasoning (>50K tokens)
  • Creative work requiring originality
  • Adversarial or edge-case handling

Decision Algorithm

function selectModel(task):
    # Rule 1: Vision override (Tier 1/2 includes text-only models)
    if task.requiresImageInput or task.requiresVision:
        return VISION_CAPABLE_MODEL  # e.g. Kimi K2.5, GPT-4o, Gemini, Claude; do not use GLM 5 or other text-only
    
    # Rule 2: Escalation override
    if task.previousAttemptFailed:
        return nextTierUp(task.previousModel)
    
    # Rule 3: Explicit complexity signals
    if task.hasSignal("debug", "architect", "design", "security"):
        return TIER_3
    
    if task.hasSignal("write", "code", "summarize", "analyze"):
        return TIER_2
    
    # Rule 4: Default classification
    complexity = classifyTask(task)
    
    if complexity == ROUTINE:
        return TIER_1
    elif complexity == MODERATE:
        return TIER_2
    else:
        return TIER_3

Behavioral Rules

For Main Session

1. Default to Tier 2 for interactive work 2. Suggest downgrade when doing routine work: "This is routine - I can handle this on a cheaper model or spawn a sub-agent." 3. Request upgrade when stuck: "This needs more reasoning power. Switching to [premium model]."

For Sub-Agents

1. Default to Tier 1 unless task is clearly moderate+ 2. Batch similar tasks to amortize overhead 3. Report failures back to parent for escalation

For Automated Tasks

1. Heartbeats/monitoring → Always Tier 1 2. Scheduled reports → Tier 1 or 2 based on complexity 3. Alert responses → Start Tier 2, escalate if needed

Communication Patterns

When suggesting model changes, use clear language:

Downgrade suggestion:

"This looks like routine file work. Want me to spawn a sub-agent on DeepSeek for this? Same result, fraction of the cost."

Upgrade request:

"I'm hitting the limits of what I can figure out here. This needs Opus-level reasoning. Switching up."

Explaining hierarchy:

"I'm running the heavy analysis on Sonnet while sub-agents fetch the data on DeepSeek. Keeps costs down without sacrificing quality where it matters."

Cost Impact

Assuming 100K tokens/day average usage:

StrategyMonthly CostNotes
Pure Opus~$225Maximum capability, maximum spend
Pure Sonnet~$45Good default for most work
Pure DeepSeek~$8Cheap but limited on hard problems
Hierarchy (80/15/5)~$19Best of all worlds

The 80/15/5 split:

  • 80% routine tasks on Tier 1 (~$6)
  • 15% moderate tasks on Tier 2 (~$7)
  • 5% complex tasks on Tier 3 (~$6)

Result: 10x cost reduction vs pure premium, with equivalent quality on complex tasks.

Integration Examples

OpenClaw

# config.yml - set default model
model: anthropic/claude-sonnet-4

# In session, switch models
/model opus  # upgrade for complex task
/model deepseek  # downgrade for routine

# Spawn sub-agent on cheap model
sessions_spawn:
  task: "Fetch and parse these 50 URLs"
  model: deepseek

OpenRouter (Tier 1 with vision or text-only):

# Tier 1 with vision — Kimi K2.5 (multimodal)
model: openrouter/moonshotai/kimi-k2.5
# Heartbeats, cron, image-involving tasks: K2.5 handles text and vision.

# Tier 1 text-only — GLM 5 (no vision)
# model: openrouter/z-ai/glm-5  # exact ID TBD on OpenRouter Z.AI
# Routine text-only only; for image tasks use Kimi K2.5 or another vision-capable model.

Claude Code

# In CLAUDE.md or project instructions
When spawning background agents, use claude-3-haiku for:
- File operations
- Simple searches  
- Status checks

Reserve claude-sonnet-4 for:
- Code generation
- Analysis tasks

General Agent Systems

def get_model_for_task(task_description: str) -> str:
    routine_signals = ['read', 'fetch', 'check', 'list', 'format', 'status']
    complex_signals = ['debug', 'architect', 'design', 'security', 'why']
    
    desc_lower = task_description.lower()
    
    if any(signal in desc_lower for signal in complex_signals):
        return "claude-opus-4"
    elif any(signal in desc_lower for signal in routine_signals):
        return "deepseek-v3"
    else:
        return "claude-sonnet-4"

Anti-Patterns

DON'T:

  • Run heartbeats on Opus
  • Use premium models for file I/O
  • Keep expensive model when task is clearly routine
  • Spawn sub-agents on premium models by default
  • Use GLM 5 (or any text-only Tier 1/2 model) for image/vision tasks — e.g. photo analysis, screenshot understanding, image-generation skills, or any tool that takes image input

DO:

  • Start mid-tier, adjust based on task
  • Spawn helpers on cheapest viable model
  • Escalate explicitly when stuck
  • Track cost per task type to optimize further

Extending This Skill

To customize for your use case:

1. Adjust tier definitions based on your provider/budget 2. Add domain-specific signals to classification rules 3. Track actual complexity vs predicted to improve heuristics 4. Set budget alerts to catch runaway premium usage

Related skills

FAQ

How does model-hierarchy pick a model?

It classifies the task as routine, moderate, or complex and routes to the cheap, mid, or premium tier, escalating if a cheaper model already failed.

What about image tasks?

A vision override sends any image or vision task to a vision-capable model rather than a text-only one.

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.