Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
erichowens avatar

Cost Verification Auditor

  • 111 installs
  • 178 repo stars
  • Updated July 14, 2026
  • erichowens/some_claude_skills

Audit and verify API costs, token usage, and billing accuracy for agent systems.

About

Cost Verification Auditor validates billing and token usage accuracy. Ensure correct pricing application and detect billing anomalies.

  • Cost auditing and verification.
  • Token usage validation.

Cost Verification Auditor by the numbers

  • 111 all-time installs (skills.sh)
  • Ranked #549 of 1,039 Cloud & Infrastructure skills by installs in the Skillselion catalog
  • Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/erichowens/some_claude_skills --skill cost-verification-auditor

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs111
repo stars178
Last updatedJuly 14, 2026
Repositoryerichowens/some_claude_skills

What it does

Audit and verify API costs, token usage, and billing accuracy for agent systems.

Files

SKILL.mdMarkdownGitHub ↗

Cost Verification Auditor

Verify that token cost estimates are within ±20% of actual Claude API usage.

When to Use

Use for:

  • Validating token estimation systems after implementation
  • Pre-deployment cost accuracy checks
  • Debugging unexpected API bills
  • Periodic estimation drift detection

NOT for:

  • Looking up model pricing (use pricing docs)
  • Budget planning or forecasting
  • Cost optimization strategies
  • Comparing models by price

Core Audit Process

Decision Tree

Has estimator? ──No──→ Build estimator first (see Calibration Guidelines)
      │
     Yes
      ↓
Define 3+ test cases (simple/medium/complex)
      ↓
Estimate BEFORE execution (no peeking!)
      ↓
Execute against real API
      ↓
Calculate variance: (actual - estimated) / estimated
      ↓
Variance ≤ ±20%? ──Yes──→ PASS ✓
      │
     No
      ↓
Apply fixes from Anti-Patterns section
      ↓
Re-run verification

Variance Formula

const inputVariance = (actual.inputTokens - estimate.inputTokens) / estimate.inputTokens;
const outputVariance = (actual.outputTokens - estimate.outputTokens) / estimate.outputTokens;
const costVariance = (actual.totalCost - estimate.totalCost) / estimate.totalCost;

// PASS if both input AND output within ±20%
const passed = Math.abs(inputVariance) <= 0.20 && Math.abs(outputVariance) <= 0.20;

Common Anti-Patterns

Anti-Pattern: The 500-Token Overhead Myth

Novice thinking: "Claude Code adds ~500 tokens overhead, so add that to every estimate."

Reality: Direct API calls have ~10 token overhead. The 500+ overhead is ONLY when using Claude Code's full context (system prompts, tools, conversation history).

Timeline:

  • Pre-2025: Many tutorials used 500+ token estimates
  • 2025+: Direct API overhead is minimal (~10 tokens)

What to use instead:

ContextOverhead
Direct API call~10 tokens
With system prompt50-200 tokens
With tools/functions100-500 tokens
Claude Code full context500-2000 tokens

How to detect: Consistent 40-90% overestimation = overhead too high.

---

Anti-Pattern: Per-Node Accuracy Obsession

Novice thinking: "Every node must be within ±20% or the estimator is broken."

Reality: LLM output length is non-deterministic. Per-node output variance of 30-50% is normal. What matters is aggregate cost accuracy.

What to use instead:

  • Focus on total DAG cost variance (should be ±20%)
  • Accept per-node output variance up to ±40%
  • Use constrained prompts ("list exactly 3") to reduce variance

How to detect: Input estimates accurate, output varies wildly = normal LLM behavior.

---

Anti-Pattern: Peeking Before Estimating

Novice thinking: "Let me run the API call first to see what tokens we get, then build the estimator."

Reality: This produces perfectly-fitted estimates that fail on new prompts. Estimation must happen BEFORE execution.

Correct approach: 1. Estimate based on prompt length and heuristics 2. Execute API call 3. Compare variance 4. Adjust heuristics if needed

Calibration Guidelines

Input Token Estimation

// Calibrated 2026-01-30
const inputTokens = Math.ceil(prompt.length / CHARS_PER_TOKEN) + OVERHEAD;
Text TypeCHARS_PER_TOKENNotes
English prose4.0Most consistent
Code3.0-3.5Symbols tokenize differently
Mixed3.5Balanced (recommended default)
JSON/structured3.0Punctuation heavy

Output Token Estimation

Prompt ConstraintMultiplierNotes
"List exactly N items"0.8x inputHighly constrained
"Brief summary"1.0x inputModerate
"Explain in detail"2-3x inputExpansive
Unconstrained1.5x inputVariable

Always: Minimum 100 output tokens for any meaningful response.

Model Behavior

ModelOutput Tendency
Claude OpusLonger, more detailed
Claude SonnetBalanced
Claude HaikuConcise, efficient

Quick Fixes

SymptomCauseFix
Overestimating by 40%+Overhead too highReduce from 500 → 10
Underestimating inputsChars/token too highReduce from 4.0 → 3.5
Output wildly variesLLM non-determinismUse constrained prompts
Total cost accurate but per-node offNormal aggregationAccept it, focus on totals

Verification Checklist

  • [ ] 3+ test cases (simple, medium, complex)
  • [ ] Estimates run BEFORE API calls
  • [ ] Variance formula: (actual - estimated) / estimated
  • [ ] Target: ±20% for input AND output
  • [ ] Report includes actionable recommendations

References

See /references/calibration-data.md for detailed calibration tables and historical data.

Related skills

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.