Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
ats-kinoshita-iso avatar

Eval Framework

  • Updated April 1, 2026
  • ats-kinoshita-iso/agent-workshop

eval-framework offers skills for designing scoring rubrics, running structured evaluations on LLM outputs, and comparing candidate outputs to recommend a winner. It is agent-tooling for developers who need to measure and compare LLM quality systematically.

Key points

  • Scoring rubric design
  • Structured LLM evaluations
  • Candidate output comparison
  • Winner recommendation

Eval Framework by the numbers

  • Data as of Jul 7, 2026 (Skillselion catalog sync)
/plugin marketplace add ats-kinoshita-iso/agent-workshop
/plugin install eval-framework@agent-workshop

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Last updatedApril 1, 2026
Repositoryats-kinoshita-iso/agent-workshop

What it does

Design scoring rubrics, run structured evaluations on LLM outputs, and compare candidate outputs to recommend a winner.

README.md

eval-framework

Evaluation framework skills for designing scoring rubrics, running structured evaluations on LLM outputs, and comparing candidates to recommend a winner.

Skills

/eval-design

Design evaluation criteria and a 1-5 anchored scoring rubric for any task or output type. Produces a structured framework with weighted dimensions, per-level anchor descriptions, and explicit pass/fail thresholds. Reduces inter-rater variance by making quality distinctions concrete and measurable.

/eval-run

Apply an existing rubric to one or more candidate outputs and produce a scored evaluation report. Scores each output per dimension, computes weighted totals, applies pass/fail thresholds, and returns actionable feedback for failing outputs.

/eval-compare

Compare two candidate outputs (A/B eval) on shared criteria and recommend the winner with a justified explanation. Handles trade-off analysis, weighted scoring, and sensitivity to dimension priority — useful for prompt engineering and model selection.

When to use

Skill Trigger
/eval-design You need to establish quality criteria before evaluating anything
/eval-run You have a rubric and want to score one or more outputs
/eval-compare You have two candidate outputs and need to pick the better one

Workflow

The three skills form a pipeline:

/eval-design  -->  /eval-run  -->  /eval-compare
   (define)         (score)          (decide)

Sources

Derived from:

  • tool_evaluation/ patterns (anthropic-cookbook)
  • patterns/agents/evaluator_optimizer.ipynb (anthropic-cookbook)

Related skills

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.