Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
affaan-m avatar

Scholar Evaluation

  • 1.4k installs
  • 238k repo stars
  • Updated August 5, 2026
  • affaan-m/ecc

This is a copy of scholar-evaluation by affaan-m - installs and ranking accrue to the original listing.

scholar-evaluation is a Claude Code skill that runs rigorous academic-style assessments of AI agent outputs, research quality, and reasoning chains for developers who need reproducible rubric scoring before trusting or p

About

scholar-evaluation is an ECC community skill that scores academic and technical artifacts with structured 1-to-5 rubrics across methodology, evidence, citations, and limitations. It supports empirical papers, theoretical work, technical reports, systematic or narrative literature reviews, proposals, thesis chapters, and conference abstracts—with comprehensive, targeted, or comparative evaluation modes. Developers and researchers reach for scholar-evaluation when reviewing agent-generated research summaries, validating that claims are citation-backed, comparing multiple papers on the same rubric, or producing structured revision feedback. Each dimension scores from 1 (substantial revision needed) through 5 (publication-ready), making quality gaps explicit before outputs inform product or publication decisions.

  • Evaluates agent responses using scholar-grade criteria including correctness, completeness, novelty and rigor
  • Produces structured scoring with detailed feedback and improvement recommendations
  • Supports parallel evaluation runs for benchmarking multiple agent variants
  • Works across research, code generation, analysis and creative reasoning tasks

Scholar Evaluation by the numbers

  • 1,382 all-time installs (skills.sh)
  • +93 installs in the week ending Aug 4, 2026 (Skillselion tracking)
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/affaan-m/ecc --skill scholar-evaluation

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs1.4k
repo stars238k
Last updatedAugust 5, 2026
Repositoryaffaan-m/ecc

How do you evaluate AI agent research output quality?

Run rigorous academic-style assessments of AI agent outputs, research quality, and reasoning chains.

Who is it for?

Developers and researchers who need reproducible academic rubrics to judge agent outputs, literature reviews, or technical reports before publication or production use.

Skip if: Teams needing only informal code review or lint feedback without academic evidence standards should skip scholar-evaluation.

When should I use this skill?

Research papers, agent reasoning chains, proposals, or literature reviews need rubric-based quality assessment with citation and methodology checks.

What you get

Rubric scores per dimension, structured revision feedback, and comparative rankings across artifacts

  • Rubric score report
  • Structured revision feedback
  • Comparative rankings

By the numbers

  • Uses 1-to-5 scoring rubric per evaluation dimension
  • Supports comprehensive, targeted, and comparative evaluation modes

Files

SKILL.mdMarkdownGitHub ↗

Scholar Evaluation

このスキルを使用して、再現可能なルーブリックで学術的または科学的な作業を評価します。

使用するタイミング

  • 研究論文、提案書、論文章、または文献レビューのレビュー。
  • 主張が引用された証拠によって支持されているかの確認。
  • 方法論、研究デザイン、分析、または限界の評価。
  • 品質または関連性について2つ以上の論文を比較する。
  • 改訂のための構造化されたフィードバックの作成。

評価の範囲

まず成果物を特定します:

  • 実証的研究論文
  • 理論論文
  • 技術レポート
  • システマティックまたはナラティブ文献レビュー
  • 研究提案書
  • 論文または学位論文の章
  • 学会アブストラクトまたは短い論文

次に範囲を選択します:

  • 包括的: すべてのルーブリック次元
  • 的を絞った: 1つまたは2つの次元(方法や引用など)
  • 比較的: 同じルーブリックに対して複数の作品をランク付け

ルーブリック

該当する各次元を1から5でスコアリング:

  • 5: 優れている;明確、厳密で、出版の準備ができている
  • 4: 良好;軽微な改善が必要
  • 3: 適切;意味のあるギャップがあるが使用可能
  • 2: 弱い;実質的な改訂が必要
  • 1: 不十分;主要な有効性または明確性の問題

適用されない次元には N/A を使用します。

1. 問題と研究質問

  • 問題は明確かつ具体的か?
  • 貢献は意義があるか?
  • 範囲と前提は明示的か?
  • 質問は主張された貢献と一致しているか?

2. 文献とコンテキスト

  • 関連する先行研究はカバーされているか?
  • 作業はソースを単に列挙するのではなく統合しているか?
  • ギャップは正確に特定されているか?
  • 最近のソースと基礎的なソースはバランスが取れているか?

3. 方法論

  • 方法は研究質問に答えるか?
  • デザインの選択は正当化されているか?
  • 変数、データセット、参加者、または材料は明確に説明されているか?
  • 別の研究者が研究を再現できるか?
  • 倫理的および実際的な制約は認識されているか?

4. データと証拠

  • データソースは信頼性があり適切か?
  • サンプルサイズまたはコーパスのカバレッジは十分か?
  • 包含、除外、前処理の決定は文書化されているか?
  • 欠損データとバイアスのリスクは議論されているか?

5. 分析

  • 統計的、定性的、または計算的方法は適切か?
  • ベースラインとコントロールは公正か?
  • 必要に応じて不確実性、感度、またはロバスト性のチェックが含まれているか?
  • 代替の説明は考慮されているか?

6. 結果と解釈

  • 結果は明確に提示されているか?
  • 主張は証拠の範囲内に留まっているか?
  • 図、表、メトリクスは理解しやすいか?
  • 否定的または帰無的な結果は正直に扱われているか?

7. 限界と妥当性への脅威

  • 限界は一般的でなく具体的か?
  • 内部、外部、構成概念、結論の妥当性リスクは対処されているか?
  • 論文は投機的な結果と実証された結果を区別しているか?

8. 文章と構造

  • 論証は理解しやすいか?
  • セクションは研究質問を中心に構成されているか?
  • 定義と表記は明確か?
  • トーンは正確で学術的か?

9. 引用

  • 引用された論文はそれに付けられた主張を支持しているか?
  • 可能な限り一次ソースが使用されているか?
  • レビューはレビューとしてラベル付けされているか?
  • プレプリントはプレプリントとしてラベル付けされているか?
  • 引用メタデータとリンクは正しいか?

レビュープロセス

1. 主張された貢献についてアブストラクト、序論、図、結論を読む。 2. 証拠の質について方法と結果を読む。 3. 最も強い主張を引用されたソースと照合する。 4. 該当する各次元をスコアリングする。 5. 重大なブロッカーと改訂提案を分離する。 6. 具体的な次の編集で終了する。

出力テンプレート

# Scholar Evaluation: <成果物>

## 総合評価

- 総合スコア: <1-5 または N/A>
- 信頼度: <高 | 中 | 低>
- サマリー: <3-5 文>

## 次元スコア

| 次元 | スコア | 証拠 | 改訂優先度 |
| --- | ---: | --- | --- |
| 問題と質問 |  |  |  |
| 文献とコンテキスト |  |  |  |
| 方法論 |  |  |  |
| データと証拠 |  |  |  |
| 分析 |  |  |  |
| 結果と解釈 |  |  |  |
| 限界 |  |  |  |
| 文章と構造 |  |  |  |
| 引用 |  |  |  |

## 重大な問題

## 推奨される改訂

## 必要な証拠確認

落とし穴

  • スコアを具体的なフィードバックの代替として使用しない。
  • 範囲外の次元の省略で論文にペナルティを与えない。
  • 引用数、出版場所、または著者の評判を品質の証拠として扱わない。
  • アブストラクトに現れるからといって根拠のない主張を受け入れない。

Related skills

How it compares

Use scholar-evaluation when you need citation-backed academic rigor on research artifacts rather than informal summarization or generic code review.

FAQ

What artifacts does scholar-evaluation assess?

scholar-evaluation assesses empirical research papers, theoretical papers, technical reports, systematic or narrative reviews, proposals, thesis chapters, and conference abstracts.

How does scholar-evaluation score quality?

scholar-evaluation applies 1-to-5 rubric scores per dimension, where 5 is publication-ready and 1 requires substantial revision, across methodology, evidence, and citation support.

Can scholar-evaluation compare multiple papers?

scholar-evaluation offers a comparative mode that ranks multiple works against the same rubric dimensions for relevance and methodological quality.

AI & Agent Buildingagentsllmresearch

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.