Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
galaxy-dawn avatar

Results Analysis

  • 360 installs
  • 5k repo stars
  • Updated July 17, 2026
  • galaxy-dawn/claude-scholar

results-analysis is a Claude Code research skill that interprets experiment metrics, ablations, and statistical significance for developers and researchers who must decide whether hypotheses hold before drafting papers.

About

results-analysis is a research interpretation skill from galaxy-dawn/claude-scholar that helps developers and ML researchers evaluate experiment outputs including metrics, ablation comparisons, and statistical significance before committing to full paper drafts. The skill supports go-no-go decisions on whether reported results substantiate stated hypotheses or require additional runs, different baselines, or revised claims. Researchers reach for results-analysis after training runs or benchmark sweeps complete and before investing time in LaTeX manuscripts or conference submissions. Source documentation is minimal, so treat it as a structured analysis workflow rather than a domain-specific stats library integration.

  • Metric aggregation patterns
  • Baseline comparison tables
  • Significance and variance checks
  • Ablation interpretation
  • Go/no-go recommendation framing

Results Analysis by the numbers

  • 360 all-time installs (skills.sh)
  • Ranked #531 of 2,064 Data Science & ML skills by installs in the Skillselion catalog
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/galaxy-dawn/claude-scholar --skill results-analysis

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs360
repo stars5k
Last updatedJuly 17, 2026
Repositorygalaxy-dawn/claude-scholar

How do you interpret ML experiment results?

Interpret experiment outputs—metrics, ablations, significance—and decide whether hypotheses hold before committing to full paper drafts.

Who is it for?

ML researchers and engineers who completed experiment runs and need structured interpretation before committing to a full paper or report draft.

Skip if: Running new training jobs, writing LaTeX manuscripts end-to-end, or production model deployment monitoring without a research paper goal.

When should I use this skill?

A researcher has experiment outputs—metrics, ablations, or significance results—and needs a hypothesis hold-or-fail verdict before drafting.

What you get

Hypothesis verdict, significance assessment, ablation interpretation notes, and go-no-go recommendation for paper drafting.

  • hypothesis verdict
  • significance assessment
  • go-no-go paper recommendation

Files

SKILL.mdMarkdownGitHub ↗

Results Analysis

Run strict, evidence-first experimental analysis for ML/AI research.

Use this skill to produce a strict analysis bundle:

  • analysis-report.md
  • stats-appendix.md
  • figure-catalog.md
  • figures/

When the user asks for review, audit, no-write, dry-run, or when inputs are incomplete, use read-only audit mode instead of producing files or figures. In that mode, output only valid/invalid statistics, blockers, claim candidates, and what evidence is missing. If invoked by /analyze-results, the command layer may write a blocker summary, but this skill should not create figures, reports, or polished conclusions from incomplete evidence.

Do not use this skill to draft a paper Results section or a full experiment wrap-up report. Those belong to ml-paper-writing or results-report.

Core contract

This skill is responsible for

  • validating experiment artifacts and comparison units,
  • running rigorous descriptive and inferential statistics,
  • generating real scientific figures when data/logs are available,
  • writing figure purposes, caption requirements, and interpretation checklists,
  • surfacing limits, blockers, and missing evidence explicitly.

This skill is not responsible for

  • paper-ready Results prose,
  • manuscript narrative polishing,
  • paper-ready figure/table packaging with pubfig / pubtab,
  • project-level experiment retrospectives.

If the user wants the complete post-experiment summary report, hand off to results-report after this bundle is ready. If the user wants publication-grade figures/tables, export parameters, publication QA, or figure/table redesign, hand off to publication-chart-skill.

Non-negotiable quality bar

1. Prefer real figures over figure specs. If the data can be read, generate real figures. Do not stop at “recommended visualization”. Exception: in read-only audit mode, do not generate figures; describe what figure would be valid after evidence is complete. 2. Never fabricate statistics. If sample size, seeds, or raw metrics are missing, state the blocker clearly. 3. Report complete statistics. Do not report only best scores or only p-values. 4. Interpret every main figure. Every major figure must have purpose, caption requirements, and post-figure interpretation notes. 5. Separate evidence from prose. This skill produces analysis artifacts; it does not write manuscript sections.

Standard workflow

1. Inventory and validate artifacts

Start by identifying:

  • metric tables (csv, json, tsv, logs),
  • training curves and checkpoints,
  • seeds / repeated runs,
  • baselines, ablations, and comparison families,
  • evaluation protocol metadata.

Validate:

  • metric direction (higher/lower is better),
  • unit of analysis (run, subject, fold, dataset, seed),
  • number of runs / seeds,
  • missing values or silent failures,
  • comparability across methods.

If the comparison is not statistically valid, say so before continuing. Do not treat repeated subject × task rows, folds, windows, trials, or seeds as independent units unless the design justifies it. Common blocker: a subject × task summary table is usually a repeated-measure summary, not an independent subject-level sample. If subjects have multiple task rows or missing task cells, state that before any significance or winner claim.

2. Lock the comparison questions

Before running statistics, define the exact comparison questions:

  • Which method is compared to which baseline?
  • What is the primary metric?
  • What is the repeated-measure unit?
  • Which ablation or robustness questions matter?
  • Which findings are decision-changing?

Do not mix unrelated comparisons into one undifferentiated table.

3. Run strict statistics

Always produce:

  • descriptive statistics: mean ± std when appropriate,
  • 95% CI or another clearly justified interval,
  • run/seed counts,
  • significance tests with assumptions stated,
  • effect sizes,
  • multiple-comparison handling when several contrasts are reported.

Default expectation:

  • check parametric assumptions first,
  • use non-parametric fallback when assumptions fail,
  • state exactly what was tested and on what samples.

See:

  • references/statistical-methods.md
  • references/statistical-reporting.md

4. Generate real scientific figures

Produce actual figures whenever artifacts are available.

Minimum expectation for a non-trivial analysis bundle:

  • one main comparison figure,
  • one supporting figure (training dynamics / ablation / breakdown / error analysis),
  • one exact numeric summary table in markdown.

Every main figure must define:

  • figure purpose,
  • plotted variables,
  • error bar meaning,
  • caption requirements,
  • interpretation checklist.

See:

  • references/visualization-best-practices.md
  • references/figure-interpretation.md

5. Write analysis artifacts

analysis-report.md

Summarize:

  • the analysis question,
  • key findings,
  • strongest supported comparisons,
  • main caveats,
  • what changed in the experimental understanding,
  • claim candidates that may later be used in reports or manuscript writing.

Each claim candidate should use this shape:

## Claim Candidates

- Claim:
  - Source evidence:
  - Allowed wording:
  - Forbidden stronger wording:
  - Uncertainty:
  - Next check:
  - Decision: keep | weaken | revise | discard
stats-appendix.md

Record:

  • descriptive statistics,
  • test choices,
  • assumptions checked,
  • effect sizes,
  • confidence intervals,
  • multiple comparison corrections,
  • explicit blockers and limitations.
figure-catalog.md

For each figure, record:

  • filename,
  • purpose,
  • data source,
  • caption draft requirements,
  • key observation,
  • interpretation checklist,
  • known caveats.

6. Final QA gate

Do not finish until all are true:

  • [ ] the primary comparison question is explicit,
  • [ ] sample size / seed count is stated,
  • [ ] inferential tests are justified,
  • [ ] effect sizes are reported for major contrasts,
  • [ ] real figures exist when data exists,
  • [ ] each figure has an interpretation note,
  • [ ] limitations and blockers are explicit,
  • [ ] each supported or strong claim candidate has evidence, uncertainty, and allowed wording,
  • [ ] over-strong manuscript wording is explicitly blocked when evidence is insufficient,
  • [ ] no manuscript-style Results draft is included.

Output structure

analysis-output/
├── analysis-report.md
├── stats-appendix.md
├── figure-catalog.md
└── figures/
    ├── figure-01-main-comparison.pdf
    ├── figure-02-ablation.pdf
    └── ...

Figure interpretation rule

For every major figure, answer all three questions: 1. Why does this figure exist? 2. What exactly should the reader notice? 3. What does that observation change in our belief or next decision?

If a figure cannot answer question 3, it is probably decorative rather than scientific.

Read-only audit mode

Use this mode when:

  • the user asks to audit or review existing artifacts,
  • the environment is read-only,
  • the user forbids file writes or figure generation,
  • core evidence is missing.

Return:

  • analysis questions,
  • valid statistics,
  • invalid or unsafe statistics,
  • claim candidates with allowed and forbidden wording,
  • blockers before report/figure generation.

Do not create analysis-output/, figures, or reports in this mode. Quarantine any statistics file whose interpretation contradicts its own p-value, test method, unit of analysis, or comparison family. Do not reuse that file for claim wording until provenance is checked.

Failure mode policy

When inputs are incomplete, say so explicitly.

Examples:

  • no seed-level data -> descriptive summary only; inferential claims blocked,
  • no comparable baseline outputs -> no significance claim,
  • no readable logs -> cannot generate dynamics figure,
  • too few runs -> effect size may be unstable; report this limitation.
  • unclear unit of analysis -> no winner claim or significance claim,
  • analysis file with contradictory interpretation -> quarantine it until provenance is checked.

Never replace missing evidence with confident prose.

Reference files

Load only what is needed:

  • references/statistical-methods.md - test selection and assumptions
  • references/statistical-reporting.md - minimum reporting standard
  • references/visualization-best-practices.md - publication-quality figure rules
  • references/figure-interpretation.md - how to explain figures with evidence
  • references/analysis-depth.md - move from observation to mechanism and decision
  • references/common-pitfalls.md - common analysis and reporting failures
  • ../research-ideation/references/research-contract.md - shared claim candidate and claim strength contract

Example files

  • examples/example-analysis-report.md
  • examples/example-stats-appendix.md
  • examples/example-figure-catalog.md

Related skills

FAQ

What does results-analysis evaluate?

results-analysis evaluates experiment metrics, ablation comparisons, and statistical significance to determine whether stated hypotheses are supported before committing to a full research paper draft.

When should results-analysis run in a research workflow?

results-analysis runs after experiment outputs are available and before manuscript drafting, providing a structured go-no-go verdict on whether results justify paper-level claims.

Data Science & MLanalyticspipelines

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.