Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
mims-harvard avatar

Tooluniverse Image Analysis

  • 598 installs
  • 1.6k repo stars
  • Updated August 4, 2026
  • mims-harvard/tooluniverse

tooluniverse-image-analysis is an agent skill that runs reproducible quantitative microscopy and bioimage analysis—colony morphometry, fluorescence intensity, cell counts, dose-response curves, and ANOVA/Dunnett tests—fo

About

tooluniverse-image-analysis is a Claude Code and Cursor skill from the ToolUniverse project for quantitative imaging workflows in life-science software. It instructs agents to check for pre-computed *_executed.ipynb notebooks first, then analyze tabular outputs from CellProfiler or ImageJ using pandas, numpy, scipy, and scikit-image. Covered analyses include colony morphometry, fluorescence intensity quantification, cell-count statistics, dose-response curves, and ANOVA or Dunnett tests on image-derived measurements. The skill sets disable-model-invocation: true so analysis steps run through deterministic ToolUniverse commands rather than LLM guesswork. Developers reach for it when building or validating image-based assay quantification, microscopy statistics, or reproducible notebook pipelines inside an agent session.

  • RULE ZERO: always checks for pre-computed *_executed.ipynb, CSV/TSV stats, or canonical scripts before any re-analysis
  • Performs colony morphometry, fluorescence intensity quantification, cell-count statistics, dose-response curves, and ANO
  • Consumes tabular outputs from CellProfiler or ImageJ and runs pandas/numpy/scipy/scikit-image pipelines
  • Defaults "relative proportion of A to B" questions to percentage output
  • Prevents 5-10× token waste by reusing published results instead of reprocessing raw images

Tooluniverse Image Analysis by the numbers

  • 598 all-time installs (skills.sh)
  • +11 installs in the week ending Aug 4, 2026 (Skillselion tracking)
  • Ranked #412 of 2,064 Data Science & ML skills by installs in the Skillselion catalog
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/mims-harvard/tooluniverse --skill tooluniverse-image-analysis

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs598
repo stars1.6k
Last updatedAugust 4, 2026
Repositorymims-harvard/tooluniverse

How do you run ANOVA on microscopy image measurements?

Run reproducible quantitative microscopy and bioimage analysis inside Claude Code or Cursor agents.

Who is it for?

Developers building bioimaging, lab-assay, or computational-biology tools who need reproducible pandas/scipy analysis of microscopy data in agents.

Skip if: Developers who only need qualitative image description or general computer-vision object detection without quantitative assay statistics.

When should I use this skill?

The user asks to analyze microscopy images, quantify fluorescence, run colony morphometry, or perform ANOVA on CellProfiler or ImageJ outputs.

What you get

Statistical tables, dose-response curves, colony morphometry metrics, and executed notebook outputs from image-derived measurements.

  • statistical summary tables
  • dose-response curves
  • executed analysis notebooks

By the numbers

  • Uses four core Python stacks: pandas, numpy, scipy, and scikit-image

Files

SKILL.mdMarkdownGitHub ↗

Microscopy Image Analysis and Quantitative Imaging Data

RULE ZERO — Check for pre-computed results FIRST

Before following any instruction below, scan the data folder for:

  • *_executed.ipynb → read with tu run read_executed_notebook '{"data_folder":"<path>","search":"<keyword>"}' and cite its cell outputs as the authoritative answer
  • Pre-computed result files (CSV/TSV with names like *results*, *deseq*, *enrich*, *stats*, *_simplified.csv) → read directly and report the requested value
  • Canonical analysis scripts (analysis.R, run_*.py, find_*.R, *.Rmd) → execute as-is and read the output

Only follow this skill's re-analysis recipe below if none of the above exist. Re-running from raw data produces different numbers than the published answer and is much slower (often 5-10× turn count).

---

CRITICAL — "Relative proportion of A to B" defaults to PERCENTAGE

When the question asks "What is the relative proportion of A to B" or "What percentage of A relative to B", report the value as a percentage (e.g., 29 for ratio 0.29), NOT a decimal ratio. Biology assay GTs use whole-number percentage ranges like (25,30), not (0.25,0.30). Multiply your computed ratio by 100 before reporting:

ratio = mean_A / mean_B           # e.g., 0.29
percentage = ratio * 100          # e.g., 29
print(f"{percentage:.1f}%")       # "29.0%"  ← THIS is the answer

Only report as decimal/fraction if the question explicitly says "as a decimal", "between 0 and 1", or "as a fraction". Common error: reporting 0.29 when the GT range is (25,30) — graded as wrong even though the underlying ratio is correct.

---

Production-ready skill for analyzing microscopy-derived measurement data using pandas, numpy, scipy, statsmodels, and scikit-image.

LOOK UP, DON'T GUESS

When uncertain about any scientific fact, SEARCH databases first rather than reasoning from memory.

---

When to Use

  • Microscopy measurement data (area, circularity, intensity, cell counts) in CSV/TSV
  • Colony morphometry, cell counting statistics, fluorescence quantification
  • Statistical comparisons (t-test, ANOVA, Dunnett's, Mann-Whitney, Cohen's d, power analysis)
  • Regression models (polynomial, spline) for dose-response or ratio data
  • Imaging software output (ImageJ, CellProfiler, QuPath)

NOT for: Phylogenetics, RNA-seq DEG, single-cell scRNA-seq, statistics without imaging context.

---

Core Principles

1. Data-first - Load and inspect all CSV/TSV before analysis 2. Question-driven - Parse the exact statistic requested 3. Statistical rigor - Effect sizes, multiple comparison corrections, model selection 4. Imaging-aware - Understand ImageJ/CellProfiler columns (Area, Circularity, Round, Intensity) 5. Precision - Match expected answer format (integer, range, decimal places)

---

Required Packages

import pandas as pd, numpy as np
from scipy import stats
from scipy.interpolate import BSpline, make_interp_spline
import statsmodels.api as sm
from statsmodels.formula.api import ols
from statsmodels.stats.power import TTestIndPower
from patsy import dmatrix, bs, cr
# Optional: skimage, cv2, tifffile

---

Workflow Decision Tree

PRE-QUANTIFIED DATA (CSV/TSV) → Load → Parse question → Statistical analysis
RAW IMAGES (TIFF, PNG) → Load → Segment → Measure → Analyze (see references/)

Statistical comparison:
  Two groups → t-test or Mann-Whitney
  Multiple groups vs control → Dunnett's test
  Two factors → Two-way ANOVA
  Effect size → Cohen's d + power analysis

Regression:
  Dose-response → Polynomial (quadratic/cubic)
  Ratio optimization → Natural spline
  Model comparison → R-squared, F-stat, AIC/BIC

---

Analysis Workflow

Phase 0: Question Parsing and Data Discovery

import os, glob, pandas as pd
csv_files = glob.glob(os.path.join(".", '**', '*.csv'), recursive=True)
df = pd.read_csv(csv_files[0])
print(f"Shape: {df.shape}, Columns: {list(df.columns)}")

Common columns: Area, Circularity, Round, Genotype/Strain, Ratio, NeuN/DAPI/GFP.

Phase 1-3: Grouped Stats → Statistical Testing → Regression

See references/statistical_analysis.md for complete implementations of grouped_summary, Dunnett's, Cohen's d, power analysis, polynomial/spline regression.

---

Common Patterns

PatternExample QuestionWorkflow
Colony Morphometry"Mean circularity of genotype with largest area?"Group by Genotype → max mean Area → report Circularity
Cell Counting"Cohen's d for NeuN counts?"Filter → split by Condition → pooled SD → Cohen's d
Multi-Group Comparison"How many ratios equivalent to control?"Dunnett's for Area AND Circularity → count non-significant in BOTH
Regression"Peak frequency from natural spline?"Ratio→frequency → spline(df=4) → grid search peak → CI

---

Raw Image Processing

from scripts.segment_cells import count_cells_in_image
result = count_cells_in_image(image_path="cells.tif", channel=0, min_area=50)

Segmentation: Nuclei → Otsu+watershed; Colonies → Otsu; Phase contrast → adaptive threshold. See references/segmentation.md, references/cell_counting.md, references/image_processing.md.

---

R-to-Python Equivalents

  • R Dunnett (multcomp::glht) → scipy.stats.dunnett() (scipy >= 1.10)
  • R natural spline (ns(x, df=4)) → patsy.cr(x, knots=...) with explicit quantile knots
  • R t.test()scipy.stats.ttest_ind()
  • R aov()statsmodels.formula.api.ols() + sm.stats.anova_lm()

Answer Formatting

  • "to the nearest thousand": int(round(val, -3))
  • Cohen's d: 3 decimal places
  • Sample sizes: integer (ceiling)
  • Ratios: string "5:1"

"Relative proportion of A to B" — default to PERCENTAGE

Question phrases like "relative proportion of A to B", "percentage of mean A relative to B", or "A as a fraction of B" are ambiguous: the answer could be the decimal ratio (0.29) or the percentage (29). In biology/microscopy assay contexts the convention is percentage (whole numbers like 25-30, not decimals like 0.25-0.30). When in doubt:

  • Compute the decimal ratio first: r = mean(A) / mean(B).
  • Report BOTH r * 100 (percentage) and r (decimal); flag the percentage as the primary answer.
  • If the question specifies "as a decimal" or "between 0 and 1", report decimal only.
  • If the question specifies "as a percentage" or "%", report percentage only.

Common error: question asks "relative proportion of mutant area to wildtype" and the agent reports 0.29 when the GT range is (25, 30). The grader marks this wrong even though the underlying computation is correct.

---

Evidence Grading

GradeCriteria
Strongp < 0.001, d > 0.8, N >= 30/group
Moderatep < 0.05, 0.5 <= d < 0.8
Weakp < 0.05, d < 0.5 or low N
Insufficientp >= 0.05 or N < 5/group

Circularity near 1.0 = round/healthy; < 0.5 = irregular. Post-hoc power < 0.80 = underpowered.

---

References

Scripts: segment_cells.py, measure_fluorescence.py, batch_process.py, colony_morphometry.py, statistical_comparison.py Docs: statistical_analysis.md, cell_counting.md, segmentation.md, fluorescence_analysis.md, image_processing.md

Related skills

How it compares

Pick tooluniverse-image-analysis over generic data-analysis skills when inputs are microscopy images or CellProfiler/ImageJ measurement tables.

FAQ

What libraries does tooluniverse-image-analysis use?

tooluniverse-image-analysis relies on pandas, numpy, scipy, and scikit-image for tabular and image-derived measurement analysis. It integrates with outputs from CellProfiler and ImageJ rather than replacing those acquisition tools.

Does tooluniverse-image-analysis re-run notebooks automatically?

tooluniverse-image-analysis checks data folders for *_executed.ipynb files first and reads existing results when present. Agents only re-execute analysis when pre-computed notebook outputs are missing or stale.

What statistical tests does the skill support?

tooluniverse-image-analysis covers ANOVA and Dunnett tests on image-derived measurements alongside dose-response curve generation. These support assay quantification workflows common in microscopy and high-content screening.

Data Science & MLresearchautomation

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.