Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
mims-harvard avatar

Tooluniverse Metagenomics Analysis

  • 203 installs
  • 1.6k repo stars
  • Updated August 4, 2026
  • mims-harvard/tooluniverse

Microbiome and environmental shotgun metagenomics profiling—taxonomy, functional annotation, and comparative community analysis via agent-orchestrated ToolUniverse calls.

About

ToolUniverse metagenomics analysis skill equipping Claude Code agents to run shotgun metagenomics workflows: sample QC, taxonomic classification, functional annotation, and cross-cohort comparisons for microbiome and environmental sequencing projects.

  • Taxonomic and functional metagenomics profiling
  • Supports comparative microbiome community analysis
  • Agent-callable ToolUniverse scientific endpoints
  • Useful for environmental and clinical microbiome studies
  • Orchestrates QC-to-annotation analysis chains

Tooluniverse Metagenomics Analysis by the numbers

  • 203 all-time installs (skills.sh)
  • +6 installs in the week ending Aug 4, 2026 (Skillselion tracking)
  • Ranked #647 of 2,064 Data Science & ML skills by installs in the Skillselion catalog
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/mims-harvard/tooluniverse --skill tooluniverse-metagenomics-analysis

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs203
repo stars1.6k
Last updatedAugust 4, 2026
Repositorymims-harvard/tooluniverse

What it does

Microbiome and environmental shotgun metagenomics profiling—taxonomy, functional annotation, and comparative community analysis via agent-orchestrated ToolUniverse calls.

Files

SKILL.mdMarkdownGitHub ↗

Metagenomics & Microbiome Analysis

Integrated pipeline for exploring microbiome studies, classifying taxa, assessing genome quality, linking microbial composition to clinical phenotypes, and interpreting findings through pathway analysis and literature context.

Guiding principles: 1. Study context first -- understand biome, sequencing method, and metadata before diving into taxa 2. Taxonomic consistency -- GTDB taxonomy as reference standard; reconcile NCBI where needed 3. Genome quality matters -- CheckM completeness/contamination thresholds determine trustworthy MAGs 4. Interpretation over enumeration -- explain what taxa mean for the biological question 5. English-first queries -- use English terms in tool calls

LOOK UP, DON'T GUESS

When uncertain about any scientific fact, SEARCH databases first rather than reasoning from memory.

---

COMPUTE, DON'T DESCRIBE

When analysis requires computation (statistics, data processing, scoring, enrichment), write and run Python code via Bash. Don't describe what you would do — execute it and report actual results. Use ToolUniverse tools to retrieve data, then Python (pandas, scipy, statsmodels, matplotlib) to analyze it.

Core Databases

DatabaseBest For
MGnifyProcessed metagenomics studies, taxonomic/functional results
GTDBStandardized bacterial/archaeal taxonomy, species-level resolution
GMrepoGut species-to-human-health phenotype associations
ENARaw sequencing datasets and study metadata
KEGGPathway mapping for microbial functional annotations
PubMed/EuropePMCPublished microbiome-disease studies
CTDChemical-microbiome-disease relationships

---

Workflow

Phase 0: Parse query → organism, biome, phenotype, or accession
Phase 1: Study Discovery → MGnify_search_studies, ENAPortal_search_studies
Phase 2: Taxonomic Classification → GTDB_search_genomes, GTDB_get_species, GTDB_search_taxon
Phase 3: Genome Quality → MGnify_search_genomes, MGnify_get_genome (CheckM metrics)
Phase 4: Functional Annotation → MGnify GO terms + KEGG pathway mapping
Phase 5: Clinical Associations → GMrepo species-phenotype links
Phase 6: Literature → PubMed/EuropePMC + CTD gene-disease
Phase 7: Interpretation & Report Synthesis

---

Key Phase Notes

Phase 1: ENA requires structured queries (e.g., study_title="*IBD*"), not free text. If ENA fails, fall back to MGnify.

Phase 2: GTDB uses its own naming (e.g., s__Bacteroides_A fragilis vs NCBI Bacteroides fragilis). Always note discrepancies. Use GTDB_search_taxon(operation="search_taxon", query=name).

Phase 3 - Quality tiers (MIMAG):

  • High: >= 90% complete, <= 5% contamination, rRNA + >= 18 tRNAs
  • Medium: >= 50% complete, <= 10% contamination
  • Low: below medium -- flag but don't exclude

Phase 4 - Functional interpretation: Don't just list GO terms. Connect to biology:

Functional CategoryKey KEGG PathwaysSignificance
SCFA productionmap00650, map00640Gut barrier, anti-inflammatory
LPS biosynthesismap00540Pro-inflammatory, endotoxemia
Bile acid metabolismmap00120Fat absorption, FXR signaling
Tryptophan metabolismmap00380Serotonin, AhR, immune
Vitamin biosynthesismap00730/740/760Host nutritional contribution

Use kegg_search_pathway(keyword=...) (NOT query). Pathway IDs need organism prefix (hsa, ko, eco), NOT bare map.

Phase 5: GMrepo uses MeSH terms: "Crohn Disease" not "IBD", "Colitis, Ulcerative" not "UC", "Colorectal Neoplasms" not "colorectal cancer". Try NCBI taxon IDs if species name fails.

Phase 6 - Evidence grading:

  • Strong: Meta-analysis or >5 studies, consistent direction
  • Moderate: 2-5 studies consistent, or 1 large cohort
  • Preliminary: Single study or conflicting
  • Mechanistic only: In vitro/animal, no human epidemiology

Phase 7 - Report: Executive summary, study landscape, GTDB taxonomy, functional interpretation (not GO term lists), clinical relevance with evidence grades, mechanistic model, genome catalog with quality tiers, data gaps.

---

Edge Cases & Fallbacks

  • Taxon not in GTDB: Try partial search or fall back to MGnify (NCBI taxonomy)
  • No GMrepo data: Normal for non-gut organisms; use literature
  • GMrepo 0 results: Use formal MeSH terms or NCBI taxon IDs
  • No KEGG match: Check MetaCyc or literature

Limitations

  • GMrepo: Gut-only
  • GTDB: Bacteria/Archaea only
  • ENA: Raw data only, strict query syntax
  • No sequence analysis: Queries databases, not raw FASTQ/FASTA

Related skills

Data Science & MLanalyticspipelines

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.