Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
mims-harvard avatar

Tooluniverse Literature Deep Research

  • 600 installs
  • 1.6k repo stars
  • Updated August 4, 2026
  • mims-harvard/tooluniverse

tooluniverse-literature-deep-research is an agent research skill that conducts systematic, citation-backed literature reviews across PubMed, EuropePMC, and bioRxiv for developers who need evidence-graded scientific answe

About

tooluniverse-literature-deep-research is a ToolUniverse skill that runs systematic literature workflows with disable-model-invocation set true so retrieval stays tool-driven. It disambiguates queries first, executes collision-aware searches across PubMed, EuropePMC, bioRxiv preprints, and citation networks, then grades every claim on a T1–T4 evidence scale before assembling structured reports with mandatory citations in all sections. Developers invoke it for systematic literature reviews, meta-analysis evidence collection, and detailed answer-with-citations tasks where fabricated references would invalidate output. The skill right-sizes deliverables to the question, enforces evidence grading on each claim, and targets biomedical and scientific domains where database-specific query collisions are common. Reach for it when a coding or research agent must ground recommendations in verifiable papers rather than model memory.

  • Disambiguates queries before searching
  • Collision-aware searches on PubMed, EuropePMC, and bioRxiv
  • Grades every claim using T1-T4 evidence levels
  • Produces mandatory structured reports with full source attribution
  • 7 core principles including 'LOOK UP, DON'T GUESS' and 'COMPUTE, DON'T DESCRIBE'

Tooluniverse Literature Deep Research by the numbers

  • 600 all-time installs (skills.sh)
  • +9 installs in the week ending Aug 4, 2026 (Skillselion tracking)
  • Ranked #1,582 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/mims-harvard/tooluniverse --skill tooluniverse-literature-deep-research

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs600
repo stars1.6k
Last updatedAugust 4, 2026
Repositorymims-harvard/tooluniverse

How do you run systematic literature reviews with citations?

Conduct systematic, citation-backed deep literature reviews across scientific databases without hallucinating sources.

Who is it for?

Developers and researchers building agents that must synthesize biomedical or scientific literature with graded, verifiable citations.

Skip if: Quick coding answers or domains without peer-reviewed database coverage where citation grading adds no value.

When should I use this skill?

A user requests systematic literature review, meta-analysis evidence, or detailed scientific answers that require PubMed-grade citations.

What you get

Structured literature reports with T1–T4 evidence grades and PubMed, EuropePMC, or bioRxiv citations.

  • Structured literature report
  • T1–T4 graded evidence sections
  • Citation-backed claim inventory

Files

SKILL.mdMarkdownGitHub ↗

Literature Deep Research

Systematic literature research: disambiguate, search with collision-aware queries, grade evidence, produce structured reports.

KEY PRINCIPLES: (1) Disambiguate first (2) Right-size deliverable (3) Grade every claim T1-T4 (4) All sections mandatory even if "limited evidence" (5) Source attribution for every claim (6) English-first queries, respond in user's language (7) Report = deliverable, not search log

---

LOOK UP, DON'T GUESS

Search PubMed/EuropePMC FIRST before reasoning. A published paper beats memory.

Factoid search strategy: 1. Extract KEY TERMS (most specific nouns/verbs) 2. EuropePMC_search_articles(query="term1 term2 term3", limit=5) 3. No results -> BROADEN (remove most restrictive term) 4. Too many -> NARROW (add specific terms) 5. Answer usually in abstract of top results 6. Failed query -> try DIFFERENT TERMS/synonyms, don't repeat

---

COMPUTE, DON'T DESCRIBE

When analysis requires computation (statistics, data processing, scoring, enrichment), write and run Python code via Bash. Don't describe what you would do — execute it and report actual results. Use ToolUniverse tools to retrieve data, then Python (pandas, scipy, statsmodels, matplotlib) to analyze it.

Workflow

Phase 0: Clarify + Mode Select → Phase 1: Disambiguate + Profile → Phase 2: Literature Search → Phase 3: Report

---

Phase 0: Mode Selection

ModeWhenDeliverable
FactoidSingle concrete question1-page fact-check report + bibliography
Mini-reviewNarrow topic1-3 page narrative
Full Deep-ResearchComprehensive overview15-section report + bibliography

Factoid Mode (Fast Path)

# [TOPIC]: Fact-check Report
## Question / ## Answer (with evidence rating) / ## Source(s) / ## Verification Notes / ## Limitations

Domain Detection

PatternDomainAction
Gene/protein symbolBiological targetFull bio disambiguation
Drug nameDrugDrug disambiguation (1.5)
Disease nameDiseaseDisease disambiguation (1.6)
CS/ML topicGeneral academicSkip bio tools, literature-only
Cross-domainInterdisciplinaryResolve each entity in its domain

Cross-Skill Delegation

  • Gene/protein deep-dive: tooluniverse-target-research
  • Drug profile: tooluniverse-drug-research
  • Disease profile: tooluniverse-disease-research

Use this skill for literature synthesis. Use specialized skills for entity profiling. For max depth, run both.

---

Phase 1: Subject Disambiguation + Profile

1.1 Biological Target Resolution

UniProt_search → UniProt_get_entry_by_accession → UniProt_id_mapping
ensembl_lookup_gene → MyGene_get_gene_annotation

1.2 Naming Collision Detection

Check first 20 results. If >20% off-topic, build negative filter: NOT [collision1] NOT [collision2]. Gene family: "ADAR" NOT "ADAR2" NOT "ADARB1". Cross-domain: add context terms.

1.3 Baseline Profile (Bio Targets)

InterPro_get_protein_domains, UniProt_get_ptm_processing_by_accession, HPA_get_subcellular_location,
GTEx_get_median_gene_expression, GO_get_annotations_for_gene, Reactome_map_uniprot_to_pathways,
STRING_get_protein_interactions, intact_get_interactions, OpenTargets_get_target_tractability_by_ensemblID

GPCR targets: delegate to tooluniverse-target-research.

1.5 Drug Disambiguation

Identity: OpenTargets_get_drug_chembId_by_generic_name, ChEMBL_get_drug, PubChem_get_CID_by_compound_name, drugbank_get_drug_basic_info_by_drug_name_or_id Targets: ChEMBL_get_drug_mechanisms, OpenTargets_get_associated_targets_by_drug_chemblId, DGIdb_get_drug_gene_interactions Safety: OpenTargets_get_drug_adverse_events_by_chemblId, OpenTargets_get_drug_indications_by_chemblId, search_clinical_trials

1.6 Disease Disambiguation

OpenTargets disease search → EFO/MONDO IDs
DisGeNET_get_disease_genes, DisGeNET_search_disease
CTD_get_disease_chemicals

1.7 Compound Queries (e.g., "metformin in breast cancer")

Resolve both entities, then cross-reference via CTD_get_chemical_gene_interactions, CTD_get_chemical_diseases, OpenTargets drug-target/drug-disease tools. Intersect shared targets/pathways.

1.8 General Academic / 1.9 Interdisciplinary

Non-bio: skip bio tools, use ArXiv/DBLP/OSF. Cross-domain: resolve bio entities with 1.1-1.3, search CS/general in parallel, merge and cross-reference.

---

Phase 2: Literature Search

Methodology stays internal. Report shows findings, not process.

2.1 Query Strategy

Step 1: Seeds (15-30 core papers): domain-specific title searches with date/sort filters. Step 2: Citation expansion: PubMed_get_cited_by, EuropePMC_get_citations/references, PubMed_get_related, SemanticScholar_get_recommendations, OpenCitations_get_citations Step 3: Collision-filtered broader queries: "[TERM]" AND ([context]) NOT [collision]

2.2 Literature Tools — core set + adaptive by domain

Run the core multi-field set on every review (catches what any single index misses), then add the domain rows that match the subject. Don't fire every source blindly — 6–10 well-chosen indexes beat 20 noisy ones.

ALWAYS run (core, all disciplines): PubMed_search_articles, EuropePMC_search_articles, openalex_search_works (query param search/query) or openalex_literature_search (query param search_keywords) — pick one and match its param; mixing them silently returns off-topic results — and SemanticScholar_search_papers

Then add by domain:

DomainAdd theseNotes
Biomedical / clinicalPMC_search_papers (full text), PubTator3_LiteratureSearch (entity & relations: queries), PubMed_Guidelines_Search (clinical guidelines)PubTator normalizes gene/drug/disease entities
Biology (ecology/evolution/plant)EuropePMC as PRIMARY + OpenAlexPubMed returns 0–1 for non-clinical biology
CS / ML / AIArXiv_search_papers, DBLP_search_publicationsarXiv + CS bibliography
Physics / HEP / astroInspireHEP_search_papers1.6M+ particle/astro records
Broad / hard-to-find / OACrossref_search_works, CORE_search_papers, DOAJ_search_articles, Fatcat_search_scholarDOI registry + OA aggregators + Internet Archive Scholar
Regional / EU-fundedOpenAIRE_search_publications, HAL_search_archiveEU open science + French national archive
Datasets / software / outputsFigshare_search_articles, Zenodo_search_recordsCitable DOIs for data & code
Preprints (latest)EuropePMC_search_articles(source='PPR'), OSF_search_preprints, BioRxiv_get_preprint/MedRxiv_get_preprint (DOI lookup)bioRxiv/medRxiv/PsyArXiv etc.

Multi-source: advanced_literature_search_agent (12+ DBs; needs Azure key -- fallback: query the core set individually). Citation impact: iCite_search_publications (RCR/APT), iCite_get_publications (by PMID), scite_get_tallies (support/contradict). PubMed-only; for CS use SemanticScholar.

A domain-specific index returning 0 (e.g. ArXiv on a pure-clinical topic) is normal — only worry if the whole core set is empty.

2.3-2.4 Full-Text & PubMed Zero-Result Fallback

Full-text: see FULLTEXT_STRATEGY.md for three-tier strategy.

CRITICAL: PubMed returns 0 for ~30% of valid queries. Always retry with EuropePMC when PubMed returns empty. This is not optional.

2.5 Tool Failure / OA Handling

Retry once -> fallback tool. Key fallbacks: PubMed_get_cited_by -> EuropePMC_get_citations -> OpenCitations. OA: Unpaywall if configured, else Europe PMC/PMC/OpenAlex flags.

---

Phase 3: Evidence Grading

TierLabelBio ExampleCS/ML Example
T1MechanisticCRISPR KO + rescue, RCTFormal proof, controlled ablation
T2FunctionalsiRNA knockdown phenotypeBenchmark with baselines
T3AssociationGWAS, screen hitObservational, case study
T4MentionReview articleSurvey, workshop abstract

Inline: Target X regulates Y [T1: PMID:12345678]. Per theme: summarize evidence distribution.

---

Report Output

FileMode
[topic]_report.mdFull
[topic]_factcheck_report.mdFactoid
[topic]_bibliography.json + .csvAll

Progressive update: create report with all section headers immediately. Fill after each phase. Write Executive Summary LAST.

Use 15-section template from REPORT_TEMPLATE.md. Domain adaptations: bio (architecture/expression/GO/disease), drug (properties/MOA/PK/safety), disease (epi/patho/genes/treatments), general (history/theories/evidence/applications).

---

Communication

Brief progress updates only: "Resolving identifiers...", "Building paper set...", "Grading evidence..." Do NOT expose: raw tool outputs, dedup counts, search round details.

---

References

  • TOOL_NAMES_REFERENCE.md -- 123 tools with parameters
  • REPORT_TEMPLATE.md -- template, domain adaptations, bibliography, completeness checklist
  • FULLTEXT_STRATEGY.md -- three-tier full-text verification
  • WORKFLOW.md -- compact cheat-sheet
  • EXAMPLES.md -- worked examples

Related skills

How it compares

Use tooluniverse-literature-deep-research when citations and evidence tiers must be enforced across scientific databases, not for general web summarization.

FAQ

Which databases does tooluniverse-literature-deep-research search?

tooluniverse-literature-deep-research queries PubMed, EuropePMC, bioRxiv preprints, and citation networks using collision-aware searches after query disambiguation.

What does T1–T4 evidence grading mean in this skill?

tooluniverse-literature-deep-research assigns every claim a T1–T4 evidence tier so structured reports show how strongly each statement is supported by retrieved literature rather than model speculation.

AI & Agent Buildingresearchautomation

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.