Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
mims-harvard avatar

Tooluniverse Sequence Retrieval

  • 1.6k installs
  • 1.6k repo stars
  • Updated August 4, 2026
  • mims-harvard/tooluniverse

The tooluniverse-sequence-retrieval skill fetches biological sequences from NCBI and ENA with explicit quality hierarchy preferring RefSeq NM and NP accessions over predicted XM and XP and GenBank submissions.

About

The tooluniverse-sequence-retrieval skill fetches biological sequences from NCBI and ENA with explicit quality hierarchy preferring RefSeq NM and NP accessions over predicted XM and XP and GenBank submissions. It supports accession lookups, gene-symbol disambiguation, transcript isoform selection, and curated versus raw submission choices. Agents explain accession types, report retrieval provenance, and handle ambiguous gene symbols with user confirmation. Use for bioinformatics pipelines needing authoritative sequence sources rather than scraped web pages. NCBI and ENA sequence retrieval with quality hierarchy. Prefers RefSeq NM_/NP_ over predicted XM_/XP_ accessions. Gene-symbol disambiguation and isoform selection. Supports accession, transcript, and curated submission paths. ToolUniverse integration for bioinformatics agents. Retrieve DNA, RNA, and protein sequences from NCBI and ENA with RefSeq quality hierarchy and disambiguation.

  • NCBI and ENA sequence retrieval with quality hierarchy.
  • Prefers RefSeq NM_/NP_ over predicted XM_/XP_ accessions.
  • Gene-symbol disambiguation and isoform selection.
  • Supports accession, transcript, and curated submission paths.
  • ToolUniverse integration for bioinformatics agents.

Tooluniverse Sequence Retrieval by the numbers

  • 1,573 all-time installs (skills.sh)
  • +13 installs in the week ending Aug 5, 2026 (Skillselion tracking)
  • Ranked #151 of 2,064 Data Science & ML skills by installs in the Skillselion catalog
  • Security screen: MEDIUM risk (skills.sh audit)
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
At a glance

tooluniverse-sequence-retrieval capabilities & compatibility

Capabilities
ncbi and ena sequence retrieval with quality hie · prefers refseq nm_/np_ over predicted xm_/xp_ ac · gene symbol disambiguation and isoform selection · supports accession, transcript, and curated subm
Use cases
data analysis · research
From the docs

What tooluniverse-sequence-retrieval says it does

Retrieve DNA/RNA/protein sequences from NCBI and ENA with disambiguation.
SKILL.md
npx skills add https://github.com/mims-harvard/tooluniverse --skill tooluniverse-sequence-retrieval

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs1.6k
repo stars1.6k
Security audit2 / 3 scanners passed
Last updatedAugust 4, 2026
Repositorymims-harvard/tooluniverse

How do I apply tooluniverse-sequence-retrieval for the workflow described in SKILL.md?

Retrieve DNA, RNA, and protein sequences from NCBI and ENA with RefSeq quality hierarchy and disambiguation.

Who is it for?

Teams using tooluniverse-sequence-retrieval as documented in the skill repository.

Skip if: Tasks outside the tooluniverse-sequence-retrieval scope defined in SKILL.md.

When should I use this skill?

User mentions tooluniverse-sequence-retrieval or related skill triggers from the description.

What you get

Structured deliverables and steps from the tooluniverse-sequence-retrieval skill workflow.

  • validated sequence profile
  • database routing decision

By the numbers

  • Uses a 5-level curation scale from ●●●● to ○○○○ per sequence
  • Covers 4 RefSeq prefix families: NC_, NM_, NP_, XM_

Files

SKILL.mdMarkdownGitHub ↗

Biological Sequence Retrieval

Retrieve DNA, RNA, and protein sequences with proper disambiguation and cross-database handling.

IMPORTANT: Always use English terms in tool calls. Only try original-language terms as fallback. Respond in the user's language.

LOOK UP DON'T GUESS: Never assume accession numbers or sequence versions. Always retrieve and verify from NCBI or ENA.

Domain Reasoning

Sequence quality hierarchy: RefSeq (NM_/NP_ = curated) > RefSeq predicted (XM_/XP_) > GenBank (submitted). Prefer the MANE Select transcript for human canonical isoforms. Check version numbers -- annotations improve across versions.

Workflow

Phase 0: Clarify (if needed) → Phase 1: Disambiguate Gene/Organism → Phase 2: Search & Retrieve → Phase 3: Report

---

Phase 0: Clarification (When Needed)

Ask ONLY if: gene exists in multiple organisms, sequence type unclear, or strain matters. Skip for: specific accessions, clear organism+gene combos, complete genome requests with organism.

---

Phase 1: Gene/Organism Disambiguation

Accession Type Decision Tree

PrefixTypeUse With
NC_/NM_/NR_/NP_/XM_RefSeqNCBI only
U/M/K/X/CP*/NZ_GenBankNCBI or ENA
EMBL formatEMBLENA preferred

CRITICAL: Never try ENA tools with RefSeq accessions -- they return 404.

Identity Checklist

  • Organism confirmed (scientific name)
  • Gene symbol/name identified
  • Sequence type determined (genomic/mRNA/protein)
  • Accession prefix identified for tool selection

---

Phase 2: Data Retrieval (Internal)

Retrieve silently. Do NOT narrate the search process.

# Search NCBI Nucleotide
result = tu.tools.NCBI_search_nucleotide(
    operation="search", organism=organism, gene=gene,
    strain=strain, keywords=keywords, seq_type=seq_type, limit=10
)

# Get accessions from UIDs
accessions = tu.tools.NCBI_fetch_accessions(operation="fetch_accession", uids=result["data"]["uids"])

# Retrieve sequence (FASTA or GenBank format)
sequence = tu.tools.NCBI_get_sequence(operation="fetch_sequence", accession=accession, format="fasta")

# ENA alternative (non-RefSeq accessions only)
entry = tu.tools.ena_get_entry(accession=accession)
fasta = tu.tools.ena_get_sequence_fasta(accession=accession)

Fallback Chains

PrimaryFallbackNotes
NCBI_get_sequenceENA (if GenBank format)NCBI unavailable
ena_get_entryNCBI_get_sequenceENA doesn't have RefSeq
NCBI_search_nucleotideTry broader keywordsNo results

---

Phase 3: Report Sequence Profile

Present as a Sequence Profile Report. Hide search process. Include:

1. Search Summary: query, database, result count 2. Primary Sequence: accession, type (RefSeq/GenBank), organism, strain, length, molecule, topology, curation level 3. Sequence Preview: first lines of FASTA (truncated) 4. Annotations Summary: CDS/tRNA/rRNA/regulatory feature counts (from GenBank format) 5. Alternative Sequences: ranked by relevance and curation, with ENA compatibility 6. Cross-Database References: RefSeq, GenBank, ENA/EMBL, BioProject, BioSample 7. Download Options: FASTA (for BLAST/alignment), GenBank (for annotation)

Curation Level Tiers

TierPrefixDescription
RefSeq Reference (best)NC_, NM_, NP_NCBI-curated, gold standard
RefSeq PredictedXM_, XP_, XR_Computationally predicted
GenBank ValidatedVariousSubmitted, some curation
GenBank DirectVariousDirect submission
Third PartyTPA_Third-party annotation

---

Reasoning Framework

Sequence quality: Prefer RefSeq over GenBank. Check version numbers. Sequences with "PREDICTED" in definition are not experimentally validated.

Accession guidance: RefSeq = NCBI-only. GenBank = mirrored in ENA/EMBL. Default to RefSeq mRNA (NM_) for human/model organisms; most complete genome assembly for microbial queries.

Cross-database reconciliation: Same sequence may have different accessions (e.g., GenBank U00096 = RefSeq NC_000913 for E. coli K-12). Always report both when available. Discrepancies between GenBank/RefSeq typically indicate RefSeq curation corrected submission errors.

Synthesis Questions

1. What is the highest-quality accession available? 2. Are there alternative accessions in other databases? 3. What is the annotation completeness? 4. Is the sequence from the expected organism/strain? 5. What download format suits the user's downstream analysis?

---

Error Handling

ErrorResponse
"No search criteria provided"Add organism, gene, or keywords
"ENA 404 error"Likely RefSeq -- use NCBI only
"No results found"Broaden search, check spelling, try synonyms
"Sequence too large"Note size, provide download link instead

---

Tool Reference

NCBI Tools: NCBI_search_nucleotide (search), NCBI_fetch_accessions (UID→accession), NCBI_get_sequence (retrieve) ENA Tools (GenBank/EMBL only): ena_get_entry (metadata), ena_get_sequence_fasta (FASTA), ena_get_entry_summary (summary)

---

Search Parameters Reference

NCBI_search_nucleotide: operation="search", organism (scientific name), gene (symbol), strain, keywords, seq_type (complete_genome/mrna/refseq), limit

NCBI_get_sequence: operation="fetch_sequence", accession, format (fasta/genbank)

Related skills

How it compares

Use when agentic sequence retrieval needs a pre-flight metadata gate; skip when working with pre-validated accession lists.

FAQ

What does tooluniverse-sequence-retrieval do?

Retrieve DNA, RNA, and protein sequences from NCBI and ENA with RefSeq quality hierarchy and disambiguation.

When should I invoke tooluniverse-sequence-retrieval?

Use when you need Retrieve DNA, RNA, and protein sequences from NCBI and ENA with RefSeq quality hierarchy and disambiguation.

What outcome does tooluniverse-sequence-retrieval produce?

The tooluniverse-sequence-retrieval skill fetches biological sequences from NCBI and ENA with explicit quality hierarchy preferring RefSeq NM and NP accessions over predicted XM and XP and GenBank sub.

Is Tooluniverse Sequence Retrieval safe to install?

skills.sh reports 2 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.

Data Science & MLanalyticsdatabases

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.