
Tooluniverse Sequence Retrieval
- 1.6k installs
- 1.6k repo stars
- Updated August 4, 2026
- mims-harvard/tooluniverse
The tooluniverse-sequence-retrieval skill fetches biological sequences from NCBI and ENA with explicit quality hierarchy preferring RefSeq NM and NP accessions over predicted XM and XP and GenBank submissions.
About
The tooluniverse-sequence-retrieval skill fetches biological sequences from NCBI and ENA with explicit quality hierarchy preferring RefSeq NM and NP accessions over predicted XM and XP and GenBank submissions. It supports accession lookups, gene-symbol disambiguation, transcript isoform selection, and curated versus raw submission choices. Agents explain accession types, report retrieval provenance, and handle ambiguous gene symbols with user confirmation. Use for bioinformatics pipelines needing authoritative sequence sources rather than scraped web pages. NCBI and ENA sequence retrieval with quality hierarchy. Prefers RefSeq NM_/NP_ over predicted XM_/XP_ accessions. Gene-symbol disambiguation and isoform selection. Supports accession, transcript, and curated submission paths. ToolUniverse integration for bioinformatics agents. Retrieve DNA, RNA, and protein sequences from NCBI and ENA with RefSeq quality hierarchy and disambiguation.
- NCBI and ENA sequence retrieval with quality hierarchy.
- Prefers RefSeq NM_/NP_ over predicted XM_/XP_ accessions.
- Gene-symbol disambiguation and isoform selection.
- Supports accession, transcript, and curated submission paths.
- ToolUniverse integration for bioinformatics agents.
Tooluniverse Sequence Retrieval by the numbers
- 1,573 all-time installs (skills.sh)
- +13 installs in the week ending Aug 5, 2026 (Skillselion tracking)
- Ranked #151 of 2,064 Data Science & ML skills by installs in the Skillselion catalog
- Security screen: MEDIUM risk (skills.sh audit)
- Data as of Aug 5, 2026 (Skillselion catalog sync)
tooluniverse-sequence-retrieval capabilities & compatibility
- Capabilities
- ncbi and ena sequence retrieval with quality hie · prefers refseq nm_/np_ over predicted xm_/xp_ ac · gene symbol disambiguation and isoform selection · supports accession, transcript, and curated subm
- Use cases
- data analysis · research
What tooluniverse-sequence-retrieval says it does
Retrieve DNA/RNA/protein sequences from NCBI and ENA with disambiguation.
npx skills add https://github.com/mims-harvard/tooluniverse --skill tooluniverse-sequence-retrievalAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 1.6k |
|---|---|
| repo stars | ★ 1.6k |
| Security audit | 2 / 3 scanners passed |
| Last updated | August 4, 2026 |
| Repository | mims-harvard/tooluniverse ↗ |
How do I apply tooluniverse-sequence-retrieval for the workflow described in SKILL.md?
Retrieve DNA, RNA, and protein sequences from NCBI and ENA with RefSeq quality hierarchy and disambiguation.
Who is it for?
Teams using tooluniverse-sequence-retrieval as documented in the skill repository.
Skip if: Tasks outside the tooluniverse-sequence-retrieval scope defined in SKILL.md.
When should I use this skill?
User mentions tooluniverse-sequence-retrieval or related skill triggers from the description.
What you get
Structured deliverables and steps from the tooluniverse-sequence-retrieval skill workflow.
- validated sequence profile
- database routing decision
By the numbers
- Uses a 5-level curation scale from ●●●● to ○○○○ per sequence
- Covers 4 RefSeq prefix families: NC_, NM_, NP_, XM_
Files
Biological Sequence Retrieval
Retrieve DNA, RNA, and protein sequences with proper disambiguation and cross-database handling.
IMPORTANT: Always use English terms in tool calls. Only try original-language terms as fallback. Respond in the user's language.
LOOK UP DON'T GUESS: Never assume accession numbers or sequence versions. Always retrieve and verify from NCBI or ENA.
Domain Reasoning
Sequence quality hierarchy: RefSeq (NM_/NP_ = curated) > RefSeq predicted (XM_/XP_) > GenBank (submitted). Prefer the MANE Select transcript for human canonical isoforms. Check version numbers -- annotations improve across versions.
Workflow
Phase 0: Clarify (if needed) → Phase 1: Disambiguate Gene/Organism → Phase 2: Search & Retrieve → Phase 3: Report---
Phase 0: Clarification (When Needed)
Ask ONLY if: gene exists in multiple organisms, sequence type unclear, or strain matters. Skip for: specific accessions, clear organism+gene combos, complete genome requests with organism.
---
Phase 1: Gene/Organism Disambiguation
Accession Type Decision Tree
| Prefix | Type | Use With |
|---|---|---|
| NC_/NM_/NR_/NP_/XM_ | RefSeq | NCBI only |
| U/M/K/X/CP*/NZ_ | GenBank | NCBI or ENA |
| EMBL format | EMBL | ENA preferred |
CRITICAL: Never try ENA tools with RefSeq accessions -- they return 404.
Identity Checklist
- Organism confirmed (scientific name)
- Gene symbol/name identified
- Sequence type determined (genomic/mRNA/protein)
- Accession prefix identified for tool selection
---
Phase 2: Data Retrieval (Internal)
Retrieve silently. Do NOT narrate the search process.
# Search NCBI Nucleotide
result = tu.tools.NCBI_search_nucleotide(
operation="search", organism=organism, gene=gene,
strain=strain, keywords=keywords, seq_type=seq_type, limit=10
)
# Get accessions from UIDs
accessions = tu.tools.NCBI_fetch_accessions(operation="fetch_accession", uids=result["data"]["uids"])
# Retrieve sequence (FASTA or GenBank format)
sequence = tu.tools.NCBI_get_sequence(operation="fetch_sequence", accession=accession, format="fasta")
# ENA alternative (non-RefSeq accessions only)
entry = tu.tools.ena_get_entry(accession=accession)
fasta = tu.tools.ena_get_sequence_fasta(accession=accession)Fallback Chains
| Primary | Fallback | Notes |
|---|---|---|
| NCBI_get_sequence | ENA (if GenBank format) | NCBI unavailable |
| ena_get_entry | NCBI_get_sequence | ENA doesn't have RefSeq |
| NCBI_search_nucleotide | Try broader keywords | No results |
---
Phase 3: Report Sequence Profile
Present as a Sequence Profile Report. Hide search process. Include:
1. Search Summary: query, database, result count 2. Primary Sequence: accession, type (RefSeq/GenBank), organism, strain, length, molecule, topology, curation level 3. Sequence Preview: first lines of FASTA (truncated) 4. Annotations Summary: CDS/tRNA/rRNA/regulatory feature counts (from GenBank format) 5. Alternative Sequences: ranked by relevance and curation, with ENA compatibility 6. Cross-Database References: RefSeq, GenBank, ENA/EMBL, BioProject, BioSample 7. Download Options: FASTA (for BLAST/alignment), GenBank (for annotation)
Curation Level Tiers
| Tier | Prefix | Description |
|---|---|---|
| RefSeq Reference (best) | NC_, NM_, NP_ | NCBI-curated, gold standard |
| RefSeq Predicted | XM_, XP_, XR_ | Computationally predicted |
| GenBank Validated | Various | Submitted, some curation |
| GenBank Direct | Various | Direct submission |
| Third Party | TPA_ | Third-party annotation |
---
Reasoning Framework
Sequence quality: Prefer RefSeq over GenBank. Check version numbers. Sequences with "PREDICTED" in definition are not experimentally validated.
Accession guidance: RefSeq = NCBI-only. GenBank = mirrored in ENA/EMBL. Default to RefSeq mRNA (NM_) for human/model organisms; most complete genome assembly for microbial queries.
Cross-database reconciliation: Same sequence may have different accessions (e.g., GenBank U00096 = RefSeq NC_000913 for E. coli K-12). Always report both when available. Discrepancies between GenBank/RefSeq typically indicate RefSeq curation corrected submission errors.
Synthesis Questions
1. What is the highest-quality accession available? 2. Are there alternative accessions in other databases? 3. What is the annotation completeness? 4. Is the sequence from the expected organism/strain? 5. What download format suits the user's downstream analysis?
---
Error Handling
| Error | Response |
|---|---|
| "No search criteria provided" | Add organism, gene, or keywords |
| "ENA 404 error" | Likely RefSeq -- use NCBI only |
| "No results found" | Broaden search, check spelling, try synonyms |
| "Sequence too large" | Note size, provide download link instead |
---
Tool Reference
NCBI Tools: NCBI_search_nucleotide (search), NCBI_fetch_accessions (UID→accession), NCBI_get_sequence (retrieve) ENA Tools (GenBank/EMBL only): ena_get_entry (metadata), ena_get_sequence_fasta (FASTA), ena_get_entry_summary (summary)
---
Search Parameters Reference
NCBI_search_nucleotide: operation="search", organism (scientific name), gene (symbol), strain, keywords, seq_type (complete_genome/mrna/refseq), limit
NCBI_get_sequence: operation="fetch_sequence", accession, format (fasta/genbank)
Sequence Retrieval Checklist
Use this checklist to ensure complete sequence profiles.
Disambiguation
- [ ] Organism confirmed (scientific name)
- [ ] Gene symbol/name identified
- [ ] Sequence type determined (genomic/mRNA/protein)
- [ ] Strain specified (if relevant)
- [ ] Accession prefix identified → tool selection
Accession Type Handling
- [ ] RefSeq (NC_, NM_, NP_, XM_) → NCBI tools only
- [ ] GenBank (U, M, CP*, etc.) → NCBI or ENA
- [ ] ENA tools NOT used with RefSeq accessions
Per Sequence (Required)
- [ ] Accession number
- [ ] Organism (scientific name)
- [ ] Sequence type (DNA/RNA/protein)
- [ ] Length
- [ ] Curation level (●●●●/●●●○/●●○○/●○○○/○○○○)
- [ ] Database source
Sequence Details
- [ ] Definition/title
- [ ] Molecule type (DNA/mRNA/protein)
- [ ] Topology (linear/circular)
- [ ] Sequence preview (first 100-200 bp)
Annotations (If GenBank Format)
- [ ] CDS count and examples
- [ ] Gene count
- [ ] Other features noted
Cross-References
- [ ] RefSeq accession (if exists)
- [ ] GenBank accession
- [ ] ENA compatibility noted
- [ ] BioProject/BioSample links
Download Options
- [ ] FASTA format command shown
- [ ] GenBank format command shown
- [ ] Direct database links provided
Report Quality
- [ ] Search process NOT shown in output
- [ ] Curation level tiers applied
- [ ] Alternative sequences listed
- [ ] Retrieval date included
Error Handling
- [ ] No results → broader search suggested
- [ ] ENA 404 → recognized as RefSeq, NCBI used
- [ ] Large sequences → download link instead of preview
Sequence Retrieval Examples
Example 1: Find E. coli K-12 Genome
from tooluniverse import ToolUniverse
tu = ToolUniverse()
tu.load_tools()
# Search
result = tu.tools.NCBI_search_nucleotide(
operation="search",
organism="Escherichia coli",
strain="K-12",
seq_type="complete_genome",
limit=3
)
# Get accessions
accessions = tu.tools.NCBI_fetch_accessions(
operation="fetch_accession",
uids=result["data"]["uids"]
)
# Get sequence (RefSeq reference)
sequence = tu.tools.NCBI_get_sequence(
operation="fetch_sequence",
accession="NC_000913.3",
format="fasta"
)
print(f"Genome size: {len(sequence['data'])} characters")Example 2: Get Human BRCA1 Gene
# Search for BRCA1
result = tu.tools.NCBI_search_nucleotide(
operation="search",
organism="Homo sapiens",
gene="BRCA1",
limit=5
)
print(f"Found {result['data']['count']} BRCA1 sequences")
# Get top accessions
accessions = tu.tools.NCBI_fetch_accessions(
operation="fetch_accession",
uids=result["data"]["uids"]
)
# Get mRNA sequence with annotations
genbank = tu.tools.NCBI_get_sequence(
operation="fetch_sequence",
accession=accessions["data"][0],
format="genbank"
)Example 3: SARS-CoV-2 Reference Genome
# Search for reference genome
result = tu.tools.NCBI_search_nucleotide(
operation="search",
organism="SARS-CoV-2",
keywords="reference genome Wuhan",
limit=1
)
# Get accession (NC_045512)
accessions = tu.tools.NCBI_fetch_accessions(
operation="fetch_accession",
uids=result["data"]["uids"]
)
# Download complete genome
genome = tu.tools.NCBI_get_sequence(
operation="fetch_sequence",
accession="NC_045512.2",
format="fasta"
)
print(genome["data"][:200]) # PreviewExample 4: Compare RefSeq vs GenBank
# Search returns both types
result = tu.tools.NCBI_search_nucleotide(
operation="search",
organism="Escherichia coli",
strain="K-12",
limit=5
)
accessions = tu.tools.NCBI_fetch_accessions(
operation="fetch_accession",
uids=result["data"]["uids"]
)
# Categorize
refseq = [a for a in accessions["data"] if a.startswith("NC_")]
genbank = [a for a in accessions["data"] if not a.startswith("NC_")]
print(f"RefSeq (NCBI only): {refseq}")
print(f"GenBank (ENA compatible): {genbank}")Example 5: Multi-Format Retrieval
accession = "NC_000913.3"
# FASTA (sequence only)
fasta = tu.tools.NCBI_get_sequence(
operation="fetch_sequence",
accession=accession,
format="fasta"
)
# GenBank (with annotations)
genbank = tu.tools.NCBI_get_sequence(
operation="fetch_sequence",
accession=accession,
format="genbank"
)
# EMBL format
embl = tu.tools.NCBI_get_sequence(
operation="fetch_sequence",
accession=accession,
format="embl"
)Related skills
How it compares
Use when agentic sequence retrieval needs a pre-flight metadata gate; skip when working with pre-validated accession lists.
FAQ
What does tooluniverse-sequence-retrieval do?
Retrieve DNA, RNA, and protein sequences from NCBI and ENA with RefSeq quality hierarchy and disambiguation.
When should I invoke tooluniverse-sequence-retrieval?
Use when you need Retrieve DNA, RNA, and protein sequences from NCBI and ENA with RefSeq quality hierarchy and disambiguation.
What outcome does tooluniverse-sequence-retrieval produce?
The tooluniverse-sequence-retrieval skill fetches biological sequences from NCBI and ENA with explicit quality hierarchy preferring RefSeq NM and NP accessions over predicted XM and XP and GenBank sub.
Is Tooluniverse Sequence Retrieval safe to install?
skills.sh reports 2 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.