
Tooluniverse Target Research
- 358 installs
- 1.6k repo stars
- Updated August 4, 2026
- mims-harvard/tooluniverse
tooluniverse-target-research is an agent skill that identifies and prioritizes therapeutic drug targets from biomedical literature, omics data, and pathway databases during early discovery planning.
About
tooluniverse-target-research is a Harvard ToolUniverse agent skill for comprehensive drug-target intelligence profiling. It resolves gene symbols, UniProt accessions, or Ensembl IDs, then explores parallel research paths covering Open Targets foundation data, protein structure, pathways, STRING interactions, GTEx and HPA expression, ClinVar and gnomAD variants, DGIdb and ChEMBL druggability, and collision-aware literature search. The workflow enforces report-first output, evidence grading T1 through T4, inline citations, and mandatory section completeness in a target report markdown file. Computational biologists and drug-discovery developers reach for tooluniverse-target-research when assessing targets like EGFR or KRAS for validation, druggability, and early pipeline prioritization.
- Harvard ToolUniverse biomedical tooling
- Drug target prioritization workflows
- Literature and pathway evidence synthesis
- Agent-driven scientific research queries
- Early-stage discovery decision support
Tooluniverse Target Research by the numbers
- 358 all-time installs (skills.sh)
- +5 installs in the week ending Aug 4, 2026 (Skillselion tracking)
- Ranked #536 of 2,064 Data Science & ML skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/mims-harvard/tooluniverse --skill tooluniverse-target-researchAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 358 |
|---|---|
| repo stars | ★ 1.6k |
| Last updated | August 4, 2026 |
| Repository | mims-harvard/tooluniverse ↗ |
How do you profile a drug target across omics databases?
Identify and prioritize therapeutic drug targets from biomedical literature, omics data, and pathway databases during early discovery planning.
Who is it for?
Computational biologists and drug-discovery developers who need multi-database target profiles with graded evidence before assay or pipeline investment.
Skip if: Application developers building unrelated SaaS features who do not need biomedical target validation or druggability assessment.
When should I use this skill?
The user asks to assess a gene or protein as a drug target, validate druggability, or produce a comprehensive target profile report.
What you get
Cited target intelligence report with expression, pathways, interactions, variant landscape, druggability assessment, and T1-T4 evidence grades.
- target intelligence report
- druggability assessment
- evidence-graded citations
By the numbers
- Covers 9 parallel research paths for target intelligence
- Uses T1-T4 evidence grading tiers for all factual claims
Files
Comprehensive Target Intelligence Gatherer
Gather complete target intelligence by exploring 9 parallel research paths. Supports targets identified by gene symbol, UniProt accession, Ensembl ID, or gene name.
KEY PRINCIPLES: 1. Report-first approach - Create report file FIRST, then populate progressively 2. Tool parameter verification - Verify params via get_tool_info before calling unfamiliar tools 3. Evidence grading - Grade all claims by evidence strength (T1-T4) 4. Citation requirements - Every fact must have inline source attribution 5. Mandatory completeness - All sections must exist with data minimums or explicit "No data" notes 6. Disambiguation first - Resolve all identifiers before research 7. Negative results documented - "No drugs found" is data; empty sections are failures 8. Collision-aware literature search - Detect and filter naming collisions 9. English-first queries - Always use English terms in tool calls, even if the user writes in another language. Translate gene names, disease names, and search terms to English. Only try original-language terms as a fallback if English returns no results. Respond in the user's language
---
LOOK UP, DON'T GUESS
When asked about a specific protein or gene target, look it up in UniProt/Ensembl/OpenTargets BEFORE reasoning about it. Verify the gene name, function, and disease associations from databases. When you're not sure about a fact, your first instinct should be to SEARCH for it using tools, not to reason harder from memory.
---
When to Use This Skill
Apply when users:
- Ask about a drug target, protein, or gene
- Need target validation or assessment
- Request druggability analysis
- Want comprehensive target profiling
- Ask "what do we know about [target]?"
- Need target-disease associations
- Request safety profile for a target
When NOT to use: Simple protein lookup, drug-only queries, disease-centric queries, sequence retrieval, structure download — use specialized skills instead.
---
Target Evaluation Reasoning Framework
Evaluating a drug target requires reasoning across four interconnected questions. Answer all four before forming a recommendation.
1. Is there genetic evidence linking this target to the disease? Genetic evidence is the strongest predictor of drug success — targets with human genetic support have approximately twice the clinical success rate as those without (Nelson et al. 2015). Ask: Are there GWAS associations connecting this gene to the disease? Do rare loss-of-function or gain-of-function variants cause or protect against the disease? Does the mouse knockout phenotype match the human disease (from OpenTargets mouse models)? OpenTargets assigns genetic evidence scores; a score > 0.7 indicates strong support. ClinVar rare variant evidence and DisGeNET curated gene-disease association scores add complementary layers. A target with no genetic link to the disease of interest carries a fundamental validation risk that cannot be resolved by downstream data.
2. Is the target druggable? Druggability has two components: structural accessibility and prior chemical matter. Structural accessibility means the target has a binding pocket where a small molecule or biologic can engage — surface-exposed receptors, enzymes with well-defined active sites, and protein-protein interaction interfaces with hot spots are tractable. Intrinsically disordered proteins and transcription factors with flat, featureless binding surfaces are typically harder. Pharos TDL classification provides a tiered assessment: Tclin (approved drug), Tchem (known active compounds), Tbio (biological function known but no drugs), Tdark (poorly characterized). If ChEMBL or BindingDB have compounds with IC50 < 1μM, the target is chemically tractable. Chemical probes (from OpenTargets chemical probes endpoint) confirm a target can be modulated, which is distinct from drug-like compounds. For GPCRs, check GPCRdb for curated agonists and antagonists.
3. Is the target safe to modulate? Safety concerns arise from two sources. First, on-target effects: if the target is essential in normal tissues (mouse KO is lethal, or gnomAD pLI is high / LOEUF is low), full inhibition will produce toxicity — the question becomes whether a partial agonist or tissue-targeted delivery can provide a therapeutic window. Second, off-target effects: does the gene have family members that could be inadvertently hit? The OpenTargets safety profile aggregates known toxicity annotations, and DepMap essentiality scores tell you which cancer cell lines require this gene for survival (useful but not directly translatable to normal tissues). Expression specificity matters: a target expressed only in the disease-relevant tissue is far safer than one expressed ubiquitously in critical organs (heart, kidney, brain).
4. What is the competitive landscape? A target with approved drugs may already be validated but competitive; a target with clinical-stage programs from competitors establishes feasibility while creating IP barriers. An entirely novel target with no drug history requires more extensive internal validation. Assess: number of ChEMBL bioactivity records (chemical matter depth), approved drugs from OpenTargets drug associations, and literature activity trends (recent paper count and key research groups). A dark target (Tdark) with strong genetic evidence but no chemical matter is a high-risk, high-reward opportunity.
Synthesizing the four dimensions: The ideal target has strong genetic evidence (GWAS + rare variant), a tractable binding site (Tclin or Tchem), acceptable safety profile (tissue-specific expression, non-lethal KO), and manageable competition. Gaps in any dimension represent validation tasks, not disqualifiers — but they must be acknowledged. A target with perfect druggability but no genetic link to disease is a tractability exercise, not a validated therapeutic hypothesis.
---
Phase 0: Tool Parameter Verification (CRITICAL)
BEFORE calling ANY tool for the first time, verify its parameters:
tool_info = tu.tools.get_tool_info(tool_name="Reactome_map_uniprot_to_pathways")
# Reveals: takes `id` not `uniprot_id`Known parameter corrections:
Reactome_map_uniprot_to_pathways: param isid(notuniprot_id)ensembl_get_xrefs: param isid(notgene_id)GTEx_get_median_gene_expression: requiresgencode_id+operation="median"; try versioned Ensembl ID if emptyOpenTargets_*: param isensemblId(camelCase, notensemblID)STRING_get_protein_interactions: takesprotein_ids(list) +speciesintact_get_interactions: takesidentifier(UniProt accession, not gene symbol)
---
Critical Workflow Requirements
Report-First (MANDATORY): Create [TARGET]_target_report.md with all section headers and [Researching...] placeholders before starting research. Update progressively. Do not show raw tool outputs to the user.
Evidence Grading (MANDATORY): Grade every claim T1-T4. T1 = clinical/genetic data; T2 = curated databases or multiple studies; T3 = computational or single study; T4 = annotation or catalog entry.
---
Core Strategy: 9 Research Paths
Target Query (e.g., "EGFR" or "P00533")
|
+- IDENTIFIER RESOLUTION (always first)
| +- Check if GPCR -> GPCRdb_get_protein
|
+- PATH 0: Open Targets Foundation (ALWAYS FIRST - fills gaps in all other paths)
|
+- PATH 1: Core Identity (names, IDs, sequence, organism)
| +- InterProScan_scan_sequence for novel domain prediction
+- PATH 2: Structure & Domains (3D structure, domains, binding sites)
| +- If GPCR: GPCRdb_get_structures (active/inactive states)
+- PATH 3: Function & Pathways (GO terms, pathways, biological role)
+- PATH 4: Protein Interactions (PPI network, complexes)
+- PATH 5: Expression Profile (tissue expression, single-cell)
+- PATH 6: Variants & Disease (mutations, clinical significance)
| +- DisGeNET_search_gene for curated gene-disease associations
+- PATH 7: Drug Interactions (known drugs, druggability, safety)
| +- Pharos_get_target for TDL classification (Tclin/Tchem/Tbio/Tdark)
| +- BindingDB_get_ligands_by_uniprot for known ligands
| +- PubChem_search_assays_by_target_gene for HTS data
| +- If GPCR: GPCRdb_get_ligands (curated agonists/antagonists)
| +- DepMap_get_gene_dependencies for target essentiality
+- PATH 8: Literature & Research (publications, trends)For detailed code implementations of each path, see IMPLEMENTATION.md.
---
Identifier Resolution (Phase 1)
Resolve ALL identifiers before any research path. Required IDs:
- UniProt accession (for protein data, structure, interactions)
- Ensembl gene ID + versioned ID (for Open Targets, GTEx)
- Gene symbol (for DGIdb, gnomAD, literature)
- Entrez gene ID (for KEGG, MyGene)
- ChEMBL target ID (for bioactivity)
- Synonyms/full name (for collision-aware literature search)
After resolution, check if target is a GPCR via GPCRdb_get_protein. See IMPLEMENTATION.md for resolution and GPCR detection code.
---
PATH 0: Open Targets Foundation (ALWAYS FIRST)
Run OpenTargets endpoints first to populate baseline data before specialized queries:
OpenTargets_get_diseases_phenotypes_by_target_ensembl→ disease associations (Section 8)OpenTargets_get_target_tractability_by_ensemblID→ druggability assessment (Section 9)OpenTargets_get_target_safety_profile_by_ensemblID→ safety liabilities (Section 10)OpenTargets_get_target_interactions_by_ensemblID→ PPI network (Section 6)OpenTargets_get_target_gene_ontology_by_ensemblID→ GO annotations (Section 5)OpenTargets_get_publications_by_target_ensemblID→ literature (Section 11)OpenTargets_get_biological_mouse_models_by_ensemblID→ mouse KO phenotypes (Sections 8/10)OpenTargets_get_chemical_probes_by_target_ensemblID→ chemical probes (Section 9)OpenTargets_get_associated_drugs_by_target_ensemblID→ known drugs (Section 9)
---
PATH 1: Core Identity
Tools: UniProt_get_entry_by_accession, UniProt_get_function_by_accession, UniProt_get_recommended_name_by_accession, UniProt_get_alternative_names_by_accession, UniProt_get_subcellular_location_by_accession, MyGene_get_gene_annotation
Populates: Sections 2 (Identifiers), 3 (Basic Information)
---
PATH 2: Structure & Domains
Use 3-step structure search chain (do NOT rely solely on PDB text search): 1. UniProt PDB cross-references (most reliable) 2. Sequence-based PDB search (catches missing annotations) 3. Domain-based search (for multi-domain proteins) 4. AlphaFold (always check)
Tools: UniProt_get_entry_by_accession (PDB xrefs), RCSBData_get_entry, PDB_search_similar_structures, alphafold_get_prediction, InterPro_get_protein_domains, UniProt_get_ptm_processing_by_accession
GPCR targets: Also query GPCRdb_get_structures for active/inactive state data.
Populates: Section 4 (Structural Biology)
---
PATH 3: Function & Pathways
Tools: GO_get_annotations_for_gene, Reactome_map_uniprot_to_pathways, kegg_get_gene_info, WikiPathways_search, enrichr_gene_enrichment_analysis
Populates: Section 5 (Function & Pathways)
---
PATH 4: Protein Interactions
Tools: STRING_get_protein_interactions, intact_get_interactions, intact_get_complex_details, BioGRID_get_interactions, HPA_get_protein_interactions_by_gene
Minimum: 20 interactors OR documented explanation.
Populates: Section 6 (Protein-Protein Interactions)
---
PATH 5: Expression Profile
GTEx with versioned ID fallback + HPA as backup.
Tools: GTEx_get_median_gene_expression, HPA_get_rna_expression_by_source, HPA_get_comprehensive_gene_details_by_ensembl_id, HPA_get_subcellular_location, HPA_get_cancer_prognostics_by_gene, HPA_get_comparative_expression_by_gene_and_cellline, CELLxGENE_get_expression_data
Reasoning: Expression specificity directly informs safety. Note whether expression is enriched in the disease-relevant tissue vs. critical organs. Ubiquitous essential expression narrows the therapeutic window.
Populates: Section 7 (Expression Profile)
---
PATH 6: Variants & Disease
Separate SNVs from CNVs in ClinVar results. Integrate DisGeNET for curated gene-disease association scores.
Tools: gnomad_get_gene_constraints, ClinVar_search_variants, OpenTargets_get_diseases_phenotypes_by_target_ensembl, DisGeNET_search_gene, civic_get_variants_by_gene, cBioPortal_get_mutations
Required constraint scores: pLI (probability of loss-of-function intolerance), LOEUF (loss-of-function observed/expected upper bound), missense Z-score, pRec (recessive probability). High pLI (> 0.9) or low LOEUF (< 0.35) indicates the gene is intolerant to loss-of-function — a major safety flag for inhibitory therapeutic strategies.
Populates: Section 8 (Genetic Variation & Disease)
---
PATH 7: Druggability & Target Validation
Tools: OpenTargets_get_target_tractability_by_ensemblID, DGIdb_get_gene_druggability, DGIdb_get_drug_gene_interactions, ChEMBL_search_targets, ChEMBL_get_target_activities, Pharos_get_target, BindingDB_get_ligands_by_uniprot, PubChem_search_assays_by_target_gene, DepMap_get_gene_dependencies, OpenTargets_get_target_safety_profile_by_ensemblID, OpenTargets_get_biological_mouse_models_by_ensemblID
GPCR targets: Also query GPCRdb_get_ligands.
Reasoning: Pharos TDL tells you where the target sits in the knowledge landscape. BindingDB Ki/IC50/Kd values tell you whether the target has been demonstrated tractable experimentally. DepMap essentiality tells you whether cancer cells require this gene (proxy for toxicity risk, not a definitive answer).
Populates: Sections 9 (Druggability), 10 (Safety), 12 (Competitive Landscape)
---
PATH 8: Literature & Research (Collision-Aware)
1. Detect collisions - Check if gene symbol has non-biological meanings 2. Build seed queries - Symbol in title with bio context, full name, UniProt accession 3. Apply collision filter - Add NOT terms for off-topic meanings 4. Expand via citations - For sparse targets (<30 papers), use citation network 5. Classify by evidence tier - T1-T4 based on title/abstract keywords
Tools: PubMed_search_articles, PubMed_get_related, EuropePMC_search_articles, EuropePMC_get_citations, PubTator3_LiteratureSearch, OpenTargets_get_publications_by_target_ensemblID
Populates: Section 11 (Literature & Research Landscape)
---
Retry Logic & Fallback Chains
ChEMBL_get_target_activitiesfails →GtoPdb_search_ligands→OpenTargets drugsintact_get_interactionsfails →STRING_get_protein_interactions→OpenTargets interactionsGO_get_annotations_for_genefails →OpenTargets GO→MyGene GOGTEx_get_median_gene_expressionfails →HPA_get_rna_expression_by_source→ document as unavailablegnomad_get_gene_constraintsfails →OpenTargets constraintendpointDGIdb_get_drug_gene_interactionsfails →OpenTargets drugs→GtoPdb_search_ligands
NEVER silently skip failed tools. Always document failures and fallbacks in the report.
---
Completeness Audit (REQUIRED before finalizing)
Before finalizing any report:
- Data minimums met for PPIs, expression, diseases, constraints, druggability
- Negative results documented explicitly
- T1-T4 grades in Executive Summary, Disease Associations, Key Papers, Recommendations
- Every data point has source attribution
---
Report Template
Create [TARGET]_target_report.md with all 15 sections initialized. See REPORT_FORMAT.md for the full template.
## 1. Executive Summary ## 9. Druggability & Pharmacology
## 2. Target Identifiers ## 10. Safety Profile
## 3. Basic Information ## 11. Literature & Research
## 4. Structural Biology ## 12. Competitive Landscape
## 5. Function & Pathways ## 13. Summary & Recommendations
## 6. Protein-Protein Interactions ## 14. Data Sources & Methodology
## 7. Expression Profile ## 15. Data Gaps & Limitations
## 8. Genetic Variation & Disease---
Synthesis: Target Assessment Framework
After completing all 9 PATHs, synthesize findings into a GO/NO-GO recommendation in the Executive Summary. Score each dimension:
- Genetic evidence: Strong (GWAS + rare variant + functional) / Moderate (GWAS or rare variant only) / Weak (expression change only) / None
- Disease association: Based on OpenTargets score (> 0.7 strong, 0.3-0.7 moderate, < 0.3 weak)
- Druggability: Approved drug exists / Tractable (known binding site, chemical probes) / Predicted tractable (structural pocket) / Undruggable
- Safety: Non-essential gene (viable KO, low pLI) / Essential with phenotype / Lethal KO or high pLI / Known toxicity target
- Selectivity: Disease-specific or enriched expression / Ubiquitous / Expressed in critical organs
- Structural data: High-res crystal with ligand / AlphaFold confident (pLDDT > 80) / Homology model / No structural info
Total score guides recommendation: strong target (all dimensions favorable), promising with defined validation tasks (2-3 gaps), speculative (multiple critical gaps), or deprioritize (no genetic link and poor druggability).
---
Reference Files
| File | Contents |
|---|---|
| IMPLEMENTATION.md | Detailed code for identifier resolution, GPCR detection, each PATH implementation, retry logic |
| EVIDENCE_GRADING.md | T1-T4 tier definitions, citation format, completeness audit checklist, data minimums |
| REPORT_FORMAT.md | Full report template with all 15 sections, table formats, section-specific guidance |
| REFERENCE.md | Complete tool reference (225+ tools) organized by category with parameters |
| EXAMPLES.md | Worked examples: EGFR full profile, KRAS druggability, target comparison, CDK4 validation, Alzheimer's targets |
Evidence Grading & Completeness Audit
Evidence grading system and completeness requirements for target intelligence reports.
Evidence Tiers
| Tier | Symbol | Criteria | Examples |
|---|---|---|---|
| T1 | three stars | Direct mechanistic evidence, human genetic proof | CRISPR KO, patient mutations, crystal structure with mechanism |
| T2 | two stars | Functional studies, model organism validation | siRNA phenotype, mouse KO, biochemical assay |
| T3 | one star | Association, screen hits, computational | GWAS hit, DepMap essentiality, expression correlation |
| T4 | no stars | Mention, review, text-mined, predicted | Review article, database annotation, computational prediction |
Required Evidence Grading Locations
Evidence grades MUST appear in: 1. Executive Summary - Key disease claims graded 2. Section 8.2 Disease Associations - Every disease link graded with source type 3. Section 11 Literature - Key papers table with evidence tier 4. Section 13 Recommendations - Scorecard items reference evidence quality
Per-Section Evidence Summary Format
---
**Evidence Quality for this Section**: Strong
- Mechanistic (T1): 12 papers
- Functional (T2): 8 papers
- Association (T3): 15 papers
- Mention (T4): 23 papers
**Data Gaps**: No CRISPR data; mouse KO phenotypes limited
---Citation Format
Every piece of information MUST include its source:
EGFR mutations cause lung adenocarcinoma [three stars: PMID:15118125, activating mutations
in patients]. *Source: ClinVar, CIViC*ClinVar SNV vs CNV Separation
Always separate single nucleotide variants from copy number variants:
### 8.3 Clinical Variants (ClinVar)
#### Single Nucleotide Variants (SNVs)
| Variant | Clinical Significance | Condition | Review Status | PMID |
|---------|----------------------|-----------|---------------|------|
| p.L858R | Pathogenic | Lung cancer | 4 stars | 15118125 |
**Total Pathogenic SNVs**: 47
#### Copy Number Variants (CNVs) - Reported Separately
| Type | Region | Clinical Significance | Frequency |
|------|--------|----------------------|-----------|
| Amplification | 7p11.2 | Pathogenic | Common in cancer |
*Note: CNV data separated as it represents different mutation mechanism*DisGeNET Evidence Tier Assignment
- DisGeNET Score >= 0.7 -> Consider T2 evidence (multiple validated sources)
- DisGeNET Score 0.4-0.7 -> Consider T3 evidence
- DisGeNET Score < 0.4 -> T4 evidence only
---
Minimum Data Requirements (Enforced)
| Section | Minimum Data | If Not Met |
|---|---|---|
| 6. PPIs | >= 20 interactors | Document which tools failed + why |
| 7. Expression | Top 10 tissues with TPM + HPA RNA summary | Note "limited data" with specific gaps |
| 8. Disease | Top 10 OT diseases + gnomAD constraints + ClinVar summary | Separate SNV/CNV; note if constraint unavailable |
| 9. Druggability | OT tractability + probes + drugs + DGIdb + GtoPdb fallback | "No drugs/probes" is valid data |
| 11. Literature | Total count + 5-year trend + 3-5 key papers with evidence tiers | Note if sparse (<50 papers) |
Post-Run Completeness Audit
Before finalizing the report, run this checklist:
Data Minimums Check
- [ ] PPIs: >= 20 interactors OR explanation why fewer
- [ ] Expression: Top 10 tissues with values OR explicit "unavailable"
- [ ] Diseases: Top 10 associations with scores OR "no associations"
- [ ] Constraints: All 4 scores (pLI, LOEUF, missense Z, pRec) OR "unavailable"
- [ ] Druggability: All modalities assessed; probes + drugs listed OR "none"
Negative Results Documented
- [ ] Empty tool results noted explicitly (not left blank)
- [ ] Failed tools with fallbacks documented
- [ ] "No data" sections have implications noted
Evidence Quality
- [ ] T1-T4 grades in Executive Summary disease claims
- [ ] T1-T4 grades in Disease Associations table
- [ ] Key papers table has evidence tiers
- [ ] Per-section evidence summaries included
Source Attribution
- [ ] Every data point has source tool/database cited
- [ ] Section-end source summaries present
Data Gap Table (Required if minimums not met)
## 15. Data Gaps & Limitations
| Section | Expected Data | Actual | Reason | Alternative Source |
|---------|---------------|--------|--------|-------------------|
| 6. PPIs | >= 20 interactors | 8 | Novel target, limited studies | Literature review needed |
| 7. Expression | GTEx TPM | None | Versioned ID not recognized | See HPA data |
| 9. Probes | Chemical probes | None | No validated probes exist | Consider tool compound dev |
**Recommendations for Data Gaps**:
1. For PPIs: Query BioGRID with broader parameters; check yeast-2-hybrid studies
2. For Expression: Query GEO directly for tissue-specific datasetsTarget Intelligence Examples
Detailed examples showing multi-step workflows for comprehensive target analysis.
Example 1: Complete EGFR Target Profile
Query: "Tell me everything about EGFR as a drug target"
Step 1: Resolve Identifiers
from tooluniverse import ToolUniverse
tu = ToolUniverse(use_cache=True)
tu.load_tools()
# Resolve EGFR to all IDs
# Search UniProt for human EGFR
search_result = tu.tools.UniProt_search(
query='gene:EGFR AND organism_id:9606',
limit=1
)
uniprot_id = search_result['results'][0]['primaryAccession'] # P00533
# Map to Ensembl
mapping = tu.tools.UniProt_id_mapping(
ids=['P00533'],
from_db='UniProtKB_AC-ID',
to_db='Ensembl'
)
ensembl_id = mapping['results'][0]['to'] # ENSG00000146648
ids = {
'symbol': 'EGFR',
'uniprot': 'P00533',
'ensembl': 'ENSG00000146648'
}Step 2: Core Identity (PATH 1)
# Get full UniProt entry
entry = tu.tools.UniProt_get_entry_by_accession(accession='P00533')
# Extract key info
identity = {
'name': entry.get('proteinDescription', {}).get('recommendedName', {}).get('fullName', {}).get('value'),
'length': entry.get('sequence', {}).get('length'),
'organism': entry.get('organism', {}).get('scientificName'),
'function': tu.tools.UniProt_get_function_by_accession(accession='P00533')
}
# Get gene annotation
gene_info = tu.tools.MyGene_get_gene_annotation(
gene_id='1956', # Entrez ID for EGFR
fields='symbol,name,summary,alias,genomic_pos'
)Output:
Name: Epidermal growth factor receptor
Symbol: EGFR
Length: 1210 amino acids
Organism: Homo sapiens
Function: Receptor tyrosine kinase binding ligands of the EGF family...Step 3: Structure & Domains (PATH 2)
# Get PDB structures from UniProt entry
pdb_refs = [xref for xref in entry.get('uniProtKBCrossReferences', [])
if xref.get('database') == 'PDB']
# Get details for best structure
best_pdb = '1M17' # Example: EGFR kinase domain
pdb_info = tu.tools.get_protein_metadata_by_pdb_id(pdb_id='1M17')
# Get AlphaFold prediction
alphafold = tu.tools.alphafold_get_prediction(qualifier='P00533')
# Get domain architecture
domains = tu.tools.InterPro_get_protein_domains(protein_id='P00533')
# Get PTMs and active sites
ptms = tu.tools.UniProt_get_ptm_processing_by_accession(accession='P00533')Output:
PDB Structures: 150+ entries
Best Resolution: 1.8Å (1M17)
AlphaFold: Available, high confidence
Key Domains:
- Receptor L domain (EGF binding)
- Furin-like domain
- Growth factor receptor domain
- Protein kinase domain
PTMs: Multiple phosphorylation sites (Y992, Y1068, Y1086...)Step 4: Function & Pathways (PATH 3)
# GO annotations
go_terms = tu.tools.GO_get_annotations_for_gene(gene_id='UniProtKB:P00533')
# Reactome pathways
reactome_pathways = tu.tools.Reactome_map_uniprot_to_pathways(id='P00533')
# KEGG pathways
kegg_info = tu.tools.kegg_get_gene_info(gene_id='hsa:1956')
# Open Targets GO
ot_go = tu.tools.OpenTargets_get_target_gene_ontology_by_ensemblID(
ensemblID='ENSG00000146648'
)Output:
GO Biological Process:
- Signal transduction (GO:0007165)
- Cell proliferation (GO:0008283)
- MAPK cascade (GO:0000165)
GO Molecular Function:
- Protein tyrosine kinase activity (GO:0004713)
- ATP binding (GO:0005524)
- Receptor binding (GO:0005102)
Key Pathways:
- EGFR signaling pathway (Reactome)
- PI3K-Akt signaling (KEGG hsa04151)
- MAPK signaling (KEGG hsa04010)
- ErbB signaling pathway (KEGG hsa04012)Step 5: Protein Interactions (PATH 4)
# STRING interactions
string_ppi = tu.tools.STRING_get_protein_interactions(
protein_ids=['EGFR'],
species=9606,
confidence_score=0.9,
limit=50
)
# IntAct experimental interactions
intact_ppi = tu.tools.intact_get_interactions(
identifier='P00533',
format='json'
)
# Open Targets interactions
ot_ppi = tu.tools.OpenTargets_get_target_interactions_by_ensemblID(
ensemblID='ENSG00000146648'
)Output:
Top STRING Interactors (score > 0.9):
1. GRB2 (0.999) - Adapter protein
2. SHC1 (0.999) - Signal transduction
3. ERBB2 (0.998) - Receptor family
4. SRC (0.996) - Kinase
5. STAT3 (0.994) - Transcription factor
IntAct Complexes:
- EGFR-GRB2-SOS1 complex
- EGFR-PI3K complexStep 6: Expression Profile (PATH 5)
# GTEx expression
gtex = tu.tools.GTEx_get_median_gene_expression(
gencode_id='ENSG00000146648.11'
)
# HPA expression
hpa = tu.tools.HPA_get_comprehensive_gene_details_by_ensembl_id(
ensembl_id='ENSG00000146648'
)
# Subcellular location
subcell = tu.tools.HPA_get_subcellular_location(
ensembl_id='ENSG00000146648'
)
# Cancer prognostics
cancer = tu.tools.HPA_get_cancer_prognostics_by_gene(
gene_symbol='EGFR'
)Output:
Top Expression Tissues (GTEx TPM):
1. Skin (150+ TPM)
2. Esophagus mucosa (120+ TPM)
3. Kidney cortex (100+ TPM)
4. Lung (80+ TPM)
Tissue Specificity: Low (broadly expressed)
Subcellular: Plasma membrane, Cytoplasm
Cancer Relevance:
- Overexpressed in: NSCLC, Glioblastoma, Colorectal
- Prognostic: Unfavorable in lung cancerStep 7: Variants & Disease (PATH 6)
# gnomAD constraint scores
constraints = tu.tools.gnomad_get_gene_constraints(gene_symbol='EGFR')
# UniProt disease variants
disease_vars = tu.tools.UniProt_get_disease_variants_by_accession(accession='P00533')
# ClinVar variants
clinvar = tu.tools.ClinVar_search_variants(gene='EGFR', max_results=100)
# Open Targets disease associations
diseases = tu.tools.OpenTargets_get_diseases_phenotypes_by_target_ensembl(
ensemblId='ENSG00000146648'
)Output:
Constraint Scores:
- pLI: 0.99 (highly loss-of-function intolerant)
- LOEUF: 0.18
- Missense Z: 3.5
Disease Associations (Open Targets):
1. Non-small cell lung carcinoma (0.95)
2. Glioblastoma multiforme (0.89)
3. Colorectal cancer (0.82)
4. Pancreatic cancer (0.75)
ClinVar Pathogenic Variants: 45
- L858R (common activating mutation)
- T790M (resistance mutation)
- Exon 19 deletionsStep 8: Drug Interactions (PATH 7)
# Open Targets tractability
tractability = tu.tools.OpenTargets_get_target_tractability_by_ensemblID(
ensemblID='ENSG00000146648'
)
# DGIdb druggability
druggability = tu.tools.DGIdb_get_gene_druggability(genes=['EGFR'])
# Known drugs
drugs = tu.tools.OpenTargets_get_associated_drugs_by_target_ensemblID(
ensemblID='ENSG00000146648'
)
# ChEMBL bioactivity
# First get ChEMBL target ID
chembl_target = tu.tools.ChEMBL_search_targets(
pref_name__contains='EGFR',
organism='Homo sapiens',
limit=1
)
target_chembl_id = chembl_target['targets'][0]['target_chembl_id'] # CHEMBL203
activities = tu.tools.ChEMBL_get_target_activities(
target_chembl_id__exact='CHEMBL203',
limit=100
)
# Safety profile
safety = tu.tools.OpenTargets_get_target_safety_profile_by_ensemblID(
ensemblID='ENSG00000146648'
)
# Chemical probes
probes = tu.tools.OpenTargets_get_chemical_probes_by_target_ensemblID(
ensemblID='ENSG00000146648'
)Output:
Tractability:
- Small Molecule: HIGH (clinical precedence)
- Antibody: HIGH (approved antibodies)
- PROTAC: HIGH (structural data available)
- Other Modalities: MEDIUM
Approved Drugs:
1. Erlotinib (Tarceva) - TKI
2. Gefitinib (Iressa) - TKI
3. Afatinib (Gilotrif) - TKI
4. Osimertinib (Tagrisso) - 3rd gen TKI
5. Cetuximab (Erbitux) - mAb
6. Panitumumab (Vectibix) - mAb
ChEMBL Activities:
- 50,000+ bioactivity records
- Best IC50: 0.5 nM (osimertinib)
Safety Liabilities:
- Skin toxicity (class effect)
- Diarrhea (common)
- Interstitial lung disease (rare)
Chemical Probes:
- Gefitinib (SGC probe)Step 9: Literature (PATH 8)
# PubMed publications
pubmed_total = tu.tools.PubMed_search_articles(
query='EGFR[Gene Name]',
limit=0 # Just count
)
pubmed_recent = tu.tools.PubMed_search_articles(
query='EGFR[Gene Name] AND "2024"[Date - Publication]',
limit=0
)
pubmed_drug = tu.tools.PubMed_search_articles(
query='EGFR AND drug AND cancer',
limit=10
)
# Open Targets publications
ot_pubs = tu.tools.OpenTargets_get_publications_by_target_ensemblID(
ensemblID='ENSG00000146648',
size=10
)Output:
Literature Summary:
- Total publications: 180,000+
- Recent (2024): 8,000+
- Drug-related: 45,000+
- Trend: Stable (mature, well-studied target)
Recent Focus Areas:
- Resistance mechanisms
- Third-generation inhibitors
- Combination therapies
- Liquid biopsy for monitoringFinal Synthesized Report
# Target Intelligence Report: EGFR
## Quick Facts
| Property | Value |
|----------|-------|
| Symbol | EGFR |
| UniProt | P00533 |
| Ensembl | ENSG00000146648 |
| Name | Epidermal growth factor receptor |
| Length | 1210 amino acids |
| Organism | Homo sapiens |
## Druggability Assessment
| Modality | Score | Evidence |
|----------|-------|----------|
| Small molecule | ✅ HIGH | 4+ approved TKIs |
| Antibody | ✅ HIGH | 2+ approved mAbs |
| PROTAC | ✅ HIGH | Structural data |
## Summary by Domain
### Identity & Function
- Receptor tyrosine kinase of the ErbB family
- Regulates cell proliferation, survival, differentiation
- Activates RAS-MAPK and PI3K-AKT pathways
### Structure
- 150+ PDB structures available
- Key domains: Kinase (for TKIs), Extracellular (for mAbs)
- AlphaFold model: High confidence
### Expression
- Broadly expressed (skin, lung, kidney highest)
- Overexpressed in multiple cancers
- Subcellular: Plasma membrane
### Variants & Disease
- Highly constrained (pLI=0.99)
- Major cancer associations: NSCLC, glioblastoma, CRC
- Key mutations: L858R, T790M, exon 19 del
### Drugs
- 6+ approved drugs
- 200+ clinical trials
- Well-characterized safety profile (skin toxicity)
### Research
- Mature target (180K+ publications)
- Active research on resistance mechanisms
## Recommendations
1. [HIGH] Excellent target validation - multiple approved therapies
2. [MEDIUM] Consider resistance mutations in drug design
3. [INFO] Extensive structural data for SBDD available---
Example 2: Novel Target Assessment (KRAS G12C)
Query: "Is KRAS druggable? What's the current state?"
Multi-Step Workflow
from tooluniverse import ToolUniverse
from concurrent.futures import ThreadPoolExecutor
tu = ToolUniverse(use_cache=True)
tu.load_tools()
# Resolve KRAS
ids = {
'symbol': 'KRAS',
'uniprot': 'P01116',
'ensembl': 'ENSG00000133703'
}
# Parallel execution
def assess_druggability():
results = {}
# 1. Tractability
results['tractability'] = tu.tools.OpenTargets_get_target_tractability_by_ensemblID(
ensemblID=ids['ensembl']
)
# 2. DGIdb assessment
results['dgidb'] = tu.tools.DGIdb_get_gene_druggability(genes=['KRAS'])
# 3. Known drugs
results['drugs'] = tu.tools.OpenTargets_get_associated_drugs_by_target_ensemblID(
ensemblID=ids['ensembl']
)
# 4. ChEMBL activities
chembl = tu.tools.ChEMBL_search_targets(
pref_name__contains='KRAS',
organism='Homo sapiens',
limit=1
)
if chembl.get('targets'):
target_id = chembl['targets'][0]['target_chembl_id']
results['activities'] = tu.tools.ChEMBL_get_target_activities(
target_chembl_id__exact=target_id,
limit=50
)
# 5. Structures
results['structures'] = tu.tools.alphafold_get_prediction(qualifier='P01116')
# 6. Safety
results['safety'] = tu.tools.OpenTargets_get_target_safety_profile_by_ensemblID(
ensemblID=ids['ensembl']
)
return results
druggability = assess_druggability()Output:
KRAS Druggability Assessment:
Tractability:
- Small Molecule: MEDIUM → HIGH (recent breakthroughs!)
- Previously "undruggable" - now validated
Approved Drugs (G12C-specific):
- Sotorasib (Lumakras) - 2021
- Adagrasib (Krazati) - 2022
ChEMBL Activities:
- 5000+ bioactivity records
- G12C-specific covalent inhibitors
Structure:
- Multiple G12C-bound structures
- Switch II pocket (key for covalent inhibitors)
Recent Breakthrough:
- Covalent inhibitors targeting G12C mutant
- Switch II pocket discovered as druggable site
- Active clinical development for other mutations (G12D, G12V)---
Example 3: Target Comparison
Query: "Compare EGFR vs HER2 as drug targets"
Parallel Analysis
targets = [
{'symbol': 'EGFR', 'uniprot': 'P00533', 'ensembl': 'ENSG00000146648'},
{'symbol': 'ERBB2', 'uniprot': 'P04626', 'ensembl': 'ENSG00000141736'} # HER2
]
def analyze_target(target):
result = {
'symbol': target['symbol'],
'tractability': tu.tools.OpenTargets_get_target_tractability_by_ensemblID(
ensemblID=target['ensembl']
),
'drugs': tu.tools.OpenTargets_get_associated_drugs_by_target_ensemblID(
ensemblID=target['ensembl']
),
'diseases': tu.tools.OpenTargets_get_diseases_phenotypes_by_target_ensembl(
ensemblId=target['ensembl']
),
'safety': tu.tools.OpenTargets_get_target_safety_profile_by_ensemblID(
ensemblID=target['ensembl']
),
'ppi': tu.tools.STRING_get_protein_interactions(
protein_ids=[target['symbol']],
species=9606,
confidence_score=0.9,
limit=20
)
}
return result
# Parallel analysis
with ThreadPoolExecutor(max_workers=2) as executor:
results = list(executor.map(analyze_target, targets))Comparison Output:
| Property | EGFR | HER2/ERBB2 |
|----------|------|------------|
| **Approved Drugs** | 6+ | 8+ |
| **TKI Tractability** | HIGH | HIGH |
| **mAb Tractability** | HIGH | HIGH |
| **ADC** | N/A | HIGH (T-DM1, T-DXd) |
| **Primary Indication** | NSCLC | Breast cancer |
| **Key Mutations** | L858R, T790M | Amplification |
| **Safety Concerns** | Skin toxicity | Cardiotoxicity |---
Example 4: Target Validation Pipeline
Query: "Validate CDK4 as a potential drug target for cancer"
Systematic Validation
def validate_target(gene_symbol, disease_area='cancer'):
"""
Systematic target validation following industry best practices.
"""
tu = ToolUniverse(use_cache=True)
tu.load_tools()
validation_results = {}
# 1. GENETIC EVIDENCE
# Disease associations
ensembl_id = 'ENSG00000135446' # CDK4
disease_assoc = tu.tools.OpenTargets_get_diseases_phenotypes_by_target_ensembl(
ensemblId=ensembl_id
)
validation_results['genetic_evidence'] = {
'disease_associations': disease_assoc,
'score': 'HIGH' if any(d.get('score', 0) > 0.5 for d in disease_assoc.get('data', [])) else 'LOW'
}
# Constraint score
constraint = tu.tools.gnomad_get_gene_constraints(gene_symbol='CDK4')
validation_results['constraint'] = constraint
# 2. EXPRESSION EVIDENCE
# Cancer expression
hpa_cancer = tu.tools.HPA_get_cancer_prognostics_by_gene(gene_symbol='CDK4')
validation_results['expression'] = hpa_cancer
# 3. FUNCTIONAL EVIDENCE
# Pathways
pathways = tu.tools.Reactome_map_uniprot_to_pathways(id='P11802')
go_terms = tu.tools.GO_get_annotations_for_gene(gene_id='UniProtKB:P11802')
validation_results['function'] = {
'pathways': pathways,
'go_terms': go_terms
}
# 4. DRUGGABILITY
tractability = tu.tools.OpenTargets_get_target_tractability_by_ensemblID(
ensemblID=ensembl_id
)
existing_drugs = tu.tools.OpenTargets_get_associated_drugs_by_target_ensemblID(
ensemblID=ensembl_id
)
validation_results['druggability'] = {
'tractability': tractability,
'existing_drugs': existing_drugs
}
# 5. SAFETY
safety = tu.tools.OpenTargets_get_target_safety_profile_by_ensemblID(
ensemblID=ensembl_id
)
mouse_models = tu.tools.OpenTargets_get_biological_mouse_models_by_ensemblID(
ensemblID=ensembl_id
)
validation_results['safety'] = {
'profile': safety,
'mouse_models': mouse_models
}
# 6. COMPETITIVE LANDSCAPE
lit_count = tu.tools.PubMed_search_articles(
query=f'{gene_symbol} AND drug AND cancer',
limit=0
)
validation_results['competitive'] = {
'publication_count': lit_count.get('count', 0),
'maturity': 'mature' if lit_count.get('count', 0) > 1000 else 'emerging'
}
return validation_results
results = validate_target('CDK4')Validation Report:
# Target Validation Report: CDK4
## Validation Scorecard
| Criterion | Score | Evidence |
|-----------|-------|----------|
| Genetic Evidence | ✅ HIGH | Strong cancer associations |
| Expression | ✅ HIGH | Overexpressed in multiple cancers |
| Functional Role | ✅ HIGH | Cell cycle regulator |
| Druggability | ✅ HIGH | Approved inhibitors |
| Safety | ⚠️ MEDIUM | On-target effects expected |
| Competitive | 🔴 HIGH | Mature, crowded space |
## Existing Drugs
- Palbociclib (Ibrance)
- Ribociclib (Kisqali)
- Abemaciclib (Verzenio)
## Recommendation
CDK4 is a **validated target** with approved drugs.
New entrants would need differentiation (selectivity, CNS penetration, etc.)---
Example 5: Finding Drug Targets for a Disease
Query: "What are the best drug targets for Alzheimer's disease?"
Disease-to-Target Discovery
def find_targets_for_disease(disease_name):
tu = ToolUniverse(use_cache=True)
tu.load_tools()
# 1. Get disease ID
disease_search = tu.tools.OpenTargets_get_disease_ids_by_name(
diseaseName=disease_name
)
efo_id = disease_search.get('id') # e.g., EFO_0000249 for Alzheimer's
# 2. Get associated targets
targets = tu.tools.OpenTargets_get_associated_targets_by_disease_efoId(
efoId=efo_id
)
# 3. For top targets, assess druggability
top_targets = targets.get('data', [])[:10]
target_assessments = []
for target in top_targets:
ensembl_id = target.get('target_id')
symbol = target.get('gene_symbol')
# Druggability
tract = tu.tools.OpenTargets_get_target_tractability_by_ensemblID(
ensemblID=ensembl_id
)
# Existing drugs
drugs = tu.tools.OpenTargets_get_associated_drugs_by_target_ensemblID(
ensemblID=ensembl_id
)
# Safety
safety = tu.tools.OpenTargets_get_target_safety_profile_by_ensemblID(
ensemblID=ensembl_id
)
target_assessments.append({
'symbol': symbol,
'ensembl_id': ensembl_id,
'disease_score': target.get('score'),
'tractability': tract,
'drug_count': len(drugs.get('data', [])),
'safety_flags': len(safety.get('data', []))
})
return target_assessments
alzheimer_targets = find_targets_for_disease("Alzheimer's disease")Output:
# Top Drug Targets for Alzheimer's Disease
| Rank | Target | Score | Druggability | Drugs | Safety |
|------|--------|-------|--------------|-------|--------|
| 1 | APP | 0.95 | MEDIUM | 0 | ⚠️ |
| 2 | PSEN1 | 0.92 | LOW | 0 | ⚠️ |
| 3 | APOE | 0.88 | LOW | 0 | ✅ |
| 4 | MAPT | 0.85 | MEDIUM | 2 | ⚠️ |
| 5 | BACE1 | 0.82 | HIGH | 0* | ⚠️ |
| 6 | GSK3B | 0.75 | HIGH | 5 | ⚠️ |
| 7 | ACE | 0.70 | HIGH | 10+ | ✅ |
| 8 | TREM2 | 0.68 | MEDIUM | 2 | ✅ |
*BACE1 inhibitors failed in trials
## Recommendations
1. **TREM2** - Emerging target, favorable safety, antibody approaches
2. **GSK3B** - Druggable, existing tool compounds
3. **ACE** - Repurposing opportunity (existing drugs)---
Quick Reference: Common Multi-Step Patterns
Pattern: Gene Symbol → Full Profile
# 1. Symbol → UniProt
search = tu.tools.UniProt_search(query=f'gene:{symbol} AND organism_id:9606', limit=1)
uniprot = search['results'][0]['primaryAccession']
# 2. UniProt → Ensembl
mapping = tu.tools.UniProt_id_mapping(ids=[uniprot], from_db='UniProtKB_AC-ID', to_db='Ensembl')
ensembl = mapping['results'][0]['to']
# 3. Get all info
entry = tu.tools.UniProt_get_entry_by_accession(accession=uniprot)
tractability = tu.tools.OpenTargets_get_target_tractability_by_ensemblID(ensemblID=ensembl)Pattern: PDB → Ligand Analysis
# 1. Get PDB info
pdb_info = tu.tools.get_protein_metadata_by_pdb_id(pdb_id='1M17')
# 2. Get ligands
ligands = pdb_info.get('rcsb_binding_affinity', [])
# 3. For each ligand, get ChEMBL data
for lig in ligands:
comp_id = lig.get('comp_id')
smiles = tu.tools.get_ligand_smiles_by_chem_comp_id(chem_comp_id=comp_id)
# Search ChEMBL for similar moleculesPattern: Disease → Target → Drug
# 1. Disease → EFO ID
efo = tu.tools.OpenTargets_get_disease_ids_by_name(diseaseName='lung cancer')
# 2. EFO → Targets
targets = tu.tools.OpenTargets_get_associated_targets_by_disease_efoId(efoId=efo['id'])
# 3. Target → Drugs
for target in targets['data'][:5]:
drugs = tu.tools.OpenTargets_get_associated_drugs_by_target_ensemblID(
ensemblID=target['target_id']
)Target Intelligence Implementation Details
Detailed code implementations for each research path. See SKILL.md for the workflow overview.
Identifier Resolution
CRITICAL: Resolve ALL identifiers before any research path.
def resolve_target_ids(tu, query):
"""
Resolve target query to ALL needed identifiers.
Returns dict with: query, uniprot, ensembl, ensembl_version, symbol,
entrez, chembl_target, hgnc
"""
ids = {
'query': query,
'uniprot': None,
'ensembl': None,
'ensembl_versioned': None, # For GTEx
'symbol': None,
'entrez': None,
'chembl_target': None,
'hgnc': None,
'full_name': None,
'synonyms': []
}
# [Resolution logic based on input type]
# ... (see current implementation)
# CRITICAL: Get versioned Ensembl ID for GTEx
if ids['ensembl']:
gene_info = tu.tools.ensembl_lookup_gene(id=ids['ensembl'], species="human")
if gene_info and gene_info.get('version'):
ids['ensembl_versioned'] = f"{ids['ensembl']}.{gene_info['version']}"
# Also get synonyms for literature collision detection
ids['full_name'] = gene_info.get('description', '').split(' [')[0]
# Get UniProt alternative names for synonyms
if ids['uniprot']:
alt_names = tu.tools.UniProt_get_alternative_names_by_accession(accession=ids['uniprot'])
if alt_names:
ids['synonyms'].extend(alt_names)
return idsGPCR Target Detection
~35% of approved drugs target GPCRs. After identifier resolution, check if target is a GPCR:
def check_gpcr_target(tu, ids):
"""
Check if target is a GPCR and retrieve specialized data.
Call after identifier resolution.
"""
symbol = ids.get('symbol', '')
# Build GPCRdb entry name
entry_name = f"{symbol.lower()}_human"
gpcr_info = tu.tools.GPCRdb_get_protein(
operation="get_protein",
protein=entry_name
)
if gpcr_info.get('status') == 'success':
# Target is a GPCR - get specialized data
structures = tu.tools.GPCRdb_get_structures(
operation="get_structures",
protein=entry_name
)
ligands = tu.tools.GPCRdb_get_ligands(
operation="get_ligands",
protein=entry_name
)
mutations = tu.tools.GPCRdb_get_mutations(
operation="get_mutations",
protein=entry_name
)
return {
'is_gpcr': True,
'gpcr_family': gpcr_info['data'].get('family'),
'gpcr_class': gpcr_info['data'].get('receptor_class'),
'structures': structures.get('data', {}).get('structures', []),
'ligands': ligands.get('data', {}).get('ligands', []),
'mutations': mutations.get('data', {}).get('mutations', []),
'ballesteros_numbering': True
}
return {'is_gpcr': False}GPCRdb Report Section (add to Section 2 for GPCR targets):
### 2.x GPCR-Specific Data (GPCRdb)
**Receptor Class**: Class A (Rhodopsin-like)
**GPCR Family**: Adrenoceptors
**Structures by State**:
| PDB ID | State | Resolution | Ligand | Year |
|--------|-------|------------|--------|------|
| 3SN6 | Active | 3.2A | Agonist (BI-167107) | 2011 |
| 2RH1 | Inactive | 2.4A | Antagonist (carazolol) | 2007 |
**Known Ligands**: 45 agonists, 32 antagonists, 8 allosteric modulators
**Key Binding Site Residues** (Ballesteros-Weinstein): 3.32, 5.42, 6.48, 7.39Collision Detection for Literature Search
Before literature search, detect naming collisions:
def detect_collisions(tu, symbol, full_name):
"""
Detect if gene symbol has naming collisions in literature.
Returns negative filter terms if collisions found.
"""
results = tu.tools.PubMed_search_articles(
query=f'"{symbol}"[Title]',
limit=20
)
off_topic_terms = []
for paper in results.get('articles', []):
title = paper.get('title', '').lower()
bio_terms = ['protein', 'gene', 'cell', 'expression', 'mutation', 'kinase', 'receptor']
if not any(term in title for term in bio_terms):
pass # Extract potential collision terms
collision_filter = ""
if off_topic_terms:
collision_filter = " NOT " + " NOT ".join(off_topic_terms)
return collision_filterPATH 0: Open Targets Foundation
def path_0_open_targets(tu, ids):
"""
Open Targets foundation data - fills gaps for sections 5, 6, 8, 9, 10, 11.
ALWAYS run this first.
"""
ensembl_id = ids['ensembl']
if not ensembl_id:
return {'status': 'skipped', 'reason': 'No Ensembl ID'}
results = {}
# 1. Diseases & Phenotypes (Section 8)
diseases = tu.tools.OpenTargets_get_diseases_phenotypes_by_target_ensemblId(
ensemblId=ensembl_id
)
results['diseases'] = diseases if diseases else {'note': 'No disease associations returned'}
# 2. Tractability (Section 9)
tractability = tu.tools.OpenTargets_get_target_tractability_by_ensemblID(
ensemblId=ensembl_id
)
results['tractability'] = tractability if tractability else {'note': 'No tractability data returned'}
# 3. Safety Profile (Section 10)
safety = tu.tools.OpenTargets_get_target_safety_profile_by_ensemblID(
ensemblId=ensembl_id
)
results['safety'] = safety if safety else {'note': 'No safety liabilities identified'}
# 4. Interactions (Section 6)
interactions = tu.tools.OpenTargets_get_target_interactions_by_ensemblID(
ensemblId=ensembl_id
)
results['interactions'] = interactions if interactions else {'note': 'No interactions returned'}
# 5. GO Annotations (Section 5)
go_terms = tu.tools.OpenTargets_get_target_gene_ontology_by_ensemblID(
ensemblId=ensembl_id
)
results['go_terms'] = go_terms if go_terms else {'note': 'No GO annotations returned'}
# 6. Publications (Section 11)
publications = tu.tools.OpenTargets_get_publications_by_target_ensemblID(
ensemblId=ensembl_id
)
results['publications'] = publications if publications else {'note': 'No publications returned'}
# 7. Mouse Models (Section 8/10)
mouse_models = tu.tools.OpenTargets_get_biological_mouse_models_by_ensemblID(
ensemblId=ensembl_id
)
results['mouse_models'] = mouse_models if mouse_models else {'note': 'No mouse model data returned'}
# 8. Chemical Probes (Section 9)
probes = tu.tools.OpenTargets_get_chemical_probes_by_target_ensemblID(
ensemblId=ensembl_id
)
results['chemical_probes'] = probes if probes else {'note': 'No chemical probes available'}
# 9. Associated Drugs (Section 9)
drugs = tu.tools.OpenTargets_get_associated_drugs_by_target_ensemblID(
ensemblId=ensembl_id
)
results['drugs'] = drugs if drugs else {'note': 'No approved/trial drugs found'}
return resultsNegative Results Are Data
Always document when a query returns empty:
### 9.3 Chemical Probes
**Status**: No validated chemical probes available for this target.
*Source: OpenTargets_get_chemical_probes_by_target_ensemblID returned empty*
**Implication**: Tool compound development would be needed for chemical biology studies.PATH 2: Structure & Domains (3-Step Chain)
Do NOT rely solely on PDB text search. Use this chain:
def path_structure_robust(tu, ids):
"""Robust structure search using 3-step chain."""
structures = {'pdb': [], 'alphafold': None, 'domains': [], 'method_notes': []}
# STEP 1: UniProt PDB Cross-References (most reliable)
if ids['uniprot']:
entry = tu.tools.UniProt_get_entry_by_accession(accession=ids['uniprot'])
pdb_xrefs = [x for x in entry.get('uniProtKBCrossReferences', [])
if x.get('database') == 'PDB']
for xref in pdb_xrefs:
pdb_id = xref.get('id')
pdb_info = tu.tools.get_protein_metadata_by_pdb_id(pdb_id=pdb_id)
if pdb_info:
structures['pdb'].append(pdb_info)
structures['method_notes'].append(f"Step 1: {len(pdb_xrefs)} PDB cross-refs from UniProt")
# STEP 2: Sequence-based PDB Search (catches missing annotations)
if ids['uniprot'] and len(structures['pdb']) < 5:
sequence = tu.tools.UniProt_get_sequence_by_accession(accession=ids['uniprot'])
if sequence and len(sequence) < 1000:
similar = tu.tools.PDB_search_similar_structures(
sequence=sequence[:500],
identity_cutoff=0.7
)
if similar:
for hit in similar[:10]:
if hit['pdb_id'] not in [s.get('pdb_id') for s in structures['pdb']]:
structures['pdb'].append(hit)
structures['method_notes'].append(f"Step 2: Sequence search (identity >= 70%)")
# STEP 3: Domain-based Search (for multi-domain proteins)
if ids['uniprot']:
domains = tu.tools.InterPro_get_protein_domains(uniprot_accession=ids['uniprot'])
structures['domains'] = domains if domains else []
# AlphaFold (always check)
alphafold = tu.tools.alphafold_get_prediction(uniprot_accession=ids['uniprot'])
structures['alphafold'] = alphafold if alphafold else {'note': 'No AlphaFold prediction'}
# Document limitations
if not structures['pdb']:
structures['limitation'] = "No direct PDB hit does NOT mean no structure exists. Check: (1) structures under different UniProt entries, (2) homolog structures, (3) domain-only structures."
return structuresPATH 5: Expression Profile (GTEx Versioned ID Fallback)
def path_expression(tu, ids):
"""Expression data with GTEx versioned ID fallback."""
results = {'gtex': None, 'hpa': None, 'failed_tools': []}
ensembl_id = ids['ensembl']
versioned_id = ids.get('ensembl_versioned')
# Try unversioned first
gtex_result = tu.tools.GTEx_get_median_gene_expression(
gencode_id=ensembl_id,
operation="median"
)
# Fallback to versioned if empty
if not gtex_result or gtex_result.get('data') == []:
if versioned_id:
gtex_result = tu.tools.GTEx_get_median_gene_expression(
gencode_id=versioned_id,
operation="median"
)
if gtex_result and gtex_result.get('data'):
results['gtex'] = gtex_result
results['gtex_note'] = f"Used versioned ID: {versioned_id}"
if not results.get('gtex'):
results['failed_tools'].append({
'tool': 'GTEx_get_median_gene_expression',
'tried': [ensembl_id, versioned_id],
'fallback': 'See HPA data below'
})
else:
results['gtex'] = gtex_result
# HPA (always query as backup)
hpa_result = tu.tools.HPA_get_rna_expression_by_source(ensembl_id=ensembl_id)
results['hpa'] = hpa_result if hpa_result else {'note': 'No HPA RNA data'}
return resultsHPA Extended Expression
def get_hpa_comprehensive_expression(tu, gene_symbol):
"""
Get comprehensive expression data from Human Protein Atlas.
Provides tissue expression, subcellular localization, cell line comparison, tissue specificity.
"""
gene_info = tu.tools.HPA_search_genes_by_query(search_query=gene_symbol)
if not gene_info:
return {'error': f'Gene {gene_symbol} not found in HPA'}
tissue_search = tu.tools.HPA_generic_search(
search_query=gene_symbol,
columns="g,gs,rnat,rnatsm,scml,scal",
format="json"
)
cell_lines = ['a549', 'mcf7', 'hela', 'hepg2', 'pc3']
cell_line_expression = {}
for cell_line in cell_lines:
try:
expr = tu.tools.HPA_get_comparative_expression_by_gene_and_cellline(
gene_name=gene_symbol,
cell_line=cell_line
)
cell_line_expression[cell_line] = expr
except:
continue
return {
'gene_info': gene_info,
'tissue_data': tissue_search,
'cell_line_expression': cell_line_expression,
'source': 'Human Protein Atlas'
}PATH 6: Variants & Disease
DisGeNET Integration
DisGeNET provides curated gene-disease associations with evidence scores. Requires: DISGENET_API_KEY
def get_disgenet_associations(tu, ids):
"""Get gene-disease associations from DisGeNET."""
symbol = ids.get('symbol')
if not symbol:
return {'status': 'skipped', 'reason': 'No gene symbol'}
gda = tu.tools.DisGeNET_search_gene(
operation="search_gene",
gene=symbol,
limit=50
)
if gda.get('status') != 'success':
return {'status': 'error', 'message': 'DisGeNET query failed'}
associations = gda.get('data', {}).get('associations', [])
strong, moderate, weak = [], [], []
for assoc in associations:
score = assoc.get('score', 0)
entry = {
'disease': assoc.get('disease_name', ''),
'umls_cui': assoc.get('disease_id', ''),
'score': score,
'evidence_index': assoc.get('ei'),
'dsi': assoc.get('dsi'),
'dpi': assoc.get('dpi')
}
if score >= 0.7:
strong.append(entry)
elif score >= 0.4:
moderate.append(entry)
else:
weak.append(entry)
return {
'total_associations': len(associations),
'strong_associations': strong,
'moderate_associations': moderate,
'weak_associations': weak[:10],
'disease_pleiotropy': len(associations)
}Evidence Tier Assignment:
- DisGeNET Score >= 0.7 -> Consider T2 evidence (multiple validated sources)
- DisGeNET Score 0.4-0.7 -> Consider T3 evidence
- DisGeNET Score < 0.4 -> T4 evidence only
PATH 7: Druggability & Target Validation
Pharos/TCRD - Target Development Level
def get_pharos_target_info(tu, ids):
"""Get Pharos/TCRD target development level and druggability."""
gene_symbol = ids.get('symbol')
uniprot = ids.get('uniprot')
if gene_symbol:
result = tu.tools.Pharos_get_target(gene=gene_symbol)
elif uniprot:
result = tu.tools.Pharos_get_target(uniprot=uniprot)
else:
return {'status': 'error', 'message': 'Need gene symbol or UniProt'}
if result.get('status') == 'success' and result.get('data'):
target = result['data']
return {
'name': target.get('name'),
'symbol': target.get('sym'),
'tdl': target.get('tdl'),
'family': target.get('fam'),
'novelty': target.get('novelty'),
'description': target.get('description'),
'publications': target.get('publicationCount'),
'interpretation': interpret_tdl(target.get('tdl'))
}
return None
def interpret_tdl(tdl):
interpretations = {
'Tclin': 'Approved drug target - highest confidence for druggability',
'Tchem': 'Small molecule active - good chemical tractability',
'Tbio': 'Biologically characterized - may require novel modalities',
'Tdark': 'Understudied - limited data, high novelty potential'
}
return interpretations.get(tdl, 'Unknown')
def search_disease_targets(tu, disease_name):
"""Find targets associated with a disease via Pharos."""
result = tu.tools.Pharos_get_disease_targets(disease=disease_name, top=50)
if result.get('status') == 'success':
targets = result['data'].get('targets', [])
by_tdl = {'Tclin': [], 'Tchem': [], 'Tbio': [], 'Tdark': []}
for t in targets:
tdl = t.get('tdl', 'Unknown')
if tdl in by_tdl:
by_tdl[tdl].append(t)
return by_tdl
return NoneDepMap - Target Essentiality
def assess_target_essentiality(tu, ids):
"""
Is this target essential for cancer cell survival?
Negative effect scores = gene is essential (cells die upon KO)
"""
gene_symbol = ids.get('symbol')
if not gene_symbol:
return {'status': 'error', 'message': 'Need gene symbol'}
deps = tu.tools.DepMap_get_gene_dependencies(gene_symbol=gene_symbol)
if deps.get('status') == 'success':
return {
'gene': gene_symbol,
'data': deps.get('data', {}),
'interpretation': 'Negative scores indicate gene is essential for cell survival',
'note': 'Score < -0.5 is strongly essential, < -1.0 is extremely essential'
}
return NoneInterProScan - Novel Domain Prediction
For uncharacterized proteins, run InterProScan to predict domains and function:
def predict_protein_domains(tu, sequence, title="Query protein"):
"""
Run InterProScan for de novo domain prediction.
Use when: Novel/uncharacterized proteins, custom sequences, sparse InterPro annotations.
"""
result = tu.tools.InterProScan_scan_sequence(
sequence=sequence,
title=title,
go_terms=True,
pathways=True
)
if result.get('status') == 'success':
data = result.get('data', {})
if data.get('job_status') == 'RUNNING':
return {
'job_id': data.get('job_id'),
'status': 'running',
'note': 'Use InterProScan_get_job_results to retrieve when ready'
}
return {
'domains': data.get('domains', []),
'domain_count': data.get('domain_count', 0),
'go_annotations': data.get('go_annotations', []),
'pathways': data.get('pathways', []),
'sequence_length': data.get('sequence_length')
}
return NoneBindingDB - Known Ligands & Binding Data
def get_bindingdb_ligands(tu, uniprot_id, affinity_cutoff=10000):
"""
Get ligands with measured binding affinities from BindingDB.
Critical for identifying chemical starting points and assessing tractability.
"""
result = tu.tools.BindingDB_get_ligands_by_uniprot(
uniprot=uniprot_id,
affinity_cutoff=affinity_cutoff
)
if result:
ligands = []
for entry in result:
ligands.append({
'smiles': entry.get('smile'),
'affinity_type': entry.get('affinity_type'),
'affinity_nM': entry.get('affinity'),
'monomer_id': entry.get('monomerid'),
'pmid': entry.get('pmid')
})
ligands.sort(key=lambda x: float(x['affinity_nM']) if x['affinity_nM'] else float('inf'))
return {
'total_ligands': len(ligands),
'ligands': ligands[:20],
'best_affinity': ligands[0]['affinity_nM'] if ligands else None
}
return {'total_ligands': 0, 'ligands': [], 'note': 'No ligands found in BindingDB'}Affinity Interpretation:
| Range | Level | Drug Potential |
|---|---|---|
| <1 nM | Ultra-potent | Clinical candidate |
| 1-10 nM | Highly potent | Drug-like |
| 10-100 nM | Potent | Good starting point |
| 100-1000 nM | Moderate | Needs optimization |
| >1000 nM | Weak | Early hit only |
PubChem BioAssay - Screening Data
def get_pubchem_assays_for_target(tu, gene_symbol):
"""Get bioassays targeting a gene from PubChem."""
assays = tu.tools.PubChem_search_assays_by_target_gene(gene_symbol=gene_symbol)
assay_info = []
if assays.get('data', {}).get('aids'):
for aid in assays['data']['aids'][:10]:
summary = tu.tools.PubChem_get_assay_summary(aid=aid)
targets = tu.tools.PubChem_get_assay_targets(aid=aid)
assay_info.append({
'aid': aid,
'summary': summary.get('data', {}),
'targets': targets.get('data', {})
})
return {
'total_assays': len(assays.get('data', {}).get('aids', [])),
'assay_details': assay_info
}PATH 8: Literature (Collision-Aware)
def path_literature_collision_aware(tu, ids):
"""Literature search with collision detection and filtering."""
symbol = ids['symbol']
full_name = ids.get('full_name', '')
uniprot = ids['uniprot']
synonyms = ids.get('synonyms', [])
# Step 1: Detect collisions
collision_filter = detect_collisions(tu, symbol, full_name)
# Step 2: Build high-precision seed queries
seed_queries = [
f'"{symbol}"[Title] AND (protein OR gene OR expression)',
f'"{full_name}"[Title]' if full_name else None,
f'"UniProt:{uniprot}"' if uniprot else None,
]
seed_queries = [q for q in seed_queries if q]
for syn in synonyms[:3]:
seed_queries.append(f'"{syn}"[Title]')
# Step 3: Execute seed queries and collect PMIDs
seed_pmids = set()
for query in seed_queries:
if collision_filter:
query = f"({query}){collision_filter}"
results = tu.tools.PubMed_search_articles(query=query, limit=30)
for article in results.get('articles', []):
seed_pmids.add(article.get('pmid'))
# Step 4: Expand via citation network (for sparse targets)
if len(seed_pmids) < 30:
expanded_pmids = set()
for pmid in list(seed_pmids)[:10]:
related = tu.tools.PubMed_get_related(pmid=pmid, limit=20)
for r in related.get('articles', []):
expanded_pmids.add(r.get('pmid'))
citing = tu.tools.EuropePMC_get_citations(pmid=pmid, limit=20)
for c in citing.get('citations', []):
expanded_pmids.add(c.get('pmid'))
seed_pmids.update(expanded_pmids)
# Step 5: Classify papers by evidence tier
papers_by_tier = {'T1': [], 'T2': [], 'T3': [], 'T4': []}
return {
'total_papers': len(seed_pmids),
'collision_filter_applied': collision_filter if collision_filter else 'None needed',
'seed_queries': seed_queries,
'papers_by_tier': papers_by_tier
}Retry Logic & Fallback Chains
Retry Policy
def call_with_retry(tu, tool_name, params, max_retries=3):
"""Call tool with retry logic."""
for attempt in range(max_retries):
try:
result = getattr(tu.tools, tool_name)(**params)
if result and not result.get('error'):
return result
except Exception as e:
if attempt < max_retries - 1:
time.sleep(2 ** attempt)
else:
return {'error': str(e), 'tool': tool_name, 'attempts': max_retries}
return NoneFallback Chains
| Primary Tool | Fallback 1 | Fallback 2 | Failure Action |
|---|---|---|---|
ChEMBL_get_target_activities | GtoPdb_search_ligands | OpenTargets drugs | Note in report |
intact_get_interactions | STRING_get_protein_interactions | OpenTargets interactions | Note in report |
GO_get_annotations_for_gene | OpenTargets GO | MyGene GO | Note in report |
GTEx_get_median_gene_expression | HPA_get_rna_expression_by_source | Note as unavailable | Document in report |
gnomad_get_gene_constraints | OpenTargets constraint | - | Note in report |
DGIdb_get_drug_gene_interactions | OpenTargets drugs | GtoPdb | Note in report |
Failure Surfacing Rule
NEVER silently skip failed tools. Always document:
### 7.1 Tissue Expression
**GTEx Data**: Unavailable (API timeout after 3 attempts)
**Fallback Data (HPA)**:
| Tissue | Expression Level | Specificity |
|--------|-----------------|-------------|
| Liver | High | Enhanced |
| Kidney | Medium | - |
*Note: For complete GTEx data, query directly at gtexportal.org*Target Intelligence Tool Reference
Complete reference of 225+ ToolUniverse tools for target research, organized by category.
1. Core Protein Information (UniProt)
| Tool | Parameters | Returns |
|---|---|---|
UniProt_get_entry_by_accession | accession | Complete protein entry |
UniProt_get_function_by_accession | accession | Functional annotations |
UniProt_get_recommended_name_by_accession | accession | Official protein name |
UniProt_get_alternative_names_by_accession | accession | Aliases and synonyms |
UniProt_get_organism_by_accession | accession | Species info |
UniProt_get_subcellular_location_by_accession | accession | Cellular localization |
UniProt_get_disease_variants_by_accession | accession | Disease variants |
UniProt_get_ptm_processing_by_accession | accession | PTMs, active sites |
UniProt_get_sequence_by_accession | accession | Amino acid sequence |
UniProt_get_isoform_ids_by_accession | accession | Splice isoforms |
UniProt_search | query, organism, limit, fields | Search results |
UniProt_id_mapping | ids, from_db, to_db | ID mappings |
UniProt_get_proteome | proteome_id | Proteome info |
UniProt_get_uniref_cluster | cluster_id | UniRef cluster |
UniProt_search_uniref | query, cluster_type, limit | UniRef search |
UniProt_get_uniparc_entry | upi | UniParc entry |
UniProt_search_uniparc | query, limit | UniParc search |
EBI Proteins API
| Tool | Parameters | Returns |
|---|---|---|
proteins_api_get_protein | accession, format | Comprehensive protein info |
proteins_api_get_features | accession | Protein features |
proteins_api_get_variants | accession | Protein variants |
proteins_api_get_comments | accession | Annotations/comments |
proteins_api_get_epitopes | accession | Epitope data |
proteins_api_get_proteomics | accession | Proteomics data |
proteins_api_get_xrefs | accession | Cross-references |
proteins_api_get_publications | accession | Related publications |
proteins_api_get_genome_mappings | accession | Genome mappings |
proteins_api_search | query | Search proteins |
2. Gene Information
MyGene (BioThings)
| Tool | Parameters | Returns |
|---|---|---|
MyGene_get_gene_annotation | gene_id, fields | Detailed gene annotation |
MyGene_query_genes | query, species, fields, size | Gene search |
MyGene_batch_query | gene_ids, species, fields | Batch gene query |
Ensembl
| Tool | Parameters | Returns |
|---|---|---|
ensembl_lookup_gene | gene_id, species | Gene lookup |
ensembl_get_sequence | id, type, species | DNA/protein sequence |
ensembl_get_variants | region, species | Variants in region |
ensembl_get_variation | id, species | Variation details |
ensembl_get_variation_phenotypes | id, species | Phenotype associations |
ensembl_get_xrefs | id, external_db | Cross-references |
ensembl_get_xrefs_by_name | name, species | Xrefs by gene name |
ensembl_get_regulatory_features | region, species | Regulatory features |
ensembl_get_genetree | id, prune_species | Gene tree |
ensembl_get_homology | species, symbol, target_species | Orthologs/paralogs |
ensembl_get_alignment | species, region | Genomic alignments |
ensembl_get_taxonomy | id | Taxonomy info |
ensembl_vep_region | species, region, allele | Variant effect prediction |
Other Gene Resources
| Tool | Parameters | Returns |
|---|---|---|
kegg_get_gene_info | gene_id | KEGG gene info |
kegg_find_genes | keyword, organism | KEGG gene search |
cBioPortal_get_genes | keyword | Cancer gene search |
civic_search_genes | gene_symbol | CIViC gene info |
gnomad_get_gene | gene_symbol | gnomAD gene data |
gnomad_search_variants | query | gnomAD gene search |
gnomad_get_gene_constraints | gene_symbol | Constraint scores |
3. Drug-Target Interactions
DGIdb
| Tool | Parameters | Returns |
|---|---|---|
DGIdb_get_drug_gene_interactions | genes, interaction_sources, interaction_types | Drug-gene interactions |
DGIdb_get_gene_druggability | genes | Druggability categories |
DGIdb_get_gene_info | genes | Gene info from DGIdb |
DGIdb_get_drug_info | drugs | Drug info from DGIdb |
ChEMBL
| Tool | Parameters | Returns |
|---|---|---|
ChEMBL_get_target | target_chembl_id, format | Target details |
ChEMBL_search_targets | pref_name__contains, organism, target_type, limit | Target search |
ChEMBL_get_target_activities | target_chembl_id__exact, limit | Bioactivity data |
ChEMBL_get_target_assays | target_chembl_id__exact, limit | Target assays |
ChEMBL_get_molecule_targets | molecule_chembl_id__exact, limit | Molecule targets |
ChEMBL_search_binding_sites | target_chembl_id | Binding sites |
ChEMBL_search_mechanisms | molecule_chembl_id, target_chembl_id | Mechanisms of action |
ChEMBL_get_molecule | chembl_id, format | Molecule details |
ChEMBL_search_molecules | pref_name__contains, limit | Molecule search |
ChEMBL_get_assay | assay_chembl_id | Assay details |
ChEMBL_search_activities | molecule_chembl_id, target_chembl_id, standard_type | Activity search |
DrugBank & GtoPdb
| Tool | Parameters | Returns |
|---|---|---|
drugbank_get_targets_by_drug_name_or_drugbank_id | query, exact_match, limit | Drug targets |
drugbank_get_drug_name_and_description_by_target_name | target_name | Drugs for target |
GtoPdb_search_targets | target_id | GtoPdb target info |
GtoPdb_search_targets | family_id | List targets |
GtoPdb_search_ligands | target_id | Target-ligand interactions |
GtoPdb_get_interactions | query | Interaction search |
STITCH
| Tool | Parameters | Returns |
|---|---|---|
STITCH_get_chemical_protein_interactions | identifiers, species, required_score, limit | Chemical-protein links |
STITCH_get_interaction_partners | identifiers, species | Interaction network |
STITCH_resolve_identifier | identifier, species | ID resolution |
GPCRdb (NEW - for GPCR Targets)
~35% of approved drugs target GPCRs. GPCRdb provides specialized data for G protein-coupled receptors.
| Tool | Parameters | Returns |
|---|---|---|
GPCRdb_get_protein | operation="get_protein", protein (entry name) | GPCR family, class, sequence info |
GPCRdb_list_proteins | operation="list_proteins", family (optional) | List GPCR families/proteins |
GPCRdb_get_structures | operation="get_structures", protein, state (optional) | Structures with receptor state (active/inactive) |
GPCRdb_get_ligands | operation="get_ligands", protein | Known ligands (agonists/antagonists) |
GPCRdb_get_mutations | operation="get_mutations", protein | Mutation effects on binding/signaling |
Entry name format: {gene_lower}_human (e.g., adrb2_human, drd2_human)
Key advantages:
- Active vs. inactive state structures
- Ballesteros-Weinstein residue numbering
- Curated ligand binding data
- Experimental mutation effects
Pharos/TCRD (NEW - Target Development Level)
NIH's Illuminating the Druggable Genome (IDG) portal provides TDL classification.
| Tool | Parameters | Returns |
|---|---|---|
Pharos_get_target | gene OR uniprot | TDL, family, novelty, description |
Pharos_search_targets | query, top | Target list with TDL |
Pharos_get_tdl_summary | - | TDL level descriptions |
Pharos_get_disease_targets | disease, top | Targets for disease with TDL |
TDL Classification:
| Level | Description | Druggability |
|---|---|---|
| Tclin | Approved drug targets | Highest |
| Tchem | Small molecule activities (IC50 < 30nM) | Good |
| Tbio | Biological annotations only | Moderate |
| Tdark | Understudied proteins | Unknown |
Example:
result = tu.tools.Pharos_get_target(gene="EGFR")
# Returns: tdl="Tclin", fam="Kinase", novelty=0.2, publicationCount=45000DepMap (NEW - Target Essentiality)
CRISPR knockout essentiality data from cancer cell lines.
| Tool | Parameters | Returns |
|---|---|---|
DepMap_get_gene_dependencies | gene_symbol | Gene essentiality data |
DepMap_get_cell_lines | tissue, cancer_type, page_size | Cell line metadata |
DepMap_search_cell_lines | query | Search cell lines |
DepMap_get_cell_line | model_id OR model_name | Detailed cell line info |
| Drug sensitivity (GDSC) | drug / cell-line / target | No TU tool — run the precision-oncology skill's scripts/gdsc_drug_response.py (GDSC bulk data, IC50/AUC) |
Effect Score Interpretation:
| Score | Meaning |
|---|---|
| < -1.0 | Strongly essential |
| -0.5 to -1.0 | Essential |
| -0.5 to 0 | Weakly essential |
| > 0 | Not essential |
Example:
deps = tu.tools.DepMap_get_gene_dependencies(gene_symbol="KRAS")
# Returns: Gene info, note about essentiality
cells = tu.tools.DepMap_get_cell_lines(cancer_type="Lung Cancer", page_size=10)
# Returns: Cell line names, cancer types, MSI statusInterProScan (NEW - Domain Prediction)
De novo domain/family prediction for novel sequences.
| Tool | Parameters | Returns |
|---|---|---|
InterProScan_scan_sequence | sequence, go_terms, pathways | Domains, GO terms, pathways |
InterProScan_get_job_status | job_id | Job status |
InterProScan_get_job_results | job_id | Completed results |
When to use: Novel proteins, Tdark targets, custom sequences.
Example:
# Submit sequence for analysis
result = tu.tools.InterProScan_scan_sequence(
sequence="MVLSPADKTNVKAAWGKVGAHAGEYGAEALERMFLSFPTTKTYFPHFDLSH...",
go_terms=True,
pathways=True
)
# Returns: job_id (if running) or domains/GO/pathways (if complete)
# Check job if still running
status = tu.tools.InterProScan_get_job_status(job_id="iprscan5-xxx")
results = tu.tools.InterProScan_get_job_results(job_id="iprscan5-xxx")BindingDB (NEW - Ligand Binding Data)
Experimental binding affinity data (Ki, IC50, Kd) for target-ligand pairs.
| Tool | Parameters | Returns |
|---|---|---|
BindingDB_get_ligands_by_uniprot | uniprot, affinity_cutoff | Ligands with affinities |
BindingDB_get_ligands_by_uniprots | uniprots, affinity_cutoff | Multi-target ligands |
BindingDB_get_ligands_by_pdb | pdb_ids, affinity_cutoff, sequence_identity | Structure-based ligands |
BindingDB_get_targets_by_compound | smiles, similarity_cutoff | Polypharmacology |
Example:
# Get ligands for EGFR
ligands = tu.tools.BindingDB_get_ligands_by_uniprot(
uniprot="P00533",
affinity_cutoff=100 # Only potent ligands <100 nM
)
# Returns: SMILES, affinity_type (Ki/IC50/Kd), affinity value, PMID
# Find targets for a compound
targets = tu.tools.BindingDB_get_targets_by_compound(
smiles="CC(=O)Nc1ccc(cc1)O",
similarity_cutoff=0.85
)
# Returns: proteins with similar compound activitiesAffinity Interpretation:
| Range | Level | Drug Potential |
|---|---|---|
| <1 nM | Ultra-potent | Clinical candidate |
| 1-30 nM | Tchem threshold | Drug-like |
| 30-100 nM | Potent | Good start |
| 100-1000 nM | Moderate | Needs optimization |
Human Protein Atlas (NEW - Expression)
Protein and RNA expression across tissues and cell lines.
| Tool | Parameters | Returns |
|---|---|---|
HPA_search_genes_by_query | search_query | Gene info, Ensembl ID |
HPA_generic_search | search_query, columns | Custom data fields |
HPA_get_comparative_expression_by_gene_and_cellline | gene_name, cell_line | Cancer vs normal |
Example:
# Search gene
gene = tu.tools.HPA_search_genes_by_query(search_query="EGFR")
# Returns: Gene name, Ensembl ID, synonyms
# Compare cancer cell line vs normal tissue
expr = tu.tools.HPA_get_comparative_expression_by_gene_and_cellline(
gene_name="EGFR",
cell_line="a549" # Lung cancer
)
# Returns: expression comparisonSupported Cell Lines: a549, mcf7, hela, hepg2, pc3, jurkat, rh30, siha, u251, ishikawa
PubChem BioAssay (NEW - Screening Data)
HTS screening data and dose-response curves.
| Tool | Parameters | Returns |
|---|---|---|
PubChem_search_assays_by_target_gene | gene_symbol | AIDs for gene |
PubChem_get_assay_summary | aid | Assay statistics |
PubChem_get_assay_targets | aid | Target info |
PubChem_get_assay_active_compounds | aid | Active CIDs |
PubChem_get_assay_dose_response | aid | IC50/EC50 data |
Example:
# Find assays for target
assays = tu.tools.PubChem_search_assays_by_target_gene(gene_symbol="EGFR")
# Returns: list of AIDs
# Get assay summary
summary = tu.tools.PubChem_get_assay_summary(aid=504526)
# Returns: active/inactive counts, target info
# Get active compounds
actives = tu.tools.PubChem_get_assay_active_compounds(aid=504526)
# Returns: CIDs of active compounds4. Open Targets Platform
Target-Centric
| Tool | Parameters | Returns |
|---|---|---|
OpenTargets_get_target_id_description_by_name | targetName | Target ID lookup |
OpenTargets_get_associated_drugs_by_target_ensemblID | ensemblID | Drugs for target |
OpenTargets_get_diseases_phenotypes_by_target_ensembl | ensemblID | Diseases for target |
OpenTargets_get_target_safety_profile_by_ensemblID | ensemblID | Safety info |
OpenTargets_get_target_tractability_by_ensemblID | ensemblID | Tractability |
OpenTargets_get_target_interactions_by_ensemblID | ensemblID | PPI via Open Targets |
OpenTargets_get_target_gene_ontology_by_ensemblID | ensemblID | GO terms |
OpenTargets_get_target_synonyms_by_ensemblID | ensemblID | Target synonyms |
OpenTargets_get_target_classes_by_ensemblID | ensemblID | Classifications |
OpenTargets_get_target_constraint_info_by_ensemblID | ensemblID | Constraint data |
OpenTargets_get_target_genomic_location_by_ensemblID | ensemblID | Genomic location |
OpenTargets_get_target_subcell_locations_by_ensembl_ID | ensemblID | Subcellular location |
OpenTargets_get_target_homologues_by_ensemblID | ensemblID | Homologs |
OpenTargets_get_target_enabling_packages_by_ensemblID | ensemblID | TEP info |
OpenTargets_get_chemical_probes_by_target_ensemblID | ensemblID | Chemical probes |
OpenTargets_get_biological_mouse_models_by_ensemblID | ensemblID | Mouse models |
OpenTargets_get_publications_by_target_ensemblID | ensemblID | Publications |
OpenTargets_get_similar_entities_by_target_ensemblID | ensemblID | Similar targets |
Disease-Target Evidence
| Tool | Parameters | Returns |
|---|---|---|
OpenTargets_get_associated_targets_by_disease_efoId | efoId | Targets for disease |
OpenTargets_target_disease_evidence | ensemblID, efoId | Evidence details |
disease_target_score | efoId, datasourceId | Disease-target scores |
5. Protein Structure
RCSB PDB
| Tool | Parameters | Returns |
|---|---|---|
get_protein_metadata_by_pdb_id | pdb_id | Basic metadata |
get_protein_classification_by_pdb_id | pdb_id | Protein classification |
get_sequence_by_pdb_id | pdb_id | PDB sequence |
get_binding_affinity_by_pdb_id | pdb_id | Binding affinity data |
get_target_cofactor_info | pdb_id | Cofactor info |
get_polymer_entity_annotations | entity_id | Polymer annotations |
get_uniprot_accession_by_entity_id | entity_id | UniProt from PDB |
get_gene_name_by_entity_id | entity_id | Gene name from PDB |
PDB_search_similar_structures | pdb_id | Similar structures |
get_polymer_entity_ids_by_pdb_id | pdb_id | Polymer entity IDs |
get_source_organism_by_pdb_id | pdb_id | Source organism |
get_citation_info_by_pdb_id | pdb_id | Citation info |
get_mutation_annotations_by_pdb_id | pdb_id | Mutation annotations |
get_assembly_info_by_pdb_id | pdb_id | Biological assembly |
get_taxonomy_by_pdb_id | pdb_id | Taxonomy |
get_crystallographic_properties_by_pdb_id | pdb_id | Crystal properties |
get_structure_validation_metrics_by_pdb_id | pdb_id | Validation metrics |
get_ligand_smiles_by_chem_comp_id | chem_comp_id | Ligand SMILES |
visualize_protein_structure_3d | pdb_id | 3D visualization |
PDBe
| Tool | Parameters | Returns |
|---|---|---|
pdbe_get_entry_summary | pdb_id | Entry summary |
pdbe_get_entry_quality | pdb_id | Quality metrics |
pdbe_get_entry_publications | pdb_id | Publications |
pdbe_get_entry_assemblies | pdb_id | Biological assemblies |
pdbe_get_entry_secondary_structure | pdb_id | Secondary structure |
pdbe_get_entry_molecules | pdb_id | Molecule info |
pdbe_get_entry_status | pdb_id | Entry status |
pdbe_get_entry_experiment | pdb_id | Experimental details |
AlphaFold
| Tool | Parameters | Returns |
|---|---|---|
alphafold_get_prediction | qualifier (UniProt) | Full 3D predictions |
alphafold_get_summary | qualifier | Summary/metadata |
alphafold_get_annotations | qualifier | Annotations |
EMDB
| Tool | Parameters | Returns |
|---|---|---|
EMDB_search_structures | query | EM structure search |
EMDB_get_structure | emdb_id | EM structure details |
6. Protein-Protein Interactions
STRING
| Tool | Parameters | Returns |
|---|---|---|
STRING_get_protein_interactions | protein_ids, species, confidence_score, network_type, limit | PPI network |
IntAct
| Tool | Parameters | Returns |
|---|---|---|
intact_get_interactions | identifier, format | Interactions |
intact_search_interactions | query, first, max | Interaction search |
intact_get_interactor | identifier, format | Interactor details |
intact_get_interaction_network | identifier, depth | Interaction network |
intact_get_interaction_details | interaction_id | Interaction details |
intact_get_interactions_by_organism | taxid, size | Organism interactions |
intact_get_interactions_by_complex | complex_id | Complex interactions |
intact_get_complex_details | complex_ac | Complex details |
Other PPI Sources
| Tool | Parameters | Returns |
|---|---|---|
BioGRID_get_interactions | gene_names, organism, interaction_type, limit | BioGRID PPI |
HPA_get_protein_interactions_by_gene | gene_symbol | HPA interactions |
humanbase_ppi_analysis | genes, tissue | HumanBase PPI |
Reactome_get_interactor | id | Reactome interactors |
pc_get_interactions | source, target | Pathway Commons |
7. Functional Annotations
Gene Ontology
| Tool | Parameters | Returns |
|---|---|---|
GO_get_annotations_for_gene | gene_id | GO annotations |
GO_get_genes_for_term | go_id, taxon, rows | Genes for GO term |
GO_search_terms | query | GO term search |
GO_get_term_details | id | GO term details |
GO_get_term_by_id | id | GO term info |
OpenTargets_get_gene_ontology_terms_by_goID | goId | GO term via OT |
InterPro & Pfam
| Tool | Parameters | Returns |
|---|---|---|
InterPro_get_protein_domains | protein_id | Domain annotations |
InterPro_search_domains | query, page_size | Domain search |
InterPro_get_domain_details | accession | Domain details |
Gene Set Enrichment
| Tool | Parameters | Returns |
|---|---|---|
enrichr_gene_enrichment_analysis | genes, gene_set_library | Enrichment analysis |
8. Pathways
Reactome
| Tool | Parameters | Returns |
|---|---|---|
Reactome_map_uniprot_to_pathways | id (UniProt) | Pathways for protein |
Reactome_map_uniprot_to_reactions | id | Reactions for protein |
Reactome_get_pathway | stId | Pathway details |
Reactome_get_pathway_reactions | stId | Pathway reactions |
Reactome_get_pathway_hierarchy | stId | Parent pathways |
Reactome_list_top_pathways | species | Top-level pathways |
Reactome_get_participants | stId | Reaction participants |
Reactome_get_reaction | stId | Reaction details |
Reactome_get_complex | stId | Complex details |
Reactome_list_species | - | All species |
Reactome_query_by_ids | ids, species | ID query |
Reactome_get_events_hierarchy | species | Full hierarchy |
Reactome_get_diseases | - | Disease pathways |
KEGG
| Tool | Parameters | Returns |
|---|---|---|
kegg_get_pathway_info | pathway_id | Pathway details |
kegg_search_pathway | keyword, org | Pathway search |
kegg_list_organisms | - | All organisms |
WikiPathways
| Tool | Parameters | Returns |
|---|---|---|
WikiPathways_get_pathway | wpid, format | Pathway content |
WikiPathways_search | query, organism | Pathway search |
Pathway Commons
| Tool | Parameters | Returns |
|---|---|---|
pc_search_pathways | query | Pathway search |
9. Gene Expression
GTEx
| Tool | Parameters | Returns |
|---|---|---|
GTEx_get_gene_expression | gencode_id, tissue_site_detail_id | Expression data |
GTEx_get_median_gene_expression | gencode_id | Median by tissue |
GTEx_get_top_expressed_genes | tissue_id | Top genes in tissue |
GTEx_get_expression_summary | gencode_id | Expression summary |
GTEx_get_eqtl_genes | tissue_id | eQTL genes |
GTEx_get_single_tissue_eqtls | gencode_id, tissue_id | Single tissue eQTL |
GTEx_get_multi_tissue_eqtls | gencode_id | Multi-tissue eQTL |
GTEx_calculate_eqtl | gencode_id, variant_id | Calculate eQTL |
Human Protein Atlas (HPA)
| Tool | Parameters | Returns |
|---|---|---|
HPA_search_genes_by_query | search_query | Gene search |
HPA_get_gene_basic_info_by_ensembl_id | ensembl_id | Basic gene info |
HPA_get_comprehensive_gene_details_by_ensembl_id | ensembl_id | Comprehensive details |
HPA_get_rna_expression_in_specific_tissues | ensembl_id, tissue | Tissue RNA expression |
HPA_get_rna_expression_by_source | ensembl_id | Expression by source |
HPA_get_subcellular_location | ensembl_id | Subcellular location |
HPA_get_disease_expression_by_gene_tissue_disease | ensembl_id, tissue, disease | Disease expression |
HPA_get_cancer_prognostics_by_gene | gene_symbol | Cancer prognostics |
HPA_get_biological_processes_by_gene | gene_symbol | Biological processes |
Single-Cell
| Tool | Parameters | Returns |
|---|---|---|
CELLxGENE_get_expression_data | gene_id, dataset_id | Single-cell expression |
CELLxGENE_get_gene_metadata | gene_id | Gene metadata |
10. Variants & Mutations
ClinVar
| Tool | Parameters | Returns |
|---|---|---|
ClinVar_search_variants | gene, condition, variant_id, max_results | Variant search |
ClinVar_get_variant_details | variant_id | Variant details |
ClinVar_get_clinical_significance | variant_id | Clinical significance |
dbSNP
| Tool | Parameters | Returns |
|---|---|---|
dbsnp_get_variant_by_rsid | rsid | dbSNP variant |
dbsnp_search_by_gene | gene_symbol | dbSNP by gene |
dbsnp_get_frequencies | rsid | Allele frequencies |
gnomAD
| Tool | Parameters | Returns |
|---|---|---|
gnomad_get_variant | variant_id | gnomAD variant |
gnomad_search_variants | query | Variant search |
gnomad_get_region | chrom, start, stop | Variants in region |
CIViC
| Tool | Parameters | Returns |
|---|---|---|
civic_get_variant | variant_id | CIViC variant |
civic_get_variants_by_gene | gene_symbol | Variants for gene |
civic_search_variants | query | Variant search |
Other Variant Sources
| Tool | Parameters | Returns |
|---|---|---|
MyVariant_get_variant_annotation | variant_id | MyVariant annotation |
MyVariant_query_variants | query | Variant query |
PharmGKB_search_variants | query | PharmGKB variants |
cBioPortal_get_mutations | gene_symbol, study_id | Cancer mutations |
RegulomeDB_query_variant | variant_id | Regulatory annotation |
gwas_search_snps | query | GWAS SNPs |
gwas_get_snp_by_id | snp_id | GWAS SNP details |
gwas_get_snps_for_gene | gene_symbol | GWAS SNPs for gene |
11. Literature
PubMed
| Tool | Parameters | Returns |
|---|---|---|
PubMed_search_articles | query, limit, api_key | Article search |
PubMed_get_article | pmid, api_key | Article metadata |
PubMed_get_related | pmid, limit | Related articles |
PubMed_get_cited_by | pmid, limit | Citing articles |
PubMed_get_links | pmid | External links |
Europe PMC
| Tool | Parameters | Returns |
|---|---|---|
EuropePMC_search_articles | query, limit | Article search |
EuropePMC_get_citations | source, article_id | Citations |
EuropePMC_get_references | source, article_id | References |
Other Literature
| Tool | Parameters | Returns |
|---|---|---|
PMC_search_papers | query | PMC full-text search |
PubTator3_LiteratureSearch | query | PubTator with NER |
PubTator3_EntityAutocomplete | query | Entity autocomplete |
openalex_search_works | query | OpenAlex publications |
openalex_literature_search | query | Literature search |
12. Pharmacogenomics
PharmGKB
| Tool | Parameters | Returns |
|---|---|---|
PharmGKB_get_gene_details | gene_symbol | Gene info |
PharmGKB_search_genes | query | Gene search |
PharmGKB_get_drug_details | drug_name | Drug details |
PharmGKB_search_drugs | query | Drug search |
PharmGKB_get_clinical_annotations | gene_symbol | Clinical annotations |
PharmGKB_get_dosing_guidelines | gene_symbol | Dosing guidelines |
OpenTargets_drug_pharmacogenomics_data | chemblId | OT pharmacogenomics |
fda_pharmacogenomic_biomarkers | - | FDA biomarkers |
13. Disease Associations
| Tool | Parameters | Returns |
|---|---|---|
OpenTargets_get_disease_ids_by_name | diseaseName | Disease ID lookup |
OpenTargets_get_disease_description_by_efoId | efoId | Disease description |
OpenTargets_get_associated_drugs_by_disease_efoId | efoId | Drugs for disease |
OpenTargets_get_publications_by_disease_efoId | efoId | Disease publications |
OpenTargets_get_disease_therapeutic_areas_by_efoId | efoId | Therapeutic areas |
gwas_search_studies | query | GWAS studies |
gwas_get_studies_for_trait | trait | Studies for trait |
gwas_search_associations | query | GWAS associations |
gwas_get_associations_for_trait | trait | Associations for trait |
GtoPdb_search_diseases | - | GtoPdb diseases |
GtoPdb_search_diseases | disease_id | Disease details |
Reactome_get_diseases | - | Reactome diseases |
OSL_get_efo_id_by_disease_name | disease_name | EFO ID lookup |
DisGeNET (NEW - Gene-Disease Associations)
DisGeNET integrates gene-disease associations from curated repositories, GWAS catalogs, animal models, and literature. Requires: DISGENET_API_KEY
| Tool | Parameters | Returns |
|---|---|---|
DisGeNET_search_gene | operation="search_gene", gene (symbol/ID), limit | Diseases associated with gene |
DisGeNET_search_disease | operation="search_disease", disease (name/UMLS CUI), limit | Genes associated with disease |
DisGeNET_get_gda | operation="get_gda", gene, disease, min_score | Gene-disease association details |
DisGeNET_get_vda | operation="get_vda", variant (rsID), limit | Variant-disease associations |
DisGeNET_get_disease_genes | operation="get_disease_genes", disease, limit | All genes for a disease |
Key metrics:
- GDA Score: 0-1 confidence score for gene-disease association
- Evidence Index: Number and diversity of sources
- Disease Specificity Index: How specific is gene to this disease
- Disease Pleiotropy Index: How many diseases gene is linked to
Interpretation:
- Score ≥0.7: Strong association (consider T2 evidence)
- Score 0.4-0.7: Moderate association
- Score <0.4: Weak/limited evidence
14. ID Conversion & Cross-References
| Tool | Parameters | Returns |
|---|---|---|
UniProt_id_mapping | ids, from_db, to_db | ID conversion |
OpenTargets_map_any_disease_id_to_all_other_ids | diseaseId | Disease ID mapping |
ebi_cross_reference_search | identifier, source | EBI cross-refs |
Reactome_query_by_ids | ids | Reactome ID lookup |
Common ID Mapping Combinations
| From | To | Tool Call |
|---|---|---|
| Gene Symbol → UniProt | UniProt_search(query='gene:EGFR AND organism_id:9606') | |
| UniProt → Ensembl | UniProt_id_mapping(ids=['P00533'], from_db='UniProtKB_AC-ID', to_db='Ensembl') | |
| Ensembl → UniProt | UniProt_id_mapping(ids=['ENSG00000146648'], from_db='Ensembl', to_db='UniProtKB') | |
| UniProt → PDB | Extract from UniProt entry cross-references | |
| Gene Symbol → Entrez | MyGene_query_genes(query='symbol:EGFR', species='human') | |
| UniProt → ChEMBL Target | ChEMBL_search_targets(pref_name__contains='EGFR', organism='Homo sapiens') |
Target Intelligence Report Format Specification
This document defines the comprehensive report format for target intelligence reports. ALL sections are REQUIRED unless marked optional.
Critical Requirements
1. Report File Creation
- Create the report file FIRST before any data collection
- File name:
[TARGET]_target_report.md(e.g.,EGFR_target_report.md) - Initialize all 14 section headers with
[Researching...]placeholders - Update sections progressively as data is retrieved
2. Citation Requirements (MANDATORY)
Every piece of data MUST cite its source. Include:
- The database name (UniProt, PDB, ChEMBL, etc.)
- The tool used (
UniProt_get_entry_by_accession, etc.) - The query parameters (accession, gene symbol, etc.)
3. Source Block Format
At the end of each major section, add a source block:
---
**Sources:**
- [Database]: `tool_name` (query_parameter)
- [Database]: `tool_name` (query_parameter)
---Report Template
# Target Intelligence Report: [FULL PROTEIN NAME]
**Generated**: [Date] | **Query**: [Original query] | **Completeness**: [X/8 paths successful]
---
## 1. Executive Summary
[2-3 sentence overview of the target covering: what it is, its primary function, druggability status, and key clinical relevance]
**Bottom Line**: [One sentence: Is this a good drug target? Why/why not?]
---
## 2. Target Identifiers
| Identifier Type | Value | Database |
|-----------------|-------|----------|
| Gene Symbol | [SYMBOL] | HGNC |
| UniProt Accession | [P#####] | UniProtKB |
| Ensembl Gene ID | [ENSG###] | Ensembl |
| Entrez Gene ID | [#####] | NCBI Gene |
| ChEMBL Target ID | [CHEMBL###] | ChEMBL |
| HGNC ID | [HGNC:####] | HGNC |
**Aliases**: [List all known aliases/synonyms]
---
## 3. Basic Information
### 3.1 Protein Description
- **Recommended Name**: [Full protein name]
- **Alternative Names**: [List]
- **Gene Name**: [Symbol] ([Full gene name])
- **Organism**: [Species] (Taxonomy ID: [####])
- **Protein Length**: [###] amino acids
- **Molecular Weight**: [###] kDa
- **Isoforms**: [Number] known isoforms
### 3.2 Protein Function
[Detailed description of protein function - at least 3-4 sentences covering:
- Primary molecular function
- Biological process involvement
- Cellular role
- Signaling pathway context]
### 3.3 Subcellular Localization
- **Primary Location**: [e.g., Plasma membrane]
- **Additional Locations**: [List]
- **Topology**: [e.g., Single-pass type I membrane protein]
---
## 4. Structural Biology
### 4.1 Experimental Structures (PDB)
| PDB ID | Resolution | Method | Ligand | Description |
|--------|------------|--------|--------|-------------|
| [####] | [#.#Å] | [X-ray/Cryo-EM/NMR] | [Ligand or Apo] | [Brief description] |
[List top 5-10 most relevant structures]
**Total PDB Entries**: [###]
**Best Resolution**: [#.#Å] ([PDB ID])
**Structure Coverage**: [Complete/Partial - which domains?]
### 4.2 AlphaFold Prediction
- **Available**: [Yes/No]
- **Confidence**: [High/Medium/Low - pLDDT scores]
- **Model URL**: [AlphaFold DB link]
### 4.3 Domain Architecture
| Domain | Position | InterPro ID | Description |
|--------|----------|-------------|-------------|
| [Domain name] | [Start-End] | [IPR######] | [Function] |
[List all domains]
### 4.4 Key Structural Features
- **Active Sites**: [List with positions]
- **Binding Sites**: [List - substrate, cofactor, drug binding]
- **PTM Sites**: [Phosphorylation, glycosylation, etc. with positions]
- **Disulfide Bonds**: [List]
### 4.5 Structural Druggability Assessment
- **Binding Pockets**: [Identified pockets suitable for small molecules]
- **Allosteric Sites**: [Known or predicted]
- **Antibody Epitopes**: [Surface accessibility for biologics]
---
## 5. Function & Pathways
### 5.1 Gene Ontology Annotations
**Molecular Function (MF)**:
| GO Term | GO ID | Evidence |
|---------|-------|----------|
| [Term] | [GO:#######] | [IDA/IEA/etc.] |
[List top 5-10]
**Biological Process (BP)**:
| GO Term | GO ID | Evidence |
|---------|-------|----------|
| [Term] | [GO:#######] | [IDA/IEA/etc.] |
[List top 5-10]
**Cellular Component (CC)**:
| GO Term | GO ID | Evidence |
|---------|-------|----------|
| [Term] | [GO:#######] | [IDA/IEA/etc.] |
[List top 5]
### 5.2 Pathway Involvement
| Pathway | Database | Pathway ID |
|---------|----------|------------|
| [Pathway name] | [Reactome/KEGG/WikiPathways] | [ID] |
[List top 10 pathways]
### 5.3 Functional Summary
[Paragraph describing the target's role in cellular signaling, disease mechanisms, and biological importance]
---
## 6. Protein-Protein Interactions
### 6.1 Interaction Network Summary
- **Total Interactors (STRING, score >0.7)**: [###]
- **Experimentally Validated (IntAct)**: [###]
- **Complex Membership**: [List complexes]
### 6.2 Top Interacting Partners
| Partner | Score | Interaction Type | Evidence | Biological Context |
|---------|-------|------------------|----------|-------------------|
| [Gene] | [0.###] | [Physical/Functional] | [Experimental/Predicted] | [Context] |
[List top 15-20 interactors]
### 6.3 Protein Complexes
| Complex Name | Members | Function |
|--------------|---------|----------|
| [Complex] | [List] | [Function] |
### 6.4 Interaction Network Implications
[Paragraph on network topology, hub status, and implications for drugging]
---
## 7. Expression Profile
### 7.1 Tissue Expression (GTEx/HPA)
| Tissue | Expression Level (TPM) | Specificity |
|--------|------------------------|-------------|
| [Tissue] | [###] | [High/Medium/Low] |
[List top 10 expressing tissues]
**Tissue Specificity Score**: [Score] ([Broadly expressed/Tissue-specific/Tissue-enriched])
### 7.2 Cell Type Expression
[Single-cell data if available - top cell types]
### 7.3 Disease-Relevant Expression
| Cancer/Disease | Expression Change | Prognostic Value |
|----------------|-------------------|------------------|
| [Disease] | [Up/Down/Unchanged] | [Favorable/Unfavorable/None] |
### 7.4 Expression-Based Druggability
- **Tumor vs Normal**: [Differential expression ratio]
- **Therapeutic Window**: [Assessment based on expression pattern]
---
## 8. Genetic Variation & Disease
### 8.1 Genetic Constraint Scores
| Metric | Value | Interpretation |
|--------|-------|----------------|
| pLI | [0.##] | [Highly constrained/Tolerant] |
| LOEUF | [0.##] | [Interpretation] |
| Missense Z-score | [#.##] | [Interpretation] |
| pRec | [0.##] | [Interpretation] |
### 8.2 Disease Associations (Open Targets)
| Disease | Association Score | Evidence Types | EFO ID |
|---------|-------------------|----------------|--------|
| [Disease] | [0.##] | [Genetic/Literature/etc.] | [EFO_#######] |
[List top 10 diseases]
### 8.3 Pathogenic Variants (ClinVar)
| Variant | Clinical Significance | Condition | Review Status |
|---------|----------------------|-----------|---------------|
| [p.XXX###YYY] | [Pathogenic/Likely pathogenic] | [Condition] | [Stars] |
[List notable pathogenic variants]
**Total ClinVar Entries**: [###]
**Pathogenic/Likely Pathogenic**: [###]
### 8.4 Cancer Mutations (COSMIC/cBioPortal)
| Mutation | Frequency | Cancer Types | Functional Impact |
|----------|-----------|--------------|-------------------|
| [Mutation] | [#%] | [Cancers] | [Activating/Inactivating/Unknown] |
[List recurrent cancer mutations]
### 8.5 Genetic Evidence Summary
[Paragraph summarizing genetic validation of the target]
---
## 9. Druggability & Pharmacology
### 9.1 Tractability Assessment (Open Targets)
| Modality | Tractability | Bucket | Evidence |
|----------|--------------|--------|----------|
| Small Molecule | [✅/⚠️/❌] | [1-10] | [Clinical/Predicted] |
| Antibody | [✅/⚠️/❌] | [1-10] | [Clinical/Predicted] |
| PROTAC | [✅/⚠️/❌] | [1-10] | [Structural feasibility] |
| Other Modalities | [✅/⚠️/❌] | [1-10] | [Evidence] |
### 9.2 Approved Drugs
| Drug Name | Brand Name | Mechanism | Indication | Approval Year |
|-----------|------------|-----------|------------|---------------|
| [Drug] | [Brand] | [Inhibitor/Agonist/etc.] | [Indication] | [Year] |
[List all approved drugs]
### 9.3 Clinical Pipeline
| Drug | Phase | Indication | Trial Count | Status |
|------|-------|------------|-------------|--------|
| [Drug] | [Phase I/II/III] | [Indication] | [###] | [Active/Completed] |
[List drugs in clinical development]
**Total Clinical Trials**: [###]
**Active Trials**: [###]
### 9.4 Bioactivity Data (ChEMBL)
- **Total Bioactivity Records**: [###]
- **Compounds Tested**: [###]
- **Most Potent Compound**: [Name] (IC50/Ki: [###] nM)
| Compound | Activity Type | Value | Target |
|----------|---------------|-------|--------|
| [Compound] | [IC50/Ki/Kd] | [###] nM | [Target form] |
[List top 5 most potent compounds]
### 9.5 Chemical Probes
| Probe | Selectivity | Use | Source |
|-------|-------------|-----|--------|
| [Probe] | [Selective/Broad] | [Recommended use] | [SGC/etc.] |
### 9.6 Drug Resistance
[Known resistance mechanisms, mutations, or mechanisms if applicable]
---
## 10. Safety Profile
### 10.1 Target Safety Liabilities (Open Targets)
| Safety Concern | Evidence | Severity | Organ System |
|----------------|----------|----------|--------------|
| [Concern] | [Animal/Human/Both] | [High/Medium/Low] | [System] |
### 10.2 Mouse Knockout Phenotypes
| Phenotype | Zygosity | Viability | Source |
|-----------|----------|-----------|--------|
| [Phenotype] | [Homo/Hetero] | [Viable/Lethal] | [IMPC/MGI] |
### 10.3 Known Drug Adverse Events
| Adverse Event | Frequency | Drug Class | Mechanism |
|---------------|-----------|------------|-----------|
| [Event] | [Common/Uncommon/Rare] | [Class] | [On-target/Off-target] |
### 10.4 Safety Summary
[Paragraph summarizing safety considerations for targeting this protein]
---
## 11. Literature & Research Landscape
### 11.1 Publication Metrics
| Metric | Value |
|--------|-------|
| Total Publications | [###,###] |
| Publications (Last 5 Years) | [###,###] |
| Publications (Last Year) | [###,###] |
| Drug-Related Publications | [###,###] |
| Clinical Publications | [###,###] |
### 11.2 Research Trend
- **Trend**: [Increasing/Stable/Declining]
- **Peak Year**: [Year]
- **Current Activity**: [High/Medium/Low]
### 11.3 Key Research Areas
[List current hot topics in research for this target]
### 11.4 Notable Recent Publications
| PMID | Title | Year | Key Finding |
|------|-------|------|-------------|
| [PMID] | [Title] | [Year] | [Finding] |
[List 3-5 important recent papers]
---
## 12. Competitive Landscape
### 12.1 Market Status
- **First-in-Class Approved**: [Yes/No - Drug name if yes]
- **Best-in-Class Status**: [Assessment]
- **Patent Landscape**: [Crowded/Moderate/Open]
### 12.2 Differentiation Opportunities
[List potential differentiation strategies for new entrants]
---
## 13. Summary & Recommendations
### 13.1 Target Validation Scorecard
| Criterion | Score (1-5) | Evidence Level |
|-----------|-------------|----------------|
| Genetic Evidence | [#] | [Strong/Moderate/Weak] |
| Expression Relevance | [#] | [Strong/Moderate/Weak] |
| Functional Understanding | [#] | [Strong/Moderate/Weak] |
| Druggability | [#] | [Strong/Moderate/Weak] |
| Safety Profile | [#] | [Favorable/Caution/Concern] |
| Competitive Position | [#] | [Open/Moderate/Crowded] |
**Overall Target Score**: [##/30]
### 13.2 Key Strengths
1. [Strength 1]
2. [Strength 2]
3. [Strength 3]
### 13.3 Key Challenges/Risks
1. [Challenge 1]
2. [Challenge 2]
3. [Challenge 3]
### 13.4 Recommendations
| Priority | Category | Recommendation |
|----------|----------|----------------|
| 🔴 HIGH | [Category] | [Recommendation] |
| 🟡 MEDIUM | [Category] | [Recommendation] |
| 🟢 LOW | [Category] | [Recommendation] |
| ℹ️ INFO | [Category] | [Recommendation] |
### 13.5 Next Steps
[Suggested follow-up analyses or experiments]
---
## 14. Data Sources & Methodology
### 14.1 Databases Queried
| Database | Section(s) | Queries | Status |
|----------|------------|---------|--------|
| UniProtKB | 2, 3, 4, 8 | [accession] | ✅ Success |
| RCSB PDB | 4 | [PDB IDs queried] | ✅ Success |
| AlphaFold DB | 4 | [accession] | ✅ Success |
| InterPro | 4 | [accession] | ✅ Success |
| Gene Ontology | 5 | [gene_id] | ✅ Success |
| Reactome | 5 | [accession] | ✅ Success |
| KEGG | 5 | [gene_id] | ✅ Success |
| STRING | 6 | [protein_ids] | ✅ Success |
| IntAct | 6 | [accession] | ✅ Success |
| GTEx | 7 | [gencode_id] | ✅ Success |
| Human Protein Atlas | 7 | [ensembl_id] | ✅ Success |
| gnomAD | 8 | [gene_symbol] | ✅ Success |
| ClinVar | 8 | [gene] | ✅ Success |
| Open Targets | 8, 9, 10 | [ensembl_id] | ✅ Success |
| ChEMBL | 9 | [target_chembl_id] | ✅ Success |
| DGIdb | 9 | [genes] | ✅ Success |
| PubMed | 11 | [query] | ✅ Success |
### 14.2 Tools Used by Section
| Section | Tools Used |
|---------|------------|
| 2. Identifiers | `UniProt_search`, `UniProt_id_mapping`, `MyGene_get_gene_annotation` |
| 3. Basic Info | `UniProt_get_entry_by_accession`, `UniProt_get_function_by_accession`, `UniProt_get_subcellular_location_by_accession` |
| 4. Structure | `get_protein_metadata_by_pdb_id`, `alphafold_get_prediction`, `InterPro_get_protein_domains`, `UniProt_get_ptm_processing_by_accession` |
| 5. Function | `GO_get_annotations_for_gene`, `Reactome_map_uniprot_to_pathways`, `kegg_get_gene_info` |
| 6. Interactions | `STRING_get_protein_interactions`, `intact_get_interactions`, `intact_get_complex_details` |
| 7. Expression | `GTEx_get_median_gene_expression`, `HPA_get_comprehensive_gene_details_by_ensembl_id`, `HPA_get_subcellular_location` |
| 8. Variants | `gnomad_get_gene_constraints`, `ClinVar_search_variants`, `OpenTargets_get_diseases_phenotypes_by_target_ensembl` |
| 9. Druggability | `OpenTargets_get_target_tractability_by_ensemblID`, `OpenTargets_get_associated_drugs_by_target_ensemblID`, `DGIdb_get_gene_druggability`, `ChEMBL_get_target_activities` |
| 10. Safety | `OpenTargets_get_target_safety_profile_by_ensemblID`, `OpenTargets_get_biological_mouse_models_by_ensemblID` |
| 11. Literature | `PubMed_search_articles`, `EuropePMC_search_articles`, `OpenTargets_get_publications_by_target_ensemblID` |
### 14.3 Data Freshness
- **Report Generated**: [YYYY-MM-DD HH:MM UTC]
- **UniProt Release**: [Release number if available]
- **PDB Last Updated**: [Date]
- **GTEx Version**: v8
- **gnomAD Version**: v4.0
### 14.4 Limitations & Data Gaps
[Document any issues encountered:]
- Tools that returned errors or empty results
- Sections with incomplete data
- Known data quality issues
- Databases that were unavailable
Example:
- ⚠️ IntAct returned no experimentally validated interactions
- ⚠️ Some ChEMBL activity data may include non-human orthologs
- ✅ All primary data sources queried successfully
---
## Appendix (Optional)
### A. Full Sequence
[Protein sequence in FASTA format if requested]
### B. Additional Structures
[Extended PDB list if many available]
### C. Complete Interaction List
[Full PPI list if requested]
### D. All Variants
[Complete variant list if requested]---
Section Completeness Checklist
For each report, verify ALL items:
Required Sections
- [ ] Section 1: Executive Summary with bottom line
- [ ] Section 2: All identifier types (UniProt, Ensembl, Entrez, ChEMBL)
- [ ] Section 3: Basic info with function description (3-4 sentences)
- [ ] Section 4: Structural data (PDB count, domains, AlphaFold)
- [ ] Section 5: GO terms (5-10 per category) and pathways (top 10)
- [ ] Section 6: PPI (15-20 interactors, complexes)
- [ ] Section 7: Expression (top 10 tissues, specificity)
- [ ] Section 8: Variants (constraint scores, top 10 diseases, pathogenic variants)
- [ ] Section 9: Druggability (tractability, all drugs, clinical pipeline)
- [ ] Section 10: Safety (liabilities, mouse KO, adverse events)
- [ ] Section 11: Literature (5 metrics, trend, key papers)
- [ ] Section 12: Competitive landscape
- [ ] Section 13: Scorecard and recommendations
- [ ] Section 14: Data sources and methodology
Data Minimums
- [ ] At least 5 PDB structures listed (if available)
- [ ] All protein domains included
- [ ] Top 10 GO terms per category
- [ ] Top 10 pathways
- [ ] Top 15-20 protein interactors
- [ ] Expression in top 10 tissues
- [ ] All 4 constraint scores (pLI, LOEUF, missense Z, pRec)
- [ ] Top 10 disease associations
- [ ] All approved drugs listed
- [ ] All safety concerns documented
- [ ] 5 publication metrics
- [ ] 3-5 recent key publications
- [ ] Scorecard with all 6 criteria
- [ ] At least 3 prioritized recommendations
Quality Checks
- [ ] Executive summary is 2-3 sentences
- [ ] Function description is 3-4 sentences
- [ ] All tables have data (no empty tables)
- [ ] Paragraphs provide synthesis, not just lists
- [ ] Recommendations are actionable and prioritized
- [ ] Limitations section is honest about data gaps
---
Section-Specific Guidance
Section 1: Executive Summary
Purpose: Give reader the key takeaways in 30 seconds
Include: 1. What the target IS (protein class, function) 2. Clinical relevance (disease associations) 3. Druggability status (has drugs? tractable?) 4. One-line recommendation
Example:
EGFR is a receptor tyrosine kinase that drives cell proliferation and is overexpressed in multiple cancers including NSCLC and glioblastoma. Multiple approved TKIs and monoclonal antibodies validate this as a highly druggable target, though resistance mutations remain a challenge.
>
Bottom Line: Well-validated, druggable target with clinical precedence; new programs should focus on resistance-breaking or novel modalities.
Section 4: Structural Biology
Purpose: Enable structure-based drug design decisions
Must include:
- Total PDB count and best resolution
- Coverage (which domains have structures?)
- AlphaFold availability and confidence
- Complete domain list with positions
- Key binding sites for drug design
Section 9: Druggability
Purpose: Assess feasibility and competitive landscape
Must include:
- Tractability for ALL modalities (SM, Ab, PROTAC, other)
- Complete list of approved drugs
- Clinical pipeline (phase, indication, status)
- ChEMBL bioactivity summary
- Chemical probes if available
Section 13: Recommendations
Purpose: Provide actionable next steps
Requirements:
- Use priority levels (HIGH/MEDIUM/LOW/INFO)
- Each recommendation must be actionable
- Include both opportunities and risks
- Suggest specific follow-up analyses
Example:
| Priority | Category | Recommendation |
|---|---|---|
| 🔴 HIGH | Validation | Strong genetic and clinical validation supports target |
| 🔴 HIGH | Competition | Crowded space - need clear differentiation strategy |
| 🟡 MEDIUM | Safety | Monitor for cardiotoxicity based on KO phenotype |
| 🟢 LOW | Structure | Additional cryo-EM of full-length protein would help |
| ℹ️ INFO | Literature | Review recent resistance mechanism papers |
Related skills
How it compares
Pick tooluniverse-target-research over single-database lookup skills when discovery teams need a cited, multi-path target dossier in one agent workflow.
FAQ
What inputs does tooluniverse-target-research accept?
tooluniverse-target-research accepts targets identified by gene symbol, UniProt accession, Ensembl ID, or gene name. It resolves identifiers first, then populates a report across expression, pathway, interaction, variant, and druggability databases.
What deliverable does tooluniverse-target-research produce?
tooluniverse-target-research creates a report-first TARGET_target_report.md file with graded T1-T4 evidence, inline source citations, and sections for structure, pathways, expression, variants, druggability, and literature.