
Tooluniverse Chemical Safety
- 349 installs
- 1.6k repo stars
- Updated August 4, 2026
- mims-harvard/tooluniverse
tooluniverse-chemical-safety is an agent skill that runs an 8-phase chemical and drug safety workflow integrating ADMET-AI predictions, CTD toxicogenomics, FDA labels, DrugBank, and STITCH for hazard assessment from SMIL
About
tooluniverse-chemical-safety is an agent skill from mims-harvard/tooluniverse that orchestrates comprehensive chemical and drug safety assessment across eight research phases. Phase 0 disambiguates compounds to SMILES, PubChem CID, and ChEMBL IDs; phases 1–2 run ADMET-AI predictive toxicology and ADMET profiling requiring `pip install tooluniverse[ml]` across nine ADMETAI tools; phases 3–7 query CTD toxicogenomics, FDA label warnings, DrugBank safety profiles, STITCH chemical-protein interactions, and ChEMBL structural alerts. A mandatory four-tier evidence grading system labels findings T1 through T4, requiring computational T3 predictions to be anchored by experimental T1 or T2 data when available. Developers and computational chemists reach for this skill when scoping lab protocols, evaluating formulation safety, profiling environmental contaminants, or building agent workflows that need AMES, DILI, LD50, hERG, GHS, and carcinogenicity endpoints from SMILES strings.
- Hazard classification lookup
- Exposure limit verification
- Regulatory list cross-check
- Lab protocol risk scoping
- ToolUniverse safety endpoints
Tooluniverse Chemical Safety by the numbers
- 349 all-time installs (skills.sh)
- +6 installs in the week ending Aug 4, 2026 (Skillselion tracking)
- Ranked #584 of 2,203 Security skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/mims-harvard/tooluniverse --skill tooluniverse-chemical-safetyAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 349 |
|---|---|
| repo stars | ★ 1.6k |
| Last updated | August 4, 2026 |
| Repository | mims-harvard/tooluniverse ↗ |
How do you assess chemical toxicity from SMILES?
Let agents check chemical hazard classes, exposure limits, and regulatory lists when scoping lab protocols, formulation safety, or environmental risk for new compounds.
Who is it for?
Computational developers and chemists scoping compound hazard, ADMET risk, or environmental toxicity before lab work or formulation decisions.
Skip if: Teams needing clinical trial design or regulatory submission authoring without compound-level SMILES toxicity evidence gathering.
When should I use this skill?
The user asks about chemical toxicity, drug safety profiling, ADMET properties, environmental health risks, toxicogenomics, or hazard assessment for a SMILES string or compound name.
What you get
Integrated risk assessment report with evidence-graded toxicity predictions, ADMET profiles, CTD gene-disease mappings, and regulatory safety extractions per compound.
- integrated risk assessment report
- evidence-graded toxicity endpoints
- chemical-gene-disease interaction maps
By the numbers
- Runs 8 structured research phases from disambiguation to risk synthesis
- Integrates 9 ADMETAI prediction tools from admetai_tools.json
- Applies 4-tier T1–T4 mandatory evidence grading on all findings
Files
Chemical Safety & Toxicology Assessment
Toxicity assessment: identify the chemical, check known hazards (GHS, IARC), then look for ADMET predictions. Dose makes the poison — always consider exposure level, as a compound that is toxic at high doses may be safe at relevant exposures. Distinguish between acute toxicity (LD50, GHS category) and chronic hazards (carcinogenicity, endocrine disruption) — they require different risk management approaches. Computational predictions (ADMETAI) are T3 evidence and must be anchored by experimental data from PubChemTox or FDA labels wherever available. When evidence conflicts between prediction and experiment, always defer to the experimental finding.
LOOK UP DON'T GUESS: never assume GHS categories, IARC classification, or CTD disease links — always call PubChemTox and CTD tools to retrieve current classifications before reporting.
Comprehensive chemical safety analysis integrating predictive AI models, curated toxicogenomics databases, regulatory safety data, and chemical-biological interaction networks.
When to Use This Skill
Triggers:
- "Is this chemical toxic?" / "Assess the safety profile of [drug/chemical]"
- "What are the ADMET properties of [SMILES]?"
- "What genes does [chemical] interact with?" / "What diseases are linked to [chemical] exposure?"
- "Drug safety assessment" / "Environmental health risk" / "Chemical hazard profiling"
Use Cases: 1. Predictive Toxicology: AI-predicted endpoints (AMES, DILI, LD50, carcinogenicity, hERG) via SMILES 2. ADMET Profiling: Absorption, distribution, metabolism, excretion, toxicity 3. Toxicogenomics: Chemical-gene-disease mapping from CTD 4. Regulatory Safety: FDA label warnings, contraindications, adverse reactions 5. Drug Safety: DrugBank safety + FDA labels combined 6. Chemical-Protein Interactions: STITCH-based interaction networks 7. Environmental Toxicology: Chemical-disease associations for contaminants
---
COMPUTE, DON'T DESCRIBE
When analysis requires computation (statistics, data processing, scoring, enrichment), write and run Python code via Bash. Don't describe what you would do — execute it and report actual results. Use ToolUniverse tools to retrieve data, then Python (pandas, scipy, statsmodels, matplotlib) to analyze it.
KEY PRINCIPLES
1. Report-first approach - Create report file FIRST, then populate progressively 2. Tool parameter verification - Verify params via get_tool_info before calling unfamiliar tools 3. Evidence grading - Grade all safety claims by evidence strength (T1-T4) 4. Citation requirements - Every toxicity finding must have inline source attribution 5. Mandatory completeness - All sections must exist with data or explicit "No data" notes 6. Disambiguation first - Resolve compound identity (name -> SMILES, CID, ChEMBL ID) before analysis 7. Negative results documented - "No toxicity signals found" is data; empty sections are failures 8. Conservative risk assessment - When evidence is ambiguous, flag as "requires further investigation" 9. English-first queries - Always use English chemical/drug names in tool calls
---
Evidence Grading System (MANDATORY)
| Tier | Symbol | Criteria | Examples |
|---|---|---|---|
| T1 | [T1] | Direct human evidence, regulatory finding | FDA boxed warning, clinical trial toxicity |
| T2 | [T2] | Animal studies, validated in vitro | Nonclinical toxicology, AMES positive, animal LD50 |
| T3 | [T3] | Computational prediction, association data | ADMET-AI prediction, CTD association |
| T4 | [T4] | Database annotation, text-mined | Literature mention, unvalidated database entry |
Evidence grades MUST appear in: Executive Summary, Toxicity Predictions, Regulatory Safety, Chemical-Gene Interactions, Risk Assessment.
---
Core Strategy: 8 Research Phases
Chemical/Drug Query
|
+-- PHASE 0: Compound Disambiguation (ALWAYS FIRST)
| Resolve name -> SMILES, PubChem CID, ChEMBL ID, formula, weight
|
+-- PHASE 1: Predictive Toxicology (ADMET-AI)
| AMES, DILI, ClinTox, carcinogenicity, LD50, hERG, skin reaction
| Stress response pathways, nuclear receptor activity
|
+-- PHASE 2: ADMET Properties
| BBB penetrance, bioavailability, clearance, CYP interactions, physicochemical
|
+-- PHASE 3: Toxicogenomics (CTD)
| Chemical-gene interactions, chemical-disease associations
|
+-- PHASE 4: Regulatory Safety (FDA Labels)
| Boxed warnings, contraindications, adverse reactions, nonclinical tox
|
+-- PHASE 5: Drug Safety Profile (DrugBank)
| Toxicity data, contraindications, drug interactions
|
+-- PHASE 6: Chemical-Protein Interactions (STITCH)
| Direct binding, off-target effects, interaction confidence
|
+-- PHASE 7: Structural Alerts (ChEMBL)
| PAINS, Brenk, Glaxo structural alerts
|
+-- SYNTHESIS: Integrated Risk Assessment
Risk classification, evidence summary, data gaps, recommendationsSee phase-procedures-detailed.md for complete tool parameters, decision logic, output templates, and fallback strategies for each phase.
---
Tool Summary by Phase
Phase 0: Compound Disambiguation
PubChem_get_CID_by_compound_name(name: str)PubChem_get_compound_properties_by_CID(cid: int)ChEMBL_get_molecule(if ChEMBL ID available)
Phase 1: Predictive Toxicology
Dependency: ADMET-AI tools require pip install tooluniverse[ml]. If unavailable, skip to Phase 3 and use CTD + PubChemTox as alternatives.ADMETAI_predict_toxicity(smiles: list[str]) - AMES, DILI, ClinTox, LD50, hERG, etc.ADMETAI_predict_stress_response(smiles: list[str])ADMETAI_predict_nuclear_receptor_activity(smiles: list[str])
Phase 2: ADMET Properties
ADMETAI_predict_BBB_penetrance/_bioavailability/_clearance_distribution/_CYP_interactions/_physicochemical_properties/_solubility_lipophilicity_hydration(all takesmiles: list[str])
Phase 3: Toxicogenomics
CTD_get_chemical_gene_interactions(input_terms: str) — chemical name, returns gene interactions across speciesCTD_get_chemical_diseases(input_terms: str) — chemical-disease associations with evidence type
Phase 3.5: PubChem Toxicity Data
PubChemTox_get_toxicity_values(cid: int) — LD50, LC50, NOAEL reference valuesPubChemTox_get_ghs_classification(cid: int) — GHS hazard classification and pictogramsPubChemTox_get_carcinogen_classification(cid: int) — NTP/IARC carcinogenicity assessmentsPubChemTox_get_acute_effects(cid: int) — acute toxicity by route/speciesPubChemTox_get_toxicity_summary(cid: int) — integrated toxicity overview
Phase 3.6: Adverse Outcome Pathways
AOPWiki_list_aops(keyword: str) — search for relevant AOPs by chemical/mechanismAOPWiki_get_aop(aop_id: int) — full AOP detail: MIE, key events, adverse outcome
Phase 3.7: Environmental Exposure Context (US facilities)
Use for exposure/environmental-justice screening — locate regulated facilities near a community before assessing population-level exposure.
EPA_search_tri_facilities(state,city,limit) — Toxics Release Inventory facilities reporting toxic chemical releasesEPA_search_frs_facilities(state,city,limit) — Facility Registry Service (all EPA-regulated facilities) for broader siting/permitting context
Phase 4: Regulatory Safety (for pharmaceuticals only)
Environmental chemicals: Skip Phases 4-5 (no FDA labels/DrugBank). Use CTD + PubChemTox + AOPWiki instead.
FDA_get_boxed_warning_info_by_drug_name/_contraindications_/_adverse_reactions_/_warnings_(all takedrug_name: str)
Phase 5: Drug Safety (for pharmaceuticals only)
drugbank_get_safety_by_drug_name_or_drugbank_id(query,case_sensitive,exact_match,limit- all 4 required)
Phase 6: Chemical-Protein Interactions
STITCH_get_chemical_protein_interactions(identifiers: list[str],species: int)- Fallback (if STITCH fails for industrial chemicals):
STRING_get_interaction_partnersfor key target genes (e.g., ESR1 for endocrine disruptors) DGIdb_get_drug_gene_interactions(genes: list[str]) — for target druggability context
Phase 7: Structural Alerts
ChEMBL_search_compound_structural_alerts(molecule_chembl_id: str)
---
Risk Classification Matrix
| Risk Level | Criteria |
|---|---|
| CRITICAL | FDA boxed warning OR multiple [T1] toxicity findings OR active DILI + active hERG |
| HIGH | FDA warnings OR [T2] animal toxicity OR multiple active ADMET endpoints |
| MEDIUM | Some [T3] predictions positive OR CTD disease associations OR structural alerts |
| LOW | All ADMET endpoints negative AND no FDA/DrugBank flags AND no CTD concerns |
| INSUFFICIENT DATA | Fewer than 3 phases returned data |
---
Report Structure
# Chemical Safety & Toxicology Report: [Compound Name]
**Generated**: YYYY-MM-DD | **SMILES**: [...] | **CID**: [...]
## Executive Summary (risk classification + key findings, all graded)
## 1. Compound Identity (disambiguation table)
## 2. Predictive Toxicology (ADMET-AI endpoints)
## 3. ADMET Profile (absorption, distribution, metabolism, excretion)
## 4. Toxicogenomics (CTD chemical-gene-disease)
## 5. Regulatory Safety (FDA label data)
## 6. Drug Safety Profile (DrugBank)
## 7. Chemical-Protein Interactions (STITCH network)
## 8. Structural Alerts (ChEMBL)
## 9. Integrated Risk Assessment (classification, evidence summary, gaps, recommendations)
## Appendix: Methods and Data SourcesSee report-templates.md for full section templates with example tables.
---
Mandatory Completeness Checklist
- [ ] Phase 0: Compound disambiguated (SMILES + CID minimum)
- [ ] Phase 1: At least 5 toxicity endpoints or "prediction unavailable"
- [ ] Phase 2: ADMET A/D/M/E sections or "not available"
- [ ] Phase 3: CTD queried; results or "no data in CTD"
- [ ] Phase 4: FDA labels queried; results or "not FDA-approved"
- [ ] Phase 5: DrugBank queried; results or "not found"
- [ ] Phase 6: STITCH queried; results or "no data available"
- [ ] Phase 7: Structural alerts checked or "ChEMBL ID not available"
- [ ] Synthesis: Risk classification with evidence summary
- [ ] Evidence Grading: All findings have [T1]-[T4] annotations
- [ ] Data Gaps: Explicitly listed
---
Common Use Patterns
1. Novel Compound: SMILES -> Phase 0 (resolve) -> Phase 1 (toxicity) -> Phase 2 (ADMET) -> Phase 7 (structural alerts) -> Synthesis 2. Approved Drug Review: Drug name -> All phases (0-7) -> Complete safety dossier 3. Environmental Chemical: Chemical name -> Phase 0 -> Phase 1-2 -> Phase 3 (CTD, key) -> Phase 6 (STITCH) -> Synthesis 4. Batch Screening: Multiple SMILES -> Phase 0 -> Phase 1-2 (batch) -> Comparative table -> Synthesis 5. Toxicogenomic Deep-Dive: Chemical + gene/disease interest -> Phase 0 -> Phase 3 (expanded CTD) -> Literature -> Synthesis
---
Limitations
- ADMET-AI: Computational [T3]; should not replace experimental testing
- CTD: May lag behind latest literature by 6-12 months
- FDA: Only covers FDA-approved drugs; not applicable to environmental chemicals
- DrugBank: Primarily drugs; limited industrial chemical coverage
- STITCH: Lower score thresholds increase false positives
- ChEMBL: Structural alerts require ChEMBL ID; not all compounds have one
- Novel compounds: May only have ADMET-AI predictions (no database evidence)
- SMILES validity: Invalid SMILES cause ADMET-AI failures
---
Reference Files
- phase-procedures-detailed.md - Complete tool parameters, decision logic, output templates, fallback strategies per phase
- evidence-grading.md - Evidence grading details and examples
- report-templates.md - Full report section templates with example tables
- phase-details.md - Additional phase context
- test_skill.py - Test suite
---
Summary
Total tools integrated: 25+ tools across 6 databases (ADMET-AI, CTD, FDA, DrugBank, STITCH, ChEMBL)
Best for: Drug safety assessment, chemical hazard profiling, environmental toxicology, ADMET characterization, toxicogenomic analysis
Outputs: Structured markdown report with risk classification (Critical/High/Medium/Low), evidence grading [T1-T4], and actionable recommendations
Evidence Grading System
Grade every toxicity claim by evidence strength.
Tier Definitions
| Tier | Symbol | Criteria | Examples |
|---|---|---|---|
| T1 | [T1] | Direct human evidence, regulatory finding | FDA boxed warning, clinical trial toxicity, human case reports |
| T2 | [T2] | Animal studies, validated in vitro | Nonclinical toxicology, AMES positive, animal LD50 |
| T3 | [T3] | Computational prediction, association data | ADMET-AI prediction, CTD association, QSAR model |
| T4 | [T4] | Database annotation, text-mined | Literature mention, database entry without validation |
Required Evidence Grading Locations
Evidence grades MUST appear in: 1. Executive Summary - Key toxicity findings graded 2. Toxicity Predictions - Every ADMET-AI endpoint with confidence note 3. Regulatory Safety - FDA findings marked [T1] 4. Chemical-Gene Interactions - CTD data marked by curation status 5. Risk Assessment - Final risk classification with supporting evidence tiers
Risk Classification Matrix
| Risk Level | Criteria |
|---|---|
| CRITICAL | FDA boxed warning present OR multiple [T1] toxicity findings OR active DILI + active hERG |
| HIGH | FDA warnings present OR [T2] animal toxicity OR multiple active ADMET endpoints |
| MEDIUM | Some [T3] predictions positive OR CTD disease associations OR structural alerts |
| LOW | All ADMET endpoints negative AND no FDA/DrugBank safety flags AND no CTD concerns |
| INSUFFICIENT DATA | Fewer than 3 phases returned data; cannot make confident assessment |
Phase Details: Chemical Safety Assessment
Phase 0: Compound Disambiguation (ALWAYS FIRST)
CRITICAL: Resolve compound identity before any analysis.
Input Types Handled
| Input Format | Resolution Strategy |
|---|---|
| Drug name (e.g., "Aspirin") | PubChem_get_CID_by_compound_name -> get SMILES from properties |
| SMILES string | Use directly for ADMET-AI; resolve to CID for other tools |
| PubChem CID | PubChem_get_compound_properties_by_CID -> get SMILES + name |
| ChEMBL ID | ChEMBL_get_molecule -> get SMILES + properties |
Resolution Steps
1. Input detection: Determine if input is name, SMILES, CID, or ChEMBL ID
- SMILES: contains typical SMILES characters (=, #, [, ], (, ), c, n, o and no spaces in middle)
- CID: numeric only
- ChEMBL: starts with "CHEMBL"
- Otherwise: treat as compound name
2. Name to CID: PubChem_get_CID_by_compound_name(name=<compound_name>) 3. CID to properties: PubChem_get_compound_properties_by_CID(cid=<cid>) 4. Extract SMILES: Get SMILES from PubChem properties (field: ConnectivitySMILES, CanonicalSMILES, or IsomericSMILES) 5. Store resolved IDs: Maintain dict with name, smiles, cid, formula, weight, inchi
---
Phase 1: Predictive Toxicology (ADMET-AI)
When: SMILES is available
| Tool | Predicted Endpoints | Parameter |
|---|---|---|
ADMETAI_predict_toxicity | AMES, Carcinogens_Lagunin, ClinTox, DILI, LD50_Zhu, Skin_Reaction, hERG | smiles: list[str] |
ADMETAI_predict_stress_response | Stress response pathway activation (ARE, ATAD5, HSE, MMP, p53) | smiles: list[str] |
ADMETAI_predict_nuclear_receptor_activity | AhR, AR, ER, PPARg, Aromatase nuclear receptor activity | smiles: list[str] |
Workflow
1. Call all 3 tools with smiles=[resolved_smiles] 2. Classification endpoints: Active (1) = toxic signal, Inactive (0) = no signal 3. Regression endpoints (LD50): Report numerical value with context 4. All predictions graded [T3]
Decision Logic
- hERG = Active: Flag prominently (cardiac safety risk)
- AMES = Active: Flag prominently (mutagenicity concern)
- DILI = Active: Flag prominently (liver toxicity concern)
- Multiple SMILES: Can batch up to ~10 SMILES in single call
- Failed prediction: Note "prediction unavailable" (don't fail entire report)
---
Phase 2: ADMET Properties
When: SMILES is available
| Tool | Properties Predicted | Parameter |
|---|---|---|
ADMETAI_predict_BBB_penetrance | Blood-brain barrier crossing probability | smiles: list[str] |
ADMETAI_predict_bioavailability | Oral bioavailability (F20%, F30%) | smiles: list[str] |
ADMETAI_predict_clearance_distribution | Clearance, VDss, half-life, PPB | smiles: list[str] |
ADMETAI_predict_CYP_interactions | CYP1A2, 2C9, 2C19, 2D6, 3A4 inhibition/substrate | smiles: list[str] |
ADMETAI_predict_physicochemical_properties | LogP, LogD, LogS, MW, pKa | smiles: list[str] |
ADMETAI_predict_solubility_lipophilicity_hydration | Aqueous solubility, lipophilicity, hydration free energy | smiles: list[str] |
Workflow
1. Call all 6 ADMET tools in parallel 2. Compile into Absorption / Distribution / Metabolism / Excretion sections 3. Assess Lipinski Rule of 5 compliance 4. Flag drug-drug interaction risks from CYP inhibition profiles
Decision Logic
- BBB penetrant + toxicity: Flag as neurotoxicity risk
- Low bioavailability: Note absorption concerns
- CYP3A4 inhibitor: Flag high DDI risk
- Lipinski violations: Report drug-likeness assessment
---
Phase 3: Toxicogenomics (CTD)
When: Compound name is resolved
| Tool | Function | Parameter |
|---|---|---|
CTD_get_chemical_gene_interactions | Genes affected by chemical | input_terms: str |
CTD_get_chemical_diseases | Diseases linked to chemical exposure | input_terms: str |
Workflow
1. Call both CTD tools with input_terms=compound_name 2. Parse gene interactions: extract gene symbols, interaction types 3. Parse disease associations: extract disease names, evidence types (marker/mechanism/therapeutic) 4. Grade curated as [T2], inferred as [T3]
Decision Logic
- Direct vs inferred: CTD separates curated direct evidence from inferred associations
- Therapeutic vs toxic: Disease associations can be therapeutic or adverse
- Prioritize marker/mechanism: Stronger causal evidence than simple associations
---
Phase 4: Regulatory Safety (FDA Labels)
When: Compound has an approved drug name
| Tool | Information Retrieved | Parameter |
|---|---|---|
FDA_get_boxed_warning_info_by_drug_name | Black box warnings (most serious) | drug_name: str |
FDA_get_contraindications_by_drug_name | Absolute contraindications | drug_name: str |
FDA_get_adverse_reactions_by_drug_name | Known adverse reactions | drug_name: str |
FDA_get_warnings_by_drug_name | Warnings and precautions | drug_name: str |
FDA_get_nonclinical_toxicology_info_by_drug_name | Animal toxicology data | drug_name: str |
FDA_get_carcinogenic_mutagenic_fertility_by_drug_name | Carcinogenicity/mutagenicity/fertility | drug_name: str |
Workflow
1. Call all 6 FDA tools in parallel 2. Prioritize: Boxed Warnings > Contraindications > Warnings > Adverse Reactions 3. All FDA label data is [T1] evidence 4. Boxed warning present: Flag as CRITICAL in executive summary 5. No FDA data: Note "Not an FDA-approved drug" and continue
---
Phase 5: Drug Safety Profile (DrugBank)
When: Compound is a known drug
| Tool | Information | Parameters |
|---|---|---|
drugbank_get_safety_by_drug_name_or_drugbank_id | Toxicity, contraindications | query: str, case_sensitive: bool, exact_match: bool, limit: int |
Workflow
1. Call with query=drug_name, case_sensitive=False, exact_match=False, limit=5 2. Parse toxicity information, overdose data, contraindications 3. Cross-reference with FDA data; if conflict, defer to FDA [T1]
---
Phase 6: Chemical-Protein Interactions (STITCH)
When: Compound can be identified by name or SMILES
| Tool | Function | Parameters |
|---|---|---|
STITCH_resolve_identifier | Resolve to STITCH ID | identifier: str, species: int (9606=human) |
STITCH_get_chemical_protein_interactions | Get interactions | identifiers: list[str], species: int, required_score: int |
STITCH_get_interaction_partners | Get interaction network | identifiers: list[str], species: int, limit: int |
Workflow
1. Resolve compound: STITCH_resolve_identifier(identifier=compound_name, species=9606) 2. Get interactions with required_score=700 3. Flag safety-relevant targets: hERG (cardiac), CYP enzymes (metabolism), nuclear receptors (endocrine)
Confidence Levels
- >900: Well-established interaction [T2]
- 700-900: Probable interaction [T3]
- 400-700: Possible interaction, needs validation [T4]
---
Phase 7: Structural Alerts (ChEMBL)
When: ChEMBL molecule ID is available
| Tool | Function | Parameters |
|---|---|---|
ChEMBL_search_compound_structural_alerts | Find structural alert matches | molecule_chembl_id: str, limit: int |
Workflow
1. Call with molecule_chembl_id=chembl_id, limit=20 2. Parse alert types: PAINS, Brenk, Glaxo 3. No ChEMBL ID: Skip gracefully; note "structural alert analysis not available"
Chemical Safety: Detailed Phase Procedures
Referenced from SKILL.md. Contains detailed tool parameters, output templates, decision logic, and fallback strategies.
---
Phase 0: Compound Disambiguation (ALWAYS FIRST)
Input Types Handled
| Input Format | Resolution Strategy |
|---|---|
| Drug name (e.g., "Aspirin") | PubChem_get_CID_by_compound_name -> get SMILES from properties |
| SMILES string | Use directly for ADMET-AI; resolve to CID for other tools |
| PubChem CID | PubChem_get_compound_properties_by_CID -> get SMILES + name |
| ChEMBL ID | ChEMBL_get_molecule -> get SMILES + properties |
Resolution Steps
1. Input detection: Determine if input is name, SMILES, CID, or ChEMBL ID
- SMILES: contains typical SMILES characters (=, #, [, ], (, ), c, n, o and no spaces in middle)
- CID: numeric only
- ChEMBL: starts with "CHEMBL"
- Otherwise: treat as compound name
2. Name to CID: PubChem_get_CID_by_compound_name(name=<compound_name>) 3. CID to properties: PubChem_get_compound_properties_by_CID(cid=<cid>) 4. Extract SMILES: Get from PubChem properties (ConnectivitySMILES, CanonicalSMILES, or IsomericSMILES) 5. Store resolved IDs: Maintain dict with name, smiles, cid, formula, weight, inchi
Output Table
## Compound Identity
| Property | Value |
|----------|-------|
| **Name** | Acetaminophen |
| **PubChem CID** | 1983 |
| **SMILES** | CC(=O)Nc1ccc(O)cc1 |
| **Formula** | C8H9NO2 |
| **Molecular Weight** | 151.16 |---
Phase 1: Predictive Toxicology (ADMET-AI)
When: SMILES is available
Tools (all take smiles: list[str])
| Tool | Predicted Endpoints |
|---|---|
ADMETAI_predict_toxicity | AMES, Carcinogens_Lagunin, ClinTox, DILI, LD50_Zhu, Skin_Reaction, hERG |
ADMETAI_predict_stress_response | Stress response pathways (ARE, ATAD5, HSE, MMP, p53) |
ADMETAI_predict_nuclear_receptor_activity | AhR, AR, ER, PPARg, Aromatase |
Decision Logic
- Multiple SMILES: Can batch up to ~10 SMILES in single call
- Failed prediction: Note "prediction unavailable" (don't fail entire report)
- Confidence: All AI predictions are [T3], not definitive
- Critical flags: hERG Active = cardiac risk, AMES Active = mutagenicity, DILI Active = liver toxicity
Output Table
### Toxicity Predictions [T3]
| Endpoint | Prediction | Interpretation | Concern Level |
|----------|-----------|---------------|---------------|
| AMES Mutagenicity | Inactive | No mutagenic signal | Low |
| DILI | Active | Drug-induced liver injury risk | HIGH |
| hERG Inhibition | Active | Cardiac arrhythmia risk | HIGH |
*All predictions from ADMET-AI. Evidence tier: [T3]*---
Phase 2: ADMET Properties
When: SMILES is available
Tools (all take smiles: list[str])
| Tool | Properties |
|---|---|
ADMETAI_predict_BBB_penetrance | Blood-brain barrier crossing |
ADMETAI_predict_bioavailability | Oral bioavailability (F20%, F30%) |
ADMETAI_predict_clearance_distribution | Clearance, VDss, half-life, PPB |
ADMETAI_predict_CYP_interactions | CYP1A2, 2C9, 2C19, 2D6, 3A4 inhibition/substrate |
ADMETAI_predict_physicochemical_properties | LogP, LogD, LogS, MW, pKa |
ADMETAI_predict_solubility_lipophilicity_hydration | Aqueous solubility, lipophilicity, hydration free energy |
Decision Logic
- BBB penetrant + toxicity: Flag as neurotoxicity risk
- Low bioavailability: F20% = Low -> absorption concerns
- CYP inhibitor: CYP3A4 inhibitor = Yes -> high DDI risk
- Lipinski violations: Count and report drug-likeness assessment
Output Format
### ADMET Profile [T3]
#### Absorption
| Property | Value | Interpretation |
|----------|-------|----------------|
| BBB Penetrance | Yes | Crosses blood-brain barrier |
| Bioavailability (F20%) | 85% | Good oral absorption |
#### Metabolism
| CYP Enzyme | Substrate | Inhibitor |
|------------|-----------|-----------|
| CYP3A4 | Yes | Yes (DDI risk) |---
Phase 3: Toxicogenomics (CTD)
When: Compound name is resolved
Tools
| Tool | Parameter | Notes |
|---|---|---|
CTD_get_chemical_gene_interactions | input_terms: str | Chemical name, MeSH name, CAS RN, or MeSH ID |
CTD_get_chemical_diseases | input_terms: str | Same |
Decision Logic
- Direct evidence vs inferred: CTD separates curated (grade [T2]) from inferred ([T3])
- Therapeutic vs toxic: Disease associations can be therapeutic or adverse
- Prioritize marker/mechanism: Stronger causal evidence than simple associations
---
Phase 4: Regulatory Safety (FDA Labels)
When: Compound has an approved drug name
Tools (all take drug_name: str)
| Tool | Information |
|---|---|
FDA_get_boxed_warning_info_by_drug_name | Black box warnings (most serious) |
FDA_get_contraindications_by_drug_name | Absolute contraindications |
FDA_get_adverse_reactions_by_drug_name | Known adverse reactions |
FDA_get_warnings_by_drug_name | Warnings and precautions |
FDA_get_nonclinical_toxicology_info_by_drug_name | Animal toxicology data |
FDA_get_carcinogenic_mutagenic_fertility_by_drug_name | Carcinogenicity/mutagenicity/fertility |
Decision Logic
- Boxed warning present: Flag as CRITICAL in executive summary
- No FDA data: Note "Not an FDA-approved drug" and continue
- All FDA findings: Grade as [T1]; nonclinical toxicology as [T2]
---
Phase 5: Drug Safety Profile (DrugBank)
When: Compound is a known drug
Tool
drugbank_get_safety_by_drug_name_or_drugbank_id(query=drug_name, case_sensitive=False, exact_match=False, limit=5)
All 4 parameters required. Cross-reference with FDA data; defer to FDA [T1] on conflicts.
---
Phase 6: Chemical-Protein Interactions (STITCH)
Tools
| Tool | Parameters |
|---|---|
STITCH_resolve_identifier | identifier: str, species: int (9606=human) |
STITCH_get_chemical_protein_interactions | identifiers: list[str], species: int, required_score: int |
STITCH_get_interaction_partners | identifiers: list[str], species: int, limit: int |
Confidence Scoring
- >900: Well-established [T2]
- 700-900: Probable [T3]
- 400-700: Possible, needs validation [T4]
- Flag safety-relevant targets: hERG (cardiac), CYP enzymes (metabolism), nuclear receptors (endocrine)
---
Phase 7: Structural Alerts (ChEMBL)
When: ChEMBL molecule ID is available
Tool
ChEMBL_search_compound_structural_alerts(molecule_chembl_id=chembl_id, limit=20)
Alert types: PAINS (pan-assay interference), Brenk (medicinal chemistry), Glaxo (GSK). No ChEMBL ID -> skip gracefully.
---
Fallback Strategies
Compound Resolution
- Primary: PubChem by name -> CID -> properties -> SMILES
- Fallback 1: ChEMBL search by name -> molecule -> SMILES
- Fallback 2: If SMILES provided directly, skip name resolution
Toxicity Prediction
- Primary: All 9 ADMET-AI endpoints
- Fallback: If fails, note "prediction failed" and continue with database evidence
Regulatory Data
- Primary: FDA labels by drug name
- Fallback: Try alternative names (brand vs generic)
CTD Data
- Primary: Search by common chemical name
- Fallback: Try MeSH name
---
Tool Parameter Quick Reference
| Tool | Parameter Name | Type | Notes |
|---|---|---|---|
| All ADMETAI tools | smiles | list[str] | Always a list, even for single compound |
| All CTD tools | input_terms | str | Chemical name, MeSH name, CAS RN, or MeSH ID |
| All FDA tools | drug_name | str | Brand or generic drug name |
| drugbank_get_safety_* | query, case_sensitive, exact_match, limit | str, bool, bool, int | All 4 required |
| STITCH_resolve_identifier | identifier, species | str, int | species=9606 for human |
| STITCH_get_chemical_protein_interactions | identifiers, species, required_score | list[str], int, int | required_score=400 default |
| PubChem_get_CID_by_compound_name | name | str | Compound name (not SMILES) |
| PubChem_get_compound_properties_by_CID | cid | int | Numeric CID |
| ChEMBL_search_compound_structural_alerts | molecule_chembl_id | str | e.g., "CHEMBL112" |
Response Format Notes
- ADMET-AI:
{status: "success", data: {...}} - CTD: List of interaction/association objects
- FDA:
{status, data}with label text - DrugBank:
{data: [...]}with drug records - STITCH: List of interaction objects with scores
- PubChem CID lookup:
{IdentifierList: {CID: [...]}} - PubChem properties: Dict with
CID,MolecularWeight,ConnectivitySMILES,IUPACName
Report Templates: Chemical Safety Assessment
Output Report Structure
All analyses generate a structured markdown report:
# Chemical Safety & Toxicology Report: [Compound Name]
**Generated**: YYYY-MM-DD HH:MM
**Compound**: [Name] | SMILES: [SMILES] | CID: [CID]
## Executive Summary
[2-3 sentence overview with risk classification and key findings, all graded]
## 1. Compound Identity
[Phase 0 results - disambiguation table]
## 2. Predictive Toxicology
[Phase 1 results - ADMET-AI toxicity endpoints]
## 3. ADMET Profile
[Phase 2 results - absorption, distribution, metabolism, excretion]
## 4. Toxicogenomics
[Phase 3 results - CTD chemical-gene-disease relationships]
## 5. Regulatory Safety
[Phase 4 results - FDA label information]
## 6. Drug Safety Profile
[Phase 5 results - DrugBank data]
## 7. Chemical-Protein Interactions
[Phase 6 results - STITCH network]
## 8. Structural Alerts
[Phase 7 results - ChEMBL alerts]
## 9. Integrated Risk Assessment
[Synthesis - risk classification, evidence summary, data gaps, recommendations]
## Appendix: Methods and Data Sources
[Tool versions, databases queried, date of access]---
Completeness Checklist
Before finalizing any report, verify:
- [ ] Phase 0: Compound fully disambiguated (SMILES + CID at minimum)
- [ ] Phase 1: At least 5 toxicity endpoints reported or "prediction unavailable" noted
- [ ] Phase 2: ADMET profile with A/D/M/E sections or "not available" noted
- [ ] Phase 3: CTD queried; gene interactions and disease associations reported or "no data in CTD"
- [ ] Phase 4: FDA labels queried; results or "not an FDA-approved drug" noted
- [ ] Phase 5: DrugBank queried; results or "not found in DrugBank" noted
- [ ] Phase 6: STITCH queried; results or "no STITCH data available" noted
- [ ] Phase 7: Structural alerts checked or "ChEMBL ID not available" noted
- [ ] Synthesis: Risk classification provided with evidence summary
- [ ] Evidence Grading: All findings have [T1]-[T4] annotations
- [ ] Data Gaps: Explicitly listed in synthesis section
---
Example Tables
Compound Identity
| Property | Value |
|----------|-------|
| **Name** | Acetaminophen |
| **PubChem CID** | 1983 |
| **SMILES** | CC(=O)Nc1ccc(O)cc1 |
| **Formula** | C8H9NO2 |
| **Molecular Weight** | 151.16 |Toxicity Predictions [T3]
| Endpoint | Prediction | Interpretation | Concern Level |
|----------|-----------|---------------|---------------|
| AMES Mutagenicity | Inactive | No mutagenic signal | Low |
| ClinTox | Active | Clinical toxicity signal | HIGH |
| DILI | Active | Drug-induced liver injury risk | HIGH |
| LD50 (Zhu) | 2.45 log(mg/kg) | ~282 mg/kg (moderate) | Medium |
| hERG Inhibition | Active | Cardiac arrhythmia risk | HIGH |ADMET Profile
#### Absorption
| Property | Value | Interpretation |
|----------|-------|----------------|
| BBB Penetrance | Yes | Crosses blood-brain barrier |
| Bioavailability (F20%) | 85% | Good oral absorption |
#### Metabolism
| CYP Enzyme | Substrate | Inhibitor |
|------------|-----------|-----------|
| CYP1A2 | No | No |
| CYP3A4 | Yes | Yes (DDI risk) |Integrated Risk Assessment
### Overall Risk Classification: [HIGH]
| Dimension | Finding | Evidence Tier | Concern |
|-----------|---------|--------------|---------|
| ADMET Toxicity | DILI active, hERG active | [T3] | HIGH |
| FDA Label | Boxed warning for hepatotoxicity | [T1] | CRITICAL |
| CTD Toxicogenomics | 156 gene interactions | [T2] | HIGH |
### Key Safety Concerns
1. **Hepatotoxicity** [T1]: FDA boxed warning + ADMET-AI DILI + CTD liver associations
2. **Cardiac Risk** [T3]: hERG prediction + STITCH hERG interaction
### Data Gaps
- [ ] No in vivo genotoxicity data
- [ ] STITCH scores moderate (700-900)
### Recommendations
1. Avoid doses >4g/day [T1]
2. Monitor liver function in chronic use [T1]
3. Screen for CYP3A4 interactions [T3]---
Common Use Patterns
| Pattern | Input | Key Phases |
|---|---|---|
| Novel Compound | SMILES string | 0, 1, 2, 7, Synthesis |
| Approved Drug Review | Drug name | All phases (0-7) |
| Environmental Chemical | Chemical name | 0, 1, 2, 3 (CTD key), 6, Synthesis |
| Batch Screening | Multiple SMILES | 0, 1 (batch), 2 (batch), Synthesis |
| Toxicogenomic Deep-Dive | Chemical + gene/disease | 0, 3 (expanded), Synthesis |
#!/usr/bin/env python3
"""
Comprehensive Test Suite for Chemical Safety & Toxicology Skill
Tests all 8 phases of the skill workflow using real tool calls.
Validates tool parameters, response structures, and documentation accuracy.
"""
import sys
import time
import traceback
# Track results
test_results = []
total_tests = 0
passed_tests = 0
failed_tests = 0
def record_result(test_name, passed, details=""):
global total_tests, passed_tests, failed_tests
total_tests += 1
if passed:
passed_tests += 1
status = "PASS"
else:
failed_tests += 1
status = "FAIL"
test_results.append({"name": test_name, "status": status, "details": details})
print(f" [{status}] {test_name}")
if details and not passed:
print(f" Details: {details}")
def test_tool_loading():
"""Test 1: Verify all required tools load into ToolUniverse"""
print("\n=== Test 1: Tool Loading ===")
try:
from tooluniverse import ToolUniverse
tu = ToolUniverse()
tu.load_tools()
required_tools = [
# ADMET-AI tools (Phase 1 & 2)
"ADMETAI_predict_toxicity",
"ADMETAI_predict_BBB_penetrance",
"ADMETAI_predict_bioavailability",
"ADMETAI_predict_clearance_distribution",
"ADMETAI_predict_CYP_interactions",
"ADMETAI_predict_nuclear_receptor_activity",
"ADMETAI_predict_physicochemical_properties",
"ADMETAI_predict_solubility_lipophilicity_hydration",
"ADMETAI_predict_stress_response",
# CTD tools (Phase 3)
"CTD_get_chemical_gene_interactions",
"CTD_get_chemical_diseases",
# FDA tools (Phase 4)
"FDA_get_boxed_warning_info_by_drug_name",
"FDA_get_contraindications_by_drug_name",
"FDA_get_adverse_reactions_by_drug_name",
"FDA_get_warnings_by_drug_name",
"FDA_get_nonclinical_toxicology_info_by_drug_name",
# DrugBank (Phase 5)
"drugbank_get_safety_by_drug_name_or_drugbank_id",
# STITCH (Phase 6)
"STITCH_resolve_identifier",
"STITCH_get_chemical_protein_interactions",
# ChEMBL (Phase 7)
"ChEMBL_search_compound_structural_alerts",
# PubChem (Phase 0)
"PubChem_get_CID_by_compound_name",
"PubChem_get_compound_properties_by_CID",
]
all_tools = set(tu.all_tool_dict.keys())
missing = [t for t in required_tools if t not in all_tools]
if missing:
record_result("All required tools loaded", False,
f"Missing tools: {missing}")
else:
record_result("All required tools loaded", True,
f"All {len(required_tools)} tools found in {len(all_tools)} total tools")
return tu
except Exception as e:
record_result("Tool loading", False, f"Exception: {e}")
return None
def test_phase0_disambiguation(tu):
"""Test 2: Phase 0 - Compound disambiguation"""
print("\n=== Test 2: Phase 0 - Compound Disambiguation ===")
if tu is None:
record_result("Phase 0 skipped (no TU)", False, "ToolUniverse not loaded")
return None, None
# Test 2a: Name to CID resolution
try:
result = tu.tools.PubChem_get_CID_by_compound_name(name="Acetaminophen")
if isinstance(result, dict) and "data" in result:
cid_data = result["data"]
if isinstance(cid_data, dict) and "IdentifierList" in cid_data:
cids = cid_data["IdentifierList"].get("CID", [])
if len(cids) > 0 and cids[0] == 1983:
record_result("PubChem name->CID (Acetaminophen=1983)", True)
else:
record_result("PubChem name->CID", True,
f"CID returned: {cids[0] if cids else 'none'} (expected 1983)")
else:
record_result("PubChem name->CID", False,
f"Unexpected data structure: {list(cid_data.keys()) if isinstance(cid_data, dict) else type(cid_data)}")
else:
# Try direct access pattern
if isinstance(result, dict):
record_result("PubChem name->CID", True,
f"Non-standard response, keys: {list(result.keys())[:5]}")
else:
record_result("PubChem name->CID", False,
f"Unexpected response type: {type(result)}")
except Exception as e:
record_result("PubChem name->CID", False, f"Exception: {e}")
# Test 2b: CID to properties
smiles = None
try:
result = tu.tools.PubChem_get_compound_properties_by_CID(cid=1983)
if isinstance(result, dict):
# Try to extract SMILES from response
data = result.get("data", result)
if isinstance(data, dict):
props = data.get("PropertyTable", {}).get("Properties", [{}])
if props and isinstance(props, list) and len(props) > 0:
smiles = props[0].get("CanonicalSMILES") or props[0].get("IsomericSMILES")
if smiles:
record_result("PubChem CID->properties (SMILES extracted)", True,
f"SMILES: {smiles}")
else:
record_result("PubChem CID->properties", True,
f"Properties found but no SMILES field. Keys: {list(props[0].keys())[:8]}")
else:
record_result("PubChem CID->properties", True,
f"Data returned, structure: {list(data.keys())[:5]}")
else:
record_result("PubChem CID->properties", True,
f"Data type: {type(data)}")
else:
record_result("PubChem CID->properties", False,
f"Unexpected response type: {type(result)}")
except Exception as e:
record_result("PubChem CID->properties", False, f"Exception: {e}")
# Fallback SMILES for Acetaminophen if extraction failed
if not smiles:
smiles = "CC(=O)Nc1ccc(O)cc1"
print(f" [INFO] Using fallback SMILES: {smiles}")
return 1983, smiles
def test_phase1_toxicity(tu, smiles):
"""Test 3: Phase 1 - Predictive toxicity (ADMET-AI)"""
print("\n=== Test 3: Phase 1 - Predictive Toxicity ===")
if tu is None or smiles is None:
record_result("Phase 1 skipped", False, "Prerequisites not met")
return
# Test 3a: Toxicity prediction
try:
result = tu.tools.ADMETAI_predict_toxicity(smiles=[smiles])
if isinstance(result, dict):
data = result.get("data", result)
record_result("ADMETAI_predict_toxicity", True,
f"Response keys: {list(data.keys())[:8] if isinstance(data, dict) else type(data)}")
else:
record_result("ADMETAI_predict_toxicity", True,
f"Response type: {type(result)}")
except Exception as e:
record_result("ADMETAI_predict_toxicity", False, f"Exception: {e}")
# Test 3b: Stress response prediction
try:
result = tu.tools.ADMETAI_predict_stress_response(smiles=[smiles])
if isinstance(result, dict):
record_result("ADMETAI_predict_stress_response", True)
else:
record_result("ADMETAI_predict_stress_response", True,
f"Type: {type(result)}")
except Exception as e:
record_result("ADMETAI_predict_stress_response", False, f"Exception: {e}")
# Test 3c: Nuclear receptor activity
try:
result = tu.tools.ADMETAI_predict_nuclear_receptor_activity(smiles=[smiles])
if isinstance(result, dict):
record_result("ADMETAI_predict_nuclear_receptor_activity", True)
else:
record_result("ADMETAI_predict_nuclear_receptor_activity", True,
f"Type: {type(result)}")
except Exception as e:
record_result("ADMETAI_predict_nuclear_receptor_activity", False, f"Exception: {e}")
def test_phase2_admet(tu, smiles):
"""Test 4: Phase 2 - ADMET Properties"""
print("\n=== Test 4: Phase 2 - ADMET Properties ===")
if tu is None or smiles is None:
record_result("Phase 2 skipped", False, "Prerequisites not met")
return
admet_tools = [
("ADMETAI_predict_BBB_penetrance", "BBB"),
("ADMETAI_predict_bioavailability", "Bioavailability"),
("ADMETAI_predict_clearance_distribution", "Clearance/Distribution"),
("ADMETAI_predict_CYP_interactions", "CYP Interactions"),
("ADMETAI_predict_physicochemical_properties", "Physicochemical"),
("ADMETAI_predict_solubility_lipophilicity_hydration", "Solubility/Lipophilicity"),
]
for tool_name, label in admet_tools:
try:
tool_fn = getattr(tu.tools, tool_name)
result = tool_fn(smiles=[smiles])
if result is not None:
record_result(f"ADMET: {label}", True)
else:
record_result(f"ADMET: {label}", False, "None returned")
except Exception as e:
record_result(f"ADMET: {label}", False, f"Exception: {e}")
def test_phase3_toxicogenomics(tu):
"""Test 5: Phase 3 - CTD Toxicogenomics"""
print("\n=== Test 5: Phase 3 - CTD Toxicogenomics ===")
if tu is None:
record_result("Phase 3 skipped", False, "Prerequisites not met")
return
# Test 5a: Chemical-gene interactions
try:
result = tu.tools.CTD_get_chemical_gene_interactions(input_terms="Acetaminophen")
if isinstance(result, (list, dict)):
if isinstance(result, list):
count = len(result)
elif isinstance(result, dict):
data = result.get("data", [])
count = len(data) if isinstance(data, list) else 1
else:
count = 0
record_result("CTD chemical-gene interactions", True,
f"Results: {count} interactions")
else:
record_result("CTD chemical-gene interactions", True,
f"Response type: {type(result)}")
except Exception as e:
record_result("CTD chemical-gene interactions", False, f"Exception: {e}")
# Test 5b: Chemical-disease associations
try:
result = tu.tools.CTD_get_chemical_diseases(input_terms="Acetaminophen")
if isinstance(result, (list, dict)):
if isinstance(result, list):
count = len(result)
elif isinstance(result, dict):
data = result.get("data", [])
count = len(data) if isinstance(data, list) else 1
else:
count = 0
record_result("CTD chemical-disease associations", True,
f"Results: {count} associations")
else:
record_result("CTD chemical-disease associations", True,
f"Response type: {type(result)}")
except Exception as e:
record_result("CTD chemical-disease associations", False, f"Exception: {e}")
def test_phase4_fda(tu):
"""Test 6: Phase 4 - FDA Regulatory Safety"""
print("\n=== Test 6: Phase 4 - FDA Regulatory Safety ===")
if tu is None:
record_result("Phase 4 skipped", False, "Prerequisites not met")
return
fda_tools = [
("FDA_get_boxed_warning_info_by_drug_name", "Boxed Warning"),
("FDA_get_contraindications_by_drug_name", "Contraindications"),
("FDA_get_adverse_reactions_by_drug_name", "Adverse Reactions"),
("FDA_get_warnings_by_drug_name", "Warnings"),
("FDA_get_nonclinical_toxicology_info_by_drug_name", "Nonclinical Toxicology"),
]
for tool_name, label in fda_tools:
try:
tool_fn = getattr(tu.tools, tool_name)
result = tool_fn(drug_name="Acetaminophen")
if result is not None:
record_result(f"FDA: {label}", True)
else:
record_result(f"FDA: {label}", False, "None returned")
except Exception as e:
record_result(f"FDA: {label}", False, f"Exception: {e}")
def test_phase5_drugbank(tu):
"""Test 7: Phase 5 - DrugBank Safety"""
print("\n=== Test 7: Phase 5 - DrugBank Safety ===")
if tu is None:
record_result("Phase 5 skipped", False, "Prerequisites not met")
return
try:
result = tu.tools.drugbank_get_safety_by_drug_name_or_drugbank_id(
query="Acetaminophen",
case_sensitive=False,
exact_match=False,
limit=5
)
if result is not None:
record_result("DrugBank safety profile", True,
f"Response type: {type(result).__name__}")
else:
record_result("DrugBank safety profile", False, "None returned")
except Exception as e:
record_result("DrugBank safety profile", False, f"Exception: {e}")
def test_phase6_stitch(tu):
"""Test 8: Phase 6 - STITCH Chemical-Protein Interactions"""
print("\n=== Test 8: Phase 6 - STITCH Interactions ===")
if tu is None:
record_result("Phase 6 skipped", False, "Prerequisites not met")
return
# Test 8a: Resolve identifier
try:
result = tu.tools.STITCH_resolve_identifier(
identifier="acetaminophen", species=9606
)
if result is not None:
record_result("STITCH resolve identifier", True)
else:
record_result("STITCH resolve identifier", False, "None returned")
except Exception as e:
record_result("STITCH resolve identifier", False, f"Exception: {e}")
# Test 8b: Chemical-protein interactions
try:
result = tu.tools.STITCH_get_chemical_protein_interactions(
identifiers=["CIDm01983"], # STITCH format for PubChem CID
species=9606,
required_score=400
)
if result is not None:
record_result("STITCH chemical-protein interactions", True)
else:
record_result("STITCH chemical-protein interactions", False, "None returned")
except Exception as e:
record_result("STITCH chemical-protein interactions", False, f"Exception: {e}")
def test_phase7_structural_alerts(tu):
"""Test 9: Phase 7 - ChEMBL Structural Alerts"""
print("\n=== Test 9: Phase 7 - Structural Alerts ===")
if tu is None:
record_result("Phase 7 skipped", False, "Prerequisites not met")
return
try:
# CHEMBL112 is Acetaminophen
result = tu.tools.ChEMBL_search_compound_structural_alerts(
molecule_chembl_id="CHEMBL112",
limit=20
)
if result is not None:
record_result("ChEMBL structural alerts", True,
f"Response type: {type(result).__name__}")
else:
record_result("ChEMBL structural alerts", False, "None returned")
except Exception as e:
record_result("ChEMBL structural alerts", False, f"Exception: {e}")
def test_environmental_chemical(tu):
"""Test 10: Environmental chemical (non-drug) workflow"""
print("\n=== Test 10: Environmental Chemical (Bisphenol A) ===")
if tu is None:
record_result("Environmental test skipped", False, "Prerequisites not met")
return
# Test BPA: a key environmental chemical
try:
result = tu.tools.CTD_get_chemical_gene_interactions(input_terms="bisphenol A")
if isinstance(result, (list, dict)):
data = result if isinstance(result, list) else result.get("data", [])
count = len(data) if isinstance(data, list) else 0
record_result("CTD: BPA gene interactions", True,
f"Found {count} gene interactions")
else:
record_result("CTD: BPA gene interactions", True,
f"Response: {type(result)}")
except Exception as e:
record_result("CTD: BPA gene interactions", False, f"Exception: {e}")
try:
result = tu.tools.CTD_get_chemical_diseases(input_terms="bisphenol A")
if isinstance(result, (list, dict)):
data = result if isinstance(result, list) else result.get("data", [])
count = len(data) if isinstance(data, list) else 0
record_result("CTD: BPA disease associations", True,
f"Found {count} disease associations")
else:
record_result("CTD: BPA disease associations", True,
f"Response: {type(result)}")
except Exception as e:
record_result("CTD: BPA disease associations", False, f"Exception: {e}")
def test_batch_smiles(tu):
"""Test 11: Batch SMILES processing"""
print("\n=== Test 11: Batch SMILES Processing ===")
if tu is None:
record_result("Batch test skipped", False, "Prerequisites not met")
return
# Test batch ADMET prediction with multiple compounds
smiles_batch = [
"CC(=O)Nc1ccc(O)cc1", # Acetaminophen
"CC(=O)Oc1ccccc1C(=O)O", # Aspirin
"CC(C)Cc1ccc(cc1)C(C)C(=O)O", # Ibuprofen
]
try:
result = tu.tools.ADMETAI_predict_toxicity(smiles=smiles_batch)
if result is not None:
record_result("Batch toxicity prediction (3 compounds)", True)
else:
record_result("Batch toxicity prediction", False, "None returned")
except Exception as e:
record_result("Batch toxicity prediction", False, f"Exception: {e}")
def generate_report():
"""Generate test report"""
print("\n" + "=" * 60)
print("CHEMICAL SAFETY SKILL - TEST REPORT")
print("=" * 60)
print(f"\nTotal Tests: {total_tests}")
print(f"Passed: {passed_tests}")
print(f"Failed: {failed_tests}")
print(f"Pass Rate: {passed_tests}/{total_tests} ({100*passed_tests/max(total_tests,1):.1f}%)")
print()
if failed_tests > 0:
print("FAILED TESTS:")
for r in test_results:
if r["status"] == "FAIL":
print(f" - {r['name']}: {r['details']}")
print()
print("ALL RESULTS:")
for r in test_results:
print(f" [{r['status']}] {r['name']}")
if r["details"]:
print(f" {r['details']}")
return failed_tests == 0
def main():
print("Chemical Safety & Toxicology Skill - Comprehensive Test Suite")
print("=" * 60)
print(f"Testing against: Acetaminophen (primary), Bisphenol A (environmental)")
print()
start_time = time.time()
# Test 1: Tool loading
tu = test_tool_loading()
# Test 2: Phase 0 - Disambiguation
cid, smiles = test_phase0_disambiguation(tu)
# Test 3: Phase 1 - Toxicity predictions
test_phase1_toxicity(tu, smiles)
# Test 4: Phase 2 - ADMET properties
test_phase2_admet(tu, smiles)
# Test 5: Phase 3 - CTD toxicogenomics
test_phase3_toxicogenomics(tu)
# Test 6: Phase 4 - FDA regulatory safety
test_phase4_fda(tu)
# Test 7: Phase 5 - DrugBank safety
test_phase5_drugbank(tu)
# Test 8: Phase 6 - STITCH interactions
test_phase6_stitch(tu)
# Test 9: Phase 7 - Structural alerts
test_phase7_structural_alerts(tu)
# Test 10: Environmental chemical workflow
test_environmental_chemical(tu)
# Test 11: Batch processing
test_batch_smiles(tu)
elapsed = time.time() - start_time
print(f"\nTotal execution time: {elapsed:.1f} seconds")
# Generate report
all_passed = generate_report()
return 0 if all_passed else 1
if __name__ == "__main__":
sys.exit(main())
Related skills
How it compares
Pick tooluniverse-chemical-safety for multi-database hazard dossiers; use tooluniverse-admet-prediction when you only need a pass/fail ADMET scorecard for drug candidates.
FAQ
What input does tooluniverse-chemical-safety require?
tooluniverse-chemical-safety starts with a compound name or SMILES string. Phase 0 disambiguates the compound to SMILES, PubChem CID, and ChEMBL ID before ADMET-AI and database queries run in subsequent phases.
How many phases does tooluniverse-chemical-safety run?
tooluniverse-chemical-safety executes 8 research phases—disambiguation, predictive toxicology, ADMET profiling, toxicogenomics, FDA regulatory safety, DrugBank profiles, STITCH interactions, and structural alerts—followed by an integrated risk synthesis.
What dependency enables ADMET-AI predictions?
ADMET-AI tools in tooluniverse-chemical-safety require `pip install tooluniverse[ml]`. If unavailable, the skill falls back to CTD toxicogenomics and PubChemTox experimental data in later phases.