
Tooluniverse Clinical Trial Design
- 404 installs
- 1.6k repo stars
- Updated August 4, 2026
- mims-harvard/tooluniverse
tooluniverse-clinical-trial-design is a ToolUniverse agent skill that produces Phase 1/2 clinical trial feasibility reports with enrollment projections and FDA pathway analysis for developers who must validate trial desi
About
tooluniverse-clinical-trial-design is a ToolUniverse agent skill (version 1.0.0, compatible with ToolUniverse 0.5+) that assesses early-phase clinical trial feasibility across six parallel research paths before teams commit protocol and engineering resources. The skill queries OpenTargets, ClinVar, gnomAD, DrugBank, FDA Orange Book, OpenFDA, FAERS, PubMed, and ClinicalTrials.gov to ground endpoint, population, comparator, effect-size, and regulatory decisions in precedent trials and FDA guidance rather than first-principles guessing. A mandatory report-first workflow writes a 14-section `[INDICATION]_trial_feasibility_report.md` with a weighted 0–100 feasibility score, A–D evidence grades, enrollment projections, biomarker strategy, and go/no-go recommendations. Developers in biotech, healthtech, or computational drug discovery reach for the skill when scoping inclusion criteria, trial arms, biomarker-selected populations, comparator selection, or IND submission strategy. The skill integrates with related ToolUniverse skills for drug, disease, target, and pharmacovigilance research.
- Endpoint and arm scoping
- Eligibility criteria drafting
- Feasibility framing
- Agent-assisted protocol planning
- Translational research support
Tooluniverse Clinical Trial Design by the numbers
- 404 all-time installs (skills.sh)
- +6 installs in the week ending Aug 4, 2026 (Skillselion tracking)
- Ranked #804 of 3,282 Productivity & Planning skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/mims-harvard/tooluniverse --skill tooluniverse-clinical-trial-designAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 404 |
|---|---|
| repo stars | ★ 1.6k |
| Last updated | August 4, 2026 |
| Repository | mims-harvard/tooluniverse ↗ |
How do you assess clinical trial feasibility before protocol design?
Scope inclusion criteria, endpoints, arms, and feasibility for clinical trials using agent-guided design checks before committing protocol and engineering resources.
Who is it for?
Developers and computational scientists in biotech or healthtech who must validate Phase 1/2 trial endpoints, populations, and FDA pathways before engineering investment.
Skip if: Teams needing biostatistical power calculations from first principles or post-market pharmacovigilance monitoring without trial design context.
When should I use this skill?
The user asks about clinical trial design, trial feasibility, enrollment projections, endpoint selection, Phase 1/2 planning, or biomarker trial scoping.
What you get
A 14-section `[INDICATION]_trial_feasibility_report.md` with 0–100 feasibility score, evidence grades, enrollment projections, endpoint recommendations, and regulatory pathway analysis.
- 14-section trial feasibility markdown report
- Feasibility scorecard with go/no-go criteria
By the numbers
- Analyzes 6 parallel research dimensions for trial feasibility
- Produces a 14-section feasibility report with 0–100 weighted score
- Version 1.0.0 compatible with ToolUniverse 0.5+
Files
Clinical Trial Design Feasibility Assessment
Systematically assess clinical trial feasibility by analyzing 6 research dimensions. Produces comprehensive feasibility reports with quantitative enrollment projections, endpoint recommendations, and regulatory pathway analysis.
IMPORTANT: Always use English terms in tool calls (drug names, disease names, biomarker names), even if the user writes in another language. Only try original-language terms as a fallback if English returns no results. Respond in the user's language.
Reasoning Before Searching
Trial design starts with the question, not the methods. Answer these four questions before running any tools — they determine everything else:
1. What is the primary endpoint? Is it overall survival (gold standard but slow), PFS (faster but surrogate), ORR (single-arm friendly but not always accepted), or a biomarker (needs validation as surrogate first)? The endpoint determines FDA pathway, statistical design, and duration. 2. Who is the population? Broad unselected vs. biomarker-enriched. Enriched populations have higher response rates, allowing smaller trials — but require a validated companion diagnostic and reduce the eligible patient pool. 3. What is the comparator? Placebo (only if no standard of care exists), active control (requires non-inferiority or superiority framing), or single-arm with historical control (acceptable for rare diseases or breakthrough designations, but FDA scrutiny is high). 4. Is the effect size realistic given the mechanism? A 20% improvement in ORR over SOC requires ~100 patients per arm. A 50% improvement requires ~30. If the mechanism only justifies a 10% improvement, the trial may be underpowered regardless of design. Check precedent effect sizes in similar trials before committing to an endpoint.
These four answers determine sample size, duration, and trial design. Look them up from precedent trials and FDA guidance — do not derive them from first principles.
LOOK UP DON'T GUESS: Never assume what the standard of care is for an indication — look it up with DrugBank and FDA tools. Never assume an endpoint is FDA-accepted — verify with search_clinical_trials precedents and OpenFDA_get_approval_history. Never estimate prevalence from memory — use OpenTargets, gnomAD, or COSMIC.
Core Principles
1. Report-First Approach (MANDATORY)
DO NOT show tool outputs to user. Instead: 1. Create [INDICATION]_trial_feasibility_report.md FIRST 2. Initialize with all section headers 3. Progressively update as data arrives 4. Present only the final report
2. Evidence Grading System
| Grade | Symbol | Criteria | Examples |
|---|---|---|---|
| A | 3-star | Regulatory acceptance, multiple precedents | FDA-approved endpoint in same indication |
| B | 2-star | Clinical validation, single precedent | Phase 3 trial in related indication |
| C | 1-star | Preclinical or exploratory | Phase 1 use, biomarker validation ongoing |
| D | 0-star | Proposed, no validation | Novel endpoint, no precedent |
3. Feasibility Score (0-100)
Weighted composite score:
- Patient Availability (30%): Population size x biomarker prevalence x geography
- Endpoint Precedent (25%): Historical use, regulatory acceptance
- Regulatory Clarity (20%): Pathway defined, precedents exist
- Comparator Feasibility (15%): Standard of care availability
- Safety Monitoring (10%): Known risks, monitoring established
Interpretation: >=75 HIGH (proceed), 50-74 MODERATE (additional validation), <50 LOW (de-risking required)
---
When to Use This Skill
Apply when users:
- Plan early-phase trials (Phase 1/2 emphasis)
- Need enrollment feasibility assessment
- Design biomarker-selected trials
- Evaluate endpoint strategies
- Assess regulatory pathways
- Compare trial design options
- Need safety monitoring plans
Trigger phrases: "clinical trial design", "trial feasibility", "enrollment projections", "endpoint selection", "trial planning", "Phase 1/2 design", "basket trial", "biomarker trial"
---
Core Strategy: 6 Research Paths
Execute 6 parallel research dimensions. See STUDY_DESIGN_PROCEDURES.md for detailed steps per path.
Trial Design Query
|
+-- PATH 1: Patient Population Sizing
| Disease prevalence, biomarker prevalence, geographic distribution,
| eligibility criteria impact, enrollment projections
|
+-- PATH 2: Biomarker Prevalence & Testing
| Mutation frequency, testing availability, turnaround time,
| cost/reimbursement, alternative biomarkers
|
+-- PATH 3: Comparator Selection
| Standard of care, approved comparators, historical controls,
| placebo appropriateness, combination therapy
|
+-- PATH 4: Endpoint Selection
| Primary endpoint precedents, FDA acceptance history,
| measurement feasibility, surrogate vs clinical endpoints
|
+-- PATH 5: Safety Endpoints & Monitoring
| Mechanism-based toxicity, class effects, organ-specific monitoring,
| DLT history, safety monitoring plan
|
+-- PATH 6: Regulatory Pathway
Regulatory precedents (505(b)(1), 505(b)(2)), breakthrough therapy,
orphan drug, fast track, FDA guidance---
Report Structure (14 Sections)
Create [INDICATION]_trial_feasibility_report.md with all 14 sections. See REPORT_TEMPLATE.md for full templates with fillable fields.
1. Executive Summary - Feasibility score, key findings, go/no-go recommendation 2. Disease Background - Prevalence, incidence, SOC, unmet need 3. Patient Population Analysis - Base population, biomarker selection, eligibility funnel, enrollment projections 4. Biomarker Strategy - Primary biomarker, alternatives, testing logistics 5. Endpoint Selection & Justification - Primary/secondary/exploratory endpoints, statistical considerations 6. Comparator Analysis - SOC, trial design options (single-arm vs randomized vs non-inferiority), drug sourcing 7. Safety Endpoints & Monitoring Plan - DLT definition, mechanism-based toxicities, organ monitoring, SMC 8. Study Design Recommendations - Phase, design type, schema, eligibility, treatment plan, assessment schedule 9. Enrollment & Site Strategy - Site selection, enrollment projections, recruitment strategies 10. Regulatory Pathway - FDA pathway, precedents, pre-IND meeting, IND timeline 11. Budget & Resource Considerations - Cost drivers, timeline, FTE requirements 12. Risk Assessment - Feasibility risks, scientific risks, mitigation strategies 13. Success Criteria & Go/No-Go Decision - Phase 1/2 criteria, interim analysis, feasibility scorecard 14. Recommendations & Next Steps - Final recommendation, critical path to IND, alternative designs
---
Tool Reference by Research Path
PATH 1: Patient Population Sizing
OpenTargets_get_disease_id_description_by_name- Disease lookupOpenTargets_get_diseases_phenotypes_by_target_ensembl- Prevalence dataClinVar_search_variants- Biomarker mutation frequencygnomad_search_variants- Population allele frequenciesPubMed_search_articles- Epidemiology literaturesearch_clinical_trials- Enrollment feasibility from past trials
PATH 2: Biomarker Prevalence & Testing
ClinVar_get_variant_details- Variant pathogenicityCOSMIC_search_mutations- Cancer-specific mutation frequenciesgnomad_get_variant- Population geneticsPubMed_search_articles- CDx test performance, guidelines
PATH 3: Comparator Selection
drugbank_get_drug_basic_info_by_drug_name_or_id- Drug infodrugbank_get_indications_by_drug_name_or_drugbank_id- Approved indicationsdrugbank_get_pharmacology_by_drug_name_or_drugbank_id- MechanismFDA_OrangeBook_search_drug- Generic availabilityOpenFDA_get_approval_history- Approval detailssearch_clinical_trials- Historical control data
PATH 4: Endpoint Selection
search_clinical_trials- Precedent trials, endpoints usedPubMed_search_articles- FDA acceptance history, endpoint validationOpenFDA_get_approval_history- Approved endpoints by indication
PATH 5: Safety Endpoints & Monitoring
drugbank_get_pharmacology_by_drug_name_or_drugbank_id- Mechanism toxicityFDA_get_warnings_and_cautions_by_drug_name- FDA black box warningsFAERS_search_reports_by_drug_and_reaction- Real-world adverse eventsFAERS_count_reactions_by_drug_event- AE frequencyFAERS_count_death_related_by_drug- Serious outcomesPubMed_search_articles- DLT definitions, monitoring strategies
PATH 6: Regulatory Pathway
OpenFDA_get_approval_history- Precedent approvalsPubMed_search_articles- Breakthrough designations, FDA guidancesearch_clinical_trials- Regulatory precedents (accelerated approval)
---
Quick Start Example
from tooluniverse import ToolUniverse
tu = ToolUniverse(use_cache=True)
tu.load_tools()
# Example: EGFR+ NSCLC trial feasibility
# Step 1: Disease prevalence
disease_info = tu.tools.OpenTargets_get_disease_id_description_by_name(
diseaseName="non-small cell lung cancer"
)
prevalence = tu.tools.OpenTargets_get_diseases_phenotypes(
efoId=disease_info['data']['id']
)
# Step 2: Biomarker prevalence
variants = tu.tools.ClinVar_search_variants(gene="EGFR", significance="pathogenic")
# Step 3: Precedent trials
trials = tu.tools.search_clinical_trials(
condition="EGFR positive non-small cell lung cancer",
status="completed", phase="2"
)
# Step 4: Standard of care comparator
soc = tu.tools.FDA_OrangeBook_search_drug(ingredient="osimertinib")
# Compile into feasibility report...See WORKFLOW_DETAILS.md for the complete 6-path Python workflow and use case examples.
---
Integration with Other Skills
- tooluniverse-drug-research: Investigate mechanism, preclinical data
- tooluniverse-disease-research: Deep dive on disease biology
- tooluniverse-target-research: Validate drug target, essentiality
- tooluniverse-pharmacovigilance: Post-market safety for comparator drugs
- tooluniverse-precision-oncology: Biomarker biology, resistance mechanisms
---
Programmatic Access (Beyond Tools)
When ToolUniverse tools return limited trial metadata, use the ClinicalTrials.gov v2 API directly:
import requests, pandas as pd
# Search with pagination (all lung cancer immunotherapy trials with results)
all_studies = []
token = None
while True:
params = {"query.cond": "lung cancer", "query.intr": "immunotherapy",
"filter.overallStatus": "COMPLETED", "filter.results": "WITH_RESULTS", "pageSize": 100}
if token: params["pageToken"] = token
resp = requests.get("https://clinicaltrials.gov/api/v2/studies", params=params).json()
all_studies.extend(resp.get("studies", []))
token = resp.get("nextPageToken")
if not token: break
# Extract structured data
rows = []
for s in all_studies:
proto = s.get("protocolSection", {})
rows.append({
"nctId": proto.get("identificationModule", {}).get("nctId"),
"title": proto.get("identificationModule", {}).get("briefTitle"),
"enrollment": proto.get("designModule", {}).get("enrollmentInfo", {}).get("count"),
"phase": proto.get("designModule", {}).get("phases", [None])[0] if proto.get("designModule", {}).get("phases") else None,
})
df = pd.DataFrame(rows)
# FDA drug approval history
drug = "pembrolizumab"
fda = requests.get(f"https://api.fda.gov/drug/drugsfda.json?search=openfda.brand_name:{drug}&limit=10").json()See tooluniverse-data-wrangling skill for pagination, error handling, and bulk download patterns.
---
Reference Files
| File | Content |
|---|---|
REPORT_TEMPLATE.md | Full 14-section report template with fillable fields |
STUDY_DESIGN_PROCEDURES.md | Detailed steps for each of the 6 research paths |
WORKFLOW_DETAILS.md | Complete Python example workflow and 5 use case summaries |
BEST_PRACTICES.md | Best practices, common pitfalls, output format requirements |
EXAMPLES.md | Additional examples |
QUICK_START.md | Quick start guide |
---
Version Information
- Version: 1.0.0
- Last Updated: February 2026
- Compatible with: ToolUniverse 0.5+
- Focus: Phase 1/2 early clinical development
# API Keys for ToolUniverse
# Copy this file to .env and fill in your actual API keys
BIOGRID_API_KEY=your_api_key_here
BOLTZ_MCP_SERVER_HOST=your_api_key_here
BRENDA_EMAIL=your_api_key_here
BRENDA_PASSWORD=your_api_key_here
DISGENET_API_KEY=your_api_key_here
EXPERT_FEEDBACK_MCP_SERVER_URL=your_api_key_here
NVIDIA_API_KEY=your_api_key_here
OMIM_API_KEY=your_api_key_here
TXAGENT_MCP_SERVER_HOST=your_api_key_here
USPTO_API_KEY=your_api_key_here
USPTO_MCP_SERVER_HOST=your_api_key_here
Best Practices & Output Requirements
Best Practices
1. Start with Report Template
Create full report structure FIRST, then populate:
# Clinical Trial Feasibility Report: [INDICATION]
## 1. Executive Summary
[Researching...]
## 2. Disease Background
[Researching...]
[...all 14 sections...]2. Use English for All Tool Calls
Even if user asks in another language:
- "EGFR+ NSCLC" not local-language equivalents
- "breast cancer" not translations
- Translate results back to user's language
3. Validate Biomarker Prevalence Across Sources
Cross-check ClinVar, gnomAD, COSMIC, and literature:
- ClinVar: Clinical significance
- gnomAD: Population frequency (for germline)
- COSMIC: Somatic mutation frequency in cancers
- Literature: Geographic/ethnic variation
4. Calculate Enrollment Funnel Explicitly
Show math for patient availability:
US NSCLC incidence: 200,000/year
x EGFR+ prevalence: 15% = 30,000
x L858R within EGFR+: 45% = 13,500
x Eligible (age, PS, prior Tx): 60% = 8,100
/ Competing trials: 3 = 2,700 available/year
For N=43, need 43/2,700 = 1.6% capture rate -> Achievable5. Evidence Grade Every Key Claim
EGFR L858R prevalence is 45% of EGFR+ NSCLC [A: PMID:12345, large
sequencing study n=1,500]. *Source: ClinVar, COSMIC*6. Provide Regulatory Precedent Details
Not just "ORR is accepted" but:
ORR is FDA-accepted for accelerated approval in NSCLC [A: FDA approvals]:
- Osimertinib (2015): ORR 57%, n=411, Tx-resistant EGFR+ (NCT01802632)
- Dacomitinib (2018): ORR 45%, n=452, 1L EGFR+ (NCT01774721)
- [3 more examples]7. Address Feasibility Risks Proactively
For each HIGH risk, provide mitigation:
Risk: Biomarker screen failure rate >70%
-> Mitigation: Liquid biopsy pre-screening (ctDNA EGFR, 7-day turnaround)8. Separate Phase 1 and Phase 2 Components
If combined Phase 1/2:
- Phase 1: Safety, DLT, RP2D (N=12-18, 3+3 or BOIN)
- Phase 2: Efficacy, ORR (N=43, Simon 2-stage)
- Distinct success criteria for each phase
---
Common Pitfalls to Avoid
Don't: Show Tool Outputs to User
# BAD
OpenTargets returned:
{
"data": {
"id": "EFO_0003060",
"name": "non-small cell lung carcinoma"
}
}Do: Present Synthesized Report
# GOOD
## Disease Background
Non-small cell lung cancer (NSCLC) represents 85% of lung cancers, with
~200,000 new cases annually in the US [A: CDC WONDER]. EGFR mutations
occur in 15% of Caucasian and 50% of Asian patients [A: PMID:23816960].
*Source: OpenTargets, ClinVar*Don't: Make Unsupported Claims
# BAD
ORR of 60% is expected based on preclinical data.Do: Ground in Evidence
# GOOD
ORR of 30-40% is projected [B] based on:
- Similar EGFR TKI (erlotinib): 32% ORR in EGFR+ NSCLC (NCT00949650)
- Our drug's 2x IC50 potency vs. erlotinib (preclinical)
*Source: ClinicalTrials.gov, internal data*Don't: Ignore Geographic Variation
# BAD
EGFR L858R prevalence: 7% of NSCLCDo: Specify Geography
# GOOD
EGFR L858R prevalence [A: COSMIC, ClinVar]:
- Caucasian (US/EU): 6-7% of NSCLC
- East Asian: 20-25% of NSCLC
-> Trial site strategy: Include Asian sites for 2x enrollment---
Output Format Requirements
Report File Naming
[INDICATION]_trial_feasibility_report.md- Example:
EGFR_L858R_NSCLC_trial_feasibility_report.md
Section Completeness
All 14 sections MUST be present (see REPORT_TEMPLATE.md for details): 1. Executive Summary 2. Disease Background 3. Patient Population Analysis (with funnel) 4. Biomarker Strategy 5. Endpoint Selection & Justification 6. Comparator Analysis 7. Safety Endpoints & Monitoring Plan 8. Study Design Recommendations 9. Enrollment & Site Strategy 10. Regulatory Pathway 11. Budget & Resource Considerations 12. Risk Assessment 13. Success Criteria & Go/No-Go Decision (with scorecard) 14. Recommendations & Next Steps
Evidence Grading Required In
- Section 1 (Executive Summary): Key findings
- Section 4 (Biomarker): Prevalence claims
- Section 5 (Endpoints): Regulatory precedents
- Section 6 (Comparator): SOC efficacy data
- Section 7 (Safety): Toxicity frequencies
- Section 10 (Regulatory): Approval precedents
- Section 13 (Scorecard): All dimensions
Feasibility Score Transparency
Show calculation:
| Dimension | Weight | Raw Score | Weighted | Evidence |
|-----------|--------|-----------|----------|----------|
| Patient Availability | 30% | 8/10 | 24 | A: Epi data |
| Endpoint Precedent | 25% | 9/10 | 22.5 | A: FDA approvals |
| Regulatory Clarity | 20% | 7/10 | 14 | B: Pre-IND advised |
| Comparator Feasibility | 15% | 9/10 | 13.5 | A: Generic avail |
| Safety Monitoring | 10% | 8/10 | 8 | B: Class effects |
| **TOTAL** | **100%** | - | **82/100** | **HIGH** |Clinical Trial Design Feasibility Examples
Concrete examples of trial feasibility assessments using ToolUniverse.
---
Example 1: Biomarker-Selected Oncology Trial (EGFR+ NSCLC)
Scenario: Assess feasibility of Phase 1/2 trial for novel EGFR inhibitor in EGFR L858R+ NSCLC patients who progressed on osimertinib.
Setup
from tooluniverse import ToolUniverse
tu = ToolUniverse(use_cache=True)
tu.load_tools()
# Trial parameters
indication = "EGFR L858R+ non-small cell lung cancer, osimertinib-resistant"
phase = "Phase 1/2"
primary_endpoint = "Objective Response Rate (ORR)"
biomarker = "EGFR L858R"Step 1: Patient Population Sizing
# 1.1: Get disease prevalence
disease_info = tu.tools.OpenTargets_get_disease_id_description_by_name(
diseaseName="non-small cell lung cancer"
)
print(f"Disease: {disease_info['data']['name']}")
print(f"EFO ID: {disease_info['data']['id']}")
# Get phenotype/prevalence data
phenotypes = tu.tools.OpenTargets_get_diseases_phenotypes(
efoId=disease_info['data']['id']
)
# 1.2: Get biomarker prevalence from ClinVar
egfr_variants = tu.tools.ClinVar_search_variants(
gene="EGFR",
significance="pathogenic,likely_pathogenic"
)
# Filter to L858R
l858r_variants = [v for v in egfr_variants['data']
if 'L858R' in v.get('name', '')]
print(f"\nEGFR L858R variants found: {len(l858r_variants)}")
# 1.3: Cross-reference with population genetics
gnomad_egfr = tu.tools.gnomad_search_variants(
gene="EGFR"
)
# Filter to L858R (c.2573T>G)
l858r_gnomad = [v for v in gnomad_egfr['data']
if v.get('hgvs_c', '').startswith('c.2573T>G')]
# 1.4: Search literature for epidemiology
epi_papers = tu.tools.PubMed_search_articles(
query="EGFR L858R prevalence NSCLC epidemiology United States Asia",
max_results=30
)
print(f"\nEpidemiology papers found: {len(epi_papers['data'])}")
# Extract key papers
key_epi_papers = []
for paper in epi_papers['data'][:5]:
key_epi_papers.append({
'title': paper.get('title'),
'pmid': paper.get('pmid'),
'year': paper.get('pub_year')
})
# 1.5: Calculate patient availability
# From literature: NSCLC ~200K/year US, EGFR+ 15%, L858R 45% of EGFR+
us_nsclc_annual = 200000
egfr_positive_rate = 0.15
l858r_within_egfr = 0.45
l858r_annual = us_nsclc_annual * egfr_positive_rate * l858r_within_egfr
print(f"\nEstimated L858R+ NSCLC patients/year (US): {l858r_annual:.0f}")
# For osimertinib-resistant: assume ~80% progress on osimertinib
osimertinib_resistant = l858r_annual * 0.80
print(f"Osimertinib-resistant L858R+ patients/year: {osimertinib_resistant:.0f}")
# Apply eligibility criteria
eligibility_factors = {
'age_18_75': 0.85,
'ecog_0_1': 0.70,
'adequate_organ': 0.90,
'measurable_disease': 0.75,
'no_brain_mets': 0.60
}
eligible_pool = osimertinib_resistant
for criterion, factor in eligibility_factors.items():
eligible_pool *= factor
print(f" After {criterion}: {eligible_pool:.0f} ({factor*100:.0f}%)")
print(f"\nFinal eligible pool: {eligible_pool:.0f} patients/year")
# Enrollment projection for N=50
target_n = 50
sites = 15
capture_rate = 0.03 # 3% of eligible patients
monthly_enrollment = (eligible_pool * capture_rate) / 12 / sites
print(f"\nTarget enrollment: {target_n}")
print(f"Sites: {sites}")
print(f"Patients per site per month: {monthly_enrollment:.2f}")
print(f"Enrollment timeline: {target_n / (monthly_enrollment * sites):.1f} months")Expected Output:
Disease: non-small cell lung cancer
EFO ID: EFO_0003060
EGFR L858R variants found: 12
Epidemiology papers found: 28
Estimated L858R+ NSCLC patients/year (US): 13500
Osimertinib-resistant L858R+ patients/year: 10800
After age_18_75: 9180 (85%)
After ecog_0_1: 6426 (70%)
After adequate_organ: 5783 (90%)
After measurable_disease: 4338 (75%)
After no_brain_mets: 2603 (60%)
Final eligible pool: 2603 patients/year
Target enrollment: 50
Sites: 15
Patients per site per month: 0.43
Enrollment timeline: 7.7 monthsStep 2: Biomarker Testing Strategy
# 2.1: Search for FDA-approved companion diagnostics
cdx_papers = tu.tools.PubMed_search_articles(
query="FDA approved companion diagnostic EGFR L858R liquid biopsy",
max_results=20
)
print("FDA-approved CDx landscape:")
for paper in cdx_papers['data'][:5]:
print(f" - {paper.get('title')} (PMID: {paper.get('pmid')})")
# 2.2: Literature on testing turnaround time
tat_papers = tu.tools.PubMed_search_articles(
query="EGFR mutation testing turnaround time NGS liquid biopsy",
max_results=15
)
# From literature: NGS turnaround 7-14 days, liquid biopsy 7-10 days
testing_strategy = {
'primary_method': 'NGS (tissue)',
'turnaround': '10-14 days',
'cost': '$500-800',
'alternative': 'Liquid biopsy (ctDNA)',
'alternative_tat': '7-10 days',
'alternative_cost': '$300-500'
}
print(f"\nRecommended biomarker testing:")
for key, value in testing_strategy.items():
print(f" {key}: {value}")Step 3: Comparator Selection
# 3.1: Get standard of care info (osimertinib)
comparator = "osimertinib"
comparator_info = tu.tools.drugbank_get_drug_basic_info_by_drug_name_or_id(
drug_name_or_drugbank_id=comparator
)
comparator_indications = tu.tools.drugbank_get_indications_by_drug_name_or_drugbank_id(
drug_name_or_drugbank_id=comparator
)
print(f"\nComparator: {comparator}")
print(f"Status: {comparator_info['data']['groups']}")
print(f"Indications: {len(comparator_indications['data'])}")
# 3.2: Get FDA approval details
fda_approval = tu.tools.OpenFDA_get_approval_history(
drug_name=comparator
)
if fda_approval and 'data' in fda_approval:
for approval in fda_approval['data'][:3]:
print(f"\n Approval: {approval.get('date', 'N/A')}")
print(f" Indication: {approval.get('indication', 'N/A')}")
# 3.3: Find historical control data from clinical trials
historical_trials = tu.tools.search_clinical_trials(
condition="EGFR positive non-small cell lung cancer",
intervention=comparator,
status="completed",
phase="2|3"
)
print(f"\n{comparator} trials found: {len(historical_trials['data'])}")
# Extract ORR data from key trials
for trial in historical_trials['data'][:3]:
print(f"\n NCT: {trial.get('nct_number')}")
print(f" Title: {trial.get('title')}")
print(f" Status: {trial.get('status')}")
# Note: Would parse results for ORR in real analysis
# 3.4: Single-arm vs. randomized design decision
print("\n" + "="*80)
print("COMPARATOR ANALYSIS SUMMARY")
print("="*80)
print(f"Standard of Care: {comparator} (FDA approved)")
print(f"Historical ORR: ~50-60% in osimertinib-naive, ~30% in T790M-resistant")
print(f"Comparator availability: Commercial supply available")
print(f"\nDesign recommendation: SINGLE-ARM PHASE 2")
print(f"Rationale:")
print(f" - Robust historical control data available (multiple trials)")
print(f" - Patient population narrowly defined (L858R, osimertinib-resistant)")
print(f" - Faster enrollment (no randomization)")
print(f" - Lower cost")
print(f" - Acceptable for Phase 2; randomized Phase 3 if successful")Step 4: Endpoint Selection
# 4.1: Search for precedent trials using ORR
orr_precedent = tu.tools.search_clinical_trials(
condition="EGFR positive non-small cell lung cancer",
phase="2",
status="completed"
)
orr_trials_count = 0
pfs_trials_count = 0
for trial in orr_precedent['data']:
primary_outcome = trial.get('primary_outcome', '').lower()
if 'response rate' in primary_outcome or 'orr' in primary_outcome:
orr_trials_count += 1
if 'progression' in primary_outcome or 'pfs' in primary_outcome:
pfs_trials_count += 1
print("PRIMARY ENDPOINT ANALYSIS")
print("="*80)
print(f"Phase 2 trials in EGFR+ NSCLC: {len(orr_precedent['data'])}")
print(f" - Using ORR as primary: {orr_trials_count}")
print(f" - Using PFS as primary: {pfs_trials_count}")
# 4.2: FDA approval precedents using ORR
orr_approval_papers = tu.tools.PubMed_search_articles(
query="FDA approval objective response rate NSCLC EGFR inhibitor accelerated",
max_results=25
)
print(f"\nFDA approval papers with ORR: {len(orr_approval_papers['data'])}")
# Sample key approvals (from literature/knowledge):
fda_orr_approvals = [
{'drug': 'osimertinib', 'year': 2015, 'orr': '57%', 'n': 411, 'indication': 'T790M+'},
{'drug': 'dacomitinib', 'year': 2018, 'orr': '75%', 'n': 452, 'indication': '1L EGFR+'},
{'drug': 'mobocertinib', 'year': 2021, 'orr': '28%', 'n': 114, 'indication': 'exon20ins'}
]
print("\nFDA Approvals Using ORR (EGFR+ NSCLC):")
for approval in fda_orr_approvals:
print(f" {approval['drug']} ({approval['year']}): ORR {approval['orr']}, n={approval['n']}")
# 4.3: Statistical design for ORR
print("\n" + "="*80)
print("RECOMMENDED PRIMARY ENDPOINT: Objective Response Rate (ORR)")
print("="*80)
print("Evidence Grade: ★★★ (Regulatory precedent, multiple approvals)")
print("\nJustification:")
print(" 1. FDA-accepted for accelerated approval in EGFR+ NSCLC")
print(" 2. Feasible in Phase 2 (rapid readout, smaller N)")
print(" 3. Clinically meaningful (patient benefit)")
print(" 4. Standard assessment (RECIST 1.1, CT imaging)")
print("\nStatistical Design (Simon 2-stage):")
print(" - Null hypothesis (H0): ORR ≤ 15% (below clinically meaningful)")
print(" - Target ORR (H1): ORR ≥ 35% (clinically significant improvement)")
print(" - Alpha: 0.05 (one-sided), Beta: 0.20 (80% power)")
print(" - Stage 1: Enroll 13 patients")
print(" → If ≥ 2 responses, proceed to Stage 2")
print(" → If < 2 responses, stop for futility")
print(" - Stage 2: Enroll 30 additional patients (N=43 total)")
print(" → Declare success if ≥ 11 responses overall (ORR ≥ 25.6%)")
print("\nSecondary Endpoints:")
print(" - Duration of Response (DoR)")
print(" - Progression-Free Survival (PFS)")
print(" - Overall Survival (OS, with long-term follow-up)")
print(" - Safety (AEs per CTCAE v5.0)")
print("\nExploratory Endpoints:")
print(" - ctDNA clearance (liquid biopsy)")
print(" - Biomarkers of resistance (T790M acquisition, C797S)")
print(" - Quality of life (EORTC QLQ-C30)")Step 5: Safety Monitoring
# 5.1: Get class effect toxicities from similar EGFR inhibitors
reference_drug = "erlotinib" # Earlier-generation EGFR TKI
reference_pharmacology = tu.tools.drugbank_get_pharmacology_by_drug_name_or_drugbank_id(
drug_name_or_drugbank_id=reference_drug
)
reference_warnings = tu.tools.FDA_get_warnings_and_cautions_by_drug_name(
drug_name=reference_drug
)
print("SAFETY MONITORING PLAN")
print("="*80)
print(f"Reference drug for class effects: {reference_drug}")
print(f"FDA warnings: {len(reference_warnings.get('data', []))}")
# 5.2: FAERS data for real-world AEs
faers_egfr = tu.tools.FAERS_search_reports_by_drug_and_reaction(
drug_name=reference_drug,
limit=1000
)
ae_counts = tu.tools.FAERS_count_reactions_by_drug_event(
medicinalproduct=reference_drug.upper()
)
print(f"\nFAERS reports for {reference_drug}: {len(faers_egfr.get('data', []))}")
print("\nTop 10 Adverse Events (FAERS):")
for i, ae in enumerate(ae_counts.get('results', [])[:10], 1):
print(f" {i}. {ae['term']}: {ae['count']} reports")
# 5.3: Define mechanism-based toxicity monitoring
print("\n" + "="*80)
print("MECHANISM-BASED TOXICITY MONITORING")
print("="*80)
toxicity_monitoring = {
'Dermatologic': {
'toxicities': ['Rash (acneiform)', 'Dry skin', 'Paronychia'],
'incidence': '60-80% (any grade), 10-15% (Grade 3+)',
'monitoring': 'Dermatology assessment q cycle, patient diary',
'management': 'Topical steroids, doxycycline, dose modification'
},
'Gastrointestinal': {
'toxicities': ['Diarrhea', 'Nausea', 'Decreased appetite'],
'incidence': '50-60% (any grade), 5-10% (Grade 3+)',
'monitoring': 'Symptom diary, electrolytes if Grade 2+',
'management': 'Loperamide, hydration, dose hold if Grade 3+'
},
'Hepatic': {
'toxicities': ['ALT/AST elevation', 'Hyperbilirubinemia'],
'incidence': '20-30% (any grade), 3-5% (Grade 3+)',
'monitoring': 'LFTs weekly (Cycle 1), then q3 weeks',
'management': 'Dose hold if ALT >5× ULN, discontinue if Hy\'s Law'
},
'Pulmonary': {
'toxicities': ['Interstitial lung disease (ILD)', 'Pneumonitis'],
'incidence': '1-3% (any grade), 0.5-1% (Grade 3+)',
'monitoring': 'CT chest at baseline, q12 weeks; symptoms q visit',
'management': 'Discontinue if ILD, systemic steroids'
}
}
for organ, details in toxicity_monitoring.items():
print(f"\n{organ}:")
for key, value in details.items():
if key == 'toxicities':
print(f" {key}: {', '.join(value)}")
else:
print(f" {key}: {value}")
# 5.4: DLT definition (Phase 1 component)
print("\n" + "="*80)
print("DOSE-LIMITING TOXICITY (DLT) DEFINITION - PHASE 1")
print("="*80)
print("Assessment Period: Cycle 1 (28 days)")
print("\nDLTs include:")
print(" - Grade ≥3 non-hematologic toxicity (except manageable with Rx)")
print(" - Grade 4 hematologic toxicity >7 days")
print(" - Any toxicity causing >2 week dose delay in Cycle 1")
print(" - Any Grade 5 (death) related to study drug")
print("\nExceptions (NOT considered DLTs):")
print(" - Grade 3 rash if resolves to ≤Grade 1 within 7 days with treatment")
print(" - Grade 3 diarrhea if resolves within 2 days with loperamide")
print(" - Grade 3 nausea/vomiting if controlled with antiemetics")
print("\nDose Escalation Design: 3+3")
print(" Starting dose: [X] mg QD (10% predicted human efficacious dose)")
print(" Dose levels: [X], [1.5X], [2X], [3X] mg QD")
print(" Rule: ≥2 DLTs at a level → MTD exceeded; previous level is MTD")Step 6: Regulatory Pathway
# 6.1: Search for breakthrough therapy designations
bt_papers = tu.tools.PubMed_search_articles(
query="FDA breakthrough therapy designation NSCLC EGFR inhibitor",
max_results=20
)
print("REGULATORY PATHWAY ANALYSIS")
print("="*80)
print(f"Breakthrough therapy papers: {len(bt_papers['data'])}")
# 6.2: Check orphan drug eligibility
us_prevalence = osimertinib_resistant # From Step 1: ~10,800/year
print(f"\nOrphan Drug Designation Eligibility:")
print(f" Indication: EGFR L858R+ NSCLC, osimertinib-resistant")
print(f" Estimated US patients: {us_prevalence:.0f}/year")
print(f" Orphan threshold: <200,000 total prevalence")
print(f" Assessment: LIKELY NOT ELIGIBLE (too prevalent)")
# 6.3: FDA guidance documents
guidance_papers = tu.tools.PubMed_search_articles(
query="FDA guidance clinical trial endpoints NSCLC oncology",
max_results=15
)
print(f"\nFDA guidance documents: {len(guidance_papers['data'])} papers")
# 6.4: Regulatory recommendations
print("\n" + "="*80)
print("RECOMMENDED REGULATORY STRATEGY")
print("="*80)
print("\nPathway: 505(b)(1) (New Drug Application)")
print(" - Novel molecular entity")
print(" - No published safety data to rely on")
print("\nPotential Designations:")
print(" 1. Breakthrough Therapy [POSSIBLE]")
print(" Criteria: Preliminary evidence of substantial improvement")
print(" Threshold: ORR >50% in osimertinib-resistant (vs. ~10-15% SOC)")
print(" Benefit: Rolling NDA submission, frequent FDA meetings")
print(" Timing: Apply after Phase 1/2a data (n≥20, ORR clear)")
print("\n 2. Fast Track [LIKELY]")
print(" Criteria: Treats serious condition, addresses unmet need")
print(" Benefit: Rolling review, more frequent FDA interaction")
print(" Timing: Apply at IND or early Phase 1")
print("\n 3. Accelerated Approval [TARGET]")
print(" Endpoint: ORR (surrogate for OS)")
print(" Requirement: Confirmatory Phase 3 trial (OS primary)")
print(" Timing: After positive Phase 2 (ORR ≥35%, n=43)")
print("\nRegulatory Milestones:")
print(" Month -4: Pre-IND meeting request")
print(" Month -3: Pre-IND meeting (discuss endpoint, design)")
print(" Month 0: IND submission")
print(" Month 1: First patient dosed (if no clinical hold)")
print(" Month 9: Phase 1 complete, RP2D determined")
print(" Month 16: Phase 2 interim (Simon Stage 1, n=13)")
print(" Month 24: Phase 2 complete (n=43)")
print(" Month 27: End-of-Phase 2 meeting, Phase 3 design discussion")Final Feasibility Report Generation
# Compile all data into feasibility score
print("\n" + "="*80)
print("FEASIBILITY SCORECARD")
print("="*80)
dimensions = {
'Patient Availability': {
'weight': 0.30,
'raw_score': 8, # 8/10
'evidence': '★★★',
'rationale': '~2,600 eligible/year, 7.7-month enrollment for n=50'
},
'Endpoint Precedent': {
'weight': 0.25,
'raw_score': 9, # 9/10
'evidence': '★★★',
'rationale': 'ORR accepted for accelerated approval, 10+ precedents'
},
'Regulatory Clarity': {
'weight': 0.20,
'raw_score': 8, # 8/10
'evidence': '★★☆',
'rationale': 'Clear 505(b)(1) path, breakthrough potential, pre-IND advised'
},
'Comparator Feasibility': {
'weight': 0.15,
'raw_score': 9, # 9/10
'evidence': '★★★',
'rationale': 'Robust historical data (osimertinib ORR 30-60%), single-arm viable'
},
'Safety Monitoring': {
'weight': 0.10,
'raw_score': 8, # 8/10
'evidence': '★★☆',
'rationale': 'EGFR TKI class effects well-characterized, manageable'
}
}
feasibility_score = 0
print(f"{'Dimension':<30} {'Weight':<10} {'Score':<10} {'Weighted':<10} {'Evidence':<10}")
print("-" * 80)
for dimension, data in dimensions.items():
weighted = data['weight'] * data['raw_score'] * 10
feasibility_score += weighted
print(f"{dimension:<30} {data['weight']*100:.0f}%{'':<7} "
f"{data['raw_score']}/10{'':<5} "
f"{weighted:.1f}{'':<7} "
f"{data['evidence']:<10}")
print(f" Rationale: {data['rationale']}")
print("-" * 80)
print(f"{'TOTAL FEASIBILITY SCORE':<30} {'100%':<10} {'':<10} "
f"{feasibility_score:.0f}/100{'':<7} {'HIGH':<10}")
print("\n" + "="*80)
print("FINAL RECOMMENDATION: RECOMMEND PROCEED")
print("="*80)
print("""
This Phase 1/2 trial demonstrates HIGH feasibility (Score: 82/100).
Key Strengths:
1. Patient availability is strong with ~2,600 eligible patients/year
2. ORR is FDA-accepted with robust regulatory precedent
3. Single-arm design is defensible with strong historical control data
4. Safety monitoring is well-established for EGFR TKI class
Critical Path:
1. Pre-IND meeting (Month -3) to confirm single-arm design acceptability
2. Secure CDx partnership for EGFR testing (liquid biopsy preferred)
3. IND submission (Month 0)
4. First patient dosed (Month 1)
5. Phase 2 interim analysis (Month 16, Simon Stage 1)
6. Phase 2 completion (Month 24, n=43)
Key Risk: Screen failure rate may be higher if liquid biopsy false-negative
Mitigation: Tissue re-biopsy for liquid biopsy-negative but clinically suspected
Budget Estimate: $3.5-5.0M (Phase 1/2 combined, 15 sites)
Timeline: 24 months (first patient to primary analysis)
""")---
Example 2: Rare Disease Trial (Niemann-Pick Type C)
Scenario: Assess feasibility of Phase 2 trial for novel cholesterol transport modifier in Niemann-Pick Type C (NPC), a lysosomal storage disorder with prevalence ~1:120,000.
Setup
from tooluniverse import ToolUniverse
tu = ToolUniverse(use_cache=True)
tu.load_tools()
indication = "Niemann-Pick Type C"
phase = "Phase 2"
primary_endpoint = "Change in NPC Clinical Severity Score"Step 1: Ultra-Rare Disease Population Sizing
# 1.1: Search OpenTargets
disease_info = tu.tools.OpenTargets_get_disease_id_description_by_name(
diseaseName="Niemann-Pick disease type C"
)
print(f"Disease: {disease_info['data']['name']}")
print(f"Description: {disease_info['data']['description'][:200]}...")
# 1.2: Literature for prevalence
prevalence_papers = tu.tools.PubMed_search_articles(
query="Niemann-Pick type C prevalence incidence epidemiology",
max_results=30
)
print(f"\nPrevalence papers: {len(prevalence_papers['data'])}")
# From literature: 1:120,000 births
us_population = 330000000
npc_prevalence = us_population / 120000
annual_births_us = 3700000
annual_incidence = annual_births_us / 120000
print(f"\nEstimated US Prevalence:")
print(f" Total NPC patients: {npc_prevalence:.0f}")
print(f" Annual incidence: {annual_incidence:.0f} new cases/year")
# 1.3: Get genetic basis
npc_genes = ["NPC1", "NPC2"]
for gene in npc_genes:
variants = tu.tools.ClinVar_search_variants(
gene=gene,
significance="pathogenic,likely_pathogenic"
)
print(f"\n{gene} pathogenic variants: {len(variants['data'])}")
# 1.4: Eligibility criteria impact (stricter for rare disease)
eligibility_factors = {
'confirmed_genetic_diagnosis': 0.90, # Molecular diagnosis required
'age_5_50': 0.60, # Exclude infantile and very late-onset
'ambulatory': 0.40, # Neurologic impairment common
'no_liver_transplant': 0.85,
'willing_travel': 0.30 # Only specialized centers
}
eligible_pool = npc_prevalence
print(f"\nEligibility Funnel:")
for criterion, factor in eligibility_factors.items():
eligible_pool *= factor
print(f" After {criterion}: {eligible_pool:.0f} ({factor*100:.0f}%)")
print(f"\nFinal eligible pool (US): {eligible_pool:.0f} patients")
print(f" Percentage of total NPC population: {(eligible_pool/npc_prevalence)*100:.1f}%")
# 1.5: Enrollment projection
target_n = 30 # Small N for rare disease
us_sites = 5 # Only specialized centers (NIH, Mayo, etc.)
international_sites = 10 # Need EU sites
total_eligible_global = eligible_pool * 3 # Assume 3× for US+EU+other
months_to_enroll = target_n / (total_eligible_global * 0.20 / 12) # 20% capture
print(f"\nEnrollment Projection:")
print(f" Target N: {target_n}")
print(f" US sites: {us_sites}")
print(f" International sites: {international_sites}")
print(f" Global eligible pool: ~{total_eligible_global:.0f}")
print(f" Estimated enrollment: {months_to_enroll:.1f} months ({months_to_enroll/12:.1f} years)")
print(f" ⚠ CHALLENGE: Multi-year enrollment for small study")Key Finding: Enrollment is a MAJOR challenge (36+ months for n=30)
Step 2: Natural History & Endpoint Selection
# 2.1: Search for natural history studies
nh_papers = tu.tools.PubMed_search_articles(
query="Niemann-Pick type C natural history clinical severity progression",
max_results=40
)
print("ENDPOINT SELECTION FOR RARE DISEASE")
print("="*80)
print(f"Natural history papers: {len(nh_papers['data'])}")
# 2.2: Search for existing clinical trials (learn from precedents)
npc_trials = tu.tools.search_clinical_trials(
condition="Niemann-Pick Disease Type C",
status="completed|active"
)
print(f"\nNPC clinical trials: {len(npc_trials['data'])}")
endpoints_used = {}
for trial in npc_trials['data']:
primary = trial.get('primary_outcome', '')
if primary:
key = primary[:50] # Truncate for grouping
endpoints_used[key] = endpoints_used.get(key, 0) + 1
print("\nEndpoints used in prior NPC trials:")
for endpoint, count in sorted(endpoints_used.items(), key=lambda x: x[1], reverse=True)[:5]:
print(f" - {endpoint}... (n={count} trials)")
# 2.3: Endpoint analysis
print("\n" + "="*80)
print("PRIMARY ENDPOINT ASSESSMENT")
print("="*80)
endpoint_options = [
{
'name': 'NPC Clinical Severity Score (17-domain)',
'evidence_grade': '★★☆',
'pros': 'Validated, used in prior trials, captures multi-organ impact',
'cons': 'Slow progression (need 18-24 month trial), subjective components',
'sample_size': 'n=30 per arm for 2-point change (90% power)',
'feasibility': 'MODERATE (long trial duration)'
},
{
'name': 'Swallowing function (videofluoroscopy)',
'evidence_grade': '★☆☆',
'pros': 'Objective, sensitive to change, clinically meaningful',
'cons': 'Not validated as primary, specialized equipment',
'sample_size': 'n=20 per arm (if validated)',
'feasibility': 'LOW (needs prospective validation)'
},
{
'name': 'Biomarker: plasma oxysterol (7-KC, cholestane-triol)',
'evidence_grade': '★★☆',
'pros': 'Mechanistic, rapid readout, validated in NPC',
'cons': 'Surrogate (not clinical benefit), FDA may not accept for approval',
'sample_size': 'n=15 per arm for 50% reduction',
'feasibility': 'HIGH (short trial, small N)'
}
]
for i, endpoint in enumerate(endpoint_options, 1):
print(f"\nOption {i}: {endpoint['name']} {endpoint['evidence_grade']}")
print(f" Pros: {endpoint['pros']}")
print(f" Cons: {endpoint['cons']}")
print(f" Sample size: {endpoint['sample_size']}")
print(f" Feasibility: {endpoint['feasibility']}")
print("\n" + "="*80)
print("RECOMMENDED ENDPOINT STRATEGY")
print("="*80)
print("Primary: NPC Clinical Severity Score (17-domain)")
print(" - Evidence grade: ★★☆ (used in prior trials)")
print(" - Design: Change from baseline to Month 18")
print(" - Sample size: n=30 (single-arm, vs. natural history)")
print("\nKey Secondary:")
print(" - Biomarker: Plasma oxysterols (7-KC, cholestane-triol)")
print(" - Swallowing function (videofluoroscopy)")
print(" - Ambulation index")
print(" - Quality of life (caregiver-rated)")
print("\n⚠ Challenge: 18-24 month trial duration compounds enrollment challenge")Step 3: Regulatory Pathway (Orphan Drug)
# 3.1: Orphan drug designation
print("REGULATORY PATHWAY: ORPHAN DRUG")
print("="*80)
print(f"Disease Prevalence: {npc_prevalence:.0f} patients in US")
print(f"Orphan Threshold: <200,000 affected individuals")
print(f"Status: ✓ QUALIFIES for Orphan Drug Designation")
print("\nOrphan Drug Benefits:")
print(" 1. 7-year market exclusivity")
print(" 2. Tax credits for clinical trial costs (25%)")
print(" 3. Waiver of PDUFA fees (~$3M)")
print(" 4. Protocol assistance from FDA")
print(" 5. Expedited review")
# 3.2: Search for orphan drug approvals in similar indications
orphan_papers = tu.tools.PubMed_search_articles(
query="FDA orphan drug approval lysosomal storage disease",
max_results=30
)
print(f"\nOrphan approvals in lysosomal storage: {len(orphan_papers['data'])} papers")
# 3.3: Regulatory precedents
print("\n" + "="*80)
print("REGULATORY PRECEDENTS (Similar Rare Diseases)")
print("="*80)
precedents = [
{
'drug': 'Miglustat (Zavesca)',
'disease': 'Gaucher disease type 1',
'year': 2003,
'endpoint': 'Organ volume reduction',
'design': 'Single-arm (n=28)',
'approval': 'Regular approval (not accelerated)'
},
{
'drug': 'Cerliponase alfa (Brineura)',
'disease': 'CLN2 Batten disease',
'year': 2017,
'endpoint': 'Motor-language score',
'design': 'Single-arm vs. natural history (n=24)',
'approval': 'Regular approval'
}
]
for p in precedents:
print(f"\n{p['drug']} ({p['year']}):")
print(f" Disease: {p['disease']}")
print(f" Endpoint: {p['endpoint']}")
print(f" Design: {p['design']}")
print(f" Approval: {p['approval']}")
print("\n" + "="*80)
print("RECOMMENDATION: Orphan Drug Designation + Natural History Control")
print("="*80)
print("Design: Single-arm Phase 2, n=30")
print("Comparator: Natural history cohort (published data + registry)")
print("Rationale:")
print(" - Ethical concerns with placebo in rare, progressive disease")
print(" - Well-characterized natural history available")
print(" - Regulatory precedent (Cerliponase alfa approved with NH control)")
print("\nKey Risk: FDA may still prefer randomized placebo-controlled")
print("Mitigation: Pre-IND meeting to discuss and gain alignment")Final Feasibility Assessment
print("\n" + "="*80)
print("FEASIBILITY SCORECARD: NIEMANN-PICK TYPE C TRIAL")
print("="*80)
dimensions = {
'Patient Availability': {
'weight': 0.30,
'raw_score': 3, # 3/10 - MAJOR CHALLENGE
'evidence': '★★★',
'rationale': 'Only ~500 eligible in US, 36+ months enrollment for n=30'
},
'Endpoint Precedent': {
'weight': 0.25,
'raw_score': 6, # 6/10
'evidence': '★★☆',
'rationale': 'Severity score used in trials, but slow progression'
},
'Regulatory Clarity': {
'weight': 0.20,
'raw_score': 8, # 8/10
'evidence': '★★★',
'rationale': 'Clear orphan path, precedents for NH control, FDA supportive'
},
'Comparator Feasibility': {
'weight': 0.15,
'raw_score': 7, # 7/10
'evidence': '★★☆',
'rationale': 'Natural history data available, registries exist'
},
'Safety Monitoring': {
'weight': 0.10,
'raw_score': 7, # 7/10
'evidence': '★☆☆',
'rationale': 'Novel mechanism, some preclinical safety data'
}
}
feasibility_score = sum(d['weight'] * d['raw_score'] * 10 for d in dimensions.values())
print(f"{'Dimension':<30} {'Weight':<10} {'Score':<10} {'Weighted':<10}")
print("-" * 70)
for dimension, data in dimensions.items():
weighted = data['weight'] * data['raw_score'] * 10
print(f"{dimension:<30} {data['weight']*100:.0f}%{'':<7} {data['raw_score']}/10{'':<5} {weighted:.1f}")
print(f" {data['rationale']}")
print("-" * 70)
print(f"TOTAL FEASIBILITY SCORE: {feasibility_score:.0f}/100 - MODERATE-LOW")
print("\n" + "="*80)
print("FINAL RECOMMENDATION: CONDITIONAL GO")
print("="*80)
print("""
This Phase 2 trial demonstrates MODERATE-LOW feasibility (Score: 58/100).
CRITICAL CHALLENGE: Patient recruitment
- Only ~500 eligible patients in US (after eligibility criteria)
- 36-48 months to enroll n=30, even with international sites
- Competing trials and natural history studies reduce available pool
STRENGTHS:
- Clear regulatory path (orphan drug, natural history control accepted)
- Significant unmet need (no approved therapies)
- Supportive patient advocacy and registry infrastructure
REQUIRED DE-RISKING STEPS:
1. Partnership with NPC patient registry (pre-identify patients)
2. Investigator consortium (NIH, Mayo, International NPC Consortium)
3. Pre-IND meeting to confirm natural history comparator acceptability
4. Biomarker enrichment (e.g., focus on NPC1 variants, exclude NPC2)
5. Adaptive design (allow enrollment extension if slow)
ALTERNATIVE DESIGN:
- If enrollment remains infeasible: Expanded Access Protocol
- Collect real-world data for future regulatory submission
- Smaller n=15-20 with biomarker primary endpoint (faster readout)
BUDGET: $5-8M (higher per-patient costs, longer duration)
TIMELINE: 48-60 months (enrollment + follow-up)
""")---
Example 3: Superiority Trial vs. Standard of Care (Checkpoint Inhibitor)
Scenario: Design Phase 2b randomized trial for novel PD-1 inhibitor vs. pembrolizumab in PD-L1 high (TPS ≥50%) NSCLC, first-line.
Setup
from tooluniverse import ToolUniverse
tu = ToolUniverse(use_cache=True)
tu.load_tools()
indication = "PD-L1 high (TPS ≥50%) non-small cell lung cancer, first-line"
design = "Phase 2b, randomized 1:1"
primary_endpoint = "Objective Response Rate (ORR)"
comparator = "pembrolizumab"Step 1: Patient Population (Biomarker-Selected)
# 1.1: Disease prevalence
disease_info = tu.tools.OpenTargets_get_disease_id_description_by_name(
diseaseName="non-small cell lung cancer"
)
# From literature: ~40% of NSCLC is PD-L1 TPS ≥50%
us_nsclc_annual = 200000
pdl1_high_rate = 0.40
pdl1_high_annual = us_nsclc_annual * pdl1_high_rate
print("PATIENT POPULATION SIZING: PD-L1 HIGH NSCLC")
print("="*80)
print(f"US NSCLC incidence: {us_nsclc_annual:,}/year")
print(f"PD-L1 TPS ≥50%: {pdl1_high_rate*100:.0f}% = {pdl1_high_annual:,}/year")
# 1.2: Eligibility criteria (first-line, no EGFR/ALK)
eligibility_factors = {
'no_egfr_alk': 0.80, # Exclude oncogene-driven
'age_18_plus': 0.98,
'ecog_0_1': 0.75,
'adequate_organ': 0.90,
'no_autoimmune': 0.85,
'no_prior_io': 1.00 # First-line
}
eligible_pool = pdl1_high_annual
print(f"\nEligibility Funnel:")
for criterion, factor in eligibility_factors.items():
eligible_pool *= factor
print(f" After {criterion}: {eligible_pool:,.0f} ({factor*100:.0f}%)")
print(f"\nFinal eligible pool: {eligible_pool:,.0f} patients/year (US)")
# 1.3: Enrollment projection for randomized trial
target_n = 120 # 60 per arm
sites = 30 # Large Phase 2b
capture_rate = 0.05 # 5% of eligible
monthly_enrollment = (eligible_pool * capture_rate) / 12 / sites
print(f"\nEnrollment Projection (Randomized 1:1):")
print(f" Target N: {target_n} ({target_n//2} per arm)")
print(f" Sites: {sites}")
print(f" Patients per site per month: {monthly_enrollment:.2f}")
print(f" Enrollment timeline: {target_n / (monthly_enrollment * sites):.1f} months")Step 2: Comparator Analysis (Pembrolizumab SOC)
# 2.1: Get pembrolizumab info
pembro_info = tu.tools.drugbank_get_drug_basic_info_by_drug_name_or_id(
drug_name_or_drugbank_id=comparator
)
pembro_indications = tu.tools.drugbank_get_indications_by_drug_name_or_drugbank_id(
drug_name_or_drugbank_id=comparator
)
print("\nCOMPARATOR DRUG: PEMBROLIZUMAB")
print("="*80)
print(f"Status: {pembro_info['data']['groups']}")
print(f"Indications: {len(pembro_indications['data'])}")
# 2.2: Get FDA approval history
pembro_approval = tu.tools.OpenFDA_get_approval_history(
drug_name=comparator
)
# 2.3: Get pivotal trial data
keynote_trials = tu.tools.search_clinical_trials(
intervention="pembrolizumab",
condition="non-small cell lung cancer",
phase="3",
status="completed"
)
print(f"\nPembrolizumab Phase 3 trials in NSCLC: {len(keynote_trials['data'])}")
# Key trial: KEYNOTE-024 (1L, PD-L1 ≥50%)
print("\n" + "="*80)
print("KEYNOTE-024: Pembrolizumab vs. Chemotherapy (PD-L1 ≥50%)")
print("="*80)
print("Design: Randomized 1:1, N=305")
print("Population: 1L NSCLC, PD-L1 TPS ≥50%")
print("Primary: PFS")
print("\nResults:")
print(" Pembrolizumab:")
print(" - ORR: 44.8%")
print(" - Median PFS: 10.3 months")
print(" - Median OS: 30.0 months (30-month landmark)")
print(" Chemotherapy:")
print(" - ORR: 27.8%")
print(" - Median PFS: 6.0 months")
print(" - Median OS: 14.2 months")
print("\nConclusion: Pembrolizumab is SOC for PD-L1 ≥50% NSCLC")
# 2.4: Comparator drug sourcing
print("\n" + "="*80)
print("COMPARATOR SOURCING")
print("="*80)
print("Drug: Pembrolizumab (Keytruda)")
print("Availability: Commercial supply")
print("Cost: ~$150-180K/year per patient")
print("Dosing: 200 mg IV q3w")
print("Stability: Refrigerated, 24-hour room temp after reconstitution")
print("Sourcing: Purchase from Merck or specialty pharmacy")
print("IRB Consideration: Standard of care, no IND required for comparator arm")Step 3: Sample Size Calculation (Superiority Design)
print("\n" + "="*80)
print("SAMPLE SIZE CALCULATION: SUPERIORITY TRIAL (ORR)")
print("="*80)
# Assumptions
pembro_orr = 0.45 # 45% from KEYNOTE-024
target_orr = 0.60 # 60% for novel drug (15% absolute improvement)
alpha = 0.05 # Two-sided
power = 0.80
dropout = 0.10
# Calculate (using normal approximation for proportions)
import math
def sample_size_two_proportions(p1, p2, alpha=0.05, power=0.80):
"""Calculate sample size for comparing two proportions"""
z_alpha = 1.96 # Two-sided alpha=0.05
z_beta = 0.84 # Power=0.80
p_avg = (p1 + p2) / 2
n = ((z_alpha * math.sqrt(2 * p_avg * (1 - p_avg)) +
z_beta * math.sqrt(p1 * (1 - p1) + p2 * (1 - p2)))**2 /
(p1 - p2)**2)
return math.ceil(n)
n_per_arm = sample_size_two_proportions(target_orr, pembro_orr)
n_total = n_per_arm * 2
n_with_dropout = math.ceil(n_total / (1 - dropout))
print(f"Null Hypothesis (H0): ORR_novel = ORR_pembro = {pembro_orr*100:.0f}%")
print(f"Alternative (H1): ORR_novel = {target_orr*100:.0f}% (15% absolute improvement)")
print(f"Alpha: {alpha} (two-sided)")
print(f"Power: {power*100:.0f}%")
print(f"\nSample Size:")
print(f" Per arm: {n_per_arm}")
print(f" Total: {n_total}")
print(f" With {dropout*100:.0f}% dropout: {n_with_dropout} ({n_with_dropout//2}/arm)")
print(f"\n⚠ NOTE: This is Phase 2b, not pivotal")
print(f" - Powering for hypothesis generation, not definitive proof")
print(f" - N={n_with_dropout} reasonable for Phase 2b go/no-go decision")
print(f" - Successful Phase 2b → Phase 3 with PFS/OS primary (N=400-600)")
# 2.5: Alternative: Non-inferiority design (if aiming for better safety)
print("\n" + "="*80)
print("ALTERNATIVE DESIGN: NON-INFERIORITY (If Better Safety)")
print("="*80)
print("Rationale: If novel drug has lower toxicity (e.g., no pneumonitis)")
print("Non-inferiority margin: Δ = -10% (ORR novel ≥ 35% if pembro 45%)")
print("Sample size: ~200/arm (larger N for non-inferiority)")
print("Conclusion: STICK WITH SUPERIORITY for Phase 2b (smaller N)")Step 4: Safety Comparison
# 4.1: Pembrolizumab safety profile
pembro_warnings = tu.tools.FDA_get_warnings_and_cautions_by_drug_name(
drug_name=comparator
)
pembro_faers = tu.tools.FAERS_count_reactions_by_drug_event(
medicinalproduct="PEMBROLIZUMAB"
)
print("\n" + "="*80)
print("SAFETY PROFILE: PEMBROLIZUMAB (COMPARATOR)")
print("="*80)
print(f"FDA warnings: {len(pembro_warnings.get('data', []))}")
print("\nTop 10 Adverse Events (FAERS):")
for i, ae in enumerate(pembro_faers.get('results', [])[:10], 1):
print(f" {i}. {ae['term']}: {ae['count']} reports")
print("\nKey Immune-Related AEs (irAEs) from Label:")
irAEs = [
{'event': 'Pneumonitis', 'incidence': '3.4%', 'grade_3_4': '0.8%', 'fatal': '0.1%'},
{'event': 'Colitis', 'incidence': '1.7%', 'grade_3_4': '0.4%', 'fatal': '0%'},
{'event': 'Hepatitis', 'incidence': '0.7%', 'grade_3_4': '0.2%', 'fatal': '0.1%'},
{'event': 'Endocrinopathies (all)', 'incidence': '8.0%', 'grade_3_4': '0.8%', 'fatal': '0%'},
{'event': 'Nephritis', 'incidence': '0.7%', 'grade_3_4': '0.2%', 'fatal': '0%'}
]
print("\n" + "="*80)
print("IMMUNE-RELATED ADVERSE EVENTS (irAEs)")
print("="*80)
print(f"{'Event':<25} {'Any Grade':<12} {'Grade 3-4':<12} {'Fatal':<10}")
print("-" * 60)
for ae in irAEs:
print(f"{ae['event']:<25} {ae['incidence']:<12} {ae['grade_3_4']:<12} {ae['fatal']:<10}")
print("\n" + "="*80)
print("SAFETY MONITORING PLAN (Both Arms)")
print("="*80)
print("Standard immune-related AE monitoring:")
print(" - Baseline: CXR, PFTs (if smoker), TSH, LFTs, Cr")
print(" - Every cycle: CBC, CMP, LFTs, TSH")
print(" - PRN: CT chest for respiratory symptoms, cortisol/ACTH for fatigue")
print("\nStopping rules for irAEs:")
print(" - Grade 2 pneumonitis: Hold drug, imaging, pulmonology")
print(" - Grade 3-4 irAE: Discontinue drug, high-dose steroids (1-2 mg/kg)")
print(" - Grade 4 or fatal irAE: Report to FDA ASAP (IND safety report)")Final Feasibility & Design Recommendation
print("\n" + "="*80)
print("FEASIBILITY SCORECARD: PEMBROLIZUMAB SUPERIORITY TRIAL")
print("="*80)
dimensions = {
'Patient Availability': {
'weight': 0.30,
'raw_score': 9,
'evidence': '★★★',
'rationale': '~40,000 eligible/year, rapid enrollment (6-8 months for N=120)'
},
'Endpoint Precedent': {
'weight': 0.25,
'raw_score': 9,
'evidence': '★★★',
'rationale': 'ORR standard for Phase 2, precedent in KEYNOTE trials'
},
'Regulatory Clarity': {
'weight': 0.20,
'raw_score': 8,
'evidence': '★★☆',
'rationale': 'Active control acceptable, Phase 3 will need PFS/OS'
},
'Comparator Feasibility': {
'weight': 0.15,
'raw_score': 9,
'evidence': '★★★',
'rationale': 'Pembrolizumab commercial supply, known efficacy (ORR 45%)'
},
'Safety Monitoring': {
'weight': 0.10,
'raw_score': 8,
'evidence': '★★★',
'rationale': 'PD-1 class effects well-known, irAE protocols established'
}
}
feasibility_score = sum(d['weight'] * d['raw_score'] * 10 for d in dimensions.values())
print(f"{'Dimension':<30} {'Weight':<10} {'Score':<10} {'Weighted':<10}")
print("-" * 70)
for dimension, data in dimensions.items():
weighted = data['weight'] * data['raw_score'] * 10
print(f"{dimension:<30} {data['weight']*100:.0f}%{'':<7} {data['raw_score']}/10{'':<5} {weighted:.1f}")
print("-" * 70)
print(f"TOTAL FEASIBILITY SCORE: {feasibility_score:.0f}/100 - HIGH")
print("\n" + "="*80)
print("FINAL RECOMMENDATION: RECOMMEND PROCEED")
print("="*80)
print(f"""
This Phase 2b randomized trial demonstrates HIGH feasibility (Score: {feasibility_score:.0f}/100).
RECOMMENDED DESIGN:
- Phase 2b, randomized 1:1, open-label
- N=120 (60 per arm) with 10% dropout buffer → Total N=132
- Population: 1L NSCLC, PD-L1 TPS ≥50%, no EGFR/ALK
- Treatment:
* Arm A: Novel PD-1 inhibitor [dose] IV q3w
* Arm B: Pembrolizumab 200 mg IV q3w
- Primary endpoint: ORR (RECIST 1.1, iRECIST for pseudoprogression)
- Secondary: PFS, OS (long-term FU), DoR, safety
- Duration: Until progression, toxicity, or 24 months
STATISTICAL PLAN:
- Power: 80% to detect 15% absolute ORR improvement (60% vs 45%)
- Analysis: Chi-square test, two-sided alpha=0.05
- Interim: 50% information (n=60), futility only (no early efficacy stop)
SUCCESS CRITERIA (Advance to Phase 3):
- ORR ≥ 55% (vs. pembro 45%, p<0.05)
- Safety profile non-inferior (no new safety signals)
- PFS trend favorable (HR <0.85)
- Duration of response ≥12 months (median)
ENROLLMENT:
- Timeline: 6-8 months (30 sites, ~0.5 patients/site/month)
- Primary analysis: Month 14 (6-month follow-up for ORR assessment)
BUDGET: $6-9M (higher cost due to comparator drug purchase + 2× monitoring)
""")---
Example 4: Non-Inferiority Trial (Oral Anticoagulant)
Scenario: Design Phase 3 non-inferiority trial for novel oral Factor XIa inhibitor vs. apixaban in atrial fibrillation, aiming for lower bleeding risk.
from tooluniverse import ToolUniverse
tu = ToolUniverse(use_cache=True)
tu.load_tools()
indication = "Atrial fibrillation, stroke prevention"
design = "Phase 3, randomized, double-blind, non-inferiority"
primary_endpoint = "Stroke or systemic embolism (composite)"
comparator = "apixaban"
# Step 1: Patient Population (Very Large)
print("PATIENT POPULATION: ATRIAL FIBRILLATION")
print("="*80)
# AFib prevalence: ~6M in US, ~33% on oral anticoagulation
us_afib_prevalence = 6000000
on_anticoagulation = us_afib_prevalence * 0.33
print(f"US AFib prevalence: {us_afib_prevalence:,}")
print(f"On oral anticoagulation: {on_anticoagulation:,}")
print(f"Eligible for trial: ~50% = {on_anticoagulation * 0.5:,.0f}")
# Step 2: Comparator (Apixaban - Standard of Care)
apixaban_info = tu.tools.drugbank_get_drug_basic_info_by_drug_name_or_id(
drug_name_or_drugbank_id=comparator
)
print(f"\nComparator: {comparator}")
print(f"Status: {apixaban_info['data']['groups']}")
# From ARISTOTLE trial: Apixaban vs. warfarin
print("\nARISTOTLE Trial (Apixaban Pivotal):")
print(" N=18,201")
print(" Stroke/SE rate: 1.27%/year (apixaban) vs. 1.60%/year (warfarin)")
print(" Major bleeding: 2.13%/year (apixaban) vs. 3.09%/year (warfarin)")
# Step 3: Non-Inferiority Margin
print("\n" + "="*80)
print("NON-INFERIORITY MARGIN DETERMINATION")
print("="*80)
print("Regulatory Guidance (FDA/EMA):")
print(" - NI margin = Fraction of comparator effect vs. placebo")
print(" - Apixaban vs. warfarin: 21% relative risk reduction in stroke")
print(" - Acceptable to preserve ≥50% of effect → NI margin = HR <1.38")
print("\nProposed NI Margin: HR <1.44 (upper bound of 95% CI)")
print("Rationale: Conservative, if novel drug has 50% lower bleeding")
# Step 4: Sample Size (LARGE)
print("\n" + "="*80)
print("SAMPLE SIZE CALCULATION: NON-INFERIORITY (TIME-TO-EVENT)")
print("="*80)
# Assumptions
stroke_rate_annual = 0.0127 # 1.27%/year from ARISTOTLE
ni_margin_hr = 1.44
alpha = 0.025 # One-sided for NI
power = 0.90
enrollment_period = 24 # months
follow_up = 36 # months
# Events needed (conservative estimate)
events_needed = 450 # Assuming HR=1.0, NI margin 1.44, 90% power
# Sample size (based on event rate)
n_per_arm = events_needed / (2 * stroke_rate_annual * (enrollment_period + follow_up) / 12)
n_total = n_per_arm * 2
print(f"Assumptions:")
print(f" Annual event rate: {stroke_rate_annual*100:.2f}%")
print(f" NI margin: HR <{ni_margin_hr}")
print(f" Power: {power*100:.0f}%")
print(f" Alpha: {alpha} (one-sided)")
print(f"\nRequired events: {events_needed}")
print(f"Sample size: {n_total:,.0f} ({n_per_arm:,.0f} per arm)")
print(f"Duration: {enrollment_period} months enrollment + {follow_up} months follow-up = {enrollment_period + follow_up} months")
print(f"\n⚠ CHALLENGE: VERY LARGE TRIAL (N={n_total:,.0f})")
print(f" - Phase 3 only (skip Phase 2 or run small Phase 2 for dose)")
print(f" - Multi-national (US, EU, Asia, LatAm)")
print(f" - Budget: $150-250M")
# Step 5: Regulatory Considerations
print("\n" + "="*80)
print("REGULATORY PATHWAY: CARDIOVASCULAR OUTCOMES TRIAL")
print("="*80)
print("FDA Requirements:")
print(" - Pre-IND: Type C meeting for NI margin discussion")
print(" - Phase 3 design: Randomized, double-blind, event-driven")
print(" - Primary: Stroke/SE (non-inferiority)")
print(" - Key secondary: Major bleeding (superiority)")
print(" - DSMB: Independent, frequent reviews (every 200 events)")
print(" - Interim: 1-2 analyses (efficacy + futility)")
print("\nApproval Pathway: Regular NDA (not accelerated)")
# Feasibility Score
print("\n" + "="*80)
print("FEASIBILITY SCORE: NON-INFERIORITY TRIAL")
print("="*80)
dimensions = {
'Patient Availability': {'weight': 0.30, 'raw_score': 10, 'rationale': 'Huge population (>1M eligible)'},
'Endpoint Precedent': {'weight': 0.25, 'raw_score': 10, 'rationale': 'Stroke/SE standard, regulatory accepted'},
'Regulatory Clarity': {'weight': 0.20, 'raw_score': 7, 'rationale': 'NI margin needs FDA agreement, precedent exists'},
'Comparator Feasibility': {'weight': 0.15, 'raw_score': 9, 'rationale': 'Apixaban generic, widely used'},
'Safety Monitoring': {'weight': 0.10, 'raw_score': 9, 'rationale': 'Bleeding monitoring standard, established'}
}
feasibility_score = sum(d['weight'] * d['raw_score'] * 10 for d in dimensions.values())
print(f"TOTAL FEASIBILITY SCORE: {feasibility_score:.0f}/100 - HIGH (but expensive)")
print(f"\nRECOMMENDATION: CONDITIONAL GO (if funded)")
print(f" - Feasibility: HIGH for patient access, endpoints, regulatory")
print(f" - Challenge: $150-250M budget, 5-year duration")
print(f" - Strategy: Partner with large pharma or seek CV outcomes specialist CRO")---
Example 5: Basket Trial (Multiple Cancers, One Biomarker)
Scenario: Design basket trial for NTRK fusion-positive solid tumors (tissue-agnostic), following larotrectinib precedent.
from tooluniverse import ToolUniverse
tu = ToolUniverse(use_cache=True)
tu.load_tools()
indication = "NTRK fusion-positive solid tumors (basket trial)"
design = "Phase 2, single-arm, basket (multiple histologies)"
primary_endpoint = "Objective Response Rate (ORR) by histology"
biomarker = "NTRK1/2/3 gene fusion"
# Step 1: Biomarker Prevalence (Rare but Pan-Cancer)
print("BIOMARKER PREVALENCE: NTRK FUSIONS")
print("="*80)
# NTRK fusions: rare (<1% most cancers, enriched in pediatric/rare tumors)
ntrk_prevalence_by_cancer = [
{'cancer': 'Secretory breast carcinoma', 'prevalence': 0.90, 'annual_us': 100},
{'cancer': 'Mammary analogue secretory carcinoma', 'prevalence': 0.95, 'annual_us': 50},
{'cancer': 'Infantile fibrosarcoma', 'prevalence': 0.90, 'annual_us': 30},
{'cancer': 'NSCLC', 'prevalence': 0.01, 'annual_us': 2000}, # 1% of 200K
{'cancer': 'Colorectal cancer', 'prevalence': 0.005, 'annual_us': 700}, # 0.5% of 140K
{'cancer': 'Thyroid cancer', 'prevalence': 0.05, 'annual_us': 250},
{'cancer': 'Glioblastoma', 'prevalence': 0.01, 'annual_us': 130},
{'cancer': 'Salivary gland', 'prevalence': 0.03, 'annual_us': 10},
{'cancer': 'Sarcoma (other)', 'prevalence': 0.02, 'annual_us': 250},
{'cancer': 'Melanoma', 'prevalence': 0.001, 'annual_us': 90}
]
print(f"{'Cancer Type':<40} {'NTRK+ Rate':<15} {'Annual US Cases':<20}")
print("-" * 80)
total_ntrk_patients = 0
for cancer in ntrk_prevalence_by_cancer:
print(f"{cancer['cancer']:<40} {cancer['prevalence']*100:>6.1f}% {cancer['annual_us']:>15,}")
total_ntrk_patients += cancer['annual_us']
print("-" * 80)
print(f"{'TOTAL NTRK+ PATIENTS/YEAR (US)':<40} {'':<15} {total_ntrk_patients:>15,}")
print(f"\nKey Insight: NTRK fusions are RARE (~{total_ntrk_patients:,}/year across all cancers)")
print(f" → Basket trial required to aggregate sufficient patients")
# Step 2: Basket Trial Design
print("\n" + "="*80)
print("BASKET TRIAL DESIGN")
print("="*80)
print("Concept: Enroll patients across MULTIPLE tumor types with NTRK fusions")
print("Rationale: NTRK inhibitor mechanism is tumor-agnostic")
print("\nDesign:")
print(" - Single-arm, open-label")
print(" - Primary endpoint: ORR per histology (≥15 patients/basket)")
print(" - Secondary: DoR, PFS, OS, safety (across all baskets)")
print(" - Enrollment: ~80-120 patients across 8-12 tumor types")
print("\nInclusion Criteria:")
print(" - NTRK1/2/3 fusion confirmed (NGS, FISH, IHC→FISH)")
print(" - Locally advanced or metastatic disease")
print(" - Measurable disease (RECIST 1.1)")
print(" - No effective standard therapy OR progressed on SOC")
print(" - Age ≥12 years (include pediatric)")
# Step 3: Biomarker Testing Strategy
print("\n" + "="*80)
print("BIOMARKER TESTING STRATEGY")
print("="*80)
# Search for NGS panels
print("NGS-based comprehensive genomic profiling (CGP):")
print(" - FoundationOne CDx (FDA-approved CDx for larotrectinib)")
print(" - Guardant360 (liquid biopsy, ctDNA)")
print(" - Institutional NGS (CLIA-certified)")
print("\nTesting Algorithm:")
print(" 1. IHC screening (pan-TRK antibody) - Fast, cheap")
print(" → If positive: Confirm with NGS or FISH")
print(" 2. NGS (if tumor profiling already done)")
print(" 3. FISH (if NGS unavailable or IHC+)")
print("\nTurnaround: 10-14 days (NGS), 5-7 days (FISH)")
print("Cost: $500-800 (IHC), $3,000-5,000 (NGS)")
# Search for testing guidelines
testing_papers = tu.tools.PubMed_search_articles(
query="NTRK fusion testing guidelines NCCN NGS",
max_results=20
)
print(f"\nNTRK testing guideline papers: {len(testing_papers['data'])}")
# Step 4: Regulatory Precedent (Larotrectinib)
print("\n" + "="*80)
print("REGULATORY PRECEDENT: LAROTRECTINIB (Vitrakvi)")
print("="*80)
# Search for larotrectinib approval
laro_papers = tu.tools.PubMed_search_articles(
query="larotrectinib FDA approval NTRK fusion basket trial",
max_results=25
)
print(f"Larotrectinib papers: {len(laro_papers['data'])}")
print("\nLarotrectinib FDA Approval (2018):")
print(" Indication: NTRK fusion-positive solid tumors (TISSUE-AGNOSTIC)")
print(" Trial Design: Basket trial, single-arm, n=55")
print(" Primary Endpoint: ORR")
print(" Results:")
print(" - Overall ORR: 75% (95% CI: 61-85%)")
print(" - Complete response: 13%")
print(" - Partial response: 62%")
print(" - Median DoR: NOT REACHED (73% at 12 months)")
print(" Histologies enrolled: 17 different tumor types")
print(" Approval: ACCELERATED (tissue-agnostic, first of its kind)")
print("\nKey Regulatory Insights:")
print(" ✓ Single-arm acceptable (no SOC for NTRK+ tumors)")
print(" ✓ ORR primary endpoint sufficient")
print(" ✓ Small N (n=55) acceptable given tumor rarity")
print(" ✓ Tissue-agnostic approval precedent SET")
# Step 5: Enrollment Feasibility
print("\n" + "="*80)
print("ENROLLMENT STRATEGY")
print("="*80)
print(f"Total NTRK+ patients/year (US): ~{total_ntrk_patients:,}")
print(f"Target enrollment: 100 patients")
print(f"Capture rate needed: {100/total_ntrk_patients*100:.1f}%")
print("\nChallenges:")
print(" 1. BROAD SCREENING: Need to test 10,000+ patients to find 100 NTRK+")
print(" 2. Competing trials: Larotrectinib, entrectinib already approved")
print(" 3. Limited testing: Not all centers do NGS routinely")
print("\nMitigation Strategies:")
print(" 1. PARTNER WITH CGP COMPANIES:")
print(" - Foundation Medicine, Guardant, Tempus")
print(" - Flag NTRK+ patients in CGP reports → refer to trial")
print(" 2. NATIONAL SCREENING PROGRAM:")
print(" - Provide free NGS testing for suspected rare fusions")
print(" 3. PATIENT ADVOCACY:")
print(" - Partner with LUNGevity, CRC Alliance, sarcoma foundations")
print(" 4. PEDIATRIC SITES:")
print(" - Children's hospitals (infantile fibrosarcoma enriched)")
print(" 5. INTERNATIONAL EXPANSION:")
print(" - EU, Asia sites (3× enrollment pool)")
print("\nProjected Enrollment Timeline:")
print(" - US sites: 25-30 (comprehensive cancer centers)")
print(" - International: 20-30")
print(" - Monthly enrollment: 2-3 patients (across all sites)")
print(" - Duration: 36-48 months to reach n=100")
# Step 6: Statistical Design
print("\n" + "="*80)
print("STATISTICAL DESIGN (BASKET TRIAL)")
print("="*80)
print("Primary Analysis: ORR per tumor type")
print(" - Minimum 15 patients per basket for analysis")
print(" - Success threshold: ORR ≥40% (vs. null ≤10%)")
print(" - If 6/15 respond (40%), 95% CI: 16-68% → Clinically meaningful")
print("\nOverall Analysis (All Tumor Types Combined):")
print(" - Secondary analysis")
print(" - Target overall ORR: ≥60% (following larotrectinib)")
print("\nInterim Analysis:")
print(" - After 30 patients: Assess safety, futility")
print(" - If ORR <20%, consider stopping")
# Feasibility Score
print("\n" + "="*80)
print("FEASIBILITY SCORECARD: NTRK BASKET TRIAL")
print("="*80)
dimensions = {
'Patient Availability': {
'weight': 0.30,
'raw_score': 4,
'evidence': '★★★',
'rationale': 'VERY RARE (~3.6K/year US), need broad screening'
},
'Endpoint Precedent': {
'weight': 0.25,
'raw_score': 10,
'evidence': '★★★',
'rationale': 'ORR accepted, larotrectinib precedent (tissue-agnostic approval)'
},
'Regulatory Clarity': {
'weight': 0.20,
'raw_score': 9,
'evidence': '★★★',
'rationale': 'Clear path (larotrectinib precedent), accelerated approval'
},
'Comparator Feasibility': {
'weight': 0.15,
'raw_score': 8,
'evidence': '★★☆',
'rationale': 'Single-arm acceptable (no SOC for NTRK+), larotrectinib approved'
},
'Safety Monitoring': {
'weight': 0.10,
'raw_score': 8,
'evidence': '★★☆',
'rationale': 'TRK inhibitor class known, manageable AEs'
}
}
feasibility_score = sum(d['weight'] * d['raw_score'] * 10 for d in dimensions.values())
print(f"{'Dimension':<30} {'Weight':<10} {'Score':<10} {'Weighted':<10}")
print("-" * 70)
for dimension, data in dimensions.items():
weighted = data['weight'] * data['raw_score'] * 10
print(f"{dimension:<30} {data['weight']*100:.0f}%{'':<7} {data['raw_score']}/10{'':<5} {weighted:.1f}")
print("-" * 70)
print(f"TOTAL FEASIBILITY SCORE: {feasibility_score:.0f}/100 - MODERATE")
print("\n" + "="*80)
print("FINAL RECOMMENDATION: CONDITIONAL GO")
print("="*80)
print("""
This basket trial demonstrates MODERATE feasibility (Score: 68/100).
CRITICAL SUCCESS FACTOR: Screening Partnership
- NTRK fusions are ultra-rare (<0.5% across cancers)
- MUST partner with CGP companies (Foundation, Guardant) to identify patients
- Alternative: Provide sponsored NGS testing program
STRENGTHS:
- Clear regulatory path (larotrectinib precedent)
- ORR endpoint accepted for tissue-agnostic approval
- High unmet need (no effective SOC for NTRK+ tumors)
CHALLENGES:
- Slow enrollment (36-48 months for n=100)
- Competing drugs (larotrectinib, entrectinib already approved)
- High screening costs (need to test 10,000+ patients)
RECOMMENDED STRATEGY:
1. Phase 1 dose escalation (n=20-30, multiple tumor types)
2. Phase 2 basket expansion (n=80-100, ≥15 per histology)
3. Concurrent: Pediatric arm (infantile fibrosarcoma, few alternatives)
4. Regulatory: Breakthrough therapy designation (after Phase 1 signals)
5. Screening: Partner with 2-3 CGP companies, patient registries
BUDGET: $15-25M (high screening costs, slow enrollment)
TIMELINE: 48-60 months (enrollment) + 12-18 months (follow-up for DoR)
""")---
Summary Table: All 5 Examples
| Example | Phase | Design | Indication | Biomarker | Primary Endpoint | Feasibility Score | Recommendation | Key Challenge |
|---|---|---|---|---|---|---|---|---|
| 1. EGFR+ NSCLC | 1/2 | Single-arm | Osimertinib-resistant NSCLC | EGFR L858R | ORR | 82/100 (HIGH) | PROCEED | None major |
| 2. Niemann-Pick C | 2 | Single-arm vs. NH | Rare lysosomal storage | Genetic (NPC1/2) | Clinical severity score | 58/100 (MOD-LOW) | CONDITIONAL GO | Slow enrollment (36+ mo) |
| 3. PD-L1 High NSCLC | 2b | Randomized 1:1 | First-line NSCLC | PD-L1 TPS ≥50% | ORR | 87/100 (HIGH) | PROCEED | Comparator cost |
| 4. Atrial Fibrillation | 3 | Non-inferiority | AFib stroke prevention | None | Stroke/SE | 90/100 (HIGH) | CONDITIONAL GO | Large N, expensive ($150-250M) |
| 5. NTRK Basket | 2 | Basket, single-arm | Pan-cancer | NTRK fusion | ORR by histology | 68/100 (MODERATE) | CONDITIONAL GO | Ultra-rare (broad screening) |
Key Learnings:
- Biomarker-selected oncology trials (Ex 1, 3) have HIGH feasibility
- Rare diseases (Ex 2) face enrollment challenges → need registries
- Non-inferiority trials (Ex 4) are feasible but expensive (large N)
- Basket trials (Ex 5) require broad screening partnerships
#!/usr/bin/env python3
"""
CLINICAL TRIAL DESIGN FEASIBILITY - COMPLETE WORKING PIPELINE
Fixed version using correct ToolUniverse tools.
Assesses trial feasibility across 6 research dimensions.
"""
from tooluniverse import ToolUniverse
from datetime import datetime
class TrialFeasibilityAnalyzer:
"""Complete trial design feasibility pipeline."""
def __init__(self):
"""Initialize ToolUniverse."""
print("Initializing ToolUniverse...")
self.tu = ToolUniverse()
self.tu.load_tools()
print(f"✅ Loaded {len(self.tu.all_tool_dict)} tools\n")
def analyze(self, indication, drug_name, phase="Phase 2", output_file=None):
"""
Complete trial feasibility analysis.
Args:
indication: Disease/indication (e.g., "EGFR-mutant NSCLC")
drug_name: Drug/intervention name
phase: Trial phase
output_file: Optional report file
Returns:
dict with feasibility analysis
"""
if output_file is None:
output_file = f"Trial_Feasibility_{drug_name.replace(' ', '_')}.md"
print("=" * 80)
print(f"TRIAL FEASIBILITY ANALYSIS")
print(f"Indication: {indication}")
print(f"Drug: {drug_name}")
print(f"Phase: {phase}")
print("=" * 80)
report = {
'indication': indication,
'drug': drug_name,
'phase': phase,
'timestamp': datetime.now().isoformat(),
'disease_info': {},
'drug_info': {},
'precedent_trials': [],
'safety_data': {},
'feasibility_score': 0
}
# Create report file
self._create_report(output_file, indication, drug_name, phase)
print("\n🔬 Running Feasibility Analysis...")
print("-" * 80)
# PATH 1: Patient Population Sizing
report['disease_info'] = self._analyze_patient_population(indication)
self._update_report(output_file, "## 1. Patient Population", report['disease_info'])
# PATH 2: Drug/Intervention Profile
report['drug_info'] = self._analyze_drug(drug_name)
self._update_report(output_file, "## 2. Drug Profile", report['drug_info'])
# PATH 3: Precedent Trials
report['precedent_trials'] = self._find_precedent_trials(indication, drug_name)
self._update_report(output_file, "## 3. Precedent Trials", report['precedent_trials'])
# PATH 4: Safety Assessment
report['safety_data'] = self._assess_safety(drug_name)
self._update_report(output_file, "## 4. Safety Profile", report['safety_data'])
# PATH 5: Literature Evidence
literature = self._search_literature(indication, drug_name)
self._update_report(output_file, "## 5. Literature Evidence", literature)
# PATH 6: Feasibility Scoring
report['feasibility_score'] = self._calculate_feasibility(report)
self._update_report(output_file, "## 6. Feasibility Assessment", {
'score': report['feasibility_score'],
'interpretation': self._interpret_feasibility(report['feasibility_score'])
})
print(f"\n✅ Analysis complete! Report saved to: {output_file}")
print(f"📊 Feasibility Score: {report['feasibility_score']}/100")
return report
def _create_report(self, filename, indication, drug, phase):
"""Create initial report file."""
with open(filename, 'w') as f:
f.write(f"# Clinical Trial Feasibility Report\n\n")
f.write(f"**Indication**: {indication}\n")
f.write(f"**Drug/Intervention**: {drug}\n")
f.write(f"**Phase**: {phase}\n")
f.write(f"**Analysis Date**: {datetime.now().strftime('%Y-%m-%d')}\n\n")
f.write("---\n\n")
def _update_report(self, filename, section, data):
"""Update report with new section."""
with open(filename, 'a') as f:
f.write(f"\n{section}\n\n")
if isinstance(data, dict):
for key, value in data.items():
f.write(f"**{key}**: {value}\n\n")
elif isinstance(data, list):
for item in data:
if isinstance(item, dict):
f.write(f"- {item}\n")
else:
f.write(f"- {item}\n")
f.write("\n")
else:
f.write(f"{data}\n\n")
def _analyze_patient_population(self, indication):
"""Analyze patient population size and characteristics."""
print("\n1️⃣ Patient Population Analysis")
population = {}
# Search for disease information (Open Targets)
print(f" Searching disease database for: {indication}")
try:
result = self.tu.tools.OpenTargets_get_disease_id_description_by_name(
disease_name=indication
)
if result.get('data', {}).get('diseases'):
diseases = result['data']['diseases']
if diseases:
disease = diseases[0]
population['disease_id'] = disease.get('id', 'N/A')
population['disease_name'] = disease.get('name', indication)
population['description'] = disease.get('description', 'N/A')[:200]
print(f" ✅ Found disease: {disease.get('name')}")
except Exception as e:
print(f" ⚠️ Error: {e}")
population['disease_name'] = indication
population['description'] = 'Could not retrieve disease information'
# Estimate prevalence (using literature search as proxy)
print(f" Estimating prevalence...")
try:
result = self.tu.tools.PubMed_search_articles(
query=f'"{indication}"[Title/Abstract] AND "prevalence"[Title/Abstract]',
max_results=5
)
if isinstance(result, dict) and result.get('data', {}).get('articles'):
articles = result['data']['articles']
population['prevalence_literature'] = len(articles)
print(f" ✅ Found {len(articles)} prevalence articles")
else:
population['prevalence_literature'] = 0
except Exception as e:
print(f" ⚠️ Error: {e}")
population['prevalence_literature'] = 0
return population
def _analyze_drug(self, drug_name):
"""Get drug information from DrugBank."""
print("\n2️⃣ Drug Profile Analysis")
drug_info = {}
print(f" Querying DrugBank for: {drug_name}")
try:
result = self.tu.tools.drugbank_get_drug_basic_info_by_drug_name_or_id(
query=drug_name, # ✅ CORRECT parameter
case_sensitive=False,
exact_match=False,
limit=1
)
if result.get('data', {}).get('drugs'):
drug = result['data']['drugs'][0]
drug_info['drugbank_id'] = drug.get('drugbank_id', 'N/A')
drug_info['name'] = drug.get('drug_name', drug_name)
drug_info['description'] = drug.get('description', 'N/A')[:200]
drug_info['approval_groups'] = drug.get('approval_groups', [])
print(f" ✅ Found: {drug.get('drug_name')}")
else:
print(f" ℹ️ Not found in DrugBank (may be novel compound)")
drug_info['name'] = drug_name
drug_info['status'] = 'Novel or not in database'
except Exception as e:
print(f" ⚠️ Error: {e}")
drug_info['name'] = drug_name
drug_info['error'] = str(e)
# Get pharmacology
print(f" Getting pharmacology...")
try:
result = self.tu.tools.drugbank_get_pharmacology_by_drug_name_or_drugbank_id(
query=drug_name, # ✅ CORRECT
case_sensitive=False,
exact_match=False,
limit=1
)
if result.get('data', {}).get('drugs'):
pharm = result['data']['drugs'][0]
drug_info['mechanism'] = pharm.get('mechanism_of_action', 'N/A')[:150]
print(f" ✅ Retrieved mechanism")
except Exception as e:
print(f" ⚠️ Error: {e}")
return drug_info
def _find_precedent_trials(self, indication, drug_name):
"""Search ClinicalTrials.gov for precedent trials."""
print("\n3️⃣ Precedent Trial Search")
precedents = []
print(f" Searching ClinicalTrials.gov...")
try:
result = self.tu.tools.search_clinical_trials(
condition=indication,
intervention=drug_name,
max_results=10
)
if result.get('data', {}).get('trials'):
trials = result['data']['trials']
precedents = [
{
'nct_id': trial.get('nct_id', 'N/A'),
'title': trial.get('title', 'N/A')[:80],
'status': trial.get('status', 'N/A'),
'phase': trial.get('phase', 'N/A')
}
for trial in trials[:5]
]
print(f" ✅ Found {len(trials)} precedent trials")
else:
print(f" ℹ️ No precedent trials found")
except Exception as e:
print(f" ⚠️ Error: {e}")
return precedents
def _assess_safety(self, drug_name):
"""Assess drug safety profile."""
print("\n4️⃣ Safety Assessment")
safety = {}
# Check DrugBank safety data
print(f" Checking safety data...")
try:
result = self.tu.tools.drugbank_get_safety_by_drug_name_or_drugbank_id(
query=drug_name, # ✅ CORRECT
case_sensitive=False,
exact_match=False,
limit=1
)
if result.get('data', {}).get('drugs'):
safety_data = result['data']['drugs'][0]
safety['toxicity'] = safety_data.get('toxicity', 'N/A')[:150]
print(f" ✅ Retrieved safety data")
except Exception as e:
print(f" ⚠️ Error: {e}")
# Check FDA warnings
print(f" Checking FDA warnings...")
try:
result = self.tu.tools.FDA_get_warnings_and_cautions_by_drug_name(
drug_name=drug_name
)
if result.get('data'):
warnings = result['data'].get('warnings', [])
safety['fda_warnings_count'] = len(warnings)
print(f" ✅ Found {len(warnings)} FDA warnings")
except Exception as e:
print(f" ⚠️ Error: {e}")
safety['fda_warnings_count'] = 0
return safety
def _search_literature(self, indication, drug_name):
"""Search PubMed for evidence."""
print("\n5️⃣ Literature Evidence")
literature = {}
query = f'"{indication}"[Title/Abstract] AND "{drug_name}"[Title/Abstract]'
print(f" PubMed search: {query[:60]}...")
try:
result = self.tu.tools.PubMed_search_articles(
query=query,
max_results=20
)
if isinstance(result, dict) and result.get('data', {}).get('articles'):
articles = result['data']['articles']
literature['article_count'] = len(articles)
literature['top_articles'] = [
{
'title': art.get('title', 'N/A')[:80],
'pmid': art.get('pmid', 'N/A')
}
for art in articles[:3]
]
print(f" ✅ Found {len(articles)} articles")
else:
literature['article_count'] = 0
print(f" ℹ️ No articles found")
except Exception as e:
print(f" ⚠️ Error: {e}")
literature['article_count'] = 0
return literature
def _calculate_feasibility(self, report):
"""Calculate feasibility score (0-100)."""
print("\n6️⃣ Feasibility Scoring")
score = 0
# Disease information available: +20
if report['disease_info'].get('disease_id'):
score += 20
print(f" ✅ Disease identified: +20")
# Drug information available: +20
if report['drug_info'].get('drugbank_id'):
score += 20
print(f" ✅ Drug in database: +20")
# Precedent trials exist: +30
if len(report['precedent_trials']) > 0:
score += 30
print(f" ✅ Precedent trials found: +30")
# Safety data available: +15
if report['safety_data']:
score += 15
print(f" ✅ Safety data available: +15")
# Literature evidence: +15
if report.get('literature', {}).get('article_count', 0) > 5:
score += 15
print(f" ✅ Strong literature: +15")
print(f" 📊 Total Score: {score}/100")
return score
def _interpret_feasibility(self, score):
"""Interpret feasibility score."""
if score >= 75:
return "HIGH FEASIBILITY - Strong precedent and data available"
elif score >= 50:
return "MODERATE FEASIBILITY - Some gaps but viable"
elif score >= 25:
return "LOW FEASIBILITY - Significant challenges"
else:
return "VERY LOW FEASIBILITY - Major gaps in data/precedent"
def main():
"""Run trial feasibility examples."""
print("=" * 80)
print("CLINICAL TRIAL FEASIBILITY PIPELINE - FIXED VERSION")
print("=" * 80)
print()
analyzer = TrialFeasibilityAnalyzer()
# Example 1: EGFR inhibitor in NSCLC
print("\n" + "=" * 80)
print("EXAMPLE: EGFR Inhibitor for EGFR-mutant NSCLC")
print("=" * 80)
report = analyzer.analyze(
indication="EGFR-mutant non-small cell lung cancer",
drug_name="osimertinib",
phase="Phase 2"
)
print("\n" + "=" * 80)
print("✅ PIPELINE COMPLETE")
print("=" * 80)
print(f"\n📄 Report: Trial_Feasibility_osimertinib.md")
print(f"📊 Feasibility Score: {report['feasibility_score']}/100")
print(f"\n💡 Trial design skill is now functional!")
if __name__ == "__main__":
main()
Clinical Trial Design - Quick Start Guide
Status: ✅ WORKING - Pipeline fixed and tested Last Updated: 2026-02-09
---
Choose Your Implementation
Python SDK
Option 1: Use the Working Pipeline (RECOMMENDED)
# Import from either file (both work)
from python_implementation import TrialFeasibilityAnalyzer
# or: from trial_pipeline import TrialFeasibilityAnalyzer
# Initialize analyzer
analyzer = TrialFeasibilityAnalyzer()
# Analyze trial feasibility
report = analyzer.analyze(
indication="EGFR-mutant non-small cell lung cancer",
drug_name="osimertinib",
phase="Phase 2"
)
# Report automatically saved to: Trial_Feasibility_osimertinib.md
print(f"Feasibility Score: {report['feasibility_score']}/100")Option 2: Use Individual Tools
from tooluniverse import ToolUniverse
tu = ToolUniverse()
tu.load_tools()
# Disease information (Open Targets)
result = tu.tools.OpenTargets_get_disease_id_description_by_name(
disease_name="non-small cell lung cancer"
)
# Drug profile (DrugBank)
result = tu.tools.drugbank_get_drug_basic_info_by_drug_name_or_id(
query="osimertinib", # ✅ Correct parameter
case_sensitive=False,
exact_match=False,
limit=1
)
# Pharmacology (DrugBank)
result = tu.tools.drugbank_get_pharmacology_by_drug_name_or_drugbank_id(
query="osimertinib", # ✅ Correct parameter
case_sensitive=False,
exact_match=False,
limit=1
)
# Safety data (DrugBank)
result = tu.tools.drugbank_get_safety_by_drug_name_or_drugbank_id(
query="osimertinib", # ✅ Correct parameter
case_sensitive=False,
exact_match=False,
limit=1
)
# Precedent trials (ClinicalTrials.gov)
result = tu.tools.search_clinical_trials(
condition="EGFR-mutant non-small cell lung cancer",
intervention="osimertinib",
max_results=10
)
# FDA warnings (FDA)
result = tu.tools.FDA_get_warnings_and_cautions_by_drug_name(
drug_name="osimertinib"
)
# Literature evidence (PubMed)
result = tu.tools.PubMed_search_articles(
query='"EGFR-mutant NSCLC" AND "osimertinib"',
max_results=20
)
# Prevalence data (PubMed)
result = tu.tools.PubMed_search_articles(
query='"EGFR-mutant NSCLC" AND "prevalence"',
max_results=5
)---
MCP (Model Context Protocol)
Option 1: Conversational (Claude Desktop or Compatible Client)
Tell Claude:
"Analyze clinical trial feasibility for osimertinib in EGFR-mutant NSCLC using ToolUniverse"
Claude will follow the workflow from SKILL.md and use these tools: 1. OpenTargets_get_disease_id_description_by_name - Disease identification 2. drugbank_get_drug_basic_info_by_drug_name_or_id - Drug profile 3. drugbank_get_pharmacology_by_drug_name_or_drugbank_id - Mechanism 4. search_clinical_trials - Precedent trials 5. drugbank_get_safety_by_drug_name_or_drugbank_id - Safety profile 6. FDA_get_warnings_and_cautions_by_drug_name - FDA warnings 7. PubMed_search_articles - Literature evidence
Option 2: Direct Tool Calls
Step 1: Disease Identification
Tool: OpenTargets_get_disease_id_description_by_name
Parameters:
{
"disease_name": "EGFR-mutant non-small cell lung cancer"
}Step 2: Drug Profile
Tool: drugbank_get_drug_basic_info_by_drug_name_or_id
Parameters:
{
"query": "osimertinib",
"case_sensitive": false,
"exact_match": false,
"limit": 1
}Step 3: Pharmacology & Mechanism
Tool: drugbank_get_pharmacology_by_drug_name_or_drugbank_id
Parameters:
{
"query": "osimertinib",
"case_sensitive": false,
"exact_match": false,
"limit": 1
}Step 4: Precedent Trials
Tool: search_clinical_trials
Parameters:
{
"condition": "EGFR-mutant non-small cell lung cancer",
"intervention": "osimertinib",
"max_results": 10
}Step 5: Safety Assessment (DrugBank)
Tool: drugbank_get_safety_by_drug_name_or_drugbank_id
Parameters:
{
"query": "osimertinib",
"case_sensitive": false,
"exact_match": false,
"limit": 1
}Step 6: FDA Warnings
Tool: FDA_get_warnings_and_cautions_by_drug_name
Parameters:
{
"drug_name": "osimertinib"
}Step 7: Literature Evidence
Tool: PubMed_search_articles
Parameters:
{
"query": "\"EGFR-mutant NSCLC\" AND \"osimertinib\"",
"max_results": 20
}Step 8: Prevalence Data
Tool: PubMed_search_articles
Parameters:
{
"query": "\"EGFR-mutant NSCLC\" AND \"prevalence\"",
"max_results": 5
}---
Run Examples (Python SDK)
# Run the working pipeline
python trial_pipeline.py
# Generates report:
# - Trial_Feasibility_osimertinib.md---
What Works ✅
- ✅ Disease identification (Open Targets)
- ✅ Drug profiling (DrugBank - correct parameters)
- ✅ Pharmacology data (DrugBank)
- ✅ Safety assessment (DrugBank + FDA)
- ✅ Precedent trial search (ClinicalTrials.gov)
- ✅ Literature evidence (PubMed)
- ✅ Prevalence estimation (PubMed proxy)
- ✅ Feasibility scoring (0-100 scale)
- ✅ Report generation (markdown)
- ✅ Clinical interpretation
---
Pipeline Analysis Steps
The pipeline performs 6-step analysis:
1. Patient Population Analysis
- Disease identification (Open Targets)
- Prevalence estimation (PubMed literature)
2. Drug Profile Analysis
- Drug identification (DrugBank)
- Mechanism of action (DrugBank pharmacology)
3. Precedent Trial Search
- Similar trials (ClinicalTrials.gov)
- Phase/status information
4. Safety Assessment
- Toxicity data (DrugBank)
- FDA warnings (FDA labels)
5. Literature Evidence
- Published studies (PubMed)
- Research support
6. Feasibility Scoring
- 0-100 score based on data availability
- Clinical interpretation
---
Feasibility Score Interpretation
- 75-100: HIGH FEASIBILITY - Strong precedent and data available
- 50-74: MODERATE FEASIBILITY - Some gaps but viable
- 25-49: LOW FEASIBILITY - Significant challenges
- 0-24: VERY LOW FEASIBILITY - Major gaps in data/precedent
---
Known Limitations
⚠️ Data Availability: Some tools return empty results:
- DrugBank may not include very new drugs (e.g., osimertinib)
- ClinicalTrials.gov API may have limited search results
- This is a data availability issue, not a code issue
⚠️ Novel Compounds: Experimental drugs may show low feasibility scores simply due to lack of historical data, not actual infeasibility
---
Tool Parameters (All Implementations)
These parameter names apply to both Python SDK and MCP:
| Tool | Parameter | Correct Name | Notes |
|---|---|---|---|
| drugbank_get_drug_basic_info_by_drug_name_or_id | Drug query | query | All DrugBank tools use this |
| drugbank_get_pharmacology | Drug query | query | NOT drug_name_or_drugbank_id |
| drugbank_get_safety | Drug query | query | NOT drug_name_or_drugbank_id |
| OpenTargets_get_disease_id | Disease name | disease_name | Returns Ensembl disease IDs |
| search_clinical_trials | Condition | condition | Separate from intervention |
| FDA_get_warnings_and_cautions | Drug name | drug_name | Simple string parameter |
| PubMed_search_articles | Search query | query | Supports PubMed query syntax |
Note: Whether using Python SDK or MCP, the parameter names are the same
---
Files
trial_pipeline.py- Complete working pipeline ✅python_implementation.py- Same pipeline (for consistency with other skills) ✅SKILL.md- Skill documentation (framework)EXAMPLES.md- Clinical scenarios (documentation)README.md- Original readmeQUICK_START.md- This file
---
Fixed: 2026-02-09 - Pipeline now uses correct ToolUniverse tool parameters
Clinical Trial Feasibility Report Template
All 14 sections below MUST be present in every feasibility report. File naming: [INDICATION]_trial_feasibility_report.md
---
1. Executive Summary
# Clinical Trial Feasibility Report: [INDICATION]
**Date**: [YYYY-MM-DD]
**Trial Type**: [Phase 1/2, biomarker-selected, basket, etc.]
**Primary Endpoint**: [ORR, PFS, DLT, etc.]
**Feasibility Score**: [0-100] - [LOW/MODERATE/HIGH]
## Key Findings
- **Patient Availability**: [Est. enrollable patients/year in US]
- **Enrollment Timeline**: [Months to target N]
- **Endpoint Precedent**: [Grade A/B/C/D] - [Description]
- **Regulatory Pathway**: [505(b)(1), breakthrough, orphan, etc.]
- **Critical Risks**: [Top 3 feasibility risks]
## Go/No-Go Recommendation
[RECOMMEND PROCEED / RECOMMEND ADDITIONAL VALIDATION / DO NOT RECOMMEND]
Rationale: [2-3 sentence summary]2. Disease Background
- Indication definition
- Prevalence and incidence (with sources)
- Current standard of care
- Unmet medical need
- Disease biology relevant to trial design
3. Patient Population Analysis
## 3.1 Base Population Size
- **US Incidence**: [X per 100,000] [evidence grade: Source]
- **Prevalence**: [Y total patients in US] [evidence grade: CDC/NCI data]
- **Annual new cases**: [Z patients/year]
## 3.2 Biomarker Selection Impact
- **Biomarker**: [e.g., EGFR L858R mutation]
- **Prevalence in disease**: [%] [evidence grade: ClinVar/COSMIC]
- **Geographic variation**: [Asian vs. Caucasian, etc.]
- **Testing availability**: [FDA-approved tests, CLIA labs]
## 3.3 Eligibility Criteria Funnel
| Criterion | Remaining Patients | % Retained |
|-----------|-------------------|------------|
| Base disease population | [N] | 100% |
| Biomarker positive | [N x biomarker %] | [%] |
| Age 18-75 | [N x age factor] | [%] |
| No prior therapy | [N x treatment-naive %] | [%] |
| ECOG 0-1 | [N x performance factor] | [%] |
| Adequate organ function | [N x eligibility factor] | [%] |
| **FINAL ELIGIBLE POOL** | **[N]** | **[%]** |
## 3.4 Geographic Distribution
- High-incidence regions: [e.g., Asia 50%, US 15% for EGFR+]
- Trial site implications
- Recruitment strategy recommendations
## 3.5 Enrollment Projections
**Assumptions**:
- Eligible pool: [N patients/year in US]
- Site activation: [M sites]
- Screening success rate: [%]
- Patients per site per month: [X]
**Target Enrollment**: [Total N]
**Projected Timeline**: [Months]
**Sites Required**: [Minimum M sites]4. Biomarker Strategy
## 4.1 Primary Biomarker
- **Biomarker**: [Gene mutation, protein expression, etc.]
- **Prevalence**: [%] [evidence grade: ClinVar data]
- **Assay Type**: [NGS, IHC, PCR, etc.]
- **FDA-Approved Tests**: [List CDx tests]
- **Turnaround Time**: [Days]
- **Cost**: [$X per test]
## 4.2 Alternative/Complementary Biomarkers
| Biomarker | Prevalence | Correlation | Testing |
|-----------|------------|-------------|---------|
| [Alt 1] | [%] | [R-squared] | [Method] |
| [Alt 2] | [%] | [R-squared] | [Method] |
## 4.3 Biomarker Testing Logistics
- Pre-screening vs. screening approach
- Central lab vs. local testing
- Tissue vs. liquid biopsy (ctDNA)
- Quality control requirements5. Endpoint Selection & Justification
## 5.1 Primary Endpoint
**Proposed**: [e.g., Objective Response Rate (ORR)]
**Regulatory Precedent** [evidence grade]:
- [N] FDA approvals in [indication] using ORR (2015-2024)
- Recent example: [Drug] approved [Year] (ORR XX%, n=YY)
- Source: search_clinical_trials, FDA_get_approval_history
**Measurement Feasibility**:
- Assessment method: [RECIST 1.1, irRECIST, etc.]
- Imaging modality: [CT, MRI, PET]
- Assessment frequency: [Every X weeks]
- Independent review: [Yes/No, cost]
**Statistical Considerations**:
- Expected ORR: [%] (based on [source])
- Null hypothesis: [%]
- Sample size: [N] (alpha=0.05, beta=0.20, two-sided)
- Response duration: [Median months]
## 5.2 Secondary Endpoints
| Endpoint | Evidence Grade | Feasibility | Rationale |
|----------|----------------|-------------|-----------|
| Progression-Free Survival (PFS) | A | High | FDA-accepted, precedent in [trials] |
| Duration of Response (DoR) | B | High | Standard in oncology |
| Overall Survival (OS) | A | Low (early phase) | Follow-up for long-term |
| [Biomarker response] | C | Medium | Exploratory, mechanistic |
## 5.3 Exploratory Endpoints
- Pharmacodynamic biomarkers (proof-of-mechanism)
- ctDNA clearance (liquid biopsy)
- Quality of life (PRO-CTCAE)
- Correlative science (tumor profiling)
## 5.4 Endpoint Risks & Mitigation
- Risk: [Low response rate -> sample size inflation]
- Mitigation: [Adaptive design, interim analysis]6. Comparator Analysis
## 6.1 Standard of Care
**Current SOC**: [Drug name(s)]
- FDA approval: [Year] [evidence grade: FDA_OrangeBook]
- Efficacy: [ORR/PFS from pivotal trial]
- Limitations: [Resistance, toxicity, access]
**SOC Comparator Feasibility**: [HIGH/MEDIUM/LOW]
## 6.2 Trial Design Options
### Option A: Single-Arm vs. SOC
- **Design**: Phase 2, single-arm, N=[X]
- **Comparator**: Historical SOC data (ORR=[%])
- **Pros**: Faster enrollment, smaller N
- **Cons**: Selection bias, regulatory skepticism
- **Feasibility Score**: [0-100]
### Option B: Randomized vs. SOC
- **Design**: Phase 2, 1:1 randomization, N=[X] per arm
- **Comparator**: Active control ([SOC drug])
- **Pros**: Robust comparison, regulatory preferred
- **Cons**: 2x enrollment, comparator sourcing
- **Feasibility Score**: [0-100]
### Option C: Non-Inferiority Design
- **Rationale**: [If aiming for better safety with similar efficacy]
- **Non-inferiority margin**: [delta = X%]
- **Sample size**: [N] (larger than superiority)
## 6.3 Comparator Drug Sourcing
- Commercial availability: [Yes/No]
- Patent status: [Generic available?]
- Cost: [$X per course]
- Stability and storage: [Requirements]7. Safety Endpoints & Monitoring Plan
## 7.1 Primary Safety Endpoint
**Dose-Limiting Toxicity (DLT)** [for Phase 1 component]:
- DLT definition: [Grade 3+ non-hematologic, Grade 4+ hematologic]
- DLT assessment period: [Cycle 1, 28 days]
- Dose escalation rule: [3+3, BOIN, mTPI]
## 7.2 Mechanism-Based Toxicities
**Drug Class**: [Kinase inhibitor, checkpoint inhibitor, etc.]
**Expected Toxicities** [evidence grade: FAERS, label data]:
| Toxicity | Incidence | Grade 3+ | Monitoring |
|----------|-----------|----------|------------|
| Diarrhea | 60% | 10% | Symptom diary, hydration |
| Rash | 40% | 5% | Dermatology consult PRN |
| Hepatotoxicity | 20% | 3% | LFTs weekly (cycle 1), then q3w |
| [Specific AE] | [%] | [%] | [Plan] |
**Data Source**: FAERS_search_reports (similar drugs), drugbank_get_pharmacology
## 7.3 Organ-Specific Monitoring
### Hepatic
- Baseline: LFTs, hepatitis panel
- Monitoring: AST/ALT/bili weekly (cycle 1), then q3w
- Stopping rule: ALT >5x ULN or bili >3x ULN
### Cardiac
- Baseline: ECG, ECHO if anthracycline history
- Monitoring: ECG q cycle, ECHO if symptoms
- Stopping rule: QTcF >500 ms, LVEF drop >15%
### Renal
- Baseline: Cr, eGFR, urinalysis
- Monitoring: Cr/eGFR q cycle
- Stopping rule: CrCl <30 mL/min
## 7.4 Safety Monitoring Committee (SMC)
- Composition: [3 independent experts: oncologist, toxicologist, biostatistician]
- Review frequency: [After every 6 patients, then quarterly]
- Stopping rules: [>=3 DLTs at dose level, >=2 drug-related deaths]8. Study Design Recommendations
## 8.1 Recommended Design
**Phase**: [1/2, 1b/2, 2]
**Design Type**: [Single-arm, randomized, basket, umbrella]
**Primary Objective**: [Assess safety and preliminary efficacy]
**Schema**:
[Indication + Biomarker]
-> Screening (Biomarker testing)
-> Enrollment
|-- [Phase 1 dose escalation: 3+3 design, N=12-18]
| Dose Levels: [X mg, Y mg, Z mg QD]
| DLT assessment: Cycle 1 (28 days)
+-- [Phase 2 expansion: Simon 2-stage, N=43]
Stage 1: N=13 (>=2 responses to proceed)
Stage 2: N=30 additional
Target ORR: 30% (H0: 10%, alpha=0.05, beta=0.20)
## 8.2 Eligibility Criteria
**Inclusion**:
- Age >=18 years
- Histologically confirmed [disease]
- [Biomarker] positive (central lab confirmed)
- Measurable disease per RECIST 1.1
- ECOG PS 0-1
- Adequate organ function
- [<=1 prior line for advanced disease]
**Exclusion**:
- Brain metastases (unless treated and stable)
- Prior [drug class] therapy
- Active infection, immunodeficiency
- Pregnancy/nursing
- Significant cardiovascular disease
## 8.3 Treatment Plan
- **Dosing**: [X mg PO QD, 28-day cycles]
- **Dose modifications**: [20% reductions for Grade 2+]
- **Duration**: Until progression, toxicity, or 24 months
- **Concomitant meds**: Supportive care allowed, restrictions on CYP3A4 inhibitors
## 8.4 Assessment Schedule
| Assessment | Screening | Cycle 1 | Cycles 2-6 | Cycles 7+ | EOT |
|------------|-----------|---------|------------|-----------|-----|
| History & PE | X | X | X | X | X |
| ECOG PS | X | X | X | X | X |
| Labs (CBC, CMP, LFT) | X | Weekly | q3w | q3w | X |
| Tumor imaging | X | - | q6w | q9w | X |
| ECG | X | - | q3w (if abnormal) | - | X |
| Biomarker (ctDNA) | X | C1D15 | q6w | - | X |
| AE assessment | - | Continuous | Continuous | Continuous | X |9. Enrollment & Site Strategy
## 9.1 Site Selection Criteria
**Required Capabilities**:
- [Biomarker] testing (or central lab partnership)
- Phase 1/2 experience
- GCP compliance, IRB approval
- Access to [patient population]
- Investigator publications in [indication]
**Geographic Distribution**:
- US sites: [N] (target regions: [high-incidence areas])
- International: [Consider Asia if biomarker enriched there]
## 9.2 Enrollment Projections
**Assumptions**:
- Screening rate: [X patients/site/month]
- Screen failure rate: [30%] (biomarker negative, eligibility)
- Enrollment rate: [Y patients/site/month]
**Timeline** (N=[total]):
| Milestone | Month | Cumulative Enrolled |
|-----------|-------|---------------------|
| First site activated | 0 | 0 |
| First patient enrolled | 1 | 1 |
| 25% enrollment | [M1] | [0.25N] |
| 50% enrollment | [M2] | [0.5N] |
| 75% enrollment | [M3] | [0.75N] |
| Last patient enrolled | [M4] | [N] |
| Primary analysis | [M4 + follow-up] | - |
**Sites Required**: [Minimum M sites to achieve timeline]
## 9.3 Recruitment Strategies
- Physician outreach: Academic consortia, tumor boards
- Patient advocacy groups: [Organization names]
- ClinicalTrials.gov listing (prominent, lay summary)
- Social media: Targeted ads in [indication] communities
- Referral network: Community oncologists10. Regulatory Pathway
## 10.1 FDA Pathway Selection
**Recommended**: [505(b)(1) / 505(b)(2) / Breakthrough / Orphan]
**Rationale**:
- [505(b)(1)]: New molecular entity, full development program
- [505(b)(2)]: [If relying on published safety data for similar drugs]
- **Breakthrough Therapy**: [If preliminary evidence of substantial improvement]
- Criteria: [X-fold ORR vs. SOC in early data]
- Benefits: Rolling review, frequent FDA meetings
- **Orphan Designation**: [If prevalence <200,000 in US]
- Benefits: 7-year exclusivity, tax credits, fee waivers
## 10.2 Regulatory Precedents
**Similar Approvals** [evidence grade]:
- [Drug A]: [Indication], [Year], [Endpoint used], [N=X], [ORR=Y%]
- [Drug B]: [Indication], [Year], [Accelerated approval -> full]
- Source: FDA_get_approval_history, drug labels
**FDA Guidance Documents**:
- [Relevant guidance title] (Year)
- Key recommendations: [e.g., ORR acceptable for Phase 2, confirmatory trial needed]
## 10.3 Pre-IND Meeting
**Recommended Topics**:
1. Primary endpoint acceptability (ORR vs. PFS)
2. Biomarker test qualification (CDx plan)
3. Comparator arm (single-arm acceptable?)
4. Pediatric study plan waiver
5. Safety monitoring plan
**Timing**: [3-4 months before IND submission]
## 10.4 IND Timeline
| Milestone | Month | Deliverable |
|-----------|-------|-------------|
| Pre-IND meeting request | -4 | Briefing package |
| Pre-IND meeting | -3 | FDA feedback |
| IND submission | 0 | Complete IND package |
| FDA 30-day review | 1 | Clinical hold or proceed |
| First patient dosed | 1-2 | After IND clearance |11. Budget & Resource Considerations
## 11.1 Cost Drivers
| Item | Cost Estimate | Notes |
|------|---------------|-------|
| Protocol development | $50-100K | CRO or internal |
| IND preparation | $100-200K | CMC, toxicology reports |
| Site activation | $50K/site x [M sites] | IRB, contracts |
| Patient recruitment | $200-500K | Advertising, patient navigation |
| [Biomarker] testing | $[X]/patient | Central lab, CDx |
| Imaging (RECIST) | $3-5K/scan x [N scans] | CT, independent review |
| Drug supply | [Depends on sponsor] | If not sponsor-provided |
| CRO monitoring | $100-300/hour | Site visits, SDV |
| Data management | $150-300K | EDC, database lock |
| Statistical analysis | $50-100K | SAP, CSR |
| **TOTAL (Phase 1/2)** | **$[X-Y]M** | [N patients, M sites] |
## 11.2 Timeline & FTE Requirements
**Duration**: [X months] (enrollment) + [Y months] (follow-up)
**Team**:
- Medical monitor: 0.5 FTE
- Project manager: 0.8 FTE
- Clinical operations: 0.3 FTE
- Data manager: 0.3 FTE
- Biostatistician: 0.2 FTE12. Risk Assessment
## 12.1 Feasibility Risks (High Priority)
| Risk | Likelihood | Impact | Mitigation |
|------|------------|--------|------------|
| Slow enrollment (biomarker screen fail) | HIGH | HIGH | Expand sites, allow alternative biomarkers, liquid biopsy screening |
| Low response rate (ORR <10%) | MEDIUM | CRITICAL | Interim futility analysis, lower null hypothesis, pivot to combination |
| Unexpected toxicity (>33% DLT rate) | LOW | CRITICAL | Conservative starting dose, adaptive escalation (BOIN), close SMC oversight |
| Comparator drug supply issues | MEDIUM | MEDIUM | Secure commercial supply early, generic sourcing |
| Regulatory pushback on single-arm design | MEDIUM | HIGH | Pre-IND meeting, plan for randomized Phase 2b if needed |
## 12.2 Scientific Risks
- Biomarker hypothesis unvalidated: [Correlative studies to de-risk]
- Patient heterogeneity: [Stratification by [factor]]
- Resistance mechanisms: [Serial biopsies for molecular profiling]13. Success Criteria & Go/No-Go Decision
## 13.1 Phase 1 Success Criteria (Go to Phase 2)
- [ ] <=33% DLT rate at RP2D
- [ ] >=50% patients achieve [PD biomarker response]
- [ ] No unexpected safety signals (Grade 5 AEs, new class effects)
- [ ] PK supports QD dosing
## 13.2 Phase 2 Interim Analysis (Simon Stage 1)
- **Enrollment**: 13 patients
- **Decision Rule**:
- >=2 responses (ORR >=15%) -> Proceed to Stage 2
- <2 responses -> Stop for futility
## 13.3 Phase 2 Final Success Criteria (Advance to Phase 3)
- [ ] ORR >=30% (95% CI lower bound >10%)
- [ ] Median DoR >=6 months
- [ ] PFS signal (HR <0.7 vs. historical SOC)
- [ ] Safety profile manageable (Grade >=3 AE <40%)
- [ ] Biomarker correlation with response (enrichment signal)
## 13.4 Feasibility Scorecard
| Dimension | Weight | Score (0-10) | Weighted | Grade |
|-----------|--------|--------------|----------|-------|
| **Patient Availability** | 30% | [X] | [0.30xX] | [grade] |
| - Base population size | - | [X] | - | [Source] |
| - Biomarker prevalence | - | [X] | - | [ClinVar data] |
| - Site access | - | [X] | - | [N sites feasible] |
| **Endpoint Precedent** | 25% | [X] | [0.25xX] | [grade] |
| - Regulatory acceptance | - | [X] | - | [FDA approvals using ORR] |
| - Measurement feasibility | - | [X] | - | [RECIST standard] |
| **Regulatory Clarity** | 20% | [X] | [0.20xX] | [grade] |
| - Pathway defined | - | [X] | - | [Breakthrough potential] |
| - Precedent approvals | - | [X] | - | [Similar indications] |
| **Comparator Feasibility** | 15% | [X] | [0.15xX] | [grade] |
| - SOC availability | - | [X] | - | [FDA-approved, generic] |
| - Historical data | - | [X] | - | [Published ORR: X%] |
| **Safety Monitoring** | 10% | [X] | [0.10xX] | [grade] |
| - Known toxicities | - | [X] | - | [FAERS, class effects] |
| - Monitoring plan | - | [X] | - | [Defined, feasible] |
| **TOTAL FEASIBILITY SCORE** | **100%** | - | **[XX/100]** | - |
**Interpretation**:
- **>=75**: HIGH feasibility - Recommend proceed to protocol development
- **50-74**: MODERATE feasibility - Additional validation recommended
- **<50**: LOW feasibility - Significant de-risking required14. Recommendations & Next Steps
## 14.1 Final Recommendation
**GO / CONDITIONAL GO / NO-GO**: [Decision]
**Rationale**: [2-3 paragraphs synthesizing feasibility analysis]
## 14.2 Critical Path to IND
**Immediate Next Steps** (Months 0-3):
- [ ] Request pre-IND meeting with FDA
- [ ] Initiate CDx partnership for [biomarker] test
- [ ] Secure drug supply (GMP manufacturing, stability)
- [ ] Draft protocol (v1.0) and ICF
- [ ] Site feasibility surveys (target [M] sites)
**IND Preparation** (Months 3-6):
- [ ] Complete CMC section
- [ ] Finalize preclinical package
- [ ] Prepare clinical protocol (incorporate FDA feedback)
- [ ] Develop CRFs and EDC database
- [ ] IND submission (Month 6)
**Post-IND** (Months 6-9):
- [ ] IRB submissions (central IRB for multi-site)
- [ ] Site contracts and budgets
- [ ] Investigator meeting
- [ ] First patient enrolled (Month 7-8)
## 14.3 Alternative Designs (If Current Design Infeasible)
**Plan B**: Broaden biomarker criteria, add international sites, basket design
**Plan C**: Randomized Phase 2 (if single-arm rejected by FDA)
## 14.4 Long-Term Development Strategy
- Phase 3 design, companion diagnostic, commercial readiness, patent strategy
- Market considerations: addressable market, competitive landscape, differentiation, pricingStudy Design Procedures
Detailed procedures for each of the 6 research paths in the clinical trial design feasibility assessment.
---
PATH 1: Patient Population Sizing
Steps
1. Disease prevalence lookup: Use OpenTargets_get_disease_id_description_by_name then OpenTargets_get_diseases_phenotypes_by_target_ensembl 2. Biomarker prevalence: Use ClinVar_search_variants (gene + significance) and gnomad_search_variants 3. Literature epidemiology: Use PubMed_search_articles for prevalence, incidence, geographic distribution 4. Enrollment feasibility: Use search_clinical_trials to find past trials and enrollment rates
Enrollment Funnel Calculation
Base disease population (incidence/year)
x Biomarker prevalence (%)
x Eligibility factors (age, PS, prior therapy, organ function): ~60%
/ Competing trials factor
= Available patients/yearGeographic Considerations
- Some biomarkers have ethnic variation (e.g., EGFR mutations: 15% Caucasian, 50% Asian)
- Consider international sites for biomarker-enriched populations
---
PATH 2: Biomarker Prevalence & Testing
Steps
1. Variant pathogenicity: ClinVar_get_variant_details 2. Cancer-specific frequency: COSMIC_search_mutations 3. Population genetics: gnomad_get_variant 4. CDx and testing: PubMed_search_articles for FDA-approved companion diagnostics, NCCN guidelines
Testing Logistics to Assess
- Pre-screening vs. screening approach
- Central lab vs. local testing
- Tissue vs. liquid biopsy (ctDNA)
- Turnaround time, cost, quality control
---
PATH 3: Comparator Selection
Steps
1. Drug info: drugbank_get_drug_basic_info_by_drug_name_or_id 2. Indications: drugbank_get_indications_by_drug_name_or_drugbank_id 3. Pharmacology: drugbank_get_pharmacology_by_drug_name_or_drugbank_id 4. Generic availability: FDA_OrangeBook_search_drug 5. Approval details: OpenFDA_get_approval_history 6. Historical controls: search_clinical_trials for completed trials with SOC
Design Options to Evaluate
- Single-arm vs. historical SOC: Faster, smaller N, but selection bias
- Randomized vs. SOC: Robust, regulatory preferred, but 2x enrollment
- Non-inferiority: When aiming for better safety with similar efficacy
---
PATH 4: Endpoint Selection
Steps
1. Precedent trials: search_clinical_trials (condition + phase + status=completed) 2. FDA acceptance: PubMed_search_articles for accelerated approval endpoints 3. Approval history: OpenFDA_get_approval_history for endpoints used in approvals
Endpoint Hierarchy (Oncology)
| Endpoint | Evidence Grade | Use Case |
|---|---|---|
| Overall Survival (OS) | A | Phase 3, confirmatory |
| Progression-Free Survival (PFS) | A | Phase 2/3 |
| Objective Response Rate (ORR) | A | Phase 2, accelerated approval |
| Duration of Response (DoR) | B | Secondary in Phase 2 |
| Biomarker response | C | Exploratory |
Statistical Considerations
- Expected effect size (from precedent trials)
- Null hypothesis (from SOC data)
- Sample size calculation (alpha=0.05, beta=0.20)
- Adaptive designs (Simon 2-stage, BOIN for dose escalation)
---
PATH 5: Safety Endpoints & Monitoring
Steps
1. Mechanism-based toxicity: drugbank_get_pharmacology_by_drug_name_or_drugbank_id (class drug) 2. FDA warnings: FDA_get_warnings_and_cautions_by_drug_name 3. Real-world AEs: FAERS_search_reports_by_drug_and_reaction, FAERS_count_reactions_by_drug_event 4. DLT definitions: PubMed_search_articles for Phase 1 DLT in similar drug class
Safety Monitoring Plan Components
- DLT definition and assessment period
- Organ-specific monitoring (hepatic, cardiac, renal)
- Safety Monitoring Committee composition and review frequency
- Stopping rules
Dose Escalation Designs
- 3+3: Simple, standard, conservative
- BOIN: Bayesian, more efficient, adaptive
- mTPI: Model-based, flexible
---
PATH 6: Regulatory Pathway
Steps
1. Breakthrough designations: PubMed_search_articles for FDA breakthrough in indication 2. Orphan drug eligibility: Calculate US prevalence (<200,000 total) 3. FDA guidance: PubMed_search_articles for relevant FDA guidance documents 4. Approval precedents: OpenFDA_get_approval_history for similar drugs
FDA Pathway Options
| Pathway | Criteria | Benefits |
|---|---|---|
| 505(b)(1) | New molecular entity | Full exclusivity |
| 505(b)(2) | Relies on published safety data | Faster, smaller safety package |
| Breakthrough Therapy | Substantial improvement on serious condition | Rolling review, frequent FDA meetings |
| Orphan Drug | Prevalence <200,000 in US | 7-year exclusivity, tax credits, fee waivers |
| Fast Track | Serious condition, unmet need | Rolling review |
| Accelerated Approval | Surrogate endpoint reasonably likely to predict benefit | Earlier approval, confirmatory trial required |
Pre-IND Meeting Topics
1. Primary endpoint acceptability 2. Biomarker test qualification (CDx plan) 3. Comparator arm (single-arm acceptable?) 4. Pediatric study plan waiver 5. Safety monitoring plan
IND Timeline
| Milestone | Month | Deliverable |
|---|---|---|
| Pre-IND meeting request | -4 | Briefing package |
| Pre-IND meeting | -3 | FDA feedback |
| IND submission | 0 | Complete IND package |
| FDA 30-day review | 1 | Clinical hold or proceed |
| First patient dosed | 1-2 | After IND clearance |
Related skills
How it compares
Pick this over general research skills when the deliverable must be a structured clinical trial feasibility report with FDA-grounded endpoint and enrollment analysis.
FAQ
What does tooluniverse-clinical-trial-design output?
tooluniverse-clinical-trial-design outputs a 14-section markdown feasibility report named `[INDICATION]_trial_feasibility_report.md`. The report includes a weighted 0–100 feasibility score, A–D evidence grades, enrollment projections, endpoint recommendations, and regulatory path
Which data sources does the clinical trial design skill use?
tooluniverse-clinical-trial-design queries ToolUniverse tools including OpenTargets, ClinVar, gnomAD, DrugBank, FDA Orange Book, OpenFDA, FAERS, PubMed, and ClinicalTrials.gov. The skill also supports direct ClinicalTrials.gov v2 API and FDA open-data calls when tool metadata is
What ToolUniverse version does clinical trial design require?
tooluniverse-clinical-trial-design version 1.0.0 requires ToolUniverse 0.5 or higher. The skill focuses on Phase 1/2 early clinical development and uses precedent-based reasoning rather than first-principles statistical derivation.