
Find Cohort Gap
- 47 installs
- 236 repo stars
- Updated August 3, 2026
- aperivue/medsci-skills
Find-cohort-gap is a Claude Code skill that mines a longitudinal cohort database for novel research topics by combining variable profiling, PI expertise matching, and literature-saturation scanning.
About
Find-cohort-gap discovers novel, publishable research topics from a longitudinal cohort database. A researcher uses it to profile cohort variables, match PI expertise, scan literature saturation, and produce ranked topic proposals with gap evidence. Unlike PICO or Elicit tools that work from literature to gaps, it works from the data outward.
- Research-gap finder that works from DB variables outward to a research question
- Profiles cohort strengths, matches PI expertise, and scans literature saturation
- Outputs ranked PICO topic proposals with novelty evidence and a discipline-alignment filter
Find Cohort Gap by the numbers
- 47 all-time installs (skills.sh)
- Ranked #944 of 2,064 Data Science & ML skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
find-cohort-gap capabilities & compatibility
- Capabilities
- design study · define variables · find journal
- Works with
- github
- Use cases
- research
- Pricing
- Free
What find-cohort-gap says it does
Research gap finder for longitudinal cohort databases.
This skill fills a gap that no existing tool addresses: **DB variables -> literature gap -> research question**.
Existing tools (PICO, FINER, SciSpace, Elicit) work from literature to gaps. This skill works from the data outward.
npx skills add https://github.com/aperivue/medsci-skills --skill find-cohort-gapAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 47 |
|---|---|
| repo stars | ★ 236 |
| Last updated | August 3, 2026 |
| Repository | aperivue/medsci-skills ↗ |
What it does
Discover novel, publishable research topics from a cohort database by scoring variable-by-literature gaps.
Who is it for?
Researchers with a cohort DB who need ranked, defensible study topics backed by novelty evidence.
Skip if: Working from a literature question toward a gap, which is what PICO or Elicit-style tools do.
When should I use this skill?
A researcher wants to systematically find new publishable topics from an existing cohort database.
What you get
Ranked PICO-format topic proposals with literature-saturation evidence and a first-author discipline-alignment filter.
- ranked PICO topic proposals with gap evidence
By the numbers
- generates 20-40 candidate topic statements
- user selects 8-12 candidates for saturation scanning
- 0-3 intersection-matrix scoring
Files
Find-Cohort-Gap Skill
You are assisting a medical researcher in systematically discovering novel, publishable research topics from a cohort database. Your approach combines cohort variable profiling, PI expertise matching, literature saturation scanning, and multi-pattern gap scoring to produce ranked topic proposals with evidence of novelty.
This skill fills a gap that no existing tool addresses: DB variables -> literature gap -> research question. Existing tools (PICO, FINER, SciSpace, Elicit) work from literature to gaps. This skill works from the data outward.
Communication Rules
- Communicate with the user in their preferred language.
- All literature citations, variable names, and medical terminology in English.
- Be direct about weak topics — kill early, save time.
Key Directories
- Output: User-specified directory (default: current working directory)
- References:
${CLAUDE_SKILL_DIR}/references/for templates and rubrics
---
Phase 0: Cohort Intake
Collect cohort metadata. Use the template at ${CLAUDE_SKILL_DIR}/references/cohort_profile_template.md.
Required information: 1. Cohort name and setting (institution, country, population type) 2. Sample size (N at baseline, N with follow-up) 3. Time span (enrollment period, follow-up duration, measurement intervals) 4. Variable categories (demographics, labs, imaging, questionnaires, medications, procedures) 5. Endpoints available (mortality, cancer incidence, cardiovascular events, hospitalization) 6. Special strengths (serial measurements, linkage to national registries, unique population) 7. Known limitations (healthy volunteer bias, attrition, missing data patterns) 8. Existing publications from this cohort (if known — to avoid duplication)
If the user provides a data dictionary file (Excel/CSV), read it to extract variable categories and construct the variable cluster map automatically.
Gate: Present the cohort profile summary. Confirm before proceeding.
---
Phase 1: PI/CA Profiling
Profile the intended PI or corresponding author to find topic-expertise alignment.
1. Search PubMed for the PI's recent publications (last 5 years).
- Use
/search-litE-utilities:bash "$EUTILS" search "AuthorLastName AuthorFirstInitial[Author]" 30 - Extract top keyword clusters from titles/abstracts.
2. Identify specialty signals:
- Academic society positions (president, board member, editor)
- Subspecialty focus areas
- Preferred journal tiers
3. Build a PI keyword map: 5-10 keyword clusters ranked by publication frequency.
If no PI is specified, skip this phase and use variable clusters alone in Phase 2.
Output: PI profile card (name, affiliation, top keywords, society roles, preferred journals).
---
Phase 2: Intersection Matrix
Cross cohort variable clusters with PI expertise to generate candidate topics.
Method
Create a matrix: rows = DB variable clusters, columns = PI keyword clusters. Score each cell 0-3:
- 3: PI has published in this exact intersection (direct match)
- 2: PI's subspecialty covers this area (strong relevance)
- 1: Tangential connection (possible but needs framing)
- 0: No connection
Candidate Generation
1. Extract all cells scoring 2-3 as primary candidates. 2. For cells scoring 1, apply the A-B substitution test: "Has someone published [this analysis] with [a different exposure/outcome] in a similar cohort?" If yes, substituting the PI's specialty variable creates a viable candidate. 3. Generate 20-40 candidate topic statements in PICO format:
- P: Population from the cohort
- E: Exposure/predictor variable(s)
- C: Comparison group
- O: Outcome (preferably hard endpoint)
Discipline Alignment Filter
Before advancing candidates to saturation scanning, apply a discipline filter:
- Who is the intended first author? Identify their department/specialty.
- Does the primary exposure variable belong to that discipline? The first
author's specialty must align with the study's core variable. For example:
- Radiology first author → imaging variable must be the primary exposure
- Cardiology first author → cardiac biomarker or ECG finding as exposure
- Neurology first author → neurological variable or brain imaging as exposure
- **Kill candidates where the primary exposure is outside the first author's
discipline.** A strong PI match alone is insufficient if the first author cannot claim ownership of the core variable.
This filter prevents generating topics where the first author's contribution is not defensible at the variable level.
Gate: Present the intersection matrix and top 20 candidates (post-discipline filter). User selects 8-12 for saturation scanning.
---
Phase 3: Literature Saturation Scan
For each selected candidate, determine how saturated the literature is.
Search Strategy
For each candidate: 1. Build a PubMed query: (exposure terms) AND (outcome terms) AND (cohort OR longitudinal OR prospective) 2. Execute search via /search-lit E-utilities. 3. Count total results and classify:
| Grade | Count | Longitudinal? | Interpretation |
|---|---|---|---|
| Blue Ocean | 0-2 papers | N/A | First report possible. Verify the topic has audience interest. |
| Green Field | 3-10 papers, all cross-sectional | No longitudinal | Optimal zone — established interest, longitudinal gap wide open. |
| Yellow | 10-30 papers | Some longitudinal | Viable only with very specific angle (unique population, novel endpoint). |
| Red | 30+ papers or MA exists | Yes | Avoid unless doing NMA or using truly unique data. |
Critical Filter
For each candidate in Green/Yellow, ask: "Has anyone published this with serial/repeated measurements?" If no — automatic upgrade by one grade.
"So What" Test
For each candidate, articulate 2-3 potential clinical implications of the findings. If you cannot state why a clinician or policymaker would care about the result, the topic fails regardless of gap score.
Output: Saturation table with grade, paper count, longitudinal gap status, and "So What" statement for each candidate.
Gate: Present saturation results. User selects 3-5 finalists for deep scoring.
---
Phase 4: 6-Pattern Scoring + Comparison Table
Apply the 6-Pattern framework to each finalist. Score each pattern 0 or 1.
6 Patterns (Universal)
Read the detailed rubric at ${CLAUDE_SKILL_DIR}/references/pattern_scoring_rubric.md.
| # | Pattern | Question | Score 1 if... |
|---|---|---|---|
| P1 | Longitudinal Advantage | Does the cohort's serial/repeated measurement structure create a clear edge over existing cross-sectional studies? | Cohort has 3+ timepoints for key variables AND no prior study used serial data for this topic. |
| P2 | Endpoint Upgrade | Can we escalate to a harder endpoint than existing studies? | Cohort links to mortality/cancer/CVD registries AND existing studies stop at surrogate endpoints. |
| P3 | Cohort Uniqueness | Is the cohort's population, scale, or setting distinctive? | Largest in this population, unique ethnic group, screening-based (no referral bias), or novel linkage. |
| P4 | PI-Topic Alignment | Does the PI's expertise and reputation strengthen this topic? | PI has society role or 5+ papers directly in this domain. Skip if no PI specified. |
| P5 | Comparison Table Gaps | Does the THIS STUDY column show 3+ differences vs existing papers? | Build comparison table (see below). 3+ checkmarks in THIS STUDY that are absent in all prior papers. |
| P6 | Complementary Design | Can this topic pair with another study from the same cohort? | Two studies using the same DB but different populations or complementary variables (e.g., viral vs non-viral). |
Comparison Table Construction
For each finalist, build a table comparing the top 3-5 existing papers against THIS STUDY:
| Feature | Author1 (Year) | Author2 (Year) | Author3 (Year) | THIS STUDY |
|---------|----------------|----------------|----------------|------------|
| Design | Cross-sectional | Cohort (5yr) | Cross-sectional | Cohort (20yr) |
| N | 3,200 | 8,500 | 12,000 | ~200,000 |
| Serial data | No | No | No | Yes (avg 5 visits) |
| Hard endpoint | Surrogate | Surrogate | All-cause mortality | CVD + all-cause mortality |
| Population | Referral | General | Screening | Health checkup (no referral bias) |
| Ethnicity | Western | Western | Asian (Japan) | Asian (Korea) |
| Subgroup analysis | No | Age only | No | Age + sex + comorbidity |Score Interpretation
| Total Score | Recommendation |
|---|---|
| 5-6 | Top-tier journal target (Lancet sub, JACC, J Hepatol level) |
| 3-4 | Specialty journal target (solid publication) |
| 1-2 | Restructure or kill — find a stronger angle before proceeding |
Gate: Present scoring results and comparison tables. User approves final ranking.
---
Phase 5: Feasibility Gate
For each scored finalist, verify practical feasibility.
Checks
1. Sample size adequacy:
- Cox regression: minimum 10 events per predictor variable (EPV rule)
- Logistic regression: same EPV rule
- For large cohorts (N>100K): warn about p-value inflation — statistically
significant results are nearly guaranteed, so focus on effect size thresholds (e.g., HR >1.2 or <0.8 for clinical relevance)
- Consider negative control strategy (EPCV) for very large samples
2. Missing data:
- Key exposure variable: <20% missing acceptable
- Key outcome: <5% missing
- If serial data: assess attrition pattern (MCAR/MAR/MNAR)
3. Follow-up adequacy:
- Outcome must have plausible latency within available follow-up
- Cancer outcomes: minimum 5 years
- CVD events: minimum 3 years
- Mortality: minimum 5 years
4. Operational definition:
- Can the exposure be defined from available variables?
- For claims data: ICD codes alone = 40-60% accuracy. Require combination
strategy (diagnosis + prescription + visit frequency + special codes)
- Cross-check expected prevalence against known epidemiological data
5. IRB/ethics:
- Is the data already IRB-approved for this type of analysis?
- Any additional approvals needed for data linkage?
6. Disease Novelty Bonus (informational, not Go/No-Go):
- Idiopathic etiology or debated mechanism → higher journal interest
- Established mechanism → needs stronger methodological novelty
Decision
- Go: All checks pass.
- Conditional Go: Minor issues solvable (e.g., missing data manageable with imputation).
- No-Go: Fatal flaw (insufficient events, no valid endpoint, key variable unavailable).
Output: Feasibility report for each finalist with Go/Conditional/No-Go status.
---
Phase 6: Output — Ranked Proposals + One-Pagers
Generate the final deliverables.
Ranked Summary Table
| Rank | Topic (PICO) | Saturation | 6-Pattern Score | Feasibility | Target Journal | Timeline |
|------|--------------|------------|-----------------|-------------|----------------|----------|
| 1 | ... | Green (0 longitudinal) | 5/6 | Go | JACC | 6 months |
| 2 | ... | Green (1 longitudinal) | 4/6 | Go | Eur Heart J | 6 months |
| 3 | ... | Blue (0 papers) | 3/6 | Conditional | Radiology | 8 months |One-Pager for Each Finalist
Use the template at ${CLAUDE_SKILL_DIR}/references/onepager_template.md.
Each one-pager includes: 1. Title: Working title for the study 2. Background: 3-4 sentences establishing the gap (with "Zero Papers" claim if applicable) 3. Comparison Table: THIS STUDY vs existing papers 4. Objective: Primary research question in PICO format 5. Methods Summary: Study design, key variables, statistical approach 6. PI Role: Why this PI is the right corresponding author 7. Target Journal: With rationale (PI alignment, scope match, gap fit) 8. Timeline: Realistic estimate (data preparation → analysis → drafting → submission) 9. 6-Pattern Score Card: Visual breakdown of each pattern
Save one-pagers as markdown files: {output_dir}/gap_proposal_{rank}_{short_topic}.md
---
Skill Integration
| Phase | Calls to other skills |
|---|---|
| Phase 1 (PI profiling) | /search-lit E-utilities for PubMed author search |
| Phase 3 (Saturation scan) | /search-lit E-utilities for topic searches |
| Phase 4 (Comparison table) | /search-lit for retrieving paper metadata |
| Downstream | Output feeds into /design-study → /write-paper pipeline |
What This Skill Does NOT Do
- Does not perform the actual statistical analysis (use
/analyze-stats) - Does not write the full manuscript (use
/write-paper) - Does not validate study design (use
/design-study) - Does not generate references (use
/search-lit) - Does not make publication-ready figures (use
/make-figures)
Anti-Hallucination
- Never fabricate references. All citations must be verified via
/search-litwith confirmed DOI or PMID. Mark unverified references as[UNVERIFIED - NEEDS MANUAL CHECK]. - Never invent clinical definitions, diagnostic criteria, or guideline recommendations. If uncertain, flag with
[VERIFY]and ask the user. - Never fabricate numerical results — compliance percentages, scores, effect sizes, or sample sizes must come from actual data or analysis output.
- If a reporting guideline item, journal policy, or clinical standard is uncertain, state the uncertainty rather than guessing.
Cohort Profile Template
Fill in the sections below to describe the cohort database. This profile drives the intersection matrix and feasibility checks.
---
Basic Information
- Cohort name:
- Institution/Organization:
- Country:
- Population type: (general population / health checkup / disease registry / claims data / hospital EMR)
- Enrollment period: (e.g., 2002-2019)
- Total N at baseline:
- N with follow-up data:
- Mean/median follow-up duration:
- Measurement intervals: (e.g., annual, biennial, at-event)
Variable Categories
Check all that apply and list key variables in each category:
- [ ] Demographics: (age, sex, BMI, smoking, alcohol, exercise, income, education)
- [ ] Laboratory: (CBC, metabolic panel, lipid panel, liver function, kidney function, tumor markers, HbA1c, ...)
- [ ] Imaging: (chest X-ray, CT, ultrasound, DEXA, mammography, ...)
- [ ] Questionnaires: (PHQ-9, IPAQ, diet, sleep, quality of life, ...)
- [ ] Vital signs: (BP, heart rate, ...)
- [ ] Anthropometry: (height, weight, waist circumference, body composition, ...)
- [ ] Medications: (prescription records, drug categories, ...)
- [ ] Procedures: (surgery codes, intervention records, ...)
- [ ] Diagnoses: (ICD codes, physician diagnosis, ...)
Endpoints Available
Check all that apply:
- [ ] All-cause mortality (linkage to: ___)
- [ ] Cause-specific mortality (categories: ___)
- [ ] Cancer incidence (linkage to: ___)
- [ ] Cardiovascular events (definition: ___)
- [ ] Hospitalization (source: ___)
- [ ] Disease incidence (ICD-based / physician-confirmed / registry)
- [ ] Other: ___
Special Strengths
What makes this cohort unique? (check all that apply)
- [ ] Serial measurements (same variables measured repeatedly over time)
- [ ] Large scale (>100K participants)
- [ ] Long follow-up (>10 years)
- [ ] National registry linkage (mortality, cancer, insurance claims)
- [ ] Screening-based (no referral bias — general population health checkups)
- [ ] Unique population (ethnicity, occupation, geography not well-studied)
- [ ] Rich phenotyping (imaging + labs + questionnaires)
- [ ] Biobank/genetic data available
- [ ] Other: ___
Known Limitations
- [ ] Healthy volunteer bias (participants may be healthier than general population)
- [ ] Attrition (estimated dropout rate: ___%)
- [ ] Missing data (key variables with >20% missing: ___)
- [ ] Limited demographics (e.g., single sex, narrow age range, single institution)
- [ ] Claims-only diagnoses (no clinical validation of ICD codes)
- [ ] No imaging data
- [ ] No medication data
- [ ] Other: ___
Existing Publications
List known papers already published from this cohort (to avoid topic duplication):
1. (Author, Year, Topic, Journal) 2. ...
Data Access
- IRB status: (approved / needs application)
- Access method: (on-site analysis center / remote access / direct download)
- Estimated turnaround: (application to data receipt)
- Cost: (if applicable)
---
Variable Cluster Map (Auto-generated)
If a data dictionary is provided, the skill will auto-generate clusters below:
| Cluster | Variables | Serial? | Endpoint Link? |
|---|---|---|---|
| (auto-filled) |
Research Topic Proposal — One-Pager
[Working Title]
Rank: #X of Y | 6-Pattern Score: X/6 | Saturation Grade: Green Field Target Journal: [Journal Name] | Estimated Timeline: X months
---
Background
[3-4 sentences establishing the clinical problem, current evidence gaps, and why this topic matters now. End with the "Zero Papers" claim if applicable: "To our knowledge, no study has examined [specific gap] using serial measurements in a [population type]."]
Comparison Table
| Feature | Author1 (Year) | Author2 (Year) | Author3 (Year) | THIS STUDY |
|---|---|---|---|---|
| Design | ||||
| N | ||||
| Serial data | ||||
| Hard endpoint | ||||
| Population | ||||
| Ethnicity | ||||
| Follow-up | ||||
| Key gap |
Unique differentiators: X (minimum 3 required)
Objective
Primary: [PICO format research question]
Secondary (optional): [1-2 secondary questions]
Methods Summary
- Study design: Retrospective cohort study
- Population: [Inclusion/exclusion criteria]
- Exposure: [Key variable(s) and operational definition]
- Outcome: [Primary endpoint and ascertainment method]
- Statistical approach: [Key methods — Cox regression, trajectory analysis, etc.]
- Sample size justification: [N eligible, expected events, EPV ratio]
6-Pattern Score Card
| Pattern | Status | Evidence |
|---|---|---|
| P1 Longitudinal Advantage | [+/-] | [one-line justification] |
| P2 Endpoint Upgrade | [+/-] | [one-line justification] |
| P3 Cohort Uniqueness | [+/-] | [one-line justification] |
| P4 PI-Topic Alignment | [+/-] | [one-line justification] |
| P5 Comparison Table 3+ | [+/-] | [one-line justification] |
| P6 Complementary Design | [+/-] | [one-line justification] |
PI Role
Corresponding Author: [Name, Title, Affiliation] Relevance: [Why this PI is the right CA — society role, expertise, journal connections]
Feasibility
- Go / Conditional Go / No-Go
- Sample size: N = X, expected events = Y, EPV = Z
- Key variables: Available / needs derivation / missing
- IRB: Covered / needs new application
- Data access: Ready / X weeks to obtain
Timeline
| Phase | Duration | Milestone |
|---|---|---|
| Data preparation | X weeks | Clean dataset, operational definitions |
| Analysis | X weeks | Primary + sensitivity analyses |
| Drafting | X weeks | Full manuscript |
| Internal review | X weeks | Co-author feedback |
| Submission | Target date | [Journal] |
Clinical Implications ("So What")
1. [Implication for clinical practice] 2. [Implication for screening/prevention policy] 3. [Implication for future research directions]
6-Pattern Scoring Rubric
Overview
Score each pattern 0 (absent) or 1 (present). Total: 0-6 points. Interpretation: 5-6 = top-tier, 3-4 = specialty journal, 1-2 = restructure or kill.
---
P1: Longitudinal Advantage
Question: Does the cohort's serial/repeated measurement structure create a clear edge over existing cross-sectional studies?
Score 1 if ALL of:
- The cohort has 3+ measurement timepoints for the key exposure variable
- No prior study on this topic used serial/trajectory data
- The research question benefits from temporal modeling (change over time, trajectory
clusters, time-to-event with time-varying exposure)
Score 0 if ANY of:
- The exposure is a one-time measurement (e.g., genetic variant, birth weight)
- Prior longitudinal studies already exist for this topic
- Serial data adds no interpretive value (e.g., stable demographic variable)
Examples:
- Score 1: Serial body composition → sarcopenia trajectory → mortality (no prior serial study)
- Score 0: Blood type → cancer risk (blood type doesn't change over time)
Theoretical basis: Repeated measures increase statistical efficiency by reducing within-subject variance and enabling trajectory-based phenotyping that cross-sectional designs cannot achieve (Lee et al., 2014, PMID 25464127).
---
P2: Endpoint Upgrade
Question: Can we escalate to a harder endpoint than existing studies?
Score 1 if BOTH of:
- The cohort links to mortality, cancer, or major cardiovascular event registries
- Existing studies on this topic used only surrogate endpoints (biomarkers, imaging
findings, composite scores) without hard clinical outcomes
Score 0 if ANY of:
- The cohort lacks hard endpoint linkage
- Prior studies already reported hard endpoints for this topic
- The research question is inherently about a surrogate (e.g., mechanism study)
Examples:
- Score 1: Existing studies link fatty liver to liver enzymes only; our cohort links to
liver-related mortality and HCC incidence
- Score 0: Existing studies already report all-cause mortality for this exposure
Endpoint hierarchy (strongest to weakest): 1. All-cause mortality 2. Cause-specific mortality 3. Major adverse events (MACE, cancer diagnosis) 4. Hospitalization 5. Disease incidence (physician diagnosis) 6. Surrogate markers (lab values, imaging scores)
---
P3: Cohort Uniqueness
Question: Is the cohort's population, scale, or setting distinctive?
Score 1 if ANY of:
- Largest cohort for this topic (>5x larger than existing studies)
- First study in this ethnic/geographic population
- Screening-based population (no referral bias) when prior studies used hospital cohorts
- Unique data linkage not available elsewhere (e.g., national registry + health checkup)
- Community-dwelling general population when prior studies used disease-specific cohorts
Score 0 if:
- Similar-sized cohorts with the same population type have published on this topic
Examples:
- Score 1: 486K health checkup participants vs existing studies of 3-9K referral patients
- Score 0: Another 500K cohort from the same country already published on this topic
---
P4: PI-Topic Alignment
Question: Does the PI's expertise and reputation strengthen this topic?
Score 1 if ANY of:
- PI holds a society leadership role directly relevant to the topic
- PI has 5+ first/corresponding author papers in this specific domain
- PI is an editorial board member of a target journal in this field
Score 0 if:
- PI's expertise is only tangentially related
- No specific PI identified (skip this pattern; score out of 5 instead)
Why this matters: A PI with society standing in the topic area signals that the study has expert oversight. Editors recognize this. The PI's name also guides target journal selection (e.g., hepatology society president -> J Hepatol).
When no PI is specified: Remove this pattern from scoring. Interpret: 4-5/5 = top-tier, 2-3/5 = specialty, 0-1/5 = restructure.
---
P5: Comparison Table Gaps (3+)
Question: Does the THIS STUDY column show 3+ unique features vs all existing papers?
Score 1 if:
- The comparison table has at least 3 rows where THIS STUDY has a checkmark/advantage
that NO prior paper has
Score 0 if:
- Fewer than 3 unique differentiators
Common differentiator categories: 1. Study design (longitudinal vs cross-sectional) 2. Sample size (order of magnitude larger) 3. Serial measurements (multiple timepoints vs single) 4. Hard endpoints (mortality vs surrogate) 5. Population type (screening vs referral) 6. Ethnicity/geography (first in this population) 7. Subgroup analyses (age/sex/comorbidity stratification) 8. Adjustment for key confounders (missing in prior studies) 9. Exposure definition (validated operational definition vs ICD-only) 10. Follow-up duration (significantly longer)
Construction method: 1. Identify 3-5 most relevant existing papers from saturation scan 2. Create table with Feature rows and Paper columns + THIS STUDY column 3. For each feature, check whether each paper and THIS STUDY address it 4. Count features unique to THIS STUDY
---
P6: Complementary Design
Question: Can this topic pair with another study from the same cohort?
Score 1 if ANY of:
- A complementary analysis using the same DB but different population subset is
feasible (e.g., diabetic vs non-diabetic; viral vs non-viral liver disease)
- The same exposure can be studied against a different outcome in a companion paper
- The topic creates a "series" with a previously published paper from the same cohort
Score 0 if:
- The topic is standalone with no natural complement
- The complementary analysis would be trivially similar (not publishable separately)
Why this matters: Paired papers from the same cohort strengthen both: the second paper can reference the first as "in this cohort, we previously showed..." and reviewers see a programmatic research line, not a one-off analysis.
---
Quick Reference Card
Pattern | Key Signal
----------------|------------------------------------------
P1 Longitudinal | "No prior study used serial data for this"
P2 Endpoint | "We add mortality/cancer to surrogate-only literature"
P3 Uniqueness | "Largest / first in this population / no referral bias"
P4 PI Alignment | "PI is society president in this exact field"
P5 Comparison | "3+ checkmarks unique to THIS STUDY"
P6 Complement | "Natural pair study exists in same DB"Literature Saturation Query Templates
Purpose
These templates help construct PubMed queries for the Phase 3 saturation scan. Adapt the bracketed terms to the specific topic.
---
Basic Saturation Query
([exposure MeSH] OR [exposure free text]) AND ([outcome MeSH] OR [outcome free text])
AND (cohort OR longitudinal OR prospective OR "follow-up")Filters: English, Humans, last 20 years (to capture the full landscape)
Longitudinal-Specific Query
To check if anyone has used serial/repeated measurements for this topic:
([exposure] OR [exposure synonym]) AND ([outcome] OR [outcome synonym])
AND ("repeated measure*" OR "serial" OR "trajectory" OR "longitudinal change"
OR "time-varying" OR "growth curve" OR "latent class trajectory")Meta-Analysis Check Query
To verify if a meta-analysis already exists:
([exposure] OR [exposure synonym]) AND ([outcome] OR [outcome synonym])
AND ("meta-analysis"[Publication Type] OR "systematic review"[Publication Type])If a meta-analysis exists → Red grade (avoid unless doing NMA).
Population-Specific Queries
Korean/Asian population filter
AND (Korea* OR Korean OR "Republic of Korea" OR Asia* OR Japan* OR China OR Chinese
OR Taiwan*)Health checkup / screening population filter
AND ("health checkup" OR "health screening" OR "health examination" OR "medical checkup"
OR "periodic health exam*" OR "annual exam*")Large cohort filter (to find comparator studies)
AND ("national health insurance" OR "claims data" OR "administrative data"
OR "population-based" OR "nationwide" OR "registry")---
Saturation Grading Protocol
After running the basic saturation query:
Step 1: Count total results
- 0-2: Blue Ocean
- 3-10: Possible Green Field (proceed to Step 2)
- 10-30: Possible Yellow (proceed to Step 2)
- 30+: Likely Red (check for MA in Step 3)
Step 2: Check longitudinal gap
Run the longitudinal-specific query.
- 0 results with serial/trajectory data → upgrade one grade
- 1-2 results → maintain current grade
- 3+ results → no upgrade
Step 3: Check meta-analysis existence
Run the MA check query.
- MA exists and is comprehensive → Red (firm)
- MA exists but outdated (>5 years) or limited scope → Yellow (update MA possible)
- No MA → maintain current grade
Step 4: Final grade assignment
| Base Count | Longitudinal Papers | MA Exists? | Final Grade |
|---|---|---|---|
| 0-2 | 0 | No | Blue Ocean |
| 3-10 | 0 | No | Green Field |
| 3-10 | 1-2 | No | Yellow |
| 10-30 | 0 | No | Green Field (upgraded) |
| 10-30 | 1-2 | No | Yellow |
| 10-30 | 3+ | No | Yellow |
| 30+ | Any | No | Yellow (borderline Red) |
| Any | Any | Yes (recent) | Red |
| Any | Any | Yes (outdated) | Yellow |
---
Example: Fatty Liver and Cardiovascular Mortality
Basic query
("fatty liver" OR "hepatic steatosis" OR NAFLD OR MASLD) AND
("cardiovascular mortality" OR "cardiac death" OR "MACE")
AND (cohort OR longitudinal OR prospective)Result: ~45 papers → base grade Red
Longitudinal check
("fatty liver" OR "hepatic steatosis") AND ("cardiovascular mortality")
AND ("trajectory" OR "serial" OR "repeated measure*" OR "longitudinal change")Result: 2 papers → no upgrade
MA check
("fatty liver" OR NAFLD) AND ("cardiovascular mortality")
AND ("meta-analysis"[PT] OR "systematic review"[PT])Result: 3 MAs → confirmed Red
Conclusion: Avoid this topic unless using truly unique data angle.
---
Tips for Effective Saturation Scanning
1. Start broad, then narrow. If the broad query returns >30, add population or design filters to find the exact niche.
2. Check the "last 3 years" subset. A topic with 20 total papers but 15 in the last 3 years is trending (good for timeliness, bad for novelty).
3. Read the most recent review article. It maps the field faster than scanning individual papers. Look for "future research directions" sections.
4. Check for registered protocols. Search PROSPERO or ClinicalTrials.gov for ongoing studies that haven't published yet — these are invisible competitors.
5. Use Semantic Scholar for citation network analysis. A paper with 200+ citations on this exact topic means the field is well-established.
schema_version: 2
name: find-cohort-gap
layer: D
owner_domain: research_gap_analysis
maturity: official
when_to_use: "Find research gaps in a longitudinal cohort DB by profiling its strengths, matching PI expertise, and scanning literature saturation."
when_NOT_to_use: "Finding meta-analysis topics (use ma-scout); designing a chosen study (use design-study)."
inputs:
- "cohort profile / data dictionary"
- "PI expertise"
- "literature landscape"
outputs:
- "ranked topic proposals with gap evidence"
side_effects:
- writes_report_artifacts
- network_access_literature
downstream_consumers:
- design-study
- define-variables
forbidden_actions:
- claim_a_gap_without_literature_evidence
- fabricate_saturation_counts
# v2.1 quality card
purpose: "Rank under-studied topics for a specific cohort, each backed by literature-saturation evidence and feasibility."
safety_boundaries:
- "Each proposed gap cites the literature scan that supports it; saturation counts come from real searches."
- "Advisory report only; does not design or execute studies."
known_limitations:
- "Literature scans are point-in-time; a gap can close between scan and submission."
- "No standalone demo; proposals require domain judgement."
validation_commands:
- "re-run the saturation search before committing to a topic"
evidence_surface: manual_workflow
Related skills
FAQ
How is this different from Elicit or SciSpace?
Those tools work from literature to gaps; find-cohort-gap works from the data outward, going DB variables to literature gap to research question.
What is the discipline-alignment filter?
It kills candidates where the primary exposure variable is outside the intended first author's specialty, so the first author can defend ownership of the core variable.