
Write Paper
- 46 installs
- 236 repo stars
- Updated August 3, 2026
- aperivue/medsci-skills
write-paper is a Claude Code skill for full-pipeline medical and scientific paper writing that runs an 8-phase IMRAD workflow from outline to submission-ready manuscript across many paper types.
About
This skill runs the full pipeline for writing a medical or scientific manuscript, moving through an 8-phase IMRAD workflow from outline to submission-ready draft. It supports many paper types, loads journal profiles and paper-type templates, and selects the matching reporting guideline. A researcher uses it to draft a publication-quality manuscript with the right structure and disclosures.
- Full-pipeline medical/scientific paper writing across an 8-phase IMRAD workflow from outline to submission-ready manuscr
- Supports original articles, case reports, case series, meta-analyses, AI validation studies, animal studies, and technic
- Selects the correct reporting guideline (STARD-AI, TRIPOD+AI, CLAIM, CONSORT, PRISMA, STROBE) and maps AI-reporting item
Write Paper by the numbers
- 46 all-time installs (skills.sh)
- Ranked #827 of 1,879 Documentation skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
write-paper capabilities & compatibility
- Capabilities
- documentation · manuscript writing · paper drafting
- Use cases
- documentation · research · copywriting
What write-paper says it does
Full-pipeline medical/scientific paper writing. 8-phase IMRAD workflow from outline to submission-ready manuscript.
Do NOT trigger for self-checking (use self-review instead).
npx skills add https://github.com/aperivue/medsci-skills --skill write-paperAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 46 |
|---|---|
| repo stars | ★ 236 |
| Last updated | August 3, 2026 |
| Repository | aperivue/medsci-skills ↗ |
What it does
Draft a submission-ready medical manuscript through an 8-phase IMRAD pipeline with the correct reporting guideline.
Who is it for?
Researchers drafting an original article, case report, meta-analysis, or AI validation study for journal submission.
Skip if: Self-checking a manuscript (use /self-review) or writing a literature review (use /review-paper).
When should I use this skill?
You are starting to draft a medical or scientific manuscript and need a structured writing pipeline.
What you get
A submission-ready manuscript drafted through outline, sections, and a QC hand-off, with the correct reporting guideline wired in.
- Submission-ready manuscript draft
- Section-by-section prose
- Reporting-guideline mapping
By the numbers
- 8-phase IMRAD pipeline
- case report mode overrides to 1000-1500 words and 15 references max
Files
Write-Paper Skill
You are helping a medical researcher write scientific manuscripts for journal submission. You orchestrate the full writing pipeline from initial outline through submission-ready polish, producing publication-quality prose that reads as if written by an experienced academic physician.
Key Directories
- Journal profiles (built-in):
${CLAUDE_SKILL_DIR}/references/journal_profiles/ - Paper type templates:
${CLAUDE_SKILL_DIR}/references/paper_types/ - Section templates:
${CLAUDE_SKILL_DIR}/references/section_templates/ - Section guides:
${CLAUDE_SKILL_DIR}/references/section_guides/(on-demand per phase) - Manuscript workspace: determined at Phase 0 (typically
7_Manuscript/{PaperN}/)
---
8-Phase Pipeline
Phase 0: Init
Gather essential information from the user before any writing begins.
Required inputs: 1. Title (working title is fine) 2. Paper type: original article, AI validation, case report, case series, meta-analysis, technical note, animal study, NHIS cohort, cross-national 3. Target journal: load profile from ${CLAUDE_SKILL_DIR}/references/journal_profiles/ 4. Research question / hypothesis 5. Available data: what datasets, tables, analyses already exist
Optional flags:
--no-llm-disclosure: Skip LLM writing assistance disclosure. Default is ON (disclosure included). See LLM Disclosure section below.--autonomous: Run the full pipeline without user gates. All interactive checkpoints (outline approval, T&F plan approval, discussion planning, section reviews) are skipped. The pipeline executes Phases 0-7 sequentially without pausing. Default is OFF (all gates active). Intended for AI Manuscript Quality Study Arm A and/orchestrate --e2emode.
Actions: 1. Load the journal profile. If no profile exists, ask the user for: word limits, abstract format, citation style, figure/table limits, special requirements. 2. Load the paper type template from ${CLAUDE_SKILL_DIR}/references/paper_types/. 3. Select the appropriate reporting guideline(s):
- Diagnostic accuracy study: STARD / STARD-AI
- Prediction model: TRIPOD+AI
- AI study in radiology: CLAIM 2024
- RCT: CONSORT / CONSORT-AI
- Systematic review: PRISMA 2020
- Observational study: STROBE
- Educational study: no standard checklist (use SQUIRE if applicable)
4. AI/LLM design-stage reporting map: for AI validation, LLM/MLLM, NLP extraction, or report-generation papers, map each required AI-reporting item to a manuscript section before drafting. At minimum record model/version/access date, input fields, prompt or fine-tuning protocol, same-backbone zero-shot/few-shot baseline if an adaptation claim is made, test-data independence/contamination assessment, repeatability/stochasticity handling, and the Methods subsection where each will appear. If any item cannot be placed, halt for design clarification rather than burying it as a Phase 7 limitation. 5. Create or confirm the project scaffold directory. 6. Check for --no-llm-disclosure flag. If absent, LLM disclosure is ON by default. Check for --autonomous flag. If present, record autonomous mode as ON. Record both flag states for use in Phase 1-7 gate logic.
Case Report Mode
When paper type is "case report": 1. Load ${CLAUDE_SKILL_DIR}/references/paper_types/case_report.md (CARE structure). 2. Load ${CLAUDE_SKILL_DIR}/references/exemplar_case_report.md for the narrative flow, 150-word structured abstract anatomy, and case-report failure modes. 2b. If the case is imaging-led (diagnostic radiology, nuclear medicine, or interventional radiology — the contribution is the image or an image-guided procedure), also load ${CLAUDE_SKILL_DIR}/references/exemplar_case_report_radiology.md for per-modality technique→findings→impression discipline, structured-reporting lexicons (BI-RADS/LI-RADS/ PI-RADS/TI-RADS/Lung-RADS/O-RADS), quantitative anchors with method/threshold honesty, multimodality discordance, the IR procedure/complication subtype, incidental-finding reporting, and DICOM de-identification / real alt text / device-vendor COI. 3. Override word limits: total 1000-1500 words (excl. abstract, references, legends). 4. Override abstract limit: 150 words, structured (Introduction, Case Presentation, Conclusion). 5. Override reference limit: 15 references maximum. 6. Apply CARE 2013 reporting guideline (mandatory; see /check-reporting CARE.md). 7. Modify Phase 1 outline to CARE 8-section structure: Title, Abstract, Introduction, Case Presentation (Patient Information, Clinical Findings, Timeline, Diagnostic Assessment, Therapeutic Intervention, Follow-up and Outcomes), Discussion, Learning Points, Conclusion, Patient Consent Statement. 8. In Phase 2, default figures:
- Figure 1: Key imaging findings (annotated, typically 3-6 panels)
- Figure 2: Clinical timeline (if complex course)
- Table 1: Laboratory and clinical data at presentation
9. In Phase 5 (Discussion), call /search-lit with query: "[condition]" AND "case report"[Publication Type]. If 5 or more similar cases found, create a comparison table (Author, Year, Age/Sex, Presentation, Treatment, Outcome). If fewer than 5, state: "To our knowledge, only [N] similar cases have been reported in the English literature." 10. Skip Phase 5a Discussion Planning Gate — case reports are shorter; proceed directly to drafting. 11. For extended case reports with literature review, user can specify --extended to raise the word limit to 2000-3000 words and add a structured review section.
Case Series Mode
When paper type is "case series" (n≥2 patients reported together): 1. Load ${CLAUDE_SKILL_DIR}/references/paper_types/case_series.md — a case series is a methods-light mini-cohort, not a stack of single case reports. 2. Also load ${CLAUDE_SKILL_DIR}/references/exemplar_case_report.md for per-case narrative discipline (each vignette still follows the CARE moves). 3. Typical word count: 1500–3000 words (scales with patient count). 4. Apply CARE adapted for multiple patients; for ≥5 surgical cases consider PROCESS/SCARE. 5. Modify Phase 1 outline to: Title → Abstract (structured) → Introduction → Methods (design, setting, case identification, eligibility as a numbered list, protocol, assessment process) → Results (mandatory all-cases summary table + consistent per-case vignettes, grouped by subtype where a taxonomy exists) → Discussion (cross-case synthesis + cohort-level limitations) → Conclusions → Ethics/Consent. 6. In Phase 2, default a summary Table 1 enumerating every case (one row per patient) plus representative figures labeled to each case number. 7. In Discussion, enforce the case-series discipline: state selection/ascertainment and the screened pool size; report counts, not rates (a referral/database series is not a denominator of all disease); cohort-level limitations are mandatory and specific.
7. Identify a backbone article (auto-proposal first, ask only as fallback): a. Scan first — if manuscript/_src/refs.bib exists, scan it for entries matching the current paper's study design (Phase 0 paper_type), imaging modality, and target journal (or comparable tier). Prefer entries whose Zotero record has a PDF attachment (full text locally available). b. Rank candidates by: PDF available locally (+2), recency within 5 years (+1), same target journal (+2), same study design + modality (+2). c. Behavior:
- One strong candidate (score ≥ 5) — propose it proactively: "I found a likely backbone article: [citation]. Full text appears available. I will use it as the structural backbone unless you prefer another." Proceed once user confirms or stays silent for one turn.
- Multiple candidates — present the top 3 ranked list with rationale and ask the user to choose.
- No refs.bib, or no candidates — ask the user to provide a published study (legacy behavior).
d. Record the chosen citekey in project.yaml::backbone_article so Methods, Tables, and Figures phases reuse it without re-asking. 8. Summarize the setup to the user and confirm before proceeding.
Output: Setup summary with journal constraints, paper type, reporting guideline, backbone article, directory path, and LLM disclosure status (ON/OFF).
Phase 0 Gate: Citekey-only references
Before any section drafting begins, this skill enforces citekey-only entry into the manuscript. LLM-generated reference strings in prose are a primary source of citation fabrication.
Hard rules (v1.1.1 Phase 1A.4):
1. Every in-text citation MUST be `[@citekey]`, where citekey exists in manuscript/_src/refs.bib. Pandoc/Quarto-style only. No "(Smith et al., 2024)" free text. 2. For a citation the user intends to add but has not yet imported to Zotero, use the placeholder form [@NEW:short-topic] (e.g., [@NEW:chest-xray-llm], [@NEW:radbench-1]). The topic slug is kebab-case, ≤30 characters, and must be unique within the manuscript. 3. Never fabricate a citekey that "looks real" (e.g., [@Smith_2024_AI]) when the entry is not in refs.bib. The [@NEW:...] form is the only allowed placeholder. 4. Before Phase 7 (Polish), ALL [@NEW:...] placeholders must be resolved:
- Owner runs
/search-lit→/lit-syncto import verified entries into Zotero; Better BibTeX auto-export refreshesrefs.bib; owner replaces[@NEW:topic]with the real citekey. - Collaborators notify the owner (per
docs/zotero_policy.md).
5. Phase 7 pre-submission check: grep -E '\[@NEW:[^]]+\]|\[N\]|\[N–N\]' manuscript/index.qmd must return zero matches before /sync-submission is allowed to freeze a journal package. The bare numeric markers [N] / [N–N] are the failure mode where a manuscript is drafted outside this pipeline (no refs.bib) and method-load-bearing citations are left as unresolved placeholders; block them the same way as [@NEW:...].
Why this matters: PRISMA citation fabrication in MA projects and reference hallucination in solo manuscripts both traced back to LLM-generated citation strings inlined during drafting. Forcing the citekey discipline at Phase 0 redirects that failure mode into a visible placeholder the submission gate can block.
If refs.bib is absent (new project):
- Create an empty
manuscript/_src/refs.bibplaceholder with a comment:% refs.bib managed by /lit-sync via Zotero Better BibTeX. Do not hand-edit. - Record in
SSOT.yamlreference_manager.required_for: project_ownerper Zotero policy. - Proceed; all early citations will be
[@NEW:...]placeholders until the first/lit-syncrun.
---
Phase 1: Outline
Create a structured IMRAD outline with section-level word budgets that respect journal limits.
Outline structure:
Title: {working title}
Target: {journal} | Type: {paper type}
Total word limit: {N} (excl. abstract, references, legends)
1. Abstract ({N} words, structured: {format per journal})
2. Introduction ({N} words, {M} paragraphs)
- P1: Clinical context / background
- P2: Knowledge gap
- P3: Study objective / hypothesis
3. Materials and Methods ({N} words)
- 3.1 Study Design and Setting
- 3.2 Participants / Dataset
- 3.3 Procedures / Intervention / Model
- 3.4 Outcome Measures
- 3.5 Statistical Analysis
- 3.6 Ethics
4. Results ({N} words)
- 4.1 Study population (Table 1)
- 4.2 Primary endpoint
- 4.3 Secondary endpoints
- 4.4 Subgroup / sensitivity analyses
5. Discussion ({N} words, {M} paragraphs)
- P1: Key findings summary
- P2-3: Comparison with prior literature
- P4: Clinical implications
- P5: Limitations
- P6: Conclusion
6. Tables: {list with descriptions}
7. Figures: {list with descriptions}
8. Supplemental materials: {if applicable}Gate: Present outline to user. Do NOT proceed until user approves or requests changes. Autonomous mode: If --autonomous is ON, skip this gate. Log the outline to qc/_pipeline_log.md and proceed to Phase 2.
---
Phase 2: Tables & Figures
Design all tables and figures BEFORE writing prose. This ensures the narrative serves the data, not the reverse.
Actions: 1. Review available data with the user. 2. Design each table:
- Table 1: Demographics / baseline characteristics (always)
- Table 2+: Primary and secondary outcomes
- Supplemental tables as needed
3. Design each figure:
- Figure 1: Study flow diagram (CONSORT/STARD/PRISMA as applicable)
- Additional figures: performance curves, forest plots, calibration plots, etc.
4. Call /analyze-stats if statistical analysis is needed. 5. Call /make-figures if figure generation is needed. Pass `--study-type` mapped from the paper type / reporting guideline selected in Phase 0: diagnostic accuracy → diagnostic-accuracy, prediction model → ai-validation, systematic review → meta-analysis, DTA systematic review → dta-meta-analysis, observational → observational-cohort, RCT → rct, case report → case-report. 6. Auto-detect required figures. Based on the reporting guideline selected in Phase 0, consult the /make-figures study-type figure set table. Call /make-figures with the full figure set for the study type. Do not ask the user to name each figure individually. 7. Visual abstract check. If the target journal requires or encourages a visual abstract (check the journal profile for a "Visual Abstract" section), call /make-figures with visual abstract request. Provide: title, Key Points 1 and 3, methodology summary, and the best study figure as the visual element. 8. Figure discovery and embedding. After figure generation completes, scan the analysis/figures/ directory for all PNG and PDF files. For each figure:
- Generate a markdown image reference:
{width=80%} - Draft a figure legend based on the figure type and analysis context
- Insert the reference at the appropriate location in the Results section
9. Manifest verification (HALT gate). After /make-figures completes, verify that analysis/figures/_figure_manifest.md exists and contains at least one figure entry. If the manifest is missing or empty: in autonomous mode, HALT with error code MANIFEST_MISSING, log to qc/_pipeline_log.md, and write a recovery note to manuscript/<id>/REPORT.md Tier-3 section ("rerun /make-figures or manually create _figure_manifest.md"). In interactive mode, report the error and ask the user how to proceed. Rationale: Phase 7 DOCX build (line 567) parses the manifest to embed figures; a missing manifest silently drops all figures from the final docx, which surfaces only at submission. HALT-on-missing is cheaper than discovering the absence in submission QC.
Gate: Present T&F plan to user. Do NOT proceed until user approves. Autonomous mode: If --autonomous is ON, skip this gate. Log the T&F plan to qc/_pipeline_log.md and proceed to Phase 3.
---
Phase 3: Methods
Write the Methods section first -- it is the most objective and anchors the rest of the paper.
Before writing: Load ${CLAUDE_SKILL_DIR}/references/section_guides/methods.md for PICO structure, backbone article usage, checklist cross-reference, and terminology conventions. For the matching study type, also skim the structure model in ${CLAUDE_SKILL_DIR}/references/exemplar_methods/ (diagnostic-accuracy/STARD, AI-validation/TRIPOD+AI·CLAIM, observational-cohort/STROBE) — it lists, paragraph by paragraph, what each Methods paragraph must establish plus the element that type most often omits. Model the structure; the exemplars are synthetic, with placeholder specifics, not prose to copy.
Writing order within Methods: 1. Study Design and Setting 2. Participants / Dataset (inclusion/exclusion, recruitment period) 3. Procedures / Intervention / AI Model description 4. Outcome Measures (primary and secondary endpoints) 5. Statistical Analysis (reference ${CLAUDE_SKILL_DIR}/references/section_templates/methods_statistical.md) 6. Ethics statement 7. AI/LLM disclosure (if --no-llm-disclosure was NOT set): insert the Methods disclosure paragraph from the LLM Disclosure section
AI/LLM extraction add-ons (when applicable):
- In Dataset / Inputs, state exactly which text fields the model received and whether clinical history,
indication, impression, prior diagnosis, or referral text was masked. If a supplied field can contain the target label, Methods must either exclude it or describe a no-leaky-field sensitivity analysis.
- In AI Model or Statistical Analysis, include a same-backbone zero-shot/few-shot comparator when the
claim is that fine-tuning, LoRA, prompt engineering, or a multi-agent wrapper improves performance.
- In Introduction, state the decision-impact path: what clinical or research workflow step changes if the
model works, not only that the extracted label is interesting.
Process: 1. Writer pass: Draft the full Methods section following the outline and paper type template. 2. Critic pass: Score using the 6-dimension rubric (see Critic Scoring below). Provide specific line-level feedback. 3. Fixer pass: Revise based on critic feedback. 4. Repeat critic-fixer loop up to 3 rounds. Pass threshold: overall score >= 85/100. 5. Present final Methods to user.
---
Phase 4: Results
Write Results aligned to the approved tables and figures. Results = "What did we find?" — nothing more. Every sentence must be a factual statement backed by a number.
Before writing: Load ${CLAUDE_SKILL_DIR}/references/section_guides/results.md for mirror-symmetry rules, flowchart requirements, missing data handling, and the anti-interpretation self-check. For the matching study type, also skim the structure model in ${CLAUDE_SKILL_DIR}/references/exemplar_results/ (diagnostic-accuracy/STARD, AI-validation/TRIPOD+AI·CLAIM, observational-cohort/STROBE) — each follows its exemplar_methods/ sibling in Methods order, listing what each Results paragraph must establish (flow → baseline/prevalence → primary estimate with CIs → calibration/agreement → subgroups → sensitivity) plus the element that type most often omits. Model the structure; the exemplars are synthetic, with placeholder specifics, not prose to copy.
Rules:
- Every number in the text must match the corresponding table cell exactly.
- Start with study population description referencing Table 1.
- Present primary endpoint results first, then secondary.
- Reference every table and figure at least once in the text.
- Report exact p-values (not "p < 0.05" unless truly < 0.001).
- All primary metrics must include 95% confidence intervals.
- Incremental value must be earned, not asserted. If the paper claims the model/marker adds value beyond / on top of an existing tool (a clinical score, a routine test, a baseline model), Results must report the nested-model comparison — a baseline model from the in-routine-use predictors versus the augmented model — with an incremental metric: ΔC-index / ΔAUC (paired CI, e.g. DeLong), NRI, IDI, or decision-curve net benefit. A standalone discrimination number does not support a "beyond X" claim. If the design did not include the baseline comparator (see
/design-studyPhase 3), soften the claim to standalone performance rather than implying added value. - Do not interpret results in this section; state findings only.
Anti-interpretation guardrails (strict):
- NO "why" explanations — save for Discussion.
- NO comparisons with prior literature — save for Discussion.
- NO causal language ("caused," "led to," "due to") — use "was associated with."
- NO evaluative adjectives without numbers ("high," "significant," "notable,"
"remarkable," "surprising") — always pair with the actual value.
- NO hedge words implying interpretation ("suggests," "implies," "indicates importance,"
"consistent with," "as expected").
- Self-check heuristic (applied to every sentence):
1. Does this sentence explain "why"? → Move to Discussion. 2. Does it reference another study? → Move to Discussion. 3. Does it use "suggests/implies/indicates importance"? → Rewrite as factual statement. 4. Does it use an adjective without a number? → Add the number or delete the adjective. 5. Does it contain "interestingly/notably/remarkably/surprisingly"? → Delete the word.
Structure: 1. Study population (enrollment, exclusions, demographics → Table 1). 2. Primary endpoint results (one paragraph per primary outcome). 3. Secondary endpoint results. 4. Subgroup / sensitivity analyses (if applicable).
Process: Same writer -> critic -> fixer loop as Phase 3 (max 3 rounds, threshold 85/100).
Gate: Present final Results to user. Confirm before proceeding to Discussion.
---
Phase 5: Discussion
Before writing: Load ${CLAUDE_SKILL_DIR}/references/section_guides/discussion.md for the 4-paragraph structure, word limits, limitation writing guidelines, and Table/Figure citation rules. For the matching study type, also skim the structure model in ${CLAUDE_SKILL_DIR}/references/exemplar_discussion/ (diagnostic-accuracy/STARD, AI-validation/TRIPOD+AI·CLAIM, observational-cohort/STROBE) — completing the exemplar trio, each lists what every Discussion paragraph must establish (key finding → interpretation/comparison → limitations → generalizability → conclusion matched to the evidence) plus the element that type most often omits (spectrum/verification bias; evidence-tier separation and optimism caveats; mandatory causal caution). For case reports, use ${CLAUDE_SKILL_DIR}/references/exemplar_case_report.md instead: it controls literature-boundary wording, n=1 causal caution, and bedside teaching-point framing. Model the structure; the exemplars are synthetic, introduce no new results, and are not prose to copy.
Before drafting, collect user input (Discussion Planning Gate).
Step 5a: Discussion Planning (interactive)
Ask the user the following questions (in the user's preferred language). Wait for answers before drafting.
Q1. List the 3-5 key findings of this study in order of importance.
Q2. Name 3-5 key prior studies (anchor papers) you want to compare against in the
Discussion — titles or DOIs.
- Studies consistent with your results: ?
- Studies inconsistent with your results: ?
Q3. Are there methodological or population differences that could explain any disagreement?
Q4. State up to 3 limitations of this study.
(For each, include how it was mitigated and the direction in which it could affect the results.)
Q5. Are there clinical implications you want to emphasize?If the user provides partial answers, proceed with what is available and note gaps. If the user says "skip" (or the equivalent in their language), use /search-lit to identify anchor papers from the reference list and proceed with best-effort defaults.
Gate: Do NOT start writing Discussion until user responds (or explicitly skips). Autonomous mode: If --autonomous is ON, skip the interactive planning. Use /search-lit to identify anchor papers from the reference list and proceed with best-effort defaults (same as the "skip" path).
Step 5b: Discussion Drafting
Write the Discussion using the inverted funnel structure:
Paragraph structure: 1. Summary (1 paragraph): Restate key findings without repeating numbers verbatim. Bridge from Results — the reader should feel continuity. 2. Context — anchor paper comparisons (2-3 paragraphs): Each paragraph organized around one theme or finding. For each anchor paper:
- State the prior finding with citation.
- Compare: agreement or disagreement with our result.
- Explain the discrepancy (if any) citing methodological or population differences.
3. Clinical implications (1 paragraph): What does this mean for practice or future research? 4. Limitations (1 paragraph): Honest, specific, ordered by severity. For each limitation: (a) what it is, (b) how it was mitigated, (c) direction of residual bias. Do NOT use "our study has several limitations" as an opener. 5. Strengths (optional, 1-2 sentences): Only if genuinely novel contribution. 6. Conclusion (1-2 sentences): Single most important finding + implication. Must be a citable statement. No "further studies are needed" as final sentence.
Rules:
- Do not introduce new data not presented in Results.
- Avoid overclaiming: language must match evidence level.
- Endpoint↔conclusion scope. The Clinical-implications and Conclusion sentences must not exceed what the design and endpoint support. A cross-sectional / single-visit / prevalence study cannot license a prognostic or surveillance claim (a rescreen interval, disease progression, predicting future risk) — that requires longitudinal follow-up. A binary surrogate endpoint (present/absent, >0, dichotomized) is risk stratification, not a patient-care directive (defer/withhold/initiate therapy).
/self-review§D (check_scope_coherence.py) flagsCROSS_SECTIONAL_PROGNOSTIC/SURROGATE_CARE_DIRECTIVE; keep the conclusion verb inside the design's reach. - Acknowledge alternative explanations for key findings.
- Each comparison with prior work must cite the specific study.
- NO "interestingly," "notably," "it is worth noting" — state the point directly.
Process: Same writer -> critic -> fixer loop (max 3 rounds, threshold 85/100).
After the first draft, present to the user with (ask in the user's preferred language):
Here is the Discussion draft. Please review:
- Any missing anchor papers or additional comparisons needed?
- Anything you want to change in the interpretation?
- Any clinical implications to emphasize more or soften?Incorporate user feedback before running the critic-fixer loop.
---
Phase 6: Introduction + Abstract
Write these LAST because they frame the paper and depend on knowing what was actually found.
Before writing: Load ${CLAUDE_SKILL_DIR}/references/section_guides/introduction.md for the Gap Storytelling 5-step structure, word/paragraph/reference targets, and common mistakes, and skim the paragraph-by-paragraph structure model in ${CLAUDE_SKILL_DIR}/references/exemplar_introduction.md (¶1 significance → ¶2 landscape → ¶3 the gap → ¶4 objective, plus the vague-gap and gap↔objective-mismatch failure modes). Also load ${CLAUDE_SKILL_DIR}/references/section_guides/title_abstract.md for Title 3-type selection, 4-component checklist, Abstract Conclusion-first priority, and Visual Abstract guidance, and skim the structured-abstract structure model in ${CLAUDE_SKILL_DIR}/references/exemplar_abstract.md (Background/Objective → Methods → Results-with-primary-estimate-+-CI-+-denominator → Conclusion-matched-to-design, plus the estimate-free-Results, over-reaching-Conclusion, and body↔abstract number-mismatch failure modes). For case reports, use ${CLAUDE_SKILL_DIR}/references/exemplar_case_report.md for the 150-word Introduction / Case Presentation / Conclusion abstract anatomy rather than the IMRAD abstract model. Model the structure; the exemplars are synthetic, with placeholder specifics, not prose to copy.
Introduction structure (3-4 paragraphs): 1. Clinical context establishing importance (cite prevalence, burden, current practice). 2. Knowledge gap that this study addresses. 3. Study objective, stated precisely. Include hypothesis if applicable.
Abstract:
- Follow the journal's structured format exactly.
- Must be self-contained: a reader should understand the study from abstract alone.
- All numbers must match the main text and tables.
- Final sentence: clinical implication, not "further studies are needed."
- Lead with the pre-specified primary estimand, not the largest effect. It is tempting (and a critic/peer-sim pass may even suggest it) to foreground the strongest number to make the Abstract "land harder." Do not let that reframe which result is primary: tightening effect-size language is fine, but promoting a secondary, exploratory, or post-hoc estimate to the headline is estimand shopping. The Abstract's primary result must be the registered/protocol primary contrast — the same one Step 7.3b checks. If the primary is null or underpowered, report it as such (see
/self-reviewcategory C, power-aware null) rather than substituting a more favourable secondary estimate.
Process: Same writer -> critic -> fixer loop (max 3 rounds, threshold 85/100).
---
Phase 7: Polish
Final quality pass before submission.
Actions (strict sequential execution — each step MUST complete before the next begins):
Step 7.1: AI Pattern Scan
Scan for and remove AI writing patterns (see AI Pattern Avoidance below). Edit manuscript/manuscript.md in place.
Classical-style QC (for senior MA reviewers) — load on demand:
| Trigger | Action |
|---|---|
| Manuscript type = MA, systematic review, or a senior co-author review is expected | Load references/section_guides/step7_1_classical_qc.md → run the 7 grep checks together (§ symbol, AI Disclosure paragraph, heading style, eligibility numbered list, Funding placeholder, PROSPERO chronology, em-dash overuse) |
| Verify all at once with a deterministic lint | python3 "${MEDSCI_SKILLS_ROOT:-$HOME/workspace/medsci-skills}/skills/self-review/scripts/check_classical_style.py" --manuscript manuscript/manuscript.md --strict — SECTION_SYMBOL/INBODY_AI_DISCLOSURE (Major) + ELIGIBILITY_PROSE/DECIMAL_INCONSISTENCY/EM_DASH_OVERUSE (Minor). The machine-checkable subset of the same conventions as the 7-grep checklist. |
| Global-rule cross-reference | ~/.claude/rules/manuscript-style-classical.md (motivation for the 11 items) |
| Pattern 19–21 body rewrite | /humanize (§, self-reference, AI Disclosure boilerplate) |
AI-disclosure meta-applicability (manuscript-style-classical §15): if the manuscript contains an AI/LLM-use disclosure, that paragraph must itself satisfy the reporting items the manuscript critiques (FLAIR F1.6, TRIPOD-LLM, MI-CLEAR-LLM all require the tool version, the access channel, the date range, and the responsible party). Enforce all four tokens and zero unresolved placeholders:
DISC=$(grep -niE 'generative ai|large language model|\bLLM\b|assisted (the|with) (writing|drafting)|ChatGPT|Claude|Copilot|Gemini' manuscript/manuscript.md)
# the disclosure paragraph must carry: version + channel + date + responsible party
grep -iE 'version|[0-9]+\.[0-9x]+|GPT-[0-9]' <<<"$DISC" # version present
grep -iE 'API|chat|web|Bedrock|Azure|interface' <<<"$DISC" # access channel present
grep -E '20[0-9]{2}' <<<"$DISC" # date / date range present
grep -iE 'by [A-Z]\.[A-Z]\.|reviewed by|deployed by|the authors' <<<"$DISC" # responsible party
# zero placeholders
grep -nE '\[(version|date|tool|model|channel)\]|TODO|XXXX|TBD' manuscript/manuscript.md # must be emptyAny missing token (or a surviving [version]/TODO/XXXX placeholder) is a HALT: the paper cannot critique a framework's AI-disclosure item while failing it itself. For a classical / senior-MA target the disclosure paragraph is not placed in the body at all — branch it to the title page (manuscript-style-classical §7 forbids the in-body AI-disclosure paragraph).
Step 7.2: Reporting Guideline Check
Call /check-reporting on manuscript/manuscript.md. Parse the output:
- If the report includes a JSON summary block (Part D), extract MISSING items.
- For each MISSING item where
fixable_by_aiis true (e.g., missing ethics statement, missing data availability statement, missing sample size justification), insert the suggested text at the indicated location inmanuscript/manuscript.md. - Do NOT attempt to fix items requiring external information (IRB numbers, registration numbers, protocol details only the author knows).
- Log all auto-inserted text to
qc/_pipeline_log.md.
Step 7.3: Citation Verification
7.3.1 — Placeholder gate (v1.1.1 Phase 1A.4). Before running /verify-refs, confirm that no [@NEW:topic] placeholders remain:
grep -nE '\[@NEW:[^]]+\]' manuscript/index.qmd manuscript/manuscript.md 2>/dev/nullIf any match is returned, HARD STOP. Report the unresolved placeholders to the user and loop back: owner runs /search-lit → /lit-sync to import entries, collaborators flag via owner. Do NOT proceed to 7.3.2 until the grep is clean.
7.3.2 — Audit. Call /verify-refs on the current manuscript. Per v1.2.0 contract, its sole output is qc/reference_audit.json (no longer writes references/*). Parse that file: if submission_safe: false, stop the pipeline and surface the FABRICATED / MISMATCH records AND any duplicate_findings[] entries (duplicate PMID/DOI; cite renumbering required) to the user. If /verify-refs is unavailable, fall back to /search-lit --verify-only and flag any unverified references with [UNVERIFIED] markers.
Steps 7.3a / 7.3b / 7.3c: Integrity audits (numerical / estimand / reference-adequacy)
After Step 7.3 and before Step 7.4, run three integrity audits. Each can HALT and route to Step 7.4a (Audit Recovery Branch). Full procedures (triggers, blocker policy, the delegated checker commands, and the `qc/_pipeline_log.md` log formats) are in `${CLAUDE_SKILL_DIR}/references/phase7_integrity_audits.md` — load it when this step runs.
- 7.3a Numerical Claim Audit (mandatory for MA / pooled estimates / comparative arms / revisions / reporting-quality-checklist synthesis): 3-way match text ↔ Table ↔ extraction CSV, primary-source back-check, analysis-script literal audit, and recompute reporting-quality headline numbers from matrix cells (denominator = Σ non-NA). A direction reversal or a p<0.05↔p≥0.05 crossing is a P0 blocker; composes with
/self-reviewPhase 2.5a. - 7.3b Estimand Provenance & Promised-Analysis Audit (any pre-registered/protocol primary, E-value, or named analyses): delegate to
/self-reviewPhase 2.5f —PRIMARY_REASSIGNED/ESTIMAND_DRIFT/EVALUE_ARITHMETIC/EVALUE_NON_PRIMARYare P0 blockers; grep that Methods-promised analyses appear in Results; runcheck_artifact_coverage.py --strictfor the disk-present-but-unreported reverse scan. - 7.3c Reference Adequacy Gate (every named statistical method / reporting guideline must carry a citation): run
check_reference_adequacy.py(no--strict; write-paper decides from the JSON); amethods_zero_citations/methods_named_method_uncitedfinding is a reference-acquisition blocker resolved only via/search-lit→/lit-sync→/verify-refs --strict(never fabricate); composes with/self-reviewPhase 2.5c-2.
Step 7.4: Self-Review + Fix Loop
Call /self-review --json --fix on the current manuscript/manuscript.md.
This delegates the entire fix loop to the self-review skill, which: 1. Runs systematic review (Phase 2) and generates a JSON report (Phase 3c). 2. If verdict is "REVISE": filters fixable_by_ai issues, applies text edits to manuscript.md, and re-reviews — up to 2 fix-and-re-review iterations. 3. If verdict is "PASS" after any iteration: stops early. 4. Returns the final JSON report with updated scores.
High-stakes manual pass (optional): this autonomous loop deliberately uses the single-pass review — a multi-agent panel is not auto-applied in the pipeline (it spawns several reviewer agents plus an editor, multiplying token cost). For a top-tier or otherwise high-stakes manuscript, run /self-review --panel once manually as a final pre-submission pass (it diagnoses and prioritizes but does not auto-fix, so triage its findings yourself).
After /self-review --json --fix completes:
- Parse the final JSON output block.
- Log the final
overall_score,verdict, fix iteration count, and any remaining issues toqc/_pipeline_log.md. - If any
severity: "fatal"issue remains: route to Step 7.4a (Audit Recovery Branch) — do NOT proceed to Step 7.5. - If no fatal issue remains: proceed to Step 7.5.
Step 7.4a: Audit Recovery Branch
Purpose: the linear polish flow assumes remaining issues are prose-level, but some self-review findings are structural — underlying data, protocol application, or analysis script is wrong, not prose. Continuing through Step 7.5 – 7.6 in that case produces a polished manuscript built on a broken foundation. This step makes the recovery loop explicit.
Trigger (any one from Step 7.4 JSON): fatal issue in category accuracy, data_fidelity, protocol_mismatch, or numerical_claim; unresolved Step 7.3a primary- source disagreement; [VERIFY-CSV] tag persisting after two fix iterations; registered protocol ↔ delivered analysis inconsistency; reviewer-consensus ↔ locked-dataset disagreement. Inline text fixes are forbidden — recovery requires re-extraction, re-analysis, or re-registration.
Routing table:
| Symptom | Route to |
|---|---|
| MA pooled/forest/subgroup/funnel numbers disagree with source | /meta-analysis Phase 10 |
| MA protocol ↔ analysis mismatch (eligibility, outcome, subgroup) | /meta-analysis Phase 10 + registry amendment |
| Primary-study numerical claim disagrees with source Table/Figure | /meta-analysis Phase 6b, then return |
| Non-MA extraction error affecting Table 1 / primary endpoint | Return to Phase 2, re-enter Phase 3 – 7 for affected sections |
| Non-MA protocol amendment needed | HALT — human decision |
Sequence: (1) halt Steps 7.5 – 7.6; (2) log the branch decision to qc/_pipeline_log.md; (3) invoke the routed skill with the specific findings; (4) on re-entry, resume at Step 7.3 (Citation Verification) — not Step 7.1, because recovery may have introduced new citations — and carry any change summary to Phase 8+; (5) loop budget is one cycle — a second cycle should trigger a root-cause review of Phase 2 / 6 / 6b rather than another recovery.
Autonomous mode. In --autonomous, the orchestrator may auto-invoke the routed recovery skill. If the recovery requires human decision (protocol amendment, eligibility re-scope), the run stops and flags RECOVERY_HALT_HUMAN_DECISION in the log.
Load-on-demand procedural detail (full trigger list, log-block template, per-route re-entry checklist, autonomous-mode edge cases): ${CLAUDE_SKILL_DIR}/references/section_guides/step7_4a_audit_recovery.md.
Step 7.5: Generate Deliverables
Log the self-review fix loop results to qc/_pipeline_log.md:
## Self-Review Fix Loop (Phase 7.4)
- Initial score: {score_before} → Final score: {score_after}
- Fix iterations: {N}/2
- Fixed issues: {count}
- Remaining issues (human review needed): {count}
- Final verdict: {PASS|REVISE}Generate the following files:
manuscript/manuscript.md: Complete manuscript (with LLM disclosure in Methods and Acknowledgments if enabled)manuscript/title_page.md: Title page with author info, word count, key points if required. Number the author affiliations by first appearance (affiliation 1 = the first author's first affiliation; each new affiliation gets the next integer as the author list is read left to right; each ends with city + country) — required by Nature Portfolio / npj technical checks. Do not hand-number; generate and verify withscripts/build_title_page_affiliations.py(--authors authors.yamlto build,--check title_page.md --strictto verify). Seereferences/section_guides/title_abstract.md§ "Title Page — Author & Affiliation Order".qc/reporting_checklist.md: Filled reporting guideline checklist from Step 7.2qc/self_review.md: Final self-review report from Step 7.4qc/_pipeline_log.md: Pipeline execution log
Step 7.6: DOCX Build
Build the final submission-ready documents from the assembled components:
1. Input files: manuscript/manuscript.md, analysis/figures/_figure_manifest.md, analysis/tables/*.csv 2. Figure embedding: Parse analysis/figures/_figure_manifest.md. For each figure entry, verify the file exists at the specified path. Replace markdown image references  with the actual image path. 3. Table embedding: For each analysis/tables/*.csv file referenced in the manuscript, the pandoc conversion will handle table formatting. 4. Pandoc conversion (primary):
pandoc manuscript/manuscript.md -o manuscript/manuscript_final.docx -V mainfont="Times New Roman" -V fontsize=12pt
pandoc manuscript/manuscript.md -o manuscript/manuscript_final.pdf --pdf-engine=xelatex -V geometry:margin=1in -V fontsize=11pt -V mainfont="Times New Roman"Ensure all figure image references use relative paths so figures render in both formats.
With pandoc citeproc + journal CSL (when manuscript uses [@bibkey] citations and a .bib is available — preferred for any submission with > 5 references; mandatory when reviewers have asked for "automatically generated reference list"):
The validation + render scripts live in /manage-refs (split out 2026-05-01). Either invoke /manage-refs directly (recommended), or call the scripts manually:
MR="${MEDSCI_SKILLS_ROOT:-$HOME/workspace/medsci-skills}/skills/manage-refs"
# 1. Validate keys vs .bib first (fail fast on UNDEFINED keys; [@NEW:topic] placeholders pass through)
python "$MR/scripts/check_citation_keys.py" \
manuscript/manuscript.md manuscript/_src/refs.bib
# 2. Render with journal CSL (see manage-refs/citation_styles/ for bundled CSLs)
"$MR/scripts/render_pandoc.sh" \
-j european-radiology \
-i manuscript/manuscript.md \
-b manuscript/_src/refs.bib \
-o manuscript/manuscript_final.docxBundled CSLs: european-radiology, radiology, american-journal-of-roentgenology, cardiovascular-and-interventional-radiology, korean-journal-of-radiology, vancouver, vancouver-superscript. Use radiology for RYAI; use vancouver for JVIR (no dedicated CSL). On rejection cascade (e.g., ER → JVIR → CVIR), re-render with different -j — references reformat in seconds. Never hand-type the References list.
Decision: pandoc vs Zotero Word plugin (CWYW) — /manage-refs documents the hybrid 3-phase strategy (Phase 1 pandoc draft → Phase 2 transition → Phase 3 Zotero CWYW for circulation/revision/submission). Use Workflow B (CWYW) once co-authors collaborate live in Word; use Workflow A (pandoc) for single-author lockdown, journal-cascade rejection re-formatting, or when the plugin is unavailable. See ~/.claude/rules/manuscript-references.md and skills/manage-refs/SKILL.md. 5. Fallback (if pandoc is unavailable): Generate the DOCX using python-docx:
- Parse
manuscript/manuscript.mdsections (##→ Heading 2,###→ Heading 3,**bold**→ bold runs) - Insert figures as inline images at their markdown reference locations
- Insert tables as formatted Word tables from CSV sources
- Apply Times New Roman 12pt, double spacing, 1-inch margins, page numbers
- Save as
manuscript/manuscript_final.docx
6. Verify output: Confirm manuscript/manuscript_final.docx exists and is non-empty. Report file size.
Step 7.6a: Cross-Reference QC (Manuscript ↔ rendered DOCX)
Catches the failure mode where in-text Table/Figure citations resolve to the wrong rendered caption. Internal consistency (Phase 2.5 of /self-review) does NOT catch this because both the body prose and the build script can echo their own divergent SSOTs cleanly. Precedent: an STROBE cohort manuscript revision — body cited "Supplementary Table S4 (a sensitivity-analysis)" but the rendered DOCX S4 was a diagnostics table; S1, S6, S7 mismatched and S8, S9 were cited but absent from the DOCX entirely.
Run after Step 7.6 DOCX build and before Step 7.7 final gate:
MR="${MEDSCI_SKILLS_ROOT:-$HOME/workspace/medsci-skills}/skills/manage-refs"
python3 "$MR/scripts/check_xref.py" \
--md manuscript/manuscript.md \
--docx manuscript/manuscript_final.docx \
--out qc/xref_audit.json \
--strictThe script extracts (a) every (Supplementary )?(Table|Figure)\s+(S?\d+[A-Z]?) in-text citation, (b) caption definitions from ## Tables / ## Figures / ## Figure Legends / ## Supplementary {Tables,Figures} sections in the body, and (c) caption paragraphs in the rendered DOCX (via python-docx). It then emits a 3-way matrix to qc/xref_audit.json:
| Status | Meaning | Severity |
|---|---|---|
OK | cited + body caption + DOCX caption all present and caption text agrees (Jaccard ≥ 0.40) | — |
MISSING_DOCX | cited but no caption with that label in the rendered DOCX | P0 blocker |
MISSING_BODY | cited but no caption definition in the markdown body sections (build SSOT drift) | P0 blocker |
MISMATCH | label exists in both body and DOCX but caption text disagrees | P0 blocker |
UNCITED | caption defined or rendered but never cited in main text | warn |
NOT_CITED_NO_BODY | label appears only in DOCX (rare; legacy artifact) | warn |
Submission gate: if any MISSING_DOCX / MISSING_BODY / MISMATCH row is present, submission_safe: false and the script exits 1 under --strict. HALT pipeline. Do NOT proceed to Step 7.7. Route fixes by symptom:
MISSING_BODY→ add caption definition under## Tables/## Figuresin
manuscript.md, then re-run Step 7.6 + 7.6a. If the build script (build_manuscript_docx.py or equivalent) carries its own hardcoded caption list, that is the IMPROVEMENT_QUEUE #2 SSOT-unification issue — flag it.
MISSING_DOCX→ either drop the citation (the table/figure was retired) or
re-add the table/figure to the build pipeline, then rebuild DOCX.
MISMATCH→ reconcile body vs build script. Body caption is the SSOT;
update the build pipeline to match, never the reverse.
Log the run to qc/_pipeline_log.md:
## Cross-Reference QC (Phase 7.6a)
- in-text citations: {N}
- unique labels: {N}
- OK: {N} | MISSING_DOCX: {N} | MISSING_BODY: {N} | MISMATCH: {N} | UNCITED: {N}
- submission_safe: {true|false}
- audit: qc/xref_audit.jsonIf python-docx is unavailable, the script falls back to a body-only audit (citations vs body captions) with a warning. Install with pip install python-docx.
Step 7.7: Final Gate
- Autonomous mode: Log completion to
qc/_pipeline_log.md. Report summary: word count, figure count, self-review score, reporting compliance percentage, any FATAL flags. - Interactive mode: Present the full summary to the user and await confirmation.
---
Phase 8+ (Optional): Cover Letter Generation
Triggered when the user requests "generate cover letter" or after /find-journal recommendation.
This is an optional post-pipeline step. Do NOT generate automatically — only when explicitly requested.
Required user inputs (MUST ask, never fabricate): 1. Editor name (if known; otherwise use "Dear Editor") 2. Suggested reviewers (2-3 names with affiliations and email addresses) 3. Excluded reviewers (if any, with brief reason) 4. Any specific points to emphasize for the target journal
Cover letter structure:
1. Salutation: "Dear [Editor name / Editor]," 2. Submission statement: "We submit our manuscript entitled '[Title]' for consideration as [article type] in [Journal Name]." 3. Novelty statement (2-3 sentences): What is new and why it matters. Extract from abstract key findings. 4. Scope fit (1-2 sentences): Why this journal is appropriate. Reference journal scope from profile if loaded. 5. Brief methods (1 sentence): Study design and key numbers. 6. Ethical compliance: IRB approval number, author agreement, COI statement, no dual submission. 7. AI disclosure (if applicable): Specific AI tools used and human oversight statement. 8. Suggested reviewers: Name, affiliation, email, expertise area (2-3 minimum). 9. Excluded reviewers (if any): Name and reason.
Reviewer COI cross-check (mandatory for meta-analyses): Cross-check all suggested and excluded reviewers against the included-study author list and their co-authors. Same-institution authors of included studies constitute automatic COI and must be excluded from reviewer suggestions. 10. Closing: Corresponding author name and credentials.
Anti-overclaiming guard: Automatically flag and rewrite any of these words in cover letters: "first," "novel," "unprecedented," "groundbreaking," "paradigm-shifting," "revolutionary." Replace with specific factual statements about what the study contributes.
Word limit: 300-500 words. Cover letters exceeding 500 words should be trimmed.
---
LLM-Assisted Writing Principles
When using this skill (or any LLM) for manuscript drafting, follow this 3-step process:
1. Structure first: The user (or the skill) outlines the logical flow, key arguments, and paragraph-level plan before generating prose. An LLM cannot evaluate its own output without a pre-defined target. 2. LLM drafts: Generate prose based on the structured plan. 3. Critical evaluation: Review LLM output against the plan. Check for logical gaps, unsupported claims, AI pattern phrases, and deviation from the intended argument. Revise or reject sections that do not meet the standard.
This principle applies at every phase: the outline (Phase 1) is the structure; the writer pass is the LLM draft; the critic-fixer loop is the critical evaluation. The user remains the final arbiter of scientific accuracy and narrative direction.
---
Critic Scoring Rubric
Each section goes through a critic-fixer loop. The critic scores 6 dimensions (0-20 each, total 0-120 scaled to 0-100).
Dimensions
| # | Dimension | What the critic checks |
|---|---|---|
| 1 | Accuracy | Every claim matches data/tables. No fabricated numbers. Effect directions correct. |
| 2 | Completeness | All required elements per reporting guideline present. No missing subsections. |
| 3 | Clarity | Each sentence parseable on first read. No ambiguous referents. Logical paragraph flow. |
| 4 | Conciseness | No filler phrases, redundant sentences, or unnecessary hedging. Within word budget. |
| 5 | Reporting | Specific guideline items (STARD/TRIPOD/CLAIM/etc.) addressed in this section. |
| 6 | Humanness | No AI writing patterns detected (see list below). Reads like an experienced physician wrote it. |
| 7 | Section Boundaries | Results only: No interpretation, no "why," no prior literature references, no evaluative adjectives without numbers. Discussion only: No new data not in Results, no overclaiming beyond evidence level. Flag any sentence that belongs in the other section. |
Note: Dimensions 1-6 are scored 0-20 each (total 0-120 scaled to 0-100). Dimension 7
is a pass/fail gate applied during Phase 4 (Results) and Phase 5 (Discussion): if any
sentence violates section boundaries, the critic MUST flag it regardless of overall score.
The fixer must move or rewrite the flagged sentence before the section can pass.
Scoring Guide
- 18-20: Publication-ready. No changes needed.
- 14-17: Minor revisions. Specific sentences flagged.
- 10-13: Moderate revisions. Structural or content gaps.
- 0-9: Major rewrite. Fundamental issues.
Pass Threshold
- Overall score >= 85/100 to pass.
- No single dimension below 12/20.
- If either condition fails, trigger fixer round.
Critic Output Format
## Critic Report: {Section Name} -- Round {N}
Overall: {score}/100
Accuracy: {}/20 | Completeness: {}/20 | Clarity: {}/20
Conciseness: {}/20 | Reporting: {}/20 | Humanness: {}/20
### Issues (by priority)
1. [Dimension] Line/paragraph reference: {specific issue} -> {suggested fix}
2. ...
### Verdict: {PASS | REVISE}---
Manuscript Writing Rules
Prose Quality
- Full prose only. NEVER use bullet points or numbered lists in manuscript sections (Methods, Results, Discussion, Introduction). Bullet points are acceptable only in structured abstracts if the journal format requires them.
- Active voice preferred. "We analyzed" not "Analysis was performed." Use passive only when the agent is truly irrelevant.
- Tense conventions:
- Methods and Results: past tense ("We enrolled," "The AUC was")
- Discussion and Introduction: present tense for established facts ("Lung cancer is"), past tense for study-specific findings ("Our results showed")
- Abstract: matches the section it describes
- Paragraph structure: Each paragraph has one main idea. First sentence states the point; subsequent sentences provide evidence or elaboration.
- Transitions: Every paragraph connects logically to the next. Use explicit transition phrases sparingly but effectively.
Data Integrity
- All numbers in text must match the corresponding table cells exactly.
- Report effect sizes with 95% confidence intervals for all primary endpoints.
- Use exact p-values (p = 0.032) rather than thresholds (p < 0.05), except when p < 0.001.
- Percentages must match: if 23 of 150, write "23 (15.3%)" -- verify the math.
- Never round numbers differently between text and tables.
AI Pattern Avoidance
The manuscript must NOT contain these patterns commonly flagged as AI-generated:
Forbidden phrases:
- "In conclusion" (use "In summary" or rephrase)
- "It is worth noting that"
- "It is important to note that"
- "Notably,"
- "Interestingly,"
- "Importantly,"
- "Furthermore," at sentence start (use "In addition," or restructure)
- "Moreover," at sentence start
- "plays a crucial role"
- "a comprehensive analysis"
- "delve into"
- "leverage" (use "use" or "apply")
- "utilize" (use "use")
- "in the realm of"
- "underscores the importance of"
- "sheds light on"
- "paves the way for"
- "a nuanced understanding"
- "the landscape of"
- "a paradigm shift"
- "robust" (unless describing a statistical method)
Forbidden structural patterns:
- Three-part list sentences ("X, Y, and Z" repeated across paragraphs)
- Excessive hedging chains ("may potentially be associated with possible")
- Mirror-structure paragraphs (same template repeated with different content)
- Grandstanding opening sentences ("In the rapidly evolving landscape of...")
Preferred alternatives:
- Vary sentence structure and length within paragraphs.
- Use specific, concrete language over abstract generalizations.
- Let data speak: "The AUC was 0.92" rather than "The model demonstrated remarkable performance."
Journal Compliance
- Respect all word limits from the loaded journal profile.
- Follow the journal's structured abstract format exactly.
- Use the journal's citation style (Vancouver numbered for most radiology journals).
- Include all journal-specific required elements (e.g., "Key Points" for AJR, CLAIM checklist for RYAI AI studies).
---
Skill Interactions
This skill orchestrates other skills at specific phases:
| Phase | Skill called | Purpose |
|---|---|---|
| 2 | /analyze-stats | Statistical analysis for tables |
| 2 | /make-figures --study-type | Figure generation with study-type auto-detection |
| 7.1 | (built-in) | AI pattern removal |
| 7.2 | /check-reporting | Reporting guideline compliance + auto-fix MISSING items |
| 7.3 | /verify-refs | Citation verification and reference artifact audit |
| 7.4 | /self-review --json | Self-review with auto-fix loop (max 2 iterations) |
| 7.4a | /meta-analysis Phase 10 (MA manuscripts) | Audit recovery branch — rebuild extraction/analysis/figures/body when self-review surfaces structural data or protocol issues |
| 7.5 | /humanize | AI-pattern density sweep (<2.0 / 1000 words) |
| 7.5a | /academic-aio (optional, off by default) | AI-search-engine and RAG visibility checklist — run after humanize so QC-confirmed claims and human-readable text anchor the PASS/PARTIAL/FAIL report. Opt-in via --aio or when preparing preprint / GitHub README / CITATION.cff / HF card alongside submission. Silent pipeline execution is explicitly prohibited by the skill's Communication Rules. |
| 7.6 | /manage-refs (pandoc citeproc / Zotero CWYW) | DOCX build from manuscript/manuscript.md + analysis/figures + analysis/tables. Bibliography rendering delegated to /manage-refs scripts/render_pandoc.sh since 2026-05-01. |
| 7.6a | /manage-refs scripts/check_xref.py --strict | Cross-reference QC: in-text Table/Figure citations ↔ body captions ↔ rendered DOCX captions (3-way matrix). Submission gate. |
| 8+ | /find-journal | Journal scope for cover letter (optional) |
If a called skill is not available, perform that step inline using the relevant section of this skill document as guidance.
---
LLM Writing Disclosure
When LLM disclosure is enabled (default), the skill generates transparency statements compliant with ICMJE 2025 and COPE guidelines. The user can disable this with --no-llm-disclosure.
Why Default ON
Major journals (Nature, Lancet, Radiology, JAMA) and the ICMJE (2025 update) require disclosure of AI writing assistance. Omitting disclosure risks rejection or retraction. The default-on design protects the user; they can opt out for journals with no such policy or when LLM assistance was minimal.
Disclosure Locations (3 places)
1. Methods Section — Last Paragraph
Insert at the end of the Methods section, after the ethics statement:
Template (adapt to specifics):
[AI-Assisted Writing Disclosure]
An artificial intelligence language model (Claude, Anthropic) was used to assist with
manuscript drafting, including structuring sections, refining prose, and verifying
internal consistency of reported statistics. All content was critically reviewed,
verified against source data, and approved by all authors. The AI tool was not involved
in study design, data collection, data analysis, or interpretation of results.Customization rules:
- Replace "Claude, Anthropic" with the actual tool(s) used.
- List specific tasks the LLM performed (drafting, editing, literature search, statistical code).
- If the LLM was also used for data analysis (e.g., statistical code generation via
/analyze-stats), state this explicitly: "was also used to generate statistical analysis code, which was reviewed and validated by [statistician/author]."
- Keep to 2-3 sentences. Do not over-explain.
2. Acknowledgments Section
Template:
The authors acknowledge the use of [Claude/tool name] ([Anthropic/developer]) for
writing assistance in preparing this manuscript. The authors retain full responsibility
for the content.3. Cover Letter — AI Disclosure Paragraph (Phase 8+)
Template:
In accordance with [Journal Name]'s policy on AI-assisted writing, we disclose that
[Claude/tool name] was used to assist with manuscript preparation, specifically
[list tasks: drafting, language editing, statistical code review]. All authors have
reviewed and take responsibility for the final content. The AI tool was not listed
as an author and did not contribute to study conception, design, or data interpretation.What NOT to Disclose
- Do not disclose routine use of grammar checkers (Grammarly, Word spell-check) — these
are not considered generative AI under current ICMJE guidance.
- Do not disclose use of reference managers (Zotero, EndNote) or statistical software
(R, Python) unless the LLM generated the analysis code.
Journal-Specific Overrides
When a journal profile is loaded in Phase 0, check for the ## AI Writing Disclosure Policy section in the profile. Tier 1 profiles now include structured fields:
- Requirement level (Required / Recommended / Not specified)
- Permitted scope (All tasks / Language editing only / Not permitted)
- Disclosure location (Methods / Acknowledgments / Cover letter / Submission form)
- AI-generated images (Allowed / Banned / Not specified)
- Policy URL
Use these fields to adjust disclosure language automatically. Key known policies:
- Radiology/RSNA: Required; language editing only; Methods + Acknowledgments; AI images banned.
- RYAI/RSNA: Required; language editing only; Methods + Acknowledgments; AI images banned.
- JAMA/AMA: Required; language editing only; Methods + Cover letter.
- Lancet: Required; language editing only ("readability and language"); Acknowledgments + prompts disclosed.
- BMJ: Required; all tasks permitted but must disclose; Methods + Acknowledgments; applies to text, images, data, diagrams.
- Nature/Springer Nature: Required; language editing only; Methods; AI images banned.
- Science/AAAS: Most restrictive. LLM use limited; treated as potential misconduct if undisclosed.
If the loaded journal profile has no AI Writing Disclosure Policy section, fall back to ICMJE 2025 defaults (disclose in Methods + Acknowledgments, language editing scope).
---
Error Handling
- If the user provides incomplete data for a table, flag specific missing values rather than inventing data.
- If word count exceeds the journal limit after a section draft, report the overage and suggest specific cuts.
- If the critic-fixer loop reaches 3 rounds without passing, present the best version to the user with the remaining issues listed, and ask for guidance.
- Never fabricate references. If a citation is needed, describe the type of reference needed and ask the user to provide it, or call
/search-litto find a real one.
Resumption
If the user returns to a partially completed manuscript: 1. Check the workspace directory for existing drafts. 2. Identify which phase was last completed. 3. Summarize progress and ask the user where to resume.
Anti-Hallucination
- Never fabricate references. All citations must be verified via
/search-litwith confirmed DOI or PMID. Mark unverified references as[UNVERIFIED - NEEDS MANUAL CHECK]. - Never invent clinical definitions, diagnostic criteria, or guideline recommendations. If uncertain, flag with
[VERIFY]and ask the user. - Never fabricate numerical results — compliance percentages, scores, effect sizes, or sample sizes must come from actual data or analysis output.
- If a reporting guideline item, journal policy, or clinical standard is uncertain, state the uncertainty rather than guessing.
---
Gates
Severity levels: ENFORCED = pipeline halts on failure (cannot proceed to next phase). ADVISORY = warning logged, user may override. OPT-IN = runs only when explicitly invoked.
| Phase | Gate | Severity | Trigger | Action on fail |
|---|---|---|---|---|
| 0 | Backbone-article auto-proposal (Phase 0 "Identify a backbone article" action) | ADVISORY | refs.bib has methodologically similar candidate | Surface to user; user accepts/declines |
| 7.0 | Citekey resolution (delegate /manage-refs scripts/check_citation_keys.py) | ENFORCED | UNDEFINED keys present | Halt; resolve via /lit-sync then re-run |
| 7.0 | NEW_PLACEHOLDER drain (delegate /manage-refs) | ENFORCED at 7.6 entry | [@NEW:topic] markers remain | Resolve each before DOCX render |
| 7.1 | Classical-style QC (manuscript-style-classical 11 items) | ENFORCED | § symbol > 0 OR AI Disclosure paragraph in body OR em-dash > 25 | Auto-fix or HALT for senior MA reviewer prep |
| 7.2 | Reporting guideline compliance (/check-reporting) | ENFORCED at submission | <100% mandatory items present | Auto-fix MISSING; ADVISORY for partial |
| 7.3 | Reference audit (/verify-refs --strict) | ENFORCED | FABRICATED or HIGH_MISMATCH_FIRST_AUTHOR > 0 | Halt; fix in Zotero, re-render refs.bib via /lit-sync |
| 7.4 | Self-review fix loop (/self-review --json --fix) | ENFORCED | score below threshold after 2 iterations | Route to Step 7.4a Audit Recovery |
| 7.4a | Audit Recovery branch (route to /meta-analysis Phase 10 for MA manuscripts) | ENFORCED in --e2e | self-review surfaced structural data issue | HALT with RECOVERY_HALT_HUMAN_DECISION if recovery validation fails twice |
| 7.5 | Humanize density (/humanize) | ADVISORY | AI patterns > 2.0 / 1000 words | Sweep + flag remaining; user reviews |
| 7.5a | AIO checklist (/academic-aio --aio) | OPT-IN | user supplies --aio flag | PASS/PARTIAL/FAIL report; never auto-applies |
| 7.6 | DOCX build (delegate /manage-refs scripts/render_pandoc.sh) | ENFORCED | render exits non-zero | Halt; report stderr to user |
| 7.6a | Cross-reference QC (delegate /manage-refs scripts/check_xref.py --strict) | ENFORCED — submission gate | MISSING_DOCX / MISSING_BODY / MISMATCH > 0 | Halt; route fixes per references/check_xref_symptoms.md |
| 7.7 | Final submission gate | ENFORCED | any of 7.0–7.6a above failed | Refuse to mark submission_safe: true |
| 8+ | Cover letter generation | OPT-IN | user invokes --cover-letter | Renders against journal profile |
Cross-cutting global rules applied during 7.x QC:
manuscript-style-classical.md(11 items, Phase 7.1 — ENFORCED)manuscript-references.md(hand-typed References list — ENFORCED via Phase 7.6 delegation)numerical-safety.md,data-integrity.md,citation-safety.md(Phase 7.3 + 7.6a — ENFORCED)senior-mentor-circulation.md(post-7.7, when round 1 begins — ADVISORY)ai-drafted-document-policy.md(Phase 0 if AI-draft attached — ENFORCED)
Abstract structure — structured-abstract anatomy
A structure model for a structured journal abstract, complementing section_guides/title_abstract.md (the section-by-section rules and word limits). Each heading is one structured field; each bullet is what that field must establish. Fill the [brackets]; do not copy this text. The abstract is self-contained — a reader must understand the study from it alone — and every number in it must match the body, the flow diagram, and the tables. Use the journal's exact field labels (some use Purpose for Background, some merge Background/Objective); the moves below are the same regardless of labelling.
Background / Objective
- One or two sentences of the specific gap — not a disease overview ("Disease X is a major health
concern" is a wasted opener). Go straight to what is unknown or insufficient about current methods.
- A final Objective sentence stating the precise question the study answers, with a verb that
matches the design (evaluate / develop and validate / compare / estimate) and does not pre-state the result. Name the primary contrast or estimand if the design is confirmatory.
Methods
- Design in the first clause: retrospective or prospective, single- or multicentre, and the
population (who, and the [YYYY–YYYY] window).
- The index test / exposure / model and the comparator or reference standard; key technical
parameters only if load-bearing.
- The primary outcome named explicitly, and the representative analysis method used to estimate
it (the same primary outcome the Results then report — no swap).
Results
- Open with the final analysed denominator (the number that matches the flow diagram), then a
one-line cohort descriptor (e.g., mean age [X] years, [n] [%] female).
- The primary effect estimate with a 95% CI and its denominator — for example,
"sensitivity [78.4%] (95% CI [71.2–84.5]; [152/194] lesions)" or "adjusted OR [1.62] (95% CI [1.18–2.23])". A bare "was significantly higher (p<0.05)" with no point estimate is not a result.
- At most one or two key secondary outcomes, each with its estimate — not a list of every number in
the paper. Internal consistency is mandatory: a percentage and its [k/N] count must agree, and must agree with the body.
Conclusion
- One or two sentences, matched to the evidence and the design: state the core finding and its
clinical implication, scoped to what was actually shown.
- Stay within the design's reach — a retrospective single-centre accuracy study supports
"showed [high specificity] for [task] in this cohort", not "is ready for clinical deployment" or any causal claim. No "further studies are needed" filler (add only if a reviewer asks).
- Invest the most effort here: it is read first and most often, and it must be a directly citable
statement.
Common omission / failure modes
- The estimate-free Results field — a "significant" claim with **no point estimate, CI, or
denominator — and a Conclusion that over-reaches the design (deployment, causal, or generalization language a retrospective/single-site study cannot license). Also avoid: a body↔abstract number mismatch (a percentage or N that disagrees with the flow diagram or tables), spinning a non-significant secondary outcome into the headline while burying or omitting the primary estimate, introducing data not in the paper, and exceeding the journal's word limit** (see the limits table in section_guides/title_abstract.md). Cross-reference section_guides/title_abstract.md (field labels, word limits, lead-with-the-primary-estimand rule) and exemplar_introduction.md (the Objective the Abstract restates).
Radiology case-report anatomy — when the image is the case
A structure model for imaging-led case reports (diagnostic radiology, nuclear medicine, and interventional radiology), complementing exemplar_case_report.md and paper_types/case_report.md. Load it when the teaching point is an imaging finding, a cross-modality discordance, an incidental finding, a structured-reporting decision, or an image-guided procedure/complication. Synthetic anatomy model: required moves, failure modes, cross-checks — not prose to copy.
What makes a radiology case report different
The contribution is the image and how it was read, so the discipline lives in the imaging description, not just the clinical narrative. Three rules govern the whole report:
1. Per-modality, in clinical order. Describe each modality the way it was acquired and read — for each: technique → findings → impression, kept separate. Walk modalities in the order they were obtained (e.g., radiograph → ultrasound → CT → MRI → PET/CT → histopathology), not lumped. 2. Reproducible technique. State the parameters another reader would need: sequence/phase, field strength, contrast agent + dose + injection rate, CT kV/keV-reconstruction/CTDIvol, PET tracer dose + uptake time + fasting, transducer frequency. For interventional cases, name devices with sizes/gauge and the step sequence. 3. Findings vs impression. Report the observation ("circumferential aortic wall thickening, iodine X mg/mL") separately from the interpretation ("favoring active vascular inflammation").
Structured reporting lexicons
When a standardized system applies, use the category and state what it means — a bare "BI-RADS 4" is weaker than "BI-RADS 4b (moderate suspicion, ~10–50% malignancy risk)."
- Breast — BI-RADS (give the sub-category and risk band)
- Liver — LI-RADS; Prostate — PI-RADS; Thyroid — TI-RADS; Lung nodule — Lung-RADS;
Adnexal — O-RADS; incidental findings — the relevant ACR Incidental Findings/white-paper guidance.
- If no system applies, describe with the standard descriptors of that modality (margin, echotexture,
signal on each sequence, enhancement kinetics/curve type, attenuation).
Quantitative anchors (and their honesty)
- Give the measurement with method and units: lesion size + location (e.g., o'clock + distance
from nipple/skin), SUVmax, iodine concentration (mg/mL) with the ROI placement rule, time–signal intensity curve type, degree of stenosis / peak systolic velocity.
- When a value has no validated threshold, say so and label it exploratory; optionally give
institutional comparison values. Do not present a number as diagnostic when no cutoff exists.
Multimodality correlation and discordance
- When modalities disagree, make the discordance the explicit teaching point and state how it was
resolved (the decisive modality, or histopathology/IHC). Example pattern: ultrasound suspicious vs MRI benign kinetic curve → resolved by core-needle biopsy.
- Modality-completeness self-critique: if a standard modality was not performed, name it as a
limitation (e.g., "mammography was not obtained, precluding complete radiologic correlation") and say why (compliance, availability, pain).
Subtype: interventional radiology (procedure / complication)
- Reproducible procedure log: access, devices with sizes (needle gauge, wire, balloon, catheter,
coil/embolic, ablation probe + protocol), imaging guidance, and the step sequence.
- Complication recognition and management in context: state latency (e.g., a delayed post-procedural bleed),
the recognition trigger, and the diagnostic-then-therapeutic pathway (angiography → embolization), with pre/post outcome. Quantify (blood loss, transfusion, preserved organ function) and give the background complication rate.
Subtype: incidental finding
- Frame imaging as a problem-solving / staging tool and describe how the incidental lesion was
characterized (the discriminating CT/US/MR features), the functional-imaging pitfall if relevant (FDG-avid benign mimic; FDG-occult malignancy), and the reporting action — what should be explicitly stated in the report to prevent mis-staging or over-treatment.
Image and figure discipline (do this before submission)
- De-identify every image at the DICOM level: remove burned-in annotations, accession numbers,
dates, institution banners, and faces; confirm panels carry no identifiers.
- Write real alt text for each figure (not a placeholder) — modality, plane, and the labeled
finding; pair complex cases with make-figures exemplar_plots/imaging_panel.md.
- Disclose device/vendor relationships: advanced-technique or device cases (spectral/photon-counting
CT, ablation systems) must state vendor research agreements, employment, or speaker arrangements.
Common failure modes
- Modality soup — findings from several modalities merged into one paragraph instead of
technique → findings → impression per modality.
- Bare structured category — "BI-RADS 4" / "PI-RADS 4" with no sub-category or risk meaning.
- Unreproducible technique — no sequence/phase, contrast, dose, or device specifics.
- Number without a threshold — a quantitative value presented as diagnostic when no validated
cutoff exists.
- Placeholder alt text / identifiable images — the most common reviewer/production stop for an
imaging case report.
- Undisclosed device/vendor COI in an advanced-technique case.
- Overclaiming from n=1 — a feasibility or detection case framed as performance/effectiveness.
Case-report anatomy — CARE narrative + 150-word abstract
A structure model for case reports, complementing paper_types/case_report.md and the CARE checklist. Use it when /write-paper Phase 0 identifies the paper type as case report. This is a synthetic anatomy model: it describes the required moves, failure modes, and cross-checks; it is not prose to copy.
Narrative spine
Strong case reports read as a clinically disciplined story, not as a miniature original article. Build the manuscript around the sequence below, keeping each move tied to the patient's course.
1. Why this case matters — the rare presentation, diagnostic trap, management lesson, adverse event, or unexpected response. The reason must be specific enough that a reader understands why a single case deserves publication. 2. Who the patient is, de-identified — age range/sex and clinically relevant background only. Remove dates, institutions, initials, locations, unique occupations, and unnecessary demographic detail. 3. What happened in time order — symptoms, examination, tests, diagnostic reasoning, intervention, follow-up, and outcome. A timeline figure or table should let the reader reconstruct the course without rereading the prose. 4. How the diagnosis was reasoned through — include alternatives considered, why they were less likely, key imaging/laboratory/pathology findings, and any diagnostic limitation. 5. What was done and what changed — treatment, dose/procedure/device details if load-bearing, response, adverse events, adherence/tolerability, and follow-up duration. 6. What the reader should learn — a narrowly scoped teaching point, anchored to the literature and the evidence level of a single case.
150-word structured abstract
Most short case reports need a compact abstract with Introduction / Case Presentation / Conclusion headings. Allocate words deliberately; do not import the IMRAD abstract model.
Introduction
- One sentence naming the condition/presentation and the precise reason the case is reportable.
- Avoid broad disease background. The abstract's first sentence should already point to the novelty
or teaching value.
Case Presentation
- Two to four sentences covering the patient's de-identified presentation, key findings, diagnostic
reasoning, intervention, and follow-up outcome.
- Include the decisive imaging/laboratory/pathology finding if it is the reason the case matters.
- Keep chronology clear; do not compress the case into an unexplained list of diagnoses and tests.
Conclusion
- One sentence stating the teaching point, scoped to a single case.
- Use cautious verbs: "may", "should prompt consideration", "is consistent with", or "highlights".
Do not claim incidence, efficacy, safety, causality, or practice-changing proof.
Case Presentation section
Patient information
- De-identified demographics and relevant clinical background.
- Main concern/symptom in the patient's sequence, not as a retrospective diagnosis.
- Past medical, family, psychosocial, medication, and exposure history only when they affect the
differential, intervention, or interpretation.
Clinical findings
- Physical examination and bedside findings that altered diagnostic reasoning.
- State relevant negatives when they narrow the differential; omit routine normal findings that do
not move the case.
Timeline
- Use a compact figure or table when the course has more than two clinically meaningful time points.
- Include onset, presentation, key tests, diagnosis, intervention changes, complications, and final
follow-up/outcome.
- Use relative time (
Day 0,Week 6,Month 3) unless exact dates are essential and approved for
publication.
Diagnostic assessment
- Name the test modality and the load-bearing finding, then the interpretation.
- For imaging cases, describe the finding and impression separately: modality/sequence, lesion or
anatomical location, discriminating feature, and how it affected the differential.
- Document diagnostic challenges: atypical presentation, unavailable tests, delayed diagnosis,
discordant results, or uncertainty that remained.
Therapeutic intervention
- State what was done, why, and when it changed.
- Include dose, route, procedure, device, duration, or surgical detail only when needed for
reproducibility or interpretation.
Follow-up and outcomes
- Report the follow-up interval, patient- or clinician-assessed outcome, objective response where
available, adverse events, and residual deficits.
- Avoid "the patient improved" unless the text specifies how improvement was assessed.
Discussion
- Open with the one-sentence lesson, not a second case summary.
- Compare against the nearest reported cases or mechanisms. If five or more similar cases are found,
use a brief comparison table; if fewer, state the search boundary and avoid implying a definitive global count.
- Separate temporal association from causality. A single case can raise a hypothesis or illustrate a
diagnostic clue; it cannot estimate treatment effect or risk.
- State what is uncertain: alternative explanations, incomplete testing, short follow-up, missing
patient perspective, or limited generalizability.
- End with a practical teaching point that a clinician can use at the bedside.
Subtype: adverse drug / device / contrast reaction (pharmacovigilance)
When the case is the adverse event (drug reaction, contrast-agent extravasation, device complication), the attribution is the contribution — make it rigorous, not narrative.
- Apply a named causality instrument, do not just assert it: the Naranjo Adverse Drug Reaction
Probability Scale or WHO-UMC categories for drugs; report the score and the resulting tier (doubtful / possible / probable / definite). State the inputs that drove it.
- Dechallenge and (only if ethical) rechallenge: document that withdrawal was followed by
resolution, and note the temporal latency from exposure to onset. Rechallenge is usually withheld on safety grounds — say so rather than leaving it unexplained.
- Severity and preventability where instruments exist (e.g., Modified Hartwig–Siegel severity;
Schumock–Thornton preventability) — these turn an anecdote into a structured safety report.
- Locate the event against a denominator: an institutional rate (events / total procedures) or a
pharmacovigilance database count (e.g., prior reports in a spontaneous-reporting system) anchors rarity without overclaiming incidence from n=1.
- Separate the index event from downstream confounders: if patient self-management or a
comorbidity worsened the course, attribute each step rather than blaming the agent for the whole trajectory.
- Close the safety loop: state that the reaction was (or should be) reported to the relevant
pharmacovigilance programme.
- A compact instrument table (tool | criteria applied | result | interpretation) is the clearest
way to present causality/severity/preventability.
Subtype: diagnostic pitfall / mimic
When the lesson is "this entity was mistaken for another," the differential is the spine.
- Name the trap explicitly in the framing ("X masquerading as Y") and state why the
misclassification was plausible (overlapping imaging/clinical features, an obscured primary site).
- Structured differential: list the realistic competitors and, for each, the feature that argued
for or against it (morphology, immunohistochemistry, virology, distribution). A differential that only names alternatives without adjudicating them is incomplete.
- Trace the resolving pathway: which test finally settled origin vs extent (e.g., inconclusive
cross-sectional imaging → tissue diagnosis → functional imaging for staging), and acknowledge where a modality was limited or omitted and why.
- Diagnostic-delay framing (optional but strong): when delay shaped the outcome, separate
patient-level from healthcare-pathway-level contributors rather than a single "late presentation."
- Self-critical mechanism reasoning: when proposing a mechanism, test it against this case's own
data and reject candidates the case does not support (e.g., a hyperperfusion mechanism is unlikely when the recorded blood pressure stayed below the autoregulatory threshold).
Required cross-checks
- Consent / anonymization: confirm written consent or the applicable waiver statement before
drafting submission-ready text; image panels must be stripped of identifiers.
- CARE coverage: Title, Keywords, Abstract, Patient Information, Clinical Findings, Timeline,
Diagnostic Assessment, Therapeutic Intervention, Follow-up/Outcomes, Discussion, Patient Perspective (if available), and Informed Consent.
- Literature boundary: search strategy, number of similar cases found, and whether a comparison
table is warranted.
- Figure anatomy: for complex courses, pair the text with
/make-figures
exemplar_plots/clinical_timeline.md.
- Imaging-led case: if the contribution is the image or an image-guided procedure, use
exemplar_case_report_radiology.md (per-modality technique→findings→impression, structured-reporting lexicons, quantitative-threshold honesty, IR procedure/complication, DICOM de-identification).
Common failure modes
- Rarity without justification — the Introduction says "rare" but gives no clinical reason,
epidemiologic anchor, or teaching value.
- Consent or de-identification gap — no consent statement, identifiable dates/institutions, or
unmasked imaging metadata.
- Chronology collapse — the diagnosis, intervention, and outcome are present but not in a
reconstructable sequence.
- Diagnostic reasoning missing — tests are listed, but alternatives and why the final diagnosis
was favored are absent.
- Causal overclaim — the Discussion treats a temporal association as proof of treatment effect,
adverse-event causality, or mechanism.
- Literature absence mishandled — "first case" or "only case" is asserted without a transparent
search boundary.
- Teaching point too broad — the conclusion asks clinicians to change practice rather than
recognize a clue, consider a diagnosis, or report similar cases.
- Causality by assertion — an adverse-event case calls the agent "causative" without a named
instrument (Naranjo/WHO-UMC), documented dechallenge, or exclusion of alternatives.
- Differential without adjudication — a mimic/pitfall case lists competing diagnoses but never
says which feature ruled each in or out.
Discussion structure — AI/ML model development + validation (TRIPOD+AI / CLAIM)
A structure model for the Discussion of a clinical AI/ML model study, completing the trio with the exemplar_methods/ and exemplar_results/ siblings. Each heading is a paragraph; each bullet is what it must establish. Fill the [brackets]; do not copy this text. Follows section_guides/discussion.md. Introduces no new results, cites no tables/figures, and the claim must not exceed the validation evidence.
Paragraph 1 — Key finding and why it matters
- Restate the primary discrimination (and that it held on the external set), plus calibration
in one phrase — not a single best number.
- One sentence on the decision the model could support, hedged to the evidence level.
Paragraphs 2–3 — Interpretation and comparison
- Compare to existing models / current practice: what the model adds (e.g., extends to a
group current criteria exclude), and frame it as complementary, not a replacement for the clinician or standard pathway.
- Biological / clinical plausibility for why the inputs carry signal (named mechanism), so
the model is not a black box asserted to work.
- Separate the evidence tiers explicitly: a retrospective validation shows discrimination;
decision-impact and outcome benefit require prospective / trial evidence — and say which is and is not yet available.
Limitations
- Name overfitting / optimistic validation candidly (any test-set leakage, internal-only
validation, thin subgroups), calibration limits, single-site/scanner data, and the absence of prospective or decision-curve evidence if a use claim is made.
Generalizability and future work
- Where the model would and would not transfer (sites, scanners, demographics); the
prospective/external validation and, ultimately, the trial needed before deployment.
Conclusion
- Matched to the evidence — "may support / warrants prospective evaluation", never "ready for
clinical deployment" or "outperforms clinicians" unless a same-task, difference-tested, prospective comparison supports it.
Common omission
- An explicit evidence-tier separation (discrimination ≠ clinical benefit) and a candid
optimism/calibration caveat — the Discussion elements AI drafts most often skip, and the gateway to deployment over-reach. Cross-reference section_guides/discussion.md, the AO5 probe and peer-review/references/exemplar_reviews/optimistic_validation_reporting.md, and the TRIPOD+AI / CLAIM critical items in peer-review/references/reviewer_calibration/compliance_floor.md.
Discussion structure — diagnostic-accuracy study (STARD)
A structure model for the Discussion of an index-test-vs-reference-standard accuracy study, completing the trio with the exemplar_methods/ and exemplar_results/ siblings. Each heading is a paragraph; each bullet is what it must establish. Fill the [brackets]; do not copy this text. Follows section_guides/discussion.md (4-paragraph base). Discussion introduces no new results and cites no tables/figures.
Paragraph 1 — Key finding and why it matters
- Restate, in 2–3 sentences, what was evaluated and the headline accuracy (the primary
sensitivity/specificity at the operating point) without re-dumping every number.
- One sentence on the clinical decision the test informs (e.g., who can avoid
[procedure]).
Paragraphs 2–3 — Interpretation and comparison to prior tests
- Compare to the existing test / prior tools head to head: agree or differ, and why
(different reference standard, spectrum, threshold, or population) — neutrally, not by disparaging prior work.
- Place the test in the clinical pathway: what it adds over current practice, and the
decision threshold's rationale (cite a decision-curve / net-benefit argument only if it was pre-specified and reported in Results — do not introduce a new analysis here, just avoid an unjustified arbitrary cut-point).
Limitations
- Tie each to design: spectrum bias (case-mix vs the intended-use population),
verification / differential-verification bias (reference standard applied unequally), reference-standard imperfection, single-setting/reader, retrospective sampling, and any data-derived threshold's optimism.
Generalizability and implications
- The populations/settings the accuracy does and does not transfer to; the implementation
context (and, for an AI index test, the prospective evaluation still required).
Conclusion
- One paragraph matched to the evidence — "supports/triages", not "replaces [the comparator]"
unless a non-inferiority comparison and prospective data support replacement.
Common omission
- Spectrum and verification bias and a threshold/clinical-pathway justification — the
Discussion points DTA drafts most often skip, and the two that most often turn a strong headline into an over-reach. Cross-reference section_guides/discussion.md and the STARD critical items in peer-review/references/reviewer_calibration/compliance_floor.md.
Discussion structure — observational cohort (STROBE)
A structure model for the Discussion of an exposure→outcome cohort study, completing the trio with the exemplar_methods/ and exemplar_results/ siblings. Each heading is a paragraph; each bullet is what it must establish. Fill the [brackets]; do not copy this text. Follows section_guides/discussion.md. Introduces no new results and cites no tables/figures.
Paragraph 1 — Key finding and why it matters
- Restate the primary association (adjusted effect with its direction/magnitude) in 2–3
sentences (only restate a comparison to conventional predictors if that comparison was pre-specified and reported in Results — introduce no new comparison here).
- One sentence on the clinical/public-health relevance — as risk stratification, not proven
causation.
- If the primary result is null, frame it by precision, not absence. Do not write "X was
not associated with Y" flatly; state what the confidence interval excludes — e.g. "the 95% CI excluded an eGFR difference larger than ~1.7" or cite the minimum detectable effect that the study was powered for. "No effect" and "could not exclude an effect of size X" are different claims; an underpowered null is "not yet established," and the MDE must be computed under the full primary-model adjustment set (a 2-covariate power calc overstates precision). Watch for bilateral over-correction — a prior "independently associated" overclaim swinging to an equally unsupported "not associated" claim during revision.
Paragraphs 2–3 — Interpretation and comparison
- Compare to prior evidence (agree/differ and why: design, population, exposure measurement).
- Causal caution is mandatory: state plainly that the estimate is an association, not a
demonstrated causal effect; discuss reverse causation and the limits a single-timepoint or cross-sectional exposure places on directionality.
- Offer a mechanism/biological plausibility as hypothesis, flagged as such.
Limitations
- Residual and unmeasured confounding with the adjustment set named and its gaps stated
(and, ideally, an E-value / tipping-point referenced from Results); selection bias; exposure misclassification; missing-data handling; outcome ascertainment completeness.
- Adjustment-set choice both ways: state that confounders were chosen by a causal rationale
(DAG / prior literature), not because they differed in Table 1, and that no mediator or consequence of the outcome was adjusted (over-adjustment — e.g. a renally-excreted lab in an eGFR model); if a borderline covariate was included, reference the drop-it sensitivity model from Results.
Generalizability
- The population the cohort represents and where the estimate may not transfer (single
ancestry/region/registry), and how external validation mitigates but does not remove the concern.
Conclusion
- Matched to the evidence — "associated with", "may identify higher-risk individuals", never a
care directive (defer/withhold/initiate) or a prognostic/surveillance claim a cross-sectional or single-visit design cannot license.
Common omission
- An explicit causation caveat with reverse-causation/unmeasured-confounding treatment —
the Discussion element cohort drafts most often soften, and the one a methods reviewer checks first. Cross-reference section_guides/discussion.md, the peer-review/references/domain-probes/observational_confounding.md probe, and the STROBE critical items in peer-review/references/reviewer_calibration/compliance_floor.md.
Exemplar Discussion sections — structure models, not prose to copy
write-paper teaches Discussion rules (section_guides/discussion.md: the 4-paragraph structure, no new results, no table/figure citations, claim-to-evidence matching) but had no worked example of what a complete Discussion looks like end to end. This directory completes the exemplar trio (exemplar_methods/ → exemplar_results/ → exemplar_discussion/) with structure models: each shows the Discussion paragraph order for a study type and, for every paragraph, what it must establish — with placeholder specifics, never copied text.
These are authored from scratch as teaching models, not extracted from any published paper. Use them to check that a draft's Discussion has every load-bearing paragraph and stays matched to the evidence; do not copy wording.
How they are used
In write-paper Phase 5 (Discussion), after loading section_guides/discussion.md: pick the exemplar matching the study type and confirm the draft restates the key finding, compares to prior work without disparagement, gives design-tied limitations, treats generalizability, and closes with a conclusion matched to the evidence — introducing no new results.
Contents
diagnostic_accuracy_stard.md— comparison vs prior test, clinical-pathway/threshold,
spectrum/verification-bias limitations (STARD).
ai_validation_tripod_claim.md— complementary-not-replacement framing, evidence-tier
separation (discrimination ≠ clinical benefit), optimism/calibration caveats (TRIPOD+AI / CLAIM).
observational_cohort_strobe.md— mandatory causal caution, reverse-causation and
unmeasured-confounding limitations, no care-directive conclusion (STROBE).
Curator guidelines (for adding more)
- Synthetic only. Write the structure with placeholder specifics; never paste a real
paper's Discussion, use no real citations, no PII, English only.
- Complete the trio for the study type (mirror the
exemplar_methods/+exemplar_results/
siblings), each line stating what the paragraph must establish, not finished prose.
- Enforce claim-to-evidence matching and no new results — the two Discussion failure
modes; name the study type's most common omission.
- Anchor to the reporting-guideline critical items in
peer-review/references/reviewer_calibration/compliance_floor.md.
- Keep each file ~45–75 lines. Cross-reference
section_guides/discussion.mdrather than
duplicating it.
Introduction structure — gap-storytelling model
A structure model for an IMRAD Introduction, complementing section_guides/introduction.md (the 5-step rules). Each heading is a paragraph move; each bullet is what it must establish. Fill the [brackets]; do not copy this text. Keep the whole Introduction to 3–4 short paragraphs — it narrows from the problem to this study, and ends with the objective. No results, no method detail, no literature-review sprawl.
¶1 — Why this matters (burden / significance)
- The clinical problem and its scale in 2–3 sentences (`[incidence / mortality / cost / decision
it affects]`) — enough to motivate, not a textbook chapter.
- Name the specific clinical decision or population the study will speak to, so the funnel has a
target.
¶2 — What is known (the landscape, then the most relevant prior work)
- The current approach / standard and what prior studies established — grouped by idea, not a
mechanical "A did X; B did Y; C did Z" list.
- Narrow to the 1–2 most relevant prior works (the ones the gap is defined against), stated
neutrally (no disparagement).
¶3 — The gap (the load-bearing paragraph)
- State concretely what is still missing or unresolved — a specific limitation of the prior
work (single-centre, no external validation, no head-to-head comparison, surrogate endpoint, unaddressed population), not a generic "few studies have examined…".
- The gap must be specific and falsifiable, and must logically license the objective in ¶4
— if filling this gap would not require this study, the gap is mis-stated.
¶4 — Objective (and, if applicable, hypothesis)
- One sentence: "We therefore [did what] to [answer the gap], in [design/population]." The verb
matches the design (evaluate / develop and validate / compare / estimate) and does not pre-state the result.
- A directional hypothesis or the primary estimand if the design is confirmatory; for an
exploratory study, say so.
Common omission / failure modes
- A vague gap ("few studies have examined…") that does not point to a specific deficiency, and
a gap↔objective mismatch (the objective does not follow from the gap, or over-reaches it) — the two faults that most often weaken an Introduction. Also avoid: starting too broad (a textbook ¶1), reviewing the whole field (¶2 sprawl), citing tangential background, and leaking results or conclusions into the objective. Cross-reference section_guides/introduction.md (5-step rules, word/paragraph targets) and section_guides/title_abstract.md.
Methods structure — AI/ML model development + validation (TRIPOD+AI / CLAIM)
A structure model for a clinical AI/ML model study. Each heading is a paragraph; each bullet is what it must establish. Fill the [brackets]; do not copy this text.
Study design and data sources
- Design (retrospective/prospective), the prediction task, and the intended use (where in
the workflow the output acts, and which decision it informs).
- Data sources, sites, and dates for each split; how the cohort was assembled.
Participants, inputs, and outcome
- Eligibility as a numbered list; the population the model is meant to serve.
- Input data types and exactly which fields the model sees (flag any field that could carry
the label — report a no-leaky-field sensitivity analysis or mask it).
- Outcome/label definition, who assigned it, and the reference standard for the label.
Data partition and leakage control
- Train / validation / test split by patient (not by image/record), and how
independence was enforced; any external (geographic/temporal) test set.
- Explicit statement that no test data informed development (preprocessing, feature/threshold
selection, model choice) — leakage inflates every metric and is invisible in the results.
Model development
- Architecture, key hyperparameters, training procedure, and how the operating threshold was
chosen (and on which split).
- Whether training was from scratch or fine-tuned; for an adaptation claim, a same-backbone
zero-shot/few-shot baseline.
- Reproducibility: hardware/software/versions, random-seed handling, code/model availability.
Evaluation
- Primary metric(s) with 95% CIs; calibration (slope/intercept or a calibration plot),
not discrimination alone, when probabilities drive a decision.
- Comparator: if vs clinicians, the **same task on the same inputs under the same
constraints**, with a test of the difference (not two standalone AUCs); decision-curve / net-benefit if a deployment/utility claim is made.
Sample size
- Events-per-variable / minimum test-set events for the target precision, or the
available-sample justification.
Reporting-guideline fit
- TRIPOD+AI (name base TRIPOD + the AI extension) and/or CLAIM 2024. Critical items: input
data + outcome, partition/leakage control, model + training details, calibration.
Common omission
- Calibration and a patient-level (not image-level) split — the two most commonly
missing, and both can make excellent-looking metrics misleading.
Methods structure — diagnostic-accuracy study (STARD)
A structure model for an index-test-vs-reference-standard accuracy study. Each heading is a paragraph; each bullet is what the paragraph must establish. Placeholders in [brackets] stand for the study's real specifics — fill them; do not copy this text.
Study design and oversight
- State the design (prospective / retrospective, consecutive / case-control sampling — and
why; case-control sampling inflates accuracy and must be named).
- Setting and dates:
[single/multi-center],[YYYY–YYYY], recruitment source. - Ethics approval + consent (or waiver) and registration if applicable.
Participants and eligibility
- Eligibility as a numbered list: inclusion
(1)…(2)…, exclusion(1)…(2)…. - The clinical question and the intended-use population the index test is meant to serve.
- Flow that lets the reader rebuild a 2×2 (enrolled → tested → analyzed; exclusions with
reasons) — this is what the STARD diagram renders.
Index test
- What it is, who performed/read it, their experience, and blinding to the reference
standard and to clinical data.
- Acquisition specifics:
[scanner/vendor/field strength], protocol, software/version. - The positivity threshold/criterion and whether it was pre-specified or data-derived.
Reference standard
- The reference standard and its rationale (the validity hierarchy: pathology > imaging
panel consensus > clinical follow-up at [interval]).
- Who applied it, their blinding to the index test, and the index↔reference time interval
(verification timing).
- How indeterminate/non-diagnostic results were handled (ITD vs per-protocol — state which).
Sample size
- The accuracy target and precision (expected sensitivity/specificity + CI half-width) or
the available-sample justification; the prevalence assumed.
Statistical analysis
- Sensitivity/specificity with 95% CIs; PPV/NPV with the prevalence they assume.
- Comparison method if two tests (paired McNemar / difference in AUC with a test of the
difference — not two standalone AUCs), and any subgroup/multiplicity handling.
- Missing/indeterminate handling and any sensitivity analysis.
- Software + version.
Reporting-guideline fit
- STARD 2015 (or STARD-AI for an AI index test — name BOTH base and extension). Critical
items: reference standard + rationale, reader blinding, participant flow, indeterminate handling (see peer-review/references/reviewer_calibration/compliance_floor.md).
Common omission
- Reader blinding to the reference standard and the index↔reference time interval
— the two details most often missing, and both bias the headline accuracy if unaddressed.
Methods structure — observational cohort (STROBE)
A structure model for an exposure→outcome cohort study. Each heading is a paragraph; each bullet is what it must establish. Fill the [brackets]; do not copy this text.
Study design and setting
- Design (prospective / retrospective cohort) and why; data source (`[registry / EMR /
health-screening DB]`), sites, and the enrollment and follow-up dates (keep enrollment distinct from follow-up end — do not conflate them).
- Ethics approval / consent or waiver; registration if applicable.
Participants and eligibility
- Eligibility as a numbered list (inclusion / exclusion).
- The assembly cascade with counts (source → eligible → analytic), so the STROBE flow
reconciles (start − Σ exclusions = analytic N).
Variables — exposure, outcome, covariates
- Exposure definition, how it was measured/classified, and the cut-points with their source
(cite the canonical definition; quote the data-dictionary mapping for a DB variable).
- Outcome definition, ascertainment, and timing.
- Covariates/confounders and how each was measured; pre-specify the adjustment set and its
rationale (a DAG), not a data-driven selection.
Bias, missing data, and study size
- Sources of bias and how they were addressed (selection, measurement, confounding).
- How missing data were handled (complete-case vs imputation — state which; complete-case
must reconcile: total − missing = complete).
- Study size: the precision / event-budget justification (events per variable for the model).
Statistical analysis
- The primary model and estimand (the contrast and effect measure), with 95% CIs; the
pre-registered primary if applicable (do not reassign the primary after seeing results).
- Subgroup/interaction and sensitivity analyses, multiplicity handling, and how continuous
variables were modeled (linearity assumption).
- Software + version.
Reporting-guideline fit
- STROBE. Critical items: eligibility/selection, exposure/outcome definitions, missing-data
handling, participant flow (see peer-review/references/reviewer_calibration/compliance_floor.md).
Common omission
- A pre-specified adjustment set with rationale and explicit missing-data handling —
the two paragraphs cohort drafts most often skip, and both are first-line reviewer targets.
Exemplar Methods sections — structure models, not prose to copy
write-paper teaches Methods structure (PICO, the reporting-guideline cross-reference, the backbone-article strategy) but had no worked example of what a complete, well-ordered Methods section looks like end to end. This directory fills that gap with structure exemplars: each shows the paragraph order for a study type and, for every paragraph, what it must establish — with placeholder specifics, never copied text.
These are authored from scratch as teaching models, not extracted from any published paper. Use them to check that a draft's Methods has every load-bearing paragraph in a sane order; do not copy wording.
How they are used
In write-paper Phase 3 (Methods), after loading section_guides/methods.md: pick the exemplar matching the study type and confirm the draft establishes each element the exemplar lists (design, setting/dates, eligibility, the index test/exposure, the reference standard/outcome, sample-size justification, the statistical plan, and the reporting-guideline fit). A missing element is a gap to fill before drafting Results.
Contents
diagnostic_accuracy_stard.md— an index-test vs reference-standard accuracy study (STARD).ai_validation_tripod_claim.md— an AI/ML model development + validation study (TRIPOD+AI / CLAIM).observational_cohort_strobe.md— an exposure→outcome observational cohort (STROBE).
Curator guidelines (for adding more)
- Synthetic only. Write the structure with placeholder specifics (
[scanner/vendor],
[N], [YYYY–YYYY]); never paste a real paper's Methods, and use no real citations, no PII, English only.
- One study type per file, paragraph by paragraph, each line stating *what the paragraph
must establish* (not finished prose).
- Anchor to the reporting guideline the study type maps to, naming the critical items
(see peer-review/references/reviewer_calibration/compliance_floor.md).
- Name the common omission for each study type (the paragraph drafts most often skip).
- Keep each file ~50–90 lines. Cross-reference
section_guides/methods.mdand the relevant
paper_types/ template rather than duplicating them.
Results structure — AI/ML model development + validation (TRIPOD+AI / CLAIM)
A structure model for the Results of a clinical AI/ML model study. It follows its Methods AI-validation sibling in Methods order. Each heading is a paragraph; each bullet is what it must establish. Fill the [brackets]; do not copy this text. Report findings only — no interpretation.
Cohort assembly and data partition (Figure 1)
- The flow as Figure 1: source → eligible → analyzed, with the train / validation / test
split shown and the counts for each set (and each external set).
- Splits are by patient; state the counts so a reader can see no patient spans two sets.
- Numbers reconcile with the Abstract, Methods, and Table 1.
Baseline characteristics, by set (Table 1)
- Population characteristics for development and each test set, so distribution shift between
development and external data is visible (one decimal convention).
- The outcome/label prevalence in each set (it conditions PPV/NPV and calibration).
Reference points before the headline
- Report the relevant baselines first — a simpler model (clinical-only / single-modality) and,
where applicable, the unaided clinician — so the model's increment is legible rather than asserted.
Primary discrimination
- The primary metric (AUROC, and AUPRC under class imbalance) with 95% CIs, reported
separately for internal and each external set (not a single pooled number).
- Where a gain over a baseline is claimed, the difference with a test (DeLong / paired),
not two standalone estimates.
Calibration and operating point
- Calibration (slope/intercept or a calibration plot; Brier/Hosmer–Lemeshow) on the test
set — discrimination alone is insufficient when a probability drives a decision.
- The decision threshold used for any sensitivity/specificity, and that it was fixed on the
development/training data, not the test set.
Clinical utility (when a use claim is made)
- Decision-curve / net-benefit across plausible thresholds, or a reader study reporting the
change in reader performance (e.g., false-positive reduction, agreement) with the model.
Interpretability and error analysis
- If interpretability is claimed, quantify it (attribution overlap, importance ranking),
and report failure modes as counts/rates by pre-specified error category, subgroup, or input condition (e.g., lesion size, image quality) — report-only; reserve causal explanation of the errors for the Discussion.
Common omission
- Calibration and per-set (not pooled) external metrics with CIs — the elements most
often missing, and both can make excellent discrimination misleading. Cross-reference section_guides/results.md, the peer-review/references/exemplar_reviews/optimistic_validation_reporting.md phrasing model, and the TRIPOD+AI / CLAIM critical items in peer-review/references/reviewer_calibration/compliance_floor.md.
Results structure — diagnostic-accuracy study (STARD)
A structure model for the Results of an index-test-vs-reference-standard accuracy study. It follows its Methods STARD sibling in Methods order. Each heading is a paragraph; each bullet is what it must establish. Fill the [brackets]; do not copy this text. Report findings only — no interpretation (that is the Discussion).
Participant flow (Figure 1)
- The STARD flow as Figure 1: enrolled → received index test → received reference standard →
analyzed, with exclusions and counts and reasons at each step.
- State how many had the reference standard by each route (e.g.,
[N]by pathology,[N]by
[interval] follow-up) and how indeterminate/non-diagnostic results were counted (ITD vs per-protocol — match what Methods declared).
- Numbers must reconcile with the Abstract, Methods, Table 1, and the 2×2.
Cohort characteristics and prevalence (Table 1)
- Baseline demographics/clinical features of the analyzed participants (one decimal
convention, held throughout).
- The prevalence of the target condition by the reference standard — every PPV/NPV is
read against it, so state it before the predictive values.
Primary accuracy
- Sensitivity and specificity with 95% CIs, at the pre-specified positivity threshold
(name the operating point; state it was fixed in advance or derived only on a training/derivation set, not optimized on the analysis/test set).
- PPV/NPV with the prevalence they assume; accuracy if reported.
- State the analysis unit and keep it fixed (per-patient / per-lesion / per-segment); when
units are nested, report each level and give the lesion-selection rule, so every metric's denominator is unambiguous.
Agreement / measurement (if a quantitative index)
- For a continuous index read against the reference, report agreement (Bland–Altman mean
difference + limits, or ICC), not only thresholded accuracy.
Subgroups and comparator
- Pre-specified subgroups (e.g., by
[size / calcification / PSA stratum]) showing where
accuracy degrades; report the subgroup 2×2 or its metrics, not just a p-value.
- If two tests are compared, the difference with a paired test (McNemar / difference in
AUC), not two standalone estimates.
Common omission
- The explicit flow diagram with per-step counts and the **handling of indeterminate
results, plus differential-verification bias** when the reference standard was applied more often after a positive index test — the elements most often missing, and each inflates the headline accuracy. Cross-reference section_guides/results.md and the STARD critical items in peer-review/references/reviewer_calibration/compliance_floor.md.
Results structure — observational cohort (STROBE)
A structure model for the Results of an exposure→outcome cohort study. It follows its Methods STROBE sibling in Methods order. Each heading is a paragraph; each bullet is what it must establish. Fill the [brackets]; do not copy this text. Report findings only — no interpretation.
Cohort assembly (Figure 1)
- The assembly cascade as Figure 1: source → eligible → analytic, with exclusion **counts and
reasons**, and the exposed/unexposed counts.
- The cascade reconciles (start − Σ exclusions = analytic N); follow-up time / person-years
stated. Numbers match the Abstract, Methods, and Table 1.
Baseline characteristics, by exposure (Table 1)
- Table 1 stratified by exposure status, with a balance metric (standardized mean
differences) and the imbalance threshold used; if matched/weighted, show pre- and post-balance.
- Show every pre-specified adjustment covariate in this table, so the reader can see whether
any covariate imbalanced by exposure is missing from the adjustment set.
Primary association — crude then adjusted
- The primary contrast and effect measure (OR / RR / HR) with 95% CIs, reported **crude
then adjusted** (or matched), so the effect of confounding control is visible.
- Report the pre-specified primary estimand; do not reassign the primary after seeing
results. State events/N for each arm.
Secondary, subgroup, and dose-response
- Subgroup / interaction estimates with CIs (report the interaction term, not two separate
stratum estimates, when effect-modification is the question).
- Dose-response across exposure levels if available; keep multiplicity in view.
Sensitivity analyses
- The robustness checks declared in Methods (alternative adjustment, complete-case vs
imputation, tipping-point / E-value for unmeasured confounding), reported as results, not deferred.
Missing data
- How much was missing and how the analytic set was reached (complete-case reconciles:
total − missing = complete), with the footnote N matching the flow.
Common omission
- Crude-and-adjusted side by side and an explicit **sensitivity analysis for unmeasured
confounding** (E-value / tipping-point) — the elements cohort Results most often skip. Cross-reference section_guides/results.md and the STROBE critical items in peer-review/references/reviewer_calibration/compliance_floor.md.
Exemplar Results sections — structure models, not prose to copy
write-paper teaches Results rules (section_guides/results.md: mirror-symmetry with Methods, the flow diagram, no interpretation) but had no worked example of what a complete, well-ordered Results section looks like end to end. This directory fills that gap with structure exemplars that follow the exemplar_methods/ siblings in Methods order: each shows the Results paragraph order for a study type and, for every paragraph, what it must establish — with placeholder specifics, never copied text.
These are authored from scratch as teaching models, not extracted from any published paper. Use them to check that a draft's Results presents every load-bearing element in the same order as its Methods; do not copy wording.
How they are used
In write-paper Phase 4 (Results), after loading section_guides/results.md: pick the exemplar matching the study type and confirm the draft presents each element in Methods order (flow → baseline/prevalence → primary estimate with CIs → calibration/agreement → subgroups/comparator → sensitivity/missing-data). A missing element is a gap to fix; an element that interprets rather than reports belongs in the Discussion.
Contents
diagnostic_accuracy_stard.md— flow, prevalence, sensitivity/specificity (+ PPV/NPV) with
CIs at a fixed operating point, agreement, comparator (STARD).
ai_validation_tripod_claim.md— partition/flow, per-set discrimination with CIs,
calibration, operating point, clinical utility (TRIPOD+AI / CLAIM).
observational_cohort_strobe.md— assembly, Table 1 by exposure, crude-then-adjusted
estimate with CIs, subgroups, sensitivity analyses (STROBE).
Curator guidelines (for adding more)
- Synthetic only. Write the structure with placeholder specifics (
[N],[YYYY–YYYY],
[metric]); never paste a real paper's Results, use no real citations, no PII, English only.
- Follow the matching `exemplar_methods/` file in Methods order (same study type), so
Results and Methods read in the same sequence.
- One study type per file, each line stating what the paragraph must establish (not
finished prose), and report-only — flag interpretation as belonging to the Discussion.
- Name the common omission for each study type, and anchor to the reporting-guideline
critical items in peer-review/references/reviewer_calibration/compliance_floor.md.
- Keep each file ~50–90 lines. Cross-reference
section_guides/results.mdrather than
duplicating it.
Journal Profile: Abdominal Radiology
Journal Identity
- Full name: Abdominal Radiology
- Abbreviation: Abdom Radiol
- Publisher: Springer Nature (Society of Abdominal Radiology, SAR)
- ISSN: 2366-004X (print), 2366-0058 (online)
- Frequency: Monthly (12 issues/year)
- Impact Factor: ~3.1 (JCR 2023)
- Open Access: Hybrid (Springer transformative agreements may cover OA)
- Acceptance rate: ~30%
- Peer review: Single-blind; typically 2 reviewers
Manuscript Types and Word Limits
| Type | Body Word Limit | Abstract | References | Figures/Tables |
|---|---|---|---|---|
| Original Article | 3500 words | 250 words (structured) | 35 | 6 |
| Review Article | 5000 words | 250 words | 60 | 10 |
| Technical Innovation | 2000 words | 150 words (unstructured) | 15 | 4 |
| Brief Report (incl. case reports) | 1500 words | 150 words | 10 | 4 |
| Pictorial Essay | 3000 words | 200 words | 20 | 12 |
| Letter to the Editor | 500 words | None | 5 | 1 |
Word counts exclude abstract, references, tables, and figure legends.
---
Abstract Requirements
Structured abstract for Original Articles, 250 words maximum:
Purpose: [Study aim]
Methods: [Design, population, imaging, analysis]
Results: [Key findings with statistics]
Conclusion: [Main conclusion — 1-2 sentences]Unstructured abstract for Technical Innovations and Brief Reports, 150 words.
---
Required Sections (Original Article)
1. Introduction — clinical context, gap, study purpose (2-3 paragraphs) 2. Methods
- Study Design: IRB, retrospective/prospective
- Patient Population: inclusion/exclusion criteria, time period
- Imaging Protocol: scanner, sequence parameters, contrast protocol
- Image Analysis: readers, blinding, measurement methods, structured reporting systems used
- Statistical Analysis: software, tests
3. Results — demographics, imaging findings, performance metrics 4. Discussion — interpretation, comparison with literature, clinical relevance, limitations 5. Conclusion
---
Statistical Reporting
- Report exact p-values; use p < 0.001 below that threshold.
- 95% CI for primary outcomes.
- For diagnostic accuracy: sensitivity, specificity, PPV, NPV, AUC with 95% CI.
- Inter-reader agreement: ICC or kappa with model specification.
- For structured reporting studies (LI-RADS, PI-RADS): report per-category distribution and diagnostic performance per category.
- Multiple comparisons: correction method stated.
- Statistical software and version must be identified.
---
Abdominal Imaging-Specific Requirements
Structured Reporting Systems
Abdominal Radiology is a key venue for studies using standardized reporting:
- LI-RADS (Liver Imaging Reporting and Data System)
- PI-RADS (Prostate Imaging Reporting and Data System)
- Bosniak classification (renal cysts)
- O-RADS (Ovarian-Adnexal Reporting and Data System)
- TI-RADS (Thyroid Imaging Reporting and Data System — when thyroid US)
When using these systems, report the version used and per-category results.
Imaging Protocol
Report for each modality:
- CT: scanner, kVp, mAs, slice thickness, reconstruction kernel, contrast protocol (volume, rate, timing)
- MRI: field strength, coil, sequences with parameters, contrast agent and dose
- US: transducer frequency, CEUS contrast agent if applicable
---
Figures
- Maximum 6 figures/tables for original articles; 12 for Pictorial Essays
- Resolution: 300 DPI minimum
- Format: TIFF, EPS, JPEG
- Color: Free online
- Annotations: arrows with legend, window/level appropriate for lesion visibility
---
Common Rejection Reasons
1. Topic outside abdominal/pelvic scope — chest, MSK, and neuro papers redirected 2. Structured reporting studies without per-category analysis — reporting LI-RADS categories without stratified performance is PARTIAL 3. Missing imaging protocol detail — essential for reproducibility in abdominal imaging 4. Small sample without novel technique — needs sufficient cases per diagnostic category 5. Overlap with themed issue — check SAR disease-focused panels schedule to align or avoid
---
Cover Letter
Should include:
- Relevance to abdominal/pelvic radiology
- Key finding summary
- Alignment with SAR focus areas (if applicable)
- Statement of originality
---
Author Guidelines URL
https://www.springer.com/journal/261/submission-guidelines
---
Positioning
Abdominal Radiology is appropriate when:
- Abdominal/pelvic imaging study (CT, MRI, US) with clinical relevance
- LI-RADS, PI-RADS, or other structured reporting validation study
- AI applied to abdominal imaging with diagnostic performance evaluation
- Contrast-enhanced imaging technique comparison (hepatobiliary, CEUS)
- Pictorial essay on abdominal pathology
Not appropriate for: non-abdominal imaging, pure methodology without abdominal imaging data, interventional procedures (use CVIR/JVIR).
Academic Radiology
Journal Identity
- Full name: Academic Radiology
- Abbreviation: Acad Radiol
- Publisher: Elsevier (Association of Academic Radiology)
- ISSN: Not specified in guidelines
- Frequency: Monthly
- Impact Factor: Not specified in guidelines
- Open Access: Hybrid (subscription + OA option)
- APC: See Elsevier OA policies
Manuscript Types and Word Limits
| Type | Abstract | Manuscript Body | References | Figures/Tables |
|---|---|---|---|---|
| Original Investigation | 200 w (structured) | No strict cap | No cap | No strict cap |
| Clinical Review / Systematic Review | 250 w | 5,000 | 100 | 5 T + 10 F |
| Innovations | 200 w | 2,500 | 15 | 5 |
| Perspectives | None | 2,000 | 20 | 2 |
| Beyond Academic Radiology | None | 1,000 | 10 | 2 |
| White Papers | None | 5,000 | 50 | 5 |
| Letter to the Editor | None | 500 | 5 | 0 |
Word counts exclude abstract, references, figure legends, and tables.
Abstract Format
Structured with four headings (for Original Investigations): 1. Background 2. Methods 3. Results 4. Conclusion
Maximum 200 words (Original Investigation) or 250 words (Clinical Review). No abbreviations unless established. No references in abstract.
Keywords
3-5 keywords on the title page.
Required Sections
1. Introduction 2. Material and Methods 3. Results 4. Conclusions
Additional required elements:
- Take-Home Messages: 1-2 bullet points (max 50 words total) for Original Investigations; up to 3 for Clinical Reviews
- Title page (separate): title, authors, affiliations, corresponding author, keywords, disclosures, funding, AI declaration
- Blinded manuscript (separate): no author-identifying information (double-blind review)
- Acknowledgements (separate file)
Citation Style
- Vancouver (numbered) style
- Numbered sequentially in order of first appearance
- DOI highly encouraged
- Journal abbreviations per ISSN List of Title Word Abbreviations
- Example:
Boyko A, Qureshi MM, Fishman MDC, Slanetz PJ. Predictors of breast cancer outcome... Acad Radiol. 2024;31(5):1727-1734. doi: 10.1016/j.acra.2023.11.037.
Reporting Guidelines
| Study Type | Required Guideline |
|---|---|
| Systematic reviews / meta-analyses | PRISMA |
| Observational studies | STROBE |
| Quality improvement | SQUIRE |
| AI/ML studies | CLAIM |
| Randomized trials | CONSORT |
Statistical Reporting
Follows standard biomedical statistical reporting conventions:
- Report exact p-values; use P < .001 only when value is below that threshold.
- 95% CI required for all primary outcomes.
- For AI/ML studies: CLAIM checklist compliance required; discrimination and calibration both expected.
- Effect sizes with units; identify specific statistical tests used.
- Statistical software and version must be named in Methods.
- No journal-specific statistical reporting requirements beyond standard practice; SAMPL guidelines recommended as reference.
Special Notes
- Double-blind review: author information must be completely removed from the blinded manuscript file.
- AI declaration mandatory: generative AI use must be declared; failure to disclose leads to automatic rejection and possible submission ban.
- Co-first authors: up to two allowed; no co-senior authors accepted.
- Proof turnaround: 48 hours.
- Drug/instrument names: use generic name first, then trade name with manufacturer and location in parentheses.
Related skills
FAQ
Which paper types are supported?
Original articles, case reports, case series, meta-analyses, AI validation studies, animal studies, technical notes, and NHIS cohort / cross-national papers.
Does it pick a reporting guideline?
Yes. It maps the paper type to STARD/STARD-AI, TRIPOD+AI, CLAIM 2024, CONSORT/CONSORT-AI, PRISMA 2020, or STROBE at Phase 0.