
Academic Aio
- 46 installs
- 236 repo stars
- Updated August 3, 2026
- aperivue/medsci-skills
Academic AIO is a skill that optimizes medical AI papers and code releases for citation by AI search engines using generative engine optimization and reporting-guideline rules.
About
Academic AIO reviews titles, abstracts, structured summary boxes, manuscripts, README/CITATION.cff, and model cards so a medical AI paper is surfaced and cited accurately by AI search engines and RAG literature tools. A researcher uses it to apply generative-engine-optimization principles alongside reporting-guideline requirements. It outputs a visible pass/fail checklist with concrete edit suggestions rather than editing silently.
- Optimizes medical AI papers for AI search engines like Perplexity, Elicit, and Consensus
- Integrates TRIPOD+AI, CLAIM 2024, STARD-AI, and DECIDE-AI reporting rules with GEO principles
- Produces a visible PASS/PARTIAL/FAIL checklist instead of silent rewrites
Academic Aio by the numbers
- 46 all-time installs (skills.sh)
- Ranked #1,348 of 1,879 Marketing & SEO skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
academic-aio capabilities & compatibility
- Capabilities
- aeo optimization · seo audit · content optimization
- Use cases
- seo · research · documentation
What academic-aio says it does
Your output is a visible pass/fail checklist with concrete edit suggestions, not silent rewrites.
content structured for LLM extraction receives up to 40 % more visibility in generative engines
Surface the checklist in the response. Never apply AIO edits silently.
npx skills add https://github.com/aperivue/medsci-skills --skill academic-aioAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 46 |
|---|---|
| repo stars | ★ 236 |
| Last updated | August 3, 2026 |
| Repository | aperivue/medsci-skills ↗ |
What it does
Optimize a medical AI paper, preprint, or code release so it is discoverable and cited accurately by AI search engines and RAG literature tools.
Who is it for?
Making medical AI papers discoverable and accurately cited by AI search and RAG tools
Skip if: General web SEO for non-academic sites
When should I use this skill?
when drafting or reviewing a paper title, abstract, summary box, README, or model card for AI-search visibility
What you get
A pass/fail AIO checklist and edits that make a paper extractable and correctly cited by AI search engines
- AIO pass/fail checklist
- title and abstract edit suggestions
- structured summary box
By the numbers
- cites up to 40% more visibility in generative engines (Aggarwal 2024)
- reports 50-90% of medical LLM answers not fully supported by cited sources
Files
Academic AIO Skill — Medical AI Paper Visibility for AI Search Engines
You are helping a medical-AI researcher optimize a paper, preprint, README, or code release so that it is surfaced and cited accurately by AI search engines (Perplexity, ChatGPT web, Elicit, Consensus, SciSpace), RAG-based literature tools, and traditional scholarly indexes (Semantic Scholar, Google Scholar, PubMed). Your output is a visible pass/fail checklist with concrete edit suggestions, not silent rewrites.
Communication Rules
- Surface the checklist in the response. Never apply AIO edits silently.
- Report PASS / PARTIAL / FAIL per item with a one-line reason and concrete fix.
- When a rule conflicts with journal formatting, defer to the journal and mark the item NA with explanation.
- Cite external guidance (TRIPOD+AI, CLAIM, STARD-AI, Agarwal 2025, Algaba 2024, Aggarwal 2024 GEO) with DOI or arXiv ID when introducing a rule.
- Do not hallucinate citations. If unsure, mark as
[VERIFY].
When to Invoke
Run this skill when the user is working on any of:
- Drafting or revising a title, abstract, structured-summary box, or plain-language summary.
- Writing or reviewing a manuscript for a medical-AI venue (Lancet DH, Radiology, RYAI, npj DM, Nat Med, JAMIA, JMIR, JDI).
- Preparing a preprint (medRxiv, arXiv, bioRxiv, Research Square).
- Composing a GitHub README,
CITATION.cff, Zenodo archive metadata, Hugging Face model card, or dataset card. - Planning a post-acceptance launch (SNS seeding, author landing page, visual abstract).
- Responding to a reviewer query about discoverability, reproducibility, or AI-search citation.
Pairs with (do not duplicate):
write-paper— Phase 6 (draft) and Phase 7 (QC). AIO rules extend the title/abstract/discussion sections.check-reporting— reporting-guideline item audit (TRIPOD+AI, CLAIM, etc.). AIO requires guideline adherence but does not reproduce the audit.self-review— adversarial review. Run AIO after self-review so QC-confirmed claims anchor the checklist.humanize— AI-pattern removal. Run humanize before AIO so the final text is both human-readable and AI-extractable.
Core Thesis
Generative engine optimization research (Aggarwal 2024, arXiv:2311.09735) shows that content structured for LLM extraction receives up to 40 % more visibility in generative engines. In medicine this effect is mediated by three gates:
1. Open-access full text — tools like Elicit and Consensus cannot extract columns from paywalled PDFs; Perplexity Academic favors OA citations. 2. Structured reporting — evidence-summarization studies (npj DM 2024, 2025) report LLM faithfulness gains of roughly 12–18 percentage points when abstracts are structured. 3. Machine-readable artifacts — CITATION.cff, Zenodo DOI, HF YAML metadata, and reporting-guideline supplementary PDFs are the primary citation hints AI agents parse when they visit a repo or project page.
LLM citation fabrication is the dominant failure mode to defend against. Agarwal et al. (Nat Commun 2025, doi:10.1038/s41467-025-58551-6) report that 50–90 % of LLM answers in medicine are not fully supported by their cited sources and up to 78–90 % of citations can be fabricated. The defensive strategy is to surface a paper's DOI and PMID in easy-to-copy form so that LLMs substitute the correct identifier instead of confabulating one.
Section 1 — Title and Abstract Optimization
1.1 Title three-slot rule
Structure: [Task] + [Modality or anatomy] + [Model family or method class]. Include one concrete differentiator (dataset scale, new benchmark, "first …") when defensible. Avoid keyword stuffing (penalized as spam by AI overviews).
Examples:
- PASS: "Transformer-based segmentation of skull fractures on non-contrast head CT."
- FAIL: "A novel advanced deep-learning AI machine-learning framework for medical image analysis."
1.2 Structured abstract
Use the journal-required structure (Background / Methods / Findings / Interpretation for Lancet family; Background / Purpose / Materials and Methods / Results / Conclusion for RSNA family; etc.). If the journal allows unstructured, still use an internally structured form. Each section stands alone as a semantic chunk of ≤ 3 sentences so that chunk-boundary splits in RAG indexes do not break the claim.
1.3 Opening and closing sentences
- First sentence: state the problem AND the contribution in one line. LLM summarizers extract this disproportionately.
- Last sentence: explicit interpretation ("we show that …", "this implies …"). No hedging-only closes.
1.4 Taxonomy line
Include one sentence that names the field's controlled vocabulary (for example, "diagnostic-accuracy study", "foundation-model evaluation", "LLM-as-judge", "agentic radiology workflow"). Entity linkers in AI indexes use this line.
1.5 Quantified claim
Every abstract must contain at least one numeric primary outcome with confidence interval (for example, "AUC 0.94 [95 % CI 0.91–0.96]" or "sensitivity 88.2 % [95 % CI 85.1–91.0]"). LLM retrievers weight papers with concrete numbers.
1.6 Reporting-guideline anchor
Place the guideline name in the abstract or the opening sentence of Methods: "Reported following TRIPOD+AI (Collins 2024) and CLAIM 2024 (Tejani 2024)". When applicable add STARD-AI 2025, DECIDE-AI, TRIPOD-LLM. This signals structure to LLMs and satisfies reviewer checklists.
AIO-rule ↔ guideline-item mapping: references/reporting_guideline_mapping.md.
1.7 Keyword, MeSH, and RadLex coverage
Title, abstract, and keywords together should cover ≥ 3× the surface area of the concept — no redundancy. Include:
- Core MeSH terms (verify against the NLM MeSH browser).
- Radiology-specific RadLex terms where applicable.
- Modality-synonym coverage ("chest radiograph (CXR)", "non-contrast CT (NCCT)").
- Both US and UK spellings when relevant.
Royal Society 2024 (doi:10.1098/rspb.2024.1222) reports that 92 % of papers waste keyword real estate by repeating title terms in abstract and keywords; avoid this.
Section 2 — Manuscript-Level AIO
2.1 Summary box
Include the journal-specific summary box verbatim when supported:
- Lancet family: "Research in context" (Evidence before this study / Added value / Implications).
- RSNA Radiology and RYAI: "Key Points" — 3 bullets, one claim each.
- npj Digital Medicine: "Plain-language summary" (150–200 words, 8th-grade reading level).
- Nature Medicine: editor's summary (supplied by editorial, but draft one proactively).
Deterministic format check. Validate the drafted box against its journal spec with python3 ${CLAUDE_SKILL_DIR}/scripts/check_summary_box.py --manuscript <file> --journal <stem> --strict (reads references/summary_box_specs.json: Key Points bullet count + one-claim-per-bullet, Research-in-context's three sub-blocks, plain-language word band). It catches the wrong-format / wrong-bullet-count box that a production technical check rejects.
These boxes are the fragments Perplexity and ChatGPT web most often copy or paraphrase verbatim; treat them as the paper's canonical citation surface.
Journal-specific templates (USER MUST VERIFY against current IFA): references/journal_summarybox_templates.yaml.
2.2 Declarative section headings
Section and subsection headings should state a claim, not a generic label. "Model underperforms on rare-finding subset" beats "Subgroup analysis".
2.3 Numeric claim compression
In the Methods and in at least one Results paragraph, compress primary-outcome statistics into a single sentence pattern: "On the internal test set (n = 842), the model achieved AUC 0.94 (95 % CI 0.91–0.96), sensitivity 88.2 % (85.1–91.0), specificity 91.4 % (88.7–93.6), at an operating point of 0.37."
This pattern is the canonical shape LLM extractors parse first.
2.4 Reproducibility block
Include a labeled block (typically end of Methods or a standalone Data/Code Availability section) listing: data availability and license, code availability with DOI, model weights and checkpoints, prompts and configuration files, random seeds, compute environment. This block is disproportionately scraped by AI agents when they cite a paper as reproducible.
2.5 Limitations enumeration
List limitations explicitly and name each one (generalizability, spectrum bias, dataset shift, single-center training, label noise). Papers with enumerated limitations score higher for trustworthiness in LLM summarization benchmarks.
2.6 Standalone figure captions
Each caption should re-state the claim, the dataset, and the metric. Captions survive in vector databases and image-retrieval indexes when surrounding body text is lost.
Section 3 — Preprint, Channel, and Indexing Strategy
3.1 Preprint versus fast-track
- Default: post to medRxiv (clinical), arXiv (methods, cs.CV / eess.IV), or bioRxiv on the day of journal submission. Rapid preprinting puts the paper into Semantic Scholar within 24–72 hours and into Perplexity's web index immediately.
- Exception: if the target journal offers a fast-track review cycle (acceptance → online within roughly 30–60 days) AND the authors prefer a single canonical version, a preprint may be skipped. In that case, compensate by aggressive post-acceptance SNS seeding and PMC deposit.
- Never skip preprint AND fast-track — this is the discoverability deadzone.
3.2 Journal preprint-policy table (verify before submission)
Most medical-AI venues allow preprints (Radiology, RYAI, Lancet DH, npj DM, Nature Medicine, JAMIA, JMIR, Cell Reports Medicine, Cell Patterns). A few have restrictions or require disclosure. Always verify the current policy on Sherpa Romeo or the journal's instructions-for-authors page before posting.
3.3 Indexing time-lag (2025 baseline)
- Perplexity Academic / ChatGPT web: real-time web crawl, citable on publication day.
- Semantic Scholar: 24–72 hours from DOI or preprint.
- Google Scholar: 1–7 days.
- PMC (NIH deposit): 2–6 weeks for accepted manuscripts; longer for CC-BY-NC.
- Elicit and Consensus: follow Semantic Scholar / OpenAlex.
- LLM training corpora (next model generation): 6–18 months.
Plan launch activities around these windows.
3.4 Open-access choice
Prefer gold OA with CC-BY when budget allows. If not, green OA via preprint plus author-accepted manuscript is acceptable. Closed-access papers without preprint lose roughly 30–50 % of AI-tool citations because Elicit, Consensus, and Perplexity Academic cannot extract from paywalled PDFs.
Funder OA-policy decision tree (Plan S, NIH, UKRI, Gates, Wellcome, NRF, MoHW): references/oac_funding_checklist.yaml.
3.5 Post-acceptance channel checklist
- Deposit AAM to PMC or Europe PMC.
- Update ORCID and Google Scholar profile.
- Post to Threads / X / BlueSky with DOI, one-sentence claim, and figure.
- Long-form post on LinkedIn (targets different LLM training corpora).
- Submit to Papers with Code if the paper reports a benchmark.
- Upload model or dataset to Hugging Face with a model/dataset card.
Section 4 — Review-Paper Strategy
Review articles function as hub nodes in knowledge graphs and accrue "lookup citations" when readers need a canonical reference for a taxonomy. For researchers building a portfolio in medical AI:
- Target at least one review or taxonomy paper per year in a top-tier venue.
- Include 5 or more taxonomy tables (model class, dataset, task type, evaluation metric, failure mode). Each table becomes a lookup target.
- Cite 100 or more primary references for breadth; 150+ for canonical status.
- Co-author with a consortium of 10+ investigators from multiple institutions when possible — this multiplies social-network reach and citation dispersal.
- Pair the review with a companion dataset, benchmark, or code artifact on Zenodo or Hugging Face to anchor AI-tool citations.
Empirically, review papers with these properties outperform original research on short-term FWCI while feeding traffic to the authors' original papers through reverse citation.
Section 5 — GitHub, CITATION.cff, Zenodo, Hugging Face
5.1 README canonical 10-slot order
1. Title + one-line description + badges (license, DOI, arXiv, Hugging Face, paper link). 2. Paper reference block — BibTeX + APA + two-sentence abstract. 3. TL;DR — at most 5 bullets: problem, approach, key result, intended users. 4. Quickstart — pip install or git clone && make demo. Should work in under 5 minutes. 5. Reproducibility — exact commands that regenerate every figure and table. Pin package versions. 6. Project structure — a tree with one-line folder descriptions. 7. Data access — license, download scripts, DUA notes. 8. FAQ — "How is this different from X?", "Can this be used clinically?", "How do I cite this?". High-value retrieval content. 9. Acknowledgements, funding, and COI. 10. License (prefer Apache-2.0 for research code).
5.2 CITATION.cff
Add a CITATION.cff file at repository root. GitHub renders it as a "Cite this repository" button, and AI agents treat it as the primary citation hint. Include authors with ORCID, title, version, DOI (post-Zenodo-archive), repository URL, and license.
5.3 Zenodo DOI
Enable GitHub–Zenodo integration for each release. Cite the version-specific DOI in the paper's Data/Code Availability section. Zenodo deposits appear in Google Scholar and OpenAlex, creating an independent citable artifact.
5.4 Hugging Face model card YAML
Required keys: license, library_name, tags, datasets, base_model (when fine-tuning), pipeline_tag. Required prose sections: Intended use, Training data, Evaluation, Limitations, Ethical considerations, and a clinical-use disclaimer ("This model is not approved for clinical diagnostic use; it is provided for research purposes only").
5.5 Hugging Face dataset card
Required prose: license, PHI and re-identification risk, task, language, splits, annotation process, known biases, ethical review status.
5.6 Web-crawler-friendly formatting
- Markdown headings are declarative claims.
- Code blocks are fenced and language-tagged.
- Tables are plain Markdown, not HTML (survive Markdown-to-vector chunking).
- Images have descriptive alt text (vision-LLMs read alt text when image retrieval fails).
- Each README section is under about 300 words to survive fixed-size chunking.
- Use question-style subheadings when natural ("Why another benchmark?", "How fast is inference?").
- Embed JSON-LD
ScholarlyArticle/SoftwareSourceCode/Dataset/Personmarkup in repository pages and author landing pages — templates inreferences/schema_markup_templates/, validated withpython scripts/validate_schema.py path/to/file.jsonld.
Section 6 — Authority and E-E-A-T Signals
- Maintain a personal author landing page (GitHub Pages, personal domain, or institutional page) that lists all papers with DOIs and open-access links. AI indexes weight author-entity pages.
- Use one consistent affiliation string across papers. Inconsistency fragments the author entity in knowledge graphs and loses citation velocity.
- Keep ORCID complete and linked to Google Scholar. Re-run author-disambiguation on Semantic Scholar every 6 months.
- Cross-link related papers by the same group in Discussion sections when defensible. Within-group self-citation increases co-retrieval probability in RAG.
- Refresh repository and model cards quarterly — articles updated quarterly outperform single-publish articles in AI-overview retention (Conductor 2026 benchmark).
Section 7 — LLM-Citation Fabrication Defense
Given Agarwal et al. Nat Commun 2025 (doi:10.1038/s41467-025-58551-6) findings that up to 78–90 % of LLM medical citations can be fabricated, take the following defensive steps:
- Surface DOI and PMID in copy-friendly text at the top of the paper's landing page and README (for example,
DOI: 10.xxxx/yyyy • PMID: 12345678). - Add a "How to cite" section with BibTeX, APA, Vancouver, and the plain-text line in one place.
- Monitor incorrect citations. Set a Google Scholar alert for the paper's title variant; periodically query Perplexity and ChatGPT web for the paper and record hallucinated bibliographic errors.
- When responding to a reviewer who cites an LLM-generated reference, verify the DOI and PMID yourself before accepting.
Section 8 — Red Flags
- Closed code described as "available on reasonable request" — scrapers treat this as "not reproducible" and AI tools demote the paper.
- Paywall-only with no preprint and no fast-track — invisible to most RAG pipelines.
- Keyword-stuffed titles ("A deep-learning artificial-intelligence machine-learning system for …") — penalized as spam.
- Abstracts opening with filler ("In recent years, AI has revolutionized …") — burns the chunk most likely to be extracted.
- Walls of theory before the README quickstart.
- "Clinical grade" or "replaces radiologists" overclaims — demoted by LLM trust heuristics and may trigger reviewer rejection.
- PHI leakage in Hugging Face dataset samples.
- Inconsistent author affiliations across co-authored papers.
Section 9 — Per-Project Application and Pipeline Integration
When invoked, run in this order:
1. Read the target artifact (title, abstract, manuscript section, README, or card) and identify its lifecycle phase: pre-draft / drafting / pre-submission / post-acceptance / post-publication. 2. Apply Sections 1–5 and 10 relevant to that artifact, filtering each rule by its `applies_to_phase` field in `references/checklists/AIO_GENERAL.md`. Out-of-phase rules become NA rather than FAIL (e.g., do not surface §11.5 multi-disciplinary roster or §12 launch sequencing as FAIL on a pre-submission audit). Produce a PASS / PARTIAL / FAIL table sorted by `expected_lift` (high → medium → low). Render via templates/aio_audit_checklist.md.j2 when programmatic. 3. Honour `defers_to` annotations to avoid duplicate audits. Items annotated with a defers_to field record only present/absent status here; item-level detail belongs to the linked skill or reference (§1.6 → /check-reporting; §3.4 / §11.3 → references/oac_funding_checklist.yaml). Cross-check reporting-guideline anchor (§1.6) by invoking /check-reporting first when the manuscript has not been audited; the AIO ↔ guideline-item mapping is in references/reporting_guideline_mapping.md. 4. Apply Section 6 author-authority audit once per submission cycle. Sections 11.1–11.5 are pre-draft rules — applies_to_phase filter auto-NAs them once drafting is complete. Section 12 launch sequencing fires only at post-acceptance / post-publication. 5. Surface Section 7 citation-defense recommendations at post-acceptance time. For multi-repo or Hugging-Face-card team audits, run scripts/batch_metadata_audit.py. 6. Output:
- The phase-filtered checklist (visible).
- A short deferred-item list with one-line status per
defers_torule. - At most 5 concrete edits ranked by `expected_lift` (high first, then medium, then low). Edits whose underlying rule is
low-lift should not appear in the Top 5 unless nohigh/mediumitems remain open.
Integration with write-paper
- Phase 4 (Title and abstract drafting) → apply Section 1 as an inline filter.
- Phase 6 (Discussion) → apply Section 2.5 (limitations) and Section 6 (cross-linking).
- Phase 7 (QC) → run AIO after reporting-guideline check and numerical-claim audit.
Output template
## Academic AIO Checklist — [Artifact type]
| # | Item | Status | Note |
|---|------|--------|------|
| 1.1 | Title three-slot | PASS/PARTIAL/FAIL | … |
| 1.2 | Structured abstract | PASS/PARTIAL/FAIL | … |
| ... | ... | ... | ... |
## Top 5 suggested edits
1. …
2. …Section 10 — Q&A and Entity-Extraction Optimization
Modern RAG indexes parse Q&A blocks more reliably than free-form prose; LLM citation engines preferentially extract claim-restatement pairs. Section 10 augments retrievability by structuring how claims are restated and how entities are linked.
10.1 Four-question Q&A block (Discussion or Appendix)
Add a labeled Q&A block — either as the closing subsection of Discussion, or as a Supplementary Box. Pattern:
- What was known before this study? — two-sentence restatement of the prior state.
- What does this study add? — two-sentence statement of the contribution.
- How might this change clinical practice or research? — one-sentence interpretation; avoid overclaim.
- Why does this matter? — one-sentence "so what" framing for non-specialists.
This block is the canonical fragment that AI-overview systems extract and cite. Lancet Digital Health "Research in context" already encodes the first two questions; the Q&A block extends them and is parseable by LLM web-search agents.
10.2 Glossary block with entity IDs
Define each domain-specific acronym inline on first use AND list them in a Glossary subsection at end of Methods or Supplementary. Attach the canonical entity ID where possible:
- MeSH term ID for clinical concepts.
- RadLex ID for radiology-specific terms.
- UMLS CUI for cross-vocabulary mapping.
- Hugging Face model ID for named models.
- arXiv ID for cited methods.
Entity linkers in Elicit, Consensus, and SciSpace use this metadata to connect a paper to knowledge graphs.
10.3 Inline citation anchor text
Avoid bare reference numbers. Use semantic anchor patterns so LLM extractors bind the citation to the specific claim:
- WEAK: "Prior work [12] showed efficacy."
- STRONG: "Smith et al. (DOI: 10.xxxx/yyyy) reported a 12 % accuracy gain on the MIMIC-CXR test set [12]."
When citing one's own prior work, name the cohort or dataset explicitly to enable cross-paper retrieval.
10.4 Explicit challenge statement
Beyond Section 2.5 (limitations enumeration), include a single-paragraph "Why this is hard" challenge statement near the start of Discussion. Pattern:
"Building accurate [task] for [modality/anatomy] is constrained by [data scarcity / label noise / dataset shift / regulatory uncertainty / interpretability]. Each of these has been documented [refs], and our results address [subset]."
LLM web-search systems quote challenge statements as authoritative summaries of field state. The 2025 KJR multimodal-LLM review used this pattern (e.g., "lack of large-scale high-quality multimodal datasets") and was preferentially extracted by Perplexity and ChatGPT web (see references/case_studies/kjr_mllm_2025.md).
Section 11 — First-Mover Timing and Citation-Graph Density
Topic timing is the most under-discussed AIO lever. Reviews and original research published at the peak of a topic's hype curve accrue citations disproportionately; reviews that lag the peak by 6–12 months under-perform regardless of quality.
11.1 Topic peak detection
Signals that a topic is approaching peak (write now, publish in ~6 months):
- arXiv/medRxiv monthly deposit rate growing > 20 % month-over-month for 3+ consecutive months.
- Major model release (GPT-4o, Claude 3.5 multimodal, MedGemini) introducing a capability not previously available.
- Funding agency Request-for-Applications (RFA) addressing the topic.
- Society guidelines (RSNA, ACR, ESR) calling for evaluation studies.
- Sustained > 1,000 weekly impressions on Twitter/X/LinkedIn for related papers.
Plan submission so publication lands at peak, not after.
11.2 Editorial-board leverage
If a corresponding author serves on the target journal's editorial board, review-process median time often drops noticeably (KJR: ~4–6 weeks faster; varies by journal). Editor's-pick or issue-highlight selection can also drive Google News indexing within 24 hours of publication.
When recruiting senior co-authors for a review paper, prefer those who hold an editorial role at the target venue. This is a legitimate editorial signal, not a conflict-of-interest issue, provided board members recuse themselves from review of their own submissions per ICMJE guidance.
11.3 PMC-auto-deposit journal preference
Open-access journals that automatically deposit to PubMed Central (PMC) reach LLM crawlers within 4–6 weeks of publication; non-PMC OA journals can take 3–6 months. PMC-auto-deposit journals in radiology/medical-AI (verify per submission, policies change):
- Korean Journal of Radiology (KJR) — auto-deposit confirmed.
- Lancet Digital Health — author-funded green OA, PMC-eligible after embargo.
- Radiology and Radiology: AI — selected articles auto-deposit.
- npj Digital Medicine — auto-deposit (Nature OA).
- JAMIA — author-funded OA route.
- JMIR — auto-deposit (PMC-indexed).
When all else is equal, prefer PMC-auto-deposit journals to compress the LLM-discoverability window.
11.4 Citation-graph anchor strategy
Discussion sections should anchor the paper in 5–10 high-visibility prior works that LLM training corpora already index well. This raises co-citation probability and makes the paper retrievable when users query the seminal works.
- Identify seminal references via Semantic Scholar's "Highly Influential Citations" filter for the topic.
- Cite them with semantic predicates (Section 10.3), not as bare lists.
- Mix recent preprints (currency signal) with 2018–2022 seminal papers (graph anchoring) — corpora-cutoff means 2024–2025-only citation profiles have low LLM retrieval weight.
11.5 Multi-disciplinary author roster
Author-affiliation diversity multiplies indexing entry points. A 10–15 author team spanning 3+ institutions and 2+ disciplines (clinical + computational) creates more author-entity nodes in Google Scholar and Semantic Scholar, each acting as a discovery surface. The 2025 KJR MLLM review used a 15-author team spanning resident + engineer + medical student + faculty across 5 institutions and accrued 64 citations within 7 months (see case study).
Section 12 — Cross-Platform Launch Sequencing
Section 3.5 (post-acceptance channel checklist) is unordered; Section 12 prescribes the timing. The first 30 days after publication are the primary discoverability window for AI-search engines and LLM training-data harvesters.
12.1 Day 0 — publication day (execute simultaneously)
- GitHub release (tag a stable version; let Zenodo mint a version-specific DOI).
- Hugging Face model card + dataset card (if applicable); link arXiv ID and DOI.
- Twitter/X + Threads + Bluesky: 1-sentence claim + key figure + DOI in copy-friendly format.
- LinkedIn announcement (long-form): hook line + structured claim block + DOI.
- Author landing-page update with PDF link (OA) or AAM.
12.2 Day 1 — propagation
- Update ORCID with DOI, abstract, and authorship role.
- Update Google Scholar (verify auto-detection within 24h; manual add if delayed).
- Update preprint server with "Accepted" version note + link to published version.
- Update institutional profile / department news page.
12.3 Week 1 — depth posts
- LinkedIn second post: long-form interpretation or methods spotlight.
- Papers with Code submission (if benchmark or model with public weights).
- ResearchGate upload of AAM (per journal policy).
- Reddit/Hacker News post if the work has broad appeal (assess fit honestly).
12.4 Weeks 2–4 — refresh signals
- README and HF card minor update (new badges, new FAQ entries).
- Follow-up blog or Substack post expanding on one figure or limitation.
- Respond to reader questions on social platforms — those answers themselves become indexed content.
12.5 Month 1 — monitoring
- Google Scholar alert for the paper title.
- Semantic Scholar / Scite citation alerts.
- Quarterly probe: query Perplexity, ChatGPT web, Elicit, Consensus, SciSpace with 3–5 expected discovery queries; record retrieval position and any hallucinated bibliographic errors.
- If a fabricated citation appears, update the README "How to cite" block (Section 7) to maximize copy-friendliness of the correct identifier.
External References
- GEO: Generative Engine Optimization — Aggarwal et al., KDD 2024, arXiv:2311.09735.
- LLM medical citation fabrication — Agarwal et al., Nat Commun 2025, doi:10.1038/s41467-025-58551-6.
- LLM citation bias — Algaba et al., 2024, arXiv:2405.15739.
- ExpertQA attribution — Malaviya et al., 2024, arXiv:2309.07852.
- TRIPOD+AI — Collins et al., BMJ 2024. EQUATOR Network.
- CLAIM 2024 — Tejani et al., Radiology: AI 2024, doi:10.1148/ryai.240300.
- STARD-AI — Sounderajah et al., Nat Med 2025, doi:10.1038/s41591-025-03953-8.
- TRIPOD-LLM — Gallifant et al., Nat Med 2024, doi:10.1038/s41591-024-03425-5.
- DECIDE-AI — Vasey et al., Nat Med 2022, doi:10.1038/s41591-022-01772-9.
- Title, abstract, keywords guide — Royal Society Proc B 2024, doi:10.1098/rspb.2024.1222.
- GitHub repository citation advantage — Yan et al., Inf Process Manag 2024, doi:10.1016/j.ipm.2023.103569.
- Semantic Scholar Open Data Platform — Kinney et al., arXiv:2301.10140.
Anti-Hallucination
- Never fabricate citations, DOIs, arXiv IDs, or reporting-guideline item numbers. Every cited reporting framework (TRIPOD+AI, CLAIM, STARD-AI, TRIPOD-LLM, DECIDE-AI) must map to a verifiable DOI or EQUATOR Network entry. Mark unverified items as
[UNVERIFIED - NEEDS MANUAL CHECK]. - Never invent journal-specific summary-box rules (Lancet Digital Health "Research in context", Radiology "Key Points", npj Digital Medicine). Verify current instructions-to-authors from the journal's website before applying.
- Never fabricate discoverability metrics (Perplexity/Elicit/Consensus retrieval scores) — only report observed behavior from a recorded probe.
- Never auto-complete author lists, ORCIDs, or affiliations in CITATION.cff or Zenodo metadata; surface empty slots to the user.
- If a compliance item, journal policy, or AI-search platform behavior is uncertain, state the uncertainty rather than guessing.
Case Study — KJR Multimodal LLM Review (2025)
Post-mortem of an unexpectedly high-engagement medical-AI review. Used to validate Sections 10–12 of the academic-aio skill. Identifying co-author and institutional details are abstracted; the paper itself is publicly indexed.
Paper
- Title: Multimodal Large Language Models in Medical Imaging: Current State and Future Directions
- Venue: Korean Journal of Radiology, Vol 26, Issue 10, pp. 900–923 (October 2025)
- DOI: 10.3348/kjr.2025.0599
- PMC ID: PMC12479233
- OA status: Creative Commons (gold OA), KJR + PMC dual indexing
- Author roster: 15-author multi-disciplinary team — resident + medical-AI researcher (first author), software engineer, medical student, faculty across 5 Asian academic medical institutions and one university computational biology department, anchored by a mid-career corresponding author who serves on the target journal's editorial board (Technology section).
- Timeline: Received 2025-05-14 → revised 2025-07-03 → accepted 2025-07-08 → published October 2025 (~5 months end-to-end, fast for a review article).
Engagement metrics (May 2026, ~7 months post-publication)
- Page views: 3,016
- PDF downloads: 628
- Citations: 64
- First-author Google Scholar impact (since 2021): h-index 4, i10-index 2, total 101 citations — of which ~63 % are attributable to this single review.
These figures place the paper in the top decile of KJR articles by 12-month citation accrual.
Driver analysis (which AIO levers actually operated)
The author's initial hypothesis space included (a) journal visibility, (b) hot topic, (c) corresponding author's reach, (d) multi-disciplinary team, (e) fast turnaround, (f) OA + PMC indexing, and (g) early SNS exposure. Post-hoc evidence:
| Hypothesis | Verdict | Notes |
|---|---|---|
| AI-search retrievability (Perplexity / ChatGPT web / Elicit / Consensus / SciSpace) | Strong (primary driver) | Paper appears at #3–4 for "MLLM medical imaging review 2025" on Google web; Perplexity preferentially cites it for queries on multimodal radiology AI. |
| OA + PMC dual indexing | Strong | Full text crawled by AI agents within 4–6 weeks; corresponds to Section 11.3. |
| Topic peak timing | Strong | Submitted shortly after GPT-4o and Claude 3.5 multimodal launches; published at the peak of the 2025 MLLM hype cycle. Corresponds to Section 11.1. |
| First-mover review | Strong | Competing reviews (in JBI, Archives of Comp Methods, ScienceDirect) appeared later or in less-discoverable venues. |
| Editorial-board signal | Moderate | Corresponding author's editorial role plausibly accelerated review and raised editor's-pick probability. Corresponds to Section 11.2. |
| Multi-disciplinary 15-author team | Moderate | Author-entity diversification across institutions multiplied Google Scholar entry points. Corresponds to Section 11.5. |
| Early SNS exposure | Weak | No clear evidence that SNS drove the bulk of traffic; KJR/PMC pathway dominated. |
Structural features that AI-search systems extracted
The article body exhibits several patterns that map to Sections 10–11 of this skill:
1. Explicit taxonomy (2D vs 3D MLLM; Applications vs Barriers) — RAG systems chunk-cite each branch independently. 2. Verbatim challenge statements — phrases such as "lack of large-scale high-quality multimodal datasets" and "hallucinated findings" are quoted by Perplexity and ChatGPT web as authoritative summaries of field state. Corresponds to Section 10.4. 3. Declarative subsection headings — claim-style rather than generic labels (Section 2.2). 4. 10 figures + 5 tables — rich pull-quotable structures for retrieval. 5. DOI + PMID surfaced cleanly in KJR landing page header — minimizes citation fabrication (Section 7).
Lessons codified into the skill
| Skill update | Source observation |
|---|---|
| §10.4 Explicit challenge statement | Field-state quotes were the most common Perplexity extraction. |
| §11.1 Topic peak detection | Submission timing aligned with model-release shockwave. |
| §11.2 Editorial-board leverage | Review-process speed plausibly editor-mediated. |
| §11.3 PMC-auto-deposit preference | KJR-PMC pathway delivered LLM crawl in 4–6 weeks. |
| §11.4 Citation-graph anchor | Discussion seeded with a mix of seminal multimodal-LLM works. |
| §11.5 Multi-disciplinary roster | 15-author 5-institution composition multiplied entry points. |
Replication checklist (apply to next medical-AI review)
- [ ] Identify a topic 6–12 months ahead of expected peak using Section 11.1 signals.
- [ ] Recruit a corresponding author with editorial-board affiliation at a PMC-auto-deposit OA journal (Section 11.2 + 11.3).
- [ ] Assemble a 10–15 author team spanning ≥ 3 institutions and ≥ 2 disciplines (Section 11.5).
- [ ] Draft taxonomy headings as declarative claims (Section 2.2) with explicit 2D-vs-3D-style subdivisions.
- [ ] Insert a "Why this is hard" challenge paragraph at the start of Discussion (Section 10.4).
- [ ] Add a four-question Q&A block in Discussion or Appendix (Section 10.1).
- [ ] Anchor Discussion in 5–10 highly-cited prior works (Section 11.4).
- [ ] Execute Day-0/Day-1/Week-1 launch sequencing (Section 12).
- [ ] Probe Perplexity / ChatGPT web / Elicit / Consensus / SciSpace at +30, +60, +90 days post-publication; correct any fabricated citations encountered (Section 7).
Caveats
- Single-paper case study — not a controlled experiment. Hot-topic timing alone could explain a large fraction of the engagement; the AIO mechanisms identified here are necessary but not provably sufficient.
- Citation accrual at 7 months is a leading indicator, not a final one. A second readout at 24 months will distinguish landscape-review citation behavior from short-term hype.
- Editorial-board involvement is institution-specific. Replicating Section 11.2 requires evaluating the target journal's recusal policy and disclosing it appropriately.
References
- KJR landing page:
https://www.kjronline.org/DOIx.php?id=10.3348/kjr.2025.0599 - PMC full text:
https://pmc.ncbi.nlm.nih.gov/articles/PMC12479233/ - GEO framework: Aggarwal et al., KDD 2024, arXiv:2311.09735.
- LLM medical citation fabrication: Agarwal et al., Nat Commun 2025, doi:10.1038/s41467-025-58551-6.
Academic AIO General Checklist
Machine-readable checklist mirroring SKILL.md sections 1–12. Use as the canonical pass/fail audit table for any medical-AI artifact (manuscript, preprint, README, CITATION.cff, HF card). Each item maps to a SKILL.md section so audit findings can be traced back to the rule.How to use
1. Copy this checklist into the working artifact directory as qc/aio_audit.md. 2. Determine the artifact's lifecycle phase (pre-draft, drafting, pre-submission, post-acceptance, post-publication) and filter rules whose applies_to_phase does not include the current phase — those become NA (do not surface as FAIL). 3. Mark each remaining item PASS / PARTIAL / FAIL with a one-line reason. 4. For FAIL items, generate a concrete edit suggestion ranked by expected_lift (high first, then medium, then low). 5. Re-run after edits until ≥ 90 % of applicable items pass. 6. Where an item carries a defers_to link, item-level detail belongs to the linked skill — record only the high-level status here to avoid duplicate audits.
Schema (v2)
items[].id : §-numbered rule id
items[].rule : one-line description
items[].applies_to : artifact types where the rule fires
items[].applies_to_phase: lifecycle phases where the rule is actionable
items[].priority : H | M | L (severity if missing)
items[].expected_lift : high | medium | low (KJR-case-validated visibility gain)
items[].defers_to : optional skill / file owning item-level detailChecklist
checklist:
schema_version: 2
last_updated: "2026-05-11"
metadata:
artifact_path: ""
artifact_type: "" # manuscript | preprint | readme | citation_cff | hf_card | dataset_card
artifact_phase: "" # pre-draft | drafting | pre-submission | post-acceptance | post-publication
journal_target: ""
audit_date: "" # YYYY-MM-DD
reviewer: ""
items:
# Section 1 — Title and Abstract Optimization
- id: 1.1
rule: Title three-slot ([Task] + [Modality/anatomy] + [Model class])
applies_to: [title]
applies_to_phase: [drafting, pre-submission]
priority: H
expected_lift: high
- id: 1.2
rule: Structured abstract per journal template; chunk-friendly ≤3-sentence sub-blocks
applies_to: [abstract]
applies_to_phase: [drafting, pre-submission]
priority: H
expected_lift: high
- id: 1.3
rule: First sentence states problem + contribution; last sentence is explicit interpretation
applies_to: [abstract]
applies_to_phase: [drafting, pre-submission]
priority: H
expected_lift: high
- id: 1.4
rule: Taxonomy line names controlled vocabulary (e.g., DTA, foundation-model evaluation)
applies_to: [abstract]
applies_to_phase: [pre-submission]
priority: M
expected_lift: low
- id: 1.5
rule: Quantified primary outcome with 95 % CI in abstract
applies_to: [abstract]
applies_to_phase: [drafting, pre-submission]
priority: H
expected_lift: high
- id: 1.6
rule: Reporting-guideline anchor present (TRIPOD+AI / CLAIM / STARD-AI / TRIPOD-LLM / DECIDE-AI / PRISMA-DTA)
applies_to: [abstract, methods]
applies_to_phase: [drafting, pre-submission]
priority: H
expected_lift: high
defers_to: /check-reporting # item-level audit; AIO records anchor-present status only
- id: 1.7
rule: MeSH and RadLex coverage; consistent UK/US spelling; no abstract-keyword redundancy
applies_to: [keywords, abstract]
applies_to_phase: [pre-submission]
priority: M
expected_lift: medium
# Section 2 — Manuscript-Level
- id: 2.1
rule: Journal-specific summary box present (Research in context / Key Points / PLS / Main Points)
applies_to: [manuscript]
applies_to_phase: [drafting, pre-submission]
priority: H
expected_lift: high # KJR case: top extracted fragment by AI overviews
- id: 2.2
rule: Section headings are declarative claims, not generic labels
applies_to: [manuscript]
applies_to_phase: [drafting, pre-submission]
priority: M
expected_lift: medium
- id: 2.3
rule: Numeric claim compression sentence in Methods and Results
applies_to: [methods, results]
applies_to_phase: [drafting, pre-submission]
priority: M
expected_lift: medium
- id: 2.4
rule: Reproducibility block (data, code, weights, prompts, seeds, env)
applies_to: [methods, data_availability]
applies_to_phase: [pre-submission]
priority: H
expected_lift: high
- id: 2.5
rule: Limitations enumerated and named
applies_to: [discussion]
applies_to_phase: [drafting, pre-submission]
priority: M
expected_lift: medium
- id: 2.6
rule: Standalone figure captions restate claim, dataset, and metric
applies_to: [figures]
applies_to_phase: [pre-submission]
priority: M
expected_lift: low
# Section 3 — Preprint, Channel, Indexing
- id: 3.1
rule: Preprint posted on submission day, OR fast-track justified
applies_to: [submission_strategy]
applies_to_phase: [pre-submission]
priority: H
expected_lift: high
- id: 3.2
rule: Journal preprint policy verified (Sherpa Romeo or IFA)
applies_to: [submission_strategy]
applies_to_phase: [pre-submission]
priority: H
expected_lift: medium
- id: 3.4
rule: Open-access route selected (gold OA preferred; green OA acceptable)
applies_to: [submission_strategy]
applies_to_phase: [pre-draft, pre-submission]
priority: H
expected_lift: high
defers_to: references/oac_funding_checklist.yaml
- id: 3.5
rule: Post-acceptance channel checklist scheduled (PMC, ORCID, Scholar, SNS, HF)
applies_to: [post_acceptance]
applies_to_phase: [post-acceptance]
priority: M
expected_lift: high
# Section 5 — GitHub / CITATION.cff / Zenodo / HF
- id: 5.1
rule: README follows 10-slot canonical order
applies_to: [readme]
applies_to_phase: [post-acceptance, post-publication]
priority: H
expected_lift: medium
- id: 5.2
rule: CITATION.cff present at repo root with ORCID, DOI, version
applies_to: [citation_cff]
applies_to_phase: [post-acceptance, post-publication]
priority: H
expected_lift: medium
- id: 5.3
rule: Zenodo DOI minted via GitHub integration; cited in paper
applies_to: [zenodo, manuscript]
applies_to_phase: [pre-submission, post-acceptance]
priority: H
expected_lift: high
- id: 5.4
rule: Hugging Face model card YAML keys + required prose sections complete
applies_to: [hf_card]
applies_to_phase: [post-acceptance, post-publication]
priority: H
expected_lift: medium
- id: 5.5
rule: Hugging Face dataset card includes PHI/re-identification disclosure
applies_to: [dataset_card]
applies_to_phase: [post-acceptance, post-publication]
priority: H
expected_lift: medium
- id: 5.6
rule: Web-crawler-friendly markdown (declarative headings, alt text, fenced code, JSON-LD)
applies_to: [readme, hf_card]
applies_to_phase: [post-acceptance, post-publication]
priority: M
expected_lift: medium
# Section 6 — Authority / E-E-A-T
- id: 6.1
rule: Personal author landing page lists all papers with DOIs
applies_to: [author_profile]
applies_to_phase: [pre-draft, drafting, pre-submission, post-acceptance, post-publication]
priority: M
expected_lift: medium
- id: 6.2
rule: Affiliation string consistent across papers
applies_to: [author_profile]
applies_to_phase: [drafting, pre-submission]
priority: H
expected_lift: high
- id: 6.3
rule: ORCID complete and linked to Google Scholar
applies_to: [author_profile]
applies_to_phase: [pre-draft, drafting, pre-submission, post-acceptance, post-publication]
priority: H
expected_lift: high
# Section 7 — LLM-Citation Fabrication Defense
- id: 7.1
rule: DOI + PMID surfaced in copy-friendly text on landing page and README
applies_to: [readme, manuscript]
applies_to_phase: [post-acceptance, post-publication]
priority: H
expected_lift: high
- id: 7.2
rule: "How to cite" section with BibTeX, APA, Vancouver, plain-text
applies_to: [readme]
applies_to_phase: [post-acceptance, post-publication]
priority: H
expected_lift: medium
- id: 7.3
rule: Scholar alert configured; Perplexity / ChatGPT probes scheduled
applies_to: [post_acceptance]
applies_to_phase: [post-publication]
priority: M
expected_lift: low
# Section 10 — Q&A and Entity-Extraction
- id: 10.1
rule: Four-question Q&A block (What was known / What this adds / How this changes practice / Why it matters)
applies_to: [discussion, supplement]
applies_to_phase: [drafting, pre-submission]
priority: H
expected_lift: high # KJR case: directly correlated with AI-overview extraction
- id: 10.2
rule: Glossary block with MeSH / RadLex / UMLS / arXiv IDs
applies_to: [methods, supplement]
applies_to_phase: [pre-submission]
priority: L
expected_lift: low # over-engineering for non-NLP / non-foundation-model papers; apply selectively
- id: 10.3
rule: Inline citation anchor text uses semantic predicates, not bare numbers
applies_to: [manuscript]
applies_to_phase: [drafting, pre-submission]
priority: M
expected_lift: medium
- id: 10.4
rule: Explicit "Why this is hard" challenge statement at start of Discussion
applies_to: [discussion]
applies_to_phase: [drafting, pre-submission]
priority: H
expected_lift: high # KJR case: top quoted fragment by Perplexity / ChatGPT web
# Section 11 — First-Mover Timing and Citation-Graph
- id: 11.1
rule: Topic-peak signals checked before submission timing decision
applies_to: [submission_strategy]
applies_to_phase: [pre-draft]
priority: H
expected_lift: high
- id: 11.2
rule: Editorial-board leverage considered (corresponding author affiliation)
applies_to: [submission_strategy]
applies_to_phase: [pre-draft]
priority: M
expected_lift: medium
- id: 11.3
rule: Target journal verified as PMC-auto-deposit when applicable
applies_to: [submission_strategy]
applies_to_phase: [pre-draft, pre-submission]
priority: H
expected_lift: high
defers_to: references/oac_funding_checklist.yaml
- id: 11.4
rule: Discussion anchored in 5–10 high-visibility seminal works
applies_to: [discussion]
applies_to_phase: [drafting]
priority: M
expected_lift: medium
- id: 11.5
rule: Author roster spans ≥ 3 institutions and ≥ 2 disciplines (review papers)
applies_to: [authorship]
applies_to_phase: [pre-draft]
priority: M
expected_lift: medium # not actionable once drafting begins; auto-NA in later phases
# Section 12 — Cross-Platform Launch Sequencing
- id: 12.1
rule: Day-0 simultaneous launch (GitHub release, HF card, X/Threads, LinkedIn, landing page)
applies_to: [post_acceptance]
applies_to_phase: [post-acceptance]
priority: H
expected_lift: high
- id: 12.2
rule: Day-1 propagation (ORCID, Scholar, preprint update, institutional page)
applies_to: [post_acceptance]
applies_to_phase: [post-acceptance]
priority: M
expected_lift: medium
- id: 12.3
rule: Week-1 depth posts (LinkedIn long-form, Papers with Code, ResearchGate)
applies_to: [post_acceptance]
applies_to_phase: [post-acceptance]
priority: M
expected_lift: medium
- id: 12.4
rule: Weeks 2–4 refresh signals scheduled
applies_to: [post_acceptance]
applies_to_phase: [post-publication]
priority: L
expected_lift: low
- id: 12.5
rule: Month-1 monitoring (alerts + AI-search probes)
applies_to: [post_acceptance]
applies_to_phase: [post-publication]
priority: M
expected_lift: mediumOutput template (paste into qc/aio_audit.md)
# AIO Audit — {artifact_path}
**Phase**: {pre-draft | drafting | pre-submission | post-acceptance | post-publication}
**Top 5 ranked by expected_lift (high → medium → low):**
| Rank | ID | Rule | Status | Reason | Suggested edit |
|------|----|------|--------|--------|----------------|
| 1 | {id} | … | FAIL | … | … |
| ... |
## Full table (applicable items only — out-of-phase rules auto-NA)
| ID | Rule | Lift | Status | Reason | Suggested edit |
|----|------|------|--------|--------|----------------|
| ... |
## Deferred items (audit elsewhere)
- §1.6 → run `/check-reporting`
- §3.4 / §11.3 → see `references/oac_funding_checklist.yaml`
## Summary
- PASS: X / Y applicable items
- High-lift fixes pending: …Migration from schema v1
If a prior qc/aio_audit.md was generated under schema v1 (no applies_to_phase / expected_lift), regenerate from this v2 file. The id list is unchanged so prior PASS/FAIL state can be re-imported by id.
---
schema_version: 1
last_updated: "2026-05-10"
purpose: |
Templates for journal-specific summary boxes used in medical-AI manuscripts.
Each entry mirrors the journal's Instructions for Authors (IFA) at the date
shown. USER MUST VERIFY the current IFA before applying — journal policies
change without notice.
verification_protocol: |
Before applying any template:
1. Open the journal's current IFA URL (ifa_url field).
2. Compare structure / labels / word targets to this YAML.
3. If any field has changed since `last_verified`, update this file before
deriving the manuscript's box.
4. Treat `last_verified` older than 6 months as STALE — re-verify before
submission.
journals:
- id: lancet-digital-health
name: The Lancet Digital Health
box_label: Research in context
ifa_url: https://www.thelancet.com/journals/landig/about
last_verified: "2026-05-10"
placement: After Introduction in main text, or as Panel 1.
structure:
- id: evidence-before
label: Evidence before this study
word_target: 100-200
notes: |
Describe the systematic search method (databases, dates, terms) and
what was known prior to this study.
- id: added-value
label: Added value of this study
word_target: 100-200
notes: |
State the contribution and how it advances the field beyond the
evidence summarized above.
- id: implications
label: Implications of all the available evidence
word_target: 100-200
notes: |
How findings (combined with prior evidence) should change clinical
practice, policy, or future research.
user_must_verify: true
- id: rsna-radiology
name: Radiology (RSNA)
box_label: Key Results
ifa_url: https://pubs.rsna.org/page/radiology/author-instructions
last_verified: "2026-05-10"
placement: First page of article, after Abstract.
structure:
- id: key-results
label: Key Results
format: 3 bullets, claim-centric, one finding each
notes: |
Each bullet should state a primary outcome with point estimate and
confidence interval where applicable. Avoid hedging language.
user_must_verify: true
- id: rsna-radiology-ai
name: Radiology Artificial Intelligence (RSNA)
box_label: Key Points
ifa_url: https://pubs.rsna.org/page/ai/author-instructions
last_verified: "2026-05-10"
placement: After Abstract.
structure:
- id: key-points
label: Key Points
format: 3 bullets
notes: |
Same conventions as Radiology Key Results — claim-centric, numeric
where possible.
user_must_verify: true
- id: npj-digital-medicine
name: npj Digital Medicine
box_label: Plain Language Summary
ifa_url: https://www.nature.com/npjdigitalmed/submission-guidelines
last_verified: "2026-05-10"
placement: After Abstract; required for all primary research.
structure:
- id: pls
label: Plain Language Summary
word_target: 150-200
reading_level: 8th-grade
notes: |
Avoid jargon, define acronyms inline, prefer everyday vocabulary.
Should be readable by a non-specialist physician or informed lay
reader. No quantitative thresholds; describe magnitude in plain terms.
user_must_verify: true
- id: nature-medicine
name: Nature Medicine
box_label: Editor's Summary
ifa_url: https://www.nature.com/nm/for-authors
last_verified: "2026-05-10"
placement: Editorial-supplied at acceptance; not author-drafted in main text.
structure:
- id: editor-summary
label: Editor's Summary
word_target: 100-150
notes: |
Supplied by the editorial team after acceptance. Authors should
proactively draft a 100–150 word version for the cover letter and
post-acceptance launch materials, so the editor has anchor wording
to refine.
user_must_verify: true
- id: jamia
name: JAMIA (Journal of the American Medical Informatics Association)
box_label: (none mandated; structured abstract used)
ifa_url: https://academic.oup.com/jamia/pages/General_Instructions
last_verified: "2026-05-10"
placement: Structured abstract serves the role of summary box.
structure:
- id: structured-abstract
label: Structured Abstract
format: Objective / Materials and Methods / Results / Discussion / Conclusion
notes: Treat each section as a chunk-friendly semantic block (≤ 3 sentences).
user_must_verify: true
related:
- skills/academic-aio/SKILL.md Section 2.1 (manuscript-level summary box)
- skills/academic-aio/SKILL.md Section 1.6 (reporting-guideline anchor)
---
schema_version: 1
last_updated: "2026-05-10"
purpose: |
Open-access compliance matrix for major funders that medical-AI researchers
may encounter. Use to decide gold OA vs green OA vs preprint route per paper.
Funder policies change; verify the policy_url before applying.
verification_protocol: |
All funder policies change without notice.
- Verify policy_url before applying decisions.
- Treat `last_verified` older than 6 months as STALE.
- When in doubt, contact the institution's research-administration office.
funders:
- id: plan-s
name: cOAlition S (Plan S)
region: Europe (multiple national funders)
policy_url: https://www.coalition-s.org/plan-s-principles/
last_verified: "2026-05-10"
requirement: Immediate OA on publication; no embargo allowed.
accepted_routes:
- Gold OA with CC-BY (preferred)
- Green OA via repository deposit at publication (no embargo)
apc_cap: ~ EUR 2,500 typical, varies by funder
notes: Strict; transformative agreements widely available.
user_must_verify: true
- id: nih-public-access
name: US NIH Public Access Policy
region: United States
policy_url: https://publicaccess.nih.gov/
last_verified: "2026-05-10"
requirement: |
From 2025 onward NIH policy requires manuscript deposit in PMC at
publication (no 12-month embargo). Earlier grants may still operate
under the legacy 12-month rule.
accepted_routes:
- Gold OA (publisher deposits to PMC)
- Green OA (author deposits AAM to PMC via NIHMS)
apc_cap: No cap; APCs are eligible expenses on most grants.
user_must_verify: true
- id: ukri
name: UK Research and Innovation (UKRI)
region: United Kingdom
policy_url: https://www.ukri.org/manage-your-award/publishing-your-research-findings/making-your-research-open/
last_verified: "2026-05-10"
requirement: Immediate OA; CC-BY preferred.
accepted_routes:
- Gold OA via transformative agreement (preferred)
- Green OA in approved repository at publication
user_must_verify: true
- id: gates-foundation
name: Bill & Melinda Gates Foundation
region: Global
policy_url: https://openaccess.gatesfoundation.org/
last_verified: "2026-05-10"
requirement: Immediate OA, CC-BY, no embargo.
accepted_routes:
- Gold OA (mandatory)
notes: Strictest among major funders. Closed-access publication is non-compliant.
user_must_verify: true
- id: wellcome
name: Wellcome Trust
region: UK / global
policy_url: https://wellcome.org/grant-funding/guidance/open-access-policy
last_verified: "2026-05-10"
requirement: Immediate OA, CC-BY, including for preprints.
accepted_routes:
- Gold OA (preferred)
- Green OA at publication
user_must_verify: true
- id: nrf-korea
name: National Research Foundation of Korea (NRF)
region: South Korea
policy_url: https://www.nrf.re.kr/
last_verified: "2026-05-10"
requirement: |
No formal mandatory immediate-OA policy as of 2026-Q2. Gold OA is
encouraged and APCs are eligible expenses on most grants.
accepted_routes:
- Gold OA (encouraged)
- Green OA (acceptable)
user_must_verify: true
- id: hira-mohw
name: Korean Ministry of Health & Welfare research grants
region: South Korea
policy_url: https://www.mohw.go.kr/
last_verified: "2026-05-10"
requirement: Public benefit / repository deposit encouraged; no strict immediate-OA mandate.
accepted_routes:
- Gold OA (encouraged)
- Green OA (acceptable)
user_must_verify: true
decision_tree:
- step: 1
question: Does any funder on this paper require immediate OA?
yes_action: |
Choose gold OA at a Plan-S compliant journal, OR green OA with
repository deposit at publication. Closed-access is not allowed.
no_action: Continue to step 2.
- step: 2
question: Is APC budget available (grant line, institutional waiver, or transformative agreement)?
yes_action: |
Gold OA preferred — lower friction, faster PMC deposit, cited
30-50 % more by AI-search tools (Section 3.4).
no_action: |
Green OA via preprint at submission + AAM deposit at acceptance.
Verify journal's green-OA timeline.
- step: 3
question: Does the target journal auto-deposit to PMC (Section 11.3)?
yes_action: Confirm and proceed. Deposit window typically 4-6 weeks.
no_action: |
Plan green OA pathway and timeline. Budget the AAM submission step
(NIHMS or Europe PMC).
- step: 4
question: Is a transformative agreement available via the institution?
yes_action: Gold OA is effectively free for the author — use it.
no_action: Apply APC budget directly or fall back to green OA.
related:
- skills/academic-aio/SKILL.md Section 3.4 (open-access choice)
- skills/academic-aio/SKILL.md Section 11.3 (PMC-auto-deposit journal preference)
Reporting-Guideline ↔ AIO Rule Mapping
This table maps each AIO rule (sections 1-12 of SKILL.md) to the corresponding item(s) in the major medical-AI reporting guidelines. Use it to align the /check-reporting audit with the /academic-aio audit so the same evidence covers both.
Core mapping
| AIO rule | TRIPOD+AI 2024 | CLAIM 2024 | STARD-AI 2025 | TRIPOD-LLM 2024 | DECIDE-AI 2022 | Notes |
|---|---|---|---|---|---|---|
| §1.1 Title three-slot | item 1 (title) | item 1 (title and abstract) | item 1 (title) | item 1a (title) | — | All require keyword presence and study-type identification. |
| §1.2 Structured abstract | item 2 (abstract) | item 1 (title and abstract) | item 2 (abstract) | item 1b (abstract) | — | Each guideline mandates structured form. |
| §1.5 Quantified primary outcome with CI | item 16 (model performance) | items 28-30 (performance metrics) | items 23-26 (diagnostic estimates) | item 17 (performance with CI) | item 8 (clinical effect) | CI mandatory for all. |
| §1.6 Reporting-guideline anchor | (compliance declaration) | (compliance declaration) | (compliance declaration) | (compliance declaration) | (compliance declaration) | Cite guideline + checklist in Methods or supplement. |
| §2.4 Reproducibility block | item 24 (data sharing) + item 25 (code sharing) | items 33-34 (data and code availability) | item 28 (data and code) | items 22-23 (artifacts) | item 12 (artifacts) | All require explicit data/code statement. |
| §2.5 Limitations enumeration | item 23 (limitations) | item 41 (limitations) | item 27 (limitations) | item 19 (limitations) | item 11 (limitations) | Enumerate, do not narrate generally. |
| §10.4 Challenge statement | (implicit in Background) | item 4 (rationale) | item 5 (rationale) | item 4 (rationale) | item 4 (rationale) | "Why this is hard" overlaps with rationale items. |
Workflow
1. Run /check-reporting first — produces a PRESENT/PARTIAL/MISSING audit per guideline item. 2. Run /academic-aio — produces the AIO PASS/PARTIAL/FAIL checklist. 3. For each AIO FAIL or PARTIAL row, check this mapping. If the underlying reporting-guideline item is also MISSING/PARTIAL, fix it once and both audits update. 4. Items present in /check-reporting audit but not in this mapping (e.g., randomization details for RCTs, domain-specific safety items) do not have an AIO consequence and can be addressed independently.
When the two audits disagree
- AIO PASS + reporting-guideline MISSING — the manuscript looks discoverable but is not formally compliant. Reviewers may still reject. Always fix the reporting-guideline gap.
- AIO FAIL + reporting-guideline PRESENT — a rule was recorded in compliance form but rendered in a way that LLM extractors cannot parse (e.g., reporting CIs in supplementary instead of inline in the abstract). Move the content into a chunk-friendly location.
Source
- TRIPOD+AI: Collins et al. BMJ 2024.
- CLAIM 2024: Tejani et al. Radiology: AI 2024, doi:10.1148/ryai.240300.
- STARD-AI 2025: Sounderajah et al. Nat Med 2025, doi:10.1038/s41591-025-03953-8.
- TRIPOD-LLM 2024: Gallifant et al. Nat Med 2024, doi:10.1038/s41591-024-03425-5.
- DECIDE-AI 2022: Vasey et al. Nat Med 2022, doi:10.1038/s41591-022-01772-9.
Anti-hallucination
Item numbers above are derived from the most recent published version of each guideline. Verify against the EQUATOR Network entry before citing item numbers in a manuscript — guideline updates renumber items.
{
"@context": "https://schema.org",
"@type": "SoftwareSourceCode",
"name": "<repo name>",
"description": "<one-line description of what this repository contains>",
"codeRepository": "https://github.com/<org>/<repo>",
"programmingLanguage": ["Python", "R"],
"license": "https://spdx.org/licenses/Apache-2.0.html",
"datePublished": "YYYY-MM-DD",
"version": "v1.0.0",
"author": [
{
"@type": "Person",
"name": "<First Last>",
"identifier": "https://orcid.org/0000-0000-0000-0000"
}
],
"isPartOf": {
"@type": "ScholarlyArticle",
"@id": "https://doi.org/10.xxxx/yyyy"
},
"identifier": {
"@type": "PropertyValue",
"propertyID": "DOI",
"value": "10.5281/zenodo.<id>"
},
"url": "https://doi.org/10.5281/zenodo.<id>",
"softwareRequirements": "Python 3.10+, PyTorch 2.0+",
"operatingSystem": ["Linux", "macOS"],
"applicationCategory": "Medical AI Research",
"keywords": ["<task>", "<modality>", "<model family>"]
}
{
"@context": "https://schema.org",
"@type": "Dataset",
"name": "<dataset name>",
"description": "<dataset description: modality, n cases, anatomy, task>",
"url": "https://doi.org/10.5281/zenodo.<id>",
"identifier": "10.5281/zenodo.<id>",
"license": "https://creativecommons.org/licenses/by/4.0/",
"creator": [
{
"@type": "Person",
"name": "<First Last>",
"identifier": "https://orcid.org/0000-0000-0000-0000"
}
],
"datePublished": "YYYY-MM-DD",
"isPartOf": {
"@type": "ScholarlyArticle",
"@id": "https://doi.org/10.xxxx/yyyy"
},
"variableMeasured": [
{"@type": "PropertyValue", "name": "<measurement 1, e.g., AUC>"},
{"@type": "PropertyValue", "name": "<measurement 2, e.g., sensitivity>"}
],
"distribution": {
"@type": "DataDownload",
"encodingFormat": "application/zip",
"contentUrl": "<direct download URL>",
"contentSize": "<bytes>"
},
"keywords": ["<task>", "<modality>", "<anatomy>", "<dataset family>"],
"spatialCoverage": "<country/region of provenance>",
"temporalCoverage": "YYYY/YYYY",
"citation": "<APA or Vancouver citation of the parent paper>",
"isAccessibleForFree": true
}
{
"@context": "https://schema.org",
"@type": "Person",
"name": "<First Last>",
"givenName": "<First>",
"familyName": "<Last>",
"identifier": "https://orcid.org/0000-0000-0000-0000",
"url": "<personal landing page URL>",
"jobTitle": "<Resident / Researcher / Faculty>",
"affiliation": [
{
"@type": "Organization",
"name": "<Department, Institution>",
"address": "<City, Country>"
}
],
"alumniOf": {
"@type": "Organization",
"name": "<medical school or PhD institution>"
},
"knowsAbout": ["Radiology", "Medical AI", "Deep learning"],
"sameAs": [
"https://orcid.org/0000-0000-0000-0000",
"https://scholar.google.com/citations?user=<id>",
"https://www.semanticscholar.org/author/<id>",
"https://github.com/<username>",
"https://huggingface.co/<username>",
"https://www.linkedin.com/in/<username>"
]
}
Schema.org JSON-LD Templates
Embed-ready JSON-LD markup for academic-aio Section 5 (GitHub / CITATION.cff / Zenodo / Hugging Face) and Section 6 (author landing page). Schema.org markup is read by Google Scholar's structured-data parser, by AI overview engines, and by RAG systems that crawl repository pages.
Files
- `ScholarlyArticle.jsonld` — embed in the paper landing page or repository README. Pairs each paper with DOI, PMID, abstract, license, OA flag, and citation graph anchors.
- `CodeRepository.jsonld` — embed in the repository README or as
.well-known/scholarly.jsonld. Links code to its parent article viaisPartOf. - `Dataset.jsonld` — embed at the Zenodo or Hugging Face dataset landing page. Includes provenance, license, and variable-level metadata.
- `Person.jsonld` — embed at the author landing page or institutional profile. Cross-links ORCID, Scholar, Semantic Scholar, GitHub, Hugging Face, LinkedIn.
How to embed
In a Markdown README (GitHub renders the HTML script tag inside HTML blocks)
<script type="application/ld+json">
{ ... contents of ScholarlyArticle.jsonld ... }
</script>In a personal landing page
Place the script tag in the <head> of the HTML page. For static-site generators (Hugo, Jekyll, Next.js), generate the JSON-LD at build time from frontmatter.
Validation
Run scripts/validate_schema.py path/to/file.jsonld to verify syntactic validity and required-field presence before deploy.
python scripts/validate_schema.py references/schema_markup_templates/*.jsonldAnti-hallucination
- Do not auto-fill placeholder values (
<First Last>,0000-0000-0000-0000,10.xxxx/yyyy). Mark unknown fields asnullor remove the field — partial fabricated identifiers are worse than missing ones. - Verify DOI format
10.{prefix}/{suffix}and ORCID format before commit.
Related
SKILL.mdSection 5 — README, CITATION.cff, Zenodo, Hugging FaceSKILL.mdSection 6 — Authority and E-E-A-T signalsSKILL.mdSection 7 — Citation-fabrication defense
{
"@context": "https://schema.org",
"@type": "ScholarlyArticle",
"headline": "<paper title verbatim>",
"alternativeHeadline": "<short version or subtitle, optional>",
"datePublished": "YYYY-MM-DD",
"author": [
{
"@type": "Person",
"name": "<First Last>",
"givenName": "<First>",
"familyName": "<Last>",
"identifier": "https://orcid.org/0000-0000-0000-0000",
"affiliation": {
"@type": "Organization",
"name": "<Department, Institution>",
"address": "<City, Country>"
}
}
],
"publisher": {
"@type": "Organization",
"name": "<Journal publisher>"
},
"isPartOf": {
"@type": "PublicationVolume",
"name": "<Journal name>",
"volumeNumber": "<Volume>",
"issueNumber": "<Issue>",
"pageStart": "<start>",
"pageEnd": "<end>"
},
"identifier": [
{"@type": "PropertyValue", "propertyID": "DOI", "value": "10.xxxx/yyyy"},
{"@type": "PropertyValue", "propertyID": "PMID", "value": "<PMID>"},
{"@type": "PropertyValue", "propertyID": "PMC", "value": "PMC<id>"}
],
"url": "https://doi.org/10.xxxx/yyyy",
"license": "https://creativecommons.org/licenses/by/4.0/",
"abstract": "<200-300 word abstract verbatim>",
"keywords": ["<MeSH term 1>", "<MeSH term 2>", "<RadLex term>", "<topic term>"],
"isAccessibleForFree": true,
"citation": [
{
"@type": "ScholarlyArticle",
"@id": "https://doi.org/10.xxxx/zzzz",
"name": "<seminal reference title>"
}
],
"subjectOf": {
"@type": "Dataset",
"@id": "https://doi.org/10.5281/zenodo.<id>",
"name": "<companion dataset name>"
}
}
{
"_comment": "Synthesis of PUBLIC, journal-documented structured-summary-box facts (bullet counts, sub-block labels, word bands) — NOT verbatim journal text. Mirrors the facts academic-aio/SKILL.md already states (Lancet 'Research in context' 3 sub-blocks; Radiology/RYAI 'Key Points' 3 one-claim bullets; npj 'Plain-language summary' 150-200 words). Verify current instructions-to-authors at the journal site before relying on a format; these are conventions, not guarantees.",
"schema_version": 1,
"formats": {
"key_points": {
"label": "Key Points",
"bullets": 3,
"one_claim_per_bullet": true,
"journals": ["radiology", "radiology-ai", "ryai", "rsna"]
},
"research_in_context": {
"label": "Research in context",
"subblocks": [
"Evidence before this study",
"Added value of this study",
"Implications of all the available evidence"
],
"journals": ["lancet-digital-health", "lancet", "lancet-oncology"]
},
"plain_language_summary": {
"label": "Plain-language summary",
"word_min": 150,
"word_max": 200,
"journals": ["npj-digital-medicine", "npj"]
}
}
}
#!/usr/bin/env python3
"""
batch_metadata_audit.py — Audit multiple medical-AI repos and Hugging Face cards
for AIO compliance.
Per repository:
- README.md present, with DOI link / badge and a How-to-cite or Citation section
- CITATION.cff present at root with title/authors/version + at least one ORCID
- LICENSE present
Per Hugging Face card (model or dataset):
- YAML front matter present with license / library_name / tags
- Required prose sections: intended use, training data, evaluation, limitations,
ethical considerations
- No PHI patterns matched
Usage:
python batch_metadata_audit.py /path/to/repo1 /path/to/repo2 \\
--hf-card model_card.md --output qc/aio_batch.json
python batch_metadata_audit.py --hf-card dataset_card.md --fail-on-issue
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from pathlib import Path
PHI_PATTERNS = [
re.compile(r"\b\d{3}-\d{2}-\d{4}\b"), # US SSN
re.compile(r"\b\d{6}-\d{7}\b"), # KR resident registration number
re.compile(r"MRN[:\s]*\d+", re.I), # medical record number
re.compile(r"\bpatient[\s_]?id[:\s]*\d+", re.I), # patient id
re.compile(r"\b\d{4}-\d{2}-\d{2}\b.*\bDOB\b", re.I), # DOB pattern
]
CITATION_CFF_REQUIRED = ("title", "authors", "version")
HF_CARD_REQUIRED_YAML = ("license", "library_name", "tags")
HF_CARD_REQUIRED_SECTIONS = (
"intended use",
"training data",
"evaluation",
"limitations",
"ethical considerations",
)
def _phi_hits(text: str) -> list[str]:
hits: list[str] = []
for pat in PHI_PATTERNS:
m = pat.search(text)
if m:
hits.append(f"Possible PHI pattern matched: {m.group(0)[:30]!r}")
return hits
def check_readme(path: Path) -> dict:
if not path.exists():
return {"present": False, "issues": ["README.md missing"]}
text = path.read_text()
issues: list[str] = []
if "DOI" not in text and "doi.org" not in text.lower():
issues.append("No DOI badge or link in README")
if "## How to cite" not in text and "## Citation" not in text:
issues.append('No "How to cite" / "Citation" section')
if not any(s in text.lower() for s in ("quickstart", "getting started", "installation")):
issues.append("No quickstart / installation section")
return {"present": True, "issues": issues}
def check_citation_cff(path: Path) -> dict:
if not path.exists():
return {"present": False, "issues": ["CITATION.cff missing at repo root"]}
text = path.read_text()
issues: list[str] = []
for key in CITATION_CFF_REQUIRED:
if f"{key}:" not in text:
issues.append(f"CITATION.cff missing key: {key}")
if "orcid" not in text.lower():
issues.append("CITATION.cff has no ORCID identifier(s)")
return {"present": True, "issues": issues}
def check_license(path: Path) -> dict:
return {
"present": path.exists(),
"issues": [] if path.exists() else ["LICENSE missing"],
}
def check_hf_card(path: Path) -> dict:
if not path.exists():
return {"present": False, "issues": ["HF card missing"]}
text = path.read_text()
issues: list[str] = []
if not text.startswith("---"):
issues.append("HF card missing YAML front matter")
else:
front_parts = text.split("---", 2)
front = front_parts[1] if len(front_parts) >= 3 else ""
for key in HF_CARD_REQUIRED_YAML:
if f"{key}:" not in front:
issues.append(f"HF card YAML missing key: {key}")
body_lower = text.lower()
for section in HF_CARD_REQUIRED_SECTIONS:
if section not in body_lower:
issues.append(f"HF card missing section: {section}")
issues.extend(_phi_hits(text))
return {"present": True, "issues": issues}
def audit_repo(repo: Path) -> dict:
return {
"repo": str(repo),
"readme": check_readme(repo / "README.md"),
"citation_cff": check_citation_cff(repo / "CITATION.cff"),
"license": check_license(repo / "LICENSE"),
}
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(
description="Audit multiple repos and HF cards for AIO compliance."
)
parser.add_argument("paths", nargs="*", type=Path,
help="Repository directories to audit.")
parser.add_argument("--hf-card", action="append", type=Path, default=[],
help="Path to a Hugging Face model/dataset card markdown file.")
parser.add_argument("--output", type=Path, default=None,
help="Write JSON report to this path (default: stdout).")
parser.add_argument("--fail-on-issue", action="store_true",
help="Exit 1 if any issues are detected.")
args = parser.parse_args(argv)
if not args.paths and not args.hf_card:
parser.error("Provide at least one repo path or --hf-card.")
report: dict = {"repos": [], "hf_cards": []}
for path in args.paths:
if path.is_dir():
report["repos"].append(audit_repo(path))
else:
report["repos"].append({"repo": str(path),
"issues": ["Not a directory"]})
for card in args.hf_card:
report["hf_cards"].append({"path": str(card), **check_hf_card(card)})
out = json.dumps(report, indent=2, ensure_ascii=False)
if args.output:
args.output.parent.mkdir(parents=True, exist_ok=True)
args.output.write_text(out)
else:
print(out)
has_issues = any(
(r.get("readme", {}).get("issues") or
r.get("citation_cff", {}).get("issues") or
r.get("license", {}).get("issues") or
r.get("issues"))
for r in report["repos"]
) or any(c.get("issues") for c in report["hf_cards"])
return 1 if (args.fail_on_issue and has_issues) else 0
if __name__ == "__main__":
sys.exit(main())
#!/usr/bin/env python3
"""Structured-summary-box conformance detector (academic-aio).
High-impact medical-AI journals require a structured summary box whose *format*
is journal-specific, and a production/technical check rejects the wrong one:
- Radiology / Radiology:AI (RSNA): "Key Points" — exactly 3 bullets, one claim each.
- Lancet family: "Research in context" — three labelled sub-blocks
(Evidence before this study / Added value of this study / Implications of all
the available evidence).
- npj Digital Medicine: "Plain-language summary" — ~150-200 words.
academic-aio already *generates* these boxes; this detector makes the spec
deterministic so a wrong-bullet-count, missing-sub-block, or over/under-length
box is caught before submission instead of at the technical check. The spec is
read from references/summary_box_specs.json (public facts, journal-keyed).
INPUTS
--manuscript markdown file containing the summary box (required).
--journal journal stem to pick the format (e.g. radiology, lancet-digital-health,
npj-digital-medicine). Optional if --format is given.
--format force a format: key_points | research_in_context | plain_language_summary.
--specs path to summary_box_specs.json (default: alongside this script's skill).
--out write a JSON report here (default: qc/summary_box_report.json).
--strict exit 1 if the box is non-conformant.
VERDICT
CONFORMANT the box matches its format's spec.
NONCONFORMANT a hard rule failed (wrong bullet count, missing sub-block,
word count outside the band, box absent).
ADVISORY only soft rules fired (e.g. a bullet carries >1 claim).
Exit: 0 conformant/advisory or report-only; 1 NONCONFORMANT under --strict;
2 input/usage error.
Stdlib-only (csv-free: json / argparse / re / pathlib).
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from pathlib import Path
def _default_specs() -> Path:
return Path(__file__).resolve().parent.parent / "references" / "summary_box_specs.json"
def _err(msg: str) -> int:
print(f"ERROR: {msg}", file=sys.stderr)
return 2
def load_specs(path: Path) -> dict:
data = json.loads(path.read_text(encoding="utf-8"))
return data["formats"]
def pick_format(formats: dict, journal: str | None, fmt: str | None) -> str | None:
if fmt:
return fmt if fmt in formats else None
if journal:
j = journal.strip().lower()
for key, spec in formats.items():
if j in [x.lower() for x in spec.get("journals", [])]:
return key
return None
def extract_block(text: str, label: str) -> str | None:
"""Return the lines under a heading/bold label matching `label`, up to the
next markdown heading or a blank-line-separated next bold label section."""
lines = text.splitlines()
label_re = re.compile(
r"^\s*(?:#{1,6}\s*|\*\*\s*|\*\s*)?" + re.escape(label) + r"\b", re.IGNORECASE
)
start = None
for i, ln in enumerate(lines):
if label_re.search(ln):
start = i
break
if start is None:
return None
out: list[str] = []
for ln in lines[start + 1:]:
if re.match(r"^\s*#{1,6}\s+\S", ln): # next heading ends the block
break
out.append(ln)
return "\n".join(out).strip()
def count_bullets(block: str) -> list[str]:
bullets: list[str] = []
for ln in block.splitlines():
m = re.match(r"^\s*(?:[-*+]|\d+[.)])\s+(.*\S)", ln)
if m:
bullets.append(m.group(1).strip())
return bullets
def multi_claim(bullet: str) -> bool:
"""Heuristic: a one-claim bullet should not pack two independent assertions.
Flags a sentence-final period followed by a capitalized new sentence, or a
semicolon joining two clauses."""
if ";" in bullet:
return True
return bool(re.search(r"[.!?]\s+[A-Z0-9]", bullet.rstrip(".")))
def word_count(block: str) -> int:
# strip the label line if it leaked in; count remaining words.
return len(re.findall(r"\b[\w'-]+\b", block))
def check(text: str, fmt: str, spec: dict) -> dict:
label = spec["label"]
block = extract_block(text, label)
findings: list[dict] = []
if block is None:
return {
"format": fmt, "label": label, "verdict": "NONCONFORMANT",
"findings": [{"rule": "box_present", "severity": "hard",
"detail": f"no '{label}' box found in the manuscript"}],
}
if fmt == "key_points":
bullets = count_bullets(block)
want = spec["bullets"]
if len(bullets) != want:
findings.append({"rule": "bullet_count", "severity": "hard",
"detail": f"found {len(bullets)} bullets, expected {want}"})
if spec.get("one_claim_per_bullet"):
for b in bullets:
if multi_claim(b):
findings.append({"rule": "one_claim_per_bullet", "severity": "soft",
"detail": f"bullet packs >1 claim: {b[:80]}"})
elif fmt == "research_in_context":
low = block.lower()
for sub in spec["subblocks"]:
if sub.lower() not in low:
findings.append({"rule": "subblock_present", "severity": "hard",
"detail": f"missing sub-block: '{sub}'"})
elif fmt == "plain_language_summary":
wc = word_count(block)
lo, hi = spec["word_min"], spec["word_max"]
if wc < lo or wc > hi:
findings.append({"rule": "word_band", "severity": "hard",
"detail": f"{wc} words, expected {lo}-{hi}"})
else:
return {"format": fmt, "label": label, "verdict": "NONCONFORMANT",
"findings": [{"rule": "unknown_format", "severity": "hard",
"detail": f"unknown format '{fmt}'"}]}
hard = any(f["severity"] == "hard" for f in findings)
verdict = "NONCONFORMANT" if hard else ("ADVISORY" if findings else "CONFORMANT")
return {"format": fmt, "label": label, "verdict": verdict, "findings": findings}
def main() -> int:
ap = argparse.ArgumentParser(description="Check a structured summary box against its journal format spec.")
ap.add_argument("--manuscript", required=True)
ap.add_argument("--journal")
ap.add_argument("--format")
ap.add_argument("--specs")
ap.add_argument("--out")
ap.add_argument("--strict", action="store_true")
args = ap.parse_args()
man = Path(args.manuscript)
if not man.is_file():
return _err(f"manuscript not found: {man}")
specs_path = Path(args.specs) if args.specs else _default_specs()
if not specs_path.is_file():
return _err(f"specs not found: {specs_path}")
formats = load_specs(specs_path)
fmt = pick_format(formats, args.journal, args.format)
if fmt is None:
return _err("could not select a format — pass --format or a --journal listed in the specs")
text = man.read_text(encoding="utf-8")
report = check(text, fmt, formats[fmt])
out_path = Path(args.out) if args.out else Path("qc") / "summary_box_report.json"
out_path.parent.mkdir(parents=True, exist_ok=True)
out_path.write_text(json.dumps(report, indent=2) + "\n", encoding="utf-8")
print("=" * 41)
print(" Summary-Box Conformance")
print("=" * 41)
print(f"format: {fmt} ({report['label']})")
print(f"verdict: {report['verdict']}")
for f in report["findings"]:
print(f" [{f['severity']}] {f['rule']}: {f['detail']}")
print(f"report: {out_path}")
if report["verdict"] == "NONCONFORMANT" and args.strict:
print("\nSUMMARY_BOX_NONCONFORMANT", file=sys.stderr)
return 1
return 0
if __name__ == "__main__":
sys.exit(main())
#!/usr/bin/env python3
"""
validate_schema.py — JSON-LD validator for academic-aio schema markup files.
Validates:
- JSON-LD syntactic validity
- @context = "https://schema.org"
- @type matches one of the supported types
- Required fields present (per schema.org minimal recommendations + medsci-skills policy)
- Identifier format (DOI, ORCID)
Usage:
python validate_schema.py path/to/file.jsonld [path/to/another.jsonld ...]
python validate_schema.py --strict references/schema_markup_templates/*.jsonld
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from pathlib import Path
REQUIRED_BY_TYPE: dict[str, list[str]] = {
"ScholarlyArticle": ["headline", "datePublished", "author", "identifier", "url"],
"SoftwareSourceCode": ["name", "codeRepository", "license", "datePublished", "author"],
"Dataset": ["name", "description", "license", "creator", "datePublished"],
"Person": ["name", "identifier"],
}
DOI_RE = re.compile(r"^10\.\d{4,9}/[-._;()/:A-Za-z0-9]+$")
ORCID_RE = re.compile(r"^https://orcid\.org/\d{4}-\d{4}-\d{4}-\d{3}[\dX]$")
PLACEHOLDER_TOKENS = ("<", "xxxx", "yyyy", "0000-0000-0000-0000")
def _is_placeholder(value: str) -> bool:
"""Return True for template placeholder strings that should skip strict checks."""
if not isinstance(value, str):
return False
lowered = value.lower()
return any(tok in lowered for tok in PLACEHOLDER_TOKENS)
def validate(path: Path) -> list[str]:
errors: list[str] = []
try:
data = json.loads(path.read_text())
except FileNotFoundError:
return [f"File not found: {path}"]
except json.JSONDecodeError as exc:
return [f"Invalid JSON: {exc}"]
ctx = data.get("@context")
if ctx != "https://schema.org":
errors.append(f'@context must be "https://schema.org" (got {ctx!r})')
typ = data.get("@type")
if typ not in REQUIRED_BY_TYPE:
errors.append(
f"@type {typ!r} not recognized "
f"(expected one of {sorted(REQUIRED_BY_TYPE)})"
)
return errors
for field in REQUIRED_BY_TYPE[typ]:
if field not in data or data[field] in (None, "", []):
errors.append(f"Missing required field: {field}")
ident = data.get("identifier")
if isinstance(ident, list):
for entry in ident:
if isinstance(entry, dict) and entry.get("propertyID") == "DOI":
value = entry.get("value", "")
if value and not _is_placeholder(value) and not DOI_RE.match(value):
errors.append(f"DOI does not match canonical format: {value!r}")
if typ == "Person" and isinstance(ident, str):
if not _is_placeholder(ident) and not ORCID_RE.match(ident):
errors.append(f"Person identifier should be an ORCID URL (got {ident!r})")
authors = data.get("author") or data.get("creator") or []
if isinstance(authors, list):
for i, a in enumerate(authors):
if isinstance(a, dict) and not a.get("name"):
errors.append(f"author[{i}] missing 'name'")
return errors
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(
description="Validate academic-aio Schema.org JSON-LD markup files."
)
parser.add_argument("files", nargs="+", type=Path)
parser.add_argument(
"--strict",
action="store_true",
help="Exit 1 on any error (default behaviour; flag retained for clarity).",
)
args = parser.parse_args(argv)
overall_ok = True
for path in args.files:
errors = validate(path)
if errors:
overall_ok = False
print(f"FAIL {path}")
for e in errors:
print(f" - {e}")
else:
print(f"PASS {path}")
return 0 if overall_ok else 1
if __name__ == "__main__":
sys.exit(main())
schema_version: 2
name: academic-aio
layer: C
owner_domain: manuscript_optimization
maturity: official
when_to_use: "Optimize a medical-AI manuscript's title, abstract, and structured summary boxes for AI search engines and RAG tools, integrating TRIPOD+AI/CLAIM/STARD-AI reporting requirements."
when_NOT_to_use: "Drafting a manuscript from scratch (use write-paper); removing AI writing tells (use humanize)."
inputs:
- "draft manuscript / abstract / title (Markdown)"
outputs:
- "AIO-optimized title, abstract, structured summary boxes"
- "visible pass/fail checklist"
deterministic_scripts:
- scripts/validate_schema.py
- scripts/batch_metadata_audit.py
side_effects:
- writes_project_artifacts
downstream_consumers:
- self-review
- check-reporting
forbidden_actions:
- silently_rewrite_without_user_review
- fabricate_reporting_guideline_compliance
# v2.1 quality card
purpose: "Make a medical-AI manuscript discoverable and citable by AI search/RAG tools without sacrificing reporting-guideline compliance."
safety_boundaries:
- "Off by default in autonomous pipelines; rewrites require user review (no silent rewrite)."
- "Reporting-guideline claims are checked, never asserted without the underlying item being present."
known_limitations:
- "GEO/AIO heuristics evolve with each engine; recommendations are point-in-time."
- "Does not draft new scientific content; optimizes existing approved text only."
validation_commands:
- "python3 scripts/validate_schema.py <summary-box-file>"
- "python3 scripts/check_summary_box.py --manuscript <file> --journal <stem> --strict"
- "bash tests/test_summary_box.sh"
evidence_surface: bundled_script
{# Jinja2 template — render with `aio_audit_checklist.md.j2` + variables.
Variables expected:
artifact_path : str absolute path of the artifact under audit
artifact_type : str manuscript | preprint | readme | citation_cff | hf_card | dataset_card
artifact_phase : str pre-draft | drafting | pre-submission | post-acceptance | post-publication
journal_target : str | None optional journal id from journal_summarybox_templates.yaml
audit_date : str YYYY-MM-DD
reviewer : str short name of the auditor
items : list[dict] loaded from references/checklists/AIO_GENERAL.md "items" list (schema v2)
reporting_guideline_pending : bool
numerical_audit_pending : bool
humanize_pending : bool
Render hint (Python):
from jinja2 import Environment, FileSystemLoader
env = Environment(loader=FileSystemLoader("templates"))
tpl = env.get_template("aio_audit_checklist.md.j2")
rendered = tpl.render(artifact_path="...", artifact_type="manuscript",
artifact_phase="pre-submission", ...)
#}
# AIO Audit — {{ artifact_path }}
**Artifact type**: {{ artifact_type }}
**Phase**: {{ artifact_phase }}
{% if journal_target -%}
**Journal target**: {{ journal_target }}
{% endif -%}
**Audit date**: {{ audit_date }}
**Reviewer**: {{ reviewer }}
{# Filter: artifact_type in applies_to, artifact_phase in applies_to_phase #}
{% set in_scope = items
| selectattr('applies_to', 'contains', artifact_type)
| selectattr('applies_to_phase', 'contains', artifact_phase)
| list -%}
{% set deferred = in_scope | selectattr('defers_to', 'defined') | list -%}
{% set actionable = in_scope | rejectattr('defers_to', 'defined') | list -%}
{# Sort actionable by expected_lift: high → medium → low #}
{% set high = actionable | selectattr('expected_lift', 'equalto', 'high') | list -%}
{% set med = actionable | selectattr('expected_lift', 'equalto', 'medium') | list -%}
{% set low = actionable | selectattr('expected_lift', 'equalto', 'low') | list -%}
## Top 5 — ranked by expected_lift (high → medium → low)
| Rank | ID | Rule | Lift | Status | Reason | Suggested edit |
|------|----|------|------|--------|--------|----------------|
{% set ordered = (high + med + low)[:5] -%}
{% for item in ordered -%}
| {{ loop.index }} | {{ item.id }} | {{ item.rule }} | {{ item.expected_lift }} | | | |
{% endfor %}
## Full applicable checklist
| ID | Rule | Lift | Priority | Status | Reason | Suggested edit |
|----|------|------|----------|--------|--------|----------------|
{% for item in (high + med + low) -%}
| {{ item.id }} | {{ item.rule }} | {{ item.expected_lift }} | {{ item.priority }} | | | |
{% endfor %}
## Deferred items (audit elsewhere)
| ID | Rule | Status | Defers to |
|----|------|--------|-----------|
{% for item in deferred -%}
| {{ item.id }} | {{ item.rule }} | | {{ item.defers_to }} |
{% endfor %}
## Out-of-phase items (NA — surface only if phase changes)
{% set out_of_scope = items
| selectattr('applies_to', 'contains', artifact_type)
| rejectattr('applies_to_phase', 'contains', artifact_phase)
| list -%}
{% if out_of_scope -%}
| ID | Rule | Applies in phase |
|----|------|------------------|
{% for item in out_of_scope -%}
| {{ item.id }} | {{ item.rule }} | {{ item.applies_to_phase | join(', ') }} |
{% endfor %}
{% else -%}
None.
{% endif %}
## Summary
- PASS: __ / {{ actionable | length }} actionable items
- Deferred items: {{ deferred | length }} (status only, detail elsewhere)
- Out-of-phase NA: {{ out_of_scope | length }}
- Top 5 fixes (ranked by expected_lift):
1.
2.
3.
4.
5.
## Cross-skill follow-up
{% if reporting_guideline_pending -%}
- [ ] Run `/check-reporting` (Section 1.6 anchor not yet verified). See `references/reporting_guideline_mapping.md` for the AIO ↔ guideline crosswalk.
{% endif -%}
{% if numerical_audit_pending -%}
- [ ] Run numerical-claim audit (data-integrity rule).
{% endif -%}
{% if humanize_pending -%}
- [ ] Run `/humanize` before submission (AI-pattern check).
{% endif %}
#!/usr/bin/env bash
# Regression test for academic-aio/scripts/batch_metadata_audit.py.
# Builds synthetic repos / HF cards (no committed data) and asserts: a clean
# repo reports no issues (exit 0), a repo missing README/CITATION/LICENSE fails
# under --fail-on-issue (exit 1), and an HF card carrying a PHI-shaped string is
# flagged. Stdlib-only (json/re), network-free, ASCII-only.
set -u
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
SCRIPT="$HERE/../scripts/batch_metadata_audit.py"
TMP="$(mktemp -d -t aio_audit_XXXX)"
trap 'rm -rf "$TMP"' EXIT
fail=0
check_exit() { local label="$1" want="$2"; shift 2
"$@" >/dev/null 2>&1; local got=$?
if [[ "$got" -eq "$want" ]]; then printf ' PASS %s\n' "$label"
else printf ' FAIL %s (exit %s, want %s)\n' "$label" "$got" "$want"; fail=$((fail+1)); fi
}
check() { local label="$1"; shift
if "$@" >/dev/null 2>&1; then printf ' PASS %s\n' "$label"
else printf ' FAIL %s\n' "$label"; fail=$((fail+1)); fi
}
[[ -f "$SCRIPT" ]] || { echo "ENV-ERR: batch_metadata_audit.py missing" >&2; exit 2; }
# --- Clean repo: README (DOI + citation + quickstart), CITATION.cff, LICENSE ---
CLEAN="$TMP/clean_repo"; mkdir -p "$CLEAN"
cat > "$CLEAN/README.md" <<'MD'
# Synthetic Tool
[](https://doi.org/10.5281/zenodo.0000000)
## Installation
pip install synthetic-tool
## How to cite
See CITATION.cff.
MD
cat > "$CLEAN/CITATION.cff" <<'CFF'
cff-version: 1.2.0
title: Synthetic Tool
version: 1.0.0
authors:
- family-names: Kim
given-names: Alice
orcid: https://orcid.org/0000-0002-1825-0097
CFF
echo "MIT License" > "$CLEAN/LICENSE"
check_exit "clean repo -> exit 0 (no issues, --fail-on-issue)" 0 \
python3 "$SCRIPT" "$CLEAN" --fail-on-issue
# --- Broken repo: directory exists but empty (all three artifacts missing) ---
BROKEN="$TMP/broken_repo"; mkdir -p "$BROKEN"
check_exit "repo missing README/CITATION/LICENSE -> exit 1" 1 \
python3 "$SCRIPT" "$BROKEN" --fail-on-issue
# Without --fail-on-issue the same audit still exits 0 (report-only).
check_exit "report-only mode -> exit 0 even with issues" 0 \
python3 "$SCRIPT" "$BROKEN"
# --- HF card with a PHI-shaped string (KR resident registration number) ---
CARD="$TMP/model_card.md"
cat > "$CARD" <<'MD'
---
license: mit
library_name: transformers
tags:
- medical
---
# Model
## Intended use
Research only.
## Training data
Synthetic records, e.g. subject 900101-1234567 was excluded.
## Evaluation
AUC reported.
## Limitations
Small sample.
## Ethical considerations
De-identified.
MD
JSON_OUT="$TMP/report.json"
python3 "$SCRIPT" --hf-card "$CARD" --output "$JSON_OUT" >/dev/null 2>&1
check "HF card report written" test -s "$JSON_OUT"
check "PHI pattern flagged in HF card" python3 -c "
import json
d=json.load(open('$JSON_OUT'))
issues=' '.join(d['hf_cards'][0]['issues'])
assert 'PHI' in issues, d['hf_cards'][0]['issues']"
check_exit "HF card with PHI -> exit 1 under --fail-on-issue" 1 \
python3 "$SCRIPT" --hf-card "$CARD" --fail-on-issue
echo "fail=$fail"; [[ "$fail" -eq 0 ]] && echo "ALL PASS" || echo "FAILURES: $fail"
exit "$fail"
#!/usr/bin/env bash
# Regression/challenge test for check_summary_box.py — deterministic, network-free,
# synthetic fixtures built at runtime. Covers each journal format + each hard rule.
set -u
HERE="$(cd "$(dirname "$0")" && pwd)"
SCRIPT="$HERE/../scripts/check_summary_box.py"
TMP="$(mktemp -d)"
trap 'rm -rf "$TMP"' EXIT
pass=0
fail=0
ck() {
local label="$1" expected="$2" actual="$3"
if [ "$expected" = "$actual" ]; then
printf ' PASS %-50s exit=%s\n' "$label" "$actual"
pass=$((pass + 1))
else
printf ' FAIL %-50s expected=%s actual=%s\n' "$label" "$expected" "$actual"
fail=$((fail + 1))
fi
}
run() { python3 "$SCRIPT" --out "$TMP/r.json" "$@" > /dev/null 2>&1; echo $?; }
# 1) conformant Key Points (3 one-claim bullets) -> exit 0
cat > "$TMP/kp_ok.md" <<'EOF'
## Key Points
- The model improved detection sensitivity in an internal test set.
- Specificity was preserved at the chosen operating threshold.
- External validation is still required before deployment.
EOF
ck "key_points conformant" 0 "$(run --manuscript "$TMP/kp_ok.md" --journal radiology --strict)"
# 2) wrong bullet count (2) -> NONCONFORMANT under --strict
cat > "$TMP/kp_bad.md" <<'EOF'
## Key Points
- The model improved detection sensitivity.
- Specificity was preserved.
EOF
ck "key_points wrong bullet count fails" 1 "$(run --manuscript "$TMP/kp_bad.md" --journal radiology --strict)"
# 3) Research in context missing a sub-block -> NONCONFORMANT
cat > "$TMP/ric_bad.md" <<'EOF'
## Research in context
**Evidence before this study** We searched PubMed for prior work.
**Added value of this study** This study adds an external cohort.
EOF
ck "research_in_context missing subblock fails" 1 "$(run --manuscript "$TMP/ric_bad.md" --journal lancet-digital-health --strict)"
# 4) Research in context complete -> CONFORMANT
cat > "$TMP/ric_ok.md" <<'EOF'
## Research in context
**Evidence before this study** We searched PubMed for prior work.
**Added value of this study** This study adds an external cohort.
**Implications of all the available evidence** Findings support a prospective trial.
EOF
ck "research_in_context complete conformant" 0 "$(run --manuscript "$TMP/ric_ok.md" --journal lancet-digital-health --strict)"
# 5) Plain-language summary over the band -> NONCONFORMANT
{ echo "## Plain-language summary"; for i in $(seq 1 260); do printf 'word '; done; echo; } > "$TMP/pls_bad.md"
ck "plain_language over-length fails" 1 "$(run --manuscript "$TMP/pls_bad.md" --journal npj-digital-medicine --strict)"
# 6) absent box -> NONCONFORMANT
echo "## Abstract" > "$TMP/none.md"
ck "absent box fails" 1 "$(run --manuscript "$TMP/none.md" --format key_points --strict)"
# 7) without --strict, a nonconformant box is reported but tolerated (exit 0)
ck "nonconformant tolerated without --strict" 0 "$(run --manuscript "$TMP/kp_bad.md" --journal radiology)"
echo "----"
echo "test_summary_box: $pass passed, $fail failed"
[ "$fail" -eq 0 ]
#!/usr/bin/env bash
# Regression test for academic-aio/scripts/validate_schema.py.
# Builds synthetic JSON-LD fixtures (no committed data) and asserts the
# validator's contract: a complete ScholarlyArticle passes; wrong @context,
# unknown @type, a missing required field, and a malformed DOI each fail.
# Stdlib-only (json/re), network-free, ASCII-only.
set -u
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
SCRIPT="$HERE/../scripts/validate_schema.py"
TMP="$(mktemp -d -t aio_schema_XXXX)"
trap 'rm -rf "$TMP"' EXIT
fail=0
check() { local label="$1" want="$2"; shift 2
"$@" >/dev/null 2>&1; local got=$?
if [[ "$got" -eq "$want" ]]; then printf ' PASS %s\n' "$label"
else printf ' FAIL %s (exit %s, want %s)\n' "$label" "$got" "$want"; fail=$((fail+1)); fi
}
[[ -f "$SCRIPT" ]] || { echo "ENV-ERR: validate_schema.py missing" >&2; exit 2; }
# Valid ScholarlyArticle (all required fields, canonical DOI).
cat > "$TMP/ok.jsonld" <<'JSON'
{
"@context": "https://schema.org",
"@type": "ScholarlyArticle",
"headline": "Synthetic Diagnostic Study",
"datePublished": "2026-01-01",
"author": [{"@type": "Person", "name": "Alice Kim"}],
"identifier": [{"@type": "PropertyValue", "propertyID": "DOI", "value": "10.1000/synthetic.2026.001"}],
"url": "https://example.org/article"
}
JSON
check "valid ScholarlyArticle -> exit 0" 0 python3 "$SCRIPT" "$TMP/ok.jsonld"
# Wrong @context.
sed 's#https://schema.org#https://example.com#' "$TMP/ok.jsonld" > "$TMP/bad_ctx.jsonld"
check "wrong @context -> exit 1" 1 python3 "$SCRIPT" "$TMP/bad_ctx.jsonld"
# Unknown @type.
sed 's/ScholarlyArticle/UnicornType/' "$TMP/ok.jsonld" > "$TMP/bad_type.jsonld"
check "unknown @type -> exit 1" 1 python3 "$SCRIPT" "$TMP/bad_type.jsonld"
# Missing required field (no "url"; still valid JSON).
cat > "$TMP/missing.jsonld" <<'JSON'
{
"@context": "https://schema.org",
"@type": "ScholarlyArticle",
"headline": "Synthetic Diagnostic Study",
"datePublished": "2026-01-01",
"author": [{"@type": "Person", "name": "Alice Kim"}],
"identifier": [{"@type": "PropertyValue", "propertyID": "DOI", "value": "10.1000/synthetic.2026.001"}]
}
JSON
check "missing required field -> exit 1" 1 python3 "$SCRIPT" "$TMP/missing.jsonld"
# Malformed DOI.
sed 's#10.1000/synthetic.2026.001#not-a-doi#' "$TMP/ok.jsonld" > "$TMP/bad_doi.jsonld"
check "malformed DOI -> exit 1" 1 python3 "$SCRIPT" "$TMP/bad_doi.jsonld"
# Mixed batch (one bad file) still fails overall.
check "batch with one bad file -> exit 1" 1 python3 "$SCRIPT" "$TMP/ok.jsonld" "$TMP/bad_ctx.jsonld"
echo "fail=$fail"; [[ "$fail" -eq 0 ]] && echo "ALL PASS" || echo "FAILURES: $fail"
exit "$fail"
Related skills
FAQ
Does academic-aio rewrite my paper?
No, it surfaces a visible PASS/PARTIAL/FAIL checklist with concrete fixes and never applies AIO edits silently.
What reporting guidelines does it integrate?
It integrates TRIPOD+AI, CLAIM 2024, STARD-AI, TRIPOD-LLM, and DECIDE-AI requirements with GEO principles.