
Nature Article Writer
- 72 installs
- 3 repo stars
- Updated June 29, 2026
- tristanmanchester/agent-skills
Helps with ai & agent building tasks during AI-assisted development.
About
nature-article-writer is a Claude Code skill for ai & agent building. It helps solo builders move faster with AI-assisted coding.
- nature-article-writer
- AI & Agent Building
- AI-coding skill
Nature Article Writer by the numbers
- 72 all-time installs (skills.sh)
- +1 installs in the week ending Jul 27, 2026 (Skillselion tracking)
- Ranked #5,605 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
- Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/tristanmanchester/agent-skills --skill nature-article-writerAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 72 |
|---|---|
| repo stars | ★ 3 |
| Last updated | June 29, 2026 |
| Repository | tristanmanchester/agent-skills ↗ |
What it does
Helps with ai & agent building tasks during AI-assisted development.
Files
Nature Article Writer
Write and revise primary research manuscripts so they feel editorially mature: precise, proportionate, detailed where detail matters, and genuinely pleasurable for a scientist to read. Beautiful Nature-style prose is not ornate prose. It is clear, load-bearing prose with strong logic, good sentence movement, and no wasted claims.
This skill optimises for editorial quality, reader trust, and human-sounding scientific prose. It does not optimise for AI-detector evasion.
Do not imitate a named living author. Emulate journal expectations, the user's own prior writing if supplied, and the specific paper's evidence profile.
When to activate this skill
Use this skill when the user:
- names Nature or a Nature Portfolio journal
- asks for "Nature-style", "Nature journal", or "high-impact journal" scientific writing
- wants a title, summary paragraph, abstract, introduction, results, discussion, methods, figure legend, presubmission package, cover letter, or reviewer response
- wants a scientific draft to sound more natural, less generic, less formulaic, or less obviously machine-written
- wants to convert notes, figures, bullet points, or a rough draft into a submission-ready manuscript
- wants a diagnostic pass on manuscript structure, claim calibration, prose quality, or compliance
Success standard
A strong output from this skill should feel like it was written by a careful scientist-editor who understands both the data and the journal:
- the central claim is evident early and never overstated
- adjacent-field readers can follow the logic without drowning in jargon
- each paragraph has a job
- each sentence earns its place
- results progress by question and answer, not lab chronology
- the prose is varied but restrained
- limitations are surfaced before reviewers must drag them out
- end matter and policy-sensitive statements are present or explicitly marked as missing
Non-negotiables
- Never invent data, methods, figures, ethics approvals, accession numbers, references, software versions, statistical results, or journal-specific limits.
- Never strengthen a claim beyond the evidence actually supplied.
- Never hide uncertainty. Mark missing facts explicitly with
[confirm],[insert ref],[insert accession], or a shortIssues to confirmlist. - Never use AI-generated figures or image content for publication.
- Never copy distinctive phrasing from published papers. Use exemplars for structure, rhythm, and level-setting, not sentence theft.
- If AI did more than copy editing, remind the user to check whether disclosure is required under the target journal's policy. Human authors remain accountable for the final text.
Default workflow
1. Build the manuscript brief
Infer or assemble the minimum brief:
- target journal and content type
- one-sentence central claim
- why it matters outside the immediate subfield
- evidence ladder: 3-6 concrete results, figures, or analyses that support the claim
- strongest prior work and the precise gap
- strongest limitation or boundary condition
- data, code, and materials availability
- ethics or compliance facts if humans, animals, clinical samples, or sensitive data are involved
If the user has scattered notes, use assets/manuscript-brief-template.md.
2. Calibrate before you draft
Use both of these calibration layers when possible.
A. Journal calibration
Consult references/modes.md and references/journal-calibration.md.
- Choose the closest bundled mode:
nature-articlenature-letterportfolio-articleportfolio-letter- If the user names a specific journal and web access is available, verify the live guide and inspect 2-4 recent primary research papers from that journal.
- Build a short internal style card: title texture, opening-paragraph shape, heading policy, legend density, end-matter order, and how aggressively claims are hedged.
B. Exemplar anchoring
If the user supplies their own accepted papers, lab style guides, or a high-quality draft they want to sound like, use references/exemplar-anchoring.md and optionally run:
python3 scripts/prose_fingerprint.py --candidate draft.md --reference exemplar1.md exemplar2.md --format textImitate broad habits such as sentence length range, paragraph density, degree of overt signposting, and tolerance for technical detail. Do not imitate distinctive turns of phrase.
3. Build the editorial architecture
Before line-level drafting, create:
- a one-sentence paper promise
- a figure-claim matrix
- a paragraph map for the major sections
Use:
- assets/editorial-blueprint-template.md
- assets/figure-claim-matrix-template.md
- assets/paragraph-map-template.md
- references/editorial-architecture.md
This is the main upgrade over a generic "write the paper" prompt. The prose improves when the structure is load-bearing before wording starts.
4. Draft in evidence order, not display order
Default drafting order: 1. figure plan and one-sentence take-home message for each figure 2. Results 3. Methods 4. Discussion or concluding synthesis 5. opening context paragraph or Introduction 6. summary paragraph or abstract 7. title 8. figure legends 9. availability statements and other end matter 10. cover letter or presubmission material if requested
Starting from figures and claims produces more grounded prose than starting from the title or abstract.
5. Shape paragraphs deliberately
Every paragraph needs:
- a topic sentence that names the paragraph's job
- evidence or reasoning that advances the job
- a final stress position that lands the important point or hands the reader to the next paragraph
Use references/sentence-craft.md and assets/paragraph-map-template.md. Prefer old-to-new information flow, concrete verbs, and sentences that end on the point that matters.
6. Run a human-voice pass tuned for scientific prose
Consult references/voice-and-variation.md.
Target common instruction-tuned LLM artefacts without turning the paper chatty:
- overuse of present-participial clause chains
- noun-heavy nominalized phrasing
- conveyor-belt transitions (
Additionally,Moreover,Importantly,Taken together) - inflated significance language
- generic concluding sentences that claim importance without stating the implication
- repeated weak sentence openings (
This,These,It,We) - flat sentence rhythm and uniform paragraph shape
Do not blindly ban passive voice, repetition, or technical compounds. Scientific prose needs all three sometimes. The aim is selective repair.
7. Run integrity and compliance checks
Use references/integrity-and-compliance.md and, if Python 3 is available:
python3 scripts/nature_preflight.py --input draft.md --mode nature-article --format textor the relevant mode:
python3 scripts/nature_preflight.py --input draft.md --mode portfolio-article --format textUse the report to fix:
- title length and title texture
- missing required or expected sections
- opening paragraph length or structure
- missing Data Availability or Code Availability sections
- bracket citations that need conversion
- figure legends missing title sentences or statistical detail
- hype words, generic AI-ish phrases, rhythm flatness, or repeated weak openers
- obvious overclaim or unsupported forward-looking claims
If scripting is unavailable, do the same checks manually.
Section guidance
Use references/section-rubric.md for detailed section-by-section repair rules. High-level rules:
Title
- clear, searchable, and readable outside the narrow subfield
- avoid hype, puns, slogans, rhetorical questions, and vague grandeur
- for main Nature, aim for roughly 75 characters and avoid numbers, acronyms, abbreviations, and punctuation unless essential
Summary paragraph or abstract
- broad context first, then the specific gap
- state the main finding once, cleanly
- end on the most defensible implication
- use references in main Nature-style summary paragraphs when appropriate
- avoid stuffing it with data scraps, acronyms, or methodological clutter
Introduction or opening
- move quickly from field context to unresolved problem
- do not write a mini-review
- finish with what the paper does, why the approach is appropriate, and what kind of answer the paper delivers
Results
- organise by conceptual question or figure logic, not by when experiments happened
- make every subsection claim-bearing
- distinguish observation from interpretation
- mention only the numbers that advance the story
Discussion
- say what the work establishes, what it suggests, and where it stops
- surface the main limitation before the reviewer does
- end with the most defensible field-level implication, not a cinematic future vision
Methods
- concise but genuinely informative
- include the details that govern interpretability and reproducibility
- use short subsection headings and concrete labels
Figure legends
- begin with a brief title sentence
- describe panels in sequence
- define statistics, sample sizes, centre values, and error bars where relevant
- stand on their own as far as reasonable
Working modes
A. Full manuscript from notes
Deliver:
- a short manuscript brief
- a figure-led blueprint
- the full draft in the chosen template
- an unresolved-gap list
- optional preflight and fingerprint reports
B. Rewrite an existing draft
Process: 1. identify the actual claim 2. preserve data and meaning 3. repair structure 4. line-edit for Nature-style clarity and reader movement 5. flag claims that need verification
C. Abstract or summary paragraph only
Where useful, provide 2-3 versions:
- conservative
- balanced or default
- slightly bolder but still defensible
Label the trade-off in claim strength.
D. Reviewer response or rebuttal
Use assets/reviewer-response-template.md.
Rules:
- answer every point directly
- quote the reviewer briefly, then answer
- specify exactly what changed and where
- concede valid points plainly
- when declining a request, explain why and offer the nearest rigorous alternative
E. Presubmission enquiry or cover letter
Use:
- assets/presubmission-enquiry-template.md
- assets/cover-letter-template.md
Write for editors, not reviewers. Explain:
- the central advance
- why broad readers should care
- why the evidence is strong enough
- why the paper fits this journal
- what the paper is not claiming
Output style
When returning a draft or rewrite:
- state the chosen journal mode and key assumptions
- give the manuscript in clean, journal-ready prose
- include a brief
Issues to confirmlist only when necessary - do not pad the answer with generic writing advice unless the user asked for it
When returning diagnostics:
- prioritise the 5-10 issues that will most improve reader trust and editorial fit
- separate structural issues from line-edit issues
- suggest concrete rewrites, not vague criticism
Fast heuristics for excellent Nature-style prose
- A clear limitation usually makes the paper sound stronger.
- If a phrase could appear in almost any paper, cut or replace it.
- The sentence ending matters. Land on the point that earns emphasis.
- Replace abstract importance language with the exact implication.
- Good scientific prose can be vivid without being promotional.
- Detail is welcome when it is the detail that lets the reader trust the claim.
- "Human sounding" here means precise, calm, varied, and specific, not casual.
Bundled references
- references/modes.md
- references/section-rubric.md
- references/editorial-architecture.md
- references/sentence-craft.md
- references/voice-and-variation.md
- references/journal-calibration.md
- references/exemplar-anchoring.md
- references/integrity-and-compliance.md
- references/research-notes.md
Bundled assets
- assets/manuscript-brief-template.md
- assets/editorial-blueprint-template.md
- assets/figure-claim-matrix-template.md
- assets/paragraph-map-template.md
- assets/nature-article-template.md
- assets/nature-letter-template.md
- assets/portfolio-article-template.md
- assets/portfolio-letter-template.md
- assets/presubmission-enquiry-template.md
- assets/cover-letter-template.md
- assets/reviewer-response-template.md
Cover letter template
Dear Editors,
Please consider our manuscript, "[Title]," for publication as a [article type] in [journal].
The manuscript addresses [broad problem] by showing that [central answer]. The advance is significant for the journal's readership because [fit and relevance]. The evidence rests on [2-3 concise evidential pillars].
We believe the paper is well suited to [journal] because:
- [fit reason 1]
- [fit reason 2]
- [fit reason 3]
The manuscript is deliberately bounded in scope. It does not claim [non-claim / limitation], and we state [important caveat] explicitly in the manuscript.
All authors have approved the manuscript and [declare no competing interests / declare the following competing interests: ...]. Data and code availability statements are included in the manuscript [or: will be finalised with repository accessions before acceptance].
Thank you for your consideration.
Sincerely, [Corresponding author]
Editorial blueprint template
1. Paper promise
- One-sentence paper promise:
- Reader takeaway in plain scientific English:
- Why this is not just an incremental result:
2. Journal calibration card
- Target journal / article type:
- Opening form:
- Heading policy:
- Title texture:
- Legend density:
- Methods placement:
- End-matter order:
- Degree of claim restraint expected:
3. Story spine
Fill one line for each major move in the paper.
1. Broad problem: 2. Precise unresolved question: 3. Why the chosen approach can answer it: 4. First load-bearing result: 5. Second load-bearing result: 6. Mechanistic or explanatory step: 7. Limitation or boundary condition: 8. Most defensible implication:
4. Figure order
For each figure:
- Figure number:
- Question:
- Key observation:
- Allowed claim:
- Transition to next figure:
5. Opening architecture
- First sentence broad frame:
- Second sentence sharper context:
- Gap sentence:
- Main answer sentence:
- Implication sentence:
- What to leave out of the opening:
6. Discussion commitments
- What the paper establishes:
- What it only suggests:
- Limitation to surface explicitly:
- Broader implication that is justified:
- Broader implication that is too much:
7. Editorial cuts
If the paper must be shortened by 10-20%, cut:
- Paragraph / section 1:
- Paragraph / section 2:
- Paragraph / section 3:
Figure-claim matrix
Use one row per figure or conceptual unit.
| Figure | Scientific question | Core observation | Allowed claim | Main uncertainty / caveat | Key statistical details for legend | Methods detail the reader must know |
|---|---|---|---|---|---|---|
| Fig. 1 | ||||||
| Fig. 2 | ||||||
| Fig. 3 | ||||||
| Fig. 4 | ||||||
| Fig. 5 | ||||||
| Extended Data / Supplementary |
Notes
- A figure earns only the claim that the evidence supports.
- If the figure's job is unclear, the prose around it will usually become generic.
- Write legends from this matrix, not from memory.
Manuscript brief template
Use this template when the source material is messy, incomplete, or spread across notes, figures, or slides.
Target
- Journal:
- Article type:
- Closest bundled mode:
- Specific author-guide details already verified:
- Details that still need checking:
Paper promise
- One-sentence central claim:
- What the paper is not claiming:
- Why it matters outside the immediate subfield:
- Why this journal's readers should care now:
Evidence ladder
List 3-6 load-bearing results.
1. Figure / analysis:
- Question answered:
- Observation:
- Allowed claim:
- Main caveat:
2. Figure / analysis: 3. Figure / analysis: 4. Figure / analysis: 5. Figure / analysis: 6. Figure / analysis:
Gap and prior work
- Strongest prior work:
- Exact unresolved problem:
- Why existing approaches were insufficient:
Limitations and risk points
- Strongest limitation:
- Most likely reviewer concern:
- Most dangerous overclaim to avoid:
- Alternative explanation that must be addressed:
Methods and compliance
- System / samples / cohort:
- Key methods:
- Statistics:
- Ethics / approvals / consent:
- Reporting checklist likely required:
- Data availability status:
- Code availability status:
- Materials availability status:
Editorial notes
- Broad-reader bridge sentence:
- Candidate title keywords:
- Acronyms that are essential:
- Acronyms to avoid:
- Unknown facts to confirm before submission:
Title
[Concise, broad-reader-facing title]
Summary paragraph
[2-3 sentences of broad context and rationale with references where appropriate.] [Specific gap or unresolved problem.] [Here we show / equivalent main answer sentence.] [Immediate implication, scaled to the evidence.]
Introduction
[Move quickly from field context to the exact unresolved question.] [State why the chosen system, method, or perturbation can answer it.] [End with what the paper does.]
Results
[Claim-led subheading]
[Question, approach, observation, and what it allows the paper to claim.]
[Claim-led subheading]
[Question, approach, observation, and what it allows the paper to claim.]
[Claim-led subheading]
[Question, approach, observation, and what it allows the paper to claim.]
Discussion
[What the study establishes.] [How it compares with prior work.] [Main limitation or boundary condition.] [Most defensible broader implication.]
Figure legends
Figure 1 | [Brief title sentence]
[Panel sequence, statistics, and symbols.]
Figure 2 | [Brief title sentence]
[Panel sequence, statistics, and symbols.]
Methods
[Short bold subsection]
[Operational detail.]
[Short bold subsection]
[Operational detail.]
Data Availability
[Repository / accession / controlled-access statement.]
Code Availability
[Repository / release condition.]
References
1. [Insert references]
Acknowledgements
[Insert]
Funding
[Insert]
Author Contributions
[Insert]
Competing Interests
[Insert]
Title
[Concise title]
Introductory paragraph
[Referenced opening paragraph with broad context, precise gap, main answer, and immediate implication.]
[Continue directly into the main narrative without main-text headings.]
[Results-driven continuous prose.] [Move from observation to mechanism or explanation.] [Surface limitations before the close.] [End on the strongest defensible implication.]
Figure legends
Figure 1 | [Brief title sentence]
[Panel sequence and statistics.]
Figure 2 | [Brief title sentence]
[Panel sequence and statistics.]
Methods
[Concise methods with clear operational detail.]
Data Availability
[Insert statement.]
Code Availability
[Insert statement if relevant.]
References
1. [Insert references]
Acknowledgements
[Insert]
Author Contributions
[Insert]
Competing Interests
[Insert]
Paragraph map template
Use this before drafting or when repairing a bloated section.
| Section | Paragraph job | Topic sentence subject | Evidence / reasoning to include | Best sentence-ending emphasis | Figure / table tie | Overclaim risk |
|---|---|---|---|---|---|---|
| Opening / Intro P1 | ||||||
| Opening / Intro P2 | ||||||
| Results P1 | ||||||
| Results P2 | ||||||
| Results P3 | ||||||
| Results P4 | ||||||
| Discussion P1 | ||||||
| Discussion P2 |
Checks
- Can you say each paragraph's job in 5-8 words?
- Does the topic sentence start from what the reader already knows?
- Does the paragraph end on the point that matters?
- Does each paragraph prepare the next one?
Title
[Journal-appropriate title]
Abstract
[Context.] [Rationale or gap.] [Main result.] [Implication.]
Introduction
[Focused framing and gap.]
Results
[Claim-led subheading]
[Evidence.]
[Claim-led subheading]
[Evidence.]
Discussion
[Meaning, scope, comparison with prior work, limitation.]
Methods
[Operational detail.]
Data Availability
[Insert statement.]
Code Availability
[Insert statement if relevant.]
References
1. [Insert references]
Acknowledgements
[Insert]
Funding
[Insert]
Author Contributions
[Insert]
Competing Interests
[Insert]
Title
[Journal-appropriate concise title]
Introductory paragraph
[Concise context, gap, main answer, and implication.]
[Continue the main narrative with minimal section scaffolding.] [Keep the argument compact and evidence-led.]
Methods
[Concise but complete enough for interpretation.]
Data Availability
[Insert statement.]
Code Availability
[Insert statement if relevant.]
References
1. [Insert references]
Acknowledgements
[Insert]
Author Contributions
[Insert]
Competing Interests
[Insert]
Presubmission enquiry template
Dear Editors,
We wish to enquire whether our manuscript, tentatively titled "[Title]", may be suitable for consideration as a [article type] in [journal].
Our study addresses [broad problem]. Specifically, we show that [central finding]. This matters because [broad-reader significance], and because the work resolves [precise gap or controversy].
The manuscript's main evidential strengths are: 1. [Load-bearing result] 2. [Independent supporting result] 3. [Mechanistic / validation / scope result]
The paper fits [journal] because [fit argument tied to readership], while remaining appropriately bounded in its claims: we do not claim [important non-claim or limitation]. We believe this combination of [advance] and [broad relevance] may be of interest to the journal's readership.
A short manuscript summary is below: [3-5 sentence summary, not pasted from the abstract]
We would be grateful for your advice on suitability.
Sincerely, [Corresponding author]
Reviewer response template
We thank the reviewers for their careful reading and constructive comments. Below we address each point in turn. Reviewer comments are reproduced in bold, followed by our responses and the changes made in the manuscript.
Reviewer 1
Comment 1
[Paste or paraphrase the reviewer comment briefly.]
Response: [Answer directly. State what the reviewer is right about, what was changed, and why.]
Changes in manuscript:
- Page / section / figure:
- New text or analysis:
- If not adopted, why not:
Comment 2
[Paste or paraphrase briefly.]
Response: [Direct answer.]
Changes in manuscript:
- Page / section / figure:
- New text or analysis:
- If not adopted, why not:
Reviewer 2
[Repeat.]
Closing
We hope that these revisions and clarifications have addressed the reviewers' concerns. Where we were unable to perform a requested analysis or experiment, we now state the limitation explicitly and explain why the request exceeds the present study's scope.
{
"skill_name": "nature-article-writer",
"evals": [
{
"id": 1,
"prompt": "Turn the following rough study notes and figure bullets into a full Nature-style drafting pack: manuscript brief, editorial blueprint, figure-claim matrix, referenced summary paragraph, and a figure-led Results outline. Do not invent missing data. End with the unresolved facts that must be confirmed before submission.",
"expected_output": "A structured drafting pack with brief, blueprint, figure-claim matrix, broad-reader summary paragraph, figure-led outline, and a clear unresolved-facts list.",
"assertions": [
"The output includes a one-sentence central claim and a statement of what the paper is not claiming.",
"The output includes a figure-claim matrix or equivalent structure tying each figure to an allowed claim.",
"The summary paragraph is written for broad scientific readers rather than only niche specialists.",
"The model does not invent accession numbers, ethics approvals, or numerical results that were not supplied.",
"The output ends with unresolved factual gaps or assumptions that still need confirmation."
]
},
{
"id": 2,
"prompt": "Rewrite this jargon-heavy, slightly generic abstract and opening for Nature. Preserve every factual claim, but make the prose more precise, more pleasurable to read, and less obviously AI-shaped. Then explain the 5 most important edits you made and why.",
"expected_output": "A rewritten abstract/opening with preserved facts, clearer reader movement, fewer generic phrases, and a short, concrete explanation of the most important edits.",
"assertions": [
"The rewrite preserves the original factual content.",
"The rewrite reduces hype words and generic significance phrases.",
"The rewrite improves sentence variety or reduces repetitive sentence openings.",
"The explanation cites concrete edits rather than generic advice.",
"The tone remains scientific rather than becoming casual or promotional."
]
},
{
"id": 3,
"prompt": "Using these three accepted papers from our group as style anchors, draft the Introduction and first two Results subsections for our new manuscript. Match the restraint, sentence movement, and paragraph density of the exemplars without copying their wording. Then run a prose fingerprint comparison and tell me where the draft is still misaligned.",
"expected_output": "A draft aligned to supplied exemplars at the level of broad habits, followed by a prose-fingerprint style comparison and revision priorities.",
"assertions": [
"The response uses the exemplars as style anchors without copying distinctive phrasing.",
"The draft identifies and respects the paper's central claim and evidence ladder.",
"The follow-up comparison reports measurable stylistic alignment or mismatch.",
"The revision priorities focus on broad habits such as rhythm, nominalization, signposting, or opener variety."
]
},
{
"id": 4,
"prompt": "Convert this chronological experiment log into a Nature-style Results section organised by question and answer rather than by lab order. Make the argument easy for an adjacent-field reader to follow, and flag anywhere the evidence does not yet support the interpretation.",
"expected_output": "A Results section reorganised by conceptual logic, with explicit flagging of claims that outrun the evidence.",
"assertions": [
"The Results section is organised around conceptual questions or figure logic rather than chronology.",
"The prose distinguishes observation from interpretation.",
"The output flags overclaim or missing-evidence points explicitly.",
"The rewritten Results remain specific and evidence-led rather than generic."
]
},
{
"id": 5,
"prompt": "Draft a Discussion for this Nature-style manuscript. Start by stating what the study establishes, compare it with the strongest prior work, surface the key limitation before the reviewer does, and end on the most defensible implication rather than a generic future-work line.",
"expected_output": "A Discussion section with calibrated claims, comparison to prior work, explicit limitation handling, and a restrained but meaningful ending.",
"assertions": [
"The Discussion separates what is established from what is only suggested.",
"A real limitation or scope boundary is included before the ending.",
"The final paragraph does not rely on generic phrases like 'opens new avenues' or 'highlights the importance of'.",
"The implication is scaled to the evidence rather than inflated."
]
},
{
"id": 6,
"prompt": "Write figure legends for Figures 1-4 from these panel notes, including brief title sentences, panel descriptions, sample sizes, centre values, error bars, and statistical tests where provided. Use placeholders where the information is still missing rather than inventing it.",
"expected_output": "A complete figure-legend block with title sentences, panel-wise descriptions, and explicit placeholders for missing statistical details.",
"assertions": [
"Each legend begins with a brief title sentence.",
"Panel descriptions appear in sequence and are easy to follow.",
"Sample-size or statistics information is included when provided.",
"Missing statistical information is marked with placeholders rather than invented."
]
},
{
"id": 7,
"prompt": "Turn this draft into a presubmission enquiry package for Nature: an editor-facing cover paragraph, a referenced summary paragraph in Nature format, and a concise journal-fit note. Do not simply paste the abstract into the cover paragraph.",
"expected_output": "A presubmission package that separates editor framing from manuscript framing and makes an explicit fit argument.",
"assertions": [
"The cover paragraph is editor-facing and not just a copied abstract.",
"The Nature summary paragraph is broad, concise, and appropriately referenced in form.",
"The fit note explains why the work matters to an interdisciplinary readership.",
"The output states at least one meaningful non-claim or boundary condition."
]
},
{
"id": 8,
"prompt": "Run a full Nature-style preflight on this draft. Report structural issues, title problems, weak availability statements, legend gaps, and any AI-ish prose patterns. Then rewrite only the three paragraphs that would most improve reader trust.",
"expected_output": "A prioritised preflight report plus targeted paragraph rewrites focused on high-value fixes.",
"assertions": [
"The report identifies missing required or expected sections when present.",
"The report checks for data and code availability statements.",
"The report flags generic phrases, hype, or rhythm/opener problems when present.",
"The follow-up rewrites target the highest-value problems rather than making broad superficial edits."
]
}
]
}Editorial architecture
Strong Nature-style papers are usually won at the architecture stage, not the adjective stage. Before drafting sentences, design the route the reader will take.
The paper promise
Write a one-sentence promise:
- what question the paper answers
- what the main answer is
- why that answer matters beyond the narrowest subfield
If the promise is fuzzy, the paper will be fuzzy.
Figure-first logic
Build the story from figures, not from chronology.
For each figure or conceptual block, write: 1. the question it answers 2. the observation or result 3. the allowed claim 4. the uncertainty or limitation attached to that claim 5. the next question it unlocks
If a figure has no job in this chain, it probably belongs in supplementary or must be reframed.
Preferred result progression
A robust sequence is often: 1. establish the phenomenon or descriptive pattern 2. show that the measurement or perturbation is trustworthy 3. identify mechanism, heterogeneity, or explanatory structure 4. test the key alternative explanation 5. show scope, limitation, or generality
Not every paper needs all five steps, but papers that jump directly from observation to sweeping implication often read weakly.
Paragraph jobs
Every paragraph should have one dominant job:
- set up a question
- state an observation
- interpret a result
- compare with prior work
- delimit scope
- transition to the next evidential need
Paragraphs fail when they try to do all jobs at once.
Question-answer cadence
A useful rhythm inside Results is:
- opening sentence: what is being asked or tested
- middle: what was done and what was found
- closing sentence: what this now allows the paper to claim or test next
This cadence produces momentum without overt theatricality.
The broad-reader bridge
Nature-style writing often needs one additional layer beyond specialist writing: a bridge for intelligent readers from nearby fields.
Build that bridge by:
- naming the scientific object and the problem early
- unpacking dense concepts the first time they appear
- limiting acronym load
- stating why the specific measurement, model, or system matters
- avoiding methods detail in the very sentences where readers are trying to learn why the result matters
Where beauty comes from
A paper is pleasurable to read when:
- the reader is rarely forced to re-orient
- claims arrive at the moment the evidence earns them
- important sentences end with the important word or idea
- there is no fake suspense and no inflated payoff
- the prose feels confident because the structure is confident
Structural repairs that usually help
If the draft feels generic
Rebuild the figure-claim matrix. Generic prose often signals generic structure.
If the draft feels dense but slippery
Move from abstraction to noun-level specificity. Ask which sentence actually carries the evidence.
If the Results sag in the middle
Identify whether two subsections are really doing the same argumentative work. Merge or re-order.
If the Discussion inflates
Split the ending into:
- what is established
- what remains uncertain
- what the clearest implication is
If the opening is flat
Check whether the first three sentences name:
- the broad problem
- the precise gap
- the kind of answer the paper provides
Blueprint checklist
Before you line-edit, ensure you can answer:
- What is the one-sentence claim?
- Which figure most directly earns it?
- Which result most changed the story when it was discovered?
- Which limitation is serious enough that readers will look for it?
- Why does this journal's readership care?
- Which two paragraphs would you cut first if you had to shorten the paper by 15%?
If you cannot answer these quickly, the paper needs architectural work before stylistic work.
Exemplar anchoring
The best way to make a manuscript feel human and specific is often to anchor it to writing the user already owns or genuinely admires. Use exemplars carefully.
Best exemplars
Prefer, in order: 1. the user's own accepted papers 2. lab or group style guides 3. internal documents the user authorises as tone anchors 4. recent papers from the target journal used only for calibration, not imitation
Avoid using a named outside author's paper as a phrasing model.
What to learn from exemplars
Extract habits, not sentences:
- average sentence length and variation
- how often paragraphs begin with explicit signposts
- whether the prose tolerates long synthesis sentences
- how dense the technical detail is before explanation arrives
- how aggressively claims are hedged
- how titles are shaped
- whether conclusions end narrowly or expansively
What not to copy
Do not transplant:
- distinctive metaphors
- signature turns of phrase
- unusual clause patterns repeated from the exemplar
- uncommon title formulas
The point is to align texture and discipline, not to ventriloquise another writer.
Suggested workflow
1. Read the exemplar(s). 2. Build a short style card in plain language. 3. Draft the manuscript. 4. Compare draft and exemplars with:
python3 scripts/prose_fingerprint.py --candidate draft.md --reference exemplar1.md exemplar2.md --format text5. Revise only the dimensions that are clearly misaligned and matter for readability.
When not to anchor heavily
Do not anchor too strongly when:
- the exemplar is from a different article type
- the exemplar is older and the journal has changed style
- the exemplar is exceptionally idiosyncratic
- the current paper needs more explanatory scaffolding than the exemplar did
Productive framing
Good:
- "Match the restraint and sentence movement of these two accepted papers."
Less good:
- "Make this sound exactly like Author X."
Useful style-card fields
- title texture
- opening strategy
- paragraph density
- mean sentence length
- sentence-length variation
- tolerance for explicit signposting
- main sources of emphasis
- common ending shapes
- hedge density
- acronym density
Integrity and compliance
Use this file to keep the manuscript publication-ready and policy-aware.
Core publication disciplines
- Do not invent missing facts. Use placeholders and a clear
Issues to confirmlist. - Keep the wording of claims proportionate to the evidence.
- Separate direct results from interpretation.
- If a statement depends on a reference, a figure, or a method detail, make that dependency visible.
AI-related policy discipline
- Large language models are not authors.
- AI-assisted copy editing for readability and language is generally treated differently from deeper scientific use, but humans remain responsible for the manuscript.
- If AI contributed beyond surface editing, the user should verify the current policy for disclosure in the target journal.
- Do not generate or submit AI-generated research figures or images.
Data and code availability
For original research, expect a Data Availability statement. Where custom code is central to the findings, expect a Code Availability statement.
Data Availability template
Data Availability The data that support the findings of this study are available from [repository / accession / controlled-access route]. [If restrictions apply, explain them clearly rather than hiding them.]
Code Availability template
Code Availability Custom code used for [analysis / modelling / figure generation] is available at [repository URL / DOI / accession]. [If release will occur on publication or on reasonable request, say so precisely.]
Do not fabricate repository links or accession numbers.
Reporting summaries and methods discipline
Some Nature and Nature Portfolio areas require reporting summaries or field-specific checklists. Where relevant:
- note that the reporting summary must be completed for submission
- ensure Methods include the operational details that govern interpretation
- make sure statistical analysis is described concretely, not only with stock phrases
Figure-legend checklist
Each figure legend should usually:
- begin with a brief title sentence
- describe panels in order
- define sample size
n - define centre values and error bars
- state statistical tests and exact P values where appropriate
- decode colours, symbols, and abbreviations
Title checklist
For main Nature-style submissions, titles should usually:
- be concise
- be understandable outside the immediate subfield
- avoid acronyms and abbreviations unless truly essential
- avoid decorative punctuation and slogans
End-matter checklist
Where relevant, check for:
- Methods
- Data Availability
- Code Availability
- References
- Acknowledgements
- Funding
- Author Contributions
- Competing Interests
- Additional Information / Correspondence
- Extended Data legends
- Ethics statement language if required by the study design
Reviewer-response discipline
Never promise analyses, experiments, or deposits that have not been done or approved. If the manuscript will be revised in a specific way, say exactly what will change and where.
Practical rule
A manuscript sounds more credible when it openly names:
- what the study can support
- what remains unresolved
- what supporting files or statements are still pending
Journal calibration
When a specific journal is named, do not rely only on generic Nature-like instincts. Calibrate to the live journal and to recent exemplars.
Why calibration matters
Two manuscripts can both be good scientific prose but differ in:
- title texture
- abstract versus opening-paragraph expectations
- tolerance for subheadings
- how much context appears in Results
- how detailed legends tend to be
- whether discussion is integrated or separate
- how much of the story is expected in Extended Data or Supplementary Information
Step 1: read the live guide
Extract only the details that affect execution:
- article type
- title rules
- abstract or summary requirements
- heading policy
- methods placement
- figure or legend expectations
- end-matter requirements
- data, code, and reporting requirements
Do not hard-code guessed word limits if the guide is not in hand.
Step 2: inspect recent papers
Read 2-4 recent primary research papers in the exact journal and article type.
Build a compact style card with:
- average title length and texture
- how the opening paragraph is built
- how soon the paper reaches the core problem
- whether Results subsections are claim-led or method-led
- how cautious the discussion sounds
- how much detail is pushed into legends, Methods, Extended Data, or Supplementary Information
Step 3: abstract structure, not wording
Use recent papers to calibrate:
- the density of context
- how broad the final implication can be
- how explicit limitations are
- whether the journal tolerates sentence-level stylistic flourishes
Never reuse distinctive phrasing from exemplar papers.
Step 4: reconcile journal style with the user's voice
Priority order: 1. scientific accuracy 2. the target journal's visible house expectations 3. the user's own established writing habits 4. local elegance preferences
If the user's preferred style conflicts with the journal's visible expectations, follow the journal for externally visible features and keep the user's voice in sentence rhythm and paragraph texture.
Step 5: state what you calibrated
When it matters, tell the user:
- which journal mode you chose
- which live features were verified
- which details still need checking
A good skill should make its assumptions legible rather than invisible.
Modes
Use these bundled modes as defaults when the user does not name a specific journal, or when live checking is unavailable. They are deliberately conservative. If the user names a specific Nature Portfolio journal and web access is available, verify the live guide before finalising exact limits, heading policy, and required end matter.
Quick mode table
| Mode | Best use | Opening | Main-text headings | Methods | End matter |
|---|---|---|---|---|---|
nature-article | Main Nature Article | Referenced summary paragraph aimed at broad readers | Usually yes | After main text and figure legends | Separate Data Availability and usually Code Availability |
nature-letter | Nature-style Letter / concise report | Referenced introductory paragraph integrated into the main text | Usually no | After main text and legends | Same policy-sensitive end matter as above |
portfolio-article | Typical Nature Portfolio research article | Abstract, often unreferenced | Usually yes | Usually after Discussion or near the end | Data Availability and Code Availability as required by journal |
portfolio-letter | Concise Nature Portfolio Letter / brief report | Introductory paragraph, often concise and referenced where applicable | Often no or minimal | Near the end | Same as above |
nature-article
Shape
- broad, editor-facing title that non-specialists can parse
- referenced summary paragraph before the main text
- main text often organised with informative subheadings
- methods follow main text and legends
- separate Data Availability and Code Availability sections near the end
Use when
- the user explicitly says Nature Article
- the story supports a broad summary paragraph and several conceptual steps
- the narrative benefits from heading-based wayfinding
Common mistakes
- title too technical or too long
- summary paragraph written like a mini abstract full of numbers
- introduction too review-like
- discussion repeating results instead of framing meaning and limits
nature-letter
Shape
- concise title
- referenced opening paragraph integrated into the main narrative
- continuous main text, usually without main-text headings
- logic must travel paragraph-to-paragraph without section labels
- methods and legends appear after the main text
Use when
- the paper is compact and the story works as a continuous argument
- the user wants a high-density narrative with minimal heading scaffolding
Common mistakes
- too many internal subclaims in the opening paragraph
- heading-style Results sections smuggled into a Letter
- generic closing paragraph that inflates impact rather than setting scope
portfolio-article
Shape
- abstract, then Introduction / Results / Discussion / Methods or a close variant
- subheadings common and often helpful
- technical detail tolerance is usually higher than in main Nature
- journal-specific limits vary substantially
Use when
- the user names a Nature Portfolio journal outside the main Nature flagship
- the paper needs standard article scaffolding
- the abstract should do more of the context and rationale work
Common mistakes
- abstract too general for a specialist portfolio journal
- Introduction bloated with literature summary
- Results read as figure captions pasted into prose
portfolio-letter
Shape
- concise opening, often closer to a Letter than a full article
- lower tolerance for long scene-setting
- narrative economy matters
- live guide checking is especially important because letter-like formats vary across journals
Use when
- the journal has a brief report or Letter format
- the core story is sharp and does not require extended narrative scaffolding
Choosing among modes
Choose the mode by asking: 1. Does the journal expect a separate abstract or a referenced opening paragraph? 2. Are main-text headings normal or discouraged? 3. Is the paper's logic best served by a continuous argument or by explicit sections? 4. Does the journal expect distinct Data Availability and Code Availability sections? 5. Is the paper broad-reader-facing enough for a main Nature style summary paragraph?
When uncertain, say which bundled mode you used and which live details still need verification.
Research notes behind v2
This file records the main evidence sources that motivated the v2 design. It is not a substitute for checking live journal pages when the target journal is known.
Official Nature and Nature Portfolio guidance
Nature writing guidance
Key takeaways used in this skill:
- read the current author pages and recent issues before drafting to a specific journal
- write for readers outside the immediate discipline
- prefer active voice where it clarifies agency
- unpack dense concepts and minimise jargon and acronyms
- keep the paper focused on a concise message
Nature formatting guidance
Key takeaways used in this skill:
- main Nature Articles use a referenced summary paragraph aimed at broad readers
- titles should stay concise and accessible
- Methods, Data Availability, and Code Availability need explicit handling
- figure legends should begin with a brief title and define statistics clearly
AI and reporting policy
Key takeaways used in this skill:
- LLMs cannot be authors
- AI-assisted copy editing is treated differently from deeper scientific use
- reporting summaries and availability statements are often required for original research
Human-like writing research incorporated into the skill
Comparative stylistics work
Recent comparative work on LLM and human writing reports that instruction-tuned models often overuse:
- present-participial clauses
- nominalizations
- noun-heavy informational density
The study explicitly argues these are useful revision cues rather than detector rules. This is why v2 uses them for diagnosis, not for "beating detectors".
Biomedical abstract vocabulary analyses
Large-scale analyses of recent biomedical abstracts suggest that LLM-affected prose often shows excess style words and a recognisable shift in lexical texture. This motivated the stronger anti-hype and anti-generic-language layers.
Reader-guidance principles
Writing guidance from the Nature / Springer ecosystem emphasises that readers take cues not only from linking words but also from where information is placed in a sentence. This motivated the sentence-craft emphasis on topic position, stress position, and old-to-new information flow.
Why this matters for skill design
The main failure mode of many "humanizer" prompts is that they push prose toward generic warmth or casualness. That is the wrong target for scientific manuscripts.
For Nature-style papers, the relevant improvements are:
- better structural architecture
- stronger paragraph jobs
- better sentence guidance
- less adjective-led importance
- fewer generic transitions
- more explicit limitation handling
- better alignment with the user's own exemplar prose
That is why v2 adds:
- journal calibration
- exemplar anchoring
- editorial blueprint templates
- a prose fingerprinting script
- stronger preflight checks
Section rubric
Use this file when drafting or repairing specific sections. For each section, first ask what job it must perform. Then repair whichever part of the section prevents it from doing that job.
Title
Job
Make the paper legible, searchable, and credible in one line.
It succeeds when
- an adjacent-field scientist can infer what the paper is about
- the title names the right entity, process, system, or intervention
- the claim is neither mushy nor overstated
Common failure modes
- too much mechanism packed into one line
- a generality with no identifiable subject
- promotional adjectives doing the work of evidence
- excessive punctuation or a two-part slogan structure
Repair moves
- cut the weakest noun phrase
- replace evaluative adjectives with the actual object or process
- move from "what this means" back to "what this paper shows"
Summary paragraph or abstract
Job
Orient a broad reader, define the gap, state the principal answer, and end on the nearest real implication.
It succeeds when
- the first lines give a reader enough footing to care
- the gap is specific, not just "little is known"
- the main finding is stated once, plainly
- the last sentence names an implication that truly follows
Common failure modes
- starts too narrowly
- sounds like a compressed Results section
- too many numbers or acronyms
- final sentence leaps into unsupported future claims
Repair moves
- widen the first two sentences
- name the specific unresolved issue
- strip out non-load-bearing numeric detail
- replace general importance claims with the exact consequence of the result
Introduction or opening
Job
Move from field context to the paper's problem and promise without becoming a literature dump.
It succeeds when
- readers know why the problem matters
- the exact unresolved issue is named
- the paper's response feels necessary and proportionate
Common failure modes
- mini-review sprawl
- a gap framed so broadly that any paper could fill it
- final paragraph that merely repeats the abstract
- transition sentences that say "however" without sharpening the problem
Repair moves
- cut citations and background that do not set up the problem
- replace generic "however" moves with the actual unresolved point
- end with what the paper does and why that design can answer the question
Results
Job
Carry the reader through the evidence in the order needed to support the paper's central claim.
It succeeds when
- each subsection answers a question
- each paragraph earns its place in the argument
- interpretation is separated from observation
- the reader always knows why the next figure matters
Common failure modes
- chronological lab notebook structure
- paragraphs that narrate procedures rather than findings
- every sentence starts from "we"
- excessive number dumping
- figure captions transplanted into prose
Repair moves
- reorganise by conceptual dependency
- start subsections with the question they answer
- keep only the numbers that alter interpretation
- end paragraphs on the finding that enables the next step
Discussion
Job
State what the work establishes, what it suggests, and where its scope ends.
It succeeds when
- the main conclusion is re-stated with better perspective, not just repeated
- competing explanations or limitations are acknowledged
- implications are scaled to the evidence
- the ending leaves the reader with a precise field-level consequence
Common failure modes
- generic "these findings highlight the importance of" ending
- sudden new literature review
- limit-free victory lap
- future-work paragraph with no real synthesis
Repair moves
- identify the strongest limitation and address it directly
- replace generic significance language with specific consequences
- separate what is proved, supported, and still unresolved
Methods
Job
Give the reader the details that govern interpretability and reproducibility.
It succeeds when
- the reader can see what was measured, how, on what, and under what rules
- statistical and preprocessing choices are explicit
- software, pipelines, exclusion rules, and sample definitions are clear where material
Common failure modes
- beautiful prose but missing operational detail
- result-critical decisions buried in supplement-only material
- vague "standard methods were used" wording
- statistical analysis described only in slogans
Repair moves
- add the information a sceptical reader would need to judge the claim
- separate acquisition, preprocessing, modelling, and statistics
- name software versions and parameters when they affect interpretation
Figure legends
Job
Let the reader understand the figure without searching the main text for basic decoding.
It succeeds when
- the legend starts with a short title sentence
- panels are described in order
- sample size, statistics, centre values, and error bars are defined where relevant
- specialised symbols are explained
Common failure modes
- legend begins directly with panel letters and no framing
- no statistical definitions
- half the information is hidden in the Methods
- verbose retelling of the main text
Repair moves
- add a title sentence naming the figure's point
- define symbols and statistics explicitly
- keep only the minimal methodological detail needed to read the figure correctly
Data Availability
Job
Tell readers where the underlying data are and how they can obtain them.
It succeeds when
- a repository, accession, or access condition is named
- restrictions are explained, not obscured
Common failure modes
- vague "available upon request"
- statement absent entirely
- multiple data types mixed together without clarity
Repair moves
- separate public deposition from controlled access
- say what is already deposited and what will be released on acceptance if that is the case
- use placeholders rather than making anything up
Code Availability
Job
Tell readers whether custom code exists, where it is, and under what conditions it can be accessed.
It succeeds when
- custom analysis or modelling code is mentioned when central
- repository or access condition is stated
- versioning or release timing is clear where relevant
Cover letter or presubmission enquiry
Job
Help an editor see fit, importance, and evidential strength quickly.
It succeeds when
- it reads like an editor memo, not a reheated abstract
- it explains why the paper belongs in this journal rather than in science generally
- it names the main evidential reasons to trust the claim
Common failure modes
- abstract pasted into letter form
- overclaiming without editorial fit logic
- no mention of scope or limitations
Repair moves
- write for a time-poor editor
- make the fit argument explicit
- state what the paper is not claiming
Sentence craft
Scientific prose becomes easier and more pleasurable to read when sentence structure does more of the guidance work. Do not rely only on transition words.
Topic position and stress position
Readers take cues from where information appears in a sentence.
Topic position
The start of the sentence should connect to what the reader is already tracking.
Useful starts:
- the object of study
- the experimental condition
- the specific comparison
- the previously established result that the new sentence extends
Weak starts:
- empty scene-setting (
Importantly,Notably,In this context) - vague pronouns (
This,These,It) when the referent is not obvious - long subordinate clauses before the reader knows the subject
Stress position
The end of the sentence naturally carries emphasis. Put the sentence's point there when possible.
Weak:
We compared treated and control samples, which were processed using our standard pipeline.- The sentence lands on logistics.
Stronger:
Using the same processing pipeline for both groups, we found that treatment selectively expanded the resistant subpopulation.- The sentence lands on the actual finding.
Old-to-new flow
Whenever possible, start with known information and end with the new contribution. This reduces re-reading.
Weak:
A previously unrecognised state transition emerged in cells exposed to stress, following our high-frame-rate imaging analysis.- The new analytic method arrives late and blurs the reader's route.
Stronger:
Using high-frame-rate imaging, we detected a previously unrecognised state transition in cells exposed to stress.
Prefer verbs over nominalizations
Noun-heavy prose is one of the fastest ways to make a manuscript sound machine-written or bureaucratic.
Weak:
This observation provides support for the interpretation that...The regulation of differentiation was mediated by...
Stronger:
This observation suggests that...X regulated differentiation by...
Do not chase this mechanically. Some noun phrases are the correct technical labels. The goal is to remove unnecessary hidden verbs.
Break present-participial clause chains
Instruction-tuned models often overuse -ing clause chains.
Weak:
Using live imaging, revealing a previously hidden transition, and suggesting a mechanism for resistance, we...
Stronger:
- split into 2-3 sentences
- give the main clause a clear verb
- keep the
-ingform only when it truly improves flow
Active and passive voice
Prefer active voice when it clarifies agency.
We measured,we trained,we compared,we estimated
Use passive voice when agency is irrelevant or the object deserves sentence-initial focus.
Samples were fixed in 4% paraformaldehydeCells were imaged every 5 min
The problem is not passive voice itself. The problem is habitual passivity that hides who did what or drains energy from claim-bearing sentences.
Tense discipline
A workable default:
- past tense for what you did and observed
- present tense for what figures show in the paper now, and for general truths
- cautious present or modal language for claims that extend beyond direct observation
Do not let tense drift randomly.
Manage acronyms and symbol load
- define only the acronyms that truly save space or prevent repetition
- avoid stacking multiple undefined abbreviations in a single sentence
- if the paper needs heavy symbol use, slow down the first appearance and orient the reader explicitly
Rhythm and variation
Human scientific prose usually mixes:
- short claim sentences
- medium explanatory sentences
- occasional longer synthesis sentences
A draft where every sentence has nearly the same length feels automated even when each sentence is individually fine.
Good variation does not mean ornamental style. It means the sentence length matches the job.
Punctuation
- use commas to reduce ambiguity, not to create faux complexity
- use semicolons sparingly and only when they genuinely compress related material
- use em dashes rarely; overuse reads mannered
- colons work well when a sentence sets up a list, contrast, or consequence
Line-edit questions
When a sentence feels wrong, ask: 1. What is this sentence trying to do? 2. Does the subject appear early enough? 3. Does the main verb do real work? 4. Does the sentence land on the point that matters? 5. Can one clause become its own sentence? 6. Is the key noun hidden inside an abstract phrase?
Quick repair patterns
Replace empty significance language
These findings highlight the importance of XThese results show that X constrains Y under Z conditions
Replace throat-clearing
It is important to note that- delete it, or say the actual point
Replace vague demonstratives
This suggestsThis increase in chromatin accessibility suggests
Shorten stacked modifiers
high-resolution time-resolved single-cell transcriptional profiling framework- distribute the information across the sentence instead of forcing it into one noun pile
Voice and variation
This file adapts anti-slop and "humanizer" principles to scientific prose. The goal is not to make papers casual. The goal is to remove the machine-like habits that make scientific text feel generic, over-smoothed, or oddly promotional.
What "human" usually means here
For research manuscripts, human-sounding prose usually has:
- a specific claim at the centre
- proportion between evidence and language
- local sentence variety
- paragraph endings that actually do argumentative work
- selective signposting rather than transition spam
- some friction and texture where the thinking is difficult, instead of frictionless generic phrasing
Common machine-like habits to cut
1. Adjective-led importance
Weak:
novel,groundbreaking,remarkable,critical,compelling,transformative
Repair:
- replace the adjective with the thing the evidence actually shows
2. Generic significance endings
Weak:
These findings highlight the importance of ...Together, our results open new avenues for ...
Repair:
- say what follows from the result, under what conditions, and where the uncertainty remains
3. Conveyor-belt signposting
Weak:
AdditionallyMoreoverImportantlyNotablyTaken togetherOverall
Repair:
- use sentence structure and paragraph order to carry the logic
- keep only the few transitions that genuinely prevent misreading
4. Formulaic contrast frames
Weak:
not only ... but alsothis is not merely ...rather than simply ...
Repair:
- state the actual contrast directly
5. Abstract noun haze
Weak:
landscape,interplay,framework,paradigm,hallmark,cornerstone,suite,robust approach
Repair:
- name the system, variable, assay, model, or process
6. Over-smoothed sentence openings
Watch for high concentrations of sentence starts with:
ThisTheseItWe- or the same transition adverb
Repair by varying the opening around:
- the object being measured
- the contrast being made
- the figure or condition under discussion
- the actual new fact
7. Repetitive paragraph shapes
A suspicious draft often has paragraph after paragraph that:
- starts with a generic topic sentence
- gives a blob of middle detail
- ends with a vague importance line
Repair:
- give each paragraph one sharper job
- end on the exact point that enables the next move
What not to overcorrect
Passive voice
Do not purge all passive voice. Methods often need it, and some result sentences benefit from object-first framing.
Repetition
Technical prose sometimes needs repeated terms for clarity. Replacing every repeated noun with a synonym can make the manuscript worse.
Caution
Some fields require careful hedging. Do not "humanize" by stripping out all restraint.
Technical density
Some papers really are dense. The aim is not simplification at all costs, but guided density.
Productive revision passes
Pass 1: remove generic importance language
Delete or replace every sentence that claims importance without specifying a consequence.
Pass 2: reduce hidden verbs
Turn abstract noun phrases back into actions where possible.
Pass 3: diversify sentence openings
Check the first word or first 2-3 words of consecutive sentences. If many repeat, vary them.
Pass 4: restore stress position
Look at sentence endings. Too many weak endings on procedural detail, acronyms, or empty nouns make the prose sag.
Pass 5: preserve the paper's intellectual temperature
The manuscript should sound calm and exact. Do not inject personality; inject precision.
Dangerous rewrites
Avoid these common overcorrections:
- making the paper sound like a press release
- replacing all technical terms with vague simpler ones
- adding literary metaphors
- removing caveats that are scientifically necessary
- forcing "variety" by using synonyms for core technical entities
A good test
After revision, ask:
- Could this sentence appear in almost any paper?
- Is the claim supported right here or in the next sentence?
- Does the paragraph feel inevitable rather than templated?
- Does the conclusion sound earned rather than staged?
If the answer is no, revise again.
\
#!/usr/bin/env python3
"""
Nature / Nature Portfolio manuscript preflight checker.
This script performs structural, stylistic, and policy-aware checks on a draft.
It is lightweight, stdlib-only, and intended for non-interactive agent use.
Examples:
python3 scripts/nature_preflight.py --input draft.md --mode nature-article --format text
python3 scripts/nature_preflight.py --input draft.md --mode portfolio-article --format json
cat draft.md | python3 scripts/nature_preflight.py --mode nature-letter
Exit codes:
0 success
2 usage error or unreadable input
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from dataclasses import dataclass, asdict
from pathlib import Path
from typing import Dict, List, Tuple
from prose_metrics import (
as_json,
metrics as prose_metrics,
read_text,
strip_frontmatter,
sentences,
words,
)
SECTION_PATTERNS = {
"abstract": re.compile(r"^\s{0,3}(?:#+\s*)?abstract\s*$", re.I),
"summary_paragraph": re.compile(r"^\s{0,3}(?:#+\s*)?(?:summary paragraph|summary)\s*$", re.I),
"introductory_paragraph": re.compile(r"^\s{0,3}(?:#+\s*)?introductory paragraph\s*$", re.I),
"introduction": re.compile(r"^\s{0,3}(?:#+\s*)?introduction\s*$", re.I),
"results": re.compile(r"^\s{0,3}(?:#+\s*)?results\s*$", re.I),
"discussion": re.compile(r"^\s{0,3}(?:#+\s*)?discussion\s*$", re.I),
"methods": re.compile(r"^\s{0,3}(?:#+\s*)?(?:online methods|methods)\s*$", re.I),
"data_availability": re.compile(r"^\s{0,3}(?:#+\s*)?data availability\s*$", re.I),
"code_availability": re.compile(r"^\s{0,3}(?:#+\s*)?code availability\s*$", re.I),
"references": re.compile(r"^\s{0,3}(?:#+\s*)?references\s*$", re.I),
"acknowledgements": re.compile(r"^\s{0,3}(?:#+\s*)?acknowledg(?:e)?ments\s*$", re.I),
"funding_statement": re.compile(r"^\s{0,3}(?:#+\s*)?funding(?: statement)?\s*$", re.I),
"author_contributions": re.compile(r"^\s{0,3}(?:#+\s*)?author contributions\s*$", re.I),
"competing_interests": re.compile(r"^\s{0,3}(?:#+\s*)?(?:competing interests|conflict of interest|conflicts of interest)\s*$", re.I),
"additional_information": re.compile(r"^\s{0,3}(?:#+\s*)?additional information\s*$", re.I),
"figure_legends": re.compile(r"^\s{0,3}(?:#+\s*)?figure legends\s*$", re.I),
"extended_data_legends": re.compile(r"^\s{0,3}(?:#+\s*)?extended data(?: figure| table)? legends\s*$", re.I),
}
MODE_REQUIREMENTS = {
"nature-article": {
"sections_required": ["methods", "data_availability", "references", "figure_legends"],
"sections_recommended": ["code_availability", "funding_statement", "author_contributions", "competing_interests"],
"opening_type": "summary",
"opening_target_words_max": 220,
"title_chars_max": 75,
"main_headings_allowed": True,
},
"nature-letter": {
"sections_required": ["methods", "data_availability", "references", "figure_legends"],
"sections_recommended": ["code_availability", "funding_statement", "author_contributions", "competing_interests"],
"opening_type": "introductory",
"opening_target_words_max": 220,
"title_chars_max": 85,
"main_headings_allowed": False,
},
"portfolio-article": {
"sections_required": ["methods", "data_availability", "references"],
"sections_recommended": ["results", "discussion", "code_availability", "funding_statement", "author_contributions", "competing_interests", "figure_legends"],
"opening_type": "abstract",
"opening_target_words_max": 250,
"title_chars_max": 120,
"main_headings_allowed": True,
},
"portfolio-letter": {
"sections_required": ["methods", "data_availability", "references"],
"sections_recommended": ["code_availability", "funding_statement", "author_contributions", "competing_interests", "figure_legends"],
"opening_type": "introductory",
"opening_target_words_max": 220,
"title_chars_max": 120,
"main_headings_allowed": False,
},
}
TITLE_BAD_WORDS = {"novel", "groundbreaking", "transformative", "remarkable", "unprecedented"}
TITLE_ACRONYM_RE = re.compile(r"\b[A-Z][A-Z0-9-]{1,}\b")
TITLE_PUNCT_RE = re.compile(r"[:;!?]")
BRACKET_CITATION_RE = re.compile(r"\[(?:\d+(?:\s*[-,]\s*\d+)*)\]")
FIGURE_HEADING_RE = re.compile(r"^\s{0,3}(?:#+\s*)?(?:figure|fig\.)\s*\d+\b", re.I)
STATS_HINT_RE = re.compile(r"\b(?:n\s*=|P\s*[<=>]|Student'?s t-test|Mann-Whitney|ANOVA|Wilcoxon|Kruskal|error bars|s\.d\.|s\.e\.m\.|median|mean)\b", re.I)
LEGEND_TITLE_SENTENCE_RE = re.compile(r"^[A-Z].{10,200}[.!?]$")
@dataclass
class Issue:
severity: str
code: str
message: str
evidence: str = ""
fix: str = ""
def build_parser() -> argparse.ArgumentParser:
p = argparse.ArgumentParser(
description="Run a structural and stylistic preflight check on a Nature-style manuscript draft."
)
p.add_argument("--input", help="Draft file to analyse. If omitted, read from stdin.")
p.add_argument("--mode", choices=sorted(MODE_REQUIREMENTS), required=True, help="Target manuscript mode.")
p.add_argument("--format", choices=["text", "json"], default="text", help="Output format.")
return p
def load_text(path: str | None) -> str:
if path:
try:
return read_text(path)
except FileNotFoundError:
raise SystemExit(f"Error: input file not found: {path}")
except OSError as exc:
raise SystemExit(f"Error: could not read {path}: {exc}")
data = sys.stdin.read()
if not data.strip():
raise SystemExit("Error: no input provided. Use --input FILE or pipe text via stdin.")
return data
def nonempty_lines(text: str) -> List[str]:
return [line for line in text.splitlines() if line.strip()]
def normalize_heading(line: str) -> str:
return re.sub(r"^\s*#+\s*", "", line).strip()
def first_title_line(text: str) -> str:
for line in text.splitlines():
s = line.strip()
if not s:
continue
if s.startswith("#"):
return normalize_heading(s)
return s
return ""
def paragraphs(text: str) -> List[str]:
return [p.strip() for p in re.split(r"\n\s*\n", text) if p.strip()]
def detect_sections(text: str) -> Dict[str, Tuple[int, int]]:
lines = text.splitlines()
hits: List[Tuple[str, int]] = []
for idx, line in enumerate(lines):
for name, pat in SECTION_PATTERNS.items():
if pat.match(line):
hits.append((name, idx))
break
ranges: Dict[str, Tuple[int, int]] = {}
for i, (name, start) in enumerate(hits):
end = len(lines) if i + 1 >= len(hits) else hits[i + 1][1]
ranges[name] = (start, end)
return ranges
def section_text(text: str, ranges: Dict[str, Tuple[int, int]], name: str) -> str:
if name not in ranges:
return ""
start, end = ranges[name]
lines = text.splitlines()
return "\n".join(lines[start + 1:end]).strip()
def get_opening_text(text: str, mode: str, sections: Dict[str, Tuple[int, int]]) -> tuple[str, str]:
if mode == "nature-article" and "summary_paragraph" in sections:
return "summary_paragraph", section_text(text, sections, "summary_paragraph")
if mode == "nature-letter" and "introductory_paragraph" in sections:
return "introductory_paragraph", section_text(text, sections, "introductory_paragraph")
if mode == "portfolio-article" and "abstract" in sections:
return "abstract", section_text(text, sections, "abstract")
if mode == "portfolio-letter" and "introductory_paragraph" in sections:
return "introductory_paragraph", section_text(text, sections, "introductory_paragraph")
title = first_title_line(text)
paras = [p for p in paragraphs(text) if p != title and len(words(p)) >= 6]
return "first_paragraph", paras[0] if paras else ""
def contains_reference_like_markers(text: str) -> bool:
return bool(BRACKET_CITATION_RE.search(text))
def title_checks(title: str, mode: str) -> List[Issue]:
issues: List[Issue] = []
limit = MODE_REQUIREMENTS[mode]["title_chars_max"]
if not title:
issues.append(Issue("error", "title_missing", "No title line detected.", fix="Add a manuscript title at the start of the document."))
return issues
if len(title) > limit:
issues.append(Issue("warning", "title_too_long", f"Title length is {len(title)} characters; target is <= {limit} for this mode.", evidence=title, fix="Shorten the title and remove non-essential modifiers."))
lower_words = {w.lower() for w in words(title)}
bad = sorted(lower_words & TITLE_BAD_WORDS)
if bad:
issues.append(Issue("warning", "title_hype", "Title contains evaluative or hype language.", evidence=", ".join(bad), fix="Replace evaluative adjectives with the actual object, process, or finding."))
if mode.startswith("nature"):
if TITLE_ACRONYM_RE.search(title):
issues.append(Issue("warning", "title_acronym", "Nature-style titles usually avoid acronyms unless essential.", evidence=title, fix="Spell out the term or remove the acronym if possible."))
if TITLE_PUNCT_RE.search(title):
issues.append(Issue("warning", "title_punctuation", "Nature-style titles usually avoid decorative punctuation such as colons or question marks.", evidence=title, fix="Simplify the title into a single clear clause if possible."))
return issues
def section_checks(text: str, mode: str, sections: Dict[str, Tuple[int, int]]) -> List[Issue]:
issues: List[Issue] = []
reqs = MODE_REQUIREMENTS[mode]
for sec in reqs["sections_required"]:
if sec not in sections:
issues.append(Issue("error", f"missing_{sec}", f"Missing required section: {sec.replace('_', ' ').title()}.", fix=f"Add a {sec.replace('_', ' ').title()} section or mark it as pending."))
for sec in reqs["sections_recommended"]:
if sec not in sections:
issues.append(Issue("note", f"missing_{sec}", f"Recommended section not found: {sec.replace('_', ' ').title()}.", fix=f"Confirm whether a {sec.replace('_', ' ').title()} section is required by the target journal."))
if not reqs["main_headings_allowed"]:
for forbidden in ("results", "discussion", "introduction"):
if forbidden in sections:
issues.append(Issue("warning", f"heading_policy_{forbidden}", f"{mode} usually reads as a continuous narrative; explicit '{forbidden.title()}' heading detected.", fix="Check whether the target format discourages main-text headings."))
return issues
def opening_checks(text: str, mode: str, sections: Dict[str, Tuple[int, int]]) -> List[Issue]:
issues: List[Issue] = []
label, opening = get_opening_text(text, mode, sections)
expected_heading = {
"nature-article": "summary_paragraph",
"nature-letter": "introductory_paragraph",
"portfolio-article": "abstract",
"portfolio-letter": "introductory_paragraph",
}[mode]
if expected_heading not in sections:
issues.append(Issue("warning", "opening_heading_missing", f"Expected opening section not found for {mode}: {expected_heading.replace('_', ' ').title()}.", evidence=label, fix="Add the expected opening section heading or confirm that the fallback first paragraph is acceptable for the target journal."))
if not opening:
issues.append(Issue("error", "opening_missing", "Could not find an opening summary/abstract/introduction paragraph.", fix="Add the required opening section for the chosen mode."))
return issues
n_words = len(words(opening))
target_max = MODE_REQUIREMENTS[mode]["opening_target_words_max"]
if n_words > target_max:
issues.append(Issue("warning", "opening_too_long", f"Opening section is {n_words} words; target is <= {target_max} for this mode.", evidence=label, fix="Cut low-value background, repeated claims, or non-essential numeric detail."))
if mode.startswith("nature") and not contains_reference_like_markers(opening):
issues.append(Issue("note", "opening_references", "Nature-style opening paragraph may need reference markers if this is the final manuscript form.", evidence=label, fix="Add references if the target journal/article type expects them."))
if mode == "nature-article":
if sum(1 for t in words(opening) if t.isupper() and len(t) > 1) > 2:
issues.append(Issue("note", "opening_acronym_load", "Opening summary paragraph contains several acronyms; Nature-style openings are usually lighter on abbreviations.", fix="Spell out or remove non-essential abbreviations in the opening paragraph."))
return issues
def availability_checks(text: str, sections: Dict[str, Tuple[int, int]]) -> List[Issue]:
issues: List[Issue] = []
if "data_availability" in sections:
dat = section_text(text, sections, "data_availability")
if "upon request" in dat.lower() and not any(x in dat.lower() for x in ("repository", "accession", "controlled", "available from")):
issues.append(Issue("warning", "data_vague", "Data Availability statement relies on vague 'upon request' wording.", evidence=dat[:180], fix="State repository, accession, or transparent access conditions if possible."))
if "code_availability" in sections:
code = section_text(text, sections, "code_availability")
if "upon request" in code.lower() and "github" not in code.lower() and "gitlab" not in code.lower() and "zenodo" not in code.lower():
issues.append(Issue("note", "code_vague", "Code Availability statement may be too vague if custom code is central.", evidence=code[:180], fix="State repository, DOI, or clear release conditions if possible."))
return issues
def figure_legend_checks(text: str, sections: Dict[str, Tuple[int, int]]) -> List[Issue]:
issues: List[Issue] = []
legends = section_text(text, sections, "figure_legends")
if not legends:
return issues
lines = [line.rstrip() for line in legends.splitlines()]
current_figure = None
current_block: List[str] = []
blocks: List[Tuple[str, str]] = []
for line in lines:
if FIGURE_HEADING_RE.match(line.strip()):
if current_figure is not None:
blocks.append((current_figure, "\n".join(current_block).strip()))
current_figure = normalize_heading(line)
current_block = []
else:
current_block.append(line)
if current_figure is not None:
blocks.append((current_figure, "\n".join(current_block).strip()))
for fig, block in blocks:
if not block:
issues.append(Issue("warning", "legend_empty", f"{fig} has no legend text.", fix="Add a title sentence and panel description."))
continue
first_line = next((ln.strip() for ln in block.splitlines() if ln.strip()), "")
if first_line and not LEGEND_TITLE_SENTENCE_RE.match(first_line):
issues.append(Issue("note", "legend_title_sentence", f"{fig} does not obviously begin with a brief title sentence.", evidence=first_line[:180], fix="Start the legend with one sentence that names the figure's point."))
if not STATS_HINT_RE.search(block):
issues.append(Issue("note", "legend_stats", f"{fig} legend contains no obvious statistics or sample-size cues.", evidence=fig, fix="Check whether n, error bars, centre values, or statistical tests need to be defined."))
return issues
def citation_checks(text: str) -> List[Issue]:
issues: List[Issue] = []
hits = BRACKET_CITATION_RE.findall(text)
if hits:
issues.append(Issue("note", "bracket_citations", "Bracket-style citations detected.", evidence=", ".join(hits[:8]), fix="Check whether the target journal wants superscript or another citation format."))
return issues
def style_checks(text: str) -> List[Issue]:
issues: List[Issue] = []
m = prose_metrics(text)
if m["hype_word_total"] > 0:
examples = ", ".join(f"{k}({v})" for k, v in list(m["hype_words"].items())[:8])
issues.append(Issue("warning", "hype_words", "Evaluative or hype words detected.", evidence=examples, fix="Replace adjective-led importance language with concrete claims."))
if m["generic_phrase_total"] > 0:
examples = ", ".join(f"{k}({v})" for k, v in list(m["generic_phrases"].items())[:8])
issues.append(Issue("warning", "generic_phrases", "Generic AI-ish manuscript phrases detected.", evidence=examples, fix="Replace generic phrases with the exact implication or delete them."))
if m["nominalization_rate"] > 0.045:
issues.append(Issue("note", "nominalization_rate", f"Nominalization rate is high ({m['nominalization_rate']}).", fix="Turn hidden verbs back into actions where possible."))
if m["participial_clause_rate"] > 0.18:
issues.append(Issue("note", "participial_rate", f"Many sentences contain '-ing' clause cues ({m['participial_clause_rate']}).", fix="Break multi-clause sentences into clearer main clauses."))
if m["transition_opener_rate"] > 0.08:
issues.append(Issue("note", "transition_rate", f"Transition-opener rate is high ({m['transition_opener_rate']}).", fix="Cut conveyor-belt transitions and let structure carry more of the logic."))
if m["weak_opener_rate"] > 0.45:
issues.append(Issue("note", "weak_openers", f"Many sentences begin with weak openers ({m['weak_opener_rate']}).", evidence=", ".join(f"{k}:{v}" for k, v in list(m["top_sentence_openers"].items())[:5]), fix="Vary sentence openings around the object, comparison, or condition rather than repeating pronouns."))
if m["flatness_score"] > 0.18 and m["sentence_count"] >= 8:
issues.append(Issue("note", "rhythm_flat", f"Sentence-length variation appears low (flatness score {m['flatness_score']}).", fix="Mix shorter claim sentences with longer explanatory ones."))
return issues
def ending_checks(text: str) -> List[Issue]:
issues: List[Issue] = []
sents = sentences(text)
if not sents:
return issues
last = sents[-1].lower()
cliches = [
"opens new avenues",
"future work",
"highlight the importance of",
"underscores the importance of",
"will be needed to fully understand",
]
for phrase in cliches:
if phrase in last:
issues.append(Issue("note", "generic_ending", "Final sentence ends on a generic future-work or importance phrase.", evidence=sents[-1][:220], fix="End on the most defensible implication or boundary condition instead."))
break
return issues
def prioritise(issues: List[Issue]) -> List[Issue]:
order = {"error": 0, "warning": 1, "note": 2}
return sorted(issues, key=lambda x: (order.get(x.severity, 9), x.code, x.message))
def render_text(mode: str, title: str, opening_label: str, metrics: dict, issues: List[Issue]) -> str:
lines = []
lines.append("Nature-style preflight report")
lines.append("============================")
lines.append(f"Mode: {mode}")
lines.append(f"Title: {title or '[missing]'}")
lines.append("")
lines.append("Document profile")
lines.append("----------------")
lines.append(f"- words: {metrics['word_count']}")
lines.append(f"- sentences: {metrics['sentence_count']}")
lines.append(f"- paragraphs: {metrics['paragraph_count']}")
lines.append(f"- mean sentence length: {metrics['mean_sentence_words']} words")
lines.append(f"- sentence-length variation: {metrics['stdev_sentence_words']}")
lines.append(f"- nominalization rate: {metrics['nominalization_rate']}")
lines.append(f"- participial-clause rate: {metrics['participial_clause_rate']}")
lines.append(f"- transition-opener rate: {metrics['transition_opener_rate']}")
lines.append(f"- weak-opener rate: {metrics['weak_opener_rate']}")
lines.append(f"- opening source: {opening_label}")
lines.append("")
lines.append("Prioritised issues")
lines.append("------------------")
if not issues:
lines.append("- No major issues detected by the bundled heuristics. Review journal-specific details manually.")
else:
for issue in issues:
line = f"- [{issue.severity.upper()}] {issue.message}"
lines.append(line)
if issue.evidence:
lines.append(f" evidence: {issue.evidence}")
if issue.fix:
lines.append(f" fix: {issue.fix}")
lines.append("")
lines.append("Interpretation")
lines.append("--------------")
lines.append("Use this report to fix structure first, then claim calibration, then line-level prose. These heuristics are advisory; they are meant to surface likely weak spots, not to replace scientific judgement.")
return "\n".join(lines)
def main() -> int:
parser = build_parser()
args = parser.parse_args()
raw = load_text(args.input)
text = strip_frontmatter(raw)
title = first_title_line(text)
sections = detect_sections(text)
opening_label, _ = get_opening_text(text, args.mode, sections)
m = prose_metrics(text)
issues: List[Issue] = []
issues.extend(title_checks(title, args.mode))
issues.extend(section_checks(text, args.mode, sections))
issues.extend(opening_checks(text, args.mode, sections))
issues.extend(availability_checks(text, sections))
issues.extend(figure_legend_checks(text, sections))
issues.extend(citation_checks(text))
issues.extend(style_checks(text))
issues.extend(ending_checks(text))
issues = prioritise(issues)
payload = {
"mode": args.mode,
"title": title,
"opening_source": opening_label,
"metrics": m,
"issues": [asdict(i) for i in issues],
"summary": {
"errors": sum(1 for i in issues if i.severity == "error"),
"warnings": sum(1 for i in issues if i.severity == "warning"),
"notes": sum(1 for i in issues if i.severity == "note"),
"total": len(issues),
},
}
if args.format == "json":
print(as_json(payload))
else:
print(render_text(args.mode, title, opening_label, m, issues))
return 0
if __name__ == "__main__":
raise SystemExit(main())
\
#!/usr/bin/env python3
"""
Compare a candidate manuscript against one or more reference texts and report
stylistic alignment cues. This is intended to help an agent match broad habits
of sentence movement and paragraph density without copying phrasing.
Examples:
python3 scripts/prose_fingerprint.py --candidate draft.md --reference paper1.md paper2.md
python3 scripts/prose_fingerprint.py --input draft.md exemplar.md --format json
Exit codes:
0 success
2 usage error or unreadable file
"""
from __future__ import annotations
import argparse
import sys
from pathlib import Path
from typing import List
from prose_metrics import (
aggregate,
as_json,
compare,
metrics,
rank_revision_priorities,
read_text,
)
def build_parser() -> argparse.ArgumentParser:
p = argparse.ArgumentParser(
description="Compare a candidate manuscript against reference texts and report high-level prose alignment cues."
)
p.add_argument(
"--candidate",
help="Candidate draft to analyse."
)
p.add_argument(
"--reference",
nargs="+",
help="One or more reference texts to compare against."
)
p.add_argument(
"--input",
nargs="+",
help="Alternative interface: pass multiple files, with the first treated as candidate and the rest as references."
)
p.add_argument(
"--format",
choices=["text", "json"],
default="text",
help="Output format."
)
return p
def resolve_args(args: argparse.Namespace) -> tuple[str, List[str]]:
if args.input:
if len(args.input) < 2:
raise SystemExit("Error: --input requires at least two files: candidate followed by one or more references.")
return args.input[0], args.input[1:]
if not args.candidate or not args.reference:
raise SystemExit("Error: provide either --candidate with --reference, or use --input.")
return args.candidate, list(args.reference)
def load(path: str) -> str:
try:
return read_text(path)
except FileNotFoundError:
raise SystemExit(f"Error: file not found: {path}")
except OSError as exc:
raise SystemExit(f"Error: could not read {path}: {exc}")
def text_report(candidate_path: str, candidate_metrics: dict, reference_paths: List[str], reference_metrics: dict, deltas: list, priorities: list) -> str:
lines = []
lines.append("Prose fingerprint report")
lines.append("======================")
lines.append(f"Candidate: {candidate_path}")
lines.append(f"References ({len(reference_paths)}): " + ", ".join(reference_paths))
lines.append("")
lines.append("Candidate profile")
lines.append("-----------------")
lines.append(f"- words: {candidate_metrics['word_count']}")
lines.append(f"- sentences: {candidate_metrics['sentence_count']}")
lines.append(f"- mean sentence length: {candidate_metrics['mean_sentence_words']} words")
lines.append(f"- sentence-length variation: {candidate_metrics['stdev_sentence_words']}")
lines.append(f"- mean paragraph length: {candidate_metrics['mean_paragraph_words']} words")
lines.append(f"- nominalization rate: {candidate_metrics['nominalization_rate']}")
lines.append(f"- participial-clause rate: {candidate_metrics['participial_clause_rate']}")
lines.append(f"- transition-opener rate: {candidate_metrics['transition_opener_rate']}")
lines.append(f"- weak-opener rate: {candidate_metrics['weak_opener_rate']}")
lines.append(f"- acronym rate: {candidate_metrics['acronym_rate']}")
lines.append(f"- rhythm flatness score: {candidate_metrics['flatness_score']}")
lines.append("")
lines.append("Reference aggregate")
lines.append("-------------------")
lines.append(f"- reference count: {reference_metrics.get('reference_count', len(reference_paths))}")
lines.append(f"- mean sentence length: {reference_metrics['mean_sentence_words']} words")
lines.append(f"- sentence-length variation: {reference_metrics['stdev_sentence_words']}")
lines.append(f"- mean paragraph length: {reference_metrics['mean_paragraph_words']} words")
lines.append(f"- nominalization rate: {reference_metrics['nominalization_rate']}")
lines.append(f"- participial-clause rate: {reference_metrics['participial_clause_rate']}")
lines.append(f"- transition-opener rate: {reference_metrics['transition_opener_rate']}")
lines.append(f"- weak-opener rate: {reference_metrics['weak_opener_rate']}")
lines.append(f"- acronym rate: {reference_metrics['acronym_rate']}")
lines.append(f"- rhythm flatness score: {reference_metrics['flatness_score']}")
lines.append("")
lines.append("Largest divergences")
lines.append("------------------")
for item in sorted(deltas, key=lambda x: abs(x["delta"]), reverse=True)[:8]:
if abs(item["delta"]) < 1e-9:
direction = "matches"
lines.append(f"- {item['label']}: candidate {direction} the references ({item['candidate']} vs {item['reference']})")
else:
direction = "higher" if item["delta"] > 0 else "lower"
lines.append(f"- {item['label']}: candidate {direction} than references ({item['candidate']} vs {item['reference']})")
lines.append("")
lines.append("Frequent candidate openers")
lines.append("--------------------------")
for k, v in candidate_metrics["top_sentence_openers"].items():
lines.append(f"- {k}: {v}")
lines.append("")
if candidate_metrics["generic_phrases"]:
lines.append("Generic phrases detected")
lines.append("------------------------")
for k, v in candidate_metrics["generic_phrases"].items():
lines.append(f"- {k}: {v}")
lines.append("")
if candidate_metrics["hype_words"]:
lines.append("Hype / evaluative words detected")
lines.append("-------------------------------")
for k, v in candidate_metrics["hype_words"].items():
lines.append(f"- {k}: {v}")
lines.append("")
lines.append("Revision priorities")
lines.append("-------------------")
if priorities:
for p in priorities:
lines.append(f"- {p}")
else:
lines.append("- Candidate broadly aligns with the reference set on the measured dimensions. Review paragraph jobs manually for finer stylistic fit.")
return "\n".join(lines)
def main() -> int:
parser = build_parser()
args = parser.parse_args()
candidate_path, reference_paths = resolve_args(args)
candidate_text = load(candidate_path)
reference_texts = [load(p) for p in reference_paths]
cand = metrics(candidate_text)
refs = [metrics(t) for t in reference_texts]
ref_agg = aggregate(refs)
deltas = compare(cand, ref_agg)
priorities = rank_revision_priorities(deltas)
payload = {
"candidate_path": candidate_path,
"reference_paths": reference_paths,
"candidate_metrics": cand,
"reference_metrics": ref_agg,
"comparisons": deltas,
"revision_priorities": priorities,
}
if args.format == "json":
print(as_json(payload))
else:
print(text_report(candidate_path, cand, reference_paths, ref_agg, deltas, priorities))
return 0
if __name__ == "__main__":
raise SystemExit(main())
\
#!/usr/bin/env python3
"""
Utility functions for lightweight manuscript-style diagnostics.
This module is stdlib-only and is intended for non-interactive use by
scripts/prose_fingerprint.py and scripts/nature_preflight.py.
"""
from __future__ import annotations
import json
import math
import re
import statistics
from collections import Counter
from pathlib import Path
from typing import Any, Dict, Iterable, List, Sequence
WORDS_RE = re.compile(r"\b[\w'-]+\b")
SENTENCE_SPLIT_RE = re.compile(r"(?<=[.!?])\s+")
PARA_SPLIT_RE = re.compile(r"\n\s*\n", re.M)
ACRONYM_RE = re.compile(r"\b[A-Z][A-Z0-9-]{1,}\b")
BE_VERB_RE = re.compile(r"\b(?:am|is|are|was|were|be|been|being)\b", re.I)
PAST_PARTICIPLE_RE = re.compile(r"\b\w+(?:ed|en)\b", re.I)
NOMINALIZATION_RE = re.compile(r"\b\w+(?:tion|sion|ment|ance|ence|ity|ness)\b", re.I)
TRANSITION_OPENERS = {
"additionally", "moreover", "furthermore", "importantly", "notably", "overall",
"therefore", "however", "thus", "consequently", "meanwhile", "taken", "together",
"in summary", "in conclusion"
}
WEAK_OPENERS = {"this", "these", "it", "we", "here"}
HYPE_WORDS = {
"novel", "groundbreaking", "transformative", "remarkable", "unprecedented",
"critical", "crucial", "important", "compelling", "robust", "exciting"
}
GENERIC_PHRASES = [
"highlights the importance of",
"underscores the importance of",
"plays a crucial role",
"taken together",
"in the broader context",
"opens new avenues",
"paves the way",
"it is important to note that",
"provides valuable insights",
"state-of-the-art",
"not only",
"not merely",
"in this landscape",
"interplay",
"leverages",
"fosters",
]
HEDGE_WORDS = {
"may", "might", "could", "suggest", "suggests", "suggested", "appear", "appears",
"likely", "possibly", "consistent", "indicate", "indicates", "associated"
}
def strip_frontmatter(text: str) -> str:
if text.startswith("---\n"):
parts = text.split("\n---\n", 1)
if len(parts) == 2:
return parts[1]
return text
def read_text(path: str | Path) -> str:
return Path(path).read_text(encoding="utf-8")
def paragraphs(text: str) -> List[str]:
return [p.strip() for p in PARA_SPLIT_RE.split(strip_frontmatter(text)) if p.strip()]
def sentences(text: str) -> List[str]:
clean = strip_frontmatter(text).strip()
if not clean:
return []
raw = [s.strip() for s in SENTENCE_SPLIT_RE.split(clean) if s.strip()]
return raw
def words(text: str) -> List[str]:
return WORDS_RE.findall(text)
def ratio(count: int, total: int) -> float:
return 0.0 if total <= 0 else count / total
def safe_mean(xs: Sequence[float]) -> float:
return float(statistics.mean(xs)) if xs else 0.0
def safe_median(xs: Sequence[float]) -> float:
return float(statistics.median(xs)) if xs else 0.0
def safe_stdev(xs: Sequence[float]) -> float:
if len(xs) <= 1:
return 0.0
return float(statistics.pstdev(xs))
def sentence_lengths(sents: Sequence[str]) -> List[int]:
return [len(words(s)) for s in sents if words(s)]
def paragraph_lengths(paras: Sequence[str]) -> List[int]:
return [len(words(p)) for p in paras if words(p)]
def first_word(sentence: str) -> str:
toks = words(sentence.lower())
return toks[0] if toks else ""
def first_two_words(sentence: str) -> str:
toks = words(sentence.lower())
return " ".join(toks[:2]) if toks else ""
def transition_opener_count(sents: Sequence[str]) -> int:
count = 0
for s in sents:
opener = first_two_words(s)
opener1 = first_word(s)
if opener in TRANSITION_OPENERS or opener1 in TRANSITION_OPENERS:
count += 1
return count
def weak_opener_count(sents: Sequence[str]) -> int:
return sum(1 for s in sents if first_word(s) in WEAK_OPENERS)
def repeated_openers(sents: Sequence[str], n: int = 5) -> Dict[str, int]:
c = Counter(first_word(s) for s in sents if first_word(s))
return dict(c.most_common(n))
def repeated_bigrams(sents: Sequence[str], n: int = 5) -> Dict[str, int]:
c = Counter(first_two_words(s) for s in sents if first_two_words(s))
return dict(c.most_common(n))
def passive_cues(sents: Sequence[str]) -> int:
count = 0
for s in sents:
if BE_VERB_RE.search(s) and PAST_PARTICIPLE_RE.search(s):
count += 1
return count
def participial_clause_cues(sents: Sequence[str]) -> int:
count = 0
pattern = re.compile(r"(?:^|,\s+)\w+(?:ing)\b", re.I)
for s in sents:
if pattern.search(s):
count += 1
return count
def nominalizations(tokens: Sequence[str]) -> int:
return sum(1 for t in tokens if NOMINALIZATION_RE.fullmatch(t))
def acronym_count(tokens: Sequence[str]) -> int:
return sum(1 for t in tokens if ACRONYM_RE.fullmatch(t))
def phrase_hits(text: str, phrases: Sequence[str]) -> Dict[str, int]:
lower = text.lower()
hits: Dict[str, int] = {}
for phrase in phrases:
n = lower.count(phrase.lower())
if n:
hits[phrase] = n
return hits
def word_hits(tokens: Sequence[str], vocab: Iterable[str]) -> Dict[str, int]:
vocab_set = {v.lower() for v in vocab}
c = Counter(t.lower() for t in tokens)
return {w: c[w] for w in sorted(vocab_set) if c[w]}
def endings(text: str) -> List[str]:
sents = sentences(text)
enders = []
for s in sents:
toks = words(s.lower())
if toks:
enders.append(toks[-1])
return enders
def flatness_score(lengths: Sequence[int]) -> float:
"""
Low variance in sentence lengths is one proxy for rhythm flatness.
Returns a score in [0, 1], where higher means flatter.
"""
if len(lengths) < 3:
return 0.0
mean = safe_mean(lengths)
stdev = safe_stdev(lengths)
if mean <= 0:
return 0.0
cv = stdev / mean
score = max(0.0, min(1.0, 0.5 - cv))
return round(score, 3)
def metrics(text: str) -> Dict[str, Any]:
clean = strip_frontmatter(text)
sents = sentences(clean)
paras = paragraphs(clean)
toks = words(clean)
slens = sentence_lengths(sents)
plens = paragraph_lengths(paras)
hype = word_hits(toks, HYPE_WORDS)
generic = phrase_hits(clean, GENERIC_PHRASES)
hedge = word_hits(toks, HEDGE_WORDS)
m: Dict[str, Any] = {
"word_count": len(toks),
"sentence_count": len(sents),
"paragraph_count": len(paras),
"mean_sentence_words": round(safe_mean(slens), 2),
"median_sentence_words": round(safe_median(slens), 2),
"stdev_sentence_words": round(safe_stdev(slens), 2),
"mean_paragraph_words": round(safe_mean(plens), 2),
"median_paragraph_words": round(safe_median(plens), 2),
"stdev_paragraph_words": round(safe_stdev(plens), 2),
"flatness_score": flatness_score(slens),
"acronym_count": acronym_count(toks),
"acronym_rate": round(ratio(acronym_count(toks), len(toks)), 4),
"nominalization_count": nominalizations(toks),
"nominalization_rate": round(ratio(nominalizations(toks), len(toks)), 4),
"participial_clause_cues": participial_clause_cues(sents),
"participial_clause_rate": round(ratio(participial_clause_cues(sents), len(sents)), 4),
"passive_cue_sentences": passive_cues(sents),
"passive_cue_rate": round(ratio(passive_cues(sents), len(sents)), 4),
"transition_opener_count": transition_opener_count(sents),
"transition_opener_rate": round(ratio(transition_opener_count(sents), len(sents)), 4),
"weak_opener_count": weak_opener_count(sents),
"weak_opener_rate": round(ratio(weak_opener_count(sents), len(sents)), 4),
"hype_words": hype,
"hype_word_total": sum(hype.values()),
"generic_phrases": generic,
"generic_phrase_total": sum(generic.values()),
"hedge_words": hedge,
"hedge_word_total": sum(hedge.values()),
"top_sentence_openers": repeated_openers(sents, n=8),
"top_sentence_opener_bigrams": repeated_bigrams(sents, n=8),
"ending_words_top": dict(Counter(endings(clean)).most_common(8)),
}
return m
def aggregate(metrics_list: Sequence[Dict[str, Any]]) -> Dict[str, Any]:
if not metrics_list:
return {}
numeric_keys = [
"word_count", "sentence_count", "paragraph_count",
"mean_sentence_words", "median_sentence_words", "stdev_sentence_words",
"mean_paragraph_words", "median_paragraph_words", "stdev_paragraph_words",
"flatness_score",
"acronym_count", "acronym_rate",
"nominalization_count", "nominalization_rate",
"participial_clause_cues", "participial_clause_rate",
"passive_cue_sentences", "passive_cue_rate",
"transition_opener_count", "transition_opener_rate",
"weak_opener_count", "weak_opener_rate",
"hype_word_total", "generic_phrase_total", "hedge_word_total",
]
out: Dict[str, Any] = {}
for key in numeric_keys:
values = [float(m.get(key, 0.0)) for m in metrics_list]
out[key] = round(safe_mean(values), 4)
# Merge counters by summed counts
counter_keys = ["top_sentence_openers", "top_sentence_opener_bigrams", "ending_words_top", "hype_words", "generic_phrases", "hedge_words"]
for key in counter_keys:
merged = Counter()
for m in metrics_list:
merged.update(m.get(key, {}))
out[key] = dict(merged.most_common(10))
out["reference_count"] = len(metrics_list)
return out
def compare(candidate: Dict[str, Any], reference: Dict[str, Any]) -> List[Dict[str, Any]]:
"""
Compare candidate metrics against aggregated reference metrics.
Returns a list of deltas with suggested interpretation.
"""
fields = [
("mean_sentence_words", "sentence length"),
("stdev_sentence_words", "sentence-length variation"),
("mean_paragraph_words", "paragraph length"),
("nominalization_rate", "nominalization rate"),
("participial_clause_rate", "participial-clause rate"),
("transition_opener_rate", "transition-opener rate"),
("weak_opener_rate", "weak-opener rate"),
("acronym_rate", "acronym rate"),
("flatness_score", "rhythm flatness score"),
("passive_cue_rate", "passive-cue rate"),
("hedge_word_total", "hedge-word density"),
("hype_word_total", "hype-word density"),
("generic_phrase_total", "generic-phrase density"),
]
results = []
sentence_count = max(1, int(candidate.get("sentence_count", 1)))
ref_sentence_count = max(1, int(reference.get("sentence_count", sentence_count)))
cand_word_count = max(1, int(candidate.get("word_count", 1)))
ref_word_count = max(1, int(reference.get("word_count", cand_word_count)))
for key, label in fields:
cand = float(candidate.get(key, 0.0))
ref = float(reference.get(key, 0.0))
# Normalise totals to rates when totals are used.
if key in {"hedge_word_total", "hype_word_total", "generic_phrase_total"}:
cand = cand / cand_word_count
ref = ref / ref_word_count
delta = cand - ref
results.append({
"metric": key,
"label": label,
"candidate": round(cand, 4),
"reference": round(ref, 4),
"delta": round(delta, 4),
})
return results
def rank_revision_priorities(comparisons: Sequence[Dict[str, Any]]) -> List[str]:
priorities: List[str] = []
for item in sorted(comparisons, key=lambda x: abs(x["delta"]), reverse=True):
label = item["label"]
d = item["delta"]
if label == "nominalization rate" and d > 0.01:
priorities.append("Reduce noun-heavy nominalizations; turn hidden verbs back into actions where possible.")
elif label == "participial-clause rate" and d > 0.05:
priorities.append("Break long '-ing' clause chains into clearer main clauses and shorter sentences.")
elif label == "transition-opener rate" and d > 0.03:
priorities.append("Cut conveyor-belt transitions and let structure carry more of the logic.")
elif label == "weak-opener rate" and d > 0.05:
priorities.append("Vary sentence openings; too many sentences start with weak pronouns or generic 'we' statements.")
elif label == "sentence-length variation" and d < -2:
priorities.append("Sentence rhythm is flatter than the references; vary sentence length more deliberately.")
elif label == "sentence length" and d > 4:
priorities.append("Average sentences are much longer than the references; split overloaded sentences.")
elif label == "acronym rate" and d > 0.01:
priorities.append("Acronym density is high relative to the reference set; unpack or remove non-essential abbreviations.")
elif label == "hype-word density" and d > 0:
priorities.append("Cut adjective-led importance language and let the evidence carry the emphasis.")
elif label == "generic-phrase density" and d > 0:
priorities.append("Replace generic manuscript phrases with precise claims or implications.")
elif label == "rhythm flatness score" and d > 0.1:
priorities.append("Sentence rhythm is unusually flat; mix short claim sentences with longer explanatory ones.")
seen = set()
unique = []
for p in priorities:
if p not in seen:
unique.append(p)
seen.add(p)
return unique[:8]
def as_json(data: Any) -> str:
return json.dumps(data, indent=2, ensure_ascii=False)
__all__ = [
"read_text",
"strip_frontmatter",
"paragraphs",
"sentences",
"words",
"metrics",
"aggregate",
"compare",
"rank_revision_priorities",
"as_json",
]