Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
ar9av avatar

Paper Writing Bench

  • 44 installs
  • 628 repo stars
  • Updated July 9, 2026
  • ar9av/paperorchestra

paper-writing-bench is a Claude skill that reverse-engineers idea and experimental-log inputs from an existing paper to build pipeline benchmark cases.

About

Reverse-engineers raw materials (a Sparse idea, a Dense idea, and an experimental log) from an existing research paper to build a benchmark case for evaluating paper-writing pipelines. It replicates the PaperWritingBench dataset-construction procedure by stripping narrative flow from a paper PDF. A developer uses it to create test inputs, run their pipeline, and compare output against the original paper.

  • Reverse-engineers raw materials from an existing paper
  • Builds Sparse/Dense idea + experimental-log benchmark cases
  • Replicates the PaperWritingBench 200-paper dataset procedure

Paper Writing Bench by the numbers

  • 44 all-time installs (skills.sh)
  • Ranked #7,794 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
At a glance

paper-writing-bench capabilities & compatibility

Capabilities
testing · research · data analysis
Use cases
testing · research · data analysis
From the docs

What paper-writing-bench says it does

The original benchmark contains 200 papers (100 CVPR 2025 + 100 ICLR 2025). For each paper, the authors reverse-engineer the (I, E) tuple by stripping narrative flow from the original PDF
SKILL.md
**100% numeric accuracy in experimental log.** This becomes the ground truth for the section-writing-agent and content-refinement-agent's hallucination check.
SKILL.md
npx skills add https://github.com/ar9av/paperorchestra --skill paper-writing-bench

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs44
repo stars628
Last updatedJuly 9, 2026
Repositoryar9av/paperorchestra

What it does

Reverse-engineer benchmark idea and experiment-log inputs from a paper to evaluate a paper-writing pipeline.

Who is it for?

Creating benchmark test cases to evaluate a paper-writing pipeline against ground truth

Skip if: Writing a new paper from your own experiments

When should I use this skill?

The user asks to build a benchmark case from a paper, reverse-engineer raw materials, or evaluate a pipeline against PaperWritingBench

What you get

Produces idea_sparse.md, idea_dense.md, and experimental_log.md benchmark cases from a paper

  • idea_sparse.md
  • idea_dense.md
  • experimental_log.md

By the numbers

  • original benchmark has 200 papers (100 CVPR 2025 + 100 ICLR 2025)
  • 3 LLM calls per paper (sparse, dense, log)
  • produces 3 output files per paper

Files

SKILL.mdMarkdownGitHub ↗

PaperWritingBench (§3)

Faithful implementation of the PaperWritingBench dataset construction procedure from PaperOrchestra (Song et al., 2026, arXiv:2604.05018, §3 and App. C, F.2).

The original benchmark contains 200 papers (100 CVPR 2025 + 100 ICLR 2025). For each paper, the authors reverse-engineer the (I, E) tuple by stripping narrative flow from the original PDF using the three prompts in App. F.2. You can use this skill to reverse-engineer your own benchmark cases from any paper PDF.

What this skill does

Given an existing AI research paper (PDF or markdown extract), produce:

  • idea.md (Sparse variant) — high-level concept note, no math, no

experimental results

  • idea.md (Dense variant) — detailed technical proposal with LaTeX

equations and variable definitions, but still no experimental results

  • experimental_log.md — exhaustive raw experimental setup, numeric data,

and qualitative observations, with all narrative references stripped

These three files form a complete (I, E) input pair for the paper-orchestra pipeline. You can then run the pipeline and compare its output to the original paper using paper-autoraters.

Inputs

  • A paper PDF or extracted markdown text. The paper uses MinerU

(Wang et al., 2024) for PDF→markdown extraction; you (the host agent) should use whatever PDF extractor your environment provides.

  • For controlled experiments, you may also extract figures separately

(PDFFigures 2.0 in the paper).

Outputs

  • bench/<paper_id>/idea_sparse.md — Sparse variant
  • bench/<paper_id>/idea_dense.md — Dense variant
  • bench/<paper_id>/experimental_log.md — Experimental log

Workflow

For each paper, run three independent LLM calls using the verbatim prompts below:

1. Sparse idea generation

Load references/sparse-idea-prompt.md. Pass the paper text (or markdown extract) as {paper_content}. The prompt instructs the model to:

  • Stop extracting at empirical verification (no Experiments / Results / Comparisons)
  • Use first-person future tense ("We propose to explore...")
  • Avoid LaTeX math; describe components by function
  • Anonymize authors and titles

Output: idea_sparse.md with the four sections (Problem Statement, Core Hypothesis, Proposed Methodology high-level, Expected Contribution).

2. Dense idea generation

Load references/dense-idea-prompt.md. Same input. The prompt instructs the model to:

  • Preserve mathematical formulations using LaTeX
  • Define every variable used in equations
  • Include specific architectural choices and dimensions
  • Same exclusion zone (no experiments)

Output: idea_dense.md with the four sections (Problem Statement, Core Hypothesis, Proposed Methodology detailed, Expected Contribution).

3. Experimental log generation

Load references/experimental-log-prompt.md. Same input. The prompt instructs the model to:

  • Use past-tense persona ("We ran...", "The results were...")
  • Strip all references to figure/table numbers
  • Deconstruct tables into raw numeric data
  • Log figure findings as factual observations
  • Anonymize authors

Output: experimental_log.md with sections for Setup, Raw Numeric Data, and Qualitative Observations.

Critical rules from the prompts

These are excerpted from App. F.2. The host agent MUST honor them:

  • No citations. None of the three outputs may contain \cite,

reference numbers, or author names from the source paper.

  • No URLs. Strip all hyperlinks.
  • Anonymize. Author identities, affiliations, acknowledgements all

removed.

  • Self-contained. Each file must make sense without the original paper.
  • No experimental leakage in idea files. The Sparse and Dense ideas

must stop where empirical verification begins. They describe what will be done, not what was done.

  • No table/figure references in experimental log. No "as shown in

Table 1", "see Fig. 5". The downstream paper-orchestra pipeline will generate its own figures and tables — the log must not assume any particular ones exist.

  • 100% numeric accuracy in experimental log. This becomes the ground

truth for the section-writing-agent and content-refinement-agent's hallucination check.

How the bench is used

After producing (idea_sparse.md, idea_dense.md, experimental_log.md) for a paper:

1. Pick a variant (Sparse or Dense) — the paper ablates both, with Dense producing more rigorous methodology and Sparse exercising the system's robustness on under-specified inputs. 2. Drop the chosen idea.md, plus experimental_log.md, plus a template.tex for the target conference, plus a conference_guidelines.md, into a paper-orchestra workspace. 3. Run the pipeline. 4. Compare the generated paper against the original using paper-autoraters (citation F1, lit review quality, SxS paper quality).

Resources

  • references/bench-overview.md — the 200-paper bench, venue cutoffs, sizes
  • references/sparse-idea-prompt.md — verbatim from App. F.2
  • references/dense-idea-prompt.md — verbatim from App. F.2
  • references/experimental-log-prompt.md — verbatim from App. F.2

Related skills

AI & Agent Buildingresearchagentsautomation

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.