Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
beita6969 avatar

Arxiv Database

  • 18 installs
  • 869 repo stars
  • Updated June 8, 2026
  • beita6969/scienceclaw

arxiv-database is a Claude skill that searches and retrieves arXiv preprints via the public Atom API using bundled Python tools.

About

This skill provides Python tools to search and retrieve arXiv preprints via the public Atom API. It supports keyword, author, category, and arXiv ID search, date filtering, and PDF download, returning structured JSON with titles, abstracts, authors, and links. A researcher uses it to find papers in fields such as CS, ML, physics, math, and statistics or to build literature-review datasets.

  • Python tools search arXiv by keyword, author, category, or ID
  • Returns structured JSON and can download PDFs for full-text analysis
  • Covers CS, ML, physics, math, statistics, and more

Arxiv Database by the numbers

  • 18 all-time installs (skills.sh)
  • Ranked #10,674 of 16,556 AI & Agent Building skills by installs in the Skillselion catalog
  • Data as of Aug 2, 2026 (Skillselion catalog sync)
At a glance

arxiv-database capabilities & compatibility

Capabilities
arxiv search · literature review · pdf retrieval
Use cases
research · data analysis
Pricing
Free
From the docs

What arxiv-database says it does

This skill provides Python tools for searching and retrieving preprints from arXiv.org via its public Atom API.
SKILL.md
Building literature review datasets for AI/ML research
SKILL.md
npx skills add https://github.com/beita6969/scienceclaw --skill arxiv-database

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs18
repo stars869
Last updatedJune 8, 2026
Repositorybeita6969/scienceclaw

What it does

Search and retrieve arXiv preprints by keyword, author, ID, date, or category, and download PDFs via a Python API wrapper.

Who is it for?

Building literature-review datasets and finding preprints in CS, ML, physics, and math

Skip if: Biomedical literature, citation counts, or peer-reviewed journal-only searches

When should I use this skill?

You need to search arXiv or download preprint PDFs programmatically

What you get

Structured JSON of arXiv preprints with abstracts and links, plus downloaded PDFs for analysis.

  • structured JSON of arXiv results
  • downloaded paper PDFs
  • literature-review dataset

By the numbers

  • 5 core search capabilities
  • 8 query field prefixes (ti/au/abs/cat/all/co/jr/id)

Files

SKILL.mdMarkdownGitHub ↗

arXiv Database

Overview

This skill provides Python tools for searching and retrieving preprints from arXiv.org via its public Atom API. It supports keyword search, author search, category filtering, arXiv ID lookup, and PDF download. Results are returned as structured JSON with titles, abstracts, authors, categories, and links.

When to Use This Skill

Use this skill when:

  • Searching for preprints in CS, ML, AI, physics, math, statistics, q-bio, q-fin, or economics
  • Looking up specific papers by arXiv ID (e.g., 2309.10668)
  • Tracking an author's recent preprints
  • Filtering papers by arXiv category (e.g., cs.LG, cs.CL, stat.ML)
  • Downloading PDFs for full-text analysis
  • Building literature review datasets for AI/ML research
  • Monitoring new submissions in a subfield

Consider alternatives when:

  • Searching for biomedical literature specifically -> Use pubmed-database or biorxiv-database
  • You need citation counts or impact metrics -> Use openalex-database
  • You need peer-reviewed journal articles only -> Use pubmed-database

Core Search Capabilities

1. Keyword Search

Search for papers by keywords in titles, abstracts, or all fields.

python scripts/arxiv_search.py \
  --keywords "sparse autoencoders" "mechanistic interpretability" \
  --max-results 20 \
  --output results.json

With category filter:

python scripts/arxiv_search.py \
  --keywords "transformer" "attention mechanism" \
  --category cs.LG \
  --max-results 50 \
  --output transformer_papers.json

Search specific fields:

# Title only
python scripts/arxiv_search.py \
  --keywords "GRPO" \
  --search-field ti \
  --max-results 10

# Abstract only
python scripts/arxiv_search.py \
  --keywords "reward model" "RLHF" \
  --search-field abs \
  --max-results 30

2. Author Search

python scripts/arxiv_search.py \
  --author "Anthropic" \
  --max-results 50 \
  --output anthropic_papers.json
python scripts/arxiv_search.py \
  --author "Ilya Sutskever" \
  --category cs.LG \
  --max-results 20

3. arXiv ID Lookup

Retrieve metadata for specific papers:

python scripts/arxiv_search.py \
  --ids 2309.10668 2406.04093 2310.01405 \
  --output sae_papers.json

Full arXiv URLs also accepted:

python scripts/arxiv_search.py \
  --ids "https://arxiv.org/abs/2309.10668"

4. Category Browsing

List recent papers in a category:

python scripts/arxiv_search.py \
  --category cs.AI \
  --max-results 100 \
  --sort-by submittedDate \
  --output recent_cs_ai.json

5. PDF Download

python scripts/arxiv_search.py \
  --ids 2309.10668 \
  --download-pdf papers/

Batch download from search results:

import json
from scripts.arxiv_search import ArxivSearcher

searcher = ArxivSearcher()

# Search first
results = searcher.search(query="ti:sparse autoencoder", max_results=5)

# Download all
for paper in results:
    arxiv_id = paper["arxiv_id"]
    searcher.download_pdf(arxiv_id, f"papers/{arxiv_id.replace('/', '_')}.pdf")

arXiv Categories

Computer Science (cs.*)

CategoryDescription
cs.AIArtificial Intelligence
cs.CLComputation and Language (NLP)
cs.CVComputer Vision
cs.LGMachine Learning
cs.NENeural and Evolutionary Computing
cs.RORobotics
cs.CRCryptography and Security
cs.DSData Structures and Algorithms
cs.IRInformation Retrieval
cs.SESoftware Engineering

Statistics & Math

CategoryDescription
stat.MLMachine Learning (Statistics)
stat.MEMethodology
math.OCOptimization and Control
math.STStatistics Theory

Other Relevant Categories

CategoryDescription
q-bio.BMBiomolecules
q-bio.GNGenomics
q-bio.QMQuantitative Methods
q-fin.STStatistical Finance
eess.SPSignal Processing
physics.comp-phComputational Physics

Full list: see references/api_reference.md.

Query Syntax

The arXiv API uses prefix-based field searches combined with Boolean operators.

Field prefixes:

  • ti: - Title
  • au: - Author
  • abs: - Abstract
  • cat: - Category
  • all: - All fields (default)
  • co: - Comment
  • jr: - Journal reference
  • id: - arXiv ID

Boolean operators (must be UPPERCASE):

ti:transformer AND abs:attention
au:bengio OR au:lecun
cat:cs.LG ANDNOT cat:cs.CV

Grouping with parentheses:

(ti:sparse AND ti:autoencoder) AND cat:cs.LG
au:anthropic AND (abs:interpretability OR abs:alignment)

Examples:

from scripts.arxiv_search import ArxivSearcher

searcher = ArxivSearcher()

# Papers about SAEs in ML
results = searcher.search(
    query="ti:sparse autoencoder AND cat:cs.LG",
    max_results=50,
    sort_by="submittedDate"
)

# Specific author in specific field
results = searcher.search(
    query="au:neel nanda AND cat:cs.LG",
    max_results=20
)

# Complex boolean query
results = searcher.search(
    query="(abs:RLHF OR abs:reinforcement learning from human feedback) AND cat:cs.CL",
    max_results=100
)

Output Format

All searches return structured JSON:

{
  "query": "ti:sparse autoencoder AND cat:cs.LG",
  "result_count": 15,
  "results": [
    {
      "arxiv_id": "2309.10668",
      "title": "Towards Monosemanticity: Decomposing Language Models With Dictionary Learning",
      "authors": ["Trenton Bricken", "Adly Templeton", "..."],
      "abstract": "Full abstract text...",
      "categories": ["cs.LG", "cs.AI"],
      "primary_category": "cs.LG",
      "published": "2023-09-19T17:58:00Z",
      "updated": "2023-10-04T14:22:00Z",
      "doi": "10.48550/arXiv.2309.10668",
      "pdf_url": "http://arxiv.org/pdf/2309.10668v1",
      "abs_url": "http://arxiv.org/abs/2309.10668v1",
      "comment": "42 pages, 30 figures",
      "journal_ref": ""
    }
  ]
}

Common Usage Patterns

Literature Review Workflow

from scripts.arxiv_search import ArxivSearcher
import json

searcher = ArxivSearcher()

# 1. Broad search
results = searcher.search(
    query="abs:mechanistic interpretability AND cat:cs.LG",
    max_results=200,
    sort_by="submittedDate"
)

# 2. Save results
with open("interp_papers.json", "w") as f:
    json.dump({"result_count": len(results), "results": results}, f, indent=2)

# 3. Filter and analyze
import pandas as pd
df = pd.DataFrame(results)
print(f"Total papers: {len(df)}")
print(f"Date range: {df['published'].min()} to {df['published'].max()}")
print(f"\nTop categories:")
print(df["primary_category"].value_counts().head(10))

Track a Research Group

searcher = ArxivSearcher()

groups = {
    "anthropic": "au:anthropic AND (cat:cs.LG OR cat:cs.CL)",
    "openai": "au:openai AND cat:cs.CL",
    "deepmind": "au:deepmind AND cat:cs.LG",
}

for name, query in groups.items():
    results = searcher.search(query=query, max_results=50, sort_by="submittedDate")
    print(f"{name}: {len(results)} recent papers")

Monitor New Submissions

searcher = ArxivSearcher()

# Most recent ML papers
results = searcher.search(
    query="cat:cs.LG",
    max_results=50,
    sort_by="submittedDate",
    sort_order="descending"
)

for paper in results[:10]:
    print(f"[{paper['published'][:10]}] {paper['title']}")
    print(f"  {paper['abs_url']}\n")

Python API

from scripts.arxiv_search import ArxivSearcher

searcher = ArxivSearcher(verbose=True)

# Free-form query (uses arXiv query syntax)
results = searcher.search(query="...", max_results=50)

# Lookup by ID
papers = searcher.get_by_ids(["2309.10668", "2406.04093"])

# Download PDF
searcher.download_pdf("2309.10668", "paper.pdf")

# Build query from components
query = ArxivSearcher.build_query(
    title="sparse autoencoder",
    author="anthropic",
    category="cs.LG"
)
results = searcher.search(query=query, max_results=20)

Best Practices

1. Respect rate limits: The API requests 3-second delays between calls. The script handles this automatically. 2. Use category filters: Dramatically reduces noise. cs.LG is where most ML papers live. 3. Cache results: Save to JSON to avoid re-fetching. 4. Use `sort_by=submittedDate` for recent papers, relevance for keyword searches. 5. Max 300 results per query: arXiv API caps at this. For larger sets, paginate with start parameter. 6. arXiv IDs: Use bare IDs (2309.10668), not full URLs, in programmatic code. 7. Combine with openalex-database: For citation counts and impact metrics arXiv doesn't provide.

Limitations

  • No full-text search: Only searches metadata (title, abstract, authors, comments)
  • No citation data: Use openalex-database or Semantic Scholar for citations
  • Max 300 results: Per query. Use pagination for larger sets.
  • Rate limited: ~1 request per 3 seconds recommended
  • Atom XML responses: The script parses these into JSON automatically
  • Search lag: New papers may take hours to appear in API results

Reference Documentation

  • API Reference: See references/api_reference.md for full endpoint specs, all categories, and response schemas

Related skills

FAQ

What search types are supported?

Keyword, author, arXiv ID, and category search, with date filtering and boolean query syntax.

When should I use a different skill?

Use pubmed or biorxiv for biomedical literature and openalex for citation counts or impact metrics.

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.