Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
oimiragieo avatar

Arxiv Mcp

  • 109 installs
  • 36 repo stars
  • Updated July 14, 2026
  • oimiragieo/agent-studio

Query arXiv papers, summaries, and citations inside agent workflows when validating ML ideas, drafting literature reviews, or grounding RAG answers in primary sources.

About

arxiv-mcp exposes arXiv as an MCP tool so coding agents can search, fetch, and reason over academic papers during early research. It supports literature reviews, method comparison, and citation gathering without manual browsing, keeping ML and AI product exploration evidence-based from the first conversation.

  • MCP-native arXiv search and retrieval
  • Literature review inside agent chats
  • Citation-ready paper metadata
  • Grounds AI features in published work
  • Speeds validate-before-build decisions

Arxiv Mcp by the numbers

  • 109 all-time installs (skills.sh)
  • Ranked #4,092 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
  • Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/oimiragieo/agent-studio --skill arxiv-mcp

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs109
repo stars36
Last updatedJuly 14, 2026
Repositoryoimiragieo/agent-studio

What it does

Query arXiv papers, summaries, and citations inside agent workflows when validating ML ideas, drafting literature reviews, or grounding RAG answers in primary sources.

Files

SKILL.mdMarkdownGitHub ↗

Mode: Cognitive/Prompt-Driven — No standalone utility script; use via agent context.

arXiv Search Skill

<identity> arXiv Search Skill - Search and retrieve academic papers from arXiv.org using existing tools (WebFetch, Exa). No MCP server installation required. </identity>

✅ No Installation Required

This skill uses existing tools to access arXiv:

  • WebFetch - Direct access to arXiv API
  • Exa - Semantic search with arXiv filtering

Works immediately - no MCP server, no restart needed.

<capabilities>

  • Search academic papers by keywords, authors, categories, or date ranges
  • Retrieve detailed paper metadata (title, authors, abstract, categories, PDF link)
  • Get specific papers by arXiv ID
  • Find related papers based on categories and keywords
  • Filter by arXiv categories (cs.AI, cs.LG, cs.CV, math., physics., etc.)
  • No API key required - uses public arXiv API

</capabilities>

Result Limits (Memory Safeguard)

arxiv-mcp returns academic papers. To prevent memory exhaustion:

  • max_results: 20 (HARD LIMIT)
  • Each paper metadata ~300 bytes
  • 20 papers × 300 bytes = ~6 KB metadata
  • Papers can be 100+ KB each if fetched - DON'T fetch full papers

Why the limit?

  • Previous limit: 100 results → 30 KB+ metadata → context explosion
  • New limit: 20 results → 6 KB metadata → memory safe
  • 20 papers is usually enough to find your target

<instructions> <execution_process>

Method 1: WebFetch with arXiv API (Recommended for specific queries)

The arXiv API is publicly accessible at http://export.arxiv.org/api/query.

Recommended Pattern

// ✓ GOOD: Limit results to 20
WebFetch({
  url: 'http://export.arxiv.org/api/query?search_query=all:transformer+attention&max_results=20&sortBy=relevance',
  prompt: 'Extract paper titles, authors, abstracts, arXiv IDs, and PDF links from these results',
});

// ✓ GOOD: Use specific filters to reduce result set
WebFetch({
  url: 'http://export.arxiv.org/api/query?search_query=all:transformer+attention+2025&max_results=20&sortBy=submittedDate',
  prompt: 'Extract recent papers on transformer attention',
});

// ✗ BAD: Old behavior - unlimited or >20 results
WebFetch({
  url: 'http://export.arxiv.org/api/query?search_query=all:neural+networks',
  // Too broad - will get 100s of results
});

// ✗ BAD: Exceeds memory limit
WebFetch({
  url: 'http://export.arxiv.org/api/query?search_query=all:deep+learning&max_results=100',
  // Over limit - memory risk
});

Search by Keywords

WebFetch({
  url: 'http://export.arxiv.org/api/query?search_query=all:transformer+attention&max_results=20&sortBy=relevance',
  prompt: 'Extract paper titles, authors, abstracts, arXiv IDs, and PDF links from these results',
});

Search by Author

WebFetch({
  url: 'http://export.arxiv.org/api/query?search_query=au:LeCun&max_results=10&sortBy=submittedDate',
  prompt: 'Extract paper titles, authors, abstracts, and arXiv IDs',
});

Search by Category

WebFetch({
  url: 'http://export.arxiv.org/api/query?search_query=cat:cs.LG&max_results=15&sortBy=submittedDate',
  prompt: 'Extract paper titles, authors, abstracts, categories, and arXiv IDs',
});

Get Specific Paper by ID

WebFetch({
  url: 'http://export.arxiv.org/api/query?id_list=2301.07041',
  prompt:
    'Extract full details: title, all authors, abstract, categories, published date, PDF link',
});

API Query Parameters

ParameterDescriptionExample
search_querySearch terms with field prefixesall:transformer, au:LeCun, ti:attention
id_listComma-separated arXiv IDs2301.07041,2302.13971
max_resultsNumber of results (default 10, max 100)max_results=20
startOffset for paginationstart=10
sortBySort order: relevance, lastUpdatedDate, submittedDatesortBy=submittedDate
sortOrderascending or descendingsortOrder=descending

Field Prefixes for search_query

PrefixFieldExample
all:All fieldsall:machine+learning
ti:Titleti:transformer
au:Authorau:Vaswani
abs:Abstractabs:attention+mechanism
cat:Categorycat:cs.LG
co:Commentco:accepted

Boolean Operators

Combine terms with AND, OR, ANDNOT:

search_query=ti:transformer+AND+abs:attention
search_query=au:LeCun+OR+au:Bengio
search_query=cat:cs.LG+ANDNOT+ti:survey

When NOT to Use arxiv-mcp

  • General web research → Use WebSearch/WebFetch instead
  • Implementation examples → Use pnpm search:code or ripgrep skill on codebase (Grep/Glob as fallback)
  • Product research → Use WebSearch with news filter
  • Community discussions → Use WebSearch for forums/Stack Overflow

arxiv-mcp is best for:

  • Finding academic papers on specific topics
  • Understanding theoretical foundations
  • Citing research in documentation
  • Quick literature review (20 papers max)

---

Method 2: Exa Search (Better for semantic/natural language queries)

Use Exa for more natural language queries with arXiv filtering:

Semantic Search

mcp__Exa__web_search_exa({
  query: 'site:arxiv.org transformer architecture attention mechanism deep learning',
  numResults: 10,
});

Recent Papers in a Field

mcp__Exa__web_search_exa({
  query: 'site:arxiv.org large language model scaling laws 2024',
  numResults: 15,
});

Author-Focused Search

mcp__Exa__web_search_exa({
  query: 'site:arxiv.org author:"Yann LeCun" deep learning',
  numResults: 10,
});

---

Common arXiv Categories

CategoryField
cs.AIArtificial Intelligence
cs.LGMachine Learning
cs.CLComputation and Language (NLP)
cs.CVComputer Vision
cs.SESoftware Engineering
cs.CRCryptography and Security
stat.MLMachine Learning (Statistics)
math.\*Mathematics (all subcategories)
physics.\*Physics (all subcategories)
q-bio.\*Quantitative Biology
econ.\*Economics

---

Workflow: Complete Research Process

Step 1: Initial Search

// Start with broad Exa search for semantic matching
mcp__Exa__web_search_exa({
  query: 'site:arxiv.org transformer attention mechanism neural networks',
  numResults: 10,
});

Step 2: Get Specific Papers

// Get details for interesting papers by ID
WebFetch({
  url: 'http://export.arxiv.org/api/query?id_list=2301.07041,2302.13971',
  prompt: 'Extract full metadata for each paper: title, authors, abstract, categories, PDF URL',
});

Step 3: Find Related Work

// Search by category of interesting paper
WebFetch({
  url: 'http://export.arxiv.org/api/query?search_query=cat:cs.LG+AND+ti:attention&max_results=10&sortBy=submittedDate',
  prompt: 'Find related papers, extract titles and abstracts',
});

Step 4: Get Recent Papers

// Latest papers in the field
WebFetch({
  url: 'http://export.arxiv.org/api/query?search_query=cat:cs.LG&max_results=20&sortBy=submittedDate&sortOrder=descending',
  prompt: 'Extract the 20 most recent machine learning papers',
});

</execution_process>

<best_practices>

1. Use Exa for discovery: Natural language queries find semantically related papers 2. Use WebFetch for precision: Specific IDs, categories, or API queries 3. Combine approaches: Exa to discover, WebFetch to deep-dive 4. Use specific queries: "transformer attention mechanism" > "machine learning" 5. Check multiple categories: Papers often span cs.AI + cs.LG + cs.CL 6. Sort by date for recent work: sortBy=submittedDate&sortOrder=descending

</best_practices> </instructions>

<examples> <usage_example> Example 1: Search for transformer papers:

WebFetch({
  url: 'http://export.arxiv.org/api/query?search_query=ti:transformer+AND+abs:attention&max_results=10&sortBy=relevance',
  prompt: 'Extract paper titles, authors, abstracts, and arXiv IDs',
});

Example 2: Find papers by researcher:

WebFetch({
  url: 'http://export.arxiv.org/api/query?search_query=au:Vaswani&max_results=15',
  prompt: 'List all papers by this author with titles and dates',
});

Example 3: Get recent ML papers:

WebFetch({
  url: 'http://export.arxiv.org/api/query?search_query=cat:cs.LG&max_results=20&sortBy=submittedDate&sortOrder=descending',
  prompt: 'Extract the 20 most recent machine learning papers with titles and abstracts',
});

Example 4: Semantic search with Exa:

mcp__Exa__web_search_exa({
  query: 'site:arxiv.org multimodal large language models vision 2024',
  numResults: 10,
});

Example 5: Get specific paper details:

WebFetch({
  url: 'http://export.arxiv.org/api/query?id_list=1706.03762',
  prompt: "Extract complete details for the 'Attention Is All You Need' paper",
});

</usage_example> </examples>

Agent Integration

This skill is automatically assigned to:

  • researcher - Academic research, literature review
  • scientific-research-expert - Deep scientific analysis
  • developer - Finding technical papers for implementation

Iron Laws

1. ALWAYS enforce max_results=20 — never allow unlimited or >20 result queries; context explosion from 100+ papers is a known failure mode that stalls agent pipelines. 2. NEVER fetch full paper PDFs during literature review — extract metadata and abstracts only; full papers are 100KB+ each and will exhaust context budget in minutes. 3. ALWAYS use Exa for semantic discovery, WebFetch for precision retrieval — Exa finds semantically related papers; WebFetch gets specific IDs or category feeds; use both in sequence, not interchangeably. 4. NEVER use broad queries without field prefixessearch_query=neural+networks returns thousands of results; always scope with ti:, au:, cat:, or abs: prefixes to target the query. 5. ALWAYS cite arXiv IDs (e.g., 2301.07041) when referencing papers — titles alone are ambiguous and change; IDs are stable, machine-readable, and enable instant retrieval.

Anti-Patterns

Anti-PatternWhy It FailsCorrect Approach
Using max_results=100 or no limitContext explosion; 100 papers × 300 bytes = 30KB+ metadataAlways set max_results=20 (hard limit)
Fetching full paper PDFsSingle paper can be 100KB+; kills context budgetExtract abstract + metadata only via API
Broad query without field prefixReturns irrelevant results across all fieldsUse ti:, au:, cat:, or abs: prefix
Using only WebFetch for discoveryMisses semantically related papers not matching exact termsUse Exa for semantic discovery first
Citing paper titles instead of arXiv IDsTitles can be ambiguous or duplicatedAlways include the arXiv ID (e.g., 1706.03762)

Memory Protocol (MANDATORY)

Before starting:

cat .claude/context/memory/learnings.md

After completing:

  • New pattern -> .claude/context/memory/learnings.md
  • Issue found -> .claude/context/memory/issues.md
  • Decision made -> .claude/context/memory/decisions.md
ASSUME INTERRUPTION: Your context may reset. If it's not in memory, it didn't happen.

Related skills

AI & Agent Buildingagentsresearchllm

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.