Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
bighardperson avatar

Pdf Reader

  • 8 installs
  • 33 repo stars
  • Updated April 26, 2026
  • bighardperson/computer-science-skills-collection

pdf-reader is a Claude skill that extracts plain text and metadata from PDF files using PyMuPDF.

About

pdf-reader extracts plain text and retrieves metadata from PDF files using PyMuPDF. It exposes two commands: extract, which pulls text with an optional max-pages limit, and metadata, which returns a JSON object of document fields like title, author, and creation date. A developer uses it to pull text or metadata out of PDFs from the command line.

  • Extracts plain text and metadata from PDFs using PyMuPDF
  • Two commands: extract (with optional --max_pages) and metadata (JSON output)
  • Handles encrypted and large PDFs gracefully

Pdf Reader by the numbers

  • 8 all-time installs (skills.sh)
  • Ranked #500 of 687 Office & Documents skills by installs in the Skillselion catalog
  • Data as of Jul 30, 2026 (Skillselion catalog sync)
At a glance

pdf-reader capabilities & compatibility

Free; requires python3 and the PyMuPDF pip package.

Capabilities
pdf parsing · documentation
Use cases
pdf parsing · documentation
Pricing
Free
From the docs

What pdf-reader says it does

Extract text, search inside PDFs, and produce summaries.
SKILL.md
Uses **PyMuPDF** (imported as `pymupdf`) for fast, reliable PDF processing
SKILL.md
npx skills add https://github.com/bighardperson/computer-science-skills-collection --skill pdf-reader

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs8
repo stars33
Last updatedApril 26, 2026
Repositorybighardperson/computer-science-skills-collection

What it does

Extract text or structured metadata from a PDF file from the command line using PyMuPDF.

Who is it for?

Extracting plain text and metadata from PDFs via a simple CLI.

Skip if: Password-protected PDFs, which return an error unless a password is supplied.

When should I use this skill?

You need to pull text or metadata out of a PDF file.

What you get

Plain text or a JSON metadata object extracted from the target PDF.

  • Extracted plain text from a PDF
  • A JSON metadata object for a PDF

By the numbers

  • Exposes 2 commands: extract and metadata
  • metadata returns 9 fields (title, author, subject, creator, producer, and more)

Files

SKILL.mdMarkdownGitHub ↗

PDF Reader Skill

The pdf-reader skill provides functionality to extract text and retrieve metadata from PDF files using PyMuPDF (fitz).

Tool API

The skill provides two commands:

extract

Extracts plain text from the specified PDF file.

  • Parameters:
  • file_path (string, required): Path to the PDF file to extract text from.
  • --max_pages (integer, optional): Maximum number of pages to extract.

Usage:

python3 skills/pdf-reader/reader.py extract /path/to/document.pdf
python3 skills/pdf-reader/reader.py extract /path/to/document.pdf --max_pages 5

Output: Plain text content from the PDF.

metadata

Retrieve metadata about the document.

  • Parameters:
  • file_path (string, required): Path to the PDF file.

Usage:

python3 skills/pdf-reader/reader.py metadata /path/to/document.pdf

Output: JSON object with PDF metadata including:

  • title: Document title
  • author: Document author
  • subject: Document subject
  • creator: Application that created the PDF
  • producer: PDF producer
  • creationDate: Creation date
  • modDate: Modification date
  • format: PDF format version
  • encryption: Encryption info (if any)

Implementation Notes

  • Uses PyMuPDF (imported as pymupdf) for fast, reliable PDF processing
  • Supports encrypted PDFs (will return error if password required)
  • Handles large PDFs efficiently with max_pages option
  • Returns structured JSON for metadata command

Example

# Extract text from first 3 pages
python3 skills/pdf-reader/reader.py extract report.pdf --max_pages 3

# Get document metadata
python3 skills/pdf-reader/reader.py metadata report.pdf
# Output:
# {
#   "title": "Annual Report 2024",
#   "author": "John Doe",
#   "creationDate": "D:20240115120000",
#   ...
# }

Error Handling

  • Returns error message if file not found or not a valid PDF
  • Returns error if PDF is encrypted and requires password
  • Gracefully handles corrupted or malformed PDFs

Related skills

FAQ

What commands does pdf-reader provide?

extract, which pulls plain text (optionally limited by --max_pages), and metadata, which returns a JSON object of PDF metadata.

Does pdf-reader handle encrypted PDFs?

It supports encrypted PDFs but returns an error if a password is required.

Office & Documentsworkflownotes

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.