
Pdf Reader
- 8 installs
- 33 repo stars
- Updated April 26, 2026
- bighardperson/computer-science-skills-collection
pdf-reader is a Claude skill that extracts plain text and metadata from PDF files using PyMuPDF.
About
pdf-reader extracts plain text and retrieves metadata from PDF files using PyMuPDF. It exposes two commands: extract, which pulls text with an optional max-pages limit, and metadata, which returns a JSON object of document fields like title, author, and creation date. A developer uses it to pull text or metadata out of PDFs from the command line.
- Extracts plain text and metadata from PDFs using PyMuPDF
- Two commands: extract (with optional --max_pages) and metadata (JSON output)
- Handles encrypted and large PDFs gracefully
Pdf Reader by the numbers
- 8 all-time installs (skills.sh)
- Ranked #500 of 687 Office & Documents skills by installs in the Skillselion catalog
- Data as of Jul 30, 2026 (Skillselion catalog sync)
pdf-reader capabilities & compatibility
Free; requires python3 and the PyMuPDF pip package.
- Capabilities
- pdf parsing · documentation
- Use cases
- pdf parsing · documentation
- Pricing
- Free
What pdf-reader says it does
Extract text, search inside PDFs, and produce summaries.
Uses **PyMuPDF** (imported as `pymupdf`) for fast, reliable PDF processing
npx skills add https://github.com/bighardperson/computer-science-skills-collection --skill pdf-readerAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 8 |
|---|---|
| repo stars | ★ 33 |
| Last updated | April 26, 2026 |
| Repository | bighardperson/computer-science-skills-collection ↗ |
What it does
Extract text or structured metadata from a PDF file from the command line using PyMuPDF.
Who is it for?
Extracting plain text and metadata from PDFs via a simple CLI.
Skip if: Password-protected PDFs, which return an error unless a password is supplied.
When should I use this skill?
You need to pull text or metadata out of a PDF file.
What you get
Plain text or a JSON metadata object extracted from the target PDF.
- Extracted plain text from a PDF
- A JSON metadata object for a PDF
By the numbers
- Exposes 2 commands: extract and metadata
- metadata returns 9 fields (title, author, subject, creator, producer, and more)
Files
PDF Reader Skill
The pdf-reader skill provides functionality to extract text and retrieve metadata from PDF files using PyMuPDF (fitz).
Tool API
The skill provides two commands:
extract
Extracts plain text from the specified PDF file.
- Parameters:
file_path(string, required): Path to the PDF file to extract text from.--max_pages(integer, optional): Maximum number of pages to extract.
Usage:
python3 skills/pdf-reader/reader.py extract /path/to/document.pdf
python3 skills/pdf-reader/reader.py extract /path/to/document.pdf --max_pages 5Output: Plain text content from the PDF.
metadata
Retrieve metadata about the document.
- Parameters:
file_path(string, required): Path to the PDF file.
Usage:
python3 skills/pdf-reader/reader.py metadata /path/to/document.pdfOutput: JSON object with PDF metadata including:
title: Document titleauthor: Document authorsubject: Document subjectcreator: Application that created the PDFproducer: PDF producercreationDate: Creation datemodDate: Modification dateformat: PDF format versionencryption: Encryption info (if any)
Implementation Notes
- Uses PyMuPDF (imported as
pymupdf) for fast, reliable PDF processing - Supports encrypted PDFs (will return error if password required)
- Handles large PDFs efficiently with
max_pagesoption - Returns structured JSON for metadata command
Example
# Extract text from first 3 pages
python3 skills/pdf-reader/reader.py extract report.pdf --max_pages 3
# Get document metadata
python3 skills/pdf-reader/reader.py metadata report.pdf
# Output:
# {
# "title": "Annual Report 2024",
# "author": "John Doe",
# "creationDate": "D:20240115120000",
# ...
# }Error Handling
- Returns error message if file not found or not a valid PDF
- Returns error if PDF is encrypted and requires password
- Gracefully handles corrupted or malformed PDFs
{
"ownerId": "kn7ahjjpbd05yyj8nvxpdekzd181303t",
"slug": "iyeque-pdf-reader",
"version": "1.1.0",
"publishedAt": 1771326951970
}{
"version": 1,
"registry": "https://clawhub.ai",
"slug": "iyeque-pdf-reader",
"installedVersion": "1.1.0",
"installedAt": 1776071282257
}
import sys
import argparse
import json
import pymupdf
def extract_text(file_path, max_pages=None):
try:
doc = pymupdf.open(file_path)
text = ""
for i, page in enumerate(doc):
if max_pages and i >= int(max_pages):
break
text += page.get_text() + "\n"
return text
except Exception as e:
return f"Error: {str(e)}"
def get_metadata(file_path):
try:
doc = pymupdf.open(file_path)
return doc.metadata
except Exception as e:
return {"error": str(e)}
if __name__ == "__main__":
parser = argparse.ArgumentParser()
parser.add_argument("command", choices=["extract", "metadata"])
parser.add_argument("file_path")
parser.add_argument("--max_pages", type=int, default=None)
args = parser.parse_args()
if args.command == "extract":
print(extract_text(args.file_path, args.max_pages))
elif args.command == "metadata":
print(json.dumps(get_metadata(args.file_path), indent=2))
Related skills
FAQ
What commands does pdf-reader provide?
extract, which pulls plain text (optionally limited by --max_pages), and metadata, which returns a JSON object of PDF metadata.
Does pdf-reader handle encrypted PDFs?
It supports encrypted PDFs but returns an error if a password is required.