Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
89jobrien avatar

Pdf Processing

  • 26 installs
  • 4 repo stars
  • Updated April 11, 2026
  • 89jobrien/steve

pdf-processing is a Claude Code skill for extracting text and tables from PDFs and merging, splitting, and filling PDF documents in Python.

About

pdf-processing is a Claude Code skill for working with PDF files in Python. It shows how to extract text and tables with pdfplumber, merge and split documents with pypdf, fill forms, run OCR on scanned pages, and handle common errors. A developer uses it when they need to programmatically read, transform, or assemble PDF files.

  • Text and table extraction with pdfplumber, including tables-to-CSV
  • Merge and split PDFs with pypdf; form filling via FORMS.md
  • Notes OCR path (pytesseract) for scanned PDFs and error handling for empty pages

Pdf Processing by the numbers

  • 26 all-time installs (skills.sh)
  • Ranked #423 of 688 Office & Documents skills by installs in the Skillselion catalog
  • Data as of Jul 28, 2026 (Skillselion catalog sync)
At a glance

pdf processing capabilities & compatibility

Capabilities
pdf extraction · pdf merge split · pdf form filling · pdf ocr
Use cases
pdf parsing · documentation
Pricing
Free
From the docs

What pdf processing says it does

Extract text and tables from PDF files, fill forms, merge documents.
SKILL.md
Use pdfplumber to extract text from PDFs:
SKILL.md
pytesseract** - OCR for scanned PDFs (requires tesseract)
SKILL.md
npx skills add https://github.com/89jobrien/steve --skill pdf-processing

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs26
repo stars4
Last updatedApril 11, 2026
Repository89jobrien/steve

What it does

Extract text and tables from PDFs and merge, split, or fill PDF documents programmatically in Python.

Who is it for?

Developers who need to programmatically extract, transform, or assemble PDF files.

Skip if: Non-PDF document formats or users wanting a no-code GUI PDF editor.

When should I use this skill?

Working with PDF files, extracting text or tables, filling forms, or merging and splitting documents.

What you get

Extracted PDF text/tables or assembled PDF documents produced with the appropriate Python library.

  • extracted PDF text
  • extracted tables (CSV)
  • merged or split PDF

By the numbers

  • 4 documented PDF packages (pdfplumber, pypdf, pdf2image, pytesseract)
  • 5 performance tips

Files

SKILL.mdMarkdownGitHub ↗

PDF Processing

Quick start

Use pdfplumber to extract text from PDFs:

import pdfplumber

with pdfplumber.open("document.pdf") as pdf:
    text = pdf.pages[0].extract_text()
    print(text)

Extracting tables

Extract tables from PDFs with automatic detection:

import pdfplumber

with pdfplumber.open("report.pdf") as pdf:
    page = pdf.pages[0]
    tables = page.extract_tables()

    for table in tables:
        for row in table:
            print(row)

Extracting all pages

Process multi-page documents efficiently:

import pdfplumber

with pdfplumber.open("document.pdf") as pdf:
    full_text = ""
    for page in pdf.pages:
        full_text += page.extract_text() + "\n\n"

    print(full_text)

Form filling

For PDF form filling, see FORMS.md for the complete guide including field analysis and validation.

Merging PDFs

Combine multiple PDF files:

from pypdf import PdfMerger

merger = PdfMerger()

for pdf in ["file1.pdf", "file2.pdf", "file3.pdf"]:
    merger.append(pdf)

merger.write("merged.pdf")
merger.close()

Splitting PDFs

Extract specific pages or ranges:

from pypdf import PdfReader, PdfWriter

reader = PdfReader("input.pdf")
writer = PdfWriter()

# Extract pages 2-5
for page_num in range(1, 5):
    writer.add_page(reader.pages[page_num])

with open("output.pdf", "wb") as output:
    writer.write(output)

Available packages

  • pdfplumber - Text and table extraction (recommended)
  • pypdf - PDF manipulation, merging, splitting
  • pdf2image - Convert PDFs to images (requires poppler)
  • pytesseract - OCR for scanned PDFs (requires tesseract)

Common patterns

Extract and save text:

import pdfplumber

with pdfplumber.open("input.pdf") as pdf:
    text = "\n\n".join(page.extract_text() for page in pdf.pages)

with open("output.txt", "w") as f:
    f.write(text)

Extract tables to CSV:

import pdfplumber
import csv

with pdfplumber.open("tables.pdf") as pdf:
    tables = pdf.pages[0].extract_tables()

    with open("output.csv", "w", newline="") as f:
        writer = csv.writer(f)
        for table in tables:
            writer.writerows(table)

Error handling

Handle common PDF issues:

import pdfplumber

try:
    with pdfplumber.open("document.pdf") as pdf:
        if len(pdf.pages) == 0:
            print("PDF has no pages")
        else:
            text = pdf.pages[0].extract_text()
            if text is None or text.strip() == "":
                print("Page contains no extractable text (might be scanned)")
            else:
                print(text)
except Exception as e:
    print(f"Error processing PDF: {e}")

Performance tips

  • Process pages in batches for large PDFs
  • Use multiprocessing for multiple files
  • Extract only needed pages rather than entire document
  • Close PDF objects after use

Related skills

FAQ

What libraries does pdf-processing use?

pdfplumber for text and table extraction, pypdf for manipulation, pdf2image for image conversion, and pytesseract for OCR of scanned PDFs.

Can it handle scanned PDFs?

Yes, it detects pages with no extractable text and points to pytesseract OCR (which requires tesseract).

Can it fill PDF forms?

Yes, form filling is covered in the bundled FORMS.md guide using PyPDF2 and pdfrw.

Office & Documentsworkflownotes

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.