
Pdf Processing
- 246 installs
- 30.1k repo stars
- Updated August 4, 2026
- davila7/claude-code-templates
pdf-processing is a Claude Code skill that teaches pypdf, pdfplumber, reportlab, and CLI tools to extract, merge, split, OCR, and transform PDFs for developers who ingest documents in backends or automation jobs.
About
pdf-processing is a document-handling skill from davila7/claude-code-templates that guides Python and command-line workflows for reading, creating, merging, splitting, rotating, watermarking, encrypting, and OCR-processing PDF files. It covers pypdf for basic operations, pdfplumber for text and table extraction to Excel, reportlab for PDF generation, plus qpdf, pdftotext, and pytesseract for CLI and scanned-document tasks. Developers reach for pdf-processing when a feature touches .pdf uploads, invoice parsing, form filling, or compliance document pipelines. Companion references FORMS.md and REFERENCE.md cover advanced form fill and JavaScript pdf-lib patterns. The quick-reference table maps eight common tasks to the best library or command.
- Text and table extraction patterns
- Merge, split, and rotate operations
- OCR and scanned-document handling
- Metadata and page-level processing
- Backend ingestion pipeline templates
Pdf Processing by the numbers
- 246 all-time installs (skills.sh)
- Ranked #212 of 688 Office & Documents skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/davila7/claude-code-templates --skill pdf-processingAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 246 |
|---|---|
| repo stars | ★ 30.1k |
| Last updated | August 4, 2026 |
| Repository | davila7/claude-code-templates ↗ |
How do you extract tables from PDFs in Python?
Extract, merge, split, OCR, and transform PDFs in app backends or automation jobs when documents are ingestion sources for search, billing, or compliance flows.
Who is it for?
Python developers building document ingestion, billing parsers, or compliance workflows that must read, transform, or generate PDF files programmatically.
Skip if: Teams needing only quick PDF viewing in a browser or production batch pipelines that require the separate pdf-processing-pro scripted toolkit.
When should I use this skill?
User mentions .pdf files, asks to extract text or tables, merge documents, fill PDF forms, or add OCR to scanned uploads.
What you get
Extracted text files, CSV or Excel tables, merged or split PDFs, filled forms, OCR text from scanned documents
- extracted text
- CSV or Excel tables
- transformed PDF files
By the numbers
- Quick-reference table covers 8 common PDF tasks across libraries and CLI tools
- Documents pypdf, pdfplumber, reportlab, qpdf, pdftotext, and pytesseract workflows
Files
PDF Processing
Quick start
Use pdfplumber to extract text from PDFs:
import pdfplumber
with pdfplumber.open("document.pdf") as pdf:
text = pdf.pages[0].extract_text()
print(text)Extracting tables
Extract tables from PDFs with automatic detection:
import pdfplumber
with pdfplumber.open("report.pdf") as pdf:
page = pdf.pages[0]
tables = page.extract_tables()
for table in tables:
for row in table:
print(row)Extracting all pages
Process multi-page documents efficiently:
import pdfplumber
with pdfplumber.open("document.pdf") as pdf:
full_text = ""
for page in pdf.pages:
full_text += page.extract_text() + "\n\n"
print(full_text)Form filling
For PDF form filling, see FORMS.md for the complete guide including field analysis and validation.
Merging PDFs
Combine multiple PDF files:
from pypdf import PdfMerger
merger = PdfMerger()
for pdf in ["file1.pdf", "file2.pdf", "file3.pdf"]:
merger.append(pdf)
merger.write("merged.pdf")
merger.close()Splitting PDFs
Extract specific pages or ranges:
from pypdf import PdfReader, PdfWriter
reader = PdfReader("input.pdf")
writer = PdfWriter()
# Extract pages 2-5
for page_num in range(1, 5):
writer.add_page(reader.pages[page_num])
with open("output.pdf", "wb") as output:
writer.write(output)Available packages
- pdfplumber - Text and table extraction (recommended)
- pypdf - PDF manipulation, merging, splitting
- pdf2image - Convert PDFs to images (requires poppler)
- pytesseract - OCR for scanned PDFs (requires tesseract)
Common patterns
Extract and save text:
import pdfplumber
with pdfplumber.open("input.pdf") as pdf:
text = "\n\n".join(page.extract_text() for page in pdf.pages)
with open("output.txt", "w") as f:
f.write(text)Extract tables to CSV:
import pdfplumber
import csv
with pdfplumber.open("tables.pdf") as pdf:
tables = pdf.pages[0].extract_tables()
with open("output.csv", "w", newline="") as f:
writer = csv.writer(f)
for table in tables:
writer.writerows(table)Error handling
Handle common PDF issues:
import pdfplumber
try:
with pdfplumber.open("document.pdf") as pdf:
if len(pdf.pages) == 0:
print("PDF has no pages")
else:
text = pdf.pages[0].extract_text()
if text is None or text.strip() == "":
print("Page contains no extractable text (might be scanned)")
else:
print(text)
except Exception as e:
print(f"Error processing PDF: {e}")Performance tips
- Process pages in batches for large PDFs
- Use multiprocessing for multiple files
- Extract only needed pages rather than entire document
- Close PDF objects after use
PDF Form Filling Guide
Overview
This guide covers filling PDF forms programmatically using PyPDF2 and pdfrw libraries.
Analyzing form fields
First, identify all fillable fields in a PDF:
from pypdf import PdfReader
reader = PdfReader("form.pdf")
fields = reader.get_fields()
for field_name, field_info in fields.items():
print(f"Field: {field_name}")
print(f" Type: {field_info.get('/FT')}")
print(f" Value: {field_info.get('/V')}")
print()Filling form fields
Fill fields with values:
from pypdf import PdfReader, PdfWriter
reader = PdfReader("form.pdf")
writer = PdfWriter()
writer.append_pages_from_reader(reader)
# Fill form fields
writer.update_page_form_field_values(
writer.pages[0],
{
"name": "John Doe",
"email": "john@example.com",
"address": "123 Main St"
}
)
with open("filled_form.pdf", "wb") as output:
writer.write(output)Flattening forms
Remove form fields after filling (make non-editable):
from pypdf import PdfReader, PdfWriter
reader = PdfReader("filled_form.pdf")
writer = PdfWriter()
for page in reader.pages:
writer.add_page(page)
# Flatten all form fields
writer.flatten_form_fields()
with open("flattened.pdf", "wb") as output:
writer.write(output)Validation
Validate field values before filling:
def validate_email(email):
return "@" in email and "." in email
def validate_form_data(data, required_fields):
errors = []
for field in required_fields:
if field not in data or not data[field]:
errors.append(f"Missing required field: {field}")
if "email" in data and not validate_email(data["email"]):
errors.append("Invalid email format")
return errors
# Usage
data = {"name": "John Doe", "email": "john@example.com"}
required = ["name", "email", "address"]
errors = validate_form_data(data, required)
if errors:
print("Validation errors:")
for error in errors:
print(f" - {error}")
else:
# Proceed with filling
passCommon field types
Text fields:
writer.update_page_form_field_values(
writer.pages[0],
{"text_field": "Some text"}
)Checkboxes:
# Check a checkbox
writer.update_page_form_field_values(
writer.pages[0],
{"checkbox_field": "/Yes"}
)
# Uncheck a checkbox
writer.update_page_form_field_values(
writer.pages[0],
{"checkbox_field": "/Off"}
)Radio buttons:
writer.update_page_form_field_values(
writer.pages[0],
{"radio_group": "/Option1"}
)Best practices
1. Always validate input data before filling 2. Check field names match exactly (case-sensitive) 3. Test with small files first 4. Keep originals - work on copies 5. Flatten after filling for distribution
Related skills
How it compares
Pick pdf-processing for library snippets and common Python PDF tasks; choose pdf-processing-pro in the same repo when you need pre-built CLI scripts with validation for production batch jobs.
FAQ
Which Python library extracts PDF tables?
pdf-processing recommends pdfplumber with page.extract_tables(), optionally exporting combined results to Excel via pandas, while pypdf handles merge, split, rotate, and encryption tasks.
How does pdf-processing handle scanned PDFs?
pdf-processing converts scanned pages to images with pdf2image, then runs pytesseract OCR per page to produce searchable text output from otherwise image-only documents.
What CLI tools does pdf-processing document?
pdf-processing documents qpdf for merge and split, pdftotext from poppler-utils for layout-preserving extraction, and pdftk when available for burst and rotate operations.