Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
beita6969 avatar

Pdf Processing Pro

  • 18 installs
  • 869 repo stars
  • Updated June 8, 2026
  • beita6969/scienceclaw

pdf-processing-pro is a skill that provides pre-built PDF-processing scripts for forms, tables, OCR, validation, and batch operations in production workflows.

About

This skill provides a production-oriented PDF processing toolkit with pre-built scripts for forms, tables, OCR, validation, and batch operations. A developer uses it for complex or high-volume PDF workflows needing robust error handling. Scripts expose CLI interfaces with consistent exit codes and logging suitable for automation pipelines.

  • Ships pre-built CLI scripts for form analysis, filling, table and text extraction, merge, split, and validation
  • Consistent exit codes (0 success, 4 validation error) for automation
  • Covers OCR, batch processing, and multi-page form workflows

Pdf Processing Pro by the numbers

  • 18 all-time installs (skills.sh)
  • Ranked #443 of 687 Office & Documents skills by installs in the Skillselion catalog
  • Data as of Aug 2, 2026 (Skillselion catalog sync)
At a glance

pdf-processing-pro capabilities & compatibility

Capabilities
pdf form fill · pdf table extraction · pdf text extraction · pdf ocr · pdf batch processing · pdf validation
Use cases
pdf parsing · documentation · ci cd
Pricing
Free
From the docs

What pdf-processing-pro says it does

Production-ready PDF processing with forms, tables, OCR, validation, and batch operations.
SKILL.md
Production-ready PDF processing toolkit with pre-built scripts, comprehensive error handling, and support for complex workflows.
SKILL.md
npx skills add https://github.com/beita6969/scienceclaw --skill pdf-processing-pro

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs18
repo stars869
Last updatedJune 8, 2026
Repositorybeita6969/scienceclaw

What it does

Run production PDF workflows with pre-built scripts for forms, tables, OCR, validation, and batch processing.

Who is it for?

Developers running complex or high-volume PDF pipelines that need validation and error handling

Skip if: Simple one-off text extraction where inline snippets suffice (use pdf-processing)

When should I use this skill?

The user works with complex PDF workflows in production, processes large volumes of PDFs, or needs robust validation

What you get

Validated, batch-processed PDFs and extracted data via CLI scripts fit for automation.

  • form field JSON
  • filled PDFs
  • extracted tables (CSV/Excel)

By the numbers

  • 10+ included scripts
  • 5 documented exit codes

Files

SKILL.mdMarkdownGitHub ↗

PDF Processing Pro

Production-ready PDF processing toolkit with pre-built scripts, comprehensive error handling, and support for complex workflows.

Quick start

Extract text from PDF

import pdfplumber

with pdfplumber.open("document.pdf") as pdf:
    text = pdf.pages[0].extract_text()
    print(text)

Analyze PDF form (using included script)

python scripts/analyze_form.py input.pdf --output fields.json
# Returns: JSON with all form fields, types, and positions

Fill PDF form with validation

python scripts/fill_form.py input.pdf data.json output.pdf
# Validates all fields before filling, includes error reporting

Extract tables from PDF

python scripts/extract_tables.py report.pdf --output tables.csv
# Extracts all tables with automatic column detection

Features

✅ Production-ready scripts

All scripts include:

  • Error handling: Graceful failures with detailed error messages
  • Validation: Input validation and type checking
  • Logging: Configurable logging with timestamps
  • Type hints: Full type annotations for IDE support
  • CLI interface: --help flag for all scripts
  • Exit codes: Proper exit codes for automation

✅ Comprehensive workflows

  • PDF Forms: Complete form processing pipeline
  • Table Extraction: Advanced table detection and extraction
  • OCR Processing: Scanned PDF text extraction
  • Batch Operations: Process multiple PDFs efficiently
  • Validation: Pre and post-processing validation

Advanced topics

PDF Form Processing

For complete form workflows including:

  • Field analysis and detection
  • Dynamic form filling
  • Validation rules
  • Multi-page forms
  • Checkbox and radio button handling

See FORMS.md

Table Extraction

For complex table extraction:

  • Multi-page tables
  • Merged cells
  • Nested tables
  • Custom table detection
  • Export to CSV/Excel

See TABLES.md

OCR Processing

For scanned PDFs and image-based documents:

  • Tesseract integration
  • Language support
  • Image preprocessing
  • Confidence scoring
  • Batch OCR

See OCR.md

Included scripts

Form processing

analyze_form.py - Extract form field information

python scripts/analyze_form.py input.pdf [--output fields.json] [--verbose]

fill_form.py - Fill PDF forms with data

python scripts/fill_form.py input.pdf data.json output.pdf [--validate]

validate_form.py - Validate form data before filling

python scripts/validate_form.py data.json schema.json

Table extraction

extract_tables.py - Extract tables to CSV/Excel

python scripts/extract_tables.py input.pdf [--output tables.csv] [--format csv|excel]

Text extraction

extract_text.py - Extract text with formatting preservation

python scripts/extract_text.py input.pdf [--output text.txt] [--preserve-formatting]

Utilities

merge_pdfs.py - Merge multiple PDFs

python scripts/merge_pdfs.py file1.pdf file2.pdf file3.pdf --output merged.pdf

split_pdf.py - Split PDF into individual pages

python scripts/split_pdf.py input.pdf --output-dir pages/

validate_pdf.py - Validate PDF integrity

python scripts/validate_pdf.py input.pdf

Common workflows

Workflow 1: Process form submissions

# 1. Analyze form structure
python scripts/analyze_form.py template.pdf --output schema.json

# 2. Validate submission data
python scripts/validate_form.py submission.json schema.json

# 3. Fill form
python scripts/fill_form.py template.pdf submission.json completed.pdf

# 4. Validate output
python scripts/validate_pdf.py completed.pdf

Workflow 2: Extract data from reports

# 1. Extract tables
python scripts/extract_tables.py monthly_report.pdf --output data.csv

# 2. Extract text for analysis
python scripts/extract_text.py monthly_report.pdf --output report.txt

Workflow 3: Batch processing

import glob
from pathlib import Path
import subprocess

# Process all PDFs in directory
for pdf_file in glob.glob("invoices/*.pdf"):
    output_file = Path("processed") / Path(pdf_file).name

    result = subprocess.run([
        "python", "scripts/extract_text.py",
        pdf_file,
        "--output", str(output_file)
    ], capture_output=True)

    if result.returncode == 0:
        print(f"✓ Processed: {pdf_file}")
    else:
        print(f"✗ Failed: {pdf_file} - {result.stderr}")

Error handling

All scripts follow consistent error patterns:

# Exit codes
# 0 - Success
# 1 - File not found
# 2 - Invalid input
# 3 - Processing error
# 4 - Validation error

# Example usage in automation
result = subprocess.run(["python", "scripts/fill_form.py", ...])

if result.returncode == 0:
    print("Success")
elif result.returncode == 4:
    print("Validation failed - check input data")
else:
    print(f"Error occurred: {result.returncode}")

Dependencies

All scripts require:

pip install pdfplumber pypdf pillow pytesseract pandas

Optional for OCR:

# Install tesseract-ocr system package
# macOS: brew install tesseract
# Ubuntu: apt-get install tesseract-ocr
# Windows: Download from GitHub releases

Performance tips

  • Use batch processing for multiple PDFs
  • Enable multiprocessing with --parallel flag (where supported)
  • Cache extracted data to avoid re-processing
  • Validate inputs early to fail fast
  • Use streaming for large PDFs (>50MB)

Best practices

1. Always validate inputs before processing 2. Use try-except in custom scripts 3. Log all operations for debugging 4. Test with sample PDFs before production 5. Set timeouts for long-running operations 6. Check exit codes in automation 7. Backup originals before modification

Troubleshooting

Common issues

"Module not found" errors:

pip install -r requirements.txt

Tesseract not found:

# Install tesseract system package (see Dependencies)

Memory errors with large PDFs:

# Process page by page instead of loading entire PDF
with pdfplumber.open("large.pdf") as pdf:
    for page in pdf.pages:
        text = page.extract_text()
        # Process page immediately

Permission errors:

chmod +x scripts/*.py

Getting help

All scripts support --help:

python scripts/analyze_form.py --help
python scripts/extract_tables.py --help

For detailed documentation on specific topics, see:

  • FORMS.md - Complete form processing guide
  • TABLES.md - Advanced table extraction
  • OCR.md - Scanned PDF processing

Related skills

FAQ

What do the exit codes mean?

0 is success, 1 file not found, 2 invalid input, 3 processing error, and 4 validation error.

Does it support OCR?

Yes, via tesseract integration for scanned and image-based documents.

Office & Documentsbackenddevops

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.