
Pdf Reader
- 524 installs
- Updated January 23, 2026
- childbamboo/claude-code-marketplace-sample
pdf-reader is a Claude Code marketplace sample skill that helps coding agents extract and work with PDF document content during AI and agent-building tasks.
About
Reads multi-page PDFs and extracts text and tables, converting to Markdown via a pdfplumber script. Used when a developer needs to read or convert PDF documents.
- Handles tables and multi-page PDFs
- Extracts text via pdfplumber to Markdown
Pdf Reader by the numbers
- 524 all-time installs (skills.sh)
- Ranked #148 of 688 Office & Documents skills by installs in the Skillselion catalog
- Data as of Jul 28, 2026 (Skillselion catalog sync)
npx skills add https://github.com/childbamboo/claude-code-marketplace-sample --skill pdf-readerAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 524 |
|---|---|
| Last updated | January 23, 2026 |
| Repository | childbamboo/claude-code-marketplace-sample ↗ |
How do agents read and parse PDF documents?
Extract text and tables from PDF files and convert them to Markdown using a pdfplumber script in WSL.
Who is it for?
Developers prototyping document-aware Claude Code agents that must ingest PDF specs, reports, or contracts.
Skip if: Production OCR pipelines, large-scale batch document indexing, or teams that only handle plain text or Markdown inputs.
When should I use this skill?
A developer asks an agent to read, summarize, or extract data from a PDF file during an AI or agent-building session.
What you get
Parsed PDF text, structured excerpts, and agent-ready document context from uploaded or referenced PDF files.
- Extracted PDF text
- Agent-ready document summaries
Files
PDF Reader
PDF ファイルをテキスト抽出して Markdown 形式に変換するスキルです。
クイックスタート
基本的な使い方
# WSL環境でPythonスクリプトを実行
wsl python3 scripts/read_pdf.py "/mnt/c/path/to/file.pdf"Markdown形式で保存
1. スクリプトでテキスト抽出 2. Write ツールで .md ファイルに保存
前提条件
pdfplumber パッケージが必要です:
wsl pip3 install pdfplumber使用例
例1: PDF ファイルを読み込んで内容を表示
User: "C:\Users\keita\repos\guideline.pdf を読み込んで"
Assistant:
1. Windowsパスを WSL パスに変換: /mnt/c/Users/keita/repos/guideline.pdf
2. wsl python3 scripts/read_pdf.py を実行
3. 抽出されたテキストを Markdown 形式で表示例2: PDF を Markdown に変換して保存
User: "ガイドライン.pdf を Markdown に変換して保存"
Assistant:
1. scripts/read_pdf.py でテキスト抽出
2. Markdown形式で構造化(ページごとに見出し、テーブルも含む)
3. Write ツールで ガイドライン.md に保存
4. 保存完了を報告ワークフロー
単一ファイルの読み込み
1. ユーザーが PDF ファイルパスを指定 2. Windows パスを WSL パス形式に変換 (C:\ → /mnt/c/) 3. wsl python3 scripts/read_pdf.py を実行 4. 抽出されたテキストを Markdown 形式で表示または保存
複数ファイルの一括処理
1. Glob で .pdf ファイルを検索 2. 各ファイルに対してスクリプトを実行 3. 結果をまとめて報告
出力形式
Markdown 構造
# [PDFファイル名]
**Total Pages:** 10
---
## Page 1
[ページ1のテキスト内容]
### Tables
**Table 1:**
| 列1 | 列2 | 列3 |
| --- | --- | --- |
| データ1 | データ2 | データ3 |
---
## Page 2
[ページ2のテキスト内容]
---スクリプト詳細
Python スクリプトは scripts/read_pdf.py に配置されています。
主な機能:
- ページごとのテキスト抽出
- テーブルの Markdown 化
- 複数ページの構造化
- エラーハンドリング
使い方:
python scripts/read_pdf.py <file_path>対応機能
- ✅ テキスト抽出(全ページ)
- ✅ テーブルの Markdown 化
- ✅ ページ番号の保持
- ✅ 構造化された出力
- ⚠️ 画像からのテキスト抽出(OCR未対応)
- ⚠️ 複雑なレイアウトは簡略化
制限事項
- スキャンされた PDF(画像のみ)からはテキスト抽出不可
- OCR 機能は含まれません
- 複雑なレイアウトは簡略化されます
- フォント情報、色などのスタイルは失われます
- 埋め込みオブジェクトは抽出されません
トラブルシューティング
pdfplumber がインストールされていない
wsl pip3 install pdfplumberテキストが抽出されない
- PDF がスキャン画像の可能性があります(OCR が必要)
- PDF が暗号化されている可能性があります
- テキストレイヤーがない PDF かもしれません
文字化けする
# 日本語対応の確認
wsl locale
# UTF-8 が含まれていることを確認メモリ不足エラー
大きな PDF ファイルの場合、ページごとに分割して処理することを検討してください。
パス変換
Windows パスから WSL パスへの変換:
C:\Users\...→/mnt/c/Users/...D:\Projects\...→/mnt/d/Projects/...- バックスラッシュ
\をスラッシュ/に変換
関連ツール
- PyPDF2: 軽量な代替ライブラリ
- pdfminer.six: より詳細な制御が必要な場合
- Camelot: テーブル抽出特化
- OCRmyPDF: スキャン PDF に OCR を適用
高度な使い方
特定のページのみ抽出
スクリプトを修正して pdf.pages[0:5] のようにスライスを使用できます。
テーブルのみ抽出
スクリプト内の extract_tables() 部分のみを使用します。
OCR が必要な場合
pytesseract と pdf2image を組み合わせて使用します(別スキルとして作成推奨)。
バージョン履歴
- v1.0.0 (2026-01-06): 初期リリース
- 基本的なテキスト抽出機能
- テーブル Markdown 化対応
- WSL環境での動作
- ページごとの構造化
PDF Reader Skill
PDF ファイルをテキスト抽出して Markdown 形式に変換するスキルです。
ファイル構成
pdf-reader/
├── SKILL.md # メインスキル定義(Claude が読む)
├── README.md # このファイル(人間向けドキュメント)
└── scripts/
└── read_pdf.py # テキスト抽出用 Python スクリプトインストール
前提条件
- WSL (Windows Subsystem for Linux)
- Python 3.x
- pdfplumber パッケージ
セットアップ
# pdfplumber のインストール
wsl pip3 install pdfplumber使い方
Claude に以下のように依頼します:
「C:\Users\keita\repos\guideline.pdf を読み込んで」Claude が自動的に: 1. Windows パスを WSL パスに変換 2. スクリプトを実行してテキスト抽出 3. Markdown 形式で構造化 4. 結果を表示または Markdown ファイルとして保存
スクリプトの直接実行
# 基本的な使い方
wsl python3 scripts/read_pdf.py "/mnt/c/path/to/file.pdf"
# 出力をファイルに保存
wsl python3 scripts/read_pdf.py "/mnt/c/path/to/file.pdf" > output.md機能
テキスト抽出
- ✅ 全ページからテキスト抽出
- ✅ ページごとに構造化
- ✅ 見出しとして整理
- ❌ OCR 機能(未対応)
テーブル処理
- ✅ テーブルの自動検出
- ✅ Markdown テーブル形式に変換
- ✅ 複数テーブルの処理
- ⚠️ 複雑なレイアウトは簡略化
出力形式
- ✅ Markdown 形式
- ✅ ページ番号付き
- ✅ 階層構造
- ✅ テーブル対応
出力例
# document.pdf
**Total Pages:** 3
---
## Page 1
これはページ1のテキスト内容です。
### Tables
**Table 1:**
| ヘッダー1 | ヘッダー2 | ヘッダー3 |
| --- | --- | --- |
| データ1 | データ2 | データ3 |
---
## Page 2
これはページ2のテキスト内容です。
---トラブルシューティング
pdfplumber が見つからない
wsl pip3 install pdfplumberテキストが抽出できない
原因1: スキャン画像 PDF
- OCR が必要です
- 対策: OCR ツールで前処理するか、OCR 対応の別スキルを使用
原因2: 暗号化された PDF
- パスワード保護されている
- 対策: パスワードを解除してから実行
原因3: テキストレイヤーがない
- PDF にテキスト情報が埋め込まれていない
- 対策: OCR を使用
文字化けする
# ロケール設定を確認
wsl locale
# 必要に応じて UTF-8 を設定
export LANG=ja_JP.UTF-8大きな PDF でメモリ不足
- ページ数が多い場合は分割処理を検討
- スクリプトを修正してページ範囲を指定
開発
スクリプトの修正
scripts/read_pdf.py を編集して機能を追加・修正できます。
カスタマイズ例
特定ページのみ抽出
# ページ 1-5 のみ抽出
for page_num, page in enumerate(pdf.pages[0:5], start=1):
...テーブルのみ抽出
# テキストを無視してテーブルのみ
tables = page.extract_tables()テスト
# テスト用の PDF ファイルで動作確認
wsl python3 scripts/read_pdf.py "/mnt/c/path/to/test.pdf"関連スキル
- docx-reader: Word 文書の読み込み
- ocr-reader (未作成): スキャン画像からのテキスト抽出
技術詳細
使用ライブラリ
- pdfplumber: PDF からのテキスト・テーブル抽出
- 高精度なテキスト抽出
- テーブル構造の保持
- ページレイアウトの解析
アルゴリズム
1. PDF ファイルを開く 2. 各ページを順に処理 3. テキストを抽出 4. テーブルを検出・抽出 5. Markdown 形式に整形 6. 全体を結合
制限事項
- スキャン画像 PDF は非対応(OCR 必要)
- 複雑なレイアウトは簡略化される
- 画像、図形は抽出されない
- フォント・色などのスタイル情報は失われる
- 巨大な PDF はメモリ制約に注意
ライセンス
このスキルは個人プロジェクト用です。
バージョン
- v1.0.0 (2026-01-06)
- 初期リリース
- 基本的なテキスト抽出機能
- テーブル Markdown 化
- WSL環境での動作確認済み
- ページごとの構造化出力
#!/usr/bin/env python3
"""
PDF Reader Script
Extracts text content from PDF files and converts to Markdown format.
"""
import sys
import os
try:
import pdfplumber
except ImportError:
print("Error: pdfplumber is not installed.")
print("Please install it with: pip install pdfplumber")
sys.exit(1)
def read_pdf(file_path):
"""
Read PDF file and extract text content.
Args:
file_path (str): Path to the PDF file
Returns:
str: Extracted text content in Markdown format
"""
try:
markdown_content = []
with pdfplumber.open(file_path) as pdf:
# Add document title
markdown_content.append(f"# {os.path.basename(file_path)}\n")
markdown_content.append(f"**Total Pages:** {len(pdf.pages)}\n")
markdown_content.append("---\n")
# Extract text from each page
for page_num, page in enumerate(pdf.pages, start=1):
markdown_content.append(f"## Page {page_num}\n")
# Extract text
text = page.extract_text()
if text:
markdown_content.append(text)
else:
markdown_content.append("*(No text content on this page)*")
# Extract tables
tables = page.extract_tables()
if tables:
markdown_content.append("\n### Tables\n")
for table_num, table in enumerate(tables, start=1):
markdown_content.append(f"\n**Table {table_num}:**\n")
markdown_content.append(format_table_as_markdown(table))
markdown_content.append("\n---\n")
return '\n'.join(markdown_content)
except FileNotFoundError:
return f"Error: File not found: {file_path}"
except Exception as e:
return f"Error reading PDF: {str(e)}"
def format_table_as_markdown(table):
"""
Format a table as Markdown.
Args:
table (list): Table data as list of lists
Returns:
str: Markdown formatted table
"""
if not table or not table[0]:
return ""
markdown_table = []
# Header row
header = table[0]
markdown_table.append("| " + " | ".join(str(cell) if cell else "" for cell in header) + " |")
# Separator row
markdown_table.append("| " + " | ".join("---" for _ in header) + " |")
# Data rows
for row in table[1:]:
markdown_table.append("| " + " | ".join(str(cell) if cell else "" for cell in row) + " |")
return "\n".join(markdown_table)
def main():
"""Main entry point for the script."""
if len(sys.argv) < 2:
print("Usage: python read_pdf.py <file_path>")
sys.exit(1)
file_path = sys.argv[1]
if not os.path.exists(file_path):
print(f"Error: File not found: {file_path}")
sys.exit(1)
if not file_path.lower().endswith('.pdf'):
print("Warning: File does not have .pdf extension")
content = read_pdf(file_path)
print(content)
if __name__ == "__main__":
main()
Related skills
FAQ
What does the pdf-reader skill do?
pdf-reader is a Claude Code marketplace sample skill that lets coding agents open PDF files, extract readable text and structure, and use that content during AI and agent-building tasks instead of ignoring binary attachments.
When should developers install pdf-reader?
Developers install pdf-reader when building document-aware agents that routinely consume PDF contracts, specifications, or reports and need a lightweight marketplace skill rather than a custom parser integration.