
Markitdown
- 77 installs
- 16 repo stars
- Updated November 20, 2025
- jackspace/claudeskillz
Convert PDFs, Office files, images, audio, and web content into clean Markdown optimized for LLM processing and RAG.
About
Uses the MarkItDown utility to convert 20+ file formats into structured Markdown for LLM pipelines. A developer uses it to extract text from documents, OCR images, transcribe audio, or prepare files for RAG.
- Supports DOCX, XLSX, PPTX, PDF, HTML, EPUB, CSV, and more
- OCR, audio transcription, and YouTube transcript extraction
Markitdown by the numbers
- 77 all-time installs (skills.sh)
- +1 installs in the week ending Aug 2, 2026 (Skillselion tracking)
- Ranked #344 of 687 Office & Documents skills by installs in the Skillselion catalog
- Data as of Aug 2, 2026 (Skillselion catalog sync)
npx skills add https://github.com/jackspace/claudeskillz --skill markitdownAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 77 |
|---|---|
| repo stars | ★ 16 |
| Last updated | November 20, 2025 |
| Repository | jackspace/claudeskillz ↗ |
What it does
Convert PDFs, Office files, images, audio, and web content into clean Markdown optimized for LLM processing and RAG.
Files
MarkItDown
Overview
MarkItDown is a Python utility that converts various file formats into Markdown format, optimized for use with large language models and text analysis pipelines. It preserves document structure (headings, lists, tables, hyperlinks) while producing clean, token-efficient Markdown output.
When to Use This Skill
Use this skill when users request:
- Converting documents to Markdown format
- Extracting text from PDF, Word, PowerPoint, or Excel files
- Performing OCR on images to extract text
- Transcribing audio files to text
- Extracting YouTube video transcripts
- Processing HTML, EPUB, or web content to Markdown
- Converting structured data (CSV, JSON, XML) to readable Markdown
- Batch converting multiple files or ZIP archives
- Preparing documents for LLM analysis or RAG systems
Core Capabilities
1. Document Conversion
Convert Office documents and PDFs to Markdown while preserving structure.
Supported formats:
- PDF files (with optional Azure Document Intelligence integration)
- Word documents (DOCX)
- PowerPoint presentations (PPTX)
- Excel spreadsheets (XLSX, XLS)
Basic usage:
from markitdown import MarkItDown
md = MarkItDown()
result = md.convert("document.pdf")
print(result.text_content)Command-line:
markitdown document.pdf -o output.mdSee references/document_conversion.md for detailed documentation on document-specific features.
2. Media Processing
Extract text from images using OCR and transcribe audio files to text.
Supported formats:
- Images (JPEG, PNG, GIF, etc.) with EXIF metadata extraction
- Audio files with speech transcription (requires speech_recognition)
Image with OCR:
from markitdown import MarkItDown
md = MarkItDown()
result = md.convert("image.jpg")
print(result.text_content) # Includes EXIF metadata and OCR textAudio transcription:
result = md.convert("audio.wav")
print(result.text_content) # Transcribed speechSee references/media_processing.md for advanced media handling options.
3. Web Content Extraction
Convert web-based content and e-books to Markdown.
Supported formats:
- HTML files and web pages
- YouTube video transcripts (via URL)
- EPUB books
- RSS feeds
YouTube transcript:
from markitdown import MarkItDown
md = MarkItDown()
result = md.convert("https://youtube.com/watch?v=VIDEO_ID")
print(result.text_content)See references/web_content.md for web extraction details.
4. Structured Data Handling
Convert structured data formats to readable Markdown tables.
Supported formats:
- CSV files
- JSON files
- XML files
CSV to Markdown table:
from markitdown import MarkItDown
md = MarkItDown()
result = md.convert("data.csv")
print(result.text_content) # Formatted as Markdown tableSee references/structured_data.md for format-specific options.
5. Advanced Integrations
Enhance conversion quality with AI-powered features.
Azure Document Intelligence: For enhanced PDF processing with better table extraction and layout analysis:
from markitdown import MarkItDown
md = MarkItDown(docintel_endpoint="<endpoint>", docintel_key="<key>")
result = md.convert("complex.pdf")LLM-Powered Image Descriptions: Generate detailed image descriptions using GPT-4o:
from markitdown import MarkItDown
from openai import OpenAI
client = OpenAI()
md = MarkItDown(llm_client=client, llm_model="gpt-4o")
result = md.convert("presentation.pptx") # Images described with LLMSee references/advanced_integrations.md for integration details.
6. Batch Processing
Process multiple files or entire ZIP archives at once.
ZIP file processing:
from markitdown import MarkItDown
md = MarkItDown()
result = md.convert("archive.zip")
print(result.text_content) # All files converted and concatenatedBatch script: Use the provided batch processing script for directory conversion:
python scripts/batch_convert.py /path/to/documents /path/to/outputSee scripts/batch_convert.py for implementation details.
Installation
Full installation (all features):
pip install 'markitdown[all]'Modular installation (specific features):
pip install 'markitdown[pdf]' # PDF support
pip install 'markitdown[docx]' # Word support
pip install 'markitdown[pptx]' # PowerPoint support
pip install 'markitdown[xlsx]' # Excel support
pip install 'markitdown[audio]' # Audio transcription
pip install 'markitdown[youtube]' # YouTube transcriptsRequirements:
- Python 3.10 or higher
Output Format
MarkItDown produces clean, token-efficient Markdown optimized for LLM consumption:
- Preserves headings, lists, and tables
- Maintains hyperlinks and formatting
- Includes metadata where relevant (EXIF, document properties)
- No temporary files created (streaming approach)
Common Workflows
Preparing documents for RAG:
from markitdown import MarkItDown
md = MarkItDown()
# Convert knowledge base documents
docs = ["manual.pdf", "guide.docx", "faq.html"]
markdown_content = []
for doc in docs:
result = md.convert(doc)
markdown_content.append(result.text_content)
# Now ready for embedding and indexingDocument analysis pipeline:
# Convert all PDFs in directory
for file in documents/*.pdf; do
markitdown "$file" -o "markdown/$(basename "$file" .pdf).md"
donePlugin System
MarkItDown supports extensible plugins for custom conversion logic. Plugins are disabled by default for security:
from markitdown import MarkItDown
# Enable plugins if needed
md = MarkItDown(enable_plugins=True)Resources
This skill includes comprehensive reference documentation for each capability:
- references/document_conversion.md - Detailed PDF, DOCX, PPTX, XLSX conversion options
- references/media_processing.md - Image OCR and audio transcription details
- references/web_content.md - HTML, YouTube, and EPUB extraction
- references/structured_data.md - CSV, JSON, XML conversion formats
- references/advanced_integrations.md - Azure Document Intelligence and LLM integration
- scripts/batch_convert.py - Batch processing utility for directories
{
"description": "Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing. Use when converting documents to markdown, extracting text from PDFs/Office files, transcribing audio, performing OCR on images, extracting YouTube transcripts, or processing batches of files. Supports 20+ formats including DOCX, XLSX, PPTX, PDF, HTML, EPUB, CSV, JSON, images with OCR, and audio with transcription.",
"references": {
"files": [
"references/advanced_integrations.md",
"references/document_conversion.md",
"references/media_processing.md",
"references/structured_data.md",
"references/web_content.md"
]
},
"content": "**Preparing documents for RAG:**\r\n```python\r\nfrom markitdown import MarkItDown\r\n\r\nmd = MarkItDown()\r\n\r\ndocs = [\"manual.pdf\", \"guide.docx\", \"faq.html\"]\r\nmarkdown_content = []\r\n\r\nfor doc in docs:\r\n result = md.convert(doc)\r\n markdown_content.append(result.text_content)\r\n\r\n```\r\n\r\n**Document analysis pipeline:**\r\n```bash\r\n\r\nMarkItDown supports extensible plugins for custom conversion logic. Plugins are disabled by default for security:\r\n\r\n```python\r\nfrom markitdown import MarkItDown",
"name": "markitdown",
"id": "scientific-pkg-markitdown",
"sections": {
"Output Format": "MarkItDown produces clean, token-efficient Markdown optimized for LLM consumption:\r\n- Preserves headings, lists, and tables\r\n- Maintains hyperlinks and formatting\r\n- Includes metadata where relevant (EXIF, document properties)\r\n- No temporary files created (streaming approach)",
"Installation": "**Full installation (all features):**\r\n```bash\r\npip install 'markitdown[all]'\r\n```\r\n\r\n**Modular installation (specific features):**\r\n```bash\r\npip install 'markitdown[pdf]' # PDF support\r\npip install 'markitdown[docx]' # Word support\r\npip install 'markitdown[pptx]' # PowerPoint support\r\npip install 'markitdown[xlsx]' # Excel support\r\npip install 'markitdown[audio]' # Audio transcription\r\npip install 'markitdown[youtube]' # YouTube transcripts\r\n```\r\n\r\n**Requirements:**\r\n- Python 3.10 or higher",
"Overview": "MarkItDown is a Python utility that converts various file formats into Markdown format, optimized for use with large language models and text analysis pipelines. It preserves document structure (headings, lists, tables, hyperlinks) while producing clean, token-efficient Markdown output.",
"When to Use This Skill": "Use this skill when users request:\r\n- Converting documents to Markdown format\r\n- Extracting text from PDF, Word, PowerPoint, or Excel files\r\n- Performing OCR on images to extract text\r\n- Transcribing audio files to text\r\n- Extracting YouTube video transcripts\r\n- Processing HTML, EPUB, or web content to Markdown\r\n- Converting structured data (CSV, JSON, XML) to readable Markdown\r\n- Batch converting multiple files or ZIP archives\r\n- Preparing documents for LLM analysis or RAG systems",
"Resources": "This skill includes comprehensive reference documentation for each capability:\r\n\r\n- **references/document_conversion.md** - Detailed PDF, DOCX, PPTX, XLSX conversion options\r\n- **references/media_processing.md** - Image OCR and audio transcription details\r\n- **references/web_content.md** - HTML, YouTube, and EPUB extraction\r\n- **references/structured_data.md** - CSV, JSON, XML conversion formats\r\n- **references/advanced_integrations.md** - Azure Document Intelligence and LLM integration\r\n- **scripts/batch_convert.py** - Batch processing utility for directories",
"Core Capabilities": "### 1. Document Conversion\r\n\r\nConvert Office documents and PDFs to Markdown while preserving structure.\r\n\r\n**Supported formats:**\r\n- PDF files (with optional Azure Document Intelligence integration)\r\n- Word documents (DOCX)\r\n- PowerPoint presentations (PPTX)\r\n- Excel spreadsheets (XLSX, XLS)\r\n\r\n**Basic usage:**\r\n```python\r\nfrom markitdown import MarkItDown\r\n\r\nmd = MarkItDown()\r\nresult = md.convert(\"document.pdf\")\r\nprint(result.text_content)\r\n```\r\n\r\n**Command-line:**\r\n```bash\r\nmarkitdown document.pdf -o output.md\r\n```\r\n\r\nSee `references/document_conversion.md` for detailed documentation on document-specific features.\r\n\r\n### 2. Media Processing\r\n\r\nExtract text from images using OCR and transcribe audio files to text.\r\n\r\n**Supported formats:**\r\n- Images (JPEG, PNG, GIF, etc.) with EXIF metadata extraction\r\n- Audio files with speech transcription (requires speech_recognition)\r\n\r\n**Image with OCR:**\r\n```python\r\nfrom markitdown import MarkItDown\r\n\r\nmd = MarkItDown()\r\nresult = md.convert(\"image.jpg\")\r\nprint(result.text_content) # Includes EXIF metadata and OCR text\r\n```\r\n\r\n**Audio transcription:**\r\n```python\r\nresult = md.convert(\"audio.wav\")\r\nprint(result.text_content) # Transcribed speech\r\n```\r\n\r\nSee `references/media_processing.md` for advanced media handling options.\r\n\r\n### 3. Web Content Extraction\r\n\r\nConvert web-based content and e-books to Markdown.\r\n\r\n**Supported formats:**\r\n- HTML files and web pages\r\n- YouTube video transcripts (via URL)\r\n- EPUB books\r\n- RSS feeds\r\n\r\n**YouTube transcript:**\r\n```python\r\nfrom markitdown import MarkItDown\r\n\r\nmd = MarkItDown()\r\nresult = md.convert(\"https://youtube.com/watch?v=VIDEO_ID\")\r\nprint(result.text_content)\r\n```\r\n\r\nSee `references/web_content.md` for web extraction details.\r\n\r\n### 4. Structured Data Handling\r\n\r\nConvert structured data formats to readable Markdown tables.\r\n\r\n**Supported formats:**\r\n- CSV files\r\n- JSON files\r\n- XML files\r\n\r\n**CSV to Markdown table:**\r\n```python\r\nfrom markitdown import MarkItDown\r\n\r\nmd = MarkItDown()\r\nresult = md.convert(\"data.csv\")\r\nprint(result.text_content) # Formatted as Markdown table\r\n```\r\n\r\nSee `references/structured_data.md` for format-specific options.\r\n\r\n### 5. Advanced Integrations\r\n\r\nEnhance conversion quality with AI-powered features.\r\n\r\n**Azure Document Intelligence:**\r\nFor enhanced PDF processing with better table extraction and layout analysis:\r\n```python\r\nfrom markitdown import MarkItDown\r\n\r\nmd = MarkItDown(docintel_endpoint=\"<endpoint>\", docintel_key=\"<key>\")\r\nresult = md.convert(\"complex.pdf\")\r\n```\r\n\r\n**LLM-Powered Image Descriptions:**\r\nGenerate detailed image descriptions using GPT-4o:\r\n```python\r\nfrom markitdown import MarkItDown\r\nfrom openai import OpenAI\r\n\r\nclient = OpenAI()\r\nmd = MarkItDown(llm_client=client, llm_model=\"gpt-4o\")\r\nresult = md.convert(\"presentation.pptx\") # Images described with LLM\r\n```\r\n\r\nSee `references/advanced_integrations.md` for integration details.\r\n\r\n### 6. Batch Processing\r\n\r\nProcess multiple files or entire ZIP archives at once.\r\n\r\n**ZIP file processing:**\r\n```python\r\nfrom markitdown import MarkItDown\r\n\r\nmd = MarkItDown()\r\nresult = md.convert(\"archive.zip\")\r\nprint(result.text_content) # All files converted and concatenated\r\n```\r\n\r\n**Batch script:**\r\nUse the provided batch processing script for directory conversion:\r\n```bash\r\npython scripts/batch_convert.py /path/to/documents /path/to/output\r\n```\r\n\r\nSee `scripts/batch_convert.py` for implementation details.",
"Common Workflows": "for file in documents/*.pdf; do\r\n markitdown \"$file\" -o \"markdown/$(basename \"$file\" .pdf).md\"\r\ndone\r\n```",
"Plugin System": "md = MarkItDown(enable_plugins=True)\r\n```"
}
}---
name: markitdown
description: Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing. Use when converting documents to markdown, extracting text from PDFs/Office files, transcribing audio, performing OCR on images, extracting YouTube transcripts, or processing batches of files. Supports 20+ formats including DOCX, XLSX, PPTX, PDF, HTML, EPUB, CSV, JSON, images with OCR, and audio with transcription.
---
# MarkItDown
## Overview
MarkItDown is a Python utility that converts various file formats into Markdown format, optimized for use with large language models and text analysis pipelines. It preserves document structure (headings, lists, tables, hyperlinks) while producing clean, token-efficient Markdown output.
## When to Use This Skill
Use this skill when users request:
- Converting documents to Markdown format
- Extracting text from PDF, Word, PowerPoint, or Excel files
- Performing OCR on images to extract text
- Transcribing audio files to text
- Extracting YouTube video transcripts
- Processing HTML, EPUB, or web content to Markdown
- Converting structured data (CSV, JSON, XML) to readable Markdown
- Batch converting multiple files or ZIP archives
- Preparing documents for LLM analysis or RAG systems
## Core Capabilities
### 1. Document Conversion
Convert Office documents and PDFs to Markdown while preserving structure.
**Supported formats:**
- PDF files (with optional Azure Document Intelligence integration)
- Word documents (DOCX)
- PowerPoint presentations (PPTX)
- Excel spreadsheets (XLSX, XLS)
**Basic usage:**
```python
from markitdown import MarkItDown
md = MarkItDown()
result = md.convert("document.pdf")
print(result.text_content)
```
**Command-line:**
```bash
markitdown document.pdf -o output.md
```
See `references/document_conversion.md` for detailed documentation on document-specific features.
### 2. Media Processing
Extract text from images using OCR and transcribe audio files to text.
**Supported formats:**
- Images (JPEG, PNG, GIF, etc.) with EXIF metadata extraction
- Audio files with speech transcription (requires speech_recognition)
**Image with OCR:**
```python
from markitdown import MarkItDown
md = MarkItDown()
result = md.convert("image.jpg")
print(result.text_content) # Includes EXIF metadata and OCR text
```
**Audio transcription:**
```python
result = md.convert("audio.wav")
print(result.text_content) # Transcribed speech
```
See `references/media_processing.md` for advanced media handling options.
### 3. Web Content Extraction
Convert web-based content and e-books to Markdown.
**Supported formats:**
- HTML files and web pages
- YouTube video transcripts (via URL)
- EPUB books
- RSS feeds
**YouTube transcript:**
```python
from markitdown import MarkItDown
md = MarkItDown()
result = md.convert("https://youtube.com/watch?v=VIDEO_ID")
print(result.text_content)
```
See `references/web_content.md` for web extraction details.
### 4. Structured Data Handling
Convert structured data formats to readable Markdown tables.
**Supported formats:**
- CSV files
- JSON files
- XML files
**CSV to Markdown table:**
```python
from markitdown import MarkItDown
md = MarkItDown()
result = md.convert("data.csv")
print(result.text_content) # Formatted as Markdown table
```
See `references/structured_data.md` for format-specific options.
### 5. Advanced Integrations
Enhance conversion quality with AI-powered features.
**Azure Document Intelligence:**
For enhanced PDF processing with better table extraction and layout analysis:
```python
from markitdown import MarkItDown
md = MarkItDown(docintel_endpoint="<endpoint>", docintel_key="<key>")
result = md.convert("complex.pdf")
```
**LLM-Powered Image Descriptions:**
Generate detailed image descriptions using GPT-4o:
```python
from markitdown import MarkItDown
from openai import OpenAI
client = OpenAI()
md = MarkItDown(llm_client=client, llm_model="gpt-4o")
result = md.convert("presentation.pptx") # Images described with LLM
```
See `references/advanced_integrations.md` for integration details.
### 6. Batch Processing
Process multiple files or entire ZIP archives at once.
**ZIP file processing:**
```python
from markitdown import MarkItDown
md = MarkItDown()
result = md.convert("archive.zip")
print(result.text_content) # All files converted and concatenated
```
**Batch script:**
Use the provided batch processing script for directory conversion:
```bash
python scripts/batch_convert.py /path/to/documents /path/to/output
```
See `scripts/batch_convert.py` for implementation details.
## Installation
**Full installation (all features):**
```bash
pip install 'markitdown[all]'
```
**Modular installation (specific features):**
```bash
pip install 'markitdown[pdf]' # PDF support
pip install 'markitdown[docx]' # Word support
pip install 'markitdown[pptx]' # PowerPoint support
pip install 'markitdown[xlsx]' # Excel support
pip install 'markitdown[audio]' # Audio transcription
pip install 'markitdown[youtube]' # YouTube transcripts
```
**Requirements:**
- Python 3.10 or higher
## Output Format
MarkItDown produces clean, token-efficient Markdown optimized for LLM consumption:
- Preserves headings, lists, and tables
- Maintains hyperlinks and formatting
- Includes metadata where relevant (EXIF, document properties)
- No temporary files created (streaming approach)
## Common Workflows
**Preparing documents for RAG:**
```python
from markitdown import MarkItDown
md = MarkItDown()
# Convert knowledge base documents
docs = ["manual.pdf", "guide.docx", "faq.html"]
markdown_content = []
for doc in docs:
result = md.convert(doc)
markdown_content.append(result.text_content)
# Now ready for embedding and indexing
```
**Document analysis pipeline:**
```bash
# Convert all PDFs in directory
for file in documents/*.pdf; do
markitdown "$file" -o "markdown/$(basename "$file" .pdf).md"
done
```
## Plugin System
MarkItDown supports extensible plugins for custom conversion logic. Plugins are disabled by default for security:
```python
from markitdown import MarkItDown
# Enable plugins if needed
md = MarkItDown(enable_plugins=True)
```
## Resources
This skill includes comprehensive reference documentation for each capability:
- **references/document_conversion.md** - Detailed PDF, DOCX, PPTX, XLSX conversion options
- **references/media_processing.md** - Image OCR and audio transcription details
- **references/web_content.md** - HTML, YouTube, and EPUB extraction
- **references/structured_data.md** - CSV, JSON, XML conversion formats
- **references/advanced_integrations.md** - Azure Document Intelligence and LLM integration
- **scripts/batch_convert.py** - Batch processing utility for directories