Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
kreuzberg-dev avatar

Format Specific Extraction

  • 11 installs
  • 8.9k repo stars
  • Updated August 4, 2026
  • kreuzberg-dev/kreuzberg

Helps with ai & agent building tasks.

About

format-specific-extraction is a Claude Code skill for ai & agent building. It helps solo builders move faster with AI-assisted development.

  • format-specific-extraction
  • AI & Agent Building
  • AI-coding skill

Format Specific Extraction by the numbers

  • 11 all-time installs (skills.sh)
  • Ranked #11,769 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/kreuzberg-dev/kreuzberg --skill format-specific-extraction

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs11
repo stars8.9k
Last updatedAugust 4, 2026
Repositorykreuzberg-dev/kreuzberg

What it does

Helps with ai & agent building tasks.

Files

SKILL.mdMarkdownGitHub ↗

Format-Specific Extraction Workflows

Office XML (DOCX/PPTX/ODT)

ZIP archive → Security validation → XML parsing → Text + tables + metadata

1. ZipBombValidator::new(limits).validate(&mut archive)? 2. Extract XML files from archive (word/document.xml, ppt/slides/*.xml, content.xml) 3. Parse with quick-xml::Reader (streaming) + DepthValidator + StringGrowthValidator 4. Extract metadata via crate::extraction::office_metadata::extract_metadata() 5. See: extractors/docx.rs, extractors/pptx.rs, extractors/odt.rs

PDF

Bytes → pdf_oxide → Per-page text + OCR fallback → Tables → Metadata

1. pdf_oxide::PdfDocument::from_bytes(content)? 2. Check if needs OCR: config.force_ocr || !has_searchable_text() 3. Extract text per page, tables if config.pages enabled 4. Feature-gated: #[cfg(feature = "pdf")] 5. See: extractors/pdf/mod.rs

Archives (ZIP/TAR/7z/GZIP)

Validate → Extract metadata → Extract plaintext files only

1. ZipBombValidator BEFORE any extraction 2. Extract metadata (file list, sizes) 3. Extract text content from plaintext files 4. Use build_archive_result() helper 5. See: extractors/archive.rs, extraction/archive/*.rs

Structured Text (JSON/YAML/TOML/XML)

Detect format from MIME → Parse → Pretty-print → Metadata

Single StructuredExtractor handles multiple MIME types. Parse with format-specific library, pretty-print to text. See: extractors/structured.rs

Email (EML/MSG)

Parse headers → Extract body (text/html) → Process attachments

See: extraction/email.rs, extractors/email.rs

Common Helpers

HelperLocationPurpose
office_metadata::extract_metadata()extraction/office.rsOffice XML metadata
cells_to_markdown()extraction/mod.rsConvert cell grid to GFM table
build_archive_result()extraction/archive/mod.rsStandard archive result

Adding a New Format

1. Add MIME type to EXT_TO_MIME in core/mime.rs 2. Create extractor implementing DocumentExtractor trait 3. Set supported_mime_types() and priority() (default: 50) 4. Register in extractors/mod.rsregister_default_extractors() 5. Feature-gate if optional: #[cfg(feature = "my-format")] 6. Apply security validators for user content 7. Add tests with fixture files

Related skills

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.