Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
kreuzberg-dev avatar

Picking A Format

  • 1 installs
  • 26 repo stars
  • Updated July 27, 2026
  • kreuzberg-dev/plugins

Maps the intended consumer (LLM, RAG store, parser, archive) to the right Kreuzberg --format and --content-format output combination.

About

Explains Kreuzberg's two orthogonal format knobs (--format and --content-format) plus token-reduction, with a decision tree by consumer. A developer uses it to choose text, markdown, djot, html, or JSON output for extracted documents.

  • Decision tree pairs --format and --content-format to LLM, RAG, parser, or archive
  • Token-reduction levels strip whitespace and boilerplate for token-tight pipelines

Picking A Format by the numbers

  • 1 all-time installs (skills.sh)
  • Ranked #565 of 688 Office & Documents skills by installs in the Skillselion catalog
  • Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/kreuzberg-dev/plugins --skill picking-a-format

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs1
repo stars26
Last updatedJuly 27, 2026
Repositorykreuzberg-dev/plugins

What it does

Maps the intended consumer (LLM, RAG store, parser, archive) to the right Kreuzberg --format and --content-format output combination.

Files

SKILL.mdMarkdownGitHub ↗

Picking a format

Kreuzberg has two orthogonal format knobs. Get them right up front and the downstream code stays simple.

KnobWhat it controlsValuesDefault
--formatHow the CLI prints the resulttext, jsontext (extract), json (batch)
--content-formatHow extracted content is rendered inside resultplain, markdown, djot, htmlplain
--token-reductionStrip whitespace / boilerplate for LLM contextsoff, light, moderate, aggressiveoff

--format json always returns the full ExtractionResult (content + metadata + tables + images). --format text prints just content. --content-format is what shows up inside that content field.

Decision tree

Who consumes the output?
├── LLM (Claude, GPT, Gemini, local) — embed/prompt context
│       --format text --content-format markdown
├── Vector store / RAG indexer
│       --format json --content-format markdown
│       (markdown preserves structure for chunking)
├── Downstream parser that expects machine-readable JSON
│       --format json --content-format plain
│       (cleanest text + structured metadata)
├── Human review / archival
│       --format text --content-format markdown
├── HTML re-rendering / web display
│       --format json --content-format html
├── Lossless intermediate for pandoc / academic tooling
│       --format json --content-format djot
└── Token-budget-constrained pipeline
        --format text --content-format plain
        (drops markup; add --token-reduction moderate for further savings)

Examples

Feed a PDF directly into an LLM:

kreuzberg extract paper.pdf --content-format markdown

Index a corpus into a RAG store with tables and headings preserved:

kreuzberg batch docs/*.pdf --format json --content-format markdown \
  | jq -c '.[] | {path: .metadata.path, content: .content, tables: .tables}'

Strip a file to bare text for a token-tight summarizer:

kreuzberg extract long.pdf \
  --content-format plain \
  --token-reduction moderate

Pull metadata only, ignore content:

kreuzberg extract file.pdf --format json | jq '.metadata'

When in doubt

  • Default to `markdown` as the content format. It is the best

compromise across LLMs, RAG, and human review, and Kreuzberg has the most faithful renderer for it.

  • Reach for plain only when downstream cannot tolerate any markup.
  • Reach for djot only if you're already in a djot/pandoc pipeline.
  • Reach for html only when re-rendering for the web.

Token-reduction (orthogonal)

--token-reduction collapses whitespace, strips repeated headers/footers, and trims boilerplate. It composes with any --content-format:

  • off (default), light, moderate, aggressive, maximum.

Use moderate as a safe starting point for LLM context windows. maximum is lossy — verify before relying on it.

See references/cli-reference.md for the full flag set and references/configuration.md for the equivalent output_format and token_reduction keys in kreuzberg.toml.

Related skills

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.