
Picking A Format
- 1 installs
- 26 repo stars
- Updated July 27, 2026
- kreuzberg-dev/plugins
Maps the intended consumer (LLM, RAG store, parser, archive) to the right Kreuzberg --format and --content-format output combination.
About
Explains Kreuzberg's two orthogonal format knobs (--format and --content-format) plus token-reduction, with a decision tree by consumer. A developer uses it to choose text, markdown, djot, html, or JSON output for extracted documents.
- Decision tree pairs --format and --content-format to LLM, RAG, parser, or archive
- Token-reduction levels strip whitespace and boilerplate for token-tight pipelines
Picking A Format by the numbers
- 1 all-time installs (skills.sh)
- Ranked #565 of 688 Office & Documents skills by installs in the Skillselion catalog
- Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/kreuzberg-dev/plugins --skill picking-a-formatAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 1 |
|---|---|
| repo stars | ★ 26 |
| Last updated | July 27, 2026 |
| Repository | kreuzberg-dev/plugins ↗ |
What it does
Maps the intended consumer (LLM, RAG store, parser, archive) to the right Kreuzberg --format and --content-format output combination.
Files
Picking a format
Kreuzberg has two orthogonal format knobs. Get them right up front and the downstream code stays simple.
| Knob | What it controls | Values | Default |
|---|---|---|---|
--format | How the CLI prints the result | text, json | text (extract), json (batch) |
--content-format | How extracted content is rendered inside result | plain, markdown, djot, html | plain |
--token-reduction | Strip whitespace / boilerplate for LLM contexts | off, light, moderate, aggressive | off |
--format json always returns the full ExtractionResult (content + metadata + tables + images). --format text prints just content. --content-format is what shows up inside that content field.
Decision tree
Who consumes the output?
├── LLM (Claude, GPT, Gemini, local) — embed/prompt context
│ --format text --content-format markdown
├── Vector store / RAG indexer
│ --format json --content-format markdown
│ (markdown preserves structure for chunking)
├── Downstream parser that expects machine-readable JSON
│ --format json --content-format plain
│ (cleanest text + structured metadata)
├── Human review / archival
│ --format text --content-format markdown
├── HTML re-rendering / web display
│ --format json --content-format html
├── Lossless intermediate for pandoc / academic tooling
│ --format json --content-format djot
└── Token-budget-constrained pipeline
--format text --content-format plain
(drops markup; add --token-reduction moderate for further savings)Examples
Feed a PDF directly into an LLM:
kreuzberg extract paper.pdf --content-format markdownIndex a corpus into a RAG store with tables and headings preserved:
kreuzberg batch docs/*.pdf --format json --content-format markdown \
| jq -c '.[] | {path: .metadata.path, content: .content, tables: .tables}'Strip a file to bare text for a token-tight summarizer:
kreuzberg extract long.pdf \
--content-format plain \
--token-reduction moderatePull metadata only, ignore content:
kreuzberg extract file.pdf --format json | jq '.metadata'When in doubt
- Default to `markdown` as the content format. It is the best
compromise across LLMs, RAG, and human review, and Kreuzberg has the most faithful renderer for it.
- Reach for
plainonly when downstream cannot tolerate any markup. - Reach for
djotonly if you're already in a djot/pandoc pipeline. - Reach for
htmlonly when re-rendering for the web.
Token-reduction (orthogonal)
--token-reduction collapses whitespace, strips repeated headers/footers, and trims boilerplate. It composes with any --content-format:
off(default),light,moderate,aggressive,maximum.
Use moderate as a safe starting point for LLM context windows. maximum is lossy — verify before relying on it.
See references/cli-reference.md for the full flag set and references/configuration.md for the equivalent output_format and token_reduction keys in kreuzberg.toml.