Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
mvanhorn avatar

Pp Firecrawl

  • 475 installs
  • 1.9k repo stars
  • Updated August 4, 2026
  • mvanhorn/printing-press-library

pp-firecrawl is an agent skill that drives the firecrawl-pp-cli Go binary to crawl sites and extract clean markdown through Firecrawl for developer web-ingestion workflows.

About

pp-firecrawl is a Printing Press Library agent skill authored by Hiten Shah under Apache-2.0 license that wraps the firecrawl-pp-cli command-line tool for Firecrawl scraping and crawling. Agents must verify the CLI is installed—via npx @mvanhorn/printing-press install firecrawl or go install of the firecrawl-pp-cli module—before executing firecrawl-pp-cli commands with the --agent flag. The skill routes empty input to --help, install mcp subcommands to MCP server setup, and all other queries to direct CLI invocation with subcommand help drilling. Allowed tools are Read and Bash, matching a shell-driven integration pattern for Claude Code, Codex, Cursor, and OpenClaw harnesses. Developers reach for pp-firecrawl when agents need site crawls, structured markdown extraction, or Firecrawl API tasks beyond brittle single-page HTTP fetches, optionally installing the Firecrawl MCP server for editor integrations.

  • Site-wide crawling
  • Markdown extraction
  • LLM-ready ingestion
  • Scrape automation

Pp Firecrawl by the numbers

  • 475 all-time installs (skills.sh)
  • +20 installs in the week ending Aug 4, 2026 (Skillselion tracking)
  • Ranked #1,832 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/mvanhorn/printing-press-library --skill pp-firecrawl

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs475
repo stars1.9k
Last updatedAugust 4, 2026
Repositorymvanhorn/printing-press-library

How do agents crawl websites with Firecrawl CLI?

Crawl and extract clean markdown from arbitrary sites via Firecrawl when agents need reliable web ingestion beyond single-page fetches.

Who is it for?

Agent workflows that need multi-page web crawling and markdown extraction through a verified Firecrawl CLI wrapper.

Skip if: Simple one-URL fetches or projects without Firecrawl API credentials—pp-firecrawl assumes Firecrawl service access and CLI installation.

When should I use this skill?

An agent needs to install firecrawl-pp-cli, crawl arbitrary sites, extract markdown, or set up the Firecrawl MCP server from Printing Press.

What you get

Installed firecrawl-pp-cli binary, executed crawl or scrape command output, and optional Firecrawl MCP server configuration.

  • Crawled markdown output
  • CLI command invocations
  • Optional MCP server install

By the numbers

  • Requires firecrawl-pp-cli Go binary as declared dependency
  • Licensed Apache-2.0 per skill front matter

Files

SKILL.mdMarkdownGitHub ↗

<!-- GENERATED FILE — DO NOT EDIT. This file is a verbatim mirror of library/developer-tools/firecrawl/SKILL.md, regenerated post-merge by tools/generate-skills/. Hand-edits here are silently overwritten on the next regen. Edit the library/ source instead. See the repository agent guide, section "Generated artifacts: registry.json, cli-skills/". -->

Firecrawl — Printing Press CLI

Prerequisites: Install the CLI

This skill drives the firecrawl-pp-cli binary. You must verify the CLI is installed before invoking any command from this skill. If it is missing, install it first:

1. Install via the Printing Press installer. It defaults binaries to $HOME/.local/bin on macOS/Linux and %LOCALAPPDATA%\Programs\PrintingPress\bin on Windows:

   npx -y @mvanhorn/printing-press-library install firecrawl --cli-only

2. Verify: firecrawl-pp-cli --version 3. Ensure the reported install directory is on $PATH for the agent/runtime that will invoke this skill.

If the npx install fails (no Node, offline, etc.), fall back to a direct Go install (requires Go 1.26.4 or newer):

go install github.com/mvanhorn/printing-press-library/library/developer-tools/firecrawl/cmd/firecrawl-pp-cli@latest

If --version reports "command not found" after install, the runtime cannot see the binary directory on $PATH. Do not proceed with skill commands until verification succeeds.

Command Reference

batch — Manage batch

  • firecrawl-pp-cli batch cancel-scrape — Cancel a batch scrape job
  • firecrawl-pp-cli batch get-scrape-errors — Get the errors of a batch scrape job
  • firecrawl-pp-cli batch get-scrape-status — Get the status of a batch scrape job
  • firecrawl-pp-cli batch scrape-and-extract-from-urls — Scrape multiple URLs and optionally extract information using an LLM

crawl — Manage crawl

  • firecrawl-pp-cli crawl cancel — Cancel a crawl job
  • firecrawl-pp-cli crawl get-active — Get all active crawls for the authenticated team
  • firecrawl-pp-cli crawl get-status — Get the status of a crawl job
  • firecrawl-pp-cli crawl urls — Crawl multiple URLs based on options

deep-research — Manage deep research

  • firecrawl-pp-cli deep-research get-status — Get the status and results of a deep research operation
  • firecrawl-pp-cli deep-research start — Start a deep research operation on a query

extract — Manage extract

  • firecrawl-pp-cli extract data — Extract structured data from pages using LLMs
  • firecrawl-pp-cli extract get-status — Get the status of an extract job

firecrawl-search — Manage firecrawl search

  • firecrawl-pp-cli firecrawl-search — Search and optionally scrape search results

llmstxt — Manage llmstxt

  • firecrawl-pp-cli llmstxt generate-llms-txt — Generate LLMs.txt for a website
  • firecrawl-pp-cli llmstxt get-llms-txt-status — Get the status and results of an LLMs.txt generation job

map — Manage map

  • firecrawl-pp-cli map — Map multiple URLs based on options

scrape — Manage scrape

  • firecrawl-pp-cli scrape — Scrape a single URL and optionally extract information using an LLM

team — Manage team

  • firecrawl-pp-cli team get-credit-usage — Get remaining credits for the authenticated team
  • firecrawl-pp-cli team get-token-usage — Get remaining tokens for the authenticated team (Extract only)

Finding the right command

When you know what you want to do but not which command does it, ask the CLI directly:

firecrawl-pp-cli which "<capability in your own words>"

which resolves a natural-language capability query to the best matching command from this CLI's curated feature index. Exit code 0 means at least one match; exit code 2 means no confident match — fall back to --help or use a narrower query.

Auth Setup

Store your access token:

firecrawl-pp-cli auth set-token YOUR_TOKEN_HERE

Or set FIRECRAWL_BEARER_AUTH as an environment variable.

Run firecrawl-pp-cli doctor to verify setup.

Agent Mode

Add --agent to any command. Expands to: --json --compact --no-input --no-color --yes.

  • Pipeable — JSON on stdout, errors on stderr
  • Filterable--select keeps a subset of fields. Dotted paths descend into nested structures; arrays traverse element-wise. Critical for keeping context small on verbose APIs:
  firecrawl-pp-cli batch cancel-scrape mock-value --agent --select id,name,status
  • Previewable--dry-run shows the request without sending
  • Offline-friendly — sync/search commands can use the local SQLite store when available
  • Non-interactive — never prompts, every input is a flag

Response envelope

Commands that read from the local store or the API wrap output in a provenance envelope:

{
  "meta": {"source": "live" | "local", "synced_at": "...", "reason": "..."},
  "results": <data>
}

Parse .results for data and .meta.source to know whether it's live or local. A human-readable N results (live) summary is printed to stderr only when stdout is a terminal — piped/agent consumers get pure JSON on stdout.

Agent Feedback

When you (or the agent) notice something off about this CLI, record it:

firecrawl-pp-cli feedback "the --since flag is inclusive but docs say exclusive"
firecrawl-pp-cli feedback --stdin < notes.txt
firecrawl-pp-cli feedback list --json --limit 10

Entries are stored locally at ~/.firecrawl-pp-cli/feedback.jsonl. They are never POSTed unless FIRECRAWL_FEEDBACK_ENDPOINT is set AND either --send is passed or FIRECRAWL_FEEDBACK_AUTO_SEND=true. Default behavior is local-only.

Write what surprised you, not a bug report. Short, specific, one line: that is the part that compounds.

Output Delivery

Every command accepts --deliver <sink>. The output goes to the named sink in addition to (or instead of) stdout, so agents can route command results without hand-piping. Three sinks are supported:

SinkEffect
stdoutDefault; write to stdout only
file:<path>Atomically write output to <path> (tmp + rename)
webhook:<url>POST the output body to the URL (application/json or application/x-ndjson when --compact)

Unknown schemes are refused with a structured error naming the supported set. Webhook failures return non-zero and log the URL + HTTP status on stderr.

Named Profiles

A profile is a saved set of flag values, reused across invocations. Use it when a scheduled agent calls the same command every run with the same configuration - HeyGen's "Beacon" pattern.

firecrawl-pp-cli profile save briefing --json
firecrawl-pp-cli --profile briefing batch cancel-scrape mock-value
firecrawl-pp-cli profile list --json
firecrawl-pp-cli profile show briefing
firecrawl-pp-cli profile delete briefing --yes

Explicit flags always win over profile values; profile values win over defaults. agent-context lists all available profiles under available_profiles so introspecting agents discover them at runtime.

Exit Codes

CodeMeaning
0Success
2Usage error (wrong arguments)
3Resource not found
4Authentication required
5API error (upstream issue)
7Rate limited (wait and retry)
10Config error

Argument Parsing

Parse $ARGUMENTS:

1. Empty, `help`, or `--help` → show firecrawl-pp-cli --help output 2. Starts with `install` → ends with mcp → MCP installation; otherwise → see Prerequisites above 3. Anything else → Direct Use (execute as CLI command with --agent)

MCP Server Installation

1. Install the MCP server:

   go install github.com/mvanhorn/printing-press-library/library/other/firecrawl-pp-cli/cmd/firecrawl-pp-mcp@latest

2. Register with Claude Code:

   claude mcp add firecrawl-pp-mcp -- firecrawl-pp-mcp

3. Verify: claude mcp list

Direct Use

1. Check if installed: which firecrawl-pp-cli If not found, offer to install (see Prerequisites at the top of this skill). 2. Match the user query to the best command from the Unique Capabilities and Command Reference above. 3. Execute with the --agent flag:

   firecrawl-pp-cli <command> [subcommand] [args] --agent

4. If ambiguous, drill into subcommand help: firecrawl-pp-cli <command> --help.

Related skills

FAQ

How do you install pp-firecrawl dependencies?

pp-firecrawl requires the firecrawl-pp-cli binary, installable with npx @mvanhorn/printing-press install firecrawl or go install github.com/mvanhorn/printing-press-library/.../firecrawl-pp-cli@latest, then verified with firecrawl-pp-cli --version.

How does pp-firecrawl invoke Firecrawl commands?

pp-firecrawl maps user intent to firecrawl-pp-cli subcommands and executes firecrawl-pp-cli <command> [args] --agent after confirming the CLI is on PATH, drilling into --help when arguments are ambiguous.

Can pp-firecrawl install an MCP server?

Yes—pp-firecrawl treats install commands ending with mcp as MCP server installation flows, separate from CLI-only setup via npx @mvanhorn/printing-press install firecrawl --cli-only.

AI & Agent Buildingagentsautomation

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.