
Apify Link Prospecting Outreach
- 146 installs
- 239 repo stars
- Updated June 29, 2026
- apify/awesome-skills
apify-link-prospecting-outreach is a Claude skill that finds SERP prospects, scores them with Ahrefs, and drafts matched link-building outreach emails and placements.
About
This skill turns a goal, a target keyword, and a URL into a tiered link-building outreach list. It finds sites ranking for the keyword, scores each prospect with Ahrefs authority and traffic, picks the strongest pitch angle per row, drafts an outreach-type-matched email, and proposes an in-article link placement. A developer uses it to run cold-email link-building or recover unlinked brand mentions.
- Finds SERP-ranking prospects and scores each with Ahrefs domain and page-level metrics
- Assigns the strongest pitch angle per prospect and drafts a matched outreach email
- Proposes a concrete in-article link placement as three artifacts and exports to xlsx
Apify Link Prospecting Outreach by the numbers
- 146 all-time installs (skills.sh)
- Ranked #1,058 of 1,879 Marketing & SEO skills by installs in the Skillselion catalog
- Data as of Aug 4, 2026 (Skillselion catalog sync)
apify-link-prospecting-outreach capabilities & compatibility
Requires an Apify account plus Ahrefs MCP; actor runs and Ahrefs API units consume credits per campaign.
- Capabilities
- seo outreach · link building · web scraping · copywriting
- Use cases
- seo · marketing · email · web scraping · copywriting
- Pricing
- Bring your own API key
What apify-link-prospecting-outreach says it does
Turn a goal + a target keyword + a URL the user wants to promote into a tiered, ready-to-send outreach list: SERP-ranking prospects with Ahrefs-scored authority
Ahrefs MCP available (the skill calls `mcp__claude_ai_Ahrefs__*` tools for prospect scoring)
npx skills add https://github.com/apify/awesome-skills --skill apify-link-prospecting-outreachAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 146 |
|---|---|
| repo stars | ★ 239 |
| Last updated | June 29, 2026 |
| Repository | apify/awesome-skills ↗ |
What it does
Build a tiered, Ahrefs-scored link-building prospect list with matched outreach emails and placements.
Who is it for?
Building a tiered outreach list for SEO link building, unlinked-mention recovery, or competitor-link replacement.
Skip if: General web scraping unrelated to link building or SEO outreach.
When should I use this skill?
The user asks to find link-building opportunities, prospect link partners, recover unlinked mentions, or run cold-email SEO outreach.
What you get
A tiered xlsx of SERP prospects with Ahrefs scores, a pitch angle, a matched email, and a link placement per row.
- Tiered prospect xlsx with pitch angle, outreach email, and link placement per row
By the numbers
- 9-step workflow checklist
- 3-artifact link placement per row
- 5 campaign goal presets
Files
Link Prospecting Outreach
Turn a goal + a target keyword + a URL the user wants to promote into a tiered, ready-to-send outreach list: SERP-ranking prospects with Ahrefs-scored authority, the strongest pitch angle per prospect, an outreach-type-matched email draft, and a copy-paste-ready link placement.
Prerequisites
(No need to check it upfront)
.envfile withAPIFY_TOKEN- Ahrefs MCP available (the skill calls
mcp__claude_ai_Ahrefs__*tools for prospect scoring) - Node.js 20.6+ (for native
--env-filesupport) - One-time setup inside the skill's
scripts/folder:npm install
Helper scripts (one config, four steps)
After Step 1–2 inputs are collected, write them to a single `campaign.json` (schema in `campaign.json.example`). Every downstream script reads --config campaign.json, so the agent doesn't fork per-campaign copies. Sequence:
# 1. Run the Actor (writes {base}.json + sub-Actor sidecars when --fetch-sub-datasets)
node --env-file=.env scripts/run_actor.js --actor "apify/link-prospecting-tool" --input '<json>' --timeout 1800 --fetch-sub-datasets --output {base}.json --format json
# 2. Build unified prospect table from the sidecars
python3 scripts/build_prospects.py --config campaign.json
# 3. (After Step 5 Ahrefs MCP calls → save to {base}_ahrefs_domain.json + {base}_ahrefs_page.json)
python3 scripts/enrich_prospects.py --config campaign.json
# 4. (After Step 8 sub-agents write outputs to /tmp/placement_outputs/row_*.json)
python3 scripts/merge_subagent_outputs.py --config campaign.json --outputs-dir /tmp/placement_outputs
# 5. Write the final xlsx + metadata sidecar
python3 scripts/write_xlsx.py --config campaign.jsonIf the runner's client-side wait elapses with the Actor still running on Apify, use scripts/fetch_run_artifacts.js --run-id <id> --output {base}.json instead of restarting. If the parent run is missing SUB_ACTOR_RESULTS (post-2026-05-20 Actor schema), scripts/fetch_subactors_from_log.js resolves sub-Actor runIds from the parent log.
Workflow
Copy this checklist and track progress:
Task Progress:
- [ ] Step 1: Collect required anchor inputs incl. goal (block on these)
- [ ] Step 2: Collect brand voice, partnership type, output format
- [ ] Step 3: Run apify/link-prospecting-tool
- [ ] Step 4: Pull leads, mentions, authors, and sub-Actor datasets
- [ ] Step 5: Enrich every domain with Ahrefs metrics, assign Prospect Tier
- [ ] Step 6: Run skip pass — flag rows to drop before drafting
- [ ] Step 7: Compute "Why This Prospect" tag per surviving row
- [ ] Step 8: Compose per-row 3-artifact placement + outreach-type-aware email
- [ ] Step 9: Render output in chosen formatStep 1: Required Anchor Inputs (ask FIRST, before anything else)
Do NOT proceed to Step 2 until every required input is answered. Surface them as the very first interaction. The dedup input (#7) is optional but must still be explicitly asked.
1. Concrete goal for this campaign — pick one preset or supply custom text. The goal drives skip-pass filtering, outreach-type template selection, and Prospect Tier thresholds. Required.
| Preset | Effect downstream |
|---|---|
Recover unlinked brand mentions | Skip pass drops every row where brand_mentioned_in_source is false. Default outreach type = unlinked-mention-claim. |
Replace competitor links | Skip pass drops every row not tagged Links to competitor. Default outreach type = competitor-link-replacement. |
Topical authority links to specific URL | No filter. Tier thresholds tighten (DR ≥ 50 for tier A). Default outreach type chosen per-row from Why This Prospect. |
Maximum link volume from any relevant site | No filter. Tier thresholds relax (DR ≥ 30 for tier A). Default outreach type chosen per-row. |
Custom | User-supplied paragraph; biases email tone and tier weights. No automatic skip filter. |
2. Target keyword(s) — one or more keywords the user wants their link to appear next to. The skill prospects the SERP for each. At least one required. 3. Brand name — the user's brand or product name. The Actor will not run without this (it is the brand input field). 4. Product/category description — one or two sentences describing what the user sells, who they sell to, and what category their product fits in. Example: "Apify — web scraping platform that runs serverless scrapers as APIs. We sell to developers and data teams who need scraped data without managing infrastructure." Required. Used in Step 6 (topical-fit gate) and Step 7 (adversarial-mention detection) to recognise prospects who are in the same product category — those won't link no matter the pitch. Without this, the skill cannot distinguish a genuine editorial opportunity from a competitor's blog. 5. URL of content to link to — the destination URL that will be inserted into partner articles. Required. 6. Competitors — anyone in the user's product category who would publish a "ours vs theirs" comparison page on their own site. Frame the ask this way explicitly: "List every company that would write an X-vs-YourBrand comparison page. These won't link to you no matter what — small competitors count too." Encourage 10+ entries; most users default to listing 3–5 obvious ones and miss the long tail. Mapped to competitorDomains on the Actor and reused in Steps 6 (adversarial-mention skip) and 7 (Links to competitor Why-tag).
After the user answers, offer (do not push) an Ahrefs auto-pull of organic competitors: "Want me to pull your top organic competitors from Ahrefs and add them to this list? Adds ~50 API units and surfaces smaller competitors you may have missed." If the user says yes and Ahrefs MCP is available, call mcp__claude_ai_Ahrefs__site-explorer-organic-competitors on the user's domain (extracted from input #5) and merge results into competitorDomains. If Ahrefs is unavailable or the user declines, proceed with the user-supplied list only. 7. Already-pitched domains (optional) — domains the user has already contacted in past campaigns. Accept a comma-separated list, a CSV/Sheet path, or "none". The skill drops these in the skip pass so the user doesn't double-pitch. Not required to proceed. 8. Number of organic results per keyword — how many Google organic SERP results to prospect per keyword. Default 10 if the user is unsure, but ask the question so the user knows the lever exists. Mapped to organicResult. 9. LLM sources to track — multi-select. Each enabled engine queries an additional AI search/chat surface and adds Google Search Scraper sub-Actor cost per result fetched. Default: all enabled. Mapping to Actor input flags:
| Option | Actor flag | Cost impact |
|---|---|---|
| ChatGPT Search | enableChatGpt | Per-result Google Search Scraper cost |
| Gemini | enableGemini | Per-result Google Search Scraper cost |
| Copilot (Microsoft / Bing) | enableCopilot | Per-result Google Search Scraper cost |
| Perplexity | enablePerplexity | Per-result Google Search Scraper cost |
| Google AI Mode | enableAiMode | Per-result Google Search Scraper cost |
| Google AI Overviews | enableAiOverviews | Free — parsed from the SERP already fetched. Keep on regardless of budget. |
Surface the multi-select to the user with all six pre-checked. Disabling individual engines is the main cost-cutting lever short of dropping organicResult — recommend keeping ChatGPT + Gemini on at minimum (they capture the largest share of LLM-driven discovery traffic in 2026). 10. Run email verification? — boolean. Default: yes. Mapped to enableEmailVerification on the Actor. When enabled, the Actor verifies every email returned by the Contact Details Scraper sub-Actor and tags each lead with a verification status (verified / catch-all / risky / invalid / unknown). The skill uses the status in Step 6 (invalid emails get auto-skipped) and surfaces it as the Email Verification column in the output. Disable only if the user is rate-limited on verification quota or running cost-tight smoke tests.
Once 1–6 and 8–10 are captured (7 is optional), move on.
Step 2: Secondary Inputs
Ask these next:
1. Brand info and voice — a short paragraph describing the product/brand and the tone for outreach (e.g., "casual and helpful", "formal B2B", "founder-led"). Used verbatim to shape every generated email. 2. Partnership type — the offer the user is willing to make. Determines the offer paragraph substituted into the per-row email. Outreach-type template selection happens separately, per-row, in Step 8.
| Option | What it offers |
|---|---|
| ABC link exchange | Three-way link swap: partner links to user, user links to a third party, third party links to partner. |
| Direct A B link exchange | Two-way link swap: partner links to user, user links to partner. |
| Resource page / list inclusion | Ask to be added to an existing curated list or roundup. No reciprocal link offered. |
| Unilateral ask (no reciprocal) | User asks for the link without offering anything in return — appropriate for unlinked-mention claims and broken-link replacements. |
| Other | User types their own offer (paid placement, free product, co-authored content, etc.). |
3. Output format:
| Format | Behavior |
|---|---|
xlsx | run_actor.js writes a styled spreadsheet to disk. |
markdown | Agent renders the table inline in chat with email drafts beneath each row. |
Step 3: Run the Actor
The Actor ID is apify/link-prospecting-tool. Full input schema lives in reference/apify-actor-usage.md.
Recommended call payload for this skill (defaults chosen for outreach-first workflow):
{
"queries": "<keyword 1>\n<keyword 2>",
"brand": "<user's brand name>",
"ownDomains": ["<user-domain.com>"],
"competitorDomains": [],
"ignoreDomains": [
"wikipedia.org", "github.com", "stackoverflow.com", "stackexchange.com",
"reddit.com", "quora.com", "youtube.com", "twitter.com", "x.com",
"linkedin.com", "facebook.com", "medium.com", "archive.org",
"chromewebstore.google.com", "addons.mozilla.org", "apps.apple.com",
"play.google.com", "microsoftedge.microsoft.com", "marketplace.visualstudio.com"
],
"organicResult": 10,
"maxContactsPerDomain": 3,
"department": ["marketing"],
"searchAuthorName": true,
"includeMention": true,
"enableChatGpt": true,
"enableGemini": true,
"enableCopilot": true,
"enablePerplexity": true,
"enableAiMode": true,
"enableAiOverviews": true,
"enableEmailVerification": true
}The six enable* LLM-source flags map 1:1 to the user's Step 1 input #9 multi-select. Pass false for any engine the user deselected. enableEmailVerification maps to Step 1 input #10.
The ignoreDomains default includes two groups:
- Giants and UGC (wikipedia, github, stackoverflow, reddit, etc.) — too broad to pitch as editorial partners.
- App / extension marketplaces (Chrome Web Store, Firefox Add-ons, Apple/Google Play, VS Code Marketplace, etc.) — product directory listings, no editorial decision-makers.
Do NOT auto-add to ignoreDomains (let the user decide):
- UGC/community sites like
kaggle.com,dev.to,substack.com,producthunt.com,g2.com,capterra.com,trustpilot.com— some users get real value pitching these. - API directories like
rapidapi.com,programmableweb.com,publicapis.dev— relevant for some products (especially developer-tool brands), irrelevant for others. Surface these as candidates only if the user wants to add them.
The URL-pattern skip rules in Step 6 catch the per-row noise (subdomain prefixes, path patterns) that ignoreDomains can't express.
department defaults to ["marketing"] only. The skill prioritises editorial-leaning contacts within the returned marketing department during row composition (see Step 8). Only add sales if the user explicitly wants BD-style partnership pitches. Only add c_suite if the prospect domains are very small (1–5 person shops) where the founder may also be the editor.
Call the runner script:
node --env-file=.env ${CLAUDE_PLUGIN_ROOT}/scripts/run_actor.js \
--actor "apify/link-prospecting-tool" \
--input 'JSON_INPUT' \
--timeout 1800 \
--fetch-sub-datasets \
--output YYYY-MM-DD_outreach.json \
--format jsonNotes:
--timeout 1800is the recommended client-side wait. The Actor itself runs 15-50+ min depending on keyword count, LLM-engine fan-out, andenableEmailVerification. Past calibration runs land in the 20–55 min range. Bumping the default avoids the partial-result situation where the runner gives up but the Actor keeps going.- If the client-side wait still elapses with the Actor still running on Apify (status
RUNNINGorREADYwhen the runner exits), do not restart the Actor. Usescripts/fetch_run_artifacts.js --run-id <id> --output <file>to poll the existing run and download all artifacts — same output shape asrun_actor.js --fetch-sub-datasets. --fetch-sub-datasetsdownloads sibling files alongside the main output:*_mentions.json,*_authors.json,*_serp.json,*_wcc.json. You need all of them to populate every output column.
Step 4: Access All Datasets
The Actor's output schema changed on or before 2026-05-20. The build_prospects script must handle the new shape; older skill versions that joined a separate MENTIONS dataset are broken.
Current schema (verified 2026-05-20):
| File written by runner / fetcher | Source | Populates |
|---|---|---|
*_output.json (main) | "All leads" dataset | Contact Full Name, Contact Job Title, Department, Seniority, Contact Email, Email Verification (when enableEmailVerification: true), Contact LinkedIn, Company, Domain. Each lead's `source_url[]` array contains the article URLs that produced this contact, each with a `brand_mentioned_in_source` boolean — this is the new home of the per-(URL, contact) mention data. |
*_serp.json | Google Search Results Scraper sub-Actor (one item per (query × engine) combination) | SERP Position, Article Title, Publish Date (via organicResults[]), and engine attribution per URL (Google Organic, ChatGPT, Gemini, Copilot, Perplexity, Google AI Mode) by joining aiModeResult.sources[], perplexitySearchResult.sources[], chatGptSearchResult.sources[], geminiSearchResult.sources[], copilotSearchResult.sources[]. URLs from ChatGPT carry a ?utm_source=chatgpt.com query suffix — normalise URLs (strip tracking params) before joining. |
*_wcc.json | Website Content Crawler sub-Actor | Placement Source Sentence, Placement With Link, Placement New Insertion, Article Author cross-check, outbound-link inspection for Links to competitor and Resource / roundup page tags. Canonical URL list for building rows — every URL that got body-crawled appears here, including ones that didn't yield a lead. |
*_authors.json | AI Web Scraper sub-Actor (when searchAuthorName: true) | Article Author, Author Source (set to searchAuthorName). Note: this sub-Actor frequently TIMES-OUT at its 300s default — partial results are still saved. |
What changed (vs. pre-2026-05-20 runs): 1. No separate MENTIONS / AUTHORS / DOMAINS_WITH_LEADS named datasets — mention info is folded into main_leads[i].source_url[]. 2. No SUB_ACTOR_RESULTS record in the parent run's key-value store. Sub-Actor runIds are now only discoverable from the parent run log via regex \[apify\.<slug> runId:([A-Za-z0-9]+)\]. The runner script's --fetch-sub-datasets flag now falls back to log-parsing when the KV index is missing; the standalone scripts/fetch_subactors_from_log.js does the same for runs whose runner already exited. 3. The mentions schema reduced: source_url[i] carries only {domain, brand_mentioned_in_source, url} — no per-engine flags like the old ChatGPT_mention / Perplexity_mention. Engine attribution must be reconstructed from the SERP sub-dataset's LLM-result sub-fields (see SERP row above).
If a column's source is missing, write "Not found" and add a manual-lookup hint in Notes. Never fabricate.
Step 5: Ahrefs Enrichment and Prospect Tier
For every unique domain that survived the Actor's filtering, fetch authority and traffic metrics via Ahrefs MCP. Call all three tools in parallel per domain (and across domains — batch parallelise to keep this step under a minute for typical 20–50 prospect lists):
| Ahrefs tool | Used for | Column it populates |
|---|---|---|
mcp__claude_ai_Ahrefs__site-explorer-domain-rating (target = domain) | Domain Rating | Domain DR |
mcp__claude_ai_Ahrefs__site-explorer-metrics (target = article URL, mode = exact) | Page-level organic traffic (last 30 days) | Page Traffic |
mcp__claude_ai_Ahrefs__site-explorer-backlinks-stats (target = domain) | Referring domains count | Referring Domains |
If Ahrefs returns no data (domain not indexed, page too new), set the column to "-" and add a Notes hint "Ahrefs has no data — verify manually before pitching". Do not fabricate values.
Assign Prospect Tier using the thresholds matching the user's goal:
| Goal | Tier A | Tier B | Tier C |
|---|---|---|---|
Topical authority links to specific URL | DR ≥ 50 AND Page Traffic ≥ 300/mo | DR 30–49 OR Page Traffic 50–299 | everything below |
Maximum link volume from any relevant site | DR ≥ 30 AND Page Traffic ≥ 100/mo | DR 15–29 OR Page Traffic 20–99 | everything below |
Recover unlinked brand mentions | irrelevant — every mention is worth claiming; tier by DR alone (≥ 40 = A, 20–39 = B, < 20 = C) | ||
Replace competitor links | tier by DR (≥ 50 = A, 30–49 = B, < 30 = C) | ||
Custom | use the Topical authority thresholds |
Surface tier breakdown to the user before Step 8 — let them confirm whether to draft emails for all tiers or only A/B.
Step 6: Skip Pass
Before drafting any email, walk every row and apply skip rules. Skipped rows get Outreach Status = "Skip", a one-line reason in Notes, and no email or placement is generated (saves tokens and user review time).
Skip rules (in order):
1. Goal mismatch. If the goal is Recover unlinked brand mentions and the row's Mentions data shows brand_mentioned_in_source: false, skip. If the goal is Replace competitor links and the row's WCC body has no outbound link to any competitorDomains entry, skip. 2. Already pitched. If the row's domain matches an entry in the optional already-pitched list from Step 1 input #7, skip. 3. Own / competitor domain leak. The Actor should already filter these, but double-check — if the row's domain matches ownDomains or competitorDomains, skip. 4. Stale content. If Publish Date is older than 5 years, skip (low chance the editor will update the post). 5. URL-pattern skip. Skip rows whose URL matches any of these patterns:
- Subdomain prefixes:
developers.*,docs.*,support.*,helpcenter.*,legacy.*,dsarequests.*,connectivity.*,community.*,dev.*(when used as a doc subdomain — e.g.dev.example.com/api/),api.*only when followed by a path that's clearly documentation (/reference/,/docs/,/spec/). Do NOT skiprapidapi.comor other API-directory domains by this rule alone —api.*is a subdomain check, not a substring check. - Path patterns:
/api-docs/,/reference/,/marketplace/,/extensions/,/profile/,/users/,/free-tools/,/spec/,/content/privacy,/content/terms,/content/dma,/content/how_we_work,/legal/,/_redirects,/sitemap. - Vendor product page patterns: URL ends in
-scraper.php,-scraping.php, contains-data-scraper.,-data-scraping.,/bots/,/extension/,/detail/(extension detail pages).
6. Non-editorial page type. Inspect the WCC page body. Skip vendor product pages, pricing pages, login walls, sign-up pages, terms/legal pages, and pages with fewer than 400 words of body text. Word count <400 is the threshold — most editorial articles are 800+ words. 7. UGC slipped through. If the page URL contains /forum/, /thread/, /comments/, /answers/, /q/, /topic/, /discussion/, or the WCC body is structured as discussion replies, skip. 8. Category-fit gate (loose). Extract 4–6 category keywords from the user's product description (Step 1 input #4) — these describe the product category, not the specific subject of the user's URL. Examples for a web-scraping product: scrape, scraping, scraper, crawl, extract, data extraction. For a CMS product: cms, headless, content, editorial. The row's WCC body must contain at least 1 of these category keywords. If not, skip with reason Article isn't in user's product category (no '<kw>' match) — kills recipe blogs, finance articles, and other off-category content that slipped through SERP filtering.
For non-English campaigns, include both source-language and English keywords in the category set — many Czech/German/French articles cite English brand names and product categories inline. Example for a Czech water-filtration brand: {filtr, filtrace, vod, voda, filter, filtration, water}. A pure-Czech keyword set would miss articles by Czech authors who write in mixed CS/EN.
Known false negatives this rule can't catch (the per-row sub-agent in Step 8 must catch them):
- Local e-commerce competitors selling the exact same product category. Past campaigns have seen multiple regional e-shops survive the mechanical pass — typically platform-based stores (e.g. Shoptet, Shopify) with "add to cart" buttons embedded in the article body. The sub-agents correctly skipped them, but the wasted compute is a smell. Future versions of this rule should detect platform fingerprints (platform bundle URLs, locale-specific add-to-cart strings,
/eshop/, embedded product cards with prices in body) and pre-skip. - Category-name homonyms. "filtr" in Czech also means "filter" in the photography or coffee sense — a coffee-filter or camera-filter blog would pass this gate but isn't a real fit. Sub-agent catches these by reading the body context.
The category gate is intentionally loose. It is a category check, not a subject check — fine-grained "does this specific article fit my specific URL?" is delegated to the per-row sub-agent in Step 8. Example: for a user URL specifically about scraping a single travel site, a general "python web scraping" guide that never mentions that travel site passes this gate because it's in the user's category. The Step 8 sub-agent then decides whether to draft a placement (e.g., an additive line that names the specific travel site) or to recommend a content-based skip.
Surface the extracted category keyword list to the user at the start of Step 6 and let them add/remove before the pass runs. 9. Adversarial-mention detection. When brand_mentioned_in_source: true, scan ±100 characters around the brand mention in the WCC body for negative-context tokens: vs, versus, alternative to, alternatives to, compared to, compared with, instead of, better than, worse than, pros and cons, comparison, review of. If any of those appear within the window, skip with reason Adversarial mention (likely competitor comparison page) — won't link. This catches "ScrapeHero vs Apify" footer mentions, "alternatives to YourBrand" listicles, and similar non-link contexts. Critical — without this rule, the unlinked-mention-claim outreach type fires on dozens of false positives. 10. No contact AND no editorial path. If Contact Email = "Not found" AND no Article Author AND no domain-level contact page found in WCC outbound links, skip — there is no one to pitch. 11. Invalid email (from verification). When enableEmailVerification: true ran, inspect each row's Email Verification status. If the primary contact's status is invalid, try the alternate contacts first before skipping the row — past runs have lost Tier A candidates because the primary contact's email was invalid but a verified alternate existed on the same domain. Only skip with reason Email failed verification (invalid address) when no alternate has a verified or unchecked email. Statuses catch-all, risky, unknown are informational only (not auto-skipped) — surface them in the Email Verification column. Status verified ships as-is. If verification didn't run for this campaign, the column shows - and this rule is a no-op.
Never suggest external lookup services or workaround tools in Notes — no hunter.io, no LinkedIn search, no third-party verification services. The skill's job is to surface what we found, factually. When information is missing (no email, no author, etc.), state the gap and stop. The user knows where to look; suggesting their tools back at them is condescending and clutters the output.
A row failing any rule above is skipped before Step 7. Skipped rows still appear in the final output (so the user can see what was filtered) but with empty placement and email cells and Outreach Status = "Skip" plus the reason in Notes.
Step 7: "Why This Prospect" Tags
For every surviving row, compute one or two Why This Prospect tags, prioritised by which makes the strongest pitch. These tags drive the outreach-type template selection in Step 8.
| Tag | Trigger | Source of truth |
|---|---|---|
Mentions brand, no backlink | brand_mentioned_in_source: true AND backlink_in_source: false in Mentions dataset | Mentions dataset |
Links to competitor [domain] | WCC page body contains an outbound link whose host matches any competitorDomains entry | WCC dataset |
Top-3 SERP for [keyword] | SERP Position is 1, 2, or 3 for any keyword | Google Search Scraper sub-dataset |
Resource / roundup page | WCC page body has 10+ outbound links AND the page title or H1 matches `/(best\ | top\ |
Outdated content | Publish Date is older than 24 months AND newer than 5 years (5+ years was already skipped) | Google Search Scraper / WCC |
A row may carry up to two tags. Order them by pitch strength using this priority: Mentions brand, no backlink > Links to competitor > Resource / roundup page > Top-3 SERP > Outdated content. If no tag fits, leave the column as "-" — the row still gets pitched, just without a special angle.
Step 8: Compose Per-Row Placement and Email
Each surviving row gets three placement artifacts plus one email draft. Apply these quality rules without exception:
1. No fabrication. If the article author or contact email is unknown, set the field to "Not found" and leave a one-line factual note (e.g., "No email found for this contact" or "No author detected"). Do not suggest external lookup tools or workarounds in Notes — see Step 6 rule 11 for the rationale. Just state the fact and stop.
2. Prioritise editorial-leaning contacts. When the All leads dataset returned multiple contacts for the same domain, prefer the one whose jobTitle matches /editor|content|writer|managing|editorial|blog|copy/i, demote anyone whose jobTitle matches /ceo|cfo|cto|founder|chief|vp\b|president/i unless the company is a 1–5 person shop. Surface the chosen contact in the row; keep alternates in Notes as "Alternate contacts: <name1> (<title>), <name2> (<title>)".
3. Three placement artifacts — try strategies in this priority order. Use the WCC sub-dataset's page text. Try strategies 1 → 2 → 3 in order; stop at the first one that produces a clean fit. Record which strategy was used by prepending the Notes field with Placement: drop-in / Placement: additive / Placement: new insertion.
Strategy 1 — drop-in (preferred). Find a sentence in the article where the user's URL can be added to existing words without changing any of the surrounding prose. The link goes on an existing word or short phrase the author already wrote. Output:
Placement Source Sentence= the verbatim sentence as it appears in the article.Placement With Link= the same sentence with the link inserted on an existing word/phrase. No new prose, no rewording, no deletions. Example: source ="Tools like Octoparse and BeautifulSoup work well for hotel data."→ with-link ="Tools like Octoparse and **[BeautifulSoup](URL)** work well for hotel data."(link added to existing word). The editor doesn't have to approve any new wording — just a hyperlink.Placement New Insertion="-".
Drop-in works when the article already names a brand, tool, or technique that maps cleanly to the user's URL. It's the lowest-friction ask of any outreach pattern: "could you add a hyperlink to a word you already wrote?"
Strategy 2 — additive (second choice). When no drop-in target exists but the article has a sentence the user's URL would naturally follow, keep the original sentence intact and add one new sentence after it. The new sentence introduces an adjacent reader-need that the article doesn't already cover and that the user's URL addresses. Output:
Placement Source Sentence= the verbatim original sentence.Placement With Link= original sentence kept verbatim, followed by→and a one-sentence follow-on containing the link. Example: source ="By integrating with Acme Travel's APIs, developers can enrich their platforms with hotel data."→ with-link ="By integrating with Acme Travel's APIs, developers can enrich their platforms with hotel data. → ...with hotel data. In need of competitor pricing data the API doesn't expose? Then you need a [hotel-data scraper](URL)."Keep the original sentence verbatim; the follow-on is the only new prose.Placement New Insertion="-".
The follow-on must (a) raise a reader need the existing sentence doesn't address, (b) connect that need to the user's URL, (c) be one sentence, ≤25 words, written in the article's voice.
Strategy 3 — new insertion (last resort). Only when neither drop-in nor additive works (e.g., no relevant sentence exists in the article body). Draft a fully new 1–2 sentence paragraph in the article's voice with a precise insertion location:
Placement Source Sentence="-".Placement With Link="-".Placement New Insertion= the drafted paragraph + the exact anchor ("insert as a new paragraph immediately after the sentence ending in '…X.' in the section under H2 'Y'.").
If even a new insertion can't be drafted (the article is the wrong topic for the user's URL), set Outreach Status = "Skip" and add Notes: "No natural placement — article topic mismatch". However, the topical-fit gate in Step 6 rule 8 should have caught this case already; if a row makes it to Step 8 and can't get a placement, treat that as a hint that the gate needs more keywords.
Every surviving row goes through a sub-agent — not just Tier A/B/mention-only. The mechanical skip pass (Step 6) cuts the obviously bad prospects (competitor domains, doc subdomains, policy pages, dead-contact rows, off-category articles). Everything that passes is by definition a candidate worth real consideration, and the sub-agent makes the final fit call: read the article, attempt a placement (drop-in → additive → new insertion), and either draft email or return placement_strategy = "skip" with a content-specific reason. Python templates / regex / keyword scoring are not acceptable for the final draft — they produce mechanical splices that read awkward in context (we've seen this fail in practice on real campaigns).
Spawn sub-agents in parallel: one per surviving row, each given the WCC text, user URL context, contact info, brand voice, and partnership offer. The output schema (placement strategy, the three placement column values, email subject + body, skip recommendation, notes) is what gets merged back into the spreadsheet row.
A row that the sub-agent decides to skip after content review gets Outreach Status = "Skip" and the agent's reason in Notes — same shape as a Step 6 mechanical skip, just with a more nuanced rationale.
4. Determine `Outreach Type` per row from the Why This Prospect tags + user goal:
| Trigger | Outreach Type |
|---|---|
Tag Mentions brand, no backlink present | unlinked-mention-claim |
Tag Links to competitor present | competitor-link-replacement |
Tag Resource / roundup page present | resource-page-inclusion |
Tag Outdated content present | outdated-content-replacement |
None of the above OR only Top-3 SERP tag | topical-niche-edit |
Pull the matching template from reference/email-templates.md. The user's Step 2 Partnership type answer substitutes into the {{offer_paragraph}} placeholder inside the template — the outreach type determines structure and opening hook, the partnership type determines the offer.
5. `Suggested Email Copy` must use the user's brand voice. Apply the voice paragraph verbatim per the voice substitution rules in reference/email-templates.md. If voice input was skipped, use the generic-professional default and note this in Notes.
5a. The email MUST include the exact placement wording — verbatim. The recipient should never have to ask "what's the wording you're suggesting?" or click through to a separate cell to see the proposed text. Embed the proposal directly:
- For drop-in: quote both the source sentence and the linked version inline. Example:
"In your line 'Tools like X and Y work well for Z', would you turn 'Y' into a hyperlink to <URL>?" - For additive: quote the anchor sentence verbatim AND the exact follow-on sentence you're proposing. Example:
"Right after your sentence 'X happens because Y.', would you add: 'For the Z case specifically, see <URL>.'?" - For new insertion: quote the anchor sentence the new paragraph should follow, then the full proposed paragraph inline. Example:
"In the 'Honorable mentions' section, after 'each platform has its own trade-offs.', would you add this paragraph: '<full paragraph with link>'?"
The email is the ask. If the wording isn't in the email, the ask is incomplete. Vague phrasing like "happy to draft it for you" / "happy to send exact wording" / "a follow-on sentence linking to..." is a content-skill bug — always rewrite to include the verbatim proposal.
6. Word cap: emails are 150 words or less (subject + body combined).
7. Personalisation is mandatory. Every email must open with a concrete reference to the specific article (title + a one-line takeaway from its content). No generic "I loved your article" openers.
Step 9: Render Output
Markdown format — agent renders inline in chat: 1. A header line with the Apify run ID and tier breakdown (A: 8, B: 15, C: 7, Skipped: 12). 2. One Markdown table row per prospect with the most actionable columns (tier, why, contact, placement summary). Skipped rows render in a separate collapsed section at the bottom. 3. Below the table, one fenced code block per non-skipped row containing the email draft (subject + body), labeled with the row index, tier, and outreach type.
xlsx format — scripts/write_xlsx.py --config campaign.json writes a 2-sheet workbook after Steps 5–8 finish.
xlsx is written as two sheets:
- `Outreach` — active rows only, full 30 columns. This is the send-ready deliverable. Sorted by
Prospect Tierascending (A first), then byDomain DRdescending, then bySERP Positionascending. - `Skipped` — skipped rows with reduced columns:
Domain,Article URL,Article Title,Skip Reason(extracted from Notes),Source Engines,Why This Prospect. This sheet exists for auditing what was filtered without cluttering the main view. Missing Ahrefs columns aren't visible here, so the empty-data confusion goes away.
The user opens the file and lands on Outreach by default — only actionable prospects. They can switch to Skipped to audit. This pattern replaces the older single-sheet-with-red-rows approach.
Both formats also produce a sidecar run_metadata.json: { runId, actorId, startedAt, finishedAt, inputs, datasetIds, tierCounts, skipCounts }. Drop it next to the main output file.
Output Row Schema (30 columns)
SERP Position, Source Engines, Keyword, Article Title, Article URL, Domain, Domain DR, Page Traffic, Referring Domains, Prospect Tier, Why This Prospect, Article Author, Author Source, Publish Date, Contact Full Name, Contact Job Title, Department, Seniority, Contact Email, Email Verification (one of verified / catch-all / risky / invalid / unknown / -), Contact LinkedIn, Company, Outreach Type, Partnership Offer, Placement Source Sentence, Placement With Link, Placement New Insertion, Suggested Email Copy, Outreach Status (default "Not started", "Skip" for skipped rows), Notes (auto-flags + skip reason + manual hints + alternate contacts).
Full schema with types and source datasets per column is in reference/output-formats.md.
Error Handling
| Error / symptom | What to do |
|---|---|
APIFY_TOKEN not found | Ask user to create .env with APIFY_TOKEN=your_token. Get one at console.apify.com/account/integrations. |
| Ahrefs MCP unavailable | Skip Step 5. Set Domain DR, Page Traffic, Referring Domains, Prospect Tier to "-" and add a one-line note in the output header explaining tiers were not computed. Continue with the rest of the workflow. |
Cannot find module 'xlsx' | Run npm install inside the skill's scripts/ folder. |
Error: 'brand' is required | The user skipped Step 1 anchor #3. Re-ask brand name. |
Actor run TIMED-OUT (client-side, Actor still running on Apify) | Do not restart. Use node --env-file=.env scripts/fetch_run_artifacts.js --run-id <runId> --output YYYY-MM-DD_outreach.json --timeout 1800 to poll the existing run and download all datasets when it terminates. Same output shape as the runner. |
Actor run TIMED-OUT (Actor itself ran past its timeoutSecs) | See reference/troubleshooting.md. Lower organicResult, cut keywords, or disable some LLM engines. Almost never the cause — usually it's the client-side wait. |
Author = Not found | Expected for ~30% of pages without bylines. Skill writes "Not found" and adds a manual-lookup hint to Notes. Do not fabricate. The AI Web Scraper sub-Actor frequently TIMES-OUT at its 300s default mid-crawl (past campaigns have lost author data for ~7 high-DR sites at a time this way); when this happens the _authors.json dataset still contains the partial results that finished before the timeout. WCC metadata.author / openGraph article:author / JSON-LD Person.name are the fallbacks the agent should try before writing "Not found". |
Sub-Actor datasets missing (no _mentions.json / _authors.json / _serp.json / _wcc.json after --fetch-sub-datasets) | Actor's output schema changed on or before 2026-05-20: SUB_ACTOR_RESULTS is no longer in the KV store, and MENTIONS/AUTHORS/DOMAINS_WITH_LEADS named datasets are gone. The runner script now falls back to log-parsing automatically — but if it didn't, run node --env-file=.env scripts/fetch_subactors_from_log.js --run-id <id> --base <prefix> to populate _serp.json, _wcc.json, _authors.json from the sub-Actor runIds visible in the parent log. Mention data is now embedded in main_leads[i].source_url[], not a separate file. |
Contact Email = Not found | Contact Details Scraper sub-Actor missed the site. Set the field to "Not found" and leave a one-line factual Notes entry ("No email found for this contact"). Do not fabricate, do not suggest external tools. If both contact and author are missing, the skip pass (Step 6, rule 7) will drop the row. |
| All rows skipped by goal filter | The user's goal is too narrow for the SERP results. Suggest broadening the goal (e.g., Topical authority instead of Recover unlinked brand mentions) or expanding keywords. |
0 leads returned | Keyword too narrow, or ownDomains / competitorDomains filtered out all SERP results. Broaden keyword, narrow exclusions, raise organicResult. |
| Costs higher than expected | Sub-Actor fan-out (Google Search Scraper, WCC, Contact Details Scraper, AI Web Scraper) stacks. See cost section in reference/apify-actor-usage.md. To shrink: drop searchAuthorName, disable AI platforms (enableChatGpt: false, etc.), lower maxContactsPerDomain to 1, lower organicResult to 5. |
{
"base": "YYYY-MM-DD_<short-name>_outreach",
"goal": "Topical authority links to specific URL | Maximum link volume from any relevant site | Recover unlinked brand mentions | Replace competitor links | Custom",
"today": "YYYY-MM-DD",
"brand": {
"name": "",
"aliases": []
},
"user_url": "",
"own_domains": [],
"competitor_domains": [],
"already_pitched_domains": [],
"category_keywords": [],
"subject_keywords": [],
"brand_voice": "",
"partnership_type": "ABC link exchange | Direct A B link exchange | Resource page / list inclusion | Unilateral ask (no reciprocal) | Other"
}
Example: two brand voices for a topical-niche-edit row with ABC partnership offer
The same prospect (an editor at a SaaS comparison blog) approached twice — once in a casual founder-led voice, once in a formal B2B voice. Note what changes (voice surface) and what stays constant (outreach type structure, placement artifacts, ABC offer).
The row in question is a typical topical-niche-edit: the article is topically relevant, ranks top-3 for the user's keyword, and has a clean existing sentence the link can be spliced into. No special angle like an unlinked mention or competitor link — just a strong topical fit.
Shared context across both examples
- Prospect article:
"The 6 best customer feedback tools for product teams"atfeedbackdaily.example. - Article author (from
searchAuthorName):Sara Bittencourt. - Outreach contact (from All leads, picked over a higher-ranking VP Marketing because of editorial-leaning job title):
Daniel Reyes, Head of Content, daniel@feedbackdaily.example. - User's brand:
AcmeFeedback. - User's content URL:
acmefeedback.example/in-app-vs-email-feedback-2026. - Goal:
Topical authority links to specific URL. - Partnership type:
ABC link exchangewith a partner siteproductloop.example. - Ahrefs metrics:
Domain DR = 64,Page Traffic = 1,800/mo,Referring Domains = 47. - Prospect Tier:
A. - Why This Prospect:
Top-3 SERP for customer feedback tools. - Outreach Type:
topical-niche-edit.
Placement artifacts (same for both versions)
The article contains a clean existing sentence where the link fits naturally, so the `Placement With Link` column is filled and Placement New Insertion is "-":
- Placement Source Sentence:
"In-app prompts beat email surveys for activated users; the inverse holds for churned ones." - Placement With Link:
"~~In-app prompts beat email surveys for activated users;~~ → "In-app prompts beat email surveys for activated users — see **[AcmeFeedback's breakdown](https://acmefeedback.example/in-app-vs-email-feedback-2026)** for the activation-stage cutoff — the inverse holds for churned ones." - Placement New Insertion:
"-"
The diff shows the splice: the existing sentence is kept; an em-dash clause carrying the link is inserted mid-sentence.
Version 1: casual, founder-led voice
User's voice paragraph:
Casual but precise. Contractions are fine. We sound like one person, not a marketing team. Lead with what we noticed, not what we want. Open "Hey {first_name}," and sign off "— Maya, AcmeFeedback".
Email:
Subject: Addition to your "6 best customer feedback tools" piece?
Hey Daniel,
Just read Sara's "The 6 best customer feedback tools for product teams" — the call that in-app prompts beat email for activated users but lose for churned ones is the cleanest framing on this I've seen.
I'm Maya from AcmeFeedback. We published acmefeedback.example/in-app-vs-email-feedback-2026 which goes deep on the activation-stage cutoff (the "when does in-app beat email?" cutoff specifically). Natural place for it: a quick em-dash clause inside Sara's "In-app prompts beat email surveys for activated users" sentence — happy to suggest exact wording.
In exchange, I'd link to your piece from ours, and a partner site (productloop.example) would add a link to yours too.
Worth a look?
— Maya, AcmeFeedbackWord count: 138. Voice traits hit:
- "Hey Daniel," opener (per voice paragraph).
- Contractions ("I'm", "I'd", "it's").
- One-person tone — no "we at AcmeFeedback" corporate plural.
- Leads with what was noticed about Sara's piece, not what Maya wants.
- "— Maya, AcmeFeedback" close.
Version 2: formal B2B voice
User's voice paragraph:
Formal, polished, no contractions. We represent a company, not an individual. Use full sentences. Open with "Dear {first_name}," and close with "Regards, the AcmeFeedback team".
Email:
Subject: Addition to "The 6 best customer feedback tools for product teams"
Dear Daniel,
We read with interest the recent article by Sara Bittencourt, "The 6 best customer feedback tools for product teams," and noted the distinction drawn between in-app prompts and email surveys, particularly regarding their differential effectiveness for activated versus churned users.
The AcmeFeedback team has published a complementary analysis at acmefeedback.example/in-app-vs-email-feedback-2026, which examines the activation-stage cutoff in detail. A natural placement would be as a parenthetical reference within the existing sentence about in-app prompts beating email for activated users.
In exchange, we would link to your article from our publication, and a partner publication, productloop.example, would extend an inbound link to your piece as well.
We welcome the opportunity to discuss.
Regards,
the AcmeFeedback teamWord count: 140. Voice traits hit:
- "Dear Daniel," opener.
- No contractions anywhere.
- Plural corporate voice ("we", "the AcmeFeedback team").
- Full sentences, no em-dashes, no fragments.
- "Regards, the AcmeFeedback team" close.
What stayed constant across both versions
- The opening reference to a specific subsection (the in-app vs email distinction) — both emails earn the recipient's attention by demonstrating they read the piece.
- The article author (Sara) is named in v2 because
searchAuthorNamereturned her, but the outreach is addressed to Daniel (the actual contact). Author mention is optional in v1 because the casual tone doesn't require attribution. - The ABC
{{offer_paragraph}}substitution: I link to you, partner links to you, you link to me. Both versions describe the trade clearly. - The placement summary maps directly to the row's
Placement With Linkcell — no "somewhere in the piece" hand-waving. - Both are under 150 words.
- Both use the same outreach-type template (
topical-niche-edit) and the same offer paragraph (ABC). Only the voice surface differs.
What changed
| Element | Casual | Formal |
|---|---|---|
| Opener | "Hey Daniel," | "Dear Daniel," |
| Pronoun | "I" | "We" / "the AcmeFeedback team" |
| Contractions | yes | no |
| Sentence length | mixed, with em-dashes | longer, formal |
| Author mention | implicit (just the article title) | explicit ("by Sara Bittencourt") |
| Closing CTA | "Worth a look?" | "We welcome the opportunity to discuss." |
| Sign-off | "— Maya, AcmeFeedback" | "Regards, the AcmeFeedback team" |
The skill must produce both correctly. If the user's voice paragraph is closer to v1, the agent must not regress to v2 — even though v2 is the "safer" generic default for outreach.
What would change if the row's outreach type were different
This row uses topical-niche-edit because no stronger angle was detected. If the same article had instead:
- Mentioned AcmeFeedback without linking →
unlinked-mention-claimtemplate, subject would be"Quick fix on your '6 best customer feedback tools' piece — missed link?", opener would lead with the unlinked mention itself, not with the in-app/email distinction. The ABC offer paragraph would still substitute in. - Linked to a competitor (e.g., Hotjar) →
competitor-link-replacementtemplate, subject would be"Updated alternative to hotjar.com in your '6 best customer feedback tools' piece", opener would acknowledge the Hotjar link and offer AcmeFeedback as an addition (not a replacement). ABC offer paragraph still substitutes in.
The partnership type (ABC) stays constant across the user's whole campaign; the outreach type varies per-row based on Why This Prospect.
Example: resource-page-inclusion outreach with the Resource page partnership type
Resource-page outreach is a different beast from the other outreach types. The prospect page is a curated list (roundup, "best of", "top 10"). The pitch is to be added as a new entry on the list, not to splice a link into prose. In the new skill model:
- Outreach Type =
resource-page-inclusion(auto-detected from theWhy This ProspecttagResource / roundup page, which is set in Step 7 when the WCC body has 10+ outbound links and a list-flavored title). - Partnership Type =
Resource page / list inclusion(selected by the user at Step 2). - Placement is always a new insertion (column
Placement New Insertion), never a splice into an existing sentence. Resource pages are list-structured; you're adding a new list item.
Context
- Prospect article:
"50 best open-source developer tools we use in 2026"atdevstackpicks.example. - Article author (from
searchAuthorName):Linus Mehta. - Outreach contact (from All leads):
Aria Chen, Editor-in-Chief, aria@devstackpicks.example— picked over the listed Head of Marketing becauseEditor-in-Chiefmatches the editorial-leaning regex in Step 8 rule 2. - User's brand:
AcmeMonitor— an open-source uptime monitor. - User's content URL:
acmemonitor.example. - Goal:
Topical authority links to specific URL. - Partnership type:
Resource page / list inclusion. - Ahrefs metrics:
Domain DR = 71,Page Traffic = 2,400/mo,Referring Domains = 89. - Prospect Tier:
A. - Why This Prospect:
Resource / roundup page. - Outreach Type:
resource-page-inclusion.
Placement artifacts
The article is a numbered list of 50 tools grouped by subsection. There's no existing sentence to splice into — the natural placement is a brand-new entry alongside Uptime Kuma in the "Monitoring and observability" subsection. So Placement New Insertion is filled and the other two are "-".
- Placement Source Sentence:
"-" - Placement With Link:
"-" - Placement New Insertion:
Insert as a new list item in the "Monitoring and observability" subsection, immediately after the Uptime Kuma entry. Format to match the existing list entries (bold tool name, one-paragraph description, link in parentheses):>
AcmeMonitor — Lightweight self-hosted uptime monitor. Single Go binary, MIT-licensed, no database required, runs on a $4/mo VPS. Trade-off vs. Uptime Kuma: smaller surface area, no built-in JS notifications dashboard. (acmemonitor.example)
The inserted draft matches the article's list format (bold name → one-paragraph description → link in parentheses) so the editor can drop it in without rewriting.
Voice and email
User's voice paragraph:
Direct, technical, no fluff. We're a small open-source project run by two engineers. We don't oversell. Open "Hi {first_name}," and close with "— the AcmeMonitor maintainers".
Email:
Subject: Suggestion for your 2026 open-source dev tools list
Hi Aria,
Just went through "50 best open-source developer tools we use in 2026" — really tight selection. The Plausible Analytics entry is the first place I've seen anyone explain the GoatCounter trade-off honestly instead of just saying "it's lighter".
I'm one of the maintainers of AcmeMonitor (acmemonitor.example) — an open-source uptime monitor, MIT-licensed, single Go binary. It sits naturally in your "Monitoring and observability" subsection, right next to Uptime Kuma. Would you consider adding it in your list? Happy to discuss potential ways of collaboration.
Best,
— the AcmeMonitor maintainersWord count: 144. Voice traits hit:
- "Hi Aria," opener.
- "— the AcmeMonitor maintainers" close (plural, since the voice paragraph said "two engineers").
- The
{{offer_paragraph}}substitution fromResource page / list inclusionpartnership type is the line"I'm not asking for a reciprocal link — happy to send this in as a straight suggestion for your list."
What's different from other outreach-type templates
| Element | topical-niche-edit / competitor-link-replacement / outdated-content-replacement | resource-page-inclusion |
|---|---|---|
| Placement mode | Either splice into existing sentence (Placement With Link) or new insertion (Placement New Insertion) | Always new insertion — you're adding a list item, not editing prose |
| Reference | A section heading, claim, or competitor link | A specific listed item (the entry adjacent to where yours would go) |
| Tone | Transactional — focus on what the addition gains the reader | Editorial — "would you consider adding it?" |
| Risk if you skip the merit | Acceptable for ABC/AB — the exchange motivates them anyway | Fatal — there's nothing in it for them otherwise |
| Format match | Doesn't matter | The new insertion must match the article's existing list entry format (bold name + paragraph + link, or whatever the page uses) |
What's still constant across all outreach types
- Opens with a concrete reference to something in the article (here: the Plausible / GoatCounter call).
- Cites a real adjacent item / claim (here: Uptime Kuma) so the recipient can see the context.
- Names the placement spot precisely ("Monitoring and observability" subsection, immediately after Uptime Kuma).
- No fabricated stats, no fabricated user counts, no fabricated "as featured in TechCrunch".
- Under 150 words.
- The drafted new insertion is the same text used in the
Placement New Insertioncolumn and in the email body — agent must keep them in sync.
Outreach-type override on the fly
The user's Step 2 partnership type and the per-row outreach type are normally chosen independently. But the Why This Prospect tag Resource / roundup page should auto-set Outreach Type = resource-page-inclusion regardless of partnership type — splicing a link into a list page reads as spam.
If the user's partnership type is ABC link exchange or Direct A B link exchange but the row is a resource page, the agent should:
1. Still use the resource-page-inclusion template structure. 2. Substitute the {{offer_paragraph}} for whatever partnership type the user picked — but flag in Notes: "Resource page detected — partnership type ABC may weaken the pitch on a list page. Consider switching this row to Unilateral ask before sending."
The user can revert if they disagree; the override is a default, not a lock.
Example: full input walkthrough + sample output
A worked example of one full run through the skill. The data below is illustrative, not from a real Actor run — domain names, contacts, Ahrefs metrics, and article titles are placeholders.
Step 1: Required anchor inputs
| Field | Value |
|---|---|
| Concrete goal | Topical authority links to specific URL |
| Target keyword(s) | headless cms for ecommerce, best headless cms 2026 |
| Brand name | AcmeCMS |
| Product/category description | AcmeCMS — headless CMS purpose-built for D2C ecommerce. We sell to merchants and storefront developers who need Stripe-readiness scoring and Shopify-equivalent ergonomics from their CMS. |
| Content URL to link to | https://acmecms.example/headless-cms-comparison-2026 |
| Competitors | contentful.com, sanity.io, strapi.io, prismic.io, crystallize.com, storyblok.com, hygraph.com, dato.cms, kontent.ai, directus.io (user-supplied — the prompt explicitly asks for "anyone who'd write a vs-AcmeCMS comparison page on their site, even small competitors"; agent offered Ahrefs auto-pull but user declined for this run) |
| Already-pitched domains (optional) | crystallize.com, prismic.io |
| Organic results per keyword | 10 |
| LLM sources to track | All 6 enabled: ChatGPT Search, Gemini, Copilot, Perplexity, Google AI Mode, Google AI Overviews |
| Run email verification? | yes (enableEmailVerification: true) |
The product/category description is used in Step 6 (topical-fit gate) to extract specific topical keywords (headless cms, ecommerce, d2c, stripe, merchant, comparison) that prospect articles must contain. It's also used in Step 6 (adversarial-mention detection) to recognise when a "Mentions AcmeCMS" page is actually a competitor's comparison footer.
The LLM-source multi-select maps directly to the Actor's enableChatGpt / enableGemini / enableCopilot / enablePerplexity / enableAiMode / enableAiOverviews flags. Email verification maps to enableEmailVerification — every email returned in the leads dataset gets an email_verification status (verified, catch-all, risky, invalid, or unknown) which feeds the Email Verification output column and the Step 6 invalid-email skip rule.
Step 2: Secondary inputs
Brand voice:
We're a small, founder-led team. Casual but precise — we use contractions, em-dashes, and we don't pretend we're a Fortune 500. Bias for honesty over hype: if a competitor does something well we'd say so. Open emails with "Hey {first_name}," and close with "— Nadia".
Partnership type: ABC link exchange (with partner site decoupled.example)
Output format: Markdown
Step 3: Actor call
node --env-file=.env ${CLAUDE_PLUGIN_ROOT}/scripts/run_actor.js \
--actor "apify/link-prospecting-tool" \
--input '{
"queries": "headless cms for ecommerce\nbest headless cms 2026",
"brand": "AcmeCMS",
"ownDomains": ["acmecms.example", "docs.acmecms.example"],
"competitorDomains": ["contentful.com", "sanity.io", "strapi.io"],
"ignoreDomains": [
"wikipedia.org","github.com","stackoverflow.com","stackexchange.com",
"reddit.com","quora.com","youtube.com","twitter.com","x.com",
"linkedin.com","facebook.com","medium.com","archive.org"
],
"organicResult": 10,
"maxContactsPerDomain": 3,
"department": ["marketing"],
"searchAuthorName": true,
"includeMention": true,
"enableChatGpt": true,
"enableAiMode": true,
"enableAiOverviews": true,
"enablePerplexity": true
}' \
--timeout 1200 \
--fetch-sub-datasets \
--output 2026-05-13_acmecms_outreach.json \
--format jsonNote that department is ["marketing"] only — c_suite is dropped because CEOs don't edit articles. The competitor list passed to the Actor is also reused in Step 7 to detect the Links to competitor tag.
Step 5: Ahrefs enrichment (per surviving domain)
For each unique domain in the Actor output, fetch:
mcp__claude_ai_Ahrefs__site-explorer-domain-rating→Domain DRmcp__claude_ai_Ahrefs__site-explorer-metrics(target=article URL,mode: "exact") →Page Traffic(last 30 days, organic)mcp__claude_ai_Ahrefs__site-explorer-backlinks-stats→Referring Domains
Compute Prospect Tier using the Topical authority thresholds (since that's the goal):
- A: DR ≥ 50 AND Page Traffic ≥ 300/mo
- B: DR 30–49 OR Page Traffic 50–299
- C: everything below
Step 6: Skip pass results
Of the 14 leads the Actor returned, 3 were skipped:
| Domain | Skip reason |
|---|---|
crystallize.com | Already pitched (Step 1 input #6) |
oldcms-blog.example | Stale content (published 2019-03-22, > 5 years old) |
pricing-page.example | Non-editorial page type (vendor pricing page) |
Surviving rows: 11. We'll show 4 of them below.
Step 7-8: Sample 4-row Markdown output
# Link prospecting results — 2026-05-13
Run ID: aB12cDeFgHiJk
Goal: Topical authority links to specific URL
Keywords: headless cms for ecommerce, best headless cms 2026
Brand: AcmeCMS
Content URL: https://acmecms.example/headless-cms-comparison-2026
Partnership: ABC link exchange
Tier breakdown: A: 4, B: 5, C: 2, Skipped: 3
| # | Tier | Why | SERP | Engines | Domain | DR | Traffic | Article | Contact | Email Verif | Outreach Type | Placement |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | A | Top-3 SERP for headless cms for ecommerce | 2 | Google Organic, ChatGPT, Gemini | saasdigest.example | 78 | 4,200 | [Headless CMS in 2026: a buyer's guide](https://saasdigest.example/headless-cms-2026-guide) | Elena Vargas, Content Editor — elena@saasdigest.example | verified | topical-niche-edit | Strategy: additive — original sentence kept verbatim + follow-on naming AcmeCMS |
| 2 | B | Links to competitor (sanity.io) | 5 | Google Organic, Perplexity, Copilot | merchantsblog.example | 41 | 180 | [Why we moved off Shopify to a headless setup](https://merchantsblog.example/moved-off-shopify) | Tom Park, Founder — tom@merchantsblog.example | catch-all | competitor-link-replacement | Strategy: new insertion — after the "Picking the CMS" subsection |
| 3 | A | Mentions brand, no backlink | 4 | Google Organic, ChatGPT, Gemini, Copilot | dtcweekly.example | 62 | 1,100 | [The 8 best headless CMSes for D2C in 2026](https://dtcweekly.example/headless-cms-d2c-2026) | Mark Lee, Head of Content — mark.lee@dtcweekly.example | verified | unlinked-mention-claim | Strategy: drop-in — link added onto the existing word "AcmeCMS" |
| 4 | B | Top-3 SERP for best headless cms 2026 | 3 | Google Organic, Gemini | ecombytes.example | 35 | 240 | [Ranking the best headless CMS platforms for 2026](https://ecombytes.example/best-headless-cms-2026) | Priya Raman, Marketing Director — priya@ecombytes.example | risky | topical-niche-edit | Strategy: new insertion — Honorable Mentions section |
## Skipped (3)
| Domain | Reason |
|---|---|
| crystallize.com | Already pitched (Step 1 input #6) |
| oldcms-blog.example | Stale content (published 2019-03-22) |
| pricing-page.example | Non-editorial page type (vendor pricing page) |Per-row placement detail
Row 1 — saasdigest.example — Tier A — topical-niche-edit (Strategy: additive)
- Placement Source Sentence:
"For teams scaling across many storefronts, some teams need a more API-first approach." - Placement With Link:
"For teams scaling across many storefronts, some teams need a more API-first approach. → For teams scaling across many storefronts, some teams need a more API-first approach. If your storefront also needs Stripe-readiness scoring, [AcmeCMS](https://acmecms.example/headless-cms-comparison-2026) is purpose-built for that."(original sentence kept verbatim; one follow-on sentence added raising a need the original doesn't address) - Placement New Insertion:
"-" Notesprepended withPlacement: additive.
Row 2 — merchantsblog.example — Tier B — competitor-link-replacement (Strategy: new insertion)
- Placement Source Sentence:
"-"(no existing sentence fits naturally) - Placement With Link:
"-" - Placement New Insertion:
"Insert as a new paragraph immediately after the sentence ending in '…we shortlisted Sanity.' in the 'Picking the CMS' subsection: 'If you're specifically optimizing for D2C ecommerce, it's worth comparing against [AcmeCMS](https://acmecms.example/headless-cms-comparison-2026), which scores merchants on Stripe-readiness — something Sanity doesn't surface directly.'" Notesprepended withPlacement: new insertion.
Row 3 — dtcweekly.example — Tier A — unlinked-mention-claim (Strategy: drop-in)
- Placement Source Sentence:
"For D2C brands going headless this year, tools like AcmeCMS and Contentful are taking this approach." - Placement With Link:
"For D2C brands going headless this year, tools like **[AcmeCMS](https://acmecms.example/headless-cms-comparison-2026)** and Contentful are taking this approach."(the link is added onto the existing word "AcmeCMS" — no other text changes, no new prose) - Placement New Insertion:
"-" Notesprepended withPlacement: drop-in. This is the lowest-friction ask of all three strategies — the editor just turns one of their own words into a hyperlink.
Row 4 — ecombytes.example — Tier B — topical-niche-edit (Strategy: new insertion)
- Placement Source Sentence:
"-" - Placement With Link:
"-" - Placement New Insertion:
"Insert as a new paragraph at the end of the 'Honorable mentions' section, after the sentence ending in '…each platform has its own trade-offs.': 'AcmeCMS is worth a look if you're specifically headless-for-D2C — they score CMSes by Stripe-readiness rather than treating ecommerce as a generic content domain. See their 2026 comparison: [acmecms.example/headless-cms-comparison-2026](https://acmecms.example/headless-cms-comparison-2026).'" Notesprepended withPlacement: new insertion.
Email drafts
Row 1 — Tier A — saasdigest.example — topical-niche-edit — Elena Vargas
Subject: Addition to "Headless CMS in 2026: a buyer's guide"?
Hey Elena,
>
Your "Headless CMS in 2026: a buyer's guide" is one of the clearer pieces on this I've found — particularly the scoring on Stripe-readiness instead of just "good for ecommerce".
>
I'm Nadia from AcmeCMS. We recently published acmecms.example/headless-cms-comparison-2026 — same topic, with a D2C-merchant-fit angle. If it fits, the natural place is alongside your existing "some teams need a more API-first approach" line in the comparison.
>
In exchange, I'd link to your guide from our piece, and a partner site (decoupled.example) would add a link to yours too.
>
Open to it?
>
— Nadia
Row 2 — Tier B — merchantsblog.example — competitor-link-replacement — Tom Park
Subject: Updated alternative to sanity.io in your "Why we moved off Shopify to a headless setup"
Hey Tom,
>
Read your write-up — the part on ditching the storefront API in favor of a CMS-first pipeline is exactly the trap most teams fall into. You shortlisted Sanity, which is a solid pick, though it doesn't surface Stripe-readiness directly.
>
I'm Nadia from AcmeCMS. We just published acmecms.example/headless-cms-comparison-2026 which scores CMSes by D2C-merchant fit, including the Stripe-readiness gap.
>
The natural place: a new paragraph after your "we shortlisted Sanity" line.
>
In exchange, I'd link to your piece from ours and a partner site (decoupled.example) would too.
>
Worth a look?
>
— Nadia
Row 3 — Tier A — dtcweekly.example — unlinked-mention-claim — Mark Lee
Subject: Quick fix on your "The 8 best headless CMSes for D2C in 2026" — missed link?
Hey Mark,
>
Noticed your piece "The 8 best headless CMSes for D2C in 2026" mentions AcmeCMS in the line "tools like AcmeCMS and Contentful are taking this approach" — thanks for the shout-out.
>
Looks like the mention isn't linked. Would you be open to adding a link to acmecms.example/headless-cms-comparison-2026? Helps readers who want the full D2C-merchant scoring methodology.
>
Happy to do an ABC swap in return if useful — I'd link your piece from ours and a partner site (decoupled.example) would too.
>
Either way, appreciate the mention.
>
— Nadia
Row 4 — Tier B — ecombytes.example — topical-niche-edit — Priya Raman
Subject: Addition to "Ranking the best headless CMS platforms for 2026"?
Hey Priya,
>
Just went through your ranking — the "honorable mentions" section is the most useful part because it's where the niche-fit picks actually surface.
>
I'm Nadia from AcmeCMS. We score CMSes by Stripe-readiness for D2C merchants specifically — acmecms.example/headless-cms-comparison-2026. Could fit naturally as a new entry at the end of your honorable mentions, after the "each platform has its own trade-offs" line.
>
In exchange, I'd link to your ranking from our piece, and a partner site (decoupled.example) would link to yours too.
>
Open to it?
>
— Nadia
Notes on this example
- All names, domains, articles, contacts, and Ahrefs metrics are placeholders. Do not use them as templates for real outreach.
- Each row uses a different outreach-type template even though every row uses the same partnership type (ABC). The outreach type is determined per-row from
Why This Prospect; the partnership type substitutes into{{offer_paragraph}}in whichever template was chosen. - Row 1 illustrates the with-link diff placement: an existing sentence is a clean fit, the diff column shows the splice.
- Rows 2 and 4 illustrate the new insertion placement: no existing sentence fits, so a drafted 1–2 sentence paragraph is provided with a precise location.
- Row 3 illustrates the simplest splice: the brand is already named in a sentence, so the link goes on the existing word — no other text changes.
- The contact in Row 1 is "Content Editor" (not "Head of Marketing") because Step 8 rule 2 prioritises editorial-leaning job titles when the All leads dataset returns multiple contacts for one domain.
- Row 3's email uses
unlinked-mention-claimeven though the user picked ABC partnership at Step 2 — the offer paragraph still uses the ABC offer, but the template structure is the mention-claim opener and the ask is gentler. - Tier counts in the header (
A: 4, B: 5, C: 2) and skip counts (3) are surfaced in therun_metadata.jsonsidecar'stierCountsandskipCountsblocks.
Apify Actor Usage: apify/link-prospecting-tool
Reference for how this skill calls the link-prospecting Actor: full input schema, recommended payload, dataset structure, sub-Actor access, billing, and the timeout gotcha.
Actor at a glance
- Actor ID:
apify/link-prospecting-tool - Apify Store page: https://apify.com/apify/link-prospecting-tool
- What it does: Runs each query against Google Search, ChatGPT Search, Perplexity, and Google AI Mode / AI Overviews; filters out the user's own and competitor domains; crawls each remaining source page to detect brand mentions and backlinks; enriches the domains that don't already link to the user with contact details; and optionally identifies article authors via AI.
- Live latest build at time of skill authoring: 0.0.5
Required vs optional inputs
| Field | Type | Required | Default | What it does |
|---|---|---|---|---|
queries | string (newline-separated) | Yes | — | One search query per line. Join keywords with \n before passing. |
brand | string | Yes | — | The user's brand name. Used case-insensitively to detect mentions and backlinks on crawled pages. The Actor will not start without this. |
organicResult | integer | No | 10 | Number of Google organic SERP results per query. Asking for more than 10 slows the run and raises Google Search Scraper cost ($4.5 per 1k results). Max 500. |
enableChatGpt | boolean | No | true | Include ChatGPT Search results. Adds Google Search Scraper cost per result. |
enableGemini | boolean | No | true | Include Google Gemini results. Adds Google Search Scraper cost per result. |
enableCopilot | boolean | No | true | Include Microsoft Copilot (Bing AI) results. Adds Google Search Scraper cost per result. |
enablePerplexity | boolean | No | true | Include Perplexity Sonar results. Adds Google Search Scraper cost per result. |
enableAiMode | boolean | No | true | Include Google AI Mode results. Adds Google Search Scraper cost per result. |
enableAiOverviews | boolean | No | true | Process Google AI Overviews surfaced in the SERP. No extra cost — parsed from the SERP already fetched for organic results. |
enableEmailVerification | boolean | No | true | Verify every email returned by the Contact Details Scraper sub-Actor. Adds Email Verifier sub-Actor cost per email (typically much cheaper than the other sub-Actors). Each lead gets an email_verification field with one of: verified, catch-all, risky, invalid, unknown. |
ownDomains | string[] | No | [] | The user's own domains, skipped during source analysis. |
competitorDomains | string[] | No | [] | Competitor domains, skipped during source analysis. |
ignoreDomains | string[] | No | [] | Other domains to exclude. The skill defaults to the standard prefill list (UGC + giants) — see "Recommended payload" below. |
maxContactsPerDomain | integer | No | 1 | Max contacts per source domain (range 1-10). The skill defaults to 3. |
includeMention | boolean | No | true | When true, sources that mention the brand without linking are also routed to the contact-enrichment step. |
department | string[] (enum) | No | ["marketing","c_suite"] | Departments to target for contact enrichment. Valid values: c_suite, engineering_technical, product, design, finance, education, human_resources, information_technology, legal, marketing, medical_health, operations, sales, consulting. |
searchAuthorName | boolean | No | false | Identify the article author via AI Web Scraper sub-Actor ($25/1k sources). The skill enables this by default for personalisation. |
Recommended payload for this skill
{
"queries": "<keyword 1>\n<keyword 2>",
"brand": "<user's brand name>",
"ownDomains": ["<user-domain.com>"],
"competitorDomains": [],
"ignoreDomains": [
"wikipedia.org", "github.com", "stackoverflow.com", "stackexchange.com",
"reddit.com", "quora.com", "youtube.com", "twitter.com", "x.com",
"linkedin.com", "facebook.com", "medium.com", "archive.org",
"chromewebstore.google.com", "addons.mozilla.org", "apps.apple.com",
"play.google.com", "microsoftedge.microsoft.com", "marketplace.visualstudio.com"
],
"organicResult": 10,
"maxContactsPerDomain": 3,
"department": ["marketing"],
"searchAuthorName": true,
"includeMention": true,
"enableChatGpt": true,
"enableGemini": true,
"enableCopilot": true,
"enablePerplexity": true,
"enableAiMode": true,
"enableAiOverviews": true,
"enableEmailVerification": true
}Override department only if the user has a specific outreach angle (e.g., add sales if they want BD-style partnerships). Override the LLM-source booleans (enableChatGpt, enableGemini, enableCopilot, enablePerplexity, enableAiMode) only to cut cost in tight budgets. enableAiOverviews is free — keep on. enableEmailVerification is recommended on; turn off only when verification quota is constrained or for cost-tight smoke tests.
Timeout gotcha (critical)
The Actor's own run-level timeoutSecs defaults to 60000 (~16 hours) — that is fine.
The trap is the API client wait timeout — how long the calling code polls for completion before giving up. Apify's JS client and most lightweight runners default to a few minutes. The link-prospecting Actor typically runs 5-15 minutes per query, much longer for big keyword lists, and 18 of 24 recent public runs ended in TIMED-OUT status with default polling.
In this skill, scripts/run_actor.js exposes --timeout (in seconds) for the client-side wait, not the Actor's runtime. Always pass at least --timeout 900. Raise to 1800 or 3600 for runs with more than three queries.
Datasets produced by one run
Every run emits one default dataset plus several others accessible from the run's Storage tab:
| Dataset | What it contains |
|---|---|
| Default (All leads) | One row per enriched contact. Split into three batches by the Actor — one contact per domain per batch. |
| Mentions | One row per source URL. Records which engines surfaced it and whether the brand was mentioned / backlinked. |
| Domains with leads | Distinct list of domains that yielded a contact. Useful to feed into the next run's competitorDomains after a successful outreach round. |
| Author list | One row per source URL with searchAuthorName: true. Contains author name and (when discoverable) email. |
| Sub-Actor results | Index of sub-runs (Google Search Scraper, Website Content Crawler, Contact Details Scraper, AI Web Scraper). Each entry links to its own dataset and run page. |
All leads row shape (from Actor README)
{
"firstName": "first name",
"lastName": "last name",
"linkedinProfile": "http://www.linkedin.com/in/fullName",
"email": "firstName@example.com",
"email_verification": "verified",
"mobileNumber": null,
"jobTitle": "SEO specialist",
"industry": null,
"city": "Prague",
"country": "Czechia",
"companyName": "example",
"companyWebsite": "example.com",
"companySize": null,
"companyLinkedin": "http://www.linkedin.com/company/example",
"companyCity": "Prague",
"departments": ["marketing"],
"seniority": "vp",
"twitter": null,
"domain": "example.com",
"source_url": [],
"brand_mentioned": false
}The email_verification field is only populated when enableEmailVerification: true ran for the run. Possible values: verified (deliverable to a real inbox), catch-all (domain accepts everything — uncertain), risky (role-based, free-mail, or other low-confidence pattern), invalid (no MX / bounces / known bad), unknown (verifier didn't return a determination). If verification didn't run, the field is absent and the skill treats it as -.
Mentions row shape
{
"url": "https://www.example.com/article",
"domain": "https://www.example.com",
"brand_mentioned_in_source": true,
"backlink_in_source": false,
"Perplexity_mention": false,
"ChatGPT_mention": true,
"Gemini_mention": false,
"Copilot_mention": false,
"AIOverview_mention": false,
"AIMode": false,
"OrganicResult_mention": true,
"queries": ["best search engine"]
}Gemini_mention and Copilot_mention are present only when the corresponding enable* flag was true for the run. The skill comma-joins all true engine flags into the Source Engines column (with friendly labels: Google Organic, ChatGPT, Gemini, Copilot, Perplexity, Google AI Mode, Google AI Overview).
Sub-Actor access
The Actor orchestrates four sub-Actors. The skill needs two of them for output enrichment; the runner script's --fetch-sub-datasets flag walks the Sub-Actor results index and downloads them.
| Sub-Actor | Used to populate |
|---|---|
| Google Search Scraper | SERP Position (rank within organic results), Article Title, Publish Date. Also drives the per-engine result fetch when enableChatGpt, enableGemini, enableCopilot, enablePerplexity, enableAiMode are true. Join on url + query. |
| Website Content Crawler | Placement Source Sentence, Placement With Link, Placement New Insertion (needs page body text). Cross-check for Article Author via metadata.author / openGraph article:author / JSON-LD Person.name. |
| Contact Details Scraper | Already merged into the All leads dataset. Rarely needed directly. |
| Email Verifier (sub-Actor) | Runs only when enableEmailVerification: true. Populates email_verification on each lead. |
| AI Web Scraper | Only runs when searchAuthorName: true. Already merged into the Author list dataset. |
Output column to dataset map
| Output column | Source |
|---|---|
SERP Position | Google Search Scraper sub-dataset (rank within organic results, joined by URL + keyword). "-" if the row didn't appear in organic SERP. |
Source Engines | Mentions dataset. Comma-join the engines where the corresponding flag is true (e.g., OrganicResult_mention → "Google Organic", ChatGPT_mention → "ChatGPT", AIMode → "Google AI Mode", AIOverview_mention → "Google AI Overview", Perplexity_mention → "Perplexity"). |
Keyword | Mentions dataset queries[0] (or join row to the query that surfaced it). |
Article Title | Google Search Scraper sub-dataset title. |
Article URL | Mentions dataset url. |
Domain | All leads dataset domain. |
Article Author | Author list dataset (primary). Fallback to WCC sub-dataset metadata.author / openGraph / JSON-LD when missing. |
Author Source | searchAuthorName if filled from Author list; metadata.author / openGraph / jsonld for fallback paths; "not found" otherwise. |
Publish Date | Google Search Scraper sub-dataset publishDate or WCC sub-dataset metadata.publishedAt. |
Contact Full Name | All leads dataset firstName + ' ' + lastName. |
Contact Job Title | All leads dataset jobTitle. |
Department | All leads dataset departments[0]. |
Seniority | All leads dataset seniority. |
Contact Email | All leads dataset email. |
Contact LinkedIn | All leads dataset linkedinProfile. |
Company | All leads dataset companyName. |
Partnership Offer | User-supplied at Step 2. Same value for every row in a run. |
Suggested Link Placement | Agent-generated using WCC page body. |
Suggested Email Copy | Agent-generated using brand voice + article context. |
Outreach Status | Constant default "Not started". |
Notes | Agent-generated flags (own-domain, competitor, vendor page, UGC) plus manual-lookup hints. |
Cost notes
Per the Actor's README, billing fan-outs through the sub-Actors. Rough order-of-magnitude per single-query run with organicResult: 10:
- Parent Actor: ~$0.02 per query.
- Google Search Scraper: $4.5 per 1k organic results + per-query for each enabled LLM source (ChatGPT, Gemini, Copilot, Perplexity, AI Mode). AI Overviews is free.
- Website Content Crawler: runs once per surviving source.
- Contact Details Scraper: runs only for sources without a backlink. Bounded by
maxContactsPerDomain. - Email Verifier (only when
enableEmailVerification: true): typically much cheaper than the other sub-Actors (~$1 per 1k emails), but adds up at scale. - AI Web Scraper (only when
searchAuthorName: true): $25 per 1k sources.
To cut cost during testing: one or two queries, organicResult: 5, maxContactsPerDomain: 1, disable some LLM sources (keep enableAiOverviews since it's free), leave searchAuthorName: false. Keep enableEmailVerification: true unless quota is tight — bad emails skipped early save more than the verification cost.
Email Templates
Templates are keyed by outreach type (determined per-row from Why This Prospect tags + the user's goal — see SKILL.md Step 8 rule 4). The user's partnership type answer at Step 2 substitutes into the {{offer_paragraph}} placeholder inside each template.
Outreach type controls the opening hook and pitch structure. Partnership type controls what the user is offering in return.
Hard rules (apply to every template)
- Word cap: 150 words total, including subject line.
- Open with a concrete reference to the article — title plus one specific takeaway from its content. No generic "I loved your article" openers.
- Include the verbatim placement wording in the body. This is non-negotiable. Pull the source sentence (quoted, verbatim from the article) and the proposed change (linked version for drop-in, full follow-on sentence for additive, full drafted paragraph for new insertion) directly into the email. The recipient must see the exact text on first read, not be asked to click through to a separate document or reply for "details". Vague phrasing like "happy to send exact wording" / "happy to draft for you" / "a follow-on linking to..." is a content-skill bug. Always rewrite to embed the proposal.
- Never fabricate the author's name. If
Article Author = "Not found", address the contact by their first name from the All leads dataset. - Never fabricate stats or quotes about the user's product. If the user didn't give you a number, don't invent one.
- Never suggest external lookup tools or workarounds (hunter.io, LinkedIn search, third-party verifiers, etc.) — in the email body or in Notes. State facts; don't coach the user on tools they already know about.
Placement priority (try in order)
Every per-row email body references one of three placement strategies, tried in this order. The placement is drafted in Step 8 and surfaced in the output's three placement columns; the email body should match the strategy that was used.
1. Drop-in (preferred) — The user's URL is added to a word/phrase the author already wrote. No prose changes, no new sentences. The email's "natural place for it" line is something like "as a hyperlink on the word 'BeautifulSoup' in your tools-comparison paragraph". This is the lowest-friction ask: "could you turn one of your existing words into a hyperlink?"
2. Additive (second choice) — Keep an existing sentence verbatim; add one new sentence after it that introduces an adjacent reader-need the article doesn't already address. The email's placement line is something like "as a follow-on after your sentence about API integration — 'In need of competitor data the API doesn't expose? Then you need [Brand]'". The new sentence must raise a need the existing prose doesn't, in the article's voice, ≤25 words.
3. New insertion (last resort) — A fully drafted 1–2 sentence paragraph at a precise anchor. Use only when no relevant sentence exists in the article body. The email's placement line names the H2/section and the anchor sentence: "as a new paragraph after the sentence ending in '…each platform has its own trade-offs.' in the Honorable Mentions section".
Match the email's placement language to the strategy. A drop-in pitch ("could you hyperlink one existing word") reads very differently from a new-insertion pitch ("here's a paragraph I drafted in your voice"). Editors notice when the ask matches the effort.
Outreach-type templates
unlinked-mention-claim
Used when Why This Prospect includes Mentions brand, no backlink. The page already references the user's brand — they just forgot the link. Easiest ask of the five types; reply rates are highest. Keep it short.
Subject: Quick fix on your "{{article_title}}" — missed link?
Body skeleton:
Hi {{first_name}},
I noticed your "{{article_title}}" mentions {{user_brand}} in {{specific_mention_location}} — thanks for the shout-out.
Looks like the mention isn't linked. Would you be open to adding a link to {{user_content_url}}? It would help readers who want more on {{specific_topic}}.
{{offer_paragraph}}
Either way, appreciated the mention.
{{user_first_name}}Substitution rules:
{{specific_mention_location}}must quote or paraphrase the actual sentence where the brand is mentioned — pull from the WCC page body. If you can't find the mention text in the WCC body, flag the row withNotes: "Mentions dataset says brand was mentioned but WCC body doesn't contain it — verify manually".- This is the one outreach type where a unilateral ask (no reciprocal link) is perfectly normal — many publishers fix unlinked mentions without expecting anything. If the partnership type is
Unilateral ask, leave{{offer_paragraph}}as a single line:It's a small change and I'm not asking for anything in return.
competitor-link-replacement
Used when Why This Prospect includes Links to competitor. The page already links to a similar resource — the pitch is to add (or swap to) the user's URL.
Subject: Updated alternative to {{competitor_domain}} in your "{{article_title}}"
Body skeleton:
Hi {{first_name}},
In your "{{article_title}}" you link to {{competitor_domain}} in the {{specific_subsection_or_claim}} section — solid pick, though I noticed it's missing {{specific_gap_or_angle}}.
I'm {{user_first_name}} from {{user_brand}}. We published {{user_content_title}} at {{user_content_url}} which covers {{specific_gap_or_angle}} directly.
{{offer_paragraph}}
The natural place for it: {{placement_summary}}.
Worth a look?
{{user_first_name}}Substitution rules:
{{competitor_domain}}is the specific competitor URL the page links to — pull from the WCC body's outbound link list.{{specific_gap_or_angle}}should be one concrete way the user's content differs from the competitor's. Ask the user for this during Step 2 (brand voice) if not already supplied. Never invent a "gap" — leave the placeholder and flag inNotesif the user didn't give you anything.{{placement_summary}}is a one-line description of the placement (e.g.,"alongside your existing X mention in the comparison list"). Pull from thePlacement With LinkorPlacement New Insertioncell.- Do NOT say "remove the link to {{competitor_domain}}" — asking the publisher to delete an existing link is a much harder ask and usually kills the reply. Frame as addition, not replacement.
resource-page-inclusion
Used when Why This Prospect includes Resource / roundup page. The page is a curated list (best-of, top-10, tool roundup). The ask is inclusion on the list.
Subject: Suggestion for your "{{article_title}}" roundup
Body skeleton:
Hi {{first_name}},
I just went through your list "{{article_title}}" — tight selection, especially {{specific_listed_item_or_section}}.
I'm {{user_first_name}} from {{user_brand}}. We built {{user_content_title}} ({{user_content_url}}) which covers {{specific_gap_or_angle}} — I think it'd sit naturally alongside {{adjacent_listed_item}}.
{{offer_paragraph}}
Happy to send a one-line description if it'd save you time.
Thanks for keeping the list updated,
{{user_first_name}}Substitution rules:
{{specific_listed_item_or_section}}and{{adjacent_listed_item}}must come from the actual page body. If the WCC dataset doesn't have enough body text to identify list items, setOutreach Status = "Skip"withNotes: "Resource page detected but list structure unreadable — manual review needed".- Roundup pages typically get unilateral asks. If the partnership type is
ABC link exchangeorDirect A B link exchange, the offer often weakens the pitch — note this inNotesand consider downgrading toUnilateral askfor the row.
outdated-content-replacement
Used when Why This Prospect includes Outdated content. Article was published 2+ years ago and could use a refresh. The pitch is an updated resource the editor can plug into the post.
Subject: Refresh idea for "{{article_title}}" (published {{publish_year}})
Body skeleton:
Hi {{first_name}},
Re-read your "{{article_title}}" — the {{specific_subsection_or_claim}} section still holds up, but a few of the data points are from {{publish_year}} and {{specific_outdated_claim}}.
I'm {{user_first_name}} from {{user_brand}}. We just published {{user_content_title}} ({{user_content_url}}) with current numbers on {{specific_topic}}.
{{offer_paragraph}}
Suggested swap: {{placement_summary}}.
Would a refresh be useful?
{{user_first_name}}Substitution rules:
{{publish_year}}is the four-digit year fromPublish Date.{{specific_outdated_claim}}should reference something concrete in the article that has plausibly changed — pricing, a tool name, a stat. If the WCC body doesn't give you a clear outdated claim, drop the second sentence and pivot to a softer "could use a refresh" framing. Do NOT invent specific outdated claims.
topical-niche-edit
Default fallback when Why This Prospect has no specific tag (or only Top-3 SERP). The article is topically relevant — the pitch is a clean addition of the user's URL in an existing section.
Subject: Addition to "{{article_title}}"?
Body skeleton:
Hi {{first_name}},
Your "{{article_title}}" is one of the clearer pieces on {{topic}} I've found — particularly {{specific_subsection_or_claim}}.
I'm {{user_first_name}} from {{user_brand}}. We recently published {{user_content_title}} at {{user_content_url}} — same topic, {{angle_one_liner}}.
If it fits, the natural place for it is {{placement_summary}}.
{{offer_paragraph}}
Open to it?
{{user_first_name}}Substitution rules:
{{angle_one_liner}}should be one sentence the user supplied at Step 2 describing how their content differs from the prospect's. Ask if not supplied. Do not invent.{{placement_summary}}is one line drawn from thePlacement With LinkorPlacement New Insertioncell.
Partnership-type offer paragraphs
The user's Step 2 partnership type answer substitutes into {{offer_paragraph}}. Use the matching block verbatim, with the listed substitutions:
ABC link exchange
In exchange, I'd link to your article from {{user_content_url}}, and a partner site I work with ({{partner_domain}}) would add a link to {{user_article_url}}.Substitution rules:
{{partner_domain}}is filled out-of-band by the user. The skill leaves the placeholder and adds toNotes:"Replace {{partner_domain}} before sending — three-way deal needs a confirmed partner."
Direct A B link exchange
In exchange, I'd link to your article from {{user_content_url}} where it fits. Straight two-way swap, no third party.If the prospect domain has Domain DR ≥ 70, flag in Notes: "High-DR prospect — direct exchange may be seen as old-school SEO. Consider switching this row's partnership type to Unilateral ask."
Resource page / list inclusion
I'm not asking for a reciprocal link — happy to send this in as a straight suggestion for your list.(This offer paragraph essentially becomes unilateral. Used only when the user explicitly picks Resource page / list inclusion at Step 2 across all rows. Per-row overrides happen in the outreach-type templates themselves.)
Unilateral ask (no reciprocal)
I'm not asking for anything in return — just thought it'd be useful for your readers.Other
The user typed their own offer at Step 2. Use the literal language verbatim. Do not soften, summarise, or substitute synonyms. If they wrote "we pay $300 per placement", the email must say "we pay $300 per placement" — not "we offer competitive compensation".
If the user gave no custom offer text, fall back to the Unilateral ask paragraph and flag in Notes: "No custom offer supplied — defaulted to unilateral ask."
Brand voice substitution
The user's brand voice paragraph (from Step 2 input #1) governs every word that isn't a template placeholder. Apply these rules:
1. Keep their adjectives. If they wrote "casual, helpful, slightly nerdy", use "casual" not "informal", "helpful" not "useful", "nerdy" not "technical". 2. Mirror sentence length. If their voice paragraph uses short sentences, generate short sentences. If they write in long flowing sentences, do the same. 3. Mirror formality register. "Founder-led casual" allows "Hey" openings, contractions, em-dashes. "Formal B2B" rules out contractions and "Hey". Match what they used in their own paragraph. 4. Mirror idiom and slang. Don't substitute their phrases with generic synonyms. If they say "we're not in the AI hype train business", that phrasing must appear in the email — verbatim or with minimal adjustment to fit grammar. 5. Match opener style. If their voice paragraph implies they'd open with "Hey there," use that. If it implies "Hi {{first_name}},", use that. 6. Match close style. Same — pull the close phrasing from the voice paragraph. If absent, use a neutral close (Best, or just the first name) and surface a Notes flag so the user can replace it.
When the user skipped voice input entirely, use generic-professional defaults (Hi {{first_name}},, no contractions, neutral adjectives, Best, close) and add a Notes row flag: "Voice not specified — generic-professional default used".
Output Formats
The skill supports two output formats: xlsx (spreadsheet file) and markdown (table + email drafts inline in chat). Both share the same 30-column row schema and produce a run_metadata.json sidecar.
Row schema (30 columns)
| # | Column | Type | Source | Notes |
|---|---|---|---|---|
| 1 | SERP Position | int or "-" | Google Search Scraper sub-dataset | Rank within organic results for the row's keyword. "-" for rows that surfaced only via AI engines. |
| 2 | Source Engines | string | Mentions dataset | Comma-joined list of engines (only those enabled in Step 1 input #9): "Google Organic", "ChatGPT", "Gemini", "Copilot", "Perplexity", "Google AI Mode", "Google AI Overview". |
| 3 | Keyword | string | Mentions dataset queries[0] | The keyword that surfaced the source. If a source appears for multiple keywords, emit one row per keyword. |
| 4 | Article Title | string | Google Search Scraper sub-dataset title (primary); WCC metadata.title (fallback). | |
| 5 | Article URL | URL | Mentions dataset url | |
| 6 | Domain | string | All leads dataset domain | |
| 7 | Domain DR | int 0–100 or "-" | Ahrefs site-explorer-domain-rating | "-" if Ahrefs has no data; do not fabricate. |
| 8 | Page Traffic | int or "-" | Ahrefs site-explorer-metrics (mode=exact, target=Article URL) | Monthly organic visits to the specific article URL, last 30 days. "-" if Ahrefs has no data. |
| 9 | Referring Domains | int or "-" | Ahrefs site-explorer-backlinks-stats (target=Domain) | Domain-level refdomain count. "-" if Ahrefs has no data. |
| 10 | Prospect Tier | enum | Computed from DR + Page Traffic + user goal | One of: A, B, C. See SKILL.md Step 5 for thresholds. Empty for rows where Ahrefs failed. |
| 11 | Why This Prospect | string | Computed (see SKILL.md Step 7) | One or two comma-joined tags, ordered by pitch strength. "-" when no tag fits. |
| 12 | Article Author | string | Author list dataset (primary); WCC metadata fallback. | "Not found" if unknown. Never fabricate. |
| 13 | Author Source | enum | how the author was sourced | One of: searchAuthorName, metadata.author, openGraph, jsonld, not found. |
| 14 | Publish Date | ISO date or "Not found" | Google Search Scraper sub-dataset publishDate (primary); WCC metadata.publishedAt (fallback). | |
| 15 | Contact Full Name | string | All leads dataset firstName + ' ' + lastName, prioritised by editorial-leaning job title (see SKILL.md Step 8 rule 2) | |
| 16 | Contact Job Title | string | All leads dataset jobTitle | |
| 17 | Department | string | All leads dataset departments[0] | |
| 18 | Seniority | string | All leads dataset seniority | e.g., vp, head, manager. |
| 19 | Contact Email | string | All leads dataset email | "Not found" if unknown. Never fabricate. |
| 20 | Email Verification | enum | All leads dataset email_verification (when enableEmailVerification: true) | One of: verified, catch-all, risky, invalid, unknown, - (when verification didn't run). invalid triggers an auto-skip in Step 6 rule 11; catch-all / risky / unknown are informational and surface a Notes hint. |
| 21 | Contact LinkedIn | URL | All leads dataset linkedinProfile | |
| 22 | Company | string | All leads dataset companyName | |
| 23 | Outreach Type | enum | Computed from Why This Prospect + goal (see SKILL.md Step 8 rule 4) | One of: unlinked-mention-claim, competitor-link-replacement, resource-page-inclusion, outdated-content-replacement, topical-niche-edit. Empty for skipped rows. |
| 24 | Partnership Offer | string | User-supplied at Step 2 | Same value for every row. |
| 25 | Placement Source Sentence | string | Agent-generated from WCC page body | The verbatim sentence from the article where the link will go. Filled in strategies 1 (drop-in) and 2 (additive); "-" for strategy 3 (new insertion). |
| 26 | Placement With Link | string | Agent-generated | For drop-in: the source sentence with the link added on an existing word (no other text change). For additive: the source sentence kept verbatim + one new follow-on sentence containing the link. "-" for new insertion. |
| 27 | Placement New Insertion | string | Agent-generated | A drafted 1–2 sentence paragraph in the article's voice with a precise insertion location. Used only when no existing sentence is a natural fit (strategy 3); "-" otherwise. |
| 28 | Suggested Email Copy | string | Agent-generated | Subject + body, separated by \n---\n. ≤150 words including subject. Brand-voice matched. Outreach-type template per Outreach Type column. Empty for skipped rows. |
| 29 | Outreach Status | enum | Default "Not started"; "Skip" for rows dropped in the skip pass | |
| 30 | Notes | string | Agent-generated | Placement strategy tag (Placement: drop-in / additive / new insertion), skip reason if skipped, email-verification hints if status is non-verified, auto-flags, manual-lookup hints when fields are "Not found", alternate contacts. |
Placement column rule (critical)
Exactly one of columns 24, 25, 26 is filled per non-skipped row:
- If the article already contains a sentence that the link fits naturally → fill 25 (
Placement With Link) AND 24 (Placement Source Sentence); leave 26 as"-". - If no existing sentence is a natural fit but the article topic still supports a link → fill 26 (
Placement New Insertion); leave 24 and 25 as"-". - If the article topic is wrong for the user's URL entirely → set
Outreach Status = "Skip"withNotes: "No natural placement — article topic mismatch"and leave all three placement columns as"-".
xlsx rendering
The runner script writes a styled .xlsx to the path passed via --output with two sheets:
Sheet 1: Outreach (active rows, full schema)
This is the send-ready deliverable. Only rows with Outreach Status != "Skip" appear here, with the full 30-column row schema documented above.
Sheet 2: Skipped (filtered rows, reduced columns)
For auditing what the pipeline filtered. Reduced 6-column schema so missing Ahrefs / placement / email data doesn't visually clutter the view:
| # | Column | Source |
|---|---|---|
| 1 | Domain | Same as Outreach sheet col 6 |
| 2 | Article URL | Same as Outreach sheet col 5 |
| 3 | Article Title | Same as Outreach sheet col 4 |
| 4 | Skip Reason | Extracted from Outreach sheet col 30 (Notes) — the part after SKIP: and before any ` |
| 5 | Source Engines | Same as Outreach sheet col 2 |
| 6 | Why This Prospect | Same as Outreach sheet col 11 |
User opens the file → lands on Outreach by default → sees only actionable rows. Switches to Skipped when they want to audit what was filtered or recover a borderline row manually.
Common Outreach-sheet styling:
- Header row in bold, frozen.
- Column widths (in order, columns 1–30):
- 12 (SERP Position), 32 (Source Engines), 24 (Keyword), 60 (Article Title), 70 (Article URL),
- 24 (Domain), 10 (Domain DR), 14 (Page Traffic), 14 (Referring Domains), 12 (Prospect Tier),
- 40 (Why This Prospect), 24 (Article Author), 14 (Author Source), 12 (Publish Date),
- 24 (Contact Full Name), 30 (Contact Job Title), 16 (Department), 14 (Seniority),
- 30 (Contact Email), 16 (Email Verification), 50 (Contact LinkedIn), 24 (Company), 24 (Outreach Type), 30 (Partnership Offer),
- 80 (Placement Source Sentence), 80 (Placement With Link), 80 (Placement New Insertion),
- 100 (Suggested Email Copy), 16 (Outreach Status), 60 (Notes).
- Cell wrap-text enabled for the long columns (Article Title, Why This Prospect, Placement Source Sentence, Placement With Link, Placement New Insertion, Suggested Email Copy, Notes).
"Not found"and"-"rendered in italic.- Row sort:
Prospect Tierascending (A first), thenDomain DRdescending, thenSERP Positionascending. Skipped rows render last. - Conditional formatting on
Prospect Tier: green = A, yellow = B, grey = C, red strikethrough = skipped.
The runner writes only the columns it can populate from the Actor datasets (columns 1–6, 12–21, 28). The agent must then: 1. Load the .json after the Actor run finishes. 2. Run Step 5 (Ahrefs enrichment) → fill columns 7, 8, 9, 10. 3. Run Step 6 (skip pass) → fill column 28 + 29 for skipped rows. 4. Run Step 7 (Why This Prospect) → fill column 11. 5. Run Step 8 (placement + email + outreach type) → fill columns 22, 24, 25, 26, 27. 6. Re-save the .xlsx.
See scripts/run_actor.js for the round-trip pattern.
Markdown rendering
The agent renders this directly in chat. Structure:
# Link prospecting results — <date>
Run ID: <apify-run-id>
Keywords: <comma-joined>
Brand: <brand name>
Content URL: <user url>
Goal: <user's goal>
Partnership: <partnership type>
Tier breakdown: A: <n>, B: <n>, C: <n>, Skipped: <n>
| # | Tier | Why | SERP | Domain | DR | Traffic | Article | Contact | Outreach Type | Placement |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | A | Links to competitor (X.com) | 3 | example.com | 72 | 1,200 | [Title](url) | Mark Lee (Head of Content) — mark@example.com | competitor-link-replacement | (sentence diff or insertion preview, truncated to 80 chars) |
| ... |
## Email drafts
### Row 1 — Tier A — example.com — competitor-link-replacement — Mark Lee
**Subject:** ...
> Hi Mark,
>
> ...
Placement:
- Source: "Several tools handle this well, including X and Y."
- With link: "Several tools handle this well, including X, Y, and **[Brand](URL)**."
### Row 2 — ...Skipped rows render in a separate collapsed section at the bottom:
## Skipped (<n>)
| Domain | Reason |
|---|---|
| example.org | Stale content (published 2018-04-11) |
| ... |Keep the main Markdown table to the most actionable columns (the 11 above) and surface the rest via the JSON sidecar if the user asks. Output ALL columns in xlsx.
run_metadata.json sidecar
Always emit this alongside the main output, even in Markdown mode (write it to disk in the same directory the user ran from):
{
"runId": "abc123...",
"actorId": "apify/link-prospecting-tool",
"startedAt": "2026-05-13T10:15:00Z",
"finishedAt": "2026-05-13T10:23:11Z",
"inputs": {
"goal": "Topical authority links to specific URL",
"queries": "best search engine\nalternative to google",
"brand": "Acme",
"ownDomains": ["acme.com"],
"competitorDomains": [],
"alreadyPitchedDomains": [],
"organicResult": 10,
"maxContactsPerDomain": 3,
"department": ["marketing"],
"searchAuthorName": true,
"includeMention": true,
"enableChatGpt": true,
"enableAiMode": true,
"enableAiOverviews": true,
"enablePerplexity": true
},
"datasetIds": {
"default": "...",
"mentions": "...",
"domainsWithLeads": "...",
"authors": "...",
"subActors": "...",
"googleSearch": "...",
"websiteContentCrawler": "..."
},
"tierCounts": { "A": 8, "B": 15, "C": 7 },
"skipCounts": {
"goalMismatch": 5,
"alreadyPitched": 2,
"staleContent": 3,
"nonEditorialPage": 1,
"ugc": 0,
"noContactNoAuthor": 1,
"topicMismatch": 0
}
}This is what makes the outreach work reproducible. Reference it when the user asks "where did row 7 come from?" — they (or you) can re-fetch each dataset from Apify storage by the IDs.
Troubleshooting
Common failures and how to fix them. Listed roughly in order of how often they hit.
Actor run TIMED-OUT
This is the single most common failure for apify/link-prospecting-tool. Public stats at time of skill authoring show 18 of 24 recent runs ending in TIMED-OUT.
There are two different timeouts and only one of them is the problem:
- Actor run-level `timeoutSecs` — Apify-side, defaults to 60000 (~16h). Almost never the cause.
- API client wait timeout — your-side, how long
run_actor.jspolls the run before giving up. This is what trips. The skill exposes it via--timeout(seconds).
Fixes, cheapest first:
1. If the Actor is still RUNNING on Apify when the client gives up — don't restart. Use scripts/fetch_run_artifacts.js --run-id <id> --output <file> --timeout 1800 to poll the existing run and download all artifacts when it terminates. Same output shape as run_actor.js --fetch-sub-datasets. This is the most common case — runs frequently take 20-50 min on multi-keyword + multi-LLM-engine campaigns. 2. Raise --timeout for the next run. The default is now 1800; on big campaigns (5+ keywords, all LLM engines on, email verification on) push to 2700 or 3600. 3. Lower organicResult to 5. Each result triggers a WCC + Contact Details Scraper fan-out — halving this halves the bottleneck. 4. Cut the number of queries. Run two keywords at a time, not ten. 5. Disable AI platforms that you don't need (enableChatGpt: false, enableAiMode: false, enablePerplexity: false). AI Overviews stays on — it's free. 6. Lower maxContactsPerDomain to 1 for the first run; raise it only on subsequent runs against a known-good keyword.
Past run durations (for calibration):
- 3 keywords, all LLM engines off: ~20 min.
- 5 keywords, partial LLM coverage: ~52 min.
- 1 keyword, all LLM engines + email verification: exceeded 900s client wait at the 16-min mark with the default dataset still empty; recovered via
fetch_run_artifacts.js.
Error: 'brand' is required (or similar 400 from the Actor)
You called the Actor without the brand field. Re-prompt the user for their brand name (Step 1 anchor input #2) and rerun.
The Actor's required fields are queries and brand. Everything else has a default.
Author = Not found
Expected for ~30% of pages. Many publishers don't expose author bylines in machine-readable form, and the AI Web Scraper sub-Actor only finds names when they're textually visible.
The skill writes "Not found" to the column and "No author detected" to Notes. Never invent a name. Do not suggest external lookup tools or workarounds in Notes — see SKILL.md Step 6 rule 11 for the rationale (the user knows where to look; the skill's job is to state what it found, not coach the user on third-party tools).
A secondary check the skill can run automatically: look in the WCC sub-dataset for the page's metadata.author, openGraph article:author, and JSON-LD Person.name. If any of those are populated, use them and set Author Source accordingly. The runner script's --fetch-sub-datasets flag already downloads the WCC dataset; the agent just needs to join on URL.
Contact Email = Not found
The Contact Details Scraper sub-Actor missed this domain. Reasons it misses:
- The domain has no public team page or contact page.
- The published contacts are in roles outside the configured
departmentfilter (the skill defaults tomarketing— see SKILL.md Step 3). - The site uses contact forms only.
Skill response: write "Not found" and add Notes: "No email found for this contact". Do not suggest external lookup tools or workarounds. Optional automatic fix: rerun with department widened — e.g., ["marketing", "sales", "operations", "consulting"].
0 leads returned
The whole run completed but the All leads dataset is empty. Causes:
- Keyword too narrow (no SERP coverage).
ownDomains+competitorDomains+ignoreDomainsfiltered out every result.- Every surviving source already mentions or backlinks the user's brand (the Actor de-prioritises these for outreach).
Diagnostic steps:
1. Check the Mentions dataset — if it has rows, the filtering is the issue. 2. Loosen ignoreDomains (remove some entries). 3. Drop competitorDomains for one diagnostic run. 4. Raise organicResult to 20. 5. Broaden the keyword (drop modifiers, try plurals, try the head term).
Cannot find module 'xlsx'
The xlsx dependency wasn't installed. Run npm install inside the skill's scripts/ folder once. The runner script's package.json declares the dep.
APIFY_TOKEN not found
The runner expects a .env file in the current working directory with APIFY_TOKEN=.... Get one from https://console.apify.com/account/integrations and create the file.
node --env-file=.env requires Node.js 20.6+.
Costs higher than expected
The Actor's billing fan-outs through sub-Actors. A single run of organicResult: 10 with all four AI platforms enabled and searchAuthorName: true can easily cost 5-10x what the parent Actor's $0.02/query suggests.
Cost-shaving levers, ordered by impact:
1. searchAuthorName: false — saves $25 per 1k sources (this is the biggest single cost in many runs). 2. organicResult: 5 instead of 10 — halves the WCC + Contact Details Scraper cost. 3. enableChatGpt: false, enablePerplexity: false — each one adds a per-result Google Search Scraper cost. 4. maxContactsPerDomain: 1 — fewer Contact Details Scraper calls per domain. 5. Tighter ignoreDomains — exclude obvious non-targets before they get crawled.
enableAiOverviews is free (parsed from the SERP that's already fetched), so keep it on regardless.
Run completed but only some sub-datasets were fetched
The --fetch-sub-datasets flag walks the Sub-Actor results index and downloads four sibling files: *_mentions.json, *_authors.json, *_serp.json, *_wcc.json. If one is missing:
*_authors.jsonmissing →searchAuthorName: falsefor this run. Re-enable and rerun if you need author names.*_wcc.jsonmissing → all sources were filtered out before crawling. Almost always paired with0 leads returned. See that section.*_serp.jsonmissing → no organic SERP scraping happened. Should not occur unless the Actor's behavior changed; check the parent run's Sub-Actor index manually in the Apify console.
"Brand mentioned everywhere — no outreach targets"
If the user's brand is already widely cited, the All leads dataset will be small even on a successful run. This is a feature, not a bug — the Actor specifically routes leads to sources that don't yet link back. To find sources that mention but don't link:
- Confirm
includeMention: true(default). This pulls in mention-only sources for outreach. - If the user wants pure cold (no mentions yet), set
includeMention: false— but then mention-only sites are also excluded.
Domain in ownDomains still shows up as a prospect
The Actor's domain matching is case-insensitive but exact-suffix. If the user owns acme.com and the source is on blog.acme.com, the filter catches it. If the user owns acme.io and the source is on getacme.com, the filter misses — add the variant explicitly.
#!/usr/bin/env python3
"""Build the unified prospect table from the Actor's datasets.
Reads campaign config + 4 Actor sidecar files; writes one row per WCC URL.
Usage:
python3 scripts/build_prospects.py --config campaign.json
Reads:
{base}.json (main leads)
{base}_serp.json (Google + LLM-engine SERP results)
{base}_wcc.json (Website Content Crawler bodies)
{base}_authors.json (AI Web Scraper author results)
Writes:
{base}_prospects.json
"""
import argparse
import json
import re
from collections import defaultdict
from urllib.parse import urlparse, parse_qsl, urlencode, urlunparse
EDIT_RE = re.compile(r'editor|content|writer|managing|editorial|blog|copy|journalist|head of content|redaktor|šéfredaktor', re.I)
DEMOTE_RE = re.compile(r'\bceo\b|\bcfo\b|\bcto\b|\bcoo\b|founder|chief|\bvp\b|president|jednatel|majitel', re.I)
SENIORITY_RANK = {"head": 5, "director": 4, "vp": 3, "senior": 4, "manager": 3,
"c_suite": 1, "entry": 2, "intern": 1, None: 0, "": 0}
TRACKING_PARAMS = {"utm_source", "utm_medium", "utm_campaign", "utm_term", "utm_content",
"fbclid", "gclid", "msclkid", "yclid", "ref", "ref_src"}
ENGINE_ORDER = [
"Google Organic", "ChatGPT", "Gemini", "Copilot", "Perplexity",
"Google AI Mode", "Google AI Overview",
]
def norm_domain(d: str) -> str:
if not d:
return ""
d = d.replace("https://", "").replace("http://", "").rstrip("/")
if d.startswith("www."):
d = d[4:]
return d.lower().split("/")[0]
def norm_url(u: str) -> str:
if not u:
return ""
try:
p = urlparse(u)
host = p.netloc.lower()
if host.startswith("www."):
host = host[4:]
q = [(k, v) for k, v in parse_qsl(p.query, keep_blank_values=True) if k.lower() not in TRACKING_PARAMS]
new = p._replace(scheme="https", netloc=host, query=urlencode(q), fragment="")
out = urlunparse(new)
if out.endswith("/") and out.count("/") > 3:
out = out[:-1]
return out
except Exception:
return u
def url_domain(u: str) -> str:
try:
return norm_domain(urlparse(u).netloc)
except Exception:
return ""
def score_contact(c):
title = c.get("jobTitle") or ""
sen = c.get("seniority") or ""
s = SENIORITY_RANK.get(sen, 0)
if EDIT_RE.search(title):
s += 100
elif DEMOTE_RE.search(title):
s -= 50
if c.get("email"):
s += 10
return s
def email_verification_status(contact):
if not contact:
return "-"
ev = contact.get("emailVerification") or {}
result = (ev.get("result") or "").lower()
return {
"ok": "verified", "valid": "verified", "invalid": "invalid",
"catch-all": "catch-all", "catch_all": "catch-all", "catchall": "catch-all",
"risky": "risky", "unknown": "unknown",
}.get(result, result or "-")
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--config", required=True)
args = ap.parse_args()
cfg = json.load(open(args.config))
base = cfg["base"]
own_domains = set(cfg.get("own_domains", []))
brand_aliases = cfg.get("brand", {}).get("aliases") or [cfg.get("brand", {}).get("name", "")]
brand_aliases = [a for a in brand_aliases if a]
leads = json.load(open(f"{base}.json"))
serp = json.load(open(f"{base}_serp.json"))
wcc = json.load(open(f"{base}_wcc.json"))
authors = json.load(open(f"{base}_authors.json"))
# Dedupe leads
seen = set()
dedup_leads = []
for l in leads:
k = l.get("personId") or l.get("email") or f"{l.get('firstName')}-{l.get('lastName')}-{l.get('domain')}"
if k in seen:
continue
seen.add(k)
dedup_leads.append(l)
print(f"Leads: {len(leads)} → {len(dedup_leads)} after dedupe")
# Index leads by domain + URL
leads_by_domain = defaultdict(list)
leads_by_url = defaultdict(list)
for l in dedup_leads:
d = norm_domain(l.get("domain", ""))
if d:
leads_by_domain[d].append(l)
for s in (l.get("source_url") or []):
u_norm = norm_url(s.get("url", ""))
if u_norm:
leads_by_url[u_norm].append(l)
# Authors by URL
authors_by_url = {}
for a in authors:
u_norm = norm_url(a.get("url", ""))
data = a.get("data") or {}
authors_by_url[u_norm] = {
"authorName": data.get("authorName"),
"publishDate": data.get("date"),
"postName": data.get("postName"),
}
# SERP + engine attribution
serp_lookup = {}
engines_by_url = defaultdict(set)
keyword_per_run = (serp[0].get("searchQuery", {}).get("term", "") if serp else "")
for entry in serp:
kw = (entry.get("searchQuery") or {}).get("term", "")
for r in entry.get("organicResults") or []:
u = norm_url(r.get("url", ""))
if not u:
continue
engines_by_url[u].add("Google Organic")
existing = serp_lookup.get(u)
if not existing or (r.get("position", 999) < existing["position"]):
serp_lookup[u] = {
"position": r.get("position"),
"title": r.get("title"),
"lastUpdated": r.get("lastUpdated"),
"keyword": kw,
}
for field, label in (
("aiModeResult", "Google AI Mode"),
("aiOverviewResult", "Google AI Overview"),
("perplexitySearchResult", "Perplexity"),
("chatGptSearchResult", "ChatGPT"),
("geminiSearchResult", "Gemini"),
("copilotSearchResult", "Copilot"),
):
block = entry.get(field) or {}
for src in (block.get("sources") or block.get("citationUrls") or []):
u = src.get("url", "") if isinstance(src, dict) else (src if isinstance(src, str) else "")
u_norm = norm_url(u)
if u_norm:
engines_by_url[u_norm].add(label)
# WCC by URL (canonical row list)
wcc_by_url = {}
for w in wcc:
u_norm = norm_url(w.get("url", ""))
if not u_norm:
continue
wcc_by_url[u_norm] = {
"original_url": w.get("url"),
"text": w.get("text") or "",
"markdown": w.get("markdown") or "",
"metadata": w.get("metadata") or {},
}
rows = []
for u_norm, wcc_entry in wcc_by_url.items():
original_url = wcc_entry["original_url"]
domain = url_domain(original_url)
if not domain:
continue
engine_set = engines_by_url.get(u_norm, set())
engines = [e for e in ENGINE_ORDER if e in engine_set]
s = serp_lookup.get(u_norm, {})
wcc_text = wcc_entry["text"]
wcc_meta = wcc_entry["metadata"]
# Brand mention: source_url[].brand_mentioned_in_source OR body-level match against any alias
brand_mentioned = False
for l in leads_by_url.get(u_norm, []):
for sl in (l.get("source_url") or []):
if norm_url(sl.get("url", "")) == u_norm and sl.get("brand_mentioned_in_source"):
brand_mentioned = True
if brand_aliases and any(re.search(r'\b' + re.escape(a) + r'\b', wcc_text, re.I) for a in brand_aliases):
brand_mentioned = True
# Author cascade: ai-web-scraper → wcc metadata.author → openGraph → jsonLd Person
author_entry = authors_by_url.get(u_norm, {})
article_author = author_entry.get("authorName")
author_source = "searchAuthorName" if article_author else None
if not article_author and wcc_meta.get("author"):
article_author = wcc_meta["author"]
author_source = "metadata.author"
if not article_author:
og = (wcc_meta.get("openGraph") or {}).get("article:author") if isinstance(wcc_meta.get("openGraph"), dict) else None
if og:
article_author = og
author_source = "openGraph"
if not article_author and wcc_meta.get("jsonLd"):
try:
for item in wcc_meta["jsonLd"]:
if isinstance(item, dict) and item.get("@type") == "Person" and item.get("name"):
article_author = item["name"]
author_source = "jsonld"
break
except Exception:
pass
if not article_author:
article_author = "Not found"
author_source = "not found"
article_title = s.get("title") or wcc_meta.get("title") or author_entry.get("postName") or "Not found"
publish_date = (wcc_meta.get("publishedAt") or wcc_meta.get("publishedTime")
or s.get("lastUpdated") or author_entry.get("publishDate") or "Not found")
# Contact pick: URL-level match first, then domain-level fallback
url_leads = leads_by_url.get(u_norm, [])
cs = sorted(url_leads or leads_by_domain.get(domain, []), key=score_contact, reverse=True)
contact = cs[0] if cs else None
alternates = cs[1:] if len(cs) > 1 else []
row = {
"Article URL": original_url,
"Domain": domain,
"Article Title": article_title,
"Article Author": article_author,
"Author Source": author_source,
"Publish Date": publish_date,
"SERP Position": s.get("position") if s.get("position") is not None else "-",
"Source Engines": ", ".join(engines) if engines else "-",
"Keyword": keyword_per_run,
"Keywords List": [keyword_per_run] if keyword_per_run else [],
"Brand Mentioned": brand_mentioned,
"Has Backlink": False, # set below from WCC outbound links
"Contact Full Name": (f"{contact.get('firstName','')} {contact.get('lastName','')}".strip()
if contact else "Not found"),
"Contact Job Title": (contact.get("jobTitle") if contact else "Not found") or "Not found",
"Department": (contact.get("departments") or ["-"])[0] if contact else "-",
"Seniority": (contact.get("seniority") if contact else "") or "-",
"Contact Email": (contact.get("email") if contact else "Not found") or "Not found",
"Email Verification": email_verification_status(contact),
"Contact LinkedIn": (contact.get("linkedinProfile") if contact else "") or "",
"Company": (contact.get("companyName") if contact else "") or domain,
"Alternate Contacts": [
f"{a.get('firstName','')} {a.get('lastName','')} ({a.get('jobTitle','')}, {a.get('email','no-email')})"
for a in alternates[:3]
],
"WCC Text": wcc_text,
"WCC OutboundLinks": [],
}
rows.append(row)
# Outbound links + Has-Backlink
LINK_RE = re.compile(r'\[([^\]]*)\]\((https?://[^)]+)\)')
for r in rows:
wcc_entry = wcc_by_url.get(norm_url(r["Article URL"]), {})
md = wcc_entry.get("markdown", "")
links = []
for m in LINK_RE.findall(md or ""):
anchor, link_url = m
dom = url_domain(link_url)
if not dom or dom == r["Domain"]:
continue
links.append({"anchor": anchor.strip(), "url": link_url, "domain": dom})
if dom in own_domains:
r["Has Backlink"] = True
r["WCC OutboundLinks"] = links
r["WCC OutboundLinkCount"] = len(links)
out_path = f"{base}_prospects.json"
with open(out_path, "w") as f:
json.dump(rows, f, indent=2, default=str)
print(f"Wrote {len(rows)} prospect rows → {out_path}")
no_contact = sum(1 for r in rows if r['Contact Email'] == 'Not found' and r['Contact Full Name'] == 'Not found')
print(f" No contact: {no_contact} | Brand mention: {sum(1 for r in rows if r['Brand Mentioned'])} | Backlink: {sum(1 for r in rows if r['Has Backlink'])}")
domains = sorted({r['Domain'] for r in rows})
print(f" Unique domains: {len(domains)}")
if __name__ == "__main__":
main()
{
"type": "module",
"dependencies": {
"xlsx": "^0.18.5"
}
}
Related skills
FAQ
What external tools does it need?
It runs the apify/link-prospecting-tool actor and calls Ahrefs MCP tools to score domain and page authority per prospect.
What campaign presets are available?
Recover unlinked brand mentions, Replace competitor links, Topical authority links to a URL, Maximum link volume, and Custom.