
Healthcare Providers Extract
- 57 installs
- 50 repo stars
- Updated July 29, 2026
- nimbleway/agent-skills
Helps with ai & agent building tasks.
About
healthcare-providers-extract is a Claude Code skill for ai & agent building. It helps solo builders move faster with AI-assisted development.
- healthcare-providers-extract
- AI & Agent Building
- AI-coding skill
Healthcare Providers Extract by the numbers
- 57 all-time installs (skills.sh)
- +4 installs in the week ending Aug 4, 2026 (Skillselion tracking)
- Ranked #6,669 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
- Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/nimbleway/agent-skills --skill healthcare-providers-extractAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 57 |
|---|---|
| repo stars | ★ 50 |
| Last updated | July 29, 2026 |
| Repository | nimbleway/agent-skills ↗ |
What it does
Helps with ai & agent building tasks.
Files
Healthcare Providers Extract
Structured practitioner extraction from healthcare practice websites, powered by Nimble's web data APIs.
User request: $ARGUMENTS
Before running any commands, read references/nimble-playbook.md for Claude Code constraints (no shell state, no &/wait, sub-agent permissions, communication style).
---
Instructions
Step 0: Preflight + WSA Discovery
Follow the transport selection + standard preflight from references/nimble-playbook.md — pick CLI or MCP at session start, then run the standard preflight calls (date calc, today, profile, memory index) in parallel.
Also simultaneously — run WSA discovery and setup:
mkdir -p ~/.nimble/memory/{reports,healthcare-providers-extract/checkpoints}ls ~/.nimble/memory/healthcare-providers-extract/checkpoints/ 2>/dev/null- Run Layer 1 (vertical) and Layer 3 (general tools) WSA discovery from
references/wsa-reference.md. Layer 2 (session-specific) runs after Step 1 when you know the user's specialty.
Classify discovered agents into phases and validate with nimble agent get per references/wsa-reference.md.
From the preflight results:
- CLI missing or API key unset ->
references/profile-and-onboarding.md, stop - Tag all
nimbleCLI calls:nimble --client-source skill-healthcare-providers-extract <subcommand>. MCP path: not yet supported — seereferences/nimble-playbook.mdfor status. - Profile exists -> note it for context. Determine mode using smart date windowing
from references/nimble-playbook.md:
- Full mode: first run OR last run > 14 days ago
- Quick refresh: last run < 14 days ago (re-extract only new/changed pages)
- Same-day repeat: if
last_runs.healthcare-providers-extractis today, check
for existing report at ~/.nimble/memory/reports/healthcare-providers-extract-*[today].md. If found, ask: "Already ran today. Run again for fresh data?"
- No profile -> that's fine. This skill doesn't require onboarding. Proceed to Step 1.
Step 1: Parse Input & Starting Questions
Parse $ARGUMENTS for input type using the Input Parsing Pattern from references/nimble-playbook.md. Key routing:
- URLs detected -> proceed to Step 3
- Specialty + location (no URLs) -> proceed to Step 2 (practice discovery)
- Unclear -> ask (counts as 1 of max 2 prompts)
If input is clear, confirm and ask one shaping question (plain text, not AskUserQuestion):
"Extracting providers from N practice sites. Quick questions:
1. Healthcare vertical? (ophthalmology, dental, dermatology, general, or other)
2. Quick scan (names + credentials only) or full extraction (all 5 fields)?"
If input is ambiguous, use AskUserQuestion (counts as 1 of max 2 prompts):
What practice sites should I extract providers from?
- Paste URLs directly (one per line)
- Provide a CSV file path or Google Sheet URL with practice URLs
- Or describe what you're looking for (e.g., "ophthalmologists in Austin, TX")
and I'll find practices first
Skip questions the user already answered in their initial message.
Step 2: Practice Discovery (Optional)
Only if the user provided a specialty + location instead of URLs.
Two input paths into discovery:
Path A — Fresh discovery. User gave specialty + location. Run Layer 2 WSA discovery for session-specific agents:
nimble agent list --limit 50 --search "[specialty]"
nimble agent list --limit 50 --search "[directory-user-mentioned]"See references/wsa-reference.md for the full discovery strategy, agent evaluation criteria, and healthcare discovery prioritization.
Run all discovery-phase agents simultaneously. Validate params with nimble agent get first.
Path B — Market-finder handoff. User ran market-finder first and wants to extract providers from those results. Read the market-finder output:
cat ~/.nimble/memory/market-finder/{slug}/entities.json 2>/dev/nullExtract practice records. Note: Google Maps results contain place_url (a Maps link) but not the practice's actual website URL. Proceed to Step 2b to resolve real website URLs before site mapping.
After either path: Deduplicate by domain. Present discovered practices:
"Found N practices for [specialty] in [location] across [M] data sources.
Proceeding to extract providers from these sites..."
Fallback — if no discovery WSAs were found, or results are sparse (< 3):
nimble search --query "[specialty] in [location]" --max-results 20 --search-depth liteStep 2b: Resolve Practice Website URLs
Discovery sources (Google Maps, Yelp, BBB) return listing URLs, not practice website URLs. Before site mapping, resolve the actual website for each practice:
1. Check structured data first — Google Maps results often include a website field in the structured output. Use it if present. 2. Extract from listing page — if no website field, extract the Maps listing to find the practice website link:
nimble extract --url "[maps-listing-url]" --format markdown3. Search fallback — if extraction fails:
nimble search --query "[practice-name] [city] official website" --max-results 3 --search-depth liteSkip practices where no website URL can be resolved — note them in the "Data Quality Summary" output section.
Step 3: Site Mapping
Follow the Site Mapping Pattern from references/nimble-playbook.md for each practice URL. Skill-specific settings:
- Keyword weight table:
references/provider-extraction-patterns.md - Page cap: 15 per site
- Fallback query:
site:[domain] doctors OR providers OR team
For 6+ practices, use sub-agents (see Sub-Agent Strategy below).
Save checkpoint: ~/.nimble/memory/healthcare-providers-extract/checkpoints/{slug}/mapping.json
Step 4: Page Extraction
WSA shortcuts first: If WSA discovery found agents that extract provider data from healthcare directories, use those for matching practices — structured WSA output is higher quality than parsed markdown.
For all other practices, follow the Page Extraction with Retry pattern from references/nimble-playbook.md. Scale using the Scaled Execution pattern from the same reference.
Save checkpoint: ~/.nimble/memory/healthcare-providers-extract/checkpoints/{slug}/extraction.json
Step 5: Structured Parsing
Parse extracted markdown to identify providers and their fields. Read references/provider-extraction-patterns.md for the 5 core fields, credential regex patterns, and specialty keywords.
For each extracted page: 1. Scan for provider name patterns (Dr. prefix, heading patterns, bold text near credential suffixes) 2. Match credentials using the regex patterns from references/provider-extraction-patterns.md 3. Match specialty using keywords for the detected healthcare vertical 4. Extract contact info (phone regex, appointment URLs, email) 5. Extract education/training mentions
Build structured records:
{
"name": "Dr. Jane Smith",
"credentials": "MD, FACS",
"specialty": "Retinal Surgery",
"contact": {"phone": "(555) 123-4567", "scheduling_url": "..."},
"education": "Fellowship: Bascom Palmer Eye Institute",
"source_url": "https://practice.com/our-doctors",
"practice_name": "Shore Center for Eye Care",
"practice_url": "https://practice.com",
"confidence": "High"
}Step 6: Deduplication & Confidence Scoring
Follow the Entity Deduplication and Entity Confidence Scoring patterns from references/nimble-playbook.md. Skill-specific dedup rules and the 5-field confidence criteria are in references/provider-extraction-patterns.md.
Step 7: Output
Present results grouped by practice, sorted by confidence within each practice.
# Provider Extraction: [N] Providers from [M] Practices
*[Date] | [H] High, [M] Medium, [L] Low confidence*
## TL;DR
Extracted [N] providers from [M] practice websites. [H] with complete profiles,
[L] with partial data. [Key finding: e.g., "12 of 15 providers are board-certified"].
## [Practice Name] ([domain])
| # | Name | Credentials | Specialty | Contact | Education | Confidence |
|---|------|------------|-----------|---------|-----------|------------|
| 1 | Dr. Jane Smith | MD, FACS | Retinal Surgery | (555) 123-4567 | Fellowship: Bascom Palmer | High |
| 2 | Dr. John Doe | OD | General Ophthalmology | [Book](url) | Residency: Wills Eye | Medium |
[Repeat per practice]
## Data Quality Summary
- **Complete profiles (High):** [N] providers
- **Partial profiles (Medium):** [N] providers — missing: [list common gaps]
- **Minimal profiles (Low):** [N] providers — missing: [list common gaps]
## Sources
[Clickable URL for every page extracted, grouped by practice]Source links are mandatory. Every provider record must trace back to a source URL.
Step 8: Save to Memory
Make all Write calls simultaneously:
- Report ->
~/.nimble/memory/reports/healthcare-providers-extract-{slug}-{date}.md - Provider data ->
~/.nimble/memory/healthcare-providers-extract/{slug}/providers.json - Profile -> update
last_runs.healthcare-providers-extractin
~/.nimble/business-profile.json (only if profile exists)
- Follow the wiki update pattern from
references/memory-and-distribution.md: update
index.md rows for all affected entity files, append a log.md entry for this run.
- Clean up checkpoint (complete run) or keep (partial run)
Step 9: Share & Distribute
Always offer distribution -- do not skip. Follow references/memory-and-distribution.md for connector detection and sharing flow.
Notion: full provider table as a dated subpage. Slack: TL;DR with provider count and confidence breakdown only.
Step 10: Follow-ups
- "Tell me more about Dr. X" -> show full extracted profile
- "Export as CSV" -> generate CSV from providers.json
- "Run on more sites" -> append new practice URLs, extract and merge
- "What's missing?" -> detail the data gaps per provider
Enrichment from discovered WSAs: If Step 0 found enrichment-phase agents (reviews, regulatory, practice details), offer them as immediate follow-ups:
"I also found [N] WSAs that could enrich this data: [brief list]. Want me to
run reputation checks or regulatory lookups on these providers/practices?"
See references/wsa-reference.md for enrichment phase mapping and fallback chains.
Sibling skill suggestions:
Next steps:
- Run healthcare-providers-enrich to fill data gaps (NPI lookup, boardcertification verification, additional contact info)
- Run healthcare-providers-verify to validate credentials and license status- Run market-finder to discover more practice URLs in this area---
Sub-Agent Strategy
For batch extraction (6+ practices), use nimble-researcher agents (agents/nimble-researcher.md) to parallelize site mapping and extraction.
Follow the sub-agent spawning rules from references/nimble-playbook.md (bypassPermissions, batch max 4, explicit Bash instruction, fallback on failure).
Spawn pattern: One agent per practice (or per batch of 3 practices for large jobs). Each agent runs Steps 3-5 for its assigned practices and returns structured provider records.
Single-practice optimization: If only 1-2 practices, run directly from the main context instead of spawning agents.
Fallback: If any agent fails, run those extractions directly from the main context. Never leave gaps in the output.
---
Error Handling
See references/nimble-playbook.md for the standard error table (missing API key, 429, 401, empty results, extraction garbage). Skill-specific errors:
- No provider pages found: "Couldn't find provider/team pages on [domain].
The site may list staff differently. Want me to try extracting from the homepage or search for this practice on healthcare directories?"
- All extractions returned garbage: "The practice sites appear to be heavily
JavaScript-rendered. Retrying with browser rendering..." (auto-retry with --render per the shared pattern)
- Ambiguous practice name: If a URL fails and the user provided a name instead,
search for the practice: nimble search --query "[practice name] [location] doctors" --max-results 5 --search-depth lite
- CSV/Sheet parse error: "Couldn't parse the input file. Expected a column with
practice URLs. Can you paste the URLs directly instead?"
Memory & Distribution
How skills persist knowledge across sessions and distribute reports to external tools.
---
Memory Architecture
All persistence lives under ~/.nimble/ — never touch user project files.
~/.nimble/
├── business-profile.json # Tier 1: Hot cache (see profile-and-onboarding.md)
└── memory/ # Tier 2: Deep storage (loaded on demand)
├── index.md # Global index (one line per directory)
├── log.md # Chronological activity log (append-only)
├── backlog.md # Research questions and knowledge gaps
├── synthesis/ # Cross-entity analysis pages
├── competitors/ # Accumulated intel per competitor
│ └── index.md # Per-directory entity catalog
├── people/ # Contact profiles for meeting prep
│ └── index.md
├── companies/ # Deep-dive research results
│ └── index.md
├── reports/ # Timestamped full skill outputs
├── positioning/ # Per-competitor positioning snapshots
│ └── index.md
└── glossary.md # Industry terms and jargonTier 1 (business-profile.json) — loaded every session. See references/profile-and-onboarding.md for the full schema and update patterns.
Tier 2 (memory/) — loaded on demand when a skill needs deeper context.
Wiki Primitives
The memory directory includes wiki-level files that make the knowledge base navigable, queryable, and self-maintaining:
~/.nimble/memory/
├── index.md # Global summary (one line per directory)
├── log.md # Chronological activity log (append-only)
├── backlog.md # Research questions and knowledge gaps
├── synthesis/ # Cross-entity analysis pages
│ ├── index.md # Per-directory catalog (same format as others)
│ └── competitive-landscape.md # (created dynamically when patterns emerge)
├── competitors/
│ ├── index.md # Per-directory entity catalog
│ ├── widgetco.md
│ └── gizmotech.md
├── people/
│ ├── index.md
│ └── alex-kim.md
├── companies/
│ ├── index.md
│ └── ...
├── reports/
├── positioning/
│ ├── index.md
│ └── ...
└── glossary.mdSkills create index files, log.md, and synthesis/ on first write if missing.
---
Wiki Content Index (Two-Tier)
Indexes live at two levels: a lightweight global index for cross-directory navigation, and per-directory indexes for detailed entity catalogs.
Global Index (~/.nimble/memory/index.md)
One line per directory — entity count and last-updated date. Never lists individual entities. Stays under 30 lines forever.
# Knowledge Index
| Directory | Entities | Last Updated |
|-----------|----------|-------------|
| [[competitors/index]] | 5 | 2026-03-20 |
| [[people/index]] | 3 | 2026-03-15 |
| [[companies/index]] | 8 | 2026-03-18 |
| [[positioning/index]] | 5 | 2026-03-20 |
| [[synthesis/index]] | 2 | 2026-03-20 |Per-Directory Index ({dir}/index.md)
One row per entity file with summary and last-updated date. Owned by the skills that write to that directory. Scales independently — each directory can grow without affecting other indexes.
# Competitors Index
| File | Summary | Updated |
|------|---------|---------|
| [[competitors/widgetco]] | Enterprise SaaS competitor, Series C | 2026-03-20 |
| [[competitors/gizmotech]] | API-first competitor, growing fast | 2026-03-20 |Rules
- Skills read only their directory's index in preflight. competitor-intel reads
competitors/index.md; meeting-prep reads people/index.md. Cross-directory lookups go through the global index first, then the relevant directory index.
- Skills update only their directory's index on write. When a skill creates or
updates an entity file, update the row in that directory's index. Use the entity file's first # Heading as the summary if none exists.
- Global index is updated after directory index changes. Bump the entity count
and last-updated date for the affected directory.
- Created on first skill run if missing. Skills should not fail if an index
doesn't exist — create it with whatever entities are written in that run.
- Obsidian-compatible.
[[path/entity]]links (without.mdextension) work as
wiki links in Obsidian. Path is relative to ~/.nimble/memory/.
---
Chronological Wiki Log (log.md)
~/.nimble/memory/log.md is an append-only timestamped record of skill runs and findings. Grep-friendly format for answering "what did I learn this week?"
# Activity Log
## [2026-03-15] meeting-prep
- Updated: [[people/alex-kim]], [[companies/widgetco]]
- Key findings:
- Alex Kim moved to VP Engineering role
- Interested in API performance benchmarks
## [2026-03-18] company-deep-dive
- Created: [[companies/target-corp]]
- Key findings:
- Series B closed at $30M, Sep 2025
- Expanding into EU market Q2 2026
## [2026-03-20] competitor-intel
- Created: [[competitors/widgetco]], [[competitors/gizmotech]]
- Updated: [[competitors/acme-rival]]
- Key findings:
- WidgetCo launched enterprise tier pricing
- GizmoTech hired new CTO from CloudCorpRules
- Append at the end of the file (oldest first, newest last). Normal writes are
pure appends — no read-insert-rewrite needed. LLMs read the whole file; humans use grep "^## \[" log.md | tail -10 for recent entries.
- Format:
## [YYYY-MM-DD] skill-name— enablesgrep "^## \[" log.md | tail -10. - Content: List entities created/updated (as
[[path/entity]]links), then 2-3
bullet points of key findings. Keep entries concise — this is a log, not a report.
- Rotate entries older than 90 days as a separate maintenance step. After
appending the new entry, check the oldest entries (at the top). If older than 90 days, remove them. This rotation is not part of the normal append — it's a periodic cleanup that triggers during writes. The full reports in reports/ are the permanent record; log.md is for recent activity scanning.
- Created on first skill run if missing.
---
Cross-Entity References
Entity files use Obsidian-compatible [[path/entity]] wiki links to connect related entities across directories.
Format
# Alex Kim
## Current Role
VP of Engineering at [[competitors/widgetco]] (since 2024)
## Related
- Employer: [[competitors/widgetco]]
- Previous: [[companies/cloudcorp]]# WidgetCo
## Key People
- [[people/alex-kim]] — VP Engineering
- [[people/jane-smith]] — CEO
## Related Competitors
- [[competitors/gizmotech]] — overlapping market segmentRules
- Link format:
[[directory/entity-slug]]— no.mdextension, path relative
to ~/.nimble/memory/. Obsidian resolves these as wiki links.
- Add cross-references when relationships are discovered. When a skill finds
that a person works at a tracked company, or two competitors share a market segment, add links in both directions.
- Handle missing targets gracefully. A cross-reference to a file that doesn't
exist yet is fine — it becomes a valid link once that entity is created. Skills should not fail on dangling links.
- Skills follow links to enrich output. When meeting-prep finds
[[competitors/widgetco]] in a person's file, it loads that competitor file for additional context. When competitor-intel finds [[people/alex-kim]] in a competitor file, it can surface that relationship in the briefing.
When to Add Cross-References
| Relationship discovered | Link from | Link to |
|---|---|---|
| Person works at company | people/{name} → competitors/{company} or companies/{company} | Reverse link too |
| Companies compete | competitors/{a} → competitors/{b} | Reverse link too |
| Person previously at company | people/{name} → companies/{company} | — (one-way is fine) |
| Synthesis cites entity | synthesis/{topic} → entity files | — (one-way) |
---
Ad-Hoc Insights
When a user signals "save this", "remember that", "note this down", or similar intent during a conversation, file the insight into the relevant entity file(s) instead of letting it vanish into chat history.
Filing Pattern
1. Identify the relevant entity file(s). If the insight is about a competitor, file it in competitors/{name}.md. If it spans multiple entities (e.g., "WidgetCo is partnering with GizmoTech"), update all relevant files. 2. Append under a dated `## Insights` section:
## Insights
### 2026-03-22
- User noted: WidgetCo's enterprise pricing is 2x ours — [[competitors/gizmotech]]
is closer to our price point [ad-hoc]3. Add cross-references if the insight connects entities (as shown above). 4. Update the directory's `index.md` — bump the last-updated date for the affected file(s), and update the global index.md counts. 5. Append to `log.md` (at the end of the file):
## [2026-03-22] ad-hoc-insight
- Updated: [[competitors/widgetco]], [[competitors/gizmotech]]
- Key findings:
- WidgetCo enterprise pricing is 2x user's, GizmoTech closer to parityRules
- Tag with `[ad-hoc]` so skills can distinguish user-contributed insights from
skill-generated findings during dedup.
- Multi-entity insights update all relevant files with cross-references between
them.
- Don't create entity files for throwaway comments. If the user says "remember
that meetings on Fridays are bad", that's a preference (update business-profile.json), not an entity insight.
---
Cross-Entity Synthesis Pages
~/.nimble/memory/synthesis/ contains pages that analyze patterns across multiple entity files. Unlike entity files (which accumulate facts about one entity), synthesis pages draw conclusions across the knowledge base.
Page Creation
Synthesis pages are created dynamically when patterns emerge across entities — not from a pre-defined list. Common examples:
| Page | Purpose | Typical trigger |
|---|---|---|
competitive-landscape.md | Market positioning, feature gaps, pricing comparison | competitor-intel after 3+ competitors |
pricing-trends.md | Pricing pattern analysis across competitors | Pricing signals recur across 3+ competitor runs |
Page names should be slug-formatted topic labels (not skill names). Skills create synthesis pages when a pattern recurs across 3+ entities — this keeps synthesis data-driven rather than speculative.
Format
Synthesis pages use YAML frontmatter to track which entity files they were built from and when. This makes staleness deterministic — compare current file timestamps against the recorded ones.
---
confidence: high
sources:
- path: competitors/widgetco.md
updated: 2026-03-20
- path: competitors/gizmotech.md
updated: 2026-03-20
- path: competitors/acme-rival.md
updated: 2026-03-18
generated_by: competitor-intel
generated_at: 2026-03-20
---
# Competitive Landscape
## Market Map
[Positioning of each competitor by segment, size, strategy]
## Feature Comparison
| Capability | Us | [[competitors/widgetco]] | [[competitors/gizmotech]] |
|---|---|---|---|
| Real-time data | ✅ | ❌ | Partial |
## Pricing Comparison
[Tier-by-tier comparison where known]
## Key Patterns
- Trend 1 with evidence from multiple competitors
- Trend 2 with cross-entity citations
## What This Means
[Strategic implications — what the patterns suggest for the user's company]Rules
- Cite source entity files with
[[path/entity]]links. Every claim must trace
back to an entity file.
- Track sources and confidence in frontmatter.
confidence: high|medium|low
reflects data completeness (high = all key sources available, low = sparse data). The sources: block lists every entity file used and its last-modified date at generation time. To check staleness, compare current file timestamps against the recorded ones — if any source was updated since generation, the page is stale.
- Refresh when sources are stale. If a skill adds major new signals to 2+ source
entities since the synthesis was generated, regenerate. Don't regenerate on every run — only when the source timestamps diverge.
- Use `nimble-analyst` agent for synthesis generation. The analyst has the right
model (Sonnet) for cross-entity pattern recognition and strategic analysis.
Generation Trigger
competitor-intel generates competitive-landscape.md when:
- 3+ competitors have been researched in the current run, OR
- The existing synthesis page's source timestamps are stale (source entities were
updated since generation)
Other synthesis pages are created by the relevant skills when patterns emerge, or on user request.
---
Research Backlog (backlog.md)
~/.nimble/memory/backlog.md tracks knowledge gaps and research questions — things to investigate in future skill runs. This is not synthesis (derived, read-only output) — it's imperative (drives future action).
# Research Backlog
## Open
- [ ] WidgetCo pricing for enterprise tier — couldn't find public pricing [2026-03-20, competitor-intel]
- [ ] GizmoTech Series B details — rumored but unconfirmed [2026-03-20, competitor-intel]
- [ ] Alex Kim's LinkedIn activity — profile was private [2026-03-15, meeting-prep]
## Resolved
- [x] WidgetCo new CTO name — confirmed: Sarah Chen [2026-03-22, competitor-intel]Rules
- Any skill can append questions to the
## Opensection when it encounters
gaps during research. Tag each with date and skill name.
- Users can add questions via ad-hoc insights ("find out about X next time").
- Skills check backlog before running to avoid re-researching resolved questions
and to prioritize open ones relevant to the current run.
- Resolved questions get moved to
## Resolvedwith a resolution date — not
deleted. This preserves the audit trail.
---
Deep Storage Formats
competitors/
One file per competitor. Append new findings under dated headers — never overwrite.
# WidgetCo
## Key Facts
- Domain: widgetco.com
- HQ: San Francisco
- Funding: Series C ($45M, Jan 2026)
- CEO: Jane Smith
## Signals
### 2026-03-20
- Launched new enterprise tier pricing — [source URL]
- Hired VP of Sales from CRMHub — [source URL]
### 2026-03-13
- Announced partnership with AWS — [source URL]people/
One file per contact. Used by meeting-prep skill.
# Alex Kim
## Current Role
VP of Engineering at WidgetCo (since 2024)
## Background
- Previously: Senior Director at CloudCorp (2019-2024)
- Education: MS Computer Science, top-10 program
## Notes from Previous Meetings
### 2026-03-15
- Interested in our API performance benchmarks
- Prefers technical depth over high-level summariescompanies/
Detailed company profiles from deep-dive research.
# Target Corp
## Overview
- Industry: Enterprise SaaS | Founded: 2015 | HQ: Austin, TX | ~500 employees
## Financials
- Last funding: Series B ($30M, Sep 2025) | Revenue: Est. $40M ARR
## Recent News
(dated entries, same format as competitors/)reports/
Timestamped full skill outputs. Save the complete briefing, not a summary.
Naming: {skill-name}-{YYYY-MM-DD}.md — if a skill may produce multiple reports per day (e.g., meeting-prep for different companies), add a qualifier: {skill-name}-{qualifier}-{YYYY-MM-DD}.md. The qualifier is defined in each skill's SKILL.md (e.g., company slug for meeting-prep).
glossary.md
Industry terms and jargon. Updated when the user uses unfamiliar terms.
Bootstrapping (First Run)
mkdir -p ~/.nimble/memory/{competitors,people,companies,reports,positioning,synthesis}Create stub files for each competitor from the onboarding flow.
index.md and log.md are created automatically on the first skill run that writes to memory — no need to create empty stubs during bootstrapping.
Differential Analysis
The key feature across all skills — only surface what's genuinely new.
Dedup Lifecycle
Memory loading happens at two points in every skill:
1. Step 0 (Preflight): Load relevant memory files for context. This tells the skill what's already known so it can pass known signals to sub-agents for dedup during research. For example, competitor-intel loads ~/.nimble/memory/competitors/*.md; meeting-prep loads ~/.nimble/memory/people/*.md.
2. Analysis step (before report generation): Final dedup check. Compare all findings from research against loaded memory. Only signals classified as NEW or UPDATED (per the freshness classification in nimble-playbook.md) make it into the report.
What "new" means
- "WidgetCo raised a Series C" is noise if already in memory
- "WidgetCo just hired a new CTO" is a new signal worth highlighting
- "WidgetCo raised a Series C" with a new detail (amount, lead investor) is an UPDATE
Learning from Corrections
When the user corrects the skill, update both tiers:
| Correction | Profile update | Deep storage update |
|---|---|---|
| "Skip CompanyX" | preferences.skip_competitors | Archive file |
| "Track CompanyY" | competitors list | Create stub file |
| "That info is wrong" | — | Update the file |
| "ARR means Annual Recurring Revenue" | — | Add to glossary.md |
| "I prefer bullet points" | preferences.output_format | — |
Always confirm the update to the user.
Checkpointing & Resume
For multi-phase pipelines (map → extract → enrich → score), save intermediate results so failed or interrupted runs can resume without re-doing completed work.
Storage
~/.nimble/memory/{skill-name}/checkpoints/{slug}/
├── map.json # Phase 1 output
├── extract.json # Phase 2 output
└── enrich.json # Phase 3 output{slug} is a stable identifier derived from the run's input parameters (e.g., URL domain, search query hash). Same input = same slug = resumable.
Checkpoint format
Each phase file is JSON:
{
"phase": "extract",
"status": "complete",
"timestamp": "2026-04-03T15:30:00Z",
"record_count": 47,
"data": [ ... ]
}status is "complete" or "partial" (interrupted mid-phase).
Resume logic
On re-run with the same parameters: 1. Detect existing checkpoint directory for the slug 2. Offer: "Found previous run (47 records from Apr 3). Resume and fill gaps, or start fresh?" 3. If resume: skip phases where status = "complete", re-run where status = "partial" or file is missing 4. If start fresh: delete the checkpoint directory and begin from phase 1
Rules
- One checkpoint directory per unique run (keyed by slug)
- Clean up checkpoints older than 30 days on skill startup
- Don't checkpoint trivial runs (< 5 records) — the overhead isn't worth it
Rules
- Never touch user project files. All persistence under
~/.nimble/. - Append, don't overwrite. Deep storage grows over time with dated sections.
- Read on demand. Only load files when the skill actually needs them.
- Update profile after every run. At minimum,
last_runstimestamp. - Update wiki files after every memory write. Update the directory's
index.md
for affected entities, bump the global index.md counts, append a log.md entry for the run, and add cross-references where relationships are discovered.
- Handle missing gracefully. If a file doesn't exist, create it. This includes
index files, log.md, backlog.md, and cross-reference targets.
---
Source Links Enforcement
Every signal in every report must include a clickable source URL. This is a hard requirement — reports without source links are incomplete and must not be distributed.
What counts as a source link:
- A direct URL to the article, press release, or page where the signal was found
- The URL returned by
nimble searchin the result'surlfield - For extracted content, the URL passed to
nimble extract --url
What does NOT count:
- A company's homepage (unless the signal is specifically about homepage content)
- A generic domain without a path (e.g.,
https://widgetco.com) - "Source: web search" or any non-clickable attribution
If a signal has no source URL after research and extraction, drop it from the report. An unsourced signal is worse than a missing one — it can't be verified and erodes trust.
---
Report Distribution
After presenting output, offer sharing based on available MCP connectors.
Connector Detection
Check before presenting options:
- Notion:
mcp__plugin_Notion_notion__notion-create-pages - Slack: Any Slack MCP tool
Sharing Flow
Use AskUserQuestion with only the available options:
Share this report?
- Save to Notion — full report as a page
- Send to Slack — TL;DR to a channel
- Both
- Skip
Notion: Create a dated subpage. If integrations.notion.reports_page_id exists in the profile, use it as parent. Otherwise ask and save the ID for next time.
Slack: Post TL;DR only — Slack is for alerts, not full reports. If integrations.slack.channel exists, use it. Otherwise ask and save.
Neither available (first run only):
Tip: If you connect a Notion or Slack MCP server, I can save reports or post
TL;DRs to your team automatically.
Don't repeat this tip on subsequent runs.
Nimble Playbook
How to run Nimble CLI commands in Claude Code. Read this before executing any commands.
---
Claude Code Execution Rules
- No shell state persistence. Variables set in one Bash call are gone in the next.
Inline all values (dates, paths, names) directly into every command.
- No `&` + `wait` parallelism. It breaks in Claude Code. Instead, make **multiple
Bash tool calls in a single response** — they run in parallel natively.
- Search returns JSON —
--output-formatdoesn't change this. With `--search-depth
lite`, the JSON is small (title, description, URL per result). Parse it directly.
- Extract returns JSON with `data.markdown` — use
--format markdownto get clean
content in the data.markdown field.
Preflight Pattern
Transport selection (run once per session)
Skills work via two transports — CLI (preferred, full surface area) or MCP (fallback, curated tool set covering the same operations). Pick one at the start of every session and stick with it; don't re-probe on every command.
| Check | If it works | What to use |
|---|---|---|
nimble --version (>= 0.12.0) and NIMBLE_API_KEY is set | CLI is ready | Bash nimble ... commands |
| `claude mcp list 2>/dev/null \ | grep -q "nimble" (or first mcp__plugin_nimble_nimble__*` call succeeds) | Plugin MCP is connected |
mcp__plugin_nimble_nimble__* tools are listed, but a read-only nimble_agents_list probe returns an auth / not-connected error or an OAuth authorization URL | Plugin is installed but the connector isn't connected (typical Cowork / claude.ai state) | Stop — guide connector connection (below). Never invent an auth-completion flow. |
| None of the above | Stop — guide install (below) | — |
Connector not connected (Cowork / claude.ai) — verify BEFORE working
In Cowork / claude.ai the plugin is often installed while its connector is not yet connected, so live data calls fail. Confirming the connection is a required preflight step — not an error to react to mid-task. When mcp__plugin_nimble_nimble__* tools are listed but you haven't confirmed the connector is live, run one read-only probe before any real work:
- A single
nimble_agents_listcall is the cheapest confirmation. Success →
connected, proceed. Auth / not-connected error, or a response containing an OAuth authorization URL → not connected.
When not connected, surface this verbatim and stop — do not fall back to WebFetch, WebSearch, curl, or any other tool, and do not guess at data:
Your Nimble plugin is installed, but its connector isn't connected yet — that's
why I can't fetch live data. To connect it:
>
1. Open Customize → Connectors
2. Find Nimble and click Connect
3. Complete the login in your browser. No Nimble account? You can create one
right there during login.
4. Once it shows Connected, re-run your request and I'll continue.
If a tool hands back an OAuth "Authorize" URL
A not-connected tool call may return an authorization link (e.g. "Authorize Nimble MCP →") instead of data. Present that link to the user exactly as given, then stop and wait. Hard rules:
- Never invent a completion flow. There is no "paste the URL from your address
bar back to me" step, and you cannot "complete the connection" yourself. Claiming either is a hallucination.
- Never say the tools "will activate" and then call them in the same turn. Wait
for the user to confirm they've authorized, then retry.
- To check whether authorization succeeded, run one read-only
nimble_agents_list
probe — don't assume.
No plugin and no CLI
If neither path works at all (no plugin installed, no CLI installed), surface this hint verbatim and stop:
Nimble isn't installed. Pick the path for your environment:
>
Any Claude product (Claude Code, Claude Cowork, claude.ai) — recommended:
```
/plugin install nimble
```
Installs the Nimble plugin. The.mcp.jsoninside the plugin auto-registers as a Connector inCustomize → Connectors. First tool call triggers the OAuth flow — no API key needed.
>
Codex CLI or other terminal agents (shell access, no `/plugin`):
```
npm i -g @nimble-way/nimble-cli
```
Thenexport NIMBLE_API_KEY=<key>and re-run. Seereferences/profile-and-onboarding.mdfor the full install flow.
>
Cursor, VS Code, or any other MCP client:
Paste this into your MCP settings (.cursor/mcp.json or host equivalent):```json
{
"mcpServers": {
"nimble": { "type": "http", "url": "https://mcp.nimbleway.com/mcp" }
}
}
```
The plugin path (/plugin install nimble) is the easiest onboarding everywhere it works — one command, OAuth handles auth, no API key to manage. Use the CLI path only when shell access is available but /plugin install isn't (Codex, raw terminal agents). Use the manual mcp.json path only for MCP clients outside the Claude family.
Standard preflight (run in parallel after transport is selected)
Every skill kicks off with these simultaneous calls:
python3 -c "from datetime import datetime, timedelta; print((datetime.now() - timedelta(days=14)).strftime('%Y-%m-%d'))"(14 days ago)date +%Y-%m-%d(today)cat ~/.nimble/business-profile.json 2>/dev/null(profile — fall back to MCP filesystem tool if shell unavailable)cat ~/.nimble/memory/index.md 2>/dev/null(global wiki index — know what directories have data)
Don't skip the transport check — running CLI commands when only MCP is available (or vice versa) wastes a turn and confuses the user.
Request Attribution
All Nimble API calls must carry a client_source tag so usage can be tracked per skill. The value is always skill- followed by the exact SKILL.md name field (e.g. skill-competitor-intel, skill-seo-intel, skill-nimble-web-expert).
CLI path — add --client-source skill-{name} as the global flag on every nimble command. Place it immediately after nimble, before the subcommand. No shell state persistence means this must be inlined on every individual call:
nimble --client-source skill-{name} search --query "..."
nimble --client-source skill-{name} extract --url "..."
nimble --client-source skill-{name} agent run --agent <name> --params '{...}'
nimble --client-source skill-{name} map --url "..."
nimble --client-source skill-{name} crawl run --url "..."MCP path — per-skill client source tracking is not yet supported by the MCP server (it currently sends X-Client-Source: nimble_mcp_server for all calls regardless of skill). This will be enabled once the MCP server adds CLIENT_SOURCE support — no action needed here until then.
Sibling Handoff
When skills in the same family chain together (e.g., extract → enrich → verify), the second skill can skip redundant preflight work. Detect a sibling handoff by checking for same-day output from the upstream skill:
ls ~/.nimble/memory/reports/{upstream-skill}-*$(date +%Y-%m-%d).md 2>/dev/nullUse the dated report as the recency signal — data files under memory/{skill}/ may not have dates in their filenames, so always verify via the report timestamp. If a same-day report exists, parse the slug from the filename and load the corresponding data files.
If same-day sibling output exists:
- Skip CLI check and profile load — they were validated minutes ago
- Reuse WSA Layer 1 and Layer 3 inventory — the catalog hasn't changed. Only
re-run Layer 2 if the specialty or context changed.
- Use the sibling's structured output directly — if the upstream skill produced
data files with domains and page URLs, don't re-search for what's already known. Construct URLs from known patterns instead of running N web searches.
If no same-day sibling output exists: Run full preflight as normal.
This pattern is optional — skills MUST still work standalone without sibling output. The handoff is a fast path, not a requirement.
Smart Date Windowing
For any skill using --start-date based on previous runs:
- First run: 14 days ago → full mode
- Last run < 3 days ago: use 7 days ago (too narrow = empty results) → quick refresh
- Last run 3-14 days ago: use the last run date → quick refresh
- Last run > 14 days ago: 14 days ago → full mode
- Same-day repeat: if
last_runs.{skill-name}is today, check if a report already
exists at ~/.nimble/memory/reports/{skill-name}*[today].md. If it does, ask the user before re-running: "Already ran today. Run again for fresh data?" Don't silently re-run — it wastes API credits and produces near-identical output. Exception — meeting-prep: Skip the same-day report check. Meeting-prep is per-meeting, not per-day — users may prep for multiple meetings in a single day. Instead, meeting-prep checks freshness at the entity level: load cached profiles from ~/.nimble/memory/people/ and ~/.nimble/memory/companies/ and offer to reuse recent research rather than blocking the run.
---
Search
# Standard search (always use --search-depth lite for discovery)
nimble search --query "company name news" --max-results 10 --search-depth lite
# News-focused search
nimble search --query "company name" --focus news --max-results 10 --search-depth lite
# Date-filtered search (inline the date — don't use variables)
nimble search --query "company funding" --focus news --start-date "2026-03-11" --max-results 10 --search-depth lite
# Social signals from X/LinkedIn
nimble search --query "Company" --include-domain '["x.com", "linkedin.com"]' --max-results 10 --search-depth lite --time-range week
# Deep search (full page content — only for comprehensive analysis, costs more)
nimble search --query "company name" --search-depth deep --max-results 5
# Fast search (premium tier — not used by default)
# nimble search --query "company name" --search-depth fast --max-results 10Key flags:
--query— search query string (required)--focus—general,news,shopping,social,coding,academic.
`social` searches social platform people indices directly (LinkedIn, X) — best for finding specific people. If it errors, use --include-domain '["linkedin.com"]' as an alternative approach.
--max-results— max results to return--start-date/--end-date— date filters (YYYY-MM-DD)--search-depth—lite(1 credit),deep(1 + 1/page)--include-domain— JSON array of domains, e.g.,'["x.com", "linkedin.com"]'--time-range— e.g.,week--country— geo-targeted results (e.g., "US", "IL")--include-answer— LLM-powered answer summary
Date range strategy:
- First run: 14 days ago
- Subsequent runs:
last_runstimestamp from business profile - If < 3 results: retry without
--start-date
Extract
# Extract article content as markdown (default for content analysis)
nimble extract --url "https://example.com/article" --format markdown
# Extract raw HTML (required for <head> metadata: canonical, schema, og, meta tags)
nimble extract --url "https://example.com" --format html
# Extract with JavaScript rendering (for dynamic/SPA pages)
nimble extract --url "https://example.com/spa" --render --format markdownResponse is JSON. The field returned depends on --format:
--format markdown→data.markdown(clean body content)--format html→data.html(raw HTML including<head>)--format plain_text→data.plain_text--format simplified_html→data.simplified_html
Format selection by use case:
| Need | Format | Why |
|---|---|---|
| Article body content, word count, headings | markdown | Clean text, no nav/footer noise |
| Meta tags (title, description, canonical, og, twitter) | html | Markdown strips <head> |
| Schema markup (JSON-LD) | html | Script tags not in markdown |
hreflang, <html lang> | html | Attributes not in markdown |
| Structured field extraction | --parse --parser '{...}' | LLM extracts specific fields |
| Both body and head | markdown + html | Two calls or parse html for both |
Key flags:
--url— target URL (required)--format—markdown,html,simplified_html,plain_text(pick based on table above)--render— render JavaScript using a browser--parse --parser '{...}'— structured extraction via LLM parser schema
Extraction fallback (if data.markdown is mostly JavaScript/boilerplate): 1. Garbage check: If data.markdown has < 100 characters of meaningful content (after stripping nav/footer boilerplate), treat it as garbage. 2. Retry with --render --format markdown (handles JS-heavy/SPA pages) 3. If still garbage: search for the same article title on a different domain 4. If still nothing: skip and log — never abort a batch for a single extraction failure
Extract async & batch
# Async — submit single URL, get task_id, poll for results
nimble extract-async --url "https://example.com/page" --render --format markdown
# Batch — up to 1,000 URLs in one request
nimble extract-batch \
--shared-inputs 'render: true' --shared-inputs 'format: markdown' \
--input '{"url": "https://example.com/page-1"}' \
--input '{"url": "https://example.com/page-2"}'Poll async tasks with nimble tasks get --task-id <id> and fetch results with nimble tasks results --task-id <id>. Poll batches with nimble batches progress --batch-id <id>.
Map & Site Mapping
nimble map --url "https://example.com/blog" --limit 20Site Mapping Pattern
Use nimble map to discover a site's page structure, then score and filter pages by relevance before extracting.
1. Discover: nimble map --url {url} --limit {cap} — returns a list of URLs 2. Score: Each skill defines a keyword/weight table for URL path segments (e.g., /providers = High, /about = Medium, /blog = Low). Score each discovered page against the table. 3. Filter: Keep pages scoring above the skill's threshold. Always include the homepage as a fallback. 4. Fallback: If nimble map returns < 3 candidates, use nimble search --query "site:{domain} {keywords}" --max-results 10 --search-depth lite
Each skill provides its own keyword/weight table in SKILL.md — the pattern here is the discover → score → filter → fallback flow.
Agents
Pre-built extraction templates for structured data from specific sites (Amazon, LinkedIn, Google, etc.). Use when you need structured fields rather than raw page content.
# List available agents (search by domain or vertical)
nimble agent list --limit 100
nimble agent list --limit 100 --search "amazon"
# Inspect an agent's schema (input params + output fields)
nimble agent get --template-name <agent_name>
# Run an agent (sync — waits for result)
nimble agent run --agent <agent_name> --params '{"key": "value"}'
# Run an agent (async — returns task_id, poll for results)
nimble agent run-async --agent <agent_name> --params '{"key": "value"}' \
--callback-url "https://your.server/callback"Key flags for `run` / `run-async`:
--agent— agent name fromnimble agent list(required)--params— JSON object with agent input parameters (required)--localization— enable zip_code/store_id localization (agent-dependent)
Additional flags for `run-async`:
--callback-url— POST callback when task completes--storage-type—s3orgs--storage-url— destination bucket URL--storage-compress— gzip the stored output--storage-object-name— custom filename instead of task_id
Response: data.parsing contains the structured output. Shape depends on agent type:
- PDP (product/profile/detail) → flat dict
- SERP / list → array of objects
- Google Search →
{"entities": {"OrganicResult": [...], ...}}
Async task states: pending → success or error. Poll with nimble tasks results --task-id <task_id>.
Fallback rule: If no agent exists for the target domain, fall back to nimble search + nimble extract. Don't fail silently — log which domains lacked agent coverage so agent-builder can fill gaps later.
Agent batch
# Up to 1,000 agent requests in one call
nimble agent run-batch \
--shared-inputs 'agent: amazon_serp' \
--input '{"params": {"keyword": "iphone 15"}}' \
--input '{"params": {"keyword": "iphone 16"}}'Returns a batch_id. Poll with nimble batches progress --batch-id <id>, then fetch individual results with nimble tasks results --task-id <id>.
Tasks & batches polling
# Single async task
nimble tasks get --task-id <task_id> # check status
nimble tasks results --task-id <task_id> # fetch results
# Batch
nimble batches progress --batch-id <batch_id> # lightweight progress check
nimble batches get --batch-id <batch_id> # all task IDs + states
nimble batches list --limit 20 # list all batches
nimble tasks list --limit 20 # list all tasksWorkflow: Always nimble agent get before nimble agent run to understand the expected input params and output fields.
Agent Creation (generate → poll → iterate → publish)
Create custom extraction agents for any website. The full lifecycle is available via CLI.
# Generate a new agent
nimble agent generate \
--agent-name niche_store_pdp \
--prompt "Extract product name, price, rating, and first 5 reviews" \
--url "https://example.com/products/widget-pro"
# Refine an existing agent (clone + apply new prompt)
nimble agent generate \
--agent-name niche_store_pdp \
--from-agent niche_store_pdp \
--prompt "Add a discount_percentage field"
# Poll generation status (async — typically 1-3 min)
nimble agent get-generation --generation-id <generation_id>
# Publish when satisfied
nimble agent publish --agent-name niche_store_pdp --version-id <version_id>Key flags for `generate`:
--agent-name— name for the agent (required)--prompt— natural language description of what to extract (required)--url— sample URL to analyze (required)--from-agent— existing agent to clone and refine (for iteration)--input-schema— custom input schema (optional, inferred if omitted)--output-schema— custom output schema (optional, inferred if omitted)--metadata— additional metadata (optional)
Generation response: returns id (generation ID), status (queued → in_progress → success / failed), and generated_version_id on success.
Workflow: Generate → poll with get-generation until success → optionally iterate with --from-agent → publish with version-id.
Polling: Generation takes 1-3 minutes. Run the generate → poll → publish loop as a background Task agent so the user isn't blocked waiting. The Task agent should poll nimble agent get-generation every 10 seconds until status is success or failed, then publish automatically (or report failure). Present results to the user when done.
MCP Fallback (when CLI is not installed)
If nimble --version returns "command not found", fall back to the Nimble MCP server. All CLI commands have MCP equivalents — discover them via the MCP tool list. MCP tools accept the same parameters as CLI flags, passed as tool arguments instead of flags.
Parallel Execution
Make multiple Bash tool calls in a single response. Claude Code runs them in parallel automatically:
- Call 1:
nimble search --query "CompanyA news" --max-results 5 --search-depth lite - Call 2:
nimble search --query "CompanyB news" --max-results 5 --search-depth lite - Call 3:
nimble search --query "CompanyC news" --max-results 5 --search-depth lite
Sub-Agent Spawning
When using the Agent tool for parallel research:
- Always `mode: "bypassPermissions"` — sub-agents don't inherit Bash permissions.
- Batch max 4 agents. More risk hitting rate limits. For 5+, batch in groups.
- Tell agents to use Bash — explicitly say "Use the Bash tool to execute nimble
commands." Some agents try WebSearch instead.
- Fallback on failure — if any agent returns without results, run those searches
directly from the main context. Don't leave gaps.
Communication Style
Inform the user at phase transitions only with concrete numbers:
- "Researching Acme Corp + 5 competitors since Mar 12..."
- "Found 12 new signals. Pulling top 4 articles..."
- "All data collected. Building your briefing..."
Don't narrate individual tool calls.
Rate Limits & Common Errors
- Rate limit: 10 req/sec per API key
- Retry on 429: Reduce simultaneous calls
- Timeout: 30 seconds per request
| Error | Cause | Fix |
|---|---|---|
NIMBLE_API_KEY not set | Missing API key | See profile-and-onboarding.md |
401 Unauthorized | Expired key | Regenerate at app.nimbleway.com |
429 Too Many Requests | Rate limit | Fewer simultaneous calls |
timeout | Slow response | Retry once, then skip |
500 Server Error | Transient server failure | Retry once without --focus; if persistent, simplify query |
empty results | No matches | Remove --start-date, broaden query |
Signal Date Validation
High-quality intelligence requires distinguishing between when a page was published and when the underlying event occurred. This matters because:
- Syndicated or republished content may carry a different publication date than the
original source
- Secondary coverage (regulatory filings, recap articles, industry roundups) can
report on events that happened weeks or months earlier
Article Date vs Event Date
Every signal has two dates:
| What it is | |
|---|---|
| Article date | When the page was published |
| Event date | When the underlying event actually happened |
A signal is "new" only if its event date falls within the freshness window.
Event Date Extraction Rules
Sub-agents must determine the event date from content:
1. Explicit past reference — "launched in Q3", "appointed last October" → event date is in the past, regardless of the article date 2. Temporal language — "last quarter", "months ago", "earlier this year" → resolve relative to the article date 3. Present tense announcement — "today announces", "is launching" → event date ≈ article date 4. Dateline — "NEW YORK, March 15 —" → event date = that dateline date 5. If ambiguous — extract the source URL and check the on-page date
Source Type Hierarchy
When the same event appears from multiple sources, prefer those closest to the event:
1. Primary — the company's own domain, official press release, regulatory filing 2. Wire service — AP, Reuters, Bloomberg 3. Major outlet — original reporting with bylines 4. Derivative — syndicated copies, aggregator sites, recap articles, or content that attributes its information to another source
If the only source for a signal is derivative, corroborate against a primary or major source before reporting.
Freshness Classification
After determining the event date, classify each signal:
| Classification | Meaning | Action |
|---|---|---|
| NEW | Event date within freshness window, not in memory | Include in report |
| UPDATED | Known event with genuinely new information | Include as update |
| STALE | Old event covered by a recent article | DROP — do not include |
| UNCERTAIN | Can't determine event date from snippet alone | Extract URL to verify; if still uncertain after extraction, DROP |
Hard rule: Only signals classified as NEW or UPDATED may appear in reports. STALE and UNCERTAIN signals must be dropped entirely — not downgraded, not footnoted, not included as "background context." If a signal can't be verified as genuinely recent, it doesn't exist as far as the report is concerned.
--start-date Best Practices
--start-date is a useful filter for reducing noise, but always validate event dates from the content itself:
- For news queries (
--focus news), consider running a parallel undated query to
surface original sources alongside recent coverage
- The existing fallback ("If < 3 results, retry without
--start-date") remains useful
Verification Budget
Not every signal needs full verification — budget extract calls by priority:
| Priority | Examples | Verification |
|---|---|---|
| P1 (high impact) | Funding, M&A, leadership changes | Always extract + corroborate (see below) |
| P2 (medium impact) | Product launches, partnerships, major hires | Extract if date is UNCERTAIN or source is derivative |
| P3 (low impact) | Blog posts, minor hires, event appearances | Trust if date looks plausible; drop if obviously stale |
Skills define their own P1/P2/P3 signal types in their SKILL.md. The verification budget above applies universally regardless of which signals a skill classifies at each level.
P1 Corroboration (Mandatory)
Any P1 signal sourced from derivative or aggregator sites must be corroborated before it can appear in a report. This is a hard gate, not a suggestion.
For each P1 signal that needs corroboration:
nimble search --query "[Company] [event summary]" --max-results 5 --search-depth liteLook for the primary source (company blog, press release, official filing, regulatory document). Check the primary source's date:
- Primary source dates the event within the freshness window → signal is NEW, include it
- Primary source dates the event outside the freshness window → reclassify as STALE, drop
- No primary source found → reclassify as UNCERTAIN, drop
Do not report P1 signals that fail corroboration. It's better to miss a real signal than to report a stale one as new — trust is the product.
---
Entity Deduplication
When a skill collects entity records from multiple sources (directories, search results, extracted pages), deduplicate before reporting. This is distinct from signal-level differential analysis (see memory-and-distribution.md) — entity dedup merges records for the same entity across sources within a single run.
Three-layer pattern (generic — each skill customizes the specifics):
1. Exact ID match — If the entity type has a canonical ID (place_id, NPI number, domain), match on that first. Exact match = same entity, merge fields. 2. Domain normalization — Strip www., trailing slashes, protocol. Compare root domains. www.acme.com/ and acme.com are the same entity. 3. Fuzzy name + location — Normalize names before comparing:
- Lowercase all characters
- Strip titles and honorifics (
Dr.,Mr.,Ms., etc.) - Strip credential suffixes (
MD,DDS,Inc,LLC,Corp, etc.) - Strip common noise words (
The,and,of,&) - Collapse whitespace and punctuation
- Compare normalized names with location context if available
This catches cross-source variations like "Dr. Jane Smith, MD" (Maps) vs "Jane Smith" (Yelp) vs "Smith Eye Care LLC" (BBB). Each source formats names differently — always normalize before comparing.
Track source_count per entity — entities confirmed by multiple sources are higher quality. Each skill defines which layers apply and any domain-specific matching rules in its reference files.
---
Entity Confidence Scoring
Rate each entity's data completeness so users know what to trust.
Generic formula — each skill defines its own target field list (N fields):
- High — All target fields found + confirmed by 2+ sources (
source_count >= 2) - Medium — >50% of target fields found
- Low — ≤50% of target fields found
Display the confidence level in output (e.g., ⬤⬤⬤ High, ⬤⬤○ Medium, ⬤○○ Low). Each skill defines its field list and may add criteria (e.g., requiring a verified phone number for High in a provider directory skill).
---
Input Parsing Pattern
Skills that accept batch input (lists of URLs, companies, locations) should detect the input type automatically:
| Input signature | Type | Action |
|---|---|---|
Contains docs.google.com/spreadsheets | Google Sheet URL | Read sheet directly |
Path ends in .csv and file exists | CSV file | Read and parse as CSV |
| Contains multiple URLs (one per line or comma-separated) | Inline URL list | Parse directly |
| Otherwise | Unknown | Ask user for input |
Normalize all inputs to a uniform list of records before batch processing. Don't assume a specific format — detect and adapt.
---
Scaled Execution
When a skill needs to run multiple WSA or API calls, choose the execution tier based on the estimated number of requests. Each skill calculates its own estimate from input size and operations per record.
| Estimated calls | Strategy | How |
|---|---|---|
| 1–10 | Individual calls | Parallel Bash calls (max 4 concurrent) |
| 11–100 | Single batch | extract-batch or agent run-batch — one API call, server-side parallelism, poll for results |
| 100–1,000 | Multiple batches | Split into batches of up to 1,000. Use sub-agents to prepare inputs and process results |
| >1,000 | Confirmation gate + batches | Show estimate, ask user to confirm before proceeding, then execute via batches |
Individual calls (1–10)
Run up to 4 concurrent Bash calls per the Parallel Execution rules above.
Batch calls (11+)
For page extraction (11+ URLs):
nimble extract-batch \
--shared-inputs 'format: markdown' \
--input '{"url": "https://example.com/page-1"}' \
--input '{"url": "https://example.com/page-2"}'Add --shared-inputs 'render: true' if pages need JavaScript rendering.
For WSA/agent calls (11+ entities):
nimble agent run-batch \
--shared-inputs 'agent: {agent_name}' \
--input '{"params": {...}}' \
--input '{"params": {...}}'Both return a batch_id. Poll progress:
nimble batches progress --batch-id {batch_id}Fetch results when complete:
nimble batches get --batch-id {batch_id}
nimble tasks results --task-id {task_id}Batch API handles up to 1,000 requests per call with server-side orchestration. For >1,000 requests, split into multiple batch calls.
Sub-agents should also batch. When spawning sub-agents for parallel work, tell each agent to use extract-batch or agent run-batch for its assigned items rather than making individual calls. One batch call per agent is faster and more reliable than 5-6 sequential calls.
Large job confirmation (>1,000)
Before executing, show the estimate and ask the user to confirm:
Estimated API calls: ~2,400 (120 locations × 3 WSAs per location × ~7 enrichment)
This is a large job. Proceed? [Y/n]Pattern: estimate → display → gate → execute
Why batch over individual calls
Individual nimble agent run calls each require a separate HTTP round-trip and Bash tool invocation. At scale (dozens+) this is slow, unreliable, and wasteful on a local machine. Batch APIs move orchestration server-side — one API call triggers all requests, and you poll for results. Always prefer batch when above the individual threshold.
---
Query Construction Tips
- Be specific: "Acme Corp product launch 2026" > "Acme Corp"
- Use `--include-domain '["domain"]'` for companies with generic names
- Fallback on empty: If < 3 results, retry without
--start-date - Combine focus modes: news + general in parallel for broader coverage
- Try variations: "CompanyName" → "Company Name" → domain
Profile & Onboarding
The business profile at ~/.nimble/business-profile.json and first-run setup flow.
---
Profile Schema
{
"company": {
"name": "Acme Corp",
"domain": "acme.com",
"description": "Enterprise SaaS platform for project management"
},
"industry_keywords": ["project management software", "team collaboration SaaS"],
"competitors": [
{ "name": "WidgetCo", "domain": "widgetco.com", "category": "project-mgmt" },
{ "name": "GizmoTech", "domain": "gizmotech.io", "category": "project-mgmt" }
],
"preferences": {
"skip_competitors": [],
"output_format": "bullet-points"
},
"integrations": {
"notion": { "reports_page_id": "" },
"slack": { "channel": "" }
},
"sales_context": {
"key_differentiators": [
"Only platform with real-time web data access",
"Sub-second API response times"
],
"integration_partners": [
{ "name": "DataStack", "type": "data warehouse" },
{ "name": "CRMHub", "type": "CRM" }
],
"case_studies": [
{ "customer": "Large enterprise retailer", "industry": "retail", "outcome": "3x faster competitive intel" }
],
"common_objections": [
{ "objection": "We already use [competitor]", "response": "Our real-time data is fresher — most competitors cache for 24h+" }
]
},
"last_runs": {
"competitor-intel": "2026-03-20T14:30:00Z",
"meeting-prep": "2026-03-22T09:00:00Z"
},
"setup_completed": true
}Reading the Profile
At the start of every skill run:
cat ~/.nimble/business-profile.json 2>/dev/nullIf missing or empty → trigger onboarding (see below).
Key fields:
company.name/company.domain— the user's companycompetitors— tracked competitors with domains and categoriesindustry_keywords— for industry-level searchespreferences.skip_competitors— competitors to excludelast_runs.{skill-name}— timestamp for time-aware searchessales_context— value positioning data (differentiators, integrations, case studies, objections)integrations— Notion/Slack config for report distribution
Updating the Profile
After every skill run — update last_runs:
import json, datetime, os
path = os.path.expanduser("~/.nimble/business-profile.json")
with open(path, "r") as f:
profile = json.load(f)
profile["last_runs"]["skill-name"] = datetime.datetime.now(datetime.timezone.utc).isoformat()
with open(path, "w") as f:
json.dump(profile, f, indent=2)On user correction — apply immediately:
| User says | Action |
|---|---|
| "Don't include CompanyX" | Add to preferences.skip_competitors |
| "Also track CompanyY" | Add to competitors (with domain + category) |
| "I moved to NewCompany" | Update company |
| "Show me more detail" | Update preferences.output_format |
Always confirm: "Got it — removed CompanyX from tracking."
Rules:
- Never overwrite the whole file. Read → modify → write.
- Preserve unknown fields.
- Handle missing file gracefully → trigger onboarding.
- JSON only, always valid.
---
First-Run Onboarding
Prerequisite Checks
The transport selection in nimble-playbook.md determines whether CLI or MCP is active. This section covers the install/upgrade/auth flow when neither is ready.
Minimum CLI version: 0.12.0
Preferred path — any Claude product (Claude Code, Claude Cowork, claude.ai)
The plugin install is one command and handles MCP registration + OAuth automatically:
"Run /plugin install nimble to install the Nimble plugin. The plugin's MCPserver auto-registers as a Connector you can see in Customize → Connectors.On first use, the OAuth flow runs in your browser — no API key needed."
This works in every Claude product (Code, Cowork, claude.ai) — they share the plugin + connector mechanism.
Plugin installed but connector not connected (Cowork / claude.ai)
The most common Cowork / claude.ai failure: the plugin is installed (mcp__plugin_nimble_nimble__* tools are listed) but its connector isn't connected, so live data calls fail. Check this before doing any work — don't fire a data call and react to the error. A single read-only nimble_agents_list probe confirms it: success = connected, proceed; auth/not-connected error or a response containing an OAuth authorization URL = not connected.
When not connected, tell the user verbatim and stop — never fall back to WebFetch, WebSearch, or any other tool, and never guess at data:
Your Nimble plugin is installed, but its connector isn't connected yet — that's
why live data isn't working. To connect it:
>
1. Open Customize → Connectors
2. Find Nimble and click Connect
3. Complete the login in your browser. No Nimble account? You can create one
right there during login.
4. Once it shows Connected, re-run your request.
If a tool returns an OAuth "Authorize" link instead of data, present the link as-is and stop. Do not invent a completion step ("paste the URL back", "I'll complete the connection") — no such step exists. Do not claim the tools will activate and then call them in the same turn. Wait for the user to authorize, then retry (or run one nimble_agents_list probe to confirm).
Codex CLI or other terminal agents (shell available, no /plugin install)
When /plugin install isn't available but the user has shell access, install the CLI directly — it exposes the full Nimble surface area:
1. Check if npm is available: npm --version 2. If npm exists:
"The Nimble CLI is required. I'll install it now."
>
Run: npm install -g @nimble-way/nimble-cli3. If npm is not available:
"The Nimble CLI requires Node.js/npm. Install Node.js first from
nodejs.org, then run: npm install -g @nimble-way/nimble-cli"4. After install, verify: nimble --version 5. If verification fails, stop and ask the user to check their PATH.
Cursor, VS Code, or other MCP clients outside the Claude family
When neither /plugin install nor shell access is workable, have the user paste this into their MCP settings (e.g., .cursor/mcp.json or the host's equivalent):
{
"mcpServers": {
"nimble": {
"type": "http",
"url": "https://mcp.nimbleway.com/mcp"
}
}
}After install, the first tool call triggers the OAuth flow automatically.
CLI outdated (version < 0.12.0)
Parse the version from nimble --version. If below 0.12.0:
"Your Nimble CLI is version [current] — version 0.12.0+ is required
for these skills. Upgrading now..."
>
Run: npm update -g @nimble-way/nimble-cliVerify after upgrade: nimble --version. If still outdated, suggest: npm uninstall -g @nimble-way/nimble-cli && npm install -g @nimble-way/nimble-cli
API key not set
You need a Nimble API key.
1. Go to app.nimbleway.com → API Keys
2. Generate a new key
3. Run: export NIMBLE_API_KEY=your_key_here4. Add to~/.zshrcor~/.bashrcto make permanent.
After the user sets it, verify: echo "NIMBLE_API_KEY=${NIMBLE_API_KEY:+set}"
API key expired (401)
Your key may have expired (72h TTL). Regenerate at app.nimbleway.com > API Keys.
All prerequisites met
Only proceed to Company Setup once CLI is installed, version is >= 0.12.0, and API key is set. Don't silently skip any check.
Company Setup (2 prompts max)
Prompt 1 — ask in plain text (NOT AskUserQuestion with options):
"What's your company's website domain? (e.g., acme.com)"
Verify — make two Bash calls simultaneously:
nimble search --query "[domain]" --include-domain '["[domain]"]' --max-results 3 --search-depth litenimble search --query "[domain] company" --max-results 5 --search-depth lite
Present what you found and confirm: "I found that [Company] ([domain]) is [brief description]. Is this the right company?"
Prompt 2 — skill-specific setup:
- competitor-intel: Offer choice via
AskUserQuestion: - Find for me — search and suggest competitors
- I'll list them — user provides names
If "Find for me", make three Bash calls simultaneously:
nimble search --query "[Company] competitors" --max-results 10 --search-depth litenimble search --query "[Company] vs" --max-results 10 --search-depth litenimble search --query "[Company] alternatives" --max-results 5 --search-depth lite
- meeting-prep: No extra setup — context comes per-meeting
- company-deep-dive: No extra setup — target company comes per-request
Create Profile
mkdir -p ~/.nimble/memory/{competitors,people,companies,reports,positioning,synthesis}Write ~/.nimble/business-profile.json using the schema above.
When setting up competitors, infer or ask for each competitor's domain and category. Also infer industry keywords from the company description.
Profile Exists
Skip onboarding. Greet with context: "Running competitor intel for Acme Corp — tracking WidgetCo, GizmoTech."
---
Error Recovery
If any step fails: 1. Tell the user what went wrong in plain language 2. Provide the exact command to fix it 3. Offer to retry
Never silently skip setup steps.
Provider Extraction Patterns
Skill-specific patterns for extracting practitioner data from healthcare practice websites. For general extraction and site mapping rules, see nimble-playbook.md.
---
Page URL Scoring
After nimble map discovers a site's pages, score each URL by path keywords to identify provider-relevant pages. See Site Mapping Pattern in nimble-playbook.md for the generic discover/score/filter/fallback flow.
Keyword Weight Table
| Weight | Path Segments | Examples |
|---|---|---|
| High | /providers, /doctors, /physicians, /our-team, /staff, /our-providers, /our-doctors, /our-physicians, /surgeons, /specialists | /our-providers, /meet-our-doctors |
| High | /dr-*, /doctor-* (individual provider pages) | /dr-jane-smith, /doctor-john-doe |
| Medium | /team, /about, /people, /about-us, /meet-the-team, /faculty, /clinicians | /about-us/team, /our-people |
| Low | Homepage (/), /services, /locations | /, /services/cataract-surgery |
| Skip | /blog, /news, /careers, /jobs, /privacy, /terms, /patient-portal, /pay-bill, /faq, /testimonials, /reviews, /gallery, /media | /blog/eye-health-tips |
Scoring Rules
- Cap at 15 pages per practice site
- Always include at least one High-weight page (or homepage as fallback)
- Individual provider pages (
/dr-*) are high value but cap at 10 per site to
avoid over-extraction on large multi-provider practices
- If
nimble mapreturns < 3 scored candidates, fall back to:
nimble search --query "site:{domain} doctors OR providers OR team" --max-results 10 --search-depth lite---
Core Extraction Fields (5 Fields)
Every provider record targets these 5 fields. Confidence scoring uses this as the N-field list (see Entity Confidence Scoring in nimble-playbook.md).
| # | Field | Key | Detection Patterns |
|---|---|---|---|
| 1 | Full Name | name | Dr. prefix, <h2>/<h3> heading patterns, bold text near credential suffixes, structured bio sections |
| 2 | Credentials | credentials | Regex patterns (see below) found adjacent to names |
| 3 | Specialty | specialty | Keywords per vertical (see below), often near name or in bio paragraph |
| 4 | Contact / Scheduling | contact | Phone regex \(?\d{3}\)?[-.\s]?\d{3}[-.\s]?\d{4}, appointment URLs (/book, /schedule, /request-appointment), email addresses |
| 5 | Education / Training | education | "Residency", "Fellowship", "Medical School", "Board Certified", university names, graduation years |
Confidence Scoring (from shared pattern)
- High -- 5/5 fields found + confirmed by 2+ pages or sources
- Medium -- 3-4/5 fields found
- Low -- 1-2/5 fields found
Display as: High, Medium, Low
---
Credential Regex Patterns
Match these suffixes adjacent to provider names. Case-insensitive. Allow comma, space, or period separators between multiple credentials.
Medical Doctors
MD|M\.D\.|DO|D\.O\.Eye Care
OD|O\.D\.|FAAODental
DDS|D\.D\.S\.|DMD|D\.M\.D\.Advanced Practice
NP|PA|PA-C|ARNP|APRN|CNS|CRNA|DNP|D\.N\.P\.Therapy & Allied Health
PT|DPT|OT|OTR|SLP|CCC-SLP|RD|RDN|LCSW|LPC|PhD|Ph\.D\.|PsyD|Psy\.D\.Board Certifications (commonly listed)
FACS|FACP|FACC|FACOG|FAAP|FACEP|FAAOS|FASRS|FRCSCCombined Pattern
When scanning extracted markdown, look for names followed by credential clusters:
[Name],?\s*((?:MD|DO|OD|DDS|DMD|NP|PA|PA-C|ARNP|APRN|PT|DPT|PhD|FACS|FAAO|FACP|FACC|FACOG|FAAP|FAAOS|FASRS)[,.\s]*)+---
Specialty Keywords by Healthcare Vertical
Ophthalmology
ophthalmology, ophthalmologist, retina, retinal, cataract, glaucoma, cornea,
corneal, LASIK, refractive surgery, oculoplastics, neuro-ophthalmology,
pediatric ophthalmology, vitreoretinal, anterior segment, posterior segment,
strabismus, ocular oncology, uveitisDental
dentist, dentistry, general dentistry, cosmetic dentistry, orthodontics,
orthodontist, periodontics, periodontist, endodontics, endodontist,
oral surgery, oral surgeon, prosthodontics, prosthodontist, pediatric
dentistry, implants, dental implants, TMJ, sedation dentistryDermatology
dermatology, dermatologist, Mohs surgery, cosmetic dermatology, skin cancer,
melanoma, psoriasis, eczema, acne, rosacea, laser treatment, botox,
fillers, chemical peel, phototherapy, patch testingGeneral / Primary Care
family medicine, internal medicine, primary care, general practice,
preventive medicine, geriatrics, urgent care, walk-in clinic,
physical exam, wellness, annual checkupOrthopedics
orthopedics, orthopedic surgery, sports medicine, joint replacement,
spine surgery, hand surgery, foot and ankle, shoulder, knee,
arthroscopy, fracture care, physical therapy, rehabilitation---
Entity Deduplication (Skill-Specific)
Apply the shared 3-layer dedup pattern from nimble-playbook.md with these skill-specific rules:
1. Exact match -- Same name + same practice domain = same provider 2. Credential match -- Same name + same credentials + same city = likely same provider (even across different practice sites) 3. Fuzzy match -- Normalize names (strip "Dr.", middle initials, suffixes), compare with Levenshtein distance <= 2 + same specialty = possible match, flag for review rather than auto-merging
Cross-source name normalization
Different sources format provider names very differently. Exact string matching across sources produces near-zero matches. Always normalize before comparing:
| Source | Raw format | After normalization |
|---|---|---|
| Google Maps | "Dr. Jane A. Smith, MD - Retina Specialist" | "jane smith" |
| Yelp | "Jane Smith" | "jane smith" |
| BBB | "Smith Eye Care LLC" | "smith eye care" |
| Practice website | "Jane A. Smith, M.D., F.A.C.S." | "jane smith" |
Normalization steps (apply in order): 1. Strip titles: Dr., Mr., Ms., Prof. 2. Strip credentials: all patterns from the Credential Regex section above 3. Strip business suffixes: LLC, Inc, Corp, PC, PLLC, PA, Associates 4. Strip specialty descriptors: "- Retina Specialist", "- Ophthalmologist" 5. Strip middle initials (single letters with optional period) 6. Lowercase, collapse whitespace, strip remaining punctuation
After normalization, match with location context (same city or same zip code). For practice-level dedup (not provider-level), also try matching the practice name against provider last names ("Smith Eye Care" → likely matches "Dr. Smith").
Track source_count -- providers found across multiple sources are higher confidence than those from a single source.
WSA Discovery for Healthcare Providers Extract
How to find and evaluate WSAs for each phase of provider extraction. The WSA catalog evolves constantly — new agents get added for healthcare directories, review sites, and regulatory databases. This skill discovers relevant agents at runtime rather than relying on a static list.
For general WSA execution rules (invocation, parsing, batch, fallback), see nimble-playbook.md.
---
Discovery Strategy
Three search layers
Run these searches at the start of the skill (during or right after preflight) to build a session-specific WSA inventory. Run all searches simultaneously:
Layer 1 — Vertical search:
nimble agent list --limit 100 --search "healthcare"Returns all agents tagged with the Healthcare vertical (clinicaltrials.gov, FDA, and any newly added healthcare agents).
Layer 2 — Session-specific search: Search for terms derived from the user's input — their specialty, location, or specific domains they mentioned:
# If user said "ophthalmology in Austin":
nimble agent list --limit 50 --search "ophthalmology"
nimble agent list --limit 50 --search "eye"
# If user mentioned specific directories:
nimble agent list --limit 50 --search "zocdoc"
nimble agent list --limit 50 --search "healthgrades"
nimble agent list --limit 50 --search "vitals"Adapt search terms to whatever the user provided. Include the specialty, common directory names for that specialty, and any domains the user mentioned.
Layer 3 — General discovery tools: These WSAs are useful across verticals for practice discovery, reputation, and verification:
nimble agent list --limit 50 --search "google_maps"
nimble agent list --limit 50 --search "yelp"
nimble agent list --limit 50 --search "bbb"
nimble agent list --limit 50 --search "review"Evaluating discovered agents
For each discovered agent, read its description and entity_type to classify it into a phase:
| If the agent description mentions... | Assign to phase |
|---|---|
| Search, listings, directory, discovery, local businesses | Discovery — finding practice URLs |
| Reviews, ratings, patient feedback, reputation | Enrichment: Reputation |
| Clinical trials, FDA, regulatory, compliance, licensing | Enrichment: Regulatory |
| Profile, detail page, business info, contact | Enrichment: Practice details |
Validate each relevant agent's params before using it:
nimble agent get --template-name [agent_name]Skip agents that don't fit — not every healthcare-tagged agent is useful for provider extraction. An FDA drug label agent, for example, isn't relevant unless the user specifically asked about pharmaceuticals.
Healthcare discovery prioritization
Recommended approach for healthcare practice discovery:
- Google Maps WSA — Primary source. Rich structured data (name, address, rating,
reviews, phone, place_id, coordinates). Process these results first. Note: the place_url field links to Maps — resolve practice website URLs in a separate step (see SKILL.md Step 2b).
- Yelp WSA — Supplementary source. Run alongside Maps but don't block on it.
Filter results by specialty keywords before merging, as Yelp categories are broader than medical specialty searches.
- BBB WSA — Best suited for enrichment (accreditation lookup on known practices)
rather than discovery. For broader BBB-based discovery, use nimble search --query "[specialty] site:bbb.org" which returns multiple results.
Building the session WSA plan
After discovery, present what you found to the user inline:
"Found N relevant WSAs for this run: [list by phase]. Using these alongside
direct site extraction."
This transparency helps the user understand what data sources are available and lets them suggest additional search terms if something is missing.
---
Phase Mapping
Discovery phase (finding practice URLs)
When: User provided a specialty + location instead of URLs.
Useful agent types: Map search, directory search, local business listings.
Search terms to try: google_maps, yelp, bing_maps, bbb, plus any healthcare directory the user mentions.
How to use: Run discovered search/listing agents with the user's specialty + location as query params. Extract practice website URLs from results. Deduplicate by domain.
Fallback (always available):
nimble search --query "[specialty] in [location]" --max-results 20 --search-depth liteExtraction phase (pulling provider data from sites)
When: Always — this is the core pipeline.
No WSA needed. This phase uses nimble map + nimble extract directly on practice websites. See provider-extraction-patterns.md for page scoring and field detection.
However, if Layer 2 discovery found a WSA that extracts provider data from a specific healthcare directory (e.g., a future zocdoc_provider_profile or healthgrades_doctor_profile agent), use it for practices listed on that directory instead of scraping their website directly — structured WSA output is higher quality than parsed markdown.
Enrichment phase (optional, on request)
When: User asks for practice reputation, regulatory data, or you're suggesting next steps.
Reputation — search terms: review, google_maps_reviews, yelp, bbb
- Use review agents with
place_idor practice URL from the discovery phase - Use BBB agents for business credibility checks
Regulatory — search terms: healthcare vertical, clinicaltrials, fda, npi, license
- Use any discovered regulatory agents to cross-reference providers or practices
- Particularly valuable for clinical trial involvement, device clearances, drug
research
Practice details — search terms: directory-specific agents found in Layer 2
- If an agent provides structured practice profiles (hours, insurance, staff count),
use it to supplement extracted data
---
Scaling WSA Calls
Follow the Scaled Execution pattern from nimble-playbook.md — it covers individual calls, batching, and the confirmation gate for large jobs.
---
Fallback Chain
If WSA discovery returns nothing useful for a phase, fall back to nimble search + nimble extract (the core Nimble tools always work):
1. Discovery fallback: nimble search --query "[specialty] in [location]" --max-results 20 --search-depth lite 2. Enrichment fallback: nimble search --query "[practice-name] reviews" --max-results 5 --search-depth lite + nimble extract on results 3. Regulatory fallback: nimble search --query "[provider-name] [credentials] NPI OR license OR board certification" --max-results 5 --search-depth lite
The skill always produces results even if zero WSAs are found — WSAs accelerate and enrich, but are never required.