
Procurer
- 4 installs
- 3 repo stars
- Updated August 5, 2026
- broomva/skills
procurer is a Claude Code skill that turns a real-world need into 3-5 ranked alternatives with cited price bands, confidence scores, and a final recommendation with a budget.
About
procurer is a Claude Code skill for grounded procurement research on any real-world need. It turns a problem into 3-5 ranked alternatives across DIY-retail, mid-retail, specialty, contractor, and consultant/turnkey tiers, each with cited low/typical/high price bands, confidence scores, and locale-aware currency. Developers use it when a decision needs calibrated cost estimates before committing. The output is decision-shaped, ending in a recommendation and budget envelope rather than a knowledge artifact.
- Turns a real-world need into 3-5 ranked alternatives with cited price bands
- Spans DIY-retail to contractor to consultant/turnkey provider tiers
- Output is decision-shaped: a recommendation plus a budget envelope
Procurer by the numbers
- 4 all-time installs (skills.sh)
- Ranked #2,331 of 3,282 Productivity & Planning skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
procurer capabilities & compatibility
- Capabilities
- research · web search · data analysis
- Use cases
- research · web search · data analysis
- Pricing
- Free
What procurer says it does
Grounded procurement research for any real-world need.
Output is decision-shaped (recommendation + budget envelope), not research-shaped (knowledge artifact).
Produce a list of **3–5 alternatives** ordered by cost and disruption.
npx skills add https://github.com/broomva/skills --skill procurerAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 4 |
|---|---|
| repo stars | ★ 3 |
| Last updated | August 5, 2026 |
| Repository | broomva/skills ↗ |
What it does
Turn a real-world need into 3-5 ranked, cited cost alternatives and a final recommendation with a budget envelope.
Who is it for?
Making a spending decision that needs calibrated cost estimates across DIY, service, and managed alternatives
Skip if: Pure knowledge research with no decision (use deep-research) or technology/library/API selection (use technical-research)
When should I use this skill?
The user asks how much something costs, what their options are, whether to DIY or hire, or describes a problem that implicitly needs a budget
What you get
A decision-shaped report: 3-5 ranked alternatives with cited price bands and a final recommendation plus budget envelope.
- 3-5 ranked cost alternatives with cited price bands
- confidence scores per alternative
- final recommendation with budget envelope
By the numbers
- 3-5 ranked alternatives per pass
- 5 provider-archetype tiers (T1-T5)
- 5-stage procedure
Files
procurer — Grounded Procurement Research
What this skill is
The procurer skill turns a real-world need into a grounded cost-calibration report with cited alternatives, price bands, confidence, and a recommendation. It is the agent's reflex when the user is about to decide how to spend money or effort on a problem and needs calibration before committing.
Concretely, the output of one procurer pass is:
1. A clear restatement of the need (and the underlying problem if the user named only a symptom). 2. 3–5 ranked alternatives, each placed on a provider-archetype tier (DIY-retail → mid-retail → specialty product → contractor → consultant / turnkey). 3. For each alternative: a cost band (low / typical / high) with the currency normalized to the user's locale, plus cited sources (provider URL + page title + fetched-at timestamp) and a confidence score (0–1). 4. Cross-cutting notes: tax/VAT handling, locale-specific suppliers, lead times, hidden costs, deal-breakers. 5. A recommendation — which alternative(s) to pursue, in what order, with the budget envelope.
The output is decision-shaped: the user should be able to read it and act, not need to do further synthesis.
---
When to invoke (the reflex)
The procurer reflex fires on any of these signals:
| Signal | Example |
|---|---|
| Explicit cost question | "How much does X cost?" / "Give me a budget for Y." |
| Options request | "What are my options for fixing the noise?" / "Should I get a mini-split or a window unit?" |
| Hire-someone question | "Who could install this?" / "Should I get a consultant?" |
| Problem with no cost asked | User describes a problem, no budget request — surface alternatives + bands anyway so they can decide. |
| Multiple-vendor comparison | "Compare suppliers for varilla #4." |
| Build-vs-buy / DIY-vs-pro | "Should I do this myself or hire someone?" |
Anti-trigger: if the user is researching a topic with no decision attached (e.g., "explain how acoustic windows work"), use deep-research instead. The line: is there a budget envelope at the end?
---
The 5-stage procedure
Stage 1 — Decompose the need into ranked alternatives
Before any search, restate the need in plain terms and separate the symptom from the failure mode. A "noisy window" is the symptom; the failure mode is usually seal leakage (~90% of the energy on residential sliders) before it's glass mass deficiency. Cheap fixes target the dominant failure mode; expensive fixes replace the system.
Use one of the canonical decomposition patterns (see references/decomposition-patterns.md):
- Incremental → augmentation → replacement (most physical / fix-it problems).
- DIY → service → managed (when responsibility transfer is a real lever).
- Standard → custom → bespoke (when specificity drives cost).
- Single-vendor → multi-vendor → integrator (sourcing complexity).
Produce a list of 3–5 alternatives ordered by cost and disruption. State the thesis of each — what problem it actually solves, not just what it is.
Stage 2 — Map provider archetypes per alternative
For each alternative, identify which provider archetypes apply (see references/provider-taxonomy.md):
| Tier | Archetype | Examples |
|---|---|---|
| T1 | DIY-retail | Big-box / hardware store / marketplace / online retail. Consumables and parts the user installs. |
| T2 | Mid-retail / specialty product | Specialty store with installation optional. |
| T3 | Specialty product / fabricator | Manufacturer / branded supplier / custom-fab shop. |
| T4 | Contractor / installer | Service that does the work end-to-end with materials it sources. |
| T5 | Consultant / engineer / turnkey | Advisory or full-service end-to-end management. |
Not every alternative spans every tier. Capture which tiers are relevant per alternative.
Stage 3 — Choose the search mode
Based on the user's urgency, budget headroom, and decision stakes, choose fast / standard / deep. See references/mode-tiers.md for full contracts.
| Mode | Searches | Providers cited per alt | Latency target | When |
|---|---|---|---|---|
| fast | 1 per alt | 1–2 (T1 only) | < 1 min total | "Just give me a rough number." |
| standard | ≥3 per alt | 3–5 (cover ≥2 tiers) | ~3–5 min | Default for most decisions. |
| deep | ≥6 per alt | 5–7 (cover ≥3 tiers) | best-effort | "I'm actually deciding now, need the full picture." |
Default to standard unless the user explicitly signals "quick" or "thorough."
Stage 4 — Search with grounding discipline
For each (alternative, provider archetype) pair, run grounded web searches. The discipline (see references/grounding-discipline.md) is non-negotiable:
1. Citation required. Every price must carry source_url, source_title, fetched_at (ISO-8601 UTC). No bare numbers. 2. No fabrication. If no public price exists for that pair, say so explicitly and leave it empty. Never fill from training data. 3. Confidence (0–1).
≥ 0.90— exact product page, SKU + unit explicit.0.70 – 0.89— listing exists but unit/spec requires interpretation, or quote-on-request signals.< 0.70— category/inferred from comparable item; flag explicitly.
4. Unit & currency normalization. Quote in the user's locale currency with the right thousand/decimal conventions. Be explicit about tax (IVA/VAT inclusive vs. add). 5. Sanity bands. If a price falls > 2× the median for that category, flag it in notes — don't drop it (the human decides). 6. Diversity bias. For standard and deep, prefer at least two tiers — a Tier-1 retail anchor plus a Tier-3+ specialty or contractor benchmark. 7. Locale-aware suppliers. Use locale-appropriate domains/brands. Default to user's stated region; ask if unclear.
When the agent has a WebSearch / WebFetch tool, use it. When it doesn't, mark the report as unsourced calibration — provide ranges from prior knowledge but explicitly note the absence of fresh citations and recommend the user run a sourced pass before committing.
Stage 5 — Render the report
Produce the report using references/report-template.md. The skeleton:
# <Need restated in one line>
## Problem framing
<2-4 sentences: what's the actual failure mode, not just the symptom>
## Alternatives
### Alternative A — <name> (Tier T1 → T3)
**Thesis**: <what this actually solves>
**Cost band (locale)**: low – typical – high
**Confidence**: 0.X
**Providers cited**: <N> — see footnotes [1] [2] [3]
**Notes**: <lead time, hidden costs, deal-breakers, tax handling>
### Alternative B — ...
### Alternative C — ...
## Cross-cutting notes
- Tax / VAT / IVA treatment
- Locale-specific supplier shortlist
- Common hidden costs
- Lead times
## Recommendation
**Start with**: <Alternative X>
**Total budget envelope**: <low – high>
**Rationale**: <2-3 sentences>
**If that doesn't work**: <Alternative Y as fallback>
## Sources
[1] <url> — <title> — fetched <iso8601>
[2] ...Then optionally call scripts/validate_report.py <report.md> to lint structural completeness (≥3 alternatives, every alternative has cost band + confidence + ≥1 citation, recommendation present, currency consistent).
---
Grounding rules (binding)
These rules bind every procurer run regardless of mode:
1. No price without a citation. A number alone is not procurement research — it's a hallucination risk. 2. No alternative without a thesis. Don't list options; list options with the problem each solves. 3. No recommendation without a budget envelope. "Go with X" is incomplete; "Go with X, $A–$B inclusive of installation" is decision-shaped. 4. No locale assumption. If the user hasn't stated their region, ask before searching — supplier networks and tax handling diverge sharply. 5. Flag the dominant failure mode early. If 80% of the user's problem can be solved by a 5% intervention (Tier-1 fix), say so before they spend Tier-4 money. Honesty about what actually causes the problem is the most valuable output.
---
Resources
references/
decomposition-patterns.md— four canonical patterns for breaking a need into alternatives.provider-taxonomy.md— the 5-tier provider archetype model, with examples across domains.grounding-discipline.md— citation / confidence / locale / tax rules in full.mode-tiers.md— fast / standard / deep search contracts with budgets and SLA targets.report-template.md— the report skeleton and a fully-filled exemplar.
scripts/
validate_report.py— structural linter for a generated procurement report. Checks alternatives count, citation completeness, confidence range, recommendation presence, currency normalization. Exit non-zero on failure.
assets/examples/
window-noise-attenuation.md— full worked example: bedroom window noise on a Bogotá avenue, three tiers of remediation.construction-materials-co.md— generalized Colombian construction materials reference (CO suppliers, families, IVA handling) — lifted from the materiales-intel.v1 rules-package pattern.
---
Compounding with other skills
- `deep-research` — when the user wants to learn about a topic before deciding, run deep-research first, then procurer for the cost layer.
- `technical-research` — for software/library choice with a cost dimension, do technical-research for the technical evaluation, then procurer for SaaS pricing / consultancy / implementation costs.
- `bookkeeping` (P8) — procurement reports that produce reusable knowledge (e.g., "the CO porcelanato market spans $35k–$120k/m²") should be filed into the entity graph via
bookkeeping file.
---
Closing handoff
The skill ends when the report renders. The agent's last message to the user is the report itself plus a one-line action prompt: "Want me to deep-dive any alternative, refresh citations, or proceed with a specific provider?"
Procurer never spends the money. It calibrates the spend.
Colombian Construction Materials — Locale Reference
A locale-specific reference for procurer runs covering CO construction materials. Lifted and generalized from the materiales-intel.v1 rules-package skeleton at freelance/_pending-constructora/rules-package/.
This file is reference data — supplier shortlists, family taxonomy, sanity bands, tax notes — to compound on top of references/grounding-discipline.md when a procurer run touches CO construction.
Note. This reference is informational — it doesn't replace running grounded web searches per Rule 1. Use it to pre-narrow the supplier search space and validate sanity bands against historical CO market knowledge.
---
Supplier shortlist by tier (CO)
Tier 1 — Grandes superficies (national retail chains)
| Provider | Domain | Notes |
|---|---|---|
| Homecenter | homecenter.com.co | Largest national chain; full catalog; reliable price page. |
| Sodimac | sodimac.com.co | Same parent as Homecenter; some store-by-store price variation. |
| Easy | easy.com.co | Smaller footprint; pricing competitive on bulk. |
| Constructor | constructor.com.co | Cemex's retail brand; strong on basics. |
Tier 2 — Fabricantes / marcas líderes (manufacturer direct, branded)
| Provider | Domain | Categories |
|---|---|---|
| Argos | argos.co | Cemento gris/blanco. |
| Cemex | cemex.com | Cemento; ready-mix. |
| Holcim | holcim.com.co | Cemento. |
| Ultracem | (search) | Cemento — regional. |
| Corona | corona.co | Cerámica, sanitarios, grifería. |
| Alfagres | alfagres.com | Pisos cerámicos / porcelanatos. |
| Grival | grival.com.co | Grifería. |
| Pavco | pavco.com.co | Tubería PVC/CPVC. |
| Pintuco | pintuco.com.co | Pintura. |
Tier 3 — Especializados (branded, narrower catalog)
| Provider | Domain | Categories |
|---|---|---|
| Cerámica Italia | ceramicaitalia.com.co | Cerámicas, baldosas. |
| FV | fv.com.co | Grifería premium. |
| Acquagrif | acquagrif.com | Grifería. |
| Gerfor | gerfor.com | Tubería. |
| Tigre-Celta | tigrecelta.com.co | Tubería. |
| Sherwin-Williams | sherwin.com.co | Pintura premium. |
| Gerdau Diaco | (search) | Hierro / acero estructural. |
| Acerías Paz del Río | (search) | Hierro. |
| Ternium | (search) | Hierro. |
| Sidoc | (search) | Hierro. |
| Tecnoglass | tecnoglass.com | Vidrio (DVH, laminado, acústico). |
| Vitelsa | (search) | Ventanería de aluminio. |
Tier 4 — Contratistas / instaladores
- Constructora local (Bogotá / Medellín / Cali).
- Instalador de vidrio / aluminios independiente.
- Contratista de plomería, electricidad, pintura por gremios separados.
Tier 5 — Consultores / arquitectos / ingenieros
- Arquitecto independiente.
- Ingeniero estructural (cuando el alcance toca estructura).
- Ingeniero acústico (e.g., AcustiCo para ventanería acústica).
- Gerencia de proyecto (para obras > 200 M COP).
---
Familia taxonomy (canonical units + synonyms)
For each family the agent uses these canonical units when normalizing prices. Synonyms help search-query construction.
| Familia | Canonical unit | Synonyms / search terms |
|---|---|---|
| hierro | unidad (varilla); kg; ton | varilla, hierro figurado, hierro liso, hierro corrugado, acero de refuerzo |
| cemento | saco (50 kg or 42.5 kg) | cemento gris, cemento blanco, cemento portland, Argos, Cemex, Holcim, Ultracem |
| pisos | m² | porcelanato, cerámica, gres porcelánico, piso mate, piso pulido |
| baldosas | m² | baldosa, enchape muro, enchape piso, revestimiento cerámico |
| griferia | unidad | grifo, grifería, llave, lavamanos, sanitario, ducha, monomando |
| tuberia_pvc | metro lineal (ml) | tubería PVC, tubo PVC, PVC presión, PVC sanitaria, Pavco, Gerfor |
| tuberia_cpvc | metro lineal (ml) | tubería CPVC, agua caliente, Pavco CPVC |
| pintura | galón / cuñeta | pintura, vinilo, esmalte, anticorrosivo, Pintuco, Sherwin, Corona |
| electrico | metro lineal (ml) | cable AWG, canaleta, caja de paso, tomacorriente, breaker, interruptor |
Sanity-band multipliers (flag outliers at > N× the family median)
| Familia | Multiplier |
|---|---|
| hierro | 1.5× |
| cemento | 1.3× |
| pisos | 2.0× |
| baldosas | 2.0× |
| griferia | 2.5× |
| tuberia_pvc | 2.0× |
| tuberia_cpvc | 2.0× |
| pintura | 1.8× |
| electrico | 2.0× |
Lower multiplier = tighter market (commodity-shaped, e.g. cemento). Higher multiplier = wider market (spec-driven, e.g. grifería).
---
Tax handling (CO)
- IVA = 19% on most goods and services. Flag in every band:
"Precio con IVA 19% incluido"if the source page shows tax-inclusive."Precio sin IVA — agregar 19%"if the source page shows the bare commercial price.- Retenciones (informational for corporate buyers):
- ReteFuente: 2.5% on goods, 4% on services (above DIAN umbral).
- ReteIVA: 15% of the IVA, when retainer applies.
- ReteICA: varies by municipio (Bogotá: 6.96‰ to 13.8‰ depending on actividad económica).
- Personas naturales generally don't apply retenciones on small purchases.
- Régimen simple / Régimen común — affects RUT-side handling, not the procurement price itself.
State tax treatment explicitly per band; never silently assume IVA-inclusive or IVA-exclusive.
---
Regional notes
| Region | Notes |
|---|---|
| Bogotá | Largest market; best supplier diversity; most price discovery available online. |
| Medellín | Strong on aluminum/glass (proximity to manufacturers); some category leadership (e.g. Tecnoglass headquarters). |
| Cali | Strong on cement (Cemex, Argos regional); lighter on premium specialty. |
| Eje Cafetero (Pereira/Manizales/Armenia) | Mid-market depth; some online presence; bias toward local distributors. |
| Costa Caribe (Barranquilla/Cartagena/Santa Marta) | Stronger import / international flow; premium specialty available via Barranquilla port; humidity affects spec choices (e.g. acabados marinos). |
| Oriente (Bucaramanga + Cúcuta) | Mid-tier; commodity strong; specialty thinner online. |
For regions outside Bogotá, ask the user whether to: 1. Search the user's region first (locale-pure). 2. Use Bogotá as a price anchor (broader supplier base, may not deliver to user's region without freight).
---
Common procurement-flow patterns for CO construction
These are the canonical decompositions for the most-asked CO procurer questions:
"How much will it cost to build/renovate <X>?"
Pattern 4 — Single-vendor → Multi-vendor → Integrator. Default to multi-vendor + arquitecto supervision for residential renovations $20M–$200M; recommend integrator for >$200M.
"What's the price of <material> for my obra?"
Pattern 3 — Standard → Custom → Bespoke. Standard SKU at T1 retail anchors the floor; T2/T3 manufacturer-direct often beats T1 above a volume threshold (~5M COP).
"Should I buy <material> retail or get the manufacturer?"
Pattern 4 in miniature — usually T1 retail wins under 1M COP, T2/T3 manufacturer wins above 5M COP (volume discount), tie zone in between.
"Who can install <X>?"
Pattern 4 directly — Single-vendor (retailer-bundled install) vs. Multi-vendor (separate trades) vs. Integrator (contratista general). Most residential users default-pick Multi-vendor and discover Integrator's value too late.
---
When this reference applies
Use this file when:
- User is in CO and the procurement involves construction materials, ventanería, plomería, eléctrico, pintura, pisos, baldosas, sanitarios.
- User asks for CO supplier shortlists or "where do I buy X in Bogotá / Medellín / etc."
- Tax treatment is ambiguous and needs explicit IVA / retención framing.
- Cross-referencing against the materiales-intel.v1 rules-package is useful (the canonical taxonomy lives here, not the rules-package — the rules-package is a deployed tenant artifact derived from this).
---
Composition with the broader procurer skill
This file is consumed at Stage 1 (decomposition) and Stage 4 (grounded search) of the procurer procedure:
- Stage 1 — When the family maps to one of the 9 CO construction families, the decomposition pattern is usually Pattern 1 (fix-it problems) or Pattern 3 (acquire-something problems).
- Stage 4 — Use the supplier shortlist as the allowed-domain set for
WebSearch; pre-filter results to known providers before broadening.
For a CO constructora tenant operating in production on the Life-runtime engine, the canonical artifact is the deployed rules-package, not this reference. The reference is for ad-hoc procurer runs and skill-internal calibration.
Window Noise Attenuation — Bedroom on Bogotá Avenue
Worked example. Output shape of onestandard-mode procurer run on a real need surfaced in conversation 2026-05-13. Treats the live answer Claude gave in chat as the unsourced-calibration baseline; a real run would replace[citation pending]placeholders with grounded sources from Homecenter, Sodimac, Tecnoglass, Vitelsa, etc.
>
Use this as a reference shape, not as final price data. The bands here are calibration ranges informed by CO market familiarity, not citations.
---
Reduce traffic noise from bedroom window on a Bogotá avenue
Locale: CO Bogotá Currency: COP Mode: standard (calibration; replace placeholders with sourced run before committing) Generated: 2026-05-13T17:00Z
Problem framing
The bedroom faces an avenue. Existing window is an aluminum slider with brush-pile (felpa) seals. The user reports that the windows attenuate "some" noise but traffic is still audible.
Dominant failure mode (Grounding Rule 8): sliding windows leak sound primarily through seal path (brush felpa is a wind-stop, not a sound-stop — air gaps between bristles transmit ~30–50% of acoustic energy on a 1% gap area). The glass path (monolithic 4–6 mm pane has a coincidence dip in the 1–4 kHz traffic band) is the secondary leak. Flanking path (aluminum frame) is third-order.
This means the cheapest Tier-1 intervention captures the largest share of the available improvement. The user should not commit Tier-3 money before trying Tier-1.
Alternatives
Alternative A — Tier-1 fix: replace seals, adjust rollers, add curtains (Tiers: T1)
Thesis. Close the dominant failure mode (felpa-seal leakage) without modifying the window system. Captures 30–50% of available perceived-noise reduction at 5% of the Tier-3 cost.
Cost band (COP).
- Low: $ 80.000 (felpa only, DIY)
- Typical: $ 250.000 (felpa + acoustic curtains + roller-adjust labor)
- High: $ 600.000 (premium blackout/acoustic curtain + installer)
Confidence: 0.65 (unsourced calibration; needs sourced refresh — felpa pricing is well-known per-meter, curtain pricing varies widely).
Providers cited (N=4 expected after sourced run):
| Tier | Provider | Price | Unit | Confidence | Source |
|---|---|---|---|---|---|
| T1 | Homecenter | $ 5.000–15.000 | metro felpa siliconada | — | [citation pending] |
| T1 | Sodimac | $ 4.500–12.000 | metro felpa siliconada | — | [citation pending] |
| T1 | Easy | $ 4.500–14.000 | metro felpa siliconada | — | [citation pending] |
| T1 | Mercado Libre — local ferretería | $ 80.000–250.000 | cortina acústica 2×2m | — | [citation pending] |
Notes.
- Felpa "siliconada" / "con aleta" (dual-fin with center plastic film) is the correct upgrade; replace the brush-only stock seal.
- Roller adjustment is a screw on the bottom-rail (visible in user photo IMG_7082). Free if DIY; ~$80k if you hire an aluminios installer for 1 hour.
- Curtains help mid-high frequencies only; expect 3–5 dB perceived reduction.
- Excludes IVA where listed; CO buyer should budget IVA 19% on materials.
- Lead time: same-day to 48h.
Alternative B — Tier-2/3 augmentation: internal acoustic secondary window (Tiers: T3 → T4)
Thesis. Add a second glazed barrier inside the existing window with a 5–10 cm air gap. The air gap kills low-frequency rumble that single-pane mass solutions miss. Typical Rw improvement: 15–25 dB on top of existing window.
Cost band (COP).
- Low: $ 600.000 / m² (acrylic magnetic insert, DIY-spec'd / custom-fab)
- Typical: $ 900.000 / m² (casement secondary window, 6mm laminated, installed)
- High: $ 1.300.000 / m² (premium acoustic interior, dual-laminated, gasketed)
Confidence: 0.60 (unsourced calibration; CO secondary-window market is fragmented across local talleres de aluminios, prices vary 1.5× across providers).
Providers cited (N=3–5 expected after sourced run):
| Tier | Provider | Price | Unit | Confidence | Source |
|---|---|---|---|---|---|
| T3 | Local taller de aluminios (Bogotá centro) | cotización a solicitud | m² instalado | — | [citation pending] |
| T3 | Vitelsa | cotización a solicitud | m² instalado | — | [citation pending] |
| T4 | Constructora Acústica CO (varios) | cotización a solicitud | m² instalado | — | [citation pending] |
Notes.
- "Ventana interior" / "contraventana acústica" — terminology to use when calling talleres.
- Single most effective intervention short of full replacement.
- Reversible — can be removed if the user moves.
- Excludes IVA; install includes labor + materials in cited bands.
- Lead time: 2–4 weeks for fabrication + 1 day install.
- Outlier-low flag: any quote below $ 500k/m² is suspect (verify SKU includes laminated glass, not just monolithic).
Alternative C — Tier-3/4 replacement: full acoustic DVH + casement frame (Tiers: T3 → T4)
Thesis. Replace the slider with a casement (compression-seal) frame and laminated DVH (asymmetric panes, PVB acoustic interlayer). Typical Rw 38–42 dB vs ~25–28 dB current. Only justified if the user is already renovating or wants the thermal/aesthetic side-benefits.
Cost band (COP).
- Low: $ 1.200.000 / m² (entry acoustic DVH, casement, no premium brand)
- Typical: $ 1.700.000 / m² (Tecnoglass / Vitelsa / Alfa with acoustic-rated DVH)
- High: $ 2.500.000 / m² (premium PVB Saflex Q / Trosifol SC, tilt-turn frame)
Confidence: 0.70 (CO acoustic DVH market priced via manufacturer distributors; bands well-anchored).
Providers cited (N=4–6 expected after sourced run):
| Tier | Provider | Price | Unit | Confidence | Source |
|---|---|---|---|---|---|
| T3 | Tecnoglass distribuidor Bogotá | cotización a solicitud | m² instalado | — | [citation pending] |
| T3 | Vitelsa | cotización a solicitud | m² instalado | — | [citation pending] |
| T3 | Alfa Arquitectura y Concreto sub | cotización a solicitud | m² instalado | — | [citation pending] |
| T4 | Constructora local con ventanería | cotización a solicitud | proyecto | — | [citation pending] |
Notes.
- Ask for the Rw rating with documentation. Most CO suppliers quote DVH without acoustic certification — that's a red flag.
- A real acoustic window will have a test certificate showing Rw ≥ 35 dB.
- Casement (batiente) or tilt-turn (oscilobatiente) only — sliders disqualify the acoustic spec.
- Includes demolition + install + interior re-finishing (~20% of bare DVH cost typically rolled in).
- IVA 19% applies; corporate buyers add ReteFuente 4% on services.
- Lead time: 6–10 weeks (fabrication-limited, especially on imported acoustic laminate).
Alternative D — Tier-5 advisory: acoustic engineer assessment (optional pre-step) (Tiers: T5)
Thesis. Before committing Tier-3 money, hire 2–4 hours of an acoustic engineer to measure current Rw, identify the dominant leak path empirically, and confirm whether B or C is actually warranted. Pays for itself if it prevents an unnecessary C-tier purchase.
Cost band (COP).
- Low: $ 350.000 (single-visit measurement, no report)
- Typical: $ 800.000 (visit + 2-page report with recommendations)
- High: $ 2.500.000 (full residential acoustic audit with 3D modeling)
Confidence: 0.55 (acoustic consultancy in CO is small, prices vary widely; CR1 / AcustiCo / independent consultants).
Providers cited (N=2–3 expected after sourced run):
| Tier | Provider | Price | Unit | Confidence | Source |
|---|---|---|---|---|---|
| T5 | AcustiCo (Bogotá) | cotización a solicitud | visita | — | [citation pending] |
| T5 | Consultor acústico independiente | cotización a solicitud | visita | — | [citation pending] |
Notes.
- Skip if user is committing < $ 2M total; the consultant fee is too high a share.
- Worth it if user is contemplating C (>$ 5M total).
Cross-cutting notes
- Tax treatment. All bands exclude IVA 19% unless noted. Corporate buyers add applicable retenciones (ReteFuente 4% on services + ReteICA per municipio).
- Bogotá supplier shortlist.
- T1 retail: Homecenter (multiple stores), Sodimac (Av 68 + Calle 80), Easy.
- T3 fabricators: Paloquemao / 7 de Agosto area concentrates the talleres de aluminios.
- T3 manufacturers: Tecnoglass (representante Bogotá), Vitelsa (Calle 80), Alfa.
- T5 consultants: AcustiCo, independent acoustic engineers (search "ingeniero acústico Bogotá").
- Hidden costs. Interior re-finishing (paint, drywall touch-up) often 15–20% above bare-installed window cost on a replacement. Curtain rod / hardware for Alternative A's curtain step is ~$ 50–150k extra not in the band.
- Lead times. A: same-day. B: 2–4 weeks. C: 6–10 weeks. D: 1–2 weeks for the consultant visit.
- Locale calibration. Bogotá apartments above floor 12 on major avenues (Calle 80, Carrera 7, Av Boyacá, Calle 26) show consistent 60–72 dBA daytime traffic noise — the user's experience is typical for the geography.
Recommendation
Start with: Alternative A (Tier-1 felpa + rollers + curtains).
Total budget envelope: $ 80.000 – $ 600.000 COP (one-time, weekend execution).
Rationale (3 sentences). The dominant failure mode on a brush-felpa aluminum slider is seal leakage, which Alternative A addresses directly at ~5% of Alternative C's cost. The user should run A first, give it 2–3 weeks of lived experience, and only escalate if the residual noise is still meaningful. The cost of A is small enough that running it as a diagnostic is rational even if the user later commits to C.
If that doesn't work: Alternative B (internal acoustic secondary window) — $ 600k–1.3M COP/m² installed. Escalate to B only if Alternative A's perceived noise reduction is < 50% after 2–3 weeks of daily use, or if low-frequency rumble (trucks, motos with modified exhausts) is the dominant residual complaint.
Skip Alternative C unless the user is already renovating that wall or wants the thermal/aesthetic side-benefits. C is over-spec'd for noise alone after B is in place.
Alternative D (consultant) is worth it only if the user is contemplating C without renovation; for the A → B path, A's diagnostic value substitutes for it.
Sources
[citation pending] — Sourced refresh required before committing. The bands in this report are calibration ranges from prior CO market knowledge and conversation context, not live citations. Recommend a standard-mode procurer run with WebSearch enabled to convert placeholders into Homecenter / Sodimac / Tecnoglass / Vitelsa product pages.
---
What this example demonstrates
- Pattern 1 decomposition (Incremental → Augmentation → Replacement) plus an optional Tier-5 consultant rung as a diagnostic.
- Honest framing of the dominant failure mode (felpa-seal leakage) before the alternatives, per Grounding Rule 8.
- Cost-disruption span across three orders of magnitude (A: $80k–$600k; B: $600k–$1.3M/m²; C: $1.2M–$2.5M/m²).
- Explicit recommendation with a fallback condition ("if A doesn't reduce ≥50% in 2–3 weeks, escalate to B").
- Locale-aware suppliers, tax handling, lead times.
- Confidence transparency — every band carries a confidence reflecting that this is unsourced calibration, not a live grounded run.
Changelog
All notable changes to the procurer skill are documented here. The format follows Keep a Changelog, and this project adheres to Semantic Versioning.
[Unreleased]
[0.1.0] — 2026-05-13
Added
- Initial release of the
procurerskill — grounded procurement research for any real-world need. - SKILL.md with the 5-stage procedure (decompose → map providers → choose mode → search with grounding → render report).
- References — five load-on-demand documents:
decomposition-patterns.md— four canonical patterns for breaking a need into alternatives (Incremental/Augmentation/Replacement; DIY/Service/Managed; Standard/Custom/Bespoke; Single/Multi/Integrator).provider-taxonomy.md— the 5-tier provider archetype model (DIY-retail → consultant/turnkey).grounding-discipline.md— eight binding rules (citation, confidence 0–1, locale, tax handling, sanity bands, dominant-failure-mode honesty).mode-tiers.md—fast/standard/deepmode contracts with concrete budgets.report-template.md— output skeleton with validator contract.- `scripts/validate_report.py` — structural linter for procurement reports. Exit non-zero on missing sections, < 3 alternatives, malformed confidence values, currency inconsistency, unresolved footnotes, or missing recommendation fields. Stdlib-only.
- Worked examples in
assets/examples/: window-noise-attenuation.md— bedroom-window noise on a Bogotá avenue, three-tier remediation (felpa replacement → secondary window → acoustic DVH) with optional Tier-5 consultant rung.construction-materials-co.md— Colombian construction materials locale reference (supplier shortlist by tier, family taxonomy, IVA/retención handling, sanity-band multipliers per family).- Tests — pytest suite for the validator covering canonical pass + structural failure modes.
- CI — GitHub Actions workflow running pytest on Python 3.11, 3.12, 3.13.
- OSS packaging — MIT LICENSE, README, CONTRIBUTING, SECURITY, .gitignore.
Lineage
- Extracted the reusable abstractions of the
materiales-intel.v1rules-package (CO construction-materials intelligence module for Broomva Life Agent OS) and generalized them into a domain-agnostic procurement skill. The original construction-specific skeleton remains at~/broomva/freelance/_pending-constructora/rules-package/as a tenant deployment artifact.
[Unreleased]: https://github.com/broomva/procurer/compare/v0.1.0...HEAD [0.1.0]: https://github.com/broomva/procurer/releases/tag/v0.1.0
Contributing to procurer
Thank you for considering contributing. This document explains what's in scope, what to expect from a review, and how to develop locally.
---
What's in scope
The procurer skill exists to do one thing well: turn a procurement need into a decision-shaped report with grounded citations. Contributions should compound on that mission.
Welcome contributions:
- New worked examples in
assets/examples/covering non-construction domains (services, technology, professional services, consumer goods) and locales other than Colombia. - Refinements to the references in
references/— sharper decomposition patterns, additional provider archetypes for specific categories, locale-specific tax/regulatory notes. - Validator improvements in
scripts/validate_report.py— additional structural checks, better error messages, performance. - Tests in
tests/— covering currently-untested validator paths or new structural rules. - CI improvements — packaging, linting, security scanning.
Out of scope:
- Replacing the 5-stage procedure with a different orchestration shape. The procedure is the skill's mechanism.
- Adding runtime dependencies. The validator is stdlib-only on purpose; new dependencies require strong justification.
- Domain-specific business logic that belongs in a downstream consumer (e.g., tenant-specific rules-packages for the Broomva Life Agent OS — those live alongside the tenant, not in this skill).
---
Development setup
git clone https://github.com/broomva/procurer.git
cd procurer
# Optional: create a venv (only needed if you want pytest)
python3 -m venv .venv
source .venv/bin/activate
pip install pytest
# Run the test suite
python3 -m pytest tests/ -v
# Validate the canonical worked example
python3 scripts/validate_report.py assets/examples/window-noise-attenuation.mdPython 3.11+ required. No other runtime dependencies.
---
How to add a new worked example
1. Create assets/examples/<topic>-<locale>.md. 2. Follow the template in references/report-template.md exactly — the validator must pass against it. 3. Include 3–5 alternatives spanning multiple tiers, the dominant-failure-mode framing, a recommendation with budget envelope, and at least placeholder citations. 4. Add a one-line entry to SKILL.md under ### assets/examples/. 5. Run python3 scripts/validate_report.py assets/examples/<your-file>.md — it must exit 0. 6. Open a PR with a short description of the procurement context (locale, decision stakes, mode used).
---
How to add a new reference
References in references/ are load-on-demand depth. They get pulled into the agent's context when the procedure needs them.
1. Create the file under references/. 2. Link it from SKILL.md in the appropriate stage of the procedure (Decompose / Map providers / Choose mode / Search with grounding / Render). 3. Keep each reference focused on one concern — if it's spanning multiple, split it. 4. Cross-link to other references at the bottom.
---
How to change the validator
1. Add a failing test first in tests/test_validator_failures.py (or canonical pass behavior in tests/test_validator_canonical.py). 2. Update scripts/validate_report.py to make the test pass. 3. If the change adds a new structural rule, update references/report-template.md "Validator contract" section. 4. If the change is breaking (an existing valid report would now fail), bump the minor version in CHANGELOG.md and call it out explicitly.
---
Coding conventions
- Python style: PEP 8, type hints on every public function.
- Markdown style: ATX headings (
#, not===), code fences with language tags, no trailing whitespace. - Commit messages: short imperative subject (≤ 72 chars), blank line, optional body. Conventional Commits prefixes (
feat:,fix:,docs:,test:,chore:,refactor:) welcome but not required. - PRs: focused, one concern per PR. Link to a Linear ticket if the change is part of a tracked initiative.
---
Releasing
Maintainers: bump version in CHANGELOG.md, tag the commit, push.
# Once CHANGELOG is updated and merged:
git tag -a v0.X.Y -m "v0.X.Y"
git push origin v0.X.Y
gh release create v0.X.Y --notes-from-tagThe skills.sh registry picks up new tags automatically via the GitHub topic agent-skill.
---
Questions?
Open an issue at <https://github.com/broomva/procurer/issues> or reach the maintainer at <contact@broomva.tech>.
procurer
Grounded procurement research for any real-world need. Turn a problem — "reduce bedroom noise from the avenue", "add air conditioning", "hire an accountant", "buy a TV mount" — into a decision-shaped report with 3–5 ranked alternatives spanning DIY-retail → consultant/turnkey, cited price bands, confidence scores, locale-aware currency, and a clear recommendation with budget envelope.
  
The procurer skill is the agent's reflex when the user is about to spend money or effort on a problem and needs calibration before committing. It enforces grounding discipline (no price without a citation; no fabrication from training data; flag the dominant failure mode early) and produces a report you can act on.
---
Quick start
Install (Claude Code via skills.sh)
# Project-level install:
npx skills add broomva/procurer
# Or globally for your user:
npx skills add -g broomva/procurerYou can also discover it through the interactive finder:
npx skills find procurerInvoke
Once installed, the skill auto-fires on procurement-shaped prompts:
"How much would it cost to fix the noise from my bedroom window?"
"What are my options for adding A/C?"
"Should I hire an accountant or do my taxes myself?"
"Give me a budget for renovating the bathroom."
You can also invoke explicitly with /procurer if your agent supports slash commands.
Validate a generated report
python3 scripts/validate_report.py path/to/report.mdExits non-zero with a list of structural issues if the report doesn't conform to the procurement-report template (≥ 3 alternatives, citations resolve, recommendation has a budget envelope, etc.).
---
What you get
Every procurer run produces a decision-shaped report with this structure:
# <Need restated>
**Locale**: CO Bogotá · **Currency**: COP · **Mode**: standard
## Problem framing
<Dominant failure mode named — separates symptom from cause>
## Alternatives
### Alternative A — <name> (Tiers: T1 → T3)
**Thesis** · **Cost band (low / typical / high)** · **Confidence (0–1)** · **Providers cited (table)** · **Notes**
### Alternative B — ...
### Alternative C — ...
## Cross-cutting notes
<Tax treatment · supplier shortlist · hidden costs · lead times>
## Recommendation
**Start with**: <X>
**Total budget envelope**: <low – high>
**Rationale**: <2–3 sentences>
**If that doesn't work**: <Y> as fallback
## Sources
[1] <url> — "<title>" — fetched <iso8601>
[2] ...See `assets/examples/window-noise-attenuation.md` for a fully-worked example.
---
The 5-stage procedure
1. Decompose the need into 3–5 ranked alternatives using one of four canonical patterns:
- Incremental → Augmentation → Replacement (fix-it problems)
- DIY → Service → Managed (accountability questions)
- Standard → Custom → Bespoke (spec-driven cost)
- Single-vendor → Multi-vendor → Integrator (sourcing complexity)
2. Map provider archetypes — each alternative gets placed on a 5-tier model (DIY-retail → mid-retail → specialty/fabricator → contractor → consultant/turnkey). 3. Choose mode — fast (rough number, < 1 min) / standard (default, ~3–5 min) / deep (executive report, multi-vendor sensitivity). 4. Search with grounding discipline — every price carries source_url, source_title, fetched_at, provider_tier, and a confidence score (0–1). No fabrication; quote-on-request signals captured explicitly. 5. Render the report — using the template above. Validate with scripts/validate_report.py.
Full procedure: `SKILL.md`. Detailed references in `references/`.
---
Why this skill is different
| Skill | Output shape | Use when |
|---|---|---|
procurer (this) | Decision-shaped — recommendation + budget envelope | About to spend money or effort |
deep-research | Research-shaped — knowledge artifact with citations | Understanding a topic |
technical-research | Spike-shaped — technology selection memo | Choosing libraries / APIs / tools |
The differentiator: procurer ends with an actionable recommendation and a budget. The user can read it once and proceed. No follow-up synthesis required.
---
Grounding discipline (binding)
These rules bind every run, regardless of mode:
1. No price without a citation. Every cost band carries source_url, source_title, fetched_at. 2. No alternative without a thesis. Don't list options; list options with the problem each solves. 3. No recommendation without a budget envelope. "Go with X" is incomplete; "Go with X, $A–$B inclusive of installation" is decision-shaped. 4. No locale assumption. Ask the user's region before searching — supplier networks and tax handling diverge sharply. 5. Flag the dominant failure mode early. If 80% of the problem can be solved by a 5% intervention, say so before recommending the 100% solution.
Full rules: `references/grounding-discipline.md`.
---
Repository layout
procurer/
├── SKILL.md # Agent entry point — frontmatter + procedure
├── README.md # This file — user-facing intro
├── LICENSE # MIT
├── CHANGELOG.md # Release history
├── CONTRIBUTING.md # How to contribute
├── SECURITY.md # Vulnerability disclosure
├── references/ # Load-on-demand depth (read by the agent)
│ ├── decomposition-patterns.md
│ ├── provider-taxonomy.md
│ ├── grounding-discipline.md
│ ├── mode-tiers.md
│ └── report-template.md
├── scripts/
│ └── validate_report.py # Structural linter for procurement reports
├── tests/ # pytest suite for the validator
│ ├── conftest.py
│ ├── test_validator_canonical.py
│ ├── test_validator_failures.py
│ └── fixtures/
└── assets/examples/ # Worked examples (treated as output references)
├── window-noise-attenuation.md
└── construction-materials-co.md---
Development
# Run the validator's test suite
python3 -m pytest tests/ -v
# Validate the canonical worked example
python3 scripts/validate_report.py assets/examples/window-noise-attenuation.md
# → OK — passes all procurer-report structural checks.No runtime dependencies — the validator uses only the Python 3.11+ standard library.
---
Compounds with
- [`bookkeeping`](https://github.com/broomva/bookkeeping) (P8) — reusable knowledge from procurement reports (market bands, supplier shortlists) is filed into the entity graph for future runs to compound on.
- `deep-research` — for needs where the user must learn before deciding, run
deep-researchfirst, thenprocurerfor the cost layer. - `technical-research` — for software/library choice with a cost dimension, combine technical evaluation with SaaS/consultancy pricing.
---
License
MIT © 2026 Carlos D. Escobar-Valbuena (broomva). See LICENSE.
---
Acknowledgments
Procurer extracts and generalizes the supplier/taxonomy/grounding discipline of the materiales-intel.v1 rules-package — the construction-materials intelligence module shipped as part of the Broomva Life Agent OS. The construction-specific reference at `assets/examples/construction-materials-co.md` preserves that lineage.
Decomposition Patterns
A procurement need rarely arrives pre-decomposed. The user names a symptom ("bedroom is noisy", "we need A/C", "I want a new kitchen") and expects you to surface the alternatives. How you decompose the need determines the quality of every downstream stage. This document is the menu of canonical patterns.
Cardinal rule. Before applying any pattern, separate the symptom from the failure mode. The cheapest, highest-value alternative usually addresses the dominant failure mode head-on, not the symptom.
---
Pattern 1 — Incremental → Augmentation → Replacement
Use for: Physical / fix-it problems where an existing system is underperforming. The default pattern for "fix the X".
The three rungs
1. Incremental — Treat the dominant failure mode of the existing system. Cheap, fast, reversible. Often DIY-retail (T1). The high-value option 80% of the time. 2. Augmentation — Add a new component alongside the existing system, keeping the original in place. Mid-cost, mid-disruption. Usually T2–T3. 3. Replacement — Tear out and replace the system entirely. Expensive, disruptive, often-overkill. T3–T5.
Worked example — Window noise
| Rung | Alternative | Thesis |
|---|---|---|
| Incremental | Replace felpa seals, adjust rollers, add acoustic curtain | The dominant leak is the brush-pile seal; closing it captures most of the available improvement. |
| Augmentation | Internal secondary window (5–10 cm air gap, casement gasket) | Adds a second barrier with an air gap — the air gap kills low-frequency rumble that mass-only solutions miss. |
| Replacement | Full DVH with laminated PVB acoustic glass + casement frame | Replaces the whole system; only worth it during a broader renovation. |
Why this pattern works
The cost of each rung roughly 10×s the previous one. So does the disruption. But the acoustic improvement often saturates after the augmentation rung — meaning replacement is rarely the right answer unless other constraints (thermal, age, aesthetics) ride along.
Other domains
- Roof leak: patch the membrane (T1) → add a torch-on overlay (T3) → re-roof (T5).
- Slow laptop: clean + reinstall (T1) → SSD/RAM upgrade (T2) → new machine (T2 with disposal of old).
- Hot bedroom: window film + curtain (T1) → portable A/C (T2) → split system installed (T4).
---
Pattern 2 — DIY → Service → Managed
Use for: Recurring tasks or services where the lever is who owns the responsibility, not what the work is.
The three rungs
1. DIY — User does the work, buying only consumables. 2. Service (point) — User hires for the specific task, retains accountability for the outcome. 3. Managed (relationship) — User outsources accountability for the whole category; provider owns the outcome on a retainer/subscription.
Worked example — Personal accounting (Colombia)
| Rung | Alternative | Thesis |
|---|---|---|
| DIY | Use Siigo Lite + DIAN portal for monthly declarations | If your structure is simple (one PN, no employees) the software handles it. |
| Service | Hire a contador independiente for annual reporting | When the structure adds complexity (rentas extranjeras, dividends from Broomva Tech Corp, criptos), a per-engagement contador catches what software misses. |
| Managed | Retainer with a contaduría firm (e.g., Crowe / BDO local) | When you have multiple LLCs / international structure / IRS exposure, a firm on retainer takes the whole category off your plate. |
Other domains
- Cleaning: own supplies (T1) → cleaner once a week (T3) → full facilities mgmt (T5).
- IT support: troubleshoot yourself (T1) → call a freelancer (T3) → MSP on retainer (T5).
- Legal: templates online (T1) → engage per-matter (T4) → outside counsel on retainer (T5).
Common trap
Don't jump straight to Managed. The pricing premium is steep, and the relationship's value only kicks in above a complexity threshold the user may not yet have hit.
---
Pattern 3 — Standard → Custom → Bespoke
Use for: Goods or services where the spec is variable and specificity drives cost.
The three rungs
1. Standard — Off-the-shelf SKU. Discoverable on a retailer's site, often Tier 1 or Tier 2. 2. Custom — Standard spec with options (size, finish, configuration). Manufacturer/fabricator does light customization. Tier 3. 3. Bespoke — Designed to fit a specific situation; no SKU exists ahead of time. Tier 4 or Tier 5.
Worked example — Kitchen sink
| Rung | Alternative | Thesis |
|---|---|---|
| Standard | Single-basin stainless 50×40 cm from Homecenter | Fits most counters, low cost, available immediately. |
| Custom | Two-basin with drainboard from Tramontina / Franke catalogue | Spec'd to the user's counter dimensions, longer lead time. |
| Bespoke | Concrete or hand-formed copper basin commissioned from a fabricator | Designed in; only worth it for a kitchen renovation with strong aesthetic intent. |
Other domains
- Furniture: IKEA → catalog from a local carpentry shop → bespoke piece by a designer.
- Suit: off-the-rack → made-to-measure → fully bespoke.
- Website: template (Squarespace) → customized template (developer) → bespoke design+build.
Cost curve
Standard → Custom is usually a 1.5–3× cost jump. Custom → Bespoke is often 5–10×. The user should know which jump they're contemplating before they ask "what's the price."
---
Pattern 4 — Single-vendor → Multi-vendor → Integrator
Use for: Complex projects with multiple sub-categories of spend, where the lever is how many vendors the user manages.
The three rungs
1. Single-vendor — One supplier handles the whole scope. Lowest user-side complexity, often higher unit cost (the vendor charges a single-source premium). T4. 2. Multi-vendor — User splits the scope across specialists and coordinates them. Lowest unit cost, highest user-side complexity. T2 + T3 + T4 combined. 3. Integrator — User hires a generalist (architect, GC, systems integrator) to coordinate the specialists. T5.
Worked example — Bathroom remodel
| Rung | Alternative | Thesis |
|---|---|---|
| Single-vendor | "Remodelaciones Andrés" handles everything | Cheapest in management overhead; you'll likely pay 15–25% over the multi-vendor sum. |
| Multi-vendor | Buy fixtures (Homecenter/Corona) + hire plumber + hire enchapador + hire pintor | Cheapest by line-item; you coordinate timing, returns, and rework. |
| Integrator | Hire an arquitecto to spec + supervise + contract subs on your behalf | Architect fee 8–15% of project cost; quality and timing risk drop sharply. |
Other domains
- Wedding: single venue does everything → DIY vendors → wedding planner.
- Software stack: monolithic SaaS → best-of-breed SaaS → systems integrator.
---
Choosing a pattern
The four patterns aren't exclusive — a real need often blends them. Use this decision tree:
Is the existing system underperforming?
├── Yes → Pattern 1 (Incremental → Augmentation → Replacement)
└── No → Is it a recurring service or accountability question?
├── Yes → Pattern 2 (DIY → Service → Managed)
└── No → Is the spec variable across users?
├── Yes → Pattern 3 (Standard → Custom → Bespoke)
└── No → Pattern 4 (Single-vendor → Multi-vendor → Integrator)Most procurement reports land on Pattern 1 or 3. Patterns 2 and 4 mostly come up for service and project scoping.
---
Anti-patterns
- Naming the symptom as the alternative. "The window is noisy" is not an alternative — "replace the felpa" is.
- Three identical alternatives. If A, B, and C all sit in Tier 3 with slightly different brands, you didn't decompose, you shopped. Span the cost-disruption axis.
- Skipping the cheapest rung. Always include the Tier-1 fix even when you think the user wants Tier-4. The Tier-1 option is the honest anchor.
Grounding Discipline
The non-negotiable rules for sourcing every price, citation, and confidence score in a procurement report. These rules exist to prevent the failure mode that destroys procurement research: plausible-sounding numbers with no provenance.
If the report violates these rules, it isn't procurement research — it's a hallucinated wishlist. The user will commit money based on it. Take the rules seriously.
---
Rule 1 — Every price has a citation
No bare numbers. Every cost band that appears in the report carries an explicit source attribution.
For each price-bearing observation, capture:
source_url— URL where the price was read. Prefer product/SKU pages over category pages, and category pages over homepages.source_title— the page title or product name as the provider displays it. Disambiguates when URLs are session-encoded.fetched_at— ISO-8601 UTC timestamp of when the price was read. Prices age fast (commodity materials, currency volatility); the timestamp is the freshness signal.provider_name— the human-readable supplier name (e.g., "Homecenter", "Tecnoglass").provider_tier— T1 / T2 / T3 / T4 / T5 fromprovider-taxonomy.md.
If the agent has no web-fetching tool available, mark every band as unsourced calibration and tell the user explicitly: "These ranges are from prior knowledge; before committing, run a sourced pass with web search."
---
Rule 2 — No fabrication
If no public price exists for an (alternative, provider) pair, leave it empty. Never fill from training data.Specifically:
- Quote-on-request products (most contractors, most T5 services) — record the quote-on-request signal explicitly:
"Tecnoglass — cotización a solicitud, no public price page; estimate range from prior CO market data: $1.5M–2.5M COP/m²; confidence 0.6." - Out-of-stock products — note the listing exists but the price isn't currently shown.
- Discontinued products — note and propose the closest current equivalent.
- Region-mismatched products — if the only public price is from a different region, note the locale gap and discount confidence.
Filling empty cells with training-data numbers is the most damaging failure mode of this skill. Better to ship an incomplete report than a polished one with invented numbers.
---
Rule 3 — Confidence (0–1) on every band
Each cost band carries a confidence score reflecting source quality:
| Confidence | Source quality |
|---|---|
| ≥ 0.90 | Exact product page. SKU + unit + currency + tax-treatment all explicit. Fetched within last 7 days. |
| 0.75 – 0.89 | Listing or quote page exists, but one of: unit needs interpretation; spec doesn't fully match; tax treatment ambiguous; > 30 days old. |
| 0.50 – 0.74 | Category-level price, comparable product, or quote-on-request market estimate from a recent transaction. Locale matches. |
| < 0.50 | Inferred from market familiarity / training data / different locale. Flag in report and recommend a sourced refresh before committing. |
Confidence is not a vibe — it's a defensible position the agent should be able to explain when asked.
---
Rule 4 — Unit and currency normalization
All bands in one report must use the same currency and the same unit conventions per category.
Currency
- Default to the user's locale currency (COP for CO, USD for US, EUR for Eurozone, etc.).
- If the user is multi-locale (e.g., Broomva operates CO + DE), ask. Don't assume.
- Use the locale's thousand/decimal conventions (CO:
$ 1.250.000; US:$1,250,000). - When converting from a foreign-priced source, capture the FX rate used and timestamp it:
"USD 850 → COP 3,400,000 (FX 4000, fetched 2026-05-13)".
Tax / VAT / IVA
- Be explicit about tax inclusion. Three states:
"Precio con IVA 19%"— price already includes tax."Precio sin IVA — agregar 19%"— add tax to the figure."IVA no aplica"— exempt category (e.g., medical, some services).- For CO procurement: also flag ReteFuente / ReteIVA / ReteICA when the user is a corporate buyer above DIAN thresholds, but mark these as informational (the buyer's accountant computes them).
- For US procurement: sales tax varies by state; flag as
"+ state sales tax"unless the source page shows tax-inclusive.
Units
- Normalize to category-canonical units (see
provider-taxonomy.md): - Sheet goods → per m² (CO) or per ft² (US).
- Pipes / wires / cables → per linear meter (ml) or per linear foot.
- Bulk materials → per kg, per ton, per saco (50 kg or 42.5 kg).
- Services → per hour, per visit, or per project (state which).
- When a source uses a different unit, convert and show both:
"$25,000/galón ≈ $6,600/litro".
---
Rule 5 — Sanity bands
For each category, the agent knows roughly what the market median is. If an observed price falls outside [median / 2, median × 2], don't drop it — flag it in `notes` and let the user decide:
"Outlier-high: 2.3× market median. Provider may include premium service or specification beyond standard.""Outlier-low: 0.4× market median. Verify SKU matches and that provider is not a clearance / open-box."
Sanity bands per category live in the locale-specific examples (see assets/examples/construction-materials-co.md for the CO seeds). For new categories, the agent computes a rough median from the cited observations and applies 2× on each side.
---
Rule 6 — Diversity bias
For standard and deep modes (see mode-tiers.md):
- Cite ≥ 2 tiers per report — typically a T1 anchor and a T3+ benchmark.
- For each individual alternative spanning multiple tiers, cite ≥ 1 provider per tier spanned.
- For a
standardprice comparison within one alternative, ≥ 3 providers total.
The reason: a single-tier report hides the structural choice. A report citing only T4 contractors hides the option of going T2; a report citing only T1 hides the option of professional install. Diversity is the report's honesty signal.
---
Rule 7 — Locale-aware suppliers
Use suppliers/domains appropriate to the user's market.
- Default the locale to the user's stated region. If the user didn't state it, ask before searching.
- For CO: prefer
homecenter.com.co,sodimac.com.co,easy.com.co,pintuco.com.co,corona.co, etc. Seeassets/examples/construction-materials-co.mdfor the CO allowlist. - For US: prefer
homedepot.com,lowes.com,amazon.com,bestbuy.com, etc. - Marketplaces are second-best. A Mercado Libre / Amazon listing is fine as a fallback but loses confidence vs. a manufacturer or branded retailer's product page.
- Never cite forums, blogs, or AI-summarized prices as primary. Use them only for context, never for the band numbers.
---
Rule 8 — Flag the dominant failure mode
This is grounding discipline at the problem level, not the data level.
Before the cost-band table, name the dominant failure mode of the user's existing situation.
If 80% of the user's problem can be solved by a Tier-1 intervention, say so loudly in the Problem framing section. Don't let the user commit Tier-4 money to solve a problem a Tier-1 fix would handle.
Examples:
- Window noise: "80–90% of perceived noise on a sliding-window installation comes from seal leakage, not glass mass. The Tier-1 felpa replacement captures most of the available improvement at 5% of the Tier-3 glass replacement cost."
- Slow laptop: "Most performance loss on laptops > 3 years old is storage saturation, not CPU. The Tier-1 SSD swap captures most of the speed at 10% of a new-machine purchase."
- Hot bedroom: "Most heat ingress on Bogotá apartments above floor 10 is direct sun on west-facing windows. Tier-1 reflective film addresses a larger share than Tier-4 mini-split installation in many cases."
Honesty about the failure mode is the procurer's most valuable single output. It's also what separates the skill from a marketplace search.
---
Composite checklist
Before rendering the report, the agent verifies:
- [ ] Every cost band has at least one citation with
source_url,source_title,fetched_at,provider_name,provider_tier. - [ ] Confidence scores assigned per Rule 3 criteria, not guesses.
- [ ] All amounts in the same currency with locale conventions.
- [ ] Tax treatment explicit for each band.
- [ ] Units normalized to category canonical.
- [ ] At least 2 tiers cited (for standard/deep).
- [ ] Suppliers locale-appropriate.
- [ ] Dominant failure mode named in Problem framing.
If any box is unchecked, fix before rendering. If a check is impossible (e.g., no web tool), mark the report as unsourced calibration and tell the user.
---
See also: mode-tiers.md for how many searches to run per band, and report-template.md for the final shape.
Mode Tiers — Fast / Standard / Deep
Procurement runs vary by stakes. A quick "roughly what does this cost" deserves a different budget than a $50M-COP decision. This file defines the three search modes with concrete budgets and latency targets.
The mode is chosen at Stage 3 of the procurer procedure (see SKILL.md). Default to standard unless the user signals otherwise.
---
Mode contracts
| Dimension | fast | standard | deep |
|---|---|---|---|
| Use when | Rough order of magnitude needed; user is exploring | Default for real decisions | High-stakes, multi-vendor decisions |
| Web searches per alternative | 1 | ≥ 3 | ≥ 6 |
| Providers cited per alternative | 1–2 | 3–5 | 5–7 |
| Tier coverage per report | 1 tier OK (T1 anchor only) | ≥ 2 tiers (T1 + ≥ 1 higher) | ≥ 3 tiers |
| Confidence floor for any band | 0.50 (with note) | 0.70 | 0.80 |
| Cross-locale fallbacks | OK with note | Discount confidence | Not allowed; fail explicitly |
| Latency target | < 1 min total | 3–5 min | best-effort, no SLA |
| Report length | 1 page / 200–400 words | 2–3 pages / 600–1,000 words | 4+ pages / 1,500+ words |
| Recommendation depth | Picks one alternative, single rationale line | Picks one + names fallback + budget envelope | Picks one + names fallback + budget envelope + scenario sensitivity (worst case, best case, lead-time tradeoffs) |
---
Mode triggers
Use fast when
- The user opens with "roughly" / "ballpark" / "give me a number" / "quickly".
- The decision is < $200 USD-equivalent or < $1M COP.
- The user already knows what they want, just wants to know the price.
- A previous procurer run on the same need ran in
standardordeeprecently; this is a refresh.
Use standard when (default)
- The user is making an actual buy/hire decision.
- The decision is $200–$10,000 USD-equivalent ($1M–$40M COP).
- The user wants to compare alternatives.
- No explicit signal of urgency or thoroughness.
Use deep when
- The user says "thorough" / "executive report" / "I need the full picture" / "presenting this to <board / client>".
- The decision is > $10,000 USD ($40M COP).
- Multi-vendor / multi-trade scope.
- Implementation will take > 2 weeks (long lead times warrant deeper diligence).
- The user is allocating a procurement budget for a team or a project.
Auto-escalation triggers
Within a run, auto-escalate to the next mode if:
fast→standard: only 1 provider found across all alternatives; user explicitly asked to compare; price spread > 3× between the 1 cited and any second-hand reference.standard→deep: price spread > 2× between cited providers; only 1 tier covered after 3+ searches per alternative; user explicitly asked for thoroughness mid-run.
Always notify the user when escalating: "Standard search surfaced only one provider; escalating to deep to widen coverage. Latency budget extends accordingly."
---
Budget caps
Procurement search isn't free — it costs latency, sometimes API calls, and the user's attention. Soft caps per mode:
| Mode | Max wall-clock | Max web searches | Max words in report |
|---|---|---|---|
fast | 90 s | 5 | 500 |
standard | 6 min | 20 | 1,200 |
deep | best-effort, but warn at 20 min | 50 | 2,500 |
These caps are the agent's reflex. When a cap is hit:
- Stop searching.
- Render the report with what's available.
- Mark any unfilled cells explicitly: "No public price located in budget; recommend operator outreach to provider X for direct quote."
---
Mode-specific search strategies
fast strategy
1. Identify the single most-likely provider tier (usually T1 retail). 2. One search per alternative: <sku> <T1 provider name>. 3. Read the price off the first product-page hit. Don't second-source. 4. Render the report with 1–2 providers cited; recommend standard if user wants more.
standard strategy
1. For each alternative, plan 3 searches: one T1 retail anchor + one T2/T3 specialty + one T4 contractor (when applicable). 2. Run them in sequence; if any return no result, swap to a sibling provider in the same tier. 3. Capture all 3 prices with confidence per grounding-discipline.md Rule 3. 4. If after 3 searches per alternative only 1 tier is covered, do a 4th search to widen tier coverage before rendering.
deep strategy
1. For each alternative, plan ≥ 6 searches: T1 anchor + 2–3 specialty/manufacturer + 2–3 contractor/integrator/consultant quotes. 2. Include quote-on-request signals: when a manufacturer's page says "cotización a solicitud", note this and capture a representative market range from prior transactions if available. 3. For multi-trade scopes, search the integrator tier (T5) separately and price the bundled cost vs. the sum of sub-tier prices to surface the integrator's overhead. 4. Compute sensitivity: what does the budget envelope look like at worst-case (highest cited prices), expected (median), and best-case (lowest cited prices)? 5. Render the report with cited providers, scenario sensitivity, and a clear recommendation including fallback paths.
---
Anti-modes (don't do this)
- `xfast` — running without any web search, citing only training data, calling it "fast". Never. Either run web searches or mark the report as unsourced calibration.
- `deeper` — exceeding the deep cap to hunt for one more provider. Diminishing returns; the marginal vendor rarely changes the recommendation. Render and ship.
- `bespoke` — designing a special mode for one report. The three tiers cover the space; pick one.
---
See also: grounding-discipline.md for the citation rules each mode enforces, and report-template.md for the mode-specific report shapes.
Provider Taxonomy
A 5-tier model for classifying who can supply each alternative in a procurement report. The tier isn't about quality or trust — it's about how much of the work the provider does, which is the dominant cost driver.
---
The 5 tiers
Tier 1 — DIY-Retail
What they sell: Consumables, parts, off-the-shelf SKUs. The user installs / uses / consumes.
Cost structure: Catalog price; volume discounts; low margin on commodities, higher margin on accessories.
When this tier is the answer: when the dominant failure mode is a worn part or a missing consumable, and the user is comfortable installing.
Examples by domain:
- Construction: Homecenter (Sodimac), Easy, Constructor (CO); Home Depot, Lowe's (US).
- Tech: Best Buy, Mercado Libre, Amazon.
- Office: Falabella, Mercado Libre.
- Auto: AutoZone, AutoParts.
Signal it's the right tier: the user said "I'll do it myself" or "what tool do I need" or the failure mode is well-understood and the part is cheap.
---
Tier 2 — Mid-Retail / Specialty Product
What they sell: Specialty products with optional installation. Often manufacturer-branded but sold through resellers.
Cost structure: Product margin + optional install fee. Install is often subcontracted.
When this tier is the answer: when the spec is more demanding than a commodity, but the install is straightforward.
Examples by domain:
- Construction: Pavco showroom, Pintuco store, Corona brand store. (CO)
- Tech: Apple Store, brand reseller (Dell Premier).
- Furniture: Tugó, Tramontina, Crate & Barrel.
- Auto: dealer accessories department.
Signal it's the right tier: the user wants a specific brand or spec, doesn't mind paying for the brand premium, and either installs themselves or has it installed by the seller.
---
Tier 3 — Fabricator / Manufacturer / Specialty Service
What they sell: Custom products built to spec, or one-step-up specialty services. Often regional or independent businesses.
Cost structure: Setup fee + materials + labor; lead time meaningful (1–8 weeks typical).
When this tier is the answer: when the standard SKU doesn't fit and customization is real.
Examples by domain:
- Construction: aluminum workshop (taller de aluminios), carpenter, locksmith, wrought-iron shop.
- Tech: contract software house, custom hardware fab.
- Furniture: independent carpenter, upholsterer.
- Auto: independent mechanic, custom shop.
Signal it's the right tier: the user described constraints that don't map to a SKU (irregular sizes, specific aesthetics, integration with existing system).
---
Tier 4 — Contractor / Installer
What they sell: End-to-end service for a defined scope. Sources materials, brings labor, delivers a working result.
Cost structure: Fixed-price or T&M; markup on materials (10–25%) + labor. Liability on the result.
When this tier is the answer: when the user wants a working result without managing the parts.
Examples by domain:
- Construction: general contractor, HVAC installer, electrician, plumber.
- Tech: managed-services provider (MSP), implementation partner.
- Events: catering company, AV company.
Signal it's the right tier: the user said "I just want it done" or the scope crosses trades the user can't coordinate themselves.
---
Tier 5 — Consultant / Engineer / Turnkey / Integrator
What they sell: Advisory, design, supervision, or full-scope project management. The most strategic and the most expensive per hour.
Cost structure: Hourly / day-rate, fixed-fee for scoped engagements, or % of project cost for integration roles.
When this tier is the answer: when the decision itself is hard, the scope is uncertain, or the project is large enough that 10–25% overhead on coordination saves more than it costs.
Examples by domain:
- Construction: architect, structural engineer, acoustic consultant.
- Tech: management consultant, enterprise architect, fractional CTO.
- Finance: tax advisor, wealth manager, M&A advisor.
- Health: specialist physician, second-opinion service.
Signal it's the right tier: the user is uncertain about what to do, not just who to hire; the project is multi-trade and high-stakes; or the user explicitly wants supervised end-to-end execution.
---
Tier picking — rules of thumb
| Situation | Tier likely most useful |
|---|---|
| Existing system, known failure mode, cheap part | T1 |
| Spec'd product, install simple, brand matters | T2 |
| Standard doesn't fit, customization real | T3 |
| Multi-trade or "just want it working" | T4 |
| Decision itself is hard, or project is large + multi-trade | T5 |
Diversity bias for grounded reports
For a standard or deep procurement run, always cite at least two tiers in the report. The Tier-1 anchor gives the user a price floor; the higher tier gives them a benchmark for what they're really comparing against. A report with only Tier-4 quotes hides the alternative of doing it cheaper.
---
Locale-specific examples
Provider names map differently per market. Examples below for the two most common procurer locales.
Colombia (CO)
| Tier | Examples |
|---|---|
| T1 | Homecenter, Sodimac, Easy, Constructor, Mercado Libre, Falabella |
| T2 | Corona, Alfagres, Pintuco, Pavco, Grival store, Tramontina showroom |
| T3 | Tallér de aluminios (local), carpintero (local), Vitelsa, Tecnoglass-distribuidor |
| T4 | Constructora local, instalador-de-vidrio, contratista de plomería |
| T5 | Arquitecto independiente, ingeniero acústico (e.g. AcustiCo), gerencia de proyecto |
United States (US)
| Tier | Examples |
|---|---|
| T1 | Home Depot, Lowe's, Amazon, Best Buy |
| T2 | Apple Store, IKEA, manufacturer-direct online (Andersen, Pella) |
| T3 | Local fabricator, custom millwork shop |
| T4 | General contractor, MEP subs |
| T5 | Architect, structural engineer, consulting firm |
For other locales, ask the user for their region before searching — supplier networks diverge sharply.
---
What changes the tier mapping
- Project size. A T5 architect on a $1,000 job is overkill; on a $100,000 job is cheap.
- User's bandwidth. A user with no time should bias toward T4/T5 even when T2/T3 would be cheaper.
- Repeat factor. If the user will need this category again, a T5 consultant who teaches them how to do it themselves pays for itself.
- Regulatory load. Some scopes (electrical, gas, structural) require a licensed provider — that pushes T4 minimum.
---
See also: decomposition-patterns.md for what the alternatives are, and mode-tiers.md for how many providers to cite per tier.
Report Template
The output of every procurer run conforms to this template. The structure is non-negotiable; the length per section scales with mode (fast / standard / deep).
A report that violates structure fails scripts/validate_report.py. Run the validator before delivering to the user.
---
The skeleton
# <Need restated in one line>
**Locale:** <CO Bogotá | US California | …>
**Currency:** <COP | USD | …>
**Mode:** <fast | standard | deep>
**Generated:** <ISO-8601 UTC>
## Problem framing
<2–4 sentences. Name the dominant failure mode (Grounding Rule 8). Separate symptom from cause. State the constraint or context the user gave (region, urgency, budget headroom, aesthetic prefs).>
## Alternatives
### Alternative A — <name> (Tiers: T1 → T3)
**Thesis.** <One sentence: what problem this actually solves. Not what the alternative *is*.>
**Cost band (in user currency).**
- Low: <number>
- Typical: <number>
- High: <number>
- Includes / excludes: <tax treatment, installation, materials breakdown if material>
**Confidence:** <0.0–1.0>. <One-line rationale: e.g., "Three exact product pages cited; unit fully specified.">
**Providers cited (N):**
| Tier | Provider | Price | Unit | Confidence | Source |
|---|---|---|---|---|---|
| T1 | Homecenter | $ 12.450 | unidad | 0.95 | [1] |
| T2 | Pavco | $ 14.200 | unidad | 0.88 | [2] |
| T3 | Tecnoglass | cotización a solicitud | m² | 0.55 | [3] |
**Notes.** <Lead time, hidden costs, deal-breakers, sanity-band flags, what's NOT in the band.>
---
### Alternative B — <name> (Tiers: T2 → T4)
… same structure …
### Alternative C — <name> (Tiers: T4 → T5)
… same structure …
(3–5 alternatives total)
## Cross-cutting notes
- **Tax treatment.** <e.g., "All prices in this report exclude IVA 19% unless noted; CO buyer should add IVA + applicable retenciones.">
- **Supplier shortlist by region.** <e.g., "Bogotá: Homecenter Av 68 / Tecnoglass distrito sur / Vitelsa Calle 80.">
- **Common hidden costs.** <e.g., "Casement window install includes demolition + waste removal + interior re-finishing — budget ~20% above the bare DVH cost.">
- **Lead times.** <e.g., "Standard DVH: 3–4 weeks. Acoustic DVH with laminated PVB: 6–8 weeks.">
## Recommendation
**Start with:** <Alternative X>
**Total budget envelope:** <low – high in user currency>
**Rationale (2–3 sentences).** <Why this alternative resolves the dominant failure mode at the best cost-per-problem-solved ratio; what's the risk if it doesn't.>
**If that doesn't work:** <Alternative Y as fallback> — <one-line condition that would trigger escalation: "if Tier-1 fix improves perceived noise by < 50% after 2 weeks, escalate to Alternative B">.
## Sources
[1] https://www.homecenter.com.co/… — "Felpa siliconada para ventana corrediza 5m" — fetched 2026-05-13T14:22Z
[2] https://www.pavco.com.co/… — "Pavco Felpa con Aleta" — fetched 2026-05-13T14:23Z
[3] https://www.tecnoglass.com/contacto — "Cotización ventanería acústica — solicitud por correo" — fetched 2026-05-13T14:25Z
…---
Filled exemplar — window noise (Bogotá)
See ../assets/examples/window-noise-attenuation.md for the worked example with full citations placeholders.
---
Mode-specific tailoring
fast mode
- Same structure, but
Providers citedhas 1–2 rows. Cross-cutting notescan be 1–2 bullets.Recommendation.Rationaleis one sentence.- Total length 200–500 words.
standard mode
- The full template above.
- 3–5 alternatives, each with 3–5 providers cited.
- Cross-cutting notes covers tax, supplier shortlist, hidden costs, lead times.
- Recommendation includes fallback.
- 600–1,200 words.
deep mode
- Full template + a Scenario sensitivity subsection in Recommendation:
- Worst case (highest cited): budget envelope.
- Expected (median): budget envelope.
- Best case (lowest cited): budget envelope.
- Add a Lead-time map: which alternatives can ship in 2 weeks vs. 8+.
- Add Risk register: 3–5 risks with mitigations.
- 1,500–3,000 words.
---
Failure modes the template prevents
| If the template were missing… | The failure mode |
|---|---|
| Problem framing | Report jumps to alternatives without diagnosing — user spends Tier-4 money on a Tier-1 problem. |
| Thesis per alternative | Alternatives become a feature list, not a decision aid. |
| Confidence column | User can't tell which numbers are firm and which are guesses. |
| Sources block | Numbers are uncheckable; report is hallucination-shaped. |
| Recommendation with budget envelope | User reads, doesn't know what to do, asks follow-up. The skill failed to be decision-shaped. |
| Cross-cutting notes | Tax treatment ambiguous, user under-budgets by 19%+. |
---
Validator contract
scripts/validate_report.py <report.md> checks (exit non-zero on failure):
1. Headers present: Problem framing, Alternatives, Recommendation, Sources. 2. ≥ 3 alternatives (sections under ## Alternatives matching ### Alternative prefix). 3. Each alternative has Thesis, Cost band, Confidence, Providers cited. 4. Each cost band has 3 numbers (low / typical / high) in consistent currency. 5. Each confidence value is in [0, 1]. 6. Every footnote [N] referenced in the body appears in ## Sources. 7. Every source has URL, title, fetched-at timestamp. 8. Recommendation block has Start with, Total budget envelope, Rationale.
---
See also: decomposition-patterns.md for choosing what the alternatives are, provider-taxonomy.md for choosing who cites them, grounding-discipline.md for the citation rules.
#!/usr/bin/env python3
"""
validate_report.py — structural linter for a procurer-skill report.
Usage:
python3 validate_report.py <report.md>
Checks (exit 1 on any failure):
1. Required top-level headers present: Problem framing, Alternatives,
Recommendation, Sources.
2. >= 3 alternatives (sections matching '### Alternative ' prefix).
3. Each alternative has Thesis, Cost band, Confidence, Providers cited.
4. Each cost band block contains Low / Typical / High numbers in a
consistent currency prefix (e.g. '$', 'COP', 'USD').
5. Each confidence value parses as a float in [0, 1].
6. Every footnote [N] referenced in the body has a matching entry in
'## Sources'.
7. Every source entry has a URL, a title, and a 'fetched' timestamp.
8. Recommendation block has 'Start with', 'Total budget envelope',
'Rationale'.
Exit codes:
0 — all checks pass.
1 — one or more structural failures.
2 — usage error (no file, file unreadable).
This validator does NOT verify price accuracy, confidence calibration,
or that the recommendation is sound. Those are the agent's responsibility.
"""
from __future__ import annotations
import re
import sys
from pathlib import Path
REQUIRED_TOP_HEADERS = [
"Problem framing",
"Alternatives",
"Recommendation",
"Sources",
]
ALTERNATIVE_REQUIRED_FIELDS = [
"Thesis",
"Cost band",
"Confidence",
"Providers cited",
]
RECOMMENDATION_REQUIRED_FIELDS = [
"Start with",
"Total budget envelope",
"Rationale",
]
MIN_ALTERNATIVES = 3
def parse_sections(text: str) -> dict[str, str]:
"""Split a markdown doc into top-level `## ` sections.
Returns a dict of `header_text -> body_text` for each `## Header` block.
"""
sections: dict[str, str] = {}
current_header: str | None = None
current_body: list[str] = []
for line in text.splitlines():
match = re.match(r"^##\s+(?!#)(.+?)\s*$", line)
if match:
if current_header is not None:
sections[current_header] = "\n".join(current_body).strip()
current_header = match.group(1).strip()
current_body = []
else:
if current_header is not None:
current_body.append(line)
if current_header is not None:
sections[current_header] = "\n".join(current_body).strip()
return sections
def parse_alternatives(alternatives_body: str) -> list[tuple[str, str]]:
"""Split the body of `## Alternatives` into individual `### Alternative …` blocks.
Returns a list of `(name, body)` tuples.
"""
blocks: list[tuple[str, str]] = []
current_name: str | None = None
current_body: list[str] = []
for line in alternatives_body.splitlines():
match = re.match(r"^###\s+(Alternative\s+.+?)\s*$", line)
if match:
if current_name is not None:
blocks.append((current_name, "\n".join(current_body).strip()))
current_name = match.group(1).strip()
current_body = []
else:
if current_name is not None:
current_body.append(line)
if current_name is not None:
blocks.append((current_name, "\n".join(current_body).strip()))
return blocks
def find_confidence_values(body: str) -> list[float]:
"""Extract confidence values from an alternative body.
Looks for patterns like `Confidence: 0.85` or `**Confidence:** 0.85`.
Returns parsed float values.
"""
values: list[float] = []
for match in re.finditer(
r"Confidence[:\*\s]+([01](?:\.\d+)?)",
body,
flags=re.IGNORECASE,
):
try:
values.append(float(match.group(1)))
except ValueError:
continue
return values
def find_cost_band_numbers(body: str) -> dict[str, list[str]]:
"""Find Low / Typical / High lines in a cost-band block.
Returns a dict like {"Low": ["$ 80.000"], "Typical": [...], "High": [...]}.
Each value is a list of raw matched strings (the prefix + number).
"""
results: dict[str, list[str]] = {"Low": [], "Typical": [], "High": []}
for label in results:
for match in re.finditer(
rf"{label}[:\s]+([^\n]+)",
body,
flags=re.IGNORECASE,
):
results[label].append(match.group(1).strip())
return results
def find_footnotes_referenced(text: str) -> set[int]:
"""Extract all `[N]` footnote references from the body."""
refs: set[int] = set()
for match in re.finditer(r"\[(\d+)\]", text):
refs.add(int(match.group(1)))
return refs
def find_sources_defined(sources_body: str) -> dict[int, str]:
"""Parse the `## Sources` body into a dict of `id -> source_line`.
Expects lines like:
[1] https://example.com — "Title" — fetched 2026-05-13T14:22Z
"""
sources: dict[int, str] = {}
for line in sources_body.splitlines():
match = re.match(r"^\s*\[(\d+)\]\s+(.+)$", line)
if match:
sources[int(match.group(1))] = match.group(2).strip()
return sources
def source_line_has_url_title_fetched(line: str) -> tuple[bool, list[str]]:
"""Verify a source line has a URL, title text, and a fetched timestamp."""
missing: list[str] = []
if not re.search(r"https?://\S+", line):
missing.append("URL")
# Title: anything between em-dashes or quoted, after URL
if not re.search(r"[—\-]\s*\S+", line) and not re.search(r'"[^"]+"', line):
missing.append("title")
if not re.search(r"fetched\s+\S+", line, flags=re.IGNORECASE):
missing.append("fetched-at timestamp")
return (len(missing) == 0, missing)
def detect_currency_prefix(s: str) -> str | None:
"""Identify a currency prefix from a price string (e.g. '$', 'COP', 'USD')."""
match = re.search(r"(\$|COP|USD|EUR|GBP|MXN|CLP)", s)
if match:
return match.group(1)
return None
def validate(report_path: Path) -> list[str]:
"""Validate a procurer report. Returns a list of error messages (empty if OK)."""
errors: list[str] = []
try:
text = report_path.read_text(encoding="utf-8")
except OSError as exc:
return [f"Cannot read file: {exc}"]
sections = parse_sections(text)
# Check 1: required top-level headers
for header in REQUIRED_TOP_HEADERS:
if header not in sections:
errors.append(f"Missing required top-level header: ## {header}")
# If Alternatives section is missing, downstream checks are moot
if "Alternatives" not in sections:
return errors
# Check 2: >= MIN_ALTERNATIVES
alternatives = parse_alternatives(sections["Alternatives"])
if len(alternatives) < MIN_ALTERNATIVES:
errors.append(
f"Found {len(alternatives)} alternative(s); need >= {MIN_ALTERNATIVES}."
)
# Check 3 + 4 + 5: per-alternative structure
all_currencies: set[str] = set()
for name, body in alternatives:
for field in ALTERNATIVE_REQUIRED_FIELDS:
if field.lower() not in body.lower():
errors.append(f"[{name}] missing field: {field}")
# Cost band — needs Low/Typical/High
bands = find_cost_band_numbers(body)
for label, matches in bands.items():
if not matches:
errors.append(f"[{name}] cost band missing '{label}'.")
else:
for raw in matches:
cur = detect_currency_prefix(raw)
if cur:
all_currencies.add(cur)
# Confidence — must parse as float in [0,1]
confs = find_confidence_values(body)
if not confs:
errors.append(f"[{name}] no parseable Confidence value found.")
else:
for c in confs:
if not (0.0 <= c <= 1.0):
errors.append(
f"[{name}] confidence {c} outside [0, 1]."
)
# Check 4 (continued): currency consistency
if len(all_currencies) > 1:
errors.append(
f"Inconsistent currency across alternatives: found {sorted(all_currencies)}."
)
# Check 6: footnotes referenced match sources defined
refs = find_footnotes_referenced(sections["Alternatives"])
sources = find_sources_defined(sections.get("Sources", ""))
undefined = refs - set(sources.keys())
if undefined:
errors.append(
f"Footnote(s) referenced but not in ## Sources: {sorted(undefined)}"
)
# Check 7: each source has URL + title + fetched-at
for sid, line in sources.items():
ok, missing = source_line_has_url_title_fetched(line)
if not ok:
errors.append(
f"Source [{sid}] missing: {', '.join(missing)} — line: {line!r}"
)
# Check 8: recommendation block fields
rec_body = sections.get("Recommendation", "")
for field in RECOMMENDATION_REQUIRED_FIELDS:
if field.lower() not in rec_body.lower():
errors.append(f"## Recommendation missing field: {field}")
return errors
def main(argv: list[str]) -> int:
if len(argv) != 2:
print("Usage: validate_report.py <report.md>", file=sys.stderr)
return 2
report_path = Path(argv[1])
if not report_path.exists():
print(f"File not found: {report_path}", file=sys.stderr)
return 2
errors = validate(report_path)
if not errors:
print(f"OK — {report_path} passes all procurer-report structural checks.")
return 0
print(f"FAIL — {report_path} has {len(errors)} structural issue(s):\n")
for err in errors:
print(f" - {err}")
return 1
if __name__ == "__main__":
sys.exit(main(sys.argv))
Security Policy
Supported Versions
| Version | Supported |
|---|---|
0.1.x (current) | Yes |
| Older | No |
This is a pre-1.0 skill. The latest tag on main is always the supported line.
Reporting a Vulnerability
If you discover a security issue, please do not open a public GitHub issue.
Email <contact@broomva.tech> with:
- A description of the issue.
- Reproduction steps if applicable.
- The version (tag) you observed it on.
- Your suggested severity (low / medium / high / critical).
You should receive an initial acknowledgment within 72 hours. We will work with you on a coordinated disclosure timeline.
Threat model — what counts as a security issue
The procurer skill itself does not handle credentials, network sockets, persistent state, or untrusted code execution. Realistic security concerns are narrow:
- Validator bypass:
scripts/validate_report.pydeclaring a malformed report as valid (e.g., accepting non-numeric prices, acceptingconfidence: 2.0, accepting unresolved footnotes). This is the most likely real issue and we treat it seriously because users downstream rely on the validator to gate procurement decisions. - Path traversal in the validator: the script reads a single file argument; if a future change introduces directory traversal or arbitrary file reads, that's a bug we want to know about.
- Prompt injection through example content: the worked examples in
assets/examples/are read into agent context. If a maliciously-crafted example could exfiltrate data or override safety policy from a downstream user's workspace, that's a concern.
Out of scope (please don't report):
- Hallucinated prices in agent output. The grounding discipline is recommended by the skill; the skill cannot prevent an underlying LLM from violating it. Hallucinated numbers are a research-quality issue, not a security issue.
- Outdated supplier shortlists in
assets/examples/construction-materials-co.md. Suppliers go in and out of business; we accept staleness and welcome PRs. - Validator strictness disagreements (e.g., "I think confidence should be allowed at 1.5"). Open a regular issue.
Updates
When a security issue is fixed, we publish:
1. A patch release with a bumped version in CHANGELOG.md. 2. A GitHub Security Advisory at <https://github.com/broomva/procurer/security/advisories>.
Subscribe to repo releases to be notified.
"""Pytest configuration for procurer validator tests."""
from __future__ import annotations
import sys
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parent.parent
SCRIPTS_DIR = REPO_ROOT / "scripts"
# Make scripts/ importable so tests can import validate_report directly
sys.path.insert(0, str(SCRIPTS_DIR))
pytest>=8.0
"""Validator must accept canonical worked examples and reject obvious violations."""
from __future__ import annotations
from pathlib import Path
import validate_report
REPO_ROOT = Path(__file__).resolve().parent.parent
EXAMPLES = REPO_ROOT / "assets" / "examples"
def test_window_noise_example_passes() -> None:
"""The canonical window-noise worked example must validate clean."""
errors = validate_report.validate(EXAMPLES / "window-noise-attenuation.md")
assert errors == [], f"Canonical example failed validation:\n - " + "\n - ".join(errors)
def test_validator_exits_zero_on_canonical_example(tmp_path, capsys) -> None:
"""CLI entry point exits 0 when the canonical example is passed."""
rc = validate_report.main(["validate_report.py", str(EXAMPLES / "window-noise-attenuation.md")])
assert rc == 0
out = capsys.readouterr().out
assert "OK" in out
def test_validator_exits_two_on_missing_file(capsys) -> None:
"""Missing file returns exit code 2 (usage error)."""
rc = validate_report.main(["validate_report.py", "/nonexistent/path.md"])
assert rc == 2
"""Validator must reject structural violations and explain what failed."""
from __future__ import annotations
from pathlib import Path
import pytest
import validate_report
REPO_ROOT = Path(__file__).resolve().parent.parent
FIXTURES = REPO_ROOT / "tests" / "fixtures"
def _validate_string(content: str, tmp_path: Path) -> list[str]:
"""Helper: write content to tmp_path and run the validator on it."""
report = tmp_path / "report.md"
report.write_text(content, encoding="utf-8")
return validate_report.validate(report)
def test_missing_top_level_headers_fails(tmp_path: Path) -> None:
content = "# A need\n\nNo sections at all.\n"
errors = _validate_string(content, tmp_path)
# Should flag all four missing top-level headers
assert any("Problem framing" in e for e in errors)
assert any("Alternatives" in e for e in errors)
assert any("Recommendation" in e for e in errors)
assert any("Sources" in e for e in errors)
def test_fewer_than_three_alternatives_fails(tmp_path: Path) -> None:
content = (
"# A need\n\n"
"## Problem framing\nSomething.\n\n"
"## Alternatives\n\n"
"### Alternative A — first\n"
"**Thesis.** x.\n**Cost band.**\n- Low: $ 100\n- Typical: $ 200\n- High: $ 300\n"
"**Confidence:** 0.8\n**Providers cited:** —\n\n"
"### Alternative B — second\n"
"**Thesis.** y.\n**Cost band.**\n- Low: $ 100\n- Typical: $ 200\n- High: $ 300\n"
"**Confidence:** 0.8\n**Providers cited:** —\n\n"
"## Recommendation\n**Start with:** A\n**Total budget envelope:** $100 – $200\n**Rationale:** because.\n\n"
"## Sources\n[1] https://x.com — \"Title\" — fetched 2026-05-13T00:00Z\n"
)
errors = _validate_string(content, tmp_path)
assert any("alternative" in e.lower() and ">=" in e for e in errors), (
f"Expected '< MIN_ALTERNATIVES' failure; got: {errors}"
)
def test_confidence_out_of_range_fails(tmp_path: Path) -> None:
content = _three_alternatives_template().replace(
"**Confidence:** 0.85", "**Confidence:** 1.5", 1
)
errors = _validate_string(content, tmp_path)
assert any("outside [0, 1]" in e for e in errors), (
f"Expected out-of-range confidence failure; got: {errors}"
)
def test_missing_confidence_fails(tmp_path: Path) -> None:
content = _three_alternatives_template().replace(
"**Confidence:** 0.85", "Confidence-not-here", 1
)
errors = _validate_string(content, tmp_path)
# The "missing field: Confidence" check should fire
assert any("Confidence" in e for e in errors)
def test_missing_cost_band_low_fails(tmp_path: Path) -> None:
content = _three_alternatives_template().replace(
"- Low: $ 100", "- LOW-MISSING: $ 100", 1
)
errors = _validate_string(content, tmp_path)
assert any("'Low'" in e or "Low" in e for e in errors)
def test_unresolved_footnote_fails(tmp_path: Path) -> None:
"""A footnote referenced in the body but not in ## Sources is flagged."""
# Template uses [1]; replace one occurrence with [99] (unresolved) so we get
# a body reference that has no matching ## Sources entry.
content = _three_alternatives_template().replace(
"**Providers cited:** see [1]", "**Providers cited:** see [99]", 1
)
errors = _validate_string(content, tmp_path)
assert any("[99]" in e or "99" in e for e in errors), (
f"Expected unresolved-footnote failure; got: {errors}"
)
def test_source_missing_url_fails(tmp_path: Path) -> None:
content = _three_alternatives_template().replace(
"https://x.com — \"Title\" — fetched 2026-05-13T00:00Z",
"no-url-here — \"Title\" — fetched 2026-05-13T00:00Z",
1,
)
errors = _validate_string(content, tmp_path)
assert any("URL" in e for e in errors), (
f"Expected source-missing-URL failure; got: {errors}"
)
def test_source_missing_fetched_at_fails(tmp_path: Path) -> None:
content = _three_alternatives_template().replace(
"fetched 2026-05-13T00:00Z", "no-timestamp", 1
)
errors = _validate_string(content, tmp_path)
assert any("fetched" in e.lower() for e in errors)
def test_inconsistent_currency_fails(tmp_path: Path) -> None:
"""Mixing $ and EUR across alternatives should flag a currency-consistency error."""
base = _three_alternatives_template()
# Replace one alternative's currency to EUR
content = base.replace("Low: $ 100", "Low: EUR 100", 1)
content = content.replace("Typical: $ 200", "Typical: EUR 200", 1)
content = content.replace("High: $ 300", "High: EUR 300", 1)
errors = _validate_string(content, tmp_path)
assert any("currency" in e.lower() for e in errors), (
f"Expected currency-consistency failure; got: {errors}"
)
def test_missing_recommendation_fields_fails(tmp_path: Path) -> None:
content = _three_alternatives_template().replace(
"**Total budget envelope:** $100 – $200\n", "", 1
)
errors = _validate_string(content, tmp_path)
assert any("Total budget envelope" in e for e in errors)
# ----------------------------------------------------------------------
# Template helper
# ----------------------------------------------------------------------
def _three_alternatives_template() -> str:
"""A minimal but valid 3-alternative report. Tests mutate this to exercise failures."""
return (
"# A need\n\n"
"## Problem framing\nSomething.\n\n"
"## Alternatives\n\n"
"### Alternative A — first\n"
"**Thesis.** x.\n"
"**Cost band.**\n- Low: $ 100\n- Typical: $ 200\n- High: $ 300\n"
"**Confidence:** 0.85\n"
"**Providers cited:** see [1]\n\n"
"### Alternative B — second\n"
"**Thesis.** y.\n"
"**Cost band.**\n- Low: $ 100\n- Typical: $ 200\n- High: $ 300\n"
"**Confidence:** 0.85\n"
"**Providers cited:** see [1]\n\n"
"### Alternative C — third\n"
"**Thesis.** z.\n"
"**Cost band.**\n- Low: $ 100\n- Typical: $ 200\n- High: $ 300\n"
"**Confidence:** 0.85\n"
"**Providers cited:** see [1]\n\n"
"## Recommendation\n"
"**Start with:** A\n"
"**Total budget envelope:** $100 – $200\n"
"**Rationale:** because.\n\n"
"## Sources\n"
"[1] https://x.com — \"Title\" — fetched 2026-05-13T00:00Z\n"
)
def test_template_helper_itself_validates_clean(tmp_path: Path) -> None:
"""Sanity check — the template helper must produce a passing report."""
errors = _validate_string(_three_alternatives_template(), tmp_path)
assert errors == [], f"Template helper unexpectedly failed validation: {errors}"
Related skills
FAQ
How many alternatives does procurer produce?
It produces 3-5 ranked alternatives, each placed on a provider-archetype tier from DIY-retail up to consultant/turnkey, with low/typical/high price bands.
When should I use deep-research instead?
Use deep-research for pure knowledge research with no decision attached; procurer fires only when there is a budget envelope at the end.