Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
pproenca avatar

Marketplace Recsys Feature Engineering

  • 145 installs
  • 191 repo stars
  • Updated July 24, 2026
  • pproenca/dot-skills

marketplace-recsys-feature-engineering: A skill for development. This provides functionality for development workflows.

Key points

  • marketplace-recsys-feature-engineering

Marketplace Recsys Feature Engineering by the numbers

  • 145 all-time installs (skills.sh)
  • +6 installs in the week ending Aug 4, 2026 (Skillselion tracking)
  • Ranked #2,587 of 4,347 Backend & APIs skills by installs in the Skillselion catalog
  • Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/pproenca/dot-skills --skill marketplace-recsys-feature-engineering

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs145
repo stars191
Last updatedJuly 24, 2026
Repositorypproenca/dot-skills

How do I use marketplace-recsys-feature-engineering for development tasks?

Use marketplace-recsys-feature-engineering for development tasks

Who is it for?

Best when you're working on backend & apis and need structured help with marketplace-recsys-feature-engineering.

Skip if: Teams with no backend & apis needs, or anyone wanting a generic chat assistant without this specific workflow.

When should I use this skill?

When you need to use marketplace-recsys-feature-engineering for development tasks, or when marketplace-recsys-feature-engineering: a skill for development. this provides functionality for development workflows.

What you get

Structured output aligned to marketplace-recsys-feature-engineering: marketplace-recsys-feature-engineering.

Files

SKILL.mdMarkdownGitHub ↗

Marketplace Engineering Recsys Feature Engineering Best Practices

Comprehensive first-principles guide for deriving usable recommender features from the raw assets of a two-sided trust marketplace — listing photos, owner-supplied listing metadata, and sitter wizard responses — for item-to-item, user-to-item, and user-to-user solutions. Contains 44 rules across 8 categories ordered by cascade impact on the feature-engineering lifecycle, plus one playbook that composes the rules into an end-to-end feature discovery workflow.

This skill is the upstream precursor to marketplace-personalisation (AWS Personalize) and marketplace-search-recsys-planning (OpenSearch retrieval). Those skills treat features as inputs they already have; this skill is about deciding what features to build from the raw assets, which decisions they serve, and how to prove each one is worth its maintenance cost.

When to Apply

Reference this skill when:

  • Planning what to extract from listing photos, descriptions, or amenity lists to power i2i similarity or u2i ranking
  • Designing or revising the sitter onboarding wizard with recsys features as the primary output
  • Deciding whether to build a vision embedding pipeline, a text encoder, or neither — and in what order
  • Composing existing base features into item-to-item, user-to-item, or user-to-user scoring
  • Auditing an existing feature store for coverage, drift, PII, duplication, or orphan features
  • Choosing a ship/kill criterion for a new recsys feature and designing the ablation A/B test
  • Answering the question: "we want to improve the similar-homes shelf — what feature should we build?"

Setup

This skill has no user-specific configuration — it is self-contained. References are live URLs to engineering blogs from Airbnb, Pinterest, DoorDash, Uber, Netflix, and Google, to open-source libraries (Feast, Sentence-Transformers, Hugging Face CLIP, H3), to foundational academic papers (Airbnb KDD 2018, Pinterest ItemSage, YouTube Semantic IDs, PinSage), and to Google's Rules of Machine Learning.

Rule Categories

Categories are ordered by cascade impact on the feature-engineering lifecycle: auditing mistakes build features on data that does not exist, first-principles mistakes produce features that do not map to real decisions, extraction mistakes poison everything downstream, and so on. Fix earlier-stage problems before later-stage problems.

#CategoryPrefixImpact
1Asset Audit and Inventoryaudit-CRITICAL
2First-Principles Feature Decompositionfirstp-CRITICAL
3Image Feature Extractionvision-HIGH
4Listing Text and Metadata Extractionlisting-HIGH
5Sitter Wizard and Profile Extractionwizard-HIGH
6Derived Similarity and Affinityderive-MEDIUM-HIGH
7Feature Quality and Governancequality-MEDIUM-HIGH
8Incremental Rollout and Value Proofprove-MEDIUM

Quick Reference

1. Asset Audit and Inventory (CRITICAL)

  • `audit-measure-coverage-before-modelling` — reject fields below 80% coverage from the feature plan
  • `audit-sample-every-asset-type-end-to-end` — pull 100 real instances through the real fetch path before planning
  • `audit-verify-rights-and-privacy-before-extraction` — ToS, GDPR, consent, face blur before encoding
  • `audit-quantify-freshness-per-asset` — age distribution + expiry + refresh bucket
  • `audit-separate-raw-assets-from-derived-features` — raw immutable in object store, derived versioned in feature store

2. First-Principles Feature Decomposition (CRITICAL)

  • `firstp-start-from-the-decision-not-the-algorithm` — decision first, sub-judgments second, tools last
  • `firstp-ask-what-signal-a-human-uses` — interview 8-12 owners and sitters; features trace back to quotes
  • `firstp-tie-every-feature-to-a-specific-solution` — no feature without a named i2i/u2i/u2u consumer
  • `firstp-prefer-directly-observed-over-learned` — observed columns first, learned embeddings second
  • `firstp-reject-features-you-cannot-serve-at-inference` — training-serving parity starts at design time
  • `firstp-kill-features-a-popularity-baseline-already-captures` — correlation screen before registration

3. Image Feature Extraction (HIGH)

  • `vision-use-clip-for-zero-shot-listing-embeddings` — zero-shot CLIP ships in a week
  • `vision-detect-room-types-before-detecting-amenities` — room prior conditions the amenity threshold
  • `vision-quantify-image-quality-separately-from-content` — blur, lighting, aesthetic as their own features
  • `vision-extract-per-object-counts-not-just-presence`n_bed = 4 beats has_bed = true
  • `vision-pool-embeddings-across-a-listings-photo-set` — pooled listing vector; per-photo stored alongside
  • `vision-fine-tune-on-your-domain-when-clip-underperforms` — contrastive fine-tune only after zero-shot plateaus

4. Listing Text and Metadata Extraction (HIGH)

  • `listing-declare-categorical-fields-for-bounded-vocabularies` — bounded vocab → categorical, validated on write
  • `listing-multi-hot-encode-amenity-lists` — fixed amenity vocabulary → multi-hot vector
  • `listing-hash-geo-to-hierarchies-not-raw-lat-lon` — H3 at multiple resolutions
  • `listing-embed-description-with-pretrained-sentence-encoder` — all-MiniLM-L6-v2 for cheap semantic text features
  • `listing-extract-stay-duration-shape-not-just-length` — bin + holiday overlap + flexibility, not raw day count
  • `listing-encode-pet-requirements-as-structured-triples`(species, count, special_needs) triples plus free text alongside

5. Sitter Wizard and Profile Extraction (HIGH)

  • `wizard-order-questions-by-information-gain` — discriminative questions first, narrative last
  • `wizard-prefer-multiple-choice-over-free-text` — categorical features by construction
  • `wizard-make-skips-genuine-and-log-them` — skip is signal; defaults destroy it
  • `wizard-capture-experience-as-counts-and-dates` — numbers, not adjectives; platform history overrides self-declaration
  • `wizard-separate-hard-constraints-from-soft-preferences` — filters vs ranking features

6. Derived Similarity and Affinity (MEDIUM-HIGH)

  • `derive-precompute-i2i-nearest-neighbours-offline` — ANN shelf built nightly, served from KV in <5ms
  • `derive-fuse-modalities-before-item-similarity` — vision + text + structured, weighted and normalised
  • `derive-use-two-tower-for-user-item-affinity` — dual encoder trained on interactions; ANN-retrieval-ready
  • `derive-score-u2u-as-symmetric-mutual-fit`min(P(owner), P(sitter)); one-sided scoring produces wasted requests
  • `derive-decompose-affinity-into-interpretable-subscores` — fit/safety/logistics/price subscores + blend
  • `derive-cache-user-embedding-with-short-ttl` — session-level cache, 60-300s TTL

7. Feature Quality and Governance (MEDIUM-HIGH)

  • `quality-version-feature-definitions-in-one-registry` — one name, one implementation, one owner
  • `quality-serve-training-and-inference-from-one-store` — feature store as the single source of truth
  • `quality-gate-features-on-coverage-and-drift` — coverage floor + PSI alarm
  • `quality-scrub-pii-before-features-leave-secure-zone` — face blur and regex scrubbing before encoding
  • `quality-freeze-feature-schemas-per-model-version` — schema hash pinned to model artifact

8. Incremental Rollout and Value Proof (MEDIUM)

  • `prove-ship-one-feature-at-a-time` — one feature, one experiment, one decision
  • `prove-measure-lift-against-feature-ablated-variant` — ablation isolates the feature from incidental changes
  • `prove-kill-features-that-dont-earn-maintenance` — quarterly kill review on attributed lift
  • `prove-dedicate-random-exploration-slice-to-new-features` — 3-5% slice catches offline-close-to-tied winners
  • `prove-retain-feature-free-baseline-permanently` — popularity baseline as drift anchor

Discovering New Features

One playbook composes the rules into an end-to-end workflow:

  • `references/playbooks/discovering.md` — Discover new features from raw marketplace assets: a seven-step workflow that starts with an asset audit and a decision decomposition and ends with a shipped ablation A/B against a feature-ablated baseline. Use when the task is "what should we build next?" rather than "fix this specific feature."

Read the playbook first when the task is an open-ended "how do we extract more signal from X?" Read individual rules when a specific implementation question arises.

How to Use

  • Read `references/_sections.md` for category structure and cascade rationale
  • Read `gotchas.md` for accumulated diagnostic lessons before suggesting interventions
  • Read `references/playbooks/discovering.md` to plan a new feature discovery cycle
  • Read individual rule files under references/ when a specific task matches the rule title
  • Use `assets/templates/_template.md` to author new rules as the skill grows

Related Skills

  • `marketplace-personalisation` — Post-extraction personalisation on AWS Personalize: event tracking, schema design, two-sided matching, cold start, feedback loops. Hand off once your features are in the store and you are ready to train a ranker.
  • `marketplace-search-recsys-planning` — OpenSearch retrieval planning: query understanding, index design, ranking, search-plus-recs blending. Hand off when the bottleneck is retrieval rather than feature availability.
  • `marketplace-pre-member-personalisation` — Pre-member journey from anonymous visit to paid membership: anonymous signal inference, onboarding intent capture, pre-member measurement. Hand off at the paid-member boundary.

Reference Files

FileDescription
references/_sections.mdCategory definitions, impact ordering, cascade rationale
references/playbooks/discovering.mdEnd-to-end feature discovery playbook
gotchas.mdAccumulated feature-engineering diagnostic lessons (living)
assets/templates/_template.mdTemplate for authoring new rules
metadata.jsonVersion, discipline, authoritative references

Related skills

FAQ

What does marketplace-recsys-feature-engineering do?

marketplace-recsys-feature-engineering: A skill for development. This provides functionality for development workflows.

When should I use marketplace-recsys-feature-engineering?

When you need to use marketplace-recsys-feature-engineering for development tasks, or when marketplace-recsys-feature-engineering: a skill for development. this provides functionality for development workflows.

What are the main capabilities?

marketplace-recsys-feature-engineering.

Backend & APIsbackendintegrations

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.