Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
affaan-m avatar

Recsys Pipeline Architect

  • 1.4k installs
  • 238k repo stars
  • Updated August 5, 2026
  • affaan-m/ecc

This is a copy of recsys-pipeline-architect by affaan-m - installs and ranking accrue to the original listing.

recsys-pipeline-architect is an agent skill that designs and scaffolds modular six-stage recommendation and ranking pipelines for developers who need top-K item selection for feeds, search rerankers, or notification tria

About

recsys-pipeline-architect in affaan-m/ecc is a spec-and-scaffold skill for composable recommendation, ranking, and feed pipelines encoding a six-stage pattern: Source, Hydrator, Filter, Scorer, Selector, and SideEffect. The framework generalizes beyond social feeds to any system picking the top K items for a user-context pair—content CMS ranking, RAG rerankers, task prioritizers, notification triage, search reranking, and ad ranking. Each stage has a defined responsibility: sourcing candidates, hydrating features, filtering ineligible items, scoring relevance, selecting the final set, and executing side effects like logging or writes. The pattern draws from xAI's open-sourced For You algorithm structure. Developers reach for recsys-pipeline-architect when greenfield ranking services need clear stage boundaries instead of monolithic recommenders. Output includes pipeline specs and scaffold code organized by stage interfaces. Triggers include feed ranking architecture, recommendation pipeline design, top-K selector patterns, and modular reranker scaffolding.

  • Encodes the exact six-stage Source→Hydrator→Filter→Scorer→Selector→SideEffect pattern from xAI's For You algorithm
  • Generates complete runnable pipeline scaffolds in TypeScript, Go, or Python
  • Helps migrate from single-score relevance to multi-action prediction with tunable weights
  • Produces composable, production-ready architecture for any top-K selection system
  • Independent MIT reimplementation with no copied code from upstream

Recsys Pipeline Architect by the numbers

  • 1,360 all-time installs (skills.sh)
  • +82 installs in the week ending Aug 5, 2026 (Skillselion tracking)
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/affaan-m/ecc --skill recsys-pipeline-architect

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs1.4k
repo stars238k
Last updatedAugust 5, 2026
Repositoryaffaan-m/ecc

How do you architect a modular recommendation pipeline?

Quickly design and scaffold modular recommendation, ranking, and personalized feed systems.

Who is it for?

Backend and ML engineers building feed, search reranking, or notification systems who need a composable six-stage top-K pipeline architecture.

Skip if: Simple static sort lists, one-off SQL ORDER BY queries, or teams with an existing monolithic recommender and no pipeline refactor planned.

When should I use this skill?

User asks to design a recommendation pipeline, scaffold a feed ranker, build a six-stage Source-Hydrator-Filter-Scorer-Selector system, or architect top-K item selection.

What you get

Six-stage pipeline specification and scaffold code with stage interfaces for sourcing, hydration, filtering, scoring, selection, and side effects.

  • Pipeline stage specification
  • Scaffold code per stage
  • Top-K selection interface definitions

By the numbers

  • Defines 6 pipeline stages: Source, Hydrator, Filter, Scorer, Selector, SideEffect

Files

SKILL.mdMarkdownGitHub ↗

recsys-pipeline-architect

A spec-and-scaffold skill for building composable recommendation, ranking, and feed pipelines. It encodes the six-stage pattern — Source → Hydrator → Filter → Scorer → Selector → SideEffect — popularized by xAI's open-sourced For You algorithm (Apache 2.0). This skill is an independent reimplementation of the pattern (MIT) — no code copied from the original.

Upstream: <https://github.com/mturac/recsys-pipeline-architect>

When to Use

  • User wants to build any system that picks "the top K items for a user/context"
  • User asks "how should I rank X" or describes a feed/personalization problem
  • User has a scoring function and needs the pipeline plumbing around it
  • User wants to migrate from a single relevance score to multi-action prediction with tunable weights
  • User is wrapping an LLM/ML scorer and needs filters, hydrators, side-effects, and a runnable scaffold in their stack (TypeScript / Go / Python)
  • Triggers: "recommendation system", "feed algorithm", "ranking pipeline", "for you feed", "candidate pipeline", "content recommender", "pipeline architecture for recsys", "RAG retrieval reranker"

When NOT to Use

  • Model architecture work (transformer design, two-tower retrieval, embedding training) — this skill is plumbing around the model, not the model itself
  • Pure ML training pipelines — the scoring function is the user's responsibility
  • Operating a deployed pipeline (monitoring, autoscaling) — out of scope

The six-stage framework

#StageJobParallel?
1SourceFetch candidates from one or more originsYes — multiple sources run in parallel
2HydratorEnrich each candidate with metadata needed for filtering and scoringYes — independent hydrators run in parallel
3FilterDrop candidates that should never be shown (blocked, expired, duplicate, ineligible)Sequential — each filter sees fewer items
4ScorerAssign each surviving candidate one or more scoresSequential — later scorers see earlier scores
5SelectorSort by final score, return top KSingle op
6SideEffectCache served IDs, log impressions, emit events, update countersAsync — must never block the response

Why this exact order

  • Sources before hydration: know what candidates exist before paying to enrich them
  • Hydration before filtering: many filters need metadata the source did not provide
  • Filtering before scoring: scoring is the expensive stage; drop the ineligible first
  • Scorer chain (not single scorer): real systems compose ML scoring + diversity reranking + business rules
  • Selector after scoring: keeps scoring deterministic and cacheable
  • SideEffects last and async: side effects must never block the user response

Workflow when invoked

Walk the user through these eight steps:

1. Clarify the use case (one round, three questions): items being ranked? input context? language/runtime? 2. Identify the candidate sources: usually in-network (followed/owned/subscribed) + out-of-network (ML retrieval / trending / similar-to-liked) 3. List required hydrations: for each filter and scorer, what data does it need that the source did not provide? 4. List the filters: duplicate, self, age, block/mute, previously-served, eligibility. Order matters — cheap before expensive. 5. Design the scorer chain: primary (ML) → combiner (multi-action with weights) → diversity → business rules 6. Selector: sort descending by final score, take top K (or stratified mix for in-network/out-of-network) 7. SideEffects: cache served IDs, emit impression events, update counters, log analytics — all fire-and-forget 8. Generate the scaffold in the user's stack

Key trade-offs to surface (don't default silently)

1. Single score vs multi-action prediction

  • Single score: train one model to predict relevance. To change behavior → retrain.
  • Multi-action: predict P(action) for many actions (read, like, share, skip, report), combine with weights at serving time. To change behavior → change weights. No retraining.

The X For You system uses multi-action with both positive and negative weights. Recommend multi-action when the user expects to tune frequently.

2. Candidate isolation in scoring

  • Isolated: each candidate scored independently. Deterministic, cacheable.
  • Joint: candidates attend to each other during scoring (e.g., transformer over batch). More expressive but non-deterministic across batches.

Default to isolation. Joint only when there's a specific reason (e.g., explicit batch-aware diversity).

3. Online vs offline

  • Request-time (online): pipeline runs on each request. Latency budget: 100–300ms. Default.
  • Pre-computed (offline batch): pipeline runs periodically, results cached. Lower latency, lower freshness.
  • Hybrid: candidate retrieval offline, ranking online.

Hard rules

1. Do not invent benchmark numbers. "How much faster?" → "depends on workload, run it yourself." 2. Attribution discipline. When the pattern is referenced, attribute as "popularized by xAI's open-sourced For You algorithm" / github.com/xai-org/x-algorithm (Apache 2.0). 3. No trademark use. Do not name the user's artifact "X-like" or use "For You" branding. Pattern is free; brand is not. Suggested naming: "candidate pipeline", "feed pipeline", "ranking pipeline", "recsys pipeline". 4. Surface trade-offs. Multi-action vs single, isolation vs joint, online vs offline — never default silently. 5. The generated scaffold must run. No pseudocode passing as code. 6. Filter order matters. Cheap before expensive. Universal before user-specific. 7. Side effects never block. Wrap in fire-and-forget patterns (goroutines / promises without await / asyncio tasks).

Anti-Patterns

  • Scoring before filtering (wastes compute on candidates that will be dropped anyway)
  • Synchronous side effects (cache writes / impression emits blocking the response)
  • A single "relevance" score when the product needs to tune for multiple objectives (engagement vs safety vs diversity vs ads)
  • Joint scoring as default (non-deterministic, harder to cache, doesn't compose with reranking stages)
  • Generating pseudocode "for illustration" — the scaffold must actually run

Upstream contents

The upstream repository at <https://github.com/mturac/recsys-pipeline-architect> ships:

  • Full SKILL.md with the complete 8-step workflow
  • 5 load-on-demand reference docs: interfaces in 4 languages (TS/Go/Python/Rust), multi-action scoring pattern, candidate isolation, filter cookbook (12 patterns), scorer cookbook (weighted sum, MMR, diversity penalty, position debiasing)
  • 3 runnable example scaffolds, every one green on its test suite:
  • Strapi v5 plugin (TypeScript / Jest — 3/3 pass)
  • Zentra-compatible pipeline (Go with generics — 3/3 pass)
  • PMAI task prioritizer (Python / FastAPI / pytest — 3/3 pass)
  • v0.1.0 release tagged
  • MIT license; pattern attributed to xAI X For You algorithm (Apache 2.0)

Install via skills.sh: npx skills add mturac/recsys-pipeline-architect

Related skills

How it compares

Pick recsys-pipeline-architect for staged feed or reranker architecture; pick model-training skills when the gap is offline ML not serving pipeline design.

FAQ

What six stages does recsys-pipeline-architect define?

recsys-pipeline-architect scaffolds Source, Hydrator, Filter, Scorer, Selector, and SideEffect stages for top-K item selection—covering feeds, RAG rerankers, search, ads, and notification triage pipelines.

When should developers use recsys-pipeline-architect?

recsys-pipeline-architect fits any system picking top K items for a user-context pair—social feeds, CMS ranking, task prioritizers, or search reranking—when modular stage boundaries beat monolithic recommenders.

AI & Agent Buildingagentsautomationllm

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.