Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
borghei avatar

Ab Test Setup

  • 79 installs
  • 451 repo stars
  • Updated July 21, 2026
  • borghei/claude-skills

ab-test-setup is a Claude skill that calculates A/B test sample sizes, designs test plans, and analyzes results for statistical significance.

About

ab-test-setup is a Claude skill that helps design and analyze A/B tests. A developer or growth team uses its Python scripts to calculate required sample size, generate a test plan from a JSON config, and analyze collected results for statistical significance. It supports segment-level analysis and batch review of past experiments to inform ship or no-ship decisions.

  • Calculates required sample size from baseline rate, minimum detectable effect, and statistical power
  • Generates a complete A/B test plan from a JSON config
  • Analyzes results for statistical significance (p-value, confidence interval, effect size)

Ab Test Setup by the numbers

  • 79 all-time installs (skills.sh)
  • Ranked #866 of 2,064 Data Science & ML skills by installs in the Skillselion catalog
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
At a glance

ab-test-setup capabilities & compatibility

Free; local Python scripts, no API keys.

Capabilities
sample size calculator · experiment designer · significance testing
Use cases
data analysis · marketing
Pricing
Free
From the docs

What ab-test-setup says it does

toolkit for calculating sample sizes, designing rigorous test plans, and analyzing results with statistical significance testing
SKILL.md
Designed for growth teams, product managers, and marketers who need to make data-driven decisions from controlled experiments.
SKILL.md
Review confidence interval, p-value, and effect size
SKILL.md
npx skills add https://github.com/borghei/claude-skills --skill ab-test-setup

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs79
repo stars451
Last updatedJuly 21, 2026
Repositoryborghei/claude-skills

What it does

Design A/B tests with proper sample sizing and analyze results for statistical significance before shipping a change.

Who is it for?

Growth teams and product managers designing rigorous conversion-rate experiments with proper sample sizing.

Skip if: Implementing the tracking or feature-flag infrastructure that runs the test.

When should I use this skill?

When you need to set up an A/B test, calculate sample size, design an experiment, or analyze A/B test results and significance.

What you get

A statistically grounded test plan and a significance-tested ship/no-ship recommendation.

  • required sample size
  • complete test plan document
  • statistical analysis with recommendation

By the numbers

  • 3 scripts: sample_size_calculator, test_designer, results_analyzer
  • 3 workflows (new test setup, results analysis, program review)

Files

SKILL.mdMarkdownGitHub ↗

A/B Test Setup Skill

Overview

Production-ready A/B testing toolkit for calculating sample sizes, designing rigorous test plans, and analyzing results with statistical significance testing. Designed for growth teams, product managers, and marketers who need to make data-driven decisions from controlled experiments.

Quick Start

# Calculate required sample sizes for a test
python scripts/sample_size_calculator.py --baseline 0.05 --mde 0.10 --power 0.80

# Design a complete A/B test plan
python scripts/test_designer.py test_config.json

# Analyze A/B test results
python scripts/results_analyzer.py results.json

Tools Overview

ToolPurposeInputOutput
sample_size_calculator.pySample size calculationBaseline rate, MDE, powerRequired samples + duration
test_designer.pyTest plan designJSON test configComplete test plan document
results_analyzer.pyResults analysisJSON with test resultsStatistical analysis + recommendation

Workflows

Workflow 1: New A/B Test Setup

1. Define hypothesis and success metric 2. Run sample_size_calculator.py with baseline conversion and minimum detectable effect 3. Create test configuration JSON (see Common Patterns) 4. Run test_designer.py to generate complete test plan 5. Share plan with stakeholders for alignment before launch

Workflow 2: Test Results Analysis

1. Collect test results into JSON format 2. Run results_analyzer.py to get statistical significance 3. Review confidence interval, p-value, and effect size 4. Check for segment-level effects if overall result is inconclusive 5. Make ship/no-ship decision based on analysis

Workflow 3: Experimentation Program Review

1. Compile results from multiple past tests 2. Run results_analyzer.py --batch on all results 3. Review win rate, average effect size, and velocity 4. Identify patterns in winning vs losing tests 5. Optimize test pipeline based on learnings

Reference Documentation

See references/ab-testing-guide.md for comprehensive methodology covering:

  • Statistical foundations (z-tests, confidence intervals)
  • Sample size theory and trade-offs
  • Common experimentation pitfalls
  • Multi-variant and sequential testing
  • Bayesian vs frequentist approaches

Common Patterns

Pattern: Test Configuration JSON

{
  "test_name": "Homepage CTA Button Color",
  "hypothesis": "Changing the CTA button from blue to green will increase click-through rate",
  "metric_primary": "cta_click_rate",
  "metric_secondary": ["signup_rate", "bounce_rate"],
  "baseline_rate": 0.045,
  "minimum_detectable_effect": 0.10,
  "significance_level": 0.05,
  "power": 0.80,
  "variants": [
    {"name": "control", "description": "Current blue CTA button"},
    {"name": "treatment", "description": "Green CTA button"}
  ],
  "daily_traffic": 5000,
  "allocation": {"control": 0.50, "treatment": 0.50}
}

Pattern: Test Results JSON

{
  "test_name": "Homepage CTA Button Color",
  "variants": {
    "control": {"visitors": 12500, "conversions": 563},
    "treatment": {"visitors": 12500, "conversions": 625}
  },
  "metric": "cta_click_rate",
  "significance_level": 0.05
}

Quick Reference: Common Effect Sizes

ContextSmall EffectMedium EffectLarge Effect
Conversion Rate2-5% relative5-15% relative> 15% relative
Revenue per User1-3%3-8%> 8%
Engagement Rate3-5%5-10%> 10%

Related skills

FAQ

What inputs does the sample-size calculator need?

Baseline conversion rate, minimum detectable effect (MDE), and statistical power, e.g. --baseline 0.05 --mde 0.10 --power 0.80.

How does it decide a winner?

results_analyzer.py returns statistical significance, confidence interval, p-value, and effect size to support a ship or no-ship decision.

Data Science & MLcontentlifecycle

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.