
Content Experimentation Best Practices
- 2.5k installs
- 171 repo stars
- Updated July 29, 2026
- sanity-io/agent-toolkit
content-experimentation-best-practices is a Sanity agent-toolkit skill for designing, running, and interpreting CMS and frontend content experiments.
About
The content-experimentation-best-practices skill from Sanity agent-toolkit documents how to run content experiments that improve conversion, engagement, and user experience. Core concepts define A/B testing, multivariate testing, statistical significance, and experimentation culture that favors data over opinion. Reference guides split experiment design, statistical foundations, CMS integration for field-level variants, and common pitfalls across statistics, design, execution, and interpretation. Agents apply it when setting up testing infrastructure, designing hypotheses and success metrics, wiring CMS-managed variants, or interpreting results without falling into peeking or underpowered tests. The skill points to detailed references for sample size, p-values, confidence intervals, power analysis, Bayesian methods, and platform-specific CMS patterns. Use it when planning landing page tests, headline variants, or frontend experimentation workflows where variant delivery and analysis must stay aligned with sound experiment design practices overall.
- Covers A/B and multivariate testing with statistical significance framing.
- References split experiment design, statistics, CMS integration, and pitfalls.
- Targets CMS-managed variants and frontend experimentation workflows.
- Emphasizes data-driven decisions and HiPPO avoidance culture.
- Links detailed guidance on metrics, sample size, and analysis mistakes.
Content Experimentation Best Practices by the numbers
- 2,508 all-time installs (skills.sh)
- +98 installs in the week ending Aug 2, 2026 (Skillselion tracking)
- Ranked #232 of 1,879 Marketing & SEO skills by installs in the Skillselion catalog
- Security screen: LOW risk (skills.sh audit)
- Data as of Aug 2, 2026 (Skillselion catalog sync)
content-experimentation-best-practices capabilities & compatibility
- Capabilities
- a/b and multivariate experiment framing · hypothesis and metrics design references · cms variant integration patterns · statistical foundations pointers · common pitfalls catalog across experiment lifecy
- Use cases
- seo · marketing · copywriting
What content-experimentation-best-practices says it does
Comparing two variants (A vs B) to determine which performs better.
Making decisions based on data rather than opinions (HiPPO avoidance).
npx skills add https://github.com/sanity-io/agent-toolkit --skill content-experimentation-best-practicesAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 2.5k |
|---|---|
| repo stars | ★ 171 |
| Security audit | 3 / 3 scanners passed |
| Last updated | July 29, 2026 |
| Repository | sanity-io/agent-toolkit ↗ |
How do I set up a content A/B test with sound metrics, variants, and statistical interpretation in a CMS stack?
Plan CMS-backed A/B and multivariate content experiments with hypothesis design, metrics, sample size, and statistical interpretation guidance.
Who is it for?
Planning CMS-backed content experiments, variant design, and interpreting A/B or multivariate test results.
Skip if: Skip for backend API load testing, paid ad campaign setup, or analytics pipeline engineering without content variants.
When should I use this skill?
User asks about A/B testing content, CMS experiment variants, sample size, or interpreting experiment statistics.
What you get
A documented experiment plan with hypotheses, metrics, variant setup guidance, and references for statistically sound analysis.
- Experiment design plan
- Metric and sample-size specification
- Variant analysis checklist
Files
Content Experimentation Best Practices
Principles and patterns for running effective content experiments to improve conversion rates, engagement, and user experience.
When to Apply
Reference these guidelines when:
- Setting up A/B or multivariate testing infrastructure
- Designing experiments for content changes
- Analyzing and interpreting test results
- Building CMS integrations for experimentation
- Deciding what to test and how
Core Concepts
A/B Testing
Comparing two variants (A vs B) to determine which performs better.
Multivariate Testing
Testing multiple variables simultaneously to find optimal combinations.
Statistical Significance
The confidence level that results aren't due to random chance.
Experimentation Culture
Making decisions based on data rather than opinions (HiPPO avoidance).
References
Start with the reference that matches the current problem, such as design, statistics, CMS integration, or pitfalls. See references/ for detailed guidance:
references/experiment-design.md— Hypothesis framework, metrics, sample size, and what to testreferences/statistical-foundations.md— p-values, confidence intervals, power analysis, Bayesian methodsreferences/cms-integration.md— CMS-managed variants, field-level variants, external platformsreferences/common-pitfalls.md— 17 common mistakes across statistics, design, execution, and interpretation
CMS Integration Patterns
Integrating experimentation with your CMS enables content teams to run tests without developer intervention.
Architecture Options
1. CMS-Managed Variants
Store experiment variants as content in the CMS.
Pros: Content team autonomy, version controlled Cons: More complex queries, potential publish coordination
// Experiment document
defineType({
name: 'experiment',
type: 'document',
fields: [
defineField({ name: 'name', type: 'string' }),
defineField({ name: 'status', type: 'string', options: {
list: ['draft', 'running', 'paused', 'concluded']
}}),
defineField({
name: 'variants',
type: 'array',
of: [{
type: 'object',
fields: [
defineField({ name: 'name', type: 'string' }),
defineField({ name: 'weight', type: 'number' }),
defineField({ name: 'content', type: 'reference', to: [{ type: 'page' }] }),
]
}]
}),
defineField({ name: 'startDate', type: 'datetime' }),
defineField({ name: 'endDate', type: 'datetime' }),
]
})2. Field-Level Variants
Store variants as fields on the content document.
Pros: Simpler queries, content stays together Cons: Less flexible, schema complexity
defineType({
name: 'landingPage',
fields: [
defineField({ name: 'headline', type: 'string' }),
defineField({
name: 'headlineVariantB',
type: 'string',
description: 'A/B test variant (leave empty if not testing)'
}),
defineField({ name: 'activeExperiment', type: 'string' }),
]
})3. External Experimentation Platform
Use dedicated tools (Optimizely, LaunchDarkly, VWO) with CMS content.
Pros: Robust analytics, proven platforms Cons: Additional cost, integration complexity
// CMS stores experiment IDs, platform handles assignment
defineField({
name: 'experimentId',
type: 'string',
description: 'Optimizely experiment ID'
})Implementation Pattern (CMS-Managed)
1. Experiment Schema
defineType({
name: 'experiment',
type: 'document',
fields: [
defineField({ name: 'name', type: 'string', validation: r => r.required() }),
defineField({ name: 'hypothesis', type: 'text' }),
defineField({
name: 'status',
type: 'string',
options: { list: ['draft', 'running', 'concluded'] },
initialValue: 'draft'
}),
defineField({
name: 'variants',
type: 'array',
of: [{
type: 'object',
name: 'variant',
fields: [
defineField({ name: 'id', type: 'string' }),
defineField({ name: 'name', type: 'string' }),
defineField({ name: 'weight', type: 'number', initialValue: 50 }),
]
}],
validation: r => r.min(2).error('Need at least 2 variants')
}),
defineField({ name: 'targetPage', type: 'reference', to: [{ type: 'page' }] }),
defineField({ name: 'targetField', type: 'string' }),
]
})2. Variant Content
// On the page being tested
defineField({
name: 'experimentVariants',
type: 'array',
of: [{
type: 'object',
fields: [
defineField({ name: 'experimentId', type: 'reference', to: [{ type: 'experiment' }] }),
defineField({ name: 'variantId', type: 'string' }),
defineField({ name: 'headline', type: 'string' }),
// Other variant-specific fields
]
}]
})3. Frontend Assignment
// Middleware or server-side
function assignVariant(experimentId: string, variants: Variant[]): string {
// Check for existing assignment in cookie
const cookieKey = `exp_${experimentId}`
const existing = getCookie(cookieKey)
if (existing) return existing
// Random assignment based on weights
const rand = Math.random() * 100
let cumulative = 0
for (const variant of variants) {
cumulative += variant.weight
if (rand <= cumulative) {
setCookie(cookieKey, variant.id, { maxAge: 30 * 24 * 60 * 60 })
return variant.id
}
}
return variants[0].id
}4. Query with Variant
*[_type == "page" && slug.current == $slug][0]{
...,
"experiment": experimentVariants[experimentId->status == "running"][0]{
experimentId->{name, _id},
variantId,
headline
}
}Analytics Integration
Event Tracking
// Track experiment exposure
function trackExposure(experimentId: string, variantId: string) {
analytics.track('Experiment Viewed', {
experimentId,
variantId,
timestamp: new Date().toISOString()
})
}
// Track conversion
function trackConversion(experimentId: string, variantId: string, metric: string) {
analytics.track('Experiment Conversion', {
experimentId,
variantId,
metric,
timestamp: new Date().toISOString()
})
}Data Layer
// Push to data layer for analytics tools
window.dataLayer.push({
event: 'experiment_assignment',
experiment_id: experimentId,
variant_id: variantId
})Best Practices
Content Team Workflow
1. Create experiment document with hypothesis 2. Create variant content 3. Set status to "running" 4. Monitor results 5. Set status to "concluded" and record winner
Avoid Flicker
- Assign variants server-side when possible
- Use CSS to hide content until variant determined
- Pre-render both variants, show based on assignment
Clean Up
- Archive concluded experiments
- Remove losing variant content
- Implement winner as default
Common Experimentation Pitfalls
Avoid these mistakes that invalidate results or lead to wrong conclusions.
Statistical Mistakes
1. Stopping Early (Peeking)
The problem: Checking results daily and stopping when you see significance.
Why it's wrong: Statistical significance fluctuates. At any point during a test, you might see "significance" that disappears with more data. This is called the "peeking problem" or "repeated significance testing."
The fix:
- Pre-calculate required sample size
- Commit to running until you reach it
- If you must peek, use sequential testing methods that account for multiple looks
2. Underpowered Tests
The problem: Running tests without enough traffic to detect realistic effect sizes.
Why it's wrong: You'll conclude "no difference" when there actually is one—you just couldn't detect it.
The fix:
- Calculate required sample size before starting
- Be realistic about minimum detectable effect (can you act on a 0.5% improvement?)
- If traffic is low, test bigger changes
3. Multiple Comparisons
The problem: Testing many variants or metrics and celebrating any that reach significance.
Why it's wrong: With 20 metrics, you expect 1 false positive at 95% confidence—by chance alone.
The fix:
- Define ONE primary metric before starting
- Use Bonferroni correction or similar for multiple comparisons
- Treat secondary metrics as directional, not conclusive
4. Ignoring Segments
The problem: Only looking at aggregate results.
Why it's wrong: Simpson's Paradox—overall winner might be loser for your key segments.
The fix:
- Always segment by device, traffic source, user type
- Check if results are consistent across segments
- If segments differ dramatically, investigate why
Design Mistakes
5. Testing Too Many Things
The problem: Changing headline, image, CTA, and layout simultaneously.
Why it's wrong: You won't know which change caused the result. And each variable multiplies required sample size.
The fix:
- Test one variable at a time (A/B testing)
- If testing multiple, use proper multivariate testing with adequate sample size
- Prioritize highest-impact changes first
6. Vague Hypothesis
The problem: "Let's see if this new design is better."
Why it's wrong: Without a hypothesis, you can't learn WHY something worked (or didn't).
The fix:
- State: "We believe [change] will [impact metric] because [reasoning]"
- Even if you're wrong, you learn something
7. No Control
The problem: Changing the control during the test, or not having one.
Why it's wrong: You need a stable baseline to compare against.
The fix:
- Never modify the control mid-test
- If you must change it, start a new test
- Document exactly what the control is
Execution Mistakes
8. External Contamination
The problem: Running a test during a sale, holiday, or major event.
Why it's wrong: External factors affect both variants differently, contaminating results.
The fix:
- Avoid tests during unusual periods
- If unavoidable, note it and extend the test past the event
- Compare to the same period historically
9. Selection Bias
The problem: Testing on a non-representative sample (e.g., only logged-in users).
Why it's wrong: Results won't generalize to your full audience.
The fix:
- Test on representative traffic
- Be explicit about who's included/excluded
- Note limitations when reporting results
10. Implementation Bugs
The problem: Variants don't render correctly, tracking fires incorrectly, assignment is biased.
Why it's wrong: You're not testing what you think you're testing.
The fix:
- QA both variants thoroughly before launch
- Verify tracking events fire correctly
- Check assignment distribution matches weights
Interpretation Mistakes
11. Celebrating Trivial Wins
The problem: Implementing a change because it was "statistically significant" even though the effect was tiny.
Why it's wrong: Statistical significance ≠ practical significance. A 0.01% improvement isn't worth the complexity.
The fix:
- Define minimum meaningful effect before starting
- Consider implementation cost vs. benefit
- Don't over-optimize
12. Ignoring Confidence Intervals
The problem: Only reporting point estimates ("5% improvement!").
Why it's wrong: The true effect could be anywhere in the confidence interval.
The fix:
- Report confidence intervals: "5% improvement (95% CI: 2%-8%)"
- Base decisions on the lower bound for conservative estimates
- Wider intervals = more uncertainty
13. Not Documenting Learnings
The problem: Running tests but not recording what you learned.
Why it's wrong: You'll repeat mistakes, forget context, lose institutional knowledge.
The fix:
- Document every test: hypothesis, results, learnings
- Include what surprised you
- Build a searchable knowledge base
Organizational Mistakes
14. HiPPO (Highest Paid Person's Opinion)
The problem: Running experiments but ignoring results when leadership disagrees.
Why it's wrong: Defeats the purpose of data-driven decision making.
The fix:
- Get buy-in before testing that results will be honored
- Present data clearly to stakeholders
- Frame as "learning" not "winning/losing"
15. Testing Everything
The problem: Running experiments on trivial changes that don't matter.
Why it's wrong: Wastes resources, creates testing fatigue, delays important experiments.
The fix:
- Prioritize tests by potential impact
- Not everything needs a test—use judgment for low-risk changes
- Focus experimentation resources on high-value decisions
16. Sample Ratio Mismatch (SRM)
The problem: The actual traffic split doesn't match the intended split (e.g., you expect 50/50 but observe 52/48).
Why it's wrong: SRM is a strong signal of an implementation bug — broken randomization, bot contamination, or redirect issues. Results from experiments with SRM cannot be trusted.
The fix:
- Check the actual split ratio against expected before analyzing results
- Use a chi-squared test to detect statistically significant mismatches
- If SRM is detected, investigate the root cause before drawing any conclusions
- Common causes: bot traffic, browser redirects dropping users, bucketing bugs
17. Novelty and Primacy Effects
The problem: Users react differently to new designs initially, and the effect fades over time.
Why it's wrong: Short experiments may show inflated effects that don't persist. Returning users may click more simply because something looks new.
The fix:
- Run experiments for at least 2 full business cycles
- Segment results by new vs. returning users
- If possible, check whether the effect holds in the second week vs. the first
Experiment Design Principles
Well-designed experiments produce actionable insights. Poorly designed ones waste time and can mislead.
The Experiment Framework
1. Hypothesis
State what you believe and why.
Bad: "Let's test a new headline" Good: "We believe a benefit-focused headline will increase signup rate by 10% because users are currently confused about our value proposition"
Structure: "We believe [change] will [impact metric] because [reasoning]"
2. Success Metric
Define primary and guardrail metrics.
Primary metric: The main thing you're trying to improve (conversion rate, engagement time) Guardrail metrics: Things that shouldn't get worse (bounce rate, page load time)
3. Sample Size
Calculate required sample size before starting.
Factors:
- Baseline conversion rate
- Minimum detectable effect (MDE)
- Statistical significance level (usually 95%)
- Statistical power (usually 80%)
Use calculators like Evan Miller's.
4. Duration
Run tests for full business cycles.
- Minimum: 1-2 weeks (capture weekly patterns)
- Include weekends
- Avoid holidays and major events
- Don't stop early when you see "winning" results
What to Test
High-Impact Areas
- Headlines and value propositions
- Call-to-action text and placement
- Form length and fields
- Pricing presentation
- Social proof placement
Lower-Impact (Usually)
- Button colors
- Minor copy tweaks
- Image variations (unless hero)
- Footer changes
Test Priority Matrix
| Impact | Effort | Priority |
|---|---|---|
| High | Low | Do first |
| High | High | Plan carefully |
| Low | Low | Quick wins |
| Low | High | Avoid |
Sanity Integration Pattern
// Experiment variant schema
defineType({
name: 'experimentVariant',
type: 'object',
fields: [
defineField({ name: 'name', type: 'string' }),
defineField({ name: 'weight', type: 'number', description: 'Traffic allocation (0-100)' }),
defineField({ name: 'content', type: 'reference', to: [{ type: 'page' }] }),
]
})
// Experiment document
defineType({
name: 'experiment',
type: 'document',
fields: [
defineField({ name: 'name', type: 'string' }),
defineField({ name: 'hypothesis', type: 'text' }),
defineField({ name: 'status', type: 'string', options: {
list: ['draft', 'running', 'concluded']
}}),
defineField({ name: 'variants', type: 'array', of: [{ type: 'experimentVariant' }] }),
defineField({ name: 'startDate', type: 'datetime' }),
defineField({ name: 'endDate', type: 'datetime' }),
defineField({ name: 'winner', type: 'string' }),
defineField({ name: 'learnings', type: 'text' }),
]
})Avoiding Common Mistakes
Don't peek and stop early
Statistical significance can fluctuate. Commit to your sample size.
Don't test too many things at once
Each variable multiplies required sample size.
Don't ignore segmentation
Winners may differ by device, traffic source, or user type.
Document everything
Future you (and your team) will thank you.
Statistical Foundations
Understanding basic statistics prevents misinterpreting experiment results.
Table of Contents
- Key concepts
- Sample size calculation
- Common statistical mistakes
- Interpreting results
- Alternative approaches
- When to trust results
Key Concepts
Statistical Significance
A measure of whether observed differences are likely real or due to chance.
- p-value < 0.05: "Statistically significant" at 95% confidence
- Means: If there were no real difference, there's less than a 5% chance of seeing results this extreme
- Does NOT mean: The change is important or meaningful
- Common misconception: The p-value is NOT "the probability the result is due to chance." It's the probability of observing data this extreme assuming the null hypothesis is true.
Confidence Interval
A range of plausible values for the true effect.
Example: "Conversion rate increased by 5% (95% CI: 2% to 8%)"
- Best estimate: 5% improvement
- Could be as low as 2% or as high as 8%
- Narrower intervals = more certainty
Statistical Power
The ability to detect a real effect when it exists.
- Standard: 80% power
- Higher power = larger sample size needed
- Low power = might miss real improvements
Minimum Detectable Effect (MDE)
The smallest improvement worth detecting.
- Smaller MDE = larger sample size needed
- Be realistic: Can you act on a 0.5% improvement?
Sample Size Calculation
Before running a test, calculate required sample size:
Required per variant = 16 × σ² / MDE²
Where:
- σ² = variance (for conversion rate: p × (1-p))
- MDE = minimum detectable effect (absolute)For a 5% baseline conversion rate, detecting a 1% absolute lift (5% → 6%):
- σ² = 0.05 × 0.95 = 0.0475
- MDE² = 0.01² = 0.0001
- n = 16 × 0.0475 / 0.0001 = 7,600 per variant
- Total: ~15,200 visitors minimum
Common Statistical Mistakes
Multiple Comparisons Problem
Testing 10 variants increases false positive rate.
Solution: Adjust significance threshold (Bonferroni correction) or use sequential testing methods.
Peeking Problem
Checking results daily and stopping when significant.
Why it's wrong: Significance fluctuates. Early "winners" often regress.
Solution: Pre-commit to sample size and duration. Use sequential testing if you must peek.
Simpson's Paradox
Overall results hide segmented truths.
Example:
- Overall: Variant B wins
- Mobile users: Variant A wins
- Desktop users: Variant A wins
- How? Different traffic mix per variant
Solution: Always segment by major factors (device, traffic source).
Survivorship Bias
Only analyzing users who completed the funnel.
Solution: Include all visitors, not just converters.
Interpreting Results
Significant + Meaningful
Clear win. Implement the change.
Significant + Trivial
Statistically different but tiny effect. Consider if worth the complexity.
Not Significant + Large Effect
Might be real but underpowered. Extend the test or accept uncertainty.
Not Significant + Small Effect
No detectable difference. Either no real effect or test was underpowered.
Alternative Approaches
Bayesian A/B Testing
An alternative to traditional (frequentist) hypothesis testing. Bayesian methods provide:
- Direct probability statements: "There's a 95% probability Variant B is better" (more intuitive than p-values)
- No peeking problem: Continuous monitoring is built in — you can check results at any time
- Credible intervals: Directly interpretable as "the true value falls in this range with X% probability"
Bayesian methods are offered by platforms like VWO and are useful when you need to make decisions with limited traffic or want more intuitive reporting for stakeholders.
Multi-Armed Bandits
Dynamically allocate more traffic to winning variants while still learning:
- Thompson Sampling: Balances exploration (learning) with exploitation (serving the best variant)
- Best for: Ongoing optimization where you want to minimize regret during the test
- Trade-off: Faster convergence to the winner, but less statistical rigor than fixed-allocation A/B tests
Consider bandits for content recommendations, personalization, or situations where the cost of showing a losing variant is high.
Sequential Testing
For teams that need to monitor experiments continuously:
- Group sequential designs (O'Brien-Fleming, Lan-DeMets) allow pre-planned interim analyses
- Always-valid p-values let you check results at any time without inflating false positive rates
- Use when you must balance the peeking problem with business pressure to act on results quickly
When to Trust Results
Checklist before declaring a winner:
- [ ] Reached pre-calculated sample size
- [ ] Ran for full business cycle (1-2 weeks minimum)
- [ ] p-value < 0.05 (or your chosen threshold)
- [ ] Effect size is meaningful for business
- [ ] Results consistent across major segments
- [ ] No external factors contaminated results
Related skills
How it compares
Pick content-experimentation-best-practices for experiment methodology and statistical guardrails rather than skills that only generate UI variants without analysis guidance.
FAQ
What experiment types does it cover?
A/B tests comparing two variants and multivariate tests exploring multiple variable combinations.
Where is detailed statistical guidance?
See references/statistical-foundations.md for p-values, confidence intervals, power analysis, and Bayesian methods.
Does it include CMS integration patterns?
Yes. references/cms-integration.md covers CMS-managed variants and field-level variant strategies.
Is Content Experimentation Best Practices safe to install?
skills.sh reports 3 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.