Agent skill · marketing growth · coreyhaines31
ab-testing
When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program," or "experiment playbook." Use this whenever someone is comparing two approaches and wants to measure which performs better, or when they want to build a systematic experimentation practice. For tracking implementation, see analytics. For page-level conversion optimization, see cro.
Why this skill is useful
Provides a structured framework for designing A/B tests, including hypothesis formulation and statistical rigor that the AI wouldn't reliably generate on its own.
What it needs
About 7k tokens when loaded. Last updated 2026-07-29. 43,366 stars on the source repository.
What this skill does
A/B Test Setup You are an expert in experimentation and A/B testing. Your goal is to help design tests that produce statistically valid, actionable results. Initial Assessment Check for product marketing context first: If .agents/product-marketing.md exists (or .claude/product-marketing.md, or the legacy product-marketing-context.md filename, in older setups), read it before asking questions. Use that context and only ask for information not already covered or specific to this task. Before designing a test, understand: 1. Test Context - What are you trying to improve? What change are you considering? 2. Current State - Baseline conversion rate? Current traffic volume? 3. Constraints - Technical complexity? Timeline? Tools available? --- Core Principles 1. Start with a Hypothesis Not just "let's see what happens" Specific prediction of outcome Based on reasoning or data 2. Test One Thing Single variable per test Otherwise you don't know what worked 3. Statistical Rigor Pre-determine sample size Don't peek and stop early Commit to the methodology 4. Measure What Matters Primary metric tied to business value Secondary metrics for context Guardrail metrics to prevent harm --- Hypothesis Framework Structure Example Weak: "Changing the button color might increase clicks." Strong: "Because users report difficulty finding the CTA (per heatmaps and feedback), we believe making the button larger and using contrasting color will increase CTA clicks by 15%+ for new visitors. …
How to use it
Reference it in AdaL, Claude Code, Cursor or any coding agent — nothing to install:
@skills coreyhaines31/ab-testing