Free tool
A/B test calculator
Enter what each variant got and see whether the difference is likely real, or work out how many visitors a test needs before you start. The formulas are shown, and nothing leaves your browser.
Enter both variants to see the result.
Enter a rate and a lift.
How it is calculated
With conversion rates p₁ = c₁ ÷ n₁ and p₂ = c₂ ÷ n₂, and the pooled rate p = (c₁ + c₂) ÷ (n₁ + n₂):
z = (p₂ − p₁) ÷ √( p (1 − p) (1/n₁ + 1/n₂) ) p-value = 2 × (1 − Φ(|z|)) Φ is the standard normal distribution relative lift = (p₂ − p₁) ÷ p₁ 95% range for the difference = (p₂ − p₁) ± 1.96 × √( p₁(1−p₁)/n₁ + p₂(1−p₂)/n₂ )
Visitors needed in each variant, for a two-sided test with significance level α and power 1 − β, where p̄ = (p₁ + p₂) ÷ 2:
n = ( z(1−α/2) × √(2 p̄ (1 − p̄)) + z(power) × √(p₁(1−p₁) + p₂(1−p₂)) )² ÷ (p₂ − p₁)²
All of it runs in your browser; nothing you enter is sent anywhere.
Count the conversions in trckable
Mark the action you test as a goal and trckable shows its conversion rate for any page or source, so the numbers for this calculator come from your own data. Start free. To see what a lift is worth in money, use the conversion revenue calculator.
An A/B test shows two versions of a page to two groups of visitors and compares how many in each group convert. Chance alone makes two identical pages differ a little, so the question is never "which number is bigger" but "is this gap bigger than chance would produce". This calculator answers that with a standard two-proportion z-test.
What the numbers mean
- Conversion rate is conversions divided by visitors for each variant. See conversion rate.
- Relative lift is how much better B is than A, as a share of A's rate. Going from 3.0% to 3.6% is a 20% lift, even though the gap is 0.6 points.
- p-value is the chance of seeing a gap at least this large if the two versions were really the same. A small value, below 0.05 by the usual convention, means chance is an unlikely explanation.
- Confidence is shown as one minus the p-value. Read it as "how surprising this gap would be if nothing changed", not as the probability that B is better.
How to run a test you can trust
- Decide the size first. Use the second tab to find how many visitors each variant needs, and plan to run until you have them.
- Do not stop when it looks good. Checking the result every day and stopping at the first significant moment makes false wins far more common than the p-value suggests. Fix the sample size, or the end date, in advance.
- Change one thing. If you change the headline, the picture and the price at once, you learn only that the bundle won or lost.
- Cover whole weeks. Weekday and weekend visitors behave differently. Run for complete weeks.
- Count the same thing for both. Use one goal, with the same definition of a visitor and a conversion, for A and B.
What this calculator cannot tell you
A significant result says the gap is unlikely to be chance. It does not say the gap is large, that it will last, or that it holds for visitors you did not test. A result that is not significant does not prove the variants are equal: it often means the test was too small to see a difference. The second tab shows how small a lift you can hope to detect with the traffic you have; for a small site, that may be only large changes, and the honest plan is to test bolder ideas.
Testing several variants at once, or looking at many goals, raises the odds of a false win. This tool compares two variants on one goal.
Questions
- What is a good p-value?
- By convention, below 0.05 counts as significant, meaning chance would produce a gap this large less than one time in twenty. It is a convention, not a law: a lower threshold suits costly mistakes, a higher one suits cheap, reversible tests.
- Why does the sample size tab ask for a minimum lift?
- The smaller the lift you want to detect, the more visitors you need, and the need grows quickly: halving the lift needs about four times the visitors. Pick the smallest lift that would be worth acting on.
- My result is not significant. Did the test fail?
- Not necessarily. It may be too small to detect the change. Check how many visitors you would need for the lift you hoped for, and either collect them or test a bigger change.
- Can I stop the test early if B is clearly winning?
- Only if you decided so before starting. Stopping when the result first crosses the line, after peeking many times, inflates the chance of a false win.
The newsletter
Notes on measuring what matters.
Occasional emails from Albi: new posts, one chart worth reading, what changed in trckable.
We send one email to confirm your address, and nothing more until you confirm. Every newsletter has a link to leave. What we keep, and who sends it: privacy.
Counted, never watched.