← All free tools

Free tool

A/B test significance calculator

Enter your visitors and conversions for A and B and get a proper two-proportion z-test: p-value, confidence interval and a plain-English verdict -- not just a number to squint at. It loads with a realistic example that is not yet significant, because knowing when you cannot call a winner yet is the most useful thing a calculator like this does. A sample-size planner is built in too, for fixing your numbers before you start.

Try it nowHow to read the result ↓

A/B test calculator

Control (A)
Variant (B)
Test settings

Two-tailed asks "is B different from A, in either direction?" -- the honest default, because before a test runs you rarely know for certain which way it will move. One-tailed only asks "is B better?" and makes it slightly easier to call a win, which is exactly why it is easy to misuse.

Conversion rate A 10.00%, B 11.50%. P-value 0.13. Not statistically significant at 95% confidence.

B is not yet statistically significant at 95% confidence.

You would see a difference this large by chance about 1 time in 8, if there were really no difference between A and B. That is too often to call it a real effect.

Conversion rate A

10.00%

Conversion rate B

11.50%

Absolute uplift

+1.50%

Relative uplift

+15.0%

P-value

0.13

Z-score

1.53

95% two-sided confidence interval on the difference (B minus A)

-0.42% to 3.42%

This is the plausible range for the true difference. If it spans zero, that is another way of seeing that the test has not yet ruled out "no difference" -- watch this alongside the p-value, not instead of it.

How to read it

What the numbers actually mean.

P-value

Roughly, how often you would see a gap this large between A and B purely by chance, if the two versions truly performed identically. A p-value of 0.03 means about 1 time in 33. Smaller is stronger evidence of a real difference -- it is not the probability that B is better, which is a common and understandable misreading.

Confidence level

95% confidence means you are prepared to accept being wrong about 1 time in 20 when you call a result significant. 99% is stricter and needs more evidence; 90% is looser. There is no universally "correct" level -- it is a judgement call about how costly a false positive would be for you.

Confidence interval

The plausible range for the true difference between A and B, given your data. A wide interval means your sample is too small to say much with precision, even if the point estimate looks impressive. If the interval spans zero, the data has not ruled out "no real difference".

Statistical vs business significance

A result can be statistically significant and still not worth shipping -- a genuine but tiny uplift on a low-traffic, low-value page may not justify the engineering or design cost of maintaining it. Significance answers "is this real?", not "is this worth doing?".

Be honest with yourself

Most under-powered tests cannot answer the question you are asking.

Fix your sample size before you start, using the planner tab above, not after you have already run the test. A test that never reaches its planned sample size is not a smaller version of a valid test -- it is usually just unable to tell a real effect from noise, in either direction. Checking a live test daily and stopping the moment it first looks significant inflates your real false-positive rate well past the 5% a single planned check implies, sometimes by several times over. And on genuinely low-traffic pages, the honest answer is often that no reasonable test length will produce a trustworthy result at all -- that is a real constraint, not a tooling problem this calculator can solve for you.

FAQ

Common questions

What maths does this calculator use?

A standard two-proportion z-test. The p-value uses a pooled standard error (it assumes, for the sake of the test, that there is truly no difference between A and B, then asks how surprising your actual result would be under that assumption). The confidence interval on the difference uses an unpooled standard error, because it is describing the observed gap rather than testing a hypothesis about it. The normal distribution itself is computed properly via the Abramowitz-Stegun approximation to the error function, not read off a lookup table -- so the result should match any other correctly-built calculator, including Evan Miller's, to within rounding.

Should I use one-tailed or two-tailed?

Two-tailed is the honest default and this tool starts there. A two-tailed test asks "is B different from A, in either direction?" -- which is the true state of your knowledge before a test runs, because a change can make things worse as easily as better. A one-tailed test only asks "is B better?", which makes it slightly easier to call a result significant. That is exactly why it is easy to misuse: switching to one-tailed after you have already seen the data, because it happens to turn a borderline result significant, is a form of p-hacking.

Why does this say my test is not significant when the conversion rate went up?

A higher observed rate is not the same as a real, repeatable difference. With small numbers of visitors or conversions, a run of ordinary luck can easily produce a gap this size even when the two versions perform identically in the long run. The p-value is an estimate of how often that would happen by chance alone -- if it is not small enough, the honest conclusion is "we do not yet know", not "B lost".

Why does the sample size calculator matter if I already have results?

Because it tells you what you needed before you started, which is the only time it can actually help. If your test never reached that sample size, a significant-looking p-value part-way through is much more likely to be noise, and a non-significant one does not mean B failed -- it may mean the test never had a real chance to detect the effect you were hoping for (the test was "underpowered"). Most A/B tests run on low-traffic sites or pages fall into exactly this trap: the traffic simply is not there to answer the question being asked of it.

Is a statistically significant result the same as a result worth shipping?

No. Statistical significance only tells you the difference is probably real, not that it is large enough to matter commercially, not that it will hold up next month, and not that it is worth the engineering cost of keeping two versions running. A tiny but "significant" uplift on a low-value page may not be worth acting on; a promising but not-yet-significant result on a high-value page may be worth extending the test rather than abandoning it.

What is "peeking" and why does the tool warn about it?

Peeking is checking a test's significance repeatedly while it is still running and stopping as soon as it first looks significant. Each look is a fresh chance for random noise to cross the threshold, so a test checked daily and stopped on the first green result has a real false-positive rate far higher than the 5% a single, planned check implies -- often several times higher. The fix is deciding your sample size in advance (the second calculator on this page) and waiting until you reach it before treating the result as final.

Take it with you

Running tests as part of a bigger CRO or research programme?

A calculator like this checks one test at a time. If you want help setting up a testing programme properly -- prioritisation, sample size planning up front, or training your team to run and read tests themselves -- tell us a little about what you are working on.

By submitting, you agree to receive course updates and our occasional newsletter. Unsubscribe any time. Privacy policy.

Learning UX yourself

Beginner UX (AI) Design starts 21 September 2026.

Tools help you do the work. They do not teach you the judgement behind it. Beginner UX (AI) Design is taught live online by working UX professionals, built for career-changers with no design background. Get the brochure and see the full curriculum, the schedule and what the weeks actually involve.

Sent to designers across the UK.

Cohort 1 starts 21 September 2026. Limited to 15 students.

By submitting, you agree to receive course updates and our occasional newsletter. Unsubscribe any time. Privacy policy.

Working in a team rather than learning on your own? We run in-house UX training built around your own product, or you can book a call to talk it through.