← All free tools

Free tool

A/B test significance calculator

Enter visitors and conversions for A and B and get a proper two-proportion z-test: p-value, confidence interval and a plain-English verdict. A sample-size planner is built in too.

Try it nowHow to read the result ↓

A/B test calculator

Control (A)
Variant (B)
Test settings

Two-tailed asks "is B different from A, in either direction?" -- the honest default, because before a test runs you rarely know for certain which way it'll move. One-tailed only asks "is B better?" and makes it slightly easier to call a win, which is exactly why it's easy to misuse.

Conversion rate A 10.00%, B 11.50%. P-value 0.13. Not statistically significant at 95% confidence.

B isn't yet statistically significant at 95% confidence.

You'd see a difference this large by chance about 1 time in 8, if there were really no difference between A and B. That's too often to call it a real effect.

Conversion rate A

10.00%

Conversion rate B

11.50%

Absolute uplift

+1.50%

Relative uplift

+15.0%

P-value

0.13

Z-score

1.53

95% two-sided confidence interval on the difference (B minus A)

-0.42% to 3.42%

This is the plausible range for the true difference. If it spans zero, that's another way of seeing that the test hasn't yet ruled out "no difference" -- watch this alongside the p-value, not instead of it.

How to use this calculator

Two steps, before and after.

1.

Pre-test planning. Before you launch anything, switch to the sample size tab and enter your baseline conversion rate and the smallest effect you care about detecting. It tells you how many visitors per variant you need, so you fix your numbers before the test starts rather than guessing partway through.

2.

Post-test evaluation. Once the test has run, enter the actual visitor and conversion counts for A and B into the calculator above. It runs a proper two-proportion z-test and returns a p-value, confidence interval and a plain-English verdict.

How to read it

What the numbers actually mean.

P-value
It isHow often you'd see a gap this large purely by chance.
It isn'tThe probability that B is better.
Confidence level
It isYour tolerance for being wrong — about 1 in 20 at 95%.
It isn'tA universally "correct" level.
Confidence interval
It isThe plausible range for the true difference, given your data.
It isn'tPrecise — wide means the sample is too small to say much.
Statistical vs business significance
It isThe answer to "is this real?".
It isn'tThe answer to "is this worth doing?" — tiny uplifts may not justify the cost.
Statistical power
It isHow likely the test is to detect a real effect, if one exists.
It isn'tA guarantee — low power risks missing a real difference.
Be honest with yourself

Most under-powered tests can't answer the question you're asking.

Fix your sample size before you start, using the planner tab above, not after you have already run the test. A test that never reaches its planned sample size isn't a smaller version of a valid test -- it's usually just unable to tell a real effect from noise, in either direction. Checking a live test daily and stopping the moment it first looks significant inflates your real false-positive rate well past the 5% a single planned check implies, sometimes by several times over. And on genuinely low-traffic pages, the honest answer is often that no reasonable test length will produce a trustworthy result at all -- that's a real constraint, not a tooling problem this calculator can solve for you.

FAQ

Common questions

What maths does this calculator use?

A standard two-proportion z-test. The p-value uses a pooled standard error (it assumes, for the sake of the test, that there's truly no difference between A and B, then asks how surprising your actual result would be under that assumption). The confidence interval on the difference uses an unpooled standard error, because it's describing the observed gap rather than testing a hypothesis about it. The normal distribution itself is computed properly via the Abramowitz-Stegun approximation to the error function, not read off a lookup table -- so the result should match any other correctly-built calculator, including Evan Miller's, to within rounding.

Should I use one-tailed or two-tailed?

Two-tailed is the honest default and this tool starts there. A two-tailed test asks "is B different from A, in either direction?" -- which is the true state of your knowledge before a test runs, because a change can make things worse as easily as better. A one-tailed test only asks "is B better?", which makes it slightly easier to call a result significant. That's exactly why it's easy to misuse: switching to one-tailed after you have already seen the data, because it happens to turn a borderline result significant, is a form of p-hacking.

Why does this say my test isn't significant when the conversion rate went up?

A higher observed rate isn't the same as a real, repeatable difference. With small numbers of visitors or conversions, a run of ordinary luck can easily produce a gap this size even when the two versions perform identically in the long run. The p-value is an estimate of how often that would happen by chance alone -- if it isn't small enough, the honest conclusion is "we don't yet know", not "B lost".

Why does the sample size calculator matter if I already have results?

Because it tells you what you needed before you started, which is the only time it can actually help. If your test never reached that sample size, a significant-looking p-value part-way through is much more likely to be noise, and a non-significant one doesn't mean B failed -- it may mean the test never had a real chance to detect the effect you were hoping for (the test was "underpowered"). Most A/B tests run on low-traffic sites or pages fall into exactly this trap: the traffic simply isn't there to answer the question being asked of it.

Is a statistically significant result the same as a result worth shipping?

No. Statistical significance only tells you the difference is probably real, not that it's large enough to matter commercially, not that it'll hold up next month, and not that it's worth the engineering cost of keeping two versions running. A tiny but "significant" uplift on a low-value page may not be worth acting on; a promising but not-yet-significant result on a high-value page may be worth extending the test rather than abandoning it.

What's "peeking" and why does the tool warn about it?

Peeking is checking a test's significance repeatedly while it's still running and stopping as soon as it first looks significant. Each look is a fresh chance for random noise to cross the threshold, so a test checked daily and stopped on the first green result has a real false-positive rate far higher than the 5% a single, planned check implies -- often several times higher. The fix is deciding your sample size in advance (the second calculator on this page) and waiting until you reach it before treating the result as final.

Take it with you

Running tests as part of a bigger CRO or research programme?

A calculator like this checks one test at a time. If you want help setting up a testing programme properly -- prioritisation, sample size planning, or training your team to run and read tests themselves -- get in touch.

Get in touch about testingSee in-house training
Learning UX yourself

Beginner UX (AI) Design starts 14 October 2026.

Tools help you do the work. They don't teach you the judgement behind it. Beginner UX (AI) Design is taught live online by working UX professionals, built for career-changers with no design background.

Book a callSee the course

Working in a team rather than learning on your own? We run in-house UX training built around your own product, and we can talk it through on the same call.