Free A/B Test Significance Calculator

← All free tools
FREE TOOL · EXPERIMENTATION

A/B Test Significance Calculator

Drop in two variants and instantly know whether your winner is statistically real — or just random noise you would be unwise to ship.

—
Variant A CVR
—
Variant B CVR
—
Relative Lift
—
Confidence
Enter your numbers to see significance.

Confidence below 95% means the result could be random chance — do not ship it yet. This uses a two-tailed two-proportion z-test, the same standard real testing tools use.

What the confidence number means

This calculator runs a two-tailed two-proportion z-test, the standard check for comparing two conversion rates. Confidence is the probability that a difference this large would not have appeared by chance alone if the two variants were genuinely identical. At 95% confidence there is roughly a one-in-twenty chance the result is noise, which is the threshold most teams treat as the minimum before shipping.

Two-tailed matters. It tests whether B differs from A in either direction, rather than assuming in advance that B is the winner. A one-tailed test reaches significance sooner, which is tempting and is exactly why it is so often misused.

What significance does not tell you

Confidence answers whether a difference is real, not whether it is worth having. A statistically significant 0.3% lift on a low-value page can be real and still not worth the engineering to ship it, while a large lift at 80% confidence may be worth running longer rather than discarding. Significance is also not a stopping rule: checking daily and shipping the moment the number crosses 95% inflates false positives badly, because you have effectively run many tests. Decide the sample size and duration before you start, run through at least one full business cycle so weekday and weekend behaviour are both represented, and read the result once.

Frequently asked questions

What is statistical significance in A/B testing?

It is the confidence that a difference between two variants is real and not random chance. 95% confidence is the standard threshold for calling a winner.

How large should an A/B test sample be?

Large enough to reach 95% confidence. Small samples produce noisy results that look like wins but disappear at scale.

What does relative lift mean?

Relative lift is the percentage improvement of variant B over variant A. A jump from 5% to 6% conversion is a 20% relative lift.

Why should I not stop a test the moment it hits 95%?

Because repeatedly checking and stopping at the first significant moment is a form of multiple testing, and it produces false winners far more often than the 5% the threshold implies. Random fluctuation crosses 95% temporarily in a surprising number of tests that end up flat. Fix the duration in advance and read the result at the end.
Run tests that actually compound.

DigiJaws runs disciplined experimentation programs that turn small wins into a permanent conversion advantage.

Deploy CRO →
Keep going · related free tools
See your site’s full SEO score — free.Instant audit, 10 checks, no signup to see it.
Run free audit →
Free tools →My toolkit →Locations →Comparisons →
Verified Agent-Ready by DigiJaws
Start free with every tool — no credit card, cancel anytimeStart Free