Skip to main content
Conversion Optimisation

CRO Testing: How to Run A/B Tests That Actually Mean Something

Most A/B tests declare a winner that does not survive contact with reality. The cause is almost never the tool — it is stopping early, testing something too small to detect, or running on traffic that cannot support a conclusion.

By Samit Dinesh Shah 2 min read

Work out the sample size before you start

The single most common failure is running a test without knowing how much traffic it needs. Required sample size depends on your baseline conversion rate, the smallest improvement worth detecting, and how confident you want to be.

The relationship is brutal: detecting a small lift needs enormously more traffic than detecting a large one. Halving the effect you want to detect roughly quadruples the sample needed. A site converting at 2 percent that wants to detect a 10 percent relative improvement typically needs tens of thousands of visitors per variant.

Calculate it first with our A/B Test Calculator. If the answer exceeds your realistic traffic, do not run the test — test something bolder instead, because a bigger change needs less traffic to prove.

Peeking is what invalidates most results

Checking a running test and stopping when it looks significant is called peeking, and it is the reason so many wins evaporate. Significance calculations assume a single check at a predetermined sample size. Checking repeatedly and stopping at the first favourable moment means you will eventually hit a random fluctuation and call it a result.

With enough peeking, a test between two identical pages will produce a significant result surprisingly often. The fix is to decide your sample size and duration in advance and not stop early regardless of what the dashboard shows.

Duration matters independently of sample size

Even if you hit your sample in three days, run for at least one full week, and preferably two. Behaviour varies systematically by day: weekday and weekend traffic convert differently, and a test that only saw Monday to Wednesday is measuring a slice of your audience.

Full weeks also avoid a subtler distortion. Ending mid-week means one variant may have received more weekend traffic than the other, which is a difference in audience rather than in the thing you tested.

Test things big enough to matter

Button colour tests are popular because they are easy and almost always inconclusive. The effect is too small to detect on normal traffic, so you spend three weeks learning nothing.

Test changes substantial enough to move behaviour: the offer itself, the headline and value proposition, page structure and what appears above the fold, form length, and pricing presentation. These produce effects large enough to detect and teach you something about your customers even when they lose.

Form length is often the highest-value test available. Every additional required field costs completions, and most forms ask for things nobody uses.

Record what you learn, including losses

A losing test is not a wasted test. It tells you something about your audience that the winning variant would not have. Teams that keep a record of every test, hypothesis and outcome compound knowledge; teams that only remember wins repeat the same experiments.

Write the hypothesis before running: what you are changing, what you expect to happen, and why. It forces clarity and prevents the retrospective storytelling that turns a random fluctuation into a fictional insight.

Key takeaways

  • Calculate sample size before running. If you cannot reach it, test a bolder change instead.
  • Peeking invalidates significance. Decide the stopping point in advance and hold to it.
  • Run full weeks — behaviour varies by day and partial weeks bias the comparison.
  • Test the offer, headline, structure and form length, not button colours.
  • Record every test including losses, and write the hypothesis before you start.

Samit Dinesh Shah

Founder · EmproIT

Samit founded EmproIT and spends most of his time on the uncomfortable question of which marketing spend is genuinely producing revenue.

Frequently asked questions

How long should I run an A/B test?

Until you reach your pre-calculated sample size, and at least one full week regardless. Two full weeks is safer because it covers day-of-week variation twice and smooths one-off events. Always end on the same day of the week you started, so both variants have seen the same mix of weekdays and weekends. Ending mid-week can leave one variant with more weekend traffic than the other, which is a difference in audience rather than a difference caused by what you tested.

What sample size do I need?

It depends on your baseline conversion rate, the smallest lift worth detecting, and your confidence level, and the numbers are usually larger than people expect. Detecting small effects requires disproportionately more traffic: roughly speaking, halving the effect you want to detect quadruples the sample needed. A page converting at 2 percent looking for a 10 percent relative improvement will typically need tens of thousands of visitors per variant. Calculate it before starting, because discovering this afterwards means the test was never capable of concluding.

What is the peeking problem?

Peeking is repeatedly checking a running test and stopping as soon as it shows significance. Standard significance calculations assume you look once, at a predetermined sample size. Checking daily and stopping at the first favourable moment dramatically inflates your false positive rate, because random fluctuation will eventually produce a significant-looking result even between two identical pages. This is the single biggest reason A/B test wins fail to replicate. Set your stopping rule in advance and do not deviate from it.

What is a good confidence level?

Ninety-five percent is the usual standard, meaning roughly a one in twenty chance of a false positive. Higher confidence requires more traffic, so there is a genuine trade-off. For decisions that are cheap to reverse, such as a headline change, 90 percent may be acceptable. For expensive or hard-to-undo changes such as a pricing structure or a full redesign, 99 percent is worth the extra traffic. What matters most is choosing the level before the test rather than picking whichever threshold your result happens to clear.

Can I test more than two variants at once?

You can, but each additional variant increases the traffic required and the chance of a false positive, because you are effectively running several comparisons at once. For most sites, sequential two-variant tests produce cleaner answers faster. Multivariate testing, which tests combinations of several elements, needs very high traffic to be meaningful and is rarely justified outside large ecommerce. If you have several ideas, rank them by expected impact and test the biggest one first.

What should I test first?

Start where the potential effect is largest and the traffic requirement is therefore smallest: the offer itself, the headline and value proposition, what appears above the fold, and form length. Form length is frequently the highest-value test available, because every additional required field costs completions and most forms request information nobody actually uses. Leave small cosmetic changes such as button colour alone; the effects are too small to detect on normal traffic and you will spend weeks learning nothing.

Want tests that hold up?

Our CRO team designs experiments with the statistics settled up front, so the wins you ship are wins that survive.

Talk About CRO