How to Determine Your A/B Test Winner

Faisal HouraniFaisal Hourani· Founder & eCommerce Growth Strategist
September 22, 20269 min read

Is your store leaking revenue?

Find out exactly where you're losing sales — takes 2 minutes.

Find Your Revenue Leaks

The four-question check that separates a real lift from a lucky week

How Do You Determine an A/B Test Winner?

Most teams call it too soon.

You determine an A/B test winner by checking three things together: statistical significance at 95%+, a sample size that hits your pre-calculated minimum per variant, and a test duration covering at least one full business cycle. A "winner" on only one of these three isn't a winner. It's a guess with a green checkmark next to it.

A dashboard showing "Variant B is winning, 97% confidence" after two days feels like proof. It usually isn't. Here's how to tell the difference, and the decision matrix WebMedic runs before calling any test for a client.

Person reviewing a web analytics dashboard with revenue and traffic charts on a laptop

What Is Statistical Significance in A/B Testing?

Significance is the number everyone stares at. It's also the most misread one.

Statistical significance is the probability that your result isn't due to random chance. When a testing tool reports a 95% confidence level, there's only a 5% chance the difference happened by luck. Most ecommerce teams use 95% as the standard threshold before declaring a winner (Shopify, 2025).

Here's what significance does not tell you: it doesn't confirm your sample size was big enough, and it doesn't confirm you ran the test long enough. A test can hit 95% significance on day two with 200 visitors per variant, and that number is still meaningless. Small samples swing wildly. One good afternoon of email traffic can push a losing variant into "significant" territory for a few hours.

This is why significance alone is the wrong stopping rule. It's one of three checks, not the whole test.

Why Does Sample Size Matter More Than Significance?

Sample size is the check most stores skip entirely.

Sample size determines whether your significance number is trustworthy in the first place. Shopify's own testing guidance sets the floor at a minimum of 1,000 visitors per variation before results are reliable enough to act on (Shopify, 2025). Below that, ordinary day-to-day variation can dress up a loser as a winner.

Run the math before the test starts, not after. Use a pre-test sample size calculator (Evan Miller's calculator and VWO's are both free) and enter your current conversion rate, the minimum lift you'd consider worth keeping (usually 10-20% relative lift for a single-element test), and your desired confidence level.

That calculator gives you a number. Don't check results before you hit it.

Close-up of hands using a calculator to work out A/B test sample size math

In our audits across 80+ Shopify stores in Malaysia and Singapore, the single most common CRO mistake we find is a test stopped the moment the dashboard turned green, often well short of the visitor count the store's own traffic volume would require for a valid read.

Here's what most stores miss: a small store doing 5,000 monthly sessions can still run valid tests. It just takes longer. Three weeks instead of five days. That's not a flaw in the process. That's the process working correctly.

How Long Should an A/B Test Run?

Duration is the third leg, and it fixes a problem significance and sample size can't catch alone.

An A/B test should run for a minimum of two weeks, long enough to cover weekday-versus-weekend behavior swings, promotions, and seasonality in one clean read (Shopify, 2025). A test that stops after four days captures one slice of your traffic pattern, not your real mix.

Weekday shoppers and weekend shoppers behave differently. Payday timing shifts buying intent. A flash sale from a competitor mid-test can distort a week's numbers. Running a full two-week cycle, even after you've technically hit significance and sample size, catches these swings before you build a permanent decision on top of them.

Fix duration and volume problems together. A low-traffic page that can't hit sample size in two weeks needs either a longer test window (go to three or four weeks) or a bigger swing (test a bigger change, not a button color, so the effect is large enough to detect faster).

Does this sound like your store? Find out where you're leaking revenue. Take the free Revenue Score. 3 minutes. Free. No pitch.

What's the Decision Matrix for Calling a Winner?

This is the part most guides skip. Here's the actual framework.

A test result only qualifies as a winner when it clears all three gates at once: 95%+ statistical significance, sample size at or above your pre-test calculation, and a minimum two-week run covering a full business cycle. Clear two gates and you have a lead, not a winner. Clear one gate and you have noise.

Score every test against this table before you touch the "declare winner" button.

Business team discussing A/B test charts and results in a meeting

Gate Threshold Pass Signal Fail Signal
Statistical significance 95%+ confidence Green in your testing tool, held stable across the last 3 days Fluctuating between 80-95% day to day
Sample size Meets or beats pre-test calculator minimum Calculator run before launch, minimum hit and logged No calculator used, or minimum not yet reached
Duration 2+ weeks, full business cycle Test includes at least one full weekend and one full week Stopped mid-week, or under 7 days total
Practical significance Lift is worth the engineering cost to ship Projected annual revenue impact calculated Lift is real but under 2-3%, not worth a rebuild

Source: WebMedic CRO audit framework, applied across 80+ Shopify stores in Malaysia, Singapore, and the UAE (2025-2026).

Three passes and one fail on "practical significance" is still worth documenting. Sometimes a statistically real, small lift is worth shipping if it costs nothing to implement (a headline swap, for instance). Three passes and a fail on significance, sample size, or duration means: keep the test running. Don't ship yet.

If a test sits at two-out-of-three for more than a week past its planned end date, that's a signal on its own. Either the effect is too small to matter, or the page doesn't get enough traffic to test that element reliably. Move on to a bigger change rather than extending the same test indefinitely. An indefinite test is not a disciplined test. It's a dashboard nobody wants to close.

What Mistakes Cause False Positive Winners?

This is where most in-house CRO programs quietly bleed revenue.

The most common cause of a false-positive A/B test winner is "peeking," checking results daily and stopping the moment a variant looks significant. Evan Miller's widely cited analysis found that repeated peeking can push a real false-positive rate as high as 26.1%, more than five times the 5% a testing dashboard implies (Evan Miller, "How Not To Run an A/B Test").

Three mistakes account for most of the bad calls we see in client audits:

  1. Peeking and stopping early. Checking the dashboard every day and hitting "declare winner" the first time it flips green. Fix: set the sample size and duration before launch, and don't look at results until both are met.
  2. Sample Ratio Mismatch (SRM). Traffic split unevenly between variants (51/49 becomes 58/42 over time) due to a bug in the testing tool or a caching issue. This silently invalidates results. Fix: check your traffic split ratio before reading any significance number. Most testing tools flag SRM automatically if you look for it.
  3. Testing during an anomaly. Running a test through a flash sale, a viral TikTok moment, or a site outage. Fix: pause or restart the test if traffic patterns shift more than 30% from baseline mid-test.

WebMedic's testing cadence with clients runs 25-day page-level tests followed by 2-week element tests (headline, UGC placement, section order), with every test logged against this exact matrix before a result gets rolled into a template. That cadence exists specifically to stop false positives from becoming permanent site changes. It's the backbone of our conversion rate optimization programs for Shopify brands in Malaysia and Singapore.

Which A/B Testing Tool Detects a Real Winner Fastest?

Not every platform calculates significance the same way, and that matters for how fast you can trust a result.

The testing tool you use changes how fast a real winner surfaces. Fixed-horizon tools (Shopify native) need a pre-calculated sample size and won't correct for peeking. Sequential and Bayesian tools (Optimizely, VWO) are built to tolerate daily checking without inflating the false-positive rate the way fixed-horizon tools do.

Tool Significance Method Best For Notes
Shopify Native (Shopify Plus) Fixed-horizon Simple page-level tests Requires manual sample size math beforehand
VWO Bayesian + frequentist Mid-size Shopify stores Built-in sample size calculator, SRM detection
Optimizely Sequential testing High-traffic stores (50k+ sessions/mo) Reduces peeking risk by design
GA4 + custom setup Frequentist Stores already deep in GA4 Needs more manual setup, more error-prone

Sources: Vendor documentation, 2025-2026.

Sequential testing tools like Optimizely are built to tolerate peeking without inflating false positives, worth the upgrade if your team checks results daily regardless of what any guide tells them to do.

Customer completing a contactless payment on a smartphone, the moment a real A/B test winner is decided

Frequently Asked Questions

How do you determine an A/B test winner?

You determine a winner by checking three things at once: 95%+ statistical significance, a sample size that hits your pre-test calculated minimum, and a test duration of at least two weeks covering a full business cycle. All three must pass together. One or two passing is not enough to call a result.

What is a good sample size for an A/B test?

Shopify's own testing guidance sets a minimum of 1,000 visitors per variation before results are reliable enough to act on. Use a pre-test sample size calculator with your current conversion rate and minimum detectable lift before launching, not after, so you know the target number going in.

How long should you wait before ending an A/B test?

Wait a minimum of two weeks so the test covers at least one full business cycle, including a weekend. Weekday and weekend shoppers behave differently, and promotions or seasonality can distort a shorter window, so a test cut off after a few days risks capturing one pattern instead of your real traffic mix.

What causes a false positive in A/B testing?

The leading cause is "peeking," checking results daily and stopping the moment a variant looks significant. Evan Miller's research on repeated significance testing found this can push the real false-positive rate to 26.1%, more than five times what a 5% significance threshold implies. Sample Ratio Mismatch and testing through traffic anomalies are the other common causes.

Is a 90% confidence level enough to call an A/B test winner?

No. 90% confidence means a 10% chance the result is random noise, double the risk of the standard 95% threshold. For a test that will permanently change a page template across your store, that extra 5-point margin is cheap insurance against shipping a false positive.

Keep Reading

Share this article

#how to determine a/b test winner #a/b test statistical significance #sample size calculator #conversion rate optimization

Ready to grow?

Find out exactly where your store is leaking revenue.

Answer a quick set of multiple-choice questions and we'll pinpoint your biggest revenue leaks — and whether we can help plug them.

Find Your Revenue Leaks

Free · No obligation · 2 minutes

Faisal Hourani

Faisal Hourani

Founder & eCommerce Growth Strategist

19 years building for the web, 9+ focused on ecommerce. Faisal founded WebMedic in 2016 to help DTC brands fix the conversion problems that hold them back. He has worked with brands across Malaysia and Singapore — from first-store launches to 8-figure scaling.

Ready to Boost Your Conversion Rates?

Book a quick strategy call. We'll analyze your store, identify your biggest revenue leaks, and show you exactly how we can plug them.

Book Your Strategy Call

Score your store

Find Your Revenue Leaks