A/B Test Prioritization Framework: Score Before You Test

Faisal HouraniFaisal Hourani· Founder & eCommerce Growth Strategist
September 23, 202610 min read

Is your store leaking revenue?

Find out exactly where you're losing sales — takes 2 minutes.

Find Your Revenue Leaks

Forty ideas, four test slots a month. The gate that decides which ones live

What Is an A/B Test Prioritization Framework?

Your backlog is a wishlist.

An A/B test prioritization framework is a system that ranks test ideas before any traffic is spent on them. WebMedic's version has two layers: an intake gate that rejects ideas with no named money metric, then a 12-point score across six dimensions. Ideas scoring 10 to 12 run next.

Say you have 40 ideas on a Notion board. A few came from a heatmap. A few came from the founder's commute. Every one is tagged "high impact."

A prioritization framework fixes that. It puts a gate in front of the board and a score behind it, so the ideas that reach a test slot got there on evidence.

This post gives you both. The gate is a form your team fills in. The score is a six-column sheet. Our list of 49 things to A/B test on a Shopify store becomes the raw material.

Sticky notes on a whiteboard, the kind of idea board an A/B test backlog starts as

Why Does Test Order Matter So Much?

Slots are the scarce thing.

Each page can host one A/B test at a time, so every test is a slot you cannot get back. On WebMedic's 25-day page-test cadence, a page yields roughly one test a month. A store with three testable pages has about three slots a month, and a wasted slot costs a month.

Overlapping tests on one page contaminate each other. Shopify's own guide calls comparing one change at a time the simplest testing method. So the page, not the idea, is your unit of capacity.

Here is the arithmetic:

Monthly test capacity = testable pages × (30 ÷ days per test)

Testable pages Days per test Tests per month
2 25 About 2
3 25 3 to 4
3 14 About 6
5 14 10 to 11

Illustrative arithmetic, not a benchmark. The cadence is WebMedic's: 25-day page-level tests and 2-week element tests. A page only counts as testable if it has the traffic for the method you propose, so check it in Evan Miller's sample size calculator first.

Now go back to the 40 ideas. At four slots a month, that is ten months of testing, and the site changes underneath the last idea long before it runs.

Most of that board will never run. Either you choose what dies, or the calendar does. Most teams choose with a score sheet, and that fails.

Two colleagues working through printed figures with a calculator beside two laptops

Why Do ICE and PIE Scores Fail on a Real Backlog?

Opinions score themselves.

ICE (Impact, Confidence, Ease) and PIE (Potential, Importance, Ease) both rank ideas on the scorer's own estimate. Neither requires evidence, a named metric, or a reject rule. WebMedic's gate adds all three, so a score of 10 out of 12 means a source backs the idea, not that someone felt strongly.

ICE and PIE are fine for a fast sort. They break when the person who proposed the idea also scores it. A confidence of 8 costs nothing to type.

Three things go wrong:

  1. Nobody names the money. "Improve conversion" survives every meeting. Then the test wins on orders, and finance asks why revenue per session fell.
  2. Cheap ideas float up. Ease carries a third of the score, so easy ideas get a head start over the ones that matter.
  3. Nothing dies. A score ranks. It does not reject. The board only grows.

The fix is not a better formula. It is a gate.

Does this sound like your store? Find out where you're leaking revenue. Take the free Revenue Score. 3 minutes. Free. No pitch.

What Is the Intake Gate, and What Does It Reject?

The gate is a form.

WebMedic's intake gate is a six-field form every idea completes before it enters the backlog: one owner, one page or component, one primary business metric, one business mode, one evidence source, and one decision rule. A blank field means the idea is not scored. It goes back to its author.

Copy this into your board as the required template:

TEST INTAKE FORM

Owner:            One name, not a team
Page/component:   One page, flow, or component
Primary metric:   Revenue, profit, CAC, or customer quality
Business mode:    Profit, acquisition, or LTV expansion
Evidence source:  Customer feedback, heatmap revenue, funnel data,
                  pricing economics, or paid-media performance
Decision rule:    What you ship, revert, iterate, or escalate, and on what result

Pick the business mode before the test runs, never after you see the result. It decides which metric counts as a win, and how to read a result against each mode gets its own post.

Which Ideas Does the Gate Reject?

Send an idea back when any of these is true:

  • The hypothesis says "improve conversion rate" without naming revenue, profit, CAC, or customer quality.
  • The idea is only a design preference.
  • The mechanism is unclear.
  • The page has too little traffic for the method proposed.
  • The team cannot measure margin and new-customer mix for an offer or price test.

The first rule kills the most ideas. Conversion rate can rise while profit falls. A test that lifts orders by leaning on discounts wins on the dashboard and loses in the bank.

What Does a Rewrite Look Like?

Same idea, before and after the gate:

Rejected: "Improve conversion rate on product pages."

Accepted: "Raise revenue per session on the hero product page by showing the shipping and returns promise beside the add-to-cart button, because exit surveys name delivery cost as the top hesitation."

The second version names a page, a component, a money metric, and a reason. Someone can now disagree with it using data, which is the point. Where does the evidence come from? Our week 1-2 diagnose sprint covers the heatmap, funnel, and survey work that feeds the evidence field.

Hands arranging index cards in rows on a wooden table

How Do You Score the Ideas That Pass Intake?

Six columns, three possible marks.

Score each surviving idea 0, 1, or 2 on six equal-weight dimensions: economic upside, evidence strength, traffic availability, implementation effort, risk, and measurement clarity. The maximum is 12. Effort and risk invert, so a 2 means low effort or low risk. WebMedic runs ideas scoring 10 to 12 next.

Dimension 0 1 2
Economic upside Low Moderate High
Evidence strength Opinion only One data source Survey plus heatmap, funnel, or economic data
Traffic availability Low Moderate High
Implementation effort High Moderate Low
Risk High Moderate Low
Measurement clarity Ambiguous Somewhat clear Clear primary metric and guardrails

Source: WebMedic scoring model.

Read the effort and risk rows twice. A 2 means low effort and low risk. It is the easiest place to get the sheet wrong, and it flips your ranking when you do.

How Do You Keep the Evidence Score Honest?

Evidence is the column teams inflate. Run three yes-or-no checks from Carl Weische's DTC test patterns before you write the number:

  • Placement: Is the change above the fold, or in an area visitors reach early?
  • Data and engagement: Do quantitative data, a known drop-off, user testing, or a prior test result back it?
  • Category: Does it change a high-leverage category, such as how social proof is shown, rather than something cosmetic like button color?

These checks carry no weights. Use them to argue for the evidence score, not to add a seventh column. If the data box stays empty, the idea is opinion only, and evidence scores 0.

What Do the Four Bands Mean?

Score Call What happens next
10 to 12 Run next Goes into the next open slot with a written brief
7 to 9 Next cycle Stays on the board and is re-scored when a slot opens
4 to 6 Refine Goes back for a sharper hypothesis or more evidence
0 to 3 Discard or monitor Comes off the board unless new data arrives

Shopper scrolling a clothing product list on a smartphone

How Do You Write a Testable Hypothesis?

One sentence, fixed shape.

A test hypothesis names four things in one sentence: the change, the audience, the metric, and the evidence. Shopify's testing guide says a good hypothesis must be measurable so you can prove it after the test (Shopify). WebMedic's template adds guardrails and a decision rule to every hypothesis.

Use this fill-in line:

If we change [page/component] from [current state] to [new state],
for [audience],
then [primary metric] should improve
because [customer, heatmap, funnel, or economic evidence].

Here is a filled-in version for a Shopify product page. It is an illustration, not a client result.

If we change the variant selector from a dropdown to visual swatches,
for first-time mobile visitors,
then revenue per session should improve
because the heatmap shows few taps on the dropdown and the
post-purchase survey keeps asking "which shade is mine?"

Primary metric:    Revenue per session
Secondary metrics: Add-to-cart rate, variant selection rate
Guardrail metrics: AOV, return rate, new-customer share
Decision rule:     Ship if revenue per session rises and no guardrail
                   worsens. Revert if it falls. Iterate if it is flat
                   but add-to-cart rises. Escalate if revenue per session
                   rises while a guardrail worsens.

The primary metric is money per visitor, not conversion rate. Conversion rate stays in the readout as context. We explain why in profit per visitor.

What Does a Scored Backlog Look Like?

Here is one board, scored.

A scored backlog ranks every surviving idea from 0 to 12 and sorts by total, with rejected ideas held outside the list. In the example below, eight of ten ideas from a hypothetical Shopify skincare store pass intake. Two run next, three wait for the next cycle, two get refined, and one is discarded.

This is an illustrative backlog, not client data. Effort and risk are scored so that 2 means low.

Rank Idea Upside Evidence Traffic Effort Risk Measure Score Call
1 Visual variant swatches on the product page 2 2 2 1 2 2 11 Run next
2 Shipping and returns promise beside add-to-cart 2 1 2 2 2 1 10 Run next
3 Bestseller cards in the main menu 1 2 2 1 2 1 9 Next cycle
4 Three-pack bundle offer on the product page 2 1 2 0 1 2 8 Next cycle
5 Founder section on the homepage 1 1 2 1 2 0 7 Next cycle
6 Reorder collection page tiles by revenue 1 0 1 2 2 0 6 Refine
7 Ingredient video on the ingredients page 1 1 0 0 2 1 5 Refine
8 Replace the cart drawer with a full-page cart 1 0 1 0 0 1 3 Discard

Two more ideas never made the table. The gate rejected them:

  • "Make the homepage feel more premium." A design preference with no mechanism and no metric.
  • "Test a higher price on the hero serum." The store cannot measure margin per session or new-customer mix, so a price test would be a guess. It comes back when cost of goods is in the reporting.

Read past the ranking. Swatches lead because the heatmap and the survey agree. The bundle has the top upside but loses three points to effort and risk. The founder section is cheap and safe, yet scores 0 on measurement because nobody can name the metric a homepage section moves cleanly.

Method follows scope. The menu cards are a sitewide, reversible change, so they suit a live test with comparable pre and post windows. The swatches are page-specific, so they get a split test with a simultaneous control. The sample size and duration math for both lives in our A/B testing guide.

That is the whole system. Gate, score, brief, slot. If you want the wider program around it, see our conversion rate optimization service and the systematic approach to ecommerce conversion optimization.

Frequently Asked Questions

What is the best A/B test prioritization framework?

The best framework requires evidence and a named money metric before an idea is scored. WebMedic's version uses a six-field intake gate, then a 12-point score across six dimensions. Ideas scoring 10 to 12 run next, and ideas scoring 0 to 3 are discarded or monitored.

How do you score A/B test ideas?

Score each idea 0, 1, or 2 on economic upside, evidence strength, traffic availability, implementation effort, risk, and measurement clarity. The maximum is 12. Effort and risk are inverted, so low effort and low risk earn a 2. Ideas that score 10 to 12 go into the next open slot.

How many A/B tests can you run at once?

You can run one A/B test per page at a time, because overlapping tests on a page contaminate each other. Monthly capacity equals testable pages multiplied by 30 divided by days per test. On a 25-day page-test cadence, three testable pages give roughly three to four tests a month.

When should you live test instead of split test?

Live test a sitewide or page-type-wide change that is low risk and easy to reverse, using comparable pre and post windows. Split test a page-specific change that could move purchase behavior materially, so a simultaneous control isolates the effect. Never live test a major landing page rebuild.

Keep Reading

Share this article

#a/b test prioritization framework #ab test prioritization #cro test backlog #test intake form #conversion rate optimization

Ready to grow?

Find out exactly where your store is leaking revenue.

Answer a quick set of multiple-choice questions and we'll pinpoint your biggest revenue leaks — and whether we can help plug them.

Find Your Revenue Leaks

Free · No obligation · 2 minutes

Faisal Hourani

Faisal Hourani

Founder & eCommerce Growth Strategist

19 years building for the web, 9+ focused on ecommerce. Faisal founded WebMedic in 2016 to help DTC brands fix the conversion problems that hold them back. He has worked with brands across Malaysia and Singapore — from first-store launches to 8-figure scaling.

Ready to Boost Your Conversion Rates?

Book a quick strategy call. We'll analyze your store, identify your biggest revenue leaks, and show you exactly how we can plug them.

Book Your Strategy Call

Score your store

Find Your Revenue Leaks