Free Shipping Threshold A/B Tests: Finding Your Number
Free shipping threshold test: how to pick candidate values, which metric decides (not conversion), how much traffic it needs and the margin guardrails.

📚 This article is part of the guide Ecommerce Checkout Optimization: The A/B Testing Playbook.
A free shipping threshold test compares two minimum order values above which the store stops charging for delivery, to find out which one delivers more profit per visitor. It is one of the few ecommerce tests where the obvious metric, conversion rate, almost always points the wrong way, because the lower threshold tends to convert better and to deliver less margin. This article is part of the complete checkout optimization playbook and covers how to pick the candidates, which metric decides, how much traffic the test needs and which guardrails keep a false winner out of production.
Why shipping decides so many purchases
Extra costs are the number one reason people abandon checkout once they already intended to buy. According to the Baymard Institute, extra costs that are too high (shipping, fees and taxes) lead with 40% of mentions among people who abandoned for a reason other than just browsing, ahead of delivery being too slow (20%), not trusting the site with card details (19%), being forced to create an account (18%) and a long or complicated checkout (17%). The same institute documents an average cart abandonment rate of 70.22%, computed from 50 studies.
That puts shipping among the first-order conversion levers, and this is where the trap lives: because shipping weighs heavily in the decision, it is easy to conclude “less shipping cost is always better” and zero out the threshold without looking at the other side of the ledger. The right question is not whether free shipping converts better, it almost always does, it is from which order value it still pays for itself.
The threshold is a number, and numbers get tested
The most common conceptual error is treating free shipping as a binary switch (on or off). In practice it is a value: above $X the customer pays nothing for delivery. That X governs three things at once:
- Conversion rate. The lower the threshold, the more people reach the benefit effortlessly, and the higher the conversion.
- Average order value. The higher the threshold, the more people add an extra item to reach it, and the larger the order.
- The shipping cost the store absorbs. The lower the threshold, the more small orders arrive with delivery paid by the store, and the thinner the margin per order.
The three move in different directions, and that is exactly why the decision cannot come from a single metric.
How to pick the candidate thresholds
Do not pick round numbers. Pick by looking at the real distribution of cart values in your store over the last few months. The practical procedure:
- Pull the order value histogram (or the cart value histogram, including abandoned ones) in bands of $20 or $25.
- Find the band with the highest density. That is where most of the movement will come from.
- Place the candidate threshold just above that dense band, at a distance one average catalog item can cover. If the typical cart closes at $130 and the average item costs $45, a $160 threshold is reachable with one more item; a $260 threshold is not.
- Pick the second candidate on the other side of the band, so the test measures a real behavior difference instead of two versions of the same value.
| Signal in the cart distribution | Candidate threshold | What to expect |
|---|---|---|
| High density just below the current threshold | Keep the threshold and test how you communicate it | Conversion gain with no margin cost, the cheapest test in this family |
| High density well below the current threshold | Test a lower threshold | Conversion rises, margin per order falls, the call depends on profit per visitor |
| Meaningful tail above the current threshold | Test a higher threshold | Average order value rises, conversion falls, and the risk is losing the small recurring buyer |
| Flat distribution with no clear peak | Test free shipping with no threshold against the current one | A legitimate hypothesis, but bring a margin guardrail from day one |
Note the first row: in many stores the best effort-to-return test is not changing the value, it is communicating the value that already exists. A progress bar showing “$23 away from free shipping” in the cart does not change the economics of the order, it changes how many people realize they are close to the benefit. Test that before touching the number, because the margin downside is zero.
The metric that decides is not conversion
This is the point that separates a well read free shipping test from a badly read one. Conversion is the metric every dashboard shows first and the one that almost always favors the lower threshold. The honest verdict comes from profit per visitor, which combines three numbers:
- the conversion rate of the variation,
- the average order value (and, inside it, the gross margin of the items),
- the shipping cost the store absorbed per order.
A variation can win on conversion, lose on average order value and still win in aggregate, or the opposite. The practical rule: decide by profit per visitor, diagnose with conversion and average order value. If your dashboard does not compute profit, the minimum acceptable approximation is revenue per visitor net of absorbed shipping, which already captures most of the trade-off. The revenue per visitor calculator helps you assemble that math.
One honest statistical caveat: the two-proportion test behind the significance calculator on this blog is built for rates (how many out of how many converted). Average order value and revenue per visitor are means of continuous, long-tailed values, and should not be read with the same test. A few very large orders move a mean on their own, and it is common for a basket size difference that looks impressive to vanish once you look at the median. Treat the conversion difference with the rigor of the proportion test and the basket difference with the caution of someone who knows the mean is being pulled by a few points.
Sizing the test
Set your current conversion rate, the minimum gain that would justify changing the threshold, and your real weekly traffic:
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
A worked example with real numbers
A store converts 2.5% of visitors into orders and gets 30,000 visits per week. The team wants to test lowering the free shipping threshold from $75 to $50.
If the ambition is to detect a 10% relative improvement in conversion (from 2.5% to 2.75%) at 95% confidence and 80% power, the math asks for 64,199 visitors per variation, which at 30,000 weekly visits takes 30 days. Aiming at 15% relative (from 2.5% to 2.875%), the requirement drops to 29,193 per variation and the test closes in about 14 days. Since lowering the threshold is a large change to the offer, aiming at 15% is defensible; aiming at 5% would ask for two and a half months to detect an effect this change would hardly produce so subtly.
Suppose the test ran to 30,000 visitors per variation and finished with 750 orders in control (2.50%, $75 threshold) against 840 in the variation (2.80%, $50 threshold). Running those four numbers through the same significance engine used across this blog: z = 2.29, p-value about 0.0222, with a 95% confidence interval for the difference between +0.04 and +0.56 percentage points, and an observed relative improvement of +12%. The interval does not cross zero, so the conversion gain is real.
Does the conversion gain pay for the shipping the store took on?
This is the second half of the math, and it is where many shipping tests stop too early. Across 30,000 weekly visits, moving conversion from 2.50% to 2.80% with an average order value of $70 generates $6,300 of incremental revenue per week (30,000 × 0.003 × 70), the same arithmetic the conversion rate impact calculator runs.
Now the cost column, which the testing dashboard never shows: dropping the threshold from $75 to $50 makes the store absorb shipping on the entire band of orders between those two values, which used to pay for delivery. If 900 orders per week fall in that band, the absorbed cost is $8,100 per week at an average of $9 per shipment, or $4,500 at $5 per shipment.
Put the confidence interval and the cost column in the same table and the decision stops being a matter of taste:
| Reading of the gain | Extra orders per week | Incremental revenue | Net at $9 absorbed | Net at $5 absorbed |
|---|---|---|---|---|
| Lower bound, +0.04 pt | 12.9 | $903 | −$7,197 | −$3,597 |
| Point estimate, +0.30 pt | 90 | $6,300 | −$1,800 | +$1,800 |
| Upper bound, +0.56 pt | 167.1 | $11,697 | +$3,597 | +$7,197 |
Read the table honestly and the same significant result supports two opposite decisions depending on one operational number you already have: your average absorbed shipping cost. And this is still revenue, not profit: subtract gross margin on the goods and every cell moves down. None of this appears on the conversion chart, which is exactly why absorbed shipping cost has to enter the reading at design time, not at the following meeting.
Mandatory guardrails
| Guardrail | Why watch it | Warning sign |
|---|---|---|
| Margin per order | The lower threshold transfers the shipping cost to the store | Conversion rises and margin falls enough to cancel the gain |
| Average order value and median | A high threshold pushes the basket up; a low one lets it drop | Mean rises and median stays flat: a few large orders are fooling the reading |
| Absorbed shipping cost per order | It is the cost column the test creates | Cost per order grows faster than incremental revenue |
| Return rate | Impulse items added to hit the threshold come back more often | Returns rise in the higher threshold variation |
| Regional mix | Shipping cost varies a lot by region, so the same threshold has different economics | Conversion rises in one region while margin sinks in another |
The last row deserves special attention in any country with long delivery distances: a single national threshold means very different margins by region, and a test read only in aggregate can approve a threshold that is profitable in dense metro areas and destructive in remote ones. Declare the reading segments (region, basket band, new vs returning) before running, and treat any cut discovered afterwards as a hypothesis for the next test, never as a conclusion. That is the same discipline that avoids Simpson’s paradox, detailed in common A/B testing mistakes.
Traps specific to free shipping tests
- Changing the threshold mid-test. It looks harmless (“we only moved it $10”) and invalidates everything: visitors in the first weeks saw a different offer from visitors in the last weeks, inside the same variation.
- Testing during a promotional campaign. Black Friday, an aggressive coupon or a parallel free shipping promotion contaminate the comparison, because the customer is already seeing another discount on screen.
- Ignoring returning customers. A returning customer knows the old threshold and reacts to the change differently from a first-time visitor. If your returning base is large, that cut has to be declared upfront.
- Reading the result in week one. The average order value effect shows up fast; the repurchase and return effects show up weeks later, and they decide whether the new threshold holds.
- Confusing communication with economics. Testing the cart progress bar is a communication test with no margin cost. Testing the threshold value is an economics test. Both are valid, and mixing them in the same variation makes it impossible to know which one caused the result.
Make this automatic with Donnu
A free shipping test needs three things at once: enough sample for a low base rate, a reading by profit per visitor instead of isolated conversion, and an honest confidence interval instead of a winner declared on day four.
Donnu A/B delivers that on your site: a light snippet that does not slow the cart down, automatic sample sizing and Bayesian statistics that do not invent certainty. Start a 14-day free trial and find your threshold with data, instead of copying the number a competitor picked on intuition.
Read also: Ecommerce Checkout Optimization: The A/B Testing Playbook · Cart Abandonment A/B Testing · Product Page A/B Testing · Leia em português: Teste A/B de Frete Grátis
References
- Baymard Institute. Cart Abandonment Rate Statistics. Average of 50 studies on cart abandonment and the distribution of abandonment reasons, including the weight of extra costs. baymard.com/lists/cart-abandonment-rate.
- Baymard Institute. Checkout Usability Research. Base of usability issues documented in large-scale ecommerce research. baymard.com/research/checkout-usability.
- Kohavi, R., Tang, D. and Xu, Y. Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press, 2020, on guardrail metrics and reading long-tailed metrics. Companion material at experimentguide.com.
- Google Search Central. A/B testing best practices for Search. On running site tests without harming indexing, which applies when threshold variations change page content. developers.google.com/search/docs/crawling-indexing/website-testing.
Frequently asked questions
- Which metric decides a free shipping threshold test?
- Profit per visitor, not conversion rate. A lower threshold almost always converts better and almost always delivers less margin per order, because the store starts absorbing shipping on smaller baskets. A higher threshold usually lifts average order value and pushes conversion down. Since the two metrics move in opposite directions, only the combination of both with your real shipping cost answers which value is better for the business. Use conversion and average order value as diagnosis, profit per visitor as the verdict.
- How do I pick the candidate threshold values to test?
- By looking at the real distribution of cart values in your store, not by picking round numbers on intuition. Pull the order value histogram for the last few months and find where the density is; a threshold sitting just above a dense band is the one with the best chance of pushing orders up without alienating most of your base. Testing 149 against 199 when almost every cart closes below 120 burns weeks of traffic in a range where almost nobody is.
- How much traffic does a free shipping test need?
- A lot, because the base order rate is low. With the math used across this blog, a store converting at 2.5% that wants to detect a 10% relative improvement needs about 64,199 visitors per variation. Aiming at 15% relative drops the requirement to about 29,193 per variation, which at 30,000 weekly visits closes in around 14 days. The average order value effect is usually larger and faster to see than the conversion effect, but you still have to wait for the full sample before deciding.
- Is free shipping for everyone a good idea?
- It depends entirely on your margin and your average shipping cost, which is exactly why it deserves a test instead of an opinion. Free shipping with no threshold removes the top stated reason for abandonment and at the same time removes any incentive for the customer to grow the basket, while making the store absorb delivery on small orders, which usually carry the worst margin. It is a legitimate hypothesis, as long as it enters the experiment with a margin guardrail and not only a conversion metric.
- Can I use a significance calculator to compare average order value between variations?
- Not directly. The significance calculator on this blog runs a two-proportion test, built to compare rates (how many out of how many converted). Average order value and revenue per visitor are means of continuous, skewed, long-tailed values and need different treatment, such as a test for means or resampling. Use the proportion test for conversion and treat the basket size difference with the caution of someone who knows a handful of large orders can move a mean on their own.