CRO

Trust Badges A/B Testing: When They Actually Help

Trust badges A/B testing: where to place them in checkout and SaaS, what to measure, why results vary, and how not to credit a famous logo for a win.

Flat illustration of a dark green shield with a padlock, next to a shopping basket, a rosette seal and a balance scale, on a mint green background

Trust badges are visual marks telling visitors that a third party has checked something: the site certificate, the payment method, the returns policy, the store’s reputation or, in SaaS, the company’s security controls. They help when they answer a real doubt at the moment it comes up, and almost always with a small effect that needs a large sample and the right metric. In the worked example in this guide, a recognized badge next to the card fields lifts completed orders by 3.1 percent with a p-value of 0.0006, and a generic padlock with truthful copy in the same spot reaches 2.7 percent: the difference between the two is not significant. This guide is part of our complete guide to conversion rate optimization (CRO) and covers the badge types, placement, what the research shows, what to measure, the risk of unknown or fake badges, and two worked examples, one ecommerce and one SaaS.

What trust badges are and why their effect varies so much

A trust badge is outsourced reassurance. Instead of the store saying “we are safe”, an outside brand shows up saying, in theory, that it checked. The catch is “in theory”: visitors rarely know what each badge verifies, and the Baymard Institute sums this up in an observation that runs through years of usability testing: the average user judges a site’s security by gut feeling, driven by how much they trust the brand and how visually robust the page looks, not by its actual technical security.

That gives this guide its first rule: a badge works on perception, so its effect depends on how much doubt exists and on who is looking. A well-known store has little doubt to remove. According to Baymard, when testing sites from large brands, participants show few security concerns even with minimal visual reinforcement, while on newer or more niche sites concerns come up easily when there are no visual cues.

The five badge types and what each one proves

type examples what it actually verifies what visitors tend to read into it main risk
Site certificate (SSL/TLS) certificate authority seal that the connection is encrypted and, depending on the certificate, who the organization is “this site is safe” nearly every site already has HTTPS; the badge says nothing about the store’s honesty
Payment card networks, digital wallets, payment providers that the method is accepted “my money is protected” a network that is not actually accepted in that flow
Guarantee and returns “free 30-day returns”, “satisfaction guaranteed” a policy the store sets itself “I am not taking a risk” hidden conditions; a promise support does not keep
Reputation and reviews rating on a review platform, complaints bodies aggregated third-party opinion “other people trusted them” stale ratings, or a platform the visitor does not recognize
B2B compliance SOC 2, ISO/IEC 27001, GDPR or LGPD controls audited by a third party, for a defined scope and period “my data will be protected” a logo without a current report or certificate; a scope narrower than the product

The middle column and the one next to it rarely match, and that gap is where the A/B test lives. A certificate seal proves very little to an ordinary shopper; a guarantee badge proves only what the store commits to; a compliance badge proves a lot, but only to someone who knows how to read it.

What a badge verifies versus what visitors recognizeTwo-axis map. Horizontal axis: how much the badge actually verifies, from little to a lot. Vertical axis: how well an ordinary shopper recognizes it, from little to a lot. A famous security brand sits high and in the middle. Payment network logos sit high and to the left. The store’s own padlock icon sits in the middle and to the left. A certificate seal from a lesser-known brand sits low and in the middle. SOC 2 and ISO 27001 sit low and to the right for ordinary shoppers, and move up for IT buyers.the most recognized badge is not always the one that verifies mosthow much the badge actually verifieslittlea lotrecog-nitionpayment network logoonly says it is acceptedfamous security brandcertificate plus brandstore’s own padlockno third party checkedcertificate seal from alesser-known brandSOC 2, ISO 27001ordinary shopperSOC 2, ISO 27001IT buyer
Qualitative positions, meant to guide hypotheses. The same compliance badge that means nothing to an ordinary shopper is decisive for whoever runs a vendor security review, which is why a badge’s effect changes so much from one audience to another.

What the evidence on trust badges shows (and what it does not)

Almost all public evidence on badges is about stated perception, not conversion measured in an experiment. That does not make it useless: it tells you where the doubt is and which brands the audience recognizes. It does not tell you how much your conversion rate will move.

Baymard: the doubt shows up at the card fields

In its article on perceived security during checkout, Baymard brings together four observations that work as a starting point for any hypothesis:

Which badge people say they trust: brand beats function

From 2013 to 2023, Baymard asked “Which badge gives you the best sense of trust when paying online?”. In the 2013 round, with 2,510 responses from US adults, 49 percent chose “Don’t know or no preference”. Among those who picked a seal, Norton took about 36 percent and McAfee about 23 percent, both from antivirus brands. The Norton seal was itself a certificate (SSL) seal, but the other certificate seals on the list, which also have a real technical function, scored well below.

The 2023 update tells, in Baymard’s words, largely the same story: Norton leads, followed by three business trust seals, with several certificate seals scoring well below. And there is the finding that matters most for testers: since 2016 the surveys have included a seal made up by the researchers, meaning nothing at all, and it performed significantly better than certificate seals from established vendors, except Norton. Baymard’s takeaway is that, beyond a very recognizable brand, which seal you pick matters less than having some visual reinforcement near the card fields.

A 2016 CXL Institute survey of 2,100 people in the US found the same pattern (the more familiar the brand, the higher the perceived security) and added audience differences: people aged 20 to 30 picked the Google badge more often than average (19 versus 15 percent), and only 3 percent of those over 50 did. The authors themselves warn that what people say in a survey is not what they do at checkout.

source design what carries over to your test what does not
Baymard, perceived security in checkout usability testing and abandonment survey where the doubt appears (card) and how to reinforce (encapsulation plus 1 or 2 icons) any conversion lift figure
Baymard, seal surveys 2013 to 2023 stated choice between seals, US adults a recognized brand matters; a generic icon already helps the ranking holds for that audience and those brands
CXL Institute, 2016 survey of 2,100 people in the US the convincing badge changes with age and gender stated familiarity is not conversion
vendor case studies tests published by whoever sells the badge or the tool hypothesis ideas effect size, because of publication bias

On that last row: published “the badge lifted conversions by X percent” cases are usually the tests that worked, told by someone with a stake in the result. The ones that moved nothing never became blog posts. Use them to generate hypotheses, never as effect estimates. Baymard itself recommends testing badges and visual reinforcement with your own audience and taking other people’s case studies with a grain of salt.

The browser padlock changed, and so did the certificate seal

Two changes in 2023 aged a lot of what used to be written about security badges:

Timeline of the browser padlock and the certificate sealTimeline with four milestones. 2013: only 14 percent of the top one million sites supported HTTPS, according to Google. 2021: only 11 percent of participants in a Chrome study understood the precise meaning of the lock icon. May to September 2023: Chrome announces replacing the lock with a tune icon, with over 95 percent of page loads on Windows already over HTTPS; the switch is scheduled for Chrome 117, due in September 2023. October 17, 2023: DigiCert automatically replaces the Norton site seal image with the DigiCert seal.HTTPS became the default, and the security symbol stopped standing out20132021May to Sep 2023Oct 17, 202314% of the top 1Msupported HTTPS11% understood whatthe lock icon meantover 95% on HTTPS;Chrome 117 swaps thelock for a neutral iconDigiCert seal replacesthe Norton seal imagefirst three milestones from the Chromium Blog; the last from Akamai’s changelog on DigiCert
The seal that topped the trust surveys stopped being displayed as such, for example on sites with Akamai-managed DigiCert OV and EV certificates. Any old “most trusted seal” ranking needs to be reread with that in mind.

The practical consequence: “the site has HTTPS” is no longer a differentiator. A certificate seal communicates, at most, that the store bothered to display one. That does not make it useless (visual reinforcement near the card still helps perception, according to Baymard), but it changes the hypothesis: you are testing visual reinforcement and brand recognition, not “security”.

Where to place trust badges: checkout, forms and pricing pages

The rule that holds across contexts is proximity: the badge goes where the doubt appears, and the doubt appears where the visitor hands something over (a card, personal data, company data, a subscription commitment). A badge in the footer is decoration; the same badge next to the card field is a hypothesis.

Where the badge works: ecommerce checkout and SaaS pricing pageTwo side-by-side sketches. Left, a checkout payment step: name and address fields with no emphasis, and the card section with a highlighted background, a padlock icon and a seal inside it, above the place order button. Right, a SaaS pricing page with three plans; on the Enterprise plan, near the request a demo button, a row with SOC 2 and ISO 27001 badges and a link to the trust center.the badge works where the visitor hands something overcheckout: payment stepfull nameshipping addresscard numberexpirysecurity codeencrypted payment✓place orderSaaS: pricing pageStarterProEnterpriseget a demoSOC 2 Type II · ISO/IEC 27001:2022see the report in our trust center
Left, what Baymard describes: encapsulation reserved for the card area, with one or two icons inside it. Right, the B2B equivalent: compliance proof next to the action that triggers a security review, with a path to the document behind it.
context where to test the badge doubt it answers primary metric follow-up read
Ecommerce checkout inside the card area, next to the place order button “can I trust this site with my card?” completed orders among checkout starters paid (not declined) orders, chargebacks
Product page near the price and buy button (guarantee, returns, payment methods) “what if it does not fit? can I pay my way?” add to cart and orders returns within 30 days
Cart next to the total and the checkout button “are the final price and returns clear?” checkout starts completed orders
SaaS demo form beside the work email and company fields “what will they do with my data?” form submissions qualified meetings
SaaS pricing page on the Enterprise plan and near the demo or trial button “will this pass my company’s security review?” demo requests or trial starts qualified opportunities, sales cycle
Self-serve signup near the create account button and the card field, if any “will I be charged without noticing?” accounts created activation, paid accounts

Two notes on the table. First, a money-back guarantee is the badge most likely to move the follow-up metric: it can lift orders and returns together, so a 30-day guardrail is mandatory. Second, some markets already give buyers a statutory withdrawal right (Brazil, for instance, grants 7 days for purchases made outside a store under article 49 of its Consumer Defense Code); a badge that repeats the law describes the law, not a differentiator, and should be presented that way. The guides on checkout optimization, product pages, cart abandonment, SaaS demo forms and SaaS pricing pages cover the rest of each step.

What to measure in a badge test

A badge is a small element, and small elements have small effects. Three measurement decisions separate a useful test from one that “wins” by chance.

1. Step completion, not badge clicks

Almost nobody clicks a badge. Clicks are a diagnostic (they confirm the element was seen and whether anyone wanted to check what it says), never the primary metric: a badge can get few clicks and still change the decision, and a badge can get lots of clicks because it looks like a button and distracts from the order. The primary metric is completion of the step the badge is meant to unlock: completed orders, submitted forms, requested demos.

2. Measure among people who reached the step, not across the whole site

If the badge only appears on the payment step, only people who reach that step can be affected. Measuring conversion over all site visitors mixes in thousands of people who never saw the badge and drowns the effect in noise. The sample difference is huge, as the math from the blog’s calculator engine shows for a 3 percent relative effect, 95 percent confidence and 80 percent power:

where you measure baseline visitors per variation why it changes
completion among checkout starters 45% 21,371 only people who can see the badge are counted
orders across all site visitors 2.7% 637,718 same relative effect, diluted in a baseline nearly 17 times smaller

Both rows describe the same scenario (6 percent of visitors start checkout and 45 percent of them finish, which gives 2.7 percent). For exposure to work this way, randomization or counting has to start at the step where the badge appears, or the analysis has to filter to people who got there, with the same rule in both arms.

3. One variable at a time, and an arm to separate brand from form

“The badge won” can mean four things: the encapsulation won, the icon won, the badge brand won, or the written promise won. A two-arm test (no badge versus a famous badge) cannot tell them apart. If the question is “is a branded badge worth paying for?”, the design needs a third arm with neutral reinforcement in the same spot and at the same size. The first worked example does exactly that.

Three-arm design to separate visual reinforcement from badge brandThree boxes. A, control: card area with no emphasis and no icon. B, neutral reinforcement: encapsulated area with a generic padlock and the copy encrypted payment. C, recognized badge: the same encapsulated area with a well-known brand’s badge in the same place and size. Lines show the comparisons: B versus A measures the reinforcement; C versus A measures the branded badge; C versus B measures what the brand adds.only C versus B answers “does the badge brand matter?”A · controlcard area styled likeevery other field,no iconB · neutral reinforcementencapsulated area,generic padlock andtruthful copyC · recognized badgesame encapsulation,well-known brand badgesame place and sizeB versus A: reinforcement effectC versus A: branded badge effectC versus B: what the brand adds
With three arms and two comparisons against control, each comparison’s significance level drops to 2.5 percent (Bonferroni correction), and the sample per arm goes up. That is the price of a sharper question.

The risk of unknown and fake badges

There are two different risks, and only one of them shows up in the test report.

An unknown badge may do nothing, or hurt. A badge the audience does not recognize is, in practice, one more image near the button. Baymard recorded certificate seals from established vendors scoring below a made-up seal; CXL recorded stated trust in each seal shifting with age. A badge that looks amateurish, or breaks the layout on mobile, can backfire: Baymard observed participants suspecting a hack when they hit layout bugs on the payment step. The test does measure this risk.

A fake or expired badge: the test cannot see it, regulators can. A badge claiming a check that never happened can “win” the test, because the report measures behavior, not truthfulness.

jurisdiction rule or case what it says, in short what it means for a test
United States FTC, Guides Concerning the Use of Endorsements and Testimonials (16 CFR 255), updated June 29, 2023 the name or seal of an organization can be an endorsement; an organization’s endorsement must come from a process reflecting its collective judgment and, if it is presented as expert, using experts or suitable standards a badge with no real verification is not a testable variation
United States FTC v. ControlScan, 2010 according to the FTC, seals conveying verification were handed out with “little or no verification” and showed a current date stamp, although no site was reviewed daily (some weekly, some with no ongoing review); under the settlement, the founder gave up 102,000 dollars and the company had to ask sites to take the seals down a “verified today” badge needs a check today
United States FTC and TRUSTe, 2014 according to the FTC complaint, from 2006 to January 2013 the promised annual recertification was not done in over 1,000 instances; the settlement included a 200,000 dollar payment a badge without the periodic review it promises can mislead
United States FTC, Guides for the Advertising of Warranties and Guarantees (16 CFR 239.3) use “satisfaction guarantee” or “money back guarantee” only if the full purchase price is refunded on request, with material limitations clearly disclosed a guarantee variation needs its conditions visible
Brazil Consumer Defense Code, articles 37, 38 and 67 prohibits advertising that is wholly or partly false, including by omission; the advertiser bears the burden of proof; running advertising one knows or should know is misleading carries three months to one year of detention plus a fine badges, guarantees and certifications on display must be provable

For SaaS, the rule is the same with more technical names. The AICPA offers SOC logos to service organizations that have received at least one SOC 1, SOC 2 or SOC 3 report from a licensed, independent CPA. A SOC 2 report is generally restricted use, while a SOC 3 can be distributed freely, and a Type 2 report evaluates how controls operated over a period, not just their design on a date. For ISO/IEC 27001, ISO points you to the certification body that issued the certificate before using any logo, and asks for the full reference, as in “certified to ISO/IEC 27001:2022”. A “SOC 2” badge without a report, or an “ISO 27001” badge without a current certificate covering the product, is the B2B version of a fake seal. This is context, not legal advice.

Worked example 1: a checkout badge, with an arm that isolates the brand

Scenario (illustrative). A fashion store with a little-known brand gets 30,000 checkout starts a week, and 45 percent of them become orders. The payment step currently styles the card fields like any other field. The hypothesis comes from Baymard: encapsulating the card area and placing a reinforcement inside it increases completion. The business question is twofold: does reinforcement work? And is a recognized-brand badge worth paying for, or does a generic padlock with truthful copy do the job?

Design. Three arms with equal traffic: A (control), B (encapsulated area, generic padlock and the copy “encrypted payment”, which is true), C (the same area with the seal of the certificate the store actually holds, in the same place and size). Unit: visitor who starts checkout. Primary: completed order. Diagnostic: interaction with the badge. Guardrail: orders declined by fraud screening. Pre-declared segment: new and returning customers. Planned comparisons: B versus A and C versus A, each at 2.5 percent (Bonferroni); C versus B as a descriptive read.

Sample size. The team wants to detect a 3 percent relative lift. In the calculator below, enter: current conversion rate 45; minimum detectable effect 3, relative; confidence 95; power 80; weekly visitors 30000; two-sided.

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

The calculator shows 21,371 visitors per variation, 42,742 in total and 10 days. Those 10 days assume two variations; with three arms, the same 21,371 per arm take 15 days. The calculator also works at 95 percent: with the correction for two comparisons (2.5 percent each), this guide’s math, using the same formula, calls for 25,881 per arm, 19 days with three arms. The team runs three full weeks, 21 days, which gives 30,000 checkout starts per arm and about 86 percent power to detect a 3 percent effect in each comparison.

Sensitivity to the minimum effect shows why badge tests are high-traffic tests (two variations, 95 percent, 80 percent power, 30,000 a week):

relative effect you want to detect visitors per variation total days
1% 192,031 384,062 90
2% 48,048 96,096 23
3% 21,371 42,742 10
5% 7,705 15,410 4

The first-week peek. At 10,000 per arm, C had 4,680 orders against 4,500 in control: plus 4.0 percent, p-value 0.0106. It looked like a bigger lift than planned. In weeks 2 and 3, with 20,000 per arm, C versus A came in at 9,240 against 9,000, plus 2.7 percent. A stronger first week does not justify a decision, and the cost of stopping there is covered in the peeking problem.

The result after 21 days.

arm checkout starts orders completion extra orders vs A
A, control 30,000 13,500 45.00%
B, neutral reinforcement 30,000 13,860 46.20% 360
C, recognized badge 30,000 13,920 46.40% 420

The split is clean (30,000 in each arm does not trip the SRM checker). Now paste the first comparison into the significance calculator: control with 30000 visitors and 13500 conversions; variation with 30000 and 13920; confidence 95.

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

The calculator shows 45.00% versus 46.40%, a relative lift of +3.1%, a p-value of 0.0006, a 95 percent interval for the difference of +0.6% … +2.2% (pp) and Significant winner · B wins (the calculator calls whatever sits in the second column B; here it is arm C). With more decimals, the lift is 3.11 percent, the p-value is 0.000577 and the interval runs from plus 0.60 to plus 2.20 percentage points, roughly plus 1.3 to plus 4.9 percent in relative terms. The p-value is below 0.025, so C also clears the corrected bar.

Switch the variation to arm B (30000 and 13860): the calculator shows 46.20%, a lift of +2.7%, a p-value of 0.0032, an interval of +0.4% … +2.0% (pp) and the same verdict. With more decimals, roughly plus 0.9 to plus 4.4 percent relative. Also below 0.025.

The read that walks into the trap. “The recognized badge won with 3.1 percent, the biggest lift: let’s buy the branded seal.” Now compare C against B in the same calculator, with B as control: 30000 and 13860 versus 30000 and 13920. The calculator shows 46.20% versus 46.40%, a lift of +0.4%, a p-value of 0.6233, an interval of -0.6% … +1.0% (pp) and Not significant yet.

Confidence intervals for the three badge test comparisonsThree 95 percent intervals on a relative effect axis from minus 2 to plus 6 percent. Neutral reinforcement versus control: plus 2.7 percent, from plus 0.9 to plus 4.4. Recognized badge versus control: plus 3.1 percent, from plus 1.3 to plus 4.9. Recognized badge versus neutral reinforcement: plus 0.4 percent, from minus 1.3 to plus 2.2, crossing zero.both reinforcements win; the brand showed no difference-2%0%+2%+4%+6%B versus Aneutral reinforcement+2.7%, p = 0.0032C versus Arecognized badge+3.1%, p = 0.0006C versus Bwhat the brand adds+0.4%, p = 0.6233approximate relative interval (difference divided by the reference arm’s rate); illustrative scenario
The C versus B interval runs from about minus 1.3 to plus 2.2 percent relative: the test does not rule out that the brand adds something, but it does not show that it does.

What the result says, and what it does not. Visual reinforcement near the card worked, with or without a brand, which is the pattern Baymard describes. The recognized badge did not separate from the generic padlock. And “did not separate” has a size: with 30,000 per arm and a rate near 46 percent, the test had only about 20 percent power to detect a 1 percent relative gap between C and B; seeing it would take 182,925 per arm. The honest conclusion is “if the brand adds anything, it adds little”, and that goes into the cost case for the badge.

The pre-declared cut: new versus returning (C versus A).

segment A: orders / starts C: orders / starts lift p-value approximate relative interval
new customer 7,200 / 18,000 (40.00%) 7,560 / 18,000 (42.00%) +5.0% 0.0001 +2.5% to +7.5%
returning customer 6,300 / 12,000 (52.50%) 6,360 / 12,000 (53.00%) +1.0% 0.4379 -1.5% to +3.4%

In the calculator, 18000 and 7200 versus 18000 and 7560 show +5.0%, p-value 0.0001 and an interval of +1.0% … +3.0% (pp); 12000 and 6300 versus 12000 and 6360 show +1.0%, p-value 0.4379, -0.8% … +1.8% (pp) and “Not significant yet”. The direction fits Baymard’s observation (people who already know the store need less reinforcement), but the right question is whether the two effects differ from each other, and the calculator does not do that math: the gap between the absolute lifts is 1.5 percentage points, with a p-value of 0.0697 in the interaction test this guide calculated. That is a signal, not proof. More in heterogeneous treatment effects.

What to do with it. Ship the visual reinforcement to everyone (it costs almost nothing and hurt no one), keep the seal of the certificate the store already holds if it is cheap, and do not buy a pricier branded badge on the strength of this test. The diagnostic backs that read: only 1.7 percent of visitors in arm C hovered over or tapped the badge; the effect came from people who never checked what it said.

Worked example 2: SOC 2 and ISO 27001 badges on a SaaS pricing page

Scenario (illustrative). A B2B SaaS company holds a current SOC 2 Type II report and an ISO/IEC 27001:2022 certificate, but only mentions them on a security page. The pricing page gets 10,000 visitors a week, and 3.0 percent request a demo. The hypothesis: a row with both badges and a link to the trust center, next to the Enterprise plan and the demo button, reduces the “will you pass our security review?” doubt.

The planning mistake. The team decided to run four weeks because that was what the quarter allowed. Go back to the sample size calculator and enter: rate 3; minimum effect 10, relative; confidence 95; power 80; weekly visitors 10000; two-sided. It shows 53,211 per variation, 106,422 in total and 75 days. Four weeks give 20,000 per arm.

Days needed by minimum effect on the SaaS pricing pageHorizontal bars of test duration at 10,000 visitors a week and a 3 percent baseline. 5 percent effect: 292 days. 10 percent: 75 days. 15 percent: 34 days. 20 percent: 20 days. 30 percent: 10 days. A vertical line marks the 28 days the test actually ran.at a 3% baseline, a small effect costs months5% relative292 days10% relative75 days15% relative34 days20% relative20 days30% relative10 daysthe test ran 28 days95% confidence, 80% power, two-sided, two variations; scale of 1.5 pixels per day
A badge rarely moves demo requests by 30 percent. With 28 days, the test could only reliably see effects of about 16 percent relative or more.

The result. In the significance calculator, enter 20000 and 600 for control and 20000 and 642 for the variation. It shows 3.00% versus 3.21%, a lift of +7.0%, a p-value of 0.2260, an interval of -0.1% … +0.5% (pp) and Not significant yet. With more decimals, the interval runs from minus 0.13 to plus 0.55 percentage points, roughly minus 4.3 to plus 18.3 percent relative.

The wrong read, in both directions. “Badges don’t work” is wrong: the test had about 40 percent power to detect a 10 percent relative lift (this guide’s math, same formula), and the smallest effect 20,000 per arm can see with 80 percent power is about 16 percent. “Badges lifted demos by 7 percent” is wrong too: the interval includes zero and includes an 18 percent gain. The test does not know. And computing “observed power” from this result does not help, for the reasons in observed power.

The tempting cut. The pre-declared segment was “visitors who opened the Enterprise plan”:

segment A B lift p-value read
opened the Enterprise plan 180 of 3,000 (6.00%) 219 of 3,000 (7.30%) +21.7% 0.0433 below 0.05, with about 46% power for 20%
all other visitors 420 of 17,000 (2.47%) 423 of 17,000 (2.49%) +0.7% 0.9167 nothing

In the calculator, 3000 and 180 versus 3000 and 219 show +21.7%, p-value 0.0433 and an interval of +0.0% … +2.6% (pp): the calculator rounds the lower bound to zero, but with more decimals it is plus 0.04 percentage points (the whole interval sits above zero, just barely, which is why the p-value lands just under 0.05), and the relative interval runs from about plus 0.7 to plus 42.7 percent. The direction makes sense (the security doubt lives with people evaluating the corporate plan), but with three reads (total and two segments) and a p-value of 0.0433, this is a hypothesis, not a result. See multiple metrics and FDR.

What to do with it. Keep the badges, because they are true, cheap and answer a question corporate buyers ask anyway, without announcing any lift. And, if the decision matters, run the right test: only for people who open the Enterprise plan, with a 6 percent baseline and a 20 percent minimum effect, the calculator asks for 6,719 per variation; at 1,500 of those visitors a week, that is 63 days. The metric that decides for that audience is qualified opportunities, which take weeks to mature, as the SaaS demo form guide explains.

How to read a badge test without falling for “it won because it’s famous”

A badge test that wins invites a ready-made story: the brand is well known, so it builds trust, so it converts. The story may be true, but a two-arm test does not support it. Six questions before you credit the win:

  1. What changed besides the badge? Encapsulation, whitespace, nearby copy, button position. If they changed together, the win belongs to the bundle.
  2. Is there a neutral arm? Without one, “brand” and “visual reinforcement” are confounded.
  3. Does the effect show up where the doubt lives? A bigger effect among new customers and on little-known brands fits the hypothesis; an equal effect everywhere suggests another explanation.
  4. Is the metric step completion? Badge clicks do not count as a win.
  5. Did the test have power for the effect you are celebrating? If the sample per arm fell short of what the calculator asks for that effect, a significant 2 to 3 percent lift is likely overstated.
  6. Was the first week stronger? A new badge gets attention because it is new; see novelty effect.

Pre-launch checklist

  1. Truth first. Every badge shown maps to a current check, certificate or report whose scope covers what you sell, and every guarantee shows its conditions.
  2. Doubt and audience mapped. Where does the visitor hand something over and hesitate, and who hesitates? SOC 2 on a clothing store checkout answers nobody’s question.
  3. Proximity. One or two badges in the doubt zone, not a wall of logos in the footer.
  4. Sample sized for a small effect. Plan for 2 to 5 percent relative, not 20, with a correction when you add a neutral arm.
  5. Guardrail for guarantees: returns and refund requests within 30 days.
  6. Layout check of the badge on mobile and major browsers before you start the test.

Automate this with Donnu

The specific pain of testing trust badges is that the effect is small, the business question often needs more than two arms, and the temptation to read segments after the fact is strong. A few configuration choices make those mistakes less likely.

In Donnu, a conversion goal can be confirmed from your own server, in addition to clicks, form submissions and page visits, which measures the order that was actually paid instead of the click on “place order”. On the Pro plan, an experiment can include a variation C, which supports the control, neutral reinforcement and branded badge design from the first example, and can be restricted by device, traffic source and new or returning visitor, which helps you declare the audience before running. The report is Bayesian, warns when the visitor split drifts from what you configured, and only declares a winner with at least 200 visitors per variation and 7 days of testing.

What stays on you: making sure the badge is true, sizing for a small effect and waiting for the sample. For that, the sample size calculator, the significance calculator and the statistical power calculator run the math in this guide with your numbers, for free.

References

Read next: Ecommerce checkout optimization · Social proof A/B testing · Product page A/B testing · SaaS pricing page optimization · SaaS demo request form optimization · Cart abandonment A/B testing · Testing multiple variants · Leia em português

Frequently asked questions

Do trust badges increase conversions?
They can, mostly where visitors hesitate and the brand is not well known. The Baymard Institute observes in usability testing that security concerns surface when people reach the card fields, that big brands need less visual reinforcement, and that 19 percent of users abandoned a checkout in the past three months because they did not trust the site with their card (2025 survey, 1,026 respondents). That is evidence about perception, not a promised lift: the only way to know for your site is to test, measuring completed orders.
Where should a security badge go on the checkout page?
Next to the card fields, not just in the footer. Baymard recommends visually encapsulating the card section (border, background, shading) and observes that one or two icons inside that area work well to reinforce perceived security. In SaaS, the equivalent spot is next to the demo request button or the Enterprise plan, where the data security question comes up.
What metric should a trust badge test use?
Completion of the step the badge is meant to unlock, measured among people who reached that step. In ecommerce, completed orders among checkout starters; in SaaS, demo requests or trial starts, with qualified opportunities as the follow-up read. Badge clicks are a diagnostic: in this guide, only 1.7 percent of visitors interacted with the badge, while completed orders rose 3.1 percent with a p-value of 0.0006.
Does a well-known badge beat a generic padlock?
Not necessarily, and a two-arm test cannot tell you. In Baymard surveys, the Norton seal led on stated trust, but a seal the researchers made up scored better than almost all certificate seals from established vendors. In the worked example here, both the recognized badge and a padlock with truthful copy beat the control, but the gap between the two is 0.4 percent with a p-value of 0.6233. To credit the win to the badge brand, you need a comparison arm with a neutral icon in the same spot.
How much traffic does a trust badge test need?
It depends on where you measure. At 45 percent completion among checkout starters, 95 percent confidence and 80 percent power, detecting a 3 percent relative lift takes 21,371 checkout starts per variation. Measured over all site visitors, at 2.7 percent orders per visit, the same effect takes 637,718 per variation. On a SaaS pricing page with a 3 percent demo request rate, detecting 10 percent relative takes 53,211 per variation, 75 days at 10,000 visitors a week.
Can I show a badge without the certification?
No. In the United States, the FTC Endorsement Guides (16 CFR 255) say the name or seal of an organization can be an endorsement, and the FTC has charged seal providers with promising checks that, according to the agency, they did not perform, such as ControlScan in 2010 and TRUSTe in 2014 (both cases ended in settlements). In Brazil, the Consumer Defense Code prohibits advertising that is wholly or partly false and puts the burden of proof on the advertiser. For SOC 2 and ISO/IEC 27001, the logo depends on a report from an independent auditor or a certificate from a certification body.