Long Copy vs Short Copy: How to Actually Test It
Long copy vs short copy in A/B testing: why the question is badly framed, how to isolate length from everything else, and a worked example read by segment.

📚 This article is part of the guide Conversion Rate Optimization (CRO): The Complete 2026 Guide.
“Long copy or short copy” is one of the oldest questions in digital marketing and one of the worst to test, because almost nobody tests what they think they are testing. Swapping a 380-word page for a 1,900-word page almost never changes length alone: it changes the arguments, the proof, the order and the structure all at once. This guide covers what reading research actually says (visitors read at most 28 percent of the words on an average page, per the Nielsen Norman Group), how to build a test that isolates length, why scroll depth is a diagnostic and never a verdict, and a worked example where the long version wins with a p-value of 0.0006 while nearly all of the gain comes from a single segment. This guide is part of our complete guide to conversion rate optimization.
Why the question is badly framed
When someone asks “long or short copy?”, the implicit question is usually a different one: “do I need to explain more, or do I need to get to the point?”. Those are not the same thing, and confusing them is what makes the test come out crooked.
Length is not a design variable, it is a consequence. Page length is the result of three separate decisions:
- How many questions you answer. A page that answers 3 objections is shorter than one that answers 9. That is scope, not style.
- How much proof you show per question. One number, versus a number plus a chart plus a testimonial. That is evidence density.
- How many words you spend per idea. The same answer in 40 or in 140 words. That is the only item that is genuinely “copy length”.
Public debate almost always treats the three as one. That is where the legend comes from: someone swaps a short, generic page for a long, well-argued one, watches the number go up, and concludes that “long copy converts”. What went up was the number of objections answered, not the number of words.
What reading research shows
Two datasets from the Nielsen Norman Group anchor any conversation about page length, and both are public and citable.
How much gets read. In “How Little Do Users Read?” (Jakob Nielsen, May 5, 2008), an analysis of 45,237 page views from 25 users with instrumented browsers (data collected in 2005 by Weinreich, Obendorf, Herder and Mayer) leads to the conclusion still quoted today: on an average web page, users have time to read at most 28 percent of the words during an average visit, and 20 percent is more likely. The same article records that visitors read half the information only on pages of 111 words or fewer, and that the average page in the set held 593 words.
Where attention lands. In “Scrolling and Attention” (Therese Fessenden, April 15, 2018), an eyetracking study of more than 130,000 fixations from 120 participants measured that 57 percent of page-viewing time stays above the fold and 74 percent within the first two screenfuls. Compared with the same group’s 2010 study, which put 80 percent above the fold, people scroll more than they used to, but the article still says they rarely go past the third screenful.
The correct reading of both studies is the same: the problem was never length, it was dependency. A 2,000-word page where the offer only becomes clear in paragraph 14 is a bad page. A 2,000-word page where the offer is settled in the first 80 words and the other 1,920 serve readers who want depth has no length problem at all.
The confounder: changing length changes everything at once
This is why most length tests teach nothing. The table shows what actually changes when someone “tests the long version”.
| what differs between the versions | is it length? | plausible effect on conversion |
|---|---|---|
| number of objections answered | no, that is scope | high, and probably what carries the result |
| amount of proof (numbers, cases, testimonials) | no, that is evidence | high; see social proof |
| number of FAQ entries on the page | no, that is scope | medium, and it also moves search traffic |
| placement of the action button and how often it repeats | no, that is layout | medium |
| page weight and load time | no, that is performance | medium; see page speed and conversion |
| words spent saying the same idea | yes | usually the smallest effect on this list |
There are two honest ways out, and they answer different questions:
The package test. You compare today’s short page against the new long page, whole, and accept that the answer is “this page beats that page”. It is a legitimate test and often the most useful one commercially. What it does not license is the sentence “long copy converts better”: it licenses “this long page beat that short page”.
The density test. You hold exactly the same information, in the same order, with the same proof, and vary only how many words each block spends. Cutting from 1,900 to 900 words without losing a single argument is hard to write and is the only design that answers the question in the title. If you go that route, freeze the number of blocks, the number of images and the position of every button, and use the word counter to record each arm’s exact length in the test plan.
Mixing the two designs and then concluding about length is what produces the categorical claims that circulate on this topic. If you want to separate more than one variable at once, the right tool is not a two-page A/B test: see A/B, multivariate and split URL testing.
What to measure, and what is only a diagnostic
| signal | role in the test | trap |
|---|---|---|
| page conversion (primary action) | primary metric | none, as long as it is declared up front |
| lead or order quality | secondary primary metric, declared up front | a long page can filter and convert less with better leads |
| scroll depth | diagnostic | rises on the long version by construction, since the page is taller |
| time on page | diagnostic | rises by construction, and also rises when the visitor is confused |
| bounce rate | diagnostic | falls when you add any interaction, with nothing having improved |
| clicks on the action button | diagnostic | more buttons produce more clicks without producing more conversions |
The rule that settles most meeting-room debates: if a metric rises by construction when the page gets longer, it cannot be the verdict. Scroll depth and time on page are the canonical examples. They are there to explain why the result came out as it did, not to decide who won. The primary metric has to be written down first, as we argue in primary metric and OEC.
The lead quality case deserves special attention in SaaS and services. A long page that explains price, scope and what the product does not do tends to reduce demo request volume and raise the share that turns into an opportunity. If your verdict is form volume alone, you will kill the version that makes money. That is covered in SaaS demo request form optimization.
Sample size: the arithmetic before the opinion
| page conversion | detect +10% relative | detect +15% relative | detect +20% relative |
|---|---|---|---|
| 2.0% | 80,682 per variant | 36,693 per variant | 21,109 per variant |
| 4.2% | 37,513 per variant | 17,050 per variant | 9,803 per variant |
| 8.0% | 18,872 per variant | 8,568 per variant | 4,921 per variant |
At 4.2 percent conversion and 8,000 visitors a week, detecting 10 percent relative takes 66 days. At 15,000 a week, 36 days. If your traffic does not get there, the answer is not to run it anyway and call a winner on day 10: it is to test a bigger change, or to read CRO for low-traffic sites.
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
Worked example: a B2B SaaS landing page
A software company runs a demo-capture landing page. The current version has 380 words: the offer, three benefits and the form. The new one has 1,900 words: the same offer and the same three benefits, plus a how-it-works section, a comparison table against the manual alternative, two use cases, pricing explained, and eight FAQ entries. It is a package test, and the report will say so.
The test runs at 38,000 visitors per arm (above the 37,513 the calculator asks for at 10 percent relative) over 36 days at 15,000 visitors a week. The traffic-source analysis was declared before launch, with two segments: cold traffic (paid media to new audiences) and warm traffic (brand search, email and referral).
| cut | version | visitors | demo requests | rate | relative lift | p-value |
|---|---|---|---|---|---|---|
| Overall | short (380 words) | 38,000 | 1,602 | 4.216% | baseline | baseline |
| Overall | long (1,900 words) | 38,000 | 1,798 | 4.732% | +12.2% | 0.0006 |
| Cold traffic | short | 26,000 | 858 | 3.30% | baseline | baseline |
| Cold traffic | long | 26,000 | 910 | 3.50% | +6.1% | 0.2083 |
| Warm traffic | short | 12,000 | 744 | 6.20% | baseline | baseline |
| Warm traffic | long | 12,000 | 888 | 7.40% | +19.4% | 0.0002 |
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
Paste 38000 and 1602 into side A and 38000 and 1798 into side B to reproduce the overall row: p-value 0.0006, relative lift 12.2 percent, confidence interval of the difference from 0.22 to 0.81 percentage points. For the warm segment, use 12000 and 744 against 12000 and 888: p-value 0.0002 and an interval from 0.56 to 1.84 percentage points.
What the segment reading licenses and what it does not. It licenses this reading because it was declared before the test ran, with only two cuts and a written hypothesis (“we expect a larger effect among people who already know the brand”). Had the traffic-source split been invented after seeing the overall result, it would be fishing, and the segment p-value would not mean what it appears to mean: that is the subject of heterogeneous treatment effects and of the pre-registered analysis plan.
Note also that the cold segment carries 26,000 visitors per arm and still does not conclude: at a 3.30 percent baseline, that sample detects roughly 12 percent relative, and the observed effect there was 6.1. Failing to conclude is not the same as showing there is no effect. Read observed power before writing “long copy does not work for cold traffic”.
The thing that usually beats both: the layered page
The strongest alternative to this debate is not picking a side, it is changing the format. Jakob Nielsen, in “Long vs. Short Articles as Content Strategy” (November 11, 2007), argues that the two reader profiles, the hurried one and the researcher, are usually the same person at different moments, and recommends starting with overviews and short, simplified pages and then linking to long, in-depth coverage on other pages. In the article’s illustrative arithmetic, the mixed diet outperforms a short-only diet.
On a landing page, that becomes a concrete design:
- The first screenful settles the entire offer. What it is, who it is for, what changes, and the action. Someone who will only read 20 percent of the words needs those 20 percent to be right here.
- Immediately below, your audience’s number one objection, answered with proof rather than adjectives.
- Depth in self-contained blocks, each with a heading that states what it answers, so diagonal reading works.
- FAQ at the end, absorbing the remaining objections while also earning search traffic and AI citations.
- The action repeated at the end of every large block, for people who decided halfway down.
This format makes the “long or short” debate nearly irrelevant, because the page is long for those who want it and short for those who do not. And it changes what is worth testing: instead of “1,900 against 380 words”, you test “which objection moves up to the first screenful”, which is a much smaller question, much cheaper, and usually more profitable.
The part an A/B test does not measure: search and AI citation
An A/B test measures what happens to people who already arrived at the page. Copy length has a second effect that happens before that and that no two-week test captures: how much surface the page offers to search engines and to AI answer engines.
A 380-word page answers one question. A well-structured 1,900-word page answers eight, and every block with its own heading is a candidate passage for an answer. This is not an argument for writing more for its own sake: it is an argument for writing in self-contained blocks with a citable claim at the top of each, which happens to be the same format that lets a hurried visitor understand the page diagonally. The two requirements coincide, which is convenient.
Three practical consequences for anyone deciding on length:
- The search gain shows up in months, the test result shows up in weeks. If the short version wins narrowly and the long one carries eight indexable blocks, the decision may be to keep the long one despite losing the test. That is not ignoring data, it is acknowledging that the test never measured that part. Record the choice in the report.
- FAQ blocks pull more weight than the rest. They answer in question form, which is exactly what an answer engine looks for, and they become structured markup on the page. See structuring FAQ content so LLMs cite you.
- Cutting long copy can drop traffic you were not measuring. Before shipping the lean winning version, check which queries bring people to the page today and whether the matching answers survive the cut.
The reverse holds too: none of this rescues a long page that restates the same idea in different words. Volume without new information earns neither search nor conversion, and it adds page weight. The criterion is always the same, applied block by block: does this block answer a question someone actually asks? If not, it is helping neither conversion nor traffic.
When long copy is a reasonable bet and when it is a symptom
Before spending 36 days of traffic, check whether the hypothesis makes sense for your case. The table maps situations to the more plausible bet.
| situation | more plausible bet | why |
|---|---|---|
| new product in a category the customer does not know | more explanation | there is no ready mental reference; every unanswered objection becomes an exit |
| expensive purchase, with a committee or internal approval | more explanation | the visitor needs material to convince other people, and they copy it from your page |
| simple product in a mature category | less copy | the visitor already knows what it is; extra words only delay the action |
| search traffic with stated purchase intent | less copy | someone searching “buy X” is past the stage of learning what X is |
| cold paid traffic, first exposure | it depends, and this is where testing pays most | it may need context, or it may just need less friction |
| high volume of “please explain” questions in support | more explanation, of the right doubt | read the actual questions first; the objection you imagine is rarely the one that shows up |
| the page is long because nobody can delete anything | neither, this is a decision problem | length by internal politics is not a content strategy |
That last row is more common than it looks. Long pages usually grow by accretion: every department asks for a paragraph, nobody has the authority to remove one, and the result is a page that is long by sedimentation rather than by choice. In that case the right test is not “long against short”: it is to rewrite the lean version with the same information and see whether anyone misses anything. The landing page grader helps map what the page is trying to do before deciding what to cut.
One warning about novelty. A large rewrite usually shows a stronger effect in the first days, especially with returning visitors, and then regresses. The test therefore has to cover at least one full weekly cycle; see the novelty effect and weekly cycles in A/B tests.
How to read the result without fooling yourself
- What changed besides the word count? List everything. If the list runs past three items, the report is about the package, not about length.
- Was the primary metric declared up front? If the verdict appeared after someone looked at scroll depth and time on page, the result is already contaminated.
- Is the effect uniform across declared segments? If not, the decision may be to serve different versions by traffic source rather than to crown a single winner.
- Did the quality of what converted change? On a form, compare the share that becomes an opportunity, not just the volume.
- Did the long version load more slowly? Measure it. A difference of seconds between arms is a confounder that has nothing to do with copy.
- Was the first week stronger than the second? That is a novelty signal, not a content one.
- Did the segment that failed to conclude have power for the observed effect? Failing to conclude is not evidence of no effect.
- Can someone rewrite the losing version incorporating what worked? The best outcome of the test is usually not the winning arm, it is the third version born from reading it.
Pre-launch checklist
- Write down which of the two tests you are running, package or density, and what the report will be allowed to claim at the end.
- Count the words in both arms and record it in the plan, along with the number of blocks, images and buttons in each.
- Declare the primary metric and state explicitly that scroll depth and time on page are diagnostics.
- Add lead quality as a declared metric whenever conversion is a form submission.
- Declare the segments up front, two or three at most, with a written hypothesis for each.
- Measure weight and load time in both arms: a long version with five extra images is also a performance test.
- Check the long version on a real phone, not just a narrowed desktop browser window.
- Size for 10 to 15 percent relative and accept whatever timeline the calculator gives back.
Automate this with Donnu
The specific pain of testing copy length is that the whole page changes, the effect is rarely uniform across audiences, and the temptation to decide on scroll depth and time on page is strong.
In Donnu, a conversion goal can be a form submission, a click or a page visit, and it can also be confirmed by your own server, which lets you use the qualified request as the metric rather than just the submitted form. An experiment can be restricted by device, traffic source and new versus returning visitor, which helps you declare the segment before running instead of discovering it afterwards. The report is Bayesian, flags when the visitor split drifts from what was configured, and only calls a winner with at least 200 visitors per variant and 7 days of data.
What stays with you: deciding which question you are testing and writing the long version that deserves to be read. For the rest, the sample size calculator, the significance calculator, the word counter and the reading time calculator run every number in this guide on your own data, free.
References
- Nielsen Norman Group (Jakob Nielsen). How Little Do Users Read? May 5, 2008. Article read in full. Source for the at most 28 percent of words read on an average visit, the 20 percent as the likelier case, the 111-word threshold for half the information being read, the 593-word average page, and the dataset of 45,237 page views from 25 users (2005 data from Weinreich, Obendorf, Herder and Mayer). Checked on September 21, 2026. nngroup.com.
- Nielsen Norman Group (Therese Fessenden). Scrolling and Attention. April 15, 2018. Article read in full. Source for the 57 percent of viewing time above the fold, the 74 percent within the first two screenfuls, the study of more than 130,000 fixations from 120 participants, and the comparison with the 80 percent measured in 2010. Checked on September 21, 2026. nngroup.com.
- Nielsen Norman Group (Jakob Nielsen). Long vs. Short Articles as Content Strategy. November 11, 2007. Article read in full. Source for the recommendation to start with overviews and short pages and then link to long coverage, the observation that the hurried reader and the researcher are usually the same person, and the illustrative arithmetic in which a mixed diet outperforms short content alone (the article’s numbers are hypothetical and declared as such by the author). Checked on September 21, 2026. nngroup.com.
Read next: Conversion rate optimization · B2B SaaS landing page A/B testing · Social proof A/B testing · Heterogeneous treatment effects · Pre-registered analysis plan · SaaS demo request form optimization · Page speed and conversion · Leia em português
Frequently asked questions
- Does long copy convert better than short copy?
- There is no general answer, and anyone who claims otherwise is selling a writing style. What reading research does show is that visitors read very little at any length: according to the Nielsen Norman Group, on an average page users have time to read at most 28 percent of the words during a visit, and 20 percent is more likely. That does not condemn long copy. It condemns long copy that has to be read in full to work.
- Why is a long-versus-short copy test almost always inconclusive?
- Because it almost never tests length alone. A long version usually arrives with new arguments, new proof, one more testimonial, an extra FAQ block and a different page structure. If it wins, you cannot tell whether the gain came from the length or from the argument that only exists in the long version. Isolating length means holding the information constant and varying only the density, which is a different and far more tedious test to build.
- Which metric should a copy-length test use?
- Page conversion, with scroll depth and time on page as diagnostics, never as the verdict. Deeper scrolling on the long version is expected by construction, since the page is taller, and proves nothing about quality. In the worked example in this guide the long version wins overall with a p-value of 0.0006, but nearly all of the gain comes from one segment.
- How much of a page do people actually read?
- Very little, and it falls with length. The Nielsen Norman Group, analyzing 45,237 page views from 25 instrumented users, estimated that visitors read half the information only on pages of 111 words or fewer, while the average page in the dataset held 593 words. In an eyetracking study of 120 participants and more than 130,000 fixations, the same group measured that 57 percent of viewing time stays above the fold and 74 percent within the first two screenfuls.
- How much traffic does a copy-length test need?
- Enough for a 10 to 15 percent relative effect, which is the plausible range when the change is real. At 4.2 percent conversion, 95 percent confidence and 80 percent power, detecting 10 percent relative needs 37,513 visitors per variant, which is 66 days at 8,000 visitors a week. Detecting 15 percent relative drops that to 17,050 per variant.
- Is there a better option than choosing between long and short?
- Almost always yes: the layered page. Jakob Nielsen recommends starting with overviews and short, simplified pages and then linking to long, in-depth coverage on other pages, arguing that a mixed diet outperforms short content alone. In landing-page practice that becomes a short answer above the fold with the long material available below and in expandable blocks, for the same visitor who is sometimes in a hurry and sometimes in research mode.