LLM A/B Testing: Testing Prompts and Models in Production
LLM A/B testing: how to test prompts and models with real users, measure acceptance, cost and latency, and avoid the LLM judge traps.

📚 This article is part of the guide AI Personalization and A/B Testing: How They Work Together.
LLM A/B testing is a controlled experiment in which real users are randomly split between two versions of the AI behavior inside your product (a different prompt, a different model, a different temperature, with or without RAG) to measure the causal effect on what they actually do. Offline evaluation, with a question set and an LLM as judge, filters candidates; only the online test tells you whether the change helps users get their task done, and it brings three traps of its own: non-deterministic output, per-message metrics under per-user randomization, and a cost per request that differs between variants. This guide is part of the guide to AI personalization and A/B testing and covers the division of labor between offline and online, the measured biases of LLM judges, a metrics and guardrails table, the simulation showing why the naive interval is wrong, a worked example of current prompt against new prompt with both calculators embedded, and the cost math that decides whether the winner pays for itself.
What changes when the variant is AI behavior
The statistics are the same as in any A/B test: two groups made equivalent by randomization, a primary metric declared upfront, a sample size computed for a minimum effect, and a readout with a confidence interval. What changes is the nature of the variant and the kind of damage it can do.
On a web page, the variant is a fixed piece of copy or layout, identical for everyone in the group. In an LLM feature, the variant is a configuration that produces different outputs on every call. That has four practical consequences:
- The same input produces different outputs. Anthropic’s Messages API documentation, checked in September 2026, says that even with a temperature of 0.0 the results will not be fully deterministic (the same page marks the parameter as deprecated: models released after Claude Opus 4.6 only accept a value of 1.0). Your metric has an extra source of variance that a button never had.
- The variant has a variable cost. A prompt with more instructions or with passages retrieved from a knowledge base consumes more input tokens, so every request costs more. That never happens when you change the color of a CTA.
- The variant can fail in new ways. Refusing a legitimate task, making up a fact, taking too long to start answering, hitting the provider’s rate limit.
- The variant can change without you touching it. A model alias that points to the latest version, or a change in the serving infrastructure, can shift behavior mid-test.
The difference from the piece on AI-generated variations matters: there, AI writes versions of a page and the test is an ordinary page test. Here, AI is the feature under test, and the variation is its behavior.
Offline evaluation vs online experiment: who decides what
Offline evaluation runs the candidate configuration over a fixed set of inputs and scores the outputs, with human labels or another model acting as judge. An online experiment exposes real users and measures what they do. You need both, and each one answers a different question.
The piece by Widad Machmouchi and Somit Gupta, from Microsoft’s Experimentation Platform, published in September 2023, draws the line plainly: offline evaluation is suitable for early feature development, but it cannot assess how model changes benefit or degrade the user experience in production. The same piece describes the experiment designs it recommends at feature launch and after launch.
| Question | Offline (set + judge) | Shadow | Online A/B test |
|---|---|---|---|
| Is the answer correct and on tone? | Yes, on the questions in the set | Partly, with no user reaction | Indirectly, through behavior |
| How much does it cost and how long does it take? | Estimate in a test environment | Yes, on real traffic | Yes, on real traffic |
| Does the user accept, apply or regenerate? | No | No | Yes |
| Does the user come back to the feature? | No | No | Yes |
| Does the change cause the observed effect? | No | No | Yes, because of randomization |
| Cost of getting it wrong | Low | Low | High, users are exposed |
An evaluation set has a structural limit that no judge fixes: it is fixed, and your users are not. The questions arriving in production shift with the season, the audience and the product itself, and the set ages quietly. That is why a healthy workflow uses offline evaluation to cut ten candidates down to two and lets the online test choose between those two.
LLM judges and the biases that have already been measured
Using another model to grade answers is cheap and scales, and there is evidence it agrees with people quite often. The paper by Zheng and coauthors, “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena” (2023), reports that strong judges such as GPT-4 reached over 80 percent agreement with human preferences, the same level of agreement found between humans. The same paper measures three biases that make a judge risky as a primary metric without calibration:
- Position bias. The authors built pairs of similar answers and swapped their order. With the default prompt, GPT-4 kept its verdict in 65.0 percent of cases, GPT-3.5 in 46.2 percent and Claude-v1 in 23.8 percent; Claude-v1 favored the first position in 75.0 percent of cases.
- Verbosity bias. In a “repetitive list” attack, where an answer was padded with rephrased items that added no new information, the judge preferred the inflated version in 91.3 percent of 23 cases with Claude-v1 and with GPT-3.5, and in 8.7 percent with GPT-4.
- Self-enhancement bias. Compared with humans, GPT-4 gave itself a 10 percent higher win rate and Claude-v1 a 25 percent higher one. The authors caution that, with limited data and small differences, the study cannot determine whether the bias actually exists.
The paper itself suggests simple mitigations: call the judge twice with the order swapped and only declare a win when the same answer wins both times, and provide a reference answer for math questions, which cut GPT-4’s failures from 14 in 20 to 3 in 20 in the authors’ test. The operating rule that follows is this: an LLM judge is a measuring instrument, and instruments get calibrated. Set aside a sample of answers, label it with people, measure the judge’s agreement on that sample, and only then use its score, and even then as a filter or a secondary metric. Treating the judge’s score as a surrogate metric for user value requires the same validation as any other surrogate.
LLM A/B testing metrics: success, usage and guardrails of their own
The primary metric should be tied to task value and measured per user, the unit you randomized. Among the metrics of an LLM feature, the Microsoft piece describes a funnel running from the opportunity to suggest all the way to a response the user accepts and keeps, and it adds metric families that a page test never needs: token consumption, truncated responses, 429 overload errors, time to first token measured at several percentiles, content-filtered responses, and retention.
| Metric | Type | Natural unit | Trap |
|---|---|---|---|
| Task success (did what they came to do) | Primary | User | Defining “success” after seeing the data |
| User accepted an answer without regenerating | Primary or secondary | User | Changing the measurement window mid-test |
| Acceptance rate per response (copied, applied) | Secondary | Message | It is a ratio metric: needs the delta method |
| Regenerate and rephrase rate | Secondary, dissatisfaction signal | Message | A slower variant cuts regeneration through fatigue, not quality |
| Thumbs up or down | Secondary | Message | Only a fraction responds and it is not randomized; responders differ from non-responders |
| Feature retention (came back next week) | Long-term primary | User | Needs a long window; novelty inflates the first weeks |
| Tokens and cost per request | Guardrail | Request | Comparing per-request averages when one variant drives more requests per user |
| Latency (p95 time to first token) | Guardrail | Request | The mean hides the tail; use quantile metrics |
| Error, timeout and 429 rate | Guardrail | Request | A quota shared between variants creates interference |
| Refusal rate on legitimate tasks | Guardrail | Message | Needs consistent classification across both variants |
| Safety incidents (filtered content, leakage) | Guardrail with a zero or near-zero limit | Event | Rare events: the test rarely has power, so they need review outside the statistics |
Three notes on that table. First, accepted without regenerating is a good per-user primary because it combines two signals (the user used the answer and did not need another one) without depending on anyone clicking a thumb. Second, thumbs ratings are not a primary metric: responders select themselves, and a variant can change who responds without changing quality. Third, LLM guardrails get their limits declared upfront, like any other guardrail metric, and cost is a guardrail, not a billing footnote.
Where the extra variance in an LLM test comes from
An LLM test has four sources of variance that a page test either lacks or has at a much smaller scale. Two of them affect how you randomize, two affect how you read the result.
Randomize by user, not by request
The first design decision is the randomization unit. Randomizing each request independently looks like it gives you more sample, and that is exactly the trap: the same user gets variant A on the first message, B on the second and A again on the third, inside the same conversation.
Randomizing by user fixes the contamination and creates a statistical cost: all of a user’s messages are correlated (a user who likes the feature accepts many answers, one who does not accepts few). Assignment has to be deterministic on the user identifier, exactly as in server-side A/B testing, so the same user always lands in the same arm on any server.
Per-message metric, per-user randomization: the naive interval is wrong
When the metric is “acceptance rate per response” and the randomized unit is the user, the metric becomes a ratio of two per-user sums (accepted responses over generated responses). Treating each message as an independent observation underestimates the standard error, and the size of the error grows with the number of messages per user. The fix is the delta method, explained in ratio metrics.
To size the problem for LLM features, we ran a simulation built for this guide (a pseudorandom generator with a fixed seed): 2,000 A/A tests, with no real difference between arms, 2,000 users per arm, a skewed number of requests per user and acceptance propensity varying across users. In each test, the difference in per-message acceptance rate was tested at 95 percent in two ways.
| Simulated scenario | Requests per user (mean) | Delta method standard error divided by naive | False positives with naive interval | False positives with delta method |
|---|---|---|---|---|
| Occasional-use feature | about 3.7 | 1.28 | 12.10% | 4.55% |
| Heavy-use chat | about 20 | 2.83 | 49.15% | 4.50% |
The nominal level is 5 percent. With the delta method, both simulations land close to it. With the naive interval, the occasional-use feature already more than doubles the false alarm rate, and the heavy-use chat declares a winner in almost half of the tests where nothing changed. That is why an LLM dashboard showing “per-message acceptance, statistically significant” without saying how it computed the standard error deserves immediate suspicion.
Novelty, shared caches and versions that change on their own
Novelty effect. New AI behavior can attract curiosity: users try it more, regenerate to see what happens, accept out of enthusiasm. The novelty effect makes the first weeks exaggerate the gain or the loss. Run full weeks, look at the effect by week of exposure, and be suspicious of a gain that only exists in the first one.
Interference through shared resources. If both variants use the same response cache, a user in arm B can receive an answer that prompt A generated and stored. If they share the same provider rate limit, a variant that consumes more can trigger 429 errors in the other. Both situations violate the assumption that one user’s treatment does not affect another user’s outcome, the subject of interference between variants. The fix is to key the cache by variant and to monitor errors per arm.
A model version that changes mid-test. Anthropic’s model IDs page, checked in September 2026, explains that for models before the 4.6 generation, the API’s dateless aliases point to the most recent dated snapshot for that version, while each model ID is a pinned version. The same page warns that even with the ID and weights unchanged, changes in serving infrastructure (request routing, safety classifiers, sampling logic) can produce minor differences in observable behavior. The practical rule applies to any provider: pin the model version in both variants, record that version in the experiment configuration, and log any known provider change during the window.
Worked example: current prompt vs new prompt
An ecommerce SaaS has a feature that writes a product description from the attributes a merchant entered. The merchant can accept the text, edit it or regenerate it. The team wrote a new prompt with more style instructions and examples retrieved from the descriptions merchants approved most often. The new prompt passed offline evaluation and a shadow experiment with no rise in errors. What is left is the online question.
Declared before launch:
- Randomization unit: merchant (user), deterministic assignment by identifier.
- Primary metric: share of users who accepted at least one description without regenerating during the test.
- Historical baseline: 30 percent.
- Minimum effect that justifies the new prompt’s cost: plus 5 percent relative.
- Guardrails: cost per thousand requests, p95 time to first token, error rate and refusal rate, with written limits.
- Model version: pinned and identical in both arms.
How many users and how many days
Enter a current rate of 30, a minimum effect of 5 relative, 95 confidence, 80 power, 10,000 visitors per week (here, new users entering the test each week) and a two-sided test in the calculator below:
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
It returns 14,856 users per variation, 29,712 in total and 21 days. Three full weeks is good news in this case, because the window leaves room to see whether any first-week novelty effect holds up in the weeks that follow.
The result
After 21 days, the numbers from the simulation built for this guide (a generator with a fixed seed, with a small true effect built into the new prompt) were:
- Control (current prompt): 14,962 users, 4,371 accepted a description without regenerating.
- Variant (new prompt): 15,038 users, 4,584 accepted.
Before reading the effect, check the split: 14,962 against 15,038 on a configured 50/50 split gives a chi-square of 0.1925 and a p-value of 0.6608 in the sample ratio mismatch (SRM) test, so there is no sign of a randomization problem. Now paste the numbers into the significance calculator:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
The screen shows a rate of 29.21% for control and 30.48% for the variant, a relative lift of +4.3%, a p-value of 0.0163, a 95% CI of the difference of +0.2% to +2.3% (pp) and the verdict “Significant winner”, with B winning. The exact values behind the screen are a difference of 1.2688 percentage points, an interval from 0.2333 to 2.3043 points and a z of 2.4012.
The per-message metric tells a different story
In the same test, per-message acceptance was 11.84 percent for control (6,587 accepted out of 55,643 requests) and 12.07 percent for the variant (6,834 out of 56,638), a difference of +0.23 points. With the delta method, the interval runs from −0.26 to +0.72 points and the p-value is 0.3623: inconclusive. The naive calculation would give an interval from −0.15 to +0.61 points, about 23 percent narrower than it should be. The right reading is that the per-message secondary metric lacks the power to confirm or contradict the primary, and the decision stays anchored on the per-user metric declared upfront.
Cost belongs in the decision: what each additional user costs
The new prompt uses more tokens. The prices below are illustrative, chosen to keep the math clear, and do not correspond to any specific provider: $3.00 per million input tokens and $15.00 per million output tokens.
| Item | Current prompt | New prompt | Difference |
|---|---|---|---|
| Input tokens per request | 1,200 | 2,000 | +66.7% |
| Output tokens per request | 300 | 320 | +6.7% |
| Cost per thousand requests | $8.10 | $10.80 | +$2.70 (+33.3%) |
| Requests per user during the test | 3.72 | 3.77 | n/a |
| Extra cost per 10,000 users | n/a | $101.69 | n/a |
The extra cost per 10,000 users is 3.7663 requests per user in the new arm, times the $0.0027 difference per request, times 10,000. The gain, on the same base of 10,000 users, is the difference in the primary metric: 126.9 additional users who accept a description without regenerating, at the point estimate. Divide one by the other, and do the same at both ends of the interval:
| Reading of the effect | Additional users per 10,000 | Cost per additional user |
|---|---|---|
| CI lower bound (+0.23 pp) | 23.3 | $4.36 |
| Point estimate (+1.27 pp) | 126.9 | $0.80 |
| CI upper bound (+2.30 pp) | 230.4 | $0.44 |
The decision is now a business one, and it has a clear shape. If an additional merchant who adopts the feature is worth more than $4.36 (for example, through its effect on plan retention), the new prompt pays for itself even in the plausible worst case. If that merchant is worth between $0.80 and $4.36, it pays off at the point estimate, but with real risk. If it is worth less than $0.80, it does not pay off even at the point estimate. The same math applies, with bigger numbers, to switching models: a larger model usually moves the cost column far more than a prompt does.
LLM A/B testing launch checklist
- Filter offline first. An up-to-date evaluation set, a judge whose agreement with human labels has been measured, and order-swapped evaluation whenever the judge compares pairs.
- Run dark mode or shadow when the change affects capacity. Model swaps and much longer prompts change cost, latency and error rate before they change any behavior.
- Randomize by user, with deterministic assignment. The same user always gets the same variant, on any server and in any conversation.
- Declare the per-user primary metric and the window. In writing, in a pre-registered analysis plan, before the first data point.
- Declare guardrails with limits. Cost per thousand requests, p95 time to first token, errors and 429s per arm, refusals, safety incidents.
- Pin the model version and record the full configuration. Model ID, versioned prompt, temperature, retrieval parameters and a per-variant cache key.
- Compute sample and duration in full weeks. Use the sample size calculator and round up to whole weeks.
- Check SRM before reading the effect. The SRM checker takes seconds.
- Read per-message metrics with the delta method. Never with a per-message binomial interval.
- Close the decision with the cost math over the whole interval. And if the feature is central to the product, keep a long-term holdout to measure retention after launch.
Common mistakes
| Mistake | Why it misleads | Fix |
|---|---|---|
| Using the LLM judge’s score as the primary metric without human calibration | The judge has measured position and verbosity bias and can reward the longest answer instead of the most useful one | Judge as an offline filter; primary based on user behavior |
| Randomizing per request | Mixes variants inside one conversation and inflates the apparent sample | Deterministic per-user assignment |
| Looking only at thumbs ratings | Responders are self-selected and rare | Acceptance, regeneration and task success as the main signals |
| Forgetting cost and latency | An acceptance winner can be more expensive and slower than the gain is worth | Declared guardrails and cost math over the CI |
| Not pinning the model version | An alias points to a new version mid-window | Fixed model ID recorded in the configuration |
| Ignoring provider changes during the test | Serving infrastructure can change behavior without changing the ID | Log the window, compare the effect by week, rerun if needed |
| Reading per-message acceptance with a naive interval | Up to 49.15% false positives in this guide’s simulation | Delta method |
| Declaring victory in week one | Novelty can inflate early usage | Full weeks and effect by week of exposure |
| Sharing a cache between variants | Arm B receives answers from prompt A | Cache key includes the variant |
Make this automatic with Donnu
Where LLM A/B testing usually breaks is not the statistics, it is the record. Three weeks later, nobody remembers which model version was pinned, which primary metric was declared, or whether the cost limit was written before or after the result came in. Without that record, the cost math over the interval turns into a debate of opinions.
To be direct about what Donnu A/B is today: a client-side A/B testing tool for web pages. Assigning a prompt or a model happens in your backend, so assigning the AI variant itself is work for your server or for a server-side feature flag platform, as explained in feature flags vs A/B testing. What Donnu does in this workflow is what it does for any experiment: it records the experiment configuration at the moment it is created, with the declared primary metric, and keeps the history per experiment. And tests on the visible layer of the feature (where the AI button appears, how the suggestion is presented, which prompt invites people to use it) are exactly the snippet’s use case.
For the numbers, the calculators in this guide are free: the sample size calculator to plan the window and the significance calculator to read the per-user primary metric. If you want to see Donnu on page tests, start a 14-day free trial.
References
- Machmouchi, W. and Gupta, S. How to Evaluate LLMs: A Complete Metric Framework. Microsoft Research, Experimentation Platform (ExP), September 27, 2023. Read in this session. Supports the statement that offline evaluation is suitable for early development but cannot assess how model changes benefit or degrade the user experience in production; the metric funnel running from the opportunity to suggest to a response that is accepted and kept; the metrics for tokens, 429 errors, truncated and content-filtered responses, time to first token at multiple percentiles, thumbs feedback and retention; and the dark mode and 0-1 (at launch), shadow (post launch, where metrics that need a user response cannot be measured) and 1-N experiment designs. microsoft.com.
- Zheng, L., Chiang, W.-L., Sheng, Y. and coauthors. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv 2306.05685, June 9, 2023. PDF read in this session. Supports the over 80 percent agreement between strong judges and human preferences, the same level as agreement between humans; table 2 on position bias (consistency of 65.0, 46.2 and 23.8 percent for GPT-4, GPT-3.5 and Claude-v1 with the default prompt, and Claude-v1 favoring the first position 75.0 percent of the time); table 3 on the repetitive list attack (91.3, 91.3 and 8.7 percent failure on 23 answers); the self-enhancement observation (10 and 25 percent higher win rates, with the caveat that the study cannot determine the bias); and the mitigations of swapping positions and using a reference answer (failures down from 14 in 20 to 3 in 20). arxiv.org.
- Anthropic. Model IDs and versioning. Claude platform documentation, checked September 13, 2026. Supports that each model ID is a pinned version, that aliases for models before the 4.6 generation point to the most recent dated snapshot for that version, and that serving infrastructure updates can produce minor behavioral differences even when the model ID and weights have not changed. platform.claude.com.
- Anthropic. Create a Message, temperature parameter. Claude API reference, checked September 13, 2026. Supports that even with a temperature of 0.0 the results will not be fully deterministic, and that the parameter is deprecated: models released after Claude Opus 4.6 only accept a value of 1.0. platform.claude.com.
- LaunchDarkly. AgentControl. Documentation (the old /docs/home/ai-configs address redirects to this page), checked September 13, 2026. An example of an experimentation platform that treats prompts, instructions and model parameters as configuration variations and compares variations on cost, latency, satisfaction and other metrics; cited to illustrate the design, not as a recommendation. launchdarkly.com.
Read next: AI personalization and A/B testing · AI-generated A/B test variations · AI copilots in A/B testing tools · Ratio metrics and the delta method · Randomization unit · Guardrail metrics · Novelty effect · Significance calculator · Leia em português
Frequently asked questions
- What is LLM A/B testing?
- It is a controlled experiment in which real users are randomly assigned to one of two configurations of a feature built on a language model, for example two prompts, two models, two temperatures, or the same answer with and without retrieval from a knowledge base (RAG). The difference from offline evaluation is that the test measures the causal effect on user behavior, such as accepting the answer, finishing the task or coming back to the feature, rather than a score given by a grader on a fixed set of questions.
- Does offline evaluation with an LLM judge replace an A/B test?
- No. It is for filtering candidates before any user is exposed. The Microsoft Research article on evaluating LLM features states that offline evaluation suits early development but cannot assess how model changes benefit or degrade the user experience in production. And the judge has documented biases: in the study by Zheng and coauthors (2023), the most consistent of the three judges evaluated kept the same verdict after the answer order was swapped, with the default prompt, in only 65.0 percent of cases.
- What should the randomization unit be in a prompt or model test?
- The user, almost always. Randomizing per request makes a single conversation mix two behaviors, contaminates the per-user metric and makes retention impossible to measure. The statistical consequence is that per-message metrics, such as acceptance rate, become ratio metrics and require the delta method. In the simulation built for this guide, the naive per-message interval produced a false positive in 49.15 percent of 2,000 A/A tests in a heavy chat scenario, against 4.50 percent with the delta method.
- Which metrics should an LLM A/B test use?
- A per-user primary metric tied to task value, such as the share of users who accepted an answer without regenerating, plus usage signals (copied, applied, regenerated, rephrased) and LLM-specific guardrails: tokens and cost per request, high-percentile latency, error and refusal rates, and safety incidents. Explicit thumbs ratings work as a secondary signal, because only a fraction of users respond and that fraction is not randomized.
- How long does a prompt A/B test take?
- It depends on the baseline rate and the effect you want to detect. In this guide example, with 30 percent of users accepting an answer without regenerating and a minimum effect of plus 5 percent relative, at 95 percent confidence and 80 percent power, you need 14,856 users per variation. With 10,000 new eligible users per week, that is 21 days, three full weeks, which also helps you get past the novelty effect.
- How do you decide whether a prompt that performs better but costs more is worth it?
- Put gain and cost in the same unit and use the interval, not just the point estimate. In the illustrative example in this guide, the new prompt costs $101.69 more per 10,000 users and produces 126.9 additional users who accept an answer, about $0.80 per additional user. At the lower bound of the confidence interval, the same cost buys only 23.3 additional users, about $4.36 each. If an additional user is worth less than that to your business, the decision comes down to how much risk you are willing to take.