A/B Testing

LLM A/B Testing: Testing Prompts and Models in Production

LLM A/B testing: how to test prompts and models with real users, measure acceptance, cost and latency, and avoid the LLM judge traps.

Flat illustration of two speech bubbles built from geometric blocks side by side, each connected to a small bar chart

LLM A/B testing is a controlled experiment in which real users are randomly split between two versions of the AI behavior inside your product (a different prompt, a different model, a different temperature, with or without RAG) to measure the causal effect on what they actually do. Offline evaluation, with a question set and an LLM as judge, filters candidates; only the online test tells you whether the change helps users get their task done, and it brings three traps of its own: non-deterministic output, per-message metrics under per-user randomization, and a cost per request that differs between variants. This guide is part of the guide to AI personalization and A/B testing and covers the division of labor between offline and online, the measured biases of LLM judges, a metrics and guardrails table, the simulation showing why the naive interval is wrong, a worked example of current prompt against new prompt with both calculators embedded, and the cost math that decides whether the winner pays for itself.

What changes when the variant is AI behavior

The statistics are the same as in any A/B test: two groups made equivalent by randomization, a primary metric declared upfront, a sample size computed for a minimum effect, and a readout with a confidence interval. What changes is the nature of the variant and the kind of damage it can do.

On a web page, the variant is a fixed piece of copy or layout, identical for everyone in the group. In an LLM feature, the variant is a configuration that produces different outputs on every call. That has four practical consequences:

The difference from the piece on AI-generated variations matters: there, AI writes versions of a page and the test is an ordinary page test. Here, AI is the feature under test, and the variation is its behavior.

Offline evaluation vs online experiment: who decides what

Offline evaluation runs the candidate configuration over a fixed set of inputs and scores the outputs, with human labels or another model acting as judge. An online experiment exposes real users and measures what they do. You need both, and each one answers a different question.

The piece by Widad Machmouchi and Somit Gupta, from Microsoft’s Experimentation Platform, published in September 2023, draws the line plainly: offline evaluation is suitable for early feature development, but it cannot assess how model changes benefit or degrade the user experience in production. The same piece describes the experiment designs it recommends at feature launch and after launch.

From evaluation set to an A/B test with real usersFour stages in sequence. First, offline evaluation with a question set and a judge, which filters candidates. Second, a dark mode experiment, which loads the feature without showing it to users and measures performance and reliability. Third, a shadow experiment, which computes both versions’ responses for the same user but shows only the control response, measuring cost, latency and safety without measuring behavior. Fourth, an A/B test with real users, the only stage that measures a causal effect on behavior.each stage answers one question, and only the last one measures behavior1. offlinequestion set+ LLM judge+ human labels2. dark modeloads without showingperformance andreliability3. shadowgenerates both answersshows only controlcost, latency, safety4. A/B testrandomized userssee the variantcausal effectstages 1 to 3: nobody sees the variant, so no acceptance or retentionmeasures what mattersStage 1 cheaply removes bad candidates. Stages 2 and 3 catch cost and capacity failures.Stage 4 is the only one that answers whether users get the task done better with the variant.
Stages 2 and 3 follow the designs described by Machmouchi and Gupta (Microsoft Research, 2023); the four-stage sequence is this guide’s synthesis, because in the original piece dark mode belongs to launching a new feature and shadow experiments to changes after launch. In a shadow experiment no user sees the variant’s response, so metrics that depend on the user’s reaction cannot be measured.
Question Offline (set + judge) Shadow Online A/B test
Is the answer correct and on tone? Yes, on the questions in the set Partly, with no user reaction Indirectly, through behavior
How much does it cost and how long does it take? Estimate in a test environment Yes, on real traffic Yes, on real traffic
Does the user accept, apply or regenerate? No No Yes
Does the user come back to the feature? No No Yes
Does the change cause the observed effect? No No Yes, because of randomization
Cost of getting it wrong Low Low High, users are exposed

An evaluation set has a structural limit that no judge fixes: it is fixed, and your users are not. The questions arriving in production shift with the season, the audience and the product itself, and the set ages quietly. That is why a healthy workflow uses offline evaluation to cut ten candidates down to two and lets the online test choose between those two.

LLM judges and the biases that have already been measured

Using another model to grade answers is cheap and scales, and there is evidence it agrees with people quite often. The paper by Zheng and coauthors, “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena” (2023), reports that strong judges such as GPT-4 reached over 80 percent agreement with human preferences, the same level of agreement found between humans. The same paper measures three biases that make a judge risky as a primary metric without calibration:

Position and verbosity bias measured in three LLM judgesOn the left, verdict consistency when the answer order is swapped: Claude-v1 twenty-three point eight percent, GPT-3.5 forty-six point two percent, GPT-4 sixty-five percent. On the right, failure rate under the repetitive list attack: Claude-v1 and GPT-3.5 ninety-one point three percent, GPT-4 eight point seven percent. Data from Zheng and coauthors, 2023.consistency when order is swappedfailure under verbosity attackdefault prompt, higher is betterlower is betterClaude-v123.8%GPT-3.546.2%GPT-465.0%Claude-v191.3%GPT-3.591.3%GPT-48.7%even the most consistent judge tested in 2023 flipped on 35% of pairs from order alonecurrent models may behave differently: measure YOUR judge’s agreement with human labels
Source: Zheng and coauthors, arXiv 2306.05685, tables 2 and 3. These numbers come from the models evaluated in 2023 and show that the bias exists and can be measured, not how the judge you use today will behave.

The paper itself suggests simple mitigations: call the judge twice with the order swapped and only declare a win when the same answer wins both times, and provide a reference answer for math questions, which cut GPT-4’s failures from 14 in 20 to 3 in 20 in the authors’ test. The operating rule that follows is this: an LLM judge is a measuring instrument, and instruments get calibrated. Set aside a sample of answers, label it with people, measure the judge’s agreement on that sample, and only then use its score, and even then as a filter or a secondary metric. Treating the judge’s score as a surrogate metric for user value requires the same validation as any other surrogate.

LLM A/B testing metrics: success, usage and guardrails of their own

The primary metric should be tied to task value and measured per user, the unit you randomized. Among the metrics of an LLM feature, the Microsoft piece describes a funnel running from the opportunity to suggest all the way to a response the user accepts and keeps, and it adds metric families that a page test never needs: token consumption, truncated responses, 429 overload errors, time to first token measured at several percentiles, content-filtered responses, and retention.

Metric Type Natural unit Trap
Task success (did what they came to do) Primary User Defining “success” after seeing the data
User accepted an answer without regenerating Primary or secondary User Changing the measurement window mid-test
Acceptance rate per response (copied, applied) Secondary Message It is a ratio metric: needs the delta method
Regenerate and rephrase rate Secondary, dissatisfaction signal Message A slower variant cuts regeneration through fatigue, not quality
Thumbs up or down Secondary Message Only a fraction responds and it is not randomized; responders differ from non-responders
Feature retention (came back next week) Long-term primary User Needs a long window; novelty inflates the first weeks
Tokens and cost per request Guardrail Request Comparing per-request averages when one variant drives more requests per user
Latency (p95 time to first token) Guardrail Request The mean hides the tail; use quantile metrics
Error, timeout and 429 rate Guardrail Request A quota shared between variants creates interference
Refusal rate on legitimate tasks Guardrail Message Needs consistent classification across both variants
Safety incidents (filtered content, leakage) Guardrail with a zero or near-zero limit Event Rare events: the test rarely has power, so they need review outside the statistics

Three notes on that table. First, accepted without regenerating is a good per-user primary because it combines two signals (the user used the answer and did not need another one) without depending on anyone clicking a thumb. Second, thumbs ratings are not a primary metric: responders select themselves, and a variant can change who responds without changing quality. Third, LLM guardrails get their limits declared upfront, like any other guardrail metric, and cost is a guardrail, not a billing footnote.

Where the extra variance in an LLM test comes from

An LLM test has four sources of variance that a page test either lacks or has at a much smaller scale. Two of them affect how you randomize, two affect how you read the result.

Randomize by user, not by request

The first design decision is the randomization unit. Randomizing each request independently looks like it gives you more sample, and that is exactly the trap: the same user gets variant A on the first message, B on the second and A again on the third, inside the same conversation.

Per-request randomization versus per-user randomizationOn the top row, per-request randomization: a single user’s conversation receives alternating responses from variants A, B, A and B, and the user’s behavior on the third message already depends on response B from the second. On the bottom row, per-user randomization: user one receives only A responses and user two only B responses, and each conversation is coherent.per-request randomization: one conversation, two behaviorsuser 1response Aresponse Bresponse Aresponse Bthe 3rd message already reacts to response B: one variant’s effect leaks into the otherper-user randomization: every conversation is coherentuser 1AAAAuser 2BBBBper-user metrics and retention now make sense, at the price of correlated messages within each user
Per-request randomization mixes experiences inside one conversation and makes it impossible to attribute acceptance or retention to a variant. The price of per-user randomization is statistical, and the delta method pays it.

Randomizing by user fixes the contamination and creates a statistical cost: all of a user’s messages are correlated (a user who likes the feature accepts many answers, one who does not accepts few). Assignment has to be deterministic on the user identifier, exactly as in server-side A/B testing, so the same user always lands in the same arm on any server.

Per-message metric, per-user randomization: the naive interval is wrong

When the metric is “acceptance rate per response” and the randomized unit is the user, the metric becomes a ratio of two per-user sums (accepted responses over generated responses). Treating each message as an independent observation underestimates the standard error, and the size of the error grows with the number of messages per user. The fix is the delta method, explained in ratio metrics.

To size the problem for LLM features, we ran a simulation built for this guide (a pseudorandom generator with a fixed seed): 2,000 A/A tests, with no real difference between arms, 2,000 users per arm, a skewed number of requests per user and acceptance propensity varying across users. In each test, the difference in per-message acceptance rate was tested at 95 percent in two ways.

Simulated scenario Requests per user (mean) Delta method standard error divided by naive False positives with naive interval False positives with delta method
Occasional-use feature about 3.7 1.28 12.10% 4.55%
Heavy-use chat about 20 2.83 49.15% 4.50%

The nominal level is 5 percent. With the delta method, both simulations land close to it. With the naive interval, the occasional-use feature already more than doubles the false alarm rate, and the heavy-use chat declares a winner in almost half of the tests where nothing changed. That is why an LLM dashboard showing “per-message acceptance, statistically significant” without saying how it computed the standard error deserves immediate suspicion.

Novelty, shared caches and versions that change on their own

Novelty effect. New AI behavior can attract curiosity: users try it more, regenerate to see what happens, accept out of enthusiasm. The novelty effect makes the first weeks exaggerate the gain or the loss. Run full weeks, look at the effect by week of exposure, and be suspicious of a gain that only exists in the first one.

Interference through shared resources. If both variants use the same response cache, a user in arm B can receive an answer that prompt A generated and stored. If they share the same provider rate limit, a variant that consumes more can trigger 429 errors in the other. Both situations violate the assumption that one user’s treatment does not affect another user’s outcome, the subject of interference between variants. The fix is to key the cache by variant and to monitor errors per arm.

A model version that changes mid-test. Anthropic’s model IDs page, checked in September 2026, explains that for models before the 4.6 generation, the API’s dateless aliases point to the most recent dated snapshot for that version, while each model ID is a pinned version. The same page warns that even with the ID and weights unchanged, changes in serving infrastructure (request routing, safety classifiers, sampling logic) can produce minor differences in observable behavior. The practical rule applies to any provider: pin the model version in both variants, record that version in the experiment configuration, and log any known provider change during the window.

Worked example: current prompt vs new prompt

An ecommerce SaaS has a feature that writes a product description from the attributes a merchant entered. The merchant can accept the text, edit it or regenerate it. The team wrote a new prompt with more style instructions and examples retrieved from the descriptions merchants approved most often. The new prompt passed offline evaluation and a shadow experiment with no rise in errors. What is left is the online question.

Declared before launch:

How many users and how many days

Enter a current rate of 30, a minimum effect of 5 relative, 95 confidence, 80 power, 10,000 visitors per week (here, new users entering the test each week) and a two-sided test in the calculator below:

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

It returns 14,856 users per variation, 29,712 in total and 21 days. Three full weeks is good news in this case, because the window leaves room to see whether any first-week novelty effect holds up in the weeks that follow.

The result

After 21 days, the numbers from the simulation built for this guide (a generator with a fixed seed, with a small true effect built into the new prompt) were:

Before reading the effect, check the split: 14,962 against 15,038 on a configured 50/50 split gives a chi-square of 0.1925 and a p-value of 0.6608 in the sample ratio mismatch (SRM) test, so there is no sign of a randomization problem. Now paste the numbers into the significance calculator:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

The screen shows a rate of 29.21% for control and 30.48% for the variant, a relative lift of +4.3%, a p-value of 0.0163, a 95% CI of the difference of +0.2% to +2.3% (pp) and the verdict “Significant winner”, with B winning. The exact values behind the screen are a difference of 1.2688 percentage points, an interval from 0.2333 to 2.3043 points and a z of 2.4012.

Confidence interval of the difference between the new and current promptHorizontal axis in percentage points from minus zero point five to three. The point estimate is plus one point twenty-seven points, with an interval from plus zero point twenty-three to plus two point thirty. The zero line lies outside the interval, so the result is significant. The planned effect line, plus one point five points, lies inside the interval, to the right of the point estimate.difference in the share of users who accepted without regeneratingzeroplanned effect +1.5 pp+1.27 pp+0.23+2.30−0.50.51.02.03.0zero is outside: the new prompt has a positive effect at 95% confidencebut the interval runs from an almost-zero gain to a gain above plan, and that weighs on the cost math
Significant does not mean large. The point estimate, +1.27 points, is actually below the planned effect of +1.5 points, and the lower bound, +0.23, is a gain that may not pay for a more expensive prompt.

The per-message metric tells a different story

In the same test, per-message acceptance was 11.84 percent for control (6,587 accepted out of 55,643 requests) and 12.07 percent for the variant (6,834 out of 56,638), a difference of +0.23 points. With the delta method, the interval runs from −0.26 to +0.72 points and the p-value is 0.3623: inconclusive. The naive calculation would give an interval from −0.15 to +0.61 points, about 23 percent narrower than it should be. The right reading is that the per-message secondary metric lacks the power to confirm or contradict the primary, and the decision stays anchored on the per-user metric declared upfront.

Cost belongs in the decision: what each additional user costs

The new prompt uses more tokens. The prices below are illustrative, chosen to keep the math clear, and do not correspond to any specific provider: $3.00 per million input tokens and $15.00 per million output tokens.

Item Current prompt New prompt Difference
Input tokens per request 1,200 2,000 +66.7%
Output tokens per request 300 320 +6.7%
Cost per thousand requests $8.10 $10.80 +$2.70 (+33.3%)
Requests per user during the test 3.72 3.77 n/a
Extra cost per 10,000 users n/a $101.69 n/a

The extra cost per 10,000 users is 3.7663 requests per user in the new arm, times the $0.0027 difference per request, times 10,000. The gain, on the same base of 10,000 users, is the difference in the primary metric: 126.9 additional users who accept a description without regenerating, at the point estimate. Divide one by the other, and do the same at both ends of the interval:

Reading of the effect Additional users per 10,000 Cost per additional user
CI lower bound (+0.23 pp) 23.3 $4.36
Point estimate (+1.27 pp) 126.9 $0.80
CI upper bound (+2.30 pp) 230.4 $0.44
Cost per additional user who accepts an answer, across the confidence intervalThree horizontal bars. At the lower bound of the interval, each additional user costs four dollars and thirty-six cents. At the point estimate, eighty cents. At the upper bound, forty-four cents. The lower bound bar is about five times longer than the point estimate bar.incremental cost per additional user (illustrative values)lower bound$4.36point estimate$0.80upper bound$0.44the extra cost is the same on every row ($101.69 per 10,000 users); the gain it buys is notif an additional user is worth less than $4.36 to you, the plausible worst case does not pay off
Net value math uses the whole interval. A statistical winner with a lower bound close to zero can be a financial loser when the variant costs more per request.

The decision is now a business one, and it has a clear shape. If an additional merchant who adopts the feature is worth more than $4.36 (for example, through its effect on plan retention), the new prompt pays for itself even in the plausible worst case. If that merchant is worth between $0.80 and $4.36, it pays off at the point estimate, but with real risk. If it is worth less than $0.80, it does not pay off even at the point estimate. The same math applies, with bigger numbers, to switching models: a larger model usually moves the cost column far more than a prompt does.

LLM A/B testing launch checklist

  1. Filter offline first. An up-to-date evaluation set, a judge whose agreement with human labels has been measured, and order-swapped evaluation whenever the judge compares pairs.
  2. Run dark mode or shadow when the change affects capacity. Model swaps and much longer prompts change cost, latency and error rate before they change any behavior.
  3. Randomize by user, with deterministic assignment. The same user always gets the same variant, on any server and in any conversation.
  4. Declare the per-user primary metric and the window. In writing, in a pre-registered analysis plan, before the first data point.
  5. Declare guardrails with limits. Cost per thousand requests, p95 time to first token, errors and 429s per arm, refusals, safety incidents.
  6. Pin the model version and record the full configuration. Model ID, versioned prompt, temperature, retrieval parameters and a per-variant cache key.
  7. Compute sample and duration in full weeks. Use the sample size calculator and round up to whole weeks.
  8. Check SRM before reading the effect. The SRM checker takes seconds.
  9. Read per-message metrics with the delta method. Never with a per-message binomial interval.
  10. Close the decision with the cost math over the whole interval. And if the feature is central to the product, keep a long-term holdout to measure retention after launch.

Common mistakes

Mistake Why it misleads Fix
Using the LLM judge’s score as the primary metric without human calibration The judge has measured position and verbosity bias and can reward the longest answer instead of the most useful one Judge as an offline filter; primary based on user behavior
Randomizing per request Mixes variants inside one conversation and inflates the apparent sample Deterministic per-user assignment
Looking only at thumbs ratings Responders are self-selected and rare Acceptance, regeneration and task success as the main signals
Forgetting cost and latency An acceptance winner can be more expensive and slower than the gain is worth Declared guardrails and cost math over the CI
Not pinning the model version An alias points to a new version mid-window Fixed model ID recorded in the configuration
Ignoring provider changes during the test Serving infrastructure can change behavior without changing the ID Log the window, compare the effect by week, rerun if needed
Reading per-message acceptance with a naive interval Up to 49.15% false positives in this guide’s simulation Delta method
Declaring victory in week one Novelty can inflate early usage Full weeks and effect by week of exposure
Sharing a cache between variants Arm B receives answers from prompt A Cache key includes the variant

Make this automatic with Donnu

Where LLM A/B testing usually breaks is not the statistics, it is the record. Three weeks later, nobody remembers which model version was pinned, which primary metric was declared, or whether the cost limit was written before or after the result came in. Without that record, the cost math over the interval turns into a debate of opinions.

To be direct about what Donnu A/B is today: a client-side A/B testing tool for web pages. Assigning a prompt or a model happens in your backend, so assigning the AI variant itself is work for your server or for a server-side feature flag platform, as explained in feature flags vs A/B testing. What Donnu does in this workflow is what it does for any experiment: it records the experiment configuration at the moment it is created, with the declared primary metric, and keeps the history per experiment. And tests on the visible layer of the feature (where the AI button appears, how the suggestion is presented, which prompt invites people to use it) are exactly the snippet’s use case.

For the numbers, the calculators in this guide are free: the sample size calculator to plan the window and the significance calculator to read the per-user primary metric. If you want to see Donnu on page tests, start a 14-day free trial.

References

Read next: AI personalization and A/B testing · AI-generated A/B test variations · AI copilots in A/B testing tools · Ratio metrics and the delta method · Randomization unit · Guardrail metrics · Novelty effect · Significance calculator · Leia em português

Frequently asked questions

What is LLM A/B testing?
It is a controlled experiment in which real users are randomly assigned to one of two configurations of a feature built on a language model, for example two prompts, two models, two temperatures, or the same answer with and without retrieval from a knowledge base (RAG). The difference from offline evaluation is that the test measures the causal effect on user behavior, such as accepting the answer, finishing the task or coming back to the feature, rather than a score given by a grader on a fixed set of questions.
Does offline evaluation with an LLM judge replace an A/B test?
No. It is for filtering candidates before any user is exposed. The Microsoft Research article on evaluating LLM features states that offline evaluation suits early development but cannot assess how model changes benefit or degrade the user experience in production. And the judge has documented biases: in the study by Zheng and coauthors (2023), the most consistent of the three judges evaluated kept the same verdict after the answer order was swapped, with the default prompt, in only 65.0 percent of cases.
What should the randomization unit be in a prompt or model test?
The user, almost always. Randomizing per request makes a single conversation mix two behaviors, contaminates the per-user metric and makes retention impossible to measure. The statistical consequence is that per-message metrics, such as acceptance rate, become ratio metrics and require the delta method. In the simulation built for this guide, the naive per-message interval produced a false positive in 49.15 percent of 2,000 A/A tests in a heavy chat scenario, against 4.50 percent with the delta method.
Which metrics should an LLM A/B test use?
A per-user primary metric tied to task value, such as the share of users who accepted an answer without regenerating, plus usage signals (copied, applied, regenerated, rephrased) and LLM-specific guardrails: tokens and cost per request, high-percentile latency, error and refusal rates, and safety incidents. Explicit thumbs ratings work as a secondary signal, because only a fraction of users respond and that fraction is not randomized.
How long does a prompt A/B test take?
It depends on the baseline rate and the effect you want to detect. In this guide example, with 30 percent of users accepting an answer without regenerating and a minimum effect of plus 5 percent relative, at 95 percent confidence and 80 percent power, you need 14,856 users per variation. With 10,000 new eligible users per week, that is 21 days, three full weeks, which also helps you get past the novelty effect.
How do you decide whether a prompt that performs better but costs more is worth it?
Put gain and cost in the same unit and use the interval, not just the point estimate. In the illustrative example in this guide, the new prompt costs $101.69 more per 10,000 users and produces 126.9 additional users who accept an answer, about $0.80 per additional user. At the lower bound of the confidence interval, the same cost buys only 23.3 additional users, about $4.36 each. If an additional user is worth less than that to your business, the decision comes down to how much risk you are willing to take.