CalculatorMasters

A/B Test Token Usage: Per Request vs Per User

Compare per-request, per-user, and traffic-weighted token metrics to choose the right view for an A/B testing analysis.

Token metrics answer different questions. Per-request usage shows model efficiency for a single interaction, per-user usage reflects behavior over time, and weighted usage summarizes the entire experiment allocation.

  • 100% Free
  • No Sign-Up Required
  • Private & Secure
  • Mobile Friendly

About A/B Test Token Usage: Per Request vs Per User

Token metrics answer different questions. Per-request usage shows model efficiency for a single interaction, per-user usage reflects behavior over time, and weighted usage summarizes the entire experiment allocation.

3

Comparisons

5

Key Factors

Instant

Results

100%

Free to Use

1

Per-Request vs Per-User Token Usage

Compare a single interaction metric with a participant-level metric.

FactorOption A: Tokens per RequestOption B: Tokens per UserWhat It Means
Primary unitOne model requestOne participating user over a periodThe appropriate unit depends on whether the question concerns individual requests or participant behavior.
Includes request frequencyNoYesPer-user usage multiplies average request tokens by average requests per user.
Best for prompt comparisonDirect view of prompt and response sizeCan include usage frequency effectsPer-request metrics isolate request design more clearly when frequency is unchanged.
Best for participant planningLess representativeMore representativeParticipant-level totals reflect the cumulative effect of repeated usage.
Effect of engagement changesNot shownShown through requests per userA variant that changes how often users interact can have a different per-user total even with similar request size.

Use tokens per request to inspect model interaction size and tokens per user to estimate typical participant-level consumption.

2

Variant-Level vs Traffic-Weighted Token Usage

Compare the usage of an assigned participant with the average usage across the full test allocation.

FactorOption A: Variant Tokens per UserOption B: Weighted Average Tokens per UserWhat It Means
Population representedUsers assigned to A or B onlyAll participating usersVariant results describe a cohort, while weighted usage describes the entire test.
Uses traffic allocationNoYesThe weighted result applies the percentage of traffic allocated to each variant.
Useful for comparing variantsYesIndirectlyThe A and B totals reveal the direct per-user usage gap.
Useful for overall experiment forecastRequires separate weightingYesThe weighted figure can be multiplied by expected participants for a high-level volume estimate.
Effect of changing a 50/50 split to 90/10No change within each variantChanges toward the majority variantAllocation affects the test-wide average but not the usage profile of either variant.

Use variant-level results to understand the A-versus-B difference and the weighted average to describe expected usage across the complete experiment.

3

Input Token Reduction vs Output Token Reduction

Compare two common ways a variant can reduce total model token volume.

FactorOption A: Reduce Input TokensOption B: Reduce Output TokensWhat It Means
Main source affectedInstructions, prompts, context, or retrieved contentGenerated response lengthThe relevant source depends on where most tokens occur in the request.
Per-request formula effectLowers input token componentLowers output token componentBoth reduce the same combined tokens-per-request total.
Potential behavioral impactMay change available model contextMay change answer detailToken reduction should be evaluated alongside the intended user experience and experimental outcome.
Measurement inputAverage input tokens per requestAverage output tokens per requestTrack each component separately to identify the source of change.

Both approaches lower total token usage when the reduced component is reflected in observed averages; the calculator adds the two components before calculating per-user totals.

Key Differences at a Glance

Tokens per request measures one interaction, while tokens per user includes average request frequency.

Variant-level usage describes a cohort; weighted usage describes the full traffic allocation.

Traffic split changes the blended experiment average but not a variant's own per-user total.

Input and output token differences both contribute directly to the combined per-request total.

Absolute token differences are measured in tokens, while percentage differences are relative to Variant B in this calculator.

How to Decide

Choose this if: Use a consistent measurement period for request frequency and token averages.
Choose this if: Compare variant tokens per user when assessing the size and direction of the A-versus-B usage difference.
Choose this if: Use the traffic-weighted result when summarizing expected average use across all test participants.
Choose this if: Review input and output tokens separately when diagnosing why one variant uses more tokens.
Choose this if: Consider high-usage cohorts separately when average values may hide long-tail behavior.

Assumptions

  • The comparison assumes the entered token figures are average values for comparable requests.
  • The per-user calculation applies one average request frequency to both variants.
  • Traffic shares sum to 100%, with Variant B receiving the remainder after Variant A.
  • All comparisons concern token volume rather than billing, quality, or business outcomes.

Related Comparisons

Frequently Asked Questions

Should I compare A/B variants by tokens per request or tokens per user?

Use per request to compare one interaction and per user to include how often a typical participant uses the feature.

Which metric should I use for the whole experiment?

Use weighted average tokens per user because it applies the traffic split to both variants.

Can a lower-token variant have a higher weighted average?

No, not if the same request frequency is used for both variants and the weighted result is calculated from only those two variants. The weighted average lies between their totals.

Does shifting traffic to a lower-token variant reduce the weighted average?

Yes. Increasing the share assigned to the lower-token variant moves the blended average toward that variant's per-user total.

Do lower token totals mean a variant is better?

Not by themselves. Token volume is one operational metric and does not measure response quality or experiment success.

Ready to calculate your result?

Try the calculator and compare options with your own inputs.

Try Calculator Free →