WHY IT MATTERS
Earn while you learn.
In a campaign, every exploratory visit is a chance to discover a better variant — and a chance to miss a conversion. Bandits make that trade-off explicit.
Uniform random never learns. ε-greedy explores on 10% of visits. UCB1 adds an uncertainty bonus to each observed mean. Thompson sampling draws a plausible conversion rate from each arm’s Beta posterior, then chooses the largest.
This is an educational comparison, not evidence of uplift for an actual client. A business rollout also needs delayed-outcome handling, audience constraints and monitoring.
HOW TO READ THIS
Evidence, with boundaries.
Every run generates one Bernoulli outcome table indexed by visitor and variant. All policies face that same table but only observe the outcome of their chosen variant. Separate seeded random streams drive each policy.
Pseudo-regret sums the gap between the best true conversion rate and the chosen rate. It is non-negative and measured in expected conversions. It is not realised revenue loss. Conversion counts are actual simulated successes.
The interval is mean ± 1.96 standard errors across independent runs, clipped at zero for regret. It describes uncertainty in the mean, not the range of individual campaign outcomes. Overlapping bands alone do not establish a tie.
Rates are fixed, feedback is immediate, and parameters are not tuned. Equal-rate arms should produce exactly zero pseudo-regret. No algorithm wins in every environment.
Russo et al., A Tutorial on Thompson Sampling ↗ · Bandit Algorithms ↗