Three arms, each with a hidden win probability. Every observed pull updates a Beta posterior; the policy balances explore (try uncertain arms) against exploit (lean into winners). Drag the true rates to change the world, or add feedback delay — the failure mode that makes a live bandit misbehave while still looking stationary.
Push Arm C's delay past ~100 and watch the policy under-explore the arm that is actually best — its evidence keeps arriving too late to matter.
Everything here is seeded: the same seed replays the same run, and all three policies see the same reward draws, so the regret gap is the policy rather than luck. Read the engine — pure JS, no framework — demos/thompson-sampling/engine.js.