← Back to home

Five small labs, opened right in front of you.

Each lab is a core Marketing Science problem — running live in your browser, no backend, no login. Every click updates a real model parameter. The last one plays back.

✦ Live · Lab 01

Thompson Sampling, running live in your browser.

Built for intuition · seeded, so a reload replays the same run

Three arms, each with a hidden win probability. Every observed pull updates a Beta posterior; the policy balances explore (try uncertain arms) against exploit (lean into winners). Drag the true rates to change the world, or add feedback delay — the failure mode that makes a live bandit misbehave while still looking stationary.

Round 0
Reward 0
Regret 0.0
In flight 0
0.0 0.25 0.5 0.75 1.0 POSTERIOR · BETA(α,β) OVER WIN-RATE
The world True rates are hidden from the policy. Changing either restarts the run.
A
B
C

Push Arm C's delay past ~100 and watch the policy under-explore the arm that is actually best — its evidence keeps arriving too late to matter.

Cumulative regret All three policies race on identical reward draws — lower is better.
press play to start the race
Policy
Fatigue capmax share per arm
off

Everything here is seeded: the same seed replays the same run, and all three policies see the same reward draws, so the regret gap is the policy rather than luck. Read the engine — pure JS, no framework — demos/thompson-sampling/engine.js.

✦ Live · Lab 02

Two-tower retrieval, right in your browser.

User tower · profile embedding

Who is this customer?

Toggle behaviours to build the customer embedding. Each tag is one dimension of u ∈ ℝ⁶.

u = [ 0.00, 0.00, 0.00, 0.00, 0.00, 0.00 ] ‖u‖ = 0.00
Try a persona
Item tower · NBA actions

Next-best action ranking

Score = cosine(u, item) — the whole catalog re-ranks live as you tweak the user tower.

    In production, item embeddings are trained jointly with user embeddings via contrastive loss on click signal — here the item vectors are hand-crafted so the demo runs purely client-side.

    ✦ Live · Lab 03

    The difference between predicting and causing.

    Predictive vs causal · the distinction budgets are lost on

    Every dot is a customer, placed by how likely they are to convert without treatment (across) versus with it (up). A response model targets whoever is most likely to convert — which puts Sure Things at the top of the list, people who would have converted anyway. An uplift model targets the only group whose behaviour actually changes.

    Targeted 0
    Incremental 0.0
    Wasted spend 0%
    Who gets targeted Share of each true segment inside the selection.
    Incremental conversions
    Response 0.0
    Uplift 0.0
    Targeting policyResponse model
    Budgetshare of customers treated
    30%
    Effect heterogeneityhow much treatment effects vary
    1.00
    Budget saved 0%

    Drag effect heterogeneity to zero and the two policies converge — with no variation in treatment effect there is nothing for uplift modelling to find. That caveat matters as much as the win. demos/uplift-quadrants/engine.js.

    ✦ Live · Lab 04

    What the model looks at, and why.

    Behavior Sequence Transformer · single attention head

    A synthetic card-transaction history, oldest on the left. The prediction query asks "what comes next?" and attends back over the sequence — taller bar, more weight. Click any transaction to change its theme and watch the weights and the forecast move. Hover one to see what that token itself attends to.

    Next theme Softmax over the vocabulary, from the attention-weighted context vector.
    Recency priorhow much position beats content
    1.00

    This is genuine scaled dot-product attention — softmax(QKᵀ/√d)V — over hand-crafted theme embeddings rather than trained ones, so the behaviour is legible instead of accurate. Drop the recency prior to zero and position stops mattering entirely; push it up and only the last transaction survives. demos/attention/engine.js.

    ✦ End of labs

    Want to see the whole NBA system?

    These four labs are small slices. Head back to the main page for the full stack — RL, causal, MMM, uplift — running in production at Techcombank.

    ✦ Playable · Lab 05

    You allocate the budget. Thompson Sampling allocates the same budget.

    Bandit duel

    Forty-eight sends, three offers, hidden conversion rates. You choose one offer per round; the policy plays the same world in parallel and you do not get to watch it. Two rules make this the real problem rather than a guessing game: your results arrive three rounds late, and no offer may take more than 60% of sends. Keys 1 2 3 send, R restarts.

    Round 0 / 48
    Conversions seen 0
    Your regret —

    Cumulative conversions The policy's curve stays hidden until the game ends.
    Your allocation An offer goes dead once one more send would breach the cap.
    Results feed Three rounds behind, on purpose.

      Both sides run on the same reward streams, so a loss is a decision rather than luck. The engine reuses the Thompson Sampling lab's own BanditRun unchanged — game logic in demos/bandit-duel/engine.js.

      More to explore

      From playgrounds to reproducible experiments.

      Eleven additional field notes cover bandits, pricing, recommendation, causal inference, MMM, MLOps and release readiness. They use synthetic data or deterministic fixtures, with methods and downloadable evidence.

      Open the experiment library ↗