✦ Technical deep dive

How the models actually work.

The main page describes what the decisioning system does. This page is the layer underneath: the architectures, the numbers they were trained and served at, the trade-offs behind each choice — and, more usefully, the alternatives that were considered and rejected, with the reasoning. Written for someone who will disagree with parts of it.

More sections — two-tower retrieval, uplift modelling, off-policy evaluation — are being written. The live, interactive versions of both models below are on the labs page.

01

Behavior Sequence Transformer

Next-theme prediction over longitudinal card-transaction histories

The problem, stated precisely

A customer's card history is a chronologically ordered sequence of transactions. The task is to predict the next spending theme — not the next merchant — because themes are the unit the business can actually act on with an offer.

The naive framing treats this as tabular classification: aggregate the history into features (spend in category X over the last 30/60/90 days, counts, recency) and fit a gradient-boosted tree. That works, and it is the right baseline. What it throws away is order. "Three coffee purchases then a flight booking" and "a flight booking then three coffee purchases" produce identical aggregates and very different intent. Sequence models exist to keep that distinction.

Framing

Given a sequence of the most recent L transactions S = (s1, …, sL), each with a theme, an amount and a timestamp, estimate P(themeL+1 = c | S) for every theme c in the vocabulary.

Input representation

Every position in the sequence is a sum of embeddings, which is the standard transformer trick: the model sees one vector per transaction, but that vector carries several independent facts.

Sequence length100most recent transactions; older history truncated
Vocabulary89level-2 spending themes, extensible to ~1,200 merchant identifiers
Position signallearnedorder matters; absolute position embedding added to each token
Side featuresper-tokenamount bucket and time gap, embedded and summed into the token
Objectivesoftmax CEcross-entropy over the theme vocabulary at the prediction position

The choice of 89 themes rather than ~1,200 merchants is the single most consequential decision in this list. Merchant-level vocabulary is far more expressive and far sparser — a long tail where most identifiers appear too rarely to learn a useful embedding, and where the head is dominated by a handful of supermarkets and e-wallets. Themes trade resolution for density. The vocabulary is deliberately built so it can be swapped to merchant level without changing the architecture.

Architecture

A transformer encoder over the transaction sequence, with the prediction taken from a pooled representation concatenated with non-sequential customer features.

TRANSACTION SEQUENCE · L = 100 s₁coffee s₂grocery s₃fuel · · · s₁₀₀dining EMBEDDING theme + position + amount bucket + time gap → d_model TRANSFORMER ENCODER × N multi-head self-attention add & layer-norm position-wise FFN add & layer-norm POOL · last position CUSTOMER FEATURES NON-SEQUENTIAL tenure · product holding · demographics CONCATENATE MLP SOFTMAX P(next theme) over 89 classes

The part that earns the architecture is self-attention. A recurrent model compresses the whole history into a fixed hidden state and, in practice, discounts distant events. Attention scores every pair of positions directly, so a transaction 60 steps back can dominate the prediction if it is semantically relevant — a mortgage payment, a first-time FX transaction — while the eleven coffees in between contribute almost nothing. That is the behaviour the attention lab makes visible.

Training, promotion and serving

Splits6 / 1 / 1months train / validation / test, on automated rolling windows
Retrainmonthlyrolling window moves forward; no manual split curation
Promotion gate3ppchampion–challenger; a challenger ships only on a 3-point absolute gain
Inferencedaily batch~5M customers scored per day
Registry & lineageMLflowon Databricks, with Unity Catalog for data lineage

Splitting on time rather than at random is not a detail. A random split leaks the future into training: the model sees a customer's March behaviour while predicting their February, and offline metrics come back beautiful and meaningless. Rolling time-based windows are the only split that matches how the model is actually used.

The signal that ended the project

After several successive retrains the champion stopped being displaced — challengers kept landing inside the 3pp gate. Read correctly, that is not stability to celebrate; it is saturation. The architecture had extracted what this representation of the data could give, and further tuning was going to buy decimal places. That is why the research agenda moved to retrieval and long-sequence modelling rather than continuing to tune this model.

Alternatives considered

Roughly in order of increasing sophistication. The right-hand column is the honest reason each was not the choice for this problem, at this data scale — most of them are better models in other settings.

ApproachWhat it buysWhy not here
GBDT on aggregates Cheap, robust, interpretable; a genuinely strong baseline Discards order entirely. Kept as the baseline the sequence model has to beat.
Matrix factorisation / ALS Mature, fast, well understood collaborative signal Static user factors. No notion of "recently", which is most of the signal in card data.
GRU4Rec First strong session-based sequence model; light to serve Recurrence compresses history into one state and decays distant events — exactly what we needed to keep.
DIN Attention over history conditioned on a candidate item — very strong for CTR Built for scoring a known candidate. Our task is generating a distribution over all themes, not re-ranking a given one.
DIEN Adds an interest-evolution GRU layer on top of DIN Materially more complex to train and serve for a gain we could not justify at this vocabulary size.
SASRec Unidirectional self-attention; close cousin of what we built A reasonable alternative — the practical difference at L=100 with side features was small. BST's explicit slot for non-sequential features fitted the feature set better.
BERT4Rec Bidirectional context via a masked objective; usually stronger offline The bidirectional objective does not match causal serving, and cloze-style training added cost for an offline gain that did not survive the promotion gate.
TiSASRec Models the time interval between events explicitly Genuinely attractive for card data, where gaps carry meaning. Approximated with time-gap bucket embeddings instead; the full treatment is on the list.
SIM / ETA Two-stage retrieval over lifelong histories — thousands of events, not 100 The natural next step once L=100 is the binding constraint. It was not yet.
HSTU / generative recommenders Reframes recommendation as generative sequence transduction; scales impressively Infrastructure cost is a different order of magnitude. On the list to explore, not to ship next quarter.
TIGER (semantic IDs) Replaces item IDs with learned semantic codes via RQ-VAE; strong cold-start Most compelling at merchant-level vocabulary. Directly relevant if the vocabulary is expanded to ~1,200.

References

02

Thompson Sampling

Exploration under uncertainty, with constraints and late feedback

The problem, stated precisely

Once you know what to recommend, you still have to decide when and where. Each delivery option — a send-time window, a placement on a surface — has an unknown conversion rate. Every send is simultaneously a chance to earn and a chance to learn, and the two goals conflict.

The usual answer is an A/B test: split traffic, wait for significance, ship the winner. That is the right tool when you will make the decision once. It is the wrong tool here, because the decision recurs every campaign, the arms' true rates drift, and a fixed 50/50 split keeps paying full price for the losing arm long after the answer is obvious. A bandit reallocates continuously instead.

The algorithm

For binary outcomes, the Beta distribution is the conjugate prior of the Bernoulli, which makes the update arithmetic rather than inference. Keep one Beta per arm, sample from each, play the winner of the sample, update.

Prior θk ~ Beta(αk, βk)
Sample θ̃k ~ Beta(αk, βk)   for each arm k
Act at = arg maxk θ̃k
Update αa ← αa + rt  ·  βa ← βa + (1 − rt)
for t = 1, 2, 3, …
    for each arm k:
        theta[k] ~ Beta(alpha[k], beta[k])     # posterior sample
    a = argmax(theta)                          # probability matching
    r = pull(a)                                # observe 0/1 reward
    if r: alpha[a] += 1
    else: beta[a]  += 1

The elegance is that no exploration parameter appears anywhere. Thompson Sampling explores because the posteriors are wide, and stops exploring because they narrow — arm k is played with exactly the probability that it is optimal given the evidence so far. This is called probability matching, and it means the exploration schedule anneals itself. Compare ε-greedy, where ε is a number a human has to pick and then decay by hand.

On regret

Thompson Sampling attains O(log T) cumulative regret and is asymptotically optimal for Bernoulli bandits — it matches the Lai–Robbins lower bound. UCB1 achieves the same order with a deterministic optimism bonus rather than sampling. ε-greedy with a fixed ε suffers regret linear in T, because it never stops paying ε to explore arms it has already ruled out. The bandit lab races all three on identical reward draws so the gap is the policy, not luck.

What it looks like in production

Timing arms3send-time windows, one posterior each
Priorsseededfrom historical open/click conversion, not uniform Beta(1,1)
Update cadenceweeklyposterior refresh, matched to campaign rhythm
Entry-point surfaces3 × 5placements × carousel positions, plus business allocation rules
Result+7ppabsolute CTR lift, from a baseline where treatment and control were statistically indistinguishable
Hard constraint2 / daymaximum direct messages per customer per day

Seeding the priors matters more than it looks. A uniform Beta(1,1) start means the system spends its first weeks re-deriving things the business already knew, on live customers. Seeding from historical conversion turns the prior into a place to start arguing from, and the posterior overrides it as soon as evidence accumulates.

Two constraints that change the problem

Combinatorial action space. Three placements and five carousel positions is not fifteen independent arms — the slots interact, and business rules forbid many combinations outright. This is the slate/combinatorial bandit setting, and the neural variant is used here because a linear model cannot express the interaction between placement and position.

The per-customer cap. A hard limit of two messages per customer per day sounds like a guardrail bolted on at the end. It is not — it changes the mathematical object. Campaigns are no longer independent optimisation problems; they compete for a shared, scarce resource, which makes the daily decision a constrained assignment across all live campaigns rather than a per-campaign ranking. That is a knapsack-shaped problem sitting underneath the bandit, and pretending otherwise is how a well-tuned per-campaign optimiser ends up sending the same customer two messages that cannibalise each other. The fatigue-cap slider in the bandit lab is a toy version of this.

Delayed feedback — the failure mode that matters

Textbook bandits assume the reward arrives immediately. Ours does not. Push converges within days; in-app and email lag substantially. The system looks stationary the whole time.

Why this breaks a bandit quietly

A pull whose reward has not landed yet is indistinguishable, to the posterior, from a pull that returned nothing. Slow-feedback arms therefore look worse than they are, get sampled less, generate less feedback, and confirm the mistake. It is a feedback loop that looks like convergence: regret climbs, the posteriors are confident, and nothing in the dashboards says "wrong".

The bandit lab reproduces this exactly: push the best arm's feedback delay past ~100 rounds and its share of pulls collapses from ~87% to ~44%, with cumulative regret roughly tripling — while a worse, faster-reporting arm takes the traffic. Known treatments include modelling the conversion-lag distribution explicitly and treating unobserved outcomes as censored rather than negative, crediting delayed rewards retroactively when they land, and — cheapest of all — refusing to compare arms whose feedback horizons differ by an order of magnitude.

Alternatives considered

ApproachWhat it buysWhy not here
Fixed A/B test Unbiased, universally understood, trivially auditable Pays full price for the losing arm for the whole test, and answers once for a decision that recurs weekly.
ε-greedy Two lines of code; a fine sanity baseline Fixed ε gives regret linear in T. Requires a hand-tuned decay schedule to be competitive.
UCB1 Deterministic, same O(log T) order, easy to explain to risk Optimism bonus is tuned for worst-case, so it over-explores when priors are informative — and we had informative priors.
LinUCB Adds context through a linear reward model; strong, well-understood Linearity cannot express placement × position interaction. Kept as the contextual baseline.
NeuralUCB Neural reward model with a UCB-style bonus Confidence bonus needs a gradient-based approximation that is costly and fiddly at our update cadence.
Neural Thompson Sampling Posterior sampling over a neural reward model; handles interactions This is what shipped for entry-point surfaces.
EXP3 Adversarial guarantees; no stochastic assumption at all Our environment is stochastic and drifting, not adversarial. Paying the adversarial regret rate buys nothing.
Full RL (contextual MDP) Optimises long-horizon value, not immediate reward — the honest framing of customer lifetime Needs credit assignment over long horizons with sparse, late signal. The delayed-feedback problem above is the same obstacle, one level harder. Where this is going, not where it is.

Getting a policy past risk: off-policy evaluation

A bandit that explores on live customers needs a story for "what happens if the new policy is bad". Off-policy evaluation is that story: estimate the value of a candidate policy from logged data collected under the current one, before any traffic moves.

IPS V̂ = (1/n) Σi   ri · πnew(ai|xi) / πlog(ai|xi)

Unbiased, and violently high-variance when the logging policy assigned small probability to an action the new policy likes. Clipping the ratio trades a little bias for a lot of variance; doubly-robust estimators combine IPS with a reward model so that either one being right is enough. All of this depends on one unglamorous prerequisite: propensities must be logged at decision time. They cannot be reconstructed afterwards, which makes propensity logging an infrastructure decision, not a modelling one.

References

✦ More to come

Two-tower retrieval, uplift and OPE are next.

Both models on this page have a live, interactive version on the labs page — with the engines written as plain, testable JavaScript you can read.