Thoughts··9 min
The Cost of Not Running Experiments
In banking, the decision that feels free, shipping a change without a controlled experiment, is usually the expensive one. And the most expensive version of it hides inside credit underwriting.
A bank’s costliest decisions rarely look costly on the day they’re made. They look obvious.
Picture the chart everyone nods at. Customers who get a payment reminder three days before the due date miss fewer payments. The gap is huge, the line is clean, and the room agrees in under a minute: send the reminder to everyone.
The chart is real. The correlation is real. The conclusion is almost certainly wrong. Customers who opt into reminders, or who happen to be reachable three days out, are already more organized, more liquid, and more likely to pay on time no matter what you do. The reminder didn’t cause the good behavior; they share a common cause. Roll it out to everyone and you may spend real money moving nothing, while congratulating yourself on a number that was always going to be there.
After enough years around credit, fraud, and growth models, I’ve come to think this is the shape of almost every expensive mistake in banking. Not a dramatic blow-up, but a confident decision built on a comparison that was never fair. The antidote is boring, and it works: run the experiment.
What an experiment actually is
An experiment is not a dashboard, a backtest, or a pilot on a friendly branch. It’s a deliberate, randomized comparison where the only systematic difference between two groups is the decision you’re testing. In banking that usually takes one of a few shapes:
- A randomized holdout. A small random slice of customers kept on the old treatment on purpose, so you always have a clean baseline to measure against.
- Champion/challenger. The current model or policy keeps most of the traffic while a challenger earns promotion on measured results instead of a meeting.
- An A/B test on a digital surface, where the flow or the offer is randomly assigned so uplift reads directly.
The common thread is randomization. It’s the one thing that lets you say “caused” without lying. Everything else, however sophisticated the model, is a story about a comparison the data never actually ran.
The cost is invisible, which is what makes it dangerous
When a trading desk loses money, it’s on the P&L. When you skip an experiment, nothing shows up anywhere. There’s no line item for “value we’d have captured if we’d known.” That’s the trap: the cost is an opportunity cost, and opportunity costs don’t send invoices. They pile up quietly, in a few ways that feed each other:
- You ship things that don’t work and never find out. Without a holdout, every launch looks successful, because the world keeps moving and something always changes afterward. Collections calls go out and repayment ticks up, but it was seasonal anyway.
- You kill things that would have worked. A challenger has one bad quarter for reasons that have nothing to do with the challenger, gets shelved, and nobody ever learns it would have lifted margin at scale. Good ideas die of bad luck.
- Your models decay and you fly blind. The only honest way to know a challenger model beats the incumbent is to run them side by side on randomized traffic. Skip that and model selection becomes a debate about offline metrics on data that no longer looks like today.
- The organization stops being able to learn. When decisions are settled by seniority instead of evidence, being confidently wrong carries no penalty, and the institution’s beliefs drift away from reality one unchallenged assumption at a time.
That last one is the one I worry about most, because you can’t point to the day it happened.
The death spiral hiding in your underwriting
Selection bias in a marketing test costs you a campaign budget. Selection bias in credit underwriting can quietly kill an entire lending product, and every metric will look healthy right up until it does. It’s the most expensive version of this mistake I know, so it’s worth walking through slowly.
When you underwrite a loan, you only ever learn the outcome of the applicants you approve. The ones you decline never take the loan, so you never find out whether they’d have repaid. Your model is trained, forever, on the survivors of its own past decisions. This is the selective labels problem, and it has teeth.
Watch what happens when losses tick up and nobody is running an experiment. Risk tightens the cutoff. It’s the obvious move, and it half-works, because the excluded band did contain some bad borrowers. But a cutoff is blunt: it drops good marginal borrowers along with the bad, because at the margin the model can’t actually tell them apart, and the margin is exactly where its labels are thinnest.
Approvals shrink and the book concentrates in the safest, most prime borrowers, who are also the most price-sensitive and the first to refinance away. Volume falls, fixed costs spread over fewer loans, unit economics worsen. Then the loop closes: the model retrains on the approved population, which now contains only the narrow slice it already trusted. It has no counterfactual from the region it just cut, so from inside its world the excluded band looks purely bad. The rational next step is to tighten again. And again.
losses rise on the book
|
v
tighten the cutoff <----------+
| |
v |
good borrowers cut too |
| |
v |
book narrows, volume falls |
| |
v |
retrain on survivors only |
| |
v |
rejects look all-bad ----------+
Each turn looks locally responsible. Losses on the approved book stay contained, so every quarterly review passes. Nobody sees the profitable borrowers being turned away, because rejected applicants don’t show up on any dashboard. The product strangles itself while every metric anyone is watching says the underwriting is working. No single step is wrong. The loop is.
The second gate: approved, mispriced, and gone
Rejection is only the first place selection bias hides. There’s a second gate right behind it, and it’s easier to miss because on paper it looks like success.
Approval isn’t a booking. It’s an offer, and the customer decides whether to accept it. When your pricing isn’t competitive, the customers with the best outside options walk, and those are usually your best risks. The ones who accept a mediocre offer skew toward people with nowhere better to go, who skew riskier. So the loans you actually book come in worse than the population you approved, even though nothing went wrong at the underwriting gate. This is adverse selection on take-up, and it feeds the same loop.
You only observe repayment on disbursed loans, a doubly-selected sample: approved, then self-selected by who was willing to accept your price. You never see how the approved-but-declined would have performed, and you never learn what price would have won them back. Booked losses come in high, risk reads it as “the population is riskier,” raises price or tightens further, and the next cohort of customers-with-options walks too. The pricing looks like it never works, and in a sense it can’t, because it’s being set from a book the price itself keeps poisoning. The tell is a product where approvals look healthy but the booking rate is quietly collapsing, and the loans that do disburse underperform.
Breaking the spiral is a strategy, not a model tweak
The trap is structural, so the fix has to be too. You can’t out-model biased data, because the bias is baked into which data you’re allowed to collect. The only way out is to deliberately collect the data the loop refuses to give you. That’s an experiment.
In underwriting it has a well-worn form: approve a small, randomly chosen fraction of applicants who fall just below the cutoff, and fund them on purpose. People call it a below-the-line test, a randomized swap-in, or just test lending, and disciplined lenders have run versions of it for decades. You accept a small, budgeted amount of expected loss on that random slice. In exchange you get the one thing the champion policy can never produce on its own: unbiased outcome data from the region you’d otherwise reject blind. More often than not the band just below the line turns out not to be uniformly bad, and a sharper policy can swap some of them in while dropping weak approvals sitting above the line.
Price needs the same treatment, for the same reason. The only way to tell a price that’s simply too high from a product nobody wants is to randomize the offered rate within a governed band and watch what take-up does, measuring two things at once: how many accept, and how the ones who accept go on to repay. That’s the only honest way to see the take-up curve and its interaction with risk, instead of inferring both from a book your current price already bent out of shape.
Run all of it as champion/challenger and it stops being a one-off study and becomes a permanent learning engine, with losses inside a governed envelope the whole time. Which, conveniently, is exactly the story risk and compliance want to hear.
“But we’re a bank, we can’t just randomize people”
It’s the most common objection, and it’s fair. You can’t randomly deny someone credit to satisfy your curiosity, and every policy change lives inside a governance process for good reason. But those constraints shape experiments, they don’t forbid them. You randomize the treatments you’re already allowed to offer. You run policy changes as champion/challenger inside approved bounds, so the challenger only ever operates in a risk envelope governance has already signed off. You hold out on interventions, not on rights.
And here’s the part risk-averse institutions miss: a well-run experiment isn’t a compliance risk, it’s a compliance asset. When a regulator asks whether a decision was fair, effective, and evidence-based, “we ran a controlled experiment and here’s the measured effect” beats “our most senior person was confident.” The discipline regulators want and the discipline experimentation provides are the same discipline.
Where I keep landing
Every key business decision is a bet. You’re going to make the bet either way, so the only real choice is whether you make it with evidence or with opinion. Not running the experiment doesn’t save you the cost of the decision; it just moves that cost somewhere you can’t see it, where it grows, unaudited and uncorrected. The experiment that feels expensive is almost always cheaper than the confident rollout that felt free.
The rule of thumb I keep coming back to: if a decision is big enough to argue about, it’s big enough to hold out a slice and measure. Measure the thing. Then decide.