lately but it’s truly a topic that fascinates me and that’s why I keep on doing it.
In today’s post I want to see how it affects and fools us in our A/B tests, by using a fictional example to illustrate better what I’m trying to share.
Yes, it’ll be football-based, but stay with me because this applies to every single field where A/B testing is possible (and that’s every field that exists). Plus, at the end, I try to generalize it so we don’t talk football all the time.
Hope you enjoy it!
A Surprising Win
There’s a new coach in our favorite team and he loves data. So much so, that every decision he makes is based on it, and none is uninformed.
The team is famous for being the slowest in the league, which has horrible consequences: they receive the most counterattacks (and goals from those situations). That’s the main reason why they lose most matches, because they do well tactically but can’t stop those fast breaks.
So the new coach, very methodical and experienced, thinks that a good warm up is key to make people run faster. But he wants to prove it and he decides to run a typical A/B test.
The A/B test is simple: the squad is divided in two groups where one keeps on warming up as usual (group A) while the other is instructed a new warm-up routine (group B).
After only four weeks, group B’s sprint times are 8% faster. Clear win? Or maybe just randomness.
It’s just like the monkeys and typewriter analogy: gather an infinite number of monkeys with typewriters and you’ll be certain at least one will come up with the Iliad.
Hence that lucky monkey achieving that seemingly impossible outcome will be perceived as a genius—yet it will most probably be pure randomness.
In the case of the 8% improvement in sprint time, the same thing can happen: the coach might have been fooled by randomness by believing the new warm-up justifies the improvement (without doing any other check).
The Problem: Random Noise Looks Like a Win
On paper our coach’s test looks convincing. The sprinting performance has increased as the group B’s average has improved by 8% in only four weeks.
The team is willing to adhere to the new warm-up routine as soon as possible. They believe it can save them from relegation.
But small datasets like the team’s can be really dangerous. After all, there are only 24 members in the squad and one unusually good or tired session can swing the averages dramatically.
Add in the intangibles like mood, sleep quality, motivation, weather and even the time of day. With such a randomness-prone environment, the odds of finding something “significant” by chance shoot up.
This is exactly the same trap online marketers fall into when they test dozens of ad variants and crown whichever one looks best after a few days. It was all luck, probably.
Just like the monkeys: the more monkeys (ad variants), the more chances of having one outperforming monkey (variation).
Now, I’m not stating that the new warm-up routine isn’t working, or that the winning ad variant is mediocre. What I mean with all this is that without careful design and analysis, a one-off spike can masquerade as a breakthrough.
What looks like a “winner” may simply be random noise.
How to Tell Signal from Noise
Once the coach was aware of the potential problem, he came to us. He wanted to learn how to tell if the results were reliable or not.
The short answer to tell signal from noise is to make your analyses more complex. But let’s see some ways to do so:
- Pre-define your hypothesis and metric. Don’t do like the coach who just performed the tests without defining what success meant for him. It’s when he saw an 8% that he decided it was good… But what if it’d had been a 5% or 3%? Would he had considered them good enough to accept the hypothesis?
- Randomize fairly. Both groups have to be properly balanced. In the case of our team, they were balanced for age, position, and injury history so one side isn’t advantaged from the start.
- Use crossover or repeated-measures design.
- Track context variables. The coach failed to record fatigue scores, weather, and workload so he couldn’t to adjust for confounders.
- Apply appropriate statistics. Yup, don’t stick to the basic stuff. Correct for multiple comparisons, or use Bayesian or hierarchical models that handle small, variable datasets more gracefully.
- Look for replication. This is one of the most important points: if the result holds when the coach repeats the test in another block or season, it’s more likely to be real (yet not enough to determine it).
So our advice to the coach after telling him all these tips would be to alternate routines every month and analysing performance across several cycles, rather than declaring victory after one single block.
Generalizing Beyond Sports
The warm-up story is just a vivid case study, but the same pitfalls show up in every single A/B test one performs.
Just like the one we mentioned in marketing, where one ad variant outperforms the rest after a few thousand impressions but, once chosen and implemented, we see that it’s no longer performing as expected.
Or another example, now in the healthcare world: experts often see small pilot trials produce dramatic effects that vanish in larger randomized controlled trials (that’s why they do them, btw).
The pattern is always the same: random variation creates illusions of success. And the antidote is also the same: careful experimental design, appropriate statistical corrections, and replication.
Please, don’t confuse a monkey’s lucky keystrokes with Shakespeare.
Closing
Group B’s gains looked like magic. But without proper controls, that magic can vanish faster than a greased pig. A/B testing is powerful, but only if you treat randomness as an opponent to outsmart and not a fluke to celebrate.
Don’t be fooled.
As for the coach: he listened us and performed the tests properly, failing to see an improvement with the new warm-up and therefore being unable to make the team faster.
They avoided relegation, though, by playing an extremely-defensive style that just didn’t create fast break opportunities for the opponents.
Happy ending, I guess.