← Data Literacy for Everyone
Module 10 Free 6 min

"Is That Lift Significant?" — Statistical Significance in Plain English

What people really mean by "statistically significant," why sample size decides whether a result is trustworthy, and why "not significant" doesn't mean "no effect." A coin-flip explanation, no formulas.

What you'll learn

  • Explain "statistically significant" in plain English — bigger than ordinary random variation
  • See why a larger sample gives more confidence, and a tiny one can't be trusted
  • Know that "not significant" doesn't mean "no effect," and significant doesn't mean "big"

The A/B test from last lesson came back: “Create Account” converted at 4.0%, “Start Free” at 4.6%. Someone asks the question that gets asked in a thousand meetings a day — “is that lift significant?” — a data person says “yes, p was 0.03,” and the room nods as if that settled it. This lesson is the translation. By the end, “significant” will be a word you can use correctly, and — more usefully — question when someone leans on it too hard. And you’ll do it with zero formulas, because the idea underneath is just about luck.

Luck is noisier than you think

Flip a fair coin ten times and you won’t always get five heads. Small samples wobble — and statistical significance is how we ask whether a difference is bigger than that ordinary wobble.

First, the word lift: it just means the improvement of one result over another — 4.0% to 4.6% is a lift of 0.6 points, or 15% relative. Now the real question. If you flip two fair coins 100 times each, one will usually “beat” the other, sometimes by a lot — not because it’s a better coin, but because randomness is lumpy. Website visitors are the same: give the identical button to two random groups and their sign-up rates will still differ a little, purely by chance. So when B beats A by 0.6 points, the honest question isn’t “did B win?” — it’s “is this gap bigger than the wobble we’d expect from luck alone at this sample size?” That question, in plain English, is all “statistically significant” means: the difference is larger than the kind of difference ordinary random variation would produce.

The words that matter

Lift
The improvement of one result over another — the size of the gap you’re testing.
Random variation
The natural wobble in any measurement — two identical groups still differ a little, just by chance.
Statistically significant
The observed difference is bigger than ordinary random variation would plausibly produce. It does not mean the difference is large or important.
Sample size
How many people (or events) are in each group. Bigger samples wobble less, so real differences are easier to see.
p-value
In plain terms: if the change actually did nothing, how often would luck alone produce a gap this big? A small p (often under 0.05) means luck is a poor explanation.
Confidence
How sure we are that a result isn’t just luck — it rises with sample size and with the size of the effect.

Sample size is the whole game

Here’s the fact that makes significance click: the same gap can be meaningless or convincing depending only on how many people you tested. With a few hundred users per side, a 4.0% vs 4.6% gap sits comfortably inside the luck-wobble — flip the week and it might reverse. With tens of thousands per side, that same gap is far too large for luck to explain. Nothing about the effect changed; the sample size shrank the wobble until the gap stood out.

The same conversion gap with wide overlapping uncertainty at small sample and clear separation at large sampleTwo panels. Left: with 400 users per side, control at 4.0 percent and variant at 4.6 percent each carry tall uncertainty whiskers that overlap heavily, so the gap could be luck. Right: with 40,000 users per side, the same two values carry tiny whiskers that no longer overlap, so the gap is hard to explain by luck.Small sample: 400 per sidewobble overlaps — could be luckA: 4.0%B: 4.6%Large sample: 40,000 per sideclear separation — luck won't stretch this farA: 4.0%B: 4.6%

Same 4.0% vs 4.6% gap in both panels. Sample size is what shrinks the luck-wobble until the difference means something.

Text description of this diagram

Both panels show the identical result — control A at 4.0%, variant B at 4.6% — with vertical “whiskers” drawn around each dot to show the range luck alone could produce. In the left panel (400 users per side), the whiskers are tall and overlap heavily, so the amber note reads wobble overlaps — could be luck: at this size, the two values are within each other’s random range, and next week’s test might flip them. In the right panel (40,000 per side), the same two dots carry tiny whiskers that no longer come close to touching, and the green note reads clear separation — luck won’t stretch this far. Nothing about the effect changed between panels — only the sample size. That’s the core lesson: significance is a statement about the gap and the sample size together, never the gap alone.

Two traps significance sets

“Not significant” does not mean “no effect.” A real 40% improvement tested on 30 people will often come back “not significant” — not because the improvement is fake, but because 30 people is far too few to see through the wobble. That’s a verdict on the test’s power, not on the idea. The honest reading of a non-significant small test is “we couldn’t tell,” not “it doesn’t work.” Kill an idea on that basis and you may be throwing away something real that you simply measured too weakly.

Significant does not mean big. This is the one that trips up executives, and it’s the whole of the next lesson: with a huge enough sample, a laughably tiny difference becomes “statistically significant” — real, provable, and possibly worth nothing. Significance answers “is it luck?”; it stays completely silent on “is it worth doing?” Confidence, the size of the effect, and the sample size have to be discussed together — any one of them alone can mislead.

Common misunderstanding

“Statistically significant means the result is big and important.” It means only that the difference is unlikely to be pure luck at this sample size — nothing about its size or its value. A 0.02% lift across ten million users can be significant and trivial; a 40% lift across thirty users can be huge and non-significant. Significance is a luck-check, not an importance-check, and treating the two as the same is the single most common misuse of the word in business.

Try this at work

When someone says a result is significant, ask the disarming follow-up: “How many were in each group?” A big sample earns the claim; a tiny one means “we couldn’t really tell” whichever way it landed. And if a result is not significant on a small test, resist “it doesn’t work” — the honest phrase is “we haven’t tested it at a size that could show an effect.” Reflect: have you ever seen a promising idea killed by a test that was simply too small to detect it?

The bottom line

“Statistically significant” means a difference is bigger than ordinary random variation would produce — a luck-check, decided by the gap and the sample size together. Bigger samples wobble less; “not significant” on a tiny test means “couldn’t tell,” not “no effect”; and significant never means large or important.

Why it matters

“Is that lift significant?” is asked constantly and answered with vibes — you can now answer it with a straight face, and question a shaky yes.

You’ve got the luck half of the picture. The other half is the one significance can’t touch: even a rock-solid, luck-proof result might not be worth the cost of doing anything about it. Separating “real” from “worth it” is where a lot of money is won and lost — and it’s the whole of the next lesson.

Quick check

1. In plain English, "the lift is statistically significant" means…

2. A real 40% improvement on 30 users comes back "not significant." The right reading is…

3. Why does a bigger sample give more confidence?

Answers explained
  1. B is correct — significance is a luck-check: the difference is larger than random variation would plausibly produce at that sample size. (If you picked A: importance is a separate question the next lesson tackles. If you picked C: a manager’s review isn’t what the word means.)
  2. C is correct — 30 people is too few to see through the wobble, so “not significant” means the test lacked the power to detect the effect, not that the effect is absent. (If you picked A: absence of significance isn’t evidence of absence, especially in tiny tests. If you picked B: the threshold is about the p-value, not the size of the lift.)
  3. A is correct — larger samples reduce random variation, so a genuine difference stands out from the noise. (If you picked B: size makes differences detectable, not important. If you picked C: defining success is a separate discipline from sample size.)