Most experimentation interviews don't ask you to design a test. They hand you a result and watch what you do with it.
The result is always designed to look like a win, and it always has something wrong with it — not because interviewers are cruel, but because that's the actual job. Designing a test is a day's work that a template can mostly do. Noticing that a statistically significant result should not be shipped is the skill that's hard to hire.
This page is about that half. For the design questions — sizing a test, what to do when you can't run one, handling a loss — the set is here: growth manager interview questions.
The worked result
"The new checkout button finished last night. Ship it?"
Users Conversions Rate Control 40,000 4,000 10.00% Variant 41,200 4,326 10.50% Relative lift +5.0%. Two-sided p = 0.019. Ran for six days.
Significant at the 5% level, a healthy-looking lift, plenty of traffic. Most candidates say ship.
Before any of the statistics:
"Before I look at whether the lift is real — the split is 40,000 against 41,200. If assignment is meant to be 50/50, that's off by 600 from the expected 40,600 in each arm. With this much traffic that's about a four-sigma imbalance, which happens by chance roughly once in forty thousand tests. So it almost certainly isn't chance, which means something in the assignment or the logging is broken.
And if assignment is broken, I can't interpret the conversion difference at all — whatever caused the imbalance may also have decided who ended up where."
That check is called sample ratio mismatch, it takes ten seconds, and it invalidates the entire result before any of the other analysis matters. Candidates who run it unprompted are rare and it is the single strongest signal available in this round.
A sample ratio mismatch is not a small problem to note and move past. It means the randomisation you're relying on didn't happen, so the two groups are no longer comparable and no amount of statistical machinery downstream repairs that. The common causes are mundane — a redirect that fails more often on one variant, bot filtering applied to one arm, an SDK that logs the exposure event before the variant renders — and all of them are the kind of bug that also correlates with conversion. The correct answer is to fix and re-run, not to analyse around it.
The second problem, if the split were clean
Suppose assignment checks out. The result still has an issue, and it's the one that catches strong candidates.
"At a 10% baseline, detecting a half-point absolute lift at 80% power needs about 56,000 users per arm. We have 40,000. So the test was underpowered for the effect we're claiming.
That doesn't mean the result is wrong. It means that among underpowered tests, the ones that reach significance are systematically the ones where noise happened to push in the same direction as the effect — so the measured 5% lift is very likely an overestimate of the real one. If I ship this expecting 5%, I should expect to be disappointed."
This is the winner's curse, and naming it separates people who have run experiments from people who have read about them.
The five traps
| Trap | The question that probes it | The answer in one line |
|---|---|---|
| Sample ratio mismatch | "Anything you'd check first?" | Compare the split to the intended ratio — a large deviation means randomisation failed |
| Peeking | "We saw significance on day two — could we have stopped?" | Not with a fixed-horizon test; repeated looks inflate the false positive rate far above 5% |
| Underpowering | "Is this enough traffic?" | Compute the sample size for the effect you care about before running, not after |
| Novelty and primacy | "Why six days?" | New things get clicked because they're new; established habits resist change. Both fade |
| Segment fishing | "It won on mobile — ship it there?" | Twenty segments means one false winner at p < 0.05 by construction |
On peeking, the answer that scores does more than say "don't": "If we want to stop early, that has to be designed in — sequential testing or a group-sequential boundary, decided before the test starts. What we can't do is run a fixed-horizon test and check it daily, because by day six we've taken six looks and our real false-positive rate is roughly double what we think it is."
On segments, the strong answer offers the alternative rather than just refusing: "I'd treat the mobile result as a hypothesis for the next test, not a finding from this one. If we genuinely expect the effect to differ by device, that's a pre-registered subgroup with the sample size planned for it."
Six days, and why the number matters
"Six days is an odd length and it worries me slightly. Weekly cycles are strong in checkout behaviour — weekends convert differently from weekdays — and a six-day test contains one weekend and a partial week, so the arms may not be balanced across days even if they're balanced overall.
I'd run in whole weeks. It's a small discipline and it removes an entire category of argument about whether Tuesday was unusual."
Reading the p-value
“p is 0.019, so we're confident at the 5% level. It's a 5% lift on checkout — that's a big number, let's roll it out.”
Everything stated is arithmetically correct and the conclusion is still wrong. The p-value answers one narrow question and says nothing about whether the two groups were comparable in the first place, which is the thing that's actually broken here.
Checking the setup before the result
“The split is off by four sigma. Before anything else, I'd want to know why — because if assignment is broken the conversion number isn't interpretable.”
Ten seconds of arithmetic that invalidates the whole analysis. It also demonstrates the habit the role needs: suspecting the instrument before trusting the reading.
The question underneath all of them
Interviewers are checking one disposition: would you be the person who stops a launch everyone wants?
Every trap on this page produces a result that somebody is excited about. The sample ratio mismatch appears on a test that won. The underpowered result is a lift someone has already told their manager about. Segment fishing usually happens because the overall test lost and someone needs a win.
So the answers that land aren't the technically complete ones. They're the ones that say what you'd do next, and to whom:
"I'd go back to whoever owns this and say: the split is broken, we can't read the result, and I'd like a week to re-run it clean. That's an annoying message to deliver and it's cheaper than shipping something we can't defend when the effect doesn't show up in the quarterly numbers."
Preparing for this round
Take a test you've actually run and check it for all five traps after the fact. Most people find at least one, and the specific memory of "we called a winner off nine days of data and the effect vanished" is worth more in an interview than any amount of theory.
Then practise the sentence that stops a launch. It's short, it's uncomfortable, and being fluent in it is most of what this role is.
For the case-study format — where you're given a funnel and asked what to do rather than a result and asked whether to ship — the worked version is here: growth case study interview.
Practical target: take the table at the top of this page and say the first thirty seconds of your answer out loud. If the first number you mention is the p-value or the lift, you've started in the wrong place. The split, the duration and the power calculation all come before the result — and saying them first is what makes the rest of the answer sound like someone who has been burned.




