A first-draft write-up — the simulation is real; the words are a starting point to make your own. This one's your field, so treat my framing as a straw man to improve.
Every intro-to-trials course says the same thing: blind the trial. Patients shouldn't know what they're taking, and — just as importantly — the people measuring the outcome shouldn't know either. It's easy to file that under "good practice" without feeling why it matters. This simulation is an attempt to make it felt.
The setup: a drug that does nothing
Set true drug effect to zero. Now the drug is a dud — no better than placebo. In a perfectly run trial, it should clear the significance bar about 5% of the time, purely by chance. That 5% is the false-positive rate we agree to live with.
Now drag unmasking up. Each notch assumes the outcome assessors are a little more aware of who's on the drug — and, being human, they rate those patients a little more generously. The biology hasn't changed. Only the measurement has. Watch the damage curve climb off the 5% line.
Why the distribution tells the story
The second chart is the real mechanism. Each simulated trial produces one estimated effect; across thousands of trials those estimates form a distribution. Perfectly blinded, that distribution sits centred on zero, and only its thin tail pokes past the significance threshold — that's your 5%.
Unmasking doesn't widen the distribution. It slides the whole thing to the right. A bias that looks modest — a few tenths of a standard deviation — marches a large share of the distribution across the line. The trial isn't noisier; it's aimed wrong, and no sample size fixes a target that's off-centre.
Bigger trials don't rescue you here. More patients shrink the noise but not the bias — so a large, unblinded trial is a precise measurement of the wrong number.
Where the sample-size slider surprises people
Turn the sample size up with a nonzero bias in play. The false-positive rate doesn't fall — it often rises, because a larger trial is better at detecting the bias as if it were a real effect. Precision without validity is a trap, and it's one that a bias like this walks you straight into.
The honest caveats
This is a deliberately simple model, and it's built to make a point, not to be a trial-planning tool. A few things worth stating plainly:
- It models a subjective endpoint. The whole mechanism is assessor bias on a rating. Hard, objective endpoints — mortality, an automated lab value — are far more resistant, though never fully immune (unblinding still shifts behaviour, dropouts, and co-interventions).
- The bias magnitude is an assumption you set. Real-world unblinding bias isn't a dial; it varies enormously by endpoint, disease, and design. The slider lets you explore, not predict.
- Unmasking is treated as a smooth fraction. Reality is lumpier — a blind either holds or breaks in specific ways. Read the x-axis as "how compromised," not a literal percentage.
Even with those simplifications, the qualitative lesson is robust and, I think, worth internalizing: blinding is not a formality layered on top of a good trial. For a subjective endpoint, it's part of what makes the trial's headline number mean anything at all.