The Lab

What unblinding does to a clinical trial

Blinding isn't bureaucratic box-ticking — it's what keeps a trial's false-positive rate where it belongs. This simulation runs thousands of trials of a drug that secretly does nothing, and shows how quickly a broken blind turns it into a "success." Drag the sliders and watch the damage.

A Monte Carlo simulation run live in your browser: two arms (drug vs placebo), a continuous outcome in standard-deviation units, ~2,000 trials per setting. Unmasking adds a measurement bias to the drug arm proportional to how unblinded the trial is. Significance is a two-sided test at 5%. This models bias in a subjective endpoint — see the caveats below.

A first-draft write-up — the simulation is real; the words are a starting point to make your own. This one's your field, so treat my framing as a straw man to improve.

Every intro-to-trials course says the same thing: blind the trial. Patients shouldn't know what they're taking, and — just as importantly — the people measuring the outcome shouldn't know either. It's easy to file that under "good practice" without feeling why it matters. This simulation is an attempt to make it felt.

The setup: a drug that does nothing

Set true drug effect to zero. Now the drug is a dud — no better than placebo. In a perfectly run trial, it should clear the significance bar about 5% of the time, purely by chance. That 5% is the false-positive rate we agree to live with.

Now drag unmasking up. Each notch assumes the outcome assessors are a little more aware of who's on the drug — and, being human, they rate those patients a little more generously. The biology hasn't changed. Only the measurement has. Watch the damage curve climb off the 5% line.

Why the distribution tells the story

The second chart is the real mechanism. Each simulated trial produces one estimated effect; across thousands of trials those estimates form a distribution. Perfectly blinded, that distribution sits centred on zero, and only its thin tail pokes past the significance threshold — that's your 5%.

Unmasking doesn't widen the distribution. It slides the whole thing to the right. A bias that looks modest — a few tenths of a standard deviation — marches a large share of the distribution across the line. The trial isn't noisier; it's aimed wrong, and no sample size fixes a target that's off-centre.

Bigger trials don't rescue you here. More patients shrink the noise but not the bias — so a large, unblinded trial is a precise measurement of the wrong number.

Where the sample-size slider surprises people

Turn the sample size up with a nonzero bias in play. The false-positive rate doesn't fall — it often rises, because a larger trial is better at detecting the bias as if it were a real effect. Precision without validity is a trap, and it's one that a bias like this walks you straight into.

The honest caveats

This is a deliberately simple model, and it's built to make a point, not to be a trial-planning tool. A few things worth stating plainly:

  • It models a subjective endpoint. The whole mechanism is assessor bias on a rating. Hard, objective endpoints — mortality, an automated lab value — are far more resistant, though never fully immune (unblinding still shifts behaviour, dropouts, and co-interventions).
  • The bias magnitude is an assumption you set. Real-world unblinding bias isn't a dial; it varies enormously by endpoint, disease, and design. The slider lets you explore, not predict.
  • Unmasking is treated as a smooth fraction. Reality is lumpier — a blind either holds or breaks in specific ways. Read the x-axis as "how compromised," not a literal percentage.

Even with those simplifications, the qualitative lesson is robust and, I think, worth internalizing: blinding is not a formality layered on top of a good trial. For a subjective endpoint, it's part of what makes the trial's headline number mean anything at all.

Work together

Designing a study you need to get right?

Blinding, endpoints, power — the choices that decide whether your result means anything happen before data collection. That's exactly where I can help.