Category: Science Explained
A randomised controlled trial is the strongest single way to test whether something actually works, because it is built to defeat the ways we fool ourselves. Here is how randomisation, blinding and placebo control do that, and where even a good RCT falls short.
Category: Science Decoded | Reading time: ~9 min | Level: Intermediate
People believe supplements work for reasons that have nothing to do with the supplement. They start taking something when they feel at their worst, so they improve partly because they were always going to. They expect a pill to help, so it does, a little, whatever is in it. And they remember the times it seemed to work and forget the times it did not. Every one of these is a real, powerful way to be fooled, and none of them requires the product to do anything.
The randomised controlled trial exists to strip all of that away. It is not a fancier version of an ordinary study; it is a machine engineered specifically to catch the ways we mislead ourselves. That is why it sits at the top of the evidence hierarchy for a single study, and why phrases like shown in a randomised trial carry real weight. This is a plain guide to how an RCT works, why each of its parts matters, and, just as important, the limits that mean even a good one is not the final word.
A randomised controlled trial is a study that tests whether an intervention causes an effect by randomly assigning participants to receive it or a comparison, then measuring the difference between the groups [1][2]. The intervention might be a supplement, a drug, or a programme; the comparison is often a placebo.
The idea in one sentence: give the treatment to one randomly chosen group and not to another, keep everything else as equal as possible, and any difference in outcome points to the treatment [1]. The power lies in that comparison. Because the two groups are formed at random and treated identically except for the intervention, a difference between them is hard to explain by anything other than the intervention itself. This is what lets an RCT do something weaker studies cannot: move from this is associated with that toward this causes that.
An RCT earns its status through three design features, each aimed at a specific way of being fooled.
Randomisation spreads confounding factors evenly [1]. If researchers chose who received the treatment, the groups might differ in age, health or motivation, and those differences could masquerade as an effect. Assigning people at random means that, on average, both known and unknown factors are balanced across the groups, so the comparison is fair. This is the feature that most distinguishes an RCT from a weaker observational study.
Blinding stops expectation from tilting the result [2]. In a single-blind trial the participants do not know which group they are in; in a double-blind trial neither the participants nor the assessors know. Because expectation shapes both how people feel and how results are judged, blinding ensures the measured difference reflects the intervention rather than everyone's hopes about it.
Placebo control reveals the intervention's own effect [2]. People often improve after taking something they believe will help, through expectation and natural recovery. Comparing against an inactive placebo made to look identical shows how much of the improvement is due to the intervention and how much would have happened anyway. Remove the placebo and an apparent benefit could be nothing but the placebo effect plus natural recovery.
Together, these three defeat the trio of natural recovery, the placebo effect, and unconscious bias, the exact traps that make anecdotes so unreliable.
Among single studies, the RCT is the design best equipped to establish that an intervention actually causes an effect rather than merely being linked to one [3]. Observational studies can show that two things occur together, but they struggle to prove cause, because the people who do one thing often differ from those who do not. The RCT is built specifically to test causation, which is why it ranks so high in the evidence hierarchy [3].
Gold standard is a fair description, but it should not be read as flawless. It means that, for the question does this work, a well-designed RCT is the most trustworthy single tool available. It does not mean any study calling itself an RCT is automatically strong, nor that one trial settles a matter. Those qualifications lead straight to the limitations.
A randomised trial is a powerful instrument, and like any instrument it has a range outside which it tells you less.
A trial can be too small, so its result may be down to chance or too imprecise to trust. It can be too short, missing long-term benefits or harms. It may enrol a narrow group of people, so its finding might not transfer to those who differ in age, sex, health or background. It answers one specific question, with one dose, one preparation and one chosen outcome, not the broader question you might actually care about. And trials are not immune to influence: funding source, how outcomes are selected, and how they are measured can all shape the result.
The most important limitation is simply that one trial is one trial. Any single study can land on a chance result or carry quirks in its design, which is why researchers prize replication and why systematic reviews and meta-analyses, which pool many trials, sit above individual RCTs in the hierarchy [1][3]. A single RCT is strong evidence. Several good RCTs agreeing is stronger.
When a product cites a randomised trial, treat it as an invitation to ask better questions rather than a closed case. How many people took part, and for how long? Who were they, and do they resemble you? What dose and preparation were used? Was it blinded and placebo-controlled? Who paid for it? A small, short, industry-funded trial in an unrepresentative group is far weaker than a large, independent, well-blinded one, though both wear the same RCT label.
And hold single results loosely. If one trial says a herb helps and another says it does not, that is not the method failing; it is a signal to look at the whole body of evidence rather than the most eye-catching study. The phrase backed by a randomised trial hides a wide range of quality, and knowing what to ask is what turns it from a marketing line into real information.
Understanding trial design is about judging efficacy claims, but it connects to safety too. Short trials in narrow populations may not reveal uncommon or long-term harms, so an intervention shown to work in a trial has not necessarily been proven safe for everyone over time. Efficacy evidence and safety evidence are separate, and both deserve scrutiny.
Pregnant, breastfeeding, or on medication? Check with a healthcare professional first.
The RCT is the standard we hold ingredients to, and it is why our pages name study sizes, durations and designs rather than waving at studies show. When the best evidence for a herb is a single small trial, we say so, because the difference between one small RCT and a body of solid ones is the difference between a hint and a conclusion.
Our Remedy Library grades herbs on the strength of their human evidence, and Remy can explain what a given trial does and does not establish for a specific ingredient. To see how single trials fit into the wider picture, our guides to grading evidence and to meta-analysis are the natural companions to this one.
1. Higgins JPT, Thomas J, et al. (eds), Cochrane (2023). Cochrane Handbook for Systematic Reviews of Interventions. Methods reference on trial design, randomisation, blinding, control groups and the value of replication. 2. Zabor EC, Kaizer AM, Hobbs BP (2020). Randomized Controlled Trials. Chest, 158(1S), S79-S87. Methods review on RCT structure, blinding and placebo control. 3. OCEBM Levels of Evidence Working Group, University of Oxford (2011). The Oxford Levels of Evidence. Framework placing randomised trials and their systematic reviews at the top of the hierarchy for questions of treatment effect.
It is a study designed to test whether an intervention, such as a supplement, drug or programme, actually causes an effect. Participants are randomly assigned to either receive the intervention or a comparison, often a placebo, and then followed to see whether the groups differ on a chosen outcome. The random assignment and the comparison group are what make it powerful: they let researchers separate the intervention's real effect from everything else that might make people improve. It is the strongest single study design for answering the question does this work.
Because it spreads the differences between people evenly across the groups. If researchers chose who received the treatment, the groups might differ in ways that affect the result, such as age, health or motivation, and those differences could masquerade as an effect. Randomly assigning people means that, on average, both known and unknown factors are balanced between the groups. So any difference in outcome is more likely to be due to the intervention itself rather than to a lopsided comparison. Randomisation is the single feature that most distinguishes an RCT from a weaker study.
Blinding means keeping people unaware of who is receiving the intervention and who is receiving the comparison. In a single-blind trial the participants do not know; in a double-blind trial neither the participants nor the researchers assessing them know. This matters because expectation shapes both how people feel and how researchers judge results. If someone knows they are taking the active treatment, they may report feeling better simply because they expect to. Blinding removes that influence, so the measured difference reflects the intervention rather than everyone's hopes about it.
A placebo is an inactive comparison, such as a dummy pill, designed to look identical to the real intervention. It is used because people often improve after taking something they believe will help, an effect driven by expectation and by the natural course of many conditions. By comparing the intervention against a placebo, researchers can see how much of the improvement is due to the intervention itself and how much would have happened anyway. Without a placebo, an apparent benefit could be nothing more than the placebo effect and natural recovery combined.
Because, among single studies, it is the design best equipped to establish that an intervention actually causes an effect rather than merely being associated with one. Its combination of randomisation, blinding and a control group defeats the main ways people are misled: natural recovery, the placebo effect, and conscious or unconscious bias in who gets what and how results are judged. Observational studies can show that two things go together but struggle to prove cause. A well-designed RCT is built specifically to test causation, which is why it sits so high in the evidence hierarchy.
Several. A trial can be too small to detect or trust an effect, or too short to reveal long-term benefits or harms. It may be run in a narrow group of people, so the result might not apply to everyone. It answers one specific question with one specific dose and outcome, not every question you might have. Trials can also be affected by funding bias and by how outcomes are chosen and measured. So a single RCT is strong evidence, but it is one study, and confidence grows when several good trials agree.
Usually not on its own. A single well-run RCT is meaningful evidence, but any one study can produce a result by chance, be limited to a particular population, or have quirks in its design. This is why researchers value replication and why systematic reviews and meta-analyses, which pool multiple trials, sit above individual RCTs in the evidence hierarchy. A confident conclusion rests on several good trials pointing the same way, not on one striking result, however well designed that single trial may be.
It is a good sign, but it is the start of the questions, not the end. Ask how many people were in it, how long it ran, who they were, what dose and preparation were used, whether it was blinded and placebo-controlled, and who funded it. A small, short, industry-funded trial in an unrepresentative group is far weaker than a large, independent, well-blinded one, even though both are RCTs. The label backed by a randomised trial can hide a wide range of quality, so the design details are where the real judgement lies.
Because individual trials differ in size, population, dose, duration, outcome measures and quality, and any one can also land on a chance result. Two honest RCTs can reach different conclusions simply because they studied slightly different things or different people, or because one was underpowered. This is not a failure of the method; it is why replication matters and why meta-analyses exist to weigh the whole body of trials together. Contradiction between single trials is a reason to look at the overall evidence, not to dismiss trials altogether.
Because RCTs are not always possible, ethical or practical, and they answer narrow questions. You cannot randomise people to lifelong habits easily, and some questions about long-term or rare outcomes are better addressed by observational studies. RCTs also test average effects in defined groups, which may not capture individual variation. So other designs contribute real information, especially where trials are impractical. The evidence hierarchy ranks RCTs highly for testing whether something works, without claiming they are the only useful form of evidence.