How We Grade Evidence: The Hierarchy, Our Scale, and Why We Grade Down

Category: Science Explained

A single study proves less than it seems, and the type of study matters as much as the result. Here is the evidence hierarchy, the scale we use to grade herbs from strong to anecdotal, and why grading down is the honest choice.

The bottom line

Category: Science Decoded | Reading time: ~9 min | Level: Intermediate

Here is a sentence you will read a hundred times on supplement labels: studies show. It is doing a lot of quiet work. Studies show could mean a thousand people in a rigorous trial, or a handful of cells in a dish, and the label is counting on you not asking which. The difference between those two is the difference between a claim you can lean on and one that will not hold your weight.

This page is about that difference. It explains the evidence hierarchy, the ladder that ranks research by how much trust it has earned, and then it lays out the plain scale we use to grade every herb we write about. Most importantly, it explains why we so often grade a herb down, and why that is a feature rather than a flaw. If you understand how we grade, you will read every other page on this site with a sharper eye.

What Grading Evidence Actually Means

Grading evidence means judging how much confidence a claim deserves, based not just on the result but on how it was produced [1]. A finding is not simply true or false; it comes with a level of certainty, and that certainty depends on the design of the study, its size, its length, and whether other studies agree.

The idea in one sentence: the type of study tells you how likely the result is to survive contact with reality [1]. Formal systems in medicine, such as the GRADE approach used in clinical guidelines, do exactly this: they start from the study design and then adjust confidence up or down based on quality [1]. Our scale is a plainer version of the same logic, built for someone deciding whether a herb is worth their time and money.

The Hierarchy, From Test Tube to Meta-Analysis

Picture a ladder. Each rung represents a type of evidence, and the higher the rung, the better the design guards against the ways researchers and patients fool themselves [3].

In-vitro studies sit at the base. These are experiments on cells or molecules in a dish. They are excellent for working out how something might act, but a compound that changes a cell in a laboratory faces a gauntlet in a real body: digestion, absorption, distribution, and breakdown. Many never make it through. An in-vitro result is a hypothesis, not a human effect.

Animal studies come next. They add a living system, which is a step closer, but animals differ from people in metabolism and physiology, and a great many promising animal findings fail when tested in humans.

Human observational and open-label studies move us into people, but with weak controls. Open-label means everyone knows who is getting the treatment, which lets expectation shape the result. These can generate signals worth chasing, but they cannot separate a real effect from a placebo one.

Randomised controlled trials are the strongest single design. Randomly assigning participants to the treatment or a comparison, ideally with neither side knowing which, controls for bias and for the natural tendency to improve anyway [3]. A well-run RCT is where a claim earns real credibility.

Systematic reviews and meta-analyses sit at the top. They gather all the trials on a question and, in the case of a meta-analysis, pool them statistically for a more precise answer [2]. The important caveat is that a review is only as good as the trials inside it: pool weak studies and you get a confident-looking answer on a shaky base [2].

Our Four-Point Scale

We translate that ladder into a scale you can act on. Every herb we cover gets one of four grades, based on the best available human evidence and how consistent it is.

Strong. Multiple well-designed human trials, usually pooled in a meta-analysis, point consistently to the same effect for a specific outcome. This is rare for botanicals, and we do not hand it out lightly.

Moderate. There is real human trial evidence, but it is limited in quantity, quality, or consistency. The effect is plausible and supported, but not settled. A herb graded moderate is one where the honest phrasing is appears to help, in this group, over this timescale.

Limited. The human evidence is thin: one small trial, or several with conflicting results or methodological problems. A limited grade does not mean a herb does nothing. It means the evidence cannot yet support a confident claim, and you should hold your expectations loosely.

Anecdotal. The support is mostly traditional use, personal reports, or preclinical work, with little or no human trial evidence for the effect in question. Anecdotal is not an insult; some anecdotally graded herbs are worth a cautious trial. It is simply an accurate description of where the evidence stands.

Why We Grade Down

Here is the part that separates honest grading from marketing. When the studies behind a positive result have weaknesses, we lower our confidence even if the headline looks good. Formal systems do the same [1].

We grade down for small sample sizes, because a handful of participants can produce a chance result. For short durations, because a two-week effect may not last. For high risk of bias, because open-label designs and industry funding can tilt findings. For inconsistency, because if trials disagree, the truth is unsettled. For publication bias, because positive results get published and null ones quietly do not, which inflates the apparent effect. And for reliance on subjective measures, because how someone rates their own sleep is easier to move than what a sleep tracker records.

Grading down is not pessimism. It is precision. A positive trial with serious flaws is genuinely less reliable than a clean one, and pretending otherwise would sell you certainty that does not exist. When we say the evidence here is weak, or nobody has properly tested this yet, that is us doing our job. A cited null result buys more trust than ten benefit claims, and we would rather earn your trust than your impulse purchase.

What This Means for You

You can use this ladder yourself, on any product, in about thirty seconds. When you meet a claim, ask three questions. What kind of study supports it, in what species? How big and how long was it? Do independent reviews agree? If the only support is in-vitro or animal work, the claim is a hypothesis. If it is one tiny open-label trial, treat it as a signal, not a settled fact.

When you read our pages, the grade in the summary tells you how much confidence to invest before you spend a penny. A moderate grade invites a thoughtful trial with realistic expectations. A limited or anecdotal grade is a flag to weigh cost and safety carefully and to expect little. None of these grades is a yes or a no; they are a measure of how much the evidence has actually earned.

Safety Note

Evidence grade and safety are separate questions. A herb can have limited evidence for a benefit yet still carry real interaction or contraindication risks, and a well-evidenced herb can still be wrong for you. We cover safety under its own heading on every ingredient page for exactly this reason, and grading a benefit down never means a herb is automatically safe.

Pregnant, breastfeeding, or on medication? Check with a healthcare professional first.

The PlantRx Angle

This scale is the backbone of everything we publish. It is why our ingredient pages state a grade in plain words, why we name study sizes and durations, and why we will tell you an ingredient is barely tested rather than let a label imply otherwise. Honest grading is the product.

Our Remedy Library carries these grades so you can see at a glance how much the evidence supports each herb, and Remy can walk you through what a grade means for a specific ingredient. If you want to go further on spotting reliable information, our guide to finding supplement information that is not an advert is the natural next read.

References

1. Guyatt GH, Oxman AD, et al. (GRADE Working Group) (2008). GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ, 336, 924-926. Foundational description of grading confidence in evidence and the reasons to rate it down. 2. Higgins JPT, Thomas J, et al. (eds), Cochrane (2023). Cochrane Handbook for Systematic Reviews of Interventions. Methods reference on systematic reviews, meta-analysis, and the limits of pooling weak studies. 3. OCEBM Levels of Evidence Working Group, University of Oxford (2011). The Oxford Levels of Evidence. Framework ranking study designs from mechanistic and observational work up to randomised trials and systematic reviews.

Frequently asked questions

What is the evidence hierarchy?

It is a ranking of research types by how much confidence their results deserve. At the base sit test-tube (in vitro) studies and animal studies, which are useful for understanding mechanisms but do not tell you what happens in a person. Above them come human studies, rising through observational work and small open-label trials to randomised controlled trials, and finally systematic reviews and meta-analyses that pool many trials. The higher the tier, the better the design guards against bias and chance, so the more weight the finding carries.

Why is a test-tube study not enough to prove a supplement works?

Because a compound that does something to cells in a dish faces a completely different situation in the body. It has to survive digestion, be absorbed, reach the right tissue at a high enough concentration, and do so without being broken down first. Many compounds that look impressive in a test tube never clear those hurdles. In-vitro results are valuable for generating ideas about how something might work, but they are the beginning of the evidence, not the end, and treating them as proof is one of the most common ways supplement claims mislead.

What is the difference between strong and limited evidence?

Strong evidence means multiple well-designed human trials, ideally pooled in a meta-analysis, point consistently in the same direction for a specific outcome. Limited evidence means the human data is thin: perhaps one small trial, or several with conflicting results or methodological problems. A limited grade does not mean a herb does nothing; it means the evidence cannot yet support a confident claim. The gap between the two grades is about the quantity and quality of the studies, not about how popular or traditional the herb is.

What does it mean to grade evidence down?

Grading down means lowering our confidence rating when the underlying studies have weaknesses, even if the headline result looks positive. Reasons to grade down include small sample sizes, short durations, high risk of bias, inconsistent results across trials, publication bias, or reliance on subjective rather than objective measures. It mirrors how formal systems like GRADE work. We do it because an honest lower grade protects you better than an inflated one, and because a positive result with serious flaws is not the same as a reliable one.

Why do you sometimes say the evidence is weak?

Because it often is, and hiding that would make us untrustworthy. A large share of supplement claims rest on preclinical work, tiny trials, or traditional use rather than solid human evidence. Saying so plainly is the whole point of grading. A cited null result, an admission that nobody has properly tested something, or a limited grade tells you exactly where you stand. We would rather lose a sale to honesty than earn one by dressing a weak herb up as clinically supported.

Is traditional use a form of evidence?

It is a form of information, but a weak one on its own. Long traditional use can suggest a plant is broadly tolerated and worth studying, and it carries cultural value. What it cannot do is show that a herb produces a specific effect, because tradition does not control for placebo, natural recovery, or the many other reasons someone might feel better. We treat traditional use as a reason to investigate and sometimes as context, not as a substitute for human trials, and we grade accordingly.

What is a meta-analysis and why is it near the top?

A meta-analysis statistically combines the results of multiple studies on the same question to produce a single, more precise estimate. Because it draws on more participants than any one trial, it can detect effects and inconsistencies that individual studies miss, which is why it sits near the top of the hierarchy. The caveat is that a meta-analysis is only as good as the trials inside it: pooling weak or biased studies produces a confident-looking answer built on shaky foundations, so the quality of the included trials still matters.

A product says it is backed by science. How do I judge that claim?

Ask what kind of science, in what species, at what quality. Backed by science can honestly mean a large well-run human trial, or it can mean a single test-tube study the brand is stretching. Look for human randomised trials, sample sizes, durations, and whether independent reviews agree. If the only support is in-vitro or animal work, the claim is a hypothesis wearing a lab coat. The evidence hierarchy is the exact tool for turning a vague science claim into a specific question you can answer.

If a herb has a limited grade, should I avoid it?

Not necessarily. A limited grade is a statement about the evidence, not a verdict that the herb is useless or unsafe. It means you should hold your expectations loosely, watch for a real effect rather than assume one, and weigh cost and safety accordingly. Some limited-evidence herbs are gentle and low-risk enough to trial thoughtfully; others are not worth it. The grade helps you decide how much confidence to invest, which is different from a simple yes or no.

Why do different sources give the same herb different grades?

Because grading involves judgement about which studies to include and how to weigh their flaws, and reasonable reviewers can differ. One source might emphasise a positive trial while another discounts it for bias. Formal systems reduce but do not eliminate this. The practical response is to look at what the highest-quality evidence says, whether independent institutional reviews agree, and how transparent each source is about its reasoning. A source that shows its working is more trustworthy than one that just hands you a rating.

Sources

Related articles

Back to the PlantRx article library