How we rate evidence.
A strength rating is a judgement, and a judgement is only useful if you can see how it was made. This page describes the process, including the parts of it that are unavoidably a matter of interpretation.
What gets included
We start from the highest level of evidence available for a question and work down only as far as we need to. In practice that means systematic reviews and meta-analyses first, then randomized controlled trials, then prospective cohort studies, then everything else.
Guidelines from major public health and clinical bodies are treated as a strong input, particularly where they represent a formal consensus process. They are not treated as infallible, and where a guideline lags recent trial evidence the brief says so.
- Human research, in a population reasonably close to the reader the brief is written for.
- Peer-reviewed publication, or a preprint clearly labelled as one.
- Enough methodological detail published to judge how the study was run.
What is excluded, or heavily discounted
- Animal and cell studies as a basis for a human recommendation. They can explain a mechanism; they cannot establish an outcome.
- Studies where the intervention is not something a reader could plausibly do.
- Conference abstracts and press releases with no accompanying paper.
- Research funded by a party with a direct commercial interest in the result, where there is no independent replication.
- Single studies presented as settling a question, regardless of how large they are.
How study quality is judged
Study design is the starting point, not the conclusion. A large, well-run cohort study can be more informative than a small, poorly controlled trial.
Beyond design, we look at whether participants were randomly allocated, whether anyone was blinded and to what, how many people dropped out and whether the analysis accounted for them, how long the study ran relative to the outcome it claims, and whether the outcome measured is the one that matters or a surrogate standing in for it.
- Sample size relative to the size of the effect being claimed.
- Duration relative to how long the outcome takes to appear.
- Whether the primary outcome was declared before the study began.
- Whether the population studied resembles the reader being advised.
- Whether an independent group has reproduced the result.
How conflicting results are handled
Conflicting findings are the normal state of health research, not an anomaly. When trials disagree, we look first for a reason: different doses, different populations, different durations, or different definitions of the outcome. An apparent contradiction often resolves into two studies answering slightly different questions.
Where the disagreement is genuine and unexplained, the rating goes down. A question with strong evidence on both sides is not a strong-evidence question — it is an unresolved one, and the brief says that plainly rather than picking the side that makes a better recommendation.
We do not resolve a conflict by counting studies. One well-designed trial can outweigh several weak ones pointing the other way.
What changes a rating
Ratings are expected to move. A rating that has never changed in an active field usually means nobody has looked at it recently.
- A new systematic review or meta-analysis that changes the overall picture.
- A large, well-conducted trial that does not replicate an earlier result.
- A retraction, a serious methodological criticism, or a failure to reproduce.
- A change in guidance from a major clinical or public health body.
- A reader pointing out something we missed. This happens and it is welcome.
A rating change is logged as a correction, with the reason stated.
Review cadence
Every research brief carries a review date and is scheduled for reassessment at least once every twelve months. Briefs on fast-moving questions are reviewed more often, and any brief can be pulled forward when something significant is published.
A review can end in three ways: the rating is confirmed and the date updated, the rating changes and the change is logged, or the brief is withdrawn because the question has been overtaken. All three are visible.
The limits of this method
This process reduces error; it does not eliminate it. Judging evidence involves interpretation at several points, and two careful people can read the same literature and land one level apart.
We are also constrained by what has been studied. Questions about older adults, women, and people with multiple conditions are frequently under-researched, which means an "insufficient" rating sometimes reflects a gap in the field rather than a genuine absence of effect. Where that is the case, the brief says so.
Last reviewed June 2026 · This page is reviewed alongside the editorial standards
How to read a strength rating.
Endurvity rates the weight of evidence behind a claim on four levels. The rating describes the state of the research, not how enthusiastic anyone is about the finding.
- Strong
Consistent enough to act on. New studies are unlikely to reverse the direction of the finding.
Several randomized trials, or systematic reviews of them, agreeing across different research groups and populations.
- Moderate
Reasonable to act on if the action is low-risk. Expect the size of the effect to move as more work is published.
A smaller body of trials, or larger trials with limitations in size, duration, or how the outcome was measured.
- Emerging
Interesting, not yet a reason to change your routine. Worth watching rather than adopting.
Early trials, observational data, or mechanistic work that has not yet been tested against a real outcome in people.
- Insufficient
Not enough evidence to say either way. A claim at this level should not be driving a decision.
Conflicting results, very small studies, or a question that has not been directly studied in the population being advised.