Methodology

How we rate evidence.

A strength rating is a judgement, and a judgement is only useful if you can see how it was made. This page describes the process, including the parts of it that are unavoidably a matter of interpretation.

Six labels, and why not four

Every research brief carries one of six strength labels, printed in full below. Four of them describe evidence that exists and says something. The other two describe the two ways evidence can fail to answer a question, and they are kept apart on purpose.

"Insufficient" means the research is too thin, too small, or too indirect to support a conclusion. "Not studied adequately" means nobody has run the study — the question has not been asked of people in a controlled way, however good the mechanism looks in a laboratory. Collapsing those into one hedge tells a reader nothing about whether to expect an answer later.

The sixth label, "mixed", covers the case where the research is not thin at all and still does not settle the question: well-conducted studies point in different directions and the difference is not explained by dose, population, or how the outcome was measured. A brief at this level names who disagrees and why rather than picking the side that reads better.

  • Strong — consistent, replicated in humans, unlikely to reverse.
  • Moderate — the effect is real; its size, population, or durability is unsettled.
  • Limited — thin: small, short, single-centre, or surrogate endpoints only.
  • Mixed — good studies disagree, and the disagreement is unresolved.
  • Insufficient — not enough evidence to say either way.
  • Not studied adequately — nobody has properly looked.

This is an Endurvity scale, shared with our sister publications. It is not GRADE, and it does not claim to follow that methodology.

What gets included

We start from the highest level of evidence available for a question and work down only as far as we need to. In practice that means systematic reviews and meta-analyses first, then randomized controlled trials, then prospective cohort studies, then everything else.

Guidelines from major public health and clinical bodies are treated as a strong input, particularly where they represent a formal consensus process. They are not treated as infallible, and where a guideline lags recent trial evidence the brief says so.

  • Human research, in a population reasonably close to the reader the brief is written for.
  • Peer-reviewed publication, or a preprint clearly labelled as one.
  • Enough methodological detail published to judge how the study was run.

What is excluded, or heavily discounted

  • Animal and cell studies as a basis for a human recommendation. They can explain a mechanism; they cannot establish an outcome.
  • Studies where the intervention is not something a reader could plausibly do.
  • Conference abstracts and press releases with no accompanying paper.
  • Research funded by a party with a direct commercial interest in the result, where there is no independent replication.
  • Single studies presented as settling a question, regardless of how large they are.

How study quality is judged

Study design is the starting point, not the conclusion. A large, well-run cohort study can be more informative than a small, poorly controlled trial.

Beyond design, we look at whether participants were randomly allocated, whether anyone was blinded and to what, how many people dropped out and whether the analysis accounted for them, how long the study ran relative to the outcome it claims, and whether the outcome measured is the one that matters or a surrogate standing in for it.

  • Sample size relative to the size of the effect being claimed.
  • Duration relative to how long the outcome takes to appear.
  • Whether the primary outcome was declared before the study began.
  • Whether the population studied resembles the reader being advised.
  • Whether an independent group has reproduced the result.

How conflicting results are handled

Conflicting findings are the normal state of health research, not an anomaly. When trials disagree, we look first for a reason: different doses, different populations, different durations, or different definitions of the outcome. An apparent contradiction often resolves into two studies answering slightly different questions.

Where the disagreement is genuine and unexplained, the brief is rated "mixed". A question with strong evidence on both sides is not a strong-evidence question, and it is not a thin-evidence one either — it is an unresolved one, which is why it gets its own label rather than being filed under a weaker rating that would misdescribe it.

We do not resolve a conflict by counting studies. One well-designed trial can outweigh several weak ones pointing the other way.

What changes a rating

Ratings are expected to move. A rating that has never changed in an active field usually means nobody has looked at it recently.

  • A new systematic review or meta-analysis that changes the overall picture.
  • A large, well-conducted trial that does not replicate an earlier result.
  • A retraction, a serious methodological criticism, or a failure to reproduce.
  • A change in guidance from a major clinical or public health body.
  • A reader pointing out something we missed. This happens and it is welcome.

A rating change is logged as a correction, with the reason stated.

Review cadence

Every research brief carries a review date and is scheduled for reassessment at least once every twelve months. Briefs on fast-moving questions are reviewed more often, and any brief can be pulled forward when something significant is published.

A review can end in three ways: the rating is confirmed and the date updated, the rating changes and the change is logged, or the brief is withdrawn because the question has been overtaken. All three are visible.

The limits of this method

This process reduces error; it does not eliminate it. Judging evidence involves interpretation at several points, and two careful people can read the same literature and land one level apart.

We are also constrained by what has been studied. Questions about older adults, women, and people with multiple conditions are frequently under-researched. That is what the bottom two labels are for: where the gap is in the field rather than in the finding, the brief says "not studied adequately" and names who was left out, instead of implying the answer is no.

Last reviewed June 2026 · This page is reviewed alongside the editorial standards

The six ratings

How to read a strength rating.

Endurvity rates the weight of evidence behind a claim on six levels. The rating describes the state of the research, not how enthusiastic anyone is about the finding. The bottom three levels are kept apart deliberately: studies disagreeing, evidence being too thin to judge, and nobody having run the study are three different situations, and a reader deserves to know which one applies.

  • Strong

    Consistent enough to act on. New studies are unlikely to reverse the direction of the finding.

    Several randomized trials, or systematic reviews of them, replicated in people across different research groups and populations.

    How we write it"reduces", "improves", "increases" — stated with the size of the effect.

  • Moderate

    Reasonable to act on if the action is low-risk. The effect is real; how large it is, in whom, and for how long are still unsettled.

    A smaller body of trials, or larger trials with limitations in size, duration, population, or a surrogate standing in for the outcome that matters.

    How we write it"appears to", "is associated with".

  • Limited

    Interesting, not yet a reason to change your routine. Worth watching rather than adopting.

    Thin evidence: small, short, single-centre trials, or studies reporting only a surrogate marker.

    How we write it"early data suggest", "one small trial found".

  • Mixed

    Good studies disagree, and the disagreement has not been resolved. Someone quoting only one side of it is not describing the evidence.

    Well-conducted studies pointing in different directions, where the difference is not explained by dose, population, or how the outcome was measured.

    How we write it"trials disagree" — followed by who disagrees and why.

  • Insufficient

    Not enough evidence to say either way. A claim at this level should not be driving a decision.

    Very few studies, very small ones, or research that never directly asked the question being answered.

    How we write it"there is not enough evidence to say".

  • Not studied

    Nobody has properly looked. The absence of a finding here is an absence of research, not a finding of no effect.

    No controlled human research on the question, or only animal, laboratory, or mechanistic work standing in for it.

    How we write it"this has not been studied in people".