Supplement Studies  /  evidence, scored
Method v0.1 · last revised July 2026

How we
score

The method is the product. Everything else is presentation. If you disagree with a score, this page tells you exactly where to argue.

The principle


We score claims, never supplements. A supplement has no single truth value. Creatine is excellent for strength in trained young men and unremarkable for cognition in rested adults. Collapsing those into one rating destroys the only information that matters.

A scorable claim names an intervention, a population, an outcome, and where possible a dose. If any of the four is missing, the claim is not yet scorable.

“Magnesium helps sleep” is not scorable. “Magnesium glycinate at 250 mg elemental, taken nightly by adults reporting poor sleep, reduces insomnia severity” is.

Four dimensions


Each is scored 0 to 10 against the evidence supporting the claim as stated. For a claim that good trials have refuted, the supporting evidence is thin and the score is low — regardless of how well those trials were run.

Quantity

× 0.25

How many participants across how many independent trials? Under 200 people or a single trial scores low. Multiple well-powered trials plus independent meta-analyses score high.

Quality

× 0.35

Randomised? Blinded? Placebo-controlled? Preregistered? Are attrition, funding source and conflicts declared? Industry-funded trials are recorded and flagged, never silently excluded.

Consistency

× 0.25

Do independent groups find the same direction of effect? Unexplained heterogeneity caps this at 5. Direct disagreement between competent syntheses caps it lower still.

Directness

× 0.15

How close is the measured outcome to the claim as people hear it? A biomarker is not the outcome. Serum cortisol is not feeling calm. Serum testosterone is not muscle. A rodent is not a person.

Score = (Quantity × 0.25) + (Quality × 0.35)
        + (Consistency × 0.25) + (Directness × 0.15)

Quality carries the heaviest weight deliberately. Twenty poorly conducted trials do not outrank two good ones, and a literature can be large and still be weak — which is precisely what a formal appraisal found for curcumin, where 19 of 25 meta-analyses were rated very low quality.

Bands


8.0 – 10.0Strong

Unlikely to reverse.

6.0 – 7.9Moderate

Direction consistent, magnitude uncertain.

4.0 – 5.9Emerging

Real signal, thin base.

2.0 – 3.9Weak

Preliminary, or contradicted as often as supported.

0.0 – 1.9Against

Good trials looked and found nothing, or found the opposite.

Every score also carries an uncertainty band — the hatched range beside the bar on each page. It is the plausible spread if the next well-powered trial reports. Wide means the score is likely to move. A claim whose band spans two tiers is reported as the lower tier.

Where two competent syntheses reach opposite conclusions, we mark the claim Contested and widen the band rather than picking a winner. Vitamin E and mortality is the clearest current example.

Non-negotiables


  1. Every citation is opened and read by a human before publication. A reference produced by a language model and not verified is never published.
  2. Every score records who scored it, when, and against which version of this method.
  3. Changes to the method trigger a re-score of affected claims, with the previous score kept visible.
  4. No payment of any kind may influence a score. No exceptions, ever.
  5. Null and negative findings are given the same prominence as positive ones.
  6. Where a claim cannot be scored as written — because it collapses several populations with different answers — it is split rather than averaged.