Nutrition claims, graded
Twenty claims, the verdict on each, and the rule that produced it
All twenty have their own investigation: what the evidence says, which papers I read and why those, how the certainty grade was reached, and what would change my mind. Every verdict on this page is dated, because a claim that is false this year can be true next year and the reverse happens too.
Two separate judgments sit on every claim here, and they are easy to confuse. The verdict says whether the sentence survives contact with the evidence. The certainty grade says how much evidence there is. A claim can be confidently false on thin evidence, and a claim can be genuinely open on excellent evidence.
Both have a written rule. The verdict matrix sets out what lands a claim in each category. The certainty thresholds set out what counts as imprecise, inconsistent or indirect, in numbers rather than impressions.
This library has a shelf life
Every verdict here is a reading of the evidence as it stood on the date next to it. New trials arrive, reviews get updated, papers get retracted. A claim graded low certainty today is the one most likely to move. Nothing on this page is permanent, and anything that stops being reviewed should stop being trusted.
The library
Filter by topic, verdict or certainty, then open any claim for the full investigation.
How a verdict is chosen
A verdict answers three questions in order, and the answers land it in exactly one category. The questions are the same for every claim, which is what stops the verdict being an opinion with a label on it.
One. Is the observation or mechanism inside the claim real? Two. Does the conclusion follow, at the size and in the population asserted? Three. Does acting on the claim carry a named harm?
| Verdict | Grain real? | Conclusion follows? | What it means |
|---|---|---|---|
| Mostly false | Often yes | No | Something true sits inside it, and the conclusion drawn is wrong about direction, magnitude or cause |
| Partly true | Yes | Only under stated conditions | It holds in a defined group or at a defined dose, and fails outside that. The claim omits the boundary |
| True but reductive | Yes | Yes, as far as it goes | Accurate and incomplete. Something material is left out that changes what a person would do |
| False, and risky | Sometimes | No | The conclusion does not follow, and acting on it carries a harm this entry names |
| Misleading framing | Both poles partly | Neither pole | The claim forces a binary that the evidence does not support in either direction |
| Evidence is mixed | Unclear | Cannot be established | Studies genuinely disagree, or there are too few to support any verdict. This is a statement about the field |
Each graded claim shows which of these applied to it, and why the neighbouring verdict was rejected, in the panel that opens under it.
What high, moderate and low actually mean
These are the four GRADE levels, used as GRADE defines them. The starting point depends on design, and five domains can pull it down. What follows is the operational threshold I use for each, so the judgment can be checked rather than taken on trust.
| Level | Interpretation | What it takes |
|---|---|---|
| High | Further research is very unlikely to change the estimate | Randomized evidence with no domain triggered |
| Moderate | Further research is likely to matter and may change the estimate | Randomized evidence with one domain triggered |
| Low | Further research is very likely to change the estimate | Randomized evidence with two domains, or observational evidence with none |
| Very low | Any estimate is very uncertain | Three or more domains, or observational evidence with one |
| Not graded | No judgment offered | I have not read the literature under this claim yet |
The five domains, with the threshold each one uses
| Domain | Threshold I apply | Why that threshold |
|---|---|---|
| Imprecision | Total participants across the pooled studies fall below roughly n = 400, or the 95 percent interval includes both a meaningful benefit and a meaningful harm | The optimal information size conventionally used in GRADE. Below it, a confidence interval is wide for reasons of sample size alone |
| Inconsistency | I² > 50% with no explanation, or point estimates falling on opposite sides of no effect | The conventional cut for substantial heterogeneity. Above it, pooling hides a disagreement rather than resolving one |
| Indirectness | The outcome measured is a surrogate rather than the thing people care about, or the population differs from the one the claim addresses | Lipids are not heart attacks. A trial in people over 45 with existing disease does not describe a healthy thirty-year-old |
| Risk of bias | More than a quarter of the pooled weight comes from studies the review rated at high risk, or from unblinded studies where blinding was possible | At that share, the biased studies are driving the estimate rather than sitting at its edge |
| Publication bias | Funnel asymmetry reported by the review, or an evidence base made entirely of small positive studies with industry funding and no negative counterparts | The pattern that appears when unflattering results were run and never published |
What can push observational evidence back up
Three things, each worth one level, and only for observational bodies of evidence: a large effect at RR > 2 or RR < 0.5 with no plausible confounder that size, a dose-response gradient, and the case where every plausible unmeasured confounder would have pushed the result toward no effect rather than away from it.
How the evidence base is counted
Each graded claim states the number of studies and the number of participants that a named review was able to pool for that specific question. Nineteen studies found and two poolable for muscle is reported as two, because two is what the estimate rests on. Counts from different reviews are not compared with each other, which is why there is no bar putting them on a shared scale.
How the papers get chosen
Nobody can read every paper on carbohydrate, and anybody claiming to has not looked at how many there are. So the honest thing is to say what gets read and in what order, and to be clear that this is a reading of the best available synthesis rather than a systematic review of my own.
The order of preference
| Look for | Why it comes first |
|---|---|
| 1. A recent systematic review on this exact question | Somebody has already done the searching, stated their method, and reported what they could not pool. That last part is usually the most useful sentence in the paper |
| 2. The largest randomized trial measuring the outcome that matters | Where no review exists, one well-powered trial with a hard endpoint beats several small ones with a surrogate |
| 3. Controlled feeding or metabolic ward work | For questions about what a diet does at matched intake, this is the only design that removes self-report |
| 4. The best observational evidence, labelled as such | Used where the question cannot be randomized, and the entry then starts at low certainty by design |
| 5. A guideline or consensus statement | A reading of the evidence rather than evidence. Useful for what a body concluded, and never cited as the finding itself |
What gets excluded, and why
Every PubMed search here filters out letters, comments, editorials, news items and errata, because those get returned alongside the paper they discuss and are easy to cite by accident. Beyond that filter, three things are excluded deliberately.
Single-arm studies used to support a magnitude. A study with no control group can tell you what happened and not what would have happened anyway. Several of the percentages in circulation about GLP-1 and muscle come from exactly this, which is why those numbers are not on this page.
Mechanistic work standing in for an outcome. Animal and cell work shows a pathway is possible. It appears here labelled as mechanism and never as the reason to do anything.
Papers I have not read past the abstract. Where only the abstract was available, the entry says so. Where a full text could not be reached, the claim rests on something else or stays ungraded.
What this is not
This is not a de novo systematic review. I have not searched every database from inception for every claim, screened the results in duplicate, or extracted data independently. What I have done is find the best existing synthesis, read it, check it against its own record on PubMed, and report what it says and what it says it cannot say.
Each graded claim names its papers and says why those and not others, inside the panel. Where the honest answer is that only one relevant study exists, that is what the entry says, and the certainty grade reflects it.
The review policy
Every claim carries two dates: when it was last reviewed, and when it is next due. The default cycle is twelve months. Four things pull a review forward: a new systematic review on the question, a trial large enough to change a pooled estimate, a retraction or expression of concern on a source, and a guideline change from a body that reads the same literature.
When a verdict changes, the old verdict and the date it changed stay visible rather than being quietly replaced. A library that revises itself without saying so is asking for the same trust it tells you not to give.
Where this method is weakest
GRADE was built for panels, and this is one person. A panel argues about whether a domain was triggered, and that argument is part of the method. Here the argument is missing, so the thresholds above are doing the work a second reader would normally do.
The domain judgments are also the softest part. Whether a population counts as indirect enough to cost a level is a judgment with a number attached, not a measurement. Where you would have graded differently, the arithmetic is on the page so you can say exactly where.
