Why it matters

Developed using GRADE is a claim about method, and it is checkable. GRADE publishes seven suggested criteria that “should be met when saying that the GRADE approach was used” — a checklist, not a licensing condition, and only criterion 7 is wholly aspirational. Criterion 4 is often misread as aspirational because it contains the word “ideally”, but it states a requirement and then an explicit floor: evidence tables should be used as the basis for judgements”, Ideally, full evidence profiles…”, At a minimum, the evidence that was assessed and the methods that were used… should be clearly described.” Only the evidence-profile clause is an aspiration; the surrounding requirement binds. It exists because “modified” variants proliferate — the Working Group “discourage[s] the use of ‘modified’ GRADE approaches that differ substantially from the approach described.” (Schünemann et al., n.d.)

This turns a vague suspicion (this guideline seems under-argued) into a specific, auditable finding — which is the standard a process-defect claim has to meet before it can be used to discount a recommendation.

Tests / indicators — the seven criteria

  1. Definition of quality used consistently with GRADE’s (the guideline-context or review-context definition — they differ).
  2. All eight criteria explicitly considered: risk of bias · directness · consistency · precision · publication bias · magnitude of effect · dose-response gradient · residual plausible confounding. Different terminology is acceptable; skipping a criterion is not.
  3. Certainty rated per outcome, not per study or per document, in four categories (or three, if justified, collapsing low and very low), with GRADE-consistent interpretation.
  4. Evidence summaries — evidence tables or detailed narrative summaries that transparently describe the judgments behind the rating; ideally full evidence profiles based on systematic reviews. At minimum, what evidence was assessed and how it was identified and appraised. Reasons for downgrading and upgrading described transparently.
  5. All four strength criteria explicitly considered: balance of consequences · quality of evidence · values and preferences of those affected · resource use — with the general approach reported (whether and how costs entered, whose values were assumed).
  6. Two strength categories (strong / weak, or conditional for weak), with interpretation and implications preserved even if wording differs.
  7. Judgments about strength transparently reported. (Schünemann et al., n.d.)

Red flags

  • A single certainty grade for the whole document or for a recommendation, rather than per outcome (fails 3). Whether this also implies 2 was skipped is a reasonable suspicion, not a frequency the handbook states.
  • Downgrade or upgrade decisions asserted without stated reasons (fails 4) — the footnotes are where the judgment lives, so their absence is not a presentational shortfall but a missing method.
  • Strength stated with no indication of whose values were assumed (fails 5). Values-variability is what makes a recommendation weak despite high-quality evidence, so omitting it removes the only way to interpret the rating. (inferred from Schünemann et al., n.d.)
  • Upgrade factors never mentioned at all — consistent with observational evidence having been floored at low without the three exits being considered, though the handbook makes no claim about how often that happens (Upgrading Observational Evidence).

Green flags

  • An evidence profile with a row per outcome, including empty rows for outcomes with no evidence (Rating Outcome Importance).
  • Footnotes recording close calls, including factors considered and declined.
  • An explicit statement of perspective and of how resource use was treated.

Decision relevance

  • A failed criterion is a specific, citable findingcertainty was not rated per outcome is checkable and arguable; this guideline is unrigorous is neither.
  • Failure is bounded, not global. A document failing criterion 5 has an under-argued strength rating; that says nothing about whether its effect estimates are right. Scope the finding to what the failed criterion covers.
  • Meeting all seven does not make a recommendation correct — it means the judgments are visible enough to be argued with. That is what GRADE claims for itself (Rating Certainty of Evidence), and no more.

Limits

  • The criteria test conformance and transparency, not correctness. A document can satisfy all seven and still reach a conclusion the evidence does not support; conversely, a non-conforming document can be right.
  • GRADE is judging its own use, so the criteria embed GRADE’s premises — they cannot adjudicate a dispute with a body that rejects the framework rather than misapplying it.
  • Source currency: §8 is flagged in-source as rewritten in the 2024 GRADE Book (as requirements for claiming the use of GRADE).

Run against a real guideline for the first time [2026-07-28, WHO SFA 2023 Annexes 6-7]

This page held seven criteria and no worked application. WHO’s SFA guideline is now readable against them, because Annex 6 (evidence profiles) and Annex 7 (evidence-to-recommendations) have both been extracted.

#CriterionVerdictEvidence
1Definition of quality consistent with GRADE’sPASSfour categories, GRADE’s own ㊉ notation
2All eight criteria explicitly consideredPASS — 7 of 8 evidenced [revised 2026-07-28]see below
3Certainty rated per outcomePASS, stronglyseparate ratings for all-cause mortality, CVD mortality, CVD events, CHD mortality, CHD events, stroke, T2D, LDL
4Evidence summaries with downgrade reasons transparentPASSper-cell judgements (Not serious / Serious / Very serious) each carrying a numbered footnote giving the reason
5All four strength criteria consideredPASSAnnex 7 covers balance of effects, overall certainty, values/variability, and resources + cost-effectiveness
6Two strength categoriesPASSstrong / conditional
7Strength judgements transparently reportedPASS on the judgements — see the gap belowevery EtD domain carries a ticked box plus reasoning

Criterion 2 in detail — this cell was revised twice, both times because my search string was wrong rather than the document.

Evidenced as considered: risk of bias, inconsistency, indirectness, imprecision, publication bias (named rows in every profile), dose-response (57 occurrences across 7 files; it appears as an assessment-column entry, and was used to upgrade — «A dose-response relationship was observed when replacing 5% of energy intake as SFA with the equivalent (i.e. isocaloric) amount of monounsaturated fatty acids. Upgraded once»), and residual confounding (8 occurrences). Thin: magnitude of effect — 1 occurrence. [searched with bin/srcgrep.py, which folds dash/ligature/line-break variance: "dose-response", "magnitude of effect", "plausible confounding", "residual confounding", "large effect" across all 7 files of the source] (World Health Organization, 2023)

Two false absences, from one cause, recorded because the pattern matters more than the cell. (1) dose-response searched with a hyphen returned zero; WHO writes it with an en-dash, and there are 57 occurrences. (2) plausible confounding — GRADE’s phrasing — returned zero; WHO writes «residual confounding», 8 occurrences. Both would have failed WHO on a criterion it meets, and each was caught only by re-searching with a different string. The general lesson: an absence claim against a PDF-extracted source is unreliable from grep. This corpus breaks words across lines (rec- ommended, carbohy- drate), varies dash forms, and carries ligatures — none of which a dash-class regex fixes. bin/srcgrep.py normalizes all of it and prints the denominator; it exists because of these two misses.

The result is conformance — and that is the reportable finding

WHO SFA 2023 substantially conforms to GRADE’s own checklist. The telos is explicit that “convention held here” is a reportable, valued finding, and this page’s own framing says the diagnostic exists to turn vague suspicion into an auditable finding — which cuts both ways, and here it cuts toward the guideline.

The one real gap is in the CHECKLIST, not in WHO

Annex 7 publishes every domain judgement and the reasoning inside each domain. It publishes no rule for how the domains were combined. WHO issued a strong recommendation on SFA with:

EtD domainWHO’s judgement
Balance of effectsFavours interventions
Overall desirable effectsModerate
Health inequityProbably reduced
Acceptability (SFA)Varies
FeasibilityProbably yes
Cost-effectivenessDon’t know

(World Health Organization, 2023)

Two unfavourable-or-unknown non-health domains did not prevent a strong recommendation. WHO does state a per-recommendation rationale in its Evidence-to-recommendations narrative — Recommendation 1 «was assessed as strong because evidence of moderate certainty… suggested reduced risk of CVDs with lower SFA intake. No undesirable effects or other mitigating factors were identified that would argue against a lower SFA intake» (World Health Organization, 2023). But that rationale argues from health evidence plus the absence of a countervailing factor — it is a qualitative why, not a published rule for how the Acceptability Varies and Cost-effectiveness Don’t know judgements were weighed against the health benefit. So the corrected claim is narrow: WHO states that the balance favoured the health evidence, but publishes no formal weight or combination rule for the non-health domains — and the annex ends on feasibility with no concluding synthesis.

Here is the part that matters, and it relocates a standing question. Criterion 7 asks that strength judgements be transparently reported. It does not ask for a combination rule. So WHO’s silence about weighting is not a conformance failure — GRADE does not require it. A process-defect charge on this ground would fail this page’s own bar, which demands the defect be documented against the standard the body is held to. (Schünemann et al., n.d.; inferred from World Health Organization, 2023)

Bearing on the standing [PRIOR — test me] about undisclosed weighting, recorded and NOT scored (adjudication is out of ingest scope). The prior’s surviving form holds that considerations are disclosed and weights are not. This case supports the description and reassigns the cause: the absent weight is what the instrument asks for, not what a body chose to withhold. Handle: the weighting [PRIOR] in CLAUDE.md.

The prior question this page skips: what if the body does NOT use GRADE at all [2026-07-31, USPSTF]

This diagnostic checks whether a body claiming GRADE actually conformed. USPSTF is the case the page did not hold: a major guideline body that runs a fully specified appraisal system and never claims GRADE. So the seven-criteria conformance check does not apply to it — the right audit for a USPSTF recommendation is against its own manual (per-outcome EPC strength-of-evidence; six critical- appraisal questions; certainty of net benefit; the A/B/C/D/I grid), not against GRADE’s checklist. -> Net Benefit and the USPSTF Recommendation Grid, GRADE vs USPSTF - Two Appraisal Systems

The trap this page must add: a “GRADE” label that is not the GRADE system. USPSTF’s own adequacy/certainty tool (Appendix XI) labels its recommendation-grade cell «GRADE (A, B, C, D, or I)» — USPSTF calls its letter-grade output “GRADE”. A reader auditing a USPSTF document who sees “GRADE” and reaches for this page’s seven criteria is checking the wrong thing: the word names the A-through-I output, not Grading of Recommendations Assessment, Development and Evaluation. (US Preventive Services Task Force, 2022)

So the diagnostic gains a precondition: before running the seven criteria, confirm the document means the GRADE method. A body can (a) use GRADE and claim it (run the check), (b) use a different system and say so (audit against that system — USPSTF), or (c) use the token “GRADE” for something else entirely (a naming collision, not a method claim). Only (a) is this page’s job. (inferred from US Preventive Services Task Force, 2022)

This checks GRADE conformance; the SR PROCESS has its own standard now [2026-07-31, IOM]

This page audits whether a body that claims GRADE conformed — a check on the certainty/strength grading step. It says nothing about whether the underlying systematic review was trustworthy: the search, screening, extraction, and synthesis that produced the body of evidence GRADE then grades. That is a different object with its own admissible institutional bar — What a Trustworthy Systematic Review Requires (IOM 2011, 21 standards / 82 elements).

The two audits compose and do not overlap. A recommendation can conform to GRADE (this page’s seven criteria) while resting on an SR that failed the IOM search or dual-screening standards — a GRADE-conformant grade computed over an untrustworthy evidence base. Conversely a rigorous SR can be graded by a non-GRADE system (USPSTF) and pass the IOM bar while this page’s check does not apply. Run both: IOM for the review process, this page for the GRADE claim. Same shared lineage (IOM built its Chapter 4 body-of-evidence standard from GRADE), so their agreement is F, not independent corroboration. (inferred from Institute of Medicine, 2011)

References

Institute of Medicine. (2011). Finding What Works in Health Care: Standards for Systematic Reviews (J. Eden, L. Levit, A. Berg, & S. Morton, Eds.). National Academies Press. https://doi.org/10.17226/13059
Schünemann, H., Brożek, J., Guyatt, G., & Oxman, A. (n.d.). GRADE Handbook: for grading quality of evidence and strength of recommendations. https://gradepro.org/handbook/
US Preventive Services Task Force. (2022). U.S. Preventive Services Task Force Procedure Manual. https://www.uspreventiveservicestaskforce.org/uspstf/sites/default/files/inline-files/procedure-manual-2022.pdf
World Health Organization. (2023). Saturated fatty acid and trans-fatty acid intake for adults and children: WHO guideline. https://www.who.int/publications/i/item/9789240073630