Strata Academy
GRADE Checklist: How to Rate Certainty of Evidence
Rate certainty from high to very low — five downgrade factors, Summary of Findings tables, and when GRADE applies after ROB 2
GRADE — five reasons to downgrade
- Risk of bias — From RoB 2 / ROBINS-I domain judgements
- Inconsistency — Unexplained heterogeneity (I² + clinical sense)
- Indirectness — PICO differs from the question of interest
- Imprecision — Wide CI or Optimal Information Size not met
- Publication bias — Funnel asymmetry + selective availability
Work the downgrade interactive below, then try a Summary of Findings–style appraisal on a real PDF.
Try GRADE downgrade interactive Appraise with GRADE context
Related frameworks
RoB 2 · PRISMA 2020 · AMSTAR 2 · Forest plots
Quick answer
GRADE rates certainty of evidence from high to very low for a body of evidence — start high for RCTs (low for observational), then downgrade for risk of bias, inconsistency, indirectness, imprecision, or publication bias.
- RCT bodies start at high certainty; observational at low — then downgrade or upgrade.
- Each outcome in a review gets its own GRADE rating in a Summary of Findings table.
- Statistical significance alone does not mean high certainty.
- Pair GRADE with ROB 2 / ROBINS-I on included studies first.
- SoF footnotes must name downgrade reasons — not icons alone.
1. What is GRADE?
GRADE (Grading of Recommendations Assessment, Development and Evaluation) is a system for rating certainty in a body of evidence and for framing clinical recommendations. In student appraisal, GRADE most often appears in systematic reviews and clinical guidelines as certainty ratings: high, moderate, low, or very low.
GRADE does not replace study-level risk-of-bias tools. You typically apply ROB 2 or ROBINS-I to individual trials first, then GRADE considers the whole body of evidence for each outcome.
Certainty answers: how confident are we that the true effect lies close to the estimated effect? It is distinct from effect size (how big) and from recommendation strength (how strong the guideline suggestion is).
NICE, SIGN, and WHO guidelines use GRADE or GRADE-compatible frameworks. When reading a NICE technology appraisal, locate the committee's certainty discussion before the recommendation — it explains why strong vs conditional wording was chosen.
The GRADEpro Guideline Development Tool (GDT) is the standard software for producing SoF tables in Cochrane and guideline work. Student projects may use simplified Word tables — but column structure should mirror official SoF format.
Surrogate outcomes: when trials report biomarkers instead of patient-important endpoints, GRADE indirectness downgrade often applies even if the meta-analysis p-value is highly significant. Name the surrogate explicitly in appraisal.
Tip: Certainty is about how confident we are that the true effect lies close to the estimate – not about how important the outcome is clinically.
2. When to use GRADE
Use GRADE when synthesising evidence across studies for one or more outcomes – especially in systematic reviews, network meta-analyses, and clinical practice guideline development.
GRADE is not usually applied to a single small RCT in isolation for coursework unless you are explicitly practising the framework on one study as a teaching exercise.
Diagnostic test reviews may use GRADE-PRO or field-specific adaptations. Qualitative evidence has separate GRADE-CERQual approaches – do not apply intervention GRADE blindly.
When reading guidelines, locate the SoF table before the recommendation strength – certainty and effect magnitude underpin 'strong' vs 'conditional' recommendations.
- Systematic review with pooled or narrative synthesis → GRADE per outcome
- Clinical practice guideline → GRADE underpins recommendation strength
- Single RCT critical appraisal → ROB 2 first; GRADE optional unless part of a review
- Diagnostic review → GRADE-PRO or field-specific adaptations
| Context | Use GRADE? | Notes |
|---|---|---|
| Systematic review with synthesis | Yes — per outcome | SoF table expected |
| Clinical practice guideline | Yes | Links to recommendation strength |
| Single RCT journal club | Usually no | Use ROB 2 + CONSORT instead |
| Qualitative synthesis | No — use CERQual | Different framework |
| Diagnostic test review | GRADE-PRO or adapted | Check field guidance |
Work the downgrade interactive below, then try a Summary of Findings–style appraisal on a real PDF. Try GRADE downgrade interactive · Appraise with GRADE context
3. Starting certainty by study design
GRADE starts from a default certainty level depending on the dominant study design in the evidence body. Randomised trials start at high certainty; observational studies start at low certainty – then factors upgrade or downgrade from there.
This default surprises students: a single RCT does not automatically mean 'high certainty' in GRADE terms if risk of bias, imprecision, or other factors downgrade the rating.
Mixed bodies of RCTs and observational studies require explicit judgement about which design dominates for each outcome. Cochrane reviews often present separate certainty for RCT-only vs observational supplement evidence.
Very low certainty is common for rare outcomes with few events — even when several RCTs exist — because imprecision downgrades dominate.
- RCTs → start high, then downgrade as needed
- Observational intervention studies → start low, may upgrade for large effects or dose–response
- Case series and mechanistic studies → rarely form high-certainty bodies alone
- Mixed evidence → document which design drives the rating per outcome
4. Five reasons to downgrade
Reviewers downgrade certainty when methodological weaknesses or inconsistency reduce confidence that the true effect lies close to the estimate.
Each downgrade typically reduces certainty by one level (e.g. high to moderate), with a maximum of three downgrades from the starting level for most bodies of evidence.
Downgrades should be explicit in the SoF table footnotes – 'downgraded for serious risk of bias' is not sufficient without naming which studies or domains drove the decision.
Multiple downgrade factors often co-occur: imprecision and publication bias frequently accompany small review bodies with selective industry sponsorship.
Optimal Information Size (OIS): when event counts are below OIS, GRADE imprecision downgrade applies even if the confidence interval excludes null — students often miss this when fixated on p-values.
- Risk of bias – serious limitations in included studies (from ROB 2 / ROBINS-I judgements).
- Inconsistency – unexplained heterogeneity in direction or size of effects (see I², prediction intervals).
- Indirectness – PICO differs from the question of interest (population, intervention, comparator, outcome).
- Imprecision – wide confidence intervals or few events; Optimal Information Size not met.
- Publication bias – asymmetry in funnel plots, grey literature suggesting selective availability.
Note: Downgrades should be justified in plain language in the review – not only as icons in a Summary of Findings table.
5. Upgrading observational evidence
Observational bodies of evidence can be upgraded in specific circumstances – large magnitude of effect, dose–response gradient, or when plausible confounding would reduce rather than inflate the apparent effect.
Upgrades are less common than downgrades in practice. Do not upgrade simply because the p-value is very small.
The 'confounding would reduce the effect' upgrade applies when unmeasured confounding is likely to bias toward the null — for example, sicker patients receiving less aggressive treatment in an observational study showing benefit.
Large magnitude (e.g. RR >2 or <0.5 with narrow CI) may support one upgrade — but only when bias from other sources is not serious.
| Upgrade factor | When it applies | Caution |
|---|---|---|
| Large effect | Very large RR/HR unlikely from bias alone | Not if serious ROB remains |
| Dose–response | Clear gradient across exposure levels | Confounding by indication may mimic dose effect |
| Confounding would reduce effect | Bias direction plausible toward null | Requires explicit argument |
6. Summary of Findings (SoF) tables
SoF tables present, per outcome: the number of studies and participants, relative and absolute effects, and the certainty rating with a short rationale for downgrades.
When appraising a review, check that every patient-important outcome has a certainty rating – not only the outcome that reached statistical significance.
Absolute risk differences (or absolute event rates per 1,000) support shared decision-making. Relative risk alone can exaggerate importance when baseline risk is low.
Footnotes should name downgrade domains (bias, inconsistency, imprecision, indirectness, publication bias) and upgrade factors if observational evidence was upgraded.
NICE public summaries sometimes simplify SoF tables — read the full technical appendix for downgrade footnotes when preparing guideline-based coursework.
- Absolute risk differences matter for shared decision-making.
- If authors pool without assessing heterogeneity, GRADE inconsistency downgrade may apply.
- Compare SoF to full risk-of-bias tables – ratings should be coherent.
- Check whether all pre-specified outcomes appear — not only favourable ones.
7. Worked example – reading a SoF table
Open any recent Cochrane review with a Summary of Findings table. Trace how ROB 2 judgements, imprecision, and inconsistency appear in footnotes for each outcome.
Compare two outcomes in the same review — one may be high certainty and another very low despite sharing included studies. Per-outcome GRADE is a feature, not an inconsistency.
8. GRADE and recommendation strength
GRADE certainty feeds into guideline recommendation strength — but they are not the same rating. Certainty is about evidence; strength is about whether benefits outweigh harms and whether the recommendation should apply broadly.
Strong recommendations: most informed patients would want the intervention; effects outweigh harms; high or moderate certainty usually required, though exceptions exist when large effects meet critical outcomes.
Conditional (weak) recommendations: trade-offs are closer; values and preferences differ; lower certainty or imprecision often applies. NICE uses 'offer' vs 'consider' language that maps loosely to this framework.
When reading SIGN or NICE guidance, locate both the certainty rating and the recommendation wording. A conditional recommendation with low certainty is not a contradiction — it reflects appropriate uncertainty for shared decision-making.
Students writing mock guideline sections should pair SoF tables with explicit certainty and recommendation strength, citing GRADE handbook definitions rather than inventing 'Grade A/B/C' labels from undergraduate teaching.
Network meta-analysis certainty uses GRADE with additional inconsistency considerations for transitivity — do not copy pairwise GRADE ratings from component comparisons without NMA-specific judgement.
- High certainty does not automatically mean strong recommendation
- Conditional recommendations are appropriate when patient values differ
- Cost and equity may affect recommendation strength beyond GRADE certainty
- EtD (evidence-to-decision) frameworks extend GRADE in WHO and NICE processes
9. GRADE in the synthesis workflow
A sensible order: (1) define outcomes a priori in the protocol, (2) assess risk of bias per study with design-appropriate tools, (3) synthesise (meta-analysis or narrative), (4) assess inconsistency and imprecision, (5) assign GRADE certainty per outcome, (6) interpret clinical implications.
Risk-of-bias judgements feed the GRADE 'risk of bias' downgrade domain – traffic-light plots should not be decorative.
Imprecision assessment uses optimal information size and confidence interval width, not only p-values. Few events in rare outcomes often force serious imprecision downgrades.
StrataResearch aligns meta-analysis manuscripts with GRADE-related statistics feedback and AMSTAR 2 / PRISMA domains so you can see whether the published review followed this sequence.
- Pre-specify outcomes in PROSPERO protocol.
- Complete ROB 2 / ROBINS-I per included study.
- Synthesise with heterogeneity investigation.
- Assess imprecision (events, CI width, OIS).
- Assign GRADE per outcome with footnoted rationale.
11. Journal club checklist (GRADE)
Locate the Summary of Findings table before discussing clinical implications. Read certainty ratings per outcome — not only the abstract's headline outcome.
For each outcome, trace downgrade footnotes to ROB plots and forest plots. Ask: 'Do I agree with the authors' inconsistency and imprecision judgements?'
Convert relative effects to absolute for one outcome using a plausible baseline risk from UK data — NICE often uses NHS baseline rates in public summaries.
Distinguish recommendation strength from certainty when the paper is a guideline — students often conflate 'conditional recommendation' with 'bad evidence'.
Very low certainty does not always mean 'do not treat' — it means the true effect may differ materially from the estimate. Shared decision-making language belongs in your appraisal conclusion.
- SoF table before clinical discussion
- Per-outcome certainty not one global label
- Footnotes linked to ROB and heterogeneity
- Absolute effects for shared decision-making
- Certainty vs recommendation strength separated
12. Common GRADE mistakes
GRADE errors in student work often reflect skipping earlier appraisal steps. These mistakes are predictable and avoidable.
- Confusing statistical significance with high certainty.
- Applying GRADE without prior risk-of-bias assessment of included studies.
- Ignoring indirectness when trials use surrogate outcomes.
- Treating GRADE as optional decoration in the discussion rather than tied to each outcome.
- Upgrading observational evidence because the association was 'significant'.
- Single certainty rating for all outcomes when only one was analysed rigorously.
13. StrataResearch and GRADE
Meta-analysis and systematic review uploads receive GRADE-aligned certainty commentary alongside AMSTAR 2 and heterogeneity feedback.
Compare automated GRADE prompts to your manual SoF reading — useful when preparing dissertation discussion sections on evidence certainty.
Pair StrataResearch output with official GRADE handbook tables when learning downgrade rules for coursework.
SoF-style certainty commentary on review uploads helps you draft footnotes for your own dissertation SoF table — mirror the phrasing style of Cochrane downgrade statements rather than inventing vague 'quality was poor' language.
When imprecision and risk of bias both apply, GRADE allows multiple downgrades — footnotes should list both explicitly rather than collapsing to a single generic 'low quality evidence' phrase.
SIGN guideline summaries use plain-language certainty labels — trace each back to the technical GRADE SoF table in the full guideline PDF when preparing GP placement presentations.
Build one SoF table from scratch in Word before your dissertation — the column structure makes sense only after you populate a real outcome row yourself.
Practise footnotes early — they carry most of the GRADE reasoning.
Frequently asked questions
How do you rate certainty of evidence with GRADE?
Start from the study-design default (RCTs high; observational low), then downgrade for risk of bias, inconsistency, indirectness, imprecision, or publication bias. Rate each important outcome separately in a Summary of Findings table with footnotes naming each downgrade.
What are the five reasons to downgrade certainty in GRADE?
Risk of bias in included studies; inconsistency (unexplained heterogeneity); indirectness of evidence; imprecision (wide CIs or few events); and publication bias. Each serious concern typically lowers certainty by one level.
What is a GRADE Summary of Findings table?
A Summary of Findings (SoF) table presents each important outcome with the effect estimate, number of participants, certainty rating (⊕⊕⊕⊕ to ⊕○○○), and footnotes explaining downgrades. It is the standard way Cochrane and guidelines communicate evidence certainty to clinicians.
When should I use GRADE?
Use GRADE when synthesising evidence across studies for one or more outcomes — especially in systematic reviews and clinical guidelines. Appraise included studies with ROB 2 or ROBINS-I first; GRADE is not a substitute for study-level risk of bias.
Does a significant p-value mean high GRADE certainty?
No. Statistical significance is about precision of the estimate relative to the null; GRADE certainty reflects confidence that the true effect is close to the estimate after considering bias, inconsistency, indirectness, imprecision, and publication bias.
Try GRADE downgrade interactive
Work the downgrade interactive below, then try a Summary of Findings–style appraisal on a real PDF.
Interactive walkthroughs and quizzes load when JavaScript is enabled — the checklist and tables above are fully readable without it.