

This number answers one question: how much can you trust the study's own design? A 1 means real doubt about whether the result reflects what actually happened. A 5 means the design, execution, and reporting hold up under close scrutiny.

There's no single yardstick for every study : a trial and a cohort study fail in different ways, so we apply the tool built for each design:
Randomized controlled trial
Cohort / case-control
Cross-sectional
Systematic review / meta-analysis
Case reports & expert opinion
Jaded Scale
Newcastle-Ottawa Scale
AXIS-based criteria
AMSTAR-2
Floor of the scale
Whether randomization and blinding happened and were described, and whether every participant who dropped out is accounted for.
How the comparison groups were selected, whether confounders like age and comorbidities were controlled for, and how outcomes were measured.
Whether the sampling method was appropriate, the sample size was justified, and confounding was addressed.
A registered protocol, honest handling of differences between pooled studies, and the size of the evidence base behind the conclusion.
No scoring tool applies. Without a comparison group, there's no way to rule out other explanations — so these sit at the bottom regardless of how well the story is told.
Where the quality score asks "can I trust this study," this rating asks "should this study change what I do." A flawless trial answering a low-stakes question can rate low here. A messier study on a life-or-death question can still rate meaningfully — though never higher than its own evidence quality allows.
01
02
03
04
05
Clinical importance
Is this something clinicians and patients genuinely need to know?
Size of the effect
Clinically meaningful, not just statistically significant — a huge relative reduction with a tiny absolute benefit doesn't score high here.
Methodological rigor
Pulled directly from the Quality-of-Evidence score above. It won’t be judged twice.
Generalizability
Do the results apply to patients you actually see, or only to a narrow research population?
Decision impact
If this were accepted as true, would it actually change what happens in the exam room?

A study with serious design flaws can't out-score its own limitations. If the underlying evidence is weak, the rating is capped — no matter how striking the finding looks on the surface.
Scoring only starts once a study clears a few basic checks.
RETRACTIONS
Studies that have been retracted, or are under a formal expression of concern, are removed entirely.
PREPRINTS
Research posted before peer review is marked down for that reason alone. No one has independently checked the work yet.
JOURNAL PRESTIGE
A journal's name or reputation never influences the score. A study is judged on what it actually did, not on where it ran.