Reading a medical research paper well means asking a consistent set of questions: what was the question, who was studied, how were they compared, what was measured, how large and how certain is the effect, and what could have produced this result other than the intervention itself. A reader who works through those questions in order can appraise most papers without specialist statistical training.
This guide sets out that sequence, explains the statistics that appear most often, and identifies the signals that should raise or lower confidence in a finding.
Start With the Question, Not the Conclusion
Identify the research question in structured terms: the population studied, the intervention or exposure, the comparator, and the outcome measured. If any of these is vague in the paper itself, that is a finding. A study of "cardiovascular health" that turns out to have measured a cholesterol fraction over eight weeks in a small volunteer sample is answering a much narrower question than its title suggests.
Then ask whether the question asked is the question you care about. A trial may be well conducted and still be irrelevant to a particular patient population, dose, or setting.
Identify the Study Design
Design determines the strength of the inference.
- Randomized controlled trial. Participants are assigned by chance. This is the design best able to support a causal claim, because randomization balances known and unknown differences between groups.
- Cohort study. Groups defined by exposure are followed over time. Useful for long-term outcomes and harms, but vulnerable to confounding.
- Case-control study. Starts from the outcome and looks back at exposures. Efficient for rare outcomes; susceptible to recall and selection bias.
- Cross-sectional study. Measures exposure and outcome simultaneously. Good for prevalence, poor for causation, because temporal sequence is unknown.
- Case series or case report. Descriptive. Valuable for generating hypotheses and flagging unexpected harms; it has no control group and cannot establish effect.
- Systematic review and meta-analysis. Synthesizes existing studies under a prespecified method. Its reliability is bounded by the quality of the included studies.
Two further checks matter. Was the study preclinical — conducted in cells or animals — rather than in humans? Findings from such work do not establish clinical benefit. And was it a preprint, posted before peer review? Preprints can be sound, but they have not yet been independently appraised.
Examine the Participants
Look at who was enrolled and who was excluded. Age range, sex distribution, disease severity, comorbidities, geography, and prior treatment all shape whether the findings apply beyond the trial. A treatment tested in adults under sixty-five with no kidney impairment tells you little about an eighty-year-old with reduced renal function.
Check attrition as well. If a substantial proportion of participants left before the end, the remaining sample may no longer be comparable across groups, and results should be treated more cautiously. Well-reported papers include a flow diagram showing how many were screened, enrolled, allocated, and analysed.
Check the Comparison
What were participants compared against? Placebo, an active drug, standard care, or nothing at all? A treatment that beats placebo has not been shown to beat the treatment already in use. Equally, if the comparator was given at an unusually low dose or in a non-standard way, an apparent advantage may be an artefact of the comparison rather than a property of the intervention.
Look at the Outcomes
Primary and Secondary Outcomes
The primary outcome is the one the study was designed and powered to detect. Secondary outcomes are exploratory. A study that misses its primary outcome but reports a significant secondary result has not demonstrated benefit; it has generated a hypothesis. Compare the reported primary outcome with the one recorded in the trial registry before enrolment. A discrepancy — outcome switching — is a serious warning sign.
Clinical Versus Surrogate Outcomes
Clinical outcomes describe how patients feel, function, or survive. Surrogate outcomes are markers assumed to predict them: blood pressure, tumour shrinkage, a laboratory value. Surrogates make trials faster and smaller, but improvements in a marker do not always deliver improvements in survival or quality of life, and there are well-documented cases where they diverged.
Composite Outcomes
A composite outcome counts several events together — for example death, hospitalization, and revascularization. If the effect is driven mainly by the least serious component, the headline result overstates the clinical importance. Look for the breakdown by component.
Understanding the Statistics
P-Values
A p-value expresses how compatible the observed data are with the assumption that there is no true effect. A small p-value means such data would be unusual if there were no effect. It does not give the probability that the hypothesis is correct, does not measure effect size, and does not indicate clinical importance. Conventional thresholds such as 0.05 are arbitrary conventions, not scientific boundaries.
Confidence Intervals
A confidence interval gives a range of values compatible with the data, and its width reflects precision. A narrow interval indicates a well-estimated effect; a wide interval indicates substantial uncertainty. An interval that includes the value representing no effect is conventionally read as not statistically significant, but its upper and lower bounds still carry information about what effects remain plausible.
Absolute and Relative Risk
This distinction is the most common source of misleading impressions. A relative risk reduction describes the proportional change between groups; an absolute risk reduction describes the change in actual event rates. Halving a risk sounds substantial, but if the underlying risk is very small, the absolute benefit is correspondingly small. Number needed to treat — how many people must receive the intervention for one to benefit — expresses the same information in a form that is harder to misread.
Multiple Comparisons and Subgroups
Testing many outcomes, time points, or subgroups increases the chance that some comparison will appear significant purely by chance. Subgroup findings that were not prespecified should be treated as hypothesis-generating, especially when they concern a group in which no overall effect was found.
Statistical Versus Clinical Significance
Very large studies can detect differences too small to matter to any patient. Ask whether the size of the effect would change a clinical decision, not merely whether it cleared a statistical threshold.
Look for Bias
- Selection bias. The groups compared differed at the outset in ways related to the outcome.
- Performance bias. Groups received different co-interventions or attention beyond the treatment under study.
- Detection bias. Outcomes were assessed differently depending on group, most likely where assessors were unblinded.
- Attrition bias. Dropout was substantial or differed between groups.
- Reporting bias. Some measured outcomes were not reported, or the reported ones were chosen after seeing the data.
- Publication bias. At the level of a literature rather than a single study: positive results are more likely to be published, so the visible evidence can overstate an effect.
For randomized trials, intention-to-treat analysis — analysing participants in the group to which they were assigned, regardless of what they actually received — preserves the protection randomization provides. Per-protocol analyses, restricted to those who complied, can reintroduce the imbalances randomization removed.
Check Reporting Standards and Registration
Established reporting guidelines make omissions easier to spot: CONSORT for randomized trials, STROBE for observational studies, PRISMA for systematic reviews. Papers following them include the information a reader needs to appraise the work. Prospective registration in a public registry such as ClinicalTrials.gov allows the published report to be checked against the original plan. Systems such as GRADE are used to rate the overall certainty of evidence across studies.
Consider Funding and Conflicts of Interest
Funding source does not invalidate a study, and industry funds much of the highest-quality clinical research. It is, however, relevant context, alongside author affiliations, roles in study design and analysis, and access to the underlying data. Look for statements on who designed the study, who analysed the data, and who had the right to publish. Non-financial interests — intellectual commitment to a prior position — matter too and are less often disclosed.
A Practical Checklist
- What is the population, intervention, comparator, and outcome?
- What is the study design, and does it support the claim being made?
- Were participants randomized, and was allocation concealed?
- Were participants, clinicians, and outcome assessors blinded?
- Does the reported primary outcome match the registered one?
- Is the outcome clinical or surrogate; is it composite?
- What is the absolute effect, not just the relative one?
- How wide is the confidence interval?
- How much attrition occurred, and how was it handled?
- Do the participants resemble the patients this would be applied to?
- Has the finding been replicated independently?
- Who funded the work, and what interests are declared?
Sources
- CONSORT Statement — reporting guideline for randomized trials
- STROBE Statement — reporting guideline for observational studies
- PRISMA Statement — reporting guideline for systematic reviews
- Cochrane — risk of bias assessment and systematic review methodology
- GRADE Working Group — rating certainty of evidence
- American Statistical Association — statement on statistical significance and p-values
- ClinicalTrials.gov — trial registration and results records
- International Committee of Medical Journal Editors — disclosure and authorship standards