← Back
Interpretation of Medical Literature · Class of 2028

Medical Literature Cram Sheet

Sessions 1–4 and 7–10 condensed: the bias taxonomy, the evidence hierarchy and USPSTF grades, every study design, rates, data and variation, diagnostic test interpretation, risk and ratios, prognosis, prevention and screening, trials, statistics and causation, and systematic reviews.

How to use this: this is a condensed, night-before-the-exam reference, not a replacement for the full study guide — it assumes you've already learned the material and just need the highest-yield facts at a glance. If a term feels unfamiliar, go back to the full guide for the explanation.

Bias, Chance & Validity

TermWhat you need to know
BiasSystematic error producing conclusions that differ from the truth. Can enter at conception/design, data collection, analysis or interpretation. Potential for bias does not mean bias is present — but it is often detectable when reading the article.
ChanceRandom variation. Equally likely above or below the truth. Reduced by good design, estimated statistically, never eliminated. May coexist with bias.
Telling them apartReadings that all fall on one side of the true value = bias. Random variation would scatter both ways.
Selection (sampling) biasCompared groups differ in more ways affecting the outcome than the study's focus. Consider at study design. Example: surgical arm healthy enough to tolerate surgery, medical arm sicker.
Measurement biasThe method of measurement gives incorrect results — blood pressure taken without a standardized procedure every time.
ConfoundingThe covariate's effect cannot be separated from the variable studied. Associated with both exposure and outcome; need not cause the disease. Consider at data analysis. May result from selection bias or chance.
Procedural biasInterview/questionnaire results affected by method of delivery — a post-op questionnaire while anesthesia is wearing off.
Attrition effectParticipants drop out; reduces sample size.
Non-response biasResponders differ from non-responders; may not reflect the population.
Performance biasCare differs between groups other than the topic of interest — controls seeking other treatment. Blinding participants protects against it.
Reporting vs publication biasReporting = only a subset of the data is made available (the study appears, not all its findings). Publication = non-significant research is less likely to be published, via editor preference and sponsors avoiding unfavorable results.
Conflict of interestFinancial or personal considerations affecting objectivity; impact may be conscious or unconscious; should be disclosed by authors.
The statistics limitAnalysis estimates the effect of chance. No analysis corrects for unknown bias. Statistics on a poor design give a false sense of respectability.
Internal validityDegree results are correct for the sample studied. “Are the conclusions valid for these people?”
External validityDegree results are true in other settings — generalisability. “Does the sample represent my patients?” Rarely determined from a single study.
VariablesIndependent = possible cause. Dependent = possible effect. Extraneous (covariate) = may impact the relationship.
Primary vs secondary researchPrimary collects its own data (questionnaires, surveys, measurements, labs) → original articles. Secondary analyses existing data → meta-analyses, systematic reviews.

Evidence & the Hierarchy

TermWhat you need to know
OriginsMcMaster University, 1970s. Sackett gave the core concepts and definition; Cochrane advocated randomized trials and systematic reviews to verify treatment effectiveness.
DefinitionThe conscientious, explicit and judicious use of current best evidence in decisions about individual patients — integrating clinical expertise with external evidence, and considering the patient's situation and preferences.
Five challengesSearching takes time · evidence unavailable · conflicting · outdated · written in language clinicians struggle with.
Three types of evidenceDescriptive (characteristics of patients with a diagnosis) · Assessment (which tests identify/predict it) · Treatment (what works, in whom).
Higher levels are…Least vulnerable to bias · more generalizable · outcomes more likely attributable to the construct studied.
Pyramid, base to apexExpert opinion/background → case-control, case series/reports → cohort → randomized controlled trials → critically-appraised articles → critically-appraised topics → systematic reviews. (The laboratory version adds in vitro and animal research below case reports.)
Filtered vs unfilteredFiltered = already appraised for you (systematic reviews, critically-appraised topics and articles). Unfiltered = primary studies you appraise yourself (randomized trials, cohort, case-control).
Systematic review vs meta-analysisReview identifies and evaluates multiple studies for one specific, focused question. Meta-analysis pools their data into a single large study — greater n usually means greater power.
USPSTF certainty= the likelihood the net benefit assessment is correct. Net benefit = benefit minus harm in a general primary care population. High = consistent results, representative populations, future studies unlikely to change it. Moderate = sufficient but constrained (number/size/quality, inconsistency, generalisability, coherence). Low = insufficient to assess effects.
USPSTF gradesA high certainty, substantial benefit → offer. B high certainty moderate, or moderate certainty moderate-to-substantial → offer. C at least moderate certainty of small benefit → offer selectively. D moderate/high certainty of no net benefit or net harm → discourage. I insufficient evidence → balance cannot be determined.
D vs ID is a finding against the service. I is the absence of a finding either way.
Judging sourcesPeer-reviewed and government databases; be cautious of websites. ABCD: Author (credentials) · Bias · Content (level of detail) · Date (still current?). The library's CRAAP test does the same job.
DatabasesMedline/PubMed (National Library of Medicine, original articles) · Cochrane Library (systematic reviews) · Evidence Based Medicine Reviews · Up-To-Date, TripPro, Clinical Key, Essential Evidence Plus · ACP Journal Club. Guidelines: AHRQ (links to USPSTF), MMWR, Task Force on Community Preventive Services.
Steps to EBPFocused clinical question → systematic review or an evidence-based synopsis from a credible resource → evaluate validity → apply to the patient → evaluate outcome.
Joint decision makingDisclosure of risks and benefits → exploration of patient values → the actual decision.

Study Design

TermWhat you need to know
Qualitative vs quantitativeQualitative: naturalistic, inductive, develops descriptions; data are observation, field notes, recordings; case study, biography, ethnography, phenomenology. Quantitative: positivistic, deductive, develops predictions; numerical, objective, measurable; trials, cohort, case-control.
Study designThe organization and plan for data collection and analysis. Strong design → strong data → valid, reliable results.
Descriptive vs analyticDescriptive describes one group (case report, case series). Analytic compares two or more groups to draw inference for a population (randomized trials, cohort, case-control).
Experimental vs observationalExperimental: the investigator assigns and manipulates the intervention (randomized trial). Observational: the exposure is observed (cohort, case-control). Cross-sectional is observational and appears under both descriptive and analytic.
Other methodsLongitudinal (changes over time) · survey · policy research/evaluation — all observational. Quasi-experimental: controls treatment but not assignment. Field research (epidemiologic investigations, new screening tests, vaccine trials): either.
Design goalAn unbiased comparison between a group with the factor and one without. Good design permits group comparison, quantification of risk difference, determination of temporal sequence, and minimized bias.
Cross-sectionalObservational, generates hypotheses. Single point in time; risk factor and disease measured together, so temporal sequence cannot be established. Ecologic versions use population-level data — beware the ecologic fallacy (applying aggregate data to an individual).
Case-controlObservational; generates or tests hypotheses. Groups defined by outcome, then assessed for a difference in exposure. Looks backward.
CohortObservational; tests hypotheses. Groups defined by exposure, followed forward. Comparison group as similar as possible except for the exposure. Follow long enough for the outcome and completely, so attrition does not skew it. Prospective (concurrent) assembles now; retrospective (historical) defines a past risk group and follows it to the present.
Randomized controlled trialExperimental; tests hypotheses. Clear hypothesis, protocol, results applicable to a larger population, random assignment, single or double blinding, adequate sample size, accounts for withdrawals and drop-outs.
Design hierarchySome designs are stronger: they prove cause and effect; weaker ones prove relationships or predictability. All studies serve a purpose and can guide practice.
Typical sequenceClinical observations → available data → case-control → cohort → randomized trials.

Rates, Incidence & Prevalence

TermWhat you need to know
RateThe number of times an event occurs during a fixed time period. Rates allow comparisons — that is why they exist rather than raw counts.
Crude death rateAll deaths ÷ mid-interval population ×1,000 or 100,000. Age-adjusted removes the effect of age structure (US 2009: crude ~794, adjusted 741 per 100,000).
Cause- and age-specificCause-specific: deaths from one cause ÷ mid-interval population ×100,000. Age-specific: limited to one age group.
Proportionate mortalityDeaths from a cause ÷ total deaths ×100. NOT a rate — the denominator is deaths, not population. All causes must sum to 100%. (CVD ~24.5% of US deaths.)
Death-to-case ratioDeaths from a cause ÷ new cases of the same disease ×100.
Maternal & child ratesAll take live births as the denominator. Neonatal <28 days · postneonatal 28–364 days · infant <1 year (×1,000) · maternal pregnancy-related deaths (×100,000). Infant mortality is the usual international comparison.
Case fatality / YPLLCase fatality = disease severity: proportion diagnosed who die from it. Years of potential life lost = dying young loses more productive years; used to target interventions.
PrevalencePortion of a group with a condition at a point in time — a snapshot. Point (single moment per person, not necessarily the same date) vs period (cases present during a span). Surveys collect it. Prevalence study = cross-sectional (SchoolNuts).
IncidencePortion who develop a condition over a period — new cases in previously healthy people. Cumulative incidence = fixed-size group. Incidence density = changing population, denominator in person-time (1 person × 10 years = 10 person-years); error arises from varying follow-up duration. Incidence study = cohort (Framingham, 1948–).
Which is bigger?Chronic disease: incidence LOWER than prevalence (cases accumulate). Acute illness: prevalence may be LOWER than incidence (people recover — the common cold).
Prevalence clinicallyPrioritizes the differential (horses, not zebras) and sets the pretest probability.
What distorts measurementCase definition — most clinical information is on a continuum, so where the line falls changes frequency. Population at risk — only the susceptible (no men in an ovarian cancer denominator); narrowing it may curtail generalisability.
SamplingRandom = equal chance. Probability = known chance, not necessarily equal (useful for subgroups). Convenience = non-random, easy to obtain. Grab = whoever researchers can find.
Distribution: person, place, timePerson = profile of those affected, demographics and behavior. Place = endemic, limited by geography. Time = epidemic (new cases concentrated in time) and pandemic (widespread; rapid like 1900s influenza or slow like HIV/AIDS).
SurveillanceActive or passive; national, state, county. Establishes a baseline. Functions: magnitude · natural history · detect outbreaks · document geographic spread · test etiology (via case-control or cohort) · evaluate control · monitor agent change. Strategies: screen high-risk groups · symptoms/syndrome · mortality (some cases detected only at death) · provider reporting · laboratory reporting.
Outbreak investigationEpidemic vs endemic → attack rate → establish diagnosis → case definition → confirm an epidemic truly exists → characterize by time, place, person → epi curve (cases vs time) → hypotheses about source/type/route → test → control measures → solve source → decide it is over. (Philadelphia Legionnaires', 1976.)

Data, Validity & Variation

TermWhat you need to know
Three types of dataNominal — categories, no order (blood group, eye color); dichotomous is nominal split in two. Ordinal — order, no defined intervals (cancer stage I–IV, edema 1+–4+); you can say more/less, not by how much. Interval — order with equal defined intervals; continuous (blood pressure, weight, limited by measurement precision) or discrete (number of children).
Validity vs reliabilityValidity (accuracy) = measures what it is supposed to; ensured by calibration against known values, or by structured questions for subjective measures (CAGE, pain scales). Reliability (reproducibility, repeatability, precision) = repeated attempts give similar results. A measure can be reliably wrong.
Three validitiesContent — includes all relevant dimensions, excludes irrelevant ones. Criterion — predicts/associates with an observable event. Construct — relates to other measures of the same phenomenon.
Types of variationIntrasubject (within one person) · intraobserver (one reader, two occasions) · interobserver (two readers — murmur graded 3 vs 4).
Biologic vs systematicTrue biologic variation = many unknown factors each with a small random effect. Systematic variation = known conditions affecting the value (time of day, after a cigarette, position, white coat).
Measurement errorRandom (chance; equally above or below) → lack of reliability. Systematic (a flaw in the process) → lack of validity; determines a range for the true score.
Intra-patient variationMay be biologic variation, measurement error, or biologic change representing disease. Intervene only for a true pathological change.
Effects of variationRandom balances out — reduce by averaging, estimate statistically; individual measurements can still mislead. Systematic will bias results, and averaging does not help.
Central tendencyMean = sum ÷ n; best estimate, but affected by outliers. Median = middle score. Mode = most frequent value.
DispersionRange = lowest to highest. Interquartile range = limited index, e.g. 25th–75th percentile. Standard deviation = variability about the mean.
Normal distributionGaussian / bell curve. ~2/3 within 1 SD, 95% within 2 SD. Based on mathematical theory, not actual measurement.
Normal vs abnormalNo clear-cut point: susceptibility → pre-symptomatic → clinical disease → dysfunction → death. Statistical abnormality ≠ clinical abnormality (cholesterol 201). You can be abnormal and well (sickle cell trait). Extremes may be desirable (low systolic, high HDL). A normal test does not rule out disease.
Defining abnormalityUse values associated with disease, disability or death; a clinically recognizable difference from healthy. Repeat abnormal tests — regression to the mean. Cutoffs are often arbitrary. Be careful of labeling.

Diagnostic Tests

TermWhat you need to know
Before orderingHorses, not zebras. Uncommon presentations of common disease beat common presentations of rare disease. No disease is rare to the patient who has it. Avoid indiscriminate ordering · risk:benefit · know the limits · do not repeat without indication · consider cost.
Diagnostic vs screeningDiagnostic = symptomatic individuals. Screening = asymptomatic individuals.
Gold standardA procedure believed to identify the disease with certainty — culture for strep, biopsy for cancer, catheterization for coronary disease, endoscopy for peptic ulcer. Often expensive, invasive/risky, unacceptable, or not easily completed, so alternatives are usually the initial choice.
The 2×2, lettereda = true positives · b = false positives · c = false negatives · d = true negatives.
SensitivityProbability a person with disease tests positive. TP/(TP+FN) = a/(a+c). Read down the disease column → fixed property of the test.
SpecificityProbability a person without disease tests negative. TN/(TN+FP) = d/(b+d). Read down the no-disease column → fixed property of the test.
Predictive valuesPPV = probability a positive result means disease = TP/(TP+FP) = a/(a+b). NPV = probability a negative result means no disease = TN/(TN+FN) = d/(c+d). Read across the test rows → depend on prevalence. Prevalence = (a+c)/(a+b+c+d).
SnOUT / SpINSn-OUT: a negative on a highly sensitive test rules OUT (few false negatives) — matters when missing disease is dangerous. Sp-IN: a positive on a highly specific test rules IN (few false positives) — matters when false positives trigger risky or costly work-ups.
Sensitivity vs specificityInversely related. The ROC curve plots true positive rate against false positive rate (1−specificity); the better test's curve is closest to the upper left with the greatest area beneath, and it also sets the cutoff point. Screening strategy: screen sensitive, confirm specific.
Prevalence effectSame test, sensitivity 99% / specificity 95%: at 1% prevalence PPV = 17%, NPV 99.9%; at 5% prevalence PPV = 51%, NPV 97%. As prevalence falls, PPV falls and NPV rises (more false positives). Spectrum bias.
Pre- and post-test probabilityPretest = chance of disease before the test, often the prevalence. Posttest = chance after it; informs treatment.
Likelihood ratiosHow much a result changes the probability. Not affected by prevalence. LR+ = Sn/(1−Sp) · LR− = (1−Sn)/Sp. LR+ of 4 → 4× more likely; LR− of 0.5 → half as likely.
LR bands>10 strong rule-in (+45%) · 5–10 moderate rule-in (+30%) · 0.5–2 indeterminate (±15%) · 0.2–0.5 weak rule-out (−15%) · 0.1–0.2 moderate rule-out (−30%) · <0.1 strong rule-out (−45%). The shift is added to the pretest probability. A very high LR rules in when pretest probability is 30–70%. Even a high LR misleads at low prevalence.
Getting to posttestConvert pretest probability → odds, multiply by the LR → posttest odds, convert back. Or the nomogram: line from pretest probability through the LR to the right-hand column.
Worked exampleRapid strep antigen, Sn 65% / Sp 96%. LR+ = 0.65/0.04 = 16.25; LR− = 0.35/0.96 = 0.36. Pretest from the modified Centor score (fever >38°C, absent cough, tender anterior nodes, tonsillar swelling/exudate = 1 each; age 3–14 = +1, 15–44 = 0, >45 = −1). Posttest ~93% positive, ~22% negative.
Parallel vs serialParallel = all at once, when speed matters; ↑sensitivity and NPV; gain depends on how well the tests complement each other. Serial = one after another, when speed matters less or tests are risky/costly; ↑specificity and PPV, ↓sensitivity and NPV; fewer tests, more time. Do the most specific first unless another is cheaper or safer (exercise tolerance test → catheterization).
Appraising a new testDisease of interest? Spectrum of patients? What is the test? Did every subject undergo both it and the gold standard? Were the testers blinded to the other result? Benefits and risks reasonable? Sn, Sp, PPV, NPV? Similar prevalence in your population?

Risk & Ratios

TermWhat you need to know
Risk factorA factor associated with an increase in the probability of developing the disease later. Raises probability, never certainty — it describes groups. Example, heart failure: smoking · hypertension · hypercholesterolemia · family history · diabetes · valvular disease.
Recognizing riskLarge risk with rapid effects = easy. Harder: long latency obscures results · frequent exposure hidden unless exposed is compared with unexposed · low incidence (a small risk needs a larger affected population) · common disease with known risk factors hides new ones. Rare disease: cohort needs a very large group, so case-control is better.
Uses of riskPrediction (probabilities from groups) · Cause (may predict without causing — a disease marker) · Diagnosis (best predictor in an individual = early clinical disease) · Prevention (remove the factor to lower probability: control hypertension, avoid second-hand smoke).
Studying risk in humansExperiments are limited by ethical constraints. Use observational designs: concurrent (prospective) cohort, historical (retrospective) cohort, case-control.
Absolute riskThe incidence of disease in the population: new cases over a given time ÷ people in the group (18 of 600 = 0.03). Does not show whether exposure raises risk.
Attributable riskExcess incidence in the exposed: incidence in exposed minus incidence in non-exposed. A difference, exposed first (0.03 − 0.01 = 0.02).
Relative riskIncidence in exposed ÷ incidence in non-exposed: how many times more likely, and the strength of association. Use incidences, never raw counts (36/300 ÷ 20/400 = 2.4, not 36/20). Exposed incidence = relative risk × unexposed incidence.
Reading relative risk= 1 null value, equal risk, no association · > 1 higher risk in exposed (2.5 = 2.5 times as likely) · < 1 negative association, may be protective (0.6 = 40 percent lower; effective vaccine). Prevention expects < 1. It is a ratio, so never negative.
Relative risk vs attributable riskRelative risk cancels out incidence, so it hides the size of the risk. Same relative risk of 2.0 gives attributable risk 0.001 at baseline 1 in 1,000 but 0.1 at 1 in 10 — higher baseline = greater attributable risk.
Cohort studyGroup with a common characteristic (age, birth date, exposure, disease, intervention) assembled at a similar point and followed over time. Needs enough follow-up time and complete follow-up so attrition does not skew the outcome. Divide the disease-free cohort by exposure, follow, compare proportions who develop disease.
Historical vs concurrentHistorical (retrospective): assembled in the past from records, followed forward to the present — data may be incomplete. Concurrent (prospective): assembled now, followed into the future. A case-control study grouped by disease is not a historical cohort. Random selection is not random assignment, so it is not a randomized trial.
Cohort pros and cons+ establishes incidence directly · logical (exposed, do they get disease?) · exposure recorded without knowing the outcome · many outcomes per exposure. − inefficient (large cohort) · expensive · years to results · only exposures recorded at the outset · susceptible to confounding.
Prospective vs retrospective vs case-cohortProspective: wide range of risk factors, lifestyle data, standardized exposure; but not for rare diseases. Retrospective: faster, cheaper, more efficient; but narrower risk factors, cannot examine characteristics missing from the data, exposure not standardized. Case-cohort: even more efficient (only a sample of the unexposed analyzed); weighting hard to follow.
ConfoundingSystematic error (bias) from an additional independent factor related to both exposure and outcome. Example: smoking, coffee and pancreatic cancer.
Controlling confoundingRandomization (equal chance; design) · Restriction (limit patient characteristics; design) · Matching (same characteristics; design and analysis) · Stratification (rates within strata; analysis) · Simple adjustment (one or few characteristics) · Multivariable adjustment (many factors, modeling) · Best-case/worst-case analysis (extreme selection-bias assumptions). Design-only: randomization and restriction.
Effect modificationA variable's presence or absence changes the effect of the exposure: synergistic (strengthens) or antagonistic (weakens). NOT a bias — confounding is. Example: aspirin, age and prior history and upper gastrointestinal bleeding.
Case-control studyCases have the condition, controls do not; controls from the same cohort, community/hospital, or matched. Groups chosen by outcome; looks backward to past exposure; yields an estimate of relative risk.
Case-control pros and cons+ smaller samples (cases picked for having the condition) · no waiting for the condition · suits rare disease with long latency (cohorts are poor at this: latency, large sample, few cases). − recall bias — exposure is often asked, and memory is imperfect · no incidence.
Relative risk vs odds ratioCohort → relative risk (incidence can be calculated directly). Case-control → odds ratio, an estimated relative risk. Lettered table (a exposed cases, b exposed noncases, c unexposed cases, d unexposed noncases): relative risk = [a/(a+b)] ÷ [c/(c+d)]; odds ratio = (a/c) ÷ (b/d) = ad/bc.
Odds ratio calculationsOdds = exposed ÷ unexposed (48 exposed, 12 unexposed = 4, not 0.8). Odds ratio = odds in cases ÷ odds in controls, cases on top. 60/100 cases vs 30/100 controls exposed: 1.5 ÷ 0.43 = 3.5. Below 1 (20/100 vs 40/100 = 0.38) = lower odds of disease.
Stating an odds ratioRead either way: cases have 11 times greater odds of prior exposure, so the exposed have 11 times greater odds of disease (the version to use). It is a ratio of odds, not a percentage or a one-in-N frequency.

Prognosis & Outcomes

TermWhat you need to know
PrognosisThe prediction for the course of disease. Expressed as rates (five-year survival, case fatality, response, remission, recurrence) or with cure / death as endpoints. Natural history = what occurs without intervention; prognosis can often be improved by clinical intervention. Detect earlier → maximize effectiveness of treatment.
Five Ds & severityDeath, disease, discomfort, disability, dissatisfaction. Severity establishes priorities for clinical services and public health programs. Compare new treatments with the natural (untreated) outcome.
Clinical courseBiologic onset → pathology evident → signs and symptoms → medical care sought → diagnosis given → treatment recommended and implemented → outcome. Preclinical phase = disease present, no signs or symptoms yet (where screening works); clinical phase = from signs and symptoms on.
Prognostic factorsIdentify patients with the same disease but different prognoses. Prognosis studies look for protective or progressing factors. Multiple factors predict better than a single data point. Clinical prediction rules = probability of an outcome from a set of patient characteristics. Adjustment asks whether some other factor explains the difference.
Risk vs prognosis studiesRisk: healthy participants · outcome is diagnosis · rates low (diagnosis rare). Prognosis: sick participants · outcome is death, complications, disability · rates high (outcomes common) · protective or progressing factors differ from risk factors.
Risk factors ≠ prognostic factorsRisk factors do not necessarily mean a poorer prognosis. Infarction: on both lists = increasing age, cigarette smoking. Risk only = male, hypertension, low-density / high-density lipoprotein pattern, inactivity. Poor prognosis only = female, hypotension, anterior infarction, congestive heart failure, ventricular arrhythmia.
Prognosis rates5-year survival = percent alive 5 years from some point in the disease · case fatality = percent of patients who die of it · disease-specific mortality = per 10,000 / 100,000 population · response = some improvement after an intervention · remission = disease no longer detectable · recurrence = return after a disease-free interval. Observation is assumed long enough for all events to occur.
Outcomes to includeThe full range of manifestations: biological and clinical outcomes, plus quality of life.
Case-fatality ratioPeople who die ÷ people who have the disease. Suits acute, short-term disease where death follows diagnosis quickly. Beware negative framing (number dying vs number surviving) — it sways how risk is perceived.
Five-year survivalFrequent in cancer. Rates can mask data: dissecting aneurysm (steep first-year drop, then plateau), lung cancer (steady), amyotrophic lateral sclerosis (slow then faster) and age 100 can share nearly the same five-year rate. Cannot evaluate disease of <5 years' duration or a new therapy introduced within 5 years.
Screening artifactDisease found earlier but death at a similar point → patients just live longer with a diagnosis. Measured survival = year of death − year of diagnosis (1989 → 1995 = 6 years vs 1991 → 1995 = 4, same death date). Earlier diagnosis adds that many years.
Survival analysisFollow a cohort over time; person-years contributed by people added at different times. Survival to a point = cumulative probability of surviving each preceding interval (multiply them: 0.95 × 0.90 × 0.60 = 0.513).
Kaplan-MeierA new interval begins at the exact time of each death → a stepwise curve. Censored patients leave the number at risk without counting as deaths. Interval survival = survivors ÷ at risk (7 of 8 = 0.875).
Cox regression & hazard ratioLike Kaplan-Meier but adjusts for confounders and handles multiple predictors (treatment, age, dose); finds prognostic indicators. Hazard ratio = relative risk: 2 = twice the risk; <1 protective; 1 = same risk; 0.25 = 25% of the comparison risk.
Survival curvesEstimate survival for a hypothetical, not actual, cohort. Precision is greatest early (left; larger population). A flat plateau can deceive: fewer participants remain followed.
Median & relative survivalMedian = time half the population survives; less affected by extremes, can be found before all have died (2, 4, 7, 9, 13, 18, 30 → 9; the mean, 11.9, is pulled up). Relative survival = observed ÷ expected survival without the disease (63 ÷ 90 = 70%).
Inception cohort & zero timeIdentify all members at a similar point (zero time); assemble at onset, when disease first becomes clinically apparent, as early and uniform as possible; state inclusion criteria; consider referral pattern (academic center vs community).
Selection biasSources: convenience samples, volunteer samples, declines from those selected, dropouts. Ask whether the prognosis difference is due to the factor studied or another factor. Controls: randomization, restriction, matching, stratification, adjustment, best/worst case variation. Randomization alone also covers unmeasured factors.
Cohort biasesSusceptibility: groups differ at entry (disease course, concurrent disease, prior treatment). Migration: dropouts or group changes (death, recovery, disability); more dropouts = more bias; are they like those who stay? Measurement: outcome detected more in one group → blind, strict criteria, equal effort. Bias may falsely create or obscure differences.
Appraising a prognosis studyInception cohort? · Entry criteria objective and reasonable? · Follow-up adequate? · Patients similar to yours (generalizability limited)? · Adjusted for prognostic factors? · Outcomes assessed objectively and blindly?

Prevention & Screening

TermWhat you need to know
Prevention definedInterventions for people without complaints to identify and modify risk factors to avoid the onset of disease, or to find disease early in its course.
Four typesImmunization (vaccine) · screening (testing people without symptoms) · behavioral / lifestyle counseling (advice about habits) · chemoprevention (a drug given to a healthy person to lower disease risk). These are types, not levels.
Three levelsPrimary = remove causes or risk factors before disease exists; individual or community level. Secondary = detect disease while asymptomatic (screening). Tertiary = treatment after diagnosis to reduce complications.
Levels on the timelineNo disease (primary) → onset → asymptomatic disease (secondary) → clinical diagnosis → clinical course (tertiary). Influenza vaccine in a healthy adult = primary; treatment after a diagnosed heart attack = tertiary.
Including a conditionThree tests: burden of suffering (death, disease, disability, discomfort, dissatisfaction, destitution; mortality, hospitalizations, incidence) · quality of the screening test (sensitivity, specificity, simplicity, cost, safety, acceptability) · treatment (effectiveness, safety, cost-effectiveness; for secondary prevention, early treatment must beat later treatment). No effective treatment → screening cannot help.
Research by levelPrimary: immunization and counseling studied with randomized controlled trials; observational when randomization is not possible or ethical; longitudinal studies for safety. Secondary: outcomes must be better when disease is found and treated early; cohort design because effectiveness can take years. Tertiary: randomized trials, but they usually test therapeutic, not preventive, effects.
Prevalence vs incidence screensFirst screening detects prevalent cases (already present). Second and subsequent screenings detect incident cases (newly developed).
Lead-time biasLead time = time by which screening advances diagnosis. Survival from diagnosis looks longer only because the diagnosis is known for longer. Age at death minus age at diagnosis: diagnosed at 60 instead of 64, death at 72 either way = 12 vs 8 years, a 4-year gain that is only lead time. Real gain only if early treatment is effective and delays death.
Length-time biasScreening detects more slow-developing, better-prognosis conditions; rapid ones cause symptoms between screens and escape. Screening works best for slowly developing conditions.
Compliance biasScreened people are more compliant, more invested in their health, so they do better for reasons unrelated to screening.
Screening test propertiesHigh sensitivity early in disease (does not miss disease) · high specificity (fewer false positives) · consider positive and negative predictive values because prevalence is low · simple, inexpensive, safe.
Unintended consequencesFalse positives: expense, inconvenience, unnecessary interventions. Negative labeling effect: ethical concern. Overdiagnosis: finding disease that does not help the patient; an extreme example of length-time bias. Incidentalomas: abnormalities outside the area examined.
Community screening: aim and objectivesTest large numbers at little or no cost. Objectives: find disease so treatment cures (sexually transmitted infections, cancers) · so treatment prevents complications (hypertension) · change a behavioral risk factor (smoking) · change a metabolic risk factor (diet to lower cholesterol).
Establishing a programImportant public health problem · effective treatment available · natural history well understood · prevalent but not found in everyone · positive screens have access to care · clear who follows up results. Benefits must outweigh costs and harms; test acceptable to public and providers; sensitivity, specificity and both predictive values known.
Evidence standardConclusive evidence that screening can alter the natural history of disease in a significant proportion of those screened. Longer survival after diagnosis is not enough (lead time).
Low prevalenceHealthy people far outnumber cases, so even a highly specific test gives more false positives than true positives. Screen higher-risk populations: prevalence and positive predictive value rise. False positives = healthy × (1 − specificity); true positives = diseased × sensitivity; positive predictive value = true positives ÷ (true + false positives).
Number needed to treatPatients treated for one to benefit. = 1 ÷ absolute risk reduction (control risk − treated risk, as a decimal): 10% to 6% = 0.04 → 25. Large = largely ineffective; small = largely effective. Cases prevented = patients ÷ number needed to treat.
Number needed to harmPatients given the intervention for one to be harmed. = 1 ÷ absolute increase in risk: 1% to 4% = 0.03 → about 33. Small = frequent harm; large = safer. Harms = patients ÷ number needed to harm.
Ideal pairing and trapsIdeally low number needed to treat, high number needed to harm. A number needed to harm below the number needed to treat means harm is more frequent than benefit. Use the risk difference, not either risk alone or the difference in counts. Baseline risk matters: a treatment that halves risk gives a bigger number in a low-risk group.

Research & Trials

TermWhat you need to know
Testing ideasAvoid decisions based on coincidence, biased comparisons and wishful thinking; use evidence on the relationship between treatments and outcomes.
Treatment & hypothesesTreatment = an intervention to improve the course of disease, applicable at any point in its natural course (medication, surgery, lifestyle change). Hypotheses come from mechanism/etiology, clinical observation, epidemiologic studies or trial and error; most require testing.
Comparative effectivenessDirect comparison of two or more interventions believed to improve outcomes; identify any harms and benefits and compare end results. Cohort studies are limited by systematic differences between groups; the randomized controlled trial is the gold standard for treatment efficacy.
RCT design, in orderResearch question → sample from the population → randomly assign to treatment groups → follow both groups forward, measure the outcome → calculate risk in each group → compare with relative risk, risk difference or hazard ratio.
Risk arithmeticRisk = patients with the outcome ÷ patients in the group. Relative risk = treatment-group risk ÷ comparison-group risk. Risk difference = comparison risk − treatment risk (percentage points). Example: 15/100 vs 30/100 → RR 0.5, difference 15 points.
Sampling vs assignmentSampling picks who enters the study; random assignment picks each subject’s group and limits systematic differences between groups. Stratified randomization: randomize separately within each stratum, then combine treatment and control patients across strata.
Sampling methodsSample = participants who represent the larger population (an entire population is rarely studied; defining it too wide or narrow causes problems). Random = same chance for every member, the best way to a representative sample. Stratified = subgroups mirroring the population (or equal groups by diagnosis). Cluster = whole groups by membership, e.g. all patients in designated hospital units.
Inclusion & exclusionInclusion = characteristics participants must have (the diagnosis, unstable glucose). Exclusion = factors that prevent inclusion (comorbidities, contraindication to the treatment under evaluation). Each added criterion shrinks the eligible pool, sometimes to a small fraction; multiply the percentages in sequence.
Trial ethicsRandomize only when no conclusive evidence favors one treatment. Primary outcome = benefits, not harms. Participants give consent. Evidence of harm, futility or effectiveness ends the study.
Comparison groups & placeboStandard comparisons: treatment vs placebo, vs another treatment, vs non-treatment. Placebo = indistinguishable from the treatment with no specific mechanism of action; helps blinding. Placebo effect = condition changes because change is expected. Hawthorne effect = behavior changes because of special attention.
ComplianceDegree participants act in accordance with instructions. Not always intentional: misunderstandings, lack of transportation, forgetfulness.
BlindingBlinding = masking: participants unaware of the treatment. Single-blind = subjects unaware, researchers aware. Double-blind = both unaware. Allocation concealment = researchers who assign patients are blinded. Open label = no blinding.
Biases in a trialParticipants without the disease · compliance · cross-over (switching from one randomly assigned treatment to another) · unacknowledged co-interventions (unrecorded extra treatment) · biased outcome assessment (fix with objective criteria and blinding). Side effects can unmask the group.
Four possible outcomesConclusion vs reality: correct when they match. Two errors: concluding a difference that does not exist, and missing a difference that does.
Efficacy vs effectivenessEfficacy = ideal trial circumstances. Effectiveness = real, ordinary circumstances (takes in compliance, practicality, cost). Internal validity = results correct for the sample studied. Generalizability = results apply to other settings.
Superiority vs non-inferioritySuperiority = is one treatment better? Non-inferiority = is a new treatment unlikely to be less effective than an established one (it may be cheaper, safer, easier)? Non-inferiority usually needs a larger sample because differences may be small.
Limits of RCTsConfounding may exist in smaller trials; expensive; logistical issues: sample selection, randomization, blinding, consistency across participants, time.
Observational studiesCollect information on patients currently being treated. Real-world patients (benefit and drawback). Confounding and effect modification are hard to avoid; findings often agree with RCTs but can differ dramatically; conclusions from uncontrolled trials are hard to draw.
Confounding vs effect modificationConfounding = additional independent factor related to both exposure and outcome (sicker patients get the drug and die more). Effect modification = presence or absence of a variable changes the effect of the exposure.
Phases & N = 1Phase I: small, no control group, safe tolerated dose. Phase II: small, preliminary effect and dose relationship. Phase III: larger, randomized, definitive evidence and common side effects (not uncommon ones). Postmarketing surveillance: large numbers after release, finds uncommon side effects. N = 1: one patient, one treatment at a time, provider and patient blinded.
Appraising a treatment studyControlled and randomized? Patients similar to yours? Complete follow-up (best- and worst-case scenarios for dropouts: if every lost treated patient had the outcome, 25/88 becomes 37/100)? Blinding? Groups similar? Clinical as well as statistical significance? Other factors considered?

Statistics & Causation

TermWhat you need to know
ProbabilityP(E) = times E occurs ÷ times E can occur (all opportunities, not just non-events — that would be odds). Lies between 0 and 1: 0 cannot occur, 1 definitely occurs, 0.5 = equal chance. All probabilities in the sample sum to 1, so P(not E) = 1 − P(E). Example: 36 of 240 = 0.15; not = 0.85.
Statistical significanceHow likely the result is to have happened by chance. More significant = less likely to be chance. Two kinds matter: statistical and clinical.
P valueProbability of an observed difference happening by chance — calculated assuming no true difference. Not the probability the null is true, that the treatment works, or that the study is biased; not the effect size. 0.5 = 50/50 · 0.05 = 1/20 (<0.05 = significant, an arbitrary convention) · 0.01 = 1/100 highly significant · 0.001 = 1/1000 very highly significant.
Reading a p valuep = 0.03: significant, but not a 3% chance the drug fails and not 97% of patients responding. p = 0.20: not significant — does not prove no effect (a low-power study can miss a real one). A lower p value does not mean a larger difference.
Null hypothesisThe hypothesis that there is no difference; usually the opposite of what researchers hope to find (drug and placebo lower pressure equally). Evidence against it → more support for a difference.
Type I error (α)False positive: stating there is a difference when there is not. What most medical research worries about; expressed by the p value (judged against the threshold, usually 0.05).
Type II error (β)False negative: stating there is no difference when there is one; the result looks not significant. Follows from study limits — short follow-up, poor compliance. When no true difference exists, the correct conclusion is “not significant”.
Statistical testsChi square: two or more proportions, many observations · Fisher exact: two proportions, few observations · Mann-Whitney U: two medians · Student t: two means · F test: two or more means · Pearson's r: association between two variables · Logistic regression: dichotomous outcome · Cox proportional hazards: time-to-event outcome · Propensity score: select matching controls.
Confidence intervalStatistical precision: a range (confidence limits) that should contain the true value; 95% = 95% of such intervals would contain it. Narrow = precise, wide = imprecise. Not the range holding 95% of patients, and does not fix bias.
Reading an intervalRelative risk: no effect = 1; interval containing 1 = not significant (e.g. 1.2, 0.8–1.8); entirely below 1 = significant reduction (0.70, 0.55–0.90). Difference: no effect = 0. Excluding the no-effect value matches p < 0.05.
Statistical powerProbability a study finds a significant difference when one really exists. Power = 1 − Pβ (type II error probability). Type II 0.20 → power 0.80; power 0.90 → type II 0.10. Low power (small trial) → likely to miss a real difference.
Multiple comparisonsStrength of tests depends on the number of questions and how focused they are; often unknown how many were made. Choosing interesting comparisons after analysis distorts the data. At p 0.05 (1/20): 60 null comparisons → about 3 chance hits; at 0.01 (1/100): 200 → about 2. Misleading recipe: many comparisons → test each → pick the significant ones → hide how many were examined.
Multivariable modelingEffect of two or more variables together (an equation): one variable's effect while accounting for others (confounding, effect modifiers), or their combined effect predicting an event. Steps: identify variables → select (strongest association first) → identify assumptions (normality, homoscedasticity, independent errors) → create the model.
Modeling limitsValidity rests on assumptions about the data; poor at identifying effect modification; may omit rare variables with a strong relationship to the outcome. Cannot prove cause or remove random error.
Statistical vs clinical significanceChance versus relevance to patients. Statistical significance does not imply clinical significance: p < 0.0001 can still be irrelevant to practice (1 millimeter of mercury drop in 40,000 patients). Larger differences tend to be more clinically significant, but small ones can still be significant. A big, important difference in a small trial (30% lower deaths, p 0.12) may be non-significant from low power.
Types of causeSufficient: if present, disease always occurs. Necessary: if absent, disease cannot occur (present → may or may not). Risk factor: raises probability; neither. Direct vs indirect (through intervening variables, e.g. obesity → blood pressure, cholesterol, glucose). Non-causal association: statistically significant, no causal relationship. Koch's postulates = single cause; also multiple causes and a web of causation.
Association to causeExclude, in order: bias (selection or measurement) → chance → confounding. Only then is a cause likely. Cause cannot be proved beyond all doubt; the more criteria met, the stronger the case.
Criteria for causalityTemporality (cause first) · strength (large relative risk) · dose-response (more exposure, more disease) · reversibility (less exposure, less disease) · consistency (replicable) · biologic plausibility · specificity (one cause, one effect — demanding) · analogy (similar exposure and disease).
Pitfalls in causal researchMeasurement bias (operator error, calibration) · recall bias (difficulty remembering) · selection bias (how people get in, poor randomization) · random error (by chance, accounted for through p values) · confounding · effect modification.
Design and causationStrongest to weakest: systematic review → randomized controlled trial → multiple time series → non-randomized trial → cohort → case-control → time series → cross-sectional → case series → case report. Random assignment balances confounders. Cross-sectional cannot show which came first.
Findings for and against causeFor: temporal sequence, small effect, specificity, analogy, biologic plausibility, consistency, large effect, dose-response, reversibility. Against: incorrect temporal sequence, no effect, not biologically plausible, no analogy, no dose-response, not reversible, not specific. A strong design finding no effect = strong evidence against cause.

Clinical Questions & Reviews

TermWhat you need to know
Narrative reviewExpert summary of evidence and recommendations. Pros: broad topics, range of issues, authors usually treat the disease. Cons: influenced by authors' values/experience · unstructured, validity questionable · cited articles may be picked to prove a point of view · may include no critique of the original research.
Systematic reviewAnswers a specific, predetermined clinical question (a narrative review is broad and open ended). Question depends on: are strong studies available, how many, do they agree. Summarized as effect size, precision, validity (internal and external).
Appraising a systematic reviewHow were studies located · inclusion/exclusion criteria · validity of original studies · biases in them · results and variation.
Clinical question formatPatient/population/problem (who) · Intervention, prognostic factor or exposure (what) · Comparison (the alternative) · Outcome (what you hope to measure). A test-accuracy question can use sensitivity and specificity as its outcome; a question asked “for all patients” needs a more specific population.
Steps of a systematic review1 Define a specific question (protocol sets question and analysis prospectively) → 2 Select studies → 3 Describe characteristics and risk of bias → 4 Summarize (forest plots, tables) → 5 Pool if justified → 6 Explain heterogeneity.
Selecting studiesExplicit criteria from the question and design · at least two reviewers confirm consistency (also for risk of bias) · list excluded studies and the rationale · assess publication and reporting bias (funnel plots).
Finding the best studiesDesign strength · sufficient sample size · kinds of bias, incl. publication bias · limitations · peer-reviewed · still current? Develop inclusion/exclusion criteria.
Publication biasResearch without significant results is less likely to be published (editor preference; sponsors avoiding unfavorable results). Always consider it: missing studies skew a review, so a pooled benefit may be overstated. A statistical model cannot fix it. Reporting bias = only a subset of a study's data is made available.
Meta-analysisCombines similar studies to examine effect size quantitatively. Weighted in proportion to study size (share = a study's participants ÷ all participants pooled). Gives a more precise estimate. Inclusion/exclusion criteria are very important.
HeterogeneityExtent to which trial results differ from one another. Test's null hypothesis: no difference between study results. Tests: Cochran's Q and the I-squared statistic. Explored by subgroup analysis, sensitivity analysis, metaregression.
Fixed vs random effectsFixed effect: no significant heterogeneity; assumes one fixed effect size, variation is chance. Random effects: heterogeneity present; allows variation not due to chance; rests on value judgment about similarity. Both give a summary effect and confidence interval; summary shown as a diamond on the forest plot.
Meta-analysis: pros and consPros: stronger, more precise estimate · detects complications · detects subgroup differences · identifies benefits vs harms. Cons: only as good as the included studies · misleading if the studies are not similar. Ask: all studies the same quality? publication bias? differences significant?
Reading a gradeNumber = strength (1 strong, 2 weak). Letter = quality of evidence (A high, B moderate, C low). Grading is a quick way to assess clinical significance, mostly for treatment and screening, and still involves value judgment.
The six grades1A consistent well-performed trials · 1B trials with important limitations · 1C observational, unsystematic experience or flawed trials · 2A high-quality evidence, benefits closely balanced · 2B moderate quality · 2C low quality, very weak.
Acting on a gradeStrong: follow unless a clear, compelling rationale for an alternative; benefits clearly outweigh risks or vice versa. Weak (2A): best action may differ with circumstances or patient/societal values. 2B: alternatives likely better for some patients. 2C: alternatives may be equally reasonable.
Guideline standardsTransparency (development and funding explicit) · Conflict of interest (financial, intellectual, institutional, patient/public disclosed) · Group composition (multidisciplinary, balanced) · Systematic review · Evidence and strength of recommendation · Description of recommendations · External review · Updating. Good guidelines give the evidence and rationale; apply clinical judgment.
Conflict of interestFinancial (who funded it, who profits) · personal/intellectual/career (personal biases, professional benefit) · institutional (benefits one organization more). Acknowledge it so readers can judge whether bias is present. Shows as selective reporting, study selection in reviews, exaggeration or fabrication.
Peer reviewSubmitted → screened by editors → forwarded to peer reviewers (reject / recommend revision / accept) → editors review language and grammar → publication. Peer review helps with conflict of interest.
Keeping currentProfessional organizations · journal articles · colleagues · patients.
Screening an articleTitle and keywords match? If not, skip. If so, read the abstract, summary and conclusion (clear aims, hypothesis, precise conclusions). Useful and relevant? If not, skip; if so, read the entire article.
Abstract vs beyondAbstract: what is it about, what was concluded, are the conclusions valid, who does it apply to, can/should I apply it. Beyond: how big the effect, how strong the methods, limitations, bias, impact, author's recommendations.
Where to lookCursory: what about → title · concluded → conclusions · likely true → design · applies to whom → patients and setting · found → results. In depth: importance → introduction · effect size → figures and tables · strength → methods · context → discussion.