Sessions 1–4 and 7–10 condensed: the bias taxonomy, the evidence hierarchy and USPSTF grades, every study design, rates, data and variation, diagnostic test interpretation, risk and ratios, prognosis, prevention and screening, trials, statistics and causation, and systematic reviews.
| Term | What you need to know |
|---|---|
| Bias | Systematic error producing conclusions that differ from the truth. Can enter at conception/design, data collection, analysis or interpretation. Potential for bias does not mean bias is present — but it is often detectable when reading the article. |
| Chance | Random variation. Equally likely above or below the truth. Reduced by good design, estimated statistically, never eliminated. May coexist with bias. |
| Telling them apart | Readings that all fall on one side of the true value = bias. Random variation would scatter both ways. |
| Selection (sampling) bias | Compared groups differ in more ways affecting the outcome than the study's focus. Consider at study design. Example: surgical arm healthy enough to tolerate surgery, medical arm sicker. |
| Measurement bias | The method of measurement gives incorrect results — blood pressure taken without a standardized procedure every time. |
| Confounding | The covariate's effect cannot be separated from the variable studied. Associated with both exposure and outcome; need not cause the disease. Consider at data analysis. May result from selection bias or chance. |
| Procedural bias | Interview/questionnaire results affected by method of delivery — a post-op questionnaire while anesthesia is wearing off. |
| Attrition effect | Participants drop out; reduces sample size. |
| Non-response bias | Responders differ from non-responders; may not reflect the population. |
| Performance bias | Care differs between groups other than the topic of interest — controls seeking other treatment. Blinding participants protects against it. |
| Reporting vs publication bias | Reporting = only a subset of the data is made available (the study appears, not all its findings). Publication = non-significant research is less likely to be published, via editor preference and sponsors avoiding unfavorable results. |
| Conflict of interest | Financial or personal considerations affecting objectivity; impact may be conscious or unconscious; should be disclosed by authors. |
| The statistics limit | Analysis estimates the effect of chance. No analysis corrects for unknown bias. Statistics on a poor design give a false sense of respectability. |
| Internal validity | Degree results are correct for the sample studied. “Are the conclusions valid for these people?” |
| External validity | Degree results are true in other settings — generalisability. “Does the sample represent my patients?” Rarely determined from a single study. |
| Variables | Independent = possible cause. Dependent = possible effect. Extraneous (covariate) = may impact the relationship. |
| Primary vs secondary research | Primary collects its own data (questionnaires, surveys, measurements, labs) → original articles. Secondary analyses existing data → meta-analyses, systematic reviews. |
| Term | What you need to know |
|---|---|
| Origins | McMaster University, 1970s. Sackett gave the core concepts and definition; Cochrane advocated randomized trials and systematic reviews to verify treatment effectiveness. |
| Definition | The conscientious, explicit and judicious use of current best evidence in decisions about individual patients — integrating clinical expertise with external evidence, and considering the patient's situation and preferences. |
| Five challenges | Searching takes time · evidence unavailable · conflicting · outdated · written in language clinicians struggle with. |
| Three types of evidence | Descriptive (characteristics of patients with a diagnosis) · Assessment (which tests identify/predict it) · Treatment (what works, in whom). |
| Higher levels are… | Least vulnerable to bias · more generalizable · outcomes more likely attributable to the construct studied. |
| Pyramid, base to apex | Expert opinion/background → case-control, case series/reports → cohort → randomized controlled trials → critically-appraised articles → critically-appraised topics → systematic reviews. (The laboratory version adds in vitro and animal research below case reports.) |
| Filtered vs unfiltered | Filtered = already appraised for you (systematic reviews, critically-appraised topics and articles). Unfiltered = primary studies you appraise yourself (randomized trials, cohort, case-control). |
| Systematic review vs meta-analysis | Review identifies and evaluates multiple studies for one specific, focused question. Meta-analysis pools their data into a single large study — greater n usually means greater power. |
| USPSTF certainty | = the likelihood the net benefit assessment is correct. Net benefit = benefit minus harm in a general primary care population. High = consistent results, representative populations, future studies unlikely to change it. Moderate = sufficient but constrained (number/size/quality, inconsistency, generalisability, coherence). Low = insufficient to assess effects. |
| USPSTF grades | A high certainty, substantial benefit → offer. B high certainty moderate, or moderate certainty moderate-to-substantial → offer. C at least moderate certainty of small benefit → offer selectively. D moderate/high certainty of no net benefit or net harm → discourage. I insufficient evidence → balance cannot be determined. |
| D vs I | D is a finding against the service. I is the absence of a finding either way. |
| Judging sources | Peer-reviewed and government databases; be cautious of websites. ABCD: Author (credentials) · Bias · Content (level of detail) · Date (still current?). The library's CRAAP test does the same job. |
| Databases | Medline/PubMed (National Library of Medicine, original articles) · Cochrane Library (systematic reviews) · Evidence Based Medicine Reviews · Up-To-Date, TripPro, Clinical Key, Essential Evidence Plus · ACP Journal Club. Guidelines: AHRQ (links to USPSTF), MMWR, Task Force on Community Preventive Services. |
| Steps to EBP | Focused clinical question → systematic review or an evidence-based synopsis from a credible resource → evaluate validity → apply to the patient → evaluate outcome. |
| Joint decision making | Disclosure of risks and benefits → exploration of patient values → the actual decision. |
| Term | What you need to know |
|---|---|
| Qualitative vs quantitative | Qualitative: naturalistic, inductive, develops descriptions; data are observation, field notes, recordings; case study, biography, ethnography, phenomenology. Quantitative: positivistic, deductive, develops predictions; numerical, objective, measurable; trials, cohort, case-control. |
| Study design | The organization and plan for data collection and analysis. Strong design → strong data → valid, reliable results. |
| Descriptive vs analytic | Descriptive describes one group (case report, case series). Analytic compares two or more groups to draw inference for a population (randomized trials, cohort, case-control). |
| Experimental vs observational | Experimental: the investigator assigns and manipulates the intervention (randomized trial). Observational: the exposure is observed (cohort, case-control). Cross-sectional is observational and appears under both descriptive and analytic. |
| Other methods | Longitudinal (changes over time) · survey · policy research/evaluation — all observational. Quasi-experimental: controls treatment but not assignment. Field research (epidemiologic investigations, new screening tests, vaccine trials): either. |
| Design goal | An unbiased comparison between a group with the factor and one without. Good design permits group comparison, quantification of risk difference, determination of temporal sequence, and minimized bias. |
| Cross-sectional | Observational, generates hypotheses. Single point in time; risk factor and disease measured together, so temporal sequence cannot be established. Ecologic versions use population-level data — beware the ecologic fallacy (applying aggregate data to an individual). |
| Case-control | Observational; generates or tests hypotheses. Groups defined by outcome, then assessed for a difference in exposure. Looks backward. |
| Cohort | Observational; tests hypotheses. Groups defined by exposure, followed forward. Comparison group as similar as possible except for the exposure. Follow long enough for the outcome and completely, so attrition does not skew it. Prospective (concurrent) assembles now; retrospective (historical) defines a past risk group and follows it to the present. |
| Randomized controlled trial | Experimental; tests hypotheses. Clear hypothesis, protocol, results applicable to a larger population, random assignment, single or double blinding, adequate sample size, accounts for withdrawals and drop-outs. |
| Design hierarchy | Some designs are stronger: they prove cause and effect; weaker ones prove relationships or predictability. All studies serve a purpose and can guide practice. |
| Typical sequence | Clinical observations → available data → case-control → cohort → randomized trials. |
| Term | What you need to know |
|---|---|
| Rate | The number of times an event occurs during a fixed time period. Rates allow comparisons — that is why they exist rather than raw counts. |
| Crude death rate | All deaths ÷ mid-interval population ×1,000 or 100,000. Age-adjusted removes the effect of age structure (US 2009: crude ~794, adjusted 741 per 100,000). |
| Cause- and age-specific | Cause-specific: deaths from one cause ÷ mid-interval population ×100,000. Age-specific: limited to one age group. |
| Proportionate mortality | Deaths from a cause ÷ total deaths ×100. NOT a rate — the denominator is deaths, not population. All causes must sum to 100%. (CVD ~24.5% of US deaths.) |
| Death-to-case ratio | Deaths from a cause ÷ new cases of the same disease ×100. |
| Maternal & child rates | All take live births as the denominator. Neonatal <28 days · postneonatal 28–364 days · infant <1 year (×1,000) · maternal pregnancy-related deaths (×100,000). Infant mortality is the usual international comparison. |
| Case fatality / YPLL | Case fatality = disease severity: proportion diagnosed who die from it. Years of potential life lost = dying young loses more productive years; used to target interventions. |
| Prevalence | Portion of a group with a condition at a point in time — a snapshot. Point (single moment per person, not necessarily the same date) vs period (cases present during a span). Surveys collect it. Prevalence study = cross-sectional (SchoolNuts). |
| Incidence | Portion who develop a condition over a period — new cases in previously healthy people. Cumulative incidence = fixed-size group. Incidence density = changing population, denominator in person-time (1 person × 10 years = 10 person-years); error arises from varying follow-up duration. Incidence study = cohort (Framingham, 1948–). |
| Which is bigger? | Chronic disease: incidence LOWER than prevalence (cases accumulate). Acute illness: prevalence may be LOWER than incidence (people recover — the common cold). |
| Prevalence clinically | Prioritizes the differential (horses, not zebras) and sets the pretest probability. |
| What distorts measurement | Case definition — most clinical information is on a continuum, so where the line falls changes frequency. Population at risk — only the susceptible (no men in an ovarian cancer denominator); narrowing it may curtail generalisability. |
| Sampling | Random = equal chance. Probability = known chance, not necessarily equal (useful for subgroups). Convenience = non-random, easy to obtain. Grab = whoever researchers can find. |
| Distribution: person, place, time | Person = profile of those affected, demographics and behavior. Place = endemic, limited by geography. Time = epidemic (new cases concentrated in time) and pandemic (widespread; rapid like 1900s influenza or slow like HIV/AIDS). |
| Surveillance | Active or passive; national, state, county. Establishes a baseline. Functions: magnitude · natural history · detect outbreaks · document geographic spread · test etiology (via case-control or cohort) · evaluate control · monitor agent change. Strategies: screen high-risk groups · symptoms/syndrome · mortality (some cases detected only at death) · provider reporting · laboratory reporting. |
| Outbreak investigation | Epidemic vs endemic → attack rate → establish diagnosis → case definition → confirm an epidemic truly exists → characterize by time, place, person → epi curve (cases vs time) → hypotheses about source/type/route → test → control measures → solve source → decide it is over. (Philadelphia Legionnaires', 1976.) |
| Term | What you need to know |
|---|---|
| Three types of data | Nominal — categories, no order (blood group, eye color); dichotomous is nominal split in two. Ordinal — order, no defined intervals (cancer stage I–IV, edema 1+–4+); you can say more/less, not by how much. Interval — order with equal defined intervals; continuous (blood pressure, weight, limited by measurement precision) or discrete (number of children). |
| Validity vs reliability | Validity (accuracy) = measures what it is supposed to; ensured by calibration against known values, or by structured questions for subjective measures (CAGE, pain scales). Reliability (reproducibility, repeatability, precision) = repeated attempts give similar results. A measure can be reliably wrong. |
| Three validities | Content — includes all relevant dimensions, excludes irrelevant ones. Criterion — predicts/associates with an observable event. Construct — relates to other measures of the same phenomenon. |
| Types of variation | Intrasubject (within one person) · intraobserver (one reader, two occasions) · interobserver (two readers — murmur graded 3 vs 4). |
| Biologic vs systematic | True biologic variation = many unknown factors each with a small random effect. Systematic variation = known conditions affecting the value (time of day, after a cigarette, position, white coat). |
| Measurement error | Random (chance; equally above or below) → lack of reliability. Systematic (a flaw in the process) → lack of validity; determines a range for the true score. |
| Intra-patient variation | May be biologic variation, measurement error, or biologic change representing disease. Intervene only for a true pathological change. |
| Effects of variation | Random balances out — reduce by averaging, estimate statistically; individual measurements can still mislead. Systematic will bias results, and averaging does not help. |
| Central tendency | Mean = sum ÷ n; best estimate, but affected by outliers. Median = middle score. Mode = most frequent value. |
| Dispersion | Range = lowest to highest. Interquartile range = limited index, e.g. 25th–75th percentile. Standard deviation = variability about the mean. |
| Normal distribution | Gaussian / bell curve. ~2/3 within 1 SD, 95% within 2 SD. Based on mathematical theory, not actual measurement. |
| Normal vs abnormal | No clear-cut point: susceptibility → pre-symptomatic → clinical disease → dysfunction → death. Statistical abnormality ≠ clinical abnormality (cholesterol 201). You can be abnormal and well (sickle cell trait). Extremes may be desirable (low systolic, high HDL). A normal test does not rule out disease. |
| Defining abnormality | Use values associated with disease, disability or death; a clinically recognizable difference from healthy. Repeat abnormal tests — regression to the mean. Cutoffs are often arbitrary. Be careful of labeling. |
| Term | What you need to know |
|---|---|
| Before ordering | Horses, not zebras. Uncommon presentations of common disease beat common presentations of rare disease. No disease is rare to the patient who has it. Avoid indiscriminate ordering · risk:benefit · know the limits · do not repeat without indication · consider cost. |
| Diagnostic vs screening | Diagnostic = symptomatic individuals. Screening = asymptomatic individuals. |
| Gold standard | A procedure believed to identify the disease with certainty — culture for strep, biopsy for cancer, catheterization for coronary disease, endoscopy for peptic ulcer. Often expensive, invasive/risky, unacceptable, or not easily completed, so alternatives are usually the initial choice. |
| The 2×2, lettered | a = true positives · b = false positives · c = false negatives · d = true negatives. |
| Sensitivity | Probability a person with disease tests positive. TP/(TP+FN) = a/(a+c). Read down the disease column → fixed property of the test. |
| Specificity | Probability a person without disease tests negative. TN/(TN+FP) = d/(b+d). Read down the no-disease column → fixed property of the test. |
| Predictive values | PPV = probability a positive result means disease = TP/(TP+FP) = a/(a+b). NPV = probability a negative result means no disease = TN/(TN+FN) = d/(c+d). Read across the test rows → depend on prevalence. Prevalence = (a+c)/(a+b+c+d). |
| SnOUT / SpIN | Sn-OUT: a negative on a highly sensitive test rules OUT (few false negatives) — matters when missing disease is dangerous. Sp-IN: a positive on a highly specific test rules IN (few false positives) — matters when false positives trigger risky or costly work-ups. |
| Sensitivity vs specificity | Inversely related. The ROC curve plots true positive rate against false positive rate (1−specificity); the better test's curve is closest to the upper left with the greatest area beneath, and it also sets the cutoff point. Screening strategy: screen sensitive, confirm specific. |
| Prevalence effect | Same test, sensitivity 99% / specificity 95%: at 1% prevalence PPV = 17%, NPV 99.9%; at 5% prevalence PPV = 51%, NPV 97%. As prevalence falls, PPV falls and NPV rises (more false positives). Spectrum bias. |
| Pre- and post-test probability | Pretest = chance of disease before the test, often the prevalence. Posttest = chance after it; informs treatment. |
| Likelihood ratios | How much a result changes the probability. Not affected by prevalence. LR+ = Sn/(1−Sp) · LR− = (1−Sn)/Sp. LR+ of 4 → 4× more likely; LR− of 0.5 → half as likely. |
| LR bands | >10 strong rule-in (+45%) · 5–10 moderate rule-in (+30%) · 0.5–2 indeterminate (±15%) · 0.2–0.5 weak rule-out (−15%) · 0.1–0.2 moderate rule-out (−30%) · <0.1 strong rule-out (−45%). The shift is added to the pretest probability. A very high LR rules in when pretest probability is 30–70%. Even a high LR misleads at low prevalence. |
| Getting to posttest | Convert pretest probability → odds, multiply by the LR → posttest odds, convert back. Or the nomogram: line from pretest probability through the LR to the right-hand column. |
| Worked example | Rapid strep antigen, Sn 65% / Sp 96%. LR+ = 0.65/0.04 = 16.25; LR− = 0.35/0.96 = 0.36. Pretest from the modified Centor score (fever >38°C, absent cough, tender anterior nodes, tonsillar swelling/exudate = 1 each; age 3–14 = +1, 15–44 = 0, >45 = −1). Posttest ~93% positive, ~22% negative. |
| Parallel vs serial | Parallel = all at once, when speed matters; ↑sensitivity and NPV; gain depends on how well the tests complement each other. Serial = one after another, when speed matters less or tests are risky/costly; ↑specificity and PPV, ↓sensitivity and NPV; fewer tests, more time. Do the most specific first unless another is cheaper or safer (exercise tolerance test → catheterization). |
| Appraising a new test | Disease of interest? Spectrum of patients? What is the test? Did every subject undergo both it and the gold standard? Were the testers blinded to the other result? Benefits and risks reasonable? Sn, Sp, PPV, NPV? Similar prevalence in your population? |
| Term | What you need to know |
|---|---|
| Risk factor | A factor associated with an increase in the probability of developing the disease later. Raises probability, never certainty — it describes groups. Example, heart failure: smoking · hypertension · hypercholesterolemia · family history · diabetes · valvular disease. |
| Recognizing risk | Large risk with rapid effects = easy. Harder: long latency obscures results · frequent exposure hidden unless exposed is compared with unexposed · low incidence (a small risk needs a larger affected population) · common disease with known risk factors hides new ones. Rare disease: cohort needs a very large group, so case-control is better. |
| Uses of risk | Prediction (probabilities from groups) · Cause (may predict without causing — a disease marker) · Diagnosis (best predictor in an individual = early clinical disease) · Prevention (remove the factor to lower probability: control hypertension, avoid second-hand smoke). |
| Studying risk in humans | Experiments are limited by ethical constraints. Use observational designs: concurrent (prospective) cohort, historical (retrospective) cohort, case-control. |
| Absolute risk | The incidence of disease in the population: new cases over a given time ÷ people in the group (18 of 600 = 0.03). Does not show whether exposure raises risk. |
| Attributable risk | Excess incidence in the exposed: incidence in exposed minus incidence in non-exposed. A difference, exposed first (0.03 − 0.01 = 0.02). |
| Relative risk | Incidence in exposed ÷ incidence in non-exposed: how many times more likely, and the strength of association. Use incidences, never raw counts (36/300 ÷ 20/400 = 2.4, not 36/20). Exposed incidence = relative risk × unexposed incidence. |
| Reading relative risk | = 1 null value, equal risk, no association · > 1 higher risk in exposed (2.5 = 2.5 times as likely) · < 1 negative association, may be protective (0.6 = 40 percent lower; effective vaccine). Prevention expects < 1. It is a ratio, so never negative. |
| Relative risk vs attributable risk | Relative risk cancels out incidence, so it hides the size of the risk. Same relative risk of 2.0 gives attributable risk 0.001 at baseline 1 in 1,000 but 0.1 at 1 in 10 — higher baseline = greater attributable risk. |
| Cohort study | Group with a common characteristic (age, birth date, exposure, disease, intervention) assembled at a similar point and followed over time. Needs enough follow-up time and complete follow-up so attrition does not skew the outcome. Divide the disease-free cohort by exposure, follow, compare proportions who develop disease. |
| Historical vs concurrent | Historical (retrospective): assembled in the past from records, followed forward to the present — data may be incomplete. Concurrent (prospective): assembled now, followed into the future. A case-control study grouped by disease is not a historical cohort. Random selection is not random assignment, so it is not a randomized trial. |
| Cohort pros and cons | + establishes incidence directly · logical (exposed, do they get disease?) · exposure recorded without knowing the outcome · many outcomes per exposure. − inefficient (large cohort) · expensive · years to results · only exposures recorded at the outset · susceptible to confounding. |
| Prospective vs retrospective vs case-cohort | Prospective: wide range of risk factors, lifestyle data, standardized exposure; but not for rare diseases. Retrospective: faster, cheaper, more efficient; but narrower risk factors, cannot examine characteristics missing from the data, exposure not standardized. Case-cohort: even more efficient (only a sample of the unexposed analyzed); weighting hard to follow. |
| Confounding | Systematic error (bias) from an additional independent factor related to both exposure and outcome. Example: smoking, coffee and pancreatic cancer. |
| Controlling confounding | Randomization (equal chance; design) · Restriction (limit patient characteristics; design) · Matching (same characteristics; design and analysis) · Stratification (rates within strata; analysis) · Simple adjustment (one or few characteristics) · Multivariable adjustment (many factors, modeling) · Best-case/worst-case analysis (extreme selection-bias assumptions). Design-only: randomization and restriction. |
| Effect modification | A variable's presence or absence changes the effect of the exposure: synergistic (strengthens) or antagonistic (weakens). NOT a bias — confounding is. Example: aspirin, age and prior history and upper gastrointestinal bleeding. |
| Case-control study | Cases have the condition, controls do not; controls from the same cohort, community/hospital, or matched. Groups chosen by outcome; looks backward to past exposure; yields an estimate of relative risk. |
| Case-control pros and cons | + smaller samples (cases picked for having the condition) · no waiting for the condition · suits rare disease with long latency (cohorts are poor at this: latency, large sample, few cases). − recall bias — exposure is often asked, and memory is imperfect · no incidence. |
| Relative risk vs odds ratio | Cohort → relative risk (incidence can be calculated directly). Case-control → odds ratio, an estimated relative risk. Lettered table (a exposed cases, b exposed noncases, c unexposed cases, d unexposed noncases): relative risk = [a/(a+b)] ÷ [c/(c+d)]; odds ratio = (a/c) ÷ (b/d) = ad/bc. |
| Odds ratio calculations | Odds = exposed ÷ unexposed (48 exposed, 12 unexposed = 4, not 0.8). Odds ratio = odds in cases ÷ odds in controls, cases on top. 60/100 cases vs 30/100 controls exposed: 1.5 ÷ 0.43 = 3.5. Below 1 (20/100 vs 40/100 = 0.38) = lower odds of disease. |
| Stating an odds ratio | Read either way: cases have 11 times greater odds of prior exposure, so the exposed have 11 times greater odds of disease (the version to use). It is a ratio of odds, not a percentage or a one-in-N frequency. |
| Term | What you need to know |
|---|---|
| Prognosis | The prediction for the course of disease. Expressed as rates (five-year survival, case fatality, response, remission, recurrence) or with cure / death as endpoints. Natural history = what occurs without intervention; prognosis can often be improved by clinical intervention. Detect earlier → maximize effectiveness of treatment. |
| Five Ds & severity | Death, disease, discomfort, disability, dissatisfaction. Severity establishes priorities for clinical services and public health programs. Compare new treatments with the natural (untreated) outcome. |
| Clinical course | Biologic onset → pathology evident → signs and symptoms → medical care sought → diagnosis given → treatment recommended and implemented → outcome. Preclinical phase = disease present, no signs or symptoms yet (where screening works); clinical phase = from signs and symptoms on. |
| Prognostic factors | Identify patients with the same disease but different prognoses. Prognosis studies look for protective or progressing factors. Multiple factors predict better than a single data point. Clinical prediction rules = probability of an outcome from a set of patient characteristics. Adjustment asks whether some other factor explains the difference. |
| Risk vs prognosis studies | Risk: healthy participants · outcome is diagnosis · rates low (diagnosis rare). Prognosis: sick participants · outcome is death, complications, disability · rates high (outcomes common) · protective or progressing factors differ from risk factors. |
| Risk factors ≠ prognostic factors | Risk factors do not necessarily mean a poorer prognosis. Infarction: on both lists = increasing age, cigarette smoking. Risk only = male, hypertension, low-density / high-density lipoprotein pattern, inactivity. Poor prognosis only = female, hypotension, anterior infarction, congestive heart failure, ventricular arrhythmia. |
| Prognosis rates | 5-year survival = percent alive 5 years from some point in the disease · case fatality = percent of patients who die of it · disease-specific mortality = per 10,000 / 100,000 population · response = some improvement after an intervention · remission = disease no longer detectable · recurrence = return after a disease-free interval. Observation is assumed long enough for all events to occur. |
| Outcomes to include | The full range of manifestations: biological and clinical outcomes, plus quality of life. |
| Case-fatality ratio | People who die ÷ people who have the disease. Suits acute, short-term disease where death follows diagnosis quickly. Beware negative framing (number dying vs number surviving) — it sways how risk is perceived. |
| Five-year survival | Frequent in cancer. Rates can mask data: dissecting aneurysm (steep first-year drop, then plateau), lung cancer (steady), amyotrophic lateral sclerosis (slow then faster) and age 100 can share nearly the same five-year rate. Cannot evaluate disease of <5 years' duration or a new therapy introduced within 5 years. |
| Screening artifact | Disease found earlier but death at a similar point → patients just live longer with a diagnosis. Measured survival = year of death − year of diagnosis (1989 → 1995 = 6 years vs 1991 → 1995 = 4, same death date). Earlier diagnosis adds that many years. |
| Survival analysis | Follow a cohort over time; person-years contributed by people added at different times. Survival to a point = cumulative probability of surviving each preceding interval (multiply them: 0.95 × 0.90 × 0.60 = 0.513). |
| Kaplan-Meier | A new interval begins at the exact time of each death → a stepwise curve. Censored patients leave the number at risk without counting as deaths. Interval survival = survivors ÷ at risk (7 of 8 = 0.875). |
| Cox regression & hazard ratio | Like Kaplan-Meier but adjusts for confounders and handles multiple predictors (treatment, age, dose); finds prognostic indicators. Hazard ratio = relative risk: 2 = twice the risk; <1 protective; 1 = same risk; 0.25 = 25% of the comparison risk. |
| Survival curves | Estimate survival for a hypothetical, not actual, cohort. Precision is greatest early (left; larger population). A flat plateau can deceive: fewer participants remain followed. |
| Median & relative survival | Median = time half the population survives; less affected by extremes, can be found before all have died (2, 4, 7, 9, 13, 18, 30 → 9; the mean, 11.9, is pulled up). Relative survival = observed ÷ expected survival without the disease (63 ÷ 90 = 70%). |
| Inception cohort & zero time | Identify all members at a similar point (zero time); assemble at onset, when disease first becomes clinically apparent, as early and uniform as possible; state inclusion criteria; consider referral pattern (academic center vs community). |
| Selection bias | Sources: convenience samples, volunteer samples, declines from those selected, dropouts. Ask whether the prognosis difference is due to the factor studied or another factor. Controls: randomization, restriction, matching, stratification, adjustment, best/worst case variation. Randomization alone also covers unmeasured factors. |
| Cohort biases | Susceptibility: groups differ at entry (disease course, concurrent disease, prior treatment). Migration: dropouts or group changes (death, recovery, disability); more dropouts = more bias; are they like those who stay? Measurement: outcome detected more in one group → blind, strict criteria, equal effort. Bias may falsely create or obscure differences. |
| Appraising a prognosis study | Inception cohort? · Entry criteria objective and reasonable? · Follow-up adequate? · Patients similar to yours (generalizability limited)? · Adjusted for prognostic factors? · Outcomes assessed objectively and blindly? |
| Term | What you need to know |
|---|---|
| Prevention defined | Interventions for people without complaints to identify and modify risk factors to avoid the onset of disease, or to find disease early in its course. |
| Four types | Immunization (vaccine) · screening (testing people without symptoms) · behavioral / lifestyle counseling (advice about habits) · chemoprevention (a drug given to a healthy person to lower disease risk). These are types, not levels. |
| Three levels | Primary = remove causes or risk factors before disease exists; individual or community level. Secondary = detect disease while asymptomatic (screening). Tertiary = treatment after diagnosis to reduce complications. |
| Levels on the timeline | No disease (primary) → onset → asymptomatic disease (secondary) → clinical diagnosis → clinical course (tertiary). Influenza vaccine in a healthy adult = primary; treatment after a diagnosed heart attack = tertiary. |
| Including a condition | Three tests: burden of suffering (death, disease, disability, discomfort, dissatisfaction, destitution; mortality, hospitalizations, incidence) · quality of the screening test (sensitivity, specificity, simplicity, cost, safety, acceptability) · treatment (effectiveness, safety, cost-effectiveness; for secondary prevention, early treatment must beat later treatment). No effective treatment → screening cannot help. |
| Research by level | Primary: immunization and counseling studied with randomized controlled trials; observational when randomization is not possible or ethical; longitudinal studies for safety. Secondary: outcomes must be better when disease is found and treated early; cohort design because effectiveness can take years. Tertiary: randomized trials, but they usually test therapeutic, not preventive, effects. |
| Prevalence vs incidence screens | First screening detects prevalent cases (already present). Second and subsequent screenings detect incident cases (newly developed). |
| Lead-time bias | Lead time = time by which screening advances diagnosis. Survival from diagnosis looks longer only because the diagnosis is known for longer. Age at death minus age at diagnosis: diagnosed at 60 instead of 64, death at 72 either way = 12 vs 8 years, a 4-year gain that is only lead time. Real gain only if early treatment is effective and delays death. |
| Length-time bias | Screening detects more slow-developing, better-prognosis conditions; rapid ones cause symptoms between screens and escape. Screening works best for slowly developing conditions. |
| Compliance bias | Screened people are more compliant, more invested in their health, so they do better for reasons unrelated to screening. |
| Screening test properties | High sensitivity early in disease (does not miss disease) · high specificity (fewer false positives) · consider positive and negative predictive values because prevalence is low · simple, inexpensive, safe. |
| Unintended consequences | False positives: expense, inconvenience, unnecessary interventions. Negative labeling effect: ethical concern. Overdiagnosis: finding disease that does not help the patient; an extreme example of length-time bias. Incidentalomas: abnormalities outside the area examined. |
| Community screening: aim and objectives | Test large numbers at little or no cost. Objectives: find disease so treatment cures (sexually transmitted infections, cancers) · so treatment prevents complications (hypertension) · change a behavioral risk factor (smoking) · change a metabolic risk factor (diet to lower cholesterol). |
| Establishing a program | Important public health problem · effective treatment available · natural history well understood · prevalent but not found in everyone · positive screens have access to care · clear who follows up results. Benefits must outweigh costs and harms; test acceptable to public and providers; sensitivity, specificity and both predictive values known. |
| Evidence standard | Conclusive evidence that screening can alter the natural history of disease in a significant proportion of those screened. Longer survival after diagnosis is not enough (lead time). |
| Low prevalence | Healthy people far outnumber cases, so even a highly specific test gives more false positives than true positives. Screen higher-risk populations: prevalence and positive predictive value rise. False positives = healthy × (1 − specificity); true positives = diseased × sensitivity; positive predictive value = true positives ÷ (true + false positives). |
| Number needed to treat | Patients treated for one to benefit. = 1 ÷ absolute risk reduction (control risk − treated risk, as a decimal): 10% to 6% = 0.04 → 25. Large = largely ineffective; small = largely effective. Cases prevented = patients ÷ number needed to treat. |
| Number needed to harm | Patients given the intervention for one to be harmed. = 1 ÷ absolute increase in risk: 1% to 4% = 0.03 → about 33. Small = frequent harm; large = safer. Harms = patients ÷ number needed to harm. |
| Ideal pairing and traps | Ideally low number needed to treat, high number needed to harm. A number needed to harm below the number needed to treat means harm is more frequent than benefit. Use the risk difference, not either risk alone or the difference in counts. Baseline risk matters: a treatment that halves risk gives a bigger number in a low-risk group. |
| Term | What you need to know |
|---|---|
| Testing ideas | Avoid decisions based on coincidence, biased comparisons and wishful thinking; use evidence on the relationship between treatments and outcomes. |
| Treatment & hypotheses | Treatment = an intervention to improve the course of disease, applicable at any point in its natural course (medication, surgery, lifestyle change). Hypotheses come from mechanism/etiology, clinical observation, epidemiologic studies or trial and error; most require testing. |
| Comparative effectiveness | Direct comparison of two or more interventions believed to improve outcomes; identify any harms and benefits and compare end results. Cohort studies are limited by systematic differences between groups; the randomized controlled trial is the gold standard for treatment efficacy. |
| RCT design, in order | Research question → sample from the population → randomly assign to treatment groups → follow both groups forward, measure the outcome → calculate risk in each group → compare with relative risk, risk difference or hazard ratio. |
| Risk arithmetic | Risk = patients with the outcome ÷ patients in the group. Relative risk = treatment-group risk ÷ comparison-group risk. Risk difference = comparison risk − treatment risk (percentage points). Example: 15/100 vs 30/100 → RR 0.5, difference 15 points. |
| Sampling vs assignment | Sampling picks who enters the study; random assignment picks each subject’s group and limits systematic differences between groups. Stratified randomization: randomize separately within each stratum, then combine treatment and control patients across strata. |
| Sampling methods | Sample = participants who represent the larger population (an entire population is rarely studied; defining it too wide or narrow causes problems). Random = same chance for every member, the best way to a representative sample. Stratified = subgroups mirroring the population (or equal groups by diagnosis). Cluster = whole groups by membership, e.g. all patients in designated hospital units. |
| Inclusion & exclusion | Inclusion = characteristics participants must have (the diagnosis, unstable glucose). Exclusion = factors that prevent inclusion (comorbidities, contraindication to the treatment under evaluation). Each added criterion shrinks the eligible pool, sometimes to a small fraction; multiply the percentages in sequence. |
| Trial ethics | Randomize only when no conclusive evidence favors one treatment. Primary outcome = benefits, not harms. Participants give consent. Evidence of harm, futility or effectiveness ends the study. |
| Comparison groups & placebo | Standard comparisons: treatment vs placebo, vs another treatment, vs non-treatment. Placebo = indistinguishable from the treatment with no specific mechanism of action; helps blinding. Placebo effect = condition changes because change is expected. Hawthorne effect = behavior changes because of special attention. |
| Compliance | Degree participants act in accordance with instructions. Not always intentional: misunderstandings, lack of transportation, forgetfulness. |
| Blinding | Blinding = masking: participants unaware of the treatment. Single-blind = subjects unaware, researchers aware. Double-blind = both unaware. Allocation concealment = researchers who assign patients are blinded. Open label = no blinding. |
| Biases in a trial | Participants without the disease · compliance · cross-over (switching from one randomly assigned treatment to another) · unacknowledged co-interventions (unrecorded extra treatment) · biased outcome assessment (fix with objective criteria and blinding). Side effects can unmask the group. |
| Four possible outcomes | Conclusion vs reality: correct when they match. Two errors: concluding a difference that does not exist, and missing a difference that does. |
| Efficacy vs effectiveness | Efficacy = ideal trial circumstances. Effectiveness = real, ordinary circumstances (takes in compliance, practicality, cost). Internal validity = results correct for the sample studied. Generalizability = results apply to other settings. |
| Superiority vs non-inferiority | Superiority = is one treatment better? Non-inferiority = is a new treatment unlikely to be less effective than an established one (it may be cheaper, safer, easier)? Non-inferiority usually needs a larger sample because differences may be small. |
| Limits of RCTs | Confounding may exist in smaller trials; expensive; logistical issues: sample selection, randomization, blinding, consistency across participants, time. |
| Observational studies | Collect information on patients currently being treated. Real-world patients (benefit and drawback). Confounding and effect modification are hard to avoid; findings often agree with RCTs but can differ dramatically; conclusions from uncontrolled trials are hard to draw. |
| Confounding vs effect modification | Confounding = additional independent factor related to both exposure and outcome (sicker patients get the drug and die more). Effect modification = presence or absence of a variable changes the effect of the exposure. |
| Phases & N = 1 | Phase I: small, no control group, safe tolerated dose. Phase II: small, preliminary effect and dose relationship. Phase III: larger, randomized, definitive evidence and common side effects (not uncommon ones). Postmarketing surveillance: large numbers after release, finds uncommon side effects. N = 1: one patient, one treatment at a time, provider and patient blinded. |
| Appraising a treatment study | Controlled and randomized? Patients similar to yours? Complete follow-up (best- and worst-case scenarios for dropouts: if every lost treated patient had the outcome, 25/88 becomes 37/100)? Blinding? Groups similar? Clinical as well as statistical significance? Other factors considered? |
| Term | What you need to know |
|---|---|
| Probability | P(E) = times E occurs ÷ times E can occur (all opportunities, not just non-events — that would be odds). Lies between 0 and 1: 0 cannot occur, 1 definitely occurs, 0.5 = equal chance. All probabilities in the sample sum to 1, so P(not E) = 1 − P(E). Example: 36 of 240 = 0.15; not = 0.85. |
| Statistical significance | How likely the result is to have happened by chance. More significant = less likely to be chance. Two kinds matter: statistical and clinical. |
| P value | Probability of an observed difference happening by chance — calculated assuming no true difference. Not the probability the null is true, that the treatment works, or that the study is biased; not the effect size. 0.5 = 50/50 · 0.05 = 1/20 (<0.05 = significant, an arbitrary convention) · 0.01 = 1/100 highly significant · 0.001 = 1/1000 very highly significant. |
| Reading a p value | p = 0.03: significant, but not a 3% chance the drug fails and not 97% of patients responding. p = 0.20: not significant — does not prove no effect (a low-power study can miss a real one). A lower p value does not mean a larger difference. |
| Null hypothesis | The hypothesis that there is no difference; usually the opposite of what researchers hope to find (drug and placebo lower pressure equally). Evidence against it → more support for a difference. |
| Type I error (α) | False positive: stating there is a difference when there is not. What most medical research worries about; expressed by the p value (judged against the threshold, usually 0.05). |
| Type II error (β) | False negative: stating there is no difference when there is one; the result looks not significant. Follows from study limits — short follow-up, poor compliance. When no true difference exists, the correct conclusion is “not significant”. |
| Statistical tests | Chi square: two or more proportions, many observations · Fisher exact: two proportions, few observations · Mann-Whitney U: two medians · Student t: two means · F test: two or more means · Pearson's r: association between two variables · Logistic regression: dichotomous outcome · Cox proportional hazards: time-to-event outcome · Propensity score: select matching controls. |
| Confidence interval | Statistical precision: a range (confidence limits) that should contain the true value; 95% = 95% of such intervals would contain it. Narrow = precise, wide = imprecise. Not the range holding 95% of patients, and does not fix bias. |
| Reading an interval | Relative risk: no effect = 1; interval containing 1 = not significant (e.g. 1.2, 0.8–1.8); entirely below 1 = significant reduction (0.70, 0.55–0.90). Difference: no effect = 0. Excluding the no-effect value matches p < 0.05. |
| Statistical power | Probability a study finds a significant difference when one really exists. Power = 1 − Pβ (type II error probability). Type II 0.20 → power 0.80; power 0.90 → type II 0.10. Low power (small trial) → likely to miss a real difference. |
| Multiple comparisons | Strength of tests depends on the number of questions and how focused they are; often unknown how many were made. Choosing interesting comparisons after analysis distorts the data. At p 0.05 (1/20): 60 null comparisons → about 3 chance hits; at 0.01 (1/100): 200 → about 2. Misleading recipe: many comparisons → test each → pick the significant ones → hide how many were examined. |
| Multivariable modeling | Effect of two or more variables together (an equation): one variable's effect while accounting for others (confounding, effect modifiers), or their combined effect predicting an event. Steps: identify variables → select (strongest association first) → identify assumptions (normality, homoscedasticity, independent errors) → create the model. |
| Modeling limits | Validity rests on assumptions about the data; poor at identifying effect modification; may omit rare variables with a strong relationship to the outcome. Cannot prove cause or remove random error. |
| Statistical vs clinical significance | Chance versus relevance to patients. Statistical significance does not imply clinical significance: p < 0.0001 can still be irrelevant to practice (1 millimeter of mercury drop in 40,000 patients). Larger differences tend to be more clinically significant, but small ones can still be significant. A big, important difference in a small trial (30% lower deaths, p 0.12) may be non-significant from low power. |
| Types of cause | Sufficient: if present, disease always occurs. Necessary: if absent, disease cannot occur (present → may or may not). Risk factor: raises probability; neither. Direct vs indirect (through intervening variables, e.g. obesity → blood pressure, cholesterol, glucose). Non-causal association: statistically significant, no causal relationship. Koch's postulates = single cause; also multiple causes and a web of causation. |
| Association to cause | Exclude, in order: bias (selection or measurement) → chance → confounding. Only then is a cause likely. Cause cannot be proved beyond all doubt; the more criteria met, the stronger the case. |
| Criteria for causality | Temporality (cause first) · strength (large relative risk) · dose-response (more exposure, more disease) · reversibility (less exposure, less disease) · consistency (replicable) · biologic plausibility · specificity (one cause, one effect — demanding) · analogy (similar exposure and disease). |
| Pitfalls in causal research | Measurement bias (operator error, calibration) · recall bias (difficulty remembering) · selection bias (how people get in, poor randomization) · random error (by chance, accounted for through p values) · confounding · effect modification. |
| Design and causation | Strongest to weakest: systematic review → randomized controlled trial → multiple time series → non-randomized trial → cohort → case-control → time series → cross-sectional → case series → case report. Random assignment balances confounders. Cross-sectional cannot show which came first. |
| Findings for and against cause | For: temporal sequence, small effect, specificity, analogy, biologic plausibility, consistency, large effect, dose-response, reversibility. Against: incorrect temporal sequence, no effect, not biologically plausible, no analogy, no dose-response, not reversible, not specific. A strong design finding no effect = strong evidence against cause. |
| Term | What you need to know |
|---|---|
| Narrative review | Expert summary of evidence and recommendations. Pros: broad topics, range of issues, authors usually treat the disease. Cons: influenced by authors' values/experience · unstructured, validity questionable · cited articles may be picked to prove a point of view · may include no critique of the original research. |
| Systematic review | Answers a specific, predetermined clinical question (a narrative review is broad and open ended). Question depends on: are strong studies available, how many, do they agree. Summarized as effect size, precision, validity (internal and external). |
| Appraising a systematic review | How were studies located · inclusion/exclusion criteria · validity of original studies · biases in them · results and variation. |
| Clinical question format | Patient/population/problem (who) · Intervention, prognostic factor or exposure (what) · Comparison (the alternative) · Outcome (what you hope to measure). A test-accuracy question can use sensitivity and specificity as its outcome; a question asked “for all patients” needs a more specific population. |
| Steps of a systematic review | 1 Define a specific question (protocol sets question and analysis prospectively) → 2 Select studies → 3 Describe characteristics and risk of bias → 4 Summarize (forest plots, tables) → 5 Pool if justified → 6 Explain heterogeneity. |
| Selecting studies | Explicit criteria from the question and design · at least two reviewers confirm consistency (also for risk of bias) · list excluded studies and the rationale · assess publication and reporting bias (funnel plots). |
| Finding the best studies | Design strength · sufficient sample size · kinds of bias, incl. publication bias · limitations · peer-reviewed · still current? Develop inclusion/exclusion criteria. |
| Publication bias | Research without significant results is less likely to be published (editor preference; sponsors avoiding unfavorable results). Always consider it: missing studies skew a review, so a pooled benefit may be overstated. A statistical model cannot fix it. Reporting bias = only a subset of a study's data is made available. |
| Meta-analysis | Combines similar studies to examine effect size quantitatively. Weighted in proportion to study size (share = a study's participants ÷ all participants pooled). Gives a more precise estimate. Inclusion/exclusion criteria are very important. |
| Heterogeneity | Extent to which trial results differ from one another. Test's null hypothesis: no difference between study results. Tests: Cochran's Q and the I-squared statistic. Explored by subgroup analysis, sensitivity analysis, metaregression. |
| Fixed vs random effects | Fixed effect: no significant heterogeneity; assumes one fixed effect size, variation is chance. Random effects: heterogeneity present; allows variation not due to chance; rests on value judgment about similarity. Both give a summary effect and confidence interval; summary shown as a diamond on the forest plot. |
| Meta-analysis: pros and cons | Pros: stronger, more precise estimate · detects complications · detects subgroup differences · identifies benefits vs harms. Cons: only as good as the included studies · misleading if the studies are not similar. Ask: all studies the same quality? publication bias? differences significant? |
| Reading a grade | Number = strength (1 strong, 2 weak). Letter = quality of evidence (A high, B moderate, C low). Grading is a quick way to assess clinical significance, mostly for treatment and screening, and still involves value judgment. |
| The six grades | 1A consistent well-performed trials · 1B trials with important limitations · 1C observational, unsystematic experience or flawed trials · 2A high-quality evidence, benefits closely balanced · 2B moderate quality · 2C low quality, very weak. |
| Acting on a grade | Strong: follow unless a clear, compelling rationale for an alternative; benefits clearly outweigh risks or vice versa. Weak (2A): best action may differ with circumstances or patient/societal values. 2B: alternatives likely better for some patients. 2C: alternatives may be equally reasonable. |
| Guideline standards | Transparency (development and funding explicit) · Conflict of interest (financial, intellectual, institutional, patient/public disclosed) · Group composition (multidisciplinary, balanced) · Systematic review · Evidence and strength of recommendation · Description of recommendations · External review · Updating. Good guidelines give the evidence and rationale; apply clinical judgment. |
| Conflict of interest | Financial (who funded it, who profits) · personal/intellectual/career (personal biases, professional benefit) · institutional (benefits one organization more). Acknowledge it so readers can judge whether bias is present. Shows as selective reporting, study selection in reviews, exaggeration or fabrication. |
| Peer review | Submitted → screened by editors → forwarded to peer reviewers (reject / recommend revision / accept) → editors review language and grammar → publication. Peer review helps with conflict of interest. |
| Keeping current | Professional organizations · journal articles · colleagues · patients. |
| Screening an article | Title and keywords match? If not, skip. If so, read the abstract, summary and conclusion (clear aims, hypothesis, precise conclusions). Useful and relevant? If not, skip; if so, read the entire article. |
| Abstract vs beyond | Abstract: what is it about, what was concluded, are the conclusions valid, who does it apply to, can/should I apply it. Beyond: how big the effect, how strong the methods, limitations, bias, impact, author's recommendations. |
| Where to look | Cursory: what about → title · concluded → conclusions · likely true → design · applies to whom → patients and setting · found → results. In depth: importance → introduction · effect size → figures and tables · strength → methods · context → discussion. |