1 · Introduction, Bias & Validity
Megan B. Finck, MMS, PA-C · Session 1
Instructional Objectives
- Identify epidemiologic data sources
- Distinguish peer reviewed literature sources from non-peer reviewed
- Identify trustworthy web-based resources of scientific evidence
- Discuss the purposes of research
- Discuss the various types of bias in medical research including: Selection/sampling bias, Measurement bias, Confounding, Procedural bias, Attrition effect, Non-response bias, Reporting Bias, Publication bias
- Describe chance and how it affects medical research
- Discuss internal and external validity
What epidemiology is
The study of the distribution and determinants of health-related states among specified populations, and the application of that study to the control of health problems. The second half of the definition is the part people forget: it is an applied discipline, not a descriptive one.
Purposes in public health practice: discover agent, host and environmental factors affecting health; determine the relative importance of causes of illness, disability and death; identify the population segments at greatest risk; evaluate whether health programs and services actually improve population health.
Contributions to medicine: investigating new disease transmission · determining preventable causes · determining natural history · studying the biological spectrum of disease · evaluating community interventions · setting disease control priorities · improving health services research · providing expert testimony.
IO 1 — Epidemiologic data sources
| Source | What it holds |
|---|---|
| United Nations Demographic Yearbook | Demographic and social statistics, international |
| World Health Organization World Health Statistics Manual | Health statistics for member states |
| Population census | National, state and county level data |
| Vital Statistics System | Births, deaths, causes of death, marriages, divorces |
| National Center for Health Statistics | Periodic surveys of disease and disability (National Health Interview Survey, National Health and Nutrition Examination Survey) |
| Disease reporting system | Centers for Disease Control and Prevention with state and county health departments — reportable diseases |
| Behavioral Risk Factor Surveillance Survey | Telephone survey: risk behaviors, chronic conditions, preventive service use |
| Disease registries | State-run, disease-specific; the value is trends over time |
| Other | Hospital data (emergency, admissions, discharges) and clinic data |
Also tested
- Disease registry. Registries are usually established by state government to collect disease-specific data; because collection continues, they provide information about trends over time, showing how a disease behaves across years.
- Behavioral Risk Factor Surveillance Survey. A telephone survey of health-related risk behaviors that also captures chronic health conditions among United States residents and preventive service use.
Clinical epidemiology and evidence-based medicine
Clinical epidemiology makes predictions about individual patients from events in groups of similar patients. Strong scientific methods are what make those predictions accurate.
Evidence-based medicine applies clinical epidemiology to patient care: using clinical judgment to assess available research for validity and applicability, then applying it to decisions.
Also tested
- Judging good information. A clinician judges information by critically analyzing it and distinguishing strong from weak. A trusting patient-provider relationship and the quality of the source material behind the provider's knowledge also bear on it.
- Good medical decision making. It depends on good information from history, examination and knowledge, and also relies on patient goals and values, provider beliefs, and the risks, benefits, effectiveness, adverse reactions and cost of any intervention.
- Diagnosis question. It asks how accurate the tests used to identify the disease are, so test properties are studied as a topic in their own right rather than assumed.
- Abnormality question. It asks whether the patient is sick or well, and sits alongside Frequency, which asks how often a disease occurs, as a question to settle before any answer about a patient is possible.
- Sources of clinical answers. Answers to clinical questions are drawn from experience, knowledge and clinical research. Leaning on experience alone is the eminence-based reasoning that evidence-based practice was formulated against.
The alternatives to evidence (know these by name)
| Name | What is substituted for evidence |
|---|---|
| Eminence-based | Senior colleagues who hold experience to trump evidence |
| Vehemence-based | Volume and stridency |
| Eloquence (elegance)-based | Sartorial elegance and verbal eloquence |
| Providence-based | The decision is left in the hands of the Almighty |
| Diffidence-based | Too timid to make any decision |
| Nervousness-based | Fear of litigation → over-investigation and overtreatment |
| Confidence-based | Bravado |
The five As
Ask (define the clinical question) → Acquire (systematic retrieval) → Appraise (validity and applicability) → Apply (change practice) → Assess (evaluate outcomes on patient care).
IO 4 — Purposes of research
- Exploration — new or understudied topic; satisfy curiosity; test feasibility of a more careful study; develop its methods.
- Description — careful deliberate observation and a description of what is found (disease rates, demographics of those affected).
- Explanation — attempt to explain why.
Primary research uses data the researcher collects personally (questionnaires, surveys, clinical measurements, laboratory results) — e.g. original research articles. Secondary research analyses existing data — meta-analyses, systematic reviews.
Also tested
- Secondary research. Analysis of existing data, with meta-analyses and systematic review articles as examples; primary research uses data the investigator collected personally.
Basic vocabulary
- Independent variable — the possible cause.
- Dependent variable — the possible effect.
- Extraneous variable (covariate) — may impact the relationship between the two.
- Population — all people in a defined setting or with defined characteristics. Sample — a subset of them.
Also tested
- Population versus sample. A population is everyone in a defined setting; a sample is a subset of it.
IO 5 — The bias taxonomy
Bias is systematic error producing conclusions that differ from the true results. It can enter at any stage: conception and design, data collection, data analysis, data interpretation. Potential for bias does not always mean bias is present — but it should be considered, and it is often detectable when reading the article.
| Bias | Definition | Example / note |
|---|---|---|
| Selection (sampling) | Compared groups differ in more ways affecting the outcome than the study's focus | Surgical arm healthy enough to tolerate surgery; medical arm has more comorbidities. Consider at study design. |
| Measurement | The method of measurement leads to incorrect results | Blood pressure taken without standardized procedure every time |
| Confounding | The covariate's effect cannot be separated from the variable being studied | Covariate is associated with exposure and outcome; need not cause the disease. May result from selection bias or chance. Consider at data analysis. |
| Procedural | Interview/questionnaire results affected by the method of delivery | Post-operative questionnaire while still under anesthesia's after-effects |
| Attrition effect | Participants drop out | Reduces sample size |
| Non-response | Responders differ from non-responders | May not reflect the whole population |
| Performance | Care differs between groups other than the topic of interest | Control group seeks other treatment knowing they are on placebo. Blinding participants protects against it. |
| Reporting | Only a subset of relevant data is made available | The study appears; not all of what it found does |
| Publication | Research without statistically significant results is less likely to be published | Journal editor preference; sponsors less likely to support unfavorable research |
| Conflict of interest | Financial or personal considerations affect objectivity | Impact may be conscious or unconscious; should be disclosed by authors |
Also tested
- Same amount of exercise. A difference in coronary events cannot be explained by exercise if the two groups exercised the same amount, whereas differences between the groups before the study (selection bias) or in how the outcome was detected (measurement bias) can produce such a finding.
- Recall bias is measurement bias. Recall bias, a measurement bias, occurs when people with an illness remember or report earlier exposures differently because they suspect those exposures caused their illness.
- Measurement bias. Measurement bias occurs when information is not obtained the same way in both groups, for example when clinicians question and record contraceptive use more carefully in women who have the disease.
- Measurement bias and unequal follow-up. Measurement bias means participants in one group are more likely than another to have their outcome detected, as when all surgically treated patients but only some medically treated patients remain under care, so pain is not measured with equal completeness.
- Selection bias from different neighborhoods. Selection bias occurs when compared groups differ in more than the focus of the study. Patients admitted by doctors in different neighborhoods likely differ in socioeconomic and other characteristics related to health outcomes, so the groups were not comparable to begin with.
- Publication bias. A sponsoring company may contribute because it is less likely to support research unfavorable to its product; journal editor preference is the other force keeping results that show nothing out of the literature.
- Confounding variable. If a placebo control group is older on average than the experimental group, age acts as a confounding variable, because it is plausibly associated with both group membership and outcome, so its effect cannot be separated from the medication's.
IO 6 — Chance
Random variation producing conclusions that differ from the true results. Statistical analysis can estimate the degree; good design reduces it but never eliminates it; it may coexist with bias.
Also tested
- Chance and small numbers. Chance is random variation producing conclusions that differ from the true results. Small numbers of participants increase the possibility that chance accounts for differences between groups and can give a misleading impression; statistical analysis can estimate chance, but it is never totally eliminated.
IO 7 — Internal and external validity
Two questions to ask of any study:
- Are the conclusions valid for the people in the sample? → internal validity, the degree results are correct for the sample studied.
- Does the sample represent my patients? → external validity, the degree results are also true in other clinical settings. Also called generalisability, and rarely determined from a single study.
Also tested
- External validity and other patients. External validity (generalizability) is the degree to which results are also true in other clinical settings and patients, so results from a study of men may not apply to a woman.
- Sampling and external validity. Even a correctly drawn sample can misrepresent its population by chance, especially when the sample is small, so correct sampling does not ensure external validity.
- Older patients and generalizability. Applying results from studies of younger patients to an older patient is a question of external validity (generalizability), which is a matter of personal judgment based on whatever is known about whether older patients respond the same way.
- Questions for any study. A provider should ask: are the conclusions valid for the sample (internal validity), and does the sample represent my patients (external validity)?
2 · Evidence
Megan B. Finck, MMS, PA-C · Session 2
Instructional Objectives
- Discuss the principles of evidence-based medicine
- Discuss the importance and value of evidence-based medicine
- Analyze resources for bias and validity
- Discuss the use of common databases to access medical literature
Origins and definition
Evidence-based medicine originated at McMaster University in the 1970s. Dr David Sackett's seminal work provided the core concepts and definition; Dr Archie Cochrane advocated randomized controlled trials and systematic reviews as tools for verifying treatment effectiveness.
The framing question: how do you know that what you do and how you do it really works?
Also tested
- Evidence-based practice. Beyond expertise and external evidence, it also considers the individual patient's situation and preferences.
Challenges of implementing it
- Searching and finding the best evidence is time consuming
- Scientific evidence may not be available
- Conflicting evidence
- Outdated evidence
- Findings presented in language clinicians may have difficulty understanding
Also tested
- Scale of the published literature. Volume alone does not settle clinical questions: with more than two million studies, thirty thousand journals and over nineteen million citations, the answers are still incomplete.
- Challenges to evidence-based medicine. Searching takes time, and evidence may be absent or conflicting; outdated evidence is another, and findings are often presented in language clinicians may have difficulty understanding.
Three types of evidence
| Type | The question it answers |
|---|---|
| Descriptive | What are the characteristics of clients with a specific diagnosis? |
| Assessment | What tests are most effective in identifying or predicting a diagnosis? |
| Treatment | What treatments are most effective against certain diagnoses in certain populations? |
Also tested
- Assessment evidence. Assessment evidence answers which tests best identify or predict a diagnosis. It sits between descriptive evidence, which characterizes patients with a diagnosis, and treatment evidence, which asks what works against it.
The evidence hierarchy
Studies at higher levels are least vulnerable to bias, more generalizable, and their outcomes are more likely attributable to the construct being studied.
Also tested
- A single recent article. Good systematic reviews provide the highest level of evidence. A single recent article taken out of context with other studies on the same question is of marginal value unless it is much stronger than all the studies before it.
Clinical studies pyramid — bottom to top
| Version with laboratory tiers | Version with filtered / unfiltered split |
|---|---|
|
In vitro (test tube) research Animal research Ideas, editorials, opinions Case reports Case series Case control studies Cohort studies Randomized controlled double blind studies Systematic reviews and meta-analyses |
Background information / expert opinion Case-controlled studies, case series / reports Cohort studies Randomized controlled trials FILTERED Critically-appraised individual articles (article synopses) FILTERED Critically-appraised topics (evidence syntheses and guidelines) FILTERED Systematic reviews |
Also tested
- Critically appraised individual articles. Also called article synopses, they sit above randomized trials because the appraisal has already been done, and below critically appraised topics and systematic reviews.
- Pyramid levels below case reports. In the version of the pyramid that includes laboratory work, animal research and in vitro research sit below case reports, with in vitro research at the very bottom and ideas, editorials and opinions between animal research and case reports.
- Clinical studies pyramid order. Cohort studies sit immediately below randomized controlled trials, with case-control studies and then case series and reports below; the order rises from case reports through case-control and cohort to randomized trials.
- Filtered information. Systematic reviews and critically appraised topics are filtered, as are critically appraised individual articles, meaning someone has already appraised the primary studies. Randomized trials, cohort and case-control studies are unfiltered, appraised only by the reader.
- Base of the clinical studies pyramid. Expert opinion and background information sit at the very base; one version puts in vitro and animal research below ideas, editorials and opinions, but unappraised opinion is the weakest footing.
USPSTF levels of certainty
Certainty = the likelihood that the assessment of net benefit is correct. Net benefit = benefit minus harm of the service as implemented in a general, primary care population.
| Level | Description |
|---|---|
| High | Consistent results from well-designed, well-conducted studies in representative primary care populations. Conclusion unlikely to be strongly affected by future studies. |
| Moderate | Evidence sufficient to determine effects, but confidence constrained by: number/size/quality of studies · inconsistency across studies · limited generalisability · lack of coherence in the chain of evidence. Magnitude or direction could change. |
| Low | Evidence insufficient to assess effects, because of: limited number or size of studies · important design flaws · inconsistency · gaps in the chain of evidence · findings not generalizable · lack of information on important health outcomes. More information may allow estimation. |
USPSTF grades of recommendation
| Grade | Certainty / benefit | What to do |
|---|---|---|
| A | High certainty, net benefit substantial | Offer or provide this service |
| B | High certainty of moderate benefit, OR moderate certainty of moderate-to-substantial benefit | Offer or provide this service |
| C | At least moderate certainty that net benefit is small | Offer selectively, by professional judgment and patient preference |
| D | Moderate or high certainty of no net benefit, or harms outweigh benefits | Discourage the use of this service |
| I | Evidence insufficient — lacking, poor quality, or conflicting | Read the clinical considerations. If offered, the patient should understand the uncertainty |
Also tested
- USPSTF Grade I. Grade I means the balance of benefits and harms cannot be determined; evidence is lacking, poor or conflicting, so if the service is offered the patient should understand that uncertainty. It does not mean the balance is unfavorable.
- USPSTF Grades A and B. Both carry the same instruction to offer or provide the service, because both rest on a net benefit judged at least moderate. Grade C offers selectively, Grade D discourages, and Grade I calls for the clinical considerations to be read.
Synthesis
- Systematic review — systematically identifying and evaluating multiple studies to answer a specific, focused question.
- Meta-analysis — qualitative or quantitative synthesis; “pools” data from multiple studies to produce a single large study. Greater n usually means greater power.
Also tested
- Meta-analysis versus systematic review. A meta-analysis combines the data rather than only appraising the studies: the review identifies and evaluates the studies, and the meta-analysis pools their data to function as one larger study.
- Pooling data across studies. A greater number of participants usually means greater power, but pooling poor studies produces a larger poor study rather than a reliable answer.
- Meta-analysis. Pooling data produces what amounts to a single large study; a greater number of participants usually means greater power, but pooling weak studies gives a larger weak study.
IO 3 — Judging sources
Quality markers: peer-reviewed · government databases · be cautious of websites.
Also tested
- Named sources of medical information. Pharmaceutical representatives appear among the named sources. Their presence is not an endorsement: the list is deliberately broad, and quality has to be judged separately by peer review and appraisal.
- Quality sources of medical information. Peer review marks a quality source, with government databases a further category. Be cautious of websites, which carry no equivalent guarantee of scrutiny before publication.
- Breadth of information sources. Information reaches clinicians from many channels of unequal quality, running from original research journals and evidence-based databases through experts and guidelines to newsletters, representatives, media and patients.
The ABCD method for web resources
| Letter | Ask |
|---|---|
| Author | Credentials, qualifications, expert in the field? |
| Bias | Is there any? |
| Content | Level of detail |
| Date | Still current? |
The library also provides the CRAAP test.
Also tested
- Fame does not ensure unbiased information. An author's fame alone is the least reassuring sign of a web site's quality, because a well-known expert may still be biased. Look instead for the sponsoring organization, facts, and stated sources.
- ABCD method, A. In the ABCD method for appraising a website, A stands for author, including credentials and expertise. The other three are bias, content (level of detail) and date (whether the material is still current).
- ABCD method, B. B asks about bias. It sits between the author's credentials and the level of detail in the content, so the method moves from who wrote it to why, then to what and when.
- Retracted vaccine study. Publication in a respected journal does not guarantee integrity: the study appeared in a major journal and was eventually retracted, so appraisal of the author, the interests behind the work and the methods still matters.
- Ethical breaches in a retracted study. Improperly obtained blood and unapproved invasive tests: blood was taken from children at a birthday party for payment, and autistic children underwent colonoscopy and lumbar puncture without research review board approval.
IO 4 — Databases and guidelines
| Resource | Note |
|---|---|
| Evidence Based Medicine Reviews | 7 leading EBM resources + Medline + 200+ full text journals |
| Cochrane Library | Systematic reviews |
| Medline / PubMed | US National Library of Medicine — original articles |
| Up-To-Date, TripPro, Clinical Key | Point-of-care and search resources |
| Essential Evidence Plus | Incorporates InfoPOEMS / InfoRetriever |
| ACP Journal Club | American College of Physicians; many specialty journal clubs exist |
| Guidelines: Agency for Healthcare Research and Quality (links to USPSTF) · MMWR Recommendations and Reports (public health policies for prevention & treatment) · Task Force on Community Preventive Services | |
Also tested
- Searching MEDLINE. MEDLINE can be searched by both content and methods terms at once, which is a strength that helps narrow a search. Its limitations are missed articles, incomplete coverage of the world's journals, and publication bias.
- MEDLINE for rare events. Searching MEDLINE and PubMed is indispensable for finding whether a rare event has been reported, but it is too inefficient for most questions at the point of care.
- Practice guidelines. They are produced at the level of the hospital or clinic, nationally, and by professional bodies, such as the Centers for Disease Control and Prevention (national) and the American Academy of Pediatrics (professional).
- Medline search tutorials. Found at the National Library of Medicine and the university library; searching and finding the best evidence is named as one of the challenges of evidence-based practice.
- Database for literature reviews. The Cochrane Database of Systematic Reviews is named as the source for literature reviews, while original articles are found through MEDLINE and PubMed.
Joint decision making and the five steps
Three primary elements of joint clinical decision making: disclosure of risks and benefits of therapeutic alternatives → exploration of patient values about the therapy and potential outcomes → the actual decision.
Steps to evidence-based practice: create a focused clinical question → systematic review of the literature or an evidence-based synopsis from a credible resource → evaluate the validity (truth) of the research → apply to individual patient encounters → evaluate performance or outcome.
Also tested
- Validity step of evidence-based practice. This step assesses whether the research is true. It comes before application, because research whose truth has not been established should not be carried into a patient encounter.
3 · Study Design & Research Methods
Megan B. Finck, MMS, PA-C · Session 2
Instructional Objectives
- Compare and contrast the following study designs: Cross-sectional, Cohort, Case-control, Randomized controlled clinical trials, Systematic reviews, Meta-analysis
- Differentiate between qualitative and quantitative research
- Describe the various research methods utilized in medical research
IO 2 — Research paradigms
| Qualitative | Quantitative | |
|---|---|---|
| Approach | Naturalistic, inductive | Positivistic, deductive |
| Develops | Descriptions | Predictions |
| Data | Observation, field notes, audio recordings; from researchers & participants | Numerical, objective, measurable |
| Examples | Case study, biography, ethnography, phenomenology | Randomized controlled trials, cohort studies, case control studies |
Also tested
- Qualitative research. It is naturalistic and inductive, developing descriptions. Its data are observation, field notes and audio recordings, drawn from both researchers and participants rather than from measurement.
What study design is, and why it matters
The organization and plan for data collection and analysis. It is directly related to the validity and reliability of data:
Also tested
- Why study design matters. Knowing each design's strengths and limits improves evaluation of evidence. Strong design produces strong data, which produces valid and reliable results, which supports better clinical decisions.
- Why study design matters. Knowing each design's strengths and limits improves the evaluation of evidence, which makes the resulting clinical decision more informed.
- Study design chain. Strong design gives strong data, which gives valid results; data gathered under a weak design cannot be rescued by careful analysis afterwards.
The two axes
| Axis | Category | Definition | Examples |
|---|---|---|---|
| Purpose | Descriptive | Describe characteristics of a group of people | Case report, case series |
| Analytic | Compare characteristics in two or more groups to draw inference for a population | Randomized controlled trials, cohort, case control | |
| Control | Experimental | The intervention (independent variable) is assigned and manipulated by the investigator | Randomized controlled trial |
| Observational | The exposure (independent variable) is observed rather than assigned | Cohort, case control |
Cross-sectional appears under both descriptive and analytic, and is always observational.
Also tested
- Case series. A case series is a descriptive design that reports the characteristics of a group of patients, such as those with an uncommon disease seen at a referral center, without a comparison group.
- Observational designs. Cohort and case-control studies are observational because the investigator watches an exposure that occurred independently rather than allocating it; the randomized trial is the experimental example.
- Analytic designs. Randomized trials, cohort and case-control studies are examples of analytic work; what unites them is comparison, each setting one group against another, which supports an inference about a population.
IO 3 — Other research methods
| Method | What it does | Class |
|---|---|---|
| Longitudinal | Measures and reports changes over time | Observational |
| Survey | Overview of the research question via purpose-built questions | Observational |
| Policy research / evaluation | Determines if a program works or a policy is helpful | Observational |
| Experimental | Researcher controls treatment and assignment to groups | Experimental |
| Quasi-experimental | Controls treatment but not assignment — less control | Between |
| Field research | Common in epidemiology during investigations (new screening tests, vaccine trials) | Either |
The goal of epidemiologic and clinical research
To permit an unbiased comparison between a group with the factor and a group without it. Good design allows: comparisons between two or more groups · quantification of the risk difference · determination of temporal sequence · minimization of biases.
Typical study sequence: Clinical observations → Available data → Case-control studies → Cohort studies → Randomized trials.
Trajectory toward EBM: Ideas (from biology, epidemiology, clinical observation, imagination, reasoning) → Hypotheses → Observational studies → Experimental studies → Evidence-based information.
Also tested
- Temporal sequence. A cohort study, starting from exposure, establishes it most directly because the exposure is recorded before the outcome occurs. A cross-sectional design measuring both at once cannot do this.
- Good study design. Beyond comparing groups, it allows quantifying the risk difference and determining temporal sequence, which bears on causation.
- Cross-sectional design. A cross-sectional design is for generating hypotheses; it is observational, and because risk factor and disease are measured at the same moment, it cannot establish which came first.
IO 1 — The designs in detail
Cross-sectional — observational, for generating hypotheses
- Cross-sectional survey — data at a single point in time; risk factor and disease measured at the same time; individual level data. Examples: obesity in a population at a point in time; colorectal cancer screening by fecal occult blood; varicella titres for population immunity.
- Cross-sectional ecologic — relates the frequency of a characteristic to an outcome; population level data. Example: the 1996 Atlanta Olympics — streets blocked off, less car emission, fewer emergency visits for childhood asthma.
- Longitudinal ecologic survey — ongoing surveillance or frequent cross-sectional surveys; measures trends. Example: as smoking rises among women, does lung cancer incidence rise?
Also tested
- Cross-sectional ecologic study. It uses population rather than individual level data, relating the frequency of a characteristic to an outcome of interest across populations rather than within individuals.
- Cross-sectional survey examples. Varicella titres showing what proportion of a population is immune, determining obesity in a population at a point in time, and colorectal cancer screening by fecal occult blood testing.
- Longitudinal ecologic survey. Ongoing surveillance or repeated surveys measuring trends; it may reveal an association, such as whether lung cancer incidence rises as smoking increases among women over time.
Case-control — observational, for generating or testing hypotheses
Groups defined by outcome (disease or no disease), then assessed for a difference in exposure or risk factors. It looks backward: from outcome to cause.
Also tested
- Case-control study. A case-control study is observational and can be used for generating or testing hypotheses, unlike the cross-sectional design, which generates hypotheses only.
Cohort — observational, for testing hypotheses
A cohort is a group of people with a common characteristic assembled at a similar point and followed over time. The characteristic may be age, date of birth, exposure to a risk factor, occupation, religious affiliation, having a specific disease, or participation in an intervention.
- Subjects grouped on the basis of exposure rather than outcome.
- The comparison group should be as similar as possible except for the exposure.
- Must be followed long enough for the expected outcome to appear, and completely, so attrition does not falsely skew the outcome data.
- Prospective (concurrent) — assemble groups in the present by exposure, collect baseline data, follow forward.
- Retrospective (historical) — go back in history to define a risk group, then follow its members forward to the present.
Also tested
- Withdrawals and drop-outs. A randomized trial must account for them because losing participants can distort the outcome. The same concern requires complete follow-up of a cohort, since participants who leave may differ from those who stay.
- Retrospective cohort study. The investigator identifies a past cohort, exposures are determined at that historical point, and the group is followed forward to the present to determine outcome. It is also called a historical cohort study.
- Cohort versus case-control study. Cohort groups by exposure, case-control groups by outcome; the cohort looks forward from exposure to outcome, the case-control backward from outcome to exposure.
- Randomized controlled trial design. A randomized controlled trial is experimental, for testing hypotheses, because the investigator assigns and manipulates the intervention rather than observing an exposure that arose on its own.
Randomized controlled trial — experimental, for testing hypotheses
- Clearly formed hypothesis · development of a study protocol
- Results must be applicable to a larger population
- Subjects randomly assigned to treatment and control
- Can be double or single blinded · needs adequate sample size
- Should account for withdrawals and drop-outs
Also tested
- Applicability of trial results. Results must be applicable to a larger population because a finding confined to the sample cannot guide practice. This is external validity; adequate sample size and a clear protocol are listed alongside it.
- Requirements of a randomized controlled trial. It needs a clear hypothesis, a protocol and random assignment, plus adequate sample size, results applicable to a larger population, and an accounting for withdrawals and drop-outs.
Systematic review and meta-analysis
As in section 2. The review identifies and evaluates; the meta-analysis pools the data into a single larger study.
4 · Rates & Disease Measurement
Megan B. Finck, MMS, PA-C · Session 3 · Fletcher ch. 2–3
Instructional Objectives
- Compare and contrast incidence and prevalence rates
- Describe considerations for disease surveillance and outbreak investigations
- Describe strategies to determine a course of action in the face of new disease
- Identify and describe ways to measure disease
- Discuss factors that impact or influence disease measurement
- Discuss the concepts of rate and rate determination
Rates
A rate is the number of times an event occurs during a fixed time period. Rates allow comparisons — that is why they exist rather than raw counts.
| Measure | Numerator | Denominator | ×10n |
|---|---|---|---|
| Crude death rate | All deaths in the interval | Mid-interval population | 1,000 or 100,000 |
| Cause-specific death rate | Deaths from a specific cause | Mid-interval population | 100,000 |
| Proportionate mortality | Deaths from a specific cause | Total deaths, all causes | 100 or 1,000 |
| Death-to-case ratio | Deaths from a specific cause | New cases of the same disease | 100 |
| Neonatal mortality rate | Deaths < 28 days | Live births | 1,000 |
| Postneonatal mortality rate | Deaths 28–364 days | Live births | 1,000 |
| Infant mortality rate | Deaths < 1 year | Live births | 1,000 |
| Maternal mortality rate | Deaths from pregnancy-related causes | Live births | 100,000 |
Age-specific mortality is limited to a particular age group. Infant mortality is commonly used to compare health status between nations. Age-adjusted rates remove the effect of a population's age structure — the US crude death rate of about 794 per 100,000 in 2009 corresponded to an age-adjusted 741.
Case fatality = a measure of disease severity: the proportion diagnosed who die from it. Years of potential life lost = death at a younger age loses more productive years; used to target interventions.
Also tested
- A raw count is not a rate. A rate requires a denominator, so a raw count such as 800,000 deaths in a year is not an incidence or prevalence rate.
- Age-adjusted death rate. It is the right measure for comparing a county with its state because it removes differences in age structure; on the crude rate alone, a county with an older population would appear less healthy.
- Proportionate mortality. The causes must sum to one hundred percent, because each cause is expressed as a share of all deaths in the population over the period.
- Maternal and child health rates. The rates named as specific to maternal and child health are live birth rates and fetal, infant and maternal death rates, with neonatal and perinatal death rates alongside them; all take live births rather than population as their denominator.
- Morbidity rates. Incidence and prevalence rates are the morbidity rates, concerning illness in the living; the mortality group covers crude, age-specific, cause-specific, adjusted and case fatality rates.
IO 1 — Incidence and prevalence
| Prevalence | Incidence | |
|---|---|---|
| Measures | Portion of a group with a condition at a point in time — a snapshot | Portion who develop a condition in a period — new cases in previously healthy people |
| Collected by | Counting those with and without; surveys are useful | Following a population over time |
| Subtypes | Point (single point per person, not necessarily the same date) · Period (cases present during a time period) | Cumulative incidence (fixed-size group) · Incidence density (changing population) |
| Study name | Prevalence study = cross-sectional | Incidence study = cohort |
| Example study | SchoolNuts — food allergy in 10–14 year olds; ~1 in 20 Melbourne children | Framingham Heart Study — began 1948, multiple generations, identified cardiovascular risk factors |
Incidence density is used in dynamic populations and often in clinical trials, where people move in and out. Its denominator is person-time: one person followed 10 years contributes 10 person-years. Error can arise from variation in duration of follow-up.
Why prevalence matters clinically: it prioritizes the differential (horses, not zebras) and it sets the pretest probability, which changes how a test result should be read.
Also tested
- Prevalence and follow-up. A prevalence (cross-sectional) study assesses a sample of a defined population for the presence of a condition at one point in time, so it involves no follow-up period.
- Average duration of disease. When incidence and prevalence are stable, average duration of disease equals prevalence divided by incidence, so a prevalence of 1.5% with an incidence of 50 per 100,000 per year gives 30 years.
- What a prevalence study describes. A prevalence (cross-sectional) study assesses a defined population for a condition and reports the proportion that has it, which is why it helps plan health services.
- Cumulative incidence versus incidence density. Cumulative incidence uses a group of fixed size, whereas incidence density measures new cases in a changing population, so its denominator is person-time rather than people.
- Prevalence studies. They are good for generating hypotheses and planning services but poor for showing cause and effect. They are also relatively inexpensive.
- Incidence example. In a town of 1,000 with 80 cases in 2015, 20 previously healthy people developing disease between 2016 and 2018 give an incidence of two percent; the 100 cases present in 2018 give a prevalence of ten percent.
- Incidence. New cases arising in previously healthy people over a specified time period; prevalence, by contrast, counts everyone who has the condition.
IO 5 — What influences measurement
- Cases on a continuum — most clinical information is not binary. The case definition has a marked impact on measured frequency.
- Population at risk — should include only those susceptible (men are not in a study of ovarian cancer). Limiting the population may curtail generalisability.
Also tested
- Observed incidence. Case definition and detection efforts change the observed frequency of disease, whereas a larger sample only reduces random error in estimating the incidence and does not systematically raise it.
- Case definition. Most clinical information lies on a continuum, not in two boxes, so where the line is drawn between case and non-case has a marked effect on how much disease the measurement finds.
Sampling
| Sample | Definition |
|---|---|
| Random | Each individual has an equal chance of selection |
| Probability | Each has a known probability, which may or may not be equal — useful for investigating subgroups (age, ethnicity) |
| Convenience | Non-random, convenient to obtain |
| Grab | Researchers “grab” whoever they can find |
Also tested
- Probability sample. A probability sample (each person has a known probability of selection, not necessarily equal) represents the population sampled only if enough people are sampled.
- Convenience sample. A sample chosen by a convenient rule with no random element, such as enrolling every 10th admitted patient, is a convenience sample; because selection is not by chance, it may misrepresent the population.
- Probability sample. Useful when subgroups such as age or ethnicity must be investigated, because a random sample of the whole population may not deliver enough of a small subgroup.
- Random sample. Every individual has an equal chance of selection; in a probability sample the chance of selection is known but need not be equal.
- Grab sample. Researchers take whoever they can find. It is a non-random approach, alongside the convenience sample, where availability decides who ends up in the study.
Distribution of disease — person, place, time
| Level | What it asks |
|---|---|
| Person | The profile of those affected — demographics and behavior: where they eat, shop, what products they use, their drinking water, whether they attended the same event |
| Place | Endemic — disease limited by geographic location |
| Time | Epidemic — concentration of new cases in a given time (awareness makes diagnosis of new cases easier). Pandemic — widespread cases; may be rapid (influenza, early 1900s) or slow (HIV/AIDS) |
Also tested
- Endemic disease. Endemic refers to disease limited by geographic location; coccidioidomycosis is confined to the deserts of the southwestern United States and Mexico.
- Distribution by person. It involves profiling those affected, including behavior as well as demographics: where they eat, where they shop, what products they use, their drinking water, and whether they attended the same event.
- Awareness of an epidemic. It can make new cases easier to diagnose, as with early COVID-19, so measured frequency can rise partly because clinicians started looking.
- Levels of disease distribution. The distribution of disease is evaluated by person, place and time, and looking for patterns across those three is a large part of what epidemiologists do.
IO 2 — Surveillance
May be active or passive. Helps establish a baseline. Operates at national, state and county level.
Also tested
- Surveillance baseline. Without a baseline, a rise in cases cannot be recognized. Establishing it is among surveillance's core purposes, alongside determining the magnitude of a health problem.
Epidemiologic functions
- Determine magnitude of a health problem
- Understand the natural history of a disease
- Detect outbreaks/epidemics — foodborne gastrointestinal disease; Vibrio after Hurricane Katrina floods
- Document geographic distribution and spread — tuberculosis on an international flight
- Allow hypothesis testing for etiology — via case control or cohort studies
- Evaluate control strategies
- Monitor changes in an infectious agent — emerging antibiotic resistance
Also tested
- Surveillance and etiology. Case-control or cohort studies allow surveillance to support hypothesis testing for etiology: the data identify a question that a comparative design is then built to answer.
- Monitoring changes in an infectious agent. An example is resistance emerging to an antibiotic it was once susceptible to.
- Surveillance. One epidemiologic function is detecting outbreaks and epidemics, alongside determining the magnitude of a problem, understanding natural history, documenting geographic spread, testing etiologic hypotheses, evaluating control strategies and monitoring changes in an agent.
Strategies
- Screening high-risk persons or groups — influenza in a nursing home resident or worker; restaurant worker with hepatitis A
- Clinical symptoms/syndrome — ask about persons with similar problems
- Mortality — some cases are detected only at death
- Provider reporting of morbidity — allows tracking of contacts and exposures
- Laboratory reporting of positive results
Also tested
- Provider reporting of morbidity. It allows tracking of possible contacts and exposures.
IO 3 — Investigating an outbreak
- Determine epidemic vs endemic
- Determine the attack rate
- Establish the diagnosis
- Establish the case definition
- Determine if an epidemic truly exists
- Characterize by time, place and person
- Characteristics: type of exposure · route of spread · timing of exposure
- Develop an epi curve — number of cases plotted against time
- Develop hypotheses about source, type and route of spread
- Test hypotheses
- Initiate control measures
The investigation arc: detect possible outbreak → find cases → generate hypotheses → test hypotheses → solve original source → control outbreak → decide the outbreak is over.
Classic epi curve: Legionnaires' disease by date of onset, Philadelphia, 1 July – 18 August 1976, following an American Legion convention.
Also tested
- Case-control study in an outbreak. Comparing the exposures of ill and well people, a case-control study is especially useful for identifying the mode of transmission, how the disease is being spread, which allows the outbreak to be controlled and the source traced.
- Confirming an epidemic exists. An apparent rise may reflect detection rather than occurrence, so a baseline and a settled case definition come first: both awareness and definition change how many cases are found.
- 1976 Philadelphia Legionnaires' disease outbreak. It illustrates an epidemic curve tied to a single gathering: cases were plotted by date of onset across July and August, following an American Legion convention.
- Starting an outbreak investigation. Determine whether the situation is epidemic or endemic, alongside the attack rate, the diagnosis and case definition, confirming that an epidemic truly exists, and characterizing it by time, place and person.
- Outbreak investigation sequence. Initiating control measures follows the testing of hypotheses; the sequence then runs on to solving the original source, controlling the outbreak and deciding that it is over.
5 · Data, Validity & Variation
Megan B. Finck, MMS, PA-C · Session 3 · Fletcher ch. 2–3
Instructional Objectives
- Discuss the types of data and their relationship to medical research
- Identify different levels of measurement in data
- Discuss the concepts of validity and reliability and their relationship to measurement
- Discuss variation, distribution, and dispersion in data
- Differentiate between the types of variation in disease measurement
- Explain the implications of normal versus skewed distribution
- Discuss the limitations of differentiating between normal and abnormal
IO 2 — Levels of measurement
| Type | Definition | Examples |
|---|---|---|
| Nominal | Categories without specified order | Blood group (ABO), eye color |
| ↳ Dichotomous | Nominal divided into two categories | Yes/no; present/absent |
| Ordinal | Some order, but no defined intervals between categories. You can say more/less/equally desirable — not by how much | Cancer staging I–IV; pitting edema 1+ to 4+; patient status; drug birth-defect risk categories A–D, X |
| Interval | Order with defined, equal intervals | — |
| ↳ Continuous | Any value in a continuum, limited by the precision of the measurement | Blood pressure, weight |
| ↳ Discrete | Only specific values | Number of children, number of heart attacks |
Also tested
- Level of measurement. It determines what can legitimately be said about two values: nominal categories can only be counted, ordinal values can be ranked, and only interval data supports statements about how much two values differ.
- Continuous versus discrete data. Continuous data can take any value in a range, though limited by the precision of the measurement; discrete data takes only specific values such as a number of children.
- Ordinal values. Two ordinal values can be compared as more, less or equally desirable, but not by how much, because the intervals between adjacent categories are not defined or equal.
IO 3 — Validity and reliability
| Validity (accuracy) | Reliability (precision) |
|---|---|
| The degree to which data measure what they are supposed to measure | Consistency of measurement — repeated attempts produce similar results |
| Ensured by calibrating devices and comparing results against known values; for subjective measures, by structured questions (CAGE questionnaires, pain scales) | Also called reproducibility, repeatability, precision |
Also tested
- Reliability without validity. A measurement can be reliable without being valid, because it can be consistently wrong. Low validity with high reliability is the combination to watch for, since repeated agreement looks like confirmation.
- Low validity, high reliability. Results are tightly clustered, all away from the true value. This is the most dangerous combination, because consistency is easily mistaken for correctness when nothing is checked against a known standard.
Validity without physical verification
| Type | Definition | Example (depression screening) |
|---|---|---|
| Content | Includes all dimensions of the relevant information while excluding irrelevant information | Are the questions representative of accepted symptoms of clinical depression? (sad or blue > 2 weeks, irritable, difficulty sleeping) |
| Criterion | The measurements predict / are associated with an observable response or event | Does a high scorer have the affect of a clinically depressed person? Does chronically high cholesterol come with xanthelasma? |
| Construct | Related to other measures believed to be part of the same phenomenon | Would a PA interviewing the patient reach the same determination as the screening test? |
Also tested
- Construct validity. A measure has construct validity when it relates to others believed part of the same phenomenon, as in whether a clinician interviewing a patient for depression reaches the same determination as a screening test.
- Content validity. The measure covers all relevant dimensions and excludes irrelevant ones.
IO 5 — Types of variation
| Type | Definition | Example |
|---|---|---|
| Intrasubject | Within one person | Blood pressure varying by biological variation, time of day, position, stress |
| Intraobserver | One observer, two occasions | A radiologist reading the same radiograph differently at two times |
| Interobserver | Two observers | One clinician grades a murmur 3, another grades it 4 |
Fluctuation reflects two phenomena combined: biologic differences and variation in measurement.
| Source | Definition |
|---|---|
| True biologic variation | The sum of many unknown factors, each contributing a small random effect. Consecutive systolic pressures under the same conditions will not be identical. |
| Systematic variation | Occurs when the conditions of measurement are known to affect the values — time of day, after a cigarette, patient position, white coat phenomenon. |
| Random measurement error | Due to chance; equal probability of being above or below the true value. → lack of reliability |
| Systematic measurement error | A flaw in the process. → lack of validity. Determines a range for the true score. |
Inter-patient variation may be due to inherent biological differences, systematic differences in the conditions of measurement, and/or measurement error. Answers to clinical questions about individuals are most often derived from groups of patients with similar conditions.
Also tested
- Inter-patient variation. Variation among different patients (inter-patient variation) does not explain differences between repeated measurements on one patient, which arise from biologic variation, measurement error, observer differences or true biologic change.
- Random and systematic error. Random error (chance) has equal probability of being above or below the true value, and variation between observers also spreads readings; a consistent shift in one direction, such as staff recording abnormal fetal heart rates as more normal, is systematic error (bias).
- Intraobserver variation. One observer interprets the same thing differently twice, such as a radiologist reviewing the same radiograph at two different times and reaching different interpretations.
- Systematic variation. Known conditions of measurement affect the values obtained, for example blood pressure varying by time of day, after a cigarette, by patient position, or through the white coat phenomenon.
IO 4 & 6 — Distribution and dispersion
Frequency distributions portray the quantity of each category, class or interval — the number or proportion of people with different values. Usable on nominal and on interval data (more commonly the latter in clinical studies). Described by central tendency (the middle) and dispersion (the spread).
| Central tendency | Definition |
|---|---|
| Mean | Best estimate of central tendency; sum of all values ÷ number of observations. Can be affected by extreme outliers. |
| Median | The middle score — half above, half below |
| Mode | The most frequently occurring value |
| Dispersion | Definition |
|---|---|
| Range | All values from lowest to highest |
| Interquartile range | A limited range index — e.g. data between the 25th and 75th percentiles |
| Standard deviation | Variability of the scores about the mean |
Also tested
- The normal curve is a model. The normal distribution is a mathematical model that naturally occurring clinical measurements may or may not resemble, so it does not describe the distribution of most naturally occurring phenomena.
- Right-skewed distribution. In a right-skewed distribution with a long right tail, the mode is smallest, the median is intermediate and the mean is largest, because extreme high values pull the mean toward the tail.
- Mean. The sum of all values divided by the number of observations; it is the best estimate of central tendency, though extreme outliers can pull it away from the bulk of the data.
- Normal distribution. Also known as the Gaussian distribution or bell curve, a symmetric shape in which mean, median and mode coincide.
IO 7 — Normal versus abnormal
There is no clear-cut point at which pathology or disease exists. Disease is often acquired by degrees:
Susceptibility → Pre-symptomatic → Clinical Disease → Dysfunction → Death
- Statistical abnormality does not necessarily mean clinical abnormality. A serum cholesterol of 201 is above “desired” but not in itself clinically significant — though being high-normal, the risk is higher than at low-normal.
- You can have an abnormal finding and not be sick — a sickle cell trait carrier.
- Value extremes may be desirable — the lower end of the systolic blood pressure range; the higher end of the high-density lipoprotein range.
- A normal test does not necessarily rule out disease.
Also tested
- Sickle cell trait carrier. An abnormal finding can exist without the person being sick; the converse is that a normal test does not necessarily rule disease out.
Defining abnormality
- Determine it in the context of values known to be associated with disease, disability or death
- A clinically recognizable difference from healthy
- Abnormal tests should be confirmed by repeating the test — regression to the mean
- Cutoff points for normal/abnormal are often arbitrary
- Be careful of labeling.
Also tested
- Abnormal is not defined by shape. Abnormality is commonly defined by statistically unusual values or by levels associated with disease, disability, death, symptoms or treatment benefit, but not by whether values fit a normal distribution, since clinical distributions are not necessarily normal.
- Defining abnormality. Be careful of labeling, because cutoffs are often arbitrary and a statistically abnormal value need not mean the patient is unwell.
6 · Diagnostic Tests
Megan B. Finck, MMS, PA-C · Session 4 · Fletcher ch. 4 + the 2×2 / likelihood ratio reading
Instructional Objectives
- Identify and describe limitations of diagnostic tests
- Discuss accuracy of diagnostic tests
- Describe the following: Sensitivity, Specificity, Positive predictive value, Negative predictive value, Pretest odds, Posttest odds, Likelihood ratios, Parallel testing, Serial testing
Before any test is ordered
- Common diseases occur commonly — horses, not zebras
- Uncommon presentations of common disease occur more frequently than common presentations of uncommon disease
- No disease is rare to the patient who has it
Considerations for ordering: avoid ordering tests indiscriminately · risk : benefit ratio · know the limitations of the test · avoid repeating a test without an indication · consider cost.
Diagnostic tests are for symptomatic individuals; screening tests for asymptomatic ones.
Also tested
- Requirement for diagnosis. Diagnosing disease requires a reliable method of detecting its presence or absence, because the task is to identify who has the disease and who does not.
- Ordering a diagnostic test. Weigh risk against benefit, the test's limits, and its cost; avoid ordering tests indiscriminately and avoid repeating a test without an indication.
The gold standard
Any diagnostic procedure believed to identify the disease with certainty: culture for streptococcal throat infection · biopsy for cancer · cardiac catheterization for coronary artery disease · endoscopy for peptic ulcer disease.
Limitations: often expensive · often invasive or risky · may not be acceptable to patients · may not be easily completed. So alternative tests are often the initial choice — and are judged by how they compare to the gold standard.
The 2×2 table
| Disease (per gold standard) | No disease (per gold standard) | Total | |
|---|---|---|---|
| New test positive | True positives (TP) — a | False positives (FP) — b | All test positives TP + FP |
| New test negative | False negatives (FN) — c | True negatives (TN) — d | All test negatives FN + TN |
| Total | All with disease TP + FN | All without disease TN + FP | All subjects |
| Measure | Meaning | Formula | Lettered |
|---|---|---|---|
| Sensitivity (Sn) | Probability a person with disease has a positive result | TP / (TP + FN) | a / (a + c) |
| Specificity (Sp) | Probability a person without disease has a negative result | TN / (TN + FP) | d / (b + d) |
| Positive predictive value (PPV) | Probability a person with a positive result has disease | TP / (TP + FP) | a / (a + b) |
| Negative predictive value (NPV) | Probability a person with a negative result is disease free | TN / (TN + FN) | d / (c + d) |
| Prevalence (pretest probability) | Probability of disease in the population of interest | — | (a + c) / (a + b + c + d) |
Also tested
- Prevalence in a study. Prevalence is the proportion of the studied population that has the disease, and the test result does not affect it. Here 95 patients with sinusitis out of 247 total is 38%.
- Validity of a test. Validity is a test's ability to determine correctly who has the disease, measured by sensitivity and specificity. Reliability, by contrast, is getting a similar result when the test is repeated under the same conditions.
- Prevalence in the 2×2 table. Prevalence is a plus c, over all four cells: everyone with the disease over everyone tested. It is also the pre-test probability in the population of interest.
- Prevalence in reasoning. Prevalence prioritizes the differential diagnosis (think horses, not zebras). It also affects the interpretation of test results, because it sets the pretest probability.
- False positive. In a 2x2 table, a false positive is when the new test is positive and the gold standard negative; a highly specific test produces few of them, so a positive on such a test rules disease in.
Worked example — HUGIT vs endoscopy for peptic ulcer disease
10,000 patients had both. 25 ulcers by the gold standard, 20 correctly identified by the new test. The new test also suggested ulcer in 80 with a negative endoscopy.
| Disease | No disease | Total | |
|---|---|---|---|
| HUGIT positive | 20 TP | 80 FP | 100 |
| HUGIT negative | 5 FN | 9,895 TN | 9,900 |
| Total | 25 | 9,975 | 10,000 |
Sensitivity = 20/(20+5) = 0.8 = 80%. Specificity = 9,895/9,975 = 0.99 = 99%. Prevalence = 25/10,000 = 0.25% (low) — which is why the 80 false positives outnumber the 20 true positives.
Also tested
- Specificity in the ulcer-testing example. 9,895 of 9,975 patients without ulcers tested negative, a specificity of 99 percent; the 80 false positives still outnumber the 20 true positives, which wrecks the positive predictive value despite this high specificity.
SnOUT and SpIN
Sn-OUT — a negative result on a highly sensitive test rules disease OUT. High sensitivity → very few false negatives. Matters when the consequences of missing the disease are serious.
Sp-IN — a positive result on a highly specific test rules disease IN. High specificity → very few false positives. Matters when false positives would go on to unnecessary, risky or expensive work-ups.
Sensitivity and specificity are inversely related. Ideally maximize both; the receiver operating characteristic (ROC) curve plots true positive rate against false positive rate (1 − specificity) and lets clinicians compare alternative tests for the same diagnosis. The better test's curve lies closest to the upper left and has the greatest area beneath it. ROC curves also help determine where to set the cutoff point. To avoid the trade-off altogether, use serial testing.
Screening strategy: screen with a highly sensitive test and confirm with a highly specific one.
Also tested
- Receiver operating characteristic curve. It is used for comparing alternative tests for the same diagnosis, since it displays the whole trade-off between true and false positives rather than a single pair of values.
Predictive values and prevalence
PPV and NPV are the numbers that matter in the exam room: “if PPV = 85%, my patient with a positive result has an 85% chance of really having the disease.”
| Same test throughout: sensitivity 99%, specificity 95%, population 10,000 | |||
|---|---|---|---|
| Prevalence | With disease | PPV | NPV |
| 1% | 100 | 99/594 = 17% | 9,405/9,406 = 99.9% |
| 5% | 500 | 495/970 = 51% | 9,025/9,030 = 97% |
Also tested
- Positive predictive value. Positive predictive value is the probability that a person with a positive test result really has disease: true positives divided by all positives (true positives + false positives). For facial pain, 49/128 = 38%.
- Negative predictive value. Negative predictive value is the probability that a person with a negative test result really is disease free: true negatives divided by all negatives (true negatives + false negatives). For no facial pain, 73/119 = 61%.
- Predictive value and false positives. Positive predictive value is true positives divided by all positives. A test with higher sensitivity but more false positives has lower specificity, and at a given prevalence its positive predictive value is lower; its negative predictive value is better.
- Predictive value and prevalence. With sensitivity of 99 percent and specificity of 95 percent, positive predictive value rises from 17 to 51 percent as prevalence rises from one to five percent, so predictive value depends heavily on prevalence.
- Why prevalence affects predictive values. Predictive values include both diseased and healthy people: each reads across a test-result row, which mixes the two disease columns in a proportion set by how common the disease is.
- False positives at low prevalence. There are usually more false positives, because the pool of healthy people is so much larger that even a small false positive rate produces many relative to the few true cases.
- Positive predictive value. The probability that someone testing positive has the disease, calculated as true positives over true positives plus false positives.
Pretest and posttest probability
- Pretest probability — the chance the patient has the disease based on what is known before the test; often the prevalence.
- Posttest probability — the chance after the test. This is what informs treatment decisions.
Also tested
- Likelihood ratio of about 1.0. Posttest probability is the pretest probability updated by the likelihood ratio. Facial pain has a positive likelihood ratio of about 1.0, so the probability of sinusitis stays at 38%, equal to the prevalence and the positive predictive value.
- Pre-test probability. The chance of disease before the test is done, often simply the prevalence of the disease in the population the patient comes from.
Likelihood ratios
How much the probability that a patient has a condition changes once the test result is available. Not affected by prevalence.
| Ratio | Formula | Reading |
|---|---|---|
| LR+ | Sensitivity / (1 − specificity) | How much more likely the condition is with a positive result. LR+ of 4 → 4× more likely than the pretest probability |
| LR− | (1 − sensitivity) / specificity | How much less likely with a negative result. LR− of 0.5 → half as likely |
| LR | Meaning | Rule of thumb (added to pretest probability) |
|---|---|---|
| > 10 | Strong evidence to rule in | LR 10 → +45% |
| 5–10 | Moderate evidence to rule in | LR 5 → +30% |
| 0.5–2 | Indeterminate | LR 2 → +15% |
| 0.2–0.5 | Weak evidence to rule out | LR 0.5 → −15% |
| 0.1–0.2 | Moderate evidence to rule out | LR 0.2 → −30% |
| < 0.1 | Strong evidence to rule out | LR 0.1 → −45% |
Getting to posttest probability. Mathematically: convert pretest probability → pretest odds; multiply by the LR → posttest odds; convert back to posttest probability. Or use the nomogram: find the pretest probability in the left column, the LR in the center, draw a straight line through both — the posttest probability is where the line crosses the right-hand column.
- LR > 1 → posttest probability higher than pretest. LR < 1 → lower.
- A very low LR (below ~0.1) virtually rules out the disease.
- When pretest probability is 30–70%, a very high LR (above ~10) rules disease in.
Also tested
- High probability assignment. A “high probability” assignment by clinicians has a likelihood ratio of 4.7. Multiplying pretest odds 0.61 by 4.7 gives posttest odds of about 2.9, and 2.9/(1 + 2.9) gives a probability of about 74%, roughly 75%.
- Low probability assignment. A “low probability” assignment by clinicians has a likelihood ratio of 0.4, weak evidence to rule out disease. Pretest odds 0.61 x 0.4 = 0.24 lowers the probability from 38% to about 20%.
- Sensitivity and the negative test. Higher sensitivity lowers the negative likelihood ratio, so the probability of disease after a negative test decreases. Lower specificity adds false positives and lowers the positive predictive value.
- Indeterminate likelihood ratio. A likelihood ratio from 0.5 to 2 is indeterminate, so an intermediate impression of 1.4 moves a 38% pretest probability only to about 45% (odds 0.61 x 1.4 = 0.86, or 46%).
- Meaning of sensitivity values. A sensitivity of 0.5 means an equal chance of identifying a diseased patient. A value of one identifies every patient with the disease, and a value of zero identifies none of them.
- Positive likelihood ratio of 4. A positive result makes the disease four times more likely, relative to the pre-test probability; the ratio is a multiplier rather than a probability in its own right.
- Likelihood ratio formulas. The positive ratio is sensitivity divided by one minus specificity (positive result in disease over positive result without it). The negative ratio is one minus sensitivity, divided by specificity (negative result in disease over negative result without it).
- Likelihood ratio of 0.1. As a rule of thumb, it lowers the probability of disease by about 45 percentage points, mirroring the 45 point rise given by a ratio of 10; below 0.1 counts as strong evidence to rule the disease out.
- Likelihood ratio of 10. As a rule of thumb, it raises the probability of disease by about 45 percentage points; the shift is added to the pre-test probability rather than replacing it, so the starting point still matters.
- Probability. A probability lies between zero and one, with zero meaning the event cannot occur and one meaning it certainly will; all the probabilities of events in a sample must sum to one.
- Likelihood ratio effect. A ratio greater than one raises the post-test probability above the pre-test value, and a ratio below one lowers it; a very low ratio, below about 0.1, virtually rules the disease out.
Worked example — the 8-year-old with a sore throat
Oral temperature 101°F, enlarged tender anterior cervical nodes, exudative tonsils. Rapid strep antigen test: sensitivity 65%, specificity 96%. Throat culture is the gold standard but takes 24–48 hours.
1. Pretest probability — modified Centor score:
| Criterion | Points |
|---|---|
| Temperature > 38°C | 1 |
| Absence of cough | 1 |
| Tender anterior cervical adenopathy | 1 |
| Tonsillar swelling or exudate | 1 |
| Age 3–14 / 15–44 / > 45 | 1 / 0 / −1 |
| Score | 0 | 1 | 2 | 3 | ≥ 4 |
|---|---|---|---|---|---|
| Risk of streptococcal infection | 2–3% | 4–6% | 10–12% | 23–28% | 38–63% |
2. Calculate the likelihood ratios:
LR+ = 0.65 / (1 − 0.96) = 0.65 / 0.04 = 16.25 — a positive result makes strep 16× more likely.
LR− = (1 − 0.65) / 0.96 = 0.35 / 0.96 = 0.36 — a negative result makes it about ⅓ as likely.
3. Read the posttest probability off the nomogram: ~93% with a positive result, ~22% with a negative one.
Using multiple tests
If multiple tests all agree, interpretation is easy. If they disagree, consider how they were conducted.
| Parallel testing | Serial testing | |
|---|---|---|
| Timing | All tests conducted at once | One after another |
| Used when | Rapid diagnosis is preferred | Rapid assessment is less critical; tests are expensive and/or risky; after simpler and safer tests are inconclusive |
| Effect | ↑ sensitivity and NPV (disease less likely to be missed); ↓ specificity and PPV | ↑ specificity and PPV; ↓ sensitivity and NPV |
| Note | Degree of sensitivity gain depends on how well the tests complement each other — one detecting cases the other misses | Less utilization of laboratory tests, but takes more time. Particularly useful when individual tests are not highly specific |
Also tested
- Independent information from each test. When diagnostic tests are used in parallel or series, each test should contribute information independently. Combining tests helps only if each adds information that the others do not already provide, detecting cases the other test misses.
- Serial testing. Serial testing requires a positive result on the earlier test before the next is done, which increases specificity and positive predictive value (fewer false positives) but lowers sensitivity and negative predictive value.
- Disagreeing tests. When two tests disagree, consider whether they ran in parallel or serially, because the two strategies do different things to the combined sensitivity and specificity.
- Serial versus parallel testing. Serial testing uses fewer tests but takes more time, so the choice comes down to how urgent the diagnosis is against how costly or risky the tests are.
- Serial testing. Used when speed matters less, or tests are risky; it follows after simpler and safer tests have proved inconclusive, and costly tests are another reason.
- Parallel testing. How much it improves sensitivity depends on how well each test catches what the other misses: two tests that fail on the same patients add little, whereas complementary tests catch far more cases.
IO 1 — Questions to ask about a new test
- What is the disease of interest?
- Did the study include a spectrum of patients?
- What is the new diagnostic test?
- Did every subject undergo both the new test and the gold standard?
- Were those conducting the tests blinded to the results of the other test?
- Benefits and risks of the new test — is it reasonable?
- Sensitivity, specificity, PPV, NPV?
- Is the prevalence of disease similar in your patient population?
Also tested
- Blinding in test comparison. Those conducting each test should be blinded to the other's result because one result could influence the reading of the other, a form of measurement bias.
- Prevalence in test studies. Predictive values will differ if prevalence differs, so a published value may not transfer to your setting; check whether a study's prevalence matches your own population.
- Limit on diagnosis. The diagnosis of disease is limited by how good the tests used to measure it are. Sensitivity and specificity guide test selection, while predictive values describe what a given result means.
7 · Risk & Ratios
Megan B. Finck, MMS, PA-C · Session 7 · Fletcher ch. 6-7
Instructional Objectives
- Describe the importance of risk and risk factors in studies of disease
- Interpret relative risks
- Discuss methods of predicting risk
- Describe cohort studies and the advantages and limitations of this type of study design in medical research
- Discuss methods of comparing risk
- Explain confounding variables and methods for controlling confounding
- Describe effect modification
- Describe case control studies and the advantages and limitations of this type of study design in medical research
- Interpret relative risks and odds ratios
IO 1 — What a risk factor is
A risk factor is a factor associated with an increase in the probability of developing the disease later. It raises the probability; it does not make the disease certain, so risk factors describe groups of people rather than guaranteeing an outcome for one person. A factor that is present only after the disease has begun is not a risk factor, and a treatment that lowers the chance of disease is a protective intervention rather than a risk factor.
Example — congestive heart failure: smoking · hypertension · hypercholesterolemia · family history · diabetes · valvular disease. For other diseases the risk factors may be less well known.
IO 1 — Recognizing risk
| Situation | Effect on recognizing the risk |
|---|---|
| Large risk with rapid effects | Easy to recognize: a big effect that appears soon after exposure is hard to miss |
| Long latency between exposure and disease | Can obscure the results: the disease appears long after the exposure, so the two can look unrelated unless people are followed for a sufficient period of time |
| Frequent exposure that results in disease | May be obscured unless “exposed” is compared with “unexposed”; when most people are exposed the disease may seem ordinary |
| Low disease incidence | More difficult. If the risk is small, a larger population must be affected to detect a difference between exposed and unexposed; a small risk produces few extra cases |
| Common disease with known risk factors | Difficult to distinguish new risk factors |
| One risk factor, many outcomes | A given risk factor may contribute to a number of different outcomes (multiple causes and effects) |
Also tested
- Risk factors easier to recognize. Risk factors are most readily recognized when exposure is associated with a new disease, because no established risk factors compete and a large effect appearing soon after exposure is easy to recognize. Long latency, low disease incidence, or a common disease with known risk factors make recognition harder.
IO 1 — What risk is used for
| Use | What it means |
|---|---|
| Prediction of a future event | Probabilities are based on groups; relative risk is calculated to determine the probability |
| Cause | Risk factors may predict disease but may not be a cause; a factor that predicts without causing is a disease marker. Prediction does not prove causation |
| Diagnosis | Risk factors increase the probability of getting the disease, but the best predictor in an individual is the presence of early clinical disease |
| Prevention | Removal of the risk factor lowers the probability of developing the disease. Examples: controlling hypertension to prevent heart failure; controlling dietary intake to prevent obesity and so decrease the likelihood of type II diabetes mellitus; avoiding second-hand cigarette smoke to avoid cancer. Advising a patient to stop smoking to lower the chance of lung cancer is prevention |
IO 1 — How risk is studied
Experimental studies of risk in humans are limited by ethical constraints: investigators cannot ethically assign people to a harmful exposure. So observational studies are used: a concurrent (prospective) cohort, a historical (retrospective) cohort, or a case-control study.
Also tested
- Why a trial of cell phone use is unlikely. Randomly assigning people to a potentially harmful exposure such as heavy cell phone use is an ethical problem, and people would not accept long-term restrictions on using their phones, so observational designs are used instead.
IO 2 — Relative risk and what its values mean
Relative risk asks how many times more likely exposed persons are than non-exposed persons. It is expressed as a ratio of the incidence in exposed persons to the incidence in non-exposed persons, and it examines the strength of association between the independent variable (exposure) and the dependent variable (outcome).
| Relative risk | Meaning |
|---|---|
| = 1 (null value) | Absolute risk in the exposed equals that in the unexposed: no association between exposure and disease |
| > 1 | Absolute risk in the exposed is greater than in the unexposed persons; risk is higher with the exposure |
| < 1 | Absolute risk in the exposed is less than in the unexposed persons: a negative association that may be protective, for example protection through exposure to an effective vaccine |
- For prevention of disease we expect the relative risk to be less than one.
- Relative risk is a ratio of two incidences, and incidences cannot be negative, so relative risk can take any value above zero but cannot be negative (and is not capped at 2 or clustered near 1).
- Reading the number: a relative risk of 2.5 means exposed persons are 2.5 times as likely to develop the disease. Above 1, the excess is read as a percentage (1.25 is 25 percent higher, 1.4 is 40 percent higher); below 1, the shortfall is read the same way, so a relative risk of 0.6 means the risk in the exposed is 0.6 times that in the unexposed, that is 40 percent lower.
- Relative risk is a ratio, not a percentage: a 2.5 percent increase in risk would be a relative risk of 1.025.
- Raw case counts cannot be compared directly, because the groups differ in size; each count is first turned into an incidence (cases divided by the people in the group) and then the two incidences are divided.
Worked example — birth control pills and thrombophlebitis
For 10 years, 2,000 women were followed: 1,000 randomly selected young women on birth control pills and 1,000 randomly selected young women who were nonusers. 30 of the women on the pills developed thrombophlebitis over the 10-year period; 3 of the nonusers developed it over the same period.
| Outcome — yes | Outcome — no | Total | |
|---|---|---|---|
| Exposed (took the pills) | exposed with the outcome | exposed without the outcome | all who were exposed |
| Unexposed (did not take the pills) | unexposed with the outcome | unexposed without the outcome | all who were unexposed |
| Total | all with the outcome | all without the outcome | everyone in the study |
| Thrombophlebitis | No thrombophlebitis | Total | |
|---|---|---|---|
| Pill users | 30 | 970 | 1,000 |
| Pill non-users | 3 | 997 | 1,000 |
Risk in users = 30/1,000 = 0.03. Risk in non-users = 3/1,000 = 0.003. Relative risk = risk in users ÷ risk in non-users = 0.03 ÷ 0.003 = 10. Expressed: women on the pill are 10 times more likely to develop thrombophlebitis than those not on the pill. Had the relative risk equaled 1, the pill would not increase the risk of thrombophlebitis: there would be no association between pill use and the development of thrombophlebitis.
Also tested
- Attributable risk without the mutation. Attributable risk is the incidence in exposed persons minus the incidence in non-exposed persons. Among women without the mutation, oral contraceptive users (3.0) minus non-users (0.8) gives 2.2 extra cases per 10,000 women per year.
- Attributable risk of factor V Leiden. Among oral contraceptive users, the incidence in carriers of factor V Leiden (28.5) minus the incidence in non-carriers (3.0) gives an attributable risk of 25.5 extra cases per 10,000 women per year.
IO 3 — Predicting risk from groups
Risk is used for prediction of a future event, and the probabilities are based on groups of people, not on one patient followed in detail. Relative risk is calculated to determine the probability. For an individual patient, the best predictor is the presence of early clinical disease, not the presence of a single risk factor and not the relative risk or the incidence of the wider population. A probability quoted from cohort data, such as the chance a smoker develops lung disease within ten years, is prediction.
Also tested
- Calibration plot. Calibration compares the predicted probability with the observed proportion. Points lying consistently above the line of perfect agreement mean predicted rates are consistently higher than observed rates, so calibration is poor.
- Risk estimate is a group probability. A risk estimate is a probability based on groups. A 2% chance of colorectal cancer in the next 5 years means the patient is a member of a group in which a very small number will develop it.
- Risk model discrimination. A risk model does not discriminate well when the estimated-risk curves of people who did and did not develop the disease heavily overlap, because both groups get nearly the same estimates.
- Calibration versus discrimination. Good calibration, meaning predicted rates match observed rates, does not impair discrimination. A model can be well calibrated yet discriminate poorly, so good calibration is not a reason risk models fail to identify who will develop disease.
- Discrimination and calibration. A risk model with poor discrimination but excellent calibration can still provide useful risk estimates, because its predicted probabilities match observed rates even if it cannot separate individuals well.
- Risk prediction models. Risk prediction models are useful for all three: predicting onset of disease, diagnosis, and predicting prognosis.
IO 3 — Absolute risk, and the cohort that measures it
Absolute risk is the possibility of an event in the population under study: the incidence of disease in that population. It does not indicate whether exposure is associated with an increased risk of disease.
Among the observational designs, the cohort study establishes incidence directly: following exposed and unexposed persons forward shows who develops the disease (the printed table of cohort study advantages calls it the only way of establishing incidence, that is absolute risk, directly). A case-control study starts from persons who already have the condition, a case report describes a single patient, and expert opinion is not a followed group, so none of them supplies incidence.
Predicting cases from a relative risk. Because relative risk = exposed incidence ÷ unexposed incidence, the exposed incidence = relative risk × unexposed incidence. Multiply that incidence by the number of people in the group to get the expected number of cases. Example: unexposed incidence 0.05 and relative risk 3 gives an exposed incidence of 0.15, so 200 exposed persons are expected to produce 0.15 × 200 = 30 cases. The relative risk multiplies the incidence, not the head count.
IO 4 — What a cohort is, and how it is followed
A cohort is a group of people with a common characteristic assembled at a similar point and followed over time. The characteristic may be age, date of birth, exposure to a risk factor, having a specific disease, or participation in a preventive or therapeutic intervention. It can be any characteristic that is followed for a significant period of time.
Design of a cohort study of risk: persons without disease are divided into two groups — those exposed to a risk factor and those not exposed. Both groups are followed over time to determine what proportion of each group develops disease. Relative risk can be calculated.
- The cohort must be followed over a sufficient period of time to observe the expected outcome.
- It must have complete follow-up so that attrition (people dropping out) does not falsely skew the outcome data. A larger control group does not replace lost participants; the safeguard is following everyone through to the outcome.
- Check what was actually measured in the follow-up: were data points collected that relate to the study?
Also tested
- Not a cohort study. A cohort is a group with a common characteristic assembled at a similar point and followed forward over time; reviewing the earlier course of patients who are already under care, with no common starting point, is not a cohort study.
IO 4 — Historical and concurrent cohorts
| Historical (retrospective) cohort | Concurrent (prospective) cohort | |
|---|---|---|
| When the cohort is assembled | In the past, identified from records, then followed forward over time to the present | In the present, then followed into the future |
| Outcomes | Have already occurred | Occur during follow-up |
| Weakness | Utility is limited since data may be incomplete: records were not set up for all the data points the research needs | Needs a long wait and many subjects (see the table below) |
Be careful: a case-control study grouped by disease is sometimes called “retrospective”. A historical cohort is grouped by exposure and followed forward, so it is not a case-control study. It is also not a randomized trial, which assigns the exposure by chance.
IO 4 — Advantages and limitations of cohort studies
| Advantages | Limitations |
|---|---|
| Establishes incidence directly | Inefficient: a large cohort must be enrolled, since only a limited number develop the condition |
| Logical approach: if persons are exposed, do they develop disease? | Expensive |
| Exposure can be elicited without bias from already knowing the outcome (recall bias: exposure information can be biased if participants know the disease studied and can recall) | Results are not available for many years |
| Can assess the relationship between the exposure and a number of outcomes or diseases | Only the risk factor (exposure) recorded at the outset can be followed, so other contributing exposures may be missed when a disease has multiple causes |
| Cohort type | Advantages | Disadvantages |
|---|---|---|
| All cohort study types | The only way of establishing incidence (absolute risk) directly · follows the same logic as the clinical question · exposure elicited without the bias that might occur if the outcome were known before exposure was documented · can assess the relationship between exposure and many diseases | Susceptible to confounding and other biases |
| Prospective | Can study a wide range of possible risk factors · can collect lifestyle and demographic data not available in most medical records · can set up standardized ways of measuring exposure and degree of exposure | Inefficient, because many more subjects must be enrolled than experience the event of interest, so it cannot be used for rare diseases · expensive · results not available for a long time · assesses only the relatively few factors recorded at the outset |
| Retrospective (historical) | More efficient, cheaper and faster than prospective cohort studies, because the data were already collected for another purpose and the patient outcomes have already occurred | Narrower range of possible risk factors · cannot examine patient characteristics not available in the data set · measurement of exposure and degree of exposure may not be standardized |
| Case-cohort | All advantages of retrospective cohort studies apply; even more efficient because only a sample of the unexposed group is analyzed | All disadvantages of retrospective cohort studies apply; difficult for readers to understand the weighting procedures used in the analysis |
IO 5 — Three types of risk
| Measure | What it is | Formula |
|---|---|---|
| Absolute risk | Possibility of an event in a population; the incidence of disease in a population. Does not indicate if exposure is associated with increased risk | New cases over a given time ÷ number of people in the group |
| Attributable risk | Possibility of disease following exposure; takes exposure to a risk factor into account. It is the excess incidence in the exposed group | Incidence in exposed persons minus incidence in non-exposed persons |
| Relative risk | How many times more likely exposed persons are than non-exposed persons; examines strength of association | Incidence in exposed persons divided by incidence in non-exposed persons |
- Attributable risk is a difference (subtraction, exposed minus unexposed, in that order); relative risk is a ratio of two incidences. Adding the incidences gives nothing meaningful.
- Example: incidence 15 per 500 (0.03) in exposed persons and 5 per 500 (0.01) in unexposed persons gives an attributable risk of 0.03 − 0.01 = 0.02, while the relative risk is 0.03 ÷ 0.01 = 3.0.
- Example: 36 of 300 exposed (0.12) and 20 of 400 unexposed (0.05) persons developed disease. The relative risk is 0.12 ÷ 0.05 = 2.4 and the attributable risk is 0.07. Dividing the raw counts (36 by 20) ignores the group sizes.
- Example: an unexposed incidence of 0.02 and a relative risk of 4 give an exposed incidence of 4 × 0.02 = 0.08; the attributable risk is 0.08 − 0.02 = 0.06.
- Example: an incidence of 8 percent in exposed and 2 percent in unexposed persons gives a relative risk of 8 ÷ 2 = 4.0; the difference of 6 percentage points is the attributable risk, and the sum of the two incidences (10) has no meaning. In a cohort where 10 of 500 vaccinated and 40 of 500 unvaccinated persons developed influenza, the relative risk with vaccination is 0.02 ÷ 0.08 = 0.25.
Also tested
- Population attributable risk. Population attributable risk is the attributable risk multiplied by the prevalence of the risk factor in the population: 25.5 per 10,000 x 0.05 = 1.3 per 10,000 women per year.
IO 5 — Same relative risk, different attributable risk
The same relative risk can hide very different numbers of excess cases, because relative risk cancels out incidence while attributable risk depends on it. The higher the baseline incidence, the greater the attributable risk for the same relative risk.
| Incidence in unexposed | Incidence in exposed | Relative risk | Attributable risk (exposed − unexposed) |
|---|---|---|---|
| 1/10,000 | 2/10,000 | 2.0 | 1/10,000 (0.1 per 1,000) |
| 1/1,000 | 2/1,000 | 2.0 | 1/1,000 |
| 1/100 | 2/100 | 2.0 | 10/1,000 |
| 1/10 | 2/10 | 2.0 | 100/1,000 |
For example, with a relative risk of 2.0 the attributable risk is 0.001 when the baseline incidence is 1 in 1,000 (0.002 minus 0.001) but 0.1 when the baseline is 1 in 10 (0.2 minus 0.1). The exposure with the higher baseline incidence has the far greater attributable risk; the relative risks are equal, so equal relative risks do not mean equal attributable risks.
IO 6 — Confounding variables
Confounding causes systematic error (bias). A confounder is an additional independent factor related to both the exposure and the outcome, so it can distort the apparent association and produce a false one. It is not random error: chance produces variation equally likely above or below the truth, small samples add random error, and dropout is the attrition effect.
Example: if smoking is more common among coffee drinkers and also raises the risk of pancreatic cancer, smoking is related to both the exposure (coffee) and the outcome (cancer) and is a confounder of the coffee-cancer association.
IO 6 — Methods for controlling confounding
| Method | Description | Design | Analysis |
|---|---|---|---|
| Randomization | Assign patients to groups in a way that gives each patient an equal chance of falling into one group or the other, so confounders are spread between the groups by chance | + | |
| Restriction | Limit the range of characteristics of patients in the study | + | |
| Matching | For each patient in one group, select one or more patients with the same characteristics (except the one under study) for a comparison group | + | + |
| Stratification | Compare rates within subgroups (strata) with otherwise similar probability of the outcome | + | |
| Simple adjustment | Mathematically adjust crude rates for one or a few characteristics so that equal weight is given to strata of similar risk | + | |
| Multivariable adjustment | Adjust for differences in a large number of factors related to outcome using mathematical modeling techniques | + | |
| Best-case / worst-case analysis | Describe how different the results could be under the most extreme (or simply very unlikely) assumption about selection bias | + |
- Only randomization and restriction are design-phase-only methods.
- Matching works at either phase, design or analysis.
- Stratification, simple adjustment, multivariable adjustment and best-case/worst-case analysis are applied only in the analysis phase.
IO 7 — Effect modification
Effect modification means the presence or absence of a variable changes the effect of the exposure. The change may be synergistic (the variable strengthens the effect) or antagonistic (it weakens or opposes the effect). It is NOT a bias: it is a real change in the effect, whereas confounding is a systematic error to be controlled.
Example: aspirin, age and prior history and their effect on the development of upper gastrointestinal bleeding — age and prior history modify how strongly aspirin affects the chance of bleeding. Likewise, if a drug lowers blood pressure far more in younger than in older adults, age is an effect modifier of the drug's effect, not a confounder.
| Confounding | Effect modification | |
|---|---|---|
| What it is | Additional factor related to both exposure and outcome | Variable whose presence or absence changes the effect of exposure |
| Bias? | Yes, systematic error, controlled by methods such as matching and stratification | No, not a bias; it may be synergistic or antagonistic and can arise in observational studies as well as trials |
IO 8 — Case-control design
In a case-control study the case group has the condition and the control group does not. The groups are chosen on the basis of whether they have the condition (outcome), not on the basis of exposure. Controls may be drawn from the same cohort, from the community or a hospital, or matched to the cases.
Direction: cases and controls are identified first and each is then compared for earlier exposure to the risk factor (exposed or not). Time runs forward for events but the research runs backward, from disease to past exposure — the reverse of a cohort study, which looks forward from exposure to disease. The design ends in an estimate of relative risk. The exposure is the factor compared between the groups (for example, past sunburns in a study comparing patients with and without melanoma); the presence or absence of the condition defines the groups.
Exposure information. Exposure is ideally measured, but oftentimes subjects are asked about their exposure, which brings recall bias: the accuracy of memory and the ability to recall specific information may not be perfect, especially once subjects know the disease studied.
Example of the design: a case-control study of helmet use and head injury among skiers and snowboarders. Cases had a head injury and controls did not; each was compared for whether a helmet had been used, with age, sex, nationality, skill level, equipment used, ski school attendance and rented-or-owned equipment controlled.
Also tested
- Common exposures and comparison. When an exposure is frequent in the population, a high frequency among patients alone can be misleading, because many people would have it regardless of their illness. It must be compared between people with and without the disease (exposed versus unexposed).
- Biased exposure information. In a case-control study, exposure can be artificially higher in cases because of recall bias or more careful questioning, recording or abstraction, which falsely raises the odds ratio. Exposure recorded at the time of an event came before it, so the event cannot have led to it.
- Sampling cases and controls. Sampling them from a defined population or cohort tends to give cases and controls representative of the source population. A cohort representative of the population is best, rather than a dynamic population, especially if exposure or disease changes rapidly over time.
- Multiple control groups. The main control group may differ systematically from the cases in ways other than the exposure of interest. A second control group shows whether results change substantially; if they do, the findings are called into question.
- Differences between cases and controls. A difference between cases and controls on a characteristic such as sports participation does not by itself show that control selection was biased, because it may be a real association or a confounder.
IO 8 — Advantages and limitations of case-control studies
Cohort studies are not ideal for capturing causation: there is a latency period between exposure and pathological effect, and a large sample is required although only a limited number of participants get the condition. Case-control studies are practical for this situation, and better suited to a rare disease with a long latency:
| Advantages | Limitations |
|---|---|
| Smaller samples, because the cases are picked for already having the condition | Recall bias when subjects are asked about past exposure |
| No need to wait for the condition to develop, so results do not take years | Cannot establish incidence directly; the measure of association is the odds ratio rather than relative risk |
| Suits a rare disease with a long latency, which a cohort could study only with a huge group followed for years | Depends on a control group without the condition, which must be chosen (from the same cohort, community or hospital, or matched) |
A case-control study still needs a defined case group and a control group, and still asks about earlier exposure; what it avoids is waiting for the condition to develop. Once the exposures that increase the risk of disease are understood, the exposure is followed to determine causation.
Also tested
- Association is not cause. A case-control study, even a well-designed one nested in a cohort, can show an association and identify a risk factor but cannot claim to have identified a cause: risk factors may predict disease without causing it, and unmeasured confounding is always possible.
- Case-control analysis within a cohort. When the cohort dataset lacks confounders that should be controlled for, a case-control analysis is more efficient, because the extra data are collected only for the cases and a sample of controls, not for the entire cohort.
- Case-control studies and incidence. Case-control studies cannot provide incidence, because the numbers of cases and controls are selected by the investigator. For a question such as how often complications occur, a cohort study is needed.
IO 9 — Which measure for which design
The odds ratio (relative odds) is the measure of association for case-control studies. For cohort studies, the incidence of disease in the exposed versus non-exposed can be calculated directly, so relative risk is usually used rather than the odds ratio. Relative risk and odds ratio allow numerical expression of the likelihood of a risk outcome: relative risk for cohort studies, odds ratio for case-control studies. The odds ratio is an estimated relative risk.
Example: researchers enroll 150 patients with a rare cancer and 300 without it and ask about past exposure. This is a case-control study, so the measure of association is the odds ratio, not the relative risk, attributable risk or absolute risk.
Also tested
- Odds ratio and relative risk. The odds ratio approximates relative risk only when the disease is rare (rule of thumb: under about 1 in 100 to 5 in 100). For a common outcome with controls sampled from noncases after the study ends, an odds ratio above 1 overestimates relative risk.
- Prevalence odds ratio. An odds ratio based on prevalent cases is a rough measure of the association between a patient characteristic and having the disease at that point in time, not a comparison of risk, which requires incident (new-onset) cases.
IO 9 — The lettered 2×2 table and the formulas
| Cases | Noncases | Total | |
|---|---|---|---|
| Exposed | a | b | a + b |
| Not exposed | c | d | c + d |
| Total | a + c | b + d |
| Study | Measure | Formula |
|---|---|---|
| Cohort | Relative risk | [ a / (a + b) ] ÷ [ c / (c + d) ] — the incidence in the exposed, a over a plus b, divided by the incidence in the unexposed, c over c plus d |
| Case-control | Odds ratio | [ a / c ] ÷ [ b / d ] = ad / bc — the odds of exposure in cases, a over c, divided by the odds of exposure in controls, b over d, which simplifies to the cross-products ad over bc |
- Odds compare those with the exposure to those without it: exposed ÷ not exposed (48 exposed and 12 unexposed cases give odds of 48/12 = 4, not the proportion 0.8).
- Odds ratio = odds of exposure in cases ÷ odds of exposure in controls. The case odds go on top; inverting the ratio is a slip.
- Example: 60 of 100 cases and 30 of 100 controls exposed. Odds in cases 60/40 = 1.5; odds in controls 30/70 = about 0.43; odds ratio = 1.5 ÷ 0.43 = 3.5, matching ad/bc = (60 × 70)/(40 × 30) = 4,200/1,200.
- Example: 20 of 100 cases and 40 of 100 controls exposed. Odds in cases 20/80 = 0.25; odds in controls 40/60 = about 0.67; odds ratio = about 0.38, which is below 1, so exposure is associated with lower odds of disease.
- Dividing the exposed counts alone (60 by 30) ignores the unexposed counts and is not the odds ratio.
IO 9 — A worked odds ratio, and how to state it
200 women were selected: 100 young women with thrombophlebitis and 100 without. History of prior birth control pill use in the past 5 years was carefully obtained. 90 of the 100 women with thrombophlebitis were using pills; 45 of the 100 women without thrombophlebitis had used them.
| Thrombophlebitis | No thrombophlebitis | |
|---|---|---|
| Pill users | 90 | 45 |
| Pill non-users | 10 | 55 |
| Total | 100 cases | 100 controls |
Odds of pill use in cases = 90/10 = 9. Odds of pill use in controls = 45/55 = 0.82. Odds ratio = 9 ÷ 0.82 = 11.0.
Expressing the result. Women with thrombophlebitis have 11 times greater odds of having taken birth control pills than women without thrombophlebitis. The version to use, and the one we actually want to know, is: women who take birth control pills have 11 times greater odds of developing thrombophlebitis. The odds ratio is the same whichever way it is read, so a result such as “patients with lung disease had six times the odds of prior asbestos exposure” is stated as: exposed persons have six times the odds of the disease.
- An odds ratio is a ratio of odds. It is not a percentage of exposed persons who are affected, and an odds ratio of six is not a one-in-six frequency.
- An odds ratio above 1 means the exposed, not the unexposed, have the greater odds.
Definitions and rules of thumb
- Risk factor. A risk factor is a factor associated with a higher chance of developing a disease later. It raises the probability of disease rather than making it certain, so risk factors describe groups, not a guaranteed outcome for one person.
- Who takes part in a risk study. Healthy people take part: a study of risk follows people who do not yet have the disease, because the question is who goes on to develop it, so the outcome of interest is diagnosis.
- Long latency. A long latency between exposure and disease can obscure the association: the disease appears long after the exposure, so the two can look unrelated unless people are followed for a sufficient period of time.
- Relative risk of 1.0. When the risk in the exposed group equals the risk in the unexposed group, the relative risk is 1.0, the null value, and this indicates no association between exposure and disease. For example, 20 of 500 exposed and 16 of 400 unexposed are both 0.04.
- Recall bias. Recall bias arises from participants' difficulty remembering past exposures, which distorts the reported exposure or history; it is a particular weakness of case-control studies.
- Effect modifier. A variable whose presence or absence changes the effect of the exposure is an effect modifier. If an exposure raises the risk much more in older patients than in younger ones, age is an effect modifier, which differs from a confounder that only distorts the association.
8 · Prognosis & Outcomes
Megan B. Finck, MMS, PA-C · Session 7 · Fletcher ch. 5-8
Instructional Objectives
- Describe prognosis
- Describe prognostic factors
- Compare and contrast risk and prognostic factors
- Discuss the clinical course and natural history of disease
- Discuss methods to measure outcomes
- Discuss survival analysis
- Discuss bias in cohort studies
IO 1 — Describe prognosis
Prognosis is the prediction for the course of disease. It is what a clinician tries to estimate once a diagnosis has been made: how the disease will unfold in terms of survival, complications, disability and similar outcomes.
- Prognosis is expressed quantitatively as a rate, such as five-year survival, the case-fatality rate, the response rate, the remission rate and the recurrence rate. Each states how often an outcome occurs.
- Prognosis can also use cure or death as its endpoints.
- The natural history of a disease is what occurs without intervention. Prognosis can often be improved through clinical intervention, which is how clinicians change the natural history to improve the outcome.
- The reason to detect disease at an earlier point is to maximize the effectiveness of treatment. Earlier detection helps only if treatment given sooner actually changes the outcome.
IO 2 — Prognostic factors
Prognostic factors help identify groups of patients who share a similar disease but have different prognoses. They work within one disease, separating the patients who have it into groups that fare better or worse.
- Prognosis studies look for protective or progressing factors — factors that protect against, or promote, progression of the disease. These are different from the factors that raise the chance of getting the disease.
- Multiple factors are a better predictor of outcome than a single data point.
- Clinical prediction rules estimate the probability of an outcome based on a set of patient characteristics, combining several factors into one estimate for an individual patient.
- Adjustment for important prognostic factors addresses whether other factors explain a difference in outcome: are the differences in the prognosis of the groups related to the factor under study or to some other factor?
Also tested
- Choosing a prediction rule. The strongest reason for choosing a clinical prediction rule is that it has been validated in different settings, that is, it predicts outcomes in patients other than those used to develop it.
- Clinical prediction rules. A clinical prediction rule estimates the probability of an outcome from a set of multiple patient characteristics, is simple to score using readily available clinical data, and is used to guide further management.
- Prevalence study cannot identify prognostic factors. Prognosis concerns outcomes of disease over time, but a prevalence (cross-sectional) study assesses a population for a condition at a single point in time and measures no events over time, so it cannot identify prognostic factors.
Example — acute myocardial infarction
In a patient who has had an acute myocardial infarction, the poor prognostic factors are: increasing age, female sex, cigarette smoking, hypotension, anterior infarction, congestive heart failure and ventricular arrhythmia. The outcomes are death, reinfarction and other complications.
Ventricular arrhythmia is a poor prognostic factor that arises after the infarction; it is not a risk factor for developing coronary disease in a healthy person.
IO 3 — Risk factors versus prognostic factors
Risk factors do not necessarily represent a poorer prognosis. Being at risk of getting a disease and faring badly once you have it are separate questions, so the two sets of factors can differ, although some factors appear on both lists.
| Risk studies | Prognosis studies | |
|---|---|---|
| Participants | Healthy participants | Sick participants |
| Outcome | Outcome is diagnosis | Outcome is death, complications, disability… |
| Rates | Rates are low — diagnosis is rare | Rates are high — outcomes are common |
| Factors | Risk factors may predict diagnosis | Protective or progressing factors are different from risk factors |
The coronary disease example
| Risk factors (predict who gets the infarction) | Poor prognostic factors (predict the outcome once it has happened) |
|---|---|
| Increasing age | Increasing age |
| Male sex | Female sex |
| Cigarette smoking | Cigarette smoking |
| Hypertension | Hypotension |
| Increased low-density lipoprotein / decreased high-density lipoprotein | Anterior infarction |
| Inactivity | Congestive heart failure |
| — | Ventricular arrhythmia |
Increasing age and cigarette smoking appear on both lists: they raise the chance of an infarction and worsen the outlook after one. Inactivity, hypertension and the lipid pattern are risk factors for developing coronary disease. Anterior infarction, congestive heart failure, hypotension and ventricular arrhythmia are poor prognostic factors that apply once the infarction has occurred. Hypertension is the blood pressure abnormality on the risk list; hypotension is the one that marks a poor prognosis.
IO 4 — Clinical course and natural history
The natural history of disease is what occurs without intervention. It is the untreated benchmark: as new treatments arise, they are compared with the natural (untreated) outcome of disease.
The five Ds
The course of disease is described and quantified by the five Ds: death, disease, discomfort, disability, dissatisfaction.
Severity of disease matters because it establishes priorities for clinical services and public health programs.
Stages of the clinical course
| Order | Stage |
|---|---|
| 1 | Biologic onset |
| 2 | Pathology evident |
| 3 | Signs and symptoms present |
| 4 | Medical care sought |
| 5 | Diagnosis given |
| 6 | Treatment recommended and implemented |
| 7 | Outcome |
The first stage is biologic onset. Diagnosis is given immediately after medical care is sought, and treatment is recommended and implemented only after the diagnosis. The outcome comes last.
- The preclinical phase is the period in which the disease is present but signs and symptoms have not yet appeared.
- The clinical phase begins once signs and symptoms of disease are present, and runs through care sought, diagnosis and treatment to the outcome.
- Screening can find disease before signs and symptoms appear, that is, during the preclinical phase. Detecting disease at an earlier point is what screening aims to do.
IO 5 — Rates that describe prognosis
| Rate | Definition |
|---|---|
| 5-year survival | Percent of patients surviving 5 years from some point in the course of their disease |
| Case fatality | Percent of patients with a disease who die of it |
| Disease-specific mortality | Number of people per 10,000 (or 100,000) population dying of a specific disease |
| Response | Percent of patients showing some evidence of improvement following an intervention |
| Remission | Percent of patients entering a phase in which disease is no longer detectable |
| Recurrence | Percent of patients who have return of disease after a disease-free interval |
Time under observation is either stated or assumed to be sufficiently long that all events that will occur have been observed.
Outcomes of disease should include the full range of manifestations: consider both biological and clinical outcomes, and also consider quality of life measures.
Also tested
- Clinical outcome matters most. Outcomes should include both biologic and clinical outcomes; for a new drug the effect on the clinical outcome, such as the rate of sudden death, matters most, because an intermediate biologic outcome such as fewer arrhythmias can be an unreliable marker of it.
The arithmetic, with fresh numbers
| Rate | Formula | Example |
|---|---|---|
| Case-fatality ratio | People who die ÷ people who have the disease | 12 deaths among 150 patients = 8% |
| Disease-specific mortality | Deaths from the disease ÷ population × 100,000 | 24 deaths in a population of 80,000 = 30 per 100,000 |
| Five-year survival | Patients alive at 5 years ÷ patients followed | 90 alive of 200 = 45% |
| Response rate | Patients improved ÷ patients treated | 36 improved of 60 = 60% |
| Recurrence rate | Patients whose disease returned ÷ patients who were disease free | 8 of 40 = 20% |
Case-fatality ratio
The case-fatality ratio is the number of people who die divided by the number of people who have the disease. It is usually used for acute, short-term disease where death occurs shortly after diagnosis, because the deaths can be counted without long follow-up. For chronic disease with survival of decades, survival measures are used instead.
From a risk communication standpoint, be careful with negative framing — presenting the number dying rather than the number surviving. The data are the same, but the framing can sway perception: it changes how the risk is perceived.
Five-year survival
Five-year survival is frequently used in clinical medicine, particularly to describe outcomes of cancer patients.
- Rates can mask important data. Four conditions can have nearly identical five-year survival — each of the four curves ends with roughly 5–10% of patients alive at five years — while the survival curves differ greatly. A single rate does not show when the deaths occur within the five years.
| Condition | Shape of the survival curve over five years |
|---|---|
| Dissecting aneurysm | Falls steeply in the first year (most patients die early), then levels off at a low plateau |
| Lung cancer | Falls steadily and roughly evenly over the years |
| Amyotrophic lateral sclerosis | Falls slowly at first, then more quickly |
| Age 100 years | Falls gradually and continuously |
Screening can inflate five-year survival
The five-year time line is problematic when considering screening programs: the disease may be found earlier, but death still occurs at a similar point. Patients simply live longer with a diagnosis, so measured survival rises without anyone living longer. This is a potential bias if disease is detected earlier but the natural history of disease is not impacted favorably by treatment.
Reading the timeline: survival from diagnosis is the year of death minus the year of diagnosis. In the standard timeline used to illustrate this, the disease begins in 1987, is detected by screening in 1989 or diagnosed without screening in 1991, and death occurs in 1995 either way. Survival from diagnosis is therefore 6 years with screening (1989 to 1995) against 4 years without (1991 to 1995), yet death did not come any later. Moving the diagnosis earlier by some number of years adds that many years to the measured survival when the date of death is unchanged.
Another limitation: those with a history of disease duration of less than five years cannot be evaluated by this measure, nor can patients treated with a new therapy introduced within the past five years, because no one has yet been followed for five years.
IO 6 — Survival analysis
Survival analysis follows a cohort over time and records the outcomes at designated time intervals. People can be added at various times in the course of their disease, and each contributes the time they are followed, giving the person-years contributed to the study.
The probability of surviving to any point in time is estimated from the cumulative probability of surviving each of the time intervals that preceded it — the chances of surviving each earlier interval are multiplied together. For example, surviving three successive intervals with probabilities of 0.95, 0.90 and 0.60 gives 0.95 × 0.90 × 0.60 = 0.513, about 51%.
Note on scope. The course material describes the outcome of survival analysis as limited to living versus dying; the same method is also applied to any single event that happens once, such as recurrence or remission.
Also tested
- The event in time-to-event analysis. The event is dichotomous, it occurs or it does not, such as living versus dying, and it can occur only once.
Kaplan-Meier method
- Rather than using pre-determined intervals, such as death at one year, the exact point in time when each death occurs is noted and a new interval begins. This generates the curve.
- Because each death starts a new interval, the curve is a series of steps: it holds level and then steps down at each death.
- Patients who are censored (no longer followed, and alive as far as is known) leave the number at risk without being counted as deaths. In each interval the probability of surviving is the number who survive divided by the number still at risk, for example 7 of 8 at risk surviving gives 0.875.
Also tested
- Who is censored. A patient is censored when removed from the study or not followed for the full time period. Developing another disease that could be fatal does not cause censoring while the patient is alive and under observation.
Cox proportional hazards survival regression
- Similar to Kaplan-Meier, but it can be used to study the effect of an exposure on outcome after adjusting for confounders.
- It can also be used to find prognostic indicators for survival in a given disease, taking into account the effects of multiple predictors on the shape of the survival curve — for example treatment differences, age differences and differences in drug dosages.
- The hazard ratio that results can be interpreted as a relative risk (risk ratio). A hazard ratio of 2 means the group with the factor has twice the risk of the outcome (dying) as the group without it. A ratio of less than 1 is considered protective; a ratio of 1 means the same risk in both groups.
- A hazard ratio of 0.25 means the group has 25% of the risk of the comparison group, that is, 75% lower.
Reading a survival curve
- Survival curves estimate the survival for a hypothetical, rather than actual, cohort.
- Precision is greater earlier, on the left side of the curve, because the population still being followed is larger.
- The shape of the curve can be deceptive: it tends to flatten or plateau when the event of interest is occurring to fewer numbers of participants being followed, so the flat tail rests on few participants and does not mean deaths have stopped.
Other survival measures
| Measure | Definition | Notes |
|---|---|---|
| Median survival time | The length of time that half the population survives | Advantages: less affected by extremes, and it can be calculated earlier, before all the population have died |
| Relative survival time | Based on the survival expected at a given age without the disease: observed survival in people with the disease ÷ expected survival if the disease were absent | Compares people with the disease with what would be expected without it |
The median is the middle value when survival times are put in order: for seven ordered times of 2, 4, 7, 9, 13, 18 and 30 months it is 9 months, whereas the mean (11.9 months) is pulled upward by the longest survivor. On a curve, the median survival is read where survival crosses 50%. For relative survival, an observed survival of 63% against an expected 90% gives 63 ÷ 90 = 70%.
IO 7 — Bias in cohort studies
Bias in survival cohort studies may falsely create or obscure differences that do exist.
Inception cohorts and zero time
- Cohorts should all be identified at a similar point in the disease course, called zero time — a shared point in the disease course.
- Inception cohorts are assembled at onset of disease, when it first becomes clinically apparent — as early as possible and as uniform as possible. Patients assembled at different points in the course cannot be compared fairly.
- Specific inclusion criteria need to be specified.
- Consider the referral pattern of the patients in the study: patients at an academic medical center may differ from those in the community.
- Sources of selection bias include convenience samples, volunteer samples, declines from those selected for study, and dropouts.
Also tested
- Zero time defined. Zero time is the point in the disease course at which every member of an inception cohort is assembled, ideally as early and as uniform as possible: the diagnosis of the condition that defines the cohort, not the time of the outcome.
- Choosing a prognosis sample. The most appropriate sample for a prognosis study depends on who will use the results, because referral pattern (academic medical center versus community) affects who is in the study and results apply only to patients similar to those the reader sees.
The named biases
| Bias | What it is | Examples and how it is handled |
|---|---|---|
| Susceptibility bias | Those assembled for the study differ in some way that affects the outcome | Different points in the disease course, other concurrent disease, prior treatment |
| Migration bias | Patients drop out of the study or change groups | Reasons include death, recovery, disability. Ask whether dropouts are similar to those who stay. As the proportion of dropouts increases, the potential for bias increases. |
| Measurement bias | Participants in one group have an increased likelihood of having their outcome detected than those in another group, that is, outcomes in one group are more likely to be detected | May be due to differences in methods used to classify the outcome. Minimize by blinding, establishing strict criteria for whether the outcome has occurred, and applying efforts to discover outcomes equally in all study groups |
| Selection bias | The compared groups differ in ways unrelated to the factor under study | Controlled by randomization, restriction, matching, stratification, adjustment and best/worst case variation |
The question to ask about selection bias: are the differences in the prognosis of the groups related to the factor under study or to some other factor(s)?
Matching, restriction and stratification handle only the factors chosen for consideration. Randomization is the only method of dealing with differences between groups whose benefits are not limited to the factors selected for consideration, so it also handles factors nobody measured. Studies typically use several methods to control bias. Blinding, matching, restriction and stratification each address a different bias: blinding addresses measurement bias, while the others address selection bias.
Generalizability is often limited because the patients in the study may not be representative of other populations or of individual patients.
Also tested
- Migration bias and prognosis. Dropouts distort results only if dropping out is related to prognosis, that is, if those who leave are not similar to those who stay. The more dropouts, the greater the potential for bias.
- A higher rate is not measurement bias. Measurement bias means one group is more likely to have its outcome detected or classified differently. A higher rate of incontinence in men who had surgery is a study result that can reflect a true effect, not itself measurement bias.
- Multiple imputation. Multiple imputation estimates missing values by statistical modeling of patients with complete data, taking their other characteristics into account. It assumes the measured characteristics can reliably predict the missing values.
What to look for in a prognosis study
- Was an inception cohort assembled? — so that patients start at a similar point in their disease.
- Were the entry criteria objective and reasonable?
- Was the follow-up of subjects adequate?
- Are the study patients similar to yours?
- Was there adjustment for important prognostic factors?
- Were outcomes assessed objectively and blindly?
Definitions and formulas
- Prognostic factors. Prognostic factors identify patients with the same disease who have different prognoses. They separate patients who share a diagnosis into groups that fare differently.
- Risk factors and prognosis. Risk factors do not necessarily represent a poorer prognosis, so being at risk of developing a disease and faring badly with it are separate questions; a risk factor need not worsen prognosis.
- Case-fatality ratio. Case-fatality ratio = number of deaths from a disease ÷ number of people with the disease. For example, 30 deaths among 400 people with the disease is 30/400 = 0.075, or 7.5 percent. It is most useful for acute, short-term disease where death follows quickly after diagnosis, because the deaths can be counted without long follow-up.
- Five-year survival rate. Five-year survival rate = number of patients alive at five years ÷ number of patients in the cohort when no one is lost to follow-up. For example, 90 of 150 patients alive is 90/150 = 0.60, or 60 percent.
- New treatments. As new treatments arise, they are compared with the natural (untreated) outcome of the disease, which is the baseline that shows whether a treatment changes what would otherwise happen.
9 · Prevention & Screening
Megan B. Finck, MMS, PA-C · Session 8 · Fletcher ch. 9–10
Instructional Objectives
- Discuss the types of clinical prevention
- Discuss the levels of prevention
- Discuss intervention in terms of prevention
- Describe advantages and disadvantages of screening exams
- Discuss the principles of community screening programs
- Describe number-needed-to-treat (NNT) and number-needed-to-harm (NNH)
Objective 1 — The types of clinical prevention
Prevention is defined as interventions for people without complaints, to identify and modify risk factors to avoid the onset of disease, or to find disease early in its course. Preventive care is built into clinical practice in four ways.
| Type | What it is | Example |
|---|---|---|
| Immunization | Giving a vaccine so that protection exists before infection occurs | A hepatitis B or influenza vaccine given to a healthy adult |
| Screening | Testing people who have no symptoms in order to find disease early | Measuring blood pressure in adults who feel well |
| Behavioral / lifestyle counseling | Advice aimed at changing habits; the intervention is advice, not a drug or a test | Counseling a smoker to quit, or advice about diet |
| Chemoprevention | Giving a medication to a healthy person solely to lower the chance of future disease | A daily medication started by a healthy patient at raised risk |
Objective 2 — The levels of prevention
| Level | What it does | Where it acts on the timeline of disease |
|---|---|---|
| Primary | Prevents disease from occurring by removing underlying causes or risk factors. Works at the individual or community level | Before any disease exists (“no disease”) |
| Secondary | Detects disease while it is asymptomatic; this is the level of screening | During asymptomatic disease, after onset and before clinical diagnosis |
| Tertiary | Clinical intervention (treatment) that reduces complications following diagnosis | During the clinical course, after clinical diagnosis |
The timeline runs: no disease → onset → asymptomatic disease → clinical diagnosis → clinical course. Primary prevention is labeled “remove risk factors”, secondary “early detection and treatment”, and tertiary “reduce complications”. The three levels are primary, secondary and tertiary; the four types of Objective 1 are different things and are not levels.
Also tested
- Prevention questions. Two questions fall under it: whether prevention stops disease arising, and whether early detection helps. They are different claims, since stopping disease from arising is not the same as finding it sooner once it has.
Objective 3 — Intervention in terms of prevention
When should a condition be included in preventive care?
Three questions decide whether a condition belongs in preventive care.
| Criterion | What is asked |
|---|---|
| 1. Burden of suffering | How great is the burden the condition causes? Measured by mortality rates, hospitalizations and incidence, and framed as six D’s: death, disease, disability, discomfort, dissatisfaction and destitution |
| 2. Quality of the screening test (if one is to be used) | Sensitivity, specificity, simplicity, cost, safety and acceptability |
| 3a. Treatment, for primary and tertiary prevention | How good is the therapeutic intervention: effectiveness, safety and cost-effectiveness |
| 3b. Treatment, for secondary prevention | How good is the treatment once the condition is found: effectiveness, safety, cost-effectiveness, and early treatment after screening being more effective than later treatment without screening, once the patient becomes symptomatic |
Each component belongs to only one criterion: disability, death and hospitalizations describe burden; sensitivity, specificity, simplicity and acceptability describe the test; effectiveness and cost-effectiveness of therapy describe the treatment.
Also tested
- Judging a vaccine. Efficacy, safety and the burden of the disease matter most when assessing a vaccine. Cost is only one component of a cost-effectiveness analysis, so on its own it is the least important.
How each level is studied
| Level | Typical research |
|---|---|
| Primary | Immunization studies are typically randomized controlled trials, and counseling approaches are studied with randomized controlled trials. Observational studies are used when randomization is not possible or not ethical. Longitudinal studies are used to study safety (adverse effects) |
| Secondary | Similar to studies in curative medicine, assessing outcomes following treatment. For screening to have value, outcomes must be better when disease is identified and treated early. A cohort design may be used because it can take years to determine effectiveness |
| Tertiary | Treatments are similar to those used in disease intervention and are tested with randomized controlled trials. Caution: such studies are typically done to prove the therapeutic effects of a treatment, not its preventive effects, so the effect on prevention is often identified later |
Cohort design. A cohort study starts with a defined population and divides it, without randomization, into an exposed and a not exposed group, then follows each forward to see who develops disease and who does not. The diagram shows the retrospective form, in which the groups are defined in the past (1980) and followed to the present (2000). Groups are defined by exposure and followed for the outcome; a case-control study does the reverse, starting from the outcome.
Objective 4 — Advantages and disadvantages of screening exams
The advantage
The advantage of screening is that disease can be found early, while it is asymptomatic, and so treated earlier in its course. It pays off when outcomes are better with early detection and treatment than when disease is found later.
Prevalence screens and incidence screens
| Screening round | Detects |
|---|---|
| First screening | Prevalent cases: disease already present in the population. Prevalence is the number affected at a given point in time |
| Second and subsequent screenings | Incident cases: disease that has newly developed since the last screen. Incidence is the number who develop the disease over a given time period |
Three biases in observational studies of screening
The three biases are lead-time bias, length-time bias and compliance bias. Each can make screening look better than it is.
Lead-time bias
Lead time is the time by which screening advances diagnosis: the time between detection from screening and when the condition would have been diagnosed from symptoms. The apparent improvement in survival time is not due to improved survival; it is due to increased time knowing the diagnosis.
| Situation | Diagnosis | Death | Survival measured from diagnosis |
|---|---|---|---|
| Unscreened | Usual (later, from symptoms) | At its usual time | Shorter |
| Screened, early treatment not effective | Earlier | Unchanged | Longer, but only because diagnosis moved earlier |
| Screened, early treatment is effective | Earlier | Later | Longer, and genuinely improved survival |
Survival from diagnosis is the age at death minus the age at diagnosis. If a cancer is found by screening at age 60 and would have been diagnosed from symptoms at age 64, and the patient dies at age 72 either way, survival from diagnosis is 12 years with screening and 8 years without. The apparent gain is 4 years, which equals the lead time, and the patient has not lived a day longer. Earlier diagnosis happens in every screened group; a longer survival is a real improvement only when early treatment is effective and delays death.
Length-time bias
Outcomes appear better with screening because more slow-developing conditions with better prognoses are detected. Slow-developing conditions are present for longer, so a screen is likely to catch them, and they have better prognoses than fast-developing conditions. Screening works best for conditions that develop slowly. Rapidly developing conditions usually cause symptoms before screening is done, so they tend to escape it and are diagnosed clinically instead. In a diagram of tumor growth against time, the rapidly growing tumors reach the point of diagnosis after symptoms between screens, while a single screen mostly catches the slower-growing ones.
Compliance bias
The outcome in the screened group appears better because those people are more compliant. Compliant patients are more invested in their health and have better outcomes for reasons unrelated to the screening itself. For example, volunteers who attend screening may have lower death rates than people who decline even if screening offers no benefit.
Properties of tests used in screening
- High sensitivity early in the disease, so the test does not miss disease
- High specificity, to reduce false positives. A screened population is mostly healthy, so specificity is what keeps healthy people from being wrongly flagged. (Missing disease is a false negative, which high sensitivity reduces.)
- Consider positive and negative predictive values because disease prevalence is low
- Simple to administer and inexpensive
- Safe to administer
Also tested
- Incidence method for sensitivity. The incidence method can be used to calculate the sensitivity of screening tests: people who test negative are followed over an interval, and those diagnosed with disease during that interval are the false negatives.
The disadvantages: unintended consequences of screening
| Consequence | Meaning |
|---|---|
| False-positive results | A healthy person is wrongly flagged: expense, inconvenience and unnecessary interventions |
| Negative labeling effect | Harm from being labeled as having a condition; an ethical concern |
| Overdiagnosis | Finding disease when the finding does not help the patient. It is an extreme example of length-time bias: screening picks up the slowest conditions, and at the extreme they are so indolent that finding them helps no one |
| Incidentalomas | An abnormality outside the area being examined, found unexpectedly while looking elsewhere |
A false-positive result flags someone who is healthy; a repeat-screen finding is an incident case, and neither is an incidentaloma. The biases above (lead-time, length-time, compliance) are further disadvantages, because they make screening appear better than it is.
Also tested
- Overdiagnosis in screening. Overdiagnosis is finding disease that would not have caused symptoms or become clinically significant. More cancers in a screened group that stays ahead, or more small tumors with no change in large tumors, means screening found cancers that would not have come to medical attention without it.
What to look for in prevention studies
- The rationale for the prevention: the burden, the quality of the screening test and the effectiveness of the intervention
- Identification of the level of prevention
- The properties of the tests used
- Strategies to reduce bias
Also tested
- Costs in cost-effectiveness analysis. A cost-effectiveness analysis of a preventive activity takes a societal perspective and counts all costs: delivery of the preventive activity, diagnostic follow-up and treatment with and without prevention, and indirect costs for patients and caregivers.
Objective 5 — Principles of community screening programs
A community screening program takes place in a community setting. Its objective is to test large numbers of individuals at little or no cost to the participants.
Possible objectives of a program
| Objective | Example |
|---|---|
| Target disease so treatment can cure patients | Sexually transmitted infections, cancers |
| Target disease so treatment can prevent complications | Hypertension |
| Target a behavioral risk factor to change the behavior | Trying to get people to stop smoking |
| Target a metabolic risk factor to change diet and avoid disease | Modified diet to reduce cholesterol levels and avoid coronary artery disease |
Establishing a screening program
- The condition should be an important or serious public health problem
- There should be effective treatment available
- The natural history of the disease should be well understood
- The disease must be prevalent but not found in everyone
- Those with positive screens must have access to care
- It must be clear who is responsible for following up the results of the screening: what counts as a positive test, how findings become part of a participant’s medical record, and so on
Before a program starts, the benefits of being screened must outweigh the costs and harms. The test should be safe with minimal side effects, quick and easy to administer, inexpensive, and acceptable to the public and to their health care providers. Its sensitivity, specificity, positive predictive value and negative predictive value must be known and be acceptable; the predictive values must be known as well because they depend on how common the disease is in the people screened.
Also tested
- Cost-effectiveness analysis. Cost-effectiveness analysis describes financial costs per clinical effect, such as dollars per year of life saved, for alternative tests or treatments, for example fecal occult blood testing compared with no screening.
Low prevalence: why high-risk groups are screened
With low prevalence, even highly specific tests generate more false positives than true positive results, because healthy people far outnumber cases: a small error rate applied to a very large healthy group produces more wrong positives than the few true cases produce right ones. Specificity and sensitivity themselves do not change with prevalence; the predictive values do. Therefore programs screen “higher risk” populations, where prevalence is higher. In a higher-risk group the positive predictive value rises (a larger share of positive results are true), while screening only people without risk factors would lower prevalence further and make the ratio of false to true positives worse. A test with lower specificity would generate still more false positives. The positive and negative predictive values can be calculated.
Also tested
- Prevalence in diagnostic and screening situations. Diagnostic tests are used in symptomatic people, and screening tests in asymptomatic people, so disease prevalence is higher in the diagnostic situation. Higher prevalence raises positive predictive value and lowers negative predictive value.
The arithmetic
Build the counts from the number screened, the prevalence, the sensitivity and the specificity (the same two-by-two table as in Session 4):
- People with disease = number screened × prevalence. True positives = people with disease × sensitivity; false negatives = the rest of them.
- Healthy people = number screened minus people with disease. False positives = healthy people × (1 − specificity); only healthy people can be false positives.
- Positive predictive value = true positives ÷ (true positives + false positives).
Example of the method with fresh figures: 5,000 people screened, prevalence 2 percent, sensitivity 80 percent, specificity 90 percent. People with disease = 100, true positives = 80. Healthy = 4,900, false positives = 490. Positive predictive value = 80 ÷ 570, about 14 percent. False positives outnumber true positives about six to one although the test is 90 percent specific.
Objective 6 — Number needed to treat and number needed to harm
| Number needed to treat (NNT) | Number needed to harm (NNH) | |
|---|---|---|
| Meaning | The number of patients who need to be treated for one to get a benefit (for example, prevent disease). Often discussed in studies of immunizations | The number of patients who need to have a given intervention for one to experience a harm |
| Calculation | 1 ÷ absolute risk reduction, where the absolute risk reduction = risk in the control group − risk in the treated group | 1 ÷ absolute increase in risk, where the increase = risk in the treated group − risk in the control group |
| Reading it | Large number → the intervention is largely ineffective (many treated to help one). Small number → largely effective | Small number → harm is frequent. Large number → harm is rare, so the intervention is safer |
Working with the numbers
- The standard calculation is not printed on the slides, which define the ideas. Both numbers are the reciprocal of an absolute change in risk, expressed as a decimal: an absolute risk reduction of 5 percentage points is 0.05, so the number needed to treat is 1 ÷ 0.05 = 20. A risk that falls from 10 percent to 6 percent is a reduction of 4 points (0.04), giving a number needed to treat of 25. A side effect that rises from 1 percent to 4 percent is an increase of 0.03, giving a number needed to harm of about 33.
- Use the difference between the two groups’ risks, not either group’s risk alone, and not the difference in event counts: when 50 of 1,000 untreated and 30 of 1,000 treated patients have an event, the risk falls from 0.05 to 0.03, an absolute reduction of 0.02, and the number needed to treat is 50, not 20.
- Comparing treatments for the same outcome over the same period: the smaller number needed to treat is the more effective treatment. For harm, the larger number needed to harm is the safer.
- Expected effect in a group = number of patients ÷ the number needed to treat (cases prevented) or ÷ the number needed to harm (patients harmed). With a number needed to treat of 50, treating 500 patients prevents 10 cases; with a number needed to harm of 100, treating 500 patients harms 5. Dividing is correct; multiplying is not.
- Comparing benefit and harm: one patient in every number-needed-to-treat benefits, and one in every number-needed-to-harm is harmed. A number needed to harm that is smaller than the number needed to treat means harm is more frequent than benefit (for example a number needed to treat of 30 and a number needed to harm of 20 for equally serious outcomes).
- Baseline risk matters. The number needed to treat depends on the baseline risk as well as on how much the treatment cuts it, so a treatment that halves risk gives a larger number in a low-risk group: halving a risk of 4 percent leaves 2 percent (a reduction of 0.02, number needed to treat 50), whereas halving a risk of 1 percent leaves 0.5 percent (a reduction of 0.005, number needed to treat 200). A large number therefore does not always mean a poor drug; it can also reflect a low-risk group.
Definitions and formulas
- Practical features of a screening test. A screening test is applied to many healthy people, so it must be simple to give, safe and inexpensive.
- Expected false positives. Expected false positives = (1 − specificity) × number of people without the disease; true positives = sensitivity × number with the disease. For example, with specificity 90 percent among 5,000 healthy people, 10 percent test positive, giving 500 false positives.
- Positive predictive value in screening. Positive predictive value = true positives ÷ (true positives + false positives). When prevalence is low, false positives outnumber true positives even for an accurate test, so the positive predictive value is low; for example 50 true positives and 450 false positives give 50/500, or 10 percent.
- Length-time bias in a program. Screening catches more slow, favorable conditions, so screened patients look better even if screening changed nothing. A program that finds mostly slow-growing cancers, with screened patients apparently outliving unscreened patients whose cancers were found by symptoms, illustrates length-time bias.
10 · Research & Trials
Megan B. Finck, MMS, PA-C · Session 8 · Fletcher ch. 9
Instructional Objectives
- Discuss following concepts and how they apply to medical research: Hypotheses, Sampling, Treatment, Comparative effectiveness, Clinical trials, Inclusion and exclusion criteria, Placebo, Randomization, Compliance, Blinding, Efficacy and effectiveness
- Describe randomized controlled trials and the advantages and limitations of this type of study design in medical research
- List considerations for treatments or interventions that undergo controlled trials
- Describe observational studies and the advantages and limitations of this type of study design in medical research
IO 1 — Testing ideas: coincidence, bias and wishful thinking
Decisions about treatment need to avoid resting on coincidence, biased comparisons and wishful thinking. The better basis is evidence-based information on the relationships between treatments and outcomes.
IO 1 — Treatment and treatment hypotheses
A treatment is an intervention to improve the course of disease. It may be applied at any point in the natural course of disease, and it may be a medication, surgery, a lifestyle change and so on.
A treatment hypothesis may develop from any of four sources:
- the mechanism or etiology of the disease
- clinical observation
- epidemiologic studies
- trial and error
However a hypothesis arises, most require testing to evaluate their usefulness and outcomes. A plausible mechanism, wide use or a clinician's approval does not replace that testing.
Also tested
- A promising hypothesis. That histamine mediates inflammation in allergic rhinitis supports a promising hypothesis that antihistamines will relieve symptoms, but the hypothesis must be tested by studying effects on symptoms in patients with allergic rhinitis.
IO 1 — Comparative effectiveness
Comparative effectiveness is the direct comparison of two or more interventions believed to improve outcomes. The comparison identifies any harms and benefits and compares the end results.
Treatments can be studied with observational and experimental designs. Cohort studies are limited by systematic differences in the groups other than the treatment factor. The randomized controlled trial is the gold standard for evaluating treatment efficacy.
IO 1 — Sampling
Sampling is the process of selecting participants for a study, trying to represent the larger population. An entire population is rarely studied, which is why a sample is used. The population has to be defined, and defining it too wide or too narrow causes problems.
| Method | What it is | Points |
|---|---|---|
| Random sampling | All members of the population have the same chance of being selected | The best way to get a representative sample; randomization can be done manually or by computer |
| Stratified sampling | Subgroups are created to represent how the groups exist in the population (by gender, age, education level and so on) | Can also provide equal groups that are not equal in the population, for example by diagnosis or diagnostic category |
| Cluster sampling | Clusters of participants are selected based on membership in a group | Examples: all patients in designated hospital units, all seniors participating in support groups, all students in designated special education classrooms |
Mirroring the population. If a clinic population is 60 percent women and a stratified sample of 100 is meant to mirror it, the sample holds 0.60 × 100 = 60 women.
Also tested
- Narrowing the study population. Doing so may curtail generalizability. The trade is often worth making, as when a study is limited to those in a service region who need the service, but the limit has to be acknowledged.
IO 1 — Inclusion and exclusion criteria
- Inclusion criteria are the characteristics participants must have.
- Exclusion criteria are the factors that prevent inclusion.
| Trial of a diabetes drug | Kind of criterion |
|---|---|
| Diagnosis of type II diabetes | Inclusion |
| History of unstable blood glucose levels | Inclusion |
| Comorbidities | Exclusion |
| Contraindication to the treatments under evaluation | Exclusion |
The same logic applies to any drug: having the disease being studied, or poor control of it, is an inclusion criterion, while a contraindication to the treatment under evaluation (an allergy to an ingredient, for example) is an exclusion criterion.
Each additional set of criteria shrinks the eligible pool. A published chart of patients with chronic obstructive pulmonary disease shows the eligible share falling step by step, as the percentage of all patients still eligible after each set of criteria:
| Criteria applied | Share of all patients still eligible |
|---|---|
| All patients | 100% |
| Pulmonary function tests not within range | 65.8% |
| Other inclusion and exclusion (symptoms, smoking) | 32.9% |
| Other aspects of the disease (oxygen, oral steroids) | 28.2% |
| Other lung disease (asthma) | 20.0% |
| Other medical conditions | 13.2% = final eligibility |
The pool shrinks with every added set of criteria, sometimes to a small fraction of all patients with the condition. To work out how many remain, multiply through the steps in sequence. For example, 1,000 patients with 50 percent meeting a first criterion and 40 percent of those meeting a second leave 1,000 × 0.5 × 0.4 = 200.
Also tested
- Strict inclusion and exclusion criteria. Strict inclusion and exclusion criteria make trial patients more alike, so a real difference between treatments is easier to detect. This improves internal validity but reduces generalizability.
IO 1 — Comparison groups, placebo and the Hawthorne effect
Comparison groups are the standard in trials. They may be treatment versus placebo, treatment versus treatment, or treatment versus non-treatment.
- A placebo is indistinguishable from the treatment but has no specific mechanism of action. It helps blinding.
- The placebo effect is the change in condition experienced by individuals given a placebo because they expect a change.
- The Hawthorne effect is the tendency of participants to change their behavior when given special interest or attention during a study.
IO 1 — Randomization
Random sampling and random assignment are different steps. Sampling picks who enters the study from the population. Assignment picks which group each of those subjects joins. Random assignment is preferred for testing a treatment because it limits systematic differences between the groups other than the treatment factor.
Also tested
- Randomized comparisons. Only characteristics that exist at the time of randomization, such as disease severity, are randomly allocated, so comparing the severely affected patients in each group remains a randomized comparison. Characteristics arising afterward (retention, response, compliance) are not.
- Per-protocol analysis. In a per-protocol analysis, outcomes are attributed to the treatment patients actually received rather than the group they were randomized to, which is what an intention-to-treat analysis uses. Who takes treatment is no longer random, so it loses randomization's protection against confounding.
- Baseline differences between groups. Baseline differences between randomized groups can result from chance (bad luck in randomization) or from a breakdown in allocation concealment. Small differences are expected.
Stratified randomization
In a stratified randomized trial the eligible patients are first divided into strata (for example strata 1, 2 and 3). Randomization to treatment or control takes place separately within each stratum. The final study groups are then formed by combining the treatment patients from every stratum into one group and the control patients into the other.
Also tested
- Stratified randomization and prognosis. Stratified randomization is particularly helpful when a baseline variable is strongly related to prognosis, especially in smaller trials, where simple randomization might leave groups with different prognosis.
IO 1 — Compliance
Compliance is the degree to which participants act in accordance with instructions (or patients with medical advice). Noncompliance is not always intentional; unintentional causes include misunderstandings, lack of transportation and forgetfulness. Deliberately skipping doses or stopping a drug by choice is intentional noncompliance.
IO 1 — Blinding
Blinding (also called masking) means making participants unaware of which treatment is being received.
| Term | Who is unaware |
|---|---|
| Single-blind | The subjects are unaware; the researchers are aware |
| Double-blind | The researchers and the subjects are both unaware |
| Allocation concealment | The researchers who assign patients to treatment groups are blinded, so they cannot steer which group a patient enters |
| Open label | No blinding is done |
Also tested
- Blinding and side effects. Side effects such as a slower pulse and lethargy can reveal who is taking the active drug to patients, to physicians, and to anyone who observes patients after treatment starts, but they cannot affect allocation concealment, because assignment is made before treatment begins.
- Blinding in randomized controlled trials. Single or double blinding is possible. It protects against performance bias, because participants who know they are on placebo may seek other treatment.
IO 1 — Efficacy and effectiveness
| Efficacy | Effectiveness | |
|---|---|---|
| Circumstances | Ideal trial circumstances | Real, ordinary circumstances |
| Takes into account | — | Compliance, practicality and cost |
A drug can therefore have high efficacy but low effectiveness: it works when everyone takes every dose under close monitoring, but poorly in clinics where doses are missed.
- Internal validity — the degree to which results are correct for the sample of patients studied.
- Generalizability — the degree to which results can be applied to other settings.
Also tested
- Intention-to-treat analysis. Patients are analyzed by the group they were randomized to, so it shows the consequences of offering a treatment, whether or not it is taken (effectiveness in ordinary circumstances), not the effect of actually taking treatment.
- Pragmatic trial. A pragmatic trial tests treatments under real, ordinary circumstances (broad eligibility, usual care, no blinding after enrollment, patient-centered outcomes), in contrast to an efficacy trial, which tests them under ideal circumstances.
IO 2 — Design of a randomized controlled trial
A randomized controlled trial is built in this order:
- Start with a research question about a population.
- Draw a sample of people (subjects) from the population to be in the study.
- Randomly assign the subjects to treatment groups.
- Follow both groups forward in time and measure the outcome of interest.
- Calculate the risk of the outcome in each of the treatment groups.
- Compare risk between groups using the relative risk, the risk difference or the hazard ratio.
Also tested
- Composite primary outcome. A primary outcome that is a composite of clinically important, related outcomes produces more outcome events than any single outcome, which increases the trial's ability to detect an effect if one is present.
Comparing risk between groups
| Measure | How it is worked out |
|---|---|
| Risk in a group | Patients with the outcome ÷ patients in that group (a count of events is not itself a risk) |
| Relative risk | Risk in the treatment group ÷ risk in the comparison group |
| Risk difference | Risk in the comparison group − risk in the treatment group, in percentage points |
Example (invented numbers). If 15 of 100 treated patients and 30 of 100 comparison patients had the outcome, the risks are 15 percent and 30 percent, the relative risk is 0.15 ÷ 0.30 = 0.5, and the risk difference is 15 percentage points.
Also tested
- Relative risk reduction. The relative risk reduction is the absolute risk reduction divided by the risk in the control group, which equals 1 minus the relative risk. Here it is 0.0026/0.0079, about 33%.
IO 2 — Superiority and non-inferiority trials
- Superiority trials determine whether one treatment is better than another.
- Non-inferiority trials determine whether a new treatment is unlikely to be less effective than an established treatment. The new treatment may be cheaper, safer or easier.
- Non-inferiority trials usually require a larger sample size than superiority trials, because the differences may be small.
Also tested
- Noninferiority margin. The margin is the largest difference judged clinically unimportant. When the true difference is significantly less than the margin, the new treatment is unlikely to be worse, so there is no clinically important difference.
IO 2 — Advantages and limitations of randomized controlled trials
Advantage: random assignment limits systematic differences between groups, which is why the randomized controlled trial is the gold standard for evaluating treatment efficacy.
Limitations:
- Confounding may exist in smaller randomized trials, because chance imbalances between the groups are more likely when few participants are enrolled.
- They are expensive.
- They raise logistical issues: sample selection, randomization, blinding, consistency with all participants, and time.
Also tested
- Randomized trials and confounding. The main advantage of randomized controlled trials over observational studies of treatment effects is prevention of confounding, which is why they are the gold standard for evaluating treatment efficacy.
- Trials cannot answer everything. Randomized controlled trials are the strongest evidence for cause and effect, but many suspected causes cannot be studied with one, because experimental studies of harm are limited by ethics and willingness to participate.
- Observing a rare side effect. To have a good (about 95 percent) chance of observing at least one event, the number of people observed should be about 3 divided by the event rate. With about 15,000 people receiving chemoprevention (half of 30,000), a side effect occurring in about 1 in 5,000 can be detected.
IO 3 — Is a trial ethical?
- Randomization is ethical only when no conclusive evidence favors one treatment: there is no reason to believe the treatments under evaluation are unequal. Once conclusive evidence exists that a treatment prevents death, for example, there is no real uncertainty left and assigning patients to placebo would be unethical.
- The primary outcome must be benefits, not harms.
- Participants must give consent.
- If there is evidence of harm, futility or effectiveness, the study is ended.
IO 3 — Methodological issues and possible bias
Five issues can bias a controlled trial:
- Participants without the disease in question being enrolled.
- Compliance problems.
- Cross-over: switching from one randomly assigned treatment to another, so participants no longer stay in the group to which they were assigned.
- Unacknowledged co-interventions: an extra treatment that goes unrecorded and can affect the outcome.
- The assessment of the outcome may be biased.
Even with blinding it may be possible to determine the treatment group, because side effects (and similar clues) can reveal it. Biased assessment of outcomes is avoided by using specific objective criteria and blinding.
The four possible outcomes of a study
| Conclusion ↓ / Reality → | Treatments are NOT different | Treatments ARE different |
|---|---|---|
| Study concludes not different | Correct | ERROR — a real difference was missed |
| Study concludes different | ERROR — a difference was found that does not exist | Correct |
Two of the four pairings are correct and two are errors. A conclusion is correct only when it matches reality; blinding does not turn a wrong conclusion into a correct one.
IO 3 — Trials of N = 1 (single subject design)
A rare type of clinical trial in which one patient is given one treatment at a time. The provider and the patient are blinded to which treatment is given when, and the outcomes are assessed.
IO 3 — Phases of clinical trials
| Phase | Size and design | Aim |
|---|---|---|
| Phase I | Small number of participants; no control group | Identify a safe, tolerated dose |
| Phase II | Small number of participants | Identify a preliminary effect and any relationship to dose |
| Phase III | Larger number of participants than phases I and II; randomized; often published | Identify definitive evidence of effect and some common side effects; it does not usually identify uncommon side effects |
| Postmarketing surveillance | Following large numbers of patients after the drug is released | Helps identify more uncommon side effects |
A side effect that occurs in about 1 in 20,000 patients is far too rare to show up in the small numbers of phases I and II or usually even in phase III; it is postmarketing surveillance of large numbers that detects it.
Also tested
- Phase III trials. Phase III trials are larger randomized trials that give definitive evidence of efficacy or effectiveness and identify some common side effects, but usually do not identify uncommon side effects.
IO 3 — Evaluating a treatment study
- Is it a controlled trial with random assignment?
- Are the study patients similar to yours? (This tells you whether the results apply to your patient.)
- Was there complete follow-up? Look at best- and worst-case scenarios for dropouts.
- Was blinding used?
- Are the intervention and control groups similar?
- Is there clinical significance as well as statistical significance?
- Are other factors affecting the outcome considered?
Best- and worst-case scenarios for dropouts. The analysis assumes the best and worst possible outcomes for the patients lost to follow-up and checks whether the study’s conclusion would change. It shows whether the conclusions survive the lost patients.
Example (invented numbers). In a treated group of 100 patients, 12 were lost and 25 of the 88 who finished had the outcome. If every lost patient is assumed to have had the outcome, the treated group’s risk is (25 + 12) ÷ 100 = 37 percent, not 25 ÷ 88 = 28 percent.
IO 4 — Observational studies: advantages and limitations
Observational studies of treatment collect information on patients currently being treated for a disorder, instead of assigning the treatment.
- Advantage: they include real-world patients, which brings both a benefit and drawbacks.
- Limitation: it is difficult to avoid confounding and effect modification.
- Their findings often agree with the findings of randomized controlled trials, but can be dramatically different.
- It is difficult to draw conclusions from uncontrolled trials, because there is nothing to compare the treated patients against.
| Term | Meaning |
|---|---|
| Confounding | An additional independent factor related to both the exposure and the outcome |
| Effect modification | The presence or absence of a variable changes the effect of the exposure |
Definitions and formulas
- Efficacy. Efficacy is concerned with results under ideal trial circumstances, whereas effectiveness reflects ordinary, real circumstances.
- Compliance. Compliance is the degree to which participants act in accordance with instructions, or patients with medical advice.
- Unintentional noncompliance. Noncompliance is not always intentional: misunderstandings, lack of transportation to visits and forgetfulness are unintentional causes.
- Sources of treatment hypotheses. Treatment hypotheses come from the mechanism or etiology of the disease, clinical observation, epidemiologic studies, and trial and error.
- Consent. Participants must give consent to take part in a trial; consent is one of the conditions for an ethical trial, alongside genuine uncertainty about which treatment is better.
- Basis of treatment decisions. Treatment decisions should rest on evidence about the relationship between a treatment and its outcomes, rather than on coincidence, biased comparisons and wishful thinking.
- Limitations of randomized controlled trials. They are expensive, and they raise logistical issues in sample selection, randomization, blinding and consistency across participants and time.
- Risk in a group. Risk = number of patients with the outcome ÷ number of patients in the group. For example, 30 of 200 is 0.15, or 15 percent.
- Relative risk. Relative risk = risk in the treated group ÷ risk in the comparison group. For example, risks of 0.05 and 0.20 give a relative risk of 0.25.
- Risk difference. Risk difference = risk in the comparison group − risk in the treated group, in percentage points. For example, risks of 30 percent and 18 percent give a difference of 12 percentage points.
- Eligibility criteria shrink the pool. Each eligibility criterion removes patients, so the pool shrinks step by step: multiply the number screened by the proportion meeting each criterion in turn. For example, 1,000 screened, with 50 percent meeting the first criterion and 40 percent of those meeting the second, leaves 200.
11 · Statistics, Correlation & Causation
Megan B. Finck, MMS, PA-C · Session 9 · Fletcher ch. 11-12
Instructional Objectives
- Explain the following statistical concepts: Significance, P value, Null hypothesis, Confidence interval, Statistical power, Multiple comparisons, Multivariable modeling
- Compare and contrast statistical significance and clinical importance
- Define criteria to determine causation
- Discuss the impact of study design on the ability to determine causation
Objective 1 — Review of probability
Probability is the likelihood that any one event will occur, given all the possible outcomes. The probability (P) that an event (E) will occur is estimated by:
The denominator is every opportunity for the event, including the ones where it occurred. Dividing by the number of times the event does not occur gives the odds, which is a different quantity, and inverting the fraction gives a value above one, which cannot be a probability.
- Probability must lie between 0 and 1. A value of 0 means the event cannot occur; a value of 1 means it will definitely occur.
- A value of 0.5 means an equal probability of occurring or not occurring, so the event is as likely to occur as not.
- The sum of all probabilities of all events in the sample must equal 1. So the probability that an event does not occur is 1 minus the probability that it does: if the probability of having a condition is 0.35, the probability of not having it is 0.65.
- Counting works the same way: if 36 of 240 patients in a clinic have a condition, the probability that a randomly chosen patient has it is 36 ÷ 240 = 0.15, and the probability of not having it is 204 ÷ 240 = 0.85.
Also tested
- Bayesian reasoning. Bayesian reasoning is about how new information changes a prior belief. The results of each new study change belief in a treatment effect from what it was before the study, just as a test result changes pretest probability into post-test probability.
Objective 1 — Significance and the p value
Two kinds of significance matter when reading a study: statistical significance and clinical significance.
Statistical significance describes how likely it is for the result to have happened by chance, so it describes how unlikely chance is as the explanation. More statistically significant = less likely to be chance. Less statistically significant = more likely to be chance. It does not describe how much the treatment helps an individual patient, how carefully the study was designed, or whether the result would change practice.
The p value is the probability of an observed difference happening by chance. Read precisely, it is calculated on the assumption that no true difference exists: it is the probability of a difference at least this large arising by chance alone. It is not the probability that the null hypothesis is true, not the probability that the treatment truly works, not the proportion of patients who respond, not the size of the difference between groups (the effect size), and not the probability that the study is biased. Bias is systematic error, and a p value addresses random chance only.
| p value | Chance of a result this extreme | How it is generally described |
|---|---|---|
| 0.5 | 50/50 | Chance is as likely as not |
| 0.05 | 1/20 | Values less than 0.05 are generally accepted to be statistically significant |
| 0.01 | 1/100 | Generally accepted to indicate highly significant |
| 0.001 | 1/1000 | Generally accepted to indicate very highly significant |
- Reading a p value of 0.03 (new drug against placebo): 0.03 is below 0.05, so the difference is statistically significant. It does not mean there is a 3 percent chance the drug does not work, and it does not mean the drug works in 97 percent of patients. Statistical significance also does not imply that the difference is large enough to change practice.
- Reading a p value of 0.20: 0.20 is above 0.05, so the result is not statistically significant and chance remains a reasonable explanation. It does not show that the treatment has no effect (the study may lack power, see below), it is not the proportion of patients who improved, and it is nowhere near very highly significant.
- A lower p value shows only that chance is less likely as an explanation. It does not show that the difference is larger or more important.
Also tested
- Meaning of the P value. The P value is the probability that a difference as big as or bigger than the one observed could have arisen by chance alone; P < 0.001 means less than 1 chance in 1,000. It says nothing about bias, importance, or generalizability.
Objective 1 — The null hypothesis
The null hypothesis is the hypothesis that there is no difference. It is usually the opposite of what the researchers are interested in finding, and it is used in statistical analysis: the test asks how well the data fit it. The hypothesis that a difference is present is the alternative to the null hypothesis. The null hypothesis is not a statement about bias or about whether the sample is representative.
- Example. Researchers hope to show that a new drug lowers blood pressure more than placebo. The null hypothesis is that the drug and placebo lower pressure equally. Drug better than placebo is the researchers' hoped-for alternative, and placebo better than drug states a difference in the wrong direction.
- If there is not evidence to support the null hypothesis, there is more support for a difference being present. Evidence against the null hypothesis, which states no difference, is what lends support to a difference. Identical groups is what the null hypothesis itself states, so weak support for it points away from that idea.
Objective 1 — Type I and type II error (four possible statistical conclusions)
| Conclusion of the statistical test | True difference | |
|---|---|---|
| Present | Absent | |
| Significant | Correct | Type I (α) error |
| Not significant | Type II (β) error | Correct |
- Type I error (α error, also called an alpha error) = false POSITIVE — stating there is a difference, or concluding that a relationship exists, when there is not. Most medical research is concerned with the possibility of type I error, which is expressed by the p value (judged against the chosen threshold, usually 0.05).
- Type II error (β error, also called a beta error) = false NEGATIVE — stating there is no difference, or concluding that no relationship exists, when there is one. The real difference is missed and the statistic appears not to be significant. (The precise wording is that the test failed to detect a real difference.)
- Type II error sometimes occurs as a result of study limitations, such as length of follow-up and subject compliance: too brief a follow-up, or subjects not adhering to treatment, can hide a real difference. A p value below 0.001 is a significant result, so it cannot be the cause of a false negative, and random assignment and blinding are strengths of design rather than causes of type II error.
- When no true difference exists, the correct conclusion is that the result is not significant. A significant result would be a type I error, the null hypothesis is true in that case (it should not be rejected), and power is defined only for the case where a difference does exist.
- Examples. A trial concludes that a supplement lowers cholesterol but in truth it has no effect: type I error (a false positive). A short trial finds no significant benefit from a drug that truly works: type II error (a false negative). Measuring the outcome incorrectly is measurement bias and choosing the wrong test is a different problem; neither is what defines type I or type II error.
Objective 1 — Common statistical tests
| Test | When used |
|---|---|
| To test the statistical significance of a difference | |
| Chi square | Between two or more proportions (when there is a large number of observations) |
| Fisher exact | Between two proportions (when there is a small number of observations) |
| Mann-Whitney U | Between two medians |
| Student t | Between two means |
| F test | Between two or more means |
| To describe the extent of association | |
| Regression coefficient | Between an independent (predictor) variable and a dependent (outcome) variable |
| Pearson's r | Between two variables |
| To model the effects of multiple variables | |
| Logistic regression | With a dichotomous outcome |
| Cox proportional hazards | With a time-to-event outcome |
| Propensity score | Select matching controls |
Means are averages of continuous values, medians are middle values, and proportions are shares of a group. The Fisher exact test takes over from the chi-square test when the number of observations is small.
Objective 1 — Confidence intervals
A confidence interval expresses statistical precision. It is a range of scores with boundaries, or confidence limits, that should contain the true value. A 95% confidence interval means that there is a 95% chance that the confidence interval contains the true value; put another way, if the study were repeated many times, about 95 percent of such intervals would contain the true value.
- A narrow interval means greater statistical precision; a wide interval means a less precise one. Two studies can report the same point estimate with different precision: a relative risk of 0.6 with an interval of 0.5 to 0.7 is more precise than the same 0.6 with an interval of 0.2 to 1.8.
- An interval is not the range that holds 95 percent of patients' individual values, it is not a range for the p value, it does not measure or correct bias (an interval can be narrow and still be wrong if the study is biased), and it is not the same as clinical importance or statistical power.
Reading an interval around a relative risk or a difference
For a relative risk, 1 means no difference, and values below 1 mean the risk is lower in the treated or exposed group. In a plot of relative risk the outcomes with intervals lying entirely on one side of 1 are statistically significant, and an interval that spans 1 is not. For a difference (for example in mean weight loss) the value of no effect is 0.
| Result | Reading |
|---|---|
| Relative risk 0.70, 95% confidence interval 0.55 to 0.90 | Whole interval below 1: a statistically significant reduction in risk |
| Relative risk 1.2, 95% confidence interval 0.8 to 1.8 | Interval contains 1: no statistically significant difference. The point estimate suggests a 20 percent increase, but an increase is not established because the interval runs from below 1 to well above it, which also makes the estimate imprecise |
| Difference in weight loss, 95% confidence interval 1.0 to 3.5 kilograms | Interval entirely above 0: statistically significant |
| Difference in weight loss, 95% confidence interval -0.5 to 3.0 kilograms (or -2.0 to 1.5, or -1.5 to 0.5) | Interval contains 0: not statistically significant |
A 95% confidence interval that excludes the value of no effect corresponds to a p value below 0.05 for the same comparison, and an interval that contains it corresponds to a p value above 0.05. The interval adds what the p value alone cannot: the range of plausible effect sizes.
Worked graph (systematic review of corticosteroids for pneumonia). Each outcome is drawn as a point (the relative risk) with a line (its confidence interval) on a scale from 0 to 2, with 1 as the line of no effect. Mortality and mechanical ventilation lie to the left of 1 with intervals that do not reach 1 (corticosteroids benefit). Hyperglycemia lies to the right of 1 with an interval that does not reach 1 (corticosteroids harm). Gastrointestinal bleed has a point just under 1 and a very wide interval that crosses 1, so it is imprecise and not significant.
Also tested
- Wide interval, low power. A relative risk interval of 0.6 to 1.5 is wide, running from clear benefit to clear harm, which shows imprecision and too little statistical power. A study with adequate power would give a narrower interval, so a null result cannot be trusted as showing no association.
Objective 1 — Statistical power
Statistical power is the probability that a study will find a statistically significant difference when the difference really exists. It is calculated in terms of type II error:
- Statistical power is high when studies have a high probability of detecting true differences. Power is the complement of the type II error probability, not the type I error probability, not the p value, not the size of the difference and not the chance that the null hypothesis is true.
- Calculation. A study with a 20 percent (0.20) chance of a type II error has power 1 − 0.20 = 0.80, or 80 percent. A study with power 0.90 has a type II error probability of 1 − 0.90 = 0.10. Ninety-five percent is the level used for confidence intervals and 5 percent (0.05) is the usual threshold for type I error; neither is derived from the type II error probability.
- Low power means a high probability of type II error, so a real difference is likely to be missed. A small trial of a drug that truly works can therefore report a non-significant p value (for example 0.30) simply because it had too little power. A non-significant result does not prove there is no effect and does not prove the null hypothesis true; it fails to reject it. That is a type II error, not a type I error.
- Larger samples give more power, which is why pooling data across studies in a meta-analysis usually helps.
Also tested
- Determinants of statistical power. Statistical power depends mainly on effect size, event rate, the type I error level and the type II error level, which together set sample size. The statistical test used matters least; it changes calculated power only slightly.
Objective 1 — Multiple comparisons
The strength of statistical tests depends on the number of research questions in a study and how focused those questions are. It is not always possible to know how many comparisons were made. Selecting more interesting comparisons after the analysis is complete can distort the data. When many comparisons are made in one study, some may look significant by chance alone; more comparisons do not make any p value more accurate, do not make type II error impossible and do not narrow confidence intervals.
- Expected chance findings. A p value of 0.05 corresponds to a chance of 1 in 20. If a study makes 60 comparisons and none reflects a true difference, about 60 ÷ 20 = 3 will appear significant by chance. At a p value of 0.01 (1 in 100), 200 comparisons with no true difference give about 200 ÷ 100 = 2 chance findings. Using the 0.05 threshold instead of 0.01 would predict 10, which is the wrong rate.
| How multiple comparisons can be misleading (Fletcher Table 11.3) | |
|---|---|
| 1 | Make multiple comparisons within a study |
| 2 | Apply tests of statistical significance to each comparison |
| 3 | Find a few comparisons that are “interesting” (statistically significant) |
| 4 | Build an article around one of these interesting findings |
| 5 | Do not mention the context of the individual comparison (how many questions were examined and which question was considered primary before the data were examined) |
| 6 | Construct a post hoc argument for the plausibility of the isolated finding |
The practice that misleads is picking interesting results after analysis and hiding how many were examined. Naming the primary question beforehand, reporting every comparison that was made, and limiting the study to a few focused questions all protect against the distortion.
Also tested
- Subgroup analyses. Subgroup analyses risk false-positive and false-negative conclusions because they are multiple comparisons in smaller groups. They still give clinicians useful information when their limitations are kept in mind, so they should not be banned.
Objective 1 — Multivariable modeling
Multivariable modeling examines the effect of two or more variables together, as opposed to a single variable in isolation. It is a mathematical construct (an equation), and it is used in two ways:
- To look at the effect of one variable on the outcome while taking the effect of other variables into account (confounding or effect modifiers), which allows one variable's effect to be isolated.
- To look at the combined effect of multiple variables acting together to predict a clinical event.
Also tested
- Advantage of multivariable modeling. Multivariable modeling lets investigators examine the effect of one variable on the outcome while taking the effects of many other variables, confounders and effect modifiers, into account at the same time.
- Risk factors in a prediction model. A multivariable prediction model keeps the variables that each improve prediction of the outcome, so those variables are the risk factors for it. Variables that do not affect the prediction are not.
Process
- Identify variables related to the outcome (confounding, effect modifiers).
- Select variables for inclusion: consider those with the strongest association with the outcome first (not those least associated, easiest to measure or with the most missing data).
- Identify assumptions. They commonly involve probability distributions, with underlying assumptions such as normality, homoscedasticity and independence of errors.
- Create the model.
Limitations
- Validity is based on assumptions about the data: a model is valid only if its assumptions hold. The number of patients recruited affects precision and power, and adding more variables or requiring every variable to be significant does not secure validity.
- Poor at identifying effect modification. (Taking confounders into account is what a model is designed to do.)
- May omit rare variables with a strong relationship with the outcome.
- A model cannot prove that the outcome is caused by a variable, cannot remove random error (which is estimated statistically and never eliminated) and does not replace the need for a comparison group.
Objective 2 — Statistical significance versus clinical importance
Statistical significance asks how likely chance is as an explanation; clinical significance (importance) asks whether the difference is relevant to practice and to patients. A short contrast: chance versus relevance to patients. (Random error versus systematic error, precision versus accuracy, and validity versus reliability are different pairs and do not contrast the two kinds of significance.)
- Statistical significance does not imply clinical significance. Differences can be present with a p value below 0.0001 and still not have any relevance to clinical practice. So a very small p value tells you nothing about relevance: it does not make a difference relevant, and it does not make it irrelevant either.
- Larger differences tend to be more clinically significant, but small differences can still be statistically significant (especially in very large samples). A lower p value is not the same as a larger difference, and sample size affects statistical significance and precision rather than how much a difference matters clinically.
- The reverse also happens: an important difference can fail to reach statistical significance, for example in a small study with low power.
- Question test. “Would it matter to the patient?” is a clinical question. “Is the p value below 0.05?”, “Does the interval exclude no effect?” and “Was the null hypothesis rejected?” are all statistical questions.
| Situation | Interpretation |
|---|---|
| 40,000 patients; systolic blood pressure lowered by 1 millimeter of mercury; p value below 0.0001 | Statistically significant but of doubtful clinical importance: a fall this small is unlikely to matter to patients |
| Small trial; 30 percent lower death rate; p value 0.12 | Possibly important but not statistically significant; 0.12 does not reach 0.05, and a small study may simply lack power, so it does not prove the treatment has no effect |
Objective 3 — The study of causation and the types of cause
Causality exists if a factor produces or contributes to the production of a specified outcome. It is not the same as association (any factor associated with an outcome), nor as merely occurring before an outcome, nor as having a significant p value.
- Single causes, multiple causes, and the web of causation (many factors acting together).
- Koch's postulates for infectious disease are the example of a single cause: the organism must be present in every case of disease; must be isolated and grown in pure culture; must cause a specific disease; and must be recovered from an inoculated animal and be identified.
| Type of causal relationship | Definition |
|---|---|
| Sufficient cause | Precedes the disease and, if present, the disease always occurs (it guarantees the disease whenever present) |
| Necessary cause | Precedes the disease and, if absent, the disease cannot occur. If the cause is present, disease may or may not occur (necessary but not sufficient) |
| Risk factor | Increases the probability of a disease in a group of persons who have the factor; neither sufficient nor necessary. It makes disease more likely without guaranteeing it, and disease can still occur without it |
| Direct causality | The factor under consideration exerts its effects without intermediate factors |
| Indirect causality | One factor influences one or more other factors (intervening variables) that are in turn directly causal. Example: cardiovascular disease associated with obesity is in part due to the higher blood pressure and higher levels of cholesterol and glucose that often occur with obesity. Likewise, poverty raises tuberculosis risk partly through crowded housing and malnutrition. Intervening variables lie on the pathway between the cause and the outcome, unlike a confounder, which is a separate factor associated with both exposure and outcome |
| Non-causal association | The relationship between the two variables is statistically significant, but no causal relationship exists (the temporal relationship may be incorrect, and a shared factor such as summer weather can drive both). Example: ice cream sales and drowning deaths both rise in summer and the link is statistically significant, yet neither causes the other |
Applying the definitions. If disease X never occurs without exposure Y but many exposed people stay healthy, Y is a necessary but not sufficient cause. Y would have to always produce disease when present to be sufficient.
Also tested
- Multiple interacting causes. Most diseases, such as coronary heart disease, have multiple interacting causes, a web of causation with several risk factors acting together. A risk factor is neither necessary nor sufficient by itself.
From association to cause
Before concluding that a statistical association is causal, three alternative explanations must be excluded, taken in this order: bias in selection or measurement first, then chance, then confounding. If bias is likely, or chance is likely, or confounding is present, the association is not causal; only when none of them explains it is a cause likely. (Dose-response, reversibility, specificity, temporality, strength and analogy are criteria that support cause, not alternative explanations to be excluded, and precision, power and validity are properties of a study.) It is not possible to prove a causal relationship beyond all doubt; what is possible is to increase the evidence.
Objective 3 — Criteria for causality
| Criterion | Meaning | Example |
|---|---|---|
| Temporality | The cause precedes the effect | Exposure comes before the disease appears |
| Strength of the association | Judged by the size of the relative risk: a large relative risk is a strong association (it is not judged by how many studies report it, by whether a mechanism is known, or by whether exposure can be reversed) | A very high relative risk |
| Dose-response | Larger exposure → higher rate of disease | Lung cancer rates rise with the number of cigarettes smoked each day |
| Reversibility | Decreased exposure → less disease | Lung cancer risk falls in people who stop smoking |
| Consistency | Findings are replicable in other studies | Several studies in different populations find the same association |
| Biologically plausible | Makes sense biologically | A known mechanism explains how the exposure could injure tissue |
| Specificity | One cause leads to one effect. It is a demanding criterion, because a single exposure often produces several effects and a cause need not be specific to be real | — |
| Analogy | Similar cause and effect pattern for a similar exposure and disease | A newly studied virus produces a disease pattern like that of a well-established similar virus |
Also tested
- Weight of specificity. Specificity (one cause leading to one effect) counts for less than the other criteria. An exposure associated with many different cancers is only weakly inconsistent with cause, because many established causes, such as cigarette smoking, produce several diseases.
- Statistical significance is not a causal criterion. Statistical significance only shows that an association is unlikely to be due to chance. The Bradford Hill criteria for causation are temporality, strength, dose-response, reversibility, consistency, biologic plausibility, specificity and analogy.
The criteria applied: Helicobacter pylori and duodenal ulcer
- Temporal: Helicobacter pylori is linked to chronic gastritis, which is associated with the development of duodenal ulcer.
- Strength: H. pylori is found in 90% of patients with duodenal ulcer. Dose-response: its concentration is higher in the mucosa of patients with duodenal ulcer than in ulcer-free patients.
- Replication: the observations have been repeatedly replicated. Biologic plausibility: there are binding sites for H. pylori on antral cells, and it can follow these cells into the duodenum.
- Alternative explanations: smoking increases risk but is not a risk factor in patients without H. pylori. Cessation of exposure: eradication of H. pylori results in healed ulcers.
- Specificity: prevalence in patients with duodenal ulcer is 90-100%, but the organism is also found in some asymptomatic patients and those with gastric ulcers, so specificity is only partly met.
- Consistency with other knowledge: prevalence is the same in men as in women, and duodenal ulcer prevalence peaked in the late 19th century when H. pylori was more prevalent because of poor living conditions.
Objective 3 — Pitfalls in causal research
| Pitfall | What it is |
|---|---|
| Measurement bias | Possible operator error, calibration problems and the like |
| Recall bias | Difficulty remembering: participants' difficulty in recalling past exposures distorts the reported exposure or history |
| Selection bias | How people get into the study, such as poor randomization |
| Random error | Occurs by chance and is accounted for through the p values used in the study. It is not handled by blinding, matching cases with controls, or excluding older participants: blinding protects against performance and measurement bias, and matching addresses confounding |
| Confounding | A covariate whose effect cannot be separated from the variable studied |
| Effect modification | The effect of an exposure differs according to another variable; a multivariable model is poor at identifying it |
Objective 4 — The impact of study design on determining causation
Evidence for or against cause depends both on the design of the study and on the findings. The hierarchy of designs for causal evidence runs from strongest to weakest:
- Systematic review
- Randomized controlled trial
- Multiple time series
- Non-randomized trial
- Cohort
- Case-control
- Time series
- Cross-sectional
- Case series
- Case report (the weakest, at the bottom of the hierarchy)
| Findings that argue FOR a causal relationship | Findings that argue AGAINST a causal relationship |
|---|---|
| Temporal sequence · small effect · specificity · analogy · biologic plausibility · consistency · large effect · dose-response · reversibility | Incorrect temporal sequence (the disease appears before the exposure) · no effect · not biologically plausible · no analogy · no dose-response · not reversible · not specific |
- The randomized controlled trial gives the strongest single-study evidence for causation (systematic reviews, which pool trials, sit at the very top with it). Random assignment balances confounders between the groups (it spreads them evenly), which is why it strengthens causal conclusions; it does not remove random error (still accounted for through p values), does not guarantee a significant result and does not make measurements more reliable.
- A cohort study ranks above a case-control study, which looks backward from the outcome. Cross-sectional studies, case series and a case report rank below it; a case series describes patients without a comparison group.
- A cross-sectional study is weak evidence for causation because exposure and outcome are measured together at a single point in time, so it cannot show which came first; the temporal sequence, an essential criterion for cause, cannot be established. It can still detect an association and does measure the outcome.
- Strong design plus a finding. A well-designed randomized trial that finds no effect of an exposure on an outcome provides strong evidence against a causal link (provided the trial had adequate power). A design at the top of the hierarchy makes the evidence strong rather than weak, and a finding of no effect argues against cause, however strong the design.
Also tested
- Time-series caution. In a time-series study, a change in disease rate after an intervention may be caused by other changes at about the same time, such as better antibiotics, so alternative explanations must be ruled out before concluding the intervention caused it.
Definitions
- Biologic plausibility. Biologic plausibility is met when the relationship makes sense biologically, for example when a known mechanism explains how an exposure could injure tissue; it strengthens the case that the relationship is causal.
- Evidence for causation. The more criteria for causation that are met, the stronger the case for a cause grows, though causation cannot be proved beyond all doubt.
12 · Clinical Questions & Reviews
Megan B. Finck, MMS, PA-C · Session 10 · Fletcher ch. 13–14
Instructional Objectives
- Discuss the advantages and limitations of narrative reviews
- Discuss the advantages and limitations of systemic reviews
- Identify the components of a PICO question: Patients, Intervention, Comparison, Outcomes
- Identify methods to conduct a systematic review
- Discuss publication bias
- Discuss the principles of meta-analysis
- Discuss the advantages and limitations of meta-analysis
- Discuss grading recommendations for treatment
- Discuss conflict of interest
- Identify resources for answering clinical questions
IO 1 — Narrative reviews
A narrative review is an expert summary of the evidence and recommendations. It is written by an author who chooses what to cite and how to describe it, with no set method for finding, choosing or appraising the studies.
| Advantages | Limitations |
|---|---|
|
It can cover broad topics and a range of issues. Its authors usually have experience caring for patients with the disease. |
The information may be influenced by the values and personal experience of the authors. It is unstructured, so its validity may be questionable. Cited articles may have been selected to prove a point of view. It may not include a critique of the original research. |
Where it fits. Because a narrative review is broad and unstructured, it suits an overview of many aspects of a disease. A question with a defined population, intervention and outcome — whether one drug lowers mortality more than another, whether a test detects a disease more accurately, or whether screening lowers deaths in one age group — is a focused question, and answering it is the job of a systematic review.
IO 2 — Systematic reviews: the question, the studies and how to appraise one
A systematic review is a review of evidence with the goal of answering a specific, predetermined clinical question. That is the contrast with a narrative review, which is broad and open ended: the question of a systematic review is specific and is set in advance, before the studies are found, so the answer cannot be shaped around what the reviewers happen to find.
Defining the question
The first task is to define a focused question. The considerations are:
- Are strong studies available?
- How many studies are available?
- Do all the studies agree?
Put together, whether a review question can be answered depends on whether strong studies exist and whether they agree. Wanting a particular answer, sharing a publisher, or the existence of earlier expert commentary does not show that answerable studies exist.
Summarizing the studies
Once studies are found, the review asks how many quality studies were found and what the results were. The results are summarized as effect size, precision and validity — both internal validity (are the conclusions valid for the sample studied) and external validity (do they hold in other settings).
Evaluating a systematic review article
A reader appraising a systematic review asks five things:
- How were the studies located?
- What inclusion and exclusion criteria were used?
- How valid were the original studies?
- What biases existed in the original studies?
- What were the results of the original studies, and was there variation among them?
The search and the criteria are read together with the validity of the original studies. The layout of tables, the length of the reference list and who drafted a section are not appraisal questions.
IO 3 — The four-part clinical question (PICO)
A focused question is organized into four components, abbreviated PICO: Patients, Intervention, Comparison and Outcomes.
| Letter | Component | The question it answers |
|---|---|---|
| P | Patient, population or problem — who | How would I describe a group of patients similar to mine? |
| I | Intervention, prognostic factor or exposure — what | Which main intervention, prognostic factor or exposure am I considering? |
| C | Comparison — what is the alternative | What is the main alternative to compare with the intervention? |
| O | Outcome you would like to measure | What can I hope to accomplish, measure, improve or affect? |
The intervention element is wider than a treatment. It can be a prognostic factor or an exposure, so the same format organizes risk and prognosis questions as well as treatment questions. The comparison is the alternative set against the intervention — another treatment, or none. The outcome is what is measured, and the patient element sets the group to which the answer will apply.
| Question | P | I | C | O |
|---|---|---|---|---|
| Vitamin C for a cold | Adults with a cold | Vitamin C | No vitamin C | Resolution of common cold symptoms |
| High fiber diet or regular exercise and colon cancer risk | Older adults | High fiber diet | Regular exercise | Reduction in risk for colon cancer |
| Virtual colonoscopy in ulcerative colitis | Patients with ulcerative colitis | Virtual colonoscopy | Standard colonoscopy | Sensitivity and specificity for detecting colon cancer |
| Angiotensin-converting enzyme inhibitors in congestive heart failure | Patients with congestive heart failure | Angiotensin-converting enzyme inhibitors | Patients who do not receive them | The death rate |
| Skin adhesives or sutures for superficial lacerations | All patients? A more specific population may be needed | Skin adhesives | Sutures | Closing superficial skin lacerations |
- A question comparing two diagnostic tests can name the accuracy of each test — sensitivity and specificity — as its outcome, as in the virtual colonoscopy row above.
- A question that is too broad needs refining. If a question is asked “for all patients”, the patient element is the one to make more specific; a more specific population may need to be defined.
- A complete question has all four parts. A question that omits the comparison does not say what the intervention is being measured against; one that names no group, or no measurable outcome (“better in some way”), cannot be answered from studies. A well-built example: in adults with migraine, does a daily preventive drug, compared with placebo, reduce the number of headache days?
Also tested
- What PICO covers. A specific clinical question is defined by PICO (Patients, Intervention, Comparison, Outcomes). Covariates are not part of a specific clinical question.
IO 4 — Methods for conducting a systematic review
| Step | Methodologic considerations |
|---|---|
| 1. Define a specific question | The PICO format organizes the question. A protocol identifies the question and the analysis prospectively. |
| 2. Select studies for the review | Find all relevant studies based on the study question and design. Determine the final studies from the initial list with explicit criteria based on the study question and study design. At least two reviewers confirm consistency of the selection process. List excluded studies and the rationale. Assess for publication and reporting bias (for example, funnel plots). |
| 3. Describe each study's characteristics and risk of bias | Scoring systems can be used. At least two reviewers confirm consistency of the assessment. |
| 4. Summarize the studies | Figures — forest plots. Tables of characteristics. |
| 5. Pooling (meta-analysis) if justified | Calculate the summary effect and confidence interval using a fixed effect model or a random effects model. On the forest plot the summary is shown as a diamond. |
| 6. Identify reasons for differences among studies (heterogeneity) if present | Statistical tests: Cochran's Q and the I-squared statistic. Subgroup analysis (based on study quality and other factors) and sensitivity analysis. Metaregression. Discussion in the text of the role of bias and study factors. |
Read the steps as an order: the question is defined first, the studies are then selected, described and summarized, pooling comes after the summary and only if it is justified, and the exploration of heterogeneity is last.
- Protocol. Writing the question and the planned analysis down before searching stops the reviewers from shaping either around the results they find. A forest plot is a figure made after analysis; a funnel plot checks for publication and reporting bias; sensitivity analysis explores differences after the results are in.
- Consistency. A single reviewer, the lead author alone, or the journal editor cannot show that the criteria were applied consistently; at least two reviewers do, both when selecting studies and when assessing each study's risk of bias.
- Excluded studies. The reason each was excluded is listed, so a reader can judge whether the criteria were applied fairly.
- Forest plots and tables summarize the studies. Funnel plots belong to the selection step, where publication and reporting bias are assessed. A nomogram and a receiver operating characteristic curve are tools for interpreting diagnostic tests, not for summarizing or checking a review.
- Methods that explore differences among pooled studies are subgroup analysis, sensitivity analysis and metaregression. Blinding and randomization protect a single trial from bias and do not explore differences between studies.
- Heterogeneity tests. Cochran's Q and the I-squared statistic quantify how much study results differ; survival-analysis tests, correlation coefficients and rank tests serve other purposes.
Finding studies
To sample the best studies, the reviewers ask:
- How strong was the study design?
- Was the sample size sufficient? (A trial with only a handful of participants is weighed against the others for exactly this reason.)
- What kind(s) of bias was present, and are there concerns for publication bias?
- What were the limitations?
- Was the research peer-reviewed?
- Is the information still current, or have more recent studies been published?
The reviewers also develop inclusion and exclusion criteria; the criteria are explicit and based on the study question and design.
IO 5 — Publication bias
Publication bias: research without statistically significant results is less likely to be published. Two causes are named: journal editor preference, and a sponsoring organization or company being less likely to support research that is unfavorable to them or their product. Research with significant results is more likely to appear in print, whatever the rigor or sample size of the studies.
Reporting bias is different: only a subset of all relevant data is made available. Publication bias concerns whole studies failing to reach print; reporting bias concerns a study that appears but shows only part of its data.
- It should always be considered. A review can include only what has been published, so studies that were never published leave the review looking at a skewed set. It threatens every kind of review that relies on published studies, systematic reviews and meta-analyses included.
- The direction of the distortion. If trials showing no benefit of a drug are never published, a meta-analysis of the journal articles sees only the trials favoring the drug, so the pooled benefit may be overstated.
- What does not fix it. A statistical model, such as a random effects model, cannot restore studies that were never published. Two reviewers selecting studies makes the selection consistent but cannot include studies that do not exist in print. Pooling does not reveal them either.
- How it is assessed. Assessing for publication and reporting bias (for example, with funnel plots) is part of selecting studies for a systematic review. A reader of a meta-analysis asks whether publication bias is possible for the selected topic.
Also tested
- Small negative studies. Research without statistically significant results is less likely to be published, and small negative studies are the least likely to be published, which produces publication bias; large studies are usually published regardless of their findings.
IO 6 — Principles of meta-analysis
A meta-analysis combines the results of similar studies to examine effect size quantitatively. Its principles:
- Weighted analysis of numerical results, proportionally to study size. A larger study counts for more. If weight is in proportion to size, each study's share of the weight is its participants divided by all the participants pooled: with trials of 200 and 800 participants, the smaller carries 200 of 1,000, which is 20 percent, and the larger 80 percent. Equal weighting would give each 50 percent, and picking only the largest study or counting positive studies is not pooling.
- Pooling gives a more precise estimate of effect size than any single study.
- Inclusion and exclusion criteria are very important, because they decide which studies are pooled.
- The studies must be similar. A statistical test for heterogeneity asks how different the reported effect sizes are.
Also tested
- Network meta-analysis. A network meta-analysis can use observational or randomized data and study-level or patient-level data. The interventions can be diverse as long as the studies address similar patients and outcomes, allowing direct and indirect comparisons.
Heterogeneity
Heterogeneity is the extent to which trial results are different from one another. A statistical test for it takes as its null hypothesis that there is no difference between the study results, and looks for statistically significant differences among them. This is not the null hypothesis of the pooled effect (no effect of the intervention), and it is not about bias or agreement between reviewers. The tests named are Cochran's Q and the I-squared statistic.
Also tested
- Severity in study-level analysis. Severity of illness cannot be controlled for in a study-level meta-analysis, because only summary data from each study are available. It can be explored by subgroup analysis, sensitivity analysis, or metaregression across studies, or by patient-level analyses.
Fixed effect and random effects models
| Fixed effect model | Random effects model | |
|---|---|---|
| When | There is no significant heterogeneity. | Heterogeneity is present. |
| Assumption | Assumes one fixed magnitude for the effect size; any variation is due to chance. | Allows for variation that is not due to chance. |
| Judgment | — | Depends on value judgment rather than statistical analysis to determine similarity. |
| Disadvantage | Studies with many similarities could be considered duplication rather than new research. | — |
Both models calculate the summary effect and its confidence interval. So a statistically significant test of heterogeneity across pooled trials favors the random effects model, while a fixed effect model suits pooled trials with no significant heterogeneity. A model that assumes a different effect in every study would not justify pooling to one fixed magnitude, and counting positive studies is not a model.
Also tested
- Why random effects. The random effects model allows for variation among studies that is not due to chance, so it is used when heterogeneity is present. It gives wider, more realistic confidence intervals than the fixed effect model.
IO 7 — Meta-analysis: advantages and limitations
| Advantages | Limitations |
|---|---|
|
A stronger, more precise estimate of effect size. Can detect complications. Can detect differences in subgroups. Helps identify benefits versus harms. |
Only as good as the included studies. May be misleading if the included studies are not similar — consider the strength of the value judgment. |
- Only as good as the studies. Combining poor studies produces a larger, more confident poor estimate; the flaws of the included studies are carried into the pooled result. A meta-analysis is not immune to them, cannot escape publication bias, and still needs inclusion and exclusion criteria.
- Pooling adds participants together, so the pooled sample is larger than any single study's; it does not reduce sample size. Combining numerical results is exactly what it does, and it is not restricted to rare complications: it can examine benefits, harms, complications and subgroups.
- When it misleads. A meta-analysis may be misleading if the studies combined are not similar. Similar studies, for example those measuring the same outcome, are what pooling calls for, and a larger pooled sample is an advantage, not a hazard.
- What pooling cannot do. It cannot find studies that were never published, and it does not remove bias inside the individual studies.
Also tested
- Patient-level pooling. Pooling data on individual patients rather than whole trials makes subgroup analyses possible. Individual trials are usually too small to examine subgroups, but combined patient-level data give enough people in each subgroup.
- When to pool. Studies can be combined into a single summary effect when the patients, interventions, and outcomes are relatively similar. A meta-analysis may be misleading if the included studies are not similar.
Evaluating a meta-analysis article
A reader appraising a meta-analysis asks:
- What is the quality of the included studies — are they all at the same level? (A meta-analysis that combines one high-quality trial with several seriously flawed studies raises exactly this question.)
- Is there a possibility of publication bias for the selected topic?
- What differences existed between the included studies, and were they statistically significant?
IO 8 — Grading recommendations for treatment
Grading evidence is a quick way to assess clinical significance. It is most often applied to treatment and screening. It still involves some degree of value judgment, because the pros and cons must be weighed. The scheme is based on the Grading of Recommendations Assessment, Development and Evaluation (GRADE) guidelines.
| Grade | Clarity of risk and benefit | Quality of supporting evidence | Implications |
|---|---|---|---|
| 1A Strong recommendation, high-quality evidence | Benefits clearly outweigh risks and burdens, or vice versa | Consistent evidence from well-performed randomized controlled trials, or overwhelming evidence in some other form. Further research is unlikely to change confidence in the estimates of benefits and risks. | Applies to most patients in most circumstances without reservation. Clinicians should follow a strong recommendation unless there is a clear and compelling rationale for an alternative approach. |
| 1B Strong recommendation, moderate-quality evidence | Benefits clearly outweigh risks and burdens, or vice versa | Randomized controlled trials with important limitations (inconsistent results, methodologic flaws, or imprecision), or very strong evidence of some other research design. Further research is likely to change confidence in the estimates. | Strong recommendation that applies to most patients. Clinicians should follow it unless there is a clear and compelling rationale for an alternative approach. |
| 1C Strong recommendation, low-quality evidence | Benefits appear to outweigh risks and burdens, or vice versa | Evidence from observational studies, unsystematic clinical experience, or randomized controlled trials with serious flaws. Any estimate of effect is uncertain. | Strong recommendation that applies to most patients. Some of the evidence base supporting the recommendation is of low quality. |
| 2A Weak recommendation, high-quality evidence | Benefits closely balanced with risks and burdens | Consistent evidence from well-performed randomized controlled trials, or overwhelming evidence of some other form. Further research is unlikely to change confidence. | Weak recommendation. The best action may differ depending on circumstances or patient or societal values. |
| 2B Weak recommendation, moderate-quality evidence | Benefits closely balanced with risks and burdens, with some uncertainty in the estimates of benefits, risks and burdens | Randomized controlled trials with important limitations, or very strong evidence from some other research design. Further research is likely to change confidence. | Weak recommendation. Alternative approaches are likely to be better for some patients under some circumstances. |
| 2C Weak recommendation, low-quality evidence | Uncertainty in the estimates of benefits, risks and burdens; benefits may be closely balanced with risks and burdens | Evidence from observational studies, unsystematic clinical experience, or randomized controlled trials with serious flaws. Any estimate of effect is uncertain. | Very weak recommendation. Other alternatives may be equally reasonable. |
- 1A versus 1B. Both are strong recommendations, in which benefits clearly outweigh risks and burdens. 1B is graded lower because its randomized trials have important limitations, so further research may change confidence.
- What 1C rests on. Observational studies, unsystematic clinical experience, or seriously flawed trials. 1A rests on consistent well-performed trials; 1B on trials with important limitations.
- What makes a recommendation weak. Benefits closely balanced with risks and burdens, or uncertainty in the estimates. Risks clearly outweighing benefits is “or vice versa” and gives a strong recommendation (against the treatment).
- Acting on a grade. Follow a strong recommendation unless there is a compelling rationale for an alternative; the choice with a weak recommendation varies with circumstances and values, as 2A shows. A grade of 2A has high-quality evidence but is weak because benefits and risks are closely balanced.
Also tested
- What grading covers. The Grading of Recommendations Assessment, Development and Evaluation (GRADE) table rates recommendations for treatment by the quality of the evidence and the strength of the recommendation. Deciding whether to use a diagnostic test is not a component of the grading.
Clinical practice guidelines and their standards
Good guidelines give the recommendations along with the supporting evidence and rationale. Apply your best clinical judgment when deciding to use guidelines in your practice; a guideline supports judgment and does not replace it.
| Standard | Explanation |
|---|---|
| Transparency | How the guideline was developed and funded has been made explicit and is publicly accessible. |
| Conflict of interest | Group members' conflicts of interest related to financial, intellectual, institutional, and patient or public activities bearing on the guideline are disclosed. |
| Group composition | Group membership was multidisciplinary and balanced, comprising a variety of methodologic experts and clinicians, and populations expected to be affected by the guideline. |
| Systematic review | Recommendations are based on systematic reviews that met high standards for quality. |
| Evidence and strength of recommendation | Each recommendation is accompanied by an explanation for its underlying reasoning, the level of confidence in the evidence, and the strength of the recommendation. |
| Description of recommendations | The guideline states precisely what the recommended action is and under what circumstances it should be performed. |
| External review | The guideline has been reviewed by the full spectrum of relevant stakeholders (for example, scientific and clinical experts, organizations, and patients). |
| Updating | The guideline reports the date of publication and evidence review, and plans for updating when there is new evidence that would substantially change the guideline. |
Tell the standards apart by what each one checks. Transparency is about how the guideline was developed and funded; conflict of interest is about disclosure of members' interests; group composition is about who sat on the panel that wrote it; external review is about stakeholders assessing the finished guideline; updating is about publication dates and plans to revise it.
IO 9 — Conflict of interest
Accepting a recommendation means asking whether the people behind it stand to gain. Conflict of interest is when financial or personal considerations affect objectivity; its impact may be conscious or unconscious. There are three kinds:
| Kind | Questions to ask |
|---|---|
| Financial | Who funded the research? Who benefits financially — the authors or their families? |
| Personal, intellectual or career | Are there personal biases involved? Are there professional benefits for the author? (A researcher whose promotion depends on publishing a positive result has this kind.) |
| Institutional | Does this benefit one organization more than others? (A trade group's guideline that endorses only products sold by its member companies has this kind.) |
- Why acknowledge it. It is important to acknowledge conflicts of interest so readers can determine whether bias (unintended or otherwise) is present. Disclosure does not exclude a study or change its data; it informs the reader.
- How it manifests: selective reporting of results; study selection for systematic reviews and meta-analyses; exaggeration or fabrication of results. Random assignment, confidence intervals and blinding of assessors are safeguards or statistical tools, not signs of conflict.
- Peer review helps, in addition to disclosure; neither replaces the other.
Also tested
- Conflict is specific. Conflict of interest can be financial, personal or intellectual, or institutional, and it exists in relation to a specific activity. Investing in medical products in general, not tied to the topic studied, is not a conflict.
The peer review process
- The manuscript is submitted.
- It is screened by the editors.
- It is forwarded to peer reviewers, who reject, recommend revision or accept for publication.
- The editors review language, grammar and so on.
- Publication.
So the editors screen a manuscript first, and peer reviewers see it only after that; after acceptance the editors still review language and grammar before publication, so an accepted article is neither published with no further review nor corrected for grammar by the peer reviewers.
Also tested
- Purpose of peer review. Peer review helps limit bias and makes articles more accurate, readable, and complete, and it is an accepted marker of quality. Published articles remain imperfect, so readers must still judge their validity.
IO 10 — Resources for answering clinical questions
How to keep current
Suggested ways to keep current: professional organizations, journal articles, colleagues and patients. Sales representatives, advertising and press releases are not among them; they are promotion, not appraised evidence. Systematic reviews, meta-analyses and clinical practice guidelines (IO 2, 6 and 8) are the resources that summarize the evidence for a clinical question.
Also tested
- General medical journals. The best general medical journals expose readers to the many dimensions of the profession, including history, politics, science, ideas, and experiences. They do not reliably answer immediate clinical questions or guarantee that you are current.
- Journals at the point of care. Subscribing to and keeping journals at hand is the least useful way to look up answers at the point of care, whereas guidelines, the Cochrane Library, and continually updated electronic references are valuable.
- Keeping up with new research. It is virtually impossible to keep up with all strong, clinically relevant new research on your own, so a publication or service that reviews the literature in your field is the most useful tool for keeping current.
- Comprehensive knowledge management. A comprehensive knowledge management plan combines browsing journals, a plan for point-of-care lookup, a publication that keeps you current, and reliable web sites for patients; all of these are basic elements.
Screening an article
- Is the title related to the topic, and does it have the keywords you have in mind? If no, skip the article and go to the next.
- If yes, read the abstract, summary and conclusion. Look for clear-cut aims and objectives, a well-defined research hypothesis, and precise conclusions.
- Is that useful or relevant to what you are looking for? If no, skip the article and go to the next. If yes, read the entire article.
Reading an abstract, and beyond it
| Quick reading — the abstract | Beyond the abstract |
|---|---|
|
What is the study about? What was concluded? Are the conclusions valid? Who does it apply to? Can I apply it to my patients? Should I apply it to my patients? |
How big was the effect? How strong were the research methods? What limitations existed, and were they addressed? Any potential for bias, and was it addressed? What impact will the results of the study have? What are the author's recommendations? |
Where to look: cursory and in-depth review
| Review | Your question | Where to look |
|---|---|---|
| Cursory (title and abstract) | What is this study about? | Title |
| What was concluded? | Conclusions | |
| Is it likely to be true? | Design | |
| To whom does it apply? | Patients and setting | |
| What was found? | Results | |
| In depth (the article) | Importance of the research question? | Introduction |
| How big was the effect? | Figures and tables | |
| How strong were the methods? | Methods | |
| Context | Discussion |
The figure shows an option to stop after each step, so the reader need not go on to the in-depth review. The size of the effect is read from the figures and tables, whereas the methods show how strong the study was, the introduction shows the importance of the question, and the discussion gives context.
Definitions and questions to ask
- Publication bias always matters. Publication bias should always be considered when reading a review: a review can only include what has been published, so unpublished null results leave it looking at a skewed set of studies, and missing studies can distort the evidence.
- Quick reading of an abstract. Quick reading asks what the study is about, whether the conclusions are valid, who it applies to, and whether it can and should be applied to your patients.
- Appraising the included studies. When a meta-analysis combines studies, ask whether the studies are all of equal quality: the quality of the included studies should be examined, and mixing very different levels of quality weakens the pooled result.
- Missing studies in a review. A review that searches only published articles is most likely to be missing studies without significant results, because research without significant results is less likely to be published.
- Weighting in a meta-analysis. A meta-analysis gives larger, more precise studies more weight, and the pooled result is a more precise estimate of the effect than any single study gives.
★ Quick-Reference Tables
Bias at a glance
| If the problem is… | …the bias is |
|---|---|
| Who ended up in each group | Selection (sampling) |
| How the values were obtained | Measurement |
| A third variable tangled with the exposure | Confounding |
| When or how the questions were asked | Procedural |
| People leaving the study | Attrition effect |
| People not answering at all | Non-response |
| Different care between the groups | Performance |
| Not all the findings disclosed | Reporting |
| The study never reaching print | Publication |
| Money or relationships behind the work | Conflict of interest |
Bias versus chance
| Bias | Chance | |
|---|---|---|
| Nature | Systematic error | Random variation |
| Direction | Consistently one way | Equally likely above or below |
| Fix | Good design. No analysis corrects unknown bias. | Reduce by design and averaging; estimate statistically; never eliminate |
Also tested
- Statistics on poorly designed research. Statistical analysis may give a false sense of respectability to the research, because no analysis can correct for unknown biases.
Which design answers which question
| Design | Groups defined by | Direction | Hypotheses | Class |
|---|---|---|---|---|
| Cross-sectional | Nothing — all measured at once | Neither | Generates | Observational |
| Case-control | Outcome | Backward | Generates or tests | Observational |
| Cohort | Exposure | Forward | Tests | Observational |
| Randomized controlled trial | Random assignment | Forward | Tests | Experimental |
Which property does not move with prevalence
| Fixed by the test | Moves with prevalence |
|---|---|
| Sensitivity · Specificity · Likelihood ratios | Positive predictive value · Negative predictive value |
Named examples worth remembering
| Example | What it illustrates |
|---|---|
| Framingham Heart Study (1948–) | Incidence / cohort study; identified cardiovascular risk factors |
| SchoolNuts study | Prevalence / cross-sectional study; ~1 in 20 Melbourne children with food allergy |
| 1996 Atlanta Olympics and childhood asthma | Cross-sectional ecologic study |
| Legionnaires' disease, Philadelphia 1976 | The classic epi curve |
| Vibrio after Hurricane Katrina | Surveillance detecting an outbreak |
| Tuberculosis on an international flight | Surveillance documenting geographic spread |
| Retracted Lancet vaccine and autism study | Undeclared conflicts of interest and ethical breaches; peer review did not catch it |
Also tested
- Population at risk. Only those susceptible to the disease or outcome are included; for example, men are not included in the denominator of a study of ovarian cancer.