Press ESC to close

What is selection bias? Definition, types, examples

Getting your Trinity Audio player ready...
Summarize this Blog with AI

Contents

 

Glossary of Key Terms

 

Term

Definition

Selection Bias

A systematic error that arises when the individuals included in a study are not representative of the population the researcher intends to study.

Sampling Bias

A subset of selection bias in which the flaw lies specifically in the method used to draw a sample from a population.

Attrition Bias

Distortion in study results caused by the non-random dropout of participants over the course of a study.

Confirmation Bias

A cognitive tendency to search for, interpret, or recall information in a way that confirms one’s pre-existing beliefs.

Survivorship Bias

The error of focusing only on subjects that passed a selection process, while overlooking those that did not.

Berkson’s Bias

A form of selection bias that arises in hospital-based studies because patients admitted to hospital are not representative of the general population.

Self-Selection Bias

Bias introduced when individuals choose for themselves whether to participate in a study, rather than being randomly assigned.

Convenience Sampling

A sampling method in which participants are chosen based on their easy availability rather than random selection.

Randomization

A technique that assigns participants to groups by chance, eliminating systematic differences between groups.

Pre-registration

The practice of publicly documenting study hypotheses, design, and analysis plan before data collection begins.

DAG (Directed Acyclic Graph)

A visual tool used in causal inference to map relationships between variables and identify potential sources of bias.

Intention-to-Treat

An analysis approach in which all participants are analyzed in the group to which they were originally assigned, regardless of dropout.

CONSORT

Consolidated Standards of Reporting Trials: a reporting guideline for randomized controlled trials.

STROBE

Strengthening the Reporting of Observational Studies in Epidemiology: a reporting guideline for observational research.

 

Key Takeaways

  • Selection bias is a structural flaw: it occurs when the people included in a study do not represent the population of interest, producing results that cannot be generalized.
  • Selection bias, attrition bias, and sampling bias are all methodological in nature and occur at different stages of research; confirmation bias is cognitive and can amplify any of the three.
  • The most reliable defense against selection bias is randomization at the design stage, combined with transparent reporting of all inclusion and exclusion criteria.
  • Students at every level can reduce selection bias by pre-registering studies, using DAGs to map assumptions, and following established reporting guidelines such as CONSORT and STROBE.

 

What Is Selection Bias?

Selection bias is a systematic error that occurs when the individuals studied are not representative of the population the researcher intends to study. It distorts findings and limits the degree to which conclusions can be generalized beyond the sample. Unlike random error, selection bias is directional: it consistently pushes results in a particular direction, making conclusions misleading rather than merely imprecise.

Selection bias matters because it can invalidate otherwise rigorous research. A study may have strong internal procedures but still produce conclusions that do not apply to the real world if the participants were selected in a non-random or systematically distorted way.

Selection bias appears across fields, including:

  • Clinical medicine and public health, where eligibility criteria or volunteer recruitment skew who is studied
  • Social science, where survey samples often over-represent educated, urban, or digitally active populations
  • Machine learning and artificial intelligence, where training data often reflects historical inequalities
  • Economics, where studies of labor market outcomes may exclude those who left the labor force
  • Journalism, where only available or willing sources are quoted

 

What Are the Main Types of Selection Bias?

Selection bias is an umbrella concept that covers several distinct mechanisms. Each arises through a different process, though all share the same core problem: the sample does not accurately represent the target population.

 

Type

Description

Common Field

Self-Selection Bias

Participants opt in voluntarily, skewing toward motivated or interested individuals

Surveys, psychology, public health

Survivorship Bias

Only those who completed or survived a process are observed

Finance, military research, business

Berkson’s Bias

Hospital or clinical populations differ systematically from the general public

Epidemiology, medicine

Exclusion Bias

Eligibility criteria inadvertently exclude meaningful subgroups

Clinical trials, social science

Healthy Worker Effect

Employed populations tend to be healthier than the general population

Occupational health, economics

Referral Bias

The referral process determines who is studied, distorting who appears in the sample

Medicine, psychology

 

Understanding which type of selection bias is present in a given study determines which remedies are most appropriate. Survivorship bias, for example, requires a fundamentally different research design fix than Berkson’s bias.

 

Selection Bias vs. Confirmation Bias: What Is the Difference?

Selection bias is a structural, methodological flaw; confirmation bias is a cognitive, psychological one. Selection bias arises from how a study is designed and executed. Confirmation bias arises from how a researcher thinks and decides. Both can distort findings, but they operate through entirely different mechanisms.

 

Defining Confirmation Bias

Confirmation bias is the tendency to search for, interpret, favor, and recall information in a way that confirms or supports one’s pre-existing beliefs or values. It affects how researchers choose which studies to read, how they interpret ambiguous data, and which findings they choose to report or emphasize.

 

How They Differ

The table below summarizes the key distinctions between selection bias and confirmation bias:

 

Bias Type

When It Occurs

Nature

Example

Selection Bias

At study design or recruitment

Structural/methodological

Recruiting only online users for a general population survey

Confirmation Bias

During any research phase

Cognitive/psychological

Ignoring studies that contradict a favored hypothesis

Attrition Bias

During or at exit from a study

Structural/methodological

Sicker patients dropping out of a drug trial early

Sampling Bias

During sample selection

Structural/methodological (subset of selection bias)

Using convenience sampling of college students

 

Can One Cause the Other?

Yes. A researcher with strong prior beliefs about a hypothesis may unconsciously design a study in a way that produces selection bias. For example, a researcher who believes that a new intervention works may set eligibility criteria that exclude the patients most likely to fail to respond to treatment. The confirmation bias in thinking leads to a structural selection bias in design.

This feedback loop makes confirmation bias particularly dangerous as a precursor to selection bias. Pre-registration is the most effective tool for breaking the loop, because it forces researchers to document hypotheses and design choices before they have seen the data.

 

Detection and Correction

Detection and correction approaches differ for each type:

  • Selection bias: Audit study design, eligibility criteria, and recruitment method for structural non-randomness
  • Confirmation bias: Use pre-registration, blinded analysis, adversarial collaboration, and structured peer review
  • Both together: Require transparency in reporting who was excluded and why, and invite external replication

 

Selection Bias vs. Attrition Bias: How Are They Related?

Selection bias occurs at the beginning of a study, when participants are recruited; attrition bias occurs during or at the end, when participants drop out in a non-random way. Both distort results by making the sample unrepresentative, but they do so at different stages and require different remedies.

 

Defining Attrition Bias

Attrition bias, also called loss-to-follow-up bias, arises when the participants who leave a study are systematically different from those who remain. If dropout is random, it reduces statistical power but does not bias results. If dropout is non-random, the remaining sample no longer represents the original population, and findings may be misleading.

Common examples of attrition bias include:

  • Clinical trials where patients who experience side effects or who are not improving stop taking the medication and withdraw, leaving only those who respond well
  • Longitudinal studies where lower-income or less-educated respondents disengage over time, leaving a wealthier and more educated sample than intended
  • Educational studies where struggling students drop out of a program, making the program appear more effective than it is

 

How Do Selection Bias and Attrition Bias Compare?

 

Feature

Selection Bias

Attrition Bias

When it happens

Before or at study entry

After study has begun

Who is affected

Who gets into the sample

Who remains in the sample

Primary remedy

Randomization, broad eligibility

Intention-to-treat analysis, rigorous follow-up

Can they co-occur?

Yes

Yes

 

Can Both Occur in the Same Study?

Yes, and this is common. A study may begin with a recruitment process that introduces selection bias and then lose participants in a non-random way that introduces attrition bias. When both are present, findings can be severely distorted. Longitudinal studies and randomized controlled trials are particularly vulnerable to this combination.

 

How Is Attrition Bias Addressed?

  • Intention-to-treat analysis: All participants are analyzed in their originally assigned group regardless of whether they completed the study
  • Multiple imputation: Statistical techniques are used to estimate plausible values for missing data from participants who dropped out
  • Rigorous follow-up protocols: Researchers invest heavily in re-contacting lapsed participants
  • Dropout documentation: The study reports who dropped out, when, and for what reason, allowing readers to assess potential bias

 

Selection Bias vs. Sampling Bias: Is There a Real Difference?

Sampling bias is a specific subtype of selection bias, not a separate concept. All sampling bias is selection bias, but not all selection bias is sampling bias. Selection bias is the broader category; sampling bias refers specifically to flaws in how the sample is drawn from the population.

 

Defining Sampling Bias

Sampling bias occurs when the method used to draw a sample systematically favors certain members of the population over others. The flaw lies in the selection procedure itself, not in who chooses to participate after being approached.

Common forms of sampling bias include:

  • Convenience sampling: Participants are chosen because they are easy to access, such as using university students for general population psychology studies
  • Undercoverage bias: Certain groups have little or no chance of being selected, such as telephone surveys that miss households without landlines
  • Voluntary response bias: Only those who feel strongly enough to respond do so, as in call-in polls or online comment sections

 

How Selection Bias and Sampling Bias Compare

 

Feature

Selection Bias

Sampling Bias

Scope

Broad umbrella term

Specific subtype of selection bias

Cause

Any non-random inclusion/exclusion process

Flawed method of drawing the sample

Relationship

Parent concept

Child concept

Example

Excluding non-English speakers from a global study

Using a phone book to sample a city’s population

 

A Canonical Example of Sampling Bias

The 1936 United States presidential election poll conducted by the Literary Digest magazine is one of the most cited examples of sampling bias in the history of research. The magazine surveyed over two million people using lists drawn from telephone directories and automobile registration records. At a time when many Americans could not afford telephones or cars, this method over-sampled wealthier, Republican-leaning citizens. The poll predicted a landslide victory for Alf Landon; Franklin D. Roosevelt won by a historic margin.

 

When the Terms Are Used Interchangeably

In everyday academic usage, the terms selection bias and sampling bias are often used interchangeably, particularly in social science and public health literature. This can cause confusion when readers try to distinguish the specific mechanism at work. A practical approach is to use selection bias as the general term and sampling bias when the flaw can be traced specifically to the method of drawing the sample.

 

How Do You Detect Selection Bias in Research?

Selection bias can be detected by systematically examining who is in a study, who was excluded, and whether those decisions were made in a non-random or systematically skewed way. The table below lists the most common red flags and the questions researchers and readers should ask.

 

Red Flag

What to Ask

Narrow eligibility criteria

Who was excluded, and why? Could exclusions distort the findings?

Volunteer or opt-in recruitment

Are those who chose to participate systematically different from those who did not?

Convenience sampling

Does the sample reflect the target population in age, income, education, or geography?

High dropout rate

Did more participants drop out of one group than another, and does the paper explain why?

Single-institution data

Does the institution serve a population that differs from the broader public?

No reporting of excluded participants

Does the paper follow CONSORT or STROBE guidelines on exclusion documentation?

 

Additional detection approaches include:

  • Comparing sample demographics to known population statistics from census or administrative data
  • Using directed acyclic graphs to map potential selection processes and identify which variables may have influenced inclusion
  • Running sensitivity analyses to test whether conclusions change under different assumptions about who might have been excluded
  • Checking whether the study has been replicated, since selection bias often produces findings that do not replicate

 

How Can Researchers Minimize Selection Bias?

Selection bias is best addressed at the design stage, before data collection begins. The strategies below range from the gold standard of randomization to practical reporting practices that increase transparency.

  • Randomization: Assign participants to study conditions using a random mechanism, which ensures that no systematic difference exists between groups before the study begins
  • Pre-registration: Document study hypotheses, eligibility criteria, and analysis plans publicly before data collection begins, reducing the risk of post-hoc design changes
  • Stratified sampling: Divide the target population into subgroups based on relevant characteristics and sample from each subgroup in proportion to its representation in the population
  • Broad eligibility criteria: Avoid narrow exclusion criteria unless they are scientifically necessary; each criterion restricts generalizability
  • Administrative and population-level datasets: Use data sources that cover entire populations rather than selected subgroups, reducing selection at the source
  • Transparent reporting: Follow CONSORT for randomized trials or STROBE for observational studies, both of which require detailed documentation of who was excluded and why

 

Real-World Examples of Selection Bias

Selection bias is not a theoretical concern: it has produced concrete harms in medicine, technology, economics, journalism, and social research. The table below presents documented examples across fields.

 

Field

Example

Bias at Work

Medicine

Historical exclusion of women from clinical trials

Findings from male-only samples were generalized to all patients, masking sex-specific drug effects

Technology

Facial recognition systems trained on non-representative image datasets

Underrepresentation of certain demographic groups led to higher error rates for those groups

Economics

Wage gap studies that do not account for labor market participation

Excluding non-participants ignores the selection process determining who works, distorting wage comparisons

Journalism

Person-on-the-street interviews during news broadcasts

Only those willing to speak on camera are sampled, skewing perceived public opinion

Social Media Research

Studies using platform data to draw conclusions about the general public

Platform users differ from non-users in age, education, income, and political engagement

 

These examples share a common structure: a study or system drew conclusions from a sample that did not represent the relevant population, and those conclusions were then applied to people who were never adequately studied.

 

Tips for Students: How to Handle Selection Bias in Your Own Research

 

For Undergraduate Students

Undergraduates are often conducting their first formal research projects and may not yet have developed the methodological intuition to spot selection bias before it becomes a problem. The tips below are designed to build that intuition and avoid the most common undergraduate errors.

 

Tip

Why It Matters

Ask who is missing from every study you read

Trains the habit of spotting selection bias before it affects your own work

Label convenience samples honestly in limitations sections

Reviewers and instructors expect this; omitting it signals inexperience

Avoid opt-in online surveys without flagging the self-selection problem

Voluntary respondents often differ from the broader population in attitudes and behaviors

Use your library’s methodology guides before finalizing study design

Catching bias problems early costs far less time than correcting them after data collection

Discuss design choices with your professor before collecting data

Faculty can identify bias risks that are not visible from inside the project

 

For Graduate Students

Graduate students are expected to demonstrate methodological sophistication. Reviewers, committees, and journal editors will look critically at selection processes in theses, dissertations, and submitted manuscripts. The tips below reflect the standards of professional research practice.

 

Tip

Why It Matters

Pre-register your study on the Open Science Framework or AsPredicted

Reduces unconscious bias in analysis and signals methodological rigor to reviewers

Draw a DAG before finalizing your design

Forces explicit mapping of selection processes and potential confounders

Justify every inclusion and exclusion criterion explicitly

Each criterion is a potential source of selection bias and must be defensible

Follow CONSORT or STROBE reporting guidelines

Ensures transparent documentation of who was and was not included

Address selection bias in your limitations chapter, do not avoid it

Committees and peer reviewers will raise it; a thoughtful response signals maturity

Run sensitivity analyses to test robustness of conclusions

Demonstrates that your findings are not an artifact of who was selected into the sample

 

Summary and Quick-Reference Comparison

The table below provides a concise comparison of the four types of bias discussed in this guide, summarizing when each enters research, the core problem it creates, and the primary way to address it.

 

Bias

Core Problem

When It Enters

Primary Fix

Selection

Non-representative inclusion

Design or recruitment

Randomization, transparent criteria

Confirmation

Motivated cognition

Any research phase

Pre-registration, peer review

Attrition

Non-random dropout

During the study

Intention-to-treat, rigorous follow-up

Sampling

Flawed sample drawing

Sample selection

Probability sampling, stratification

 

The most important practical takeaway is that bias awareness is not a one-time checklist: it requires sustained attention at every stage of research design, data collection, analysis, and reporting.

Recommended further reading for students and researchers who want to deepen their understanding of selection bias and causal inference includes works by Miguel Hernán and James Robins on causal inference, and by William Shadish, Thomas Cook, and Donald Campbell on experimental and quasi-experimental designs.

 

Frequently Asked Questions

 

What is the difference between selection bias and sampling bias in research?

Sampling bias is a subtype of selection bias. Selection bias is the broad term for any systematic error that makes a sample unrepresentative of the target population. Sampling bias refers specifically to a flaw in the method used to draw the sample from that population. Every case of sampling bias is also a case of selection bias, but selection bias can also arise from other sources, such as exclusion criteria, self-selection by participants, or attrition during a study.

 

How does selection bias affect the validity of a study?

Selection bias threatens external validity, meaning the degree to which findings can be generalized beyond the study sample. When participants are not representative of the target population, results may accurately describe the sample but say nothing meaningful about the broader group the study was intended to understand. In some cases, selection bias also threatens internal validity if the comparison groups within a study differ in systematic ways before the intervention begins.

 

Can selection bias occur in randomized controlled trials?

Yes. Randomized controlled trials are the gold standard for reducing selection bias, but they are not immune. Selection bias can enter through narrow eligibility criteria that exclude large segments of the population, through volunteer recruitment methods that attract a particular type of participant, or through attrition that occurs non-randomly during the trial. Even a perfectly randomized study can produce biased findings if the eligible population is systematically different from the general population.

 

What is survivorship bias and how is it related to selection bias?

Survivorship bias is a form of selection bias in which only the subjects that passed a selection process, or survived to the end of an observation period, are studied. The subjects that failed or dropped out are omitted from the analysis, often invisibly. Classic examples include studying only successful companies to draw lessons about business strategy, or analyzing only soldiers who survived a war to study the effects of battlefield conditions. In both cases, the missing data are not random, and conclusions drawn from survivors alone will be systematically skewed.

 

How do you identify selection bias in a published study?

The most reliable way to identify selection bias in a published study is to examine the eligibility criteria, the recruitment method, the dropout rate, and the demographic profile of the final sample. Ask whether the people who were excluded or who chose not to participate might differ systematically from those who were included. Compare the sample characteristics to population-level statistics. Check whether the study follows CONSORT or STROBE reporting guidelines, which require explicit documentation of exclusions. If the exclusion criteria are narrow, the recruitment method is based on voluntary participation, or the dropout rate is high and unexplained, selection bias is likely.

 

What is the best way to avoid selection bias in survey research?

The most effective strategies for avoiding selection bias in survey research are probability sampling and high response rates. Probability sampling ensures that every member of the target population has a known, non-zero chance of being selected. High response rates reduce the risk that non-respondents differ systematically from respondents. Where probability sampling is not feasible, researchers should use stratified or quota sampling to ensure key demographic groups are represented in proportion to their share of the population, and should clearly document the limitations of their sampling method in the published report.

 

Is selection bias the same as volunteer bias?

Volunteer bias, also called self-selection bias, is one specific type of selection bias. It occurs when individuals who choose to participate in a study differ systematically from those who do not. Volunteers tend to be more motivated, more health-conscious, more educated, or more interested in the topic being studied than the general population. This makes them an unrepresentative sample. While volunteer bias is one of the most common forms of selection bias in behavioral and health research, selection bias is a broader concept that encompasses many other mechanisms as well.

 

How does selection bias differ from information bias?

Selection bias concerns who is in the study: it arises when the sample is not representative of the target population. Information bias concerns what is measured: it arises when data on exposure, outcome, or covariates are measured inaccurately or inconsistently. Both can distort study findings, but they operate at different stages and require different remedies. Selection bias is addressed through study design and recruitment; information bias is addressed through measurement standardization, blinding, and validated instruments. A study can suffer from both simultaneously, which is why comprehensive bias assessments consider both who was included and how data were collected.

This article was originally published on May 15, 2024, and updated on June 28, 2026.