What is selection bias?
A type of error that occurs when a sample is collected in a way that is not representative of the population
selection bias explained in plain English
Selection bias happens when the way you choose or collect data influences the outcome, making it unfair or inaccurate. This can lead to incorrect conclusions or decisions.
Analogy
Imagine you're trying to determine the average height of people in a city, but you only measure the heights of people who play basketball. Your sample would be biased towards taller people, giving you an incorrect average height for the entire city.
Example
A company conducts a survey about customer satisfaction, but only sends it to customers who have made a purchase in the last month. The results may not accurately represent the opinions of all customers, including those who haven't made a recent purchase.
How is selection bias used?
Selection bias can occur in various fields, including medicine, social sciences, and business, whenever data is collected or sampled. It's essential to recognize and address selection bias to ensure the validity and reliability of research findings or business decisions.
Common misconceptions about selection bias
One common misconception is that selection bias only occurs in intentional or malicious ways. However, it can also occur unintentionally, such as when a survey is only available online, excluding people who don't have internet access.
History
The concept of selection bias has been recognized for centuries, with early statisticians and researchers acknowledging the potential for biased samples. However, it wasn't until the 20th century that selection bias became a widely recognized and studied phenomenon in statistics and research methodology.
People also read
- attribute
A characteristic or feature of an object or concept
- automation bias
The tendency to over-rely on automated systems and ignore or underweight human judgment
- bias
A systematic error or distortion in a machine learning model's results
- bias (math) or bias term
A constant added to a linear combination of inputs in a machine learning model
- calibration layer
A component in a neural network that adjusts the output to match the true probabilities of a task
- Confabulation
When an AI produces a confident, fluent answer that sounds true but is factually wrong — generating plausible language without a reliable link to reality.
- confirmation bias
The tendency to favor information that confirms existing beliefs or expectations
- counterfactual fairness
A fairness metric in AI that ensures decisions are fair by comparing actual outcomes with hypothetical outcomes where a sensitive attribute is different
- coverage bias
A type of bias that occurs when the data used to train a model does not accurately represent the population or phenomenon being studied
- demographic parity
A fairness metric in machine learning that ensures equal outcomes for different demographic groups