Entry
Reader's guide
Entries A-Z
Selection Bias
Selection bias is an important concern in any social science research design because its presence generally leads to inaccurate estimates. Selection bias occurs when the presence of observations in the sample depends on the value of the variable of interest. When this happens, the sample is no longer randomly drawn from the population being studied, and any inferences about that population that are based on the selected sample will be biased. Although researchers should take care to design studies in ways that mitigate nonrandom selection, in many cases, the problem is unavoidable, particularly if the data-generating process is out of their control. For this reason, then, many methods have been developed that attempt to correct for the problems associated with selection bias. In general, these approaches involve modeling the selection process and then controlling for it when evaluating the outcome variable.
For an example of the consequences of selection on inference, consider the effect of SAT scores on college grades. Under the assumption that higher SAT scores are related to success in college, one would expect to find a strong and positive relationship between the two among students admitted to a top university. The students with lower SAT scores who are admitted, however, must have had some other strong qualities to gain admittance. Therefore, these students will achieve unusually high grades compared to other students with similar SAT scores who were not admitted. This will bias the estimated effect of SAT scores on grades toward zero. Note that the selection bias does not occur because the sample is over representative of students with high SAT scores, but because the students with low scores are not representative of all students with low SAT scores in that they score high on some unmeasured component that is associated with success.
A frequent misperception of selection bias is that it can be dismissed as a concern if one is merely trying to make inferences about the sample being analyzed rather than the population of interest, which would be akin to claiming that we have correctly estimated the relationship between SAT scores and grades among admittees only. This is false: The grades of the students in our example are representative of neither all college applicants nor students with similar SAT scores. Put another way, our sample is not a random sample of students conditional on their SAT scores. The estimated relationship is therefore inaccurate. In this example, the regression coefficient will be biased toward zero because of the selection process. In general, however, it may be biased in either direction, and in more complicated (multivariable) regressions, the direction of the bias may not be known.
Although selection bias can take many different forms, it is generally separated into two different types: censoring and truncation. Censoring occurs when all of the characteristics of each observation are observed, but the dependent variable of interest is missing. Truncation occurs when the characteristics of individuals that would be censored are also not observed; all observations are either complete or completely missing. In general, censoring is easier to deal with because data exist that can be used to directly estimate the selection process. Models with truncation rely more on assumptions about how the selection process is related to the process affecting the outcome variable.
...
- Analysis of Variance
- Association and Correlation
- Association
- Association Model
- Asymmetric Measures
- Biserial Correlation
- Canonical Correlation Analysis
- Correlation
- Correspondence Analysis
- Intraclass Correlation
- Multiple Correlation
- Part Correlation
- Partial Correlation
- Pearson's Correlation Coefficient
- Semipartial Correlation
- Simple Correlation (Regression)
- Spearman Correlation Coefficient
- Strength of Association
- Symmetric Measures
- Basic Qualitative Research
- Basic Statistics
- F Ratio
- N(n)
- t-Test
- X¯
- Y Variable
- z-Test
- Alternative Hypothesis
- Average
- Bar Graph
- Bell-Shaped Curve
- Bimodal
- Case
- Causal Modeling
- Cell
- Covariance
- Cumulative Frequency Polygon
- Data
- Dependent Variable
- Dispersion
- Exploratory Data Analysis
- Frequency Distribution
- Histogram
- Hypothesis
- Independent Variable
- Measures of Central Tendency
- Median
- Null Hypothesis
- Pie Chart
- Regression
- Standard Deviation
- Statistic
- Causal Modeling
- DISCOURSE/CONVERSATION ANALYSIS
- Econometrics
- Epistemology
- Ethnography
- Evaluation
- Event History Analysis
- Experimental Design
- Factor Analysis and Related Techniques
- Feminist Methodology
- Generalized Linear Models
- HISTORICAL/COMPARATIVE
- Interviewing in Qualitative Research
- Latent Variable Model
- LIFE HISTORY/BIOGRAPHY
- LOG-LINEAR MODELS (CATEGORICAL DEPENDENT VARIABLES)
- Longitudinal Analysis
- Mathematics and Formal Models
- Measurement Level
- Measurement Testing and Classification
- Multilevel Analysis
- Multiple Regression
- Qualitative Data Analysis
- Sampling in Qualitative Research
- Sampling in Surveys
- Scaling
- Significance Testing
- Simple Regression
- Survey Design
- Time Series
- ARIMA
- Box-Jenkins Modeling
- Cointegration
- Detrending
- Durbin-Watson Statistic
- Error Correction Models
- Forecasting
- Granger Causality
- Interrupted Time-Series Design
- Intervention Analysis
- Lag Structure
- Moving Average
- Periodicity
- Serial Correlation
- Spectral Analysis
- Time-Series Cross-Section (TSCS) Models
- Time-Series Data (Analysis/Design)
- Trend Analysis
Get a 30 day FREE TRIAL
-
Watch videos from a variety of sources bringing classroom topics to life
-
Read modern, diverse business cases
-
Explore hundreds of books and reference titles
Sage Recommends
We found other relevant content for you on other Sage platforms.
Have you created a personal profile? Login or create a profile so that you can save clips, playlists and searches