Skip to main content icon/video/no-internet

Regression is the most widely used data analysis technique in the nonexperimental social sciences. In a regression model, a DEPENDENT VARIABLE, commonly labeled Y, is a function of one or more INDEPENDENT VARIABLES, commonly labeled X1, X2, and so on. The independent variables are assumed to explain, or at least predict, the phenomenon measured by the dependent variable. The simplest of regression models posits only one independent variable: Y = a + b1X1 + e. This bivariate regression says variable Y is a linear additive function of variable X1, plus a constant (the fixed value represented by a) plus error (represented by e). Of principal interest is the influence of X1 on Y, the regression coefficient represented by b1. To obtain numerical estimates for the values a and b1, a straight line is fitted to the observations on X1 and Y by the method of ORDINARY LEAST SQUARES (OLS). A more complicated regression model posits two independent variables: Y = a + b1X1 + b2X2 + e. This multiple regression model says Y is a linear function of X1 and X2 (plus a constant term and an error term). A multiple regression analysis equation always has two or more independent variables. Below, the mechanics of bivariate and multivariate regression are reviewed in a data example. Then, assumptions, history, and recent developments are considered.

Bivariate Regression

Social scientists often want to examine the relationship between two variables. To illustrate, an empirical example is unfolded. Imagine that educational sociologists in a small midwestern college seek information on the annual earnings of their students well after graduation. They want to know many things, including the influence of basic demographic forces, such as the socioeconomic background of the students’ families. In particular, they seek to establish the link, if any, between the educational attainment of parents and the later job earnings of the child. To gather the necessary data on these and other research questions, they draw a RANDOM SAMPLE from the alumni list of 350 students who graduated 15 years ago. Because of cost considerations, they sample 1 out of 10, yielding a sample size of 35. Table 1 contains some of the data from the survey administered to them. In column 1 is the code number given the student. In column 2 is the parent education variable, measured as the number of years of formal schooling completed by the most educated parent (mother or father). In column 3 is the student income variable, measured as the reported gross annual income (in thousands of dollars) the student earned 15 years after graduation. In column 4 is the gender of the student, scored 1 = male or 0 = female. These data are listed as they would appear when entered into a typical computer STATISTICAL PACKAGE.

The research question at hand is whether parent education (variable X1) helps account for later student income (variable Y). The dominant HYPOTHESIS is that they are positively related. The more educated the parental environment, the more child opportunities and incentives for learning and advancement and, ultimately, higher income on the job. Does an analysis of the data support the hypothesis? In bivariate regression, the first step is inspection of a SCATTERPLOT, asin Figure 1.

...

locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading