Skip to main content icon/video/no-internet

Weighting

The term weighting refers to the process of incorporating sampling weights into analyses when using educational data collected through a complex sampling procedure, one in which the sample cannot be assumed to be from a simple random sampling process. The use of sampling weights is often a challenge for data analysts who are new to sampling theory. This entry introduces different types of sampling weights and the importance of incorporating sampling weights into analyses when using data collected through a multistage sampling procedure.

Introduction

For data selected following a complex sampling design, the sample itself may not reflect the population characteristics in known ways. Data analysts therefore must incorporate sampling weights to more appropriately mimic the population and thus to obtain unbiased estimates of population parameters. For example, suppose a population of 20 individuals consisted of 10 females and 10 males. In the sample, two males and four females were randomly selected using stratified sampling, reflecting probabilities of selection of πmale = 2/10= .2 and πfemale = 4/10 = .4. If the population average height is calculated based on this sample, which contains proportionally more females than males, the parameter estimate will be likely biased toward the population mean for males.

The sampling weight is generally defined as the inverse of the selection probability and can be considered the number of units each observation in the sample represents. In our example, the sampling weight for males wmale is 1/.2 = 5, representing 5 males in the population. Similarly, the sampling weight wfemale is 1/.4 = 2.5 for females and represents 2.5 females in the population.

In addition to this simple sampling weight based on stratification, other types of weights are available with large-scale data in education. The definitions and use of a variety of weights, as well as weight adjustment approaches, are presented in the following sections.

Types of Weights

A number of different weight variables may be available in any large-scale data set. A researcher needs to fully understand the differences across these weights to be able to select the appropriate ones for analyses of interest.

Base Weights

In large-scale educational survey, observations are often selected following a multistage framework, within which weights are provided for each stage. A typical three-stage survey, for example, may include the selection of geographic areas, schools in those selected areas, and then students within the selected schools. Two-stage surveys are also common in education, including the selection of schools, followed by the selection of students. The units being selected at each stage (e.g., the county, the school, or the student) are based on some predefined probability of selection.

Consider a two-stage sample as an example. It is rare that schools are selected using a simple random design; instead, the school has a given selection of probability and thus has a sampling weight that may differ from other schools. Researchers frequently use a sampling procedure called probability proportional to size sampling. In particular, the selection probabilities for schools vary depending on, and are proportional to, the schools’ size. In other words, large schools have relatively larger selection probabilities, πj, as compared to smaller schools, and vice versa. Therefore, larger schools tend to have smaller sampling weights (taking the inverse 1/πj = wj) than smaller schools. The school sampling weight for a given selected school is the number of schools it represents in the sampling frame. Once schools are selected, in the second stage, students are selected perhaps using categories of a given characteristic (e.g., age and race). These categories are referred to as strata. Because the student selection probability, πi|j, is the selection probability of the student within the selected school, it reflects the conditional selection probability. Thus, the second-stage within-school conditional sampling weight is 1/πi|j = wi|j.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading