Skip to main content icon/video/no-internet

Dependent observations are observations that are somehow linked or clustered; they are observations that have a systematic relationship to each other. Clustering typically results from research and sampling design strategies, and the clustering can occur within space, across time, or both.

Single and multistage cluster samples produce samples that are clustered in space, typically geographical space. Identifying and selecting known and defined clusters, and then selecting observations from the chosen clusters, produce observations that are not independent of each other. Rather, the observations are connected by the cluster, whether it is a geographic entity (e.g., city), organization (e.g., school), or household.

Longitudinal research designs produce observations that are linked across time. For example, even if individuals are selected randomly, a researcher may obtain an observation (e.g., attitudeit) on each individual i at various points in time, t. The resultant observations (e.g., attitudei1 and attitudei2) are not independent because they are both attached to person i.

The occurrence of dependent observations violates a key assumption of the general linear model: that the error terms are independent and identically distributed. Observations that are linked in space or across time are typically more homogeneous than independent observations. Because linked observations do not provide as much “new” information as independently selected observations might, the resultant sample characteristics are less heterogeneous. In qualitative research, this means that researchers should make inferences cautiously (King, Keohane, & Verba, 1994). In quantitative research, this means that standard errors of estimates will be deflated (Kish, 1965).

To correct for dependent observations resulting from cluster sampling techniques, sampling weights should be used if they are a function of the dependent variable, and thus the error term. In addition, Winship and Radbill (1994) recommend using the White heteroskedastic consistent estimator for standard errors instead of trying to explicitly model the structure of the error variances. This estimator can also be used to correct for multiplicity in the sample and is available in standard statistical packages.

To correct for temporally dependent observations, researchers typically assess and model the order of the temporal autocorrelation and introduce lagged endogenous and/or exogenous variables. Such techniques have been studied extensively in the time-series literature.

Instead of simply correcting for dependence, researchers are increasingly interested in modeling and explaining the nature of the dependence. Links between individuals or organizations, and clusters of such units, are of primary interest to social science researchers. For decades, network theorists and analysts have focused their efforts on describing and explaining the nature of dependence observed within a cluster or clusters; recent advancements extend these methods to include affiliation networks (Skvoretz & Faust, 1999) and methods for studying networks over time (Snijders, 2001). More recently, Bryk and Raudenbush (1992) have developed a framework for modeling clustered, or what they call “nested,” data. For such hierarchical (non)linear models (also known as mixed models, because they typically contain both fixed and random effects), cross-level interactions in multilevel analysis permit an investigation of contextual effects: how higher-level units (schools, organizations) can influence lower-level relationships.

...

locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading