Skip to main content icon/video/no-internet

Although the term association is used broadly, association model has a specific meaning in the literature on categorical data analysis. By association model, we refer to a class of statistical models that fit observed frequencies in a cross-classified table with the objective of measuring the strength of association between two or more ordered categorical variables. For a two-way table, the strength of association being measured is between the two categorical variables that comprise the cross-classified table. For three-way or higher-way tables, the strength of association being measured can be between any pair of ordered categorical variables that comprise the cross-classified table. Although some association models make use of the a priori ordering of the categories, other models do not begin with such an assumption and indeed reveal the ordering of the categories through estimation. The association model is a special case of a LOG-LINEAR MODEL or log-bilinear model.

Leo Goodman should be given credit for having developed association models. His 1979 paper, published in the Journal of the American Statistical Association, set the foundation for the field. This seminal paper was included along with other relevant papers in his 1984 book, The Analysis of Cross-Classified Data Having Ordered Categories. Here I first present the canonical case for a two-way table before discussing extensions for three-way and higher-way tables. I will also give three examples in sociology and demography to illustrate the usefulness of association models.

GENERAL SETUP FOR A TWO-WAY CROSS-CLASSIFIED TABLE

For the cell of the ith row and the jth column (i = 1,…, I, and j = 1,…, J) in a two-way table of R and C, let fij denote the observed frequency and Fij the expected frequency under some model. Without loss of generality, a log-linear model for the table can be written as follows:

None

where μ is the main effect, μR is the row effect, μC is the column effect, and μRC is the interaction effect on the logarithm of the expected frequency. All the parameters in equation (1) are subject to ANOVA-type normalization constraints (see Powers & Xie, 2000, pp. 108–110). It is common to leave μR and μC unconstrained and estimated nonparametrically. This practice is also called the “saturation” of the marginal distributions of the row and column variables. What is of special interest is μRC: At one extreme, μRC may all be zero, resulting in an independence model. At another extreme, μRC may be “saturated,” taking (I − 1)(J − 1) degrees of freedom, yielding exact predictions (Fij = fij for all i and j).

Typically, the researcher is interested in fitting models between the two extreme cases by altering specifications for μRC. It is easy to show that all odds ratios in a two-way table are functions of the interaction parameters (μRC). Let θij denote a local log-odds ratio for a 2 × 2 subtable formed from four adjacent cells obtained from two adjacent row categories and two adjacent column categories:

None

Let us assume that the row and column variables are ordinal on some scales x and y. The scales may be observed or latent. A linear-by-linear association model is as follows:

None

...

locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading