Skip to main content icon/video/no-internet

Test–Retest Reliability

Reliability estimates utilizing a test–retest approach measure the degree to which the same testing instrument produces similar results when administered to the same individual in as similar a manner as possible over a period of time. Test–retest reliability is a popular form of reliability estimation for the development and validation of test instruments and is based on correlation. Test–retest reliability falls behind only internal consistency estimates (e.g., coefficient α) in popularity for the evaluation of reliability. Test–retest reliability is a measure of test consistency and score fluctuation emphasizing the psychometric assessment of test form stability over a period of time. For instance, if an intelligence test is administered twice to the same individual within a short period of time, then a high test–retest reliability coefficient would be expected because of the general stability of the measured intellectual functioning and the standardized testing procedures. In contrast with other consistency estimates, such as those that examine internal consistency (e.g., coefficient α or split-half reliability), test–retest reliability is a measure of temporal stability. Because the instrument used during the calculation of test–retest reliability is the same during both administrations, this approach to reliability estimation assesses measurement error as the degree to which changes happen across administrations. If different but supposedly related testing instruments were administered in the same manner over a period of time, this administration would depend on an alternative form of reliability. If reliability is calculated using responses from a single test administration based on how similar responses are to one another, this administration would depend on the coefficient α. This entry discusses the theoretical approach and assumptions of as well as issues associated with test–retest reliability and then provides information on notation and interpretation.

Theoretical Approach and Underlying Assumptions

Test–retest is a measure of reliability as seen through the lens of classical test theory. In this approach to classical test theory, the closer obtained scores are to one another over two administrations, the higher the test–retest reliability coefficient. Higher reliability coefficients indicate a greater portion of true score measurement and lesser amount of error. Thus, higher reliability coefficients indicate more precise and stable measurement. For instance, if 80% of variability in test scores is attributable to systematic performance, then the instrument would have a .80 test–retest reliability coefficient, indicating 20% of variability being the result of error. Some examples of error that may occur causing variability between scores include variations in attention and concentration to the task at hand, learning as a result of test exposure, approaches to testing that are indicative of haphazard or of careless responding, and problems with item comprehension. Although it is impossible to remove all variability from measurement, the expectation is that well-designed tests for a stable trait will be able to obtain a consistently reliable measurement of the underlying true score.

Test–retest reliability relies on two underlying assumptions. Test–retest assumes that true scores of the measured characteristic for an individual do not change over time and that all variation in an observed score is due to either random or systematic error. Not all characteristics are ideal for this assumption because some are expected to change over time. Whereas major personality characteristics (e.g., the Big Five personality traits such as extraversion and agreeableness) are generally conceptually stable over the lifetime and thus appropriate for test–retest reliability measurement, other state-based attributes are not. For instance, depression is a mood state and would be expected to fluctuate over a course of time, therefore use of test–retest coefficients to demonstrate evidence of reliability would be less appropriate. In the interim between separate administrations of a depression test, individuals are likely to experience a change in their stress (e.g., receive parking tickets, have disagreements with loved ones, enjoy a rewarding day at work) and may even experience major life events. All of these would be expected to impact the amount of depression the person reports because the underlying level of depression experienced would have changed.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading