Some cautionary notes regarding the use of aggregated scores as a measure of behavioral stability
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
The utility of the method of aggregation as a measure of behavioral consistency was examined in 26 studies involving computer-generated, repeated-measurement data. The first series of studies involved rectangular distributions in which score constant. A second series of studies used normally distributed z scores, and score consistency was manipulated by inducing a desired correlation between the scores in adjacent trials. In both sets of studies, the aggregate stability coefficient was a strictly increasing function of the number of aggregated trials, and even trivial amounts of score stability resulted in large stability coefficients. In a third series of studies, high stability coefficients occurred when computed on combined unstable subsamples which differed from each other only in central tendency. Terminal aggregate coefficients were compared with Spearman-Brown prophecy and Cronbach's alpha reliability coefficients computed on the experimental data. It was concluded that the method of aggregation produces spuriously high estimates of behavioral consistency. It was further shown that the Spearman-Brown prophecy formula and coefficient alpha accurately predict the results of the aggregation method, suggesting that aggregation is an internal consistency reliability procedure. The equating of stability with traditional notions of reliability was questioned.
