Scores and scales for school achievement
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
In his handbook chapter on scales and norms, Angoff (1971, p. 508) observed that "without a meaningful score to transmit the value of the test performance, the test ceases to be a measuring instrument and becomes merely a preactice exercise for the student on a collection of items." A multitude of scores are commonly used in the schools of the United States to give meaning to scores, from the largely undocumented and unexamined scales of teacher tests to the venerable and meticulously equated 200-to-800 scale of the SAT. Competition among the publishers of school achievement batteries has resulted in the creation of numerous computer-generated score reports for these instruments, tailored to parents, teachers, principals, district administrators, and other audiences. Both norm-referenced and criterion-referenced interpretations are typically provided, the latter often less defensible than the former. One fairly recent norm-referenced score that may be preferable to the percentile for some applications is the NCE. Recent educational assessments at the state and national levels, using the methods of item response theory, are advancing the state of the art in the scaling of educational achievement. Meaningful descriptions and illustrative items can be used to anchor specific scores, and to describe just what examinees at different levels know or can do. Further research is needed on statistical models that can capture the complexities of multidimensional content domains, in which the relative difficulties of items change as a function of curriculum and instruction from one classroom to another. Nonetheless, significant progress is being made, and exemplars have already appeared to IRT-based criterion-referenced scales that may point the direction of educational measurement for the future.
