Test scores are fundamental in assessing an individual's performance on an examination. Essentially, a test score is a numerical representation that summarizes the evidence gathered from an examinee's responses to test items, specifically those related to the constructs being measured. This numerical value provides insight into how well a person performed, but its true meaning often depends on how it's categorized and interpreted. It's crucial to distinguish
between different types of scores and their intended uses to fully grasp what a test score communicates.
Raw Scores Versus Scaled Scores
At its most basic level, a test score can be a raw score. A raw score is a direct, untransformed count, such as the simple number of questions answered correctly on a test. This raw count provides an immediate, straightforward measure of performance. However, raw scores alone can sometimes be difficult to interpret, especially when comparing performance across different test forms or over time.This is where scaled scores come into play. A scaled score is the result of one or more transformations applied to a raw score. The primary purpose of scaled scores is to report results on a consistent scale for all examinees. For instance, if a test has two forms, and one is more challenging than the other, equating processes can determine that a 65% on the easier form is equivalent to a 68% on the harder form. Both of these raw scores can then be converted to the same scaled score, perhaps 350 on a scale of 100 to 500, ensuring fair comparison. This transformation occurs after the assessment and any equating processes are complete, making it an issue of interpretability rather than psychometrics itself.
Interpreting Test Scores: Norm-Referenced and Criterion-Referenced
Test scores are typically interpreted using either a norm-referenced or a criterion-referenced approach, or sometimes a combination of both. A norm-referenced interpretation provides meaning about an examinee's standing relative to other examinees. For example, knowing a student scored in the 90th percentile tells you they performed better than 90% of the norm group. This type of interpretation is useful for ranking and comparing individuals within a specific population.In contrast, a criterion-referenced interpretation focuses on what the examinee knows or can do with regard to a specific subject matter, irrespective of how other examinees performed. This approach measures performance against a predetermined standard or set of criteria. For example, a score indicating that a student has mastered 80% of the learning objectives in a course is a criterion-referenced interpretation. Both interpretation methods offer valuable, yet distinct, insights into an examinee's performance.
Examples of Scaled Scores: ACT and SAT
In the United States, two widely recognized tests that utilize scaled scores are the ACT and the SAT. The ACT's scale ranges from 0 to 36, while the SAT's scale ranges from 200 to 800 per section. These scales were chosen to represent specific statistical properties, such as a mean and standard deviation, allowing for consistent reporting and comparison of scores over time and across different test administrations. The upper and lower bounds of these scales are designed to encompass over 99% of the population, as scores outside this range are considered difficult to measure and offer little practical value.These standardized scaled scores are crucial for college admissions, providing a common metric to evaluate applicants from diverse educational backgrounds. They help admissions officers place local data, such as coursework and grades, into a national perspective. The use of scaled scores ensures that a score of, for example, 30 on the ACT or 700 on an SAT section consistently represents a particular level of performance, regardless of the specific test form taken or the particular group of students who took it.













