Scoring Models and Reporting Flashcards
6 cards from real CBT practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 Scoring Models and Reporting flashcards as text
A certification body administers multiple forms of a CBT exam throughout the year. To ensure that a score of 450 on Form A represents the same level of proficiency as a score of 450 on the more difficult Form B, which of the following psychometric procedures is essential?
Answer: Equating
Equating is the statistical process used to adjust scores on different forms of a test to account for variations in difficulty. This ensures that scores are comparable and have the same meaning regardless of which form a candidate takes. Standard setting determines the passing score, item analysis evaluates individual question performance, and score validation is a broader process of ensuring scores are used appropriately.
A candidate's score report for a certification exam indicates their scaled score is 750, with a passing score of 700. The report also shows a Standard Error of Measurement (SEM) of 25. What does the SEM indicate?
Answer: The range within which the candidate's 'true score' likely falls.
The Standard Error of Measurement (SEM) quantifies the amount of error in a test score. It is used to create a confidence interval, or a range of scores, that likely contains the candidate's true score (the score they would get if there were no measurement error). It does not directly relate to the average score, percentage correct, or percentile rank.
Which of the following BEST describes the primary advantage of using scaled scores instead of raw scores for reporting results of a high-stakes CBT program with multiple test forms?
Answer: They allow for fair comparison of candidate performance across different, potentially unequally difficult, test forms.
The main purpose of scaling scores is to enable fair and direct comparisons of performance even when candidates take different versions (forms) of an exam that may have minor variations in difficulty. A raw score (number correct) on an easier form is not equivalent to the same raw score on a harder form. Scaling adjusts for this, ensuring that a specific scaled score signifies the same level of proficiency regardless of the form taken.
A panel of Subject Matter Experts (SMEs) is convened to determine the passing score for a new certification exam. Each SME independently reviews every test item and estimates the probability that a 'minimally competent candidate' would answer it correctly. The average of these probabilities across all items and all SMEs is then calculated to set the initial pass mark. This process is an application of the:
Answer: Angoff method
This scenario precisely describes the Angoff method, a widely used, item-centered standard-setting procedure. It relies on SME judgments about item difficulty for a hypothetical minimally competent candidate to establish a defensible, criterion-referenced cut score.
When designing a score report for candidates who did not pass a CBT exam, which component is MOST valuable for guiding their future study efforts?
Answer: Diagnostic feedback showing performance by major content domains or sections.
While the overall score indicates the result, diagnostic feedback broken down by content areas (e.g., domains, objectives) provides actionable information. It helps candidates identify their specific areas of weakness so they can focus their remediation efforts effectively. Providing specific incorrect questions is often avoided to protect item security, and percentile rank is less useful for targeted study.
A testing organization needs to provide its board of directors with a high-level overview of exam performance trends from the past year, including pass rates by demographic group and geographic region. To do this, individual candidate score data must be compiled and summarized. This process is known as:
Answer: Data aggregation
Data aggregation is the process of collecting and combining data from multiple sources to present it in a summary format. In this scenario, individual candidate results are being grouped and summarized to show trends and patterns for stakeholders, which is a classic use case for data aggregation.