Scoring Models and Reporting Flashcards
7 cards from real CBT practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Scoring Models and Reporting flashcards as text
What does a scaled score accomplish in CBT reporting?
Answer: It converts raw scores to a common scale across different test forms
Scaled scores equate different test forms to a common scale so that scores are comparable regardless of which version an examinee received.
In Item Response Theory (IRT), the 'b' parameter refers to:
Answer: Item difficulty
The 'b' parameter in IRT represents item difficulty, indicating the ability level at which an examinee has a 50% chance of answering correctly.
Which standard-setting method requires panelists to estimate the probability that a borderline candidate answers each item correctly?
Answer: Angoff method
The Angoff method asks panelists to estimate the probability that a minimally competent (borderline) candidate would answer each item correctly, and the average becomes the cut score.
A score report shows a candidate scored at the 78th percentile. This means:
Answer: The candidate scored higher than 78% of the norm group
A percentile rank of 78 means the candidate performed better than 78% of the reference (norm) group.
What is the primary purpose of vertical scaling in educational CBT programs?
Answer: To compare scores across different grade levels over time
Vertical scaling links tests across different grade levels so student growth can be measured on a single developmental score scale.
In CBT adaptive testing, what does the test information function (TIF) describe?
Answer: How precisely the test measures ability at different levels
The TIF shows the precision (inverse of measurement error) of the test at different ability levels, helping designers ensure accurate measurement where it matters most.
Which score report element is most useful for a candidate who failed a CBT to plan remediation?
Answer: Domain or section subscores
Domain or section subscores show where performance was weakest, enabling targeted remediation rather than re-studying everything.