CODESP Test Validation and Reliability 2 — Questions and Answers
Question 1: Which type of validity evidence examines whether test scores correlate with an external criterion measured at the same time?
- Predictive validity
- Concurrent validity (Correct answer)
- Content validity
- Construct validity
Correct answer: Concurrent validity
Concurrent validity is established when test scores correlate with a criterion measure collected simultaneously.
Question 2: A reliability coefficient of 0.50 for a selection test would generally be considered:
- Excellent and exceeds standards
- Acceptable for high-stakes decisions
- Too low for most selection purposes (Correct answer)
- Appropriate only for cognitive tests
Correct answer: Too low for most selection purposes
A reliability coefficient of 0.50 is generally considered too low for employment selection tests, which typically require 0.70 or higher.
Question 3: What does the standard error of measurement (SEM) primarily tell us?
- The average score on the test
- The expected variation in an individual's scores across repeated administrations (Correct answer)
- The correlation between two test forms
- The difference between observed and true scores for a group
Correct answer: The expected variation in an individual's scores across repeated administrations
SEM estimates how much an individual's observed score is likely to vary if the test were administered multiple times.
Question 4: When a test is used to predict future job performance measured six months after hiring, this is an example of:
- Concurrent validity
- Content validity
- Predictive validity (Correct answer)
- Face validity
Correct answer: Predictive validity
Predictive validity involves collecting criterion data after a time delay following test administration.
Question 5: Which reliability method requires only a single test administration?
- Test-retest reliability
- Parallel forms reliability
- Internal consistency reliability (Correct answer)
- Inter-rater reliability
Correct answer: Internal consistency reliability
Internal consistency methods (e.g., Cronbach's alpha, KR-20) assess reliability using data from a single test administration.
Question 6: Differential item functioning (DIF) analysis is used to identify items that:
- Are too difficult for all test takers
- Perform differently for subgroups after matching on ability (Correct answer)
- Have low point-biserial correlations
- Are outside the content domain
Correct answer: Perform differently for subgroups after matching on ability
DIF flags items that favor or disadvantage a particular subgroup after controlling for overall ability level.
Question 7: In classical test theory, a person's observed score is defined as:
- True score minus measurement error
- True score plus measurement error (Correct answer)
- True score divided by reliability
- Error score minus true score
Correct answer: True score plus measurement error
Classical test theory states that Observed Score = True Score + Error Score.
Which type of validity evidence examines whether test scores correlate with an external criterion measured at the same time?