CODESP Test Validation and Reliability 5 — Questions and Answers
Question 1: According to the Society for Industrial-Organizational Psychology (SIOP) Principles, which of the following is the LEAST acceptable justification for using a test with adverse impact?
- Evidence of strong criterion-related validity
- Evidence of strong content validity for critical job tasks
- A belief that the test 'looks' job-related to applicants (Correct answer)
- Validity generalization evidence from similar jobs
Correct answer: A belief that the test 'looks' job-related to applicants
Face validity — the perception that a test looks job-related — is not a sufficient scientific or legal justification for using a test with adverse impact.
Question 2: Coefficient alpha will be artificially INFLATED when:
- Items are negatively correlated with each other
- The test contains many items measuring distinct facets of a construct
- Items have very high inter-correlations due to item overlap or redundancy (Correct answer)
- The test is administered under speeded conditions
Correct answer: Items have very high inter-correlations due to item overlap or redundancy
When items are redundant or essentially paraphrase one another, alpha rises artificially without reflecting true reliability.
Question 3: A test manual reports a test-retest reliability of 0.85 with a 6-month interval. Compared to a 2-week interval, the 6-month coefficient is MOST likely:
- Higher, because more practice effects accumulate
- Lower, because real changes in the construct may occur over time (Correct answer)
- The same, because reliability is a fixed property of a test
- Higher, because memory of previous answers fades
Correct answer: Lower, because real changes in the construct may occur over time
Longer time intervals allow true changes in the measured attribute to occur, which lowers test-retest reliability coefficients.
Question 4: Which approach to establishing validity is MOST appropriate for a newly developed test when no local criterion data are yet available?
- Concurrent validity study
- Validity generalization / transportability argument (Correct answer)
- Predictive validity study with incumbents
- Differential item functioning analysis
Correct answer: Validity generalization / transportability argument
When local criterion data are unavailable, validity generalization allows practitioners to use accumulated evidence from prior studies to justify test use.
Question 5: The four-fifths (4/5) rule in the Uniform Guidelines is used to evaluate:
- Whether a test has acceptable reliability
- Whether adverse impact exists in selection rates between groups (Correct answer)
- Whether content validity evidence meets minimum standards
- Whether a criterion measure is job-relevant
Correct answer: Whether adverse impact exists in selection rates between groups
The 4/5 rule compares the selection rate of a protected group to that of the highest-selected group to detect adverse impact.
Question 6: A personnel psychologist corrects an observed validity coefficient for criterion unreliability. The corrected coefficient will be:
- Lower than the observed coefficient
- The same as the observed coefficient
- Higher than the observed coefficient (Correct answer)
- Meaningless without range restriction correction
Correct answer: Higher than the observed coefficient
Correcting for criterion unreliability removes attenuation caused by criterion measurement error, yielding a higher estimated true validity.
Question 7: Which of the following BEST describes the concept of test bias in the context of employment testing?
- Any test that results in different pass rates across demographic groups
- Systematic over- or under-prediction of criterion scores for a subgroup (Correct answer)
- A test that applicants perceive as unfair or irrelevant
- A test with low content validity for a specific subgroup
Correct answer: Systematic over- or under-prediction of criterion scores for a subgroup
Test bias in psychometrics refers specifically to differential prediction — when a test systematically over- or under-predicts job performance for a particular group.
According to the Society for Industrial-Organizational Psychology (SIOP) Principles, which of the following is the LEAST acceptable justification for using a test with adverse impact?