CODESP - Cooperative Organization for the Development of Employee Selection Procedures Test Validation and Reliability Questions and Answers — Questions and Answers
Question 1: A personnel commission uses a panel of three subject matter experts (SMEs) to score candidates' oral interview responses for a "Fire Captain" promotion. After the interviews, a review of the scores reveals a very low correlation among the ratings provided by the three SMEs for the same candidates. This low correlation indicates a problem with which type of reliability?
- Test-retest reliability
- Inter-rater reliability (Correct answer)
- Internal consistency
- Parallel forms reliability
Correct answer: Inter-rater reliability
Inter-rater reliability refers to the degree of agreement among different raters or judges. Since the scores from the three SMEs are inconsistent, it points to a lack of consensus in how they are applying the scoring criteria, which is a direct issue of inter-rater reliability.
Question 2: A public agency wants to conduct a criterion-related validity study for a new pre-employment test for budget analysts. They administer the test to all current budget analysts and simultaneously collect their most recent performance appraisal scores. The agency then correlates the test scores with the performance scores. What specific type of validation strategy is this?
- Predictive validity
- Content validity
- Concurrent validity (Correct answer)
- Construct validity
Correct answer: Concurrent validity
This is an example of concurrent validity, a type of criterion-related validation. It is "concurrent" because the test data (predictor) and the performance data (criterion) are collected at or around the same time from current employees. Predictive validity would involve testing applicants, hiring them, and then collecting performance data at a later date.
Question 3: A public works department uses a test to select equipment operators. The test consistently produces the same scores for individuals who take it multiple times (high test-retest reliability). However, there is no correlation between scores on the test and subsequent measures of on-the-job performance. Which statement best describes this selection test?
- The test is reliable, but not valid. (Correct answer)
- The test is valid, but not reliable.
- The test is both reliable and valid.
- The test is neither reliable nor valid.
Correct answer: The test is reliable, but not valid.
Reliability refers to the consistency of a measure. The scenario states the test is consistent, so it is reliable. Validity refers to whether the test measures what it is supposed to measure and predicts job performance. Since there is no correlation between test scores and job performance, the test is not valid for its intended purpose. A test can be reliable without being valid, but it cannot be valid without being reliable.
Question 4: A police department wants to develop a selection test to measure the abstract characteristic of "integrity," which is difficult to observe directly. To validate this test, they show that scores on it are highly correlated with scores on other established integrity tests and are not correlated with measures of unrelated traits like cognitive ability. This approach is primarily focused on establishing which type of validity?
- Content validity
- Predictive validity
- Concurrent validity
- Construct validity (Correct answer)
Correct answer: Construct validity
Construct validity is the extent to which a test measures the theoretical construct or trait it claims to measure (e.g., integrity, intelligence, leadership). Establishing construct validity often involves demonstrating convergent validity (correlation with similar measures) and discriminant validity (no correlation with dissimilar measures), as described in the scenario.
Question 5: An analyst develops a new 100-item multiple-choice test for a clerical position and calculates a reliability coefficient of .55 using Cronbach's alpha. What is the most accurate interpretation of this result?
- The test is highly reliable and ready for operational use.
- The test scores are strongly correlated with job performance.
- The test is not sufficiently reliable for making selection decisions, as scores contain a large amount of measurement error. (Correct answer)
- 55% of the applicants will pass the test.
Correct answer: The test is not sufficiently reliable for making selection decisions, as scores contain a large amount of measurement error.
A reliability coefficient indicates the extent to which a test is free from random error. A coefficient of .55 is generally considered too low for a selection instrument. Professional standards typically require reliability coefficients of .70 or higher, and preferably .80 or higher, for tests used to make decisions about individuals. A low coefficient means the scores are inconsistent.
Question 6: Which of the following is the most critical step in establishing the content validity of a written examination for a journey-level electrician position?
- Ensuring the test items are a representative sample of the important knowledge, skills, and abilities (KSAs) identified in a thorough job analysis. (Correct answer)
- Correlating test scores with scores on a different, already-validated electrician test.
- Administering the test to applicants and later correlating their scores with their on-the-job performance ratings.
- Demonstrating that the test measures the theoretical construct of "electrical aptitude."
Correct answer: Ensuring the test items are a representative sample of the important knowledge, skills, and abilities (KSAs) identified in a thorough job analysis.
Content validity is established by demonstrating that the content of a selection procedure (e.g., test questions) is a representative sample of the important work behaviors and/or KSAs required for the job. This link is established directly through a comprehensive job analysis.
A personnel commission uses a panel of three subject matter experts (SMEs) to score candidates' oral interview responses for a "Fire Captain" promotion.
After the interviews, a review of the scores reveals a very low correlation among the ratings provided by the three SMEs for the same candidates.
This low correlation indicates a problem with which type of reliability?