NCE Assessment 2 — Questions and Answers
Question 1: Which of the following best describes the concept of 'standard error of measurement' (SEM) in psychological testing?
- The average score on a standardized test
- An estimate of the amount of error in an individual's obtained score (Correct answer)
- The difference between parallel forms of a test
- The range of scores obtained by the normative sample
Correct answer: An estimate of the amount of error in an individual's obtained score
The standard error of measurement (SEM) estimates the degree to which an individual's obtained score may deviate from their true score due to measurement error. Smaller SEM values indicate greater precision.
The SEM is calculated as: SEM = SD x sqrt(1 - reliability coefficient). It provides a confidence interval around an obtained score. For example, if a client scores 70 with a SEM of 3, their true score likely falls between 67 and 73. Understanding SEM is critical for interpreting test scores accurately in counseling contexts.
Question 2: A counselor administers the Beck Depression Inventory (BDI-II) to a client. This instrument is best classified as a:
- Projective test
- Performance-based measure
- Self-report inventory (Correct answer)
- Structured clinical interview
Correct answer: Self-report inventory
The BDI-II is a self-report inventory in which clients rate the severity of their own depressive symptoms. It is widely used as a screening and outcome measure.
Self-report inventories like the BDI-II require clients to answer questions about their own thoughts, feelings, and behaviors. Advantages include ease of administration, standardized scoring, and low cost. Limitations include susceptibility to response bias (e.g., social desirability, malingering). The BDI-II measures 21 symptom clusters aligned with DSM criteria for major depressive disorder.
Question 3: When a test measures what it is supposed to measure, this is referred to as:
- Reliability
- Normability
- Validity (Correct answer)
- Standardization
Correct answer: Validity
Validity refers to the degree to which a test actually measures what it claims to measure. It is the most fundamental quality of a psychological test.
Validity has several types: content validity (items represent the domain), criterion-related validity (correlates with external criteria -- concurrent and predictive), and construct validity (measures the theoretical construct). Validity evidence is gathered over time through multiple studies. A test can be reliable without being valid, but a valid test must have some degree of reliability.
Question 4: Which of the following is an example of a norm-referenced assessment?
- A driving test requiring 70% correct to pass
- A CPR certification checklist
- A state licensure exam with a passing cut score
- The Wechsler Adult Intelligence Scale (WAIS) (Correct answer)
Correct answer: The Wechsler Adult Intelligence Scale (WAIS)
The WAIS is norm-referenced because scores are interpreted by comparing an individual's performance to a standardization sample (the norm group), typically yielding standard scores with a mean of 100 and SD of 15.
Norm-referenced assessments rank individuals relative to a representative sample. Common derived scores include percentile ranks, standard scores, stanines, and T-scores. In contrast, criterion-referenced tests (like a driving test) measure mastery of specific skills against a predetermined standard. Intelligence tests, personality inventories, and many achievement tests are norm-referenced.
Question 5: A school counselor is selecting an achievement test for students whose primary language is not English. Which psychometric consideration is MOST important?
- Test-retest reliability
- Split-half reliability
- Cultural and linguistic fairness (Correct answer)
- Number of items on the test
Correct answer: Cultural and linguistic fairness
Cultural and linguistic fairness is the most critical consideration when assessing individuals whose primary language differs from the test's language. Bias can invalidate results and lead to harmful decisions.
Assessment bias occurs when test items, administration procedures, or norms systematically favor or disadvantage specific groups. Differential item functioning (DIF) analysis identifies biased items. For English language learners, counselors should consider using translated or adapted instruments with appropriate norms, non-verbal tests, or interpreter-assisted assessments. Ethical codes require that assessment practices be culturally competent and nondiscriminatory.
Question 6: Which scale of measurement allows for meaningful calculation of ratios (e.g., twice as much)?
- Nominal
- Ordinal
- Interval
- Ratio (Correct answer)
Correct answer: Ratio
Ratio scales have all properties of interval scales plus an absolute zero point, which allows meaningful ratio comparisons. Examples include height, weight, reaction time, and number of correct answers.
The four scales of measurement are: Nominal (categories only, no order), Ordinal (ranked order, unequal intervals), Interval (equal intervals, no true zero -- e.g., temperature in Celsius), and Ratio (equal intervals plus true zero -- e.g., weight, height, income). Most psychological constructs use interval-level scoring via standardized tests, while ratio-level measurement is rare in counseling assessment.
Which of the following best describes the concept of 'standard error of measurement' (SEM) in psychological testing?