TKT Assessment and Testing Types 2 — Questions and Answers
Question 1: What is the difference between 'norm-referenced' and 'criterion-referenced' assessment?
- Norm-referenced assessment measures a learner's performance against fixed criteria; criterion-referenced measures against other learners
- Norm-referenced assessment compares a learner's performance to that of a reference group; criterion-referenced measures performance against defined standards or objectives (Correct answer)
- Norm-referenced is used only for speaking; criterion-referenced is used only for writing
- Norm-referenced is formative; criterion-referenced is summative
Correct answer: Norm-referenced assessment compares a learner's performance to that of a reference group; criterion-referenced measures performance against defined standards or objectives
Norm-referenced tests rank learners relative to a comparison group; criterion-referenced tests measure whether learners have achieved specific learning outcomes, regardless of how others perform.
Norm-referenced assessment ranks learners relative to a reference group — results are meaningful only in comparison to other test takers. For example, an IQ test or a standardised national examination that places learners in percentile ranks is norm-referenced. Criterion-referenced assessment, by contrast, measures learner performance against predetermined criteria or standards — for example, a driving test (pass or fail based on specific competencies) or a CEFR-based language test (A1, A2, B1, etc. levels). Most proficiency examinations (including Cambridge qualifications) have criterion-referenced elements. TKT Module 3 covers assessment types, and candidates should be able to distinguish between these two approaches.
Question 2: What does 'diagnostic assessment' aim to do?
- Assign learners to a proficiency level at the end of a course
- Identify specific areas of strength and weakness in learners' language knowledge or skills before or during a course (Correct answer)
- Certify that a learner has reached a particular proficiency level
- Rank learners in order of ability within a class
Correct answer: Identify specific areas of strength and weakness in learners' language knowledge or skills before or during a course
Diagnostic assessment identifies learners' specific strengths and weaknesses in order to inform teaching planning and provide targeted support where needed.
Diagnostic assessment is designed to identify specific gaps in learners' language knowledge or skills so that teaching can be tailored to address those weaknesses. Unlike placement tests (which assign learners to a level) or achievement tests (which measure what has been learned), diagnostic tests reveal fine-grained information about what learners can and cannot do within a given area. For example, a diagnostic grammar test might reveal that intermediate learners consistently confuse the present perfect with the simple past. This information enables teachers to plan targeted remedial work. TKT Module 3 covers the purpose and use of different assessment types.
Question 3: What is 'continuous assessment' and what is its advantage over a single end-of-course exam?
- Assessment that only occurs at the very end of a course in one formal examination
- Assessment that takes place throughout a course, building a picture of learner progress over time through multiple tasks and observations (Correct answer)
- An assessment method where learners are assessed every single day
- A type of assessment where the teacher continually adjusts marks upward
Correct answer: Assessment that takes place throughout a course, building a picture of learner progress over time through multiple tasks and observations
Continuous assessment gathers evidence of learning through multiple assessments across a course, giving a more rounded and valid picture of learner ability than a single high-stakes examination.
Continuous assessment (also called coursework assessment or portfolio-based assessment) involves collecting evidence of learning through a range of tasks and activities throughout the course — such as written assignments, project work, oral presentations, and in-class tests. Its advantages over a single end-of-course exam include: a fuller and more valid picture of learner ability across different skill areas; reduced test anxiety; the opportunity for learners to demonstrate improvement over time; and the provision of regular feedback during the learning process. Many language qualifications now incorporate elements of continuous assessment alongside final examinations.
Question 4: What is a 'portfolio' as an assessment tool in language learning?
- A formal written examination portfolio completed under timed conditions
- A collection of a learner's work, selected and organised to demonstrate their achievements, progress, and reflections over a period of time (Correct answer)
- A government-issued record of a learner's national examination results
- A textbook exercise book used to record and practise grammar
Correct answer: A collection of a learner's work, selected and organised to demonstrate their achievements, progress, and reflections over a period of time
A language portfolio is a purposeful collection of selected learner work that provides evidence of learning, progress, and self-reflection across a course or period of study.
A learner portfolio is a purposefully organised collection of evidence of learning, including samples of written work, audio or video recordings of spoken performance, self-assessments, teacher feedback, and reflective commentary. Portfolios are particularly associated with the European Language Portfolio (ELP), which encourages learners to take ownership of their learning through self-assessment and reflection. Advantages of portfolio assessment include: authenticity; demonstration of process as well as product; learner engagement in self-reflection; and provision of evidence across a range of skills. Portfolios are considered an example of alternative assessment and are increasingly used in ELT programmes. TKT Module 3 addresses alternative assessment tools including portfolios.
Question 5: What does 'face validity' mean in the context of language testing?
- The degree to which a test accurately measures what it claims to measure
- The degree to which a test appears, to test takers and stakeholders, to be relevant and appropriate for its intended purpose (Correct answer)
- The degree to which a test produces consistent results across different test-taking occasions
- The degree to which a test is easy to administer and mark
Correct answer: The degree to which a test appears, to test takers and stakeholders, to be relevant and appropriate for its intended purpose
Face validity refers to whether a test looks valid to its users — whether learners and teachers perceive it as a fair and relevant measure of the skills being assessed.
Face validity is the extent to which a test appears to measure what it is supposed to measure from the perspective of the test takers, teachers, and other stakeholders — even if this is not technically a psychometric property in the same way as construct validity. A test with high face validity is perceived as fair, relevant, and appropriate by those who take or administer it. For example, a speaking test that requires learners to engage in a realistic conversation has higher face validity than one that requires them to read isolated sentences aloud. Face validity is important for maintaining the motivation and cooperation of test takers. TKT Module 3 covers the qualities of a good language test including face validity.
Question 6: What is 'inter-rater reliability' in language assessment?
- The consistency of marks given by the same examiner across different test occasions
- The degree to which different examiners or markers assign the same or similar scores to the same piece of work (Correct answer)
- The fairness of a test in terms of its difficulty level
- The ability of a test to distinguish between high and low achievers
Correct answer: The degree to which different examiners or markers assign the same or similar scores to the same piece of work
Inter-rater reliability measures the consistency of scores given by two or more different examiners to the same test performance, ensuring that marks reflect actual ability rather than examiner subjectivity.
Inter-rater reliability (or inter-marker reliability) refers to the extent to which two or more examiners arrive at similar scores when assessing the same piece of learner work. It is particularly relevant in the assessment of productive skills (speaking and writing) where subjective judgment is involved. High inter-rater reliability is achieved through clear, detailed marking criteria; calibration sessions (where markers discuss and agree on standards); and moderation (cross-checking a sample of marks). Low inter-rater reliability indicates that marks are inconsistent and may not fairly reflect learner ability. Cambridge examinations use detailed mark schemes and systematic moderation to ensure high inter-rater reliability.
What is the difference between 'norm-referenced' and 'criterion-referenced' assessment?