Machine Translation Flashcards
7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Machine Translation flashcards as text
Which evaluation metric for MT is specifically designed to correlate better with human judgments by incorporating synonyms and paraphrases?
Answer: METEOR
METEOR uses stemming, synonym matching, and paraphrase tables to better match human evaluations compared to BLEU's strict n-gram overlap.
What is the 'exposure bias' problem in sequence-to-sequence MT training?
Answer: Training uses ground-truth tokens as input, but inference uses model predictions, causing a distribution mismatch
Exposure bias arises because teacher-forcing during training always provides correct previous tokens, whereas at test time the model conditions on its own (possibly wrong) outputs.
In phrase-based SMT, what does a 'distortion model' control?
Answer: The reordering (permutation) of translated phrases relative to source order
The distortion model assigns a cost to jumping between non-adjacent source phrases, controlling how much the target word order can differ from the source.
Which post-editing metric measures the minimum number of edit operations needed to correct an MT output into an acceptable translation?
Answer: TER (Translation Edit Rate)
TER counts the number of shifts, insertions, deletions, and substitutions needed to convert the MT hypothesis into a reference translation.
What is 'pivot translation' and when is it used?
Answer: Using an intermediate language to translate between two languages that lack parallel data
Pivot (bridge) translation routes low-resource language pairs through a high-resource pivot language (often English) when direct parallel data is unavailable.
Which component of a neural MT system is responsible for generating a fixed-length context vector in older encoder-decoder architectures (pre-attention)?
Answer: Final encoder hidden state
In early seq2seq models, the last encoder hidden state compressed the entire source sentence into a single context vector passed to the decoder.
What does 'document-level MT' aim to improve over sentence-level MT?
Answer: Coherence, coreference resolution, and consistency across sentences
Document-level MT models consider inter-sentence context to correctly resolve pronouns, maintain consistent terminology, and improve discourse coherence.