โ† All NLP Flashcard Decks

Machine Translation Flashcards

7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Machine Translation flashcards as text
  1. Which evaluation metric for MT is specifically designed to correlate better with human judgments by incorporating synonyms and paraphrases?

    Answer: METEOR

    METEOR uses stemming, synonym matching, and paraphrase tables to better match human evaluations compared to BLEU's strict n-gram overlap.

  2. What is the 'exposure bias' problem in sequence-to-sequence MT training?

    Answer: Training uses ground-truth tokens as input, but inference uses model predictions, causing a distribution mismatch

    Exposure bias arises because teacher-forcing during training always provides correct previous tokens, whereas at test time the model conditions on its own (possibly wrong) outputs.

  3. In phrase-based SMT, what does a 'distortion model' control?

    Answer: The reordering (permutation) of translated phrases relative to source order

    The distortion model assigns a cost to jumping between non-adjacent source phrases, controlling how much the target word order can differ from the source.

  4. Which post-editing metric measures the minimum number of edit operations needed to correct an MT output into an acceptable translation?

    Answer: TER (Translation Edit Rate)

    TER counts the number of shifts, insertions, deletions, and substitutions needed to convert the MT hypothesis into a reference translation.

  5. What is 'pivot translation' and when is it used?

    Answer: Using an intermediate language to translate between two languages that lack parallel data

    Pivot (bridge) translation routes low-resource language pairs through a high-resource pivot language (often English) when direct parallel data is unavailable.

  6. Which component of a neural MT system is responsible for generating a fixed-length context vector in older encoder-decoder architectures (pre-attention)?

    Answer: Final encoder hidden state

    In early seq2seq models, the last encoder hidden state compressed the entire source sentence into a single context vector passed to the decoder.

  7. What does 'document-level MT' aim to improve over sentence-level MT?

    Answer: Coherence, coreference resolution, and consistency across sentences

    Document-level MT models consider inter-sentence context to correctly resolve pronouns, maintain consistent terminology, and improve discourse coherence.