Machine Translation Flashcards
7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Machine Translation flashcards as text
Which approach to MT does NOT require any parallel bilingual data during training?
Answer: Unsupervised MT using monolingual corpora only
Unsupervised MT methods (e.g., using denoising autoencoders and back-translation on monolingual data) require no parallel sentences.
What is the main advantage of using subword tokenization (e.g., BPE) in NMT?
Answer: Reduces vocabulary size while handling rare and OOV words
Byte-pair encoding splits rare words into frequent subword units, giving the model coverage of open-vocabulary words without an extremely large vocabulary.
In multilingual NMT, what is the 'language token' prepended to the source sentence used for?
Answer: Signaling the desired target language to a single shared model
A target-language tag (e.g., ) prepended to the input tells a universal NMT model which language to generate.
Which IBM Model introduced the concept of word fertility in MT alignment?
Answer: IBM Model 3
IBM Model 3 introduced fertility, the number of target words that a single source word generates.
What does 'zero-shot translation' mean in a multilingual NMT system?
Answer: Translating between a language pair never seen together in training
Zero-shot translation is the ability to translate between two languages that were never paired together in training data.
Which technique allows an NMT model to incorporate external knowledge such as a domain-specific glossary at inference time?
Answer: Hard lexical constraints via constrained beam search
Constrained decoding forces the beam search to include specified terms, implementing hard lexical constraints without retraining.
In domain adaptation for MT, what is 'catastrophic forgetting'?
Answer: Fine-tuning on in-domain data degrades performance on the original general domain
Catastrophic forgetting occurs when fine-tuning on a new domain overwrites general knowledge learned during pretraining, hurting out-of-domain performance.