Advanced Topics & Theory Flashcards
7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Advanced Topics & Theory flashcards as text
What is the primary purpose of the Transformer's multi-head attention mechanism?
Answer: To allow the model to jointly attend to information from different representation subspaces
Multi-head attention runs attention in parallel across multiple subspaces, letting the model capture different types of relationships simultaneously.
Which decoding strategy samples from the top-k most probable next tokens at each step?
Answer: Top-k sampling
Top-k sampling restricts the sampling pool to the k highest-probability tokens, balancing diversity and coherence.
In the context of NLP, what does 'perplexity' measure?
Answer: How well a probability model predicts a sample
Perplexity is the exponentiated average negative log-likelihood of a test set, measuring how uncertain the model is about each token.
What is 'catastrophic forgetting' in continual learning for NLP models?
Answer: The tendency of a neural network to lose previously learned information when trained on new tasks
Catastrophic forgetting occurs when optimizing for a new task overwrites the weights that encoded prior knowledge.
Which technique uses a smaller 'student' model to mimic the output distribution of a larger 'teacher' model?
Answer: Knowledge distillation
Knowledge distillation trains a compact student model to match the soft probability outputs of a pre-trained teacher, compressing knowledge without large accuracy loss.
What is the role of the key-query-value (K-Q-V) structure in self-attention?
Answer: Queries match against keys to produce attention weights, then weighted sum of values forms the output
Each token's query attends over all keys; the resulting attention weights are applied to values to produce a context-aware representation.
Which regularization technique randomly masks input tokens during pre-training, requiring the model to reconstruct them?
Answer: Masked language modeling (MLM)
MLM, used in BERT, masks 15% of tokens and trains the model to predict the originals, learning bidirectional context.