Natural Language Processing Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Natural Language Processing flashcards as text
In transformer-based models, what is the role of the feed-forward sublayer that follows multi-head attention in each encoder block?
Answer: To apply a position-wise non-linear transformation independently to each token
The position-wise feed-forward network applies two linear transformations with a ReLU (or GELU) activation to each token representation independently, adding non-linearity.
What problem does label smoothing address during training of NLP classifiers?
Answer: Overconfident softmax predictions that hurt generalization
Label smoothing replaces hard 0/1 targets with soft distributions (e.g., 0.9 / 0.1/n), preventing the model from becoming overconfident and improving calibration.
Which technique is used in ELMo to create context-sensitive word representations?
Answer: Bidirectional LSTM language models whose hidden states are combined
ELMo (Embeddings from Language Models) uses a two-layer bidirectional LSTM trained as a language model; representations are task-specific linear combinations of all layer states.
In dependency parsing, what does a 'head' and 'dependent' relationship capture?
Answer: A directed binary grammatical relation between a governing word and a word it governs
In dependency grammar, each word (except the root) has exactly one head; the arc direction and label (e.g., nsubj, dobj) encode the syntactic relation.
What is the vanishing gradient problem in the context of training recurrent neural networks on long sequences?
Answer: Gradients shrink exponentially through time steps, preventing learning of long-range dependencies
When backpropagating through many time steps, repeated multiplication of small Jacobians causes gradients to vanish, making it difficult to capture distant dependencies.
Which of the following best describes the concept of 'zero-shot' learning in NLP?
Answer: Applying a pre-trained model to a task it was never explicitly fine-tuned on
Zero-shot learning leverages a model's pre-trained knowledge and natural language task descriptions to perform tasks without any task-specific training examples.
In the Transformer decoder, why is masked self-attention used during training?
Answer: To prevent the model from attending to future tokens, preserving autoregressive generation
During training, future token positions are masked with -∞ before softmax so the model cannot 'cheat' by looking ahead, mirroring inference-time generation.