Word Embeddings Flashcards
7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Word Embeddings flashcards as text
How do contextualized embeddings from ELMo differ fundamentally from word2vec embeddings?
Answer: ELMo produces a different vector for the same word depending on its sentence context using a biLSTM
ELMo generates dynamic, context-sensitive embeddings by passing the full sentence through a bidirectional LSTM, so 'bank' gets different vectors in different sentences.
What is 'transfer learning' in the context of pre-trained word embeddings?
Answer: Using embeddings learned on a large general corpus as the starting point for a domain-specific task
Pre-trained embeddings encode general language knowledge that can be transferred to downstream tasks, especially when task-specific data is scarce.
In a bag-of-words document representation, what information is lost compared to using averaged word embeddings?
Answer: Word order and compositionality are discarded in both, but BoW also ignores semantic similarity between different words
BoW treats vocabulary as a sparse orthogonal space where different words have zero similarity, whereas averaged embeddings encode semantic relatedness between words.
What is 'domain adaptation' when applied to word embeddings?
Answer: Further training or fine-tuning general embeddings on in-domain text to capture domain-specific vocabulary and meaning
General-purpose embeddings may not capture specialized terminology well, so continuing training on domain text (e.g., biomedical papers) adapts them to the target domain.
Which technique allows word embeddings from two different languages to be mapped into a shared cross-lingual vector space?
Answer: Learning a linear transformation (rotation) using bilingual anchor word pairs
Cross-lingual alignment methods like VecMap find a rotation matrix that maps monolingual embeddings into a shared space using bilingual seed lexicons.
What does the term 'out-of-vocabulary (OOV)' mean for a word embedding model, and how does fastText mitigate it?
Answer: OOV words were not seen during training; fastText mitigates this by composing embeddings from character n-grams
FastText can approximate embeddings for OOV words by summing the n-gram vectors of their character substrings, which often share morphemes with known words.
What bias problem has been documented in word embeddings trained on large web corpora?
Answer: Embeddings reflect societal biases, e.g., associating 'programmer' more closely with male terms than female terms
Bolukbasi et al. (2016) showed that word2vec embeddings encode gender stereotypes present in training corpora, such as gender-occupation associations.