Text Preprocessing Flashcards
7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Text Preprocessing flashcards as text
Which stemming algorithm is known for being the most aggressive and producing the shortest stems?
Answer: Lancaster Stemmer
The Lancaster (Paice-Husk) stemmer is the most aggressive English stemmer, often over-stemming words to very short roots.
What does TF-IDF stand for in NLP?
Answer: Term Frequency - Inverse Document Frequency
TF-IDF stands for Term Frequency–Inverse Document Frequency, a statistical measure used to evaluate word importance in a document relative to a corpus.
In text preprocessing, what is the purpose of removing hapax legomena?
Answer: To remove words appearing only once in the corpus
Hapax legomena are words occurring only once in a corpus; removing them reduces noise and vocabulary size without significant information loss.
Which of the following best describes the 'bag-of-words' representation after text preprocessing?
Answer: An unordered collection of word frequencies ignoring grammar
Bag-of-words represents text as an unordered set of word counts, discarding grammar and word order information.
What is byte pair encoding (BPE) primarily used for in text preprocessing?
Answer: Subword tokenization to handle out-of-vocabulary words
BPE iteratively merges frequent character pairs to build a subword vocabulary, enabling models to handle rare and unknown words.
When lowercasing text, which scenario presents the greatest risk of information loss?
Answer: Converting named entity 'US' to 'us'
Converting 'US' (United States) to 'us' (pronoun) conflates two entirely different meanings, making disambiguation impossible downstream.
What is the main advantage of using a lemmatizer over a stemmer?
Answer: Lemmatizers return valid dictionary words based on morphological analysis
Lemmatizers use morphological analysis and a dictionary to return actual base forms (lemmas), while stemmers use heuristic rules that may produce non-words.