Text Preprocessing Flashcards
7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Text Preprocessing flashcards as text
Which regular expression pattern correctly matches any whitespace character in Python's `re` module?
Answer: \s
\s matches any whitespace character including spaces, tabs, newlines, and other Unicode whitespace.
What is Unicode normalization form NFC used for in text preprocessing?
Answer: Composing characters into their canonical composed form
NFC (Canonical Decomposition followed by Canonical Composition) ensures characters with diacritics are stored in a single composed code point rather than multiple characters.
In the context of text preprocessing for social media data, what does 'denoising' typically involve?
Answer: Cleaning hashtags, URLs, emojis, and slang from text
Denoising social media text involves removing or normalizing noisy elements like URLs, hashtags, @ mentions, emojis, and informal spelling.
What does the 'max_features' parameter control in scikit-learn's CountVectorizer?
Answer: The size of the vocabulary built from top frequent terms
max_features limits the vocabulary to the top N most frequent terms across the corpus, reducing dimensionality.
Which tokenization approach handles contractions like "don't" most correctly for downstream NLP tasks?
Answer: Splitting into 'do' and "n't" as a negation marker
Splitting into 'do' and "n't" preserves the negation information as a distinct token, which is linguistically motivated and useful for sentiment analysis.
What is the purpose of applying a minimum document frequency (min_df) threshold in text vectorization?
Answer: To remove rare terms that appear in fewer than N documents
min_df removes terms that appear in fewer than the specified number (or fraction) of documents, eliminating noise from very rare words.
Which of the following is an example of a morphological inflection that stemming is designed to handle?
Answer: 'run', 'runs', 'running', 'ran'
Stemming normalizes morphological variants like run/runs/running/ran to a common stem, reducing vocabulary size.