Text Preprocessing Flashcards
7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Text Preprocessing flashcards as text
What is the 'vocabulary explosion' problem in NLP and how does subword tokenization address it?
Answer: Unbounded vocabulary from rare/novel words; subword units cap vocabulary at a fixed size
Subword tokenization (BPE, WordPiece) splits rare words into known subword units, enabling a fixed vocabulary to handle any new word.
Which of the following describes the difference between type and token in corpus linguistics?
Answer: A token is each individual word occurrence; a type is a unique word form
Tokens are all word occurrences in text (including repeats), while types are the distinct unique words; 'the cat sat on the mat' has 6 tokens but 5 types.
In preprocessing pipeline design, why is it generally recommended to fit vectorizers only on training data?
Answer: Fitting on test data causes data leakage by letting the model learn test set statistics
Fitting on test data leaks information about the test distribution into the model, causing overly optimistic evaluation metrics.
What preprocessing technique is specifically designed to handle the informal spelling variations in user-generated text (e.g., 'goooood', 'pleaseeeee')?
Answer: Character repetition normalization
Character repetition normalization reduces elongated words (e.g., 'goooood' → 'good') by collapsing repeated characters to a standard form.
How does the Punkt sentence tokenizer in NLTK handle abbreviations like 'Dr.' and 'Mr.'?
Answer: It learns abbreviations from training data and treats their periods as non-sentence-final
Punkt is an unsupervised algorithm that learns abbreviation patterns from text, avoiding false sentence splits on titles and abbreviations.
What is the role of a 'text normalization pipeline' in an NLP system?
Answer: To apply a sequence of preprocessing steps transforming raw text into a consistent, clean representation
A text normalization pipeline chains steps like lowercasing, tokenization, stop word removal, and stemming/lemmatization into a reproducible preprocessing workflow.
Why might preserving case information be important when preprocessing text for a Named Entity Recognition (NER) task?
Answer: Capitalization is a strong signal for identifying proper nouns and named entities
In English, named entities like person names, places, and organizations are typically capitalized, making case a valuable feature for NER models.