← All NLP Flashcard Decks

Text Preprocessing Flashcards

7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Text Preprocessing flashcards as text
  1. Which stemming algorithm is known for being the most aggressive and producing the shortest stems?

    Answer: Lancaster Stemmer

    The Lancaster (Paice-Husk) stemmer is the most aggressive English stemmer, often over-stemming words to very short roots.

  2. What does TF-IDF stand for in NLP?

    Answer: Term Frequency - Inverse Document Frequency

    TF-IDF stands for Term Frequency–Inverse Document Frequency, a statistical measure used to evaluate word importance in a document relative to a corpus.

  3. In text preprocessing, what is the purpose of removing hapax legomena?

    Answer: To remove words appearing only once in the corpus

    Hapax legomena are words occurring only once in a corpus; removing them reduces noise and vocabulary size without significant information loss.

  4. Which of the following best describes the 'bag-of-words' representation after text preprocessing?

    Answer: An unordered collection of word frequencies ignoring grammar

    Bag-of-words represents text as an unordered set of word counts, discarding grammar and word order information.

  5. What is byte pair encoding (BPE) primarily used for in text preprocessing?

    Answer: Subword tokenization to handle out-of-vocabulary words

    BPE iteratively merges frequent character pairs to build a subword vocabulary, enabling models to handle rare and unknown words.

  6. When lowercasing text, which scenario presents the greatest risk of information loss?

    Answer: Converting named entity 'US' to 'us'

    Converting 'US' (United States) to 'us' (pronoun) conflates two entirely different meanings, making disambiguation impossible downstream.

  7. What is the main advantage of using a lemmatizer over a stemmer?

    Answer: Lemmatizers return valid dictionary words based on morphological analysis

    Lemmatizers use morphological analysis and a dictionary to return actual base forms (lemmas), while stemmers use heuristic rules that may produce non-words.