โ† All NLP Flashcard Decks

Text Preprocessing Flashcards

7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Text Preprocessing flashcards as text
  1. What preprocessing challenge arises specifically when handling Chinese or Japanese text that doesn't exist with English?

    Answer: Word boundary segmentation since words are not space-delimited

    Chinese and Japanese text lacks spaces between words, so word segmentation algorithms are required before any word-level processing.

  2. What is the effect of applying log normalization to term frequency (TF) values?

    Answer: It reduces the impact of very high frequency terms by compressing the scale

    Log normalization (e.g., 1 + log(tf)) prevents very frequent terms from dominating by compressing large frequency differences.

  3. In preprocessing a legal or medical corpus, why might domain-specific stop words be added beyond the standard list?

    Answer: Common domain terms like 'whereas' or 'patient' appear frequently but carry little discriminative value

    High-frequency domain terms that appear in nearly every document (like 'patient' in medical text) behave like stop words and can be added to a custom list.

  4. What is a 'character n-gram' and why is it useful in text preprocessing?

    Answer: A sequence of n consecutive characters, useful for handling misspellings and morphology

    Character n-grams capture subword patterns, making models robust to typos, morphological variants, and out-of-vocabulary words.

  5. Which text normalization step would convert '5 kilometers' and '5 km' to a consistent representation?

    Answer: Unit normalization or entity normalization

    Unit normalization maps different representations of the same measurement (km, kilometers, kilometre) to a canonical form.

  6. When preprocessing text for a sentiment analysis model, removing which of the following would likely hurt performance the most?

    Answer: Negation words like 'not', 'never', 'hardly'

    Negation words fundamentally flip sentiment polarity (e.g., 'not good' = negative), so removing them severely degrades sentiment classification.

  7. What is the primary purpose of sentence boundary detection (SBD) in text preprocessing?

    Answer: To split a document into individual sentences for downstream processing

    SBD (also called sentence segmentation) correctly identifies where one sentence ends and another begins, enabling sentence-level analysis.