← All NLP Flashcard Decks

Tokenization Flashcards

7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Tokenization flashcards as text
  1. What is the primary purpose of Byte Pair Encoding (BPE) in tokenization?

    Answer: To iteratively merge the most frequent adjacent byte pairs into a single token

    BPE starts with individual characters and iteratively merges the most frequent adjacent pairs to build a vocabulary of subword units.

  2. Which tokenization strategy is best suited for handling out-of-vocabulary (OOV) words?

    Answer: Subword tokenization

    Subword tokenization breaks unknown words into known subword pieces, effectively eliminating the OOV problem.

  3. In the context of tokenization, what does the term 'vocabulary size' refer to?

    Answer: The total number of unique tokens the tokenizer can produce

    Vocabulary size is the count of distinct tokens (words, subwords, or characters) that a tokenizer's model recognizes.

  4. Which tokenizer is used by the original BERT model?

    Answer: WordPiece

    BERT uses WordPiece tokenization, which builds a vocabulary by merging token pairs that maximize the likelihood of the training data.

  5. What distinguishes WordPiece from BPE in subword tokenization?

    Answer: BPE merges based on frequency, WordPiece merges based on likelihood of the training data

    BPE selects the most frequent pair to merge, while WordPiece selects the pair whose merge maximizes the training corpus likelihood.

  6. What is a 'special token' in the context of modern NLP tokenizers?

    Answer: A reserved token with a specific role such as [CLS], [SEP], or [PAD]

    Special tokens like [CLS], [SEP], and [PAD] are reserved markers added by tokenizers to encode structural information for models.

  7. Which of the following best describes 'tokenization normalization'?

    Answer: Preprocessing steps like lowercasing or accent removal applied before splitting text into tokens

    Normalization refers to preprocessing transformations—such as Unicode normalization, lowercasing, or accent stripping—applied to text before tokenization.