← All NLP Flashcard Decks

Tokenization Flashcards

7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Tokenization flashcards as text
  1. What is the 'unigram language model' tokenization algorithm used in SentencePiece?

    Answer: A probabilistic approach that finds the tokenization maximizing the likelihood under a unigram language model, iteratively pruning the vocabulary

    The unigram algorithm starts with a large vocabulary and iteratively removes tokens that least reduce the corpus likelihood until the target vocabulary size is reached.

  2. What is 'truncation' in the context of tokenizer preprocessing for transformer models?

    Answer: Cutting token sequences that exceed the model's maximum input length

    Truncation clips token sequences to fit within the model's maximum context window, discarding tokens beyond the limit.

  3. What is 'padding' in batch tokenization and why is it necessary?

    Answer: Adding [PAD] tokens to shorter sequences so all sequences in a batch have the same length, enabling efficient tensor operations

    Padding equalizes sequence lengths within a batch so they can be stacked into rectangular tensors required by GPU-accelerated matrix operations.

  4. Which of the following is a key advantage of character-level tokenization over word-level tokenization?

    Answer: No out-of-vocabulary (OOV) problem since any text can be represented

    Character-level tokenization can represent any string using a small fixed alphabet, completely eliminating OOV issues.

  5. What does 'token alignment' refer to when tokenizing text for tasks like Named Entity Recognition (NER)?

    Answer: Mapping subword tokens back to their original word boundaries to correctly assign labels

    In NER, word-level labels must be aligned to subword tokens; typically only the first subword of each word receives the label while others get a special ignore label.

  6. What is the effect of choosing a larger vocabulary size in subword tokenization?

    Answer: Shorter token sequences but more unique token types to learn embeddings for

    A larger vocabulary allows more complete words and longer subwords, reducing sequence length but requiring the model to learn more embedding vectors.

  7. Why might a tokenizer produce different numbers of tokens for the same English word depending on whether it appears at the start of a sentence or mid-sentence in GPT-style BPE?

    Answer: GPT-style BPE treats a leading space as part of the token, so 'dog' and ' dog' are different tokens

    GPT tokenizers prepend a space to most words as a prefix (e.g., 'Ġdog'), distinguishing a word at the start of a sentence from a mid-sentence occurrence.