โ† All NLP Flashcard Decks

Tokenization Flashcards

7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Tokenization flashcards as text
  1. What is 'offset mapping' in tokenizer output and what is it used for?

    Answer: A list of (start, end) character positions in the original text corresponding to each token, used for span extraction tasks

    Offset mapping records the character-level start and end positions of each token in the original input string, enabling tasks like question answering to return answer spans.

  2. In the Hugging Face Tokenizers library, what does 'fast' vs. 'slow' tokenizer refer to?

    Answer: Fast tokenizers are implemented in Rust and support features like offset mapping; slow tokenizers are pure Python implementations

    Hugging Face 'fast' tokenizers are backed by the Rust-based tokenizers library, offering speed improvements and additional features like offset mapping not available in Python-based 'slow' tokenizers.

  3. What is 'multilingual tokenization' and what challenge does it address?

    Answer: Building a shared vocabulary that covers multiple languages so a single model can process text in many languages

    Multilingual tokenization trains a single subword vocabulary on text from many languages, enabling cross-lingual transfer in models like mBERT and XLM-RoBERTa.

  4. What does 'tokenizer overfitting' mean in practice?

    Answer: A vocabulary trained on a domain-specific corpus that performs poorly on general text because its subwords are too specialized

    A tokenizer trained on narrow domain text may learn highly domain-specific subword units that fragment out-of-domain text inefficiently.

  5. Which statement correctly describes how Chinese text is typically tokenized in BERT-based models?

    Answer: Each Chinese character is treated as an individual token since Chinese doesn't use whitespace between words

    Chinese BERT tokenization inserts spaces around every character before applying WordPiece, effectively treating each character as a basic unit.

  6. What is 'token budget' or 'token limit' and why does it matter for applications using large language models?

    Answer: The maximum number of tokens (input + output) a model can process in one request, affecting cost and what fits in context

    LLM APIs charge per token and enforce context window limits, so understanding token counts is critical for managing cost and ensuring inputs fit within the model's context.

  7. What is a 'tokenizer mismatch' and why is it a critical issue in NLP deployments?

    Answer: Using a different tokenizer at inference time than was used during model training, causing the model to receive unexpected token ID sequences

    A tokenizer mismatch means the model receives token IDs that don't correspond to the embeddings it learned, causing degraded or nonsensical outputs.