โ† All NLP Flashcard Decks

Tokenization Flashcards

7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Tokenization flashcards as text
  1. What problem does the SentencePiece library solve that earlier tokenizers did not?

    Answer: It performs tokenization without relying on language-specific whitespace rules, treating raw text as input

    SentencePiece treats the input as a raw stream of Unicode characters, making it language-agnostic and not dependent on whitespace-delimited words.

  2. In tokenization, what is a 'token ID' (or input ID)?

    Answer: The integer index of a token in the tokenizer's vocabulary

    A token ID is the integer index that maps a token to its row in the model's embedding matrix.

  3. What is 'detokenization' in NLP?

    Answer: Converting a sequence of tokens back into a human-readable string

    Detokenization reconstructs the original (or near-original) text string from a list of tokens, reversing the tokenization process.

  4. Which tokenization approach is used by GPT-2 and GPT-3?

    Answer: Byte-level BPE

    GPT-2 and GPT-3 use byte-level BPE, which operates on UTF-8 bytes rather than Unicode characters, ensuring every string can be tokenized.

  5. What does 'attention mask' indicate in the output of a Hugging Face tokenizer?

    Answer: Which token positions the model should attend to (1) vs. ignore as padding (0)

    The attention mask is a binary tensor marking real tokens with 1 and padding tokens with 0, telling the model which positions to attend to.

  6. What is the 'unknown token' ([UNK]) used for in word-level tokenization?

    Answer: To replace any word not found in the vocabulary during inference

    The [UNK] token is a fallback that replaces any input word absent from the fixed vocabulary, grouping all OOV words into a single representation.

  7. When using subword tokenization, the word 'unbelievably' might be split into which of the following?

    Answer: ['un', '##believ', '##ably']

    WordPiece-style subword tokenizers break words into frequent subword pieces, using '##' to indicate continuation of the previous token.