Tokenization Flashcards
7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Tokenization flashcards as text
What problem does the SentencePiece library solve that earlier tokenizers did not?
Answer: It performs tokenization without relying on language-specific whitespace rules, treating raw text as input
SentencePiece treats the input as a raw stream of Unicode characters, making it language-agnostic and not dependent on whitespace-delimited words.
In tokenization, what is a 'token ID' (or input ID)?
Answer: The integer index of a token in the tokenizer's vocabulary
A token ID is the integer index that maps a token to its row in the model's embedding matrix.
What is 'detokenization' in NLP?
Answer: Converting a sequence of tokens back into a human-readable string
Detokenization reconstructs the original (or near-original) text string from a list of tokens, reversing the tokenization process.
Which tokenization approach is used by GPT-2 and GPT-3?
Answer: Byte-level BPE
GPT-2 and GPT-3 use byte-level BPE, which operates on UTF-8 bytes rather than Unicode characters, ensuring every string can be tokenized.
What does 'attention mask' indicate in the output of a Hugging Face tokenizer?
Answer: Which token positions the model should attend to (1) vs. ignore as padding (0)
The attention mask is a binary tensor marking real tokens with 1 and padding tokens with 0, telling the model which positions to attend to.
What is the 'unknown token' ([UNK]) used for in word-level tokenization?
Answer: To replace any word not found in the vocabulary during inference
The [UNK] token is a fallback that replaces any input word absent from the fixed vocabulary, grouping all OOV words into a single representation.
When using subword tokenization, the word 'unbelievably' might be split into which of the following?
Answer: ['un', '##believ', '##ably']
WordPiece-style subword tokenizers break words into frequent subword pieces, using '##' to indicate continuation of the previous token.