Natural Language Processing Flashcards
6 cards from real Artificial Intelligence practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Natural Language Processing flashcards as text
What architecture is BERT (Bidirectional Encoder Representations from Transformers) based on?
Answer: Transformer encoder
BERT uses only the encoder portion of the Transformer and is pre-trained with bidirectional context, reading the entire sentence at once.
What does 'fine-tuning' a pre-trained language model involve?
Answer: Continuing training on a task-specific labeled dataset to adapt the model to a new task
Fine-tuning continues gradient-based training of a pre-trained model on a smaller task-specific dataset, adapting general representations to the target task.
What is the primary difference between extractive and abstractive text summarization?
Answer: Extractive selects existing sentences from the source; abstractive generates new sentences
Extractive summarization picks and ranks existing sentences from the document, while abstractive summarization generates novel sentences that may not appear verbatim in the source.
What is a 'language model' in NLP?
Answer: A model that assigns probabilities to sequences of words or predicts the next word in a sequence
A language model learns the probability distribution over sequences of words, enabling tasks like text generation, completion, and scoring sentence fluency.
Which NLP task determines whether the relationship between two sentences is entailment, contradiction, or neutral?
Answer: Natural language inference
Natural language inference (NLI) classifies the logical relationship between a premise and a hypothesis sentence as entailment, contradiction, or neutral.
What does the subword tokenization algorithm BPE (Byte Pair Encoding) do?
Answer: Iteratively merges the most frequent character pairs to build a vocabulary of subword units
BPE starts with individual characters and repeatedly merges the most frequent adjacent pair until a target vocabulary size is reached, balancing between word and character tokenization.