Natural Language Processing Flashcards
6 cards from real Artificial Intelligence practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 Natural Language Processing flashcards as text
What is tokenization in natural language processing?
Answer: Splitting text into meaningful units such as words or subwords
Tokenization breaks raw text into tokens (words, subwords, or characters) that serve as the basic units for NLP model input.
What does TF-IDF measure in information retrieval and NLP?
Answer: The importance of a word in a document relative to a corpus
TF-IDF multiplies term frequency (how often a word appears in a document) by inverse document frequency (penalizing words common across many documents).
Which word embedding model learns vector representations by predicting surrounding words (skip-gram) or predicting a word from context (CBOW)?
Answer: Word2Vec
Word2Vec offers two architectures: skip-gram predicts context words from a target, and CBOW predicts a target word from its context window.
What NLP task involves labeling each token in a sentence with its grammatical role (noun, verb, etc.)?
Answer: Part-of-speech tagging
Part-of-speech (POS) tagging assigns grammatical categories to each token, such as noun, verb, adjective, or adverb.
What is the 'attention mechanism' in neural NLP models?
Answer: A method that allows the model to weight the importance of different input tokens when producing an output
Attention lets the model dynamically focus on relevant parts of the input sequence when generating each output token, enabling better handling of long-range dependencies.
What is named entity recognition (NER) in NLP?
Answer: Identifying and classifying real-world entities such as people, organizations, and locations in text
NER identifies spans of text that refer to specific entity categories (persons, organizations, dates, etc.) and tags them accordingly.