← All NLP Flashcard Decks

Text Classification Flashcards

7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Text Classification flashcards as text
  1. Which regularization technique is commonly applied in neural text classifiers to prevent overfitting by randomly deactivating neurons during training?

    Answer: Dropout

    Dropout randomly sets a fraction of neuron activations to zero during each training pass, forcing the network to learn redundant representations and reducing overfitting.

  2. What is the key advantage of fine-tuning a pre-trained language model (e.g., BERT) for text classification compared to training from scratch?

    Answer: It leverages general language knowledge learned from large corpora, needing far less task-specific data

    Pre-trained models encode rich linguistic knowledge from massive unlabeled corpora, so fine-tuning requires only a small labeled dataset to adapt them to a specific classification task.

  3. In text classification, what problem does 'class imbalance' refer to?

    Answer: One class having significantly more examples than others in the training set

    Class imbalance occurs when training data has a disproportionate number of examples for some classes, causing the model to be biased toward predicting the majority class.

  4. Which technique involves creating synthetic minority-class examples to address class imbalance in text classification?

    Answer: SMOTE (Synthetic Minority Over-sampling Technique)

    SMOTE generates new synthetic examples for the minority class by interpolating between existing minority-class instances, helping balance the training distribution.

  5. What is hierarchical text classification?

    Answer: Organizing categories into a tree structure where documents are classified from broad to specific

    Hierarchical classification leverages a taxonomy (e.g., Science → Biology → Genetics) to make classification decisions at multiple levels of specificity, improving accuracy on fine-grained categories.

  6. When using cross-validation for evaluating a text classifier, what does 'k-fold' mean?

    Answer: The data is divided into k equal parts, each used once as the validation set while the rest train the model

    In k-fold cross-validation, the dataset is split into k folds; the model trains on k-1 folds and validates on the remaining fold, rotating through all k folds to get a robust performance estimate.

  7. Which loss function is most commonly used when training a neural network for multi-class text classification?

    Answer: Categorical Cross-Entropy

    Categorical cross-entropy measures the divergence between the predicted probability distribution over all classes and the true one-hot encoded label, making it ideal for multi-class problems.