โ† All Artificial Intelligence Flashcard Decks

Computer Vision Flashcards

6 cards from real Artificial Intelligence practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 Computer Vision flashcards as text
  1. What is 'transfer learning' typically used for in computer vision applications?

    Answer: Starting with weights from a model pre-trained on ImageNet and fine-tuning on a smaller domain-specific dataset

    Transfer learning leverages rich feature representations learned from large image datasets (e.g., ImageNet) and adapts them to new tasks with less data and compute.

  2. What does the Intersection over Union (IoU) metric measure in object detection?

    Answer: The overlap between predicted and ground-truth bounding boxes divided by their union

    IoU measures how well a predicted bounding box overlaps the ground-truth box; a higher IoU (closer to 1.0) indicates a more accurate localization.

  3. Which technique allows a network to locate the image regions most responsible for a classification decision?

    Answer: Grad-CAM (Gradient-weighted Class Activation Mapping)

    Grad-CAM uses gradients of the target class score flowing into the final convolutional layer to produce a heatmap highlighting discriminative image regions.

  4. What is 'optical flow' in computer vision?

    Answer: The pattern of apparent motion of objects between consecutive video frames

    Optical flow estimates the vector field of pixel motion between frames, used for video analysis, action recognition, and object tracking.

  5. What is 'non-maximum suppression' (NMS) used for in object detection?

    Answer: Eliminating redundant overlapping bounding boxes by keeping only the highest-confidence detection per object

    NMS suppresses duplicate detections by discarding bounding boxes that overlap significantly with a higher-scoring box, ensuring each object has one final detection.

  6. What does a Vision Transformer (ViT) do differently from a CNN when processing images?

    Answer: It splits images into patches and processes them as sequences using self-attention, like a Transformer for text

    ViT divides an image into fixed-size patches, flattens and linearly embeds them, then applies Transformer self-attention over the sequence of patch embeddings.