Computer Vision Flashcards
6 cards from real Artificial Intelligence practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Computer Vision flashcards as text
What is the primary purpose of a convolutional neural network (CNN) in computer vision tasks?
Answer: To automatically learn hierarchical spatial features from images for recognition tasks
CNNs use learnable convolutional filters to extract local spatial features at multiple scales, building hierarchical representations from edges to complex objects.
What does 'image segmentation' mean in computer vision?
Answer: Assigning a class label to every pixel in the image
Image segmentation partitions an image into regions, assigning class labels to individual pixels for semantic segmentation or instance boundaries for instance segmentation.
Which computer vision task outputs both a class label and a bounding box for each detected object?
Answer: Object detection
Object detection identifies what objects are present and where they are by predicting class labels and bounding box coordinates for each detected instance.
What does data augmentation in computer vision typically involve?
Answer: Applying random transformations like flips, rotations, and crops to training images to improve generalization
Data augmentation artificially expands training data by applying label-preserving transformations, helping models become robust to variations in orientation, scale, and lighting.
What is the role of anchor boxes in object detection models like YOLO and Faster R-CNN?
Answer: They serve as pre-defined reference bounding boxes of various aspect ratios that the model adjusts to fit detected objects
Anchor boxes are a set of predefined bounding box shapes; the detector predicts offsets from these anchors to localize objects of different sizes and aspect ratios.
What is 'image classification' in computer vision?
Answer: Assigning a single label to an entire image based on its content
Image classification takes an image as input and outputs a single class label representing the dominant subject or scene category.