Deep Learning & Neural Networks Flashcards
7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Deep Learning & Neural Networks flashcards as text
What is 'gradient checkpointing' used for in deep learning training?
Answer: Trading compute for memory by recomputing activations during the backward pass
Gradient checkpointing discards intermediate activations and recomputes them during backpropagation, reducing memory at the cost of extra computation.
In the LSTM architecture, which gate controls how much of the previous cell state is retained?
Answer: Forget gate
The forget gate applies a sigmoid-activated mask to the previous cell state, determining what proportion of historical information to carry forward.
What distinguishes 'instance normalization' from 'batch normalization'?
Answer: Instance normalization normalizes each sample's spatial dimensions independently of other samples in the batch
Instance normalization computes mean and variance per sample per channel, making it batch-size-independent and preferred in style transfer tasks.
Which concept in neural architecture search (NAS) allows gradient-based optimization of the architecture by making the search space continuous?
Answer: DARTS (Differentiable Architecture Search)
DARTS relaxes the discrete architecture choice into a continuous softmax over candidate operations, enabling end-to-end gradient optimization.
In contrastive learning (e.g., SimCLR), what is the role of the projection head?
Answer: To map representations to a space where the contrastive loss is applied, then discarded at downstream fine-tuning
The projection head is a small MLP used only during pretraining to compute the contrastive loss; the backbone encoder is used for downstream tasks.
What problem does 'dying ReLU' refer to in deep networks?
Answer: Neurons whose pre-activations are consistently negative, causing them to always output zero and stop learning
When a ReLU neuron's input is always negative its gradient is zero, permanently preventing weight updates and effectively killing that neuron.
Which technique is used to visualize what spatial patterns a specific CNN filter has learned to detect?
Answer: Activation maximization via gradient ascent on the input
Activation maximization synthesizes an input image by ascending the gradient of a target neuron's activation, revealing its preferred pattern.