Deep Learning & Neural Networks Flashcards
7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Deep Learning & Neural Networks flashcards as text
What is 'weight tying' in language models such as GPT, and why is it used?
Answer: Sharing the input embedding matrix with the output projection layer to reduce parameters and improve generalization
Weight tying reuses the token embedding matrix as the output softmax projection, significantly reducing parameters and linking input and output token representations.
In knowledge distillation, what does the 'dark knowledge' contained in the teacher's soft labels refer to?
Answer: The probability mass the teacher assigns to incorrect classes, encoding inter-class similarity
Soft targets reveal that the teacher considers some wrong classes more plausible than others, encoding rich similarity structure the student can exploit.
Which property of sinusoidal positional encodings in the original transformer makes them potentially generalizable to sequence lengths unseen during training?
Answer: They use fixed frequencies that can represent any integer position deterministically
Fixed sinusoidal functions with geometric frequency progressions can encode any position as a unique vector, unlike learned embeddings limited to seen positions.
What distinguishes a variational autoencoder (VAE) from a standard autoencoder in terms of the latent space?
Answer: The VAE encodes inputs as probability distributions and samples from them, imposing a prior on the latent space
A VAE encoder outputs mean and variance parameters; sampling introduces stochasticity and the KL term regularizes the latent space toward a prior (e.g., standard Gaussian).
In self-supervised learning with masked autoencoders (MAE), what percentage of image patches is typically masked during pretraining, and why is such a high ratio used?
Answer: 75%, to prevent the model from relying on local texture interpolation and force semantic understanding
Masking 75% of patches forces the model to infer semantically meaningful content rather than simply interpolating from adjacent pixels.
What is 'mode collapse' in generative adversarial networks?
Answer: The generator produces only a limited variety of outputs, ignoring the full target data distribution
Mode collapse occurs when the generator learns to fool the discriminator with a small subset of realistic outputs, failing to capture the data distribution's diversity.
Which of the following best describes the purpose of 'layer-wise adaptive rate scaling' (LARS) in training large batch deep learning models?
Answer: It computes a per-layer learning rate by scaling with the ratio of weight norm to gradient norm, enabling large-batch stability
LARS adapts the learning rate per layer so that the update magnitude stays proportional to the weight magnitude, preventing instability when using very large batch sizes.