Unsupervised Machine Learning Models Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Unsupervised Machine Learning Models flashcards as text
Which linkage criterion in hierarchical clustering tends to produce compact, spherical clusters?
Answer: Ward's linkage
Ward's linkage minimizes the total within-cluster variance at each merge step, producing compact and spherical clusters.
In DBSCAN, a point is classified as a 'border point' when it:
Answer: Has fewer than MinPts neighbors within epsilon but is reachable from a core point
A border point has fewer than MinPts neighbors within epsilon but falls within the epsilon-neighborhood of a core point.
What does the reconstruction error in an autoencoder measure?
Answer: The difference between input and decoder output
Reconstruction error measures how well the decoder can reproduce the original input from the compressed latent representation.
Which of the following best describes the 'manifold hypothesis' as it relates to dimensionality reduction?
Answer: High-dimensional data lies on or near a lower-dimensional manifold embedded in the high-dimensional space
The manifold hypothesis states that real-world high-dimensional data concentrates near a lower-dimensional curved surface (manifold).
In topic modeling with Latent Dirichlet Allocation (LDA), what does the Dirichlet prior control?
Answer: The sparsity of topic-document and word-topic distributions
The Dirichlet hyperparameters alpha and beta control the sparsity of per-document topic distributions and per-topic word distributions respectively.
What is the primary advantage of using UMAP over t-SNE for dimensionality reduction?
Answer: UMAP better preserves global structure and is significantly faster on large datasets
UMAP preserves more global structure than t-SNE and scales better computationally due to its graph-based approximation approach.
In a Gaussian Mixture Model, the E-step of the EM algorithm computes:
Answer: The posterior probability that each data point belongs to each Gaussian component
The E-step computes the responsibilities — the posterior probabilities (soft assignments) of each point belonging to each mixture component given current parameters.