← All MS-DS Master of Data science Flashcard Decks

Unsupervised Learning Techniques Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Unsupervised Learning Techniques flashcards as text
  1. What is the computational complexity of a single iteration of k-means on n data points with k clusters and d dimensions?

    Answer: O(nkd)

    Each of the n points must compute distances to each of the k cluster centroids in d dimensions, giving O(nkd) per iteration.

  2. Which method is used by UMAP to preserve both local and global structure, distinguishing it from t-SNE?

    Answer: UMAP constructs a fuzzy topological representation and optimizes its cross-entropy with a low-dimensional analog

    UMAP is grounded in Riemannian geometry and algebraic topology, building a fuzzy simplicial complex and minimizing its cross-entropy with a low-dimensional representation, preserving more global structure than t-SNE.

  3. What is the 'elbow method' used to determine in unsupervised learning?

    Answer: The appropriate number of clusters k in k-means by plotting inertia vs. k

    The elbow method plots within-cluster sum of squares (inertia) against k; the 'elbow' point where improvement diminishes suggests the optimal k.

  4. A researcher applies hierarchical agglomerative clustering with single linkage to data with two elongated, touching clusters. What artifact is likely to occur?

    Answer: Chaining: the two clusters will merge prematurely into one elongated cluster

    Single linkage uses the minimum pairwise distance between clusters, making it susceptible to chaining where a bridge of close points merges two distinct elongated clusters early.

  5. In a Variational Autoencoder (VAE), what is the purpose of the reparameterization trick?

    Answer: To allow gradients to flow through the stochastic sampling step during backpropagation

    The reparameterization trick expresses z = μ + σ·ε (ε ~ N(0,I)), making the random node deterministic w.r.t. ε and allowing gradient flow through μ and σ.

  6. Which metric is specifically designed to evaluate clustering quality when ground-truth labels are available?

    Answer: Adjusted Rand Index (ARI)

    The Adjusted Rand Index measures the similarity between predicted cluster assignments and ground-truth labels, adjusted for chance agreement.

  7. What distinguishes a self-organizing map (SOM) from k-means clustering?

    Answer: SOM preserves topological relationships of data on a low-dimensional grid; k-means does not

    SOMs arrange neurons on a 2D grid and update neighboring neurons during training, preserving the topological structure of the input space in the output map.