โ† All MS-DS Master of Data science Flashcard Decks

Unsupervised Learning Techniques Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Unsupervised Learning Techniques flashcards as text
  1. What problem does the k-means++ initialization scheme address compared to random initialization?

    Answer: It reduces the chance of poor convergence to local optima by spreading initial centroids

    k-means++ selects each subsequent initial centroid with probability proportional to its squared distance from the nearest existing centroid, leading to better spread and faster convergence.

  2. In spectral clustering, what is the graph Laplacian used for?

    Answer: To embed data into a low-dimensional space where standard clustering is applied

    Spectral clustering computes the eigenvectors of the graph Laplacian to embed the data into a space that reveals cluster structure, then applies k-means on this embedding.

  3. Which of the following is a core limitation of PCA for dimensionality reduction in the context of unsupervised learning?

    Answer: PCA only captures linear relationships and misses nonlinear structure in the data

    PCA finds orthogonal directions of maximum variance using linear projections, so it cannot capture curved manifolds or other nonlinear structure in high-dimensional data.

  4. A data scientist applies k-means to customer transaction data and observes that the algorithm assigns all points to one cluster in the first iteration. What is the most likely cause?

    Answer: All initial centroids were placed in the same region of the feature space

    If all centroids are initialized very close together, all points become nearest to one centroid, causing a degenerate solution; k-means++ initialization mitigates this.

  5. In autoencoders used for anomaly detection, what serves as the anomaly score for a data point?

    Answer: The reconstruction error between input and decoder output

    An autoencoder trained on normal data learns to reconstruct normal patterns well; anomalies have high reconstruction error because the model cannot encode and decode them accurately.

  6. Which statement best describes the difference between hard and soft clustering?

    Answer: Hard clustering assigns each point to exactly one cluster; soft clustering assigns fractional memberships

    In hard clustering (e.g., k-means), each point belongs to exactly one cluster; in soft clustering (e.g., GMM), each point has a probability of belonging to each cluster.

  7. What is the primary difference between agglomerative and divisive hierarchical clustering?

    Answer: Agglomerative starts with each point as its own cluster and merges; divisive starts with one cluster and splits

    Agglomerative (bottom-up) begins with n singleton clusters and iteratively merges the most similar pair, while divisive (top-down) begins with all points in one cluster and recursively splits.