← All MS-DS Master of Data science Flashcard Decks

Unsupervised Machine Learning Models Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Unsupervised Machine Learning Models flashcards as text
  1. What does the Hopkins statistic measure in cluster analysis?

    Answer: The spatial randomness of data — whether data is clusterable at all

    The Hopkins statistic tests whether data has a significantly non-random spatial structure, indicating whether it is worth clustering.

  2. Which of the following is a key limitation of Principal Component Analysis (PCA) when applied to image data?

    Answer: PCA assumes linear relationships and cannot capture nonlinear structure in image manifolds

    PCA is a linear method and fails to capture nonlinear manifold structure present in image data, which is why kernel PCA or autoencoders are preferred.

  3. In the context of self-organizing maps (SOMs), what does the neighborhood function determine?

    Answer: How much surrounding neurons are updated when a best-matching unit is found

    The neighborhood function (typically Gaussian) controls how strongly neurons adjacent to the best-matching unit are pulled toward the input during each update.

  4. What is the 'cold start' problem in collaborative filtering?

    Answer: Inability to make recommendations for new users or items with no interaction history

    The cold start problem occurs when the system cannot generate recommendations for new users or items because no historical interaction data exists.

  5. Which property distinguishes agglomerative from divisive hierarchical clustering?

    Answer: Agglomerative starts with each point as its own cluster and merges upward; divisive starts with one cluster and splits

    Agglomerative clustering is bottom-up (merging), while divisive is top-down (splitting); both produce dendrograms but differ in direction.

  6. In Independent Component Analysis (ICA), what assumption distinguishes it from PCA?

    Answer: ICA assumes statistically independent non-Gaussian sources, whereas PCA assumes uncorrelated Gaussian components

    ICA exploits statistical independence and non-Gaussianity to separate sources, going beyond the second-order decorrelation that PCA achieves.

  7. What does a high average silhouette score (close to 1) indicate about a clustering solution?

    Answer: Points are well-matched to their own cluster and poorly matched to neighboring clusters

    A silhouette score near 1 means each point is much closer to its own cluster's centroid than to any other cluster, indicating dense, well-separated clusters.