← All Apache Spark Flashcard Decks

MLlib and Machine Learning Flashcards

6 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 MLlib and Machine Learning flashcards as text
  1. Which class in Spark ML is used for hyperparameter tuning via cross-validation?

    Answer: CrossValidator

    CrossValidator performs k-fold cross-validation across a parameter grid to select the best model hyperparameters.

  2. What does the BinaryClassificationEvaluator use by default to evaluate model performance in Spark ML?

    Answer: Area Under ROC Curve (AUC-ROC)

    BinaryClassificationEvaluator defaults to Area Under ROC Curve (areaUnderROC) as its evaluation metric.

  3. Which Spark MLlib algorithm performs dimensionality reduction?

    Answer: PCA (Principal Component Analysis)

    PCA (Principal Component Analysis) reduces the dimensionality of feature vectors by projecting them onto principal components.

  4. How do you save a trained Spark ML model to disk?

    Answer: model.write().save('path')

    Spark ML models are saved using model.write().save('path'), which persists the model in Parquet format.

  5. What is the purpose of StandardScaler in Spark MLlib?

    Answer: Normalizes feature vectors to have zero mean and/or unit standard deviation

    StandardScaler standardizes features by removing the mean and scaling to unit variance, improving convergence for gradient-based algorithms.

  6. Which evaluation metric does MulticlassClassificationEvaluator use by default in Spark ML?

    Answer: F1 Score

    MulticlassClassificationEvaluator defaults to the F1 score as the evaluation metric for multi-class classification.