MLlib and Machine Learning Flashcards
6 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 MLlib and Machine Learning flashcards as text
Which class in Spark ML is used for hyperparameter tuning via cross-validation?
Answer: CrossValidator
CrossValidator performs k-fold cross-validation across a parameter grid to select the best model hyperparameters.
What does the BinaryClassificationEvaluator use by default to evaluate model performance in Spark ML?
Answer: Area Under ROC Curve (AUC-ROC)
BinaryClassificationEvaluator defaults to Area Under ROC Curve (areaUnderROC) as its evaluation metric.
Which Spark MLlib algorithm performs dimensionality reduction?
Answer: PCA (Principal Component Analysis)
PCA (Principal Component Analysis) reduces the dimensionality of feature vectors by projecting them onto principal components.
How do you save a trained Spark ML model to disk?
Answer: model.write().save('path')
Spark ML models are saved using model.write().save('path'), which persists the model in Parquet format.
What is the purpose of StandardScaler in Spark MLlib?
Answer: Normalizes feature vectors to have zero mean and/or unit standard deviation
StandardScaler standardizes features by removing the mean and scaling to unit variance, improving convergence for gradient-based algorithms.
Which evaluation metric does MulticlassClassificationEvaluator use by default in Spark ML?
Answer: F1 Score
MulticlassClassificationEvaluator defaults to the F1 score as the evaluation metric for multi-class classification.