Data Analysis and Interpretation Flashcards
7 cards from real CAS practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Data Analysis and Interpretation flashcards as text
An actuary applies a one-way analysis of an insurance rating factor and finds that relativities are not monotone. Before making adjustments, which of the following is the most important consideration?
Answer: Whether the non-monotone pattern reflects genuine risk or is due to sampling variability or sparse data
Non-monotone relativities often result from thin data in certain cells; credibility weighting or smoothing may be appropriate before forcing monotonicity.
The coefficient of variation (CV) of a loss distribution is defined as:
Answer: Standard deviation divided by the mean
The CV = σ/μ measures relative dispersion; it is dimensionless and allows comparison of variability across distributions with different scales.
When using gradient boosting machines (GBM) for loss cost modeling, the learning rate (shrinkage) hyperparameter primarily controls:
Answer: The contribution of each successive tree to the overall prediction
The learning rate scales each tree's contribution, trading off slower learning (smaller rate, more trees needed) for better generalization.
A p-p plot (probability-probability plot) is used to compare an empirical distribution to a theoretical one. Unlike a Q-Q plot, the p-p plot compares:
Answer: Empirical CDFs evaluated at each data point against theoretical CDF values
A p-p plot maps the empirical CDF value at each observation against the theoretical CDF value at the same point; deviations from the diagonal indicate distributional misfit.
In a stochastic claims reserving model using bootstrapping of Pearson residuals, the main purpose is to:
Answer: Generate a distribution of reserve estimates to quantify reserve uncertainty
Bootstrapping the residuals of a chain-ladder model produces many simulated triangles, yielding a distribution of possible ultimate losses that quantifies reserve risk.
The Cramér-von Mises statistic is used in goodness-of-fit testing. Compared to the Kolmogorov-Smirnov statistic, it:
Answer: Measures the integrated squared difference between empirical and theoretical CDFs
The Cramér-von Mises statistic integrates the squared deviation between the empirical and theoretical CDFs over all observed values, giving global rather than maximum deviation.
An insurer builds a classification tree to predict whether a policy will generate a large loss. The tree is grown to full depth on training data (zero misclassification error). The primary problem with deploying this model is:
Answer: The model is overfit and will likely perform poorly on new, unseen data
A fully grown tree memorizes training data noise (overfitting), resulting in high variance and degraded predictive performance on hold-out or future data.