Data Analysis and Interpretation Flashcards
7 cards from real CAS practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Data Analysis and Interpretation flashcards as text
When fitting a generalized linear model (GLM) to insurance loss data, which link function is most commonly used for modeling claim frequency?
Answer: Log link
The log link is standard for Poisson-distributed claim frequency models because it ensures predicted frequencies remain positive.
A claims analyst observes that residuals from a regression model exhibit a funnel shape when plotted against fitted values. This pattern most likely indicates:
Answer: Heteroscedasticity in the error terms
A funnel-shaped residual plot is the classic diagnostic for heteroscedasticity, where error variance is not constant across fitted values.
In credibility theory, the Bühlmann credibility factor Z is defined as n/(n+k). As the number of observations n increases toward infinity, Z approaches:
Answer: 1
As n → ∞, the credibility factor Z = n/(n+k) approaches 1, meaning full weight is given to observed experience.
Which of the following best describes the use of a Q-Q plot in actuarial data analysis?
Answer: Assessing whether data follow a specified theoretical distribution
A Q-Q plot compares empirical quantiles of data against theoretical quantiles of a reference distribution to assess distributional fit.
An insurer uses principal component analysis (PCA) on 20 rating variables. The first three principal components explain 85% of total variance. The primary benefit of using these three components instead of all 20 is:
Answer: Reducing dimensionality while retaining most variance
PCA's main benefit is dimensionality reduction — capturing most variance with far fewer uncorrelated components, which simplifies modeling.
A dataset of 500 claims has a sample mean of $12,000 and sample standard deviation of $8,000. The 95% confidence interval for the population mean is constructed using the t-distribution rather than the normal because:
Answer: The population standard deviation is unknown
The t-distribution is used when the population standard deviation is unknown and must be estimated from sample data.
In a loss development triangle, the volume-weighted average link ratio for a development period is calculated by:
Answer: Dividing total cumulative losses at the later age by total cumulative losses at the earlier age
The volume-weighted (chain-ladder) link ratio divides the sum of later-age cumulative losses by the sum of earlier-age cumulative losses, giving more weight to larger values.