← All MS-DS Master of Data science Flashcard Decks

Data Visualization and Communication Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Data Visualization and Communication flashcards as text
  1. Which color encoding principle states that hue should be used for nominal data while luminance/saturation should encode ordered data?

    Answer: Bertin's retinal variables principle

    Jacques Bertin established that hue differentiates categories (nominal), while value/lightness encodes order or quantity (ordinal/quantitative).

  2. A dashboard shows a KPI card with a large number and no trend context. What is the primary communication failure?

    Answer: Absence of a reference baseline or comparison

    A single number without a baseline, target, or historical trend gives the viewer no way to judge whether the value is good or bad.

  3. When visualizing a distribution with heavy outliers, which chart type preserves the most information without distortion?

    Answer: Box-and-whisker plot

    A box-and-whisker plot explicitly shows median, IQR, whiskers, and individual outlier points without collapsing the distribution into a single mean.

  4. The 'lie factor' metric introduced by Edward Tufte measures what?

    Answer: The ratio of the effect size shown in the graphic to the effect size in the data

    Tufte defined Lie Factor = (size of effect shown in graphic) / (size of effect in data); a value far from 1.0 indicates visual distortion.

  5. In a connected scatterplot (path chart), what does the path between points encode that a standard scatterplot omits?

    Answer: Temporal sequence or ordering

    The path connects observations in order (usually time), revealing how two variables co-evolved — information a static scatterplot cannot convey.

  6. Which technique is most appropriate for visualizing the joint distribution of two continuous variables when the dataset contains 500,000 points?

    Answer: 2D kernel density estimation (contour plot) or hexbin plot

    At large scale, overplotting makes scatterplots unreadable; hexbin or KDE contour plots aggregate point density into readable cells or iso-lines.

  7. A presenter wants to compare part-to-whole relationships across five product categories simultaneously. Which chart type is most appropriate?

    Answer: 100% stacked bar chart

    A 100% stacked bar chart normalizes each bar to 100%, making it easy to compare proportional compositions across categories.