课题:选对图表,让数据自己开口说话
一、引言:为什么可视化是分析师的必修课?
L101 我们学会了用频数分布和直方图把杂乱数据「分组打包」。但直方图只是可视化工具之一。在真实工作场景中,不同类型的数据、不同的分析目的,需要不同的图表。
一个场景: 你的基金经理要求你「用一张图说明我们组合过去一年的表现」。你会用什么图?饼图?柱状图?折线图?
选错图表 = 数据被误解。CFA 一级对可视化的考查核心就是:什么数据用什么图。
二、核心图表类型与适用场景
2.1 柱状图(Bar Chart)
定义: 用矩形高度表示各类别的数值,类别为离散型数据。
| 变体 | 特点 | 适用场景 |
|---|---|---|
| 垂直柱状图 | 矩形纵向排列 | 对比不同类别的数值大小 |
| 水平柱状图 | 矩形横向排列 | 类别名称较长时更易读 |
| 分组柱状图 | 同组内多个矩形并列 | 跨类别、跨组比较(如不同行业 × 不同年份) |
| 堆叠柱状图 | 同组内矩形叠加 | 展示总量及各部分构成 |
适用数据: 名义数据、有序分类数据 典型场景: 各行业 PE 估值对比、不同基金规模排名、各资产类别配置比例
⚠️ CFA 陷阱: 柱状图用于离散类别,直方图用于连续数据分组。柱状图的柱子之间有间隙,直方图的柱子之间无间隙(连续)。
2.2 折线图(Line Chart)
定义: 用线段连接连续时间点上的数据值,展示趋势变化。
| 特点 | 说明 |
|---|---|
| X 轴 | 通常为时间变量(日期、季度、年份) |
| Y 轴 | 连续数值变量 |
| 核心优势 | 展示趋势、识别拐点、发现周期性规律 |
| 可叠加 | 多条折线可画在同一张图上进行对比 |
典型场景: - 股价走势图(最经典的折线图) - 基金净值曲线对比 - 宏观经济指标(GDP、CPI)时间序列 - 移动平均线叠加(MA5、MA10、MA20)
⚠️ CFA 陷阱: 折线图要求 X 轴变量有自然顺序(通常是时间)。如果用折线图连接无序类别(如不同行业),就是误导!
2.3 散点图(Scatter Plot)
定义: 用二维坐标系中的点表示两个连续变量之间的关系。
| 要素 | 说明 |
|---|---|
| X 轴 | 自变量(解释变量) |
| Y 轴 | 因变量(被解释变量) |
| 每个点 | 一个观测值(Xᵢ, Yᵢ) |
| 视觉信号 | 点的整体走向反映相关性方向和强度 |
读图要点: - 正相关: 点云从左下到右上 - 负相关: 点云从左上到右下 - 无相关: 点云呈圆形/随机分布 - 非线性关系: 点云呈 U 形、倒 U 形等曲线模式
典型场景: - 风险 vs 收益散点图(每只基金一个点) - 股票 Beta vs 预期收益率 - 市盈率 vs 盈利增长率
散点图的升级版:气泡图(Bubble Chart) - 在散点图基础上用气泡大小表示第三个变量 - 例如:X = 风险,Y = 收益,气泡大小 = 基金规模
2.4 饼图(Pie Chart)与环形图(Donut Chart)
定义: 用扇形角度表示各部分占整体的比例。
| 特点 | 说明 |
|---|---|
| 适用数据 | 构成比例,所有部分之和 = 100% |
| 核心优势 | 直观展示「谁占大头」 |
| 最大弱点 | 人眼不擅长精确比较角度大小 |
使用规则(CFA 考点): 1. 类别不宜过多:超过 5-6 个扇区 → 拥挤难读 2. 占比从大到小排列:12 点方向开始顺时针 3. 不宜用于精确比较:≈2% vs ≈3% 的扇区人眼很难区分 4. 不能用于趋势对比:两个饼图并排比较不同时期 → 低效
典型场景: 投资组合资产配置比例、行业市值分布
⚠️ CFA 陷阱: 饼图不适合展示趋势变化或精确数值对比。如果需要精确比较,用柱状图。
2.5 热力图(Heat Map)
定义: 用颜色深浅编码数值大小,适用于矩阵型数据的可视化。
| 要素 | 说明 |
|---|---|
| 行 × 列 | 两个分类维度 |
| 颜色 | 颜色越深(或越亮)= 数值越大 |
| 核心优势 | 快速识别「热点」和异常模式 |
典型场景: - 相关性矩阵热力图(最经典应用) - 不同行业 × 不同月份的收益率矩阵 - 全球市场涨跌热力图
2.6 箱线图(Box Plot / Box-and-Whisker Plot)
定义: 用五个统计量(最小值、Q1、中位数、Q3、最大值)概括数据分布的图形。
最小值 Q1 中位数 Q3 最大值
|———[ |====| ]———|
whisker box whisker
| 要素 | 含义 |
|---|---|
| 箱子底部 | Q1(第 25 百分位) |
| 箱内横线 | 中位数(Q2 / 第 50 百分位) |
| 箱子顶部 | Q3(第 75 百分位) |
| 箱高(IQR) | Q3 − Q1,反映中间 50% 数据的离散程度 |
| 下须线 | 最小值(或 Q1 − 1.5×IQR,超出为异常值) |
| 上须线 | 最大值(或 Q3 + 1.5×IQR,超出为异常值) |
| 异常值 | 须线之外的点,通常用圆点标注 |
核心优势: - 一张图同时展示中心位置、离散程度、偏度、异常值 - 便于多组数据并排比较(如不同行业 PE 分布箱线图)
三、图表选择决策树
问题1:要展示什么?
├── 数据分布 → 直方图、箱线图、频数多边形
├── 数据比较 → 柱状图(离散类别)
├── 趋势变化 → 折线图(时间序列)
├── 变量关系 → 散点图(两变量连续)
├── 构成比例 → 饼图/环形图(整体=100%)
└── 矩阵模式 → 热力图
四、常见可视化误区
| 误区 | 正确做法 |
|---|---|
| 用饼图展示趋势变化 | 用折线图 |
| 用折线图连接无序类别 | 用柱状图 |
| Y 轴不从 0 开始(柱状图) | 柱状图 Y 轴必须从 0 开始,否则视觉比例被扭曲 |
| 3D 效果图表 | 避免!3D 扭曲视觉感知,增加误读风险 |
| 颜色过多、图例混乱 | 颜色不超过 5-6 种,保持简洁 |
| 双 Y 轴滥用 | 谨慎使用,注明刻度,避免视觉误导 |
五、实战案例
场景: 你是某基金的分析师,老板要求你用数据可视化展示以下信息:
- 过去 5 年 A 股、港股、美股三大市场的年化收益率走势
- 当前组合的资产配置比例
- 持仓 20 只股票的风险-收益散点分布
- 各行业板块 PE 估值的分布区间对比
你应该选什么图?
| 需求 | 推荐图表 | 理由 |
|---|---|---|
| 三市场 5 年收益率走势 | 折线图(3 条线叠加) | 时间序列趋势对比 |
| 资产配置比例 | 饼图或环形图 | 展示各部分占整体比例 |
| 20 只股票风险-收益 | 散点图(X=波动率, Y=收益率) | 两连续变量关系探索 |
| 各行业 PE 分布对比 | 并排箱线图 | 多组数据的五数概括 + 异常值 |
六、核心公式与术语速查
| 英文术语 | 中文 | 关键特征 |
|---|---|---|
| Bar Chart | 柱状图 | 离散类别比较 |
| Histogram | 直方图 | 连续数据分组,无间隙 |
| Line Chart | 折线图 | 时间序列趋势 |
| Scatter Plot | 散点图 | 两变量关系 |
| Pie Chart | 饼图 | 构成比例 |
| Heat Map | 热力图 | 颜色编码矩阵 |
| Box Plot | 箱线图 | 五数概括 + 异常值 |
| Bubble Chart | 气泡图 | 散点图 + 第三维(气泡大小) |
| Frequency Polygon | 频数多边形 | 直方图的折线版,连接各区间中点 |
七、练习题
Q1(概念辨析)
以下哪种图表最适合展示某基金过去 36 个月的月度净值变化趋势? - A. 饼图 - B. 柱状图 - C. 折线图 - D. 箱线图
Q2(场景判断)
分析师想同时展示 5 个不同行业的 PE(市盈率)中位数、离散程度和异常值。最佳选择是: - A. 分组柱状图 - B. 并排箱线图 - C. 饼图 - D. 散点图
Q3(陷阱识别)
以下关于数据可视化的说法,哪一项是错误的? - A. 柱状图用于离散类别,直方图用于连续数据分组 - B. 饼图的扇区数量不宜超过 5-6 个 - C. 柱状图的 Y 轴可以不从 0 开始,只要标注清楚即可 - D. 散点图可以用来发现两个变量间的非线性关系
Q4(术语匹配)
「气泡图」相比普通散点图,额外的维度是通过什么表示的? - A. X 轴刻度 - B. Y 轴刻度 - C. 点的颜色 - D. 点的大小
八、答案与解析
| 题号 | 答案 | 解析 |
|---|---|---|
| Q1 | C | 月度净值变化 = 时间序列趋势,折线图是最佳选择。柱状图可行但不是最优;饼图严重不合适。 |
| Q2 | B | 箱线图一张图就能同时展示中位数、四分位距(离散程度)、异常值。5 个行业并排箱线图 → 一目了然。 |
| Q3 | C | 柱状图 Y 轴必须从 0 开始,否则矩形高度比例会被视觉扭曲。这是 CFA 常考的雷区。 |
| Q4 | D | 气泡大小表示第三个连续变量,是散点图向三维的扩展。 |
九、记住三句话
| # | 金句 |
|---|---|
| 1 | 柱状图看类别,折线图看趋势,散点图看关系,饼图看比例——选错图表,数据白费。 |
| 2 | 柱状图的 Y 轴必须从 0 开始,直方图的柱子之间没有间隙——两个经典 CFA 陷阱。 |
| 3 | 箱线图是「数据分布的 X 光片」:一张图同时展示中心、离散、偏态和异常值。 |
下一课 L103:数据整理与清洗——如何把脏数据变成可分析的金数据。
Topic: Choosing the Right Chart to Let Data Speak for Itself
1. Introduction: Why Visualization is an Essential Analyst Skill
In L101, we learned how to use frequency distributions and histograms to "bundle" messy data. But histograms are just one tool in the visualization toolkit. In real-world work, different data types and analytical objectives demand different charts.
A scenario: Your portfolio manager asks you to "use one chart to illustrate our portfolio's performance over the past year." What chart would you choose? Pie chart? Bar chart? Line chart?
Choosing the wrong chart = data gets misinterpreted. The CFA Level 1 visualization focus is exactly this: what chart for what data.
2. Core Chart Types and Their Use Cases
2.1 Bar Chart
Definition: Uses rectangular heights to represent values across discrete categories.
| Variant | Features | Use Case |
|---|---|---|
| Vertical Bar Chart | Bars arranged vertically | Comparing values across categories |
| Horizontal Bar Chart | Bars arranged horizontally | Better readability for long category names |
| Grouped Bar Chart | Multiple bars side-by-side per group | Cross-category, cross-group comparison (e.g., sectors × years) |
| Stacked Bar Chart | Bars stacked within each group | Showing total and its component breakdown |
Suitable data: Nominal data, ordinal categorical data Typical scenarios: Sector PE valuation comparison, fund size rankings, asset class allocation
⚠️ CFA Trap: Bar charts are for discrete categories; histograms are for continuous data grouped into bins. Bar chart bars have gaps between them; histogram bars have no gaps (continuous).
2.2 Line Chart
Definition: Connects data points across a continuous sequence (usually time) with line segments to show trends.
| Feature | Description |
|---|---|
| X-axis | Usually a time variable (dates, quarters, years) |
| Y-axis | Continuous numeric variable |
| Core strength | Revealing trends, inflection points, and cyclical patterns |
| Overlay capability | Multiple lines can be plotted on the same chart for comparison |
Typical scenarios: - Stock price charts (the classic line chart) - Fund NAV curve comparison - Macroeconomic indicators (GDP, CPI) time series - Moving average overlays (MA5, MA10, MA20)
⚠️ CFA Trap: Line charts require the X-axis variable to have a natural order (usually time). Using a line chart to connect unordered categories (e.g., different industries) is misleading!
2.3 Scatter Plot
Definition: Uses points in a two-dimensional coordinate system to represent the relationship between two continuous variables.
| Element | Description |
|---|---|
| X-axis | Independent variable (explanatory variable) |
| Y-axis | Dependent variable (response variable) |
| Each point | One observation (Xᵢ, Yᵢ) |
| Visual signal | The overall pattern reflects the direction and strength of correlation |
Reading guide: - Positive correlation: Point cloud runs from bottom-left to top-right - Negative correlation: Point cloud runs from top-left to bottom-right - No correlation: Point cloud appears circular / randomly distributed - Non-linear relationship: Point cloud shows U-shaped, inverted U-shaped, or other curved patterns
Typical scenarios: - Risk vs. return scatter plot (each fund = one point) - Stock Beta vs. expected return - P/E ratio vs. earnings growth rate
Upgrade: Bubble Chart - Adds bubble size as a third variable on top of the scatter plot - Example: X = risk, Y = return, bubble size = fund AUM
2.4 Pie Chart & Donut Chart
Definition: Uses sector angles to represent each part's proportion of the whole.
| Feature | Description |
|---|---|
| Suitable data | Composition proportions; all parts must sum to 100% |
| Core strength | Intuitive display of "who dominates" |
| Key weakness | Human eyes are poor at precisely comparing angular sizes |
Usage rules (CFA exam points): 1. Limit categories: More than 5–6 sectors → cluttered and hard to read 2. Sort by size descending: Start from 12 o'clock, proceed clockwise 3. Not for precise comparison: ≈2% vs. ≈3% sectors are nearly indistinguishable by eye 4. Not for trend comparison: Two pie charts side-by-side to compare time periods → inefficient
Typical scenarios: Portfolio asset allocation, sector market cap distribution
⚠️ CFA Trap: Pie charts are unsuitable for showing trends or precise numerical comparisons. When precision is needed, use a bar chart.
2.5 Heat Map
Definition: Encodes numerical values using color intensity; ideal for matrix-format data visualization.
| Element | Description |
|---|---|
| Rows × Columns | Two categorical dimensions |
| Color | Darker (or brighter) color = higher value |
| Core strength | Rapid identification of "hot spots" and anomaly patterns |
Typical scenarios: - Correlation matrix heat map (the classic application) - Industry × month return matrix - Global market performance heat map
2.6 Box Plot (Box-and-Whisker Plot)
Definition: Summarizes data distribution using five statistics (minimum, Q1, median, Q3, maximum).
Minimum Q1 Median Q3 Maximum
|———[ |====| ]———|
whisker box whisker
| Element | Meaning |
|---|---|
| Bottom of box | Q1 (25th percentile) |
| Line inside box | Median (Q2 / 50th percentile) |
| Top of box | Q3 (75th percentile) |
| Box height (IQR) | Q3 − Q1, reflects dispersion of the middle 50% |
| Lower whisker | Minimum (or Q1 − 1.5×IQR; beyond this = outlier) |
| Upper whisker | Maximum (or Q3 + 1.5×IQR; beyond this = outlier) |
| Outliers | Points beyond whiskers, typically marked with dots |
Core strengths: - A single chart simultaneously shows central tendency, dispersion, skewness, and outliers - Excellent for side-by-side multi-group comparison (e.g., box plots of PE ratios across sectors)
3. Chart Selection Decision Tree
Question 1: What do you want to show?
├── Data distribution → Histogram, Box Plot, Frequency Polygon
├── Data comparison → Bar Chart (discrete categories)
├── Trend over time → Line Chart (time series)
├── Variable relationship → Scatter Plot (two continuous variables)
├── Composition / proportion → Pie Chart / Donut Chart (whole = 100%)
└── Matrix pattern → Heat Map
4. Common Visualization Mistakes
| Mistake | Correct Approach |
|---|---|
| Using a pie chart to show trends | Use a line chart |
| Using a line chart to connect unordered categories | Use a bar chart |
| Y-axis not starting at zero (bar chart) | Bar chart Y-axis MUST start at zero, or visual proportions are distorted |
| 3D-effect charts | Avoid! 3D distorts visual perception and increases misreading risk |
| Too many colors, messy legend | Limit colors to 5–6, keep it clean |
| Dual Y-axis abuse | Use sparingly; label scales clearly; avoid visual deception |
5. Real-World Case Study
Scenario: You are an analyst at a fund. Your boss asks you to visualize the following:
- Annualized return trends of A-shares, H-shares, and US equities over the past 5 years
- Current portfolio asset allocation breakdown
- Risk-return scatter distribution of 20 portfolio holdings
- PE valuation range comparison across industry sectors
What chart should you choose?
| Requirement | Recommended Chart | Rationale |
|---|---|---|
| 3-market 5-year return trends | Line chart (3 lines overlaid) | Time series trend comparison |
| Asset allocation breakdown | Pie chart or donut chart | Show each part's proportion of the whole |
| 20 stocks risk-return | Scatter plot (X = volatility, Y = return) | Exploring relationship between two continuous variables |
| Sector PE distribution comparison | Side-by-side box plots | Five-number summary + outliers across groups |
6. Key Terminology Quick Reference
| Term | Key Characteristic |
|---|---|
| Bar Chart | Discrete category comparison |
| Histogram | Continuous data binned, no gaps between bars |
| Line Chart | Time series trend |
| Scatter Plot | Two-variable relationship |
| Pie Chart | Composition proportion |
| Heat Map | Color-coded matrix |
| Box Plot | Five-number summary + outliers |
| Bubble Chart | Scatter plot + third dimension (bubble size) |
| Frequency Polygon | Line-chart version of histogram, connecting bin midpoints |
7. Practice Questions
Q1 (Concept)
Which chart is most appropriate for showing the monthly NAV change of a fund over the past 36 months? - A. Pie chart - B. Bar chart - C. Line chart - D. Box plot
Q2 (Scenario)
An analyst wants to simultaneously display the median P/E, dispersion, and outliers for 5 different industry sectors. The best choice is: - A. Grouped bar chart - B. Side-by-side box plots - C. Pie chart - D. Scatter plot
Q3 (Trap Identification)
Which of the following statements about data visualization is incorrect? - A. Bar charts are for discrete categories; histograms are for continuous data grouped into bins - B. A pie chart should not have more than 5–6 sectors - C. A bar chart's Y-axis does not need to start at zero, as long as the scale is clearly labeled - D. Scatter plots can reveal non-linear relationships between two variables
Q4 (Terminology)
Compared to a regular scatter plot, what represents the additional dimension in a "bubble chart"? - A. X-axis scale - B. Y-axis scale - C. Point color - D. Point size
8. Answers & Explanations
| Q# | Answer | Explanation |
|---|---|---|
| Q1 | C | Monthly NAV change = time series trend. A line chart is optimal. A bar chart is possible but suboptimal; a pie chart is completely inappropriate. |
| Q2 | B | A box plot simultaneously displays median, IQR (dispersion), and outliers in one chart. Five sectors side-by-side → clear at a glance. |
| Q3 | C | A bar chart's Y-axis MUST start at zero; otherwise, bar height proportions are visually distorted. This is a classic CFA exam trap. |
| Q4 | D | Bubble size represents the third continuous variable, extending the scatter plot into a third dimension. |
9. Three Takeaways
| # | Key Insight |
|---|---|
| 1 | Bar chart for categories, line chart for trends, scatter plot for relationships, pie chart for proportions — wrong chart, wasted data. |
| 2 | Bar chart Y-axis must start at zero; histogram bars have no gaps between them — two classic CFA traps. |
| 3 | A box plot is an "X-ray of data distribution": one chart reveals center, spread, skewness, and outliers simultaneously. |
Next up L103: Data Cleaning & Wrangling — turning dirty data into analysis-ready gold.