课题:数据对中心的"锚点"——三种均值与它们的江湖地位
一、引言:一个投资分析师的三道考题
假设你面前有三组数据,每组代表一个投资领域的关键指标:
| 场景 | 数据 | 你要回答的问题 |
|---|---|---|
| A | 某板块 10 只股票的市盈率 | 这个板块的"典型"估值是多少? |
| B | 某基金过去 5 年的年化回报率:-20%, +15%, +30%, +10%, +5% | 该基金 5 年的实际复合增长率? |
| C | 你要计算一个涵盖 100 只股票的指数的平均 PE | 大盘 PE 怎么算? |
三个场景,三个答案——分别对应算数平均、几何平均、加权平均。
如果你只用一种"平均"去套所有场景,你可能买到的是"虚假的便宜",或者"夸大的回报"。
本课就是让你彻底搞清楚:集中趋势的每种度量方式,应该在什么时候用、绝不能用在什么时候。
二、核心概念:中心在哪里?
2.1 集中趋势(Central Tendency)的定义
集中趋势是指一组数据"聚拢"在哪个值周围。它是描述性统计的第一步——在任何分析之前,先问:数据的"典型值"是多少?
三个经典指标: - 均值(Mean):算术平均值,受每个数据点影响 - 中位数(Median):排序后的中间值,不受极端值影响 - 众数(Mode):出现频率最高的值,可用于定性数据
📌 CFA 一级考点:知道什么尺度数据能用什么集中趋势指标。
| 测量尺度 | 可用指标 | 最佳选择 |
|---|---|---|
| 名义(Nominal) | 众数 | 众数 |
| 序数(Ordinal) | 中位数、众数 | 中位数 |
| 间隔(Interval) | 均值、中位数、众数 | 均值(无异常值)/ 中位数(有异常值) |
| 比率(Ratio) | 均值、中位数、众数、几何平均 | 视场景而定 |
三、三种均值(Mean)深度解析
3.1 算术平均(Arithmetic Mean)——"求平均"的本能反应
公式:
$$\bar{x} = \frac{x_1 + x_2 + \cdots + x_n}{n} = \frac{\sum_{i=1}^{n} x_i}{n}$$
特征: - 最直观、最常用的均值 - 每个数据点等权重 - 致命弱点:对异常值(Outliers)极度敏感
实际案例: 你分析 10 位基金经理的业绩。9 位年化回报 5%-12%,第 10 位是 -60%(爆仓)。算数平均可能被拉低到 2%,但 90% 的经理表现其实不错。
适用场景: - 没有严重偏斜的正态分布数据 - 需要用到方差和标准差的计算(因为方差基于算术平均) - 预测下一期的期望值
不适用的场景: - 有极端异常值(被扭曲)→ 用中位数 - 计算一段时间内的增长率 → 用几何平均
3.2 几何平均(Geometric Mean)——"复利的朋友"
公式:
$$G = \sqrt[n]{x_1 \times x_2 \times \cdots \times x_n} = (x_1 \times x_2 \times \cdots \times x_n)^{1/n}$$
对于回报率数据,每个 $x_i = (1 + r_i)$:
$$G = \sqrt[n]{(1+r_1)(1+r_2)\cdots(1+r_n)} - 1$$
特征: - 反映复合增长的真实速度 - 永远 ≤ 算数平均(除非所有数据相等) - 对负值或零值敏感(出现负值或零,几何平均可能无意义)
实际案例——为什么基金宣传"算数平均"而不用"几何平均":
某基金的三年回报率: - 第 1 年:+50% - 第 2 年:-30% - 第 3 年:+50%
算数平均: (50% + (-30%) + 50%) / 3 = 23.3% → 看起来很诱人 😍
几何平均: ³√(1.5 × 0.7 × 1.5) - 1 = ³√1.575 - 1 ≈ 16.3% → 这才是真实增速 😑
验证: 投资 100 元: - 第 1 年末:100 × 1.5 = 150 - 第 2 年末:150 × 0.7 = 105 - 第 3 年末:105 × 1.5 = 157.5
如果用算术平均 23.3% 复利:100 × 1.233³ = 187.5(高估了 30 元!)
🔑 CFA 必考:时间序列的回报率,用几何平均衡量复合增长率;横截面的数据,用算数平均衡量集中位置。
实战陷阱识别: - 基金宣传材料中"平均年化回报"如果是算术平均 → 可能被夸大 - 看业绩时找 CAGR(Compound Annual Growth Rate)→ 这才是几何平均 - 波动越大,算术平均和几何平均的差距越大
3.3 调和平均(Harmonic Mean)——"定投的法宝"
公式:
$$H = \frac{n}{\frac{1}{x_1} + \frac{1}{x_2} + \cdots + \frac{1}{x_n}} = \frac{n}{\sum_{i=1}^{n} \frac{1}{x_i}}$$
特征: - 永远 ≤ 几何平均 ≤ 算数平均 - 对极小值非常敏感 - 主要用于"固定金额购买"的场景
实际案例——定投中的调和平均:
你每月定投 1000 元购买某 ETF: - 第一个月:价格 20 元/份 → 买入 50 份 - 第二个月:价格 25 元/份 → 买入 40 份 - 第三个月:价格 10 元/份 → 买入 100 份
总投入:3000 元,总份额:190 份
调和平均价格: 3 / (1/20 + 1/25 + 1/10) = 3 / (0.05 + 0.04 + 0.10) = 3 / 0.19 = 15.79 元
算数平均价格: (20 + 25 + 10) / 3 = 18.33 元
实际每股成本 = 3000 / 190 = 15.79 元 ← 正好等于调和平均!
🔑 调和平均在你的实际成本计算中自然出现——当你用固定金额而非固定份额购买时。
CFA 一级调和平均的主要考点: 1. 计算"每股平均成本"时用调和平均 2. 调和平均 ≤ 几何平均 ≤ 算数平均(除非所有值相等) 3. 调和平均对极端小值非常敏感
3.4 加权平均(Weighted Mean)——"大盘 PE 的真相"
公式:
$$\bar{x}_w = \frac{w_1 x_1 + w_2 x_2 + \cdots + w_n x_n}{w_1 + w_2 + \cdots + w_n} = \frac{\sum w_i x_i}{\sum w_i}$$
特征: - 每个数据点赋予不同的重要性(权重) - 最简单的权重 = 相等的权重 → 回归算数平均
实际案例——为什么大盘 PE 不能直接用算数平均?
上证 50 指数中: - 贵州茅台:PE = 30,市值 2 万亿 - 某小盘股:PE = 50,市值 200 亿
如果简单算术平均:(30 + 50) / 2 = 40 × ——误导!
加权平均(按市值权重):(30 × 20000 + 50 × 200) / (20000 + 200) ≈ 30.20 ✓
正确的指数 PE 就是按市值加权的!!
📌 CFA 考点:指数计算、组合收益、WACC 中的权重概念,全部基于加权平均。
3.5 三种均值关系总结
| 均值类型 | 公式关键 | 关系 | 典型应用 |
|---|---|---|---|
| 算术平均 | $\sum x_i / n$ | 最大(通常) | 横截面数据、期望值 |
| 几何平均 | $(\prod x_i)^{1/n}$ | 中间 | 时间序列回报率、CAGR |
| 调和平均 | $n / \sum (1/x_i)$ | 最小(通常) | 定投成本、等额购买 |
三者关系:调和平均 ≤ 几何平均 ≤ 算数平均
等式成立的条件:所有数据完全相等。
四、中位数(Median)—— 不随波逐流的稳健派
4.1 定义与计算
中位数:将数据从小到大排序后,位于正中间的值。
- n 为奇数: 第 (n+1)/2 个值
- n 为偶数: 中间两个值的算术平均
4.2 中位数 vs 均值——什么时候该"背弃"均值
经典案例:深圳平均工资 vs 中位数工资
假设一个 10 人小公司的月薪(万元):
1.2, 1.3, 1.5, 1.5, 1.6, 1.7, 1.8, 2.0, 2.2, 35.0(CEO)
- 算术平均: 49.8 / 10 = 4.98 万
- 中位数: 第 5 和第 6 的平均 = (1.6 + 1.7) / 2 = 1.65 万
哪个更能代表"普通员工"的工资?显然是中位数 1.65 万。
当数据呈正偏态(右偏)时:均值 > 中位数 > 众数 当数据呈负偏态(左偏)时:均值 < 中位数 < 众数
4.3 中位数的优势
- ✅ 不受极端值影响(稳健性)
- ✅ 适合序数数据(如晨星评级)
- ✅ 适合偏态分布
- ❌ 不便于数学运算(无法用中位数直接推导方差)
- ❌ 小样本时可能不稳定
五、众数(Mode)——"最多人的选择"
5.1 定义
众数:数据中出现频率最高的值。
- 可以有 0 个众数(所有值出现次数相同)
- 可以有 1 个众数(单峰 Unimodal)
- 可以有 多个众数(双峰 Bimodal / 多峰 Multimodal)
5.2 众数的独特地位
唯一可以用于名义数据的集中趋势指标!
例子: - 投资者最偏好的 ETF 类别 → 众数是 "Diversified Emerging Mkts" - 最常被推荐的基金评级 → 众数是 4 星
5.3 众数的局限
- 对数据的微小变化可能非常敏感(bimodal 分布中,多一个数据点可能让众数跳变)
- 不能反映数据的整体分布
- 连续数据中,众数的意义有限(除非分组后看 modal class)
六、分位数(Quantiles)—— 比中位数更细的视角
6.1 常用分位数
| 分位数 | 含义 | 位置公式 |
|---|---|---|
| 中位数(Q2) | 50% 的数据小于此值 | L = (n+1) × 0.50 |
| 第一四分位数(Q1) | 25% 的数据小于此值 | L = (n+1) × 0.25 |
| 第三四分位数(Q3) | 75% 的数据小于此值 | L = (n+1) × 0.75 |
| 第 k 百分位数 | k% 的数据小于此值 | L = (n+1) × (k/100) |
位置公式: $L_y = (n+1) \times \frac{y}{100}$
- 如果 L 是整数 → 直接取第 L 个值
- 如果 L 不是整数 → 线性插值
6.2 实际案例——用分位数判断基金表现
某基金经理告诉你:"我的基金回报率处于前 25%。"
他怎么证明?——他的基金回报 > Q3(75% 分位数),即超过 75% 的同行。
- Q1 → 前 25% 最差的
- Q2(中位数)→ 刚好在中间
- Q3 → 前 25% 最好的 → 这就是"前 25%"的含义
七、综合对比:什么时候用什么?
| 场景 | 推荐指标 | 理由 |
|---|---|---|
| 投资组合上一季度的平均日回报 | 算术平均 | 横截面,期望值 |
| 过去 5 年你的基金年化回报 | 几何平均 | 时间序列,复合增长 |
| 每月定额买入 ETF 的平均成本 | 调和平均 | 定投场景 |
| 沪深 300 指数的 PE | 加权平均 | 按市值加权 |
| 深圳的"典型"收入水平 | 中位数 | 偏态分布,有异常值 |
| 最受投资者欢迎的板块 | 众数 | 分类数据 |
| 你的基金在同行的排位 | 分位数 | 相对位置 |
八、测试题
选择题
Q1:一只基金过去 4 年的回报率分别为 +10%、-5%、+20%、+15%。该基金的年化复合增长率(CAGR)应使用哪个均值? - A. 算术平均 - B. 几何平均 - C. 调和平均 - D. 加权平均
Q2:以下哪类数据可以用中位数但不能用均值作为集中趋势指标? - A. 正态分布的收益率 - B. 严重右偏的收入数据 - C. 标准正态分布的 z-score - D. 某 ETF 每日收盘价
Q3:调和平均、几何平均、算术平均三者大小关系是? - A. 算术 ≥ 几何 ≥ 调和(除非所有数据相等) - B. 调和 ≥ 几何 ≥ 算术(除非所有数据相等) - C. 几何 ≥ 算术 ≥ 调和(除非所有数据相等) - D. 三者没有固定关系
Q4:名义数据(如投资风格标签:成长/价值/平衡)唯一可以用哪个集中趋势指标? - A. 算术平均 - B. 中位数 - C. 众数 - D. 几何平均
Q5:某分析师要计算沪深 300 指数成分股的平均市盈率。以下哪种方法最正确? - A. 将 300 只股票的 PE 简单算数平均 - B. 按市值加权平均 - C. 按流通股本加权平均 - D. 用中位数 PE
Q6:一个投资组合中,如果 Q1 回报 = 2%,Q3 回报 = 12%,则下列哪项正确? - A. 25% 的时期回报超过 12% - B. 50% 的时期回报在 2% 到 12% 之间 - C. 中位数回报一定等于 7% - D. 最小值一定小于 2%
Q7:在正偏态(右偏)分布中,均值、中位数、众数的关系是? - A. 均值 < 中位数 < 众数 - B. 均值 > 中位数 > 众数 - C. 均值 = 中位数 = 众数 - D. 中位数 > 均值 > 众数
答案
Q1:B — 时间序列的回报率衡量复合增长,必须用几何平均。CAGR = [(1.10)(0.95)(1.20)(1.15)]^(1/4) - 1 ≈ 9.47%。算术平均= (10% - 5% + 20% + 15%) / 4 = 10%,高于真实复合增长率。
Q2:B — 收入数据通常严重右偏(少数极高收入拉高均值),中位数更能代表"典型"收入水平。正态分布的数据均值和中位数相等。
Q3:A — 算术平均 ≥ 几何平均 ≥ 调和平均,等式仅在所有数据相等时成立。这是 CFA 一级的经典考点。
Q4:C — 名义数据(纯分类标签)没有数值、没有排序,只能用众数统计哪类最常出现。你不能对"成长"和"价值"求平均。
Q5:B — 指数 PE 按市值加权是正确的行业标准。简单算术平均会把小盘股和大盘股同等对待,扭曲指数的真实估值水平。
Q6:B — Q1 和 Q3 之间包含了中间 50% 的数据(从第 25 百分位到第 75 百分位),所以 50% 的时期回报在 2%-12% 之间。A 说的是 Q3 以上的 25%,但表述相反了。C 中位数不一定等于 Q1 和 Q3 的平均。
Q7:B — 正偏态(右偏)分布中,右边有一条长尾,拉高了均值。众数在峰值(最高),中位数在中间,均值在最右边:均值 > 中位数 > 众数。
九、备考要点
| 优先级 | 考点 | 出现概率 | 关键记忆 |
|---|---|---|---|
| ⭐⭐⭐ | 算术/几何/调和平均的适用场景 | 极高 | 横截面→算术,时间序列回报→几何,定投成本→调和 |
| ⭐⭐⭐ | 正偏态中均值>中位数>众数 | 极高 | 想到收入分布——长尾在右边 |
| ⭐⭐⭐ | 中位数 vs 均值的选择 | 高 | 有异常值/偏态→中位数 |
| ⭐⭐⭐ | 加权平均的权重概念 | 高 | 指数计算、组合收益 |
| ⭐⭐ | 分位数计算公式 L = (n+1)×(y/100) | 中 | 注意整数/非整数的不同处理 |
| ⭐⭐ | 名义数据只能用众数 | 中 | 不能做加减乘除的数据 |
| ⭐⭐ | 调和≤几何≤算术的关系 | 中 | 除非所有值相等 |
| ⭐ | 众数的多峰性 | 低 | Unimodal/Bimodal |
十、快速记忆卡
📌 一句话总结:
算术平均 → 横截面期望值(最常用)
几何平均 → 复合增长率 / CAGR(时间序列必用)
调和平均 → 定投成本(等额购买)
中位数 → 偏态数据的"典型值"
众数 → 分类数据的"最流行"
加权平均 → 指数/组合的市值加权
分位数 → 相对排名(你打败了多少人)
CFA 一级 · 量化方法 · 模块 2.3 开始 | L105 集中趋势 · 2026-07-11
Topic: The "Anchors" of Data — Three Types of Means and Their Rightful Place
I. Introduction: Three Questions for an Investment Analyst
Imagine you're faced with three datasets, each representing a key investment metric:
| Scenario | Data | The Question |
|---|---|---|
| A | P/E ratios of 10 stocks in a sector | What is the "typical" valuation of this sector? |
| B | A fund's annual returns over 5 years: -20%, +15%, +30%, +10%, +5% | What is the fund's actual compound growth rate? |
| C | You need to calculate the average P/E of a 100-stock index | How should index-level P/E be computed? |
Three scenarios, three answers — corresponding to arithmetic mean, geometric mean, and weighted mean.
If you use only one type of "average" for every situation, you may end up buying "fake cheapness" or "exaggerated returns."
The goal of this lesson: thoroughly understand which measure of central tendency to use when, and when absolutely not to use it.
II. Core Concept: Where Is the Center?
2.1 Definition of Central Tendency
Central tendency refers to the value around which a dataset clusters. It is the first step in descriptive statistics — before any analysis, ask: what is the "typical value" of the data?
Three classic measures: - Mean — the arithmetic average, influenced by every data point - Median — the middle value after sorting, unaffected by extreme values - Mode — the most frequently occurring value, usable for qualitative data
📌 CFA Level I Focus: know which measure of central tendency can be used for each measurement scale.
| Measurement Scale | Available Measures | Best Choice |
|---|---|---|
| Nominal | Mode | Mode |
| Ordinal | Median, Mode | Median |
| Interval | Mean, Median, Mode | Mean (no outliers) / Median (with outliers) |
| Ratio | Mean, Median, Mode, Geometric Mean | Depends on context |
III. Deep Dive: The Three Types of Means
3.1 Arithmetic Mean — The "Default" Average
Formula:
$$\bar{x} = \frac{x_1 + x_2 + \cdots + x_n}{n} = \frac{\sum_{i=1}^{n} x_i}{n}$$
Characteristics: - The most intuitive and commonly used mean - Every data point carries equal weight - Fatal flaw: extremely sensitive to outliers
Real-World Example: You analyze the performance of 10 fund managers. Nine have annual returns of 5%-12%, but the tenth blew up at -60%. The arithmetic mean may be dragged down to 2%, yet 90% of managers performed decently.
When to Use: - Normally distributed data without severe skewness - When computing variance and standard deviation (variance is based on arithmetic mean) - Forecasting expected values for the next period
When NOT to Use: - Data with extreme outliers (the mean gets distorted) → Use median - Measuring growth rates over time → Use geometric mean
3.2 Geometric Mean — "The Friend of Compound Returns"
Formula:
$$G = \sqrt[n]{x_1 \times x_2 \times \cdots \times x_n} = (x_1 \times x_2 \times \cdots \times x_n)^{1/n}$$
For return data, each $x_i = (1 + r_i)$:
$$G = \sqrt[n]{(1+r_1)(1+r_2)\cdots(1+r_n)} - 1$$
Characteristics: - Reflects the true speed of compound growth - Always ≤ Arithmetic Mean (unless all values are equal) - Sensitive to negative or zero values (geometric mean may become meaningless)
Real-World Example — Why funds advertise "arithmetic mean" instead of "geometric mean":
A fund's three-year returns: - Year 1: +50% - Year 2: -30% - Year 3: +50%
Arithmetic Mean: (50% + (-30%) + 50%) / 3 = 23.3% → Looks very attractive 😍
Geometric Mean: ³√(1.5 × 0.7 × 1.5) - 1 = ³√1.575 - 1 ≈ 16.3% → This is the real growth rate 😑
Verification: Invest $100: - End of Year 1: 100 × 1.5 = 150 - End of Year 2: 150 × 0.7 = 105 - End of Year 3: 105 × 1.5 = 157.5
Using arithmetic mean 23.3% compounded: 100 × 1.233³ = 187.5 (overstated by $30!)
🔑 CFA Must-Know: For time-series returns, use geometric mean for compound growth rates. For cross-sectional data, use arithmetic mean for central location.
Real-World Pitfalls: - Fund marketing materials showing "average annual return" as arithmetic mean → likely inflated - When reviewing performance, look for CAGR (Compound Annual Growth Rate) → that's geometric mean - The greater the volatility, the larger the gap between arithmetic and geometric mean
3.3 Harmonic Mean — "The DCA (Dollar-Cost Averaging) Ally"
Formula:
$$H = \frac{n}{\frac{1}{x_1} + \frac{1}{x_2} + \cdots + \frac{1}{x_n}} = \frac{n}{\sum_{i=1}^{n} \frac{1}{x_i}}$$
Characteristics: - Always ≤ Geometric Mean ≤ Arithmetic Mean - Highly sensitive to very small values - Primarily used in "fixed-dollar-amount purchasing" scenarios
Real-World Example — Harmonic Mean in Dollar-Cost Averaging:
You invest $1,000 per month into an ETF: - Month 1: price $20/share → buy 50 shares - Month 2: price $25/share → buy 40 shares - Month 3: price $10/share → buy 100 shares
Total investment: $3,000; Total shares: 190
Harmonic Mean Price: 3 / (1/20 + 1/25 + 1/10) = 3 / (0.05 + 0.04 + 0.10) = 3 / 0.19 = $15.79
Arithmetic Mean Price: (20 + 25 + 10) / 3 = $18.33
Actual cost per share = 3000 / 190 = $15.79 ← Exactly the harmonic mean!
🔑 The harmonic mean emerges naturally when you buy with a fixed dollar amount rather than a fixed number of shares.
CFA Level I Harmonic Mean Focus Areas: 1. Use harmonic mean for calculating "average cost per share" 2. Harmonic ≤ Geometric ≤ Arithmetic (unless all values equal) 3. Harmonic mean is highly sensitive to extremely small values
3.4 Weighted Mean — "The Truth About Index-Level P/E"
Formula:
$$\bar{x}_w = \frac{w_1 x_1 + w_2 x_2 + \cdots + w_n x_n}{w_1 + w_2 + \cdots + w_n} = \frac{\sum w_i x_i}{\sum w_i}$$
Characteristics: - Each data point is assigned a different level of importance (weight) - Simplest case: equal weights → reduces to arithmetic mean
Real-World Example — Why index P/E cannot use simple arithmetic mean:
In the S&P 500: - Apple: P/E = 30, market cap $3 trillion - A small-cap stock: P/E = 50, market cap $20 billion
Simple arithmetic mean: (30 + 50) / 2 = 40 ✗ — Misleading!
Weighted mean (by market cap): (30 × 3000 + 50 × 20) / (3000 + 20) ≈ 30.13 ✓
Correct index P/E is always market-cap weighted!!
📌 CFA Focus: Index calculations, portfolio returns, WACC — all rely on the concept of weighted averages.
3.5 Summary: Three Means Relationship
| Mean Type | Formula Key | Relationship | Typical Application |
|---|---|---|---|
| Arithmetic | $\sum x_i / n$ | Largest (usually) | Cross-sectional data, expected value |
| Geometric | $(\prod x_i)^{1/n}$ | Middle | Time-series returns, CAGR |
| Harmonic | $n / \sum (1/x_i)$ | Smallest (usually) | DCA cost, equal-dollar purchases |
Relationship: Harmonic ≤ Geometric ≤ Arithmetic
Equality holds only when: all values are identical.
IV. Median — The Robust Measure That Ignores Outliers
4.1 Definition and Calculation
Median: the middle value after sorting data from smallest to largest.
- n is odd: the (n+1)/2 th value
- n is even: arithmetic mean of the two middle values
4.2 Median vs. Mean — When to "Betray" the Mean
Classic Example: Average income vs. median income
Consider the monthly salaries (in $ thousands) of a 10-person company:
1.2, 1.3, 1.5, 1.5, 1.6, 1.7, 1.8, 2.0, 2.2, 35.0 (CEO)
- Arithmetic Mean: 49.8 / 10 = $49,800
- Median: average of 5th and 6th = (1.6 + 1.7) / 2 = $16,500
Which better represents a "typical employee's" salary? Clearly the median at $16,500.
When data is positively skewed (right-skewed): Mean > Median > Mode When data is negatively skewed (left-skewed): Mean < Median < Mode
4.3 Advantages of the Median
- ✅ Unaffected by extreme values (robustness)
- ✅ Suitable for ordinal data (e.g., Morningstar ratings)
- ✅ Suitable for skewed distributions
- ❌ Difficult to use in mathematical derivations (cannot directly derive variance from median)
- ❌ May be unstable with small samples
V. Mode — "The Most Popular Choice"
5.1 Definition
Mode: the value that appears most frequently in a dataset.
- Can have 0 modes (all values appear equally often)
- Can have 1 mode (Unimodal)
- Can have multiple modes (Bimodal / Multimodal)
5.2 The Unique Position of the Mode
The only measure of central tendency that can be used with nominal data!
Examples: - The most popular ETF category among investors → Mode is "Diversified Emerging Mkts" - The most commonly assigned fund rating → Mode is 4 stars
5.3 Limitations of the Mode
- Can be highly sensitive to small data changes (in a bimodal distribution, one extra data point may shift the mode)
- Does not reflect the overall distribution
- Limited meaning for continuous data (use modal class after grouping)
VI. Quantiles — A Finer Lens Than the Median
6.1 Common Quantiles
| Quantile | Meaning | Position Formula |
|---|---|---|
| Median (Q2) | 50% of data fall below this value | L = (n+1) × 0.50 |
| First Quartile (Q1) | 25% of data fall below this value | L = (n+1) × 0.25 |
| Third Quartile (Q3) | 75% of data fall below this value | L = (n+1) × 0.75 |
| k-th Percentile | k% of data fall below this value | L = (n+1) × (k/100) |
Position formula: $L_y = (n+1) \times \frac{y}{100}$
- If L is an integer → take the L-th value directly
- If L is not an integer → use linear interpolation
6.2 Real-World Example — Using Quantiles to Assess Fund Performance
A fund manager tells you: "My fund's return is in the top quartile."
How to verify? — His fund return > Q3 (75th percentile), meaning it beat 75% of peers.
- Q1 → worst 25%
- Q2 (Median) → right in the middle
- Q3 → best 25% → This is what "top quartile" means
VII. Comprehensive Comparison: When to Use What?
| Scenario | Recommended Measure | Reason |
|---|---|---|
| Average daily return of a portfolio last quarter | Arithmetic Mean | Cross-sectional, expected value |
| Your fund's annualized return over the past 5 years | Geometric Mean | Time series, compound growth |
| Average cost when buying ETF with fixed monthly amount | Harmonic Mean | Dollar-cost averaging |
| S&P 500 index P/E | Weighted Mean | Market-cap weighted |
| "Typical" income level in a city | Median | Skewed distribution with outliers |
| Most popular sector among investors | Mode | Categorical data |
| Your fund's rank among peers | Quantile | Relative position |
VIII. Practice Questions
Multiple Choice
Q1: A fund's returns over the past 4 years are +10%, -5%, +20%, +15%. To calculate the fund's compound annual growth rate (CAGR), which mean should be used? - A. Arithmetic Mean - B. Geometric Mean - C. Harmonic Mean - D. Weighted Mean
Q2: For which type of data can the median, but NOT the mean, be used as a measure of central tendency? - A. Normally distributed returns - B. Severely right-skewed income data - C. Z-scores from a standard normal distribution - D. Daily closing prices of an ETF
Q3: The relationship among harmonic, geometric, and arithmetic means is: - A. Arithmetic ≥ Geometric ≥ Harmonic (unless all values are equal) - B. Harmonic ≥ Geometric ≥ Arithmetic (unless all values are equal) - C. Geometric ≥ Arithmetic ≥ Harmonic (unless all values are equal) - D. There is no fixed relationship
Q4: Nominal data (e.g., investment style labels: Growth/Value/Blend) can only use which measure of central tendency? - A. Arithmetic Mean - B. Median - C. Mode - D. Geometric Mean
Q5: An analyst wants to calculate the average P/E ratio of S&P 500 constituents. Which method is most correct? - A. Simple arithmetic average of all 500 P/E ratios - B. Market-cap weighted average - C. Free-float weighted average - D. Median P/E
Q6: In a portfolio, if Q1 return = 2% and Q3 return = 12%, which of the following is correct? - A. 25% of periods had returns exceeding 12% - B. 50% of periods had returns between 2% and 12% - C. The median return must equal 7% - D. The minimum must be less than 2%
Q7: In a positively skewed (right-skewed) distribution, the relationship among mean, median, and mode is: - A. Mean < Median < Mode - B. Mean > Median > Mode - C. Mean = Median = Mode - D. Median > Mean > Mode
Answers
Q1: B — Time-series returns measure compound growth and must use geometric mean. CAGR = [(1.10)(0.95)(1.20)(1.15)]^(1/4) - 1 ≈ 9.47%. Arithmetic mean = (10% - 5% + 20% + 15%) / 4 = 10%, which overstates the true compound growth rate.
Q2: B — Income data is typically severely right-skewed (a few extremely high incomes pull the mean upward). The median better represents the "typical" income level. For normally distributed data, the mean and median are equal.
Q3: A — Arithmetic ≥ Geometric ≥ Harmonic, with equality only when all values are identical. This is a classic CFA Level I concept.
Q4: C — Nominal data (pure categorical labels) has no numerical values and no ordering. Only the mode can count which category appears most often. You cannot "average" Growth and Value.
Q5: B — Market-cap weighted P/E is the industry standard for index valuation. Simple arithmetic average would treat small-cap and large-cap stocks equally, distorting the index's true valuation level.
Q6: B — The interquartile range (Q1 to Q3) contains the middle 50% of the data (from the 25th to the 75th percentile), so 50% of periods had returns between 2% and 12%. A is incorrect — 25% exceed Q3. C is not necessarily true. D is not guaranteed.
Q7: B — In a positively skewed distribution, the long tail is on the right, pulling the mean upward. The mode sits at the peak (highest point), the median in the middle, and the mean furthest right: Mean > Median > Mode.
IX. Key Exam Points
| Priority | Concept | Probability | Memory Cue |
|---|---|---|---|
| ⭐⭐⭐ | When to use Arithmetic/Geometric/Harmonic Mean | Very High | Cross-section → Arithmetic; Time-series returns → Geometric; DCA cost → Harmonic |
| ⭐⭐⭐ | Positive skew: Mean > Median > Mode | Very High | Think of income distribution — long tail on the right |
| ⭐⭐⭐ | Median vs. Mean selection | High | Outliers/skewness → Median |
| ⭐⭐⭐ | Weighted mean concept | High | Index calculation, portfolio returns |
| ⭐⭐ | Quantile position formula L = (n+1)×(y/100) | Medium | Different handling for integer vs. non-integer L |
| ⭐⭐ | Nominal data can only use Mode | Medium | No arithmetic operations possible on labels |
| ⭐⭐ | Harmonic ≤ Geometric ≤ Arithmetic | Medium | Unless all values are equal |
| ⭐ | Unimodal/Bimodal distinction | Low | Multimodal distributions |
X. Quick Memory Card
📌 One-Line Summary:
Arithmetic Mean → Cross-sectional expected value (most common)
Geometric Mean → Compound growth rate / CAGR (mandatory for time series)
Harmonic Mean → Dollar-cost averaging cost (equal-dollar purchases)
Median → "Typical value" for skewed data
Mode → "Most popular" for categorical data
Weighted Mean → Market-cap weighting for indices/portfolios
Quantiles → Relative rank (how many peers you beat)
CFA Level I · Quantitative Methods · Module 2.3 Begins | L105 Central Tendency · 2026-07-11