课题:数据不会说话?分位数是你的翻译官
一、引言:为什么基金经理总能说"我们跑赢了一半同行"?
你见过这样的基金宣传吗: - "本基金业绩超过 75% 的同类产品" - "我们的回报率处于行业前 25%"
这些陈述背后的数学工具,就是分位数(Quantile)。
分位数是统计学中最实用、最直观的工具之一——它不关心平均值,只关心你在群体中的相对位置。
🔥 核心直觉:均值告诉你"大家平均赚多少";分位数告诉你"你比别人强多少"。在投资领域,后者往往更重要。
二、什么是分位数?
2.1 定义
分位数(Quantile) 是将一个有序数据集分割为等份的点。
给你一个已排序的数据集,第 p 分位数是这样一个值:有 p 比例的数据 ≤ 该值。
关键前提:数据必须先排序(从小到大)。 未排序的分位数没有意义。
2.2 通用公式
设数据量为 n,要求第 p 分位数:
位置指数:L_y = (n + 1) × y
其中 y 是分位点(百分位数时 y = p/100)。
三、三大核心分位数类型
3.1 四分位数(Quartiles)
将数据分成 4 等份。
| 四分位数 | 含义 | 分位点 | 别称 |
|---|---|---|---|
| Q₁(第一四分位数) | 25% 的数据 ≤ 此值 | 0.25 | 下四分位数 |
| Q₂(第二四分位数) | 50% 的数据 ≤ 此值 | 0.50 | 中位数 |
| Q₃(第三四分位数) | 75% 的数据 ≤ 此值 | 0.75 | 上四分位数 |
实战案例:基金评级
某平台将 1000 只基金按年化回报排序: - Q₁ = 第 250 名的回报:落后基金分界线 - Q₂ = 第 500 名的回报:市场平均水平 - Q₃ = 第 750 名的回报:优秀基金分界线
如果你的基金超过 Q₃ → 你在前 25%,可以上宣传材料了 📈
3.2 百分位数(Percentiles)
将数据分成 100 等份。
第 k 百分位数 = 有 k% 的数据 ≤ 该值。
关键对应关系: - P₂₅ = Q₁ - P₅₀ = Q₂ = 中位数 - P₇₅ = Q₃
实战案例:考试成绩
CFA 考试的通过线并非固定分数,而是基于最低及格分数(MPS)——本质上是一个经过校准的百分位数概念: - 你的得分如果超过 MPS 对应的百分位 → 通过 - 这是你和同期考生的相对竞赛,而非绝对分数
3.3 五分位数(Quintiles)与十分位数(Deciles)
| 类型 | 等份数 | 关键分位点 | 应用场景 |
|---|---|---|---|
| 五分位数 | 5 等份 | 20%, 40%, 60%, 80% | 行业分组(如前 20%) |
| 十分位数 | 10 等份 | 10%, 20%, ..., 90% | 收入分配、风险分层 |
四、计算方法(核心考点)
4.1 位置公式
对于 n 个已排序数据,第 y 分位数(0 < y < 1)的位置:
L_y = (n + 1) × y
4.2 两步计算法
- 计算 L_y = (n+1) × y
- 如果 L_y 是整数 → 直接取第 L_y 个数据
- 如果 L_y 不是整数 → 线性插值
4.3 线性插值法
设 L_y = a + d,其中 a 是整数部分,d 是小数部分(0 < d < 1):
分位数 = X_a + d × (X_{a+1} − X_a)
4.4 完整计算示例
数据集(已排序): 3, 5, 7, 8, 9, 11, 13, 15
n = 8
求 Q₁(第 25 百分位): - L = (8+1) × 0.25 = 9 × 0.25 = 2.25 - a = 2, d = 0.25 - Q₁ = X₂ + 0.25 × (X₃ − X₂) = 5 + 0.25 × (7−5) = 5 + 0.5 = 5.5
求 Q₂(中位数 / 第 50 百分位): - L = (8+1) × 0.50 = 9 × 0.50 = 4.5 - a = 4, d = 0.5 - Q₂ = X₄ + 0.5 × (X₅ − X₄) = 8 + 0.5 × (9−8) = 8.5
求 Q₃(第 75 百分位): - L = (8+1) × 0.75 = 9 × 0.75 = 6.75 - a = 6, d = 0.75 - Q₃ = X₆ + 0.75 × (X₇ − X₆) = 11 + 0.75 × (13−11) = 11 + 1.5 = 12.5
4.5 四分位距(IQR)
IQR = Q₃ − Q₁
上例 IQR = 12.5 − 5.5 = 7.0
📌 IQR 是衡量数据中间 50% 离散程度的核心指标,不受极端值影响。
五、分位数的三大应用
5.1 箱线图(Box Plot)——可视化利器
下边界 Q₁ Q₂ Q₃ 上边界
| |─────|─────| |
o---+----------+=====+=====+----------+---o
异常值 └── IQR ──┘ 异常值
箱线图构造规则: - 箱体:Q₁ 到 Q₃ - 箱内线:Q₂(中位数) - 下须(Lower Whisker):Q₁ − 1.5 × IQR 范围内的最小值 - 上须(Upper Whisker):Q₃ + 1.5 × IQR 范围内的最大值 - 须外的点:异常值(Outliers)
5.2 异常值识别
温和异常值: 距离 Q₁ 或 Q₃ 超过 1.5 × IQR 极端异常值: 距离 Q₁ 或 Q₃ 超过 3.0 × IQR
📌 这是 CFA 一级常见考点:基于 IQR 的异常值判断 vs 基于标准差的判断。
5.3 相对位置评估
应用场景: 你想知道某只股票的表现到底处于什么水平。
- 将同类 500 只股票排序
- 你的股票 P₈₀ 分位 → 超过 80% 的同类
- 这是一个比你报"平均涨了 15%"更有说服力的数字
六、注意事项与常见陷阱
6.1 陷阱一:忘记排序
❌ 错误做法:直接对原始数据取某个位置的值 ✅ 正确做法:必须先排序,再计算分位数
6.2 陷阱二:公式混淆
不同教材用的 (n+1) 系数可能略有差异(有些用 n,有些用 n-1),CFA 一级统一使用 L_y = (n+1) × y。
6.3 陷阱三:分位数 vs 分位点
- 分位点(y): 一个比例值,如 0.25、0.75
- 分位数(Quantile): 该分位点对应的数据值
- 例:P₂₅(第 25 百分位数) = 5.5,其中 0.25 是分位点,5.5 是分位数
6.4 陷阱四:中位数的韧性
中位数不受极端值影响,均值会被极端值严重扭曲。
| 数据集 | 均值 | 中位数 |
|---|---|---|
| [3, 5, 7, 8, 9, 11, 13, 15] | 8.875 | 8.5 |
| [3, 5, 7, 8, 9, 11, 13, 150] | 25.75 | 8.5 |
均值从 8.875 暴涨到 25.75,中位数纹丝不动。 这就是为什么居民收入统计"平均数被平均"——要用中位数看真相。
七、总结:分位数工具箱备忘
| 工具 | 公式/定义 | 用途 |
|---|---|---|
| 位置指数 | L_y = (n+1) × y | 定位分位数值 |
| Q₁ | L_0.25 | 下四分位 |
| Q₂ | L_0.50 | 中位数 |
| Q₃ | L_0.75 | 上四分位 |
| IQR | Q₃ − Q₁ | 中间 50% 离散度 |
| 异常值边界 | Q₁ − 1.5×IQR / Q₃ + 1.5×IQR | 识别异常值 |
| 线性插值 | X_a + d × (X_{a+1} − X_a) | 非整数位置计算 |
🔥 记住:统计学不是算数字,是讲故事。分位数帮你讲好"我在哪里"这个故事。
八、测试题
题目 1
数据集(已排序):[2, 4, 6, 8, 10, 12, 14, 16, 18, 20],n = 10。
求 Q₁。
A. 5.5 B. 5.25 C. 6.0 D. 5.75
题目 2
同上数据集,求第 80 百分位数 P₈₀。
A. 16.8 B. 16.0 C. 17.2 D. 17.8
题目 3
以下关于分位数的说法,哪一个是错误的?
A. 中位数是第 50 百分位数,也是第二四分位数 Q₂ B. 计算分位数前必须先对数据排序 C. IQR 对极端值非常敏感,因此不适合用于异常值检测 D. P₂₅ 和 Q₁ 是等价的
九、答案与解析
题目 1 答案:D. 5.75
解析:
Q₁ → y = 0.25 L = (10+1) × 0.25 = 11 × 0.25 = 2.75 a = 2, d = 0.75 Q₁ = X₂ + 0.75 × (X₃ − X₂) = 4 + 0.75 × (6 − 4) = 4 + 1.5 = 5.75
题目 2 答案:A. 16.8
解析:
P₈₀ → y = 0.80 L = (10+1) × 0.80 = 11 × 0.80 = 8.8 a = 8, d = 0.8 P₈₀ = X₈ + 0.8 × (X₉ − X₈) = 16 + 0.8 × (18−16) = 16 + 1.6 = 16.8
题目 3 答案:C
解析:
C 是错误的。IQR 的设计初衷正是不受极端值影响(因为它只看 Q₁ 和 Q₃,即中间 50% 数据)。基于 IQR 的异常值检测方法(箱线图法)是统计学中识别异常值的标准工具之一,正是因为 IQR 具有对极端值的稳健性。
A ✅ 中位数 = P₅₀ = Q₂ B ✅ 排序是计算分位数的前提 C ❌ 刚好说反——IQR 对极端值不敏感,天然适合异常值检测 D ✅ P₂₅ 和 Q₁ 都是 25% 分位点
十、关键公式速记
| 公式 | 记忆口诀 |
|---|---|
| L_y = (n+1) × y | "位置指数 = 总数加一乘分位" |
| IQR = Q₃ − Q₁ | "箱体宽度就是四分位距" |
| 插值 = X_a + d × (X_{a+1}−X_a) | "基数加比例乘间距" |
| 异常值下界 = Q₁ − 1.5×IQR | "1.5 倍箱宽定异常" |
CFA 一级 · L107 · 分位数:四分位数、百分位数 · 中文版 · 2026-07-13
Topic: Data Doesn't Speak — Quantiles Are Your Interpreter
1. Introduction: Why Every Fund Manager Claims "We Beat Half Our Peers"
You've seen this before: - "This fund outperforms 75% of peers" - "Our returns are in the top quartile"
Behind these statements lies a single mathematical tool: quantiles.
Quantiles are among the most practical and intuitive tools in statistics — they don't care about averages, only about your relative position in the group.
🔥 Core intuition: The mean tells you "how much everyone earns on average"; quantiles tell you "how much better you are than others." In investing, the latter often matters more.
2. What Is a Quantile?
2.1 Definition
A quantile is a value that divides a sorted dataset into equal-sized groups.
Given a sorted dataset, the p-th quantile is the value below which a proportion p of the data falls.
Critical prerequisite: data must be sorted (ascending) first. Unsorted quantiles are meaningless.
2.2 General Formula
For n data points, to find the y-th quantile:
Position index: L_y = (n + 1) × y
where y is the quantile point (for percentiles, y = k/100).
3. Three Core Quantile Types
3.1 Quartiles
Divide data into 4 equal parts.
| Quartile | Meaning | Quantile Point | Also Known As |
|---|---|---|---|
| Q₁ (First Quartile) | 25% of data ≤ this value | 0.25 | Lower Quartile |
| Q₂ (Second Quartile) | 50% of data ≤ this value | 0.50 | Median |
| Q₃ (Third Quartile) | 75% of data ≤ this value | 0.75 | Upper Quartile |
Real-World Example: Fund Ratings
A platform ranks 1,000 funds by annualized return: - Q₁ = return of the 250th fund: laggard cutoff - Q₂ = return of the 500th fund: market median - Q₃ = return of the 750th fund: top-performer cutoff
If your fund exceeds Q₃ → you're in the top 25%. Time for the marketing brochure 📈
3.2 Percentiles
Divide data into 100 equal parts.
The k-th percentile = k% of data ≤ this value.
Key Relationships: - P₂₅ = Q₁ - P₅₀ = Q₂ = Median - P₇₅ = Q₃
Real-World Example: Exam Scores
The CFA exam passing score is not a fixed number — it is based on the Minimum Passing Score (MPS), which is essentially a calibrated percentile concept: - If your score exceeds the MPS-aligned percentile → you pass - It's a relative competition against your cohort, not an absolute score threshold
3.3 Quintiles and Deciles
| Type | Equal Parts | Key Points | Application |
|---|---|---|---|
| Quintiles | 5 | 20%, 40%, 60%, 80% | Industry groupings (e.g., top 20%) |
| Deciles | 10 | 10%, 20%, ..., 90% | Income distribution, risk stratification |
4. Calculation Method (Key Exam Topic)
4.1 Position Formula
For n sorted data points, the position of the y-th quantile (0 < y < 1):
L_y = (n + 1) × y
4.2 Two-Step Procedure
- Compute L_y = (n+1) × y
- If L_y is an integer → directly take the L_y-th data value
- If L_y is not an integer → linear interpolation
4.3 Linear Interpolation
Let L_y = a + d, where a is the integer part and d is the fractional part (0 < d < 1):
Quantile = X_a + d × (X_{a+1} − X_a)
4.4 Full Worked Example
Dataset (sorted): 3, 5, 7, 8, 9, 11, 13, 15
n = 8
Find Q₁ (25th percentile): - L = (8+1) × 0.25 = 9 × 0.25 = 2.25 - a = 2, d = 0.25 - Q₁ = X₂ + 0.25 × (X₃ − X₂) = 5 + 0.25 × (7−5) = 5 + 0.5 = 5.5
Find Q₂ (Median / 50th percentile): - L = (8+1) × 0.50 = 9 × 0.50 = 4.5 - a = 4, d = 0.5 - Q₂ = X₄ + 0.5 × (X₅ − X₄) = 8 + 0.5 × (9−8) = 8.5
Find Q₃ (75th percentile): - L = (8+1) × 0.75 = 9 × 0.75 = 6.75 - a = 6, d = 0.75 - Q₃ = X₆ + 0.75 × (X₇ − X₆) = 11 + 0.75 × (13−11) = 11 + 1.5 = 12.5
4.5 Interquartile Range (IQR)
IQR = Q₃ − Q₁
From the example: IQR = 12.5 − 5.5 = 7.0
📌 IQR is the core measure of dispersion for the middle 50% of data — it is unaffected by extreme values.
5. Three Key Applications of Quantiles
5.1 Box Plot — A Visualization Powerhouse
Lower Fence Q₁ Q₂ Q₃ Upper Fence
| |─────|─────| |
o---+----------+=====+=====+----------+---o
Outlier └── IQR ──┘ Outlier
Box Plot Construction Rules: - Box: Q₁ to Q₃ - Line inside box: Q₂ (median) - Lower Whisker: minimum value within Q₁ − 1.5 × IQR - Upper Whisker: maximum value within Q₃ + 1.5 × IQR - Points beyond whiskers: outliers
5.2 Outlier Detection
Mild outliers: more than 1.5 × IQR from Q₁ or Q₃ Extreme outliers: more than 3.0 × IQR from Q₁ or Q₃
📌 This is a common CFA Level 1 exam topic: IQR-based outlier detection vs. standard deviation-based detection.
5.3 Relative Position Assessment
Application: You want to know where a particular stock stands among its peers.
- Sort 500 peer stocks by return
- Your stock at the P₈₀ quantile → it beats 80% of peers
- This is far more compelling than saying "the average return is 15%."
6. Pitfalls and Common Traps
6.1 Trap 1: Forgetting to Sort
❌ Wrong: taking a value at some position from raw data ✅ Correct: sort first, then compute quantiles
6.2 Trap 2: Formula Confusion
Different textbooks may use slightly different coefficients (some use n, some use n−1). CFA Level 1 uniformly uses L_y = (n+1) × y.
6.3 Trap 3: Quantile Point vs. Quantile Value
- Quantile point (y): a proportion, e.g., 0.25, 0.75
- Quantile value: the data value at that quantile point
- Example: P₂₅ (25th percentile) = 5.5, where 0.25 is the quantile point, 5.5 is the quantile value
6.4 Trap 4: The Robustness of the Median
The median is unaffected by extreme values; the mean gets severely distorted.
| Dataset | Mean | Median |
|---|---|---|
| [3, 5, 7, 8, 9, 11, 13, 15] | 8.875 | 8.5 |
| [3, 5, 7, 8, 9, 11, 13, 150] | 25.75 | 8.5 |
The mean skyrockets from 8.875 to 25.75, while the median stays rock-solid. This is why income statistics often use the median to reveal the truth rather than the "average" that gets distorted by the ultra-rich.
7. Summary: Quantile Toolkit Cheat Sheet
| Tool | Formula/Definition | Purpose |
|---|---|---|
| Position Index | L_y = (n+1) × y | Locate the quantile value |
| Q₁ | L_0.25 | Lower quartile |
| Q₂ | L_0.50 | Median |
| Q₃ | L_0.75 | Upper quartile |
| IQR | Q₃ − Q₁ | Middle 50% dispersion |
| Outlier Boundaries | Q₁ − 1.5×IQR / Q₃ + 1.5×IQR | Detect outliers |
| Linear Interpolation | X_a + d × (X_{a+1} − X_a) | Non-integer position calculation |
🔥 Remember: statistics isn't about crunching numbers — it's about telling stories. Quantiles help you tell the story of "where I stand."
8. Practice Questions
Question 1
Dataset (sorted): [2, 4, 6, 8, 10, 12, 14, 16, 18, 20], n = 10.
Find Q₁.
A. 5.5 B. 5.25 C. 6.0 D. 5.75
Question 2
Using the same dataset, find the 80th percentile P₈₀.
A. 16.8 B. 16.0 C. 17.2 D. 17.8
Question 3
Which of the following statements about quantiles is incorrect?
A. The median is the 50th percentile and also the second quartile Q₂ B. Data must be sorted before computing quantiles C. IQR is highly sensitive to extreme values and therefore unsuitable for outlier detection D. P₂₅ and Q₁ are equivalent
9. Answers and Explanations
Question 1 Answer: D. 5.75
Explanation:
Q₁ → y = 0.25 L = (10+1) × 0.25 = 11 × 0.25 = 2.75 a = 2, d = 0.75 Q₁ = X₂ + 0.75 × (X₃ − X₂) = 4 + 0.75 × (6 − 4) = 4 + 1.5 = 5.75
Question 2 Answer: A. 16.8
Explanation:
P₈₀ → y = 0.80 L = (10+1) × 0.80 = 11 × 0.80 = 8.8 a = 8, d = 0.8 P₈₀ = X₈ + 0.8 × (X₉ − X₈) = 16 + 0.8 × (18−16) = 16 + 1.6 = 16.8
Question 3 Answer: C
Explanation:
C is incorrect. The IQR is specifically designed to be insensitive to extreme values (because it only considers Q₁ and Q₃ — the middle 50% of data). IQR-based outlier detection (the box plot method) is one of the standard tools in statistics for identifying outliers, precisely because of IQR's robustness to extreme values.
A ✅ Median = P₅₀ = Q₂ B ✅ Sorting is a prerequisite for quantile calculation C ❌ Gets it exactly backwards — IQR is insensitive to extremes, making it naturally suited for outlier detection D ✅ P₂₅ and Q₁ are both the 25% quantile point
10. Key Formula Quick Reference
| Formula | Memory Aid |
|---|---|
| L_y = (n+1) × y | "Position = sample size plus one times quantile" |
| IQR = Q₃ − Q₁ | "Box width equals interquartile range" |
| Interpolation = X_a + d × (X_{a+1}−X_a) | "Base plus fraction times gap" |
| Outlier Lower Bound = Q₁ − 1.5×IQR | "1.5 box-widths defines outliers" |
CFA Level 1 · L107 · Quantiles: Quartiles and Percentiles · English Version · 2026-07-13