定量方法(Quantitative Methods)— 概率论模块 · 第七课
一、本课定位
L120 学了正态分布:钟形对称曲线,由 μ 和 σ 完全定义,68-95-99.7 经验法则快速估算概率。但经验法则只能给 μ±1σ、μ±2σ、μ±3σ 三个整数倍的结果。想精确知道 P(收益 > 2.3%),就需要标准正态分布和 z 分数。
| 项目 | 说明 |
|---|---|
| 模块 | 2.4 概率论 |
| 前置知识 | L120 正态分布(μ、σ、68-95-99.7 规则) |
| 后续衔接 | L122 对数正态分布 |
| 难度 | ★★★★☆ |
| 考试权重 | 高(z 表查值 + 计算,3-4 题) |
| 阅读时间 | 约 15 分钟 |
二、核心概念
1. 为什么需要标准化?
正态分布 N(μ, σ²) 有无穷多种——不同的 μ,不同的 σ,生成无数条不同钟形曲线。为每种 (μ, σ) 组合都编一本概率表是不可能的。
解法:把所有正态分布转化到同一个"标尺"上。 就像把美元、欧元、日元全部换成同一种基准货币——z 分数就是"统一换算工具"。
2. 标准正态分布(Standard Normal Distribution)
$$\text{记作: } Z \sim N(0, 1)$$
$$f(z) = \frac{1}{\sqrt{2\pi}} \cdot e^{-\frac{z^2}{2}}$$
| 性质 | 标准正态 |
|---|---|
| 均值 μ | 0 |
| 标准差 σ | 1 |
| 68-95-99.7 | 68% 在 [-1,1];95% 在 [-2,2];99.7% 在 [-3,3] |
🧠 标准正态分布只有一套 z 表——查一次,管天下。
3. Z 分数(Z-Score)
Z 分数的含义:"X 距离均值多少倍标准差?"
$$z = \frac{X - \mu}{\sigma}$$
| 符号 | 含义 |
|---|---|
| X | 原始数据点 |
| μ | 总体均值 |
| σ | 总体标准差 |
| z | 标准化值(偏离均值几个 σ) |
案例 1:标准化计算
X ~ N(100, 25),即 μ = 100,σ = 5:
| X 值 | z = (X−100)/5 | z 值 | 含义 |
|---|---|---|---|
| 110 | (110−100)/5 | +2.0 | 比均值高 2σ |
| 92.5 | (92.5−100)/5 | −1.5 | 比均值低 1.5σ |
| 100 | (100−100)/5 | 0 | 正好在均值上 |
📊 标准化"抹去"了原始单位(元、%、点),只剩下无单位的 σ 倍数。
4. Z 表的四种查表场景(考试核心)
以下是精简 z 表(CFA 考试常用值):
| z | 0.00 | 0.05 | 0.10 |
|---|---|---|---|
| 0.0 | 0.5000 | 0.5199 | 0.5398 |
| 0.5 | 0.6915 | 0.7088 | 0.7257 |
| 1.0 | 0.8413 | 0.8531 | 0.8643 |
| 1.5 | 0.9332 | 0.9394 | 0.9452 |
| 1.96 | 0.9750 | — | — |
| 2.0 | 0.9772 | 0.9798 | 0.9821 |
| 2.5 | 0.9938 | 0.9946 | 0.9953 |
| 3.0 | 0.9987 | 0.9989 | 0.9990 |
四种场景:
| 场景 | 公式 | 示例 |
|---|---|---|
| ① P(X ≤ a) 左尾 | 直接查 P(Z ≤ z) | P(Z ≤ 1.2) = 0.8849 |
| ② P(X ≥ a) 右尾 | 1 − P(Z ≤ z) | P(Z ≥ 1.2) = 1 − 0.8849 = 0.1151 |
| ③ P(a ≤ X ≤ b) 中间 | F(b) − F(a) | P(−1 ≤ Z ≤ 1) = 0.8413 − 0.1587 = 0.6826 |
| ④ 已知概率求临界值 | 反查 z 表 | 95% 双尾 → z = 1.96 |
负 z 处理——利用对称性:
$$P(Z \leq -z) = 1 - P(Z \leq z)$$
案例 2:实际金融问题
某基金月收益 ~ N(0.8%, 2.5%²)。求亏损超过 2% 的概率。
$$z = \frac{-2\% - 0.8\%}{2.5\%} = -1.12$$
P(Z ≤ −1.12) = 1 − P(Z ≤ 1.12) = 1 − 0.8686 = 0.1314 ≈ 13.14%
📊 该基金每月有约 13% 的概率亏损超过 2%。
5. 关键 z 值——务必背诵
| z | P(Z ≤ z) | 应用 |
|---|---|---|
| 0 | 0.5000 | 均值点 |
| 1.00 | 0.8413 | 1σ 右侧 |
| 1.28 | ≈0.90 | 90% 单尾 |
| 1.645 | 0.95 | 95% 单尾 |
| 1.96 | 0.975 | 95% 双尾 |
| 2.00 | 0.9772 | 2σ 右侧 |
| 2.33 | ≈0.99 | 99% 单尾 |
| 2.58 | 0.995 | 99% 双尾 |
| 3.00 | 0.9987 | 3σ 右侧 |
🎯 口诀:「单尾 95→1.645,双尾 95→1.96,单尾 99→2.33,双尾 99→2.58」
6. 逆标准化:从概率反推 X
$$X = \mu + z \cdot \sigma$$
案例 3:VaR 逆向计算
投资组合日收益 ~ N(0.05%, 1.1%²)。95% 置信水平日 VaR:
- 95% 置信 → P(Z ≤ z) = 0.05 → z = −1.645
- VaR = 0.05% + (−1.645)(1.1%) = −1.76%
📊 95% 置信水平下,单日最大损失不超过 1.76%。
7. z 分数的金融应用矩阵
| 应用 | z 分数角色 |
|---|---|
| VaR 计算 | z = (损失阈值 − μ)/σ |
| 异常检测 | |
| 跨资产比较 | 不同收益标准化后可比 |
| 信用评分 | Altman Z-Score 破产预测 |
| 业绩归因 | α/σ(α) > 1.96 → 统计显著 |
三、核心公式速记
| 公式 | 用途 |
|---|---|
| z = (X − μ) / σ | X 标准化为 z 分数 |
| X = μ + z·σ | z 反推原始 X |
| P(Z ≤ −z) = 1 − P(Z ≤ z) | 负 z 用对称性查表 |
| P(a ≤ X ≤ b) = F(z₂) − F(z₁) | 区间概率 |
| CI: μ ± z_{α/2}·σ | 置信区间 |
四、常见陷阱
| ❌ 错误 | ✅ 正确 |
|---|---|
| 忘记 X−μ 的正负号 | X < μ 时 z 为负,必须用对称性 |
| z 分数和原始 X 混淆 | z = 2 是偏离 2σ,不是 X = 2 |
| 方差当标准差用 | z 的分母是 σ,不是 σ² |
| 区间概率用错 | P(a ≤ X ≤ b) = F(b) − F(a),不是加 |
| 1.96 做单尾 | 1.96 是双尾 95%;单尾 95% 用 1.645 |
| 不画图直接查表 | 先画示意图确认要左尾/右尾/中间 |
五、实战测试
【测试题】
Q1(概念题) X ~ N(50, 25)。计算 z 分数时正确的做法是:
A. 计算 z = (X − 50) / 5,因为 σ = 5
B. 计算 z = (X − 50) / 25,因为 σ² = 25
C. 计算 z = X − 50 / 5,因为减法优先级高于除法
D. 不需要计算 z,直接用 X 查标准正态表
Q2(计算题) X ~ N(80, 144)。求 P(X > 95) 最接近的值:
A. 0.1056
B. 0.1250
C. 0.1587
D. 0.2112
Q3(计算题) X ~ N(200, 400)。若 P(X < k) = 0.8413,则 k 最接近: (已知 P(Z ≤ 1.00) = 0.8413)
A. 210
B. 220
C. 240
D. 260
Q4(综合题) 三只股票标准化收益(z 分数)计算中:
| 股票 | 收益 | μ | σ | z 公式 |
|---|---|---|---|---|
| A | 5.2% | 3.0% | 1.5% | (5.2−3.0)/1.5 |
| B | −1.8% | 0.5% | 0.8% | (−1.8−0.5)/0.8 |
| C | 8.0% | 4.0% | 2.0% | (8.0−4.0)/2.0 |
按 z 分数从大到小排列(最佳 → 最差):
A. C > A > B
B. A > C > B
C. C > B > A
D. A > B > C
Q5(进阶概念题) 关于 z 分数,哪一项错误?
A. z 分数衡量偏离均值的标准差倍数
B. 即使原始分布非正态,z 分数依然可计算
C. z > 0 说明 X 低于均值
D. 标准化后任何数据的样本均值和样本标准差变为 0 和 1
【答案与解析】
A1:A
- A ✅ σ² = 25 ⇒ σ = 5,z = (X − 50)/5 正确
- B ❌ z 的分母必须是标准差 σ = 5,不是方差 σ² = 25
- C ❌ 先做减法再做除法,需要括号:(X − 50)/5
- D ❌ 原始 X 不能直接用标准正态表,必须先标准化
A2:A — 0.1056
计算过程: - σ² = 144 ⇒ σ = 12 - z = (95 − 80) / 12 = 15/12 = 1.25 - P(Z ≤ 1.25) = 0.8944 - P(Z > 1.25) = 1 − 0.8944 = 0.1056
⚠️ 注意:先准确计算 z 值,区分左尾/右尾的转换。
A3:B — 220
- σ² = 400 ⇒ σ = 20
- P(Z ≤ 1.00) = 0.8413 ⇒ z = 1.00
- k = μ + z·σ = 200 + 1.00 × 20 = 220
A4:B — A > C > B
逐只计算 z: - 股票 A:z_A = (5.2 − 3.0)/1.5 = 2.2/1.5 = 1.467 - 股票 B:z_B = (−1.8 − 0.5)/0.8 = −2.3/0.8 = −2.875 - 股票 C:z_C = (8.0 − 4.0)/2.0 = 4.0/2.0 = 2.000
排名:C(z=2.0) > A(z=1.47) > B(z=−2.88)?等等……
仔细看:z_C = 2.000,z_A = 1.467,顺序是 C > A > B。
🧠 等一下——按照这个计算,应该是 C > A > B,对应选项 A。
但发现问题出在这里: z_A = (5.2 − 3.0)/1.5 = 2.2/1.5 ≈ 1.467 z_C = (8.0 − 4.0)/2.0 = 4.0/2.0 = 2.000 z_B = (−1.8 − 0.5)/0.8 = −2.3/0.8 = −2.875
C(2.00) > A(1.47) > B(−2.88) → 选项 A:C > A > B
📊 z 分数不仅能排序,还能量化"好多少":C 比均值多 2σ,A 多 1.47σ,B 比均值少 2.88σ。
A5:C
- A ✅ z 分数的定义就是"偏离均值的标准差的个数"
- B ✅ z 分数只是数学变换:(X−μ)/σ,不需要正态分布假设
- C ❌ z > 0 意味着 X 高于均值(分子 X−μ > 0)
- D ✅ 标准化后样本均值变为 0,样本标准差变为 1(数学性质)
这道题容易选错 D,因为"任何数据"听起来太绝对。但标准化是纯数学操作——把数据减去均值除以标准差后,新数据的均值必然是 0,标准差必然是 1。这和分布形状无关。
📚 下一课 L122:对数正态分布——为什么资产价格用对数正态而不是正态建模
Quantitative Methods — Probability Module · Lesson 7
I. Lesson Positioning
L120 covered the normal distribution: a bell-shaped symmetric curve fully defined by μ and σ, with the 68-95-99.7 empirical rule for quick probability estimates. But that rule only gives you the three integer σ multiples. To precisely compute P(return > 2.3%), you need the standard normal distribution and z-scores.
| Item | Description |
|---|---|
| Module | 2.4 Probability |
| Prerequisite | L120 Normal Distribution (μ, σ, 68-95-99.7 rule) |
| Next | L122 Log-normal Distribution |
| Difficulty | ★★★★☆ |
| Exam Weight | High (z-table lookup + calculation, 3–4 questions) |
| Reading Time | ~15 minutes |
II. Core Concepts
1. Why Standardize?
The normal distribution N(μ, σ²) has infinitely many possibilities — different μ, different σ, generating countless different bell curves. Publishing a separate probability table for every (μ, σ) combination is impossible.
Solution: transform all normal distributions onto a single "ruler." Just as you'd convert USD, EUR, and JPY into a single base currency to compare them — the z-score is that "universal conversion tool."
2. Standard Normal Distribution
$$\text{Denoted: } Z \sim N(0, 1)$$
$$f(z) = \frac{1}{\sqrt{2\pi}} \cdot e^{-\frac{z^2}{2}}$$
| Property | Standard Normal |
|---|---|
| Mean μ | 0 |
| Std Dev σ | 1 |
| 68-95-99.7 | 68% in [−1,1]; 95% in [−2,2]; 99.7% in [−3,3] |
🧠 The standard normal has only one z-table — look it up once, use it everywhere.
3. Z-Score
The z-score answers: "How many standard deviations is X away from the mean?"
$$z = \frac{X - \mu}{\sigma}$$
| Symbol | Meaning |
|---|---|
| X | Raw data point |
| μ | Population mean |
| σ | Population standard deviation |
| z | Standardized value (# of σ away from μ) |
Example 1: Standardization
X ~ N(100, 25), i.e., μ = 100, σ = 5:
| X | z = (X−100)/5 | z-value | Meaning |
|---|---|---|---|
| 110 | (110−100)/5 | +2.0 | 2σ above mean |
| 92.5 | (92.5−100)/5 | −1.5 | 1.5σ below mean |
| 100 | (100−100)/5 | 0 | exactly at mean |
📊 Standardization strips away the original units (%, dollars, points), leaving only unit-free σ multiples.
4. Four Z-Table Lookup Scenarios (Exam Core)
Abbreviated z-table (CFA exam common values):
| z | 0.00 | 0.05 | 0.10 |
|---|---|---|---|
| 0.0 | 0.5000 | 0.5199 | 0.5398 |
| 0.5 | 0.6915 | 0.7088 | 0.7257 |
| 1.0 | 0.8413 | 0.8531 | 0.8643 |
| 1.5 | 0.9332 | 0.9394 | 0.9452 |
| 1.96 | 0.9750 | — | — |
| 2.0 | 0.9772 | 0.9798 | 0.9821 |
| 2.5 | 0.9938 | 0.9946 | 0.9953 |
| 3.0 | 0.9987 | 0.9989 | 0.9990 |
Four scenarios:
| Scenario | Formula | Example |
|---|---|---|
| ① P(X ≤ a) left tail | Direct lookup P(Z ≤ z) | P(Z ≤ 1.2) = 0.8849 |
| ② P(X ≥ a) right tail | 1 − P(Z ≤ z) | P(Z ≥ 1.2) = 1 − 0.8849 = 0.1151 |
| ③ P(a ≤ X ≤ b) middle | F(b) − F(a) | P(−1 ≤ Z ≤ 1) = 0.8413 − 0.1587 = 0.6826 |
| ④ Find cutoff from probability | Reverse lookup | 95% two-tailed → z = 1.96 |
Handling negative z — use symmetry:
$$P(Z \leq -z) = 1 - P(Z \leq z)$$
Example 2: Real Finance Problem
A fund's monthly return ~ N(0.8%, 2.5%²). Probability of a loss exceeding 2%:
$$z = \frac{-2\% - 0.8\%}{2.5\%} = -1.12$$
P(Z ≤ −1.12) = 1 − P(Z ≤ 1.12) = 1 − 0.8686 = 0.1314 ≈ 13.14%
📊 This fund has roughly a 13% chance of losing more than 2% in any given month.
5. Critical Z-Values — Must Memorize
| z | P(Z ≤ z) | Application |
|---|---|---|
| 0 | 0.5000 | Mean point |
| 1.00 | 0.8413 | 1σ to the right |
| 1.28 | ≈0.90 | 90% one-tailed |
| 1.645 | 0.95 | 95% one-tailed |
| 1.96 | 0.975 | 95% two-tailed |
| 2.00 | 0.9772 | 2σ to the right |
| 2.33 | ≈0.99 | 99% one-tailed |
| 2.58 | 0.995 | 99% two-tailed |
| 3.00 | 0.9987 | 3σ to the right |
🎯 Mnemonic: "One-tail 95 → 1.645, two-tail 95 → 1.96, one-tail 99 → 2.33, two-tail 99 → 2.58"
6. Reverse Standardization: From Probability Back to X
$$X = \mu + z \cdot \sigma$$
Example 3: VaR Calculation
Portfolio daily return ~ N(0.05%, 1.1%²). Daily VaR at 95% confidence:
- 95% confidence → P(Z ≤ z) = 0.05 → z = −1.645
- VaR = 0.05% + (−1.645)(1.1%) = −1.76%
📊 At 95% confidence, the max single-day loss is 1.76%.
7. Financial Applications of Z-Scores
| Application | Z-Score Role |
|---|---|
| VaR | z = (loss threshold − μ) / σ |
| Anomaly Detection | |
| Cross-asset Comparison | Different returns become comparable after standardization |
| Credit Scoring | Altman Z-Score for bankruptcy prediction |
| Performance Attribution | α / σ(α) > 1.96 → statistically significant |
III. Key Formulas
| Formula | Purpose |
|---|---|
| z = (X − μ) / σ | Standardize X to a z-score |
| X = μ + z·σ | Convert z back to original X |
| P(Z ≤ −z) = 1 − P(Z ≤ z) | Negative z via symmetry |
| P(a ≤ X ≤ b) = F(z₂) − F(z₁) | Interval probability |
| CI: μ ± z_{α/2}·σ | Confidence interval |
IV. Common Pitfalls
| ❌ Mistake | ✅ Correct |
|---|---|
| Ignoring sign of X − μ | When X < μ, z is negative — must use symmetry |
| Confusing z-score with raw X | z = 2 means 2σ away, not X = 2 |
| Using variance instead of std dev | Denominator is σ, not σ² |
| Wrong operation for interval prob | P(a ≤ X ≤ b) = F(b) − F(a), not addition |
| Using 1.96 for one-tailed test | 1.96 is two-tailed 95%; use 1.645 for one-tailed |
| Looking up z-table without sketching | Always draw a quick diagram to confirm left/right/middle |
V. Practice Questions
【Questions】
Q1 (Concept) X ~ N(50, 25). When computing the z-score, the correct approach is:
A. Compute z = (X − 50) / 5, because σ = 5
B. Compute z = (X − 50) / 25, because σ² = 25
C. Compute z = X − 50 / 5, because subtraction takes priority over division
D. No z needed; look up X directly in the standard normal table
Q2 (Calculation) X ~ N(80, 144). The value of P(X > 95) is closest to:
A. 0.1056
B. 0.1250
C. 0.1587
D. 0.2112
Q3 (Calculation) X ~ N(200, 400). If P(X < k) = 0.8413, then k is closest to: (Given: P(Z ≤ 1.00) = 0.8413)
A. 210
B. 220
C. 240
D. 260
Q4 (Comprehensive) Standardized returns (z-scores) for three stocks:
| Stock | Return | μ | σ | z Formula |
|---|---|---|---|---|
| A | 5.2% | 3.0% | 1.5% | (5.2−3.0)/1.5 |
| B | −1.8% | 0.5% | 0.8% | (−1.8−0.5)/0.8 |
| C | 8.0% | 4.0% | 2.0% | (8.0−4.0)/2.0 |
Rank by z-score from highest to lowest (best → worst performer):
A. C > A > B
B. A > C > B
C. C > B > A
D. A > B > C
Q5 (Advanced Concept) Which statement about z-scores is incorrect?
A. A z-score measures the number of standard deviations from the mean
B. Z-scores can be computed even when the underlying distribution is non-normal
C. z > 0 means X is below the mean
D. After standardization, any dataset's sample mean becomes 0 and sample standard deviation becomes 1
【Answers & Explanations】
A1: A
- A ✅ σ² = 25 ⇒ σ = 5, so z = (X − 50)/5 is correct
- B ❌ The denominator must be standard deviation σ = 5, not variance σ² = 25
- C ❌ Subtraction first, then division — need parentheses: (X − 50)/5
- D ❌ Raw X values cannot be used directly with the standard normal table — must standardize first
A2: A — 0.1056
Calculation: - σ² = 144 ⇒ σ = 12 - z = (95 − 80) / 12 = 15/12 = 1.25 - P(Z ≤ 1.25) = 0.8944 - P(Z > 1.25) = 1 − 0.8944 = 0.1056
⚠️ Note: Compute z accurately first, then distinguish between left-tail and right-tail conversions.
A3: B — 220
- σ² = 400 ⇒ σ = 20
- P(Z ≤ 1.00) = 0.8413 ⇒ z = 1.00
- k = μ + z·σ = 200 + 1.00 × 20 = 220
A4: A — C > A > B
Compute each z-score: - Stock A: z_A = (5.2 − 3.0)/1.5 = 2.2/1.5 = 1.467 - Stock B: z_B = (−1.8 − 0.5)/0.8 = −2.3/0.8 = −2.875 - Stock C: z_C = (8.0 − 4.0)/2.0 = 4.0/2.0 = 2.000
Ranking: C (2.00) > A (1.47) > B (−2.88)
📊 Z-scores not only rank performance but quantify "how much better": C is 2σ above its mean, A is 1.47σ above, B is 2.88σ below.
A5: C
- A ✅ Definition of z-score: number of standard deviations from the mean
- B ✅ Z-score is purely a mathematical transformation: (X − μ)/σ, no normality assumption required
- C ❌ z > 0 means X is above the mean (numerator X − μ > 0)
- D ✅ After standardization, the sample mean always becomes 0 and sample standard deviation becomes 1 (mathematical property)
D may sound "too absolute," but it's a mathematical certainty: subtracting the mean makes the new mean 0, and dividing by σ makes the new standard deviation 1. This holds regardless of distribution shape.
📚 Next Lesson L122: Log-normal Distribution — why asset prices use log-normal instead of normal modeling