Standard II — Integrity of Capital Markets Module 1 · 15-20% Weight Lesson 133

📖 Type I 与 Type II 错误

CFA Level I · L133 · Type I and Type II Errors

定量方法(Quantitative Methods)— 假设检验 · 第 3 课


一、上节课回顾:你已经有了判断工具

L132 教会了我们完整的一套判断流程:

设立 H₀ / Hₐ → 选择 α → 收集样本 → 计算检验统计量 → 求 p 值 → 比较 p 与 α → 做出决策

但关键问题是:

🎯 你的决策一定正确吗?不。假设检验的每一条结论都有可能出错。问题是——你会犯哪种错?

本节课将把一个「确定性判断」的幻觉拆开,让你看到背后两种不可同时消除的风险。


二、四种可能的结果:2×2 决策矩阵

2.1 现实 × 决策的交叉

在假设检验中,有两个层面: - 现实层面:H₀ 到底是真的还是假的?(我们永远不知道真相) - 决策层面:我们选择拒绝还是不拒绝 H₀?

两者交叉得到 4 种可能结果:

H₀ 实际为真 H₀ 实际为假
不拒绝 H₀ ✅ 正确决策(1 - α) ❌ Type II 错误(β)
拒绝 H₀ ❌ Type I 错误(α) ✅ 正确决策(1 - β = Power)

2.2 用两个故事来记

故事 Type I 错误 Type II 错误
🏛️ 法庭审判 冤枉好人(无罪被判有罪) 放过坏人(有罪被判无罪)
🏥 医学检测 假阳性(没病被诊为有病) 假阴性(有病被诊为没病)
📈 投资判断 误判策略有效(其实无效) 错过真正有效的策略
🚨 火灾警报 误报(没火灾但警报响了) 漏报(有火灾但警报没响)

💡 记忆法则:Type I = 冤枉好人(False Positive),Type II = 放过坏人(False Negative)


三、Type I 错误(α 错误)——「宁可错杀」

3.1 定义

Type I 错误 = 当 H₀ 实际为真时,我们错误地拒绝了 H₀。

3.2 发生概率

P(Type I 错误) = α(显著性水平)

这个概率由我们主动控制。 我们设定 α = 0.05,就等于说:「我最多接受 5% 的概率冤枉 H₀。」

3.3 金融场景实例

场景: 你发现了一个股票交易信号。

H₀: 该信号没有预测能力(平均收益 = 0)
Hₐ: 该信号有正预测能力(平均收益 > 0)
决策 现实 结果
回测 p = 0.03,α = 0.05 → 拒绝 H₀,认为信号有效 但其实信号纯属随机 ❌ Type I 错误:你上了一个假信号,真金白银去交易,结果亏钱

🔥 Type I 错误在金融中的代价:过度交易、资金损耗、对回测的过度信任。

3.4 为什么 α 就是我们主动设置的 Type I 错误率?

α = 0.05 的含义:
如果 H₀ 为真,检验统计量仍有 5% 的几率落入拒绝域。
→ 即使 H₀ 完全正确,每 100 次检验中,平均会有 5 次「冤枉」H₀。

这就是为什么 α 是「显著性水平」——它同时定义了 Type I 错误率。


四、Type II 错误(β 错误)——「永远错过」

4.1 定义

Type II 错误 = 当 H₀ 实际为假时,我们未能拒绝 H₀(错误地保留了 H₀)。

4.2 发生概率

P(Type II 错误) = β

β 不由我们直接设定,它取决于: - 效应量(真实值与 H₀ 假设值的差距):效应越大,β 越小 - 样本量 n:n 越大,β 越小 - α 的大小:α 越小 → β 越大(两者此消彼长) - 总体变异性 σ:σ 越大 → β 越大

4.3 金融场景实例

场景: 继续上面的交易信号。

H₀: 该信号没有预测能力
Hₐ: 该信号有正预测能力
决策 现实 结果
回测 p = 0.12,α = 0.05 → 不拒绝 H₀,放弃该信号 但这信号实际上是有效的 ❌ Type II 错误:你错过了一个真正能赚钱的策略

🔥 Type II 错误在金融中的代价:错失 Alpha、资金闲置、被竞争对手抢先。


五、Type I 与 Type II 的此消彼长(The Trade-off)

5.1 核心规律

⚖️ 在样本量固定时,降低 Type I 错误率(减小 α)必然增加 Type II 错误率(增大 β)。

5.2 为什么?

α = 0.10 → 拒绝域大 → 容易拒绝 H₀ → Type I 高,Type II 低
α = 0.05 → 拒绝域适中 → 适中
α = 0.01 → 拒绝域小 → 很难拒绝 H₀ → Type I 低,Type II 高
α Type I 风险 Type II 风险 适用场景
0.10 高(10%冤案率) 低 不想错过任何可能的信号
0.05 中 中 常规研究标准
0.01 低(1%冤案率) 高 冤枉代价极高(医学)

5.3 唯一同时降低两者的方法

📈 增大样本量 n 可以在不改变 α 的情况下减小 β(同时提高检验的 Power)。

为什么?n ↑ → 标准误 ↓ → 分布更集中 → 更容易区分 H₀ 和 Hₐ → β ↓


六、检验的 Power(检验功效)

6.1 定义

Power(检验功效) = 当 H₀ 实际为假时,正确拒绝 H₀ 的概率。

Power = 1 - β = P(拒绝 H₀ | H₀ 为假)

6.2 Power 的含义

Power 解读
0.80(80%) 如果效应真的存在,有 80% 的机会将其检测出来
0.50(50%) 抛硬币一样的检测能力 — 太弱了
0.95(95%) 几乎不会漏过真实效应 — 很强

📌 行业标准:Power ≥ 0.80(即 β ≤ 0.20)被认为是可接受的检验功效。

6.3 影响 Power 的四大因素

因素 变化方向 Power 变化 原因
样本量 n ↑ ↑ 标准误 ↓,分布更集中
效应量(真实偏差) ↑ ↑ 信号更强,更容易检测
α ↑ ↑ 拒绝域扩大
总体标准差 σ ↑ ↓ 噪声变大,信号更难分辨

6.4 Power 的直观图解

假设 H₀: μ = 100,真实 μ = 105,σ/√n = 2

       H₀ 分布(μ=100)        真实分布(μ=105)
           ╱ ╲                     ╱ ╲
          ╱   ╲                   ╱   ╲
         ╱     ╲                 ╱     ╲
    ────┴───────┴───────┬───────┴───────┴────
       97  99  101 103  105 107 109 111

            ←── α = Type I 区域
                           ←── Power = 1 - β
                           ←── β = Type II 区域

💡 两个分布重合越多 → Power 越低;重合越少 → Power 越高。


七、实战案例:用 Power 做投资决策

案例背景

某量化研究员发现了一个新的因子信号,声称月均超额收益为 0.30%(即 μ > 0)。已知: - 样本月均 alpha = 0.25% - 样本标准差 s = 0.80% - n = 48 个月 - α = 0.05(右尾检验) - 假设真实效应 = 0.30%(研究员的声称)

Step 1:计算检验

H₀: μ ≤ 0
Hₐ: μ > 0

SE = 0.80% / √48 ≈ 0.1155%

临界 t₀.₀₅,₄₇ ≈ 1.678
→ 当 x̄ > 0 + 1.678 × 0.1155% = 0.194% 时,拒绝 H₀

Step 2:计算 Power

如果真实 μ = 0.30%,则:

P(拒绝 H₀ | μ = 0.30%)
= P(x̄ > 0.194% | μ = 0.30%)
= P(Z > (0.194% - 0.30%) / 0.1155%)
= P(Z > -0.92)
≈ 0.8212

Power ≈ 82% → 如果该因子真的有 0.30% 的月超额收益,这个回测有 82% 的概率检测出来。

Step 3:如果样本只有 12 个月呢?

SE = 0.80% / √12 ≈ 0.231%

临界值 = 0 + 1.796 × 0.231% = 0.415%

P(拒绝 H₀ | μ = 0.30%)
= P(Z > (0.415% - 0.30%) / 0.231%)
= P(Z > 0.50)
≈ 0.3085

Power ≈ 31% → 只有 31% 的概率检测到真正有效的信号!

🔥 结论:样本量不足时,即使信号确实有效,也极大概率检测不出来!很多「回测无效」的结论可能只是 Power 太低。


八、α 与 β 的权重:情境决定

8.1 「宁可错杀」(控制 Type II)

场景 为什么优先控制 β?
流行病早期筛查 宁可让一些人虚惊一场,也别放过一例
反洗钱监控 宁可多查几笔正常交易,也别漏掉一笔可疑的
投资早期信号监控 宁可多看几个假信号,也别错过真的

🔑 这些场景中 Type II 的代价远大于 Type I。

8.2 「宁可放过」(控制 Type I)

场景 为什么优先控制 α?
临床试验审批 宁可错过一些有效药,也不能让无效药伤人
死刑判决 证据必须铁证如山,宁可放过也不冤枉
高频交易上线 宁可少上一个信号,也不能让假信号烧钱

🔑 这些场景中 Type I 的代价远大于 Type II。

8.3 CFA 职业判断框架

面对一个假设检验问题时,CFA 候选人应问:

  1. 如果错误拒绝 H₀(Type I),后果有多严重?
  2. 如果错误保留 H₀(Type II),后果有多严重?
  3. 哪个后果更不可接受?
  4. 据此选择 α 水平(更小的 α 保护 Type I,更大的 α 保护 Type II)

九、多组检验中的 Type I 错误累积

9.1 多重比较问题

单次检验的 Type I 错误率 = α = 0.05

进行 20 次独立检验时:
P(至少一次 Type I 错误)
= 1 - P(所有都正确)
= 1 - (1 - 0.05)^20
= 1 - 0.358
= 0.642

🚨 64.2%! 做 20 次检验就有 64% 概率至少出现一次假阳性!

9.2 数据挖掘的陷阱

「我在 100 个因子中找到了 5 个显著的(p < 0.05)」
→ 在 α = 0.05 下,100 个随机因子平均会有 5 个「显著」
→ 这些很可能全是 Type I 错误!

🔥 这就是为什么量化界常说:「如果你折磨数据足够久,它会招供任何事。」

9.3 Bonferroni 校正(CFA 了解即可)

调整后 α* = α / k(k = 检验次数)

例如做 10 次检验,原始 α = 0.05:
α* = 0.05 / 10 = 0.005
→ 只有 p < 0.005 才算显著

十、课堂练习

📝 Part A:基础概念

Q1. Type I 错误是指:

A. H₀ 为假时,不拒绝 H₀ B. H₀ 为真时,拒绝 H₀ C. Hₐ 为真时,拒绝 H₀ D. H₀ 为假时,拒绝 H₀

Q2. Type II 错误发生的概率为:

A. α B. 1 - α C. β D. 1 - β

Q3. 检验的 Power 是:

A. 不拒绝 H₀ 时结论正确的概率 B. 拒绝 H₀ 时结论正确的概率 C. H₀ 为假时正确拒绝 H₀ 的概率 D. H₀ 为真时不拒绝 H₀ 的概率


📝 Part B:Trade-off 理解

Q4. 在样本量固定时,将 α 从 0.05 降低到 0.01 会:

A. 同时降低 Type I 和 Type II 错误 B. 降低 Type I 错误,增加 Type II 错误 C. 增加 Type I 错误,降低 Type II 错误 D. 同时增加 Type I 和 Type II 错误

Q5. 以下哪项可以同时降低 Type I 和 Type II 错误?

A. 增大 α B. 减小 α C. 增大样本量 D. 减小效应量

Q6. Power = 0.90 意味着:

A. H₀ 有 90% 的概率为真 B. 如果 H₀ 为假,有 90% 的概率正确拒绝 H₀ C. Type I 错误率是 10% D. 结论有 90% 的概率正确


📝 Part C:应用判断

Q7. 一款新药的临床试验中,H₀: 药物无效。以下哪种错误更严重?

A. Type I 错误 — 批准无效药 B. Type II 错误 — 不批准有效药 C. 两种一样严重 D. 无法判断

Q8. 基金经理宣称策略有 Alpha。回测后 p = 0.07,α = 0.05。不拒绝 H₀。但实际上策略真有 Alpha。这属于:

A. Type I 错误 B. Type II 错误 C. 正确决策 D. 无法判断

Q9. 增大样本量的影响是:

A. 增大 α B. 减小 Power C. 增大 Power D. 增大 β

Q10. 做 50 次独立假设检验(α = 0.05),至少犯一次 Type I 错误的概率约为:

A. 5% B. 50% C. 92% D. 100%


十一、答案与解析

Part A:基础概念

Q1. 答案:B

Type I 错误 = H0 为真时错误地拒绝了 H0(冤枉好人)。A 是 Type II 错误,C/D 描述的是错误类型的混淆。

Q2. 答案:C

P(Type II 错误) = beta。alpha 是 Type I 错误概率,1-alpha 是 H0 为真时不拒绝的概率(正确),1-beta 是 Power。

Q3. 答案:C

Power = P(拒绝 H0 | H0 为假) = 1 - beta。D 描述的是 1-alpha(置信水平)。


Part B:Trade-off 理解

Q4. 答案:B

alpha 降低 → 拒绝域缩小 → 更难拒绝 H0 → Type I 下降,Type II 上升。这是此消彼长的经典体现。

Q5. 答案:C

只有增大样本量 n 可以同时降低 Type I 和 Type II 错误。n 增大 → 标准误减小 → 分布更集中 → 区分 H0 和 Ha 的能力增强。

Q6. 答案:B

Power = P(拒绝 H0 | H0 为假)。Power = 0.90 = 如果 H0 确实为假,有 90% 的概率检测出来(正确拒绝)。


Part C:应用判断

Q7. 答案:A

药物审批中 Type I 错误(批准无效药)的代价是患者受到伤害甚至死亡,远比 Type II 错误(暂时错过有效药)严重。因此 FDA 等机构会选择极小的 alpha。

Q8. 答案:B

p = 0.07 > alpha = 0.05 → 不拒绝 H0。但策略真有 Alpha(H0 实际为假)→ 未能拒绝一个假的 H0 = Type II 错误。

Q9. 答案:C

n 增大 → 标准误减小 → Power 增大。alpha 不受样本量影响,beta 减小。

Q10. 答案:C

P(至少一次 Type I) = 1 - (1 - 0.05)^50 = 1 - 0.077 ≈ 0.923,即约 92%。


十二、本课小结

              ┌─────────────────────────────────┐
              │   假设检验的两种错误              │
              └─────────────────────────────────┘
                             │
          ┌──────────────────┼──────────────────┐
          │                                     │
    Type I 错误(alpha)               Type II 错误(beta)
    H0 真 → 拒绝 H0                     H0 假 → 不拒绝 H0
    「冤枉好人」                         「放过坏人」
    可控(我设 alpha)                   不可直接设,受 n/效应量/sigma 影响
          │                                     │
          │         此消彼长(固定 n)            │
          │◄────────────────────────────────►│
          │                                     │
          └──────────────────┬──────────────────┘
                             │
                    Power = 1 - beta
                 (正确拒绝假 H0 的能力)
                    标准:Power >= 0.80
                             │
                    增大 n 是同时优化二者的唯一方法

核心公式速记

概念 概率 条件
Type I 错误 alpha P(拒绝 H0
Type II 错误 beta P(不拒绝 H0
置信水平 1 - alpha P(不拒绝 H0
Power 1 - beta P(拒绝 H0

十三、下节课预告

L134:影响 Power 的因素与样本量计算

  • 深入 Power 分析的定量计算
  • 在给定 alpha 和 beta 下,如何确定最小需要的样本量
  • Power 曲线与效应的关系
  • CFA 考试中的典型 Power 计算题

📊 Sindy姐说:记住 2x2 矩阵,理解 alpha 和 beta 的 trade-off,知道 Power >= 0.80 是黄金标准,这一课的核心思想就掌握了。真正的难点在于「当下」——你手里拿着一个 p 值,你怎么知道这个决策有没有犯错误?答案是:你不知道,但你知道概率。这才是理性决策的本质。

🔜 下一课 L134 将教你如何在实际研究中计算需要的样本量——这是 CFA 考试的高频考点,也是一个优秀分析师的基本功。


CFA 一级 · 定量方法 · L133 · Type I 与 Type II 错误 2026-08-08 · 中文版

Quantitative Methods — Hypothesis Testing · Session 3


I. Previous Lesson Review: You Now Have the Decision Toolkit

L132 taught us a complete decision process:

Set H0/Ha → Choose alpha → Collect Sample → Compute Test Statistic → Find p-value → Compare p with alpha → Make Decision

But the critical question is:

🎯 Is your decision always correct? No. Every conclusion from hypothesis testing could be wrong. The question is — which type of error will you make?

This lesson deconstructs the illusion of certainty and reveals two types of risk that cannot be simultaneously eliminated.


II. Four Possible Outcomes: The 2x2 Decision Matrix

2.1 Reality x Decision Cross

In hypothesis testing, there are two layers: - Reality layer: Is H0 actually true or false? (We never know the truth) - Decision layer: Do we choose to reject or not reject H0?

The cross of the two yields 4 possible outcomes:

H0 is Actually True H0 is Actually False
Do Not Reject H0 ✅ Correct Decision (1-alpha) ❌ Type II Error (beta)
Reject H0 ❌ Type I Error (alpha) ✅ Correct Decision (1-beta = Power)

2.2 Remember With Two Stories

Story Type I Error Type II Error
🏛️ Court Trial Convict the innocent (false conviction) Acquit the guilty (false acquittal)
🏥 Medical Test False positive (test says sick, but healthy) False negative (test says healthy, but sick)
📈 Investment Decision Adopt a strategy that doesn't work Miss a strategy that actually works
🚨 Fire Alarm False alarm (no fire but alarm rings) Missed alarm (fire but no alarm)

💡 Memory Rule: Type I = False Positive, Type II = False Negative


III. Type I Error (Alpha Error) — False Positive

3.1 Definition

Type I Error = Rejecting H0 when H0 is actually true.

3.2 Probability of Occurrence

P(Type I Error) = alpha (significance level)

This probability is under our direct control. Setting alpha = 0.05 means: "I accept at most a 5% chance of wrongfully rejecting H0."

3.3 Financial Scenario

Scenario: You discovered a stock trading signal.

H0: The signal has no predictive power (mean return = 0)
Ha: The signal has positive predictive power (mean return > 0)
Decision Reality Outcome
Backtest p = 0.03, alpha = 0.05 → Reject H0, adopt the signal But the signal is purely random ❌ Type I Error: You deploy a false signal, trade real money, and lose

🔥 Cost of Type I Error in Finance: Overtrading, capital erosion, overconfidence in backtests.

3.4 Why Alpha IS the Type I Error Rate?

alpha = 0.05 means:
If H0 is true, the test statistic still has a 5% chance of falling into the rejection region.
→ Even if H0 is perfectly correct, on average 5 out of 100 tests will falsely reject H0.

This is why alpha is called the "significance level" — it simultaneously defines the Type I error rate.


IV. Type II Error (Beta Error) — False Negative

4.1 Definition

Type II Error = Failing to reject H0 when H0 is actually false.

4.2 Probability of Occurrence

P(Type II Error) = beta

Beta is not directly set by us; it depends on: - Effect size (the gap between true value and H0 hypothesized value): larger effect → smaller beta - Sample size n: larger n → smaller beta - Alpha level: smaller alpha → larger beta (trade-off) - Population variability sigma: larger sigma → larger beta

4.3 Financial Scenario

Scenario: Continuing the trading signal example.

H0: The signal has no predictive power
Ha: The signal has positive predictive power
Decision Reality Outcome
Backtest p = 0.12, alpha = 0.05 → Do not reject H0, abandon the signal But the signal actually works ❌ Type II Error: You missed a genuinely profitable strategy

🔥 Cost of Type II Error in Finance: Lost Alpha, idle capital, competitor gets there first.


V. The Type I vs Type II Trade-off

5.1 Core Principle

⚖️ With fixed sample size, reducing Type I error rate (smaller alpha) necessarily increases Type II error rate (larger beta).

5.2 Why?

alpha = 0.10 → Large rejection region → Easy to reject H0 → Type I high, Type II low
alpha = 0.05 → Moderate rejection region → Balanced
alpha = 0.01 → Small rejection region → Hard to reject H0 → Type I low, Type II high
alpha Type I Risk Type II Risk When to Use
0.10 High (10% false conviction) Low Don't want to miss any possible signal
0.05 Medium Medium Standard research practice
0.01 Low (1% false conviction) High Cost of false conviction is extreme (medicine)

5.3 The Only Way to Reduce Both Simultaneously

📈 Increasing sample size n can reduce beta without changing alpha (simultaneously increasing Power).

Why? n ↑ → standard error ↓ → distributions become more concentrated → easier to distinguish H0 from Ha → beta ↓


VI. The Power of a Test

6.1 Definition

Power = Probability of correctly rejecting H0 when H0 is actually false.

Power = 1 - beta = P(Reject H0 | H0 is False)

6.2 Interpreting Power

Power Interpretation
0.80 (80%) If an effect truly exists, 80% chance of detecting it
0.50 (50%) Coin-flip detection ability — too weak
0.95 (95%) Almost never misses a true effect — very strong

📌 Industry Standard: Power >= 0.80 (i.e., beta <= 0.20) is considered acceptable.

6.3 Four Factors Affecting Power

Factor Direction Power Change Reason
Sample size n ↑ ↑ SE ↓, distributions more concentrated
Effect size (true deviation) ↑ ↑ Stronger signal, easier to detect
alpha ↑ ↑ Rejection region widens
Population SD sigma ↑ ↓ More noise, harder to distinguish signal

6.4 Visual Illustration of Power

Assume H0: mu = 100, True mu = 105, sigma/sqrt(n) = 2

       H0 Distribution (mu=100)    True Distribution (mu=105)
           ╱ ╲                     ╱ ╲
          ╱   ╲                   ╱   ╲
         ╱     ╲                 ╱     ╲
    ────┴───────┴───────┬───────┴───────┴────
       97  99  101 103  105 107 109 111

            ←── alpha = Type I region
                           ←── Power = 1 - beta
                           ←── beta = Type II region

💡 Greater overlap between distributions → Lower Power; Less overlap → Higher Power.


VII. Case Study: Using Power for Investment Decisions

Background

A quant researcher claims to have discovered a new factor with monthly excess return of 0.30% (i.e., mu > 0). Given: - Sample monthly alpha = 0.25% - Sample standard deviation s = 0.80% - n = 48 months - alpha = 0.05 (right-tailed test) - Assume true effect = 0.30% (researcher's claim)

Step 1: Set Up the Test

H0: mu <= 0
Ha: mu > 0

SE = 0.80% / sqrt(48) ≈ 0.1155%

Critical t_0.05,47 ≈ 1.678
→ Reject H0 when x-bar > 0 + 1.678 * 0.1155% = 0.194%

Step 2: Calculate Power

If true mu = 0.30%, then:

P(Reject H0 | mu = 0.30%)
= P(x-bar > 0.194% | mu = 0.30%)
= P(Z > (0.194% - 0.30%) / 0.1155%)
= P(Z > -0.92)
≈ 0.8212

Power ≈ 82% → If the factor truly generates 0.30% monthly excess return, this backtest has 82% probability of detecting it.

Step 3: What if Only 12 Months of Data?

SE = 0.80% / sqrt(12) ≈ 0.231%

Critical value = 0 + 1.796 * 0.231% = 0.415%

P(Reject H0 | mu = 0.30%)
= P(Z > (0.415% - 0.30%) / 0.231%)
= P(Z > 0.50)
≈ 0.3085

Power ≈ 31% → Only 31% chance of detecting a genuinely effective signal!

🔥 Conclusion: With insufficient sample size, even a truly effective signal is extremely unlikely to be detected! Many "failed backtests" may simply be due to low Power.


VIII. Weighting Alpha vs Beta: Context Determines

8.1 "Better to Overreact" (Control Type II)

Scenario Why prioritize minimizing beta?
Early epidemic screening Better false alarms than missing a case
Anti-money laundering Better to investigate some clean transactions than miss a suspicious one
Early investment signal monitoring Better to screen some false signals than miss a real one

🔑 In these scenarios, the cost of Type II far exceeds Type I.

8.2 "Better to Let Go" (Control Type I)

Scenario Why prioritize minimizing alpha?
Clinical trial approval Better to miss some effective drugs than approve a harmful one
Death penalty sentencing Evidence must be ironclad; better to acquit than convict the innocent
High-frequency trading deployment Better to skip a signal than let a false signal burn capital

🔑 In these scenarios, the cost of Type I far exceeds Type II.

8.3 CFA Professional Judgment Framework

When facing a hypothesis testing problem, a CFA candidate should ask:

  1. If I wrongly reject H0 (Type I), how severe are the consequences?
  2. If I wrongly retain H0 (Type II), how severe are the consequences?
  3. Which outcome is more unacceptable?
  4. Choose alpha level accordingly (smaller alpha protects against Type I, larger alpha protects against Type II)

IX. Type I Error Accumulation in Multiple Testing

9.1 The Multiple Comparisons Problem

Single test Type I error rate = alpha = 0.05

When conducting 20 independent tests:
P(at least one Type I error)
= 1 - P(all correct)
= 1 - (1 - 0.05)^20
= 1 - 0.358
= 0.642

🚨 64.2%! With 20 tests, there is a 64% probability of at least one false positive!

9.2 The Data Mining Trap

"I found 5 significant factors out of 100 tested (p < 0.05)"
→ At alpha = 0.05, 100 random factors would average 5 "significant" results
→ These are very likely all Type I Errors!

🔥 This is why the quant world says: "If you torture the data long enough, it will confess to anything."

9.3 Bonferroni Correction (CFA awareness level)

Adjusted alpha* = alpha / k (k = number of tests)

Example: 10 tests, original alpha = 0.05:
alpha* = 0.05 / 10 = 0.005
→ Only p < 0.005 is considered significant

X. Practice Questions

Part A: Basic Concepts

Q1. A Type I error occurs when:

A. H0 is false and we do not reject H0 B. H0 is true and we reject H0 C. Ha is true and we reject H0 D. H0 is false and we reject H0

Q2. The probability of a Type II error is:

A. alpha B. 1 - alpha C. beta D. 1 - beta

Q3. The power of a test is:

A. The probability of a correct decision when we do not reject H0 B. The probability of a correct decision when we reject H0 C. The probability of correctly rejecting H0 when H0 is false D. The probability of not rejecting H0 when H0 is true


Part B: Trade-off Understanding

Q4. With fixed sample size, reducing alpha from 0.05 to 0.01 will:

A. Reduce both Type I and Type II errors B. Reduce Type I error, increase Type II error C. Increase Type I error, reduce Type II error D. Increase both Type I and Type II errors

Q5. Which of the following can reduce BOTH Type I and Type II errors simultaneously?

A. Increase alpha B. Decrease alpha C. Increase sample size D. Decrease effect size

Q6. Power = 0.90 means:

A. H0 has a 90% probability of being true B. If H0 is false, there is a 90% probability of correctly rejecting H0 C. The Type I error rate is 10% D. The conclusion has a 90% probability of being correct


Part C: Application & Judgment

Q7. In a clinical trial for a new drug, H0: drug is ineffective. Which error is more severe?

A. Type I Error — approving an ineffective drug B. Type II Error — not approving an effective drug C. Both equally severe D. Cannot determine

Q8. A fund manager claims a strategy has Alpha. Backtest yields p = 0.07, alpha = 0.05. We do not reject H0. But the strategy truly has Alpha. This is:

A. Type I Error B. Type II Error C. Correct decision D. Cannot determine

Q9. Increasing sample size will:

A. Increase alpha B. Decrease Power C. Increase Power D. Increase beta

Q10. Conducting 50 independent hypothesis tests (alpha = 0.05), the probability of at least one Type I error is approximately:

A. 5% B. 50% C. 92% D. 100%


XI. Answers and Explanations

Part A: Basic Concepts

Q1. Answer: B

Type I Error = rejecting H0 when H0 is true (false positive). A is Type II Error. C/D describe confused error types.

Q2. Answer: C

P(Type II Error) = beta. alpha is Type I error probability, 1-alpha is probability of correct non-rejection when H0 is true, 1-beta is Power.

Q3. Answer: C

Power = P(Reject H0 | H0 is False) = 1 - beta. D describes 1-alpha (confidence level).


Part B: Trade-off Understanding

Q4. Answer: B

alpha ↓ → rejection region shrinks → harder to reject H0 → Type I ↓, Type II ↑. Classic trade-off.

Q5. Answer: C

Only increasing sample size n can simultaneously reduce both Type I and Type II errors. n ↑ → SE ↓ → distributions more concentrated → improved ability to distinguish H0 from Ha.

Q6. Answer: B

Power = P(Reject H0 | H0 is False). Power = 0.90 = if H0 is truly false, there is a 90% chance of detecting it (correctly rejecting).


Part C: Application & Judgment

Q7. Answer: A

In drug approval, Type I Error (approving ineffective/harmful drug) costs patient harm or death, far more severe than Type II (temporarily missing an effective drug). This is why agencies like the FDA choose extremely small alpha.

Q8. Answer: B

p = 0.07 > alpha = 0.05 → do not reject H0. But the strategy truly has Alpha (H0 is actually false) → failing to reject a false H0 = Type II Error.

Q9. Answer: C

n ↑ → SE ↓ → Power ↑. alpha is unaffected by sample size. beta ↓.

Q10. Answer: C

P(at least one Type I) = 1 - (1 - 0.05)^50 = 1 - 0.077 ≈ 0.923, approximately 92%.


XII. Lesson Summary

              ┌─────────────────────────────────┐
              │  Two Types of Errors in Testing  │
              └─────────────────────────────────┘
                             │
          ┌──────────────────┼──────────────────┐
          │                                     │
    Type I Error (alpha)              Type II Error (beta)
    H0 True → Reject H0               H0 False → Don't Reject H0
    "False Positive"                  "False Negative"
    Controllable (set alpha)           Not directly set; depends on n/effect/sigma
          │                                     │
          │         Trade-off (fixed n)          │
          │◄────────────────────────────────►│
          │                                     │
          └──────────────────┬──────────────────┘
                             │
                    Power = 1 - beta
            (Ability to correctly reject false H0)
                 Standard: Power >= 0.80
                             │
            Increasing n is the only way to improve both

Quick Reference Formulas

Concept Probability Condition
Type I Error alpha P(Reject H0
Type II Error beta P(Don't Reject H0
Confidence Level 1 - alpha P(Don't Reject H0
Power 1 - beta P(Reject H0

XIII. Next Lesson Preview

L134: Factors Affecting Power & Sample Size Calculation

  • In-depth quantitative Power analysis
  • How to determine the minimum required sample size given alpha and beta
  • Power curves and their relationship with effect size
  • Typical Power calculation questions on the CFA exam

📊 Sindy says: Memorize the 2x2 matrix, understand the alpha-beta trade-off, know that Power >= 0.80 is the gold standard — master these and you have grasped the core of this lesson. The real difficulty lies in the "present moment" — holding a p-value in your hand, how do you know if this decision is an error? Answer: you don't, but you know the probabilities. That is the essence of rational decision-making.

🔜 Next up: L134 will teach you how to calculate required sample sizes in real research — a high-frequency CFA exam topic and a fundamental skill for any good analyst.


CFA Level I · Quantitative Methods · L133 · Type I and Type II Errors 2026-08-08 · English Version

🔜 下一课 · L134

CFA 一级 · L134 · 单尾 vs 双尾检验 — 一、上节课回顾:你知道检验会出错 · 二、为什么检验要分「单尾」和「双尾」? · 三、双尾检验(Two-Tailed Test)