Standard II — Integrity of Capital Markets Module 1 · 15-20% Weight Lesson 137

📖 简单线性回归模型

CFA Level 1 · L137 · Simple Linear Regression Model

定量方法(Quantitative Methods)— 简单线性回归 · 第 1 课


一、本课定位

课次 主题 核心能力
L131-L136 假设检验 H0/Ha、p 值、Type I/II、Z/T 检验
L137 简单线性回归模型 理解 Y = b0 + b1X + ε 的结构与含义
L138 最小二乘法(OLS) 估计 b0、b1
L139 R² 与 F 检验 模型整体解释力
L140 回归假设与诊断 检验假设是否成立
L141 定量模块终测 15 题综合测试

🎯 L137 是整个回归模块的基石——先读懂模型在说什么,后面才能算得对、判得准。


二、为什么要学回归?

先看几个真实的金融场景:

场景 X(自变量) Y(因变量)
💰 资产定价 CAPM 市场超额收益 个股超额收益
🏠 房价预测 房屋面积 房屋价格
📈 风险分析 GDP 增长率 股票指数收益率
🏦 信用评估 资产负债率 违约概率(Logit 变体)
💵 汇率预测 利差 汇率变动

回归 = 用一个变量的信息去预测/解释另一个变量。金融最核心的「预测」工具。


三、简单线性回归模型

3.1 基本公式

$$ Y_i = b_0 + b_1 X_i + \varepsilon_i $$

符号 名称 含义
Y_i 因变量(Dependent Variable) 我们要预测/解释的量
X_i 自变量(Independent Variable) 用来预测 Y 的量
b_0 截距(Intercept) X=0 时 Y 的期望值
b_1 斜率(Slope Coefficient) X 每变动 1 单位,Y 的预期变动量
ε_i 误差项(Error Term) 模型无法解释的随机部分

3.2 「简单」vs「多元」

简单线性回归 多元线性回归
自变量个数 1 个 X 多个 X(X1, X2, X3...)
公式 Y = b0 + b1X + ε Y = b0 + b1X1 + b2X2 + ... + ε
图形 一条直线 一个超平面(无法画)
CFA 一级 ✅ 重点 ✅ 后续会学(L2 重点)

3.3 模型的结构——拆解理解

把每一个 Y 值拆成两部分:

$$\text{实际值} = \underbrace{b_0 + b_1X_i}{\text{系统部分(可解释)}} + \underbrace{\varepsilon_i}{\text{随机部分(不可解释)}}$$

系统部分 = 回归线:所有(Xi, Yi)点沿着的理想直线

随机部分 = 误差:每个点到回归线的垂直距离

Y ↑
 |    ·  ← 实际点
 |   /|   ↑ 竖线 = ε_i(残差)
 |  / | 
 | ·——· ← 回归线上的点(拟合值)
 |/
 +──────────→ X

四、回归系数的经济含义

4.1 b₁(斜率)——最关键的系数

b₁ = X 每增加 1 单位,Y 预期变化 b₁ 单位

金融案例:

回归模型 b₁ 的含义
股票收益 ~ 市场收益(CAPM β) 市场涨 1%,该股预期涨 β%
房价 ~ 面积 每多 1 平方米,房价涨 b₁ 万元
销售额 ~ 广告费 每多花 1 万广告费,预期多卖 b₁ 万
利率变动 ~ 通胀率 通胀涨 1%,利率预期变 b₁%

4.2 b₀(截距)——什么也不做时的起点

b₀ = X = 0 时 Y 的预期值

⚠️ 注意:X=0 不一定有意义! - 房价~面积:b₀ = 0平方米的房价 → 没意义,不要做经济解释 - 超额收益~市场超额收益:b₀ = 市场超额为0时的个股超额 → 有实际意义(Alpha!)

4.3 b₀ 与 b₁ 的实际判断

CAPM 回归:R_stock - Rf = b₀ + b₁(R_market - Rf) + ε

系数 名称 如果为正值 如果与 0 无显著差异
b₀ Jensen's Alpha 股票有超额收益能力 没有显著 Alpha
b₁ Beta 顺周期股(大盘涨就跟涨) 与市场无关

五、误差项 ε 的六大经典假设(Gauss-Markov 基础)

这些假设在 L140 会详细展开,L137 先理解基本概念。

编号 假设 含义 违反后果
① E(ε X) = 0 给定 X,误差的条件期望为 0
② 同方差性 Var(ε|X) = σ² 恒定不变 标准误不准,t 检验失效
③ 无自相关 Cov(εi, εj) = 0 (i ≠ j) 标准误偏小,容易拒绝 H0
④ X 非随机/外生 X 与 ε 不相关 估计有偏且不一致
⑤ ε ~ 正态分布 误差正态分布(小样本检验需要) 大样本问题不大(CLT)
⑥ 线性关系 Y 与 X 是线性关系 模型设定错误

🔑 满足①-④ → OLS 是最佳线性无偏估计量(BLUE)


六、回归分析的标准输出——读懂一张回归表

CFA 一级考试经常给回归输出表,你需要会读:

Dependent Variable: Stock_Return
Method: Least Squares
Sample: 60 months

Variable      Coefficient   Std. Error   t-Statistic   Prob.
────────────────────────────────────────────────────────────
Intercept      0.0025        0.0038        0.658        0.513
Mkt_Return     1.2000        0.1500        8.000        0.000
────────────────────────────────────────────────────────────
R-squared      0.5200        Durbin-Watson: 1.95
F-statistic    64.00        Prob(F): 0.000

读表要点:

数据 怎么读
Coefficient b₀=0.0025(月 Alpha≈0.25%),b₁=1.20(Beta=1.2,进攻型)
Std. Error 系数估计的精度(越小越好)
t-Statistic Coeff/Std.Error,检验「系数=0」的统计量
Prob. p 值,<0.05 → 系数显著不为 0
R² 市场收益解释了 52% 的个股收益波动
F-statistic 整体检验 H0: b₁=0,Prob(F)<0.05 → 模型有意义

本课只要求会读表。OLS 如何算出这些数 → L138 见。


七、练习题(10 题)

基础概念(Q1–Q5)

Q1. 简单线性回归模型中,自变量是: A. 被预测的变量 B. 误差项 C. 用于预测的变量 D. 截距

Q2. 回归模型 Y = 3 + 2X + ε 中,斜率系数表示: A. Y 的期望值为 3 B. X 每增加 1 单位,Y 预期增加 2 单位 C. X 每增加 1 单位,Y 一定增加 2 单位 D. 误差的期望值为 3

Q3. 关于截距 b₀,正确的是: A. CAPM 回归中 b₀ 代表 Beta B. b₀ 始终有经济学含义 C. X=0 时 Y 的期望值 D. b₀ 通常等于因变量的均值

Q4. 误差项 ε 代表: A. 回归线预测的值 B. 自变量 X 的测量误差 C. 实际 Y 值与回归线预测值之差 D. 回归系数的标准误

Q5. 简单线性回归中「简单」指: A. 模型容易理解 B. 只有一个自变量 C. 不需要假设检验 D. 线性关系一定成立

应用理解(Q6–Q8)

Q6. 分析师回归:月股票收益 = 0.001 + 1.5 × 月市场收益。这支股票的 Beta 是: A. 0.001 B. 1.0 C. 1.5 D. 无法确定

Q7. 接 Q6,如果市场本月涨 2%,预期该股票涨: A. 0.001% B. 2.0% C. 3.0% D. 3.001%

Q8. 分析师估计 b₁ = -0.8(p=0.03),Ha: b₁ ≠ 0,alpha=0.05。结论: A. b₁ 显著为正 B. b₁ 显著为负 C. b₁ 与 0 无显著差异 D. 样本量不足,不能下结论

综合推理(Q9–Q10)

Q9. 回归输出:b₁=2.5, Std.Error=1.0。检验 H0: b₁=0 的 t 统计量是: A. 0.4 B. 2.5 C. 1.5 D. 3.5

Q10. CAPM 回归中 b₀(Alpha)p=0.42, b₁(Beta)p=0.001。alpha=0.05,正确结论: A. 基金经理有显著选股能力 B. 股票 Beta 显著,但 Alpha 不显著 C. Alpha 和 Beta 都不显著 D. 模型整体无效


八、答案与详解

题号 答案 详解
Q1 C 自变量(X)= 用来预测的变量。A 是因变量 Y。B 是随机误差 ε。D 是参数 b₀。
Q2 B b₁=2 意味着 X 每增 1,Y 预期增 2。C 错在「一定」——有 ε 存在,实际值可能偏离。
Q3 C b₀ = E(Y|X=0)。A 错:CAPM 中 b₀ 是 Alpha(不是 Beta)。B 错:X=0 不在数据范围内时截距无经济意义。
Q4 C ε = Y_actual - Y_predicted = 实际值减拟合值。D 是系数标准差,不是误差项。
Q5 B 「简单」= 只有一个自变量。多元回归有多个 X。A/C/D 都不是统计学术语的正确定义。
Q6 C b₁ 就是 Beta 系数,回归中 β = 1.5。截距 0.001 是 Alpha(月度)。
Q7 D E(Y) = 0.001 + 1.5 × 2% = 0.001 + 3.0% = 3.001%。
Q8 B b₁=-0.8, p=0.03<0.05 → 拒绝 H0: b₁=0 → b₁ 显著不等于 0。因 b₁ 为负 → 显著为负。
Q9 B t = Coefficient / Std.Error = 2.5 / 1.0 = 2.5。这是检验「系数是否为 0」的标准公式。
Q10 B Beta: p=0.001<0.05 → 显著。Alpha: p=0.42>0.05 → 不显著。结论:Beta 有意义,但无证据表明该基金有选股能力(Alpha 不显著)。

九、CFA 一级核心考点

考点 记忆要点
模型结构 Y = b₀ + b₁X + ε,三部分:截距 + 斜率×X + 误差
b₁ 含义 X 每变 1 单位,Y 预期变 b₁ 单位(最关键的数字)
b₀ 含义 X=0 时 Y 的预期值(不一定有经济意义)
ε 本质 模型解释不了的随机噪音
读回归表 Coeff → t = Coeff/SE → p 值 → 判断显著性
CAPM 回归 b₀=Alpha(选股能力),b₁=Beta(市场敏感度)
不做的事 不要用 X=0 不在样本内的截距做经济解释!

📌 下一课 L138:最小二乘法(OLS)—— 如何算出 b₀ 和 b₁?我们要学会让误差平方和最小。

Quantitative Methods — Simple Linear Regression · Lesson 1


I. Where This Lesson Fits

Lesson Topic Core Competency
L131–L136 Hypothesis Testing H0/Ha, p-value, Type I/II errors, Z/T tests
L137 Simple Linear Regression Model Understand the structure and meaning of Y = b₀ + b₁X + ε
L138 Ordinary Least Squares (OLS) Estimating b₀ and b₁
L139 R² and F-Test Overall model explanatory power
L140 Regression Assumptions & Diagnostics Testing whether assumptions hold
L141 Quantitative Methods Final Test 15-question comprehensive exam

🎯 L137 is the foundation of the entire regression module — understand what the model says before we learn to calculate and evaluate.


II. Why Learn Regression?

Real-world finance applications:

Scenario X (Independent Variable) Y (Dependent Variable)
💰 CAPM Asset Pricing Market excess return Individual stock excess return
🏠 Housing Price Prediction Floor area House price
📈 Risk Analysis GDP growth rate Stock index return
🏦 Credit Assessment Debt-to-asset ratio Default probability (Logit variant)
💵 Exchange Rate Forecast Interest rate differential Exchange rate movement

Regression = using information from one variable to predict/explain another. The most essential "prediction" tool in finance.


III. The Simple Linear Regression Model

3.1 Basic Formula

$$ Y_i = b_0 + b_1 X_i + \varepsilon_i $$

Symbol Name Meaning
Yᵢ Dependent Variable The variable we are trying to predict/explain
Xᵢ Independent Variable The variable used to predict Y
b₀ Intercept Expected value of Y when X = 0
b₁ Slope Coefficient Expected change in Y per 1-unit change in X
εᵢ Error Term The random, unexplained portion of the model

3.2 "Simple" vs. "Multiple"

Simple Linear Regression Multiple Linear Regression
Number of X variables 1 X Multiple X (X₁, X₂, X₃...)
Formula Y = b₀ + b₁X + ε Y = b₀ + b₁X₁ + b₂X₂ + ... + ε
Visual A straight line A hyperplane (cannot draw)
CFA Level 1 ✅ Key topic ✅ Covered later (L2 focus)

3.3 Model Structure — Deconstructed

Decompose every Y value into two parts:

$$\text{Actual Value} = \underbrace{b_0 + b_1X_i}{\text{Systematic Part (Explainable)}} + \underbrace{\varepsilon_i}{\text{Random Part (Unexplainable)}}$$

Systematic Part = Regression Line: the ideal straight line that all (Xᵢ, Yᵢ) points follow

Random Part = Error: the vertical distance from each point to the regression line

Y ↑
 |    ·  ← actual point
 |   /|   ↑ vertical line = εᵢ (residual)
 |  / | 
 | ·——· ← point on regression line (fitted value)
 |/
 +──────────→ X

IV. Economic Interpretation of Regression Coefficients

4.1 b₁ (Slope) — The Most Critical Coefficient

b₁ = for each 1-unit increase in X, Y is expected to change by b₁ units

Finance examples:

Regression Model Meaning of b₁
Stock Return ~ Market Return (CAPM β) When the market rises 1%, the stock is expected to rise β%
House Price ~ Floor Area Each additional square meter adds b₁ to the price
Sales ~ Advertising Spend Each additional 10k in ad spend yields b₁ more in expected sales
Interest Rate Change ~ Inflation Rate Each 1% rise in inflation changes the interest rate by b₁%

4.2 b₀ (Intercept) — Starting Point When Nothing Else Matters

b₀ = Expected value of Y when X = 0

⚠️ Caution: X = 0 may not be meaningful! - House Price ~ Area: b₀ = price of a 0 m² house → meaningless, do not interpret economically - Excess Return ~ Market Excess Return: b₀ = stock's excess return when market excess is 0 → has real meaning (Alpha!)

4.3 Practical Judgment of b₀ and b₁

CAPM Regression: R_stock − Rf = b₀ + b₁(R_market − Rf) + ε

Coefficient Name If Positive If Not Significantly Different from 0
b₀ Jensen's Alpha The stock has excess return capability No significant Alpha
b₁ Beta Cyclical stock (moves with the market) Unrelated to market

V. The Six Classical Assumptions of the Error Term ε (Gauss-Markov Foundation)

These assumptions will be explored in detail in L140. In L137, grasp the basic concepts.

# Assumption Meaning Consequence if Violated
① E(ε|X) = 0 Conditional expectation of errors is zero given X Biased intercept
② Homoskedasticity Var(ε|X) = σ² is constant Incorrect standard errors, t-tests invalid
③ No Autocorrelation Cov(εᵢ, εⱼ) = 0 (i ≠ j) Standard errors too small, too easy to reject H₀
④ X is Non-random / Exogenous X is uncorrelated with ε Biased and inconsistent estimates
⑤ ε ~ Normal Distribution Errors are normally distributed (needed for small-sample tests) Less of a problem in large samples (CLT)
⑥ Linear Relationship Y and X have a linear relationship Model misspecification

🔑 Satisfying ①–④ → OLS is the Best Linear Unbiased Estimator (BLUE)


VI. Standard Regression Output — How to Read a Regression Table

The CFA Level 1 exam frequently presents regression output tables. You must know how to read them:

Dependent Variable: Stock_Return
Method: Least Squares
Sample: 60 months

Variable      Coefficient   Std. Error   t-Statistic   Prob.
────────────────────────────────────────────────────────────
Intercept      0.0025        0.0038        0.658        0.513
Mkt_Return     1.2000        0.1500        8.000        0.000
────────────────────────────────────────────────────────────
R-squared      0.5200        Durbin-Watson: 1.95
F-statistic    64.00        Prob(F): 0.000

How to read the table:

Item How to Interpret
Coefficient b₀ = 0.0025 (monthly Alpha ≈ 0.25%), b₁ = 1.20 (Beta = 1.2, aggressive stock)
Std. Error Precision of coefficient estimate (smaller is better)
t-Statistic Coeff / Std. Error — tests "coefficient = 0"
Prob. p-value; if < 0.05 → coefficient is significantly different from 0
R² Market returns explain 52% of the variation in individual stock returns
F-statistic Overall test H₀: b₁ = 0; Prob(F) < 0.05 → the model is meaningful

In this lesson, the goal is only to read the table. How OLS computes these numbers → L138.


VII. Practice Questions (10 Questions)

Basic Concepts (Q1–Q5)

Q1. In a simple linear regression model, the independent variable is: A. The variable being predicted B. The error term C. The variable used for prediction D. The intercept

Q2. In the regression model Y = 3 + 2X + ε, the slope coefficient indicates: A. The expected value of Y is 3 B. For each 1-unit increase in X, Y is expected to increase by 2 units C. For each 1-unit increase in X, Y will definitely increase by 2 units D. The expected value of the error term is 3

Q3. Regarding the intercept b₀, which statement is correct? A. In a CAPM regression, b₀ represents Beta B. b₀ always has an economic interpretation C. It is the expected value of Y when X = 0 D. b₀ is usually equal to the mean of the dependent variable

Q4. The error term ε represents: A. The value predicted by the regression line B. Measurement error in the independent variable X C. The difference between the actual Y value and the regression line prediction D. The standard error of the regression coefficient

Q5. The term "simple" in simple linear regression refers to: A. The model is easy to understand B. There is only one independent variable C. No hypothesis testing is required D. A linear relationship is guaranteed to hold

Applied Understanding (Q6–Q8)

Q6. An analyst runs the regression: Monthly Stock Return = 0.001 + 1.5 × Monthly Market Return. The Beta of this stock is: A. 0.001 B. 1.0 C. 1.5 D. Cannot be determined

Q7. Continuing from Q6, if the market rises 2% this month, the expected return of this stock is: A. 0.001% B. 2.0% C. 3.0% D. 3.001%

Q8. An analyst estimates b₁ = −0.8 (p = 0.03), Ha: b₁ ≠ 0, α = 0.05. The conclusion is: A. b₁ is significantly positive B. b₁ is significantly negative C. b₁ is not significantly different from 0 D. Sample size is insufficient to draw a conclusion

Integrated Reasoning (Q9–Q10)

Q9. Regression output: b₁ = 2.5, Std. Error = 1.0. The t-statistic for testing H₀: b₁ = 0 is: A. 0.4 B. 2.5 C. 1.5 D. 3.5

Q10. In a CAPM regression, b₀ (Alpha) has p = 0.42, and b₁ (Beta) has p = 0.001. With α = 0.05, the correct conclusion is: A. The fund manager has significant stock-picking ability B. The stock's Beta is significant, but Alpha is not significant C. Neither Alpha nor Beta is significant D. The overall model is invalid


VIII. Answers and Explanations

Q Answer Explanation
Q1 C The independent variable (X) = the variable used for prediction. A is the dependent variable Y. B is the random error ε. D is the parameter b₀.
Q2 B b₁ = 2 means that for each 1-unit increase in X, Y is expected to increase by 2. C is wrong because of "definitely" — with ε present, actual values may deviate from expectation.
Q3 C b₀ = E(Y|X = 0). A is wrong: in CAPM, b₀ is Alpha (not Beta). B is wrong: when X = 0 lies outside the data range, the intercept has no economic meaning.
Q4 C ε = Y_actual − Y_predicted = actual value minus fitted value. D is the standard error of coefficients, not the error term.
Q5 B "Simple" = only one independent variable. Multiple regression has multiple X variables. A/C/D are not the correct statistical definitions of the term.
Q6 C b₁ is the Beta coefficient; here β = 1.5. The intercept of 0.001 is the monthly Alpha.
Q7 D E(Y) = 0.001 + 1.5 × 2% = 0.001 + 3.0% = 3.001%.
Q8 B b₁ = −0.8, p = 0.03 < 0.05 → reject H₀: b₁ = 0 → b₁ is significantly different from 0. Since b₁ is negative → significantly negative.
Q9 B t = Coefficient / Std. Error = 2.5 / 1.0 = 2.5. This is the standard formula for testing "coefficient = 0."
Q10 B Beta: p = 0.001 < 0.05 → significant. Alpha: p = 0.42 > 0.05 → not significant. Conclusion: Beta is meaningful, but there is no evidence of stock-picking ability (Alpha not significant).

IX. CFA Level 1 Key Takeaways

Key Point Memory Aid
Model Structure Y = b₀ + b₁X + ε — three parts: intercept + slope × X + error
Meaning of b₁ For each 1-unit change in X, Y is expected to change by b₁ units (the most critical number)
Meaning of b₀ Expected Y when X = 0 (may not have economic meaning)
Nature of ε Random noise that the model cannot explain
Reading a Regression Table Coeff → t = Coeff/SE → p-value → determine significance
CAPM Regression b₀ = Alpha (stock-picking skill), b₁ = Beta (market sensitivity)
What NOT to Do Do not interpret the intercept economically when X = 0 is outside the sample range!

📌 Next lesson L138: Ordinary Least Squares (OLS) — How do we compute b₀ and b₁? We will learn to minimize the sum of squared errors.

🔜 下一课 · L138

最小二乘法(OLS)— 学会计算 b0、b1,让误差平方和最小