定量方法(Quantitative Methods)— 简单线性回归 · 第 1 课
一、本课定位
| 课次 | 主题 | 核心能力 |
|---|---|---|
| L131-L136 | 假设检验 | H0/Ha、p 值、Type I/II、Z/T 检验 |
| L137 | 简单线性回归模型 | 理解 Y = b0 + b1X + ε 的结构与含义 |
| L138 | 最小二乘法(OLS) | 估计 b0、b1 |
| L139 | R² 与 F 检验 | 模型整体解释力 |
| L140 | 回归假设与诊断 | 检验假设是否成立 |
| L141 | 定量模块终测 | 15 题综合测试 |
🎯 L137 是整个回归模块的基石——先读懂模型在说什么,后面才能算得对、判得准。
二、为什么要学回归?
先看几个真实的金融场景:
| 场景 | X(自变量) | Y(因变量) |
|---|---|---|
| 💰 资产定价 CAPM | 市场超额收益 | 个股超额收益 |
| 🏠 房价预测 | 房屋面积 | 房屋价格 |
| 📈 风险分析 | GDP 增长率 | 股票指数收益率 |
| 🏦 信用评估 | 资产负债率 | 违约概率(Logit 变体) |
| 💵 汇率预测 | 利差 | 汇率变动 |
回归 = 用一个变量的信息去预测/解释另一个变量。金融最核心的「预测」工具。
三、简单线性回归模型
3.1 基本公式
$$ Y_i = b_0 + b_1 X_i + \varepsilon_i $$
| 符号 | 名称 | 含义 |
|---|---|---|
| Y_i | 因变量(Dependent Variable) | 我们要预测/解释的量 |
| X_i | 自变量(Independent Variable) | 用来预测 Y 的量 |
| b_0 | 截距(Intercept) | X=0 时 Y 的期望值 |
| b_1 | 斜率(Slope Coefficient) | X 每变动 1 单位,Y 的预期变动量 |
| ε_i | 误差项(Error Term) | 模型无法解释的随机部分 |
3.2 「简单」vs「多元」
| 简单线性回归 | 多元线性回归 | |
|---|---|---|
| 自变量个数 | 1 个 X | 多个 X(X1, X2, X3...) |
| 公式 | Y = b0 + b1X + ε | Y = b0 + b1X1 + b2X2 + ... + ε |
| 图形 | 一条直线 | 一个超平面(无法画) |
| CFA 一级 | ✅ 重点 | ✅ 后续会学(L2 重点) |
3.3 模型的结构——拆解理解
把每一个 Y 值拆成两部分:
$$\text{实际值} = \underbrace{b_0 + b_1X_i}{\text{系统部分(可解释)}} + \underbrace{\varepsilon_i}{\text{随机部分(不可解释)}}$$
系统部分 = 回归线:所有(Xi, Yi)点沿着的理想直线
随机部分 = 误差:每个点到回归线的垂直距离
Y ↑
| · ← 实际点
| /| ↑ 竖线 = ε_i(残差)
| / |
| ·——· ← 回归线上的点(拟合值)
|/
+──────────→ X
四、回归系数的经济含义
4.1 b₁(斜率)——最关键的系数
b₁ = X 每增加 1 单位,Y 预期变化 b₁ 单位
金融案例:
| 回归模型 | b₁ 的含义 |
|---|---|
| 股票收益 ~ 市场收益(CAPM β) | 市场涨 1%,该股预期涨 β% |
| 房价 ~ 面积 | 每多 1 平方米,房价涨 b₁ 万元 |
| 销售额 ~ 广告费 | 每多花 1 万广告费,预期多卖 b₁ 万 |
| 利率变动 ~ 通胀率 | 通胀涨 1%,利率预期变 b₁% |
4.2 b₀(截距)——什么也不做时的起点
b₀ = X = 0 时 Y 的预期值
⚠️ 注意:X=0 不一定有意义! - 房价~面积:b₀ = 0平方米的房价 → 没意义,不要做经济解释 - 超额收益~市场超额收益:b₀ = 市场超额为0时的个股超额 → 有实际意义(Alpha!)
4.3 b₀ 与 b₁ 的实际判断
CAPM 回归:R_stock - Rf = b₀ + b₁(R_market - Rf) + ε
| 系数 | 名称 | 如果为正值 | 如果与 0 无显著差异 |
|---|---|---|---|
| b₀ | Jensen's Alpha | 股票有超额收益能力 | 没有显著 Alpha |
| b₁ | Beta | 顺周期股(大盘涨就跟涨) | 与市场无关 |
五、误差项 ε 的六大经典假设(Gauss-Markov 基础)
这些假设在 L140 会详细展开,L137 先理解基本概念。
| 编号 | 假设 | 含义 | 违反后果 |
|---|---|---|---|
| ① | E(ε | X) = 0 | 给定 X,误差的条件期望为 0 |
| ② | 同方差性 | Var(ε|X) = σ² 恒定不变 | 标准误不准,t 检验失效 |
| ③ | 无自相关 | Cov(εi, εj) = 0 (i ≠ j) | 标准误偏小,容易拒绝 H0 |
| ④ | X 非随机/外生 | X 与 ε 不相关 | 估计有偏且不一致 |
| ⑤ | ε ~ 正态分布 | 误差正态分布(小样本检验需要) | 大样本问题不大(CLT) |
| ⑥ | 线性关系 | Y 与 X 是线性关系 | 模型设定错误 |
🔑 满足①-④ → OLS 是最佳线性无偏估计量(BLUE)
六、回归分析的标准输出——读懂一张回归表
CFA 一级考试经常给回归输出表,你需要会读:
Dependent Variable: Stock_Return
Method: Least Squares
Sample: 60 months
Variable Coefficient Std. Error t-Statistic Prob.
────────────────────────────────────────────────────────────
Intercept 0.0025 0.0038 0.658 0.513
Mkt_Return 1.2000 0.1500 8.000 0.000
────────────────────────────────────────────────────────────
R-squared 0.5200 Durbin-Watson: 1.95
F-statistic 64.00 Prob(F): 0.000
读表要点:
| 数据 | 怎么读 |
|---|---|
| Coefficient | b₀=0.0025(月 Alpha≈0.25%),b₁=1.20(Beta=1.2,进攻型) |
| Std. Error | 系数估计的精度(越小越好) |
| t-Statistic | Coeff/Std.Error,检验「系数=0」的统计量 |
| Prob. | p 值,<0.05 → 系数显著不为 0 |
| R² | 市场收益解释了 52% 的个股收益波动 |
| F-statistic | 整体检验 H0: b₁=0,Prob(F)<0.05 → 模型有意义 |
本课只要求会读表。OLS 如何算出这些数 → L138 见。
七、练习题(10 题)
基础概念(Q1–Q5)
Q1. 简单线性回归模型中,自变量是: A. 被预测的变量 B. 误差项 C. 用于预测的变量 D. 截距
Q2. 回归模型 Y = 3 + 2X + ε 中,斜率系数表示: A. Y 的期望值为 3 B. X 每增加 1 单位,Y 预期增加 2 单位 C. X 每增加 1 单位,Y 一定增加 2 单位 D. 误差的期望值为 3
Q3. 关于截距 b₀,正确的是: A. CAPM 回归中 b₀ 代表 Beta B. b₀ 始终有经济学含义 C. X=0 时 Y 的期望值 D. b₀ 通常等于因变量的均值
Q4. 误差项 ε 代表: A. 回归线预测的值 B. 自变量 X 的测量误差 C. 实际 Y 值与回归线预测值之差 D. 回归系数的标准误
Q5. 简单线性回归中「简单」指: A. 模型容易理解 B. 只有一个自变量 C. 不需要假设检验 D. 线性关系一定成立
应用理解(Q6–Q8)
Q6. 分析师回归:月股票收益 = 0.001 + 1.5 × 月市场收益。这支股票的 Beta 是: A. 0.001 B. 1.0 C. 1.5 D. 无法确定
Q7. 接 Q6,如果市场本月涨 2%,预期该股票涨: A. 0.001% B. 2.0% C. 3.0% D. 3.001%
Q8. 分析师估计 b₁ = -0.8(p=0.03),Ha: b₁ ≠ 0,alpha=0.05。结论: A. b₁ 显著为正 B. b₁ 显著为负 C. b₁ 与 0 无显著差异 D. 样本量不足,不能下结论
综合推理(Q9–Q10)
Q9. 回归输出:b₁=2.5, Std.Error=1.0。检验 H0: b₁=0 的 t 统计量是: A. 0.4 B. 2.5 C. 1.5 D. 3.5
Q10. CAPM 回归中 b₀(Alpha)p=0.42, b₁(Beta)p=0.001。alpha=0.05,正确结论: A. 基金经理有显著选股能力 B. 股票 Beta 显著,但 Alpha 不显著 C. Alpha 和 Beta 都不显著 D. 模型整体无效
八、答案与详解
| 题号 | 答案 | 详解 |
|---|---|---|
| Q1 | C | 自变量(X)= 用来预测的变量。A 是因变量 Y。B 是随机误差 ε。D 是参数 b₀。 |
| Q2 | B | b₁=2 意味着 X 每增 1,Y 预期增 2。C 错在「一定」——有 ε 存在,实际值可能偏离。 |
| Q3 | C | b₀ = E(Y|X=0)。A 错:CAPM 中 b₀ 是 Alpha(不是 Beta)。B 错:X=0 不在数据范围内时截距无经济意义。 |
| Q4 | C | ε = Y_actual - Y_predicted = 实际值减拟合值。D 是系数标准差,不是误差项。 |
| Q5 | B | 「简单」= 只有一个自变量。多元回归有多个 X。A/C/D 都不是统计学术语的正确定义。 |
| Q6 | C | b₁ 就是 Beta 系数,回归中 β = 1.5。截距 0.001 是 Alpha(月度)。 |
| Q7 | D | E(Y) = 0.001 + 1.5 × 2% = 0.001 + 3.0% = 3.001%。 |
| Q8 | B | b₁=-0.8, p=0.03<0.05 → 拒绝 H0: b₁=0 → b₁ 显著不等于 0。因 b₁ 为负 → 显著为负。 |
| Q9 | B | t = Coefficient / Std.Error = 2.5 / 1.0 = 2.5。这是检验「系数是否为 0」的标准公式。 |
| Q10 | B | Beta: p=0.001<0.05 → 显著。Alpha: p=0.42>0.05 → 不显著。结论:Beta 有意义,但无证据表明该基金有选股能力(Alpha 不显著)。 |
九、CFA 一级核心考点
| 考点 | 记忆要点 |
|---|---|
| 模型结构 | Y = b₀ + b₁X + ε,三部分:截距 + 斜率×X + 误差 |
| b₁ 含义 | X 每变 1 单位,Y 预期变 b₁ 单位(最关键的数字) |
| b₀ 含义 | X=0 时 Y 的预期值(不一定有经济意义) |
| ε 本质 | 模型解释不了的随机噪音 |
| 读回归表 | Coeff → t = Coeff/SE → p 值 → 判断显著性 |
| CAPM 回归 | b₀=Alpha(选股能力),b₁=Beta(市场敏感度) |
| 不做的事 | 不要用 X=0 不在样本内的截距做经济解释! |
📌 下一课 L138:最小二乘法(OLS)—— 如何算出 b₀ 和 b₁?我们要学会让误差平方和最小。
Quantitative Methods — Simple Linear Regression · Lesson 1
I. Where This Lesson Fits
| Lesson | Topic | Core Competency |
|---|---|---|
| L131–L136 | Hypothesis Testing | H0/Ha, p-value, Type I/II errors, Z/T tests |
| L137 | Simple Linear Regression Model | Understand the structure and meaning of Y = b₀ + b₁X + ε |
| L138 | Ordinary Least Squares (OLS) | Estimating b₀ and b₁ |
| L139 | R² and F-Test | Overall model explanatory power |
| L140 | Regression Assumptions & Diagnostics | Testing whether assumptions hold |
| L141 | Quantitative Methods Final Test | 15-question comprehensive exam |
🎯 L137 is the foundation of the entire regression module — understand what the model says before we learn to calculate and evaluate.
II. Why Learn Regression?
Real-world finance applications:
| Scenario | X (Independent Variable) | Y (Dependent Variable) |
|---|---|---|
| 💰 CAPM Asset Pricing | Market excess return | Individual stock excess return |
| 🏠 Housing Price Prediction | Floor area | House price |
| 📈 Risk Analysis | GDP growth rate | Stock index return |
| 🏦 Credit Assessment | Debt-to-asset ratio | Default probability (Logit variant) |
| 💵 Exchange Rate Forecast | Interest rate differential | Exchange rate movement |
Regression = using information from one variable to predict/explain another. The most essential "prediction" tool in finance.
III. The Simple Linear Regression Model
3.1 Basic Formula
$$ Y_i = b_0 + b_1 X_i + \varepsilon_i $$
| Symbol | Name | Meaning |
|---|---|---|
| Yᵢ | Dependent Variable | The variable we are trying to predict/explain |
| Xᵢ | Independent Variable | The variable used to predict Y |
| b₀ | Intercept | Expected value of Y when X = 0 |
| b₁ | Slope Coefficient | Expected change in Y per 1-unit change in X |
| εᵢ | Error Term | The random, unexplained portion of the model |
3.2 "Simple" vs. "Multiple"
| Simple Linear Regression | Multiple Linear Regression | |
|---|---|---|
| Number of X variables | 1 X | Multiple X (X₁, X₂, X₃...) |
| Formula | Y = b₀ + b₁X + ε | Y = b₀ + b₁X₁ + b₂X₂ + ... + ε |
| Visual | A straight line | A hyperplane (cannot draw) |
| CFA Level 1 | ✅ Key topic | ✅ Covered later (L2 focus) |
3.3 Model Structure — Deconstructed
Decompose every Y value into two parts:
$$\text{Actual Value} = \underbrace{b_0 + b_1X_i}{\text{Systematic Part (Explainable)}} + \underbrace{\varepsilon_i}{\text{Random Part (Unexplainable)}}$$
Systematic Part = Regression Line: the ideal straight line that all (Xᵢ, Yᵢ) points follow
Random Part = Error: the vertical distance from each point to the regression line
Y ↑
| · ← actual point
| /| ↑ vertical line = εᵢ (residual)
| / |
| ·——· ← point on regression line (fitted value)
|/
+──────────→ X
IV. Economic Interpretation of Regression Coefficients
4.1 b₁ (Slope) — The Most Critical Coefficient
b₁ = for each 1-unit increase in X, Y is expected to change by b₁ units
Finance examples:
| Regression Model | Meaning of b₁ |
|---|---|
| Stock Return ~ Market Return (CAPM β) | When the market rises 1%, the stock is expected to rise β% |
| House Price ~ Floor Area | Each additional square meter adds b₁ to the price |
| Sales ~ Advertising Spend | Each additional 10k in ad spend yields b₁ more in expected sales |
| Interest Rate Change ~ Inflation Rate | Each 1% rise in inflation changes the interest rate by b₁% |
4.2 b₀ (Intercept) — Starting Point When Nothing Else Matters
b₀ = Expected value of Y when X = 0
⚠️ Caution: X = 0 may not be meaningful! - House Price ~ Area: b₀ = price of a 0 m² house → meaningless, do not interpret economically - Excess Return ~ Market Excess Return: b₀ = stock's excess return when market excess is 0 → has real meaning (Alpha!)
4.3 Practical Judgment of b₀ and b₁
CAPM Regression: R_stock − Rf = b₀ + b₁(R_market − Rf) + ε
| Coefficient | Name | If Positive | If Not Significantly Different from 0 |
|---|---|---|---|
| b₀ | Jensen's Alpha | The stock has excess return capability | No significant Alpha |
| b₁ | Beta | Cyclical stock (moves with the market) | Unrelated to market |
V. The Six Classical Assumptions of the Error Term ε (Gauss-Markov Foundation)
These assumptions will be explored in detail in L140. In L137, grasp the basic concepts.
| # | Assumption | Meaning | Consequence if Violated |
|---|---|---|---|
| ① | E(ε|X) = 0 | Conditional expectation of errors is zero given X | Biased intercept |
| ② | Homoskedasticity | Var(ε|X) = σ² is constant | Incorrect standard errors, t-tests invalid |
| ③ | No Autocorrelation | Cov(εᵢ, εⱼ) = 0 (i ≠ j) | Standard errors too small, too easy to reject H₀ |
| ④ | X is Non-random / Exogenous | X is uncorrelated with ε | Biased and inconsistent estimates |
| ⑤ | ε ~ Normal Distribution | Errors are normally distributed (needed for small-sample tests) | Less of a problem in large samples (CLT) |
| ⑥ | Linear Relationship | Y and X have a linear relationship | Model misspecification |
🔑 Satisfying ①–④ → OLS is the Best Linear Unbiased Estimator (BLUE)
VI. Standard Regression Output — How to Read a Regression Table
The CFA Level 1 exam frequently presents regression output tables. You must know how to read them:
Dependent Variable: Stock_Return
Method: Least Squares
Sample: 60 months
Variable Coefficient Std. Error t-Statistic Prob.
────────────────────────────────────────────────────────────
Intercept 0.0025 0.0038 0.658 0.513
Mkt_Return 1.2000 0.1500 8.000 0.000
────────────────────────────────────────────────────────────
R-squared 0.5200 Durbin-Watson: 1.95
F-statistic 64.00 Prob(F): 0.000
How to read the table:
| Item | How to Interpret |
|---|---|
| Coefficient | b₀ = 0.0025 (monthly Alpha ≈ 0.25%), b₁ = 1.20 (Beta = 1.2, aggressive stock) |
| Std. Error | Precision of coefficient estimate (smaller is better) |
| t-Statistic | Coeff / Std. Error — tests "coefficient = 0" |
| Prob. | p-value; if < 0.05 → coefficient is significantly different from 0 |
| R² | Market returns explain 52% of the variation in individual stock returns |
| F-statistic | Overall test H₀: b₁ = 0; Prob(F) < 0.05 → the model is meaningful |
In this lesson, the goal is only to read the table. How OLS computes these numbers → L138.
VII. Practice Questions (10 Questions)
Basic Concepts (Q1–Q5)
Q1. In a simple linear regression model, the independent variable is: A. The variable being predicted B. The error term C. The variable used for prediction D. The intercept
Q2. In the regression model Y = 3 + 2X + ε, the slope coefficient indicates: A. The expected value of Y is 3 B. For each 1-unit increase in X, Y is expected to increase by 2 units C. For each 1-unit increase in X, Y will definitely increase by 2 units D. The expected value of the error term is 3
Q3. Regarding the intercept b₀, which statement is correct? A. In a CAPM regression, b₀ represents Beta B. b₀ always has an economic interpretation C. It is the expected value of Y when X = 0 D. b₀ is usually equal to the mean of the dependent variable
Q4. The error term ε represents: A. The value predicted by the regression line B. Measurement error in the independent variable X C. The difference between the actual Y value and the regression line prediction D. The standard error of the regression coefficient
Q5. The term "simple" in simple linear regression refers to: A. The model is easy to understand B. There is only one independent variable C. No hypothesis testing is required D. A linear relationship is guaranteed to hold
Applied Understanding (Q6–Q8)
Q6. An analyst runs the regression: Monthly Stock Return = 0.001 + 1.5 × Monthly Market Return. The Beta of this stock is: A. 0.001 B. 1.0 C. 1.5 D. Cannot be determined
Q7. Continuing from Q6, if the market rises 2% this month, the expected return of this stock is: A. 0.001% B. 2.0% C. 3.0% D. 3.001%
Q8. An analyst estimates b₁ = −0.8 (p = 0.03), Ha: b₁ ≠ 0, α = 0.05. The conclusion is: A. b₁ is significantly positive B. b₁ is significantly negative C. b₁ is not significantly different from 0 D. Sample size is insufficient to draw a conclusion
Integrated Reasoning (Q9–Q10)
Q9. Regression output: b₁ = 2.5, Std. Error = 1.0. The t-statistic for testing H₀: b₁ = 0 is: A. 0.4 B. 2.5 C. 1.5 D. 3.5
Q10. In a CAPM regression, b₀ (Alpha) has p = 0.42, and b₁ (Beta) has p = 0.001. With α = 0.05, the correct conclusion is: A. The fund manager has significant stock-picking ability B. The stock's Beta is significant, but Alpha is not significant C. Neither Alpha nor Beta is significant D. The overall model is invalid
VIII. Answers and Explanations
| Q | Answer | Explanation |
|---|---|---|
| Q1 | C | The independent variable (X) = the variable used for prediction. A is the dependent variable Y. B is the random error ε. D is the parameter b₀. |
| Q2 | B | b₁ = 2 means that for each 1-unit increase in X, Y is expected to increase by 2. C is wrong because of "definitely" — with ε present, actual values may deviate from expectation. |
| Q3 | C | b₀ = E(Y|X = 0). A is wrong: in CAPM, b₀ is Alpha (not Beta). B is wrong: when X = 0 lies outside the data range, the intercept has no economic meaning. |
| Q4 | C | ε = Y_actual − Y_predicted = actual value minus fitted value. D is the standard error of coefficients, not the error term. |
| Q5 | B | "Simple" = only one independent variable. Multiple regression has multiple X variables. A/C/D are not the correct statistical definitions of the term. |
| Q6 | C | b₁ is the Beta coefficient; here β = 1.5. The intercept of 0.001 is the monthly Alpha. |
| Q7 | D | E(Y) = 0.001 + 1.5 × 2% = 0.001 + 3.0% = 3.001%. |
| Q8 | B | b₁ = −0.8, p = 0.03 < 0.05 → reject H₀: b₁ = 0 → b₁ is significantly different from 0. Since b₁ is negative → significantly negative. |
| Q9 | B | t = Coefficient / Std. Error = 2.5 / 1.0 = 2.5. This is the standard formula for testing "coefficient = 0." |
| Q10 | B | Beta: p = 0.001 < 0.05 → significant. Alpha: p = 0.42 > 0.05 → not significant. Conclusion: Beta is meaningful, but there is no evidence of stock-picking ability (Alpha not significant). |
IX. CFA Level 1 Key Takeaways
| Key Point | Memory Aid |
|---|---|
| Model Structure | Y = b₀ + b₁X + ε — three parts: intercept + slope × X + error |
| Meaning of b₁ | For each 1-unit change in X, Y is expected to change by b₁ units (the most critical number) |
| Meaning of b₀ | Expected Y when X = 0 (may not have economic meaning) |
| Nature of ε | Random noise that the model cannot explain |
| Reading a Regression Table | Coeff → t = Coeff/SE → p-value → determine significance |
| CAPM Regression | b₀ = Alpha (stock-picking skill), b₁ = Beta (market sensitivity) |
| What NOT to Do | Do not interpret the intercept economically when X = 0 is outside the sample range! |
📌 Next lesson L138: Ordinary Least Squares (OLS) — How do we compute b₀ and b₁? We will learn to minimize the sum of squared errors.