Regression Analysis: Simple and Multiple Linear Regression, R² and Adjusted R², Quadratic Forms and Fisher–Cochran, Gauss–Markov, Tests and Confidence Intervals
1. Simple linear regression
The model yᵢ = β₀ + β₁xᵢ + εᵢ with E εᵢ = 0, Var εᵢ = σ², uncorrelated. Least squares minimises Σ(yᵢ − β₀ − β₁xᵢ)² and gives b₁ = Sxy/Sxx, b₀ = ȳ − b₁x̄, with Sxx = Σ(xᵢ − x̄)², Sxy = Σ(xᵢ − x̄)(yᵢ − ȳ), Syy = Σ(yᵢ − ȳ)². The residuals eᵢ = yᵢ − ŷᵢ satisfy Σeᵢ = 0 and Σxᵢeᵢ = 0, the line passes through (x̄, ȳ), SSE = Syy − b₁Sxy, and σ̂² = SSE/(n − 2) is unbiased. In terms of correlation, b₁ = r s_y/s_x and r² = Sxy²/(SxxSyy); the two regression slopes satisfy b_yx · b_xy = r².
| Quantity | Value |
|---|---|
| x̄, ȳ | 3, 4 |
| Sxx, Sxy, Syy | 10, 6, 6 |
| b₁ = Sxy/Sxx | 0.6 |
| b₀ = ȳ − b₁x̄ | 2.2 |
| SSE = Syy − b₁Sxy | 6 − 3.6 = 2.4 |
| σ̂² = SSE/(n − 2) | 0.8 |
| R² = Sxy²/(SxxSyy) | 36/60 = 0.6 |
2. Multiple regression in matrix form
y = Xβ + ε with X of size n × p and full column rank p (p counts the intercept column). Least squares solves the normal equations XᵀXβ̂ = Xᵀy, so β̂ = (XᵀX)⁻¹Xᵀy. With XᵀX = [[5, 10], [10, 30]] and Xᵀy = (20, 50), (XᵀX)⁻¹ = (1/50)[[30, −10], [−10, 5]] and β̂ = (1/50)(100, 50) = (2, 1). The fitted values are ŷ = Hy with the hat matrix H = X(XᵀX)⁻¹Xᵀ: symmetric, idempotent, the orthogonal projection onto the column space of X, with tr H = rank H = p. Residuals e = (I − H)y are orthogonal to every column of X.
- E β̂ = β and Cov β̂ = σ²(XᵀX)⁻¹; σ̂² = eᵀe/(n − p) is unbiased.
- With normal errors, β̂ ~ N_p(β, σ²(XᵀX)⁻¹), (n − p)σ̂²/σ² ~ χ²n−p, and the two are independent.
3. The Gauss–Markov theorem
If E ε = 0 and Cov ε = σ²I (zero-mean, homoscedastic, uncorrelated errors — no normality), then for every estimable cᵀβ the least-squares estimator cᵀβ̂ is the best linear unbiased estimator (BLUE): its variance is the smallest among all estimators that are linear in y and unbiased. The proof writes any other linear unbiased estimator as (cᵀ(XᵀX)⁻¹Xᵀ + dᵀ)y with dᵀX = 0, whose variance exceeds that of cᵀβ̂ by σ²‖d‖².
4. Quadratic forms and the Fisher–Cochran theorem
- Mean: for any random Y with mean μ and dispersion Σ, E[YᵀAY] = tr(AΣ) + μᵀAμ — no normality needed. With Σ = I, μ = (1, 2) and A = [[2, 1], [1, 1]]: tr A + μᵀAμ = 3 + 10 = 13.
- Chi-square: if Y ~ N_n(0, I) and A is symmetric, YᵀAY ~ χ²_r iff A is idempotent of rank r. More generally, for Y ~ N(0, Σ), YᵀAY ~ χ²_r iff AΣ is idempotent of rank r.
- Independence (Craig): for Y ~ N(0, I), YᵀAY and YᵀBY are independent iff AB = 0; a linear form LY and YᵀAY are independent iff LA = 0.
Fisher–Cochran theorem: let Y ~ N_n(0, I) and YᵀY = Σᵢ YᵀAᵢY with Aᵢ symmetric of ranks rᵢ. Then the forms YᵀAᵢY are independent χ²rᵢ if and only if Σrᵢ = n. This is why an ANOVA table works: with normal errors, σ = 1 for simplicity and H₀ (all slopes zero) true, the regression and residual sums of squares are quadratic forms in H − (1/n)J and I − H, of ranks p − 1 and n − p adding to n − 1, so they are independent chi-squares and their ratio is F.
5. R², adjusted R² and the ANOVA F test
SST = Σ(yᵢ − ȳ)² splits as SST = SSR + SSE (for a model with an intercept). R² = SSR/SST = 1 − SSE/SST is the proportion of variation explained, and equals the squared correlation between y and ŷ. With k regressors and n observations, adjusted R² = 1 − (1 − R²)(n − 1)/(n − k − 1). R² never decreases when a regressor is added; adjusted R² falls unless the new regressor’s t-statistic exceeds 1 in absolute value, which is why it is used to compare models of different sizes. With n = 25, k = 4 and R² = 0.8, adjusted R² = 1 − 0.2 × 24/20 = 0.76.
| Source | df | Sum of squares | Mean square |
|---|---|---|---|
| Regression | k | SSR | SSR/k |
| Residual | n − k − 1 | SSE | SSE/(n − k − 1) |
| Total | n − 1 | SST | — |
The overall F test of H₀: β₁ = … = β_k = 0 uses F = [SSR/k]/[SSE/(n − k − 1)] = [R²/k]/[(1 − R²)/(n − k − 1)] ~ F(k, n − k − 1): with n = 25, k = 4, R² = 0.8, F = 0.2/0.01 = 20. With n = 30 and k = 3, the degrees of freedom are 3, 26 and 29.
6. Tests for regression coefficients and confidence intervals
- t test for one coefficient: t = β̂ⱼ/SE(β̂ⱼ), SE² = σ̂²[(XᵀX)⁻¹]ⱼⱼ, compared with tn−p. In simple regression SE(b₁) = σ̂/√Sxx, and t² equals the overall F.
- Confidence interval: β̂ⱼ ± tn−p, α/2 SE(β̂ⱼ). With β̂₁ = 2.5, SE = 0.4 and t₂₀,₀.₀₂₅ = 2.086, the upper limit is 2.5 + 0.834 = 3.33.
- Partial F test for a subset of q coefficients: F = [(SSE_reduced − SSE_full)/q]/[SSE_full/(n − p)] ~ F(q, n − p).
- Mean response vs a new observation at x₀: the interval for E(y | x₀) uses Var(ŷ₀) = σ²x₀ᵀ(XᵀX)⁻¹x₀; the prediction interval for a new y₀ adds σ² for the new error, so it is always wider and does not shrink to zero as n → ∞.
Key takeaways
- b₁ = Sxy/Sxx, b₀ = ȳ − b₁x̄, σ̂² = SSE/(n − 2); residuals sum to zero and are orthogonal to x; b_yx b_xy = r².
- β̂ = (XᵀX)⁻¹Xᵀy, Cov β̂ = σ²(XᵀX)⁻¹; H is a symmetric idempotent projection with trace p.
- Gauss–Markov: zero-mean, homoscedastic, uncorrelated errors make OLS the BLUE — no normality, and nothing about biased or non-linear estimators.
- E[YᵀAY] = tr(AΣ) + μᵀAμ; YᵀAY ~ χ²_r iff A idempotent of rank r (Y ~ N(0, I)); Cochran: ranks adding to n ⇔ independent χ².
- Adjusted R² = 1 − (1 − R²)(n − 1)/(n − k − 1); F = [R²/k]/[(1 − R²)/(n − k − 1)] on (k, n − k − 1) df.
- t = β̂ⱼ/SE on n − p df; CI β̂ⱼ ± t SE; prediction intervals are wider than intervals for the mean response.
Practice questions (15)
Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.
For the data x = 1, 2, 3, 4, 5 and y = 2, 4, 5, 4, 5, the least-squares intercept b₀ of the regression of y on x, correct to one decimal place, is ____.
Numerical answer — type the value.
Show answer
Answer: 2.2
x̄ = 3, ȳ = 4, Sxx = 10, Sxy = (−2)(−2) + (−1)(0) + 0 + (1)(0) + (2)(1) = 6, so b₁ = 0.6 and b₀ = 4 − 0.6 × 3 = 2.2. Reporting ȳ = 4 as the intercept forgets that the line passes through (x̄, ȳ), not (0, ȳ).For the same data (x = 1, 2, 3, 4, 5; y = 2, 4, 5, 4, 5), the coefficient of determination R², correct to one decimal place, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.6
Syy = 4 + 0 + 1 + 0 + 1 = 6, so R² = Sxy²/(SxxSyy) = 36/60 = 0.6; equivalently SSR = b₁Sxy = 3.6 and R² = 3.6/6. The correlation itself is √0.6 = 0.775 — reporting r answers a different question.For the same data (x = 1, 2, 3, 4, 5; y = 2, 4, 5, 4, 5), the unbiased estimate of the error variance σ², correct to one decimal place, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.8
SSE = Syy − b₁Sxy = 6 − 0.6 × 6 = 2.4, and σ̂² = SSE/(n − 2) = 2.4/3 = 0.8. Check with the residuals: fitted values 2.8, 3.4, 4.0, 4.6, 5.2 give residuals −0.8, 0.6, 1.0, −0.6, −0.2 and squares summing to 2.4. Dividing by n gives the biased 0.48.A regression with k = 4 regressors fitted to n = 25 observations has R² = 0.8. The adjusted R², correct to two decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.76
1 − (1 − R²)(n − 1)/(n − k − 1) = 1 − 0.2 × 24/20 = 1 − 0.24 = 0.76. Using n − k in the denominator gives 1 − 0.2 × 24/21 = 0.771.For the same regression (n = 25, k = 4, R² = 0.8), the F statistic for H₀: all four slopes are zero is ____.
Numerical answer — type the value.
Show answer
Answer: 20
F = [R²/k]/[(1 − R²)/(n − k − 1)] = (0.8/4)/(0.2/20) = 0.2/0.01 = 20, on (4, 20) degrees of freedom. Using n − 1 = 24 residual df gives 24.A regression with an intercept and 3 regressors is fitted to 20 observations. The trace of the hat matrix H = X(XᵀX)⁻¹Xᵀ is ____.
Numerical answer — type the value.
Show answer
Answer: 4
tr H = tr[(XᵀX)⁻¹XᵀX] = tr I_p = p = 4 (three slopes plus the intercept). H is 20 × 20, but as a projection its trace is its rank. Answering 3 forgets the intercept column; 20 is the trace of the identity.In the linear model y = Xβ + ε with X of full column rank, which statements are true?
Show answer
Answer: B — If E ε = 0 and Cov ε = σ²I, the least-squares estimator is the best linear unbiased estimator of β; C — A biased estimator of β can have smaller mean squared error than least squares; D — Cov(β̂) = σ²(XᵀX)⁻¹
(A) False: only the first two moments are assumed. (B) The theorem. (C) Gauss–Markov compares only unbiased linear estimators; ridge regression trades a little bias for a large cut in variance when XᵀX is near-singular. (D) β̂ = (XᵀX)⁻¹Xᵀy, so Cov β̂ = (XᵀX)⁻¹Xᵀ(σ²I)X(XᵀX)⁻¹ = σ²(XᵀX)⁻¹.Y ~ N_n(0, I) and A, B are n × n symmetric matrices. Which statements are true?
Show answer
Answer: A — If A is idempotent of rank r, then YᵀAY ~ χ²_r; B — If AB = 0, then YᵀAY and YᵀBY are independent; C — For the hat matrix H of an n × p design of rank p, Yᵀ(I − H)Y ~ χ²_{n−p}
(A) A = PDPᵀ with D having r ones, so YᵀAY is a sum of r squared independent N(0, 1). (B) Craig’s theorem. (C) I − H is symmetric idempotent of rank n − p. (D) False: A = diag(2, 0, …, 0) is PSD, and YᵀAY = 2Y₁² is 2χ²₁, which has mean 2 and variance 8, not a chi-square.A regression coefficient is estimated as β̂₁ = 2.5 with standard error 0.4 on 20 residual degrees of freedom. Using t₂₀,₀.₀₂₅ = 2.086, the upper limit of the 95% confidence interval for β₁, correct to two decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 3.33
2.5 + 2.086 × 0.4 = 2.5 + 0.8344 = 3.3344 → 3.33. Using 1.96 gives 3.28, which treats σ as known; the t quantile belongs to the residual degrees of freedom.Y ~ N₂(μ, I) with μ = (1, 2), and A = [[2, 1], [1, 1]]. E[YᵀAY] is ____.
Numerical answer — type the value.
Show answer
Answer: 13
E[YᵀAY] = tr(AΣ) + μᵀAμ = tr A + μᵀAμ = 3 + (2 × 1 + 2 × 1 × 2 + 1 × 4) = 3 + 10 = 13. Dropping the trace term gives 10 (the value at the mean); dropping μᵀAμ gives 3 (the value for a zero mean).An extra regressor is added to a multiple regression model. Which statement is always true?
Show answer
Answer: A — R² does not decrease
Least squares over a larger column space can only reduce (or keep) the residual sum of squares, so R² = 1 − SSE/SST cannot fall. Adjusted R² charges for the lost degree of freedom and falls when the new regressor’s |t| < 1 — which is the point of using it.For a regression with an intercept and one regressor, XᵀX = [[5, 10], [10, 30]] and Xᵀy = (20, 50). The least-squares estimate of the slope is ____.
Numerical answer — type the value.
Show answer
Answer: 1
det XᵀX = 150 − 100 = 50, (XᵀX)⁻¹ = (1/50)[[30, −10], [−10, 5]], β̂ = (1/50)(600 − 500, −200 + 250) = (2, 1): intercept 2, slope 1. Check the normal equations: 5 × 2 + 10 × 1 = 20 and 10 × 2 + 30 × 1 = 50. Dividing 50 by 30 entry by entry gives 1.67, which is not how a matrix inverse works.The regression coefficient of y on x is 1.6 and that of x on y is 0.4. The correlation coefficient between x and y, correct to one decimal place, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.8
b_yx b_xy = r² = 1.6 × 0.4 = 0.64, so |r| = 0.8, and r takes the common sign of the two slopes: +0.8. The arithmetic mean (1.6 + 0.4)/2 = 1 cannot be a correlation — r is their geometric mean.At a point x₀, compare the 95% confidence interval for the mean response E(y | x₀) with the 95% prediction interval for a new observation y₀. Which is true?
Show answer
Answer: A — The prediction interval is always wider
Var(y₀ − ŷ₀) = σ²[1 + x₀ᵀ(XᵀX)⁻¹x₀] exceeds Var(ŷ₀) = σ²x₀ᵀ(XᵀX)⁻¹x₀ by the variance σ² of the new error. As n grows, the second term vanishes but the 1 does not, so only the confidence interval shrinks to zero width.A multiple regression with an intercept and 3 regressors is fitted to 30 observations. The degrees of freedom for regression, residual and total in its ANOVA table are:
Show answer
Answer: A — 3, 26, 29
Total is n − 1 = 29 (one degree of freedom goes to the mean), regression is k = 3, residual is n − k − 1 = 26, and 3 + 26 = 29 — the Cochran rank condition in action. Counting the intercept as a regression degree of freedom gives 4, which is only right for the uncentred total of 30.