Engineering Mathematics II: Regression, Correlation, Statistical Significance and the Chi-Square Test
1. Regression analysis: the least-squares line
The regression of y on x fits y = a + bx by minimising the sum of squared vertical deviations Σ(yᵢ − a − bxᵢ)². With Sxx = Σ(x − x̄)², Syy = Σ(y − ȳ)² and Sxy = Σ(x − x̄)(y − ȳ), the normal equations give b = Sxy/Sxx and a = ȳ − b x̄, so the line always passes through (x̄, ȳ). The regression of x on y, x = a′ + b′y, minimises horizontal deviations and has b′ = Sxy/Syy; the two lines are different unless the points lie exactly on a line. The coefficient of determination R² = r² is the fraction of the variance of y the line explains.
| x | y | x − x̄ | y − ȳ | product |
|---|---|---|---|---|
| 1 | 2 | −2 | −2 | 4 |
| 2 | 4 | −1 | 0 | 0 |
| 3 | 5 | 0 | 1 | 0 |
| 4 | 4 | 1 | 0 | 0 |
| 5 | 5 | 2 | 1 | 2 |
Here x̄ = 3, ȳ = 4, Sxx = 10, Syy = 6 and Sxy = 6. So b = 6/10 = 0.6, a = 4 − 0.6 × 3 = 2.2, and the line is y = 2.2 + 0.6x; the regression of x on y has b′ = 6/6 = 1.0. The residual standard error, √(Σe²/(n − 2)), uses n − 2 because two parameters were estimated.
2. The correlation coefficient
Pearson's correlation coefficient r = Sxy/√(Sxx·Syy) = cov(x, y)/(σₓσ_y) measures the strength of the linear relation, −1 ≤ r ≤ 1. For the data above, r = 6/√(10 × 6) = 6/√60 = 0.775. Three identities are examined often: r² = b·b′ (the product of the two regression slopes: 0.6 × 1.0 = 0.6, so r = 0.775); r has the sign of both slopes, which always share a sign; and r is unchanged by a change of origin or scale of either variable (a positive scale factor), so correlating DNs or radiances gives the same r.
- r = 0 means no linear relation; y = x² over a symmetric range has r = 0 and a perfect relation.
- Correlation is not causation: two bands can correlate because both respond to illumination.
- Spearman's rank correlation rₛ = 1 − 6Σd²/(n(n² − 1)) uses ranks, for monotonic but non-linear relations.
3. Statistical significance: levels, critical values and the two errors
A significance test states a null hypothesis H₀ (no difference, no correlation, the model fits), computes a test statistic whose distribution under H₀ is known, and rejects H₀ if the statistic falls beyond the critical value at the chosen level of significance α (commonly 5 % or 1 %). The p-value is the probability, under H₀, of a statistic at least as extreme as the one observed; reject when p < α. Rejecting a true H₀ is a Type I error (probability α); accepting a false H₀ is a Type II error (probability β; the power of the test is 1 − β).
| Confidence | z | Use |
|---|---|---|
| 50 % | 0.6745 | the probable error |
| 95 % | 1.96 | the usual 5 % test |
| 99 % | 2.576 | the 1 % test |
| 99.73 % | 3 | the surveyor's blunder rejection limit |
Testing a correlation. Under H₀: ρ = 0, t = r√(n − 2)/√(1 − r²) follows Student's t with n − 2 degrees of freedom. For r = 0.6 from n = 12 pairs, t = 0.6 × √10/√0.64 = 0.6 × 3.162/0.8 = 2.37, which exceeds the two-tailed 5 % value of t with 10 degrees of freedom (2.228), so the correlation is significant. Testing a mean against a known σ: z = (x̄ − μ₀)/(σ/√n); with an estimated s and small n, use t with n − 1 degrees of freedom.
4. The chi-square test
The chi-square statistic χ² = Σ(O − E)²/E compares observed frequencies O with those expected under H₀. It is always ≥ 0 and grows with disagreement; H₀ is rejected when χ² exceeds the critical value for the right degrees of freedom. Every expected count should be at least about 5 for the approximation to hold. Upper 5 % critical values: 3.841 (1 degree of freedom), 5.991 (2), 7.815 (3), 9.488 (4).
| Use | Statistic | Degrees of freedom |
|---|---|---|
| Goodness of fit to a distribution with k classes | Σ(O − E)²/E | k − 1 − (number of parameters estimated from the data) |
| Independence in an r × c contingency table | Σ(O − E)²/E with E = row total × column total/N | (r − 1)(c − 1) |
| Global test of an adjustment's variance factor | vᵀWv/σ₀² (a-priori σ₀²) | redundancy r = n − u; two-tailed |
Worked: 200 checked pixels are expected to fall equally into four land-cover classes (50 each) but are observed as 60, 45, 40 and 55. χ² = (10² + 5² + 10² + 5²)/50 = 250/50 = 5.0 with 4 − 1 = 3 degrees of freedom. Since 5.0 < 7.815, the departure from equal shares is not significant at 5 %.
Key takeaways
- Regression of y on x: b = Sxy/Sxx, a = ȳ − b x̄; the line passes through the means, and the regression of x on y is a different line.
- r = Sxy/√(SxxSyy), r² = b·b′, and r is unchanged by linear rescaling with a positive factor.
- A test rejects H₀ beyond the critical value at level α; Type I error has probability α, Type II has β.
- The significance of r uses t = r√(n − 2)/√(1 − r²) with n − 2 degrees of freedom.
- χ² = Σ(O − E)²/E with k − 1 (minus estimated parameters) or (r − 1)(c − 1) degrees of freedom; in an adjustment, vᵀWv/σ₀² with n − u.
Practice questions (14)
Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.
For the data (x, y) = (1, 2), (2, 4), (3, 5), (4, 4), (5, 5), the slope b of the least-squares regression line of y on x is ____.
Numerical answer — type the value.
Show answer
Answer: 0.6
x̄ = 3, ȳ = 4. Sxy = (−2)(−2) + (−1)(0) + 0(1) + 1(0) + 2(1) = 6; Sxx = 4 + 1 + 0 + 1 + 4 = 10. b = Sxy/Sxx = 6/10 = 0.6. Using Syy (= 6) as the divisor gives 1.0, which is the slope of x on y.For the same data (1, 2), (2, 4), (3, 5), (4, 4), (5, 5), the intercept a of the regression line y = a + bx is ____.
Numerical answer — type the value.
Show answer
Answer: 2.2
The regression line passes through (x̄, ȳ) = (3, 4). With b = 0.6, a = ȳ − b x̄ = 4 − 0.6 × 3 = 2.2. Check at x = 5: 2.2 + 3.0 = 5.2, close to the observed 5.For the data (1, 2), (2, 4), (3, 5), (4, 4), (5, 5), Pearson's correlation coefficient r (to three decimal places) is ____.
Numerical answer — type the value.
Show answer
Answer: 0.775
Sxy = 6, Sxx = 10, Syy = (−2)² + 0² + 1² + 0² + 1² = 6. r = 6/√(10 × 6) = 6/7.746 = 0.775. Check through the slopes: b = 0.6, b′ = 1.0, r² = 0.6, r = √0.6 = 0.775. Reporting r² = 0.6 as r is the usual slip.The regression coefficient of y on x is 0.8 and that of x on y is 0.45. The correlation coefficient between x and y is
Show answer
Answer: A — 0.6
r² = b·b′ = 0.8 × 0.45 = 0.36, so r = +0.6, positive because both coefficients are positive. 0.36 is r² itself; 1.25 cannot be a correlation at all.Every observation of x is converted from metres to feet and 100 is added to every value of y. The correlation coefficient between x and y
Show answer
Answer: A — is unchanged
r is a ratio of the covariance to the product of the standard deviations; a positive scale factor multiplies both numerator and denominator alike, and a shift of origin changes neither. So r is invariant — which is why correlating raw DNs or calibrated radiances gives the same value.A sample of 12 pairs gives a correlation coefficient r = 0.6. The t statistic for testing H₀: ρ = 0 (to two decimal places) is ____.
Numerical answer — type the value.
Show answer
Answer: 2.37
t = r√(n − 2)/√(1 − r²) = 0.6 × √10/√(1 − 0.36) = 0.6 × 3.1623/0.8 = 2.37, with 10 degrees of freedom. It exceeds the two-tailed 5 % critical value 2.228, so the correlation is significant. Using √n instead of √(n − 2) gives 2.60.Rejecting a null hypothesis that is in fact true is called
Show answer
Answer: A — a Type I error
A Type I error is a false rejection and its probability is the significance level α. A Type II error is failing to reject a false H₀ (probability β), and the power is 1 − β. A gross error is a surveying blunder, a different idea.Residuals in a large set of observations are normally distributed with standard deviation σ. A surveyor rejects any observation whose residual exceeds 3σ. The probability that a good observation is wrongly rejected is about
Show answer
Answer: A — 0.27 %
±3σ contains 99.73 % of a normal population, so 100 − 99.73 = 0.27 % of good observations fall outside it. 4.55 % is the tail outside ±2σ and 31.7 % the tail outside ±1σ.Of 200 check pixels expected in equal numbers in four land-cover classes, 60, 45, 40 and 55 are observed. The chi-square statistic for goodness of fit is ____.
Numerical answer — type the value.
Show answer
Answer: 5
E = 200/4 = 50 in each class. χ² = (60 − 50)²/50 + (45 − 50)²/50 + (40 − 50)²/50 + (55 − 50)²/50 = (100 + 25 + 100 + 25)/50 = 5.0. With 3 degrees of freedom the 5 % critical value is 7.815, so equal shares are not rejected.A chi-square test of independence is applied to a contingency table of 4 land-cover classes by 3 soil types. The number of degrees of freedom is
Show answer
Answer: A — 6
For an r × c table, degrees of freedom = (r − 1)(c − 1) = 3 × 2 = 6, because row and column totals are fixed by the data. 11 is the goodness-of-fit count for 12 cells, which ignores those constraints.A level network adjustment has 15 observations and 6 unknown heights. Its weighted sum of squared residuals is vᵀWv = 27.0 with a-priori σ₀² = 1. The a-posteriori variance factor σ̂₀² is ____.
Numerical answer — type the value.
Show answer
Answer: 3
Redundancy r = n − u = 15 − 6 = 9, and σ̂₀² = vᵀWv/r = 27.0/9 = 3.0. The chi-square statistic vᵀWv/σ₀² = 27.0 on 9 degrees of freedom is far above what is expected (its mean is 9), which points to a blunder or to weights that were too optimistic. Dividing by 15 gives 1.8.Which of the following statements about least-squares regression and correlation are correct?
Show answer
Answer: A — Both regression lines pass through the point (x̄, ȳ); B — The two regression coefficients always have the same sign; D — The coefficient of determination equals the square of r for a straight-line fit
Each line has intercept fixed by the means, so both pass through (x̄, ȳ); both slopes are Sxy divided by a positive sum of squares, so they share the sign of Sxy; and R² = r² for a straight line. r = 0 rules out only a linear relation: y = x² on a symmetric range has r = 0 and a perfect relation.Which of the following are true of the chi-square test?
Show answer
Answer: A — The statistic Σ(O − E)²/E can never be negative; B — Estimating a parameter of the fitted distribution from the data reduces the degrees of freedom by one; C — The approximation is doubtful when expected counts fall below about 5
Every term is a square over a positive expected count, so χ² ≥ 0; each estimated parameter removes a degree of freedom (k − 1 − p); and small expected counts make the chi-square approximation unreliable, so classes are merged. A larger χ² means worse agreement, not better.A test is carried out at the 5 % level of significance. Which of the following are correct?
Show answer
Answer: A — The probability of a Type I error is 0.05; B — H₀ is rejected when the p-value is less than 0.05; C — For a two-tailed z test the critical values are ±1.96
α = 0.05 is by definition the Type I error probability; the decision rule is p < α; and 2.5 % in each tail of the standard normal lies beyond ±1.96. Not rejecting H₀ only means the data are consistent with it — the test may simply lack power.