Engineering Mathematics II: Regression, Correlation, Statistical Significance and the Chi-Square Test

The second half of Part A's Engineering Mathematics names regression analysis, the correlation coefficient, the “statistical significant value” and the chi-square test. They are the statistics a geomatics engineer uses every day: fitting a line to calibration data or a trend to ground-control residuals, measuring how strongly two spectral bands move together, deciding whether an apparent difference or a residual is real or chance, and testing whether an adjustment's residuals or a class distribution agree with what was expected. This chapter fits the least-squares regression line and reads its coefficients, computes Pearson's r and its relation to the two regression slopes, sets up a significance test with its level, critical value and two kinds of error, and applies the chi-square statistic to goodness of fit, to contingency tables and to the variance factor of an adjustment. The traps are nearly all about which variable is regressed on which, and how many degrees of freedom a test has.

1. Regression analysis: the least-squares line

The regression of y on x fits y = a + bx by minimising the sum of squared vertical deviations Σ(yᵢ − a − bxᵢ)². With Sxx = Σ(x − x̄)², Syy = Σ(y − ȳ)² and Sxy = Σ(x − x̄)(y − ȳ), the normal equations give b = Sxy/Sxx and a = ȳ − b x̄, so the line always passes through (x̄, ȳ). The regression of x on y, x = a′ + b′y, minimises horizontal deviations and has b′ = Sxy/Syy; the two lines are different unless the points lie exactly on a line. The coefficient of determination R² = r² is the fraction of the variance of y the line explains.

Worked data set used throughout this chapter
xyx − x̄y − ȳproduct
12−2−24
24−100
35010
44100
55212

Here x̄ = 3, ȳ = 4, Sxx = 10, Syy = 6 and Sxy = 6. So b = 6/10 = 0.6, a = 4 − 0.6 × 3 = 2.2, and the line is y = 2.2 + 0.6x; the regression of x on y has b′ = 6/6 = 1.0. The residual standard error, √(Σe²/(n − 2)), uses n − 2 because two parameters were estimated.

⚠️ You cannot invert one regression line to get the other
Solving y = 2.2 + 0.6x for x gives a slope of 1/0.6 = 1.67, not the 1.0 of the regression of x on y. Predict y from x with the first line and x from y with the second.

2. The correlation coefficient

Pearson's correlation coefficient r = Sxy/√(Sxx·Syy) = cov(x, y)/(σₓσ_y) measures the strength of the linear relation, −1 ≤ r ≤ 1. For the data above, r = 6/√(10 × 6) = 6/√60 = 0.775. Three identities are examined often: r² = b·b′ (the product of the two regression slopes: 0.6 × 1.0 = 0.6, so r = 0.775); r has the sign of both slopes, which always share a sign; and r is unchanged by a change of origin or scale of either variable (a positive scale factor), so correlating DNs or radiances gives the same r.

  • r = 0 means no linear relation; y = x² over a symmetric range has r = 0 and a perfect relation.
  • Correlation is not causation: two bands can correlate because both respond to illumination.
  • Spearman's rank correlation rₛ = 1 − 6Σd²/(n(n² − 1)) uses ranks, for monotonic but non-linear relations.
ℹ️ Where geomatics meets it
The inter-band correlation matrix of an image (Part B2) is exactly this r computed for every pair of bands; values near 1 between visible bands are the reason principal components compress multispectral data so well.

3. Statistical significance: levels, critical values and the two errors

A significance test states a null hypothesis H₀ (no difference, no correlation, the model fits), computes a test statistic whose distribution under H₀ is known, and rejects H₀ if the statistic falls beyond the critical value at the chosen level of significance α (commonly 5 % or 1 %). The p-value is the probability, under H₀, of a statistic at least as extreme as the one observed; reject when p < α. Rejecting a true H₀ is a Type I error (probability α); accepting a false H₀ is a Type II error (probability β; the power of the test is 1 − β).

Standard normal critical values (two-tailed unless stated)
ConfidencezUse
50 %0.6745the probable error
95 %1.96the usual 5 % test
99 %2.576the 1 % test
99.73 %3the surveyor's blunder rejection limit

Testing a correlation. Under H₀: ρ = 0, t = r√(n − 2)/√(1 − r²) follows Student's t with n − 2 degrees of freedom. For r = 0.6 from n = 12 pairs, t = 0.6 × √10/√0.64 = 0.6 × 3.162/0.8 = 2.37, which exceeds the two-tailed 5 % value of t with 10 degrees of freedom (2.228), so the correlation is significant. Testing a mean against a known σ: z = (x̄ − μ₀)/(σ/√n); with an estimated s and small n, use t with n − 1 degrees of freedom.

⚠️ Significant is not the same as large
With enough pairs a small r becomes significant, and with few pairs a large r may not be. Significance says the relation is unlikely to be chance; r² says how much it explains. They answer different questions.

4. The chi-square test

The chi-square statistic χ² = Σ(O − E)²/E compares observed frequencies O with those expected under H₀. It is always ≥ 0 and grows with disagreement; H₀ is rejected when χ² exceeds the critical value for the right degrees of freedom. Every expected count should be at least about 5 for the approximation to hold. Upper 5 % critical values: 3.841 (1 degree of freedom), 5.991 (2), 7.815 (3), 9.488 (4).

Three uses, three ways of counting degrees of freedom
UseStatisticDegrees of freedom
Goodness of fit to a distribution with k classesΣ(O − E)²/Ek − 1 − (number of parameters estimated from the data)
Independence in an r × c contingency tableΣ(O − E)²/E with E = row total × column total/N(r − 1)(c − 1)
Global test of an adjustment's variance factorvᵀWv/σ₀² (a-priori σ₀²)redundancy r = n − u; two-tailed

Worked: 200 checked pixels are expected to fall equally into four land-cover classes (50 each) but are observed as 60, 45, 40 and 55. χ² = (10² + 5² + 10² + 5²)/50 = 250/50 = 5.0 with 4 − 1 = 3 degrees of freedom. Since 5.0 < 7.815, the departure from equal shares is not significant at 5 %.

🎯 Why an adjustment is tested at both ends
A variance factor far above 1 means the residuals are bigger than the weights promised — a blunder or optimistic weights. One far below 1 means the observations were better than the weights assumed. Both are model errors, so the chi-square test of an adjustment is two-tailed.

Key takeaways

  • Regression of y on x: b = Sxy/Sxx, a = ȳ − b x̄; the line passes through the means, and the regression of x on y is a different line.
  • r = Sxy/√(SxxSyy), r² = b·b′, and r is unchanged by linear rescaling with a positive factor.
  • A test rejects H₀ beyond the critical value at level α; Type I error has probability α, Type II has β.
  • The significance of r uses t = r√(n − 2)/√(1 − r²) with n − 2 degrees of freedom.
  • χ² = Σ(O − E)²/E with k − 1 (minus estimated parameters) or (r − 1)(c − 1) degrees of freedom; in an adjustment, vᵀWv/σ₀² with n − u.

Practice questions (14)

Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.

  1. For the data (x, y) = (1, 2), (2, 4), (3, 5), (4, 4), (5, 5), the slope b of the least-squares regression line of y on x is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.6

    x̄ = 3, ȳ = 4. Sxy = (−2)(−2) + (−1)(0) + 0(1) + 1(0) + 2(1) = 6; Sxx = 4 + 1 + 0 + 1 + 4 = 10. b = Sxy/Sxx = 6/10 = 0.6. Using Syy (= 6) as the divisor gives 1.0, which is the slope of x on y.
  2. For the same data (1, 2), (2, 4), (3, 5), (4, 4), (5, 5), the intercept a of the regression line y = a + bx is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 2.2

    The regression line passes through (x̄, ȳ) = (3, 4). With b = 0.6, a = ȳ − b x̄ = 4 − 0.6 × 3 = 2.2. Check at x = 5: 2.2 + 3.0 = 5.2, close to the observed 5.
  3. For the data (1, 2), (2, 4), (3, 5), (4, 4), (5, 5), Pearson's correlation coefficient r (to three decimal places) is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.775

    Sxy = 6, Sxx = 10, Syy = (−2)² + 0² + 1² + 0² + 1² = 6. r = 6/√(10 × 6) = 6/7.746 = 0.775. Check through the slopes: b = 0.6, b′ = 1.0, r² = 0.6, r = √0.6 = 0.775. Reporting r² = 0.6 as r is the usual slip.
  4. The regression coefficient of y on x is 0.8 and that of x on y is 0.45. The correlation coefficient between x and y is

    1. 0.6
    2. 0.36
    3. 0.625
    4. 1.25
    Show answer

    Answer: A — 0.6

    r² = b·b′ = 0.8 × 0.45 = 0.36, so r = +0.6, positive because both coefficients are positive. 0.36 is r² itself; 1.25 cannot be a correlation at all.
  5. Every observation of x is converted from metres to feet and 100 is added to every value of y. The correlation coefficient between x and y

    1. is unchanged
    2. is multiplied by 3.28
    3. increases by 100
    4. changes sign
    Show answer

    Answer: A — is unchanged

    r is a ratio of the covariance to the product of the standard deviations; a positive scale factor multiplies both numerator and denominator alike, and a shift of origin changes neither. So r is invariant — which is why correlating raw DNs or calibrated radiances gives the same value.
  6. A sample of 12 pairs gives a correlation coefficient r = 0.6. The t statistic for testing H₀: ρ = 0 (to two decimal places) is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 2.37

    t = r√(n − 2)/√(1 − r²) = 0.6 × √10/√(1 − 0.36) = 0.6 × 3.1623/0.8 = 2.37, with 10 degrees of freedom. It exceeds the two-tailed 5 % critical value 2.228, so the correlation is significant. Using √n instead of √(n − 2) gives 2.60.
  7. Rejecting a null hypothesis that is in fact true is called

    1. a Type I error
    2. a Type II error
    3. the power of the test
    4. a gross error
    Show answer

    Answer: A — a Type I error

    A Type I error is a false rejection and its probability is the significance level α. A Type II error is failing to reject a false H₀ (probability β), and the power is 1 − β. A gross error is a surveying blunder, a different idea.
  8. Residuals in a large set of observations are normally distributed with standard deviation σ. A surveyor rejects any observation whose residual exceeds 3σ. The probability that a good observation is wrongly rejected is about

    1. 0.27 %
    2. 4.55 %
    3. 31.7 %
    4. 5 %
    Show answer

    Answer: A — 0.27 %

    ±3σ contains 99.73 % of a normal population, so 100 − 99.73 = 0.27 % of good observations fall outside it. 4.55 % is the tail outside ±2σ and 31.7 % the tail outside ±1σ.
  9. Of 200 check pixels expected in equal numbers in four land-cover classes, 60, 45, 40 and 55 are observed. The chi-square statistic for goodness of fit is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 5

    E = 200/4 = 50 in each class. χ² = (60 − 50)²/50 + (45 − 50)²/50 + (40 − 50)²/50 + (55 − 50)²/50 = (100 + 25 + 100 + 25)/50 = 5.0. With 3 degrees of freedom the 5 % critical value is 7.815, so equal shares are not rejected.
  10. A chi-square test of independence is applied to a contingency table of 4 land-cover classes by 3 soil types. The number of degrees of freedom is

    1. 6
    2. 12
    3. 11
    4. 5
    Show answer

    Answer: A — 6

    For an r × c table, degrees of freedom = (r − 1)(c − 1) = 3 × 2 = 6, because row and column totals are fixed by the data. 11 is the goodness-of-fit count for 12 cells, which ignores those constraints.
  11. A level network adjustment has 15 observations and 6 unknown heights. Its weighted sum of squared residuals is vᵀWv = 27.0 with a-priori σ₀² = 1. The a-posteriori variance factor σ̂₀² is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 3

    Redundancy r = n − u = 15 − 6 = 9, and σ̂₀² = vᵀWv/r = 27.0/9 = 3.0. The chi-square statistic vᵀWv/σ₀² = 27.0 on 9 degrees of freedom is far above what is expected (its mean is 9), which points to a blunder or to weights that were too optimistic. Dividing by 15 gives 1.8.
  12. Which of the following statements about least-squares regression and correlation are correct?

    1. Both regression lines pass through the point (x̄, ȳ)
    2. The two regression coefficients always have the same sign
    3. r = 0 proves that x and y are unrelated
    4. The coefficient of determination equals the square of r for a straight-line fit
    Show answer

    Answer: A — Both regression lines pass through the point (x̄, ȳ); B — The two regression coefficients always have the same sign; D — The coefficient of determination equals the square of r for a straight-line fit

    Each line has intercept fixed by the means, so both pass through (x̄, ȳ); both slopes are Sxy divided by a positive sum of squares, so they share the sign of Sxy; and R² = r² for a straight line. r = 0 rules out only a linear relation: y = x² on a symmetric range has r = 0 and a perfect relation.
  13. Which of the following are true of the chi-square test?

    1. The statistic Σ(O − E)²/E can never be negative
    2. Estimating a parameter of the fitted distribution from the data reduces the degrees of freedom by one
    3. The approximation is doubtful when expected counts fall below about 5
    4. A larger chi-square value means a better fit
    Show answer

    Answer: A — The statistic Σ(O − E)²/E can never be negative; B — Estimating a parameter of the fitted distribution from the data reduces the degrees of freedom by one; C — The approximation is doubtful when expected counts fall below about 5

    Every term is a square over a positive expected count, so χ² ≥ 0; each estimated parameter removes a degree of freedom (k − 1 − p); and small expected counts make the chi-square approximation unreliable, so classes are merged. A larger χ² means worse agreement, not better.
  14. A test is carried out at the 5 % level of significance. Which of the following are correct?

    1. The probability of a Type I error is 0.05
    2. H₀ is rejected when the p-value is less than 0.05
    3. For a two-tailed z test the critical values are ±1.96
    4. Failing to reject H₀ proves that H₀ is true
    Show answer

    Answer: A — The probability of a Type I error is 0.05; B — H₀ is rejected when the p-value is less than 0.05; C — For a two-tailed z test the critical values are ±1.96

    α = 0.05 is by definition the Type I error probability; the decision rule is p < α; and 2.5 % in each tail of the standard normal lies beyond ±1.96. Not rejecting H₀ only means the data are consistent with it — the test may simply lack power.