Statistics for Environmental Data: Shape, Risk and Reliability, Correlation, Regression and Hypothesis Tests

The second chapter written for the Mathematics Foundation section of the GATE Environmental Science and Engineering (ES) paper. Its Probability and Statistics heading is longer than any engineering paper’s: after descriptive statistics, central tendency, dispersion, probability, conditional probability, Bayes’ theorem and the distributions — which the shared Probability and Statistics chapter teaches — it names skewness and kurtosis, risk and reliability, correlation, single and multiple regression models, and hypothesis testing by the t-test, the F-test and the chi-square test. Those are gathered here, with the environmental reading of each: a skewed set of coliform counts, the risk that a 100-year flood arrives within a design life, the reliability of duplicate pumps, the regression of BOD on COD, and the test of whether an effluent mean really exceeds its limit.

1. Dispersion, skewness and kurtosis

Beyond the mean, median, mode and standard deviation, three descriptive measures recur in environmental data. The coefficient of variation, CV = σ/x̄ × 100, compares the variability of sets with different means — a treatment plant whose effluent BOD has a CV of 15% is steadier than one with 40%, whatever their averages. The geometric mean, (x₁x₂…xₙ)^(1/n), is the natural average of data that span orders of magnitude, which is why bacterial counts and many water-quality criteria are stated as geometric means: counts of 10, 100 and 1000 have a geometric mean of 100 and an arithmetic mean of 370, which one bad sample dominates. The interquartile range, Q₃ − Q₁, measures spread without being moved by extremes.

Skewness measures asymmetry. Karl Pearson’s first coefficient is Sk = (mean − mode)/σ; his second, used when the mode is ill-defined, is Sk = 3(mean − median)/σ; the moment coefficient is γ₁ = μ₃/σ³, where μ₃ is the third central moment. For a positively skewed (right-tailed) distribution mean > median > mode — the usual shape of pollutant concentrations, which cannot fall below zero but can spike far above the typical value. Kurtosis measures the weight of the tails relative to the peak: β₂ = μ₄/μ₂², equal to 3 for the normal distribution (mesokurtic); β₂ > 3 is leptokurtic (sharper peak, heavier tails) and β₂ < 3 platykurtic. The excess kurtosis β₂ − 3 is zero for the normal curve.

🧠 Read the order of mean, median and mode
Right-skewed: mode < median < mean, and the mean is dragged toward the long tail. Left-skewed: the order reverses. Symmetric: all three coincide. The empirical relation mean − mode ≈ 3(mean − median) holds for moderately skewed sets and is why Pearson’s two coefficients agree.

2. Risk and reliability

Reliability R(t) is the probability that a component performs without failure up to time t. With a constant failure rate λ the time to failure is exponential, R(t) = e^(−λt), and the mean time to failure is MTTF = 1/λ; the memoryless property of the exponential means a pump that has run 1000 hours is as reliable for the next 100 as a new one. Systems combine by the rules of independent events. Components in series all must work, so R_s = R₁R₂…Rₙ, always less than the weakest component. Components in parallel (redundancy) fail only if all fail, so R_p = 1 − (1 − R₁)(1 − R₂)…(1 − Rₙ). Two standby pumps each of reliability 0.9 give 1 − 0.1² = 0.99.

In hydrology and environmental design, risk is the probability that a design event is exceeded during a design life. An event of return period T has annual exceedance probability p = 1/T, so the probability that it occurs at least once in n years is R = 1 − (1 − 1/T)ⁿ, and the reliability of the design is 1 − R. A 100-year flood has a 22% chance of occurring within a 25-year design life, which is why a return period much longer than the design life is chosen when failure is costly. In health and hazard assessment risk is also written as probability × consequence, and the environmental-management chapter builds human-health risk on that footing.

⚠️ The chance of a T-year event in T years is not 1
Set n = T: R = 1 − (1 − 1/T)^T, which tends to 1 − 1/e ≈ 0.63 for large T. A 50-year storm has about a 64% chance, not a certainty, of occurring in the next 50 years — and a 36% chance of not occurring at all.

3. Correlation

The covariance of x and y is cov(x, y) = Σ(x − x̄)(y − ȳ)/n, and Karl Pearson’s correlation coefficient normalises it: r = cov(x, y)/(σₓσᵧ) = [nΣxy − ΣxΣy]/√{[nΣx² − (Σx)²][nΣy² − (Σy)²]}. It lies between −1 and +1, is unchanged by a change of origin or scale of either variable (apart from sign if a scale factor is negative), and measures only linear association: y = x² over a range symmetric about zero has r = 0 despite a perfect relation. Correlation also says nothing about cause — dissolved oxygen and fish kills correlate with temperature in summer because warm water holds less oxygen, not because the thermometer kills fish.

When the data are ranks, or when a monotonic but non-linear relation is suspected, Spearman’s rank correlation is used: ρ = 1 − 6Σd²/[n(n² − 1)], where d is the difference between the two ranks of each item and n the number of items. Tied values receive the average of the ranks they would occupy.

4. Single and multiple regression

The regression line of y on x, fitted by least squares, is y − ȳ = b_yx(x − x̄) with b_yx = r σᵧ/σₓ = cov(x, y)/σₓ²; the line of x on y is x − x̄ = b_xy(y − ȳ) with b_xy = r σₓ/σᵧ. Three consequences carry most of the marks: both lines pass through (x̄, ȳ), so they intersect at the means; r² = b_yx × b_xy, with r taking the common sign of the two coefficients; and the two lines coincide only when r = ±1, becoming perpendicular when r = 0. The coefficient of determination r² is the fraction of the variance of y explained by the regression — an r of 0.8 explains 64% of it, not 80%.

Multiple regression models one response on several predictors: y = b₀ + b₁x₁ + b₂x₂ + … + bₖxₖ + ε — BOD on COD and TSS, or a pollutant concentration on wind speed, temperature and traffic count. The coefficients are found by least squares from the normal equations, which in matrix form are (XᵀX)b = Xᵀy, so b = (XᵀX)⁻¹Xᵀy. Each bᵢ is the change in y per unit xᵢ with the other predictors held constant. The multiple R² never decreases when a predictor is added, even a useless one, so models are compared by the adjusted R², 1 − (1 − R²)(n − 1)/(n − k − 1), which penalises extra predictors. Strongly correlated predictors (multicollinearity) make the individual coefficients unstable even when the fit is good.

⚠️ A regression coefficient above 1 is allowed; two of them are not
One of b_yx and b_xy may exceed 1 in magnitude, but their product is r² ≤ 1, so both cannot. A question offering b_yx = 1.6 and b_xy = 0.9 describes an impossible data set, since 1.44 > 1.

5. Hypothesis testing: t, F and chi-square

A test sets a null hypothesis H₀ (no difference, no effect) against an alternative H₁, computes a statistic from the sample, and rejects H₀ if the statistic falls in the critical region fixed by the significance level α. Rejecting a true H₀ is a Type I error, with probability α; failing to reject a false H₀ is a Type II error, with probability β, and 1 − β is the power. The p-value is the probability, if H₀ were true, of a statistic at least as extreme as the one observed; reject H₀ when p < α. Failing to reject is not proof that H₀ is true — a small sample may simply lack power.

The three tests the syllabus names
TestStatisticDegrees of freedom and use
One-sample tt = (x̄ − μ₀)/(s/√n), s the sample standard deviationn − 1; does the mean effluent BOD differ from a limit μ₀ when σ is unknown and n is small?
Two-sample t (pooled)t = (x̄₁ − x̄₂)/[s_p√(1/n₁ + 1/n₂)]n₁ + n₂ − 2; upstream against downstream, before against after; paired data use the differences instead
F-testF = s₁²/s₂², larger variance in the numerator(n₁ − 1, n₂ − 1); are two variances equal? Also the basis of ANOVA
Chi-squareχ² = Σ(O − E)²/E over all categoriesk − 1 − (parameters estimated) for goodness of fit; (r − 1)(c − 1) for an r × c contingency table, with E = row total × column total/grand total
ℹ️ z or t?
Use z when the population standard deviation is known or the sample is large (roughly n > 30); use t when σ is estimated from a small sample. The t distribution has heavier tails than the normal and approaches it as the degrees of freedom grow.

Key takeaways

  • CV = σ/x̄ compares variability across different means; the geometric mean is the right average for counts spanning orders of magnitude.
  • Pearson skewness is (mean − mode)/σ or 3(mean − median)/σ; right skew puts the mean above the median. Kurtosis β₂ = μ₄/μ₂² is 3 for the normal curve.
  • Series reliability multiplies; parallel reliability is one minus the product of the failure probabilities. The risk of a T-year event within n years is 1 − (1 − 1/T)ⁿ.
  • Both regression lines pass through (x̄, ȳ); r² = b_yx × b_xy; r² is the fraction of variance explained, and adjusted R² is how multiple-regression models are compared.
  • t = (x̄ − μ₀)/(s/√n) with n − 1 degrees of freedom; F is the ratio of variances, larger on top; χ² = Σ(O − E)²/E. α is the Type I error probability, and not rejecting H₀ does not prove it.

Practice questions (17)

Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.

  1. A set of daily PM₁₀ readings has mean 50 μg/m³, mode 44 μg/m³ and standard deviation 12 μg/m³. What is Karl Pearson’s first coefficient of skewness?

    Numerical answer — type the value.

    Show answer

    Answer: 0.5

    Sk = (mean − mode)/σ = (50 − 44)/12 = 0.5, a moderate positive (right) skew, the usual shape for a pollutant that cannot go below zero but can spike. Dividing by the variance, 144, gives 0.042 and confuses σ with σ².
  2. For a set of effluent COD values the mean is 30 mg/L, the median 27 mg/L and the standard deviation 6 mg/L. What is Pearson’s second coefficient of skewness?

    Numerical answer — type the value.

    Show answer

    Answer: 1.5

    Sk = 3(mean − median)/σ = 3 × (30 − 27)/6 = 9/6 = 1.5. Omitting the factor 3 gives 0.5, which is the first coefficient’s form applied to the median.
  3. A distribution whose moment coefficient of kurtosis β₂ is 4.2 is:

    1. platykurtic
    2. mesokurtic
    3. leptokurtic
    4. negatively skewed
    Show answer

    Answer: C — leptokurtic

    The normal distribution has β₂ = 3 (mesokurtic). A value above 3 means a sharper peak and heavier tails, which is leptokurtic; below 3 is platykurtic. Kurtosis says nothing about asymmetry, so option D answers a different question.
  4. Three samples from a bathing beach give faecal coliform counts of 10, 100 and 1000 per 100 mL. What is their geometric mean, per 100 mL?

    Numerical answer — type the value.

    Show answer

    Answer: 100

    GM = (10 × 100 × 1000)^(1/3) = (10⁶)^(1/3) = 100. The arithmetic mean, 370, is dominated by the single highest sample, which is why bacteriological criteria use the geometric mean.
  5. A dosing line has a pump, a flow controller and a valve in series, with reliabilities 0.90, 0.95 and 0.99. What is the reliability of the line? Give the answer to three decimal places.

    Numerical answer — type the value.

    Show answer

    Answer: 0.846

    In series every component must work: R = 0.90 × 0.95 × 0.99 = 0.8465, i.e. 0.846 to three places — lower than the weakest component, 0.90. Adding the reliabilities or taking the minimum overstates it.
  6. Two identical pumps, each with reliability 0.9, are installed in parallel so that either can meet the demand. What is the reliability of the pair?

    Numerical answer — type the value.

    Show answer

    Answer: 0.99

    The pair fails only if both fail: R = 1 − (1 − 0.9)(1 − 0.9) = 1 − 0.01 = 0.99. Multiplying 0.9 × 0.9 = 0.81 is the series rule and describes a system that needs both pumps.
  7. A levee is designed for the 100-year flood. What is the probability that this flood is equalled or exceeded at least once during a 25-year design life? Give the answer to three decimal places.

    Numerical answer — type the value.

    Show answer

    Answer: 0.222

    R = 1 − (1 − 1/T)ⁿ = 1 − 0.99²⁵ = 1 − 0.7778 = 0.222. The shortcut n/T = 0.25 overstates it by counting years with two or more floods more than once.
  8. A blower has a constant failure rate of 0.002 per hour. What is the probability that it runs 100 hours without failure? Give the answer to three decimal places.

    Numerical answer — type the value.

    Show answer

    Answer: 0.819

    With a constant rate, R(t) = e^(−λt) = e^(−0.002 × 100) = e^(−0.2) = 0.819. The linear approximation 1 − λt = 0.8 is the first two terms of the Taylor series and slightly underestimates it; the MTTF is 1/λ = 500 h.
  9. For paired BOD and COD data the regression coefficient of y on x is 0.8 and that of x on y is 0.45, both positive. What is the correlation coefficient r?

    Numerical answer — type the value.

    Show answer

    Answer: 0.6

    r² = b_yx × b_xy = 0.8 × 0.45 = 0.36, so r = +0.6, positive because both coefficients are positive. Taking their mean, 0.625, is a common wrong shortcut: r is their geometric mean.
  10. Five monitoring stations are ranked by two analysts, and the sum of the squared differences between their ranks is 4. What is Spearman’s rank correlation coefficient?

    Numerical answer — type the value.

    Show answer

    Answer: 0.8

    ρ = 1 − 6Σd²/[n(n² − 1)] = 1 − 6 × 4/(5 × 24) = 1 − 24/120 = 0.8. Using n² instead of n(n² − 1) in the denominator gives 0.04, which is meaningless for a correlation.
  11. The two regression lines of a bivariate data set intersect at:

    1. the origin
    2. the point of means (x̄, ȳ)
    3. the point of medians
    4. nowhere, since they are parallel
    Show answer

    Answer: B — the point of means (x̄, ȳ)

    Each least-squares line satisfies its normal equations, the first of which says the line passes through (x̄, ȳ); both do, so they meet there. They are parallel never, coincide when r = ±1, and are perpendicular when r = 0.
  12. A second predictor is added to a regression of effluent BOD on COD. The multiple R² rises from 0.700 to 0.702 while the adjusted R² falls. The correct conclusion is:

    1. the new predictor explains a substantial share of the variance and should be kept
    2. the calculation is wrong, because R² and adjusted R² always move together
    3. the new predictor adds almost nothing, and the simpler model is preferable
    4. the correlation between BOD and COD has become negative
    Show answer

    Answer: C — the new predictor adds almost nothing, and the simpler model is preferable

    R² can never fall when a predictor is added, so a tiny rise is expected even from noise. Adjusted R² = 1 − (1 − R²)(n − 1)/(n − k − 1) charges for the lost degree of freedom and falls when the gain is too small to pay for it — the signal to keep the simpler model. Option B is false because the two measures are designed to diverge.
  13. Sixteen effluent samples have a mean BOD of 52 mg/L and a sample standard deviation of 8 mg/L. What is the t statistic for testing whether the mean exceeds a limit of 48 mg/L?

    Numerical answer — type the value.

    Show answer

    Answer: 2

    t = (x̄ − μ₀)/(s/√n) = (52 − 48)/(8/√16) = 4/2 = 2, with 15 degrees of freedom. Dividing by s alone, 4/8 = 0.5, forgets that the standard error of a mean shrinks with √n.
  14. A survey expects wet, dry and domestic-hazardous waste in the proportions 25 : 50 : 25 and finds 30, 50 and 20 bins in a sample of 100. What is the chi-square statistic?

    Numerical answer — type the value.

    Show answer

    Answer: 2

    E = 25, 50, 25. χ² = (30 − 25)²/25 + (50 − 50)²/50 + (20 − 25)²/25 = 1 + 0 + 1 = 2, with 3 − 1 = 2 degrees of freedom. Dividing by the observed counts instead of the expected ones is the usual slip.
  15. Two laboratories analyse the same reference sample; their sample variances are 16 and 9 (mg/L)². What is the F statistic for testing equality of their variances? Give the answer to two decimal places.

    Numerical answer — type the value.

    Show answer

    Answer: 1.78

    F = larger variance/smaller variance = 16/9 = 1.78, so that F ≥ 1 and the upper-tail table applies. Using the standard deviations, 4/3 = 1.33, is not an F statistic.
  16. In a contingency table of 200 households classified by income group and by whether they segregate waste, one row totals 60 and one column totals 40. What is the expected frequency in the cell where they meet, under independence?

    Numerical answer — type the value.

    Show answer

    Answer: 12

    E = row total × column total/grand total = 60 × 40/200 = 12. This is the expected count if the two classifications were independent; χ² then compares each observed cell with its E.
  17. Which of the following statements about hypothesis testing are correct?

    1. The significance level α is the probability of rejecting a true null hypothesis
    2. If the p-value is less than α, the null hypothesis is rejected
    3. For a fixed α, a larger sample increases the power of the test
    4. A result that is not significant proves the null hypothesis true
    Show answer

    Answer: A — The significance level α is the probability of rejecting a true null hypothesis; B — If the p-value is less than α, the null hypothesis is rejected; C — For a fixed α, a larger sample increases the power of the test

    (a) defines the Type I error probability. (b) is the p-value decision rule. (c) holds because the standard error falls as √n grows, so a real difference is more likely to be detected and β falls. (d) is false: failing to reject H₀ may mean only that the sample was too small to detect a real effect.