Mathematics and Statistics in Ecology II: Study Design, Hypothesis Testing, t and Chi-square Tests, Correlation, Regression and ANOVA
1. Study design: replication, randomisation, controls and pseudoreplication
An observational study measures variables as they occur and can show association; a manipulative experiment applies treatments to units and, with randomisation, can show cause. Three principles make an experiment interpretable. Replication — several independent experimental units per treatment — estimates the natural variation against which a treatment effect is judged. Randomisation of treatments to units prevents systematic differences (a wetter corner of a field, the observer’s choice) from being confounded with the treatment. Controls show what happens without the treatment, including procedural controls (a cage that excludes nothing, to test the effect of the cage itself).
Pseudoreplication (Hurlbert, 1984) is treating non-independent measurements as independent replicates: taking 20 soil cores from one fertilised plot and 20 from one unfertilised plot gives n = 1 per treatment, not 20, because any difference between the plots — not only the fertiliser — could explain the result. The same trap arises with repeated measurements of the same animals, several fish in one tank, or several quadrats in one burnt and one unburnt forest. Where a single impact site cannot be replicated, the BACI design (Before–After, Control–Impact) compares the change at the impact site with the change at control sites over the same period. Sample size is chosen in advance by a power analysis, from the effect size worth detecting, the expected variance, α and the power wanted.
2. Hypothesis testing, the p-value, errors and power
A test sets a null hypothesis H₀ (no effect, no difference) against an alternative H₁, computes a test statistic from the data, and asks how surprising it would be if H₀ were true. The p-value is the probability, assuming H₀ is true, of obtaining a test statistic at least as extreme as the one observed. If p ≤ α, the chosen significance level (conventionally 0.05), H₀ is rejected. A two-tailed test counts extremes in both directions; a one-tailed test is justified only if a difference in the other direction would be treated exactly like no difference.
| Decision | H₀ true | H₀ false |
|---|---|---|
| Reject H₀ | Type I error — probability α | correct — probability 1 − β, the power |
| Do not reject H₀ | correct — probability 1 − α | Type II error — probability β |
Lowering α to avoid false positives raises β for a fixed sample; the way to lower both is a larger sample or a less variable measurement. Power (1 − β) rises with the effect size, the sample size and α, and falls with the variance; 0.8 is a common target. Multiple testing inflates the Type I error: with k independent tests at α = 0.05, the chance of at least one false positive is 1 − 0.95^k, which is 0.226 for k = 5. The Bonferroni correction tests each at α/k (0.005 for ten tests).
3. The t-tests
| Test | Statistic | Degrees of freedom | Use |
|---|---|---|---|
| One-sample | t = (x̄ − μ₀)/(s/√n) | n − 1 | does a mean differ from a stated value? |
| Two-sample (pooled) | t = (x̄₁ − x̄₂)/[s_p √(1/n₁ + 1/n₂)] | n₁ + n₂ − 2 | do two independent groups differ? |
| Paired | t = d̄/(s_d/√n) | n − 1 (n pairs) | the same units measured twice, or matched pairs |
The pooled standard deviation is s_p = √{[(n₁ − 1)s₁² + (n₂ − 1)s₂²]/(n₁ + n₂ − 2)}. Worked example: two groups of 10 lizards with mean masses 12 g and 10 g and s_p = 2 g give t = 2/(2 × √0.2) = 2/0.894 = 2.24, on 18 degrees of freedom; the two-tailed 5 % critical value is 2.101, so the difference is significant. The t-tests assume approximately normal data (or large samples) and independent observations; the pooled test also assumes equal variances, and Welch’s test drops that assumption. A paired design removes variation between units: if 9 trees are measured before and after a treatment with mean difference 1.5 and s_d = 3, t = 1.5/(3/3) = 1.5.
4. Chi-square tests: goodness of fit and contingency tables
The chi-square statistic compares observed counts O with expected counts E: χ² = Σ(O − E)²/E, always on counts, never on percentages. In a goodness-of-fit test the expected counts come from a hypothesis. A monohybrid cross expected to give 3:1 that yields 290 dominant and 110 recessive offspring (E = 300 and 100) gives χ² = 10²/300 + 10²/100 = 0.33 + 1.00 = 1.33, on 1 degree of freedom; the 5 % critical value is 3.84, so the 3:1 ratio is not rejected. Degrees of freedom are the number of categories minus 1, minus one more for each parameter estimated from the data.
A Hardy–Weinberg test is a goodness-of-fit test with one estimated parameter. For 50 AA, 30 Aa and 20 aa: p = (100 + 30)/200 = 0.65 and q = 0.35, so the expected counts are 42.25, 45.5 and 12.25. χ² = 7.75²/42.25 + 15.5²/45.5 + 7.75²/12.25 = 1.42 + 5.28 + 4.90 = 11.60, with 3 − 1 − 1 = 1 degree of freedom, far above 3.84: the population departs from Hardy–Weinberg proportions, with a deficit of heterozygotes. A contingency test asks whether two categorical variables are independent — say, habitat type and presence of a species. Expected counts are (row total × column total)/grand total, and the degrees of freedom are (r − 1)(c − 1); a 3 × 4 table has 6. The test is unreliable when expected counts are small (a common rule is that none should be below 5).
5. Correlation and linear regression
Pearson’s correlation coefficient r = S_xy/√(S_xx S_yy), where S_xy = Σ(x − x̄)(y − ȳ), S_xx = Σ(x − x̄)² and S_yy = Σ(y − ȳ)², measures the strength of a linear association on a scale from −1 to +1, with no distinction between the variables. Least-squares linear regression fits y = a + bx by minimising the sum of squared vertical residuals, giving b = S_xy/S_xx and a = ȳ − b x̄; it treats x as the predictor measured without error and y as the response. The coefficient of determination r² is the fraction of the variance in y explained by the line, and b = r × s_y/s_x links the two.
| Quantity | Working | Value |
|---|---|---|
| Means | x̄ = 15/5, ȳ = 20/5 | 3 and 4 |
| S_xy | (−2)(−2) + (−1)(0) + (0)(1) + (1)(0) + (2)(1) | 6 |
| S_xx and S_yy | 4 + 1 + 0 + 1 + 4; 4 + 0 + 1 + 0 + 1 | 10 and 6 |
| Slope and intercept | b = 6/10; a = 4 − 0.6 × 3 | 0.6 and 2.2 |
| r and r² | 6/√(10 × 6) = 6/7.746 | 0.775 and 0.6 |
6. One-way analysis of variance
ANOVA tests whether the means of k groups are all equal by comparing variation between groups with variation within them. The total sum of squares partitions exactly: SS_total = SS_between + SS_within, with degrees of freedom N − 1 = (k − 1) + (N − k). The mean squares are MS_between = SS_between/(k − 1) and MS_within = SS_within/(N − k), and F = MS_between/MS_within. If H₀ is true both estimate the same error variance and F ≈ 1; a large F means the group means differ. With k = 3 groups of 5 (N = 15), SS_between = 40 and SS_within = 60: MS_between = 20, MS_within = 5 and F = 4.0 on (2, 12) degrees of freedom, above the 5 % critical value of 3.89.
- A significant F says that at least one mean differs, not which one; post hoc comparisons (Tukey’s test, for example) find which pairs differ while controlling the overall Type I error.
- Running all pairwise t-tests instead inflates the Type I error; with three groups that is three tests.
- For two groups, one-way ANOVA and the pooled two-sample t-test are equivalent: F = t².
- ANOVA assumes independent observations, roughly normal residuals and equal variances among groups; non-parametric alternatives (Kruskal–Wallis) relax normality.
Key takeaways
- Replicate the experimental unit, randomise treatments and include controls; many samples from one plot per treatment are pseudoreplicates, n = 1.
- The p-value is P(data this extreme | H₀); Type I error has probability α, Type II β, power 1 − β; k tests give 1 − (1 − α)^k false-positive risk, and Bonferroni uses α/k.
- Two-sample t = (x̄₁ − x̄₂)/[s_p√(1/n₁ + 1/n₂)] on n₁ + n₂ − 2 df; paired t = d̄/(s_d/√n) on n − 1.
- χ² = Σ(O − E)²/E on counts; df = categories − 1 − estimated parameters (1 for a Hardy–Weinberg test) or (r − 1)(c − 1) for a contingency table.
- b = S_xy/S_xx, a = ȳ − bx̄, r² is the variance explained; ANOVA’s F = MS_between/MS_within on (k − 1, N − k) df, and F = t² for two groups.
Practice questions (15)
Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.
To test the effect of fertiliser on soil microbes, a researcher fertilises one field, leaves a neighbouring field unfertilised, and takes 30 soil cores from each. The main flaw is
Show answer
Answer: A — pseudoreplication: there is only one experimental unit per treatment
The treatment was applied to fields, so the field is the experimental unit and n = 1 per treatment. The 30 cores are subsamples; any pre-existing difference between the two fields is confounded with the fertiliser. More cores do not fix this; more fields, randomly assigned, do.Which statements about Type I and Type II errors are correct?
Show answer
Answer: A — A Type I error is rejecting a true null hypothesis; B — The probability of a Type II error is β, and power is 1 − β; C — For a fixed sample size, lowering α tends to raise β
Type I is a false rejection (probability α), Type II a failure to reject a false H₀ (probability β), and power is 1 − β. With the sample fixed, demanding stronger evidence to reject (smaller α) makes misses more likely. A p-value is computed assuming H₀ is true, so it cannot be the probability that H₀ is true.Five independent tests are carried out, each at α = 0.05, and every null hypothesis is true. The probability of at least one false positive, to three decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.226
P(no false positive) = 0.95⁵ = 0.774, so P(at least one) = 1 − 0.774 = 0.226. Adding 5 × 0.05 = 0.25 overstates it slightly because the events can overlap; the Bonferroni level for five tests would be 0.01.A study makes 10 independent comparisons and wants an overall Type I error rate of 0.05. Under the Bonferroni correction, each comparison should be tested at a significance level of ____.
Numerical answer — type the value.
Show answer
Answer: 0.005
Bonferroni divides α by the number of tests: 0.05/10 = 0.005. Multiplying, 0.5, would make false positives far more likely rather than less.Two independent groups of 10 lizards have mean masses 12 g and 10 g, and the pooled standard deviation is 2 g. The two-sample t statistic, to two decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 2.24
t = (12 − 10)/[2 × √(1/10 + 1/10)] = 2/(2 × 0.4472) = 2/0.894 = 2.24, on 18 df. Omitting the √(1/n₁ + 1/n₂) term gives 1.0, and using √(1/10) alone gives 3.16.Nine trees are measured before and after a treatment. The mean of the nine differences is 1.5 and their standard deviation is 3. The paired t statistic, to one decimal place, is ____.
Numerical answer — type the value.
Show answer
Answer: 1.5
t = d̄/(s_d/√n) = 1.5/(3/√9) = 1.5/1 = 1.5, on 8 df. Dividing by s_d alone gives 0.5, which forgets that the standard error of the mean difference shrinks with √n.A cross expected to give a 3:1 ratio produces 290 dominant and 110 recessive offspring. The chi-square statistic for goodness of fit, to two decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 1.33
Expected 300 and 100; χ² = (290 − 300)²/300 + (110 − 100)²/100 = 0.333 + 1.000 = 1.33, below the 1-df critical value 3.84. Dividing by the observed instead of the expected counts gives 0.345 + 0.909 = 1.25.A sample of 100 individuals has 50 AA, 30 Aa and 20 aa. The chi-square statistic for departure from Hardy–Weinberg proportions (expected counts from the sample allele frequency), to one decimal place, is ____.
Numerical answer — type the value.
Show answer
Answer: 11.6
p = 130/200 = 0.65, q = 0.35; expected 42.25, 45.5 and 12.25. χ² = 7.75²/42.25 + 15.5²/45.5 + 7.75²/12.25 = 1.42 + 5.28 + 4.90 = 11.6, on 3 − 1 − 1 = 1 df — a highly significant heterozygote deficit. Assuming p = 0.5 instead of estimating it gives different expected counts and a different value.A contingency table cross-classifies sites by three habitat types and four abundance classes. The number of degrees of freedom for the chi-square test of independence is ____.
Numerical answer — type the value.
Show answer
Answer: 6
df = (r − 1)(c − 1) = (3 − 1)(4 − 1) = 2 × 3 = 6. The total number of cells minus one, 11, is the goodness-of-fit rule applied to the wrong kind of test.A Hardy–Weinberg chi-square test is carried out on counts of the three genotypes at a two-allele locus, with the allele frequency estimated from the same data. The degrees of freedom are
Show answer
Answer: A — 1
Three genotype classes give 3 − 1 = 2, and estimating p from the data removes one more: 3 − 1 − 1 = 1. Answering 2 is the commonest error in this question; it would be right only if p were specified in advance rather than estimated.For the data x = 1, 2, 3, 4, 5 and y = 2, 4, 5, 4, 5, the least-squares slope of y on x, to one decimal place, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.6
x̄ = 3, ȳ = 4, S_xy = 4 + 0 + 0 + 0 + 2 = 6 and S_xx = 10, so b = 6/10 = 0.6 and the intercept a = 4 − 1.8 = 2.2. Dividing by S_yy = 6 gives 1.0, the slope of x on y inverted.For the same data (x = 1, 2, 3, 4, 5; y = 2, 4, 5, 4, 5), the coefficient of determination r², to one decimal place, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.6
r² = S_xy²/(S_xx S_yy) = 36/(10 × 6) = 0.6, so the line explains 60 % of the variance in y; r = 0.775. Reporting r instead of r² gives 0.8, and it is a coincidence of these data that r² equals the slope.Which statements about correlation and regression are correct?
Show answer
Answer: A — r² is the proportion of variance in y explained by the linear regression on x; B — The slope of y on x equals r × s_y/s_x; C — A correlation near zero can occur with a strong hump-shaped relationship
r² is the explained fraction and b = r s_y/s_x. Pearson’s r measures linear association only, so a symmetric hump can give r ≈ 0. The two regression slopes multiply to r², so they are inverses only when r² = 1.A one-way ANOVA compares 3 groups of 5 observations each. The between-group sum of squares is 40 and the within-group sum of squares is 60. The F statistic, to one decimal place, is ____.
Numerical answer — type the value.
Show answer
Answer: 4.0
MS_between = 40/(3 − 1) = 20 and MS_within = 60/(15 − 3) = 5, so F = 20/5 = 4.0 on (2, 12) df. Dividing the sums of squares directly, 40/60 = 0.67, ignores the degrees of freedom.An ecologist compares mean leaf nitrogen among five sites. Why is one-way ANOVA preferred to running every pairwise t-test?
Show answer
Answer: A — Ten separate t-tests would inflate the overall Type I error rate
Five sites give ⁵C₂ = 10 pairwise tests, and at α = 0.05 each the chance of at least one false positive is far above 5 %. ANOVA tests all means at once at α. It still assumes independence, and it does not identify which pairs differ — that needs post hoc tests.