Mathematics and Statistics in Ecology I: Functions and Rates, Probability Models for Field Data, and Descriptive Statistics
1. Simple functions as ecological models
| Function | Form | Ecological use | How to linearise |
|---|---|---|---|
| Linear | y = a + bx | per-capita growth rate against density in the logistic model; b is the slope | already linear |
| Quadratic | y = a + bx + cx² | hump-shaped responses: population growth rate against N, species performance along a gradient | vertex at x = −b/(2c) |
| Exponential | N = N₀ert | unrestricted growth, radioactive decay, first-order decomposition | ln N = ln N₀ + rt: a semi-log plot with slope r |
| Logarithmic | y = a + b ln x | diminishing returns; richness against sampling effort | plot y against ln x |
| Power | y = ax^b | species–area S = cA^z; allometry, e.g. metabolic rate ∝ M3/4 | log y = log a + b log x: a log–log plot with slope b |
| Saturating | f = aN/(1 + ahN) | Holling Type II functional response; Michaelis–Menten uptake | 1/f against 1/N (Lineweaver–Burk form) |
The log of a quantity turns multiplication into addition, which is why ecological data are so often plotted on log axes. On a semi-log plot (ln N against t) exponential growth is a straight line whose slope is r: a population whose ln N rises from 4.0 to 5.2 in 6 years has r = 1.2/6 = 0.2 per year. On a log–log plot a power law is a straight line whose slope is the exponent. In allometry, Kleiber’s law gives basal metabolic rate roughly ∝ M0.75, so a 16-fold heavier animal uses 160.75 = 8 times the energy, but only half as much per gram — which is why small mammals must eat so much relative to their size.
The Holling Type II response f = aN/(1 + ahN) has a searching efficiency (attack rate) a and a handling time h. At low prey density f ≈ aN, rising linearly; at high density f approaches 1/h, because a predator cannot handle more than one prey per h. With a = 0.5, h = 0.1 and N = 20: f = 0.5 × 20/(1 + 0.5 × 0.1 × 20) = 10/2 = 5 prey per predator per unit time, half of the ceiling of 10 — the density at which f reaches half its maximum is N = 1/(ah) = 20.
2. The derivative as a rate: slopes of ecological curves
The derivative dy/dx is the slope of the tangent to a curve, the instantaneous rate of change of y with x; the rules of differentiation are in the shared derivatives chapter. In ecology the derivative is usually a rate: dN/dt is the population growth rate, and (1/N)dN/dt the per-capita rate. For the logistic model, dN/dt = rN − rN²/K is a quadratic in N; its derivative with respect to N, r − 2rN/K, is zero at N = K/2, where the growth rate is at its maximum rK/4 — positive slope below K/2, negative above. A maximum or minimum of any smooth curve is found in the same way: set the first derivative to zero, then check the sign of the second derivative (negative for a maximum).
3. Counting and probability in field sampling
| Situation | Count | Field example |
|---|---|---|
| Order all n distinct items | n! | 4 transects can be walked in 4! = 24 orders |
| Order r of n (permutations) | ⁿPᵣ = n!/(n − r)! | a first, second and third site to visit from 10: ¹⁰P₃ = 720 |
| Choose r of n, order irrelevant (combinations) | ⁿCᵣ = n!/[r!(n − r)!] | 3 ponds to sample from 10: ¹⁰C₃ = 120 |
Order matters for a permutation and not for a combination, so ⁿPᵣ = r! × ⁿCᵣ: each of the 120 possible sets of three ponds can be visited in 3! = 6 orders, giving 720. A "random sample of k sites" means that each of the ⁿCₖ possible sets is equally likely. The addition, multiplication and conditional-probability rules are in the shared probability chapter. Their most-used ecological application is imperfect detection. If a species present at a site is detected on any one visit with probability p, independently, then after k visits the probability it is detected at least once is 1 − (1 − p)^k: with p = 0.3 and four visits, 1 − 0.7⁴ = 1 − 0.2401 = 0.76. A site where the species was never seen may still be occupied, and Bayes’ theorem gives the probability that it is.
Worked Bayes example: suppose 40 % of sites are occupied (prior ψ = 0.4) and the species is missed on all four visits at an occupied site with probability 0.7⁴ = 0.2401. P(not detected) = 0.4 × 0.2401 + 0.6 × 1 = 0.696, and P(occupied | not detected) = 0.0960/0.696 = 0.138. So 14 % of "empty" sites are in fact occupied, and a naive count of detections underestimates occupancy. This is the logic of occupancy modelling.
4. Probability distributions for ecological data
The binomial distribution gives the probability of k "successes" in n independent trials each with probability p: P(k) = ⁿCₖ p^k (1 − p)n−k, with mean np and variance np(1 − p). If each of four seedlings survives with probability 0.5, the chance that exactly two survive is ⁴C₂ × 0.5⁴ = 6/16 = 0.375. The Poisson distribution describes counts of independent, randomly placed events per unit — individuals per quadrat under a random spatial pattern: P(k) = e−m m^k/k!, with mean = variance = m. The proportion of empty quadrats is e−m; with a mean of 1.5 individuals per quadrat, e−1.5 = 0.223 of quadrats should be empty if the pattern is random. The normal distribution describes continuous measurements such as body size: about 68 % of values lie within one standard deviation of the mean, 95 % within 1.96 SD and 99.7 % within 3 SD.
| Index of dispersion I = s²/x̄ | Pattern | Typical cause |
|---|---|---|
| I ≈ 1 | random (Poisson) | individuals settle independently in a uniform habitat |
| I > 1 | clumped (aggregated) | patchy resources, social grouping, limited seed dispersal |
| I < 1 | uniform (regular) | territoriality, competition between neighbours, allelopathy |
5. Descriptive statistics, standard error and confidence intervals
For a sample x₁, …, xₙ: the mean x̄ = Σxᵢ/n; the median is the middle value, robust to outliers; the sample variance s² = Σ(xᵢ − x̄)²/(n − 1), dividing by n − 1 because one degree of freedom is used to estimate the mean; and the standard deviation s = √s². For 2, 4, 4, 5, 5: x̄ = 4, the squared deviations are 4, 0, 0, 1, 1, their sum is 6 and s² = 6/4 = 1.5. The coefficient of variation CV = s/x̄ × 100 % compares variability between variables with different means or units: a mean body mass of 25 g with s = 5 g gives CV = 20 %. CV is also the standard way to compare the variability of population sizes through time.
The standard error of the mean SE = s/√n measures how precisely x̄ estimates the population mean μ; it shrinks as √n, so quadrupling the sample halves it. A 95 % confidence interval is x̄ ± t(0.025, n − 1) × SE, where t is from the t distribution with n − 1 degrees of freedom (1.96 from the normal for large n). With n = 16, x̄ = 50 and s = 8: SE = 8/4 = 2, t(0.025, 15) = 2.131, and the interval is 50 ± 4.26, from 45.74 to 54.26. Its meaning is about the procedure: 95 % of intervals built this way from repeated samples would contain μ; it is not a 95 % probability that μ lies in this particular interval.
Key takeaways
- Exponential growth is a straight line on a semi-log plot with slope r; a power law is a straight line on a log–log plot with slope equal to its exponent.
- The derivative is a rate; set it to zero to find the logistic maximum at N = K/2, and read optima off the slope of a curve.
- Detection after k visits is 1 − (1 − p)^k; Bayes’ theorem gives the chance that an undetected site is occupied.
- Poisson counts have variance equal to the mean; s²/x̄ ≈ 1 is random, > 1 clumped, < 1 uniform; the empty fraction is e−m.
- s² divides by n − 1; CV = s/x̄; SE = s/√n; a 95 % CI is x̄ ± t(0.025, n − 1) × SE and describes the procedure, not this one interval.
Practice questions (16)
Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.
A population’s natural logarithm, ln N, rises linearly from 4.0 to 5.2 over 6 years. Its intrinsic rate of increase r, per year, to one decimal place, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.2
On a semi-log plot exponential growth has slope r: (5.2 − 4.0)/6 = 1.2/6 = 0.2 per year. Dividing 5.2 by 4.0 and then by 6 treats the log values as the population sizes themselves.Basal metabolic rate scales as body mass to the power 0.75. An animal 16 times heavier than another has a basal metabolic rate greater by a factor of ____.
Numerical answer — type the value.
Show answer
Answer: 8
160.75 = (2⁴)3/4 = 2³ = 8. The heavier animal uses 8 times the energy but only 8/16 = half as much per gram. Answering 12 multiplies 16 by 0.75 instead of raising it to that power.A predator has a Holling Type II functional response f = aN/(1 + ahN) with attack rate a = 0.5 and handling time h = 0.1. At a prey density N = 20, the number of prey eaten per predator per unit time is ____.
Numerical answer — type the value.
Show answer
Answer: 5
f = 0.5 × 20/(1 + 0.5 × 0.1 × 20) = 10/2 = 5. Ignoring handling gives the Type I value aN = 10, which is also the ceiling 1/h that the curve only approaches at high density.Data on island area A and species number S give a straight line with slope 0.27 when plotted as log S against log A. This means that
Show answer
Answer: A — S = cA^{0.27}, a power law with exponent 0.27
A straight line on log–log axes is a power function, and its slope is the exponent: log S = log c + 0.27 log A gives S = cA0.27. A constant increase per km² would be linear on arithmetic axes, and an exponential would be linear on a semi-log plot of log S against A.For the logistic model, dN/dt = rN − rN²/K. The derivative of dN/dt with respect to N is zero at
Show answer
Answer: A — N = K/2
d/dN (rN − rN²/K) = r − 2rN/K = 0 gives N = K/2, the density of fastest growth. N = K and N = 0 are where dN/dt itself is zero, and rK/4 is the maximum value of the growth rate, not a density.A survey team must choose 3 of 10 ponds to sample, and the order of visiting them does not matter. The number of different sets of ponds it can choose is ____.
Numerical answer — type the value.
Show answer
Answer: 120
Order is irrelevant, so this is a combination: ¹⁰C₃ = 10 × 9 × 8/(3 × 2 × 1) = 120. The permutation ¹⁰P₃ = 720 counts each set six times, once for every order of visiting it.A frog species present at a pond is detected on any one survey with probability 0.3, independently between surveys. The probability that it is detected at least once in four surveys, to two decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.76
P(never detected) = 0.7⁴ = 0.2401, so P(at least once) = 1 − 0.2401 = 0.76. Adding 4 × 0.3 = 1.2 gives an impossible probability, because the events "detected on survey i" are not mutually exclusive.Forty per cent of ponds in a region are occupied by a frog. At an occupied pond the frog is missed on all four surveys with probability 0.24; at an unoccupied pond it is never detected. A pond where the frog was not detected in four surveys is occupied with probability, to two decimal places, ____.
Numerical answer — type the value.
Show answer
Answer: 0.14
P(not detected) = 0.4 × 0.24 + 0.6 × 1 = 0.096 + 0.6 = 0.696. P(occupied | not detected) = 0.096/0.696 = 0.14. Answering 0.24 gives P(not detected | occupied), the reverse conditional; 0.4 ignores the survey evidence altogether.Each of four tagged seedlings survives the dry season with probability 0.5, independently. The probability that exactly two survive, to three decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.375
Binomial: ⁴C₂ × 0.5² × 0.5² = 6 × 0.0625 = 0.375. Forgetting the ⁴C₂ = 6 arrangements gives 0.0625, the probability of one particular pair surviving and the other two dying.Individuals of a shrub are distributed at random, with a mean of 1.5 per quadrat. The expected proportion of empty quadrats, to three decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.223
A random pattern gives Poisson counts, and P(0) = e−m = e−1.5 = 0.223. Using 1/1.5 = 0.667 or 1 − e−1.5 = 0.777 (the proportion occupied) are the usual slips.Counts of a beetle in 50 quadrats have a mean of 4 and a variance of 12. The spatial pattern is best described as
Show answer
Answer: A — clumped, since the variance-to-mean ratio is 3
s²/x̄ = 12/4 = 3, well above 1, so the beetles are aggregated. A Poisson (random) pattern would have variance equal to the mean, and a uniform pattern variance below the mean, so a variance exceeding the mean points to clumping, not uniformity.The body masses of a sample of lizards have a mean of 25 g and a standard deviation of 5 g. The coefficient of variation, in per cent, is ____.
Numerical answer — type the value.
Show answer
Answer: 20
CV = s/x̄ × 100 = 5/25 × 100 = 20 %. It is unit-free, so it can compare the variability of masses with that of lengths; 500 is the inverse ratio times 100.The numbers of eggs in five nests are 2, 4, 4, 5 and 5. The sample variance (dividing by n − 1), to one decimal place, is ____.
Numerical answer — type the value.
Show answer
Answer: 1.5
x̄ = 20/5 = 4; squared deviations 4, 0, 0, 1, 1 sum to 6; s² = 6/(5 − 1) = 1.5. Dividing by n gives 1.2, the biased population-variance formula, and taking the square root gives the standard deviation 1.22.A sample of 16 fish has mean length 50 cm and standard deviation 8 cm. Using t(0.025, 15) = 2.131, the upper limit of the 95 % confidence interval for the mean, in cm to two decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 54.26
SE = 8/√16 = 2; half-width = 2.131 × 2 = 4.262; upper limit = 50 + 4.262 = 54.26. Using the SD instead of the SE gives 67.05, and using 1.96 instead of t gives 53.92, too narrow for n = 16.Which statements about the standard deviation (SD) and the standard error of the mean (SE) are correct?
Show answer
Answer: A — SE = SD/√n; B — Increasing the sample size fourfold roughly halves the SE; C — SD describes the spread of individual observations
SE = SD/√n, so a fourfold larger sample halves it, and it measures the precision of the mean. SD measures spread among individuals, a property of the population that a larger sample estimates better but does not shrink.A 95 % confidence interval for mean seed mass is computed as 12.1 to 14.3 mg. Which interpretation is correct?
Show answer
Answer: A — If sampling were repeated many times, about 95 % of intervals built this way would contain the true mean
The confidence level belongs to the procedure: in repeated samples, 95 % of such intervals cover μ. Once computed, the fixed interval either contains μ or not, so the second option is the common misreading. The interval is about the mean, not individual seeds, and the sample mean is always at its centre.