Estimation I: Sufficiency, Minimal Sufficiency, Completeness, Ancillarity and Basu, Unbiased Estimation, UMVUE and the Cramér–Rao Inequality
1. Sufficiency and the factorisation theorem
A statistic T(X) is sufficient for θ if the conditional distribution of the sample given T does not depend on θ: once T is known, the rest of the data carries no information about θ. The Neyman–Fisher factorisation theorem: T is sufficient iff f(x; θ) = g(T(x); θ) h(x), with h free of θ. Any one-to-one function of a sufficient statistic is sufficient.
| Model | Sufficient statistic |
|---|---|
| Bernoulli(p), Poisson(λ), Exp(λ) | ΣXᵢ |
| N(μ, σ²), σ known | ΣXᵢ |
| N(μ, σ²), both unknown | (ΣXᵢ, ΣXᵢ²), equivalently (X̄, S²) |
| U(0, θ) | X₍ₙ₎ |
| U(θ, θ + 1) | (X₍₁₎, X₍ₙ₎) |
| Gamma(α, β), both unknown | (ΣXᵢ, ∏Xᵢ) |
For U(0, θ): f(x; θ) = θ⁻ⁿ 1{x₍ₙ₎ ≤ θ} 1{x₍₁₎ ≥ 0} — the indicator of the parameter-dependent support must go into g, which is why the maximum, not the mean, is sufficient.
2. Minimal sufficiency
A sufficient statistic is minimal if it is a function of every other sufficient statistic — the greatest reduction of the data that loses nothing. The Lehmann–Scheffé criterion: T is minimal sufficient if, for all x and y, f(x; θ)/f(y; θ) is free of θ exactly when T(x) = T(y). For N(μ, σ²) the ratio is exp of a linear combination of Σxᵢ − Σyᵢ and Σxᵢ² − Σyᵢ², free of (μ, σ²) iff both differences vanish: (ΣX, ΣX²) is minimal.
The criterion also shows how little some models reduce. For the Cauchy location family, and for the Laplace location family, the minimal sufficient statistic is the whole set of order statistics — no one- or two-dimensional summary is sufficient. For U(θ, θ + 1) it is the pair (X₍₁₎, X₍ₙ₎): two numbers for one parameter.
3. Completeness, exponential families, ancillary statistics and Basu’s theorem
T is complete if E_θ[g(T)] = 0 for every θ forces g(T) = 0 with probability 1: no non-trivial function of T has mean zero for all θ, so an unbiased estimator that is a function of T is unique. For U(0, θ), X₍ₙ₎ is complete; for U(θ, θ + 1), (X₍₁₎, X₍ₙ₎) is not — the range X₍ₙ₎ − X₍₁₎ has a θ-free mean, so g = range − (n − 1)/(n + 1) has zero mean for every θ without being zero.
A k-parameter exponential family f(x; θ) = h(x) c(θ) exp(Σⱼ wⱼ(θ)Tⱼ(x)) has (ΣTⱼ(Xᵢ))ⱼ sufficient, and complete when the range of (w₁(θ), …, w_k(θ)) contains an open set in ℝᵏ (full rank). Binomial, Poisson, exponential, gamma, normal with one or both parameters unknown: all complete. N(θ, θ²) is curved — two statistics, one parameter — and not complete.
A statistic is ancillary if its distribution does not depend on θ: S² in N(μ, σ² known), X₍ₙ₎ − X₍₁₎ in a location family, X₍₁₎/X₍ₙ₎ in a scale family. Basu’s theorem: a complete sufficient statistic is independent of every ancillary statistic. Applications: X̄ and S² are independent in a normal sample (X̄ complete sufficient for μ with σ fixed, S² ancillary for μ); in U(0, θ), X₍ₙ₎ is independent of X₍₁₎/X₍ₙ₎; for i.i.d. exponentials, ΣXᵢ is independent of X₁/ΣXᵢ, which makes E[X₁/ΣXᵢ] = E X₁/E ΣXᵢ = 1/n.
4. Unbiasedness, UMVUE, Rao–Blackwell and Lehmann–Scheffé
T is unbiased for τ(θ) if E_θ T = τ(θ) for every θ; its MSE is Var T + bias², so a biased estimator can beat an unbiased one. A UMVUE has the smallest variance among unbiased estimators, uniformly in θ. Rao–Blackwell: if δ is unbiased and T sufficient, φ(T) = E[δ | T] is a statistic (sufficiency makes it free of θ), is unbiased, and Var φ ≤ Var δ. Lehmann–Scheffé: an unbiased estimator that is a function of a complete sufficient statistic is the unique UMVUE.
| Model and target | UMVUE |
|---|---|
| Poisson(λ), λ | X̄ |
| Poisson(λ), e−λ = P(X = 0) | (1 − 1/n)T, T = ΣXᵢ |
| Bernoulli(p), p(1 − p) | T(n − T)/[n(n − 1)] |
| U(0, θ), θ | (n + 1)X₍ₙ₎/n |
| N(μ, σ²), σ² | S² = Σ(Xᵢ − X̄)²/(n − 1) |
| Exp(mean θ), θ | X̄ |
The Poisson row is Rao–Blackwell in action: 1{X₁ = 0} is unbiased for e−λ, and given T = t, X₁ ~ Bin(t, 1/n), so E[1{X₁ = 0} | T] = (1 − 1/n)^T. With n = 5 and T = 10, the UMVUE is 0.8¹⁰ = 0.107; the MLE e−X̄ = e−2 = 0.135 is biased.
5. Fisher information and the Cramér–Rao inequality
Under regularity conditions (support free of θ, differentiation under the integral), the Fisher information per observation is I(θ) = E[(∂ ln f/∂θ)²] = −E[∂² ln f/∂θ²], and a sample of n carries nI(θ). The Cramér–Rao inequality: any unbiased T for τ(θ) has Var T ≥ [τ′(θ)]²/(nI(θ)). Equality holds iff the score ∂ ln L/∂θ is a linear function of T, which happens exactly in the one-parameter exponential family with T its natural statistic (up to linear transformation).
| Model | I(θ) | CR bound for unbiased estimators of θ |
|---|---|---|
| Bernoulli(p) | 1/[p(1 − p)] | p(1 − p)/n |
| Poisson(λ) | 1/λ | λ/n |
| N(μ, σ²), μ, σ known | 1/σ² | σ²/n |
| N(0, σ²), parameter σ² | 1/(2σ⁴) | 2σ⁴/n |
| Exp(mean θ) | 1/θ² | θ²/n |
Key takeaways
- Factorisation: f = g(T; θ)h(x). A parameter-dependent support puts an order statistic into T (X₍ₙ₎ for U(0, θ)).
- Minimal sufficiency: f(x)/f(y) free of θ ⇔ T(x) = T(y). Cauchy and Laplace location families need all the order statistics.
- Full-rank exponential families give complete sufficient statistics; U(θ, θ + 1) is the standard minimal-but-not-complete case.
- Basu: complete sufficient ⊥ ancillary — X̄ ⊥ S² in the normal, X₍ₙ₎ ⊥ X₍₁₎/X₍ₙ₎ in U(0, θ).
- Rao–Blackwell improves; Lehmann–Scheffé certifies: unbiased + function of complete sufficient = UMVUE. Poisson e−λ: (1 − 1/n)ΣX.
- Var T ≥ [τ′(θ)]²/(nI(θ)); I = 1/[p(1 − p)], 1/λ, 1/σ², 1/(2σ⁴); the bound needs a θ-free support.
Practice questions (14)
Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.
X₁, …, Xₙ are i.i.d. U(0, θ). A sufficient statistic for θ is:
Show answer
Answer: A — X₍ₙ₎ = max Xᵢ
L(θ) = θ⁻ⁿ 1{X₍ₙ₎ ≤ θ} depends on the data only through X₍ₙ₎, so by factorisation it is sufficient. X̄ is not: two samples with the same mean but different maxima give different likelihoods (one may even be zero).X₁, …, Xₙ (n ≥ 2) are i.i.d. U(θ, θ + 1). Which statements are true?
Show answer
Answer: A — (X₍₁₎, X₍ₙ₎) is minimal sufficient for θ; C — X₍ₙ₎ − X₍₁₎ is an ancillary statistic
(A) L(θ) = 1{X₍ₙ₎ − 1 ≤ θ ≤ X₍₁₎}; the ratio for two samples is θ-free iff both extremes agree. (C) Xᵢ − θ ~ U(0, 1), so the range has a θ-free distribution. (B) False: E[range] = (n − 1)/(n + 1) for every θ, so range − (n − 1)/(n + 1) is a non-zero function with zero mean. (D) False: the likelihood depends on the extremes, not the mean.X₁, …, X₁₀ are i.i.d. Poisson(λ). At λ = 2, the Fisher information about λ contained in the whole sample is ____.
Numerical answer — type the value.
Show answer
Answer: 5
ln f = x ln λ − λ − ln x!, ∂² ln f/∂λ² = −x/λ², so I(λ) = E X/λ² = 1/λ per observation and nI = 10/2 = 5. Reporting 0.5, the per-observation value, forgets the sample size; 0.2 is the Cramér–Rao bound λ/n, its reciprocal.X₁, …, X₂₅ are i.i.d. Bernoulli(p). At p = 0.2, the Cramér–Rao lower bound for the variance of unbiased estimators of p, correct to four decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.0064
I(p) = 1/[p(1 − p)], so the bound is p(1 − p)/n = 0.16/25 = 0.0064, attained by X̄. Forgetting n gives 0.16, the variance of a single observation.X₁, …, X₅ are i.i.d. Poisson(λ) and the observed total is ΣXᵢ = 10. The UMVUE of e−λ, evaluated at these data, correct to three decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.107
1{X₁ = 0} is unbiased for e−λ; given T = 10, X₁ ~ Bin(10, 1/5), so E[1{X₁ = 0} | T] = (4/5)¹⁰ = 0.1074 → 0.107. T is complete sufficient, so Lehmann–Scheffé makes this the UMVUE. The plug-in e−X̄ = e−2 = 0.135 is the MLE, which is biased.X₁, …, X₄ are i.i.d. U(0, θ) and the largest observation is 6. The UMVUE of θ, evaluated at these data, correct to one decimal place, is ____.
Numerical answer — type the value.
Show answer
Answer: 7.5
E X₍ₙ₎ = nθ/(n + 1), so (n + 1)X₍ₙ₎/n = (5/4) × 6 = 7.5 is unbiased and a function of the complete sufficient X₍ₙ₎. Answering 6, the MLE, is biased downwards — the maximum of a sample always falls short of θ.X₁, …, X₁₀ are i.i.d. Bernoulli(p) with ΣXᵢ = 4. The UMVUE of p(1 − p), evaluated at these data, correct to three decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.267
T = ΣXᵢ is complete sufficient, and E[T(n − T)] = n(n − 1)p(1 − p), so T(n − T)/[n(n − 1)] = 4 × 6/90 = 0.2667 → 0.267 is the UMVUE. The plug-in X̄(1 − X̄) = 0.24 is biased: its mean is (n − 1)/n × p(1 − p).Which statements about Basu’s theorem and its applications are true?
Show answer
Answer: A — In a N(μ, σ²) sample with σ² known, X̄ and S² are independent, by Basu’s theorem; B — In a U(0, θ) sample, X₍ₙ₎ and X₍₁₎/X₍ₙ₎ are independent; C — Basu’s theorem requires the sufficient statistic to be complete
(A) X̄ is complete sufficient for μ, and S² = Σ(Xᵢ − X̄)²/(n − 1) is location-invariant, hence ancillary. (B) X₍ₙ₎ is complete sufficient for θ, and X₍₁₎/X₍ₙ₎ is scale-invariant, hence ancillary. (C) Completeness is the hypothesis that makes it work. (D) False: the whole sample is sufficient, and it is not independent of, say, the range.X₁, …, Xₙ are i.i.d. N(μ, σ²) with both parameters unknown. A complete sufficient statistic for (μ, σ²) is:
Show answer
Answer: A — (ΣXᵢ, ΣXᵢ²)
The normal is a two-parameter exponential family with natural parameters (μ/σ², −1/(2σ²)), whose range is an open half-plane, so (ΣX, ΣX²) is complete and sufficient. ΣXᵢ alone is sufficient only when σ² is known; the extremes and the median are not sufficient at all.X ~ N(0, σ²). The Fisher information about the parameter σ² in one observation, at σ² = 2, correct to three decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.125
With v = σ²: ln f = −½ ln(2πv) − x²/(2v), ∂²/∂v² = 1/(2v²) − x²/v³, and −E[·] = −1/(2v²) + v/v³ = 1/(2v²) = 1/8 = 0.125. The information about σ itself is 2/σ² = 1; the two differ because the parametrisation changes the scale of the score.X₁, …, Xₙ are i.i.d. exponential with mean θ. Which unbiased estimator attains the Cramér–Rao lower bound?
Show answer
Answer: A — X̄, as an estimator of θ
The score is n(X̄ − θ)/θ², linear in X̄, so X̄ attains the bound θ²/n for θ. (n − 1)/ΣX is the UMVUE of 1/θ, but its variance exceeds the bound because it is not linear in X̄. nX₍₁₎ is unbiased (the minimum is Exp with mean θ/n) with variance θ² — n times the bound; X₁ likewise has variance θ².X₁, …, X₂₅ are i.i.d. N(μ, 25). The estimator T = X̄ + 1 is used for μ. The mean squared error of T is ____.
Numerical answer — type the value.
Show answer
Answer: 2
MSE = Var T + bias² = 25/25 + 1² = 2. Giving 1 reports only the variance and ignores the bias; giving 26 uses the variance of a single observation.δ is an unbiased estimator of τ(θ) and T is a sufficient statistic. Let φ(T) = E[δ | T]. Which statements are true?
Show answer
Answer: A — φ(T) is unbiased for τ(θ); B — Var φ(T) ≤ Var δ for every θ; C — φ(T) does not depend on θ, because T is sufficient
(A) Tower property: E φ(T) = E δ = τ(θ). (B) Var δ = Var φ + E[Var(δ | T)] ≥ Var φ. (C) Sufficiency makes the conditional distribution of the data given T, and so E[δ | T], free of θ — without it φ would not be a statistic. (D) False: that needs T to be complete as well (Lehmann–Scheffé).X₁, …, Xₙ (n ≥ 3) are i.i.d. from the Cauchy location family with density 1/(π[1 + (x − θ)²]). A minimal sufficient statistic for θ is:
Show answer
Answer: A — the vector of order statistics (X₍₁₎, …, X₍ₙ₎)
f(x; θ)/f(y; θ) = ∏[1 + (yᵢ − θ)²]/∏[1 + (xᵢ − θ)²], a ratio of polynomials in θ that is constant iff the two sets of roots {xᵢ ± i} and {yᵢ ± i} coincide, i.e. iff the samples are permutations of each other. So nothing coarser than the order statistics is sufficient; the median is a good estimator but discards information.