Estimation II: Consistency, the Method of Moments, Maximum Likelihood and Its Properties, and Confidence Intervals from Pivots

The second chapter for Section 8 of the GATE Statistics paper: consistent estimators, method of moments estimators, and maximum likelihood estimators and their properties — invariance, dependence on sufficient statistics, consistency, asymptotic normality and efficiency; then interval estimation — pivotal quantities, the confidence intervals built from them, and coverage probability. The numericals are an MLE or moment estimate computed from given data, or a confidence limit computed from a table value that the question states.

1. Consistent estimators

Tₙ is consistent for θ if Tₙ → θ in probability for every θ. A convenient sufficient condition: bias → 0 and variance → 0 (then MSE → 0, and Chebyshev gives consistency). X̄ is consistent for the mean by the weak law; S² is consistent for σ²; X₍ₙ₎ is consistent for θ in U(0, θ), since P(X₍ₙ₎ < θ − ε) = (1 − ε/θ)ⁿ → 0. If Tₙ is consistent for θ and g is continuous, g(Tₙ) is consistent for g(θ).

⚠️ Unbiased and consistent are independent properties
X₁ alone is unbiased for μ but not consistent — its distribution never tightens. X₍ₙ₎ is biased for θ in U(0, θ) but consistent. X̄ + 1/n is biased for every n and consistent. Neither property implies the other.

2. The method of moments

Equate the first k sample moments mⱼ = (1/n)ΣXᵢʲ to the population moments and solve for the k parameters. U(0, θ): E X = θ/2, so θ̂ = 2X̄. Gamma(α, rate λ): E X = α/λ and Var X = α/λ², so α̂ = X̄²/σ̂² and λ̂ = X̄/σ̂², with σ̂² the variance with divisor n — a sample with mean 6 and variance 12 gives α̂ = 3, λ̂ = 0.5. Moment estimators are consistent (continuous functions of consistent sample moments) and usually easy, but they need not be functions of sufficient statistics, can fall outside the parameter space (2X̄ can be smaller than the largest observation), and fail where moments do not exist (the Cauchy).

3. Maximum likelihood estimation and its properties

The MLE θ̂ maximises L(θ) = ∏f(xᵢ; θ), usually through the score equation ∂ ln L/∂θ = 0 — but only when the maximum is interior and smooth. Where the support depends on θ, maximise directly: for U(0, θ), L = θ⁻ⁿ for θ ≥ X₍ₙ₎ and 0 otherwise, a decreasing function on its range, so θ̂ = X₍ₙ₎. For U(θ − ½, θ + ½), L = 1 on the whole interval [X₍ₙ₎ − ½, X₍₁₎ + ½], and every point of it is an MLE — the MLE need not be unique.

Standard MLEs
ModelMLEUnbiased?
Bernoulli(p), Poisson(λ)X̄yes
Exp(rate λ)1/X̄no: E = nλ/(n − 1)
N(μ, σ²)X̄, (1/n)Σ(Xᵢ − X̄)²μ̂ yes, σ̂² no
U(0, θ)X₍ₙ₎no: E = nθ/(n + 1)
Laplace location, density ½e−|x − θ|sample medianyes, by symmetry (odd n)
  • Invariance: the MLE of g(θ) is g(θ̂) — for Poisson data with X̄ = 2, the MLE of P(X = 0) = e−λ is e−2 = 0.135.
  • Sufficiency: a unique MLE is a function of every sufficient statistic, because L depends on the data only through it.
  • Large-sample: under regularity, θ̂ is consistent and √n(θ̂ − θ) → N(0, 1/I(θ)) — asymptotically unbiased and attaining the Cramér–Rao bound, i.e. asymptotically efficient.

4. Interval estimation: pivots and coverage

A pivot Q(X, θ) is a function of the data and θ whose distribution does not depend on θ. If P(a ≤ Q ≤ b) = 1 − α, solving a ≤ Q ≤ b for θ gives a set with coverage probability P_θ(θ ∈ C(X)) = 1 − α for every θ; the infimum of coverage over θ is the confidence coefficient. The probability statement is about the random interval before the data are seen — a computed interval either contains θ or it does not.

Pivots and the intervals they give (level 1 − α)
SettingPivotInterval
Normal mean, σ known√n(X̄ − μ)/σ ~ N(0, 1)X̄ ± zα/2 σ/√n
Normal mean, σ unknown√n(X̄ − μ)/S ~ tn−1X̄ ± tn−1, α/2 S/√n
Normal variance(n − 1)S²/σ² ~ χ²n−1[(n − 1)S²/χ²upper, (n − 1)S²/χ²lower]
Exponential rate λ2λΣXᵢ ~ χ²2n[χ²2n, lower/(2ΣX), χ²2n, upper/(2ΣX)]
U(0, θ)X₍ₙ₎/θ, CDF uⁿ on (0, 1)[X₍ₙ₎, X₍ₙ₎/α1/n]

Worked: σ = 10, n = 25 and z = 1.96 give a 95% interval of length 2 × 1.96 × 2 = 7.84. With σ unknown, n = 16, X̄ = 50, S = 8 and t₁₅,₀.₀₂₅ = 2.131, the interval is 50 ± 4.262, upper limit 54.26. For U(0, θ) with n = 3, the interval [X₍₃₎, X₍₃₎/c] has coverage 1 − c³, so coverage 0.875 needs c = 0.5: with X₍₃₎ = 4, the interval is [4, 8].

Key takeaways

  • Bias → 0 and variance → 0 give consistency; unbiasedness and consistency are separate properties (X₁ vs X₍ₙ₎).
  • Method of moments: U(0, θ) gives 2X̄; gamma gives α̂ = X̄²/σ̂². Simple and consistent, not always sensible.
  • MLE: X₍ₙ₎ for U(0, θ) (not from the score), (1/n)Σ(X − X̄)² for σ², the median for a Laplace location; U(θ ± ½) has a whole interval of MLEs.
  • MLEs are invariant, depend on the data through sufficient statistics, and are asymptotically N(θ, 1/(nI(θ))).
  • Pivots: √n(X̄ − μ)/σ, √n(X̄ − μ)/S, (n − 1)S²/σ², 2λΣX, X₍ₙ₎/θ. Coverage is a property of the procedure, not of one computed interval.

Practice questions (14)

Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.

  1. A Poisson(λ) sample is 2, 0, 3, 1, 4. The maximum likelihood estimate of λ is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 2

    ln L = Σxᵢ ln λ − nλ + const; the score Σxᵢ/λ − n = 0 gives λ̂ = X̄ = 10/5 = 2. The sample variance, 2.5, is another estimate of λ (the Poisson has mean = variance) but not the MLE.
  2. For the Poisson sample 2, 0, 3, 1, 4, the maximum likelihood estimate of P(X = 0), correct to three decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.135

    By invariance, the MLE of e−λ is e−λ̂ = e−2 = 0.1353 → 0.135. The observed proportion of zeros, 1/5 = 0.2, is a different (unbiased but less efficient) estimate, and the UMVUE (1 − 1/5)¹⁰ = 0.107 is a third.
  3. A U(0, θ) sample is 1.2, 3.4, 2.0, 0.8, 2.6. Which statements are true?

    1. The method of moments estimate of θ is 4
    2. The maximum likelihood estimate of θ is 3.4
    3. The maximum likelihood estimator of θ is unbiased
    4. The method of moments estimator of θ is unbiased
    Show answer

    Answer: A — The method of moments estimate of θ is 4; B — The maximum likelihood estimate of θ is 3.4; D — The method of moments estimator of θ is unbiased

    (A) X̄ = 10/5 = 2 and θ̂ = 2X̄ = 4. (B) L = θ⁻⁵ for θ ≥ 3.4, decreasing, so θ̂ = X₍₅₎ = 3.4. (C) False: E X₍ₙ₎ = nθ/(n + 1) = 5θ/6. (D) E[2X̄] = 2 × θ/2 = θ. Here the two estimators disagree by 0.6, and the MLE is the one that can never fall below the largest observation.
  4. A N(μ, σ²) sample is 2, 4, 4, 4, 5, 5, 7, 9. The maximum likelihood estimate of σ² is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 4

    X̄ = 40/8 = 5, Σ(xᵢ − 5)² = 9 + 1 + 1 + 1 + 0 + 0 + 4 + 16 = 32, σ̂² = 32/8 = 4. Dividing by n − 1 gives S² = 4.571, the unbiased estimate, which is not the MLE.
  5. A sample from a Gamma(α, λ) distribution (rate λ) has mean 6 and variance 12 (computed with divisor n). The method of moments estimate of α is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 3

    α/λ = 6 and α/λ² = 12 give λ = 6/12 = 0.5 and α = 6 × 0.5 = 3; equivalently α̂ = X̄²/σ̂² = 36/12 = 3. Answering 0.5 gives the rate, not the shape.
  6. A random sample of 16 from a normal population has X̄ = 50 and S = 8. Using t₁₅,₀.₀₂₅ = 2.131, the upper limit of the 95% confidence interval for μ, correct to two decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 54.26

    X̄ + t S/√n = 50 + 2.131 × 8/4 = 50 + 4.262 = 54.262 → 54.26. Using z = 1.96 gives 53.92, which ignores that σ was estimated; dividing by 16 rather than √16 gives 51.07.
  7. For a normal population with σ = 10, a sample of size 25 is taken. Using z₀.₀₂₅ = 1.96, the length of the 95% confidence interval for μ, correct to two decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 7.84

    Length = 2 z σ/√n = 2 × 1.96 × 10/5 = 7.84, whatever X̄ turns out to be. The half-width 3.92 answers a different question.
  8. X₁, X₂, X₃ are i.i.d. U(0, θ). The interval [X₍₃₎, X₍₃₎/c] is to have coverage probability 0.875. If X₍₃₎ = 4, the upper endpoint of the interval is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 8

    The pivot X₍₃₎/θ has CDF u³ on (0, 1). θ ∈ [X₍₃₎, X₍₃₎/c] iff c ≤ X₍₃₎/θ ≤ 1, with probability 1 − c³ = 0.875, so c³ = 0.125, c = 0.5 and the upper endpoint is 4/0.5 = 8. Using c = 0.875 treats the coverage as the cut-off itself.
  9. Which statements about maximum likelihood estimators are true (assuming the usual regularity conditions where needed)?

    1. If θ̂ is the MLE of θ, then g(θ̂) is the MLE of g(θ)
    2. The MLE is always unbiased
    3. A unique MLE is a function of any sufficient statistic
    4. √n(θ̂ − θ) converges in distribution to N(0, 1/I(θ))
    Show answer

    Answer: A — If θ̂ is the MLE of θ, then g(θ̂) is the MLE of g(θ); C — A unique MLE is a function of any sufficient statistic; D — √n(θ̂ − θ) converges in distribution to N(0, 1/I(θ))

    (A) Invariance. (B) False: the normal-variance MLE has mean (n − 1)σ²/n and the U(0, θ) MLE has mean nθ/(n + 1). (C) By factorisation L(θ) = g(T; θ)h(x), so maximising L is maximising g(T; θ). (D) The asymptotic normality and efficiency of the MLE, with I the per-observation Fisher information.
  10. X₁, …, Xₙ are i.i.d. U(θ − 1/2, θ + 1/2). The maximum likelihood estimator of θ is:

    1. any value in [X₍ₙ₎ − 1/2, X₍₁₎ + 1/2]
    2. uniquely X̄
    3. uniquely X₍ₙ₎ − 1/2
    4. uniquely the sample median
    Show answer

    Answer: A — any value in [X₍ₙ₎ − 1/2, X₍₁₎ + 1/2]

    L(θ) = 1 when every Xᵢ ∈ [θ − 1/2, θ + 1/2], i.e. X₍ₙ₎ − 1/2 ≤ θ ≤ X₍₁₎ + 1/2, and 0 otherwise. L is flat on that interval, so every point of it maximises — the MLE is not unique. The midrange (X₍₁₎ + X₍ₙ₎)/2 is one natural choice from it; X̄ need not even lie in it.
  11. A 95% confidence interval for μ computed from one sample is (48.1, 53.9). The correct interpretation is:

    1. The procedure that produced it covers μ in 95% of repeated samples
    2. P(48.1 < μ < 53.9) = 0.95
    3. 95% of the observations lie in (48.1, 53.9)
    4. X̄ lies in (48.1, 53.9) with probability 0.95
    Show answer

    Answer: A — The procedure that produced it covers μ in 95% of repeated samples

    μ is a fixed constant, so once the endpoints are numbers, the interval either contains it or not; the 95% is the coverage probability of the random interval before the data are seen. The interval is about μ, not about individual observations, and X̄ is its centre, contained with certainty.
  12. Which statements about consistency are true?

    1. X₍ₙ₎ is a consistent estimator of θ for a U(0, θ) sample
    2. X₁ is a consistent estimator of the population mean
    3. S² is a consistent estimator of σ² when the fourth moment is finite
    4. An estimator whose bias and variance both tend to 0 is consistent
    Show answer

    Answer: A — X₍ₙ₎ is a consistent estimator of θ for a U(0, θ) sample; C — S² is a consistent estimator of σ² when the fourth moment is finite; D — An estimator whose bias and variance both tend to 0 is consistent

    (A) P(|X₍ₙ₎ − θ| > ε) = (1 − ε/θ)ⁿ → 0. (B) False: X₁ is unbiased but its distribution does not depend on n, so it never concentrates. (C) S² = n/(n − 1) × [mean of X² − X̄²], continuous in consistent sample moments (the finite fourth moment even gives Var S² → 0). (D) MSE → 0, then Chebyshev.
  13. X₁, …, X₅ are i.i.d. exponential with rate λ and ΣXᵢ = 10. Using the pivot 2λΣXᵢ ~ χ²₁₀ and the lower 2.5% point χ²₁₀,₀.₉₇₅ = 3.247, the lower limit of the equal-tailed 95% confidence interval for λ, correct to three decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.162

    P(3.247 ≤ 2λΣX ≤ χ²_upper) = 0.95, so λ ≥ 3.247/(2 × 10) = 0.16235 → 0.162. The pivot has 2n = 10 degrees of freedom because each 2λXᵢ is χ²₂; using n = 5 degrees of freedom is the usual slip.
  14. X₁, …, X₅ are i.i.d. with density ½e−|x − θ|. The observations are 3, 9, 1, 7, 4. The maximum likelihood estimate of θ is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 4

    Maximising L means minimising Σ|xᵢ − θ|, which the sample median does: ordered data 1, 3, 4, 7, 9 give θ̂ = 4. The mean, 4.8, minimises Σ(xᵢ − θ)², which is the normal likelihood, not this one.