Testing of Hypotheses I: Size, Power and p-values, the Neyman–Pearson Lemma, Monotone Likelihood Ratio and Uniformly Most Powerful Tests

Section 9 of the GATE Statistics paper is the optimality theory of tests, and it divides naturally in two, so it has two chapters. This first one builds from the definitions — hypotheses, critical regions, the two errors, size, level, power and p-values, and randomised tests for discrete data — to the Neyman–Pearson lemma and most powerful tests, the monotone likelihood ratio (MLR) property, and uniformly most powerful (UMP) tests for families with MLR. The second chapter takes the cases where no UMP test exists: UMP unbiased tests, likelihood ratio tests and large-sample tests. Every p-value here is computed from a table value that the question itself states.

1. Hypotheses, errors, size, power and p-values

A hypothesis is simple if it specifies the distribution completely (θ = θ₀) and composite otherwise (θ ≤ θ₀). A test is a critical region C, or more generally a test function φ(x) ∈ [0, 1], the probability of rejecting H₀ at x. Rejecting a true H₀ is a type I error, accepting a false one a type II error. The power function is β(θ) = E_θφ(X) = P_θ(reject H₀); the size is supθ∈Θ₀ β(θ), and a test is of level α if its size is at most α. Power at θ ∈ Θ₁ is 1 − P(type II error).

The p-value is the smallest level at which the observed data would reject: for a right-tailed Z-test, p = 1 − Φ(z_obs); for a two-sided test, p = 2[1 − Φ(|z_obs|)]. With z = 2.33 and Φ(2.33) = 0.9901, the one-sided p-value is 0.0099 and the two-sided 0.0198. A p-value is computed assuming H₀ is true; it is not the probability that H₀ is true.

Discrete statistics cannot hit every size exactly, so a randomised test rejects outright on part of the boundary and with probability γ at the boundary point. X ~ Bin(5, p), H₀: p = 1/2 against p > 1/2, size 0.05: P(X = 5) = 1/32 = 0.03125 and P(X = 4) = 5/32, so reject at X = 5 and with probability γ = (0.05 − 0.03125)/0.15625 = 0.12 at X = 4.

2. The Neyman–Pearson lemma and most powerful tests

For simple against simple, H₀: f = f₀ against H₁: f = f₁, the Neyman–Pearson lemma says the test that rejects when f₁(x)/f₀(x) > k, accepts when it is < k, and randomises on equality, with k (and γ) chosen to give size α, is most powerful of level α; and every most powerful test has this form (almost everywhere). The work is always the same: write the likelihood ratio, show it is monotone in some statistic T, turn "ratio > k" into "T > c", and choose c from the null distribution of T.

  • Normal mean: N(μ, σ²) with σ known, H₀: μ = μ₀ against μ = μ₁ > μ₀. The ratio is increasing in X̄, so the MP test rejects when X̄ > μ₀ + z_α σ/√n. With μ₀ = 10, μ₁ = 12, σ/√n = 1 and α = 0.05, c = 11.645 and the power is P(Z > −0.355) = Φ(0.355) = 0.6387.
  • A single observation: H₀: X ~ U(0, 1) against H₁: f₁(x) = 2x. The ratio 2x is increasing, so reject when X > c; size 0.05 gives c = 0.95 and power ∫0.95^1 2x dx = 1 − 0.95² = 0.0975.
  • Binomial: Bin(10, p), H₀: p = 0.5 against p = 0.8, rejecting for X ≥ 9: size = (10 + 1)/1024 = 0.0107, power = 10(0.8)⁹(0.2) + 0.8¹⁰ = 0.376.
🎯 Why a most powerful test is never below its size
The trivial test φ ≡ α has size α and power α. A most powerful test must do at least as well, so its power is at least α — an MP test is automatically unbiased. The same comparison is the first step in proving the lemma.

3. The monotone likelihood ratio property

A family {f(x; θ)} has MLR in T(x) if for every θ₁ < θ₂ the ratio f(x; θ₂)/f(x; θ₁) is a non-decreasing function of T(x). A one-parameter exponential family f = h(x)c(θ)ew(θ)T(x) has MLR in T whenever w is increasing. So the binomial (in ΣX), Poisson (ΣX), normal mean with σ known (X̄), normal variance with μ known (Σ(X − μ)²), and exponential with mean θ (ΣX) all have MLR. U(0, θ) is not an exponential family but has MLR in X₍ₙ₎: the ratio is (θ₁/θ₂)ⁿ for X₍ₙ₎ ≤ θ₁ and ∞ for θ₁ < X₍ₙ₎ ≤ θ₂. The Cauchy location family does not have MLR: its ratio rises and then falls back to 1 as x → ±∞.

4. Uniformly most powerful tests

A level-α test is UMP for H₀ against a composite H₁ if it is most powerful against every θ ∈ Θ₁ at once. Karlin–Rubin: if the family has MLR in T, the test that rejects when T > c (randomising at T = c), with Pθ₀(reject) = α, is UMP of level α for H₀: θ ≤ θ₀ against H₁: θ > θ₀, and its power function is non-decreasing. The reason: the NP test of θ₀ against any θ₁ > θ₀ is "T > c" with the same c, so one test is best against all of them, and monotonicity of the power keeps the size at θ₀.

  • U(0, θ): H₀: θ ≤ 1 against θ > 1, reject when X₍ₙ₎ > c with cⁿ = 1 − α. For n = 5 and α = 0.05, the power at θ = 2 is 1 − (c/2)⁵ = 1 − 0.95/32 = 0.970.
  • Poisson: X ~ Poisson(λ), H₀: λ ≤ 1 against λ > 1, rejecting for X ≥ 3: size = 1 − e−1(1 + 1 + 1/2) = 0.080.
  • Exponential with mean θ: H₀: θ ≤ θ₀ against θ > θ₀, reject when ΣXᵢ > c, with 2ΣXᵢ/θ₀ ~ χ²2n giving c.
⚠️ No UMP test against a two-sided alternative
For N(μ, 1), H₀: μ = 0 against H₁: μ ≠ 0, the MP test against μ > 0 rejects for large X̄ and the MP test against μ < 0 for small X̄. No single test is best on both sides, so no UMP test exists. The standard remedy is to restrict the class — to unbiased tests — which is where the next chapter begins.

Key takeaways

  • Size = sup of the power over H₀; power on H₁ = 1 − P(type II); the p-value assumes H₀ and is not P(H₀ true).
  • Discrete data need randomisation at the boundary: γ = (α − P(beyond))/P(boundary).
  • Neyman–Pearson: reject when f₁/f₀ > k. Turn the ratio into a statistic and a cut-off from its null distribution.
  • MLR: exponential families with increasing w(θ), and U(0, θ) in X₍ₙ₎; the Cauchy location family has none.
  • Karlin–Rubin: MLR in T ⇒ "reject for T > c" is UMP for one-sided hypotheses; two-sided alternatives have no UMP test.

Practice questions (13)

Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.

  1. X̄ is the mean of 25 observations from N(μ, 25). To test H₀: μ = 10 against H₁: μ = 12, the test rejects H₀ when X̄ > 11.645 (size 0.05). Using Φ(0.355) = 0.6387, the power of the test, correct to four decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.6387

    Under μ = 12, X̄ ~ N(12, 1), so power = P(X̄ > 11.645) = P(Z > −0.355) = Φ(0.355) = 0.6387. Answering 1 − 0.6387 = 0.3613 gives the type II error probability instead.
  2. X ~ Bin(10, p). The test of H₀: p = 0.5 rejects H₀ when X ≥ 9. Its size, correct to four decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.0107

    P(X ≥ 9 | p = 0.5) = [C(10, 9) + C(10, 10)]/2¹⁰ = 11/1024 = 0.01074 → 0.0107. Counting only X = 10 gives 1/1024 = 0.001.
  3. For the test that rejects H₀: p = 0.5 when X ≥ 9, with X ~ Bin(10, p), the power at p = 0.8, correct to three decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.376

    P(X ≥ 9 | p = 0.8) = 10(0.8)⁹(0.2) + (0.8)¹⁰ = 0.8⁹(2 + 0.8) = 0.13422 × 2.8 = 0.3758 → 0.376. Stopping at 0.8¹⁰ = 0.107 omits the outcome X = 9.
  4. A two-sided Z-test gives an observed statistic z = 2.33. Using Φ(2.33) = 0.9901, the p-value, correct to four decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.0198

    p = 2[1 − Φ(2.33)] = 2 × 0.0099 = 0.0198. The one-sided value 0.0099 is half of it, and 0.9901 is the probability below the observed value, not beyond it.
  5. X₁, …, X₅ are i.i.d. U(0, θ). The UMP size-0.05 test of H₀: θ ≤ 1 against H₁: θ > 1 rejects H₀ when X₍₅₎ > c. The power of this test at θ = 2, correct to three decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.970

    Size: Pθ=1(X₍₅₎ > c) = 1 − c⁵ = 0.05, so c⁵ = 0.95. Power at θ = 2: 1 − (c/2)⁵ = 1 − 0.95/32 = 0.9703 → 0.970. There is no need to compute c itself; computing (c/2) as 0.95/2 before raising to the fifth power is the slip to avoid.
  6. Which families have the monotone likelihood ratio property in the statistic given?

    1. Poisson(λ), in ΣXᵢ
    2. U(0, θ), in X₍ₙ₎
    3. Cauchy location family, in a single observation X
    4. N(0, σ²), in ΣXᵢ²
    Show answer

    Answer: A — Poisson(λ), in ΣXᵢ; B — U(0, θ), in X₍ₙ₎; D — N(0, σ²), in ΣXᵢ²

    (A) The ratio is (λ₂/λ₁)ΣX e−n(λ₂−λ₁), increasing in ΣX for λ₂ > λ₁. (B) (θ₁/θ₂)ⁿ while X₍ₙ₎ ≤ θ₁, then ∞ — non-decreasing. (C) False: [1 + (x − θ₁)²]/[1 + (x − θ₂)²] tends to 1 at both ends and peaks in between. (D) exp[(1/(2σ₁²) − 1/(2σ₂²))ΣX²] up to a constant, increasing in ΣX² for σ₂ > σ₁.
  7. X₁, …, Xₙ are i.i.d. N(μ, 1). For H₀: μ = 0 against H₁: μ ≠ 0 at level α:

    1. no UMP test exists
    2. the test rejecting for X̄ > z_α/√n is UMP
    3. the test rejecting for |X̄| > zα/2/√n is UMP
    4. every level-α test has the same power
    Show answer

    Answer: A — no UMP test exists

    Against μ > 0 the unique MP test rejects for large X̄, against μ < 0 for small X̄; a test cannot coincide with both. The one-sided test has power below α for μ < 0. The two-sided test is UMP only within the unbiased tests (UMPU), not among all level-α tests.
  8. A single observation X is used to test H₀: X ~ U(0, 1) against H₁: X has density 2x on (0, 1). The most powerful test of size 0.05 has power, correct to four decimal places, ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.0975

    f₁/f₀ = 2x is increasing, so the MP test rejects when X > c; P₀(X > c) = 1 − c = 0.05 gives c = 0.95. Power = P₁(X > 0.95) = 1 − 0.95² = 0.0975 — barely above the size, because one observation carries little information. Rejecting for small X gives power 0.0025, far below α.
  9. Which statements about tests are true?

    1. The size of a test is the supremum of its power function over the null parameter set
    2. The p-value is the probability that H₀ is true
    3. At θ in the alternative, power = 1 − P(type II error)
    4. A test of level α has size at most α
    Show answer

    Answer: A — The size of a test is the supremum of its power function over the null parameter set; C — At θ in the alternative, power = 1 − P(type II error); D — A test of level α has size at most α

    (A), (C), (D) are the definitions. (B) False: the p-value is P(data at least as extreme as observed | H₀), a probability computed under H₀, not a probability of H₀ — in the frequentist framework H₀ is not random at all.
  10. X ~ Bin(5, p). A randomised test of H₀: p = 1/2 against H₁: p > 1/2 rejects when X = 5 and rejects with probability γ when X = 4. The value of γ that makes the size exactly 0.05, correct to two decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.12

    P(X = 5) + γP(X = 4) = 1/32 + γ(5/32) = 0.05 gives γ = (0.05 − 0.03125)/0.15625 = 0.12. Rejecting at X ≥ 4 outright would give size 6/32 = 0.1875, far too large; rejecting only at 5 gives 0.03125, too small.
  11. X₁, …, Xₙ are i.i.d. exponential with mean θ. The UMP level-α test of H₀: θ ≤ θ₀ against H₁: θ > θ₀ rejects H₀ when:

    1. ΣXᵢ > c
    2. ΣXᵢ < c
    3. X₍₁₎ > c
    4. |ΣXᵢ − nθ₀| > c
    Show answer

    Answer: A — ΣXᵢ > c

    f = θ⁻ⁿ e−ΣX/θ: w(θ) = −1/θ is increasing, so the family has MLR in ΣX and Karlin–Rubin gives "reject for large ΣX", with c from 2ΣX/θ₀ ~ χ²2n. A large mean makes large observations likelier, so rejecting for small ΣX points the wrong way; the minimum is not sufficient.
  12. X ~ Poisson(λ). The test of H₀: λ = 1 that rejects when X ≥ 3 has size, correct to three decimal places, ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.080

    P(X ≥ 3) = 1 − e−1(1 + 1 + 1/2) = 1 − 2.5 × 0.36788 = 0.0803 → 0.080. Omitting the x = 2 term 1/2! gives 1 − 2e−1 = 0.264, the size of "reject when X ≥ 2".
  13. A type I error is:

    1. rejecting H₀ when H₀ is true
    2. accepting H₀ when H₀ is false
    3. rejecting H₀ when H₁ is true
    4. accepting H₁ when H₁ is true
    Show answer

    Answer: A — rejecting H₀ when H₀ is true

    A type I error is a false rejection; its probability, maximised over H₀, is the size. Accepting a false H₀ is the type II error; rejecting H₀ when H₁ is true is the correct decision whose probability is the power.