Testing of Hypotheses I: Size, Power and p-values, the Neyman–Pearson Lemma, Monotone Likelihood Ratio and Uniformly Most Powerful Tests
1. Hypotheses, errors, size, power and p-values
A hypothesis is simple if it specifies the distribution completely (θ = θ₀) and composite otherwise (θ ≤ θ₀). A test is a critical region C, or more generally a test function φ(x) ∈ [0, 1], the probability of rejecting H₀ at x. Rejecting a true H₀ is a type I error, accepting a false one a type II error. The power function is β(θ) = E_θφ(X) = P_θ(reject H₀); the size is supθ∈Θ₀ β(θ), and a test is of level α if its size is at most α. Power at θ ∈ Θ₁ is 1 − P(type II error).
The p-value is the smallest level at which the observed data would reject: for a right-tailed Z-test, p = 1 − Φ(z_obs); for a two-sided test, p = 2[1 − Φ(|z_obs|)]. With z = 2.33 and Φ(2.33) = 0.9901, the one-sided p-value is 0.0099 and the two-sided 0.0198. A p-value is computed assuming H₀ is true; it is not the probability that H₀ is true.
Discrete statistics cannot hit every size exactly, so a randomised test rejects outright on part of the boundary and with probability γ at the boundary point. X ~ Bin(5, p), H₀: p = 1/2 against p > 1/2, size 0.05: P(X = 5) = 1/32 = 0.03125 and P(X = 4) = 5/32, so reject at X = 5 and with probability γ = (0.05 − 0.03125)/0.15625 = 0.12 at X = 4.
2. The Neyman–Pearson lemma and most powerful tests
For simple against simple, H₀: f = f₀ against H₁: f = f₁, the Neyman–Pearson lemma says the test that rejects when f₁(x)/f₀(x) > k, accepts when it is < k, and randomises on equality, with k (and γ) chosen to give size α, is most powerful of level α; and every most powerful test has this form (almost everywhere). The work is always the same: write the likelihood ratio, show it is monotone in some statistic T, turn "ratio > k" into "T > c", and choose c from the null distribution of T.
- Normal mean: N(μ, σ²) with σ known, H₀: μ = μ₀ against μ = μ₁ > μ₀. The ratio is increasing in X̄, so the MP test rejects when X̄ > μ₀ + z_α σ/√n. With μ₀ = 10, μ₁ = 12, σ/√n = 1 and α = 0.05, c = 11.645 and the power is P(Z > −0.355) = Φ(0.355) = 0.6387.
- A single observation: H₀: X ~ U(0, 1) against H₁: f₁(x) = 2x. The ratio 2x is increasing, so reject when X > c; size 0.05 gives c = 0.95 and power ∫0.95^1 2x dx = 1 − 0.95² = 0.0975.
- Binomial: Bin(10, p), H₀: p = 0.5 against p = 0.8, rejecting for X ≥ 9: size = (10 + 1)/1024 = 0.0107, power = 10(0.8)⁹(0.2) + 0.8¹⁰ = 0.376.
3. The monotone likelihood ratio property
A family {f(x; θ)} has MLR in T(x) if for every θ₁ < θ₂ the ratio f(x; θ₂)/f(x; θ₁) is a non-decreasing function of T(x). A one-parameter exponential family f = h(x)c(θ)ew(θ)T(x) has MLR in T whenever w is increasing. So the binomial (in ΣX), Poisson (ΣX), normal mean with σ known (X̄), normal variance with μ known (Σ(X − μ)²), and exponential with mean θ (ΣX) all have MLR. U(0, θ) is not an exponential family but has MLR in X₍ₙ₎: the ratio is (θ₁/θ₂)ⁿ for X₍ₙ₎ ≤ θ₁ and ∞ for θ₁ < X₍ₙ₎ ≤ θ₂. The Cauchy location family does not have MLR: its ratio rises and then falls back to 1 as x → ±∞.
4. Uniformly most powerful tests
A level-α test is UMP for H₀ against a composite H₁ if it is most powerful against every θ ∈ Θ₁ at once. Karlin–Rubin: if the family has MLR in T, the test that rejects when T > c (randomising at T = c), with Pθ₀(reject) = α, is UMP of level α for H₀: θ ≤ θ₀ against H₁: θ > θ₀, and its power function is non-decreasing. The reason: the NP test of θ₀ against any θ₁ > θ₀ is "T > c" with the same c, so one test is best against all of them, and monotonicity of the power keeps the size at θ₀.
- U(0, θ): H₀: θ ≤ 1 against θ > 1, reject when X₍ₙ₎ > c with cⁿ = 1 − α. For n = 5 and α = 0.05, the power at θ = 2 is 1 − (c/2)⁵ = 1 − 0.95/32 = 0.970.
- Poisson: X ~ Poisson(λ), H₀: λ ≤ 1 against λ > 1, rejecting for X ≥ 3: size = 1 − e−1(1 + 1 + 1/2) = 0.080.
- Exponential with mean θ: H₀: θ ≤ θ₀ against θ > θ₀, reject when ΣXᵢ > c, with 2ΣXᵢ/θ₀ ~ χ²2n giving c.
Key takeaways
- Size = sup of the power over H₀; power on H₁ = 1 − P(type II); the p-value assumes H₀ and is not P(H₀ true).
- Discrete data need randomisation at the boundary: γ = (α − P(beyond))/P(boundary).
- Neyman–Pearson: reject when f₁/f₀ > k. Turn the ratio into a statistic and a cut-off from its null distribution.
- MLR: exponential families with increasing w(θ), and U(0, θ) in X₍ₙ₎; the Cauchy location family has none.
- Karlin–Rubin: MLR in T ⇒ "reject for T > c" is UMP for one-sided hypotheses; two-sided alternatives have no UMP test.
Practice questions (13)
Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.
X̄ is the mean of 25 observations from N(μ, 25). To test H₀: μ = 10 against H₁: μ = 12, the test rejects H₀ when X̄ > 11.645 (size 0.05). Using Φ(0.355) = 0.6387, the power of the test, correct to four decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.6387
Under μ = 12, X̄ ~ N(12, 1), so power = P(X̄ > 11.645) = P(Z > −0.355) = Φ(0.355) = 0.6387. Answering 1 − 0.6387 = 0.3613 gives the type II error probability instead.X ~ Bin(10, p). The test of H₀: p = 0.5 rejects H₀ when X ≥ 9. Its size, correct to four decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.0107
P(X ≥ 9 | p = 0.5) = [C(10, 9) + C(10, 10)]/2¹⁰ = 11/1024 = 0.01074 → 0.0107. Counting only X = 10 gives 1/1024 = 0.001.For the test that rejects H₀: p = 0.5 when X ≥ 9, with X ~ Bin(10, p), the power at p = 0.8, correct to three decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.376
P(X ≥ 9 | p = 0.8) = 10(0.8)⁹(0.2) + (0.8)¹⁰ = 0.8⁹(2 + 0.8) = 0.13422 × 2.8 = 0.3758 → 0.376. Stopping at 0.8¹⁰ = 0.107 omits the outcome X = 9.A two-sided Z-test gives an observed statistic z = 2.33. Using Φ(2.33) = 0.9901, the p-value, correct to four decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.0198
p = 2[1 − Φ(2.33)] = 2 × 0.0099 = 0.0198. The one-sided value 0.0099 is half of it, and 0.9901 is the probability below the observed value, not beyond it.X₁, …, X₅ are i.i.d. U(0, θ). The UMP size-0.05 test of H₀: θ ≤ 1 against H₁: θ > 1 rejects H₀ when X₍₅₎ > c. The power of this test at θ = 2, correct to three decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.970
Size: Pθ=1(X₍₅₎ > c) = 1 − c⁵ = 0.05, so c⁵ = 0.95. Power at θ = 2: 1 − (c/2)⁵ = 1 − 0.95/32 = 0.9703 → 0.970. There is no need to compute c itself; computing (c/2) as 0.95/2 before raising to the fifth power is the slip to avoid.Which families have the monotone likelihood ratio property in the statistic given?
Show answer
Answer: A — Poisson(λ), in ΣXᵢ; B — U(0, θ), in X₍ₙ₎; D — N(0, σ²), in ΣXᵢ²
(A) The ratio is (λ₂/λ₁)ΣX e−n(λ₂−λ₁), increasing in ΣX for λ₂ > λ₁. (B) (θ₁/θ₂)ⁿ while X₍ₙ₎ ≤ θ₁, then ∞ — non-decreasing. (C) False: [1 + (x − θ₁)²]/[1 + (x − θ₂)²] tends to 1 at both ends and peaks in between. (D) exp[(1/(2σ₁²) − 1/(2σ₂²))ΣX²] up to a constant, increasing in ΣX² for σ₂ > σ₁.X₁, …, Xₙ are i.i.d. N(μ, 1). For H₀: μ = 0 against H₁: μ ≠ 0 at level α:
Show answer
Answer: A — no UMP test exists
Against μ > 0 the unique MP test rejects for large X̄, against μ < 0 for small X̄; a test cannot coincide with both. The one-sided test has power below α for μ < 0. The two-sided test is UMP only within the unbiased tests (UMPU), not among all level-α tests.A single observation X is used to test H₀: X ~ U(0, 1) against H₁: X has density 2x on (0, 1). The most powerful test of size 0.05 has power, correct to four decimal places, ____.
Numerical answer — type the value.
Show answer
Answer: 0.0975
f₁/f₀ = 2x is increasing, so the MP test rejects when X > c; P₀(X > c) = 1 − c = 0.05 gives c = 0.95. Power = P₁(X > 0.95) = 1 − 0.95² = 0.0975 — barely above the size, because one observation carries little information. Rejecting for small X gives power 0.0025, far below α.Which statements about tests are true?
Show answer
Answer: A — The size of a test is the supremum of its power function over the null parameter set; C — At θ in the alternative, power = 1 − P(type II error); D — A test of level α has size at most α
(A), (C), (D) are the definitions. (B) False: the p-value is P(data at least as extreme as observed | H₀), a probability computed under H₀, not a probability of H₀ — in the frequentist framework H₀ is not random at all.X ~ Bin(5, p). A randomised test of H₀: p = 1/2 against H₁: p > 1/2 rejects when X = 5 and rejects with probability γ when X = 4. The value of γ that makes the size exactly 0.05, correct to two decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.12
P(X = 5) + γP(X = 4) = 1/32 + γ(5/32) = 0.05 gives γ = (0.05 − 0.03125)/0.15625 = 0.12. Rejecting at X ≥ 4 outright would give size 6/32 = 0.1875, far too large; rejecting only at 5 gives 0.03125, too small.X₁, …, Xₙ are i.i.d. exponential with mean θ. The UMP level-α test of H₀: θ ≤ θ₀ against H₁: θ > θ₀ rejects H₀ when:
Show answer
Answer: A — ΣXᵢ > c
f = θ⁻ⁿ e−ΣX/θ: w(θ) = −1/θ is increasing, so the family has MLR in ΣX and Karlin–Rubin gives "reject for large ΣX", with c from 2ΣX/θ₀ ~ χ²2n. A large mean makes large observations likelier, so rejecting for small ΣX points the wrong way; the minimum is not sufficient.X ~ Poisson(λ). The test of H₀: λ = 1 that rejects when X ≥ 3 has size, correct to three decimal places, ____.
Numerical answer — type the value.
Show answer
Answer: 0.080
P(X ≥ 3) = 1 − e−1(1 + 1 + 1/2) = 1 − 2.5 × 0.36788 = 0.0803 → 0.080. Omitting the x = 2 term 1/2! gives 1 − 2e−1 = 0.264, the size of "reject when X ≥ 2".A type I error is:
Show answer
Answer: A — rejecting H₀ when H₀ is true
A type I error is a false rejection; its probability, maximised over H₀, is the size. Accepting a false H₀ is the type II error; rejecting H₀ when H₁ is true is the correct decision whose probability is the power.