Joint Distributions II: Functions of Random Vectors, Order Statistics, and the Chi-square, t and F Sampling Distributions

The second chapter for Section 5 of the GATE Statistics paper, on what the section names after the bivariate normal: functions of a random vector and their distributions, by the Jacobian, convolution and MGF methods; the distributions of order statistics, one at a time and jointly; and the sampling distributions — the central chi-square, central t and central F — with the facts about a normal sample that estimation and testing rely on: the independence of X̄ and S², and the chi-square law of (n − 1)S²/σ². The Jacobian, order-statistic and sampling-distribution calculations are where Section 5’s two-mark numericals usually come from.

1. Functions of a random vector

Jacobian method: if (U, V) = g(X, Y) is one-to-one with inverse (x(u, v), y(u, v)), then fU,V(u, v) = fX,Y(x(u, v), y(u, v)) |∂(x, y)/∂(u, v)|, on the image of the support. To get the distribution of one function U, introduce a convenient companion V, transform, and integrate V out.

Example: X, Y independent Exp(1); U = X + Y, V = X/(X + Y). Then x = uv, y = u(1 − v), |J| = u, and fU,V = e−u · u on u > 0, 0 < v < 1. It factors: U ~ Gamma(2, 1) and V ~ U(0, 1), independent. So P(V ≤ 0.3, U ≤ 1) = 0.3 × (1 − 2e−1) = 0.079. The same calculation for Gamma(a) and Gamma(b) gives a Gamma(a + b) sum independent of a β₁(a, b) proportion.

  • Convolution: for independent X, Y, fX+Y(z) = ∫ f_X(x) f_Y(z − x) dx. Two U(0, 1) give the triangle min(z, 2 − z) on (0, 2), with value 0.5 at z = 1.5.
  • MGF method: MX+Y = M_X M_Y for independent X, Y, then recognise the product: n independent Exp(λ) give (1 − t/λ)−n, the Gamma(n, λ) MGF.
  • Ratios of normals: X/Y for independent N(0, 1) is standard Cauchy; Z/√(V/k) with V ~ χ²_k independent is t_k.

2. Order statistics

For an i.i.d. sample from a continuous F with density f, the ordered values X₍₁₎ < … < X₍ₙ₎ have

Distributions of order statistics
StatisticDensity or CDF
Maximum X₍ₙ₎F₍ₙ₎(x) = F(x)ⁿ
Minimum X₍₁₎F₍₁₎(x) = 1 − [1 − F(x)]ⁿ
r-th, X₍ᵣ₎n!/[(r − 1)!(n − r)!] Fr−1(1 − F)n−r f
Joint (X₍ᵣ₎, X₍ₛ₎), r < s, x < yn!/[(r−1)!(s−r−1)!(n−s)!] F(x)r−1[F(y) − F(x)]s−r−1[1 − F(y)]n−s f(x)f(y)
All n jointlyn! f(x₁)…f(xₙ) on x₁ < … < xₙ
  • Uniform samples: X₍ᵣ₎ ~ β₁(r, n − r + 1), so E X₍ᵣ₎ = r/(n + 1); the maximum of 4 has mean 4/5, the middle of 3 is β₁(2, 2) with variance 1/20. The joint density of (X₍₁₎, X₍ₙ₎) is n(n − 1)(y − x)n−2 on 0 < x < y < 1, and the range has mean (n − 1)/(n + 1).
  • Exponential samples: X₍₁₎ ~ Exp(nλ), and the spacings X₍ᵢ₎ − X₍ᵢ₋₁₎ are independent Exp((n − i + 1)λ). So E X₍ₙ₎ = (1/λ)(1 + 1/2 + … + 1/n): for three Exp(1), 11/6.
🧠 Read an order-statistic density as a count
The density of X₍ᵣ₎ at x says: r − 1 observations fall below x (Fr−1), one falls at x (f), and n − r above (1 − F)n−r, with the multinomial coefficient counting the ways. Written that way, every formula in the table is reconstructed in seconds rather than memorised.

3. The central chi-square, t and F distributions

Sampling distributions
DistributionConstructionMeanVariance
χ²_kZ₁² + … + Z_k², Zᵢ i.i.d. N(0, 1) = Gamma(k/2, rate 1/2)k2k
t_kZ/√(V/k), V ~ χ²_k independent of Z0 (k > 1)k/(k − 2) (k > 2)
F(m, n)(U/m)/(V/n), U ~ χ²_m, V ~ χ²_n independentn/(n − 2) (n > 2)2n²(m + n − 2)/[m(n − 2)²(n − 4)] (n > 4)
  • χ²_m + χ²_n = χ²m+n for independent summands; χ²₂ is exponential with mean 2, so P(χ²₂ > 4) = e−2 = 0.135.
  • t_k → N(0, 1) as k → ∞; t₁ is the standard Cauchy; T ~ t_k ⇒ T² ~ F(1, k).
  • F ~ F(m, n) ⇒ 1/F ~ F(n, m), so lower F quantiles come from upper ones: F1−α(m, n) = 1/F_α(n, m).

4. Sampling from a normal population

For X₁, …, Xₙ i.i.d. N(μ, σ²): X̄ ~ N(μ, σ²/n); X̄ and S² are independent; (n − 1)S²/σ² ~ χ²n−1; √n(X̄ − μ)/S ~ tn−1. Hence Var(S²) = 2σ⁴/(n − 1). For two independent normal samples, (S₁²/σ₁²)/(S₂²/σ₂²) ~ F(n₁ − 1, n₂ − 1). The independence of X̄ and S² characterises the normal: it holds for no other distribution.

⚠️ n − 1, not n
The degrees of freedom are n − 1 because one linear constraint, Σ(Xᵢ − X̄) = 0, is spent estimating μ. √n(X̄ − μ)/S with n = 10 is t₉, and 9S²/σ² is χ²₉. With μ known, Σ(Xᵢ − μ)²/σ² is χ²_n — the one place n itself appears.

Key takeaways

  • fU,V = fX,Y · |∂(x, y)/∂(u, v)|; for gammas with a common rate the sum is independent of the proportion, which is beta.
  • X₍ₙ₎ has CDF Fⁿ, X₍₁₎ has 1 − (1 − F)ⁿ; X₍ᵣ₎ density = count below × density × count above.
  • Uniform: X₍ᵣ₎ ~ β₁(r, n − r + 1), mean r/(n + 1). Exponential: min is Exp(nλ); E max = (1/λ)Σ1/i.
  • χ²_k: mean k, variance 2k; t_k variance k/(k − 2); F(m, n) mean n/(n − 2); t² ~ F(1, k); t₁ is Cauchy.
  • Normal sample: X̄ ⊥ S², (n − 1)S²/σ² ~ χ²n−1, √n(X̄ − μ)/S ~ tn−1.

Practice questions (14)

Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.

  1. X and Y are independent U(0, 1). The density of X + Y at the point 1.5, correct to one decimal place, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.5

    fX+Y(z) = ∫ 1{0 < x < 1} 1{0 < z − x < 1} dx = length of (z − 1, 1) = 2 − z for 1 < z < 2, so 0.5 at z = 1.5. Treating the sum as U(0, 2) gives a flat 0.5 everywhere — right here by coincidence, wrong at z = 1, where the density is 1.
  2. X₁, …, X₄ are i.i.d. U(0, 1). The expected value of max(X₁, …, X₄), correct to one decimal place, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.8

    F₍₄₎(x) = x⁴, density 4x³, mean ∫4x⁴ dx = 4/5 = 0.8 — the general r/(n + 1) with r = n = 4. Answering 1, the largest possible value, confuses the expected maximum with the supremum.
  3. X₁, X₂, X₃ are i.i.d. U(0, 1). The variance of the sample median X₍₂₎, correct to two decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.05

    X₍₂₎ has density 3!/(1!1!) x(1 − x) = 6x(1 − x), i.e. β₁(2, 2), with variance ab/[(a + b)²(a + b + 1)] = 4/(16 × 5) = 0.05. Quoting the variance of a single uniform, 1/12 = 0.083, ignores that the median of three is concentrated near 1/2.
  4. X₁, X₂, X₃ are i.i.d. exponential with rate 2. The expected value of min(X₁, X₂, X₃), correct to three decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.167

    P(min > x) = (e−2x)³ = e−6x, so the minimum is Exp(6) with mean 1/6 = 0.1667 → 0.167. Dividing the mean 1/2 by 3 gives the same number here, but the reason is that the rates add — for a non-exponential distribution that shortcut fails.
  5. X₁, …, X₄ are i.i.d. U(0, 1). The expected value of the sample range X₍₄₎ − X₍₁₎, correct to one decimal place, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.6

    E X₍₄₎ = 4/5 and E X₍₁₎ = 1/5, so E(range) = 3/5 = 0.6 = (n − 1)/(n + 1). Linearity of expectation does the work; the dependence between the maximum and the minimum does not matter for the mean.
  6. X₁, X₂, X₃ are i.i.d. exponential with mean 1. The expected value of max(X₁, X₂, X₃), correct to two decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 1.83

    Spacings: the first failure among 3 takes Exp(3) (mean 1/3), then among 2, Exp(2) (1/2), then Exp(1) (1). E max = 1/3 + 1/2 + 1 = 11/6 = 1.833 → 1.83. Answering 3 adds three means of 1 as if the variables were waited for one after another.
  7. X₁, …, X₁₀ are i.i.d. N(μ, σ²), with sample mean X̄ and sample variance S² (divisor n − 1). Which statements are true?

    1. X̄ and S² are independent
    2. 9S²/σ² has the χ²₉ distribution
    3. √10 (X̄ − μ)/S has the t₁₀ distribution
    4. Var(S²) = 2σ⁴/9
    Show answer

    Answer: A — X̄ and S² are independent; B — 9S²/σ² has the χ²₉ distribution; D — Var(S²) = 2σ⁴/9

    (A) and (B) are the normal-sample theorem, with n − 1 = 9. (C) False: the statistic is t with n − 1 = 9 degrees of freedom, not 10. (D) S² = σ²χ²₉/9, so Var S² = σ⁴ × 2 × 9/81 = 2σ⁴/9.
  8. Which statements about the t, F and chi-square distributions are true?

    1. If T ~ t_k, then T² ~ F(1, k)
    2. If F ~ F(m, n), then 1/F ~ F(n, m)
    3. The mean of F(m, n) is n/(n − 2) for n > 2
    4. t_k has variance k/(k − 2) for every k ≥ 1
    Show answer

    Answer: A — If T ~ t_k, then T² ~ F(1, k); B — If F ~ F(m, n), then 1/F ~ F(n, m); C — The mean of F(m, n) is n/(n − 2) for n > 2

    (A) T² = Z²/(V/k) = (χ²₁/1)/(χ²_k/k). (B) Invert the ratio of the two scaled chi-squares. (C) E[U/m] = 1 and E[n/V] = n/(n − 2) by independence. (D) False: the variance is finite only for k > 2 — t₁ (Cauchy) has no mean, and t₂ has infinite variance.
  9. T has the t distribution with 6 degrees of freedom. Var(T), correct to one decimal place, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 1.5

    Var(t_k) = k/(k − 2) = 6/4 = 1.5 — larger than the normal’s 1 because S is itself random. Answering 1 treats t₆ as if it were already standard normal.
  10. The mean of the F(4, 10) distribution, correct to two decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 1.25

    E F(m, n) = n/(n − 2) = 10/8 = 1.25, depending only on the denominator degrees of freedom. Using m/(m − 2) = 2 takes the wrong one.
  11. X and Y are independent Exp(1). Let U = X + Y and V = X/(X + Y). P(V ≤ 0.3, U ≤ 1), correct to three decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.079

    With x = uv, y = u(1 − v), |J| = u and fU,V = u e−u on u > 0, 0 < v < 1: U ~ Gamma(2, 1) and V ~ U(0, 1) are independent. P(U ≤ 1) = 1 − e−1(1 + 1) = 0.2642, so the answer is 0.3 × 0.2642 = 0.0793 → 0.079. Using P(U ≤ 1) = 1 − e−1, as if U were Exp(1), gives 0.190.
  12. X₁, …, Xₙ are i.i.d. exponential with rate λ. The distribution of X₁ + … + Xₙ is:

    1. Gamma with shape n and rate λ
    2. Exponential with rate nλ
    3. Exponential with rate λ/n
    4. Gamma with shape λ and rate n
    Show answer

    Answer: A — Gamma with shape n and rate λ

    The MGF of the sum is [λ/(λ − t)]ⁿ = (1 − t/λ)−n, the Gamma(n, λ) MGF. Exponential with rate nλ is the distribution of the minimum, not the sum; the sum has mean n/λ, which no exponential with the listed rates matches.
  13. X has the chi-square distribution with 2 degrees of freedom. P(X > 4), correct to three decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.135

    χ²₂ = Gamma(1, rate 1/2) is exponential with mean 2, so P(X > 4) = e−4/2 = e−2 = 0.1353 → 0.135. Using mean 1 (rate 1) gives e−4 = 0.018.
  14. For an i.i.d. sample of size n ≥ 2 from U(0, 1), the joint density of (X₍₁₎, X₍ₙ₎) at 0 < x < y < 1 is:

    1. n(n − 1)(y − x)n−2
    2. n²xⁿ⁻¹(1 − y)ⁿ⁻¹
    3. n(y − x)n−1
    4. n! (y − x)n−2
    Show answer

    Answer: A — n(n − 1)(y − x)^{n−2}

    One observation at x, one at y and the other n − 2 in between: n!/(0! (n − 2)! 0!) (y − x)n−2 = n(n − 1)(y − x)n−2. It integrates to 1 over the triangle. The product of the two marginal densities is wrong because the minimum and maximum are dependent; n! over-counts the arrangements of the middle n − 2.