Multivariate Analysis: The Multivariate Normal, Its Marginals and Conditionals, MLE of the Mean Vector and Dispersion Matrix, Wishart, Hotelling’s T² and Multiple and Partial Correlation

Section 11 of the GATE Statistics paper carries the bivariate normal of Section 5 into p dimensions. It follows the syllabus: the multivariate normal distribution and its properties; its conditional and marginal distributions; maximum likelihood estimation of the mean vector and the dispersion matrix; Hotelling’s T² test; the Wishart distribution and its basic properties; and multiple and partial correlation coefficients. It opens with the algebra of mean vectors and dispersion matrices — linear transformations, eigenvalues and the principal axes of Σ — because every numerical in the section is matrix arithmetic on a small Σ.

1. Mean vectors, dispersion matrices and their eigenvalues

For a random vector X with mean μ and dispersion (covariance) matrix Σ = E[(X − μ)(X − μ)ᵀ]: E[AX + b] = Aμ + b and Cov(AX + b) = AΣAᵀ; in particular Var(aᵀX) = aᵀΣa ≥ 0, so Σ is symmetric positive semidefinite, and singular exactly when some linear combination is constant. With Σ = [[4, 2], [2, 3]], Var(X₁ − 2X₂) = (1, −2)Σ(1, −2)ᵀ = 4 − 8 + 12 = 8.

Σ = PΛPᵀ with P orthogonal and Λ = diag(λ₁ ≥ … ≥ λ_p ≥ 0). The total variance tr Σ = Σλᵢ, the generalised variance det Σ = ∏λᵢ, and the components Y = PᵀX are uncorrelated with variances λᵢ — the principal axes of Σ, along which the constant-density ellipsoids of a normal vector are aligned (the principal components of X). The first axis carries the proportion λ₁/tr Σ of the total variance: for Σ = [[5, 2], [2, 2]], λ = 6, 1 and the proportion is 6/7 = 0.857.

2. The multivariate normal and its properties

X ~ N_p(μ, Σ) if every linear combination aᵀX is univariate normal; for Σ positive definite the density is (2π)−p/2|Σ|−1/2 exp[−½(x − μ)ᵀΣ⁻¹(x − μ)], and the MGF is exp(tᵀμ + ½tᵀΣt). The properties the paper tests:

  • Linear maps: AX + b ~ N(Aμ + b, AΣAᵀ); every marginal, and every subvector, is normal.
  • Independence: two subvectors are independent iff their cross-covariance block is zero — for the multivariate normal, uncorrelated means independent.
  • The quadratic form: (X − μ)ᵀΣ⁻¹(X − μ) ~ χ²_p — the squared Mahalanobis distance. For x − μ = (1, 1) and Σ = [[2, 1], [1, 2]], Σ⁻¹ = (1/3)[[2, −1], [−1, 2]] and the distance² is 2/3.
  • Normal marginals are not enough: a vector whose components are each normal need not be multivariate normal (the random-sign construction of Section 5).

3. Marginal and conditional distributions

Partition X = (X₁, X₂) with μ = (μ₁, μ₂) and Σ = [[Σ₁₁, Σ₁₂], [Σ₂₁, Σ₂₂]]. Then X₁ ~ N(μ₁, Σ₁₁), and X₁ | X₂ = x₂ ~ N(μ₁ + Σ₁₂Σ₂₂⁻¹(x₂ − μ₂), Σ₁₁ − Σ₁₂Σ₂₂⁻¹Σ₂₁). The conditional mean is linear in x₂ (the regression of X₁ on X₂), and the conditional dispersion Σ₁₁.₂ does not depend on x₂; it is the Schur complement, so det Σ = det Σ₂₂ · det Σ₁₁.₂.

Example: Σ = [[4, 1, 2], [1, 3, 1], [2, 1, 5]], conditioning X₁ on (X₂, X₃). Σ₂₂ = [[3, 1], [1, 5]] has inverse (1/14)[[5, −1], [−1, 3]], and Σ₁₂Σ₂₂⁻¹Σ₂₁ = (1/14)(5 − 4 + 12) = 13/14, so Var(X₁ | X₂, X₃) = 4 − 13/14 = 43/14 ≈ 3.07 — which equals det Σ/det Σ₂₂ = 43/14.

4. Maximum likelihood estimation and the Wishart distribution

For a sample X₁, …, Xₙ from N_p(μ, Σ), with A = Σ(Xᵢ − X̄)(Xᵢ − X̄)ᵀ: μ̂ = X̄ and Σ̂ = A/n — the latter biased, with S = A/(n − 1) unbiased. X̄ ~ N_p(μ, Σ/n), X̄ and A are independent, and A has the Wishart distribution W_p(n − 1, Σ).

  • Definition: if Z₁, …, Z_m are i.i.d. N_p(0, Σ), W = ΣZᵢZᵢᵀ ~ W_p(m, Σ). For p = 1 it is σ²χ²_m.
  • Mean: E W = mΣ. With m = 5 and σ₁₂ = 0.5, E W₁₂ = 2.5.
  • Additivity: independent W_p(m₁, Σ) and W_p(m₂, Σ) add to W_p(m₁ + m₂, Σ).
  • Linear maps: CWCᵀ ~ W_q(m, CΣCᵀ); in particular aᵀWa/aᵀΣa ~ χ²_m for every fixed a ≠ 0.

5. Hotelling’s T² test

To test H₀: μ = μ₀ with Σ unknown, T² = n(X̄ − μ₀)ᵀS⁻¹(X̄ − μ₀), and under H₀ (n − p)T²/[p(n − 1)] ~ F(p, n − p). For p = 1 it is t² with t the one-sample t-statistic, and T² is the likelihood ratio test. With n = 10, p = 2 and T² = 9, the F statistic is 8 × 9/(2 × 9) = 4 on (2, 8) df. T² is invariant under non-singular linear transformations of the data, so units and scales do not matter. Two samples: T² = [n₁n₂/(n₁ + n₂)](X̄₁ − X̄₂)ᵀS_p⁻¹(X̄₁ − X̄₂), with S_p the pooled dispersion.

⚠️ Why not p separate t-tests
Testing each coordinate at level α inflates the overall type I error and ignores the correlation between coordinates: two means can each look unremarkable while their combination, measured against the correlation, is extreme. T² tests the whole vector at exactly level α, and it is the maximum over all a of the squared t-statistic for aᵀμ.

6. Multiple and partial correlation

The partial correlation r₁₂.₃ is the correlation between X₁ and X₂ after removing the linear effect of X₃ from both: r₁₂.₃ = (r₁₂ − r₁₃r₂₃)/√[(1 − r₁₃²)(1 − r₂₃²)]. The multiple correlation R₁.₂₃ is the correlation between X₁ and its best linear predictor from X₂ and X₃: R²₁.₂₃ = (r₁₂² + r₁₃² − 2r₁₂r₁₃r₂₃)/(1 − r₂₃²), and in general R²₁.₂₃…ₚ = 1 − |R|/R₁₁ with R the correlation matrix and R₁₁ the cofactor of its (1, 1) entry. With r₁₂ = 0.6, r₁₃ = 0.5, r₂₃ = 0.4: r₁₂.₃ = 0.4/√0.63 = 0.504 and R²₁.₂₃ = 0.37/0.84 = 0.440, R₁.₂₃ = 0.664.

  • 0 ≤ R₁.₂₃ ≤ 1 — unlike simple and partial correlations, it is never negative.
  • R₁.₂₃ ≥ |r₁₂| and ≥ |r₁₃|: adding a predictor never lowers the multiple correlation.
  • 1 − R²₁.₂₃ = (1 − r₁₂²)(1 − r₁₃.₂²): the unexplained fraction factorises step by step.
  • A partial correlation can have the opposite sign to the simple one, and can be larger or smaller in size.

Key takeaways

  • Cov(AX + b) = AΣAᵀ; tr Σ is the total and det Σ the generalised variance; the principal axes are Σ’s eigenvectors, with variances its eigenvalues.
  • N_p: every aᵀX normal; zero cross-covariance ⇔ independence; (X − μ)ᵀΣ⁻¹(X − μ) ~ χ²_p.
  • X₁ | X₂ ~ N(μ₁ + Σ₁₂Σ₂₂⁻¹(x₂ − μ₂), Σ₁₁ − Σ₁₂Σ₂₂⁻¹Σ₂₁); the conditional dispersion does not depend on x₂.
  • μ̂ = X̄, Σ̂ = A/n (biased); X̄ ⊥ A ~ W_p(n − 1, Σ); E W = mΣ, Wisharts add.
  • T² = n(X̄ − μ₀)ᵀS⁻¹(X̄ − μ₀), (n − p)T²/[p(n − 1)] ~ F(p, n − p); p = 1 gives t².
  • r₁₂.₃ = (r₁₂ − r₁₃r₂₃)/√[(1 − r₁₃²)(1 − r₂₃²)]; R²₁.₂₃ = (r₁₂² + r₁₃² − 2r₁₂r₁₃r₂₃)/(1 − r₂₃²) ∈ [0, 1] and ≥ each r₁ⱼ².

Practice questions (14)

Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.

  1. X = (X₁, X₂) has dispersion matrix Σ = [[4, 2], [2, 3]]. Var(X₁ − 2X₂) is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 8

    aᵀΣa with a = (1, −2): 4 + 4 × 3 + 2 × 1 × (−2) × 2 = 4 + 12 − 8 = 8. Adding the covariance term instead gives 24; ignoring it gives 16.
  2. A random vector has dispersion matrix Σ = [[5, 2], [2, 2]]. The proportion of the total variance carried by its first principal axis (largest eigenvalue of Σ), correct to three decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.857

    Trace 7 and determinant 10 − 4 = 6 give eigenvalues 6 and 1, so the proportion is 6/7 = 0.857. Using the largest diagonal entry, 5/7 = 0.714, ignores the covariance that the rotation to principal axes absorbs.
  3. (X₁, X₂, X₃) is trivariate normal with dispersion matrix Σ = [[4, 1, 2], [1, 3, 1], [2, 1, 5]]. Var(X₁ | X₂, X₃), correct to two decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 3.07

    Σ₂₂ = [[3, 1], [1, 5]], Σ₂₂⁻¹ = (1/14)[[5, −1], [−1, 3]], Σ₁₂ = (1, 2): Σ₁₂Σ₂₂⁻¹Σ₂₁ = (5 − 4 + 12)/14 = 13/14. So Var = 4 − 13/14 = 43/14 = 3.071 → 3.07; check det Σ/det Σ₂₂ = 43/14. Subtracting only the separate effects 1/3 + 4/5 ignores the correlation between X₂ and X₃.
  4. (X, Y) is bivariate normal with mean (1, 2) and dispersion matrix [[2, 1], [1, 3]]. E[Y | X = 3] is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 3

    E[Y | X = x] = μ_Y + (σ_XY/σ_X²)(x − μ_X) = 2 + (1/2)(3 − 1) = 3. Dividing by σ_Y² = 3 instead gives the slope of X on Y and the answer 2.67.
  5. A sample of n = 10 from a bivariate normal population gives Hotelling’s T² = 9 for H₀: μ = μ₀. The corresponding F statistic is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 4

    F = (n − p)T²/[p(n − 1)] = (10 − 2) × 9/(2 × 9) = 4, on (p, n − p) = (2, 8) degrees of freedom. Dividing T² by p alone gives 4.5, forgetting the factor (n − p)/(n − 1).
  6. W ~ W₂(5, Σ) with Σ = [[1, 0.5], [0.5, 2]]. E(W₁₂), correct to one decimal place, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 2.5

    E W = mΣ, so E W₁₂ = 5 × 0.5 = 2.5 — each of the five terms ZᵢZᵢᵀ contributes Σ. Reporting 0.5 forgets the degrees of freedom.
  7. For three variables, r₁₂ = 0.6, r₁₃ = 0.5 and r₂₃ = 0.4. The partial correlation r₁₂.₃, correct to three decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.504

    (0.6 − 0.5 × 0.4)/√[(1 − 0.25)(1 − 0.16)] = 0.4/√0.63 = 0.4/0.7937 = 0.504. Forgetting the square root in the denominator gives 0.4/0.63 = 0.635.
  8. For the same correlations (r₁₂ = 0.6, r₁₃ = 0.5, r₂₃ = 0.4), the multiple correlation coefficient R₁.₂₃, correct to two decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.66

    R² = (0.36 + 0.25 − 2 × 0.6 × 0.5 × 0.4)/(1 − 0.16) = 0.37/0.84 = 0.4405, so R = 0.664 → 0.66. It exceeds |r₁₂| = 0.6, as it must. Reporting R² = 0.44 answers for the square.
  9. X ~ N_p(μ, Σ) with Σ positive definite. Which statements are true?

    1. Every linear combination aᵀX is univariate normal
    2. Uncorrelated components of X are independent
    3. (X − μ)ᵀΣ⁻¹(X − μ) has the χ²_p distribution
    4. Any random vector whose components are each normal is multivariate normal
    Show answer

    Answer: A — Every linear combination aᵀX is univariate normal; B — Uncorrelated components of X are independent; C — (X − μ)ᵀΣ⁻¹(X − μ) has the χ²_p distribution

    (A) This is the definition. (B) A zero covariance block makes the density factorise. (C) Z = Σ−1/2(X − μ) ~ N_p(0, I) and the form is ZᵀZ, a sum of p squared standard normals. (D) False: X ~ N(0, 1) and SX with an independent random sign S are each normal, but X + SX has an atom at 0.
  10. X₁, …, Xₙ are i.i.d. N_p(μ, Σ), with X̄ the sample mean and Σ̂ = (1/n)Σ(Xᵢ − X̄)(Xᵢ − X̄)ᵀ. Which statements are true?

    1. X̄ is the MLE of μ
    2. Σ̂ is an unbiased estimator of Σ
    3. X̄ and Σ̂ are independent
    4. nΣ̂ has the Wishart distribution W_p(n − 1, Σ)
    Show answer

    Answer: A — X̄ is the MLE of μ; C — X̄ and Σ̂ are independent; D — nΣ̂ has the Wishart distribution W_p(n − 1, Σ)

    (A) Maximising over μ for any Σ gives X̄. (B) False: E Σ̂ = (n − 1)Σ/n; the unbiased version divides by n − 1. (C) The multivariate analogue of X̄ ⊥ S². (D) nΣ̂ = A = Σ(Xᵢ − X̄)(Xᵢ − X̄)ᵀ, which is W_p(n − 1, Σ): one degree of freedom is spent on X̄.
  11. When p = 1, Hotelling’s T² statistic for H₀: μ = μ₀ reduces to:

    1. the square of the one-sample t statistic
    2. the one-sample t statistic
    3. the chi-square statistic (n − 1)S²/σ₀²
    4. the F statistic for two variances
    Show answer

    Answer: A — the square of the one-sample t statistic

    n(X̄ − μ₀)S⁻²(X̄ − μ₀) = [√n(X̄ − μ₀)/S]² = t², and F(1, n − 1) = t²n−1 agrees with (n − 1)T²/(n − 1). T² is a squared distance, so it cannot be the signed t itself.
  12. Which statements about multiple and partial correlation coefficients are true?

    1. R₁.₂₃ ≥ |r₁₂|
    2. 0 ≤ R₁.₂₃ ≤ 1
    3. r₁₂.₃ always has the same sign as r₁₂
    4. 1 − R²₁.₂₃ = (1 − r₁₂²)(1 − r₁₃.₂²)
    Show answer

    Answer: A — R₁.₂₃ ≥ |r₁₂|; B — 0 ≤ R₁.₂₃ ≤ 1; D — 1 − R²₁.₂₃ = (1 − r₁₂²)(1 − r₁₃.₂²)

    (A) The best linear predictor from (X₂, X₃) does at least as well as the best from X₂ alone. (B) R is the correlation of X₁ with its own best linear predictor, which is non-negative. (C) False: with r₁₂ = 0.3 and r₁₃ = r₂₃ = 0.7 (a valid correlation matrix, determinant 0.224), the numerator 0.3 − 0.49 is negative. (D) Explaining X₁ by X₂ leaves 1 − r₁₂², of which X₃ then explains the fraction r₁₃.₂².
  13. For x − μ = (1, 1) and Σ = [[2, 1], [1, 2]], the squared Mahalanobis distance (x − μ)ᵀΣ⁻¹(x − μ), correct to three decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.667

    Σ⁻¹ = (1/3)[[2, −1], [−1, 2]], so the form is (2 − 1 − 1 + 2)/3 = 2/3 = 0.667. (1, 1) is an eigenvector of Σ with eigenvalue 3, which gives the same 2/3 at once. The Euclidean distance² 2 ignores that the direction (1, 1) is the direction of greatest spread.
  14. W₁ ~ W_p(m₁, Σ) and W₂ ~ W_p(m₂, Σ) are independent. The distribution of W₁ + W₂ is:

    1. W_p(m₁ + m₂, Σ)
    2. W_p(m₁ + m₂, 2Σ)
    3. W2p(m₁ + m₂, Σ)
    4. W_p(max(m₁, m₂), Σ)
    Show answer

    Answer: A — W_p(m₁ + m₂, Σ)

    Each Wᵢ is a sum of mᵢ outer products ZZᵀ of independent N_p(0, Σ) vectors; together they are a sum of m₁ + m₂ such products, so W_p(m₁ + m₂, Σ) — the matrix analogue of χ²m₁ + χ²m₂ = χ²m₁+m₂. The scale matrix stays Σ, and the dimension stays p.