Multivariate Analysis: The Multivariate Normal, Its Marginals and Conditionals, MLE of the Mean Vector and Dispersion Matrix, Wishart, Hotelling’s T² and Multiple and Partial Correlation
1. Mean vectors, dispersion matrices and their eigenvalues
For a random vector X with mean μ and dispersion (covariance) matrix Σ = E[(X − μ)(X − μ)ᵀ]: E[AX + b] = Aμ + b and Cov(AX + b) = AΣAᵀ; in particular Var(aᵀX) = aᵀΣa ≥ 0, so Σ is symmetric positive semidefinite, and singular exactly when some linear combination is constant. With Σ = [[4, 2], [2, 3]], Var(X₁ − 2X₂) = (1, −2)Σ(1, −2)ᵀ = 4 − 8 + 12 = 8.
Σ = PΛPᵀ with P orthogonal and Λ = diag(λ₁ ≥ … ≥ λ_p ≥ 0). The total variance tr Σ = Σλᵢ, the generalised variance det Σ = ∏λᵢ, and the components Y = PᵀX are uncorrelated with variances λᵢ — the principal axes of Σ, along which the constant-density ellipsoids of a normal vector are aligned (the principal components of X). The first axis carries the proportion λ₁/tr Σ of the total variance: for Σ = [[5, 2], [2, 2]], λ = 6, 1 and the proportion is 6/7 = 0.857.
2. The multivariate normal and its properties
X ~ N_p(μ, Σ) if every linear combination aᵀX is univariate normal; for Σ positive definite the density is (2π)−p/2|Σ|−1/2 exp[−½(x − μ)ᵀΣ⁻¹(x − μ)], and the MGF is exp(tᵀμ + ½tᵀΣt). The properties the paper tests:
- Linear maps: AX + b ~ N(Aμ + b, AΣAᵀ); every marginal, and every subvector, is normal.
- Independence: two subvectors are independent iff their cross-covariance block is zero — for the multivariate normal, uncorrelated means independent.
- The quadratic form: (X − μ)ᵀΣ⁻¹(X − μ) ~ χ²_p — the squared Mahalanobis distance. For x − μ = (1, 1) and Σ = [[2, 1], [1, 2]], Σ⁻¹ = (1/3)[[2, −1], [−1, 2]] and the distance² is 2/3.
- Normal marginals are not enough: a vector whose components are each normal need not be multivariate normal (the random-sign construction of Section 5).
3. Marginal and conditional distributions
Partition X = (X₁, X₂) with μ = (μ₁, μ₂) and Σ = [[Σ₁₁, Σ₁₂], [Σ₂₁, Σ₂₂]]. Then X₁ ~ N(μ₁, Σ₁₁), and X₁ | X₂ = x₂ ~ N(μ₁ + Σ₁₂Σ₂₂⁻¹(x₂ − μ₂), Σ₁₁ − Σ₁₂Σ₂₂⁻¹Σ₂₁). The conditional mean is linear in x₂ (the regression of X₁ on X₂), and the conditional dispersion Σ₁₁.₂ does not depend on x₂; it is the Schur complement, so det Σ = det Σ₂₂ · det Σ₁₁.₂.
Example: Σ = [[4, 1, 2], [1, 3, 1], [2, 1, 5]], conditioning X₁ on (X₂, X₃). Σ₂₂ = [[3, 1], [1, 5]] has inverse (1/14)[[5, −1], [−1, 3]], and Σ₁₂Σ₂₂⁻¹Σ₂₁ = (1/14)(5 − 4 + 12) = 13/14, so Var(X₁ | X₂, X₃) = 4 − 13/14 = 43/14 ≈ 3.07 — which equals det Σ/det Σ₂₂ = 43/14.
4. Maximum likelihood estimation and the Wishart distribution
For a sample X₁, …, Xₙ from N_p(μ, Σ), with A = Σ(Xᵢ − X̄)(Xᵢ − X̄)ᵀ: μ̂ = X̄ and Σ̂ = A/n — the latter biased, with S = A/(n − 1) unbiased. X̄ ~ N_p(μ, Σ/n), X̄ and A are independent, and A has the Wishart distribution W_p(n − 1, Σ).
- Definition: if Z₁, …, Z_m are i.i.d. N_p(0, Σ), W = ΣZᵢZᵢᵀ ~ W_p(m, Σ). For p = 1 it is σ²χ²_m.
- Mean: E W = mΣ. With m = 5 and σ₁₂ = 0.5, E W₁₂ = 2.5.
- Additivity: independent W_p(m₁, Σ) and W_p(m₂, Σ) add to W_p(m₁ + m₂, Σ).
- Linear maps: CWCᵀ ~ W_q(m, CΣCᵀ); in particular aᵀWa/aᵀΣa ~ χ²_m for every fixed a ≠ 0.
5. Hotelling’s T² test
To test H₀: μ = μ₀ with Σ unknown, T² = n(X̄ − μ₀)ᵀS⁻¹(X̄ − μ₀), and under H₀ (n − p)T²/[p(n − 1)] ~ F(p, n − p). For p = 1 it is t² with t the one-sample t-statistic, and T² is the likelihood ratio test. With n = 10, p = 2 and T² = 9, the F statistic is 8 × 9/(2 × 9) = 4 on (2, 8) df. T² is invariant under non-singular linear transformations of the data, so units and scales do not matter. Two samples: T² = [n₁n₂/(n₁ + n₂)](X̄₁ − X̄₂)ᵀS_p⁻¹(X̄₁ − X̄₂), with S_p the pooled dispersion.
6. Multiple and partial correlation
The partial correlation r₁₂.₃ is the correlation between X₁ and X₂ after removing the linear effect of X₃ from both: r₁₂.₃ = (r₁₂ − r₁₃r₂₃)/√[(1 − r₁₃²)(1 − r₂₃²)]. The multiple correlation R₁.₂₃ is the correlation between X₁ and its best linear predictor from X₂ and X₃: R²₁.₂₃ = (r₁₂² + r₁₃² − 2r₁₂r₁₃r₂₃)/(1 − r₂₃²), and in general R²₁.₂₃…ₚ = 1 − |R|/R₁₁ with R the correlation matrix and R₁₁ the cofactor of its (1, 1) entry. With r₁₂ = 0.6, r₁₃ = 0.5, r₂₃ = 0.4: r₁₂.₃ = 0.4/√0.63 = 0.504 and R²₁.₂₃ = 0.37/0.84 = 0.440, R₁.₂₃ = 0.664.
- 0 ≤ R₁.₂₃ ≤ 1 — unlike simple and partial correlations, it is never negative.
- R₁.₂₃ ≥ |r₁₂| and ≥ |r₁₃|: adding a predictor never lowers the multiple correlation.
- 1 − R²₁.₂₃ = (1 − r₁₂²)(1 − r₁₃.₂²): the unexplained fraction factorises step by step.
- A partial correlation can have the opposite sign to the simple one, and can be larger or smaller in size.
Key takeaways
- Cov(AX + b) = AΣAᵀ; tr Σ is the total and det Σ the generalised variance; the principal axes are Σ’s eigenvectors, with variances its eigenvalues.
- N_p: every aᵀX normal; zero cross-covariance ⇔ independence; (X − μ)ᵀΣ⁻¹(X − μ) ~ χ²_p.
- X₁ | X₂ ~ N(μ₁ + Σ₁₂Σ₂₂⁻¹(x₂ − μ₂), Σ₁₁ − Σ₁₂Σ₂₂⁻¹Σ₂₁); the conditional dispersion does not depend on x₂.
- μ̂ = X̄, Σ̂ = A/n (biased); X̄ ⊥ A ~ W_p(n − 1, Σ); E W = mΣ, Wisharts add.
- T² = n(X̄ − μ₀)ᵀS⁻¹(X̄ − μ₀), (n − p)T²/[p(n − 1)] ~ F(p, n − p); p = 1 gives t².
- r₁₂.₃ = (r₁₂ − r₁₃r₂₃)/√[(1 − r₁₃²)(1 − r₂₃²)]; R²₁.₂₃ = (r₁₂² + r₁₃² − 2r₁₂r₁₃r₂₃)/(1 − r₂₃²) ∈ [0, 1] and ≥ each r₁ⱼ².
Practice questions (14)
Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.
X = (X₁, X₂) has dispersion matrix Σ = [[4, 2], [2, 3]]. Var(X₁ − 2X₂) is ____.
Numerical answer — type the value.
Show answer
Answer: 8
aᵀΣa with a = (1, −2): 4 + 4 × 3 + 2 × 1 × (−2) × 2 = 4 + 12 − 8 = 8. Adding the covariance term instead gives 24; ignoring it gives 16.A random vector has dispersion matrix Σ = [[5, 2], [2, 2]]. The proportion of the total variance carried by its first principal axis (largest eigenvalue of Σ), correct to three decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.857
Trace 7 and determinant 10 − 4 = 6 give eigenvalues 6 and 1, so the proportion is 6/7 = 0.857. Using the largest diagonal entry, 5/7 = 0.714, ignores the covariance that the rotation to principal axes absorbs.(X₁, X₂, X₃) is trivariate normal with dispersion matrix Σ = [[4, 1, 2], [1, 3, 1], [2, 1, 5]]. Var(X₁ | X₂, X₃), correct to two decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 3.07
Σ₂₂ = [[3, 1], [1, 5]], Σ₂₂⁻¹ = (1/14)[[5, −1], [−1, 3]], Σ₁₂ = (1, 2): Σ₁₂Σ₂₂⁻¹Σ₂₁ = (5 − 4 + 12)/14 = 13/14. So Var = 4 − 13/14 = 43/14 = 3.071 → 3.07; check det Σ/det Σ₂₂ = 43/14. Subtracting only the separate effects 1/3 + 4/5 ignores the correlation between X₂ and X₃.(X, Y) is bivariate normal with mean (1, 2) and dispersion matrix [[2, 1], [1, 3]]. E[Y | X = 3] is ____.
Numerical answer — type the value.
Show answer
Answer: 3
E[Y | X = x] = μ_Y + (σ_XY/σ_X²)(x − μ_X) = 2 + (1/2)(3 − 1) = 3. Dividing by σ_Y² = 3 instead gives the slope of X on Y and the answer 2.67.A sample of n = 10 from a bivariate normal population gives Hotelling’s T² = 9 for H₀: μ = μ₀. The corresponding F statistic is ____.
Numerical answer — type the value.
Show answer
Answer: 4
F = (n − p)T²/[p(n − 1)] = (10 − 2) × 9/(2 × 9) = 4, on (p, n − p) = (2, 8) degrees of freedom. Dividing T² by p alone gives 4.5, forgetting the factor (n − p)/(n − 1).W ~ W₂(5, Σ) with Σ = [[1, 0.5], [0.5, 2]]. E(W₁₂), correct to one decimal place, is ____.
Numerical answer — type the value.
Show answer
Answer: 2.5
E W = mΣ, so E W₁₂ = 5 × 0.5 = 2.5 — each of the five terms ZᵢZᵢᵀ contributes Σ. Reporting 0.5 forgets the degrees of freedom.For three variables, r₁₂ = 0.6, r₁₃ = 0.5 and r₂₃ = 0.4. The partial correlation r₁₂.₃, correct to three decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.504
(0.6 − 0.5 × 0.4)/√[(1 − 0.25)(1 − 0.16)] = 0.4/√0.63 = 0.4/0.7937 = 0.504. Forgetting the square root in the denominator gives 0.4/0.63 = 0.635.For the same correlations (r₁₂ = 0.6, r₁₃ = 0.5, r₂₃ = 0.4), the multiple correlation coefficient R₁.₂₃, correct to two decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.66
R² = (0.36 + 0.25 − 2 × 0.6 × 0.5 × 0.4)/(1 − 0.16) = 0.37/0.84 = 0.4405, so R = 0.664 → 0.66. It exceeds |r₁₂| = 0.6, as it must. Reporting R² = 0.44 answers for the square.X ~ N_p(μ, Σ) with Σ positive definite. Which statements are true?
Show answer
Answer: A — Every linear combination aᵀX is univariate normal; B — Uncorrelated components of X are independent; C — (X − μ)ᵀΣ⁻¹(X − μ) has the χ²_p distribution
(A) This is the definition. (B) A zero covariance block makes the density factorise. (C) Z = Σ−1/2(X − μ) ~ N_p(0, I) and the form is ZᵀZ, a sum of p squared standard normals. (D) False: X ~ N(0, 1) and SX with an independent random sign S are each normal, but X + SX has an atom at 0.X₁, …, Xₙ are i.i.d. N_p(μ, Σ), with X̄ the sample mean and Σ̂ = (1/n)Σ(Xᵢ − X̄)(Xᵢ − X̄)ᵀ. Which statements are true?
Show answer
Answer: A — X̄ is the MLE of μ; C — X̄ and Σ̂ are independent; D — nΣ̂ has the Wishart distribution W_p(n − 1, Σ)
(A) Maximising over μ for any Σ gives X̄. (B) False: E Σ̂ = (n − 1)Σ/n; the unbiased version divides by n − 1. (C) The multivariate analogue of X̄ ⊥ S². (D) nΣ̂ = A = Σ(Xᵢ − X̄)(Xᵢ − X̄)ᵀ, which is W_p(n − 1, Σ): one degree of freedom is spent on X̄.When p = 1, Hotelling’s T² statistic for H₀: μ = μ₀ reduces to:
Show answer
Answer: A — the square of the one-sample t statistic
n(X̄ − μ₀)S⁻²(X̄ − μ₀) = [√n(X̄ − μ₀)/S]² = t², and F(1, n − 1) = t²n−1 agrees with (n − 1)T²/(n − 1). T² is a squared distance, so it cannot be the signed t itself.Which statements about multiple and partial correlation coefficients are true?
Show answer
Answer: A — R₁.₂₃ ≥ |r₁₂|; B — 0 ≤ R₁.₂₃ ≤ 1; D — 1 − R²₁.₂₃ = (1 − r₁₂²)(1 − r₁₃.₂²)
(A) The best linear predictor from (X₂, X₃) does at least as well as the best from X₂ alone. (B) R is the correlation of X₁ with its own best linear predictor, which is non-negative. (C) False: with r₁₂ = 0.3 and r₁₃ = r₂₃ = 0.7 (a valid correlation matrix, determinant 0.224), the numerator 0.3 − 0.49 is negative. (D) Explaining X₁ by X₂ leaves 1 − r₁₂², of which X₃ then explains the fraction r₁₃.₂².For x − μ = (1, 1) and Σ = [[2, 1], [1, 2]], the squared Mahalanobis distance (x − μ)ᵀΣ⁻¹(x − μ), correct to three decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.667
Σ⁻¹ = (1/3)[[2, −1], [−1, 2]], so the form is (2 − 1 − 1 + 2)/3 = 2/3 = 0.667. (1, 1) is an eigenvector of Σ with eigenvalue 3, which gives the same 2/3 at once. The Euclidean distance² 2 ignores that the direction (1, 1) is the direction of greatest spread.W₁ ~ W_p(m₁, Σ) and W₂ ~ W_p(m₂, Σ) are independent. The distribution of W₁ + W₂ is:
Show answer
Answer: A — W_p(m₁ + m₂, Σ)
Each Wᵢ is a sum of mᵢ outer products ZZᵀ of independent N_p(0, Σ) vectors; together they are a sum of m₁ + m₂ such products, so W_p(m₁ + m₂, Σ) — the matrix analogue of χ²m₁ + χ²m₂ = χ²m₁+m₂. The scale matrix stays Σ, and the dimension stays p.