Joint Distributions I: Joint, Marginal and Conditional Distributions, Conditional Expectation, Correlation, the Multinomial and the Bivariate Normal
1. Joint distribution, mass and density functions, and marginals
The joint distribution function F(x, y) = P(X ≤ x, Y ≤ y) is non-decreasing and right-continuous in each argument, tends to 0 as either argument → −∞ and to 1 as both → ∞, and must satisfy the rectangle inequality: P(a < X ≤ b, c < Y ≤ d) = F(b, d) − F(a, d) − F(b, c) + F(a, c) ≥ 0. Marginals come from letting the other argument go to ∞: F_X(x) = F(x, ∞). For densities, f_X(x) = ∫ f(x, y) dy, integrating over the range of y allowed for that x — the support, not the whole line, is where most errors occur.
Example: f(x, y) = 2 on 0 < x < y < 1. Then f_X(x) = ∫ₓ¹ 2 dy = 2(1 − x) and f_Y(y) = ∫₀^y 2 dx = 2y. The product of the marginals is 4y(1 − x) ≠ 2, and in any case a triangular support can never factor: a support that is not a product set rules out independence at once.
2. Conditional distributions and conditional expectation
f(y | x) = f(x, y)/f_X(x) where f_X(x) > 0, and E[Y | X = x] = ∫ y f(y | x) dy. In the triangle example, Y | X = x is uniform on (x, 1) and X | Y = y is uniform on (0, y), so E[X | Y = y] = y/2 and E[Y | X = x] = (1 + x)/2. E[Y | X] is a random variable, a function of X, and it obeys:
- Tower property: E[E[Y | X]] = E Y.
- Law of total variance: Var Y = E[Var(Y | X)] + Var(E[Y | X]). If N ~ Poisson(4) and X | N ~ Bin(N, 1/2), then Var X = E[N/4] + Var(N/2) = 1 + 1 = 2 — and indeed X ~ Poisson(2).
- Taking out what is known: E[g(X)Y | X] = g(X)E[Y | X]; if X and Y are independent, E[Y | X] = E Y.
- Best predictor: E[Y | X] minimises E[(Y − g(X))²] over all functions g; the best linear predictor is its linear-regression approximation, and the two coincide for the bivariate normal.
3. Product moments, correlation, the joint MGF and independence
Cov(X, Y) = E[XY] − E X E Y, and ρ = Cov/(σ_Xσ_Y) ∈ [−1, 1] by Cauchy–Schwarz, with |ρ| = 1 iff Y is almost surely a linear function of X. Var(aX + bY) = a²Var X + b²Var Y + 2ab Cov(X, Y): with Var X = 4, Var Y = 1, Cov = 1, Var(2X − 3Y) = 16 + 9 − 12 = 13. In the triangle example, E X = 1/3, E Y = 2/3, E XY = ∫₀¹∫₀^y 2xy dx dy = 1/4, so Cov = 1/4 − 2/9 = 1/36; both variances are 1/18, and ρ = 1/2.
The joint MGF M(s, t) = E[esX + tY] gives E[XʲYᵏ] as ∂j+kM/∂sʲ∂tᵏ at (0, 0), and X, Y are independent iff M(s, t) = M_X(s)M_Y(t) (equivalently iff the joint density or CDF factorises). Independence implies zero correlation; the converse fails.
4. The multinomial distribution
n independent trials, each falling in one of k classes with probabilities p₁, …, p_k: the counts have P(X₁ = x₁, …, X_k = x_k) = n!/(x₁!…x_k!) p₁x₁…p_kx_k with Σxᵢ = n. Each marginal is binomial, Xᵢ ~ Bin(n, pᵢ); Cov(Xᵢ, Xⱼ) = −npᵢpⱼ, negative because the counts compete for the same n; the correlation is −√(pᵢpⱼ/((1 − pᵢ)(1 − pⱼ))). Conditionally on Xₖ = m, the remaining counts are multinomial with n − m trials and probabilities rescaled to sum to 1.
5. The bivariate normal distribution
(X, Y) is bivariate normal with means μ₁, μ₂, standard deviations σ₁, σ₂ and correlation ρ (|ρ| < 1) if its density is exp(−Q/2)/(2πσ₁σ₂√(1 − ρ²)), with Q = [z₁² − 2ρz₁z₂ + z₂²]/(1 − ρ²) and zᵢ the standardised values. Its MGF is exp(μ₁s + μ₂t + ½(σ₁²s² + 2ρσ₁σ₂st + σ₂²t²)), so M(s, t) = exp(s² + st + t²) is bivariate normal with σ₁² = σ₂² = 2, covariance 1 and ρ = 1/2.
| Quantity | Result |
|---|---|
| Marginals | X ~ N(μ₁, σ₁²), Y ~ N(μ₂, σ₂²) |
| Conditional mean | E[Y | X = x] = μ₂ + ρ(σ₂/σ₁)(x − μ₁) |
| Conditional variance | Var(Y | X = x) = σ₂²(1 − ρ²), the same for every x |
| Linear combinations | aX + bY is normal |
| Independence | ρ = 0 ⇔ X and Y independent |
With μ = (10, 20), σ₁ = 2, σ₂ = 3 and ρ = 0.6: E[Y | X = 12] = 20 + 0.6 × 1.5 × 2 = 21.8 and Var(Y | X) = 9 × 0.64 = 5.76. The conditional variance does not depend on the observed x, which is the homoscedasticity a regression model assumes.
Key takeaways
- Marginals integrate the other variable over its allowed range; a support that is not a rectangle already rules out independence.
- E[E[Y | X]] = E Y and Var Y = E Var(Y | X) + Var E(Y | X); E[Y | X] is the best predictor of Y from X in mean square.
- Var(aX + bY) = a²σ_X² + b²σ_Y² + 2ab Cov; independence ⇔ joint MGF factorises; zero correlation is weaker (X and X²).
- Multinomial: binomial marginals, Cov(Xᵢ, Xⱼ) = −npᵢpⱼ.
- Bivariate normal: E[Y | x] = μ₂ + ρ(σ₂/σ₁)(x − μ₁), Var(Y | x) = σ₂²(1 − ρ²); ρ = 0 ⇔ independence — but only for jointly normal pairs.
Practice questions (13)
Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.
(X, Y) has joint density f(x, y) = 2 on 0 < x < y < 1. The value of E[X | Y = 0.8], correct to one decimal place, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.4
f_Y(y) = ∫₀^y 2 dx = 2y, so f(x | y) = 2/(2y) = 1/y on (0, y): X | Y = y is uniform on (0, y) and E[X | Y = 0.8] = 0.4. Using the unconditional mean E X = 1/3 ignores the information in Y.For the same density f(x, y) = 2 on 0 < x < y < 1, the correlation coefficient of X and Y, correct to one decimal place, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.5
E X = ∫2x(1 − x) dx = 1/3, E Y = ∫2y² dy = 2/3, E XY = ∫₀¹∫₀^y 2xy dx dy = ∫y³ dy = 1/4, Cov = 1/4 − 2/9 = 1/36. E X² = 1/6 and E Y² = 1/2, so both variances are 1/18. ρ = (1/36)/(1/18) = 0.5. Stopping at the covariance 0.028 answers a different question.X and Y take values 0 and 1 with P(0, 0) = 0.1, P(0, 1) = 0.2, P(1, 0) = 0.3 and P(1, 1) = 0.4. Cov(X, Y), correct to two decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: -0.02
E X = P(X = 1) = 0.7, E Y = P(Y = 1) = 0.6, E XY = P(1, 1) = 0.4. Cov = 0.4 − 0.42 = −0.02. A slightly negative covariance is easy to report as 0.02 by subtracting the wrong way; the sign matters. Type the answer as -0.02.N ~ Poisson(4), and given N = n, X ~ Bin(n, 1/2). The variance of X is ____.
Numerical answer — type the value.
Show answer
Answer: 2
Var X = E[Var(X | N)] + Var(E[X | N]) = E[N/4] + Var(N/2) = 1 + 4/4 = 2. Indeed X is a thinned Poisson, X ~ Poisson(2). Keeping only E[Var(X | N)] = 1 forgets the variability of the number of trials.X ~ U(−1, 1) and Y = X². Which statements are true?
Show answer
Answer: A — Cov(X, Y) = 0; C — E[Y | X] = X²; D — E[X | Y] = 0
(A) Cov = E X³ − E X E X² = 0 − 0 = 0. (B) False: P(X > 1/2, Y < 1/4) = 0 but P(X > 1/2)P(Y < 1/4) = (1/4)(1/2) > 0. (C) Y is a function of X. (D) Given Y = y, X = ±√y equally likely, so the conditional mean is 0. Zero correlation coexists with complete dependence.(X₁, X₂, X₃) is multinomial with n = 10 and probabilities (0.2, 0.3, 0.5). Cov(X₁, X₂), correct to one decimal place, is ____.
Numerical answer — type the value.
Show answer
Answer: -0.6
Cov(Xᵢ, Xⱼ) = −npᵢpⱼ = −10 × 0.2 × 0.3 = −0.6. The sign is negative because a trial that lands in class 1 cannot land in class 2; a positive 0.6 misses that constraint. Type the answer as -0.6.(X, Y) is bivariate normal with E X = 10, E Y = 20, σ_X = 2, σ_Y = 3 and ρ = 0.6. E[Y | X = 12], correct to one decimal place, is ____.
Numerical answer — type the value.
Show answer
Answer: 21.8
E[Y | X = x] = μ_Y + ρ(σ_Y/σ_X)(x − μ_X) = 20 + 0.6 × 1.5 × 2 = 21.8. Using the slope ρσ_X/σ_Y = 0.4 (the regression of X on Y) gives 20.8.Which statements are true?
Show answer
Answer: A — If (X, Y) is bivariate normal, the marginal distributions of X and Y are normal; B — If (X, Y) is bivariate normal with ρ = 0, then X and Y are independent; D — If (X, Y) is bivariate normal, then X + Y is normal
(A) Set t = 0 in the joint MGF. (B) With ρ = 0 the density factorises into the two marginals. (C) False: X ~ N(0, 1) and Y = SX with an independent random sign S have normal marginals, but X + Y has an atom at 0. (D) The MGF of X + Y is M(t, t), a normal MGF.Var X = 4, Var Y = 1 and Cov(X, Y) = 1. Var(2X − 3Y) is ____.
Numerical answer — type the value.
Show answer
Answer: 13
Var(2X − 3Y) = 4 × 4 + 9 × 1 + 2 × 2 × (−3) × 1 = 16 + 9 − 12 = 13. Adding the covariance term instead of subtracting it gives 37; ignoring it gives 25.(X, Y) has joint moment generating function M(s, t) = exp(s² + st + t²). The correlation coefficient of X and Y, correct to two decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.50
Match exp(½(σ₁²s² + 2σ₁₂st + σ₂²t²)): σ₁² = 2, σ₂² = 2 and 2σ₁₂/2 = 1, so σ₁₂ = 1. ρ = 1/√(2 × 2) = 0.5. Reading the coefficient of st directly as the covariance and dividing by variances of 1 gives 1, forgetting the factor ½ in the exponent.For a joint distribution function F, P(a < X ≤ b, c < Y ≤ d) equals:
Show answer
Answer: A — F(b, d) − F(a, d) − F(b, c) + F(a, c)
Inclusion–exclusion on the quadrants: remove the strips {X ≤ a} and {Y ≤ c} from {X ≤ b, Y ≤ d}, then add back their intersection, which was removed twice. F(b, d) − F(a, c) subtracts only one corner; the product form assumes independence.(X, Y) has joint density f(x, y) = 6e−2x−3y for x, y > 0. P(X < Y), correct to one decimal place, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.4
The density factorises: X ~ Exp(rate 2) and Y ~ Exp(rate 3), independent. P(X < Y) = ∫₀^∞ 2e−2xe−3x dx = 2/5 = 0.4. Answering 3/5 swaps the rates: the variable with the larger rate tends to be smaller.Among all functions g, the one that minimises E[(Y − g(X))²] is:
Show answer
Answer: A — g(X) = E[Y | X]
E[(Y − g)²] = E[(Y − E[Y | X])²] + E[(E[Y | X] − g)²], because the cross term vanishes by the tower property; the second term is zero exactly for g = E[Y | X]. The least-squares line is best among LINEAR functions only; the conditional median minimises absolute error, not squared error.