Convergence of Random Variables: The Four Modes and Their Relations, Slutsky, Borel–Cantelli, the Laws of Large Numbers and the Central Limit Theorem

Section 6 of the GATE Statistics paper is short on the page and long in the paper, because every large-sample result later on — consistency, the asymptotic normality of the MLE, large-sample tests — is a statement about convergence. The chapter follows the syllabus: convergence in distribution, in probability, almost surely and in r-th mean, and the implications among them with the counterexample that blocks each missing arrow; Slutsky’s lemma; the Borel–Cantelli lemmas; the weak and strong laws of large numbers; and the central limit theorem for i.i.d. random variables, with the normal-approximation numericals it produces.

1. The four modes of convergence

Definitions
ModeXₙ → X means
Almost surely (a.s.)P(lim Xₙ = X) = 1
In probabilityfor every ε > 0, P(|Xₙ − X| > ε) → 0
In r-th mean (r ≥ 1)E|Xₙ − X|ʳ → 0
In distributionFₙ(x) → F(x) at every continuity point x of F

Convergence in distribution is about the distributions only — the variables need not even be defined on the same space — and the continuity-point clause matters: Xₙ = 1/n converges in distribution to X = 0, although Fₙ(0) = 0 for every n and F(0) = 1, because 0 is a discontinuity point of F. By Lévy’s continuity theorem, Xₙ → X in distribution iff the characteristic functions converge pointwise.

2. How the modes are related

  • a.s. ⇒ in probability ⇒ in distribution.
  • r-th mean ⇒ in probability (Markov: P(|Xₙ − X| > ε) ≤ E|Xₙ − X|ʳ/εʳ), and r-th mean ⇒ s-th mean for s < r (Lyapunov).
  • In distribution to a constant ⇒ in probability to that constant.
  • In probability ⇒ a.s. along some subsequence; and if Σ P(|Xₙ − X| > ε) < ∞ for every ε, then a.s. (Borel–Cantelli).
The counterexamples that block the other arrows
SequenceConvergesFails
Xₙ = n·1{U < 1/n}, one U ~ U(0, 1)a.s. and in probability to 0in mean: E Xₙ = 1
Independent Xₙ = 1 w.p. 1/n, else 0in probability and in every mean to 0a.s.: Xₙ = 1 infinitely often
Independent Xₙ = n w.p. 1/n, else 0in probability to 0a.s. and in mean
Xₙ = −X for X ~ N(0, 1)in distribution to Xin probability: |Xₙ − X| = 2|X|

3. Slutsky’s lemma and the continuous mapping theorem

Slutsky: if Xₙ → X in distribution and Yₙ → c (a constant) in probability, then Xₙ + Yₙ → X + c, XₙYₙ → cX and Xₙ/Yₙ → X/c (c ≠ 0), all in distribution. With Xₙ → N(0, 1) and Yₙ → 2, XₙYₙ + Yₙ → N(2, 4). The continuous mapping theorem: g continuous preserves each of the three modes a.s., in probability and in distribution.

The standard use: √n(X̄ − μ)/σ → N(0, 1) by the CLT and S → σ in probability by the law of large numbers, so √n(X̄ − μ)/S → N(0, 1) as well — which is why the t-statistic is asymptotically normal for any distribution with finite variance. Combined with a first-order Taylor expansion (the delta method), √n(g(X̄) − g(μ)) → N(0, g′(μ)²σ²): for Poisson(4) data, √n(√X̄ − 2) → N(0, (1/4)² × 4) = N(0, 0.25).

⚠️ Slutsky needs a constant limit
If Xₙ → X and Yₙ → Y only in distribution, Xₙ + Yₙ need not converge to X + Y: with Z ~ N(0, 1), take Xₙ = Z and Yₙ = −Z. Both converge in distribution to N(0, 1), but their sum is 0, not N(0, 2). The lemma is safe precisely because one limit is degenerate.

4. The Borel–Cantelli lemmas

For events Aₙ, {Aₙ i.o.} = lim sup Aₙ = ∩ₙ ∪k≥n A_k is the event that infinitely many occur. First lemma: if ΣP(Aₙ) < ∞, then P(Aₙ i.o.) = 0 — no independence needed. Second lemma: if the Aₙ are independent and ΣP(Aₙ) = ∞, then P(Aₙ i.o.) = 1. So for independent Aₙ, P(Aₙ i.o.) is 0 or 1 according as the series converges or diverges: P(Aₙ) = 1/n² gives 0, P(Aₙ) = 1/n gives 1.

🎯 Why the second lemma needs independence
Take one U ~ U(0, 1) and Aₙ = {U < 1/n}. ΣP(Aₙ) = Σ1/n = ∞, but the Aₙ are nested and only finitely many contain any given U > 0, so P(Aₙ i.o.) = P(U = 0) = 0. Dependence lets the events pile onto the same outcomes.

5. The laws of large numbers and the central limit theorem

  • Weak law (Chebyshev): uncorrelated Xᵢ with common mean μ and variance σ² give P(|X̄ − μ| ≥ ε) ≤ σ²/(nε²) → 0. Khintchine: for i.i.d. Xᵢ a finite mean suffices, X̄ → μ in probability.
  • Strong law (Kolmogorov): for i.i.d. Xᵢ, X̄ → μ almost surely iff E|X₁| < ∞, and then μ = E X₁. For Cauchy data neither law applies, and X̄ is Cauchy for every n.
  • Central limit theorem (Lindeberg–Lévy): for i.i.d. Xᵢ with mean μ and finite variance σ² > 0, √n(X̄ − μ)/σ → N(0, 1) in distribution. Nothing else about the distribution is needed.

Worked: with mean 2, variance 4 and n = 100, X̄ is approximately N(2, 0.04), so P(X̄ > 2.4) ≈ P(Z > 2) = 1 − 0.9772 = 0.0228. For a count, apply a continuity correction: in 400 fair tosses, S has mean 200 and SD 10, and P(S ≥ 220) ≈ P(Z ≥ (219.5 − 200)/10) = P(Z ≥ 1.95) = 0.0256. Chebyshev instead gives the sample size guaranteeing P(|X̄ − μ| ≥ 0.1) ≤ 0.05 when σ² = 1: n ≥ 1/(0.01 × 0.05) = 2000.

ℹ️ Extreme values converge too
Not every limit is normal. For i.i.d. U(0, 1), P(n X₍₁₎ > x) = (1 − x/n)ⁿ → e−x: the scaled minimum converges in distribution to Exp(1). Such limits come straight from the definition of convergence in distribution, with no central limit theorem involved.

Key takeaways

  • a.s. ⇒ P ⇒ D; rth mean ⇒ P; D to a constant ⇒ P. Every other arrow has a counterexample.
  • n·1{U < 1/n} → 0 a.s. but not in mean; independent 1{w.p. 1/n} → 0 in probability but not a.s.
  • Slutsky: Xₙ → X in D and Yₙ → c in P give Xₙ + Yₙ → X + c, XₙYₙ → cX, Xₙ/Yₙ → X/c.
  • Borel–Cantelli: ΣP(Aₙ) < ∞ ⇒ P(i.o.) = 0 always; ΣP(Aₙ) = ∞ ⇒ P(i.o.) = 1 only for independent events.
  • WLLN needs a finite mean (i.i.d.); SLLN holds iff E|X| < ∞; CLT needs a finite variance, and counts take a continuity correction.

Practice questions (13)

Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.

  1. Which implications between modes of convergence always hold?

    1. Almost sure convergence implies convergence in probability
    2. Convergence in probability implies almost sure convergence
    3. Convergence in distribution to a constant implies convergence in probability to it
    4. Convergence in second mean implies convergence in first mean
    Show answer

    Answer: A — Almost sure convergence implies convergence in probability; C — Convergence in distribution to a constant implies convergence in probability to it; D — Convergence in second mean implies convergence in first mean

    (A) Standard. (B) False: independent Xₙ with P(Xₙ = 1) = 1/n converge to 0 in probability, but by the second Borel–Cantelli lemma Xₙ = 1 infinitely often. (C) P(|Xₙ − c| > ε) ≤ 1 − Fₙ(c + ε) + Fₙ(c − ε) → 0, both c ± ε being continuity points. (D) Lyapunov (or Cauchy–Schwarz): E|Y| ≤ (E Y²)1/2.
  2. Xₙ are independent with P(Xₙ = n) = 1/n and P(Xₙ = 0) = 1 − 1/n. Then Xₙ converges to 0:

    1. in probability, but neither almost surely nor in first mean
    2. almost surely and in first mean
    3. in first mean but not in probability
    4. in no sense at all
    Show answer

    Answer: A — in probability, but neither almost surely nor in first mean

    P(|Xₙ| > ε) = 1/n → 0, so in probability. E|Xₙ| = n × 1/n = 1 for every n, so not in first mean. Σ1/n = ∞ and the events are independent, so by the second Borel–Cantelli lemma Xₙ = n infinitely often: not almost surely. Convergence in probability does not need the values to stay bounded.
  3. Let U ~ U(0, 1) be a single random variable and Xₙ = n·1{U < 1/n}. Which statements are true?

    1. Xₙ → 0 almost surely
    2. Xₙ → 0 in probability
    3. Xₙ → 0 in first mean
    4. E(Xₙ) → 0
    Show answer

    Answer: A — Xₙ → 0 almost surely; B — Xₙ → 0 in probability

    (A) For any fixed U = u > 0, Xₙ = 0 as soon as n > 1/u, and P(U = 0) = 0. (B) follows from (A). (C), (D) False: E Xₙ = n × P(U < 1/n) = 1 for every n. All the mass escapes to a shrinking set where the value explodes — the standard example that a.s. convergence does not imply convergence of means.
  4. X₁, …, X₁₀₀ are i.i.d. with mean 2 and variance 4. Using the central limit theorem and Φ(2) = 0.9772, the approximate value of P(X̄ > 2.4), correct to four decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.0228

    X̄ ≈ N(2, 4/100), SD 0.2. z = (2.4 − 2)/0.2 = 2 and P(Z > 2) = 1 − 0.9772 = 0.0228. Using the SD of a single observation, 2, gives z = 0.2 and forgets that averaging shrinks the spread by √n.
  5. A fair coin is tossed 400 times and S is the number of heads. Using the normal approximation with a continuity correction and Φ(1.95) = 0.9744, the approximate value of P(S ≥ 220), correct to four decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.0256

    E S = 200, SD = √(400 × 0.25) = 10. P(S ≥ 220) = P(S > 219.5) ≈ P(Z > 1.95) = 1 − 0.9744 = 0.0256. Without the correction, z = 2 gives 0.0228; with the correction applied the wrong way (220.5), z = 2.05.
  6. i.i.d. observations have variance 1. Using Chebyshev’s inequality, the smallest sample size n that guarantees P(|X̄ − μ| ≥ 0.1) ≤ 0.05 is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 2000

    P(|X̄ − μ| ≥ 0.1) ≤ σ²/(n × 0.01) = 100/n ≤ 0.05 needs n ≥ 2000. The normal approximation would need only about (1.96/0.1)² ≈ 385, but that is an approximation, not a guarantee, and Chebyshev is what the question asks for.
  7. Xₙ → N(0, 1) in distribution and Yₙ → 2 in probability. Then XₙYₙ + Yₙ converges in distribution to:

    1. N(2, 4)
    2. N(2, 2)
    3. N(0, 4)
    4. N(2, 1)
    Show answer

    Answer: A — N(2, 4)

    By Slutsky, XₙYₙ → 2Z and adding Yₙ → 2 gives 2Z + 2 ~ N(2, 4): the variance scales by 2² = 4. N(2, 2) scales the variance by 2 instead of 4, and N(0, 4) drops the added constant.
  8. X₁, …, Xₙ are i.i.d. Poisson(4). The limiting distribution of √n(√X̄ − 2) is N(0, v). The value of v, correct to two decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.25

    CLT: √n(X̄ − 4) → N(0, 4). With g(x) = √x, g′(4) = 1/(2√4) = 1/4, so the delta method gives variance (1/4)² × 4 = 0.25 — the square root is the variance-stabilising transform for the Poisson, and 1/4 would be the answer for every λ. Answering 4 forgets the factor g′(μ)².
  9. Which statements about events Aₙ are true?

    1. If P(Aₙ) = 1/n², then P(Aₙ infinitely often) = 0
    2. If the Aₙ are independent and P(Aₙ) = 1/n, then P(Aₙ infinitely often) = 1
    3. The first Borel–Cantelli lemma requires the Aₙ to be independent
    4. Without independence, ΣP(Aₙ) = ∞ does not force P(Aₙ infinitely often) = 1
    Show answer

    Answer: A — If P(Aₙ) = 1/n², then P(Aₙ infinitely often) = 0; B — If the Aₙ are independent and P(Aₙ) = 1/n, then P(Aₙ infinitely often) = 1; D — Without independence, ΣP(Aₙ) = ∞ does not force P(Aₙ infinitely often) = 1

    (A) Σ1/n² < ∞: first lemma. (B) Σ1/n = ∞ with independence: second lemma. (C) False: the first lemma is a union bound, P(∪k≥nA_k) ≤ Σk≥nP(A_k) → 0, with no independence. (D) Aₙ = {U < 1/n} for one uniform U has ΣP = ∞ but P(i.o.) = 0.
  10. For i.i.d. X₁, X₂, … with E|X₁| < ∞ and mean μ, the strong law of large numbers states that:

    1. X̄ₙ → μ almost surely
    2. X̄ₙ → μ in probability only
    3. √n(X̄ₙ − μ) → N(0, σ²)
    4. X̄ₙ = μ for all large n
    Show answer

    Answer: A — X̄ₙ → μ almost surely

    Kolmogorov’s SLLN gives almost sure convergence, which is stronger than the in-probability statement of the weak law. The √n statement is the CLT, which also needs a finite variance; and X̄ₙ is random for every n, so it need never equal μ exactly.
  11. X₁, …, Xₙ are i.i.d. U(0, 1) and X₍₁₎ is their minimum. The limit as n → ∞ of P(n X₍₁₎ > 2), correct to three decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.135

    P(nX₍₁₎ > 2) = P(all Xᵢ > 2/n) = (1 − 2/n)ⁿ → e−2 = 0.1353 → 0.135: nX₍₁₎ converges in distribution to Exp(1). Answering 0 reasons that X₍₁₎ → 0, but the factor n exactly compensates.
  12. Xₙ = 1/n with probability 1 and X = 0. With Fₙ and F their distribution functions, which is correct?

    1. Fₙ(0) → 0 ≠ F(0) = 1, yet Xₙ → X in distribution because 0 is not a continuity point of F
    2. Xₙ does not converge to X in distribution because Fₙ(0) does not converge to F(0)
    3. Xₙ converges to X in distribution but not in probability
    4. Fₙ(x) → F(x) for every real x
    Show answer

    Answer: A — Fₙ(0) → 0 ≠ F(0) = 1, yet Xₙ → X in distribution because 0 is not a continuity point of F

    Fₙ(x) = 1{x ≥ 1/n}, which tends to 1 for x > 0 and to 0 for x ≤ 0, so it matches F except at x = 0, where F jumps. The definition exempts discontinuity points, precisely so that a sequence of constants converging to 0 converges in distribution to 0. It converges in probability too, trivially.
  13. i.i.d. observations have mean 3 and variance 9. For n = 100, the upper bound Chebyshev’s inequality gives for P(|X̄ − 3| ≥ 1), correct to two decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.09

    Var X̄ = 9/100 = 0.09, and P(|X̄ − 3| ≥ 1) ≤ 0.09/1² = 0.09. Using the variance of one observation gives 9, a useless bound above 1.