Probability: Axioms, Conditioning and Bayes, Random Variables, Moments and Generating Functions, Transformations and Inequalities

Section 3 of the GATE Statistics paper. The engineering papers ask probability as a toolkit of formulas; this paper asks it as a theory. The chapter follows the syllabus: the axiomatic definition and the properties of a probability function; conditional probability, Bayes’ theorem and independence of events, including the difference between pairwise and mutual independence; random variables, distribution functions, mass and density functions and quantiles; expectation, variance, moments, the moment generating function and the probability generating function, with the characteristic function that always exists; the distribution of a function of a random variable; and the Markov, Chebyshev and Jensen inequalities. Every later section — distributions, estimation, testing — is built from the objects defined here.

1. The axioms and the properties of a probability function

A probability space is (Ω, 𝓕, P): a sample space, a σ-field of events (closed under complements and countable unions), and a function P with P(A) ≥ 0, P(Ω) = 1, and countable additivity — P(∪Aᵢ) = ΣP(Aᵢ) for pairwise disjoint Aᵢ. Everything else is a theorem: P(∅) = 0, P(Aᶜ) = 1 − P(A), A ⊆ B ⇒ P(A) ≤ P(B), P(A ∪ B) = P(A) + P(B) − P(A ∩ B), Boole’s inequality P(∪Aᵢ) ≤ ΣP(Aᵢ), and continuity: if Aₙ ↑ A then P(Aₙ) ↑ P(A), and if Aₙ ↓ A then P(Aₙ) ↓ P(A).

Inclusion–exclusion for three events: P(A ∪ B ∪ C) = ΣP(A) − ΣP(A ∩ B) + P(A ∩ B ∩ C). Bonferroni: P(A ∩ B) ≥ P(A) + P(B) − 1. With P(A) = 0.5, P(B) = 0.4 and P(A ∪ B) = 0.7, P(A ∩ B) = 0.2 and P(A ∩ Bᶜ) = 0.3.

2. Conditional probability, Bayes’ theorem and independence

P(A | B) = P(A ∩ B)/P(B) for P(B) > 0, and for fixed B it is itself a probability function. The multiplication rule P(A₁ ∩ … ∩ Aₙ) = P(A₁)P(A₂ | A₁)…; the total probability rule P(A) = ΣP(A | Bᵢ)P(Bᵢ) over a partition; and Bayes: P(Bⱼ | A) = P(A | Bⱼ)P(Bⱼ)/ΣP(A | Bᵢ)P(Bᵢ). A test with sensitivity 0.95 and false-positive rate 0.05 for a condition of prevalence 0.01 gives P(condition | positive) = 0.0095/(0.0095 + 0.0495) ≈ 0.16 — a low prior dominates an accurate test.

A and B are independent if P(A ∩ B) = P(A)P(B); then so are A and Bᶜ, Aᶜ and Bᶜ. Events A₁, …, Aₙ are mutually independent if the product rule holds for every subcollection — 2ⁿ − n − 1 equations, not just the pairs. Disjoint events with positive probabilities are never independent: P(A ∩ B) = 0 ≠ P(A)P(B).

⚠️ Pairwise is not mutual
Toss two fair coins. A = first is a head, B = second is a head, C = both show the same face. Each has probability 1/2 and each pair intersects with probability 1/4, so the three are pairwise independent. But A ∩ B ⊆ C, so P(A ∩ B ∩ C) = 1/4 ≠ 1/8: not mutually independent.

3. Random variables, distribution functions, mass and density functions, quantiles

A random variable is a measurable X : Ω → ℝ. Its distribution function F(x) = P(X ≤ x) is characterised by three properties: non-decreasing, right-continuous, F(−∞) = 0 and F(∞) = 1. Every such F is the CDF of some random variable. Jumps give point masses: P(X = x) = F(x) − F(x−), and P(a < X ≤ b) = F(b) − F(a). A discrete X has a mass function p(x) with Σp = 1; a continuous X has a density f ≥ 0 with ∫f = 1 and F′ = f at continuity points; a mixed X has both jumps and a continuous part.

The p-th quantile is ξ_p = inf{x : F(x) ≥ p}; the median is ξ0.5. For a continuous strictly increasing F it solves F(ξ) = p: the exponential with mean θ has F(x) = 1 − e−x/θ, so the median is θ ln 2 — for θ = 2, 1.386, below the mean, as the right skew suggests.

4. Expectation, variance, moments and generating functions

E[g(X)] = Σg(x)p(x) or ∫g(x)f(x) dx, defined when the sum or integral converges absolutely. Linearity E[aX + bY] = aE[X] + bE[Y] needs no independence; Var(aX + b) = a²Var X; Var X = E[X²] − (E X)². For X ≥ 0, E X = ∫₀^∞ P(X > x) dx (for non-negative integers, E X = Σk≥1 P(X ≥ k)) — if P(X > x) = 1/(1 + x)², E X = 1.

The three generating functions
FunctionDefinitionWhat it gives
Moment generating functionM(t) = E[etX]E[Xᵏ] = M⁽ᵏ⁾(0), when M exists near 0; MX+Y = M_X M_Y for independent X, Y
Probability generating functionG(s) = E[s^X], X ∈ {0, 1, 2, …}P(X = k) = G⁽ᵏ⁾(0)/k!; G′(1) = E X; G″(1) = E[X(X − 1)]
Characteristic functionφ(t) = E[eitX]always exists; |φ| ≤ 1, φ(0) = 1; real iff X is symmetric about 0

M(t) = (1 − 2t)−3 gives M′(0) = 6 and M″(0) = 48, so E X = 6 and Var X = 48 − 36 = 12 — the gamma distribution with shape 3 and scale 2. G(s) = (0.4 + 0.6s)⁵ gives G″(1) = 5 × 4 × 0.6² = 7.2 = E[X(X − 1)]. An MGF that exists in a neighbourhood of 0 determines the distribution uniquely. It need not exist: the Cauchy distribution has none and no moments, and the lognormal has every moment but no MGF, because E[etX] = ∞ for every t > 0.

5. The distribution of a function of a random variable

  • Distribution-function method: F_Y(y) = P(g(X) ≤ y), worked out directly — always valid, and the safest route when g is not one-to-one.
  • Change of variable for monotone differentiable g: f_Y(y) = f_X(g⁻¹(y)) |d g⁻¹(y)/dy|. For non-monotone g, sum over the roots: Y = X² gives f_Y(y) = [f_X(√y) + f_X(−√y)]/(2√y).
  • Probability integral transform: if F is continuous, F(X) ~ U(0, 1); conversely F⁻¹(U) has CDF F. So −θ ln U is exponential with mean θ, and if X has density 2x on (0, 1), Y = X² = F(X) is uniform, with variance 1/12.

6. The Markov, Chebyshev and Jensen inequalities

  • Markov: for X ≥ 0 and a > 0, P(X ≥ a) ≤ E X/a. With E X = 2, P(X ≥ 10) ≤ 0.2 — whatever the distribution.
  • Chebyshev: P(|X − μ| ≥ kσ) ≤ 1/k², Markov applied to (X − μ)². With μ = 50 and σ = 5, P(40 < X < 60) ≥ 1 − 1/4 = 0.75.
  • Jensen: for convex φ, φ(E X) ≤ E φ(X); reversed for concave φ, with equality only for degenerate X or linear φ. So E[X²] ≥ (E X)², E[e^X] ≥ eE X, E[1/X] ≥ 1/E X for X > 0, and E[ln X] ≤ ln E X.
🎯 Why these bounds matter even though they are loose
For a normal variable P(|X − μ| ≥ 2σ) ≈ 0.046, far below Chebyshev’s 0.25. The value of the inequality is that it holds for every distribution with a finite variance, which is exactly what the weak law of large numbers needs: P(|X̄ − μ| ≥ ε) ≤ σ²/(nε²) → 0.

Key takeaways

  • Three axioms — non-negativity, P(Ω) = 1, countable additivity — give every other rule, including continuity of P along monotone sequences.
  • Bayes turns P(A | B) into P(B | A); a rare condition keeps the posterior low even for an accurate test. Pairwise independence does not imply mutual independence.
  • A CDF is non-decreasing, right-continuous, 0 at −∞ and 1 at ∞; P(X = x) is the jump F(x) − F(x−).
  • E Xᵏ = M⁽ᵏ⁾(0), G′(1) = E X, G″(1) = E X(X − 1). The characteristic function always exists; the MGF may not (Cauchy; lognormal has moments but no MGF).
  • f_Y(y) = f_X(g⁻¹(y))|dg⁻¹/dy| for monotone g; F(X) ~ U(0, 1) for continuous F.
  • Markov P(X ≥ a) ≤ E X/a; Chebyshev P(|X − μ| ≥ kσ) ≤ 1/k²; Jensen φ(E X) ≤ E φ(X) for convex φ.

Practice questions (15)

Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.

  1. A condition has prevalence 1%. A test detects it with probability 0.95 when present and gives a false positive with probability 0.05 when absent. Given a positive result, the probability that the condition is present, correct to two decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.16

    P(+) = 0.95 × 0.01 + 0.05 × 0.99 = 0.0095 + 0.0495 = 0.059. P(present | +) = 0.0095/0.059 = 0.161 → 0.16. Answering 0.95 confuses P(+ | present) with P(present | +).
  2. Two fair coins are tossed. A = {first coin head}, B = {second coin head}, C = {both coins show the same face}. Which statements are true?

    1. A and B are independent
    2. A and C are independent
    3. A, B and C are mutually independent
    4. P(A ∩ B ∩ C) = 1/4
    Show answer

    Answer: A — A and B are independent; B — A and C are independent; D — P(A ∩ B ∩ C) = 1/4

    (A) P(A ∩ B) = 1/4 = (1/2)(1/2). (B) A ∩ C = {HH}, probability 1/4 = P(A)P(C). (C) False: A ∩ B = {HH} ⊆ C, so (D) P(A ∩ B ∩ C) = 1/4, not the 1/8 mutual independence would require. Checking all pairs is not enough; the triple product must hold too.
  3. A random variable has moment generating function M(t) = (1 − 2t)−3 for t < 1/2. Its variance is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 12

    M′(t) = 6(1 − 2t)−4, M″(t) = 48(1 − 2t)−5. E X = 6, E X² = 48, Var X = 48 − 36 = 12 (gamma, shape 3, scale 2: 3 × 2² = 12). Reporting M″(0) = 48 as the variance forgets to subtract (E X)².
  4. A random variable X on {0, 1, 2, …} has probability generating function G(s) = (0.4 + 0.6s)⁵. The value of E[X(X − 1)], correct to one decimal place, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 7.2

    G″(s) = 5 × 4 × 0.6² (0.4 + 0.6s)³, so G″(1) = 20 × 0.36 = 7.2. This is Bin(5, 0.6): E X = 3, Var X = 1.2, and E[X(X − 1)] = Var X + (E X)² − E X = 1.2 + 9 − 3 = 7.2. Giving E X² = 10.2 answers a different moment.
  5. X has density f(x) = 2x on (0, 1). The variance of Y = X², correct to three decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.083

    F_X(x) = x², so Y = F_X(X) ~ U(0, 1) by the probability integral transform: P(Y ≤ y) = P(X ≤ √y) = y. Var Y = 1/12 = 0.0833 → 0.083. Directly: E X² = ∫2x³ = 1/2, E X⁴ = ∫2x⁵ = 1/3, Var = 1/3 − 1/4 = 1/12. Computing Var X = 1/18 = 0.056 answers for X, not Y.
  6. X has distribution function F(x) = 0 for x < 0, x/2 for 0 ≤ x < 1, and 1 for x ≥ 1. P(X = 1) is:

    1. 1/2
    2. 0
    3. 1
    4. 1/4
    Show answer

    Answer: A — 1/2

    P(X = 1) = F(1) − F(1−) = 1 − 1/2 = 1/2: a mixed distribution, uniform with density 1/2 on [0, 1) and an atom of mass 1/2 at 1. Answering 0 applies the continuous-variable rule "points have probability zero" to a CDF that jumps.
  7. X has mean 50 and standard deviation 5. Chebyshev’s inequality gives the lower bound for P(40 < X < 60) as ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.75

    The interval is μ ± 2σ, so P(|X − μ| ≥ 2σ) ≤ 1/4 and P(|X − μ| < 2σ) ≥ 3/4 = 0.75. Quoting the normal value 0.95 assumes a distribution the question does not give.
  8. A non-negative random variable has mean 2. The best upper bound Markov’s inequality gives for P(X ≥ 10) is:

    1. 0.2
    2. 0.04
    3. 0.8
    4. 0.5
    Show answer

    Answer: A — 0.2

    P(X ≥ 10) ≤ E X/10 = 0.2. The value 0.04 = (2/10)² would need a Chebyshev-type bound with a known variance, which is not given; Markov uses only the mean and non-negativity.
  9. X is a positive, non-degenerate random variable with all the expectations below finite. Which inequalities hold?

    1. E[1/X] ≥ 1/E[X]
    2. E[ln X] ≤ ln E[X]
    3. E[X²] ≤ (E[X])²
    4. E[e^X] ≥ eE[X]
    Show answer

    Answer: A — E[1/X] ≥ 1/E[X]; B — E[ln X] ≤ ln E[X]; D — E[e^X] ≥ e^{E[X]}

    Jensen: 1/x and eˣ are convex on (0, ∞), so (A) and (D) hold; ln x is concave, so (B) holds. (C) is reversed: x² is convex, so E[X²] ≥ (E X)², the difference being Var X > 0 for a non-degenerate X.
  10. The median of an exponential distribution with mean 2, correct to two decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 1.39

    F(m) = 1 − e−m/2 = 1/2 gives m = 2 ln 2 = 1.386 → 1.39. Taking the median equal to the mean, 2, ignores the right skew; using the rate as if it were the mean gives ln 2/2 = 0.35.
  11. P(A) = 0.5, P(B) = 0.4 and P(A ∪ B) = 0.7. The value of P(A ∩ Bᶜ) is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.3

    P(A ∩ B) = 0.5 + 0.4 − 0.7 = 0.2, and P(A ∩ Bᶜ) = P(A) − P(A ∩ B) = 0.3. Multiplying P(A)P(Bᶜ) = 0.3 happens to agree here only because 0.2 = 0.5 × 0.4 — these events are independent, which the question did not say and the method should not assume.
  12. Which distribution has finite moments of every order but no moment generating function in any neighbourhood of 0?

    1. Lognormal
    2. Cauchy
    3. Exponential
    4. Poisson
    Show answer

    Answer: A — Lognormal

    If ln X ~ N(μ, σ²), E[Xᵏ] = ekμ + k²σ²/2 is finite for all k, but E[etX] = ∞ for every t > 0 because etx outgrows the lognormal tail. The Cauchy has no moments at all; the exponential’s MGF exists for t below the rate, and the Poisson’s for all t.
  13. A non-negative random variable X has P(X > x) = 1/(1 + x)² for x ≥ 0. E[X] is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 1

    E X = ∫₀^∞ P(X > x) dx = ∫₀^∞ (1 + x)−2 dx = 1. Via the density f(x) = 2/(1 + x)³: ∫₀^∞ 2x/(1 + x)³ dx = 1 as well. Integrating the survival function against x (which gives E[X²]/2) is the usual slip; here E X² is actually infinite.
  14. Which property must every cumulative distribution function F have?

    1. Right-continuity
    2. Left-continuity
    3. Differentiability
    4. Strict monotonicity
    Show answer

    Answer: A — Right-continuity

    F(x) = P(X ≤ x) and {X ≤ x + 1/n} ↓ {X ≤ x}, so continuity of probability gives F(x + 1/n) → F(x): right-continuity. It can jump from the left (at atoms), need not be differentiable (discrete X), and is constant on gaps in the support, so it need not be strictly increasing.
  15. Which statements about the characteristic function φ(t) = E[eitX] are true for every random variable X?

    1. φ(t) exists for every real t
    2. |φ(t)| ≤ 1
    3. φ(0) = 1
    4. φ(t) is real-valued
    Show answer

    Answer: A — φ(t) exists for every real t; B — |φ(t)| ≤ 1; C — φ(0) = 1

    |eitX| = 1, so the expectation always exists and |φ(t)| ≤ E|eitX| = 1; φ(0) = E[1] = 1. (D) fails in general: φ is real exactly when X and −X have the same distribution — for an exponential variable, φ(t) = 1/(1 − it) is complex.