Probability: Axioms, Conditioning and Bayes, Random Variables, Moments and Generating Functions, Transformations and Inequalities
1. The axioms and the properties of a probability function
A probability space is (Ω, 𝓕, P): a sample space, a σ-field of events (closed under complements and countable unions), and a function P with P(A) ≥ 0, P(Ω) = 1, and countable additivity — P(∪Aᵢ) = ΣP(Aᵢ) for pairwise disjoint Aᵢ. Everything else is a theorem: P(∅) = 0, P(Aᶜ) = 1 − P(A), A ⊆ B ⇒ P(A) ≤ P(B), P(A ∪ B) = P(A) + P(B) − P(A ∩ B), Boole’s inequality P(∪Aᵢ) ≤ ΣP(Aᵢ), and continuity: if Aₙ ↑ A then P(Aₙ) ↑ P(A), and if Aₙ ↓ A then P(Aₙ) ↓ P(A).
Inclusion–exclusion for three events: P(A ∪ B ∪ C) = ΣP(A) − ΣP(A ∩ B) + P(A ∩ B ∩ C). Bonferroni: P(A ∩ B) ≥ P(A) + P(B) − 1. With P(A) = 0.5, P(B) = 0.4 and P(A ∪ B) = 0.7, P(A ∩ B) = 0.2 and P(A ∩ Bᶜ) = 0.3.
2. Conditional probability, Bayes’ theorem and independence
P(A | B) = P(A ∩ B)/P(B) for P(B) > 0, and for fixed B it is itself a probability function. The multiplication rule P(A₁ ∩ … ∩ Aₙ) = P(A₁)P(A₂ | A₁)…; the total probability rule P(A) = ΣP(A | Bᵢ)P(Bᵢ) over a partition; and Bayes: P(Bⱼ | A) = P(A | Bⱼ)P(Bⱼ)/ΣP(A | Bᵢ)P(Bᵢ). A test with sensitivity 0.95 and false-positive rate 0.05 for a condition of prevalence 0.01 gives P(condition | positive) = 0.0095/(0.0095 + 0.0495) ≈ 0.16 — a low prior dominates an accurate test.
A and B are independent if P(A ∩ B) = P(A)P(B); then so are A and Bᶜ, Aᶜ and Bᶜ. Events A₁, …, Aₙ are mutually independent if the product rule holds for every subcollection — 2ⁿ − n − 1 equations, not just the pairs. Disjoint events with positive probabilities are never independent: P(A ∩ B) = 0 ≠ P(A)P(B).
3. Random variables, distribution functions, mass and density functions, quantiles
A random variable is a measurable X : Ω → ℝ. Its distribution function F(x) = P(X ≤ x) is characterised by three properties: non-decreasing, right-continuous, F(−∞) = 0 and F(∞) = 1. Every such F is the CDF of some random variable. Jumps give point masses: P(X = x) = F(x) − F(x−), and P(a < X ≤ b) = F(b) − F(a). A discrete X has a mass function p(x) with Σp = 1; a continuous X has a density f ≥ 0 with ∫f = 1 and F′ = f at continuity points; a mixed X has both jumps and a continuous part.
The p-th quantile is ξ_p = inf{x : F(x) ≥ p}; the median is ξ0.5. For a continuous strictly increasing F it solves F(ξ) = p: the exponential with mean θ has F(x) = 1 − e−x/θ, so the median is θ ln 2 — for θ = 2, 1.386, below the mean, as the right skew suggests.
4. Expectation, variance, moments and generating functions
E[g(X)] = Σg(x)p(x) or ∫g(x)f(x) dx, defined when the sum or integral converges absolutely. Linearity E[aX + bY] = aE[X] + bE[Y] needs no independence; Var(aX + b) = a²Var X; Var X = E[X²] − (E X)². For X ≥ 0, E X = ∫₀^∞ P(X > x) dx (for non-negative integers, E X = Σk≥1 P(X ≥ k)) — if P(X > x) = 1/(1 + x)², E X = 1.
| Function | Definition | What it gives |
|---|---|---|
| Moment generating function | M(t) = E[etX] | E[Xᵏ] = M⁽ᵏ⁾(0), when M exists near 0; MX+Y = M_X M_Y for independent X, Y |
| Probability generating function | G(s) = E[s^X], X ∈ {0, 1, 2, …} | P(X = k) = G⁽ᵏ⁾(0)/k!; G′(1) = E X; G″(1) = E[X(X − 1)] |
| Characteristic function | φ(t) = E[eitX] | always exists; |φ| ≤ 1, φ(0) = 1; real iff X is symmetric about 0 |
M(t) = (1 − 2t)−3 gives M′(0) = 6 and M″(0) = 48, so E X = 6 and Var X = 48 − 36 = 12 — the gamma distribution with shape 3 and scale 2. G(s) = (0.4 + 0.6s)⁵ gives G″(1) = 5 × 4 × 0.6² = 7.2 = E[X(X − 1)]. An MGF that exists in a neighbourhood of 0 determines the distribution uniquely. It need not exist: the Cauchy distribution has none and no moments, and the lognormal has every moment but no MGF, because E[etX] = ∞ for every t > 0.
5. The distribution of a function of a random variable
- Distribution-function method: F_Y(y) = P(g(X) ≤ y), worked out directly — always valid, and the safest route when g is not one-to-one.
- Change of variable for monotone differentiable g: f_Y(y) = f_X(g⁻¹(y)) |d g⁻¹(y)/dy|. For non-monotone g, sum over the roots: Y = X² gives f_Y(y) = [f_X(√y) + f_X(−√y)]/(2√y).
- Probability integral transform: if F is continuous, F(X) ~ U(0, 1); conversely F⁻¹(U) has CDF F. So −θ ln U is exponential with mean θ, and if X has density 2x on (0, 1), Y = X² = F(X) is uniform, with variance 1/12.
6. The Markov, Chebyshev and Jensen inequalities
- Markov: for X ≥ 0 and a > 0, P(X ≥ a) ≤ E X/a. With E X = 2, P(X ≥ 10) ≤ 0.2 — whatever the distribution.
- Chebyshev: P(|X − μ| ≥ kσ) ≤ 1/k², Markov applied to (X − μ)². With μ = 50 and σ = 5, P(40 < X < 60) ≥ 1 − 1/4 = 0.75.
- Jensen: for convex φ, φ(E X) ≤ E φ(X); reversed for concave φ, with equality only for degenerate X or linear φ. So E[X²] ≥ (E X)², E[e^X] ≥ eE X, E[1/X] ≥ 1/E X for X > 0, and E[ln X] ≤ ln E X.
Key takeaways
- Three axioms — non-negativity, P(Ω) = 1, countable additivity — give every other rule, including continuity of P along monotone sequences.
- Bayes turns P(A | B) into P(B | A); a rare condition keeps the posterior low even for an accurate test. Pairwise independence does not imply mutual independence.
- A CDF is non-decreasing, right-continuous, 0 at −∞ and 1 at ∞; P(X = x) is the jump F(x) − F(x−).
- E Xᵏ = M⁽ᵏ⁾(0), G′(1) = E X, G″(1) = E X(X − 1). The characteristic function always exists; the MGF may not (Cauchy; lognormal has moments but no MGF).
- f_Y(y) = f_X(g⁻¹(y))|dg⁻¹/dy| for monotone g; F(X) ~ U(0, 1) for continuous F.
- Markov P(X ≥ a) ≤ E X/a; Chebyshev P(|X − μ| ≥ kσ) ≤ 1/k²; Jensen φ(E X) ≤ E φ(X) for convex φ.
Practice questions (15)
Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.
A condition has prevalence 1%. A test detects it with probability 0.95 when present and gives a false positive with probability 0.05 when absent. Given a positive result, the probability that the condition is present, correct to two decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.16
P(+) = 0.95 × 0.01 + 0.05 × 0.99 = 0.0095 + 0.0495 = 0.059. P(present | +) = 0.0095/0.059 = 0.161 → 0.16. Answering 0.95 confuses P(+ | present) with P(present | +).Two fair coins are tossed. A = {first coin head}, B = {second coin head}, C = {both coins show the same face}. Which statements are true?
Show answer
Answer: A — A and B are independent; B — A and C are independent; D — P(A ∩ B ∩ C) = 1/4
(A) P(A ∩ B) = 1/4 = (1/2)(1/2). (B) A ∩ C = {HH}, probability 1/4 = P(A)P(C). (C) False: A ∩ B = {HH} ⊆ C, so (D) P(A ∩ B ∩ C) = 1/4, not the 1/8 mutual independence would require. Checking all pairs is not enough; the triple product must hold too.A random variable has moment generating function M(t) = (1 − 2t)−3 for t < 1/2. Its variance is ____.
Numerical answer — type the value.
Show answer
Answer: 12
M′(t) = 6(1 − 2t)−4, M″(t) = 48(1 − 2t)−5. E X = 6, E X² = 48, Var X = 48 − 36 = 12 (gamma, shape 3, scale 2: 3 × 2² = 12). Reporting M″(0) = 48 as the variance forgets to subtract (E X)².A random variable X on {0, 1, 2, …} has probability generating function G(s) = (0.4 + 0.6s)⁵. The value of E[X(X − 1)], correct to one decimal place, is ____.
Numerical answer — type the value.
Show answer
Answer: 7.2
G″(s) = 5 × 4 × 0.6² (0.4 + 0.6s)³, so G″(1) = 20 × 0.36 = 7.2. This is Bin(5, 0.6): E X = 3, Var X = 1.2, and E[X(X − 1)] = Var X + (E X)² − E X = 1.2 + 9 − 3 = 7.2. Giving E X² = 10.2 answers a different moment.X has density f(x) = 2x on (0, 1). The variance of Y = X², correct to three decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.083
F_X(x) = x², so Y = F_X(X) ~ U(0, 1) by the probability integral transform: P(Y ≤ y) = P(X ≤ √y) = y. Var Y = 1/12 = 0.0833 → 0.083. Directly: E X² = ∫2x³ = 1/2, E X⁴ = ∫2x⁵ = 1/3, Var = 1/3 − 1/4 = 1/12. Computing Var X = 1/18 = 0.056 answers for X, not Y.X has distribution function F(x) = 0 for x < 0, x/2 for 0 ≤ x < 1, and 1 for x ≥ 1. P(X = 1) is:
Show answer
Answer: A — 1/2
P(X = 1) = F(1) − F(1−) = 1 − 1/2 = 1/2: a mixed distribution, uniform with density 1/2 on [0, 1) and an atom of mass 1/2 at 1. Answering 0 applies the continuous-variable rule "points have probability zero" to a CDF that jumps.X has mean 50 and standard deviation 5. Chebyshev’s inequality gives the lower bound for P(40 < X < 60) as ____.
Numerical answer — type the value.
Show answer
Answer: 0.75
The interval is μ ± 2σ, so P(|X − μ| ≥ 2σ) ≤ 1/4 and P(|X − μ| < 2σ) ≥ 3/4 = 0.75. Quoting the normal value 0.95 assumes a distribution the question does not give.A non-negative random variable has mean 2. The best upper bound Markov’s inequality gives for P(X ≥ 10) is:
Show answer
Answer: A — 0.2
P(X ≥ 10) ≤ E X/10 = 0.2. The value 0.04 = (2/10)² would need a Chebyshev-type bound with a known variance, which is not given; Markov uses only the mean and non-negativity.X is a positive, non-degenerate random variable with all the expectations below finite. Which inequalities hold?
Show answer
Answer: A — E[1/X] ≥ 1/E[X]; B — E[ln X] ≤ ln E[X]; D — E[e^X] ≥ e^{E[X]}
Jensen: 1/x and eˣ are convex on (0, ∞), so (A) and (D) hold; ln x is concave, so (B) holds. (C) is reversed: x² is convex, so E[X²] ≥ (E X)², the difference being Var X > 0 for a non-degenerate X.The median of an exponential distribution with mean 2, correct to two decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 1.39
F(m) = 1 − e−m/2 = 1/2 gives m = 2 ln 2 = 1.386 → 1.39. Taking the median equal to the mean, 2, ignores the right skew; using the rate as if it were the mean gives ln 2/2 = 0.35.P(A) = 0.5, P(B) = 0.4 and P(A ∪ B) = 0.7. The value of P(A ∩ Bᶜ) is ____.
Numerical answer — type the value.
Show answer
Answer: 0.3
P(A ∩ B) = 0.5 + 0.4 − 0.7 = 0.2, and P(A ∩ Bᶜ) = P(A) − P(A ∩ B) = 0.3. Multiplying P(A)P(Bᶜ) = 0.3 happens to agree here only because 0.2 = 0.5 × 0.4 — these events are independent, which the question did not say and the method should not assume.Which distribution has finite moments of every order but no moment generating function in any neighbourhood of 0?
Show answer
Answer: A — Lognormal
If ln X ~ N(μ, σ²), E[Xᵏ] = ekμ + k²σ²/2 is finite for all k, but E[etX] = ∞ for every t > 0 because etx outgrows the lognormal tail. The Cauchy has no moments at all; the exponential’s MGF exists for t below the rate, and the Poisson’s for all t.A non-negative random variable X has P(X > x) = 1/(1 + x)² for x ≥ 0. E[X] is ____.
Numerical answer — type the value.
Show answer
Answer: 1
E X = ∫₀^∞ P(X > x) dx = ∫₀^∞ (1 + x)−2 dx = 1. Via the density f(x) = 2/(1 + x)³: ∫₀^∞ 2x/(1 + x)³ dx = 1 as well. Integrating the survival function against x (which gives E[X²]/2) is the usual slip; here E X² is actually infinite.Which property must every cumulative distribution function F have?
Show answer
Answer: A — Right-continuity
F(x) = P(X ≤ x) and {X ≤ x + 1/n} ↓ {X ≤ x}, so continuity of probability gives F(x + 1/n) → F(x): right-continuity. It can jump from the left (at atoms), need not be differentiable (discrete X), and is constant on gaps in the support, so it need not be strictly increasing.Which statements about the characteristic function φ(t) = E[eitX] are true for every random variable X?
Show answer
Answer: A — φ(t) exists for every real t; B — |φ(t)| ≤ 1; C — φ(0) = 1
|eitX| = 1, so the expectation always exists and |φ(t)| ≤ E|eitX| = 1; φ(0) = E[1] = 1. (D) fails in general: φ is real exactly when X and −X have the same distribution — for an exponential variable, φ(t) = 1/(1 − it) is complex.