Standard Univariate Distributions: Bernoulli to Poisson, Uniform to Cauchy, and How They Are Related
1. The discrete families
| Distribution | Mass function | Mean | Variance |
|---|---|---|---|
| Bernoulli(p) | pˣq1−x, x = 0, 1 | p | pq |
| Binomial(n, p) | C(n, x)pˣqⁿ⁻ˣ | np | npq |
| Geometric: trials to 1st success | qx−1p, x = 1, 2, … | 1/p | q/p² |
| Negative binomial: failures before r-th success | C(x + r − 1, x)pʳqˣ | rq/p | rq/p² |
| Hypergeometric: n draws, M marked of N | C(M, x)C(N − M, n − x)/C(N, n) | nM/N | (nM/N)(1 − M/N)(N − n)/(N − 1) |
| Discrete uniform on 1, …, k | 1/k | (k + 1)/2 | (k² − 1)/12 |
| Poisson(λ) | e−λλˣ/x! | λ | λ |
The hypergeometric variance carries the finite population correction (N − n)/(N − 1): drawing 5 items without replacement from 20 of which 8 are defective gives mean 2 and variance 5 × 0.4 × 0.6 × 15/19 = 0.947, smaller than the binomial 1.2 because sampling without replacement removes information as it goes.
2. Relations among the discrete families
- Poisson limit: Bin(n, p) with n → ∞, p → 0 and np = λ fixed tends to Poisson(λ). Bin(1000, 0.002) has P(X = 0) = 0.998¹⁰⁰⁰ ≈ e−2.
- Binomial limit: the hypergeometric tends to Bin(n, M/N) as N → ∞ with M/N fixed.
- Sums: independent Bin(n₁, p) + Bin(n₂, p) = Bin(n₁ + n₂, p); Poisson(λ₁) + Poisson(λ₂) = Poisson(λ₁ + λ₂); a sum of r independent geometrics (failure count) is negative binomial (r, p).
- Conditioning: if X, Y are independent Poissons, X | X + Y = n ~ Bin(n, λ₁/(λ₁ + λ₂)).
- Memorylessness: the geometric is the only discrete distribution with P(X > m + n | X > m) = P(X > n). With p = 0.2, P(X > 5 | X > 2) = P(X > 3) = 0.8³ = 0.512.
The Poisson mass function satisfies p(x)/p(x − 1) = λ/x, so it increases while x < λ and its mode is ⌊λ⌋ (both λ − 1 and λ when λ is an integer). With λ = 2: P(X = 0) = e−2 = 0.1353, P(X ≤ 1) = 3e−2 = 0.406.
3. The continuous families
| Distribution | Density | Mean | Variance |
|---|---|---|---|
| Uniform(a, b) | 1/(b − a) | (a + b)/2 | (b − a)²/12 |
| Exponential, rate λ | λe−λx, x > 0 | 1/λ | 1/λ² |
| Double exponential (Laplace) | (1/2β)e−|x − μ|/β | μ | 2β² |
| Gamma(α, rate λ) | λ^α xα−1e−λx/Γ(α) | α/λ | α/λ² |
| Beta, first kind β₁(a, b) | xa−1(1 − x)b−1/B(a, b), 0 < x < 1 | a/(a + b) | ab/[(a + b)²(a + b + 1)] |
| Beta, second kind β₂(a, b) | xa−1/[B(a, b)(1 + x)a+b], x > 0 | a/(b − 1), b > 1 | a(a + b − 1)/[(b − 1)²(b − 2)], b > 2 |
| Weibull(shape k, scale θ) | F = 1 − e−(x/θ)ᵏ | θΓ(1 + 1/k) | θ²[Γ(1 + 2/k) − Γ(1 + 1/k)²] |
| Normal(μ, σ²) | (σ√(2π))−1e−(x − μ)²/2σ² | μ | σ² |
| Cauchy(μ, σ) | σ/[π(σ² + (x − μ)²)] | does not exist | does not exist |
The generating functions to remember: Poisson eλ(eᵗ − 1); binomial (q + peᵗ)ⁿ; exponential λ/(λ − t) for t < λ; gamma (1 − t/λ)−α; normal eμt + σ²t²/2; Laplace (μ = 0, β = 1) 1/(1 − t²) for |t| < 1; Cauchy characteristic function eiμt − σ|t|, which is not differentiable at t = 0 — the signature of a missing mean.
4. Exponential, gamma, beta and Weibull: the relations
- Memoryless: P(X > s + t | X > s) = P(X > t) characterises the exponential among continuous distributions; its hazard rate f/(1 − F) = λ is constant.
- Sums: n independent Exp(λ) sum to Gamma(n, λ); independent gammas with the same rate add their shapes. The minimum of independent exponentials is exponential with the rates added.
- Ratios: X ~ Gamma(a, λ), Y ~ Gamma(b, λ) independent ⇒ X/(X + Y) ~ β₁(a, b), independent of X + Y, and X/Y ~ β₂(a, b). If U ~ β₁(a, b) then U/(1 − U) ~ β₂(a, b).
- Laplace: the difference of two independent Exp(1) variables has density ½e−|x|, variance 1 + 1 = 2.
- Weibull: if E ~ Exp(1), θE1/k is Weibull(k, θ); k = 1 is the exponential, k > 1 an increasing hazard. The median is θ(ln 2)1/k: for k = 2, θ = 1, it is 0.833.
- Uniform: the sum of two independent U(0, 1) variables is triangular on (0, 2) with peak at 1, not uniform; −ln U ~ Exp(1).
5. The normal and the Cauchy
If X ~ N(μ, σ²), Z = (X − μ)/σ ~ N(0, 1), and linear combinations of independent normals are normal. Z has odd moments 0, E Z² = 1, E Z⁴ = 3 (so the kurtosis is 3), and E|Z| = √(2/π) ≈ 0.798. P(|Z| ≤ 1) ≈ 0.683, P(|Z| ≤ 2) ≈ 0.954, P(|Z| ≤ 3) ≈ 0.997. For X ~ N(10, 4), P(X > 12) = P(Z > 1) = 1 − Φ(1) = 0.1587.
The Cauchy has symmetric, bell-shaped density but tails like 1/x², so E|X| = ∞: no mean, no variance. Its median and location are μ, and its quartiles are μ ± σ. It arises as the ratio of two independent standard normals and as tan(πU − π/2). The average of n independent standard Cauchy variables is again standard Cauchy — the characteristic function e−|t| raised to the n and evaluated at t/n gives e−|t| back — so the sample mean does not settle down however large n is.
Key takeaways
- Binomial npq, Poisson mean = variance = λ, geometric (trials) 1/p and q/p², negative binomial (failures) rq/p and rq/p²; the hypergeometric carries (N − n)/(N − 1).
- Geometric and exponential are the memoryless ones; Bin(n, λ/n) → Poisson(λ); the minimum of exponentials adds the rates.
- Gamma(α, λ) has mean α/λ and variance α/λ²; X/(X + Y) of two same-rate gammas is β₁, X/Y is β₂.
- β₁(a, b): mean a/(a + b), variance ab/[(a + b)²(a + b + 1)]; Laplace variance 2β²; Weibull median θ(ln 2)1/k.
- E|Z| = √(2/π), E Z⁴ = 3; the Cauchy has no mean, its sample mean is Cauchy again, and its characteristic function e−|t| has a corner at 0.
Practice questions (15)
Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.
The variance of a Bin(10, 0.3) random variable, correct to one decimal place, is ____.
Numerical answer — type the value.
Show answer
Answer: 2.1
Var = npq = 10 × 0.3 × 0.7 = 2.1. The mean np = 3 is the usual wrong answer, and np² = 0.9 drops the factor q.X ~ Poisson(2). P(X ≤ 1), correct to three decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.406
P(0) + P(1) = e−2(1 + 2) = 3 × 0.13534 = 0.406. Stopping at P(X = 0) = 0.135 answers P(X < 1).X is the number of Bernoulli(0.2) trials needed for the first success. P(X > 5 | X > 2), correct to three decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.512
P(X > k) = 0.8ᵏ (the first k trials fail). So P(X > 5 | X > 2) = 0.8⁵/0.8² = 0.8³ = 0.512 — memorylessness: having waited 2 trials changes nothing. Answering 0.8⁵ = 0.328 ignores the conditioning.X is the number of failures before the 3rd success in independent trials with success probability 0.5. The variance of X is ____.
Numerical answer — type the value.
Show answer
Answer: 6
X is a sum of 3 independent geometric failure counts, each with variance q/p² = 0.5/0.25 = 2, so Var X = rq/p² = 6. The mean is rq/p = 3. Using rq, the binomial-looking formula, gives 1.5.Five items are drawn without replacement from a lot of 20 of which 8 are defective. The variance of the number of defectives drawn, correct to two decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.95
Hypergeometric: n(M/N)(1 − M/N)(N − n)/(N − 1) = 5 × 0.4 × 0.6 × 15/19 = 1.2 × 0.7895 = 0.947 → 0.95. Omitting the finite population correction gives the binomial 1.2.X has the double exponential density f(x) = ½e−|x|, −∞ < x < ∞. The variance of X is ____.
Numerical answer — type the value.
Show answer
Answer: 2
By symmetry E X = 0, and E X² = ∫₀^∞ x²e−x dx = Γ(3) = 2. Equivalently X is the difference of two independent Exp(1) variables, so Var = 1 + 1 = 2. Answering 1 treats |X| as if its second moment were the exponential’s variance.X follows the beta distribution of the first kind with parameters a = 2 and b = 3. The variance of X, correct to two decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.04
ab/[(a + b)²(a + b + 1)] = 6/(25 × 6) = 0.04; the mean is 2/5 = 0.4. Dropping the factor (a + b + 1) gives 0.24, suspiciously close to 1/4 — the largest variance any variable on (0, 1) can have, and a sanity check worth running on every beta answer.Which statements are true? (All variables in each statement are independent.)
Show answer
Answer: A — If X, Y ~ Exp(1), then X − Y has the double exponential distribution; B — If X ~ Gamma(a, λ) and Y ~ Gamma(b, λ), then X/(X + Y) ~ β₁(a, b); C — With X, Y as in the previous option, X/Y ~ β₂(a, b)
(A) The convolution of e−x and ex on the appropriate ranges gives ½e−|z|. (B) The Jacobian of (x, y) → (x + y, x/(x + y)) separates the joint density into a gamma in the sum and a beta in the proportion. (C) X/Y = V/(1 − V) with V ~ β₁(a, b), which is β₂(a, b). (D) False: the sum has the triangular density z on (0, 1) and 2 − z on (1, 2).Let X₁, …, Xₙ be independent standard Cauchy variables. Which statements are true?
Show answer
Answer: A — E(X₁) does not exist; C — The sample mean X̄ has the standard Cauchy distribution
(A) ∫|x|/(π(1 + x²)) dx diverges logarithmically. (B) False: φ(t) = e−|t| has a corner at 0, and differentiability at 0 would give a finite mean. (C) φ_X̄(t) = [e−|t|/n]ⁿ = e−|t|. (D) False: no mean, so no variance; the scale 1 is the semi-interquartile range, not a standard deviation.X has the Weibull distribution with P(X ≤ x) = 1 − e−x² for x > 0. The median of X, correct to two decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.83
1 − e−m² = 1/2 gives m² = ln 2 and m = √0.6931 = 0.8326 → 0.83. X is E1/2 for E ~ Exp(1), and the median of E is ln 2 = 0.69 — forgetting the square root answers for E, not X.X ~ N(10, 4), where 4 is the variance. Using Φ(1) = 0.8413, P(X > 12), correct to four decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.1587
σ = 2, so z = (12 − 10)/2 = 1 and P(X > 12) = 1 − 0.8413 = 0.1587. Dividing by the variance 4 gives z = 0.5, the most common slip when the second parameter of N(μ, σ²) is read as σ.For Z ~ N(0, 1), the value of E|Z|, correct to two decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.80
E|Z| = 2∫₀^∞ z φ(z) dz = 2 × (1/√(2π)) × ∫₀^∞ z e−z²/2 dz = 2/√(2π) = √(2/π) = 0.7979 → 0.80. E Z = 0 by symmetry is not the answer: the absolute value removes the cancellation.X ~ Bin(1000, 0.002). The Poisson approximation to P(X = 0) is:
Show answer
Answer: A — e^{−2}
λ = np = 2, so P(X = 0) ≈ e−2 = 0.1353 (exactly 0.998¹⁰⁰⁰ = 0.1351). e−0.002 uses p in place of np, and 2e−2 is P(X = 1).X follows the beta distribution of the second kind β₂(m, n), with density ∝ xm−1/(1 + x)m+n on x > 0. For n > 1, E(X) equals:
Show answer
Answer: A — m/(n − 1)
E X = B(m + 1, n − 1)/B(m, n) = [Γ(m + 1)Γ(n − 1)]/[Γ(m)Γ(n)] = m/(n − 1), finite only for n > 1 because the tail is x−n−1. m/(m + n) is the mean of the first-kind beta; confusing the two kinds is the trap.Which continuous distribution on (0, ∞) is memoryless, i.e. P(X > s + t | X > s) = P(X > t) for all s, t > 0?
Show answer
Answer: A — Exponential
Memorylessness means the survival function satisfies S(s + t) = S(s)S(t); the only monotone solutions are S(t) = e−λt. Gamma(2) and Weibull(2) have increasing hazard (older means likelier to fail), and a uniform on (0, 1) cannot even survive past 1.