Standard Univariate Distributions: Bernoulli to Poisson, Uniform to Cauchy, and How They Are Related

Section 4 of the GATE Statistics paper names fifteen distributions: Bernoulli, binomial, geometric, negative binomial, hypergeometric, discrete uniform, Poisson, continuous uniform, exponential, double exponential, gamma, beta of the first and second kinds, Weibull, normal and Cauchy. The paper expects each one’s mass or density function, mean, variance and generating function to be known at sight, and — more often — the relations between them: the Poisson as a limit of the binomial, the memoryless property that singles out the geometric and the exponential, sums of exponentials as gammas, ratios of gammas as betas, and the Cauchy as the distribution with no mean. This chapter collects them in tables, then teaches the relations with the numbers worked.

1. The discrete families

Discrete distributions (q = 1 − p)
DistributionMass functionMeanVariance
Bernoulli(p)pˣq1−x, x = 0, 1ppq
Binomial(n, p)C(n, x)pˣqⁿ⁻ˣnpnpq
Geometric: trials to 1st successqx−1p, x = 1, 2, …1/pq/p²
Negative binomial: failures before r-th successC(x + r − 1, x)pʳqˣrq/prq/p²
Hypergeometric: n draws, M marked of NC(M, x)C(N − M, n − x)/C(N, n)nM/N(nM/N)(1 − M/N)(N − n)/(N − 1)
Discrete uniform on 1, …, k1/k(k + 1)/2(k² − 1)/12
Poisson(λ)e−λλˣ/x!λλ
⚠️ Which geometric, which negative binomial
Two conventions are in use. Counting trials, the geometric has mean 1/p; counting failures before the first success, it has mean q/p — the variance q/p² is the same for both, because the two variables differ by the constant 1. Read the stem for which count is meant before quoting a formula; the negative binomial has the same two versions, with means r/p and rq/p.

The hypergeometric variance carries the finite population correction (N − n)/(N − 1): drawing 5 items without replacement from 20 of which 8 are defective gives mean 2 and variance 5 × 0.4 × 0.6 × 15/19 = 0.947, smaller than the binomial 1.2 because sampling without replacement removes information as it goes.

2. Relations among the discrete families

  • Poisson limit: Bin(n, p) with n → ∞, p → 0 and np = λ fixed tends to Poisson(λ). Bin(1000, 0.002) has P(X = 0) = 0.998¹⁰⁰⁰ ≈ e−2.
  • Binomial limit: the hypergeometric tends to Bin(n, M/N) as N → ∞ with M/N fixed.
  • Sums: independent Bin(n₁, p) + Bin(n₂, p) = Bin(n₁ + n₂, p); Poisson(λ₁) + Poisson(λ₂) = Poisson(λ₁ + λ₂); a sum of r independent geometrics (failure count) is negative binomial (r, p).
  • Conditioning: if X, Y are independent Poissons, X | X + Y = n ~ Bin(n, λ₁/(λ₁ + λ₂)).
  • Memorylessness: the geometric is the only discrete distribution with P(X > m + n | X > m) = P(X > n). With p = 0.2, P(X > 5 | X > 2) = P(X > 3) = 0.8³ = 0.512.

The Poisson mass function satisfies p(x)/p(x − 1) = λ/x, so it increases while x < λ and its mode is ⌊λ⌋ (both λ − 1 and λ when λ is an integer). With λ = 2: P(X = 0) = e−2 = 0.1353, P(X ≤ 1) = 3e−2 = 0.406.

3. The continuous families

Continuous distributions
DistributionDensityMeanVariance
Uniform(a, b)1/(b − a)(a + b)/2(b − a)²/12
Exponential, rate λλe−λx, x > 01/λ1/λ²
Double exponential (Laplace)(1/2β)e−|x − μ|/βμ2β²
Gamma(α, rate λ)λ^α xα−1e−λx/Γ(α)α/λα/λ²
Beta, first kind β₁(a, b)xa−1(1 − x)b−1/B(a, b), 0 < x < 1a/(a + b)ab/[(a + b)²(a + b + 1)]
Beta, second kind β₂(a, b)xa−1/[B(a, b)(1 + x)a+b], x > 0a/(b − 1), b > 1a(a + b − 1)/[(b − 1)²(b − 2)], b > 2
Weibull(shape k, scale θ)F = 1 − e−(x/θ)ᵏθΓ(1 + 1/k)θ²[Γ(1 + 2/k) − Γ(1 + 1/k)²]
Normal(μ, σ²)(σ√(2π))−1e−(x − μ)²/2σ²μσ²
Cauchy(μ, σ)σ/[π(σ² + (x − μ)²)]does not existdoes not exist

The generating functions to remember: Poisson eλ(eᵗ − 1); binomial (q + peᵗ)ⁿ; exponential λ/(λ − t) for t < λ; gamma (1 − t/λ)−α; normal eμt + σ²t²/2; Laplace (μ = 0, β = 1) 1/(1 − t²) for |t| < 1; Cauchy characteristic function eiμt − σ|t|, which is not differentiable at t = 0 — the signature of a missing mean.

4. Exponential, gamma, beta and Weibull: the relations

  • Memoryless: P(X > s + t | X > s) = P(X > t) characterises the exponential among continuous distributions; its hazard rate f/(1 − F) = λ is constant.
  • Sums: n independent Exp(λ) sum to Gamma(n, λ); independent gammas with the same rate add their shapes. The minimum of independent exponentials is exponential with the rates added.
  • Ratios: X ~ Gamma(a, λ), Y ~ Gamma(b, λ) independent ⇒ X/(X + Y) ~ β₁(a, b), independent of X + Y, and X/Y ~ β₂(a, b). If U ~ β₁(a, b) then U/(1 − U) ~ β₂(a, b).
  • Laplace: the difference of two independent Exp(1) variables has density ½e−|x|, variance 1 + 1 = 2.
  • Weibull: if E ~ Exp(1), θE1/k is Weibull(k, θ); k = 1 is the exponential, k > 1 an increasing hazard. The median is θ(ln 2)1/k: for k = 2, θ = 1, it is 0.833.
  • Uniform: the sum of two independent U(0, 1) variables is triangular on (0, 2) with peak at 1, not uniform; −ln U ~ Exp(1).

5. The normal and the Cauchy

If X ~ N(μ, σ²), Z = (X − μ)/σ ~ N(0, 1), and linear combinations of independent normals are normal. Z has odd moments 0, E Z² = 1, E Z⁴ = 3 (so the kurtosis is 3), and E|Z| = √(2/π) ≈ 0.798. P(|Z| ≤ 1) ≈ 0.683, P(|Z| ≤ 2) ≈ 0.954, P(|Z| ≤ 3) ≈ 0.997. For X ~ N(10, 4), P(X > 12) = P(Z > 1) = 1 − Φ(1) = 0.1587.

The Cauchy has symmetric, bell-shaped density but tails like 1/x², so E|X| = ∞: no mean, no variance. Its median and location are μ, and its quartiles are μ ± σ. It arises as the ratio of two independent standard normals and as tan(πU − π/2). The average of n independent standard Cauchy variables is again standard Cauchy — the characteristic function e−|t| raised to the n and evaluated at t/n gives e−|t| back — so the sample mean does not settle down however large n is.

🎯 Why the Cauchy is on the syllabus
It is the standing counterexample: to the law of large numbers (no mean to converge to), to the central limit theorem (no variance), to "a symmetric distribution has mean equal to its centre", and to the method of moments. When an option claims something holds "for every distribution", the Cauchy is the first thing to test it on.

Key takeaways

  • Binomial npq, Poisson mean = variance = λ, geometric (trials) 1/p and q/p², negative binomial (failures) rq/p and rq/p²; the hypergeometric carries (N − n)/(N − 1).
  • Geometric and exponential are the memoryless ones; Bin(n, λ/n) → Poisson(λ); the minimum of exponentials adds the rates.
  • Gamma(α, λ) has mean α/λ and variance α/λ²; X/(X + Y) of two same-rate gammas is β₁, X/Y is β₂.
  • β₁(a, b): mean a/(a + b), variance ab/[(a + b)²(a + b + 1)]; Laplace variance 2β²; Weibull median θ(ln 2)1/k.
  • E|Z| = √(2/π), E Z⁴ = 3; the Cauchy has no mean, its sample mean is Cauchy again, and its characteristic function e−|t| has a corner at 0.

Practice questions (15)

Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.

  1. The variance of a Bin(10, 0.3) random variable, correct to one decimal place, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 2.1

    Var = npq = 10 × 0.3 × 0.7 = 2.1. The mean np = 3 is the usual wrong answer, and np² = 0.9 drops the factor q.
  2. X ~ Poisson(2). P(X ≤ 1), correct to three decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.406

    P(0) + P(1) = e−2(1 + 2) = 3 × 0.13534 = 0.406. Stopping at P(X = 0) = 0.135 answers P(X < 1).
  3. X is the number of Bernoulli(0.2) trials needed for the first success. P(X > 5 | X > 2), correct to three decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.512

    P(X > k) = 0.8ᵏ (the first k trials fail). So P(X > 5 | X > 2) = 0.8⁵/0.8² = 0.8³ = 0.512 — memorylessness: having waited 2 trials changes nothing. Answering 0.8⁵ = 0.328 ignores the conditioning.
  4. X is the number of failures before the 3rd success in independent trials with success probability 0.5. The variance of X is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 6

    X is a sum of 3 independent geometric failure counts, each with variance q/p² = 0.5/0.25 = 2, so Var X = rq/p² = 6. The mean is rq/p = 3. Using rq, the binomial-looking formula, gives 1.5.
  5. Five items are drawn without replacement from a lot of 20 of which 8 are defective. The variance of the number of defectives drawn, correct to two decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.95

    Hypergeometric: n(M/N)(1 − M/N)(N − n)/(N − 1) = 5 × 0.4 × 0.6 × 15/19 = 1.2 × 0.7895 = 0.947 → 0.95. Omitting the finite population correction gives the binomial 1.2.
  6. X has the double exponential density f(x) = ½e−|x|, −∞ < x < ∞. The variance of X is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 2

    By symmetry E X = 0, and E X² = ∫₀^∞ x²e−x dx = Γ(3) = 2. Equivalently X is the difference of two independent Exp(1) variables, so Var = 1 + 1 = 2. Answering 1 treats |X| as if its second moment were the exponential’s variance.
  7. X follows the beta distribution of the first kind with parameters a = 2 and b = 3. The variance of X, correct to two decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.04

    ab/[(a + b)²(a + b + 1)] = 6/(25 × 6) = 0.04; the mean is 2/5 = 0.4. Dropping the factor (a + b + 1) gives 0.24, suspiciously close to 1/4 — the largest variance any variable on (0, 1) can have, and a sanity check worth running on every beta answer.
  8. Which statements are true? (All variables in each statement are independent.)

    1. If X, Y ~ Exp(1), then X − Y has the double exponential distribution
    2. If X ~ Gamma(a, λ) and Y ~ Gamma(b, λ), then X/(X + Y) ~ β₁(a, b)
    3. With X, Y as in the previous option, X/Y ~ β₂(a, b)
    4. If X, Y ~ U(0, 1), then X + Y ~ U(0, 2)
    Show answer

    Answer: A — If X, Y ~ Exp(1), then X − Y has the double exponential distribution; B — If X ~ Gamma(a, λ) and Y ~ Gamma(b, λ), then X/(X + Y) ~ β₁(a, b); C — With X, Y as in the previous option, X/Y ~ β₂(a, b)

    (A) The convolution of e−x and ex on the appropriate ranges gives ½e−|z|. (B) The Jacobian of (x, y) → (x + y, x/(x + y)) separates the joint density into a gamma in the sum and a beta in the proportion. (C) X/Y = V/(1 − V) with V ~ β₁(a, b), which is β₂(a, b). (D) False: the sum has the triangular density z on (0, 1) and 2 − z on (1, 2).
  9. Let X₁, …, Xₙ be independent standard Cauchy variables. Which statements are true?

    1. E(X₁) does not exist
    2. The characteristic function of X₁ is differentiable at t = 0
    3. The sample mean X̄ has the standard Cauchy distribution
    4. Var(X₁) = 1
    Show answer

    Answer: A — E(X₁) does not exist; C — The sample mean X̄ has the standard Cauchy distribution

    (A) ∫|x|/(π(1 + x²)) dx diverges logarithmically. (B) False: φ(t) = e−|t| has a corner at 0, and differentiability at 0 would give a finite mean. (C) φ_X̄(t) = [e−|t|/n]ⁿ = e−|t|. (D) False: no mean, so no variance; the scale 1 is the semi-interquartile range, not a standard deviation.
  10. X has the Weibull distribution with P(X ≤ x) = 1 − e−x² for x > 0. The median of X, correct to two decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.83

    1 − e−m² = 1/2 gives m² = ln 2 and m = √0.6931 = 0.8326 → 0.83. X is E1/2 for E ~ Exp(1), and the median of E is ln 2 = 0.69 — forgetting the square root answers for E, not X.
  11. X ~ N(10, 4), where 4 is the variance. Using Φ(1) = 0.8413, P(X > 12), correct to four decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.1587

    σ = 2, so z = (12 − 10)/2 = 1 and P(X > 12) = 1 − 0.8413 = 0.1587. Dividing by the variance 4 gives z = 0.5, the most common slip when the second parameter of N(μ, σ²) is read as σ.
  12. For Z ~ N(0, 1), the value of E|Z|, correct to two decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.80

    E|Z| = 2∫₀^∞ z φ(z) dz = 2 × (1/√(2π)) × ∫₀^∞ z e−z²/2 dz = 2/√(2π) = √(2/π) = 0.7979 → 0.80. E Z = 0 by symmetry is not the answer: the absolute value removes the cancellation.
  13. X ~ Bin(1000, 0.002). The Poisson approximation to P(X = 0) is:

    1. e−2
    2. e−0.002
    3. 2e−2
    4. 1 − e−2
    Show answer

    Answer: A — e^{−2}

    λ = np = 2, so P(X = 0) ≈ e−2 = 0.1353 (exactly 0.998¹⁰⁰⁰ = 0.1351). e−0.002 uses p in place of np, and 2e−2 is P(X = 1).
  14. X follows the beta distribution of the second kind β₂(m, n), with density ∝ xm−1/(1 + x)m+n on x > 0. For n > 1, E(X) equals:

    1. m/(n − 1)
    2. m/(m + n)
    3. m/n
    4. (m − 1)/(n + 1)
    Show answer

    Answer: A — m/(n − 1)

    E X = B(m + 1, n − 1)/B(m, n) = [Γ(m + 1)Γ(n − 1)]/[Γ(m)Γ(n)] = m/(n − 1), finite only for n > 1 because the tail is x−n−1. m/(m + n) is the mean of the first-kind beta; confusing the two kinds is the trap.
  15. Which continuous distribution on (0, ∞) is memoryless, i.e. P(X > s + t | X > s) = P(X > t) for all s, t > 0?

    1. Exponential
    2. Gamma with shape 2
    3. Weibull with shape 2
    4. Uniform on (0, 1)
    Show answer

    Answer: A — Exponential

    Memorylessness means the survival function satisfies S(s + t) = S(s)S(t); the only monotone solutions are S(t) = e−λt. Gamma(2) and Weibull(2) have increasing hazard (older means likelier to fail), and a uniform on (0, 1) cannot even survive past 1.