Non-parametric Statistics: The Empirical Distribution Function, Goodness of Fit, Runs, Sign and Signed-Rank Tests, Mann–Whitney, Kruskal–Wallis and Rank Correlation

Section 10 of the GATE Statistics paper: the empirical distribution function and its properties; goodness-of-fit tests — the chi-square test and the Kolmogorov–Smirnov test; run tests; the sign test and the Wilcoxon signed-rank test; the Mann–Whitney U-test and the Kruskal–Wallis test; and the rank correlation coefficients of Spearman and Kendall. What these tests share is that their null distributions do not depend on the unknown continuous distribution of the data — they are distribution-free — so the paper asks for a statistic computed from a small data set, its null mean and variance, or a p-value from a binomial count, and every one of those is worked below.

1. The empirical distribution function

For a sample X₁, …, Xₙ, Fₙ(x) = (1/n)#{i : Xᵢ ≤ x}, a step function jumping by 1/n at each observation. For fixed x, nFₙ(x) ~ Bin(n, F(x)), so Fₙ(x) is unbiased for F(x) with variance F(x)[1 − F(x)]/n, and it is consistent (strongly, by the SLLN). The Glivenko–Cantelli theorem strengthens this to uniform convergence: supₓ|Fₙ(x) − F(x)| → 0 almost surely. For the sample 3.1, 1.2, 4.5, 2.2, 5.0, F₅(3) = 2/5 = 0.4; if F(x) = 0.3 and n = 20, Var Fₙ(x) = 0.21/20 = 0.0105.

2. Goodness of fit: the chi-square and Kolmogorov–Smirnov tests

Pearson’s chi-square: with k classes, observed Oᵢ and expected Eᵢ = npᵢ under H₀, χ² = Σ(Oᵢ − Eᵢ)²/Eᵢ ≈ χ² with k − 1 − m degrees of freedom, m the number of parameters estimated from the data. Expected counts should not be too small (the usual rule is Eᵢ ≥ 5, merging classes if needed). A die rolled 60 times with counts 8, 12, 9, 11, 6, 14 gives χ² = (4 + 4 + 1 + 1 + 16 + 16)/10 = 4.2 on 5 df — no evidence against fairness. Fitting a Poisson with λ estimated to 6 classes leaves 6 − 1 − 1 = 4 df. The same statistic tests independence in an r × c table on (r − 1)(c − 1) df.

Kolmogorov–Smirnov: Dₙ = supₓ|Fₙ(x) − F₀(x)| for a fully specified continuous F₀. The supremum is attained at a data point, so compute D⁺ = maxᵢ[i/n − F₀(x₍ᵢ₎)] and D⁻ = maxᵢ[F₀(x₍ᵢ₎) − (i − 1)/n] and take D = max(D⁺, D⁻). Because F₀(X) is uniform, the null distribution of Dₙ is the same for every continuous F₀ — the test is distribution-free. The two-sample version uses sup|Fₘ − Gₙ|.

K–S computation for H₀: U(0, 1), data 0.1, 0.4, 0.45, 0.8
ix₍ᵢ₎i/n − x₍ᵢ₎x₍ᵢ₎ − (i − 1)/n
10.10.150.10
20.40.100.15
30.450.30−0.05
40.80.200.05

So D⁺ = 0.30, D⁻ = 0.15 and D₄ = 0.30, the largest vertical gap, just after the third observation.

3. Run tests

A run is a maximal block of identical symbols. In a random arrangement of n₁ symbols of one kind and n₂ of another, the number of runs R has E R = 1 + 2n₁n₂/(n₁ + n₂) and Var R = 2n₁n₂(2n₁n₂ − n₁ − n₂)/[(n₁ + n₂)²(n₁ + n₂ − 1)]. Too few runs suggest clustering or trend, too many suggest alternation. A A B B B A B A A A B B has six runs (AA | BBB | A | B | AAA | BB) with n₁ = n₂ = 6, against E R = 1 + 72/12 = 7. The Wald–Wolfowitz two-sample test pools two samples, labels each value by its sample, and counts runs in the ordered pooled sequence: few runs mean the samples differ.

4. The sign test and the Wilcoxon signed-rank test

Sign test for H₀: median = m₀ (or, for paired data, median difference 0): drop zeros; the number S of positive differences among the n remaining is Bin(n, 1/2) under H₀. With 10 positive out of 12, the one-sided p-value is P(S ≥ 10) = (66 + 12 + 1)/4096 = 0.0193. It uses only signs, so it needs almost no assumptions — and wastes the magnitudes.

Wilcoxon signed-rank (assumes a symmetric distribution): rank |dᵢ| from 1 to n and let W⁺ be the sum of the ranks of the positive dᵢ. W⁺ + W⁻ = n(n + 1)/2, and under H₀ E W⁺ = n(n + 1)/4, Var W⁺ = n(n + 1)(2n + 1)/24. For d = 1.5, −0.5, 2.0, 3.1, −1.0, 2.6 the ranks of |d| are 3, 1, 4, 6, 2, 5, so W⁺ = 3 + 4 + 6 + 5 = 18 and W⁻ = 3, against a null mean of 10.5.

5. The Mann–Whitney U-test and the Kruskal–Wallis test

Mann–Whitney: for samples X (size n₁) and Y (size n₂), U_X = #{(i, j) : Xᵢ > Yⱼ}. Equivalently, if R_X is the sum of the X ranks in the pooled sample, U_X = R_X − n₁(n₁ + 1)/2, and U_X + U_Y = n₁n₂. Under H₀ (same distribution) E U = n₁n₂/2, Var U = n₁n₂(n₁ + n₂ + 1)/12. X = 3, 7, 9 and Y = 1, 4, 5, 8 give pooled ranks 2, 5, 7 for X, R_X = 14 and U_X = 14 − 6 = 8, against a null mean of 6 — directly, 3 beats one Y, 7 beats three and 9 beats four.

Kruskal–Wallis extends this to k samples: rank all N observations together, let Rᵢ be the rank sum of sample i (size nᵢ), and compute H = 12/[N(N + 1)] Σ Rᵢ²/nᵢ − 3(N + 1), approximately χ²k−1 under H₀. For three groups of three with rank sums 6, 15 and 24 (complete separation), H = (12/90)(12 + 75 + 192) − 30 = 37.2 − 30 = 7.2, the largest value possible with these sizes. For k = 2, H is the square of the standardised Mann–Whitney statistic.

6. Rank correlation: Spearman and Kendall

Spearman’s ρ_s is Pearson’s correlation of the ranks; without ties, ρ_s = 1 − 6Σdᵢ²/[n(n² − 1)], dᵢ the rank differences. Kendall’s τ counts pairs: τ = (C − D)/[n(n − 1)/2], with C concordant and D discordant pairs. Both lie in [−1, 1], equal ±1 for perfect monotone agreement or disagreement, and are unchanged by any increasing transformation of either variable. For x-ranks 1, 2, 3, 4, 5 and y-ranks 2, 1, 4, 3, 5: d = −1, 1, −1, 1, 0, Σd² = 4, ρ_s = 1 − 24/120 = 0.8; the discordant pairs are (1, 2) and (3, 4), so C = 8, D = 2 and τ = 0.6. They measure the same thing on different scales and are not equal in general.

🧠 Which test for which question
One sample or paired, location: sign test (signs only) or signed-rank (signs and magnitudes, symmetry assumed). Two independent samples: Mann–Whitney. k samples: Kruskal–Wallis. Whole distribution against a specified F₀: K–S (continuous, fully specified) or chi-square (grouped, parameters may be estimated). Randomness of a sequence: runs.

Key takeaways

  • nFₙ(x) ~ Bin(n, F(x)): Fₙ(x) is unbiased with variance F(1 − F)/n; Glivenko–Cantelli gives uniform a.s. convergence.
  • Chi-square GOF: Σ(O − E)²/E on k − 1 − m df. K–S: D = max over data points of i/n − F₀(x₍ᵢ₎) and F₀(x₍ᵢ₎) − (i − 1)/n; distribution-free.
  • Runs: E R = 1 + 2n₁n₂/(n₁ + n₂). Sign test: S ~ Bin(n, 1/2) after dropping zeros.
  • Signed rank: E W⁺ = n(n + 1)/4, Var = n(n + 1)(2n + 1)/24. Mann–Whitney: U = R − n₁(n₁ + 1)/2, E U = n₁n₂/2, Var = n₁n₂(n₁ + n₂ + 1)/12.
  • Kruskal–Wallis H = 12/[N(N + 1)]ΣRᵢ²/nᵢ − 3(N + 1) ≈ χ²k−1. Spearman 1 − 6Σd²/[n(n² − 1)]; Kendall (C − D)/C(n, 2).

Practice questions (16)

Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.

  1. A sample is 3.1, 1.2, 4.5, 2.2, 5.0. The value of the empirical distribution function F₅(3), correct to one decimal place, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.4

    Two observations, 1.2 and 2.2, are ≤ 3, so F₅(3) = 2/5 = 0.4. Counting 3.1 as well (it is the closest value) gives 0.6; Fₙ counts values at or below x, not near it.
  2. For a sample of size 20 from a distribution with F(x₀) = 0.3, the variance of the empirical distribution function F₂₀(x₀), correct to four decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.0105

    20F₂₀(x₀) ~ Bin(20, 0.3), so Var F₂₀(x₀) = 0.3 × 0.7/20 = 0.0105. Giving the binomial variance 20 × 0.21 = 4.2 answers for the count, not the proportion.
  3. A sample of four, 0.1, 0.4, 0.45, 0.8, is tested against H₀: U(0, 1). The Kolmogorov–Smirnov statistic D₄, correct to one decimal place, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.3

    D⁺ = max(0.25 − 0.1, 0.5 − 0.4, 0.75 − 0.45, 1 − 0.8) = 0.30 and D⁻ = max(0.1, 0.15, −0.05, 0.05) = 0.15, so D = 0.3. Checking only the gaps F₀(x₍ᵢ₎) − (i − 1)/n, or only at the left end of each step, misses the largest gap, just after 0.45.
  4. A die is rolled 60 times with face counts 8, 12, 9, 11, 6, 14. The chi-square goodness-of-fit statistic for fairness, correct to one decimal place, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 4.2

    Each Eᵢ = 10: Σ(O − E)²/E = (4 + 4 + 1 + 1 + 16 + 16)/10 = 4.2, on 5 df — well below the 5% point 11.07. Dividing by the observed counts instead of the expected (Neyman’s variant) gives a different number and is not Pearson’s statistic.
  5. A Poisson distribution, with λ estimated from the data, is fitted to counts grouped into 6 classes. The degrees of freedom of the chi-square goodness-of-fit test are ____.

    Numerical answer — type the value.

    Show answer

    Answer: 4

    df = k − 1 − m = 6 − 1 − 1 = 4: one for the constraint that the counts sum to n, one for the estimated λ. Answering 5 forgets that estimating λ uses up a degree of freedom.
  6. The sequence A A B B B A B A A A B B contains six A’s and six B’s. Under randomness, the expected number of runs is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 7

    E R = 1 + 2n₁n₂/(n₁ + n₂) = 1 + 72/12 = 7. The observed number is 6 (AA | BBB | A | B | AAA | BB), close to 7, so there is no evidence against randomness. Answering 6 reports the observed count, not the expectation.
  7. In a paired experiment there are 12 non-zero differences, 10 of them positive. For the sign test of H₀: median difference = 0 against a positive shift, the p-value, correct to four decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.0193

    S ~ Bin(12, 1/2): P(S ≥ 10) = [C(12, 10) + C(12, 11) + C(12, 12)]/2¹² = (66 + 12 + 1)/4096 = 79/4096 = 0.0193. P(S = 10) alone, 66/4096 = 0.0161, is not a p-value: it must include the more extreme outcomes 11 and 12.
  8. Paired differences are 1.5, −0.5, 2.0, 3.1, −1.0, 2.6. The Wilcoxon signed-rank statistic W⁺ (the sum of the ranks of the positive differences) is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 18

    Ranks of |d|: 0.5 → 1, 1.0 → 2, 1.5 → 3, 2.0 → 4, 2.6 → 5, 3.1 → 6. W⁺ = 3 + 4 + 6 + 5 = 18 and W⁻ = 1 + 2 = 3; check W⁺ + W⁻ = 21 = 6 × 7/2. Ranking the signed values instead of their absolute values gives the negatives the lowest ranks by construction.
  9. Samples X: 3, 7, 9 and Y: 1, 4, 5, 8 are compared by the Mann–Whitney test. The statistic U_X = #{(i, j) : Xᵢ > Yⱼ} is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 8

    3 exceeds one Y (1), 7 exceeds three (1, 4, 5), 9 exceeds all four: U_X = 8. Via ranks: X has pooled ranks 2, 5, 7, R_X = 14, U_X = 14 − 3 × 4/2 = 8. Reporting R_X = 14 forgets to subtract n₁(n₁ + 1)/2; U_Y = 12 − 8 = 4 is the other count.
  10. Three groups of three observations are ranked together; the rank sums are 6, 15 and 24. The Kruskal–Wallis statistic H, correct to one decimal place, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 7.2

    N = 9: H = 12/(9 × 10) × (6²/3 + 15²/3 + 24²/3) − 3 × 10 = (12/90)(12 + 75 + 192) − 30 = 37.2 − 30 = 7.2, referred to χ²₂. Forgetting to divide each Rᵢ² by nᵢ gives an H three times as large before the subtraction.
  11. Two judges rank five items: judge A gives 1, 2, 3, 4, 5 and judge B gives 2, 1, 4, 3, 5. Spearman’s rank correlation coefficient, correct to one decimal place, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.8

    d = −1, 1, −1, 1, 0, Σd² = 4, ρ_s = 1 − 6 × 4/(5 × 24) = 1 − 0.2 = 0.8. Using n² in place of n² − 1 gives 1 − 24/125 = 0.81.
  12. For the same rankings (A: 1, 2, 3, 4, 5; B: 2, 1, 4, 3, 5), Kendall’s τ, correct to one decimal place, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.6

    Of the C(5, 2) = 10 pairs, only (1, 2) and (3, 4) are ordered differently by B, so C = 8, D = 2 and τ = (8 − 2)/10 = 0.6. It is not equal to Spearman’s 0.8 — the two coefficients use different scales. Dividing by 20 (ordered pairs) halves the answer.
  13. Which statements are true?

    1. The Kolmogorov–Smirnov test is distribution-free under H₀ for a fully specified continuous F₀
    2. nFₙ(x) has the Bin(n, F(x)) distribution for each fixed x
    3. The sign test uses the magnitudes of the differences
    4. The Kruskal–Wallis test extends the Mann–Whitney test to more than two samples
    Show answer

    Answer: A — The Kolmogorov–Smirnov test is distribution-free under H₀ for a fully specified continuous F₀; B — nFₙ(x) has the Bin(n, F(x)) distribution for each fixed x; D — The Kruskal–Wallis test extends the Mann–Whitney test to more than two samples

    (A) F₀(Xᵢ) are i.i.d. U(0, 1) under H₀, and Dₙ is unchanged by that transformation. (B) Each 1{Xᵢ ≤ x} is Bernoulli(F(x)). (C) False: it uses only signs; the signed-rank test is the one that adds magnitudes. (D) For k = 2, H is the square of the standardised rank-sum statistic.
  14. Which statements about rank statistics (no ties) are true?

    1. U_X + U_Y = n₁n₂ for the Mann–Whitney counts
    2. W⁺ + W⁻ = n(n + 1)/2 for the signed-rank sums
    3. Under H₀, E(U_X) = n₁n₂
    4. Spearman’s ρ_s and Kendall’s τ are always equal
    Show answer

    Answer: A — U_X + U_Y = n₁n₂ for the Mann–Whitney counts; B — W⁺ + W⁻ = n(n + 1)/2 for the signed-rank sums

    (A) Every one of the n₁n₂ pairs has exactly one of Xᵢ > Yⱼ, Yⱼ > Xᵢ. (B) Together they sum all the ranks 1 to n. (C) False: by symmetry under H₀, U_X and U_Y have the same distribution, so E U_X = n₁n₂/2. (D) False: for the rankings (1, 2, 3, 4, 5) and (2, 1, 4, 3, 5), ρ_s = 0.8 but τ = 0.6.
  15. Two independent samples are to be compared for a location shift without assuming normality. The appropriate test among these is:

    1. the Mann–Whitney U-test
    2. the Wilcoxon signed-rank test
    3. the runs test for randomness of one sequence
    4. Spearman’s rank correlation test
    Show answer

    Answer: A — the Mann–Whitney U-test

    Mann–Whitney ranks the pooled samples and asks whether one sample’s ranks are systematically higher. The signed-rank test is for one sample or paired differences — independent samples have no pairing; the runs test checks the order of a single sequence, and rank correlation measures association between paired variables, not a shift.
  16. Compared with the sign test, the Wilcoxon signed-rank test for a median additionally assumes that:

    1. the distribution is symmetric about its median
    2. the distribution is normal
    3. the variance is known
    4. the sample size is at least 30
    Show answer

    Answer: A — the distribution is symmetric about its median

    Under H₀ the signs must be independent of the magnitudes, with each sign equally likely at every |d| — which is symmetry about the hypothesised median. Normality and a known variance are what the t- and z-tests need; neither test here needs a large sample, since exact null distributions are tabulated.