Basic Image Classification and Segmentation: Supervised and Unsupervised Techniques, Accuracy Analysis

This is a section of Part B2, Image Processing and Analysis, of the Geomatics Engineering (GE) paper. B2's last heading is “Basic Image Classification and Segmentation: Supervised and unsupervised techniques, accuracy analysis”. This chapter sets out feature space, spectral against information classes and the classification workflow; clusters pixels without training data by K-means and ISODATA; classifies from training areas by the parallelepiped, minimum-distance and maximum-likelihood rules, with the training-sample requirement; scores a map with the error matrix, overall, user's and producer's accuracy and kappa; and ends with segmentation — thresholding and Otsu, edges, region growing, split-and-merge and watershed — and the object-based approach built on it. The numericals are K-means updates, distances, error-matrix accuracies, kappa and Otsu's between-class variance.

1. Classification: classes, feature space and the workflow

Image classification turns a multispectral image into a thematic map by giving every pixel a class label. Each pixel is a point in feature space — a k-band pixel is a point in k dimensions — and pixels of one cover type cluster together. The spectral classes a computer can find are groups in feature space; the information classes a user wants (forest, crop, water, built-up) are not always the same: one information class may split into several spectral classes (sunlit and shaded forest), and two information classes may share one (a dry crop and bare soil). Classification is hard when each pixel takes exactly one class and soft (fuzzy) when it takes a membership in each, which suits mixed pixels that straddle two covers.

Supervised against unsupervised
AspectSupervisedUnsupervised
Starts fromclasses the analyst defines and samplesthe data alone; the computer groups pixels
Training datarequired — known areas for every classnot required
Labels are attachedbefore allocationafter clustering, by the analyst
Examplesparallelepiped, minimum distance, maximum likelihoodK-means, ISODATA
Suitswell-known area with defined classesunfamiliar area; exploring the spectral classes
  • Workflow (supervised): define classes → choose bands → collect training areas → compute class statistics → allocate every pixel → smooth (for example a 3 × 3 majority filter) → assess accuracy with independent reference data.
  • Pixel-based against object-based: a pixel classifier looks at one spectrum at a time and, on fine-resolution images, gives a speckled “salt-and-pepper” map; an object-based one segments the image first and classifies the objects (see section 5).

2. Unsupervised classification: K-means and ISODATA

K-means clusters pixels into k groups. (1) Choose k and initial cluster means (seeds). (2) Assign every pixel to the nearest mean, by Euclidean distance in feature space. (3) Recompute each mean as the average of its members. (4) Repeat 2–3 until the means stop moving or fewer than a set fraction of pixels change cluster. Worked, in one band: pixels 2, 4, 8, 10, 12 with seeds 4 and 10. Distances put 2 and 4 with the first seed and 8, 10 and 12 with the second (8 is 4 from the first and 2 from the second). New means: (2 + 4)/2 = 3 and (8 + 10 + 12)/3 = 10. A second pass changes no assignment, so the process has converged. The within-cluster sum of squares is (2 − 3)² + (4 − 3)² + (8 − 10)² + (10 − 10)² + (12 − 10)² = 2 + 8 = 10, the quantity K-means reduces at every step.

ISODATA (Iterative Self-Organising Data Analysis) is K-means with the number of clusters allowed to change. After each iteration it deletes clusters with too few members, splits a cluster whose standard deviation in some band exceeds a limit (or that has many members), and merges clusters whose means are closer than a limit. The analyst therefore sets an approximate number of clusters, a minimum cluster size, a maximum standard deviation, a minimum distance between means, a maximum number of merges and a convergence threshold, often 95% of pixels unchanged.

⚠️ Clusters are spectral, not thematic
Nothing in K-means says what a cluster is. Cluster 7 becomes “water” only when the analyst compares it with reference data, and two clusters may need to be recoded as one class. Results also depend on the seeds, and Euclidean distance favours round, equal-sized clusters.

3. Supervised classifiers

The analyst first digitises training areas for every class on the image, from field visits or maps, and the software computes each class's mean vector μ and covariance matrix Σ. Three classical rules then allocate every pixel. The parallelepiped (box) rule gives each class a range, often μ ± kσ, in every band and allocates a pixel that falls in a class's box in all bands; it is fast, but boxes overlap (a pixel falls in two) and leave gaps (a pixel falls in none and stays unclassified). The minimum-distance-to-mean rule allocates a pixel to the class whose mean is nearest in Euclidean distance; every pixel is classified, the boundaries are straight perpendicular bisectors, and it ignores the spread of the classes. The maximum-likelihood rule assumes each class is a multivariate Gaussian cloud and allocates the pixel to the class with the highest discriminant g_i(x) = ln P(ωᵢ) − ½ ln|Σᵢ| − ½(x − μᵢ)ᵀΣᵢ⁻¹(x − μᵢ); it uses both spread and correlation and gives curved (quadratic) boundaries, so it is the most accurate of the three when the classes are close to Gaussian and the training data adequate.

Worked minimum distance, bands (red, NIR): water mean (20, 10), vegetation (40, 90), soil (80, 60); the pixel (35, 70). Distances: to water √(15² + 60²) = 61.8, to vegetation √(5² + 20²) = √425 = 20.6, to soil √(45² + 10²) = 46.1, so it is vegetation. Worked maximum likelihood in one band with equal priors, class A mean 40 and σ 5, class B mean 60 and σ 15, pixel 50: it is 10 from both means, so minimum distance cannot decide, but ln-likelihood is −ln σ − (x − μ)²/2σ², which is −ln 5 − 2 = −3.61 for A and −ln 15 − 0.22 = −2.93 for B, so the pixel goes to the broader class B.

  • Training data needed: the covariance matrix of a class in k bands can be inverted only if the class has at least k + 1 training pixels, and in practice 10k to 100k pixels per class are advised, well spread over the class.
  • Other classifiers in current use: support vector machines, decision trees and random forests, k-nearest neighbours and neural networks. They assume no Gaussian shape and cope better with many bands and mixed data, but they are still supervised and need training samples.
  • Prior probabilities P(ωᵢ) shift the boundary towards the less likely class; the threshold option leaves a pixel unclassified when its likelihood is below a chosen value.
🎯 Why maximum likelihood beats minimum distance
Minimum distance treats every class as a circle of equal size. Real classes are elongated ellipses — wheat varies more than deep water — and the covariance term lets the rule respect that, at the cost of needing more training pixels and more computation.

4. Accuracy assessment

A classified map is tested against reference data from field visits or higher-resolution imagery, at sample locations chosen by (stratified) random sampling and independent of the training pixels — testing on the training pixels only measures how well the classifier remembers them. The counts go into the error (confusion) matrix, with the classified map in the rows and the reference in the columns, correct pixels on the diagonal. Overall accuracy = Σ diagonal/N. User's accuracy of a class = correct/its row total (1 − commission error): how often the map's label is right. Producer's accuracy = correct/its column total (1 − omission error): how much of the real class was found. Kappa κ = (p_o − p_e)/(1 − p_e), where p_o is the overall accuracy and p_e = Σ(row totalᵢ × column totalᵢ)/N² is the agreement expected by chance.

Error matrix (rows: classified; columns: reference)
Classified asForestCropWaterRow totalUser's
Forest70558070/80 = 87.5%
Crop46067060/70 = 85.7%
Water15445044/50 = 88%
Column total757055200—
Producer's70/75 = 93.3%60/70 = 85.7%44/55 = 80%——

From this matrix: overall accuracy = (70 + 60 + 44)/200 = 174/200 = 87%. Chance agreement p_e = (80 × 75 + 70 × 70 + 50 × 55)/200² = 13650/40000 = 0.341, so κ = (0.87 − 0.341)/(1 − 0.341) = 0.80. Water has a producer's accuracy of only 80% because 11 of the 55 real water pixels were labelled something else (omission), although a pixel labelled water is right 88% of the time. A simpler two-class case: 45 and 40 correct on the diagonal, 5 and 10 off it, gives overall accuracy 0.85, p_e = 0.5 and κ = 0.35/0.5 = 0.7.

⚠️ User's and producer's are easy to swap
Read the row for the user's accuracy (the map user asks: is the label right?) and the column for the producer's (the map maker asks: did I find it all?). Kappa is never larger than overall accuracy, is 1 for a perfect map, 0 for chance-level agreement, and can be negative.

5. Image segmentation

Segmentation divides an image into regions that are homogeneous in some property, so that later steps work on objects rather than single pixels. Thresholding splits on DN: if the histogram is bimodal, the valley between the modes is the threshold; Otsu's method picks the threshold that maximises the between-class variance σ_B² = ω₀ω₁(μ₀ − μ₁)². Worked: DNs 10, 12, 14, 16, 100, 102, 104, 106 split at the valley give ω₀ = ω₁ = 0.5, μ₀ = 13 and μ₁ = 103, so σ_B² = 0.25 × 90² = 2025. Edge-based segmentation finds boundaries with a gradient filter (Sobel) and links them into closed outlines. Region-based methods work from similarity: region growing starts from seed pixels and adds neighbours whose DN is within a tolerance of the region; split-and-merge divides the image into quadrants until each is homogeneous, then merges similar neighbours; watershed treats the gradient image as a relief and floods it from its minima.

Object-based image analysis (OBIA) joins segmentation to classification. A multi-resolution segmentation, controlled by a scale parameter (larger scale, larger and fewer objects) and by colour, shape and compactness weights, builds the objects; each object is then classified using not only its mean spectrum but its shape, texture, size and neighbours. On high-resolution images it gives cleaner maps than pixel-by-pixel classification, because a house or a field is one object rather than a scatter of unlike pixels, and it can use rules such as “a bright object next to a road and 10–20 m long is a building”. Its results depend on the scale chosen.

ℹ️ Segment first, or classify first?
Coarse pixels (tens of metres) mix covers, and per-pixel classification is usual. When pixels are much smaller than the objects (sub-metre imagery) the object is the meaningful unit and segmentation first is usual.

Key takeaways

  • Supervised classification needs training areas defined before allocation; unsupervised clustering (K-means, ISODATA) needs none and is labelled afterwards; ISODATA adds splitting, merging and deleting of clusters.
  • K-means: assign to the nearest mean, recompute means, repeat until stable; it minimises the within-cluster sum of squares and depends on the seeds.
  • Parallelepiped leaves overlaps and unclassified gaps; minimum distance ignores class spread; maximum likelihood uses μ, Σ and priors, needs at least k + 1 (advisably 10k–100k) pixels per class.
  • Error matrix: overall = Σ diagonal/N; user's = correct/row total (commission); producer's = correct/column total (omission); κ = (p_o − p_e)/(1 − p_e) ≤ overall accuracy; test on data independent of training.
  • Segmentation: thresholding (Otsu maximises ω₀ω₁(μ₀ − μ₁)²), edge, region growing, split-and-merge, watershed; OBIA classifies objects using spectrum, shape, texture and context, at a chosen scale.

Practice questions (16)

Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.

  1. Pixels of DN 2, 4, 8, 10 and 12 are clustered by K-means with initial means 4 and 10. After the first assignment and update, the mean of the second cluster is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 10

    2 and 4 are nearer to 4; 8 (2 from 10, 4 from 4), 10 and 12 are nearer to 10. The second mean is (8 + 10 + 12)/3 = 10, and the first is (2 + 4)/2 = 3. A second pass changes nothing, so the clustering has converged.
  2. For the converged clusters {2, 4} and {8, 10, 12} the total within-cluster sum of squared deviations from the cluster means is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 10

    Cluster {2, 4}: mean 3, (−1)² + 1² = 2. Cluster {8, 10, 12}: mean 10, (−2)² + 0 + 2² = 8. Total 2 + 8 = 10. K-means never increases this quantity from one iteration to the next.
  3. Compared with plain K-means, ISODATA additionally

    1. needs training areas for every class
    2. uses the Gaussian likelihood of each class
    3. splits, merges and deletes clusters, so the number of clusters can change
    4. always gives exactly the requested number of clusters
    Show answer

    Answer: C — splits, merges and deletes clusters, so the number of clusters can change

    ISODATA adds split, merge and delete rules with analyst-set thresholds, so the final number of clusters differs from the initial one. Like K-means it is unsupervised and distance-based, and it is exactly K-means that fixes k.
  4. Class means in (red, NIR) are water (20, 10), vegetation (40, 90) and soil (80, 60). The Euclidean distance from the pixel (35, 70) to the vegetation mean, to one decimal place, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 20.6

    √((35 − 40)² + (70 − 90)²) = √(25 + 400) = √425 = 20.6. To water it is 61.8 and to soil 46.1, so a minimum-distance classifier labels the pixel vegetation.
  5. A maximum-likelihood classifier is to be trained for a 6-band image. The least number of training pixels per class for which the class covariance matrix can be inverted is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 7

    With n samples the k × k sample covariance matrix has rank at most n − 1, so it is invertible only if n − 1 ≥ k, i.e. n ≥ k + 1 = 7. In practice many more (10k to 100k) are needed for stable estimates.
  6. In one band, class A has mean 40 and σ = 5 and class B has mean 60 and σ = 15; the priors are equal. A pixel of DN 50 is classified by minimum distance and by maximum likelihood. The outcome is

    1. minimum distance cannot separate them (a tie), while maximum likelihood gives class B
    2. minimum distance gives A and maximum likelihood also gives A
    3. both give class A
    4. both give class B
    Show answer

    Answer: A — minimum distance cannot separate them (a tie), while maximum likelihood gives class B

    DN 50 is 10 from both means, so minimum distance ties. The log-likelihood −ln σ − (x − μ)²/2σ² is −ln 5 − 2 = −3.61 for A and −ln 15 − 0.22 = −2.93 for B; B is larger, because the broad class B spreads its probability thinly but loses little at 10 from its mean, whereas the narrow class A is penalised heavily.
  7. A parallelepiped classifier leaves some pixels with no label because

    1. its covariance matrix cannot be inverted
    2. a pixel can fall outside every class's box in at least one band
    3. it always allocates to the nearest mean, which may be undefined
    4. it uses prior probabilities that sum to more than 1
    Show answer

    Answer: B — a pixel can fall outside every class's box in at least one band

    Each class occupies a box (often μ ± kσ in every band); a pixel outside all boxes matches no class and stays unclassified, and one inside two boxes is ambiguous. Minimum distance and maximum likelihood, by contrast, always assign a class unless a threshold is set.
  8. In the error matrix (classified rows, reference columns) Forest 70/5/5, Crop 4/60/6, Water 1/5/44, with 200 test pixels in all, the overall accuracy in percent is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 87

    The diagonal holds 70 + 60 + 44 = 174 correct pixels; 174/200 = 0.87, i.e. 87%. Overall accuracy hides differences between classes, which the user's and producer's accuracies reveal.
  9. For the same matrix, the producer's accuracy of Water (the water column holds 5 + 6 + 44 = 55 reference pixels, of which 44 were classified as water), in percent, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 80

    Producer's accuracy = correct/column total = 44/55 = 0.80, i.e. 80%; the other 11 real water pixels were omitted (5 labelled forest and 6 labelled crop). Its complement, 20%, is the omission error.
  10. From the same matrix, the user's accuracy of Crop (60 correct in a row total of 70), in percent to one decimal place, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 85.7

    User's accuracy = correct/row total = 60/70 = 0.857, i.e. 85.7%. Of the pixels the map calls crop, 14.3% (10 of 70) are commission errors: really forest (4) or water (6).
  11. A two-class map is tested with 100 pixels: rows W 45/5 and L 10/40 (columns reference W, L). Overall accuracy is 0.85 and chance agreement is 0.5. The kappa coefficient is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.7

    κ = (p_o − p_e)/(1 − p_e) = (0.85 − 0.5)/(1 − 0.5) = 0.35/0.5 = 0.7. Check p_e: rows 50 and 50, columns 55 and 45, (50 × 55 + 50 × 45)/100² = 0.5. Kappa is below the overall accuracy of 0.85 because it removes chance agreement.
  12. Which of the following statements about accuracy assessment are correct?

    1. User's accuracy is the complement of commission error
    2. Producer's accuracy is the complement of omission error
    3. Kappa can never exceed the overall accuracy
    4. The reference samples should be the same pixels used to train the classifier
    Show answer

    Answer: A — User's accuracy is the complement of commission error; B — Producer's accuracy is the complement of omission error; C — Kappa can never exceed the overall accuracy

    User's accuracy = 1 − commission error (row-based) and producer's = 1 − omission error (column-based); κ = (p_o − p_e)/(1 − p_e) ≤ p_o. Testing on the training pixels is invalid because it measures memory rather than performance — the reference sample must be independent.
  13. Which of the following are supervised classification methods that need training samples?

    1. Maximum likelihood
    2. Parallelepiped
    3. Minimum distance to mean
    4. ISODATA
    Show answer

    Answer: A — Maximum likelihood; B — Parallelepiped; C — Minimum distance to mean

    Maximum likelihood, minimum distance and parallelepiped all take their class statistics from analyst-defined training areas. ISODATA clusters the data on its own and the analyst labels the clusters afterwards, so it is unsupervised.
  14. Eight pixels of DN 10, 12, 14, 16, 100, 102, 104 and 106 are split by a threshold in the valley between the two groups. The between-class variance ω₀ω₁(μ₀ − μ₁)² is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 2025

    Each group has weight 0.5. The means are (10 + 12 + 14 + 16)/4 = 13 and (100 + 102 + 104 + 106)/4 = 103. σ_B² = 0.5 × 0.5 × (13 − 103)² = 0.25 × 8100 = 2025, the maximum Otsu could find for this data, since no other split does better.
  15. Starting from a seed pixel and adding neighbouring pixels whose DN is within a tolerance of the region's mean is called

    1. region growing
    2. K-means clustering
    3. Otsu thresholding
    4. edge linking
    Show answer

    Answer: A — region growing

    Region growing is a region-based segmentation that expands from seeds by a similarity criterion. Otsu works on the global histogram, K-means clusters in feature space without spatial contiguity, and edge linking joins boundary pixels.
  16. Which statements about object-based image analysis are correct?

    1. It removes the need to choose any parameter
    2. It can use shape, texture and neighbour relations as well as mean spectrum
    3. It segments the image into objects before classifying them
    4. A larger scale parameter gives larger and fewer objects
    Show answer

    Answer: B — It can use shape, texture and neighbour relations as well as mean spectrum; C — It segments the image into objects before classifying them; D — A larger scale parameter gives larger and fewer objects

    OBIA segments first and classifies objects using spectral, shape, texture and context features; the scale parameter controls object size (larger scale, larger objects). It does not remove parameter choices — scale, colour and shape weights are set by the analyst and change the result.