Image Transformation: Principal Component Analysis, Discriminant Analysis, Colour Transformations and Indices
1. Band ratios and indices: NDVI and NDWI
Image transformation makes new bands from old ones. The simplest is the band ratio: divide the DN or reflectance of one band by another. Illumination multiplies every band of a pixel by about the same factor — a sunlit slope is brighter than a shaded slope of the same crop — so the ratio cancels it: NIR 80 and red 20 on the sunlit side, NIR 40 and red 10 in shadow, give the ratio 4 both times. What is left is the spectral shape, which is why ratios separate cover types and suppress topographic shading. Their limits are that an additive haze offset does not cancel, and that different targets can share one ratio.
NDVI (normalised difference vegetation index) = (NIR − Red)/(NIR + Red). Healthy leaves absorb red for photosynthesis and reflect near-infrared from their internal structure, so vegetation has high NIR and low red and a high NDVI; the value always lies between −1 and +1. Roughly: dense green vegetation 0.6–0.9, sparse or stressed vegetation 0.2–0.5, bare soil about 0.1–0.2, and water, snow and cloud near zero or negative. From reflectances NIR 0.45 and red 0.05, NDVI = 0.40/0.50 = 0.8. From DNs 120 and 40 it is 80/160 = 0.5. Normalising by the sum makes it insensitive to a common illumination factor, and it is used for crop condition, biomass and drought monitoring; it saturates over very dense canopy and is disturbed by bright soil under sparse vegetation, which the soil-adjusted SAVI = (1 + L)(NIR − Red)/(NIR + Red + L), with L = 0.5, reduces: NIR 0.5, red 0.1 gives 1.5 × 0.4/1.1 = 0.55.
NDWI is a family of water indices. The one that maps open water (McFeeters) is NDWI = (Green − NIR)/(Green + NIR): water reflects more green than NIR, which it absorbs almost entirely, so it is positive over water and negative over vegetation and soil. For green 0.10 and NIR 0.05, NDWI = 0.05/0.15 = 0.33. The one that tracks the water content of vegetation (Gao) is (NIR − SWIR)/(NIR + SWIR), since short-wave infrared is absorbed by leaf water. The MNDWI, (Green − SWIR)/(Green + SWIR), separates water from built-up land better than the green–NIR form. The same construction gives the built-up index NDBI = (SWIR − NIR)/(SWIR + NIR).
2. Principal component analysis
Neighbouring bands are strongly correlated, so an n-band image carries far fewer than n independent pieces of information. Principal component analysis (PCA) rotates the axes of the band space to those of the data cloud. The steps: (1) compute the covariance matrix C of the bands; (2) find its eigenvalues λ from det(C − λI) = 0 and its eigenvectors; (3) rank the eigenvectors by eigenvalue; (4) project each pixel, y = Aᵀ(x − μ). The result is n principal components (PCs) that are uncorrelated — their covariance matrix is diagonal — with PC1 of maximum variance, PC2 the most of what remains at right angles to it, and so on. The eigenvalue λₖ is the variance of PCk, and Σλ = trace(C) = the total variance of the original bands, so nothing is created or lost; PCk carries λₖ/Σλ of it.
Worked. Two bands with variances 5 and 5 and covariance 3 have C = [5 3; 3 5], trace 10 and determinant 25 − 9 = 16, so λ² − 10λ + 16 = 0 and λ = 8 and 2. The eigenvectors are (1, 1)/√2 for λ = 8 and (1, −1)/√2 for λ = 2. PC1 holds 8/10 = 80% of the variance. A pixel (10, 6) with band means (8, 4) has deviations (2, 2), so PC1 = (2 + 2)/√2 = 2.83 and PC2 = (2 − 2)/√2 = 0. For a correlation matrix [1 r; r 1] the eigenvalues are 1 ± r: with r = 0.9, 1.9 and 0.1, so PC1 holds 95% of the total.
- Compression: keep the first few PCs; six Landsat reflective bands often need two or three PCs to hold over 95% of the variance.
- Noise removal: the last PCs hold little variance and much noise; dropping them and inverting the transform cleans the image.
- Change detection: stack two dates and take PCs; PC1 holds what is common and the higher PCs the change.
- Colour display: PC1, PC2 and PC3 as red, green and blue give a composite with little redundancy.
3. Discriminant analysis
PCA is unsupervised: it looks for maximum total variance and never asks which pixel belongs to which class. Discriminant analysis is supervised: from training pixels of known classes it finds the linear combinations of bands that separate the classes best. Fisher's criterion maximises the ratio of between-class to within-class variance; for two classes in one dimension J = (μ₁ − μ₂)²/(σ₁² + σ₂²), so means 40 and 60 with variances 10 and 10 give J = 400/20 = 20. For two classes in several bands the best direction is w ∝ S_w⁻¹(μ₁ − μ₂), with S_w the pooled within-class covariance. With c classes and k bands there are at most min(k, c − 1) discriminant axes: 6 bands and 4 classes give 3.
| Aspect | PCA | Discriminant analysis |
|---|---|---|
| Uses class labels | no — unsupervised | yes — needs training samples |
| Objective | maximum total variance | maximum class separation (between/within) |
| Number of axes | up to the number of bands | at most min(bands, classes − 1) |
| Typical use | compression, noise, display | feature reduction before classification; the classifier itself |
As a classifier, a pixel x is given to the class with the largest discriminant function. If every class has the same covariance and the same prior, the boundary between two classes is the perpendicular bisector of the segment joining their means: for means 40 and 60 with a common σ, the threshold is 50, and DN 47 goes to the class of mean 40. A larger prior for one class moves the boundary towards the other class's mean; unequal covariances curve the boundary into a quadratic surface, which is the Gaussian maximum-likelihood classifier of the classification chapter. The separation of the classes can be summarised by the Mahalanobis distance D = |μ₁ − μ₂|/σ; for σ = 10 that is 20/10 = 2.
4. Colour transformations: RGB, IHS and CMYK
A display makes colour by adding red, green and blue light: the RGB additive model, in which R + G = yellow, G + B = cyan, R + B = magenta and all three at full strength give white. Three 8-bit channels give 256³ = 16,777,216 colours (24 bits); three 4-bit channels give 16³ = 4096. A three-band image is shown by assigning bands to the three guns; the standard false colour composite puts NIR in red, red in green and green in blue, so healthy vegetation looks red. CMYK is the subtractive model of printing: cyan, magenta and yellow inks absorb red, green and blue respectively, and black (K) is added because the three do not make a true black. With R, G and B scaled to 0–1: C = 1 − R, M = 1 − G, Y = 1 − B, and K = min(C, M, Y) = 1 − max(R, G, B); for (0.8, 0.4, 0.2) that is K = 0.2.
IHS (HSI) describes colour the way people do. Intensity I = (R + G + B)/3 is the brightness; hue H is the dominant colour, an angle with 0° red, 120° green and 240° blue; saturation S = 1 − 3 min(R, G, B)/(R + G + B) is the purity, 1 for a pure colour and 0 for grey. For (0.8, 0.4, 0.2): I = 1.4/3 = 0.467 and S = 1 − 0.6/1.4 = 0.57. Yellow (1, 1, 0) has hue 60°, S = 1 and I = 0.667. The transform matters because it separates brightness from colour, so each can be handled alone.
- IHS pan-sharpening: resample the three multispectral bands to the pan pixel size, transform RGB → IHS, replace I with the high-resolution panchromatic band (after matching its histogram to I), and transform back. Spatial detail comes from the pan band, colour from H and S; it works because brightness carries the detail.
- Limitation: the pan band's spectral range differs from the three bands' intensity, so colours can be distorted; the method also handles only three bands at once.
- Contrast in IHS: stretch I and S separately, then return to RGB — this enhances brightness and colour purity without shifting the hues, which stretching R, G and B independently would do.
Key takeaways
- Ratios cancel a common illumination factor; NDVI = (NIR − Red)/(NIR + Red) lies in [−1, 1], high for vegetation, near zero or negative for water; SAVI = (1 + L)(NIR − Red)/(NIR + Red + L).
- NDWI (open water) = (Green − NIR)/(Green + NIR), positive over water; the vegetation-moisture form is (NIR − SWIR)/(NIR + SWIR). Name the bands before reading the sign.
- PCA: eigen-decomposition of the band covariance matrix; PCs are uncorrelated, PC1 has the most variance, Σλ = trace(C), PCk holds λₖ/Σλ; [a b; b a] has eigenvalues a ± b.
- Discriminant analysis is supervised and maximises between-class over within-class variance (J = Δμ²/Σσ²), with at most min(bands, classes − 1) axes; PCA is unsupervised.
- RGB adds light (24 bits = 16,777,216 colours); CMYK subtracts it (K = 1 − max(R, G, B)); IHS separates intensity (mean), hue (angle) and saturation (1 − 3 min/sum), the basis of IHS pan-sharpening.
Practice questions (17)
Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.
A vegetated pixel has reflectance 0.45 in the near-infrared and 0.05 in the red. Its NDVI is ____.
Numerical answer — type the value.
Show answer
Answer: 0.8
NDVI = (NIR − Red)/(NIR + Red) = (0.45 − 0.05)/(0.45 + 0.05) = 0.40/0.50 = 0.8. High NIR against low red is the signature of healthy green leaves.Over a lake the green reflectance is 0.10 and the NIR reflectance 0.05. Using McFeeters' NDWI = (Green − NIR)/(Green + NIR), its value to two decimal places is ____.
Numerical answer — type the value.
Show answer
Answer: 0.33
(0.10 − 0.05)/(0.10 + 0.05) = 0.05/0.15 = 0.333, so 0.33. It is positive because water reflects more green than NIR; vegetation and soil give negative values with this form.For NIR reflectance 0.5, red reflectance 0.1 and soil-adjustment factor L = 0.5, SAVI = (1 + L)(NIR − Red)/(NIR + Red + L) equals, to two decimal places, ____.
Numerical answer — type the value.
Show answer
Answer: 0.55
Numerator 1.5 × (0.5 − 0.1) = 0.6; denominator 0.5 + 0.1 + 0.5 = 1.1; 0.6/1.1 = 0.545, so 0.55. Plain NDVI would be 0.4/0.6 = 0.67; the L term damps the influence of the soil background.A hillside crop is 40% darker on its shaded side. Which transformation best removes this brightness difference?
Show answer
Answer: A — A band ratio such as NIR/Red
Shading multiplies both bands by about the same factor and the ratio divides it out. A stretch, a smoothing filter and resampling all leave the brightness difference between the two slopes in place.Two bands have variances 5 and 5 and covariance 3. The percentage of the total variance carried by the first principal component is ____.
Numerical answer — type the value.
Show answer
Answer: 80
C = [5 3; 3 5] has trace 10 and determinant 16, so λ² − 10λ + 16 = 0 and λ = 8 and 2. PC1 holds 8/(8 + 2) = 80%. Note that 8 + 2 = 10 equals the total original variance 5 + 5.Two bands with means (8, 4) and covariance matrix [5 3; 3 5] have first eigenvector (1, 1)/√2. The first-principal-component score of the pixel (10, 6), to two decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 2.83
Deviations from the means are (2, 2). Score = (1/√2)(2 + 2) = 4/1.4142 = 2.83. The second component, along (1, −1)/√2, is (2 − 2)/√2 = 0: this pixel lies exactly on the first axis.Two bands have a correlation matrix [1 0.9; 0.9 1]. The percentage of the total variance held by the first principal component of this standardised PCA is ____.
Numerical answer — type the value.
Show answer
Answer: 95
The eigenvalues of [1 r; r 1] are 1 + r = 1.9 and 1 − r = 0.1. The total is 2 (the trace), so PC1 holds 1.9/2 = 95%. The higher the correlation, the more one component summarises.Which statements about principal component analysis of a multispectral image are correct?
Show answer
Answer: B — The eigenvalues add up to the total variance of the original bands; C — The principal components are mutually uncorrelated
PCA rotates to the eigenvectors of C, giving a diagonal covariance (uncorrelated PCs) and Σλ = trace(C). It is unsupervised, and PC1 is simply the direction of greatest variance, often overall brightness, which need not be the class of interest.Two classes have mean DNs 40 and 60 in one band and each has variance 10. Fisher's criterion J = (μ₁ − μ₂)²/(σ₁² + σ₂²) is ____.
Numerical answer — type the value.
Show answer
Answer: 20
(40 − 60)² = 400 and σ₁² + σ₂² = 10 + 10 = 20, so J = 400/20 = 20. A larger J means the classes are farther apart relative to their spread and easier to separate.A six-band image is to be classified into four known classes. The maximum number of discriminant axes obtainable by linear discriminant analysis is ____.
Numerical answer — type the value.
Show answer
Answer: 3
The number of axes is at most min(bands, classes − 1) = min(6, 3) = 3: the class means of c classes span a space of only c − 1 dimensions.The essential difference between principal component analysis and discriminant analysis is that
Show answer
Answer: D — discriminant analysis uses known class labels to maximise separation, while PCA maximises total variance without them
PCA is unsupervised and looks only at the spread of the whole data; discriminant analysis is supervised and uses the training classes. Both use covariance matrices, and PCA's components are the ones that are uncorrelated.A colour has R = 0.8, G = 0.4 and B = 0.2, each on a 0–1 scale. Its saturation S = 1 − 3 min(R, G, B)/(R + G + B), to two decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 0.57
min = 0.2 and the sum is 1.4, so S = 1 − 0.6/1.4 = 1 − 0.4286 = 0.5714, so 0.57. A pure colour (one channel zero) gives S = 1; equal R, G, B give S = 0.Three colour channels of 4 bits each can represent ____ different colours.
Numerical answer — type the value.
Show answer
Answer: 4096
Each channel has 2⁴ = 16 levels, so the total is 16³ = 4096 (12 bits in all). Three 8-bit channels would give 256³ = 16,777,216.In IHS pan-sharpening, the step that brings in the fine spatial detail is
Show answer
Answer: D — replacing the intensity component with the high-resolution panchromatic band before the inverse transform
Intensity carries brightness and therefore spatial detail, whereas hue and saturation carry the colour. Substituting the pan band for I and inverting gives a sharp image with the multispectral colours. Changing H or S would alter the colours, and CMYK is only for printing.Which of the following statements about spectral indices are correct?
Show answer
Answer: B — McFeeters' NDWI, (Green − NIR)/(Green + NIR), is positive over open water; C — Multiplying both NIR and red by the same illumination factor leaves NDVI unchanged; D — NDVI always lies between −1 and +1
For non-negative reflectances |NIR − Red| ≤ NIR + Red, so NDVI is bounded by ±1; the green–NIR NDWI is positive for water; and a common factor cancels between numerator and denominator. An additive offset does not: (a + h) and (b + h) give a different quotient, which is why haze must be corrected first.Which of the following statements about colour models are correct?
Show answer
Answer: B — In IHS, hue is an angle describing the dominant colour; C — RGB is additive: red plus green light gives yellow; D — CMYK is subtractive and is the model used for printing
Screens add red, green and blue light, so R + G = yellow; printers subtract with cyan, magenta and yellow inks plus black; hue is the colour angle (0° red, 120° green, 240° blue). Saturation is purity — how far from grey — while brightness is intensity.A colour has R = 0.8, G = 0.4 and B = 0.2 on a 0–1 scale. Its black (K) component in the CMYK model, K = 1 − max(R, G, B), is ____.
Numerical answer — type the value.
Show answer
Answer: 0.2
max(R, G, B) = 0.8, so K = 1 − 0.8 = 0.2. Equivalently C = 0.2, M = 0.6, Y = 0.8 and K = min(C, M, Y) = 0.2: the black ink replaces the grey component that all three inks share.