Digital Image Characteristics: Histogram, Scattergram, and the Variance–Covariance and Correlation Matrices

This is a section of Part B2, Image Processing and Analysis, of the Geomatics Engineering (GE) paper. B2's second heading asks for the numbers that describe an image before anything is done to it: “Digital image characteristics: image histogram and scattergram and their significance, Variance-Covariance and Correlation matrix and their significance”. This chapter defines the digital image — rows, columns, bands and bit depth, and the BSQ, BIL and BIP ways of storing it; reads the histogram for brightness, contrast, saturation, modes and thresholds, and shows what it cannot tell; turns two bands into a scattergram (feature space), where land covers form clusters and soil forms a line; and computes the variance–covariance and correlation matrices, which measure inter-band redundancy and are the input to principal components and maximum-likelihood classification. The numericals are image statistics, covariance and correlation from a few pixels, the variance of a band sum and file sizes.

1. The digital image and its statistics

A multispectral digital image is a three-dimensional array of digital numbers (DN): rows × columns × bands, each DN stored in n bits. Its size is rows × columns × bands × bits/8 bytes: 1000 × 1000 pixels in 3 bands at 8 bits is 3 × 10⁶ bytes. Three interleaving formats arrange the same numbers differently: BSQ (band sequential: all of band 1, then band 2 — fast for single-band work), BIL (band interleaved by line: line 1 of every band, then line 2) and BIP (band interleaved by pixel: all bands of pixel 1, then pixel 2 — fast for per-pixel spectral work such as classification).

Univariate statistics of a band: mean μ = ΣDN/N, variance σ² = Σ(DN − μ)²/(N − 1), standard deviation σ, minimum, maximum, median and mode. For the six DNs 12, 15, 18, 15, 20, 10: μ = 90/6 = 15, the deviations −3, 0, 3, 0, 5, −5 give Σ = 68, σ² = 68/5 = 13.6 and σ = 3.69. The mean measures overall brightness; the standard deviation measures contrast — the spread of values the display has to show.

2. The image histogram and its significance

The histogram plots the number of pixels at each DN. Divided by the pixel count it is the normalised histogram p(DN), an estimate of the probability of each value, and its running sum is the cumulative histogram, the basis of histogram equalisation and matching. In a 3-bit image of 64 pixels with 16 pixels at DN 4, p(4) = 16/64 = 0.25.

Reading a histogram
ShapeWhat it meansWhat to do
Narrow, bunched in a small DN rangelow contrast: the sensor range is under-usedcontrast stretch
Shifted toward low or high DNdark or bright image; a lifted minimum in the blue band indicates hazeoffset, dark-object subtraction
Spikes at 0 or at the maximum DNsaturation (clipping): information lost therecannot be recovered; note it
Two or more peaks (bimodal)distinct classes, e.g. water and land in an NIR bandthreshold in the valley between the peaks
⚠️ A histogram has no geography
Shuffling every pixel of an image into random positions leaves its histogram unchanged. Two very different images can therefore share one histogram, and nothing about shape, texture or where a class lies can be read from it.

3. The scattergram and feature space

A scattergram (two-dimensional histogram, feature-space plot) places every pixel at the point (DN in band i, DN in band j). Its significance: the shape of the cloud shows how correlated the bands are — a narrow diagonal ellipse means they carry largely the same information, a round cloud means they are independent; clusters are spectral classes, whose separation shows how well two bands discriminate them; in the red–NIR scattergram bare soils fall along a straight soil line, vegetation lies well above it toward high NIR, and water lies near the origin. Scattergrams are how training samples are checked, how parallelepiped boxes and cluster centres are visualised, and why a pair of bands is chosen over another.

ℹ️ Why the soil line matters
Soil brightness changes with moisture and roughness, but its red and NIR reflectances change together, so soils slide along one line. Vegetation indices measure how far a pixel lies from that line toward high NIR — the idea behind NDVI and SAVI in the transformation chapter.

4. The variance–covariance and correlation matrices

For k bands the variance–covariance matrix C is k × k with Cᵢⱼ = Σ(DNᵢ − μᵢ)(DNⱼ − μⱼ)/(N − 1): the variances on the diagonal, the covariances off it, and C symmetric, so it has k(k + 1)/2 distinct elements (28 for 7 bands). The correlation matrix R standardises it: rᵢⱼ = Cᵢⱼ/(σᵢσⱼ), with 1 on the diagonal and −1 ≤ rᵢⱼ ≤ 1. Worked: bands 1 = [2, 4, 6, 8] and 2 = [1, 3, 2, 6] have means 5 and 3, deviations (−3, −2), (−1, 0), (1, −1), (3, 3), cross-products summing to 14; so C₁₂ = 14/3 = 4.667, C₁₁ = 20/3, C₂₂ = 14/3, and r₁₂ = 14/√(20 × 14) = 0.837. A covariance of 12 between bands of variance 16 and 25 is r = 12/(4 × 5) = 0.6.

  • Redundancy: visible bands are often correlated above 0.9, so a three-band display of them shows little more than one band. High correlation is the reason principal component analysis compresses multispectral data.
  • Band selection: the optimum index factor OIF = (σᵢ + σⱼ + σₖ)/(|rᵢⱼ| + |rᵢₖ| + |rⱼₖ|) favours a band triple with high variance and low mutual correlation; SDs 10, 12, 15 with correlations 0.5, 0.6, 0.7 give OIF = 37/1.8 = 20.6.
  • Classification: the class covariance matrix defines the shape of each class in the maximum-likelihood and Mahalanobis classifiers.
  • Combining bands: var(a + b) = var(a) + var(b) + 2 cov(a, b); for variances 16 and 25 and covariance 12 it is 65, and var(a − b) = 17 — which is why a difference of correlated bands has low variance.
⚠️ Covariance depends on units; correlation does not
Converting DN to radiance with a gain multiplies a covariance by the product of the two gains but leaves the correlation unchanged. So PCA on the covariance matrix is dominated by the high-variance bands, and PCA on the correlation matrix (standardised PCA) treats all bands equally.

Key takeaways

  • Image size = rows × columns × bands × bits/8; BSQ, BIL and BIP store the same DNs in different orders.
  • Histogram: mean ≈ brightness, spread ≈ contrast, spikes at the ends = saturation, bimodal = separable classes; it carries no spatial information.
  • Scattergram: elongated cloud = correlated bands, clusters = spectral classes, soil line in red–NIR space.
  • Cᵢⱼ = Σ(xᵢ − μᵢ)(xⱼ − μⱼ)/(N − 1); rᵢⱼ = Cᵢⱼ/(σᵢσⱼ); C is symmetric with k(k + 1)/2 distinct elements.
  • High inter-band correlation means redundancy (the basis of PCA); var(a ± b) = var a + var b ± 2cov; OIF picks low-correlation, high-variance band triples.

Practice questions (14)

Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.

  1. An uncompressed image has 1000 × 1000 pixels in 3 bands at 8 bits per pixel per band. Its size, in megabytes (1 MB = 10⁶ bytes), is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 3

    Size = 1000 × 1000 × 3 × 8/8 bytes = 3,000,000 bytes = 3 MB. The storage order (BSQ, BIL or BIP) changes where each value sits, not how many there are.
  2. Six pixels of a band have DNs 12, 15, 18, 15, 20 and 10. Their sample standard deviation (divisor N − 1), to two decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 3.69

    Mean = 90/6 = 15. Deviations −3, 0, 3, 0, 5, −5; squares sum to 9 + 0 + 9 + 0 + 25 + 25 = 68. Variance = 68/5 = 13.6, σ = 3.69. With divisor N = 6 it would be 3.37.
  3. In a 3-bit image of 64 pixels, 16 pixels have DN 4. The value of the normalised histogram at DN 4 is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.25

    p(4) = n₄/N = 16/64 = 0.25, the estimated probability that a pixel has DN 4. The normalised histogram of all eight levels sums to 1; the raw count, 16, is the unnormalised histogram.
  4. The histogram of a near-infrared band over a coastal area shows two well-separated peaks, one at low DN and one at high DN. The most likely interpretation is

    1. water (low DN) and land with vegetation (high DN)
    2. sensor saturation at both ends
    3. atmospheric haze
    4. a low-contrast image needing a stretch
    Show answer

    Answer: A — water (low DN) and land with vegetation (high DN)

    Water absorbs NIR and is dark; vegetated land reflects NIR strongly and is bright, so a bimodal NIR histogram separates them, and a threshold in the valley maps the shoreline. Saturation appears as spikes exactly at 0 or the maximum DN; haze lifts the minimum mainly in the blue.
  5. The scattergram of two bands shows a narrow ellipse along the 1:1 diagonal. This indicates that the two bands

    1. are highly correlated and largely redundant
    2. are statistically independent
    3. are negatively correlated
    4. separate vegetation from soil ideally
    Show answer

    Answer: A — are highly correlated and largely redundant

    Points lying close to a line of positive slope mean that when one band is bright the other is too: correlation near +1, so the second band adds little. Independent bands give a round cloud, and a negative correlation lies along a falling diagonal.
  6. Two bands have variances 16 and 25 and a covariance of 12. Their correlation coefficient is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.6

    r = cov/(σ₁σ₂) = 12/(√16 × √25) = 12/(4 × 5) = 0.6. Dividing by the product of the variances (400) instead of the standard deviations gives 0.03.
  7. Four pixels have DNs (2, 1), (4, 3), (6, 2) and (8, 6) in bands 1 and 2. The sample covariance between the bands (divisor N − 1), to two decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 4.67

    Means 5 and 3. Deviations (−3, −2), (−1, 0), (1, −1), (3, 3); cross-products 6, 0, −1, 9, sum 14. Covariance = 14/3 = 4.67. Dividing by N gives 3.50.
  8. For the same four pixels, (2, 1), (4, 3), (6, 2) and (8, 6), the correlation coefficient between bands 1 and 2 (to three decimal places) is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 0.837

    Σ cross-products = 14; Σ squared deviations are 9 + 1 + 1 + 9 = 20 for band 1 and 4 + 0 + 1 + 9 = 14 for band 2. r = 14/√(20 × 14) = 14/16.733 = 0.837. The divisor N − 1 cancels, so it does not matter here.
  9. Two bands have variances 16 and 25 and a covariance of 12. The variance of their sum (band 1 + band 2) is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 65

    var(a + b) = var(a) + var(b) + 2cov(a, b) = 16 + 25 + 24 = 65. Ignoring the covariance gives 41, which is right only for uncorrelated bands; the difference a − b would have variance 17.
  10. The variance–covariance matrix of a 7-band image is symmetric. The number of distinct elements it contains is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 28

    A symmetric k × k matrix has k(k + 1)/2 distinct entries: 7 × 8/2 = 28 — 7 variances and 21 covariances. The full matrix has 49 entries, but Cᵢⱼ = Cⱼᵢ.
  11. Three bands have standard deviations 10, 12 and 15, and their pairwise correlation coefficients are 0.5, 0.6 and 0.7. Their optimum index factor (to two decimal places) is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 20.56

    OIF = Σσ/Σ|r| = (10 + 12 + 15)/(0.5 + 0.6 + 0.7) = 37/1.8 = 20.56. A higher OIF means more total variance and less duplication, so the triple with the largest OIF makes the most informative colour composite.
  12. Which of the following can be learnt from an image histogram alone?

    1. Whether the image uses only a small part of the available DN range
    2. Whether pixels are saturated at the maximum DN
    3. A threshold that separates two classes with distinct brightness
    4. Where in the image a particular land cover is located
    Show answer

    Answer: A — Whether the image uses only a small part of the available DN range; B — Whether pixels are saturated at the maximum DN; C — A threshold that separates two classes with distinct brightness

    The histogram shows the DN range used, spikes at the extremes where values are clipped, and valleys between modes where a threshold separates classes. It discards all positional information, so where a land cover lies cannot be read from it.
  13. Which of the following statements about the variance–covariance and correlation matrices of a multispectral image are correct?

    1. The diagonal of the correlation matrix is all ones
    2. Both matrices are symmetric
    3. Rescaling one band by a positive gain changes its covariances but not its correlations
    4. A correlation of 0.95 between two bands means they carry nearly independent information
    Show answer

    Answer: A — The diagonal of the correlation matrix is all ones; B — Both matrices are symmetric; C — Rescaling one band by a positive gain changes its covariances but not its correlations

    Each band correlates perfectly with itself; cov(i, j) = cov(j, i) makes both symmetric; and a gain g multiplies cov(i, j) by g but also σᵢ by g, cancelling in r. A correlation of 0.95 means the bands are nearly redundant, the opposite of independent.
  14. In a red–near-infrared scattergram of an agricultural scene, which of the following are expected?

    1. Bare soils lie along a roughly straight soil line
    2. Dense vegetation lies at high NIR and low red
    3. Water lies near the origin, low in both bands
    4. Vegetation and soil fall on the same point
    Show answer

    Answer: A — Bare soils lie along a roughly straight soil line; B — Dense vegetation lies at high NIR and low red; C — Water lies near the origin, low in both bands

    Soil red and NIR reflectances rise together, forming the soil line; vegetation absorbs red and reflects NIR, so it sits far above that line; water is dark in both. That separation is exactly what the red–NIR pair is chosen for, so vegetation and soil do not coincide.