Image Enhancement: Noise and Spatial Filters, Contrast Enhancement (Linear and Non-linear)

This is a section of Part B2, Image Processing and Analysis, of the Geomatics Engineering (GE) paper. B2's fourth heading is “Image Enhancement: Spatial Enhancement: Noise and Spatial filters, Contrast Enhancement: Linear and Non-linear methods” — operations that make an image easier to interpret without changing what it records. This chapter separates the four kinds of noise and their remedies, and the noise reduction bought by averaging N values (σ/√N); reads kernels by their sum (smoothing 1, edges 0) and works the mean, median, Sobel and unsharp-masking operations on small windows; then turns to the histogram, with the linear min–max, saturation and piecewise stretches and the non-linear logarithmic, gamma and equalisation transforms. The numericals are window filters, stretched DNs, equalised levels and output sizes.

1. Noise and its sources

Noise is variation in DN that does not come from the scene. Four kinds are examined. Random (Gaussian) noise is electronic and thermal: DN = true value + a random error of standard deviation σ. Impulse (salt-and-pepper) noise is isolated pixels stuck at the extremes, 0 or 255, from bit errors or dropped data. Periodic noise repeats regularly — striping and banding from mismatched detectors, or a stray electrical frequency — and shows as sharp spikes in the Fourier spectrum. Speckle in radar images is multiplicative: coherent waves from the scatterers inside one resolution cell interfere, so even a uniform field looks grainy. The signal-to-noise ratio is SNR = μ/σ, often quoted in decibels as 20 log₁₀(μ/σ); μ = 100 and σ = 1 give SNR 100, or 40 dB.

Noise types and their usual remedy
NoiseLooks likeUsual remedy
Gaussianfine grain over the whole imagemean or Gaussian low-pass filter; averaging repeated frames
Impulse (salt and pepper)isolated black and white dotsmedian filter — an average would smear each dot over its neighbours
Periodicregular stripes or a grid patternnotch filter in the Fourier domain; detector destriping
Speckle (radar)multiplicative grain, worse in bright areasmultilook averaging; adaptive filters such as Lee and Frost

Averaging is the basic remedy for random noise: the mean of N independent noisy values has standard deviation σ/√N. A 3 × 3 window averages 9 pixels and cuts σ by a factor of 3; a 5 × 5 window by 5. The price is blur, because the window also averages across edges — so a larger window removes more noise and destroys more detail.

⚠️ Noise reduction and sharpness pull in opposite directions
No linear smoothing filter reduces noise without also reducing edge contrast. The median filter is the standard exception for impulse noise because it removes an isolated outlier and leaves a step edge almost intact.

2. Spatial filters: smoothing, edges and sharpening

A spatial filter replaces each pixel by a weighted sum of its neighbourhood, the weights forming a kernel that is convolved with the image (see the linear-systems chapter). The sum of the kernel weights decides the behaviour: a low-pass (smoothing) kernel sums to 1, so flat areas keep their brightness; a high-pass (edge) kernel sums to 0, so flat areas go to zero and only changes in brightness survive. A median filter is not a kernel at all: it sorts the neighbourhood and takes the middle value, which makes it non-linear.

Common 3 × 3 kernels
FilterKernel (rows)SumEffect
Meanall nine weights 1/91smooths, blurs edges
Weighted (Gaussian-like)[1 2 1; 2 4 2; 1 2 1] / 161smooths with less blur than the mean
Laplacian[0 1 0; 1 −4 1; 0 1 0]0non-directional edges and points
Sobel (horizontal gradient)[−1 0 1; −2 0 2; −1 0 1]0vertical edges; the transpose finds horizontal edges
Sharpening[−1 −1 −1; −1 9 −1; −1 −1 −1]1original plus its edges: detail boosted, brightness kept

Worked mean and median. A window holds 12 15 18 / 10 90 14 / 16 13 11, the centre 90 being an impulse. The mean is 199/9 = 22.11 — the spike drags the result far above every neighbour. Sorted, the values are 10 11 12 13 14 15 16 18 90, so the median is the fifth, 14, which sits among the neighbours and the impulse is gone. Worked Sobel: a vertical step with 10 on the left two columns and 40 on the right, in all three rows, gives Gx = 30 + 60 + 30 = 120 and Gy = −70 + 70 = 0, so the gradient magnitude is √(Gx² + Gy²) = 120, often approximated by |Gx| + |Gy|.

  • Unsharp masking / high-boost: sharpened = original + k × (original − blurred). A pixel of 50 whose 3 × 3 mean is 40, with k = 1.5, becomes 50 + 1.5 × 10 = 65. The bracket is the high-pass detail; k sets how much is added back.
  • Output size: sliding an M × M kernel over an N × N image without padding leaves (N − M + 1)² outputs — 508 × 508 for a 5 × 5 kernel on 512 × 512. Software either pads the border (zeros, replicated or mirrored pixels) or drops the outer rim.
  • In the frequency domain a smoothing kernel is a low-pass filter and an edge kernel a high-pass filter. An ideal (sharp cut-off) low-pass filter causes ringing near edges; a Gaussian cut-off does not.
🧠 Read the kernel sum before anything else
Sum 1: a smoother or a sharpener that keeps mean brightness. Sum 0: an edge detector whose output is zero on flat ground. A kernel summing to 2 would double the brightness of the whole image.

3. Contrast enhancement: linear methods

A sensor records over a wide DN range, but most scenes occupy a narrow part of it, so the histogram is bunched and the display looks flat. Contrast enhancement re-maps the DNs onto the full display range 0–255, and changes only appearance — it adds no information. The linear (min–max) stretch is DN_out = 255 × (DN − min)/(max − min), rounded. With min 40 and max 120 the gain is 255/80 = 3.19, and DN 100 becomes 255 × 60/80 = 191.25, so 191. A straight line in the transfer plot means equal DN differences stay equal.

Extreme DNs from a few noisy or unusual pixels can waste most of the range, so the practical form is the saturation (clipped) stretch: choose the limits at, say, the 2nd and 98th percentiles, or at mean ± kσ, and send everything beyond them to 0 or 255. For mean 80 and σ = 10 with k = 2 the limits are 60 and 100; DN 75 becomes 255 × 15/40 = 95.6, so 96, and DN 55 goes to 0. A piecewise linear stretch uses several segments with different slopes, spending more of the display range on the DN band that holds the class of interest. The negative (inverse) transform DN_out = 255 − DN reverses light and dark.

⚠️ Stretching does not sharpen and does not restore
A stretch spreads the same 80 DN levels over 256 display levels; it does not create new levels or remove noise, and the noise is stretched with the signal. The clipped tails are lost for good.

4. Contrast enhancement: non-linear methods

Histogram equalisation uses the image's own cumulative histogram as the transfer curve: s_k = round((L − 1) × Σⱼ≤ₖ p(j)), with L levels. It spends more output levels where many pixels are, and its result has an approximately flat histogram. Worked, for a 3-bit image (L = 8) of 64 pixels with counts 16, 24, 16 and 8 at DN 2, 3, 4 and 5: the cumulative fractions are 0.25, 0.625, 0.875 and 1, so the new levels are round(1.75) = 2, round(4.375) = 4, round(6.125) = 6 and 7. The four occupied levels have moved apart, giving a wider spread. The mapping is monotonic but discrete, so levels can merge and a perfectly flat histogram is never reached. Histogram matching (specification) forces the histogram to a chosen shape — used to make two dates look alike before comparison.

Non-linear transfer functions
MethodTransfer functionEffect
Logarithmicoutput DN = 255 × log(1 + DN)/log 256expands dark values, compresses bright; DN 3 → 255 × log 4/log 256 = 63.75 ≈ 64
Power law (gamma)output DN = 255 × (DN/255)^γγ < 1 brightens and expands the dark end; γ > 1 darkens and expands the bright end
ExponentialDN_out grows as an exponential of DNexpands bright values, compresses dark
Histogram equalisationthe cumulative histogramautomatic; favours the frequent DNs
Gaussian stretchmatches a normal histogramretains tails better than equalisation

Choosing. Linear stretches keep the relative spacing of tones and suit roughly symmetric histograms. A histogram skewed toward dark values — water, shadow, or a low-reflectance scene — is better served by a logarithmic or γ < 1 stretch, which opens the dark end that a linear stretch would leave crushed. Equalisation needs no parameters and shows subtle contrast in the most populous range, but it can exaggerate noise in flat regions and gives unnatural tones, so it is a viewing aid rather than a preparation for classification.

🎯 Why a gamma below 1 opens the shadows
For γ = 0.5 the curve is a square root: a DN one quarter of full scale is sent to one half. Small inputs get a steep slope, so the dark end is stretched, while the bright end, where the curve flattens, is compressed.

Key takeaways

  • Random noise falls as σ/√N with averaging (3 × 3 cuts σ threefold) but averaging blurs edges; impulse noise needs the median filter; periodic noise a notch filter; speckle multilook averaging.
  • Kernel sum 1 = smoothing or sharpening that keeps brightness; sum 0 = edge detection (Laplacian, Sobel); the median filter is non-linear and edge-preserving.
  • Sobel gradient magnitude = √(Gx² + Gy²); unsharp masking adds k × (original − blurred) back; a valid convolution leaves (N − M + 1)² pixels.
  • Linear stretch: DN_out = 255(DN − min)/(max − min); clip at mean ± kσ or percentiles to avoid wasting range on outliers.
  • Non-linear: equalisation s_k = round((L − 1) × cumulative fraction); logarithm and γ < 1 open the dark end, γ > 1 the bright end; enhancement changes appearance, never information.

Practice questions (17)

Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.

  1. Independent random noise of standard deviation 12 DN affects every pixel. After averaging a 4 × 4 window of a uniform area, the noise standard deviation is ____ DN.

    Numerical answer — type the value.

    Show answer

    Answer: 3

    Averaging N = 16 independent values divides σ by √16 = 4: 12/4 = 3 DN. The same average also blurs any edge inside the window, which is the price of the noise reduction.
  2. Isolated black and white pixels scattered over an image are best removed by a

    1. median filter
    2. mean filter
    3. linear contrast stretch
    4. Laplacian filter
    Show answer

    Answer: A — median filter

    The median picks the middle value of the neighbourhood, so a lone outlier is discarded. A mean would smear each dot over its neighbours, a Laplacian would highlight it, and a stretch would only make it more visible.
  3. A 3 × 3 window contains 12 15 18 / 10 90 14 / 16 13 11. The output of a median filter at the centre is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 14

    Sorted: 10, 11, 12, 13, 14, 15, 16, 18, 90. The fifth of nine values is 14. The impulse 90 is at the top of the list and never chosen; a mean filter would return 199/9 = 22.11.
  4. The same window (12 15 18 / 10 90 14 / 16 13 11) is passed through a 3 × 3 mean filter. The centre output, to two decimal places, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 22.11

    Row sums are 45, 114 and 40, total 199; 199/9 = 22.11. One impulse of 90 raises the output about 8 units above the median, showing why the mean is a poor choice against salt-and-pepper noise.
  5. A 3 × 3 window has 10 10 40 in each of its three rows. The horizontal-gradient Sobel kernel [−1 0 1; −2 0 2; −1 0 1] gives, at the centre, the value ____.

    Numerical answer — type the value.

    Show answer

    Answer: 120

    Each row gives −10 + 40 = 30 with weights 1, 2, 1: 30 + 2 × 30 + 30 = 120. The vertical Gy is 0 because the three rows are identical; the gradient magnitude is √(120² + 0²) = 120, a strong vertical edge.
  6. A pixel of DN 50 has a 3 × 3 mean of 40. Unsharp masking with sharpened = original + k × (original − blurred) and k = 1.5 gives the value ____.

    Numerical answer — type the value.

    Show answer

    Answer: 65

    Detail = 50 − 40 = 10; added back with weight 1.5: 50 + 15 = 65. With k = 0 the image is unchanged; larger k gives crisper edges and more visible noise.
  7. A 5 × 5 kernel is applied to a 512 × 512 image with no padding, keeping only positions where the kernel fully overlaps the image. The output has ____ pixels along each side.

    Numerical answer — type the value.

    Show answer

    Answer: 508

    Output side = N − M + 1 = 512 − 5 + 1 = 508; two rows and columns are lost on every edge. Padding the border with zeros or mirrored values would keep 512.
  8. A 3 × 3 kernel has weights −1 −1 −1 / −1 8 −1 / −1 −1 −1. On a perfectly uniform area of DN 90 its output is

    1. −90
    2. 90
    3. 0
    4. 720
    Show answer

    Answer: C — 0

    The weights sum to 8 − 8 = 0, so a uniform area gives 0 × 90 = 0. It is a high-pass (edge) kernel: only changes in brightness produce a response. 720 would be the centre term alone.
  9. A band has min 40 and max 120. A linear min–max stretch to 0–255 turns DN 100 into (nearest integer) ____.

    Numerical answer — type the value.

    Show answer

    Answer: 191

    DN_out = 255 × (100 − 40)/(120 − 40) = 255 × 0.75 = 191.25, which rounds to 191. Forgetting to subtract the minimum gives 255 × 100/120 = 212.
  10. A band has mean 80 and standard deviation 10. A saturation stretch sends mean − 2σ to 0 and mean + 2σ to 255, linearly between. DN 75 becomes (nearest integer) ____.

    Numerical answer — type the value.

    Show answer

    Answer: 96

    Limits are 60 and 100. DN_out = 255 × (75 − 60)/(100 − 60) = 255 × 0.375 = 95.6, so 96. Pixels below 60 saturate to 0 and above 100 to 255, which is why this stretch has more contrast than min–max.
  11. A 3-bit image (8 levels) of 64 pixels has 16, 24, 16 and 8 pixels at DN 2, 3, 4 and 5 and none elsewhere. Using s_k = round(7 × cumulative fraction), histogram equalisation maps DN 3 to ____.

    Numerical answer — type the value.

    Show answer

    Answer: 4

    The cumulative count up to DN 3 is 16 + 24 = 40, a fraction 0.625. Then 7 × 0.625 = 4.375, which rounds to 4. DN 2 → round(1.75) = 2, DN 4 → round(6.125) = 6, DN 5 → 7.
  12. A logarithmic stretch DN_out = 255 × log(1 + DN)/log 256 is applied to DN 3. The output, to the nearest integer, is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 64

    log(1 + 3) = log 4 and log 256 = 4 log 4, hence the ratio is exactly 1/4 for any base and DN_out = 255/4 = 63.75, which rounds to 64. A DN of only 3 is lifted to a quarter of the display range, opening the dark end.
  13. An image of a lake and dark forest has a histogram crowded at low DNs with a long tail to the right. Which enhancement best shows detail in the dark areas?

    1. A γ > 1 power law
    2. A 3 × 3 mean filter
    3. A logarithmic or γ < 1 stretch
    4. A negative (inverse) transform
    Show answer

    Answer: C — A logarithmic or γ < 1 stretch

    Both curves have a steep slope at small inputs and flatten at large ones, so they spread the crowded dark DNs over many display levels. A γ > 1 curve does the opposite. A negative only swaps light and dark, and a mean filter smooths detail away.
  14. Which of the following operations are linear and shift-invariant, so that they can be written as a convolution with a fixed kernel?

    1. A 3 × 3 mean filter
    2. The Laplacian filter
    3. The Sobel Gx filter
    4. The 3 × 3 median filter
    Show answer

    Answer: A — A 3 × 3 mean filter; B — The Laplacian filter; C — The Sobel Gx filter

    The mean, Laplacian and Sobel filters are fixed weighted sums, hence linear and shift-invariant. The median sorts the neighbourhood, and the median of a sum is not the sum of the medians, so it is not linear and no kernel represents it.
  15. Which of the following are true of histogram equalisation?

    1. It gives more output levels to frequently occurring DNs
    2. Its transfer function is the image's own cumulative histogram scaled to the output range
    3. It can merge different input DNs into one output DN
    4. It produces a perfectly flat histogram for any discrete image
    Show answer

    Answer: A — It gives more output levels to frequently occurring DNs; B — Its transfer function is the image's own cumulative histogram scaled to the output range; C — It can merge different input DNs into one output DN

    The CDF is the mapping, and it is steepest where pixels are numerous, so common DNs get more levels; with rounding, close DNs can land on one level. An exactly flat histogram is impossible for discrete data because all pixels of one input level move together.
  16. Which statements about a linear contrast stretch are correct?

    1. It changes the appearance but adds no new information to the data
    2. It removes the haze contribution from the histogram
    3. Noise is stretched together with the signal
    4. Clipping the tails at mean ± 2σ usually gives more contrast in the bulk of the image than a min–max stretch
    Show answer

    Answer: A — It changes the appearance but adds no new information to the data; C — Noise is stretched together with the signal; D — Clipping the tails at mean ± 2σ usually gives more contrast in the bulk of the image than a min–max stretch

    A stretch is a re-labelling of DNs for display: no information is created, clipping frees range for the majority of pixels, and noise scales with the signal. Removing haze needs a radiometric correction such as dark-object subtraction; a stretch only moves the histogram.
  17. Radar images of a uniform crop field look grainy even though the field is homogeneous. The most appropriate description and treatment is

    1. multiplicative speckle from coherent interference, reduced by multilook averaging or adaptive filtering
    2. atmospheric haze, removed by dark-object subtraction
    3. quantization error, removed by using more bits
    4. periodic noise from detector mismatch, removed by a notch filter
    Show answer

    Answer: A — multiplicative speckle from coherent interference, reduced by multilook averaging or adaptive filtering

    Speckle comes from the coherent addition of waves scattered by many elements inside a resolution cell; it is multiplicative and inherent to coherent imaging. Averaging independent looks or an adaptive filter such as Lee's reduces it. The other three describe different defects.