Image Enhancement: Noise and Spatial Filters, Contrast Enhancement (Linear and Non-linear)
1. Noise and its sources
Noise is variation in DN that does not come from the scene. Four kinds are examined. Random (Gaussian) noise is electronic and thermal: DN = true value + a random error of standard deviation σ. Impulse (salt-and-pepper) noise is isolated pixels stuck at the extremes, 0 or 255, from bit errors or dropped data. Periodic noise repeats regularly — striping and banding from mismatched detectors, or a stray electrical frequency — and shows as sharp spikes in the Fourier spectrum. Speckle in radar images is multiplicative: coherent waves from the scatterers inside one resolution cell interfere, so even a uniform field looks grainy. The signal-to-noise ratio is SNR = μ/σ, often quoted in decibels as 20 log₁₀(μ/σ); μ = 100 and σ = 1 give SNR 100, or 40 dB.
| Noise | Looks like | Usual remedy |
|---|---|---|
| Gaussian | fine grain over the whole image | mean or Gaussian low-pass filter; averaging repeated frames |
| Impulse (salt and pepper) | isolated black and white dots | median filter — an average would smear each dot over its neighbours |
| Periodic | regular stripes or a grid pattern | notch filter in the Fourier domain; detector destriping |
| Speckle (radar) | multiplicative grain, worse in bright areas | multilook averaging; adaptive filters such as Lee and Frost |
Averaging is the basic remedy for random noise: the mean of N independent noisy values has standard deviation σ/√N. A 3 × 3 window averages 9 pixels and cuts σ by a factor of 3; a 5 × 5 window by 5. The price is blur, because the window also averages across edges — so a larger window removes more noise and destroys more detail.
2. Spatial filters: smoothing, edges and sharpening
A spatial filter replaces each pixel by a weighted sum of its neighbourhood, the weights forming a kernel that is convolved with the image (see the linear-systems chapter). The sum of the kernel weights decides the behaviour: a low-pass (smoothing) kernel sums to 1, so flat areas keep their brightness; a high-pass (edge) kernel sums to 0, so flat areas go to zero and only changes in brightness survive. A median filter is not a kernel at all: it sorts the neighbourhood and takes the middle value, which makes it non-linear.
| Filter | Kernel (rows) | Sum | Effect |
|---|---|---|---|
| Mean | all nine weights 1/9 | 1 | smooths, blurs edges |
| Weighted (Gaussian-like) | [1 2 1; 2 4 2; 1 2 1] / 16 | 1 | smooths with less blur than the mean |
| Laplacian | [0 1 0; 1 −4 1; 0 1 0] | 0 | non-directional edges and points |
| Sobel (horizontal gradient) | [−1 0 1; −2 0 2; −1 0 1] | 0 | vertical edges; the transpose finds horizontal edges |
| Sharpening | [−1 −1 −1; −1 9 −1; −1 −1 −1] | 1 | original plus its edges: detail boosted, brightness kept |
Worked mean and median. A window holds 12 15 18 / 10 90 14 / 16 13 11, the centre 90 being an impulse. The mean is 199/9 = 22.11 — the spike drags the result far above every neighbour. Sorted, the values are 10 11 12 13 14 15 16 18 90, so the median is the fifth, 14, which sits among the neighbours and the impulse is gone. Worked Sobel: a vertical step with 10 on the left two columns and 40 on the right, in all three rows, gives Gx = 30 + 60 + 30 = 120 and Gy = −70 + 70 = 0, so the gradient magnitude is √(Gx² + Gy²) = 120, often approximated by |Gx| + |Gy|.
- Unsharp masking / high-boost: sharpened = original + k × (original − blurred). A pixel of 50 whose 3 × 3 mean is 40, with k = 1.5, becomes 50 + 1.5 × 10 = 65. The bracket is the high-pass detail; k sets how much is added back.
- Output size: sliding an M × M kernel over an N × N image without padding leaves (N − M + 1)² outputs — 508 × 508 for a 5 × 5 kernel on 512 × 512. Software either pads the border (zeros, replicated or mirrored pixels) or drops the outer rim.
- In the frequency domain a smoothing kernel is a low-pass filter and an edge kernel a high-pass filter. An ideal (sharp cut-off) low-pass filter causes ringing near edges; a Gaussian cut-off does not.
3. Contrast enhancement: linear methods
A sensor records over a wide DN range, but most scenes occupy a narrow part of it, so the histogram is bunched and the display looks flat. Contrast enhancement re-maps the DNs onto the full display range 0–255, and changes only appearance — it adds no information. The linear (min–max) stretch is DN_out = 255 × (DN − min)/(max − min), rounded. With min 40 and max 120 the gain is 255/80 = 3.19, and DN 100 becomes 255 × 60/80 = 191.25, so 191. A straight line in the transfer plot means equal DN differences stay equal.
Extreme DNs from a few noisy or unusual pixels can waste most of the range, so the practical form is the saturation (clipped) stretch: choose the limits at, say, the 2nd and 98th percentiles, or at mean ± kσ, and send everything beyond them to 0 or 255. For mean 80 and σ = 10 with k = 2 the limits are 60 and 100; DN 75 becomes 255 × 15/40 = 95.6, so 96, and DN 55 goes to 0. A piecewise linear stretch uses several segments with different slopes, spending more of the display range on the DN band that holds the class of interest. The negative (inverse) transform DN_out = 255 − DN reverses light and dark.
4. Contrast enhancement: non-linear methods
Histogram equalisation uses the image's own cumulative histogram as the transfer curve: s_k = round((L − 1) × Σⱼ≤ₖ p(j)), with L levels. It spends more output levels where many pixels are, and its result has an approximately flat histogram. Worked, for a 3-bit image (L = 8) of 64 pixels with counts 16, 24, 16 and 8 at DN 2, 3, 4 and 5: the cumulative fractions are 0.25, 0.625, 0.875 and 1, so the new levels are round(1.75) = 2, round(4.375) = 4, round(6.125) = 6 and 7. The four occupied levels have moved apart, giving a wider spread. The mapping is monotonic but discrete, so levels can merge and a perfectly flat histogram is never reached. Histogram matching (specification) forces the histogram to a chosen shape — used to make two dates look alike before comparison.
| Method | Transfer function | Effect |
|---|---|---|
| Logarithmic | output DN = 255 × log(1 + DN)/log 256 | expands dark values, compresses bright; DN 3 → 255 × log 4/log 256 = 63.75 ≈ 64 |
| Power law (gamma) | output DN = 255 × (DN/255)^γ | γ < 1 brightens and expands the dark end; γ > 1 darkens and expands the bright end |
| Exponential | DN_out grows as an exponential of DN | expands bright values, compresses dark |
| Histogram equalisation | the cumulative histogram | automatic; favours the frequent DNs |
| Gaussian stretch | matches a normal histogram | retains tails better than equalisation |
Choosing. Linear stretches keep the relative spacing of tones and suit roughly symmetric histograms. A histogram skewed toward dark values — water, shadow, or a low-reflectance scene — is better served by a logarithmic or γ < 1 stretch, which opens the dark end that a linear stretch would leave crushed. Equalisation needs no parameters and shows subtle contrast in the most populous range, but it can exaggerate noise in flat regions and gives unnatural tones, so it is a viewing aid rather than a preparation for classification.
Key takeaways
- Random noise falls as σ/√N with averaging (3 × 3 cuts σ threefold) but averaging blurs edges; impulse noise needs the median filter; periodic noise a notch filter; speckle multilook averaging.
- Kernel sum 1 = smoothing or sharpening that keeps brightness; sum 0 = edge detection (Laplacian, Sobel); the median filter is non-linear and edge-preserving.
- Sobel gradient magnitude = √(Gx² + Gy²); unsharp masking adds k × (original − blurred) back; a valid convolution leaves (N − M + 1)² pixels.
- Linear stretch: DN_out = 255(DN − min)/(max − min); clip at mean ± kσ or percentiles to avoid wasting range on outliers.
- Non-linear: equalisation s_k = round((L − 1) × cumulative fraction); logarithm and γ < 1 open the dark end, γ > 1 the bright end; enhancement changes appearance, never information.
Practice questions (17)
Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.
Independent random noise of standard deviation 12 DN affects every pixel. After averaging a 4 × 4 window of a uniform area, the noise standard deviation is ____ DN.
Numerical answer — type the value.
Show answer
Answer: 3
Averaging N = 16 independent values divides σ by √16 = 4: 12/4 = 3 DN. The same average also blurs any edge inside the window, which is the price of the noise reduction.Isolated black and white pixels scattered over an image are best removed by a
Show answer
Answer: A — median filter
The median picks the middle value of the neighbourhood, so a lone outlier is discarded. A mean would smear each dot over its neighbours, a Laplacian would highlight it, and a stretch would only make it more visible.A 3 × 3 window contains 12 15 18 / 10 90 14 / 16 13 11. The output of a median filter at the centre is ____.
Numerical answer — type the value.
Show answer
Answer: 14
Sorted: 10, 11, 12, 13, 14, 15, 16, 18, 90. The fifth of nine values is 14. The impulse 90 is at the top of the list and never chosen; a mean filter would return 199/9 = 22.11.The same window (12 15 18 / 10 90 14 / 16 13 11) is passed through a 3 × 3 mean filter. The centre output, to two decimal places, is ____.
Numerical answer — type the value.
Show answer
Answer: 22.11
Row sums are 45, 114 and 40, total 199; 199/9 = 22.11. One impulse of 90 raises the output about 8 units above the median, showing why the mean is a poor choice against salt-and-pepper noise.A 3 × 3 window has 10 10 40 in each of its three rows. The horizontal-gradient Sobel kernel [−1 0 1; −2 0 2; −1 0 1] gives, at the centre, the value ____.
Numerical answer — type the value.
Show answer
Answer: 120
Each row gives −10 + 40 = 30 with weights 1, 2, 1: 30 + 2 × 30 + 30 = 120. The vertical Gy is 0 because the three rows are identical; the gradient magnitude is √(120² + 0²) = 120, a strong vertical edge.A pixel of DN 50 has a 3 × 3 mean of 40. Unsharp masking with sharpened = original + k × (original − blurred) and k = 1.5 gives the value ____.
Numerical answer — type the value.
Show answer
Answer: 65
Detail = 50 − 40 = 10; added back with weight 1.5: 50 + 15 = 65. With k = 0 the image is unchanged; larger k gives crisper edges and more visible noise.A 5 × 5 kernel is applied to a 512 × 512 image with no padding, keeping only positions where the kernel fully overlaps the image. The output has ____ pixels along each side.
Numerical answer — type the value.
Show answer
Answer: 508
Output side = N − M + 1 = 512 − 5 + 1 = 508; two rows and columns are lost on every edge. Padding the border with zeros or mirrored values would keep 512.A 3 × 3 kernel has weights −1 −1 −1 / −1 8 −1 / −1 −1 −1. On a perfectly uniform area of DN 90 its output is
Show answer
Answer: C — 0
The weights sum to 8 − 8 = 0, so a uniform area gives 0 × 90 = 0. It is a high-pass (edge) kernel: only changes in brightness produce a response. 720 would be the centre term alone.A band has min 40 and max 120. A linear min–max stretch to 0–255 turns DN 100 into (nearest integer) ____.
Numerical answer — type the value.
Show answer
Answer: 191
DN_out = 255 × (100 − 40)/(120 − 40) = 255 × 0.75 = 191.25, which rounds to 191. Forgetting to subtract the minimum gives 255 × 100/120 = 212.A band has mean 80 and standard deviation 10. A saturation stretch sends mean − 2σ to 0 and mean + 2σ to 255, linearly between. DN 75 becomes (nearest integer) ____.
Numerical answer — type the value.
Show answer
Answer: 96
Limits are 60 and 100. DN_out = 255 × (75 − 60)/(100 − 60) = 255 × 0.375 = 95.6, so 96. Pixels below 60 saturate to 0 and above 100 to 255, which is why this stretch has more contrast than min–max.A 3-bit image (8 levels) of 64 pixels has 16, 24, 16 and 8 pixels at DN 2, 3, 4 and 5 and none elsewhere. Using s_k = round(7 × cumulative fraction), histogram equalisation maps DN 3 to ____.
Numerical answer — type the value.
Show answer
Answer: 4
The cumulative count up to DN 3 is 16 + 24 = 40, a fraction 0.625. Then 7 × 0.625 = 4.375, which rounds to 4. DN 2 → round(1.75) = 2, DN 4 → round(6.125) = 6, DN 5 → 7.A logarithmic stretch DN_out = 255 × log(1 + DN)/log 256 is applied to DN 3. The output, to the nearest integer, is ____.
Numerical answer — type the value.
Show answer
Answer: 64
log(1 + 3) = log 4 and log 256 = 4 log 4, hence the ratio is exactly 1/4 for any base and DN_out = 255/4 = 63.75, which rounds to 64. A DN of only 3 is lifted to a quarter of the display range, opening the dark end.An image of a lake and dark forest has a histogram crowded at low DNs with a long tail to the right. Which enhancement best shows detail in the dark areas?
Show answer
Answer: C — A logarithmic or γ < 1 stretch
Both curves have a steep slope at small inputs and flatten at large ones, so they spread the crowded dark DNs over many display levels. A γ > 1 curve does the opposite. A negative only swaps light and dark, and a mean filter smooths detail away.Which of the following operations are linear and shift-invariant, so that they can be written as a convolution with a fixed kernel?
Show answer
Answer: A — A 3 × 3 mean filter; B — The Laplacian filter; C — The Sobel Gx filter
The mean, Laplacian and Sobel filters are fixed weighted sums, hence linear and shift-invariant. The median sorts the neighbourhood, and the median of a sum is not the sum of the medians, so it is not linear and no kernel represents it.Which of the following are true of histogram equalisation?
Show answer
Answer: A — It gives more output levels to frequently occurring DNs; B — Its transfer function is the image's own cumulative histogram scaled to the output range; C — It can merge different input DNs into one output DN
The CDF is the mapping, and it is steepest where pixels are numerous, so common DNs get more levels; with rounding, close DNs can land on one level. An exactly flat histogram is impossible for discrete data because all pixels of one input level move together.Which statements about a linear contrast stretch are correct?
Show answer
Answer: A — It changes the appearance but adds no new information to the data; C — Noise is stretched together with the signal; D — Clipping the tails at mean ± 2σ usually gives more contrast in the bulk of the image than a min–max stretch
A stretch is a re-labelling of DNs for display: no information is created, clipping frees range for the majority of pixels, and noise scales with the signal. Removing haze needs a radiometric correction such as dark-object subtraction; a stretch only moves the histogram.Radar images of a uniform crop field look grainy even though the field is homogeneous. The most appropriate description and treatment is
Show answer
Answer: A — multiplicative speckle from coherent interference, reduced by multilook averaging or adaptive filtering
Speckle comes from the coherent addition of waves scattered by many elements inside a resolution cell; it is multiplicative and inherent to coherent imaging. Averaging independent looks or an adaptive filter such as Lee's reduces it. The other three describe different defects.