GIS II: Basic Spatial Analysis — Interpolation, Buffer, Overlay, Terrain Modelling and Network Analysis

The last item under Part A's GIS heading names five operations, and this chapter takes them in the syllabus's order: “Basic Spatial analysis: Interpolation, Buffer, Overlay, Terrain Modeling and Network analysis”. Interpolation estimates values between sample points — Thiessen polygons, inverse distance weighting, splines, trend surfaces and kriging with its semivariogram. Buffering builds zones at a distance from features. Overlay combines layers, by vector operations (union, intersect, identity, clip, erase) or by raster map algebra (local, focal, zonal and global). Terrain modelling derives slope, aspect, hillshade, viewsheds, drainage and volumes from a DEM or TIN. Network analysis finds shortest paths, service areas and allocations over a graph with costs. Each has a short computation — an IDW estimate, a buffer area, a slope angle, a shortest path — that GATE sets as a numerical.

1. Interpolation

Spatial interpolation estimates a value at an unsampled location from nearby samples, on Tobler's principle that near things are more alike than distant ones. An exact interpolator returns the sample values at the sample points; an inexact (smoothing) one need not. Global methods use all points (trend surfaces); local methods use a neighbourhood.

The methods
MethodHowCharacter
Thiessen (Voronoi) polygonseach location takes the value of its nearest sampleexact, stepped; used for areal rainfall
Inverse distance weighting (IDW)z = Σ(zᵢ/dᵢᵖ)/Σ(1/dᵢᵖ), usually p = 2exact; estimates never exceed the sample range; “bull's-eyes” around samples
Splinea smooth surface of minimum curvature through the pointsexact; can overshoot beyond the data range
Trend surfaceleast-squares polynomial z = f(x, y) over all pointsglobal, inexact; shows regional trend
Krigingweights from a fitted semivariogram γ(h), which rises from the nugget to the sill at the rangegeostatistical; best linear unbiased; gives an estimation variance with each value

Worked IDW (p = 2): samples 10, 20 and 30 at distances 1, 2 and 4 from the target. Weights 1, 1/4 and 1/16 sum to 1.3125; Σw·z = 10 + 5 + 1.875 = 16.875; z = 16.875/1.3125 = 12.86. With p = 1 the weights are 1, 1/2, 1/4 and z = 27.5/1.75 = 15.71 — a smaller power lets distant points count more.

2. Buffer and overlay

A buffer is the zone within a given distance of a feature: a circle around a point, a corridor with rounded ends around a line, a ring around (or inside) a polygon. Buffers may be fixed-width or variable (width from an attribute, such as a wider setback for a larger river), multiple-ring, and dissolved (overlaps merged) or not. The area of a buffer of width d on both sides of a straight line of length L is 2dL + πd²: for L = 1000 m and d = 50 m it is 100 000 + 7854 = 107 854 m².

Vector overlay operations on an input layer and an overlay layer
OperationOutput extentAttributes
Union (OR)all of both layersfrom both
Intersect (AND)only where both existfrom both
Identitythe input layer's extentfrom both, where they overlap
Clipinput features inside the overlay boundaryinput only (a cookie cutter)
Eraseinput features outside the overlayinput only

Raster overlay is map algebra: local operations work cell by cell across layers (sum, product, Boolean AND/OR, reclassify); focal (neighbourhood) operations compute from a moving window (a 3 × 3 mean or maximum); zonal operations summarise one layer within the zones of another (mean elevation per watershed); global operations use the whole grid (Euclidean distance to the nearest road). Suitability analysis reclassifies criteria to a common scale and combines them by Boolean AND or by a weighted overlay Σwᵢsᵢ. Vector overlay of boundaries digitised separately creates sliver polygons that must be removed with a tolerance.

3. Terrain modelling

Terrain is represented by a raster DEM (a grid of elevations), a TIN, or contours. From a DEM, slope at a cell comes from the elevation gradients: with ∂z/∂x and ∂z/∂y estimated by finite differences across the 3 × 3 neighbourhood (for example (z_east − z_west)/(2 × cell size)), slope = arctan √((∂z/∂x)² + (∂z/∂y)²), in degrees, or 100 × the rise in per cent. Aspect is the compass direction the slope faces — the direction of steepest descent. Worked: 10 m cells, z_east − z_west = 6 m and z_north − z_south = 8 m. ∂z/∂x = 6/20 = 0.3, ∂z/∂y = 8/20 = 0.4, gradient 0.5, slope = arctan 0.5 = 26.57° (50 %). The ground rises to the east and north, so it faces south-west: aspect = 180° + arctan(0.3/0.4) = 216.9°.

  • Hillshade simulates illumination from a chosen sun azimuth and altitude for relief display.
  • Viewshed (visibility): the cells visible from an observer point, found by line-of-sight tests; used for tower siting.
  • Hydrological modelling: fill sinks, compute flow direction (D8: each cell drains to the steepest of its eight neighbours), flow accumulation, then stream networks and watershed boundaries.
  • Cut and fill volume: Σ(z_existing − z_design) × cell area over the cells; profiles along a line; contours threaded from the grid.
⚠️ Degrees and per cent are not the same number
A 45° slope is 100 %, not 50 %; a 50 % slope is 26.6°. And slope from a coarse DEM is systematically gentler than from a fine one, because the larger cell averages out the steep parts.

4. Network analysis

A network is a graph of nodes (junctions) and edges (links) with an impedance (cost: length, travel time, toll) on each edge, plus turn restrictions and one-way rules. It needs topology — connectivity — to work, which is why an undershoot at a junction breaks a route. The core operations: shortest (least-cost) path between two nodes, solved by Dijkstra's algorithm (repeatedly fix the unvisited node with the smallest tentative cost and relax its neighbours; it requires non-negative costs); the travelling salesman problem (best order to visit many stops); service areas (all edges reachable within a cost, such as 10 minutes from a fire station); closest facility; and location–allocation (where to put facilities and which demand each serves).

Worked Dijkstra from A: edges A–B 4, A–C 2, C–B 1, B–D 5, C–D 8
StepNode fixedTentative costs after relaxing
1A (0)B 4, C 2
2C (2)B min(4, 2 + 1) = 3, D 2 + 8 = 10
3B (3)D min(10, 3 + 5) = 8
4D (8)path A–C–B–D, cost 8
🧠 The direct edge is often not the shortest
A–B costs 4 directly but 3 through C, and C–D costs 8 directly but 6 through B. GATE's shortest-path items are built around exactly such detours, so relax every edge rather than guessing.

Key takeaways

  • Interpolation: Thiessen (nearest), IDW Σ(z/dᵖ)/Σ(1/dᵖ) (exact, bounded), spline (smooth, can overshoot), trend surface (global, inexact), kriging (semivariogram: nugget, sill, range).
  • A line buffer of width d both sides has area 2dL + πd²; buffers can be fixed, variable, multi-ring, dissolved.
  • Vector overlay: union, intersect, identity, clip, erase; raster map algebra: local, focal, zonal, global; watch for sliver polygons.
  • Slope = arctan √(zₓ² + z_y²) (45° = 100 %); aspect is the downslope direction; hillshade, viewshed, D8 flow direction and cut–fill come from the DEM.
  • Network analysis needs connectivity and impedances; Dijkstra finds least-cost paths with non-negative costs; service areas, closest facility and location–allocation build on it.

Practice questions (14)

Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.

  1. Three samples with values 10, 20 and 30 lie at distances 1, 2 and 4 units from a target point. The IDW estimate with power p = 2 (to two decimal places) is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 12.86

    Weights 1/d² = 1, 0.25, 0.0625; Σw = 1.3125. Σw·z = 10 + 5 + 1.875 = 16.875. z = 16.875/1.3125 = 12.857, i.e. 12.86. The simple mean, 20, ignores distance; with p = 1 the estimate would be 15.71.
  2. For the same samples (10, 20, 30 at distances 1, 2, 4), the IDW estimate with power p = 1 (to two decimal places) is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 15.71

    Weights 1/d = 1, 0.5, 0.25; Σw = 1.75. Σw·z = 10 + 10 + 7.5 = 27.5. z = 27.5/1.75 = 15.714, i.e. 15.71. It is larger than the p = 2 value (12.86) because a lower power gives the distant high values more weight.
  3. In a semivariogram, the separation distance beyond which the semivariance stops increasing and samples are no longer spatially correlated is the

    1. range
    2. sill
    3. nugget
    4. lag
    Show answer

    Answer: A — range

    The range is the distance at which the curve levels off; the sill is the semivariance value it levels off at; the nugget is the value extrapolated to zero distance (measurement error and micro-scale variation); a lag is any separation interval used to bin the pairs.
  4. A straight pipeline 1000 m long is buffered by 50 m on both sides, with rounded ends. The buffer area, in m² (to the nearest square metre), is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 107854

    The corridor is a rectangle 1000 m × 100 m = 100,000 m² plus two half-circles of radius 50 m, i.e. one full circle π × 50² = 7853.98 m². Total = 107,853.98 ≈ 107,854 m². Leaving out the rounded ends gives 100 000 m².
  5. A land-use layer is overlaid on a district boundary so that the output keeps only land-use polygons inside the district and carries only land-use attributes. The operation is

    1. clip
    2. union
    3. identity
    4. erase
    Show answer

    Answer: A — clip

    Clip is a cookie cutter: the overlay boundary cuts the input and none of its attributes are added. Intersect would add the district's attributes; union keeps everything from both; erase keeps what lies outside.
  6. Computing, for each cell, the mean of the 3 × 3 window centred on it is an example of which kind of map-algebra operation?

    1. Focal (neighbourhood)
    2. Local
    3. Zonal
    4. Global
    Show answer

    Answer: A — Focal (neighbourhood)

    A focal operation uses a moving neighbourhood around each cell. Local operations use only the same cell in each layer, zonal operations use the cells of a zone defined by another layer, and global operations can depend on every cell.
  7. On a DEM with 10 m cells, the eastern neighbour of a cell is 6 m higher than the western one and the northern neighbour is 8 m higher than the southern one. The slope at the cell, in degrees (to two decimal places), is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 26.57

    ∂z/∂x = 6/(2 × 10) = 0.3 and ∂z/∂y = 8/20 = 0.4. Gradient = √(0.09 + 0.16) = 0.5, slope = arctan 0.5 = 26.57° (50 %). Dividing by 10 instead of 20 doubles both gradients and gives 45°.
  8. For the same cell (ground rising 0.3 m/m to the east and 0.4 m/m to the north), the aspect measured clockwise from north as the direction of steepest descent, in degrees (to one decimal place), is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 216.9

    Steepest descent points opposite the gradient: toward (east −0.3, north −0.4), i.e. south-west. Its bearing is 180° + arctan(0.3/0.4) = 180° + 36.87° = 216.9°. Reporting the uphill direction, 36.9°, is the usual slip.
  9. A slope of 45° expressed as a percentage is

    1. 100 %
    2. 50 %
    3. 45 %
    4. 70.7 %
    Show answer

    Answer: A — 100 %

    Per cent slope = 100 × rise/run = 100 × tan θ = 100 × tan 45° = 100 %. The value 50 % would be 45/90, which confuses an angle with a fraction of a right angle; 70.7 % is 100 sin 45°.
  10. In the D8 flow-direction method on a DEM, each cell drains to

    1. the one of its eight neighbours with the steepest downhill drop
    2. all lower neighbours in proportion to their drop
    3. the nearest stream cell regardless of elevation
    4. its four edge-sharing neighbours equally
    Show answer

    Answer: A — the one of its eight neighbours with the steepest downhill drop

    D8 sends all flow to a single neighbour, the one with the largest drop divided by distance (diagonal distances are √2 cells). Splitting flow among lower neighbours is the idea of multiple-flow-direction methods, not D8.
  11. A road network has edges with costs A–B 4, A–C 2, C–B 1, B–D 5 and C–D 8 (all two-way). The least cost from A to D is ____.

    Numerical answer — type the value.

    Show answer

    Answer: 8

    Dijkstra from A: fix C at 2, which improves B to 2 + 1 = 3 and sets D to 10; fix B at 3, which improves D to 3 + 5 = 8. The path is A–C–B–D at cost 8. The direct-looking A–B–D costs 9 and A–C–D costs 10.
  12. Which of the following are requirements or properties of network analysis in a GIS?

    1. The network must be topologically connected at its junctions
    2. Each edge carries an impedance such as length or travel time
    3. Dijkstra's algorithm assumes non-negative edge costs
    4. A service area is the single shortest route between two facilities
    Show answer

    Answer: A — The network must be topologically connected at its junctions; B — Each edge carries an impedance such as length or travel time; C — Dijkstra's algorithm assumes non-negative edge costs

    Routing follows connectivity, so a broken junction breaks the route; costs on edges are what is minimised; and Dijkstra's greedy fixing of the smallest tentative cost is only valid when no edge can reduce a cost. A service area is the set of edges reachable within a cost limit, not a single route.
  13. Which of the following statements about spatial interpolation are correct?

    1. IDW estimates always lie between the minimum and maximum sample values
    2. Kriging provides an estimate of the error variance at each interpolated point
    3. A spline surface can go above the largest sample value
    4. A first-order trend surface passes exactly through every sample point
    Show answer

    Answer: A — IDW estimates always lie between the minimum and maximum sample values; B — Kriging provides an estimate of the error variance at each interpolated point; C — A spline surface can go above the largest sample value

    IDW is a weighted average with positive weights, so it cannot leave the data range; kriging's variance comes from the semivariogram model; a minimum-curvature spline can overshoot between points. A trend surface is a least-squares fit and generally misses the points — it is inexact.
  14. Which of the following can be derived from a DEM alone, without any other layer?

    1. A slope map
    2. The viewshed from a proposed tower location
    3. Watershed boundaries from flow direction and accumulation
    4. Land-use classes
    Show answer

    Answer: A — A slope map; B — The viewshed from a proposed tower location; C — Watershed boundaries from flow direction and accumulation

    Slope comes from the elevation gradients, a viewshed from line-of-sight tests over the elevation surface, and watersheds from D8 flow direction and accumulation — all from elevations only. Land use is not a property of elevation; it needs imagery or survey.