Image processing
Image Gradients
How image gradients measure the direction and strength of intensity change, and why they underpin edges, corners, and motion estimation.
beginner
An image gradient describes how quickly, and in which direction, pixel intensity changes at each point in an image. Gradients are one of the most basic measurements in computer vision: edges, corners, texture, and motion are all detected by looking at how intensity changes rather than at intensity itself.
What Is an Image Gradient?
Treat a grayscale image as a function that gives the intensity at each pixel. The gradient at a point is the vector of its two partial derivatives:
- measures how intensity changes moving horizontally.
- measures how intensity changes moving vertically.
From these two components we get two useful quantities:
The magnitude is large where intensity changes sharply, such as at the boundary between a dark object and a bright background. The orientation points in the direction of steepest increase in intensity, which is perpendicular to the edge itself.
In flat, uniform regions both derivatives are close to zero, so the gradient carries almost no information there.
Why Do Gradients Matter?
Most of the structure that vision algorithms rely on is located where intensity changes:
- Edges are points of high gradient magnitude. Edge detectors such as Canny are built on gradient magnitude and orientation.
- Corners are points where the gradient is strong in more than one direction. Corner detectors such as Harris and Shi–Tomasi analyze the distribution of gradient orientations in a small window.
- Motion can be estimated from how gradients relate to changes over time. Optical flow and feature tracking both depend directly on spatial gradients.
- Descriptors such as HOG and SIFT summarize local image appearance as histograms of gradient orientations.
Gradients are also largely unaffected by adding a constant to every pixel, which makes them more robust to uniform brightness changes than raw intensities.
Computing Gradients
Images are discrete, so derivatives are approximated with small convolution kernels.
Finite differences
The simplest approximation is the central difference, which compares the pixels on either side of the current one:
As a convolution kernel this is (up to scale). It is cheap but very sensitive to noise, because differentiation amplifies high-frequency variation.
Sobel
The Sobel operator combines a central difference in one direction with smoothing in the perpendicular direction. The horizontal kernel is:
The vertical kernel is its transpose. The smoothing makes Sobel considerably more robust to noise than a plain finite difference, which is why it is the most common default.
Scharr
For 3×3 kernels, the Scharr operator gives a more accurate estimate of gradient orientation than Sobel:
Prefer Scharr when orientation accuracy matters, for example when building orientation histograms.
Smoothing first
Noise produces large spurious derivatives. A common approach is to smooth the image with a Gaussian before differentiating, or equivalently to convolve with the derivative of a Gaussian. The width of the Gaussian sets the scale of the structures the gradient responds to: a small width preserves fine detail, while a larger width suppresses noise and small texture.
Practical Example
Computing gradient magnitude and orientation with OpenCV:
import cv2
image = cv2.imread("input.png", cv2.IMREAD_GRAYSCALE)
image = cv2.GaussianBlur(image, (5, 5), 1.0)
# Use a floating-point output type: derivatives can be negative.
gx = cv2.Sobel(image, cv2.CV_32F, 1, 0, ksize=3)
gy = cv2.Sobel(image, cv2.CV_32F, 0, 1, ksize=3)
magnitude, orientation = cv2.cartToPolar(gx, gy, angleInDegrees=True)
A frequent mistake is to request an 8-bit unsigned output (cv2.CV_8U). Negative derivatives are then clipped to zero, so dark-to-bright and bright-to-dark transitions are no longer treated symmetrically.
Limitations
- Noise sensitivity. Differentiation amplifies noise. Some smoothing is almost always required, at the cost of blurring fine detail.
- Scale dependence. A gradient computed at one scale may miss structures that are only visible at another. Multi-scale approaches address this.
- No information in flat regions. Uniform or textureless areas produce near-zero gradients, which limits any method that depends on them.
- Discrete approximation. Small kernels only approximate true derivatives, and their accuracy depends on the direction of the edge.
- Saturation. In over- or under-exposed regions intensity is clipped, and the real gradient cannot be recovered.
References
- Canny, J. (1986). A Computational Approach to Edge Detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 8(6), 679–698.
- Harris, C. & Stephens, M. (1988). A Combined Corner and Edge Detector. Proceedings of the Alvey Vision Conference.