Motion analysis
Optical Flow
How optical flow estimates the apparent motion of image content between frames, and the assumptions behind the classic methods.
intermediate
Optical flow is the apparent motion of image content between two frames. For each pixel, it estimates a two-dimensional displacement that describes where that part of the image moved. Optical flow is a building block for video analysis, object tracking, video stabilization, action recognition, and visual odometry.
What Is Optical Flow?
Given two frames taken a short time apart, optical flow assigns a vector to each pixel: the horizontal and vertical displacement, in pixels, of the image content at that location.
It is important to distinguish optical flow from the true motion field, the projection of real 3D motion onto the image. The two usually agree, but not always:
- A smooth, uniformly colored sphere rotating in place has a non-zero motion field but produces almost no visible change, so its optical flow is close to zero.
- A static scene under a moving light source has a zero motion field, but changing shading produces non-zero optical flow.
Optical flow measures what changes in the image, not what moves in the world.
The Brightness Constancy Assumption
Classic optical flow methods assume that a point keeps the same intensity as it moves:
For small displacements, a first-order Taylor expansion of the right-hand side gives the optical flow constraint equation:
Here and are the spatial image gradients, and is the change in intensity between the two frames.
The Aperture Problem
The constraint equation is a single equation with two unknowns, and . At any individual pixel it only determines the component of motion along the gradient direction. Motion along an edge produces no change in intensity and cannot be observed locally.
This is the aperture problem: viewed through a small window, a moving straight edge looks the same whatever its motion along the edge. Every optical flow method must add some further assumption to resolve it.
Local methods: Lucas–Kanade
The Lucas–Kanade method assumes that the flow is constant within a small window around each pixel. Each pixel in the window contributes one constraint equation, giving an overdetermined system that is solved by least squares.
The solution depends on the 2×2 matrix built from gradients in the window:
The flow can be estimated reliably only when both eigenvalues of are large, meaning the window contains gradients in more than one direction. This is the same condition used to detect corners, which is why Lucas–Kanade is commonly applied at corner-like points in feature tracking.
Global methods: Horn–Schunck
The Horn–Schunck method instead assumes that the flow field varies smoothly across the image. It minimizes a single energy over the whole image that combines two terms:
- how well the flow satisfies the constraint equation at each pixel, and
- how much the flow changes between neighboring pixels.
The smoothness term propagates information from textured regions into uniform ones, producing a flow vector at every pixel, at the cost of blurring flow across motion boundaries.
Sparse and Dense Flow
- Sparse flow estimates motion only at selected points, usually corners. It is fast and reliable where it applies. Pyramidal Lucas–Kanade is the standard example.
- Dense flow estimates a vector at every pixel. Horn–Schunck and Farnebäck’s polynomial-expansion method are classic examples.
Large Motions and Image Pyramids
The Taylor expansion behind the constraint equation is only valid for small displacements, typically a pixel or two. Real videos often contain much larger motions.
The standard solution is a coarse-to-fine strategy. An image pyramid is built by repeatedly downsampling each frame. Flow is estimated at the coarsest level, where large motions become small, and then used to initialize the estimate at each finer level.
Practical Example
Dense optical flow with Farnebäck’s method in OpenCV:
import cv2
import numpy as np
previous = cv2.imread("frame_000.png", cv2.IMREAD_GRAYSCALE)
current = cv2.imread("frame_001.png", cv2.IMREAD_GRAYSCALE)
flow = cv2.calcOpticalFlowFarneback(
previous, current, None,
pyr_scale=0.5, levels=3, winsize=15,
iterations=3, poly_n=5, poly_sigma=1.2, flags=0,
)
# flow has shape (height, width, 2): the (dx, dy) displacement of each pixel.
magnitude, angle = cv2.cartToPolar(flow[..., 0], flow[..., 1])
A common way to visualize dense flow is to map direction to hue and speed to brightness:
hsv = np.zeros((*previous.shape, 3), dtype=np.uint8)
hsv[..., 0] = angle * 180 / np.pi / 2 # hue: direction of motion
hsv[..., 1] = 255
hsv[..., 2] = cv2.normalize(magnitude, None, 0, 255, cv2.NORM_MINMAX) # brightness: speed
visualization = cv2.cvtColor(hsv, cv2.COLOR_HSV2BGR)
Limitations
- Brightness changes. Changes in illumination, shadows, and specular highlights violate brightness constancy and produce spurious flow.
- Textureless regions. Where there are no gradients, the constraint equation carries no information and flow must be inferred from surrounding regions or not at all.
- Occlusions. Pixels that become hidden or newly visible between frames have no correct correspondence.
- Motion boundaries. Smoothness assumptions blur the flow where objects with different motions meet.
- Large displacements. Fast motion and thin structures can be lost even with coarse-to-fine estimation.
Modern learning-based methods, such as RAFT, estimate flow with deep networks trained on large datasets and handle many of these cases considerably better than the classic methods. The concepts above, including brightness constancy, the aperture problem, and coarse-to-fine estimation, remain the foundation for understanding them.
Related
- Feature Tracking
How distinctive image points are selected and followed across video frames, using the classic KLT tracker as the main example.
References
- Horn, B. K. P. & Schunck, B. G. (1981). Determining Optical Flow. Artificial Intelligence, 17(1–3), 185–203.
- Lucas, B. D. & Kanade, T. (1981). An Iterative Image Registration Technique with an Application to Stereo Vision. Proceedings of the 7th International Joint Conference on Artificial Intelligence (IJCAI), 674–679.
- Farnebäck, G. (2003). Two-Frame Motion Estimation Based on Polynomial Expansion. Scandinavian Conference on Image Analysis (SCIA), 363–370.