Motion analysis
Feature Tracking
How distinctive image points are selected and followed across video frames, using the classic KLT tracker as the main example.
intermediate
Feature tracking follows a set of distinctive image points from one video frame to the next. Instead of estimating motion everywhere, it concentrates on points that can be located reliably, and produces tracks: the positions of each point over time. Feature tracks are the input to many higher-level tasks, including visual odometry, video stabilization, structure from motion, and augmented reality.
What Is Feature Tracking?
A feature tracker repeats two steps:
- Select points in a frame that will be easy to find again.
- Track each point into the next frame by searching for the location where its surrounding image patch matches best.
Tracking is closely related to sparse optical flow: each tracked point is an optical flow estimate at a single location, followed consistently over many frames.
Choosing Good Features
Not every point can be tracked. In a uniform region, a patch looks the same wherever it is placed. Along a straight edge, a patch can slide along the edge without changing, which is the aperture problem. Only points with intensity variation in more than one direction, such as corners, can be located in both and .
This is measured with the 2×2 matrix built from the image gradients in a window around the point:
The eigenvalues and of describe how much the intensity varies in the two principal directions:
- both small: a flat region, not trackable;
- one large and one small: an edge, trackable in one direction only;
- both large: a corner-like point, trackable.
The Shi–Tomasi criterion keeps points where exceeds a threshold. Shi and Tomasi argued that features should be chosen by exactly the condition that makes the tracker’s equations well-conditioned. The earlier Harris detector uses a closely related score, , which avoids computing the eigenvalues explicitly.
Tracking Features Between Frames
The classic tracker is KLT (Kanade–Lucas–Tomasi). For each feature, it finds the displacement that best aligns the patch around the point in the previous frame with the next frame, using the Lucas–Kanade method:
- The displacement is refined iteratively by solving a small linear system built from the matrix above.
- An image pyramid handles larger motions: the displacement is estimated at a coarse resolution, then refined at each finer level.
For each feature, the tracker reports the new position and whether tracking succeeded.
Keeping Tracks Healthy
Over a long sequence, tracks degrade. Practical trackers add a few safeguards:
- Forward–backward check. Track each point from frame A to frame B and back again. If it does not return close to where it started, discard it.
- Re-detection. Points are lost to occlusion and to leaving the frame. When the number of active tracks drops too low, detect new features, usually in regions that currently have none.
- Drift. Small errors accumulate as a point is tracked frame after frame. Comparing against the patch from the frame where the feature was first detected, rather than only the previous frame, helps detect drift.
- Geometric outlier rejection. When the scene motion follows a known model, such as a homography or the epipolar geometry, a robust estimator like RANSAC can remove tracks that are inconsistent with it.
Tracking and Matching
Feature tracking assumes small motion between consecutive frames. When images are far apart in time or viewpoint, such as photos taken from different positions, the usual approach is detection and matching instead: detect features independently in each image, describe each one with a descriptor such as SIFT or ORB, and match the descriptors. Matching handles large changes in position, scale, and rotation, but is typically slower and less precise between consecutive video frames than tracking.
Practical Example
KLT tracking with a forward–backward check in OpenCV:
import cv2
import numpy as np
previous = cv2.imread("frame_000.png", cv2.IMREAD_GRAYSCALE)
current = cv2.imread("frame_001.png", cv2.IMREAD_GRAYSCALE)
points = cv2.goodFeaturesToTrack(
previous, maxCorners=200, qualityLevel=0.01, minDistance=7
)
lk_params = dict(winSize=(21, 21), maxLevel=3)
tracked, status, _ = cv2.calcOpticalFlowPyrLK(previous, current, points, None, **lk_params)
# Forward-backward check: track back and keep points that return close to their start.
back, back_status, _ = cv2.calcOpticalFlowPyrLK(current, previous, tracked, None, **lk_params)
distance = np.linalg.norm(points - back, axis=2).ravel()
good = (status.ravel() == 1) & (back_status.ravel() == 1) & (distance < 1.0)
old_points = points[good]
new_points = tracked[good]
cv2.goodFeaturesToTrack applies the Shi–Tomasi criterion, and cv2.calcOpticalFlowPyrLK performs pyramidal Lucas–Kanade tracking.
Limitations
- Small-motion assumption. Fast motion, motion blur, and large frame-to-frame changes cause tracks to be lost, even with image pyramids.
- Appearance changes. Changes in lighting, scale, rotation, and perspective alter the patch around a point and degrade tracking over time.
- Occlusion. Points hidden behind other objects cannot be tracked and may be re-attached to the wrong structure when they reappear.
- Texture dependence. Scenes with few corners, such as blank walls or repetitive patterns, provide few reliable features.
- Sparse output. Tracking describes motion only at selected points. When motion is needed everywhere, dense optical flow is required.
Related
- Optical Flow
How optical flow estimates the apparent motion of image content between frames, and the assumptions behind the classic methods.
References
- Shi, J. & Tomasi, C. (1994). Good Features to Track. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 593–600.
- Tomasi, C. & Kanade, T. (1991). Detection and Tracking of Point Features. Technical Report CMU-CS-91-132, Carnegie Mellon University.
- Lucas, B. D. & Kanade, T. (1981). An Iterative Image Registration Technique with an Application to Stereo Vision. Proceedings of the 7th International Joint Conference on Artificial Intelligence (IJCAI), 674–679.