Motion & Tracking

Motion Estimation

Recovering how image content, objects, or the camera moved from a sequence of images, from per-pixel flow to global and 3D motion.

intermediate

Motion estimation is the problem of working out, from two or more images taken at different times, how things moved between them. Depending on the application, “things” means the image content at every pixel, a set of tracked points, a whole frame, the camera, or the 3D points of the scene. The answer is a motion representation: a flow field, a set of displacement vectors, a parametric transformation, a camera trajectory, or a 3D motion field. Motion estimation underlies video compression, stabilization, visual odometry, and many video understanding systems.

This article is a hub: it describes the variants of the problem and how they relate, and points to the articles that treat each one in depth.

Problem Definition

Given images I1,…,ITI_1, \dots, I_T of a scene, estimate the motion that relates them. In the most common two-frame form, the goal is a mapping from positions in I1I_1 to corresponding positions in I2I_2:

x↦x+w(x;θ)\mathbf{x} \mapsto \mathbf{x} + \mathbf{w}(\mathbf{x}; \boldsymbol{\theta})

The variants differ in what w\mathbf{w} is allowed to be: a free vector at every pixel (dense flow), vectors at selected points only (sparse motion), one vector per block of pixels (block matching), or a transformation with a few global parameters θ\boldsymbol{\theta} (parametric motion). With calibrated cameras, the motion can instead be expressed in 3D, as the camera’s rotation and translation or as the 3D displacement of scene points.

What images measure is the movement of brightness patterns, which need not equal the projected 3D motion; optical flow explains where the two differ.

Input and Output

  • Input: two or more frames of a video or image sequence, sometimes with camera calibration, stereo pairs, or depth.
  • Output: one of the motion representations below, often with a confidence or validity mask for pixels, such as occluded ones, where no reliable estimate exists.

Variants

Dense optical flow

A 2D displacement vector for every pixel: the most general 2D representation and the most studied form of the problem. See optical flow.

Sparse feature motion

Displacements are estimated only at distinctive points, such as corners, and followed over many frames to form tracks. This is cheaper and more reliable than dense flow where it applies, and it is the basis of visual odometry and structure from motion. See feature tracking.

Block matching

The image is divided into fixed blocks, and each block is assigned the displacement that best matches it to a region of a reference frame, typically by minimizing the sum of absolute or squared differences over a search window. Jain and Jain (1981) applied displacement estimation of this kind to interframe image coding, and block-based motion compensation remains central to video codecs: the encoder transmits motion vectors and a prediction residual instead of the full frame. The goal is compression efficiency, so the chosen vectors need not match physical movement.

Parametric and global motion

All pixels of an image or region share a low-dimensional model: translation (2 parameters), similarity (4), affine (6), or homography (8). A homography exactly describes the motion of a planar scene, or of any scene viewed by a camera that only rotates about its center. Bergen et al. (1992) described a hierarchical framework that estimates such models directly from image intensities. Global motion is used in video stabilization, panorama stitching, and camera motion compensation.

Camera motion (ego-motion)

Estimating the camera’s own rotation and translation. With a calibrated camera, corresponding points in two views constrain the relative pose through epipolar geometry, up to an unknown scale for a single camera. Over a sequence this becomes visual odometry, and with map building and loop closure, visual SLAM.

3D scene flow

The 3D motion of every visible scene point, introduced as scene flow by Vedula et al. (1999); optical flow is its projection into the image. Scene flow requires depth, from stereo, depth sensors, or LiDAR, and is used in autonomous driving to separate independently moving objects from the static world.

At the edges of the task, object tracking follows whole objects and their identities rather than pixel-level motion, and long-range point tracking follows individual points through many frames, including through occlusion.

Common Assumptions

Image data alone does not determine motion uniquely, so every variant adds assumptions:

  • Brightness constancy: a point keeps its appearance as it moves, which links motion to the image gradients. One such constraint per pixel cannot fix a 2D vector, which is the aperture problem.
  • Small motion: linearized methods hold only for displacements of about a pixel, so most methods work coarse to fine over an image pyramid.
  • Spatial coherence: neighboring pixels move similarly, whether as constant motion in a window, global or piecewise smoothness, or a parametric model.
  • Rigidity: for camera motion and scene flow, the scene or each object moves as a rigid body.

Challenges

  • Large displacements: fast motion, especially of small objects, exceeds the range of local and coarse-to-fine methods.
  • Occlusion: pixels that disappear or appear between frames have no correspondence.
  • Textureless and repetitive regions: many displacements fit the data equally well.
  • Motion boundaries: motion is discontinuous where surfaces meet, and smoothness assumptions blur it.
  • Appearance changes: lighting, shadows, motion blur, and noise violate brightness constancy.
  • Independent motion: for camera motion, moving objects are outliers that must be rejected.

Approaches

Differential methods solve the linearized brightness constancy constraint with an extra assumption: constant flow in a window (Lucas–Kanade) or global smoothness (Horn & Schunck, 1981). Coarse-to-fine pyramids extended them to larger motions.

Matching methods compare patches directly over a search range. Block matching is the simplest form; phase correlation estimates a global translation in the frequency domain. Matching handles larger motions than linearization but is coarse and ambiguous in weak texture.

Feature-based methods detect and track or match distinctive points, then fit a parametric or geometric model with a robust estimator such as RANSAC. This is the standard route to homographies and camera motion.

Variational methods minimize a data term plus a regularizer. Robust penalties that preserve motion boundaries, better data terms, and occlusion reasoning made them the leading dense flow methods on benchmarks such as Middlebury (Baker et al., 2011) until the mid-2010s.

Learning-based methods train networks on large synthetic datasets with exact ground truth, beginning with FlowNet (Dosovitskiy et al., 2015). Later designs built classical ideas into the network: pyramids, warping, and cost volumes in PWC-Net (Sun et al., 2018), and iterative refinement over all-pairs correlations in RAFT (Teed & Deng, 2020). Networks took over because hand-designed energies struggled with large motions, occlusions, and weak texture and were slow to optimize, while learned models absorb priors from data and run fast on GPUs. Self-supervised variants train on real video with a photometric (brightness constancy) loss instead of ground truth.

Datasets and Benchmarks

Ground-truth motion is hard to measure for real scenes, so benchmarks rely on controlled capture, laser scanning, or computer graphics:

  • Middlebury: small, high-accuracy sequences with hidden fluorescent texture, plus synthetic scenes (Baker et al., 2011).
  • MPI Sintel: frames rendered from an open-source animated film, with large motions, motion blur, and occlusions.
  • KITTI 2012 and 2015: real driving scenes with laser-scan ground truth for flow and stereo, scene flow (2015), and a separate visual odometry benchmark.
  • FlyingChairs and FlyingThings3D: large synthetic training sets for learned flow and scene flow.
  • TUM RGB-D and KITTI odometry: camera trajectories for evaluating ego-motion.

Evaluation Metrics

Dense and sparse 2D motion is measured by endpoint error, the pixel distance between estimated and true displacement vectors, and by outlier rates built on it; the endpoint error article also covers the older angular error. Camera motion is evaluated with trajectory errors, such as the absolute trajectory error and relative pose error used by the TUM RGB-D benchmark. Video coding judges motion estimation indirectly, by rate–distortion performance.

Applications

  • Video compression: block motion vectors predict each frame from its neighbors.
  • Video stabilization and panoramas: global motion is estimated and compensated or composited.
  • Visual odometry, SLAM, and structure from motion: camera motion and scene structure from tracked features.
  • Autonomous driving and robotics: scene flow and ego-motion to detect moving obstacles.
  • Video understanding: optical flow as an input for action recognition and video segmentation.
  • Frame interpolation and video editing: synthesizing in-between frames and propagating edits along motion.
  • Medical imaging: registering images over time, for example to measure heart motion.

Open Problems

  • Generalization: networks trained on synthetic data can degrade on real footage with different motion statistics, noise, and blur.
  • Occlusions and long-range correspondence: estimating where hidden points went, and keeping correspondences over many frames rather than two.
  • Small, fast, and non-rigid motion: thin structures, fast objects, and deforming surfaces remain the main sources of error.
  • Efficiency: accurate dense methods are expensive at high resolution and on embedded hardware.
  • Ground truth for real scenes: benchmarks are limited by what can be measured, which biases evaluation toward rigid scenes and rendered data.

Related

  • Optical Flow

    The apparent motion of image content between two frames, represented as a two-dimensional displacement at every pixel.

  • Feature Tracking

    How distinctive image points are selected and followed across video frames, using the classic KLT tracker as the main example.

  • Endpoint Error

    The standard accuracy measure for optical flow, the distance in pixels between an estimated flow vector and the true one, averaged over the image.

  • Brightness Constancy Assumption

    The assumption that a scene point keeps the same image intensity as it moves between frames, which turns motion estimation into an intensity-matching problem.

  • Object Tracking

    Estimating the position, extent, or state of one or more objects in every frame of a video, keeping each object's identity over time.

  • Optical Flow Estimation

    Computing a dense field of pixel displacements between two video frames, from classical variational methods to learned networks such as RAFT.

References

  1. Horn, B. K. P. & Schunck, B. G. (1981). Determining Optical Flow. Artificial Intelligence, 17(1–3), 185–203.
  2. Jain, J. R. & Jain, A. K. (1981). Displacement Measurement and Its Application in Interframe Image Coding. IEEE Transactions on Communications, 29(12), 1799–1808.
  3. Bergen, J. R., Anandan, P., Hanna, K. J. & Hingorani, R. (1992). Hierarchical Model-Based Motion Estimation. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 237–252.
  4. Vedula, S., Baker, S., Rander, P., Collins, R. & Kanade, T. (1999). Three-Dimensional Scene Flow. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 722–729.
  5. Baker, S., Scharstein, D., Lewis, J. P., Roth, S., Black, M. J. & Szeliski, R. (2011). A Database and Evaluation Methodology for Optical Flow. International Journal of Computer Vision, 92(1), 1–31.
  6. Dosovitskiy, A., Fischer, P., Ilg, E., Häusser, P., Hazırbaş, C., Golkov, V., van der Smagt, P., Cremers, D. & Brox, T. (2015). FlowNet: Learning Optical Flow with Convolutional Networks. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2758–2766.
  7. Sun, D., Yang, X., Liu, M.-Y. & Kautz, J. (2018). PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 8934–8943.
  8. Teed, Z. & Deng, J. (2020). RAFT: Recurrent All-Pairs Field Transforms for Optical Flow. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 402–419.