Motion & Tracking

Optical Flow Estimation

Computing a dense field of pixel displacements between two video frames, from classical variational methods to learned networks such as RAFT.

intermediate

Optical flow estimation is the task of computing, from two frames of a video, where each pixel of the first frame moved in the second. The result is a dense field of two-dimensional displacement vectors, the optical flow. It is the most studied form of motion estimation, and many video applications build on it.

Problem Definition

Given two frames I0I_0 and I1I_1, estimate a displacement field w(x)=(u,v)\mathbf{w}(\mathbf{x}) = (u, v) such that the content at pixel x\mathbf{x} in I0I_0 appears at x+w(x)\mathbf{x} + \mathbf{w}(\mathbf{x}) in I1I_1. Most methods start from brightness constancy, I1(x+w)=I0(x)I_1(\mathbf{x} + \mathbf{w}) = I_0(\mathbf{x}), but this gives one equation for two unknowns at each pixel (the aperture problem) and is ambiguous in uniform regions. Every method therefore adds a prior: constant flow in a window, smoothness across the image, or regularities learned from data.

Three details make the definition precise:

  • Direction. Forward flow maps I0I_0 to I1I_1 and is defined on the grid of I0I_0; backward flow maps I1I_1 to I0I_0. The two are not simple negatives of each other, because they are sampled on different grids and differ at occlusions.
  • Occlusion. Pixels of I0I_0 that are hidden in I1I_1 have no true correspondence. Methods either extrapolate a plausible vector from visible neighbors or mark the pixels as invalid. A common test is the forward–backward consistency check: following the forward flow and then the backward flow should return to the starting pixel, and pixels where it does not are flagged as occluded or unreliable.
  • Units. Flow is measured in pixels per frame, so it scales with resolution and frame rate.

Input and Output

  • Input: two frames, usually consecutive, in grayscale or color. Multi-frame methods take a longer sequence.
  • Output: an H×W×2H \times W \times 2 array of displacements, sometimes with an occlusion mask or a per-pixel confidence.

Variants

  • Dense and sparse flow. Dense flow assigns a vector to every pixel. Sparse flow estimates motion only at selected, well-textured points and is the basis of feature tracking.
  • Two-frame and multi-frame flow. Most benchmarks evaluate two-frame flow. Multi-frame methods exploit temporal smoothness and can reason about occlusions using a third frame in which a hidden point is visible.
  • Scene flow. The 3D motion of every scene point, estimated from stereo video or depth. Optical flow is its projection into the image.
  • Supervised and unsupervised flow. Networks are trained either on ground-truth flow, mostly synthetic, or on real video with a photometric loss derived from brightness constancy.

Challenges

  • Large displacements. Linearized data terms are valid only for small motions, and fast motion of small objects is particularly hard.
  • Textureless and repetitive regions. Many displacements fit the data equally well.
  • Motion boundaries. Flow is discontinuous where surfaces meet, and smoothness priors blur it.
  • Illumination and appearance changes. Shadows, highlights, exposure changes, motion blur, and noise violate brightness constancy.

Approaches

Local methods

The Lucas–Kanade method (1981) assumes the flow is constant in a small window around each pixel and solves the constraints from that window by least squares. It is fast and accurate at corners and in texture but ill-conditioned along edges and in uniform regions, so it is most reliable at selected points, which is why it became the standard for feature tracking.

Global variational methods

Horn and Schunck (1981) estimated the whole field at once by minimizing an energy with a brightness-constancy data term and a quadratic smoothness term, so that reliable estimates propagate into regions without texture. Their quadratic penalties over-smooth motion boundaries and are dominated by outliers. Later work replaced them with robust penalties and with total-variation regularization and an L1L^1 data term (TV-L1), which tolerate outliers and allow sharp discontinuities. Brox et al. (2004) combined brightness and gradient constancy with robust penalties and minimized the energy without linearizing the data term, which made variational methods considerably more accurate.

Coarse-to-fine warping

Because linearization limits each step to small motions, variational and local methods are run on an image pyramid: flow is estimated at low resolution, upsampled, used to warp the second frame, and refined at the next finer level. Brox et al. showed that this scheme follows from minimizing the non-linearized energy. See coarse-to-fine estimation.

Matching-based large-displacement methods

Coarse-to-fine schemes miss small structures that move further than their own size, because those structures vanish at the coarse levels. Brox and Malik (2011) added descriptor matches as an extra term in the variational energy, so that long-range correspondences found by matching guide the dense solution. DeepFlow (Weinzaepfel et al., 2013) followed the same route with a dedicated matching algorithm, DeepMatching. Both still relied on a coarse-to-fine scheme. EpicFlow (Revaud et al., 2015) replaced it: it interpolates a sparse set of matches into a dense field using an edge-aware geodesic distance and refines that field with a one-level variational minimization.

Learned methods

FlowNet (Dosovitskiy et al., 2015) showed that a convolutional network can be trained end to end to predict flow. Its FlowNetC variant introduced a correlation layer that compares features of the two frames, and the authors created the synthetic Flying Chairs dataset because existing ground truth was too small for training. FlowNet 2.0 (Ilg et al., 2017) stacked several such networks, warping the second image by the intermediate flow between stages. PWC-Net (Sun et al., 2018) built the classical principles into a compact network: a learned feature pyramid, warping of second-frame features by the current flow, and a cost volume with a small search range at each level. RAFT (Teed and Deng, 2020) dropped the coarse-to-fine cascade. It computes correlations between all pairs of pixels at one-eighth resolution, pools them into a multi-scale correlation pyramid, and updates a single flow field with a recurrent unit over many iterations. Its authors reported a 30% lower endpoint error than the best published result on Sintel’s final pass, and many later networks adopted its iterative design.

Networks overtook variational methods largely because hand-designed energies were slow to optimize and struggled with large motions, occlusions, and weak texture, while learned models absorb such priors from data and run quickly on GPUs.

Datasets and Benchmarks

Real-scene ground truth is hard to measure, so benchmarks use controlled capture, laser scanning, or rendering.

  • Middlebury (Baker et al., 2011) provides a small set of high-accuracy sequences, some captured with hidden fluorescent texture, and was the main benchmark of the late 2000s.
  • MPI Sintel (Butler et al., 2012) renders frames from an open-source animated film, with long sequences, large motions, motion blur, and atmospheric effects. A clean pass omits the blur and atmospheric effects that the final pass adds, and its larger motions made errors far higher than on Middlebury.
  • KITTI 2012 and KITTI 2015 (Geiger et al., 2012; Menze and Geiger, 2015) contain real driving scenes with semi-dense ground truth from a laser scanner. The 2012 scenes are static, so all motion comes from the camera; the 2015 version adds independently moving cars, labeled by fitting 3D models to them.
  • FlyingChairs and FlyingThings3D (Dosovitskiy et al., 2015; Mayer et al., 2016) are large synthetic training sets of randomly moving objects. They are not realistic, but networks pretrained on them generalize surprisingly well, and the usual training schedule starts with them before fine-tuning on Sintel or KITTI.

Evaluation Metrics

Accuracy is measured with endpoint error, the Euclidean distance between estimated and true flow vectors, averaged over pixels with valid ground truth. Sintel reports it over all pixels and separately over matched pixels, visible in both frames, and unmatched ones. KITTI 2015 ranks methods by Fl-all, the percentage of pixels whose endpoint error exceeds both 3 pixels and 5% of the true motion. Runtime and memory matter for practical use and are usually reported alongside accuracy.

Applications

  • Video processing: stabilization, frame interpolation and slow motion, video editing that propagates changes along motion, and temporal consistency in video enhancement.
  • Video understanding: flow as an input stream for action recognition and as a cue for motion segmentation.
  • Robotics and driving: obstacle detection, visual odometry, and scene flow for separating moving objects from the static world.
  • Scientific and medical imaging: fluid velocimetry, cell motion, and tissue deformation.
  • Training signals: flow supervises or regularizes other tasks, such as video depth estimation and video generation.

Open Problems

  • Generalization: networks trained on synthetic data can degrade on real footage with different noise, blur, and motion statistics.
  • Occlusions: estimating motion for pixels with no visible match remains guesswork.
  • Small, fast, and thin objects: they remain difficult, even for networks without coarse-to-fine pyramids.
  • High resolution and efficiency: all-pairs correlation grows quadratically with the number of pixels, which limits accuracy on 4K video and on embedded hardware.
  • Long-range correspondence: two-frame flow chains accumulate drift, which has motivated multi-frame flow and point tracking through many frames.
  • Real-world ground truth: benchmarks are limited by what can be measured, which biases evaluation toward rendered scenes and rigid driving data.

Related

  • Lucas–Kanade Method

    A local, gradient-based method that estimates the displacement of an image window by assuming constant motion within it and solving a small least-squares problem, iterated with warping.

  • Horn–Schunck Method

    A global variational method that computes dense optical flow by minimizing brightness constancy errors together with a penalty on spatial variation of the flow, solved by a simple iterative averaging scheme.

  • Coarse-to-Fine Estimation

    Estimating large motions with small-motion methods by solving on an image pyramid, from the coarsest level to full resolution, and refining the estimate at each level.

  • Aperture Problem

    Why motion seen through a small window is ambiguous along edges, so that only the component of motion across an edge can be measured locally.

  • Endpoint Error

    The standard accuracy measure for optical flow, the distance in pixels between an estimated flow vector and the true one, averaged over the image.

  • RAFT

    A deep network for optical flow that matches all pairs of pixels once and refines a single flow field with a recurrent update operator.

References

  1. Horn, B. K. P. & Schunck, B. G. (1981). Determining Optical Flow. Artificial Intelligence, 17(1–3), 185–203.
  2. Lucas, B. D. & Kanade, T. (1981). An Iterative Image Registration Technique with an Application to Stereo Vision. Proceedings of the 7th International Joint Conference on Artificial Intelligence (IJCAI), 674–679.
  3. Brox, T., Bruhn, A., Papenberg, N. & Weickert, J. (2004). High Accuracy Optical Flow Estimation Based on a Theory for Warping. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science 3024, 25–36.
  4. Brox, T. & Malik, J. (2011). Large Displacement Optical Flow: Descriptor Matching in Variational Motion Estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 33(3), 500–513.
  5. Weinzaepfel, P., Revaud, J., Harchaoui, Z. & Schmid, C. (2013). DeepFlow: Large Displacement Optical Flow with Deep Matching. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 1385–1392.
  6. Revaud, J., Weinzaepfel, P., Harchaoui, Z. & Schmid, C. (2015). EpicFlow: Edge-Preserving Interpolation of Correspondences for Optical Flow. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 1164–1172.
  7. Dosovitskiy, A., Fischer, P., Ilg, E., Häusser, P., Hazırbaş, C., Golkov, V., van der Smagt, P., Cremers, D. & Brox, T. (2015). FlowNet: Learning Optical Flow with Convolutional Networks. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2758–2766.
  8. Ilg, E., Mayer, N., Saikia, T., Keuper, M., Dosovitskiy, A. & Brox, T. (2017). FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 1647–1655.
  9. Sun, D., Yang, X., Liu, M.-Y. & Kautz, J. (2018). PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 8934–8943.
  10. Teed, Z. & Deng, J. (2020). RAFT: Recurrent All-Pairs Field Transforms for Optical Flow. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 402–419.
  11. Butler, D. J., Wulff, J., Stanley, G. B. & Black, M. J. (2012). A Naturalistic Open Source Movie for Optical Flow Evaluation. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 611–625.
  12. Baker, S., Scharstein, D., Lewis, J. P., Roth, S., Black, M. J. & Szeliski, R. (2011). A Database and Evaluation Methodology for Optical Flow. International Journal of Computer Vision, 92(1), 1–31.
  13. Geiger, A., Lenz, P. & Urtasun, R. (2012). Are We Ready for Autonomous Driving? The KITTI Vision Benchmark Suite. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3354–3361.
  14. Menze, M. & Geiger, A. (2015). Object Scene Flow for Autonomous Vehicles. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3061–3070.
  15. Mayer, N., Ilg, E., Häusser, P., Fischer, P., Cremers, D., Dosovitskiy, A. & Brox, T. (2016). A Large Dataset to Train Convolutional Networks for Disparity, Optical Flow, and Scene Flow Estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4040–4048.