Motion & Tracking
RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
The ECCV 2020 best paper by Teed and Deng that replaced coarse-to-fine flow networks with an all-pairs correlation volume and a weight-tied recurrent update.
advanced
“RAFT: Recurrent All-Pairs Field Transforms for Optical Flow” by Zachary Teed and Jia Deng of Princeton University was presented at the European Conference on Computer Vision (ECCV) in 2020, after appearing on arXiv in March of that year, and received the conference’s Best Paper Award. It introduced RAFT, a network for optical flow estimation that compares all pairs of pixels once and then refines a single flow field with a small recurrent update applied many times. It set new records on the Sintel and KITTI benchmarks and generalized markedly better from synthetic training data than earlier networks. This article covers the paper as published; the RAFT article explains the model in detail.
Problem
Classical optical flow had been posed as energy minimization since Horn and Schunck: a data term aligning similar image regions traded off against a regularizer encoding plausible motion. The authors argued that hand-designing an objective robust to fast motion, occlusion, blur, and textureless surfaces had become the bottleneck. Learned methods could sidestep the objective by predicting flow directly, and by the time of the paper they matched the best classical methods while running much faster.
The leading networks, including SpyNet, PWC-Net, LiteFlowNet, and VCN, inherited the classical coarse-to-fine strategy: estimate flow at low resolution, then upsample and refine. Teed and Deng identified three costs of this cascade: errors made at coarse levels are hard to recover from, small fast-moving objects vanish at low resolution, and multi-stage cascades typically need over a million training iterations. Most refinement schemes also used separate weights at each stage, fixing the number of refinements. The only recurrent precedent they cite, IRR, reused FlowNetS (38 million parameters, at most five iterations) or PWC-Net (iterations bounded by pyramid levels). DCFlow had built a 4D cost volume over learned features, but processed it with semi-global matching, so it could not be trained end to end on flow.
Contribution
The paper’s central claim is that learned flow does not need a pyramid of warping stages. RAFT combines three ideas:
- All-pairs correlation. A 4D volume of feature similarities between every pixel of the first frame and every pixel of the second, computed once by a single matrix multiplication and pooled into a multi-scale pyramid over the second frame’s dimensions only. Each update can therefore see both small and large displacements without coarse-to-fine processing.
- A single flow field. Flow is maintained and refined at one resolution throughout, rather than passed up a pyramid, which also allows initializing it from the previous frame pair in video.
- A weight-tied recurrent update. A lightweight convolutional GRU, 2.7 million parameters, reads correlation features around the current estimate and outputs a flow increment. Because its weights are shared, it can be run for any number of steps; the authors applied it more than 100 times at inference without divergence.
They framed the design as a learned analogue of iterative optimization: the encoder learns the features, the correlation volume measures similarity, and the update operator learns to propose descent directions in place of a hand-derived Taylor approximation of a data term.
Method
Two encoders produce features at one-eighth resolution: a feature encoder applied to both frames and a context encoder applied to the first. The correlation volume is built from the feature maps and average-pooled into four levels. At each iteration, a lookup operator samples every level in a small neighborhood around each pixel’s current correspondence; the GRU combines these samples with the current flow and the context features and predicts a residual update. A learned convex upsampling layer turns the one-eighth-resolution flow into full resolution. Training supervises every intermediate estimate with an loss whose weights grow exponentially toward the last iteration, with 12 unrolled updates. The full model has about 5.3 million parameters, and a small variant, RAFT-S, about 1 million.
The paper also notes that the all-pairs volume costs for pixels, and gives an equivalent formulation that pools the second frame’s features instead and computes correlations on demand at for iterations. The RAFT article gives the equations, the GRU structure, and the training schedule.
Results
Results are reported with endpoint error (EPE) and, on KITTI, Fl-all, the percentage of outlier pixels.
- Generalization. Trained only on the synthetic FlyingChairs and FlyingThings datasets, RAFT reached an EPE of 1.43 on the Sintel training set’s clean pass (29% below FlowNet2) and 2.71 on the final pass. On the KITTI 2015 training set it reached 5.04, a 40% reduction from the best earlier network’s 8.36.
- Sintel test set. After fine-tuning, and initializing each pair from the previous frame’s flow, RAFT ranked first on both passes, with an EPE of 1.61 on clean and 2.86 on final (2.855 in the abstract), against 4.098 for the best published final-pass result, a 30% reduction.
- KITTI 2015 test set. Fl-all of 5.10%, against 6.10% for the best published method, a 16% reduction, ranking first among optical flow methods on the leaderboard at the time.
- Efficiency. With 10 updates, RAFT processed frames at 10 per second on a GTX 1080 Ti, and RAFT-S at 20 per second. The authors report about ten times fewer training iterations than other architectures, and RAFT-S outperformed PWC-Net and VCN, which are more than six times larger, after the same synthetic training.
The ablations support the design: tied weights beat untied ones while using far fewer parameters, the GRU beat plain convolutions, correlation lookups beat feature warping, and all-pairs correlation matched a 128-pixel search range on Sintel and did better on KITTI, without a range having to be chosen. Accuracy improved quickly over the first few iterations and kept improving slowly up to 32 updates.
Impact
ECCV 2020 gave RAFT its Best Paper Award, with honorable mentions to NeRF and “Towards Streaming Perception”. The paper became a common starting point for learned flow: a number of later methods kept RAFT’s correlation-plus-iterative-update structure and changed its components. Its pattern, matching once and then refining an estimate with a shared recurrent operator, also transferred beyond flow, largely through its authors’ own follow-up work. The official code was released, RAFT later became available as a pretrained model in torchvision, and the paper has been widely cited.
Limitations
The paper itself is candid about the cost of all-pairs matching: the volume grows quadratically with the number of pixels. The authors found precomputation inexpensive on a GPU, with the correlation taking 95 of the 550 milliseconds per 1080p frame, and offered the on-demand variant, but memory remains the constraint for high-resolution input. Other limitations follow from the design:
- Iterative cost. Accuracy comes from many updates. In the paper’s own timing table, RAFT with 10 updates took 0.10 s per frame pair against 0.04 s for PWC-Net+, though it was faster than VCN, IRR-PWC, and FlowNet2; the benchmark results use 24 to 32 updates.
- Not truly full resolution. The single flow field lives at one-eighth of the input resolution; fine detail at full resolution depends on the learned upsampling.
- Occlusions. Updates rely on local correlation lookups and convolutions, so pixels with no visible match get motion only through spatial propagation of context.
- Supervision. Like other supervised flow networks, RAFT depends on synthetic pretraining and benchmark-specific fine-tuning, and its best Sintel results use the warm start from the previous frame rather than two frames alone.
What Came After
The authors and their colleagues applied the same template to other correspondence problems. RAFT-Stereo (Lipson, Teed, and Deng, 3DV 2021) restricts correlation to the same image row for rectified stereo and adds recurrent units at several resolutions. RAFT-3D (Teed and Deng, CVPR 2021) iteratively updates a field of rigid 3D motions for scene flow. DROID-SLAM (Teed and Deng, NeurIPS 2021) builds monocular, stereo, and RGB-D visual SLAM on recurrent updates of camera poses and depth through a dense bundle adjustment layer.
Within optical flow, GMA (Jiang et al., ICCV 2021) adds a global motion aggregation module to RAFT to estimate motion in occluded regions. FlowFormer (Huang et al., ECCV 2022) encodes the all-pairs cost volume with transformer layers while keeping a recurrent decoder that queries costs at positions given by the current flow estimate. The RAFT article describes these variants in more detail.
Related
- RAFT
A deep network for optical flow that matches all pairs of pixels once and refines a single flow field with a recurrent update operator.
- Optical Flow Estimation
Computing a dense field of pixel displacements between two video frames, from classical variational methods to learned networks such as RAFT.
- Determining Optical Flow
The 1981 paper by Horn and Schunck that computed dense optical flow by combining the brightness change constraint with a global smoothness assumption, founding the variational approach to motion estimation.
References
- Teed, Z. & Deng, J. (2020). RAFT: Recurrent All-Pairs Field Transforms for Optical Flow. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 402–419.
- Lipson, L., Teed, Z. & Deng, J. (2021). RAFT-Stereo: Multilevel Recurrent Field Transforms for Stereo Matching. International Conference on 3D Vision (3DV), 218–227.
- Teed, Z. & Deng, J. (2021). DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras. Advances in Neural Information Processing Systems 34 (NeurIPS), 16558–16569.
- Jiang, S., Campbell, D., Lu, Y., Li, H. & Hartley, R. (2021). Learning to Estimate Hidden Motions with Global Motion Aggregation. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 9752–9761.
- Huang, Z., Shi, X., Zhang, C., Wang, Q., Cheung, K. C., Qin, H., Dai, J. & Li, H. (2022). FlowFormer: A Transformer Architecture for Optical Flow. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 668–685.