Motion & Tracking
RAFT
A deep network for optical flow that matches all pairs of pixels once and refines a single flow field with a recurrent update operator.
advanced
RAFT (Recurrent All-Pairs Field Transforms) is a deep network for optical flow estimation, introduced by Zachary Teed and Jia Deng at ECCV 2020. It is a convolutional network with a recurrent update operator: it compares every pixel of the first frame with every pixel of the second in a 4D correlation volume, then refines a single flow field over many iterations with a convolutional GRU that reads from that volume. This design, matching once and refining iteratively at one resolution, became a template for later learned flow, stereo, and scene flow models.
Problem Addressed
Before RAFT, the strongest flow networks, such as PWC-Net and LiteFlowNet, followed the classical coarse-to-fine scheme. Teed and Deng named three weaknesses of this cascade: coarse-level errors are hard to undo, small fast-moving objects are missed, and multi-stage cascades need long training, often over a million iterations.
Earlier refinement networks also mostly used separate weights for each step, which fixed the number of refinements; the one recurrent exception the authors cite, IRR, was limited to five iterations or to the number of pyramid levels. RAFT instead keeps one flow field at a fixed resolution, looks up both small and large displacements in a precomputed multi-scale correlation volume, and applies a small weight-tied update many times. The authors liken the update to a step of a first-order optimizer whose features and motion priors are learned.
Architecture
Given frames and , RAFT extracts features and builds the correlation volume once, then iteratively updates a flow field initialized to zero, with all stages trained end to end. The full model has 4.8 million parameters, or 5.3 million with the learned upsampling module; a small version, RAFT-S, has about 1 million.
Feature Encoder
A convolutional encoder maps each frame to features at one-eighth resolution, with . It consists of six residual blocks, two each at one-half, one-quarter, and one-eighth resolution, uses instance normalization, and is applied with shared weights to both frames. RAFT-S replaces the residual units with bottleneck units.
Context Encoder
A second network with the same architecture, applied only to , uses batch normalization in the full model. In the official implementation its output is split in two: one part, passed through , initializes the GRU’s hidden state, and the other is injected as context into every update step.
Correlation Volume
The visual similarity between all pairs of feature vectors forms a 4D volume, where and here denote the size of the feature maps:
computed as a single matrix multiplication (the official code also divides by ). RAFT then average-pools the last two dimensions, those of , with kernel sizes 1, 2, 4, and 8, giving a four-level correlation pyramid. Because the dimensions are never pooled, every pixel of the first frame keeps its own full-resolution correlation map, which the authors credit for recovering small, fast objects.
A lookup operator turns the volume into features. For each pixel of , it takes the estimated correspondence and samples each pyramid level bilinearly at integer offsets within radius of , scaled to that level. The paper defines this neighborhood with the distance, , but the official code samples the full square of offsets. With for the full model, the coarsest level reaches offsets of 256 pixels at input resolution. The samples from all levels are concatenated.
The volume costs for feature pixels, once per image pair; an equivalent variant pools the features of instead and computes correlations on demand, at for iterations.
Update Operator
The update operator is a convolutional GRU with weights tied across iterations. Its input concatenates encoded correlation features, encoded current flow, and context features:
The full model replaces each GRU with two in sequence, one with and one with convolutions, to enlarge the receptive field cheaply; RAFT-S uses a single GRU. Two convolutions on the hidden state predict a residual , and the estimate becomes , always at one-eighth resolution. The update operator has 2.7 million parameters.
Upsampling
Each full-resolution flow vector is a convex combination of the coarse vectors around it. Two convolutions on the hidden state predict an mask, and a softmax over the 9 neighbors makes the weights convex. The authors found this more accurate than bilinear upsampling, particularly near motion boundaries.
Training
RAFT is trained on ground-truth flow: pretraining on the synthetic FlyingChairs (100k iterations) and FlyingThings3D (100k iterations) datasets, then dataset-specific fine-tuning. For Sintel, the authors fine-tuned for 100k iterations on a mixture of Sintel, FlyingThings3D, KITTI 2015, and HD1K; for KITTI, they fine-tuned that model for another 50k iterations on KITTI 2015. Training used AdamW, photometric, spatial, and occlusion augmentation, and 12 unrolled updates.
The loss is the distance between ground truth and every intermediate prediction , upsampled to full resolution, with weights that increase exponentially toward the last iteration:
Gradients flow through each update but are stopped through the incoming estimate .
Inference
RAFT takes two RGB frames and returns a full-resolution flow field. Because the update is weight-tied, the number of iterations is chosen at test time: the authors evaluated with 32 updates on Sintel and 24 on KITTI. In their ablation, accuracy changes little beyond 32 updates, and even 200 updates do not diverge.
For video, a single-resolution design allows warm starting: the flow from the previous frame pair is forward-projected to initialize the next, with occlusion gaps filled by nearest-neighbor interpolation.
The authors reported 10 frames per second at on a GTX 1080 Ti with 10 updates (0.10 s per pair, against 0.04 s for PWC-Net+ on the same GPU), and 20 frames per second for RAFT-S. On 1080p video, 12 updates took 550 ms per frame, of which the all-pairs correlation took 95 ms.
Performance
As reported by Teed and Deng (2020), with endpoint error (EPE) and, on the KITTI 2015 test set, Fl-all, the percentage of outlier pixels:
- Generalization from synthetic data. Trained only on FlyingChairs and FlyingThings3D, RAFT reached an EPE of 1.43 on the Sintel training set’s clean pass and 2.71 on its final pass, and 5.04 on the KITTI 2015 training set, compared with 8.36 for the best earlier network trained on the same data.
- Sintel test set (2020). With warm start, RAFT reached 1.61 on the clean pass and 2.86 on the final pass (2.855 in the abstract), a 30% reduction from the best published final-pass result of 4.098. The two-frame model scored 1.94 and 3.18.
- KITTI 2015 test set (2020). Fl-all of 5.10%, against 6.10% for the best published method.
- Efficiency. After the same synthetic training, RAFT-S had a lower Sintel final-pass error than PWC-Net and VCN, which are more than six times larger.
Strengths
- Large and small displacements are handled together, without a coarse-to-fine cascade.
- Accuracy can be traded for speed at test time through the number of iterations.
- Few parameters, about ten times fewer training iterations than earlier networks by the authors’ account, and strong generalization from synthetic data.
Limitations
- The all-pairs volume grows quadratically with the number of pixels, which limits resolution on memory-constrained hardware unless the slower on-demand variant is used.
- The update operator relies on local correlation lookups and convolutions, so occluded pixels, which have no match, receive motion only through spatial propagation; GMA was designed to address this.
- Iterative refinement makes it slower than single-pass networks such as PWC-Net+.
- Like other supervised flow networks, it depends on synthetic pretraining and benchmark-specific fine-tuning.
Historical Importance
RAFT (2020) showed that learned flow does not need the pyramid of warping stages that had dominated it. Its authors carried the pattern beyond optical flow by changing the quantity being updated, to disparity, rigid 3D motion, or camera poses and depth, and other groups built on RAFT directly, as GMA did.
Variants
RAFT-Stereo (Lipson, Teed, and Deng, 2021) applies the design to rectified stereo, building a 3D correlation volume restricted to the same image row and adding GRUs at several resolutions. RAFT-3D (Teed and Deng, 2021) estimates scene flow from stereo or RGB-D frames by iteratively updating a dense field of rigid 3D motions. DROID-SLAM (Teed and Deng, 2021) builds visual SLAM on RAFT’s components, recurrently updating camera poses and depth through a dense bundle adjustment layer over many frames.
Within optical flow, GMA (Jiang et al., 2021) adds a transformer-based global motion aggregation module to RAFT to improve estimates in occluded regions. FlowFormer (Huang et al., 2022) encodes the all-pairs cost volume with transformer layers and decodes it with a recurrent transformer decoder that, as in RAFT, queries costs at positions given by the current flow estimate.
Implementations
The official PyTorch implementation, princeton-vl/RAFT, is released under the BSD 3-Clause license with pretrained models. torchvision includes the full and small models as raft_large and raft_small in torchvision.models.optical_flow, with pretrained weights. The following sketch follows the torchvision 0.29 API; image height and width must be divisible by 8:
import torch
from torchvision.models.optical_flow import raft_large, Raft_Large_Weights
weights = Raft_Large_Weights.DEFAULT
model = raft_large(weights=weights).eval()
preprocess = weights.transforms() # converts to float and maps to [-1, 1]
# frame1, frame2: uint8 tensors of shape (B, 3, H, W)
img1, img2 = preprocess(frame1, frame2)
with torch.no_grad():
flows = model(img1, img2, num_flow_updates=12) # list, one flow per update
flow = flows[-1] # (B, 2, H, W), in pixels
Related
- Coarse-to-Fine Estimation
Estimating large motions with small-motion methods by solving on an image pyramid, from the coarsest level to full resolution, and refining the estimate at each level.
- Optical Flow
The apparent motion of image content between two frames, represented as a two-dimensional displacement at every pixel.
- Horn–Schunck Method
A global variational method that computes dense optical flow by minimizing brightness constancy errors together with a penalty on spatial variation of the flow, solved by a simple iterative averaging scheme.
- Motion Estimation
Recovering how image content, objects, or the camera moved from a sequence of images, from per-pixel flow to global and 3D motion.
- RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
The ECCV 2020 best paper by Teed and Deng that replaced coarse-to-fine flow networks with an all-pairs correlation volume and a weight-tied recurrent update.
References
- Teed, Z. & Deng, J. (2020). RAFT: Recurrent All-Pairs Field Transforms for Optical Flow. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 402–419.
- Teed, Z. & Deng, J. (2021). RAFT-3D: Scene Flow using Rigid-Motion Embeddings. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 8371–8380.
- Jiang, S., Campbell, D., Lu, Y., Li, H. & Hartley, R. (2021). Learning to Estimate Hidden Motions with Global Motion Aggregation. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 9752–9761.
- Lipson, L., Teed, Z. & Deng, J. (2021). RAFT-Stereo: Multilevel Recurrent Field Transforms for Stereo Matching. International Conference on 3D Vision (3DV), 218–227.
- Teed, Z. & Deng, J. (2021). DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras. Advances in Neural Information Processing Systems 34 (NeurIPS), 16558–16569.
- Huang, Z., Shi, X., Zhang, C., Wang, Q., Cheung, K. C., Qin, H., Dai, J. & Li, H. (2022). FlowFormer: A Transformer Architecture for Optical Flow. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 668–685.