Datasets & Benchmarks
MPI Sintel
An optical flow benchmark rendered from the open-source animated film Sintel, with long sequences, large motions, motion blur, and atmospheric effects.
beginner
MPI Sintel is an optical flow dataset and benchmark rendered from Sintel, a 2010 open-source animated short film produced by the Blender Foundation. It was created by Daniel Butler, Jonas Wulff, Garrett Stanley, and Michael Black at the University of Washington, the Max Planck Institute for Intelligent Systems in Tübingen, and Georgia Tech, and was presented at ECCV 2012. Because the film’s 3D scenes are public, the authors could re-render them with exact per-pixel motion, giving dense ground truth for imagery far more complex than earlier benchmarks.
Purpose
Ground-truth flow cannot be measured directly in real scenes with natural motion, so earlier datasets were small and simple. By 2012 the top methods on the Middlebury flow benchmark were tightly grouped, and its sequences had small motions and no blur. Sintel was designed to create “room at the top” (Butler et al., 2012) with long sequences, large and nonrigid motion, specular reflections, motion and defocus blur, and atmospheric effects such as fog, and to provide enough training data for learning-based methods.
Dataset Size
The authors selected 35 clips from the film: 23 for training and 12 for testing. Most clips are 50 frames long; six are shorter. In total there are 1,064 training frames and 564 test frames, which gives 1,041 training frame pairs with ground truth. Frames are 1024 × 436 pixels, saved as 8-bit PNG files, at 24 frames per second. The complete download is about 5.3 GB.
Annotation Types
Each frame is rendered in up to three passes that add realism step by step:
- Albedo: flat, unshaded colors, so brightness constancy holds almost everywhere except at occlusions.
- Clean: adds shading, shadows, specular reflections, and inter-reflections.
- Final: adds motion blur, depth-of-field blur, atmospheric effects, and color correction, close to the released film.
Ground-truth flow comes from a modified version of Blender’s motion blur pipeline, which computes the motion vector of each pixel. For the training set, the dataset provides forward flow in the Middlebury .flo format, masks of unmatched (occluded) pixels, and invalid-pixel masks. In 2015 the authors added training data for stereo disparity, depth and camera motion, and segmentation.
To make the flow well defined, the authors changed some scenes: for example, the main character’s hair, a semi-transparent particle system in the film, was re-rendered as opaque strands, and transparency was excluded throughout. Their workshop paper (Wulff et al., 2012) documents how the dataset differs from the film.
Splits
Ground truth is public for the 23 training sequences and withheld for the 12 test sequences, which are provided in the clean and final passes only. Because the source files are public, a determined cheater could reconstruct the test flow, so two test sequences were rendered with small random perturbations of the camera and static objects. A method that does much worse on them than on the rest would suggest cheating.
The website limits each user to one upload every eight hours and three per 30 days, and asks that methods be tuned on the training set, which users can split for validation. A bundler program subsamples the flow fields before upload.
Tasks
The main task is two-frame or multi-frame optical flow estimation. The 2015 additions provide training data for stereo, depth, and segmentation.
Evaluation
Methods are ranked by average endpoint error (EPE) on the test set, separately for the clean and final passes; the final pass is the table’s default view. The table also reports EPE over matched pixels, visible in both frames, and unmatched ones, visible only in the first, as well as EPE binned by distance from the nearest occlusion boundary and by speed. The endpoint error article gives the bin definitions and the size of the unmatched region. The average is taken over all frames, not over per-sequence averages, so longer sequences count for more.
Uploaded results appear in the public table only once approved for display, and every public method must be fully described in a document such as a published paper or technical report.
License
As of October 2026, the MPI Sintel website, including its downloads page and FAQ, names no license or terms of use for the dataset; it shows only a “© 2012 Max Planck Gesellschaft” notice and asks users to cite Butler et al. (2012). The film itself is released under the Creative Commons Attribution 3.0 license, with attribution to the Blender Foundation. The terms for the derived dataset are therefore unclear; check with the authors before commercial use.
Limitations and Bias
- Synthetic appearance. Sintel is rendered, not filmed. Butler et al. found its first-order image and motion statistics similar to those of real film and video clips with similar content, but they did not claim it represents all natural video.
- Physics violations. Objects sometimes pass through each other, and lighting is often physically implausible, as in many films. The authors caution against using it for methods that rely strongly on physical laws.
- One film. All scenes share one artistic style and a small set of characters and environments, so methods tuned to Sintel may overfit to it.
- No transparency. Each pixel has a single flow vector, so transparent and semi-transparent surfaces are excluded.
- Small by deep-learning standards. About a thousand training pairs is too little to train a flow network from scratch, which is why networks are pretrained on synthetic data such as Flying Chairs.
Common Uses
Sintel and KITTI are the benchmarks on which learned flow methods such as FlowNet, FlowNet 2.0, and RAFT reported their main results, on both Sintel passes. Its training set is used for fine-tuning learned models after synthetic pretraining, and for measuring generalization when a model is evaluated on it without fine-tuning. The clean and final passes let papers separate failures caused by large motion and occlusion from failures caused by blur and fog.
Historical Milestones
- 2012: At release, methods with endpoint errors below 0.5 pixel on Middlebury scored around 10 pixels on Sintel, with errors above 40 pixels in unmatched regions (Butler et al., 2012).
- 2015: FlowNet, a convolutional network trained end to end on Flying Chairs and fine-tuned on Sintel, was competitive with classical methods but did not match the best of them, EpicFlow, on the Sintel test set (Dosovitskiy et al., 2015).
- 2017: FlowNet 2.0 reported accuracy on par with state-of-the-art methods while running at interactive frame rates (Ilg et al., 2017).
- 2020: RAFT reached an endpoint error of 2.855 on the final pass, a 30% reduction from the best published result of 4.098, as reported by Teed and Deng (2020).
Related Datasets
- Middlebury: the earlier flow benchmark, with small motions and very accurate ground truth.
- KITTI: real driving scenes with semi-dense laser-based ground truth.
- Flying Chairs and FlyingThings3D: large synthetic training sets used before fine-tuning on Sintel. FlyingThings3D’s companion dataset Monkaa is rendered, like Sintel, from an open-source Blender film.
Resources
- MPI Sintel website: downloads, the current results table, and submission.
- Project page at the Max Planck Institute for Intelligent Systems.
Related
- Endpoint Error
The standard accuracy measure for optical flow, the distance in pixels between an estimated flow vector and the true one, averaged over the image.
- Middlebury Benchmarks
Small, high-accuracy stereo and optical flow benchmarks from Middlebury College and Microsoft Research that defined how dense correspondence methods were evaluated in the 2000s.
- KITTI
A benchmark suite of real driving data from Karlsruhe, with stereo, optical flow, scene flow, odometry, object detection, and tracking benchmarks built from camera, LiDAR, and GPS/IMU recordings.
- Flying Chairs
A synthetic optical flow dataset of rendered chairs moving over Flickr photographs, created to train FlowNet and still used to pretrain flow networks.
- RAFT
A deep network for optical flow that matches all pairs of pixels once and refines a single flow field with a recurrent update operator.
- Optical Flow
The apparent motion of image content between two frames, represented as a two-dimensional displacement at every pixel.
References
- Butler, D. J., Wulff, J., Stanley, G. B. & Black, M. J. (2012). A Naturalistic Open Source Movie for Optical Flow Evaluation. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 611–625.
- Wulff, J., Butler, D. J., Stanley, G. B. & Black, M. J. (2012). Lessons and Insights from Creating a Synthetic Optical Flow Benchmark. Computer Vision – ECCV 2012. Workshops and Demonstrations, Lecture Notes in Computer Science, 168–177.
- Dosovitskiy, A., Fischer, P., Ilg, E., Häusser, P., Hazırbaş, C., Golkov, V., van der Smagt, P., Cremers, D. & Brox, T. (2015). FlowNet: Learning Optical Flow with Convolutional Networks. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2758–2766.
- Ilg, E., Mayer, N., Saikia, T., Keuper, M., Dosovitskiy, A. & Brox, T. (2017). FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 1647–1655.
- Teed, Z. & Deng, J. (2020). RAFT: Recurrent All-Pairs Field Transforms for Optical Flow. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 402–419.