Datasets & Benchmarks
Flying Chairs
A synthetic optical flow dataset of rendered chairs moving over Flickr photographs, created to train FlowNet and still used to pretrain flow networks.
beginner
Flying Chairs (written FlyingChairs in file names and code) is a synthetic optical flow dataset of 22,872 image pairs in which rendered chairs move in front of photographs from Flickr. Alexey Dosovitskiy, Philipp Fischer, and colleagues at the University of Freiburg and the Technical University of Munich created it in 2015 to train FlowNet, the convolutional network that showed optical flow could be learned end to end. The scenes look nothing like real video, yet networks trained on them predicted flow on real benchmarks with competitive accuracy.
Purpose
A flow network has to learn the whole task from examples, and in 2015 the available ground truth was far too small: Middlebury has 8 training pairs with ground truth, KITTI 194, and MPI Sintel 1,041 (Dosovitskiy et al., 2015). Since true correspondences are very hard to measure in real footage, the authors traded realism for quantity: synthetic image pairs whose flow is known exactly because the motion is generated.
Dataset Size
The dataset contains 22,872 image pairs, each with a dense ground-truth flow field, at 512 × 384 pixels. The authors note that the size was chosen arbitrarily; the generator could produce more.
The pairs are built from two ingredients (Dosovitskiy et al., 2015):
- Backgrounds: 964 Flickr photographs of 1024 × 768 pixels from the categories “city”, “landscape”, and “mountain”, each cut into four 512 × 384 quadrants. Each background is reused many times.
- Foregrounds: 809 chair models from a public collection of rendered 3D chairs, each available in 62 views. Each image contains 16 to 24 chairs of random type, view, size, and position.
Annotation Types
To create a pair, the generator samples a random affine transformation (scaling, rotation, and translation) for the background and a further one for each chair relative to the background, which mimics a moving camera and independently moving objects. Applying the transformations renders the second image and gives the exact flow at every pixel. The parameter distributions were tuned so that the histogram of displacements roughly matches Sintel’s, with many small motions and a long tail of large ones.
The released archive contains the two images of each pair as .ppm files and the forward flow in the Middlebury .flo format. Flying Chairs 2, generated later with the same settings, adds backward flow, occlusions, motion boundaries, and object IDs.
Splits
The official split, distributed as a text file on the dataset page, assigns 22,232 pairs to training and 640 to validation; the FlowNet paper calls the 640 pairs a test set and used them to monitor overfitting. There is no hidden test set or leaderboard: Flying Chairs is training data, and flow methods are compared on other benchmarks.
Tasks
Flying Chairs is used to train supervised networks for optical flow estimation. Low error on its validation pairs says little about accuracy on real footage.
Evaluation
Validation error is measured with average endpoint error. In the FlowNet paper, the networks beat classical methods such as DeepFlow and EpicFlow on the 640 held-out pairs, but not on Sintel, where the motions and appearance differ.
License
As of October 2026, the official dataset page states that the dataset is provided for research purposes only and without any warranty, and that any commercial use is prohibited; users are asked to cite the FlowNet paper. No standard license name is given. The FlowNet paper notes that the Flickr background images are under a non-commercial license.
Limitations and Bias
- Unrealistic scenes. Chairs float over unrelated photographs, with no lighting interaction, shadows, or consistent 3D layout.
- Planar motion. Both chairs and backgrounds move by 2D affine transformations, so there is no parallax, perspective change, or nonrigid motion as in real scenes.
- No camera effects. Unlike Sintel’s final pass, there is no motion blur or fog.
- Repetition. 964 backgrounds and 809 chair types are reused across 22,872 pairs, and the authors found data augmentation (geometric transformations, noise, and color changes) crucial to avoid overfitting.
- Domain gap. FlowNet’s authors observed that KITTI’s strong projective transformations differ greatly from anything in Flying Chairs; the raw predictions there were fairly good, and fine-tuning on Sintel, whose motions and images are more natural, improved them.
Common Uses
Flying Chairs is the first stage of the usual training schedule for supervised flow networks:
- Chairs: pretrain on Flying Chairs.
- Things: continue on FlyingThings3D, a more realistic synthetic dataset.
- Fine-tune on the target benchmark, such as Sintel or KITTI, often mixed with other data.
FlyingThings3D (Mayer et al., 2016), which Ilg et al. (2017) describe as a three-dimensional version of Flying Chairs, shows ShapeNet objects flying along random 3D trajectories in front of static 3D backgrounds, rendered at 960 × 540 with true 3D motion and lighting. It contains about 25,000 stereo frames, with dense flow, disparity, and disparity change, and is part of a collection with two other synthetic datasets, Monkaa and Driving.
Ilg et al. (2017) found that the order matters. Training on FlyingThings3D alone gave worse results than training on Flying Chairs alone, and the best results came from training on Flying Chairs first and only then on FlyingThings3D, which also beat training on a mixture of the two. For FlowNetS, continuing to train on Flying Chairs lowered the endpoint error on the Sintel clean training set only from 4.24 to 4.21, while continuing on FlyingThings3D lowered it to 3.79. They conjectured that the simpler data teaches the network to match colors before it develops priors for 3D motion and realistic lighting.
RAFT, for example, follows this schedule, with 100,000 iterations on each synthetic dataset before fine-tuning. RAFT also reports results after training on the two synthetic datasets alone, labeled “C+T”, as a test of how well a model generalizes to other benchmarks.
Historical Milestones
- 2015: FlowNet, trained only on Flying Chairs, beat the large-displacement method LDOF on the Sintel test set without fine-tuning (Dosovitskiy et al., 2015). Its FlowNetC variant introduced the correlation layer, an early learned cost volume.
- 2016: FlyingThings3D, Monkaa, and Driving extended the approach to disparity and scene flow (Mayer et al., 2016).
- 2017: FlowNet 2.0 established the Chairs-then-Things schedule and released ChairsSDHom, a variant in the style of Flying Chairs with smaller displacements, for real-world footage with small motions (Ilg et al., 2017).
- 2020: Trained only on Flying Chairs and FlyingThings3D, RAFT reached an endpoint error of 1.43 on the Sintel clean training set and 2.71 on the final pass, as reported by Teed and Deng (2020).
Related Datasets
- FlyingThings3D: the 3D successor, used as the second training stage.
- MPI Sintel: a rendered benchmark with realistic appearance, used for fine-tuning and evaluation.
- KITTI: real driving scenes, the main real-world target.
- Middlebury: the earlier benchmark whose
.floformat Flying Chairs uses.
Resources
- Flying Chairs dataset page: downloads of Flying Chairs, Flying Chairs 2, and ChairsSDHom, the train/validation split, and Python reading routines.
- Scene Flow datasets page: FlyingThings3D, Monkaa, and Driving.
Related
- RAFT
A deep network for optical flow that matches all pairs of pixels once and refines a single flow field with a recurrent update operator.
- MPI Sintel
An optical flow benchmark rendered from the open-source animated film Sintel, with long sequences, large motions, motion blur, and atmospheric effects.
- KITTI
A benchmark suite of real driving data from Karlsruhe, with stereo, optical flow, scene flow, odometry, object detection, and tracking benchmarks built from camera, LiDAR, and GPS/IMU recordings.
- Cost Volume
A tensor of matching costs or similarities between each pixel of one image and its candidate correspondences in another, indexed by pixel position and candidate disparity, displacement, or depth.
- Endpoint Error
The standard accuracy measure for optical flow, the distance in pixels between an estimated flow vector and the true one, averaged over the image.
- Middlebury Benchmarks
Small, high-accuracy stereo and optical flow benchmarks from Middlebury College and Microsoft Research that defined how dense correspondence methods were evaluated in the 2000s.
References
- Dosovitskiy, A., Fischer, P., Ilg, E., Häusser, P., Hazırbaş, C., Golkov, V., van der Smagt, P., Cremers, D. & Brox, T. (2015). FlowNet: Learning Optical Flow with Convolutional Networks. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2758–2766.
- Mayer, N., Ilg, E., Häusser, P., Fischer, P., Cremers, D., Dosovitskiy, A. & Brox, T. (2016). A Large Dataset to Train Convolutional Networks for Disparity, Optical Flow, and Scene Flow Estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4040–4048.
- Ilg, E., Mayer, N., Saikia, T., Keuper, M., Dosovitskiy, A. & Brox, T. (2017). FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 1647–1655.
- Teed, Z. & Deng, J. (2020). RAFT: Recurrent All-Pairs Field Transforms for Optical Flow. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 402–419.