Datasets & Benchmarks
Middlebury Benchmarks
Small, high-accuracy stereo and optical flow benchmarks from Middlebury College and Microsoft Research that defined how dense correspondence methods were evaluated in the 2000s.
beginner
The Middlebury benchmarks are two related collections of images with dense ground-truth correspondences, hosted at vision.middlebury.edu: a stereo benchmark, started by Daniel Scharstein and Richard Szeliski with their 2002 taxonomy of stereo algorithms, and an optical flow benchmark, introduced by Baker et al. at ICCV 2007 and described fully in a 2011 journal paper. Both pair a small number of carefully captured image sets with an online evaluation table, and for about a decade they were the main way stereo and optical flow estimation methods were compared.
Purpose
The stereo benchmark compares algorithms quantitatively against ground-truth disparities, within a framework that breaks stereo methods into matching cost computation, cost aggregation, disparity optimization, and refinement (Scharstein and Szeliski, 2002). The flow benchmark extended the earlier synthetic evaluations of the 1990s to real, nonrigid scenes, whose ground-truth flow is far harder to measure than disparity (Baker et al., 2011).
Dataset Size
Stereo. The data was released in batches (as listed on the official datasets page):
- 2001: six piecewise planar scenes, each with 9 views. The 2002 study used two of them, Sawtooth and Venus, with Tsukuba (from the University of Tsukuba) and Map.
- 2003: Cones and Teddy, with complex geometry, up to 1800 × 1500 pixels at full size.
- 2005 and 2006: 9 and 21 further scenes captured with the 2003 technique.
- 2014: 33 scenes of about 6 megapixels (Scharstein et al., 2014).
- 2021: 24 scenes captured with a mobile device mounted on a robot arm.
Optical flow. The flow benchmark has 24 sequences, most of them 8 frames long, split into 12 for evaluation and 12 for training. Each sequence has ground truth for one frame pair. Of the 12 evaluation sequences, 8 have hidden ground-truth flow and 4 high-speed sequences are used only for frame interpolation; the training set likewise has 8 sequences with public ground truth. Images are small: 584 × 388 for the real hidden-texture scenes and 640 × 480 for the new synthetic ones.
Annotation Types
Stereo ground truth is a disparity map for a reference view. The 2001 scenes were made of planar objects such as posters, so each image was hand-labeled into planar regions and the disparity computed from the fitted motion of each plane. From 2003, disparities were measured with structured light: projectors cast Gray-code and sine-wave patterns onto the scene, and decoding them labels each pixel with a code that can be matched between views, without calibrating the projectors (Scharstein and Szeliski, 2003). The 2014 system improved calibration and subpixel matching and reports a disparity accuracy of 0.2 pixels on most surfaces.
Optical flow ground truth came from four kinds of data (Baker et al., 2011):
- Hidden fluorescent texture. Scenes were sprayed with a fine pattern of fluorescent paint and moved in small steps by a computer-controlled stage. At each step a camera took one image under visible light and one under ultraviolet light, where the paint glows. Tracking the glowing texture in high-resolution UV images gives dense flow accurate to about 1/60 pixel, while the paint is nearly invisible in the downsampled test images. Pixels where forward and backward tracking disagree, mostly occlusions, are marked unknown.
- Synthetic scenes. Rendered rocks, trees, and buildings (Grove and Urban) with complex occlusions, plus the classic Yosemite sequence for comparison with older work.
- Stereo pairs. Teddy and Venus from the stereo benchmark, whose disparity is a purely horizontal flow, accurate to about 0.25 pixel.
- High-speed video. Sequences captured at 60 frames per second, with every other frame withheld as ground truth for frame interpolation instead of flow.
Ground-truth flow is stored in the .flo format defined in the benchmark’s code: a 4-byte tag, the width and height as integers, then interleaved little-endian 32-bit floats in row order, with values above marking unknown flow. MPI Sintel and Flying Chairs later adopted the same format. A NumPy reader, tested on a synthetic file:
import numpy as np
TAG = 202021.25 # "PIEH" read as a little-endian float32
def read_flo(path):
"""Read a Middlebury .flo file into an (H, W, 2) float32 array of (u, v)."""
with open(path, "rb") as f:
tag = np.fromfile(f, "<f4", count=1)[0]
if tag != TAG:
raise ValueError(f"{path}: not a .flo file (tag {tag})")
width, height = np.fromfile(f, "<i4", count=2)
data = np.fromfile(f, "<f4", count=2 * width * height)
return data.reshape(height, width, 2)
def write_flo(path, flow):
height, width, _ = flow.shape
with open(path, "wb") as f:
f.write(b"PIEH")
np.array([width, height], "<i4").tofile(f)
flow.astype("<f4").tofile(f)
# Round trip on a synthetic 4 x 6 flow field with one unknown pixel.
flow = np.zeros((4, 6, 2), np.float32)
flow[..., 0] = 1.5 # everything moves 1.5 px to the right
flow[1, 2] = (-3.0, 0.25)
flow[3, 5] = (1e10, 1e10) # |u| or |v| above 1e9 means "unknown"
write_flo("example.flo", flow)
f = read_flo("example.flo")
valid = (np.abs(f) <= 1e9).all(axis=-1)
print(f.shape, f.dtype)
print("valid pixels:", valid.sum(), "of", valid.size)
print("flow at row 1, col 2:", f[1, 2])
print("mean u over valid pixels:", round(float(f[valid, 0].mean()), 4))
Output:
(4, 6, 2) float32
valid pixels: 23 of 24
flow at row 1, col 2: [-3. 0.25]
mean u over valid pixels: 1.3043
Splits
Both benchmarks withhold the ground truth of their test data. The flow benchmark publishes ground truth only for its training (“other”) sequences, and Yosemite is the one evaluation sequence whose flow was already public. The current stereo benchmark, version 3, has 15 training and 15 test image pairs, mostly from the 2014 scenes, at up to 3000 × 2000 pixels with disparities up to 800 pixels. Results for both sets had to be submitted together with the same parameters, so methods got one shot at the test set.
Tasks
- Two-view stereo matching: estimating a disparity map from a rectified image pair, typically by building and optimizing a cost volume.
- Optical flow estimation between two frames, plus multi-frame methods that use the surrounding frames.
- Frame interpolation, using the high-speed flow sequences.
Evaluation
The stereo benchmark ranks methods mainly by the percentage of bad pixels, whose disparity error exceeds a threshold (1 pixel in the 2002 study), alongside RMS disparity error. Statistics are computed over all pixels, non-occluded pixels, textureless regions, and regions near depth discontinuities. Version 3 adds several thresholds, average and percentile errors, and running time.
The flow benchmark reports endpoint error and angular error for flow, and interpolation error and normalized interpolation error for predicted frames. Each measure has averages, outlier percentages, and percentile errors over all pixels, near motion discontinuities, and textureless regions: 32 tables, each sorted by average rank over the 8 sequences and 3 regions. The organizers state that they do not aim to provide one overall ranking.
License
As of October 2026, neither benchmark names a formal license. The stereo pages grant permission to use and publish the images and disparity maps, and ask users to cite the paper for each batch of datasets. The flow pages ask that anyone reporting results cite Baker et al. (2011) but state no terms of use. Check with the authors before commercial use.
Limitations and Bias
- Small. The flow training set has ground truth for only 8 frame pairs, far too little to train deep networks. Butler et al. (2012) note that it was not designed as training data; the small size of Middlebury, KITTI, and Sintel later led FlowNet’s authors to create Flying Chairs.
- Small motions. Most flow is under about 10 pixels per frame; the largest is about 35 pixels, in the synthetic Urban scene.
- Laboratory scenes. The hidden-texture scenes are laboratory setups captured in stop motion, so there is no motion blur, and all stereo scenes are static and mostly indoor.
- Saturation. By 2012 the top flow methods were tightly grouped, so small changes on one sequence could reorder the average-rank ranking; Butler et al. (2012) cited this as motivation for Sintel.
Common Uses
Middlebury was the standard benchmark for classical stereo methods and for variational flow methods descended from Horn–Schunck, usually solved coarse to fine. Beyond its .flo format, its color coding for visualizing flow was adopted by Sintel.
Historical Milestones
- 2002: The stereo taxonomy and the first online stereo evaluation.
- 2003: Structured-light ground truth, motivated by stereo methods outpacing the ability of existing data to tell the best apart (Scharstein and Szeliski, 2003). The second version of the stereo evaluation ranked methods on Tsukuba, Venus, Teddy, and Cones.
- 2007: The flow benchmark was first presented at ICCV, with a preliminary dataset and a handful of algorithms implemented by the authors.
- 2009–2011: The journal paper, first released as a Microsoft Research technical report in December 2009, argued that endpoint error should become the preferred measure of flow accuracy instead of angular error; the default results page shows average endpoint error.
- 2012: When MPI Sintel was introduced, 73 methods were listed on the flow benchmark, and Classic+NL, ranked 13th, had an average endpoint error of 0.32 pixel (Butler et al., 2012).
- 2014–2015: 33 high-resolution stereo scenes and version 3 of the stereo evaluation, which its authors describe as integrating lessons from KITTI, PASCAL, and Sintel, with public tables for both the training and test sets.
As of October 2026, neither online evaluation accepts new results: the flow submission page states that submission is no longer enabled, and the stereo version 3 page states that upload and evaluation are disabled. Published results remain viewable.
Related Datasets
- MPI Sintel: rendered flow with long sequences, large motions, and blur, designed to address Middlebury’s limits.
- KITTI: real driving scenes with laser-based ground truth for stereo and flow.
- Flying Chairs: synthetic training data for learned flow.
Resources
- Middlebury stereo: datasets, evaluation version 3 tables, and the evaluation SDK.
- Middlebury optical flow: datasets, results tables, and the
.floreading and writing code.
Related
- Endpoint Error
The standard accuracy measure for optical flow, the distance in pixels between an estimated flow vector and the true one, averaged over the image.
- MPI Sintel
An optical flow benchmark rendered from the open-source animated film Sintel, with long sequences, large motions, motion blur, and atmospheric effects.
- KITTI
A benchmark suite of real driving data from Karlsruhe, with stereo, optical flow, scene flow, odometry, object detection, and tracking benchmarks built from camera, LiDAR, and GPS/IMU recordings.
- Cost Volume
A tensor of matching costs or similarities between each pixel of one image and its candidate correspondences in another, indexed by pixel position and candidate disparity, displacement, or depth.
- Optical Flow
The apparent motion of image content between two frames, represented as a two-dimensional displacement at every pixel.
- Coarse-to-Fine Estimation
Estimating large motions with small-motion methods by solving on an image pyramid, from the coarsest level to full resolution, and refining the estimate at each level.
References
- Scharstein, D. & Szeliski, R. (2002). A Taxonomy and Evaluation of Dense Two-Frame Stereo Correspondence Algorithms. International Journal of Computer Vision, 47(1–3), 7–42.
- Scharstein, D. & Szeliski, R. (2003). High-Accuracy Stereo Depth Maps Using Structured Light. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR), Vol. 1, 195–202.
- Baker, S., Scharstein, D., Lewis, J. P., Roth, S., Black, M. J. & Szeliski, R. (2011). A Database and Evaluation Methodology for Optical Flow. International Journal of Computer Vision, 92(1), 1–31.
- Scharstein, D., Hirschmüller, H., Kitajima, Y., Krathwohl, G., Nešić, N., Wang, X. & Westling, P. (2014). High-Resolution Stereo Datasets with Subpixel-Accurate Ground Truth. German Conference on Pattern Recognition (GCPR), Lecture Notes in Computer Science, 31–42.
- Butler, D. J., Wulff, J., Stanley, G. B. & Black, M. J. (2012). A Naturalistic Open Source Movie for Optical Flow Evaluation. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 611–625.