Datasets & Benchmarks
KITTI
A benchmark suite of real driving data from Karlsruhe, with stereo, optical flow, scene flow, odometry, object detection, and tracking benchmarks built from camera, LiDAR, and GPS/IMU recordings.
beginner
KITTI is a collection of computer vision benchmarks built from real driving data, created as a joint project of the Karlsruhe Institute of Technology and the Toyota Technological Institute at Chicago. Andreas Geiger, Philip Lenz, and Raquel Urtasun introduced the suite at CVPR 2012 with benchmarks for stereo, optical flow, visual odometry, and object detection with orientation estimation. A 2013 journal paper with Christoph Stiller described the raw recordings, and in 2015 Moritz Menze and Geiger added harder stereo, flow, and scene flow benchmarks with moving objects. Later additions cover tracking (2013), road and lane estimation, 3D detection, depth, and semantic segmentation.
KITTI is a shared pool of recordings from which separate benchmarks were extracted, each with its own data, splits, metric, and evaluation server. This article describes the suite as a whole.
Purpose
When KITTI appeared, the standard stereo and flow benchmarks, notably Middlebury, were captured in controlled laboratory settings. Geiger et al. (2012) argued that such data did not reflect the difficulties of robotics, and found that methods ranking high on Middlebury performed below average on real street scenes. KITTI’s goal was real-world benchmarks with accurate ground truth for the tasks an autonomous car needs.
Dataset Size
The recording platform was a Volkswagen Passat station wagon carrying two grayscale and two color Point Grey Flea2 cameras, arranged as two stereo pairs with a baseline of about 54 cm, a Velodyne HDL-64E rotating laser scanner, and an OXTS RT 3003 GPS/IMU navigation system with RTK correction. The laser scanner spins at 10 Hz and triggers the cameras, so images and scans are synchronized at 10 frames per second. Rectified images are about 0.5 megapixels, roughly 1240 × 376 pixels.
Geiger et al. (2013) report 6 hours of recordings, made during daytime on five days in September and October 2011 in and around Karlsruhe, Germany, including highways and rural roads. The published raw data, sorted into the categories City, Residential, Road, Campus, and Person, contains about a quarter of the recordings and totals 180 GB.
Classes
The object annotations define eight classes: Car, Van, Truck, Pedestrian, Person (sitting), Cyclist, Tram, and Misc (for example, trailers and Segways). Cars and pedestrians dominate, because labeling focused on them. The object detection benchmark evaluates cars, pedestrians, and cyclists, and the tracking benchmark evaluates only cars and pedestrians, the classes with enough labeled instances. The semantic segmentation benchmark added in 2018 follows the Cityscapes label format.
Annotation Types
Most ground truth comes from the laser scanner, not from labeling images:
- Disparity and flow (2012). For each benchmark frame, the authors registered the laser scans of the 5 frames before and 5 after it with ICP, projected the accumulated points into the image, and manually removed ambiguous regions such as windows and fences. Flow comes from projecting the same 3D points into the next frame. The ground truth is not interpolated, so about 50% of pixels have a value. The scenes were chosen to be static, so all motion comes from the camera.
- Disparity and flow (2015). Menze and Geiger (2015) annotated 400 dynamic scenes from the raw data. The static background comes from 7 accumulated scans, corrected for the scanner’s rolling shutter. Moving cars could not be recovered from the laser data alone, so the authors fitted 3D CAD models, chosen from a set of 16 vehicles, to each moving vehicle using its laser points, stereo disparities, and manually selected correspondences.
- Camera poses. Odometry ground truth is the output of the GPS/IMU system.
- 3D bounding boxes. Hired annotators, not online crowd workers, labeled objects as 3D box tracklets in the laser point clouds, using a tool that also showed the camera images. Boxes carry occlusion and truncation labels.
- Later additions. Pixel masks for tracked objects (MOTS), semantic labels, manually annotated road and ego-lane areas, and depth maps.
Splits
Every benchmark has a training set with public ground truth and a test set whose ground truth is withheld. Test results come only from the evaluation server, which limits submissions and, under its current policy, accepts only novel work aimed at peer-reviewed publication; other work must be evaluated on a split of the training set. Papers therefore define their own validation splits, so “KITTI validation” numbers are not always comparable.
The official splits, from the benchmark pages:
| Benchmark | Training | Test |
|---|---|---|
| Stereo and flow 2012 | 194 image pairs | 195 image pairs |
| Stereo, flow, and scene flow 2015 | 200 scenes | 200 scenes |
| Visual odometry | 11 sequences (00–10) | 11 sequences (11–21) |
| Object detection (2D, 3D, bird’s-eye view) | 7,481 images | 7,518 images |
| Tracking, MOTS, and STEP | 21 sequences | 29 sequences |
| Road and lane | 289 images | 290 images |
The 22 odometry sequences total 39.2 km, the object benchmark contains 80,256 labeled objects, and the depth benchmark has over 93,000 depth maps aligned with the raw data.
Tasks
- Stereo matching (2012, 2015): disparity from a rectified stereo pair.
- Optical flow (2012, 2015): dense 2D motion between consecutive frames. See optical flow estimation.
- Scene flow (2015): disparity in both frames plus flow, together describing 3D motion.
- Visual odometry and SLAM (2012): the camera trajectory, from images, laser data, or both.
- Object detection (2012, with 3D and bird’s-eye-view detection added in 2017): 2D or 3D boxes with orientation.
- Multi-object tracking: car and pedestrian trajectories. MOTS (2019, Voigtlaender et al.) adds pixel masks, and KITTI-STEP (Weber et al., 2021) adds a semantic class for every pixel. See multi-object tracking.
- Road and lane estimation (2013, with Honda Research Institute Europe): the road area and the ego-lane.
- Depth completion and single-image depth prediction (2017).
- Semantic and instance segmentation (2018).
Evaluation
Each benchmark has its own metric, with evaluation code in its development kit:
- Stereo and flow. The 2012 benchmark counts pixels whose disparity error or endpoint error exceeds a threshold of 2 to 5 pixels, 3 pixels by default, separately for non-occluded pixels and for all pixels with ground truth. The 2015 benchmark counts a pixel as an outlier when its error exceeds both 3 pixels and 5% of the true value. It reports D1 and D2 for disparity in the two frames, Fl for flow, and SF for scene flow (an outlier in any of the three), each over background, foreground, and all pixels (as in Fl-all).
- Odometry. Translational error (in percent) and rotational error (in degrees per meter), averaged over all subsequences of 100 to 800 m.
- Object detection. Average precision under PASCAL-style matching, with an intersection-over-union threshold of 70% for cars and 50% for pedestrians and cyclists. Results are split into three difficulty levels by minimum box height, occlusion, and truncation: easy (at least 40 pixels tall, fully visible, at most 15% truncated), moderate (at least 25 pixels, partly occluded, at most 30%), and hard (at least 25 pixels, difficult to see, at most 50%). Methods are ranked by moderate. Geiger et al. (2012) also introduced average orientation similarity (AOS), which additionally scores each detection’s estimated orientation. Since 2019 the benchmark computes AP from 40 recall positions instead of 11.
- Tracking. Originally the CLEAR MOT metrics, including MOTA, with mostly tracked, partly tracked, and mostly lost counts. Since 25 February 2021 tracking and MOTS rank methods by HOTA, computed with the TrackEval code, and MOTA was redefined to match MOTChallenge. Reference detections are provided, so tracking-by-detection methods such as SORT can be compared on identical inputs.
- Other benchmarks. Road estimation is scored in the bird’s-eye view with the maximum F1-measure, depth completion with errors of depth and inverse depth, depth prediction with the scale-invariant logarithmic error, and semantic segmentation with intersection over union.
License
The KITTI website states that all of its datasets and benchmarks are published under the Creative Commons Attribution-NonCommercial-ShareAlike 3.0 License (as of October 2026). Use requires attribution, commercial use is not permitted, and derived works may be distributed only under the same license. Spin-off datasets such as KITTI-360 and Virtual KITTI are distributed by their own projects, so check their terms separately.
Limitations and Bias
- One region, one season, good conditions. All recordings come from the Karlsruhe area over a few days in autumn 2011, during daytime. Night, rain, snow, and other countries are absent.
- Sparse, sensor-limited ground truth. Laser-derived disparity and flow cover about half the pixels. The sky, distant structures, and hand-removed regions such as windows and fences are not evaluated, so methods are scored mainly on surfaces a laser scanner sees well.
- Rigid motion only. The 2012 scenes are static, and the 2015 moving objects are vehicles modeled as rigid CAD shapes, so non-rigid motion is not tested.
- Small test sets. A few hundred image pairs for stereo and flow and 29 tracking sequences are small by modern standards, which makes rankings sensitive to a few hard frames and invites overfitting.
- Annotation gaps. Only a few classes are evaluated, and the detection page notes that about 2% of the boxes at the hard level were not recognized by human annotators, capping recall at 98%.
- Saturation. After more than a decade of submissions and tuning on the same test sets, small differences between top entries say little about real progress.
Common Uses
- Benchmarking and fine-tuning stereo, optical flow, and scene flow networks. Flow networks such as RAFT are usually pretrained on synthetic data such as Flying Chairs, then fine-tuned on KITTI 2015 and MPI Sintel.
- Evaluating visual and LiDAR odometry and SLAM on long outdoor sequences.
- Training and evaluating 2D, 3D, and bird’s-eye-view detectors, from camera, LiDAR, or both.
- Evaluating multi-object tracking in driving scenes.
- Training self-supervised depth and ego-motion methods on the raw sequences, which need no labels.
Historical Milestones
- 2012. In the initial evaluation, the best stereo method (PCBP) left 4.72% of non-occluded pixels above 3 pixels of error, and the best flow method (TGV2CENSUS) 11.14%, errors well above those reported on Middlebury; stereo odometry with VISO2-S reached 2.2% average translation error (Geiger et al., 2012). Most flow errors occurred at large displacements.
- 2015. The 2015 stereo, flow, and scene flow benchmarks added independently moving objects (Menze and Geiger, 2015).
- 2017. 3D and bird’s-eye-view object detection and the depth completion and prediction benchmarks were added.
- 2020. Teed and Deng reported a Fl-all of 5.10% for RAFT on the KITTI 2015 test set, against 6.10% for the best published method at the time.
- 2021. Tracking and MOTS moved to HOTA as the primary metric.
Related Datasets
Cityscapes followed in 2016 with fine semantic and instance annotation of street scenes from many cities, and KITTI’s own segmentation benchmark adopted its format. nuScenes (2020) and the Waymo Open Dataset (2020) are much larger, with 360-degree sensor coverage, several cities, and more varied weather and lighting; they are now widely used for 3D detection and tracking alongside or instead of KITTI. Virtual KITTI (Gaidon et al., 2016) recreates KITTI tracking sequences in a game engine, with automatic ground truth and controlled weather and camera changes. KITTI-360, from the same group, records 73.7 km of suburban Karlsruhe with fisheye and stereo cameras and dense 2D and 3D semantic labels.
For optical flow, the main counterparts are Middlebury and the synthetic MPI Sintel; for tracking, MOTChallenge covers crowded pedestrian scenes.
Resources
- The KITTI Vision Benchmark Suite: the official site, with data downloads, development kits, sensor setup, and the leaderboards for current results.
- Raw data: synchronized and rectified sequences with tracklets and calibration (login required to download).
- KITTI-360: the successor dataset.
Related
- Optical Flow Estimation
Computing a dense field of pixel displacements between two video frames, from classical variational methods to learned networks such as RAFT.
- Multi-Object Tracking
Estimating the trajectories of a varying, unknown number of objects in a video while keeping each object's identity consistent over time.
- Endpoint Error
The standard accuracy measure for optical flow, the distance in pixels between an estimated flow vector and the true one, averaged over the image.
- Higher Order Tracking Accuracy
A multi-object tracking metric that combines detection accuracy and association accuracy through their geometric mean, averaged over localization thresholds.
- Middlebury Benchmarks
Small, high-accuracy stereo and optical flow benchmarks from Middlebury College and Microsoft Research that defined how dense correspondence methods were evaluated in the 2000s.
- MPI Sintel
An optical flow benchmark rendered from the open-source animated film Sintel, with long sequences, large motions, motion blur, and atmospheric effects.
References
- Geiger, A., Lenz, P. & Urtasun, R. (2012). Are We Ready for Autonomous Driving? The KITTI Vision Benchmark Suite. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3354–3361.
- Geiger, A., Lenz, P., Stiller, C. & Urtasun, R. (2013). Vision Meets Robotics: The KITTI Dataset. The International Journal of Robotics Research, 32(11), 1231–1237.
- Menze, M. & Geiger, A. (2015). Object Scene Flow for Autonomous Vehicles. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3061–3070.