Datasets & Benchmarks
MOTChallenge
A benchmark for multiple pedestrian tracking, released as MOT15, MOT16, MOT17, and MOT20, with hidden test annotations and a central evaluation server.
beginner
MOTChallenge is a benchmark for multi-object tracking of pedestrians in video. Laura Leal-Taixé, Anton Milan, Ian Reid, Stefan Roth, and Konrad Schindler launched it in October 2014. It combined existing and new video sequences with fixed training and test splits, a set of precomputed detections, and one evaluation server that scored every method the same way. Its pedestrian releases, MOT15, MOT16, MOT17, and MOT20, became the standard way to compare multiple-people trackers. As of October 2026 the website and evaluation server are offline, while the dataset archives and a static copy of the leaderboards remain available.
Purpose
Before MOTChallenge, tracking papers evaluated on different subsets of sequences such as PETS, trained models in different ways, and used different evaluation scripts, so their numbers could not be compared (Leal-Taixé et al., 2015). The benchmark fixed all three. Test annotations are hidden, so methods cannot be tuned on the test data. Shared public detections separate the tracker’s contribution from the detector’s, as tracking by detection requires.
Dataset Size
| Release | Sequences (train / test) | Frames | Pedestrian boxes | Public detections |
|---|---|---|---|---|
| MOT15 | 11 / 11 | 11,286 | about 101,000 | ACF |
| MOT16 | 7 / 7 | 11,235 | about 293,000 | DPM v5 |
| MOT17 | 7 / 7 (the MOT16 videos) | 11,235 | about 300,000 | DPM, Faster R-CNN, SDP |
| MOT20 | 4 / 4 | 13,410 | about 1.65 million | Faster R-CNN |
Counts are from the release papers (Dendorfer et al., 2020, 2021). Crowding grows with each release. The mean density is about 7 to 11 pedestrians per frame in MOT15 and 21 to 32 in MOT17. In MOT20 it exceeds 100, and some frames hold up to 246 people. MOT15 gathered sequences from earlier datasets (TUD, PETS 2009, ETH, TownCentre, and KITTI) and added six new ones. Frame rates range from 2.5 to 30 frames per second. Apart from two sequences taken from the ETH dataset, MOT16’s videos were recorded by the benchmark’s authors. MOT20 was filmed in three scenes, indoors and outdoors, by day and night.
Classes
The tracking target is the pedestrian. From MOT16 on, annotators also labeled people on vehicles, static (sitting or lying) people, distractors such as mannequins, posters, and reflections, cars, bicycles, motorbikes, and occluders such as pillars and trees. The person-like classes are neither rewarded nor penalized in the evaluation, so a tracker that follows a sitting person is not charged a false positive. Vehicles and occluders are used to compute how occluded each pedestrian is (Dendorfer et al., 2021).
Annotation Types
Each object has a bounding box in every frame and an identity that persists while the person remains in view. MOT15 inherited annotations of uneven quality from its sources. For MOT16, researchers annotated every sequence from scratch under one written protocol:
- Extent. The box covers the whole person, including parts that are occluded or outside the image, as tightly as possible.
- Duration. A track starts as soon as the person’s position can be determined, typically when about 10% is visible, and ends as late as possible.
- Occlusion. Annotation continues through occlusions while the position can be inferred unambiguously. A person who reappears after a long, ambiguous occlusion, or who re-enters the frame, receives a new identity.
- Visibility. A visibility ratio between 0 and 1 is computed automatically for every box from the overlapping annotations.
- Sanity check. After annotation, a pedestrian detector was run on every video, and confident detections of missed people or distractors were added.
MOT17 relabeled the same videos with more accurate boxes, added missed pedestrians and occluders, and incorporated feedback from benchmark users (Dendorfer et al., 2021). MOT20 followed the same protocol. Results and annotations use the MOTChallenge text format described in the multi-object tracking article.
Splits
Every release splits its sequences in half, balancing camera motion, viewpoint, and lighting conditions between the two halves. Training annotations are public. Test annotations were never released, and test results came only from the evaluation server. MOT20 reserves one of its three scenes for the test set to measure generalization.
To discourage tuning on the test data, each method was expected to submit once per benchmark and tune on the training sequences. A resubmission required a 72-hour wait, with at most four submissions per tracker and benchmark, and multiple accounts were forbidden (Dendorfer et al., 2021).
Tasks
The core task is 2D multiple-pedestrian tracking from monocular video. Two protocols exist:
- Public detections. The tracker must use the provided detections. In MOT17, every tracker runs on all three detection sets and the scores are averaged, which rewards robustness to detectors of different quality.
- Private detections. The tracker brings its own detector, and the leaderboard marked such entries. Joint detection-and-tracking models, which contain their own detector, report results under this protocol, and ByteTrack reports both.
MOT17Det and MOT20Det offer the same frames as pedestrian detection benchmarks.
Evaluation
Tracker output is matched to ground truth with an intersection-over-union threshold of 0.5 and the Hungarian algorithm, and metrics are computed over all test sequences concatenated rather than averaged per sequence. The original reports included the CLEAR MOT metrics, MOTA and MOTP; mostly tracked, partially tracked, and mostly lost counts; fragmentations; and IDF1. MOTA was the ranking metric, although the organizers recommended reading it alongside the identity-based measures (Dendorfer et al., 2021). On 26 March 2021 the benchmark began reporting HOTA, which balances detection and association, and TrackEval, the reference implementation of HOTA, became the official evaluation code.
License
An archived copy of motchallenge.net from 1 January 2026 states that the datasets are published under the Creative Commons Attribution-NonCommercial-ShareAlike 3.0 License (CC BY-NC-SA 3.0): attribution is required, commercial use is not allowed, and derived works must carry the same license. As of October 2026 the live site is offline and no longer displays this notice, so the archived statement is the most recent official one. Many sequences come from earlier datasets, and the benchmark asks users to cite those original sources as well.
Limitations and Bias
- Pedestrians only. The benchmark measures people tracking in urban and indoor scenes and says little about vehicles, animals, or arbitrary categories.
- Limited association difficulty. The authors of DanceTrack argued that in existing tracking datasets most objects have distinctive appearance, so re-identification usually suffices for association, and built their benchmark of look-alike dancers to stress association instead (see multi-object tracking).
- Few videos. MOT17 has seven test sequences and MOT20 four, so a few long, crowded sequences dominate the totals.
- Label noise. MOT15’s collected annotations were inaccurate in places, especially with moving cameras, and MOT17 relabeled the MOT16 videos. The organizers acknowledged remaining annotation errors and invited reports.
- Benchmark overfitting. MOT15 included well-known sequences such as PETS09-S2L1, to which methods had already been tuned (Dendorfer et al., 2021). Luiten et al. (2021) also argued that the ranking metric shapes research, and that MOTA, the primary metric for years, rewards detection far more than association.
- Detector confound. Private-detection results mix detector and tracker improvements, and strong entries often train their detectors on extra pedestrian datasets.
Common Uses
MOT17 and MOT20 are the default benchmarks for pedestrian trackers, from motion-based ones such as SORT and ByteTrack to re-identification and transformer models. The MOT17 training set also serves for ablations, often split in half by frames for validation, as in the ByteTrack paper.
Historical Milestones
- 2014–2015. MOT15 released in October 2014. At the first benchmark workshop, at WACV 2015, the best method reached a MOTA of only 12.7% on the six new sequences (Dendorfer et al., 2021).
- 2016. SORT, with a Kalman filter, the Hungarian algorithm, and a Faster R-CNN detector, reported the highest MOTA among the online trackers its authors compared on MOT15. MOT16 was released with the new annotation protocol.
- 2017. MOT17 introduced three public detection sets. Crowded sequences shown at a CVPR 2019 workshop later became MOT20 (2020). DeepSORT reported about 45% fewer identity switches than SORT on MOT16 by adding appearance features.
- 2021. The benchmark began reporting HOTA alongside MOTA and IDF1.
- 2022. ByteTrack’s authors reported 80.3 MOTA, 77.3 IDF1, and 63.1 HOTA on the MOT17 test set with private detections, which they reported as first on the leaderboard at the time.
- 2026. The website and evaluation server went offline. The leaderboards survive as a static archive of all published results as of 16 April 2026.
Related Datasets
The KITTI tracking benchmark covers cars and pedestrians from a driving vehicle. The MOTChallenge platform also hosted 3D MOT 2015, MOTS (tracking with pixel masks), Head Tracking 21, 3D zebrafish (3D-ZeF20) and cell (CTMC-v1) tracking, TAO, STEP, and the synthetic MOTSynth. DanceTrack, BDD100K, nuScenes, and the Waymo Open Dataset cover other domains.
Resources
- Official site: motchallenge.net. As of October 2026, the site returns only a notice (HTTP 410) that the service is offline: the evaluation server, submissions, and user accounts no longer operate. The dataset archives remain at their published download addresses.
- Leaderboards: a static results archive preserves all published results as of 16 April 2026, with per-sequence scores and CSV exports for each benchmark.
- Evaluation code: TrackEval computes HOTA, CLEAR MOT, and identity metrics in the MOTChallenge format, so the training sets can still be evaluated locally.
Related
- Multiple Object Tracking Accuracy
The CLEAR MOT accuracy score for multi-object tracking, which counts misses, false positives, and identity switches over a sequence and normalizes them by the number of ground-truth objects.
- IDF1
An identity-based score for multi-object tracking that matches whole ground-truth and predicted trajectories one-to-one and reports the F1 score of correctly identified detections.
- Higher Order Tracking Accuracy
A multi-object tracking metric that combines detection accuracy and association accuracy through their geometric mean, averaged over localization thresholds.
- Tracking by Detection
Building object tracks by running a detector on every frame and linking its detections over time with a motion model and data association.
- SORT
Simple Online and Realtime Tracking, a multi-object tracker that links per-frame detections into tracks using a constant-velocity Kalman filter on each box and Hungarian matching on box overlap.
- ByteTrack
A multi-object tracker that associates high-score detections first and then matches the remaining tracks to low-score detections, recovering occluded objects that score thresholds would discard.
References
- Leal-Taixé, L., Milan, A., Reid, I., Roth, S. & Schindler, K. (2015). MOTChallenge 2015: Towards a Benchmark for Multi-Target Tracking. arXiv:1504.01942.
- Milan, A., Leal-Taixé, L., Reid, I., Roth, S. & Schindler, K. (2016). MOT16: A Benchmark for Multi-Object Tracking. arXiv:1603.00831.
- Dendorfer, P., Rezatofighi, H., Milan, A., Shi, J., Cremers, D., Reid, I., Roth, S., Schindler, K. & Leal-Taixé, L. (2020). MOT20: A Benchmark for Multi Object Tracking in Crowded Scenes. arXiv:2003.09003.
- Dendorfer, P., Ošep, A., Milan, A., Schindler, K., Cremers, D., Reid, I., Roth, S. & Leal-Taixé, L. (2021). MOTChallenge: A Benchmark for Single-Camera Multiple Target Tracking. International Journal of Computer Vision, 129(4), 845–881.
- Luiten, J., Ošep, A., Dendorfer, P., Torr, P., Geiger, A., Leal-Taixé, L. & Leibe, B. (2021). HOTA: A Higher Order Metric for Evaluating Multi-Object Tracking. International Journal of Computer Vision, 129(2), 548–578.