Evaluation & Metrics
Multiple Object Tracking Accuracy
The CLEAR MOT accuracy score for multi-object tracking, which counts misses, false positives, and identity switches over a sequence and normalizes them by the number of ground-truth objects.
beginner
Multiple object tracking accuracy (MOTA) summarizes how many mistakes a multi-object tracker makes over a video: missed objects, outputs that match no object, and identity switches, relative to the number of ground-truth objects. Bernardin and Stiefelhagen (2008) introduced MOTA together with the multiple object tracking precision (MOTP) as the CLEAR MOT metrics, used in the CLEAR 2006 and 2007 evaluation workshops. Milan et al. (2016), introducing the MOT16 release of the MOTChallenge benchmark, called it perhaps the most widely used tracking metric.
Formula
In each frame , ground-truth objects are matched one-to-one to tracker outputs (below). With misses, false positives, mismatches (identity switches), and ground-truth objects present in frame ,
MOTChallenge writes , , , and and reports a percentage. The companion precision score averages the localization error of match over all matches, with matches in frame :
Matching. Bernardin and Stiefelhagen match frame by frame with a threshold on the distance between an object and an output:
- Keep every correspondence from the previous frame whose object and output are both still present and within of each other, even if another output is now closer.
- Match the remaining objects and outputs with the Hungarian algorithm, minimizing the total distance and allowing only pairs within .
- When a new correspondence contradicts the object’s recorded mapping, count a mismatch and update the mapping.
- Count unmatched outputs as false positives and unmatched objects as misses.
MOTChallenge counts an identity switch whenever a ground-truth object is matched to a track other than its last known assignment, a definition Milan et al. (2016) describe as stricter than the original. Bernardin and Stiefelhagen used a Euclidean distance on the ground plane with = 50 cm for 3D person tracking, and box overlap with a threshold of zero overlap for face tracking; MOTChallenge requires an intersection over union of at least 0.5.
Intuition
MOTA is one minus the sum of three error rates, each relative to the number of ground-truth objects: the miss rate, the false-positive rate, and the mismatch rate. A mismatch is charged once, in the frame where the identity changes, not in every later frame, so a single swap costs the same whether it happens early or late in a sequence.
Errors are summed before dividing. If four objects are missed for four frames and one remaining object is then tracked perfectly for four frames, averaging per-frame miss rates gives 50%, while the global ratio gives 16 misses out of 20 objects, 80% (Bernardin & Stiefelhagen, 2008).
Interpretation
- Range: , often given as a percentage. A value of 1 means no misses, false positives, or switches; higher is better.
- MOTA can be negative, when the tracker makes more errors than there are ground-truth objects, which false positives alone can cause. In the CLEAR 2006 single-person task, one system scored −79.78%: its estimates were mostly more than 50 cm from the person, and each such frame counted as both a miss and a false positive.
- MOTP measures localization only. It is a mean distance (lower is better) or a mean overlap (higher is better). With an overlap threshold of 0.5, as in MOTChallenge, MOTP lies between 50% and 100%, and Milan et al. note that it mostly reflects the detector’s box accuracy.
When to Use It
MOTA suits applications where finding every object in every frame matters more than keeping identities, such as collision avoidance, and for comparison with published results. Report it together with its components (misses, false positives, switches) and with an identity metric such as IDF1.
Failure Cases
- Dominated by detection. Misses and false positives usually far outnumber identity switches: in the CLEAR 3D person tracking evaluations, mismatch ratios were often below 2%, against 20–40% for misses or false positives. Luiten et al. (2021), who proposed HOTA as a replacement, found that for the top ten MOT17 trackers in April 2020, detection errors outnumbered switches by a factor of 42 to 186, and that across 175 MOT17 trackers MOTA without its switch term explained over 99% of the variation in MOTA.
- Short-term association only. A switch compares an object’s current output only with its last one. A tracker that makes an identity error and later returns to the correct identity is charged twice, once for the error and once for the correction, so it scores worse than a tracker that keeps the wrong identity (Luiten et al., 2021). Nothing rewards keeping one identity for a whole trajectory.
- Frame-rate dependence. One identity switch in a 2.5-second clip lowers MOTA to 0.99 at 40 frames per second but to 0.90 at 4 frames per second, because the number of ground-truth objects grows with the frame count (Luiten et al., 2021).
- Threshold dependence. Localization enters only through . An infinite threshold keeps every correspondence valid and hides swaps, while a threshold near zero turns every object into a miss, so must be chosen per task (Bernardin & Stiefelhagen, 2008).
Variants
- Track quality measures. MOTChallenge also reports mostly tracked (MT) targets, covered for at least 80% of their life span regardless of identity, mostly lost (ML) targets, covered for less than 20%, and fragmentations (FM), the number of times a ground-truth trajectory goes from tracked to untracked and is later resumed. Milan et al. (2016) attribute these to Wu and Nevatia.
- Dropping the switch term. Without , the score measures detection alone and is known as the multiple object detection accuracy (MODA). The acoustic speaker tracking task in CLEAR 2007 used the same idea, called A-MOTA, because speaker identities were not required.
Related Metrics
- IDF1 (Ristani et al., 2016) matches whole trajectories instead of single frames and measures how long each object keeps the correct identity.
- HOTA (Luiten et al., 2021) combines a detection accuracy and an association accuracy with equal weight and averages over localization thresholds, addressing MOTA’s detection bias and fixed threshold.
Related
- IDF1
An identity-based score for multi-object tracking that matches whole ground-truth and predicted trajectories one-to-one and reports the F1 score of correctly identified detections.
- Higher Order Tracking Accuracy
A multi-object tracking metric that combines detection accuracy and association accuracy through their geometric mean, averaged over localization thresholds.
- MOTChallenge
A benchmark for multiple pedestrian tracking, released as MOT15, MOT16, MOT17, and MOT20, with hidden test annotations and a central evaluation server.
- Hungarian Algorithm
An algorithm that finds the minimum-cost one-to-one matching between two sets, used in computer vision to match detections to tracks and predictions to ground truth.
- SORT
Simple Online and Realtime Tracking, a multi-object tracker that links per-frame detections into tracks using a constant-velocity Kalman filter on each box and Hungarian matching on box overlap.
- Tracking by Detection
Building object tracks by running a detector on every frame and linking its detections over time with a motion model and data association.
References
- Bernardin, K. & Stiefelhagen, R. (2008). Evaluating Multiple Object Tracking Performance: The CLEAR MOT Metrics. EURASIP Journal on Image and Video Processing, 2008, Article 246309.
- Milan, A., Leal-Taixé, L., Reid, I., Roth, S. & Schindler, K. (2016). MOT16: A Benchmark for Multi-Object Tracking. arXiv:1603.00831.
- Luiten, J., Ošep, A., Dendorfer, P., Torr, P., Geiger, A., Leal-Taixé, L. & Leibe, B. (2021). HOTA: A Higher Order Metric for Evaluating Multi-Object Tracking. International Journal of Computer Vision, 129(2), 548–578.
- Ristani, E., Solera, F., Zou, R., Cucchiara, R. & Tomasi, C. (2016). Performance Measures and a Data Set for Multi-Target, Multi-Camera Tracking. European Conference on Computer Vision (ECCV) Workshops, Lecture Notes in Computer Science, 17–35.