Motion & Tracking
Multi-Object Tracking
Estimating the trajectories of a varying, unknown number of objects in a video while keeping each object's identity consistent over time.
intermediate
Multi-object tracking (MOT) follows every object of the categories of interest, such as pedestrians or vehicles, through a video and assigns each one a persistent identity. Unlike single-object tracking, where one target is given in the first frame, a multi-object tracker must discover objects by itself, decide when they enter and leave, and keep them apart when they cross, occlude each other, or look alike. It is the branch of object tracking behind traffic analysis, crowd monitoring, sports statistics, and the perception stacks of autonomous vehicles. This article surveys the task; its building blocks, such as the Kalman filter, the Hungarian algorithm, and data association, have their own articles.
Problem Definition
Given a video, a multi-object tracker estimates a set of trajectories, one per object, where each trajectory is a sequence of states (usually bounding boxes) labeled with a single identity. The number of objects is unknown and changes over time, so the tracker must handle track birth, when a new object appears, and track death, when an object leaves the scene, as well as temporary gaps while an object is occluded or missed by the detector.
Trackers are either online or offline. An online tracker produces the output for frame using only frames up to , as robots and vehicles require; once it commits to an association it cannot revise it. An offline, or batch, tracker sees the whole sequence (or a window of it) before deciding, and can use future evidence to resolve ambiguities, at the cost of latency.
Input and Output
The input is a video, sometimes with depth, lidar, or radar. In the common tracking-by-detection setting, it also includes per-frame detections from an object detector, and benchmarks often provide fixed “public” detections so that trackers can be compared independently of the detector.
The output is a list of boxes, each with a frame number and a track identity. The MOTChallenge format, widely reused, is a text file with one line per box: frame, track ID, the box’s left, top, width, and height, a confidence score, and optional 3D coordinates. Identity numbers mean nothing beyond consistency.
Variants
- Pedestrian and vehicle MOT in surveillance and driving video is the classical setting, with boxes in image coordinates.
- 3D MOT tracks objects as 3D boxes in world coordinates, usually from lidar point clouds or multi-camera rigs on vehicles, where metric positions make motion models more reliable.
- Multi-camera MOT keeps identities across several cameras with overlapping or disjoint views, and must match people who leave one view and reappear in another.
- Multi-object tracking and segmentation (MOTS) replaces boxes with pixel masks, which removes the ambiguity of overlapping boxes in crowds.
- Large- and open-vocabulary tracking extends tracking beyond a few categories: the TAO (Tracking Any Object) benchmark (Dave et al., 2020) annotates objects from hundreds of categories, and open-vocabulary trackers must also handle classes unseen in training.
Challenges
- Occlusion. Objects disappear behind others or behind scenery for many frames and must be re-identified when they reappear.
- Crowded scenes. Dense groups produce overlapping boxes, missed detections, and many plausible associations.
- Identity switches. Two tracks exchange identities when their objects pass close to each other.
- Fragmentation. One object’s trajectory breaks into several tracks, typically after a long occlusion.
- Similar appearance. Team uniforms, dancers, or uniformly dressed workers make appearance cues unreliable.
- Camera motion. A moving or shaking camera breaks constant-velocity motion models defined in image coordinates.
- Detector errors. False positives create spurious tracks, and false negatives create gaps; tracking cannot recover objects the detector never finds.
Approaches
Tracking-by-detection with motion models. The dominant paradigm detects objects in every frame, predicts each track’s next position, and solves an assignment between predictions and detections. SORT (Bewley et al., 2016) reduced this to a Kalman filter with a constant-velocity model, an overlap-based cost, and the Hungarian algorithm, and showed that with a strong detector this simple pipeline was competitive with far more complex trackers.
Appearance re-identification. Motion alone fails across occlusions and crossings. DeepSORT (Wojke et al., 2017) added appearance embeddings from a person re-identification network, combining appearance and motion distances so that tracks could be recovered after longer gaps with fewer identity switches.
Batch and global methods. Offline trackers formulate association over the whole sequence, for example as a min-cost network flow in which each unit of flow is one trajectory, or as a graph partitioning or graph neural network problem over detections. Optimizing globally lets later evidence correct earlier ambiguity; the data association article covers these formulations.
Associating low-confidence detections. Most trackers discard detections below a score threshold, losing occluded objects. ByteTrack (Zhang et al., 2022) first matches high-score detections, then matches the remaining tracks to low-score detections, recovering partially occluded objects while discarding unmatched low-score boxes as background.
Joint detection and tracking. One network both detects and tracks. Tracktor (Bergmann et al., 2019) uses a detector’s box regression head to move each track’s box into the next frame, without training specifically for tracking. CenterTrack (Zhou et al., 2020) takes two consecutive frames and the previous frame’s object centers as input, and predicts detections with offsets linking them to the previous frame. JDE (Wang et al., 2020) and FairMOT (Zhang et al., 2021) add a re-identification embedding branch to the detector, so boxes and appearance features come from one forward pass.
Transformer and query-based tracking. TrackFormer (Meinhardt et al., 2022) and MOTR (Zeng et al., 2022) extend the DETR detection transformer with track queries: each query represents one object, is carried from frame to frame, and outputs that object’s box in each new frame, while separate queries detect new objects. Association becomes implicit in attention rather than a separate matching step.
The arc runs from hand-designed motion models and explicit assignment toward learned appearance, motion, and association. Simple tracking-by-detection pipelines remain strong baselines, because detector quality drives much of the final accuracy.
Datasets and Benchmarks
MOTChallenge is the standard pedestrian benchmark. MOT15 gathered mostly existing sequences, MOT16 added new, more challenging videos, MOT17 kept the MOT16 sequences with more accurate labels and public detections from three detectors (Dendorfer et al., 2021), and MOT20 added eight sequences of very crowded scenes (Dendorfer et al., 2020). The KITTI tracking benchmark covers cars and pedestrians from a driving platform, and BDD100K provides large-scale driving video with box and segmentation tracking. DanceTrack (Sun et al., 2022) targets association rather than detection: its dancers look similar and move in complex, non-linear patterns, so appearance and simple motion models are both weak. For 3D tracking, nuScenes and the Waymo Open Dataset provide lidar and camera data with annotated 3D tracks.
Evaluation Metrics
Each metric first matches tracker output to ground truth, then counts errors; they differ in what they count and so in what they reward.
CLEAR MOT. Bernardin & Stiefelhagen (2008) matched objects to hypotheses frame by frame, keeping earlier correspondences where still valid, and defined the multiple object tracking accuracy
where , , and are the misses, false positives, and mismatches (identity switches) in frame , and is the number of ground-truth objects present. MOTA can be negative. The multiple object tracking precision (MOTP) is the average distance, or overlap, between matched pairs, and measures localization only. Because misses and false positives usually far outnumber identity switches, MOTA is dominated by detection quality.
IDF1. Ristani et al. (2016) match whole ground-truth trajectories to whole predicted trajectories one-to-one, then count identity true positives, false positives, and false negatives:
IDF1 rewards keeping the same identity over a trajectory’s whole length, so it emphasizes association.
HOTA. Luiten et al. (2021) argued that MOTA overweights detection and IDF1 overweights association. HOTA is the geometric mean of a detection accuracy and an association accuracy, , computed at 19 localization thresholds and averaged over them, so detection, association, and localization all contribute. It decomposes into sub-metrics for analyzing each error type.
Applications
- Autonomous driving. 3D tracks of vehicles, cyclists, and pedestrians give the velocities and histories that motion prediction and planning need.
- Surveillance and crowd analysis. Counting people, measuring flows, and analyzing movement in public spaces.
- Traffic monitoring. Counting vehicles and measuring speeds and turning movements from roadside cameras.
- Sports analytics. Player tracks yield distance, speed, and formation statistics.
- Biology. Tracking cells or animals to study behavior and lineage.
Open Problems
- Long-term identity. Re-identifying objects after long occlusions or exits, and across cameras, without confusing similar-looking objects.
- Crowds and look-alikes. Association remains brittle when appearance carries little information.
- Open-world tracking. Tracking arbitrary categories, including ones absent from training, with consistent identities.
- End-to-end learning. Query-based trackers learn association directly but have not consistently replaced simpler pipelines built on strong detectors.
- Evaluation. Rankings change with the metric, and no single number captures every failure mode that matters to an application.
- Efficiency. Many deployments need real-time tracking on embedded hardware.
Related
- SORT
Simple Online and Realtime Tracking, a multi-object tracker that links per-frame detections into tracks using a constant-velocity Kalman filter on each box and Hungarian matching on box overlap.
- Kalman Filter
A recursive algorithm that estimates the hidden state of a linear dynamic system from a sequence of noisy measurements, widely used to smooth and predict object positions in tracking.
- Hungarian Algorithm
An algorithm that finds the minimum-cost one-to-one matching between two sets, used in computer vision to match detections to tracks and predictions to ground truth.
References
- Bernardin, K. & Stiefelhagen, R. (2008). Evaluating Multiple Object Tracking Performance: The CLEAR MOT Metrics. EURASIP Journal on Image and Video Processing, 2008, Article 246309.
- Ristani, E., Solera, F., Zou, R., Cucchiara, R. & Tomasi, C. (2016). Performance Measures and a Data Set for Multi-Target, Multi-Camera Tracking. European Conference on Computer Vision (ECCV) Workshops, Lecture Notes in Computer Science, 17–35.
- Luiten, J., Ošep, A., Dendorfer, P., Torr, P., Geiger, A., Leal-Taixé, L. & Leibe, B. (2021). HOTA: A Higher Order Metric for Evaluating Multi-Object Tracking. International Journal of Computer Vision, 129(2), 548–578.
- Dendorfer, P., Ošep, A., Milan, A., Schindler, K., Cremers, D., Reid, I., Roth, S. & Leal-Taixé, L. (2021). MOTChallenge: A Benchmark for Single-Camera Multiple Target Tracking. International Journal of Computer Vision, 129(4), 845–881.
- Bewley, A., Ge, Z., Ott, L., Ramos, F. & Upcroft, B. (2016). Simple Online and Realtime Tracking. IEEE International Conference on Image Processing (ICIP), 3464–3468.
- Wojke, N., Bewley, A. & Paulus, D. (2017). Simple Online and Realtime Tracking with a Deep Association Metric. IEEE International Conference on Image Processing (ICIP), 3645–3649.
- Bergmann, P., Meinhardt, T. & Leal-Taixé, L. (2019). Tracking Without Bells and Whistles. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 941–951.
- Zhou, X., Koltun, V. & Krähenbühl, P. (2020). Tracking Objects as Points. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 474–490.
- Zhang, Y., Sun, P., Jiang, Y., Yu, D., Weng, F., Yuan, Z., Luo, P., Liu, W. & Wang, X. (2022). ByteTrack: Multi-Object Tracking by Associating Every Detection Box. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 1–21.
- Wang, Z., Zheng, L., Liu, Y., Li, Y. & Wang, S. (2020). Towards Real-Time Multi-Object Tracking. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 107–122.
- Zhang, Y., Wang, C., Wang, X., Zeng, W. & Liu, W. (2021). FairMOT: On the Fairness of Detection and Re-identification in Multiple Object Tracking. International Journal of Computer Vision, 129(11), 3069–3087.
- Meinhardt, T., Kirillov, A., Leal-Taixé, L. & Feichtenhofer, C. (2022). TrackFormer: Multi-Object Tracking with Transformers. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 8834–8844.
- Zeng, F., Dong, B., Zhang, Y., Wang, T., Zhang, X. & Wei, Y. (2022). MOTR: End-to-End Multiple-Object Tracking with Transformer. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 659–675.
- Dendorfer, P., Rezatofighi, H., Milan, A., Shi, J., Cremers, D., Reid, I., Roth, S., Schindler, K. & Leal-Taixé, L. (2020). MOT20: A Benchmark for Multi Object Tracking in Crowded Scenes. arXiv:2003.09003.
- Dave, A., Khurana, T., Tokmakov, P., Schmid, C. & Ramanan, D. (2020). TAO: A Large-Scale Benchmark for Tracking Any Object. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 436–454.
- Sun, P., Cao, J., Jiang, Y., Yuan, Z., Bai, S., Kitani, K. & Luo, P. (2022). DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse Motion. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 20961–20970.