Motion & Tracking
Object Tracking
Estimating the position, extent, or state of one or more objects in every frame of a video, keeping each object's identity over time.
intermediate
Object tracking estimates where one or more objects are in every frame of a video, and keeps track of which object is which. Where object detection answers “what is in this image and where?”, tracking adds time: it links observations across frames into trajectories, one per object. It is called visual object tracking when following a single target and multi-object tracking when following many targets with identities, and it is a core component of surveillance, autonomous driving, robotics, sports analytics, and video editing.
This article is an overview of the task. Individual methods, such as the Kalman filter and the Hungarian algorithm, have their own articles.
Problem Definition
A tracker estimates the state of each target at each time step: its position, a bounding box, a segmentation mask, a set of keypoints, or a 3D pose with velocity. The sequence of states for one object is its trajectory. Two assumptions shape most formulations:
- Temporal coherence. Objects move and change appearance gradually, so the previous state is a strong prior for the current one.
- Persistent identity. Each object keeps one identity for as long as it exists in the scene, even while temporarily hidden.
Tracking is therefore partly localization (where is the object now?) and partly correspondence (which observation belongs to which trajectory?).
Input and Output
The input is a video, sometimes with depth, lidar, or radar. In single-object tracking, the target is specified only by a bounding box in the first frame, with no category. In multi-object tracking, the categories of interest, such as pedestrians or vehicles, are fixed in advance, and a detector usually proposes candidates in every frame.
The output is a set of trajectories with consistent identity labels, most often as bounding boxes, but also as segmentation masks (video object segmentation), keypoints (pose tracking), or oriented 3D boxes with velocity (driving). An online tracker uses only frames up to the current one, as live systems must; an offline tracker sees the whole sequence and can use future frames to resolve ambiguities.
Variants
- Single-object vs multi-object. Multi-object tracking must also decide when objects enter and leave, and keep many identities apart, which makes association the central difficulty.
- Model-free vs category-specific. Model-free trackers follow an arbitrary object from one example and learn its appearance on the fly; category-specific trackers rely on a detector for known classes.
- Short-term vs long-term. Long-term tracking allows the target to disappear for many frames, so the tracker must report its absence and re-detect it when it returns.
- 2D vs 3D. 3D trackers estimate world positions, typically from lidar, stereo, or multiple cameras, where motion models are more physically meaningful.
- Point vs object tracking. Feature tracking follows small patches with no notion of objects; object tracking follows whole objects with extent and identity. Point tracks are often an ingredient of object trackers.
Challenges
- Occlusion. The tracker must predict through the gap or recover afterwards.
- Appearance change. Rotation, deformation, scale, and illumination alter the target. An appearance model that adapts too slowly loses the target; one that adapts too quickly absorbs the background.
- Fast motion and blur. Large displacements break the small-motion assumption and blur the features the tracker relies on.
- Distractors. Similar-looking nearby objects, such as other people in a crowd, pull the tracker away.
- Identity switches. In multi-object tracking, two trajectories swap identities when objects cross, even if every box is well localized.
- Drift. Updating the appearance model with slightly wrong estimates accumulates error until the tracker follows background.
- Real-time constraints. Many applications need video-rate tracking on limited hardware.
Approaches
Template matching and mean shift
The earliest trackers stored an image patch of the target and searched the next frame for the best match, by correlation or by the iterative Lucas–Kanade alignment that also underlies feature tracking. Templates are precise but brittle under appearance change. Mean shift tracking instead represents the target by a color histogram and climbs to the nearby location whose histogram matches best, which tolerates deformation but is weak against similarly colored backgrounds.
Bayesian filtering
Filtering treats tracking as recursive estimation of a hidden state: predict with a motion model, then correct with the new observation. The Kalman filter does this exactly for linear-Gaussian models and remains the standard motion model in multi-object trackers. Particle filters, such as CONDENSATION (Isard & Blake, 1998), represent the state distribution by weighted samples, handling nonlinear motion and multiple hypotheses in clutter at higher cost. Filtering answers “where should it be?”; an appearance or detection model must still answer “where is it?”
Tracking-by-detection
As detectors became reliable, multi-object tracking shifted to tracking-by-detection: detect objects in every frame, then link detections into trajectories. The linking step, data association, is often posed as an assignment problem between predicted tracks and new detections and solved with the Hungarian algorithm. SORT showed that a Kalman filter, box-overlap costs, and Hungarian matching on top of a strong detector were competitive with far more complex systems. DeepSORT added learned appearance embeddings to reduce identity switches, and ByteTrack improved recall by also associating low-confidence detections.
Discriminative correlation filters
Correlation filter trackers, beginning with MOSSE (Bolme et al., 2010) and extended by kernelized correlation filters (KCF; Henriques et al., 2015), learn a filter whose response peaks at the target. Computing it in the Fourier domain let early versions run at hundreds of frames per second. They led short-term single-object benchmarks in the early-to-mid 2010s, later with deep features.
Siamese and transformer trackers
Siamese trackers, popularized by SiamFC (Bertinetto et al., 2016), learn offline a similarity function between a target template and a search region, so no online training is needed at test time. Later versions added region-proposal and box-regression heads. Transformer trackers replace plain correlation with attention between template and search region, and many process both in a single backbone. In multi-object tracking, joint detection and tracking methods predict detections and their links to previous frames in one network, with some transformer models carrying “track queries” from frame to frame.
The arc runs from hand-designed appearance models adapted online toward representations learned offline from large video datasets. Templates could not cope with appearance change, online-learned models drifted, and learned similarity became practical once large annotated video collections, such as ImageNet VID and later LaSOT and GOT-10k, were available for training.
Datasets and Benchmarks
For single-object tracking, the OTB benchmark (Wu et al., 2013), extended to 100 sequences in 2015, established a common protocol and attribute-based analysis. The VOT challenge has run annually since 2013, later adding long-term and real-time tracks; until 2019 it reinitialized trackers after each failure, and in 2020 it moved to an anchor-based protocol. LaSOT provides long sequences, and GOT-10k uses disjoint object classes for training and testing to measure generalization to unseen categories (Huang et al., 2021). For multi-object tracking, MOTChallenge (MOT15, MOT16/17, MOT20) is the standard pedestrian benchmark. Driving datasets such as KITTI, nuScenes, and the Waymo Open Dataset cover 2D and 3D vehicle and pedestrian tracking.
Evaluation Metrics
- Success plot. The fraction of frames whose predicted box overlaps the ground truth above a threshold, swept over thresholds and summarized by the area under the curve.
- Precision plot. The fraction of frames whose predicted center lies within a given distance of the true center, commonly reported at 20 pixels.
- Expected average overlap (EAO). VOT’s summary score, combining accuracy (overlap while tracking) and robustness (failures).
- MOTA. Combines false positives, misses, and identity switches into one score, dominated by detection quality.
- IDF1. The F1 score of correctly identified detections, measuring identity consistency over whole trajectories.
- HOTA. The geometric mean of detection accuracy and association accuracy, designed to balance the two (Luiten et al., 2021).
Applications
- Autonomous driving and robotics. 3D tracks of vehicles, cyclists, and pedestrians provide velocities for prediction and planning.
- Surveillance and crowd analysis. Counting people, measuring flow, and detecting unusual movement.
- Sports analytics. Player and ball tracks yield statistics and tactical analysis.
- Video editing and augmented reality. Tracked masks and boxes anchor effects, blur faces, or attach graphics.
- Biology and medicine. Cell tracking in microscopy and instrument tracking in surgical video.
- Human–computer interaction. Hand, face, and body tracking for gesture interfaces and motion capture.
Tracking draws on motion estimation to predict where objects move, and on dense optical flow in methods that propagate masks or boxes between frames.
Open Problems
- Long-term robustness. Recovering after long occlusions or exits from view without re-detecting distractors.
- Identity preservation in crowds. Keeping many similar, interacting objects apart remains the main error source in multi-object tracking.
- Open-world tracking. Tracking unseen categories, and unifying single-object, multi-object, and segmentation tracking in one model.
- Evaluation. Rankings change with how a metric weights detection against association, and benchmarks saturate as methods tune to them.
- Efficiency. The most accurate transformer trackers are costly, while many deployments need real-time tracking on embedded hardware.
Related
- Feature Tracking
How distinctive image points are selected and followed across video frames, using the classic KLT tracker as the main example.
- Kalman Filter
A recursive algorithm that estimates the hidden state of a linear dynamic system from a sequence of noisy measurements, widely used to smooth and predict object positions in tracking.
- Motion Estimation
Recovering how image content, objects, or the camera moved from a sequence of images, from per-pixel flow to global and 3D motion.
- Data Association
Deciding which measurements or detections belong to which tracked targets, and which are false alarms, missed detections, new targets, or targets that have disappeared.
- Multi-Object Tracking
Estimating the trajectories of a varying, unknown number of objects in a video while keeping each object's identity consistent over time.
- SORT
Simple Online and Realtime Tracking, a multi-object tracker that links per-frame detections into tracks using a constant-velocity Kalman filter on each box and Hungarian matching on box overlap.
References
- Yilmaz, A., Javed, O. & Shah, M. (2006). Object Tracking: A Survey. ACM Computing Surveys, 38(4), Article 13.
- Smeulders, A. W. M., Chu, D. M., Cucchiara, R., Calderara, S., Dehghan, A. & Shah, M. (2014). Visual Tracking: An Experimental Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(7), 1442–1468.
- Wu, Y., Lim, J. & Yang, M.-H. (2013). Online Object Tracking: A Benchmark. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2411–2418.
- Milan, A., Leal-Taixé, L., Reid, I., Roth, S. & Schindler, K. (2016). MOT16: A Benchmark for Multi-Object Tracking. arXiv:1603.00831.
- Luiten, J., Ošep, A., Dendorfer, P., Torr, P., Geiger, A., Leal-Taixé, L. & Leibe, B. (2021). HOTA: A Higher Order Metric for Evaluating Multi-Object Tracking. International Journal of Computer Vision, 129(2), 548–578.
- Wu, Y., Lim, J. & Yang, M.-H. (2015). Object Tracking Benchmark. IEEE Transactions on Pattern Analysis and Machine Intelligence, 37(9), 1834–1848.
- Isard, M. & Blake, A. (1998). CONDENSATION—Conditional Density Propagation for Visual Tracking. International Journal of Computer Vision, 29(1), 5–28.
- Bolme, D. S., Beveridge, J. R., Draper, B. A. & Lui, Y. M. (2010). Visual Object Tracking Using Adaptive Correlation Filters. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2544–2550.
- Henriques, J. F., Caseiro, R., Martins, P. & Batista, J. (2015). High-Speed Tracking with Kernelized Correlation Filters. IEEE Transactions on Pattern Analysis and Machine Intelligence, 37(3), 583–596.
- Bertinetto, L., Valmadre, J., Henriques, J. F., Vedaldi, A. & Torr, P. H. S. (2016). Fully-Convolutional Siamese Networks for Object Tracking. European Conference on Computer Vision (ECCV) Workshops, Lecture Notes in Computer Science, 850–865.
- Huang, L., Zhao, X. & Huang, K. (2021). GOT-10k: A Large High-Diversity Benchmark for Generic Object Tracking in the Wild. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(5), 1562–1577.