Motion & Tracking
Tracking by Detection
Building object tracks by running a detector on every frame and linking its detections over time with a motion model and data association.
intermediate
Tracking by detection builds object trajectories in two separate steps: an object detector finds objects independently in every frame, and a tracker links those detections over time into tracks with persistent identities. The detector answers “what is where in this frame?”; the tracker answers “which of these boxes is the same object I saw before?”. This split has been the dominant design for multi-object tracking since detectors became reliable, and it is the pattern behind SORT, DeepSORT, and ByteTrack. This article describes the technique in general; those articles cover specific trackers.
Purpose
Object tracking must solve two problems at once: localizing objects and keeping their identities. Tracking by detection decouples them. Localization is delegated to a per-frame detector, which can be trained on large image datasets, improved independently, and swapped without changing the tracker. Temporal reasoning is reduced to data association: deciding which detection in frame continues which track from frame , and which detections are new objects or false alarms.
Because the detector re-finds objects in every frame, the tracker does not accumulate localization error the way a template or appearance tracker that updates itself can drift onto the background. In exchange, the tracker inherits every detector error.
When to Use It
Tracking by detection fits when:
- The object categories are known in advance, such as pedestrians, vehicles, or cells, and a detector for them exists or can be trained.
- The number of objects varies. Detections provide a natural way to start tracks for objects that enter and to notice when objects leave.
- Real-time or online operation is required. With a simple motion model and assignment step, association costs little next to the detector.
- Components should be replaceable. A better detector usually improves tracking without retraining the tracker.
It fits poorly when the target has no category and is defined only by a box in the first frame (model-free single-object tracking, where an appearance model must be learned online), when the detector is weak in the target domain, or when objects are too small, blurred, or occluded to be detected for long stretches.
How It Works
Each track carries a state, typically a box with velocities, and a short history. For every frame:
- Detect. Run the detector, producing boxes with confidence scores. Usually only boxes above a score threshold are kept.
- Predict. Propagate each track to the current frame with a motion model, most often a constant-velocity Kalman filter.
- Associate. Compute a cost between every predicted track and every detection, discard implausible pairs (gating), and solve for a one-to-one assignment, for example with the Hungarian algorithm.
- Update, create, delete. Correct matched tracks with their detections; start tentative tracks from unmatched detections; mark unmatched tracks as lost, keep them for a few frames in case the object reappears, and delete them afterwards.
for each frame:
detections = detector(frame), filtered by score
for each track: predict its state in this frame
cost[i, j] = distance(track i, detection j) # IoU, appearance, motion
cost[i, j] = infinity where the pair fails the gate
matches = assignment(cost)
update matched tracks; create tracks from unmatched detections
age unmatched tracks; delete those lost for too long
output confirmed tracks
Steps 2–4 are the tracker. They see only boxes, scores, and optionally appearance features cropped from the image.
How to Apply It
Detector choice. Tracking quality is bounded by detection quality, so the detector is usually the most important decision. Train or fine-tune it on data resembling the deployment scenes, including crowded and occluded examples, and choose its score threshold together with the tracker’s parameters rather than in isolation.
Motion model. A constant-velocity Kalman filter on box center, scale, and aspect ratio (as in SORT) works for pedestrians and vehicles at video frame rates. Nonlinear motion calls for an extended Kalman filter or a particle filter. Under strong camera motion, compensating for global image motion before prediction matters more than refining the object model.
Cost function. The common costs are:
- Overlap, between the predicted and detected boxes: cheap and reliable when frames are close in time.
- Motion distance, such as the Mahalanobis distance under the Kalman filter’s predicted covariance, which accounts for the track’s uncertainty.
- Appearance distance between embeddings from a re-identification network, which can recover identities after occlusion or across large displacements.
Combinations are common: in its reported experiments, DeepSORT ranked pairs by appearance alone and used the Mahalanobis distance only as a gate.
Gating. Reject pairs whose overlap is below a minimum or whose Mahalanobis distance exceeds a chi-squared quantile before assignment. Gating removes impossible matches and shrinks the problem; too tight a gate breaks tracks under fast motion.
Track management. Three rules shape the output as much as association does: how many consecutive matches a new track needs before it is reported (suppressing one-frame false positives), how long a lost track survives (bridging occlusions versus drifting onto the wrong object), and whether lost tracks may be re-matched by motion alone or need appearance evidence.
Variants
Online and batch. Online trackers commit to an assignment frame by frame, as SORT does. Batch, or offline, trackers associate detections over a window or a whole sequence, for example as min-cost network flow or graph partitioning, so later evidence can fix earlier mistakes; the data association article covers these formulations. Between the two, some trackers link detections into short reliable tracklets and then join tracklets.
With appearance re-identification. DeepSORT (Wojke et al., 2017) added an embedding network trained on person re-identification data. On MOT16, the authors reported about 45% fewer identity switches than SORT with the same detections. Joint models such as JDE and FairMOT produce detections and embeddings from one network (see multi-object tracking).
Using low-confidence detections. A fixed score threshold discards occluded objects whose scores drop. ByteTrack associates high-score detections first, then matches the remaining tracks to low-score detections by overlap, and discards low-score boxes left unmatched as background.
Probabilistic tracking by detection. Breitenstein et al. (2011) fed a pedestrian detector’s continuous confidence and per-person classifiers trained online into a particle filter, rather than using only thresholded boxes, tracking several people online from a single, possibly moving, uncalibrated camera. Andriluka, Roth & Schiele (2008) combined part-based people detections with a learned model of walking articulation to improve position and pose hypotheses over several frames, a direction their title calls “people-detection-by-tracking.”
Single-object tracking by detection. For one model-free target, the detector is a classifier learned online from the first frame. Tracking-Learning-Detection (TLD; Kalal et al., 2012) runs a frame-to-frame tracker alongside a detector of the target’s appearances so far: the detector corrects the tracker after failures, and a learning component uses the tracker’s trajectory to estimate the detector’s errors and update it.
Joint detection and tracking. The alternative paradigm merges the two steps in one network. Tracktor regresses each track’s box into the next frame with the detector’s own head, CenterTrack predicts detections with offsets to the previous frame, and transformer trackers carry track queries across frames, so association is learned rather than solved explicitly. The multi-object tracking article surveys these methods.
Best Practices
- Evaluate detector and tracker separately. Benchmarks such as MOTChallenge offer public detections, so trackers can be compared on identical input; private-detection results mix detector and tracker gains.
- Tune thresholds jointly on validation sequences: the detection score threshold, the gating threshold, the confirmation count, and the lost-track lifetime interact.
- Prefer overlap for short gaps and appearance for long ones. Overlap is reliable between consecutive frames; appearance helps after occlusion but is unreliable for heavily occluded or blurred boxes.
- Account for frame rate. At low frame rates, boxes of the same object may not overlap, and lost-track lifetimes measured in frames correspond to different durations.
- Report identity metrics such as IDF1 or HOTA alongside MOTA, which is dominated by detection errors.
Pitfalls
- Detector errors dominate. Missed detections become gaps or fragmented tracks, false positives become spurious tracks, and the tracker cannot recover objects the detector never finds.
- Identity switches. When objects cross or occlude each other, predicted boxes overlap the wrong detections and identities swap. Motion-only costs are most vulnerable.
- Threshold sensitivity. A high score threshold loses occluded objects; a low one admits background boxes. Settings tuned on one dataset often transfer poorly to scenes with different density or camera motion.
- Coasting drift. A lost track keeps moving with its last velocity; kept too long, it can capture an unrelated detection.
- Image-space motion models under camera motion. Panning or shaking cameras make every object appear to move, breaking constant-velocity predictions.
Evidence and Impact
The case for the technique rests on the observation that detection quality drives tracking quality. Bewley et al. (2016) showed that replacing the ACF pedestrian detector with Faster R-CNN raised SORT’s MOTA on their validation sequences from 15.1 to 34.0, and that SORT, with only a Kalman filter and Hungarian matching on box overlap, was competitive with far more complex online trackers on the 2015 MOTChallenge benchmark.
Later work kept the paradigm and improved its parts. Zhang et al. (2022) described tracking by detection as the most effective paradigm for multi-object tracking at the time. In their ablation with YOLOX detections on a MOT17 validation split (the second half of each training video), SORT reached 74.6 MOTA, DeepSORT 75.4, and their two-stage BYTE association 76.6, so the association strategy changed results by a few points once the detector was strong. They also reported that adding BYTE’s association to nine existing trackers, including joint detection-and-tracking models, improved MOTA, IDF1, and identity switches in almost all configurations; for CenterTrack, IDF1 rose from 64.2 to 74.0 on their MOT17 validation split. The boundary between the two paradigms is therefore permeable: joint models still benefit from an explicit association step.
Related
- Multi-Object Tracking
Estimating the trajectories of a varying, unknown number of objects in a video while keeping each object's identity consistent over time.
- SORT
Simple Online and Realtime Tracking, a multi-object tracker that links per-frame detections into tracks using a constant-velocity Kalman filter on each box and Hungarian matching on box overlap.
- ByteTrack
A multi-object tracker that associates high-score detections first and then matches the remaining tracks to low-score detections, recovering occluded objects that score thresholds would discard.
- Kalman Filter
A recursive algorithm that estimates the hidden state of a linear dynamic system from a sequence of noisy measurements, widely used to smooth and predict object positions in tracking.
- Hungarian Algorithm
An algorithm that finds the minimum-cost one-to-one matching between two sets, used in computer vision to match detections to tracks and predictions to ground truth.
- Particle Filter
A sequential Monte Carlo method that represents the probability distribution of a hidden state with weighted random samples, so it can track through nonlinear models and ambiguous, multimodal beliefs.
References
- Andriluka, M., Roth, S. & Schiele, B. (2008). People-Tracking-by-Detection and People-Detection-by-Tracking. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 1–8.
- Breitenstein, M. D., Reichlin, F., Leibe, B., Koller-Meier, E. & Van Gool, L. (2011). Online Multiperson Tracking-by-Detection from a Single, Uncalibrated Camera. IEEE Transactions on Pattern Analysis and Machine Intelligence, 33(9), 1820–1833.
- Kalal, Z., Mikolajczyk, K. & Matas, J. (2012). Tracking-Learning-Detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 34(7), 1409–1422.
- Bewley, A., Ge, Z., Ott, L., Ramos, F. & Upcroft, B. (2016). Simple Online and Realtime Tracking. IEEE International Conference on Image Processing (ICIP), 3464–3468.
- Wojke, N., Bewley, A. & Paulus, D. (2017). Simple Online and Realtime Tracking with a Deep Association Metric. IEEE International Conference on Image Processing (ICIP), 3645–3649.
- Zhang, Y., Sun, P., Jiang, Y., Yu, D., Weng, F., Yuan, Z., Luo, P., Liu, W. & Wang, X. (2022). ByteTrack: Multi-Object Tracking by Associating Every Detection Box. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 1–21.