Motion & Tracking
Simple Online and Realtime Tracking with a Deep Association Metric
The 2017 paper by Wojke, Bewley, and Paulus that introduced DeepSORT, extending SORT with a CNN appearance descriptor trained for person re-identification and a matching cascade, and reporting about 45% fewer identity switches on MOT16.
intermediate
“Simple Online and Realtime Tracking with a Deep Association Metric” by Nicolai Wojke and Dietrich Paulus of the University of Koblenz-Landau and Alex Bewley of Queensland University of Technology was presented at the IEEE International Conference on Image Processing (ICIP) in 2017, after appearing on arXiv in March of that year. It introduced the tracker known as DeepSORT: SORT with an appearance descriptor learned offline from person re-identification data, a gallery of each track’s recent appearances, and a matching cascade that favors recently seen tracks. Bewley was the first author of the original SORT paper. DeepSORT has no separate article; the SORT article describes it among SORT’s variants.
Problem
SORT had shown that a Kalman filter on each box and Hungarian matching on box overlap, fed by a strong detector, could compete with far more complex online trackers. Its weakness was identity. In the SORT paper’s own benchmark table, it had more identity switches than any other online tracker listed. Wojke and colleagues traced this to the association metric, which is reliable only while the state estimate is certain. Once an object is occluded, the prediction drifts, and when the object reappears its detection may no longer match the old track. The authors describe such occlusions as typical of frontal-view camera scenes.
The authors set aside batch methods, which optimize over whole videos and cannot report identities online, and Multiple Hypothesis Tracking and joint probabilistic data association, which work frame by frame at a cost SORT had avoided. The goal was to keep identities through longer occlusions while staying online, simple, and fast.
Contribution
The paper replaces SORT’s association metric with one that combines motion and appearance:
- A learned appearance descriptor. A convolutional network, trained offline to tell pedestrians apart, maps each detected box to a unit-length 128-dimensional vector, so the expensive learning happens before tracking and online association reduces to nearest-neighbor queries.
- An appearance gallery per track. Each track stores the descriptors of its last 100 associated detections, and a detection’s appearance distance to a track is the smallest cosine distance to any of them, so an identity can be recovered after a long occlusion.
- Motion as a gate. The squared Mahalanobis distance between a detection and the track’s predicted box, thresholded at the 95% quantile of the chi-squared distribution with four degrees of freedom, rules out physically implausible pairs.
- A matching cascade. Tracks are matched in order of how recently they were last associated.
- Open code. The authors released their implementation and a pre-trained network (github.com/nwojke/deep_sort).
Method
Track handling follows SORT, with changes. The Kalman state is eight-dimensional: box center, aspect ratio, and height, plus their velocities, under a constant-velocity model. A track is deleted after frames without a match, 30 in the experiments, compared with SORT’s single frame, which gives the appearance model time to bridge an occlusion. New tracks are tentative for their first three frames and are deleted if they miss a match during that period.
The cost of pairing track with detection is a weighted sum of the Mahalanobis distance and the appearance distance . A pair is admissible only if it passes both gates; the appearance gate’s threshold was set on separate training data. The authors found a reasonable choice under substantial camera motion, and used it for the reported results, so in practice appearance alone ranked the candidates and motion served only as a gate.
The matching cascade addresses a property of the Mahalanobis distance that the authors call counterintuitive. As a track coasts unobserved, its predicted covariance grows, and the distance, measured in standard deviations, shrinks for every detection. In a competition for a detection, the most uncertain track would tend to win. The cascade therefore solves a sequence of assignment problems: first for tracks last matched one frame ago, then for those last matched two frames ago, and so on up to , each on the detections still unassigned. A final round applies SORT’s IoU matching to unconfirmed tracks and to unmatched tracks last seen one frame earlier, to absorb sudden appearance changes, for example from partial occlusion by static scene structure, and erroneous initialization.
The descriptor network is a wide residual network, two convolutional layers followed by six residual blocks, with about 2.8 million parameters. It was trained on MARS, a person re-identification dataset that the paper describes as holding over 1.1 million images of 1,261 pedestrians. The paper leaves the training procedure out of scope. One forward pass on 32 boxes took about 30 ms on a laptop GPU (GeForce GTX 1050 mobile).
Results
The evaluation used the seven test sequences of MOT16, part of the MOTChallenge benchmark, with detections from a Faster R-CNN detector that Yu et al. had trained on public and private data for their POI tracker, thresholded at a confidence of 0.3. For a fair comparison, the authors re-ran SORT on the same detections.
| Tracker | MOTA | MOTP | Mostly tracked | Mostly lost | ID switches | Fragmentations | False positives |
|---|---|---|---|---|---|---|---|
| SORT | 59.8 | 79.6 | 25.4% | 22.7% | 1,423 | 1,835 | 8,698 |
| DeepSORT | 61.4 | 79.1 | 32.8% | 18.2% | 781 | 2,008 | 12,852 |
Identity switches fell from 1,423 to 781, the roughly 45% reduction in the abstract. Mostly tracked trajectories rose and mostly lost ones fell, while fragmentations rose slightly, which the authors attributed to maintaining identities through occlusions and misses. Among the online methods in the paper’s table, DeepSORT had the fewest identity switches; its MOTA was below that of POI (66.1), the tracker whose detections it used.
The authors attributed most of the MOTA gap to false positives, mostly sporadic detections on static scene structure that the long track lifetime joined into stable, stationary tracks; they suggested that a higher detection threshold could raise MOTA. The paper’s two speed figures disagree: the text reports about 20 Hz for the full system, with roughly half the time spent computing features, while the results table lists 40 Hz for DeepSORT (and 60 Hz for SORT). The paper does not reconcile them.
Impact
DeepSORT showed that SORT’s main weakness could be addressed without giving up its structure: keep the Kalman filter and per-frame assignment, and add an appearance model trained separately on re-identification data. The released code and pre-trained model made the tracker easy to adopt, and the paper has been widely cited.
Later work treated it as a reference point. The ByteTrack authors compared against it as one of the popular association methods, and their official code reuses DeepSORT’s Kalman filter and its box parameterization; StrongSORT took DeepSORT as its starting point.
Limitations
Several limits are visible in the paper itself. The evaluation covers a single benchmark, MOT16, with one set of private detections, so the comparison with SORT isolates the association changes but the comparison with other entries mixes detectors. The weighted motion–appearance cost is presented as the general formulation, but the reported results use , so results for the combined cost were not reported. The appearance network is specific to pedestrians, the paper omits how it was trained, and the authors note that real-time operation assumes a modern GPU. The longer track lifetime that bridges occlusions also raised the false-positive count.
Later work questioned parts of the design as detectors improved. With a modern YOLOX detector on a MOT17 validation split, the ByteTrack authors found that their motion-only association beat DeepSORT on MOTA, IDF1, and identity switches, and noted that appearance features become unreliable under severe occlusion. The StrongSORT authors argued that the matching cascade’s prior constraint limits matching accuracy once the rest of the tracker is strong.
What Came After
StrongSORT (Du et al., arXiv 2022, IEEE Transactions on Multimedia 2023) revisited DeepSORT component by component. It replaced the detector with YOLOX-X, following ByteTrack, and the small CNN with a stronger re-identification model. It also updated appearance features with an exponential moving average instead of a gallery, added camera-motion compensation, combined appearance and motion in the cost, and replaced the matching cascade with a single global assignment. Two plug-in modules, appearance-free tracklet linking and Gaussian-smoothed interpolation, gave StrongSORT++.
ByteTrack (Zhang et al., ECCV 2022) went the other way for its main tracker: no appearance model, but a second association round that matches leftover tracks to low-confidence detections. BoT-SORT (Aharon, Orfaig & Bobrovsky, 2022) built on ByteTrack and brought appearance back as an option, fusing IoU and re-identification distances, alongside camera-motion compensation and a Kalman state with box width and height. The ByteTrack paper, ByteTrack, and SORT articles cover these trackers in more detail.
Related
- SORT
Simple Online and Realtime Tracking, a multi-object tracker that links per-frame detections into tracks using a constant-velocity Kalman filter on each box and Hungarian matching on box overlap.
- Simple Online and Realtime Tracking
The 2016 paper by Bewley et al. that introduced SORT, showing that a Kalman filter and Hungarian matching on strong CNN detections rival far more complex online multi-object trackers.
- Data Association
Deciding which measurements or detections belong to which tracked targets, and which are false alarms, missed detections, new targets, or targets that have disappeared.
- Multi-Object Tracking
Estimating the trajectories of a varying, unknown number of objects in a video while keeping each object's identity consistent over time.
- ByteTrack: Multi-Object Tracking by Associating Every Detection Box
The ECCV 2022 paper by Zhang et al. that introduced BYTE, a second association round for low-confidence detections, and the ByteTrack tracker built on it with a YOLOX detector.
References
- Wojke, N., Bewley, A. & Paulus, D. (2017). Simple Online and Realtime Tracking with a Deep Association Metric. IEEE International Conference on Image Processing (ICIP), 3645–3649.
- Zhang, Y., Sun, P., Jiang, Y., Yu, D., Weng, F., Yuan, Z., Luo, P., Liu, W. & Wang, X. (2022). ByteTrack: Multi-Object Tracking by Associating Every Detection Box. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 1–21.
- Du, Y., Zhao, Z., Song, Y., Zhao, Y., Su, F., Gong, T. & Meng, H. (2023). StrongSORT: Make DeepSORT Great Again. IEEE Transactions on Multimedia, 25, 8725–8737.
- Aharon, N., Orfaig, R. & Bobrovsky, B.-Z. (2022). BoT-SORT: Robust Associations Multi-Pedestrian Tracking. arXiv:2206.14651.