Motion & Tracking
ByteTrack: Multi-Object Tracking by Associating Every Detection Box
The ECCV 2022 paper by Zhang et al. that introduced BYTE, a second association round for low-confidence detections, and the ByteTrack tracker built on it with a YOLOX detector.
intermediate
“ByteTrack: Multi-Object Tracking by Associating Every Detection Box” by Yifu Zhang, Peize Sun, Yi Jiang, Dongdong Yu, Fucheng Weng, Zehuan Yuan, Ping Luo, Wenyu Liu, and Xinggang Wang, of Huazhong University of Science and Technology, the University of Hong Kong, and ByteDance, appeared on arXiv in October 2021 and was published at the European Conference on Computer Vision (ECCV) in 2022. It made two contributions: BYTE, an association rule that gives tracks left unmatched a second chance with low-confidence detections instead of discarding those boxes, and ByteTrack, a tracker that combines BYTE with a SORT-style Kalman filter and the YOLOX detector. The name reflects the authors’ view of each detection box as a basic unit of a tracklet, like a byte in a program. The ByteTrack article explains the algorithm and its official implementation in detail.
Problem
By 2021, tracking by detection was, in the authors’ words, the most effective paradigm for multi-object tracking. Almost all such trackers kept only detections above a confidence threshold, typically 0.5, because low-score boxes contain many background false positives. The paper argued that this throws away real objects at exactly the wrong moment. When a person becomes partly occluded or blurred, the detector’s score drops and the box is filtered out, interrupting the track; lowering the threshold instead admits background boxes as spurious tracks. The authors named this a dilemma that very few methods addressed.
Their observation was that similarity to existing tracks separates the two cases. A low-score box near where a track is predicted to be is probably that object; a low-score box that matches no track is probably background. Most prior work, the authors argued, had focused on better similarity measures and matching strategies, such as DeepSORT’s appearance cascade, while the way detection boxes are used sets an upper bound on what association can achieve.
Contribution
The paper’s contributions, as its introduction frames them:
- BYTE. A simple association method that uses almost every detection box: high-score boxes are matched first, and the tracks left over are then matched to low-score boxes, with unmatched low-score boxes discarded as background.
- Generality. BYTE can be inserted into existing trackers regardless of their detector or first-round similarity. The authors applied it to nine trackers of different designs.
- ByteTrack. A tracker built from the YOLOX detector and BYTE, with motion as the only association cue on the pedestrian benchmarks. The authors reported it as first on the MOT17 and MOT20 leaderboards of MOTChallenge at the time, and as the fastest tracker in their comparison tables. The code, pre-trained models, deployment versions, and tutorials for applying BYTE to other trackers were released (github.com/ifzhang/ByteTrack).
Method
The paper’s pseudocode splits each frame’s detections at a single score threshold , 0.6 by default, into high-score and low-score sets, and predicts every track’s box with a Kalman filter, as in SORT. The first association matches high-score boxes to all tracks, including lost ones, using IoU or re-identification feature distance, solved with the Hungarian algorithm; the implementation details state that assignments with IoU below 0.2 are rejected. The second association matches the tracks still unmatched to the low-score boxes, by IoU alone, since the authors found appearance features of heavily occluded or blurred boxes unreliable. Unmatched low-score boxes are deleted, unmatched tracks are kept as lost for 30 frames, and unmatched high-score boxes start new tracks. Only active tracks are output.
ByteTrack’s detector was YOLOX-X, initialized from COCO weights and trained for the MOT17 test submission on MOT17, CrowdHuman, CityPersons, and ETHZ, at an input size of 1440 × 800. In an appendix, the authors describe modifying YOLOX so that boxes are not clipped at the image border, because MOT17 annotates whole-body boxes even for people partly outside the frame. Test-set results also used linear interpolation across track gaps of up to 20 frames as post-processing. The ByteTrack article describes where the official code departs from the pseudocode, including a separate low threshold of 0.1 and score-weighted IoU in the first round.
Results
Ablations used a split of the MOT17 training set: the first half of each video for training, the second half for validation.
- Against other association methods. With the same YOLOX detections, BYTE raised SORT’s MOTA from 74.6 to 76.6 and IDF1 from 76.9 to 79.3, and cut identity switches from 291 to 159. It also scored above DeepSORT (75.4 MOTA) and MOTDT (75.8), both of which use appearance features.
- Threshold robustness. Varying the high threshold from 0.2 to 0.8, BYTE’s MOTA and IDF1 changed less than SORT’s, because objects below the threshold are still recovered in the second round.
- What the low-score boxes contain. The low-score boxes BYTE kept held notably more true positives than false positives, even in sequences where false positives dominate the low-score boxes overall.
- Nine trackers. BYTE was added to JDE, CSTrack, FairMOT, TraDes, QDTrack, CenterTrack, Chained-Tracker, TransTrack, and MOTR, which variously use re-identification, learned motion, chained pairs of frames, and attention. It was either inserted after each tracker’s own association or used to replace it on the tracker’s detections. Identity switches fell in every configuration, and MOTA and IDF1 rose in almost all, with four drops of under one point; for CenterTrack, IDF1 rose from 64.2 to 74.0 and identity switches fell from 528 to 144. The abstract describes consistent IDF1 gains of 1 to 10 points, but IDF1 fell slightly in two configurations of the table.
On the test sets, under the private-detection protocol, the authors reported:
| Benchmark | MOTA | IDF1 | HOTA | Speed (V100) |
|---|---|---|---|---|
| MOT17 | 80.3 | 77.3 | 63.1 | 29.6 FPS |
| MOT20 | 77.8 | 75.2 | 61.3 | 17.5 FPS |
The speeds include detection, measured at half precision with batch size 1; on MOT17, association took about 4 ms per frame. The authors reported that both entries ranked first on the respective leaderboards at the time, ahead of trackers that used appearance or attention. On BDD100K, a multi-class driving benchmark with low frame rates and large camera motion, they dropped the Kalman filter and used re-identification features from a UniTrack model in the first round, reporting 40.1 mMOTA and 55.8 mIDF1 on the test set, against 35.5 and 52.3 for QDTrack. They also reported ranking first on the crowded HiEve benchmark, and, under the public-detection protocol, 67.4 MOTA on the MOT17 test set.
Impact
The paper shifted attention from the similarity measure to how detection boxes are used. Its two-round, score-split association was easy to add to other trackers; BoT-SORT, for example, kept it and added appearance features to the first round. ByteTrack itself, with its released detector weights and code, served as the baseline that BoT-SORT and others extended, and the paper has been widely cited.
The results also reinforced a lesson from SORT: with a strong detector, a Kalman filter and IoU matching could match or beat trackers built on re-identification or transformers. Later trackers, including StrongSORT and OC-SORT, adopted ByteTrack’s YOLOX-X detector setup, which made it easier to compare association methods on equal detections.
Limitations
Much of ByteTrack’s benchmark accuracy comes from its detector, training data, and post-processing rather than from BYTE. The test-set comparison is between complete systems, and on validation interpolation alone raised MOTA from 76.6 to 78.3. The authors’ own ablations show how strongly results depend on the detector: with the much smaller YOLOX-Nano at a lower resolution, MOTA on their validation split fell to 64.4. MOTA, the headline metric, is itself dominated by detection errors, as the paper notes.
The motion model is the other weak point. On BDD100K the authors dropped the Kalman filter for appearance matching, citing low frame rates and large camera motion, conditions where constant-velocity prediction and IoU struggle. With the same detections, the OC-SORT authors later reported ByteTrack well behind their method on DanceTrack, a dance dataset with highly nonlinear motion. The pseudocode also omits track rebirth, which the text covers only briefly, and the official code adds rules, such as track confirmation, that the paper does not state.
What Came After
BoT-SORT (Aharon, Orfaig & Bobrovsky, 2022) built directly on ByteTrack, adding camera-motion compensation, a Kalman state that estimates box width and height, and an optional fusion of IoU and re-identification distances. OC-SORT (Cao et al., CVPR 2023) used ByteTrack’s released YOLOX weights to compare association methods on equal detections, and corrected the error that a Kalman filter accumulates while a track coasts through an occlusion, targeting nonlinear motion. StrongSORT (Du et al., 2023), which revisited DeepSORT, adopted YOLOX-X following ByteTrack. The DeepSORT paper article covers that lineage, and the ByteTrack and SORT articles describe these variants.
Related
- ByteTrack
A multi-object tracker that associates high-score detections first and then matches the remaining tracks to low-score detections, recovering occluded objects that score thresholds would discard.
- Simple Online and Realtime Tracking with a Deep Association Metric
The 2017 paper by Wojke, Bewley, and Paulus that introduced DeepSORT, extending SORT with a CNN appearance descriptor trained for person re-identification and a matching cascade, and reporting about 45% fewer identity switches on MOT16.
- Simple Online and Realtime Tracking
The 2016 paper by Bewley et al. that introduced SORT, showing that a Kalman filter and Hungarian matching on strong CNN detections rival far more complex online multi-object trackers.
- MOTChallenge
A benchmark for multiple pedestrian tracking, released as MOT15, MOT16, MOT17, and MOT20, with hidden test annotations and a central evaluation server.
- Tracking by Detection
Building object tracks by running a detector on every frame and linking its detections over time with a motion model and data association.
References
- Zhang, Y., Sun, P., Jiang, Y., Yu, D., Weng, F., Yuan, Z., Luo, P., Liu, W. & Wang, X. (2022). ByteTrack: Multi-Object Tracking by Associating Every Detection Box. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 1–21.
- Aharon, N., Orfaig, R. & Bobrovsky, B.-Z. (2022). BoT-SORT: Robust Associations Multi-Pedestrian Tracking. arXiv:2206.14651.
- Cao, J., Pang, J., Weng, X., Khirodkar, R. & Kitani, K. (2023). Observation-Centric SORT: Rethinking SORT for Robust Multi-Object Tracking. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 9686–9696.
- Du, Y., Zhao, Z., Song, Y., Zhao, Y., Su, F., Gong, T. & Meng, H. (2023). StrongSORT: Make DeepSORT Great Again. IEEE Transactions on Multimedia, 25, 8725–8737.