Evaluation & Metrics

Higher Order Tracking Accuracy

A multi-object tracking metric that combines detection accuracy and association accuracy through their geometric mean, averaged over localization thresholds.

intermediate

Higher Order Tracking Accuracy (HOTA) scores a multi-object tracker by how well it detects objects, how well it keeps their identities over time, and how precisely it localizes them. It is the geometric mean of a detection accuracy and an association accuracy, averaged over a range of localization thresholds. Luiten et al. (2021) proposed it, arguing that the two established metrics are each biased: MOTA overemphasizes detection, and IDF1 association. HOTA has been reported by MOTChallenge since 2021 and is the primary tracking metric of the KITTI benchmark.

Formula

At a localization threshold α\alpha, predicted and ground-truth detections are matched one-to-one in each frame, and a pair can match only if its similarity SS, usually the box intersection over union, is at least α\alpha. Matched pairs are true positives (TP), unmatched ground-truth detections are false negatives (FN), and unmatched predictions are false positives (FP). The detection accuracy is the Jaccard index of these counts:

DetAα=∣TP∣∣TP∣+∣FN∣+∣FP∣.\text{DetA}_\alpha = \frac{|\mathrm{TP}|}{|\mathrm{TP}| + |\mathrm{FN}| + |\mathrm{FP}|}.

Association is measured for each true positive cc. Its true positive associations, TPA(c)\mathrm{TPA}(c), are the true positives with the same ground-truth identity and the same predicted identity as cc. Its false negative associations, FNA(c)\mathrm{FNA}(c), are detections of the same ground-truth object that received a different predicted identity or were missed. Its false positive associations, FPA(c)\mathrm{FPA}(c), are predictions with the same predicted identity that belong to another object or to none. The association score of cc, and the association accuracy, are

A(c)=∣TPA(c)∣∣TPA(c)∣+∣FNA(c)∣+∣FPA(c)∣,AssAα=1∣TP∣∑c∈TPA(c).A(c) = \frac{|\mathrm{TPA}(c)|}{|\mathrm{TPA}(c)| + |\mathrm{FNA}(c)| + |\mathrm{FPA}(c)|}, \qquad \text{AssA}_\alpha = \frac{1}{|\mathrm{TP}|} \sum_{c \in \mathrm{TP}} A(c).

HOTA at threshold α\alpha combines them, and the final score averages over 19 thresholds:

HOTAα=DetAα⋅AssAα=∑c∈TPA(c)∣TP∣+∣FN∣+∣FP∣,HOTA=119∑α∈{0.05, 0.10, …, 0.95}HOTAα.\text{HOTA}_\alpha = \sqrt{\text{DetA}_\alpha \cdot \text{AssA}_\alpha} = \sqrt{\frac{\sum_{c \in \mathrm{TP}} A(c)}{|\mathrm{TP}| + |\mathrm{FN}| + |\mathrm{FP}|}}, \qquad \text{HOTA} = \frac{1}{19} \sum_{\alpha \in \{0.05,\, 0.10,\, \ldots,\, 0.95\}} \text{HOTA}_\alpha .

The average approximates the integral of HOTAα\text{HOTA}_\alpha over α∈[0,1]\alpha \in [0, 1], and the matching is recomputed at each threshold. The authors call this a “double Jaccard” formulation: a Jaccard index over detections in which each true positive is weighted by a Jaccard index over associations.

Matching. The Hungarian algorithm selects the matches that maximize, in order of priority, the number of true positives, their mean association score, and their mean similarity. Each candidate pair (i,j)(i, j) receives the score

MS(i,j)={1ϵ+Amax⁡(i,j)+ϵ S(i,j)if S(i,j)≥α,0otherwise,\mathrm{MS}(i, j) = \begin{cases} \dfrac{1}{\epsilon} + A_{\max}(i, j) + \epsilon\, S(i, j) & \text{if } S(i, j) \ge \alpha, \\[4pt] 0 & \text{otherwise,} \end{cases}

where ϵ\epsilon is small and Amax⁡(i,j)A_{\max}(i, j) is the association score the pair would have if detections could match many-to-many. A(c)A(c) depends on the matches themselves, so it cannot be optimized directly; Amax⁡A_{\max} is a proxy, and the matching is not guaranteed to maximize HOTA itself.

Sub-metrics. Detection splits into recall and precision, DetReα=∣TP∣/(∣TP∣+∣FN∣)\text{DetRe}_\alpha = |\mathrm{TP}| / (|\mathrm{TP}| + |\mathrm{FN}|) and DetPrα=∣TP∣/(∣TP∣+∣FP∣)\text{DetPr}_\alpha = |\mathrm{TP}| / (|\mathrm{TP}| + |\mathrm{FP}|). Association does the same per true positive:

AssReα=1∣TP∣∑c∈TP∣TPA(c)∣∣TPA(c)∣+∣FNA(c)∣,AssPrα=1∣TP∣∑c∈TP∣TPA(c)∣∣TPA(c)∣+∣FPA(c)∣.\text{AssRe}_\alpha = \frac{1}{|\mathrm{TP}|} \sum_{c \in \mathrm{TP}} \frac{|\mathrm{TPA}(c)|}{|\mathrm{TPA}(c)| + |\mathrm{FNA}(c)|}, \qquad \text{AssPr}_\alpha = \frac{1}{|\mathrm{TP}|} \sum_{c \in \mathrm{TP}} \frac{|\mathrm{TPA}(c)|}{|\mathrm{TPA}(c)| + |\mathrm{FPA}(c)|}.

Localization accuracy is the mean similarity of matched pairs, averaged over thresholds:

LocA=119∑α1∣TPα∣∑c∈TPαS(c).\text{LocA} = \frac{1}{19} \sum_{\alpha} \frac{1}{|\mathrm{TP}_\alpha|} \sum_{c \in \mathrm{TP}_\alpha} S(c).

Each sub-metric is averaged over the 19 thresholds in the same way as HOTA.

Intuition

A(c)A(c) asks, for one correct detection, how much of its ground-truth trajectory and its predicted trajectory coincide over the whole video. If a tracker follows one person with one identity for the first half of the video and another for the second half, every detection of that person has an association score of one half, whereas MOTA would count the switch as a single error. HOTA matches detections, like MOTA, but scores association over whole trajectories, like IDF1.

Interpretation

HOTA, DetA, AssA, and LocA all range from 0 to 1 and are usually reported as percentages. At each threshold, HOTAα\text{HOTA}_\alpha lies between DetAα\text{DetA}_\alpha and AssAα\text{AssA}_\alpha and approaches zero if either does, so a tracker cannot compensate for failing at one by excelling at the other. With a single object and a single track, AssA equals DetA and HOTA reduces to the Jaccard index.

Luiten et al. (2021) re-scored 37 published MOT17 trackers in April 2020. Tracktor++v2 ranked first by MOTA (56.3) and third by HOTA (45.1), with DetA 45.3 and AssA 45.0. SAS_MOT17 ranked 31st by MOTA (44.2) but fifth by HOTA (43.0), because its association accuracy (49.6) was the highest of the 37 while its detection accuracy (37.5) was weak. Note that DetA and AssA must be combined at each α\alpha before averaging, so the final HOTA is generally not exactly the geometric mean of the final DetA and AssA.

When to Use It

Use HOTA as the headline metric for benchmark comparisons, alongside MOTA and IDF1 for continuity with earlier work. Use the sub-metrics to diagnose a tracker or tune it for an application. A driving system may weight detection recall and precision most, while crowd or sports analysis cares about association. HOTA works for any representation with a similarity in [0,1][0, 1], including boxes, masks, 3D boxes, and points, and for multi-camera tracking. The authors also reported a user study of 230 participants in which, when two metrics disagreed, people agreed with HOTA over MOTA 61.6% of the time and over IDF1 72.0% of the time, excluding ties.

Failure Cases

  • Online use. Association is scored over the whole video, so a detection’s score depends on future frames. The authors note that this may make HOTA less than ideal for evaluating online trackers.
  • Fragmentation is ignored. By design, a track broken into many short pieces and one broken once can score the same if their global alignment is equal. The authors note this is a drawback when short-range continuity matters.
  • Fixed balance. Equal weighting of detection and association is a design choice. Valmadre et al. (2021) compared rankings of 125 MOT17 submissions and found HOTA somewhat less correlated with identity-switch rates than IDF1 and ATA, two metrics that match whole trajectories. They argued that every combined metric fixes an implicit trade-off between detection and association, while the right balance depends on the application.
  • Long tracks weigh more. AssA averages over detections, so objects visible for many frames dominate the association score.
  • Many classes and partial labels. Li et al. (2022) showed that per-class evaluation breaks down when trackers misclassify objects in long-tailed data: in their example, a tracker that follows objects correctly but assigns the wrong classes scores zero under MOTA, IDF1, and HOTA alike. Counting every unmatched prediction as a false positive also wrongly penalizes datasets in which not every object is annotated.

Variants

Luiten et al. (2021) defined several extensions:

  • Online HOTA scores each detection’s association using only frames up to the current one.
  • Fragmentation-aware HOTA combines each detection’s association score with a short-range fragment score through a geometric mean.
  • Weighted HOTA gives false negatives, false positives, and the two association errors adjustable weights.
  • Classification-aware HOTA, with class-averaged and federated forms, weights each match by the predicted probability of its true class.
  • Confidence-ranked HOTA evaluates confidence-scored outputs across recall levels.

TETA (Li et al., 2022) builds on HOTA’s association score but separates localization, association, and classification, and averages them arithmetically.

  • MOTA (Bernardin & Stiefelhagen, 2008) counts misses, false positives, and identity switches at one fixed threshold. It is unbounded below and biased toward detection.
  • IDF1 (Ristani et al., 2016) is the F1 score of a one-to-one matching between whole trajectories. It is biased toward association.
  • MOTP measures localization at one threshold. LocA is its multi-threshold counterpart.
  • ALTA (Valmadre et al., 2021) restricts association to a temporal window, so the window length sets the balance between detection and association.

The reference implementation is TrackEval, which also computes the CLEAR MOT and identity metrics and serves as the official evaluation code of MOTChallenge and KITTI tracking.

Related

  • Multi-Object Tracking

    Estimating the trajectories of a varying, unknown number of objects in a video while keeping each object's identity consistent over time.

  • MOTChallenge

    A benchmark for multiple pedestrian tracking, released as MOT15, MOT16, MOT17, and MOT20, with hidden test annotations and a central evaluation server.

  • KITTI

    A benchmark suite of real driving data from Karlsruhe, with stereo, optical flow, scene flow, odometry, object detection, and tracking benchmarks built from camera, LiDAR, and GPS/IMU recordings.

  • Hungarian Algorithm

    An algorithm that finds the minimum-cost one-to-one matching between two sets, used in computer vision to match detections to tracks and predictions to ground truth.

  • Tracking by Detection

    Building object tracks by running a detector on every frame and linking its detections over time with a motion model and data association.

  • ByteTrack

    A multi-object tracker that associates high-score detections first and then matches the remaining tracks to low-score detections, recovering occluded objects that score thresholds would discard.

References

  1. Luiten, J., Ošep, A., Dendorfer, P., Torr, P., Geiger, A., Leal-Taixé, L. & Leibe, B. (2021). HOTA: A Higher Order Metric for Evaluating Multi-Object Tracking. International Journal of Computer Vision, 129(2), 548–578.
  2. Bernardin, K. & Stiefelhagen, R. (2008). Evaluating Multiple Object Tracking Performance: The CLEAR MOT Metrics. EURASIP Journal on Image and Video Processing, 2008, Article 246309.
  3. Ristani, E., Solera, F., Zou, R., Cucchiara, R. & Tomasi, C. (2016). Performance Measures and a Data Set for Multi-Target, Multi-Camera Tracking. European Conference on Computer Vision (ECCV) Workshops, Lecture Notes in Computer Science, 17–35.
  4. Valmadre, J., Bewley, A., Huang, J., Sun, C., Sminchisescu, C. & Schmid, C. (2021). Local Metrics for Multi-Object Tracking. arXiv:2104.02631.
  5. Li, S., Danelljan, M., Ding, H., Huang, T. E. & Yu, F. (2022). Tracking Every Thing in the Wild. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 498–515.