Evaluation & Metrics

IDF1

An identity-based score for multi-object tracking that matches whole ground-truth and predicted trajectories one-to-one and reports the F1 score of correctly identified detections.

beginner

IDF1 measures how well a multi-object tracker keeps identities: for what fraction of the time each object is labeled with the right identity. It first matches every ground-truth trajectory to at most one predicted trajectory over the whole video, then scores the detections that agree with this matching as an F1 score. Ristani et al. (2016) proposed it, with identification precision (IDP) and identification recall (IDR), for multi-target, multi-camera tracking, arguing that some users care more about “who is where at all times” than about how often a tracker makes each type of wrong decision. Single-camera benchmarks such as MOTChallenge have since adopted it as a secondary metric next to MOTA.

Formula

Matching. Each ground-truth trajectory τ\tau may be matched to one predicted trajectory γ\gamma, and vice versa. In a frame where both exist, their detections agree if they overlap enough, for example an intersection over union of at least a threshold Δ\Delta in the image, or a distance below 1 m on the ground plane. Matching τ\tau with γ\gamma costs the number of frames of τ\tau not explained by γ\gamma plus the number of frames of γ\gamma not explained by τ\tau; leaving a trajectory unmatched costs all its frames. A minimum-cost bipartite matching over the whole sequence, computable with the Hungarian algorithm, gives the one-to-one assignment.

Counts. Detections where matched trajectories agree are identity true positives (IDTP). The remaining predicted detections, from unmatched trajectories or from frames where a matched pair disagrees, are identity false positives (IDFP), and the remaining ground-truth detections are identity false negatives (IDFN). Then

IDP=IDTPIDTP+IDFP,IDR=IDTPIDTP+IDFN,\text{IDP} = \frac{\text{IDTP}}{\text{IDTP} + \text{IDFP}}, \qquad \text{IDR} = \frac{\text{IDTP}}{\text{IDTP} + \text{IDFN}}, IDF1=2 IDTP2 IDTP+IDFP+IDFN.\text{IDF1} = \frac{2\,\text{IDTP}}{2\,\text{IDTP} + \text{IDFP} + \text{IDFN}}.

Since IDTP + IDFP is the number of predicted detections and IDTP + IDFN the number of ground-truth detections, minimizing IDFP + IDFN is the same as maximizing IDTP.

Intuition

The evaluator, not the tracker, chooses which predicted identity stands for which true object, and Ristani et al. make the choice most favorable to the tracker. Every frame in which an object carries a different identity is then an error.

This differs from counting identity switches. If a tracker covers one person with identity 1 for two thirds of the trajectory and identity 2 for the rest, a switch-based count charges one error if the trajectory breaks once, but seven if it alternates between the two identities in eight pieces, although identity 1 covers the same two thirds. IDF1 charges the third of the frames not explained by identity 1 in both cases, and less when identity 1 covers more of the trajectory (Ristani et al., 2016).

Interpretation

  • Range: [0,1][0, 1], often given as a percentage. 1 means every ground-truth detection is explained by its matched prediction and every prediction is matched; higher is better.
  • IDP and IDR separate the two error directions: low IDP means predicted identities that do not fit any true object, low IDR means true objects left unexplained, through misses or identities split across several tracks.
  • Ristani et al. describe IDF1 as the ratio of correctly identified detections to the average number of ground-truth and predicted detections.

When to Use It

Use IDF1 when identity matters over long periods, as in following a person through a camera network, counting unique visitors, or re-identification after occlusion. Report it next to MOTA, which is dominated by detection errors, or use HOTA, which balances the two.

Failure Cases

Luiten et al. (2021), who proposed HOTA as an alternative, criticize IDF1 on several counts:

  • Unintuitive matches. Because the matching is global, a trajectory can be paired with a partner that fits it poorly, because each one’s better partners are matched elsewhere.
  • Correct detections can lower the score. Take one object tracked for NN frames. A tracker that labels the first half with identity 1 and the second half with identity 2 gets IDTP =N/2= N/2, IDFP =N/2= N/2, and IDFN =N/2= N/2, so IDF1 =0.5= 0.5. A tracker that outputs only the first half gets IDF1 =2/3= 2/3, so detecting less scores more.
  • Association outside the matching is ignored. Unmatched predictions count only as false positives, so whether they are fragmented or consistent does not change the score.
  • Track counting. Extra or missing trajectories are penalized as a whole, so a high IDF1 depends heavily on producing about as many tracks as there are objects.
  • No localization quality. As in MOTA, a detection either passes the threshold or not.

Variants

IDF1 was designed for camera networks. Detections from all cameras enter one matching, so overlapping and disjoint fields of view are handled the same way, and a hand-over error costs the frames it mislabels instead of a separate penalty that depends on the first and last frames in each view. Ristani et al. also suggest computing the scores once per camera and once across all cameras: the drop in IDF1 from the single-camera to the multi-camera evaluation measures cross-camera association. In single-camera tracking the same definitions apply unchanged.

Example

MOTA and IDF1 on a toy sequence with two people. Tracker A detects both people in every frame but restarts both identities halfway; tracker B keeps both identities but misses each person for three frames. The MOTA function counts an identity switch when an object’s track differs from its last one, as MOTChallenge does.

import numpy as np
from scipy.optimize import linear_sum_assignment


def clear_mot(gt, pred, thresh):
    """MOTA with CLEAR MOT matching. gt, pred: one dict {id: (x, y)} per frame."""
    fn = fp = idsw = n_gt = 0
    prev, last = {}, {}  # previous frame's matches; last track seen for each gt id
    for g_t, p_t in zip(gt, pred):
        dist = lambda g, p: np.linalg.norm(np.subtract(g_t[g], p_t[p]))
        # 1. Keep last frame's correspondences that are still valid.
        matches = {g: p for g, p in prev.items()
                   if g in g_t and p in p_t and dist(g, p) < thresh}
        # 2. Hungarian matching on the rest, rejecting pairs beyond the threshold.
        gs = [g for g in g_t if g not in matches]
        ps = [p for p in p_t if p not in matches.values()]
        if gs and ps:
            cost = np.array([[dist(g, p) for p in ps] for g in gs])
            for r, c in zip(*linear_sum_assignment(np.where(cost < thresh, cost, 1e6))):
                if cost[r, c] < thresh:
                    matches[gs[r]] = ps[c]
        # 3. Count errors.
        for g, p in matches.items():
            if g in last and last[g] != p:
                idsw += 1
            last[g] = p
        fn += len(g_t) - len(matches)
        fp += len(p_t) - len(matches)
        n_gt += len(g_t)
        prev = matches
    return 1 - (fn + fp + idsw) / n_gt, dict(FN=fn, FP=fp, IDSW=idsw)


def idf1(gt, pred, thresh):
    """IDF1 with one global one-to-one matching of whole trajectories."""
    g_ids = sorted({g for f in gt for g in f})
    p_ids = sorted({p for f in pred for p in f})
    overlap = np.zeros((len(g_ids), len(p_ids)))  # frames where the pair coincides
    for g_t, p_t in zip(gt, pred):
        for i, g in enumerate(g_ids):
            for j, p in enumerate(p_ids):
                if g in g_t and p in p_t and np.linalg.norm(
                        np.subtract(g_t[g], p_t[p])) < thresh:
                    overlap[i, j] += 1
    # Minimizing IDFN + IDFP is the same as maximizing IDTP.
    rows, cols = linear_sum_assignment(-overlap)
    idtp = int(overlap[rows, cols].sum())
    idfn = sum(len(f) for f in gt) - idtp
    idfp = sum(len(f) for f in pred) - idtp
    return 2 * idtp / (2 * idtp + idfp + idfn), dict(IDTP=idtp, IDFP=idfp, IDFN=idfn)


# Two people walk side by side for 10 frames.
T = 10
gt = [{"a": (t, 0.0), "b": (t, 5.0)} for t in range(T)]

# Tracker A detects everyone in every frame, but starts new identities
# for both people halfway through, as if both tracks were dropped and restarted.
pred_a = [{(1 if t < 5 else 3): (t, 0.2), (2 if t < 5 else 4): (t, 5.2)}
          for t in range(T)]

# Tracker B keeps both identities but misses each person in three frames.
missed = {1: {3, 4, 5}, 2: {6, 7, 8}}
pred_b = [{k: (t, y) for k, y in [(1, 0.2), (2, 5.2)] if t not in missed[k]}
          for t in range(T)]

for name, pred in [("A", pred_a), ("B", pred_b)]:
    mota, c = clear_mot(gt, pred, thresh=1.0)
    f1, d = idf1(gt, pred, thresh=1.0)
    print(f"tracker {name}: MOTA = {mota:.2f} {c}   IDF1 = {f1:.2f} {d}")

Output:

tracker A: MOTA = 0.90 {'FN': 0, 'FP': 0, 'IDSW': 2}   IDF1 = 0.50 {'IDTP': 10, 'IDFP': 10, 'IDFN': 10}
tracker B: MOTA = 0.70 {'FN': 6, 'FP': 0, 'IDSW': 0}   IDF1 = 0.82 {'IDTP': 14, 'IDFP': 0, 'IDFN': 6}

MOTA charges tracker A only two switches out of 20 ground-truth detections, so it ranks A first. IDF1 matches each person to one of their two half-length tracks, so half of A’s detections carry the wrong identity, and it ranks B first. Neither ranking is wrong: they answer different questions.

  • MOTA (Bernardin & Stiefelhagen, 2008) matches frame by frame and counts misses, false positives, and identity switches; it emphasizes detection.
  • HOTA (Luiten et al., 2021) scores association over all overlapping pairs of trajectories instead of one global matching, combines it with detection accuracy, and averages over localization thresholds.

Related

  • Multiple Object Tracking Accuracy

    The CLEAR MOT accuracy score for multi-object tracking, which counts misses, false positives, and identity switches over a sequence and normalizes them by the number of ground-truth objects.

  • Higher Order Tracking Accuracy

    A multi-object tracking metric that combines detection accuracy and association accuracy through their geometric mean, averaged over localization thresholds.

  • MOTChallenge

    A benchmark for multiple pedestrian tracking, released as MOT15, MOT16, MOT17, and MOT20, with hidden test annotations and a central evaluation server.

  • Hungarian Algorithm

    An algorithm that finds the minimum-cost one-to-one matching between two sets, used in computer vision to match detections to tracks and predictions to ground truth.

  • Data Association

    Deciding which measurements or detections belong to which tracked targets, and which are false alarms, missed detections, new targets, or targets that have disappeared.

  • Tracking by Detection

    Building object tracks by running a detector on every frame and linking its detections over time with a motion model and data association.

References

  1. Ristani, E., Solera, F., Zou, R., Cucchiara, R. & Tomasi, C. (2016). Performance Measures and a Data Set for Multi-Target, Multi-Camera Tracking. European Conference on Computer Vision (ECCV) Workshops, Lecture Notes in Computer Science, 17–35.
  2. Bernardin, K. & Stiefelhagen, R. (2008). Evaluating Multiple Object Tracking Performance: The CLEAR MOT Metrics. EURASIP Journal on Image and Video Processing, 2008, Article 246309.
  3. Luiten, J., Ošep, A., Dendorfer, P., Torr, P., Geiger, A., Leal-Taixé, L. & Leibe, B. (2021). HOTA: A Higher Order Metric for Evaluating Multi-Object Tracking. International Journal of Computer Vision, 129(2), 548–578.