Evaluation & Metrics

Endpoint Error

The standard accuracy measure for optical flow, the distance in pixels between an estimated flow vector and the true one, averaged over the image.

beginner

Endpoint error (EPE), also written EE, measures how far an estimated optical flow vector lands from the true one. At each pixel it is the Euclidean distance, in pixels, between the tip of the estimated displacement vector and the tip of the ground-truth vector. Averaged over an image or a dataset, it gives the average endpoint error (AEPE, often written simply EPE), the most widely reported accuracy number for optical flow and other dense motion estimation methods.

Formula

Let w(x)=(u,v)\mathbf{w}(\mathbf{x}) = (u, v) be the estimated flow at pixel x\mathbf{x} and wgt(x)=(ugt,vgt)\mathbf{w}^{\mathrm{gt}}(\mathbf{x}) = (u^{\mathrm{gt}}, v^{\mathrm{gt}}) the ground truth. The per-pixel endpoint error is

EPE(x)=∥w(x)−wgt(x)∥2=(u−ugt)2+(v−vgt)2\mathrm{EPE}(\mathbf{x}) = \left\lVert \mathbf{w}(\mathbf{x}) - \mathbf{w}^{\mathrm{gt}}(\mathbf{x}) \right\rVert_2 = \sqrt{\left(u - u^{\mathrm{gt}}\right)^2 + \left(v - v^{\mathrm{gt}}\right)^2}

The average endpoint error over a set Ω\Omega of NN evaluated pixels is

AEPE=1N∑x∈ΩEPE(x)\mathrm{AEPE} = \frac{1}{N} \sum_{\mathbf{x} \in \Omega} \mathrm{EPE}(\mathbf{x})

Ω\Omega contains only pixels with valid ground truth, and benchmarks often restrict it further, for example to non-occluded pixels.

Intuition

Draw the estimated and true displacement vectors from the same pixel. EPE is the length of the line joining their tips. It does not distinguish errors in direction from errors in length: a vector 1 pixel too long and a vector 1 pixel off to the side both score 1.

Interpretation

  • Range: [0,∞)[0, \infty), in pixels. Zero is a perfect match; lower is better.
  • Scale depends on the data. EPE grows with image resolution and motion size, so a value is only meaningful within one benchmark. When Sintel was introduced, methods with an average EPE below 0.5 pixel on Middlebury scored roughly 10 pixels on Sintel, which has much larger motions (Butler et al., 2012).
  • Ground truth has finite accuracy. Baker et al. (2011) note that part of the Middlebury ground truth is only accurate to about 0.25 pixel, so smaller differences there are not meaningful.
  • Averages hide distributions. An AEPE of 1 pixel can mean every pixel is off by about 1 pixel, or that most are nearly perfect and a few are wildly wrong. Outlier percentages separate these cases.

When to Use It

EPE is the default whenever dense ground-truth correspondences exist: optical flow, the 2D flow component of scene flow, and point tracking with known trajectories. It is also the usual training loss for supervised flow networks.

Failure Cases

  • Small, fast objects. The average is dominated by the large background, so a method can miss a fast-moving object entirely and still report a low AEPE (see the example below).
  • Large motions dominate. A 5% error on a 100-pixel displacement contributes 5 pixels, as much as a 5-pixel error on a static pixel, so a few large motions can decide a ranking.
  • Different pixel sets. Some benchmarks score occluded pixels, others have no ground truth there. Numbers computed over different pixel sets, or averaged per frame versus over all pixels at once, are not comparable.
  • Accuracy is not usefulness. EPE ignores whether errors fall on motion boundaries, which matter for segmentation and frame interpolation, and whether flow is temporally consistent.

Variants

Outlier percentages

  • Middlebury reports RX, the percentage of pixels with endpoint error above X pixels (X = 0.5, 1.0, 2.0), and AX, the error at the Xth percentile (A50, A75, A95) (Baker et al., 2011).
  • KITTI 2015 (Fl) counts a pixel as an outlier when its endpoint error exceeds both 3 pixels and 5% of the true flow magnitude (Menze & Geiger, 2015). The relative term tolerates small errors on very large motions, where the laser-based ground truth is less accurate. The benchmark reports outliers over background (Fl-bg), foreground objects (Fl-fg), and all pixels with ground truth (Fl-all), and ranks methods by the percentage of erroneous pixels.
Fl=100%N∑x∈Ω1 ⁣[ EPE(x)>3 px ∧ EPE(x)>0.05∥wgt(x)∥ ]\mathrm{Fl} = \frac{100\%}{N} \sum_{\mathbf{x} \in \Omega} \mathbb{1}\!\left[\, \mathrm{EPE}(\mathbf{x}) > 3 \ \text{px} \ \wedge \ \mathrm{EPE}(\mathbf{x}) > 0.05 \left\lVert \mathbf{w}^{\mathrm{gt}}(\mathbf{x}) \right\rVert \,\right]

Region-restricted averages

  • MPI Sintel reports EPE over all, matched, and unmatched pixels, the last being visible in only one frame of a pair (nearly 8.5% of pixels). In the original evaluation, average errors in unmatched regions exceeded 40 pixels, roughly eight times those in matched regions. Sintel also bins EPE by distance to the nearest occlusion boundary (under 10, 10–60, over 60 pixels) and by speed (under 10, 10–40, over 40 pixels per frame) (Butler et al., 2012).
  • Middlebury computes statistics over all pixels with reliable ground truth, around motion discontinuities, and in textureless regions (Baker et al., 2011).

Angular error

The older angular error (AE) treats each flow vector as a space-time direction (u,v,1)(u, v, 1) and measures the angle between estimate and ground truth:

AE=cos⁡−1(1+u ugt+v vgt1+u2+v2 1+(ugt)2+(vgt)2)\mathrm{AE} = \cos^{-1} \left( \frac{1 + u\,u^{\mathrm{gt}} + v\,v^{\mathrm{gt}}}{\sqrt{1 + u^2 + v^2}\,\sqrt{1 + (u^{\mathrm{gt}})^2 + (v^{\mathrm{gt}})^2}} \right)

The measure dates to Fleet and Jepson (1990) and became standard through the survey of Barron, Fleet and Beauchemin (1994). It is reported in degrees and avoids dividing by zero for static pixels, but it penalizes errors in large motions less than errors in small ones, and the constant 1 is an arbitrary choice of scale. Baker et al. (2011) showed a method that failed on a fast-moving building ranked 6th by average AE but 20th by average EPE, and argued that EPE, which they trace to Otte and Nagel (1994), should be the preferred measure. Later benchmarks followed.

Example

Average EPE and a KITTI-style outlier rate in NumPy, on a synthetic case where the method misses a small, fast object:

import numpy as np


def endpoint_error(flow_pred, flow_gt):
    """Per-pixel endpoint error for flow fields of shape (H, W, 2)."""
    return np.linalg.norm(flow_pred - flow_gt, axis=-1)


def flow_metrics(flow_pred, flow_gt, valid=None, abs_thresh=3.0, rel_thresh=0.05):
    epe = endpoint_error(flow_pred, flow_gt)
    if valid is None:
        valid = np.ones(epe.shape, dtype=bool)
    gt_magnitude = np.linalg.norm(flow_gt, axis=-1)

    # KITTI 2015-style outlier: error above 3 px AND above 5% of the true motion.
    outlier = (epe > abs_thresh) & (epe > rel_thresh * gt_magnitude)

    return epe[valid].mean(), 100.0 * outlier[valid].mean()


rng = np.random.default_rng(0)
gt = np.zeros((100, 100, 2))
gt[..., 0] = 2.0                      # background moves 2 px to the right
gt[40:60, 40:60] = (20.0, -5.0)       # a small object moves fast

pred = gt + rng.normal(scale=0.3, size=gt.shape)   # small noise everywhere
pred[40:60, 40:60] = (2.0, 0.0)       # the object is missed: predicted as background

valid = np.ones(gt.shape[:2], dtype=bool)
valid[:, :5] = False                  # e.g. pixels without ground truth

aepe, outliers = flow_metrics(pred, gt, valid)
print(f"average EPE: {aepe:.2f} px")         # 1.15 px
print(f"outliers:    {outliers:.1f} %")      # 4.2 %

object_mask = np.zeros_like(valid)
object_mask[40:60, 40:60] = True
print(f"EPE on object: {endpoint_error(pred, gt)[object_mask].mean():.2f} px")  # 18.68 px

The average EPE of 1.15 pixels looks acceptable, yet every object pixel is wrong by almost 19 pixels. The outlier rate and the per-region EPE reveal the failure.

  • Outlier rate (Fl, RX): robust to a few extreme errors, but ignores how wrong the outliers are.
  • Disparity error in stereo: the one-dimensional analogue, with the same 3-pixel-and-5% rule on KITTI 2015.
  • Interpolation error: Middlebury also scores flow by how well it predicts an in-between frame, which needs no ground-truth flow.

Related

  • Motion Estimation

    Recovering how image content, objects, or the camera moved from a sequence of images, from per-pixel flow to global and 3D motion.

  • Optical Flow Estimation

    Computing a dense field of pixel displacements between two video frames, from classical variational methods to learned networks such as RAFT.

  • RAFT

    A deep network for optical flow that matches all pairs of pixels once and refines a single flow field with a recurrent update operator.

References

  1. Baker, S., Scharstein, D., Lewis, J. P., Roth, S., Black, M. J. & Szeliski, R. (2011). A Database and Evaluation Methodology for Optical Flow. International Journal of Computer Vision, 92(1), 1–31.
  2. Butler, D. J., Wulff, J., Stanley, G. B. & Black, M. J. (2012). A Naturalistic Open Source Movie for Optical Flow Evaluation. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 611–625.
  3. Menze, M. & Geiger, A. (2015). Object Scene Flow for Autonomous Vehicles. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3061–3070.
  4. Barron, J. L., Fleet, D. J. & Beauchemin, S. S. (1994). Performance of Optical Flow Techniques. International Journal of Computer Vision, 12(1), 43–77.
  5. Fleet, D. J. & Jepson, A. D. (1990). Computation of Component Image Velocity from Local Phase Information. International Journal of Computer Vision, 5(1), 77–104.
  6. Otte, M. & Nagel, H.-H. (1994). Optical Flow Estimation: Advances and Comparisons. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science, 49–60.