Motion & Tracking

Brightness Constancy Assumption

The assumption that a scene point keeps the same image intensity as it moves between frames, which turns motion estimation into an intensity-matching problem.

intermediate

The brightness constancy assumption states that a point in the scene has the same intensity in every frame in which it is visible: as it moves across the image, its brightness travels with it. It is what makes correspondence measurable from pixel values alone, and it is the starting point of most classical optical flow methods, of direct image alignment, and of stereo matching costs. It is also called the intensity constancy or photometric consistency assumption.

Definition

Brightness constancy is the assumption that the intensity recorded for a scene point does not change between frames, so that any change in the intensity at a fixed pixel is caused only by motion. It is a modeling assumption about the scene, the lighting, and the camera, not a law of image formation.

Intuition

Track a single speck of paint on a moving object. If the light falling on it and the camera’s exposure stay the same, the speck has the same gray value in every frame, and finding where it went is a search for a patch with the same values. This is reasonable for matte surfaces under steady lighting between consecutive video frames. It is wrong for a mirror, a flickering light, or a shadow sweeping across a wall, and every method built on it inherits those failure cases.

Formal Definition

For frames I0I_0 and I1I_1 and a displacement field w(x)=(u,v)\mathbf{w}(\mathbf{x}) = (u, v), brightness constancy states

I1(x+w)=I0(x).I_1(\mathbf{x} + \mathbf{w}) = I_0(\mathbf{x}).

This equation is nonlinear in w\mathbf{w}, because w\mathbf{w} appears inside the image function. Expanding I1I_1 around x\mathbf{x} with a Taylor series,

I1(x+w)=I1(x)+∇I1(x)⊤w+12 w⊤H(x) w+…I_1(\mathbf{x} + \mathbf{w}) = I_1(\mathbf{x}) + \nabla I_1(\mathbf{x})^\top \mathbf{w} + \tfrac{1}{2}\, \mathbf{w}^\top \mathbf{H}(\mathbf{x})\, \mathbf{w} + \dots

where H\mathbf{H} is the Hessian of I1I_1, and keeping only the first-order term with It=I1(x)−I0(x)I_t = I_1(\mathbf{x}) - I_0(\mathbf{x}) gives the optical flow constraint equation:

Ixu+Iyv+It=0.I_x u + I_y v + I_t = 0.

The small-motion condition

Dropping the second-order term is valid only when it is small compared with the first-order term:

∣12 w⊤H w∣≪∣∇I⊤w∣.\big|\tfrac{1}{2}\, \mathbf{w}^\top \mathbf{H}\, \mathbf{w}\big| \ll \big|\nabla I^\top \mathbf{w}\big|.

The displacement must be small compared with the distance over which the image gradient changes. For a sinusoidal pattern of wavelength λ\lambda, the linear model degrades long before ∥w∥\lVert \mathbf{w} \rVert reaches λ/2\lambda / 2, and at λ/2\lambda / 2 the motion is ambiguous outright: shifting a sinusoid by half a period forward or backward produces the same image. Fine texture therefore allows only very small displacements; smooth, low-frequency content tolerates larger ones.

This is why practical methods blur the images before differentiating, work on an image pyramid from coarse to fine, and iterate: after each linear step they warp I1I_1 by the current estimate and linearize again, so each step only explains a small residual motion. Brox et al. minimized the original, non-linearized energy with nested fixed-point iterations and showed that this numerical scheme amounts to coarse-to-fine warping, giving a theoretical justification to a technique previously motivated by experiment.

The continuous form

In continuous time, brightness constancy along a trajectory x(t)\mathbf{x}(t) means the total derivative of intensity vanishes:

ddtI(x(t),t)=Ixdxdt+Iydydt+It=0.\frac{\mathrm{d}}{\mathrm{d}t} I\big(\mathbf{x}(t), t\big) = I_x \frac{\mathrm{d}x}{\mathrm{d}t} + I_y \frac{\mathrm{d}y}{\mathrm{d}t} + I_t = 0.

Here the constraint is exact, with velocities in place of displacements; this is the form Horn and Schunck derived. The small-motion condition appears only because video is sampled at discrete times.

Properties

One equation per pixel

The constraint gives one equation for two unknowns at every pixel and fixes only the component of motion along the image gradient. This is the aperture problem, and every flow method adds a second assumption to resolve it.

When the assumption fails

Brightness constancy is violated whenever the intensity of a scene point changes for reasons other than motion:

  • Illumination changes. Lights switching on, lamp flicker, passing clouds, and automatic exposure or gain control.
  • Shading changes. A matte surface that rotates relative to the light changes its shading even under fixed lighting.
  • Specular highlights and shadows. Highlights move with the viewpoint and the light, not with the surface; a moving shadow changes the brightness of a static surface.
  • Occlusion and disocclusion. A point that becomes hidden, or newly visible, has no match in the other frame, and no photometric model can supply one.
  • Noise, blur, and compression. Sensor noise enters ItI_t directly as the difference of two noisy images; motion blur and compression artifacts change from frame to frame.
  • Transparency and atmospheric effects. Smoke, glass, rain, and fog mix the intensities of surfaces that move differently.

These violations are not small, evenly spread Gaussian noise but large errors confined to regions such as a highlight or an occluded strip along a moving edge. A quadratic penalty on the brightness residual lets these few pixels dominate the solution, which motivates the robust variants below.

Relaxed and Robust Variants

Most modern methods keep the idea of brightness constancy but change what is assumed constant or how violations are penalized. The choice appears as the data term ρ(I1(x+w)−I0(x))\rho\big(I_1(\mathbf{x} + \mathbf{w}) - I_0(\mathbf{x})\big).

Robust penalties

Black and Anandan replaced the quadratic penalty with robust functions from statistics, such as the Lorentzian ρ(r)=log⁡(1+12(r/σ)2)\rho(r) = \log\big(1 + \tfrac{1}{2}(r / \sigma)^2\big), whose influence on the solution decreases for large residuals, so that pixels violating brightness constancy are treated as outliers instead of being fitted. They applied the same idea to the smoothness term so that flow can change abruptly at motion boundaries. The Lorentzian is not convex; convex, differentiable approximations to the absolute value, such as the Charbonnier penalty r2+ε2\sqrt{r^2 + \varepsilon^2} used by Brox et al., are a common alternative because they keep the optimization well behaved.

Gradient constancy

Brox et al. added a gradient constancy assumption to the data term:

∇I1(x+w)=∇I0(x).\nabla I_1(\mathbf{x} + \mathbf{w}) = \nabla I_0(\mathbf{x}).

The gradient is unchanged by an additive change in brightness, so this term tolerates an intensity offset while still constraining motion through image structure. Their energy combines brightness and gradient constancy in a robustly penalized data term with a robust, discontinuity-preserving smoothness term.

Normalized and affine models

When intensities change by a gain and an offset, I1(x+w)=α I0(x)+βI_1(\mathbf{x} + \mathbf{w}) = \alpha\, I_0(\mathbf{x}) + \beta, zero-mean normalized cross-correlation removes both by normalizing each patch to zero mean and unit variance before comparison. Direct visual odometry takes the alternative route: Direct Sparse Odometry (DSO) estimates affine brightness parameters for each frame jointly with camera motion and depth, and can additionally use a photometric calibration of exposure time, vignetting, and the camera’s response function.

Non-parametric transforms

Zabih and Woodfill proposed the rank and census transforms, which replace each pixel by a description of the order relations between it and its neighbors. The census transform records one bit per neighbor stating whether that neighbor is darker than the center pixel. Any monotonic change of intensities leaves these bits unchanged, and an outlying neighbor changes only its own bit. Census costs are widely used in stereo, and census-based photometric losses are used to train unsupervised flow networks such as UnFlow.

Examples

  • Lucas–Kanade. Assumes brightness constancy and constant flow in a small window. Iterating the linearized solution with warping is Gauss–Newton minimization of the sum of squared brightness differences, as Baker and Matthews show in their unifying framework for Lucas–Kanade alignment.
  • Block matching in video codecs. Searches for the displacement that minimizes the sum of absolute differences between blocks, brightness constancy without linearization.
  • Stereo matching. For a rectified pair, brightness constancy becomes Iright(x−d,y)=Ileft(x,y)I_\text{right}(x - d, y) = I_\text{left}(x, y) for a disparity dd. Differences between the two cameras’ exposure and response add to the usual failure cases, one reason census costs are popular in stereo.

Practical Example

Estimating a global translation from the linear constraint, with and without re-linearization. The texture is smoothed white noise, so it varies over a few pixels.

import numpy as np
from scipy import ndimage

rng = np.random.default_rng(0)
# Smooth random texture: white noise blurred with a Gaussian (sigma = 3 pixels).
texture = ndimage.gaussian_filter(rng.standard_normal((256, 256)), sigma=3)
inner = (slice(24, -24), slice(24, -24))  # ignore borders affected by the shift

def linearized_step(I0, I1):
    """Least-squares solution of I_x u + I_y v + I_t = 0 over the whole image."""
    Iy, Ix = np.gradient((I0 + I1) / 2)
    It = I1 - I0
    A = np.stack([Ix[inner].ravel(), Iy[inner].ravel()], axis=1)
    return np.linalg.lstsq(A, -It[inner].ravel(), rcond=None)[0]

def iterated(I0, I1, steps=10):
    """Repeat the linearized step, each time warping I1 back by the current estimate."""
    w = np.zeros(2)
    for _ in range(steps):
        warped = ndimage.shift(I1, (-w[1], -w[0]), order=3, mode="nearest")
        w += linearized_step(I0, warped)
    return w

for d in [0.25, 1.0, 2.0, 4.0, 8.0, 16.0]:
    # Second frame: the same texture translated d pixels to the right (u = d, v = 0).
    I1 = ndimage.shift(texture, (0, d), order=3, mode="nearest")
    u1 = linearized_step(texture, I1)[0]
    uk = iterated(texture, I1)[0]
    print(f"true u = {d:5.2f}   one step: {u1:5.2f}   iterated: {uk:5.2f}")

Output:

true u =  0.25   one step:  0.26   iterated:  0.25
true u =  1.00   one step:  1.04   iterated:  1.00
true u =  2.00   one step:  2.16   iterated:  2.00
true u =  4.00   one step:  4.88   iterated:  4.00
true u =  8.00   one step:  4.82   iterated:  8.00
true u = 16.00   one step: -0.00   iterated: -0.01

A single linear step is accurate only for sub-pixel to one-pixel motion. Re-linearizing after warping recovers larger shifts, but once the displacement is far beyond the scale of the texture, even the iterated estimate fails; such motion must first be estimated at a coarser pyramid level.

Common Misconceptions

  • “The optical flow constraint equation is brightness constancy.” It is the first-order approximation of brightness constancy, valid only for small motion. The nonlinear equation, when it holds at all, holds for any displacement.
  • “Brightness constancy holds for static cameras.” It concerns the photometry of scene points, not camera motion. A static camera with automatic exposure violates it; a moving camera filming a matte scene under steady light may satisfy it well.
  • “Robust penalties fix occlusion.” They limit the damage occluded pixels do elsewhere, but cannot produce a match for a point that is not visible. Occluded regions are usually filled in from their surroundings.
  • “Learning-based methods no longer use it.” Supervised networks do not use it explicitly, but their cost volumes compare feature similarity, a learned form of constancy. Unsupervised flow and depth networks are trained directly with photometric losses.

Where It Is Used

  • Optical flow estimation. The data term of variational methods and the matching cost of local methods (optical flow).
  • Feature tracking. The KLT tracker minimizes the brightness difference between a patch and its displaced copy (feature tracking).
  • Direct alignment and odometry. Lucas–Kanade-style algorithms estimate a parametric warp such as a homography, and direct visual odometry and SLAM methods such as LSD-SLAM and DSO estimate camera motion and depth, by minimizing photometric error instead of matching features.
  • Stereo, multi-view stereo, and unsupervised learning of flow and depth, which rely on photo-consistency across views or frames, as do block matching and other forms of motion estimation.

Related

  • Optical Flow

    The apparent motion of image content between two frames, represented as a two-dimensional displacement at every pixel.

  • Aperture Problem

    Why motion seen through a small window is ambiguous along edges, so that only the component of motion across an edge can be measured locally.

  • Motion Estimation

    Recovering how image content, objects, or the camera moved from a sequence of images, from per-pixel flow to global and 3D motion.

  • Lucas–Kanade Method

    A local, gradient-based method that estimates the displacement of an image window by assuming constant motion within it and solving a small least-squares problem, iterated with warping.

  • Horn–Schunck Method

    A global variational method that computes dense optical flow by minimizing brightness constancy errors together with a penalty on spatial variation of the flow, solved by a simple iterative averaging scheme.

  • Optical Flow Estimation

    Computing a dense field of pixel displacements between two video frames, from classical variational methods to learned networks such as RAFT.

References

  1. Horn, B. K. P. & Schunck, B. G. (1981). Determining Optical Flow. Artificial Intelligence, 17(1–3), 185–203.
  2. Black, M. J. & Anandan, P. (1996). The Robust Estimation of Multiple Motions: Parametric and Piecewise-Smooth Flow Fields. Computer Vision and Image Understanding, 63(1), 75–104.
  3. Brox, T., Bruhn, A., Papenberg, N. & Weickert, J. (2004). High Accuracy Optical Flow Estimation Based on a Theory for Warping. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science 3024, 25–36.
  4. Zabih, R. & Woodfill, J. (1994). Non-parametric Local Transforms for Computing Visual Correspondence. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science 801, 151–158.
  5. Baker, S. & Matthews, I. (2004). Lucas-Kanade 20 Years On: A Unifying Framework. International Journal of Computer Vision, 56(3), 221–255.
  6. Engel, J., Koltun, V. & Cremers, D. (2018). Direct Sparse Odometry. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(3), 611–625.
  7. Meister, S., Hur, J. & Roth, S. (2018). UnFlow: Unsupervised Learning of Optical Flow with a Bidirectional Census Loss. AAAI Conference on Artificial Intelligence, 32(1).