Motion & Tracking
Brightness Constancy Assumption
The assumption that a scene point keeps the same image intensity as it moves between frames, which turns motion estimation into an intensity-matching problem.
intermediate
The brightness constancy assumption states that a point in the scene has the same intensity in every frame in which it is visible: as it moves across the image, its brightness travels with it. It is what makes correspondence measurable from pixel values alone, and it is the starting point of most classical optical flow methods, of direct image alignment, and of stereo matching costs. It is also called the intensity constancy or photometric consistency assumption.
Definition
Brightness constancy is the assumption that the intensity recorded for a scene point does not change between frames, so that any change in the intensity at a fixed pixel is caused only by motion. It is a modeling assumption about the scene, the lighting, and the camera, not a law of image formation.
Intuition
Track a single speck of paint on a moving object. If the light falling on it and the camera’s exposure stay the same, the speck has the same gray value in every frame, and finding where it went is a search for a patch with the same values. This is reasonable for matte surfaces under steady lighting between consecutive video frames. It is wrong for a mirror, a flickering light, or a shadow sweeping across a wall, and every method built on it inherits those failure cases.
Formal Definition
For frames and and a displacement field , brightness constancy states
This equation is nonlinear in , because appears inside the image function. Expanding around with a Taylor series,
where is the Hessian of , and keeping only the first-order term with gives the optical flow constraint equation:
The small-motion condition
Dropping the second-order term is valid only when it is small compared with the first-order term:
The displacement must be small compared with the distance over which the image gradient changes. For a sinusoidal pattern of wavelength , the linear model degrades long before reaches , and at the motion is ambiguous outright: shifting a sinusoid by half a period forward or backward produces the same image. Fine texture therefore allows only very small displacements; smooth, low-frequency content tolerates larger ones.
This is why practical methods blur the images before differentiating, work on an image pyramid from coarse to fine, and iterate: after each linear step they warp by the current estimate and linearize again, so each step only explains a small residual motion. Brox et al. minimized the original, non-linearized energy with nested fixed-point iterations and showed that this numerical scheme amounts to coarse-to-fine warping, giving a theoretical justification to a technique previously motivated by experiment.
The continuous form
In continuous time, brightness constancy along a trajectory means the total derivative of intensity vanishes:
Here the constraint is exact, with velocities in place of displacements; this is the form Horn and Schunck derived. The small-motion condition appears only because video is sampled at discrete times.
Properties
One equation per pixel
The constraint gives one equation for two unknowns at every pixel and fixes only the component of motion along the image gradient. This is the aperture problem, and every flow method adds a second assumption to resolve it.
When the assumption fails
Brightness constancy is violated whenever the intensity of a scene point changes for reasons other than motion:
- Illumination changes. Lights switching on, lamp flicker, passing clouds, and automatic exposure or gain control.
- Shading changes. A matte surface that rotates relative to the light changes its shading even under fixed lighting.
- Specular highlights and shadows. Highlights move with the viewpoint and the light, not with the surface; a moving shadow changes the brightness of a static surface.
- Occlusion and disocclusion. A point that becomes hidden, or newly visible, has no match in the other frame, and no photometric model can supply one.
- Noise, blur, and compression. Sensor noise enters directly as the difference of two noisy images; motion blur and compression artifacts change from frame to frame.
- Transparency and atmospheric effects. Smoke, glass, rain, and fog mix the intensities of surfaces that move differently.
These violations are not small, evenly spread Gaussian noise but large errors confined to regions such as a highlight or an occluded strip along a moving edge. A quadratic penalty on the brightness residual lets these few pixels dominate the solution, which motivates the robust variants below.
Relaxed and Robust Variants
Most modern methods keep the idea of brightness constancy but change what is assumed constant or how violations are penalized. The choice appears as the data term .
Robust penalties
Black and Anandan replaced the quadratic penalty with robust functions from statistics, such as the Lorentzian , whose influence on the solution decreases for large residuals, so that pixels violating brightness constancy are treated as outliers instead of being fitted. They applied the same idea to the smoothness term so that flow can change abruptly at motion boundaries. The Lorentzian is not convex; convex, differentiable approximations to the absolute value, such as the Charbonnier penalty used by Brox et al., are a common alternative because they keep the optimization well behaved.
Gradient constancy
Brox et al. added a gradient constancy assumption to the data term:
The gradient is unchanged by an additive change in brightness, so this term tolerates an intensity offset while still constraining motion through image structure. Their energy combines brightness and gradient constancy in a robustly penalized data term with a robust, discontinuity-preserving smoothness term.
Normalized and affine models
When intensities change by a gain and an offset, , zero-mean normalized cross-correlation removes both by normalizing each patch to zero mean and unit variance before comparison. Direct visual odometry takes the alternative route: Direct Sparse Odometry (DSO) estimates affine brightness parameters for each frame jointly with camera motion and depth, and can additionally use a photometric calibration of exposure time, vignetting, and the camera’s response function.
Non-parametric transforms
Zabih and Woodfill proposed the rank and census transforms, which replace each pixel by a description of the order relations between it and its neighbors. The census transform records one bit per neighbor stating whether that neighbor is darker than the center pixel. Any monotonic change of intensities leaves these bits unchanged, and an outlying neighbor changes only its own bit. Census costs are widely used in stereo, and census-based photometric losses are used to train unsupervised flow networks such as UnFlow.
Examples
- Lucas–Kanade. Assumes brightness constancy and constant flow in a small window. Iterating the linearized solution with warping is Gauss–Newton minimization of the sum of squared brightness differences, as Baker and Matthews show in their unifying framework for Lucas–Kanade alignment.
- Block matching in video codecs. Searches for the displacement that minimizes the sum of absolute differences between blocks, brightness constancy without linearization.
- Stereo matching. For a rectified pair, brightness constancy becomes for a disparity . Differences between the two cameras’ exposure and response add to the usual failure cases, one reason census costs are popular in stereo.
Practical Example
Estimating a global translation from the linear constraint, with and without re-linearization. The texture is smoothed white noise, so it varies over a few pixels.
import numpy as np
from scipy import ndimage
rng = np.random.default_rng(0)
# Smooth random texture: white noise blurred with a Gaussian (sigma = 3 pixels).
texture = ndimage.gaussian_filter(rng.standard_normal((256, 256)), sigma=3)
inner = (slice(24, -24), slice(24, -24)) # ignore borders affected by the shift
def linearized_step(I0, I1):
"""Least-squares solution of I_x u + I_y v + I_t = 0 over the whole image."""
Iy, Ix = np.gradient((I0 + I1) / 2)
It = I1 - I0
A = np.stack([Ix[inner].ravel(), Iy[inner].ravel()], axis=1)
return np.linalg.lstsq(A, -It[inner].ravel(), rcond=None)[0]
def iterated(I0, I1, steps=10):
"""Repeat the linearized step, each time warping I1 back by the current estimate."""
w = np.zeros(2)
for _ in range(steps):
warped = ndimage.shift(I1, (-w[1], -w[0]), order=3, mode="nearest")
w += linearized_step(I0, warped)
return w
for d in [0.25, 1.0, 2.0, 4.0, 8.0, 16.0]:
# Second frame: the same texture translated d pixels to the right (u = d, v = 0).
I1 = ndimage.shift(texture, (0, d), order=3, mode="nearest")
u1 = linearized_step(texture, I1)[0]
uk = iterated(texture, I1)[0]
print(f"true u = {d:5.2f} one step: {u1:5.2f} iterated: {uk:5.2f}")
Output:
true u = 0.25 one step: 0.26 iterated: 0.25
true u = 1.00 one step: 1.04 iterated: 1.00
true u = 2.00 one step: 2.16 iterated: 2.00
true u = 4.00 one step: 4.88 iterated: 4.00
true u = 8.00 one step: 4.82 iterated: 8.00
true u = 16.00 one step: -0.00 iterated: -0.01
A single linear step is accurate only for sub-pixel to one-pixel motion. Re-linearizing after warping recovers larger shifts, but once the displacement is far beyond the scale of the texture, even the iterated estimate fails; such motion must first be estimated at a coarser pyramid level.
Common Misconceptions
- “The optical flow constraint equation is brightness constancy.” It is the first-order approximation of brightness constancy, valid only for small motion. The nonlinear equation, when it holds at all, holds for any displacement.
- “Brightness constancy holds for static cameras.” It concerns the photometry of scene points, not camera motion. A static camera with automatic exposure violates it; a moving camera filming a matte scene under steady light may satisfy it well.
- “Robust penalties fix occlusion.” They limit the damage occluded pixels do elsewhere, but cannot produce a match for a point that is not visible. Occluded regions are usually filled in from their surroundings.
- “Learning-based methods no longer use it.” Supervised networks do not use it explicitly, but their cost volumes compare feature similarity, a learned form of constancy. Unsupervised flow and depth networks are trained directly with photometric losses.
Where It Is Used
- Optical flow estimation. The data term of variational methods and the matching cost of local methods (optical flow).
- Feature tracking. The KLT tracker minimizes the brightness difference between a patch and its displaced copy (feature tracking).
- Direct alignment and odometry. Lucas–Kanade-style algorithms estimate a parametric warp such as a homography, and direct visual odometry and SLAM methods such as LSD-SLAM and DSO estimate camera motion and depth, by minimizing photometric error instead of matching features.
- Stereo, multi-view stereo, and unsupervised learning of flow and depth, which rely on photo-consistency across views or frames, as do block matching and other forms of motion estimation.
Related
- Optical Flow
The apparent motion of image content between two frames, represented as a two-dimensional displacement at every pixel.
- Aperture Problem
Why motion seen through a small window is ambiguous along edges, so that only the component of motion across an edge can be measured locally.
- Motion Estimation
Recovering how image content, objects, or the camera moved from a sequence of images, from per-pixel flow to global and 3D motion.
- Lucas–Kanade Method
A local, gradient-based method that estimates the displacement of an image window by assuming constant motion within it and solving a small least-squares problem, iterated with warping.
- Horn–Schunck Method
A global variational method that computes dense optical flow by minimizing brightness constancy errors together with a penalty on spatial variation of the flow, solved by a simple iterative averaging scheme.
- Optical Flow Estimation
Computing a dense field of pixel displacements between two video frames, from classical variational methods to learned networks such as RAFT.
References
- Horn, B. K. P. & Schunck, B. G. (1981). Determining Optical Flow. Artificial Intelligence, 17(1–3), 185–203.
- Black, M. J. & Anandan, P. (1996). The Robust Estimation of Multiple Motions: Parametric and Piecewise-Smooth Flow Fields. Computer Vision and Image Understanding, 63(1), 75–104.
- Brox, T., Bruhn, A., Papenberg, N. & Weickert, J. (2004). High Accuracy Optical Flow Estimation Based on a Theory for Warping. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science 3024, 25–36.
- Zabih, R. & Woodfill, J. (1994). Non-parametric Local Transforms for Computing Visual Correspondence. European Conference on Computer Vision (ECCV), Lecture Notes in Computer Science 801, 151–158.
- Baker, S. & Matthews, I. (2004). Lucas-Kanade 20 Years On: A Unifying Framework. International Journal of Computer Vision, 56(3), 221–255.
- Engel, J., Koltun, V. & Cremers, D. (2018). Direct Sparse Odometry. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(3), 611–625.
- Meister, S., Hur, J. & Roth, S. (2018). UnFlow: Unsupervised Learning of Optical Flow with a Bidirectional Census Loss. AAAI Conference on Artificial Intelligence, 32(1).