Image Processing & Computational Photography
Image Warping
Transforming the geometry of an image by a coordinate mapping, usually computed by sampling the input image at the inverse-mapped position of every output pixel.
intermediate
Image warping changes where things are in an image rather than what they look like, moving content to new positions given by a coordinate mapping: a rotation, a perspective change, or a different displacement at every pixel. It is how images are aligned, how motion estimates are tested and refined, how panoramas are assembled, and how networks resample images and features by a predicted motion.
Definition
Image warping is the geometric transformation of an image by a coordinate mapping that takes positions in the input image to positions in the output image. The output has the input’s content at the transformed positions:
Because generally falls between pixel centers, computing a warped image requires resampling: interpolating the image at non-integer positions, and filtering it when the warp shrinks the image.
Intuition
Picture the image printed on a rubber sheet that is stretched, rotated, or pushed around locally; warping photographs the deformed sheet on a fresh pixel grid. There are two ways to do it. Forward mapping takes every input pixel and moves it to where sends it. Backward mapping, also called inverse mapping, takes every output pixel and asks where it came from, looking up the input at :
Forward mapping is natural but leaves gaps. Where the warp enlarges the image, some output pixels receive nothing (holes); where it shrinks the image, several input pixels land in the same output pixel (overlaps), and something must decide how to combine them. Backward mapping avoids both: every output pixel gets exactly one value, interpolated from the input around a non-integer source position. This is why backward mapping with interpolation is the standard way to warp. It needs the inverse mapping, which parametric warps provide and dense warps obtain by defining the displacement field on the output grid.
Formal Definition
Parametric warps
A parametric warp is described by a few parameters shared by the whole image. In homogeneous coordinates , the common planar warps are matrices, nested from the most to the least constrained:
A similarity transform preserves angles, an affine transform preserves parallel lines, and a homography (projective transform) preserves straight lines. A homography is defined only up to scale, so the transformed point must be divided by its third coordinate:
A homography relates two images of a planar scene, or two images taken by a camera rotating about its center. The backward warp of any of these matrices is the warp by its inverse.
Dense warps
A non-parametric warp gives every pixel its own displacement , as an optical flow field or a displacement map does. Warping the second frame of a pair by the flow from the first to the second brings it back onto the first frame’s grid:
Where the flow is correct and nothing is occluded, , so the residual measures what the flow fails to explain. This is a backward warp, because is defined on the grid of the image being produced.
Interpolation
The source position is real-valued; write it as with integers and . Three interpolation schemes are common:
- Nearest neighbor takes the closest pixel. It keeps values unchanged, which suits label maps, but produces jagged results and position errors of up to half a pixel in each coordinate.
- Bilinear interpolation weights the four surrounding pixels, as in the formula below. It is continuous and the usual default, but slightly blurs the image, and its gradient is discontinuous at pixel boundaries.
- Bicubic interpolation weights a neighborhood with a cubic convolution kernel; cubic B-spline interpolation, as in SciPy, uses the same support after prefiltering the whole image. Both are sharper and smoother than bilinear, at the cost of mild ringing near strong edges.
Here is the pixel in column and row , and bilinear interpolation is
Properties
- Backward warps need , or a displacement field defined on the output grid. A forward flow field cannot be used directly for backward warping of the other frame; it must be inverted or re-estimated in the other direction.
- Composition. Parametric warps of one kind compose by matrix multiplication and stay in the same family. Translations, similarities, affine transforms, and homographies each form a group, which inverse compositional alignment relies on (Baker and Matthews, 2004).
- Interpolation error accumulates. Each resampling blurs a little, so a chain of warps should be composed into one mapping and applied once.
- Minification aliases. When the warp shrinks part of the image, an output pixel spans several input pixels, and interpolating at a single point aliases fine detail. Prefilter the input by the local scale factor first, or sample from a prefiltered image pyramid. Williams’ mipmaps (1983) store such a pyramid for texture mapping and filter each lookup by interpolating bilinearly within, and linearly between, the two levels nearest the required scale.
- Borders. Source positions outside the input have no data. They can be filled with a constant, by replicating or reflecting the border, or by wrapping for periodic images, but none of these is real content. In estimation, such pixels are best marked invalid and excluded from residuals and losses.
- Occlusions. A dense backward warp fills even pixels whose content is hidden in the source frame, copying whatever is visible there instead, so occluded regions need a separate mask.
Examples
The code warps an analytic test image by an affine transform with backward mapping and compares three interpolation orders with the exact answer, counts holes and overlaps under forward mapping, warps a frame back by a known dense flow field, and checks a pure integer shift.
import numpy as np
from scipy import ndimage
H, W = 200, 240
y, x = np.mgrid[0:H, 0:W].astype(float)
def scene(x, y):
"""A smooth test image that we can evaluate at any real-valued position."""
return np.sin(x / 7.0) * np.cos(y / 5.0) + 0.5 * np.sin((x + 2 * y) / 11.0)
image = scene(x, y)
# 1. Parametric warp: rotate by 10 degrees and scale by 1.2 about the image center.
theta, s = np.deg2rad(10), 1.2
A = s * np.array([[np.cos(theta), -np.sin(theta)],
[np.sin(theta), np.cos(theta)]])
c = np.array([W / 2, H / 2])
t = c - A @ c # forward map: x_dst = A x_src + t
# Backward mapping: for every output pixel, find where it came from in the input.
dst = np.stack([x.ravel(), y.ravel()])
src = np.linalg.solve(A, dst - t[:, None])
truth = scene(src[0], src[1]).reshape(H, W) # exact value of the warped image
# Score only output pixels whose source lies at least 3 px inside the input.
inside = ((src[0] >= 3) & (src[0] <= W - 4) & (src[1] >= 3) & (src[1] <= H - 4)).reshape(H, W)
for order, name in [(0, "nearest"), (1, "bilinear"), (3, "cubic spline")]:
# map_coordinates takes (row, column) coordinates.
warped = ndimage.map_coordinates(image, [src[1], src[0]], order=order,
mode="constant", cval=np.nan).reshape(H, W)
err = np.abs(warped - truth)[inside]
print(f"affine, {name:12s}: mean abs error {err.mean():.1e}, max {err.max():.1e}")
# 2. Forward mapping: push every input pixel to its rounded destination.
def forward_hits(scale):
M = scale * A / s
tt = c - M @ c
fx = np.rint(M[0, 0] * x + M[0, 1] * y + tt[0]).astype(int)
fy = np.rint(M[1, 0] * x + M[1, 1] * y + tt[1]).astype(int)
ok = (fx >= 0) & (fx < W) & (fy >= 0) & (fy < H)
hits = np.zeros((H, W), int)
np.add.at(hits, (fy[ok], fx[ok]), 1)
core = hits[H // 2 - 40:H // 2 + 40, W // 2 - 40:W // 2 + 40] # region covered either way
return np.mean(core == 0), np.mean(core > 1)
for scale in [1.2, 0.8]:
holes, overlaps = forward_hits(scale)
print(f"forward mapping, scale {scale}: {holes:.1%} of output pixels get no value, "
f"{overlaps:.1%} get several")
# 3. Dense warp with a smooth, known flow field w = (u, v) defined on frame 0.
u = 3.25 + 1.5 * np.sin(y / 30.0)
v = -1.5 + 1.0 * np.cos(x / 40.0)
frame1 = image
frame0 = scene(x + u, y + v) # content at x in frame 0 is at x + w(x) in frame 1
# Backward warp: sample frame 1 at x + w(x) to bring it onto frame 0's grid.
back = ndimage.map_coordinates(frame1, [y + v, x + u], order=3, mode="nearest")
valid = (x + u >= 3) & (x + u <= W - 4) & (y + v >= 3) & (y + v <= H - 4)
print(f"dense flow: mean |frame0 - frame1| = {np.abs(frame0 - frame1)[valid].mean():.4f}, "
f"mean |frame0 - warped frame1| = {np.abs(frame0 - back)[valid].mean():.1e}")
# 4. Constant shift check: warping by w = (3, 0) moves content exactly 3 pixels.
shifted = ndimage.map_coordinates(image, [y, x - 3.0], order=1, mode="nearest")
print("pure shift by 3 px: output[:, 3:] equals input[:, :-3]:",
np.allclose(shifted[:, 3:], image[:, :-3]))
Output:
affine, nearest : mean abs error 3.2e-02, max 1.6e-01
affine, bilinear : mean abs error 2.3e-03, max 1.0e-02
affine, cubic spline: mean abs error 1.6e-06, max 8.3e-04
forward mapping, scale 1.2: 30.6% of output pixels get no value, 0.0% get several
forward mapping, scale 0.8: 0.0% of output pixels get no value, 48.2% get several
dense flow: mean |frame0 - frame1| = 0.2474, mean |frame0 - warped frame1| = 3.4e-06
pure shift by 3 px: output[:, 3:] equals input[:, :-3]: True
On this smooth image, each step up in interpolation order reduces the error by an order of magnitude or more; SciPy’s order=3 is cubic B-spline interpolation, which prefilters the image so that the spline passes through the samples. Forward mapping of the same rotation leaves almost a third of the output pixels empty when it enlarges by 1.2, and sends several input pixels to almost half of the output pixels when it shrinks by 0.8. Warping the second frame by the known flow removes the difference between the frames up to interpolation error, and a whole-pixel shift moves the content exactly.
Common Misconceptions
- “Warping moves pixels to their new positions.” Implementations almost always do the opposite: they visit output pixels and fetch values from the input. The mapping to implement is the inverse of the one that describes the motion.
- “Bilinear interpolation is enough for any warp.” It is adequate for magnification and mild changes of scale, but aliases when the warp shrinks the image by a large factor; a prefilter or a pyramid is needed.
- “Pixel coordinates are unambiguous.” Libraries differ in whether is a pixel center or a corner, whether coordinates are (row, column) or , and whether they are normalized to . Half-pixel offsets between these conventions are a frequent bug.
Where It Is Used
- Image registration and alignment. Aligning two images means finding the warp that makes one match the other, then applying it. Direct methods compare pixel intensities after warping; feature-based methods estimate the warp from feature matches and then resample. Szeliski’s tutorial (2007) covers both, starting from the hierarchy of motion models above.
- The Lucas–Kanade method. The Lucas–Kanade algorithm warps the image by the current parameters of a warp , interpolating at sub-pixel positions, measures the residual against the template, and solves for an update. Baker and Matthews (2004) classify its variants by how the warp is updated; the efficient inverse compositional form needs warps that form a group, which holds for homographies.
- Coarse-to-fine optical flow. At each pyramid level, coarse-to-fine estimation warps the second image by the flow upsampled from the coarser level, so that only a small residual motion remains to be estimated.
- Learned optical flow. As described in optical flow estimation, FlowNet 2.0 warps the second image by the intermediate flow between its stacked networks, and PWC-Net warps the second frame’s features at every pyramid level. RAFT does not warp images or features; it uses a related operation, bilinearly sampling its correlation volume around .
- Video frame interpolation. Super SloMo (Jiang et al., 2018) estimates bidirectional flow between the two input frames, combines and refines it into flows for the intermediate time, warps both inputs to that time, and fuses them with predicted visibility maps that exclude occluded pixels.
- Panoramas and stitching. Each photograph is warped onto a common compositing surface, such as a plane, cylinder, or sphere, using the homographies or rotations estimated during alignment, and the warped images are then blended, for example with pyramid blending.
- Differentiable warping. Bilinear sampling is differentiable with respect to both the image values and the sampling positions, so a warp can sit inside a network trained by backpropagation. Spatial Transformer Networks (Jaderberg et al., 2015) make this a module: a localization network predicts the warp parameters, a grid generator maps the regular output grid back to source positions (a backward warp), and a sampler interpolates the input there. The authors used affine, projective, and thin plate spline warps, needed sub-gradients because of the sampling kernel’s discontinuities, and noted that downsampling with a small kernel such as bilinear can alias.
- Other uses. Lens distortion correction, stereo rectification, geometric data augmentation, and texture mapping are typically implemented as backward warps with interpolation.
Related
- Image Pyramid
A multi-scale representation of an image built by repeatedly smoothing and subsampling it, so that each level holds a coarser copy at half the resolution of the one below.
- Lucas–Kanade Method
A local, gradient-based method that estimates the displacement of an image window by assuming constant motion within it and solving a small least-squares problem, iterated with warping.
- Coarse-to-Fine Estimation
Estimating large motions with small-motion methods by solving on an image pyramid, from the coarsest level to full resolution, and refining the estimate at each level.
- Optical Flow Estimation
Computing a dense field of pixel displacements between two video frames, from classical variational methods to learned networks such as RAFT.
- Feature Matching
Finding corresponding keypoints between two or more images of the same scene, from descriptor nearest-neighbor search to learned matchers.
- RAFT
A deep network for optical flow that matches all pairs of pixels once and refines a single flow field with a recurrent update operator.
References
- Szeliski, R. (2007). Image Alignment and Stitching: A Tutorial. Foundations and Trends in Computer Graphics and Vision, 2(1), 1–104.
- Baker, S. & Matthews, I. (2004). Lucas-Kanade 20 Years On: A Unifying Framework. International Journal of Computer Vision, 56(3), 221–255.
- Williams, L. (1983). Pyramidal Parametrics. ACM SIGGRAPH Computer Graphics, 17(3), 1–11.
- Jaderberg, M., Simonyan, K., Zisserman, A. & Kavukcuoglu, K. (2015). Spatial Transformer Networks. Advances in Neural Information Processing Systems (NeurIPS), 28.
- Jiang, H., Sun, D., Jampani, V., Yang, M.-H., Learned-Miller, E. & Kautz, J. (2018). Super SloMo: High Quality Estimation of Multiple Intermediate Frames for Video Interpolation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 9000–9008.