Foundations
Gaussian Distribution
The bell-shaped probability distribution defined by a mean and a covariance, the default model for noise and uncertainty in estimation and tracking.
beginner
The Gaussian distribution, also called the normal distribution, is the bell-shaped probability distribution fully described by its mean and its variance, or in several dimensions by a mean vector and a covariance matrix. It is the default model for measurement noise and for uncertainty about a quantity being estimated. In computer vision it underlies least squares, the Kalman filter, Mahalanobis gating in data association, and the standard models of image noise.
Its importance comes less from nature being exactly Gaussian than from its mathematical convenience: linear operations, conditioning, and the fusion of independent measurements all turn Gaussians into Gaussians, so estimates can be computed in closed form. It is named after Carl Friedrich Gauss, whose 1809 work on the orbits of celestial bodies derived this law of errors from the assumption that the arithmetic mean of repeated observations is the most probable value, and used it to justify least squares.
Definition
A random variable is Gaussian if its probability density is the exponential of a negative quadratic function of the variable. In one dimension the density is a symmetric bell centered on the mean, with a width set by the standard deviation. A random vector is (jointly) Gaussian if every linear combination of its components is Gaussian; its distribution is then determined by its mean vector and covariance matrix.
Intuition
A Gaussian expresses a belief of the form “the value is probably near , and errors of a few standard deviations are increasingly unlikely”. Probability falls off with the square of the distance from the mean, so it decays very quickly: in one dimension, about 68% of the probability lies within one standard deviation, 95% within two, and 99.7% within three.
In two dimensions, think of a detector reporting the position of an object. If errors are larger horizontally than vertically, or tend to be positive in when they are positive in , the cloud of reported positions forms a tilted ellipse rather than a circle. The covariance matrix describes the size, shape, and orientation of that ellipse.
Formal Definition
A scalar Gaussian with mean and variance , written , has density
A -dimensional Gaussian with mean and symmetric positive-definite covariance has density
where is the determinant. The parameters are the first two moments:
The diagonal entries of are the variances of the components and the off-diagonal entries their covariances. A singular describes a degenerate Gaussian concentrated on a lower-dimensional subspace; it has no density in , but most of the properties below carry over.
Properties
Geometry of the covariance
The density is constant on the ellipsoids . Writing the eigendecomposition , the ellipsoid’s axes point along the eigenvectors (the columns of ), and its semi-axis lengths are . The eigenvectors are therefore the directions of independent variation, and the standard deviation along each. A covariance proportional to the identity gives circles; unequal eigenvalues give elongated ellipses, such as the uncertainty of optical flow measured on an edge (aperture problem).
Mahalanobis distance
The quantity in the exponent defines the Mahalanobis distance, named after Prasanta Chandra Mahalanobis, whose 1936 paper set out this generalized distance:
It measures distance in standard deviations along each principal axis, so a point far from the mean along a direction of large variance can be closer, in this sense, than a nearby point along a direction of small variance. If , then follows a chi-squared distribution with degrees of freedom. This gives confidence regions: in two dimensions, the 95% ellipse is . Trackers use the same test to gate detections, rejecting those too unlikely to belong to a track.
Closure properties
These properties, standard results for the multivariate normal, are what make Gaussian models computable in closed form.
Linear transformations. If , then for any matrix and vector ,
The Kalman filter’s covariance prediction, , is this rule applied to the motion model, with the process noise added by the next property.
Sums of independent Gaussians. If and are independent, then . Means add and covariances add, which is why independent noise sources combine by adding variances, not standard deviations.
Marginals and conditionals. Partition into blocks and , with mean and covariance partitioned accordingly. The marginal of is Gaussian with mean and covariance : just drop the other block. The conditional is also Gaussian:
Observing shifts the mean of in proportion to how they covary and never increases its uncertainty. The Kalman filter’s update step is this formula, with the predicted state as and the measurement as : then , is the innovation covariance , is the Kalman gain, and is the innovation.
Products of Gaussian densities. The product of two Gaussian densities in the same variable is, up to normalization, another Gaussian:
This is how two independent estimates of the same quantity are fused: inverse covariances (information) add, and the fused mean is the information-weighted average. In one dimension, fusing a prediction of variance with a measurement of variance moves the estimate a fraction of the way toward the measurement, exactly the scalar Kalman gain.
Central limit theorem
The suitably normalized sum of many independent random variables with finite variance tends to a Gaussian, whatever their individual distributions. This is one reason Gaussian noise models work: an error made of many small, independent contributions is approximately Gaussian. It also explains why repeatedly applying a box filter to an image approximates Gaussian blur.
Connection to least squares and Bayesian estimation
The negative log of a Gaussian density is a quadratic form plus a constant. Maximizing a Gaussian likelihood is therefore minimizing a sum of squared, covariance-weighted errors, which is why least squares is the maximum-likelihood estimate under Gaussian noise. In Bayesian filtering, linear models with Gaussian noise keep the belief Gaussian at every step, so the recursion reduces to updating a mean and a covariance: the Kalman filter. Kalman’s 1960 paper notes that its optimality holds for Gaussian processes or, without that assumption, among linear estimators under quadratic loss.
Examples
Sampling, estimation, and Mahalanobis distance
The code samples a correlated 2D Gaussian, estimates its parameters, finds its principal axes, and compares Euclidean and Mahalanobis distances.
import numpy as np
rng = np.random.default_rng(0)
mu = np.array([2.0, 1.0])
Sigma = np.array([[4.0, 1.5],
[1.5, 1.0]])
# Sample: x = mu + L z, with L the Cholesky factor of Sigma and z ~ N(0, I).
L = np.linalg.cholesky(Sigma)
X = mu + rng.standard_normal((5000, 2)) @ L.T
mu_hat = X.mean(axis=0)
Sigma_hat = np.cov(X, rowvar=False)
print("estimated mean:", np.round(mu_hat, 2))
print("estimated covariance:\n", np.round(Sigma_hat, 2))
# Principal axes of the ellipses: eigenvectors and standard deviations.
evals, evecs = np.linalg.eigh(Sigma) # ascending eigenvalues
minor, major = evecs[:, 0], evecs[:, 1] * np.sign(evecs[0, 1])
print("axis std devs:", np.round(np.sqrt(evals), 2))
print("major axis direction:", np.round(major, 2))
def mahalanobis(x, mu, Sigma):
d = x - mu
return np.sqrt(d @ np.linalg.solve(Sigma, d))
# Two points at the same Euclidean distance (2) from the mean.
for p in [mu + 2 * major, mu + 2 * minor]:
print(f"point {np.round(p, 2)}: Euclidean 2.00, "
f"Mahalanobis {mahalanobis(p, mu, Sigma):.2f}")
# Squared Mahalanobis distance is chi-squared with 2 degrees of freedom;
# its 95% quantile is 5.991.
D = X - mu
d2 = np.einsum("ij,ij->i", D @ np.linalg.inv(Sigma), D)
print(f"fraction inside the 95% ellipse: {np.mean(d2 < 5.991):.3f}")
Output:
estimated mean: [2.02 1.01]
estimated covariance:
[[4.05 1.48]
[1.48 0.97]]
axis std devs: [0.62 2.15]
major axis direction: [0.92 0.38]
point [3.85 1.77]: Euclidean 2.00, Mahalanobis 0.93
point [ 2.77 -0.85]: Euclidean 2.00, Mahalanobis 3.25
fraction inside the 95% ellipse: 0.948
The sample estimates are close to the true parameters. Both test points are two units from the mean, but the one along the major axis is less than one standard deviation away, while the one along the minor axis is more than three. The fraction of samples inside the 95% ellipse matches the chi-squared prediction.
Image noise
The simplest image noise model adds independent, zero-mean Gaussian noise of fixed variance to every pixel, a common assumption in denoising experiments. Real sensor noise is signal-dependent: Foi and colleagues model raw data as a Poissonian part for photon sensing, whose variance grows with intensity, plus a Gaussian part for the remaining stationary disturbances. Because Poisson noise at moderate and high counts is close to Gaussian, the combination is commonly approximated by a Gaussian whose variance is an affine function of the noise-free intensity.
Gaussian filters and mixtures
The Gaussian function also appears as a kernel rather than as a model of noise. Gaussian smoothing convolves an image with a sampled Gaussian, which suppresses noise before computing image gradients and before subsampling in image pyramids. The kernel’s width is a scale parameter; no claim about the image’s statistics is implied.
A Gaussian mixture model is a weighted sum of several Gaussians. It represents multimodal distributions, such as the colors of a background pixel that alternates between leaves and sky, and is typically fitted with the expectation–maximization algorithm.
Common Misconceptions
- “Uncorrelated means independent.” True for components of a jointly Gaussian vector, false in general: two variables can each be Gaussian and uncorrelated yet dependent, because Gaussian marginals do not imply a jointly Gaussian distribution.
- “The one-sigma ellipse contains 68% of the probability.” Only in one dimension. In two dimensions the ellipse contains about 39%, and about 86%; confidence regions must use chi-squared quantiles for the right dimension.
- “Using a Gaussian filter assumes Gaussian noise.” The smoothing kernel and the noise model are separate choices.
Limitations
- Light tails. The Gaussian assigns vanishing probability to large errors, so a model built on it treats an outlier, such as a false detection or a mismatched feature, as overwhelming evidence and lets it pull the estimate far off. Heavy-tailed alternatives (Laplace, Student’s t) and robust penalties reduce the influence of such points.
- Unimodality. A single Gaussian cannot represent “the object is behind one of two occluders” or “the match is one of three repeated window corners”. Mixtures, multiple-hypothesis trackers, and particle filters, which represent a belief with weighted samples, handle such beliefs at higher cost.
- Nonlinearity. A Gaussian pushed through a nonlinear function, such as perspective projection, is no longer Gaussian. Extended and unscented Kalman filters approximate the result with a new Gaussian, which works when the nonlinearity is mild relative to the uncertainty.
- Covariance must be known or estimated. Errors in propagate into gating thresholds, filter gains, and confidence regions.
Where It Is Used
- Estimation: least squares as maximum likelihood, with covariances for the fitted parameters.
- Tracking: the Kalman filter and other Gaussian approximations to the Bayes filter; Mahalanobis gating in data association and costs for the Hungarian algorithm.
- Image processing: noise models for denoising; Gaussian kernels for smoothing, derivatives, and scale spaces.
- Machine learning: Gaussian mixture models for background subtraction and clustering, and Gaussian priors and likelihoods in probabilistic models.
Related
- Least Squares
The method of fitting a model to more measurements than unknowns by minimizing the sum of squared residuals, the workhorse of estimation in computer vision.
- Kalman Filter
A recursive algorithm that estimates the hidden state of a linear dynamic system from a sequence of noisy measurements, widely used to smooth and predict object positions in tracking.
- Bayesian Filtering
Recursive estimation of the probability distribution of a hidden, changing state from a sequence of noisy measurements, by alternating a motion-model prediction with a Bayes' rule update.
- Data Association
Deciding which measurements or detections belong to which tracked targets, and which are false alarms, missed detections, new targets, or targets that have disappeared.
- Image Pyramid
A multi-scale representation of an image built by repeatedly smoothing and subsampling it, so that each level holds a coarser copy at half the resolution of the one below.
References
- Gauss, C. F. (1809). Theoria motus corporum coelestium in sectionibus conicis solem ambientium. Hamburg: Perthes & Besser.
- Mahalanobis, P. C. (1936). On the Generalised Distance in Statistics. Proceedings of the National Institute of Sciences of India, 2(1), 49–55. Reprinted in Sankhyā A, 80(Suppl. 1), 1–7 (2018).
- Anderson, T. W. (2003). An Introduction to Multivariate Statistical Analysis (3rd ed.). Hoboken, NJ: Wiley.
- Kalman, R. E. (1960). A New Approach to Linear Filtering and Prediction Problems. Journal of Basic Engineering, 82(1), 35–45.
- Foi, A., Trimeche, M., Katkovnik, V. & Egiazarian, K. (2008). Practical Poissonian-Gaussian Noise Modeling and Fitting for Single-Image Raw-Data. IEEE Transactions on Image Processing, 17(10), 1737–1754.