Foundations

Gaussian Distribution

The bell-shaped probability distribution defined by a mean and a covariance, the default model for noise and uncertainty in estimation and tracking.

beginner

The Gaussian distribution, also called the normal distribution, is the bell-shaped probability distribution fully described by its mean and its variance, or in several dimensions by a mean vector and a covariance matrix. It is the default model for measurement noise and for uncertainty about a quantity being estimated. In computer vision it underlies least squares, the Kalman filter, Mahalanobis gating in data association, and the standard models of image noise.

Its importance comes less from nature being exactly Gaussian than from its mathematical convenience: linear operations, conditioning, and the fusion of independent measurements all turn Gaussians into Gaussians, so estimates can be computed in closed form. It is named after Carl Friedrich Gauss, whose 1809 work on the orbits of celestial bodies derived this law of errors from the assumption that the arithmetic mean of repeated observations is the most probable value, and used it to justify least squares.

Definition

A random variable is Gaussian if its probability density is the exponential of a negative quadratic function of the variable. In one dimension the density is a symmetric bell centered on the mean, with a width set by the standard deviation. A random vector is (jointly) Gaussian if every linear combination of its components is Gaussian; its distribution is then determined by its mean vector and covariance matrix.

Intuition

A Gaussian expresses a belief of the form “the value is probably near μ\mu, and errors of a few standard deviations are increasingly unlikely”. Probability falls off with the square of the distance from the mean, so it decays very quickly: in one dimension, about 68% of the probability lies within one standard deviation, 95% within two, and 99.7% within three.

In two dimensions, think of a detector reporting the position of an object. If errors are larger horizontally than vertically, or tend to be positive in xx when they are positive in yy, the cloud of reported positions forms a tilted ellipse rather than a circle. The covariance matrix describes the size, shape, and orientation of that ellipse.

Formal Definition

A scalar Gaussian with mean μ\mu and variance σ2\sigma^2, written x∼N(μ,σ2)x \sim \mathcal{N}(\mu, \sigma^2), has density

N(x; μ,σ2)=12πσ2exp⁡ ⁣(−(x−μ)22σ2).\mathcal{N}(x;\, \mu, \sigma^2) = \frac{1}{\sqrt{2\pi\sigma^2}} \exp\!\left( -\frac{(x - \mu)^2}{2\sigma^2} \right).

A dd-dimensional Gaussian with mean μ∈Rd\boldsymbol{\mu} \in \mathbb{R}^d and symmetric positive-definite covariance Σ∈Rd×d\Sigma \in \mathbb{R}^{d \times d} has density

N(x; μ,Σ)=1(2π)d/2∣Σ∣1/2exp⁡ ⁣(−12(x−μ)⊤Σ−1(x−μ)),\mathcal{N}(\mathbf{x};\, \boldsymbol{\mu}, \Sigma) = \frac{1}{(2\pi)^{d/2} \lvert \Sigma \rvert^{1/2}} \exp\!\left( -\tfrac{1}{2} (\mathbf{x} - \boldsymbol{\mu})^\top \Sigma^{-1} (\mathbf{x} - \boldsymbol{\mu}) \right),

where ∣Σ∣\lvert \Sigma \rvert is the determinant. The parameters are the first two moments:

μ=E[x],Σ=E[(x−μ)(x−μ)⊤].\boldsymbol{\mu} = \mathbb{E}[\mathbf{x}], \qquad \Sigma = \mathbb{E}\big[ (\mathbf{x} - \boldsymbol{\mu})(\mathbf{x} - \boldsymbol{\mu})^\top \big].

The diagonal entries of Σ\Sigma are the variances of the components and the off-diagonal entries their covariances. A singular Σ\Sigma describes a degenerate Gaussian concentrated on a lower-dimensional subspace; it has no density in Rd\mathbb{R}^d, but most of the properties below carry over.

Properties

Geometry of the covariance

The density is constant on the ellipsoids (x−μ)⊤Σ−1(x−μ)=c2(\mathbf{x} - \boldsymbol{\mu})^\top \Sigma^{-1} (\mathbf{x} - \boldsymbol{\mu}) = c^2. Writing the eigendecomposition Σ=VΛV⊤\Sigma = V \Lambda V^\top, the ellipsoid’s axes point along the eigenvectors (the columns of VV), and its semi-axis lengths are cλic\sqrt{\lambda_i}. The eigenvectors are therefore the directions of independent variation, and λi\sqrt{\lambda_i} the standard deviation along each. A covariance proportional to the identity gives circles; unequal eigenvalues give elongated ellipses, such as the uncertainty of optical flow measured on an edge (aperture problem).

Mahalanobis distance

The quantity in the exponent defines the Mahalanobis distance, named after Prasanta Chandra Mahalanobis, whose 1936 paper set out this generalized distance:

dM(x)=(x−μ)⊤Σ−1(x−μ).d_M(\mathbf{x}) = \sqrt{(\mathbf{x} - \boldsymbol{\mu})^\top \Sigma^{-1} (\mathbf{x} - \boldsymbol{\mu})}.

It measures distance in standard deviations along each principal axis, so a point far from the mean along a direction of large variance can be closer, in this sense, than a nearby point along a direction of small variance. If x∼N(μ,Σ)\mathbf{x} \sim \mathcal{N}(\boldsymbol{\mu}, \Sigma), then dM2d_M^2 follows a chi-squared distribution with dd degrees of freedom. This gives confidence regions: in two dimensions, the 95% ellipse is dM2≤5.991d_M^2 \le 5.991. Trackers use the same test to gate detections, rejecting those too unlikely to belong to a track.

Closure properties

These properties, standard results for the multivariate normal, are what make Gaussian models computable in closed form.

Linear transformations. If x∼N(μ,Σ)\mathbf{x} \sim \mathcal{N}(\boldsymbol{\mu}, \Sigma), then for any matrix AA and vector c\mathbf{c},

Ax+c∼N(Aμ+c,  AΣA⊤).A\mathbf{x} + \mathbf{c} \sim \mathcal{N}(A\boldsymbol{\mu} + \mathbf{c},\; A \Sigma A^\top).

The Kalman filter’s covariance prediction, Pk∣k−1=FPk−1∣k−1F⊤+QP_{k \mid k-1} = F P_{k-1 \mid k-1} F^\top + Q, is this rule applied to the motion model, with the process noise added by the next property.

Sums of independent Gaussians. If x∼N(μ1,Σ1)\mathbf{x} \sim \mathcal{N}(\boldsymbol{\mu}_1, \Sigma_1) and y∼N(μ2,Σ2)\mathbf{y} \sim \mathcal{N}(\boldsymbol{\mu}_2, \Sigma_2) are independent, then x+y∼N(μ1+μ2,Σ1+Σ2)\mathbf{x} + \mathbf{y} \sim \mathcal{N}(\boldsymbol{\mu}_1 + \boldsymbol{\mu}_2, \Sigma_1 + \Sigma_2). Means add and covariances add, which is why independent noise sources combine by adding variances, not standard deviations.

Marginals and conditionals. Partition x\mathbf{x} into blocks xa\mathbf{x}_a and xb\mathbf{x}_b, with mean and covariance partitioned accordingly. The marginal of xa\mathbf{x}_a is Gaussian with mean μa\boldsymbol{\mu}_a and covariance Σaa\Sigma_{aa}: just drop the other block. The conditional is also Gaussian:

xa∣xb∼N ⁣(μa+ΣabΣbb−1(xb−μb),  Σaa−ΣabΣbb−1Σba).\mathbf{x}_a \mid \mathbf{x}_b \sim \mathcal{N}\!\big( \boldsymbol{\mu}_a + \Sigma_{ab} \Sigma_{bb}^{-1} (\mathbf{x}_b - \boldsymbol{\mu}_b),\; \Sigma_{aa} - \Sigma_{ab} \Sigma_{bb}^{-1} \Sigma_{ba} \big).

Observing xb\mathbf{x}_b shifts the mean of xa\mathbf{x}_a in proportion to how they covary and never increases its uncertainty. The Kalman filter’s update step is this formula, with the predicted state as xa\mathbf{x}_a and the measurement as xb\mathbf{x}_b: then Σab=Pk∣k−1H⊤\Sigma_{ab} = P_{k \mid k-1} H^\top, Σbb\Sigma_{bb} is the innovation covariance SkS_k, ΣabΣbb−1\Sigma_{ab} \Sigma_{bb}^{-1} is the Kalman gain, and xb−μb\mathbf{x}_b - \boldsymbol{\mu}_b is the innovation.

Products of Gaussian densities. The product of two Gaussian densities in the same variable is, up to normalization, another Gaussian:

N(x; μ1,Σ1) N(x; μ2,Σ2)∝N(x; μ,Σ),Σ=(Σ1−1+Σ2−1)−1,μ=Σ(Σ1−1μ1+Σ2−1μ2).\mathcal{N}(\mathbf{x};\, \boldsymbol{\mu}_1, \Sigma_1)\, \mathcal{N}(\mathbf{x};\, \boldsymbol{\mu}_2, \Sigma_2) \propto \mathcal{N}(\mathbf{x};\, \boldsymbol{\mu}, \Sigma), \qquad \Sigma = (\Sigma_1^{-1} + \Sigma_2^{-1})^{-1}, \quad \boldsymbol{\mu} = \Sigma (\Sigma_1^{-1} \boldsymbol{\mu}_1 + \Sigma_2^{-1} \boldsymbol{\mu}_2).

This is how two independent estimates of the same quantity are fused: inverse covariances (information) add, and the fused mean is the information-weighted average. In one dimension, fusing a prediction of variance PP with a measurement of variance RR moves the estimate a fraction P/(P+R)P / (P + R) of the way toward the measurement, exactly the scalar Kalman gain.

Central limit theorem

The suitably normalized sum of many independent random variables with finite variance tends to a Gaussian, whatever their individual distributions. This is one reason Gaussian noise models work: an error made of many small, independent contributions is approximately Gaussian. It also explains why repeatedly applying a box filter to an image approximates Gaussian blur.

Connection to least squares and Bayesian estimation

The negative log of a Gaussian density is a quadratic form plus a constant. Maximizing a Gaussian likelihood is therefore minimizing a sum of squared, covariance-weighted errors, which is why least squares is the maximum-likelihood estimate under Gaussian noise. In Bayesian filtering, linear models with Gaussian noise keep the belief Gaussian at every step, so the recursion reduces to updating a mean and a covariance: the Kalman filter. Kalman’s 1960 paper notes that its optimality holds for Gaussian processes or, without that assumption, among linear estimators under quadratic loss.

Examples

Sampling, estimation, and Mahalanobis distance

The code samples a correlated 2D Gaussian, estimates its parameters, finds its principal axes, and compares Euclidean and Mahalanobis distances.

import numpy as np

rng = np.random.default_rng(0)
mu = np.array([2.0, 1.0])
Sigma = np.array([[4.0, 1.5],
                  [1.5, 1.0]])

# Sample: x = mu + L z, with L the Cholesky factor of Sigma and z ~ N(0, I).
L = np.linalg.cholesky(Sigma)
X = mu + rng.standard_normal((5000, 2)) @ L.T

mu_hat = X.mean(axis=0)
Sigma_hat = np.cov(X, rowvar=False)
print("estimated mean:", np.round(mu_hat, 2))
print("estimated covariance:\n", np.round(Sigma_hat, 2))

# Principal axes of the ellipses: eigenvectors and standard deviations.
evals, evecs = np.linalg.eigh(Sigma)           # ascending eigenvalues
minor, major = evecs[:, 0], evecs[:, 1] * np.sign(evecs[0, 1])
print("axis std devs:", np.round(np.sqrt(evals), 2))
print("major axis direction:", np.round(major, 2))

def mahalanobis(x, mu, Sigma):
    d = x - mu
    return np.sqrt(d @ np.linalg.solve(Sigma, d))

# Two points at the same Euclidean distance (2) from the mean.
for p in [mu + 2 * major, mu + 2 * minor]:
    print(f"point {np.round(p, 2)}: Euclidean 2.00, "
          f"Mahalanobis {mahalanobis(p, mu, Sigma):.2f}")

# Squared Mahalanobis distance is chi-squared with 2 degrees of freedom;
# its 95% quantile is 5.991.
D = X - mu
d2 = np.einsum("ij,ij->i", D @ np.linalg.inv(Sigma), D)
print(f"fraction inside the 95% ellipse: {np.mean(d2 < 5.991):.3f}")

Output:

estimated mean: [2.02 1.01]
estimated covariance:
 [[4.05 1.48]
 [1.48 0.97]]
axis std devs: [0.62 2.15]
major axis direction: [0.92 0.38]
point [3.85 1.77]: Euclidean 2.00, Mahalanobis 0.93
point [ 2.77 -0.85]: Euclidean 2.00, Mahalanobis 3.25
fraction inside the 95% ellipse: 0.948

The sample estimates are close to the true parameters. Both test points are two units from the mean, but the one along the major axis is less than one standard deviation away, while the one along the minor axis is more than three. The fraction of samples inside the 95% ellipse matches the chi-squared prediction.

Image noise

The simplest image noise model adds independent, zero-mean Gaussian noise of fixed variance to every pixel, a common assumption in denoising experiments. Real sensor noise is signal-dependent: Foi and colleagues model raw data as a Poissonian part for photon sensing, whose variance grows with intensity, plus a Gaussian part for the remaining stationary disturbances. Because Poisson noise at moderate and high counts is close to Gaussian, the combination is commonly approximated by a Gaussian whose variance is an affine function of the noise-free intensity.

Gaussian filters and mixtures

The Gaussian function also appears as a kernel rather than as a model of noise. Gaussian smoothing convolves an image with a sampled Gaussian, which suppresses noise before computing image gradients and before subsampling in image pyramids. The kernel’s width is a scale parameter; no claim about the image’s statistics is implied.

A Gaussian mixture model is a weighted sum of several Gaussians. It represents multimodal distributions, such as the colors of a background pixel that alternates between leaves and sky, and is typically fitted with the expectation–maximization algorithm.

Common Misconceptions

  • “Uncorrelated means independent.” True for components of a jointly Gaussian vector, false in general: two variables can each be Gaussian and uncorrelated yet dependent, because Gaussian marginals do not imply a jointly Gaussian distribution.
  • “The one-sigma ellipse contains 68% of the probability.” Only in one dimension. In two dimensions the ellipse dM=1d_M = 1 contains about 39%, and dM=2d_M = 2 about 86%; confidence regions must use chi-squared quantiles for the right dimension.
  • “Using a Gaussian filter assumes Gaussian noise.” The smoothing kernel and the noise model are separate choices.

Limitations

  • Light tails. The Gaussian assigns vanishing probability to large errors, so a model built on it treats an outlier, such as a false detection or a mismatched feature, as overwhelming evidence and lets it pull the estimate far off. Heavy-tailed alternatives (Laplace, Student’s t) and robust penalties reduce the influence of such points.
  • Unimodality. A single Gaussian cannot represent “the object is behind one of two occluders” or “the match is one of three repeated window corners”. Mixtures, multiple-hypothesis trackers, and particle filters, which represent a belief with weighted samples, handle such beliefs at higher cost.
  • Nonlinearity. A Gaussian pushed through a nonlinear function, such as perspective projection, is no longer Gaussian. Extended and unscented Kalman filters approximate the result with a new Gaussian, which works when the nonlinearity is mild relative to the uncertainty.
  • Covariance must be known or estimated. Errors in Σ\Sigma propagate into gating thresholds, filter gains, and confidence regions.

Where It Is Used

  • Estimation: least squares as maximum likelihood, with covariances for the fitted parameters.
  • Tracking: the Kalman filter and other Gaussian approximations to the Bayes filter; Mahalanobis gating in data association and costs for the Hungarian algorithm.
  • Image processing: noise models for denoising; Gaussian kernels for smoothing, derivatives, and scale spaces.
  • Machine learning: Gaussian mixture models for background subtraction and clustering, and Gaussian priors and likelihoods in probabilistic models.

Related

  • Least Squares

    The method of fitting a model to more measurements than unknowns by minimizing the sum of squared residuals, the workhorse of estimation in computer vision.

  • Kalman Filter

    A recursive algorithm that estimates the hidden state of a linear dynamic system from a sequence of noisy measurements, widely used to smooth and predict object positions in tracking.

  • Bayesian Filtering

    Recursive estimation of the probability distribution of a hidden, changing state from a sequence of noisy measurements, by alternating a motion-model prediction with a Bayes' rule update.

  • Data Association

    Deciding which measurements or detections belong to which tracked targets, and which are false alarms, missed detections, new targets, or targets that have disappeared.

  • Image Pyramid

    A multi-scale representation of an image built by repeatedly smoothing and subsampling it, so that each level holds a coarser copy at half the resolution of the one below.

References

  1. Gauss, C. F. (1809). Theoria motus corporum coelestium in sectionibus conicis solem ambientium. Hamburg: Perthes & Besser.
  2. Mahalanobis, P. C. (1936). On the Generalised Distance in Statistics. Proceedings of the National Institute of Sciences of India, 2(1), 49–55. Reprinted in Sankhyā A, 80(Suppl. 1), 1–7 (2018).
  3. Anderson, T. W. (2003). An Introduction to Multivariate Statistical Analysis (3rd ed.). Hoboken, NJ: Wiley.
  4. Kalman, R. E. (1960). A New Approach to Linear Filtering and Prediction Problems. Journal of Basic Engineering, 82(1), 35–45.
  5. Foi, A., Trimeche, M., Katkovnik, V. & Egiazarian, K. (2008). Practical Poissonian-Gaussian Noise Modeling and Fitting for Single-Image Raw-Data. IEEE Transactions on Image Processing, 17(10), 1737–1754.