Motion & Tracking
Bayesian Filtering
Recursive estimation of the probability distribution of a hidden, changing state from a sequence of noisy measurements, by alternating a motion-model prediction with a Bayes' rule update.
intermediate
Bayesian filtering is the framework for estimating a hidden state that changes over time, such as the position and velocity of a tracked object or the pose of a moving camera, from a stream of noisy measurements. It maintains a probability distribution over the state and updates it recursively: predict how the state has evolved, then correct the prediction with the newest measurement using Bayes’ rule. The Kalman filter, particle filters, and grid-based filters are all ways of carrying out this one recursion under different assumptions, which is why the framework underlies most probabilistic object tracking and visual state estimation.
Definition
A Bayesian filter computes, at every time step , the posterior distribution of the current state given all measurements so far, . It does so recursively: the posterior at step is computed from the posterior at step and the new measurement alone, without revisiting earlier data. This posterior is often called the belief.
Point estimates, such as the posterior mean or mode, and uncertainty measures, such as its covariance, are derived from it.
Intuition
Picture a navigator without satellite positioning. Between landmarks, dead reckoning moves the estimated position forward while the uncertainty grows; a lighthouse sighting then narrows it to the positions consistent with both. A Bayesian filter does this with probability distributions. The predict step pushes the belief through the motion model, which moves it and spreads it out. The update step multiplies the belief by how well each possible state explains the new measurement, then renormalizes, which sharpens it. If the measurement is ambiguous, for example because several places look alike, the belief can keep several separate peaks until later measurements rule all but one out.
Formal Definition
State-space model
The filter assumes a probabilistic state-space model with three parts:
- an initial distribution ;
- a transition model , describing how the state evolves;
- a measurement model, or likelihood, , describing how measurements are generated from the state.
Two conditional independence assumptions make recursion possible:
- First-order Markov dynamics. Given the previous state, the current state is independent of all earlier states and measurements: .
- Conditionally independent measurements. Given the current state, a measurement is independent of all other states and measurements: .
A common way to specify these densities is with functions and noise, and , with independent noise sequences and . The state must be chosen so that the Markov assumption holds: a position-only state is not Markov for an object that moves with momentum, while position plus velocity can be.
Predict
Given the previous posterior , the prediction for step follows from the Chapman–Kolmogorov equation, which marginalizes over the previous state:
The Markov assumption is what allows to be replaced by the transition model.
Update
When arrives, Bayes’ rule combines the prediction, which acts as the prior, with the likelihood:
with the normalizing constant
Conditional independence of the measurements is what allows to be replaced by the measurement model. The normalizer is the predictive probability of the measurement; trackers use it to score how plausible a measurement is under a track. For a discrete state, the integrals become sums. If no measurement arrives at step , the update is skipped and the prediction is the posterior.
Properties
Why the general recursion is intractable
The two equations are exact but, in general, only conceptual. For nonlinear dynamics or measurements, or non-Gaussian noise, the integrals have no closed form and the posterior has no fixed parametric shape: it can become skewed or multimodal. Representing it on a grid works in low dimensions, but the number of cells grows exponentially with the state dimension. Every practical Bayesian filter is therefore either an exact solution for a restricted model or an approximation (Arulampalam et al., 2002).
The family of filters
- Kalman filter: the exact linear-Gaussian case. If the transition and measurement models are linear with additive Gaussian noise, and the initial distribution is Gaussian, every prediction and posterior is Gaussian. The recursion then reduces to updating a mean and a covariance in closed form. Ho and Lee (1964) formulated estimation in these Bayesian terms and showed that, under linear-Gaussian assumptions, the recursion yields the Kalman filter.
- Extended and unscented Kalman filters. For nonlinear models, these filters keep a Gaussian belief but approximate how it propagates. The extended Kalman filter linearizes the models around the current estimate; the unscented Kalman filter passes a small, deterministically chosen set of points through the nonlinear functions and refits a Gaussian. Both can perform poorly when the true posterior is far from Gaussian, for example bimodal.
- Grid-based filters. If the state takes finitely many values, the recursion is computed exactly as sums over a table of probabilities. A continuous state can be discretized into cells, which approximates the posterior with a histogram of any shape, at a cost that grows quickly with dimension.
- Particle filters. Sequential Monte Carlo methods represent the posterior by a set of weighted samples: each sample is propagated through the transition model and reweighted by the likelihood, and samples are resampled so that computation concentrates where the probability is. The sampling importance resampling, or bootstrap, filter of Gordon, Salmond and Smith (1993) is the basic form. Particle filters handle arbitrary models and multimodal beliefs, but the number of samples needed can grow exponentially with the size of the problem, which limits them in high dimensions (Snyder et al., 2008).
Filtering, prediction, and smoothing
Filtering estimates the current state from measurements up to now, . Prediction estimates a future state, with , by applying the predict step repeatedly without updates. Smoothing estimates a past state from measurements that include later ones, with , typically by a backward pass after the forward filter. Smoothed estimates are usually more accurate because they use more data, but they are available only after the fact, which suits offline video analysis rather than online tracking (Särkkä, 2013).
Examples
Tracking a target’s position
With an object’s image position and velocity as the state, a constant-velocity transition model, and a detector’s position with Gaussian noise as the measurement, the problem is linear-Gaussian and the Kalman filter is the exact Bayesian filter. When the likelihood comes from appearance, such as a color-histogram or template score over the image, clutter and look-alike objects make it multimodal, and particle filters are used instead; CONDENSATION (Isard & Blake, 1998) is a classic instance in visual tracking.
A grid-based filter in one dimension
The example below tracks an object moving around a circular track of 60 cells. Each step it usually advances one cell, sometimes zero or two. The camera cannot measure position; it only reports whether the object is at one of seven identical markers, with a 5% false-detection rate and a 10% miss rate. A Kalman filter cannot represent this problem, because the belief is multimodal: seeing a marker means being near any one of seven places.
import numpy as np
n = 60 # cells on a circular 1D track
markers = np.zeros(n, dtype=bool)
markers[[4, 11, 19, 22, 34, 41, 52]] = True # identical-looking markers
p_move = np.array([0.05, 0.9, 0.05]) # moves 0, 1 or 2 cells per step
p_hit, p_false = 0.9, 0.05 # P(marker seen | at marker / not)
def predict(bel):
# Chapman-Kolmogorov sum: convolve the belief with the motion model.
return sum(p * np.roll(bel, s) for s, p in enumerate(p_move))
def update(bel, seen):
like = np.where(markers, p_hit, p_false) # p(z_k | x_k) for "seen"
if not seen:
like = 1.0 - like
post = like * bel
return post / post.sum() # Bayes' rule, normalized
def peaks(bel, thresh=0.05):
# Count separated regions of the belief holding at least `thresh` mass.
mass = bel + np.roll(bel, 1) + np.roll(bel, -1)
is_peak = (bel >= np.roll(bel, 1)) & (bel > np.roll(bel, -1))
return int(np.sum(is_peak & (mass >= thresh)))
rng = np.random.default_rng(0)
x = 0 # true cell, unknown to the filter
bel = np.full(n, 1.0 / n) # uniform prior: no idea where
for k in range(1, 61):
x = (x + rng.choice(3, p=p_move)) % n
seen = rng.random() < (p_hit if markers[x] else p_false)
bel = update(predict(bel), seen)
if k in (1, 10, 20, 25, 40, 60):
near = bel[[(x - 1) % n, x, (x + 1) % n]].sum()
print(f"k={k:2d} true={x:2d} MAP={np.argmax(bel):2d} "
f"peaks={peaks(bel)} P(within 1 cell of truth)={near:.2f}")
Output:
k= 1 true= 1 MAP= 0 peaks=0 P(within 1 cell of truth)=0.06
k=10 true= 9 MAP=14 peaks=6 P(within 1 cell of truth)=0.05
k=20 true=19 MAP=19 peaks=4 P(within 1 cell of truth)=0.33
k=25 true=24 MAP=24 peaks=2 P(within 1 cell of truth)=0.79
k=40 true=39 MAP=39 peaks=2 P(within 1 cell of truth)=0.90
k=60 true=58 MAP=59 peaks=1 P(within 1 cell of truth)=0.70
After the first sightings the belief has several peaks, one per hypothesis consistent with the observations, and the most probable cell is wrong. As the object passes more markers, only the hypothesis matching the spacing of sightings survives, and the belief collapses onto the true position. Between markers the motion noise spreads it again, which is why the probability near the truth drops at the end.
Visual odometry and SLAM
Filter-based visual odometry and SLAM take the camera pose and velocity, and sometimes map landmark positions, as the state; the transition model comes from a motion model or an inertial sensor, and the measurement model projects landmarks into the image. Because projection and rotation are nonlinear, these systems use the extended Kalman filter or related approximations, as in MonoSLAM (Davison et al., 2007) and the multi-state constraint Kalman filter for visual-inertial navigation (Mourikis & Roumeliotis, 2007).
Uncertain measurement origin
The update step assumes the filter knows which measurement belongs to its state. With several targets or clutter, that is not given, and the origin of each measurement becomes another unknown. Data association methods either commit to one assignment before the update, or, as in probabilistic data association, average the update over the possible origins weighted by their probabilities, which is itself a Bayesian computation over hypotheses.
Common Misconceptions
- “A Bayesian filter is a specific algorithm.” It is a recursion; the Kalman, particle, and grid filters are implementations of it under different belief representations.
- “The Kalman filter is an approximation of the Bayes filter.” For linear-Gaussian models it is the exact Bayes filter. It becomes an approximation only when applied to models that violate those assumptions, which is the setting of the extended and unscented variants.
- “The filter returns the state.” It returns a distribution. Reporting only its mean can be badly misleading when the belief is multimodal: the mean of two peaks can lie in a region of near-zero probability.
- “A narrow belief is a correct belief.” The posterior is only as good as the model. With a wrong motion model or underestimated noise, the filter becomes overconfident, effectively ignores measurements that disagree with it, and is confidently wrong.
Where It Is Used
- Object tracking. Motion models in object tracking and multi-object tracking, usually a Kalman filter, as in SORT; particle filters for cluttered, multimodal appearance likelihoods.
- Visual odometry, visual-inertial navigation, and SLAM. Recursive estimation of camera pose and map.
- Robot localization. Grid-based and particle-filter localization in a known map.
- Sensor fusion. Cameras combined with radar, LiDAR, or inertial sensors, each through its own measurement model.
Related
- Kalman Filter
A recursive algorithm that estimates the hidden state of a linear dynamic system from a sequence of noisy measurements, widely used to smooth and predict object positions in tracking.
- Object Tracking
Estimating the position, extent, or state of one or more objects in every frame of a video, keeping each object's identity over time.
- Gaussian Distribution
The bell-shaped probability distribution defined by a mean and a covariance, the default model for noise and uncertainty in estimation and tracking.
- Data Association
Deciding which measurements or detections belong to which tracked targets, and which are false alarms, missed detections, new targets, or targets that have disappeared.
References
- Ho, Y. C. & Lee, R. C. K. (1964). A Bayesian Approach to Problems in Stochastic Estimation and Control. IEEE Transactions on Automatic Control, 9(4), 333–339.
- Gordon, N. J., Salmond, D. J. & Smith, A. F. M. (1993). Novel Approach to Nonlinear/Non-Gaussian Bayesian State Estimation. IEE Proceedings F (Radar and Signal Processing), 140(2), 107–113.
- Isard, M. & Blake, A. (1998). CONDENSATION—Conditional Density Propagation for Visual Tracking. International Journal of Computer Vision, 29(1), 5–28.
- Arulampalam, M. S., Maskell, S., Gordon, N. & Clapp, T. (2002). A Tutorial on Particle Filters for Online Nonlinear/Non-Gaussian Bayesian Tracking. IEEE Transactions on Signal Processing, 50(2), 174–188.
- Särkkä, S. (2013). Bayesian Filtering and Smoothing. Cambridge University Press.
- Snyder, C., Bengtsson, T., Bickel, P. & Anderson, J. (2008). Obstacles to High-Dimensional Particle Filtering. Monthly Weather Review, 136(12), 4629–4640.
- Davison, A. J., Reid, I. D., Molton, N. D. & Stasse, O. (2007). MonoSLAM: Real-Time Single Camera SLAM. IEEE Transactions on Pattern Analysis and Machine Intelligence, 29(6), 1052–1067.
- Mourikis, A. I. & Roumeliotis, S. I. (2007). A Multi-State Constraint Kalman Filter for Vision-aided Inertial Navigation. IEEE International Conference on Robotics and Automation (ICRA), 3565–3572.