← New search

Other meanings of rate

Probability theory

Rate function

A rate function quantifies the exponential cost of an atypical macroscopic outcome in a large-deviation principle. If random variables become increasingly concentrated near typical values as a scale parameter grows, the probability of observing a value x away from that set often behaves like exp(−I(x) an), where an tends to infinity and I is the rate function.1 Unlike a probability density, I is not generally normalized and need not itself be a probability. Its zeros identify typical behavior, while larger values indicate exponentially rarer deviations.

I(x) ≥ 0
nonnegative cost
rate function
aₙ → ∞
scaling parameter
large-deviation scale
I(x)=0
typical-value condition
zero set
1

Definition and interpretation

A rate function assigns an exponential rarity cost to outcomes, rather than describing their ordinary probability directly. A sequence of random variables Xn satisfies a large-deviation principle with speed an and rate function I when closed sets provide an exponential upper bound and open sets provide a corresponding lower bound: roughly, P(Xn ∈ A) behaves on the logarithmic scale like exp[−an infx∈AI(x)].1 The function is typically lower semicontinuous, nonnegative, and may take the value +∞ for impossible or super-exponentially unlikely outcomes.

The zeros of I form the set of values that remain typical at the chosen scale. A unique zero indicates concentration around one macroscopic value, but multiple zeros can describe phase coexistence or competing stable states. The rate function records exponential order only: two events can have the same rate while differing substantially in polynomial prefactors, which large-deviation theory deliberately suppresses.

2

Construction and standard examples

For independent sample means, the rate function is often obtained from a cumulant-generating function by convex duality. Cramér's theorem gives I(x)=supt{tx−λ(t)}, where λ(t)=log E[etY] for an underlying random variable Y, whenever the relevant regularity conditions hold.2 This Legendre–Fenchel transform makes I convex for independent identically distributed sums and connects fluctuations to exponential tilting.

For Bernoulli trials with success probability p, the sample proportion has rate function I(x)=x log(x/p)+(1−x)log[(1−x)/(1−p)] on 0≤x≤1, with the usual continuous interpretation at the endpoints. This is the binary relative entropy, showing that information-theoretic divergence can serve as a probabilistic rarity cost. In statistical mechanics, analogous functions describe energy, magnetization, or empirical-measure fluctuations and may become nonconvex when the underlying model has interacting degrees of freedom.

3

Contraction, transforms, and applications

A rate function can be transferred through a continuous observable by the contraction principle. If Yn=f(Xn) and Xn has rate function I, then the induced cost is commonly J(y)=inf{I(x): f(x)=y}.1 This principle lets one derive fluctuation costs for sums, ratios, norms, currents, and other quantities without restarting the analysis from probability estimates.

Varadhan's lemma links rate functions with asymptotics of exponential expectations: under suitable hypotheses, a scaled logarithm of E[exp(anf(Xn))] is governed by supx{f(x)−I(x)}. The result underlies Laplace methods, importance sampling, risk-sensitive control, queueing asymptotics, random matrices, and nonequilibrium statistical mechanics. In physics, the Legendre transform of a rate function can correspond to a scaled cumulant-generating function, although differentiability and ensemble equivalence require separate checks.3

4

Lesser-known aspects

Rate functions need not be finite, smooth, or strictly convex. A hard constraint may produce I(x)=+∞ outside an admissible region; boundary points can create nondifferentiable corners; and several minimizers can coexist. A rate function may also fail to be a good rate function if its sublevel sets are not compact, a distinction that matters when passing from finite-dimensional variables to paths or measures.

The Gärtner–Ellis theorem provides a powerful route from limiting scaled cumulant-generating functions, but it does not automatically guarantee the full lower bound at points where the limiting transform is not differentiable.4 In dynamical systems and stochastic processes, the relevant object may be a path-space rate function rather than a function of one number. Large deviations can then describe rare trajectories, escape paths, empirical occupations, or histories of current fluctuations. These cases expose a useful limitation: exponential asymptotics identify dominant mechanisms, but they do not by themselves determine prefactors, fluctuation determinants, or the most probable path when minimizers are nonunique.

Glossary

Large-deviation principle
A pair of exponential upper and lower bounds describing probabilities on a scale that grows with the system size or observation time.
Speed
The diverging scale aₙ multiplying the rate function in logarithmic probability asymptotics.
Good rate function
A lower-semicontinuous rate function whose sublevel sets are compact.
Contraction principle
A rule for obtaining the rate function of a transformed random variable by minimizing the original cost over inverse images.
Legendre–Fenchel transform
The convex-duality operation that often converts a scaled cumulant-generating function into a rate function.

The notation and normalization of a rate function depend on the selected speed; multiplying the speed by a constant correspondingly rescales the rate function.