Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Fisher geometry and the meaning of L2

Euclidean distance between parameter vectors measures changes in the chosen coordinates. Fisher geometry instead measures first-order changes in the log density, averaged under the model. Its coordinate matrix changes when the parameters change, while the length assigned to the same statistical motion stays fixed.

A regular model and its scores

Let pθp_\theta be densities with respect to a fixed measure ν\nu, with θ\theta in an open subset of Rd\mathbb R^d. Assume a common positive support, twice continuously differentiable densities, and local integrable bounds that justify differentiating the normalizing integral twice. Assume the scores are square integrable. For the KL expansion below, also require that the expected log density has a second-order Taylor expansion with an integrable remainder.

Write θ(x)=logpθ(x)\ell_\theta(x)=\log p_\theta(x) and define the column score vector

sθ(x)=θθ(x).s_\theta(x)=\nabla_\theta\ell_\theta(x).

Along a curve with velocity vv, the log density changes at rate vTsθ(x)v^\mathsf{T}s_\theta(x). Differentiating pθdν=1\int p_\theta\,d\nu=1 gives

Eθ[sθ]=θpθdν=0.\mathbb E_\theta[s_\theta] =\int\nabla_\theta p_\theta\,d\nu=0.

The Fisher information matrix per observation is

Iθ=Eθ[sθsθT],gθ(v,w)=vTIθw.I_\theta=\mathbb E_\theta[s_\theta s_\theta^\mathsf{T}],\qquad g_\theta(v,w)=v^\mathsf{T}I_\theta w.

Thus gθ(v,v)g_\theta(v,v) is the mean square directional change in log density. The matrix is always positive semidefinite. It gives a Riemannian metric on a regular region where it is positive definite and smooth. A null direction means a zero first-order change in the density almost everywhere; global identifiability alone does not exclude such singular parameterizations.

Local KL divergence gives the same quadratic form

Differentiating the normalizer a second time, using ijp=p(ij+ij)\partial_i\partial_jp=p(\partial_i\partial_j\ell+ \partial_i\ell\,\partial_j\ell), gives

Eθ[θ2θ]=Iθ.\mathbb E_\theta[\nabla_\theta^2\ell_\theta]=-I_\theta.

Now hold the first distribution fixed in

DKL(pθpθ+δ)=Eθ[θ(X)θ+δ(X)].D_{\mathrm{KL}}(p_\theta\|p_{\theta+\delta}) =\mathbb E_\theta[\ell_\theta(X)-\ell_{\theta+\delta}(X)].

Taylor expansion of the second log density has a linear term with expectation zero and a quadratic term involving Iθ-I_\theta. Therefore

DKL(pθpθ+δ)=12δTIθδ+o(δ22).D_{\mathrm{KL}}(p_\theta\|p_{\theta+\delta}) =\frac12\delta^\mathsf{T}I_\theta\delta +o(\|\delta\|_2^2).

This is a local expansion as δ0\delta\to0 at a fixed interior parameter. Finite KL is generally asymmetric and is not squared Riemannian distance. For nn independent observations, scores add and their cross-covariances vanish, so the information matrix is nIθnI_\theta.

Reparameterization changes the matrix, not the length

Use a smooth invertible chart θ=θ(ϕ)\theta=\theta(\phi) and let J=θ/ϕJ=\partial\theta/\partial\phi. By the chain rule,

sϕ=JTsθ,Iϕ=JTIθJ.s_\phi=J^\mathsf{T}s_\theta,\qquad I_\phi=J^\mathsf{T}I_\theta J.

Since dθ=Jdϕd\theta=J\,d\phi, the two quadratic expressions agree:

dθTIθdθ=dϕTIϕdϕ.d\theta^\mathsf{T}I_\theta d\theta =d\phi^\mathsf{T}I_\phi d\phi.

This is exactly the metric transformation law. A Euclidean metric can also be transformed correctly. The problem arises if one declares the identity matrix to be the metric anew in every nonlinear chart: that changes the geometry.

The normal running example

For XN(μ,σ2)X\sim N(\mu,\sigma^2) with μR\mu\in\mathbb R and σ>0\sigma>0, use θ=(μ,σ)T\theta=(\mu,\sigma)^\mathsf{T}. The introductory calculation gives the Fisher matrix I(θ)=diag(σ2,2σ2)I(\theta)=\operatorname{diag}(\sigma^{-2},2\sigma^{-2}). Its transformation to mean and variance coordinates is worked out there as well.

We can check the local KL expansion directly for this family. Let θ~=(μ~,σ~)T\widetilde\theta=(\widetilde\mu,\widetilde\sigma)^\mathsf{T}, with σ~>0\widetilde\sigma>0. Taking the expected log density ratio and using Eθ[(Xμ~)2]=σ2+(μμ~)2\mathbb E_\theta[(X-\widetilde\mu)^2]=\sigma^2+(\mu-\widetilde\mu)^2 gives the exact formula

DKL(pθpθ~)=logσ~σ12+σ2+(μμ~)22σ~2.\begin{aligned} D_{\mathrm{KL}}(p_\theta\|p_{\widetilde\theta}) &=\log\frac{\widetilde\sigma}{\sigma}-\frac12\\ &\quad+\frac{\sigma^2+(\mu-\widetilde\mu)^2}{2\widetilde\sigma^2}. \end{aligned}

For a fixed velocity v=(vμ,vσ)Tv=(v^\mu,v^\sigma)^\mathsf{T}, set θ~=θ+εv\widetilde\theta=\theta+\varepsilon v. Expanding the logarithm and the reciprocal square at ε=0\varepsilon=0 cancels the linear terms and yields

DKL(pθpθ+εv)=ε22σ2((vμ)2+2(vσ)2)+o(ε2)=ε22gθ(v,v)+o(ε2).\begin{gathered} D_{\mathrm{KL}}(p_\theta\|p_{\theta+\varepsilon v})\\ =\frac{\varepsilon^2}{2\sigma^2} \bigl((v^\mu)^2+2(v^\sigma)^2\bigr)+o(\varepsilon^2)\\ =\frac{\varepsilon^2}{2}g_\theta(v,v)+o(\varepsilon^2). \end{gathered}

Thus the squared Fisher speed determines the leading KL change along either parameter direction, or any combination of them. The expansion is local at a fixed σ>0\sigma>0, with σ+εvσ>0\sigma+\varepsilon v^\sigma>0. If σ\sigma is known and only μ\mu varies, the restricted metric is ds2=dμ2/σ2ds^2=d\mu^2/\sigma^2, so a scaled Euclidean metric is appropriate for that subfamily.

Bernoulli probability and log odds

For XBernoulli(q)X\sim\operatorname{Bernoulli}(q) with 0<q<10<q<1, differentiation gives

sq(x)=xq1x1q=xqq(1q).s_q(x)=\frac{x}{q}-\frac{1-x}{1-q} =\frac{x-q}{q(1-q)}.

Because Var(X)=q(1q)\operatorname{Var}(X)=q(1-q),

Iq=1q(1q),ds2=dq2q(1q).I_q=\frac1{q(1-q)},\qquad ds^2=\frac{dq^2}{q(1-q)}.

Equal small changes in probability have different local statistical sizes: the coefficient is 4 at q=1/2q=1/2 and 10000/9910000/99 at q=1/100q=1/100. This statement uses the infinitesimal metric, not a finite-step equality for KL.

Now take the natural parameter θ=log(q/(1q))\theta=\log(q/(1-q)). Since dq/dθ=q(1q)dq/d\theta=q(1-q), the same metric is

Iθ=q(1q),ds2=q(1q)dθ2=dq2q(1q).I_\theta=q(1-q),\qquad ds^2=q(1-q)d\theta^2=\frac{dq^2}{q(1-q)}.

This agrees with Iθ=ψ(θ)I_\theta=\psi''(\theta) from the exponential-family calculation.

Three different uses of L2

The Euclidean norm on a finite parameter vector is more precisely an 2\ell^2 norm. It should be distinguished from two function-space constructions.

SpaceSquared sizeWhat determines it
Parameter coordinatesivi2\sum_i v_i^2A chosen Euclidean inner product in that chart
Scores in L2(Pθ)L^2(P_\theta)Eθ[(vTsθ)2]\mathbb E_\theta[(v^\mathsf{T}s_\theta)^2]The model law at the point; this is Fisher information
Density differences in L2(ν)L^2(\nu)(pq)2dν\int(p-q)^2\,d\nuA reference measure and densities relative to it; finiteness is an extra requirement

Raw density L2L^2 is invariant under parameter relabeling because the densities do not change. It generally depends on the measurement coordinates. For Lebesgue densities and Y=aXY=aX with a>0a>0, the transformed densities satisfy pY(y)=pX(y/a)/ap_Y(y)=p_X(y/a)/a, so

(pYqY)2dy=1a(pXqX)2dx.\int(p_Y-q_Y)^2\,dy=\frac1a\int(p_X-q_X)^2\,dx.

A fixed invertible transformation of the observation contributes a parameter-independent Jacobian to the log density, so it leaves scores and Fisher information unchanged. This explains a distinction between Fisher geometry and raw density L2L^2 that parameter invariance alone cannot explain.

Square-root densities recover Fisher geometry

There is a useful L2(ν)L^2(\nu) representation. Map a density to uθ=2pθu_\theta=2\sqrt{p_\theta}, which has norm 2. Assume this map is differentiable in L2(ν)L^2(\nu), with derivative along vv given by

vuθ=pθvTsθ.\partial_vu_\theta=\sqrt{p_\theta}\,v^\mathsf{T}s_\theta.

Then its squared L2(ν)L^2(\nu) norm is precisely vTIθvv^\mathsf{T}I_\theta v. This is an isometric realization of the tangent metric by square-root densities. It is not the raw density-difference construction above.

With the convention

H2(p,q)=1pqdν=12pqL2(ν)2,H^2(p,q)=1-\int\sqrt{pq}\,d\nu =\frac12\|\sqrt p-\sqrt q\|_{L^2(\nu)}^2,

the same differentiability gives

H2(pθ,pθ+δ)=18δTIθδ+o(δ22).H^2(p_\theta,p_{\theta+\delta}) =\frac18\delta^\mathsf{T}I_\theta\delta+o(\|\delta\|_2^2).

The factors 1/21/2 for KL and 1/81/8 for this squared Hellinger convention describe the same local Fisher metric. Finite Hellinger distance measures a chord between square-root densities; Fisher–Rao distance minimizes path length within the specified statistical model.

For background on the metric and these representations, see §§3.9–3.12 of Frank Nielsen’s An elementary introduction to information geometry. The conditional-expectation note explains projection in the other Hilbert space used here, L2(Pθ)L^2(P_\theta).

Continue with dual coordinates and KL projections, or return to the Information Geometry hub.