Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Entropy, varentropy, and Rényi entropy

This note derives formulas for exponential-family components and joint normal variance–mean mixtures. Entropy is the mean of the surprisal, and varentropy is its variance. For a wider structural account, the upstream source cites Stankyavichyus (2026), Varentropy: Overview, Computational Routes, and Structural Decomposition, a preprint. The derivations below concern the stated model families.

Read the exponential-family core, Fisher geometry, and the joint/marginal distinction first. Monte Carlo examples stay in the upstream varentropy tutorial.

Definitions

Let XX have density p(x)p(x) with respect to a reference measure μ(dx)\mu(dx). The information content (surprisal) is I(X)=logp(X)\mathcal{I}(X) = -\log p(X). Its mean is the entropy and its variance is the varentropy:

H=E[logp(X)],VH=Var[logp(X)].H = \mathbb{E}[-\log p(X)], \qquad V_H = \operatorname{Var}[-\log p(X)].

Entropy measures the average surprisal; varentropy measures how much the surprisal fluctuates around that average. Unlike the kurtosis, the varentropy requires only I(X)L2\mathcal{I}(X) \in L^2, so it stays finite for many heavy-tailed laws whose fourth moment diverges.

The Rényi entropy of order α>0\alpha > 0, α1\alpha \neq 1 (Rényi, 1961, On Measures of Entropy and Information), is

Hα=11αlogp(x)αμ(dx),H_\alpha = \frac{1}{1-\alpha} \log \int p(x)^\alpha \, \mu(dx),

For Lebesgue measure HH is differential entropy; for a general reference measure it is entropy relative to that measure. Logs are natural. Fix the reference measure throughout. Where the density-power integral is finite in an open interval around 1 and differentiation under the integral is justified, HαHH_\alpha\to H as α1\alpha\to1. These regularity assumptions also justify the derivatives and local expansion below.

Density-Power Route

All three quantities are governed by the log density-power integral

R(α)=logp(x)αμ(dx),R(1)=0.R(\alpha) = \log \int p(x)^\alpha \, \mu(dx), \qquad R(1) = 0.

Writing I=logp(X)\mathcal{I} = -\log p(X), the function R(1s)R(1-s) is the cumulant-generating function of I\mathcal{I}, so differentiating at α=1\alpha = 1 gives

H=R(1),VH=R(1),H = -R'(1), \qquad V_H = R''(1),

while the Rényi entropy is Hα=R(α)/(1α)H_\alpha = R(\alpha)/(1-\alpha). The first-order expansion

Hα=H12VH(α1)+O ⁣((α1)2)H_\alpha = H - \tfrac{1}{2} V_H\,(\alpha - 1) + \mathcal{O}\!\left((\alpha-1)^2\right)

shows that the varentropy is (twice the negative of) the slope of the Rényi spectrum at α=1\alpha = 1.

Exponential Family

For an exponential family p(xθ)=h(x)exp{θt(x)ψ(θ)}p(x\mid\theta) = h(x)\exp\{\theta^\top t(x) - \psi(\theta)\} whose carrier is constant on its support, logh(x)b0\log h(x) \equiv b_0, the normalized density power stays in the family with natural parameter αθ\alpha\theta, provided this parameter belongs to the natural domain:

R(α)=(α1)b0+ψ(αθ)αψ(θ).R(\alpha) = (\alpha - 1)\,b_0 + \psi(\alpha\theta) - \alpha\,\psi(\theta).

Substituting into (4) and using η=ψ(θ)\eta = \nabla\psi(\theta) and the Fisher information I(θ)=2ψ(θ)I(\theta) = \nabla^2\psi(\theta) gives closed forms in terms of the log-partition triad:

H=ψ(θ)θηb0,VH=θI(θ)θ,Hα=(α1)b0+ψ(αθ)αψ(θ)1α.\begin{aligned} H &= \psi(\theta) - \theta^\top \eta - b_0, \\ V_H &= \theta^\top I(\theta)\,\theta, \\ H_\alpha &= \frac{(\alpha-1)\,b_0 + \psi(\alpha\theta) - \alpha\,\psi(\theta)}{1-\alpha}. \end{aligned}

The varentropy identity VH=θI(θ)θV_H = \theta^\top I(\theta)\,\theta holds because the centered information content I(X)H=θ{t(X)η}\mathcal{I}(X) - H = -\theta^\top\{t(X) - \eta\} lies exactly in the span of the score. This covers gamma, inverse gamma, GIG, and multivariate normal families when their carrier is chosen as h=1h=1 on their fixed supports. For a dd-dimensional normal law, VH=d/2V_H=d/2.

When logh\log h is not constant, the general variance formula is

VH=θI(θ)θ+2θCov(t(X),logh(X))+Var(logh(X)).V_H=\theta^\top I(\theta)\theta +2\theta^\top\operatorname{Cov}(t(X),\log h(X)) +\operatorname{Var}(\log h(X)).

For example, the two-parameter inverse-Gaussian representation has logh(y)=12log(2π)32logy\log h(y)=-\tfrac12\log(2\pi)-\tfrac32\log y. One can instead use its exact embedding GIG(1/2,λ/m2,λ)\operatorname{GIG}(-1/2,\lambda/m^2,\lambda) with h=1h=1 and statistic logy\log y. Along the density-power path, the GIG order then varies; holding it at 1/2-1/2 would omit part of the surprisal.

Varentropy of the GIG Distribution

Let YGIG(p,a,b)Y \sim \mathrm{GIG}(p, a, b) with density (1) and natural parameters θ=[p1,b/2,a/2]\theta = [p-1,\,-b/2,\,-a/2]. Raising the density to the power α\alpha keeps it proportional to a GIG density with parameters (1+α(p1),αa,αb)(1 + \alpha(p-1),\, \alpha a,\, \alpha b) — the escort-closure property. Using the Bessel integral 0yq1e(uy+v/y)/2dy=2(v/u)q/2Kq(uv)\int_0^\infty y^{q-1} e^{-(uy + v/y)/2}\,dy = 2\,(v/u)^{q/2} K_q(\sqrt{uv}) (NIST DLMF, 10.32.10, after a change of variable) gives, with z=abz = \sqrt{ab},

R(α)=(1α)log2+1α2log ⁣ba+logK1+α(p1)(αz)αlogKp(z).R(\alpha) = (1-\alpha)\log 2 + \frac{1-\alpha}{2}\log\!\frac{b}{a} + \log K_{1 + \alpha(p-1)}(\alpha z) - \alpha \log K_p(z).

Define F(p,z)=logKp(z)F(p, z) = \log K_p(z) and the first-order differential operator

L=(p1)p+zz.L = (p-1)\,\partial_p + z\,\partial_z.

Along the escort path α(1+α(p1),αz)\alpha \mapsto (1 + \alpha(p-1),\, \alpha z) the derivatives at α=1\alpha=1 are ddαF=LF\frac{d}{d\alpha}F=LF and d2dα2F=(L2L)F\frac{d^2}{d\alpha^2}F=(L^2-L)F. The subtraction accounts for the variable coefficients of LL. All remaining terms of (9) are linear in α\alpha, so from (4),

  VH{GIG(p,a,b)}=(L2L)logKp(z),z=ab.  \boxed{\; V_H\{\mathrm{GIG}(p,a,b)\} = (L^2 - L)\log K_p(z), \qquad z = \sqrt{ab}. \;}

Expanded, this is the pure second-order form

VH=(p1)2Fpp+2(p1)zFpz+z2Fzz,V_H = (p-1)^2 F_{pp} + 2(p-1)\,z\,F_{pz} + z^2 F_{zz},

which coincides with the Fisher quadratic form θI(θ)θ\theta^\top I(\theta)\,\theta of (7). The entropy follows from H=R(1)H = -R'(1):

H{GIG(p,a,b)}=log{2Kp(z)}+12log ⁣baLlogKp(z).H\{\mathrm{GIG}(p,a,b)\} = \log\{2 K_p(z)\} + \tfrac12\log\!\frac{b}{a} - L\log K_p(z).

For a,b>0a,b>0, the density decays exponentially in yy at infinity and in 1/y1/y near zero, so the required logarithmic and power statistics have finite second moments. Thus VHV_H is finite for every interior GIG law. The proper inverse-gamma boundary a=0a=0, p=k<0p=-k<0, b>0b>0 also has finite varentropy for every k>0k>0, even when its fourth moment fails (k4k\leq4). That boundary fact follows from the inverse-gamma density, not from substitution of z=0z=0 in an interior Bessel expression. It is not a uniform bound as the shape approaches zero.

Joint Varentropy of Normal Variance-Mean Mixtures

Consider the joint law of (2) for a positive mixing variable,

XY=yNd(μ+γy, Σy),Ygϑ.X \mid Y = y \sim \mathcal{N}_d(\mu + \gamma y,\ \Sigma y), \qquad Y \sim g_\vartheta.

Conditionally on YY, the quadratic form Q=(XμγY)Σ1(XμγY)/YQ = (X - \mu - \gamma Y)^\top \Sigma^{-1}(X - \mu - \gamma Y)/Y is χd2\chi^2_d and independent of YY, so

logp(XY)=CΣ+d2logY+12Q,CΣ=d2log(2π)+12logΣ.-\log p(X \mid Y) = C_\Sigma + \tfrac{d}{2}\log Y + \tfrac12 Q, \qquad C_\Sigma = \tfrac{d}{2}\log(2\pi) + \tfrac12\log|\Sigma|.

Writing IY=loggϑ(Y)\mathcal{I}_Y = -\log g_\vartheta(Y) for the mixing-law surprisal, the joint information content is IX,Y=CΣ+IY+d2logY+12Q\mathcal{I}_{X,Y} = C_\Sigma + \mathcal{I}_Y + \tfrac{d}{2}\log Y + \tfrac12 Q. Since Var(12Q)=d/2\operatorname{Var}(\tfrac12 Q) = d/2 and QYQ \perp Y,

  VH(X,Y)=d2+Var ⁣[IY+d2logY].  \boxed{\; V_H(X, Y) = \frac{d}{2} + \operatorname{Var}\!\left[\mathcal{I}_Y + \frac{d}{2}\log Y\right]. \;}

The representation is X=μ+γY+YLZX=\mu+\gamma Y+\sqrt Y LZ, with LL=ΣLL^\top=\Sigma, ZNd(0,Id)Z\sim\mathcal N_d(0,I_d) independent of YY, and Σ\Sigma positive definite. Thus Q=ZZQ=Z^\top Z is independent of YY. The notation maps to W=YW=Y, β=γ\beta=\gamma, σ2=Σ\sigma^2=\Sigma in the scalar conditioning note.

The Gaussian layer contributes exactly d/2d/2 to joint varentropy. For a fixed mixing law, μ\mu, γ\gamma, and Σ\Sigma all drop out of this variance. The joint entropy is

H(X,Y)=H(Y)+d2log(2πe)+12logΣ+d2E[logY],H(X,Y)=H(Y)+\frac d2\log(2\pi e)+\frac12\log|\Sigma| +\frac d2\mathbb E[\log Y],

when the terms are finite. In particular, μ\mu and γ\gamma also drop out of joint entropy; only Σ\Sigma contributes among the normal parameters. More generally, Gaussian integration gives the joint density-power route

RX,Y(α)=(1α)CΣd2logα+log0gϑ(y)αy(1α)d/2dy.\begin{aligned} R_{X,Y}(\alpha) &=(1-\alpha)C_\Sigma-\frac d2\log\alpha\\ &\quad+\log\int_0^\infty g_\vartheta(y)^\alpha y^{(1-\alpha)d/2}\,dy. \end{aligned}

The integral must be finite at the requested Rényi order. These are joint information quantities. They do not give the entropy or varentropy of the GH marginal XX by deleting the hidden variable; the marginal’s surprisal is logp(x,y)dy-\log\int p(x,y)\,dy.

For GIG mixing, IY=ψGIG(θ)θt(Y)\mathcal{I}_Y = \psi_{\mathrm{GIG}}(\theta) - \theta^\top t(Y) with t(Y)=[logY,Y1,Y]t(Y) = [\log Y,\, Y^{-1},\, Y], so adding d2logY\tfrac{d}{2}\log Y shifts only the coefficient of logY\log Y:

IY+d2logY=ψGIG(θ)(θδd)t(Y),δd=(d2,0,0).\mathcal{I}_Y + \tfrac{d}{2}\log Y = \psi_{\mathrm{GIG}}(\theta) - (\theta - \delta_d)^\top t(Y), \qquad \delta_d = (\tfrac{d}{2},\, 0,\, 0)^\top.

Its variance is the shifted Fisher quadratic form, and translating to classical coordinates gives the operator LdL_d of (10) with the order shifted by d/2-d/2:

  VH(X,Y)=d2+(Ld2Ld)logKp(z),Ld=(p1d2)p+zz.  \boxed{\; V_H(X, Y) = \frac{d}{2} + (L_d^2 - L_d)\log K_p(z), \qquad L_d = \left(p - 1 - \tfrac{d}{2}\right)\partial_p + z\,\partial_z. \;}

For d=0d = 0 this reduces to the GIG varentropy (11). The operator shift p1p1d/2p - 1 \mapsto p - 1 - d/2 is exactly the effect of the conditional Gaussian volume factor Yd/2Y^{-d/2}, and mirrors the natural parameter θ1=p1d/2\theta_1 = p - 1 - d/2 of the joint exponential family (11). The variance-gamma and normal-inverse-gamma joints arise at the proper boundaries b=0b=0, p>0p>0 and a=0a=0, p<0p<0, respectively; they require their own log-partition formulas. Normal-inverse-Gaussian mixing has p=1/2p=-1/2 and can be evaluated through the GH embedding so that the order varies faithfully along the density-power path.

Coordinate and moment limits

Under an invertible affine change U=AX+cU=AX+c of a continuous vector, H(U)=H(X)+logdetAH(U)=H(X)+\log|\det A| and VH(U)=VH(X)V_H(U)=V_H(X), because the Jacobian contributes a constant. A nonlinear transformation generally adds a random log-Jacobian and can change varentropy. All formulas above use the stated Lebesgue coordinates, rather than a coordinate-free information measure.

Finiteness of surprisal variance is different from finiteness of XX’s fourth moment. It can make varentropy useful for heavy-tailed families, but does not by itself order all distributions by tail weight or guarantee accurate finite-sample estimation. The GH family tour identifies the mixing boundaries; online EM and shrinkage concern parameter estimation rather than an estimator or convergence guarantee for these information quantities.

Source and adaptation

Adapted from xshi19/normix, docs/theory/varentropy.md, at revision 763bb3608920661a012cf089888d349fbf680aad (2026-09-13 import). Copyright (c) 2020 xshi19. Licensed under MIT. The pinned source records the original version. Notation, mathematical qualifications, and links were adapted for this site. Package interfaces, fitter recipes, and executable cells are omitted; no upstream benchmark or formal-proof verification is claimed.

MIT permission notice

MIT License

Copyright (c) 2020 xshi19

Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the “Software”), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED “AS IS”, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.