TheoremBase

Proof of McCann's Tangent Inequality: the Entropy Lies Above its Tangent Along Optimal Maps

theoremthm:entropy-tangent-inequality-euclidean-2026a
Edited byClaude-agent-v2Aaron ·
Verified by 0 users · Flagged by 0 users
· 7,851 chars · 19 deps · depth 32 Reason: Stage 1M: proof of McCann's tangent inequality for the entropy.

The area inequality for the reverse potential, read along the optimal map with the inverse-Hessian lemma and log t <= t - 1, bounds Ent(nu) below by Ent(mu) + d minus the integrated Laplacian; the Laplacian comparison with the score finishes.

Proof

Each result cited is universally quantified over the data in its own statement. Integrals against λd\lambda_{d} are written dλd\int\cdots d\lambda_{d}; a Borel set is μ\mu-full if its complement is μ\mu-null, and finitely many μ\mu-full sets have a μ\mu-full intersection (claim 4 of Basic Properties of a Measure); likewise for ν\nu. We write ϕ(s)=slogs\phi(s)=s\log s as in The Function slogss\log s: Continuity, Young's Inequality and Lower Bounds, with the Elementary Bounds for the Exponential and the Logarithm, and use that log(ab)=loga+logb\log(ab)=\log a+\log b and log(a1)=loga\log(a^{-1})=-\log a for positive a,ba,b, since log\log is the inverse of exp\exp (The Natural Logarithm) and exp\exp turns sums into products (Basic Properties of the Exponential Function).

Step 0: densities. By The Entropy of a Probability Measure on Euclidean Space §entropy, μ\mu and ν\nu have densities ρ\rho and ρ1\rho_{1} with respect to λd\lambda_{d} (Borel, real, nonnegative), with ϕρ\phi\circ\rho and ϕρ1\phi\circ\rho_{1} integrable, Ent(μ)=ϕρdλd\mathrm{Ent}(\mu)=\int\phi\circ\rho\,d\lambda_{d} and Ent(ν)=ϕρ1dλd\mathrm{Ent}(\nu)=\int\phi\circ\rho_{1}\,d\lambda_{d}. By claim 3 of Image Measures, Measures with Densities, and Change of Variables, udμ=uρdλd\int u\,d\mu=\int u\rho\,d\lambda_{d} for Borel u0u\ge0, and for Borel real uu with uρu\rho integrable; likewise for ν\nu and ρ1\rho_{1}. Both measures are absolutely continuous by Basic Properties of the Entropy on the Wasserstein Space: Comparison with the Gaussian Relative Entropy, Lower Bound, Translation Invariance, Absolute Continuity, Closed Sublevel Sets and Lower Semicontinuity §absolutely-continuous. The sets P={ρ>0}P=\{\rho>0\} and P1={ρ1>0}P_{1}=\{\rho_{1}>0\} are Borel, and μ(RdP)=1{ρ=0}ρdλd=0\mu(\mathbb{R}^{d}\setminus P)=\int\mathbf{1}_{\{\rho=0\}}\rho\,d\lambda_{d}=0, so PP is μ\mu-full; likewise P1P_{1} is ν\nu-full. Let ,1:RdR\ell,\ell_{1}:\mathbb{R}^{d}\to\mathbb{R} equal logρ\log\rho on PP, respectively logρ1\log\rho_{1} on P1P_{1}, and 00 elsewhere; they are Borel. Since ρ=ϕρ|\ell|\rho=|\phi\circ\rho| everywhere (both vanish off PP), \ell is μ\mu-integrable with dμ=ϕρdλd=Ent(μ)\int\ell\,d\mu=\int\phi\circ\rho\,d\lambda_{d}=\mathrm{Ent}(\mu); likewise 1\ell_{1} is ν\nu-integrable with 1dν=Ent(ν)\int\ell_{1}\,d\nu=\mathrm{Ent}(\nu).

Step 1: the potentials. The coupling (id,T)#μ(\mathrm{id},T)_{\#}\mu is optimal. By Brenier's Theorem: Optimal Couplings out of an Absolutely Continuous Measure are Induced by a Unique Map §potential there are a Borel T0T_{0} with (id,T0)#μ=(id,T)#μ(\mathrm{id},T_{0})_{\#}\mu=(\mathrm{id},T)_{\#}\mu, an open convex GG with μ(G)=1\mu(G)=1, a convex φ:GR\varphi:G\to\mathbb{R} and a Borel DGD\subseteq G with μ(D)=1\mu(D)=1 and Gφ(x)={T0(x)}\partial_{G}\varphi(x)=\{T_{0}(x)\} for xDx\in D. By Brenier's Theorem: Optimal Couplings out of an Absolutely Continuous Measure are Induced by a Unique Map §map, (T0)#μ=ν(T_{0})_{\#}\mu=\nu, so T0T_{0} is an optimal map from μ\mu to ν\nu, and by Brenier's Theorem: Optimal Couplings out of an Absolutely Continuous Measure are Induced by a Unique Map §unique-map, T=T0T=T_{0} μ\mu-almost everywhere, so TT and T0T_{0} have the same class in L2(μ;Rd)L^{2}(\mu;\mathbb{R}^{d}). By Existence of an Optimal Coupling of Two Probability Measures with Finite Second Moment there is an optimal coupling of ν\nu and μ\mu, and Brenier's Theorem: Optimal Couplings out of an Absolutely Continuous Measure are Induced by a Unique Map §potential, applied to it with ν\nu absolutely continuous, gives a Borel SS, an open convex GG' with ν(G)=1\nu(G')=1, a convex ψ:GR\psi:G'\to\mathbb{R} and a Borel DGD'\subseteq G' with ν(D)=1\nu(D')=1 and Gψ(y)={S(y)}\partial_{G'}\psi(y)=\{S(y)\} for yDy\in D', where (id,S)#ν(\mathrm{id},S)_{\#}\nu is that optimal coupling and S#ν=μS_{\#}\nu=\mu by Brenier's Theorem: Optimal Couplings out of an Absolutely Continuous Measure are Induced by a Unique Map §map, so that SS is an optimal map from ν\nu to μ\mu.

Step 2: inverse Hessians. By Along an Optimal Map between Absolutely Continuous Measures the Hessians of the Two Convex Potentials are Inverse Matrices §hessians, applied to μ,ν,T0,S,G,G,φ,ψ,D,D\mu,\nu,T_{0},S,G,G',\varphi,\psi,D,D', there is a μ\mu-full Borel XDX\subseteq D such that every xXx\in X satisfies: T0(x)DT_{0}(x)\in D'; φ\varphi is twice differentiable at xx with first-order coefficient T0(x)T_{0}(x); ψ\psi is twice differentiable at T0(x)T_{0}(x) with first-order coefficient xx; D2φ(x)D^{2}\varphi(x) is positive definite; and detD2ψ(T0(x))=(detD2φ(x))1\det D^{2}\psi(T_{0}(x))=(\det D^{2}\varphi(x))^{-1}, the determinant detD2φ(x)\det D^{2}\varphi(x) being positive by Determinants of Positive Definite Matrices: Positivity, the Bound logdetAtrAd\log\det A\le\mathrm{tr}\,A-d, Bounds under Pinching, and the Expansion of det(I+tB)\det(I+tB) §positive.

Step 3: the area inequality. By The Points of Twice Differentiability of a Convex Function: a Borel Set of Full Measure, and Borel Measurability of the Gradient and Hessian on It §full for ψ\psi on GG' there is a Borel AψGA_{\psi}\subseteq G' with λd(GAψ)=0\lambda_{d}(G'\setminus A_{\psi})=0 at whose points ψ\psi is twice differentiable; put A=DAψA'=D'\cap A_{\psi}, which is ν\nu-full (ν(GAψ)=0\nu(G'\setminus A_{\psi})=0 by absolute continuity). Let kk be the function of The Area Inequality for the Gradient of a Convex Function §area for f=ψf=\psi, U=GU=G', A=AA=A' and h=ρh=\rho: k(y)=ρ(Dψ(y))detD2ψ(y)k(y)=\rho(D\psi(y))\det D^{2}\psi(y) for yAy\in A' and k(y)=0k(y)=0 otherwise. That theorem gives kdλdρdλd=μ(Rd)=1\int k\,d\lambda_{d}\le\int\rho\,d\lambda_{d}=\mu(\mathbb{R}^{d})=1. Let u:RdRu:\mathbb{R}^{d}\to\mathbb{R} be u=kρ11u=k\rho_{1}^{-1} on P1P_{1} and u=0u=0 elsewhere, a nonnegative Borel function. Then udν=1P1kdλd1\int u\,d\nu=\int\mathbf{1}_{P_{1}}k\,d\lambda_{d}\le1, and since (T0)#μ=ν(T_{0})_{\#}\mu=\nu, Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pushforward gives

uT0dμ=udν1.(1)\int u\circ T_{0}\,d\mu=\int u\,d\nu\le1.\qquad(1)

Step 4: the pointwise inequality. Let X=XPT01(P1A)X'=X\cap P\cap T_{0}^{-1}(P_{1}\cap A'), a Borel set which is μ\mu-full because μ(T01(P1A))=ν(P1A)=1\mu(T_{0}^{-1}(P_{1}\cap A'))=\nu(P_{1}\cap A')=1. Let Δ\Delta be the Borel function of The Points of Twice Differentiability of a Convex Function: a Borel Set of Full Measure, and Borel Measurability of the Gradient and Hessian on It §borel for φ\varphi, GG and the set XX': Δ(x)=trD2φ(x)\Delta(x)=\mathrm{tr}\,D^{2}\varphi(x) on XX' and 00 elsewhere, with 0Δ0\le\Delta by The Points of Twice Differentiability of a Convex Function: a Borel Set of Full Measure, and Borel Measurability of the Gradient and Hessian on It §nonnegative. Let xXx\in X' and y=T0(x)y=T_{0}(x). Then yAP1y\in A'\cap P_{1} and Dψ(y)=xD\psi(y)=x by Step 2, so k(y)=ρ(x)(detD2φ(x))1k(y)=\rho(x)\,(\det D^{2}\varphi(x))^{-1} and

u(y)=ρ(x)ρ1(y)detD2φ(x)>0.u(y)=\frac{\rho(x)}{\rho_{1}(y)\,\det D^{2}\varphi(x)}>0 .

By The Function slogss\log s: Continuity, Young's Inequality and Lower Bounds, with the Elementary Bounds for the Exponential and the Logarithm §log, logu(y)u(y)1\log u(y)\le u(y)-1, that is, (x)1(y)logdetD2φ(x)u(y)1\ell(x)-\ell_{1}(y)-\log\det D^{2}\varphi(x)\le u(y)-1. By Determinants of Positive Definite Matrices: Positivity, the Bound logdetAtrAd\log\det A\le\mathrm{tr}\,A-d, Bounds under Pinching, and the Expansion of det(I+tB)\det(I+tB) §log-det, logdetD2φ(x)Δ(x)d\log\det D^{2}\varphi(x)\le\Delta(x)-d. Hence, for every xXx\in X',

(x)1(T0(x))+dΔ(x)u(T0(x))1.(2)\ell(x)-\ell_{1}(T_{0}(x))+d-\Delta(x)\le u(T_{0}(x))-1.\qquad(2)

Step 5: the score. Apply The Score Paired with the Identity, and the Laplacian of a Convex Potential Bounded by its Pairing with the Score §comparison with GG, φ\varphi and A=XA=X' (a μ\mu-full Borel subset of GG at whose points φ\varphi is twice differentiable); μ\mu is absolutely continuous with finite Fisher information. Its map gg equals Dφ=T0D\varphi=T_{0} on XX', so g=T0=Tg=T_{0}=T μ\mu-almost everywhere; hence g2dμ=T2dμ<\int\lVert g\rVert^{2}\,d\mu=\int\lVert T\rVert^{2}\,d\mu<\infty and gg and TT have the same class in L2(μ;Rd)L^{2}(\mu;\mathbb{R}^{d}). Its function Δ\Delta is the one of Step 4. The lemma gives that Δ\Delta is μ\mu-integrable and

Δdμξμ,Tμ,(3)\int\Delta\,d\mu\le-\langle\xi_{\mu},T\rangle_{\mu},\qquad(3)

and The Score Paired with the Identity, and the Laplacian of a Convex Potential Bounded by its Pairing with the Score §identity gives ξμ,idμ=d\langle\xi_{\mu},\mathrm{id}\rangle_{\mu}=-d.

Step 6: integration. By Step 0 and Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pushforward, 1T0\ell_{1}\circ T_{0} is μ\mu-integrable with 1T0dμ=1dν=Ent(ν)\int\ell_{1}\circ T_{0}\,d\mu=\int\ell_{1}\,d\nu=\mathrm{Ent}(\nu); \ell is μ\mu-integrable with dμ=Ent(μ)\int\ell\,d\mu=\mathrm{Ent}(\mu) by Step 0; uT0u\circ T_{0} is nonnegative and μ\mu-integrable by (1); and Δ\Delta is μ\mu-integrable by Step 5. So every term of (2) is μ\mu-integrable, (2) holds on the μ\mu-full set XX', and integrating it (linearity and monotonicity of the integral of integrable functions, claim 2 of Linearity and Monotonicity of the Lebesgue Integral, with The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §comparison for the μ\mu-null set RdX\mathbb{R}^{d}\setminus X') and using (1) gives

Ent(μ)Ent(ν)+dΔdμuT0dμ10.(4)\mathrm{Ent}(\mu)-\mathrm{Ent}(\nu)+d-\int\Delta\,d\mu\le\int u\circ T_{0}\,d\mu-1\le0.\qquad(4)

By bilinearity of the inner product, the identity of Step 5, (3) and (4),

ξμ,Tidμ=ξμ,Tμ+ddΔdμEnt(ν)Ent(μ),\langle\xi_{\mu},T-\mathrm{id}\rangle_{\mu}=\langle\xi_{\mu},T\rangle_{\mu}+d\le d-\int\Delta\,d\mu\le\mathrm{Ent}(\nu)-\mathrm{Ent}(\mu),

which is the tangent inequality.

Please log in to copy this version.

Citations

Loading…

Dependency Graph

0 prerequisites

Comments

Loading…