Each result cited is universally quantified over the data in its own statement. Elementary order and field arithmetic in R is used without mention (Elementary Order Arithmetic in an Ordered Field, Elementary Arithmetic in an Ordered Field). Integrals against λd are written ∫⋯dλd; a Borel set is μ-full if its complement is μ-null, and finitely many μ-full sets have a μ-full intersection (claim 4 of Basic Properties of a Measure); likewise for ν. We use that exp(u+v)=exp(u)exp(v), exp(0)=1, 0<exp(u) and exp(−u)=exp(u)−1 (claims 1 and 2 of Basic Properties of the Exponential Function), and that log is the inverse of exp (The Natural Logarithm).
Step 0: densities. By Finite Fisher Information, the Score and the Fisher Information of a Probability Measure §finite, μ,ν∈P2(Rd). By The Density Cost of a Convex Lipschitz Integrand §cost, μ and ν have densities ρ and ρ1 with respect to λd (Borel, real, nonnegative), the functions Φ∘ρ and Φ∘ρ1 are Borel and λd-integrable, and GΦ(μ)=∫Φ∘ρdλd, GΦ(ν)=∫Φ∘ρ1dλd. Since μ is the measure with density ρ with respect to λd, claim 3 of Image Measures, Measures with Densities, and Change of Variables gives ∫udμ=∫uρdλd for Borel u≥0, and, for Borel real u, that u is μ-integrable exactly when uρ is λd-integrable, the same identity then holding. By the Lipschitz bound of Convex Lipschitz Integrands §integrand with a=0 and Φ(0)=0, 0≤Φ(s) for every s≥0.
Step 1: the potentials. The coupling (id,T)#μ is optimal (Optimal Transport Maps and Uniquely Mapped Pairs of Probability Measures §map). By Brenier's Theorem: Optimal Couplings out of an Absolutely Continuous Measure are Induced by a Unique Map §potential there are a Borel T0 with (id,T0)#μ=(id,T)#μ, an open convex G with μ(G)=1, a convex φ:G→R and a Borel D⊆G with μ(D)=1 and ∂Gφ(x)={T0(x)} for x∈D. By Brenier's Theorem: Optimal Couplings out of an Absolutely Continuous Measure are Induced by a Unique Map §map, (T0)#μ=ν and ∫∥T0∥2dμ<∞, so T0 is an optimal map from μ to ν, and by Brenier's Theorem: Optimal Couplings out of an Absolutely Continuous Measure are Induced by a Unique Map §unique-map, T=T0 μ-almost everywhere. Likewise (id,T′)#ν is an optimal coupling of ν and μ, and the same three clauses, applied with ν (absolutely continuous) in place of μ, give a Borel S with (id,S)#ν=(id,T′)#ν, an open convex G′ with ν(G′)=1, a convex ψ:G′→R and a Borel D′⊆G′ with ν(D′)=1 and ∂G′ψ(y)={S(y)} for y∈D′, such that S#ν=μ, ∫∥S∥2dν<∞, S is an optimal map from ν to μ, and S=T′ ν-almost everywhere.
Step 2: the good set. By McCann's Jacobian Equation Along an Optimal Map Between Absolutely Continuous Measures §jacobian, applied to μ,ν,ρ,ρ1,T0,S,G,G′,φ,ψ,D,D′, there is a μ-full Borel X⊆D such that every x∈X satisfies: φ is twice differentiable at x with first-order coefficient T0(x); ψ is twice differentiable at T0(x); D2φ(x) and D2ψ(T0(x)) are positive definite; detD2φ(x)⋅detD2ψ(T0(x))=1 (these from Along an Optimal Map between Absolutely Continuous Measures the Hessians of the Two Convex Potentials are Inverse Matrices §hessians); and ρ(x)=ρ1(T0(x))detD2φ(x) with ρ(x)>0. By Along an Optimal Map between Absolutely Continuous Measures the Hessians of the Two Convex Potentials are Inverse Matrices §hessians, applied to ν,μ,S,T0,G′,G,ψ,φ,D′,D (its hypotheses are symmetric in the two measures), there is a ν-full Borel Y⊆D′ such that at every y∈Y the function ψ is twice differentiable with first-order coefficient S(y). Let X′=X∩T0−1(Y), a Borel set which is μ-full because μ(T0−1(Y))=ν(Y)=1. Let Δφ be the function Δ of The Points of Twice Differentiability of a Convex Function: a Borel Set of Full Measure, and Borel Measurability of the Gradient and Hessian on It §borel for φ, G and the set X′, and Δψ the one for ψ, G′ and the set Y: Δφ(x)=trD2φ(x) on X′, Δψ(y)=trD2ψ(y) on Y, both 0 elsewhere, and both Borel. Let H:Rd→R be the Borel function
H=A(Δφ−d)+A(Δψ∘T0−d)+AdL2.
Step 3: the pointwise inequality. Fix x∈X′ and write r=ρ(x)>0, M=D2φ(x), N=D2ψ(T0(x)) and δ=detM, so that 0<δ by Determinants of Positive Definite Matrices: Positivity, the Bound logdetA≤trA−d, Bounds under Pinching, and the Expansion of det(I+tB) §positive, detN=δ−1, ρ1(T0(x))=rδ−1 by Step 2, and trM=Δφ(x), trN=Δψ(T0(x)) since T0(x)∈Y, so that H(x)=A(trM−d)+A(trN−d)+dL2A−1. We show
Φ(r)−δΦ(rδ−1)≤rH(x).(1)
Trace bounds. Let g=d−1logδ, t=exp(g/2) and s=t2=exp(g); by the exponential rules and induction on d, sd=exp(dg)=exp(logδ)=δ. For a positive real c and a positive definite B∈S(d), the matrix cB is symmetric and positive definite (z⋅(cBz)=c(z⋅Bz) for z∈Rd, Symmetric, Positive Semidefinite, and Positive Definite Real Matrices), tr(cB)=ctrB (Trace of a Real Square Matrix), and det(cB)=det(cId)detB=cddetB by The Determinant is Multiplicative and claim 1 of The Determinant of a Triangular Matrix is the Product of its Diagonal Entries, cId being lower triangular with all diagonal entries c. With c=s−1 and B=M this gives det(s−1M)=s−dδ=1, so Determinants of Positive Definite Matrices: Positivity, the Bound logdetA≤trA−d, Bounds under Pinching, and the Expansion of det(I+tB) §log-det yields 0=log1≤s−1trM−d, that is trM≥ds; with c=s and B=N it gives det(sN)=δδ−1=1 and likewise trN≥ds−1. Hence
(trM−d)+(trN−d)≥d(t2+t−2−2)=d(t−t−1)2≥0.(2)
Case δ≤1. Convexity in Convex Lipschitz Integrands §integrand, with a=rδ−1, b=0 and weight δ, together with Φ(0)=0, gives Φ(r)=Φ(δ⋅rδ−1+(1−δ)⋅0)≤δΦ(rδ−1), so the left side of (1) is at most 0, while rH(x)≥0 by (2), since r>0, A>0 and dL2A−1≥0.
Case 1<δ. By The Function slogs: Continuity, Young's Inequality and Lower Bounds, with the Elementary Bounds for the Exponential and the Logarithm §log, 1−δ−1≤logδ, so 0<logδ and 0<g. By Step 0, 0≤(δ−1)Φ(rδ−1), i.e. Φ(rδ−1)≤δΦ(rδ−1); and rδ−1≤r, so the Lipschitz bound of Convex Lipschitz Integrands §integrand gives
Φ(r)−δΦ(rδ−1)≤Φ(r)−Φ(rδ−1)≤Lr(1−δ−1)≤Lrlogδ=rdLg.
By The Function slogs: Continuity, Young's Inequality and Lower Bounds, with the Elementary Bounds for the Exponential and the Logarithm §exp, t≥p with p=1+g/2≥1, so t−1≤p−1 and t−t−1≥p−p−1=(p−1)(p+1)p−1≥p−1=g/2≥0; hence (t−t−1)2≥g2/4. From 0≤A(g/2−L/A)2=Ag2/4−Lg+L2/A we get Lg≤Ag2/4+L2/A≤A(t−t−1)2+L2/A, and multiplying by rd>0 and using (2),
rdLg≤r(Ad(t−t−1)2+AdL2)≤rH(x).
This proves (1) in both cases. By Step 2, φ has first-order coefficient T0(x) at x and δΦ(rδ−1)=Φ(ρ1(T0(x)))detD2φ(x).
Step 4: the scores. Apply The Score Paired with the Identity, and the Laplacian of a Convex Potential Bounded by its Pairing with the Score §comparison to μ (absolutely continuous, with finite Fisher information) with G, φ and, as the set called A there (not our constant A), the set X′, a μ-full Borel subset of G at whose points φ is twice differentiable; its function Δ is Δφ, and its map gμ equals Dφ=T0 on X′, hence gμ=T0=T μ-almost everywhere by Step 1. So ∫∥gμ∥2dμ=∫∥T0∥2dμ<∞ by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §comparison, and gμ and T have the same class in L2(μ;Rd). That clause gives that Δφ is μ-integrable with ∫Δφdμ≤−⟨ξμ,T⟩μ, and The Score Paired with the Identity, and the Laplacian of a Convex Potential Bounded by its Pairing with the Score §identity gives ⟨ξμ,id⟩μ=−d; by bilinearity of the inner product,
∫Δφdμ−d≤⟨ξμ,id−T⟩μ.(3)
In the same way, The Score Paired with the Identity, and the Laplacian of a Convex Potential Bounded by its Pairing with the Score §comparison applied to ν with G′, ψ and the set Y, whose map equals S on Y and hence T′ ν-almost everywhere, with ∫∥S∥2dν<∞ by Step 1, together with The Score Paired with the Identity, and the Laplacian of a Convex Potential Bounded by its Pairing with the Score §identity for ν, shows that Δψ is ν-integrable and
∫Δψdν−d≤⟨ξν,id−T′⟩ν.(4)
Since (T0)#μ=ν, Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pushforward shows that Δψ∘T0 is μ-integrable with ∫Δψ∘T0dμ=∫Δψdν. Constants being μ-integrable with integral equal to themselves (μ(Rd)=1), H is μ-integrable and, by claim 2 of Linearity and Monotonicity of the Lebesgue Integral, (3) and (4),
∫Hdμ=A(∫Δφdμ−d)+A(∫Δψdν−d)+AdL2≤A(⟨ξμ,id−T⟩μ+⟨ξν,id−T′⟩ν)+AdL2.(5)
Step 5: integration. Let k be the function of The Area Inequality for the Gradient of a Convex Function §area for f=φ, U=G, the set X′ in place of its set A, and h=Φ∘ρ1 (Borel and nonnegative by Step 0): by Step 3, k(x)=Φ(ρ1(T0(x)))detD2φ(x)=δΦ(rδ−1) in the notation there for x∈X′, and k(x)=0 otherwise. That theorem gives that k is Borel and nonnegative with ∫kdλd≤∫Φ∘ρ1dλd=GΦ(ν); so k is λd-integrable. Next, ∫1Rd∖X′ρdλd=μ(Rd∖X′)=0, so by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §vanishing there is a λd-null set Z with ρ(x)=0 for every x∈/X′∪Z. Outside Z, Φ∘ρ=1X′Φ∘ρ (as Φ(0)=0) and Hρ=1X′Hρ. By Step 0, Hρ is λd-integrable with ∫Hρdλd=∫Hdμ, so by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §comparison the functions 1X′Φ∘ρ and 1X′Hρ are λd-integrable with integrals GΦ(μ) and ∫Hdμ. By (1), 1X′Φ∘ρ−k≤1X′Hρ on X′, and also off X′, where both sides vanish. Integrating (claim 2 of Linearity and Monotonicity of the Lebesgue Integral) and using (5),
GΦ(μ)−GΦ(ν)≤GΦ(μ)−∫kdλd≤∫Hdμ≤A(⟨ξμ,id−T⟩μ+⟨ξν,id−T′⟩ν)+AdL2,
which is the displacement bound.