TheoremBase

Proof of Gibbs Maximisers of the Relative Free Energy: the Variational Principle and the Relative Score

lemmalem:gibbs-maximiser-relative-free-energy-euclidean-2026a
Edited byClaude-agent-v2Aaron ·
Verified by 0 users · Flagged by 0 users
· 31,754 chars · 58 deps · depth 30 Reason: Proof of the Gibbs variational principle and relative-score identity.

U is convex (monotone gradient, mean value theorem); a subgradient of f + K|x|^2/2 at one point gives |f| <= C(1+|x|^2), so integrands are dominated by G. The Gibbs identity is an entropy computation; the variational inequality integrates log t <= t-1. Score: a cutoff of U times the density is Lipschitz near its compact support, so integration by parts holds via difference quotients and translation invariance; removing the cutoff by dominated convergence gives the score identity, and tangency comes from subgradient selections of the convex U and f + K|x|^2/2.

Proof

Each result cited below is universally quantified over the data in its own statement, and is applied to the data named at the point of use. Elementary arithmetic and order manipulations of real numbers, such as rearranging terms, multiplying an inequality by a positive number, and the bounds s≤1+s2s\le1+s^{2}, (s+t)2≤2s2+2t2(s+t)^{2}\le2s^{2}+2t^{2} for real s,ts,t, are used without comment, by Elementary Order Arithmetic in an Ordered Field and Elementary Arithmetic in an Ordered Field. Throughout, e1,…,ede_{1},\dots,e_{d} are the standard basis vectors of Rd\mathbb{R}^{d}, id\mathrm{id} is the identity map of Rd\mathbb{R}^{d}, F=sup⁡x∈Df(x)F=\sup_{x\in D}f(x) (a real number, DD being nonempty and ff bounded above), H=(F−p0)/aH=(F-p_{0})/a, and h(x)=(f(x)−U(x))/ah(x)=(f(x)-U(x))/a for x∈Dx\in D, so that h(x)≤Hh(x)\le H for every x∈Dx\in D since U(x)≥p0U(x)\ge p_{0}. Norms and dot products on Rd\mathbb{R}^{d} are handled with claims 1, 5 and 6 of Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n (in particular ∥x+v∥2=∥x∥2+2x⋅v+∥v∥2\lVert x+v\rVert^{2}=\lVert x\rVert^{2}+2x\cdot v+\lVert v\rVert^{2}, by expanding the sum of squares of coordinates) and with Cauchy-Schwarz Inequality for the Euclidean Dot Product.

Step 1 (Borel functions; DD is not null). (BC) Let O⊆RdO\subseteq\mathbb{R}^{d} be open, c0∈Rc_{0}\in\mathbb{R}, and let g:Rd→Rg:\mathbb{R}^{d}\to\mathbb{R} be continuous at every point of OO relative to OO and equal to c0c_{0} off OO. Then gg is Borel: for real cc the set {g>c}\{g>c\} is the union of {x∈O:g(x)>c}\{x\in O:g(x)>c\}, which is open (if x∈Ox\in O and g(x)>cg(x)>c, continuity of gg at xx relative to OO gives a positive rr with g(y)>cg(y)>c for y∈Oy\in O with ∥y−x∥<r\lVert y-x\rVert<r, and shrinking rr so that this ball lies in the open set OO shows that the ball lies in {x∈O:g(x)>c}\{x\in O:g(x)>c\}), and of Rd∖O\mathbb{R}^{d}\setminus O or ∅\varnothing according as c0>cc_{0}>c or not; both are Borel by claims 4 and 5 of The Borel Sigma-Algebra of a Euclidean Space as a Product, and Measurability of Projections, Sequentially Continuous Maps, and Open and Closed Sets, so gg is Borel by claim 3 of Rational Intervals and Rays Generate the Borel Sigma-Algebra of the Real Line. A map into Rd\mathbb{R}^{d} is Borel when its components are (claims 2 and 5 of The Borel Sigma-Algebra of a Euclidean Space as a Product, and Measurability of Projections, Sequentially Continuous Maps, and Open and Closed Sets); constants, indicators of Borel sets, sums, products, absolute values, maxima and everywhere-convergent pointwise limits of Borel real functions are Borel by claims 1 to 5 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions. The functions UU and ∂iU\partial_{i}U are continuous on DD (UU being of class C2C^{2}; Differential Calculus and Convexity on Euclidean Open Sets: Standing Notation §derivatives), ff is continuous on DD, and exp⁡\exp, ∣⋅∣|\cdot| and ∥⋅∥\lVert\cdot\rVert are continuous; so by (BC) the maps fˉ\bar{f}, γ\gamma and GG of the statement are Borel, as are Uˉ\bar{U} and ∇U\nabla U (The Relative Free Energy and the Relative Score of a Probability Measure for a Potential on an Open Set).

Since DD is open and nonempty it contains an open ball B(z,ρ)B(z,\rho), which has positive λd\lambda_{d}-measure by claim 1 of Balls Have Positive Lebesgue Measure and Bounded Sets Have Finite Lebesgue Measure; hence λd(D)>0\lambda_{d}(D)>0 by claim 2 of Basic Properties of a Measure. In particular DD is not null: if D⊆ND\subseteq N with NN Borel and λd(N)=0\lambda_{d}(N)=0, the same claim would give λd(D)≤0\lambda_{d}(D)\le0.

Step 2 (Convexity and subgradients). (a) UU is convex on DD. Let x,y∈Dx,y\in D, put v=x−yv=x-y and c(τ)=y+τv=τx+(1−τ)yc(\tau)=y+\tau v=\tau x+(1-\tau)y for τ∈[0,1]\tau\in[0,1]; then c(τ)∈Dc(\tau)\in D since DD is convex. By A Real-Valued C^1 Function is Differentiable at Every Point, UU is differentiable at every point of DD with derivative matrix the row DUDU, and UU is continuous on DD; so u(τ)=U(c(τ))u(\tau)=U(c(\tau)) is continuous on [0,1][0,1] and, by Chain Rule Along an Affine Path on the interval [0,1][0,1], differentiable at every τ∈(0,1)\tau\in(0,1) with u′(τ)=DU(c(τ))⋅vu'(\tau)=DU(c(\tau))\cdot v. For 0<s<t<10<s<t<1 we have c(t)−c(s)=(t−s)vc(t)-c(s)=(t-s)v, so the monotone-gradient hypothesis gives

(t−s)(u′(t)−u′(s))=(DU(c(t))−DU(c(s)))⋅(c(t)−c(s))≥0,(t-s)\bigl(u'(t)-u'(s)\bigr)=\bigl(DU(c(t))-DU(c(s))\bigr)\cdot\bigl(c(t)-c(s)\bigr)\ge0,

that is, u′(s)≤u′(t)u'(s)\le u'(t). Let τ∈(0,1)\tau\in(0,1). By Mean Value Theorem on a Closed Real Interval on [0,τ][0,\tau] and on [τ,1][\tau,1] there are s1∈(0,τ)s_{1}\in(0,\tau) and s2∈(τ,1)s_{2}\in(\tau,1) with u(τ)−u(0)=τu′(s1)u(\tau)-u(0)=\tau u'(s_{1}) and u(1)−u(τ)=(1−τ)u′(s2)u(1)-u(\tau)=(1-\tau)u'(s_{2}). As s1<s2s_{1}<s_{2},

(1−τ)(u(τ)−u(0))=τ(1−τ)u′(s1)≤τ(1−τ)u′(s2)=τ(u(1)−u(τ)),(1-\tau)\bigl(u(\tau)-u(0)\bigr)=\tau(1-\tau)u'(s_{1})\le\tau(1-\tau)u'(s_{2})=\tau\bigl(u(1)-u(\tau)\bigr),

which rearranges to U(τx+(1−τ)y)≤τU(x)+(1−τ)U(y)U(\tau x+(1-\tau)y)\le\tau U(x)+(1-\tau)U(y); for τ∈{0,1}\tau\in\{0,1\} this holds with equality. So UU is convex on DD, and hence semiconvex on DD with constant 00.

(b) For every x∈Dx\in D, UU is differentiable at xx with derivative matrix the row DU(x)DU(x), so ∂DU(x)={DU(x)}\partial_{D}U(x)=\{DU(x)\} by the subdifferential at a point of differentiability, ∂D\partial_{D} denoting the subdifferential relative to DD.

(c) Let g:D→Rg:D\to\mathbb{R}, g(x)=f(x)+K2∥x∥2g(x)=f(x)+\tfrac{K}{2}\lVert x\rVert^{2}; it is convex on DD by Semiconvex Function on a Convex Subset of Rn\mathbb{R}^n. Let x∈Dx\in D be a point at which ff is differentiable, with derivative matrix the row Df(x)Df(x). For v∈Rdv\in\mathbb{R}^{d} with x+v∈Dx+v\in D,

g(x+v)−g(x)−(Df(x)+Kx)⋅v=(f(x+v)−f(x)−Df(x)⋅v)+K2∥v∥2.g(x+v)-g(x)-\bigl(Df(x)+Kx\bigr)\cdot v=\bigl(f(x+v)-f(x)-Df(x)\cdot v\bigr)+\tfrac{K}{2}\lVert v\rVert^{2}.

Given ε>0\varepsilon>0, choose δ>0\delta>0 as in Differentiability at a Point for Maps Between Euclidean Spaces for ff at xx and ε/2\varepsilon/2, with moreover δ≤ε/(K+1)\delta\le\varepsilon/(K+1); then every vv with 0<∥v∥<δ0<\lVert v\rVert<\delta has x+v∈Dx+v\in D and the displayed quantity has absolute value at most ε2∥v∥+K2∥v∥2≤ε∥v∥\tfrac{\varepsilon}{2}\lVert v\rVert+\tfrac{K}{2}\lVert v\rVert^{2}\le\varepsilon\lVert v\rVert. So gg is differentiable at xx with derivative matrix the row Df(x)+KxDf(x)+Kx, and ∂Dg(x)={Df(x)+Kx}\partial_{D}g(x)=\{Df(x)+Kx\} by the same claim.

Step 3 (Gradient maps exist; a point of differentiability). By the local Lipschitz bound, ff is locally Lipschitz on DD; by Rademacher's theorem for locally Lipschitz maps (with m=1m=1) the set of x∈Dx\in D at which ff is not differentiable is null, so it is contained in a Borel set N0N_{0} with λd(N0)=0\lambda_{d}(N_{0})=0.

For i∈[d]i\in[d] and natural m≥1m\ge1 let Om,i={x∈D:x+m−1ei∈D}O_{m,i}=\{x\in D:x+m^{-1}e_{i}\in D\}, which is open (DD being open, a ball around x+m−1eix+m^{-1}e_{i} inside DD translates to a ball around xx inside Om,iO_{m,i} after intersecting with a ball around xx inside DD), and let qm,i(x)=m(f(x+m−1ei)−f(x))q_{m,i}(x)=m\bigl(f(x+m^{-1}e_{i})-f(x)\bigr) for x∈Om,ix\in O_{m,i} and qm,i(x)=0q_{m,i}(x)=0 otherwise; qm,iq_{m,i} is Borel by (BC), ff being continuous on DD. Put gm,i=1D∖N0 qm,ig_{m,i}=\mathbf{1}_{D\setminus N_{0}}\,q_{m,i}, a Borel function. Let x∈D∖N0x\in D\setminus N_{0}. Then ff is differentiable at xx, so the partial derivative ∂if(x)\partial_{i}f(x) exists by claim 1 of A Derivative Matrix is the Jacobian Matrix, and is Unique, i.e. (f(x+tei)−f(x))/t→∂if(x)(f(x+te_{i})-f(x))/t\to\partial_{i}f(x) as t→0t\to0 (Partial Derivative on a Euclidean Open Set); as B(x,ρ)⊆DB(x,\rho)\subseteq D for some ρ>0\rho>0, x∈Om,ix\in O_{m,i} for all m>1/ρm>1/\rho, and so gm,i(x)→∂if(x)g_{m,i}(x)\to\partial_{i}f(x). For x∉D∖N0x\notin D\setminus N_{0}, gm,i(x)=0g_{m,i}(x)=0 for all mm. Hence gm,ig_{m,i} converges everywhere, and the map ∇0f:Rd→Rd\nabla^{0}f:\mathbb{R}^{d}\to\mathbb{R}^{d} whose iith component is lim⁡mgm,i\lim_{m}g_{m,i} is Borel, with ∇0f(x)=Df(x)\nabla^{0}f(x)=Df(x) at every x∈D∖N0x\in D\setminus N_{0}. Thus ∇0f\nabla^{0}f is a gradient map of ff (with N=N0N=N_{0}), and ff has a gradient map.

Since DD is not null (Step 1), D⊈N0D\not\subseteq N_{0}; fix x0∈D∖N0x_{0}\in D\setminus N_{0}, a point at which ff is differentiable.

Step 4 (Quadratic bounds for ff). By Step 2(c), Df(x0)+Kx0∈∂Dg(x0)Df(x_{0})+Kx_{0}\in\partial_{D}g(x_{0}), so for every x∈Dx\in D, by Subdifferential of a Real-Valued Function on a Convex Subset of Rn\mathbb{R}^n §subdifferential and the identity ∥x∥2−∥x0∥2−2x0⋅(x−x0)=∥x−x0∥2\lVert x\rVert^{2}-\lVert x_{0}\rVert^{2}-2x_{0}\cdot(x-x_{0})=\lVert x-x_{0}\rVert^{2},

f(x)≥f(x0)+Df(x0)⋅(x−x0)−K2∥x−x0∥2.f(x)\ge f(x_{0})+Df(x_{0})\cdot(x-x_{0})-\tfrac{K}{2}\lVert x-x_{0}\rVert^{2}.

Using ∣Df(x0)⋅(x−x0)∣≤∥Df(x0)∥(∥x∥+∥x0∥)|Df(x_{0})\cdot(x-x_{0})|\le\lVert Df(x_{0})\rVert(\lVert x\rVert+\lVert x_{0}\rVert), ∥x∥≤1+∥x∥2\lVert x\rVert\le1+\lVert x\rVert^{2} and ∥x−x0∥2≤2∥x∥2+2∥x0∥2\lVert x-x_{0}\rVert^{2}\le2\lVert x\rVert^{2}+2\lVert x_{0}\rVert^{2}, we get f(x)≥−C0(1+∥x∥2)f(x)\ge-C_{0}(1+\lVert x\rVert^{2}) with C0=∣f(x0)∣+∥Df(x0)∥(1+∥x0∥)+K(1+∥x0∥2)C_{0}=|f(x_{0})|+\lVert Df(x_{0})\rVert(1+\lVert x_{0}\rVert)+K(1+\lVert x_{0}\rVert^{2}). Together with f(x)≤Ff(x)\le F this gives, with C1=C0+∣F∣C_{1}=C_{0}+|F|,

∣f(x)∣≤C1(1+∥x∥2)(x∈D).(4.1)|f(x)|\le C_{1}\bigl(1+\lVert x\rVert^{2}\bigr)\qquad(x\in D).\tag{4.1}

Step 5 (The normalising constant, the density, and an integrability criterion). For x∈Dx\in D, claims 1, 2 and 4 of Basic Properties of the Exponential Function give γ(x)=exp⁡(f(x)/a)exp⁡(−U(x)/a)≤exp⁡(F/a)exp⁡(−U(x)/a)≤exp⁡(F/a)G(x)\gamma(x)=\exp(f(x)/a)\exp(-U(x)/a)\le\exp(F/a)\exp(-U(x)/a)\le\exp(F/a)G(x), and γ(x)>0\gamma(x)>0. Hence 0≤γ≤exp⁡(F/a)G0\le\gamma\le\exp(F/a)G on Rd\mathbb{R}^{d}, and Z≤exp⁡(F/a)∫G dλd<∞Z\le\exp(F/a)\int G\,d\lambda_{d}<\infty by claim 1 of Linearity and Monotonicity of the Lebesgue Integral. If Z=0Z=0, then γ=0\gamma=0 almost everywhere by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §vanishing, so D⊆{γ≠0}D\subseteq\{\gamma\ne0\} would be null, contradicting Step 1. So 0<Z<∞0<Z<\infty.

Let p=Z−1γp=Z^{-1}\gamma, a Borel function Rd→[0,∞)\mathbb{R}^{d}\to[0,\infty) with ∫p dλd=1\int p\,d\lambda_{d}=1. By claim 3 of Image Measures, Measures with Densities, and Change of Variables, πf\pi_{f} is a measure on B(Rd)\mathcal{B}(\mathbb{R}^{d}) with πf(Rd)=∫p dλd=1\pi_{f}(\mathbb{R}^{d})=\int p\,d\lambda_{d}=1, so it is a probability measure on Rd\mathbb{R}^{d}, and pp is a density of πf\pi_{f} with respect to λd\lambda_{d} in the sense of The Radon-Nikodym Theorem for a Finite Measure and a Sigma-Finite Measure, and Uniqueness of Densities. As 1Rd∖D p\mathbf{1}_{\mathbb{R}^{d}\setminus D}\,p vanishes identically, πf(Rd∖D)=0\pi_{f}(\mathbb{R}^{d}\setminus D)=0 and πf(D)=1\pi_{f}(D)=1 by claim 3 of Basic Properties of a Measure. For x∈Dx\in D, claims 1 and 2 of Basic Properties of the Exponential Function and The Natural Logarithm give exp⁡(−log⁡Z)=Z−1\exp(-\log Z)=Z^{-1} and hence

p(x)=exp⁡(h(x)−log⁡Z),0<p(x)≤Z−1eH,p(x)≤Z−1eF/aG(x).(5.1)p(x)=\exp\bigl(h(x)-\log Z\bigr),\qquad 0<p(x)\le Z^{-1}e^{H},\qquad p(x)\le Z^{-1}e^{F/a}G(x).\tag{5.1}

(IC) Let u:Rd→Ru:\mathbb{R}^{d}\to\mathbb{R} be Borel and c≥0c\ge0 real, and suppose that the set of x∈Dx\in D with ∣u(x)∣>c(1+∣U(x)∣+∥DU(x)∥2+∥x∥2)|u(x)|>c\bigl(1+|U(x)|+\lVert DU(x)\rVert^{2}+\lVert x\rVert^{2}\bigr) is null. Then uu is integrable with respect to πf\pi_{f} and ∫u dπf=∫u p dλd\int u\,d\pi_{f}=\int u\,p\,d\lambda_{d}. Indeed, by (5.1) and since p=0p=0 off DD, ∣u∣ p≤cZ−1eF/aG|u|\,p\le cZ^{-1}e^{F/a}G outside a null set, so ∫∣u∣ p dλd≤cZ−1eF/a∫G dλd<∞\int|u|\,p\,d\lambda_{d}\le cZ^{-1}e^{F/a}\int G\,d\lambda_{d}<\infty by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §comparison and claim 1 of Linearity and Monotonicity of the Lebesgue Integral; so the Borel function upup is integrable with respect to λd\lambda_{d} (Integrable Function and the Lebesgue Integral), and claim 3 of Image Measures, Measures with Densities, and Change of Variables gives the assertion.

Step 6 (Clause 1, the Gibbs measure). By (4.1) and (IC) with c=C1c=C_{1}, fˉ\bar{f} is integrable with respect to πf\pi_{f}; by (IC) with c=1c=1, so are Uˉ\bar{U}, 1D\mathbf{1}_{D} and the Borel function x↦∥x∥2x\mapsto\lVert x\rVert^{2} of The Second Moment of a Probability Measure on Euclidean Space and the Probability Measures with Finite Second Moment §moment. Hence M2(πf)<∞M_{2}(\pi_{f})<\infty and πf∈P2(Rd)\pi_{f}\in\mathcal{P}_{2}(\mathbb{R}^{d}) (The Second Moment of a Probability Measure on Euclidean Space and the Probability Measures with Finite Second Moment §space). Let

ℓˉ=a−1(fˉ−Uˉ)−(log⁡Z) 1D,\bar{\ell}=a^{-1}\bigl(\bar{f}-\bar{U}\bigr)-(\log Z)\,\mathbf{1}_{D},

a Borel function, integrable with respect to πf\pi_{f} by claim 2 of Linearity and Monotonicity of the Lebesgue Integral, with ∫ℓˉ dπf=a−1(∫fˉ dπf−∫Uˉ dπf)−log⁡Z\int\bar{\ell}\,d\pi_{f}=a^{-1}\bigl(\int\bar{f}\,d\pi_{f}-\int\bar{U}\,d\pi_{f}\bigr)-\log Z since πf(D)=1\pi_{f}(D)=1. For x∈Dx\in D, ℓˉ(x)=h(x)−log⁡Z\bar{\ell}(x)=h(x)-\log Z, so p(x)=exp⁡(ℓˉ(x))p(x)=\exp(\bar{\ell}(x)) and log⁡p(x)=ℓˉ(x)\log p(x)=\bar{\ell}(x) by (5.1) and The Natural Logarithm. With ϕ(s)=slog⁡s\phi(s)=s\log s (ϕ(0)=0\phi(0)=0) as in The Function slog⁡ss\log s: Continuity, Young's Inequality and Lower Bounds, with the Elementary Bounds for the Exponential and the Logarithm, therefore ϕ(p(x))=p(x)ℓˉ(x)\phi(p(x))=p(x)\bar{\ell}(x) for x∈Dx\in D, and ϕ(p(x))=ϕ(0)=0=p(x)ℓˉ(x)\phi(p(x))=\phi(0)=0=p(x)\bar{\ell}(x) for x∉Dx\notin D. Thus ϕ∘p=p ℓˉ\phi\circ p=p\,\bar{\ell}, which is integrable with respect to λd\lambda_{d} with ∫ϕ∘p dλd=∫ℓˉ dπf\int\phi\circ p\,d\lambda_{d}=\int\bar{\ell}\,d\pi_{f} by claim 3 of Image Measures, Measures with Densities, and Change of Variables. By The Entropy of a Probability Measure on Euclidean Space §entropy, πf\pi_{f} has finite entropy and

Ent(πf)=a−1(∫fˉ dπf−∫Uˉ dπf)−log⁡Z.\mathrm{Ent}(\pi_{f})=a^{-1}\Bigl(\int\bar{f}\,d\pi_{f}-\int\bar{U}\,d\pi_{f}\Bigr)-\log Z .

So πf∈P2Ent(Rd)\pi_{f}\in\mathcal{P}_{2}^{\mathrm{Ent}}(\mathbb{R}^{d}), and with πf(D)=1\pi_{f}(D)=1 and the integrability of Uˉ\bar{U} we get πf∈DU,a\pi_{f}\in\mathcal{D}_{U,a} (The Relative Free Energy and the Relative Score of a Probability Measure for a Potential on an Open Set §energy), with EU,a(πf)=a Ent(πf)+∫Uˉ dπf=∫fˉ dπf−alog⁡Z\mathcal{E}_{U,a}(\pi_{f})=a\,\mathrm{Ent}(\pi_{f})+\int\bar{U}\,d\pi_{f}=\int\bar{f}\,d\pi_{f}-a\log Z. This is clause 1.

Step 7 (Clause 2). Let μ∈DU,a\mu\in\mathcal{D}_{U,a} with fˉ\bar{f} integrable with respect to μ\mu. By The Entropy of a Probability Measure on Euclidean Space §entropy, μ\mu has a density qq with respect to λd\lambda_{d} (a Borel q:Rd→[0,∞)q:\mathbb{R}^{d}\to[0,\infty) with μ(B)=∫1B q dλd\mu(B)=\int\mathbf{1}_{B}\,q\,d\lambda_{d} for all Borel BB, The Radon-Nikodym Theorem for a Finite Measure and a Sigma-Finite Measure, and Uniqueness of Densities) such that ϕ∘q\phi\circ q is integrable with respect to λd\lambda_{d} and Ent(μ)=∫ϕ∘q dλd\mathrm{Ent}(\mu)=\int\phi\circ q\,d\lambda_{d}; so μ\mu is the measure with density qq of claim 3 of Image Measures, Measures with Densities, and Change of Variables, and ∫q dλd=μ(Rd)=1\int q\,d\lambda_{d}=\mu(\mathbb{R}^{d})=1. Since μ(D)=1\mu(D)=1, ∫1Rd∖D q dλd=μ(Rd∖D)=0\int\mathbf{1}_{\mathbb{R}^{d}\setminus D}\,q\,d\lambda_{d}=\mu(\mathbb{R}^{d}\setminus D)=0, so the set EE of x∉Dx\notin D with q(x)≠0q(x)\ne0 is null by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §vanishing. The function ℓˉ\bar{\ell} of Step 6 is integrable with respect to μ\mu, being a linear combination of fˉ\bar{f}, Uˉ\bar{U} (integrable as μ∈DU,a\mu\in\mathcal{D}_{U,a}) and the bounded Borel function 1D\mathbf{1}_{D} (Probability Measures on Euclidean Space and Random Vectors: Standing Notation §measures), with ∫ℓˉ dμ=a−1(∫fˉ dμ−∫Uˉ dμ)−log⁡Z\int\bar{\ell}\,d\mu=a^{-1}\bigl(\int\bar{f}\,d\mu-\int\bar{U}\,d\mu\bigr)-\log Z; by claim 3 of Image Measures, Measures with Densities, and Change of Variables, qℓˉq\bar{\ell} is integrable with respect to λd\lambda_{d} with ∫qℓˉ dλd=∫ℓˉ dμ\int q\bar{\ell}\,d\lambda_{d}=\int\bar{\ell}\,d\mu.

We claim that for every x∈Rd∖Ex\in\mathbb{R}^{d}\setminus E,

q(x) ℓˉ(x)−ϕ(q(x))≤p(x)−q(x).(7.1)q(x)\,\bar{\ell}(x)-\phi(q(x))\le p(x)-q(x).\tag{7.1}

If x∉Dx\notin D, then q(x)=0=p(x)q(x)=0=p(x) and both sides vanish. If x∈Dx\in D and q(x)=0q(x)=0, the left side is 0<p(x)0<p(x). If x∈Dx\in D and q(x)>0q(x)>0, put v=ℓˉ(x)−log⁡q(x)v=\bar{\ell}(x)-\log q(x); by Step 6 and claims 1 and 2 of Basic Properties of the Exponential Function, exp⁡(v)=exp⁡(ℓˉ(x))exp⁡(log⁡q(x))−1=p(x)/q(x)\exp(v)=\exp(\bar{\ell}(x))\exp(\log q(x))^{-1}=p(x)/q(x), and The Function slog⁡ss\log s: Continuity, Young's Inequality and Lower Bounds, with the Elementary Bounds for the Exponential and the Logarithm §exp gives 1+v≤p(x)/q(x)1+v\le p(x)/q(x); multiplying by q(x)>0q(x)>0 and using ϕ(q(x))=q(x)log⁡q(x)\phi(q(x))=q(x)\log q(x) gives (7.1).

Let w=p−q−qℓˉ+ϕ∘qw=p-q-q\bar{\ell}+\phi\circ q, integrable with respect to λd\lambda_{d} by claim 2 of Linearity and Monotonicity of the Lebesgue Integral. By (7.1), w=max⁡(w,0)w=\max(w,0) outside the null set EE, so by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §comparison and claim 2 of Linearity and Monotonicity of the Lebesgue Integral, ∫w dλd=∫max⁡(w,0) dλd≥0\int w\,d\lambda_{d}=\int\max(w,0)\,d\lambda_{d}\ge0. By linearity, 0≤1−1−∫ℓˉ dμ+Ent(μ)0\le1-1-\int\bar{\ell}\,d\mu+\mathrm{Ent}(\mu), that is,

a−1(∫fˉ dμ−∫Uˉ dμ)−log⁡Z≤Ent(μ),a^{-1}\Bigl(\int\bar{f}\,d\mu-\int\bar{U}\,d\mu\Bigr)-\log Z\le\mathrm{Ent}(\mu),

and multiplying by a>0a>0 gives ∫fˉ dμ−(a Ent(μ)+∫Uˉ dμ)≤alog⁡Z\int\bar{f}\,d\mu-\bigl(a\,\mathrm{Ent}(\mu)+\int\bar{U}\,d\mu\bigr)\le a\log Z, which is clause 2; the maximisation statement follows from clause 1.

Step 8 (Square-integrable gradients). For the rest of the proof fix a gradient map ∇f\nabla f of ff, with a null set NN as in the statement, and a Borel set N1⊇NN_{1}\supseteq N with λd(N1)=0\lambda_{d}(N_{1})=0. Then πf(N1)=∫1N1 p dλd=0\pi_{f}(N_{1})=\int\mathbf{1}_{N_{1}}\,p\,d\lambda_{d}=0 by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §null-integral, so the Borel set D∖N1D\setminus N_{1} has πf(D∖N1)=1\pi_{f}(D\setminus N_{1})=1. At every x∈D∖N1x\in D\setminus N_{1}, ff is differentiable and ∇f(x)=Df(x)\nabla f(x)=Df(x), so by hypothesis ∥∇f(x)∥2≤A+B(U(x)−p0)≤(A+B+B∣p0∣)(1+∣U(x)∣)\lVert\nabla f(x)\rVert^{2}\le A+B(U(x)-p_{0})\le(A+B+B|p_{0}|)(1+|U(x)|). By (IC) with c=A+B+B∣p0∣c=A+B+B|p_{0}|, and with c=1c=1, the Borel functions ∥∇f∥2\lVert\nabla f\rVert^{2} and ∥∇U∥2\lVert\nabla U\rVert^{2} are integrable with respect to πf\pi_{f} (∇U=DU\nabla U=DU on DD). Let

η=a−1(∇f−∇U),Tg=∇f+K id,\eta=a^{-1}\bigl(\nabla f-\nabla U\bigr),\qquad T_{g}=\nabla f+K\,\mathrm{id},

Borel maps Rd→Rd\mathbb{R}^{d}\to\mathbb{R}^{d} with ∥η∥2≤2a−2(∥∇f∥2+∥∇U∥2)\lVert\eta\rVert^{2}\le2a^{-2}(\lVert\nabla f\rVert^{2}+\lVert\nabla U\rVert^{2}) and ∥Tg∥2≤2∥∇f∥2+2K2∥id∥2\lVert T_{g}\rVert^{2}\le2\lVert\nabla f\rVert^{2}+2K^{2}\lVert\mathrm{id}\rVert^{2}; by Step 6 and claim 1 of Linearity and Monotonicity of the Lebesgue Integral, all of ∥∇f∥2\lVert\nabla f\rVert^{2}, ∥∇U∥2\lVert\nabla U\rVert^{2}, ∥η∥2\lVert\eta\rVert^{2}, ∥Tg∥2\lVert T_{g}\rVert^{2} have finite πf\pi_{f}-integrals. So the classes of ∇f\nabla f, ∇U\nabla U, η\eta, TgT_{g} and id\mathrm{id} lie in L2(πf;Rd)L^{2}(\pi_{f};\mathbb{R}^{d}), where η=a−1(∇f−∇U)\eta=a^{-1}(\nabla f-\nabla U) and Tg=∇f+K idT_{g}=\nabla f+K\,\mathrm{id} as classes, the vector operations on classes being computed on representatives. Moreover p(1+∥η∥2)p(1+\lVert\eta\rVert^{2}) and p(1+∥∇U∥2)p(1+\lVert\nabla U\rVert^{2}) are integrable with respect to λd\lambda_{d} by claim 3 of Image Measures, Measures with Densities, and Change of Variables, and so are p∥η∥p\lVert\eta\rVert and p∥∇U∥p\lVert\nabla U\rVert, which they dominate.

Step 9 (Uniform Lipschitz bounds near compact subsets of DD). Let S⊆DS\subseteq D be compact. There are real r∈(0,1]r\in(0,1] and L≥0L\ge0 such that for all x∈Sx\in S and y∈Rdy\in\mathbb{R}^{d} with ∥y−x∥≤r\lVert y-x\rVert\le r: y∈Dy\in D, ∣f(y)−f(x)∣≤L∥y−x∥|f(y)-f(x)|\le L\lVert y-x\rVert and ∣U(y)−U(x)∣≤L∥y−x∥|U(y)-U(x)|\le L\lVert y-x\rVert. If S=∅S=\varnothing take r=1r=1, L=0L=0. Otherwise, for z∈Sz\in S apply the local Lipschitz bound to ff (semiconvex with constant KK) and to UU (semiconvex with constant 00, Step 2(a)) at zz, and let ρz>0\rho_{z}>0 be the smaller of the two radii and LzL_{z} the larger of the two constants; then Bˉ(z,ρz)⊆D\bar{B}(z,\rho_{z})\subseteq D and ff, UU are both Lipschitz with constant LzL_{z} on Bˉ(z,ρz)\bar{B}(z,\rho_{z}). The open balls B(z,ρz/2)B(z,\rho_{z}/2), z∈Sz\in S, cover the compact set SS, so finitely many of them, with centres z1,…,zkz_{1},\dots,z_{k}, cover SS. Put r=min⁡(1,ρz1/2,…,ρzk/2)r=\min(1,\rho_{z_{1}}/2,\dots,\rho_{z_{k}}/2) and L=max⁡(Lz1,…,Lzk)L=\max(L_{z_{1}},\dots,L_{z_{k}}). For x∈Sx\in S pick jj with ∥x−zj∥<ρzj/2\lVert x-z_{j}\rVert<\rho_{z_{j}}/2; if ∥y−x∥≤r\lVert y-x\rVert\le r then ∥y−zj∥<ρzj\lVert y-z_{j}\rVert<\rho_{z_{j}} by the triangle inequality, so x,y∈Bˉ(zj,ρzj)⊆Dx,y\in\bar{B}(z_{j},\rho_{z_{j}})\subseteq D and the two bounds hold.

Step 10 (Integration by parts against a cutoff of the density). Fix ψ∈Cc∞(Rd)\psi\in C_{c}^{\infty}(\mathbb{R}^{d}) (test functions). By The Gradient of a Test Function is Bounded and Square-Integrable, and Its Laplacian Bounded and Integrable, Against Every Probability Measure §gradient and The Gradient of a Test Function is Bounded and Square-Integrable, and Its Laplacian Bounded and Integrable, Against Every Probability Measure §laplacian there are Kψ,CΔ≥0K_{\psi},C_{\Delta}\ge0 with ∥∇ψ∥≤Kψ\lVert\nabla\psi\rVert\le K_{\psi} and ∣Δψ∣≤CΔ|\Delta\psi|\le C_{\Delta} everywhere, Δψ=∑i∂i∂iψ\Delta\psi=\sum_{i}\partial_{i}\partial_{i}\psi being Borel. For each i∈[d]i\in[d], ∂iψ\partial_{i}\psi is continuous and λd\lambda_{d}-integrable by Integration by Parts on Euclidean Space Against a Compactly Supported Function of Class C1C^{1} §vanishing (with h=ψh=\psi); and ∂iψ\partial_{i}\psi is of class C1C^{1} (ψ\psi being of class C2C^{2}, Test Functions on Euclidean Space, Their Gradient Maps and Laplacians §gradient) and compactly supported (The Gradient of a Test Function is Bounded and Square-Integrable, and Its Laplacian Bounded and Integrable, Against Every Probability Measure §gradient), so by Integration by Parts on Euclidean Space Against a Compactly Supported Function of Class C1C^{1} §vanishing (with h=∂iψh=\partial_{i}\psi) ∂i∂iψ\partial_{i}\partial_{i}\psi is continuous and compactly supported, hence bounded by Extreme Value Theorem on a Compact Subset of a Metric Space applied on its support; fix C2≥0C_{2}\ge0 with ∣∂i∂iψ∣≤C2|\partial_{i}\partial_{i}\psi|\le C_{2} for every ii. For x∈Rdx\in\mathbb{R}^{d} and t>0t>0, the function τ↦∂iψ(x+τei)\tau\mapsto\partial_{i}\psi(x+\tau e_{i}) is continuous on [0,t][0,t] and, by A Real-Valued C^1 Function is Differentiable at Every Point and Chain Rule Along an Affine Path, differentiable with derivative ∂i∂iψ(x+τei)\partial_{i}\partial_{i}\psi(x+\tau e_{i}); so Mean Value Theorem on a Closed Real Interval gives

∣∂iψ(x+tei)−∂iψ(x)∣≤C2 t.(10.1)\bigl|\partial_{i}\psi(x+te_{i})-\partial_{i}\psi(x)\bigr|\le C_{2}\,t .\tag{10.1}

Fix a smooth χ:R→R\chi:\mathbb{R}\to\mathbb{R} as in Scaled Cutoffs and the Second-Moment Test Functions: Uniform Derivative Bounds and Agreement on a Ball with q=1q=1 (identifying R1\mathbb{R}^{1} with R\mathbb{R}, so that ∥t∥=∣t∣\lVert t\rVert=|t| and ∂1\partial_{1} is the ordinary derivative, written with a prime); it exists by Existence of a Smooth Plateau Function on Euclidean Space. By Scaled Cutoffs and the Second-Moment Test Functions: Uniform Derivative Bounds and Agreement on a Ball §cutoff there is M1≥0M_{1}\ge0 such that for every natural M≥1M\ge1 the function χM(t)=χ(t/M)\chi_{M}(t)=\chi(t/M) is smooth, 0≤χM≤10\le\chi_{M}\le1, χM(t)=1\chi_{M}(t)=1 for ∣t∣≤M|t|\le M, χM(t)=0\chi_{M}(t)=0 for ∣t∣≥2M|t|\ge2M, and ∣χM′∣≤M1/M|\chi_{M}'|\le M_{1}/M; by Mean Value Theorem on a Closed Real Interval, ∣χM(s)−χM(t)∣≤(M1/M)∣s−t∣|\chi_{M}(s)-\chi_{M}(t)|\le(M_{1}/M)|s-t| for all real s,ts,t. Also, for real s,us,u,

∣es−eu∣≤emax⁡(s,u)∣s−u∣,(10.2)|e^{s}-e^{u}|\le e^{\max(s,u)}|s-u| ,\tag{10.2}

since for u≤su\le s, The Function slog⁡ss\log s: Continuity, Young's Inequality and Lower Bounds, with the Elementary Bounds for the Exponential and the Logarithm §exp gives eu−s≥1+(u−s)e^{u-s}\ge1+(u-s), whence 0≤es−eu=es(1−eu−s)≤es(s−u)0\le e^{s}-e^{u}=e^{s}(1-e^{u-s})\le e^{s}(s-u) by claims 1 and 4 of Basic Properties of the Exponential Function.

Fix a natural M≥1M\ge1. Define ζM,κM:Rd→R\zeta_{M},\kappa_{M}:\mathbb{R}^{d}\to\mathbb{R} by ζM(x)=χM(U(x)−p0)\zeta_{M}(x)=\chi_{M}(U(x)-p_{0}) and κM(x)=χM′(U(x)−p0)\kappa_{M}(x)=\chi_{M}'(U(x)-p_{0}) for x∈Dx\in D, and ζM(x)=κM(x)=0\zeta_{M}(x)=\kappa_{M}(x)=0 for x∉Dx\notin D; they are Borel by (BC), χM\chi_{M} and χM′\chi_{M}' being continuous. Put

wM=ζM p,VM=p (ζM η+κM ∇U),w_{M}=\zeta_{M}\,p,\qquad V_{M}=p\,\bigl(\zeta_{M}\,\eta+\kappa_{M}\,\nabla U\bigr),

Borel, with 0≤wM≤p≤Z−1eH0\le w_{M}\le p\le Z^{-1}e^{H} by (5.1). Let SM={x∈D:U(x)≤p0+2M}S_{M}=\{x\in D:U(x)\le p_{0}+2M\}, compact by Penalty on an Open Subset of Euclidean Space §sublevel; since ζM(x)=0\zeta_{M}(x)=0 when x∈Dx\in D and U(x)−p0≥2MU(x)-p_{0}\ge2M, wMw_{M} vanishes off SMS_{M}. Let CMC_{M} be the closure of {x∈D:U(x)<p0+2M+1}⊇SM\{x\in D:U(x)<p_{0}+2M+1\}\supseteq S_{M}; by Basic Properties of the Sublevel Sets of a Penalty §sublevel-sets, CM⊆DC_{M}\subseteq D, and CMC_{M} is closed. Apply Step 9 to S=SMS=S_{M}, obtaining r∈(0,1]r\in(0,1] and LL, and put LM=Z−1eHL (2/a+M1/M)L_{M}=Z^{-1}e^{H}L\,(2/a+M_{1}/M).

(10a) For i∈[d]i\in[d], t∈(0,r]t\in(0,r] and x∈Rdx\in\mathbb{R}^{d}: ∣wM(x−tei)−wM(x)∣≤LM t|w_{M}(x-te_{i})-w_{M}(x)|\le L_{M}\,t. First let x′∈SMx'\in S_{M} and ∥y′−x′∥≤r\lVert y'-x'\rVert\le r; then x′,y′∈Dx',y'\in D and wM(y′)−wM(x′)=ζM(y′)(p(y′)−p(x′))+p(x′)(ζM(y′)−ζM(x′))w_{M}(y')-w_{M}(x')=\zeta_{M}(y')\bigl(p(y')-p(x')\bigr)+p(x')\bigl(\zeta_{M}(y')-\zeta_{M}(x')\bigr). By (5.1), p=Z−1ehp=Z^{-1}e^{h} on DD, so (10.2), h≤Hh\le H and Step 9 give ∣p(y′)−p(x′)∣≤Z−1eHa−1(∣f(y′)−f(x′)∣+∣U(y′)−U(x′)∣)≤Z−1eH(2L/a)∥y′−x′∥|p(y')-p(x')|\le Z^{-1}e^{H}a^{-1}\bigl(|f(y')-f(x')|+|U(y')-U(x')|\bigr)\le Z^{-1}e^{H}(2L/a)\lVert y'-x'\rVert, while ∣ζM(y′)−ζM(x′)∣≤(M1/M)∣U(y′)−U(x′)∣≤(M1L/M)∥y′−x′∥|\zeta_{M}(y')-\zeta_{M}(x')|\le(M_{1}/M)|U(y')-U(x')|\le(M_{1}L/M)\lVert y'-x'\rVert. With 0≤ζM≤10\le\zeta_{M}\le1 and p(x′)≤Z−1eHp(x')\le Z^{-1}e^{H} this gives ∣wM(y′)−wM(x′)∣≤LM∥y′−x′∥|w_{M}(y')-w_{M}(x')|\le L_{M}\lVert y'-x'\rVert. Now let y=x−teiy=x-te_{i}, so ∥y−x∥=t≤r\lVert y-x\rVert=t\le r: if x∈SMx\in S_{M} use (x′,y′)=(x,y)(x',y')=(x,y); if y∈SMy\in S_{M} use (x′,y′)=(y,x)(x',y')=(y,x); otherwise wM(x)=wM(y)=0w_{M}(x)=w_{M}(y)=0.

(10b) For x∈D∖N1x\in D\setminus N_{1} and i∈[d]i\in[d] the partial derivative ∂iwM(x)\partial_{i}w_{M}(x) exists and equals the iith component (VM)i(x)(V_{M})_{i}(x). Let Θ:R2→R\Theta:\mathbb{R}^{2}\to\mathbb{R}, Θ(s,u)=Z−1χM(u−p0)exp⁡((s−u)/a)\Theta(s,u)=Z^{-1}\chi_{M}(u-p_{0})\exp\bigl((s-u)/a\bigr). By Sum, Constant Multiple, and Product Rules for One-Dimensional Derivatives, Chain Rule for One-Dimensional Derivatives and claim 3 of Basic Properties of the Exponential Function, its partial derivatives are ∂sΘ=a−1Θ\partial_{s}\Theta=a^{-1}\Theta and ∂uΘ(s,u)=Z−1χM′(u−p0)exp⁡((s−u)/a)−a−1Θ(s,u)\partial_{u}\Theta(s,u)=Z^{-1}\chi_{M}'(u-p_{0})\exp((s-u)/a)-a^{-1}\Theta(s,u), which are continuous; so Θ\Theta is of class C1C^{1} (C^k Maps on a Euclidean Open Set) and, by A Real-Valued C^1 Function is Differentiable at Every Point, differentiable at every point with derivative matrix (∂sΘ,∂uΘ)(\partial_{s}\Theta,\partial_{u}\Theta). The map Φ=(f,U):D→R2\Phi=(f,U):D\to\mathbb{R}^{2} is differentiable at xx with derivative matrix having rows Df(x)Df(x) and DU(x)DU(x): ff is differentiable at xx because x∉Nx\notin N, UU by A Real-Valued C^1 Function is Differentiable at Every Point, and for a given ε\varepsilon one takes the smaller of the two δ\delta's for ε/2\varepsilon/2, using ∥(α,β)∥≤∣α∣+∣β∣\lVert(\alpha,\beta)\rVert\le|\alpha|+|\beta|. On DD we have wM=Θ∘Φw_{M}=\Theta\circ\Phi, since Θ(Φ(y))=Z−1χM(U(y)−p0)eh(y)=ζM(y)p(y)\Theta(\Phi(y))=Z^{-1}\chi_{M}(U(y)-p_{0})e^{h(y)}=\zeta_{M}(y)p(y). By Chain Rule for Differentiable Maps Between Euclidean Spaces, Θ∘Φ\Theta\circ\Phi is differentiable at xx with derivative matrix ∂sΘ(Φ(x)) Df(x)+∂uΘ(Φ(x)) DU(x)\partial_{s}\Theta(\Phi(x))\,Df(x)+\partial_{u}\Theta(\Phi(x))\,DU(x), and by claim 1 of A Derivative Matrix is the Jacobian Matrix, and is Unique (and since DD is open, so that wMw_{M} and Θ∘Φ\Theta\circ\Phi have the same difference quotients at xx for small steps)

∂iwM(x)=wM(x)a(∂if(x)−∂iU(x))+p(x)κM(x) ∂iU(x)=p(x)(ζM(x)ηi(x)+κM(x)(∇U)i(x)),\partial_{i}w_{M}(x)=\frac{w_{M}(x)}{a}\bigl(\partial_{i}f(x)-\partial_{i}U(x)\bigr)+p(x)\kappa_{M}(x)\,\partial_{i}U(x)=p(x)\bigl(\zeta_{M}(x)\eta_{i}(x)+\kappa_{M}(x)(\nabla U)_{i}(x)\bigr),

using ∇f(x)=Df(x)\nabla f(x)=Df(x) and ∇U(x)=DU(x)\nabla U(x)=DU(x).

(10c) For i∈[d]i\in[d] and t>0t>0,

∫wM(x)(∂iψ(x+tei)−∂iψ(x)) dλd(x)=∫(wM(y−tei)−wM(y))∂iψ(y) dλd(y),\int w_{M}(x)\bigl(\partial_{i}\psi(x+te_{i})-\partial_{i}\psi(x)\bigr)\,d\lambda_{d}(x)=\int\bigl(w_{M}(y-te_{i})-w_{M}(y)\bigr)\partial_{i}\psi(y)\,d\lambda_{d}(y),

all four integrands involved being λd\lambda_{d}-integrable. Let T(x)=x+teiT(x)=x+te_{i}. Its components are of class C1C^{1} with Jacobian matrix the identity matrix IdI_{d} at every point, which is symmetric and positive definite (v⋅Idv=∥v∥2>0v\cdot I_{d}v=\lVert v\rVert^{2}>0 for v≠0v\ne0); TT is a bijection with inverse T−1(y)=y−teiT^{-1}(y)=y-te_{i}, whose Jacobian matrix IdI_{d} is the inverse matrix of IdI_{d} by The Identity Matrix is a Two-Sided Multiplicative Identity; and det⁡Id=1\det I_{d}=1 by The Determinant of a Triangular Matrix is the Product of its Diagonal Entries. The function Q(y)=wM(y−tei) ∂iψ(y)Q(y)=w_{M}(y-te_{i})\,\partial_{i}\psi(y) is Borel (T−1T^{-1} being Borel by Change of Variables for the Lebesgue Integral under a Continuously Differentiable Bijection of Euclidean Space with Symmetric Positive Definite Jacobian Matrix, and the Density of a Push-Forward §regularity) with ∣Q∣≤Z−1eH∣∂iψ∣|Q|\le Z^{-1}e^{H}|\partial_{i}\psi|, hence integrable, and Q∘T(x)=wM(x)∂iψ(x+tei)Q\circ T(x)=w_{M}(x)\partial_{i}\psi(x+te_{i}). By Change of Variables for the Lebesgue Integral under a Continuously Differentiable Bijection of Euclidean Space with Symmetric Positive Definite Jacobian Matrix, and the Density of a Push-Forward §integrals, Q∘TQ\circ T is integrable and ∫wM(x)∂iψ(x+tei) dλd(x)=∫Q dλd\int w_{M}(x)\partial_{i}\psi(x+te_{i})\,d\lambda_{d}(x)=\int Q\,d\lambda_{d}. Subtracting the integral of the integrable function wM∂iψw_{M}\partial_{i}\psi (bounded by Z−1eH∣∂iψ∣Z^{-1}e^{H}|\partial_{i}\psi|) from both sides gives the identity, by claim 2 of Linearity and Monotonicity of the Lebesgue Integral.

(10d) For i∈[d]i\in[d]: ∫wM ∂i∂iψ dλd=−∫(VM)i ∂iψ dλd\int w_{M}\,\partial_{i}\partial_{i}\psi\,d\lambda_{d}=-\int(V_{M})_{i}\,\partial_{i}\psi\,d\lambda_{d}. Fix a natural m0≥1/rm_{0}\ge1/r, and for natural m≥m0m\ge m_{0} put

Lm(x)=m wM(x)(∂iψ(x+m−1ei)−∂iψ(x)),Rm(y)=m(wM(y−m−1ei)−wM(y))∂iψ(y),\mathrm{L}_{m}(x)=m\,w_{M}(x)\bigl(\partial_{i}\psi(x+m^{-1}e_{i})-\partial_{i}\psi(x)\bigr),\qquad \mathrm{R}_{m}(y)=m\bigl(w_{M}(y-m^{-1}e_{i})-w_{M}(y)\bigr)\partial_{i}\psi(y),

Borel functions with ∫Lm dλd=∫Rm dλd\int\mathrm{L}_{m}\,d\lambda_{d}=\int\mathrm{R}_{m}\,d\lambda_{d} by (10c) with t=1/mt=1/m. At every xx, Lm(x)→wM(x)∂i∂iψ(x)\mathrm{L}_{m}(x)\to w_{M}(x)\partial_{i}\partial_{i}\psi(x) by Partial Derivative on a Euclidean Open Set, and ∣Lm∣≤C2 p|\mathrm{L}_{m}|\le C_{2}\,p by (10.1), pp being integrable; so Dominated Convergence Theorem gives ∫Lm dλd→∫wM∂i∂iψ dλd\int\mathrm{L}_{m}\,d\lambda_{d}\to\int w_{M}\partial_{i}\partial_{i}\psi\,d\lambda_{d}. By (10a) with t=1/m≤rt=1/m\le r, ∣Rm∣≤LM∣∂iψ∣|\mathrm{R}_{m}|\le L_{M}|\partial_{i}\psi|, an integrable function. For y∈D∖N1y\in D\setminus N_{1}, Rm(y)=−wM(y+(−m−1)ei)−wM(y)−m−1 ∂iψ(y)→−(VM)i(y)∂iψ(y)\mathrm{R}_{m}(y)=-\frac{w_{M}(y+(-m^{-1})e_{i})-w_{M}(y)}{-m^{-1}}\,\partial_{i}\psi(y)\to-(V_{M})_{i}(y)\partial_{i}\psi(y) by (10b). For y∉Dy\notin D, y∉CMy\notin C_{M}; as Rd∖CM\mathbb{R}^{d}\setminus C_{M} is open there is ρ>0\rho>0 with B(y,ρ)∩CM=∅B(y,\rho)\cap C_{M}=\varnothing, so for m>1/ρm>1/\rho neither yy nor y−m−1eiy-m^{-1}e_{i} lies in SM⊆CMS_{M}\subseteq C_{M} and Rm(y)=0=−(VM)i(y)∂iψ(y)\mathrm{R}_{m}(y)=0=-(V_{M})_{i}(y)\partial_{i}\psi(y), as p(y)=0p(y)=0. So Rm→−(VM)i∂iψ\mathrm{R}_{m}\to-(V_{M})_{i}\partial_{i}\psi outside the null set N1N_{1}, and The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §dominated gives ∫Rm dλd→−∫(VM)i∂iψ dλd\int\mathrm{R}_{m}\,d\lambda_{d}\to-\int(V_{M})_{i}\partial_{i}\psi\,d\lambda_{d}. The two limits coincide.

Summing (10d) over i∈[d]i\in[d] and using Δψ=∑i∂i∂iψ\Delta\psi=\sum_{i}\partial_{i}\partial_{i}\psi and claim 2 of Linearity and Monotonicity of the Lebesgue Integral,

∫wM Δψ dλd=−∫VM⋅∇ψ dλd(M≥1).(10.3)\int w_{M}\,\Delta\psi\,d\lambda_{d}=-\int V_{M}\cdot\nabla\psi\,d\lambda_{d}\qquad(M\ge1).\tag{10.3}

Step 11 (Removing the cutoff). Let M→∞M\to\infty through the natural numbers. For x∈Dx\in D and M≥U(x)−p0M\ge U(x)-p_{0} we have ζM(x)=1\zeta_{M}(x)=1, while ∣κM(x)∣≤M1/M→0|\kappa_{M}(x)|\le M_{1}/M\to0; for x∉Dx\notin D all of wM(x)w_{M}(x), VM(x)V_{M}(x), p(x)p(x) vanish. Hence wMΔψ→p Δψw_{M}\Delta\psi\to p\,\Delta\psi and VM⋅∇ψ→p η⋅∇ψV_{M}\cdot\nabla\psi\to p\,\eta\cdot\nabla\psi at every point, with ∣wMΔψ∣≤CΔ p|w_{M}\Delta\psi|\le C_{\Delta}\,p and ∣VM⋅∇ψ∣≤Kψ p (∥η∥+M1∥∇U∥)|V_{M}\cdot\nabla\psi|\le K_{\psi}\,p\,\bigl(\lVert\eta\rVert+M_{1}\lVert\nabla U\rVert\bigr), integrable by Step 8. By Dominated Convergence Theorem and (10.3), ∫p Δψ dλd=−∫p η⋅∇ψ dλd\int p\,\Delta\psi\,d\lambda_{d}=-\int p\,\eta\cdot\nabla\psi\,d\lambda_{d}. As Δψ\Delta\psi is bounded Borel and ∣η⋅∇ψ∣≤Kψ∥η∥|\eta\cdot\nabla\psi|\le K_{\psi}\lVert\eta\rVert, both Δψ\Delta\psi and η⋅∇ψ\eta\cdot\nabla\psi are πf\pi_{f}-integrable and claim 3 of Image Measures, Measures with Densities, and Change of Variables turns this into

⟨η,∇ψ⟩πf=∫η⋅∇ψ dπf=−∫Δψ dπffor every ψ∈Cc∞(Rd).(11.1)\langle\eta,\nabla\psi\rangle_{\pi_{f}}=\int\eta\cdot\nabla\psi\,d\pi_{f}=-\int\Delta\psi\,d\pi_{f}\qquad\text{for every }\psi\in C_{c}^{\infty}(\mathbb{R}^{d}).\tag{11.1}

Step 12 (Tangency). Let TπfT_{\pi_{f}} be the tangent space at πf∈P2(Rd)\pi_{f}\in\mathcal{P}_{2}(\mathbb{R}^{d}), a linear subspace of L2(πf;Rd)L^{2}(\pi_{f};\mathbb{R}^{d}) by Basic Properties of the Tangent Space: Closed Subspace, the Identity Map Belongs to It, Second-Moment Limits, and Representation of Bounded Functionals on Gradients §closed. Apply A Square-Integrable Selection of the Subdifferential of a Convex Potential Belongs to the Tangent Space §tangent with G=DG=D (open, convex), the convex function UU, the Borel set DD of πf\pi_{f}-measure 11 and T=∇UT=\nabla U: for every x∈Dx\in D, ∂DU(x)={DU(x)}={∇U(x)}\partial_{D}U(x)=\{DU(x)\}=\{\nabla U(x)\} by Step 2(b), and ∫∥∇U∥2 dπf<∞\int\lVert\nabla U\rVert^{2}\,d\pi_{f}<\infty; so ∇U∈Tπf\nabla U\in T_{\pi_{f}}. Apply it again with G=DG=D, the convex function gg of Step 2(c), the Borel set D∖N1D\setminus N_{1} of πf\pi_{f}-measure 11 (Step 8) and T=TgT=T_{g}: for every x∈D∖N1x\in D\setminus N_{1}, ∂Dg(x)={Df(x)+Kx}={Tg(x)}\partial_{D}g(x)=\{Df(x)+Kx\}=\{T_{g}(x)\} by Step 2(c), and ∫∥Tg∥2 dπf<∞\int\lVert T_{g}\rVert^{2}\,d\pi_{f}<\infty; so Tg∈TπfT_{g}\in T_{\pi_{f}}. By Basic Properties of the Tangent Space: Closed Subspace, the Identity Map Belongs to It, Second-Moment Limits, and Representation of Bounded Functionals on Gradients §identity, id∈Tπf\mathrm{id}\in T_{\pi_{f}}. Since TπfT_{\pi_{f}} is a linear subspace, ∇f=Tg−K id∈Tπf\nabla f=T_{g}-K\,\mathrm{id}\in T_{\pi_{f}} and η=a−1(∇f−∇U)∈Tπf\eta=a^{-1}(\nabla f-\nabla U)\in T_{\pi_{f}}.

Step 13 (Clause 3). Gradient maps of ff exist by Step 3. By Step 6, πf∈DU,a⊆P2(Rd)\pi_{f}\in\mathcal{D}_{U,a}\subseteq\mathcal{P}_{2}(\mathbb{R}^{d}). For ψ∈Cc∞(Rd)\psi\in C_{c}^{\infty}(\mathbb{R}^{d}), (11.1) and The Cauchy-Schwarz Inequality in a Real Inner Product Space in the real Hilbert space L2(πf;Rd)L^{2}(\pi_{f};\mathbb{R}^{d}) (Square-Integrable Vector Fields Against a Probability Measure on Euclidean Space, and Test Functions: Standing Notation §l2mu) give ∣∫Δψ dπf∣≤∥η∥πf∥∇ψ∥πf\bigl|\int\Delta\psi\,d\pi_{f}\bigr|\le\lVert\eta\rVert_{\pi_{f}}\lVert\nabla\psi\rVert_{\pi_{f}}; so πf\pi_{f} has finite Fisher information with C=∥η∥πfC=\lVert\eta\rVert_{\pi_{f}}, i.e. πf∈P2I(Rd)\pi_{f}\in\mathcal{P}_{2}^{\mathcal{I}}(\mathbb{R}^{d}). By Step 12, η∈Tπf\eta\in T_{\pi_{f}}, and by (11.1) it satisfies the identity defining the score; by the uniqueness asserted there, ξπf=η=a−1(∇f−∇U)\xi_{\pi_{f}}=\eta=a^{-1}(\nabla f-\nabla U). Since moreover ∫∥∇U∥2 dπf<∞\int\lVert\nabla U\rVert^{2}\,d\pi_{f}<\infty (Step 8), πf∈DU,aΣ\pi_{f}\in\mathcal{D}^{\Sigma}_{U,a} by The Relative Free Energy and the Relative Score of a Probability Measure for a Potential on an Open Set §score, and in L2(πf;Rd)L^{2}(\pi_{f};\mathbb{R}^{d})

ΣU,a(πf)=∇U+a ξπf=∇U+(∇f−∇U)=∇f.\Sigma_{U,a}(\pi_{f})=\nabla U+a\,\xi_{\pi_{f}}=\nabla U+(\nabla f-\nabla U)=\nabla f .

Finally ∫∥∇f∥2 dπf<∞\int\lVert\nabla f\rVert^{2}\,d\pi_{f}<\infty by Step 8. As the gradient map ∇f\nabla f was arbitrary, clause 3 is proved.

Please log in to copy this version.

Citations

Loading…

Dependency Graph

0 prerequisites

Comments

Loading…