Each result cited below is universally quantified over the data in its own statement, and is applied to the data named at the point of use. Elementary arithmetic and order manipulations of real numbers, such as rearranging terms, multiplying an inequality by a positive number, and the bounds s ≤ 1 + s 2 s\le1+s^{2} s ≤ 1 + s 2 , ( s + t ) 2 ≤ 2 s 2 + 2 t 2 (s+t)^{2}\le2s^{2}+2t^{2} ( s + t ) 2 ≤ 2 s 2 + 2 t 2 for real s , t s,t s , t , are used without comment, by Elementary Order Arithmetic in an Ordered Field and Elementary Arithmetic in an Ordered Field . Throughout, e 1 , … , e d e_{1},\dots,e_{d} e 1 , … , e d are the standard basis vectors of R d \mathbb{R}^{d} R d , i d \mathrm{id} id is the identity map of R d \mathbb{R}^{d} R d , F = sup x ∈ D f ( x ) F=\sup_{x\in D}f(x) F = sup x ∈ D f ( x ) (a real number, D D D being nonempty and f f f bounded above), H = ( F − p 0 ) / a H=(F-p_{0})/a H = ( F − p 0 ) / a , and h ( x ) = ( f ( x ) − U ( x ) ) / a h(x)=(f(x)-U(x))/a h ( x ) = ( f ( x ) − U ( x )) / a for x ∈ D x\in D x ∈ D , so that h ( x ) ≤ H h(x)\le H h ( x ) ≤ H for every x ∈ D x\in D x ∈ D since U ( x ) ≥ p 0 U(x)\ge p_{0} U ( x ) ≥ p 0 . Norms and dot products on R d \mathbb{R}^{d} R d are handled with claims 1, 5 and 6 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n (in particular ∥ x + v ∥ 2 = ∥ x ∥ 2 + 2 x ⋅ v + ∥ v ∥ 2 \lVert x+v\rVert^{2}=\lVert x\rVert^{2}+2x\cdot v+\lVert v\rVert^{2} ∥ x + v ∥ 2 = ∥ x ∥ 2 + 2 x ⋅ v + ∥ v ∥ 2 , by expanding the sum of squares of coordinates) and with Cauchy-Schwarz Inequality for the Euclidean Dot Product .
Step 1 (Borel functions; D D D is not null). (BC) Let O ⊆ R d O\subseteq\mathbb{R}^{d} O ⊆ R d be open, c 0 ∈ R c_{0}\in\mathbb{R} c 0 ∈ R , and let g : R d → R g:\mathbb{R}^{d}\to\mathbb{R} g : R d → R be continuous at every point of O O O relative to O O O and equal to c 0 c_{0} c 0 off O O O . Then g g g is Borel: for real c c c the set { g > c } \{g>c\} { g > c } is the union of { x ∈ O : g ( x ) > c } \{x\in O:g(x)>c\} { x ∈ O : g ( x ) > c } , which is open (if x ∈ O x\in O x ∈ O and g ( x ) > c g(x)>c g ( x ) > c , continuity of g g g at x x x relative to O O O gives a positive r r r with g ( y ) > c g(y)>c g ( y ) > c for y ∈ O y\in O y ∈ O with ∥ y − x ∥ < r \lVert y-x\rVert<r ∥ y − x ∥ < r , and shrinking r r r so that this ball lies in the open set O O O shows that the ball lies in { x ∈ O : g ( x ) > c } \{x\in O:g(x)>c\} { x ∈ O : g ( x ) > c } ), and of R d ∖ O \mathbb{R}^{d}\setminus O R d ∖ O or ∅ \varnothing ∅ according as c 0 > c c_{0}>c c 0 > c or not; both are Borel by claims 4 and 5 of The Borel Sigma-Algebra of a Euclidean Space as a Product, and Measurability of Projections, Sequentially Continuous Maps, and Open and Closed Sets , so g g g is Borel by claim 3 of Rational Intervals and Rays Generate the Borel Sigma-Algebra of the Real Line . A map into R d \mathbb{R}^{d} R d is Borel when its components are (claims 2 and 5 of The Borel Sigma-Algebra of a Euclidean Space as a Product, and Measurability of Projections, Sequentially Continuous Maps, and Open and Closed Sets ); constants, indicators of Borel sets, sums, products, absolute values, maxima and everywhere-convergent pointwise limits of Borel real functions are Borel by claims 1 to 5 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions . The functions U U U and ∂ i U \partial_{i}U ∂ i U are continuous on D D D (U U U being of class C 2 C^{2} C 2 ; Differential Calculus and Convexity on Euclidean Open Sets: Standing Notation §derivatives ), f f f is continuous on D D D , and exp \exp exp , ∣ ⋅ ∣ |\cdot| ∣ ⋅ ∣ and ∥ ⋅ ∥ \lVert\cdot\rVert ∥ ⋅ ∥ are continuous; so by (BC) the maps f ˉ \bar{f} f ˉ , γ \gamma γ and G G G of the statement are Borel, as are U ˉ \bar{U} U ˉ and ∇ U \nabla U ∇ U (The Relative Free Energy and the Relative Score of a Probability Measure for a Potential on an Open Set ).
Since D D D is open and nonempty it contains an open ball B ( z , ρ ) B(z,\rho) B ( z , ρ ) , which has positive λ d \lambda_{d} λ d -measure by claim 1 of Balls Have Positive Lebesgue Measure and Bounded Sets Have Finite Lebesgue Measure ; hence λ d ( D ) > 0 \lambda_{d}(D)>0 λ d ( D ) > 0 by claim 2 of Basic Properties of a Measure . In particular D D D is not null: if D ⊆ N D\subseteq N D ⊆ N with N N N Borel and λ d ( N ) = 0 \lambda_{d}(N)=0 λ d ( N ) = 0 , the same claim would give λ d ( D ) ≤ 0 \lambda_{d}(D)\le0 λ d ( D ) ≤ 0 .
Step 2 (Convexity and subgradients). (a) U U U is convex on D D D . Let x , y ∈ D x,y\in D x , y ∈ D , put v = x − y v=x-y v = x − y and c ( τ ) = y + τ v = τ x + ( 1 − τ ) y c(\tau)=y+\tau v=\tau x+(1-\tau)y c ( τ ) = y + τv = τx + ( 1 − τ ) y for τ ∈ [ 0 , 1 ] \tau\in[0,1] τ ∈ [ 0 , 1 ] ; then c ( τ ) ∈ D c(\tau)\in D c ( τ ) ∈ D since D D D is convex. By A Real-Valued C^1 Function is Differentiable at Every Point , U U U is differentiable at every point of D D D with derivative matrix the row D U DU D U , and U U U is continuous on D D D ; so u ( τ ) = U ( c ( τ ) ) u(\tau)=U(c(\tau)) u ( τ ) = U ( c ( τ )) is continuous on [ 0 , 1 ] [0,1] [ 0 , 1 ] and, by Chain Rule Along an Affine Path on the interval [ 0 , 1 ] [0,1] [ 0 , 1 ] , differentiable at every τ ∈ ( 0 , 1 ) \tau\in(0,1) τ ∈ ( 0 , 1 ) with u ′ ( τ ) = D U ( c ( τ ) ) ⋅ v u'(\tau)=DU(c(\tau))\cdot v u ′ ( τ ) = D U ( c ( τ )) ⋅ v . For 0 < s < t < 1 0<s<t<1 0 < s < t < 1 we have c ( t ) − c ( s ) = ( t − s ) v c(t)-c(s)=(t-s)v c ( t ) − c ( s ) = ( t − s ) v , so the monotone-gradient hypothesis gives
( t − s ) ( u ′ ( t ) − u ′ ( s ) ) = ( D U ( c ( t ) ) − D U ( c ( s ) ) ) ⋅ ( c ( t ) − c ( s ) ) ≥ 0 , (t-s)\bigl(u'(t)-u'(s)\bigr)=\bigl(DU(c(t))-DU(c(s))\bigr)\cdot\bigl(c(t)-c(s)\bigr)\ge0, ( t − s ) ( u ′ ( t ) − u ′ ( s ) ) = ( D U ( c ( t )) − D U ( c ( s )) ) ⋅ ( c ( t ) − c ( s ) ) ≥ 0 ,
that is, u ′ ( s ) ≤ u ′ ( t ) u'(s)\le u'(t) u ′ ( s ) ≤ u ′ ( t ) . Let τ ∈ ( 0 , 1 ) \tau\in(0,1) τ ∈ ( 0 , 1 ) . By Mean Value Theorem on a Closed Real Interval on [ 0 , τ ] [0,\tau] [ 0 , τ ] and on [ τ , 1 ] [\tau,1] [ τ , 1 ] there are s 1 ∈ ( 0 , τ ) s_{1}\in(0,\tau) s 1 ∈ ( 0 , τ ) and s 2 ∈ ( τ , 1 ) s_{2}\in(\tau,1) s 2 ∈ ( τ , 1 ) with u ( τ ) − u ( 0 ) = τ u ′ ( s 1 ) u(\tau)-u(0)=\tau u'(s_{1}) u ( τ ) − u ( 0 ) = τ u ′ ( s 1 ) and u ( 1 ) − u ( τ ) = ( 1 − τ ) u ′ ( s 2 ) u(1)-u(\tau)=(1-\tau)u'(s_{2}) u ( 1 ) − u ( τ ) = ( 1 − τ ) u ′ ( s 2 ) . As s 1 < s 2 s_{1}<s_{2} s 1 < s 2 ,
( 1 − τ ) ( u ( τ ) − u ( 0 ) ) = τ ( 1 − τ ) u ′ ( s 1 ) ≤ τ ( 1 − τ ) u ′ ( s 2 ) = τ ( u ( 1 ) − u ( τ ) ) , (1-\tau)\bigl(u(\tau)-u(0)\bigr)=\tau(1-\tau)u'(s_{1})\le\tau(1-\tau)u'(s_{2})=\tau\bigl(u(1)-u(\tau)\bigr), ( 1 − τ ) ( u ( τ ) − u ( 0 ) ) = τ ( 1 − τ ) u ′ ( s 1 ) ≤ τ ( 1 − τ ) u ′ ( s 2 ) = τ ( u ( 1 ) − u ( τ ) ) ,
which rearranges to U ( τ x + ( 1 − τ ) y ) ≤ τ U ( x ) + ( 1 − τ ) U ( y ) U(\tau x+(1-\tau)y)\le\tau U(x)+(1-\tau)U(y) U ( τx + ( 1 − τ ) y ) ≤ τU ( x ) + ( 1 − τ ) U ( y ) ; for τ ∈ { 0 , 1 } \tau\in\{0,1\} τ ∈ { 0 , 1 } this holds with equality. So U U U is convex on D D D , and hence semiconvex on D D D with constant 0 0 0 .
(b) For every x ∈ D x\in D x ∈ D , U U U is differentiable at x x x with derivative matrix the row D U ( x ) DU(x) D U ( x ) , so ∂ D U ( x ) = { D U ( x ) } \partial_{D}U(x)=\{DU(x)\} ∂ D U ( x ) = { D U ( x )} by the subdifferential at a point of differentiability , ∂ D \partial_{D} ∂ D denoting the subdifferential relative to D D D .
(c) Let g : D → R g:D\to\mathbb{R} g : D → R , g ( x ) = f ( x ) + K 2 ∥ x ∥ 2 g(x)=f(x)+\tfrac{K}{2}\lVert x\rVert^{2} g ( x ) = f ( x ) + 2 K ∥ x ∥ 2 ; it is convex on D D D by Semiconvex Function on a Convex Subset of R n \mathbb{R}^n R n . Let x ∈ D x\in D x ∈ D be a point at which f f f is differentiable, with derivative matrix the row D f ( x ) Df(x) D f ( x ) . For v ∈ R d v\in\mathbb{R}^{d} v ∈ R d with x + v ∈ D x+v\in D x + v ∈ D ,
g ( x + v ) − g ( x ) − ( D f ( x ) + K x ) ⋅ v = ( f ( x + v ) − f ( x ) − D f ( x ) ⋅ v ) + K 2 ∥ v ∥ 2 . g(x+v)-g(x)-\bigl(Df(x)+Kx\bigr)\cdot v=\bigl(f(x+v)-f(x)-Df(x)\cdot v\bigr)+\tfrac{K}{2}\lVert v\rVert^{2}. g ( x + v ) − g ( x ) − ( D f ( x ) + K x ) ⋅ v = ( f ( x + v ) − f ( x ) − D f ( x ) ⋅ v ) + 2 K ∥ v ∥ 2 .
Given ε > 0 \varepsilon>0 ε > 0 , choose δ > 0 \delta>0 δ > 0 as in Differentiability at a Point for Maps Between Euclidean Spaces for f f f at x x x and ε / 2 \varepsilon/2 ε /2 , with moreover δ ≤ ε / ( K + 1 ) \delta\le\varepsilon/(K+1) δ ≤ ε / ( K + 1 ) ; then every v v v with 0 < ∥ v ∥ < δ 0<\lVert v\rVert<\delta 0 < ∥ v ∥ < δ has x + v ∈ D x+v\in D x + v ∈ D and the displayed quantity has absolute value at most ε 2 ∥ v ∥ + K 2 ∥ v ∥ 2 ≤ ε ∥ v ∥ \tfrac{\varepsilon}{2}\lVert v\rVert+\tfrac{K}{2}\lVert v\rVert^{2}\le\varepsilon\lVert v\rVert 2 ε ∥ v ∥ + 2 K ∥ v ∥ 2 ≤ ε ∥ v ∥ . So g g g is differentiable at x x x with derivative matrix the row D f ( x ) + K x Df(x)+Kx D f ( x ) + K x , and ∂ D g ( x ) = { D f ( x ) + K x } \partial_{D}g(x)=\{Df(x)+Kx\} ∂ D g ( x ) = { D f ( x ) + K x } by the same claim .
Step 3 (Gradient maps exist; a point of differentiability). By the local Lipschitz bound , f f f is locally Lipschitz on D D D ; by Rademacher's theorem for locally Lipschitz maps (with m = 1 m=1 m = 1 ) the set of x ∈ D x\in D x ∈ D at which f f f is not differentiable is null, so it is contained in a Borel set N 0 N_{0} N 0 with λ d ( N 0 ) = 0 \lambda_{d}(N_{0})=0 λ d ( N 0 ) = 0 .
For i ∈ [ d ] i\in[d] i ∈ [ d ] and natural m ≥ 1 m\ge1 m ≥ 1 let O m , i = { x ∈ D : x + m − 1 e i ∈ D } O_{m,i}=\{x\in D:x+m^{-1}e_{i}\in D\} O m , i = { x ∈ D : x + m − 1 e i ∈ D } , which is open (D D D being open, a ball around x + m − 1 e i x+m^{-1}e_{i} x + m − 1 e i inside D D D translates to a ball around x x x inside O m , i O_{m,i} O m , i after intersecting with a ball around x x x inside D D D ), and let q m , i ( x ) = m ( f ( x + m − 1 e i ) − f ( x ) ) q_{m,i}(x)=m\bigl(f(x+m^{-1}e_{i})-f(x)\bigr) q m , i ( x ) = m ( f ( x + m − 1 e i ) − f ( x ) ) for x ∈ O m , i x\in O_{m,i} x ∈ O m , i and q m , i ( x ) = 0 q_{m,i}(x)=0 q m , i ( x ) = 0 otherwise; q m , i q_{m,i} q m , i is Borel by (BC), f f f being continuous on D D D . Put g m , i = 1 D ∖ N 0 q m , i g_{m,i}=\mathbf{1}_{D\setminus N_{0}}\,q_{m,i} g m , i = 1 D ∖ N 0 q m , i , a Borel function. Let x ∈ D ∖ N 0 x\in D\setminus N_{0} x ∈ D ∖ N 0 . Then f f f is differentiable at x x x , so the partial derivative ∂ i f ( x ) \partial_{i}f(x) ∂ i f ( x ) exists by claim 1 of A Derivative Matrix is the Jacobian Matrix, and is Unique , i.e. ( f ( x + t e i ) − f ( x ) ) / t → ∂ i f ( x ) (f(x+te_{i})-f(x))/t\to\partial_{i}f(x) ( f ( x + t e i ) − f ( x )) / t → ∂ i f ( x ) as t → 0 t\to0 t → 0 (Partial Derivative on a Euclidean Open Set ); as B ( x , ρ ) ⊆ D B(x,\rho)\subseteq D B ( x , ρ ) ⊆ D for some ρ > 0 \rho>0 ρ > 0 , x ∈ O m , i x\in O_{m,i} x ∈ O m , i for all m > 1 / ρ m>1/\rho m > 1/ ρ , and so g m , i ( x ) → ∂ i f ( x ) g_{m,i}(x)\to\partial_{i}f(x) g m , i ( x ) → ∂ i f ( x ) . For x ∉ D ∖ N 0 x\notin D\setminus N_{0} x ∈ / D ∖ N 0 , g m , i ( x ) = 0 g_{m,i}(x)=0 g m , i ( x ) = 0 for all m m m . Hence g m , i g_{m,i} g m , i converges everywhere, and the map ∇ 0 f : R d → R d \nabla^{0}f:\mathbb{R}^{d}\to\mathbb{R}^{d} ∇ 0 f : R d → R d whose i i i th component is lim m g m , i \lim_{m}g_{m,i} lim m g m , i is Borel, with ∇ 0 f ( x ) = D f ( x ) \nabla^{0}f(x)=Df(x) ∇ 0 f ( x ) = D f ( x ) at every x ∈ D ∖ N 0 x\in D\setminus N_{0} x ∈ D ∖ N 0 . Thus ∇ 0 f \nabla^{0}f ∇ 0 f is a gradient map of f f f (with N = N 0 N=N_{0} N = N 0 ), and f f f has a gradient map.
Since D D D is not null (Step 1), D ⊈ N 0 D\not\subseteq N_{0} D ⊆ N 0 ; fix x 0 ∈ D ∖ N 0 x_{0}\in D\setminus N_{0} x 0 ∈ D ∖ N 0 , a point at which f f f is differentiable.
Step 4 (Quadratic bounds for f f f ). By Step 2(c), D f ( x 0 ) + K x 0 ∈ ∂ D g ( x 0 ) Df(x_{0})+Kx_{0}\in\partial_{D}g(x_{0}) D f ( x 0 ) + K x 0 ∈ ∂ D g ( x 0 ) , so for every x ∈ D x\in D x ∈ D , by Subdifferential of a Real-Valued Function on a Convex Subset of R n \mathbb{R}^n R n §subdifferential and the identity ∥ x ∥ 2 − ∥ x 0 ∥ 2 − 2 x 0 ⋅ ( x − x 0 ) = ∥ x − x 0 ∥ 2 \lVert x\rVert^{2}-\lVert x_{0}\rVert^{2}-2x_{0}\cdot(x-x_{0})=\lVert x-x_{0}\rVert^{2} ∥ x ∥ 2 − ∥ x 0 ∥ 2 − 2 x 0 ⋅ ( x − x 0 ) = ∥ x − x 0 ∥ 2 ,
f ( x ) ≥ f ( x 0 ) + D f ( x 0 ) ⋅ ( x − x 0 ) − K 2 ∥ x − x 0 ∥ 2 . f(x)\ge f(x_{0})+Df(x_{0})\cdot(x-x_{0})-\tfrac{K}{2}\lVert x-x_{0}\rVert^{2}. f ( x ) ≥ f ( x 0 ) + D f ( x 0 ) ⋅ ( x − x 0 ) − 2 K ∥ x − x 0 ∥ 2 .
Using ∣ D f ( x 0 ) ⋅ ( x − x 0 ) ∣ ≤ ∥ D f ( x 0 ) ∥ ( ∥ x ∥ + ∥ x 0 ∥ ) |Df(x_{0})\cdot(x-x_{0})|\le\lVert Df(x_{0})\rVert(\lVert x\rVert+\lVert x_{0}\rVert) ∣ D f ( x 0 ) ⋅ ( x − x 0 ) ∣ ≤ ∥ D f ( x 0 )∥ (∥ x ∥ + ∥ x 0 ∥) , ∥ x ∥ ≤ 1 + ∥ x ∥ 2 \lVert x\rVert\le1+\lVert x\rVert^{2} ∥ x ∥ ≤ 1 + ∥ x ∥ 2 and ∥ x − x 0 ∥ 2 ≤ 2 ∥ x ∥ 2 + 2 ∥ x 0 ∥ 2 \lVert x-x_{0}\rVert^{2}\le2\lVert x\rVert^{2}+2\lVert x_{0}\rVert^{2} ∥ x − x 0 ∥ 2 ≤ 2 ∥ x ∥ 2 + 2 ∥ x 0 ∥ 2 , we get f ( x ) ≥ − C 0 ( 1 + ∥ x ∥ 2 ) f(x)\ge-C_{0}(1+\lVert x\rVert^{2}) f ( x ) ≥ − C 0 ( 1 + ∥ x ∥ 2 ) with C 0 = ∣ f ( x 0 ) ∣ + ∥ D f ( x 0 ) ∥ ( 1 + ∥ x 0 ∥ ) + K ( 1 + ∥ x 0 ∥ 2 ) C_{0}=|f(x_{0})|+\lVert Df(x_{0})\rVert(1+\lVert x_{0}\rVert)+K(1+\lVert x_{0}\rVert^{2}) C 0 = ∣ f ( x 0 ) ∣ + ∥ D f ( x 0 )∥ ( 1 + ∥ x 0 ∥) + K ( 1 + ∥ x 0 ∥ 2 ) . Together with f ( x ) ≤ F f(x)\le F f ( x ) ≤ F this gives, with C 1 = C 0 + ∣ F ∣ C_{1}=C_{0}+|F| C 1 = C 0 + ∣ F ∣ ,
∣ f ( x ) ∣ ≤ C 1 ( 1 + ∥ x ∥ 2 ) ( x ∈ D ) . (4.1) |f(x)|\le C_{1}\bigl(1+\lVert x\rVert^{2}\bigr)\qquad(x\in D).\tag{4.1} ∣ f ( x ) ∣ ≤ C 1 ( 1 + ∥ x ∥ 2 ) ( x ∈ D ) . ( 4.1 )
Step 5 (The normalising constant, the density, and an integrability criterion). For x ∈ D x\in D x ∈ D , claims 1, 2 and 4 of Basic Properties of the Exponential Function give γ ( x ) = exp ( f ( x ) / a ) exp ( − U ( x ) / a ) ≤ exp ( F / a ) exp ( − U ( x ) / a ) ≤ exp ( F / a ) G ( x ) \gamma(x)=\exp(f(x)/a)\exp(-U(x)/a)\le\exp(F/a)\exp(-U(x)/a)\le\exp(F/a)G(x) γ ( x ) = exp ( f ( x ) / a ) exp ( − U ( x ) / a ) ≤ exp ( F / a ) exp ( − U ( x ) / a ) ≤ exp ( F / a ) G ( x ) , and γ ( x ) > 0 \gamma(x)>0 γ ( x ) > 0 . Hence 0 ≤ γ ≤ exp ( F / a ) G 0\le\gamma\le\exp(F/a)G 0 ≤ γ ≤ exp ( F / a ) G on R d \mathbb{R}^{d} R d , and Z ≤ exp ( F / a ) ∫ G d λ d < ∞ Z\le\exp(F/a)\int G\,d\lambda_{d}<\infty Z ≤ exp ( F / a ) ∫ G d λ d < ∞ by claim 1 of Linearity and Monotonicity of the Lebesgue Integral . If Z = 0 Z=0 Z = 0 , then γ = 0 \gamma=0 γ = 0 almost everywhere by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §vanishing , so D ⊆ { γ ≠ 0 } D\subseteq\{\gamma\ne0\} D ⊆ { γ = 0 } would be null, contradicting Step 1. So 0 < Z < ∞ 0<Z<\infty 0 < Z < ∞ .
Let p = Z − 1 γ p=Z^{-1}\gamma p = Z − 1 γ , a Borel function R d → [ 0 , ∞ ) \mathbb{R}^{d}\to[0,\infty) R d → [ 0 , ∞ ) with ∫ p d λ d = 1 \int p\,d\lambda_{d}=1 ∫ p d λ d = 1 . By claim 3 of Image Measures, Measures with Densities, and Change of Variables , π f \pi_{f} π f is a measure on B ( R d ) \mathcal{B}(\mathbb{R}^{d}) B ( R d ) with π f ( R d ) = ∫ p d λ d = 1 \pi_{f}(\mathbb{R}^{d})=\int p\,d\lambda_{d}=1 π f ( R d ) = ∫ p d λ d = 1 , so it is a probability measure on R d \mathbb{R}^{d} R d , and p p p is a density of π f \pi_{f} π f with respect to λ d \lambda_{d} λ d in the sense of The Radon-Nikodym Theorem for a Finite Measure and a Sigma-Finite Measure, and Uniqueness of Densities . As 1 R d ∖ D p \mathbf{1}_{\mathbb{R}^{d}\setminus D}\,p 1 R d ∖ D p vanishes identically, π f ( R d ∖ D ) = 0 \pi_{f}(\mathbb{R}^{d}\setminus D)=0 π f ( R d ∖ D ) = 0 and π f ( D ) = 1 \pi_{f}(D)=1 π f ( D ) = 1 by claim 3 of Basic Properties of a Measure . For x ∈ D x\in D x ∈ D , claims 1 and 2 of Basic Properties of the Exponential Function and The Natural Logarithm give exp ( − log Z ) = Z − 1 \exp(-\log Z)=Z^{-1} exp ( − log Z ) = Z − 1 and hence
p ( x ) = exp ( h ( x ) − log Z ) , 0 < p ( x ) ≤ Z − 1 e H , p ( x ) ≤ Z − 1 e F / a G ( x ) . (5.1) p(x)=\exp\bigl(h(x)-\log Z\bigr),\qquad 0<p(x)\le Z^{-1}e^{H},\qquad p(x)\le Z^{-1}e^{F/a}G(x).\tag{5.1} p ( x ) = exp ( h ( x ) − log Z ) , 0 < p ( x ) ≤ Z − 1 e H , p ( x ) ≤ Z − 1 e F / a G ( x ) . ( 5.1 )
(IC) Let u : R d → R u:\mathbb{R}^{d}\to\mathbb{R} u : R d → R be Borel and c ≥ 0 c\ge0 c ≥ 0 real, and suppose that the set of x ∈ D x\in D x ∈ D with ∣ u ( x ) ∣ > c ( 1 + ∣ U ( x ) ∣ + ∥ D U ( x ) ∥ 2 + ∥ x ∥ 2 ) |u(x)|>c\bigl(1+|U(x)|+\lVert DU(x)\rVert^{2}+\lVert x\rVert^{2}\bigr) ∣ u ( x ) ∣ > c ( 1 + ∣ U ( x ) ∣ + ∥ D U ( x ) ∥ 2 + ∥ x ∥ 2 ) is null. Then u u u is integrable with respect to π f \pi_{f} π f and ∫ u d π f = ∫ u p d λ d \int u\,d\pi_{f}=\int u\,p\,d\lambda_{d} ∫ u d π f = ∫ u p d λ d . Indeed, by (5.1) and since p = 0 p=0 p = 0 off D D D , ∣ u ∣ p ≤ c Z − 1 e F / a G |u|\,p\le cZ^{-1}e^{F/a}G ∣ u ∣ p ≤ c Z − 1 e F / a G outside a null set, so ∫ ∣ u ∣ p d λ d ≤ c Z − 1 e F / a ∫ G d λ d < ∞ \int|u|\,p\,d\lambda_{d}\le cZ^{-1}e^{F/a}\int G\,d\lambda_{d}<\infty ∫ ∣ u ∣ p d λ d ≤ c Z − 1 e F / a ∫ G d λ d < ∞ by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §comparison and claim 1 of Linearity and Monotonicity of the Lebesgue Integral ; so the Borel function u p up u p is integrable with respect to λ d \lambda_{d} λ d (Integrable Function and the Lebesgue Integral ), and claim 3 of Image Measures, Measures with Densities, and Change of Variables gives the assertion.
Step 6 (Clause 1, the Gibbs measure). By (4.1) and (IC) with c = C 1 c=C_{1} c = C 1 , f ˉ \bar{f} f ˉ is integrable with respect to π f \pi_{f} π f ; by (IC) with c = 1 c=1 c = 1 , so are U ˉ \bar{U} U ˉ , 1 D \mathbf{1}_{D} 1 D and the Borel function x ↦ ∥ x ∥ 2 x\mapsto\lVert x\rVert^{2} x ↦ ∥ x ∥ 2 of The Second Moment of a Probability Measure on Euclidean Space and the Probability Measures with Finite Second Moment §moment . Hence M 2 ( π f ) < ∞ M_{2}(\pi_{f})<\infty M 2 ( π f ) < ∞ and π f ∈ P 2 ( R d ) \pi_{f}\in\mathcal{P}_{2}(\mathbb{R}^{d}) π f ∈ P 2 ( R d ) (The Second Moment of a Probability Measure on Euclidean Space and the Probability Measures with Finite Second Moment §space ). Let
ℓ ˉ = a − 1 ( f ˉ − U ˉ ) − ( log Z ) 1 D , \bar{\ell}=a^{-1}\bigl(\bar{f}-\bar{U}\bigr)-(\log Z)\,\mathbf{1}_{D}, ℓ ˉ = a − 1 ( f ˉ − U ˉ ) − ( log Z ) 1 D ,
a Borel function, integrable with respect to π f \pi_{f} π f by claim 2 of Linearity and Monotonicity of the Lebesgue Integral , with ∫ ℓ ˉ d π f = a − 1 ( ∫ f ˉ d π f − ∫ U ˉ d π f ) − log Z \int\bar{\ell}\,d\pi_{f}=a^{-1}\bigl(\int\bar{f}\,d\pi_{f}-\int\bar{U}\,d\pi_{f}\bigr)-\log Z ∫ ℓ ˉ d π f = a − 1 ( ∫ f ˉ d π f − ∫ U ˉ d π f ) − log Z since π f ( D ) = 1 \pi_{f}(D)=1 π f ( D ) = 1 . For x ∈ D x\in D x ∈ D , ℓ ˉ ( x ) = h ( x ) − log Z \bar{\ell}(x)=h(x)-\log Z ℓ ˉ ( x ) = h ( x ) − log Z , so p ( x ) = exp ( ℓ ˉ ( x ) ) p(x)=\exp(\bar{\ell}(x)) p ( x ) = exp ( ℓ ˉ ( x )) and log p ( x ) = ℓ ˉ ( x ) \log p(x)=\bar{\ell}(x) log p ( x ) = ℓ ˉ ( x ) by (5.1) and The Natural Logarithm . With ϕ ( s ) = s log s \phi(s)=s\log s ϕ ( s ) = s log s (ϕ ( 0 ) = 0 \phi(0)=0 ϕ ( 0 ) = 0 ) as in The Function s log s s\log s s log s : Continuity, Young's Inequality and Lower Bounds, with the Elementary Bounds for the Exponential and the Logarithm , therefore ϕ ( p ( x ) ) = p ( x ) ℓ ˉ ( x ) \phi(p(x))=p(x)\bar{\ell}(x) ϕ ( p ( x )) = p ( x ) ℓ ˉ ( x ) for x ∈ D x\in D x ∈ D , and ϕ ( p ( x ) ) = ϕ ( 0 ) = 0 = p ( x ) ℓ ˉ ( x ) \phi(p(x))=\phi(0)=0=p(x)\bar{\ell}(x) ϕ ( p ( x )) = ϕ ( 0 ) = 0 = p ( x ) ℓ ˉ ( x ) for x ∉ D x\notin D x ∈ / D . Thus ϕ ∘ p = p ℓ ˉ \phi\circ p=p\,\bar{\ell} ϕ ∘ p = p ℓ ˉ , which is integrable with respect to λ d \lambda_{d} λ d with ∫ ϕ ∘ p d λ d = ∫ ℓ ˉ d π f \int\phi\circ p\,d\lambda_{d}=\int\bar{\ell}\,d\pi_{f} ∫ ϕ ∘ p d λ d = ∫ ℓ ˉ d π f by claim 3 of Image Measures, Measures with Densities, and Change of Variables . By The Entropy of a Probability Measure on Euclidean Space §entropy , π f \pi_{f} π f has finite entropy and
E n t ( π f ) = a − 1 ( ∫ f ˉ d π f − ∫ U ˉ d π f ) − log Z . \mathrm{Ent}(\pi_{f})=a^{-1}\Bigl(\int\bar{f}\,d\pi_{f}-\int\bar{U}\,d\pi_{f}\Bigr)-\log Z . Ent ( π f ) = a − 1 ( ∫ f ˉ d π f − ∫ U ˉ d π f ) − log Z .
So π f ∈ P 2 E n t ( R d ) \pi_{f}\in\mathcal{P}_{2}^{\mathrm{Ent}}(\mathbb{R}^{d}) π f ∈ P 2 Ent ( R d ) , and with π f ( D ) = 1 \pi_{f}(D)=1 π f ( D ) = 1 and the integrability of U ˉ \bar{U} U ˉ we get π f ∈ D U , a \pi_{f}\in\mathcal{D}_{U,a} π f ∈ D U , a (The Relative Free Energy and the Relative Score of a Probability Measure for a Potential on an Open Set §energy ), with E U , a ( π f ) = a E n t ( π f ) + ∫ U ˉ d π f = ∫ f ˉ d π f − a log Z \mathcal{E}_{U,a}(\pi_{f})=a\,\mathrm{Ent}(\pi_{f})+\int\bar{U}\,d\pi_{f}=\int\bar{f}\,d\pi_{f}-a\log Z E U , a ( π f ) = a Ent ( π f ) + ∫ U ˉ d π f = ∫ f ˉ d π f − a log Z . This is clause 1.
Step 7 (Clause 2). Let μ ∈ D U , a \mu\in\mathcal{D}_{U,a} μ ∈ D U , a with f ˉ \bar{f} f ˉ integrable with respect to μ \mu μ . By The Entropy of a Probability Measure on Euclidean Space §entropy , μ \mu μ has a density q q q with respect to λ d \lambda_{d} λ d (a Borel q : R d → [ 0 , ∞ ) q:\mathbb{R}^{d}\to[0,\infty) q : R d → [ 0 , ∞ ) with μ ( B ) = ∫ 1 B q d λ d \mu(B)=\int\mathbf{1}_{B}\,q\,d\lambda_{d} μ ( B ) = ∫ 1 B q d λ d for all Borel B B B , The Radon-Nikodym Theorem for a Finite Measure and a Sigma-Finite Measure, and Uniqueness of Densities ) such that ϕ ∘ q \phi\circ q ϕ ∘ q is integrable with respect to λ d \lambda_{d} λ d and E n t ( μ ) = ∫ ϕ ∘ q d λ d \mathrm{Ent}(\mu)=\int\phi\circ q\,d\lambda_{d} Ent ( μ ) = ∫ ϕ ∘ q d λ d ; so μ \mu μ is the measure with density q q q of claim 3 of Image Measures, Measures with Densities, and Change of Variables , and ∫ q d λ d = μ ( R d ) = 1 \int q\,d\lambda_{d}=\mu(\mathbb{R}^{d})=1 ∫ q d λ d = μ ( R d ) = 1 . Since μ ( D ) = 1 \mu(D)=1 μ ( D ) = 1 , ∫ 1 R d ∖ D q d λ d = μ ( R d ∖ D ) = 0 \int\mathbf{1}_{\mathbb{R}^{d}\setminus D}\,q\,d\lambda_{d}=\mu(\mathbb{R}^{d}\setminus D)=0 ∫ 1 R d ∖ D q d λ d = μ ( R d ∖ D ) = 0 , so the set E E E of x ∉ D x\notin D x ∈ / D with q ( x ) ≠ 0 q(x)\ne0 q ( x ) = 0 is null by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §vanishing . The function ℓ ˉ \bar{\ell} ℓ ˉ of Step 6 is integrable with respect to μ \mu μ , being a linear combination of f ˉ \bar{f} f ˉ , U ˉ \bar{U} U ˉ (integrable as μ ∈ D U , a \mu\in\mathcal{D}_{U,a} μ ∈ D U , a ) and the bounded Borel function 1 D \mathbf{1}_{D} 1 D (Probability Measures on Euclidean Space and Random Vectors: Standing Notation §measures ), with ∫ ℓ ˉ d μ = a − 1 ( ∫ f ˉ d μ − ∫ U ˉ d μ ) − log Z \int\bar{\ell}\,d\mu=a^{-1}\bigl(\int\bar{f}\,d\mu-\int\bar{U}\,d\mu\bigr)-\log Z ∫ ℓ ˉ d μ = a − 1 ( ∫ f ˉ d μ − ∫ U ˉ d μ ) − log Z ; by claim 3 of Image Measures, Measures with Densities, and Change of Variables , q ℓ ˉ q\bar{\ell} q ℓ ˉ is integrable with respect to λ d \lambda_{d} λ d with ∫ q ℓ ˉ d λ d = ∫ ℓ ˉ d μ \int q\bar{\ell}\,d\lambda_{d}=\int\bar{\ell}\,d\mu ∫ q ℓ ˉ d λ d = ∫ ℓ ˉ d μ .
We claim that for every x ∈ R d ∖ E x\in\mathbb{R}^{d}\setminus E x ∈ R d ∖ E ,
q ( x ) ℓ ˉ ( x ) − ϕ ( q ( x ) ) ≤ p ( x ) − q ( x ) . (7.1) q(x)\,\bar{\ell}(x)-\phi(q(x))\le p(x)-q(x).\tag{7.1} q ( x ) ℓ ˉ ( x ) − ϕ ( q ( x )) ≤ p ( x ) − q ( x ) . ( 7.1 )
If x ∉ D x\notin D x ∈ / D , then q ( x ) = 0 = p ( x ) q(x)=0=p(x) q ( x ) = 0 = p ( x ) and both sides vanish. If x ∈ D x\in D x ∈ D and q ( x ) = 0 q(x)=0 q ( x ) = 0 , the left side is 0 < p ( x ) 0<p(x) 0 < p ( x ) . If x ∈ D x\in D x ∈ D and q ( x ) > 0 q(x)>0 q ( x ) > 0 , put v = ℓ ˉ ( x ) − log q ( x ) v=\bar{\ell}(x)-\log q(x) v = ℓ ˉ ( x ) − log q ( x ) ; by Step 6 and claims 1 and 2 of Basic Properties of the Exponential Function , exp ( v ) = exp ( ℓ ˉ ( x ) ) exp ( log q ( x ) ) − 1 = p ( x ) / q ( x ) \exp(v)=\exp(\bar{\ell}(x))\exp(\log q(x))^{-1}=p(x)/q(x) exp ( v ) = exp ( ℓ ˉ ( x )) exp ( log q ( x ) ) − 1 = p ( x ) / q ( x ) , and The Function s log s s\log s s log s : Continuity, Young's Inequality and Lower Bounds, with the Elementary Bounds for the Exponential and the Logarithm §exp gives 1 + v ≤ p ( x ) / q ( x ) 1+v\le p(x)/q(x) 1 + v ≤ p ( x ) / q ( x ) ; multiplying by q ( x ) > 0 q(x)>0 q ( x ) > 0 and using ϕ ( q ( x ) ) = q ( x ) log q ( x ) \phi(q(x))=q(x)\log q(x) ϕ ( q ( x )) = q ( x ) log q ( x ) gives (7.1).
Let w = p − q − q ℓ ˉ + ϕ ∘ q w=p-q-q\bar{\ell}+\phi\circ q w = p − q − q ℓ ˉ + ϕ ∘ q , integrable with respect to λ d \lambda_{d} λ d by claim 2 of Linearity and Monotonicity of the Lebesgue Integral . By (7.1), w = max ( w , 0 ) w=\max(w,0) w = max ( w , 0 ) outside the null set E E E , so by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §comparison and claim 2 of Linearity and Monotonicity of the Lebesgue Integral , ∫ w d λ d = ∫ max ( w , 0 ) d λ d ≥ 0 \int w\,d\lambda_{d}=\int\max(w,0)\,d\lambda_{d}\ge0 ∫ w d λ d = ∫ max ( w , 0 ) d λ d ≥ 0 . By linearity, 0 ≤ 1 − 1 − ∫ ℓ ˉ d μ + E n t ( μ ) 0\le1-1-\int\bar{\ell}\,d\mu+\mathrm{Ent}(\mu) 0 ≤ 1 − 1 − ∫ ℓ ˉ d μ + Ent ( μ ) , that is,
a − 1 ( ∫ f ˉ d μ − ∫ U ˉ d μ ) − log Z ≤ E n t ( μ ) , a^{-1}\Bigl(\int\bar{f}\,d\mu-\int\bar{U}\,d\mu\Bigr)-\log Z\le\mathrm{Ent}(\mu), a − 1 ( ∫ f ˉ d μ − ∫ U ˉ d μ ) − log Z ≤ Ent ( μ ) ,
and multiplying by a > 0 a>0 a > 0 gives ∫ f ˉ d μ − ( a E n t ( μ ) + ∫ U ˉ d μ ) ≤ a log Z \int\bar{f}\,d\mu-\bigl(a\,\mathrm{Ent}(\mu)+\int\bar{U}\,d\mu\bigr)\le a\log Z ∫ f ˉ d μ − ( a Ent ( μ ) + ∫ U ˉ d μ ) ≤ a log Z , which is clause 2; the maximisation statement follows from clause 1.
Step 8 (Square-integrable gradients). For the rest of the proof fix a gradient map ∇ f \nabla f ∇ f of f f f , with a null set N N N as in the statement, and a Borel set N 1 ⊇ N N_{1}\supseteq N N 1 ⊇ N with λ d ( N 1 ) = 0 \lambda_{d}(N_{1})=0 λ d ( N 1 ) = 0 . Then π f ( N 1 ) = ∫ 1 N 1 p d λ d = 0 \pi_{f}(N_{1})=\int\mathbf{1}_{N_{1}}\,p\,d\lambda_{d}=0 π f ( N 1 ) = ∫ 1 N 1 p d λ d = 0 by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §null-integral , so the Borel set D ∖ N 1 D\setminus N_{1} D ∖ N 1 has π f ( D ∖ N 1 ) = 1 \pi_{f}(D\setminus N_{1})=1 π f ( D ∖ N 1 ) = 1 . At every x ∈ D ∖ N 1 x\in D\setminus N_{1} x ∈ D ∖ N 1 , f f f is differentiable and ∇ f ( x ) = D f ( x ) \nabla f(x)=Df(x) ∇ f ( x ) = D f ( x ) , so by hypothesis ∥ ∇ f ( x ) ∥ 2 ≤ A + B ( U ( x ) − p 0 ) ≤ ( A + B + B ∣ p 0 ∣ ) ( 1 + ∣ U ( x ) ∣ ) \lVert\nabla f(x)\rVert^{2}\le A+B(U(x)-p_{0})\le(A+B+B|p_{0}|)(1+|U(x)|) ∥ ∇ f ( x ) ∥ 2 ≤ A + B ( U ( x ) − p 0 ) ≤ ( A + B + B ∣ p 0 ∣ ) ( 1 + ∣ U ( x ) ∣ ) . By (IC) with c = A + B + B ∣ p 0 ∣ c=A+B+B|p_{0}| c = A + B + B ∣ p 0 ∣ , and with c = 1 c=1 c = 1 , the Borel functions ∥ ∇ f ∥ 2 \lVert\nabla f\rVert^{2} ∥ ∇ f ∥ 2 and ∥ ∇ U ∥ 2 \lVert\nabla U\rVert^{2} ∥ ∇ U ∥ 2 are integrable with respect to π f \pi_{f} π f (∇ U = D U \nabla U=DU ∇ U = D U on D D D ). Let
η = a − 1 ( ∇ f − ∇ U ) , T g = ∇ f + K i d , \eta=a^{-1}\bigl(\nabla f-\nabla U\bigr),\qquad T_{g}=\nabla f+K\,\mathrm{id}, η = a − 1 ( ∇ f − ∇ U ) , T g = ∇ f + K id ,
Borel maps R d → R d \mathbb{R}^{d}\to\mathbb{R}^{d} R d → R d with ∥ η ∥ 2 ≤ 2 a − 2 ( ∥ ∇ f ∥ 2 + ∥ ∇ U ∥ 2 ) \lVert\eta\rVert^{2}\le2a^{-2}(\lVert\nabla f\rVert^{2}+\lVert\nabla U\rVert^{2}) ∥ η ∥ 2 ≤ 2 a − 2 (∥ ∇ f ∥ 2 + ∥ ∇ U ∥ 2 ) and ∥ T g ∥ 2 ≤ 2 ∥ ∇ f ∥ 2 + 2 K 2 ∥ i d ∥ 2 \lVert T_{g}\rVert^{2}\le2\lVert\nabla f\rVert^{2}+2K^{2}\lVert\mathrm{id}\rVert^{2} ∥ T g ∥ 2 ≤ 2 ∥ ∇ f ∥ 2 + 2 K 2 ∥ id ∥ 2 ; by Step 6 and claim 1 of Linearity and Monotonicity of the Lebesgue Integral , all of ∥ ∇ f ∥ 2 \lVert\nabla f\rVert^{2} ∥ ∇ f ∥ 2 , ∥ ∇ U ∥ 2 \lVert\nabla U\rVert^{2} ∥ ∇ U ∥ 2 , ∥ η ∥ 2 \lVert\eta\rVert^{2} ∥ η ∥ 2 , ∥ T g ∥ 2 \lVert T_{g}\rVert^{2} ∥ T g ∥ 2 have finite π f \pi_{f} π f -integrals. So the classes of ∇ f \nabla f ∇ f , ∇ U \nabla U ∇ U , η \eta η , T g T_{g} T g and i d \mathrm{id} id lie in L 2 ( π f ; R d ) L^{2}(\pi_{f};\mathbb{R}^{d}) L 2 ( π f ; R d ) , where η = a − 1 ( ∇ f − ∇ U ) \eta=a^{-1}(\nabla f-\nabla U) η = a − 1 ( ∇ f − ∇ U ) and T g = ∇ f + K i d T_{g}=\nabla f+K\,\mathrm{id} T g = ∇ f + K id as classes, the vector operations on classes being computed on representatives. Moreover p ( 1 + ∥ η ∥ 2 ) p(1+\lVert\eta\rVert^{2}) p ( 1 + ∥ η ∥ 2 ) and p ( 1 + ∥ ∇ U ∥ 2 ) p(1+\lVert\nabla U\rVert^{2}) p ( 1 + ∥ ∇ U ∥ 2 ) are integrable with respect to λ d \lambda_{d} λ d by claim 3 of Image Measures, Measures with Densities, and Change of Variables , and so are p ∥ η ∥ p\lVert\eta\rVert p ∥ η ∥ and p ∥ ∇ U ∥ p\lVert\nabla U\rVert p ∥ ∇ U ∥ , which they dominate.
Step 9 (Uniform Lipschitz bounds near compact subsets of D D D ). Let S ⊆ D S\subseteq D S ⊆ D be compact. There are real r ∈ ( 0 , 1 ] r\in(0,1] r ∈ ( 0 , 1 ] and L ≥ 0 L\ge0 L ≥ 0 such that for all x ∈ S x\in S x ∈ S and y ∈ R d y\in\mathbb{R}^{d} y ∈ R d with ∥ y − x ∥ ≤ r \lVert y-x\rVert\le r ∥ y − x ∥ ≤ r : y ∈ D y\in D y ∈ D , ∣ f ( y ) − f ( x ) ∣ ≤ L ∥ y − x ∥ |f(y)-f(x)|\le L\lVert y-x\rVert ∣ f ( y ) − f ( x ) ∣ ≤ L ∥ y − x ∥ and ∣ U ( y ) − U ( x ) ∣ ≤ L ∥ y − x ∥ |U(y)-U(x)|\le L\lVert y-x\rVert ∣ U ( y ) − U ( x ) ∣ ≤ L ∥ y − x ∥ . If S = ∅ S=\varnothing S = ∅ take r = 1 r=1 r = 1 , L = 0 L=0 L = 0 . Otherwise, for z ∈ S z\in S z ∈ S apply the local Lipschitz bound to f f f (semiconvex with constant K K K ) and to U U U (semiconvex with constant 0 0 0 , Step 2(a)) at z z z , and let ρ z > 0 \rho_{z}>0 ρ z > 0 be the smaller of the two radii and L z L_{z} L z the larger of the two constants; then B ˉ ( z , ρ z ) ⊆ D \bar{B}(z,\rho_{z})\subseteq D B ˉ ( z , ρ z ) ⊆ D and f f f , U U U are both Lipschitz with constant L z L_{z} L z on B ˉ ( z , ρ z ) \bar{B}(z,\rho_{z}) B ˉ ( z , ρ z ) . The open balls B ( z , ρ z / 2 ) B(z,\rho_{z}/2) B ( z , ρ z /2 ) , z ∈ S z\in S z ∈ S , cover the compact set S S S , so finitely many of them, with centres z 1 , … , z k z_{1},\dots,z_{k} z 1 , … , z k , cover S S S . Put r = min ( 1 , ρ z 1 / 2 , … , ρ z k / 2 ) r=\min(1,\rho_{z_{1}}/2,\dots,\rho_{z_{k}}/2) r = min ( 1 , ρ z 1 /2 , … , ρ z k /2 ) and L = max ( L z 1 , … , L z k ) L=\max(L_{z_{1}},\dots,L_{z_{k}}) L = max ( L z 1 , … , L z k ) . For x ∈ S x\in S x ∈ S pick j j j with ∥ x − z j ∥ < ρ z j / 2 \lVert x-z_{j}\rVert<\rho_{z_{j}}/2 ∥ x − z j ∥ < ρ z j /2 ; if ∥ y − x ∥ ≤ r \lVert y-x\rVert\le r ∥ y − x ∥ ≤ r then ∥ y − z j ∥ < ρ z j \lVert y-z_{j}\rVert<\rho_{z_{j}} ∥ y − z j ∥ < ρ z j by the triangle inequality, so x , y ∈ B ˉ ( z j , ρ z j ) ⊆ D x,y\in\bar{B}(z_{j},\rho_{z_{j}})\subseteq D x , y ∈ B ˉ ( z j , ρ z j ) ⊆ D and the two bounds hold.
Step 10 (Integration by parts against a cutoff of the density). Fix ψ ∈ C c ∞ ( R d ) \psi\in C_{c}^{\infty}(\mathbb{R}^{d}) ψ ∈ C c ∞ ( R d ) (test functions ). By The Gradient of a Test Function is Bounded and Square-Integrable, and Its Laplacian Bounded and Integrable, Against Every Probability Measure §gradient and The Gradient of a Test Function is Bounded and Square-Integrable, and Its Laplacian Bounded and Integrable, Against Every Probability Measure §laplacian there are K ψ , C Δ ≥ 0 K_{\psi},C_{\Delta}\ge0 K ψ , C Δ ≥ 0 with ∥ ∇ ψ ∥ ≤ K ψ \lVert\nabla\psi\rVert\le K_{\psi} ∥ ∇ ψ ∥ ≤ K ψ and ∣ Δ ψ ∣ ≤ C Δ |\Delta\psi|\le C_{\Delta} ∣Δ ψ ∣ ≤ C Δ everywhere, Δ ψ = ∑ i ∂ i ∂ i ψ \Delta\psi=\sum_{i}\partial_{i}\partial_{i}\psi Δ ψ = ∑ i ∂ i ∂ i ψ being Borel. For each i ∈ [ d ] i\in[d] i ∈ [ d ] , ∂ i ψ \partial_{i}\psi ∂ i ψ is continuous and λ d \lambda_{d} λ d -integrable by Integration by Parts on Euclidean Space Against a Compactly Supported Function of Class C 1 C^{1} C 1 §vanishing (with h = ψ h=\psi h = ψ ); and ∂ i ψ \partial_{i}\psi ∂ i ψ is of class C 1 C^{1} C 1 (ψ \psi ψ being of class C 2 C^{2} C 2 , Test Functions on Euclidean Space, Their Gradient Maps and Laplacians §gradient ) and compactly supported (The Gradient of a Test Function is Bounded and Square-Integrable, and Its Laplacian Bounded and Integrable, Against Every Probability Measure §gradient ), so by Integration by Parts on Euclidean Space Against a Compactly Supported Function of Class C 1 C^{1} C 1 §vanishing (with h = ∂ i ψ h=\partial_{i}\psi h = ∂ i ψ ) ∂ i ∂ i ψ \partial_{i}\partial_{i}\psi ∂ i ∂ i ψ is continuous and compactly supported, hence bounded by Extreme Value Theorem on a Compact Subset of a Metric Space applied on its support; fix C 2 ≥ 0 C_{2}\ge0 C 2 ≥ 0 with ∣ ∂ i ∂ i ψ ∣ ≤ C 2 |\partial_{i}\partial_{i}\psi|\le C_{2} ∣ ∂ i ∂ i ψ ∣ ≤ C 2 for every i i i . For x ∈ R d x\in\mathbb{R}^{d} x ∈ R d and t > 0 t>0 t > 0 , the function τ ↦ ∂ i ψ ( x + τ e i ) \tau\mapsto\partial_{i}\psi(x+\tau e_{i}) τ ↦ ∂ i ψ ( x + τ e i ) is continuous on [ 0 , t ] [0,t] [ 0 , t ] and, by A Real-Valued C^1 Function is Differentiable at Every Point and Chain Rule Along an Affine Path , differentiable with derivative ∂ i ∂ i ψ ( x + τ e i ) \partial_{i}\partial_{i}\psi(x+\tau e_{i}) ∂ i ∂ i ψ ( x + τ e i ) ; so Mean Value Theorem on a Closed Real Interval gives
∣ ∂ i ψ ( x + t e i ) − ∂ i ψ ( x ) ∣ ≤ C 2 t . (10.1) \bigl|\partial_{i}\psi(x+te_{i})-\partial_{i}\psi(x)\bigr|\le C_{2}\,t .\tag{10.1} ∂ i ψ ( x + t e i ) − ∂ i ψ ( x ) ≤ C 2 t . ( 10.1 )
Fix a smooth χ : R → R \chi:\mathbb{R}\to\mathbb{R} χ : R → R as in Scaled Cutoffs and the Second-Moment Test Functions: Uniform Derivative Bounds and Agreement on a Ball with q = 1 q=1 q = 1 (identifying R 1 \mathbb{R}^{1} R 1 with R \mathbb{R} R , so that ∥ t ∥ = ∣ t ∣ \lVert t\rVert=|t| ∥ t ∥ = ∣ t ∣ and ∂ 1 \partial_{1} ∂ 1 is the ordinary derivative, written with a prime); it exists by Existence of a Smooth Plateau Function on Euclidean Space . By Scaled Cutoffs and the Second-Moment Test Functions: Uniform Derivative Bounds and Agreement on a Ball §cutoff there is M 1 ≥ 0 M_{1}\ge0 M 1 ≥ 0 such that for every natural M ≥ 1 M\ge1 M ≥ 1 the function χ M ( t ) = χ ( t / M ) \chi_{M}(t)=\chi(t/M) χ M ( t ) = χ ( t / M ) is smooth, 0 ≤ χ M ≤ 1 0\le\chi_{M}\le1 0 ≤ χ M ≤ 1 , χ M ( t ) = 1 \chi_{M}(t)=1 χ M ( t ) = 1 for ∣ t ∣ ≤ M |t|\le M ∣ t ∣ ≤ M , χ M ( t ) = 0 \chi_{M}(t)=0 χ M ( t ) = 0 for ∣ t ∣ ≥ 2 M |t|\ge2M ∣ t ∣ ≥ 2 M , and ∣ χ M ′ ∣ ≤ M 1 / M |\chi_{M}'|\le M_{1}/M ∣ χ M ′ ∣ ≤ M 1 / M ; by Mean Value Theorem on a Closed Real Interval , ∣ χ M ( s ) − χ M ( t ) ∣ ≤ ( M 1 / M ) ∣ s − t ∣ |\chi_{M}(s)-\chi_{M}(t)|\le(M_{1}/M)|s-t| ∣ χ M ( s ) − χ M ( t ) ∣ ≤ ( M 1 / M ) ∣ s − t ∣ for all real s , t s,t s , t . Also, for real s , u s,u s , u ,
∣ e s − e u ∣ ≤ e max ( s , u ) ∣ s − u ∣ , (10.2) |e^{s}-e^{u}|\le e^{\max(s,u)}|s-u| ,\tag{10.2} ∣ e s − e u ∣ ≤ e m a x ( s , u ) ∣ s − u ∣ , ( 10.2 )
since for u ≤ s u\le s u ≤ s , The Function s log s s\log s s log s : Continuity, Young's Inequality and Lower Bounds, with the Elementary Bounds for the Exponential and the Logarithm §exp gives e u − s ≥ 1 + ( u − s ) e^{u-s}\ge1+(u-s) e u − s ≥ 1 + ( u − s ) , whence 0 ≤ e s − e u = e s ( 1 − e u − s ) ≤ e s ( s − u ) 0\le e^{s}-e^{u}=e^{s}(1-e^{u-s})\le e^{s}(s-u) 0 ≤ e s − e u = e s ( 1 − e u − s ) ≤ e s ( s − u ) by claims 1 and 4 of Basic Properties of the Exponential Function .
Fix a natural M ≥ 1 M\ge1 M ≥ 1 . Define ζ M , κ M : R d → R \zeta_{M},\kappa_{M}:\mathbb{R}^{d}\to\mathbb{R} ζ M , κ M : R d → R by ζ M ( x ) = χ M ( U ( x ) − p 0 ) \zeta_{M}(x)=\chi_{M}(U(x)-p_{0}) ζ M ( x ) = χ M ( U ( x ) − p 0 ) and κ M ( x ) = χ M ′ ( U ( x ) − p 0 ) \kappa_{M}(x)=\chi_{M}'(U(x)-p_{0}) κ M ( x ) = χ M ′ ( U ( x ) − p 0 ) for x ∈ D x\in D x ∈ D , and ζ M ( x ) = κ M ( x ) = 0 \zeta_{M}(x)=\kappa_{M}(x)=0 ζ M ( x ) = κ M ( x ) = 0 for x ∉ D x\notin D x ∈ / D ; they are Borel by (BC), χ M \chi_{M} χ M and χ M ′ \chi_{M}' χ M ′ being continuous. Put
w M = ζ M p , V M = p ( ζ M η + κ M ∇ U ) , w_{M}=\zeta_{M}\,p,\qquad V_{M}=p\,\bigl(\zeta_{M}\,\eta+\kappa_{M}\,\nabla U\bigr), w M = ζ M p , V M = p ( ζ M η + κ M ∇ U ) ,
Borel, with 0 ≤ w M ≤ p ≤ Z − 1 e H 0\le w_{M}\le p\le Z^{-1}e^{H} 0 ≤ w M ≤ p ≤ Z − 1 e H by (5.1). Let S M = { x ∈ D : U ( x ) ≤ p 0 + 2 M } S_{M}=\{x\in D:U(x)\le p_{0}+2M\} S M = { x ∈ D : U ( x ) ≤ p 0 + 2 M } , compact by Penalty on an Open Subset of Euclidean Space §sublevel ; since ζ M ( x ) = 0 \zeta_{M}(x)=0 ζ M ( x ) = 0 when x ∈ D x\in D x ∈ D and U ( x ) − p 0 ≥ 2 M U(x)-p_{0}\ge2M U ( x ) − p 0 ≥ 2 M , w M w_{M} w M vanishes off S M S_{M} S M . Let C M C_{M} C M be the closure of { x ∈ D : U ( x ) < p 0 + 2 M + 1 } ⊇ S M \{x\in D:U(x)<p_{0}+2M+1\}\supseteq S_{M} { x ∈ D : U ( x ) < p 0 + 2 M + 1 } ⊇ S M ; by Basic Properties of the Sublevel Sets of a Penalty §sublevel-sets , C M ⊆ D C_{M}\subseteq D C M ⊆ D , and C M C_{M} C M is closed. Apply Step 9 to S = S M S=S_{M} S = S M , obtaining r ∈ ( 0 , 1 ] r\in(0,1] r ∈ ( 0 , 1 ] and L L L , and put L M = Z − 1 e H L ( 2 / a + M 1 / M ) L_{M}=Z^{-1}e^{H}L\,(2/a+M_{1}/M) L M = Z − 1 e H L ( 2/ a + M 1 / M ) .
(10a) For i ∈ [ d ] i\in[d] i ∈ [ d ] , t ∈ ( 0 , r ] t\in(0,r] t ∈ ( 0 , r ] and x ∈ R d x\in\mathbb{R}^{d} x ∈ R d : ∣ w M ( x − t e i ) − w M ( x ) ∣ ≤ L M t |w_{M}(x-te_{i})-w_{M}(x)|\le L_{M}\,t ∣ w M ( x − t e i ) − w M ( x ) ∣ ≤ L M t . First let x ′ ∈ S M x'\in S_{M} x ′ ∈ S M and ∥ y ′ − x ′ ∥ ≤ r \lVert y'-x'\rVert\le r ∥ y ′ − x ′ ∥ ≤ r ; then x ′ , y ′ ∈ D x',y'\in D x ′ , y ′ ∈ D and w M ( y ′ ) − w M ( x ′ ) = ζ M ( y ′ ) ( p ( y ′ ) − p ( x ′ ) ) + p ( x ′ ) ( ζ M ( y ′ ) − ζ M ( x ′ ) ) w_{M}(y')-w_{M}(x')=\zeta_{M}(y')\bigl(p(y')-p(x')\bigr)+p(x')\bigl(\zeta_{M}(y')-\zeta_{M}(x')\bigr) w M ( y ′ ) − w M ( x ′ ) = ζ M ( y ′ ) ( p ( y ′ ) − p ( x ′ ) ) + p ( x ′ ) ( ζ M ( y ′ ) − ζ M ( x ′ ) ) . By (5.1), p = Z − 1 e h p=Z^{-1}e^{h} p = Z − 1 e h on D D D , so (10.2), h ≤ H h\le H h ≤ H and Step 9 give ∣ p ( y ′ ) − p ( x ′ ) ∣ ≤ Z − 1 e H a − 1 ( ∣ f ( y ′ ) − f ( x ′ ) ∣ + ∣ U ( y ′ ) − U ( x ′ ) ∣ ) ≤ Z − 1 e H ( 2 L / a ) ∥ y ′ − x ′ ∥ |p(y')-p(x')|\le Z^{-1}e^{H}a^{-1}\bigl(|f(y')-f(x')|+|U(y')-U(x')|\bigr)\le Z^{-1}e^{H}(2L/a)\lVert y'-x'\rVert ∣ p ( y ′ ) − p ( x ′ ) ∣ ≤ Z − 1 e H a − 1 ( ∣ f ( y ′ ) − f ( x ′ ) ∣ + ∣ U ( y ′ ) − U ( x ′ ) ∣ ) ≤ Z − 1 e H ( 2 L / a ) ∥ y ′ − x ′ ∥ , while ∣ ζ M ( y ′ ) − ζ M ( x ′ ) ∣ ≤ ( M 1 / M ) ∣ U ( y ′ ) − U ( x ′ ) ∣ ≤ ( M 1 L / M ) ∥ y ′ − x ′ ∥ |\zeta_{M}(y')-\zeta_{M}(x')|\le(M_{1}/M)|U(y')-U(x')|\le(M_{1}L/M)\lVert y'-x'\rVert ∣ ζ M ( y ′ ) − ζ M ( x ′ ) ∣ ≤ ( M 1 / M ) ∣ U ( y ′ ) − U ( x ′ ) ∣ ≤ ( M 1 L / M ) ∥ y ′ − x ′ ∥ . With 0 ≤ ζ M ≤ 1 0\le\zeta_{M}\le1 0 ≤ ζ M ≤ 1 and p ( x ′ ) ≤ Z − 1 e H p(x')\le Z^{-1}e^{H} p ( x ′ ) ≤ Z − 1 e H this gives ∣ w M ( y ′ ) − w M ( x ′ ) ∣ ≤ L M ∥ y ′ − x ′ ∥ |w_{M}(y')-w_{M}(x')|\le L_{M}\lVert y'-x'\rVert ∣ w M ( y ′ ) − w M ( x ′ ) ∣ ≤ L M ∥ y ′ − x ′ ∥ . Now let y = x − t e i y=x-te_{i} y = x − t e i , so ∥ y − x ∥ = t ≤ r \lVert y-x\rVert=t\le r ∥ y − x ∥ = t ≤ r : if x ∈ S M x\in S_{M} x ∈ S M use ( x ′ , y ′ ) = ( x , y ) (x',y')=(x,y) ( x ′ , y ′ ) = ( x , y ) ; if y ∈ S M y\in S_{M} y ∈ S M use ( x ′ , y ′ ) = ( y , x ) (x',y')=(y,x) ( x ′ , y ′ ) = ( y , x ) ; otherwise w M ( x ) = w M ( y ) = 0 w_{M}(x)=w_{M}(y)=0 w M ( x ) = w M ( y ) = 0 .
(10b) For x ∈ D ∖ N 1 x\in D\setminus N_{1} x ∈ D ∖ N 1 and i ∈ [ d ] i\in[d] i ∈ [ d ] the partial derivative ∂ i w M ( x ) \partial_{i}w_{M}(x) ∂ i w M ( x ) exists and equals the i i i th component ( V M ) i ( x ) (V_{M})_{i}(x) ( V M ) i ( x ) . Let Θ : R 2 → R \Theta:\mathbb{R}^{2}\to\mathbb{R} Θ : R 2 → R , Θ ( s , u ) = Z − 1 χ M ( u − p 0 ) exp ( ( s − u ) / a ) \Theta(s,u)=Z^{-1}\chi_{M}(u-p_{0})\exp\bigl((s-u)/a\bigr) Θ ( s , u ) = Z − 1 χ M ( u − p 0 ) exp ( ( s − u ) / a ) . By Sum, Constant Multiple, and Product Rules for One-Dimensional Derivatives , Chain Rule for One-Dimensional Derivatives and claim 3 of Basic Properties of the Exponential Function , its partial derivatives are ∂ s Θ = a − 1 Θ \partial_{s}\Theta=a^{-1}\Theta ∂ s Θ = a − 1 Θ and ∂ u Θ ( s , u ) = Z − 1 χ M ′ ( u − p 0 ) exp ( ( s − u ) / a ) − a − 1 Θ ( s , u ) \partial_{u}\Theta(s,u)=Z^{-1}\chi_{M}'(u-p_{0})\exp((s-u)/a)-a^{-1}\Theta(s,u) ∂ u Θ ( s , u ) = Z − 1 χ M ′ ( u − p 0 ) exp (( s − u ) / a ) − a − 1 Θ ( s , u ) , which are continuous; so Θ \Theta Θ is of class C 1 C^{1} C 1 (C^k Maps on a Euclidean Open Set ) and, by A Real-Valued C^1 Function is Differentiable at Every Point , differentiable at every point with derivative matrix ( ∂ s Θ , ∂ u Θ ) (\partial_{s}\Theta,\partial_{u}\Theta) ( ∂ s Θ , ∂ u Θ ) . The map Φ = ( f , U ) : D → R 2 \Phi=(f,U):D\to\mathbb{R}^{2} Φ = ( f , U ) : D → R 2 is differentiable at x x x with derivative matrix having rows D f ( x ) Df(x) D f ( x ) and D U ( x ) DU(x) D U ( x ) : f f f is differentiable at x x x because x ∉ N x\notin N x ∈ / N , U U U by A Real-Valued C^1 Function is Differentiable at Every Point , and for a given ε \varepsilon ε one takes the smaller of the two δ \delta δ 's for ε / 2 \varepsilon/2 ε /2 , using ∥ ( α , β ) ∥ ≤ ∣ α ∣ + ∣ β ∣ \lVert(\alpha,\beta)\rVert\le|\alpha|+|\beta| ∥( α , β )∥ ≤ ∣ α ∣ + ∣ β ∣ . On D D D we have w M = Θ ∘ Φ w_{M}=\Theta\circ\Phi w M = Θ ∘ Φ , since Θ ( Φ ( y ) ) = Z − 1 χ M ( U ( y ) − p 0 ) e h ( y ) = ζ M ( y ) p ( y ) \Theta(\Phi(y))=Z^{-1}\chi_{M}(U(y)-p_{0})e^{h(y)}=\zeta_{M}(y)p(y) Θ ( Φ ( y )) = Z − 1 χ M ( U ( y ) − p 0 ) e h ( y ) = ζ M ( y ) p ( y ) . By Chain Rule for Differentiable Maps Between Euclidean Spaces , Θ ∘ Φ \Theta\circ\Phi Θ ∘ Φ is differentiable at x x x with derivative matrix ∂ s Θ ( Φ ( x ) ) D f ( x ) + ∂ u Θ ( Φ ( x ) ) D U ( x ) \partial_{s}\Theta(\Phi(x))\,Df(x)+\partial_{u}\Theta(\Phi(x))\,DU(x) ∂ s Θ ( Φ ( x )) D f ( x ) + ∂ u Θ ( Φ ( x )) D U ( x ) , and by claim 1 of A Derivative Matrix is the Jacobian Matrix, and is Unique (and since D D D is open, so that w M w_{M} w M and Θ ∘ Φ \Theta\circ\Phi Θ ∘ Φ have the same difference quotients at x x x for small steps)
∂ i w M ( x ) = w M ( x ) a ( ∂ i f ( x ) − ∂ i U ( x ) ) + p ( x ) κ M ( x ) ∂ i U ( x ) = p ( x ) ( ζ M ( x ) η i ( x ) + κ M ( x ) ( ∇ U ) i ( x ) ) , \partial_{i}w_{M}(x)=\frac{w_{M}(x)}{a}\bigl(\partial_{i}f(x)-\partial_{i}U(x)\bigr)+p(x)\kappa_{M}(x)\,\partial_{i}U(x)=p(x)\bigl(\zeta_{M}(x)\eta_{i}(x)+\kappa_{M}(x)(\nabla U)_{i}(x)\bigr), ∂ i w M ( x ) = a w M ( x ) ( ∂ i f ( x ) − ∂ i U ( x ) ) + p ( x ) κ M ( x ) ∂ i U ( x ) = p ( x ) ( ζ M ( x ) η i ( x ) + κ M ( x ) ( ∇ U ) i ( x ) ) ,
using ∇ f ( x ) = D f ( x ) \nabla f(x)=Df(x) ∇ f ( x ) = D f ( x ) and ∇ U ( x ) = D U ( x ) \nabla U(x)=DU(x) ∇ U ( x ) = D U ( x ) .
(10c) For i ∈ [ d ] i\in[d] i ∈ [ d ] and t > 0 t>0 t > 0 ,
∫ w M ( x ) ( ∂ i ψ ( x + t e i ) − ∂ i ψ ( x ) ) d λ d ( x ) = ∫ ( w M ( y − t e i ) − w M ( y ) ) ∂ i ψ ( y ) d λ d ( y ) , \int w_{M}(x)\bigl(\partial_{i}\psi(x+te_{i})-\partial_{i}\psi(x)\bigr)\,d\lambda_{d}(x)=\int\bigl(w_{M}(y-te_{i})-w_{M}(y)\bigr)\partial_{i}\psi(y)\,d\lambda_{d}(y), ∫ w M ( x ) ( ∂ i ψ ( x + t e i ) − ∂ i ψ ( x ) ) d λ d ( x ) = ∫ ( w M ( y − t e i ) − w M ( y ) ) ∂ i ψ ( y ) d λ d ( y ) ,
all four integrands involved being λ d \lambda_{d} λ d -integrable. Let T ( x ) = x + t e i T(x)=x+te_{i} T ( x ) = x + t e i . Its components are of class C 1 C^{1} C 1 with Jacobian matrix the identity matrix I d I_{d} I d at every point, which is symmetric and positive definite (v ⋅ I d v = ∥ v ∥ 2 > 0 v\cdot I_{d}v=\lVert v\rVert^{2}>0 v ⋅ I d v = ∥ v ∥ 2 > 0 for v ≠ 0 v\ne0 v = 0 ); T T T is a bijection with inverse T − 1 ( y ) = y − t e i T^{-1}(y)=y-te_{i} T − 1 ( y ) = y − t e i , whose Jacobian matrix I d I_{d} I d is the inverse matrix of I d I_{d} I d by The Identity Matrix is a Two-Sided Multiplicative Identity ; and det I d = 1 \det I_{d}=1 det I d = 1 by The Determinant of a Triangular Matrix is the Product of its Diagonal Entries . The function Q ( y ) = w M ( y − t e i ) ∂ i ψ ( y ) Q(y)=w_{M}(y-te_{i})\,\partial_{i}\psi(y) Q ( y ) = w M ( y − t e i ) ∂ i ψ ( y ) is Borel (T − 1 T^{-1} T − 1 being Borel by Change of Variables for the Lebesgue Integral under a Continuously Differentiable Bijection of Euclidean Space with Symmetric Positive Definite Jacobian Matrix, and the Density of a Push-Forward §regularity ) with ∣ Q ∣ ≤ Z − 1 e H ∣ ∂ i ψ ∣ |Q|\le Z^{-1}e^{H}|\partial_{i}\psi| ∣ Q ∣ ≤ Z − 1 e H ∣ ∂ i ψ ∣ , hence integrable, and Q ∘ T ( x ) = w M ( x ) ∂ i ψ ( x + t e i ) Q\circ T(x)=w_{M}(x)\partial_{i}\psi(x+te_{i}) Q ∘ T ( x ) = w M ( x ) ∂ i ψ ( x + t e i ) . By Change of Variables for the Lebesgue Integral under a Continuously Differentiable Bijection of Euclidean Space with Symmetric Positive Definite Jacobian Matrix, and the Density of a Push-Forward §integrals , Q ∘ T Q\circ T Q ∘ T is integrable and ∫ w M ( x ) ∂ i ψ ( x + t e i ) d λ d ( x ) = ∫ Q d λ d \int w_{M}(x)\partial_{i}\psi(x+te_{i})\,d\lambda_{d}(x)=\int Q\,d\lambda_{d} ∫ w M ( x ) ∂ i ψ ( x + t e i ) d λ d ( x ) = ∫ Q d λ d . Subtracting the integral of the integrable function w M ∂ i ψ w_{M}\partial_{i}\psi w M ∂ i ψ (bounded by Z − 1 e H ∣ ∂ i ψ ∣ Z^{-1}e^{H}|\partial_{i}\psi| Z − 1 e H ∣ ∂ i ψ ∣ ) from both sides gives the identity, by claim 2 of Linearity and Monotonicity of the Lebesgue Integral .
(10d) For i ∈ [ d ] i\in[d] i ∈ [ d ] : ∫ w M ∂ i ∂ i ψ d λ d = − ∫ ( V M ) i ∂ i ψ d λ d \int w_{M}\,\partial_{i}\partial_{i}\psi\,d\lambda_{d}=-\int(V_{M})_{i}\,\partial_{i}\psi\,d\lambda_{d} ∫ w M ∂ i ∂ i ψ d λ d = − ∫ ( V M ) i ∂ i ψ d λ d . Fix a natural m 0 ≥ 1 / r m_{0}\ge1/r m 0 ≥ 1/ r , and for natural m ≥ m 0 m\ge m_{0} m ≥ m 0 put
L m ( x ) = m w M ( x ) ( ∂ i ψ ( x + m − 1 e i ) − ∂ i ψ ( x ) ) , R m ( y ) = m ( w M ( y − m − 1 e i ) − w M ( y ) ) ∂ i ψ ( y ) , \mathrm{L}_{m}(x)=m\,w_{M}(x)\bigl(\partial_{i}\psi(x+m^{-1}e_{i})-\partial_{i}\psi(x)\bigr),\qquad \mathrm{R}_{m}(y)=m\bigl(w_{M}(y-m^{-1}e_{i})-w_{M}(y)\bigr)\partial_{i}\psi(y), L m ( x ) = m w M ( x ) ( ∂ i ψ ( x + m − 1 e i ) − ∂ i ψ ( x ) ) , R m ( y ) = m ( w M ( y − m − 1 e i ) − w M ( y ) ) ∂ i ψ ( y ) ,
Borel functions with ∫ L m d λ d = ∫ R m d λ d \int\mathrm{L}_{m}\,d\lambda_{d}=\int\mathrm{R}_{m}\,d\lambda_{d} ∫ L m d λ d = ∫ R m d λ d by (10c) with t = 1 / m t=1/m t = 1/ m . At every x x x , L m ( x ) → w M ( x ) ∂ i ∂ i ψ ( x ) \mathrm{L}_{m}(x)\to w_{M}(x)\partial_{i}\partial_{i}\psi(x) L m ( x ) → w M ( x ) ∂ i ∂ i ψ ( x ) by Partial Derivative on a Euclidean Open Set , and ∣ L m ∣ ≤ C 2 p |\mathrm{L}_{m}|\le C_{2}\,p ∣ L m ∣ ≤ C 2 p by (10.1), p p p being integrable; so Dominated Convergence Theorem gives ∫ L m d λ d → ∫ w M ∂ i ∂ i ψ d λ d \int\mathrm{L}_{m}\,d\lambda_{d}\to\int w_{M}\partial_{i}\partial_{i}\psi\,d\lambda_{d} ∫ L m d λ d → ∫ w M ∂ i ∂ i ψ d λ d . By (10a) with t = 1 / m ≤ r t=1/m\le r t = 1/ m ≤ r , ∣ R m ∣ ≤ L M ∣ ∂ i ψ ∣ |\mathrm{R}_{m}|\le L_{M}|\partial_{i}\psi| ∣ R m ∣ ≤ L M ∣ ∂ i ψ ∣ , an integrable function. For y ∈ D ∖ N 1 y\in D\setminus N_{1} y ∈ D ∖ N 1 , R m ( y ) = − w M ( y + ( − m − 1 ) e i ) − w M ( y ) − m − 1 ∂ i ψ ( y ) → − ( V M ) i ( y ) ∂ i ψ ( y ) \mathrm{R}_{m}(y)=-\frac{w_{M}(y+(-m^{-1})e_{i})-w_{M}(y)}{-m^{-1}}\,\partial_{i}\psi(y)\to-(V_{M})_{i}(y)\partial_{i}\psi(y) R m ( y ) = − − m − 1 w M ( y + ( − m − 1 ) e i ) − w M ( y ) ∂ i ψ ( y ) → − ( V M ) i ( y ) ∂ i ψ ( y ) by (10b). For y ∉ D y\notin D y ∈ / D , y ∉ C M y\notin C_{M} y ∈ / C M ; as R d ∖ C M \mathbb{R}^{d}\setminus C_{M} R d ∖ C M is open there is ρ > 0 \rho>0 ρ > 0 with B ( y , ρ ) ∩ C M = ∅ B(y,\rho)\cap C_{M}=\varnothing B ( y , ρ ) ∩ C M = ∅ , so for m > 1 / ρ m>1/\rho m > 1/ ρ neither y y y nor y − m − 1 e i y-m^{-1}e_{i} y − m − 1 e i lies in S M ⊆ C M S_{M}\subseteq C_{M} S M ⊆ C M and R m ( y ) = 0 = − ( V M ) i ( y ) ∂ i ψ ( y ) \mathrm{R}_{m}(y)=0=-(V_{M})_{i}(y)\partial_{i}\psi(y) R m ( y ) = 0 = − ( V M ) i ( y ) ∂ i ψ ( y ) , as p ( y ) = 0 p(y)=0 p ( y ) = 0 . So R m → − ( V M ) i ∂ i ψ \mathrm{R}_{m}\to-(V_{M})_{i}\partial_{i}\psi R m → − ( V M ) i ∂ i ψ outside the null set N 1 N_{1} N 1 , and The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §dominated gives ∫ R m d λ d → − ∫ ( V M ) i ∂ i ψ d λ d \int\mathrm{R}_{m}\,d\lambda_{d}\to-\int(V_{M})_{i}\partial_{i}\psi\,d\lambda_{d} ∫ R m d λ d → − ∫ ( V M ) i ∂ i ψ d λ d . The two limits coincide.
Summing (10d) over i ∈ [ d ] i\in[d] i ∈ [ d ] and using Δ ψ = ∑ i ∂ i ∂ i ψ \Delta\psi=\sum_{i}\partial_{i}\partial_{i}\psi Δ ψ = ∑ i ∂ i ∂ i ψ and claim 2 of Linearity and Monotonicity of the Lebesgue Integral ,
∫ w M Δ ψ d λ d = − ∫ V M ⋅ ∇ ψ d λ d ( M ≥ 1 ) . (10.3) \int w_{M}\,\Delta\psi\,d\lambda_{d}=-\int V_{M}\cdot\nabla\psi\,d\lambda_{d}\qquad(M\ge1).\tag{10.3} ∫ w M Δ ψ d λ d = − ∫ V M ⋅ ∇ ψ d λ d ( M ≥ 1 ) . ( 10.3 )
Step 11 (Removing the cutoff). Let M → ∞ M\to\infty M → ∞ through the natural numbers. For x ∈ D x\in D x ∈ D and M ≥ U ( x ) − p 0 M\ge U(x)-p_{0} M ≥ U ( x ) − p 0 we have ζ M ( x ) = 1 \zeta_{M}(x)=1 ζ M ( x ) = 1 , while ∣ κ M ( x ) ∣ ≤ M 1 / M → 0 |\kappa_{M}(x)|\le M_{1}/M\to0 ∣ κ M ( x ) ∣ ≤ M 1 / M → 0 ; for x ∉ D x\notin D x ∈ / D all of w M ( x ) w_{M}(x) w M ( x ) , V M ( x ) V_{M}(x) V M ( x ) , p ( x ) p(x) p ( x ) vanish. Hence w M Δ ψ → p Δ ψ w_{M}\Delta\psi\to p\,\Delta\psi w M Δ ψ → p Δ ψ and V M ⋅ ∇ ψ → p η ⋅ ∇ ψ V_{M}\cdot\nabla\psi\to p\,\eta\cdot\nabla\psi V M ⋅ ∇ ψ → p η ⋅ ∇ ψ at every point, with ∣ w M Δ ψ ∣ ≤ C Δ p |w_{M}\Delta\psi|\le C_{\Delta}\,p ∣ w M Δ ψ ∣ ≤ C Δ p and ∣ V M ⋅ ∇ ψ ∣ ≤ K ψ p ( ∥ η ∥ + M 1 ∥ ∇ U ∥ ) |V_{M}\cdot\nabla\psi|\le K_{\psi}\,p\,\bigl(\lVert\eta\rVert+M_{1}\lVert\nabla U\rVert\bigr) ∣ V M ⋅ ∇ ψ ∣ ≤ K ψ p ( ∥ η ∥ + M 1 ∥ ∇ U ∥ ) , integrable by Step 8. By Dominated Convergence Theorem and (10.3), ∫ p Δ ψ d λ d = − ∫ p η ⋅ ∇ ψ d λ d \int p\,\Delta\psi\,d\lambda_{d}=-\int p\,\eta\cdot\nabla\psi\,d\lambda_{d} ∫ p Δ ψ d λ d = − ∫ p η ⋅ ∇ ψ d λ d . As Δ ψ \Delta\psi Δ ψ is bounded Borel and ∣ η ⋅ ∇ ψ ∣ ≤ K ψ ∥ η ∥ |\eta\cdot\nabla\psi|\le K_{\psi}\lVert\eta\rVert ∣ η ⋅ ∇ ψ ∣ ≤ K ψ ∥ η ∥ , both Δ ψ \Delta\psi Δ ψ and η ⋅ ∇ ψ \eta\cdot\nabla\psi η ⋅ ∇ ψ are π f \pi_{f} π f -integrable and claim 3 of Image Measures, Measures with Densities, and Change of Variables turns this into
⟨ η , ∇ ψ ⟩ π f = ∫ η ⋅ ∇ ψ d π f = − ∫ Δ ψ d π f for every ψ ∈ C c ∞ ( R d ) . (11.1) \langle\eta,\nabla\psi\rangle_{\pi_{f}}=\int\eta\cdot\nabla\psi\,d\pi_{f}=-\int\Delta\psi\,d\pi_{f}\qquad\text{for every }\psi\in C_{c}^{\infty}(\mathbb{R}^{d}).\tag{11.1} ⟨ η , ∇ ψ ⟩ π f = ∫ η ⋅ ∇ ψ d π f = − ∫ Δ ψ d π f for every ψ ∈ C c ∞ ( R d ) . ( 11.1 )
Step 12 (Tangency). Let T π f T_{\pi_{f}} T π f be the tangent space at π f ∈ P 2 ( R d ) \pi_{f}\in\mathcal{P}_{2}(\mathbb{R}^{d}) π f ∈ P 2 ( R d ) , a linear subspace of L 2 ( π f ; R d ) L^{2}(\pi_{f};\mathbb{R}^{d}) L 2 ( π f ; R d ) by Basic Properties of the Tangent Space: Closed Subspace, the Identity Map Belongs to It, Second-Moment Limits, and Representation of Bounded Functionals on Gradients §closed . Apply A Square-Integrable Selection of the Subdifferential of a Convex Potential Belongs to the Tangent Space §tangent with G = D G=D G = D (open, convex), the convex function U U U , the Borel set D D D of π f \pi_{f} π f -measure 1 1 1 and T = ∇ U T=\nabla U T = ∇ U : for every x ∈ D x\in D x ∈ D , ∂ D U ( x ) = { D U ( x ) } = { ∇ U ( x ) } \partial_{D}U(x)=\{DU(x)\}=\{\nabla U(x)\} ∂ D U ( x ) = { D U ( x )} = { ∇ U ( x )} by Step 2(b), and ∫ ∥ ∇ U ∥ 2 d π f < ∞ \int\lVert\nabla U\rVert^{2}\,d\pi_{f}<\infty ∫ ∥ ∇ U ∥ 2 d π f < ∞ ; so ∇ U ∈ T π f \nabla U\in T_{\pi_{f}} ∇ U ∈ T π f . Apply it again with G = D G=D G = D , the convex function g g g of Step 2(c), the Borel set D ∖ N 1 D\setminus N_{1} D ∖ N 1 of π f \pi_{f} π f -measure 1 1 1 (Step 8) and T = T g T=T_{g} T = T g : for every x ∈ D ∖ N 1 x\in D\setminus N_{1} x ∈ D ∖ N 1 , ∂ D g ( x ) = { D f ( x ) + K x } = { T g ( x ) } \partial_{D}g(x)=\{Df(x)+Kx\}=\{T_{g}(x)\} ∂ D g ( x ) = { D f ( x ) + K x } = { T g ( x )} by Step 2(c), and ∫ ∥ T g ∥ 2 d π f < ∞ \int\lVert T_{g}\rVert^{2}\,d\pi_{f}<\infty ∫ ∥ T g ∥ 2 d π f < ∞ ; so T g ∈ T π f T_{g}\in T_{\pi_{f}} T g ∈ T π f . By Basic Properties of the Tangent Space: Closed Subspace, the Identity Map Belongs to It, Second-Moment Limits, and Representation of Bounded Functionals on Gradients §identity , i d ∈ T π f \mathrm{id}\in T_{\pi_{f}} id ∈ T π f . Since T π f T_{\pi_{f}} T π f is a linear subspace, ∇ f = T g − K i d ∈ T π f \nabla f=T_{g}-K\,\mathrm{id}\in T_{\pi_{f}} ∇ f = T g − K id ∈ T π f and η = a − 1 ( ∇ f − ∇ U ) ∈ T π f \eta=a^{-1}(\nabla f-\nabla U)\in T_{\pi_{f}} η = a − 1 ( ∇ f − ∇ U ) ∈ T π f .
Step 13 (Clause 3). Gradient maps of f f f exist by Step 3. By Step 6, π f ∈ D U , a ⊆ P 2 ( R d ) \pi_{f}\in\mathcal{D}_{U,a}\subseteq\mathcal{P}_{2}(\mathbb{R}^{d}) π f ∈ D U , a ⊆ P 2 ( R d ) . For ψ ∈ C c ∞ ( R d ) \psi\in C_{c}^{\infty}(\mathbb{R}^{d}) ψ ∈ C c ∞ ( R d ) , (11.1) and The Cauchy-Schwarz Inequality in a Real Inner Product Space in the real Hilbert space L 2 ( π f ; R d ) L^{2}(\pi_{f};\mathbb{R}^{d}) L 2 ( π f ; R d ) (Square-Integrable Vector Fields Against a Probability Measure on Euclidean Space, and Test Functions: Standing Notation §l2mu ) give ∣ ∫ Δ ψ d π f ∣ ≤ ∥ η ∥ π f ∥ ∇ ψ ∥ π f \bigl|\int\Delta\psi\,d\pi_{f}\bigr|\le\lVert\eta\rVert_{\pi_{f}}\lVert\nabla\psi\rVert_{\pi_{f}} ∫ Δ ψ d π f ≤ ∥ η ∥ π f ∥ ∇ ψ ∥ π f ; so π f \pi_{f} π f has finite Fisher information with C = ∥ η ∥ π f C=\lVert\eta\rVert_{\pi_{f}} C = ∥ η ∥ π f , i.e. π f ∈ P 2 I ( R d ) \pi_{f}\in\mathcal{P}_{2}^{\mathcal{I}}(\mathbb{R}^{d}) π f ∈ P 2 I ( R d ) . By Step 12, η ∈ T π f \eta\in T_{\pi_{f}} η ∈ T π f , and by (11.1) it satisfies the identity defining the score ; by the uniqueness asserted there, ξ π f = η = a − 1 ( ∇ f − ∇ U ) \xi_{\pi_{f}}=\eta=a^{-1}(\nabla f-\nabla U) ξ π f = η = a − 1 ( ∇ f − ∇ U ) . Since moreover ∫ ∥ ∇ U ∥ 2 d π f < ∞ \int\lVert\nabla U\rVert^{2}\,d\pi_{f}<\infty ∫ ∥ ∇ U ∥ 2 d π f < ∞ (Step 8), π f ∈ D U , a Σ \pi_{f}\in\mathcal{D}^{\Sigma}_{U,a} π f ∈ D U , a Σ by The Relative Free Energy and the Relative Score of a Probability Measure for a Potential on an Open Set §score , and in L 2 ( π f ; R d ) L^{2}(\pi_{f};\mathbb{R}^{d}) L 2 ( π f ; R d )
Σ U , a ( π f ) = ∇ U + a ξ π f = ∇ U + ( ∇ f − ∇ U ) = ∇ f . \Sigma_{U,a}(\pi_{f})=\nabla U+a\,\xi_{\pi_{f}}=\nabla U+(\nabla f-\nabla U)=\nabla f . Σ U , a ( π f ) = ∇ U + a ξ π f = ∇ U + ( ∇ f − ∇ U ) = ∇ f .
Finally ∫ ∥ ∇ f ∥ 2 d π f < ∞ \int\lVert\nabla f\rVert^{2}\,d\pi_{f}<\infty ∫ ∥ ∇ f ∥ 2 d π f < ∞ by Step 8. As the gradient map ∇ f \nabla f ∇ f was arbitrary, clause 3 is proved.