Each result cited is universally quantified over the data in its own statement.
Real arithmetic, the order of R \mathbb{R} R and absolute values are those of The Real Numbers: Standing Notation and Background §background , used without further mention. Every square root is the nonnegative one of Existence and Uniqueness of the Nonnegative Square Root , and for nonnegative reals s , t s,t s , t we use that s ≤ t s\le t s ≤ t if and only if s 2 ≤ t 2 s^{2}\le t^{2} s 2 ≤ t 2 , the weak form (claim 2) of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field ; in particular ( s + t ) 1 / 2 ≤ s 1 / 2 + t 1 / 2 (s+t)^{1/2}\le s^{1/2}+t^{1/2} ( s + t ) 1/2 ≤ s 1/2 + t 1/2 , since ( s 1 / 2 + t 1 / 2 ) 2 ≥ s + t (s^{1/2}+t^{1/2})^{2}\ge s+t ( s 1/2 + t 1/2 ) 2 ≥ s + t . For real s , t s,t s , t we also use s t ≤ 1 4 s 2 + t 2 st\le\tfrac14s^{2}+t^{2} s t ≤ 4 1 s 2 + t 2 , ( s + t ) 2 ≤ 2 s 2 + 2 t 2 (s+t)^{2}\le2s^{2}+2t^{2} ( s + t ) 2 ≤ 2 s 2 + 2 t 2 and ∣ s ∣ ≤ 1 + s 2 |s|\le1+s^{2} ∣ s ∣ ≤ 1 + s 2 , which follow from ( 1 2 s − t ) 2 ≥ 0 (\tfrac12s-t)^{2}\ge0 ( 2 1 s − t ) 2 ≥ 0 , ( s − t ) 2 ≥ 0 (s-t)^{2}\ge0 ( s − t ) 2 ≥ 0 and ( ∣ s ∣ − 1 ) 2 ≥ 0 (|s|-1)^{2}\ge0 ( ∣ s ∣ − 1 ) 2 ≥ 0 .
Data. Write V = v ∘ p d V=v\circ p_{d} V = v ∘ p d with head dimension d d d , profile v v v and semiconvexity constant K ≥ 0 K\ge0 K ≥ 0 , as in Admissible Cylindrical Potentials on a Hilbert Space in the Noise Geometry §admissible . Fix b b b as in Admissible Cylindrical Potentials on a Hilbert Space in the Noise Geometry §below and C C C as in Admissible Cylindrical Potentials on a Hilbert Space in the Noise Geometry §slope ; by Basic Properties of an Admissible Cylindrical Potential: Continuity, Growth under Noise Translations, Integrability, the Tangent Inequality and Tangency of the Noise Gradient §translation , C ≥ 0 C\ge0 C ≥ 0 , and L = C 1 / 2 ( ∣ b + 1 ∣ + 2 ) ≥ 0 L=C^{1/2}(|b+1|+2)\ge0 L = C 1/2 ( ∣ b + 1∣ + 2 ) ≥ 0 is the number defined there. Fix C ′ C' C ′ as in Admissible Cylindrical Potentials on a Hilbert Space in the Noise Geometry §curvature for ε = ( 2 β ) − 1 \varepsilon=(2\beta)^{-1} ε = ( 2 β ) − 1 , and C ′ ′ C'' C ′′ as there for ε = 1 \varepsilon=1 ε = 1 . For u ∈ R d u\in\mathbb{R}^{d} u ∈ R d put ∣ ∇ a v ( u ) ∣ 2 = ∑ k = 1 d a k ( ∂ k v ( u ) ) 2 ≥ 0 |\nabla_{a}v(u)|^{2}=\sum_{k=1}^{d}a_{k}(\partial_{k}v(u))^{2}\ge0 ∣ ∇ a v ( u ) ∣ 2 = ∑ k = 1 d a k ( ∂ k v ( u ) ) 2 ≥ 0 , with square root ∣ ∇ a v ( u ) ∣ |\nabla_{a}v(u)| ∣ ∇ a v ( u ) ∣ ; ∥ u ∥ \lVert u\rVert ∥ u ∥ is the Euclidean norm, with ∥ u ∥ 2 = ∑ k = 1 d u k 2 \lVert u\rVert^{2}=\sum_{k=1}^{d}u_{k}^{2} ∥ u ∥ 2 = ∑ k = 1 d u k 2 by Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n §square . For x ∈ X x\in X x ∈ X and k ∈ [ d ] k\in[d] k ∈ [ d ] , x k x_{k} x k is the k k k -th entry of p d ( x ) p_{d}(x) p d ( x ) and ∂ k V ( x ) = ∂ k v ( p d ( x ) ) \partial_{k}V(x)=\partial_{k}v(p_{d}(x)) ∂ k V ( x ) = ∂ k v ( p d ( x )) by Admissible Cylindrical Potentials on a Hilbert Space in the Noise Geometry §gradient , so that ∣ ∇ a V ( x ) ∣ a 2 = ∣ ∇ a v ( p d ( x ) ) ∣ 2 |\nabla_{a}V(x)|_{a}^{2}=|\nabla_{a}v(p_{d}(x))|^{2} ∣ ∇ a V ( x ) ∣ a 2 = ∣ ∇ a v ( p d ( x )) ∣ 2 by Basic Properties of an Admissible Cylindrical Potential: Continuity, Growth under Noise Translations, Integrability, the Tangent Inequality and Tangency of the Noise Gradient §continuity . Put σ = ∑ k = 1 d a k \sigma=\sum_{k=1}^{d}a_{k} σ = ∑ k = 1 d a k and κ 0 = ∑ k = 1 d a k c k − 2 \kappa_{0}=\sum_{k=1}^{d}a_{k}c_{k}^{-2} κ 0 = ∑ k = 1 d a k c k − 2 , positive real numbers (a k a_{k} a k and c k c_{k} c k are positive by Probability Measures on a Hilbert Space Transported in the Noise Norm: Standing Notation §weights and A Diagonal Gaussian Reference Measure on the Noise Wasserstein Space, Rescaled Heads and Gaussian Tails: Standing Notation §gaussian ).
Step 0 (Integrability and measurability). By Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §continuity , p d p_{d} p d is Borel and ∥ p d ( x ) − p d ( 0 X ) ∥ ≤ ∣ x − 0 X ∣ \lVert p_{d}(x)-p_{d}(0_{X})\rVert\le|x-0_{X}| ∥ p d ( x ) − p d ( 0 X )∥ ≤ ∣ x − 0 X ∣ , where p d ( 0 X ) = 0 R d p_{d}(0_{X})=0_{\mathbb{R}^{d}} p d ( 0 X ) = 0 R d ; so ∥ p d ( x ) ∥ 2 ≤ ∣ x ∣ 2 \lVert p_{d}(x)\rVert^{2}\le|x|^{2} ∥ p d ( x ) ∥ 2 ≤ ∣ x ∣ 2 , and since μ ∈ P 2 ( X ) \mu\in\mathcal{P}_{2}(X) μ ∈ P 2 ( X ) the monotonicity of Linearity and Monotonicity of the Lebesgue Integral §nonnegative gives
m : = ∫ X ∥ p d ( x ) ∥ 2 μ ( d x ) ≤ M 2 ( μ ) < ∞ , m:=\int_{X}\lVert p_{d}(x)\rVert^{2}\,\mu(dx)\le M_{2}(\mu)<\infty, m := ∫ X ∥ p d ( x ) ∥ 2 μ ( d x ) ≤ M 2 ( μ ) < ∞ ,
with M 2 M_{2} M 2 the second moment of The Second Moment of a Borel Probability Measure on a Hilbert Space and the Probability Measures with Finite Second Moment §moment . For k ∈ [ d ] k\in[d] k ∈ [ d ] , ∣ x k ∣ ≤ 1 + x k 2 ≤ 1 + ∥ p d ( x ) ∥ 2 |x_{k}|\le1+x_{k}^{2}\le1+\lVert p_{d}(x)\rVert^{2} ∣ x k ∣ ≤ 1 + x k 2 ≤ 1 + ∥ p d ( x ) ∥ 2 , so x ↦ ∣ x k ∣ x\mapsto|x_{k}| x ↦ ∣ x k ∣ is integrable with respect to μ \mu μ ; and ∂ k V \partial_{k}V ∂ k V is integrable with respect to μ \mu μ by Basic Properties of an Admissible Cylindrical Potential: Continuity, Growth under Noise Translations, Integrability, the Tangent Inequality and Tangency of the Noise Gradient §integrable , V V V being integrable. Hence the functions
S 1 ( x ) = ∑ k = 1 d a k ∣ ∂ k V ( x ) ∣ , S 2 ( x ) = ∑ k = 1 d a k c k − 1 ∣ x k ∣ S_{1}(x)=\sum_{k=1}^{d}a_{k}\,|\partial_{k}V(x)|,\qquad S_{2}(x)=\sum_{k=1}^{d}a_{k}c_{k}^{-1}\,|x_{k}| S 1 ( x ) = k = 1 ∑ d a k ∣ ∂ k V ( x ) ∣ , S 2 ( x ) = k = 1 ∑ d a k c k − 1 ∣ x k ∣
are integrable with respect to μ \mu μ by Linearity and Monotonicity of the Lebesgue Integral §integrable . Every function on X X X formed below is obtained from the coordinates x k x_{k} x k , the continuous functions ∂ k V \partial_{k}V ∂ k V of Basic Properties of an Admissible Cylindrical Potential: Continuity, Growth under Noise Translations, Integrability, the Tangent Inequality and Tangency of the Noise Gradient §continuity , and functions f ∘ p d f\circ p_{d} f ∘ p d with f f f continuous on R d \mathbb{R}^{d} R d , by finitely many sums, products and absolute values; it is therefore Borel by claims 3 and 4 of Borel Measurability and Bounded Integration on a Metric Space and claims 2 to 4 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions , a function of class C k C^{k} C k on R d \mathbb{R}^{d} R d being continuous by claim 3 of Euclidean Space is Open in Itself, and C k C^k C k Maps are Continuous . Whenever such a function is bounded pointwise in absolute value by an integrable function (or by a constant, constants being integrable by claim 6(a) of Borel Measurability and Bounded Integration on a Metric Space ), it is integrable, by the monotonicity of Linearity and Monotonicity of the Lebesgue Integral §nonnegative and the definition of integrability ; this remark is used without repetition.
Step 1 (Bounds on the profile). Put U ( u ) = v ( u ) + b + 1 U(u)=v(u)+b+1 U ( u ) = v ( u ) + b + 1 for u ∈ R d u\in\mathbb{R}^{d} u ∈ R d ; then U ( u ) ≥ 1 U(u)\ge1 U ( u ) ≥ 1 by Admissible Cylindrical Potentials on a Hilbert Space in the Noise Geometry §below . A constant function on R d \mathbb{R}^{d} R d is smooth by claim 2 of Constants, Coordinate Functions, Sums and Products of C k C^k C k Functions on a Euclidean Open Set , hence of class C 2 C^{2} C 2 by Smooth Map on a Euclidean Open Set , and its partial derivatives vanish, its difference quotients in Partial Derivative on a Euclidean Open Set being 0 0 0 ; so by claims 1 and 3 of Constants, Coordinate Functions, Sums and Products of C k C^k C k Functions on a Euclidean Open Set , U U U is of class C 2 C^{2} C 2 on R d \mathbb{R}^{d} R d , with ∂ k U = ∂ k v \partial_{k}U=\partial_{k}v ∂ k U = ∂ k v and ∂ j ∂ k U = ∂ j ∂ k v \partial_{j}\partial_{k}U=\partial_{j}\partial_{k}v ∂ j ∂ k U = ∂ j ∂ k v for j , k ∈ [ d ] j,k\in[d] j , k ∈ [ d ] .
Slope. Since v ( u ) = U ( u ) − ( b + 1 ) v(u)=U(u)-(b+1) v ( u ) = U ( u ) − ( b + 1 ) and U ( u ) ≥ 1 U(u)\ge1 U ( u ) ≥ 1 , ∣ v ( u ) ∣ ≤ U ( u ) + ∣ b + 1 ∣ ≤ ( 1 + ∣ b + 1 ∣ ) U ( u ) |v(u)|\le U(u)+|b+1|\le(1+|b+1|)U(u) ∣ v ( u ) ∣ ≤ U ( u ) + ∣ b + 1∣ ≤ ( 1 + ∣ b + 1∣ ) U ( u ) , so 0 < 1 + ∣ v ( u ) ∣ ≤ ( ∣ b + 1 ∣ + 2 ) U ( u ) 0<1+|v(u)|\le(|b+1|+2)U(u) 0 < 1 + ∣ v ( u ) ∣ ≤ ( ∣ b + 1∣ + 2 ) U ( u ) . By Admissible Cylindrical Potentials on a Hilbert Space in the Noise Geometry §slope and C ≥ 0 C\ge0 C ≥ 0 , ∣ ∇ a v ( u ) ∣ 2 ≤ C ( 1 + ∣ v ( u ) ∣ ) 2 ≤ C ( ∣ b + 1 ∣ + 2 ) 2 U ( u ) 2 = ( L U ( u ) ) 2 |\nabla_{a}v(u)|^{2}\le C(1+|v(u)|)^{2}\le C(|b+1|+2)^{2}U(u)^{2}=(L\,U(u))^{2} ∣ ∇ a v ( u ) ∣ 2 ≤ C ( 1 + ∣ v ( u ) ∣ ) 2 ≤ C ( ∣ b + 1∣ + 2 ) 2 U ( u ) 2 = ( L U ( u ) ) 2 , whence
∣ ∇ a v ( u ) ∣ ≤ L U ( u ) , a k 1 / 2 ∣ ∂ k v ( u ) ∣ ≤ ∣ ∇ a v ( u ) ∣ ( k ∈ [ d ] ) , (1.1) |\nabla_{a}v(u)|\le L\,U(u),\qquad a_{k}^{1/2}|\partial_{k}v(u)|\le|\nabla_{a}v(u)|\quad(k\in[d]),\tag{1.1} ∣ ∇ a v ( u ) ∣ ≤ L U ( u ) , a k 1/2 ∣ ∂ k v ( u ) ∣ ≤ ∣ ∇ a v ( u ) ∣ ( k ∈ [ d ]) , ( 1.1 )
the second because a k ( ∂ k v ( u ) ) 2 a_{k}(\partial_{k}v(u))^{2} a k ( ∂ k v ( u ) ) 2 is one of the nonnegative terms of ∣ ∇ a v ( u ) ∣ 2 |\nabla_{a}v(u)|^{2} ∣ ∇ a v ( u ) ∣ 2 .
Hessian. Let u ∈ R d u\in\mathbb{R}^{d} u ∈ R d and put T ( u ) = ∑ i = 1 d ( a i ∂ i ∂ i v ( u ) + K ) T(u)=\sum_{i=1}^{d}\bigl(a_{i}\,\partial_{i}\partial_{i}v(u)+K\bigr) T ( u ) = ∑ i = 1 d ( a i ∂ i ∂ i v ( u ) + K ) . Applying Admissible Cylindrical Potentials on a Hilbert Space in the Noise Geometry §semiconvex with the vector h h h whose only nonzero entry is h j = a j 1 / 2 h_{j}=a_{j}^{1/2} h j = a j 1/2 gives a j ∂ j ∂ j v ( u ) ≥ − K a_{j}\partial_{j}\partial_{j}v(u)\ge-K a j ∂ j ∂ j v ( u ) ≥ − K , so each term of T ( u ) T(u) T ( u ) is nonnegative. For j ≠ k j\ne k j = k in [ d ] [d] [ d ] , apply it with h j = a j 1 / 2 h_{j}=a_{j}^{1/2} h j = a j 1/2 , h k = ± a k 1 / 2 h_{k}=\pm a_{k}^{1/2} h k = ± a k 1/2 and all other entries 0 0 0 ; since ∂ j ∂ k v ( u ) = ∂ k ∂ j v ( u ) \partial_{j}\partial_{k}v(u)=\partial_{k}\partial_{j}v(u) ∂ j ∂ k v ( u ) = ∂ k ∂ j v ( u ) by claim 1 of Equality of Mixed Second Partial Derivatives and Symmetry of the Hessian , this gives a j ∂ j ∂ j v ( u ) + a k ∂ k ∂ k v ( u ) ± 2 ( a j a k ) 1 / 2 ∂ j ∂ k v ( u ) ≥ − 2 K a_{j}\partial_{j}\partial_{j}v(u)+a_{k}\partial_{k}\partial_{k}v(u)\pm2(a_{j}a_{k})^{1/2}\partial_{j}\partial_{k}v(u)\ge-2K a j ∂ j ∂ j v ( u ) + a k ∂ k ∂ k v ( u ) ± 2 ( a j a k ) 1/2 ∂ j ∂ k v ( u ) ≥ − 2 K , whence 2 ( a j a k ) 1 / 2 ∣ ∂ j ∂ k v ( u ) ∣ ≤ ( a j ∂ j ∂ j v ( u ) + K ) + ( a k ∂ k ∂ k v ( u ) + K ) ≤ T ( u ) 2(a_{j}a_{k})^{1/2}|\partial_{j}\partial_{k}v(u)|\le(a_{j}\partial_{j}\partial_{j}v(u)+K)+(a_{k}\partial_{k}\partial_{k}v(u)+K)\le T(u) 2 ( a j a k ) 1/2 ∣ ∂ j ∂ k v ( u ) ∣ ≤ ( a j ∂ j ∂ j v ( u ) + K ) + ( a k ∂ k ∂ k v ( u ) + K ) ≤ T ( u ) . For j = k j=k j = k , the number y = a j ∂ j ∂ j v ( u ) y=a_{j}\partial_{j}\partial_{j}v(u) y = a j ∂ j ∂ j v ( u ) satisfies − K ≤ y -K\le y − K ≤ y and y + K ≤ T ( u ) y+K\le T(u) y + K ≤ T ( u ) , so ∣ y ∣ ≤ T ( u ) + K |y|\le T(u)+K ∣ y ∣ ≤ T ( u ) + K . With Admissible Cylindrical Potentials on a Hilbert Space in the Noise Geometry §curvature for ε = 1 \varepsilon=1 ε = 1 , therefore, for all j , k ∈ [ d ] j,k\in[d] j , k ∈ [ d ] ,
( a j a k ) 1 / 2 ∣ ∂ j ∂ k v ( u ) ∣ ≤ T ( u ) + K , T ( u ) ≤ ∣ ∇ a v ( u ) ∣ 2 + C ′ ′ ( 1 + ∥ u ∥ 2 ) + d K . (1.2) (a_{j}a_{k})^{1/2}|\partial_{j}\partial_{k}v(u)|\le T(u)+K,\qquad T(u)\le|\nabla_{a}v(u)|^{2}+C''\bigl(1+\lVert u\rVert^{2}\bigr)+dK.\tag{1.2} ( a j a k ) 1/2 ∣ ∂ j ∂ k v ( u ) ∣ ≤ T ( u ) + K , T ( u ) ≤ ∣ ∇ a v ( u ) ∣ 2 + C ′′ ( 1 + ∥ u ∥ 2 ) + d K . ( 1.2 )
Step 2 (The truncations). Let p ∈ N p\in\mathbb{N} p ∈ N , read as a positive real number, and let J = ( 0 , ∞ ) J=(0,\infty) J = ( 0 , ∞ ) , an open interval , which is an open subset of R 1 \mathbb{R}^{1} R 1 (for s ∈ J s\in J s ∈ J the ball of radius s s s about s s s lies in J J J ). Define G p : J → R G_{p}:J\to\mathbb{R} G p : J → R by G p ( s ) = p s ( p + s ) − 1 = p − p 2 ( p + s ) − 1 G_{p}(s)=ps(p+s)^{-1}=p-p^{2}(p+s)^{-1} G p ( s ) = p s ( p + s ) − 1 = p − p 2 ( p + s ) − 1 . The function s ↦ p + s s\mapsto p+s s ↦ p + s is differentiable at every point of J J J with derivative 1 1 1 , its difference quotients in Derivative at an Interior Point being 1 1 1 , and it does not vanish on J J J . Hence, by claim 1 of Reciprocal Rule for One-Dimensional Derivatives and claims 2 and 3 of Sum, Constant Multiple, and Product Rules for One-Dimensional Derivatives (writing ( p + s ) − 2 (p+s)^{-2} ( p + s ) − 2 and ( p + s ) − 3 (p+s)^{-3} ( p + s ) − 3 as products of copies of ( p + s ) − 1 (p+s)^{-1} ( p + s ) − 1 ), G p G_{p} G p , G p ′ G_{p}' G p ′ and G p ′ ′ G_{p}'' G p ′′ are differentiable at every point of J J J , with
G p ′ ( s ) = p 2 ( p + s ) − 2 , G p ′ ′ ( s ) = − 2 p 2 ( p + s ) − 3 , G_{p}'(s)=p^{2}(p+s)^{-2},\qquad G_{p}''(s)=-2p^{2}(p+s)^{-3}, G p ′ ( s ) = p 2 ( p + s ) − 2 , G p ′′ ( s ) = − 2 p 2 ( p + s ) − 3 ,
and all three are continuous at every point of J J J relative to J J J by Differentiability at an Interior Point Implies Continuity There ; on R 1 \mathbb{R}^{1} R 1 the Euclidean distance is ∣ s − s ′ ∣ |s-s'| ∣ s − s ′ ∣ , so this is continuity in the sense of clause 1 of C^k Maps on a Euclidean Open Set . As the partial derivative of Partial Derivative on a Euclidean Open Set with n = 1 n=1 n = 1 is the derivative of Derivative at an Interior Point , G p G_{p} G p is of class C 2 C^{2} C 2 on J J J in the sense of C^k Maps on a Euclidean Open Set , with ∂ 1 G p = G p ′ \partial_{1}G_{p}=G_{p}' ∂ 1 G p = G p ′ and ∂ 1 ∂ 1 G p = G p ′ ′ \partial_{1}\partial_{1}G_{p}=G_{p}'' ∂ 1 ∂ 1 G p = G p ′′ . For s ≥ 1 s\ge1 s ≥ 1 put α = p ( p + s ) − 1 \alpha=p(p+s)^{-1} α = p ( p + s ) − 1 , so that α \alpha α and 1 − α = s ( p + s ) − 1 1-\alpha=s(p+s)^{-1} 1 − α = s ( p + s ) − 1 lie in ( 0 , 1 ) (0,1) ( 0 , 1 ) ; then G p ( s ) = p ( 1 − α ) G_{p}(s)=p(1-\alpha) G p ( s ) = p ( 1 − α ) , G p ′ ( s ) = α 2 G_{p}'(s)=\alpha^{2} G p ′ ( s ) = α 2 , G p ′ ( s ) s = p α ( 1 − α ) G_{p}'(s)\,s=p\,\alpha(1-\alpha) G p ′ ( s ) s = p α ( 1 − α ) , G p ′ ( s ) s 2 = p 2 ( 1 − α ) 2 G_{p}'(s)\,s^{2}=p^{2}(1-\alpha)^{2} G p ′ ( s ) s 2 = p 2 ( 1 − α ) 2 and ∣ G p ′ ′ ( s ) ∣ s 2 = 2 p α ( 1 − α ) 2 |G_{p}''(s)|\,s^{2}=2p\,\alpha(1-\alpha)^{2} ∣ G p ′′ ( s ) ∣ s 2 = 2 p α ( 1 − α ) 2 , so
0 < G p ( s ) ≤ p , 0 < G p ′ ( s ) ≤ 1 , G p ′ ( s ) s ≤ p , G p ′ ( s ) s 2 ≤ p 2 , G p ′ ′ ( s ) < 0 , ∣ G p ′ ′ ( s ) ∣ s 2 ≤ 2 p . (2.1) 0<G_{p}(s)\le p,\quad0<G_{p}'(s)\le1,\quad G_{p}'(s)\,s\le p,\quad G_{p}'(s)\,s^{2}\le p^{2},\quad G_{p}''(s)<0,\quad|G_{p}''(s)|\,s^{2}\le2p.\tag{2.1} 0 < G p ( s ) ≤ p , 0 < G p ′ ( s ) ≤ 1 , G p ′ ( s ) s ≤ p , G p ′ ( s ) s 2 ≤ p 2 , G p ′′ ( s ) < 0 , ∣ G p ′′ ( s ) ∣ s 2 ≤ 2 p . ( 2.1 )
As U U U takes values in [ 1 , ∞ ) ⊆ J [1,\infty)\subseteq J [ 1 , ∞ ) ⊆ J , we may put F p = G p ∘ U F_{p}=G_{p}\circ U F p = G p ∘ U , w p = G p ′ ∘ U w_{p}=G_{p}'\circ U w p = G p ′ ∘ U and z p = G p ′ ′ ∘ U z_{p}=G_{p}''\circ U z p = G p ′′ ∘ U on R d \mathbb{R}^{d} R d ; note w p ( u ) = ( 1 + p − 1 U ( u ) ) − 2 w_{p}(u)=(1+p^{-1}U(u))^{-2} w p ( u ) = ( 1 + p − 1 U ( u ) ) − 2 .
Monotonicity in p p p . Fix u u u . If p ≤ p ′ p\le p' p ≤ p ′ in N \mathbb{N} N , then 0 < p ′ − 1 U ( u ) ≤ p − 1 U ( u ) 0<p'^{-1}U(u)\le p^{-1}U(u) 0 < p ′ − 1 U ( u ) ≤ p − 1 U ( u ) , so 1 ≤ 1 + p ′ − 1 U ( u ) ≤ 1 + p − 1 U ( u ) 1\le1+p'^{-1}U(u)\le1+p^{-1}U(u) 1 ≤ 1 + p ′ − 1 U ( u ) ≤ 1 + p − 1 U ( u ) ; squaring, and using that reciprocals reverse the order of positive numbers (Order Reversal under Reciprocals, and Summability of the Reciprocals of the Squares §reciprocal ), w p ( u ) ≤ w p ′ ( u ) w_{p}(u)\le w_{p'}(u) w p ( u ) ≤ w p ′ ( u ) . Moreover p − 1 → 0 p^{-1}\to0 p − 1 → 0 as p → ∞ p\to\infty p → ∞ , by The Archimedean Property of the Real Numbers and Limit of a Sequence of Real Numbers , so w p ( u ) → 1 w_{p}(u)\to1 w p ( u ) → 1 by Arithmetic of Limits of Real Sequences .
Step 3 (The test functions). Apply Scaled Cutoffs and the Second-Moment Test Functions: Uniform Derivative Bounds and Agreement on a Ball with q = d q=d q = d : fix a function χ : R d → R \chi:\mathbb{R}^{d}\to\mathbb{R} χ : R d → R as in that lemma, and for positive r r r let χ r ( u ) = χ ( r − 1 u ) \chi_{r}(u)=\chi(r^{-1}u) χ r ( u ) = χ ( r − 1 u ) (u ∈ R d u\in\mathbb{R}^{d} u ∈ R d ) be its cutoffs, as defined there; let M 1 , M 2 ≥ 0 M_{1},M_{2}\ge0 M 1 , M 2 ≥ 0 be the constants of Scaled Cutoffs and the Second-Moment Test Functions: Uniform Derivative Bounds and Agreement on a Ball §cutoff . We use r ∈ N r\in\mathbb{N} r ∈ N , read as a positive real number, so r − 2 ≤ r − 1 r^{-2}\le r^{-1} r − 2 ≤ r − 1 . By that clause, χ r \chi_{r} χ r is smooth, hence of class C 2 C^{2} C 2 by Smooth Map on a Euclidean Open Set ; 0 ≤ χ r ≤ 1 0\le\chi_{r}\le1 0 ≤ χ r ≤ 1 ; χ r ( u ) = 1 \chi_{r}(u)=1 χ r ( u ) = 1 if ∥ u ∥ ≤ r \lVert u\rVert\le r ∥ u ∥ ≤ r and χ r ( u ) = 0 \chi_{r}(u)=0 χ r ( u ) = 0 if ∥ u ∥ ≥ 2 r \lVert u\rVert\ge2r ∥ u ∥ ≥ 2 r ; and ∣ ∂ i χ r ∣ ≤ M 1 r − 1 |\partial_{i}\chi_{r}|\le M_{1}r^{-1} ∣ ∂ i χ r ∣ ≤ M 1 r − 1 , ∣ ∂ j ∂ i χ r ∣ ≤ M 2 r − 2 |\partial_{j}\partial_{i}\chi_{r}|\le M_{2}r^{-2} ∣ ∂ j ∂ i χ r ∣ ≤ M 2 r − 2 on R d \mathbb{R}^{d} R d for i , j ∈ [ d ] i,j\in[d] i , j ∈ [ d ] .
For p , r ∈ N p,r\in\mathbb{N} p , r ∈ N let g = g p , r = F p χ r : R d → R g=g_{p,r}=F_{p}\,\chi_{r}:\mathbb{R}^{d}\to\mathbb{R} g = g p , r = F p χ r : R d → R . By claim 2 of A Composition of C k C^k C k Maps Between Euclidean Open Sets is of Class C k C^k C k , applied to U : R d → J U:\mathbb{R}^{d}\to J U : R d → J and G p G_{p} G p , both of class C 2 C^{2} C 2 , F p F_{p} F p is of class C 2 C^{2} C 2 on R d \mathbb{R}^{d} R d , and so is g g g by claim 3 of Constants, Coordinate Functions, Sums and Products of C k C^k C k Functions on a Euclidean Open Set . By claim 1 of A Composition of C k C^k C k Maps Between Euclidean Open Sets is of Class C k C^k C k , applied to G p G_{p} G p and to G p ′ G_{p}' G p ′ (both of class C 1 C^{1} C 1 on J J J ), ∂ k F p = w p ∂ k v \partial_{k}F_{p}=w_{p}\,\partial_{k}v ∂ k F p = w p ∂ k v and ∂ j w p = z p ∂ j v \partial_{j}w_{p}=z_{p}\,\partial_{j}v ∂ j w p = z p ∂ j v ; so the product rule (claim 1 of Constants, Coordinate Functions, Sums and Products of C k C^k C k Functions on a Euclidean Open Set ) gives, for j , k ∈ [ d ] j,k\in[d] j , k ∈ [ d ] , on R d \mathbb{R}^{d} R d ,
∂ k g = w p ∂ k v χ r + F p ∂ k χ r , (3.1) \partial_{k}g=w_{p}\,\partial_{k}v\,\chi_{r}+F_{p}\,\partial_{k}\chi_{r},\tag{3.1} ∂ k g = w p ∂ k v χ r + F p ∂ k χ r , ( 3.1 )
∂ j ∂ k g = ( z p ∂ j v ∂ k v + w p ∂ j ∂ k v ) χ r + w p ∂ k v ∂ j χ r + w p ∂ j v ∂ k χ r + F p ∂ j ∂ k χ r . (3.2) \partial_{j}\partial_{k}g=\bigl(z_{p}\,\partial_{j}v\,\partial_{k}v+w_{p}\,\partial_{j}\partial_{k}v\bigr)\chi_{r}+w_{p}\,\partial_{k}v\,\partial_{j}\chi_{r}+w_{p}\,\partial_{j}v\,\partial_{k}\chi_{r}+F_{p}\,\partial_{j}\partial_{k}\chi_{r}.\tag{3.2} ∂ j ∂ k g = ( z p ∂ j v ∂ k v + w p ∂ j ∂ k v ) χ r + w p ∂ k v ∂ j χ r + w p ∂ j v ∂ k χ r + F p ∂ j ∂ k χ r . ( 3.2 )
Bounds. Let u ∈ R d u\in\mathbb{R}^{d} u ∈ R d and write U = U ( u ) U=U(u) U = U ( u ) . By (1.1) and (2.1), w p ∣ ∂ k v ∣ ≤ a k − 1 / 2 w p L U ≤ a k − 1 / 2 L p w_{p}|\partial_{k}v|\le a_{k}^{-1/2}w_{p}LU\le a_{k}^{-1/2}Lp w p ∣ ∂ k v ∣ ≤ a k − 1/2 w p LU ≤ a k − 1/2 L p , ∣ z p ∣ ∣ ∂ j v ∣ ∣ ∂ k v ∣ ≤ ( a j a k ) − 1 / 2 ∣ z p ∣ L 2 U 2 ≤ ( a j a k ) − 1 / 2 2 p L 2 |z_{p}|\,|\partial_{j}v|\,|\partial_{k}v|\le(a_{j}a_{k})^{-1/2}|z_{p}|L^{2}U^{2}\le(a_{j}a_{k})^{-1/2}2pL^{2} ∣ z p ∣ ∣ ∂ j v ∣ ∣ ∂ k v ∣ ≤ ( a j a k ) − 1/2 ∣ z p ∣ L 2 U 2 ≤ ( a j a k ) − 1/2 2 p L 2 , and
w p ( u ) ∣ ∇ a v ( u ) ∣ 2 ≤ w p ( u ) L 2 U 2 ≤ L 2 p 2 . (3.3) w_{p}(u)\,|\nabla_{a}v(u)|^{2}\le w_{p}(u)\,L^{2}U^{2}\le L^{2}p^{2}.\tag{3.3} w p ( u ) ∣ ∇ a v ( u ) ∣ 2 ≤ w p ( u ) L 2 U 2 ≤ L 2 p 2 . ( 3.3 )
If χ r ( u ) ≠ 0 \chi_{r}(u)\ne0 χ r ( u ) = 0 , then ∥ u ∥ < 2 r \lVert u\rVert<2r ∥ u ∥ < 2 r , and by (1.2), 0 < w p ≤ 1 0<w_{p}\le1 0 < w p ≤ 1 and (3.3), w p ( a j a k ) 1 / 2 ∣ ∂ j ∂ k v ∣ χ r ≤ w p ( T ( u ) + K ) ≤ w p ∣ ∇ a v ( u ) ∣ 2 + ∣ C ′ ′ ∣ ( 1 + 4 r 2 ) + ( d + 1 ) K ≤ L 2 p 2 + ∣ C ′ ′ ∣ ( 1 + 4 r 2 ) + ( d + 1 ) K w_{p}(a_{j}a_{k})^{1/2}|\partial_{j}\partial_{k}v|\chi_{r}\le w_{p}(T(u)+K)\le w_{p}|\nabla_{a}v(u)|^{2}+|C''|(1+4r^{2})+(d+1)K\le L^{2}p^{2}+|C''|(1+4r^{2})+(d+1)K w p ( a j a k ) 1/2 ∣ ∂ j ∂ k v ∣ χ r ≤ w p ( T ( u ) + K ) ≤ w p ∣ ∇ a v ( u ) ∣ 2 + ∣ C ′′ ∣ ( 1 + 4 r 2 ) + ( d + 1 ) K ≤ L 2 p 2 + ∣ C ′′ ∣ ( 1 + 4 r 2 ) + ( d + 1 ) K ; if χ r ( u ) = 0 \chi_{r}(u)=0 χ r ( u ) = 0 this term vanishes. With 0 < F p ≤ p 0<F_{p}\le p 0 < F p ≤ p and the bounds on χ r \chi_{r} χ r and its derivatives, (3.1) and (3.2) show that g g g , all ∂ k g \partial_{k}g ∂ k g and all ∂ j ∂ k g \partial_{j}\partial_{k}g ∂ j ∂ k g are bounded on R d \mathbb{R}^{d} R d (with bounds depending on p p p and r r r ). Hence g ∈ C b 2 ( R d ) g\in C^{2}_{b}(\mathbb{R}^{d}) g ∈ C b 2 ( R d ) by Bounded Twice Continuously Differentiable Functions with Bounded First and Second Partial Derivatives on Euclidean Space §bounded , and the hypothesis of the lemma with n = d n=d n = d gives, β \beta β being positive,
β L μ a , V ( g p , r ) ≤ β ∣ L μ a , V ( g p , r ) ∣ ≤ β R ∥ ∇ a ( g p , r ∘ p d ) ∥ μ . (3.4) \beta\,L^{a,V}_{\mu}(g_{p,r})\le\beta\,\bigl|L^{a,V}_{\mu}(g_{p,r})\bigr|\le\beta R\,\lVert\nabla_{a}(g_{p,r}\circ p_{d})\rVert_{\mu}.\tag{3.4} β L μ a , V ( g p , r ) ≤ β L μ a , V ( g p , r ) ≤ βR ∥ ∇ a ( g p , r ∘ p d ) ∥ μ . ( 3.4 )
Step 4 (A lower bound for the functional). Fix p , r ∈ N p,r\in\mathbb{N} p , r ∈ N and write g = g p , r g=g_{p,r} g = g p , r . For x ∈ X x\in X x ∈ X put u = p d ( x ) u=p_{d}(x) u = p d ( x ) . By The Gibbs Ornstein-Uhlenbeck Functional of a Probability Measure at a Bounded C^2 Function of Finitely Many Coordinates §functional with n = d n=d n = d and the linearity of Linearity and Monotonicity of the Lebesgue Integral §integrable , L μ a , V ( g ) = ∫ X Φ d μ L^{a,V}_{\mu}(g)=\int_{X}\Phi\,d\mu L μ a , V ( g ) = ∫ X Φ d μ , where the integrable function Φ \Phi Φ is
Φ ( x ) = ∑ k = 1 d a k ( ( u k c k + ∂ k v ( u ) β ) ∂ k g ( u ) − ∂ k ∂ k g ( u ) ) . \Phi(x)=\sum_{k=1}^{d}a_{k}\Bigl(\Bigl(\frac{u_{k}}{c_{k}}+\frac{\partial_{k}v(u)}{\beta}\Bigr)\partial_{k}g(u)-\partial_{k}\partial_{k}g(u)\Bigr). Φ ( x ) = k = 1 ∑ d a k ( ( c k u k + β ∂ k v ( u ) ) ∂ k g ( u ) − ∂ k ∂ k g ( u ) ) .
Inserting (3.1) and (3.2) with j = k j=k j = k , and suppressing the argument u u u ,
β Φ = w p χ r ∣ ∇ a v ∣ 2 + β w p χ r ∑ k = 1 d a k c k u k ∂ k v − β w p χ r ∑ k = 1 d a k ∂ k ∂ k v − β z p χ r ∣ ∇ a v ∣ 2 + E , \beta\Phi=w_{p}\chi_{r}|\nabla_{a}v|^{2}+\beta w_{p}\chi_{r}\sum_{k=1}^{d}\frac{a_{k}}{c_{k}}u_{k}\,\partial_{k}v-\beta w_{p}\chi_{r}\sum_{k=1}^{d}a_{k}\,\partial_{k}\partial_{k}v-\beta z_{p}\chi_{r}|\nabla_{a}v|^{2}+E, β Φ = w p χ r ∣ ∇ a v ∣ 2 + β w p χ r k = 1 ∑ d c k a k u k ∂ k v − β w p χ r k = 1 ∑ d a k ∂ k ∂ k v − β z p χ r ∣ ∇ a v ∣ 2 + E ,
E = F p ∑ k = 1 d a k ∂ k v ∂ k χ r + β F p ∑ k = 1 d a k c k u k ∂ k χ r − 2 β w p ∑ k = 1 d a k ∂ k v ∂ k χ r − β F p ∑ k = 1 d a k ∂ k ∂ k χ r . E=F_{p}\sum_{k=1}^{d}a_{k}\,\partial_{k}v\,\partial_{k}\chi_{r}+\beta F_{p}\sum_{k=1}^{d}\frac{a_{k}}{c_{k}}u_{k}\,\partial_{k}\chi_{r}-2\beta w_{p}\sum_{k=1}^{d}a_{k}\,\partial_{k}v\,\partial_{k}\chi_{r}-\beta F_{p}\sum_{k=1}^{d}a_{k}\,\partial_{k}\partial_{k}\chi_{r}. E = F p k = 1 ∑ d a k ∂ k v ∂ k χ r + β F p k = 1 ∑ d c k a k u k ∂ k χ r − 2 β w p k = 1 ∑ d a k ∂ k v ∂ k χ r − β F p k = 1 ∑ d a k ∂ k ∂ k χ r .
We bound the terms from below. (i) − β z p χ r ∣ ∇ a v ∣ 2 ≥ 0 -\beta z_{p}\chi_{r}|\nabla_{a}v|^{2}\ge0 − β z p χ r ∣ ∇ a v ∣ 2 ≥ 0 , since z p < 0 z_{p}<0 z p < 0 by (2.1), χ r ≥ 0 \chi_{r}\ge0 χ r ≥ 0 and β > 0 \beta>0 β > 0 . (ii) For each k k k , with s = a k 1 / 2 ∣ ∂ k v ∣ s=a_{k}^{1/2}|\partial_{k}v| s = a k 1/2 ∣ ∂ k v ∣ and t = β a k 1 / 2 c k − 1 ∣ u k ∣ t=\beta a_{k}^{1/2}c_{k}^{-1}|u_{k}| t = β a k 1/2 c k − 1 ∣ u k ∣ , β a k c k − 1 ∣ u k ∣ ∣ ∂ k v ∣ = s t ≤ 1 4 a k ( ∂ k v ) 2 + β 2 a k c k − 2 u k 2 \beta a_{k}c_{k}^{-1}|u_{k}|\,|\partial_{k}v|=st\le\tfrac14a_{k}(\partial_{k}v)^{2}+\beta^{2}a_{k}c_{k}^{-2}u_{k}^{2} β a k c k − 1 ∣ u k ∣ ∣ ∂ k v ∣ = s t ≤ 4 1 a k ( ∂ k v ) 2 + β 2 a k c k − 2 u k 2 , and u k 2 ≤ ∥ u ∥ 2 u_{k}^{2}\le\lVert u\rVert^{2} u k 2 ≤ ∥ u ∥ 2 ; summing over k k k and multiplying by w p χ r ∈ [ 0 , 1 ] w_{p}\chi_{r}\in[0,1] w p χ r ∈ [ 0 , 1 ] , the second term of β Φ \beta\Phi β Φ is at least − 1 4 w p χ r ∣ ∇ a v ∣ 2 − β 2 κ 0 ∥ u ∥ 2 -\tfrac14w_{p}\chi_{r}|\nabla_{a}v|^{2}-\beta^{2}\kappa_{0}\lVert u\rVert^{2} − 4 1 w p χ r ∣ ∇ a v ∣ 2 − β 2 κ 0 ∥ u ∥ 2 . (iii) By the choice of C ′ C' C ′ , ∑ k = 1 d a k ∂ k ∂ k v ( u ) ≤ ( 2 β ) − 1 ∣ ∇ a v ( u ) ∣ 2 + C ′ ( 1 + ∥ u ∥ 2 ) \sum_{k=1}^{d}a_{k}\partial_{k}\partial_{k}v(u)\le(2\beta)^{-1}|\nabla_{a}v(u)|^{2}+C'(1+\lVert u\rVert^{2}) ∑ k = 1 d a k ∂ k ∂ k v ( u ) ≤ ( 2 β ) − 1 ∣ ∇ a v ( u ) ∣ 2 + C ′ ( 1 + ∥ u ∥ 2 ) ; multiplying by − β w p χ r ≤ 0 -\beta w_{p}\chi_{r}\le0 − β w p χ r ≤ 0 , and using w p χ r ∈ [ 0 , 1 ] w_{p}\chi_{r}\in[0,1] w p χ r ∈ [ 0 , 1 ] , the third term of β Φ \beta\Phi β Φ is at least − 1 2 w p χ r ∣ ∇ a v ∣ 2 − β ∣ C ′ ∣ ( 1 + ∥ u ∥ 2 ) -\tfrac12w_{p}\chi_{r}|\nabla_{a}v|^{2}-\beta|C'|(1+\lVert u\rVert^{2}) − 2 1 w p χ r ∣ ∇ a v ∣ 2 − β ∣ C ′ ∣ ( 1 + ∥ u ∥ 2 ) . (iv) Since 0 < F p ≤ p 0<F_{p}\le p 0 < F p ≤ p , 0 < w p ≤ 1 0<w_{p}\le1 0 < w p ≤ 1 , ∣ u k ∣ = ∣ x k ∣ |u_{k}|=|x_{k}| ∣ u k ∣ = ∣ x k ∣ , ∂ k v ( u ) = ∂ k V ( x ) \partial_{k}v(u)=\partial_{k}V(x) ∂ k v ( u ) = ∂ k V ( x ) and r − 2 ≤ r − 1 r^{-2}\le r^{-1} r − 2 ≤ r − 1 , the bounds on the derivatives of χ r \chi_{r} χ r give ∣ E ∣ ≤ e r ( x ) |E|\le e_{r}(x) ∣ E ∣ ≤ e r ( x ) , where
e r ( x ) = r − 1 ( M 1 ( p + 2 β ) S 1 ( x ) + β p M 1 S 2 ( x ) + β p M 2 σ ) . e_{r}(x)=r^{-1}\Bigl(M_{1}(p+2\beta)\,S_{1}(x)+\beta pM_{1}\,S_{2}(x)+\beta pM_{2}\,\sigma\Bigr). e r ( x ) = r − 1 ( M 1 ( p + 2 β ) S 1 ( x ) + βp M 1 S 2 ( x ) + βp M 2 σ ) .
Adding, β Φ ( x ) ≥ 1 4 w p ( u ) χ r ( u ) ∣ ∇ a v ( u ) ∣ 2 − Q ( x ) − e r ( x ) \beta\Phi(x)\ge\tfrac14w_{p}(u)\chi_{r}(u)|\nabla_{a}v(u)|^{2}-Q(x)-e_{r}(x) β Φ ( x ) ≥ 4 1 w p ( u ) χ r ( u ) ∣ ∇ a v ( u ) ∣ 2 − Q ( x ) − e r ( x ) , with Q ( x ) = β 2 κ 0 ∥ p d ( x ) ∥ 2 + β ∣ C ′ ∣ ( 1 + ∥ p d ( x ) ∥ 2 ) Q(x)=\beta^{2}\kappa_{0}\lVert p_{d}(x)\rVert^{2}+\beta|C'|(1+\lVert p_{d}(x)\rVert^{2}) Q ( x ) = β 2 κ 0 ∥ p d ( x ) ∥ 2 + β ∣ C ′ ∣ ( 1 + ∥ p d ( x ) ∥ 2 ) . By Step 0, Q Q Q and e r e_{r} e r are integrable, with
∫ X Q d μ = D : = ( β 2 κ 0 + β ∣ C ′ ∣ ) m + β ∣ C ′ ∣ , ∫ X e r d μ = A p r , A p : = M 1 ( p + 2 β ) ∫ X S 1 d μ + β p M 1 ∫ X S 2 d μ + β p M 2 σ , \int_{X}Q\,d\mu=D:=(\beta^{2}\kappa_{0}+\beta|C'|)\,m+\beta|C'|,\qquad\int_{X}e_{r}\,d\mu=\frac{A_{p}}{r},\quad A_{p}:=M_{1}(p+2\beta)\!\int_{X}\!S_{1}\,d\mu+\beta pM_{1}\!\int_{X}\!S_{2}\,d\mu+\beta pM_{2}\sigma, ∫ X Q d μ = D := ( β 2 κ 0 + β ∣ C ′ ∣ ) m + β ∣ C ′ ∣ , ∫ X e r d μ = r A p , A p := M 1 ( p + 2 β ) ∫ X S 1 d μ + βp M 1 ∫ X S 2 d μ + βp M 2 σ ,
by Linearity and Monotonicity of the Lebesgue Integral §integrable and claim 6(a) of Borel Measurability and Bounded Integration on a Metric Space ; D D D does not depend on p p p or r r r , and A p A_{p} A p does not depend on r r r . The function x ↦ w p ( p d ( x ) ) χ r ( p d ( x ) ) ∣ ∇ a v ( p d ( x ) ) ∣ 2 x\mapsto w_{p}(p_{d}(x))\chi_{r}(p_{d}(x))|\nabla_{a}v(p_{d}(x))|^{2} x ↦ w p ( p d ( x )) χ r ( p d ( x )) ∣ ∇ a v ( p d ( x )) ∣ 2 lies between 0 0 0 and L 2 p 2 L^{2}p^{2} L 2 p 2 by (3.3), so it is integrable; let Y p , r Y_{p,r} Y p , r be its integral. The linearity and monotonicity of Linearity and Monotonicity of the Lebesgue Integral §integrable now give
β L μ a , V ( g p , r ) ≥ 1 4 Y p , r − D − A p r . (4.1) \beta\,L^{a,V}_{\mu}(g_{p,r})\ge\tfrac14Y_{p,r}-D-\frac{A_{p}}{r}.\tag{4.1} β L μ a , V ( g p , r ) ≥ 4 1 Y p , r − D − r A p . ( 4.1 )
Step 5 (An upper bound for the gradient norm). Since g ∈ C b 2 ( R d ) g\in C^{2}_{b}(\mathbb{R}^{d}) g ∈ C b 2 ( R d ) is of class C 1 C^{1} C 1 by claim 2 of Euclidean Space is Open in Itself, and C k C^k C k Maps are Continuous , with g g g and its first partial derivatives bounded, g ∈ C b 1 ( R d ) g\in C^{1}_{b}(\mathbb{R}^{d}) g ∈ C b 1 ( R d ) (Bounded Continuously Differentiable Functions with Bounded Partial Derivatives on Euclidean Space §bounded ) and ( d , g ) (d,g) ( d , g ) is a representation of g ∘ p d g\circ p_{d} g ∘ p d . By The Noise Gradient of a Cylindrical Function and the Noise Tangent Space at a Probability Measure on a Hilbert Space §gradient and Bounded C^1 Cylindrical Functions: Linear Structure, the Gradient and the Partial Derivatives, Bounds and Integrability, and Density in the Square-Integrable Functions §partial , ∣ ∇ a ( g ∘ p d ) ( x ) ∣ a 2 = ∑ k = 1 d a k ( ∂ k g ( u ) ) 2 |\nabla_{a}(g\circ p_{d})(x)|_{a}^{2}=\sum_{k=1}^{d}a_{k}(\partial_{k}g(u))^{2} ∣ ∇ a ( g ∘ p d ) ( x ) ∣ a 2 = ∑ k = 1 d a k ( ∂ k g ( u ) ) 2 with u = p d ( x ) u=p_{d}(x) u = p d ( x ) , and by The Space of Square-Integrable Maps from a Measure Space into a Hilbert Space with an Orthonormal Basis §operations , ∥ ∇ a ( g ∘ p d ) ∥ μ 2 = ∫ X ∣ ∇ a ( g ∘ p d ) ∣ a 2 d μ \lVert\nabla_{a}(g\circ p_{d})\rVert_{\mu}^{2}=\int_{X}|\nabla_{a}(g\circ p_{d})|_{a}^{2}\,d\mu ∥ ∇ a ( g ∘ p d ) ∥ μ 2 = ∫ X ∣ ∇ a ( g ∘ p d ) ∣ a 2 d μ . By (3.1), ( ∂ k g ) 2 ≤ 2 w p 2 χ r 2 ( ∂ k v ) 2 + 2 F p 2 ( ∂ k χ r ) 2 ≤ 2 w p ( ∂ k v ) 2 + 2 p 2 M 1 2 r − 2 (\partial_{k}g)^{2}\le2w_{p}^{2}\chi_{r}^{2}(\partial_{k}v)^{2}+2F_{p}^{2}(\partial_{k}\chi_{r})^{2}\le2w_{p}(\partial_{k}v)^{2}+2p^{2}M_{1}^{2}r^{-2} ( ∂ k g ) 2 ≤ 2 w p 2 χ r 2 ( ∂ k v ) 2 + 2 F p 2 ( ∂ k χ r ) 2 ≤ 2 w p ( ∂ k v ) 2 + 2 p 2 M 1 2 r − 2 , using 0 < w p ≤ 1 0<w_{p}\le1 0 < w p ≤ 1 , 0 ≤ χ r ≤ 1 0\le\chi_{r}\le1 0 ≤ χ r ≤ 1 and 0 < F p ≤ p 0<F_{p}\le p 0 < F p ≤ p . Multiplying by a k a_{k} a k , summing, and integrating,
∥ ∇ a ( g p , r ∘ p d ) ∥ μ 2 ≤ 2 X p + 2 σ p 2 M 1 2 r − 2 , X p : = ∫ X w p ( p d ( x ) ) ∣ ∇ a v ( p d ( x ) ) ∣ 2 μ ( d x ) ≤ L 2 p 2 , \lVert\nabla_{a}(g_{p,r}\circ p_{d})\rVert_{\mu}^{2}\le2X_{p}+2\sigma p^{2}M_{1}^{2}r^{-2},\qquad X_{p}:=\int_{X}w_{p}(p_{d}(x))\,|\nabla_{a}v(p_{d}(x))|^{2}\,\mu(dx)\le L^{2}p^{2}, ∥ ∇ a ( g p , r ∘ p d ) ∥ μ 2 ≤ 2 X p + 2 σ p 2 M 1 2 r − 2 , X p := ∫ X w p ( p d ( x )) ∣ ∇ a v ( p d ( x )) ∣ 2 μ ( d x ) ≤ L 2 p 2 ,
the integrand of X p X_{p} X p lying between 0 0 0 and L 2 p 2 L^{2}p^{2} L 2 p 2 by (3.3). Taking square roots,
∥ ∇ a ( g p , r ∘ p d ) ∥ μ ≤ ( 2 X p ) 1 / 2 + ( 2 σ ) 1 / 2 p M 1 r − 1 . (5.1) \lVert\nabla_{a}(g_{p,r}\circ p_{d})\rVert_{\mu}\le(2X_{p})^{1/2}+(2\sigma)^{1/2}pM_{1}r^{-1}.\tag{5.1} ∥ ∇ a ( g p , r ∘ p d ) ∥ μ ≤ ( 2 X p ) 1/2 + ( 2 σ ) 1/2 p M 1 r − 1 . ( 5.1 )
Step 6 (Removing the cutoff). Fix p p p . By (3.4), (4.1) and (5.1), for every r ∈ N r\in\mathbb{N} r ∈ N ,
1 4 Y p , r ≤ D + β R ( 2 X p ) 1 / 2 + B p r , B p : = A p + β R ( 2 σ ) 1 / 2 p M 1 , (6.1) \tfrac14Y_{p,r}\le D+\beta R\,(2X_{p})^{1/2}+\frac{B_{p}}{r},\qquad B_{p}:=A_{p}+\beta R\,(2\sigma)^{1/2}pM_{1},\tag{6.1} 4 1 Y p , r ≤ D + βR ( 2 X p ) 1/2 + r B p , B p := A p + βR ( 2 σ ) 1/2 p M 1 , ( 6.1 )
where R ≥ 0 R\ge0 R ≥ 0 was used. For each x ∈ X x\in X x ∈ X there is, by The Archimedean Property of the Real Numbers , an r 0 ∈ N r_{0}\in\mathbb{N} r 0 ∈ N with ∥ p d ( x ) ∥ ≤ r 0 \lVert p_{d}(x)\rVert\le r_{0} ∥ p d ( x )∥ ≤ r 0 , and then χ r ( p d ( x ) ) = 1 \chi_{r}(p_{d}(x))=1 χ r ( p d ( x )) = 1 for all r ≥ r 0 r\ge r_{0} r ≥ r 0 ; so the integrands of Y p , r Y_{p,r} Y p , r converge pointwise, as r → ∞ r\to\infty r → ∞ , to that of X p X_{p} X p , and they are bounded by the integrable constant L 2 p 2 L^{2}p^{2} L 2 p 2 . By claim 3 of Dominated Convergence Theorem , Y p , r → X p Y_{p,r}\to X_{p} Y p , r → X p . Also B p / r → 0 B_{p}/r\to0 B p / r → 0 by The Archimedean Property of the Real Numbers and Limit of a Sequence of Real Numbers . Letting r → ∞ r\to\infty r → ∞ in (6.1), by Arithmetic of Limits of Real Sequences and the comparison claim (claim 1) of Order Properties of Limits of Real Sequences ,
1 4 X p ≤ D + β R ( 2 X p ) 1 / 2 . (6.2) \tfrac14X_{p}\le D+\beta R\,(2X_{p})^{1/2}.\tag{6.2} 4 1 X p ≤ D + βR ( 2 X p ) 1/2 . ( 6.2 )
Step 7 (Square-integrability of the noise gradient of V V V ). Put y = X p 1 / 2 ≥ 0 y=X_{p}^{1/2}\ge0 y = X p 1/2 ≥ 0 and c 1 = 4 β R 2 1 / 2 ≥ 0 c_{1}=4\beta R\,2^{1/2}\ge0 c 1 = 4 βR 2 1/2 ≥ 0 ; (6.2) reads y 2 ≤ 4 D + c 1 y y^{2}\le4D+c_{1}y y 2 ≤ 4 D + c 1 y . We claim y ≤ c 1 + 2 D 1 / 2 y\le c_{1}+2D^{1/2} y ≤ c 1 + 2 D 1/2 . Otherwise y > c 1 + 2 D 1 / 2 ≥ 0 y>c_{1}+2D^{1/2}\ge0 y > c 1 + 2 D 1/2 ≥ 0 , so y > 0 y>0 y > 0 , y ≥ 2 D 1 / 2 y\ge2D^{1/2} y ≥ 2 D 1/2 , and y 2 > ( c 1 + 2 D 1 / 2 ) y = c 1 y + 2 D 1 / 2 y ≥ c 1 y + 4 D y^{2}>(c_{1}+2D^{1/2})y=c_{1}y+2D^{1/2}y\ge c_{1}y+4D y 2 > ( c 1 + 2 D 1/2 ) y = c 1 y + 2 D 1/2 y ≥ c 1 y + 4 D , a contradiction. Hence, for every p ∈ N p\in\mathbb{N} p ∈ N ,
X p ≤ M ∗ : = ( c 1 + 2 D 1 / 2 ) 2 , X_{p}\le M_{*}:=(c_{1}+2D^{1/2})^{2}, X p ≤ M ∗ := ( c 1 + 2 D 1/2 ) 2 ,
a number not depending on p p p . By Step 2, for each x x x the sequence p ↦ w p ( p d ( x ) ) ∣ ∇ a v ( p d ( x ) ) ∣ 2 p\mapsto w_{p}(p_{d}(x))|\nabla_{a}v(p_{d}(x))|^{2} p ↦ w p ( p d ( x )) ∣ ∇ a v ( p d ( x )) ∣ 2 of nonnegative Borel functions is nondecreasing and converges to ∣ ∇ a v ( p d ( x ) ) ∣ 2 = ∣ ∇ a V ( x ) ∣ a 2 |\nabla_{a}v(p_{d}(x))|^{2}=|\nabla_{a}V(x)|_{a}^{2} ∣ ∇ a v ( p d ( x )) ∣ 2 = ∣ ∇ a V ( x ) ∣ a 2 (Arithmetic of Limits of Real Sequences ); a nondecreasing convergent real sequence has its limit as least upper bound (each term is at most the limit by claim 1 of Order Properties of Limits of Real Sequences applied to the tails, and every upper bound dominates the limit by the same claim). So Monotone Convergence Theorem gives
∫ X ∣ ∇ a V ∣ a 2 d μ = sup p ∈ N X p ≤ M ∗ < ∞ . (7.1) \int_{X}|\nabla_{a}V|_{a}^{2}\,d\mu=\sup_{p\in\mathbb{N}}X_{p}\le M_{*}<\infty.\tag{7.1} ∫ X ∣ ∇ a V ∣ a 2 d μ = p ∈ N sup X p ≤ M ∗ < ∞. ( 7.1 )
By Basic Properties of an Admissible Cylindrical Potential: Continuity, Growth under Noise Translations, Integrability, the Tangent Inequality and Tangency of the Noise Gradient §tangent , the class of ∇ a V \nabla_{a}V ∇ a V lies in L 2 ( μ ; X a ) L^{2}(\mu;X^{a}) L 2 ( μ ; X a ) and belongs to T μ a T^{a}_{\mu} T μ a ; write ∥ ∇ a V ∥ μ \lVert\nabla_{a}V\rVert_{\mu} ∥ ∇ a V ∥ μ for its norm.
Step 8 (Splitting the Gibbs functional). Let n ∈ N n\in\mathbb{N} n ∈ N and g ∈ C b 2 ( R n ) g\in C^{2}_{b}(\mathbb{R}^{n}) g ∈ C b 2 ( R n ) , and put ψ = g ∘ p n \psi=g\circ p_{n} ψ = g ∘ p n . As in Step 5, ( n , g ) (n,g) ( n , g ) is a representation of ψ ∈ F C b 1 ( X ) \psi\in\mathcal{F}C^{1}_{b}(X) ψ ∈ F C b 1 ( X ) , and by The Noise Gradient of a Cylindrical Function and the Noise Tangent Space at a Probability Measure on a Hilbert Space §gradient and Bounded C^1 Cylindrical Functions: Linear Structure, the Gradient and the Partial Derivatives, Bounds and Integrability, and Density in the Square-Integrable Functions §partial the k k k -th coordinate of ∇ a ψ ( x ) ∈ X a \nabla_{a}\psi(x)\in X^{a} ∇ a ψ ( x ) ∈ X a is a k ∂ k g ( p n ( x ) ) a_{k}\partial_{k}g(p_{n}(x)) a k ∂ k g ( p n ( x )) for k ≤ n k\le n k ≤ n and 0 0 0 for k > n k>n k > n . By Basic Properties of an Admissible Cylindrical Potential: Continuity, Growth under Noise Translations, Integrability, the Tangent Inequality and Tangency of the Noise Gradient §continuity , applied with h = ∇ a ψ ( x ) h=\nabla_{a}\psi(x) h = ∇ a ψ ( x ) , and since ∂ k V = 0 \partial_{k}V=0 ∂ k V = 0 for k > d k>d k > d (Admissible Cylindrical Potentials on a Hilbert Space in the Noise Geometry §gradient ),
⟨ ∇ a V ( x ) , ∇ a ψ ( x ) ⟩ a = ∑ k = 1 n a k ∂ k V ( x ) ∂ k g ( p n ( x ) ) ( x ∈ X ) . (8.1) \langle\nabla_{a}V(x),\nabla_{a}\psi(x)\rangle_{a}=\sum_{k=1}^{n}a_{k}\,\partial_{k}V(x)\,\partial_{k}g(p_{n}(x))\qquad(x\in X).\tag{8.1} ⟨ ∇ a V ( x ) , ∇ a ψ ( x ) ⟩ a = k = 1 ∑ n a k ∂ k V ( x ) ∂ k g ( p n ( x )) ( x ∈ X ) . ( 8.1 )
(The clause cited gives ∑ k = 1 d ∂ k V ( x ) h k \sum_{k=1}^{d}\partial_{k}V(x)\,h_{k} ∑ k = 1 d ∂ k V ( x ) h k , with h k h_{k} h k the k k k -th coordinate of ∇ a ψ ( x ) \nabla_{a}\psi(x) ∇ a ψ ( x ) ; this sum and the sum in (8.1) both reduce to the sum over k ≤ min ( d , n ) k\le\min(d,n) k ≤ min ( d , n ) : the terms with k > d k>d k > d vanish because ∂ k V = 0 \partial_{k}V=0 ∂ k V = 0 , and those with k > n k>n k > n because h k = 0 h_{k}=0 h k = 0 .) Each summand is integrable with respect to μ \mu μ by the preamble of The Gibbs Ornstein-Uhlenbeck Functional of a Probability Measure at a Bounded C^2 Function of Finitely Many Coordinates , so the function (8.1) is integrable by Linearity and Monotonicity of the Lebesgue Integral §integrable . The noise Ornstein-Uhlenbeck functional L μ a ( g ) L^{a}_{\mu}(g) L μ a ( g ) of The Noise Ornstein-Uhlenbeck Functional of a Probability Measure at a Bounded C^2 Function of Finitely Many Coordinates §functional is defined, μ \mu μ being in P 2 ( X ) \mathcal{P}_{2}(X) P 2 ( X ) , and for each k ∈ [ n ] k\in[n] k ∈ [ n ] the integrand of the k k k -th term of The Gibbs Ornstein-Uhlenbeck Functional of a Probability Measure at a Bounded C^2 Function of Finitely Many Coordinates §functional is the integrand of the k k k -th term of The Noise Ornstein-Uhlenbeck Functional of a Probability Measure at a Bounded C^2 Function of Finitely Many Coordinates §functional plus β − 1 ∂ k V ( ∂ k g ∘ p n ) \beta^{-1}\partial_{k}V\,(\partial_{k}g\circ p_{n}) β − 1 ∂ k V ( ∂ k g ∘ p n ) . By the linearity of Linearity and Monotonicity of the Lebesgue Integral §integrable and (8.1),
L μ a , V ( g ) = L μ a ( g ) + 1 β ∫ X ⟨ ∇ a V , ∇ a ψ ⟩ a d μ = L μ a ( g ) + 1 β ⟨ ∇ a V , ∇ a ψ ⟩ μ , (8.2) L^{a,V}_{\mu}(g)=L^{a}_{\mu}(g)+\frac{1}{\beta}\int_{X}\langle\nabla_{a}V,\nabla_{a}\psi\rangle_{a}\,d\mu=L^{a}_{\mu}(g)+\frac{1}{\beta}\,\langle\nabla_{a}V,\nabla_{a}\psi\rangle_{\mu},\tag{8.2} L μ a , V ( g ) = L μ a ( g ) + β 1 ∫ X ⟨ ∇ a V , ∇ a ψ ⟩ a d μ = L μ a ( g ) + β 1 ⟨ ∇ a V , ∇ a ψ ⟩ μ , ( 8.2 )
the second equality by The Space of Square-Integrable Maps from a Measure Space into a Hilbert Space with an Orthonormal Basis §operations , both ∇ a V \nabla_{a}V ∇ a V (Step 7) and ∇ a ψ \nabla_{a}\psi ∇ a ψ (The Noise Gradient of a Cylindrical Function and the Noise Tangent Space at a Probability Measure on a Hilbert Space §gradient ) being square-integrable with respect to μ \mu μ ; here ⟨ ⋅ , ⋅ ⟩ μ \langle\cdot,\cdot\rangle_{\mu} ⟨ ⋅ , ⋅ ⟩ μ is the inner product of L 2 ( μ ; X a ) L^{2}(\mu;X^{a}) L 2 ( μ ; X a ) .
Step 9 (Finite Fisher information relative to γ c \gamma_{c} γ c ). L 2 ( μ ; X a ) L^{2}(\mu;X^{a}) L 2 ( μ ; X a ) is a real Hilbert space by Probability Measures on a Hilbert Space Transported in the Noise Norm: Standing Notation §fields , so The Cauchy-Schwarz Inequality in a Real Inner Product Space gives ∣ ⟨ ∇ a V , ∇ a ψ ⟩ μ ∣ ≤ ∥ ∇ a V ∥ μ ∥ ∇ a ψ ∥ μ |\langle\nabla_{a}V,\nabla_{a}\psi\rangle_{\mu}|\le\lVert\nabla_{a}V\rVert_{\mu}\lVert\nabla_{a}\psi\rVert_{\mu} ∣ ⟨ ∇ a V , ∇ a ψ ⟩ μ ∣ ≤ ∥ ∇ a V ∥ μ ∥ ∇ a ψ ∥ μ . With (8.2) and the hypothesis,
∣ L μ a ( g ) ∣ ≤ ∣ L μ a , V ( g ) ∣ + β − 1 ∣ ⟨ ∇ a V , ∇ a ψ ⟩ μ ∣ ≤ ( R + β − 1 ∥ ∇ a V ∥ μ ) ∥ ∇ a ( g ∘ p n ) ∥ μ \bigl|L^{a}_{\mu}(g)\bigr|\le\bigl|L^{a,V}_{\mu}(g)\bigr|+\beta^{-1}\bigl|\langle\nabla_{a}V,\nabla_{a}\psi\rangle_{\mu}\bigr|\le\bigl(R+\beta^{-1}\lVert\nabla_{a}V\rVert_{\mu}\bigr)\lVert\nabla_{a}(g\circ p_{n})\rVert_{\mu} L μ a ( g ) ≤ L μ a , V ( g ) + β − 1 ⟨ ∇ a V , ∇ a ψ ⟩ μ ≤ ( R + β − 1 ∥ ∇ a V ∥ μ ) ∥ ∇ a ( g ∘ p n ) ∥ μ
for every n ∈ N n\in\mathbb{N} n ∈ N and every g ∈ C b 2 ( R n ) g\in C^{2}_{b}(\mathbb{R}^{n}) g ∈ C b 2 ( R n ) , with the nonnegative constant R + β − 1 ∥ ∇ a V ∥ μ R+\beta^{-1}\lVert\nabla_{a}V\rVert_{\mu} R + β − 1 ∥ ∇ a V ∥ μ . Since μ ∈ P 2 ( X ) \mu\in\mathcal{P}_{2}(X) μ ∈ P 2 ( X ) , A Bound on the Noise Ornstein-Uhlenbeck Functional by Noise Gradients Gives Finite Weighted Fisher Information §fisher shows that μ \mu μ has a relative score ( ζ k ) k ∈ N (\zeta_{k})_{k\in\mathbb{N}} ( ζ k ) k ∈ N with respect to γ c \gamma_{c} γ c and finite Fisher information relative to γ c \gamma_{c} γ c with weights a a a . Let Z μ a Z^{a}_{\mu} Z μ a be its noise score field (The Noise Score Field of a Measure of Finite Weighted Fisher Information: Existence, Norm, Pairing with Noise Gradients, Head Approximation and Tangency §field ).
Step 10 (The relative score with respect to the Gibbs measure). Now μ ∈ P 2 ( X ) \mu\in\mathcal{P}_{2}(X) μ ∈ P 2 ( X ) , V V V is integrable with respect to μ \mu μ , μ \mu μ has a relative score with respect to γ c \gamma_{c} γ c and finite Fisher information relative to γ c \gamma_{c} γ c with weights a a a (Step 9), and ∫ X ∣ ∇ a V ∣ a 2 d μ < ∞ \int_{X}|\nabla_{a}V|_{a}^{2}\,d\mu<\infty ∫ X ∣ ∇ a V ∣ a 2 d μ < ∞ by (7.1). By Entropy and Score Relative to the Gibbs Measure Split into Their Gaussian Parts and the Potential, with a Fisher Information Bound §splitting , μ \mu μ has a relative score with respect to γ β V \gamma^{V}_{\beta} γ β V and finite Fisher information relative to γ β V \gamma^{V}_{\beta} γ β V with weights a a a , which is the first assertion of the lemma. By Entropy and Score Relative to the Gibbs Measure Split into Their Gaussian Parts and the Potential, with a Fisher Information Bound §field , the element Ξ = β Z μ a + ∇ a V \Xi=\beta Z^{a}_{\mu}+\nabla_{a}V Ξ = β Z μ a + ∇ a V of L 2 ( μ ; X a ) L^{2}(\mu;X^{a}) L 2 ( μ ; X a ) satisfies
∥ Ξ ∥ μ 2 = β 2 I a ( μ ∣ γ β V ) . (10.1) \lVert\Xi\rVert_{\mu}^{2}=\beta^{2}\,\mathcal{I}_{a}(\mu\,|\,\gamma^{V}_{\beta}).\tag{10.1} ∥ Ξ ∥ μ 2 = β 2 I a ( μ ∣ γ β V ) . ( 10.1 )
Moreover Z μ a ∈ T μ a Z^{a}_{\mu}\in T^{a}_{\mu} Z μ a ∈ T μ a by The Noise Score Field of a Measure of Finite Weighted Fisher Information: Existence, Norm, Pairing with Noise Gradients, Head Approximation and Tangency §tangent , ∇ a V ∈ T μ a \nabla_{a}V\in T^{a}_{\mu} ∇ a V ∈ T μ a by Step 7, and T μ a T^{a}_{\mu} T μ a is a linear subspace of L 2 ( μ ; X a ) L^{2}(\mu;X^{a}) L 2 ( μ ; X a ) by Linearity of the Noise Gradient, and the Noise Tangent Space is a Closed Linear Subspace §subspace ; so Ξ ∈ T μ a \Xi\in T^{a}_{\mu} Ξ ∈ T μ a .
Step 11 (The bound I a ( μ ∣ γ β V ) ≤ R 2 \mathcal{I}_{a}(\mu\,|\,\gamma^{V}_{\beta})\le R^{2} I a ( μ ∣ γ β V ) ≤ R 2 ). Let n ∈ N n\in\mathbb{N} n ∈ N , g ∈ C b 2 ( R n ) g\in C^{2}_{b}(\mathbb{R}^{n}) g ∈ C b 2 ( R n ) and ψ = g ∘ p n \psi=g\circ p_{n} ψ = g ∘ p n . By bilinearity of the inner product, The Noise Score Field of a Measure of Finite Weighted Fisher Information: Existence, Norm, Pairing with Noise Gradients, Head Approximation and Tangency §pairing-functional and (8.2),
⟨ Ξ , ∇ a ψ ⟩ μ = β ⟨ Z μ a , ∇ a ψ ⟩ μ + ⟨ ∇ a V , ∇ a ψ ⟩ μ = β L μ a ( g ) + ⟨ ∇ a V , ∇ a ψ ⟩ μ = β L μ a , V ( g ) , \langle\Xi,\nabla_{a}\psi\rangle_{\mu}=\beta\langle Z^{a}_{\mu},\nabla_{a}\psi\rangle_{\mu}+\langle\nabla_{a}V,\nabla_{a}\psi\rangle_{\mu}=\beta L^{a}_{\mu}(g)+\langle\nabla_{a}V,\nabla_{a}\psi\rangle_{\mu}=\beta L^{a,V}_{\mu}(g), ⟨ Ξ , ∇ a ψ ⟩ μ = β ⟨ Z μ a , ∇ a ψ ⟩ μ + ⟨ ∇ a V , ∇ a ψ ⟩ μ = β L μ a ( g ) + ⟨ ∇ a V , ∇ a ψ ⟩ μ = β L μ a , V ( g ) ,
so by the hypothesis
⟨ Ξ , ∇ a ψ ⟩ μ ≤ ∣ ⟨ Ξ , ∇ a ψ ⟩ μ ∣ ≤ β R ∥ ∇ a ψ ∥ μ . (11.1) \langle\Xi,\nabla_{a}\psi\rangle_{\mu}\le\bigl|\langle\Xi,\nabla_{a}\psi\rangle_{\mu}\bigr|\le\beta R\,\lVert\nabla_{a}\psi\rVert_{\mu}.\tag{11.1} ⟨ Ξ , ∇ a ψ ⟩ μ ≤ ⟨ Ξ , ∇ a ψ ⟩ μ ≤ βR ∥ ∇ a ψ ∥ μ . ( 11.1 )
Since Ξ ∈ T μ a \Xi\in T^{a}_{\mu} Ξ ∈ T μ a , Noise Gradients of Bounded C^2 Cylindrical Functions Are Dense in the Noise Tangent Space §density yields ψ j ∈ F C b 2 ( X ) \psi_{j}\in\mathcal{F}C^{2}_{b}(X) ψ j ∈ F C b 2 ( X ) , j ∈ N j\in\mathbb{N} j ∈ N , with ∥ ∇ a ψ j − Ξ ∥ μ → 0 \lVert\nabla_{a}\psi_{j}-\Xi\rVert_{\mu}\to0 ∥ ∇ a ψ j − Ξ ∥ μ → 0 ; that is, ∇ a ψ j → Ξ \nabla_{a}\psi_{j}\to\Xi ∇ a ψ j → Ξ in the metric space L 2 ( μ ; X a ) L^{2}(\mu;X^{a}) L 2 ( μ ; X a ) , whose distance is ∥ ⋅ − ⋅ ∥ μ \lVert\cdot-\cdot\rVert_{\mu} ∥ ⋅ − ⋅ ∥ μ (Real Hilbert Space §topology ). By Bounded C^2 Cylindrical Functions on a Hilbert Space with an Orthonormal Basis §cylindrical , ψ j = g j ∘ p n j \psi_{j}=g_{j}\circ p_{n_{j}} ψ j = g j ∘ p n j with n j ∈ N n_{j}\in\mathbb{N} n j ∈ N and g j ∈ C b 2 ( R n j ) g_{j}\in C^{2}_{b}(\mathbb{R}^{n_{j}}) g j ∈ C b 2 ( R n j ) , so (11.1) holds for each ψ j \psi_{j} ψ j . By The Norm Metric of a Real Inner Product Space: Triangle Inequalities, Limits and Continuity §continuity , ⟨ ∇ a ψ j , Ξ ⟩ μ → ⟨ Ξ , Ξ ⟩ μ = ∥ Ξ ∥ μ 2 \langle\nabla_{a}\psi_{j},\Xi\rangle_{\mu}\to\langle\Xi,\Xi\rangle_{\mu}=\lVert\Xi\rVert_{\mu}^{2} ⟨ ∇ a ψ j , Ξ ⟩ μ → ⟨ Ξ , Ξ ⟩ μ = ∥ Ξ ∥ μ 2 and ∥ ∇ a ψ j ∥ μ → ∥ Ξ ∥ μ \lVert\nabla_{a}\psi_{j}\rVert_{\mu}\to\lVert\Xi\rVert_{\mu} ∥ ∇ a ψ j ∥ μ → ∥ Ξ ∥ μ . Letting j → ∞ j\to\infty j → ∞ in (11.1), by Arithmetic of Limits of Real Sequences and claim 1 of Order Properties of Limits of Real Sequences , ∥ Ξ ∥ μ 2 ≤ β R ∥ Ξ ∥ μ \lVert\Xi\rVert_{\mu}^{2}\le\beta R\,\lVert\Xi\rVert_{\mu} ∥ Ξ ∥ μ 2 ≤ βR ∥ Ξ ∥ μ . If ∥ Ξ ∥ μ > 0 \lVert\Xi\rVert_{\mu}>0 ∥ Ξ ∥ μ > 0 , dividing gives ∥ Ξ ∥ μ ≤ β R \lVert\Xi\rVert_{\mu}\le\beta R ∥ Ξ ∥ μ ≤ βR ; if ∥ Ξ ∥ μ = 0 \lVert\Xi\rVert_{\mu}=0 ∥ Ξ ∥ μ = 0 , then ∥ Ξ ∥ μ ≤ β R \lVert\Xi\rVert_{\mu}\le\beta R ∥ Ξ ∥ μ ≤ βR as β R ≥ 0 \beta R\ge0 βR ≥ 0 . Hence ∥ Ξ ∥ μ 2 ≤ β 2 R 2 \lVert\Xi\rVert_{\mu}^{2}\le\beta^{2}R^{2} ∥ Ξ ∥ μ 2 ≤ β 2 R 2 , and by (10.1), dividing by β 2 > 0 \beta^{2}>0 β 2 > 0 ,
I a ( μ ∣ γ β V ) ≤ R 2 . ■ \mathcal{I}_{a}(\mu\,|\,\gamma^{V}_{\beta})\le R^{2}.\qquad\blacksquare I a ( μ ∣ γ β V ) ≤ R 2 . ■