Each result cited is universally quantified over the data in its own statement. Throughout, a function of class C 1 C^{1} C 1 on R d \mathbb{R}^{d} R d and its partial derivatives are continuous (clause 1 of C^k Maps on a Euclidean Open Set , claim 3 of Euclidean Space is Open in Itself, and C k C^k C k Maps are Continuous ), hence Borel (claim 3 of The Borel Sigma-Algebra of a Euclidean Space as a Product, and Measurability of Projections, Sequentially Continuous Maps, and Open and Closed Sets ); a map R n → R d \mathbb{R}^{n}\to\mathbb{R}^{d} R n → R d is Borel when its components are, and compositions of Borel maps are Borel (Probability Measures on Euclidean Space and Random Vectors: Standing Notation §borel-maps ); a bounded Borel function is integrable with respect to every probability measure (Probability Measures on Euclidean Space and Random Vectors: Standing Notation §measures ); the integral of a constant against a probability measure is that constant (Simple Function and Its Integral ); "change of variables" refers to Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pushforward ; ∥ x ∥ 2 = ∑ i x i 2 \lVert x\rVert^{2}=\sum_{i}x_{i}^{2} ∥ x ∥ 2 = ∑ i x i 2 and ∣ x i ∣ ≤ ∥ x ∥ |x_{i}|\le\lVert x\rVert ∣ x i ∣ ≤ ∥ x ∥ (claims 1 and 4 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n ); and for u , w ∈ R d u,w\in\mathbb{R}^{d} u , w ∈ R d , ∣ u ⋅ w ∣ ≤ ∥ u ∥ ∥ w ∥ |u\cdot w|\le\lVert u\rVert\lVert w\rVert ∣ u ⋅ w ∣ ≤ ∥ u ∥ ∥ w ∥ and ∥ u − w ∥ 2 ≤ 2 ∥ u ∥ 2 + 2 ∥ w ∥ 2 \lVert u-w\rVert^{2}\le2\lVert u\rVert^{2}+2\lVert w\rVert^{2} ∥ u − w ∥ 2 ≤ 2 ∥ u ∥ 2 + 2 ∥ w ∥ 2 (Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions ). For a function g g g of class C 1 C^{1} C 1 on R d \mathbb{R}^{d} R d with ∣ ∂ i g ∣ ≤ C |\partial_{i}g|\le C ∣ ∂ i g ∣ ≤ C for all i i i , part (i) of Multivariate Taylor Expansion with Uniform Second-Order Remainder (the segment between any two points lies in R d \mathbb{R}^{d} R d , and the Euclidean distance of u u u and w w w is ∥ u − w ∥ \lVert u-w\rVert ∥ u − w ∥ by claim 2 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n ) gives
∣ g ( u ) − g ( w ) ∣ ≤ d C ∥ u − w ∥ ( u , w ∈ R d ) , (Lip) |g(u)-g(w)|\le\sqrt{d}\,C\,\lVert u-w\rVert\qquad(u,w\in\mathbb{R}^{d}),\tag{Lip} ∣ g ( u ) − g ( w ) ∣ ≤ d C ∥ u − w ∥ ( u , w ∈ R d ) , ( Lip )
and if g g g is of class C 2 C^{2} C 2 with ∣ ∂ j ∂ i g ∣ ≤ C |\partial_{j}\partial_{i}g|\le C ∣ ∂ j ∂ i g ∣ ≤ C , part (ii) gives, with D g ( u ) ⋅ w = ∑ i ∂ i g ( u ) w i Dg(u)\cdot w=\sum_{i}\partial_{i}g(u)w_{i} D g ( u ) ⋅ w = ∑ i ∂ i g ( u ) w i ,
∣ g ( u + w ) − g ( u ) − D g ( u ) ⋅ w ∣ ≤ 1 2 d C ∥ w ∥ 2 ( u , w ∈ R d ) . (Tay) |g(u+w)-g(u)-Dg(u)\cdot w|\le\tfrac12\,d\,C\,\lVert w\rVert^{2}\qquad(u,w\in\mathbb{R}^{d}).\tag{Tay} ∣ g ( u + w ) − g ( u ) − D g ( u ) ⋅ w ∣ ≤ 2 1 d C ∥ w ∥ 2 ( u , w ∈ R d ) . ( Tay )
For a Borel map Γ : R d → R d \Gamma:\mathbb{R}^{d}\to\mathbb{R}^{d} Γ : R d → R d with ∣ Γ i ∣ ≤ C |\Gamma_{i}|\le C ∣ Γ i ∣ ≤ C one has ∥ Γ ∥ 2 ≤ d C 2 \lVert\Gamma\rVert^{2}\le dC^{2} ∥ Γ ∥ 2 ≤ d C 2 (claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field , claim 1 of Comparison and Absolute Value Bounds for Finite Sums of Real Numbers ), so its class lies in L 2 ( μ ; R d ) L^{2}(\mu;\mathbb{R}^{d}) L 2 ( μ ; R d ) for every μ ∈ P ( R d ) \mu\in\mathcal{P}(\mathbb{R}^{d}) μ ∈ P ( R d ) ; for X ∈ L 2 ( Ω ; R d ) X\in L^{2}(\Omega;\mathbb{R}^{d}) X ∈ L 2 ( Ω ; R d ) with L ( X ) = μ \mathcal{L}(X)=\mu L ( X ) = μ , Γ ∘ X \Gamma\circ X Γ ∘ X is the class of Composition of a Square-Integrable Vector Field with a Random Vector, and the Lifted Score: Isometry, Norm, Weak Identity and Second-Moment Identity §composition , and by that clause and Basic Properties of Random Vectors: Coordinates, Borel Images and Arithmetic, Change of Variables, Almost Sure Equality and Pairs §expectation , ∥ Γ ∘ X ∥ L 2 2 = ∫ ∥ Γ ∥ 2 d μ \lVert\Gamma\circ X\rVert_{L^{2}}^{2}=\int\lVert\Gamma\rVert^{2}d\mu ∥ Γ ∘ X ∥ L 2 2 = ∫ ∥ Γ ∥ 2 d μ . Finally, for V ∈ L 2 ( Ω ; R d ) V\in L^{2}(\Omega;\mathbb{R}^{d}) V ∈ L 2 ( Ω ; R d ) with a representative,
E [ ∥ V ∥ ] ≤ ∥ V ∥ L 2 , (H) \mathbb{E}[\lVert V\rVert]\le\lVert V\rVert_{L^{2}},\tag{H} E [∥ V ∥] ≤ ∥ V ∥ L 2 , ( H )
by Hoelder's Inequality, for Two and for Finitely Many Factors §holder with p = q = 2 p=q=2 p = q = 2 (conjugate exponents in the sense of Conjugate Exponents and Young's Inequality §conjugate , as 1 2 + 1 2 = 1 \tfrac12+\tfrac12=1 2 1 + 2 1 = 1 ) applied to ∥ V ∥ \lVert V\rVert ∥ V ∥ and the constant 1 1 1 : E [ ∥ V ∥ ] ≤ ( E [ ∥ V ∥ 2 ] ) 1 / 2 = ∥ V ∥ L 2 \mathbb{E}[\lVert V\rVert]\le(\mathbb{E}[\lVert V\rVert^{2}])^{1/2}=\lVert V\rVert_{L^{2}} E [∥ V ∥] ≤ ( E [∥ V ∥ 2 ] ) 1/2 = ∥ V ∥ L 2 (The Space of Square-Integrable Random Vectors §inner-product ).
Iterated integrals (FB). Let p , r ∈ N p,r\in\mathbb{N} p , r ∈ N with 1 ≤ p , r 1\le p,r 1 ≤ p , r , α ∈ P ( R p ) \alpha\in\mathcal{P}(\mathbb{R}^{p}) α ∈ P ( R p ) , β ∈ P ( R r ) \beta\in\mathcal{P}(\mathbb{R}^{r}) β ∈ P ( R r ) , and let F : R p + r → R F:\mathbb{R}^{p+r}\to\mathbb{R} F : R p + r → R be Borel with ∣ F ( ι ( x , y ) ) ∣ ≤ g ( x ) + h ( y ) |F(\iota(x,y))|\le g(x)+h(y) ∣ F ( ι ( x , y )) ∣ ≤ g ( x ) + h ( y ) for all x , y x,y x , y , where g : R p → [ 0 , ∞ ) g:\mathbb{R}^{p}\to[0,\infty) g : R p → [ 0 , ∞ ) and h : R r → [ 0 , ∞ ) h:\mathbb{R}^{r}\to[0,\infty) h : R r → [ 0 , ∞ ) are Borel and integrable with respect to α \alpha α and β \beta β . Then: for every x x x the function y ↦ F ( ι ( x , y ) ) y\mapsto F(\iota(x,y)) y ↦ F ( ι ( x , y )) is Borel and β \beta β -integrable; the function x ↦ ∫ F ( ι ( x , y ) ) β ( d y ) x\mapsto\int F(\iota(x,y))\beta(dy) x ↦ ∫ F ( ι ( x , y )) β ( d y ) is Borel and α \alpha α -integrable; symmetrically in the other order; F F F is integrable with respect to α ⊠ β \alpha\boxtimes\beta α ⊠ β ; and
∫ R p + r F d ( α ⊠ β ) = ∫ ( ∫ F ( ι ( x , y ) ) β ( d y ) ) α ( d x ) = ∫ ( ∫ F ( ι ( x , y ) ) α ( d x ) ) β ( d y ) . \int_{\mathbb{R}^{p+r}}F\,d(\alpha\boxtimes\beta)=\int\Bigl(\int F(\iota(x,y))\,\beta(dy)\Bigr)\alpha(dx)=\int\Bigl(\int F(\iota(x,y))\,\alpha(dx)\Bigr)\beta(dy). ∫ R p + r F d ( α ⊠ β ) = ∫ ( ∫ F ( ι ( x , y )) β ( d y ) ) α ( d x ) = ∫ ( ∫ F ( ι ( x , y )) α ( d x ) ) β ( d y ) .
Indeed, let F + = max ( F , 0 ) F^{+}=\max(F,0) F + = max ( F , 0 ) and F − = max ( − F , 0 ) F^{-}=\max(-F,0) F − = max ( − F , 0 ) , nonnegative Borel functions (claim 4 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions ) with F = F + − F − F=F^{+}-F^{-} F = F + − F − and F ± ≤ ∣ F ∣ F^{\pm}\le|F| F ± ≤ ∣ F ∣ . By Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §product , F ± ∘ ι F^{\pm}\circ\iota F ± ∘ ι is measurable for the product σ \sigma σ -algebra and ∫ F ± d ( α ⊠ β ) = ∫ F ± ∘ ι d ( α ⊗ β ) \int F^{\pm}d(\alpha\boxtimes\beta)=\int F^{\pm}\circ\iota\,d(\alpha\otimes\beta) ∫ F ± d ( α ⊠ β ) = ∫ F ± ∘ ι d ( α ⊗ β ) ; by Tonelli and Fubini Theorems (Sections and Tonelli), the sections of F ± ∘ ι F^{\pm}\circ\iota F ± ∘ ι are measurable, the functions x ↦ ∫ F ± ( ι ( x , y ) ) β ( d y ) x\mapsto\int F^{\pm}(\iota(x,y))\beta(dy) x ↦ ∫ F ± ( ι ( x , y )) β ( d y ) and y ↦ ∫ F ± ( ι ( x , y ) ) α ( d x ) y\mapsto\int F^{\pm}(\iota(x,y))\alpha(dx) y ↦ ∫ F ± ( ι ( x , y )) α ( d x ) are measurable, and the three integrals of F ± F^{\pm} F ± (product, and the two iterated orders) coincide in [ 0 , ∞ ] [0,\infty] [ 0 , ∞ ] . By the domination and claim 1 of Linearity and Monotonicity of the Lebesgue Integral , ∫ F ± ( ι ( x , y ) ) β ( d y ) ≤ g ( x ) + ∫ h d β < ∞ \int F^{\pm}(\iota(x,y))\beta(dy)\le g(x)+\int h\,d\beta<\infty ∫ F ± ( ι ( x , y )) β ( d y ) ≤ g ( x ) + ∫ h d β < ∞ for every x x x , and the common value of the three integrals is at most ∫ g d α + ∫ h d β < ∞ \int g\,d\alpha+\int h\,d\beta<\infty ∫ g d α + ∫ h d β < ∞ . Hence every section is integrable, the functions x ↦ ∫ F ( ι ( x , y ) ) β ( d y ) = ∫ F + ( ι ( x , y ) ) β ( d y ) − ∫ F − ( ι ( x , y ) ) β ( d y ) x\mapsto\int F(\iota(x,y))\beta(dy)=\int F^{+}(\iota(x,y))\beta(dy)-\int F^{-}(\iota(x,y))\beta(dy) x ↦ ∫ F ( ι ( x , y )) β ( d y ) = ∫ F + ( ι ( x , y )) β ( d y ) − ∫ F − ( ι ( x , y )) β ( d y ) are real-valued, measurable (claim 2 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions ) and integrable, F F F is α ⊠ β \alpha\boxtimes\beta α ⊠ β -integrable (Integrable Function and the Lebesgue Integral ), and the displayed identities follow by subtracting the identities for F − F^{-} F − from those for F + F^{+} F + , using claim 2 of Linearity and Monotonicity of the Lebesgue Integral and the definition of the integral of an integrable function as the difference of the integrals of its positive and negative parts. A bounded Borel F F F with ∣ F ∣ ≤ C |F|\le C ∣ F ∣ ≤ C satisfies the hypothesis with g = C g=C g = C , h = 0 h=0 h = 0 .
Proof of claim 1. Let f f f be as in claim 1 and U U U the lift of u f u_{f} u f . Apply The Lift of a Linear Functional of the Measure: Integrability, L-Gradient, Lipschitz Gradient Map and Translation Laplacian with M 1 = M 2 = M M_{1}=M_{2}=M M 1 = M 2 = M .
(a) u f u_{f} u f is continuously L L L -differentiable by The Lift of a Linear Functional of the Measure: Integrability, L-Gradient, Lipschitz Gradient Map and Translation Laplacian §lipschitz , i.e. U ∈ C 1 ( L 2 ( Ω ; R d ) ) U\in C^{1}(L^{2}(\Omega;\mathbb{R}^{d})) U ∈ C 1 ( L 2 ( Ω ; R d )) (L-Differentiability of a Function on the Wasserstein Space via the Fréchet Derivative of Its Lift §c1 ).
(b) By The Lift of a Linear Functional of the Measure: Integrability, L-Gradient, Lipschitz Gradient Map and Translation Laplacian §derivative , D U ( X ) = D f ∘ X DU(X)=Df\circ X D U ( X ) = D f ∘ X for every X X X , the class of the composition of D f Df D f with a representative of X X X , which is the class D f ∘ X Df\circ X D f ∘ X of Composition of a Square-Integrable Vector Field with a Random Vector, and the Lifted Score: Isometry, Norm, Weak Identity and Second-Moment Identity §composition for the class D f ∈ L 2 ( L ( X ) ; R d ) Df\in L^{2}(\mathcal{L}(X);\mathbb{R}^{d}) D f ∈ L 2 ( L ( X ) ; R d ) . For μ ∈ P 2 ( R d ) \mu\in\mathcal{P}_{2}(\mathbb{R}^{d}) μ ∈ P 2 ( R d ) the class D f Df D f lies in T μ T_{\mu} T μ by Gradients of Functions with Bounded Derivatives Belong to the Tangent Space; Constants Are Tangent; the Score Identity; the Score Has Mean Zero; Translation of Tangent Fields §tangent , so property (b) holds with η = D f \eta=Df η = D f , and ∇ u f ( μ ) \nabla u_{f}(\mu) ∇ u f ( μ ) is the class of x ↦ D f ( x ) x\mapsto Df(x) x ↦ D f ( x ) (Test Functions on the Wasserstein Space: the Intrinsic Gradient and the Translation Hessian §gradient ).
(c) U U U is twice continuously differentiable along translations by The Lift of a Linear Functional of the Measure: Integrability, L-Gradient, Lipschitz Gradient Map and Translation Laplacian §translation .
(d) and the Lipschitz constant. Let X , Y ∈ L 2 ( Ω ; R d ) X,Y\in L^{2}(\Omega;\mathbb{R}^{d}) X , Y ∈ L 2 ( Ω ; R d ) with representatives. By The Lift of a Linear Functional of the Measure: Integrability, L-Gradient, Lipschitz Gradient Map and Translation Laplacian §integrable , U ( X ) − U ( Y ) = E [ f ∘ X ] − E [ f ∘ Y ] = E [ f ∘ X − f ∘ Y ] U(X)-U(Y)=\mathbb{E}[f\circ X]-\mathbb{E}[f\circ Y]=\mathbb{E}[f\circ X-f\circ Y] U ( X ) − U ( Y ) = E [ f ∘ X ] − E [ f ∘ Y ] = E [ f ∘ X − f ∘ Y ] , and by (Lip), ∣ f ( X ( ω ) ) − f ( Y ( ω ) ) ∣ ≤ d M ∥ X ( ω ) − Y ( ω ) ∥ |f(X(\omega))-f(Y(\omega))|\le\sqrt{d}\,M\lVert X(\omega)-Y(\omega)\rVert ∣ f ( X ( ω )) − f ( Y ( ω )) ∣ ≤ d M ∥ X ( ω ) − Y ( ω )∥ for every ω \omega ω , so claim 2 of Linearity and Monotonicity of the Lebesgue Integral and (H) give ∣ U ( X ) − U ( Y ) ∣ ≤ d M E [ ∥ X − Y ∥ ] ≤ d M ∥ X − Y ∥ L 2 |U(X)-U(Y)|\le\sqrt{d}\,M\,\mathbb{E}[\lVert X-Y\rVert]\le\sqrt{d}\,M\lVert X-Y\rVert_{L^{2}} ∣ U ( X ) − U ( Y ) ∣ ≤ d M E [∥ X − Y ∥] ≤ d M ∥ X − Y ∥ L 2 . Thus U U U is Lipschitz with constant d M \sqrt{d}\,M d M (Lipschitz Map Between Metric Spaces , with the distance d L 2 d_{L^{2}} d L 2 of The Wasserstein Space and Its Lift to Square-Integrable Random Vectors: Standing Notation §space ), so u f u_{f} u f is Lipschitz with constant d M \sqrt{d}\,M d M by Basic Properties of the Lift: Law Invariance, the Correspondence on a Rich Space, and Transfer of Boundedness, Lipschitz Constants and Continuity §lipschitz-converse , the space being rich, and continuous by A Lipschitz Map is Uniformly Continuous . Hence u f u_{f} u f is a test function.
The translation Hessian. Let μ ∈ P 2 ( R d ) \mu\in\mathcal{P}_{2}(\mathbb{R}^{d}) μ ∈ P 2 ( R d ) , X X X with L ( X ) = μ \mathcal{L}(X)=\mu L ( X ) = μ (The Wasserstein Distance and the Mean-Square Distance of Random Vectors §onto ), and ϕ X ( a ) = U ( X + c a ) \phi_{X}(a)=U(X+c_{a}) ϕ X ( a ) = U ( X + c a ) . By The Wasserstein Space and Its Lift to Square-Integrable Random Vectors: Standing Notation §constants , The Lift of a Linear Functional of the Measure: Integrability, L-Gradient, Lipschitz Gradient Map and Translation Laplacian §integrable , Basic Properties of Random Vectors: Coordinates, Borel Images and Arithmetic, Change of Variables, Almost Sure Equality and Pairs §expectation and change of variables, ϕ X ( a ) = ∫ f d ( τ a ) # μ = ∫ f ( x + a ) μ ( d x ) \phi_{X}(a)=\int f\,d(\tau_{a})_{\#}\mu=\int f(x+a)\,\mu(dx) ϕ X ( a ) = ∫ f d ( τ a ) # μ = ∫ f ( x + a ) μ ( d x ) . Fix a a a and i ∈ [ d ] i\in[d] i ∈ [ d ] . For t ∈ ( − 1 , 1 ) t\in(-1,1) t ∈ ( − 1 , 1 ) and x ∈ R d x\in\mathbb{R}^{d} x ∈ R d put F ( t , x ) = f ( x + a + t e i ) F(t,x)=f(x+a+te_{i}) F ( t , x ) = f ( x + a + t e i ) . Each x ↦ F ( t , x ) x\mapsto F(t,x) x ↦ F ( t , x ) is integrable (The Lift of a Linear Functional of the Measure: Integrability, L-Gradient, Lipschitz Gradient Map and Translation Laplacian §integrable for ( τ a + t e i ) # μ ∈ P 2 ( R d ) (\tau_{a+te_{i}})_{\#}\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}) ( τ a + t e i ) # μ ∈ P 2 ( R d ) , Gradients of Functions with Bounded Derivatives Belong to the Tangent Space; Constants Are Tangent; the Score Identity; the Score Has Mean Zero; Translation of Tangent Fields §translation , and change of variables); for each x x x , t ↦ F ( t , x ) t\mapsto F(t,x) t ↦ F ( t , x ) is differentiable at every t t t with derivative ∂ i f ( x + a + t e i ) \partial_{i}f(x+a+te_{i}) ∂ i f ( x + a + t e i ) by Chain Rule Along an Affine Path and A Real-Valued C^1 Function is Differentiable at Every Point ; and ∣ ∂ i f ∣ ≤ M |\partial_{i}f|\le M ∣ ∂ i f ∣ ≤ M , a constant integrable against μ \mu μ . By Differentiation under the Integral Sign , t ↦ ϕ X ( a + t e i ) t\mapsto\phi_{X}(a+te_{i}) t ↦ ϕ X ( a + t e i ) is differentiable at 0 0 0 with derivative ∫ ∂ i f ( x + a ) μ ( d x ) \int\partial_{i}f(x+a)\mu(dx) ∫ ∂ i f ( x + a ) μ ( d x ) ; its difference quotients at 0 0 0 are those of Partial Derivative on a Euclidean Open Set for ϕ X \phi_{X} ϕ X at a a a , so ∂ i ϕ X ( a ) = ∫ ∂ i f ( x + a ) μ ( d x ) \partial_{i}\phi_{X}(a)=\int\partial_{i}f(x+a)\,\mu(dx) ∂ i ϕ X ( a ) = ∫ ∂ i f ( x + a ) μ ( d x ) . Let μ ˇ = ( − i d ) # μ \check{\mu}=(-\mathrm{id})_{\#}\mu μ ˇ = ( − id ) # μ , the push-forward under the map x ↦ − x x\mapsto-x x ↦ − x , which is Borel by the componentwise criterion and claim 2 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions ; change of variables gives ∂ i ϕ X ( a ) = ∫ ∂ i f ( a − y ) μ ˇ ( d y ) = ( ( ∂ i f ) ∗ μ ˇ ) ( a ) \partial_{i}\phi_{X}(a)=\int\partial_{i}f(a-y)\,\check{\mu}(dy)=((\partial_{i}f)*\check{\mu})(a) ∂ i ϕ X ( a ) = ∫ ∂ i f ( a − y ) μ ˇ ( d y ) = (( ∂ i f ) ∗ μ ˇ ) ( a ) in the notation of The Convolution of a Bounded Continuous Function with a Probability Measure on Euclidean Space: Continuity, Boundedness and Differentiation Under the Integral . Since ∂ i f \partial_{i}f ∂ i f is of class C 1 C^{1} C 1 with ∣ ∂ i f ∣ ≤ M |\partial_{i}f|\le M ∣ ∂ i f ∣ ≤ M and ∣ ∂ j ∂ i f ∣ ≤ M |\partial_{j}\partial_{i}f|\le M ∣ ∂ j ∂ i f ∣ ≤ M , The Convolution of a Bounded Continuous Function with a Probability Measure on Euclidean Space: Continuity, Boundedness and Differentiation Under the Integral §derivative shows that ∂ i ϕ X \partial_{i}\phi_{X} ∂ i ϕ X is of class C 1 C^{1} C 1 on R d \mathbb{R}^{d} R d with ∂ j ∂ i ϕ X ( a ) = ( ( ∂ j ∂ i f ) ∗ μ ˇ ) ( a ) = ∫ ∂ j ∂ i f ( x + a ) μ ( d x ) \partial_{j}\partial_{i}\phi_{X}(a)=((\partial_{j}\partial_{i}f)*\check{\mu})(a)=\int\partial_{j}\partial_{i}f(x+a)\,\mu(dx) ∂ j ∂ i ϕ X ( a ) = (( ∂ j ∂ i f ) ∗ μ ˇ ) ( a ) = ∫ ∂ j ∂ i f ( x + a ) μ ( d x ) (change of variables back). Also ϕ X \phi_{X} ϕ X is continuous, since ∣ ϕ X ( a ) − ϕ X ( a ′ ) ∣ ≤ ∫ ∣ f ( x + a ) − f ( x + a ′ ) ∣ μ ( d x ) ≤ d M ∥ a − a ′ ∥ |\phi_{X}(a)-\phi_{X}(a')|\le\int|f(x+a)-f(x+a')|\mu(dx)\le\sqrt{d}\,M\lVert a-a'\rVert ∣ ϕ X ( a ) − ϕ X ( a ′ ) ∣ ≤ ∫ ∣ f ( x + a ) − f ( x + a ′ ) ∣ μ ( d x ) ≤ d M ∥ a − a ′ ∥ by (Lip) and claim 2 of Linearity and Monotonicity of the Lebesgue Integral , so ϕ X \phi_{X} ϕ X is of class C 2 C^{2} C 2 (C^k Maps on a Euclidean Open Set ), and by Hessian Matrix of a C^2 Function and Test Functions on the Wasserstein Space: the Intrinsic Gradient and the Translation Hessian §hessian , H u f ( μ ) = D 2 ϕ X ( 0 R d ) H_{u_{f}}(\mu)=D^{2}\phi_{X}(0_{\mathbb{R}^{d}}) H u f ( μ ) = D 2 ϕ X ( 0 R d ) has entry ∂ i ∂ j ϕ X ( 0 R d ) = ∫ ∂ i ∂ j f d μ \partial_{i}\partial_{j}\phi_{X}(0_{\mathbb{R}^{d}})=\int\partial_{i}\partial_{j}f\,d\mu ∂ i ∂ j ϕ X ( 0 R d ) = ∫ ∂ i ∂ j f d μ in row i i i and column j j j ; since this matrix belongs to S ( d ) \mathcal{S}(d) S ( d ) by Test Functions on the Wasserstein Space: the Intrinsic Gradient and the Translation Hessian §hessian , its entry in row i i i and column j j j equals its entry in row j j j and column i i i , namely ∫ ∂ j ∂ i f d μ \int\partial_{j}\partial_{i}f\,d\mu ∫ ∂ j ∂ i f d μ , as stated; each ∂ j ∂ i f \partial_{j}\partial_{i}f ∂ j ∂ i f is bounded and Borel as noted.
Proof of claim 2. Let K K K be as in claim 2.
Oddness of D K DK DK . Let x ∈ R d x\in\mathbb{R}^{d} x ∈ R d , i ∈ [ d ] i\in[d] i ∈ [ d ] . For h ≠ 0 h\ne0 h = 0 , evenness gives ( K ( x + h e i ) − K ( x ) ) / h = ( K ( − x + ( − h ) e i ) − K ( − x ) ) / h = − ( K ( − x + ( − h ) e i ) − K ( − x ) ) / ( − h ) \bigl(K(x+he_{i})-K(x)\bigr)/h=\bigl(K(-x+(-h)e_{i})-K(-x)\bigr)/h=-\bigl(K(-x+(-h)e_{i})-K(-x)\bigr)/(-h) ( K ( x + h e i ) − K ( x ) ) / h = ( K ( − x + ( − h ) e i ) − K ( − x ) ) / h = − ( K ( − x + ( − h ) e i ) − K ( − x ) ) / ( − h ) . Given ε > 0 \varepsilon>0 ε > 0 , let δ \delta δ and δ ′ \delta' δ ′ be the radii of Partial Derivative on a Euclidean Open Set for K K K at − x -x − x and at x x x with ε / 2 \varepsilon/2 ε /2 in place of ε \varepsilon ε (claim 8 of Elementary Order Arithmetic in an Ordered Field ), and let h h h satisfy 0 < ∣ h ∣ < min ( δ , δ ′ ) 0<|h|<\min(\delta,\delta') 0 < ∣ h ∣ < min ( δ , δ ′ ) (claim 9 there; then 0 < ∣ − h ∣ < δ 0<|-h|<\delta 0 < ∣ − h ∣ < δ by claim 2 of Properties of the Absolute Value in an Ordered Field ). The left side lies within ε / 2 \varepsilon/2 ε /2 of ∂ i K ( x ) \partial_{i}K(x) ∂ i K ( x ) and the right side within ε / 2 \varepsilon/2 ε /2 of − ∂ i K ( − x ) -\partial_{i}K(-x) − ∂ i K ( − x ) , so ∣ ∂ i K ( x ) + ∂ i K ( − x ) ∣ < ε |\partial_{i}K(x)+\partial_{i}K(-x)|<\varepsilon ∣ ∂ i K ( x ) + ∂ i K ( − x ) ∣ < ε by claim 5 of Properties of the Absolute Value in an Ordered Field ; as ε > 0 \varepsilon>0 ε > 0 was arbitrary, ∂ i K ( x ) = − ∂ i K ( − x ) \partial_{i}K(x)=-\partial_{i}K(-x) ∂ i K ( x ) = − ∂ i K ( − x ) by Comparison of Real Numbers with Arbitrary Positive Slack §vanishing .
Borel functions and the potential. K K K , ∂ i K \partial_{i}K ∂ i K and ∂ j ∂ i K \partial_{j}\partial_{i}K ∂ j ∂ i K are continuous, hence Borel. The map z ↦ p r 1 ( z ) − p r 2 ( z ) z\mapsto\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z) z ↦ pr 1 ( z ) − pr 2 ( z ) on R d + d \mathbb{R}^{d+d} R d + d is Borel (its components are differences of the Borel components of the projections, Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections , claim 2 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions ), so z ↦ K ( p r 1 ( z ) − p r 2 ( z ) ) z\mapsto K(\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z)) z ↦ K ( pr 1 ( z ) − pr 2 ( z )) is Borel, and bounded by M M M ; likewise y ↦ K ( x − y ) y\mapsto K(x-y) y ↦ K ( x − y ) for fixed x x x . Let μ ∈ P 2 ( R d ) \mu\in\mathcal{P}_{2}(\mathbb{R}^{d}) μ ∈ P 2 ( R d ) . In the notation of The Convolution of a Bounded Continuous Function with a Probability Measure on Euclidean Space: Continuity, Boundedness and Differentiation Under the Integral (dimension q = d q=d q = d ), K ∗ μ K*\mu K ∗ μ is the function ( H ∗ μ ) (H*\mu) ( H ∗ μ ) of that lemma with H = K H=K H = K . By The Convolution of a Bounded Continuous Function with a Probability Measure on Euclidean Space: Continuity, Boundedness and Differentiation Under the Integral §derivative with G = K G=K G = K (of class C 1 C^{1} C 1 , ∣ K ∣ ≤ M |K|\le M ∣ K ∣ ≤ M , ∣ ∂ i K ∣ ≤ M |\partial_{i}K|\le M ∣ ∂ i K ∣ ≤ M ), K ∗ μ K*\mu K ∗ μ is of class C 1 C^{1} C 1 with ∂ i ( K ∗ μ ) = ( ∂ i K ) ∗ μ \partial_{i}(K*\mu)=(\partial_{i}K)*\mu ∂ i ( K ∗ μ ) = ( ∂ i K ) ∗ μ ; by the same clause with G = ∂ i K G=\partial_{i}K G = ∂ i K (of class C 1 C^{1} C 1 since K K K is of class C 2 C^{2} C 2 , clause 2 of C^k Maps on a Euclidean Open Set , with ∣ ∂ i K ∣ ≤ M |\partial_{i}K|\le M ∣ ∂ i K ∣ ≤ M , ∣ ∂ j ∂ i K ∣ ≤ M |\partial_{j}\partial_{i}K|\le M ∣ ∂ j ∂ i K ∣ ≤ M ), ( ∂ i K ) ∗ μ (\partial_{i}K)*\mu ( ∂ i K ) ∗ μ is of class C 1 C^{1} C 1 with ∂ j ( ( ∂ i K ) ∗ μ ) = ( ∂ j ∂ i K ) ∗ μ \partial_{j}((\partial_{i}K)*\mu)=(\partial_{j}\partial_{i}K)*\mu ∂ j (( ∂ i K ) ∗ μ ) = ( ∂ j ∂ i K ) ∗ μ . Hence K ∗ μ K*\mu K ∗ μ is of class C 2 C^{2} C 2 on R d \mathbb{R}^{d} R d with the displayed formulas for ∂ i ( K ∗ μ ) \partial_{i}(K*\mu) ∂ i ( K ∗ μ ) and ∂ j ∂ i ( K ∗ μ ) \partial_{j}\partial_{i}(K*\mu) ∂ j ∂ i ( K ∗ μ ) , and the three bounds by M M M follow from The Convolution of a Bounded Continuous Function with a Probability Measure on Euclidean Space: Continuity, Boundedness and Differentiation Under the Integral §continuous . Write Γ μ = D ( K ∗ μ ) \Gamma_{\mu}=D(K*\mu) Γ μ = D ( K ∗ μ ) , a Borel map with ∣ ( Γ μ ) i ∣ ≤ M |(\Gamma_{\mu})_{i}|\le M ∣ ( Γ μ ) i ∣ ≤ M , and Γ μ ( x ) = ∫ D K ( x − y ) μ ( d y ) \Gamma_{\mu}(x)=\int DK(x-y)\mu(dy) Γ μ ( x ) = ∫ DK ( x − y ) μ ( d y ) coordinatewise.
The two expressions for K K \mathcal{K}_{K} K K . By (FB) with p = r = d p=r=d p = r = d , α = β = μ \alpha=\beta=\mu α = β = μ and the bounded Borel F ( z ) = K ( p r 1 ( z ) − p r 2 ( z ) ) F(z)=K(\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z)) F ( z ) = K ( pr 1 ( z ) − pr 2 ( z )) , for which F ( ι ( x , y ) ) = K ( x − y ) F(\iota(x,y))=K(x-y) F ( ι ( x , y )) = K ( x − y ) (Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections ),
∫ R d + d F d ( μ ⊠ μ ) = ∫ ( ∫ K ( x − y ) μ ( d y ) ) μ ( d x ) = ∫ K ∗ μ d μ , \int_{\mathbb{R}^{d+d}}F\,d(\mu\boxtimes\mu)=\int\Bigl(\int K(x-y)\,\mu(dy)\Bigr)\mu(dx)=\int K*\mu\,d\mu , ∫ R d + d F d ( μ ⊠ μ ) = ∫ ( ∫ K ( x − y ) μ ( d y ) ) μ ( d x ) = ∫ K ∗ μ d μ ,
which is the identity in the statement. More generally, for α , β ∈ P ( R d ) \alpha,\beta\in\mathcal{P}(\mathbb{R}^{d}) α , β ∈ P ( R d ) write K ( α , β ) = ∫ ( ∫ K ( x − y ) β ( d y ) ) α ( d x ) \mathcal{K}(\alpha,\beta)=\int(\int K(x-y)\beta(dy))\alpha(dx) K ( α , β ) = ∫ ( ∫ K ( x − y ) β ( d y )) α ( d x ) ; by (FB) it also equals ∫ ( ∫ K ( x − y ) α ( d x ) ) β ( d y ) \int(\int K(x-y)\alpha(dx))\beta(dy) ∫ ( ∫ K ( x − y ) α ( d x )) β ( d y ) , and K K ( μ ) = K ( μ , μ ) \mathcal{K}_{K}(\mu)=\mathcal{K}(\mu,\mu) K K ( μ ) = K ( μ , μ ) .
Translation invariance. Let a ∈ R d a\in\mathbb{R}^{d} a ∈ R d and μ a = ( τ a ) # μ \mu_{a}=(\tau_{a})_{\#}\mu μ a = ( τ a ) # μ , an element of P 2 ( R d ) \mathcal{P}_{2}(\mathbb{R}^{d}) P 2 ( R d ) by Gradients of Functions with Bounded Derivatives Belong to the Tangent Space; Constants Are Tangent; the Score Identity; the Score Has Mean Zero; Translation of Tangent Fields §translation . For every x x x , change of variables gives ( K ∗ μ a ) ( x ) = ∫ K ( x − y − a ) μ ( d y ) = ( K ∗ μ ) ( x − a ) (K*\mu_{a})(x)=\int K(x-y-a)\mu(dy)=(K*\mu)(x-a) ( K ∗ μ a ) ( x ) = ∫ K ( x − y − a ) μ ( d y ) = ( K ∗ μ ) ( x − a ) , and then K K ( μ a ) = ∫ ( K ∗ μ ) ( x − a ) μ a ( d x ) = ∫ ( K ∗ μ ) ( x + a − a ) μ ( d x ) = K K ( μ ) \mathcal{K}_{K}(\mu_{a})=\int(K*\mu)(x-a)\,\mu_{a}(dx)=\int(K*\mu)(x+a-a)\,\mu(dx)=\mathcal{K}_{K}(\mu) K K ( μ a ) = ∫ ( K ∗ μ ) ( x − a ) μ a ( d x ) = ∫ ( K ∗ μ ) ( x + a − a ) μ ( d x ) = K K ( μ ) , the function K ∗ μ K*\mu K ∗ μ being bounded and continuous, hence Borel and integrable.
Reduction to the law of a pair. Let X , H ∈ L 2 ( Ω ; R d ) X,H\in L^{2}(\Omega;\mathbb{R}^{d}) X , H ∈ L 2 ( Ω ; R d ) with representatives, μ = L ( X ) \mu=\mathcal{L}(X) μ = L ( X ) , and let π = L ( ( X , H ) ) ∈ P ( R d + d ) \pi=\mathcal{L}((X,H))\in\mathcal{P}(\mathbb{R}^{d+d}) π = L (( X , H )) ∈ P ( R d + d ) , the law of the pairing, a coupling of μ \mu μ and L ( H ) \mathcal{L}(H) L ( H ) by Basic Properties of Random Vectors: Coordinates, Borel Images and Arithmetic, Change of Variables, Almost Sure Equality and Pairs §pair , so that ( p r 1 ) # π = μ (\mathrm{pr}_{1})_{\#}\pi=\mu ( pr 1 ) # π = μ . Let θ : R d + d → R d \theta:\mathbb{R}^{d+d}\to\mathbb{R}^{d} θ : R d + d → R d , θ ( z ) = p r 1 ( z ) + p r 2 ( z ) \theta(z)=\mathrm{pr}_{1}(z)+\mathrm{pr}_{2}(z) θ ( z ) = pr 1 ( z ) + pr 2 ( z ) , a Borel map; since X + H = θ ∘ ( X , H ) X+H=\theta\circ(X,H) X + H = θ ∘ ( X , H ) pointwise, Basic Properties of Random Vectors: Coordinates, Borel Images and Arithmetic, Change of Variables, Almost Sure Equality and Pairs §composition gives L ( X + H ) = θ # π \mathcal{L}(X+H)=\theta_{\#}\pi L ( X + H ) = θ # π . For a bounded Borel G : R d → R G:\mathbb{R}^{d}\to\mathbb{R} G : R d → R and Borel S : R d + d → R d S:\mathbb{R}^{d+d}\to\mathbb{R}^{d} S : R d + d → R d , change of variables gives ∫ G d ( S # π ) = ∫ G ∘ S d π \int G\,d(S_{\#}\pi)=\int G\circ S\,d\pi ∫ G d ( S # π ) = ∫ G ∘ S d π . Applying this twice (inside, for fixed x x x , to y ↦ K ( x − y ) y\mapsto K(x-y) y ↦ K ( x − y ) , then outside, to the bounded Borel function x ↦ ∫ K ( x − S ( z ′ ) ) π ( d z ′ ) x\mapsto\int K(x-S(z'))\pi(dz') x ↦ ∫ K ( x − S ( z ′ )) π ( d z ′ ) , Borel by (FB) for the bounded Borel function ( x , z ′ ) ↦ K ( x − S ( z ′ ) ) (x,z')\mapsto K(x-S(z')) ( x , z ′ ) ↦ K ( x − S ( z ′ )) on R d + ( d + d ) \mathbb{R}^{d+(d+d)} R d + ( d + d ) ),
K K ( S # π ) = ∫ ( ∫ K ( S ( z ) − S ( z ′ ) ) π ( d z ′ ) ) π ( d z ) for S ∈ { p r 1 , θ } . (R) \mathcal{K}_{K}(S_{\#}\pi)=\int\Bigl(\int K\bigl(S(z)-S(z')\bigr)\,\pi(dz')\Bigr)\pi(dz)\qquad\text{for }S\in\{\mathrm{pr}_{1},\theta\}.\tag{R} K K ( S # π ) = ∫ ( ∫ K ( S ( z ) − S ( z ′ ) ) π ( d z ′ ) ) π ( d z ) for S ∈ { pr 1 , θ } . ( R )
Thus, with Φ \Phi Φ the lift of K K \mathcal{K}_{K} K K , Φ ( X ) = K K ( μ ) \Phi(X)=\mathcal{K}_{K}(\mu) Φ ( X ) = K K ( μ ) is (R) with S = p r 1 S=\mathrm{pr}_{1} S = pr 1 and Φ ( X + H ) = K K ( θ # π ) \Phi(X+H)=\mathcal{K}_{K}(\theta_{\#}\pi) Φ ( X + H ) = K K ( θ # π ) is (R) with S = θ S=\theta S = θ .
(a): differentiability. Keep X , H , π X,H,\pi X , H , π as above and write, for z , z ′ ∈ R d + d z,z'\in\mathbb{R}^{d+d} z , z ′ ∈ R d + d , u = p r 1 ( z ) − p r 1 ( z ′ ) u=\mathrm{pr}_{1}(z)-\mathrm{pr}_{1}(z') u = pr 1 ( z ) − pr 1 ( z ′ ) and w = p r 2 ( z ) − p r 2 ( z ′ ) w=\mathrm{pr}_{2}(z)-\mathrm{pr}_{2}(z') w = pr 2 ( z ) − pr 2 ( z ′ ) , so θ ( z ) − θ ( z ′ ) = u + w \theta(z)-\theta(z')=u+w θ ( z ) − θ ( z ′ ) = u + w . Consider the Borel functions on R ( d + d ) + ( d + d ) \mathbb{R}^{(d+d)+(d+d)} R ( d + d ) + ( d + d ) given at ι ( z , z ′ ) \iota(z,z') ι ( z , z ′ ) by F 1 = K ( u + w ) F_{1}=K(u+w) F 1 = K ( u + w ) , F 2 = K ( u ) F_{2}=K(u) F 2 = K ( u ) and F 3 = D K ( u ) ⋅ w F_{3}=DK(u)\cdot w F 3 = DK ( u ) ⋅ w (Borel by the preamble and claims 2 and 3 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions ). F 1 , F 2 F_{1},F_{2} F 1 , F 2 are bounded by M M M , and ∣ F 3 ∣ ≤ d M ∥ w ∥ ≤ d M ( ∥ p r 2 ( z ) ∥ + ∥ p r 2 ( z ′ ) ∥ ) |F_{3}|\le\sqrt{d}\,M\lVert w\rVert\le\sqrt{d}\,M(\lVert\mathrm{pr}_{2}(z)\rVert+\lVert\mathrm{pr}_{2}(z')\rVert) ∣ F 3 ∣ ≤ d M ∥ w ∥ ≤ d M (∥ pr 2 ( z )∥ + ∥ pr 2 ( z ′ )∥) (claims 5 and 6 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n ), where z ↦ ∥ p r 2 ( z ) ∥ z\mapsto\lVert\mathrm{pr}_{2}(z)\rVert z ↦ ∥ pr 2 ( z )∥ is Borel and π \pi π -integrable with ∫ ∥ p r 2 ( z ) ∥ π ( d z ) = E [ ∥ H ∥ ] ≤ ∥ H ∥ L 2 \int\lVert\mathrm{pr}_{2}(z)\rVert\pi(dz)=\mathbb{E}[\lVert H\rVert]\le\lVert H\rVert_{L^{2}} ∫ ∥ pr 2 ( z )∥ π ( d z ) = E [∥ H ∥] ≤ ∥ H ∥ L 2 (change of variables, (H)); so (FB) applies to F 1 , F 2 , F 3 F_{1},F_{2},F_{3} F 1 , F 2 , F 3 and to F 1 − F 2 − F 3 F_{1}-F_{2}-F_{3} F 1 − F 2 − F 3 with α = β = π \alpha=\beta=\pi α = β = π . By (Tay), ∣ F 1 − F 2 − F 3 ∣ ≤ 1 2 d M ∥ w ∥ 2 ≤ d M ( ∥ p r 2 ( z ) ∥ 2 + ∥ p r 2 ( z ′ ) ∥ 2 ) |F_{1}-F_{2}-F_{3}|\le\tfrac12dM\lVert w\rVert^{2}\le dM(\lVert\mathrm{pr}_{2}(z)\rVert^{2}+\lVert\mathrm{pr}_{2}(z')\rVert^{2}) ∣ F 1 − F 2 − F 3 ∣ ≤ 2 1 d M ∥ w ∥ 2 ≤ d M (∥ pr 2 ( z ) ∥ 2 + ∥ pr 2 ( z ′ ) ∥ 2 ) , and ∫ ∥ p r 2 ( z ) ∥ 2 π ( d z ) = E [ ∥ H ∥ 2 ] = ∥ H ∥ L 2 2 \int\lVert\mathrm{pr}_{2}(z)\rVert^{2}\pi(dz)=\mathbb{E}[\lVert H\rVert^{2}]=\lVert H\rVert_{L^{2}}^{2} ∫ ∥ pr 2 ( z ) ∥ 2 π ( d z ) = E [∥ H ∥ 2 ] = ∥ H ∥ L 2 2 ; hence, by (FB), claim 2 of Linearity and Monotonicity of the Lebesgue Integral and Tonelli for the dominating function,
∣ ∫ ( F 1 − F 2 − F 3 ) d ( π ⊠ π ) ∣ ≤ 2 d M ∥ H ∥ L 2 2 . \Bigl|\int(F_{1}-F_{2}-F_{3})\,d(\pi\boxtimes\pi)\Bigr|\le2dM\lVert H\rVert_{L^{2}}^{2}. ∫ ( F 1 − F 2 − F 3 ) d ( π ⊠ π ) ≤ 2 d M ∥ H ∥ L 2 2 .
By (R) and (FB), ∫ F 1 d ( π ⊠ π ) = Φ ( X + H ) \int F_{1}d(\pi\boxtimes\pi)=\Phi(X+H) ∫ F 1 d ( π ⊠ π ) = Φ ( X + H ) and ∫ F 2 d ( π ⊠ π ) = Φ ( X ) \int F_{2}\,d(\pi\boxtimes\pi)=\Phi(X) ∫ F 2 d ( π ⊠ π ) = Φ ( X ) . For F 3 F_{3} F 3 , write F 3 = F 3 ′ − F 3 ′ ′ F_{3}=F_{3}'-F_{3}'' F 3 = F 3 ′ − F 3 ′′ with F 3 ′ = D K ( u ) ⋅ p r 2 ( z ) F_{3}'=DK(u)\cdot\mathrm{pr}_{2}(z) F 3 ′ = DK ( u ) ⋅ pr 2 ( z ) and F 3 ′ ′ = D K ( u ) ⋅ p r 2 ( z ′ ) F_{3}''=DK(u)\cdot\mathrm{pr}_{2}(z') F 3 ′′ = DK ( u ) ⋅ pr 2 ( z ′ ) , both dominated as F 3 F_{3} F 3 is. Integrating F 3 ′ F_{3}' F 3 ′ first in z ′ z' z ′ : for fixed z z z , ∫ D K ( p r 1 ( z ) − p r 1 ( z ′ ) ) ⋅ p r 2 ( z ) π ( d z ′ ) = ( ∫ D K ( p r 1 ( z ) − y ) μ ( d y ) ) ⋅ p r 2 ( z ) = Γ μ ( p r 1 ( z ) ) ⋅ p r 2 ( z ) \int DK(\mathrm{pr}_{1}(z)-\mathrm{pr}_{1}(z'))\cdot\mathrm{pr}_{2}(z)\,\pi(dz')=\Bigl(\int DK(\mathrm{pr}_{1}(z)-y)\,\mu(dy)\Bigr)\cdot\mathrm{pr}_{2}(z)=\Gamma_{\mu}(\mathrm{pr}_{1}(z))\cdot\mathrm{pr}_{2}(z) ∫ DK ( pr 1 ( z ) − pr 1 ( z ′ )) ⋅ pr 2 ( z ) π ( d z ′ ) = ( ∫ DK ( pr 1 ( z ) − y ) μ ( d y ) ) ⋅ pr 2 ( z ) = Γ μ ( pr 1 ( z )) ⋅ pr 2 ( z ) (coordinatewise, claim 2 of Linearity and Monotonicity of the Lebesgue Integral and change of variables with ( p r 1 ) # π = μ (\mathrm{pr}_{1})_{\#}\pi=\mu ( pr 1 ) # π = μ ), and then ∫ Γ μ ( p r 1 ( z ) ) ⋅ p r 2 ( z ) π ( d z ) = E [ Γ μ ( X ) ⋅ H ] = ⟨ Γ μ ∘ X , H ⟩ L 2 \int\Gamma_{\mu}(\mathrm{pr}_{1}(z))\cdot\mathrm{pr}_{2}(z)\,\pi(dz)=\mathbb{E}[\Gamma_{\mu}(X)\cdot H]=\langle\Gamma_{\mu}\circ X,H\rangle_{L^{2}} ∫ Γ μ ( pr 1 ( z )) ⋅ pr 2 ( z ) π ( d z ) = E [ Γ μ ( X ) ⋅ H ] = ⟨ Γ μ ∘ X , H ⟩ L 2 (Basic Properties of Random Vectors: Coordinates, Borel Images and Arithmetic, Change of Variables, Almost Sure Equality and Pairs §expectation for the pairing ( X , H ) (X,H) ( X , H ) , The Space of Square-Integrable Random Vectors §inner-product ). Integrating F 3 ′ ′ F_{3}'' F 3 ′′ first in z z z (the other order in (FB)): for fixed z ′ z' z ′ , ∫ D K ( p r 1 ( z ) − p r 1 ( z ′ ) ) π ( d z ) = ∫ D K ( x − p r 1 ( z ′ ) ) μ ( d x ) = − ∫ D K ( p r 1 ( z ′ ) − x ) μ ( d x ) = − Γ μ ( p r 1 ( z ′ ) ) \int DK(\mathrm{pr}_{1}(z)-\mathrm{pr}_{1}(z'))\,\pi(dz)=\int DK(x-\mathrm{pr}_{1}(z'))\,\mu(dx)=-\int DK(\mathrm{pr}_{1}(z')-x)\,\mu(dx)=-\Gamma_{\mu}(\mathrm{pr}_{1}(z')) ∫ DK ( pr 1 ( z ) − pr 1 ( z ′ )) π ( d z ) = ∫ DK ( x − pr 1 ( z ′ )) μ ( d x ) = − ∫ DK ( pr 1 ( z ′ ) − x ) μ ( d x ) = − Γ μ ( pr 1 ( z ′ )) by the oddness of D K DK DK , so ∫ F 3 ′ ′ d ( π ⊠ π ) = − ∫ Γ μ ( p r 1 ( z ′ ) ) ⋅ p r 2 ( z ′ ) π ( d z ′ ) = − ⟨ Γ μ ∘ X , H ⟩ L 2 \int F_{3}''\,d(\pi\boxtimes\pi)=-\int\Gamma_{\mu}(\mathrm{pr}_{1}(z'))\cdot\mathrm{pr}_{2}(z')\,\pi(dz')=-\langle\Gamma_{\mu}\circ X,H\rangle_{L^{2}} ∫ F 3 ′′ d ( π ⊠ π ) = − ∫ Γ μ ( pr 1 ( z ′ )) ⋅ pr 2 ( z ′ ) π ( d z ′ ) = − ⟨ Γ μ ∘ X , H ⟩ L 2 . Therefore ∫ F 3 d ( π ⊠ π ) = 2 ⟨ Γ μ ∘ X , H ⟩ L 2 = ⟨ 2 Γ μ ∘ X , H ⟩ L 2 \int F_{3}\,d(\pi\boxtimes\pi)=2\langle\Gamma_{\mu}\circ X,H\rangle_{L^{2}}=\langle2\Gamma_{\mu}\circ X,H\rangle_{L^{2}} ∫ F 3 d ( π ⊠ π ) = 2 ⟨ Γ μ ∘ X , H ⟩ L 2 = ⟨ 2 Γ μ ∘ X , H ⟩ L 2 , and altogether
∣ Φ ( X + H ) − Φ ( X ) − ⟨ 2 Γ μ ∘ X , H ⟩ L 2 ∣ ≤ 2 d M ∥ H ∥ L 2 2 ( H ∈ L 2 ( Ω ; R d ) ) . \bigl|\Phi(X+H)-\Phi(X)-\langle2\,\Gamma_{\mu}\circ X,H\rangle_{L^{2}}\bigr|\le2dM\lVert H\rVert_{L^{2}}^{2}\qquad(H\in L^{2}(\Omega;\mathbb{R}^{d})). Φ ( X + H ) − Φ ( X ) − ⟨ 2 Γ μ ∘ X , H ⟩ L 2 ≤ 2 d M ∥ H ∥ L 2 2 ( H ∈ L 2 ( Ω ; R d )) .
Given ε > 0 \varepsilon>0 ε > 0 , put δ = ε / ( 2 d M + 1 ) \delta=\varepsilon/(2dM+1) δ = ε / ( 2 d M + 1 ) ; for ∥ H ∥ L 2 < δ \lVert H\rVert_{L^{2}}<\delta ∥ H ∥ L 2 < δ the right side is at most ε ∥ H ∥ L 2 \varepsilon\lVert H\rVert_{L^{2}} ε ∥ H ∥ L 2 (claim 5 of Elementary Arithmetic in an Ordered Field ). Hence Φ \Phi Φ is differentiable at X X X with gradient D Φ ( X ) = 2 Γ μ ∘ X D\Phi(X)=2\,\Gamma_{\mu}\circ X D Φ ( X ) = 2 Γ μ ∘ X (Fréchet Differentiability and the Gradient on an Open Subset of a Real Inner Product Space §differentiable , Fréchet Differentiability and the Gradient on an Open Subset of a Real Inner Product Space §gradient ), for every X X X with L ( X ) = μ \mathcal{L}(X)=\mu L ( X ) = μ .
(a): continuity of the gradient map. Let X , Y ∈ L 2 ( Ω ; R d ) X,Y\in L^{2}(\Omega;\mathbb{R}^{d}) X , Y ∈ L 2 ( Ω ; R d ) with laws μ , ν \mu,\nu μ , ν . By The Norm Metric of a Real Inner Product Space: Triangle Inequalities, Limits and Continuity §triangle and the linearity of composition (Composition of a Square-Integrable Vector Field with a Random Vector, and the Lifted Score: Isometry, Norm, Weak Identity and Second-Moment Identity §composition ), ∥ D Φ ( X ) − D Φ ( Y ) ∥ L 2 ≤ 2 ∥ Γ μ ∘ X − Γ μ ∘ Y ∥ L 2 + 2 ∥ ( Γ μ − Γ ν ) ∘ Y ∥ L 2 \lVert D\Phi(X)-D\Phi(Y)\rVert_{L^{2}}\le2\lVert\Gamma_{\mu}\circ X-\Gamma_{\mu}\circ Y\rVert_{L^{2}}+2\lVert(\Gamma_{\mu}-\Gamma_{\nu})\circ Y\rVert_{L^{2}} ∥ D Φ ( X ) − D Φ ( Y ) ∥ L 2 ≤ 2 ∥ Γ μ ∘ X − Γ μ ∘ Y ∥ L 2 + 2 ∥( Γ μ − Γ ν ) ∘ Y ∥ L 2 . Each ∂ i ( K ∗ μ ) \partial_{i}(K*\mu) ∂ i ( K ∗ μ ) is of class C 1 C^{1} C 1 with partials bounded by M M M , so (Lip) gives ∣ ( Γ μ ) i ( x ) − ( Γ μ ) i ( y ) ∣ ≤ d M ∥ x − y ∥ |(\Gamma_{\mu})_{i}(x)-(\Gamma_{\mu})_{i}(y)|\le\sqrt{d}\,M\lVert x-y\rVert ∣ ( Γ μ ) i ( x ) − ( Γ μ ) i ( y ) ∣ ≤ d M ∥ x − y ∥ and hence ∥ Γ μ ( x ) − Γ μ ( y ) ∥ ≤ d M ∥ x − y ∥ \lVert\Gamma_{\mu}(x)-\Gamma_{\mu}(y)\rVert\le dM\lVert x-y\rVert ∥ Γ μ ( x ) − Γ μ ( y )∥ ≤ d M ∥ x − y ∥ ; thus ∥ Γ μ ∘ X − Γ μ ∘ Y ∥ L 2 2 = E [ ∥ Γ μ ( X ) − Γ μ ( Y ) ∥ 2 ] ≤ d 2 M 2 ∥ X − Y ∥ L 2 2 \lVert\Gamma_{\mu}\circ X-\Gamma_{\mu}\circ Y\rVert_{L^{2}}^{2}=\mathbb{E}[\lVert\Gamma_{\mu}(X)-\Gamma_{\mu}(Y)\rVert^{2}]\le d^{2}M^{2}\lVert X-Y\rVert_{L^{2}}^{2} ∥ Γ μ ∘ X − Γ μ ∘ Y ∥ L 2 2 = E [∥ Γ μ ( X ) − Γ μ ( Y ) ∥ 2 ] ≤ d 2 M 2 ∥ X − Y ∥ L 2 2 . For every y ∈ R d y\in\mathbb{R}^{d} y ∈ R d and i i i , change of variables and Basic Properties of Random Vectors: Coordinates, Borel Images and Arithmetic, Change of Variables, Almost Sure Equality and Pairs §expectation give ( Γ μ ) i ( y ) − ( Γ ν ) i ( y ) = E [ ∂ i K ( y − X ) − ∂ i K ( y − Y ) ] (\Gamma_{\mu})_{i}(y)-(\Gamma_{\nu})_{i}(y)=\mathbb{E}[\partial_{i}K(y-X)-\partial_{i}K(y-Y)] ( Γ μ ) i ( y ) − ( Γ ν ) i ( y ) = E [ ∂ i K ( y − X ) − ∂ i K ( y − Y )] , whose absolute value is at most d M E [ ∥ X − Y ∥ ] ≤ d M ∥ X − Y ∥ L 2 \sqrt{d}\,M\,\mathbb{E}[\lVert X-Y\rVert]\le\sqrt{d}\,M\lVert X-Y\rVert_{L^{2}} d M E [∥ X − Y ∥] ≤ d M ∥ X − Y ∥ L 2 by (Lip) for ∂ i K \partial_{i}K ∂ i K , claim 2 of Linearity and Monotonicity of the Lebesgue Integral and (H); so ∥ Γ μ ( y ) − Γ ν ( y ) ∥ ≤ d M ∥ X − Y ∥ L 2 \lVert\Gamma_{\mu}(y)-\Gamma_{\nu}(y)\rVert\le dM\lVert X-Y\rVert_{L^{2}} ∥ Γ μ ( y ) − Γ ν ( y )∥ ≤ d M ∥ X − Y ∥ L 2 for every y y y , and ∥ ( Γ μ − Γ ν ) ∘ Y ∥ L 2 ≤ d M ∥ X − Y ∥ L 2 \lVert(\Gamma_{\mu}-\Gamma_{\nu})\circ Y\rVert_{L^{2}}\le dM\lVert X-Y\rVert_{L^{2}} ∥( Γ μ − Γ ν ) ∘ Y ∥ L 2 ≤ d M ∥ X − Y ∥ L 2 . Altogether ∥ D Φ ( X ) − D Φ ( Y ) ∥ L 2 ≤ 4 d M ∥ X − Y ∥ L 2 \lVert D\Phi(X)-D\Phi(Y)\rVert_{L^{2}}\le4dM\lVert X-Y\rVert_{L^{2}} ∥ D Φ ( X ) − D Φ ( Y ) ∥ L 2 ≤ 4 d M ∥ X − Y ∥ L 2 , so the gradient map is Lipschitz, hence continuous (A Lipschitz Map is Uniformly Continuous ), and Φ ∈ C 1 ( L 2 ( Ω ; R d ) ) \Phi\in C^{1}(L^{2}(\Omega;\mathbb{R}^{d})) Φ ∈ C 1 ( L 2 ( Ω ; R d )) .
(b). For μ ∈ P 2 ( R d ) \mu\in\mathcal{P}_{2}(\mathbb{R}^{d}) μ ∈ P 2 ( R d ) the class of 2 Γ μ = 2 D ( K ∗ μ ) 2\Gamma_{\mu}=2D(K*\mu) 2 Γ μ = 2 D ( K ∗ μ ) lies in T μ T_{\mu} T μ : K ∗ μ K*\mu K ∗ μ is of class C 1 C^{1} C 1 with ∣ ∂ i ( K ∗ μ ) ∣ ≤ M |\partial_{i}(K*\mu)|\le M ∣ ∂ i ( K ∗ μ ) ∣ ≤ M , so D ( K ∗ μ ) ∈ T μ D(K*\mu)\in T_{\mu} D ( K ∗ μ ) ∈ T μ by Gradients of Functions with Bounded Derivatives Belong to the Tangent Space; Constants Are Tangent; the Score Identity; the Score Has Mean Zero; Translation of Tangent Fields §tangent , and T μ T_{\mu} T μ is a linear subspace (Basic Properties of the Tangent Space: Closed Subspace, the Identity Map Belongs to It, Second-Moment Limits, and Representation of Bounded Functionals on Gradients §closed ). By the differentiability paragraph, D Φ ( X ) = ( 2 Γ μ ) ∘ X D\Phi(X)=(2\Gamma_{\mu})\circ X D Φ ( X ) = ( 2 Γ μ ) ∘ X for every X X X with L ( X ) = μ \mathcal{L}(X)=\mu L ( X ) = μ (linearity of composition, Composition of a Square-Integrable Vector Field with a Random Vector, and the Lifted Score: Isometry, Norm, Weak Identity and Second-Moment Identity §composition ). So (b) holds, and ∇ K K ( μ ) \nabla\mathcal{K}_{K}(\mu) ∇ K K ( μ ) is the class of x ↦ 2 D ( K ∗ μ ) ( x ) x\mapsto2D(K*\mu)(x) x ↦ 2 D ( K ∗ μ ) ( x ) (Test Functions on the Wasserstein Space: the Intrinsic Gradient and the Translation Hessian §gradient ).
(c) and the translation Hessian. For X X X with law μ \mu μ , ϕ X ( a ) = Φ ( X + c a ) = K K ( ( τ a ) # μ ) = K K ( μ ) \phi_{X}(a)=\Phi(X+c_{a})=\mathcal{K}_{K}((\tau_{a})_{\#}\mu)=\mathcal{K}_{K}(\mu) ϕ X ( a ) = Φ ( X + c a ) = K K (( τ a ) # μ ) = K K ( μ ) by The Wasserstein Space and Its Lift to Square-Integrable Random Vectors: Standing Notation §constants and translation invariance; a constant function is smooth (claim 2 of Constants, Coordinate Functions, Sums and Products of C k C^k C k Functions on a Euclidean Open Set ) and its partial derivatives, and theirs, vanish identically (the difference quotients of Partial Derivative on a Euclidean Open Set are 0 0 0 ). Hence Φ \Phi Φ is twice continuously differentiable along translations at every X X X , and H K K ( μ ) = D 2 ϕ X ( 0 R d ) = 0 d H_{\mathcal{K}_{K}}(\mu)=D^{2}\phi_{X}(0_{\mathbb{R}^{d}})=0_{d} H K K ( μ ) = D 2 ϕ X ( 0 R d ) = 0 d (Hessian Matrix of a C^2 Function , Test Functions on the Wasserstein Space: the Intrinsic Gradient and the Translation Hessian §hessian ).
(d) and the Lipschitz constant. Let X , Y ∈ L 2 ( Ω ; R d ) X,Y\in L^{2}(\Omega;\mathbb{R}^{d}) X , Y ∈ L 2 ( Ω ; R d ) with laws μ , ν \mu,\nu μ , ν and let π = L ( ( X , Y ) ) \pi=\mathcal{L}((X,Y)) π = L (( X , Y )) , a coupling of μ \mu μ and ν \nu ν (Basic Properties of Random Vectors: Coordinates, Borel Images and Arithmetic, Change of Variables, Almost Sure Equality and Pairs §pair ). By (R) with S = p r 1 S=\mathrm{pr}_{1} S = pr 1 and with S = p r 2 S=\mathrm{pr}_{2} S = pr 2 (the same argument, ( p r 2 ) # π = ν (\mathrm{pr}_{2})_{\#}\pi=\nu ( pr 2 ) # π = ν ), and (FB),
Φ ( X ) − Φ ( Y ) = ∫ ( ∫ [ K ( p r 1 ( z ) − p r 1 ( z ′ ) ) − K ( p r 2 ( z ) − p r 2 ( z ′ ) ) ] π ( d z ′ ) ) π ( d z ) . \Phi(X)-\Phi(Y)=\int\Bigl(\int\bigl[K(\mathrm{pr}_{1}(z)-\mathrm{pr}_{1}(z'))-K(\mathrm{pr}_{2}(z)-\mathrm{pr}_{2}(z'))\bigr]\pi(dz')\Bigr)\pi(dz). Φ ( X ) − Φ ( Y ) = ∫ ( ∫ [ K ( pr 1 ( z ) − pr 1 ( z ′ )) − K ( pr 2 ( z ) − pr 2 ( z ′ )) ] π ( d z ′ ) ) π ( d z ) .
By (Lip) and claims 5 and 6 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n , the integrand is bounded in absolute value by d M ( ∥ p r 1 ( z ) − p r 2 ( z ) ∥ + ∥ p r 1 ( z ′ ) − p r 2 ( z ′ ) ∥ ) \sqrt{d}\,M\bigl(\lVert\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z)\rVert+\lVert\mathrm{pr}_{1}(z')-\mathrm{pr}_{2}(z')\rVert\bigr) d M ( ∥ pr 1 ( z ) − pr 2 ( z )∥ + ∥ pr 1 ( z ′ ) − pr 2 ( z ′ )∥ ) , and ∫ ∥ p r 1 ( z ) − p r 2 ( z ) ∥ π ( d z ) = E [ ∥ X − Y ∥ ] ≤ ∥ X − Y ∥ L 2 \int\lVert\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z)\rVert\pi(dz)=\mathbb{E}[\lVert X-Y\rVert]\le\lVert X-Y\rVert_{L^{2}} ∫ ∥ pr 1 ( z ) − pr 2 ( z )∥ π ( d z ) = E [∥ X − Y ∥] ≤ ∥ X − Y ∥ L 2 (change of variables, (H)). Integrating twice with claim 2 of Linearity and Monotonicity of the Lebesgue Integral , ∣ Φ ( X ) − Φ ( Y ) ∣ ≤ 2 d M ∥ X − Y ∥ L 2 |\Phi(X)-\Phi(Y)|\le2\sqrt{d}\,M\lVert X-Y\rVert_{L^{2}} ∣Φ ( X ) − Φ ( Y ) ∣ ≤ 2 d M ∥ X − Y ∥ L 2 . Thus Φ \Phi Φ is Lipschitz with constant 2 d M 2\sqrt{d}\,M 2 d M , so K K \mathcal{K}_{K} K K is Lipschitz with constant 2 d M 2\sqrt{d}\,M 2 d M by Basic Properties of the Lift: Law Invariance, the Correspondence on a Rich Space, and Transfer of Boundedness, Lipschitz Constants and Continuity §lipschitz-converse and continuous by A Lipschitz Map is Uniformly Continuous . Hence K K \mathcal{K}_{K} K K is a test function with the stated gradient and translation Hessian.