Each result cited is universally quantified over the data in its own statement. Throughout, λ d \lambda_{d} λ d is Lebesgue measure on R d \mathbb{R}^{d} R d (Euclidean Space and Lebesgue Measure: Standing Notation §measure ), e i e_{i} e i is the i i i th standard basis vector of R d \mathbb{R}^{d} R d , the dot product a ⋅ x = ∑ i = 1 d a i x i a\cdot x=\sum_{i=1}^{d}a_{i}x_{i} a ⋅ x = ∑ i = 1 d a i x i is that of Difference, Dot Product, and Orthogonality in R n \mathbb{R}^n R n , so that a ⋅ e i = a i a\cdot e_{i}=a_{i} a ⋅ e i = a i by claim 7 of Properties of Finite Sums , the set R d \mathbb{R}^{d} R d is open by claim 1 of Euclidean Space is Open in Itself, and C k C^k C k Maps are Continuous , and N \mathbb{N} N is read in R \mathbb{R} R as in The Real Numbers: Standing Notation and Background §numbers . A function of class C 1 C^{1} C 1 on R d \mathbb{R}^{d} R d and its partial derivatives are continuous on R d \mathbb{R}^{d} R d by clause 1 of C^k Maps on a Euclidean Open Set and claim 3 of Euclidean Space is Open in Itself, and C k C^k C k Maps are Continuous , hence Borel by claim 3 of The Borel Sigma-Algebra of a Euclidean Space as a Product, and Measurability of Projections, Sequentially Continuous Maps, and Open and Closed Sets ; a smooth function is of class C k C^{k} C k for every k k k (Smooth Map on a Euclidean Open Set ). Two elementary facts about partial derivatives are used repeatedly and are proved here from Partial Derivative on a Euclidean Open Set .
(F1) Uniqueness. If L L L and L ′ L' L ′ both have the property of that definition at a a a for f f f and i i i , then for every ε > 0 \varepsilon>0 ε > 0 , applying the definition with ε / 2 \varepsilon/2 ε /2 (claim 8 of Elementary Order Arithmetic in an Ordered Field ) to both values and taking an h ≠ 0 h\ne0 h = 0 smaller than both radii, the same quotient lies within ε / 2 \varepsilon/2 ε /2 of L L L and of L ′ L' L ′ , so ∣ L − L ′ ∣ < ε |L-L'|<\varepsilon ∣ L − L ′ ∣ < ε by claim 5 of Properties of the Absolute Value in an Ordered Field ; hence ∣ L − L ′ ∣ ≤ ε |L-L'|\le\varepsilon ∣ L − L ′ ∣ ≤ ε for every ε > 0 \varepsilon>0 ε > 0 and L = L ′ L=L' L = L ′ by Comparison of Real Numbers with Arbitrary Positive Slack §vanishing .
(F2) Locality. If g : R d → R g:\mathbb{R}^{d}\to\mathbb{R} g : R d → R is constant on the set { x ′ : ∥ x ′ − x ∥ < r } \{x':\lVert x'-x\rVert<r\} { x ′ : ∥ x ′ − x ∥ < r } for some r > 0 r>0 r > 0 , then ∂ i g ( x ) \partial_{i}g(x) ∂ i g ( x ) exists and equals 0 0 0 : for 0 < ∣ h ∣ < r 0<|h|<r 0 < ∣ h ∣ < r one has ∥ ( x + h e i ) − x ∥ = ∣ h ∣ < r \lVert(x+he_{i})-x\rVert=|h|<r ∥( x + h e i ) − x ∥ = ∣ h ∣ < r by claim 5 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n , so the difference quotient in the definition is 0 0 0 , and L = 0 L=0 L = 0 satisfies the definition with δ = r \delta=r δ = r for every ε \varepsilon ε .
Proof of claim 1.
Growth. Let x ∈ R d x\in\mathbb{R}^{d} x ∈ R d . The segment from 0 R d 0_{\mathbb{R}^{d}} 0 R d to x x x lies in R d \mathbb{R}^{d} R d , so part (i) of Multivariate Taylor Expansion with Uniform Second-Order Remainder , applied with W = R d W=\mathbb{R}^{d} W = R d , the points 0 R d 0_{\mathbb{R}^{d}} 0 R d and x x x , and M 1 = M M_{1}=M M 1 = M , gives ∣ f ( x ) − f ( 0 R d ) ∣ ≤ d M ∥ x ∥ |f(x)-f(0_{\mathbb{R}^{d}})|\le\sqrt{d}\,M\lVert x\rVert ∣ f ( x ) − f ( 0 R d ) ∣ ≤ d M ∥ x ∥ , the distance between the two points being ∥ x ∥ \lVert x\rVert ∥ x ∥ by claim 2 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n . Hence ∣ f ( x ) ∣ ≤ ∣ f ( 0 R d ) ∣ + ∣ f ( x ) − f ( 0 R d ) ∣ ≤ ∣ f ( 0 R d ) ∣ + d M ∥ x ∥ |f(x)|\le|f(0_{\mathbb{R}^{d}})|+|f(x)-f(0_{\mathbb{R}^{d}})|\le|f(0_{\mathbb{R}^{d}})|+\sqrt{d}\,M\lVert x\rVert ∣ f ( x ) ∣ ≤ ∣ f ( 0 R d ) ∣ + ∣ f ( x ) − f ( 0 R d ) ∣ ≤ ∣ f ( 0 R d ) ∣ + d M ∥ x ∥ by claim 5 of Properties of the Absolute Value in an Ordered Field .
Square integrability. Each ∂ i f \partial_{i}f ∂ i f is Borel, so D f Df D f is Borel by the componentwise criterion of Probability Measures on Euclidean Space and Random Vectors: Standing Notation §borel-maps . For every x x x , claim 1 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n gives ∥ D f ( x ) ∥ 2 = ∑ i = 1 d ( ∂ i f ( x ) ) 2 \lVert Df(x)\rVert^{2}=\sum_{i=1}^{d}(\partial_{i}f(x))^{2} ∥ D f ( x ) ∥ 2 = ∑ i = 1 d ( ∂ i f ( x ) ) 2 , and ( ∂ i f ( x ) ) 2 = ∣ ∂ i f ( x ) ∣ 2 ≤ M 2 (\partial_{i}f(x))^{2}=|\partial_{i}f(x)|^{2}\le M^{2} ( ∂ i f ( x ) ) 2 = ∣ ∂ i f ( x ) ∣ 2 ≤ M 2 by claim 4 of Properties of the Absolute Value in an Ordered Field and claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field , so ∥ D f ( x ) ∥ 2 ≤ d M 2 \lVert Df(x)\rVert^{2}\le dM^{2} ∥ D f ( x ) ∥ 2 ≤ d M 2 by claim 1 of Comparison and Absolute Value Bounds for Finite Sums of Real Numbers . The function ∥ D f ∥ 2 \lVert Df\rVert^{2} ∥ D f ∥ 2 is Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions , and the monotonicity in claim 1 of Linearity and Monotonicity of the Lebesgue Integral together with Simple Function and Its Integral (the integral of the constant d M 2 dM^{2} d M 2 against the probability measure μ \mu μ is d M 2 dM^{2} d M 2 ) gives ∫ ∥ D f ∥ 2 d μ ≤ d M 2 < ∞ \int\lVert Df\rVert^{2}\,d\mu\le dM^{2}<\infty ∫ ∥ D f ∥ 2 d μ ≤ d M 2 < ∞ . Thus the class of D f Df D f lies in L 2 ( μ ; R d ) L^{2}(\mu;\mathbb{R}^{d}) L 2 ( μ ; R d ) (Square-Integrable Vector Fields Against a Probability Measure on Euclidean Space, and Test Functions: Standing Notation §l2mu ).
Mollification. By Existence of Mollifier Kernels of Every Radius there is a mollifier kernel ρ \rho ρ of radius 1 1 1 on R d \mathbb{R}^{d} R d ; for ε > 0 \varepsilon>0 ε > 0 let ρ ε ( y ) = ( ε − 1 ) d ρ ( ε − 1 y ) \rho_{\varepsilon}(y)=(\varepsilon^{-1})^{d}\rho(\varepsilon^{-1}y) ρ ε ( y ) = ( ε − 1 ) d ρ ( ε − 1 y ) , a mollifier kernel of radius ε \varepsilon ε by Rescaling a Mollifier Kernel : it is smooth, nonnegative, vanishes at every y y y with ∥ y ∥ > ε \lVert y\rVert>\varepsilon ∥ y ∥ > ε , and ∫ ρ ε d λ d = 1 \int\rho_{\varepsilon}\,d\lambda_{d}=1 ∫ ρ ε d λ d = 1 . Let g : R d → R g:\mathbb{R}^{d}\to\mathbb{R} g : R d → R be continuous. Since B ˉ ( x , ε ) ⊆ R d \bar B(x,\varepsilon)\subseteq\mathbb{R}^{d} B ˉ ( x , ε ) ⊆ R d for every x x x , the convolution g ∗ ρ ε g*\rho_{\varepsilon} g ∗ ρ ε is defined on all of R d \mathbb{R}^{d} R d (the set Ω ε \Omega^{\varepsilon} Ω ε of that definition with Ω = R d \Omega=\mathbb{R}^{d} Ω = R d ), ( g ∗ ρ ε ) ( x ) = ∫ g ( x − y ) ρ ε ( y ) λ d ( d y ) (g*\rho_{\varepsilon})(x)=\int g(x-y)\rho_{\varepsilon}(y)\,\lambda_{d}(dy) ( g ∗ ρ ε ) ( x ) = ∫ g ( x − y ) ρ ε ( y ) λ d ( d y ) , and it is smooth on R d \mathbb{R}^{d} R d by claim 2 of Convolution with a C k C^k C k Kernel is of Class C k C^k C k . If ∣ g ∣ ≤ C |g|\le C ∣ g ∣ ≤ C on R d \mathbb{R}^{d} R d , then ∣ g ( x − y ) ρ ε ( y ) ∣ ≤ C ρ ε ( y ) |g(x-y)\rho_{\varepsilon}(y)|\le C\rho_{\varepsilon}(y) ∣ g ( x − y ) ρ ε ( y ) ∣ ≤ C ρ ε ( y ) for all y y y , so ∣ ( g ∗ ρ ε ) ( x ) ∣ ≤ ∫ C ρ ε d λ d = C |(g*\rho_{\varepsilon})(x)|\le\int C\rho_{\varepsilon}\,d\lambda_{d}=C ∣ ( g ∗ ρ ε ) ( x ) ∣ ≤ ∫ C ρ ε d λ d = C by claim 2 of Linearity and Monotonicity of the Lebesgue Integral (the integrand is integrable by claim 1 of The Convolution Integrand is Continuous, Compactly Supported and Integrable ). Likewise, since ρ ε ( y ) = 0 \rho_{\varepsilon}(y)=0 ρ ε ( y ) = 0 for ∥ y ∥ > ε \lVert y\rVert>\varepsilon ∥ y ∥ > ε , the growth bound gives ∣ f ( x − y ) ρ ε ( y ) ∣ ≤ ( ∣ f ( 0 R d ) ∣ + d M ( ∥ x ∥ + ε ) ) ρ ε ( y ) |f(x-y)\rho_{\varepsilon}(y)|\le\bigl(|f(0_{\mathbb{R}^{d}})|+\sqrt{d}\,M(\lVert x\rVert+\varepsilon)\bigr)\rho_{\varepsilon}(y) ∣ f ( x − y ) ρ ε ( y ) ∣ ≤ ( ∣ f ( 0 R d ) ∣ + d M (∥ x ∥ + ε ) ) ρ ε ( y ) for all y y y (claim 6 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n : ∥ x − y ∥ ≤ ∥ x ∥ + ∥ y ∥ \lVert x-y\rVert\le\lVert x\rVert+\lVert y\rVert ∥ x − y ∥ ≤ ∥ x ∥ + ∥ y ∥ ), whence
∣ ( f ∗ ρ ε ) ( x ) ∣ ≤ ∣ f ( 0 R d ) ∣ + d M ( ∥ x ∥ + ε ) ( x ∈ R d ) . (G) |(f*\rho_{\varepsilon})(x)|\le|f(0_{\mathbb{R}^{d}})|+\sqrt{d}\,M(\lVert x\rVert+\varepsilon)\qquad(x\in\mathbb{R}^{d}).\tag{G} ∣ ( f ∗ ρ ε ) ( x ) ∣ ≤ ∣ f ( 0 R d ) ∣ + d M (∥ x ∥ + ε ) ( x ∈ R d ) . ( G )
The derivative of a mollification. Let g : R d → R g:\mathbb{R}^{d}\to\mathbb{R} g : R d → R be of class C 1 C^{1} C 1 on R d \mathbb{R}^{d} R d with ∣ ∂ i g ∣ ≤ C |\partial_{i}g|\le C ∣ ∂ i g ∣ ≤ C on R d \mathbb{R}^{d} R d for all i i i . We show
∂ i ( g ∗ ρ ε ) = ( ∂ i g ) ∗ ρ ε on R d , for every i ∈ [ d ] . (D) \partial_{i}(g*\rho_{\varepsilon})=(\partial_{i}g)*\rho_{\varepsilon}\qquad\text{on }\mathbb{R}^{d},\ \text{for every }i\in[d].\tag{D} ∂ i ( g ∗ ρ ε ) = ( ∂ i g ) ∗ ρ ε on R d , for every i ∈ [ d ] . ( D )
Fix x x x and i i i . By claim 2 of Differentiating a Convolution through the Kernel , ∂ i ( g ∗ ρ ε ) ( x ) = ∫ g ( x − y ) ∂ i ρ ε ( y ) λ d ( d y ) \partial_{i}(g*\rho_{\varepsilon})(x)=\int g(x-y)\,\partial_{i}\rho_{\varepsilon}(y)\,\lambda_{d}(dy) ∂ i ( g ∗ ρ ε ) ( x ) = ∫ g ( x − y ) ∂ i ρ ε ( y ) λ d ( d y ) , the integrand being integrable by claim 1 of The Convolution Integrand is Continuous, Compactly Supported and Integrable applied to the kernel ∂ i ρ ε \partial_{i}\rho_{\varepsilon} ∂ i ρ ε (claim 1 of Differentiating a Convolution through the Kernel ). By claim 3 of Translation and Reflection Invariance of Lebesgue Measure on R n \mathbb{R}^n R n with a = x a=x a = x (applied to the integrable function y ↦ g ( x − y ) ∂ i ρ ε ( y ) y\mapsto g(x-y)\partial_{i}\rho_{\varepsilon}(y) y ↦ g ( x − y ) ∂ i ρ ε ( y ) , evaluated at x − y x-y x − y ), this equals ∫ g ( y ) ∂ i ρ ε ( x − y ) λ d ( d y ) \int g(y)\,\partial_{i}\rho_{\varepsilon}(x-y)\,\lambda_{d}(dy) ∫ g ( y ) ∂ i ρ ε ( x − y ) λ d ( d y ) . Let k ( y ) = ρ ε ( x − y ) k(y)=\rho_{\varepsilon}(x-y) k ( y ) = ρ ε ( x − y ) . For y ∈ R d y\in\mathbb{R}^{d} y ∈ R d and j ∈ [ d ] j\in[d] j ∈ [ d ] , the function t ↦ k ( y + t e j ) = ρ ε ( ( x − y ) + t ( − e j ) ) t\mapsto k(y+te_{j})=\rho_{\varepsilon}\bigl((x-y)+t(-e_{j})\bigr) t ↦ k ( y + t e j ) = ρ ε ( ( x − y ) + t ( − e j ) ) is differentiable at 0 0 0 with derivative ∑ l ∂ l ρ ε ( x − y ) ( − e j ) l = − ∂ j ρ ε ( x − y ) \sum_{l}\partial_{l}\rho_{\varepsilon}(x-y)(-e_{j})_{l}=-\partial_{j}\rho_{\varepsilon}(x-y) ∑ l ∂ l ρ ε ( x − y ) ( − e j ) l = − ∂ j ρ ε ( x − y ) (claim 7 of Properties of Finite Sums ) by Chain Rule Along an Affine Path , since ρ ε \rho_{\varepsilon} ρ ε is differentiable at x − y x-y x − y by A Real-Valued C^1 Function is Differentiable at Every Point ; its difference quotients at 0 0 0 are exactly those of Partial Derivative on a Euclidean Open Set for k k k at y y y , so ∂ j k ( y ) = − ∂ j ρ ε ( x − y ) \partial_{j}k(y)=-\partial_{j}\rho_{\varepsilon}(x-y) ∂ j k ( y ) = − ∂ j ρ ε ( x − y ) by (F1). These partials are continuous (a composition of continuous maps, claim 3 of Semicontinuity and Continuity Under Composition with a Continuous Map , the map y ↦ x − y y\mapsto x-y y ↦ x − y being continuous as ∥ ( x − y ) − ( x − y ′ ) ∥ = ∥ y ′ − y ∥ \lVert(x-y)-(x-y')\rVert=\lVert y'-y\rVert ∥( x − y ) − ( x − y ′ )∥ = ∥ y ′ − y ∥ ), as is k k k , so k k k is of class C 1 C^{1} C 1 on R d \mathbb{R}^{d} R d . Moreover k ( y ) = 0 k(y)=0 k ( y ) = 0 whenever ∥ y ∥ > ∥ x ∥ + ε \lVert y\rVert>\lVert x\rVert+\varepsilon ∥ y ∥ > ∥ x ∥ + ε , since then ∥ x − y ∥ ≥ ∥ y ∥ − ∥ x ∥ > ε \lVert x-y\rVert\ge\lVert y\rVert-\lVert x\rVert>\varepsilon ∥ x − y ∥ ≥ ∥ y ∥ − ∥ x ∥ > ε (claims 5 and 6 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n : ∥ y ∥ ≤ ∥ y − x ∥ + ∥ x ∥ \lVert y\rVert\le\lVert y-x\rVert+\lVert x\rVert ∥ y ∥ ≤ ∥ y − x ∥ + ∥ x ∥ and ∥ y − x ∥ = ∥ x − y ∥ \lVert y-x\rVert=\lVert x-y\rVert ∥ y − x ∥ = ∥ x − y ∥ ); so k k k is compactly supported by claim 2 of Compact Support on R n \mathbb{R}^n R n Means Vanishing Outside a Bounded Set . Integration by Parts on Euclidean Space Against a Compactly Supported Function of Class C 1 C^{1} C 1 §parts , applied to g g g and k k k , gives ∫ g ∂ i k d λ d = − ∫ ( ∂ i g ) k d λ d \int g\,\partial_{i}k\,d\lambda_{d}=-\int(\partial_{i}g)\,k\,d\lambda_{d} ∫ g ∂ i k d λ d = − ∫ ( ∂ i g ) k d λ d , that is, ∫ g ( y ) ∂ i ρ ε ( x − y ) λ d ( d y ) = ∫ ∂ i g ( y ) ρ ε ( x − y ) λ d ( d y ) \int g(y)\,\partial_{i}\rho_{\varepsilon}(x-y)\,\lambda_{d}(dy)=\int\partial_{i}g(y)\,\rho_{\varepsilon}(x-y)\,\lambda_{d}(dy) ∫ g ( y ) ∂ i ρ ε ( x − y ) λ d ( d y ) = ∫ ∂ i g ( y ) ρ ε ( x − y ) λ d ( d y ) (claim 2 of Linearity and Monotonicity of the Lebesgue Integral with the factor − 1 -1 − 1 ). By claim 3 of Translation and Reflection Invariance of Lebesgue Measure on R n \mathbb{R}^n R n once more, the right side is ∫ ∂ i g ( x − y ) ρ ε ( y ) λ d ( d y ) = ( ( ∂ i g ) ∗ ρ ε ) ( x ) \int\partial_{i}g(x-y)\,\rho_{\varepsilon}(y)\,\lambda_{d}(dy)=((\partial_{i}g)*\rho_{\varepsilon})(x) ∫ ∂ i g ( x − y ) ρ ε ( y ) λ d ( d y ) = (( ∂ i g ) ∗ ρ ε ) ( x ) , the function ∂ i g \partial_{i}g ∂ i g being continuous. This proves (D). Consequently, by the bound above, ∣ ∂ i ( g ∗ ρ ε ) ∣ ≤ C |\partial_{i}(g*\rho_{\varepsilon})|\le C ∣ ∂ i ( g ∗ ρ ε ) ∣ ≤ C on R d \mathbb{R}^{d} R d .
Uniform convergence on balls. Let g g g be continuous on R d \mathbb{R}^{d} R d , let n ∈ N n\in\mathbb{N} n ∈ N and let K n = B ˉ ( 0 R d , n ) K_{n}=\bar B(0_{\mathbb{R}^{d}},n) K n = B ˉ ( 0 R d , n ) , which is compact by claim 2 of A Closed Euclidean Ball is Convex and Compact . Claim 2 of Mollification Converges Uniformly on Compact Subsets , applied with Ω = R d \Omega=\mathbb{R}^{d} Ω = R d , the kernel ρ \rho ρ of radius 1 1 1 and the compact set K n K_{n} K n , gives for every η > 0 \eta>0 η > 0 an ε 0 > 0 \varepsilon_{0}>0 ε 0 > 0 such that ∣ ( g ∗ ρ ε ) ( x ) − g ( x ) ∣ < η |(g*\rho_{\varepsilon})(x)-g(x)|<\eta ∣ ( g ∗ ρ ε ) ( x ) − g ( x ) ∣ < η for all x ∈ K n x\in K_{n} x ∈ K n and 0 < ε < ε 0 0<\varepsilon<\varepsilon_{0} 0 < ε < ε 0 .
Cutoff. Fix χ \chi χ as in Scaled Cutoffs and the Second-Moment Test Functions: Uniform Derivative Bounds and Agreement on a Ball for the dimension d d d , and for n ∈ N n\in\mathbb{N} n ∈ N let χ n \chi_{n} χ n be the function χ R \chi_{R} χ R of that lemma with R = n R=n R = n ; by Scaled Cutoffs and the Second-Moment Test Functions: Uniform Derivative Bounds and Agreement on a Ball §cutoff it is smooth and compactly supported, 0 ≤ χ n ≤ 1 0\le\chi_{n}\le1 0 ≤ χ n ≤ 1 , χ n ( x ) = 1 \chi_{n}(x)=1 χ n ( x ) = 1 for ∥ x ∥ ≤ n \lVert x\rVert\le n ∥ x ∥ ≤ n , χ n ( x ) = 0 \chi_{n}(x)=0 χ n ( x ) = 0 for ∥ x ∥ ≥ 2 n \lVert x\rVert\ge2n ∥ x ∥ ≥ 2 n , and ∣ ∂ i χ n ∣ ≤ M 1 n − 1 |\partial_{i}\chi_{n}|\le M_{1}n^{-1} ∣ ∂ i χ n ∣ ≤ M 1 n − 1 , ∣ ∂ j ∂ i χ n ∣ ≤ M 2 n − 2 |\partial_{j}\partial_{i}\chi_{n}|\le M_{2}n^{-2} ∣ ∂ j ∂ i χ n ∣ ≤ M 2 n − 2 on R d \mathbb{R}^{d} R d , with M 1 , M 2 M_{1},M_{2} M 1 , M 2 independent of n n n . By (F2): if ∥ x ∥ < n \lVert x\rVert<n ∥ x ∥ < n then χ n \chi_{n} χ n is constant on { x ′ : ∥ x ′ − x ∥ < n − ∥ x ∥ } \{x':\lVert x'-x\rVert<n-\lVert x\rVert\} { x ′ : ∥ x ′ − x ∥ < n − ∥ x ∥} (claim 6 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n ), so ∂ i χ n ( x ) = 0 \partial_{i}\chi_{n}(x)=0 ∂ i χ n ( x ) = 0 , and then, ∂ i χ n \partial_{i}\chi_{n} ∂ i χ n being constant on the same set, ∂ j ∂ i χ n ( x ) = 0 \partial_{j}\partial_{i}\chi_{n}(x)=0 ∂ j ∂ i χ n ( x ) = 0 ; if ∥ x ∥ > 2 n \lVert x\rVert>2n ∥ x ∥ > 2 n then χ n \chi_{n} χ n is constant on { x ′ : ∥ x ′ − x ∥ < ∥ x ∥ − 2 n } \{x':\lVert x'-x\rVert<\lVert x\rVert-2n\} { x ′ : ∥ x ′ − x ∥ < ∥ x ∥ − 2 n } (claim 7 of Properties of the Absolute Value in an Ordered Field applied through claim 6 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n : ∥ x ′ ∥ ≥ ∥ x ∥ − ∥ x − x ′ ∥ > 2 n \lVert x'\rVert\ge\lVert x\rVert-\lVert x-x'\rVert>2n ∥ x ′ ∥ ≥ ∥ x ∥ − ∥ x − x ′ ∥ > 2 n ), so χ n ( x ) = ∂ i χ n ( x ) = ∂ j ∂ i χ n ( x ) = 0 \chi_{n}(x)=\partial_{i}\chi_{n}(x)=\partial_{j}\partial_{i}\chi_{n}(x)=0 χ n ( x ) = ∂ i χ n ( x ) = ∂ j ∂ i χ n ( x ) = 0 .
For n ∈ N n\in\mathbb{N} n ∈ N and 0 < ε ≤ n 0<\varepsilon\le n 0 < ε ≤ n put
ψ n , ε = χ n ⋅ ( f ∗ ρ ε ) . \psi_{n,\varepsilon}=\chi_{n}\cdot(f*\rho_{\varepsilon}). ψ n , ε = χ n ⋅ ( f ∗ ρ ε ) .
It is smooth on R d \mathbb{R}^{d} R d by claim 3 of Constants, Coordinate Functions, Sums and Products of C k C^k C k Functions on a Euclidean Open Set , and it vanishes at every x x x with ∥ x ∥ > 2 n \lVert x\rVert>2n ∥ x ∥ > 2 n , as χ n \chi_{n} χ n does, so it is compactly supported by claim 2 of Compact Support on R n \mathbb{R}^n R n Means Vanishing Outside a Bounded Set . Hence ψ n , ε ∈ C c ∞ ( R d ) \psi_{n,\varepsilon}\in C_{c}^{\infty}(\mathbb{R}^{d}) ψ n , ε ∈ C c ∞ ( R d ) (Test Functions on Euclidean Space, Their Gradient Maps and Laplacians §space ), and ∇ ψ n , ε ∈ G μ \nabla\psi_{n,\varepsilon}\in G_{\mu} ∇ ψ n , ε ∈ G μ (The Tangent Space of the Wasserstein Space at a Probability Measure §gradients ). By claim 1 of Constants, Coordinate Functions, Sums and Products of C k C^k C k Functions on a Euclidean Open Set and (D) applied to g = f g=f g = f ,
∂ i ψ n , ε = ∂ i χ n ⋅ ( f ∗ ρ ε ) + χ n ⋅ ( ( ∂ i f ) ∗ ρ ε ) . (P) \partial_{i}\psi_{n,\varepsilon}=\partial_{i}\chi_{n}\cdot(f*\rho_{\varepsilon})+\chi_{n}\cdot\bigl((\partial_{i}f)*\rho_{\varepsilon}\bigr).\tag{P} ∂ i ψ n , ε = ∂ i χ n ⋅ ( f ∗ ρ ε ) + χ n ⋅ ( ( ∂ i f ) ∗ ρ ε ) . ( P )
The gradient estimate. Let n ∈ N n\in\mathbb{N} n ∈ N , η > 0 \eta>0 η > 0 , and let ε 0 \varepsilon_{0} ε 0 be provided by the uniform convergence paragraph for the d d d continuous functions ∂ 1 f , … , ∂ d f \partial_{1}f,\dots,\partial_{d}f ∂ 1 f , … , ∂ d f on K n K_{n} K n simultaneously (the least of the d d d radii, obtained by applying claim 9 of Elementary Order Arithmetic in an Ordered Field repeatedly). Let 0 < ε < ε 0 0<\varepsilon<\varepsilon_{0} 0 < ε < ε 0 with ε ≤ n \varepsilon\le n ε ≤ n . For ∥ x ∥ < n \lVert x\rVert<n ∥ x ∥ < n , (P) and the cutoff paragraph give ∂ i ψ n , ε ( x ) = ( ( ∂ i f ) ∗ ρ ε ) ( x ) \partial_{i}\psi_{n,\varepsilon}(x)=((\partial_{i}f)*\rho_{\varepsilon})(x) ∂ i ψ n , ε ( x ) = (( ∂ i f ) ∗ ρ ε ) ( x ) , so ∣ ∂ i ψ n , ε ( x ) − ∂ i f ( x ) ∣ < η |\partial_{i}\psi_{n,\varepsilon}(x)-\partial_{i}f(x)|<\eta ∣ ∂ i ψ n , ε ( x ) − ∂ i f ( x ) ∣ < η and, by claim 1 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n , claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field and claim 1 of Comparison and Absolute Value Bounds for Finite Sums of Real Numbers , ∥ ∇ ψ n , ε ( x ) − D f ( x ) ∥ 2 ≤ d η 2 \lVert\nabla\psi_{n,\varepsilon}(x)-Df(x)\rVert^{2}\le d\eta^{2} ∥ ∇ ψ n , ε ( x ) − D f ( x ) ∥ 2 ≤ d η 2 . For ∥ x ∥ > 2 n \lVert x\rVert>2n ∥ x ∥ > 2 n , ∂ i ψ n , ε ( x ) = 0 \partial_{i}\psi_{n,\varepsilon}(x)=0 ∂ i ψ n , ε ( x ) = 0 by (P). For n ≤ ∥ x ∥ ≤ 2 n n\le\lVert x\rVert\le2n n ≤ ∥ x ∥ ≤ 2 n , (P), (G), the cutoff bounds, the bound ∣ ( ∂ i f ) ∗ ρ ε ∣ ≤ M |(\partial_{i}f)*\rho_{\varepsilon}|\le M ∣ ( ∂ i f ) ∗ ρ ε ∣ ≤ M and ε ≤ n \varepsilon\le n ε ≤ n , 1 ≤ n 1\le n 1 ≤ n give
∣ ∂ i ψ n , ε ( x ) ∣ ≤ M 1 n − 1 ( ∣ f ( 0 R d ) ∣ + 3 n d M ) + M ≤ M 1 ∣ f ( 0 R d ) ∣ + 3 d M M 1 + M = : C 1 , |\partial_{i}\psi_{n,\varepsilon}(x)|\le M_{1}n^{-1}\bigl(|f(0_{\mathbb{R}^{d}})|+3n\sqrt{d}\,M\bigr)+M\le M_{1}|f(0_{\mathbb{R}^{d}})|+3\sqrt{d}\,MM_{1}+M=:C_{1}, ∣ ∂ i ψ n , ε ( x ) ∣ ≤ M 1 n − 1 ( ∣ f ( 0 R d ) ∣ + 3 n d M ) + M ≤ M 1 ∣ f ( 0 R d ) ∣ + 3 d M M 1 + M =: C 1 ,
using claims 4 and 5 of Properties of the Absolute Value in an Ordered Field and claim 5 of Elementary Arithmetic in an Ordered Field . Hence for ∥ x ∥ ≥ n \lVert x\rVert\ge n ∥ x ∥ ≥ n , ∥ ∇ ψ n , ε ( x ) ∥ ≤ d C 1 \lVert\nabla\psi_{n,\varepsilon}(x)\rVert\le\sqrt{d}\,C_{1} ∥ ∇ ψ n , ε ( x )∥ ≤ d C 1 and ∥ D f ( x ) ∥ ≤ d M \lVert Df(x)\rVert\le\sqrt{d}\,M ∥ D f ( x )∥ ≤ d M (as above), so ∥ ∇ ψ n , ε ( x ) − D f ( x ) ∥ 2 ≤ C 2 2 \lVert\nabla\psi_{n,\varepsilon}(x)-Df(x)\rVert^{2}\le C_{2}^{2} ∥ ∇ ψ n , ε ( x ) − D f ( x ) ∥ 2 ≤ C 2 2 with C 2 = d ( C 1 + M ) C_{2}=\sqrt{d}(C_{1}+M) C 2 = d ( C 1 + M ) , by claims 5 and 6 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n and claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field . Writing A n = { x : ∥ x ∥ ≥ n } A_{n}=\{x:\lVert x\rVert\ge n\} A n = { x : ∥ x ∥ ≥ n } , a Borel set by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions and The Borel Sigma-Algebra of a Euclidean Space as a Product, and Measurability of Projections, Sequentially Continuous Maps, and Open and Closed Sets , the pointwise bound ∥ ∇ ψ n , ε − D f ∥ 2 ≤ d η 2 + C 2 2 1 A n \lVert\nabla\psi_{n,\varepsilon}-Df\rVert^{2}\le d\eta^{2}+C_{2}^{2}\mathbf{1}_{A_{n}} ∥ ∇ ψ n , ε − D f ∥ 2 ≤ d η 2 + C 2 2 1 A n holds on R d \mathbb{R}^{d} R d , and claim 1 of Linearity and Monotonicity of the Lebesgue Integral with Simple Function and Its Integral gives
∥ ∇ ψ n , ε − D f ∥ μ 2 ≤ d η 2 + C 2 2 μ ( A n ) ≤ d η 2 + C 2 2 n − 2 M 2 ( μ ) , (E1) \lVert\nabla\psi_{n,\varepsilon}-Df\rVert_{\mu}^{2}\le d\eta^{2}+C_{2}^{2}\,\mu(A_{n})\le d\eta^{2}+C_{2}^{2}\,n^{-2}M_{2}(\mu),\tag{E1} ∥ ∇ ψ n , ε − D f ∥ μ 2 ≤ d η 2 + C 2 2 μ ( A n ) ≤ d η 2 + C 2 2 n − 2 M 2 ( μ ) , ( E1 )
where μ ( A n ) = μ ( { ∥ x ∥ 2 ≥ n 2 } ) ≤ n − 2 ∫ ∥ x ∥ 2 μ ( d x ) = n − 2 M 2 ( μ ) \mu(A_{n})=\mu(\{\lVert x\rVert^{2}\ge n^{2}\})\le n^{-2}\int\lVert x\rVert^{2}\mu(dx)=n^{-2}M_{2}(\mu) μ ( A n ) = μ ({∥ x ∥ 2 ≥ n 2 }) ≤ n − 2 ∫ ∥ x ∥ 2 μ ( d x ) = n − 2 M 2 ( μ ) by claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field , The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §markov and The Second Moment of a Probability Measure on Euclidean Space and the Probability Measures with Finite Second Moment §moment .
Conclusion of claim 1. Let k ∈ N k\in\mathbb{N} k ∈ N . By The Archimedean Property of the Real Numbers choose n k ∈ N n_{k}\in\mathbb{N} n k ∈ N with C 2 2 M 2 ( μ ) n k − 2 < ( 2 k 2 ) − 1 C_{2}^{2}M_{2}(\mu)n_{k}^{-2}<(2k^{2})^{-1} C 2 2 M 2 ( μ ) n k − 2 < ( 2 k 2 ) − 1 , then η k > 0 \eta_{k}>0 η k > 0 with d η k 2 < ( 2 k 2 ) − 1 d\eta_{k}^{2}<(2k^{2})^{-1} d η k 2 < ( 2 k 2 ) − 1 (claim 8 of Elementary Order Arithmetic in an Ordered Field and Existence and Uniqueness of the Nonnegative Square Root ), then ε k \varepsilon_{k} ε k as in the gradient estimate for ( n k , η k ) (n_{k},\eta_{k}) ( n k , η k ) , and put ψ k = ψ n k , ε k \psi_{k}=\psi_{n_{k},\varepsilon_{k}} ψ k = ψ n k , ε k . Then ∥ ∇ ψ k − D f ∥ μ 2 < k − 2 \lVert\nabla\psi_{k}-Df\rVert_{\mu}^{2}<k^{-2} ∥ ∇ ψ k − D f ∥ μ 2 < k − 2 , i.e. d μ ( ∇ ψ k , D f ) < k − 1 d_{\mu}(\nabla\psi_{k},Df)<k^{-1} d μ ( ∇ ψ k , D f ) < k − 1 (claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field ). Hence ∇ ψ k → D f \nabla\psi_{k}\to Df ∇ ψ k → D f in the metric space ( L 2 ( μ ; R d ) , d μ ) (L^{2}(\mu;\mathbb{R}^{d}),d_{\mu}) ( L 2 ( μ ; R d ) , d μ ) (Convergent Sequence in a Metric Space , with The Archimedean Property of the Real Numbers ), and since every ∇ ψ k ∈ G μ \nabla\psi_{k}\in G_{\mu} ∇ ψ k ∈ G μ , the class D f Df D f lies in the closure G μ ‾ = T μ \overline{G_{\mu}}=T_{\mu} G μ = T μ by Sequential Characterization of the Closure in a Metric Space and The Tangent Space of the Wasserstein Space at a Probability Measure §tangent .
Proof of claim 2. Let a ∈ R d a\in\mathbb{R}^{d} a ∈ R d and f a ( x ) = a ⋅ x = ∑ i = 1 d a i x i f_{a}(x)=a\cdot x=\sum_{i=1}^{d}a_{i}x_{i} f a ( x ) = a ⋅ x = ∑ i = 1 d a i x i . For x ∈ R d x\in\mathbb{R}^{d} x ∈ R d , i ∈ [ d ] i\in[d] i ∈ [ d ] and h ≠ 0 h\ne0 h = 0 , the difference quotient of Partial Derivative on a Euclidean Open Set is ( f a ( x + h e i ) − f a ( x ) ) / h = ( a ⋅ ( h e i ) ) / h = a i h / h = a i (f_{a}(x+he_{i})-f_{a}(x))/h=(a\cdot(he_{i}))/h=a_{i}h/h=a_{i} ( f a ( x + h e i ) − f a ( x )) / h = ( a ⋅ ( h e i )) / h = a i h / h = a i (Difference, Dot Product, and Orthogonality in R n \mathbb{R}^n R n ; claims 2 and 3 of Properties of Finite Sums for the additivity and homogeneity of the dot product in its second argument, and a ⋅ e i = a i a\cdot e_{i}=a_{i} a ⋅ e i = a i ), so ∂ i f a ( x ) = a i \partial_{i}f_{a}(x)=a_{i} ∂ i f a ( x ) = a i by that definition and (F1); the partials are constant, hence continuous, and f a f_{a} f a is continuous by Continuity of Sums and Products of Real-Valued Functions on a Metric Space , so f a f_{a} f a is of class C 1 C^{1} C 1 on R d \mathbb{R}^{d} R d ; and ∣ ∂ i f a ∣ = ∣ a i ∣ ≤ ∥ a ∥ |\partial_{i}f_{a}|=|a_{i}|\le\lVert a\rVert ∣ ∂ i f a ∣ = ∣ a i ∣ ≤ ∥ a ∥ by claim 4 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n . By claim 1 with M = ∥ a ∥ M=\lVert a\rVert M = ∥ a ∥ , the class of D f a Df_{a} D f a , the constant map with value a a a , lies in T μ T_{\mu} T μ ; that class is the class written a a a .
Proof of claim 3. Now f f f is of class C 2 C^{2} C 2 with ∣ ∂ j ∂ i f ∣ ≤ M |\partial_{j}\partial_{i}f|\le M ∣ ∂ j ∂ i f ∣ ≤ M . Each ∂ j ∂ i f \partial_{j}\partial_{i}f ∂ j ∂ i f is continuous (clause 2 of C^k Maps on a Euclidean Open Set ), so Δ f = ∑ i ∂ i ∂ i f \Delta f=\sum_{i}\partial_{i}\partial_{i}f Δ f = ∑ i ∂ i ∂ i f (The Laplacian of a Twice Continuously Differentiable Function §laplacian ) is continuous by Continuity of Sums and Products of Real-Valued Functions on a Metric Space , hence Borel, and ∣ Δ f ∣ ≤ d M |\Delta f|\le dM ∣Δ f ∣ ≤ d M by claim 2 of Comparison and Absolute Value Bounds for Finite Sums of Real Numbers ; being bounded and Borel it is integrable with respect to μ \mu μ (Integrable Function and the Lebesgue Integral and claim 1 of Linearity and Monotonicity of the Lebesgue Integral ).
Retain the sequence ψ k = ψ n k , ε k \psi_{k}=\psi_{n_{k},\varepsilon_{k}} ψ k = ψ n k , ε k of the proof of claim 1, choosing ε k \varepsilon_{k} ε k now also smaller than the radius provided by the uniform convergence paragraph for the d d d continuous functions ∂ i ∂ i f \partial_{i}\partial_{i}f ∂ i ∂ i f on K n k K_{n_{k}} K n k with the same η k \eta_{k} η k ; (E1) still holds. By Finite Fisher Information, the Score and the Fisher Information of a Probability Measure §score , ⟨ ξ μ , ∇ ψ k ⟩ μ = − ∫ Δ ψ k d μ \langle\xi_{\mu},\nabla\psi_{k}\rangle_{\mu}=-\int\Delta\psi_{k}\,d\mu ⟨ ξ μ , ∇ ψ k ⟩ μ = − ∫ Δ ψ k d μ for every k k k . The left side converges to ⟨ ξ μ , D f ⟩ μ \langle\xi_{\mu},Df\rangle_{\mu} ⟨ ξ μ , D f ⟩ μ by The Norm Metric of a Real Inner Product Space: Triangle Inequalities, Limits and Continuity §continuity , since ∇ ψ k → D f \nabla\psi_{k}\to Df ∇ ψ k → D f in L 2 ( μ ; R d ) L^{2}(\mu;\mathbb{R}^{d}) L 2 ( μ ; R d ) . It remains to show ∫ Δ ψ k d μ → ∫ Δ f d μ \int\Delta\psi_{k}\,d\mu\to\int\Delta f\,d\mu ∫ Δ ψ k d μ → ∫ Δ f d μ ; then ⟨ ξ μ , D f ⟩ μ = − ∫ Δ f d μ \langle\xi_{\mu},Df\rangle_{\mu}=-\int\Delta f\,d\mu ⟨ ξ μ , D f ⟩ μ = − ∫ Δ f d μ by claim 3 of Arithmetic of Limits of Real Sequences (negation of limits) and claim 1 of Uniqueness of Limits and Boundedness of Convergent Real Sequences (uniqueness of limits).
Fix n , ε n,\varepsilon n , ε as in the construction and write ψ = ψ n , ε \psi=\psi_{n,\varepsilon} ψ = ψ n , ε , F = f ∗ ρ ε F=f*\rho_{\varepsilon} F = f ∗ ρ ε . By (D) applied to g = f g=f g = f and then to g = ∂ i f g=\partial_{i}f g = ∂ i f (of class C 1 C^{1} C 1 with partials bounded by M M M ), ∂ i F = ( ∂ i f ) ∗ ρ ε \partial_{i}F=(\partial_{i}f)*\rho_{\varepsilon} ∂ i F = ( ∂ i f ) ∗ ρ ε and ∂ j ∂ i F = ( ∂ j ∂ i f ) ∗ ρ ε \partial_{j}\partial_{i}F=(\partial_{j}\partial_{i}f)*\rho_{\varepsilon} ∂ j ∂ i F = ( ∂ j ∂ i f ) ∗ ρ ε , so ∣ ∂ i F ∣ ≤ M |\partial_{i}F|\le M ∣ ∂ i F ∣ ≤ M and ∣ ∂ j ∂ i F ∣ ≤ M |\partial_{j}\partial_{i}F|\le M ∣ ∂ j ∂ i F ∣ ≤ M . Differentiating (P) with claim 1 of Constants, Coordinate Functions, Sums and Products of C k C^k C k Functions on a Euclidean Open Set ,
∂ i ∂ i ψ = ∂ i ∂ i χ n ⋅ F + 2 ∂ i χ n ⋅ ∂ i F + χ n ⋅ ∂ i ∂ i F , \partial_{i}\partial_{i}\psi=\partial_{i}\partial_{i}\chi_{n}\cdot F+2\,\partial_{i}\chi_{n}\cdot\partial_{i}F+\chi_{n}\cdot\partial_{i}\partial_{i}F, ∂ i ∂ i ψ = ∂ i ∂ i χ n ⋅ F + 2 ∂ i χ n ⋅ ∂ i F + χ n ⋅ ∂ i ∂ i F ,
and summing over i i i (Properties of Finite Sums , claim 2), Δ ψ = Δ χ n ⋅ F + 2 ∑ i ∂ i χ n ∂ i F + χ n Δ F \Delta\psi=\Delta\chi_{n}\cdot F+2\sum_{i}\partial_{i}\chi_{n}\,\partial_{i}F+\chi_{n}\,\Delta F Δ ψ = Δ χ n ⋅ F + 2 ∑ i ∂ i χ n ∂ i F + χ n Δ F . For ∥ x ∥ < n \lVert x\rVert<n ∥ x ∥ < n the cutoff paragraph gives Δ ψ ( x ) = Δ F ( x ) = ∑ i ( ( ∂ i ∂ i f ) ∗ ρ ε ) ( x ) \Delta\psi(x)=\Delta F(x)=\sum_{i}((\partial_{i}\partial_{i}f)*\rho_{\varepsilon})(x) Δ ψ ( x ) = Δ F ( x ) = ∑ i (( ∂ i ∂ i f ) ∗ ρ ε ) ( x ) , so ∣ Δ ψ ( x ) − Δ f ( x ) ∣ ≤ ∑ i ∣ ( ( ∂ i ∂ i f ) ∗ ρ ε ) ( x ) − ∂ i ∂ i f ( x ) ∣ < d η |\Delta\psi(x)-\Delta f(x)|\le\sum_{i}|((\partial_{i}\partial_{i}f)*\rho_{\varepsilon})(x)-\partial_{i}\partial_{i}f(x)|<d\eta ∣Δ ψ ( x ) − Δ f ( x ) ∣ ≤ ∑ i ∣ (( ∂ i ∂ i f ) ∗ ρ ε ) ( x ) − ∂ i ∂ i f ( x ) ∣ < d η (claim 2 of Comparison and Absolute Value Bounds for Finite Sums of Real Numbers ) when ε \varepsilon ε is below the radius chosen for these d d d functions. For ∥ x ∥ > 2 n \lVert x\rVert>2n ∥ x ∥ > 2 n , Δ ψ ( x ) = 0 \Delta\psi(x)=0 Δ ψ ( x ) = 0 , so ∣ Δ ψ ( x ) − Δ f ( x ) ∣ ≤ d M |\Delta\psi(x)-\Delta f(x)|\le dM ∣Δ ψ ( x ) − Δ f ( x ) ∣ ≤ d M . For n ≤ ∥ x ∥ ≤ 2 n n\le\lVert x\rVert\le2n n ≤ ∥ x ∥ ≤ 2 n , by (G), the cutoff bounds and 1 ≤ n 1\le n 1 ≤ n , ε ≤ n \varepsilon\le n ε ≤ n ,
∣ Δ ψ ( x ) ∣ ≤ d M 2 n − 2 ( ∣ f ( 0 R d ) ∣ + 3 n d M ) + 2 d M 1 n − 1 M + d M ≤ d M 2 ( ∣ f ( 0 R d ) ∣ + 3 d M ) + 2 d M 1 M + d M = : C 3 ′ , |\Delta\psi(x)|\le dM_{2}n^{-2}\bigl(|f(0_{\mathbb{R}^{d}})|+3n\sqrt{d}\,M\bigr)+2dM_{1}n^{-1}M+dM\le dM_{2}\bigl(|f(0_{\mathbb{R}^{d}})|+3\sqrt{d}\,M\bigr)+2dM_{1}M+dM=:C_{3}', ∣Δ ψ ( x ) ∣ ≤ d M 2 n − 2 ( ∣ f ( 0 R d ) ∣ + 3 n d M ) + 2 d M 1 n − 1 M + d M ≤ d M 2 ( ∣ f ( 0 R d ) ∣ + 3 d M ) + 2 d M 1 M + d M =: C 3 ′ ,
so ∣ Δ ψ ( x ) − Δ f ( x ) ∣ ≤ C 3 ′ + d M = : C 3 |\Delta\psi(x)-\Delta f(x)|\le C_{3}'+dM=:C_{3} ∣Δ ψ ( x ) − Δ f ( x ) ∣ ≤ C 3 ′ + d M =: C 3 , and C 3 ≥ d M C_{3}\ge dM C 3 ≥ d M . Hence ∣ Δ ψ − Δ f ∣ ≤ d η + C 3 1 A n |\Delta\psi-\Delta f|\le d\eta+C_{3}\mathbf{1}_{A_{n}} ∣Δ ψ − Δ f ∣ ≤ d η + C 3 1 A n on R d \mathbb{R}^{d} R d , and by claim 2 of Linearity and Monotonicity of the Lebesgue Integral (both functions are bounded Borel, hence integrable), claim 1 of that theorem, Simple Function and Its Integral and the Markov bound above,
∣ ∫ Δ ψ d μ − ∫ Δ f d μ ∣ ≤ ∫ ∣ Δ ψ − Δ f ∣ d μ ≤ d η + C 3 n − 2 M 2 ( μ ) . (E2) \Bigl|\int\Delta\psi\,d\mu-\int\Delta f\,d\mu\Bigr|\le\int|\Delta\psi-\Delta f|\,d\mu\le d\eta+C_{3}\,n^{-2}M_{2}(\mu).\tag{E2} ∫ Δ ψ d μ − ∫ Δ f d μ ≤ ∫ ∣Δ ψ − Δ f ∣ d μ ≤ d η + C 3 n − 2 M 2 ( μ ) . ( E2 )
With the choices made for ψ k \psi_{k} ψ k (and, if necessary, n k n_{k} n k chosen also so large that C 3 M 2 ( μ ) n k − 2 < ( 2 k ) − 1 C_{3}M_{2}(\mu)n_{k}^{-2}<(2k)^{-1} C 3 M 2 ( μ ) n k − 2 < ( 2 k ) − 1 and η k \eta_{k} η k also with d η k < ( 2 k ) − 1 d\eta_{k}<(2k)^{-1} d η k < ( 2 k ) − 1 , which does not affect (E1)), (E2) gives ∣ ∫ Δ ψ k d μ − ∫ Δ f d μ ∣ < k − 1 |\int\Delta\psi_{k}\,d\mu-\int\Delta f\,d\mu|<k^{-1} ∣ ∫ Δ ψ k d μ − ∫ Δ f d μ ∣ < k − 1 for every k k k , so ∫ Δ ψ k d μ → ∫ Δ f d μ \int\Delta\psi_{k}\,d\mu\to\int\Delta f\,d\mu ∫ Δ ψ k d μ → ∫ Δ f d μ by Limit of a Sequence of Real Numbers and The Archimedean Property of the Real Numbers . This proves claim 3.
Proof of claim 4. Let a ∈ R d a\in\mathbb{R}^{d} a ∈ R d and f a ( x ) = a ⋅ x f_{a}(x)=a\cdot x f a ( x ) = a ⋅ x as in the proof of claim 2. Its partials ∂ i f a = a i \partial_{i}f_{a}=a_{i} ∂ i f a = a i are constant, so for all i , j i,j i , j the difference quotients of ∂ i f a \partial_{i}f_{a} ∂ i f a vanish and ∂ j ∂ i f a = 0 \partial_{j}\partial_{i}f_{a}=0 ∂ j ∂ i f a = 0 by Partial Derivative on a Euclidean Open Set and (F1); these are continuous, so f a f_{a} f a is of class C 2 C^{2} C 2 on R d \mathbb{R}^{d} R d (clause 2 of C^k Maps on a Euclidean Open Set ) with ∣ ∂ i f a ∣ ≤ ∥ a ∥ |\partial_{i}f_{a}|\le\lVert a\rVert ∣ ∂ i f a ∣ ≤ ∥ a ∥ and ∣ ∂ j ∂ i f a ∣ = 0 ≤ ∥ a ∥ |\partial_{j}\partial_{i}f_{a}|=0\le\lVert a\rVert ∣ ∂ j ∂ i f a ∣ = 0 ≤ ∥ a ∥ , and Δ f a = 0 \Delta f_{a}=0 Δ f a = 0 . Claim 3 with M = ∥ a ∥ M=\lVert a\rVert M = ∥ a ∥ gives ⟨ ξ μ , a ⟩ μ = ⟨ ξ μ , D f a ⟩ μ = − ∫ 0 d μ = 0 \langle\xi_{\mu},a\rangle_{\mu}=\langle\xi_{\mu},Df_{a}\rangle_{\mu}=-\int0\,d\mu=0 ⟨ ξ μ , a ⟩ μ = ⟨ ξ μ , D f a ⟩ μ = − ∫ 0 d μ = 0 (Simple Function and Its Integral ; and − 0 = 0 -0=0 − 0 = 0 ).
Proof of claim 5. Let a ∈ R d a\in\mathbb{R}^{d} a ∈ R d and η ∈ T μ \eta\in T_{\mu} η ∈ T μ , and write μ a = ( τ a ) # μ \mu_{a}=(\tau_{a})_{\#}\mu μ a = ( τ a ) # μ , a probability measure on R d \mathbb{R}^{d} R d by Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pushforward , τ a \tau_{a} τ a being Borel by The Space of Square-Integrable Random Vectors is a Real Hilbert Space; Its Laws Have Finite Second Moment; Constants and Translations §constants . By Couplings on Euclidean Space: Product Coupling, Swap, Finiteness of the Cost, Push-Forward Couplings, Modifying One Marginal, Quantisation, Gluing over a Finitely Supported Measure, and the Lipschitz Bound §pushforward with S = i d S=\mathrm{id} S = id , T = τ a T=\tau_{a} T = τ a , the plan ( i d , τ a ) # μ (\mathrm{id},\tau_{a})_{\#}\mu ( id , τ a ) # μ is a coupling of μ \mu μ and μ a \mu_{a} μ a with cost ∫ ∥ x − ( x + a ) ∥ 2 μ ( d x ) = ∥ a ∥ 2 < ∞ \int\lVert x-(x+a)\rVert^{2}\mu(dx)=\lVert a\rVert^{2}<\infty ∫ ∥ x − ( x + a ) ∥ 2 μ ( d x ) = ∥ a ∥ 2 < ∞ (claim 5 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n and Simple Function and Its Integral ), so μ a ∈ P 2 ( R d ) \mu_{a}\in\mathcal{P}_{2}(\mathbb{R}^{d}) μ a ∈ P 2 ( R d ) by Couplings on Euclidean Space: Product Coupling, Swap, Finiteness of the Cost, Push-Forward Couplings, Modifying One Marginal, Quantisation, Gluing over a Finitely Supported Measure, and the Lipschitz Bound §cost-finite .
Transport of classes. For a Borel ζ : R d → R d \zeta:\mathbb{R}^{d}\to\mathbb{R}^{d} ζ : R d → R d write ζ a = ζ ∘ τ − a \zeta^{a}=\zeta\circ\tau_{-a} ζ a = ζ ∘ τ − a , the map x ↦ ζ ( x − a ) x\mapsto\zeta(x-a) x ↦ ζ ( x − a ) ; it is Borel as a composition of Borel maps (Probability Measures on Euclidean Space and Random Vectors: Standing Notation §borel-maps ), and ζ a ∘ τ a = ζ \zeta^{a}\circ\tau_{a}=\zeta ζ a ∘ τ a = ζ since τ − a ( τ a ( x ) ) = x \tau_{-a}(\tau_{a}(x))=x τ − a ( τ a ( x )) = x . By claim 2 of Image Measures, Measures with Densities, and Change of Variables (change of variables for the image measure μ a \mu_{a} μ a of μ \mu μ under τ a \tau_{a} τ a ), for every Borel ζ \zeta ζ ,
∫ ∥ ζ a ∥ 2 d μ a = ∫ ∥ ζ a ∘ τ a ∥ 2 d μ = ∫ ∥ ζ ∥ 2 d μ . (T) \int\lVert\zeta^{a}\rVert^{2}\,d\mu_{a}=\int\lVert\zeta^{a}\circ\tau_{a}\rVert^{2}\,d\mu=\int\lVert\zeta\rVert^{2}\,d\mu.\tag{T} ∫ ∥ ζ a ∥ 2 d μ a = ∫ ∥ ζ a ∘ τ a ∥ 2 d μ = ∫ ∥ ζ ∥ 2 d μ . ( T )
If ζ , ζ ′ \zeta,\zeta' ζ , ζ ′ are Borel with μ ( { ζ = ζ ′ } ) = 1 \mu(\{\zeta=\zeta'\})=1 μ ({ ζ = ζ ′ }) = 1 , then { ζ a = ζ ′ a } = τ a ( { ζ = ζ ′ } ) \{\zeta^{a}=\zeta'^{a}\}=\tau_{a}(\{\zeta=\zeta'\}) { ζ a = ζ ′ a } = τ a ({ ζ = ζ ′ }) has τ a \tau_{a} τ a -preimage { ζ = ζ ′ } \{\zeta=\zeta'\} { ζ = ζ ′ } , so μ a ( { ζ a = ζ ′ a } ) = μ ( { ζ = ζ ′ } ) = 1 \mu_{a}(\{\zeta^{a}=\zeta'^{a}\})=\mu(\{\zeta=\zeta'\})=1 μ a ({ ζ a = ζ ′ a }) = μ ({ ζ = ζ ′ }) = 1 by claim 1 of Image Measures, Measures with Densities, and Change of Variables (the set { ζ a = ζ ′ a } = { x : ∥ ζ a ( x ) − ζ ′ a ( x ) ∥ = 0 } \{\zeta^{a}=\zeta'^{a}\}=\{x:\lVert\zeta^{a}(x)-\zeta'^{a}(x)\rVert=0\} { ζ a = ζ ′ a } = { x : ∥ ζ a ( x ) − ζ ′ a ( x )∥ = 0 } is Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions ). Hence, for a representative η \eta η of the given class, η a \eta^{a} η a is Borel with ∫ ∥ η a ∥ 2 d μ a = ∫ ∥ η ∥ 2 d μ < ∞ \int\lVert\eta^{a}\rVert^{2}d\mu_{a}=\int\lVert\eta\rVert^{2}d\mu<\infty ∫ ∥ η a ∥ 2 d μ a = ∫ ∥ η ∥ 2 d μ < ∞ , its class in L 2 ( μ a ; R d ) L^{2}(\mu_{a};\mathbb{R}^{d}) L 2 ( μ a ; R d ) does not depend on the representative, and ∥ η a ∥ μ a = ∥ η ∥ μ \lVert\eta^{a}\rVert_{\mu_{a}}=\lVert\eta\rVert_{\mu} ∥ η a ∥ μ a = ∥ η ∥ μ by (T) and Square-Integrable Vector Fields Against a Probability Measure on Euclidean Space, and Test Functions: Standing Notation §l2mu . This proves all assertions of claim 5 except tangency.
Translates of test functions. Let ψ ∈ C c ∞ ( R d ) \psi\in C_{c}^{\infty}(\mathbb{R}^{d}) ψ ∈ C c ∞ ( R d ) and ψ a = ψ ∘ τ − a \psi^{a}=\psi\circ\tau_{-a} ψ a = ψ ∘ τ − a . For every function g g g on R d \mathbb{R}^{d} R d , every x x x and i i i , the difference quotients of Partial Derivative on a Euclidean Open Set for g a = g ∘ τ − a g^{a}=g\circ\tau_{-a} g a = g ∘ τ − a at x x x are those of g g g at x − a x-a x − a , since g a ( x + h e i ) = g ( x − a + h e i ) g^{a}(x+he_{i})=g(x-a+he_{i}) g a ( x + h e i ) = g ( x − a + h e i ) ; so ∂ i g a \partial_{i}g^{a} ∂ i g a exists at x x x if and only if ∂ i g \partial_{i}g ∂ i g exists at x − a x-a x − a , and then ∂ i ( g a ) = ( ∂ i g ) a \partial_{i}(g^{a})=(\partial_{i}g)^{a} ∂ i ( g a ) = ( ∂ i g ) a by (F1). Also g a g^{a} g a is continuous whenever g g g is (claim 3 of Semicontinuity and Continuity Under Composition with a Continuous Map , τ − a \tau_{-a} τ − a being continuous as ∥ τ − a ( x ) − τ − a ( x ′ ) ∥ = ∥ x − x ′ ∥ \lVert\tau_{-a}(x)-\tau_{-a}(x')\rVert=\lVert x-x'\rVert ∥ τ − a ( x ) − τ − a ( x ′ )∥ = ∥ x − x ′ ∥ ). By induction on k k k (Principle of Induction for the Natural Numbers , on the set of k k k such that for every g g g of class C k C^{k} C k the function g a g^{a} g a is of class C k C^{k} C k with ∂ i ( g a ) = ( ∂ i g ) a \partial_{i}(g^{a})=(\partial_{i}g)^{a} ∂ i ( g a ) = ( ∂ i g ) a ), using clauses 1 and 2 of C^k Maps on a Euclidean Open Set , ψ a \psi^{a} ψ a is of class C k C^{k} C k for every k k k , i.e. smooth. By claim 2 of Compact Support on R n \mathbb{R}^n R n Means Vanishing Outside a Bounded Set there is R > 0 R>0 R > 0 with ψ ( y ) = 0 \psi(y)=0 ψ ( y ) = 0 for ∥ y ∥ > R \lVert y\rVert>R ∥ y ∥ > R ; if ∥ x ∥ > R + ∥ a ∥ \lVert x\rVert>R+\lVert a\rVert ∥ x ∥ > R + ∥ a ∥ then ∥ x − a ∥ ≥ ∥ x ∥ − ∥ a ∥ > R \lVert x-a\rVert\ge\lVert x\rVert-\lVert a\rVert>R ∥ x − a ∥ ≥ ∥ x ∥ − ∥ a ∥ > R (claim 6 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n ), so ψ a ( x ) = ψ ( x − a ) = 0 \psi^{a}(x)=\psi(x-a)=0 ψ a ( x ) = ψ ( x − a ) = 0 , and ψ a \psi^{a} ψ a is compactly supported by the same claim. Thus ψ a ∈ C c ∞ ( R d ) \psi^{a}\in C_{c}^{\infty}(\mathbb{R}^{d}) ψ a ∈ C c ∞ ( R d ) and ∇ ( ψ a ) = ( ∇ ψ ) a \nabla(\psi^{a})=(\nabla\psi)^{a} ∇ ( ψ a ) = ( ∇ ψ ) a (Test Functions on Euclidean Space, Their Gradient Maps and Laplacians §gradient ), so ( ∇ ψ ) a ∈ G μ a (\nabla\psi)^{a}\in G_{\mu_{a}} ( ∇ ψ ) a ∈ G μ a .
Tangency. Since η ∈ T μ = G μ ‾ \eta\in T_{\mu}=\overline{G_{\mu}} η ∈ T μ = G μ , Sequential Characterization of the Closure in a Metric Space gives ψ k ∈ C c ∞ ( R d ) \psi_{k}\in C_{c}^{\infty}(\mathbb{R}^{d}) ψ k ∈ C c ∞ ( R d ) with ∥ ∇ ψ k − η ∥ μ → 0 \lVert\nabla\psi_{k}-\eta\rVert_{\mu}\to0 ∥ ∇ ψ k − η ∥ μ → 0 . By (T) applied to ζ = ∇ ψ k − η \zeta=\nabla\psi_{k}-\eta ζ = ∇ ψ k − η , and since ( ∇ ψ k − η ) a = ( ∇ ψ k ) a − η a (\nabla\psi_{k}-\eta)^{a}=(\nabla\psi_{k})^{a}-\eta^{a} ( ∇ ψ k − η ) a = ( ∇ ψ k ) a − η a pointwise, ∥ ( ∇ ψ k ) a − η a ∥ μ a = ∥ ∇ ψ k − η ∥ μ → 0 \lVert(\nabla\psi_{k})^{a}-\eta^{a}\rVert_{\mu_{a}}=\lVert\nabla\psi_{k}-\eta\rVert_{\mu}\to0 ∥( ∇ ψ k ) a − η a ∥ μ a = ∥ ∇ ψ k − η ∥ μ → 0 . As ( ∇ ψ k ) a = ∇ ( ψ k a ) ∈ G μ a (\nabla\psi_{k})^{a}=\nabla(\psi_{k}^{a})\in G_{\mu_{a}} ( ∇ ψ k ) a = ∇ ( ψ k a ) ∈ G μ a , the class η a \eta^{a} η a lies in G μ a ‾ = T μ a \overline{G_{\mu_{a}}}=T_{\mu_{a}} G μ a = T μ a by Sequential Characterization of the Closure in a Metric Space and The Tangent Space of the Wasserstein Space at a Probability Measure §tangent .