Each result cited below is universally quantified over the data in its own statement and is applied to the data named where it is used; The Gaussian Kernels on Euclidean Space: Scaling, Derivatives up to Order Three, the Convolution Identity, Moments and Tails , Gaussian Smoothing of a Probability Measure on Euclidean Space: Regularity, Mass, Duality, Approximation of a Bounded Function with a Modulus of Continuity, and the Pairing Identity and The Convolution of a Bounded Continuous Function with a Probability Measure on Euclidean Space: Continuity, Boundedness and Differentiation Under the Integral are read with q = d q=d q = d , and The Gaussian Smoothing Weight: Normalization, Derivatives, Exponential Tilting, Moments, and First-Order Remainder with m = d m=d m = d and η = s \eta=s η = s .
Conventions. The number s s s with 0 < s ≤ 1 2 0<s\le\tfrac12 0 < s ≤ 2 1 is fixed. Integrals ∫ f ( y ) d y \int f(y)\,dy ∫ f ( y ) d y are over R d \mathbb{R}^{d} R d with respect to λ d \lambda_{d} λ d , and ∫ Q Θ s d λ d \int_{Q}\Theta_{s}\,d\lambda_{d} ∫ Q Θ s d λ d stands for ∫ 1 Q Θ s d λ d \int\mathbf{1}_{Q}\Theta_{s}\,d\lambda_{d} ∫ 1 Q Θ s d λ d . Densities are those of The Radon-Nikodym Theorem for a Finite Measure and a Sigma-Finite Measure, and Uniqueness of Densities on the measure space ( R d , B ( R d ) , λ d ) (\mathbb{R}^{d},\mathcal{B}(\mathbb{R}^{d}),\lambda_{d}) ( R d , B ( R d ) , λ d ) . By Optimal Transport on the Flat Torus: Standing Notation §conventions and The Flat Torus: Standing Notation §periodic , C p e r C_{\mathrm{per}} C per and C p e r k C^{k}_{\mathrm{per}} C per k are the classes of Lattice-Periodic Functions and the Periodic Function Classes §classes . Let g s g_{s} g s be the Gaussian kernel of The Gaussian Kernels on Euclidean Space: Scaling, Derivatives up to Order Three, the Convolution Identity, Moments and Tails , and write α s \alpha_{s} α s for the positive constant called c s c_{s} c s there (the letters c s c_{s} c s and C s C_{s} C s are reserved for the bounds of clause 1), so that g s ( z ) = α s exp ( − ∥ z ∥ 2 / ( 2 s ) ) g_{s}(z)=\alpha_{s}\exp(-\lVert z\rVert^{2}/(2s)) g s ( z ) = α s exp ( − ∥ z ∥ 2 / ( 2 s )) . By The Gaussian Kernels on Euclidean Space: Scaling, Derivatives up to Order Three, the Convolution Identity, Moments and Tails §derivatives , g s g_{s} g s is smooth and even, with ∂ i g s ( z ) = − s − 1 z i g s ( z ) \partial_{i}g_{s}(z)=-s^{-1}z_{i}\,g_{s}(z) ∂ i g s ( z ) = − s − 1 z i g s ( z ) and ∂ j ∂ i g s ( z ) = ( s − 2 z i z j − s − 1 δ i j ) g s ( z ) \partial_{j}\partial_{i}g_{s}(z)=(s^{-2}z_{i}z_{j}-s^{-1}\delta_{ij})\,g_{s}(z) ∂ j ∂ i g s ( z ) = ( s − 2 z i z j − s − 1 δ ij ) g s ( z ) . As recorded there, g s g_{s} g s is the Gaussian smoothing weight of claim 1 of The Gaussian Smoothing Weight: Normalization, Derivatives, Exponential Tilting, Moments, and First-Order Remainder ; by that claim g s g_{s} g s is Borel, 0 < g s ( z ) 0<g_{s}(z) 0 < g s ( z ) for every z z z , and ∫ g s ( z − a ) d z = 1 \int g_{s}(z-a)\,dz=1 ∫ g s ( z − a ) d z = 1 for every a ∈ R d a\in\mathbb{R}^{d} a ∈ R d . For κ ∈ P ( R d ) \kappa\in\mathcal{P}(\mathbb{R}^{d}) κ ∈ P ( R d ) let g s ∗ κ g_{s}*\kappa g s ∗ κ be the Gaussian smoothing of κ \kappa κ at scale s s s , and for a bounded Borel ϕ : R d → R \phi:\mathbb{R}^{d}\to\mathbb{R} ϕ : R d → R let g s ∗ ϕ g_{s}*\phi g s ∗ ϕ be as defined in the same place. Let κ s ∈ P ( R d ) \kappa_{s}\in\mathcal{P}(\mathbb{R}^{d}) κ s ∈ P ( R d ) be the measure with density g s ∗ κ g_{s}*\kappa g s ∗ κ with respect to λ d \lambda_{d} λ d (claim 3 of Image Measures, Measures with Densities, and Change of Variables ), exactly as in The Heat Semigroup on the Probability Measures on the Torus ; thus S s κ = π # κ s S_{s}\kappa=\pi_{\#}\kappa_{s} S s κ = π # κ s , a member of P ( T d ) \mathcal{P}(\mathbb{T}^{d}) P ( T d ) , for κ ∈ P ( T d ) \kappa\in\mathcal{P}(\mathbb{T}^{d}) κ ∈ P ( T d ) by The Heat Semigroup on the Probability Measures on the Torus §heat .
Step 0 (Preliminaries).
(a) Measures on the torus. Let κ ∈ P ( T d ) \kappa\in\mathcal{P}(\mathbb{T}^{d}) κ ∈ P ( T d ) . As κ ( Q ) = 1 = κ ( R d ) \kappa(Q)=1=\kappa(\mathbb{R}^{d}) κ ( Q ) = 1 = κ ( R d ) (Probability Measures on the Flat Torus, the Torus Cost of a Coupling, the Torus Wasserstein Distance and Optimal Couplings §measures ), claim 3 of Basic Properties of a Measure gives κ ( R d ∖ Q ) = 0 \kappa(\mathbb{R}^{d}\setminus Q)=0 κ ( R d ∖ Q ) = 0 , and for Borel A A A , finite additivity and monotonicity (claims 1 and 2 there), applied to the disjoint sets A ∩ Q A\cap Q A ∩ Q and A ∖ Q ⊆ R d ∖ Q A\setminus Q\subseteq\mathbb{R}^{d}\setminus Q A ∖ Q ⊆ R d ∖ Q , give κ ( A ) = κ ( A ∩ Q ) \kappa(A)=\kappa(A\cap Q) κ ( A ) = κ ( A ∩ Q ) . Also, for x ∈ Q x\in Q x ∈ Q one has 0 ≤ x i < 1 0\le x_{i}<1 0 ≤ x i < 1 , so x i 2 ≤ 1 x_{i}^{2}\le1 x i 2 ≤ 1 , for each i ∈ [ d ] i\in[d] i ∈ [ d ] , whence ∥ x ∥ 2 = ∑ i = 1 d x i 2 ≤ d \lVert x\rVert^{2}=\sum_{i=1}^{d}x_{i}^{2}\le d ∥ x ∥ 2 = ∑ i = 1 d x i 2 ≤ d (claim 1 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n ) and ∥ x ∥ ≤ d \lVert x\rVert\le\sqrt{d} ∥ x ∥ ≤ d (claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field ).
(b) Iterated integrals. Let ι : R d × R d → R d + d \iota:\mathbb{R}^{d}\times\mathbb{R}^{d}\to\mathbb{R}^{d+d} ι : R d × R d → R d + d be the concatenation map of Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pairs . If F ^ : R d + d → [ 0 , ∞ ] \hat F:\mathbb{R}^{d+d}\to[0,\infty] F ^ : R d + d → [ 0 , ∞ ] is Borel, then ( x , y ) ↦ F ^ ( ι ( x , y ) ) (x,y)\mapsto\hat F(\iota(x,y)) ( x , y ) ↦ F ^ ( ι ( x , y )) is measurable for B ( R d ) ⊗ B ( R d ) \mathcal{B}(\mathbb{R}^{d})\otimes\mathcal{B}(\mathbb{R}^{d}) B ( R d ) ⊗ B ( R d ) , since ι \iota ι is measurable for that σ \sigma σ -algebra and B ( R d + d ) \mathcal{B}(\mathbb{R}^{d+d}) B ( R d + d ) by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §product and a composite of measurable maps is measurable; and F ^ ( ι ( x , y ) ) \hat F(\iota(x,y)) F ^ ( ι ( x , y )) is obtained by substituting x , y x,y x , y for p r 1 ( z ) , p r 2 ( z ) \mathrm{pr}_{1}(z),\mathrm{pr}_{2}(z) pr 1 ( z ) , pr 2 ( z ) in any formula for F ^ ( z ) \hat F(z) F ^ ( z ) , by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections . Lebesgue measure is σ \sigma σ -finite by Lebesgue Measure on Euclidean Space is Sigma-Finite §sigma-finite , and every member of P ( R d ) \mathcal{P}(\mathbb{R}^{d}) P ( R d ) is finite, hence σ \sigma σ -finite. So the Tonelli clause of Tonelli and Fubini Theorems applies to such integrands for any two of these measures: the iterated integrals exist in either order and are equal, and the partial integrals are measurable functions of the remaining variable. Every integrand to which this is applied below has this form, F ^ \hat F F ^ being built from p r 1 \mathrm{pr}_{1} pr 1 , p r 2 \mathrm{pr}_{2} pr 2 , π \pi π , g s g_{s} g s , continuous functions and indicators of Borel sets by composition, sums and products, and hence Borel by Probability Measures on Euclidean Space and Random Vectors: Standing Notation §borel-maps .
(c) The integer part. For y ∈ R d y\in\mathbb{R}^{d} y ∈ R d put n ( y ) = y − π ( y ) n(y)=y-\pi(y) n ( y ) = y − π ( y ) . By The Half-Open Unit Cell Tiles Euclidean Space §wrap , n ( y ) n(y) n ( y ) is the unique m ∈ Z d m\in\mathbb{Z}^{d} m ∈ Z d with y − m ∈ Q y-m\in Q y − m ∈ Q provided by The Half-Open Unit Cell Tiles Euclidean Space §tiling ; n n n is Borel, being the composite of the Borel pairing ( i d , π ) (\mathrm{id},\pi) ( id , π ) (Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §pairing ) with the map z ↦ p r 1 ( z ) − p r 2 ( z ) z\mapsto\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z) z ↦ pr 1 ( z ) − pr 2 ( z ) , which is Borel as a componentwise difference of the Borel coordinate projections (Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections , claim 2 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions ); and n ( y + k ) = y + k − π ( y ) = n ( y ) + k n(y+k)=y+k-\pi(y)=n(y)+k n ( y + k ) = y + k − π ( y ) = n ( y ) + k for k ∈ Z d k\in\mathbb{Z}^{d} k ∈ Z d , as π ( y + k ) = π ( y ) \pi(y+k)=\pi(y) π ( y + k ) = π ( y ) by the same clause. We record three facts.
(c1) For k ∈ Z d k\in\mathbb{Z}^{d} k ∈ Z d : n ( y ) = k n(y)=k n ( y ) = k if and only if y − k ∈ Q y-k\in Q y − k ∈ Q . Indeed, if n ( y ) = k n(y)=k n ( y ) = k then y − k = π ( y ) ∈ Q y-k=\pi(y)\in Q y − k = π ( y ) ∈ Q ; conversely, if y − k ∈ Q y-k\in Q y − k ∈ Q then k k k is the unique m ∈ Z d m\in\mathbb{Z}^{d} m ∈ Z d with y − m ∈ Q y-m\in Q y − m ∈ Q (The Half-Open Unit Cell Tiles Euclidean Space §tiling ), which is n ( y ) n(y) n ( y ) . Hence { y : n ( y ) = k } = Q + k \{y:n(y)=k\}=Q+k { y : n ( y ) = k } = Q + k , a Borel set with λ d ( Q + k ) = λ d ( Q ) = 1 \lambda_{d}(Q+k)=\lambda_{d}(Q)=1 λ d ( Q + k ) = λ d ( Q ) = 1 by claim 1 of Translation and Reflection Invariance of Lebesgue Measure on R n \mathbb{R}^n R n and The Half-Open Unit Cell Tiles Euclidean Space §cell . In particular (k = 0 k=0 k = 0 ) n ( y ) = 0 n(y)=0 n ( y ) = 0 for every y ∈ Q y\in Q y ∈ Q .
(c2) For z ∈ R d z\in\mathbb{R}^{d} z ∈ R d and Borel B ⊆ Q B\subseteq Q B ⊆ Q :
∫ 1 B ( z − n ( y ) ) d y = 1 B ( π ( z ) ) and ∫ 1 B ( z + n ( y ) ) d y = 1 B ( π ( z ) ) . \int\mathbf{1}_{B}\bigl(z-n(y)\bigr)\,dy=\mathbf{1}_{B}\bigl(\pi(z)\bigr)\qquad\text{and}\qquad\int\mathbf{1}_{B}\bigl(z+n(y)\bigr)\,dy=\mathbf{1}_{B}\bigl(\pi(z)\bigr). ∫ 1 B ( z − n ( y ) ) d y = 1 B ( π ( z ) ) and ∫ 1 B ( z + n ( y ) ) d y = 1 B ( π ( z ) ) .
For the first identity: n ( y ) ∈ Z d n(y)\in\mathbb{Z}^{d} n ( y ) ∈ Z d , so by (c1) applied to z z z and k = n ( y ) k=n(y) k = n ( y ) , z − n ( y ) ∈ Q z-n(y)\in Q z − n ( y ) ∈ Q holds exactly when n ( y ) = n ( z ) n(y)=n(z) n ( y ) = n ( z ) , and then z − n ( y ) = π ( z ) z-n(y)=\pi(z) z − n ( y ) = π ( z ) . As B ⊆ Q B\subseteq Q B ⊆ Q , the set { y : z − n ( y ) ∈ B } \{y:z-n(y)\in B\} { y : z − n ( y ) ∈ B } is { y : n ( y ) = n ( z ) } = Q + n ( z ) \{y:n(y)=n(z)\}=Q+n(z) { y : n ( y ) = n ( z )} = Q + n ( z ) if π ( z ) ∈ B \pi(z)\in B π ( z ) ∈ B and is empty otherwise; its indicator is the integrand, whose integral is its λ d \lambda_{d} λ d -measure (Measure Spaces and the Lebesgue Integral: Standing Notation §integral ), namely 1 1 1 or 0 0 0 by (c1). For the second: − n ( y ) ∈ Z d -n(y)\in\mathbb{Z}^{d} − n ( y ) ∈ Z d (Lattice-Periodic Functions and the Periodic Function Classes §lattice ), so by (c1) applied to z z z and k = − n ( y ) k=-n(y) k = − n ( y ) , z + n ( y ) = z − k ∈ Q z+n(y)=z-k\in Q z + n ( y ) = z − k ∈ Q holds exactly when − n ( y ) = n ( z ) -n(y)=n(z) − n ( y ) = n ( z ) , that is, n ( y ) = − n ( z ) n(y)=-n(z) n ( y ) = − n ( z ) , and then z + n ( y ) = z − n ( z ) = π ( z ) z+n(y)=z-n(z)=\pi(z) z + n ( y ) = z − n ( z ) = π ( z ) . So the set { y : z + n ( y ) ∈ B } \{y:z+n(y)\in B\} { y : z + n ( y ) ∈ B } is { y : n ( y ) = − n ( z ) } = Q − n ( z ) \{y:n(y)=-n(z)\}=Q-n(z) { y : n ( y ) = − n ( z )} = Q − n ( z ) if π ( z ) ∈ B \pi(z)\in B π ( z ) ∈ B and is empty otherwise, and its measure is 1 1 1 or 0 0 0 by (c1).
(d) Heat smoothing evaluated on a set. Let κ ∈ P ( R d ) \kappa\in\mathcal{P}(\mathbb{R}^{d}) κ ∈ P ( R d ) and B ∈ B ( R d ) B\in\mathcal{B}(\mathbb{R}^{d}) B ∈ B ( R d ) . Then
π # κ s ( B ) = ∫ ( ∫ 1 B ( π ( z ) ) g s ( z − w ) d z ) κ ( d w ) . ( W ) \pi_{\#}\kappa_{s}(B)=\int\Bigl(\int\mathbf{1}_{B}\bigl(\pi(z)\bigr)\,g_{s}(z-w)\,dz\Bigr)\kappa(dw). \qquad(\mathrm{W}) π # κ s ( B ) = ∫ ( ∫ 1 B ( π ( z ) ) g s ( z − w ) d z ) κ ( d w ) . ( W )
Indeed, φ B = 1 B ∘ π \varphi_{B}=\mathbf{1}_{B}\circ\pi φ B = 1 B ∘ π is Borel, π \pi π being Borel (The Half-Open Unit Cell Tiles Euclidean Space §wrap ), takes values in { 0 , 1 } \{0,1\} { 0 , 1 } , and is the indicator of π − 1 ( B ) \pi^{-1}(B) π − 1 ( B ) . By the definition of the push-forward (Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pushforward ), claim 3 of Image Measures, Measures with Densities, and Change of Variables , and Gaussian Smoothing of a Probability Measure on Euclidean Space: Regularity, Mass, Duality, Approximation of a Bounded Function with a Modulus of Continuity, and the Pairing Identity §duality with ϕ = φ B \phi=\varphi_{B} ϕ = φ B and M = 1 M=1 M = 1 ,
π # κ s ( B ) = κ s ( π − 1 ( B ) ) = ∫ φ B ( g s ∗ κ ) d λ d = ∫ g s ∗ φ B d κ , \pi_{\#}\kappa_{s}(B)=\kappa_{s}\bigl(\pi^{-1}(B)\bigr)=\int\varphi_{B}\,(g_{s}*\kappa)\,d\lambda_{d}=\int g_{s}*\varphi_{B}\,d\kappa, π # κ s ( B ) = κ s ( π − 1 ( B ) ) = ∫ φ B ( g s ∗ κ ) d λ d = ∫ g s ∗ φ B d κ ,
and ( g s ∗ φ B ) ( w ) = ∫ g s ( w − z ) φ B ( z ) d z = ∫ 1 B ( π ( z ) ) g s ( z − w ) d z (g_{s}*\varphi_{B})(w)=\int g_{s}(w-z)\varphi_{B}(z)\,dz=\int\mathbf{1}_{B}(\pi(z))\,g_{s}(z-w)\,dz ( g s ∗ φ B ) ( w ) = ∫ g s ( w − z ) φ B ( z ) d z = ∫ 1 B ( π ( z )) g s ( z − w ) d z by the evenness of g s g_{s} g s .
Step 1 (Kernels and a Gaussian bound). Let h 0 = g s h_{0}=g_{s} h 0 = g s and, for i , j ∈ [ d ] i,j\in[d] i , j ∈ [ d ] , h i = ∂ i g s h_{i}=\partial_{i}g_{s} h i = ∂ i g s and h i j = ∂ j ∂ i g s h_{ij}=\partial_{j}\partial_{i}g_{s} h ij = ∂ j ∂ i g s ; we call these 1 + d + d 2 1+d+d^{2} 1 + d + d 2 functions the kernels . As g s g_{s} g s is smooth, it is of class C 3 C^{3} C 3 (Smooth Map on a Euclidean Open Set ); by clause 2 of C^k Maps on a Euclidean Open Set each h i h_{i} h i is of class C 2 C^{2} C 2 and each h i j h_{ij} h ij exists and is of class C 1 C^{1} C 1 , so every kernel is continuous (claim 3 of Euclidean Space is Open in Itself, and C k C^k C k Maps are Continuous ), hence Borel, and ∂ i h 0 = h i \partial_{i}h_{0}=h_{i} ∂ i h 0 = h i , ∂ j h i = h i j \partial_{j}h_{i}=h_{ij} ∂ j h i = h ij everywhere.
Let e ( y ) = exp ( − ∥ y ∥ 2 / 2 ) e(y)=\exp(-\lVert y\rVert^{2}/2) e ( y ) = exp ( − ∥ y ∥ 2 /2 ) ; this is the function ψ 1 \psi_{1} ψ 1 of The Gaussian Smoothing Weight: Normalization, Derivatives, Exponential Tilting, Moments, and First-Order Remainder (there with η = 1 \eta=1 η = 1 ), so by claim 1 there it is Borel and positive with ∫ e d y < ∞ \int e\,dy<\infty ∫ e d y < ∞ , that is, integrable. Put K s = s − 1 + s − 2 K_{s}=s^{-1}+s^{-2} K s = s − 1 + s − 2 , so K s ≥ 1 K_{s}\ge1 K s ≥ 1 as s − 1 ≥ 2 s^{-1}\ge2 s − 1 ≥ 2 , and for real R ≥ 0 R\ge0 R ≥ 0 put b R = R + d b_{R}=R+\sqrt{d} b R = R + d and M R = 5 K s α s exp ( 3 2 b R 2 ) M_{R}=5K_{s}\alpha_{s}\exp(\tfrac32b_{R}^{2}) M R = 5 K s α s exp ( 2 3 b R 2 ) . We claim: for every kernel h h h , every v v v with ∥ v ∥ ≤ R \lVert v\rVert\le R ∥ v ∥ ≤ R and every y y y ,
∣ h ( v + n ( y ) ) ∣ ≤ M R e ( y ) . ( D ) \bigl|h\bigl(v+n(y)\bigr)\bigr|\le M_{R}\,e(y). \qquad(\mathrm{D}) h ( v + n ( y ) ) ≤ M R e ( y ) . ( D )
Put u = v + n ( y ) u=v+n(y) u = v + n ( y ) . First, ∣ h ( u ) ∣ ≤ K s ( 1 + ∥ u ∥ 2 ) g s ( u ) |h(u)|\le K_{s}(1+\lVert u\rVert^{2})g_{s}(u) ∣ h ( u ) ∣ ≤ K s ( 1 + ∥ u ∥ 2 ) g s ( u ) : for h 0 h_{0} h 0 because K s ≥ 1 K_{s}\ge1 K s ≥ 1 ; for h i h_{i} h i because ∣ u i ∣ ≤ ∥ u ∥ ≤ 1 + ∥ u ∥ 2 |u_{i}|\le\lVert u\rVert\le1+\lVert u\rVert^{2} ∣ u i ∣ ≤ ∥ u ∥ ≤ 1 + ∥ u ∥ 2 (claim 4 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n , and 0 ≤ ( 1 − ∥ u ∥ ) 2 0\le(1-\lVert u\rVert)^{2} 0 ≤ ( 1 − ∥ u ∥ ) 2 ) and s − 1 ≤ K s s^{-1}\le K_{s} s − 1 ≤ K s ; for h i j h_{ij} h ij because ∣ s − 2 u i u j − s − 1 δ i j ∣ ≤ s − 2 ∥ u ∥ 2 + s − 1 ≤ K s ( 1 + ∥ u ∥ 2 ) |s^{-2}u_{i}u_{j}-s^{-1}\delta_{ij}|\le s^{-2}\lVert u\rVert^{2}+s^{-1}\le K_{s}(1+\lVert u\rVert^{2}) ∣ s − 2 u i u j − s − 1 δ ij ∣ ≤ s − 2 ∥ u ∥ 2 + s − 1 ≤ K s ( 1 + ∥ u ∥ 2 ) . Next let t = ∥ u ∥ 2 / ( 8 s ) ≥ 0 t=\lVert u\rVert^{2}/(8s)\ge0 t = ∥ u ∥ 2 / ( 8 s ) ≥ 0 . By claim 4 of Basic Properties of the Exponential Function , exp ( t ) ≥ 1 + t \exp(t)\ge1+t exp ( t ) ≥ 1 + t , so 1 ≤ exp ( t ) 1\le\exp(t) 1 ≤ exp ( t ) and ∥ u ∥ 2 = 8 s t ≤ 8 s exp ( t ) \lVert u\rVert^{2}=8st\le8s\exp(t) ∥ u ∥ 2 = 8 s t ≤ 8 s exp ( t ) , whence 1 + ∥ u ∥ 2 ≤ ( 1 + 8 s ) exp ( t ) ≤ 5 exp ( t ) 1+\lVert u\rVert^{2}\le(1+8s)\exp(t)\le5\exp(t) 1 + ∥ u ∥ 2 ≤ ( 1 + 8 s ) exp ( t ) ≤ 5 exp ( t ) . By claims 1 and 2 there, g s ( u ) = α s exp ( − t ) exp ( − 3 t ) g_{s}(u)=\alpha_{s}\exp(-t)\exp(-3t) g s ( u ) = α s exp ( − t ) exp ( − 3 t ) and exp ( − t ) = exp ( t ) − 1 \exp(-t)=\exp(t)^{-1} exp ( − t ) = exp ( t ) − 1 , so
∣ h ( u ) ∣ ≤ 5 K s α s exp ( − 3 ∥ u ∥ 2 8 s ) ≤ 5 K s α s exp ( − 3 ∥ u ∥ 2 4 ) , |h(u)|\le5K_{s}\alpha_{s}\exp\Bigl(-\frac{3\lVert u\rVert^{2}}{8s}\Bigr)\le5K_{s}\alpha_{s}\exp\Bigl(-\frac{3\lVert u\rVert^{2}}{4}\Bigr), ∣ h ( u ) ∣ ≤ 5 K s α s exp ( − 8 s 3 ∥ u ∥ 2 ) ≤ 5 K s α s exp ( − 4 3 ∥ u ∥ 2 ) ,
the last step because 3 8 s ≥ 3 4 \frac{3}{8s}\ge\frac34 8 s 3 ≥ 4 3 and exp \exp exp is increasing (claim 4 there). Now y = n ( y ) + π ( y ) = u − v + π ( y ) y=n(y)+\pi(y)=u-v+\pi(y) y = n ( y ) + π ( y ) = u − v + π ( y ) with ∥ π ( y ) ∥ ≤ d \lVert\pi(y)\rVert\le\sqrt{d} ∥ π ( y )∥ ≤ d (Step 0(a), as π ( y ) ∈ Q \pi(y)\in Q π ( y ) ∈ Q ), so ∥ y ∥ ≤ ∥ u ∥ + b R \lVert y\rVert\le\lVert u\rVert+b_{R} ∥ y ∥ ≤ ∥ u ∥ + b R by the triangle inequality (claims 5 and 6 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n ). Since 2 ∥ u ∥ b R ≤ 1 2 ∥ u ∥ 2 + 2 b R 2 2\lVert u\rVert b_{R}\le\tfrac12\lVert u\rVert^{2}+2b_{R}^{2} 2 ∥ u ∥ b R ≤ 2 1 ∥ u ∥ 2 + 2 b R 2 (expand 0 ≤ ( 1 2 ∥ u ∥ − 2 b R ) 2 0\le(\tfrac{1}{\sqrt2}\lVert u\rVert-\sqrt2b_{R})^{2} 0 ≤ ( 2 1 ∥ u ∥ − 2 b R ) 2 ) and squaring is monotone on nonnegative numbers (claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field ), ∥ y ∥ 2 ≤ ( ∥ u ∥ + b R ) 2 ≤ 3 2 ∥ u ∥ 2 + 3 b R 2 \lVert y\rVert^{2}\le(\lVert u\rVert+b_{R})^{2}\le\tfrac32\lVert u\rVert^{2}+3b_{R}^{2} ∥ y ∥ 2 ≤ (∥ u ∥ + b R ) 2 ≤ 2 3 ∥ u ∥ 2 + 3 b R 2 , that is, 3 4 ∥ u ∥ 2 ≥ 1 2 ∥ y ∥ 2 − 3 2 b R 2 \tfrac34\lVert u\rVert^{2}\ge\tfrac12\lVert y\rVert^{2}-\tfrac32b_{R}^{2} 4 3 ∥ u ∥ 2 ≥ 2 1 ∥ y ∥ 2 − 2 3 b R 2 . Using once more that exp \exp exp is increasing and multiplicative, exp ( − 3 4 ∥ u ∥ 2 ) ≤ exp ( 3 2 b R 2 ) e ( y ) \exp(-\tfrac34\lVert u\rVert^{2})\le\exp(\tfrac32b_{R}^{2})\,e(y) exp ( − 4 3 ∥ u ∥ 2 ) ≤ exp ( 2 3 b R 2 ) e ( y ) , which proves (D).
Step 2 (The periodised Gaussian and its derivatives).
(2a) Definition, periodicity and continuity. For a kernel h h h and v ∈ R d v\in\mathbb{R}^{d} v ∈ R d , the function y ↦ h ( v + n ( y ) ) y\mapsto h(v+n(y)) y ↦ h ( v + n ( y )) is Borel (a continuous function of the Borel map y ↦ v + n ( y ) y\mapsto v+n(y) y ↦ v + n ( y ) , Step 0(c)) and dominated by M ∥ v ∥ e M_{\lVert v\rVert}e M ∥ v ∥ e by (D), hence integrable; put
H h ( v ) = ∫ h ( v + n ( y ) ) d y ∈ R , P = H h 0 , P i = H h i , P i j = H h i j . H_{h}(v)=\int h\bigl(v+n(y)\bigr)\,dy\in\mathbb{R},\qquad P=H_{h_{0}},\quad P_{i}=H_{h_{i}},\quad P_{ij}=H_{h_{ij}} . H h ( v ) = ∫ h ( v + n ( y ) ) d y ∈ R , P = H h 0 , P i = H h i , P ij = H h ij .
Periodicity. For k ∈ Z d k\in\mathbb{Z}^{d} k ∈ Z d , h ( v + k + n ( y ) ) = h ( v + n ( y + k ) ) h(v+k+n(y))=h(v+n(y+k)) h ( v + k + n ( y )) = h ( v + n ( y + k )) by Step 0(c), and claim 3 of Translation and Reflection Invariance of Lebesgue Measure on R n \mathbb{R}^n R n with a = k a=k a = k , applied to the integrable function y ↦ h ( v + n ( y ) ) y\mapsto h(v+n(y)) y ↦ h ( v + n ( y )) , gives H h ( v + k ) = ∫ h ( v + n ( y + k ) ) d y = H h ( v ) H_{h}(v+k)=\int h(v+n(y+k))\,dy=H_{h}(v) H h ( v + k ) = ∫ h ( v + n ( y + k )) d y = H h ( v ) .
Continuity. Let ( v j ) j ∈ N (v_{j})_{j\in\mathbb{N}} ( v j ) j ∈ N converge to v v v in ( R d , d E ) (\mathbb{R}^{d},d_{E}) ( R d , d E ) . Choose J ∈ N J\in\mathbb{N} J ∈ N with ∥ v j − v ∥ < 1 \lVert v_{j}-v\rVert<1 ∥ v j − v ∥ < 1 for all j ≥ J j\ge J j ≥ J (Convergent Sequence in a Metric Space , with d E ( v j , v ) = ∥ v j − v ∥ d_{E}(v_{j},v)=\lVert v_{j}-v\rVert d E ( v j , v ) = ∥ v j − v ∥ by claim 2 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n ), put R = ∥ v ∥ + 1 R=\lVert v\rVert+1 R = ∥ v ∥ + 1 and w j = v J + j w_{j}=v_{J+j} w j = v J + j , so that ∥ w j ∥ ≤ R \lVert w_{j}\rVert\le R ∥ w j ∥ ≤ R . For each y y y , d E ( w j + n ( y ) , v + n ( y ) ) = ∥ w j − v ∥ d_{E}(w_{j}+n(y),v+n(y))=\lVert w_{j}-v\rVert d E ( w j + n ( y ) , v + n ( y )) = ∥ w j − v ∥ , so w j + n ( y ) → v + n ( y ) w_{j}+n(y)\to v+n(y) w j + n ( y ) → v + n ( y ) and h ( w j + n ( y ) ) → h ( v + n ( y ) ) h(w_{j}+n(y))\to h(v+n(y)) h ( w j + n ( y )) → h ( v + n ( y )) by claim 1 of Continuity Between Metric Spaces is Equivalent to Sequential Continuity . By (D) the functions y ↦ h ( w j + n ( y ) ) y\mapsto h(w_{j}+n(y)) y ↦ h ( w j + n ( y )) are all dominated by the integrable M R e M_{R}e M R e , so Dominated Convergence Theorem gives H h ( w j ) → H h ( v ) H_{h}(w_{j})\to H_{h}(v) H h ( w j ) → H h ( v ) ; given ε > 0 \varepsilon>0 ε > 0 and L L L with ∣ H h ( w j ) − H h ( v ) ∣ < ε |H_{h}(w_{j})-H_{h}(v)|<\varepsilon ∣ H h ( w j ) − H h ( v ) ∣ < ε for j ≥ L j\ge L j ≥ L , we get ∣ H h ( v j ) − H h ( v ) ∣ < ε |H_{h}(v_{j})-H_{h}(v)|<\varepsilon ∣ H h ( v j ) − H h ( v ) ∣ < ε for j ≥ J + L j\ge J+L j ≥ J + L . By claim 3 of Continuity Between Metric Spaces is Equivalent to Sequential Continuity , H h H_{h} H h is continuous on R d \mathbb{R}^{d} R d ; this is also continuity at every point in the sense of Continuity at a Point for Maps Between Euclidean Spaces , whose conditions ∑ i ( x i − a i ) 2 < δ 2 \sum_{i}(x_{i}-a_{i})^{2}<\delta^{2} ∑ i ( x i − a i ) 2 < δ 2 and ( f ( x ) − f ( a ) ) 2 < ε 2 (f(x)-f(a))^{2}<\varepsilon^{2} ( f ( x ) − f ( a ) ) 2 < ε 2 read ∥ x − a ∥ < δ \lVert x-a\rVert<\delta ∥ x − a ∥ < δ and ∣ f ( x ) − f ( a ) ∣ < ε |f(x)-f(a)|<\varepsilon ∣ f ( x ) − f ( a ) ∣ < ε by claim 1 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n and claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field . Thus every H h H_{h} H h belongs to C p e r C_{\mathrm{per}} C per .
(2b) Derivatives. Let h h h be a kernel and m ∈ [ d ] m\in[d] m ∈ [ d ] such that ∂ m h \partial_{m}h ∂ m h is again a kernel (namely h = h 0 h=h_{0} h = h 0 , ∂ m h = h m \partial_{m}h=h_{m} ∂ m h = h m , or h = h i h=h_{i} h = h i , ∂ m h = h i m \partial_{m}h=h_{im} ∂ m h = h im ). We show that ∂ m H h ( v ) \partial_{m}H_{h}(v) ∂ m H h ( v ) exists and equals H ∂ m h ( v ) H_{\partial_{m}h}(v) H ∂ m h ( v ) for every v v v . Fix v v v ; for real t t t let v [ t ] v[t] v [ t ] be the point with m m m th coordinate t t t and the other coordinates those of v v v , as in Slice Function and the Partial Derivative , and let U U U be the open interval with endpoints v m − 1 v_{m}-1 v m − 1 and v m + 1 v_{m}+1 v m + 1 . For t ∈ U t\in U t ∈ U , v [ t ] − v v[t]-v v [ t ] − v has the single nonzero coordinate t − v m t-v_{m} t − v m , so ∥ v [ t ] ∥ ≤ ∥ v ∥ + 1 = : R \lVert v[t]\rVert\le\lVert v\rVert+1=:R ∥ v [ t ]∥ ≤ ∥ v ∥ + 1 =: R (claims 1 and 6 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n ). Apply Differentiation under the Integral Sign to ( R d , B ( R d ) , λ d ) (\mathbb{R}^{d},\mathcal{B}(\mathbb{R}^{d}),\lambda_{d}) ( R d , B ( R d ) , λ d ) and f ( t , y ) = h ( v [ t ] + n ( y ) ) f(t,y)=h(v[t]+n(y)) f ( t , y ) = h ( v [ t ] + n ( y )) on U × R d U\times\mathbb{R}^{d} U × R d . Condition (i) holds by (2a). For (ii), fix y y y and t 0 ∈ U t_{0}\in U t 0 ∈ U and put b = v [ t 0 ] + n ( y ) b=v[t_{0}]+n(y) b = v [ t 0 ] + n ( y ) ; for real η ≠ 0 \eta\ne0 η = 0 the point v [ t 0 + η ] + n ( y ) v[t_{0}+\eta]+n(y) v [ t 0 + η ] + n ( y ) is b b b with its m m m th coordinate b m b_{m} b m replaced by b m + η b_{m}+\eta b m + η , so ( f ( t 0 + η , y ) − f ( t 0 , y ) ) / η (f(t_{0}+\eta,y)-f(t_{0},y))/\eta ( f ( t 0 + η , y ) − f ( t 0 , y )) / η is exactly the difference quotient of Partial Derivative on a Euclidean Open Set for h h h at b b b in the m m m th variable; as h h h is of class C 1 C^{1} C 1 , that definition's condition with value ∂ m h ( b ) \partial_{m}h(b) ∂ m h ( b ) holds, and it is word for word the condition that t ↦ f ( t , y ) t\mapsto f(t,y) t ↦ f ( t , y ) be differentiable at t 0 t_{0} t 0 with derivative ∂ m h ( b ) \partial_{m}h(b) ∂ m h ( b ) (Derivative at an Interior Point ). For (iii), ∣ ∂ m h ( v [ t ] + n ( y ) ) ∣ ≤ M R e ( y ) |\partial_{m}h(v[t]+n(y))|\le M_{R}e(y) ∣ ∂ m h ( v [ t ] + n ( y )) ∣ ≤ M R e ( y ) for t ∈ U t\in U t ∈ U by (D), ∂ m h \partial_{m}h ∂ m h being a kernel. The theorem shows that F ( t ) = H h ( v [ t ] ) F(t)=H_{h}(v[t]) F ( t ) = H h ( v [ t ]) is differentiable on U U U with F ′ ( t ) = H ∂ m h ( v [ t ] ) F'(t)=H_{\partial_{m}h}(v[t]) F ′ ( t ) = H ∂ m h ( v [ t ]) . Since the domain is all of R d \mathbb{R}^{d} R d , the radius 1 1 1 is admissible in claim 1 of Slice Function and the Partial Derivative , F F F is the slice function of H h H_{h} H h at v v v in the m m m th variable, and claim 2 there gives ∂ m H h ( v ) = F ′ ( v m ) = H ∂ m h ( v ) \partial_{m}H_{h}(v)=F'(v_{m})=H_{\partial_{m}h}(v) ∂ m H h ( v ) = F ′ ( v m ) = H ∂ m h ( v ) .
Consequently: P P P is continuous with partial derivatives ∂ i P = P i \partial_{i}P=P_{i} ∂ i P = P i that exist everywhere and are continuous, so P P P is of class C 1 C^{1} C 1 (clause 1 of C^k Maps on a Euclidean Open Set ); likewise each P i P_{i} P i is of class C 1 C^{1} C 1 with ∂ j P i = P i j \partial_{j}P_{i}=P_{ij} ∂ j P i = P ij ; hence P P P is of class C 2 C^{2} C 2 with ∂ j ∂ i P = P i j \partial_{j}\partial_{i}P=P_{ij} ∂ j ∂ i P = P ij (clause 2 there, with k = 1 k=1 k = 1 ), and P ∈ C p e r 2 P\in C^{2}_{\mathrm{per}} P ∈ C per 2 by (2a). By (2a) and Elementary Properties of Lattice-Periodic Functions §bounded each of the finitely many functions P P P , P i P_{i} P i , P i j P_{ij} P ij lies in C p e r C_{\mathrm{per}} C per and is bounded; let C s ≥ 0 C_{s}\ge0 C s ≥ 0 be a common bound (the largest of the individual bounds), so that ∣ P ( v ) ∣ ≤ C s |P(v)|\le C_{s} ∣ P ( v ) ∣ ≤ C s , ∣ ∂ i P ( v ) ∣ ≤ C s |\partial_{i}P(v)|\le C_{s} ∣ ∂ i P ( v ) ∣ ≤ C s and ∣ ∂ j ∂ i P ( v ) ∣ ≤ C s |\partial_{j}\partial_{i}P(v)|\le C_{s} ∣ ∂ j ∂ i P ( v ) ∣ ≤ C s for all v v v and i , j ∈ [ d ] i,j\in[d] i , j ∈ [ d ] .
Step 3 (A positive lower bound). Put c s = α s exp ( − d / ( 2 s ) ) c_{s}=\alpha_{s}\exp(-d/(2s)) c s = α s exp ( − d / ( 2 s )) , a positive real number by claim 2 of Basic Properties of the Exponential Function . Let v ∈ Q v\in Q v ∈ Q . For y ∈ Q y\in Q y ∈ Q one has n ( y ) = 0 n(y)=0 n ( y ) = 0 by (c1), so 1 Q ( y ) g s ( v + n ( y ) ) = g s ( v ) 1 Q ( y ) \mathbf{1}_{Q}(y)\,g_{s}(v+n(y))=g_{s}(v)\,\mathbf{1}_{Q}(y) 1 Q ( y ) g s ( v + n ( y )) = g s ( v ) 1 Q ( y ) for every y y y ; since 0 ≤ 1 Q ( y ) g s ( v + n ( y ) ) ≤ g s ( v + n ( y ) ) 0\le\mathbf{1}_{Q}(y)g_{s}(v+n(y))\le g_{s}(v+n(y)) 0 ≤ 1 Q ( y ) g s ( v + n ( y )) ≤ g s ( v + n ( y )) , the nonnegative case of Linearity and Monotonicity of the Lebesgue Integral (claim 1) gives
P ( v ) = ∫ g s ( v + n ( y ) ) d y ≥ ∫ 1 Q ( y ) g s ( v ) d y = g s ( v ) λ d ( Q ) = g s ( v ) , P(v)=\int g_{s}\bigl(v+n(y)\bigr)\,dy\ge\int\mathbf{1}_{Q}(y)\,g_{s}(v)\,dy=g_{s}(v)\,\lambda_{d}(Q)=g_{s}(v), P ( v ) = ∫ g s ( v + n ( y ) ) d y ≥ ∫ 1 Q ( y ) g s ( v ) d y = g s ( v ) λ d ( Q ) = g s ( v ) ,
using ∫ 1 Q d λ d = λ d ( Q ) \int\mathbf{1}_{Q}\,d\lambda_{d}=\lambda_{d}(Q) ∫ 1 Q d λ d = λ d ( Q ) (Measure Spaces and the Lebesgue Integral: Standing Notation §integral ) and λ d ( Q ) = 1 \lambda_{d}(Q)=1 λ d ( Q ) = 1 (The Half-Open Unit Cell Tiles Euclidean Space §cell ). By Step 0(a), ∥ v ∥ 2 ≤ d \lVert v\rVert^{2}\le d ∥ v ∥ 2 ≤ d , so − ∥ v ∥ 2 / ( 2 s ) ≥ − d / ( 2 s ) -\lVert v\rVert^{2}/(2s)\ge-d/(2s) − ∥ v ∥ 2 / ( 2 s ) ≥ − d / ( 2 s ) and, exp \exp exp being increasing (claim 4 of Basic Properties of the Exponential Function ), g s ( v ) ≥ c s g_{s}(v)\ge c_{s} g s ( v ) ≥ c s . Now let v ∈ R d v\in\mathbb{R}^{d} v ∈ R d be arbitrary. Then π ( v ) ∈ Q \pi(v)\in Q π ( v ) ∈ Q and v = π ( v ) + n ( v ) v=\pi(v)+n(v) v = π ( v ) + n ( v ) with n ( v ) ∈ Z d n(v)\in\mathbb{Z}^{d} n ( v ) ∈ Z d (Step 0(c)), so by periodicity (2a), P ( v ) = P ( π ( v ) ) ≥ c s P(v)=P(\pi(v))\ge c_{s} P ( v ) = P ( π ( v )) ≥ c s . Together with Step 2, c s ≤ P ( v ) ≤ C s c_{s}\le P(v)\le C_{s} c s ≤ P ( v ) ≤ C s for every v v v ; in particular 0 < c s ≤ C s 0<c_{s}\le C_{s} 0 < c s ≤ C s .
Step 4 (Cell identities). For w ∈ R d w\in\mathbb{R}^{d} w ∈ R d and Borel B ⊆ Q B\subseteq Q B ⊆ Q ,
∫ 1 B ( x ) P ( x − w ) d x = ∫ 1 B ( π ( z ) ) g s ( z − w ) d z , ( E + ) \int\mathbf{1}_{B}(x)\,P(x-w)\,dx=\int\mathbf{1}_{B}\bigl(\pi(z)\bigr)\,g_{s}(z-w)\,dz, \qquad(\mathrm{E}^{+}) ∫ 1 B ( x ) P ( x − w ) d x = ∫ 1 B ( π ( z ) ) g s ( z − w ) d z , ( E + )
∫ 1 B ( x ) P ( − x ) d x = ∫ 1 B ( π ( z ) ) g s ( z ) d z . ( E − ) \int\mathbf{1}_{B}(x)\,P(-x)\,dx=\int\mathbf{1}_{B}\bigl(\pi(z)\bigr)\,g_{s}(z)\,dz. \qquad(\mathrm{E}^{-}) ∫ 1 B ( x ) P ( − x ) d x = ∫ 1 B ( π ( z ) ) g s ( z ) d z . ( E − )
All integrands here are nonnegative and Borel (P P P is continuous and positive by Steps 2 and 3).
Proof of ( E + ) (\mathrm{E}^{+}) ( E + ) . By the definition of P P P , taking the constant 1 B ( x ) \mathbf{1}_{B}(x) 1 B ( x ) inside the inner integral (claim 1 of Linearity and Monotonicity of the Lebesgue Integral ) and using Step 0(b) with λ d \lambda_{d} λ d twice, the left side is ∫ ( ∫ 1 B ( x ) g s ( x + n ( y ) − w ) d x ) d y \int\bigl(\int\mathbf{1}_{B}(x)g_{s}(x+n(y)-w)\,dx\bigr)dy ∫ ( ∫ 1 B ( x ) g s ( x + n ( y ) − w ) d x ) d y . For fixed y y y , claim 2 of Translation and Reflection Invariance of Lebesgue Measure on R n \mathbb{R}^n R n with a = n ( y ) a=n(y) a = n ( y ) , applied to the nonnegative Borel function f y ( z ) = 1 B ( z − n ( y ) ) g s ( z − w ) f_{y}(z)=\mathbf{1}_{B}(z-n(y))g_{s}(z-w) f y ( z ) = 1 B ( z − n ( y )) g s ( z − w ) , for which f y ( x + n ( y ) ) = 1 B ( x ) g s ( x + n ( y ) − w ) f_{y}(x+n(y))=\mathbf{1}_{B}(x)g_{s}(x+n(y)-w) f y ( x + n ( y )) = 1 B ( x ) g s ( x + n ( y ) − w ) , shows that the inner integral is ∫ 1 B ( z − n ( y ) ) g s ( z − w ) d z \int\mathbf{1}_{B}(z-n(y))g_{s}(z-w)\,dz ∫ 1 B ( z − n ( y )) g s ( z − w ) d z . Exchanging the order again by Step 0(b), taking the factor g s ( z − w ) g_{s}(z-w) g s ( z − w ) out of the inner integral (claim 1 of Linearity and Monotonicity of the Lebesgue Integral ) and using the first identity of Step 0(c2), the left side equals ∫ ( ∫ 1 B ( z − n ( y ) ) d y ) g s ( z − w ) d z = ∫ 1 B ( π ( z ) ) g s ( z − w ) d z \int\bigl(\int\mathbf{1}_{B}(z-n(y))\,dy\bigr)g_{s}(z-w)\,dz=\int\mathbf{1}_{B}(\pi(z))g_{s}(z-w)\,dz ∫ ( ∫ 1 B ( z − n ( y )) d y ) g s ( z − w ) d z = ∫ 1 B ( π ( z )) g s ( z − w ) d z .
Proof of ( E − ) (\mathrm{E}^{-}) ( E − ) . By the evenness of g s g_{s} g s , P ( − x ) = ∫ g s ( − x + n ( y ) ) d y = ∫ g s ( x − n ( y ) ) d y P(-x)=\int g_{s}(-x+n(y))\,dy=\int g_{s}(x-n(y))\,dy P ( − x ) = ∫ g s ( − x + n ( y )) d y = ∫ g s ( x − n ( y )) d y . As before, Step 0(b) turns the left side into ∫ ( ∫ 1 B ( x ) g s ( x − n ( y ) ) d x ) d y \int\bigl(\int\mathbf{1}_{B}(x)g_{s}(x-n(y))\,dx\bigr)dy ∫ ( ∫ 1 B ( x ) g s ( x − n ( y )) d x ) d y . For fixed y y y , claim 2 of Translation and Reflection Invariance of Lebesgue Measure on R n \mathbb{R}^n R n with a = − n ( y ) a=-n(y) a = − n ( y ) , applied to the nonnegative Borel function f y ( z ) = 1 B ( z + n ( y ) ) g s ( z ) f_{y}(z)=\mathbf{1}_{B}(z+n(y))g_{s}(z) f y ( z ) = 1 B ( z + n ( y )) g s ( z ) , for which f y ( x − n ( y ) ) = 1 B ( x ) g s ( x − n ( y ) ) f_{y}(x-n(y))=\mathbf{1}_{B}(x)g_{s}(x-n(y)) f y ( x − n ( y )) = 1 B ( x ) g s ( x − n ( y )) , shows that the inner integral is ∫ 1 B ( z + n ( y ) ) g s ( z ) d z \int\mathbf{1}_{B}(z+n(y))g_{s}(z)\,dz ∫ 1 B ( z + n ( y )) g s ( z ) d z . Exchanging the order by Step 0(b) and using the second identity of Step 0(c2), the left side equals ∫ ( ∫ 1 B ( z + n ( y ) ) d y ) g s ( z ) d z = ∫ 1 B ( π ( z ) ) g s ( z ) d z \int\bigl(\int\mathbf{1}_{B}(z+n(y))\,dy\bigr)g_{s}(z)\,dz=\int\mathbf{1}_{B}(\pi(z))g_{s}(z)\,dz ∫ ( ∫ 1 B ( z + n ( y )) d y ) g s ( z ) d z = ∫ 1 B ( π ( z )) g s ( z ) d z .
Step 5 (Identification of the kernel; clause 1). Let P − : R d → R P^{-}:\mathbb{R}^{d}\to\mathbb{R} P − : R d → R be P − ( v ) = P ( − v ) P^{-}(v)=P(-v) P − ( v ) = P ( − v ) . It is Z d \mathbb{Z}^{d} Z d -periodic, since − k ∈ Z d -k\in\mathbb{Z}^{d} − k ∈ Z d for k ∈ Z d k\in\mathbb{Z}^{d} k ∈ Z d and so P ( − v − k ) = P ( − v ) P(-v-k)=P(-v) P ( − v − k ) = P ( − v ) by (2a); and it is continuous, since for x , v ∈ R d x,v\in\mathbb{R}^{d} x , v ∈ R d one has ∥ ( − x ) − ( − v ) ∥ = ∥ x − v ∥ \lVert(-x)-(-v)\rVert=\lVert x-v\rVert ∥( − x ) − ( − v )∥ = ∥ x − v ∥ (claim 5 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n ), so a δ \delta δ witnessing the continuity of P P P at − v -v − v for a given ε \varepsilon ε witnesses that of P − P^{-} P − at v v v (Continuous Map Between Metric Spaces ). Thus P , P − ∈ C p e r P,P^{-}\in C_{\mathrm{per}} P , P − ∈ C per , and both are positive by Step 3.
Recall from the statement of Existence and Uniqueness of a Continuous Periodic Density for the Torus Heat Semigroup Started at the Origin that δ 0 ∈ P ( T d ) \delta_{0}\in\mathcal{P}(\mathbb{T}^{d}) δ 0 ∈ P ( T d ) . Let B ∈ B ( R d ) B\in\mathcal{B}(\mathbb{R}^{d}) B ∈ B ( R d ) with B ⊆ Q B\subseteq Q B ⊆ Q . By (W) with κ = δ 0 \kappa=\delta_{0} κ = δ 0 and Dirac Measures on Euclidean Space: Probability Measure, Integrals, Push-Forwards and the Coupling of Two Dirac Measures §integral (the inner integral in (W) being a nonnegative Borel function of w w w by Step 0(b)),
S s δ 0 ( B ) = π # ( δ 0 ) s ( B ) = ∫ 1 B ( π ( z ) ) g s ( z ) d z . S_{s}\delta_{0}(B)=\pi_{\#}(\delta_{0})_{s}(B)=\int\mathbf{1}_{B}\bigl(\pi(z)\bigr)\,g_{s}(z)\,dz . S s δ 0 ( B ) = π # ( δ 0 ) s ( B ) = ∫ 1 B ( π ( z ) ) g s ( z ) d z .
By ( E + ) (\mathrm{E}^{+}) ( E + ) with w = 0 w=0 w = 0 and by ( E − ) (\mathrm{E}^{-}) ( E − ) , both ∫ 1 B P d λ d \int\mathbf{1}_{B}P\,d\lambda_{d} ∫ 1 B P d λ d and ∫ 1 B P − d λ d \int\mathbf{1}_{B}P^{-}\,d\lambda_{d} ∫ 1 B P − d λ d equal this number. Now let A ∈ B ( R d ) A\in\mathcal{B}(\mathbb{R}^{d}) A ∈ B ( R d ) be arbitrary. As S s δ 0 ∈ P ( T d ) S_{s}\delta_{0}\in\mathcal{P}(\mathbb{T}^{d}) S s δ 0 ∈ P ( T d ) , Step 0(a) and the case B = A ∩ Q B=A\cap Q B = A ∩ Q give, for G ∈ { P , P − } G\in\{P,P^{-}\} G ∈ { P , P − } ,
S s δ 0 ( A ) = S s δ 0 ( A ∩ Q ) = ∫ 1 A ∩ Q G d λ d = ∫ 1 A ( 1 Q G ) d λ d . S_{s}\delta_{0}(A)=S_{s}\delta_{0}(A\cap Q)=\int\mathbf{1}_{A\cap Q}\,G\,d\lambda_{d}=\int\mathbf{1}_{A}\,(\mathbf{1}_{Q}G)\,d\lambda_{d}. S s δ 0 ( A ) = S s δ 0 ( A ∩ Q ) = ∫ 1 A ∩ Q G d λ d = ∫ 1 A ( 1 Q G ) d λ d .
The functions 1 Q P \mathbf{1}_{Q}P 1 Q P and 1 Q P − \mathbf{1}_{Q}P^{-} 1 Q P − are Borel and nonnegative, so both are densities of S s δ 0 S_{s}\delta_{0} S s δ 0 with respect to λ d \lambda_{d} λ d . By The Torus Heat Kernel §kernel , Θ s \Theta_{s} Θ s is the unique member of C p e r C_{\mathrm{per}} C per with this property, uniqueness holding by Existence and Uniqueness of a Continuous Periodic Density for the Torus Heat Semigroup Started at the Origin §kernel ; hence
Θ s = P = P − . \Theta_{s}=P=P^{-}. Θ s = P = P − .
Consequently Θ s ∈ C p e r 2 \Theta_{s}\in C^{2}_{\mathrm{per}} Θ s ∈ C per 2 (Step 2), Θ s ( − v ) = P − ( v ) = P ( v ) = Θ s ( v ) \Theta_{s}(-v)=P^{-}(v)=P(v)=\Theta_{s}(v) Θ s ( − v ) = P − ( v ) = P ( v ) = Θ s ( v ) for every v v v , and, with c s c_{s} c s and C s C_{s} C s from Steps 2 and 3, 0 < c s ≤ C s 0<c_{s}\le C_{s} 0 < c s ≤ C s , c s ≤ Θ s ( v ) ≤ C s c_{s}\le\Theta_{s}(v)\le C_{s} c s ≤ Θ s ( v ) ≤ C s , ∣ ∂ i Θ s ( v ) ∣ ≤ C s |\partial_{i}\Theta_{s}(v)|\le C_{s} ∣ ∂ i Θ s ( v ) ∣ ≤ C s and ∣ ∂ j ∂ i Θ s ( v ) ∣ ≤ C s |\partial_{j}\partial_{i}\Theta_{s}(v)|\le C_{s} ∣ ∂ j ∂ i Θ s ( v ) ∣ ≤ C s for all v v v and i , j ∈ [ d ] i,j\in[d] i , j ∈ [ d ] . Finally, taking A = R d A=\mathbb{R}^{d} A = R d above and using that S s δ 0 S_{s}\delta_{0} S s δ 0 is a probability measure, ∫ 1 Q Θ s d λ d = S s δ 0 ( R d ) = 1 \int\mathbf{1}_{Q}\Theta_{s}\,d\lambda_{d}=S_{s}\delta_{0}(\mathbb{R}^{d})=1 ∫ 1 Q Θ s d λ d = S s δ 0 ( R d ) = 1 . This proves clause 1. We write ∂ i Θ s = P i \partial_{i}\Theta_{s}=P_{i} ∂ i Θ s = P i and ∂ j ∂ i Θ s = P i j \partial_{j}\partial_{i}\Theta_{s}=P_{ij} ∂ j ∂ i Θ s = P ij as in Step 2.
Step 6 (The density of a heat-smoothed measure). Let μ ∈ P ( T d ) \mu\in\mathcal{P}(\mathbb{T}^{d}) μ ∈ P ( T d ) and let p μ p_{\mu} p μ be as in clause 2, p μ ( x ) = ∫ Θ s ( x − w ) μ ( d w ) p_{\mu}(x)=\int\Theta_{s}(x-w)\,\mu(dw) p μ ( x ) = ∫ Θ s ( x − w ) μ ( d w ) ; in the notation of The Convolution of a Bounded Continuous Function with a Probability Measure on Euclidean Space: Continuity, Boundedness and Differentiation Under the Integral §continuous this is Θ s ∗ μ \Theta_{s}*\mu Θ s ∗ μ , with Θ s \Theta_{s} Θ s continuous and ∣ Θ s ∣ ≤ C s |\Theta_{s}|\le C_{s} ∣ Θ s ∣ ≤ C s , so p μ p_{\mu} p μ is continuous, hence Borel, and it is nonnegative by monotonicity of the integral. Let B ∈ B ( R d ) B\in\mathcal{B}(\mathbb{R}^{d}) B ∈ B ( R d ) with B ⊆ Q B\subseteq Q B ⊆ Q . Taking the constant 1 B ( x ) \mathbf{1}_{B}(x) 1 B ( x ) inside the integral defining p μ ( x ) p_{\mu}(x) p μ ( x ) (claim 1 of Linearity and Monotonicity of the Lebesgue Integral ), using Step 0(b) with the measures λ d \lambda_{d} λ d and μ \mu μ , then ( E + ) (\mathrm{E}^{+}) ( E + ) (with P = Θ s P=\Theta_{s} P = Θ s ) and (W) with κ = μ \kappa=\mu κ = μ ,
∫ 1 B p μ d λ d = ∫ ( ∫ 1 B ( x ) Θ s ( x − w ) d x ) μ ( d w ) = ∫ ( ∫ 1 B ( π ( z ) ) g s ( z − w ) d z ) μ ( d w ) = π # μ s ( B ) = S s μ ( B ) . \int\mathbf{1}_{B}\,p_{\mu}\,d\lambda_{d}=\int\Bigl(\int\mathbf{1}_{B}(x)\,\Theta_{s}(x-w)\,dx\Bigr)\mu(dw)=\int\Bigl(\int\mathbf{1}_{B}\bigl(\pi(z)\bigr)g_{s}(z-w)\,dz\Bigr)\mu(dw)=\pi_{\#}\mu_{s}(B)=S_{s}\mu(B). ∫ 1 B p μ d λ d = ∫ ( ∫ 1 B ( x ) Θ s ( x − w ) d x ) μ ( d w ) = ∫ ( ∫ 1 B ( π ( z ) ) g s ( z − w ) d z ) μ ( d w ) = π # μ s ( B ) = S s μ ( B ) .
For arbitrary A ∈ B ( R d ) A\in\mathcal{B}(\mathbb{R}^{d}) A ∈ B ( R d ) , as S s μ ∈ P ( T d ) S_{s}\mu\in\mathcal{P}(\mathbb{T}^{d}) S s μ ∈ P ( T d ) , Step 0(a) and the case B = A ∩ Q B=A\cap Q B = A ∩ Q give S s μ ( A ) = S s μ ( A ∩ Q ) = ∫ 1 A ( 1 Q p μ ) d λ d S_{s}\mu(A)=S_{s}\mu(A\cap Q)=\int\mathbf{1}_{A}(\mathbf{1}_{Q}p_{\mu})\,d\lambda_{d} S s μ ( A ) = S s μ ( A ∩ Q ) = ∫ 1 A ( 1 Q p μ ) d λ d . The function 1 Q p μ \mathbf{1}_{Q}p_{\mu} 1 Q p μ is Borel and nonnegative, so it is a density of S s μ S_{s}\mu S s μ with respect to λ d \lambda_{d} λ d .
Step 7 (Clause 2). Let μ \mu μ and p μ p_{\mu} p μ be as in Step 6.
Regularity. Θ s \Theta_{s} Θ s is of class C 2 C^{2} C 2 , hence of class C 1 C^{1} C 1 (claim 2 of Euclidean Space is Open in Itself, and C k C^k C k Maps are Continuous ), with ∣ Θ s ∣ ≤ C s |\Theta_{s}|\le C_{s} ∣ Θ s ∣ ≤ C s and ∣ ∂ i Θ s ∣ ≤ C s |\partial_{i}\Theta_{s}|\le C_{s} ∣ ∂ i Θ s ∣ ≤ C s . By The Convolution of a Bounded Continuous Function with a Probability Measure on Euclidean Space: Continuity, Boundedness and Differentiation Under the Integral §derivative with G = Θ s G=\Theta_{s} G = Θ s and C = C s C=C_{s} C = C s , the function p μ = Θ s ∗ μ p_{\mu}=\Theta_{s}*\mu p μ = Θ s ∗ μ is of class C 1 C^{1} C 1 with ∂ i p μ ( y ) = ∫ ∂ i Θ s ( y − x ) μ ( d x ) \partial_{i}p_{\mu}(y)=\int\partial_{i}\Theta_{s}(y-x)\,\mu(dx) ∂ i p μ ( y ) = ∫ ∂ i Θ s ( y − x ) μ ( d x ) for every y y y and i ∈ [ d ] i\in[d] i ∈ [ d ] . By clause 2 of C^k Maps on a Euclidean Open Set , each ∂ i Θ s \partial_{i}\Theta_{s} ∂ i Θ s is of class C 1 C^{1} C 1 , and ∣ ∂ i Θ s ∣ ≤ C s |\partial_{i}\Theta_{s}|\le C_{s} ∣ ∂ i Θ s ∣ ≤ C s , ∣ ∂ j ∂ i Θ s ∣ ≤ C s |\partial_{j}\partial_{i}\Theta_{s}|\le C_{s} ∣ ∂ j ∂ i Θ s ∣ ≤ C s ; so the same clause of The Convolution of a Bounded Continuous Function with a Probability Measure on Euclidean Space: Continuity, Boundedness and Differentiation Under the Integral with G = ∂ i Θ s G=\partial_{i}\Theta_{s} G = ∂ i Θ s shows that ∂ i p μ = ( ∂ i Θ s ) ∗ μ \partial_{i}p_{\mu}=(\partial_{i}\Theta_{s})*\mu ∂ i p μ = ( ∂ i Θ s ) ∗ μ is of class C 1 C^{1} C 1 . By clause 2 of C^k Maps on a Euclidean Open Set with k = 1 k=1 k = 1 , p μ p_{\mu} p μ is of class C 2 C^{2} C 2 on R d \mathbb{R}^{d} R d .
Periodicity and bounds. For k ∈ Z d k\in\mathbb{Z}^{d} k ∈ Z d , p μ ( y + k ) = ∫ Θ s ( ( y − x ) + k ) μ ( d x ) = p μ ( y ) p_{\mu}(y+k)=\int\Theta_{s}((y-x)+k)\,\mu(dx)=p_{\mu}(y) p μ ( y + k ) = ∫ Θ s (( y − x ) + k ) μ ( d x ) = p μ ( y ) by the periodicity of Θ s \Theta_{s} Θ s ; so p μ ∈ C p e r 2 p_{\mu}\in C^{2}_{\mathrm{per}} p μ ∈ C per 2 . Since c s ≤ Θ s ( y − x ) ≤ C s c_{s}\le\Theta_{s}(y-x)\le C_{s} c s ≤ Θ s ( y − x ) ≤ C s for all x x x , the integrable case of Linearity and Monotonicity of the Lebesgue Integral (claim 2), with constants integrable against the probability measure μ \mu μ with integrals c s c_{s} c s and C s C_{s} C s , gives c s ≤ p μ ( y ) ≤ C s c_{s}\le p_{\mu}(y)\le C_{s} c s ≤ p μ ( y ) ≤ C s for every y y y .
Density. By Step 6, 1 Q p μ \mathbf{1}_{Q}p_{\mu} 1 Q p μ is a density of S s μ S_{s}\mu S s μ with respect to λ d \lambda_{d} λ d .
Score. By the three preceding paragraphs, p μ : R d → R p_{\mu}:\mathbb{R}^{d}\to\mathbb{R} p μ : R d → R is Z d \mathbb{Z}^{d} Z d -periodic and of class C 2 C^{2} C 2 on R d \mathbb{R}^{d} R d , satisfies p μ ( y ) ≥ c s > 0 p_{\mu}(y)\ge c_{s}>0 p μ ( y ) ≥ c s > 0 for every y y y , and 1 Q p μ \mathbf{1}_{Q}p_{\mu} 1 Q p μ is a density of S s μ S_{s}\mu S s μ with respect to λ d \lambda_{d} λ d ; that is, p μ p_{\mu} p μ is a function with every property required of p p p in Heat Smoothing on the Torus: Distance to the Identity, Contraction, a Smooth Positive Periodic Density, Finite Entropy, and Finite Fisher Information §density for this μ \mu μ . Clause 5 of that lemma, Heat Smoothing on the Torus: Distance to the Identity, Contraction, a Smooth Positive Periodic Density, Finite Entropy, and Finite Fisher Information §score , is stated for p p p as in clause 3, and clause 3 only asserts the existence of such functions without singling one out, so clause 5 applies to every such function, since any two functions satisfying clause 3 for the same μ \mu μ have products with 1 Q \mathbf{1}_{Q} 1 Q that are densities of S s μ S_{s}\mu S s μ with respect to λ d \lambda_{d} λ d , so they agree λ d \lambda_{d} λ d -almost everywhere on Q Q Q by The Radon-Nikodym Theorem for a Finite Measure and a Sigma-Finite Measure, and Uniqueness of Densities §uniqueness , hence S s μ S_{s}\mu S s μ -almost everywhere, S s μ S_{s}\mu S s μ being carried by Q Q Q with a density; so the class of p − 1 ∇ p p^{-1}\nabla p p − 1 ∇ p is the same for all of them; applied with p = p μ p=p_{\mu} p = p μ it gives S s μ ∈ P I ( T d ) S_{s}\mu\in\mathcal{P}^{\mathcal{I}}(\mathbb{T}^{d}) S s μ ∈ P I ( T d ) and that ξ S s μ \xi_{S_{s}\mu} ξ S s μ is the class of y ↦ p μ ( y ) − 1 ∇ p μ ( y ) y\mapsto p_{\mu}(y)^{-1}\nabla p_{\mu}(y) y ↦ p μ ( y ) − 1 ∇ p μ ( y ) . This proves clause 2.
Step 8 (A Lipschitz bound from bounded partial derivatives). Let G : R d → R G:\mathbb{R}^{d}\to\mathbb{R} G : R d → R be of class C 1 C^{1} C 1 on R d \mathbb{R}^{d} R d and let C ≥ 0 C\ge0 C ≥ 0 be real with ∣ ∂ k G ( z ) ∣ ≤ C |\partial_{k}G(z)|\le C ∣ ∂ k G ( z ) ∣ ≤ C for all z ∈ R d z\in\mathbb{R}^{d} z ∈ R d and k ∈ [ d ] k\in[d] k ∈ [ d ] . We show ∣ G ( a ) − G ( b ) ∣ ≤ d C ∥ a − b ∥ |G(a)-G(b)|\le d\,C\,\lVert a-b\rVert ∣ G ( a ) − G ( b ) ∣ ≤ d C ∥ a − b ∥ for all a , b ∈ R d a,b\in\mathbb{R}^{d} a , b ∈ R d .
Fix a , b a,b a , b . For k ∈ { 0 , 1 , … , d } k\in\{0,1,\dots,d\} k ∈ { 0 , 1 , … , d } let z ( k ) z^{(k)} z ( k ) be the point whose i i i th coordinate is b i b_{i} b i for i ≤ k i\le k i ≤ k and a i a_{i} a i for i > k i>k i > k ; so z ( 0 ) = a z^{(0)}=a z ( 0 ) = a , z ( d ) = b z^{(d)}=b z ( d ) = b , and G ( b ) − G ( a ) = ∑ k = 1 d ( G ( z ( k ) ) − G ( z ( k − 1 ) ) ) G(b)-G(a)=\sum_{k=1}^{d}\bigl(G(z^{(k)})-G(z^{(k-1)})\bigr) G ( b ) − G ( a ) = ∑ k = 1 d ( G ( z ( k ) ) − G ( z ( k − 1 ) ) ) . Fix k ∈ [ d ] k\in[d] k ∈ [ d ] . The points z ( k − 1 ) z^{(k-1)} z ( k − 1 ) and z ( k ) z^{(k)} z ( k ) differ at most in the k k k th coordinate, which is a k a_{k} a k , respectively b k b_{k} b k . If a k = b k a_{k}=b_{k} a k = b k , the k k k th term is 0 0 0 . Otherwise, for real t t t let z [ t ] z[t] z [ t ] be the point with k k k th coordinate t t t and the other coordinates those of z ( k − 1 ) z^{(k-1)} z ( k − 1 ) , so that z [ a k ] = z ( k − 1 ) z[a_{k}]=z^{(k-1)} z [ a k ] = z ( k − 1 ) and z [ b k ] = z ( k ) z[b_{k}]=z^{(k)} z [ b k ] = z ( k ) ; let p 0 = min { a k , b k } − 1 p_{0}=\min\{a_{k},b_{k}\}-1 p 0 = min { a k , b k } − 1 and q 0 = max { a k , b k } + 1 q_{0}=\max\{a_{k},b_{k}\}+1 q 0 = max { a k , b k } + 1 , and let φ : ( p 0 , q 0 ) → R \varphi:(p_{0},q_{0})\to\mathbb{R} φ : ( p 0 , q 0 ) → R be φ ( t ) = G ( z [ t ] ) \varphi(t)=G(z[t]) φ ( t ) = G ( z [ t ]) .
φ \varphi φ is differentiable on ( p 0 , q 0 ) (p_{0},q_{0}) ( p 0 , q 0 ) with φ ′ ( t ) = ∂ k G ( z [ t ] ) \varphi'(t)=\partial_{k}G(z[t]) φ ′ ( t ) = ∂ k G ( z [ t ]) . Fix t ∈ ( p 0 , q 0 ) t\in(p_{0},q_{0}) t ∈ ( p 0 , q 0 ) and put c = z [ t ] c=z[t] c = z [ t ] ; for real σ \sigma σ , the point c [ σ ] c[\sigma] c [ σ ] of Slice Function and the Partial Derivative (the k k k th coordinate of c c c replaced by σ \sigma σ ) is z [ σ ] z[\sigma] z [ σ ] . The domain being all of R d \mathbb{R}^{d} R d , the radius 1 1 1 is admissible in claim 1 of that lemma, and the slice function of G G G at c c c in the k k k th variable is σ ↦ G ( z [ σ ] ) \sigma\mapsto G(z[\sigma]) σ ↦ G ( z [ σ ]) on the interval with endpoints t − 1 t-1 t − 1 and t + 1 t+1 t + 1 . The partial derivative ∂ k G ( c ) \partial_{k}G(c) ∂ k G ( c ) exists, G G G being of class C 1 C^{1} C 1 , so by claim 2 there the slice function is differentiable at t t t with derivative ∂ k G ( c ) \partial_{k}G(c) ∂ k G ( c ) . The slice function and φ \varphi φ agree at every t + h t+h t + h with ∣ h ∣ < 1 |h|<1 ∣ h ∣ < 1 and t + h ∈ ( p 0 , q 0 ) t+h\in(p_{0},q_{0}) t + h ∈ ( p 0 , q 0 ) ; so, given ε > 0 \varepsilon>0 ε > 0 , the minimum of 1 1 1 and a δ \delta δ witnessing the condition of Derivative at an Interior Point for the slice function at t t t witnesses it for φ \varphi φ at t t t , with the same value ∂ k G ( z [ t ] ) \partial_{k}G(z[t]) ∂ k G ( z [ t ]) .
By Mean Value Theorem on an Open Interval , applied to φ \varphi φ and the two points min { a k , b k } < max { a k , b k } \min\{a_{k},b_{k}\}<\max\{a_{k},b_{k}\} min { a k , b k } < max { a k , b k } of ( p 0 , q 0 ) (p_{0},q_{0}) ( p 0 , q 0 ) , there is t ∗ t_{*} t ∗ between them with φ ( b k ) − φ ( a k ) = φ ′ ( t ∗ ) ( b k − a k ) \varphi(b_{k})-\varphi(a_{k})=\varphi'(t_{*})(b_{k}-a_{k}) φ ( b k ) − φ ( a k ) = φ ′ ( t ∗ ) ( b k − a k ) (the identity for the ordered pair being symmetric under exchanging a k a_{k} a k and b k b_{k} b k ). Hence
∣ G ( z ( k ) ) − G ( z ( k − 1 ) ) ∣ = ∣ ∂ k G ( z [ t ∗ ] ) ∣ ∣ b k − a k ∣ ≤ C ∣ b k − a k ∣ ≤ C ∥ b − a ∥ , \bigl|G(z^{(k)})-G(z^{(k-1)})\bigr|=\bigl|\partial_{k}G(z[t_{*}])\bigr|\,|b_{k}-a_{k}|\le C\,|b_{k}-a_{k}|\le C\,\lVert b-a\rVert, G ( z ( k ) ) − G ( z ( k − 1 ) ) = ∂ k G ( z [ t ∗ ]) ∣ b k − a k ∣ ≤ C ∣ b k − a k ∣ ≤ C ∥ b − a ∥ ,
by claim 4 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n applied to b − a b-a b − a , whose k k k th coordinate is b k − a k b_{k}-a_{k} b k − a k ; this bound also holds when a k = b k a_{k}=b_{k} a k = b k . Summing over k ∈ [ d ] k\in[d] k ∈ [ d ] with the triangle inequality for the absolute value, and using ∥ b − a ∥ = ∥ a − b ∥ \lVert b-a\rVert=\lVert a-b\rVert ∥ b − a ∥ = ∥ a − b ∥ (claim 5 of the same lemma), ∣ G ( a ) − G ( b ) ∣ ≤ d C ∥ a − b ∥ |G(a)-G(b)|\le d\,C\,\lVert a-b\rVert ∣ G ( a ) − G ( b ) ∣ ≤ d C ∥ a − b ∥ .
Step 9 (Clause 3). Put L s = d 2 C s L_{s}=d^{2}C_{s} L s = d 2 C s , a real number with L s ≥ d C s ≥ 0 L_{s}\ge d\,C_{s}\ge0 L s ≥ d C s ≥ 0 as 1 ≤ d 1\le d 1 ≤ d . Let μ , ν ∈ P ( T d ) \mu,\nu\in\mathcal{P}(\mathbb{T}^{d}) μ , ν ∈ P ( T d ) and y ∈ R d y\in\mathbb{R}^{d} y ∈ R d , and write W = W T ( μ , ν ) ≥ 0 W=W_{\mathbb{T}}(\mu,\nu)\ge0 W = W T ( μ , ν ) ≥ 0 .
For i ∈ [ d ] i\in[d] i ∈ [ d ] let f 0 ( x ) = Θ s ( y − x ) f_{0}(x)=\Theta_{s}(y-x) f 0 ( x ) = Θ s ( y − x ) and f i ( x ) = ∂ i Θ s ( y − x ) f_{i}(x)=\partial_{i}\Theta_{s}(y-x) f i ( x ) = ∂ i Θ s ( y − x ) . Each is Z d \mathbb{Z}^{d} Z d -periodic: Θ s \Theta_{s} Θ s is periodic by clause 1 and ∂ i Θ s \partial_{i}\Theta_{s} ∂ i Θ s by Elementary Properties of Lattice-Periodic Functions §derivative , and y − ( x + k ) = ( y − x ) + ( − k ) y-(x+k)=(y-x)+(-k) y − ( x + k ) = ( y − x ) + ( − k ) with − k ∈ Z d -k\in\mathbb{Z}^{d} − k ∈ Z d . Apply Step 8 to G = Θ s G=\Theta_{s} G = Θ s , which is of class C 1 C^{1} C 1 (Step 7) with ∣ ∂ j Θ s ∣ ≤ C s |\partial_{j}\Theta_{s}|\le C_{s} ∣ ∂ j Θ s ∣ ≤ C s , and to G = ∂ i Θ s G=\partial_{i}\Theta_{s} G = ∂ i Θ s , which is of class C 1 C^{1} C 1 (Step 7) with ∣ ∂ j ∂ i Θ s ∣ ≤ C s |\partial_{j}\partial_{i}\Theta_{s}|\le C_{s} ∣ ∂ j ∂ i Θ s ∣ ≤ C s , in each case with C = C s C=C_{s} C = C s : for x , x ′ ∈ R d x,x'\in\mathbb{R}^{d} x , x ′ ∈ R d and l ∈ { 0 } ∪ [ d ] l\in\{0\}\cup[d] l ∈ { 0 } ∪ [ d ] ,
∣ f l ( x ) − f l ( x ′ ) ∣ ≤ d C s ∥ ( y − x ) − ( y − x ′ ) ∥ = d C s ∥ x − x ′ ∥ , |f_{l}(x)-f_{l}(x')|\le d\,C_{s}\,\lVert(y-x)-(y-x')\rVert=d\,C_{s}\,\lVert x-x'\rVert, ∣ f l ( x ) − f l ( x ′ ) ∣ ≤ d C s ∥( y − x ) − ( y − x ′ )∥ = d C s ∥ x − x ′ ∥ ,
using claim 5 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n . As d E ( x , x ′ ) = ∥ x − x ′ ∥ d_{E}(x,x')=\lVert x-x'\rVert d E ( x , x ′ ) = ∥ x − x ′ ∥ (claim 2 there), each f l f_{l} f l is Lipschitz with constant d C s dC_{s} d C s as a map from ( R d , d E ) (\mathbb{R}^{d},d_{E}) ( R d , d E ) to R \mathbb{R} R , hence continuous by A Lipschitz Map is Uniformly Continuous , so f l ∈ C p e r f_{l}\in C_{\mathrm{per}} f l ∈ C per . By The Torus Wasserstein Distance: Comparison with the Euclidean Distance, Wrapping, and Integrals of Periodic Functions §integrals with L = d C s L=dC_{s} L = d C s ,
∣ ∫ f l d μ − ∫ f l d ν ∣ ≤ d C s W ( l ∈ { 0 } ∪ [ d ] ) . \Bigl|\int f_{l}\,d\mu-\int f_{l}\,d\nu\Bigr|\le d\,C_{s}\,W\qquad(l\in\{0\}\cup[d]). ∫ f l d μ − ∫ f l d ν ≤ d C s W ( l ∈ { 0 } ∪ [ d ]) .
By clause 2 (Step 7), ∫ f 0 d μ = p μ ( y ) \int f_{0}\,d\mu=p_{\mu}(y) ∫ f 0 d μ = p μ ( y ) and ∫ f i d μ = ∂ i p μ ( y ) \int f_{i}\,d\mu=\partial_{i}p_{\mu}(y) ∫ f i d μ = ∂ i p μ ( y ) , and likewise for ν \nu ν . So ∣ p μ ( y ) − p ν ( y ) ∣ ≤ d C s W ≤ L s W |p_{\mu}(y)-p_{\nu}(y)|\le dC_{s}W\le L_{s}W ∣ p μ ( y ) − p ν ( y ) ∣ ≤ d C s W ≤ L s W . For the gradients, put e i = ∂ i p μ ( y ) − ∂ i p ν ( y ) e_{i}=\partial_{i}p_{\mu}(y)-\partial_{i}p_{\nu}(y) e i = ∂ i p μ ( y ) − ∂ i p ν ( y ) , the i i i th coordinate of ∇ p μ ( y ) − ∇ p ν ( y ) \nabla p_{\mu}(y)-\nabla p_{\nu}(y) ∇ p μ ( y ) − ∇ p ν ( y ) (Optimal Transport on the Flat Torus: Standing Notation §calculus ); then ∣ e i ∣ ≤ d C s W |e_{i}|\le dC_{s}W ∣ e i ∣ ≤ d C s W , so e i 2 ≤ ( d C s W ) 2 e_{i}^{2}\le(dC_{s}W)^{2} e i 2 ≤ ( d C s W ) 2 by claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field , and by claim 1 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n ,
∥ ∇ p μ ( y ) − ∇ p ν ( y ) ∥ 2 = ∑ i = 1 d e i 2 ≤ d ( d C s W ) 2 ≤ ( d 2 C s W ) 2 = ( L s W ) 2 , \lVert\nabla p_{\mu}(y)-\nabla p_{\nu}(y)\rVert^{2}=\sum_{i=1}^{d}e_{i}^{2}\le d\,(dC_{s}W)^{2}\le(d^{2}C_{s}W)^{2}=(L_{s}W)^{2}, ∥ ∇ p μ ( y ) − ∇ p ν ( y ) ∥ 2 = i = 1 ∑ d e i 2 ≤ d ( d C s W ) 2 ≤ ( d 2 C s W ) 2 = ( L s W ) 2 ,
as d ≤ d 2 d\le d^{2} d ≤ d 2 . Both ∥ ∇ p μ ( y ) − ∇ p ν ( y ) ∥ \lVert\nabla p_{\mu}(y)-\nabla p_{\nu}(y)\rVert ∥ ∇ p μ ( y ) − ∇ p ν ( y )∥ and L s W L_{s}W L s W are nonnegative, so claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field gives ∥ ∇ p μ ( y ) − ∇ p ν ( y ) ∥ ≤ L s W \lVert\nabla p_{\mu}(y)-\nabla p_{\nu}(y)\rVert\le L_{s}W ∥ ∇ p μ ( y ) − ∇ p ν ( y )∥ ≤ L s W . This proves clause 3. ■ \blacksquare ■