Throughout, Φ = Φ ν \Phi=\Phi_{\nu} Φ = Φ ν , ∥ ⋅ ∥ \lVert\cdot\rVert ∥ ⋅ ∥ is the Euclidean norm and x ⋅ y x\cdot y x ⋅ y the dot product of R d \mathbb{R}^{d} R d . We use the following facts repeatedly.
(F1) Wrapped displacements. For z ∈ R d z\in\mathbb{R}^{d} z ∈ R d and k ∈ Z d k\in\mathbb{Z}^{d} k ∈ Z d one has z − ϖ ( z ) ∈ Z d z-\varpi(z)\in\mathbb{Z}^{d} z − ϖ ( z ) ∈ Z d and ∥ ϖ ( z ) ∥ 2 ≤ d / 4 \lVert\varpi(z)\rVert^{2}\le d/4 ∥ ϖ ( z ) ∥ 2 ≤ d /4 by The Flat Torus Distance: Minimality of the Wrapped Displacement, Periodicity, the Metric on the Unit Cell and the Lipschitz Bound §range , and ∥ ϖ ( z ) ∥ ≤ ∥ z − k ∥ \lVert\varpi(z)\rVert\le\lVert z-k\rVert ∥ ϖ ( z )∥ ≤ ∥ z − k ∥ by The Flat Torus Distance: Minimality of the Wrapped Displacement, Periodicity, the Metric on the Unit Cell and the Lipschitz Bound §minimal . For x , y ∈ R d x,y\in\mathbb{R}^{d} x , y ∈ R d , d T ( x , y ) = ∥ ϖ ( y − x ) ∥ d_{\mathbb{T}}(x,y)=\lVert\varpi(y-x)\rVert d T ( x , y ) = ∥ ϖ ( y − x )∥ by the definition of the flat torus distance . Sums and differences of points of Z d \mathbb{Z}^{d} Z d lie in Z d \mathbb{Z}^{d} Z d , their coordinates being sums and differences of integers (Lattice-Periodic Functions and the Periodic Function Classes §lattice ). The map ϖ \varpi ϖ is Borel by The Flat Torus Distance: Minimality of the Wrapped Displacement, Periodicity, the Metric on the Unit Cell and the Lipschitz Bound §lipschitz , and a difference of two Borel maps into R d \mathbb{R}^{d} R d is Borel, componentwise, by claim 2 of The Borel Sigma-Algebra of a Euclidean Space as a Product, and Measurability of Projections, Sequentially Continuous Maps, and Open and Closed Sets and claim 2 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions (exactly as in Optimal Maps on the Torus, Uniquely Mapped Pairs and the Displacement Field of a Map §displacement ); likewise for sums.
(F2) Pairings and push-forwards. For m ∈ N m\in\mathbb{N} m ∈ N with 1 ≤ m 1\le m 1 ≤ m , Borel maps u , v : R m → R d u,v:\mathbb{R}^{m}\to\mathbb{R}^{d} u , v : R m → R d and λ ∈ P ( R m ) \lambda\in\mathcal{P}(\mathbb{R}^{m}) λ ∈ P ( R m ) , the pairing ( u , v ) (u,v) ( u , v ) is Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §pairing and satisfies p r 1 ∘ ( u , v ) = u \mathrm{pr}_{1}\circ(u,v)=u pr 1 ∘ ( u , v ) = u , p r 2 ∘ ( u , v ) = v \mathrm{pr}_{2}\circ(u,v)=v pr 2 ∘ ( u , v ) = v by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections . Consequently, for B ∈ B ( R d ) B\in\mathcal{B}(\mathbb{R}^{d}) B ∈ B ( R d ) , u # λ ( B ) = λ ( ( u , v ) − 1 ( p r 1 − 1 ( B ) ) ) = ( p r 1 ) # ( ( u , v ) # λ ) ( B ) u_{\#}\lambda(B)=\lambda\bigl((u,v)^{-1}(\mathrm{pr}_{1}^{-1}(B))\bigr)=(\mathrm{pr}_{1})_{\#}\bigl((u,v)_{\#}\lambda\bigr)(B) u # λ ( B ) = λ ( ( u , v ) − 1 ( pr 1 − 1 ( B )) ) = ( pr 1 ) # ( ( u , v ) # λ ) ( B ) , and similarly v # λ = ( p r 2 ) # ( ( u , v ) # λ ) v_{\#}\lambda=(\mathrm{pr}_{2})_{\#}((u,v)_{\#}\lambda) v # λ = ( pr 2 ) # (( u , v ) # λ ) , by the definition of the push-forward . The change-of-variables formula ∫ g d ( F # λ ) = ∫ g ∘ F d λ \int g\,d(F_{\#}\lambda)=\int g\circ F\,d\lambda ∫ g d ( F # λ ) = ∫ g ∘ F d λ is used for Borel F F F and bounded Borel g g g ; bounded Borel functions are integrable against every probability measure by Probability Measures on Euclidean Space and Random Vectors: Standing Notation §measures , and integrals of such functions are combined by linearity and monotonicity, claim 2 of Linearity and Monotonicity of the Lebesgue Integral . Norms, squared norms and dot products of Borel maps into R d \mathbb{R}^{d} R d are Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions , and ∣ x ⋅ y ∣ ≤ ∥ x ∥ ∥ y ∥ |x\cdot y|\le\lVert x\rVert\,\lVert y\rVert ∣ x ⋅ y ∣ ≤ ∥ x ∥ ∥ y ∥ by the same clause.
(F3) The distance as an infimum. For μ 1 , μ 2 ∈ P ( T d ) \mu_{1},\mu_{2}\in\mathcal{P}(\mathbb{T}^{d}) μ 1 , μ 2 ∈ P ( T d ) and κ ∈ Π ( μ 1 , μ 2 ) \kappa\in\Pi(\mu_{1},\mu_{2}) κ ∈ Π ( μ 1 , μ 2 ) , W T ( μ 1 , μ 2 ) 2 ≤ I T ( κ ) W_{\mathbb{T}}(\mu_{1},\mu_{2})^{2}\le I_{\mathbb{T}}(\kappa) W T ( μ 1 , μ 2 ) 2 ≤ I T ( κ ) , since W T ( μ 1 , μ 2 ) 2 W_{\mathbb{T}}(\mu_{1},\mu_{2})^{2} W T ( μ 1 , μ 2 ) 2 is the infimum of the torus costs by Probability Measures on the Flat Torus, the Torus Cost of a Coupling, the Torus Wasserstein Distance and Optimal Couplings §distance ; and equality holds when κ \kappa κ is optimal, by Probability Measures on the Flat Torus, the Torus Cost of a Coupling, the Torus Wasserstein Distance and Optimal Couplings §optimal .
Proof of clause 1. Let μ ∈ P ( T d ) \mu\in\mathcal{P}(\mathbb{T}^{d}) μ ∈ P ( T d ) , let T T T be an optimal map from μ \mu μ to ν \nu ν , and fix the Borel map v T v_{T} v T of Optimal Maps on the Torus, Uniquely Mapped Pairs and the Displacement Field of a Map §displacement as the representative of its class; ∥ v T ( x ) ∥ ≤ d / 2 \lVert v_{T}(x)\rVert\le\sqrt d/2 ∥ v T ( x )∥ ≤ d /2 for every x x x by (F1).
Step 1.1 (the value at μ \mu μ ). By Optimal Maps on the Torus, Uniquely Mapped Pairs and the Displacement Field of a Map §map , T # μ = ν T_{\#}\mu=\nu T # μ = ν and ( i d , T ) # μ ∈ Π ( μ , ν ) (\mathrm{id},T)_{\#}\mu\in\Pi(\mu,\nu) ( id , T ) # μ ∈ Π ( μ , ν ) is optimal, so by (F3), the definition of the torus cost , (F2) and (F1),
W T ( μ , ν ) 2 = I T ( ( i d , T ) # μ ) = ∫ R d d T ( x , T ( x ) ) 2 μ ( d x ) = ∫ R d ∥ ϖ ( T ( x ) − x ) ∥ 2 μ ( d x ) = ∫ R d ∥ v T ∥ 2 d μ = ∥ v T ∥ μ 2 , W_{\mathbb{T}}(\mu,\nu)^{2}=I_{\mathbb{T}}\bigl((\mathrm{id},T)_{\#}\mu\bigr)=\int_{\mathbb{R}^{d}}d_{\mathbb{T}}\bigl(x,T(x)\bigr)^{2}\,\mu(dx)=\int_{\mathbb{R}^{d}}\bigl\lVert\varpi\bigl(T(x)-x\bigr)\bigr\rVert^{2}\,\mu(dx)=\int_{\mathbb{R}^{d}}\lVert v_{T}\rVert^{2}\,d\mu=\lVert v_{T}\rVert_{\mu}^{2}, W T ( μ , ν ) 2 = I T ( ( id , T ) # μ ) = ∫ R d d T ( x , T ( x ) ) 2 μ ( d x ) = ∫ R d ϖ ( T ( x ) − x ) 2 μ ( d x ) = ∫ R d ∥ v T ∥ 2 d μ = ∥ v T ∥ μ 2 ,
the last equality being the definition of the norm in Optimal Transport on the Flat Torus: Standing Notation §fields . Hence Φ ( μ ) = 1 2 ∥ v T ∥ μ 2 \Phi(\mu)=\tfrac12\lVert v_{T}\rVert_{\mu}^{2} Φ ( μ ) = 2 1 ∥ v T ∥ μ 2 .
Step 1.2 (a competitor coupling). Let μ ′ ∈ P ( T d ) \mu'\in\mathcal{P}(\mathbb{T}^{d}) μ ′ ∈ P ( T d ) and γ ∈ Π ( μ , μ ′ ) \gamma\in\Pi(\mu,\mu') γ ∈ Π ( μ , μ ′ ) . The map S = ( p r 2 , T ∘ p r 1 ) : R d + d → R d + d S=(\mathrm{pr}_{2},T\circ\mathrm{pr}_{1}):\mathbb{R}^{d+d}\to\mathbb{R}^{d+d} S = ( pr 2 , T ∘ pr 1 ) : R d + d → R d + d is Borel by (F2), T ∘ p r 1 T\circ\mathrm{pr}_{1} T ∘ pr 1 being a composite of Borel maps (Probability Measures on Euclidean Space and Random Vectors: Standing Notation §borel-maps ). Put σ = S # γ \sigma=S_{\#}\gamma σ = S # γ . By (F2), ( p r 1 ) # σ = ( p r 2 ) # γ = μ ′ (\mathrm{pr}_{1})_{\#}\sigma=(\mathrm{pr}_{2})_{\#}\gamma=\mu' ( pr 1 ) # σ = ( pr 2 ) # γ = μ ′ , and for B ∈ B ( R d ) B\in\mathcal{B}(\mathbb{R}^{d}) B ∈ B ( R d ) , ( p r 2 ) # σ ( B ) = γ ( p r 1 − 1 ( T − 1 ( B ) ) ) = μ ( T − 1 ( B ) ) = ν ( B ) (\mathrm{pr}_{2})_{\#}\sigma(B)=\gamma\bigl(\mathrm{pr}_{1}^{-1}(T^{-1}(B))\bigr)=\mu(T^{-1}(B))=\nu(B) ( pr 2 ) # σ ( B ) = γ ( pr 1 − 1 ( T − 1 ( B )) ) = μ ( T − 1 ( B )) = ν ( B ) . Thus σ ∈ Π ( μ ′ , ν ) \sigma\in\Pi(\mu',\nu) σ ∈ Π ( μ ′ , ν ) , and by (F3) and the change-of-variables formula,
2 Φ ( μ ′ ) = W T ( μ ′ , ν ) 2 ≤ I T ( σ ) = ∫ R d + d d T ( p r 2 ( w ) , T ( p r 1 ( w ) ) ) 2 γ ( d w ) . 2\Phi(\mu')=W_{\mathbb{T}}(\mu',\nu)^{2}\le I_{\mathbb{T}}(\sigma)=\int_{\mathbb{R}^{d+d}}d_{\mathbb{T}}\bigl(\mathrm{pr}_{2}(w),T(\mathrm{pr}_{1}(w))\bigr)^{2}\,\gamma(dw). 2Φ ( μ ′ ) = W T ( μ ′ , ν ) 2 ≤ I T ( σ ) = ∫ R d + d d T ( pr 2 ( w ) , T ( pr 1 ( w )) ) 2 γ ( d w ) .
Step 1.3 (pointwise bound). Let w ∈ R d + d w\in\mathbb{R}^{d+d} w ∈ R d + d and write x = p r 1 ( w ) x=\mathrm{pr}_{1}(w) x = pr 1 ( w ) , x ′ = p r 2 ( w ) x'=\mathrm{pr}_{2}(w) x ′ = pr 2 ( w ) . By (F1) the point k = ( T ( x ) − x − v T ( x ) ) − ( x ′ − x − ϖ ( x ′ − x ) ) k=\bigl(T(x)-x-v_{T}(x)\bigr)-\bigl(x'-x-\varpi(x'-x)\bigr) k = ( T ( x ) − x − v T ( x ) ) − ( x ′ − x − ϖ ( x ′ − x ) ) lies in Z d \mathbb{Z}^{d} Z d , since v T ( x ) = ϖ ( T ( x ) − x ) v_{T}(x)=\varpi(T(x)-x) v T ( x ) = ϖ ( T ( x ) − x ) ; and T ( x ) − x ′ − k = v T ( x ) − ϖ ( x ′ − x ) T(x)-x'-k=v_{T}(x)-\varpi(x'-x) T ( x ) − x ′ − k = v T ( x ) − ϖ ( x ′ − x ) . Hence, by (F1) and bilinearity of the dot product,
d T ( x ′ , T ( x ) ) 2 = ∥ ϖ ( T ( x ) − x ′ ) ∥ 2 ≤ ∥ v T ( x ) − ϖ ( x ′ − x ) ∥ 2 = ∥ v T ( x ) ∥ 2 − 2 v T ( x ) ⋅ ϖ ( x ′ − x ) + d T ( x , x ′ ) 2 . d_{\mathbb{T}}\bigl(x',T(x)\bigr)^{2}=\bigl\lVert\varpi\bigl(T(x)-x'\bigr)\bigr\rVert^{2}\le\bigl\lVert v_{T}(x)-\varpi(x'-x)\bigr\rVert^{2}=\lVert v_{T}(x)\rVert^{2}-2\,v_{T}(x)\cdot\varpi(x'-x)+d_{\mathbb{T}}(x,x')^{2}. d T ( x ′ , T ( x ) ) 2 = ϖ ( T ( x ) − x ′ ) 2 ≤ v T ( x ) − ϖ ( x ′ − x ) 2 = ∥ v T ( x ) ∥ 2 − 2 v T ( x ) ⋅ ϖ ( x ′ − x ) + d T ( x , x ′ ) 2 .
Step 1.4 (integration). The three functions of w w w on the right are Borel by (F1) and (F2) and bounded by d / 4 d/4 d /4 in absolute value (using ∣ a ⋅ b ∣ ≤ ∥ a ∥ ∥ b ∥ |a\cdot b|\le\lVert a\rVert\,\lVert b\rVert ∣ a ⋅ b ∣ ≤ ∥ a ∥ ∥ b ∥ ), so they are integrable against γ \gamma γ . By (F2) and ( p r 1 ) # γ = μ (\mathrm{pr}_{1})_{\#}\gamma=\mu ( pr 1 ) # γ = μ , ∫ ∥ v T ∘ p r 1 ∥ 2 d γ = ∥ v T ∥ μ 2 = 2 Φ ( μ ) \int\lVert v_{T}\circ\mathrm{pr}_{1}\rVert^{2}\,d\gamma=\lVert v_{T}\rVert_{\mu}^{2}=2\Phi(\mu) ∫ ∥ v T ∘ pr 1 ∥ 2 d γ = ∥ v T ∥ μ 2 = 2Φ ( μ ) (Step 1.1); the integral of v T ( p r 1 ( w ) ) ⋅ ϖ ( p r 2 ( w ) − p r 1 ( w ) ) v_{T}(\mathrm{pr}_{1}(w))\cdot\varpi(\mathrm{pr}_{2}(w)-\mathrm{pr}_{1}(w)) v T ( pr 1 ( w )) ⋅ ϖ ( pr 2 ( w ) − pr 1 ( w )) is J T ( v T , γ ) \mathcal{J}_{\mathbb{T}}(v_{T},\gamma) J T ( v T , γ ) by The Torus Displacement Pairing of a Vector Field Along a Coupling §pairing ; and the integral of d T ( p r 1 ( w ) , p r 2 ( w ) ) 2 d_{\mathbb{T}}(\mathrm{pr}_{1}(w),\mathrm{pr}_{2}(w))^{2} d T ( pr 1 ( w ) , pr 2 ( w ) ) 2 is I T ( γ ) I_{\mathbb{T}}(\gamma) I T ( γ ) by Probability Measures on the Flat Torus, the Torus Cost of a Coupling, the Torus Wasserstein Distance and Optimal Couplings §cost . Integrating Step 1.3 against γ \gamma γ by monotonicity and linearity (F2) and using Step 1.2,
2 Φ ( μ ′ ) ≤ 2 Φ ( μ ) − 2 J T ( v T , γ ) + I T ( γ ) , 2\Phi(\mu')\le2\Phi(\mu)-2\,\mathcal{J}_{\mathbb{T}}(v_{T},\gamma)+I_{\mathbb{T}}(\gamma), 2Φ ( μ ′ ) ≤ 2Φ ( μ ) − 2 J T ( v T , γ ) + I T ( γ ) ,
and multiplying by 1 2 \tfrac12 2 1 gives clause 1.
Proof of clause 2. Let μ \mu μ be absolutely continuous and T T T an optimal map from μ \mu μ to ν \nu ν , with v T v_{T} v T as above; write W = W T ( μ , ν ) W=W_{\mathbb{T}}(\mu,\nu) W = W T ( μ , ν ) . The class of − v T -v_{T} − v T lies in L 2 ( μ ; R d ) L^{2}(\mu;\mathbb{R}^{d}) L 2 ( μ ; R d ) , and J T ( − v T , γ ) = − J T ( v T , γ ) \mathcal{J}_{\mathbb{T}}(-v_{T},\gamma)=-\mathcal{J}_{\mathbb{T}}(v_{T},\gamma) J T ( − v T , γ ) = − J T ( v T , γ ) for every μ ′ \mu' μ ′ and γ ∈ Π ( μ , μ ′ ) \gamma\in\Pi(\mu,\mu') γ ∈ Π ( μ , μ ′ ) , by The Torus Displacement Pairing: Linearity, the Cost Bound, and Vanishing of a Field with First-Order Small Pairings §linear with a = − 1 a=-1 a = − 1 , b = 0 b=0 b = 0 . So we must show: for every real ε > 0 \varepsilon>0 ε > 0 there is a real θ > 0 \theta>0 θ > 0 with
∣ Φ ( μ ′ ) − Φ ( μ ) + J T ( v T , γ ) ∣ ≤ ε I T ( γ ) ( ∗ ) \bigl|\Phi(\mu')-\Phi(\mu)+\mathcal{J}_{\mathbb{T}}(v_{T},\gamma)\bigr|\le\varepsilon\sqrt{I_{\mathbb{T}}(\gamma)}\qquad(\ast) Φ ( μ ′ ) − Φ ( μ ) + J T ( v T , γ ) ≤ ε I T ( γ ) ( ∗ )
for all μ ′ ∈ P ( T d ) \mu'\in\mathcal{P}(\mathbb{T}^{d}) μ ′ ∈ P ( T d ) and γ ∈ Π ( μ , μ ′ ) \gamma\in\Pi(\mu,\mu') γ ∈ Π ( μ , μ ′ ) with I T ( γ ) < θ 2 I_{\mathbb{T}}(\gamma)<\theta^{2} I T ( γ ) < θ 2 .
Step 2.1 (the lower estimate). We claim: for every real ε > 0 \varepsilon>0 ε > 0 there is a real θ 1 > 0 \theta_{1}>0 θ 1 > 0 such that Φ ( μ ′ ) − Φ ( μ ) + J T ( v T , γ ) ≥ − ε I T ( γ ) \Phi(\mu')-\Phi(\mu)+\mathcal{J}_{\mathbb{T}}(v_{T},\gamma)\ge-\varepsilon\sqrt{I_{\mathbb{T}}(\gamma)} Φ ( μ ′ ) − Φ ( μ ) + J T ( v T , γ ) ≥ − ε I T ( γ ) for all μ ′ ∈ P ( T d ) \mu'\in\mathcal{P}(\mathbb{T}^{d}) μ ′ ∈ P ( T d ) and γ ∈ Π ( μ , μ ′ ) \gamma\in\Pi(\mu,\mu') γ ∈ Π ( μ , μ ′ ) with I T ( γ ) < θ 1 2 I_{\mathbb{T}}(\gamma)<\theta_{1}^{2} I T ( γ ) < θ 1 2 . Suppose not, and fix ε > 0 \varepsilon>0 ε > 0 for which it fails. For each n ∈ N n\in\mathbb{N} n ∈ N , in this order: using the failure with θ 1 = 1 / ( n + 1 ) \theta_{1}=1/(n+1) θ 1 = 1/ ( n + 1 ) , choose μ n ′ ∈ P ( T d ) \mu'_{n}\in\mathcal{P}(\mathbb{T}^{d}) μ n ′ ∈ P ( T d ) and γ n ∈ Π ( μ , μ n ′ ) \gamma_{n}\in\Pi(\mu,\mu'_{n}) γ n ∈ Π ( μ , μ n ′ ) with
I T ( γ n ) < 1 ( n + 1 ) 2 , Φ ( μ n ′ ) − Φ ( μ ) + J T ( v T , γ n ) < − ε I T ( γ n ) ; ( 1 ) I_{\mathbb{T}}(\gamma_{n})<\frac{1}{(n+1)^{2}},\qquad\Phi(\mu'_{n})-\Phi(\mu)+\mathcal{J}_{\mathbb{T}}(v_{T},\gamma_{n})<-\varepsilon\sqrt{I_{\mathbb{T}}(\gamma_{n})};\qquad(1) I T ( γ n ) < ( n + 1 ) 2 1 , Φ ( μ n ′ ) − Φ ( μ ) + J T ( v T , γ n ) < − ε I T ( γ n ) ; ( 1 )
then choose an optimal σ n ′ ∈ Π ( μ n ′ , ν ) \sigma'_{n}\in\Pi(\mu'_{n},\nu) σ n ′ ∈ Π ( μ n ′ , ν ) by The Torus Wasserstein Space is a Sequentially Compact Metric Space in which Optimal Couplings Exist §optimal ; then choose a gluing Σ n ∈ P ( R 3 d ) \Sigma_{n}\in\mathcal{P}(\mathbb{R}^{3d}) Σ n ∈ P ( R 3 d ) of γ n \gamma_{n} γ n and σ n ′ \sigma'_{n} σ n ′ by Gluing Two Couplings over a Common Middle Marginal, and the Composite Coupling §glued . That lemma is stated in Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation , whose dimension d d d is the one of The Wasserstein Space and Its Lift to Square-Integrable Random Vectors: Standing Notation §data , an arbitrary natural number with 1 ≤ d 1\le d 1 ≤ d ; we take it to be the torus dimension d d d . Its hypotheses hold: μ , μ n ′ , ν ∈ P 2 ( R d ) \mu,\mu'_{n},\nu\in\mathcal{P}_{2}(\mathbb{R}^{d}) μ , μ n ′ , ν ∈ P 2 ( R d ) by The Torus Wasserstein Distance: Comparison with the Euclidean Distance, Wrapping, and Integrals of Periodic Functions §inclusion , and its couplings are those of Couplings of Two Probability Measures on Euclidean Space and Their Quadratic Cost §coupling (Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §dimensions ), which are also the couplings of Optimal Transport on the Flat Torus: Standing Notation §measures . Thus, with q 1 , q 2 , q 3 \mathrm{q}_{1},\mathrm{q}_{2},\mathrm{q}_{3} q 1 , q 2 , q 3 the coordinate maps of R 3 d \mathbb{R}^{3d} R 3 d , ( q 1 , q 2 ) # Σ n = γ n (\mathrm{q}_{1},\mathrm{q}_{2})_{\#}\Sigma_{n}=\gamma_{n} ( q 1 , q 2 ) # Σ n = γ n and ( q 2 , q 3 ) # Σ n = σ n ′ (\mathrm{q}_{2},\mathrm{q}_{3})_{\#}\Sigma_{n}=\sigma'_{n} ( q 2 , q 3 ) # Σ n = σ n ′ . Put s n = I T ( γ n ) s_{n}=\sqrt{I_{\mathbb{T}}(\gamma_{n})} s n = I T ( γ n ) , so 0 ≤ s n < 1 / ( n + 1 ) ≤ 1 0\le s_{n}<1/(n+1)\le1 0 ≤ s n < 1/ ( n + 1 ) ≤ 1 , and W n ′ = W T ( μ n ′ , ν ) W'_{n}=W_{\mathbb{T}}(\mu'_{n},\nu) W n ′ = W T ( μ n ′ , ν ) , so 2 Φ ( μ n ′ ) = W n ′ 2 = I T ( σ n ′ ) 2\Phi(\mu'_{n})=W_{n}'^{2}=I_{\mathbb{T}}(\sigma'_{n}) 2Φ ( μ n ′ ) = W n ′ 2 = I T ( σ n ′ ) by (F3).
Step 2.2 (the composite displacement). Define Borel maps R 3 d → R d \mathbb{R}^{3d}\to\mathbb{R}^{d} R 3 d → R d (by (F1) and (F2)) by
a n = ϖ ∘ ( q 2 − q 1 ) , b n = ϖ ∘ ( q 3 − q 2 ) , Z n = a n + b n . a_{n}=\varpi\circ(\mathrm{q}_{2}-\mathrm{q}_{1}),\qquad b_{n}=\varpi\circ(\mathrm{q}_{3}-\mathrm{q}_{2}),\qquad Z_{n}=a_{n}+b_{n}. a n = ϖ ∘ ( q 2 − q 1 ) , b n = ϖ ∘ ( q 3 − q 2 ) , Z n = a n + b n .
By (F1), ∥ a n ∥ , ∥ b n ∥ ≤ d / 2 \lVert a_{n}\rVert,\lVert b_{n}\rVert\le\sqrt d/2 ∥ a n ∥ , ∥ b n ∥ ≤ d /2 and ∥ Z n ∥ ≤ d \lVert Z_{n}\rVert\le\sqrt d ∥ Z n ∥ ≤ d everywhere, and for every w w w ,
q 3 ( w ) − q 1 ( w ) − Z n ( w ) = ( q 3 ( w ) − q 2 ( w ) − ϖ ( q 3 ( w ) − q 2 ( w ) ) ) + ( q 2 ( w ) − q 1 ( w ) − ϖ ( q 2 ( w ) − q 1 ( w ) ) ) ∈ Z d . \mathrm{q}_{3}(w)-\mathrm{q}_{1}(w)-Z_{n}(w)=\bigl(\mathrm{q}_{3}(w)-\mathrm{q}_{2}(w)-\varpi(\mathrm{q}_{3}(w)-\mathrm{q}_{2}(w))\bigr)+\bigl(\mathrm{q}_{2}(w)-\mathrm{q}_{1}(w)-\varpi(\mathrm{q}_{2}(w)-\mathrm{q}_{1}(w))\bigr)\in\mathbb{Z}^{d}. q 3 ( w ) − q 1 ( w ) − Z n ( w ) = ( q 3 ( w ) − q 2 ( w ) − ϖ ( q 3 ( w ) − q 2 ( w )) ) + ( q 2 ( w ) − q 1 ( w ) − ϖ ( q 2 ( w ) − q 1 ( w )) ) ∈ Z d .
By (F2), ( q 1 ) # Σ n = ( p r 1 ) # γ n = μ (\mathrm{q}_{1})_{\#}\Sigma_{n}=(\mathrm{pr}_{1})_{\#}\gamma_{n}=\mu ( q 1 ) # Σ n = ( pr 1 ) # γ n = μ and ( q 3 ) # Σ n = ( p r 2 ) # σ n ′ = ν (\mathrm{q}_{3})_{\#}\Sigma_{n}=(\mathrm{pr}_{2})_{\#}\sigma'_{n}=\nu ( q 3 ) # Σ n = ( pr 2 ) # σ n ′ = ν . Since ∥ a n ( w ) ∥ = d T ( q 1 ( w ) , q 2 ( w ) ) \lVert a_{n}(w)\rVert=d_{\mathbb{T}}(\mathrm{q}_{1}(w),\mathrm{q}_{2}(w)) ∥ a n ( w )∥ = d T ( q 1 ( w ) , q 2 ( w )) by (F1), the change-of-variables formula through ( q 1 , q 2 ) (\mathrm{q}_{1},\mathrm{q}_{2}) ( q 1 , q 2 ) and Probability Measures on the Flat Torus, the Torus Cost of a Coupling, the Torus Wasserstein Distance and Optimal Couplings §cost give ∫ ∥ a n ∥ 2 d Σ n = I T ( γ n ) = s n 2 \int\lVert a_{n}\rVert^{2}\,d\Sigma_{n}=I_{\mathbb{T}}(\gamma_{n})=s_{n}^{2} ∫ ∥ a n ∥ 2 d Σ n = I T ( γ n ) = s n 2 ; in the same way, through ( q 2 , q 3 ) (\mathrm{q}_{2},\mathrm{q}_{3}) ( q 2 , q 3 ) , ∫ ∥ b n ∥ 2 d Σ n = I T ( σ n ′ ) = W n ′ 2 \int\lVert b_{n}\rVert^{2}\,d\Sigma_{n}=I_{\mathbb{T}}(\sigma'_{n})=W_{n}'^{2} ∫ ∥ b n ∥ 2 d Σ n = I T ( σ n ′ ) = W n ′ 2 . Also, v T ∘ q 1 ⋅ a n v_{T}\circ\mathrm{q}_{1}\cdot a_{n} v T ∘ q 1 ⋅ a n is the Borel integrand of The Torus Displacement Pairing of a Vector Field Along a Coupling §pairing composed with ( q 1 , q 2 ) (\mathrm{q}_{1},\mathrm{q}_{2}) ( q 1 , q 2 ) , so by the change-of-variables formula
∫ R 3 d ( v T ∘ q 1 ) ⋅ a n d Σ n = J T ( v T , γ n ) . ( 2 ) \int_{\mathbb{R}^{3d}}(v_{T}\circ\mathrm{q}_{1})\cdot a_{n}\,d\Sigma_{n}=\mathcal{J}_{\mathbb{T}}(v_{T},\gamma_{n}).\qquad(2) ∫ R 3 d ( v T ∘ q 1 ) ⋅ a n d Σ n = J T ( v T , γ n ) . ( 2 )
Let L 2 ( Σ n ; R d ) L^{2}(\Sigma_{n};\mathbb{R}^{d}) L 2 ( Σ n ; R d ) be the space of Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §fields read with 3 d 3d 3 d in place of q q q and d d d in place of r r r , with inner product ⟨ ⋅ , ⋅ ⟩ Σ n \langle\cdot,\cdot\rangle_{\Sigma_{n}} ⟨ ⋅ , ⋅ ⟩ Σ n and norm ∥ ⋅ ∥ Σ n \lVert\cdot\rVert_{\Sigma_{n}} ∥ ⋅ ∥ Σ n ; it is a real inner product space by The Space of Square-Integrable Random Vectors §inner-product (indeed a real Hilbert space by The Space of Square-Integrable Random Vectors is a Real Hilbert Space; Its Laws Have Finite Second Moment; Constants and Translations §hilbert ). The classes of the bounded Borel maps a n a_{n} a n , b n b_{n} b n , Z n Z_{n} Z n and v T ∘ q 1 v_{T}\circ\mathrm{q}_{1} v T ∘ q 1 belong to it. So ∥ a n ∥ Σ n = s n \lVert a_{n}\rVert_{\Sigma_{n}}=s_{n} ∥ a n ∥ Σ n = s n , ∥ b n ∥ Σ n = W n ′ \lVert b_{n}\rVert_{\Sigma_{n}}=W'_{n} ∥ b n ∥ Σ n = W n ′ and, by the triangle inequality The Norm Metric of a Real Inner Product Space: Triangle Inequalities, Limits and Continuity §triangle , ∥ Z n ∥ Σ n ≤ s n + W n ′ \lVert Z_{n}\rVert_{\Sigma_{n}}\le s_{n}+W'_{n} ∥ Z n ∥ Σ n ≤ s n + W n ′ .
Step 2.3 (near optimality of Z n Z_{n} Z n ). By the triangle inequality and symmetry of the metric W T W_{\mathbb{T}} W T (The Torus Wasserstein Space is a Sequentially Compact Metric Space in which Optimal Couplings Exist §metric ), W n ′ ≤ W T ( μ , μ n ′ ) + W W'_{n}\le W_{\mathbb{T}}(\mu,\mu'_{n})+W W n ′ ≤ W T ( μ , μ n ′ ) + W , and W T ( μ , μ n ′ ) ≤ s n W_{\mathbb{T}}(\mu,\mu'_{n})\le s_{n} W T ( μ , μ n ′ ) ≤ s n by (F3) applied to γ n \gamma_{n} γ n . Hence ∥ Z n ∥ Σ n ≤ W + 2 s n \lVert Z_{n}\rVert_{\Sigma_{n}}\le W+2s_{n} ∥ Z n ∥ Σ n ≤ W + 2 s n , and squaring these nonnegative numbers and using s n 2 ≤ s n < 1 / ( n + 1 ) s_{n}^{2}\le s_{n}<1/(n+1) s n 2 ≤ s n < 1/ ( n + 1 ) ,
∫ R 3 d ∥ Z n ∥ 2 d Σ n ≤ W 2 + 4 W s n + 4 s n 2 ≤ W 2 + 4 W + 4 n + 1 . \int_{\mathbb{R}^{3d}}\lVert Z_{n}\rVert^{2}\,d\Sigma_{n}\le W^{2}+4Ws_{n}+4s_{n}^{2}\le W^{2}+\frac{4W+4}{n+1}. ∫ R 3 d ∥ Z n ∥ 2 d Σ n ≤ W 2 + 4 W s n + 4 s n 2 ≤ W 2 + n + 1 4 W + 4 .
Given a real ε ′ > 0 \varepsilon'>0 ε ′ > 0 , choose N ∈ N N\in\mathbb{N} N ∈ N with ( 4 W + 4 ) / ( N + 1 ) ≤ ε ′ (4W+4)/(N+1)\le\varepsilon' ( 4 W + 4 ) / ( N + 1 ) ≤ ε ′ (Archimedean property); then ∫ ∥ Z n ∥ 2 d Σ n ≤ W 2 + ε ′ \int\lVert Z_{n}\rVert^{2}\,d\Sigma_{n}\le W^{2}+\varepsilon' ∫ ∥ Z n ∥ 2 d Σ n ≤ W 2 + ε ′ for every n ≥ N n\ge N n ≥ N . Together with Step 2.2, this verifies all hypotheses of Nearly Optimal Composite Displacements on the Torus Converge to the Optimal Displacement with q = 3 d q=3d q = 3 d , X n = q 1 X_{n}=\mathrm{q}_{1} X n = q 1 , Y n = q 3 Y_{n}=\mathrm{q}_{3} Y n = q 3 and the Z n Z_{n} Z n above (μ \mu μ is absolutely continuous and T T T is an optimal map from μ \mu μ to ν \nu ν ). By Nearly Optimal Composite Displacements on the Torus Converge to the Optimal Displacement §convergence , the real sequence
c n = ∫ R 3 d ∥ Z n − v T ∘ q 1 ∥ 2 d Σ n = ∥ Z n − v T ∘ q 1 ∥ Σ n 2 c_{n}=\int_{\mathbb{R}^{3d}}\lVert Z_{n}-v_{T}\circ\mathrm{q}_{1}\rVert^{2}\,d\Sigma_{n}=\lVert Z_{n}-v_{T}\circ\mathrm{q}_{1}\rVert_{\Sigma_{n}}^{2} c n = ∫ R 3 d ∥ Z n − v T ∘ q 1 ∥ 2 d Σ n = ∥ Z n − v T ∘ q 1 ∥ Σ n 2
converges to 0 0 0 .
Step 2.4 (the reverse inequality). By Gluing Two Couplings over a Common Middle Marginal, and the Composite Coupling §composite , κ n = ( q 1 , q 3 ) # Σ n ∈ Π ( μ , ν ) \kappa_{n}=(\mathrm{q}_{1},\mathrm{q}_{3})_{\#}\Sigma_{n}\in\Pi(\mu,\nu) κ n = ( q 1 , q 3 ) # Σ n ∈ Π ( μ , ν ) , and by the change-of-variables formula I T ( κ n ) = ∫ d T ( q 1 ( w ) , q 3 ( w ) ) 2 Σ n ( d w ) I_{\mathbb{T}}(\kappa_{n})=\int d_{\mathbb{T}}(\mathrm{q}_{1}(w),\mathrm{q}_{3}(w))^{2}\,\Sigma_{n}(dw) I T ( κ n ) = ∫ d T ( q 1 ( w ) , q 3 ( w ) ) 2 Σ n ( d w ) . For each w w w , taking k = q 3 ( w ) − q 1 ( w ) − Z n ( w ) ∈ Z d k=\mathrm{q}_{3}(w)-\mathrm{q}_{1}(w)-Z_{n}(w)\in\mathbb{Z}^{d} k = q 3 ( w ) − q 1 ( w ) − Z n ( w ) ∈ Z d (Step 2.2) in (F1) gives d T ( q 1 ( w ) , q 3 ( w ) ) ≤ ∥ Z n ( w ) ∥ d_{\mathbb{T}}(\mathrm{q}_{1}(w),\mathrm{q}_{3}(w))\le\lVert Z_{n}(w)\rVert d T ( q 1 ( w ) , q 3 ( w )) ≤ ∥ Z n ( w )∥ . Hence, by (F3), monotonicity, and bilinearity and symmetry of ⟨ ⋅ , ⋅ ⟩ Σ n \langle\cdot,\cdot\rangle_{\Sigma_{n}} ⟨ ⋅ , ⋅ ⟩ Σ n ,
2 Φ ( μ ) = W 2 ≤ I T ( κ n ) ≤ ∥ Z n ∥ Σ n 2 = s n 2 + 2 ⟨ a n , b n ⟩ Σ n + 2 Φ ( μ n ′ ) . 2\Phi(\mu)=W^{2}\le I_{\mathbb{T}}(\kappa_{n})\le\lVert Z_{n}\rVert_{\Sigma_{n}}^{2}=s_{n}^{2}+2\langle a_{n},b_{n}\rangle_{\Sigma_{n}}+2\Phi(\mu'_{n}). 2Φ ( μ ) = W 2 ≤ I T ( κ n ) ≤ ∥ Z n ∥ Σ n 2 = s n 2 + 2 ⟨ a n , b n ⟩ Σ n + 2Φ ( μ n ′ ) .
Since b n = Z n − a n b_{n}=Z_{n}-a_{n} b n = Z n − a n and Z n = v T ∘ q 1 + ( Z n − v T ∘ q 1 ) Z_{n}=v_{T}\circ\mathrm{q}_{1}+(Z_{n}-v_{T}\circ\mathrm{q}_{1}) Z n = v T ∘ q 1 + ( Z n − v T ∘ q 1 ) , bilinearity and (2) give
⟨ a n , b n ⟩ Σ n = J T ( v T , γ n ) + ⟨ a n , Z n − v T ∘ q 1 ⟩ Σ n − s n 2 , \langle a_{n},b_{n}\rangle_{\Sigma_{n}}=\mathcal{J}_{\mathbb{T}}(v_{T},\gamma_{n})+\langle a_{n},Z_{n}-v_{T}\circ\mathrm{q}_{1}\rangle_{\Sigma_{n}}-s_{n}^{2}, ⟨ a n , b n ⟩ Σ n = J T ( v T , γ n ) + ⟨ a n , Z n − v T ∘ q 1 ⟩ Σ n − s n 2 ,
and the middle term has absolute value at most s n c n s_{n}\sqrt{c_{n}} s n c n by the Cauchy–Schwarz inequality The Cauchy-Schwarz Inequality in a Real Inner Product Space in L 2 ( Σ n ; R d ) L^{2}(\Sigma_{n};\mathbb{R}^{d}) L 2 ( Σ n ; R d ) . Substituting,
2 Φ ( μ ) ≤ 2 Φ ( μ n ′ ) + 2 J T ( v T , γ n ) + 2 s n c n − s n 2 , 2\Phi(\mu)\le2\Phi(\mu'_{n})+2\,\mathcal{J}_{\mathbb{T}}(v_{T},\gamma_{n})+2s_{n}\sqrt{c_{n}}-s_{n}^{2}, 2Φ ( μ ) ≤ 2Φ ( μ n ′ ) + 2 J T ( v T , γ n ) + 2 s n c n − s n 2 ,
so, as s n 2 ≥ 0 s_{n}^{2}\ge0 s n 2 ≥ 0 ,
Φ ( μ n ′ ) − Φ ( μ ) + J T ( v T , γ n ) ≥ − s n c n for every n ∈ N . ( 3 ) \Phi(\mu'_{n})-\Phi(\mu)+\mathcal{J}_{\mathbb{T}}(v_{T},\gamma_{n})\ge-s_{n}\sqrt{c_{n}}\qquad\text{for every }n\in\mathbb{N}.\qquad(3) Φ ( μ n ′ ) − Φ ( μ ) + J T ( v T , γ n ) ≥ − s n c n for every n ∈ N . ( 3 )
Step 2.5 (contradiction). Since ( c n ) (c_{n}) ( c n ) converges to 0 0 0 (Limit of a Sequence of Real Numbers ), there is n n n with c n < ε 2 c_{n}<\varepsilon^{2} c n < ε 2 , hence c n ≤ ε \sqrt{c_{n}}\le\varepsilon c n ≤ ε by claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field , both c n \sqrt{c_{n}} c n and ε \varepsilon ε being nonnegative. For this n n n , (3) and s n ≥ 0 s_{n}\ge0 s n ≥ 0 give Φ ( μ n ′ ) − Φ ( μ ) + J T ( v T , γ n ) ≥ − ε s n \Phi(\mu'_{n})-\Phi(\mu)+\mathcal{J}_{\mathbb{T}}(v_{T},\gamma_{n})\ge-\varepsilon s_{n} Φ ( μ n ′ ) − Φ ( μ ) + J T ( v T , γ n ) ≥ − ε s n , contradicting (1). This proves the claim of Step 2.1.
Step 2.6 (conclusion). Let ε > 0 \varepsilon>0 ε > 0 , let θ 1 \theta_{1} θ 1 be as in Step 2.1, and let θ \theta θ be the smaller of 2 ε 2\varepsilon 2 ε and θ 1 \theta_{1} θ 1 (Elementary Properties of the Minimum of Two Elements ). Let μ ′ ∈ P ( T d ) \mu'\in\mathcal{P}(\mathbb{T}^{d}) μ ′ ∈ P ( T d ) and γ ∈ Π ( μ , μ ′ ) \gamma\in\Pi(\mu,\mu') γ ∈ Π ( μ , μ ′ ) with I T ( γ ) < θ 2 I_{\mathbb{T}}(\gamma)<\theta^{2} I T ( γ ) < θ 2 ; then I T ( γ ) < θ ≤ 2 ε \sqrt{I_{\mathbb{T}}(\gamma)}<\theta\le2\varepsilon I T ( γ ) < θ ≤ 2 ε , so 1 2 I T ( γ ) = 1 2 I T ( γ ) I T ( γ ) ≤ ε I T ( γ ) \tfrac12I_{\mathbb{T}}(\gamma)=\tfrac12\sqrt{I_{\mathbb{T}}(\gamma)}\sqrt{I_{\mathbb{T}}(\gamma)}\le\varepsilon\sqrt{I_{\mathbb{T}}(\gamma)} 2 1 I T ( γ ) = 2 1 I T ( γ ) I T ( γ ) ≤ ε I T ( γ ) . By clause 1, proved above (with the same T T T ), Φ ( μ ′ ) − Φ ( μ ) + J T ( v T , γ ) ≤ 1 2 I T ( γ ) ≤ ε I T ( γ ) \Phi(\mu')-\Phi(\mu)+\mathcal{J}_{\mathbb{T}}(v_{T},\gamma)\le\tfrac12I_{\mathbb{T}}(\gamma)\le\varepsilon\sqrt{I_{\mathbb{T}}(\gamma)} Φ ( μ ′ ) − Φ ( μ ) + J T ( v T , γ ) ≤ 2 1 I T ( γ ) ≤ ε I T ( γ ) , and by Step 2.1, since I T ( γ ) < θ 1 2 I_{\mathbb{T}}(\gamma)<\theta_{1}^{2} I T ( γ ) < θ 1 2 , the same quantity is at least − ε I T ( γ ) -\varepsilon\sqrt{I_{\mathbb{T}}(\gamma)} − ε I T ( γ ) . This is ( ∗ ) (\ast) ( ∗ ) . Hence Φ \Phi Φ is differentiable along couplings at μ \mu μ with gradient − v T -v_{T} − v T , and by the uniqueness in Differentiability Along Couplings of a Function on the Torus Wasserstein Space, and Its Gradient §gradient , ∇ Φ ( μ ) = − v T \nabla\Phi(\mu)=-v_{T} ∇Φ ( μ ) = − v T .
Proof of clause 3. Keep μ \mu μ , T T T and v T v_{T} v T as in clause 2. By McCann's Theorem on the Flat Torus: Optimal Couplings out of an Absolutely Continuous Measure are Induced by a Unique Map with a Periodic Potential §potential there are a Z d \mathbb{Z}^{d} Z d -periodic function φ : R d → R \varphi:\mathbb{R}^{d}\to\mathbb{R} φ : R d → R , Lipschitz with constant d / 2 \sqrt d/2 d /2 , such that ψ = β − φ \psi=\beta-\varphi ψ = β − φ is convex on R d \mathbb{R}^{d} R d , where β ( x ) = 1 2 ∥ x ∥ 2 \beta(x)=\tfrac12\lVert x\rVert^{2} β ( x ) = 2 1 ∥ x ∥ 2 , and a set D ∈ B ( R d ) D\in\mathcal{B}(\mathbb{R}^{d}) D ∈ B ( R d ) with μ ( D ) = 1 \mu(D)=1 μ ( D ) = 1 such that ∂ R d ψ ( x ) = { x + v T ( x ) } \partial_{\mathbb{R}^{d}}\psi(x)=\{x+v_{T}(x)\} ∂ R d ψ ( x ) = { x + v T ( x )} for every x ∈ D x\in D x ∈ D . Subdifferentials are those of Subdifferential of a Real-Valued Function on a Convex Subset of R n \mathbb{R}^n R n §subdifferential , relative to R d \mathbb{R}^{d} R d , which is open and convex. Recall that Q Q Q is the half-open unit cell of The Half-Open Unit Cell Tiles Euclidean Space §cell , so every x ∈ Q x\in Q x ∈ Q has 0 ≤ x i < 1 0\le x_{i}<1 0 ≤ x i < 1 for all i i i and hence ∥ x ∥ ≤ d \lVert x\rVert\le\sqrt d ∥ x ∥ ≤ d ; and μ ( Q ) = 1 \mu(Q)=1 μ ( Q ) = 1 by Probability Measures on the Flat Torus, the Torus Cost of a Coupling, the Torus Wasserstein Distance and Optimal Couplings §measures .
Step 3.1 (the quadratic). The zero function on R d \mathbb{R}^{d} R d satisfies the quadratic increment inequality of Quadratic Increment Characterisation of Semiconvexity with constant 1 1 1 , since its right side there is 1 2 t ( 1 − t ) ∥ x − y ∥ 2 ≥ 0 \tfrac12t(1-t)\lVert x-y\rVert^{2}\ge0 2 1 t ( 1 − t ) ∥ x − y ∥ 2 ≥ 0 ; hence it is semiconvex with constant 1 1 1 , which by Semiconvex Function on a Convex Subset of R n \mathbb{R}^n R n means that β \beta β is convex on R d \mathbb{R}^{d} R d . For x , y , p ∈ R d x,y,p\in\mathbb{R}^{d} x , y , p ∈ R d , expanding by bilinearity of the dot product, β ( y ) − β ( x ) − x ⋅ ( y − x ) = 1 2 ∥ y − x ∥ 2 ≥ 0 \beta(y)-\beta(x)-x\cdot(y-x)=\tfrac12\lVert y-x\rVert^{2}\ge0 β ( y ) − β ( x ) − x ⋅ ( y − x ) = 2 1 ∥ y − x ∥ 2 ≥ 0 , so x ∈ ∂ β ( x ) x\in\partial\beta(x) x ∈ ∂ β ( x ) ; and if p ∈ ∂ β ( x ) p\in\partial\beta(x) p ∈ ∂ β ( x ) then, taking y = p y=p y = p , 0 ≤ β ( p ) − β ( x ) − p ⋅ ( p − x ) = − 1 2 ∥ p − x ∥ 2 0\le\beta(p)-\beta(x)-p\cdot(p-x)=-\tfrac12\lVert p-x\rVert^{2} 0 ≤ β ( p ) − β ( x ) − p ⋅ ( p − x ) = − 2 1 ∥ p − x ∥ 2 , so p = x p=x p = x . Thus ∂ β ( x ) = { x } \partial\beta(x)=\{x\} ∂ β ( x ) = { x } for every x x x .
Step 3.2 (local bound on subgradients of ψ \psi ψ ). Put R = d + 2 R=\sqrt d+2 R = d + 2 and L = 2 R + d / 2 L=2R+\sqrt d/2 L = 2 R + d /2 . For z , w ∈ B ˉ ( 0 , 2 R ) z,w\in\bar{B}(0,2R) z , w ∈ B ˉ ( 0 , 2 R ) , ∣ β ( z ) − β ( w ) ∣ = 1 2 ∣ ( z − w ) ⋅ ( z + w ) ∣ ≤ 1 2 ∥ z − w ∥ ( ∥ z ∥ + ∥ w ∥ ) ≤ 2 R ∥ z − w ∥ |\beta(z)-\beta(w)|=\tfrac12|(z-w)\cdot(z+w)|\le\tfrac12\lVert z-w\rVert(\lVert z\rVert+\lVert w\rVert)\le2R\lVert z-w\rVert ∣ β ( z ) − β ( w ) ∣ = 2 1 ∣ ( z − w ) ⋅ ( z + w ) ∣ ≤ 2 1 ∥ z − w ∥ (∥ z ∥ + ∥ w ∥) ≤ 2 R ∥ z − w ∥ and ∣ φ ( z ) − φ ( w ) ∣ ≤ d 2 ∥ z − w ∥ |\varphi(z)-\varphi(w)|\le\tfrac{\sqrt d}{2}\lVert z-w\rVert ∣ φ ( z ) − φ ( w ) ∣ ≤ 2 d ∥ z − w ∥ , so ∣ ψ ( z ) − ψ ( w ) ∣ ≤ L ∥ z − w ∥ |\psi(z)-\psi(w)|\le L\lVert z-w\rVert ∣ ψ ( z ) − ψ ( w ) ∣ ≤ L ∥ z − w ∥ . By Elementary Calculus of the Subdifferential of a Convex Function §bounded (with U = R d U=\mathbb{R}^{d} U = R d , y 0 = 0 y_{0}=0 y 0 = 0 , r = R r=R r = R , M = L M=L M = L ), every subgradient of ψ \psi ψ at a point of B ˉ ( 0 , R ) \bar{B}(0,R) B ˉ ( 0 , R ) has norm at most L L L ; and ∂ ψ ( y ) ≠ ∅ \partial\psi(y)\neq\emptyset ∂ ψ ( y ) = ∅ for every y y y by The Subdifferential of a Convex Function on an Open Convex Set is Nonempty §nonempty .
Step 3.3 (Lipschitz truncations). Apply The Lipschitz Truncation of a Convex Function: a Global Lipschitz Convex Minorant Agreeing with It Where the Slope is Small with n = d n=d n = d , G = R d G=\mathbb{R}^{d} G = R d and level L L L , first to β \beta β (witness x 0 = 0 x_{0}=0 x 0 = 0 , q 0 = 0 ∈ ∂ β ( 0 ) q_{0}=0\in\partial\beta(0) q 0 = 0 ∈ ∂ β ( 0 ) , Step 3.1) and then to ψ \psi ψ (witness x 0 = 0 x_{0}=0 x 0 = 0 and any q 0 ∈ ∂ ψ ( 0 ) q_{0}\in\partial\psi(0) q 0 ∈ ∂ ψ ( 0 ) , of norm at most L L L by Step 3.2). By The Lipschitz Truncation of a Convex Function: a Global Lipschitz Convex Minorant Agreeing with It Where the Slope is Small §minorant , β L \beta^{L} β L and ψ L \psi^{L} ψ L are convex on R d \mathbb{R}^{d} R d and Lipschitz with constant L L L . Let y ∈ B ˉ ( 0 , R ) y\in\bar{B}(0,R) y ∈ B ˉ ( 0 , R ) . Since y ∈ ∂ β ( y ) y\in\partial\beta(y) y ∈ ∂ β ( y ) and ∥ y ∥ ≤ R ≤ L \lVert y\rVert\le R\le L ∥ y ∥ ≤ R ≤ L , and since ψ \psi ψ has a subgradient at y y y of norm at most L L L (Step 3.2), The Lipschitz Truncation of a Convex Function: a Global Lipschitz Convex Minorant Agreeing with It Where the Slope is Small §agreement gives β L ( y ) = β ( y ) \beta^{L}(y)=\beta(y) β L ( y ) = β ( y ) and ψ L ( y ) = ψ ( y ) \psi^{L}(y)=\psi(y) ψ L ( y ) = ψ ( y ) ; hence
φ ( y ) = β L ( y ) − ψ L ( y ) ( y ∈ B ˉ ( 0 , R ) ) . ( 4 ) \varphi(y)=\beta^{L}(y)-\psi^{L}(y)\qquad(y\in\bar{B}(0,R)).\qquad(4) φ ( y ) = β L ( y ) − ψ L ( y ) ( y ∈ B ˉ ( 0 , R )) . ( 4 )
By The Lipschitz Truncation of a Convex Function: a Global Lipschitz Convex Minorant Agreeing with It Where the Slope is Small §unique and Steps 3.1 and 3.2: ∂ β L ( y ) = { y } \partial\beta^{L}(y)=\{y\} ∂ β L ( y ) = { y } for y ∈ B ˉ ( 0 , R ) y\in\bar{B}(0,R) y ∈ B ˉ ( 0 , R ) , and ∂ ψ L ( x ) = { x + v T ( x ) } \partial\psi^{L}(x)=\{x+v_{T}(x)\} ∂ ψ L ( x ) = { x + v T ( x )} for x ∈ D ∩ B ˉ ( 0 , R ) x\in D\cap\bar{B}(0,R) x ∈ D ∩ B ˉ ( 0 , R ) , the point x + v T ( x ) x+v_{T}(x) x + v T ( x ) being a subgradient of ψ \psi ψ at a point of B ˉ ( 0 , R ) \bar{B}(0,R) B ˉ ( 0 , R ) and so of norm at most L L L .
Step 3.4 (mollification). Fix a mollifier kernel ρ \rho ρ of radius 1 1 1 on R d \mathbb{R}^{d} R d (claim 2 of Existence of Mollifier Kernels of Every Radius , with n = d n=d n = d , δ = 1 \delta=1 δ = 1 ). For m ∈ N m\in\mathbb{N} m ∈ N put ε m = 1 / ( m + 1 ) \varepsilon_{m}=1/(m+1) ε m = 1/ ( m + 1 ) , a sequence of positive reals converging to 0 0 0 by the Archimedean property, and ρ m ( y ) = ( ε m − 1 ) d ρ ( ε m − 1 y ) \rho_{m}(y)=(\varepsilon_{m}^{-1})^{d}\rho(\varepsilon_{m}^{-1}y) ρ m ( y ) = ( ε m − 1 ) d ρ ( ε m − 1 y ) , a mollifier kernel of radius ε m ≤ 1 \varepsilon_{m}\le1 ε m ≤ 1 by Rescaling a Mollifier Kernel ; in particular ρ m \rho_{m} ρ m is smooth, hence continuous, and vanishes off B ˉ ( 0 , ε m ) \bar{B}(0,\varepsilon_{m}) B ˉ ( 0 , ε m ) . The function φ \varphi φ is continuous by A Lipschitz Map is Uniformly Continuous . Let φ m = φ ∗ ρ m \varphi_{m}=\varphi*\rho_{m} φ m = φ ∗ ρ m , β m = β L ∗ ρ m \beta_{m}=\beta^{L}*\rho_{m} β m = β L ∗ ρ m and ψ m = ψ L ∗ ρ m \psi_{m}=\psi^{L}*\rho_{m} ψ m = ψ L ∗ ρ m be the convolutions with Ω = R d \Omega=\mathbb{R}^{d} Ω = R d , defined on all of R d \mathbb{R}^{d} R d ; the last two are the mollifications of Mollification of a Lipschitz Convex Function: Smooth Convex Approximations with Bounded Gradients Converging Where the Subgradient is Unique for δ = 1 \delta=1 δ = 1 and ε = ε m \varepsilon=\varepsilon_{m} ε = ε m .
(a) φ m ∈ C p e r ∞ \varphi_{m}\in C^{\infty}_{\mathrm{per}} φ m ∈ C per ∞ : it is smooth on R d \mathbb{R}^{d} R d by claim 2 of Convolution with a C k C^k C k Kernel is of Class C k C^k C k ; and for x ∈ R d x\in\mathbb{R}^{d} x ∈ R d and k ∈ Z d k\in\mathbb{Z}^{d} k ∈ Z d the integrands defining φ m ( x + k ) \varphi_{m}(x+k) φ m ( x + k ) and φ m ( x ) \varphi_{m}(x) φ m ( x ) coincide, since φ ( x + k − y ) ρ m ( y ) = φ ( x − y ) ρ m ( y ) \varphi(x+k-y)\rho_{m}(y)=\varphi(x-y)\rho_{m}(y) φ ( x + k − y ) ρ m ( y ) = φ ( x − y ) ρ m ( y ) by periodicity of φ \varphi φ (Lattice-Periodic Functions and the Periodic Function Classes §periodic ), so φ m ( x + k ) = φ m ( x ) \varphi_{m}(x+k)=\varphi_{m}(x) φ m ( x + k ) = φ m ( x ) (Lattice-Periodic Functions and the Periodic Function Classes §classes ). Hence the class of ∇ φ m \nabla\varphi_{m} ∇ φ m lies in G μ G_{\mu} G μ (The Tangent Space of the Torus Wasserstein Space at a Probability Measure §gradients ), and ∇ φ m \nabla\varphi_{m} ∇ φ m is Borel by Optimal Transport on the Flat Torus: Standing Notation §calculus .
(b) By Mollification of a Lipschitz Convex Function: Smooth Convex Approximations with Bounded Gradients Converging Where the Subgradient is Unique §regularity , applied to β L \beta^{L} β L and to ψ L \psi^{L} ψ L , β m \beta_{m} β m and ψ m \psi_{m} ψ m are smooth with ∥ D β m ( x ) ∥ ≤ L \lVert D\beta_{m}(x)\rVert\le L ∥ D β m ( x )∥ ≤ L and ∥ D ψ m ( x ) ∥ ≤ L \lVert D\psi_{m}(x)\rVert\le L ∥ D ψ m ( x )∥ ≤ L for every x x x .
(c) Let U = { x ∈ R d : ∥ x ∥ < d + 1 } U=\{x\in\mathbb{R}^{d}:\lVert x\rVert<\sqrt d+1\} U = { x ∈ R d : ∥ x ∥ < d + 1 } ; it is open, since for x ∈ U x\in U x ∈ U every z z z with ∥ z − x ∥ < d + 1 − ∥ x ∥ \lVert z-x\rVert<\sqrt d+1-\lVert x\rVert ∥ z − x ∥ < d + 1 − ∥ x ∥ lies in U U U by the triangle inequality, and Q ⊆ U Q\subseteq U Q ⊆ U . For x ′ ∈ U x'\in U x ′ ∈ U and y ∈ R d y\in\mathbb{R}^{d} y ∈ R d : if ∥ y ∥ ≤ ε m \lVert y\rVert\le\varepsilon_{m} ∥ y ∥ ≤ ε m then ∥ x ′ − y ∥ < d + 2 = R \lVert x'-y\rVert<\sqrt d+2=R ∥ x ′ − y ∥ < d + 2 = R , so by (4) φ ( x ′ − y ) ρ m ( y ) = β L ( x ′ − y ) ρ m ( y ) − ψ L ( x ′ − y ) ρ m ( y ) \varphi(x'-y)\rho_{m}(y)=\beta^{L}(x'-y)\rho_{m}(y)-\psi^{L}(x'-y)\rho_{m}(y) φ ( x ′ − y ) ρ m ( y ) = β L ( x ′ − y ) ρ m ( y ) − ψ L ( x ′ − y ) ρ m ( y ) ; otherwise all three products vanish. The three integrands are integrable by claim 1 of The Convolution Integrand is Continuous, Compactly Supported and Integrable , so linearity (claim 2 of Linearity and Monotonicity of the Lebesgue Integral ) gives φ m = β m − ψ m \varphi_{m}=\beta_{m}-\psi_{m} φ m = β m − ψ m on U U U . By Differential Calculus and Convexity on Euclidean Open Sets: Standing Notation §derivatives the partial derivatives at points of U U U of φ m \varphi_{m} φ m and of β m − ψ m \beta_{m}-\psi_{m} β m − ψ m are those of their common restriction to U U U , and by claim 1 of Constants, Coordinate Functions, Sums and Products of C k C^k C k Functions on a Euclidean Open Set those of β m − ψ m \beta_{m}-\psi_{m} β m − ψ m are the differences of those of β m \beta_{m} β m and ψ m \psi_{m} ψ m . Hence ∇ φ m ( x ) = D β m ( x ) − D ψ m ( x ) \nabla\varphi_{m}(x)=D\beta_{m}(x)-D\psi_{m}(x) ∇ φ m ( x ) = D β m ( x ) − D ψ m ( x ) and, by (b), ∥ ∇ φ m ( x ) ∥ ≤ 2 L \lVert\nabla\varphi_{m}(x)\rVert\le2L ∥ ∇ φ m ( x )∥ ≤ 2 L for every x ∈ U x\in U x ∈ U , in particular for every x ∈ Q x\in Q x ∈ Q .
(d) Let x ∈ D ∩ Q x\in D\cap Q x ∈ D ∩ Q . Since Q ⊆ B ˉ ( 0 , R ) Q\subseteq\bar{B}(0,R) Q ⊆ B ˉ ( 0 , R ) , Step 3.3 gives ∂ β L ( x ) = { x } \partial\beta^{L}(x)=\{x\} ∂ β L ( x ) = { x } and ∂ ψ L ( x ) = { x + v T ( x ) } \partial\psi^{L}(x)=\{x+v_{T}(x)\} ∂ ψ L ( x ) = { x + v T ( x )} , so by Mollification of a Lipschitz Convex Function: Smooth Convex Approximations with Bounded Gradients Converging Where the Subgradient is Unique §gradients the sequences ( D β m ( x ) ) m (D\beta_{m}(x))_{m} ( D β m ( x ) ) m and ( D ψ m ( x ) ) m (D\psi_{m}(x))_{m} ( D ψ m ( x ) ) m converge to x x x and to x + v T ( x ) x+v_{T}(x) x + v T ( x ) . By (c),
∥ ∇ φ m ( x ) + v T ( x ) ∥ ≤ ∥ D β m ( x ) − x ∥ + ∥ D ψ m ( x ) − x − v T ( x ) ∥ ; \lVert\nabla\varphi_{m}(x)+v_{T}(x)\rVert\le\lVert D\beta_{m}(x)-x\rVert+\lVert D\psi_{m}(x)-x-v_{T}(x)\rVert; ∥ ∇ φ m ( x ) + v T ( x )∥ ≤ ∥ D β m ( x ) − x ∥ + ∥ D ψ m ( x ) − x − v T ( x )∥ ;
given a real η > 0 \eta>0 η > 0 , both terms on the right are less than η / 2 \sqrt\eta/2 η /2 for all m m m beyond some index, so f m ( x ) = ∥ ∇ φ m ( x ) + v T ( x ) ∥ 2 < η f_{m}(x)=\lVert\nabla\varphi_{m}(x)+v_{T}(x)\rVert^{2}<\eta f m ( x ) = ∥ ∇ φ m ( x ) + v T ( x ) ∥ 2 < η there. Thus ( f m ( x ) ) m (f_{m}(x))_{m} ( f m ( x ) ) m converges to 0 0 0 for every x ∈ D ∩ Q x\in D\cap Q x ∈ D ∩ Q .
(e) The functions f m : R d → R f_{m}:\mathbb{R}^{d}\to\mathbb{R} f m : R d → R are Borel by (a), (F1) and (F2). For x ∈ Q x\in Q x ∈ Q , 0 ≤ f m ( x ) ≤ ( 2 L + d / 2 ) 2 0\le f_{m}(x)\le(2L+\sqrt d/2)^{2} 0 ≤ f m ( x ) ≤ ( 2 L + d /2 ) 2 by (c) and (F1), and the constant ( 2 L + d / 2 ) 2 (2L+\sqrt d/2)^{2} ( 2 L + d /2 ) 2 is integrable against μ \mu μ . The set ( R d ∖ D ) ∪ ( R d ∖ Q ) (\mathbb{R}^{d}\setminus D)\cup(\mathbb{R}^{d}\setminus Q) ( R d ∖ D ) ∪ ( R d ∖ Q ) is μ \mu μ -null by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §null-union , so the convergence in (d) and the bound hold for μ \mu μ -almost every x x x . By The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §dominated , applied on ( R d , B ( R d ) , μ ) (\mathbb{R}^{d},\mathcal{B}(\mathbb{R}^{d}),\mu) ( R d , B ( R d ) , μ ) with limit 0 0 0 , ∫ f m d μ \int f_{m}\,d\mu ∫ f m d μ converges to 0 0 0 .
(f) Since ∫ f m d μ = ∥ ∇ φ m − ( − v T ) ∥ μ 2 \int f_{m}\,d\mu=\lVert\nabla\varphi_{m}-(-v_{T})\rVert_{\mu}^{2} ∫ f m d μ = ∥ ∇ φ m − ( − v T ) ∥ μ 2 is the square of the distance in L 2 ( μ ; R d ) L^{2}(\mu;\mathbb{R}^{d}) L 2 ( μ ; R d ) between the classes of ∇ φ m \nabla\varphi_{m} ∇ φ m and − v T -v_{T} − v T (Optimal Transport on the Flat Torus: Standing Notation §fields ), that distance converges to 0 0 0 : given η > 0 \eta>0 η > 0 , eventually ∫ f m d μ < η 2 \int f_{m}\,d\mu<\eta^{2} ∫ f m d μ < η 2 . So the sequence ( ∇ φ m ) m (\nabla\varphi_{m})_{m} ( ∇ φ m ) m of elements of G μ G_{\mu} G μ converges to − v T -v_{T} − v T , and by Sequential Characterization of the Closure in a Metric Space the class of − v T -v_{T} − v T lies in the closure of G μ G_{\mu} G μ , which is T μ T_{\mu} T μ by The Tangent Space of the Torus Wasserstein Space at a Probability Measure §tangent . As T μ T_{\mu} T μ is a linear subspace of L 2 ( μ ; R d ) L^{2}(\mu;\mathbb{R}^{d}) L 2 ( μ ; R d ) by The Torus Tangent Space: Linearity of the Periodic Calculus, Closed Subspace, and Representation of Bounded Functionals on Gradients §subspace , the class of v T = ( − 1 ) ( − v T ) v_{T}=(-1)(-v_{T}) v T = ( − 1 ) ( − v T ) lies in T μ T_{\mu} T μ . This proves clause 3.