The proof adapts the competitor construction of the published proof of The Support of an Optimal Coupling is Cyclically Monotone , with the quadratic cost replaced by the torus cost. Each result cited below is universally quantified over the data in its own statement. Finite sums of real numbers are those of Finite Sum Notation in a Field , and for N ∈ N N\in\mathbb{N} N ∈ N we write Σ N \Sigma_{N} Σ N for the finite sum ∑ i = 1 N 1 \sum_{i=1}^{N}1 ∑ i = 1 N 1 of N N N copies of 1 1 1 , a positive real number by claim 6 of Properties of Finite Sums together with claims 6 and 2 of Elementary Order Arithmetic in an Ordered Field . We use once that a finite nonempty set of real numbers has a least member, which follows from claim 9 of Elementary Order Arithmetic in an Ordered Field together with the trichotomy of the order and induction on the number of members. The maps p r 1 , p r 2 : R d + d → R d \mathrm{pr}_{1},\mathrm{pr}_{2}:\mathbb{R}^{d+d}\to\mathbb{R}^{d} pr 1 , pr 2 : R d + d → R d are the coordinate projections , ι \iota ι is the concatenation map fixed there, Π ( ⋅ , ⋅ ) \Pi(\cdot,\cdot) Π ( ⋅ , ⋅ ) denotes sets of couplings , and B ( z , r ) B(z,r) B ( z , r ) is the open ball of the metric space ( R d + d , d E ) (\mathbb{R}^{d+d},d_{E}) ( R d + d , d E ) .
0. Preliminaries on the cost. Let c T : R d + d → R c_{\mathbb{T}}:\mathbb{R}^{d+d}\to\mathbb{R} c T : R d + d → R be c T ( z ) = d T ( p r 1 ( z ) , p r 2 ( z ) ) 2 c_{\mathbb{T}}(z)=d_{\mathbb{T}}(\mathrm{pr}_{1}(z),\mathrm{pr}_{2}(z))^{2} c T ( z ) = d T ( pr 1 ( z ) , pr 2 ( z ) ) 2 . By Probability Measures on the Flat Torus, the Torus Cost of a Coupling, the Torus Wasserstein Distance and Optimal Couplings §cost it is Borel with values in [ 0 , d / 4 ] [0,d/4] [ 0 , d /4 ] ; so, by Probability Measures on Euclidean Space and Random Vectors: Standing Notation §measures , it is integrable with respect to every member of P ( R d + d ) \mathcal{P}(\mathbb{R}^{d+d}) P ( R d + d ) , its integral against such a measure being a real number in [ 0 , d / 4 ] [0,d/4] [ 0 , d /4 ] by claim 1 of Linearity and Monotonicity of the Lebesgue Integral and The Integral of an Indicator Function is the Measure of the Set ; and I T ( ρ ) = ∫ c T d ρ I_{\mathbb{T}}(\rho)=\int c_{\mathbb{T}}\,d\rho I T ( ρ ) = ∫ c T d ρ for every coupling ρ \rho ρ of two members of P ( T d ) \mathcal{P}(\mathbb{T}^{d}) P ( T d ) . Put Λ = d + d \Lambda=\sqrt{d}+\sqrt{d} Λ = d + d , a nonnegative real by claim 2 of Elementary Arithmetic in an Ordered Field .
We record the following estimate. Let p , q ∈ R d p,q\in\mathbb{R}^{d} p , q ∈ R d , let r r r be a nonnegative real, and let z ∈ R d + d z\in\mathbb{R}^{d+d} z ∈ R d + d satisfy ∥ p r 1 ( z ) − p ∥ ≤ r \lVert\mathrm{pr}_{1}(z)-p\rVert\le r ∥ pr 1 ( z ) − p ∥ ≤ r and ∥ p r 2 ( z ) − q ∥ ≤ r \lVert\mathrm{pr}_{2}(z)-q\rVert\le r ∥ pr 2 ( z ) − q ∥ ≤ r . Then
∣ c T ( z ) − d T ( p , q ) 2 ∣ ≤ Λ r . ( ⋆ ) \bigl|c_{\mathbb{T}}(z)-d_{\mathbb{T}}(p,q)^{2}\bigr|\le\Lambda\,r. \qquad(\star) c T ( z ) − d T ( p , q ) 2 ≤ Λ r . ( ⋆ )
Indeed, write a = p r 1 ( z ) a=\mathrm{pr}_{1}(z) a = pr 1 ( z ) and b = p r 2 ( z ) b=\mathrm{pr}_{2}(z) b = pr 2 ( z ) . By claim 5 of Properties of the Absolute Value in an Ordered Field the left side is at most ∣ d T ( a , b ) 2 − d T ( p , b ) 2 ∣ + ∣ d T ( p , b ) 2 − d T ( p , q ) 2 ∣ |d_{\mathbb{T}}(a,b)^{2}-d_{\mathbb{T}}(p,b)^{2}|+|d_{\mathbb{T}}(p,b)^{2}-d_{\mathbb{T}}(p,q)^{2}| ∣ d T ( a , b ) 2 − d T ( p , b ) 2 ∣ + ∣ d T ( p , b ) 2 − d T ( p , q ) 2 ∣ . The first term is at most d ∥ a − p ∥ \sqrt{d}\,\lVert a-p\rVert d ∥ a − p ∥ by the second inequality of The Flat Torus Distance: Minimality of the Wrapped Displacement, Periodicity, the Metric on the Unit Cell and the Lipschitz Bound §lipschitz . Since d T ( p , b ) = d T ( b , p ) d_{\mathbb{T}}(p,b)=d_{\mathbb{T}}(b,p) d T ( p , b ) = d T ( b , p ) and d T ( p , q ) = d T ( q , p ) d_{\mathbb{T}}(p,q)=d_{\mathbb{T}}(q,p) d T ( p , q ) = d T ( q , p ) by The Flat Torus Distance: Minimality of the Wrapped Displacement, Periodicity, the Metric on the Unit Cell and the Lipschitz Bound §symmetry , the second term equals ∣ d T ( b , p ) 2 − d T ( q , p ) 2 ∣ |d_{\mathbb{T}}(b,p)^{2}-d_{\mathbb{T}}(q,p)^{2}| ∣ d T ( b , p ) 2 − d T ( q , p ) 2 ∣ , which the same inequality, applied to the points b , q , p b,q,p b , q , p in place of x , x ′ , y x,x',y x , x ′ , y , bounds by d ∥ b − q ∥ \sqrt{d}\,\lVert b-q\rVert d ∥ b − q ∥ . As 0 ≤ d 0\le\sqrt{d} 0 ≤ d , claim 5 of Elementary Arithmetic in an Ordered Field turns the two hypotheses into d ∥ a − p ∥ + d ∥ b − q ∥ ≤ d r + d r = Λ r \sqrt{d}\,\lVert a-p\rVert+\sqrt{d}\,\lVert b-q\rVert\le\sqrt{d}\,r+\sqrt{d}\,r=\Lambda r d ∥ a − p ∥ + d ∥ b − q ∥ ≤ d r + d r = Λ r .
Suppose, for contradiction, that supp γ \operatorname{supp}\gamma supp γ is not torus-cyclically monotone. By Torus-Cyclically Monotone Subset of a Doubled Euclidean Space §monotone , and since the order of R \mathbb{R} R is total, there are N ∈ N N\in\mathbb{N} N ∈ N and z 1 , … , z N ∈ supp γ z_{1},\dots,z_{N}\in\operatorname{supp}\gamma z 1 , … , z N ∈ supp γ such that, writing x i = p r 1 ( z i ) x_{i}=\mathrm{pr}_{1}(z_{i}) x i = pr 1 ( z i ) and y i = p r 2 ( z i ) y_{i}=\mathrm{pr}_{2}(z_{i}) y i = pr 2 ( z i ) for i ∈ [ N ] i\in[N] i ∈ [ N ] and x N + 1 = x 1 x_{N+1}=x_{1} x N + 1 = x 1 ,
∑ i = 1 N d T ( x i + 1 , y i ) 2 < ∑ i = 1 N d T ( x i , y i ) 2 . \sum_{i=1}^{N}d_{\mathbb{T}}(x_{i+1},y_{i})^{2}<\sum_{i=1}^{N}d_{\mathbb{T}}(x_{i},y_{i})^{2}. i = 1 ∑ N d T ( x i + 1 , y i ) 2 < i = 1 ∑ N d T ( x i , y i ) 2 .
Let Δ \Delta Δ be the right side minus the left side, a positive real by claim 1 of Elementary Order Arithmetic in an Ordered Field . For i ∈ [ N ] i\in[N] i ∈ [ N ] write i − = i − 1 i^{-}=i-1 i − = i − 1 if 2 ≤ i 2\le i 2 ≤ i and 1 − = N 1^{-}=N 1 − = N ; the map i ↦ i − i\mapsto i^{-} i ↦ i − is a permutation of [ N ] [N] [ N ] , its inverse being i ↦ i + 1 i\mapsto i+1 i ↦ i + 1 for i < N i<N i < N and N ↦ 1 N\mapsto1 N ↦ 1 .
1. Reindexing. For every i ∈ [ N ] i\in[N] i ∈ [ N ] one has x ( i − ) + 1 = x i x_{(i^{-})+1}=x_{i} x ( i − ) + 1 = x i : for 2 ≤ i 2\le i 2 ≤ i because ( i − 1 ) + 1 = i (i-1)+1=i ( i − 1 ) + 1 = i , and for i = 1 i=1 i = 1 because N + 1 N+1 N + 1 is the index of x N + 1 = x 1 x_{N+1}=x_{1} x N + 1 = x 1 . Hence Invariance of Finite Sums and Products under Reindexing by a Permutation , applied to the family a i = d T ( x i + 1 , y i ) 2 a_{i}=d_{\mathbb{T}}(x_{i+1},y_{i})^{2} a i = d T ( x i + 1 , y i ) 2 and the permutation i ↦ i − i\mapsto i^{-} i ↦ i − , gives ∑ i = 1 N d T ( x i , y i − ) 2 = ∑ i = 1 N d T ( x i + 1 , y i ) 2 \sum_{i=1}^{N}d_{\mathbb{T}}(x_{i},y_{i^{-}})^{2}=\sum_{i=1}^{N}d_{\mathbb{T}}(x_{i+1},y_{i})^{2} ∑ i = 1 N d T ( x i , y i − ) 2 = ∑ i = 1 N d T ( x i + 1 , y i ) 2 , and therefore
∑ i = 1 N d T ( x i , y i − ) 2 − ∑ i = 1 N d T ( x i , y i ) 2 = − Δ . ( 1 ) \sum_{i=1}^{N}d_{\mathbb{T}}(x_{i},y_{i^{-}})^{2}-\sum_{i=1}^{N}d_{\mathbb{T}}(x_{i},y_{i})^{2}=-\Delta. \qquad(1) i = 1 ∑ N d T ( x i , y i − ) 2 − i = 1 ∑ N d T ( x i , y i ) 2 = − Δ. ( 1 )
2. Choice of the radius. Put C = Σ N ( Λ + Λ + 1 ) C=\Sigma_{N}\,(\Lambda+\Lambda+1) C = Σ N ( Λ + Λ + 1 ) . Since 0 ≤ Λ + Λ 0\le\Lambda+\Lambda 0 ≤ Λ + Λ (claim 2 of Elementary Arithmetic in an Ordered Field ), 1 ≤ Λ + Λ + 1 1\le\Lambda+\Lambda+1 1 ≤ Λ + Λ + 1 by claim 3 there, so 0 < Λ + Λ + 1 0<\Lambda+\Lambda+1 0 < Λ + Λ + 1 by claims 6 and 2 of Elementary Order Arithmetic in an Ordered Field , and 0 < C 0<C 0 < C by claim 5 there; hence C − 1 C^{-1} C − 1 exists and 0 < Δ C − 1 0<\Delta\,C^{-1} 0 < Δ C − 1 by claims 7 and 5 there. By claim 3 of The Archimedean Property of the Real Numbers , applied to Δ C − 1 \Delta\,C^{-1} Δ C − 1 , there is a positive real ε \varepsilon ε (the reciprocal of the image of a natural number) with ε < Δ C − 1 \varepsilon<\Delta\,C^{-1} ε < Δ C − 1 ; multiplying by the positive number C C C (claim 10 of Elementary Order Arithmetic in an Ordered Field ) gives
C ε < Δ . ( 2 ) C\,\varepsilon<\Delta. \qquad(2) C ε < Δ. ( 2 )
For i ∈ [ N ] i\in[N] i ∈ [ N ] put A i = B ( z i , ε ) ∈ B ( R d + d ) A_{i}=B(z_{i},\varepsilon)\in\mathcal{B}(\mathbb{R}^{d+d}) A i = B ( z i , ε ) ∈ B ( R d + d ) and m i = γ ( A i ) m_{i}=\gamma(A_{i}) m i = γ ( A i ) , a real number with 0 < m i ≤ 1 0<m_{i}\le1 0 < m i ≤ 1 by Support of a Borel Measure on a Metric Space §support and claim 2 of Basic Properties of a Measure . Let m m m be the least of m 1 , … , m N m_{1},\dots,m_{N} m 1 , … , m N and put θ = m Σ N − 1 \theta=m\,\Sigma_{N}^{-1} θ = m Σ N − 1 , a positive real.
3. The normalised pieces and their marginals. For i ∈ [ N ] i\in[N] i ∈ [ N ] let h i = m i − 1 1 A i h_{i}=m_{i}^{-1}\mathbf{1}_{A_{i}} h i = m i − 1 1 A i , a measurable function from R d + d \mathbb{R}^{d+d} R d + d to [ 0 , ∞ ) [0,\infty) [ 0 , ∞ ) , and let γ i \gamma_{i} γ i be the measure with density h i h_{i} h i with respect to γ \gamma γ , as in claim 3 of that lemma. For B ∈ B ( R d + d ) B\in\mathcal{B}(\mathbb{R}^{d+d}) B ∈ B ( R d + d ) the product 1 B h i \mathbf{1}_{B}h_{i} 1 B h i is m i − 1 1 A i ∩ B m_{i}^{-1}\mathbf{1}_{A_{i}\cap B} m i − 1 1 A i ∩ B , so claim 1 of Linearity and Monotonicity of the Lebesgue Integral and The Integral of an Indicator Function is the Measure of the Set give
γ i ( B ) = m i − 1 γ ( A i ∩ B ) . \gamma_{i}(B)=m_{i}^{-1}\,\gamma(A_{i}\cap B). γ i ( B ) = m i − 1 γ ( A i ∩ B ) .
In particular γ i ∈ P ( R d + d ) \gamma_{i}\in\mathcal{P}(\mathbb{R}^{d+d}) γ i ∈ P ( R d + d ) and γ i ( R d + d ∖ A i ) = 0 \gamma_{i}(\mathbb{R}^{d+d}\setminus A_{i})=0 γ i ( R d + d ∖ A i ) = 0 . Put α i = ( p r 1 ) # γ i \alpha_{i}=(\mathrm{pr}_{1})_{\#}\gamma_{i} α i = ( pr 1 ) # γ i and β i = ( p r 2 ) # γ i \beta_{i}=(\mathrm{pr}_{2})_{\#}\gamma_{i} β i = ( pr 2 ) # γ i , probability measures on R d \mathbb{R}^{d} R d by claim 1 of Image Measures, Measures with Densities, and Change of Variables .
If z ∈ A i z\in A_{i} z ∈ A i then ∥ z − z i ∥ = d E ( z , z i ) < ε \lVert z-z_{i}\rVert=d_{E}(z,z_{i})<\varepsilon ∥ z − z i ∥ = d E ( z , z i ) < ε , by claim 2 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n and the symmetry of the metric d E d_{E} d E ; and p r 1 ( z ) − x i = p r 1 ( z − z i ) \mathrm{pr}_{1}(z)-x_{i}=\mathrm{pr}_{1}(z-z_{i}) pr 1 ( z ) − x i = pr 1 ( z − z i ) , because by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections both points are concatenations of their projections and concatenation is compatible with differences by claim 2 of Concatenation Identifies a Product of Euclidean Spaces with a Euclidean Space ; so the norm bound of Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections gives ∥ p r 1 ( z ) − x i ∥ ≤ ∥ z − z i ∥ < ε \lVert\mathrm{pr}_{1}(z)-x_{i}\rVert\le\lVert z-z_{i}\rVert<\varepsilon ∥ pr 1 ( z ) − x i ∥ ≤ ∥ z − z i ∥ < ε , and likewise ∥ p r 2 ( z ) − y i ∥ < ε \lVert\mathrm{pr}_{2}(z)-y_{i}\rVert<\varepsilon ∥ pr 2 ( z ) − y i ∥ < ε .
For i ∈ [ N ] i\in[N] i ∈ [ N ] let U i = { x ∈ R d : ε 2 < ∥ x − x i ∥ 2 } U_{i}=\{x\in\mathbb{R}^{d}:\varepsilon^{2}<\lVert x-x_{i}\rVert^{2}\} U i = { x ∈ R d : ε 2 < ∥ x − x i ∥ 2 } and V i = { y ∈ R d : ε 2 < ∥ y − y i ∥ 2 } V_{i}=\{y\in\mathbb{R}^{d}:\varepsilon^{2}<\lVert y-y_{i}\rVert^{2}\} V i = { y ∈ R d : ε 2 < ∥ y − y i ∥ 2 } . They are Borel: by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions , applied to the Borel identity map and a constant map, x ↦ ∥ x − x i ∥ 2 x\mapsto\lVert x-x_{i}\rVert^{2} x ↦ ∥ x − x i ∥ 2 is Borel, and the criterion of Measure Spaces and the Lebesgue Integral: Standing Notation §measurable applies. By claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field , x ∈ U i x\in U_{i} x ∈ U i exactly when ε < ∥ x − x i ∥ \varepsilon<\lVert x-x_{i}\rVert ε < ∥ x − x i ∥ , so by the previous paragraph p r 1 − 1 ( U i ) ⊆ R d + d ∖ A i \mathrm{pr}_{1}^{-1}(U_{i})\subseteq\mathbb{R}^{d+d}\setminus A_{i} pr 1 − 1 ( U i ) ⊆ R d + d ∖ A i and
α i ( U i ) = γ i ( p r 1 − 1 ( U i ) ) ≤ γ i ( R d + d ∖ A i ) = 0 \alpha_{i}(U_{i})=\gamma_{i}\bigl(\mathrm{pr}_{1}^{-1}(U_{i})\bigr)\le\gamma_{i}(\mathbb{R}^{d+d}\setminus A_{i})=0 α i ( U i ) = γ i ( pr 1 − 1 ( U i ) ) ≤ γ i ( R d + d ∖ A i ) = 0
by claim 2 of Basic Properties of a Measure ; likewise β i ( V i ) = 0 \beta_{i}(V_{i})=0 β i ( V i ) = 0 .
4. The competitor. Let g = 1 − θ ∑ i = 1 N h i g=1-\theta\sum_{i=1}^{N}h_{i} g = 1 − θ ∑ i = 1 N h i . For every z z z one has m i − 1 ≤ m − 1 m_{i}^{-1}\le m^{-1} m i − 1 ≤ m − 1 , since 0 < m ≤ m i 0<m\le m_{i} 0 < m ≤ m i and inversion reverses the order on the positive reals by claims 7 and 10 of Elementary Order Arithmetic in an Ordered Field , so h i ( z ) ≤ m − 1 h_{i}(z)\le m^{-1} h i ( z ) ≤ m − 1 by claim 5 of Elementary Arithmetic in an Ordered Field ; summing the resulting nonnegative differences m − 1 − h i ( z ) m^{-1}-h_{i}(z) m − 1 − h i ( z ) with claims 2, 3 and 5 of Properties of Finite Sums gives θ ∑ i = 1 N h i ( z ) ≤ θ Σ N m − 1 = 1 \theta\sum_{i=1}^{N}h_{i}(z)\le\theta\,\Sigma_{N}\,m^{-1}=1 θ ∑ i = 1 N h i ( z ) ≤ θ Σ N m − 1 = 1 ; thus g g g is a measurable function from R d + d \mathbb{R}^{d+d} R d + d to [ 0 , ∞ ) [0,\infty) [ 0 , ∞ ) , and the measure ω \omega ω with density g g g with respect to γ \gamma γ satisfies, by the computation of step 3,
ω ( B ) = γ ( B ) − θ ∑ i = 1 N γ i ( B ) ( B ∈ B ( R d + d ) ) . \omega(B)=\gamma(B)-\theta\sum_{i=1}^{N}\gamma_{i}(B)\qquad(B\in\mathcal{B}(\mathbb{R}^{d+d})). ω ( B ) = γ ( B ) − θ i = 1 ∑ N γ i ( B ) ( B ∈ B ( R d + d )) .
For i ∈ [ N ] i\in[N] i ∈ [ N ] let σ i \sigma_{i} σ i be the measure with constant density θ \theta θ with respect to the product measure α i ⊠ β i − \alpha_{i}\boxtimes\beta_{i^{-}} α i ⊠ β i − of Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §product , so that σ i ( B ) = θ ( α i ⊠ β i − ) ( B ) \sigma_{i}(B)=\theta\,(\alpha_{i}\boxtimes\beta_{i^{-}})(B) σ i ( B ) = θ ( α i ⊠ β i − ) ( B ) by the same computation, and put σ N + 1 = ω \sigma_{N+1}=\omega σ N + 1 = ω .
For i ∈ [ N + 1 ] i\in[N+1] i ∈ [ N + 1 ] let φ i \varphi_{i} φ i be the bijection z ↦ ( z , i ) z\mapsto(z,i) z ↦ ( z , i ) of R d + d \mathbb{R}^{d+d} R d + d onto X i = R d + d × { i } X_{i}=\mathbb{R}^{d+d}\times\{i\} X i = R d + d × { i } and let ( X i , F i , ρ i ) (X_{i},\mathcal{F}_{i},\rho_{i}) ( X i , F i , ρ i ) be the transport of ( R d + d , B ( R d + d ) , σ i ) (\mathbb{R}^{d+d},\mathcal{B}(\mathbb{R}^{d+d}),\sigma_{i}) ( R d + d , B ( R d + d ) , σ i ) along φ i \varphi_{i} φ i , as in claim 2 of Assembly of Measure Spaces: Restriction, Transport, One-Point Spaces, and Countable Disjoint Unions . The sets X i X_{i} X i are pairwise disjoint, so claim 4 of that lemma provides their countable disjoint union ( X ⊔ , F ⊔ , ρ ⊔ ) (X_{\sqcup},\mathcal{F}_{\sqcup},\rho_{\sqcup}) ( X ⊔ , F ⊔ , ρ ⊔ ) , and the map Φ : X ⊔ → R d + d \Phi:X_{\sqcup}\to\mathbb{R}^{d+d} Φ : X ⊔ → R d + d with Φ ( ( z , i ) ) = z \Phi((z,i))=z Φ (( z , i )) = z is measurable by claim 4(b) there. Let γ ~ = Φ # ρ ⊔ \tilde\gamma=\Phi_{\#}\rho_{\sqcup} γ ~ = Φ # ρ ⊔ be its image measure ; since Φ − 1 ( B ) \Phi^{-1}(B) Φ − 1 ( B ) meets X i X_{i} X i in φ i ( B ) \varphi_{i}(B) φ i ( B ) , whose ρ i \rho_{i} ρ i -measure is σ i ( B ) \sigma_{i}(B) σ i ( B ) by claim 2 of Assembly of Measure Spaces: Restriction, Transport, One-Point Spaces, and Countable Disjoint Unions , claim 4(a) there gives
γ ~ ( B ) = ∑ i = 1 N + 1 σ i ( B ) = γ ( B ) − θ ∑ i = 1 N γ i ( B ) + θ ∑ i = 1 N ( α i ⊠ β i − ) ( B ) . \tilde\gamma(B)=\sum_{i=1}^{N+1}\sigma_{i}(B)=\gamma(B)-\theta\sum_{i=1}^{N}\gamma_{i}(B)+\theta\sum_{i=1}^{N}(\alpha_{i}\boxtimes\beta_{i^{-}})(B). γ ~ ( B ) = i = 1 ∑ N + 1 σ i ( B ) = γ ( B ) − θ i = 1 ∑ N γ i ( B ) + θ i = 1 ∑ N ( α i ⊠ β i − ) ( B ) .
5. The competitor is a coupling. Taking B = R d + d B=\mathbb{R}^{d+d} B = R d + d gives γ ~ ( R d + d ) = 1 − θ Σ N + θ Σ N = 1 \tilde\gamma(\mathbb{R}^{d+d})=1-\theta\,\Sigma_{N}+\theta\,\Sigma_{N}=1 γ ~ ( R d + d ) = 1 − θ Σ N + θ Σ N = 1 . For A ∈ B ( R d ) A\in\mathcal{B}(\mathbb{R}^{d}) A ∈ B ( R d ) , using γ i ( p r 1 − 1 ( A ) ) = α i ( A ) \gamma_{i}(\mathrm{pr}_{1}^{-1}(A))=\alpha_{i}(A) γ i ( pr 1 − 1 ( A )) = α i ( A ) and ( α i ⊠ β i − ) ( p r 1 − 1 ( A ) ) = α i ( A ) (\alpha_{i}\boxtimes\beta_{i^{-}})(\mathrm{pr}_{1}^{-1}(A))=\alpha_{i}(A) ( α i ⊠ β i − ) ( pr 1 − 1 ( A )) = α i ( A ) from Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §product ,
γ ~ ( p r 1 − 1 ( A ) ) = γ ( p r 1 − 1 ( A ) ) = μ ( A ) , \tilde\gamma\bigl(\mathrm{pr}_{1}^{-1}(A)\bigr)=\gamma\bigl(\mathrm{pr}_{1}^{-1}(A)\bigr)=\mu(A), γ ~ ( pr 1 − 1 ( A ) ) = γ ( pr 1 − 1 ( A ) ) = μ ( A ) ,
and for B ∈ B ( R d ) B\in\mathcal{B}(\mathbb{R}^{d}) B ∈ B ( R d ) , using ∑ i = 1 N β i − ( B ) = ∑ i = 1 N β i ( B ) \sum_{i=1}^{N}\beta_{i^{-}}(B)=\sum_{i=1}^{N}\beta_{i}(B) ∑ i = 1 N β i − ( B ) = ∑ i = 1 N β i ( B ) by Invariance of Finite Sums and Products under Reindexing by a Permutation ,
γ ~ ( p r 2 − 1 ( B ) ) = γ ( p r 2 − 1 ( B ) ) = ν ( B ) . \tilde\gamma\bigl(\mathrm{pr}_{2}^{-1}(B)\bigr)=\gamma\bigl(\mathrm{pr}_{2}^{-1}(B)\bigr)=\nu(B). γ ~ ( pr 2 − 1 ( B ) ) = γ ( pr 2 − 1 ( B ) ) = ν ( B ) .
Hence γ ~ ∈ Π ( μ , ν ) \tilde\gamma\in\Pi(\mu,\nu) γ ~ ∈ Π ( μ , ν ) by Couplings of Two Probability Measures on Euclidean Space and Their Quadratic Cost §coupling , and its torus cost I T ( γ ~ ) I_{\mathbb{T}}(\tilde\gamma) I T ( γ ~ ) is defined by Probability Measures on the Flat Torus, the Torus Cost of a Coupling, the Torus Wasserstein Distance and Optimal Couplings §cost .
6. The cost of the competitor. Let f : R d + d → [ 0 , ∞ ] f:\mathbb{R}^{d+d}\to[0,\infty] f : R d + d → [ 0 , ∞ ] be Borel. By claim 2 of Image Measures, Measures with Densities, and Change of Variables applied to Φ \Phi Φ , ∫ f d γ ~ = ∫ X ⊔ f ∘ Φ d ρ ⊔ \int f\,d\tilde\gamma=\int_{X_{\sqcup}}f\circ\Phi\,d\rho_{\sqcup} ∫ f d γ ~ = ∫ X ⊔ f ∘ Φ d ρ ⊔ ; the function f ∘ Φ f\circ\Phi f ∘ Φ is measurable, and claim 4(c) of Assembly of Measure Spaces: Restriction, Transport, One-Point Spaces, and Countable Disjoint Unions splits this integral as the sum over i ∈ [ N + 1 ] i\in[N+1] i ∈ [ N + 1 ] of ∫ X i ( f ∘ Φ ) ∣ X i d ρ i \int_{X_{i}}(f\circ\Phi)|_{X_{i}}\,d\rho_{i} ∫ X i ( f ∘ Φ ) ∣ X i d ρ i ; since ρ i \rho_{i} ρ i is the image measure of σ i \sigma_{i} σ i under φ i \varphi_{i} φ i (claim 2 of Assembly of Measure Spaces: Restriction, Transport, One-Point Spaces, and Countable Disjoint Unions ) and ( f ∘ Φ ) ∘ φ i = f (f\circ\Phi)\circ\varphi_{i}=f ( f ∘ Φ ) ∘ φ i = f , claim 2 of Image Measures, Measures with Densities, and Change of Variables identifies the i i i th summand with ∫ f d σ i \int f\,d\sigma_{i} ∫ f d σ i . Thus
∫ f d γ ~ = ∑ i = 1 N + 1 ∫ f d σ i . \int f\,d\tilde\gamma=\sum_{i=1}^{N+1}\int f\,d\sigma_{i}. ∫ f d γ ~ = i = 1 ∑ N + 1 ∫ f d σ i .
Take f = c T f=c_{\mathbb{T}} f = c T . By claim 3 of Image Measures, Measures with Densities, and Change of Variables and claim 1 of Linearity and Monotonicity of the Lebesgue Integral , ∫ c T d σ i = θ ∫ c T d ( α i ⊠ β i − ) \int c_{\mathbb{T}}\,d\sigma_{i}=\theta\int c_{\mathbb{T}}\,d(\alpha_{i}\boxtimes\beta_{i^{-}}) ∫ c T d σ i = θ ∫ c T d ( α i ⊠ β i − ) for i ∈ [ N ] i\in[N] i ∈ [ N ] , ∫ c T d ω = ∫ c T g d γ \int c_{\mathbb{T}}\,d\omega=\int c_{\mathbb{T}}\,g\,d\gamma ∫ c T d ω = ∫ c T g d γ , and ∫ c T h i d γ = ∫ c T d γ i \int c_{\mathbb{T}}\,h_{i}\,d\gamma=\int c_{\mathbb{T}}\,d\gamma_{i} ∫ c T h i d γ = ∫ c T d γ i . Since c T = c T g + ∑ i = 1 N θ c T h i c_{\mathbb{T}}=c_{\mathbb{T}}\,g+\sum_{i=1}^{N}\theta\,c_{\mathbb{T}}\,h_{i} c T = c T g + ∑ i = 1 N θ c T h i pointwise, claim 1 of Linearity and Monotonicity of the Lebesgue Integral gives I T ( γ ) = ∫ c T d ω + θ ∑ i = 1 N ∫ c T d γ i I_{\mathbb{T}}(\gamma)=\int c_{\mathbb{T}}\,d\omega+\theta\sum_{i=1}^{N}\int c_{\mathbb{T}}\,d\gamma_{i} I T ( γ ) = ∫ c T d ω + θ ∑ i = 1 N ∫ c T d γ i . All integrals here are real numbers by step 0, so
I T ( γ ~ ) = I T ( γ ) − θ ∑ i = 1 N ∫ c T d γ i + θ ∑ i = 1 N ∫ c T d ( α i ⊠ β i − ) . ( 3 ) I_{\mathbb{T}}(\tilde\gamma)=I_{\mathbb{T}}(\gamma)-\theta\sum_{i=1}^{N}\int c_{\mathbb{T}}\,d\gamma_{i}+\theta\sum_{i=1}^{N}\int c_{\mathbb{T}}\,d(\alpha_{i}\boxtimes\beta_{i^{-}}). \qquad(3) I T ( γ ~ ) = I T ( γ ) − θ i = 1 ∑ N ∫ c T d γ i + θ i = 1 ∑ N ∫ c T d ( α i ⊠ β i − ) . ( 3 )
We estimate the two families of integrals. Let i ∈ [ N ] i\in[N] i ∈ [ N ] . For z ∈ A i z\in A_{i} z ∈ A i , step 3 and ( ⋆ ) (\star) ( ⋆ ) with ( p , q ) = ( x i , y i ) (p,q)=(x_{i},y_{i}) ( p , q ) = ( x i , y i ) and r = ε r=\varepsilon r = ε give d T ( x i , y i ) 2 ≤ c T ( z ) + Λ ε d_{\mathbb{T}}(x_{i},y_{i})^{2}\le c_{\mathbb{T}}(z)+\Lambda\varepsilon d T ( x i , y i ) 2 ≤ c T ( z ) + Λ ε (claim 6 of Properties of the Absolute Value in an Ordered Field ); as γ i ( R d + d ∖ A i ) = 0 \gamma_{i}(\mathbb{R}^{d+d}\setminus A_{i})=0 γ i ( R d + d ∖ A i ) = 0 , this holds for γ i \gamma_{i} γ i -almost every z z z . Next, ( α i ⊠ β i − ) ( p r 1 − 1 ( U i ) ) = α i ( U i ) = 0 (\alpha_{i}\boxtimes\beta_{i^{-}})(\mathrm{pr}_{1}^{-1}(U_{i}))=\alpha_{i}(U_{i})=0 ( α i ⊠ β i − ) ( pr 1 − 1 ( U i )) = α i ( U i ) = 0 and ( α i ⊠ β i − ) ( p r 2 − 1 ( V i − ) ) = β i − ( V i − ) = 0 (\alpha_{i}\boxtimes\beta_{i^{-}})(\mathrm{pr}_{2}^{-1}(V_{i^{-}}))=\beta_{i^{-}}(V_{i^{-}})=0 ( α i ⊠ β i − ) ( pr 2 − 1 ( V i − )) = β i − ( V i − ) = 0 by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §product and step 3, so by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §null-union almost every z z z for α i ⊠ β i − \alpha_{i}\boxtimes\beta_{i^{-}} α i ⊠ β i − lies outside both sets; for such z z z the order being total and claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field give ∥ p r 1 ( z ) − x i ∥ ≤ ε \lVert\mathrm{pr}_{1}(z)-x_{i}\rVert\le\varepsilon ∥ pr 1 ( z ) − x i ∥ ≤ ε and ∥ p r 2 ( z ) − y i − ∥ ≤ ε \lVert\mathrm{pr}_{2}(z)-y_{i^{-}}\rVert\le\varepsilon ∥ pr 2 ( z ) − y i − ∥ ≤ ε , and ( ⋆ ) (\star) ( ⋆ ) with ( p , q ) = ( x i , y i − ) (p,q)=(x_{i},y_{i^{-}}) ( p , q ) = ( x i , y i − ) gives c T ( z ) ≤ d T ( x i , y i − ) 2 + Λ ε c_{\mathbb{T}}(z)\le d_{\mathbb{T}}(x_{i},y_{i^{-}})^{2}+\Lambda\varepsilon c T ( z ) ≤ d T ( x i , y i − ) 2 + Λ ε . Both sides of each of these two inequalities are nonnegative Borel functions of z z z , the constants having integral equal to themselves against a probability measure by The Integral of an Indicator Function is the Measure of the Set and claim 1 of Linearity and Monotonicity of the Lebesgue Integral ; so The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §comparison and claim 1 of Linearity and Monotonicity of the Lebesgue Integral give
d T ( x i , y i ) 2 − Λ ε ≤ ∫ c T d γ i , ∫ c T d ( α i ⊠ β i − ) ≤ d T ( x i , y i − ) 2 + Λ ε . d_{\mathbb{T}}(x_{i},y_{i})^{2}-\Lambda\varepsilon\le\int c_{\mathbb{T}}\,d\gamma_{i},\qquad\int c_{\mathbb{T}}\,d(\alpha_{i}\boxtimes\beta_{i^{-}})\le d_{\mathbb{T}}(x_{i},y_{i^{-}})^{2}+\Lambda\varepsilon . d T ( x i , y i ) 2 − Λ ε ≤ ∫ c T d γ i , ∫ c T d ( α i ⊠ β i − ) ≤ d T ( x i , y i − ) 2 + Λ ε .
Summing over i ∈ [ N ] i\in[N] i ∈ [ N ] with claims 2, 3 and 5 of Properties of Finite Sums , multiplying by θ \theta θ (claim 5 of Elementary Arithmetic in an Ordered Field ) and inserting into (3), then using (1),
I T ( γ ~ ) − I T ( γ ) ≤ θ ( ∑ i = 1 N d T ( x i , y i − ) 2 − ∑ i = 1 N d T ( x i , y i ) 2 ) + θ Σ N ( Λ + Λ ) ε = − θ Δ + θ Σ N ( Λ + Λ ) ε , I_{\mathbb{T}}(\tilde\gamma)-I_{\mathbb{T}}(\gamma)\le\theta\Bigl(\sum_{i=1}^{N}d_{\mathbb{T}}(x_{i},y_{i^{-}})^{2}-\sum_{i=1}^{N}d_{\mathbb{T}}(x_{i},y_{i})^{2}\Bigr)+\theta\,\Sigma_{N}\,(\Lambda+\Lambda)\,\varepsilon=-\theta\Delta+\theta\,\Sigma_{N}\,(\Lambda+\Lambda)\,\varepsilon , I T ( γ ~ ) − I T ( γ ) ≤ θ ( i = 1 ∑ N d T ( x i , y i − ) 2 − i = 1 ∑ N d T ( x i , y i ) 2 ) + θ Σ N ( Λ + Λ ) ε = − θ Δ + θ Σ N ( Λ + Λ ) ε ,
the second term accounting for the 2 N 2N 2 N integrals, each of which deviates from the corresponding squared torus distance by at most Λ ε \Lambda\varepsilon Λ ε .
7. The contradiction. Since 0 ≤ Σ N ε 0\le\Sigma_{N}\,\varepsilon 0 ≤ Σ N ε , claim 5 of Elementary Arithmetic in an Ordered Field gives Σ N ( Λ + Λ ) ε ≤ Σ N ( Λ + Λ + 1 ) ε = C ε \Sigma_{N}(\Lambda+\Lambda)\varepsilon\le\Sigma_{N}(\Lambda+\Lambda+1)\varepsilon=C\varepsilon Σ N ( Λ + Λ ) ε ≤ Σ N ( Λ + Λ + 1 ) ε = Cε , and C ε < Δ C\varepsilon<\Delta Cε < Δ by (2); multiplying by the positive number θ \theta θ (claim 10 of Elementary Order Arithmetic in an Ordered Field and claim 5 of Elementary Arithmetic in an Ordered Field ) and combining with claims 1 and 2 of Elementary Order Arithmetic in an Ordered Field ,
I T ( γ ~ ) − I T ( γ ) < − θ Δ + θ Δ = 0. I_{\mathbb{T}}(\tilde\gamma)-I_{\mathbb{T}}(\gamma)<-\theta\Delta+\theta\Delta=0 . I T ( γ ~ ) − I T ( γ ) < − θ Δ + θ Δ = 0.
Thus I T ( γ ~ ) < I T ( γ ) I_{\mathbb{T}}(\tilde\gamma)<I_{\mathbb{T}}(\gamma) I T ( γ ~ ) < I T ( γ ) , and I T ( γ ) = W T ( μ , ν ) 2 I_{\mathbb{T}}(\gamma)=W_{\mathbb{T}}(\mu,\nu)^{2} I T ( γ ) = W T ( μ , ν ) 2 because γ \gamma γ is optimal (Probability Measures on the Flat Torus, the Torus Cost of a Coupling, the Torus Wasserstein Distance and Optimal Couplings §optimal ). This contradicts W T ( μ , ν ) 2 ≤ I T ( γ ~ ) W_{\mathbb{T}}(\mu,\nu)^{2}\le I_{\mathbb{T}}(\tilde\gamma) W T ( μ , ν ) 2 ≤ I T ( γ ~ ) , which holds by Probability Measures on the Flat Torus, the Torus Cost of a Coupling, the Torus Wasserstein Distance and Optimal Couplings §distance because W T ( μ , ν ) 2 W_{\mathbb{T}}(\mu,\nu)^{2} W T ( μ , ν ) 2 is the greatest lower bound of the torus costs of the members of Π ( μ , ν ) \Pi(\mu,\nu) Π ( μ , ν ) and γ ~ ∈ Π ( μ , ν ) \tilde\gamma\in\Pi(\mu,\nu) γ ~ ∈ Π ( μ , ν ) by step 5.
Therefore supp γ \operatorname{supp}\gamma supp γ is torus-cyclically monotone, which is claim 1 of the statement.