TheoremBase

Clause 1 compares with the competitor coupling (pr2, T o pr1)#gamma using minimality of the wrapped displacement; clause 2 gets the lower bound by contradiction, gluing near-optimal couplings and applying the composite-displacement convergence lemma; clause 3 mollifies the periodic potential, identifies the mollified gradient near the cell through Lipschitz truncations of the convex parts, and passes to the limit by dominated convergence.

Proof

Throughout, Φ=Φν\Phi=\Phi_{\nu}, ∥⋅∥\lVert\cdot\rVert is the Euclidean norm and x⋅yx\cdot y the dot product of Rd\mathbb{R}^{d}. We use the following facts repeatedly.

(F1) Wrapped displacements. For z∈Rdz\in\mathbb{R}^{d} and k∈Zdk\in\mathbb{Z}^{d} one has z−ϖ(z)∈Zdz-\varpi(z)\in\mathbb{Z}^{d} and ∥ϖ(z)∥2≤d/4\lVert\varpi(z)\rVert^{2}\le d/4 by The Flat Torus Distance: Minimality of the Wrapped Displacement, Periodicity, the Metric on the Unit Cell and the Lipschitz Bound §range, and ∥ϖ(z)∥≤∥z−k∥\lVert\varpi(z)\rVert\le\lVert z-k\rVert by The Flat Torus Distance: Minimality of the Wrapped Displacement, Periodicity, the Metric on the Unit Cell and the Lipschitz Bound §minimal. For x,y∈Rdx,y\in\mathbb{R}^{d}, dT(x,y)=∥ϖ(y−x)∥d_{\mathbb{T}}(x,y)=\lVert\varpi(y-x)\rVert by the definition of the flat torus distance. Sums and differences of points of Zd\mathbb{Z}^{d} lie in Zd\mathbb{Z}^{d}, their coordinates being sums and differences of integers (Lattice-Periodic Functions and the Periodic Function Classes §lattice). The map ϖ\varpi is Borel by The Flat Torus Distance: Minimality of the Wrapped Displacement, Periodicity, the Metric on the Unit Cell and the Lipschitz Bound §lipschitz, and a difference of two Borel maps into Rd\mathbb{R}^{d} is Borel, componentwise, by claim 2 of The Borel Sigma-Algebra of a Euclidean Space as a Product, and Measurability of Projections, Sequentially Continuous Maps, and Open and Closed Sets and claim 2 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions (exactly as in Optimal Maps on the Torus, Uniquely Mapped Pairs and the Displacement Field of a Map §displacement); likewise for sums.

(F2) Pairings and push-forwards. For m∈Nm\in\mathbb{N} with 1≤m1\le m, Borel maps u,v:Rm→Rdu,v:\mathbb{R}^{m}\to\mathbb{R}^{d} and λ∈P(Rm)\lambda\in\mathcal{P}(\mathbb{R}^{m}), the pairing (u,v)(u,v) is Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §pairing and satisfies pr1∘(u,v)=u\mathrm{pr}_{1}\circ(u,v)=u, pr2∘(u,v)=v\mathrm{pr}_{2}\circ(u,v)=v by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections. Consequently, for B∈B(Rd)B\in\mathcal{B}(\mathbb{R}^{d}), u#λ(B)=λ((u,v)−1(pr1−1(B)))=(pr1)#((u,v)#λ)(B)u_{\#}\lambda(B)=\lambda\bigl((u,v)^{-1}(\mathrm{pr}_{1}^{-1}(B))\bigr)=(\mathrm{pr}_{1})_{\#}\bigl((u,v)_{\#}\lambda\bigr)(B), and similarly v#λ=(pr2)#((u,v)#λ)v_{\#}\lambda=(\mathrm{pr}_{2})_{\#}((u,v)_{\#}\lambda), by the definition of the push-forward. The change-of-variables formula ∫g d(F#λ)=∫g∘F dλ\int g\,d(F_{\#}\lambda)=\int g\circ F\,d\lambda is used for Borel FF and bounded Borel gg; bounded Borel functions are integrable against every probability measure by Probability Measures on Euclidean Space and Random Vectors: Standing Notation §measures, and integrals of such functions are combined by linearity and monotonicity, claim 2 of Linearity and Monotonicity of the Lebesgue Integral. Norms, squared norms and dot products of Borel maps into Rd\mathbb{R}^{d} are Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions, and ∣x⋅y∣≤∥x∥ ∥y∥|x\cdot y|\le\lVert x\rVert\,\lVert y\rVert by the same clause.

(F3) The distance as an infimum. For μ1,μ2∈P(Td)\mu_{1},\mu_{2}\in\mathcal{P}(\mathbb{T}^{d}) and κ∈Π(μ1,μ2)\kappa\in\Pi(\mu_{1},\mu_{2}), WT(μ1,μ2)2≤IT(κ)W_{\mathbb{T}}(\mu_{1},\mu_{2})^{2}\le I_{\mathbb{T}}(\kappa), since WT(μ1,μ2)2W_{\mathbb{T}}(\mu_{1},\mu_{2})^{2} is the infimum of the torus costs by Probability Measures on the Flat Torus, the Torus Cost of a Coupling, the Torus Wasserstein Distance and Optimal Couplings §distance; and equality holds when κ\kappa is optimal, by Probability Measures on the Flat Torus, the Torus Cost of a Coupling, the Torus Wasserstein Distance and Optimal Couplings §optimal.

Proof of clause 1. Let μ∈P(Td)\mu\in\mathcal{P}(\mathbb{T}^{d}), let TT be an optimal map from μ\mu to ν\nu, and fix the Borel map vTv_{T} of Optimal Maps on the Torus, Uniquely Mapped Pairs and the Displacement Field of a Map §displacement as the representative of its class; ∥vT(x)∥≤d/2\lVert v_{T}(x)\rVert\le\sqrt d/2 for every xx by (F1).

Step 1.1 (the value at μ\mu). By Optimal Maps on the Torus, Uniquely Mapped Pairs and the Displacement Field of a Map §map, T#μ=νT_{\#}\mu=\nu and (id,T)#μ∈Π(μ,ν)(\mathrm{id},T)_{\#}\mu\in\Pi(\mu,\nu) is optimal, so by (F3), the definition of the torus cost, (F2) and (F1),

WT(μ,ν)2=IT((id,T)#μ)=∫RddT(x,T(x))2 μ(dx)=∫Rd∥ϖ(T(x)−x)∥2 μ(dx)=∫Rd∥vT∥2 dμ=∥vT∥μ2,W_{\mathbb{T}}(\mu,\nu)^{2}=I_{\mathbb{T}}\bigl((\mathrm{id},T)_{\#}\mu\bigr)=\int_{\mathbb{R}^{d}}d_{\mathbb{T}}\bigl(x,T(x)\bigr)^{2}\,\mu(dx)=\int_{\mathbb{R}^{d}}\bigl\lVert\varpi\bigl(T(x)-x\bigr)\bigr\rVert^{2}\,\mu(dx)=\int_{\mathbb{R}^{d}}\lVert v_{T}\rVert^{2}\,d\mu=\lVert v_{T}\rVert_{\mu}^{2},

the last equality being the definition of the norm in Optimal Transport on the Flat Torus: Standing Notation §fields. Hence Φ(μ)=12∥vT∥μ2\Phi(\mu)=\tfrac12\lVert v_{T}\rVert_{\mu}^{2}.

Step 1.2 (a competitor coupling). Let μ′∈P(Td)\mu'\in\mathcal{P}(\mathbb{T}^{d}) and γ∈Π(μ,μ′)\gamma\in\Pi(\mu,\mu'). The map S=(pr2,T∘pr1):Rd+d→Rd+dS=(\mathrm{pr}_{2},T\circ\mathrm{pr}_{1}):\mathbb{R}^{d+d}\to\mathbb{R}^{d+d} is Borel by (F2), T∘pr1T\circ\mathrm{pr}_{1} being a composite of Borel maps (Probability Measures on Euclidean Space and Random Vectors: Standing Notation §borel-maps). Put σ=S#γ\sigma=S_{\#}\gamma. By (F2), (pr1)#σ=(pr2)#γ=μ′(\mathrm{pr}_{1})_{\#}\sigma=(\mathrm{pr}_{2})_{\#}\gamma=\mu', and for B∈B(Rd)B\in\mathcal{B}(\mathbb{R}^{d}), (pr2)#σ(B)=γ(pr1−1(T−1(B)))=μ(T−1(B))=ν(B)(\mathrm{pr}_{2})_{\#}\sigma(B)=\gamma\bigl(\mathrm{pr}_{1}^{-1}(T^{-1}(B))\bigr)=\mu(T^{-1}(B))=\nu(B). Thus σ∈Π(μ′,ν)\sigma\in\Pi(\mu',\nu), and by (F3) and the change-of-variables formula,

2Φ(μ′)=WT(μ′,ν)2≤IT(σ)=∫Rd+ddT(pr2(w),T(pr1(w)))2 γ(dw).2\Phi(\mu')=W_{\mathbb{T}}(\mu',\nu)^{2}\le I_{\mathbb{T}}(\sigma)=\int_{\mathbb{R}^{d+d}}d_{\mathbb{T}}\bigl(\mathrm{pr}_{2}(w),T(\mathrm{pr}_{1}(w))\bigr)^{2}\,\gamma(dw).

Step 1.3 (pointwise bound). Let w∈Rd+dw\in\mathbb{R}^{d+d} and write x=pr1(w)x=\mathrm{pr}_{1}(w), x′=pr2(w)x'=\mathrm{pr}_{2}(w). By (F1) the point k=(T(x)−x−vT(x))−(x′−x−ϖ(x′−x))k=\bigl(T(x)-x-v_{T}(x)\bigr)-\bigl(x'-x-\varpi(x'-x)\bigr) lies in Zd\mathbb{Z}^{d}, since vT(x)=ϖ(T(x)−x)v_{T}(x)=\varpi(T(x)-x); and T(x)−x′−k=vT(x)−ϖ(x′−x)T(x)-x'-k=v_{T}(x)-\varpi(x'-x). Hence, by (F1) and bilinearity of the dot product,

dT(x′,T(x))2=∥ϖ(T(x)−x′)∥2≤∥vT(x)−ϖ(x′−x)∥2=∥vT(x)∥2−2 vT(x)⋅ϖ(x′−x)+dT(x,x′)2.d_{\mathbb{T}}\bigl(x',T(x)\bigr)^{2}=\bigl\lVert\varpi\bigl(T(x)-x'\bigr)\bigr\rVert^{2}\le\bigl\lVert v_{T}(x)-\varpi(x'-x)\bigr\rVert^{2}=\lVert v_{T}(x)\rVert^{2}-2\,v_{T}(x)\cdot\varpi(x'-x)+d_{\mathbb{T}}(x,x')^{2}.

Step 1.4 (integration). The three functions of ww on the right are Borel by (F1) and (F2) and bounded by d/4d/4 in absolute value (using ∣a⋅b∣≤∥a∥ ∥b∥|a\cdot b|\le\lVert a\rVert\,\lVert b\rVert), so they are integrable against γ\gamma. By (F2) and (pr1)#γ=μ(\mathrm{pr}_{1})_{\#}\gamma=\mu, ∫∥vT∘pr1∥2 dγ=∥vT∥μ2=2Φ(μ)\int\lVert v_{T}\circ\mathrm{pr}_{1}\rVert^{2}\,d\gamma=\lVert v_{T}\rVert_{\mu}^{2}=2\Phi(\mu) (Step 1.1); the integral of vT(pr1(w))⋅ϖ(pr2(w)−pr1(w))v_{T}(\mathrm{pr}_{1}(w))\cdot\varpi(\mathrm{pr}_{2}(w)-\mathrm{pr}_{1}(w)) is JT(vT,γ)\mathcal{J}_{\mathbb{T}}(v_{T},\gamma) by The Torus Displacement Pairing of a Vector Field Along a Coupling §pairing; and the integral of dT(pr1(w),pr2(w))2d_{\mathbb{T}}(\mathrm{pr}_{1}(w),\mathrm{pr}_{2}(w))^{2} is IT(γ)I_{\mathbb{T}}(\gamma) by Probability Measures on the Flat Torus, the Torus Cost of a Coupling, the Torus Wasserstein Distance and Optimal Couplings §cost. Integrating Step 1.3 against γ\gamma by monotonicity and linearity (F2) and using Step 1.2,

2Φ(μ′)≤2Φ(μ)−2 JT(vT,γ)+IT(γ),2\Phi(\mu')\le2\Phi(\mu)-2\,\mathcal{J}_{\mathbb{T}}(v_{T},\gamma)+I_{\mathbb{T}}(\gamma),

and multiplying by 12\tfrac12 gives clause 1.

Proof of clause 2. Let μ\mu be absolutely continuous and TT an optimal map from μ\mu to ν\nu, with vTv_{T} as above; write W=WT(μ,ν)W=W_{\mathbb{T}}(\mu,\nu). The class of −vT-v_{T} lies in L2(μ;Rd)L^{2}(\mu;\mathbb{R}^{d}), and JT(−vT,γ)=−JT(vT,γ)\mathcal{J}_{\mathbb{T}}(-v_{T},\gamma)=-\mathcal{J}_{\mathbb{T}}(v_{T},\gamma) for every μ′\mu' and γ∈Π(μ,μ′)\gamma\in\Pi(\mu,\mu'), by The Torus Displacement Pairing: Linearity, the Cost Bound, and Vanishing of a Field with First-Order Small Pairings §linear with a=−1a=-1, b=0b=0. So we must show: for every real ε>0\varepsilon>0 there is a real θ>0\theta>0 with

∣Φ(μ′)−Φ(μ)+JT(vT,γ)∣≤εIT(γ)(∗)\bigl|\Phi(\mu')-\Phi(\mu)+\mathcal{J}_{\mathbb{T}}(v_{T},\gamma)\bigr|\le\varepsilon\sqrt{I_{\mathbb{T}}(\gamma)}\qquad(\ast)

for all μ′∈P(Td)\mu'\in\mathcal{P}(\mathbb{T}^{d}) and γ∈Π(μ,μ′)\gamma\in\Pi(\mu,\mu') with IT(γ)<θ2I_{\mathbb{T}}(\gamma)<\theta^{2}.

Step 2.1 (the lower estimate). We claim: for every real ε>0\varepsilon>0 there is a real θ1>0\theta_{1}>0 such that Φ(μ′)−Φ(μ)+JT(vT,γ)≥−εIT(γ)\Phi(\mu')-\Phi(\mu)+\mathcal{J}_{\mathbb{T}}(v_{T},\gamma)\ge-\varepsilon\sqrt{I_{\mathbb{T}}(\gamma)} for all μ′∈P(Td)\mu'\in\mathcal{P}(\mathbb{T}^{d}) and γ∈Π(μ,μ′)\gamma\in\Pi(\mu,\mu') with IT(γ)<θ12I_{\mathbb{T}}(\gamma)<\theta_{1}^{2}. Suppose not, and fix ε>0\varepsilon>0 for which it fails. For each n∈Nn\in\mathbb{N}, in this order: using the failure with θ1=1/(n+1)\theta_{1}=1/(n+1), choose μn′∈P(Td)\mu'_{n}\in\mathcal{P}(\mathbb{T}^{d}) and γn∈Π(μ,μn′)\gamma_{n}\in\Pi(\mu,\mu'_{n}) with

IT(γn)<1(n+1)2,Φ(μn′)−Φ(μ)+JT(vT,γn)<−εIT(γn);(1)I_{\mathbb{T}}(\gamma_{n})<\frac{1}{(n+1)^{2}},\qquad\Phi(\mu'_{n})-\Phi(\mu)+\mathcal{J}_{\mathbb{T}}(v_{T},\gamma_{n})<-\varepsilon\sqrt{I_{\mathbb{T}}(\gamma_{n})};\qquad(1)

then choose an optimal σn′∈Π(μn′,ν)\sigma'_{n}\in\Pi(\mu'_{n},\nu) by The Torus Wasserstein Space is a Sequentially Compact Metric Space in which Optimal Couplings Exist §optimal; then choose a gluing Σn∈P(R3d)\Sigma_{n}\in\mathcal{P}(\mathbb{R}^{3d}) of γn\gamma_{n} and σn′\sigma'_{n} by Gluing Two Couplings over a Common Middle Marginal, and the Composite Coupling §glued. That lemma is stated in Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation, whose dimension dd is the one of The Wasserstein Space and Its Lift to Square-Integrable Random Vectors: Standing Notation §data, an arbitrary natural number with 1≤d1\le d; we take it to be the torus dimension dd. Its hypotheses hold: μ,μn′,ν∈P2(Rd)\mu,\mu'_{n},\nu\in\mathcal{P}_{2}(\mathbb{R}^{d}) by The Torus Wasserstein Distance: Comparison with the Euclidean Distance, Wrapping, and Integrals of Periodic Functions §inclusion, and its couplings are those of Couplings of Two Probability Measures on Euclidean Space and Their Quadratic Cost §coupling (Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §dimensions), which are also the couplings of Optimal Transport on the Flat Torus: Standing Notation §measures. Thus, with q1,q2,q3\mathrm{q}_{1},\mathrm{q}_{2},\mathrm{q}_{3} the coordinate maps of R3d\mathbb{R}^{3d}, (q1,q2)#Σn=γn(\mathrm{q}_{1},\mathrm{q}_{2})_{\#}\Sigma_{n}=\gamma_{n} and (q2,q3)#Σn=σn′(\mathrm{q}_{2},\mathrm{q}_{3})_{\#}\Sigma_{n}=\sigma'_{n}. Put sn=IT(γn)s_{n}=\sqrt{I_{\mathbb{T}}(\gamma_{n})}, so 0≤sn<1/(n+1)≤10\le s_{n}<1/(n+1)\le1, and Wn′=WT(μn′,ν)W'_{n}=W_{\mathbb{T}}(\mu'_{n},\nu), so 2Φ(μn′)=Wn′2=IT(σn′)2\Phi(\mu'_{n})=W_{n}'^{2}=I_{\mathbb{T}}(\sigma'_{n}) by (F3).

Step 2.2 (the composite displacement). Define Borel maps R3d→Rd\mathbb{R}^{3d}\to\mathbb{R}^{d} (by (F1) and (F2)) by

an=ϖ∘(q2−q1),bn=ϖ∘(q3−q2),Zn=an+bn.a_{n}=\varpi\circ(\mathrm{q}_{2}-\mathrm{q}_{1}),\qquad b_{n}=\varpi\circ(\mathrm{q}_{3}-\mathrm{q}_{2}),\qquad Z_{n}=a_{n}+b_{n}.

By (F1), ∥an∥,∥bn∥≤d/2\lVert a_{n}\rVert,\lVert b_{n}\rVert\le\sqrt d/2 and ∥Zn∥≤d\lVert Z_{n}\rVert\le\sqrt d everywhere, and for every ww,

q3(w)−q1(w)−Zn(w)=(q3(w)−q2(w)−ϖ(q3(w)−q2(w)))+(q2(w)−q1(w)−ϖ(q2(w)−q1(w)))∈Zd.\mathrm{q}_{3}(w)-\mathrm{q}_{1}(w)-Z_{n}(w)=\bigl(\mathrm{q}_{3}(w)-\mathrm{q}_{2}(w)-\varpi(\mathrm{q}_{3}(w)-\mathrm{q}_{2}(w))\bigr)+\bigl(\mathrm{q}_{2}(w)-\mathrm{q}_{1}(w)-\varpi(\mathrm{q}_{2}(w)-\mathrm{q}_{1}(w))\bigr)\in\mathbb{Z}^{d}.

By (F2), (q1)#Σn=(pr1)#γn=μ(\mathrm{q}_{1})_{\#}\Sigma_{n}=(\mathrm{pr}_{1})_{\#}\gamma_{n}=\mu and (q3)#Σn=(pr2)#σn′=ν(\mathrm{q}_{3})_{\#}\Sigma_{n}=(\mathrm{pr}_{2})_{\#}\sigma'_{n}=\nu. Since ∥an(w)∥=dT(q1(w),q2(w))\lVert a_{n}(w)\rVert=d_{\mathbb{T}}(\mathrm{q}_{1}(w),\mathrm{q}_{2}(w)) by (F1), the change-of-variables formula through (q1,q2)(\mathrm{q}_{1},\mathrm{q}_{2}) and Probability Measures on the Flat Torus, the Torus Cost of a Coupling, the Torus Wasserstein Distance and Optimal Couplings §cost give ∫∥an∥2 dΣn=IT(γn)=sn2\int\lVert a_{n}\rVert^{2}\,d\Sigma_{n}=I_{\mathbb{T}}(\gamma_{n})=s_{n}^{2}; in the same way, through (q2,q3)(\mathrm{q}_{2},\mathrm{q}_{3}), ∫∥bn∥2 dΣn=IT(σn′)=Wn′2\int\lVert b_{n}\rVert^{2}\,d\Sigma_{n}=I_{\mathbb{T}}(\sigma'_{n})=W_{n}'^{2}. Also, vT∘q1⋅anv_{T}\circ\mathrm{q}_{1}\cdot a_{n} is the Borel integrand of The Torus Displacement Pairing of a Vector Field Along a Coupling §pairing composed with (q1,q2)(\mathrm{q}_{1},\mathrm{q}_{2}), so by the change-of-variables formula

∫R3d(vT∘q1)⋅an dΣn=JT(vT,γn).(2)\int_{\mathbb{R}^{3d}}(v_{T}\circ\mathrm{q}_{1})\cdot a_{n}\,d\Sigma_{n}=\mathcal{J}_{\mathbb{T}}(v_{T},\gamma_{n}).\qquad(2)

Let L2(Σn;Rd)L^{2}(\Sigma_{n};\mathbb{R}^{d}) be the space of Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §fields read with 3d3d in place of qq and dd in place of rr, with inner product ⟨⋅,⋅⟩Σn\langle\cdot,\cdot\rangle_{\Sigma_{n}} and norm ∥⋅∥Σn\lVert\cdot\rVert_{\Sigma_{n}}; it is a real inner product space by The Space of Square-Integrable Random Vectors §inner-product (indeed a real Hilbert space by The Space of Square-Integrable Random Vectors is a Real Hilbert Space; Its Laws Have Finite Second Moment; Constants and Translations §hilbert). The classes of the bounded Borel maps ana_{n}, bnb_{n}, ZnZ_{n} and vT∘q1v_{T}\circ\mathrm{q}_{1} belong to it. So ∥an∥Σn=sn\lVert a_{n}\rVert_{\Sigma_{n}}=s_{n}, ∥bn∥Σn=Wn′\lVert b_{n}\rVert_{\Sigma_{n}}=W'_{n} and, by the triangle inequality The Norm Metric of a Real Inner Product Space: Triangle Inequalities, Limits and Continuity §triangle, ∥Zn∥Σn≤sn+Wn′\lVert Z_{n}\rVert_{\Sigma_{n}}\le s_{n}+W'_{n}.

Step 2.3 (near optimality of ZnZ_{n}). By the triangle inequality and symmetry of the metric WTW_{\mathbb{T}} (The Torus Wasserstein Space is a Sequentially Compact Metric Space in which Optimal Couplings Exist §metric), Wn′≤WT(μ,μn′)+WW'_{n}\le W_{\mathbb{T}}(\mu,\mu'_{n})+W, and WT(μ,μn′)≤snW_{\mathbb{T}}(\mu,\mu'_{n})\le s_{n} by (F3) applied to γn\gamma_{n}. Hence ∥Zn∥Σn≤W+2sn\lVert Z_{n}\rVert_{\Sigma_{n}}\le W+2s_{n}, and squaring these nonnegative numbers and using sn2≤sn<1/(n+1)s_{n}^{2}\le s_{n}<1/(n+1),

∫R3d∥Zn∥2 dΣn≤W2+4Wsn+4sn2≤W2+4W+4n+1.\int_{\mathbb{R}^{3d}}\lVert Z_{n}\rVert^{2}\,d\Sigma_{n}\le W^{2}+4Ws_{n}+4s_{n}^{2}\le W^{2}+\frac{4W+4}{n+1}.

Given a real ε′>0\varepsilon'>0, choose N∈NN\in\mathbb{N} with (4W+4)/(N+1)≤ε′(4W+4)/(N+1)\le\varepsilon' (Archimedean property); then ∫∥Zn∥2 dΣn≤W2+ε′\int\lVert Z_{n}\rVert^{2}\,d\Sigma_{n}\le W^{2}+\varepsilon' for every n≥Nn\ge N. Together with Step 2.2, this verifies all hypotheses of Nearly Optimal Composite Displacements on the Torus Converge to the Optimal Displacement with q=3dq=3d, Xn=q1X_{n}=\mathrm{q}_{1}, Yn=q3Y_{n}=\mathrm{q}_{3} and the ZnZ_{n} above (μ\mu is absolutely continuous and TT is an optimal map from μ\mu to ν\nu). By Nearly Optimal Composite Displacements on the Torus Converge to the Optimal Displacement §convergence, the real sequence

cn=∫R3d∥Zn−vT∘q1∥2 dΣn=∥Zn−vT∘q1∥Σn2c_{n}=\int_{\mathbb{R}^{3d}}\lVert Z_{n}-v_{T}\circ\mathrm{q}_{1}\rVert^{2}\,d\Sigma_{n}=\lVert Z_{n}-v_{T}\circ\mathrm{q}_{1}\rVert_{\Sigma_{n}}^{2}

converges to 00.

Step 2.4 (the reverse inequality). By Gluing Two Couplings over a Common Middle Marginal, and the Composite Coupling §composite, κn=(q1,q3)#Σn∈Π(μ,ν)\kappa_{n}=(\mathrm{q}_{1},\mathrm{q}_{3})_{\#}\Sigma_{n}\in\Pi(\mu,\nu), and by the change-of-variables formula IT(κn)=∫dT(q1(w),q3(w))2 Σn(dw)I_{\mathbb{T}}(\kappa_{n})=\int d_{\mathbb{T}}(\mathrm{q}_{1}(w),\mathrm{q}_{3}(w))^{2}\,\Sigma_{n}(dw). For each ww, taking k=q3(w)−q1(w)−Zn(w)∈Zdk=\mathrm{q}_{3}(w)-\mathrm{q}_{1}(w)-Z_{n}(w)\in\mathbb{Z}^{d} (Step 2.2) in (F1) gives dT(q1(w),q3(w))≤∥Zn(w)∥d_{\mathbb{T}}(\mathrm{q}_{1}(w),\mathrm{q}_{3}(w))\le\lVert Z_{n}(w)\rVert. Hence, by (F3), monotonicity, and bilinearity and symmetry of ⟨⋅,⋅⟩Σn\langle\cdot,\cdot\rangle_{\Sigma_{n}},

2Φ(μ)=W2≤IT(κn)≤∥Zn∥Σn2=sn2+2⟨an,bn⟩Σn+2Φ(μn′).2\Phi(\mu)=W^{2}\le I_{\mathbb{T}}(\kappa_{n})\le\lVert Z_{n}\rVert_{\Sigma_{n}}^{2}=s_{n}^{2}+2\langle a_{n},b_{n}\rangle_{\Sigma_{n}}+2\Phi(\mu'_{n}).

Since bn=Zn−anb_{n}=Z_{n}-a_{n} and Zn=vT∘q1+(Zn−vT∘q1)Z_{n}=v_{T}\circ\mathrm{q}_{1}+(Z_{n}-v_{T}\circ\mathrm{q}_{1}), bilinearity and (2) give

⟨an,bn⟩Σn=JT(vT,γn)+⟨an,Zn−vT∘q1⟩Σn−sn2,\langle a_{n},b_{n}\rangle_{\Sigma_{n}}=\mathcal{J}_{\mathbb{T}}(v_{T},\gamma_{n})+\langle a_{n},Z_{n}-v_{T}\circ\mathrm{q}_{1}\rangle_{\Sigma_{n}}-s_{n}^{2},

and the middle term has absolute value at most sncns_{n}\sqrt{c_{n}} by the Cauchy–Schwarz inequality The Cauchy-Schwarz Inequality in a Real Inner Product Space in L2(Σn;Rd)L^{2}(\Sigma_{n};\mathbb{R}^{d}). Substituting,

2Φ(μ)≤2Φ(μn′)+2 JT(vT,γn)+2sncn−sn2,2\Phi(\mu)\le2\Phi(\mu'_{n})+2\,\mathcal{J}_{\mathbb{T}}(v_{T},\gamma_{n})+2s_{n}\sqrt{c_{n}}-s_{n}^{2},

so, as sn2≥0s_{n}^{2}\ge0,

Φ(μn′)−Φ(μ)+JT(vT,γn)≥−sncnfor every n∈N.(3)\Phi(\mu'_{n})-\Phi(\mu)+\mathcal{J}_{\mathbb{T}}(v_{T},\gamma_{n})\ge-s_{n}\sqrt{c_{n}}\qquad\text{for every }n\in\mathbb{N}.\qquad(3)

Step 2.5 (contradiction). Since (cn)(c_{n}) converges to 00 (Limit of a Sequence of Real Numbers), there is nn with cn<ε2c_{n}<\varepsilon^{2}, hence cn≤ε\sqrt{c_{n}}\le\varepsilon by claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field, both cn\sqrt{c_{n}} and ε\varepsilon being nonnegative. For this nn, (3) and sn≥0s_{n}\ge0 give Φ(μn′)−Φ(μ)+JT(vT,γn)≥−εsn\Phi(\mu'_{n})-\Phi(\mu)+\mathcal{J}_{\mathbb{T}}(v_{T},\gamma_{n})\ge-\varepsilon s_{n}, contradicting (1). This proves the claim of Step 2.1.

Step 2.6 (conclusion). Let ε>0\varepsilon>0, let θ1\theta_{1} be as in Step 2.1, and let θ\theta be the smaller of 2ε2\varepsilon and θ1\theta_{1} (Elementary Properties of the Minimum of Two Elements). Let μ′∈P(Td)\mu'\in\mathcal{P}(\mathbb{T}^{d}) and γ∈Π(μ,μ′)\gamma\in\Pi(\mu,\mu') with IT(γ)<θ2I_{\mathbb{T}}(\gamma)<\theta^{2}; then IT(γ)<θ≤2ε\sqrt{I_{\mathbb{T}}(\gamma)}<\theta\le2\varepsilon, so 12IT(γ)=12IT(γ)IT(γ)≤εIT(γ)\tfrac12I_{\mathbb{T}}(\gamma)=\tfrac12\sqrt{I_{\mathbb{T}}(\gamma)}\sqrt{I_{\mathbb{T}}(\gamma)}\le\varepsilon\sqrt{I_{\mathbb{T}}(\gamma)}. By clause 1, proved above (with the same TT), Φ(μ′)−Φ(μ)+JT(vT,γ)≤12IT(γ)≤εIT(γ)\Phi(\mu')-\Phi(\mu)+\mathcal{J}_{\mathbb{T}}(v_{T},\gamma)\le\tfrac12I_{\mathbb{T}}(\gamma)\le\varepsilon\sqrt{I_{\mathbb{T}}(\gamma)}, and by Step 2.1, since IT(γ)<θ12I_{\mathbb{T}}(\gamma)<\theta_{1}^{2}, the same quantity is at least −εIT(γ)-\varepsilon\sqrt{I_{\mathbb{T}}(\gamma)}. This is (∗)(\ast). Hence Φ\Phi is differentiable along couplings at μ\mu with gradient −vT-v_{T}, and by the uniqueness in Differentiability Along Couplings of a Function on the Torus Wasserstein Space, and Its Gradient §gradient, ∇Φ(μ)=−vT\nabla\Phi(\mu)=-v_{T}.

Proof of clause 3. Keep μ\mu, TT and vTv_{T} as in clause 2. By McCann's Theorem on the Flat Torus: Optimal Couplings out of an Absolutely Continuous Measure are Induced by a Unique Map with a Periodic Potential §potential there are a Zd\mathbb{Z}^{d}-periodic function φ:Rd→R\varphi:\mathbb{R}^{d}\to\mathbb{R}, Lipschitz with constant d/2\sqrt d/2, such that ψ=β−φ\psi=\beta-\varphi is convex on Rd\mathbb{R}^{d}, where β(x)=12∥x∥2\beta(x)=\tfrac12\lVert x\rVert^{2}, and a set D∈B(Rd)D\in\mathcal{B}(\mathbb{R}^{d}) with μ(D)=1\mu(D)=1 such that ∂Rdψ(x)={x+vT(x)}\partial_{\mathbb{R}^{d}}\psi(x)=\{x+v_{T}(x)\} for every x∈Dx\in D. Subdifferentials are those of Subdifferential of a Real-Valued Function on a Convex Subset of Rn\mathbb{R}^n §subdifferential, relative to Rd\mathbb{R}^{d}, which is open and convex. Recall that QQ is the half-open unit cell of The Half-Open Unit Cell Tiles Euclidean Space §cell, so every x∈Qx\in Q has 0≤xi<10\le x_{i}<1 for all ii and hence ∥x∥≤d\lVert x\rVert\le\sqrt d; and μ(Q)=1\mu(Q)=1 by Probability Measures on the Flat Torus, the Torus Cost of a Coupling, the Torus Wasserstein Distance and Optimal Couplings §measures.

Step 3.1 (the quadratic). The zero function on Rd\mathbb{R}^{d} satisfies the quadratic increment inequality of Quadratic Increment Characterisation of Semiconvexity with constant 11, since its right side there is 12t(1−t)∥x−y∥2≥0\tfrac12t(1-t)\lVert x-y\rVert^{2}\ge0; hence it is semiconvex with constant 11, which by Semiconvex Function on a Convex Subset of Rn\mathbb{R}^n means that β\beta is convex on Rd\mathbb{R}^{d}. For x,y,p∈Rdx,y,p\in\mathbb{R}^{d}, expanding by bilinearity of the dot product, β(y)−β(x)−x⋅(y−x)=12∥y−x∥2≥0\beta(y)-\beta(x)-x\cdot(y-x)=\tfrac12\lVert y-x\rVert^{2}\ge0, so x∈∂β(x)x\in\partial\beta(x); and if p∈∂β(x)p\in\partial\beta(x) then, taking y=py=p, 0≤β(p)−β(x)−p⋅(p−x)=−12∥p−x∥20\le\beta(p)-\beta(x)-p\cdot(p-x)=-\tfrac12\lVert p-x\rVert^{2}, so p=xp=x. Thus ∂β(x)={x}\partial\beta(x)=\{x\} for every xx.

Step 3.2 (local bound on subgradients of ψ\psi). Put R=d+2R=\sqrt d+2 and L=2R+d/2L=2R+\sqrt d/2. For z,w∈Bˉ(0,2R)z,w\in\bar{B}(0,2R), ∣β(z)−β(w)∣=12∣(z−w)⋅(z+w)∣≤12∥z−w∥(∥z∥+∥w∥)≤2R∥z−w∥|\beta(z)-\beta(w)|=\tfrac12|(z-w)\cdot(z+w)|\le\tfrac12\lVert z-w\rVert(\lVert z\rVert+\lVert w\rVert)\le2R\lVert z-w\rVert and ∣φ(z)−φ(w)∣≤d2∥z−w∥|\varphi(z)-\varphi(w)|\le\tfrac{\sqrt d}{2}\lVert z-w\rVert, so ∣ψ(z)−ψ(w)∣≤L∥z−w∥|\psi(z)-\psi(w)|\le L\lVert z-w\rVert. By Elementary Calculus of the Subdifferential of a Convex Function §bounded (with U=RdU=\mathbb{R}^{d}, y0=0y_{0}=0, r=Rr=R, M=LM=L), every subgradient of ψ\psi at a point of Bˉ(0,R)\bar{B}(0,R) has norm at most LL; and ∂ψ(y)≠∅\partial\psi(y)\neq\emptyset for every yy by The Subdifferential of a Convex Function on an Open Convex Set is Nonempty §nonempty.

Step 3.3 (Lipschitz truncations). Apply The Lipschitz Truncation of a Convex Function: a Global Lipschitz Convex Minorant Agreeing with It Where the Slope is Small with n=dn=d, G=RdG=\mathbb{R}^{d} and level LL, first to β\beta (witness x0=0x_{0}=0, q0=0∈∂β(0)q_{0}=0\in\partial\beta(0), Step 3.1) and then to ψ\psi (witness x0=0x_{0}=0 and any q0∈∂ψ(0)q_{0}\in\partial\psi(0), of norm at most LL by Step 3.2). By The Lipschitz Truncation of a Convex Function: a Global Lipschitz Convex Minorant Agreeing with It Where the Slope is Small §minorant, βL\beta^{L} and ψL\psi^{L} are convex on Rd\mathbb{R}^{d} and Lipschitz with constant LL. Let y∈Bˉ(0,R)y\in\bar{B}(0,R). Since y∈∂β(y)y\in\partial\beta(y) and ∥y∥≤R≤L\lVert y\rVert\le R\le L, and since ψ\psi has a subgradient at yy of norm at most LL (Step 3.2), The Lipschitz Truncation of a Convex Function: a Global Lipschitz Convex Minorant Agreeing with It Where the Slope is Small §agreement gives βL(y)=β(y)\beta^{L}(y)=\beta(y) and ψL(y)=ψ(y)\psi^{L}(y)=\psi(y); hence

φ(y)=βL(y)−ψL(y)(y∈Bˉ(0,R)).(4)\varphi(y)=\beta^{L}(y)-\psi^{L}(y)\qquad(y\in\bar{B}(0,R)).\qquad(4)

By The Lipschitz Truncation of a Convex Function: a Global Lipschitz Convex Minorant Agreeing with It Where the Slope is Small §unique and Steps 3.1 and 3.2: ∂βL(y)={y}\partial\beta^{L}(y)=\{y\} for y∈Bˉ(0,R)y\in\bar{B}(0,R), and ∂ψL(x)={x+vT(x)}\partial\psi^{L}(x)=\{x+v_{T}(x)\} for x∈D∩Bˉ(0,R)x\in D\cap\bar{B}(0,R), the point x+vT(x)x+v_{T}(x) being a subgradient of ψ\psi at a point of Bˉ(0,R)\bar{B}(0,R) and so of norm at most LL.

Step 3.4 (mollification). Fix a mollifier kernel ρ\rho of radius 11 on Rd\mathbb{R}^{d} (claim 2 of Existence of Mollifier Kernels of Every Radius, with n=dn=d, δ=1\delta=1). For m∈Nm\in\mathbb{N} put εm=1/(m+1)\varepsilon_{m}=1/(m+1), a sequence of positive reals converging to 00 by the Archimedean property, and ρm(y)=(εm−1)dρ(εm−1y)\rho_{m}(y)=(\varepsilon_{m}^{-1})^{d}\rho(\varepsilon_{m}^{-1}y), a mollifier kernel of radius εm≤1\varepsilon_{m}\le1 by Rescaling a Mollifier Kernel; in particular ρm\rho_{m} is smooth, hence continuous, and vanishes off Bˉ(0,εm)\bar{B}(0,\varepsilon_{m}). The function φ\varphi is continuous by A Lipschitz Map is Uniformly Continuous. Let φm=φ∗ρm\varphi_{m}=\varphi*\rho_{m}, βm=βL∗ρm\beta_{m}=\beta^{L}*\rho_{m} and ψm=ψL∗ρm\psi_{m}=\psi^{L}*\rho_{m} be the convolutions with Ω=Rd\Omega=\mathbb{R}^{d}, defined on all of Rd\mathbb{R}^{d}; the last two are the mollifications of Mollification of a Lipschitz Convex Function: Smooth Convex Approximations with Bounded Gradients Converging Where the Subgradient is Unique for δ=1\delta=1 and ε=εm\varepsilon=\varepsilon_{m}.

(a) φm∈Cper∞\varphi_{m}\in C^{\infty}_{\mathrm{per}}: it is smooth on Rd\mathbb{R}^{d} by claim 2 of Convolution with a CkC^k Kernel is of Class CkC^k; and for x∈Rdx\in\mathbb{R}^{d} and k∈Zdk\in\mathbb{Z}^{d} the integrands defining φm(x+k)\varphi_{m}(x+k) and φm(x)\varphi_{m}(x) coincide, since φ(x+k−y)ρm(y)=φ(x−y)ρm(y)\varphi(x+k-y)\rho_{m}(y)=\varphi(x-y)\rho_{m}(y) by periodicity of φ\varphi (Lattice-Periodic Functions and the Periodic Function Classes §periodic), so φm(x+k)=φm(x)\varphi_{m}(x+k)=\varphi_{m}(x) (Lattice-Periodic Functions and the Periodic Function Classes §classes). Hence the class of ∇φm\nabla\varphi_{m} lies in GμG_{\mu} (The Tangent Space of the Torus Wasserstein Space at a Probability Measure §gradients), and ∇φm\nabla\varphi_{m} is Borel by Optimal Transport on the Flat Torus: Standing Notation §calculus.

(b) By Mollification of a Lipschitz Convex Function: Smooth Convex Approximations with Bounded Gradients Converging Where the Subgradient is Unique §regularity, applied to βL\beta^{L} and to ψL\psi^{L}, βm\beta_{m} and ψm\psi_{m} are smooth with ∥Dβm(x)∥≤L\lVert D\beta_{m}(x)\rVert\le L and ∥Dψm(x)∥≤L\lVert D\psi_{m}(x)\rVert\le L for every xx.

(c) Let U={x∈Rd:∥x∥<d+1}U=\{x\in\mathbb{R}^{d}:\lVert x\rVert<\sqrt d+1\}; it is open, since for x∈Ux\in U every zz with ∥z−x∥<d+1−∥x∥\lVert z-x\rVert<\sqrt d+1-\lVert x\rVert lies in UU by the triangle inequality, and Q⊆UQ\subseteq U. For x′∈Ux'\in U and y∈Rdy\in\mathbb{R}^{d}: if ∥y∥≤εm\lVert y\rVert\le\varepsilon_{m} then ∥x′−y∥<d+2=R\lVert x'-y\rVert<\sqrt d+2=R, so by (4) φ(x′−y)ρm(y)=βL(x′−y)ρm(y)−ψL(x′−y)ρm(y)\varphi(x'-y)\rho_{m}(y)=\beta^{L}(x'-y)\rho_{m}(y)-\psi^{L}(x'-y)\rho_{m}(y); otherwise all three products vanish. The three integrands are integrable by claim 1 of The Convolution Integrand is Continuous, Compactly Supported and Integrable, so linearity (claim 2 of Linearity and Monotonicity of the Lebesgue Integral) gives φm=βm−ψm\varphi_{m}=\beta_{m}-\psi_{m} on UU. By Differential Calculus and Convexity on Euclidean Open Sets: Standing Notation §derivatives the partial derivatives at points of UU of φm\varphi_{m} and of βm−ψm\beta_{m}-\psi_{m} are those of their common restriction to UU, and by claim 1 of Constants, Coordinate Functions, Sums and Products of CkC^k Functions on a Euclidean Open Set those of βm−ψm\beta_{m}-\psi_{m} are the differences of those of βm\beta_{m} and ψm\psi_{m}. Hence ∇φm(x)=Dβm(x)−Dψm(x)\nabla\varphi_{m}(x)=D\beta_{m}(x)-D\psi_{m}(x) and, by (b), ∥∇φm(x)∥≤2L\lVert\nabla\varphi_{m}(x)\rVert\le2L for every x∈Ux\in U, in particular for every x∈Qx\in Q.

(d) Let x∈D∩Qx\in D\cap Q. Since Q⊆Bˉ(0,R)Q\subseteq\bar{B}(0,R), Step 3.3 gives ∂βL(x)={x}\partial\beta^{L}(x)=\{x\} and ∂ψL(x)={x+vT(x)}\partial\psi^{L}(x)=\{x+v_{T}(x)\}, so by Mollification of a Lipschitz Convex Function: Smooth Convex Approximations with Bounded Gradients Converging Where the Subgradient is Unique §gradients the sequences (Dβm(x))m(D\beta_{m}(x))_{m} and (Dψm(x))m(D\psi_{m}(x))_{m} converge to xx and to x+vT(x)x+v_{T}(x). By (c),

∥∇φm(x)+vT(x)∥≤∥Dβm(x)−x∥+∥Dψm(x)−x−vT(x)∥;\lVert\nabla\varphi_{m}(x)+v_{T}(x)\rVert\le\lVert D\beta_{m}(x)-x\rVert+\lVert D\psi_{m}(x)-x-v_{T}(x)\rVert;

given a real η>0\eta>0, both terms on the right are less than η/2\sqrt\eta/2 for all mm beyond some index, so fm(x)=∥∇φm(x)+vT(x)∥2<ηf_{m}(x)=\lVert\nabla\varphi_{m}(x)+v_{T}(x)\rVert^{2}<\eta there. Thus (fm(x))m(f_{m}(x))_{m} converges to 00 for every x∈D∩Qx\in D\cap Q.

(e) The functions fm:Rd→Rf_{m}:\mathbb{R}^{d}\to\mathbb{R} are Borel by (a), (F1) and (F2). For x∈Qx\in Q, 0≤fm(x)≤(2L+d/2)20\le f_{m}(x)\le(2L+\sqrt d/2)^{2} by (c) and (F1), and the constant (2L+d/2)2(2L+\sqrt d/2)^{2} is integrable against μ\mu. The set (Rd∖D)∪(Rd∖Q)(\mathbb{R}^{d}\setminus D)\cup(\mathbb{R}^{d}\setminus Q) is μ\mu-null by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §null-union, so the convergence in (d) and the bound hold for μ\mu-almost every xx. By The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §dominated, applied on (Rd,B(Rd),μ)(\mathbb{R}^{d},\mathcal{B}(\mathbb{R}^{d}),\mu) with limit 00, ∫fm dμ\int f_{m}\,d\mu converges to 00.

(f) Since ∫fm dμ=∥∇φm−(−vT)∥μ2\int f_{m}\,d\mu=\lVert\nabla\varphi_{m}-(-v_{T})\rVert_{\mu}^{2} is the square of the distance in L2(μ;Rd)L^{2}(\mu;\mathbb{R}^{d}) between the classes of ∇φm\nabla\varphi_{m} and −vT-v_{T} (Optimal Transport on the Flat Torus: Standing Notation §fields), that distance converges to 00: given η>0\eta>0, eventually ∫fm dμ<η2\int f_{m}\,d\mu<\eta^{2}. So the sequence (∇φm)m(\nabla\varphi_{m})_{m} of elements of GμG_{\mu} converges to −vT-v_{T}, and by Sequential Characterization of the Closure in a Metric Space the class of −vT-v_{T} lies in the closure of GμG_{\mu}, which is TμT_{\mu} by The Tangent Space of the Torus Wasserstein Space at a Probability Measure §tangent. As TμT_{\mu} is a linear subspace of L2(μ;Rd)L^{2}(\mu;\mathbb{R}^{d}) by The Torus Tangent Space: Linearity of the Periodic Calculus, Closed Subspace, and Representation of Bounded Functionals on Gradients §subspace, the class of vT=(−1)(−vT)v_{T}=(-1)(-v_{T}) lies in TμT_{\mu}. This proves clause 3.

Citations

Loading…

Dependencies

Uses0

Loading…

Comments

Log in to comment.

Loading…