Each result cited is universally quantified over the data in its own statement. We write W W W for W 2 W_{2} W 2 , 1 / n 1/n 1/ n for the multiplicative inverse of the positive real ι ( n ) \iota(n) ι ( n ) attached to n ∈ N n\in\mathbb{N} n ∈ N (The Real Numbers: Standing Notation and Background §numbers ), and, as in The Squared Wasserstein Distance to a Fixed Measure and Functions of the Mean are Intrinsic Test Functions , a vector a ∈ R d a\in\mathbb{R}^{d} a ∈ R d also for the class of the constant map with value a a a in any L 2 ( ρ ; R d ) L^{2}(\rho;\mathbb{R}^{d}) L 2 ( ρ ; R d ) . The real sequence ( 1 / n ) n ∈ N (1/n)_{n\in\mathbb{N}} ( 1/ n ) n ∈ N converges to 0 0 0 by The Archimedean Property of the Real Numbers . Convergence in R d \mathbb{R}^{d} R d is convergence in the metric space ( R d , d E ) (\mathbb{R}^{d},d_{E}) ( R d , d E ) , where d E ( a , a ′ ) = ∥ a − a ′ ∥ d_{E}(a,a')=\lVert a-a'\rVert d E ( a , a ′ ) = ∥ a − a ′ ∥ by claim 2 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n .
Step 1 (The midpoint and two auxiliary functions). By Existence of an Optimal Coupling of Two Probability Measures with Finite Second Moment fix an optimal coupling π ^ ∈ Π ( μ ^ , ν ^ ) \hat{\pi}\in\Pi(\hat{\mu},\hat{\nu}) π ^ ∈ Π ( μ ^ , ν ^ ) , and let m ^ \hat{m} m ^ , c c c and A A A be as in The Displacement Midpoint and the Midpoint Split of the Squared Wasserstein Distance for μ = μ ^ \mu=\hat{\mu} μ = μ ^ , ν = ν ^ \nu=\hat{\nu} ν = ν ^ and π ^ \hat{\pi} π ^ . Put ζ ^ = m ( μ ^ ) \hat{\zeta}=m(\hat{\mu}) ζ ^ = m ( μ ^ ) , ω ^ = m ( ν ^ ) \hat{\omega}=m(\hat{\nu}) ω ^ = m ( ν ^ ) and M = Ψ ( μ ^ , ν ^ ) M=\Psi(\hat{\mu},\hat{\nu}) M = Ψ ( μ ^ , ν ^ ) , so that c = 1 2 ( ζ ^ + ω ^ ) c=\tfrac12(\hat{\zeta}+\hat{\omega}) c = 2 1 ( ζ ^ + ω ^ ) , m ^ ∈ P 2 ( R d ) \hat{m}\in\mathcal{P}_{2}(\mathbb{R}^{d}) m ^ ∈ P 2 ( R d ) by The Displacement Midpoint and the Midpoint Split of the Squared Wasserstein Distance §midpoint , and 0 ≤ A ( ρ ) 0\le A(\rho) 0 ≤ A ( ρ ) for every ρ \rho ρ by The Displacement Midpoint and the Midpoint Split of the Squared Wasserstein Distance §nonnegative . Fix e 0 ∈ R e_{0}\in\mathbb{R} e 0 ∈ R with e 0 ≤ E ( ρ ) e_{0}\le\mathcal{E}(\rho) e 0 ≤ E ( ρ ) for every ρ ∈ D \rho\in\mathcal{D} ρ ∈ D (Basic Properties of a Wasserstein-Coercive Penalty Pair §bounded-below ). For ρ ∈ D \rho\in\mathcal{D} ρ ∈ D put
Θ ( ρ ) = u δ − ( ρ ) − α A ( ρ ) , Ξ ( ρ ) = v δ + ( ρ ) + α A ( ρ ) . \Theta(\rho)=u^{-}_{\delta}(\rho)-\alpha A(\rho),\qquad\Xi(\rho)=v^{+}_{\delta}(\rho)+\alpha A(\rho). Θ ( ρ ) = u δ − ( ρ ) − α A ( ρ ) , Ξ ( ρ ) = v δ + ( ρ ) + α A ( ρ ) .
As u δ − = u − δ E u^{-}_{\delta}=u-\delta\mathcal{E} u δ − = u − δ E and v δ + = v + δ E v^{+}_{\delta}=v+\delta\mathcal{E} v δ + = v + δ E on D \mathcal{D} D , u ≤ b u\le b u ≤ b , b ′ ≤ v b'\le v b ′ ≤ v , and 0 ≤ α A ( ρ ) 0\le\alpha A(\rho) 0 ≤ α A ( ρ ) and δ e 0 ≤ δ E ( ρ ) \delta e_{0}\le\delta\mathcal{E}(\rho) δ e 0 ≤ δ E ( ρ ) by claim 5 of Elementary Arithmetic in an Ordered Field , for every ρ ∈ D \rho\in\mathcal{D} ρ ∈ D
Θ ( ρ ) ≤ b − δ E ( ρ ) ≤ b − δ e 0 , b ′ + δ e 0 ≤ b ′ + δ E ( ρ ) ≤ Ξ ( ρ ) . ( 1 a ) \Theta(\rho)\le b-\delta\,\mathcal{E}(\rho)\le b-\delta e_{0},\qquad b'+\delta e_{0}\le b'+\delta\,\mathcal{E}(\rho)\le\Xi(\rho).\qquad(1\mathrm{a}) Θ ( ρ ) ≤ b − δ E ( ρ ) ≤ b − δ e 0 , b ′ + δ e 0 ≤ b ′ + δ E ( ρ ) ≤ Ξ ( ρ ) . ( 1 a )
For ρ , σ ∈ D \rho,\sigma\in\mathcal{D} ρ , σ ∈ D one has Ψ ( ρ , σ ) − ( Θ ( ρ ) − Ξ ( σ ) − α 2 ∥ m ( ρ ) − m ( σ ) ∥ 2 ) = α 2 ( 2 A ( ρ ) + 2 A ( σ ) + ∥ m ( ρ ) − m ( σ ) ∥ 2 − W ( ρ , σ ) 2 ) \Psi(\rho,\sigma)-\bigl(\Theta(\rho)-\Xi(\sigma)-\tfrac{\alpha}{2}\lVert m(\rho)-m(\sigma)\rVert^{2}\bigr)=\tfrac{\alpha}{2}\bigl(2A(\rho)+2A(\sigma)+\lVert m(\rho)-m(\sigma)\rVert^{2}-W(\rho,\sigma)^{2}\bigr) Ψ ( ρ , σ ) − ( Θ ( ρ ) − Ξ ( σ ) − 2 α ∥ m ( ρ ) − m ( σ ) ∥ 2 ) = 2 α ( 2 A ( ρ ) + 2 A ( σ ) + ∥ m ( ρ ) − m ( σ ) ∥ 2 − W ( ρ , σ ) 2 ) , which is nonnegative by The Displacement Midpoint and the Midpoint Split of the Squared Wasserstein Distance §split and vanishes exactly when equality holds there, α 2 \tfrac{\alpha}{2} 2 α being positive (claims 5 and 8 of Elementary Order Arithmetic in an Ordered Field , claim 3 of Zero Products and Elementary Identities in a Field ). With the maximality of ( μ ^ , ν ^ ) (\hat{\mu},\hat{\nu}) ( μ ^ , ν ^ ) ,
Θ ( ρ ) − Ξ ( σ ) − α 2 ∥ m ( ρ ) − m ( σ ) ∥ 2 ≤ Ψ ( ρ , σ ) ≤ M ( ρ , σ ∈ D ) , ( 1 b ) \Theta(\rho)-\Xi(\sigma)-\tfrac{\alpha}{2}\lVert m(\rho)-m(\sigma)\rVert^{2}\le\Psi(\rho,\sigma)\le M\qquad(\rho,\sigma\in\mathcal{D}),\qquad(1\mathrm{b}) Θ ( ρ ) − Ξ ( σ ) − 2 α ∥ m ( ρ ) − m ( σ ) ∥ 2 ≤ Ψ ( ρ , σ ) ≤ M ( ρ , σ ∈ D ) , ( 1 b )
with equality in the first inequality exactly when equality holds in The Displacement Midpoint and the Midpoint Split of the Squared Wasserstein Distance §split for ( ρ , σ ) (\rho,\sigma) ( ρ , σ ) . By The Displacement Midpoint and the Midpoint Split of the Squared Wasserstein Distance §endpoints this is so for ( μ ^ , ν ^ ) (\hat{\mu},\hat{\nu}) ( μ ^ , ν ^ ) :
Θ ( μ ^ ) − Ξ ( ν ^ ) − α 2 ∥ ζ ^ − ω ^ ∥ 2 = M . ( 1 c ) \Theta(\hat{\mu})-\Xi(\hat{\nu})-\tfrac{\alpha}{2}\lVert\hat{\zeta}-\hat{\omega}\rVert^{2}=M.\qquad(1\mathrm{c}) Θ ( μ ^ ) − Ξ ( ν ^ ) − 2 α ∥ ζ ^ − ω ^ ∥ 2 = M . ( 1 c )
Step 2 (A A A is an intrinsic test function). Let q c : R d → R q_{c}:\mathbb{R}^{d}\to\mathbb{R} q c : R d → R , q c ( a ) = ∥ a − c ∥ 2 = d E ( a , c ) 2 q_{c}(a)=\lVert a-c\rVert^{2}=d_{E}(a,c)^{2} q c ( a ) = ∥ a − c ∥ 2 = d E ( a , c ) 2 . By A Scaled Squared Distance to a Point is of Class C 2 C^2 C 2 , with Gradient and Hessian , applied with the point c c c and the scalar 1 1 1 , q c q_{c} q c is of class C 2 C^{2} C 2 on R d \mathbb{R}^{d} R d with D q c ( a ) = 2 ( a − c ) Dq_{c}(a)=2(a-c) D q c ( a ) = 2 ( a − c ) and D 2 q c ( a ) = 2 I d D^{2}q_{c}(a)=2I_{d} D 2 q c ( a ) = 2 I d . Since A ( ρ ) = W ( ρ , m ^ ) 2 − q c ( m ( ρ ) ) A(\rho)=W(\rho,\hat{m})^{2}-q_{c}(m(\rho)) A ( ρ ) = W ( ρ , m ^ ) 2 − q c ( m ( ρ )) and D \mathcal{D} D has the map property, The Squared Wasserstein Distance to a Fixed Measure and Functions of the Mean are Intrinsic Test Functions §distance (with ν 0 = m ^ \nu_{0}=\hat{m} ν 0 = m ^ ), The Squared Wasserstein Distance to a Fixed Measure and Functions of the Mean are Intrinsic Test Functions §mean (with ϕ = q c \phi=q_{c} ϕ = q c ) and Restrictions, Sums, Real Multiples and Differences of Intrinsic Test Functions on the Wasserstein Space §difference show that A A A is an intrinsic test function on D \mathcal{D} D , with
∇ A ( ρ ) = 2 ( i d − G ρ ) − 2 ( m ( ρ ) − c ) ( ρ ∈ D ) , H A ( ρ ) = 2 I d − 2 I d = 0 d ( ρ ∈ P 2 ( R d ) ) , \nabla A(\rho)=2(\mathrm{id}-G_{\rho})-2\bigl(m(\rho)-c\bigr)\quad(\rho\in\mathcal{D}),\qquad H_{A}(\rho)=2I_{d}-2I_{d}=0_{d}\quad\bigl(\rho\in\mathcal{P}_{2}(\mathbb{R}^{d})\bigr), ∇ A ( ρ ) = 2 ( id − G ρ ) − 2 ( m ( ρ ) − c ) ( ρ ∈ D ) , H A ( ρ ) = 2 I d − 2 I d = 0 d ( ρ ∈ P 2 ( R d ) ) ,
where G ρ G_{\rho} G ρ is any optimal map from ρ \rho ρ to m ^ \hat{m} m ^ . By property (a) of Intrinsic Test Functions on the Wasserstein Space and Their Translation Hessians §test , A A A is continuous on P 2 ( R d ) \mathcal{P}_{2}(\mathbb{R}^{d}) P 2 ( R d ) .
Step 3 (Compactness). We show:
(K-u) Let ( ρ n ) n ∈ N (\rho_{n})_{n\in\mathbb{N}} ( ρ n ) n ∈ N be a sequence in D \mathcal{D} D , ζ ∈ R d \zeta\in\mathbb{R}^{d} ζ ∈ R d and ℓ ∈ R \ell\in\mathbb{R} ℓ ∈ R with ( m ( ρ n ) ) n (m(\rho_{n}))_{n} ( m ( ρ n ) ) n converging to ζ \zeta ζ and ℓ ≤ Θ ( ρ n ) \ell\le\Theta(\rho_{n}) ℓ ≤ Θ ( ρ n ) for every n n n . Then there are a strictly increasing sequence ( n j ) j ∈ N (n_{j})_{j\in\mathbb{N}} ( n j ) j ∈ N in N \mathbb{N} N and ρ ∈ D \rho\in\mathcal{D} ρ ∈ D with m ( ρ ) = ζ m(\rho)=\zeta m ( ρ ) = ζ such that ( ρ n j ) j (\rho_{n_{j}})_{j} ( ρ n j ) j converges to ρ \rho ρ in ( P 2 ( R d ) , W ) (\mathcal{P}_{2}(\mathbb{R}^{d}),W) ( P 2 ( R d ) , W ) and, for every positive ε ∈ R \varepsilon\in\mathbb{R} ε ∈ R , Θ ( ρ n j ) < Θ ( ρ ) + ε \Theta(\rho_{n_{j}})<\Theta(\rho)+\varepsilon Θ ( ρ n j ) < Θ ( ρ ) + ε for all sufficiently large j j j .
(K-v) The same with ℓ ′ ∈ R \ell'\in\mathbb{R} ℓ ′ ∈ R , Ξ ( σ n ) ≤ ℓ ′ \Xi(\sigma_{n})\le\ell' Ξ ( σ n ) ≤ ℓ ′ for every n n n in place of the lower bound, and the conclusion Ξ ( σ ) − ε < Ξ ( σ n j ) \Xi(\sigma)-\varepsilon<\Xi(\sigma_{n_{j}}) Ξ ( σ ) − ε < Ξ ( σ n j ) for all sufficiently large j j j .
For (K-u): by (1a), ℓ ≤ b − δ E ( ρ n ) \ell\le b-\delta\mathcal{E}(\rho_{n}) ℓ ≤ b − δ E ( ρ n ) , so E ( ρ n ) ≤ c 1 \mathcal{E}(\rho_{n})\le c_{1} E ( ρ n ) ≤ c 1 with c 1 = δ − 1 ( b − ℓ ) c_{1}=\delta^{-1}(b-\ell) c 1 = δ − 1 ( b − ℓ ) (claims 3 and 5 of Elementary Arithmetic in an Ordered Field , δ − 1 \delta^{-1} δ − 1 being positive by claim 7 of Elementary Order Arithmetic in an Ordered Field ). The set { ρ ′ ∈ D : E ( ρ ′ ) ≤ c 1 } \{\rho'\in\mathcal{D}:\mathcal{E}(\rho')\le c_{1}\} { ρ ′ ∈ D : E ( ρ ′ ) ≤ c 1 } is sequentially compact (Wasserstein-Coercive Penalty Pairs §coercive ), so there are a strictly increasing ( n j ) j (n_{j})_{j} ( n j ) j and a point ρ \rho ρ of that set, hence of D \mathcal{D} D , with ρ n j → ρ \rho_{n_{j}}\to\rho ρ n j → ρ (Sequentially Compact Subset of a Metric Space ). By The Mean of a Square-Integrable Probability Measure, Its Lift, Its Centring, and Functions of the Mean and Centred Integrals as Test Functions §mean , ∥ m ( ρ n j ) − m ( ρ ) ∥ ≤ W ( ρ n j , ρ ) \lVert m(\rho_{n_{j}})-m(\rho)\rVert\le W(\rho_{n_{j}},\rho) ∥ m ( ρ n j ) − m ( ρ )∥ ≤ W ( ρ n j , ρ ) , so m ( ρ n j ) → m ( ρ ) m(\rho_{n_{j}})\to m(\rho) m ( ρ n j ) → m ( ρ ) ; also m ( ρ n j ) → ζ m(\rho_{n_{j}})\to\zeta m ( ρ n j ) → ζ by A Subsequence of a Convergent Sequence Has the Same Limit , so m ( ρ ) = ζ m(\rho)=\zeta m ( ρ ) = ζ by Uniqueness of Limits in a Metric Space . Let ε > 0 \varepsilon>0 ε > 0 . Since u ≤ b u\le b u ≤ b , u u u is bounded above near each point (Upper and Lower Semicontinuous Envelopes of a Real-Valued Function §near-bounds ), so u δ − u^{-}_{\delta} u δ − is upper semicontinuous on D \mathcal{D} D relative to D \mathcal{D} D by Basic Properties of the Delta-Envelopes on the Wasserstein Space §semicontinuity ; with the continuity of A A A (Step 2) there is a positive r r r such that every ρ ′ ∈ D \rho'\in\mathcal{D} ρ ′ ∈ D with W ( ρ ′ , ρ ) < r W(\rho',\rho)<r W ( ρ ′ , ρ ) < r satisfies u δ − ( ρ ′ ) < u δ − ( ρ ) + ε 2 u^{-}_{\delta}(\rho')<u^{-}_{\delta}(\rho)+\tfrac{\varepsilon}{2} u δ − ( ρ ′ ) < u δ − ( ρ ) + 2 ε and ∣ A ( ρ ′ ) − A ( ρ ) ∣ < ε 2 α − 1 |A(\rho')-A(\rho)|<\tfrac{\varepsilon}{2}\alpha^{-1} ∣ A ( ρ ′ ) − A ( ρ ) ∣ < 2 ε α − 1 (the least of two radii, claim 9 of Elementary Order Arithmetic in an Ordered Field ), hence Θ ( ρ ′ ) < Θ ( ρ ) + ε \Theta(\rho')<\Theta(\rho)+\varepsilon Θ ( ρ ′ ) < Θ ( ρ ) + ε by claim 3 of Elementary Order Arithmetic in an Ordered Field and claim 3 of Properties of the Absolute Value in an Ordered Field . As W ( ρ n j , ρ ) < r W(\rho_{n_{j}},\rho)<r W ( ρ n j , ρ ) < r for all large j j j , (K-u) follows. (K-v) is proved in the same way: (1a) gives E ( σ n ) ≤ δ − 1 ( ℓ ′ − b ′ ) \mathcal{E}(\sigma_{n})\le\delta^{-1}(\ell'-b') E ( σ n ) ≤ δ − 1 ( ℓ ′ − b ′ ) , and v δ + v^{+}_{\delta} v δ + is lower semicontinuous on D \mathcal{D} D relative to D \mathcal{D} D by the second part of Basic Properties of the Delta-Envelopes on the Wasserstein Space §semicontinuity , applied to v v v , which is bounded below near each point since b ′ ≤ v b'\le v b ′ ≤ v .
Step 4 (The fibre functions). For ζ ∈ R d \zeta\in\mathbb{R}^{d} ζ ∈ R d let D ζ = { ρ ∈ D : m ( ρ ) = ζ } \mathcal{D}_{\zeta}=\{\rho\in\mathcal{D}:m(\rho)=\zeta\} D ζ = { ρ ∈ D : m ( ρ ) = ζ } . It is nonempty: ( τ ζ − ζ ^ ) # μ ^ ∈ D (\tau_{\zeta-\hat{\zeta}})_{\#}\hat{\mu}\in\mathcal{D} ( τ ζ − ζ ^ ) # μ ^ ∈ D by Penalty Pairs on the Wasserstein Space: the Penalty, Its Score, and Their Domains §translation-invariant , and its mean is ζ ^ + ( ζ − ζ ^ ) = ζ \hat{\zeta}+(\zeta-\hat{\zeta})=\zeta ζ ^ + ( ζ − ζ ^ ) = ζ by The Mean of a Square-Integrable Probability Measure, Its Lift, Its Centring, and Functions of the Mean and Centred Integrals as Test Functions §mean . By (1a) the set { Θ ( ρ ) : ρ ∈ D ζ } \{\Theta(\rho):\rho\in\mathcal{D}_{\zeta}\} { Θ ( ρ ) : ρ ∈ D ζ } is bounded above and { Ξ ( σ ) : σ ∈ D ζ } \{\Xi(\sigma):\sigma\in\mathcal{D}_{\zeta}\} { Ξ ( σ ) : σ ∈ D ζ } bounded below; let U ( ζ ) ∈ R U(\zeta)\in\mathbb{R} U ( ζ ) ∈ R be the supremum of the first and V ( ζ ) ∈ R V(\zeta)\in\mathbb{R} V ( ζ ) ∈ R the infimum of the second (Approximation Property of the Supremum and the Infimum in R \mathbb{R} R ).
(4a) The fibre extrema are attained. Let ζ ∈ R d \zeta\in\mathbb{R}^{d} ζ ∈ R d . By claim 3 of Approximation Property of the Supremum and the Infimum in R \mathbb{R} R choose ρ n ∈ D ζ \rho_{n}\in\mathcal{D}_{\zeta} ρ n ∈ D ζ with U ( ζ ) − 1 / n < Θ ( ρ n ) U(\zeta)-1/n<\Theta(\rho_{n}) U ( ζ ) − 1/ n < Θ ( ρ n ) for each n n n ; then U ( ζ ) − 1 ≤ Θ ( ρ n ) U(\zeta)-1\le\Theta(\rho_{n}) U ( ζ ) − 1 ≤ Θ ( ρ n ) , as 1 / n ≤ 1 1/n\le1 1/ n ≤ 1 . (K-u), with the constant sequence of means ζ \zeta ζ and ℓ = U ( ζ ) − 1 \ell=U(\zeta)-1 ℓ = U ( ζ ) − 1 , gives ( n j ) j (n_{j})_{j} ( n j ) j and ρ ∈ D ζ \rho\in\mathcal{D}_{\zeta} ρ ∈ D ζ . For ε > 0 \varepsilon>0 ε > 0 take j j j so large that Θ ( ρ n j ) < Θ ( ρ ) + ε \Theta(\rho_{n_{j}})<\Theta(\rho)+\varepsilon Θ ( ρ n j ) < Θ ( ρ ) + ε and 1 / n j < ε 1/n_{j}<\varepsilon 1/ n j < ε (the sequence ( 1 / n j ) j (1/n_{j})_{j} ( 1/ n j ) j converges to 0 0 0 by A Subsequence of a Convergent Sequence Has the Same Limit ); then U ( ζ ) < Θ ( ρ ) + 2 ε U(\zeta)<\Theta(\rho)+2\varepsilon U ( ζ ) < Θ ( ρ ) + 2 ε . By Comparison of Real Numbers with Arbitrary Positive Slack §slack-above , U ( ζ ) ≤ Θ ( ρ ) ≤ U ( ζ ) U(\zeta)\le\Theta(\rho)\le U(\zeta) U ( ζ ) ≤ Θ ( ρ ) ≤ U ( ζ ) , so Θ ( ρ ) = U ( ζ ) \Theta(\rho)=U(\zeta) Θ ( ρ ) = U ( ζ ) . In the same way, with claim 4 of Approximation Property of the Supremum and the Infimum in R \mathbb{R} R , (K-v) and Comparison of Real Numbers with Arbitrary Positive Slack §slack-below , there is σ ∈ D ζ \sigma\in\mathcal{D}_{\zeta} σ ∈ D ζ with Ξ ( σ ) = V ( ζ ) \Xi(\sigma)=V(\zeta) Ξ ( σ ) = V ( ζ ) .
(4b) U U U is upper and V V V lower semicontinuous on R d \mathbb{R}^{d} R d , in the sense of Upper Semicontinuous Function on a Subset of a Metric Space and Lower Semicontinuous Function on a Subset of a Metric Space in ( R d , d E ) (\mathbb{R}^{d},d_{E}) ( R d , d E ) , which is the reading of Second-Order Equations on Euclidean Open Sets §extrema . Suppose U U U were not upper semicontinuous at ζ \zeta ζ . Then there is ε > 0 \varepsilon>0 ε > 0 such that for every n n n some ζ n \zeta_{n} ζ n has d E ( ζ n , ζ ) < 1 / n d_{E}(\zeta_{n},\zeta)<1/n d E ( ζ n , ζ ) < 1/ n and U ( ζ ) + ε ≤ U ( ζ n ) U(\zeta)+\varepsilon\le U(\zeta_{n}) U ( ζ ) + ε ≤ U ( ζ n ) ; so ζ n → ζ \zeta_{n}\to\zeta ζ n → ζ . By (4a) choose ρ n ∈ D ζ n \rho_{n}\in\mathcal{D}_{\zeta_{n}} ρ n ∈ D ζ n with Θ ( ρ n ) = U ( ζ n ) \Theta(\rho_{n})=U(\zeta_{n}) Θ ( ρ n ) = U ( ζ n ) . (K-u) with ℓ = U ( ζ ) + ε \ell=U(\zeta)+\varepsilon ℓ = U ( ζ ) + ε gives ρ ∈ D ζ \rho\in\mathcal{D}_{\zeta} ρ ∈ D ζ and, for large j j j , U ( ζ ) + ε ≤ Θ ( ρ n j ) < Θ ( ρ ) + ε 2 ≤ U ( ζ ) + ε 2 U(\zeta)+\varepsilon\le\Theta(\rho_{n_{j}})<\Theta(\rho)+\tfrac{\varepsilon}{2}\le U(\zeta)+\tfrac{\varepsilon}{2} U ( ζ ) + ε ≤ Θ ( ρ n j ) < Θ ( ρ ) + 2 ε ≤ U ( ζ ) + 2 ε , which is impossible as ε 2 < ε \tfrac{\varepsilon}{2}<\varepsilon 2 ε < ε (claim 8 of Elementary Order Arithmetic in an Ordered Field ). The lower semicontinuity of V V V follows in the same way from (4a) and (K-v).
Step 5 (Ishii's lemma on the means; claim 2). Let ζ , ω ∈ R d \zeta,\omega\in\mathbb{R}^{d} ζ , ω ∈ R d and, by (4a), ρ ∈ D ζ \rho\in\mathcal{D}_{\zeta} ρ ∈ D ζ and σ ∈ D ω \sigma\in\mathcal{D}_{\omega} σ ∈ D ω with Θ ( ρ ) = U ( ζ ) \Theta(\rho)=U(\zeta) Θ ( ρ ) = U ( ζ ) and Ξ ( σ ) = V ( ω ) \Xi(\sigma)=V(\omega) Ξ ( σ ) = V ( ω ) . By (1b),
U ( ζ ) − V ( ω ) − α 2 ∥ ζ − ω ∥ 2 ≤ M . ( 5 a ) U(\zeta)-V(\omega)-\tfrac{\alpha}{2}\lVert\zeta-\omega\rVert^{2}\le M.\qquad(5\mathrm{a}) U ( ζ ) − V ( ω ) − 2 α ∥ ζ − ω ∥ 2 ≤ M . ( 5 a )
Since μ ^ ∈ D ζ ^ \hat{\mu}\in\mathcal{D}_{\hat{\zeta}} μ ^ ∈ D ζ ^ and ν ^ ∈ D ω ^ \hat{\nu}\in\mathcal{D}_{\hat{\omega}} ν ^ ∈ D ω ^ , Θ ( μ ^ ) ≤ U ( ζ ^ ) \Theta(\hat{\mu})\le U(\hat{\zeta}) Θ ( μ ^ ) ≤ U ( ζ ^ ) and V ( ω ^ ) ≤ Ξ ( ν ^ ) V(\hat{\omega})\le\Xi(\hat{\nu}) V ( ω ^ ) ≤ Ξ ( ν ^ ) , so (1c) and (5a) give
U ( ζ ^ ) − V ( ω ^ ) − α 2 ∥ ζ ^ − ω ^ ∥ 2 = M . ( 5 b ) U(\hat{\zeta})-V(\hat{\omega})-\tfrac{\alpha}{2}\lVert\hat{\zeta}-\hat{\omega}\rVert^{2}=M.\qquad(5\mathrm{b}) U ( ζ ^ ) − V ( ω ^ ) − 2 α ∥ ζ ^ − ω ^ ∥ 2 = M . ( 5 b )
Apply Ishii's Lemma: Test Data and Matrix Bounds at a Maximum of a Quadratically Penalised Difference with n = d n=d n = d , the open set Ω = R d \Omega=\mathbb{R}^{d} Ω = R d (claim 1 of Euclidean Space is Open in Itself, and C k C^k C k Maps are Continuous ), the functions U U U and V V V (Step 4), x ^ = ζ ^ \hat{x}=\hat{\zeta} x ^ = ζ ^ , y ^ = ω ^ \hat{y}=\hat{\omega} y ^ = ω ^ and radius 1 1 1 : its hypothesis holds by (5a) and (5b) at all points. Let X , Y ∈ S ( d ) \mathbb{X},\mathbb{Y}\in\mathcal{S}(d) X , Y ∈ S ( d ) be the matrices it provides and p = α ( ζ ^ − ω ^ ) p=\alpha(\hat{\zeta}-\hat{\omega}) p = α ( ζ ^ − ω ^ ) . Its claims Ishii's Lemma: Test Data and Matrix Bounds at a Maximum of a Quadratically Penalised Difference §ordering , Ishii's Lemma: Test Data and Matrix Bounds at a Maximum of a Quadratically Penalised Difference §norm-bound and Ishii's Lemma: Test Data and Matrix Bounds at a Maximum of a Quadratically Penalised Difference §quadratic-bound are the four conditions of The Second-Order Structure Condition at Optimally Coupled Pairs on the Lift of the Wasserstein Space §admitted ; this is claim 2. By Ishii's Lemma: Test Data and Matrix Bounds at a Maximum of a Quadratically Penalised Difference §test-data , ( ζ ^ , U ( ζ ^ ) , p , X ) (\hat{\zeta},U(\hat{\zeta}),p,\mathbb{X}) ( ζ ^ , U ( ζ ^ ) , p , X ) is approximable by test data from above for U U U and ( ω ^ , V ( ω ^ ) , p , Y ) (\hat{\omega},V(\hat{\omega}),p,\mathbb{Y}) ( ω ^ , V ( ω ^ ) , p , Y ) from below for V V V , with open set R d \mathbb{R}^{d} R d .
Step 6 (Test functions at fibre maximisers and minimisers). Let n ∈ N n\in\mathbb{N} n ∈ N . By Quadruple Approximable by Test-Function Data §above with ε = 1 / n \varepsilon=1/n ε = 1/ n there are ζ n ∈ R d \zeta_{n}\in\mathbb{R}^{d} ζ n ∈ R d , a function χ n \chi_{n} χ n of class C 2 C^{2} C 2 on R d \mathbb{R}^{d} R d and a positive r n r_{n} r n such that U ( ζ ) − χ n ( ζ ) ≤ U ( ζ n ) − χ n ( ζ n ) U(\zeta)-\chi_{n}(\zeta)\le U(\zeta_{n})-\chi_{n}(\zeta_{n}) U ( ζ ) − χ n ( ζ ) ≤ U ( ζ n ) − χ n ( ζ n ) whenever d E ( ζ , ζ n ) < r n d_{E}(\zeta,\zeta_{n})<r_{n} d E ( ζ , ζ n ) < r n , and
d E ( ζ n , ζ ^ ) < 1 n , ∣ U ( ζ n ) − U ( ζ ^ ) ∣ < 1 n , ∥ D χ n ( ζ n ) − p ∥ < 1 n , ∥ D 2 χ n ( ζ n ) − X ∥ < 1 n , d_{E}(\zeta_{n},\hat{\zeta})<\tfrac1n,\qquad|U(\zeta_{n})-U(\hat{\zeta})|<\tfrac1n,\qquad\lVert D\chi_{n}(\zeta_{n})-p\rVert<\tfrac1n,\qquad\lVert D^{2}\chi_{n}(\zeta_{n})-\mathbb{X}\rVert<\tfrac1n, d E ( ζ n , ζ ^ ) < n 1 , ∣ U ( ζ n ) − U ( ζ ^ ) ∣ < n 1 , ∥ D χ n ( ζ n ) − p ∥ < n 1 , ∥ D 2 χ n ( ζ n ) − X ∥ < n 1 ,
the last because d S ( d ) ( P , P ′ ) = ∥ P − P ′ ∥ d_{\mathcal{S}(d)}(P,P')=\lVert P-P'\rVert d S ( d ) ( P , P ′ ) = ∥ P − P ′ ∥ (Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §matrices ). By (4a) choose ρ n ∈ D ζ n \rho_{n}\in\mathcal{D}_{\zeta_{n}} ρ n ∈ D ζ n with Θ ( ρ n ) = U ( ζ n ) \Theta(\rho_{n})=U(\zeta_{n}) Θ ( ρ n ) = U ( ζ n ) , and let φ n = α A + χ n ∘ m \varphi_{n}=\alpha A+\chi_{n}\circ m φ n = α A + χ n ∘ m . By Step 2, The Squared Wasserstein Distance to a Fixed Measure and Functions of the Mean are Intrinsic Test Functions §mean and Restrictions, Sums, Real Multiples and Differences of Intrinsic Test Functions on the Wasserstein Space §linear , φ n \varphi_{n} φ n is an intrinsic test function on D \mathcal{D} D with
∇ φ n ( ρ n ) = α ∇ A ( ρ n ) + D χ n ( ζ n ) , H φ n ( ρ n ) = α 0 d + D 2 χ n ( ζ n ) = D 2 χ n ( ζ n ) . \nabla\varphi_{n}(\rho_{n})=\alpha\nabla A(\rho_{n})+D\chi_{n}(\zeta_{n}),\qquad H_{\varphi_{n}}(\rho_{n})=\alpha0_{d}+D^{2}\chi_{n}(\zeta_{n})=D^{2}\chi_{n}(\zeta_{n}). ∇ φ n ( ρ n ) = α ∇ A ( ρ n ) + D χ n ( ζ n ) , H φ n ( ρ n ) = α 0 d + D 2 χ n ( ζ n ) = D 2 χ n ( ζ n ) .
For ρ ′ ∈ D \rho'\in\mathcal{D} ρ ′ ∈ D with W ( ρ ′ , ρ n ) < r n W(\rho',\rho_{n})<r_{n} W ( ρ ′ , ρ n ) < r n we have d E ( m ( ρ ′ ) , ζ n ) ≤ W ( ρ ′ , ρ n ) < r n d_{E}(m(\rho'),\zeta_{n})\le W(\rho',\rho_{n})<r_{n} d E ( m ( ρ ′ ) , ζ n ) ≤ W ( ρ ′ , ρ n ) < r n by The Mean of a Square-Integrable Probability Measure, Its Lift, Its Centring, and Functions of the Mean and Centred Integrals as Test Functions §mean , so, by the definition of U U U and ρ ′ ∈ D m ( ρ ′ ) \rho'\in\mathcal{D}_{m(\rho')} ρ ′ ∈ D m ( ρ ′ ) ,
u δ − ( ρ ′ ) − φ n ( ρ ′ ) = Θ ( ρ ′ ) − χ n ( m ( ρ ′ ) ) ≤ U ( m ( ρ ′ ) ) − χ n ( m ( ρ ′ ) ) ≤ U ( ζ n ) − χ n ( ζ n ) = u δ − ( ρ n ) − φ n ( ρ n ) . u^{-}_{\delta}(\rho')-\varphi_{n}(\rho')=\Theta(\rho')-\chi_{n}(m(\rho'))\le U(m(\rho'))-\chi_{n}(m(\rho'))\le U(\zeta_{n})-\chi_{n}(\zeta_{n})=u^{-}_{\delta}(\rho_{n})-\varphi_{n}(\rho_{n}). u δ − ( ρ ′ ) − φ n ( ρ ′ ) = Θ ( ρ ′ ) − χ n ( m ( ρ ′ )) ≤ U ( m ( ρ ′ )) − χ n ( m ( ρ ′ )) ≤ U ( ζ n ) − χ n ( ζ n ) = u δ − ( ρ n ) − φ n ( ρ n ) .
So u δ − − φ n u^{-}_{\delta}-\varphi_{n} u δ − − φ n has a local maximum relative to D \mathcal{D} D at ρ n \rho_{n} ρ n (Local Maximum of a Function Relative to a Subset of a Metric Space ).
Symmetrically, Quadruple Approximable by Test-Function Data §below with ε = 1 / n \varepsilon=1/n ε = 1/ n gives ω n \omega_{n} ω n , χ n ′ \chi'_{n} χ n ′ of class C 2 C^{2} C 2 on R d \mathbb{R}^{d} R d and r n ′ > 0 r'_{n}>0 r n ′ > 0 with V ( ω ) − χ n ′ ( ω ) ≥ V ( ω n ) − χ n ′ ( ω n ) V(\omega)-\chi'_{n}(\omega)\ge V(\omega_{n})-\chi'_{n}(\omega_{n}) V ( ω ) − χ n ′ ( ω ) ≥ V ( ω n ) − χ n ′ ( ω n ) whenever d E ( ω , ω n ) < r n ′ d_{E}(\omega,\omega_{n})<r'_{n} d E ( ω , ω n ) < r n ′ , and the four bounds with ω n , ω ^ , V , χ n ′ , Y \omega_{n},\hat{\omega},V,\chi'_{n},\mathbb{Y} ω n , ω ^ , V , χ n ′ , Y in place of ζ n , ζ ^ , U , χ n , X \zeta_{n},\hat{\zeta},U,\chi_{n},\mathbb{X} ζ n , ζ ^ , U , χ n , X . Choose σ n ∈ D ω n \sigma_{n}\in\mathcal{D}_{\omega_{n}} σ n ∈ D ω n with Ξ ( σ n ) = V ( ω n ) \Xi(\sigma_{n})=V(\omega_{n}) Ξ ( σ n ) = V ( ω n ) and let ψ n = ( − α ) A + χ n ′ ∘ m \psi_{n}=(-\alpha)A+\chi'_{n}\circ m ψ n = ( − α ) A + χ n ′ ∘ m , an intrinsic test function on D \mathcal{D} D with ∇ ψ n ( σ n ) = − α ∇ A ( σ n ) + D χ n ′ ( ω n ) \nabla\psi_{n}(\sigma_{n})=-\alpha\nabla A(\sigma_{n})+D\chi'_{n}(\omega_{n}) ∇ ψ n ( σ n ) = − α ∇ A ( σ n ) + D χ n ′ ( ω n ) and H ψ n ( σ n ) = D 2 χ n ′ ( ω n ) H_{\psi_{n}}(\sigma_{n})=D^{2}\chi'_{n}(\omega_{n}) H ψ n ( σ n ) = D 2 χ n ′ ( ω n ) . For σ ′ ∈ D \sigma'\in\mathcal{D} σ ′ ∈ D with W ( σ ′ , σ n ) < r n ′ W(\sigma',\sigma_{n})<r'_{n} W ( σ ′ , σ n ) < r n ′ , v δ + ( σ ′ ) − ψ n ( σ ′ ) = Ξ ( σ ′ ) − χ n ′ ( m ( σ ′ ) ) ≥ V ( m ( σ ′ ) ) − χ n ′ ( m ( σ ′ ) ) ≥ V ( ω n ) − χ n ′ ( ω n ) = v δ + ( σ n ) − ψ n ( σ n ) v^{+}_{\delta}(\sigma')-\psi_{n}(\sigma')=\Xi(\sigma')-\chi'_{n}(m(\sigma'))\ge V(m(\sigma'))-\chi'_{n}(m(\sigma'))\ge V(\omega_{n})-\chi'_{n}(\omega_{n})=v^{+}_{\delta}(\sigma_{n})-\psi_{n}(\sigma_{n}) v δ + ( σ ′ ) − ψ n ( σ ′ ) = Ξ ( σ ′ ) − χ n ′ ( m ( σ ′ )) ≥ V ( m ( σ ′ )) − χ n ′ ( m ( σ ′ )) ≥ V ( ω n ) − χ n ′ ( ω n ) = v δ + ( σ n ) − ψ n ( σ n ) , so v δ + − ψ n v^{+}_{\delta}-\psi_{n} v δ + − ψ n has a local minimum relative to D \mathcal{D} D at σ n \sigma_{n} σ n (Local Minimum of a Function Relative to a Subset of a Metric Space ).
Step 7 (The limits ρ ∗ \rho^{*} ρ ∗ , σ ∗ \sigma^{*} σ ∗ ; claim 1). We have m ( ρ n ) = ζ n → ζ ^ m(\rho_{n})=\zeta_{n}\to\hat{\zeta} m ( ρ n ) = ζ n → ζ ^ and Θ ( ρ n ) = U ( ζ n ) > U ( ζ ^ ) − 1 / n ≥ U ( ζ ^ ) − 1 \Theta(\rho_{n})=U(\zeta_{n})>U(\hat{\zeta})-1/n\ge U(\hat{\zeta})-1 Θ ( ρ n ) = U ( ζ n ) > U ( ζ ^ ) − 1/ n ≥ U ( ζ ^ ) − 1 . (K-u) gives a strictly increasing ( n j ) j (n_{j})_{j} ( n j ) j and ρ ∗ ∈ D ζ ^ \rho^{*}\in\mathcal{D}_{\hat{\zeta}} ρ ∗ ∈ D ζ ^ with ρ n j → ρ ∗ \rho_{n_{j}}\to\rho^{*} ρ n j → ρ ∗ ; exactly as in (4a), Θ ( ρ ∗ ) = U ( ζ ^ ) \Theta(\rho^{*})=U(\hat{\zeta}) Θ ( ρ ∗ ) = U ( ζ ^ ) , and then ∣ Θ ( ρ n j ) − Θ ( ρ ∗ ) ∣ = ∣ U ( ζ n j ) − U ( ζ ^ ) ∣ < 1 / n j |\Theta(\rho_{n_{j}})-\Theta(\rho^{*})|=|U(\zeta_{n_{j}})-U(\hat{\zeta})|<1/n_{j} ∣Θ ( ρ n j ) − Θ ( ρ ∗ ) ∣ = ∣ U ( ζ n j ) − U ( ζ ^ ) ∣ < 1/ n j . Symmetrically, (K-v) gives ( n j ′ ) j (n'_{j})_{j} ( n j ′ ) j and σ ∗ ∈ D ω ^ \sigma^{*}\in\mathcal{D}_{\hat{\omega}} σ ∗ ∈ D ω ^ with σ n j ′ → σ ∗ \sigma_{n'_{j}}\to\sigma^{*} σ n j ′ → σ ∗ , Ξ ( σ ∗ ) = V ( ω ^ ) \Xi(\sigma^{*})=V(\hat{\omega}) Ξ ( σ ∗ ) = V ( ω ^ ) and ∣ Ξ ( σ n j ′ ) − Ξ ( σ ∗ ) ∣ < 1 / n j ′ |\Xi(\sigma_{n'_{j}})-\Xi(\sigma^{*})|<1/n'_{j} ∣Ξ ( σ n j ′ ) − Ξ ( σ ∗ ) ∣ < 1/ n j ′ . By (1b) and (5b),
M = Θ ( ρ ∗ ) − Ξ ( σ ∗ ) − α 2 ∥ ζ ^ − ω ^ ∥ 2 ≤ Ψ ( ρ ∗ , σ ∗ ) ≤ M , M=\Theta(\rho^{*})-\Xi(\sigma^{*})-\tfrac{\alpha}{2}\lVert\hat{\zeta}-\hat{\omega}\rVert^{2}\le\Psi(\rho^{*},\sigma^{*})\le M, M = Θ ( ρ ∗ ) − Ξ ( σ ∗ ) − 2 α ∥ ζ ^ − ω ^ ∥ 2 ≤ Ψ ( ρ ∗ , σ ∗ ) ≤ M ,
which is claim 1; and equality in the first inequality means, by Step 1, that equality holds in The Displacement Midpoint and the Midpoint Split of the Squared Wasserstein Distance §split for ( ρ ∗ , σ ∗ ) (\rho^{*},\sigma^{*}) ( ρ ∗ , σ ∗ ) , hence also for ( σ ∗ , ρ ∗ ) (\sigma^{*},\rho^{*}) ( σ ∗ , ρ ∗ ) , both sides of that inequality being symmetric in the pair by The Quadratic Wasserstein Distance is a Metric on the Wasserstein Space §symmetry and claim 5 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n .
Step 8 (The limiting gradients). As D \mathcal{D} D has the map property and ρ ∗ , σ ∗ ∈ D \rho^{*},\sigma^{*}\in\mathcal{D} ρ ∗ , σ ∗ ∈ D , the pairs ( ρ ∗ , m ^ ) (\rho^{*},\hat{m}) ( ρ ∗ , m ^ ) , ( ρ ∗ , σ ∗ ) (\rho^{*},\sigma^{*}) ( ρ ∗ , σ ∗ ) , ( σ ∗ , m ^ ) (\sigma^{*},\hat{m}) ( σ ∗ , m ^ ) and ( σ ∗ , ρ ∗ ) (\sigma^{*},\rho^{*}) ( σ ∗ , ρ ∗ ) are uniquely mapped (The Map Property of a Set of Probability Measures §map-property ). By The Displacement Midpoint and the Midpoint Split of the Squared Wasserstein Distance §equality for ( ρ ∗ , σ ∗ ) (\rho^{*},\sigma^{*}) ( ρ ∗ , σ ∗ ) with the optimal maps G ρ ∗ G_{\rho^{*}} G ρ ∗ and S S S , the constant e e e there has value 1 2 ( ζ ^ + ω ^ ) − c = 0 R d \tfrac12(\hat{\zeta}+\hat{\omega})-c=0_{\mathbb{R}^{d}} 2 1 ( ζ ^ + ω ^ ) − c = 0 R d , so 2 ( i d − G ρ ∗ ) = i d − S 2(\mathrm{id}-G_{\rho^{*}})=\mathrm{id}-S 2 ( id − G ρ ∗ ) = id − S in L 2 ( ρ ∗ ; R d ) L^{2}(\rho^{*};\mathbb{R}^{d}) L 2 ( ρ ∗ ; R d ) ; likewise, for ( σ ∗ , ρ ∗ ) (\sigma^{*},\rho^{*}) ( σ ∗ , ρ ∗ ) with G σ ∗ G_{\sigma^{*}} G σ ∗ and S ′ S' S ′ , 2 ( i d − G σ ∗ ) = i d − S ′ 2(\mathrm{id}-G_{\sigma^{*}})=\mathrm{id}-S' 2 ( id − G σ ∗ ) = id − S ′ in L 2 ( σ ∗ ; R d ) L^{2}(\sigma^{*};\mathbb{R}^{d}) L 2 ( σ ∗ ; R d ) . As 2 ( ζ ^ − c ) = ζ ^ − ω ^ 2(\hat{\zeta}-c)=\hat{\zeta}-\hat{\omega} 2 ( ζ ^ − c ) = ζ ^ − ω ^ and 2 ( ω ^ − c ) = ω ^ − ζ ^ 2(\hat{\omega}-c)=\hat{\omega}-\hat{\zeta} 2 ( ω ^ − c ) = ω ^ − ζ ^ , Step 2 gives ∇ A ( ρ ∗ ) = ( i d − S ) − ( ζ ^ − ω ^ ) \nabla A(\rho^{*})=(\mathrm{id}-S)-(\hat{\zeta}-\hat{\omega}) ∇ A ( ρ ∗ ) = ( id − S ) − ( ζ ^ − ω ^ ) and ∇ A ( σ ∗ ) = ( i d − S ′ ) − ( ω ^ − ζ ^ ) \nabla A(\sigma^{*})=(\mathrm{id}-S')-(\hat{\omega}-\hat{\zeta}) ∇ A ( σ ∗ ) = ( id − S ′ ) − ( ω ^ − ζ ^ ) , hence, with p = α ( ζ ^ − ω ^ ) p=\alpha(\hat{\zeta}-\hat{\omega}) p = α ( ζ ^ − ω ^ ) ,
α ∇ A ( ρ ∗ ) + p = α ( i d − S ) in L 2 ( ρ ∗ ; R d ) , − α ∇ A ( σ ∗ ) + p = α ( S ′ − i d ) in L 2 ( σ ∗ ; R d ) . ( 8 a ) \alpha\nabla A(\rho^{*})+p=\alpha(\mathrm{id}-S)\ \text{in }L^{2}(\rho^{*};\mathbb{R}^{d}),\qquad-\alpha\nabla A(\sigma^{*})+p=\alpha(S'-\mathrm{id})\ \text{in }L^{2}(\sigma^{*};\mathbb{R}^{d}).\qquad(8\mathrm{a}) α ∇ A ( ρ ∗ ) + p = α ( id − S ) in L 2 ( ρ ∗ ; R d ) , − α ∇ A ( σ ∗ ) + p = α ( S ′ − id ) in L 2 ( σ ∗ ; R d ) . ( 8 a )
Step 9 (Claim 3). For each j j j let π j ∈ Π ( ρ n j , ρ ∗ ) \pi_{j}\in\Pi(\rho_{n_{j}},\rho^{*}) π j ∈ Π ( ρ n j , ρ ∗ ) be optimal (Existence of an Optimal Coupling of Two Probability Measures with Finite Second Moment ), so I ( π j ) = W ( ρ n j , ρ ∗ ) 2 → 0 I(\pi_{j})=W(\rho_{n_{j}},\rho^{*})^{2}\to0 I ( π j ) = W ( ρ n j , ρ ∗ ) 2 → 0 (Optimal Coupling of Two Probability Measures with Finite Second Moment §optimal ). By property (c) of Intrinsic Test Functions on the Wasserstein Space and Their Translation Hessians §test for A A A on D \mathcal{D} D , the discrepancy D j D_{j} D j of ∇ A ( ρ n j ) \nabla A(\rho_{n_{j}}) ∇ A ( ρ n j ) and ∇ A ( ρ ∗ ) \nabla A(\rho^{*}) ∇ A ( ρ ∗ ) along π j \pi_{j} π j converges to 0 0 0 . Let E j E_{j} E j be the discrepancy of ∇ φ n j ( ρ n j ) \nabla\varphi_{n_{j}}(\rho_{n_{j}}) ∇ φ n j ( ρ n j ) and α ( i d − S ) \alpha(\mathrm{id}-S) α ( id − S ) along π j \pi_{j} π j , which is the integral in claim 3. It does not depend on representatives (The Discrepancy of Two Square-Integrable Vector Fields Along a Coupling of Their Base Measures §well-defined ), so by (8a) and Step 6 its integrand may be taken to be ∥ F 1 ( z ) + F 2 ( z ) ∥ 2 \lVert F_{1}(z)+F_{2}(z)\rVert^{2} ∥ F 1 ( z ) + F 2 ( z ) ∥ 2 with
F 1 ( z ) = α ( ∇ A ( ρ n j ) ( x ) − ∇ A ( ρ ∗ ) ( y ) ) , F 2 ( z ) = D χ n j ( ζ n j ) − p . F_{1}(z)=\alpha\bigl(\nabla A(\rho_{n_{j}})(x)-\nabla A(\rho^{*})(y)\bigr),\qquad F_{2}(z)=D\chi_{n_{j}}(\zeta_{n_{j}})-p . F 1 ( z ) = α ( ∇ A ( ρ n j ) ( x ) − ∇ A ( ρ ∗ ) ( y ) ) , F 2 ( z ) = D χ n j ( ζ n j ) − p .
Both are Borel and square-integrable against π j \pi_{j} π j , with ∥ F 1 ∥ π j = α D j \lVert F_{1}\rVert_{\pi_{j}}=\alpha\sqrt{D_{j}} ∥ F 1 ∥ π j = α D j (claim 5 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n ) and ∥ F 2 ∥ π j = ∥ D χ n j ( ζ n j ) − p ∥ < 1 / n j \lVert F_{2}\rVert_{\pi_{j}}=\lVert D\chi_{n_{j}}(\zeta_{n_{j}})-p\rVert<1/n_{j} ∥ F 2 ∥ π j = ∥ D χ n j ( ζ n j ) − p ∥ < 1/ n j , so the triangle inequality The Norm Metric of a Real Inner Product Space: Triangle Inequalities, Limits and Continuity §triangle in the real Hilbert space L 2 ( π j ; R d ) L^{2}(\pi_{j};\mathbb{R}^{d}) L 2 ( π j ; R d ) (Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §fields ) gives E j < α D j + 1 / n j \sqrt{E_{j}}<\alpha\sqrt{D_{j}}+1/n_{j} E j < α D j + 1/ n j . Moreover u δ − ( ρ n j ) − u δ − ( ρ ∗ ) = ( Θ ( ρ n j ) − Θ ( ρ ∗ ) ) + α ( A ( ρ n j ) − A ( ρ ∗ ) ) u^{-}_{\delta}(\rho_{n_{j}})-u^{-}_{\delta}(\rho^{*})=\bigl(\Theta(\rho_{n_{j}})-\Theta(\rho^{*})\bigr)+\alpha\bigl(A(\rho_{n_{j}})-A(\rho^{*})\bigr) u δ − ( ρ n j ) − u δ − ( ρ ∗ ) = ( Θ ( ρ n j ) − Θ ( ρ ∗ ) ) + α ( A ( ρ n j ) − A ( ρ ∗ ) ) , where the first difference has absolute value below 1 / n j 1/n_{j} 1/ n j (Step 7) and the second converges to 0 0 0 by the continuity of A A A .
Let ε > 0 \varepsilon>0 ε > 0 . Each of the real sequences ( I ( π j ) ) j (I(\pi_{j}))_{j} ( I ( π j ) ) j , ( D j ) j (\sqrt{D_{j}})_{j} ( D j ) j (claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field ), ( 1 / n j ) j (1/n_{j})_{j} ( 1/ n j ) j and ( A ( ρ n j ) − A ( ρ ∗ ) ) j (A(\rho_{n_{j}})-A(\rho^{*}))_{j} ( A ( ρ n j ) − A ( ρ ∗ ) ) j converges to 0 0 0 , so we may fix j j j with I ( π j ) < ε 2 I(\pi_{j})<\varepsilon^{2} I ( π j ) < ε 2 , 1 / n j < ε 2 1/n_{j}<\tfrac{\varepsilon}{2} 1/ n j < 2 ε , α D j < ε 2 \alpha\sqrt{D_{j}}<\tfrac{\varepsilon}{2} α D j < 2 ε and α ∣ A ( ρ n j ) − A ( ρ ∗ ) ∣ < ε 2 \alpha|A(\rho_{n_{j}})-A(\rho^{*})|<\tfrac{\varepsilon}{2} α ∣ A ( ρ n j ) − A ( ρ ∗ ) ∣ < 2 ε . Put ρ = ρ n j \rho=\rho_{n_{j}} ρ = ρ n j , φ = φ n j \varphi=\varphi_{n_{j}} φ = φ n j and π = π j \pi=\pi_{j} π = π j . By Step 6, u δ − − φ u^{-}_{\delta}-\varphi u δ − − φ has a local maximum relative to D \mathcal{D} D at ρ \rho ρ ; ∣ u δ − ( ρ ) − u δ − ( ρ ∗ ) ∣ < ε |u^{-}_{\delta}(\rho)-u^{-}_{\delta}(\rho^{*})|<\varepsilon ∣ u δ − ( ρ ) − u δ − ( ρ ∗ ) ∣ < ε by claim 5 of Properties of the Absolute Value in an Ordered Field ; E j < ε \sqrt{E_{j}}<\varepsilon E j < ε , so E j < ε 2 E_{j}<\varepsilon^{2} E j < ε 2 (claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field ); and ∥ H φ ( ρ ) − X ∥ = ∥ D 2 χ n j ( ζ n j ) − X ∥ < 1 / n j < ε \lVert H_{\varphi}(\rho)-\mathbb{X}\rVert=\lVert D^{2}\chi_{n_{j}}(\zeta_{n_{j}})-\mathbb{X}\rVert<1/n_{j}<\varepsilon ∥ H φ ( ρ ) − X ∥ = ∥ D 2 χ n j ( ζ n j ) − X ∥ < 1/ n j < ε . This is claim 3.
Step 10 (Claim 4). The same argument, with σ n j ′ \sigma_{n'_{j}} σ n j ′ , σ ∗ \sigma^{*} σ ∗ , optimal couplings γ j ∈ Π ( σ n j ′ , σ ∗ ) \gamma_{j}\in\Pi(\sigma_{n'_{j}},\sigma^{*}) γ j ∈ Π ( σ n j ′ , σ ∗ ) , ψ n j ′ \psi_{n'_{j}} ψ n j ′ , Ξ \Xi Ξ , v δ + = Ξ − α A v^{+}_{\delta}=\Xi-\alpha A v δ + = Ξ − α A , the second identity of (8a), the fields F 1 ( z ) = − α ( ∇ A ( σ n j ′ ) ( x ) − ∇ A ( σ ∗ ) ( y ) ) F_{1}(z)=-\alpha\bigl(\nabla A(\sigma_{n'_{j}})(x)-\nabla A(\sigma^{*})(y)\bigr) F 1 ( z ) = − α ( ∇ A ( σ n j ′ ) ( x ) − ∇ A ( σ ∗ ) ( y ) ) and F 2 ( z ) = D χ n j ′ ′ ( ω n j ′ ) − p F_{2}(z)=D\chi'_{n'_{j}}(\omega_{n'_{j}})-p F 2 ( z ) = D χ n j ′ ′ ( ω n j ′ ) − p , the local minimum of Step 6 and H ψ n j ′ ( σ n j ′ ) = D 2 χ n j ′ ′ ( ω n j ′ ) H_{\psi_{n'_{j}}}(\sigma_{n'_{j}})=D^{2}\chi'_{n'_{j}}(\omega_{n'_{j}}) H ψ n j ′ ( σ n j ′ ) = D 2 χ n j ′ ′ ( ω n j ′ ) , gives claim 4.