Each result cited below is universally quantified over the data in its own statement, and is applied to the data named at the point of use. The rules for adding inequalities, for multiplying them by nonnegative or positive real numbers, and for handling absolute values, from Elementary Order Arithmetic in an Ordered Field , Elementary Arithmetic in an Ordered Field and Properties of the Absolute Value in an Ordered Field , are used without further mention; so are the facts that a square of a real number is nonnegative (claim 2 of Nonnegativity of Squares in an Ordered Field ) and that for nonnegative reals a , b a,b a , b one has a < b a<b a < b , a ≤ b a\le b a ≤ b , a = b a=b a = b exactly when a 2 < b 2 a^{2}<b^{2} a 2 < b 2 , a 2 ≤ b 2 a^{2}\le b^{2} a 2 ≤ b 2 , a 2 = b 2 a^{2}=b^{2} a 2 = b 2 respectively (claims 1, 2 and 3 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field ).
Step 0 (Notation and preliminary facts). Fix the data of the statement: the penalty pair ( D , D Σ , E , Σ ) (\mathcal{D},\mathcal{D}_{\Sigma},\mathcal{E},\Sigma) ( D , D Σ , E , Σ ) , the reals λ 0 , θ \lambda_{0},\theta λ 0 , θ with 0 < λ 0 0<\lambda_{0} 0 < λ 0 and 0 < θ ≤ 1 0<\theta\le1 0 < θ ≤ 1 , the natural number p p p and the matrix Γ ∈ M p × d ( R ) \Gamma\in\mathcal{M}_{p\times d}(\mathbb{R}) Γ ∈ M p × d ( R ) , the function g g g , the operator F F F , and a real number C C C as in (Growth).
(0.1) Inner product spaces. For ν ∈ P 2 ( R d ) \nu\in\mathcal{P}_{2}(\mathbb{R}^{d}) ν ∈ P 2 ( R d ) the space L 2 ( ν ; R d ) L^{2}(\nu;\mathbb{R}^{d}) L 2 ( ν ; R d ) , with inner product ⟨ ⋅ , ⋅ ⟩ ν \langle\cdot,\cdot\rangle_{\nu} ⟨ ⋅ , ⋅ ⟩ ν and norm ∥ ⋅ ∥ ν \lVert\cdot\rVert_{\nu} ∥ ⋅ ∥ ν , is a real Hilbert space by Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §fields , in particular a real inner product space ; its norm satisfies ∥ x ∥ ν 2 = ⟨ x , x ⟩ ν \lVert x\rVert_{\nu}^{2}=\langle x,x\rangle_{\nu} ∥ x ∥ ν 2 = ⟨ x , x ⟩ ν by Real Inner Product Space §norm , and ∥ x ∥ ν 2 = ∫ R d ∥ x ∥ 2 d ν \lVert x\rVert_{\nu}^{2}=\int_{\mathbb{R}^{d}}\lVert x\rVert^{2}\,d\nu ∥ x ∥ ν 2 = ∫ R d ∥ x ∥ 2 d ν for a representative x x x , by the formula for the norm in Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §fields . Inner products are symmetric by condition (a) of Real Inner Product Space §inner-product ; bilinearity, homogeneity of the norm and the expansion of ∥ x ± y ∥ ν 2 \lVert x\pm y\rVert_{\nu}^{2} ∥ x ± y ∥ ν 2 are Elementary Identities in a Real Inner Product Space §bilinear , Elementary Identities in a Real Inner Product Space §homogeneity and Elementary Identities in a Real Inner Product Space §expansion ; and ∣ ⟨ x , y ⟩ ν ∣ ≤ ∥ x ∥ ν ∥ y ∥ ν |\langle x,y\rangle_{\nu}|\le\lVert x\rVert_{\nu}\lVert y\rVert_{\nu} ∣ ⟨ x , y ⟩ ν ∣ ≤ ∥ x ∥ ν ∥ y ∥ ν by The Cauchy-Schwarz Inequality in a Real Inner Product Space . For ν ∈ D Σ \nu\in\mathcal{D}_{\Sigma} ν ∈ D Σ the score Σ ( ν ) \Sigma(\nu) Σ ( ν ) lies in L 2 ( ν ; R d ) L^{2}(\nu;\mathbb{R}^{d}) L 2 ( ν ; R d ) by Penalty Pairs on the Wasserstein Space: the Penalty, Its Score, and Their Domains §pair , and E ( ν ) \mathcal{E}(\nu) E ( ν ) is a real number because D Σ ⊆ D \mathcal{D}_{\Sigma}\subseteq\mathcal{D} D Σ ⊆ D .
(0.2) The constant C C C is nonnegative. By Penalty Pairs on the Wasserstein Space: the Penalty, Its Score, and Their Domains §nonempty there is μ 0 ∈ D Σ ⊆ D \mu_{0}\in\mathcal{D}_{\Sigma}\subseteq\mathcal{D} μ 0 ∈ D Σ ⊆ D . Its second moment is a nonnegative real number, so (Growth) gives 0 ≤ M 2 ( μ 0 ) ≤ C ( 1 + ∣ E ( μ 0 ) ∣ ) 0\le M_{2}(\mu_{0})\le C\,(1+|\mathcal{E}(\mu_{0})|) 0 ≤ M 2 ( μ 0 ) ≤ C ( 1 + ∣ E ( μ 0 ) ∣ ) . If C < 0 C<0 C < 0 , then, as 0 < 1 + ∣ E ( μ 0 ) ∣ 0<1+|\mathcal{E}(\mu_{0})| 0 < 1 + ∣ E ( μ 0 ) ∣ , we would get C ( 1 + ∣ E ( μ 0 ) ∣ ) < 0 C\,(1+|\mathcal{E}(\mu_{0})|)<0 C ( 1 + ∣ E ( μ 0 ) ∣ ) < 0 , a contradiction. Hence 0 ≤ C 0\le C 0 ≤ C .
(0.3) A bound for g g g . By (Running cost) and Bounded Real-Valued Function on a Set , fix a real M g ≥ 0 M_{g}\ge0 M g ≥ 0 with ∣ g ( μ ) ∣ ≤ M g |g(\mu)|\le M_{g} ∣ g ( μ ) ∣ ≤ M g for every μ ∈ P 2 ( R d ) \mu\in\mathcal{P}_{2}(\mathbb{R}^{d}) μ ∈ P 2 ( R d ) .
(0.4) The common-noise term. We apply The Trace as a Sum of Quadratic Forms, its Monotonicity and a Norm Bound with its dimensions m m m and p p p read as our p p p and d d d (both at least 1 1 1 , being natural numbers ) and its matrix A A A read as Γ \Gamma Γ ; for j ∈ [ p ] j\in[p] j ∈ [ p ] let ζ j ∈ R d \zeta_{j}\in\mathbb{R}^{d} ζ j ∈ R d be the j j j th row of Γ \Gamma Γ in the sense of that lemma, and write ⋅ \cdot ⋅ for the dot product and ∥ ⋅ ∥ \lVert\cdot\rVert ∥ ⋅ ∥ for the Euclidean norm of R d \mathbb{R}^{d} R d . For Z ∈ S ( d ) Z\in\mathcal{S}(d) Z ∈ S ( d ) put Q ( Z ) = t r ( Γ ⊤ Γ Z ) Q(Z)=\mathrm{tr}(\Gamma^{\top}\Gamma Z) Q ( Z ) = tr ( Γ ⊤ Γ Z ) ; by The Trace as a Sum of Quadratic Forms, its Monotonicity and a Norm Bound §rows ,
Q ( Z ) = ∑ j = 1 p ζ j ⋅ ( Z ζ j ) . Q(Z)=\sum_{j=1}^{p}\zeta_{j}\cdot(Z\zeta_{j}). Q ( Z ) = j = 1 ∑ p ζ j ⋅ ( Z ζ j ) .
Put β Γ = t r ( Γ ⊤ Γ ) \beta_{\Gamma}=\mathrm{tr}(\Gamma^{\top}\Gamma) β Γ = tr ( Γ ⊤ Γ ) . By The Trace as a Sum of Quadratic Forms, its Monotonicity and a Norm Bound §squared-rows , β Γ = ∑ j = 1 p ∥ ζ j ∥ 2 \beta_{\Gamma}=\sum_{j=1}^{p}\lVert\zeta_{j}\rVert^{2} β Γ = ∑ j = 1 p ∥ ζ j ∥ 2 , so 0 ≤ β Γ 0\le\beta_{\Gamma} 0 ≤ β Γ by claim 5 of Properties of Finite Sums , each summand being a square. For Z ∈ S ( d ) Z\in\mathcal{S}(d) Z ∈ S ( d ) and j ∈ [ p ] j\in[p] j ∈ [ p ] , claim 2 of Properties of the Norm of a Symmetric Real Matrix (with n = d n=d n = d ) gives ∣ ζ j ⋅ ( Z ζ j ) ∣ ≤ ∥ Z ∥ ∥ ζ j ∥ 2 |\zeta_{j}\cdot(Z\zeta_{j})|\le\lVert Z\rVert\,\lVert\zeta_{j}\rVert^{2} ∣ ζ j ⋅ ( Z ζ j ) ∣ ≤ ∥ Z ∥ ∥ ζ j ∥ 2 ; hence, by claims 2 and 1 of Comparison and Absolute Value Bounds for Finite Sums of Real Numbers and claim 3 of Properties of Finite Sums ,
∣ Q ( Z ) ∣ ≤ ∑ j = 1 p ∣ ζ j ⋅ ( Z ζ j ) ∣ ≤ ∑ j = 1 p ∥ Z ∥ ∥ ζ j ∥ 2 = β Γ ∥ Z ∥ . |Q(Z)|\le\sum_{j=1}^{p}\bigl|\zeta_{j}\cdot(Z\zeta_{j})\bigr|\le\sum_{j=1}^{p}\lVert Z\rVert\,\lVert\zeta_{j}\rVert^{2}=\beta_{\Gamma}\lVert Z\rVert . ∣ Q ( Z ) ∣ ≤ j = 1 ∑ p ζ j ⋅ ( Z ζ j ) ≤ j = 1 ∑ p ∥ Z ∥ ∥ ζ j ∥ 2 = β Γ ∥ Z ∥ .
Differences and scalar multiples of members of S ( d ) \mathcal{S}(d) S ( d ) lie in S ( d ) \mathcal{S}(d) S ( d ) by claim 1 of The Positive Semidefinite Ordering is a Partial Order Compatible with the Linear Structure . For Y , H ∈ S ( d ) Y,H\in\mathcal{S}(d) Y , H ∈ S ( d ) , real δ \delta δ and j ∈ [ p ] j\in[p] j ∈ [ p ] , claim 1 of Linearity of the Matrix-Vector Product and the Quadratic Form as a Double Sum gives ( Y ± δ H ) ζ j = Y ζ j ± δ ( H ζ j ) (Y\pm\delta H)\zeta_{j}=Y\zeta_{j}\pm\delta\,(H\zeta_{j}) ( Y ± δH ) ζ j = Y ζ j ± δ ( H ζ j ) , and claim 5 of Bilinearity and Symmetry of the Dot Product on R n \mathbb{R}^n R n then gives ζ j ⋅ ( ( Y ± δ H ) ζ j ) = ζ j ⋅ ( Y ζ j ) ± δ ( ζ j ⋅ ( H ζ j ) ) \zeta_{j}\cdot((Y\pm\delta H)\zeta_{j})=\zeta_{j}\cdot(Y\zeta_{j})\pm\delta\,\bigl(\zeta_{j}\cdot(H\zeta_{j})\bigr) ζ j ⋅ (( Y ± δH ) ζ j ) = ζ j ⋅ ( Y ζ j ) ± δ ( ζ j ⋅ ( H ζ j ) ) ; summing over j j j with claims 2 and 3 of Properties of Finite Sums , Q ( Y ± δ H ) = Q ( Y ) ± δ Q ( H ) Q(Y\pm\delta H)=Q(Y)\pm\delta\,Q(H) Q ( Y ± δH ) = Q ( Y ) ± δ Q ( H ) , and in the same way Q ( Y − Y ′ ) = Q ( Y ) − Q ( Y ′ ) Q(Y-Y')=Q(Y)-Q(Y') Q ( Y − Y ′ ) = Q ( Y ) − Q ( Y ′ ) for Y , Y ′ ∈ S ( d ) Y,Y'\in\mathcal{S}(d) Y , Y ′ ∈ S ( d ) . For μ ∈ D \mu\in\mathcal{D} μ ∈ D the translation Hessian H E ( μ ) H_{\mathcal{E}}(\mu) H E ( μ ) lies in S ( d ) \mathcal{S}(d) S ( d ) by Penalty Pairs on the Wasserstein Space: the Penalty, Its Score, and Their Domains §hessian , and we write h ( μ ) = Q ( H E ( μ ) ) = t r ( Γ ⊤ Γ H E ( μ ) ) h(\mu)=Q(H_{\mathcal{E}}(\mu))=\mathrm{tr}(\Gamma^{\top}\Gamma H_{\mathcal{E}}(\mu)) h ( μ ) = Q ( H E ( μ )) = tr ( Γ ⊤ Γ H E ( μ )) ; by (Growth), ∣ h ( μ ) ∣ ≤ C ( 1 + ∣ E ( μ ) ∣ ) |h(\mu)|\le C(1+|\mathcal{E}(\mu)|) ∣ h ( μ ) ∣ ≤ C ( 1 + ∣ E ( μ ) ∣ ) . Finally 0 < 1 2 0<\tfrac{1}{2} 0 < 2 1 by claims 8 and 7 of Elementary Order Arithmetic in an Ordered Field .
(0.5) Elementary inequalities in δ \delta δ . Let δ ∈ R \delta\in\mathbb{R} δ ∈ R with 0 < δ < 1 0<\delta<1 0 < δ < 1 . Then 0 < θ δ ≤ δ < 1 0<\theta\delta\le\delta<1 0 < θ δ ≤ δ < 1 , because 0 < θ ≤ 1 0<\theta\le1 0 < θ ≤ 1 . Consequently 0 < 1 + θ δ 0<1+\theta\delta 0 < 1 + θ δ , 0 ≤ 1 − θ δ ≤ 1 0\le1-\theta\delta\le1 0 ≤ 1 − θ δ ≤ 1 , δ 2 + θ δ 2 2 > 0 \tfrac{\delta}{2}+\tfrac{\theta\delta^{2}}{2}>0 2 δ + 2 θ δ 2 > 0 , δ 2 − θ δ 2 2 = δ 2 ( 1 − θ δ ) ≥ 0 \tfrac{\delta}{2}-\tfrac{\theta\delta^{2}}{2}=\tfrac{\delta}{2}(1-\theta\delta)\ge0 2 δ − 2 θ δ 2 = 2 δ ( 1 − θ δ ) ≥ 0 , and δ − θ δ 2 2 = δ − δ 2 θ δ ≥ δ 2 > 0 \delta-\tfrac{\theta\delta^{2}}{2}=\delta-\tfrac{\delta}{2}\,\theta\delta\ge\tfrac{\delta}{2}>0 δ − 2 θ δ 2 = δ − 2 δ θ δ ≥ 2 δ > 0 . For real m , s m,s m , s one has 2 m s ≤ m 2 + s 2 2ms\le m^{2}+s^{2} 2 m s ≤ m 2 + s 2 , since 0 ≤ ( m − s ) 2 = m 2 − 2 m s + s 2 0\le(m-s)^{2}=m^{2}-2ms+s^{2} 0 ≤ ( m − s ) 2 = m 2 − 2 m s + s 2 .
(0.6) Expanded form of the shifts. Let ( ν , q ) ∈ V ( D Σ ) (\nu,q)\in\mathcal{V}(\mathcal{D}_{\Sigma}) ( ν , q ) ∈ V ( D Σ ) , r ∈ R r\in\mathbb{R} r ∈ R , Y ∈ S ( d ) Y\in\mathcal{S}(d) Y ∈ S ( d ) , δ > 0 \delta>0 δ > 0 , and write σ = Σ ( ν ) \sigma=\Sigma(\nu) σ = Σ ( ν ) . By The Bundle of Vector Fields over a Set of Measures, Second-Order Equation Operators on the Wasserstein Space, and Their Delta-Shifts §shifted , the formula of The Discounted Hamilton-Jacobi Equation with Common Noise and a Penalty Drift on the Wasserstein Space §operator , linearity of Q Q Q (0.4), and ⟨ σ , q ± δ σ ⟩ ν = ⟨ σ , q ⟩ ν ± δ ∥ σ ∥ ν 2 \langle\sigma,q\pm\delta\sigma\rangle_{\nu}=\langle\sigma,q\rangle_{\nu}\pm\delta\lVert\sigma\rVert_{\nu}^{2} ⟨ σ , q ± δ σ ⟩ ν = ⟨ σ , q ⟩ ν ± δ ∥ σ ∥ ν 2 (0.1),
F δ − ( ν , r , q , Y ) = λ 0 r + λ 0 δ E ( ν ) − 1 2 Q ( Y ) − 1 2 δ h ( ν ) + θ 2 ∥ q + δ σ ∥ ν 2 + ⟨ σ , q ⟩ ν + δ ∥ σ ∥ ν 2 − g ( ν ) , (0a) F^{-}_{\delta}(\nu,r,q,Y)=\lambda_{0}r+\lambda_{0}\delta\,\mathcal{E}(\nu)-\frac{1}{2}Q(Y)-\frac{1}{2}\delta\,h(\nu)+\frac{\theta}{2}\lVert q+\delta\sigma\rVert_{\nu}^{2}+\langle\sigma,q\rangle_{\nu}+\delta\lVert\sigma\rVert_{\nu}^{2}-g(\nu),\tag{0a} F δ − ( ν , r , q , Y ) = λ 0 r + λ 0 δ E ( ν ) − 2 1 Q ( Y ) − 2 1 δ h ( ν ) + 2 θ ∥ q + δ σ ∥ ν 2 + ⟨ σ , q ⟩ ν + δ ∥ σ ∥ ν 2 − g ( ν ) , ( 0a )
F δ + ( ν , r , q , Y ) = λ 0 r − λ 0 δ E ( ν ) − 1 2 Q ( Y ) + 1 2 δ h ( ν ) + θ 2 ∥ q − δ σ ∥ ν 2 + ⟨ σ , q ⟩ ν − δ ∥ σ ∥ ν 2 − g ( ν ) . (0c) F^{+}_{\delta}(\nu,r,q,Y)=\lambda_{0}r-\lambda_{0}\delta\,\mathcal{E}(\nu)-\frac{1}{2}Q(Y)+\frac{1}{2}\delta\,h(\nu)+\frac{\theta}{2}\lVert q-\delta\sigma\rVert_{\nu}^{2}+\langle\sigma,q\rangle_{\nu}-\delta\lVert\sigma\rVert_{\nu}^{2}-g(\nu).\tag{0c} F δ + ( ν , r , q , Y ) = λ 0 r − λ 0 δ E ( ν ) − 2 1 Q ( Y ) + 2 1 δ h ( ν ) + 2 θ ∥ q − δ σ ∥ ν 2 + ⟨ σ , q ⟩ ν − δ ∥ σ ∥ ν 2 − g ( ν ) . ( 0c )
Expanding ∥ q ± δ σ ∥ ν 2 = ∥ q ∥ ν 2 ± 2 δ ⟨ σ , q ⟩ ν + δ 2 ∥ σ ∥ ν 2 \lVert q\pm\delta\sigma\rVert_{\nu}^{2}=\lVert q\rVert_{\nu}^{2}\pm2\delta\langle\sigma,q\rangle_{\nu}+\delta^{2}\lVert\sigma\rVert_{\nu}^{2} ∥ q ± δ σ ∥ ν 2 = ∥ q ∥ ν 2 ± 2 δ ⟨ σ , q ⟩ ν + δ 2 ∥ σ ∥ ν 2 by (0.1) gives
F δ − ( ν , r , q , Y ) = λ 0 r + λ 0 δ E ( ν ) − 1 2 Q ( Y ) − 1 2 δ h ( ν ) + θ 2 ∥ q ∥ ν 2 + ( 1 + θ δ ) ⟨ σ , q ⟩ ν + ( δ + θ δ 2 2 ) ∥ σ ∥ ν 2 − g ( ν ) , (0b) F^{-}_{\delta}(\nu,r,q,Y)=\lambda_{0}r+\lambda_{0}\delta\,\mathcal{E}(\nu)-\frac{1}{2}Q(Y)-\frac{1}{2}\delta\,h(\nu)+\frac{\theta}{2}\lVert q\rVert_{\nu}^{2}+(1+\theta\delta)\langle\sigma,q\rangle_{\nu}+\Bigl(\delta+\frac{\theta\delta^{2}}{2}\Bigr)\lVert\sigma\rVert_{\nu}^{2}-g(\nu),\tag{0b} F δ − ( ν , r , q , Y ) = λ 0 r + λ 0 δ E ( ν ) − 2 1 Q ( Y ) − 2 1 δ h ( ν ) + 2 θ ∥ q ∥ ν 2 + ( 1 + θ δ ) ⟨ σ , q ⟩ ν + ( δ + 2 θ δ 2 ) ∥ σ ∥ ν 2 − g ( ν ) , ( 0b )
F δ + ( ν , r , q , Y ) = λ 0 r − λ 0 δ E ( ν ) − 1 2 Q ( Y ) + 1 2 δ h ( ν ) + θ 2 ∥ q ∥ ν 2 + ( 1 − θ δ ) ⟨ σ , q ⟩ ν − ( δ − θ δ 2 2 ) ∥ σ ∥ ν 2 − g ( ν ) . (0d) F^{+}_{\delta}(\nu,r,q,Y)=\lambda_{0}r-\lambda_{0}\delta\,\mathcal{E}(\nu)-\frac{1}{2}Q(Y)+\frac{1}{2}\delta\,h(\nu)+\frac{\theta}{2}\lVert q\rVert_{\nu}^{2}+(1-\theta\delta)\langle\sigma,q\rangle_{\nu}-\Bigl(\delta-\frac{\theta\delta^{2}}{2}\Bigr)\lVert\sigma\rVert_{\nu}^{2}-g(\nu).\tag{0d} F δ + ( ν , r , q , Y ) = λ 0 r − λ 0 δ E ( ν ) − 2 1 Q ( Y ) + 2 1 δ h ( ν ) + 2 θ ∥ q ∥ ν 2 + ( 1 − θ δ ) ⟨ σ , q ⟩ ν − ( δ − 2 θ δ 2 ) ∥ σ ∥ ν 2 − g ( ν ) . ( 0d )
Step 1 (Local strict properness). Let R > 0 R>0 R > 0 , ( ν , q ) ∈ V ( D Σ ) (\nu,q)\in\mathcal{V}(\mathcal{D}_{\Sigma}) ( ν , q ) ∈ V ( D Σ ) , Y ∈ S ( d ) Y\in\mathcal{S}(d) Y ∈ S ( d ) and − R ≤ s ≤ r ≤ R -R\le s\le r\le R − R ≤ s ≤ r ≤ R . By the formula of The Discounted Hamilton-Jacobi Equation with Common Noise and a Penalty Drift on the Wasserstein Space §operator all terms other than λ 0 r \lambda_{0}r λ 0 r and λ 0 s \lambda_{0}s λ 0 s cancel, so F ( ν , r , q , Y ) − F ( ν , s , q , Y ) = λ 0 ( r − s ) F(\nu,r,q,Y)-F(\nu,s,q,Y)=\lambda_{0}(r-s) F ( ν , r , q , Y ) − F ( ν , s , q , Y ) = λ 0 ( r − s ) . Hence the positive real λ 0 \lambda_{0} λ 0 is a properness constant for F F F at R R R , for every R > 0 R>0 R > 0 , and F F F is locally strictly proper .
Step 2 (Shift-coercivity). Let δ , R ∈ R \delta,R\in\mathbb{R} δ , R ∈ R with 0 < δ < 1 0<\delta<1 0 < δ < 1 and 0 < R 0<R 0 < R . Put
A = 2 λ 0 R + 1 2 β Γ R + 1 2 C ( 1 + R ) + M g + R 2 2 , B = R + 2 A + R 2 2 δ , C δ , R = 1 + 4 ( R + B ) δ ; A=2\lambda_{0}R+\frac{1}{2}\beta_{\Gamma}R+\frac{1}{2}C(1+R)+M_{g}+\frac{R^{2}}{2},\qquad B=R+2A+\frac{R^{2}}{2\delta},\qquad C_{\delta,R}=1+\frac{4(R+B)}{\delta}; A = 2 λ 0 R + 2 1 β Γ R + 2 1 C ( 1 + R ) + M g + 2 R 2 , B = R + 2 A + 2 δ R 2 , C δ , R = 1 + δ 4 ( R + B ) ;
these are nonnegative reals by (0.2), (0.3) and (0.4). We show that C δ , R C_{\delta,R} C δ , R is a score bound for F F F at ( δ , R ) (\delta,R) ( δ , R ) .
Let ξ = ( ν , r , q , Y ) \xi=(\nu,r,q,Y) ξ = ( ν , r , q , Y ) and η = ( ν ′ , r ′ , q ′ , Y ′ ) \eta=(\nu',r',q',Y') η = ( ν ′ , r ′ , q ′ , Y ′ ) be R R R -bounded test data with F δ − ( ξ ) − F δ + ( η ) < R F^{-}_{\delta}(\xi)-F^{+}_{\delta}(\eta)<R F δ − ( ξ ) − F δ + ( η ) < R . By Test Data for an Intrinsic Second-Order Equation Operator on the Wasserstein Space and the Admissible Sets §admissible every member of S δ , R − S^{-}_{\delta,R} S δ , R − is such a ξ \xi ξ for some η \eta η , and every member of S δ , R + S^{+}_{\delta,R} S δ , R + is such an η \eta η for some ξ \xi ξ ; so it suffices to show s ≤ C δ , R s\le C_{\delta,R} s ≤ C δ , R and s ′ ≤ C δ , R s'\le C_{\delta,R} s ′ ≤ C δ , R , where s = ∥ Σ ( ν ) ∥ ν s=\lVert\Sigma(\nu)\rVert_{\nu} s = ∥ Σ ( ν ) ∥ ν and s ′ = ∥ Σ ( ν ′ ) ∥ ν ′ s'=\lVert\Sigma(\nu')\rVert_{\nu'} s ′ = ∥ Σ ( ν ′ ) ∥ ν ′ . R R R -boundedness gives ∣ r ∣ , ∣ r ′ ∣ < R |r|,|r'|<R ∣ r ∣ , ∣ r ′ ∣ < R , ∣ E ( ν ) ∣ , ∣ E ( ν ′ ) ∣ < R |\mathcal{E}(\nu)|,|\mathcal{E}(\nu')|<R ∣ E ( ν ) ∣ , ∣ E ( ν ′ ) ∣ < R , ∥ q ∥ ν , ∥ q ′ ∥ ν ′ < R \lVert q\rVert_{\nu},\lVert q'\rVert_{\nu'}<R ∥ q ∥ ν , ∥ q ′ ∥ ν ′ < R and ∥ Y ∥ , ∥ Y ′ ∥ < R \lVert Y\rVert,\lVert Y'\rVert<R ∥ Y ∥ , ∥ Y ′ ∥ < R .
Lower bound for F δ − ( ξ ) F^{-}_{\delta}(\xi) F δ − ( ξ ) . In (0a) we have λ 0 r ≥ − λ 0 R \lambda_{0}r\ge-\lambda_{0}R λ 0 r ≥ − λ 0 R ; λ 0 δ E ( ν ) ≥ − λ 0 δ ∣ E ( ν ) ∣ ≥ − λ 0 R \lambda_{0}\delta\mathcal{E}(\nu)\ge-\lambda_{0}\delta|\mathcal{E}(\nu)|\ge-\lambda_{0}R λ 0 δ E ( ν ) ≥ − λ 0 δ ∣ E ( ν ) ∣ ≥ − λ 0 R as δ < 1 \delta<1 δ < 1 ; − 1 2 Q ( Y ) ≥ − 1 2 β Γ ∥ Y ∥ ≥ − 1 2 β Γ R -\tfrac{1}{2}Q(Y)\ge-\tfrac{1}{2}\beta_{\Gamma}\lVert Y\rVert\ge-\tfrac{1}{2}\beta_{\Gamma}R − 2 1 Q ( Y ) ≥ − 2 1 β Γ ∥ Y ∥ ≥ − 2 1 β Γ R by (0.4); − 1 2 δ h ( ν ) ≥ − 1 2 ∣ h ( ν ) ∣ ≥ − 1 2 C ( 1 + R ) -\tfrac{1}{2}\delta h(\nu)\ge-\tfrac{1}{2}|h(\nu)|\ge-\tfrac{1}{2}C(1+R) − 2 1 δ h ( ν ) ≥ − 2 1 ∣ h ( ν ) ∣ ≥ − 2 1 C ( 1 + R ) by (0.4); θ 2 ∥ q + δ Σ ( ν ) ∥ ν 2 ≥ 0 \tfrac{\theta}{2}\lVert q+\delta\Sigma(\nu)\rVert_{\nu}^{2}\ge0 2 θ ∥ q + δ Σ ( ν ) ∥ ν 2 ≥ 0 ; ⟨ Σ ( ν ) , q ⟩ ν ≥ − s R \langle\Sigma(\nu),q\rangle_{\nu}\ge-sR ⟨ Σ ( ν ) , q ⟩ ν ≥ − s R by the Cauchy-Schwarz inequality of (0.1); and − g ( ν ) ≥ − M g -g(\nu)\ge-M_{g} − g ( ν ) ≥ − M g . Hence, since δ s 2 ≥ δ 2 s 2 \delta s^{2}\ge\tfrac{\delta}{2}s^{2} δ s 2 ≥ 2 δ s 2 ,
F δ − ( ξ ) ≥ δ s 2 − R s − A ≥ φ ( s ) − A , where φ ( x ) = δ 2 x 2 − R x ( x ∈ R ) . F^{-}_{\delta}(\xi)\ \ge\ \delta s^{2}-Rs-A\ \ge\ \varphi(s)-A,\qquad\text{where }\varphi(x)=\frac{\delta}{2}x^{2}-Rx\ \ (x\in\mathbb{R}). F δ − ( ξ ) ≥ δ s 2 − R s − A ≥ φ ( s ) − A , where φ ( x ) = 2 δ x 2 − R x ( x ∈ R ) .
Upper bound for F δ + ( η ) F^{+}_{\delta}(\eta) F δ + ( η ) . In (0d) we have λ 0 r ′ ≤ λ 0 R \lambda_{0}r'\le\lambda_{0}R λ 0 r ′ ≤ λ 0 R ; − λ 0 δ E ( ν ′ ) ≤ λ 0 R -\lambda_{0}\delta\mathcal{E}(\nu')\le\lambda_{0}R − λ 0 δ E ( ν ′ ) ≤ λ 0 R ; − 1 2 Q ( Y ′ ) ≤ 1 2 β Γ R -\tfrac{1}{2}Q(Y')\le\tfrac{1}{2}\beta_{\Gamma}R − 2 1 Q ( Y ′ ) ≤ 2 1 β Γ R ; 1 2 δ h ( ν ′ ) ≤ 1 2 C ( 1 + R ) \tfrac{1}{2}\delta h(\nu')\le\tfrac{1}{2}C(1+R) 2 1 δ h ( ν ′ ) ≤ 2 1 C ( 1 + R ) ; θ 2 ∥ q ′ ∥ ν ′ 2 ≤ R 2 2 \tfrac{\theta}{2}\lVert q'\rVert_{\nu'}^{2}\le\tfrac{R^{2}}{2} 2 θ ∥ q ′ ∥ ν ′ 2 ≤ 2 R 2 as θ ≤ 1 \theta\le1 θ ≤ 1 ; ( 1 − θ δ ) ⟨ Σ ( ν ′ ) , q ′ ⟩ ν ′ ≤ ( 1 − θ δ ) s ′ R ≤ R s ′ (1-\theta\delta)\langle\Sigma(\nu'),q'\rangle_{\nu'}\le(1-\theta\delta)s'R\le Rs' ( 1 − θ δ ) ⟨ Σ ( ν ′ ) , q ′ ⟩ ν ′ ≤ ( 1 − θ δ ) s ′ R ≤ R s ′ by Cauchy-Schwarz and (0.5); − ( δ − θ δ 2 2 ) s ′ 2 ≤ − δ 2 s ′ 2 -(\delta-\tfrac{\theta\delta^{2}}{2})s'^{2}\le-\tfrac{\delta}{2}s'^{2} − ( δ − 2 θ δ 2 ) s ′ 2 ≤ − 2 δ s ′ 2 by (0.5); and − g ( ν ′ ) ≤ M g -g(\nu')\le M_{g} − g ( ν ′ ) ≤ M g . Hence F δ + ( η ) ≤ A − φ ( s ′ ) F^{+}_{\delta}(\eta)\le A-\varphi(s') F δ + ( η ) ≤ A − φ ( s ′ ) .
Conclusion. For every real x x x , 0 ≤ δ 2 ( x − R δ ) 2 = φ ( x ) + R 2 2 δ 0\le\tfrac{\delta}{2}(x-\tfrac{R}{\delta})^{2}=\varphi(x)+\tfrac{R^{2}}{2\delta} 0 ≤ 2 δ ( x − δ R ) 2 = φ ( x ) + 2 δ R 2 , so φ ( x ) ≥ − R 2 2 δ \varphi(x)\ge-\tfrac{R^{2}}{2\delta} φ ( x ) ≥ − 2 δ R 2 . From the two bounds, φ ( s ) + φ ( s ′ ) − 2 A ≤ F δ − ( ξ ) − F δ + ( η ) < R \varphi(s)+\varphi(s')-2A\le F^{-}_{\delta}(\xi)-F^{+}_{\delta}(\eta)<R φ ( s ) + φ ( s ′ ) − 2 A ≤ F δ − ( ξ ) − F δ + ( η ) < R , hence φ ( s ) < R + 2 A − φ ( s ′ ) ≤ B \varphi(s)<R+2A-\varphi(s')\le B φ ( s ) < R + 2 A − φ ( s ′ ) ≤ B and likewise φ ( s ′ ) < B \varphi(s')<B φ ( s ′ ) < B . Now let x ≥ 0 x\ge0 x ≥ 0 with φ ( x ) < B \varphi(x)<B φ ( x ) < B , and suppose x > C δ , R x>C_{\delta,R} x > C δ , R . Then x > 1 x>1 x > 1 , x > 4 R δ x>\tfrac{4R}{\delta} x > δ 4 R and x > 4 B δ x>\tfrac{4B}{\delta} x > δ 4 B , all three numbers being at most C δ , R C_{\delta,R} C δ , R . The second gives R x < δ 4 x 2 Rx<\tfrac{\delta}{4}x^{2} R x < 4 δ x 2 , so φ ( x ) > δ 4 x 2 \varphi(x)>\tfrac{\delta}{4}x^{2} φ ( x ) > 4 δ x 2 ; the first gives x 2 > x x^{2}>x x 2 > x , so δ 4 x 2 > δ 4 x \tfrac{\delta}{4}x^{2}>\tfrac{\delta}{4}x 4 δ x 2 > 4 δ x ; and the third gives δ 4 x > B \tfrac{\delta}{4}x>B 4 δ x > B . Thus φ ( x ) > B \varphi(x)>B φ ( x ) > B , a contradiction. Hence x ≤ C δ , R x\le C_{\delta,R} x ≤ C δ , R ; applied to x = s x=s x = s and x = s ′ x=s' x = s ′ this proves the claim. As δ , R \delta,R δ , R were arbitrary, F F F satisfies the shift-coercivity condition .
Step 3 (An estimate along a coupling). Claim. Let ν ′ , ν ∈ P 2 ( R d ) \nu',\nu\in\mathcal{P}_{2}(\mathbb{R}^{d}) ν ′ , ν ∈ P 2 ( R d ) , π ∈ Π ( ν ′ , ν ) \pi\in\Pi(\nu',\nu) π ∈ Π ( ν ′ , ν ) , a , b ∈ L 2 ( ν ′ ; R d ) a,b\in L^{2}(\nu';\mathbb{R}^{d}) a , b ∈ L 2 ( ν ′ ; R d ) , c ∈ L 2 ( ν ; R d ) c\in L^{2}(\nu;\mathbb{R}^{d}) c ∈ L 2 ( ν ; R d ) , and let ρ > 0 \rho>0 ρ > 0 with ∥ a ∥ ν ′ ≤ ρ \lVert a\rVert_{\nu'}\le\rho ∥ a ∥ ν ′ ≤ ρ . Write K \mathcal{K} K for the cross pairing and D = ∫ R d + d ∥ b ( x ) − c ( y ) ∥ 2 π ( d z ) D=\int_{\mathbb{R}^{d+d}}\lVert b(x)-c(y)\rVert^{2}\,\pi(dz) D = ∫ R d + d ∥ b ( x ) − c ( y ) ∥ 2 π ( d z ) for the discrepancy, a nonnegative real by The Discrepancy of Two Square-Integrable Vector Fields Along a Coupling of Their Base Measures §well-defined . Then
( ⟨ a , b ⟩ ν ′ − K ( a , c , π ) ) 2 ≤ ρ 2 D . \bigl(\langle a,b\rangle_{\nu'}-\mathcal{K}(a,c,\pi)\bigr)^{2}\le\rho^{2}D . ( ⟨ a , b ⟩ ν ′ − K ( a , c , π ) ) 2 ≤ ρ 2 D .
Proof. Put β = ⟨ a , b ⟩ ν ′ − K ( a , c , π ) \beta=\langle a,b\rangle_{\nu'}-\mathcal{K}(a,c,\pi) β = ⟨ a , b ⟩ ν ′ − K ( a , c , π ) and let t ∈ R t\in\mathbb{R} t ∈ R . The field t a + b ta+b t a + b lies in L 2 ( ν ′ ; R d ) L^{2}(\nu';\mathbb{R}^{d}) L 2 ( ν ′ ; R d ) , and its discrepancy with c c c along π \pi π is nonnegative by The Discrepancy of Two Square-Integrable Vector Fields Along a Coupling of Their Base Measures §well-defined . By The Cross Pairing of Two Square-Integrable Vector Fields Along a Coupling §polarisation that discrepancy equals ∥ t a + b ∥ ν ′ 2 − 2 K ( t a + b , c , π ) + ∥ c ∥ ν 2 \lVert ta+b\rVert_{\nu'}^{2}-2\mathcal{K}(ta+b,c,\pi)+\lVert c\rVert_{\nu}^{2} ∥ t a + b ∥ ν ′ 2 − 2 K ( t a + b , c , π ) + ∥ c ∥ ν 2 ; by (0.1), ∥ t a + b ∥ ν ′ 2 = t 2 ∥ a ∥ ν ′ 2 + 2 t ⟨ a , b ⟩ ν ′ + ∥ b ∥ ν ′ 2 \lVert ta+b\rVert_{\nu'}^{2}=t^{2}\lVert a\rVert_{\nu'}^{2}+2t\langle a,b\rangle_{\nu'}+\lVert b\rVert_{\nu'}^{2} ∥ t a + b ∥ ν ′ 2 = t 2 ∥ a ∥ ν ′ 2 + 2 t ⟨ a , b ⟩ ν ′ + ∥ b ∥ ν ′ 2 ; by The Cross Pairing of Two Square-Integrable Vector Fields Along a Coupling §linear , K ( t a + b , c , π ) = t K ( a , c , π ) + K ( b , c , π ) \mathcal{K}(ta+b,c,\pi)=t\mathcal{K}(a,c,\pi)+\mathcal{K}(b,c,\pi) K ( t a + b , c , π ) = t K ( a , c , π ) + K ( b , c , π ) ; and by The Cross Pairing of Two Square-Integrable Vector Fields Along a Coupling §polarisation again, ∥ b ∥ ν ′ 2 − 2 K ( b , c , π ) + ∥ c ∥ ν 2 = D \lVert b\rVert_{\nu'}^{2}-2\mathcal{K}(b,c,\pi)+\lVert c\rVert_{\nu}^{2}=D ∥ b ∥ ν ′ 2 − 2 K ( b , c , π ) + ∥ c ∥ ν 2 = D . Therefore
0 ≤ t 2 ∥ a ∥ ν ′ 2 + 2 t β + D ≤ t 2 ρ 2 + 2 t β + D for every t ∈ R . 0\le t^{2}\lVert a\rVert_{\nu'}^{2}+2t\beta+D\le t^{2}\rho^{2}+2t\beta+D\qquad\text{for every }t\in\mathbb{R}. 0 ≤ t 2 ∥ a ∥ ν ′ 2 + 2 tβ + D ≤ t 2 ρ 2 + 2 tβ + D for every t ∈ R .
With t = − β ρ − 2 t=-\beta\rho^{-2} t = − β ρ − 2 this reads 0 ≤ β 2 ρ − 2 − 2 β 2 ρ − 2 + D = D − β 2 ρ − 2 0\le\beta^{2}\rho^{-2}-2\beta^{2}\rho^{-2}+D=D-\beta^{2}\rho^{-2} 0 ≤ β 2 ρ − 2 − 2 β 2 ρ − 2 + D = D − β 2 ρ − 2 , and multiplying by ρ 2 > 0 \rho^{2}>0 ρ 2 > 0 gives the claim.
Consequence. If, for every n ∈ N n\in\mathbb{N} n ∈ N , ν n ′ ∈ P 2 ( R d ) \nu'_{n}\in\mathcal{P}_{2}(\mathbb{R}^{d}) ν n ′ ∈ P 2 ( R d ) , π n ∈ Π ( ν n ′ , ν ) \pi_{n}\in\Pi(\nu'_{n},\nu) π n ∈ Π ( ν n ′ , ν ) , a n , b n ∈ L 2 ( ν n ′ ; R d ) a_{n},b_{n}\in L^{2}(\nu'_{n};\mathbb{R}^{d}) a n , b n ∈ L 2 ( ν n ′ ; R d ) with ∥ a n ∥ ν n ′ ≤ ρ \lVert a_{n}\rVert_{\nu'_{n}}\le\rho ∥ a n ∥ ν n ′ ≤ ρ , and the discrepancies D n D_{n} D n of b n b_{n} b n and c c c along π n \pi_{n} π n converge to 0 0 0 , then β n = ⟨ a n , b n ⟩ ν n ′ − K ( a n , c , π n ) \beta_{n}=\langle a_{n},b_{n}\rangle_{\nu'_{n}}-\mathcal{K}(a_{n},c,\pi_{n}) β n = ⟨ a n , b n ⟩ ν n ′ − K ( a n , c , π n ) converges to 0 0 0 : given ε > 0 \varepsilon>0 ε > 0 choose N N N with D n < ε 2 ρ − 2 D_{n}<\varepsilon^{2}\rho^{-2} D n < ε 2 ρ − 2 for n ≥ N n\ge N n ≥ N ; then ∣ β n ∣ 2 = β n 2 ≤ ρ 2 D n < ε 2 |\beta_{n}|^{2}=\beta_{n}^{2}\le\rho^{2}D_{n}<\varepsilon^{2} ∣ β n ∣ 2 = β n 2 ≤ ρ 2 D n < ε 2 (claim 1 of Nonnegativity of Squares in an Ordered Field ), so ∣ β n ∣ < ε |\beta_{n}|<\varepsilon ∣ β n ∣ < ε .
Step 4 (Shift-semicontinuity). Let δ , R ∈ R \delta,R\in\mathbb{R} δ , R ∈ R with 0 < δ < 1 0<\delta<1 0 < δ < 1 and 0 < R 0<R 0 < R , let ξ n = ( ν n , r n , q n , Y n ) \xi_{n}=(\nu_{n},r_{n},q_{n},Y_{n}) ξ n = ( ν n , r n , q n , Y n ) (n ∈ N n\in\mathbb{N} n ∈ N ) and ξ = ( ν , r , q , Y ) \xi=(\nu,r,q,Y) ξ = ( ν , r , q , Y ) be test data, and let ( π n ) (\pi_{n}) ( π n ) be couplings such that ( ξ n ) (\xi_{n}) ( ξ n ) converges to ξ \xi ξ along ( π n ) (\pi_{n}) ( π n ) with score bounded by R R R . Write σ n = Σ ( ν n ) \sigma_{n}=\Sigma(\nu_{n}) σ n = Σ ( ν n ) , σ = Σ ( ν ) \sigma=\Sigma(\nu) σ = Σ ( ν ) , and K n ( a , c ) = K ( a , c , π n ) \mathcal{K}_{n}(a,c)=\mathcal{K}(a,c,\pi_{n}) K n ( a , c ) = K ( a , c , π n ) . By that clause: every ξ n \xi_{n} ξ n is R R R -bounded, so ∣ E ( ν n ) ∣ < R |\mathcal{E}(\nu_{n})|<R ∣ E ( ν n ) ∣ < R and ∥ q n ∥ ν n < R \lVert q_{n}\rVert_{\nu_{n}}<R ∥ q n ∥ ν n < R ; ∥ σ n ∥ ν n ≤ R \lVert\sigma_{n}\rVert_{\nu_{n}}\le R ∥ σ n ∥ ν n ≤ R ; ( π n ) (\pi_{n}) ( π n ) is a sequence of couplings of vanishing cost , lim I ( π n ) = 0 \lim I(\pi_{n})=0 lim I ( π n ) = 0 ; ( q n ) (q_{n}) ( q n ) converges strongly to q q q , i.e. the discrepancies D n D_{n} D n of q n q_{n} q n and q q q along π n \pi_{n} π n converge to 0 0 0 ; ( σ n ) (\sigma_{n}) ( σ n ) converges weakly to σ \sigma σ , i.e. lim K n ( σ n , w ) = ⟨ σ , w ⟩ ν \lim\mathcal{K}_{n}(\sigma_{n},w)=\langle\sigma,w\rangle_{\nu} lim K n ( σ n , w ) = ⟨ σ , w ⟩ ν for every w ∈ L 2 ( ν ; R d ) w\in L^{2}(\nu;\mathbb{R}^{d}) w ∈ L 2 ( ν ; R d ) ; r n → r r_{n}\to r r n → r ; and Y n → Y Y_{n}\to Y Y n → Y in the metric d S ( d ) ( Z , Z ′ ) = ∥ Z − Z ′ ∥ d_{\mathcal{S}(d)}(Z,Z')=\lVert Z-Z'\rVert d S ( d ) ( Z , Z ′ ) = ∥ Z − Z ′ ∥ of Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §matrices . Limits of sums, products and scalar multiples of convergent real sequences are computed by Arithmetic of Limits of Real Sequences .
(4.1) W 2 ( ν n , ν ) → 0 W_{2}(\nu_{n},\nu)\to0 W 2 ( ν n , ν ) → 0 . By The Quadratic Wasserstein Distance on Euclidean Space §distance , W 2 ( ν n , ν ) 2 ≤ I ( π n ) W_{2}(\nu_{n},\nu)^{2}\le I(\pi_{n}) W 2 ( ν n , ν ) 2 ≤ I ( π n ) . Given ε > 0 \varepsilon>0 ε > 0 , choose N N N with I ( π n ) < ε 2 I(\pi_{n})<\varepsilon^{2} I ( π n ) < ε 2 for n ≥ N n\ge N n ≥ N ; then W 2 ( ν n , ν ) 2 < ε 2 W_{2}(\nu_{n},\nu)^{2}<\varepsilon^{2} W 2 ( ν n , ν ) 2 < ε 2 , so W 2 ( ν n , ν ) < ε W_{2}(\nu_{n},\nu)<\varepsilon W 2 ( ν n , ν ) < ε . The distance is symmetric by The Quadratic Wasserstein Distance is a Metric on the Wasserstein Space §metric , so also W 2 ( ν , ν n ) < ε W_{2}(\nu,\nu_{n})<\varepsilon W 2 ( ν , ν n ) < ε for n ≥ N n\ge N n ≥ N .
(4.2) Convergent terms. (a) Q ( Y n ) → Q ( Y ) Q(Y_{n})\to Q(Y) Q ( Y n ) → Q ( Y ) , because ∣ Q ( Y n ) − Q ( Y ) ∣ = ∣ Q ( Y n − Y ) ∣ ≤ β Γ ∥ Y n − Y ∥ |Q(Y_{n})-Q(Y)|=|Q(Y_{n}-Y)|\le\beta_{\Gamma}\lVert Y_{n}-Y\rVert ∣ Q ( Y n ) − Q ( Y ) ∣ = ∣ Q ( Y n − Y ) ∣ ≤ β Γ ∥ Y n − Y ∥ by (0.4), and ∥ Y n − Y ∥ → 0 \lVert Y_{n}-Y\rVert\to0 ∥ Y n − Y ∥ → 0 . (b) h ( ν n ) → h ( ν ) h(\nu_{n})\to h(\nu) h ( ν n ) → h ( ν ) : put R ′ = R + ∣ E ( ν ) ∣ > 0 R'=R+|\mathcal{E}(\nu)|>0 R ′ = R + ∣ E ( ν ) ∣ > 0 and D R ′ = { μ ∈ D : ∣ E ( μ ) ∣ ≤ R ′ } \mathcal{D}_{R'}=\{\mu\in\mathcal{D}:|\mathcal{E}(\mu)|\le R'\} D R ′ = { μ ∈ D : ∣ E ( μ ) ∣ ≤ R ′ } , which contains ν \nu ν and every ν n \nu_{n} ν n . By (Hessian continuity) the restriction of h h h to D R ′ \mathcal{D}_{R'} D R ′ is continuous at ν \nu ν relative to D R ′ \mathcal{D}_{R'} D R ′ : for ε > 0 \varepsilon>0 ε > 0 there is γ > 0 \gamma>0 γ > 0 with ∣ h ( μ ) − h ( ν ) ∣ < ε |h(\mu)-h(\nu)|<\varepsilon ∣ h ( μ ) − h ( ν ) ∣ < ε whenever μ ∈ D R ′ \mu\in\mathcal{D}_{R'} μ ∈ D R ′ and W 2 ( ν , μ ) < γ W_{2}(\nu,\mu)<\gamma W 2 ( ν , μ ) < γ ; by (4.1) this applies to μ = ν n \mu=\nu_{n} μ = ν n for all large n n n . (c) g ( ν n ) → g ( ν ) g(\nu_{n})\to g(\nu) g ( ν n ) → g ( ν ) : by (Running cost) and Uniformly Continuous Map Between Metric Spaces , for ε > 0 \varepsilon>0 ε > 0 there is γ > 0 \gamma>0 γ > 0 with ∣ g ( μ ) − g ( μ ′ ) ∣ < ε |g(\mu)-g(\mu')|<\varepsilon ∣ g ( μ ) − g ( μ ′ ) ∣ < ε whenever W 2 ( μ , μ ′ ) < γ W_{2}(\mu,\mu')<\gamma W 2 ( μ , μ ′ ) < γ , the metric on R \mathbb{R} R being that of The Absolute Value Metric on the Real Line ; apply (4.1). (d) ∥ q n ∥ ν n 2 → ∥ q ∥ ν 2 \lVert q_{n}\rVert_{\nu_{n}}^{2}\to\lVert q\rVert_{\nu}^{2} ∥ q n ∥ ν n 2 → ∥ q ∥ ν 2 : by the consequence in Step 3 with ρ = R \rho=R ρ = R , a n = b n = q n a_{n}=b_{n}=q_{n} a n = b n = q n , c = q c=q c = q , the numbers β n = ∥ q n ∥ ν n 2 − K n ( q n , q ) \beta_{n}=\lVert q_{n}\rVert_{\nu_{n}}^{2}-\mathcal{K}_{n}(q_{n},q) β n = ∥ q n ∥ ν n 2 − K n ( q n , q ) tend to 0 0 0 ; by The Cross Pairing of Two Square-Integrable Vector Fields Along a Coupling §polarisation , D n = ∥ q n ∥ ν n 2 − 2 K n ( q n , q ) + ∥ q ∥ ν 2 = 2 β n − ∥ q n ∥ ν n 2 + ∥ q ∥ ν 2 D_{n}=\lVert q_{n}\rVert_{\nu_{n}}^{2}-2\mathcal{K}_{n}(q_{n},q)+\lVert q\rVert_{\nu}^{2}=2\beta_{n}-\lVert q_{n}\rVert_{\nu_{n}}^{2}+\lVert q\rVert_{\nu}^{2} D n = ∥ q n ∥ ν n 2 − 2 K n ( q n , q ) + ∥ q ∥ ν 2 = 2 β n − ∥ q n ∥ ν n 2 + ∥ q ∥ ν 2 , so ∥ q n ∥ ν n 2 = 2 β n − D n + ∥ q ∥ ν 2 → ∥ q ∥ ν 2 \lVert q_{n}\rVert_{\nu_{n}}^{2}=2\beta_{n}-D_{n}+\lVert q\rVert_{\nu}^{2}\to\lVert q\rVert_{\nu}^{2} ∥ q n ∥ ν n 2 = 2 β n − D n + ∥ q ∥ ν 2 → ∥ q ∥ ν 2 . (e) ⟨ σ n , q n ⟩ ν n → ⟨ σ , q ⟩ ν \langle\sigma_{n},q_{n}\rangle_{\nu_{n}}\to\langle\sigma,q\rangle_{\nu} ⟨ σ n , q n ⟩ ν n → ⟨ σ , q ⟩ ν : by the consequence in Step 3 with ρ = R \rho=R ρ = R , a n = σ n a_{n}=\sigma_{n} a n = σ n , b n = q n b_{n}=q_{n} b n = q n , c = q c=q c = q , the difference ⟨ σ n , q n ⟩ ν n − K n ( σ n , q ) \langle\sigma_{n},q_{n}\rangle_{\nu_{n}}-\mathcal{K}_{n}(\sigma_{n},q) ⟨ σ n , q n ⟩ ν n − K n ( σ n , q ) tends to 0 0 0 , and K n ( σ n , q ) → ⟨ σ , q ⟩ ν \mathcal{K}_{n}(\sigma_{n},q)\to\langle\sigma,q\rangle_{\nu} K n ( σ n , q ) → ⟨ σ , q ⟩ ν by weak convergence with w = q w=q w = q .
(4.3) Lower semicontinuous terms. (a) For every ε > 0 \varepsilon>0 ε > 0 there is N N N with E ( ν n ) > E ( ν ) − ε \mathcal{E}(\nu_{n})>\mathcal{E}(\nu)-\varepsilon E ( ν n ) > E ( ν ) − ε for n ≥ N n\ge N n ≥ N : by (Semicontinuity) and Lower Semicontinuous Function on a Subset of a Metric Space , E \mathcal{E} E is lower semicontinuous at ν ∈ D \nu\in\mathcal{D} ν ∈ D relative to D \mathcal{D} D , which gives γ > 0 \gamma>0 γ > 0 with E ( ν ) − ε < E ( μ ) \mathcal{E}(\nu)-\varepsilon<\mathcal{E}(\mu) E ( ν ) − ε < E ( μ ) for μ ∈ D \mu\in\mathcal{D} μ ∈ D with W 2 ( ν , μ ) < γ W_{2}(\nu,\mu)<\gamma W 2 ( ν , μ ) < γ ; apply (4.1), as ν n ∈ D Σ ⊆ D \nu_{n}\in\mathcal{D}_{\Sigma}\subseteq\mathcal{D} ν n ∈ D Σ ⊆ D . (b) For every ε > 0 \varepsilon>0 ε > 0 there is N N N with ∥ σ n ∥ ν n 2 > ∥ σ ∥ ν 2 − ε \lVert\sigma_{n}\rVert_{\nu_{n}}^{2}>\lVert\sigma\rVert_{\nu}^{2}-\varepsilon ∥ σ n ∥ ν n 2 > ∥ σ ∥ ν 2 − ε for n ≥ N n\ge N n ≥ N : by The Cross Pairing of Two Square-Integrable Vector Fields Along a Coupling §polarisation and The Discrepancy of Two Square-Integrable Vector Fields Along a Coupling of Their Base Measures §well-defined , 0 ≤ ∥ σ n ∥ ν n 2 − 2 K n ( σ n , σ ) + ∥ σ ∥ ν 2 0\le\lVert\sigma_{n}\rVert_{\nu_{n}}^{2}-2\mathcal{K}_{n}(\sigma_{n},\sigma)+\lVert\sigma\rVert_{\nu}^{2} 0 ≤ ∥ σ n ∥ ν n 2 − 2 K n ( σ n , σ ) + ∥ σ ∥ ν 2 , so ∥ σ n ∥ ν n 2 ≥ 2 K n ( σ n , σ ) − ∥ σ ∥ ν 2 \lVert\sigma_{n}\rVert_{\nu_{n}}^{2}\ge2\mathcal{K}_{n}(\sigma_{n},\sigma)-\lVert\sigma\rVert_{\nu}^{2} ∥ σ n ∥ ν n 2 ≥ 2 K n ( σ n , σ ) − ∥ σ ∥ ν 2 ; the right side converges to 2 ∥ σ ∥ ν 2 − ∥ σ ∥ ν 2 = ∥ σ ∥ ν 2 2\lVert\sigma\rVert_{\nu}^{2}-\lVert\sigma\rVert_{\nu}^{2}=\lVert\sigma\rVert_{\nu}^{2} 2 ∥ σ ∥ ν 2 − ∥ σ ∥ ν 2 = ∥ σ ∥ ν 2 by weak convergence with w = σ w=\sigma w = σ , so it exceeds ∥ σ ∥ ν 2 − ε \lVert\sigma\rVert_{\nu}^{2}-\varepsilon ∥ σ ∥ ν 2 − ε for large n n n .
(4.4) The lower shift. By (0b), F δ − ( ξ n ) = u n + v n F^{-}_{\delta}(\xi_{n})=u_{n}+v_{n} F δ − ( ξ n ) = u n + v n and F δ − ( ξ ) = u + v F^{-}_{\delta}(\xi)=u+v F δ − ( ξ ) = u + v , where
u n = λ 0 r n − 1 2 Q ( Y n ) − 1 2 δ h ( ν n ) + θ 2 ∥ q n ∥ ν n 2 + ( 1 + θ δ ) ⟨ σ n , q n ⟩ ν n − g ( ν n ) , v n = λ 0 δ E ( ν n ) + ( δ + θ δ 2 2 ) ∥ σ n ∥ ν n 2 , u_{n}=\lambda_{0}r_{n}-\frac{1}{2}Q(Y_{n})-\frac{1}{2}\delta\,h(\nu_{n})+\frac{\theta}{2}\lVert q_{n}\rVert_{\nu_{n}}^{2}+(1+\theta\delta)\langle\sigma_{n},q_{n}\rangle_{\nu_{n}}-g(\nu_{n}),\qquad v_{n}=\lambda_{0}\delta\,\mathcal{E}(\nu_{n})+\Bigl(\delta+\frac{\theta\delta^{2}}{2}\Bigr)\lVert\sigma_{n}\rVert_{\nu_{n}}^{2}, u n = λ 0 r n − 2 1 Q ( Y n ) − 2 1 δ h ( ν n ) + 2 θ ∥ q n ∥ ν n 2 + ( 1 + θ δ ) ⟨ σ n , q n ⟩ ν n − g ( ν n ) , v n = λ 0 δ E ( ν n ) + ( δ + 2 θ δ 2 ) ∥ σ n ∥ ν n 2 ,
and u , v u,v u , v are the same expressions at ξ \xi ξ . By (4.2) and the limit laws, u n → u u_{n}\to u u n → u . Let k = λ 0 δ + δ + θ δ 2 2 + 1 > 0 k=\lambda_{0}\delta+\delta+\tfrac{\theta\delta^{2}}{2}+1>0 k = λ 0 δ + δ + 2 θ δ 2 + 1 > 0 . Given ε > 0 \varepsilon>0 ε > 0 , (4.3) applied with ε k − 1 \varepsilon k^{-1} ε k − 1 gives, for large n n n , v n ≥ v − ( λ 0 δ + δ + θ δ 2 2 ) ε k − 1 ≥ v − ε v_{n}\ge v-(\lambda_{0}\delta+\delta+\tfrac{\theta\delta^{2}}{2})\varepsilon k^{-1}\ge v-\varepsilon v n ≥ v − ( λ 0 δ + δ + 2 θ δ 2 ) ε k − 1 ≥ v − ε , the coefficients being nonnegative by (0.5). Now let c ∈ R c\in\mathbb{R} c ∈ R be such that for every ε > 0 \varepsilon>0 ε > 0 there is N N N with F δ − ( ξ n ) ≤ c + ε F^{-}_{\delta}(\xi_{n})\le c+\varepsilon F δ − ( ξ n ) ≤ c + ε for n ≥ N n\ge N n ≥ N . Fix ε > 0 \varepsilon>0 ε > 0 and choose n n n so large that F δ − ( ξ n ) ≤ c + ε F^{-}_{\delta}(\xi_{n})\le c+\varepsilon F δ − ( ξ n ) ≤ c + ε , ∣ u n − u ∣ < ε |u_{n}-u|<\varepsilon ∣ u n − u ∣ < ε and v n ≥ v − ε v_{n}\ge v-\varepsilon v n ≥ v − ε (the largest of three thresholds). Then
F δ − ( ξ ) = u + v < u n + ε + v n + ε = F δ − ( ξ n ) + 2 ε ≤ c + 3 ε . F^{-}_{\delta}(\xi)=u+v<u_{n}+\varepsilon+v_{n}+\varepsilon=F^{-}_{\delta}(\xi_{n})+2\varepsilon\le c+3\varepsilon . F δ − ( ξ ) = u + v < u n + ε + v n + ε = F δ − ( ξ n ) + 2 ε ≤ c + 3 ε .
As ε > 0 \varepsilon>0 ε > 0 is arbitrary, Comparison of Real Numbers with Arbitrary Positive Slack §slack-above gives F δ − ( ξ ) ≤ c F^{-}_{\delta}(\xi)\le c F δ − ( ξ ) ≤ c .
(4.5) The upper shift. By (0d), F δ + ( ξ n ) = u n ′ − v n ′ F^{+}_{\delta}(\xi_{n})=u'_{n}-v'_{n} F δ + ( ξ n ) = u n ′ − v n ′ and F δ + ( ξ ) = u ′ − v ′ F^{+}_{\delta}(\xi)=u'-v' F δ + ( ξ ) = u ′ − v ′ , where
u n ′ = λ 0 r n − 1 2 Q ( Y n ) + 1 2 δ h ( ν n ) + θ 2 ∥ q n ∥ ν n 2 + ( 1 − θ δ ) ⟨ σ n , q n ⟩ ν n − g ( ν n ) , v n ′ = λ 0 δ E ( ν n ) + ( δ − θ δ 2 2 ) ∥ σ n ∥ ν n 2 , u'_{n}=\lambda_{0}r_{n}-\frac{1}{2}Q(Y_{n})+\frac{1}{2}\delta\,h(\nu_{n})+\frac{\theta}{2}\lVert q_{n}\rVert_{\nu_{n}}^{2}+(1-\theta\delta)\langle\sigma_{n},q_{n}\rangle_{\nu_{n}}-g(\nu_{n}),\qquad v'_{n}=\lambda_{0}\delta\,\mathcal{E}(\nu_{n})+\Bigl(\delta-\frac{\theta\delta^{2}}{2}\Bigr)\lVert\sigma_{n}\rVert_{\nu_{n}}^{2}, u n ′ = λ 0 r n − 2 1 Q ( Y n ) + 2 1 δ h ( ν n ) + 2 θ ∥ q n ∥ ν n 2 + ( 1 − θ δ ) ⟨ σ n , q n ⟩ ν n − g ( ν n ) , v n ′ = λ 0 δ E ( ν n ) + ( δ − 2 θ δ 2 ) ∥ σ n ∥ ν n 2 ,
and u ′ , v ′ u',v' u ′ , v ′ are the same expressions at ξ \xi ξ . As in (4.4), u n ′ → u ′ u'_{n}\to u' u n ′ → u ′ , and, the coefficients λ 0 δ \lambda_{0}\delta λ 0 δ and δ − θ δ 2 2 \delta-\tfrac{\theta\delta^{2}}{2} δ − 2 θ δ 2 being nonnegative by (0.5), for every ε > 0 \varepsilon>0 ε > 0 we have v n ′ ≥ v ′ − ε v'_{n}\ge v'-\varepsilon v n ′ ≥ v ′ − ε for large n n n . Let c ∈ R c\in\mathbb{R} c ∈ R be such that for every ε > 0 \varepsilon>0 ε > 0 there is N N N with c − ε ≤ F δ + ( ξ n ) c-\varepsilon\le F^{+}_{\delta}(\xi_{n}) c − ε ≤ F δ + ( ξ n ) for n ≥ N n\ge N n ≥ N . Fix ε > 0 \varepsilon>0 ε > 0 and choose n n n so large that c − ε ≤ F δ + ( ξ n ) c-\varepsilon\le F^{+}_{\delta}(\xi_{n}) c − ε ≤ F δ + ( ξ n ) , ∣ u n ′ − u ′ ∣ < ε |u'_{n}-u'|<\varepsilon ∣ u n ′ − u ′ ∣ < ε and v n ′ ≥ v ′ − ε v'_{n}\ge v'-\varepsilon v n ′ ≥ v ′ − ε . Then
c − ε ≤ u n ′ − v n ′ < u ′ + ε − v ′ + ε = F δ + ( ξ ) + 2 ε , c-\varepsilon\le u'_{n}-v'_{n}<u'+\varepsilon-v'+\varepsilon=F^{+}_{\delta}(\xi)+2\varepsilon, c − ε ≤ u n ′ − v n ′ < u ′ + ε − v ′ + ε = F δ + ( ξ ) + 2 ε ,
so c − 3 ε ≤ F δ + ( ξ ) c-3\varepsilon\le F^{+}_{\delta}(\xi) c − 3 ε ≤ F δ + ( ξ ) , and Comparison of Real Numbers with Arbitrary Positive Slack §slack-below gives c ≤ F δ + ( ξ ) c\le F^{+}_{\delta}(\xi) c ≤ F δ + ( ξ ) .
By (4.4) and (4.5), F F F is shift-semicontinuous at ( δ , R ) (\delta,R) ( δ , R ) ; as δ , R \delta,R δ , R were arbitrary, F F F satisfies the shift-semicontinuity condition .
Step 5 (Second-order structure at uniquely mapped pairs). Let T = { t ∈ R : 0 ≤ t } T=\{t\in\mathbb{R}:0\le t\} T = { t ∈ R : 0 ≤ t } . The pair ( ω 1 , ω 2 ) (\omega_{1},\omega_{2}) ( ω 1 , ω 2 ) below is chosen first; it depends only on g , λ 0 , θ , C g,\lambda_{0},\theta,C g , λ 0 , θ , C , and we show it is a second-order structure pair for F F F at R R R for every positive R R R .
(5.1) The modulus ω 1 \omega_{1} ω 1 . For s ∈ T s\in T s ∈ T let G ( s ) = { ∣ g ( μ ′ ) − g ( ν ′ ) ∣ : μ ′ , ν ′ ∈ P 2 ( R d ) , W 2 ( μ ′ , ν ′ ) 2 ≤ s } G(s)=\{|g(\mu')-g(\nu')|:\mu',\nu'\in\mathcal{P}_{2}(\mathbb{R}^{d}),\ W_{2}(\mu',\nu')^{2}\le s\} G ( s ) = { ∣ g ( μ ′ ) − g ( ν ′ ) ∣ : μ ′ , ν ′ ∈ P 2 ( R d ) , W 2 ( μ ′ , ν ′ ) 2 ≤ s } . It contains 0 = ∣ g ( μ 0 ) − g ( μ 0 ) ∣ 0=|g(\mu_{0})-g(\mu_{0})| 0 = ∣ g ( μ 0 ) − g ( μ 0 ) ∣ , with μ 0 \mu_{0} μ 0 from (0.2), since W 2 ( μ 0 , μ 0 ) = 0 W_{2}(\mu_{0},\mu_{0})=0 W 2 ( μ 0 , μ 0 ) = 0 by The Quadratic Wasserstein Distance is a Metric on the Wasserstein Space §metric ; and it is bounded above by 2 M g 2M_{g} 2 M g by (0.3). Hence ω 1 ( s ) = sup G ( s ) \omega_{1}(s)=\sup G(s) ω 1 ( s ) = sup G ( s ) is defined by The Real Numbers: Standing Notation and Background §bounds , and 0 ≤ ω 1 ( s ) 0\le\omega_{1}(s) 0 ≤ ω 1 ( s ) because ω 1 ( s ) \omega_{1}(s) ω 1 ( s ) is an upper bound of G ( s ) ∋ 0 G(s)\ni0 G ( s ) ∋ 0 (Upper Bound and Least Upper Bound ). Given ε > 0 \varepsilon>0 ε > 0 , uniform continuity of g g g (Uniformly Continuous Map Between Metric Spaces ) gives γ > 0 \gamma>0 γ > 0 with ∣ g ( μ ′ ) − g ( ν ′ ) ∣ < ε |g(\mu')-g(\nu')|<\varepsilon ∣ g ( μ ′ ) − g ( ν ′ ) ∣ < ε whenever W 2 ( μ ′ , ν ′ ) < γ W_{2}(\mu',\nu')<\gamma W 2 ( μ ′ , ν ′ ) < γ . Put γ 1 = γ 2 / 4 > 0 \gamma_{1}=\gamma^{2}/4>0 γ 1 = γ 2 /4 > 0 . If t ∈ T t\in T t ∈ T and t ≤ γ 1 t\le\gamma_{1} t ≤ γ 1 , every element of G ( t ) G(t) G ( t ) comes from μ ′ , ν ′ \mu',\nu' μ ′ , ν ′ with W 2 ( μ ′ , ν ′ ) 2 ≤ ( γ / 2 ) 2 W_{2}(\mu',\nu')^{2}\le(\gamma/2)^{2} W 2 ( μ ′ , ν ′ ) 2 ≤ ( γ /2 ) 2 , so W 2 ( μ ′ , ν ′ ) ≤ γ / 2 < γ W_{2}(\mu',\nu')\le\gamma/2<\gamma W 2 ( μ ′ , ν ′ ) ≤ γ /2 < γ and the element is < ε <\varepsilon < ε ; thus ε \varepsilon ε is an upper bound of G ( t ) G(t) G ( t ) and ω 1 ( t ) ≤ ε \omega_{1}(t)\le\varepsilon ω 1 ( t ) ≤ ε , the supremum being the least upper bound. So ω 1 \omega_{1} ω 1 is a modulus of continuity , and by construction ∣ g ( μ ′ ) − g ( ν ′ ) ∣ ≤ ω 1 ( s ) |g(\mu')-g(\nu')|\le\omega_{1}(s) ∣ g ( μ ′ ) − g ( ν ′ ) ∣ ≤ ω 1 ( s ) whenever W 2 ( μ ′ , ν ′ ) 2 ≤ s W_{2}(\mu',\nu')^{2}\le s W 2 ( μ ′ , ν ′ ) 2 ≤ s .
(5.2) The function ω 2 \omega_{2} ω 2 . For t ∈ T t\in T t ∈ T and real α > 1 \alpha>1 α > 1 put ω 2 ( t , α ) = ( λ 0 + C + 4 C θ 2 α 2 ) t \omega_{2}(t,\alpha)=(\lambda_{0}+C+4C\theta^{2}\alpha^{2})\,t ω 2 ( t , α ) = ( λ 0 + C + 4 C θ 2 α 2 ) t . For each α > 1 \alpha>1 α > 1 the coefficient is nonnegative by (0.2), so t ↦ ω 2 ( t , α ) t\mapsto\omega_{2}(t,\alpha) t ↦ ω 2 ( t , α ) is a modulus of continuity by Linear Moduli of Continuity §modulus .
(5.3) The inequality. Let R > 0 R>0 R > 0 , and let α , δ , μ , ν , S , S ′ , r , X , Y \alpha,\delta,\mu,\nu,S,S',r,\mathbb{X},\mathbb{Y} α , δ , μ , ν , S , S ′ , r , X , Y be as in The Second-Order Structure Condition at Uniquely Mapped Pairs on the Wasserstein Space §pair : 1 < α 1<\alpha 1 < α , 0 < δ < 1 0<\delta<1 0 < δ < 1 , μ , ν ∈ D Σ \mu,\nu\in\mathcal{D}_{\Sigma} μ , ν ∈ D Σ with both ordered pairs ( μ , ν ) (\mu,\nu) ( μ , ν ) and ( ν , μ ) (\nu,\mu) ( ν , μ ) uniquely mapped , S S S an optimal map from μ \mu μ to ν \nu ν and S ′ S' S ′ one from ν \nu ν to μ \mu μ , r ∈ [ − R , R ] r\in[-R,R] r ∈ [ − R , R ] , and ( X , Y ) (\mathbb{X},\mathbb{Y}) ( X , Y ) admitted at α \alpha α . (The condition δ ( ∣ E ( μ ) ∣ + ∣ E ( ν ) ∣ ) ≤ R \delta(|\mathcal{E}(\mu)|+|\mathcal{E}(\nu)|)\le R δ ( ∣ E ( μ ) ∣ + ∣ E ( ν ) ∣ ) ≤ R will not be needed.) Write W = W 2 ( μ , ν ) W=W_{2}(\mu,\nu) W = W 2 ( μ , ν ) , σ = Σ ( μ ) ∈ L 2 ( μ ; R d ) \sigma=\Sigma(\mu)\in L^{2}(\mu;\mathbb{R}^{d}) σ = Σ ( μ ) ∈ L 2 ( μ ; R d ) , τ = Σ ( ν ) ∈ L 2 ( ν ; R d ) \tau=\Sigma(\nu)\in L^{2}(\nu;\mathbb{R}^{d}) τ = Σ ( ν ) ∈ L 2 ( ν ; R d ) , a = α ( i d − S ) ∈ L 2 ( μ ; R d ) a=\alpha(\mathrm{id}-S)\in L^{2}(\mu;\mathbb{R}^{d}) a = α ( id − S ) ∈ L 2 ( μ ; R d ) , b = α ( S ′ − i d ) = − α ( i d − S ′ ) ∈ L 2 ( ν ; R d ) b=\alpha(S'-\mathrm{id})=-\alpha(\mathrm{id}-S')\in L^{2}(\nu;\mathbb{R}^{d}) b = α ( S ′ − id ) = − α ( id − S ′ ) ∈ L 2 ( ν ; R d ) , e = ∣ E ( μ ) ∣ + ∣ E ( ν ) ∣ e=|\mathcal{E}(\mu)|+|\mathcal{E}(\nu)| e = ∣ E ( μ ) ∣ + ∣ E ( ν ) ∣ and t = δ ( e + 1 ) t=\delta(e+1) t = δ ( e + 1 ) , and let Δ \Delta Δ be the difference F δ − ( μ , r , a , X ) − F δ + ( ν , r , b , Y ) F^{-}_{\delta}(\mu,r,a,\mathbb{X})-F^{+}_{\delta}(\nu,r,b,\mathbb{Y}) F δ − ( μ , r , a , X ) − F δ + ( ν , r , b , Y ) to be bounded below.
Norms of the displacements. By The Optimal Map as a Square-Integrable Vector Field: Integrability, Transport Cost and Uniqueness of the Class §cost , ∥ i d − S ∥ μ 2 = W 2 \lVert\mathrm{id}-S\rVert_{\mu}^{2}=W^{2} ∥ id − S ∥ μ 2 = W 2 and ∥ i d − S ′ ∥ ν 2 = W 2 ( ν , μ ) 2 = W 2 \lVert\mathrm{id}-S'\rVert_{\nu}^{2}=W_{2}(\nu,\mu)^{2}=W^{2} ∥ id − S ′ ∥ ν 2 = W 2 ( ν , μ ) 2 = W 2 (symmetry, The Quadratic Wasserstein Distance is a Metric on the Wasserstein Space §metric ); hence ∥ i d − S ∥ μ = ∥ i d − S ′ ∥ ν = W \lVert\mathrm{id}-S\rVert_{\mu}=\lVert\mathrm{id}-S'\rVert_{\nu}=W ∥ id − S ∥ μ = ∥ id − S ′ ∥ ν = W and, by homogeneity (0.1) with ∣ α ∣ = α |\alpha|=\alpha ∣ α ∣ = α , ∥ a ∥ μ = ∥ b ∥ ν = α W \lVert a\rVert_{\mu}=\lVert b\rVert_{\nu}=\alpha W ∥ a ∥ μ = ∥ b ∥ ν = α W .
A bound on W 2 W^{2} W 2 . By Optimal Transport Maps and Uniquely Mapped Pairs of Probability Measures §map , S # μ = ν S_{\#}\mu=\nu S # μ = ν , so ∥ S ∥ μ 2 = ∫ R d ∥ S ∥ 2 d μ = M 2 ( ν ) \lVert S\rVert_{\mu}^{2}=\int_{\mathbb{R}^{d}}\lVert S\rVert^{2}\,d\mu=M_{2}(\nu) ∥ S ∥ μ 2 = ∫ R d ∥ S ∥ 2 d μ = M 2 ( ν ) by The Optimal Map as a Square-Integrable Vector Field: Integrability, Transport Cost and Uniqueness of the Class §square-integrable and (0.1); and ∥ i d ∥ μ 2 = M 2 ( μ ) \lVert\mathrm{id}\rVert_{\mu}^{2}=M_{2}(\mu) ∥ id ∥ μ 2 = M 2 ( μ ) by Basic Properties of the Tangent Space: Closed Subspace, the Identity Map Belongs to It, Second-Moment Limits, and Representation of Bounded Functionals on Gradients §identity . The parallelogram law Elementary Identities in a Real Inner Product Space §parallelogram gives W 2 = ∥ i d − S ∥ μ 2 ≤ ∥ i d − S ∥ μ 2 + ∥ i d + S ∥ μ 2 = 2 M 2 ( μ ) + 2 M 2 ( ν ) W^{2}=\lVert\mathrm{id}-S\rVert_{\mu}^{2}\le\lVert\mathrm{id}-S\rVert_{\mu}^{2}+\lVert\mathrm{id}+S\rVert_{\mu}^{2}=2M_{2}(\mu)+2M_{2}(\nu) W 2 = ∥ id − S ∥ μ 2 ≤ ∥ id − S ∥ μ 2 + ∥ id + S ∥ μ 2 = 2 M 2 ( μ ) + 2 M 2 ( ν ) , and (Growth) yields W 2 ≤ 2 C ( 2 + e ) W^{2}\le2C(2+e) W 2 ≤ 2 C ( 2 + e ) . Since 2 + e ≤ 2 ( 1 + e ) 2+e\le2(1+e) 2 + e ≤ 2 ( 1 + e ) , we get δ W 2 ≤ 2 C δ ( 2 + e ) ≤ 4 C t \delta W^{2}\le2C\delta(2+e)\le4Ct δ W 2 ≤ 2 C δ ( 2 + e ) ≤ 4 Ct .
Expansion of Δ \Delta Δ . Subtracting (0c) at ( ν , r , b , Y ) (\nu,r,b,\mathbb{Y}) ( ν , r , b , Y ) from (0a) at ( μ , r , a , X ) (\mu,r,a,\mathbb{X}) ( μ , r , a , X ) , the terms λ 0 r \lambda_{0}r λ 0 r cancel and
Δ = λ 0 δ ( E ( μ ) + E ( ν ) ) + 1 2 ( Q ( Y ) − Q ( X ) ) − 1 2 δ ( h ( μ ) + h ( ν ) ) + θ 2 ( ∥ a + δ σ ∥ μ 2 − ∥ b − δ τ ∥ ν 2 ) + ( ⟨ σ , a ⟩ μ − ⟨ τ , b ⟩ ν ) + δ ∥ σ ∥ μ 2 + δ ∥ τ ∥ ν 2 + g ( ν ) − g ( μ ) . \Delta=\lambda_{0}\delta\bigl(\mathcal{E}(\mu)+\mathcal{E}(\nu)\bigr)+\frac{1}{2}\bigl(Q(\mathbb{Y})-Q(\mathbb{X})\bigr)-\frac{1}{2}\delta\bigl(h(\mu)+h(\nu)\bigr)+\frac{\theta}{2}\Bigl(\lVert a+\delta\sigma\rVert_{\mu}^{2}-\lVert b-\delta\tau\rVert_{\nu}^{2}\Bigr)+\bigl(\langle\sigma,a\rangle_{\mu}-\langle\tau,b\rangle_{\nu}\bigr)+\delta\lVert\sigma\rVert_{\mu}^{2}+\delta\lVert\tau\rVert_{\nu}^{2}+g(\nu)-g(\mu). Δ = λ 0 δ ( E ( μ ) + E ( ν ) ) + 2 1 ( Q ( Y ) − Q ( X ) ) − 2 1 δ ( h ( μ ) + h ( ν ) ) + 2 θ ( ∥ a + δ σ ∥ μ 2 − ∥ b − δ τ ∥ ν 2 ) + ( ⟨ σ , a ⟩ μ − ⟨ τ , b ⟩ ν ) + δ ∥ σ ∥ μ 2 + δ ∥ τ ∥ ν 2 + g ( ν ) − g ( μ ) .
By (0.1), ∥ a + δ σ ∥ μ 2 = ∥ a ∥ μ 2 + 2 δ ⟨ a , σ ⟩ μ + δ 2 ∥ σ ∥ μ 2 \lVert a+\delta\sigma\rVert_{\mu}^{2}=\lVert a\rVert_{\mu}^{2}+2\delta\langle a,\sigma\rangle_{\mu}+\delta^{2}\lVert\sigma\rVert_{\mu}^{2} ∥ a + δ σ ∥ μ 2 = ∥ a ∥ μ 2 + 2 δ ⟨ a , σ ⟩ μ + δ 2 ∥ σ ∥ μ 2 and ∥ b − δ τ ∥ ν 2 = ∥ b ∥ ν 2 − 2 δ ⟨ b , τ ⟩ ν + δ 2 ∥ τ ∥ ν 2 \lVert b-\delta\tau\rVert_{\nu}^{2}=\lVert b\rVert_{\nu}^{2}-2\delta\langle b,\tau\rangle_{\nu}+\delta^{2}\lVert\tau\rVert_{\nu}^{2} ∥ b − δ τ ∥ ν 2 = ∥ b ∥ ν 2 − 2 δ ⟨ b , τ ⟩ ν + δ 2 ∥ τ ∥ ν 2 ; as ∥ a ∥ μ 2 = ∥ b ∥ ν 2 = α 2 W 2 \lVert a\rVert_{\mu}^{2}=\lVert b\rVert_{\nu}^{2}=\alpha^{2}W^{2} ∥ a ∥ μ 2 = ∥ b ∥ ν 2 = α 2 W 2 ,
θ 2 ( ∥ a + δ σ ∥ μ 2 − ∥ b − δ τ ∥ ν 2 ) = θ δ ⟨ a , σ ⟩ μ + θ δ ⟨ b , τ ⟩ ν + θ δ 2 2 ∥ σ ∥ μ 2 − θ δ 2 2 ∥ τ ∥ ν 2 . \frac{\theta}{2}\Bigl(\lVert a+\delta\sigma\rVert_{\mu}^{2}-\lVert b-\delta\tau\rVert_{\nu}^{2}\Bigr)=\theta\delta\langle a,\sigma\rangle_{\mu}+\theta\delta\langle b,\tau\rangle_{\nu}+\frac{\theta\delta^{2}}{2}\lVert\sigma\rVert_{\mu}^{2}-\frac{\theta\delta^{2}}{2}\lVert\tau\rVert_{\nu}^{2}. 2 θ ( ∥ a + δ σ ∥ μ 2 − ∥ b − δ τ ∥ ν 2 ) = θ δ ⟨ a , σ ⟩ μ + θ δ ⟨ b , τ ⟩ ν + 2 θ δ 2 ∥ σ ∥ μ 2 − 2 θ δ 2 ∥ τ ∥ ν 2 .
Bounds for the individual terms. (i) Monotonicity of the score: the pair is displacement convex by (Convexity), i.e. 0 0 0 -displacement convex , so A λ \lambda λ -Displacement Convex Penalty Pair Has a λ \lambda λ -Monotone Score Along Optimal Couplings §mapped with λ = 0 \lambda=0 λ = 0 gives 0 ≤ ⟨ σ , i d − S ⟩ μ + ⟨ τ , i d − S ′ ⟩ ν 0\le\langle\sigma,\mathrm{id}-S\rangle_{\mu}+\langle\tau,\mathrm{id}-S'\rangle_{\nu} 0 ≤ ⟨ σ , id − S ⟩ μ + ⟨ τ , id − S ′ ⟩ ν . By bilinearity (0.1), ⟨ σ , a ⟩ μ − ⟨ τ , b ⟩ ν = α ( ⟨ σ , i d − S ⟩ μ + ⟨ τ , i d − S ′ ⟩ ν ) ≥ 0 \langle\sigma,a\rangle_{\mu}-\langle\tau,b\rangle_{\nu}=\alpha\bigl(\langle\sigma,\mathrm{id}-S\rangle_{\mu}+\langle\tau,\mathrm{id}-S'\rangle_{\nu}\bigr)\ge0 ⟨ σ , a ⟩ μ − ⟨ τ , b ⟩ ν = α ( ⟨ σ , id − S ⟩ μ + ⟨ τ , id − S ′ ⟩ ν ) ≥ 0 . (ii) Traces: X ⪯ Y \mathbb{X}\preceq\mathbb{Y} X ⪯ Y by The Second-Order Structure Condition at Optimally Coupled Pairs on the Lift of the Wasserstein Space §admitted , the ordering being that of Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §matrices , so Q ( X ) ≤ Q ( Y ) Q(\mathbb{X})\le Q(\mathbb{Y}) Q ( X ) ≤ Q ( Y ) by The Trace as a Sum of Quadratic Forms, its Monotonicity and a Norm Bound §monotone , applied as in (0.4) with A = Γ A=\Gamma A = Γ , and 1 2 ( Q ( Y ) − Q ( X ) ) ≥ 0 \tfrac{1}{2}(Q(\mathbb{Y})-Q(\mathbb{X}))\ge0 2 1 ( Q ( Y ) − Q ( X )) ≥ 0 . (iii) Cross terms: by Cauchy-Schwarz (0.1) and (0.5) with m = ∥ σ ∥ μ m=\lVert\sigma\rVert_{\mu} m = ∥ σ ∥ μ , s = θ α W s=\theta\alpha W s = θ α W ,
θ δ ⟨ a , σ ⟩ μ ≥ − θ δ α W ∥ σ ∥ μ = − δ 2 2 m s ≥ − δ 2 ∥ σ ∥ μ 2 − δ 2 θ 2 α 2 W 2 , \theta\delta\langle a,\sigma\rangle_{\mu}\ge-\theta\delta\,\alpha W\lVert\sigma\rVert_{\mu}=-\frac{\delta}{2}\,2ms\ge-\frac{\delta}{2}\lVert\sigma\rVert_{\mu}^{2}-\frac{\delta}{2}\theta^{2}\alpha^{2}W^{2}, θ δ ⟨ a , σ ⟩ μ ≥ − θ δ α W ∥ σ ∥ μ = − 2 δ 2 m s ≥ − 2 δ ∥ σ ∥ μ 2 − 2 δ θ 2 α 2 W 2 ,
and in the same way θ δ ⟨ b , τ ⟩ ν ≥ − δ 2 ∥ τ ∥ ν 2 − δ 2 θ 2 α 2 W 2 \theta\delta\langle b,\tau\rangle_{\nu}\ge-\tfrac{\delta}{2}\lVert\tau\rVert_{\nu}^{2}-\tfrac{\delta}{2}\theta^{2}\alpha^{2}W^{2} θ δ ⟨ b , τ ⟩ ν ≥ − 2 δ ∥ τ ∥ ν 2 − 2 δ θ 2 α 2 W 2 . (iv) Collecting the score terms: the coefficient of ∥ σ ∥ μ 2 \lVert\sigma\rVert_{\mu}^{2} ∥ σ ∥ μ 2 becomes δ + θ δ 2 2 − δ 2 = δ 2 + θ δ 2 2 \delta+\tfrac{\theta\delta^{2}}{2}-\tfrac{\delta}{2}=\tfrac{\delta}{2}+\tfrac{\theta\delta^{2}}{2} δ + 2 θ δ 2 − 2 δ = 2 δ + 2 θ δ 2 and that of ∥ τ ∥ ν 2 \lVert\tau\rVert_{\nu}^{2} ∥ τ ∥ ν 2 becomes δ − θ δ 2 2 − δ 2 = δ 2 ( 1 − θ δ ) \delta-\tfrac{\theta\delta^{2}}{2}-\tfrac{\delta}{2}=\tfrac{\delta}{2}(1-\theta\delta) δ − 2 θ δ 2 − 2 δ = 2 δ ( 1 − θ δ ) , both nonnegative by (0.5), so these terms are ≥ 0 \ge0 ≥ 0 ; the remaining contribution is − θ 2 α 2 δ W 2 ≥ − 4 C θ 2 α 2 t -\theta^{2}\alpha^{2}\delta W^{2}\ge-4C\theta^{2}\alpha^{2}t − θ 2 α 2 δ W 2 ≥ − 4 C θ 2 α 2 t . (v) Penalty terms: λ 0 δ ( E ( μ ) + E ( ν ) ) ≥ − λ 0 δ e ≥ − λ 0 t \lambda_{0}\delta(\mathcal{E}(\mu)+\mathcal{E}(\nu))\ge-\lambda_{0}\delta e\ge-\lambda_{0}t λ 0 δ ( E ( μ ) + E ( ν )) ≥ − λ 0 δe ≥ − λ 0 t . (vi) Hessian terms: by (0.4), − 1 2 δ ( h ( μ ) + h ( ν ) ) ≥ − 1 2 δ C ( 2 + e ) ≥ − 1 2 δ ⋅ 2 C ( 1 + e ) = − C t -\tfrac{1}{2}\delta(h(\mu)+h(\nu))\ge-\tfrac{1}{2}\delta\,C(2+e)\ge-\tfrac{1}{2}\delta\cdot2C(1+e)=-Ct − 2 1 δ ( h ( μ ) + h ( ν )) ≥ − 2 1 δ C ( 2 + e ) ≥ − 2 1 δ ⋅ 2 C ( 1 + e ) = − Ct , using 0 ≤ C 0\le C 0 ≤ C (0.2) and 2 + e ≤ 2 ( 1 + e ) 2+e\le2(1+e) 2 + e ≤ 2 ( 1 + e ) . (vii) Running cost: W 2 ≤ α W 2 ≤ α W 2 + α − 1 W^{2}\le\alpha W^{2}\le\alpha W^{2}+\alpha^{-1} W 2 ≤ α W 2 ≤ α W 2 + α − 1 , as 1 < α 1<\alpha 1 < α , 0 ≤ W 2 0\le W^{2} 0 ≤ W 2 and 0 < α − 1 0<\alpha^{-1} 0 < α − 1 (claim 7 of Elementary Order Arithmetic in an Ordered Field ); so (5.1) with s = α W 2 + α − 1 s=\alpha W^{2}+\alpha^{-1} s = α W 2 + α − 1 gives g ( ν ) − g ( μ ) ≥ − ∣ g ( μ ) − g ( ν ) ∣ ≥ − ω 1 ( α W 2 + α − 1 ) g(\nu)-g(\mu)\ge-|g(\mu)-g(\nu)|\ge-\omega_{1}(\alpha W^{2}+\alpha^{-1}) g ( ν ) − g ( μ ) ≥ − ∣ g ( μ ) − g ( ν ) ∣ ≥ − ω 1 ( α W 2 + α − 1 ) .
Adding (i)-(vii) to the expansion of Δ \Delta Δ ,
− ω 1 ( α W 2 ( μ , ν ) 2 + α − 1 ) − ω 2 ( δ ( ∣ E ( μ ) ∣ + ∣ E ( ν ) ∣ + 1 ) , α ) ≤ F δ − ( μ , r , α ( i d − S ) , X ) − F δ + ( ν , r , α ( S ′ − i d ) , Y ) . -\omega_{1}\bigl(\alpha W_{2}(\mu,\nu)^{2}+\alpha^{-1}\bigr)-\omega_{2}\bigl(\delta(|\mathcal{E}(\mu)|+|\mathcal{E}(\nu)|+1),\alpha\bigr)\le F^{-}_{\delta}\bigl(\mu,r,\alpha(\mathrm{id}-S),\mathbb{X}\bigr)-F^{+}_{\delta}\bigl(\nu,r,\alpha(S'-\mathrm{id}),\mathbb{Y}\bigr). − ω 1 ( α W 2 ( μ , ν ) 2 + α − 1 ) − ω 2 ( δ ( ∣ E ( μ ) ∣ + ∣ E ( ν ) ∣ + 1 ) , α ) ≤ F δ − ( μ , r , α ( id − S ) , X ) − F δ + ( ν , r , α ( S ′ − id ) , Y ) .
Hence ( ω 1 , ω 2 ) (\omega_{1},\omega_{2}) ( ω 1 , ω 2 ) is a second-order structure pair for F F F at every R > 0 R>0 R > 0 , and F F F satisfies the second-order structure condition at uniquely mapped pairs .
Steps 1, 2, 4 and 5 together prove the conclusion of the statement.