Each result cited below is universally quantified over the data in its own statement, and is applied to the data named at the point of use. The rules for adding inequalities, for multiplying them by nonnegative or positive real numbers, and for handling absolute values, from Elementary Order Arithmetic in an Ordered Field , Elementary Arithmetic in an Ordered Field and Properties of the Absolute Value in an Ordered Field , are used without further mention; so are the facts that a square of a real number is nonnegative (claim 2 of Nonnegativity of Squares in an Ordered Field ) and that for nonnegative reals a , b a,b a , b one has a < b a<b a < b , a ≤ b a\le b a ≤ b , a = b a=b a = b exactly when a 2 < b 2 a^{2}<b^{2} a 2 < b 2 , a 2 ≤ b 2 a^{2}\le b^{2} a 2 ≤ b 2 , a 2 = b 2 a^{2}=b^{2} a 2 = b 2 respectively (claims 1, 2 and 3 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field ).
Step 0 (Notation and preliminary facts). Fix the data of the statement: V V V , λ 0 , σ , θ , L \lambda_{0},\sigma,\theta,L λ 0 , σ , θ , L with 0 < λ 0 0<\lambda_{0} 0 < λ 0 , 0 < σ 0<\sigma 0 < σ , 0 < θ ≤ 1 0<\theta\le1 0 < θ ≤ 1 , 0 ≤ L 0\le L 0 ≤ L , the natural number p p p and the matrix Γ ∈ M p × d ( R ) \Gamma\in\mathcal{M}_{p\times d}(\mathbb{R}) Γ ∈ M p × d ( R ) , the functions g g g and Φ \Phi Φ , the Langevin free-energy pair ( D , D Σ , E , Σ ) (\mathcal{D},\mathcal{D}_{\Sigma},\mathcal{E},\Sigma) ( D , D Σ , E , Σ ) with potential V V V and noise intensity σ \sigma σ , and the operator F F F ; write G = G Φ \mathcal{G}=\mathcal{G}_{\Phi} G = G Φ . Let F 0 F_{0} F 0 be the Langevin Hamilton-Jacobi operator with common noise, with potential V V V , noise intensity σ \sigma σ , discount λ 0 \lambda_{0} λ 0 , common-noise matrix Γ \Gamma Γ , control cost θ \theta θ and running cost g g g , with δ \delta δ -shifts F 0 , δ − , F 0 , δ + F^{-}_{0,\delta},F^{+}_{0,\delta} F 0 , δ − , F 0 , δ + relative to the pair. The number d d d is read in R \mathbb{R} R , where 1 ≤ d 1\le d 1 ≤ d , so 0 ≤ d L 2 0\le dL^{2} 0 ≤ d L 2 .
(P1) The pair. By The Langevin Free-Energy Pair is a Wasserstein-Coercive Penalty Pair: Growth Bounds, Continuity of the Translation Hessian, and the First Variation of the Penalty §pair , the pair is a penalty pair on P 2 ( R d ) \mathcal{P}_{2}(\mathbb{R}^{d}) P 2 ( R d ) ; in particular D Σ ⊆ D \mathcal{D}_{\Sigma}\subseteq\mathcal{D} D Σ ⊆ D by Penalty Pairs on the Wasserstein Space: the Penalty, Its Score, and Their Domains §pair . By The Langevin Free-Energy Pair is a Wasserstein-Coercive Penalty Pair: Growth Bounds, Continuity of the Translation Hessian, and the First Variation of the Penalty §coercive , D \mathcal{D} D has the map property. By The Langevin Free-Energy Pair of a Confining Potential on the Wasserstein Space §pair , D ⊆ P 2 E n t ( R d ) \mathcal{D}\subseteq\mathcal{P}_{2}^{\mathrm{Ent}}(\mathbb{R}^{d}) D ⊆ P 2 Ent ( R d ) , D Σ ⊆ P 2 I ( R d ) \mathcal{D}_{\Sigma}\subseteq\mathcal{P}_{2}^{\mathcal{I}}(\mathbb{R}^{d}) D Σ ⊆ P 2 I ( R d ) , and Σ ( ν ) = ∇ V + σ 2 2 ξ ν \Sigma(\nu)=\nabla V+\tfrac{\sigma^{2}}{2}\xi_{\nu} Σ ( ν ) = ∇ V + 2 σ 2 ξ ν for ν ∈ D Σ \nu\in\mathcal{D}_{\Sigma} ν ∈ D Σ , with the score ξ ν ∈ T ν ⊆ L 2 ( ν ; R d ) \xi_{\nu}\in T_{\nu}\subseteq L^{2}(\nu;\mathbb{R}^{d}) ξ ν ∈ T ν ⊆ L 2 ( ν ; R d ) . Hence every ν ∈ D Σ \nu\in\mathcal{D}_{\Sigma} ν ∈ D Σ has finite entropy, so is absolutely continuous by Basic Properties of the Entropy on the Wasserstein Space: Comparison with the Gaussian Relative Entropy, Lower Bound, Translation Invariance, Absolute Continuity, Closed Sublevel Sets and Lower Semicontinuity §absolutely-continuous , has finite Fisher information, and satisfies 0 ≤ G ( ν ) ≤ L 0\le\mathcal{G}(\nu)\le L 0 ≤ G ( ν ) ≤ L by The Density Cost of a Convex Lipschitz Integrand §cost .
(P2) The operators. By The Hamilton-Jacobi Equation with Common Noise for Controlled Langevin Dynamics in a Confining Potential on the Wasserstein Space §operator , F 0 F_{0} F 0 is the Hamilton-Jacobi operator with common noise and penalty drift of the pair with discount λ 0 \lambda_{0} λ 0 , common-noise matrix Γ \Gamma Γ , control cost θ \theta θ and running cost g g g , so its values are given by the formula of The Discounted Hamilton-Jacobi Equation with Common Noise and a Penalty Drift on the Wasserstein Space §operator ; and by The Langevin Hamilton-Jacobi Equation with Common Noise and a Density Cost on the Wasserstein Space §operator , F ( ν , r , q , Y ) = F 0 ( ν , r , q , Y ) − G ( ν ) F(\nu,r,q,Y)=F_{0}(\nu,r,q,Y)-\mathcal{G}(\nu) F ( ν , r , q , Y ) = F 0 ( ν , r , q , Y ) − G ( ν ) for every ( ν , q ) ∈ V ( D Σ ) (\nu,q)\in\mathcal{V}(\mathcal{D}_{\Sigma}) ( ν , q ) ∈ V ( D Σ ) , r ∈ R r\in\mathbb{R} r ∈ R and Y ∈ S ( d ) Y\in\mathcal{S}(d) Y ∈ S ( d ) . By The Bundle of Vector Fields over a Set of Measures, Second-Order Equation Operators on the Wasserstein Space, and Their Delta-Shifts §shifted , the δ \delta δ -shifts (δ > 0 \delta>0 δ > 0 ) evaluate the operator at the same measure ν \nu ν and change only the arguments r , q , Y r,q,Y r , q , Y ; hence
F δ − ( ν , r , q , Y ) = F 0 , δ − ( ν , r , q , Y ) − G ( ν ) , F δ + ( ν , r , q , Y ) = F 0 , δ + ( ν , r , q , Y ) − G ( ν ) . (0e) F^{-}_{\delta}(\nu,r,q,Y)=F^{-}_{0,\delta}(\nu,r,q,Y)-\mathcal{G}(\nu),\qquad F^{+}_{\delta}(\nu,r,q,Y)=F^{+}_{0,\delta}(\nu,r,q,Y)-\mathcal{G}(\nu).\tag{0e} F δ − ( ν , r , q , Y ) = F 0 , δ − ( ν , r , q , Y ) − G ( ν ) , F δ + ( ν , r , q , Y ) = F 0 , δ + ( ν , r , q , Y ) − G ( ν ) . ( 0e )
(P3) F 0 F_{0} F 0 is degenerate elliptic, by The Hamilton-Jacobi Operator with Common Noise and Penalty Drift is Degenerate Elliptic , whose hypotheses hold: the pair is a penalty pair by (P1), λ 0 \lambda_{0} λ 0 and θ \theta θ are positive, p ∈ N p\in\mathbb{N} p ∈ N and Γ ∈ M p × d ( R ) \Gamma\in\mathcal{M}_{p\times d}(\mathbb{R}) Γ ∈ M p × d ( R ) , g g g is a function P 2 ( R d ) → R \mathcal{P}_{2}(\mathbb{R}^{d})\to\mathbb{R} P 2 ( R d ) → R , and F 0 F_{0} F 0 is the operator named there for these data by (P2).
(P4) F 0 F_{0} F 0 is locally strictly proper and satisfies the shift-coercivity condition, the shift-semicontinuity condition and the second-order structure condition at uniquely mapped pairs. This is The Hamilton-Jacobi Operator with Common Noise and Penalty Drift Satisfies the Hypotheses of the Comparison Principle for a Displacement Convex Pair §conclusion , applied to the pair (a penalty pair by (P1)), with λ 0 \lambda_{0} λ 0 , with θ \theta θ (which satisfies 0 < θ ≤ 1 0<\theta\le1 0 < θ ≤ 1 ), with p p p , Γ \Gamma Γ , g g g and F 0 F_{0} F 0 (the operator named there, by (P2)); we discharge its hypotheses one by one. (Convexity) is The Langevin Free-Energy Pair is Displacement Convex, with Closed Score Along Couplings and Regular Penalised Maxima §convex . (Semicontinuity) holds by The Langevin Free-Energy Pair is a Wasserstein-Coercive Penalty Pair: Growth Bounds, Continuity of the Translation Hessian, and the First Variation of the Penalty §growth , which gives lower semicontinuity of E \mathcal{E} E on D \mathcal{D} D and a real number C 0 C_{0} C 0 with M 2 ( μ ) ≤ C 0 ( 1 + ∣ E ( μ ) ∣ ) M_{2}(\mu)\le C_{0}(1+|\mathcal{E}(\mu)|) M 2 ( μ ) ≤ C 0 ( 1 + ∣ E ( μ ) ∣ ) and ∣ t r H E ( μ ) ∣ ≤ C 0 ( 1 + ∣ E ( μ ) ∣ ) |\mathrm{tr}\,H_{\mathcal{E}}(\mu)|\le C_{0}(1+|\mathcal{E}(\mu)|) ∣ tr H E ( μ ) ∣ ≤ C 0 ( 1 + ∣ E ( μ ) ∣ ) for every μ ∈ D \mu\in\mathcal{D} μ ∈ D .
The common-noise trace as a finite sum of entries. For k ∈ [ p ] k\in[p] k ∈ [ p ] let w k ∈ R d w_{k}\in\mathbb{R}^{d} w k ∈ R d be the k k k th row of Γ \Gamma Γ , the point whose i i i th coordinate is w k , i = Γ k i w_{k,i}=\Gamma_{ki} w k , i = Γ ki for i ∈ [ d ] i\in[d] i ∈ [ d ] . Let μ ∈ D \mu\in\mathcal{D} μ ∈ D ; then H E ( μ ) ∈ S ( d ) H_{\mathcal{E}}(\mu)\in\mathcal{S}(d) H E ( μ ) ∈ S ( d ) by Penalty Pairs on the Wasserstein Space: the Penalty, Its Score, and Their Domains §hessian . Apply The Trace as a Sum of Quadratic Forms, its Monotonicity and a Norm Bound §rows with its matrix A A A taken to be Γ \Gamma Γ and X = H E ( μ ) X=H_{\mathcal{E}}(\mu) X = H E ( μ ) ; the letters m m m and p p p of that lemma are its own dimensions, and here its m m m is read as our p p p and its p p p as our d d d (both at least 1 1 1 , being natural numbers ), so that its rows a k a_{k} a k are our w k w_{k} w k . Together with claim 4 of Linearity of the Matrix-Vector Product and the Quadratic Form as a Double Sum , applied with n = d n=d n = d , M = H E ( μ ) M=H_{\mathcal{E}}(\mu) M = H E ( μ ) and both of its points w w w and z z z taken to be w k w_{k} w k , for each k ∈ [ p ] k\in[p] k ∈ [ p ] , and the commutativity and associativity of multiplication in R \mathbb{R} R , it gives
t r ( Γ ⊤ Γ H E ( μ ) ) = ∑ k = 1 p w k ⋅ ( H E ( μ ) w k ) = ∑ k = 1 p ∑ i = 1 d ∑ l = 1 d w k , i w k , l H E ( μ ) i l . (T) \mathrm{tr}\bigl(\Gamma^{\top}\Gamma H_{\mathcal{E}}(\mu)\bigr)=\sum_{k=1}^{p}w_{k}\cdot\bigl(H_{\mathcal{E}}(\mu)w_{k}\bigr)=\sum_{k=1}^{p}\ \sum_{i=1}^{d}\ \sum_{l=1}^{d}w_{k,i}w_{k,l}\,H_{\mathcal{E}}(\mu)_{il}.\tag{T} tr ( Γ ⊤ Γ H E ( μ ) ) = k = 1 ∑ p w k ⋅ ( H E ( μ ) w k ) = k = 1 ∑ p i = 1 ∑ d l = 1 ∑ d w k , i w k , l H E ( μ ) i l . ( T )
Put s Γ = ∑ k = 1 p ∑ i = 1 d ∑ l = 1 d ∣ w k , i ∣ ∣ w k , l ∣ s_{\Gamma}=\sum_{k=1}^{p}\sum_{i=1}^{d}\sum_{l=1}^{d}|w_{k,i}|\,|w_{k,l}| s Γ = ∑ k = 1 p ∑ i = 1 d ∑ l = 1 d ∣ w k , i ∣ ∣ w k , l ∣ , a real number that does not depend on μ \mu μ and is nonnegative by claim 5 of Properties of Finite Sums (applied to the innermost sums first), each summand being a product of nonnegative numbers.
For (Growth) , let μ ∈ D \mu\in\mathcal{D} μ ∈ D and write t μ = C 0 ( 1 + ∣ E ( μ ) ∣ ) t_{\mu}=C_{0}(1+|\mathcal{E}(\mu)|) t μ = C 0 ( 1 + ∣ E ( μ ) ∣ ) . For i , l ∈ [ d ] i,l\in[d] i , l ∈ [ d ] , The Langevin Free-Energy Pair is a Wasserstein-Coercive Penalty Pair: Growth Bounds, Continuity of the Translation Hessian, and the First Variation of the Penalty §entries gives ∣ H E ( μ ) i l ∣ ≤ t r H E ( μ ) ≤ ∣ t r H E ( μ ) ∣ ≤ t μ |H_{\mathcal{E}}(\mu)_{il}|\le\mathrm{tr}\,H_{\mathcal{E}}(\mu)\le|\mathrm{tr}\,H_{\mathcal{E}}(\mu)|\le t_{\mu} ∣ H E ( μ ) i l ∣ ≤ tr H E ( μ ) ≤ ∣ tr H E ( μ ) ∣ ≤ t μ . By claim 4 of Properties of the Absolute Value in an Ordered Field , used twice, and claim 5 of Elementary Arithmetic in an Ordered Field with the nonnegative multiplier ∣ w k , i ∣ ∣ w k , l ∣ |w_{k,i}|\,|w_{k,l}| ∣ w k , i ∣ ∣ w k , l ∣ , every summand of (T) satisfies
∣ w k , i w k , l H E ( μ ) i l ∣ = ∣ w k , i ∣ ∣ w k , l ∣ ∣ H E ( μ ) i l ∣ ≤ t μ ∣ w k , i ∣ ∣ w k , l ∣ . \bigl|w_{k,i}w_{k,l}\,H_{\mathcal{E}}(\mu)_{il}\bigr|=|w_{k,i}|\,|w_{k,l}|\,|H_{\mathcal{E}}(\mu)_{il}|\le t_{\mu}\,|w_{k,i}|\,|w_{k,l}| . w k , i w k , l H E ( μ ) i l = ∣ w k , i ∣ ∣ w k , l ∣ ∣ H E ( μ ) i l ∣ ≤ t μ ∣ w k , i ∣ ∣ w k , l ∣.
Claim 2 of Comparison and Absolute Value Bounds for Finite Sums of Real Numbers bounds the absolute value of each of the three nested sums in (T) by the sum of the absolute values of its summands, claim 1 of that lemma carries these bounds through the enclosing sums, and claim 3 of Properties of Finite Sums (applied to each of the three sums) takes out the factor t μ t_{\mu} t μ ; so
∣ t r ( Γ ⊤ Γ H E ( μ ) ) ∣ ≤ ∑ k = 1 p ∑ i = 1 d ∑ l = 1 d t μ ∣ w k , i ∣ ∣ w k , l ∣ = s Γ C 0 ( 1 + ∣ E ( μ ) ∣ ) . \bigl|\mathrm{tr}\bigl(\Gamma^{\top}\Gamma H_{\mathcal{E}}(\mu)\bigr)\bigr|\le\sum_{k=1}^{p}\sum_{i=1}^{d}\sum_{l=1}^{d}t_{\mu}\,|w_{k,i}|\,|w_{k,l}|=s_{\Gamma}\,C_{0}\bigl(1+|\mathcal{E}(\mu)|\bigr). tr ( Γ ⊤ Γ H E ( μ ) ) ≤ k = 1 ∑ p i = 1 ∑ d l = 1 ∑ d t μ ∣ w k , i ∣ ∣ w k , l ∣ = s Γ C 0 ( 1 + ∣ E ( μ ) ∣ ) .
Put C = ∣ C 0 ∣ ( 1 + s Γ ) C=|C_{0}|(1+s_{\Gamma}) C = ∣ C 0 ∣ ( 1 + s Γ ) , fixed from now on. Since 0 ≤ 1 + ∣ E ( μ ) ∣ 0\le1+|\mathcal{E}(\mu)| 0 ≤ 1 + ∣ E ( μ ) ∣ , C 0 ≤ ∣ C 0 ∣ C_{0}\le|C_{0}| C 0 ≤ ∣ C 0 ∣ , s Γ C 0 ≤ s Γ ∣ C 0 ∣ s_{\Gamma}C_{0}\le s_{\Gamma}|C_{0}| s Γ C 0 ≤ s Γ ∣ C 0 ∣ , and ∣ C 0 ∣ ≤ C |C_{0}|\le C ∣ C 0 ∣ ≤ C and s Γ ∣ C 0 ∣ ≤ C s_{\Gamma}|C_{0}|\le C s Γ ∣ C 0 ∣ ≤ C (as 0 ≤ s Γ 0\le s_{\Gamma} 0 ≤ s Γ and 0 ≤ ∣ C 0 ∣ 0\le|C_{0}| 0 ≤ ∣ C 0 ∣ ), we obtain M 2 ( μ ) ≤ C ( 1 + ∣ E ( μ ) ∣ ) M_{2}(\mu)\le C(1+|\mathcal{E}(\mu)|) M 2 ( μ ) ≤ C ( 1 + ∣ E ( μ ) ∣ ) and ∣ t r ( Γ ⊤ Γ H E ( μ ) ) ∣ ≤ C ( 1 + ∣ E ( μ ) ∣ ) |\mathrm{tr}(\Gamma^{\top}\Gamma H_{\mathcal{E}}(\mu))|\le C(1+|\mathcal{E}(\mu)|) ∣ tr ( Γ ⊤ Γ H E ( μ )) ∣ ≤ C ( 1 + ∣ E ( μ ) ∣ ) for every μ ∈ D \mu\in\mathcal{D} μ ∈ D , which is the hypothesis with the constant C C C .
For (Hessian continuity) , let R R R be positive and S R = { μ ∈ D : ∣ E ( μ ) ∣ ≤ R } S_{R}=\{\mu\in\mathcal{D}:|\mathcal{E}(\mu)|\le R\} S R = { μ ∈ D : ∣ E ( μ ) ∣ ≤ R } . For i , l ∈ [ d ] i,l\in[d] i , l ∈ [ d ] the restriction of μ ↦ H E ( μ ) i l \mu\mapsto H_{\mathcal{E}}(\mu)_{il} μ ↦ H E ( μ ) i l to S R S_{R} S R is continuous by The Langevin Free-Energy Pair is a Wasserstein-Coercive Penalty Pair: Growth Bounds, Continuity of the Translation Hessian, and the First Variation of the Penalty §entries , so for k ∈ [ p ] k\in[p] k ∈ [ p ] the restriction of μ ↦ w k , i w k , l H E ( μ ) i l \mu\mapsto w_{k,i}w_{k,l}\,H_{\mathcal{E}}(\mu)_{il} μ ↦ w k , i w k , l H E ( μ ) i l to S R S_{R} S R is continuous by claim 5 of Continuity of Sums and Products of Real-Valued Functions on a Metric Space (its multiple c f cf c f with c = w k , i w k , l c=w_{k,i}w_{k,l} c = w k , i w k , l ), taken in the metric space ( P 2 ( R d ) , W 2 ) (\mathcal{P}_{2}(\mathbb{R}^{d}),W_{2}) ( P 2 ( R d ) , W 2 ) with the subset S R S_{R} S R . A finite sum of functions continuous on S R S_{R} S R is continuous on S R S_{R} S R , by induction on the number of summands along the recursion of claim 1 of Properties of Finite Sums , each step being claim 5 of Continuity of Sums and Products of Real-Valued Functions on a Metric Space for its sum f + g f+g f + g of two functions. Applying this to the three nested sums of (T), the restriction of μ ↦ t r ( Γ ⊤ Γ H E ( μ ) ) \mu\mapsto\mathrm{tr}(\Gamma^{\top}\Gamma H_{\mathcal{E}}(\mu)) μ ↦ tr ( Γ ⊤ Γ H E ( μ )) to S R S_{R} S R is continuous.
(Running cost) is the hypothesis (Running cost) of the statement, boundedness and uniform continuity of g g g being understood in the same sense in both statements.
(0.1) Inner product spaces. For ν ∈ P 2 ( R d ) \nu\in\mathcal{P}_{2}(\mathbb{R}^{d}) ν ∈ P 2 ( R d ) the space L 2 ( ν ; R d ) L^{2}(\nu;\mathbb{R}^{d}) L 2 ( ν ; R d ) , with inner product ⟨ ⋅ , ⋅ ⟩ ν \langle\cdot,\cdot\rangle_{\nu} ⟨ ⋅ , ⋅ ⟩ ν and norm ∥ ⋅ ∥ ν \lVert\cdot\rVert_{\nu} ∥ ⋅ ∥ ν , is a real Hilbert space by Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §fields , in particular a real inner product space ; its norm satisfies ∥ x ∥ ν 2 = ⟨ x , x ⟩ ν \lVert x\rVert_{\nu}^{2}=\langle x,x\rangle_{\nu} ∥ x ∥ ν 2 = ⟨ x , x ⟩ ν by Real Inner Product Space §norm , and ∥ x ∥ ν 2 = ∫ R d ∥ x ∥ 2 d ν \lVert x\rVert_{\nu}^{2}=\int_{\mathbb{R}^{d}}\lVert x\rVert^{2}\,d\nu ∥ x ∥ ν 2 = ∫ R d ∥ x ∥ 2 d ν for a representative x x x , by the formula for the norm in Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §fields . Inner products are symmetric by condition (a) of Real Inner Product Space §inner-product ; bilinearity, homogeneity of the norm and the expansion of ∥ x ± y ∥ ν 2 \lVert x\pm y\rVert_{\nu}^{2} ∥ x ± y ∥ ν 2 are Elementary Identities in a Real Inner Product Space §bilinear , Elementary Identities in a Real Inner Product Space §homogeneity and Elementary Identities in a Real Inner Product Space §expansion ; and ∣ ⟨ x , y ⟩ ν ∣ ≤ ∥ x ∥ ν ∥ y ∥ ν |\langle x,y\rangle_{\nu}|\le\lVert x\rVert_{\nu}\lVert y\rVert_{\nu} ∣ ⟨ x , y ⟩ ν ∣ ≤ ∥ x ∥ ν ∥ y ∥ ν by The Cauchy-Schwarz Inequality in a Real Inner Product Space . For ν ∈ D Σ \nu\in\mathcal{D}_{\Sigma} ν ∈ D Σ the score Σ ( ν ) \Sigma(\nu) Σ ( ν ) lies in L 2 ( ν ; R d ) L^{2}(\nu;\mathbb{R}^{d}) L 2 ( ν ; R d ) by Penalty Pairs on the Wasserstein Space: the Penalty, Its Score, and Their Domains §pair , and E ( ν ) \mathcal{E}(\nu) E ( ν ) is a real number because D Σ ⊆ D \mathcal{D}_{\Sigma}\subseteq\mathcal{D} D Σ ⊆ D .
(0.2) The constant C C C is nonnegative. By Penalty Pairs on the Wasserstein Space: the Penalty, Its Score, and Their Domains §nonempty there is μ 0 ∈ D Σ ⊆ D \mu_{0}\in\mathcal{D}_{\Sigma}\subseteq\mathcal{D} μ 0 ∈ D Σ ⊆ D . Its second moment is a nonnegative real number, so (P4) gives 0 ≤ M 2 ( μ 0 ) ≤ C ( 1 + ∣ E ( μ 0 ) ∣ ) 0\le M_{2}(\mu_{0})\le C\,(1+|\mathcal{E}(\mu_{0})|) 0 ≤ M 2 ( μ 0 ) ≤ C ( 1 + ∣ E ( μ 0 ) ∣ ) . If C < 0 C<0 C < 0 , then, as 0 < 1 + ∣ E ( μ 0 ) ∣ 0<1+|\mathcal{E}(\mu_{0})| 0 < 1 + ∣ E ( μ 0 ) ∣ , we would get C ( 1 + ∣ E ( μ 0 ) ∣ ) < 0 C\,(1+|\mathcal{E}(\mu_{0})|)<0 C ( 1 + ∣ E ( μ 0 ) ∣ ) < 0 , a contradiction. Hence 0 ≤ C 0\le C 0 ≤ C .
(0.3) A bound for g g g . By (Running cost) and Bounded Real-Valued Function on a Set , fix a real M g ≥ 0 M_{g}\ge0 M g ≥ 0 with ∣ g ( μ ) ∣ ≤ M g |g(\mu)|\le M_{g} ∣ g ( μ ) ∣ ≤ M g for every μ ∈ P 2 ( R d ) \mu\in\mathcal{P}_{2}(\mathbb{R}^{d}) μ ∈ P 2 ( R d ) .
(0.4) The common-noise term. For Z ∈ S ( d ) Z\in\mathcal{S}(d) Z ∈ S ( d ) put Q ( Z ) = t r ( Γ ⊤ Γ Z ) Q(Z)=\mathrm{tr}(\Gamma^{\top}\Gamma Z) Q ( Z ) = tr ( Γ ⊤ Γ Z ) ; by The Trace as a Sum of Quadratic Forms, its Monotonicity and a Norm Bound §rows , applied as in (P4) with A = Γ A=\Gamma A = Γ , Q ( Z ) = ∑ k = 1 p w k ⋅ ( Z w k ) Q(Z)=\sum_{k=1}^{p}w_{k}\cdot(Zw_{k}) Q ( Z ) = ∑ k = 1 p w k ⋅ ( Z w k ) . Differences and scalar multiples of members of S ( d ) \mathcal{S}(d) S ( d ) lie in S ( d ) \mathcal{S}(d) S ( d ) by claim 1 of The Positive Semidefinite Ordering is a Partial Order Compatible with the Linear Structure . For Y , H ∈ S ( d ) Y,H\in\mathcal{S}(d) Y , H ∈ S ( d ) , real δ \delta δ and k ∈ [ p ] k\in[p] k ∈ [ p ] , claim 1 of Linearity of the Matrix-Vector Product and the Quadratic Form as a Double Sum gives ( Y ± δ H ) w k = Y w k ± δ ( H w k ) (Y\pm\delta H)w_{k}=Yw_{k}\pm\delta\,(Hw_{k}) ( Y ± δH ) w k = Y w k ± δ ( H w k ) , and claim 5 of Bilinearity and Symmetry of the Dot Product on R n \mathbb{R}^n R n then gives w k ⋅ ( ( Y ± δ H ) w k ) = w k ⋅ ( Y w k ) ± δ ( w k ⋅ ( H w k ) ) w_{k}\cdot((Y\pm\delta H)w_{k})=w_{k}\cdot(Yw_{k})\pm\delta\,\bigl(w_{k}\cdot(Hw_{k})\bigr) w k ⋅ (( Y ± δH ) w k ) = w k ⋅ ( Y w k ) ± δ ( w k ⋅ ( H w k ) ) ; summing over k k k with claims 2 and 3 of Properties of Finite Sums , Q ( Y ± δ H ) = Q ( Y ) ± δ Q ( H ) Q(Y\pm\delta H)=Q(Y)\pm\delta\,Q(H) Q ( Y ± δH ) = Q ( Y ) ± δ Q ( H ) . For μ ∈ D \mu\in\mathcal{D} μ ∈ D we write h ( μ ) = Q ( H E ( μ ) ) = t r ( Γ ⊤ Γ H E ( μ ) ) h(\mu)=Q(H_{\mathcal{E}}(\mu))=\mathrm{tr}(\Gamma^{\top}\Gamma H_{\mathcal{E}}(\mu)) h ( μ ) = Q ( H E ( μ )) = tr ( Γ ⊤ Γ H E ( μ )) ; by (P4), ∣ h ( μ ) ∣ ≤ C ( 1 + ∣ E ( μ ) ∣ ) |h(\mu)|\le C(1+|\mathcal{E}(\mu)|) ∣ h ( μ ) ∣ ≤ C ( 1 + ∣ E ( μ ) ∣ ) .
(0.5) Elementary inequalities in δ \delta δ . Let δ ∈ R \delta\in\mathbb{R} δ ∈ R with 0 < δ < 1 0<\delta<1 0 < δ < 1 . Then 0 < θ δ ≤ δ < 1 0<\theta\delta\le\delta<1 0 < θ δ ≤ δ < 1 , because 0 < θ ≤ 1 0<\theta\le1 0 < θ ≤ 1 . Consequently δ 2 + θ δ 2 2 > 0 \tfrac{\delta}{2}+\tfrac{\theta\delta^{2}}{2}>0 2 δ + 2 θ δ 2 > 0 and δ 2 − θ δ 2 2 = δ 2 ( 1 − θ δ ) ≥ 0 \tfrac{\delta}{2}-\tfrac{\theta\delta^{2}}{2}=\tfrac{\delta}{2}(1-\theta\delta)\ge0 2 δ − 2 θ δ 2 = 2 δ ( 1 − θ δ ) ≥ 0 . For real m , s m,s m , s one has 2 m s ≤ m 2 + s 2 2ms\le m^{2}+s^{2} 2 m s ≤ m 2 + s 2 , since 0 ≤ ( m − s ) 2 = m 2 − 2 m s + s 2 0\le(m-s)^{2}=m^{2}-2ms+s^{2} 0 ≤ ( m − s ) 2 = m 2 − 2 m s + s 2 .
(0.6) Expanded form of the shifts. Let ( ν , q ) ∈ V ( D Σ ) (\nu,q)\in\mathcal{V}(\mathcal{D}_{\Sigma}) ( ν , q ) ∈ V ( D Σ ) , r ∈ R r\in\mathbb{R} r ∈ R , Y ∈ S ( d ) Y\in\mathcal{S}(d) Y ∈ S ( d ) , δ > 0 \delta>0 δ > 0 , and write ζ = Σ ( ν ) \zeta=\Sigma(\nu) ζ = Σ ( ν ) . By (0e), The Bundle of Vector Fields over a Set of Measures, Second-Order Equation Operators on the Wasserstein Space, and Their Delta-Shifts §shifted , the formula of The Discounted Hamilton-Jacobi Equation with Common Noise and a Penalty Drift on the Wasserstein Space §operator for F 0 F_{0} F 0 (P2), linearity of Q Q Q (0.4), and ⟨ ζ , q ± δ ζ ⟩ ν = ⟨ ζ , q ⟩ ν ± δ ∥ ζ ∥ ν 2 \langle\zeta,q\pm\delta\zeta\rangle_{\nu}=\langle\zeta,q\rangle_{\nu}\pm\delta\lVert\zeta\rVert_{\nu}^{2} ⟨ ζ , q ± δ ζ ⟩ ν = ⟨ ζ , q ⟩ ν ± δ ∥ ζ ∥ ν 2 (0.1),
F δ − ( ν , r , q , Y ) = λ 0 r + λ 0 δ E ( ν ) − 1 2 Q ( Y ) − 1 2 δ h ( ν ) + θ 2 ∥ q + δ ζ ∥ ν 2 + ⟨ ζ , q ⟩ ν + δ ∥ ζ ∥ ν 2 − g ( ν ) − G ( ν ) , (0a) F^{-}_{\delta}(\nu,r,q,Y)=\lambda_{0}r+\lambda_{0}\delta\,\mathcal{E}(\nu)-\frac{1}{2}Q(Y)-\frac{1}{2}\delta\,h(\nu)+\frac{\theta}{2}\lVert q+\delta\zeta\rVert_{\nu}^{2}+\langle\zeta,q\rangle_{\nu}+\delta\lVert\zeta\rVert_{\nu}^{2}-g(\nu)-\mathcal{G}(\nu),\tag{0a} F δ − ( ν , r , q , Y ) = λ 0 r + λ 0 δ E ( ν ) − 2 1 Q ( Y ) − 2 1 δ h ( ν ) + 2 θ ∥ q + δ ζ ∥ ν 2 + ⟨ ζ , q ⟩ ν + δ ∥ ζ ∥ ν 2 − g ( ν ) − G ( ν ) , ( 0a )
F δ + ( ν , r , q , Y ) = λ 0 r − λ 0 δ E ( ν ) − 1 2 Q ( Y ) + 1 2 δ h ( ν ) + θ 2 ∥ q − δ ζ ∥ ν 2 + ⟨ ζ , q ⟩ ν − δ ∥ ζ ∥ ν 2 − g ( ν ) − G ( ν ) . (0b) F^{+}_{\delta}(\nu,r,q,Y)=\lambda_{0}r-\lambda_{0}\delta\,\mathcal{E}(\nu)-\frac{1}{2}Q(Y)+\frac{1}{2}\delta\,h(\nu)+\frac{\theta}{2}\lVert q-\delta\zeta\rVert_{\nu}^{2}+\langle\zeta,q\rangle_{\nu}-\delta\lVert\zeta\rVert_{\nu}^{2}-g(\nu)-\mathcal{G}(\nu).\tag{0b} F δ + ( ν , r , q , Y ) = λ 0 r − λ 0 δ E ( ν ) − 2 1 Q ( Y ) + 2 1 δ h ( ν ) + 2 θ ∥ q − δ ζ ∥ ν 2 + ⟨ ζ , q ⟩ ν − δ ∥ ζ ∥ ν 2 − g ( ν ) − G ( ν ) . ( 0b )
Step 1 (Degenerate ellipticity). Let ( ν , q ) ∈ V ( D Σ ) (\nu,q)\in\mathcal{V}(\mathcal{D}_{\Sigma}) ( ν , q ) ∈ V ( D Σ ) , r ∈ R r\in\mathbb{R} r ∈ R and X , Y ∈ S ( d ) X,Y\in\mathcal{S}(d) X , Y ∈ S ( d ) with X ⪯ Y X\preceq Y X ⪯ Y . By (P3) and Degenerate Elliptic Second-Order Equation Operators on the Wasserstein Space §elliptic , F 0 ( ν , r , q , Y ) ≤ F 0 ( ν , r , q , X ) F_{0}(\nu,r,q,Y)\le F_{0}(\nu,r,q,X) F 0 ( ν , r , q , Y ) ≤ F 0 ( ν , r , q , X ) ; subtracting G ( ν ) \mathcal{G}(\nu) G ( ν ) from both sides and using (P2) gives F ( ν , r , q , Y ) ≤ F ( ν , r , q , X ) F(\nu,r,q,Y)\le F(\nu,r,q,X) F ( ν , r , q , Y ) ≤ F ( ν , r , q , X ) . Hence F F F is degenerate elliptic , which is claim 1.
Step 2 (Local strict properness). Let R > 0 R>0 R > 0 . By (P4) and Locally Strictly Proper Second-Order Equation Operator on the Wasserstein Space §strictly-proper there is a properness constant λ > 0 \lambda>0 λ > 0 for F 0 F_{0} F 0 at R R R . For ( ν , q ) ∈ V ( D Σ ) (\nu,q)\in\mathcal{V}(\mathcal{D}_{\Sigma}) ( ν , q ) ∈ V ( D Σ ) , Y ∈ S ( d ) Y\in\mathcal{S}(d) Y ∈ S ( d ) and − R ≤ s ≤ r ≤ R -R\le s\le r\le R − R ≤ s ≤ r ≤ R , the terms G ( ν ) \mathcal{G}(\nu) G ( ν ) cancel by (P2), so F ( ν , r , q , Y ) − F ( ν , s , q , Y ) = F 0 ( ν , r , q , Y ) − F 0 ( ν , s , q , Y ) ≥ λ ( r − s ) F(\nu,r,q,Y)-F(\nu,s,q,Y)=F_{0}(\nu,r,q,Y)-F_{0}(\nu,s,q,Y)\ge\lambda(r-s) F ( ν , r , q , Y ) − F ( ν , s , q , Y ) = F 0 ( ν , r , q , Y ) − F 0 ( ν , s , q , Y ) ≥ λ ( r − s ) . Hence λ \lambda λ is a properness constant for F F F at R R R ; as R > 0 R>0 R > 0 was arbitrary, F F F is locally strictly proper .
Step 3 (Shift-coercivity). Let δ , R ∈ R \delta,R\in\mathbb{R} δ , R ∈ R with 0 < δ < 1 0<\delta<1 0 < δ < 1 and 0 < R 0<R 0 < R ; then 0 < R ≤ R + L 0<R\le R+L 0 < R ≤ R + L . By (P4) and The Shift-Coercivity Condition for an Equation Operator on the Wasserstein Space §coercivity there is a score bound C ′ ≥ 0 C'\ge0 C ′ ≥ 0 for F 0 F_{0} F 0 at ( δ , R + L ) (\delta,R+L) ( δ , R + L ) ; we show that C ′ C' C ′ is a score bound for F F F at ( δ , R ) (\delta,R) ( δ , R ) . Let ξ = ( ν , r , q , Y ) \xi=(\nu,r,q,Y) ξ = ( ν , r , q , Y ) and η = ( ν ′ , r ′ , q ′ , Y ′ ) \eta=(\nu',r',q',Y') η = ( ν ′ , r ′ , q ′ , Y ′ ) be R R R -bounded test data with F δ − ( ξ ) − F δ + ( η ) < R F^{-}_{\delta}(\xi)-F^{+}_{\delta}(\eta)<R F δ − ( ξ ) − F δ + ( η ) < R . Each of the five strict inequalities defining R R R -boundedness remains true when R R R is replaced by R + L R+L R + L , so ξ \xi ξ and η \eta η are ( R + L ) (R+L) ( R + L ) -bounded; and by (0e) and G ( ν ) ≤ L \mathcal{G}(\nu)\le L G ( ν ) ≤ L , 0 ≤ G ( ν ′ ) 0\le\mathcal{G}(\nu') 0 ≤ G ( ν ′ ) (P1),
F 0 , δ − ( ξ ) − F 0 , δ + ( η ) = F δ − ( ξ ) − F δ + ( η ) + G ( ν ) − G ( ν ′ ) < R + L . F^{-}_{0,\delta}(\xi)-F^{+}_{0,\delta}(\eta)=F^{-}_{\delta}(\xi)-F^{+}_{\delta}(\eta)+\mathcal{G}(\nu)-\mathcal{G}(\nu')<R+L . F 0 , δ − ( ξ ) − F 0 , δ + ( η ) = F δ − ( ξ ) − F δ + ( η ) + G ( ν ) − G ( ν ′ ) < R + L .
By Test Data for an Intrinsic Second-Order Equation Operator on the Wasserstein Space and the Admissible Sets §admissible , every member ξ \xi ξ of S δ , R − ( F ) S^{-}_{\delta,R}(F) S δ , R − ( F ) comes with such an η \eta η , hence belongs to S δ , R + L − ( F 0 ) S^{-}_{\delta,R+L}(F_{0}) S δ , R + L − ( F 0 ) , and every member η \eta η of S δ , R + ( F ) S^{+}_{\delta,R}(F) S δ , R + ( F ) comes with such a ξ \xi ξ , hence belongs to S δ , R + L + ( F 0 ) S^{+}_{\delta,R+L}(F_{0}) S δ , R + L + ( F 0 ) . So every test datum ( ν , r , q , Y ) (\nu,r,q,Y) ( ν , r , q , Y ) in S δ , R − ( F ) S^{-}_{\delta,R}(F) S δ , R − ( F ) or S δ , R + ( F ) S^{+}_{\delta,R}(F) S δ , R + ( F ) satisfies ∥ Σ ( ν ) ∥ ν ≤ C ′ \lVert\Sigma(\nu)\rVert_{\nu}\le C' ∥ Σ ( ν ) ∥ ν ≤ C ′ , i.e. C ′ C' C ′ is a score bound for F F F at ( δ , R ) (\delta,R) ( δ , R ) . As δ , R \delta,R δ , R were arbitrary, F F F satisfies the shift-coercivity condition .
Step 4 (The density cost at uniquely mapped pairs). Claim. Let μ , ν ∈ D Σ \mu,\nu\in\mathcal{D}_{\Sigma} μ , ν ∈ D Σ be such that both ordered pairs ( μ , ν ) (\mu,\nu) ( μ , ν ) and ( ν , μ ) (\nu,\mu) ( ν , μ ) are uniquely mapped , let S S S be an optimal map from μ \mu μ to ν \nu ν and S ′ S' S ′ one from ν \nu ν to μ \mu μ (such maps exist by Optimal Transport Maps and Uniquely Mapped Pairs of Probability Measures §uniquely-mapped ), with classes i d − S ∈ L 2 ( μ ; R d ) \mathrm{id}-S\in L^{2}(\mu;\mathbb{R}^{d}) id − S ∈ L 2 ( μ ; R d ) and i d − S ′ ∈ L 2 ( ν ; R d ) \mathrm{id}-S'\in L^{2}(\nu;\mathbb{R}^{d}) id − S ′ ∈ L 2 ( ν ; R d ) as in The Optimal Map as a Square-Integrable Vector Field: Integrability, Transport Cost and Uniqueness of the Class §square-integrable , and put M = ⟨ Σ ( μ ) , i d − S ⟩ μ + ⟨ Σ ( ν ) , i d − S ′ ⟩ ν M=\langle\Sigma(\mu),\mathrm{id}-S\rangle_{\mu}+\langle\Sigma(\nu),\mathrm{id}-S'\rangle_{\nu} M = ⟨ Σ ( μ ) , id − S ⟩ μ + ⟨ Σ ( ν ) , id − S ′ ⟩ ν . Then for every positive α ∈ R \alpha\in\mathbb{R} α ∈ R
G ( μ ) − G ( ν ) ≤ α M + 4 d L 2 σ 2 α and G ( ν ) − G ( μ ) ≤ α M + 4 d L 2 σ 2 α , (K) \mathcal{G}(\mu)-\mathcal{G}(\nu)\le\alpha M+\frac{4dL^{2}}{\sigma^{2}\alpha}\qquad\text{and}\qquad\mathcal{G}(\nu)-\mathcal{G}(\mu)\le\alpha M+\frac{4dL^{2}}{\sigma^{2}\alpha},\tag{K} G ( μ ) − G ( ν ) ≤ α M + σ 2 α 4 d L 2 and G ( ν ) − G ( μ ) ≤ α M + σ 2 α 4 d L 2 , ( K )
and
M ≤ ( ∥ Σ ( μ ) ∥ μ + ∥ Σ ( ν ) ∥ ν ) W 2 ( μ , ν ) . (K’) M\le\bigl(\lVert\Sigma(\mu)\rVert_{\mu}+\lVert\Sigma(\nu)\rVert_{\nu}\bigr)W_{2}(\mu,\nu).\tag{K'} M ≤ ( ∥ Σ ( μ ) ∥ μ + ∥ Σ ( ν ) ∥ ν ) W 2 ( μ , ν ) . ( K’ )
Proof. (4a) A second pair. As σ 2 2 ≥ 0 \tfrac{\sigma^{2}}{2}\ge0 2 σ 2 ≥ 0 , Existence and Uniqueness of the Nonnegative Square Root gives a real σ ′ ≥ 0 \sigma'\ge0 σ ′ ≥ 0 with σ ′ 2 = σ 2 2 \sigma'^{2}=\tfrac{\sigma^{2}}{2} σ ′ 2 = 2 σ 2 ; since σ 2 2 ≠ 0 \tfrac{\sigma^{2}}{2}\ne0 2 σ 2 = 0 we have σ ′ ≠ 0 \sigma'\ne0 σ ′ = 0 , so σ ′ > 0 \sigma'>0 σ ′ > 0 . Let ( D ′ , D Σ ′ , E ′ , Σ ′ ) (\mathcal{D}',\mathcal{D}'_{\Sigma},\mathcal{E}',\Sigma') ( D ′ , D Σ ′ , E ′ , Σ ′ ) be the Langevin free-energy pair with potential V V V and noise intensity σ ′ \sigma' σ ′ . In The Langevin Free-Energy Pair of a Confining Potential on the Wasserstein Space §pair the conditions defining D \mathcal{D} D and D Σ \mathcal{D}_{\Sigma} D Σ do not involve the noise intensity, so D ′ = D \mathcal{D}'=\mathcal{D} D ′ = D and D Σ ′ = D Σ \mathcal{D}'_{\Sigma}=\mathcal{D}_{\Sigma} D Σ ′ = D Σ ; and for ρ ∈ D Σ \rho\in\mathcal{D}_{\Sigma} ρ ∈ D Σ , Σ ′ ( ρ ) = ∇ V + σ ′ 2 2 ξ ρ = ∇ V + σ 2 4 ξ ρ \Sigma'(\rho)=\nabla V+\tfrac{\sigma'^{2}}{2}\xi_{\rho}=\nabla V+\tfrac{\sigma^{2}}{4}\xi_{\rho} Σ ′ ( ρ ) = ∇ V + 2 σ ′ 2 ξ ρ = ∇ V + 4 σ 2 ξ ρ . Since σ 2 2 = σ 2 4 + σ 2 4 \tfrac{\sigma^{2}}{2}=\tfrac{\sigma^{2}}{4}+\tfrac{\sigma^{2}}{4} 2 σ 2 = 4 σ 2 + 4 σ 2 , computing in the vector space T ρ T_{\rho} T ρ gives Σ ( ρ ) = Σ ′ ( ρ ) + σ 2 4 ξ ρ \Sigma(\rho)=\Sigma'(\rho)+\tfrac{\sigma^{2}}{4}\xi_{\rho} Σ ( ρ ) = Σ ′ ( ρ ) + 4 σ 2 ξ ρ . The primed pair is a penalty pair by The Langevin Free-Energy Pair is a Wasserstein-Coercive Penalty Pair: Growth Bounds, Continuity of the Translation Hessian, and the First Variation of the Penalty §pair and is displacement convex, i.e. 0 0 0 -displacement convex , by The Langevin Free-Energy Pair is Displacement Convex, with Closed Score Along Couplings and Regular Penalised Maxima §convex , both applied with potential V V V and noise intensity σ ′ \sigma' σ ′ .
(4b) Monotonicity. Since μ , ν ∈ D Σ ′ \mu,\nu\in\mathcal{D}'_{\Sigma} μ , ν ∈ D Σ ′ , A λ \lambda λ -Displacement Convex Penalty Pair Has a λ \lambda λ -Monotone Score Along Optimal Couplings §mapped , applied to the primed pair with λ = 0 \lambda=0 λ = 0 and to μ , ν , S , S ′ \mu,\nu,S,S' μ , ν , S , S ′ , gives 0 = 0 ⋅ W 2 ( μ , ν ) 2 ≤ ⟨ Σ ′ ( μ ) , i d − S ⟩ μ + ⟨ Σ ′ ( ν ) , i d − S ′ ⟩ ν 0=0\cdot W_{2}(\mu,\nu)^{2}\le\langle\Sigma'(\mu),\mathrm{id}-S\rangle_{\mu}+\langle\Sigma'(\nu),\mathrm{id}-S'\rangle_{\nu} 0 = 0 ⋅ W 2 ( μ , ν ) 2 ≤ ⟨ Σ ′ ( μ ) , id − S ⟩ μ + ⟨ Σ ′ ( ν ) , id − S ′ ⟩ ν . Put M e n t = ⟨ ξ μ , i d − S ⟩ μ + ⟨ ξ ν , i d − S ′ ⟩ ν M_{\mathrm{ent}}=\langle\xi_{\mu},\mathrm{id}-S\rangle_{\mu}+\langle\xi_{\nu},\mathrm{id}-S'\rangle_{\nu} M ent = ⟨ ξ μ , id − S ⟩ μ + ⟨ ξ ν , id − S ′ ⟩ ν . By (4a) and bilinearity (0.1), M = ⟨ Σ ′ ( μ ) , i d − S ⟩ μ + ⟨ Σ ′ ( ν ) , i d − S ′ ⟩ ν + σ 2 4 M e n t ≥ σ 2 4 M e n t M=\langle\Sigma'(\mu),\mathrm{id}-S\rangle_{\mu}+\langle\Sigma'(\nu),\mathrm{id}-S'\rangle_{\nu}+\tfrac{\sigma^{2}}{4}M_{\mathrm{ent}}\ge\tfrac{\sigma^{2}}{4}M_{\mathrm{ent}} M = ⟨ Σ ′ ( μ ) , id − S ⟩ μ + ⟨ Σ ′ ( ν ) , id − S ′ ⟩ ν + 4 σ 2 M ent ≥ 4 σ 2 M ent .
(4c) Proof of (K). By (P1), μ \mu μ and ν \nu ν are absolutely continuous members of P 2 I ( R d ) \mathcal{P}_{2}^{\mathcal{I}}(\mathbb{R}^{d}) P 2 I ( R d ) . Let α > 0 \alpha>0 α > 0 and put A = α σ 2 4 > 0 A=\tfrac{\alpha\sigma^{2}}{4}>0 A = 4 α σ 2 > 0 , so that d L 2 A = 4 d L 2 σ 2 α \tfrac{dL^{2}}{A}=\tfrac{4dL^{2}}{\sigma^{2}\alpha} A d L 2 = σ 2 α 4 d L 2 . The Density Cost Along Optimal Maps is Controlled by the Monotonicity of the Score §displacement , applied with L L L , Φ \Phi Φ , μ \mu μ , ν \nu ν , T = S T=S T = S , T ′ = S ′ T'=S' T ′ = S ′ and A A A , gives G ( μ ) − G ( ν ) ≤ α σ 2 4 M e n t + 4 d L 2 σ 2 α ≤ α M + 4 d L 2 σ 2 α \mathcal{G}(\mu)-\mathcal{G}(\nu)\le\alpha\,\tfrac{\sigma^{2}}{4}M_{\mathrm{ent}}+\tfrac{4dL^{2}}{\sigma^{2}\alpha}\le\alpha M+\tfrac{4dL^{2}}{\sigma^{2}\alpha} G ( μ ) − G ( ν ) ≤ α 4 σ 2 M ent + σ 2 α 4 d L 2 ≤ α M + σ 2 α 4 d L 2 by (4b), as α > 0 \alpha>0 α > 0 . Applied instead with ν , μ \nu,\mu ν , μ in place of μ , ν \mu,\nu μ , ν and with T = S ′ T=S' T = S ′ , T ′ = S T'=S T ′ = S , it gives G ( ν ) − G ( μ ) ≤ A ( ⟨ ξ ν , i d − S ′ ⟩ ν + ⟨ ξ μ , i d − S ⟩ μ ) + d L 2 A = α σ 2 4 M e n t + 4 d L 2 σ 2 α \mathcal{G}(\nu)-\mathcal{G}(\mu)\le A\bigl(\langle\xi_{\nu},\mathrm{id}-S'\rangle_{\nu}+\langle\xi_{\mu},\mathrm{id}-S\rangle_{\mu}\bigr)+\tfrac{dL^{2}}{A}=\alpha\,\tfrac{\sigma^{2}}{4}M_{\mathrm{ent}}+\tfrac{4dL^{2}}{\sigma^{2}\alpha} G ( ν ) − G ( μ ) ≤ A ( ⟨ ξ ν , id − S ′ ⟩ ν + ⟨ ξ μ , id − S ⟩ μ ) + A d L 2 = α 4 σ 2 M ent + σ 2 α 4 d L 2 , and (4b) again gives the second inequality of (K).
(4d) Proof of (K'). By The Optimal Map as a Square-Integrable Vector Field: Integrability, Transport Cost and Uniqueness of the Class §cost , ∥ i d − S ∥ μ 2 = W 2 ( μ , ν ) 2 \lVert\mathrm{id}-S\rVert_{\mu}^{2}=W_{2}(\mu,\nu)^{2} ∥ id − S ∥ μ 2 = W 2 ( μ , ν ) 2 and ∥ i d − S ′ ∥ ν 2 = W 2 ( ν , μ ) 2 = W 2 ( μ , ν ) 2 \lVert\mathrm{id}-S'\rVert_{\nu}^{2}=W_{2}(\nu,\mu)^{2}=W_{2}(\mu,\nu)^{2} ∥ id − S ′ ∥ ν 2 = W 2 ( ν , μ ) 2 = W 2 ( μ , ν ) 2 , by symmetry of W 2 W_{2} W 2 (The Quadratic Wasserstein Distance is a Metric on the Wasserstein Space §metric ); all these numbers being nonnegative, ∥ i d − S ∥ μ = ∥ i d − S ′ ∥ ν = W 2 ( μ , ν ) \lVert\mathrm{id}-S\rVert_{\mu}=\lVert\mathrm{id}-S'\rVert_{\nu}=W_{2}(\mu,\nu) ∥ id − S ∥ μ = ∥ id − S ′ ∥ ν = W 2 ( μ , ν ) . By the Cauchy-Schwarz inequality (0.1), M ≤ ∣ ⟨ Σ ( μ ) , i d − S ⟩ μ ∣ + ∣ ⟨ Σ ( ν ) , i d − S ′ ⟩ ν ∣ ≤ ( ∥ Σ ( μ ) ∥ μ + ∥ Σ ( ν ) ∥ ν ) W 2 ( μ , ν ) M\le|\langle\Sigma(\mu),\mathrm{id}-S\rangle_{\mu}|+|\langle\Sigma(\nu),\mathrm{id}-S'\rangle_{\nu}|\le\bigl(\lVert\Sigma(\mu)\rVert_{\mu}+\lVert\Sigma(\nu)\rVert_{\nu}\bigr)W_{2}(\mu,\nu) M ≤ ∣ ⟨ Σ ( μ ) , id − S ⟩ μ ∣ + ∣ ⟨ Σ ( ν ) , id − S ′ ⟩ ν ∣ ≤ ( ∥ Σ ( μ ) ∥ μ + ∥ Σ ( ν ) ∥ ν ) W 2 ( μ , ν ) .
Step 5 (Shift-semicontinuity). Let δ , R ∈ R \delta,R\in\mathbb{R} δ , R ∈ R with 0 < δ < 1 0<\delta<1 0 < δ < 1 and 0 < R 0<R 0 < R , let ξ n = ( ν n , r n , q n , Y n ) \xi_{n}=(\nu_{n},r_{n},q_{n},Y_{n}) ξ n = ( ν n , r n , q n , Y n ) (n ∈ N n\in\mathbb{N} n ∈ N ) and ξ = ( ν , r , q , Y ) \xi=(\nu,r,q,Y) ξ = ( ν , r , q , Y ) be test data, and let ( π n ) (\pi_{n}) ( π n ) be couplings such that ( ξ n ) (\xi_{n}) ( ξ n ) converges to ξ \xi ξ along ( π n ) (\pi_{n}) ( π n ) with score bounded by R R R . This notion involves only the pair, the set of test data and R R R -boundedness, which by Test Data for an Intrinsic Second-Order Equation Operator on the Wasserstein Space and the Admissible Sets §data and Test Data for an Intrinsic Second-Order Equation Operator on the Wasserstein Space and the Admissible Sets §bounded are the same for F F F and for F 0 F_{0} F 0 , both being operators over D Σ \mathcal{D}_{\Sigma} D Σ . By that clause, ∥ Σ ( ν n ) ∥ ν n ≤ R \lVert\Sigma(\nu_{n})\rVert_{\nu_{n}}\le R ∥ Σ ( ν n ) ∥ ν n ≤ R for every n n n , and ( π n ) (\pi_{n}) ( π n ) is a sequence of couplings of vanishing cost , so I ( π n ) I(\pi_{n}) I ( π n ) converges to 0 0 0 .
(5.1) W 2 ( ν n , ν ) → 0 W_{2}(\nu_{n},\nu)\to0 W 2 ( ν n , ν ) → 0 . By The Quadratic Wasserstein Distance on Euclidean Space §distance , W 2 ( ν n , ν ) 2 ≤ I ( π n ) W_{2}(\nu_{n},\nu)^{2}\le I(\pi_{n}) W 2 ( ν n , ν ) 2 ≤ I ( π n ) . Given ε > 0 \varepsilon>0 ε > 0 , choose N N N with I ( π n ) < ε 2 I(\pi_{n})<\varepsilon^{2} I ( π n ) < ε 2 for n ≥ N n\ge N n ≥ N ; then W 2 ( ν n , ν ) 2 < ε 2 W_{2}(\nu_{n},\nu)^{2}<\varepsilon^{2} W 2 ( ν n , ν ) 2 < ε 2 , so W 2 ( ν n , ν ) < ε W_{2}(\nu_{n},\nu)<\varepsilon W 2 ( ν n , ν ) < ε . The distance is symmetric by The Quadratic Wasserstein Distance is a Metric on the Wasserstein Space §metric , so also W 2 ( ν , ν n ) < ε W_{2}(\nu,\nu_{n})<\varepsilon W 2 ( ν , ν n ) < ε for n ≥ N n\ge N n ≥ N .
(5.2) G ( ν n ) → G ( ν ) \mathcal{G}(\nu_{n})\to\mathcal{G}(\nu) G ( ν n ) → G ( ν ) . By (P1), ν \nu ν and every ν n \nu_{n} ν n lie in D Σ ⊆ D \mathcal{D}_{\Sigma}\subseteq\mathcal{D} D Σ ⊆ D and D \mathcal{D} D has the map property, so by The Map Property of a Set of Probability Measures §map-property both ordered pairs ( ν n , ν ) (\nu_{n},\nu) ( ν n , ν ) and ( ν , ν n ) (\nu,\nu_{n}) ( ν , ν n ) are uniquely mapped. Put b = R + ∥ Σ ( ν ) ∥ ν > 0 b=R+\lVert\Sigma(\nu)\rVert_{\nu}>0 b = R + ∥ Σ ( ν ) ∥ ν > 0 . For every n n n and every α > 0 \alpha>0 α > 0 , Step 4 applied to ν n , ν \nu_{n},\nu ν n , ν , with optimal maps S n S_{n} S n from ν n \nu_{n} ν n to ν \nu ν and S n ′ S'_{n} S n ′ from ν \nu ν to ν n \nu_{n} ν n (which exist by Optimal Transport Maps and Uniquely Mapped Pairs of Probability Measures §uniquely-mapped ) and with M n M_{n} M n the corresponding number M M M , gives, by (K), (K') and ∥ Σ ( ν n ) ∥ ν n ≤ R \lVert\Sigma(\nu_{n})\rVert_{\nu_{n}}\le R ∥ Σ ( ν n ) ∥ ν n ≤ R ,
∣ G ( ν n ) − G ( ν ) ∣ ≤ α M n + 4 d L 2 σ 2 α ≤ α b W 2 ( ν n , ν ) + 4 d L 2 σ 2 α . |\mathcal{G}(\nu_{n})-\mathcal{G}(\nu)|\le\alpha M_{n}+\frac{4dL^{2}}{\sigma^{2}\alpha}\le\alpha\,b\,W_{2}(\nu_{n},\nu)+\frac{4dL^{2}}{\sigma^{2}\alpha}. ∣ G ( ν n ) − G ( ν ) ∣ ≤ α M n + σ 2 α 4 d L 2 ≤ α b W 2 ( ν n , ν ) + σ 2 α 4 d L 2 .
Let ε > 0 \varepsilon>0 ε > 0 . First choose α = 1 + 8 d L 2 σ 2 ε \alpha=1+\tfrac{8dL^{2}}{\sigma^{2}\varepsilon} α = 1 + σ 2 ε 8 d L 2 ; then α > 0 \alpha>0 α > 0 and α ε 2 > 8 d L 2 σ 2 ε ⋅ ε 2 = 4 d L 2 σ 2 \alpha\tfrac{\varepsilon}{2}>\tfrac{8dL^{2}}{\sigma^{2}\varepsilon}\cdot\tfrac{\varepsilon}{2}=\tfrac{4dL^{2}}{\sigma^{2}} α 2 ε > σ 2 ε 8 d L 2 ⋅ 2 ε = σ 2 4 d L 2 , as 0 ≤ d L 2 0\le dL^{2} 0 ≤ d L 2 , so 4 d L 2 σ 2 α < ε 2 \tfrac{4dL^{2}}{\sigma^{2}\alpha}<\tfrac{\varepsilon}{2} σ 2 α 4 d L 2 < 2 ε . Then, by (5.1), choose N N N with W 2 ( ν n , ν ) < ε 2 α b W_{2}(\nu_{n},\nu)<\tfrac{\varepsilon}{2\alpha b} W 2 ( ν n , ν ) < 2 α b ε for n ≥ N n\ge N n ≥ N . For n ≥ N n\ge N n ≥ N the display gives ∣ G ( ν n ) − G ( ν ) ∣ < ε 2 + ε 2 = ε |\mathcal{G}(\nu_{n})-\mathcal{G}(\nu)|<\tfrac{\varepsilon}{2}+\tfrac{\varepsilon}{2}=\varepsilon ∣ G ( ν n ) − G ( ν ) ∣ < 2 ε + 2 ε = ε .
(5.3) Transfer from F 0 F_{0} F 0 . By (P4) and The Shift-Semicontinuity Condition for an Equation Operator on the Wasserstein Space §semicontinuity , F 0 F_{0} F 0 is shift-semicontinuous at ( δ , R ) (\delta,R) ( δ , R ) . First, let c ∈ R c\in\mathbb{R} c ∈ R be such that for every ε > 0 \varepsilon>0 ε > 0 there is N N N with F δ − ( ξ n ) ≤ c + ε F^{-}_{\delta}(\xi_{n})\le c+\varepsilon F δ − ( ξ n ) ≤ c + ε for n ≥ N n\ge N n ≥ N , and put c ′ = c + G ( ν ) c'=c+\mathcal{G}(\nu) c ′ = c + G ( ν ) . Given ε > 0 \varepsilon>0 ε > 0 , choose N 1 N_{1} N 1 with F δ − ( ξ n ) ≤ c + ε 2 F^{-}_{\delta}(\xi_{n})\le c+\tfrac{\varepsilon}{2} F δ − ( ξ n ) ≤ c + 2 ε for n ≥ N 1 n\ge N_{1} n ≥ N 1 and, by (5.2), N 2 N_{2} N 2 with ∣ G ( ν n ) − G ( ν ) ∣ < ε 2 |\mathcal{G}(\nu_{n})-\mathcal{G}(\nu)|<\tfrac{\varepsilon}{2} ∣ G ( ν n ) − G ( ν ) ∣ < 2 ε for n ≥ N 2 n\ge N_{2} n ≥ N 2 ; for n ≥ max { N 1 , N 2 } n\ge\max\{N_{1},N_{2}\} n ≥ max { N 1 , N 2 } , (0e) gives F 0 , δ − ( ξ n ) = F δ − ( ξ n ) + G ( ν n ) < c ′ + ε F^{-}_{0,\delta}(\xi_{n})=F^{-}_{\delta}(\xi_{n})+\mathcal{G}(\nu_{n})<c'+\varepsilon F 0 , δ − ( ξ n ) = F δ − ( ξ n ) + G ( ν n ) < c ′ + ε . The first implication of The Shift-Semicontinuity Condition for an Equation Operator on the Wasserstein Space §level for F 0 F_{0} F 0 , with c ′ c' c ′ , gives F 0 , δ − ( ξ ) ≤ c ′ F^{-}_{0,\delta}(\xi)\le c' F 0 , δ − ( ξ ) ≤ c ′ , that is, by (0e), F δ − ( ξ ) ≤ c F^{-}_{\delta}(\xi)\le c F δ − ( ξ ) ≤ c . Secondly, let c ∈ R c\in\mathbb{R} c ∈ R be such that for every ε > 0 \varepsilon>0 ε > 0 there is N N N with c − ε ≤ F δ + ( ξ n ) c-\varepsilon\le F^{+}_{\delta}(\xi_{n}) c − ε ≤ F δ + ( ξ n ) for n ≥ N n\ge N n ≥ N , and put c ′ = c + G ( ν ) c'=c+\mathcal{G}(\nu) c ′ = c + G ( ν ) . Given ε > 0 \varepsilon>0 ε > 0 , choosing N 1 , N 2 N_{1},N_{2} N 1 , N 2 in the same way, for n ≥ max { N 1 , N 2 } n\ge\max\{N_{1},N_{2}\} n ≥ max { N 1 , N 2 } we get F 0 , δ + ( ξ n ) = F δ + ( ξ n ) + G ( ν n ) > c ′ − ε F^{+}_{0,\delta}(\xi_{n})=F^{+}_{\delta}(\xi_{n})+\mathcal{G}(\nu_{n})>c'-\varepsilon F 0 , δ + ( ξ n ) = F δ + ( ξ n ) + G ( ν n ) > c ′ − ε ; the second implication for F 0 F_{0} F 0 gives c ′ ≤ F 0 , δ + ( ξ ) c'\le F^{+}_{0,\delta}(\xi) c ′ ≤ F 0 , δ + ( ξ ) , that is, c ≤ F δ + ( ξ ) c\le F^{+}_{\delta}(\xi) c ≤ F δ + ( ξ ) .
By (5.3), F F F is shift-semicontinuous at ( δ , R ) (\delta,R) ( δ , R ) ; as δ , R \delta,R δ , R were arbitrary, F F F satisfies the shift-semicontinuity condition .
Step 6 (Second-order structure at uniquely mapped pairs). Let T = { t ∈ R : 0 ≤ t } T=\{t\in\mathbb{R}:0\le t\} T = { t ∈ R : 0 ≤ t } . The pair ( ω 1 ′ , ω 2 ) (\omega'_{1},\omega_{2}) ( ω 1 ′ , ω 2 ) below is chosen first; it depends only on g , d , L , σ , λ 0 , θ , C g,d,L,\sigma,\lambda_{0},\theta,C g , d , L , σ , λ 0 , θ , C , and we show it is a second-order structure pair for F F F at R R R for every positive R R R .
(6.1) The modulus ω 1 ′ \omega'_{1} ω 1 ′ . For s ∈ T s\in T s ∈ T let U ( s ) = { ∣ g ( μ ′ ) − g ( ν ′ ) ∣ : μ ′ , ν ′ ∈ P 2 ( R d ) , W 2 ( μ ′ , ν ′ ) 2 ≤ s } U(s)=\{|g(\mu')-g(\nu')|:\mu',\nu'\in\mathcal{P}_{2}(\mathbb{R}^{d}),\ W_{2}(\mu',\nu')^{2}\le s\} U ( s ) = { ∣ g ( μ ′ ) − g ( ν ′ ) ∣ : μ ′ , ν ′ ∈ P 2 ( R d ) , W 2 ( μ ′ , ν ′ ) 2 ≤ s } . It contains 0 = ∣ g ( μ 0 ) − g ( μ 0 ) ∣ 0=|g(\mu_{0})-g(\mu_{0})| 0 = ∣ g ( μ 0 ) − g ( μ 0 ) ∣ , with μ 0 \mu_{0} μ 0 from (0.2), since W 2 ( μ 0 , μ 0 ) = 0 W_{2}(\mu_{0},\mu_{0})=0 W 2 ( μ 0 , μ 0 ) = 0 by The Quadratic Wasserstein Distance is a Metric on the Wasserstein Space §metric ; and it is bounded above by 2 M g 2M_{g} 2 M g by (0.3). Hence ω 1 ( s ) = sup U ( s ) \omega_{1}(s)=\sup U(s) ω 1 ( s ) = sup U ( s ) is defined by The Real Numbers: Standing Notation and Background §bounds , and 0 ≤ ω 1 ( s ) 0\le\omega_{1}(s) 0 ≤ ω 1 ( s ) because ω 1 ( s ) \omega_{1}(s) ω 1 ( s ) is an upper bound of U ( s ) ∋ 0 U(s)\ni0 U ( s ) ∋ 0 (Upper Bound and Least Upper Bound ). Given ε > 0 \varepsilon>0 ε > 0 , uniform continuity of g g g (Uniformly Continuous Map Between Metric Spaces ) gives γ > 0 \gamma>0 γ > 0 with ∣ g ( μ ′ ) − g ( ν ′ ) ∣ < ε |g(\mu')-g(\nu')|<\varepsilon ∣ g ( μ ′ ) − g ( ν ′ ) ∣ < ε whenever W 2 ( μ ′ , ν ′ ) < γ W_{2}(\mu',\nu')<\gamma W 2 ( μ ′ , ν ′ ) < γ . Put γ 1 = γ 2 / 4 > 0 \gamma_{1}=\gamma^{2}/4>0 γ 1 = γ 2 /4 > 0 . If t ∈ T t\in T t ∈ T and t ≤ γ 1 t\le\gamma_{1} t ≤ γ 1 , every element of U ( t ) U(t) U ( t ) comes from μ ′ , ν ′ \mu',\nu' μ ′ , ν ′ with W 2 ( μ ′ , ν ′ ) 2 ≤ ( γ / 2 ) 2 W_{2}(\mu',\nu')^{2}\le(\gamma/2)^{2} W 2 ( μ ′ , ν ′ ) 2 ≤ ( γ /2 ) 2 , so W 2 ( μ ′ , ν ′ ) ≤ γ / 2 < γ W_{2}(\mu',\nu')\le\gamma/2<\gamma W 2 ( μ ′ , ν ′ ) ≤ γ /2 < γ and the element is < ε <\varepsilon < ε ; thus ε \varepsilon ε is an upper bound of U ( t ) U(t) U ( t ) and ω 1 ( t ) ≤ ε \omega_{1}(t)\le\varepsilon ω 1 ( t ) ≤ ε , the supremum being the least upper bound. So ω 1 \omega_{1} ω 1 is a modulus of continuity , and by construction ∣ g ( μ ′ ) − g ( ν ′ ) ∣ ≤ ω 1 ( s ) |g(\mu')-g(\nu')|\le\omega_{1}(s) ∣ g ( μ ′ ) − g ( ν ′ ) ∣ ≤ ω 1 ( s ) whenever W 2 ( μ ′ , ν ′ ) 2 ≤ s W_{2}(\mu',\nu')^{2}\le s W 2 ( μ ′ , ν ′ ) 2 ≤ s . Put ω 1 ′ ( s ) = ω 1 ( s ) + 4 d L 2 σ 2 s \omega'_{1}(s)=\omega_{1}(s)+\tfrac{4dL^{2}}{\sigma^{2}}\,s ω 1 ′ ( s ) = ω 1 ( s ) + σ 2 4 d L 2 s for s ∈ T s\in T s ∈ T . The coefficient 4 d L 2 σ 2 \tfrac{4dL^{2}}{\sigma^{2}} σ 2 4 d L 2 is nonnegative, so s ↦ 4 d L 2 σ 2 s s\mapsto\tfrac{4dL^{2}}{\sigma^{2}}s s ↦ σ 2 4 d L 2 s is a modulus of continuity by Linear Moduli of Continuity §modulus , and ω 1 ′ \omega'_{1} ω 1 ′ is a modulus of continuity by Sums, Nonnegative Multiples, Monotonicity and Quadratic Reparametrisation of Moduli of Continuity §sum .
(6.2) The function ω 2 \omega_{2} ω 2 . For t ∈ T t\in T t ∈ T and real α > 1 \alpha>1 α > 1 put ω 2 ( t , α ) = ( λ 0 + C + 4 C θ 2 α 2 ) t \omega_{2}(t,\alpha)=(\lambda_{0}+C+4C\theta^{2}\alpha^{2})\,t ω 2 ( t , α ) = ( λ 0 + C + 4 C θ 2 α 2 ) t . For each α > 1 \alpha>1 α > 1 the coefficient is nonnegative by (0.2), so t ↦ ω 2 ( t , α ) t\mapsto\omega_{2}(t,\alpha) t ↦ ω 2 ( t , α ) is a modulus of continuity by Linear Moduli of Continuity §modulus .
(6.3) The inequality. Let R > 0 R>0 R > 0 , and let α , δ , μ , ν , S , S ′ , r , X , Y \alpha,\delta,\mu,\nu,S,S',r,\mathbb{X},\mathbb{Y} α , δ , μ , ν , S , S ′ , r , X , Y be as in The Second-Order Structure Condition at Uniquely Mapped Pairs on the Wasserstein Space §pair : 1 < α 1<\alpha 1 < α , 0 < δ < 1 0<\delta<1 0 < δ < 1 , μ , ν ∈ D Σ \mu,\nu\in\mathcal{D}_{\Sigma} μ , ν ∈ D Σ with both ordered pairs ( μ , ν ) (\mu,\nu) ( μ , ν ) and ( ν , μ ) (\nu,\mu) ( ν , μ ) uniquely mapped , S S S an optimal map from μ \mu μ to ν \nu ν and S ′ S' S ′ one from ν \nu ν to μ \mu μ , r ∈ [ − R , R ] r\in[-R,R] r ∈ [ − R , R ] , and ( X , Y ) (\mathbb{X},\mathbb{Y}) ( X , Y ) admitted at α \alpha α . (The condition δ ( ∣ E ( μ ) ∣ + ∣ E ( ν ) ∣ ) ≤ R \delta(|\mathcal{E}(\mu)|+|\mathcal{E}(\nu)|)\le R δ ( ∣ E ( μ ) ∣ + ∣ E ( ν ) ∣ ) ≤ R will not be needed.) Write W = W 2 ( μ , ν ) W=W_{2}(\mu,\nu) W = W 2 ( μ , ν ) , ζ = Σ ( μ ) ∈ L 2 ( μ ; R d ) \zeta=\Sigma(\mu)\in L^{2}(\mu;\mathbb{R}^{d}) ζ = Σ ( μ ) ∈ L 2 ( μ ; R d ) , τ = Σ ( ν ) ∈ L 2 ( ν ; R d ) \tau=\Sigma(\nu)\in L^{2}(\nu;\mathbb{R}^{d}) τ = Σ ( ν ) ∈ L 2 ( ν ; R d ) , a = α ( i d − S ) ∈ L 2 ( μ ; R d ) a=\alpha(\mathrm{id}-S)\in L^{2}(\mu;\mathbb{R}^{d}) a = α ( id − S ) ∈ L 2 ( μ ; R d ) , b = α ( S ′ − i d ) = − α ( i d − S ′ ) ∈ L 2 ( ν ; R d ) b=\alpha(S'-\mathrm{id})=-\alpha(\mathrm{id}-S')\in L^{2}(\nu;\mathbb{R}^{d}) b = α ( S ′ − id ) = − α ( id − S ′ ) ∈ L 2 ( ν ; R d ) , e = ∣ E ( μ ) ∣ + ∣ E ( ν ) ∣ e=|\mathcal{E}(\mu)|+|\mathcal{E}(\nu)| e = ∣ E ( μ ) ∣ + ∣ E ( ν ) ∣ and t = δ ( e + 1 ) t=\delta(e+1) t = δ ( e + 1 ) , and let Δ \Delta Δ be the difference F δ − ( μ , r , a , X ) − F δ + ( ν , r , b , Y ) F^{-}_{\delta}(\mu,r,a,\mathbb{X})-F^{+}_{\delta}(\nu,r,b,\mathbb{Y}) F δ − ( μ , r , a , X ) − F δ + ( ν , r , b , Y ) to be bounded below.
Norms of the displacements. The data μ , ν , S , S ′ \mu,\nu,S,S' μ , ν , S , S ′ satisfy the hypotheses of Step 4, so ∥ i d − S ∥ μ = ∥ i d − S ′ ∥ ν = W \lVert\mathrm{id}-S\rVert_{\mu}=\lVert\mathrm{id}-S'\rVert_{\nu}=W ∥ id − S ∥ μ = ∥ id − S ′ ∥ ν = W by (4d), and, by homogeneity (0.1) with ∣ α ∣ = α |\alpha|=\alpha ∣ α ∣ = α , ∥ a ∥ μ = ∥ b ∥ ν = α W \lVert a\rVert_{\mu}=\lVert b\rVert_{\nu}=\alpha W ∥ a ∥ μ = ∥ b ∥ ν = α W .
A bound on W 2 W^{2} W 2 . By Optimal Transport Maps and Uniquely Mapped Pairs of Probability Measures §map , S # μ = ν S_{\#}\mu=\nu S # μ = ν , so ∥ S ∥ μ 2 = ∫ R d ∥ S ∥ 2 d μ = M 2 ( ν ) \lVert S\rVert_{\mu}^{2}=\int_{\mathbb{R}^{d}}\lVert S\rVert^{2}\,d\mu=M_{2}(\nu) ∥ S ∥ μ 2 = ∫ R d ∥ S ∥ 2 d μ = M 2 ( ν ) by The Optimal Map as a Square-Integrable Vector Field: Integrability, Transport Cost and Uniqueness of the Class §square-integrable and (0.1); and ∥ i d ∥ μ 2 = M 2 ( μ ) \lVert\mathrm{id}\rVert_{\mu}^{2}=M_{2}(\mu) ∥ id ∥ μ 2 = M 2 ( μ ) by Basic Properties of the Tangent Space: Closed Subspace, the Identity Map Belongs to It, Second-Moment Limits, and Representation of Bounded Functionals on Gradients §identity . The parallelogram law Elementary Identities in a Real Inner Product Space §parallelogram gives W 2 = ∥ i d − S ∥ μ 2 ≤ ∥ i d − S ∥ μ 2 + ∥ i d + S ∥ μ 2 = 2 M 2 ( μ ) + 2 M 2 ( ν ) W^{2}=\lVert\mathrm{id}-S\rVert_{\mu}^{2}\le\lVert\mathrm{id}-S\rVert_{\mu}^{2}+\lVert\mathrm{id}+S\rVert_{\mu}^{2}=2M_{2}(\mu)+2M_{2}(\nu) W 2 = ∥ id − S ∥ μ 2 ≤ ∥ id − S ∥ μ 2 + ∥ id + S ∥ μ 2 = 2 M 2 ( μ ) + 2 M 2 ( ν ) , and (P4) yields W 2 ≤ 2 C ( 2 + e ) W^{2}\le2C(2+e) W 2 ≤ 2 C ( 2 + e ) . Since 2 + e ≤ 2 ( 1 + e ) 2+e\le2(1+e) 2 + e ≤ 2 ( 1 + e ) , we get δ W 2 ≤ 2 C δ ( 2 + e ) ≤ 4 C t \delta W^{2}\le2C\delta(2+e)\le4Ct δ W 2 ≤ 2 C δ ( 2 + e ) ≤ 4 Ct .
Expansion of Δ \Delta Δ . Subtracting (0b) at ( ν , r , b , Y ) (\nu,r,b,\mathbb{Y}) ( ν , r , b , Y ) from (0a) at ( μ , r , a , X ) (\mu,r,a,\mathbb{X}) ( μ , r , a , X ) , the terms λ 0 r \lambda_{0}r λ 0 r cancel and
Δ = λ 0 δ ( E ( μ ) + E ( ν ) ) + 1 2 ( Q ( Y ) − Q ( X ) ) − 1 2 δ ( h ( μ ) + h ( ν ) ) + θ 2 ( ∥ a + δ ζ ∥ μ 2 − ∥ b − δ τ ∥ ν 2 ) + ( ⟨ ζ , a ⟩ μ − ⟨ τ , b ⟩ ν + G ( ν ) − G ( μ ) ) + δ ∥ ζ ∥ μ 2 + δ ∥ τ ∥ ν 2 + g ( ν ) − g ( μ ) . \begin{aligned}
\Delta={}&\lambda_{0}\delta\bigl(\mathcal{E}(\mu)+\mathcal{E}(\nu)\bigr)+\frac{1}{2}\bigl(Q(\mathbb{Y})-Q(\mathbb{X})\bigr)-\frac{1}{2}\delta\bigl(h(\mu)+h(\nu)\bigr)+\frac{\theta}{2}\Bigl(\lVert a+\delta\zeta\rVert_{\mu}^{2}-\lVert b-\delta\tau\rVert_{\nu}^{2}\Bigr)\\
&+\bigl(\langle\zeta,a\rangle_{\mu}-\langle\tau,b\rangle_{\nu}+\mathcal{G}(\nu)-\mathcal{G}(\mu)\bigr)+\delta\lVert\zeta\rVert_{\mu}^{2}+\delta\lVert\tau\rVert_{\nu}^{2}+g(\nu)-g(\mu).
\end{aligned} Δ = λ 0 δ ( E ( μ ) + E ( ν ) ) + 2 1 ( Q ( Y ) − Q ( X ) ) − 2 1 δ ( h ( μ ) + h ( ν ) ) + 2 θ ( ∥ a + δ ζ ∥ μ 2 − ∥ b − δ τ ∥ ν 2 ) + ( ⟨ ζ , a ⟩ μ − ⟨ τ , b ⟩ ν + G ( ν ) − G ( μ ) ) + δ ∥ ζ ∥ μ 2 + δ ∥ τ ∥ ν 2 + g ( ν ) − g ( μ ) .
By (0.1), ∥ a + δ ζ ∥ μ 2 = ∥ a ∥ μ 2 + 2 δ ⟨ a , ζ ⟩ μ + δ 2 ∥ ζ ∥ μ 2 \lVert a+\delta\zeta\rVert_{\mu}^{2}=\lVert a\rVert_{\mu}^{2}+2\delta\langle a,\zeta\rangle_{\mu}+\delta^{2}\lVert\zeta\rVert_{\mu}^{2} ∥ a + δ ζ ∥ μ 2 = ∥ a ∥ μ 2 + 2 δ ⟨ a , ζ ⟩ μ + δ 2 ∥ ζ ∥ μ 2 and ∥ b − δ τ ∥ ν 2 = ∥ b ∥ ν 2 − 2 δ ⟨ b , τ ⟩ ν + δ 2 ∥ τ ∥ ν 2 \lVert b-\delta\tau\rVert_{\nu}^{2}=\lVert b\rVert_{\nu}^{2}-2\delta\langle b,\tau\rangle_{\nu}+\delta^{2}\lVert\tau\rVert_{\nu}^{2} ∥ b − δ τ ∥ ν 2 = ∥ b ∥ ν 2 − 2 δ ⟨ b , τ ⟩ ν + δ 2 ∥ τ ∥ ν 2 ; as ∥ a ∥ μ 2 = ∥ b ∥ ν 2 = α 2 W 2 \lVert a\rVert_{\mu}^{2}=\lVert b\rVert_{\nu}^{2}=\alpha^{2}W^{2} ∥ a ∥ μ 2 = ∥ b ∥ ν 2 = α 2 W 2 ,
θ 2 ( ∥ a + δ ζ ∥ μ 2 − ∥ b − δ τ ∥ ν 2 ) = θ δ ⟨ a , ζ ⟩ μ + θ δ ⟨ b , τ ⟩ ν + θ δ 2 2 ∥ ζ ∥ μ 2 − θ δ 2 2 ∥ τ ∥ ν 2 . \frac{\theta}{2}\Bigl(\lVert a+\delta\zeta\rVert_{\mu}^{2}-\lVert b-\delta\tau\rVert_{\nu}^{2}\Bigr)=\theta\delta\langle a,\zeta\rangle_{\mu}+\theta\delta\langle b,\tau\rangle_{\nu}+\frac{\theta\delta^{2}}{2}\lVert\zeta\rVert_{\mu}^{2}-\frac{\theta\delta^{2}}{2}\lVert\tau\rVert_{\nu}^{2}. 2 θ ( ∥ a + δ ζ ∥ μ 2 − ∥ b − δ τ ∥ ν 2 ) = θ δ ⟨ a , ζ ⟩ μ + θ δ ⟨ b , τ ⟩ ν + 2 θ δ 2 ∥ ζ ∥ μ 2 − 2 θ δ 2 ∥ τ ∥ ν 2 .
Bounds for the individual terms. (i) Score and density-cost terms: by bilinearity (0.1), ⟨ ζ , a ⟩ μ − ⟨ τ , b ⟩ ν = α ( ⟨ ζ , i d − S ⟩ μ + ⟨ τ , i d − S ′ ⟩ ν ) = α M \langle\zeta,a\rangle_{\mu}-\langle\tau,b\rangle_{\nu}=\alpha\bigl(\langle\zeta,\mathrm{id}-S\rangle_{\mu}+\langle\tau,\mathrm{id}-S'\rangle_{\nu}\bigr)=\alpha M ⟨ ζ , a ⟩ μ − ⟨ τ , b ⟩ ν = α ( ⟨ ζ , id − S ⟩ μ + ⟨ τ , id − S ′ ⟩ ν ) = α M , with M M M the number of Step 4 for μ , ν , S , S ′ \mu,\nu,S,S' μ , ν , S , S ′ ; as α > 0 \alpha>0 α > 0 , the first inequality of (K) gives ⟨ ζ , a ⟩ μ − ⟨ τ , b ⟩ ν + G ( ν ) − G ( μ ) ≥ − 4 d L 2 σ 2 α − 1 \langle\zeta,a\rangle_{\mu}-\langle\tau,b\rangle_{\nu}+\mathcal{G}(\nu)-\mathcal{G}(\mu)\ge-\tfrac{4dL^{2}}{\sigma^{2}}\alpha^{-1} ⟨ ζ , a ⟩ μ − ⟨ τ , b ⟩ ν + G ( ν ) − G ( μ ) ≥ − σ 2 4 d L 2 α − 1 . (ii) Traces: X ⪯ Y \mathbb{X}\preceq\mathbb{Y} X ⪯ Y by The Second-Order Structure Condition at Optimally Coupled Pairs on the Lift of the Wasserstein Space §admitted , the ordering being that of Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §matrices , so Q ( X ) ≤ Q ( Y ) Q(\mathbb{X})\le Q(\mathbb{Y}) Q ( X ) ≤ Q ( Y ) by The Trace as a Sum of Quadratic Forms, its Monotonicity and a Norm Bound §monotone , applied as in (P4) with A = Γ A=\Gamma A = Γ , and 1 2 ( Q ( Y ) − Q ( X ) ) ≥ 0 \tfrac{1}{2}(Q(\mathbb{Y})-Q(\mathbb{X}))\ge0 2 1 ( Q ( Y ) − Q ( X )) ≥ 0 . (iii) Cross terms: by Cauchy-Schwarz (0.1) and (0.5) with m = ∥ ζ ∥ μ m=\lVert\zeta\rVert_{\mu} m = ∥ ζ ∥ μ , s = θ α W s=\theta\alpha W s = θ α W ,
θ δ ⟨ a , ζ ⟩ μ ≥ − θ δ α W ∥ ζ ∥ μ = − δ 2 2 m s ≥ − δ 2 ∥ ζ ∥ μ 2 − δ 2 θ 2 α 2 W 2 , \theta\delta\langle a,\zeta\rangle_{\mu}\ge-\theta\delta\,\alpha W\lVert\zeta\rVert_{\mu}=-\frac{\delta}{2}\,2ms\ge-\frac{\delta}{2}\lVert\zeta\rVert_{\mu}^{2}-\frac{\delta}{2}\theta^{2}\alpha^{2}W^{2}, θ δ ⟨ a , ζ ⟩ μ ≥ − θ δ α W ∥ ζ ∥ μ = − 2 δ 2 m s ≥ − 2 δ ∥ ζ ∥ μ 2 − 2 δ θ 2 α 2 W 2 ,
and in the same way θ δ ⟨ b , τ ⟩ ν ≥ − δ 2 ∥ τ ∥ ν 2 − δ 2 θ 2 α 2 W 2 \theta\delta\langle b,\tau\rangle_{\nu}\ge-\tfrac{\delta}{2}\lVert\tau\rVert_{\nu}^{2}-\tfrac{\delta}{2}\theta^{2}\alpha^{2}W^{2} θ δ ⟨ b , τ ⟩ ν ≥ − 2 δ ∥ τ ∥ ν 2 − 2 δ θ 2 α 2 W 2 . (iv) Collecting the score terms: the coefficient of ∥ ζ ∥ μ 2 \lVert\zeta\rVert_{\mu}^{2} ∥ ζ ∥ μ 2 becomes δ + θ δ 2 2 − δ 2 = δ 2 + θ δ 2 2 \delta+\tfrac{\theta\delta^{2}}{2}-\tfrac{\delta}{2}=\tfrac{\delta}{2}+\tfrac{\theta\delta^{2}}{2} δ + 2 θ δ 2 − 2 δ = 2 δ + 2 θ δ 2 and that of ∥ τ ∥ ν 2 \lVert\tau\rVert_{\nu}^{2} ∥ τ ∥ ν 2 becomes δ − θ δ 2 2 − δ 2 = δ 2 ( 1 − θ δ ) \delta-\tfrac{\theta\delta^{2}}{2}-\tfrac{\delta}{2}=\tfrac{\delta}{2}(1-\theta\delta) δ − 2 θ δ 2 − 2 δ = 2 δ ( 1 − θ δ ) , both nonnegative by (0.5), so these terms are ≥ 0 \ge0 ≥ 0 ; the remaining contribution is − θ 2 α 2 δ W 2 ≥ − 4 C θ 2 α 2 t -\theta^{2}\alpha^{2}\delta W^{2}\ge-4C\theta^{2}\alpha^{2}t − θ 2 α 2 δ W 2 ≥ − 4 C θ 2 α 2 t . (v) Penalty terms: λ 0 δ ( E ( μ ) + E ( ν ) ) ≥ − λ 0 δ e ≥ − λ 0 t \lambda_{0}\delta(\mathcal{E}(\mu)+\mathcal{E}(\nu))\ge-\lambda_{0}\delta e\ge-\lambda_{0}t λ 0 δ ( E ( μ ) + E ( ν )) ≥ − λ 0 δe ≥ − λ 0 t . (vi) Hessian terms: by (0.4), − 1 2 δ ( h ( μ ) + h ( ν ) ) ≥ − 1 2 δ C ( 2 + e ) ≥ − 1 2 δ ⋅ 2 C ( 1 + e ) = − C t -\tfrac{1}{2}\delta(h(\mu)+h(\nu))\ge-\tfrac{1}{2}\delta\,C(2+e)\ge-\tfrac{1}{2}\delta\cdot2C(1+e)=-Ct − 2 1 δ ( h ( μ ) + h ( ν )) ≥ − 2 1 δ C ( 2 + e ) ≥ − 2 1 δ ⋅ 2 C ( 1 + e ) = − Ct , using 0 ≤ C 0\le C 0 ≤ C (0.2) and 2 + e ≤ 2 ( 1 + e ) 2+e\le2(1+e) 2 + e ≤ 2 ( 1 + e ) . (vii) Running cost: W 2 ≤ α W 2 ≤ α W 2 + α − 1 W^{2}\le\alpha W^{2}\le\alpha W^{2}+\alpha^{-1} W 2 ≤ α W 2 ≤ α W 2 + α − 1 , as 1 < α 1<\alpha 1 < α , 0 ≤ W 2 0\le W^{2} 0 ≤ W 2 and 0 < α − 1 0<\alpha^{-1} 0 < α − 1 (claim 7 of Elementary Order Arithmetic in an Ordered Field ); so (6.1) with s = α W 2 + α − 1 s=\alpha W^{2}+\alpha^{-1} s = α W 2 + α − 1 gives g ( ν ) − g ( μ ) ≥ − ∣ g ( μ ) − g ( ν ) ∣ ≥ − ω 1 ( α W 2 + α − 1 ) g(\nu)-g(\mu)\ge-|g(\mu)-g(\nu)|\ge-\omega_{1}(\alpha W^{2}+\alpha^{-1}) g ( ν ) − g ( μ ) ≥ − ∣ g ( μ ) − g ( ν ) ∣ ≥ − ω 1 ( α W 2 + α − 1 ) . (viii) Combining (i) and (vii): α − 1 ≤ α W 2 + α − 1 \alpha^{-1}\le\alpha W^{2}+\alpha^{-1} α − 1 ≤ α W 2 + α − 1 and 0 ≤ 4 d L 2 σ 2 0\le\tfrac{4dL^{2}}{\sigma^{2}} 0 ≤ σ 2 4 d L 2 , so the sum of the lower bounds in (i) and (vii) is at least − ω 1 ( α W 2 + α − 1 ) − 4 d L 2 σ 2 ( α W 2 + α − 1 ) = − ω 1 ′ ( α W 2 + α − 1 ) -\omega_{1}(\alpha W^{2}+\alpha^{-1})-\tfrac{4dL^{2}}{\sigma^{2}}(\alpha W^{2}+\alpha^{-1})=-\omega'_{1}(\alpha W^{2}+\alpha^{-1}) − ω 1 ( α W 2 + α − 1 ) − σ 2 4 d L 2 ( α W 2 + α − 1 ) = − ω 1 ′ ( α W 2 + α − 1 ) .
Adding (ii)-(vi) and (viii) to the expansion of Δ \Delta Δ ,
− ω 1 ′ ( α W 2 ( μ , ν ) 2 + α − 1 ) − ω 2 ( δ ( ∣ E ( μ ) ∣ + ∣ E ( ν ) ∣ + 1 ) , α ) ≤ F δ − ( μ , r , α ( i d − S ) , X ) − F δ + ( ν , r , α ( S ′ − i d ) , Y ) . -\omega'_{1}\bigl(\alpha W_{2}(\mu,\nu)^{2}+\alpha^{-1}\bigr)-\omega_{2}\bigl(\delta(|\mathcal{E}(\mu)|+|\mathcal{E}(\nu)|+1),\alpha\bigr)\le F^{-}_{\delta}\bigl(\mu,r,\alpha(\mathrm{id}-S),\mathbb{X}\bigr)-F^{+}_{\delta}\bigl(\nu,r,\alpha(S'-\mathrm{id}),\mathbb{Y}\bigr). − ω 1 ′ ( α W 2 ( μ , ν ) 2 + α − 1 ) − ω 2 ( δ ( ∣ E ( μ ) ∣ + ∣ E ( ν ) ∣ + 1 ) , α ) ≤ F δ − ( μ , r , α ( id − S ) , X ) − F δ + ( ν , r , α ( S ′ − id ) , Y ) .
Hence ( ω 1 ′ , ω 2 ) (\omega'_{1},\omega_{2}) ( ω 1 ′ , ω 2 ) is a second-order structure pair for F F F at every R > 0 R>0 R > 0 , and F F F satisfies the second-order structure condition at uniquely mapped pairs .
Step 1 proves claim 1, and Steps 2, 3, 5 and 6 together prove claim 2.