Each result cited below is universally quantified over the data in its own statement, and is applied to the data named at the point of use. The symmetry and the triangle inequality of W 2 W_{2} W 2 (The Quadratic Wasserstein Distance is a Metric on the Wasserstein Space §symmetry , The Quadratic Wasserstein Distance is a Metric on the Wasserstein Space §triangle ) and the rules of Elementary Order Arithmetic in an Ordered Field and Elementary Arithmetic in an Ordered Field for adding and scaling inequalities are used without further mention. For ν , ρ ∈ P 2 ( R d ) \nu,\rho\in\mathcal{P}_{2}(\mathbb{R}^{d}) ν , ρ ∈ P 2 ( R d ) and π ∈ Π ( ν , ρ ) \pi\in\Pi(\nu,\rho) π ∈ Π ( ν , ρ ) we have W 2 ( ν , ρ ) ≤ I ( π ) W_{2}(\nu,\rho)\le\sqrt{I(\pi)} W 2 ( ν , ρ ) ≤ I ( π ) , by The Quadratic Wasserstein Distance on Euclidean Space §distance and claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field . Each v ∈ S v\in\mathcal{S} v ∈ S , being a viscosity subsolution, is bounded above near each point of P 2 ( R d ) \mathcal{P}_{2}(\mathbb{R}^{d}) P 2 ( R d ) .
Claim 1. Let μ ∈ P 2 ( R d ) \mu\in\mathcal{P}_{2}(\mathbb{R}^{d}) μ ∈ P 2 ( R d ) and let c , r c,r c , r be as in the assumption that S \mathcal{S} S is locally uniformly bounded above, for μ \mu μ . For ν \nu ν with W 2 ( ν , μ ) ≤ r W_{2}(\nu,\mu)\le r W 2 ( ν , μ ) ≤ r , c c c is an upper bound of { v ( ν ) : v ∈ S } \{v(\nu):v\in\mathcal{S}\} { v ( ν ) : v ∈ S } , so u ( ν ) ≤ c u(\nu)\le c u ( ν ) ≤ c by Upper Bound and Least Upper Bound ; hence c ∈ A u ( μ ) c\in A_{u}(\mu) c ∈ A u ( μ ) and u u u is bounded above near each point by Upper and Lower Semicontinuous Envelopes of a Real-Valued Function §near-bounds . For v ∈ S v\in\mathcal{S} v ∈ S , v ( μ ) v(\mu) v ( μ ) lies in the set of which u ( μ ) u(\mu) u ( μ ) is an upper bound, so v ( μ ) ≤ u ( μ ) v(\mu)\le u(\mu) v ( μ ) ≤ u ( μ ) . Consequently, for positive δ \delta δ the functions v − δ E v-\delta\mathcal{E} v − δ E and u − δ E u-\delta\mathcal{E} u − δ E on D \mathcal{D} D , both bounded above near each point of D \mathcal{D} D by The Delta-Envelopes of a Function on the Wasserstein Space Relative to a Penalty Pair §minus , satisfy v − δ E ≤ u − δ E v-\delta\mathcal{E}\le u-\delta\mathcal{E} v − δ E ≤ u − δ E pointwise, and Properties of the Upper Semicontinuous Envelope §monotone gives v δ − ( ν ) ≤ u δ − ( ν ) v^{-}_{\delta}(\nu)\le u^{-}_{\delta}(\nu) v δ − ( ν ) ≤ u δ − ( ν ) for ν ∈ D \nu\in\mathcal{D} ν ∈ D .
Claim 2. By claim 1, u u u is bounded above near each point. Let δ ∈ R \delta\in\mathbb{R} δ ∈ R be positive, let φ \varphi φ be an intrinsic test function on D \mathcal{D} D , let μ ^ ∈ D \hat{\mu}\in\mathcal{D} μ ^ ∈ D be a point at which the function with value u δ − ( μ ) − φ ( μ ) u^{-}_{\delta}(\mu)-\varphi(\mu) u δ − ( μ ) − φ ( μ ) at μ ∈ D \mu\in\mathcal{D} μ ∈ D has a local maximum relative to D \mathcal{D} D , witnessed by a positive radius τ \tau τ as in Local Maximum of a Function Relative to a Subset of a Metric Space , and let ε ∈ R \varepsilon\in\mathbb{R} ε ∈ R be positive. Write M = u δ − ( μ ^ ) − φ ( μ ^ ) M=u^{-}_{\delta}(\hat{\mu})-\varphi(\hat{\mu}) M = u δ − ( μ ^ ) − φ ( μ ^ ) .
Step 1: a strict maximum. Put β = ε 8 \beta=\tfrac{\varepsilon}{8} β = 8 ε and let ψ 0 ( ν ) = W 2 ( ν , μ ^ ) 2 \psi_{0}(\nu)=W_{2}(\nu,\hat{\mu})^{2} ψ 0 ( ν ) = W 2 ( ν , μ ^ ) 2 . Since D \mathcal{D} D has the map property, ψ 0 \psi_{0} ψ 0 is an intrinsic test function on D \mathcal{D} D with H ψ 0 ( ν ) = 2 I d H_{\psi_{0}}(\nu)=2I_{d} H ψ 0 ( ν ) = 2 I d for every ν \nu ν and ∇ ψ 0 ( μ ^ ) = 2 ( i d − S ) \nabla\psi_{0}(\hat{\mu})=2(\mathrm{id}-S) ∇ ψ 0 ( μ ^ ) = 2 ( id − S ) for any optimal map S S S from μ ^ \hat{\mu} μ ^ to μ ^ \hat{\mu} μ ^ , by The Squared Wasserstein Distance to a Fixed Measure and Functions of the Mean are Intrinsic Test Functions §distance . The identity map is such an S S S : i d # μ ^ = μ ^ \mathrm{id}_{\#}\hat{\mu}=\hat{\mu} id # μ ^ = μ ^ , and by The Displacement Pairing of a Square-Integrable Vector Field Along a Coupling §displacement with S = i d S=\mathrm{id} S = id the coupling ( i d , i d ) # μ ^ (\mathrm{id},\mathrm{id})_{\#}\hat{\mu} ( id , id ) # μ ^ has cost ∥ i d − i d ∥ μ ^ 2 = 0 = W 2 ( μ ^ , μ ^ ) 2 \lVert\mathrm{id}-\mathrm{id}\rVert_{\hat{\mu}}^{2}=0=W_{2}(\hat{\mu},\hat{\mu})^{2} ∥ id − id ∥ μ ^ 2 = 0 = W 2 ( μ ^ , μ ^ ) 2 , so it is optimal by Optimal Coupling of Two Probability Measures with Finite Second Moment §optimal and i d \mathrm{id} id is an optimal map by Optimal Transport Maps and Uniquely Mapped Pairs of Probability Measures §map . Hence ∇ ψ 0 ( μ ^ ) = 0 \nabla\psi_{0}(\hat{\mu})=0 ∇ ψ 0 ( μ ^ ) = 0 . Let φ ~ = φ + β ψ 0 \tilde{\varphi}=\varphi+\beta\psi_{0} φ ~ = φ + β ψ 0 . By Restrictions, Sums, Real Multiples and Differences of Intrinsic Test Functions on the Wasserstein Space §linear , φ ~ \tilde{\varphi} φ ~ is an intrinsic test function on D \mathcal{D} D with
∇ φ ~ ( μ ^ ) = ∇ φ ( μ ^ ) , H φ ~ ( μ ^ ) = H φ ( μ ^ ) + 2 β I d , φ ~ ( μ ^ ) = φ ( μ ^ ) , \nabla\tilde{\varphi}(\hat{\mu})=\nabla\varphi(\hat{\mu}),\qquad H_{\tilde{\varphi}}(\hat{\mu})=H_{\varphi}(\hat{\mu})+2\beta I_{d},\qquad\tilde{\varphi}(\hat{\mu})=\varphi(\hat{\mu}), ∇ φ ~ ( μ ^ ) = ∇ φ ( μ ^ ) , H φ ~ ( μ ^ ) = H φ ( μ ^ ) + 2 β I d , φ ~ ( μ ^ ) = φ ( μ ^ ) ,
the last because W 2 ( μ ^ , μ ^ ) = 0 W_{2}(\hat{\mu},\hat{\mu})=0 W 2 ( μ ^ , μ ^ ) = 0 . Moreover ∥ I d ∥ ≤ 1 \lVert I_{d}\rVert\le1 ∥ I d ∥ ≤ 1 : for ξ ∈ R d \xi\in\mathbb{R}^{d} ξ ∈ R d with ∥ ξ ∥ ≤ 1 \lVert\xi\rVert\le1 ∥ ξ ∥ ≤ 1 we have ∣ ξ ⋅ ( I d ξ ) ∣ = ∥ ξ ∥ 2 ≤ 1 |\xi\cdot(I_{d}\xi)|=\lVert\xi\rVert^{2}\le1 ∣ ξ ⋅ ( I d ξ ) ∣ = ∥ ξ ∥ 2 ≤ 1 by claim 1 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n and claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field , so 1 1 1 is an upper bound of the set whose least upper bound is ∥ I d ∥ \lVert I_{d}\rVert ∥ I d ∥ by Real Matrices, Symmetric Matrices and the Semidefinite Ordering: Standing Notation §norm . By claim 5 of Properties of the Norm of a Symmetric Real Matrix , ∥ H φ ~ ( μ ^ ) − H φ ( μ ^ ) ∥ = ∥ 2 β I d ∥ ≤ 2 β = ε 4 \lVert H_{\tilde{\varphi}}(\hat{\mu})-H_{\varphi}(\hat{\mu})\rVert=\lVert2\beta I_{d}\rVert\le2\beta=\tfrac{\varepsilon}{4} ∥ H φ ~ ( μ ^ ) − H φ ( μ ^ )∥ = ∥ 2 β I d ∥ ≤ 2 β = 4 ε . For ν ∈ D \nu\in\mathcal{D} ν ∈ D with W 2 ( ν , μ ^ ) < τ W_{2}(\nu,\hat{\mu})<\tau W 2 ( ν , μ ^ ) < τ the local maximum gives
u δ − ( ν ) − φ ~ ( ν ) ≤ M − β W 2 ( ν , μ ^ ) 2 . (1) u^{-}_{\delta}(\nu)-\tilde{\varphi}(\nu)\le M-\beta\,W_{2}(\nu,\hat{\mu})^{2}. \tag{1} u δ − ( ν ) − φ ~ ( ν ) ≤ M − β W 2 ( ν , μ ^ ) 2 . ( 1 )
Step 2: radii. Using Continuous Map Between Metric Spaces for φ \varphi φ and φ ~ \tilde{\varphi} φ ~ (property (a), Intrinsic Test Functions on the Wasserstein Space and Their Translation Hessians §continuity ), Upper Semicontinuous Function on a Subset of a Metric Space for u δ − u^{-}_{\delta} u δ − , which is upper semicontinuous on D \mathcal{D} D by Basic Properties of the Delta-Envelopes on the Wasserstein Space §semicontinuity , and property (e) of φ ~ \tilde{\varphi} φ ~ (Intrinsic Test Functions on the Wasserstein Space and Their Translation Hessians §hessian-continuity ), choose positive radii θ 0 , θ 2 , θ 3 , θ 4 \theta_{0},\theta_{2},\theta_{3},\theta_{4} θ 0 , θ 2 , θ 3 , θ 4 such that for ν ∈ P 2 ( R d ) \nu\in\mathcal{P}_{2}(\mathbb{R}^{d}) ν ∈ P 2 ( R d ) : W 2 ( ν , μ ^ ) < θ 0 W_{2}(\nu,\hat{\mu})<\theta_{0} W 2 ( ν , μ ^ ) < θ 0 implies ∣ φ ( ν ) − φ ( μ ^ ) ∣ < 1 |\varphi(\nu)-\varphi(\hat{\mu})|<1 ∣ φ ( ν ) − φ ( μ ^ ) ∣ < 1 ; W 2 ( ν , μ ^ ) < θ 2 W_{2}(\nu,\hat{\mu})<\theta_{2} W 2 ( ν , μ ^ ) < θ 2 implies ∥ H φ ~ ( ν ) − H φ ~ ( μ ^ ) ∥ < ε 4 \lVert H_{\tilde{\varphi}}(\nu)-H_{\tilde{\varphi}}(\hat{\mu})\rVert<\tfrac{\varepsilon}{4} ∥ H φ ~ ( ν ) − H φ ~ ( μ ^ )∥ < 4 ε ; ν ∈ D \nu\in\mathcal{D} ν ∈ D and W 2 ( ν , μ ^ ) < θ 3 W_{2}(\nu,\hat{\mu})<\theta_{3} W 2 ( ν , μ ^ ) < θ 3 imply u δ − ( ν ) < u δ − ( μ ^ ) + ε 4 u^{-}_{\delta}(\nu)<u^{-}_{\delta}(\hat{\mu})+\tfrac{\varepsilon}{4} u δ − ( ν ) < u δ − ( μ ^ ) + 4 ε ; W 2 ( ν , μ ^ ) < θ 4 W_{2}(\nu,\hat{\mu})<\theta_{4} W 2 ( ν , μ ^ ) < θ 4 implies ∣ φ ~ ( ν ) − φ ~ ( μ ^ ) ∣ < ε 8 |\tilde{\varphi}(\nu)-\tilde{\varphi}(\hat{\mu})|<\tfrac{\varepsilon}{8} ∣ φ ~ ( ν ) − φ ~ ( μ ^ ) ∣ < 8 ε .
We also need a radius for the gradients: there is a positive θ 1 \theta_{1} θ 1 such that for every ν ∈ D \nu\in\mathcal{D} ν ∈ D and every π ∈ Π ( ν , μ ^ ) \pi\in\Pi(\nu,\hat{\mu}) π ∈ Π ( ν , μ ^ ) with I ( π ) < θ 1 2 I(\pi)<\theta_{1}^{2} I ( π ) < θ 1 2 the discrepancy of ∇ φ ~ ( ν ) \nabla\tilde{\varphi}(\nu) ∇ φ ~ ( ν ) and ∇ φ ~ ( μ ^ ) \nabla\tilde{\varphi}(\hat{\mu}) ∇ φ ~ ( μ ^ ) along π \pi π is less than ( ε 4 ) 2 (\tfrac{\varepsilon}{4})^{2} ( 4 ε ) 2 . Suppose not. Let ( h n ) n ∈ N (h_{n})_{n\in\mathbb{N}} ( h n ) n ∈ N be a sequence of positive reals with limit 0 0 0 (Existence of a Sequence of Positive Real Numbers with Limit Zero ); for each n n n there are ν n ∈ D \nu_{n}\in\mathcal{D} ν n ∈ D and π n ∈ Π ( ν n , μ ^ ) \pi_{n}\in\Pi(\nu_{n},\hat{\mu}) π n ∈ Π ( ν n , μ ^ ) with I ( π n ) < h n 2 I(\pi_{n})<h_{n}^{2} I ( π n ) < h n 2 and discrepancy D n ≥ ( ε 4 ) 2 D_{n}\ge(\tfrac{\varepsilon}{4})^{2} D n ≥ ( 4 ε ) 2 . Since 0 ≤ I ( π n ) < h n 2 0\le I(\pi_{n})<h_{n}^{2} 0 ≤ I ( π n ) < h n 2 and ( h n 2 ) (h_{n}^{2}) ( h n 2 ) has limit 0 0 0 by claim 2 of Arithmetic of Limits of Real Sequences , ( I ( π n ) ) (I(\pi_{n})) ( I ( π n )) has limit 0 0 0 by claim 2 of Order Properties of Limits of Real Sequences ; property (c) of φ ~ \tilde{\varphi} φ ~ on D \mathcal{D} D (Intrinsic Test Functions on the Wasserstein Space and Their Translation Hessians §gradient-continuity ), at the point μ ^ ∈ D \hat{\mu}\in\mathcal{D} μ ^ ∈ D , then says that ( D n ) (D_{n}) ( D n ) has limit 0 0 0 , and claim 1 of Order Properties of Limits of Real Sequences gives ( ε 4 ) 2 ≤ 0 (\tfrac{\varepsilon}{4})^{2}\le0 ( 4 ε ) 2 ≤ 0 , contradicting claim 5 of Elementary Order Arithmetic in an Ordered Field .
Let c , r c,r c , r be as in the assumption that S \mathcal{S} S is locally uniformly bounded above, for μ ^ \hat{\mu} μ ^ . Put
ρ = 1 2 min { τ , r , θ 0 } , σ = 1 2 min { ρ , θ 1 , θ 2 , θ 3 , θ 4 , ε } , η = min { β σ 2 3 , ε 24 } , \rho=\tfrac12\min\{\tau,r,\theta_{0}\},\qquad\sigma=\tfrac12\min\{\rho,\theta_{1},\theta_{2},\theta_{3},\theta_{4},\varepsilon\},\qquad\eta=\min\Bigl\{\tfrac{\beta\sigma^{2}}{3},\tfrac{\varepsilon}{24}\Bigr\}, ρ = 2 1 min { τ , r , θ 0 } , σ = 2 1 min { ρ , θ 1 , θ 2 , θ 3 , θ 4 , ε } , η = min { 3 β σ 2 , 24 ε } ,
all positive by claim 2 of Elementary Properties of the Minimum of Two Elements (applied repeatedly) and claim 8 of Elementary Order Arithmetic in an Ordered Field , and let K = { ν ∈ D : W 2 ( ν , μ ^ ) ≤ ρ } K=\{\nu\in\mathcal{D}:W_{2}(\nu,\hat{\mu})\le\rho\} K = { ν ∈ D : W 2 ( ν , μ ^ ) ≤ ρ } . Since ρ < τ \rho<\tau ρ < τ , ρ < r \rho<r ρ < r and ρ < θ 0 \rho<\theta_{0} ρ < θ 0 , every ν ∈ K \nu\in K ν ∈ K satisfies (1), v ( ν ) ≤ c v(\nu)\le c v ( ν ) ≤ c for every v ∈ S v\in\mathcal{S} v ∈ S , and ∣ φ ( ν ) − φ ( μ ^ ) ∣ < 1 |\varphi(\nu)-\varphi(\hat{\mu})|<1 ∣ φ ( ν ) − φ ( μ ^ ) ∣ < 1 .
Step 3: a member of the family and its maximiser. Let θ η \theta_{\eta} θ η be a positive radius with ∣ φ ~ ( ν ) − φ ~ ( μ ^ ) ∣ < η |\tilde{\varphi}(\nu)-\tilde{\varphi}(\hat{\mu})|<\eta ∣ φ ~ ( ν ) − φ ~ ( μ ^ ) ∣ < η whenever W 2 ( ν , μ ^ ) < θ η W_{2}(\nu,\hat{\mu})<\theta_{\eta} W 2 ( ν , μ ^ ) < θ η , and put η ′ = min { η , 1 2 θ η , ρ } \eta'=\min\{\eta,\tfrac12\theta_{\eta},\rho\} η ′ = min { η , 2 1 θ η , ρ } . By Properties of the Upper Semicontinuous Envelope §approximation , applied to u − δ E u-\delta\mathcal{E} u − δ E on D \mathcal{D} D at μ ^ \hat{\mu} μ ^ with η ′ \eta' η ′ , there is z ∈ D z\in\mathcal{D} z ∈ D with W 2 ( z , μ ^ ) ≤ η ′ W_{2}(z,\hat{\mu})\le\eta' W 2 ( z , μ ^ ) ≤ η ′ and ∣ u ( z ) − δ E ( z ) − u δ − ( μ ^ ) ∣ < η ′ ≤ η |u(z)-\delta\,\mathcal{E}(z)-u^{-}_{\delta}(\hat{\mu})|<\eta'\le\eta ∣ u ( z ) − δ E ( z ) − u δ − ( μ ^ ) ∣ < η ′ ≤ η . Then z ∈ K z\in K z ∈ K and ∣ φ ~ ( z ) − φ ~ ( μ ^ ) ∣ < η |\tilde{\varphi}(z)-\tilde{\varphi}(\hat{\mu})|<\eta ∣ φ ~ ( z ) − φ ~ ( μ ^ ) ∣ < η . By claim 3 of Approximation Property of the Supremum and the Infimum in R \mathbb{R} R there is v ∈ S v\in\mathcal{S} v ∈ S with u ( z ) − η < v ( z ) u(z)-\eta<v(z) u ( z ) − η < v ( z ) .
Let g : D → R g:\mathcal{D}\to\mathbb{R} g : D → R be g ( ν ) = v δ − ( ν ) − φ ~ ( ν ) g(\nu)=v^{-}_{\delta}(\nu)-\tilde{\varphi}(\nu) g ( ν ) = v δ − ( ν ) − φ ~ ( ν ) . It is upper semicontinuous on D \mathcal{D} D : at ν 0 ∈ D \nu_{0}\in\mathcal{D} ν 0 ∈ D and for positive ε 0 \varepsilon_{0} ε 0 , Basic Properties of the Delta-Envelopes on the Wasserstein Space §semicontinuity and Upper Semicontinuous Function on a Subset of a Metric Space give a radius within which v δ − ( ν ) < v δ − ( ν 0 ) + ε 0 2 v^{-}_{\delta}(\nu)<v^{-}_{\delta}(\nu_{0})+\tfrac{\varepsilon_{0}}{2} v δ − ( ν ) < v δ − ( ν 0 ) + 2 ε 0 , continuity of φ ~ \tilde{\varphi} φ ~ a radius within which − φ ~ ( ν ) < − φ ~ ( ν 0 ) + ε 0 2 -\tilde{\varphi}(\nu)<-\tilde{\varphi}(\nu_{0})+\tfrac{\varepsilon_{0}}{2} − φ ~ ( ν ) < − φ ~ ( ν 0 ) + 2 ε 0 (claim 3 of Properties of the Absolute Value in an Ordered Field ), and within the lesser radius the two add to g ( ν ) < g ( ν 0 ) + ε 0 g(\nu)<g(\nu_{0})+\varepsilon_{0} g ( ν ) < g ( ν 0 ) + ε 0 . By Local Bounds for the Delta-Envelope and Attained Maxima on Closed Balls, for a Wasserstein-Coercive Penalty Pair §envelope-bound , applied to v v v with the radius r r r and the constant c c c , v δ − ( ν ) ≤ c − δ E ( ν ) v^{-}_{\delta}(\nu)\le c-\delta\,\mathcal{E}(\nu) v δ − ( ν ) ≤ c − δ E ( ν ) for ν ∈ D \nu\in\mathcal{D} ν ∈ D with W 2 ( ν , μ ^ ) < r W_{2}(\nu,\hat{\mu})<r W 2 ( ν , μ ^ ) < r ; and for ν ∈ K \nu\in K ν ∈ K , − φ ~ ( ν ) ≤ − φ ( ν ) < 1 − φ ( μ ^ ) -\tilde{\varphi}(\nu)\le-\varphi(\nu)<1-\varphi(\hat{\mu}) − φ ~ ( ν ) ≤ − φ ( ν ) < 1 − φ ( μ ^ ) , since β ψ 0 ( ν ) ≥ 0 \beta\psi_{0}(\nu)\ge0 β ψ 0 ( ν ) ≥ 0 . Hence g ( ν ) ≤ ( c + 1 − φ ( μ ^ ) ) − δ E ( ν ) g(\nu)\le\bigl(c+1-\varphi(\hat{\mu})\bigr)-\delta\,\mathcal{E}(\nu) g ( ν ) ≤ ( c + 1 − φ ( μ ^ ) ) − δ E ( ν ) for ν ∈ K \nu\in K ν ∈ K , and Local Bounds for the Delta-Envelope and Attained Maxima on Closed Balls, for a Wasserstein-Coercive Penalty Pair §attained , with the radius ρ \rho ρ and the constant c + 1 − φ ( μ ^ ) c+1-\varphi(\hat{\mu}) c + 1 − φ ( μ ^ ) , provides ν ^ ∈ K \hat{\nu}\in K ν ^ ∈ K with g ( ν ) ≤ g ( ν ^ ) g(\nu)\le g(\hat{\nu}) g ( ν ) ≤ g ( ν ^ ) for every ν ∈ K \nu\in K ν ∈ K .
Step 4: the maximiser is close to μ ^ \hat{\mu} μ ^ . By Basic Properties of the Delta-Envelopes on the Wasserstein Space §semicontinuity , v ( z ) − δ E ( z ) ≤ v δ − ( z ) v(z)-\delta\,\mathcal{E}(z)\le v^{-}_{\delta}(z) v ( z ) − δ E ( z ) ≤ v δ − ( z ) , so, using z ∈ K z\in K z ∈ K , the choice of v v v , of z z z and φ ~ ( μ ^ ) = φ ( μ ^ ) \tilde{\varphi}(\hat{\mu})=\varphi(\hat{\mu}) φ ~ ( μ ^ ) = φ ( μ ^ ) ,
g ( ν ^ ) ≥ g ( z ) ≥ v ( z ) − δ E ( z ) − φ ~ ( z ) > u ( z ) − δ E ( z ) − η − φ ~ ( z ) > u δ − ( μ ^ ) − 2 η − ( φ ( μ ^ ) + η ) = M − 3 η . g(\hat{\nu})\ge g(z)\ge v(z)-\delta\,\mathcal{E}(z)-\tilde{\varphi}(z)>u(z)-\delta\,\mathcal{E}(z)-\eta-\tilde{\varphi}(z)>u^{-}_{\delta}(\hat{\mu})-2\eta-\bigl(\varphi(\hat{\mu})+\eta\bigr)=M-3\eta . g ( ν ^ ) ≥ g ( z ) ≥ v ( z ) − δ E ( z ) − φ ~ ( z ) > u ( z ) − δ E ( z ) − η − φ ~ ( z ) > u δ − ( μ ^ ) − 2 η − ( φ ( μ ^ ) + η ) = M − 3 η .
On the other hand v δ − ( ν ^ ) ≤ u δ − ( ν ^ ) v^{-}_{\delta}(\hat{\nu})\le u^{-}_{\delta}(\hat{\nu}) v δ − ( ν ^ ) ≤ u δ − ( ν ^ ) by claim 1, and ν ^ ∈ K \hat{\nu}\in K ν ^ ∈ K satisfies (1), so g ( ν ^ ) ≤ M − β W 2 ( ν ^ , μ ^ ) 2 g(\hat{\nu})\le M-\beta\,W_{2}(\hat{\nu},\hat{\mu})^{2} g ( ν ^ ) ≤ M − β W 2 ( ν ^ , μ ^ ) 2 . Therefore β W 2 ( ν ^ , μ ^ ) 2 < 3 η ≤ β σ 2 \beta\,W_{2}(\hat{\nu},\hat{\mu})^{2}<3\eta\le\beta\sigma^{2} β W 2 ( ν ^ , μ ^ ) 2 < 3 η ≤ β σ 2 , whence W 2 ( ν ^ , μ ^ ) 2 < σ 2 W_{2}(\hat{\nu},\hat{\mu})^{2}<\sigma^{2} W 2 ( ν ^ , μ ^ ) 2 < σ 2 and W 2 ( ν ^ , μ ^ ) < σ W_{2}(\hat{\nu},\hat{\mu})<\sigma W 2 ( ν ^ , μ ^ ) < σ by claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field . In particular W 2 ( ν ^ , μ ^ ) W_{2}(\hat{\nu},\hat{\mu}) W 2 ( ν ^ , μ ^ ) is less than each of ρ , θ 1 , θ 2 , θ 3 , θ 4 \rho,\theta_{1},\theta_{2},\theta_{3},\theta_{4} ρ , θ 1 , θ 2 , θ 3 , θ 4 and ε \varepsilon ε , each of which is at least 2 σ 2\sigma 2 σ . Consequently
u δ − ( μ ^ ) − ε 4 < v δ − ( ν ^ ) < u δ − ( μ ^ ) + ε 4 : (2) u^{-}_{\delta}(\hat{\mu})-\tfrac{\varepsilon}{4}<v^{-}_{\delta}(\hat{\nu})<u^{-}_{\delta}(\hat{\mu})+\tfrac{\varepsilon}{4}: \tag{2} u δ − ( μ ^ ) − 4 ε < v δ − ( ν ^ ) < u δ − ( μ ^ ) + 4 ε : ( 2 )
the right inequality from v δ − ( ν ^ ) ≤ u δ − ( ν ^ ) v^{-}_{\delta}(\hat{\nu})\le u^{-}_{\delta}(\hat{\nu}) v δ − ( ν ^ ) ≤ u δ − ( ν ^ ) and the choice of θ 3 \theta_{3} θ 3 ; the left from v δ − ( ν ^ ) = g ( ν ^ ) + φ ~ ( ν ^ ) > M − 3 η + φ ~ ( μ ^ ) − ε 8 = u δ − ( μ ^ ) − 3 η − ε 8 v^{-}_{\delta}(\hat{\nu})=g(\hat{\nu})+\tilde{\varphi}(\hat{\nu})>M-3\eta+\tilde{\varphi}(\hat{\mu})-\tfrac{\varepsilon}{8}=u^{-}_{\delta}(\hat{\mu})-3\eta-\tfrac{\varepsilon}{8} v δ − ( ν ^ ) = g ( ν ^ ) + φ ~ ( ν ^ ) > M − 3 η + φ ~ ( μ ^ ) − 8 ε = u δ − ( μ ^ ) − 3 η − 8 ε , the choice of θ 4 \theta_{4} θ 4 and 3 η ≤ ε 8 3\eta\le\tfrac{\varepsilon}{8} 3 η ≤ 8 ε .
Step 5: the witnesses for v v v . The function D → R \mathcal{D}\to\mathbb{R} D → R with value v δ − ( ν ) − φ ~ ( ν ) v^{-}_{\delta}(\nu)-\tilde{\varphi}(\nu) v δ − ( ν ) − φ ~ ( ν ) has a local maximum at ν ^ \hat{\nu} ν ^ relative to D \mathcal{D} D , with radius ρ − W 2 ( ν ^ , μ ^ ) \rho-W_{2}(\hat{\nu},\hat{\mu}) ρ − W 2 ( ν ^ , μ ^ ) , positive since W 2 ( ν ^ , μ ^ ) < ρ W_{2}(\hat{\nu},\hat{\mu})<\rho W 2 ( ν ^ , μ ^ ) < ρ : if ν ∈ D \nu\in\mathcal{D} ν ∈ D and W 2 ( ν ^ , ν ) < ρ − W 2 ( ν ^ , μ ^ ) W_{2}(\hat{\nu},\nu)<\rho-W_{2}(\hat{\nu},\hat{\mu}) W 2 ( ν ^ , ν ) < ρ − W 2 ( ν ^ , μ ^ ) then W 2 ( ν , μ ^ ) < ρ W_{2}(\nu,\hat{\mu})<\rho W 2 ( ν , μ ^ ) < ρ , so ν ∈ K \nu\in K ν ∈ K and g ( ν ) ≤ g ( ν ^ ) g(\nu)\le g(\hat{\nu}) g ( ν ) ≤ g ( ν ^ ) . Put ε ′ ′ = min { ε 8 , σ } \varepsilon''=\min\{\tfrac{\varepsilon}{8},\sigma\} ε ′′ = min { 8 ε , σ } . Applying Viscosity Subsolution, Supersolution and Solution of a Second-Order Equation on the Wasserstein Space §subsolution to the viscosity subsolution v v v , with δ \delta δ , the intrinsic test function φ ~ \tilde{\varphi} φ ~ on D \mathcal{D} D , the point ν ^ \hat{\nu} ν ^ and the tolerance ε ′ ′ \varepsilon'' ε ′′ , we obtain ν ′ ∈ D Σ \nu'\in\mathcal{D}_{\Sigma} ν ′ ∈ D Σ , π ′ ∈ Π ( ν ′ , ν ^ ) \pi'\in\Pi(\nu',\hat{\nu}) π ′ ∈ Π ( ν ′ , ν ^ ) , s ∈ R s\in\mathbb{R} s ∈ R , q ∈ L 2 ( ν ′ ; R d ) q\in L^{2}(\nu';\mathbb{R}^{d}) q ∈ L 2 ( ν ′ ; R d ) and Y ∈ S ( d ) Y\in\mathcal{S}(d) Y ∈ S ( d ) with
I ( π ′ ) < ε ′ ′ 2 , ∣ v δ − ( ν ′ ) − v δ − ( ν ^ ) ∣ < ε ′ ′ , ∣ s − v δ − ( ν ^ ) ∣ < ε ′ ′ , I(\pi')<\varepsilon''^{2},\quad|v^{-}_{\delta}(\nu')-v^{-}_{\delta}(\hat{\nu})|<\varepsilon'',\quad|s-v^{-}_{\delta}(\hat{\nu})|<\varepsilon'', I ( π ′ ) < ε ′′ 2 , ∣ v δ − ( ν ′ ) − v δ − ( ν ^ ) ∣ < ε ′′ , ∣ s − v δ − ( ν ^ ) ∣ < ε ′′ ,
∫ R d + d ∥ q ( x ) − ∇ φ ~ ( ν ^ ) ( y ) ∥ 2 π ′ ( d z ) < ε ′ ′ 2 , ∥ Y − H φ ~ ( ν ^ ) ∥ < ε ′ ′ , F δ − ( ν ′ , s , q , Y ) ≤ ε ′ ′ . \int_{\mathbb{R}^{d+d}}\lVert q(x)-\nabla\tilde{\varphi}(\hat{\nu})(y)\rVert^{2}\,\pi'(dz)<\varepsilon''^{2},\quad\lVert Y-H_{\tilde{\varphi}}(\hat{\nu})\rVert<\varepsilon'',\quad F^{-}_{\delta}(\nu',s,q,Y)\le\varepsilon'' . ∫ R d + d ∥ q ( x ) − ∇ φ ~ ( ν ^ ) ( y ) ∥ 2 π ′ ( d z ) < ε ′′ 2 , ∥ Y − H φ ~ ( ν ^ )∥ < ε ′′ , F δ − ( ν ′ , s , q , Y ) ≤ ε ′′ .
Step 6: transport to μ ^ \hat{\mu} μ ^ . By Existence of an Optimal Coupling of Two Probability Measures with Finite Second Moment there is γ ∈ Π ( ν ^ , μ ^ ) \gamma\in\Pi(\hat{\nu},\hat{\mu}) γ ∈ Π ( ν ^ , μ ^ ) with I ( γ ) = W 2 ( ν ^ , μ ^ ) 2 I(\gamma)=W_{2}(\hat{\nu},\hat{\mu})^{2} I ( γ ) = W 2 ( ν ^ , μ ^ ) 2 . Let ς \varsigma ς be a gluing of π ′ \pi' π ′ and γ \gamma γ (Gluing Two Couplings over a Common Middle Marginal, and the Composite Coupling §glued ) and π = ( q 1 , q 3 ) # ς \pi=(\mathrm{q}_{1},\mathrm{q}_{3})_{\#}\varsigma π = ( q 1 , q 3 ) # ς , which belongs to Π ( ν ′ , μ ^ ) \Pi(\nu',\hat{\mu}) Π ( ν ′ , μ ^ ) with I ( π ) ≤ I ( π ′ ) + I ( γ ) < ε ′ ′ + W 2 ( ν ^ , μ ^ ) < 2 σ \sqrt{I(\pi)}\le\sqrt{I(\pi')}+\sqrt{I(\gamma)}<\varepsilon''+W_{2}(\hat{\nu},\hat{\mu})<2\sigma I ( π ) ≤ I ( π ′ ) + I ( γ ) < ε ′′ + W 2 ( ν ^ , μ ^ ) < 2 σ by Gluing Two Couplings over a Common Middle Marginal, and the Composite Coupling §composite and claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field . We check the six conditions of Viscosity Subsolution, Supersolution and Solution of a Second-Order Equation on the Wasserstein Space §subsolution for u u u at μ ^ \hat{\mu} μ ^ with φ \varphi φ and tolerance ε \varepsilon ε , with witnesses ν ′ \nu' ν ′ , π \pi π , s s s , q q q , Y Y Y .
First, I ( π ) < 2 σ ≤ ε \sqrt{I(\pi)}<2\sigma\le\varepsilon I ( π ) < 2 σ ≤ ε , so I ( π ) < ε 2 I(\pi)<\varepsilon^{2} I ( π ) < ε 2 by claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field .
Secondly, ν ′ ∈ D \nu'\in\mathcal{D} ν ′ ∈ D and W 2 ( ν ′ , μ ^ ) ≤ I ( π ) < 2 σ ≤ θ 3 W_{2}(\nu',\hat{\mu})\le\sqrt{I(\pi)}<2\sigma\le\theta_{3} W 2 ( ν ′ , μ ^ ) ≤ I ( π ) < 2 σ ≤ θ 3 , so u δ − ( ν ′ ) < u δ − ( μ ^ ) + ε 4 u^{-}_{\delta}(\nu')<u^{-}_{\delta}(\hat{\mu})+\tfrac{\varepsilon}{4} u δ − ( ν ′ ) < u δ − ( μ ^ ) + 4 ε ; and by claim 1 and (2), u δ − ( ν ′ ) ≥ v δ − ( ν ′ ) > v δ − ( ν ^ ) − ε ′ ′ > u δ − ( μ ^ ) − ε 4 − ε 8 u^{-}_{\delta}(\nu')\ge v^{-}_{\delta}(\nu')>v^{-}_{\delta}(\hat{\nu})-\varepsilon''>u^{-}_{\delta}(\hat{\mu})-\tfrac{\varepsilon}{4}-\tfrac{\varepsilon}{8} u δ − ( ν ′ ) ≥ v δ − ( ν ′ ) > v δ − ( ν ^ ) − ε ′′ > u δ − ( μ ^ ) − 4 ε − 8 ε . So ∣ u δ − ( ν ′ ) − u δ − ( μ ^ ) ∣ < ε |u^{-}_{\delta}(\nu')-u^{-}_{\delta}(\hat{\mu})|<\varepsilon ∣ u δ − ( ν ′ ) − u δ − ( μ ^ ) ∣ < ε by claim 9 of Properties of the Absolute Value in an Ordered Field .
Thirdly, by claim 5 of Properties of the Absolute Value in an Ordered Field and (2), ∣ s − u δ − ( μ ^ ) ∣ ≤ ∣ s − v δ − ( ν ^ ) ∣ + ∣ v δ − ( ν ^ ) − u δ − ( μ ^ ) ∣ < ε 8 + ε 4 < ε |s-u^{-}_{\delta}(\hat{\mu})|\le|s-v^{-}_{\delta}(\hat{\nu})|+|v^{-}_{\delta}(\hat{\nu})-u^{-}_{\delta}(\hat{\mu})|<\tfrac{\varepsilon}{8}+\tfrac{\varepsilon}{4}<\varepsilon ∣ s − u δ − ( μ ^ ) ∣ ≤ ∣ s − v δ − ( ν ^ ) ∣ + ∣ v δ − ( ν ^ ) − u δ − ( μ ^ ) ∣ < 8 ε + 4 ε < ε .
Fourthly, since ν ^ ∈ D \hat{\nu}\in\mathcal{D} ν ^ ∈ D and I ( γ ) = W 2 ( ν ^ , μ ^ ) 2 < θ 1 2 I(\gamma)=W_{2}(\hat{\nu},\hat{\mu})^{2}<\theta_{1}^{2} I ( γ ) = W 2 ( ν ^ , μ ^ ) 2 < θ 1 2 (claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field ), the choice of θ 1 \theta_{1} θ 1 bounds the discrepancy of ∇ φ ~ ( ν ^ ) \nabla\tilde{\varphi}(\hat{\nu}) ∇ φ ~ ( ν ^ ) and ∇ φ ~ ( μ ^ ) = ∇ φ ( μ ^ ) \nabla\tilde{\varphi}(\hat{\mu})=\nabla\varphi(\hat{\mu}) ∇ φ ~ ( μ ^ ) = ∇ φ ( μ ^ ) along γ \gamma γ by ( ε 4 ) 2 (\tfrac{\varepsilon}{4})^{2} ( 4 ε ) 2 . By A Triangle Inequality for Discrepancies Along a Composite Coupling §triangle , applied with q q q , ∇ φ ~ ( ν ^ ) \nabla\tilde{\varphi}(\hat{\nu}) ∇ φ ~ ( ν ^ ) and ∇ φ ( μ ^ ) \nabla\varphi(\hat{\mu}) ∇ φ ( μ ^ ) , and claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field ,
∫ R d + d ∥ q ( x ) − ∇ φ ( μ ^ ) ( y ) ∥ 2 π ( d z ) < ε ′ ′ + ε 4 < ε , \sqrt{\int_{\mathbb{R}^{d+d}}\lVert q(x)-\nabla\varphi(\hat{\mu})(y)\rVert^{2}\,\pi(dz)}<\varepsilon''+\tfrac{\varepsilon}{4}<\varepsilon , ∫ R d + d ∥ q ( x ) − ∇ φ ( μ ^ ) ( y ) ∥ 2 π ( d z ) < ε ′′ + 4 ε < ε ,
so the discrepancy of q q q and ∇ φ ( μ ^ ) \nabla\varphi(\hat{\mu}) ∇ φ ( μ ^ ) along π \pi π is less than ε 2 \varepsilon^{2} ε 2 .
Fifthly, all matrices below lie in S ( d ) \mathcal{S}(d) S ( d ) , differences of symmetric matrices being symmetric by claim 1 of The Positive Semidefinite Ordering is a Partial Order Compatible with the Linear Structure , and claim 5 of Properties of the Norm of a Symmetric Real Matrix gives
∥ Y − H φ ( μ ^ ) ∥ ≤ ∥ Y − H φ ~ ( ν ^ ) ∥ + ∥ H φ ~ ( ν ^ ) − H φ ~ ( μ ^ ) ∥ + ∥ H φ ~ ( μ ^ ) − H φ ( μ ^ ) ∥ < ε 8 + ε 4 + ε 4 < ε , \lVert Y-H_{\varphi}(\hat{\mu})\rVert\le\lVert Y-H_{\tilde{\varphi}}(\hat{\nu})\rVert+\lVert H_{\tilde{\varphi}}(\hat{\nu})-H_{\tilde{\varphi}}(\hat{\mu})\rVert+\lVert H_{\tilde{\varphi}}(\hat{\mu})-H_{\varphi}(\hat{\mu})\rVert<\tfrac{\varepsilon}{8}+\tfrac{\varepsilon}{4}+\tfrac{\varepsilon}{4}<\varepsilon , ∥ Y − H φ ( μ ^ )∥ ≤ ∥ Y − H φ ~ ( ν ^ )∥ + ∥ H φ ~ ( ν ^ ) − H φ ~ ( μ ^ )∥ + ∥ H φ ~ ( μ ^ ) − H φ ( μ ^ )∥ < 8 ε + 4 ε + 4 ε < ε ,
the middle term by the choice of θ 2 \theta_{2} θ 2 and the last by Step 1.
Finally, F δ − ( ν ′ , s , q , Y ) ≤ ε ′ ′ ≤ ε F^{-}_{\delta}(\nu',s,q,Y)\le\varepsilon''\le\varepsilon F δ − ( ν ′ , s , q , Y ) ≤ ε ′′ ≤ ε .
As δ \delta δ , φ \varphi φ , μ ^ \hat{\mu} μ ^ and ε \varepsilon ε were arbitrary and u u u is bounded above near each point, u u u is a viscosity subsolution of F F F relative to the penalty pair.