Each result cited is universally quantified over the data in its own statement, and is applied to the data named at the point of use. The symmetry and the triangle inequality of W a W_{a} W a on P ρ a \mathcal{P}^{a}_{\rho} P ρ a (The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §symmetry , The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §triangle ) and the rules of Elementary Order Arithmetic in an Ordered Field and Elementary Arithmetic in an Ordered Field for adding and scaling inequalities are used without further mention. Any two members of P ρ a \mathcal{P}^{a}_{\rho} P ρ a , equal or not, form a noise-connected ordered pair by The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §connected , and P ρ a ⊆ P 2 ( X ) \mathcal{P}^{a}_{\rho}\subseteq\mathcal{P}_{2}(X) P ρ a ⊆ P 2 ( X ) by The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §moments . For ν , ν 0 ∈ P ρ a \nu,\nu_{0}\in\mathcal{P}^{a}_{\rho} ν , ν 0 ∈ P ρ a and π ∈ Π a ( ν , ν 0 ) \pi\in\Pi^{a}(\nu,\nu_{0}) π ∈ Π a ( ν , ν 0 ) we have W a ( ν , ν 0 ) 2 ≤ I a ( π ) W_{a}(\nu,\nu_{0})^{2}\le I^{a}(\pi) W a ( ν , ν 0 ) 2 ≤ I a ( π ) by The Noise Wasserstein Distance §distance , hence W a ( ν , ν 0 ) ≤ I a ( π ) W_{a}(\nu,\nu_{0})\le\sqrt{I^{a}(\pi)} W a ( ν , ν 0 ) ≤ I a ( π ) by claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field ; and W a ( ν , ν ) = 0 W_{a}(\nu,\nu)=0 W a ( ν , ν ) = 0 by A Toolkit for Penalised Comparison on the Noise Wasserstein Space: Constant Test Functions and Linear Combinations of Test Functions, the Identity as a Unique Noise-Optimal Map, and Discrepancies Along the Push-Forward Under (id, id) and Along Glued Couplings §identity . For ν ∈ P ρ a \nu\in\mathcal{P}^{a}_{\rho} ν ∈ P ρ a , L 2 ( ν ; X a ) L^{2}(\nu;X^{a}) L 2 ( ν ; X a ) is a real vector space (Probability Measures on a Hilbert Space Transported in the Noise Norm: Standing Notation §fields ), in which gradients are added, subtracted and scaled. Each v ∈ S v\in\mathcal{S} v ∈ S , being a viscosity subsolution, has penalty-subordinate growth from above (Viscosity Subsolution, Supersolution and Solution of a First-Order Equation on the Noise Wasserstein Space Relative to a Noise Penalty Pair §subsolution ). By Basic Properties of a Noise-Closed Noise Penalty Pair: Lower Bound, Lower Semicontinuity, Complete Sublevel Sets and Bounded Distances §lsc the penalty E \mathcal{E} E is lower semicontinuous on D \mathcal{D} D relative to D \mathcal{D} D .
Claim 1. Let δ ∈ R \delta\in\mathbb{R} δ ∈ R be positive and let C C C be as in the assumption that S \mathcal{S} S is uniformly subordinate from above (The Pointwise Supremum of a Uniformly Subordinate Family of Viscosity Subsolutions on the Noise Wasserstein Space is a Viscosity Subsolution §uniform-growth ), for δ \delta δ . For μ ∈ D \mu\in\mathcal{D} μ ∈ D , C + δ E ( μ ) C+\delta\,\mathcal{E}(\mu) C + δ E ( μ ) is an upper bound of { v ( μ ) : v ∈ S } \{v(\mu):v\in\mathcal{S}\} { v ( μ ) : v ∈ S } , so u ( μ ) ≤ C + δ E ( μ ) u(\mu)\le C+\delta\,\mathcal{E}(\mu) u ( μ ) ≤ C + δ E ( μ ) by Upper Bound and Least Upper Bound . As δ \delta δ was arbitrary, u u u has penalty-subordinate growth from above by Penalty-Subordinate Growth of a Function on the Domain of a Noise Penalty Pair §above . For v ∈ S v\in\mathcal{S} v ∈ S and μ ∈ D \mu\in\mathcal{D} μ ∈ D , v ( μ ) v(\mu) v ( μ ) lies in the set of which u ( μ ) u(\mu) u ( μ ) is an upper bound, so v ( μ ) ≤ u ( μ ) v(\mu)\le u(\mu) v ( μ ) ≤ u ( μ ) . Consequently, for positive δ \delta δ the functions v − δ E v-\delta\mathcal{E} v − δ E and u − δ E u-\delta\mathcal{E} u − δ E on D \mathcal{D} D , both bounded above near each point of D \mathcal{D} D by The Delta-Envelopes of a Function on the Domain of a Noise Penalty Pair §minus , satisfy v − δ E ≤ u − δ E v-\delta\mathcal{E}\le u-\delta\mathcal{E} v − δ E ≤ u − δ E pointwise, and Properties of the Upper Semicontinuous Envelope §monotone , in the metric space ( P ρ a , W a ) (\mathcal{P}^{a}_{\rho},W_{a}) ( P ρ a , W a ) with S = D S=\mathcal{D} S = D , gives v δ − ( ν ) ≤ u δ − ( ν ) v^{-}_{\delta}(\nu)\le u^{-}_{\delta}(\nu) v δ − ( ν ) ≤ u δ − ( ν ) for ν ∈ D \nu\in\mathcal{D} ν ∈ D .
Claim 2. By claim 1, u u u has penalty-subordinate growth from above, so its δ \delta δ -envelopes u δ − u^{-}_{\delta} u δ − are defined. Let δ ∈ R \delta\in\mathbb{R} δ ∈ R satisfy 0 < δ < 1 0<\delta<1 0 < δ < 1 , as in Viscosity Subsolution, Supersolution and Solution of a First-Order Equation on the Noise Wasserstein Space Relative to a Noise Penalty Pair §subsolution , let φ \varphi φ be a noise intrinsic test function on D \mathcal{D} D , let μ ^ ∈ D \hat{\mu}\in\mathcal{D} μ ^ ∈ D be a point at which the function with value u δ − ( μ ) − φ ( μ ) u^{-}_{\delta}(\mu)-\varphi(\mu) u δ − ( μ ) − φ ( μ ) at μ ∈ D \mu\in\mathcal{D} μ ∈ D has a local maximum relative to D \mathcal{D} D , witnessed by a positive radius τ \tau τ as in Local Maximum of a Function Relative to a Subset of a Metric Space , and let ε ∈ R \varepsilon\in\mathbb{R} ε ∈ R be positive. Write M = u δ − ( μ ^ ) − φ ( μ ^ ) M=u^{-}_{\delta}(\hat{\mu})-\varphi(\hat{\mu}) M = u δ − ( μ ^ ) − φ ( μ ^ ) . The data below are chosen in the order: β \beta β and φ ~ \tilde{\varphi} φ ~ (Step 1); the radii θ 0 , θ 1 , θ 3 , θ 4 \theta_{0},\theta_{1},\theta_{3},\theta_{4} θ 0 , θ 1 , θ 3 , θ 4 , then C C C , then r 0 , σ , α r_{0},\sigma,\alpha r 0 , σ , α (Step 2); then θ α \theta_{\alpha} θ α , ζ \zeta ζ , v v v , ℓ \ell ℓ , B B B , ε ′ \varepsilon' ε ′ , ( c k ) k ∈ N (c_{k})_{k\in\mathbb{N}} ( c k ) k ∈ N , and then ν ^ \hat{\nu} ν ^ and ( ν k ) k ∈ N (\nu_{k})_{k\in\mathbb{N}} ( ν k ) k ∈ N (Step 3); then ε ′ ′ \varepsilon'' ε ′′ and the witnesses for v v v (Step 5).
Step 1: a strict maximum. Put β = ε 8 \beta=\tfrac{\varepsilon}{8} β = 8 ε and let ψ 0 : P ρ a → R \psi_{0}:\mathcal{P}^{a}_{\rho}\to\mathbb{R} ψ 0 : P ρ a → R be ψ 0 ( ν ) = W a ( ν , μ ^ ) 2 \psi_{0}(\nu)=W_{a}(\nu,\hat{\mu})^{2} ψ 0 ( ν ) = W a ( ν , μ ^ ) 2 . Since D \mathcal{D} D has the noise map property and μ ^ ∈ D \hat{\mu}\in\mathcal{D} μ ^ ∈ D , Squared Noise Wasserstein Distances, Their Convergent Series and Linear Combinations are Noise Intrinsic Test Functions §distance , applied with Q = D Q=\mathcal{D} Q = D and μ ^ \hat{\mu} μ ^ in the role of its ν 0 \nu_{0} ν 0 , shows that ψ 0 \psi_{0} ψ 0 is a noise intrinsic test function on D \mathcal{D} D with ∇ ψ 0 ( μ ^ ) = 2 ( i d − S ) \nabla\psi_{0}(\hat{\mu})=2(\mathrm{id}-S) ∇ ψ 0 ( μ ^ ) = 2 ( id − S ) for any noise-optimal map S S S from μ ^ \hat{\mu} μ ^ to μ ^ \hat{\mu} μ ^ . By A Toolkit for Penalised Comparison on the Noise Wasserstein Space: Constant Test Functions and Linear Combinations of Test Functions, the Identity as a Unique Noise-Optimal Map, and Discrepancies Along the Push-Forward Under (id, id) and Along Glued Couplings §identity , i d \mathrm{id} id is such a map and its displacement i d − i d \mathrm{id}-\mathrm{id} id − id is the zero element 0 μ ^ 0_{\hat{\mu}} 0 μ ^ of L 2 ( μ ^ ; X a ) L^{2}(\hat{\mu};X^{a}) L 2 ( μ ^ ; X a ) ; hence ∇ ψ 0 ( μ ^ ) = 2 ⋅ ( − 0 μ ^ ) = 0 μ ^ \nabla\psi_{0}(\hat{\mu})=2\cdot(-0_{\hat{\mu}})=0_{\hat{\mu}} ∇ ψ 0 ( μ ^ ) = 2 ⋅ ( − 0 μ ^ ) = 0 μ ^ . Let φ ~ = φ + β ψ 0 \tilde{\varphi}=\varphi+\beta\psi_{0} φ ~ = φ + β ψ 0 . By Squared Noise Wasserstein Distances, Their Convergent Series and Linear Combinations are Noise Intrinsic Test Functions §linear , with Q = D Q=\mathcal{D} Q = D , φ 1 = φ \varphi_{1}=\varphi φ 1 = φ , φ 2 = ψ 0 \varphi_{2}=\psi_{0} φ 2 = ψ 0 , s = 1 s=1 s = 1 and t = β t=\beta t = β , φ ~ \tilde{\varphi} φ ~ is a noise intrinsic test function on D \mathcal{D} D with
∇ φ ~ ( μ ^ ) = ∇ φ ( μ ^ ) + β 0 μ ^ = ∇ φ ( μ ^ ) , φ ~ ( μ ^ ) = φ ( μ ^ ) , \nabla\tilde{\varphi}(\hat{\mu})=\nabla\varphi(\hat{\mu})+\beta\,0_{\hat{\mu}}=\nabla\varphi(\hat{\mu}),\qquad\tilde{\varphi}(\hat{\mu})=\varphi(\hat{\mu}), ∇ φ ~ ( μ ^ ) = ∇ φ ( μ ^ ) + β 0 μ ^ = ∇ φ ( μ ^ ) , φ ~ ( μ ^ ) = φ ( μ ^ ) ,
the last because W a ( μ ^ , μ ^ ) = 0 W_{a}(\hat{\mu},\hat{\mu})=0 W a ( μ ^ , μ ^ ) = 0 . For ν ∈ D \nu\in\mathcal{D} ν ∈ D with W a ( ν , μ ^ ) < τ W_{a}(\nu,\hat{\mu})<\tau W a ( ν , μ ^ ) < τ the local maximum gives u δ − ( ν ) − φ ( ν ) ≤ M u^{-}_{\delta}(\nu)-\varphi(\nu)\le M u δ − ( ν ) − φ ( ν ) ≤ M , that is,
u δ − ( ν ) − φ ~ ( ν ) ≤ M − β W a ( ν , μ ^ ) 2 . (1) u^{-}_{\delta}(\nu)-\tilde{\varphi}(\nu)\le M-\beta\,W_{a}(\nu,\hat{\mu})^{2}. \tag{1} u δ − ( ν ) − φ ~ ( ν ) ≤ M − β W a ( ν , μ ^ ) 2 . ( 1 )
Step 2: radii. Using Continuous Map Between Metric Spaces for φ \varphi φ and φ ~ \tilde{\varphi} φ ~ , which are continuous on P ρ a \mathcal{P}^{a}_{\rho} P ρ a by property (a) of Noise Intrinsic Test Functions on the Noise Wasserstein Space §continuity , and Upper Semicontinuous Function on a Subset of a Metric Space for u δ − u^{-}_{\delta} u δ − , which is upper semicontinuous on D \mathcal{D} D by Basic Properties of the Delta-Envelopes on the Noise Wasserstein Space, and the Envelopes of Bounded Functions for a Noise-Closed Penalty Pair §semicontinuity , choose positive radii θ 0 , θ 3 , θ 4 \theta_{0},\theta_{3},\theta_{4} θ 0 , θ 3 , θ 4 such that for ν ∈ P ρ a \nu\in\mathcal{P}^{a}_{\rho} ν ∈ P ρ a : W a ( ν , μ ^ ) < θ 0 W_{a}(\nu,\hat{\mu})<\theta_{0} W a ( ν , μ ^ ) < θ 0 implies ∣ φ ( ν ) − φ ( μ ^ ) ∣ < 1 |\varphi(\nu)-\varphi(\hat{\mu})|<1 ∣ φ ( ν ) − φ ( μ ^ ) ∣ < 1 ; ν ∈ D \nu\in\mathcal{D} ν ∈ D and W a ( ν , μ ^ ) < θ 3 W_{a}(\nu,\hat{\mu})<\theta_{3} W a ( ν , μ ^ ) < θ 3 imply u δ − ( ν ) < u δ − ( μ ^ ) + ε 4 u^{-}_{\delta}(\nu)<u^{-}_{\delta}(\hat{\mu})+\tfrac{\varepsilon}{4} u δ − ( ν ) < u δ − ( μ ^ ) + 4 ε ; W a ( ν , μ ^ ) < θ 4 W_{a}(\nu,\hat{\mu})<\theta_{4} W a ( ν , μ ^ ) < θ 4 implies ∣ φ ~ ( ν ) − φ ~ ( μ ^ ) ∣ < ε 8 |\tilde{\varphi}(\nu)-\tilde{\varphi}(\hat{\mu})|<\tfrac{\varepsilon}{8} ∣ φ ~ ( ν ) − φ ~ ( μ ^ ) ∣ < 8 ε .
We also need a radius for the gradients: there is a positive θ 1 \theta_{1} θ 1 such that for every ν ∈ D \nu\in\mathcal{D} ν ∈ D and every π ∈ Π a ( ν , μ ^ ) \pi\in\Pi^{a}(\nu,\hat{\mu}) π ∈ Π a ( ν , μ ^ ) with I a ( π ) < θ 1 2 I^{a}(\pi)<\theta_{1}^{2} I a ( π ) < θ 1 2 the discrepancy ∫ X × X ∣ ∇ φ ~ ( ν ) ( x ) − ∇ φ ~ ( μ ^ ) ( y ) ∣ a 2 π ( d z ) \int_{X\times X}|\nabla\tilde{\varphi}(\nu)(x)-\nabla\tilde{\varphi}(\hat{\mu})(y)|_{a}^{2}\,\pi(dz) ∫ X × X ∣∇ φ ~ ( ν ) ( x ) − ∇ φ ~ ( μ ^ ) ( y ) ∣ a 2 π ( d z ) is less than ( ε 4 ) 2 (\tfrac{\varepsilon}{4})^{2} ( 4 ε ) 2 . Suppose not. Let ( h n ) n ∈ N (h_{n})_{n\in\mathbb{N}} ( h n ) n ∈ N be a sequence of positive reals with limit 0 0 0 (Existence of a Sequence of Positive Real Numbers with Limit Zero ); for each n n n there are ν n ∈ D \nu_{n}\in\mathcal{D} ν n ∈ D and π n ∈ Π a ( ν n , μ ^ ) \pi_{n}\in\Pi^{a}(\nu_{n},\hat{\mu}) π n ∈ Π a ( ν n , μ ^ ) with I a ( π n ) < h n 2 I^{a}(\pi_{n})<h_{n}^{2} I a ( π n ) < h n 2 and discrepancy D n ≥ ( ε 4 ) 2 D_{n}\ge(\tfrac{\varepsilon}{4})^{2} D n ≥ ( 4 ε ) 2 . Since 0 ≤ I a ( π n ) < h n 2 0\le I^{a}(\pi_{n})<h_{n}^{2} 0 ≤ I a ( π n ) < h n 2 (Couplings of Finite Noise Cost and Their Noise Cost §cost ) and ( h n 2 ) (h_{n}^{2}) ( h n 2 ) has limit 0 0 0 by Arithmetic of Limits of Real Sequences §products , ( I a ( π n ) ) (I^{a}(\pi_{n})) ( I a ( π n )) has limit 0 0 0 by claim 2 of Order Properties of Limits of Real Sequences , so ( π n ) n ∈ N (\pi_{n})_{n\in\mathbb{N}} ( π n ) n ∈ N is a sequence of couplings of vanishing noise cost from ( ν n ) n ∈ N (\nu_{n})_{n\in\mathbb{N}} ( ν n ) n ∈ N to μ ^ \hat{\mu} μ ^ (Strong and Weak Convergence of Noise Fields Along Couplings of Vanishing Noise Cost §couplings ). Property (c) of φ ~ \tilde{\varphi} φ ~ on D \mathcal{D} D (Noise Intrinsic Test Functions on the Noise Wasserstein Space §gradient-continuity ), at the point μ ^ ∈ D \hat{\mu}\in\mathcal{D} μ ^ ∈ D , then says that ( ∇ φ ~ ( ν n ) ) (\nabla\tilde{\varphi}(\nu_{n})) ( ∇ φ ~ ( ν n )) converges strongly to ∇ φ ~ ( μ ^ ) \nabla\tilde{\varphi}(\hat{\mu}) ∇ φ ~ ( μ ^ ) along ( π n ) (\pi_{n}) ( π n ) , that is (Strong and Weak Convergence of Noise Fields Along Couplings of Vanishing Noise Cost §strong ), ( D n ) (D_{n}) ( D n ) has limit 0 0 0 ; claim 1 of Order Properties of Limits of Real Sequences gives ( ε 4 ) 2 ≤ 0 (\tfrac{\varepsilon}{4})^{2}\le0 ( 4 ε ) 2 ≤ 0 , contradicting Elementary Order Arithmetic in an Ordered Field §positive-products .
Let C C C be as in the assumption that S \mathcal{S} S is uniformly subordinate from above, for the positive number δ 2 \tfrac{\delta}{2} 2 δ (Elementary Order Arithmetic in an Ordered Field §halving ). Put
r 0 = 1 2 min { τ , θ 0 } , σ = 1 2 min { r 0 , θ 1 , θ 3 , θ 4 , ε } , α = min { β σ 2 3 , ε 24 } , r_{0}=\tfrac12\min\{\tau,\theta_{0}\},\qquad\sigma=\tfrac12\min\{r_{0},\theta_{1},\theta_{3},\theta_{4},\varepsilon\},\qquad\alpha=\min\Bigl\{\tfrac{\beta\sigma^{2}}{3},\tfrac{\varepsilon}{24}\Bigr\}, r 0 = 2 1 min { τ , θ 0 } , σ = 2 1 min { r 0 , θ 1 , θ 3 , θ 4 , ε } , α = min { 3 β σ 2 , 24 ε } ,
all positive by claim 2 of Elementary Properties of the Minimum of Two Elements (applied repeatedly) and Elementary Order Arithmetic in an Ordered Field §halving , and let K = { ν ∈ D : W a ( ν , μ ^ ) ≤ r 0 } K=\{\nu\in\mathcal{D}:W_{a}(\nu,\hat{\mu})\le r_{0}\} K = { ν ∈ D : W a ( ν , μ ^ ) ≤ r 0 } . Since r 0 < τ r_{0}<\tau r 0 < τ and r 0 < θ 0 r_{0}<\theta_{0} r 0 < θ 0 , every ν ∈ K \nu\in K ν ∈ K satisfies (1) and ∣ φ ( ν ) − φ ( μ ^ ) ∣ < 1 |\varphi(\nu)-\varphi(\hat{\mu})|<1 ∣ φ ( ν ) − φ ( μ ^ ) ∣ < 1 .
Step 3: a member of the family and its maximiser. Let θ α \theta_{\alpha} θ α be a positive radius with ∣ φ ~ ( ν ) − φ ~ ( μ ^ ) ∣ < α |\tilde{\varphi}(\nu)-\tilde{\varphi}(\hat{\mu})|<\alpha ∣ φ ~ ( ν ) − φ ~ ( μ ^ ) ∣ < α whenever W a ( ν , μ ^ ) < θ α W_{a}(\nu,\hat{\mu})<\theta_{\alpha} W a ( ν , μ ^ ) < θ α (continuity of φ ~ \tilde{\varphi} φ ~ ), and put α ′ = min { α , 1 2 θ α , r 0 } \alpha'=\min\{\alpha,\tfrac12\theta_{\alpha},r_{0}\} α ′ = min { α , 2 1 θ α , r 0 } . By Properties of the Upper Semicontinuous Envelope §approximation , applied in ( P ρ a , W a ) (\mathcal{P}^{a}_{\rho},W_{a}) ( P ρ a , W a ) to u − δ E u-\delta\mathcal{E} u − δ E on S = D S=\mathcal{D} S = D at μ ^ \hat{\mu} μ ^ with α ′ \alpha' α ′ , there is ζ ∈ D \zeta\in\mathcal{D} ζ ∈ D with W a ( ζ , μ ^ ) ≤ α ′ W_{a}(\zeta,\hat{\mu})\le\alpha' W a ( ζ , μ ^ ) ≤ α ′ and ∣ u ( ζ ) − δ E ( ζ ) − u δ − ( μ ^ ) ∣ < α ′ ≤ α |u(\zeta)-\delta\,\mathcal{E}(\zeta)-u^{-}_{\delta}(\hat{\mu})|<\alpha'\le\alpha ∣ u ( ζ ) − δ E ( ζ ) − u δ − ( μ ^ ) ∣ < α ′ ≤ α . Then ζ ∈ K \zeta\in K ζ ∈ K and ∣ φ ~ ( ζ ) − φ ~ ( μ ^ ) ∣ < α |\tilde{\varphi}(\zeta)-\tilde{\varphi}(\hat{\mu})|<\alpha ∣ φ ~ ( ζ ) − φ ~ ( μ ^ ) ∣ < α . By Approximation Property of the Supremum and the Infimum in R \mathbb{R} R §epsilon-above there is v ∈ S v\in\mathcal{S} v ∈ S with u ( ζ ) − α < v ( ζ ) u(\zeta)-\alpha<v(\zeta) u ( ζ ) − α < v ( ζ ) .
Let g : D → R g:\mathcal{D}\to\mathbb{R} g : D → R be g ( ν ) = v δ − ( ν ) − φ ~ ( ν ) g(\nu)=v^{-}_{\delta}(\nu)-\tilde{\varphi}(\nu) g ( ν ) = v δ − ( ν ) − φ ~ ( ν ) . It is upper semicontinuous on D \mathcal{D} D : at ν 0 ∈ D \nu_{0}\in\mathcal{D} ν 0 ∈ D and for positive ε 0 \varepsilon_{0} ε 0 , Basic Properties of the Delta-Envelopes on the Noise Wasserstein Space, and the Envelopes of Bounded Functions for a Noise-Closed Penalty Pair §semicontinuity and Upper Semicontinuous Function on a Subset of a Metric Space give a radius within which v δ − ( ν ) < v δ − ( ν 0 ) + ε 0 2 v^{-}_{\delta}(\nu)<v^{-}_{\delta}(\nu_{0})+\tfrac{\varepsilon_{0}}{2} v δ − ( ν ) < v δ − ( ν 0 ) + 2 ε 0 , continuity of φ ~ \tilde{\varphi} φ ~ a radius within which − φ ~ ( ν ) < − φ ~ ( ν 0 ) + ε 0 2 -\tilde{\varphi}(\nu)<-\tilde{\varphi}(\nu_{0})+\tfrac{\varepsilon_{0}}{2} − φ ~ ( ν ) < − φ ~ ( ν 0 ) + 2 ε 0 (Properties of the Absolute Value in an Ordered Field §bounds ), and within the lesser radius the two add to g ( ν ) < g ( ν 0 ) + ε 0 g(\nu)<g(\nu_{0})+\varepsilon_{0} g ( ν ) < g ( ν 0 ) + ε 0 . Since v ( ν ) ≤ C + δ 2 E ( ν ) v(\nu)\le C+\tfrac{\delta}{2}\,\mathcal{E}(\nu) v ( ν ) ≤ C + 2 δ E ( ν ) for every ν ∈ D \nu\in\mathcal{D} ν ∈ D , 0 ≤ δ 2 ≤ δ 0\le\tfrac{\delta}{2}\le\delta 0 ≤ 2 δ ≤ δ and δ − δ 2 = δ 2 \delta-\tfrac{\delta}{2}=\tfrac{\delta}{2} δ − 2 δ = 2 δ , Basic Properties of the Delta-Envelopes on the Noise Wasserstein Space, and the Envelopes of Bounded Functions for a Noise-Closed Penalty Pair §bound , applied to v v v with δ 2 \tfrac{\delta}{2} 2 δ in the role of the weight there written η \eta η and with the constant C C C , gives v δ − ( ν ) ≤ C − δ 2 E ( ν ) v^{-}_{\delta}(\nu)\le C-\tfrac{\delta}{2}\,\mathcal{E}(\nu) v δ − ( ν ) ≤ C − 2 δ E ( ν ) for every ν ∈ D \nu\in\mathcal{D} ν ∈ D ; and for ν ∈ K \nu\in K ν ∈ K , − φ ~ ( ν ) ≤ − φ ( ν ) < 1 − φ ( μ ^ ) -\tilde{\varphi}(\nu)\le-\varphi(\nu)<1-\varphi(\hat{\mu}) − φ ~ ( ν ) ≤ − φ ( ν ) < 1 − φ ( μ ^ ) , since β ψ 0 ( ν ) ≥ 0 \beta\psi_{0}(\nu)\ge0 β ψ 0 ( ν ) ≥ 0 . Hence
g ( ν ) ≤ ( C + 1 − φ ( μ ^ ) ) − δ 2 E ( ν ) ( ν ∈ K ) . (3) g(\nu)\le\bigl(C+1-\varphi(\hat{\mu})\bigr)-\tfrac{\delta}{2}\,\mathcal{E}(\nu)\qquad(\nu\in K). \tag{3} g ( ν ) ≤ ( C + 1 − φ ( μ ^ ) ) − 2 δ E ( ν ) ( ν ∈ K ) . ( 3 )
On the other hand, for ν ∈ K \nu\in K ν ∈ K we have v δ − ( ν ) ≤ u δ − ( ν ) v^{-}_{\delta}(\nu)\le u^{-}_{\delta}(\nu) v δ − ( ν ) ≤ u δ − ( ν ) by claim 1, and ν \nu ν satisfies (1), so
g ( ν ) ≤ M − β W a ( ν , μ ^ ) 2 ≤ M ( ν ∈ K ) . (4) g(\nu)\le M-\beta\,W_{a}(\nu,\hat{\mu})^{2}\le M\qquad(\nu\in K). \tag{4} g ( ν ) ≤ M − β W a ( ν , μ ^ ) 2 ≤ M ( ν ∈ K ) . ( 4 )
By Basic Properties of the Delta-Envelopes on the Noise Wasserstein Space, and the Envelopes of Bounded Functions for a Noise-Closed Penalty Pair §semicontinuity , v ( ζ ) − δ E ( ζ ) ≤ v δ − ( ζ ) v(\zeta)-\delta\,\mathcal{E}(\zeta)\le v^{-}_{\delta}(\zeta) v ( ζ ) − δ E ( ζ ) ≤ v δ − ( ζ ) , so, by the choice of v v v , of ζ \zeta ζ and φ ~ ( μ ^ ) = φ ( μ ^ ) \tilde{\varphi}(\hat{\mu})=\varphi(\hat{\mu}) φ ~ ( μ ^ ) = φ ( μ ^ ) ,
g ( ζ ) ≥ v ( ζ ) − δ E ( ζ ) − φ ~ ( ζ ) > u ( ζ ) − δ E ( ζ ) − α − φ ~ ( ζ ) > u δ − ( μ ^ ) − 2 α − ( φ ( μ ^ ) + α ) = M − 3 α . (5) g(\zeta)\ge v(\zeta)-\delta\,\mathcal{E}(\zeta)-\tilde{\varphi}(\zeta)>u(\zeta)-\delta\,\mathcal{E}(\zeta)-\alpha-\tilde{\varphi}(\zeta)>u^{-}_{\delta}(\hat{\mu})-2\alpha-\bigl(\varphi(\hat{\mu})+\alpha\bigr)=M-3\alpha . \tag{5} g ( ζ ) ≥ v ( ζ ) − δ E ( ζ ) − φ ~ ( ζ ) > u ( ζ ) − δ E ( ζ ) − α − φ ~ ( ζ ) > u δ − ( μ ^ ) − 2 α − ( φ ( μ ^ ) + α ) = M − 3 α . ( 5 )
We perturb g g g by a Borwein--Preiss argument on a complete subset of K K K . Put
ℓ = max { E ( ζ ) , 2 δ ( C + 1 − φ ( μ ^ ) − M + 3 α ) } . \ell=\max\Bigl\{\mathcal{E}(\zeta),\ \tfrac{2}{\delta}\bigl(C+1-\varphi(\hat{\mu})-M+3\alpha\bigr)\Bigr\}. ℓ = max { E ( ζ ) , δ 2 ( C + 1 − φ ( μ ^ ) − M + 3 α ) } .
Then E ( ζ ) ≤ ℓ \mathcal{E}(\zeta)\le\ell E ( ζ ) ≤ ℓ by claim 1 of Elementary Properties of the Maximum of Two Elements , which also gives 2 δ ( C + 1 − φ ( μ ^ ) − M + 3 α ) ≤ ℓ \tfrac{2}{\delta}\bigl(C+1-\varphi(\hat{\mu})-M+3\alpha\bigr)\le\ell δ 2 ( C + 1 − φ ( μ ^ ) − M + 3 α ) ≤ ℓ ; so by (3) every ν ∈ K \nu\in K ν ∈ K with E ( ν ) > ℓ \mathcal{E}(\nu)>\ell E ( ν ) > ℓ satisfies g ( ν ) < ( C + 1 − φ ( μ ^ ) ) − δ 2 ℓ ≤ M − 3 α g(\nu)<\bigl(C+1-\varphi(\hat{\mu})\bigr)-\tfrac{\delta}{2}\,\ell\le M-3\alpha g ( ν ) < ( C + 1 − φ ( μ ^ ) ) − 2 δ ℓ ≤ M − 3 α , since δ 2 \tfrac{\delta}{2} 2 δ is positive; we record this as
g ( ν ) < M − 3 α for ν ∈ K with E ( ν ) > ℓ . (6) g(\nu)<M-3\alpha\qquad\text{for }\nu\in K\text{ with }\mathcal{E}(\nu)>\ell. \tag{6} g ( ν ) < M − 3 α for ν ∈ K with E ( ν ) > ℓ . ( 6 )
Write D ℓ = { ν ∈ D : E ( ν ) ≤ ℓ } \mathcal{D}_{\ell}=\{\nu\in\mathcal{D}:\mathcal{E}(\nu)\le\ell\} D ℓ = { ν ∈ D : E ( ν ) ≤ ℓ } , let B ∈ R B\in\mathbb{R} B ∈ R be as in Noise-Closed Noise Penalty Pairs §bounded for ℓ \ell ℓ , so that W a ( ν , ρ ) ≤ B W_{a}(\nu,\rho)\le B W a ( ν , ρ ) ≤ B for every ν ∈ D ℓ \nu\in\mathcal{D}_{\ell} ν ∈ D ℓ , and let K ℓ = { ν ∈ D ℓ : W a ( ν , μ ^ ) ≤ r 0 } = K ∩ D ℓ K_{\ell}=\{\nu\in\mathcal{D}_{\ell}:W_{a}(\nu,\hat{\mu})\le r_{0}\}=K\cap\mathcal{D}_{\ell} K ℓ = { ν ∈ D ℓ : W a ( ν , μ ^ ) ≤ r 0 } = K ∩ D ℓ , which contains ζ \zeta ζ . Since ζ ∈ D ℓ \zeta\in\mathcal{D}_{\ell} ζ ∈ D ℓ , 0 ≤ W a ( ζ , ρ ) ≤ B 0\le W_{a}(\zeta,\rho)\le B 0 ≤ W a ( ζ , ρ ) ≤ B , so 0 ≤ B 0\le B 0 ≤ B . By Basic Properties of a Noise-Closed Noise Penalty Pair: Lower Bound, Lower Semicontinuity, Complete Sublevel Sets and Bounded Distances §complete , ( D ℓ , W a ) (\mathcal{D}_{\ell},W_{a}) ( D ℓ , W a ) is a complete metric space. The restriction of W a W_{a} W a to K ℓ × K ℓ K_{\ell}\times K_{\ell} K ℓ × K ℓ , again written W a W_{a} W a , is a metric on K ℓ K_{\ell} K ℓ , the conditions of Metric Space being inherited from D ℓ \mathcal{D}_{\ell} D ℓ , and ( K ℓ , W a ) (K_{\ell},W_{a}) ( K ℓ , W a ) is complete in the sense of Complete Metric Space : a Cauchy sequence ( μ n ) n ∈ N (\mu_{n})_{n\in\mathbb{N}} ( μ n ) n ∈ N in ( K ℓ , W a ) (K_{\ell},W_{a}) ( K ℓ , W a ) is one in ( D ℓ , W a ) (\mathcal{D}_{\ell},W_{a}) ( D ℓ , W a ) , the distances being the same, so it converges in ( D ℓ , W a ) (\mathcal{D}_{\ell},W_{a}) ( D ℓ , W a ) to some μ ∈ D ℓ \mu\in\mathcal{D}_{\ell} μ ∈ D ℓ (Convergent Sequence in a Metric Space ); for every positive t t t there is n n n with W a ( μ n , μ ) < t W_{a}(\mu_{n},\mu)<t W a ( μ n , μ ) < t , whence W a ( μ , μ ^ ) ≤ W a ( μ , μ n ) + W a ( μ n , μ ^ ) < t + r 0 W_{a}(\mu,\hat{\mu})\le W_{a}(\mu,\mu_{n})+W_{a}(\mu_{n},\hat{\mu})<t+r_{0} W a ( μ , μ ^ ) ≤ W a ( μ , μ n ) + W a ( μ n , μ ^ ) < t + r 0 ; were W a ( μ , μ ^ ) > r 0 W_{a}(\mu,\hat{\mu})>r_{0} W a ( μ , μ ^ ) > r 0 , the choice t = W a ( μ , μ ^ ) − r 0 t=W_{a}(\mu,\hat{\mu})-r_{0} t = W a ( μ , μ ^ ) − r 0 would give W a ( μ , μ ^ ) < W a ( μ , μ ^ ) W_{a}(\mu,\hat{\mu})<W_{a}(\mu,\hat{\mu}) W a ( μ , μ ^ ) < W a ( μ , μ ^ ) ; so W a ( μ , μ ^ ) ≤ r 0 W_{a}(\mu,\hat{\mu})\le r_{0} W a ( μ , μ ^ ) ≤ r 0 , μ ∈ K ℓ \mu\in K_{\ell} μ ∈ K ℓ , and ( μ n ) n ∈ N (\mu_{n})_{n\in\mathbb{N}} ( μ n ) n ∈ N converges to μ \mu μ in ( K ℓ , W a ) (K_{\ell},W_{a}) ( K ℓ , W a ) .
The restriction of g g g to K ℓ K_{\ell} K ℓ is upper semicontinuous on K ℓ K_{\ell} K ℓ , the radii supplied by the upper semicontinuity of g g g on D \mathcal{D} D serving at each point of K ℓ ⊆ D K_{\ell}\subseteq\mathcal{D} K ℓ ⊆ D , and it is bounded above by M M M by (4). Hence { g ( ν ) : ν ∈ K ℓ } \{g(\nu):\nu\in K_{\ell}\} { g ( ν ) : ν ∈ K ℓ } , nonempty since ζ ∈ K ℓ \zeta\in K_{\ell} ζ ∈ K ℓ , has a least upper bound (The Real Numbers: Standing Notation and Background §bounds ), which is at most M M M by Upper Bound and Least Upper Bound , and by (5), g ( ζ ) > M − 3 α ≥ sup ν ∈ K ℓ g ( ν ) − 3 α g(\zeta)>M-3\alpha\ge\sup_{\nu\in K_{\ell}}g(\nu)-3\alpha g ( ζ ) > M − 3 α ≥ sup ν ∈ K ℓ g ( ν ) − 3 α . For ν , ν ′ ∈ K ℓ \nu,\nu'\in K_{\ell} ν , ν ′ ∈ K ℓ let κ ( ν , ν ′ ) = W a ( ν , ν ′ ) 2 \kappa(\nu,\nu')=W_{a}(\nu,\nu')^{2} κ ( ν , ν ′ ) = W a ( ν , ν ′ ) 2 . Then κ ( ν , ν ) = 0 \kappa(\nu,\nu)=0 κ ( ν , ν ) = 0 , and 0 ≤ W a ( ν , ν ′ ) ≤ 2 B 0\le W_{a}(\nu,\nu')\le2B 0 ≤ W a ( ν , ν ′ ) ≤ 2 B by Basic Properties of a Noise-Closed Noise Penalty Pair: Lower Bound, Lower Semicontinuity, Complete Sublevel Sets and Bounded Distances §diameter , applied with ℓ \ell ℓ and B B B , so 0 ≤ κ ( ν , ν ′ ) ≤ 4 B 2 0\le\kappa(\nu,\nu')\le4B^{2} 0 ≤ κ ( ν , ν ′ ) ≤ 4 B 2 by claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field . For ν ′ ∈ K ℓ \nu'\in K_{\ell} ν ′ ∈ K ℓ the function κ ( ⋅ , ν ′ ) \kappa(\cdot,\nu') κ ( ⋅ , ν ′ ) is the restriction to K ℓ K_{\ell} K ℓ of the function with value W a ( ν , ν ′ ) 2 W_{a}(\nu,\nu')^{2} W a ( ν , ν ′ ) 2 at ν ∈ P ρ a \nu\in\mathcal{P}^{a}_{\rho} ν ∈ P ρ a , which is a noise intrinsic test function on D \mathcal{D} D by Squared Noise Wasserstein Distances, Their Convergent Series and Linear Combinations are Noise Intrinsic Test Functions §distance (with Q = D Q=\mathcal{D} Q = D and ν ′ \nu' ν ′ as its ν 0 \nu_{0} ν 0 ) and so continuous on P ρ a \mathcal{P}^{a}_{\rho} P ρ a (Noise Intrinsic Test Functions on the Noise Wasserstein Space §continuity ); by Continuous Map Between Metric Spaces and Properties of the Absolute Value in an Ordered Field §bounds it is therefore lower semicontinuous on K ℓ K_{\ell} K ℓ in the sense of Lower Semicontinuous Function on a Subset of a Metric Space . For every positive t t t , the positive number t 2 4 \tfrac{t^{2}}{4} 4 t 2 has the property that κ ( ν , ν ′ ) ≤ t 2 4 \kappa(\nu,\nu')\le\tfrac{t^{2}}{4} κ ( ν , ν ′ ) ≤ 4 t 2 implies W a ( ν , ν ′ ) 2 < t 2 W_{a}(\nu,\nu')^{2}<t^{2} W a ( ν , ν ′ ) 2 < t 2 and hence W a ( ν , ν ′ ) < t W_{a}(\nu,\nu')<t W a ( ν , ν ′ ) < t , by claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field . Put
ε ′ = ε 32 ( 1 + r 0 ) , c k = ε ′ ( 1 2 ) k ( k ∈ N ) , \varepsilon'=\frac{\varepsilon}{32(1+r_{0})},\qquad c_{k}=\varepsilon'\bigl(\tfrac12\bigr)^{k}\quad(k\in\mathbb{N}), ε ′ = 32 ( 1 + r 0 ) ε , c k = ε ′ ( 2 1 ) k ( k ∈ N ) ,
all positive; by Series of Nonnegative Real Numbers, Comparison, and the Geometric Series §geometric and Elementary Properties of Series of Real Numbers §linearity the series ∑ k = 1 ∞ c k \sum_{k=1}^{\infty}c_{k} ∑ k = 1 ∞ c k converges, with sum ε ′ \varepsilon' ε ′ .
We apply A Smooth Variational Principle of Borwein-Preiss Type with a Gauge on a Complete Metric Space to the nonempty complete metric space ( K ℓ , W a ) (K_{\ell},W_{a}) ( K ℓ , W a ) , the restriction of g g g to K ℓ K_{\ell} K ℓ in the role of the function there written f f f , the gauge κ \kappa κ in the role of the function there written g g g , the bound 4 B 2 4B^{2} 4 B 2 in the role of its G G G , the weights ( c k ) k ∈ N (c_{k})_{k\in\mathbb{N}} ( c k ) k ∈ N , the tolerance 3 α 3\alpha 3 α in the role of its ε \varepsilon ε , and the point ζ \zeta ζ in the role of its x 1 x_{1} x 1 . It provides ν ^ ∈ K ℓ \hat{\nu}\in K_{\ell} ν ^ ∈ K ℓ and a sequence ( ν k ) k ∈ N (\nu_{k})_{k\in\mathbb{N}} ( ν k ) k ∈ N in K ℓ K_{\ell} K ℓ with ν 1 = ζ \nu_{1}=\zeta ν 1 = ζ . Since W a ( ν k , ρ ) ≤ B W_{a}(\nu_{k},\rho)\le B W a ( ν k , ρ ) ≤ B for every k k k and 0 ≤ B 0\le B 0 ≤ B , Squared Noise Wasserstein Distances, Their Convergent Series and Linear Combinations are Noise Intrinsic Test Functions §series-convergence , with B B B , ( ν k ) k ∈ N (\nu_{k})_{k\in\mathbb{N}} ( ν k ) k ∈ N and ( c k ) k ∈ N (c_{k})_{k\in\mathbb{N}} ( c k ) k ∈ N in the roles of its B B B , ( μ k ) k ∈ N (\mu_{k})_{k\in\mathbb{N}} ( μ k ) k ∈ N and ( β k ) k ∈ N (\beta_{k})_{k\in\mathbb{N}} ( β k ) k ∈ N , shows that ψ : P ρ a → R \psi:\mathcal{P}^{a}_{\rho}\to\mathbb{R} ψ : P ρ a → R , ψ ( ν ) = ∑ k = 1 ∞ c k W a ( ν , ν k ) 2 \psi(\nu)=\sum_{k=1}^{\infty}c_{k}\,W_{a}(\nu,\nu_{k})^{2} ψ ( ν ) = ∑ k = 1 ∞ c k W a ( ν , ν k ) 2 , is well defined, and that ∑ k = 1 ∞ c k W a ( ν , ν k ) \sum_{k=1}^{\infty}c_{k}\,W_{a}(\nu,\nu_{k}) ∑ k = 1 ∞ c k W a ( ν , ν k ) converges for every ν \nu ν ; and ψ ( ν ) ≥ 0 \psi(\nu)\ge0 ψ ( ν ) ≥ 0 for every ν \nu ν , its partial sums being nonnegative (Series of Nonnegative Real Numbers, Comparison, and the Geometric Series §dominates ). The function of the theorem is g − ψ g-\psi g − ψ on K ℓ K_{\ell} K ℓ , and by its clauses A Smooth Variational Principle of Borwein-Preiss Type with a Gauge on a Complete Metric Space §value and A Smooth Variational Principle of Borwein-Preiss Type with a Gauge on a Complete Metric Space §maximum ,
g ( ν ^ ) − ψ ( ν ^ ) ≥ g ( ζ ) > M − 3 α , g ( ν ) − ψ ( ν ) < g ( ν ^ ) − ψ ( ν ^ ) for ν ∈ K ℓ with ν ≠ ν ^ , (7) g(\hat{\nu})-\psi(\hat{\nu})\ge g(\zeta)>M-3\alpha,\qquad g(\nu)-\psi(\nu)<g(\hat{\nu})-\psi(\hat{\nu})\quad\text{for }\nu\in K_{\ell}\text{ with }\nu\ne\hat{\nu}, \tag{7} g ( ν ^ ) − ψ ( ν ^ ) ≥ g ( ζ ) > M − 3 α , g ( ν ) − ψ ( ν ) < g ( ν ^ ) − ψ ( ν ^ ) for ν ∈ K ℓ with ν = ν ^ , ( 7 )
the first by (5).
By Squared Noise Wasserstein Distances, Their Convergent Series and Linear Combinations are Noise Intrinsic Test Functions §series-test , with Q = D Q=\mathcal{D} Q = D , ψ \psi ψ is a noise intrinsic test function on D \mathcal{D} D , so by Squared Noise Wasserstein Distances, Their Convergent Series and Linear Combinations are Noise Intrinsic Test Functions §linear with s = t = 1 s=t=1 s = t = 1 so is φ ^ = φ ~ + ψ \hat{\varphi}=\tilde{\varphi}+\psi φ ^ = φ ~ + ψ , with ∇ φ ^ ( ν ^ ) = ∇ φ ~ ( ν ^ ) + ∇ ψ ( ν ^ ) \nabla\hat{\varphi}(\hat{\nu})=\nabla\tilde{\varphi}(\hat{\nu})+\nabla\psi(\hat{\nu}) ∇ φ ^ ( ν ^ ) = ∇ φ ~ ( ν ^ ) + ∇ ψ ( ν ^ ) . For each k k k , since ν ^ ∈ D \hat{\nu}\in\mathcal{D} ν ^ ∈ D and D \mathcal{D} D has the noise map property, the ordered pair ( ν ^ , ν k ) (\hat{\nu},\nu_{k}) ( ν ^ , ν k ) is uniquely noise-mapped (The Noise Map Property of a Set of Probability Measures §map-property ), so there is a noise-optimal map S k S_{k} S k from ν ^ \hat{\nu} ν ^ to ν k \nu_{k} ν k (Noise-Optimal Maps and Uniquely Noise-Mapped Pairs §uniquely-mapped ). Since ν ^ , ν k ∈ K \hat{\nu},\nu_{k}\in K ν ^ , ν k ∈ K , W a ( ν ^ , ν k ) ≤ W a ( ν ^ , μ ^ ) + W a ( μ ^ , ν k ) ≤ 2 r 0 W_{a}(\hat{\nu},\nu_{k})\le W_{a}(\hat{\nu},\hat{\mu})+W_{a}(\hat{\mu},\nu_{k})\le2r_{0} W a ( ν ^ , ν k ) ≤ W a ( ν ^ , μ ^ ) + W a ( μ ^ , ν k ) ≤ 2 r 0 , so by Squared Noise Wasserstein Distances, Their Convergent Series and Linear Combinations are Noise Intrinsic Test Functions §series-gradient (with these S k S_{k} S k ), Elementary Properties of Series of Real Numbers §order and Elementary Properties of Series of Real Numbers §linearity ,
∥ ∇ ψ ( ν ^ ) ∥ ν ^ ≤ 2 ∑ k = 1 ∞ c k W a ( ν ^ , ν k ) ≤ 4 r 0 ε ′ = r 0 ε 8 ( 1 + r 0 ) < ε 8 . (8) \lVert\nabla\psi(\hat{\nu})\rVert_{\hat{\nu}}\le2\sum_{k=1}^{\infty}c_{k}\,W_{a}(\hat{\nu},\nu_{k})\le4r_{0}\,\varepsilon'=\frac{r_{0}\,\varepsilon}{8(1+r_{0})}<\frac{\varepsilon}{8}. \tag{8} ∥ ∇ ψ ( ν ^ ) ∥ ν ^ ≤ 2 k = 1 ∑ ∞ c k W a ( ν ^ , ν k ) ≤ 4 r 0 ε ′ = 8 ( 1 + r 0 ) r 0 ε < 8 ε . ( 8 )
Step 4: the maximiser is close to μ ^ \hat{\mu} μ ^ . By (7) and ψ ( ν ^ ) ≥ 0 \psi(\hat{\nu})\ge0 ψ ( ν ^ ) ≥ 0 , g ( ν ^ ) ≥ g ( ν ^ ) − ψ ( ν ^ ) > M − 3 α g(\hat{\nu})\ge g(\hat{\nu})-\psi(\hat{\nu})>M-3\alpha g ( ν ^ ) ≥ g ( ν ^ ) − ψ ( ν ^ ) > M − 3 α ; on the other hand ν ^ ∈ K ℓ ⊆ K \hat{\nu}\in K_{\ell}\subseteq K ν ^ ∈ K ℓ ⊆ K , so (4) gives g ( ν ^ ) ≤ M − β W a ( ν ^ , μ ^ ) 2 g(\hat{\nu})\le M-\beta\,W_{a}(\hat{\nu},\hat{\mu})^{2} g ( ν ^ ) ≤ M − β W a ( ν ^ , μ ^ ) 2 . Therefore β W a ( ν ^ , μ ^ ) 2 < 3 α ≤ β σ 2 \beta\,W_{a}(\hat{\nu},\hat{\mu})^{2}<3\alpha\le\beta\sigma^{2} β W a ( ν ^ , μ ^ ) 2 < 3 α ≤ β σ 2 , whence W a ( ν ^ , μ ^ ) 2 < σ 2 W_{a}(\hat{\nu},\hat{\mu})^{2}<\sigma^{2} W a ( ν ^ , μ ^ ) 2 < σ 2 and W a ( ν ^ , μ ^ ) < σ W_{a}(\hat{\nu},\hat{\mu})<\sigma W a ( ν ^ , μ ^ ) < σ by claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field . In particular W a ( ν ^ , μ ^ ) W_{a}(\hat{\nu},\hat{\mu}) W a ( ν ^ , μ ^ ) is less than each of r 0 , θ 1 , θ 3 , θ 4 r_{0},\theta_{1},\theta_{3},\theta_{4} r 0 , θ 1 , θ 3 , θ 4 and ε \varepsilon ε , each of which is at least 2 σ 2\sigma 2 σ . Consequently
u δ − ( μ ^ ) − ε 4 < v δ − ( ν ^ ) < u δ − ( μ ^ ) + ε 4 : (2) u^{-}_{\delta}(\hat{\mu})-\tfrac{\varepsilon}{4}<v^{-}_{\delta}(\hat{\nu})<u^{-}_{\delta}(\hat{\mu})+\tfrac{\varepsilon}{4}: \tag{2} u δ − ( μ ^ ) − 4 ε < v δ − ( ν ^ ) < u δ − ( μ ^ ) + 4 ε : ( 2 )
the right inequality from v δ − ( ν ^ ) ≤ u δ − ( ν ^ ) v^{-}_{\delta}(\hat{\nu})\le u^{-}_{\delta}(\hat{\nu}) v δ − ( ν ^ ) ≤ u δ − ( ν ^ ) and the choice of θ 3 \theta_{3} θ 3 ; the left from v δ − ( ν ^ ) = g ( ν ^ ) + φ ~ ( ν ^ ) > M − 3 α + φ ~ ( μ ^ ) − ε 8 = u δ − ( μ ^ ) − 3 α − ε 8 v^{-}_{\delta}(\hat{\nu})=g(\hat{\nu})+\tilde{\varphi}(\hat{\nu})>M-3\alpha+\tilde{\varphi}(\hat{\mu})-\tfrac{\varepsilon}{8}=u^{-}_{\delta}(\hat{\mu})-3\alpha-\tfrac{\varepsilon}{8} v δ − ( ν ^ ) = g ( ν ^ ) + φ ~ ( ν ^ ) > M − 3 α + φ ~ ( μ ^ ) − 8 ε = u δ − ( μ ^ ) − 3 α − 8 ε , the choice of θ 4 \theta_{4} θ 4 and 3 α ≤ ε 8 3\alpha\le\tfrac{\varepsilon}{8} 3 α ≤ 8 ε .
Step 5: the witnesses for v v v . The function D → R \mathcal{D}\to\mathbb{R} D → R with value v δ − ( ν ) − φ ^ ( ν ) = g ( ν ) − ψ ( ν ) v^{-}_{\delta}(\nu)-\hat{\varphi}(\nu)=g(\nu)-\psi(\nu) v δ − ( ν ) − φ ^ ( ν ) = g ( ν ) − ψ ( ν ) has a local maximum at ν ^ \hat{\nu} ν ^ relative to D \mathcal{D} D , with radius r 0 − W a ( ν ^ , μ ^ ) r_{0}-W_{a}(\hat{\nu},\hat{\mu}) r 0 − W a ( ν ^ , μ ^ ) , positive since W a ( ν ^ , μ ^ ) < r 0 W_{a}(\hat{\nu},\hat{\mu})<r_{0} W a ( ν ^ , μ ^ ) < r 0 : if ν ∈ D \nu\in\mathcal{D} ν ∈ D and W a ( ν ^ , ν ) < r 0 − W a ( ν ^ , μ ^ ) W_{a}(\hat{\nu},\nu)<r_{0}-W_{a}(\hat{\nu},\hat{\mu}) W a ( ν ^ , ν ) < r 0 − W a ( ν ^ , μ ^ ) then W a ( ν , μ ^ ) < r 0 W_{a}(\nu,\hat{\mu})<r_{0} W a ( ν , μ ^ ) < r 0 , so ν ∈ K \nu\in K ν ∈ K ; if E ( ν ) ≤ ℓ \mathcal{E}(\nu)\le\ell E ( ν ) ≤ ℓ , then ν ∈ K ℓ \nu\in K_{\ell} ν ∈ K ℓ and g ( ν ) − ψ ( ν ) ≤ g ( ν ^ ) − ψ ( ν ^ ) g(\nu)-\psi(\nu)\le g(\hat{\nu})-\psi(\hat{\nu}) g ( ν ) − ψ ( ν ) ≤ g ( ν ^ ) − ψ ( ν ^ ) by (7), with equality when ν = ν ^ \nu=\hat{\nu} ν = ν ^ ; if E ( ν ) > ℓ \mathcal{E}(\nu)>\ell E ( ν ) > ℓ , then by ψ ( ν ) ≥ 0 \psi(\nu)\ge0 ψ ( ν ) ≥ 0 , (6) and (7), g ( ν ) − ψ ( ν ) ≤ g ( ν ) < M − 3 α < g ( ν ^ ) − ψ ( ν ^ ) g(\nu)-\psi(\nu)\le g(\nu)<M-3\alpha<g(\hat{\nu})-\psi(\hat{\nu}) g ( ν ) − ψ ( ν ) ≤ g ( ν ) < M − 3 α < g ( ν ^ ) − ψ ( ν ^ ) . Put ε ′ ′ = min { ε 8 , σ } \varepsilon''=\min\{\tfrac{\varepsilon}{8},\sigma\} ε ′′ = min { 8 ε , σ } . Applying Viscosity Subsolution, Supersolution and Solution of a First-Order Equation on the Noise Wasserstein Space Relative to a Noise Penalty Pair §subsolution to the viscosity subsolution v v v , with the same δ \delta δ , which satisfies 0 < δ < 1 0<\delta<1 0 < δ < 1 , the noise intrinsic test function φ ^ \hat{\varphi} φ ^ on D \mathcal{D} D , the point ν ^ \hat{\nu} ν ^ and the tolerance ε ′ ′ \varepsilon'' ε ′′ , we obtain ν ′ ∈ D Σ \nu'\in\mathcal{D}_{\Sigma} ν ′ ∈ D Σ , π ′ ∈ Π a ( ν ′ , ν ^ ) \pi'\in\Pi^{a}(\nu',\hat{\nu}) π ′ ∈ Π a ( ν ′ , ν ^ ) , s ∈ R s\in\mathbb{R} s ∈ R and q ∈ L 2 ( ν ′ ; X a ) q\in L^{2}(\nu';X^{a}) q ∈ L 2 ( ν ′ ; X a ) with
I a ( π ′ ) < ε ′ ′ 2 , ∣ v δ − ( ν ′ ) − v δ − ( ν ^ ) ∣ < ε ′ ′ , ∣ s − v δ − ( ν ^ ) ∣ < ε ′ ′ , I^{a}(\pi')<\varepsilon''^{2},\quad|v^{-}_{\delta}(\nu')-v^{-}_{\delta}(\hat{\nu})|<\varepsilon'',\quad|s-v^{-}_{\delta}(\hat{\nu})|<\varepsilon'', I a ( π ′ ) < ε ′′ 2 , ∣ v δ − ( ν ′ ) − v δ − ( ν ^ ) ∣ < ε ′′ , ∣ s − v δ − ( ν ^ ) ∣ < ε ′′ ,
∫ X × X ∣ q ( x ) − ∇ φ ^ ( ν ^ ) ( y ) ∣ a 2 π ′ ( d z ) < ε ′ ′ 2 , F δ − ( ν ′ , s , q ) ≤ ε ′ ′ . \int_{X\times X}|q(x)-\nabla\hat{\varphi}(\hat{\nu})(y)|_{a}^{2}\,\pi'(dz)<\varepsilon''^{2},\quad F^{-}_{\delta}(\nu',s,q)\le\varepsilon'' . ∫ X × X ∣ q ( x ) − ∇ φ ^ ( ν ^ ) ( y ) ∣ a 2 π ′ ( d z ) < ε ′′ 2 , F δ − ( ν ′ , s , q ) ≤ ε ′′ .
We next replace π ′ \pi' π ′ by a coupling along which q q q is compared with ∇ φ ~ ( ν ^ ) \nabla\tilde{\varphi}(\hat{\nu}) ∇ φ ~ ( ν ^ ) . Let Δ = ( i d , i d ) # ν ^ \Delta=(\mathrm{id},\mathrm{id})_{\#}\hat{\nu} Δ = ( id , id ) # ν ^ ; by A Toolkit for Penalised Comparison on the Noise Wasserstein Space: Constant Test Functions and Linear Combinations of Test Functions, the Identity as a Unique Noise-Optimal Map, and Discrepancies Along the Push-Forward Under (id, id) and Along Glued Couplings §diagonal , applied with ν ^ \hat{\nu} ν ^ as its ν \nu ν , Δ ∈ Π a ( ν ^ , ν ^ ) \Delta\in\Pi^{a}(\hat{\nu},\hat{\nu}) Δ ∈ Π a ( ν ^ , ν ^ ) and I a ( Δ ) = 0 I^{a}(\Delta)=0 I a ( Δ ) = 0 . All of ν ′ , ν ^ , μ ^ \nu',\hat{\nu},\hat{\mu} ν ′ , ν ^ , μ ^ lie in P 2 ( X ) \mathcal{P}_{2}(X) P 2 ( X ) . Let ς ′ \varsigma' ς ′ be a gluing of π ′ \pi' π ′ and Δ \Delta Δ (Gluing Two Couplings on a Hilbert Space over a Common Middle Marginal, and the Triangle Inequalities for the Quadratic and Noise Costs §glued , with ν ′ , ν ^ , ν ^ \nu',\hat{\nu},\hat{\nu} ν ′ , ν ^ , ν ^ in place of its μ , λ , ν \mu,\lambda,\nu μ , λ , ν ) and π ′ ′ = ( q 1 , q 3 ) # ς ′ \pi''=(q_{1},q_{3})_{\#}\varsigma' π ′′ = ( q 1 , q 3 ) # ς ′ , which by Gluing Two Couplings on a Hilbert Space over a Common Middle Marginal, and the Triangle Inequalities for the Quadratic and Noise Costs §noise-triangle belongs to Π a ( ν ′ , ν ^ ) \Pi^{a}(\nu',\hat{\nu}) Π a ( ν ′ , ν ^ ) with I a ( π ′ ′ ) ≤ I a ( π ′ ) + I a ( Δ ) = I a ( π ′ ) < ε ′ ′ \sqrt{I^{a}(\pi'')}\le\sqrt{I^{a}(\pi')}+\sqrt{I^{a}(\Delta)}=\sqrt{I^{a}(\pi')}<\varepsilon'' I a ( π ′′ ) ≤ I a ( π ′ ) + I a ( Δ ) = I a ( π ′ ) < ε ′′ , the last by claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field . By A Toolkit for Penalised Comparison on the Noise Wasserstein Space: Constant Test Functions and Linear Combinations of Test Functions, the Identity as a Unique Noise-Optimal Map, and Discrepancies Along the Push-Forward Under (id, id) and Along Glued Couplings §discrepancy-gluing , applied with ν ′ , ν ^ , ν ^ \nu',\hat{\nu},\hat{\nu} ν ′ , ν ^ , ν ^ as its ν , λ , μ \nu,\lambda,\mu ν , λ , μ , with π ′ \pi' π ′ , Δ \Delta Δ and ς ′ \varsigma' ς ′ as its π 12 \pi_{12} π 12 , π 23 \pi_{23} π 23 and σ \sigma σ , and with q q q , ∇ φ ^ ( ν ^ ) \nabla\hat{\varphi}(\hat{\nu}) ∇ φ ^ ( ν ^ ) and ∇ φ ~ ( ν ^ ) \nabla\tilde{\varphi}(\hat{\nu}) ∇ φ ~ ( ν ^ ) as its q q q , η \eta η and θ \theta θ ; by A Toolkit for Penalised Comparison on the Noise Wasserstein Space: Constant Test Functions and Linear Combinations of Test Functions, the Identity as a Unique Noise-Optimal Map, and Discrepancies Along the Push-Forward Under (id, id) and Along Glued Couplings §diagonal , which gives ∫ X × X ∣ ∇ φ ^ ( ν ^ ) ( x ) − ∇ φ ~ ( ν ^ ) ( y ) ∣ a 2 Δ ( d z ) = ∥ ∇ φ ^ ( ν ^ ) − ∇ φ ~ ( ν ^ ) ∥ ν ^ 2 \int_{X\times X}|\nabla\hat{\varphi}(\hat{\nu})(x)-\nabla\tilde{\varphi}(\hat{\nu})(y)|_{a}^{2}\,\Delta(dz)=\lVert\nabla\hat{\varphi}(\hat{\nu})-\nabla\tilde{\varphi}(\hat{\nu})\rVert_{\hat{\nu}}^{2} ∫ X × X ∣∇ φ ^ ( ν ^ ) ( x ) − ∇ φ ~ ( ν ^ ) ( y ) ∣ a 2 Δ ( d z ) = ∥ ∇ φ ^ ( ν ^ ) − ∇ φ ~ ( ν ^ ) ∥ ν ^ 2 , whose nonnegative square root is ∥ ∇ φ ^ ( ν ^ ) − ∇ φ ~ ( ν ^ ) ∥ ν ^ \lVert\nabla\hat{\varphi}(\hat{\nu})-\nabla\tilde{\varphi}(\hat{\nu})\rVert_{\hat{\nu}} ∥ ∇ φ ^ ( ν ^ ) − ∇ φ ~ ( ν ^ ) ∥ ν ^ (Existence and Uniqueness of the Nonnegative Square Root ); by ∇ φ ^ ( ν ^ ) − ∇ φ ~ ( ν ^ ) = ∇ ψ ( ν ^ ) \nabla\hat{\varphi}(\hat{\nu})-\nabla\tilde{\varphi}(\hat{\nu})=\nabla\psi(\hat{\nu}) ∇ φ ^ ( ν ^ ) − ∇ φ ~ ( ν ^ ) = ∇ ψ ( ν ^ ) (Step 3); by claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field and by (8),
∫ X × X ∣ q ( x ) − ∇ φ ~ ( ν ^ ) ( y ) ∣ a 2 π ′ ′ ( d z ) < ε ′ ′ + ∥ ∇ ψ ( ν ^ ) ∥ ν ^ < ε ′ ′ + ε 8 . (10) \sqrt{\int_{X\times X}|q(x)-\nabla\tilde{\varphi}(\hat{\nu})(y)|_{a}^{2}\,\pi''(dz)}<\varepsilon''+\lVert\nabla\psi(\hat{\nu})\rVert_{\hat{\nu}}<\varepsilon''+\tfrac{\varepsilon}{8}. \tag{10} ∫ X × X ∣ q ( x ) − ∇ φ ~ ( ν ^ ) ( y ) ∣ a 2 π ′′ ( d z ) < ε ′′ + ∥ ∇ ψ ( ν ^ ) ∥ ν ^ < ε ′′ + 8 ε . ( 10 )
Step 6: transport to μ ^ \hat{\mu} μ ^ . By The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §optimal there is a noise-optimal coupling γ ∈ Π a ( ν ^ , μ ^ ) \gamma\in\Pi^{a}(\hat{\nu},\hat{\mu}) γ ∈ Π a ( ν ^ , μ ^ ) , so I a ( γ ) = W a ( ν ^ , μ ^ ) 2 I^{a}(\gamma)=W_{a}(\hat{\nu},\hat{\mu})^{2} I a ( γ ) = W a ( ν ^ , μ ^ ) 2 (Noise-Optimal Couplings §optimal ) and I a ( γ ) = W a ( ν ^ , μ ^ ) \sqrt{I^{a}(\gamma)}=W_{a}(\hat{\nu},\hat{\mu}) I a ( γ ) = W a ( ν ^ , μ ^ ) (Existence and Uniqueness of the Nonnegative Square Root ). Let ς \varsigma ς be a gluing of π ′ ′ \pi'' π ′′ and γ \gamma γ (Gluing Two Couplings on a Hilbert Space over a Common Middle Marginal, and the Triangle Inequalities for the Quadratic and Noise Costs §glued , with ν ′ , ν ^ , μ ^ \nu',\hat{\nu},\hat{\mu} ν ′ , ν ^ , μ ^ in place of its μ , λ , ν \mu,\lambda,\nu μ , λ , ν ) and π = ( q 1 , q 3 ) # ς \pi=(q_{1},q_{3})_{\#}\varsigma π = ( q 1 , q 3 ) # ς , which by Gluing Two Couplings on a Hilbert Space over a Common Middle Marginal, and the Triangle Inequalities for the Quadratic and Noise Costs §noise-triangle belongs to Π a ( ν ′ , μ ^ ) \Pi^{a}(\nu',\hat{\mu}) Π a ( ν ′ , μ ^ ) with I a ( π ) ≤ I a ( π ′ ′ ) + I a ( γ ) < ε ′ ′ + W a ( ν ^ , μ ^ ) < 2 σ \sqrt{I^{a}(\pi)}\le\sqrt{I^{a}(\pi'')}+\sqrt{I^{a}(\gamma)}<\varepsilon''+W_{a}(\hat{\nu},\hat{\mu})<2\sigma I a ( π ) ≤ I a ( π ′′ ) + I a ( γ ) < ε ′′ + W a ( ν ^ , μ ^ ) < 2 σ . We check the five conditions of Viscosity Subsolution, Supersolution and Solution of a First-Order Equation on the Noise Wasserstein Space Relative to a Noise Penalty Pair §subsolution for u u u at μ ^ \hat{\mu} μ ^ with φ \varphi φ and tolerance ε \varepsilon ε , with witnesses ν ′ \nu' ν ′ , π \pi π , s s s , q q q .
First, I a ( π ) < 2 σ ≤ ε \sqrt{I^{a}(\pi)}<2\sigma\le\varepsilon I a ( π ) < 2 σ ≤ ε , so I a ( π ) < ε 2 I^{a}(\pi)<\varepsilon^{2} I a ( π ) < ε 2 by claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field .
Secondly, ν ′ ∈ D \nu'\in\mathcal{D} ν ′ ∈ D and W a ( ν ′ , μ ^ ) ≤ I a ( π ) < 2 σ ≤ θ 3 W_{a}(\nu',\hat{\mu})\le\sqrt{I^{a}(\pi)}<2\sigma\le\theta_{3} W a ( ν ′ , μ ^ ) ≤ I a ( π ) < 2 σ ≤ θ 3 , so u δ − ( ν ′ ) < u δ − ( μ ^ ) + ε 4 u^{-}_{\delta}(\nu')<u^{-}_{\delta}(\hat{\mu})+\tfrac{\varepsilon}{4} u δ − ( ν ′ ) < u δ − ( μ ^ ) + 4 ε ; and by claim 1 and (2), u δ − ( ν ′ ) ≥ v δ − ( ν ′ ) > v δ − ( ν ^ ) − ε ′ ′ > u δ − ( μ ^ ) − ε 4 − ε 8 u^{-}_{\delta}(\nu')\ge v^{-}_{\delta}(\nu')>v^{-}_{\delta}(\hat{\nu})-\varepsilon''>u^{-}_{\delta}(\hat{\mu})-\tfrac{\varepsilon}{4}-\tfrac{\varepsilon}{8} u δ − ( ν ′ ) ≥ v δ − ( ν ′ ) > v δ − ( ν ^ ) − ε ′′ > u δ − ( μ ^ ) − 4 ε − 8 ε . So ∣ u δ − ( ν ′ ) − u δ − ( μ ^ ) ∣ < ε |u^{-}_{\delta}(\nu')-u^{-}_{\delta}(\hat{\mu})|<\varepsilon ∣ u δ − ( ν ′ ) − u δ − ( μ ^ ) ∣ < ε by Properties of the Absolute Value in an Ordered Field §strict-two-sided .
Thirdly, by Properties of the Absolute Value in an Ordered Field §triangle and (2), ∣ s − u δ − ( μ ^ ) ∣ ≤ ∣ s − v δ − ( ν ^ ) ∣ + ∣ v δ − ( ν ^ ) − u δ − ( μ ^ ) ∣ < ε 8 + ε 4 < ε |s-u^{-}_{\delta}(\hat{\mu})|\le|s-v^{-}_{\delta}(\hat{\nu})|+|v^{-}_{\delta}(\hat{\nu})-u^{-}_{\delta}(\hat{\mu})|<\tfrac{\varepsilon}{8}+\tfrac{\varepsilon}{4}<\varepsilon ∣ s − u δ − ( μ ^ ) ∣ ≤ ∣ s − v δ − ( ν ^ ) ∣ + ∣ v δ − ( ν ^ ) − u δ − ( μ ^ ) ∣ < 8 ε + 4 ε < ε .
Fourthly, since ν ^ ∈ D \hat{\nu}\in\mathcal{D} ν ^ ∈ D , γ ∈ Π a ( ν ^ , μ ^ ) \gamma\in\Pi^{a}(\hat{\nu},\hat{\mu}) γ ∈ Π a ( ν ^ , μ ^ ) and I a ( γ ) = W a ( ν ^ , μ ^ ) 2 < θ 1 2 I^{a}(\gamma)=W_{a}(\hat{\nu},\hat{\mu})^{2}<\theta_{1}^{2} I a ( γ ) = W a ( ν ^ , μ ^ ) 2 < θ 1 2 (claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field ), the choice of θ 1 \theta_{1} θ 1 bounds the discrepancy of ∇ φ ~ ( ν ^ ) \nabla\tilde{\varphi}(\hat{\nu}) ∇ φ ~ ( ν ^ ) and ∇ φ ~ ( μ ^ ) = ∇ φ ( μ ^ ) \nabla\tilde{\varphi}(\hat{\mu})=\nabla\varphi(\hat{\mu}) ∇ φ ~ ( μ ^ ) = ∇ φ ( μ ^ ) along γ \gamma γ by ( ε 4 ) 2 (\tfrac{\varepsilon}{4})^{2} ( 4 ε ) 2 , so its square root is less than ε 4 \tfrac{\varepsilon}{4} 4 ε (claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field ). By A Toolkit for Penalised Comparison on the Noise Wasserstein Space: Constant Test Functions and Linear Combinations of Test Functions, the Identity as a Unique Noise-Optimal Map, and Discrepancies Along the Push-Forward Under (id, id) and Along Glued Couplings §discrepancy-gluing , applied with ν ′ , ν ^ , μ ^ \nu',\hat{\nu},\hat{\mu} ν ′ , ν ^ , μ ^ as its ν , λ , μ \nu,\lambda,\mu ν , λ , μ , with π ′ ′ \pi'' π ′′ , γ \gamma γ and ς \varsigma ς as its π 12 \pi_{12} π 12 , π 23 \pi_{23} π 23 and σ \sigma σ , and with q q q , ∇ φ ~ ( ν ^ ) \nabla\tilde{\varphi}(\hat{\nu}) ∇ φ ~ ( ν ^ ) and ∇ φ ( μ ^ ) \nabla\varphi(\hat{\mu}) ∇ φ ( μ ^ ) as its q q q , η \eta η and θ \theta θ , and by (10),
∫ X × X ∣ q ( x ) − ∇ φ ( μ ^ ) ( y ) ∣ a 2 π ( d z ) < ε ′ ′ + ε 8 + ε 4 < ε , \sqrt{\int_{X\times X}|q(x)-\nabla\varphi(\hat{\mu})(y)|_{a}^{2}\,\pi(dz)}<\varepsilon''+\tfrac{\varepsilon}{8}+\tfrac{\varepsilon}{4}<\varepsilon , ∫ X × X ∣ q ( x ) − ∇ φ ( μ ^ ) ( y ) ∣ a 2 π ( d z ) < ε ′′ + 8 ε + 4 ε < ε ,
so the discrepancy of q q q and ∇ φ ( μ ^ ) \nabla\varphi(\hat{\mu}) ∇ φ ( μ ^ ) along π \pi π is less than ε 2 \varepsilon^{2} ε 2 by claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field .
Finally, F δ − ( ν ′ , s , q ) ≤ ε ′ ′ ≤ ε F^{-}_{\delta}(\nu',s,q)\le\varepsilon''\le\varepsilon F δ − ( ν ′ , s , q ) ≤ ε ′′ ≤ ε .
As δ \delta δ , φ \varphi φ , μ ^ \hat{\mu} μ ^ and ε \varepsilon ε were arbitrary and u u u has penalty-subordinate growth from above, u u u is a viscosity subsolution of F F F relative to the noise penalty pair by Viscosity Subsolution, Supersolution and Solution of a First-Order Equation on the Noise Wasserstein Space Relative to a Noise Penalty Pair §subsolution .