Each result cited is universally quantified over the data in its own statement. Elementary arithmetic and order of real numbers, square roots of nonnegative reals, the monotonicity of squaring and of square roots, limits of real sequences (their arithmetic, their order properties and the squeeze principle) and the Archimedean property are used freely; they are carried by The Real Numbers: Standing Notation and Background §background . For z ∈ X × X z\in X\times X z ∈ X × X we write x = π 1 ( z ) x=\pi_{1}(z) x = π 1 ( z ) and y = π 2 ( z ) y=\pi_{2}(z) y = π 2 ( z ) , as in Probability Measures on a Hilbert Space Transported in the Noise Norm: Standing Notation §background . The noise space X a X^{a} X a is a linear subspace of X X X and a real Hilbert space with inner product ⟨ ⋅ , ⋅ ⟩ a \langle\cdot,\cdot\rangle_{a} ⟨ ⋅ , ⋅ ⟩ a and norm ∣ ⋅ ∣ a |\cdot|_{a} ∣ ⋅ ∣ a , by The Noise Space is a Real Hilbert Space: Orthonormal Basis, Continuous Embedding, Partial Sums, Closed Balls and Borel Measurability §hilbert ; for λ ∈ P ( X ) \lambda\in\mathcal{P}(X) λ ∈ P ( X ) and γ ∈ P ( X × X ) \gamma\in\mathcal{P}(X\times X) γ ∈ P ( X × X ) the spaces L 2 ( λ ; X a ) L^{2}(\lambda;X^{a}) L 2 ( λ ; X a ) and L 2 ( γ ; X a ) L^{2}(\gamma;X^{a}) L 2 ( γ ; X a ) are real Hilbert spaces, by Probability Measures on a Hilbert Space Transported in the Noise Norm: Standing Notation §fields . In each of these spaces we use the identities Elementary Identities in a Real Inner Product Space §bilinear (bilinearity), Elementary Identities in a Real Inner Product Space §homogeneity (homogeneity of the norm) and Elementary Identities in a Real Inner Product Space §expansion (expansion), the Cauchy--Schwarz inequality The Cauchy-Schwarz Inequality in a Real Inner Product Space and the triangle inequality of The Norm Metric of a Real Inner Product Space: Triangle Inequalities, Limits and Continuity §triangle . Elements of the L 2 L^{2} L 2 spaces are denoted by the same symbols as representatives (The Space of Square-Integrable Maps from a Measure Space into a Hilbert Space with an Orthonormal Basis §convention ), and sums, real multiples, inner products and norms are computed on representatives as in The Space of Square-Integrable Maps from a Measure Space into a Hilbert Space with an Orthonormal Basis §operations . Two facts are used throughout. First, for λ , λ ′ ∈ P ρ a \lambda,\lambda'\in\mathcal{P}^{a}_{\rho} λ , λ ′ ∈ P ρ a and γ ∈ Π a ( λ , λ ′ ) \gamma\in\Pi^{a}(\lambda,\lambda') γ ∈ Π a ( λ , λ ′ ) ,
W a ( λ , λ ′ ) ≤ I a ( γ ) , (W) W_{a}(\lambda,\lambda')\le\sqrt{I^{a}(\gamma)},\tag{W} W a ( λ , λ ′ ) ≤ I a ( γ ) , ( W )
by The Noise Wasserstein Distance §distance and the monotonicity of square roots. Second, if T T T is a noise-optimal map from λ ∈ P ρ a \lambda\in\mathcal{P}^{a}_{\rho} λ ∈ P ρ a to λ ′ ∈ P ρ a \lambda'\in\mathcal{P}^{a}_{\rho} λ ′ ∈ P ρ a , then the class T − i d T-\mathrm{id} T − id is represented by x ↦ T ( x ) − x x\mapsto T(x)-x x ↦ T ( x ) − x , a Borel map from X X X to X X X with values in X a X^{a} X a (Noise-Optimal Maps and Uniquely Noise-Mapped Pairs , preamble and clause Noise-Optimal Maps and Uniquely Noise-Mapped Pairs §map ); since i d − T = − ( T − i d ) = ( − 1 ) ( T − i d ) \mathrm{id}-T=-(T-\mathrm{id})=(-1)(T-\mathrm{id}) id − T = − ( T − id ) = ( − 1 ) ( T − id ) by claim 5 of Elementary Identities in a Vector Space , the class 2 ( i d − T ) = ( − 2 ) ( T − i d ) 2(\mathrm{id}-T)=(-2)(T-\mathrm{id}) 2 ( id − T ) = ( − 2 ) ( T − id ) is represented by x ↦ ( − 2 ) ( T ( x ) − x ) x\mapsto(-2)(T(x)-x) x ↦ ( − 2 ) ( T ( x ) − x ) .
Step 1 (Elementary bounds for W a W_{a} W a ). Let λ , μ ′ , μ ′ ′ ∈ P ρ a \lambda,\mu',\mu''\in\mathcal{P}^{a}_{\rho} λ , μ ′ , μ ′′ ∈ P ρ a . By the triangle inequality and the symmetry of W a W_{a} W a (The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §triangle , The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §symmetry ), W a ( μ ′ ′ , λ ) ≤ W a ( μ ′ , μ ′ ′ ) + W a ( μ ′ , λ ) W_{a}(\mu'',\lambda)\le W_{a}(\mu',\mu'')+W_{a}(\mu',\lambda) W a ( μ ′′ , λ ) ≤ W a ( μ ′ , μ ′′ ) + W a ( μ ′ , λ ) and W a ( μ ′ , λ ) ≤ W a ( μ ′ , μ ′ ′ ) + W a ( μ ′ ′ , λ ) W_{a}(\mu',\lambda)\le W_{a}(\mu',\mu'')+W_{a}(\mu'',\lambda) W a ( μ ′ , λ ) ≤ W a ( μ ′ , μ ′′ ) + W a ( μ ′′ , λ ) . Hence (i) ∣ W a ( μ ′ ′ , λ ) − W a ( μ ′ , λ ) ∣ ≤ W a ( μ ′ , μ ′ ′ ) |W_{a}(\mu'',\lambda)-W_{a}(\mu',\lambda)|\le W_{a}(\mu',\mu'') ∣ W a ( μ ′′ , λ ) − W a ( μ ′ , λ ) ∣ ≤ W a ( μ ′ , μ ′′ ) . (ii) If moreover W a ( μ ′ , μ ′ ′ ) ≤ 1 W_{a}(\mu',\mu'')\le1 W a ( μ ′ , μ ′′ ) ≤ 1 , then W a ( μ ′ ′ , λ ) ≤ W a ( μ ′ , λ ) + 1 W_{a}(\mu'',\lambda)\le W_{a}(\mu',\lambda)+1 W a ( μ ′′ , λ ) ≤ W a ( μ ′ , λ ) + 1 , and, writing s = W a ( μ ′ ′ , λ ) s=W_{a}(\mu'',\lambda) s = W a ( μ ′′ , λ ) and t = W a ( μ ′ , λ ) t=W_{a}(\mu',\lambda) t = W a ( μ ′ , λ ) , both nonnegative, ∣ s 2 − t 2 ∣ = ∣ s − t ∣ ( s + t ) |s^{2}-t^{2}|=|s-t|\,(s+t) ∣ s 2 − t 2 ∣ = ∣ s − t ∣ ( s + t ) , so that
∣ W a ( μ ′ ′ , λ ) 2 − W a ( μ ′ , λ ) 2 ∣ ≤ ( 2 W a ( μ ′ , λ ) + 1 ) W a ( μ ′ , μ ′ ′ ) . \bigl|W_{a}(\mu'',\lambda)^{2}-W_{a}(\mu',\lambda)^{2}\bigr|\le\bigl(2W_{a}(\mu',\lambda)+1\bigr)\,W_{a}(\mu',\mu''). W a ( μ ′′ , λ ) 2 − W a ( μ ′ , λ ) 2 ≤ ( 2 W a ( μ ′ , λ ) + 1 ) W a ( μ ′ , μ ′′ ) .
Step 2 (Claim 1, property (a)). In Steps 2 to 6, ν 0 ∈ P ρ a \nu_{0}\in\mathcal{P}^{a}_{\rho} ν 0 ∈ P ρ a and ψ ( λ ) = W a ( λ , ν 0 ) 2 \psi(\lambda)=W_{a}(\lambda,\nu_{0})^{2} ψ ( λ ) = W a ( λ , ν 0 ) 2 are as in claim 1. Let μ ∈ P ρ a \mu\in\mathcal{P}^{a}_{\rho} μ ∈ P ρ a and let ε ∈ R \varepsilon\in\mathbb{R} ε ∈ R be positive; put δ = min { 1 , ε ( 2 W a ( μ , ν 0 ) + 1 ) − 1 } \delta=\min\{1,\varepsilon\,(2W_{a}(\mu,\nu_{0})+1)^{-1}\} δ = min { 1 , ε ( 2 W a ( μ , ν 0 ) + 1 ) − 1 } , which is positive. If μ ′ ∈ P ρ a \mu'\in\mathcal{P}^{a}_{\rho} μ ′ ∈ P ρ a satisfies W a ( μ , μ ′ ) < δ W_{a}(\mu,\mu')<\delta W a ( μ , μ ′ ) < δ , then Step 1(ii), read with λ = ν 0 \lambda=\nu_{0} λ = ν 0 , with μ \mu μ in the role of its μ ′ \mu' μ ′ and μ ′ \mu' μ ′ in the role of its μ ′ ′ \mu'' μ ′′ , gives ∣ ψ ( μ ′ ) − ψ ( μ ) ∣ ≤ ( 2 W a ( μ , ν 0 ) + 1 ) W a ( μ , μ ′ ) < ε |\psi(\mu')-\psi(\mu)|\le(2W_{a}(\mu,\nu_{0})+1)\,W_{a}(\mu,\mu')<\varepsilon ∣ ψ ( μ ′ ) − ψ ( μ ) ∣ ≤ ( 2 W a ( μ , ν 0 ) + 1 ) W a ( μ , μ ′ ) < ε . So ψ \psi ψ is continuous at μ \mu μ for the metric W a W_{a} W a of The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §metric and the absolute-value metric; as μ \mu μ was arbitrary, property (a) holds.
Step 3 (Claim 1, the upper estimate). In Steps 3 to 5 let μ ∈ Q \mu\in Q μ ∈ Q and let S S S be any noise-optimal map from μ \mu μ to ν 0 \nu_{0} ν 0 ; one exists because the pair ( μ , ν 0 ) (\mu,\nu_{0}) ( μ , ν 0 ) is uniquely noise-mapped by The Noise Map Property of a Set of Probability Measures §map-property and Noise-Optimal Maps and Uniquely Noise-Mapped Pairs §uniquely-mapped . Put η = 2 ( i d − S ) ∈ L 2 ( μ ; X a ) \eta=2(\mathrm{id}-S)\in L^{2}(\mu;X^{a}) η = 2 ( id − S ) ∈ L 2 ( μ ; X a ) . By Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §linear , read with s = − 2 s=-2 s = − 2 and t = 0 t=0 t = 0 , J a ( η , π ) = − 2 J a ( S − i d , π ) \mathcal{J}^{a}(\eta,\pi)=-2\,\mathcal{J}^{a}(S-\mathrm{id},\pi) J a ( η , π ) = − 2 J a ( S − id , π ) for every ν ∈ P ( X ) \nu\in\mathcal{P}(X) ν ∈ P ( X ) and π ∈ Π a ( μ , ν ) \pi\in\Pi^{a}(\mu,\nu) π ∈ Π a ( μ , ν ) . Also ψ ( μ ) = W a ( μ , ν 0 ) 2 = ∥ S − i d ∥ μ 2 = ∫ X n a ( S ( x ) − x ) μ ( d x ) \psi(\mu)=W_{a}(\mu,\nu_{0})^{2}=\lVert S-\mathrm{id}\rVert_{\mu}^{2}=\int_{X}n_{a}(S(x)-x)\,\mu(dx) ψ ( μ ) = W a ( μ , ν 0 ) 2 = ∥ S − id ∥ μ 2 = ∫ X n a ( S ( x ) − x ) μ ( d x ) , by The Noise-Optimal Map: Transport Cost, Stability Along Nearly Optimal Couplings and Stability Under Perturbation of the Source §cost and Noise-Optimal Maps and Uniquely Noise-Mapped Pairs §map . We show: for every ν ∈ P ρ a \nu\in\mathcal{P}^{a}_{\rho} ν ∈ P ρ a and every π ∈ Π a ( μ , ν ) \pi\in\Pi^{a}(\mu,\nu) π ∈ Π a ( μ , ν ) ,
ψ ( ν ) − ψ ( μ ) − J a ( η , π ) ≤ I a ( π ) . (3) \psi(\nu)-\psi(\mu)-\mathcal{J}^{a}(\eta,\pi)\le I^{a}(\pi).\tag{3} ψ ( ν ) − ψ ( μ ) − J a ( η , π ) ≤ I a ( π ) . ( 3 )
Let δ \delta δ be the displacement field of π \pi π , an element of L 2 ( π ; X a ) L^{2}(\pi;X^{a}) L 2 ( π ; X a ) with ∥ δ ∥ π 2 = I a ( π ) \lVert\delta\rVert_{\pi}^{2}=I^{a}(\pi) ∥ δ ∥ π 2 = I a ( π ) (Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §displacement-field ). The map p : X × X → X a p:X\times X\to X^{a} p : X × X → X a , p ( z ) = S ( x ) − x p(z)=S(x)-x p ( z ) = S ( x ) − x , is the composite of the Borel map π 1 \pi_{1} π 1 (Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §product-sigma ) with the Borel map x ↦ S ( x ) − x x\mapsto S(x)-x x ↦ S ( x ) − x , hence Borel by claim 4 of Borel Measurability and Bounded Integration on a Metric Space , and it takes values in X a X^{a} X a , so it is measurable into X a X^{a} X a by The Noise Space is a Real Hilbert Space: Orthonormal Basis, Continuous Embedding, Partial Sums, Closed Balls and Borel Measurability §measurable . Its square norm ∣ p ( z ) ∣ a 2 = n a ( S ( x ) − x ) |p(z)|_{a}^{2}=n_{a}(S(x)-x) ∣ p ( z ) ∣ a 2 = n a ( S ( x ) − x ) has integral ∫ X n a ( S ( x ) − x ) μ ( d x ) = ψ ( μ ) \int_{X}n_{a}(S(x)-x)\,\mu(dx)=\psi(\mu) ∫ X n a ( S ( x ) − x ) μ ( d x ) = ψ ( μ ) against π \pi π , by the change-of-variables formula of Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §pushforward with ( π 1 ) # π = μ (\pi_{1})_{\#}\pi=\mu ( π 1 ) # π = μ ; so p ∈ L 2 ( π ; X a ) p\in L^{2}(\pi;X^{a}) p ∈ L 2 ( π ; X a ) with ∥ p ∥ π 2 = ψ ( μ ) \lVert p\rVert_{\pi}^{2}=\psi(\mu) ∥ p ∥ π 2 = ψ ( μ ) . Moreover ⟨ p , δ ⟩ π = ∫ X × X ⟨ S ( x ) − x , δ ( z ) ⟩ a π ( d z ) = J a ( S − i d , π ) \langle p,\delta\rangle_{\pi}=\int_{X\times X}\langle S(x)-x,\delta(z)\rangle_{a}\,\pi(dz)=\mathcal{J}^{a}(S-\mathrm{id},\pi) ⟨ p , δ ⟩ π = ∫ X × X ⟨ S ( x ) − x , δ ( z ) ⟩ a π ( d z ) = J a ( S − id , π ) by Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §pairing , applied with the representative x ↦ S ( x ) − x x\mapsto S(x)-x x ↦ S ( x ) − x of S − i d S-\mathrm{id} S − id . By the expansion identity in L 2 ( π ; X a ) L^{2}(\pi;X^{a}) L 2 ( π ; X a ) ,
∥ δ − p ∥ π 2 = I a ( π ) − 2 J a ( S − i d , π ) + ψ ( μ ) = I a ( π ) + J a ( η , π ) + ψ ( μ ) . \lVert\delta-p\rVert_{\pi}^{2}=I^{a}(\pi)-2\,\mathcal{J}^{a}(S-\mathrm{id},\pi)+\psi(\mu)=I^{a}(\pi)+\mathcal{J}^{a}(\eta,\pi)+\psi(\mu). ∥ δ − p ∥ π 2 = I a ( π ) − 2 J a ( S − id , π ) + ψ ( μ ) = I a ( π ) + J a ( η , π ) + ψ ( μ ) .
Now let π ′ ′ = ( S ∘ π 1 , π 2 ) # π \pi''=(S\circ\pi_{1},\pi_{2})_{\#}\pi π ′′ = ( S ∘ π 1 , π 2 ) # π . By Couplings on a Hilbert Space: Product Coupling, Swap, Finiteness of the Cost, Push-Forward Couplings, Modifying One Marginal, Quantisation, Gluing over a Finitely Supported Measure, the Lipschitz Bound and the Moment Bound §modification and S # μ = ν 0 S_{\#}\mu=\nu_{0} S # μ = ν 0 (Noise-Optimal Maps and Uniquely Noise-Mapped Pairs §map ), π ′ ′ ∈ Π ( ν 0 , ν ) \pi''\in\Pi(\nu_{0},\nu) π ′′ ∈ Π ( ν 0 , ν ) . For z ∈ D a z\in D_{a} z ∈ D a one has y − S ( x ) = ( y − x ) − ( S ( x ) − x ) = δ ( z ) − p ( z ) ∈ X a y-S(x)=(y-x)-(S(x)-x)=\delta(z)-p(z)\in X^{a} y − S ( x ) = ( y − x ) − ( S ( x ) − x ) = δ ( z ) − p ( z ) ∈ X a , X a X^{a} X a being a linear subspace; so D a D_{a} D a is contained in ( S ∘ π 1 , π 2 ) − 1 ( D a ) (S\circ\pi_{1},\pi_{2})^{-1}(D_{a}) ( S ∘ π 1 , π 2 ) − 1 ( D a ) , which is Borel by Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §pairing , and π ′ ′ ( D a ) ≥ π ( D a ) = 1 \pi''(D_{a})\ge\pi(D_{a})=1 π ′′ ( D a ) ≥ π ( D a ) = 1 by Couplings of Finite Noise Cost and Their Noise Cost §finite . The function c a ∘ ( S ∘ π 1 , π 2 ) c_{a}\circ(S\circ\pi_{1},\pi_{2}) c a ∘ ( S ∘ π 1 , π 2 ) , whose value at z z z is n a ( y − S ( x ) ) n_{a}(y-S(x)) n a ( y − S ( x )) (The Noise Space is a Real Hilbert Space: Orthonormal Basis, Continuous Embedding, Partial Sums, Closed Balls and Borel Measurability §pairs ), agrees with ∣ δ − p ∣ a 2 |\delta-p|_{a}^{2} ∣ δ − p ∣ a 2 at every z ∈ D a z\in D_{a} z ∈ D a , hence π \pi π -almost everywhere; both are nonnegative and Borel (Measurable Maps into a Hilbert Space with an Orthonormal Basis: Norms, Inner Products and Linear Combinations, Synthesis from Coordinates, Square-Integrability, and Almost-Everywhere Equality §operations ), so by the change-of-variables formula of Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §pushforward and The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §comparison ,
∫ X × X c a d π ′ ′ = ∫ X × X n a ( y − S ( x ) ) π ( d z ) = ∥ δ − p ∥ π 2 < ∞ . \int_{X\times X}c_{a}\,d\pi''=\int_{X\times X}n_{a}\bigl(y-S(x)\bigr)\,\pi(dz)=\lVert\delta-p\rVert_{\pi}^{2}<\infty . ∫ X × X c a d π ′′ = ∫ X × X n a ( y − S ( x ) ) π ( d z ) = ∥ δ − p ∥ π 2 < ∞.
Thus π ′ ′ ∈ Π a ( ν 0 , ν ) \pi''\in\Pi^{a}(\nu_{0},\nu) π ′′ ∈ Π a ( ν 0 , ν ) (Couplings of Finite Noise Cost and Their Noise Cost §couplings ) with I a ( π ′ ′ ) = ∥ δ − p ∥ π 2 I^{a}(\pi'')=\lVert\delta-p\rVert_{\pi}^{2} I a ( π ′′ ) = ∥ δ − p ∥ π 2 , and by the symmetry The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §symmetry and The Noise Wasserstein Distance §distance , ψ ( ν ) = W a ( ν 0 , ν ) 2 ≤ I a ( π ′ ′ ) = I a ( π ) + J a ( η , π ) + ψ ( μ ) \psi(\nu)=W_{a}(\nu_{0},\nu)^{2}\le I^{a}(\pi'')=I^{a}(\pi)+\mathcal{J}^{a}(\eta,\pi)+\psi(\mu) ψ ( ν ) = W a ( ν 0 , ν ) 2 ≤ I a ( π ′′ ) = I a ( π ) + J a ( η , π ) + ψ ( μ ) , which is (3).
Step 4 (Claim 1, the lower estimate along a gluing). For κ ∈ Π a ( μ , ν 0 ) \kappa\in\Pi^{a}(\mu,\nu_{0}) κ ∈ Π a ( μ , ν 0 ) let e ( κ ) = ∫ X × X n a ( S ( x ) − y ) κ ( d z ) ∈ [ 0 , ∞ ] e(\kappa)=\int_{X\times X}n_{a}(S(x)-y)\,\kappa(dz)\in[0,\infty] e ( κ ) = ∫ X × X n a ( S ( x ) − y ) κ ( d z ) ∈ [ 0 , ∞ ] , the integral of the nonnegative Borel function k = c a ∘ ( π 2 , S ∘ π 1 ) k=c_{a}\circ(\pi_{2},S\circ\pi_{1}) k = c a ∘ ( π 2 , S ∘ π 1 ) , k ( z ) = n a ( S ( x ) − y ) k(z)=n_{a}(S(x)-y) k ( z ) = n a ( S ( x ) − y ) (The Noise Space is a Real Hilbert Space: Orthonormal Basis, Continuous Embedding, Partial Sums, Closed Balls and Borel Measurability §pairs , Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §pairing , claim 4 of Borel Measurability and Bounded Integration on a Metric Space ). We show: for every ν ∈ P ρ a \nu\in\mathcal{P}^{a}_{\rho} ν ∈ P ρ a and every π ∈ Π a ( μ , ν ) \pi\in\Pi^{a}(\mu,\nu) π ∈ Π a ( μ , ν ) there is κ ∈ Π a ( μ , ν 0 ) \kappa\in\Pi^{a}(\mu,\nu_{0}) κ ∈ Π a ( μ , ν 0 ) with e ( κ ) < ∞ e(\kappa)<\infty e ( κ ) < ∞ ,
I a ( κ ) ≤ 2 I a ( π ) + W a ( μ , ν 0 ) and ψ ( ν ) − ψ ( μ ) − J a ( η , π ) ≥ − 2 e ( κ ) I a ( π ) . (4) \sqrt{I^{a}(\kappa)}\le2\sqrt{I^{a}(\pi)}+W_{a}(\mu,\nu_{0})\qquad\text{and}\qquad\psi(\nu)-\psi(\mu)-\mathcal{J}^{a}(\eta,\pi)\ge-2\sqrt{e(\kappa)}\,\sqrt{I^{a}(\pi)}.\tag{4} I a ( κ ) ≤ 2 I a ( π ) + W a ( μ , ν 0 ) and ψ ( ν ) − ψ ( μ ) − J a ( η , π ) ≥ − 2 e ( κ ) I a ( π ) . ( 4 )
The pair ( ν , ν 0 ) (\nu,\nu_{0}) ( ν , ν 0 ) is noise-connected by The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §connected , so by The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §optimal there is a noise-optimal σ ∈ Π a ( ν , ν 0 ) \sigma\in\Pi^{a}(\nu,\nu_{0}) σ ∈ Π a ( ν , ν 0 ) , with I a ( σ ) = W a ( ν , ν 0 ) 2 = ψ ( ν ) I^{a}(\sigma)=W_{a}(\nu,\nu_{0})^{2}=\psi(\nu) I a ( σ ) = W a ( ν , ν 0 ) 2 = ψ ( ν ) (Noise-Optimal Couplings §optimal ). The measures μ , ν , ν 0 \mu,\nu,\nu_{0} μ , ν , ν 0 lie in P 2 ( X ) \mathcal{P}_{2}(X) P 2 ( X ) by The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §moments , and π ∈ Π ( μ , ν ) \pi\in\Pi(\mu,\nu) π ∈ Π ( μ , ν ) , σ ∈ Π ( ν , ν 0 ) \sigma\in\Pi(\nu,\nu_{0}) σ ∈ Π ( ν , ν 0 ) by Couplings of Finite Noise Cost and Their Noise Cost §couplings ; so Gluing Two Couplings on a Hilbert Space over a Common Middle Marginal, and the Triangle Inequalities for the Quadratic and Noise Costs §glued provides Σ ∈ P ( X ( 3 ) ) \Sigma\in\mathcal{P}(X_{(3)}) Σ ∈ P ( X ( 3 ) ) with ( q 1 , q 2 ) # Σ = π (q_{1},q_{2})_{\#}\Sigma=\pi ( q 1 , q 2 ) # Σ = π and ( q 2 , q 3 ) # Σ = σ (q_{2},q_{3})_{\#}\Sigma=\sigma ( q 2 , q 3 ) # Σ = σ , in the notation of that lemma. Put κ = ( q 1 , q 3 ) # Σ \kappa=(q_{1},q_{3})_{\#}\Sigma κ = ( q 1 , q 3 ) # Σ . By Gluing Two Couplings on a Hilbert Space over a Common Middle Marginal, and the Triangle Inequalities for the Quadratic and Noise Costs §noise-triangle , κ ∈ Π a ( μ , ν 0 ) \kappa\in\Pi^{a}(\mu,\nu_{0}) κ ∈ Π a ( μ , ν 0 ) and I a ( κ ) ≤ I a ( π ) + I a ( σ ) = I a ( π ) + W a ( ν , ν 0 ) \sqrt{I^{a}(\kappa)}\le\sqrt{I^{a}(\pi)}+\sqrt{I^{a}(\sigma)}=\sqrt{I^{a}(\pi)}+W_{a}(\nu,\nu_{0}) I a ( κ ) ≤ I a ( π ) + I a ( σ ) = I a ( π ) + W a ( ν , ν 0 ) ; and W a ( ν , ν 0 ) ≤ W a ( μ , ν ) + W a ( μ , ν 0 ) ≤ I a ( π ) + W a ( μ , ν 0 ) W_{a}(\nu,\nu_{0})\le W_{a}(\mu,\nu)+W_{a}(\mu,\nu_{0})\le\sqrt{I^{a}(\pi)}+W_{a}(\mu,\nu_{0}) W a ( ν , ν 0 ) ≤ W a ( μ , ν ) + W a ( μ , ν 0 ) ≤ I a ( π ) + W a ( μ , ν 0 ) by The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §triangle , The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §symmetry and (W). This is the first inequality in (4).
We work in the real Hilbert space L 2 ( Σ ; X a ) L^{2}(\Sigma;X^{a}) L 2 ( Σ ; X a ) of The Space of Square-Integrable Maps from a Measure Space into a Hilbert Space with an Orthonormal Basis §classes , read with the measure space ( X ( 3 ) , B ( X ( 3 ) ) , Σ ) (X_{(3)},\mathcal{B}(X_{(3)}),\Sigma) ( X ( 3 ) , B ( X ( 3 ) ) , Σ ) , E = X a E=X^{a} E = X a and the basis of Probability Measures on a Hilbert Space Transported in the Noise Norm: Standing Notation §noise-space , a real Hilbert space by The Space of Square-Integrable Hilbert-Valued Maps is a Real Hilbert Space: Coordinates and Synthesis §hilbert ; push-forwards by Borel maps on X ( 3 ) X_{(3)} X ( 3 ) are image measures, and integrals transform by claim 2 of Image Measures, Measures with Densities, and Change of Variables , as recorded in Gluing Two Couplings on a Hilbert Space over a Common Middle Marginal, and the Triangle Inequalities for the Quadratic and Noise Costs . Let δ π \delta_{\pi} δ π , δ σ \delta_{\sigma} δ σ and δ κ \delta_{\kappa} δ κ be the displacement fields of π \pi π , σ \sigma σ and κ \kappa κ (Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §displacement-field , read for σ \sigma σ with ν , ν 0 \nu,\nu_{0} ν , ν 0 in place of its μ , ν \mu,\nu μ , ν ), measurable maps from X × X X\times X X × X into X a X^{a} X a , and define on X ( 3 ) X_{(3)} X ( 3 )
V = δ π ∘ ( q 1 , q 2 ) , U = δ σ ∘ ( q 2 , q 3 ) , Z = δ κ ∘ ( q 1 , q 3 ) , P ( w ) = S ( q 1 ( w ) ) − q 1 ( w ) . V=\delta_{\pi}\circ(q_{1},q_{2}),\qquad U=\delta_{\sigma}\circ(q_{2},q_{3}),\qquad Z=\delta_{\kappa}\circ(q_{1},q_{3}),\qquad P(w)=S(q_{1}(w))-q_{1}(w). V = δ π ∘ ( q 1 , q 2 ) , U = δ σ ∘ ( q 2 , q 3 ) , Z = δ κ ∘ ( q 1 , q 3 ) , P ( w ) = S ( q 1 ( w )) − q 1 ( w ) .
V V V , U U U and Z Z Z are measurable into X a X^{a} X a by claim 4 of Borel Measurability and Bounded Integration on a Metric Space , the pairings being Borel (Gluing Two Couplings on a Hilbert Space over a Common Middle Marginal, and the Triangle Inequalities for the Quadratic and Noise Costs ); P P P is Borel into X X X by the same claim and takes values in X a X^{a} X a , hence is measurable into X a X^{a} X a by The Noise Space is a Real Hilbert Space: Orthonormal Basis, Continuous Embedding, Partial Sums, Closed Balls and Borel Measurability §measurable . By claim 2 of Image Measures, Measures with Densities, and Change of Variables and Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §displacement-field , ∫ ∣ V ∣ a 2 d Σ = ∫ ∣ δ π ∣ a 2 d π = I a ( π ) \int|V|_{a}^{2}\,d\Sigma=\int|\delta_{\pi}|_{a}^{2}\,d\pi=I^{a}(\pi) ∫ ∣ V ∣ a 2 d Σ = ∫ ∣ δ π ∣ a 2 d π = I a ( π ) , and likewise ∫ ∣ U ∣ a 2 d Σ = I a ( σ ) = ψ ( ν ) \int|U|_{a}^{2}\,d\Sigma=I^{a}(\sigma)=\psi(\nu) ∫ ∣ U ∣ a 2 d Σ = I a ( σ ) = ψ ( ν ) and ∫ ∣ Z ∣ a 2 d Σ = I a ( κ ) \int|Z|_{a}^{2}\,d\Sigma=I^{a}(\kappa) ∫ ∣ Z ∣ a 2 d Σ = I a ( κ ) . Since q 1 = π 1 ∘ ( q 1 , q 2 ) q_{1}=\pi_{1}\circ(q_{1},q_{2}) q 1 = π 1 ∘ ( q 1 , q 2 ) , one has ( q 1 ) # Σ = ( π 1 ) # π = μ (q_{1})_{\#}\Sigma=(\pi_{1})_{\#}\pi=\mu ( q 1 ) # Σ = ( π 1 ) # π = μ (for Borel A ⊆ X A\subseteq X A ⊆ X , Σ ( q 1 − 1 ( A ) ) = π ( π 1 − 1 ( A ) ) \Sigma(q_{1}^{-1}(A))=\pi(\pi_{1}^{-1}(A)) Σ ( q 1 − 1 ( A )) = π ( π 1 − 1 ( A )) ), so ∫ ∣ P ∣ a 2 d Σ = ∫ X n a ( S ( x ) − x ) μ ( d x ) = ψ ( μ ) \int|P|_{a}^{2}\,d\Sigma=\int_{X}n_{a}(S(x)-x)\,\mu(dx)=\psi(\mu) ∫ ∣ P ∣ a 2 d Σ = ∫ X n a ( S ( x ) − x ) μ ( d x ) = ψ ( μ ) . Hence V , U , Z , P ∈ L 2 ( Σ ; X a ) V,U,Z,P\in L^{2}(\Sigma;X^{a}) V , U , Z , P ∈ L 2 ( Σ ; X a ) with
∥ V ∥ Σ 2 = I a ( π ) , ∥ U ∥ Σ 2 = ψ ( ν ) , ∥ Z ∥ Σ 2 = I a ( κ ) , ∥ P ∥ Σ 2 = ψ ( μ ) . \lVert V\rVert_{\Sigma}^{2}=I^{a}(\pi),\qquad\lVert U\rVert_{\Sigma}^{2}=\psi(\nu),\qquad\lVert Z\rVert_{\Sigma}^{2}=I^{a}(\kappa),\qquad\lVert P\rVert_{\Sigma}^{2}=\psi(\mu). ∥ V ∥ Σ 2 = I a ( π ) , ∥ U ∥ Σ 2 = ψ ( ν ) , ∥ Z ∥ Σ 2 = I a ( κ ) , ∥ P ∥ Σ 2 = ψ ( μ ) .
The integrand of ⟨ P , V ⟩ Σ \langle P,V\rangle_{\Sigma} ⟨ P , V ⟩ Σ is h ∘ ( q 1 , q 2 ) h\circ(q_{1},q_{2}) h ∘ ( q 1 , q 2 ) with h ( z ) = ⟨ S ( x ) − x , δ π ( z ) ⟩ a h(z)=\langle S(x)-x,\delta_{\pi}(z)\rangle_{a} h ( z ) = ⟨ S ( x ) − x , δ π ( z ) ⟩ a , which is Borel and π \pi π -integrable with integral J a ( S − i d , π ) \mathcal{J}^{a}(S-\mathrm{id},\pi) J a ( S − id , π ) by Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §pairing ; so claim 2 of Image Measures, Measures with Densities, and Change of Variables gives ⟨ P , V ⟩ Σ = J a ( S − i d , π ) \langle P,V\rangle_{\Sigma}=\mathcal{J}^{a}(S-\mathrm{id},\pi) ⟨ P , V ⟩ Σ = J a ( S − id , π ) .
Let G = ( q 1 , q 2 ) − 1 ( D a ) ∩ ( q 2 , q 3 ) − 1 ( D a ) G=(q_{1},q_{2})^{-1}(D_{a})\cap(q_{2},q_{3})^{-1}(D_{a}) G = ( q 1 , q 2 ) − 1 ( D a ) ∩ ( q 2 , q 3 ) − 1 ( D a ) , a Borel set whose complement is the union of two sets of Σ \Sigma Σ -measure π ( X × X ∖ D a ) = 0 \pi(X\times X\setminus D_{a})=0 π ( X × X ∖ D a ) = 0 and σ ( X × X ∖ D a ) = 0 \sigma(X\times X\setminus D_{a})=0 σ ( X × X ∖ D a ) = 0 (Couplings of Finite Noise Cost and Their Noise Cost §finite ), hence Σ \Sigma Σ -null by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §null-union . For w ∈ G w\in G w ∈ G , writing q i = q i ( w ) q_{i}=q_{i}(w) q i = q i ( w ) , one has V ( w ) = q 2 − q 1 V(w)=q_{2}-q_{1} V ( w ) = q 2 − q 1 and U ( w ) = q 3 − q 2 U(w)=q_{3}-q_{2} U ( w ) = q 3 − q 2 , so q 3 − q 1 = U ( w ) + V ( w ) ∈ X a q_{3}-q_{1}=U(w)+V(w)\in X^{a} q 3 − q 1 = U ( w ) + V ( w ) ∈ X a , whence ( q 1 , q 3 ) ( w ) ∈ D a (q_{1},q_{3})(w)\in D_{a} ( q 1 , q 3 ) ( w ) ∈ D a and Z ( w ) = q 3 − q 1 = U ( w ) + V ( w ) Z(w)=q_{3}-q_{1}=U(w)+V(w) Z ( w ) = q 3 − q 1 = U ( w ) + V ( w ) . Thus U + V ∼ Σ Z U+V\sim_{\Sigma}Z U + V ∼ Σ Z in the sense of Measurable Maps into a Hilbert Space with an Orthonormal Basis: Norms, Inner Products and Linear Combinations, Synthesis from Coordinates, Square-Integrability, and Almost-Everywhere Equality §almost-everywhere , and U + V = Z U+V=Z U + V = Z in L 2 ( Σ ; X a ) L^{2}(\Sigma;X^{a}) L 2 ( Σ ; X a ) by The Space of Square-Integrable Maps from a Measure Space into a Hilbert Space with an Orthonormal Basis §classes ; in particular ∥ U + V ∥ Σ 2 = I a ( κ ) \lVert U+V\rVert_{\Sigma}^{2}=I^{a}(\kappa) ∥ U + V ∥ Σ 2 = I a ( κ ) . Put F = U + V − P ∈ L 2 ( Σ ; X a ) F=U+V-P\in L^{2}(\Sigma;X^{a}) F = U + V − P ∈ L 2 ( Σ ; X a ) . For w ∈ G w\in G w ∈ G , F ( w ) = q 3 − S ( q 1 ) ∈ X a F(w)=q_{3}-S(q_{1})\in X^{a} F ( w ) = q 3 − S ( q 1 ) ∈ X a , so ∣ F ( w ) ∣ a 2 = ∣ S ( q 1 ) − q 3 ∣ a 2 = k ( ( q 1 , q 3 ) ( w ) ) |F(w)|_{a}^{2}=|S(q_{1})-q_{3}|_{a}^{2}=k((q_{1},q_{3})(w)) ∣ F ( w ) ∣ a 2 = ∣ S ( q 1 ) − q 3 ∣ a 2 = k (( q 1 , q 3 ) ( w )) by the homogeneity of ∣ ⋅ ∣ a |\cdot|_{a} ∣ ⋅ ∣ a (Elementary Identities in a Real Inner Product Space §homogeneity ). By The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §comparison and claim 2 of Image Measures, Measures with Densities, and Change of Variables ,
∥ F ∥ Σ 2 = ∫ X ( 3 ) k ∘ ( q 1 , q 3 ) d Σ = ∫ X × X k d κ = e ( κ ) , \lVert F\rVert_{\Sigma}^{2}=\int_{X_{(3)}}k\circ(q_{1},q_{3})\,d\Sigma=\int_{X\times X}k\,d\kappa=e(\kappa), ∥ F ∥ Σ 2 = ∫ X ( 3 ) k ∘ ( q 1 , q 3 ) d Σ = ∫ X × X k d κ = e ( κ ) ,
so e ( κ ) < ∞ e(\kappa)<\infty e ( κ ) < ∞ and ∥ F ∥ Σ = e ( κ ) \lVert F\rVert_{\Sigma}=\sqrt{e(\kappa)} ∥ F ∥ Σ = e ( κ ) . Now U = ( U + V ) − V U=(U+V)-V U = ( U + V ) − V and U + V = F + P U+V=F+P U + V = F + P , so the expansion identity and bilinearity give
ψ ( ν ) = ∥ U ∥ Σ 2 = ∥ U + V ∥ Σ 2 − 2 ⟨ F , V ⟩ Σ − 2 ⟨ P , V ⟩ Σ + ∥ V ∥ Σ 2 = I a ( κ ) − 2 ⟨ F , V ⟩ Σ + J a ( η , π ) + I a ( π ) , \psi(\nu)=\lVert U\rVert_{\Sigma}^{2}=\lVert U+V\rVert_{\Sigma}^{2}-2\langle F,V\rangle_{\Sigma}-2\langle P,V\rangle_{\Sigma}+\lVert V\rVert_{\Sigma}^{2}=I^{a}(\kappa)-2\langle F,V\rangle_{\Sigma}+\mathcal{J}^{a}(\eta,\pi)+I^{a}(\pi), ψ ( ν ) = ∥ U ∥ Σ 2 = ∥ U + V ∥ Σ 2 − 2 ⟨ F , V ⟩ Σ − 2 ⟨ P , V ⟩ Σ + ∥ V ∥ Σ 2 = I a ( κ ) − 2 ⟨ F , V ⟩ Σ + J a ( η , π ) + I a ( π ) ,
using − 2 ⟨ P , V ⟩ Σ = − 2 J a ( S − i d , π ) = J a ( η , π ) -2\langle P,V\rangle_{\Sigma}=-2\mathcal{J}^{a}(S-\mathrm{id},\pi)=\mathcal{J}^{a}(\eta,\pi) − 2 ⟨ P , V ⟩ Σ = − 2 J a ( S − id , π ) = J a ( η , π ) (Step 3). Hence
ψ ( ν ) − ψ ( μ ) − J a ( η , π ) = ( I a ( κ ) − ψ ( μ ) ) − 2 ⟨ F , V ⟩ Σ + I a ( π ) . \psi(\nu)-\psi(\mu)-\mathcal{J}^{a}(\eta,\pi)=\bigl(I^{a}(\kappa)-\psi(\mu)\bigr)-2\langle F,V\rangle_{\Sigma}+I^{a}(\pi). ψ ( ν ) − ψ ( μ ) − J a ( η , π ) = ( I a ( κ ) − ψ ( μ ) ) − 2 ⟨ F , V ⟩ Σ + I a ( π ) .
The first bracket is nonnegative because ψ ( μ ) = W a ( μ , ν 0 ) 2 ≤ I a ( κ ) \psi(\mu)=W_{a}(\mu,\nu_{0})^{2}\le I^{a}(\kappa) ψ ( μ ) = W a ( μ , ν 0 ) 2 ≤ I a ( κ ) (The Noise Wasserstein Distance §distance ), I a ( π ) ≥ 0 I^{a}(\pi)\ge0 I a ( π ) ≥ 0 , and ∣ ⟨ F , V ⟩ Σ ∣ ≤ ∥ F ∥ Σ ∥ V ∥ Σ = e ( κ ) I a ( π ) |\langle F,V\rangle_{\Sigma}|\le\lVert F\rVert_{\Sigma}\lVert V\rVert_{\Sigma}=\sqrt{e(\kappa)}\sqrt{I^{a}(\pi)} ∣ ⟨ F , V ⟩ Σ ∣ ≤ ∥ F ∥ Σ ∥ V ∥ Σ = e ( κ ) I a ( π ) by the Cauchy--Schwarz inequality. This is the second inequality in (4).
Step 5 (Claim 1, property (b) and the gradient). We show that ψ \psi ψ is differentiable along noise couplings at μ \mu μ with gradient η \eta η . Let ε ∈ R \varepsilon\in\mathbb{R} ε ∈ R be positive. The order of choice is: first a positive θ − \theta_{-} θ − for the lower estimate, depending on ε \varepsilon ε , then θ = min { ε , θ − } \theta=\min\{\varepsilon,\theta_{-}\} θ = min { ε , θ − } .
Lower estimate. We claim there is a positive θ − \theta_{-} θ − such that ψ ( ν ) − ψ ( μ ) − J a ( η , π ) ≥ − ε I a ( π ) \psi(\nu)-\psi(\mu)-\mathcal{J}^{a}(\eta,\pi)\ge-\varepsilon\sqrt{I^{a}(\pi)} ψ ( ν ) − ψ ( μ ) − J a ( η , π ) ≥ − ε I a ( π ) for all ν ∈ P ρ a \nu\in\mathcal{P}^{a}_{\rho} ν ∈ P ρ a and π ∈ Π a ( μ , ν ) \pi\in\Pi^{a}(\mu,\nu) π ∈ Π a ( μ , ν ) with I a ( π ) < θ − 2 I^{a}(\pi)<\theta_{-}^{2} I a ( π ) < θ − 2 . Suppose not. Then, taking θ − = j − 1 \theta_{-}=j^{-1} θ − = j − 1 for j ∈ N j\in\mathbb{N} j ∈ N , there are ν j ∈ P ρ a \nu_{j}\in\mathcal{P}^{a}_{\rho} ν j ∈ P ρ a and π j ∈ Π a ( μ , ν j ) \pi_{j}\in\Pi^{a}(\mu,\nu_{j}) π j ∈ Π a ( μ , ν j ) with I a ( π j ) < j − 2 I^{a}(\pi_{j})<j^{-2} I a ( π j ) < j − 2 and ψ ( ν j ) − ψ ( μ ) − J a ( η , π j ) < − ε I a ( π j ) \psi(\nu_{j})-\psi(\mu)-\mathcal{J}^{a}(\eta,\pi_{j})<-\varepsilon\sqrt{I^{a}(\pi_{j})} ψ ( ν j ) − ψ ( μ ) − J a ( η , π j ) < − ε I a ( π j ) . For each j j j let κ j ∈ Π a ( μ , ν 0 ) \kappa_{j}\in\Pi^{a}(\mu,\nu_{0}) κ j ∈ Π a ( μ , ν 0 ) be as in Step 4 for ν j \nu_{j} ν j and π j \pi_{j} π j , and write e j = e ( κ j ) e_{j}=e(\kappa_{j}) e j = e ( κ j ) . Then − 2 e j I a ( π j ) < − ε I a ( π j ) -2\sqrt{e_{j}}\sqrt{I^{a}(\pi_{j})}<-\varepsilon\sqrt{I^{a}(\pi_{j})} − 2 e j I a ( π j ) < − ε I a ( π j ) , that is ( 2 e j − ε ) I a ( π j ) > 0 (2\sqrt{e_{j}}-\varepsilon)\sqrt{I^{a}(\pi_{j})}>0 ( 2 e j − ε ) I a ( π j ) > 0 ; so I a ( π j ) > 0 \sqrt{I^{a}(\pi_{j})}>0 I a ( π j ) > 0 and 2 e j > ε 2\sqrt{e_{j}}>\varepsilon 2 e j > ε , whence e j > ε 2 / 4 e_{j}>\varepsilon^{2}/4 e j > ε 2 /4 for every j j j . On the other hand W a ( μ , ν 0 ) ≤ I a ( κ j ) W_{a}(\mu,\nu_{0})\le\sqrt{I^{a}(\kappa_{j})} W a ( μ , ν 0 ) ≤ I a ( κ j ) by (W), and I a ( κ j ) ≤ 2 I a ( π j ) + W a ( μ , ν 0 ) < 2 j − 1 + W a ( μ , ν 0 ) \sqrt{I^{a}(\kappa_{j})}\le2\sqrt{I^{a}(\pi_{j})}+W_{a}(\mu,\nu_{0})<2j^{-1}+W_{a}(\mu,\nu_{0}) I a ( κ j ) ≤ 2 I a ( π j ) + W a ( μ , ν 0 ) < 2 j − 1 + W a ( μ , ν 0 ) by (4); by the squeeze principle ( I a ( κ j ) ) j (\sqrt{I^{a}(\kappa_{j})})_{j} ( I a ( κ j ) ) j converges to W a ( μ , ν 0 ) W_{a}(\mu,\nu_{0}) W a ( μ , ν 0 ) , and so ( I a ( κ j ) ) j (I^{a}(\kappa_{j}))_{j} ( I a ( κ j ) ) j converges to W a ( μ , ν 0 ) 2 W_{a}(\mu,\nu_{0})^{2} W a ( μ , ν 0 ) 2 . The pair ( μ , ν 0 ) (\mu,\nu_{0}) ( μ , ν 0 ) is uniquely noise-mapped and S S S is a noise-optimal map from μ \mu μ to ν 0 \nu_{0} ν 0 , so The Noise-Optimal Map: Transport Cost, Stability Along Nearly Optimal Couplings and Stability Under Perturbation of the Source §stability , applied to the sequence ( κ j ) j ∈ N (\kappa_{j})_{j\in\mathbb{N}} ( κ j ) j ∈ N in Π a ( μ , ν 0 ) \Pi^{a}(\mu,\nu_{0}) Π a ( μ , ν 0 ) , shows that ( e j ) j (e_{j})_{j} ( e j ) j converges to 0 0 0 . This contradicts e j > ε 2 / 4 > 0 e_{j}>\varepsilon^{2}/4>0 e j > ε 2 /4 > 0 for every j j j , and proves the claim.
Conclusion. Let θ = min { ε , θ − } \theta=\min\{\varepsilon,\theta_{-}\} θ = min { ε , θ − } , which is positive, and let ν ∈ P ρ a \nu\in\mathcal{P}^{a}_{\rho} ν ∈ P ρ a and π ∈ Π a ( μ , ν ) \pi\in\Pi^{a}(\mu,\nu) π ∈ Π a ( μ , ν ) satisfy I a ( π ) < θ 2 I^{a}(\pi)<\theta^{2} I a ( π ) < θ 2 . Then I a ( π ) < θ − 2 I^{a}(\pi)<\theta_{-}^{2} I a ( π ) < θ − 2 , so the lower estimate holds; and I a ( π ) < θ ≤ ε \sqrt{I^{a}(\pi)}<\theta\le\varepsilon I a ( π ) < θ ≤ ε , so I a ( π ) = I a ( π ) I a ( π ) ≤ ε I a ( π ) I^{a}(\pi)=\sqrt{I^{a}(\pi)}\sqrt{I^{a}(\pi)}\le\varepsilon\sqrt{I^{a}(\pi)} I a ( π ) = I a ( π ) I a ( π ) ≤ ε I a ( π ) and (3) gives the upper estimate. Hence ∣ ψ ( ν ) − ψ ( μ ) − J a ( η , π ) ∣ ≤ ε I a ( π ) |\psi(\nu)-\psi(\mu)-\mathcal{J}^{a}(\eta,\pi)|\le\varepsilon\sqrt{I^{a}(\pi)} ∣ ψ ( ν ) − ψ ( μ ) − J a ( η , π ) ∣ ≤ ε I a ( π ) , so ψ \psi ψ is differentiable along noise couplings at μ \mu μ with gradient η \eta η , and ∇ ψ ( μ ) = η = 2 ( i d − S ) \nabla\psi(\mu)=\eta=2(\mathrm{id}-S) ∇ ψ ( μ ) = η = 2 ( id − S ) by Differentiability of a Function on the Noise-Connected Measures Along Noise Couplings, and Its Gradient §gradient . Since S S S was an arbitrary noise-optimal map from μ \mu μ to ν 0 \nu_{0} ν 0 , this is the gradient formula of claim 1. Moreover S − i d ∈ T μ a S-\mathrm{id}\in T^{a}_{\mu} S − id ∈ T μ a by The Noise Map Property of a Set of Probability Measures §map-property , and T μ a T^{a}_{\mu} T μ a is a linear subspace of L 2 ( μ ; X a ) L^{2}(\mu;X^{a}) L 2 ( μ ; X a ) by Linearity of the Noise Gradient, and the Noise Tangent Space is a Closed Linear Subspace §subspace , so ∇ ψ ( μ ) = ( − 2 ) ( S − i d ) ∈ T μ a \nabla\psi(\mu)=(-2)(S-\mathrm{id})\in T^{a}_{\mu} ∇ ψ ( μ ) = ( − 2 ) ( S − id ) ∈ T μ a . As μ ∈ Q \mu\in Q μ ∈ Q was arbitrary, property (b) holds.
Step 6 (Claim 1, property (c)). Let μ ∈ Q \mu\in Q μ ∈ Q , let ( μ n ) n ∈ N (\mu_{n})_{n\in\mathbb{N}} ( μ n ) n ∈ N be a sequence in Q Q Q and let ( π n ) n ∈ N (\pi_{n})_{n\in\mathbb{N}} ( π n ) n ∈ N be a sequence of couplings of vanishing noise cost from ( μ n ) (\mu_{n}) ( μ n ) to μ \mu μ . Let S S S be a noise-optimal map from μ \mu μ to ν 0 \nu_{0} ν 0 and, for each n n n , S n S_{n} S n one from μ n \mu_{n} μ n to ν 0 \nu_{0} ν 0 , which exist as in Step 3; all pairs ( μ , ν 0 ) (\mu,\nu_{0}) ( μ , ν 0 ) and ( μ n , ν 0 ) (\mu_{n},\nu_{0}) ( μ n , ν 0 ) are uniquely noise-mapped by The Noise Map Property of a Set of Probability Measures §map-property . By Step 5, ∇ ψ ( μ n ) = 2 ( i d − S n ) \nabla\psi(\mu_{n})=2(\mathrm{id}-S_{n}) ∇ ψ ( μ n ) = 2 ( id − S n ) and ∇ ψ ( μ ) = 2 ( i d − S ) \nabla\psi(\mu)=2(\mathrm{id}-S) ∇ ψ ( μ ) = 2 ( id − S ) , represented by x ↦ ( − 2 ) ( S n ( x ) − x ) x\mapsto(-2)(S_{n}(x)-x) x ↦ ( − 2 ) ( S n ( x ) − x ) and y ↦ ( − 2 ) ( S ( y ) − y ) y\mapsto(-2)(S(y)-y) y ↦ ( − 2 ) ( S ( y ) − y ) (opening paragraph). For every z z z , by homogeneity in X a X^{a} X a (Elementary Identities in a Real Inner Product Space §homogeneity ),
∣ ( − 2 ) ( S n ( x ) − x ) − ( − 2 ) ( S ( y ) − y ) ∣ a 2 = 4 ∣ ( S n ( x ) − x ) − ( S ( y ) − y ) ∣ a 2 . \bigl|(-2)(S_{n}(x)-x)-(-2)(S(y)-y)\bigr|_{a}^{2}=4\,\bigl|(S_{n}(x)-x)-(S(y)-y)\bigr|_{a}^{2}. ( − 2 ) ( S n ( x ) − x ) − ( − 2 ) ( S ( y ) − y ) a 2 = 4 ( S n ( x ) − x ) − ( S ( y ) − y ) a 2 .
Both sides are nonnegative Borel functions of z z z (Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §discrepancy ), so by Linearity and Monotonicity of the Lebesgue Integral §nonnegative the discrepancy of ∇ ψ ( μ n ) \nabla\psi(\mu_{n}) ∇ ψ ( μ n ) and ∇ ψ ( μ ) \nabla\psi(\mu) ∇ ψ ( μ ) along π n \pi_{n} π n is 4 4 4 times that of S n − i d S_{n}-\mathrm{id} S n − id and S − i d S-\mathrm{id} S − id , the discrepancies not depending on the representatives (Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §discrepancy ). The latter converges to 0 0 0 by The Noise-Optimal Map: Transport Cost, Stability Along Nearly Optimal Couplings and Stability Under Perturbation of the Source §source , read with ν 0 \nu_{0} ν 0 in place of its ν \nu ν and ( π n ) (\pi_{n}) ( π n ) in place of its ( γ n ) (\gamma_{n}) ( γ n ) , and Strong and Weak Convergence of Noise Fields Along Couplings of Vanishing Noise Cost §strong ; so the former converges to 0 0 0 , which is property (c) . With Steps 2 and 5, ψ \psi ψ is a noise intrinsic test function on Q Q Q , which proves claim 1.
Step 7 (Claim 5). Let φ 1 , φ 2 \varphi_{1},\varphi_{2} φ 1 , φ 2 and s , t s,t s , t be as in claim 5, put φ = s φ 1 + t φ 2 \varphi=s\varphi_{1}+t\varphi_{2} φ = s φ 1 + t φ 2 and c = ∣ s ∣ + ∣ t ∣ + 1 c=|s|+|t|+1 c = ∣ s ∣ + ∣ t ∣ + 1 , which is positive, with ( ∣ s ∣ + ∣ t ∣ ) c − 1 ≤ 1 (|s|+|t|)\,c^{-1}\le1 ( ∣ s ∣ + ∣ t ∣ ) c − 1 ≤ 1 . We verify the three properties of Noise Intrinsic Test Functions on the Noise Wasserstein Space §test for φ \varphi φ on Q Q Q .
(a) The functions φ 1 \varphi_{1} φ 1 and φ 2 \varphi_{2} φ 2 are continuous on P ρ a \mathcal{P}^{a}_{\rho} P ρ a by property (a); by claim 5 of Continuity of Sums and Products of Real-Valued Functions on a Metric Space , read with the metric space ( P ρ a , W a ) (\mathcal{P}^{a}_{\rho},W_{a}) ( P ρ a , W a ) of The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §metric and A = P ρ a A=\mathcal{P}^{a}_{\rho} A = P ρ a , so are s φ 1 s\varphi_{1} s φ 1 , t φ 2 t\varphi_{2} t φ 2 and their sum φ \varphi φ .
(b) Let μ ∈ Q \mu\in Q μ ∈ Q and put ξ = s ∇ φ 1 ( μ ) + t ∇ φ 2 ( μ ) ∈ L 2 ( μ ; X a ) \xi=s\,\nabla\varphi_{1}(\mu)+t\,\nabla\varphi_{2}(\mu)\in L^{2}(\mu;X^{a}) ξ = s ∇ φ 1 ( μ ) + t ∇ φ 2 ( μ ) ∈ L 2 ( μ ; X a ) . Let ε ∈ R \varepsilon\in\mathbb{R} ε ∈ R be positive. The order of choice is: first positive θ 1 \theta_{1} θ 1 and θ 2 \theta_{2} θ 2 as in Differentiability of a Function on the Noise-Connected Measures Along Noise Couplings, and Its Gradient §differentiable for φ 1 \varphi_{1} φ 1 , ∇ φ 1 ( μ ) \nabla\varphi_{1}(\mu) ∇ φ 1 ( μ ) and for φ 2 \varphi_{2} φ 2 , ∇ φ 2 ( μ ) \nabla\varphi_{2}(\mu) ∇ φ 2 ( μ ) , both with ε c − 1 \varepsilon c^{-1} ε c − 1 in place of ε \varepsilon ε (property (b), the gradient ∇ φ i ( μ ) \nabla\varphi_{i}(\mu) ∇ φ i ( μ ) being the unique field of Differentiability of a Function on the Noise-Connected Measures Along Noise Couplings, and Its Gradient §gradient , which has the property of clause 1 of that definition); then θ = min { θ 1 , θ 2 } \theta=\min\{\theta_{1},\theta_{2}\} θ = min { θ 1 , θ 2 } , which is positive. Let ν ∈ P ρ a \nu\in\mathcal{P}^{a}_{\rho} ν ∈ P ρ a and π ∈ Π a ( μ , ν ) \pi\in\Pi^{a}(\mu,\nu) π ∈ Π a ( μ , ν ) satisfy I a ( π ) < θ 2 I^{a}(\pi)<\theta^{2} I a ( π ) < θ 2 ; then I a ( π ) < θ 1 2 I^{a}(\pi)<\theta_{1}^{2} I a ( π ) < θ 1 2 and I a ( π ) < θ 2 2 I^{a}(\pi)<\theta_{2}^{2} I a ( π ) < θ 2 2 . Put Y 1 = φ 1 ( ν ) − φ 1 ( μ ) − J a ( ∇ φ 1 ( μ ) , π ) Y_{1}=\varphi_{1}(\nu)-\varphi_{1}(\mu)-\mathcal{J}^{a}(\nabla\varphi_{1}(\mu),\pi) Y 1 = φ 1 ( ν ) − φ 1 ( μ ) − J a ( ∇ φ 1 ( μ ) , π ) and Y 2 = φ 2 ( ν ) − φ 2 ( μ ) − J a ( ∇ φ 2 ( μ ) , π ) Y_{2}=\varphi_{2}(\nu)-\varphi_{2}(\mu)-\mathcal{J}^{a}(\nabla\varphi_{2}(\mu),\pi) Y 2 = φ 2 ( ν ) − φ 2 ( μ ) − J a ( ∇ φ 2 ( μ ) , π ) , so that ∣ Y 1 ∣ ≤ ε c − 1 I a ( π ) |Y_{1}|\le\varepsilon c^{-1}\sqrt{I^{a}(\pi)} ∣ Y 1 ∣ ≤ ε c − 1 I a ( π ) and ∣ Y 2 ∣ ≤ ε c − 1 I a ( π ) |Y_{2}|\le\varepsilon c^{-1}\sqrt{I^{a}(\pi)} ∣ Y 2 ∣ ≤ ε c − 1 I a ( π ) . By Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §linear , J a ( ξ , π ) = s J a ( ∇ φ 1 ( μ ) , π ) + t J a ( ∇ φ 2 ( μ ) , π ) \mathcal{J}^{a}(\xi,\pi)=s\,\mathcal{J}^{a}(\nabla\varphi_{1}(\mu),\pi)+t\,\mathcal{J}^{a}(\nabla\varphi_{2}(\mu),\pi) J a ( ξ , π ) = s J a ( ∇ φ 1 ( μ ) , π ) + t J a ( ∇ φ 2 ( μ ) , π ) , so φ ( ν ) − φ ( μ ) − J a ( ξ , π ) = s Y 1 + t Y 2 \varphi(\nu)-\varphi(\mu)-\mathcal{J}^{a}(\xi,\pi)=sY_{1}+tY_{2} φ ( ν ) − φ ( μ ) − J a ( ξ , π ) = s Y 1 + t Y 2 and
∣ s Y 1 + t Y 2 ∣ ≤ ∣ s ∣ ∣ Y 1 ∣ + ∣ t ∣ ∣ Y 2 ∣ ≤ ( ∣ s ∣ + ∣ t ∣ ) ε c − 1 I a ( π ) ≤ ε I a ( π ) . |sY_{1}+tY_{2}|\le|s|\,|Y_{1}|+|t|\,|Y_{2}|\le(|s|+|t|)\,\varepsilon c^{-1}\sqrt{I^{a}(\pi)}\le\varepsilon\sqrt{I^{a}(\pi)} . ∣ s Y 1 + t Y 2 ∣ ≤ ∣ s ∣ ∣ Y 1 ∣ + ∣ t ∣ ∣ Y 2 ∣ ≤ ( ∣ s ∣ + ∣ t ∣ ) ε c − 1 I a ( π ) ≤ ε I a ( π ) .
Hence φ \varphi φ is differentiable along noise couplings at μ \mu μ with gradient ξ \xi ξ , and ∇ φ ( μ ) = ξ \nabla\varphi(\mu)=\xi ∇ φ ( μ ) = ξ by Differentiability of a Function on the Noise-Connected Measures Along Noise Couplings, and Its Gradient §gradient ; this is the gradient formula of claim 5. The fields ∇ φ 1 ( μ ) \nabla\varphi_{1}(\mu) ∇ φ 1 ( μ ) and ∇ φ 2 ( μ ) \nabla\varphi_{2}(\mu) ∇ φ 2 ( μ ) lie in T μ a T^{a}_{\mu} T μ a by property (b), and T μ a T^{a}_{\mu} T μ a is a linear subspace of L 2 ( μ ; X a ) L^{2}(\mu;X^{a}) L 2 ( μ ; X a ) by Linearity of the Noise Gradient, and the Noise Tangent Space is a Closed Linear Subspace §subspace , so ξ ∈ T μ a \xi\in T^{a}_{\mu} ξ ∈ T μ a .
(c) Let μ ∈ Q \mu\in Q μ ∈ Q , let ( μ n ) n ∈ N (\mu_{n})_{n\in\mathbb{N}} ( μ n ) n ∈ N be a sequence in Q Q Q and let ( π n ) n ∈ N (\pi_{n})_{n\in\mathbb{N}} ( π n ) n ∈ N be a sequence of couplings of vanishing noise cost from ( μ n ) (\mu_{n}) ( μ n ) to μ \mu μ . For n ∈ N n\in\mathbb{N} n ∈ N let A n A_{n} A n , B n B_{n} B n and C n C_{n} C n be the discrepancies along π n \pi_{n} π n (Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §discrepancy ) of ∇ φ 1 ( μ n ) \nabla\varphi_{1}(\mu_{n}) ∇ φ 1 ( μ n ) and ∇ φ 1 ( μ ) \nabla\varphi_{1}(\mu) ∇ φ 1 ( μ ) , of ∇ φ 2 ( μ n ) \nabla\varphi_{2}(\mu_{n}) ∇ φ 2 ( μ n ) and ∇ φ 2 ( μ ) \nabla\varphi_{2}(\mu) ∇ φ 2 ( μ ) , and of ∇ φ ( μ n ) \nabla\varphi(\mu_{n}) ∇ φ ( μ n ) and ∇ φ ( μ ) \nabla\varphi(\mu) ∇ φ ( μ ) . By property (c) for φ 1 \varphi_{1} φ 1 and φ 2 \varphi_{2} φ 2 and Strong and Weak Convergence of Noise Fields Along Couplings of Vanishing Noise Cost §strong , ( A n ) (A_{n}) ( A n ) and ( B n ) (B_{n}) ( B n ) converge to 0 0 0 . Fix representatives f n , f , g n , g f_{n},f,g_{n},g f n , f , g n , g of ∇ φ 1 ( μ n ) , ∇ φ 1 ( μ ) , ∇ φ 2 ( μ n ) , ∇ φ 2 ( μ ) \nabla\varphi_{1}(\mu_{n}),\nabla\varphi_{1}(\mu),\nabla\varphi_{2}(\mu_{n}),\nabla\varphi_{2}(\mu) ∇ φ 1 ( μ n ) , ∇ φ 1 ( μ ) , ∇ φ 2 ( μ n ) , ∇ φ 2 ( μ ) ; by (b) and The Space of Square-Integrable Maps from a Measure Space into a Hilbert Space with an Orthonormal Basis §operations , s f n + t g n sf_{n}+tg_{n} s f n + t g n and s f + t g sf+tg s f + t g , formed pointwise, represent ∇ φ ( μ n ) \nabla\varphi(\mu_{n}) ∇ φ ( μ n ) and ∇ φ ( μ ) \nabla\varphi(\mu) ∇ φ ( μ ) , and discrepancies do not depend on the representatives (Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §discrepancy ). For z ∈ X × X z\in X\times X z ∈ X × X put u = f n ( x ) − f ( y ) u=f_{n}(x)-f(y) u = f n ( x ) − f ( y ) and v = g n ( x ) − g ( y ) v=g_{n}(x)-g(y) v = g n ( x ) − g ( y ) , elements of X a X^{a} X a ; then the triangle inequality and homogeneity in X a X^{a} X a , with ( p + q ) 2 ≤ 2 p 2 + 2 q 2 (p+q)^{2}\le2p^{2}+2q^{2} ( p + q ) 2 ≤ 2 p 2 + 2 q 2 for reals, give
∣ ( s f n ( x ) + t g n ( x ) ) − ( s f ( y ) + t g ( y ) ) ∣ a 2 = ∣ s u + t v ∣ a 2 ≤ ( ∣ s ∣ ∣ u ∣ a + ∣ t ∣ ∣ v ∣ a ) 2 ≤ 2 s 2 ∣ u ∣ a 2 + 2 t 2 ∣ v ∣ a 2 . \bigl|\bigl(sf_{n}(x)+tg_{n}(x)\bigr)-\bigl(sf(y)+tg(y)\bigr)\bigr|_{a}^{2}=|su+tv|_{a}^{2}\le\bigl(|s|\,|u|_{a}+|t|\,|v|_{a}\bigr)^{2}\le2s^{2}|u|_{a}^{2}+2t^{2}|v|_{a}^{2}. ( s f n ( x ) + t g n ( x ) ) − ( s f ( y ) + t g ( y ) ) a 2 = ∣ s u + t v ∣ a 2 ≤ ( ∣ s ∣ ∣ u ∣ a + ∣ t ∣ ∣ v ∣ a ) 2 ≤ 2 s 2 ∣ u ∣ a 2 + 2 t 2 ∣ v ∣ a 2 .
All three functions of z z z are the nonnegative Borel integrands of the discrepancies above, so Linearity and Monotonicity of the Lebesgue Integral §nonnegative gives 0 ≤ C n ≤ 2 s 2 A n + 2 t 2 B n 0\le C_{n}\le2s^{2}A_{n}+2t^{2}B_{n} 0 ≤ C n ≤ 2 s 2 A n + 2 t 2 B n ; the right-hand side converges to 0 0 0 , hence so does ( C n ) (C_{n}) ( C n ) , and property (c) holds. So φ \varphi φ is a noise intrinsic test function on Q Q Q , which proves claim 5.
Step 8 (Two facts on series). (T) Tail estimate. Let E E E be a real inner product space with norm ∣ ⋅ ∣ E |\cdot|_{E} ∣ ⋅ ∣ E , let ( v k ) k ∈ N (v_{k})_{k\in\mathbb{N}} ( v k ) k ∈ N be a sequence in E E E whose series converges (Series in a Real Inner Product Space §convergent ), with partial sums s n s_{n} s n and sum v v v , and let ( b k ) k ∈ N (b_{k})_{k\in\mathbb{N}} ( b k ) k ∈ N be real numbers with ∣ v k ∣ E ≤ b k |v_{k}|_{E}\le b_{k} ∣ v k ∣ E ≤ b k for every k k k whose series converges, with partial sums t n t_{n} t n and sum b b b . Then ∣ v − s K ∣ E ≤ b − t K |v-s_{K}|_{E}\le b-t_{K} ∣ v − s K ∣ E ≤ b − t K for every K ∈ N K\in\mathbb{N} K ∈ N . Proof: fix K K K . By the recursive definition of finite sums (Finite Sum Notation in a Vector Space , and The Real Numbers: Standing Notation and Background §naturals for real numbers), s K + 1 − s K = v K + 1 s_{K+1}-s_{K}=v_{K+1} s K + 1 − s K = v K + 1 and s K + n + 1 − s K = ( s K + n − s K ) + v K + n + 1 s_{K+n+1}-s_{K}=(s_{K+n}-s_{K})+v_{K+n+1} s K + n + 1 − s K = ( s K + n − s K ) + v K + n + 1 , and likewise for ( t n ) (t_{n}) ( t n ) ; so induction on n n n , with the triangle inequality of The Norm Metric of a Real Inner Product Space: Triangle Inequalities, Limits and Continuity §triangle and the hypothesis ∣ v k ∣ E ≤ b k |v_{k}|_{E}\le b_{k} ∣ v k ∣ E ≤ b k , gives ∣ s K + n − s K ∣ E ≤ t K + n − t K |s_{K+n}-s_{K}|_{E}\le t_{K+n}-t_{K} ∣ s K + n − s K ∣ E ≤ t K + n − t K for every n ∈ N n\in\mathbb{N} n ∈ N . Since 0 ≤ ∣ v k ∣ E ≤ b k 0\le|v_{k}|_{E}\le b_{k} 0 ≤ ∣ v k ∣ E ≤ b k , Series of Nonnegative Real Numbers, Comparison, and the Geometric Series §dominates gives t K + n ≤ b t_{K+n}\le b t K + n ≤ b , whence ∣ s K + n − s K ∣ E ≤ b − t K |s_{K+n}-s_{K}|_{E}\le b-t_{K} ∣ s K + n − s K ∣ E ≤ b − t K . The sequence ( s K + n ) n ∈ N (s_{K+n})_{n\in\mathbb{N}} ( s K + n ) n ∈ N is the subsequence of ( s n ) (s_{n}) ( s n ) determined by the strictly increasing index sequence n ↦ K + n n\mapsto K+n n ↦ K + n , so it converges to v v v by A Subsequence of a Convergent Sequence Has the Same Limit , applied in the metric space of the norm of E E E ; so ( s K + n − s K ) n (s_{K+n}-s_{K})_{n} ( s K + n − s K ) n converges to v − s K v-s_{K} v − s K by The Norm Metric of a Real Inner Product Space: Triangle Inequalities, Limits and Continuity §linear-limits , and ( ∣ s K + n − s K ∣ E ) n (|s_{K+n}-s_{K}|_{E})_{n} ( ∣ s K + n − s K ∣ E ) n converges to ∣ v − s K ∣ E |v-s_{K}|_{E} ∣ v − s K ∣ E by The Norm Metric of a Real Inner Product Space: Triangle Inequalities, Limits and Continuity §continuity ; the order of limits gives ∣ v − s K ∣ E ≤ b − t K |v-s_{K}|_{E}\le b-t_{K} ∣ v − s K ∣ E ≤ b − t K . We use (T) for E = L 2 ( λ ; X a ) E=L^{2}(\lambda;X^{a}) E = L 2 ( λ ; X a ) and for E = R E=\mathbb{R} E = R , the real inner product space of Elementary Properties of Bounded Linear Maps and Functionals on Real Inner Product Spaces §real-line with norm the absolute value, in which partial sums, convergence and sums of series agree with those of Series of Real Numbers by Elementary Properties of Series in a Real Inner Product Space §real-line .
(S) Signed comparison. Let ( c k ) (c_{k}) ( c k ) and ( ω k ) (\omega_{k}) ( ω k ) be real sequences with ∣ c k ∣ ≤ ω k |c_{k}|\le\omega_{k} ∣ c k ∣ ≤ ω k for every k k k such that ∑ k = 1 ∞ ω k \sum_{k=1}^{\infty}\omega_{k} ∑ k = 1 ∞ ω k converges with sum ω \omega ω . Then ∑ k = 1 ∞ c k \sum_{k=1}^{\infty}c_{k} ∑ k = 1 ∞ c k converges and ∣ ∑ k = 1 ∞ c k ∣ ≤ ω |\sum_{k=1}^{\infty}c_{k}|\le\omega ∣ ∑ k = 1 ∞ c k ∣ ≤ ω . Proof: 0 ≤ c k + ω k ≤ 2 ω k 0\le c_{k}+\omega_{k}\le2\omega_{k} 0 ≤ c k + ω k ≤ 2 ω k , and ∑ k 2 ω k \sum_{k}2\omega_{k} ∑ k 2 ω k converges by Elementary Properties of Series of Real Numbers §linearity , so ∑ k ( c k + ω k ) \sum_{k}(c_{k}+\omega_{k}) ∑ k ( c k + ω k ) converges by Series of Nonnegative Real Numbers, Comparison, and the Geometric Series §comparison ; then ∑ k c k = ∑ k ( ( c k + ω k ) + ( − 1 ) ω k ) \sum_{k}c_{k}=\sum_{k}\bigl((c_{k}+\omega_{k})+(-1)\omega_{k}\bigr) ∑ k c k = ∑ k ( ( c k + ω k ) + ( − 1 ) ω k ) converges by Elementary Properties of Series of Real Numbers §linearity . As c k ≤ ω k c_{k}\le\omega_{k} c k ≤ ω k and − c k ≤ ω k -c_{k}\le\omega_{k} − c k ≤ ω k , and ∑ k ( − c k ) = − ∑ k c k \sum_{k}(-c_{k})=-\sum_{k}c_{k} ∑ k ( − c k ) = − ∑ k c k by the same clause, Elementary Properties of Series of Real Numbers §order gives ∑ k c k ≤ ω \sum_{k}c_{k}\le\omega ∑ k c k ≤ ω and − ∑ k c k ≤ ω -\sum_{k}c_{k}\le\omega − ∑ k c k ≤ ω .
Step 9 (Claim 2, and notation for the series). In Steps 9 to 13 let B B B , ( μ k ) k ∈ N (\mu_{k})_{k\in\mathbb{N}} ( μ k ) k ∈ N and ( β k ) k ∈ N (\beta_{k})_{k\in\mathbb{N}} ( β k ) k ∈ N be as in claim 2. For λ ∈ P ρ a \lambda\in\mathcal{P}^{a}_{\rho} λ ∈ P ρ a put R λ = W a ( λ , ρ ) + B R_{\lambda}=W_{a}(\lambda,\rho)+B R λ = W a ( λ , ρ ) + B , which is nonnegative. Since ρ ∈ P ρ a \rho\in\mathcal{P}^{a}_{\rho} ρ ∈ P ρ a by The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §reference , the triangle inequality and symmetry (The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §triangle , The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §symmetry ) give, for every k ∈ N k\in\mathbb{N} k ∈ N ,
0 ≤ W a ( λ , μ k ) ≤ W a ( λ , ρ ) + W a ( μ k , ρ ) ≤ R λ , hence 0 ≤ W a ( λ , μ k ) 2 ≤ R λ 2 . (9) 0\le W_{a}(\lambda,\mu_{k})\le W_{a}(\lambda,\rho)+W_{a}(\mu_{k},\rho)\le R_{\lambda},\qquad\text{hence}\qquad0\le W_{a}(\lambda,\mu_{k})^{2}\le R_{\lambda}^{2}.\tag{9} 0 ≤ W a ( λ , μ k ) ≤ W a ( λ , ρ ) + W a ( μ k , ρ ) ≤ R λ , hence 0 ≤ W a ( λ , μ k ) 2 ≤ R λ 2 . ( 9 )
Let μ ∈ P ρ a \mu\in\mathcal{P}^{a}_{\rho} μ ∈ P ρ a . By Series of Nonnegative Real Numbers, Comparison, and the Geometric Series §tail-bound , read with ( β k ) (\beta_{k}) ( β k ) in place of its ( μ k ) (\mu_{k}) ( μ k ) , M = R μ M=R_{\mu} M = R μ and w k = W a ( μ , μ k ) w_{k}=W_{a}(\mu,\mu_{k}) w k = W a ( μ , μ k ) , and again with M = R μ 2 M=R_{\mu}^{2} M = R μ 2 and w k = W a ( μ , μ k ) 2 w_{k}=W_{a}(\mu,\mu_{k})^{2} w k = W a ( μ , μ k ) 2 , both series of claim 2 converge; this proves claim 2. Let β = ∑ k = 1 ∞ β k \beta=\sum_{k=1}^{\infty}\beta_{k} β = ∑ k = 1 ∞ β k , let σ K = ∑ k = 1 K β k \sigma_{K}=\sum_{k=1}^{K}\beta_{k} σ K = ∑ k = 1 K β k be the partial sums (Series of Real Numbers §partial-sums ) and r K = β − σ K r_{K}=\beta-\sigma_{K} r K = β − σ K . By Series of Nonnegative Real Numbers, Comparison, and the Geometric Series §dominates , 0 ≤ σ K ≤ β 0\le\sigma_{K}\le\beta 0 ≤ σ K ≤ β , so r K ≥ 0 r_{K}\ge0 r K ≥ 0 , and β ≥ σ 1 = β 1 > 0 \beta\ge\sigma_{1}=\beta_{1}>0 β ≥ σ 1 = β 1 > 0 . By Series of Real Numbers §convergent , ( σ K ) (\sigma_{K}) ( σ K ) converges to β \beta β , so ( r K ) (r_{K}) ( r K ) , and with it ( L r K ) (L\,r_{K}) ( L r K ) and ( L r K 2 ) (L\,r_{K}^{2}) ( L r K 2 ) for every real L L L , converge to 0 0 0 ; hence for every positive δ \delta δ and every real L L L there is K ∈ N K\in\mathbb{N} K ∈ N with L r K < δ L\,r_{K}<\delta L r K < δ , and one with L r K 2 < δ L\,r_{K}^{2}<\delta L r K 2 < δ . Let ψ \psi ψ be the function of claim 3 and, for K ∈ N K\in\mathbb{N} K ∈ N , ψ K ( λ ) = ∑ k = 1 K β k W a ( λ , μ k ) 2 \psi_{K}(\lambda)=\sum_{k=1}^{K}\beta_{k}W_{a}(\lambda,\mu_{k})^{2} ψ K ( λ ) = ∑ k = 1 K β k W a ( λ , μ k ) 2 its partial sums.
Step 10 (Claim 3, property (a)). Let μ ∈ P ρ a \mu\in\mathcal{P}^{a}_{\rho} μ ∈ P ρ a and let ε ∈ R \varepsilon\in\mathbb{R} ε ∈ R be positive. Put δ = min { 1 , ε ( β ( 2 R μ + 1 ) ) − 1 } \delta=\min\{1,\varepsilon\,(\beta(2R_{\mu}+1))^{-1}\} δ = min { 1 , ε ( β ( 2 R μ + 1 ) ) − 1 } , which is positive. Let ν ∈ P ρ a \nu\in\mathcal{P}^{a}_{\rho} ν ∈ P ρ a satisfy W a ( μ , ν ) < δ W_{a}(\mu,\nu)<\delta W a ( μ , ν ) < δ and put Δ k = W a ( ν , μ k ) 2 − W a ( μ , μ k ) 2 \Delta_{k}=W_{a}(\nu,\mu_{k})^{2}-W_{a}(\mu,\mu_{k})^{2} Δ k = W a ( ν , μ k ) 2 − W a ( μ , μ k ) 2 . By Elementary Properties of Series of Real Numbers §linearity the series ∑ k β k Δ k \sum_{k}\beta_{k}\Delta_{k} ∑ k β k Δ k converges with sum ψ ( ν ) − ψ ( μ ) \psi(\nu)-\psi(\mu) ψ ( ν ) − ψ ( μ ) . Since W a ( μ , ν ) ≤ 1 W_{a}(\mu,\nu)\le1 W a ( μ , ν ) ≤ 1 , Step 1(ii), read with μ ′ = μ \mu'=\mu μ ′ = μ , μ ′ ′ = ν \mu''=\nu μ ′′ = ν and λ = μ k \lambda=\mu_{k} λ = μ k , and (9) give ∣ β k Δ k ∣ ≤ β k ( 2 R μ + 1 ) W a ( μ , ν ) |\beta_{k}\Delta_{k}|\le\beta_{k}(2R_{\mu}+1)W_{a}(\mu,\nu) ∣ β k Δ k ∣ ≤ β k ( 2 R μ + 1 ) W a ( μ , ν ) , and the series of these bounds converges with sum β ( 2 R μ + 1 ) W a ( μ , ν ) \beta(2R_{\mu}+1)W_{a}(\mu,\nu) β ( 2 R μ + 1 ) W a ( μ , ν ) by Elementary Properties of Series of Real Numbers §linearity . By Step 8(S), ∣ ψ ( ν ) − ψ ( μ ) ∣ ≤ β ( 2 R μ + 1 ) W a ( μ , ν ) < β ( 2 R μ + 1 ) δ ≤ ε |\psi(\nu)-\psi(\mu)|\le\beta(2R_{\mu}+1)W_{a}(\mu,\nu)<\beta(2R_{\mu}+1)\delta\le\varepsilon ∣ ψ ( ν ) − ψ ( μ ) ∣ ≤ β ( 2 R μ + 1 ) W a ( μ , ν ) < β ( 2 R μ + 1 ) δ ≤ ε . So ψ \psi ψ is continuous at μ \mu μ , and property (a) holds.
Step 11 (The candidate gradients). Let μ ′ ∈ Q \mu'\in Q μ ′ ∈ Q . For every k k k the pair ( μ ′ , μ k ) (\mu',\mu_{k}) ( μ ′ , μ k ) is uniquely noise-mapped by The Noise Map Property of a Set of Probability Measures §map-property ; let S k μ ′ S^{\mu'}_{k} S k μ ′ be any noise-optimal map from μ ′ \mu' μ ′ to μ k \mu_{k} μ k (Noise-Optimal Maps and Uniquely Noise-Mapped Pairs §uniquely-mapped ); then S k μ ′ − i d ∈ T μ ′ a S^{\mu'}_{k}-\mathrm{id}\in T^{a}_{\mu'} S k μ ′ − id ∈ T μ ′ a by the same clause. Put ξ k μ ′ = 2 β k ( i d − S k μ ′ ) = ( − 2 β k ) ( S k μ ′ − i d ) ∈ L 2 ( μ ′ ; X a ) \xi^{\mu'}_{k}=2\beta_{k}(\mathrm{id}-S^{\mu'}_{k})=(-2\beta_{k})(S^{\mu'}_{k}-\mathrm{id})\in L^{2}(\mu';X^{a}) ξ k μ ′ = 2 β k ( id − S k μ ′ ) = ( − 2 β k ) ( S k μ ′ − id ) ∈ L 2 ( μ ′ ; X a ) and let s K μ ′ = ∑ k = 1 K ξ k μ ′ s^{\mu'}_{K}=\sum_{k=1}^{K}\xi^{\mu'}_{k} s K μ ′ = ∑ k = 1 K ξ k μ ′ be its partial sums (Series in a Real Inner Product Space §partial-sums ).
(i) ∥ ξ k μ ′ ∥ μ ′ = 2 β k W a ( μ ′ , μ k ) \lVert \xi^{\mu'}_{k}\rVert_{\mu'}=2\beta_{k}W_{a}(\mu',\mu_{k}) ∥ ξ k μ ′ ∥ μ ′ = 2 β k W a ( μ ′ , μ k ) : indeed ∥ S k μ ′ − i d ∥ μ ′ 2 = W a ( μ ′ , μ k ) 2 \lVert S^{\mu'}_{k}-\mathrm{id}\rVert_{\mu'}^{2}=W_{a}(\mu',\mu_{k})^{2} ∥ S k μ ′ − id ∥ μ ′ 2 = W a ( μ ′ , μ k ) 2 by The Noise-Optimal Map: Transport Cost, Stability Along Nearly Optimal Couplings and Stability Under Perturbation of the Source §cost , both numbers ∥ S k μ ′ − i d ∥ μ ′ \lVert S^{\mu'}_{k}-\mathrm{id}\rVert_{\mu'} ∥ S k μ ′ − id ∥ μ ′ and W a ( μ ′ , μ k ) W_{a}(\mu',\mu_{k}) W a ( μ ′ , μ k ) are nonnegative, and ∣ − 2 β k ∣ = 2 β k |{-2\beta_{k}}|=2\beta_{k} ∣ − 2 β k ∣ = 2 β k , so homogeneity in L 2 ( μ ′ ; X a ) L^{2}(\mu';X^{a}) L 2 ( μ ′ ; X a ) (Elementary Identities in a Real Inner Product Space §homogeneity ) gives the claim.
(ii) The series ∑ k ξ k μ ′ \sum_{k}\xi^{\mu'}_{k} ∑ k ξ k μ ′ converges in L 2 ( μ ′ ; X a ) L^{2}(\mu';X^{a}) L 2 ( μ ′ ; X a ) ; we write η μ ′ \eta_{\mu'} η μ ′ for its sum. Moreover ∥ η μ ′ ∥ μ ′ ≤ 2 ∑ k = 1 ∞ β k W a ( μ ′ , μ k ) \lVert\eta_{\mu'}\rVert_{\mu'}\le2\sum_{k=1}^{\infty}\beta_{k}W_{a}(\mu',\mu_{k}) ∥ η μ ′ ∥ μ ′ ≤ 2 ∑ k = 1 ∞ β k W a ( μ ′ , μ k ) . Indeed, by (i), claim 2 and Elementary Properties of Series of Real Numbers §linearity , the series ∑ k ∥ ξ k μ ′ ∥ μ ′ \sum_{k}\lVert \xi^{\mu'}_{k}\rVert_{\mu'} ∑ k ∥ ξ k μ ′ ∥ μ ′ converges with sum 2 ∑ k β k W a ( μ ′ , μ k ) 2\sum_{k}\beta_{k}W_{a}(\mu',\mu_{k}) 2 ∑ k β k W a ( μ ′ , μ k ) , so ∑ k ξ k μ ′ \sum_{k}\xi^{\mu'}_{k} ∑ k ξ k μ ′ converges absolutely , and Elementary Properties of Series in a Real Inner Product Space §absolute in the real Hilbert space L 2 ( μ ′ ; X a ) L^{2}(\mu';X^{a}) L 2 ( μ ′ ; X a ) gives both assertions.
(iii) If M ∈ R M\in\mathbb{R} M ∈ R is nonnegative and W a ( μ ′ , μ k ) ≤ M W_{a}(\mu',\mu_{k})\le M W a ( μ ′ , μ k ) ≤ M for every k k k , then ∥ η μ ′ − s K μ ′ ∥ μ ′ ≤ 2 M r K \lVert\eta_{\mu'}-s^{\mu'}_{K}\rVert_{\mu'}\le2M\,r_{K} ∥ η μ ′ − s K μ ′ ∥ μ ′ ≤ 2 M r K for every K K K . Indeed ∥ ξ k μ ′ ∥ μ ′ ≤ 2 M β k \lVert \xi^{\mu'}_{k}\rVert_{\mu'}\le2M\beta_{k} ∥ ξ k μ ′ ∥ μ ′ ≤ 2 M β k by (i), the series ∑ k 2 M β k \sum_{k}2M\beta_{k} ∑ k 2 M β k converges with sum 2 M β 2M\beta 2 Mβ and K K K -th partial sum 2 M σ K 2M\sigma_{K} 2 M σ K (Elementary Properties of Series of Real Numbers §linearity ), and Step 8(T) in E = L 2 ( μ ′ ; X a ) E=L^{2}(\mu';X^{a}) E = L 2 ( μ ′ ; X a ) gives ∥ η μ ′ − s K μ ′ ∥ μ ′ ≤ 2 M β − 2 M σ K = 2 M r K \lVert\eta_{\mu'}-s^{\mu'}_{K}\rVert_{\mu'}\le2M\beta-2M\sigma_{K}=2Mr_{K} ∥ η μ ′ − s K μ ′ ∥ μ ′ ≤ 2 Mβ − 2 M σ K = 2 M r K .
(iv) η μ ′ ∈ T μ ′ a \eta_{\mu'}\in T^{a}_{\mu'} η μ ′ ∈ T μ ′ a . By Linearity of the Noise Gradient, and the Noise Tangent Space is a Closed Linear Subspace §subspace , T μ ′ a T^{a}_{\mu'} T μ ′ a is a closed linear subspace of L 2 ( μ ′ ; X a ) L^{2}(\mu';X^{a}) L 2 ( μ ′ ; X a ) . Each ξ k μ ′ \xi^{\mu'}_{k} ξ k μ ′ is a real multiple of S k μ ′ − i d ∈ T μ ′ a S^{\mu'}_{k}-\mathrm{id}\in T^{a}_{\mu'} S k μ ′ − id ∈ T μ ′ a , hence lies in T μ ′ a T^{a}_{\mu'} T μ ′ a ; since s 1 μ ′ = ξ 1 μ ′ s^{\mu'}_{1}=\xi^{\mu'}_{1} s 1 μ ′ = ξ 1 μ ′ and s K + 1 μ ′ = s K μ ′ + ξ K + 1 μ ′ s^{\mu'}_{K+1}=s^{\mu'}_{K}+\xi^{\mu'}_{K+1} s K + 1 μ ′ = s K μ ′ + ξ K + 1 μ ′ (Finite Sum Notation in a Vector Space ), induction gives s K μ ′ ∈ T μ ′ a s^{\mu'}_{K}\in T^{a}_{\mu'} s K μ ′ ∈ T μ ′ a for every K K K . The sequence ( s K μ ′ ) (s^{\mu'}_{K}) ( s K μ ′ ) converges to η μ ′ \eta_{\mu'} η μ ′ (Series in a Real Inner Product Space §convergent ) and T μ ′ a T^{a}_{\mu'} T μ ′ a is closed in the sense of Real Hilbert Space §topology , so η μ ′ ∈ T μ ′ a \eta_{\mu'}\in T^{a}_{\mu'} η μ ′ ∈ T μ ′ a by Sequential Characterization of Closed Subsets of a Metric Space .
(v) For every K ∈ N K\in\mathbb{N} K ∈ N , ψ K \psi_{K} ψ K is a noise intrinsic test function on Q Q Q and ∇ ψ K ( μ ′ ) = s K μ ′ \nabla\psi_{K}(\mu')=s^{\mu'}_{K} ∇ ψ K ( μ ′ ) = s K μ ′ for every μ ′ ∈ Q \mu'\in Q μ ′ ∈ Q . Indeed, for each k k k the function φ k ( λ ) = W a ( λ , μ k ) 2 \varphi_{k}(\lambda)=W_{a}(\lambda,\mu_{k})^{2} φ k ( λ ) = W a ( λ , μ k ) 2 is a noise intrinsic test function on Q Q Q with ∇ φ k ( μ ′ ) = 2 ( i d − S k μ ′ ) \nabla\varphi_{k}(\mu')=2(\mathrm{id}-S^{\mu'}_{k}) ∇ φ k ( μ ′ ) = 2 ( id − S k μ ′ ) for μ ′ ∈ Q \mu'\in Q μ ′ ∈ Q , by claim 1 (Steps 2 to 6) read with ν 0 = μ k \nu_{0}=\mu_{k} ν 0 = μ k . Now ψ 1 = β 1 φ 1 + 0 φ 1 \psi_{1}=\beta_{1}\varphi_{1}+0\,\varphi_{1} ψ 1 = β 1 φ 1 + 0 φ 1 and ψ K + 1 = 1 ψ K + β K + 1 φ K + 1 \psi_{K+1}=1\,\psi_{K}+\beta_{K+1}\varphi_{K+1} ψ K + 1 = 1 ψ K + β K + 1 φ K + 1 by the recursive definition of finite sums (The Real Numbers: Standing Notation and Background §naturals ), so induction on K K K with claim 5 (Step 7) shows that ψ K \psi_{K} ψ K is a noise intrinsic test function on Q Q Q with ∇ ψ 1 ( μ ′ ) = β 1 ⋅ 2 ( i d − S 1 μ ′ ) + 0 ⋅ 2 ( i d − S 1 μ ′ ) = ξ 1 μ ′ \nabla\psi_{1}(\mu')=\beta_{1}\cdot2(\mathrm{id}-S^{\mu'}_{1})+0\cdot2(\mathrm{id}-S^{\mu'}_{1})=\xi^{\mu'}_{1} ∇ ψ 1 ( μ ′ ) = β 1 ⋅ 2 ( id − S 1 μ ′ ) + 0 ⋅ 2 ( id − S 1 μ ′ ) = ξ 1 μ ′ and ∇ ψ K + 1 ( μ ′ ) = 1 s K μ ′ + β K + 1 ⋅ 2 ( i d − S K + 1 μ ′ ) = s K + 1 μ ′ \nabla\psi_{K+1}(\mu')=1\,s^{\mu'}_{K}+\beta_{K+1}\cdot2(\mathrm{id}-S^{\mu'}_{K+1})=s^{\mu'}_{K+1} ∇ ψ K + 1 ( μ ′ ) = 1 s K μ ′ + β K + 1 ⋅ 2 ( id − S K + 1 μ ′ ) = s K + 1 μ ′ , by the vector-space axioms of L 2 ( μ ′ ; X a ) L^{2}(\mu';X^{a}) L 2 ( μ ′ ; X a ) , claim 3 of Elementary Identities in a Vector Space and Finite Sum Notation in a Vector Space .
Step 12 (Claim 3, property (b), and claim 4). Let μ ∈ Q \mu\in Q μ ∈ Q , let S k = S k μ S_{k}=S^{\mu}_{k} S k = S k μ be the noise-optimal maps of Step 11 (an arbitrary choice), and write η = η μ \eta=\eta_{\mu} η = η μ and s K = s K μ s_{K}=s^{\mu}_{K} s K = s K μ . We show that ψ \psi ψ is differentiable along noise couplings at μ \mu μ with gradient η \eta η . Let ε ∈ R \varepsilon\in\mathbb{R} ε ∈ R be positive. The order of choice is: first K K K , then θ 1 \theta_{1} θ 1 , then θ \theta θ . Choose K ∈ N K\in\mathbb{N} K ∈ N with ( 4 R μ + 1 ) r K < ε ⋅ 2 − 1 (4R_{\mu}+1)\,r_{K}<\varepsilon\cdot2^{-1} ( 4 R μ + 1 ) r K < ε ⋅ 2 − 1 (Step 9). By Step 11(v) and property (b) for ψ K \psi_{K} ψ K , ψ K \psi_{K} ψ K is differentiable along noise couplings at μ \mu μ with gradient s K s_{K} s K (Differentiability of a Function on the Noise-Connected Measures Along Noise Couplings, and Its Gradient §gradient ); choose a positive θ 1 \theta_{1} θ 1 as in Differentiability of a Function on the Noise-Connected Measures Along Noise Couplings, and Its Gradient §differentiable for ψ K \psi_{K} ψ K , s K s_{K} s K and ε ⋅ 2 − 1 \varepsilon\cdot2^{-1} ε ⋅ 2 − 1 , and put θ = min { θ 1 , 1 } \theta=\min\{\theta_{1},1\} θ = min { θ 1 , 1 } , which is positive.
Let ν ∈ P ρ a \nu\in\mathcal{P}^{a}_{\rho} ν ∈ P ρ a and π ∈ Π a ( μ , ν ) \pi\in\Pi^{a}(\mu,\nu) π ∈ Π a ( μ , ν ) satisfy I a ( π ) < θ 2 I^{a}(\pi)<\theta^{2} I a ( π ) < θ 2 . Then I a ( π ) < θ 1 2 I^{a}(\pi)<\theta_{1}^{2} I a ( π ) < θ 1 2 and I a ( π ) < 1 I^{a}(\pi)<1 I a ( π ) < 1 , so W a ( μ , ν ) ≤ I a ( π ) < 1 W_{a}(\mu,\nu)\le\sqrt{I^{a}(\pi)}<1 W a ( μ , ν ) ≤ I a ( π ) < 1 by (W). Put Δ k = W a ( ν , μ k ) 2 − W a ( μ , μ k ) 2 \Delta_{k}=W_{a}(\nu,\mu_{k})^{2}-W_{a}(\mu,\mu_{k})^{2} Δ k = W a ( ν , μ k ) 2 − W a ( μ , μ k ) 2 .
First, the real tail. By Elementary Properties of Series of Real Numbers §linearity the series ∑ k β k Δ k \sum_{k}\beta_{k}\Delta_{k} ∑ k β k Δ k converges with sum ψ ( ν ) − ψ ( μ ) \psi(\nu)-\psi(\mu) ψ ( ν ) − ψ ( μ ) , and its K K K -th partial sum is ψ K ( ν ) − ψ K ( μ ) \psi_{K}(\nu)-\psi_{K}(\mu) ψ K ( ν ) − ψ K ( μ ) . By Step 1(ii) (read with μ ′ = μ \mu'=\mu μ ′ = μ , μ ′ ′ = ν \mu''=\nu μ ′′ = ν , λ = μ k \lambda=\mu_{k} λ = μ k ), (9) and (W), ∣ β k Δ k ∣ ≤ β k ( 2 R μ + 1 ) I a ( π ) |\beta_{k}\Delta_{k}|\le\beta_{k}(2R_{\mu}+1)\sqrt{I^{a}(\pi)} ∣ β k Δ k ∣ ≤ β k ( 2 R μ + 1 ) I a ( π ) ; the series of these bounds converges with sum β ( 2 R μ + 1 ) I a ( π ) \beta(2R_{\mu}+1)\sqrt{I^{a}(\pi)} β ( 2 R μ + 1 ) I a ( π ) and K K K -th partial sum σ K ( 2 R μ + 1 ) I a ( π ) \sigma_{K}(2R_{\mu}+1)\sqrt{I^{a}(\pi)} σ K ( 2 R μ + 1 ) I a ( π ) (Elementary Properties of Series of Real Numbers §linearity ), so Step 8(T) in E = R E=\mathbb{R} E = R gives
∣ ( ψ ( ν ) − ψ ( μ ) ) − ( ψ K ( ν ) − ψ K ( μ ) ) ∣ ≤ ( 2 R μ + 1 ) r K I a ( π ) . \bigl|\bigl(\psi(\nu)-\psi(\mu)\bigr)-\bigl(\psi_{K}(\nu)-\psi_{K}(\mu)\bigr)\bigr|\le(2R_{\mu}+1)\,r_{K}\sqrt{I^{a}(\pi)} . ( ψ ( ν ) − ψ ( μ ) ) − ( ψ K ( ν ) − ψ K ( μ ) ) ≤ ( 2 R μ + 1 ) r K I a ( π ) .
Second, the field tail. By Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §linear , J a ( η , π ) = J a ( s K , π ) + J a ( η − s K , π ) \mathcal{J}^{a}(\eta,\pi)=\mathcal{J}^{a}(s_{K},\pi)+\mathcal{J}^{a}(\eta-s_{K},\pi) J a ( η , π ) = J a ( s K , π ) + J a ( η − s K , π ) , and by Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §bound and Step 11(iii) with M = R μ M=R_{\mu} M = R μ (admissible by (9)), ∣ J a ( η − s K , π ) ∣ ≤ ∥ η − s K ∥ μ I a ( π ) ≤ 2 R μ r K I a ( π ) |\mathcal{J}^{a}(\eta-s_{K},\pi)|\le\lVert\eta-s_{K}\rVert_{\mu}\sqrt{I^{a}(\pi)}\le2R_{\mu}r_{K}\sqrt{I^{a}(\pi)} ∣ J a ( η − s K , π ) ∣ ≤ ∥ η − s K ∥ μ I a ( π ) ≤ 2 R μ r K I a ( π ) . Third, the head: ∣ ψ K ( ν ) − ψ K ( μ ) − J a ( s K , π ) ∣ ≤ ε ⋅ 2 − 1 I a ( π ) |\psi_{K}(\nu)-\psi_{K}(\mu)-\mathcal{J}^{a}(s_{K},\pi)|\le\varepsilon\cdot2^{-1}\sqrt{I^{a}(\pi)} ∣ ψ K ( ν ) − ψ K ( μ ) − J a ( s K , π ) ∣ ≤ ε ⋅ 2 − 1 I a ( π ) by the choice of θ 1 \theta_{1} θ 1 . Writing ψ ( ν ) − ψ ( μ ) − J a ( η , π ) \psi(\nu)-\psi(\mu)-\mathcal{J}^{a}(\eta,\pi) ψ ( ν ) − ψ ( μ ) − J a ( η , π ) as the sum of the head expression, the real tail difference and − J a ( η − s K , π ) -\mathcal{J}^{a}(\eta-s_{K},\pi) − J a ( η − s K , π ) , the triangle inequality for the absolute value gives
∣ ψ ( ν ) − ψ ( μ ) − J a ( η , π ) ∣ ≤ ( ε ⋅ 2 − 1 + ( 4 R μ + 1 ) r K ) I a ( π ) ≤ ε I a ( π ) . \bigl|\psi(\nu)-\psi(\mu)-\mathcal{J}^{a}(\eta,\pi)\bigr|\le\bigl(\varepsilon\cdot2^{-1}+(4R_{\mu}+1)\,r_{K}\bigr)\sqrt{I^{a}(\pi)}\le\varepsilon\sqrt{I^{a}(\pi)} . ψ ( ν ) − ψ ( μ ) − J a ( η , π ) ≤ ( ε ⋅ 2 − 1 + ( 4 R μ + 1 ) r K ) I a ( π ) ≤ ε I a ( π ) .
Hence ψ \psi ψ is differentiable along noise couplings at μ \mu μ with gradient η \eta η , and ∇ ψ ( μ ) = η \nabla\psi(\mu)=\eta ∇ ψ ( μ ) = η by Differentiability of a Function on the Noise-Connected Measures Along Noise Couplings, and Its Gradient §gradient . By Step 11(iv), ∇ ψ ( μ ) ∈ T μ a \nabla\psi(\mu)\in T^{a}_{\mu} ∇ ψ ( μ ) ∈ T μ a , so property (b) holds. Since the noise-optimal maps S k S_{k} S k were arbitrary, this also proves claim 4: for any choice of noise-optimal maps S k S_{k} S k from μ \mu μ to μ k \mu_{k} μ k , the series ∑ k 2 β k ( i d − S k ) \sum_{k}2\beta_{k}(\mathrm{id}-S_{k}) ∑ k 2 β k ( id − S k ) converges in L 2 ( μ ; X a ) L^{2}(\mu;X^{a}) L 2 ( μ ; X a ) , in the sense of Real Hilbert Spaces: Series, Products, Orthonormal Bases and Differential Calculus §series , to ∇ ψ ( μ ) \nabla\psi(\mu) ∇ ψ ( μ ) , and ∥ ∇ ψ ( μ ) ∥ μ ≤ 2 ∑ k = 1 ∞ β k W a ( μ , μ k ) \lVert\nabla\psi(\mu)\rVert_{\mu}\le2\sum_{k=1}^{\infty}\beta_{k}W_{a}(\mu,\mu_{k}) ∥ ∇ ψ ( μ ) ∥ μ ≤ 2 ∑ k = 1 ∞ β k W a ( μ , μ k ) , by Step 11(ii).
Step 13 (Claim 3, property (c)). Let μ ∈ Q \mu\in Q μ ∈ Q , let ( λ n ) n ∈ N (\lambda_{n})_{n\in\mathbb{N}} ( λ n ) n ∈ N be a sequence in Q Q Q , and let ( π n ) n ∈ N (\pi_{n})_{n\in\mathbb{N}} ( π n ) n ∈ N be a sequence of couplings of vanishing noise cost from ( λ n ) (\lambda_{n}) ( λ n ) to μ \mu μ . For each n n n fix the noise-optimal maps S k λ n S^{\lambda_{n}}_{k} S k λ n of Step 11, write g n = ∇ ψ ( λ n ) = η λ n g_{n}=\nabla\psi(\lambda_{n})=\eta_{\lambda_{n}} g n = ∇ ψ ( λ n ) = η λ n and g = ∇ ψ ( μ ) = η μ g=\nabla\psi(\mu)=\eta_{\mu} g = ∇ ψ ( μ ) = η μ (Step 12), and let D n D_{n} D n be the discrepancy of g n g_{n} g n and g g g along π n \pi_{n} π n (Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §discrepancy , read with λ n \lambda_{n} λ n and μ \mu μ in place of its ν \nu ν and μ \mu μ , as in Strong and Weak Convergence of Noise Fields Along Couplings of Vanishing Noise Cost ). For K ∈ N K\in\mathbb{N} K ∈ N write h n , K = s K λ n = ∇ ψ K ( λ n ) h_{n,K}=s^{\lambda_{n}}_{K}=\nabla\psi_{K}(\lambda_{n}) h n , K = s K λ n = ∇ ψ K ( λ n ) and h K = s K μ = ∇ ψ K ( μ ) h_{K}=s^{\mu}_{K}=\nabla\psi_{K}(\mu) h K = s K μ = ∇ ψ K ( μ ) (Step 11(v)), and D n , K D_{n,K} D n , K for the discrepancy of h n , K h_{n,K} h n , K and h K h_{K} h K along π n \pi_{n} π n .
Comparison. Fix n n n and K K K , and choose representatives of h n , K h_{n,K} h n , K , v = g n − h n , K v=g_{n}-h_{n,K} v = g n − h n , K , h K h_{K} h K and w = g − h K w=g-h_{K} w = g − h K ; the pointwise sums h n , K + v h_{n,K}+v h n , K + v and h K + w h_{K}+w h K + w represent g n g_{n} g n and g g g (The Space of Square-Integrable Maps from a Measure Space into a Hilbert Space with an Orthonormal Basis §operations ), and discrepancies do not depend on the representatives. For every z z z , g n ( x ) − g ( y ) = ( h n , K ( x ) − h K ( y ) ) + ( v ( x ) − 0 X ) + ( 0 X − w ( y ) ) g_{n}(x)-g(y)=\bigl(h_{n,K}(x)-h_{K}(y)\bigr)+\bigl(v(x)-0_{X}\bigr)+\bigl(0_{X}-w(y)\bigr) g n ( x ) − g ( y ) = ( h n , K ( x ) − h K ( y ) ) + ( v ( x ) − 0 X ) + ( 0 X − w ( y ) ) , so by the triangle inequality in X a X^{a} X a and the real inequality ( p + q + r ) 2 ≤ 3 ( p 2 + q 2 + r 2 ) (p+q+r)^{2}\le3(p^{2}+q^{2}+r^{2}) ( p + q + r ) 2 ≤ 3 ( p 2 + q 2 + r 2 ) ,
∣ g n ( x ) − g ( y ) ∣ a 2 ≤ 3 ∣ h n , K ( x ) − h K ( y ) ∣ a 2 + 3 ∣ v ( x ) − 0 X ∣ a 2 + 3 ∣ 0 X − w ( y ) ∣ a 2 . |g_{n}(x)-g(y)|_{a}^{2}\le3\,|h_{n,K}(x)-h_{K}(y)|_{a}^{2}+3\,|v(x)-0_{X}|_{a}^{2}+3\,|0_{X}-w(y)|_{a}^{2}. ∣ g n ( x ) − g ( y ) ∣ a 2 ≤ 3 ∣ h n , K ( x ) − h K ( y ) ∣ a 2 + 3 ∣ v ( x ) − 0 X ∣ a 2 + 3 ∣ 0 X − w ( y ) ∣ a 2 .
The three functions on the right are the nonnegative Borel integrands of the discrepancies along π n \pi_{n} π n of h n , K h_{n,K} h n , K and h K h_{K} h K , of v v v and the zero element of L 2 ( μ ; X a ) L^{2}(\mu;X^{a}) L 2 ( μ ; X a ) , and of the zero element of L 2 ( λ n ; X a ) L^{2}(\lambda_{n};X^{a}) L 2 ( λ n ; X a ) and w w w ; the zero elements are the classes of the constant map 0 X 0_{X} 0 X (The Space of Square-Integrable Hilbert-Valued Maps is a Real Hilbert Space: Coordinates and Synthesis §hilbert ) and have norm 0 0 0 . By Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §cross-bound the cross pairings against a zero element vanish, so Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §discrepancy gives the values ∥ v ∥ λ n 2 \lVert v\rVert_{\lambda_{n}}^{2} ∥ v ∥ λ n 2 and ∥ w ∥ μ 2 \lVert w\rVert_{\mu}^{2} ∥ w ∥ μ 2 for the last two discrepancies. Integrating against π n \pi_{n} π n by Linearity and Monotonicity of the Lebesgue Integral §nonnegative ,
D n ≤ 3 D n , K + 3 ∥ g n − h n , K ∥ λ n 2 + 3 ∥ g − h K ∥ μ 2 . D_{n}\le3D_{n,K}+3\,\lVert g_{n}-h_{n,K}\rVert_{\lambda_{n}}^{2}+3\,\lVert g-h_{K}\rVert_{\mu}^{2}. D n ≤ 3 D n , K + 3 ∥ g n − h n , K ∥ λ n 2 + 3 ∥ g − h K ∥ μ 2 .
Tails. Since ( I a ( π n ) ) (I^{a}(\pi_{n})) ( I a ( π n )) converges to 0 0 0 (Strong and Weak Convergence of Noise Fields Along Couplings of Vanishing Noise Cost §couplings ), there is N 0 ∈ N N_{0}\in\mathbb{N} N 0 ∈ N with I a ( π n ) < 1 I^{a}(\pi_{n})<1 I a ( π n ) < 1 for n ≥ N 0 n\ge N_{0} n ≥ N 0 . For such n n n , W a ( μ , λ n ) = W a ( λ n , μ ) ≤ I a ( π n ) < 1 W_{a}(\mu,\lambda_{n})=W_{a}(\lambda_{n},\mu)\le\sqrt{I^{a}(\pi_{n})}<1 W a ( μ , λ n ) = W a ( λ n , μ ) ≤ I a ( π n ) < 1 by The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §symmetry and (W), so Step 1(ii), read with μ ′ = μ \mu'=\mu μ ′ = μ , μ ′ ′ = λ n \mu''=\lambda_{n} μ ′′ = λ n and λ = μ k \lambda=\mu_{k} λ = μ k , and (9) give W a ( λ n , μ k ) ≤ R μ + 1 W_{a}(\lambda_{n},\mu_{k})\le R_{\mu}+1 W a ( λ n , μ k ) ≤ R μ + 1 for every k k k ; also W a ( μ , μ k ) ≤ R μ + 1 W_{a}(\mu,\mu_{k})\le R_{\mu}+1 W a ( μ , μ k ) ≤ R μ + 1 by (9). Step 11(iii) with M = R μ + 1 M=R_{\mu}+1 M = R μ + 1 , at λ n \lambda_{n} λ n and at μ \mu μ , gives ∥ g n − h n , K ∥ λ n ≤ 2 ( R μ + 1 ) r K \lVert g_{n}-h_{n,K}\rVert_{\lambda_{n}}\le2(R_{\mu}+1)r_{K} ∥ g n − h n , K ∥ λ n ≤ 2 ( R μ + 1 ) r K and ∥ g − h K ∥ μ ≤ 2 ( R μ + 1 ) r K \lVert g-h_{K}\rVert_{\mu}\le2(R_{\mu}+1)r_{K} ∥ g − h K ∥ μ ≤ 2 ( R μ + 1 ) r K , whence
D n ≤ 3 D n , K + 24 ( R μ + 1 ) 2 r K 2 ( n ≥ N 0 , K ∈ N ) . D_{n}\le3D_{n,K}+24\,(R_{\mu}+1)^{2}\,r_{K}^{2}\qquad(n\ge N_{0},\ K\in\mathbb{N}). D n ≤ 3 D n , K + 24 ( R μ + 1 ) 2 r K 2 ( n ≥ N 0 , K ∈ N ) .
Conclusion. Let ε ∈ R \varepsilon\in\mathbb{R} ε ∈ R be positive. The order of choice is: first K K K , then N 1 N_{1} N 1 . Choose K K K with 24 ( R μ + 1 ) 2 r K 2 < ε ⋅ 2 − 1 24(R_{\mu}+1)^{2}r_{K}^{2}<\varepsilon\cdot2^{-1} 24 ( R μ + 1 ) 2 r K 2 < ε ⋅ 2 − 1 (Step 9). By Step 11(v), property (c) of the noise intrinsic test function ψ K \psi_{K} ψ K on Q Q Q , applied to μ \mu μ , ( λ n ) (\lambda_{n}) ( λ n ) and ( π n ) (\pi_{n}) ( π n ) , shows with Strong and Weak Convergence of Noise Fields Along Couplings of Vanishing Noise Cost §strong that ( D n , K ) n ∈ N (D_{n,K})_{n\in\mathbb{N}} ( D n , K ) n ∈ N converges to 0 0 0 , so there is N 1 N_{1} N 1 with 3 D n , K < ε ⋅ 2 − 1 3D_{n,K}<\varepsilon\cdot2^{-1} 3 D n , K < ε ⋅ 2 − 1 for n ≥ N 1 n\ge N_{1} n ≥ N 1 . For n ≥ max { N 0 , N 1 } n\ge\max\{N_{0},N_{1}\} n ≥ max { N 0 , N 1 } we get 0 ≤ D n < ε 0\le D_{n}<\varepsilon 0 ≤ D n < ε . Hence ( D n ) (D_{n}) ( D n ) converges to 0 0 0 , and property (c) holds.
Conclusion. Claim 1 was proved in Steps 2 to 6 and claim 5 in Step 7. Claim 2 is Step 9. Steps 10, 12 and 13 verify properties (a), (b) and (c) of Noise Intrinsic Test Functions on the Noise Wasserstein Space §test for the function ψ \psi ψ of claim 3, which is therefore a noise intrinsic test function on Q Q Q : this is claim 3. Claim 4 was proved in Step 12.