TheoremBase

Expand the shifts of the operator; properness is immediate, coercivity follows from completing a square, semicontinuity from strong and weak convergence along couplings and lower semicontinuity of the penalty, the structure condition from score monotonicity derived from displacement convexity along optimal maps, and momentum continuity from a Lipschitz bound.

Proof

Each result cited is universally quantified over the data in its own statement.

Elementary order and arithmetic of real numbers (adding inequalities, multiplying them by nonnegative or positive reals, handling absolute values, the nonnegativity of squares) is carried by The Real Numbers: Standing Notation and Background §background and is used without further mention; so is the fact that for nonnegative reals x,yx,y one has x<yx<y, respectively x≤yx\le y, exactly when x2<y2x^{2}<y^{2}, respectively x2≤y2x^{2}\le y^{2} (claims 1 and 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field). The four assumptions of the statement are referred to as (Convexity), (Lower semicontinuity of the penalty), (Growth) and (Running cost).

Step 0 (Notation and preliminary facts). Fix the data of the statement: the noise penalty pair (D,DΣ,E,Σ)(\mathcal{D},\mathcal{D}_{\Sigma},\mathcal{E},\Sigma), the reals λ0,θ\lambda_{0},\theta with 0<λ00<\lambda_{0} and 0<θ≤10<\theta\le1, the function gg, the operator FF, a real number CC as in (Growth), and a real number M≥0M\ge0 as in (Running cost), so that ∣g(μ)∣≤M|g(\mu)|\le M for every μ∈D\mu\in\mathcal{D}.

(0.1) Hilbert spaces and distances. For ν∈Pρa\nu\in\mathcal{P}^{a}_{\rho} the space L2(ν;Xa)L^{2}(\nu;X^{a}), with inner product ⟨⋅,⋅⟩ν\langle\cdot,\cdot\rangle_{\nu} and norm ∥⋅∥ν\lVert\cdot\rVert_{\nu}, is a real Hilbert space by Probability Measures on a Hilbert Space Transported in the Noise Norm: Standing Notation §fields, in particular a real inner product space, and ∥x∥ν2=⟨x,x⟩ν\lVert x\rVert_{\nu}^{2}=\langle x,x\rangle_{\nu} by Real Inner Product Space §norm. Inner products are symmetric by condition (a) of Real Inner Product Space §inner-product; bilinearity, the zero vector, homogeneity of the norm and the expansion of ∥x±y∥ν2\lVert x\pm y\rVert_{\nu}^{2} are Elementary Identities in a Real Inner Product Space §bilinear, Elementary Identities in a Real Inner Product Space §zero, Elementary Identities in a Real Inner Product Space §homogeneity and Elementary Identities in a Real Inner Product Space §expansion; ∣⟨x,y⟩ν∣≤∥x∥ν∥y∥ν|\langle x,y\rangle_{\nu}|\le\lVert x\rVert_{\nu}\lVert y\rVert_{\nu} by The Cauchy-Schwarz Inequality in a Real Inner Product Space; and the norm satisfies the triangle inequality of The Norm Metric of a Real Inner Product Space: Triangle Inequalities, Limits and Continuity §triangle. The noise space XaX^{a}, with inner product ⟨⋅,⋅⟩a\langle\cdot,\cdot\rangle_{a}, is a real Hilbert space and a linear subspace of XX carrying the vector operations of XX, by The Noise Space is a Real Hilbert Space: Orthonormal Basis, Continuous Embedding, Partial Sums, Closed Balls and Borel Measurability §hilbert; its inner product is linear in the first argument by conditions (b) and (c) of Real Inner Product Space §inner-product. For ν∈DΣ\nu\in\mathcal{D}_{\Sigma} one has ν∈D\nu\in\mathcal{D}, so that E(ν)\mathcal{E}(\nu) is a real number, and Σ(ν)∈Tνa⊆L2(ν;Xa)\Sigma(\nu)\in T^{a}_{\nu}\subseteq L^{2}(\nu;X^{a}), by Noise Penalty Pairs on the Noise Wasserstein Space: the Penalty, Its Score, and Their Domains §pair. For μ,ν∈Pρa\mu,\nu\in\mathcal{P}^{a}_{\rho} the ordered pair (μ,ν)(\mu,\nu) is noise-connected by The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §connected; Wa(μ,ν)W_{a}(\mu,\nu) is a nonnegative real number with Wa(μ,ν)2≤Ia(π)W_{a}(\mu,\nu)^{2}\le I^{a}(\pi) for every π∈Πa(μ,ν)\pi\in\Pi^{a}(\mu,\nu), by The Noise Wasserstein Distance §distance; Wa(μ,ν)=Wa(ν,μ)W_{a}(\mu,\nu)=W_{a}(\nu,\mu) by The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §symmetry; Wa(μ,μ)=0W_{a}(\mu,\mu)=0 by The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §separation; and WaW_{a} satisfies the triangle inequality The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §triangle. Moreover ρ∈Pρa\rho\in\mathcal{P}^{a}_{\rho} by The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §reference. Every π∈Πa(μ,ν)\pi\in\Pi^{a}(\mu,\nu) belongs to Π(μ,ν)\Pi(\mu,\nu) by Couplings of Finite Noise Cost and Their Noise Cost §couplings, and its noise cost Ia(π)I^{a}(\pi) is a nonnegative real number by Couplings of Finite Noise Cost and Their Noise Cost §cost.

(0.2) Cross pairings and discrepancies. Let ν′,ν∈P(X)\nu',\nu\in\mathcal{P}(X) and π∈Π(ν′,ν)\pi\in\Pi(\nu',\nu), and let Ka\mathcal{K}^{a} be the cross pairing of Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §cross, that clause being applied with ν′,ν\nu',\nu in place of its ν,μ\nu,\mu. (a) For k∈L2(ν′;Xa)k\in L^{2}(\nu';X^{a}) and w∈L2(ν;Xa)w\in L^{2}(\nu;X^{a}), the discrepancy D=∫X×X∣k(x)−w(y)∣a2 π(dz)D=\int_{X\times X}|k(x)-w(y)|_{a}^{2}\,\pi(dz) is a nonnegative real number by Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §discrepancy, and

D=∥k∥ν′2−2 Ka(k,w,π)+∥w∥ν2D=\lVert k\rVert_{\nu'}^{2}-2\,\mathcal{K}^{a}(k,w,\pi)+\lVert w\rVert_{\nu}^{2}

by the same clause. (b) For h,k∈L2(ν′;Xa)h,k\in L^{2}(\nu';X^{a}), w∈L2(ν;Xa)w\in L^{2}(\nu;X^{a}) and t∈Rt\in\mathbb{R},

Ka(th+k,w,π)=t Ka(h,w,π)+Ka(k,w,π).\mathcal{K}^{a}(th+k,w,\pi)=t\,\mathcal{K}^{a}(h,w,\pi)+\mathcal{K}^{a}(k,w,\pi).

Indeed, fix representatives h,k,wh,k,w. By The Space of Square-Integrable Maps from a Measure Space into a Hilbert Space with an Orthonormal Basis §operations the pointwise map th+kth+k is a representative of the class th+kth+k, and by Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §cross the cross pairing may be computed with any representatives, the functions z↦⟨h(x),w(y)⟩az\mapsto\langle h(x),w(y)\rangle_{a} and z↦⟨k(x),w(y)⟩az\mapsto\langle k(x),w(y)\rangle_{a} being integrable against π\pi. At every zz, linearity of ⟨⋅,⋅⟩a\langle\cdot,\cdot\rangle_{a} in its first argument (0.1) gives ⟨th(x)+k(x),w(y)⟩a=t⟨h(x),w(y)⟩a+⟨k(x),w(y)⟩a\langle th(x)+k(x),w(y)\rangle_{a}=t\langle h(x),w(y)\rangle_{a}+\langle k(x),w(y)\rangle_{a}, and the linearity of the integral, Linearity and Monotonicity of the Lebesgue Integral §integrable, gives the identity.

(0.3) Displacements of noise-optimal maps. Let μ,ν∈Pρa\mu,\nu\in\mathcal{P}^{a}_{\rho} and let SS be a noise-optimal map from μ\mu to ν\nu; by Noise-Optimal Maps and Uniquely Noise-Mapped Pairs §map, S#μ=νS_{\#}\mu=\nu and the displacement S−id∈L2(μ;Xa)S-\mathrm{id}\in L^{2}(\mu;X^{a}) is the class of the measurable map x↦S(x)−xx\mapsto S(x)-x into XaX^{a}. Apply Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §displacement with this μ\mu, with h=S−idh=S-\mathrm{id} and this representative, and with t=1t=1. Since the vector operations of XaX^{a} are those of XX (0.1), the map S1S_{1} there is x↦x+(S(x)−x)=S(x)x\mapsto x+(S(x)-x)=S(x), so S1=SS_{1}=S, and the coupling πS=(id,S)#μ∈Πa(μ,ν)\pi_{S}=(\mathrm{id},S)_{\#}\mu\in\Pi^{a}(\mu,\nu) satisfies

Ia(πS)=∥S−id∥μ2,Ja(ζ,πS)=⟨ζ,S−id⟩μfor every ζ∈L2(μ;Xa),I^{a}(\pi_{S})=\lVert S-\mathrm{id}\rVert_{\mu}^{2},\qquad\mathcal{J}^{a}(\zeta,\pi_{S})=\langle\zeta,S-\mathrm{id}\rangle_{\mu}\quad\text{for every }\zeta\in L^{2}(\mu;X^{a}),

where Ja\mathcal{J}^{a} is the noise displacement pairing of Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §pairing. The coupling πS\pi_{S} is noise-optimal by Noise-Optimal Maps and Uniquely Noise-Mapped Pairs §map, so Ia(πS)=Wa(μ,ν)2I^{a}(\pi_{S})=W_{a}(\mu,\nu)^{2} by Noise-Optimal Couplings §optimal; as both numbers are nonnegative, ∥S−id∥μ=Wa(μ,ν)\lVert S-\mathrm{id}\rVert_{\mu}=W_{a}(\mu,\nu).

(0.4) The constant CC is nonnegative. By Noise Penalty Pairs on the Noise Wasserstein Space: the Penalty, Its Score, and Their Domains §nonempty there is μ0∈DΣ⊆D\mu_{0}\in\mathcal{D}_{\Sigma}\subseteq\mathcal{D}. Then (Growth) gives 0≤Wa(μ0,ρ)2≤C (1+∣E(μ0)∣)0\le W_{a}(\mu_{0},\rho)^{2}\le C\,(1+|\mathcal{E}(\mu_{0})|). If C<0C<0, then, as 0<1+∣E(μ0)∣0<1+|\mathcal{E}(\mu_{0})|, we would get C (1+∣E(μ0)∣)<0C\,(1+|\mathcal{E}(\mu_{0})|)<0, a contradiction. Hence 0≤C0\le C.

(0.5) Elementary inequalities in δ\delta. Let δ∈R\delta\in\mathbb{R} with 0<δ<10<\delta<1. Then 0<θδ≤δ<10<\theta\delta\le\delta<1, because 0<θ≤10<\theta\le1. Consequently 0<1+θδ0<1+\theta\delta, 0≤1−θδ≤10\le1-\theta\delta\le1, δ2+θδ22>0\tfrac{\delta}{2}+\tfrac{\theta\delta^{2}}{2}>0, δ2−θδ22=δ2(1−θδ)≥0\tfrac{\delta}{2}-\tfrac{\theta\delta^{2}}{2}=\tfrac{\delta}{2}(1-\theta\delta)\ge0, and δ−θδ22=δ−δ2 θδ≥δ2>0\delta-\tfrac{\theta\delta^{2}}{2}=\delta-\tfrac{\delta}{2}\,\theta\delta\ge\tfrac{\delta}{2}>0. For real m,sm,s one has 2ms≤m2+s22ms\le m^{2}+s^{2}, since 0≤(m−s)2=m2−2ms+s20\le(m-s)^{2}=m^{2}-2ms+s^{2}.

(0.6) Expanded form of the shifts. Let (ν,q)∈Va(DΣ)(\nu,q)\in\mathcal{V}^{a}(\mathcal{D}_{\Sigma}), r∈Rr\in\mathbb{R} and δ>0\delta>0, and write σ=Σ(ν)\sigma=\Sigma(\nu). By The Bundle of Noise Fields over a Set of Measures, First-Order Equation Operators on the Noise Wasserstein Space, and Their Delta-Shifts §shifted, the formula of The Discounted Hamilton-Jacobi Equation with a Penalty Drift on the Noise Wasserstein Space §operator, and ⟨σ,q±δσ⟩ν=⟨σ,q⟩ν±δ∥σ∥ν2\langle\sigma,q\pm\delta\sigma\rangle_{\nu}=\langle\sigma,q\rangle_{\nu}\pm\delta\lVert\sigma\rVert_{\nu}^{2} (0.1),

Fδ−(ν,r,q)=λ0r+λ0δ E(ν)+θ2∥q+δσ∥ν2+⟨σ,q⟩ν+δ∥σ∥ν2−g(ν),(0a)F^{-}_{\delta}(\nu,r,q)=\lambda_{0}r+\lambda_{0}\delta\,\mathcal{E}(\nu)+\frac{\theta}{2}\lVert q+\delta\sigma\rVert_{\nu}^{2}+\langle\sigma,q\rangle_{\nu}+\delta\lVert\sigma\rVert_{\nu}^{2}-g(\nu),\tag{0a} Fδ+(ν,r,q)=λ0r−λ0δ E(ν)+θ2∥q−δσ∥ν2+⟨σ,q⟩ν−δ∥σ∥ν2−g(ν).(0c)F^{+}_{\delta}(\nu,r,q)=\lambda_{0}r-\lambda_{0}\delta\,\mathcal{E}(\nu)+\frac{\theta}{2}\lVert q-\delta\sigma\rVert_{\nu}^{2}+\langle\sigma,q\rangle_{\nu}-\delta\lVert\sigma\rVert_{\nu}^{2}-g(\nu).\tag{0c}

Expanding ∥q±δσ∥ν2=∥q∥ν2±2δ⟨σ,q⟩ν+δ2∥σ∥ν2\lVert q\pm\delta\sigma\rVert_{\nu}^{2}=\lVert q\rVert_{\nu}^{2}\pm2\delta\langle\sigma,q\rangle_{\nu}+\delta^{2}\lVert\sigma\rVert_{\nu}^{2} by (0.1) gives

Fδ−(ν,r,q)=λ0r+λ0δ E(ν)+θ2∥q∥ν2+(1+θδ)⟨σ,q⟩ν+(δ+θδ22)∥σ∥ν2−g(ν),(0b)F^{-}_{\delta}(\nu,r,q)=\lambda_{0}r+\lambda_{0}\delta\,\mathcal{E}(\nu)+\frac{\theta}{2}\lVert q\rVert_{\nu}^{2}+(1+\theta\delta)\langle\sigma,q\rangle_{\nu}+\Bigl(\delta+\frac{\theta\delta^{2}}{2}\Bigr)\lVert\sigma\rVert_{\nu}^{2}-g(\nu),\tag{0b} Fδ+(ν,r,q)=λ0r−λ0δ E(ν)+θ2∥q∥ν2+(1−θδ)⟨σ,q⟩ν−(δ−θδ22)∥σ∥ν2−g(ν).(0d)F^{+}_{\delta}(\nu,r,q)=\lambda_{0}r-\lambda_{0}\delta\,\mathcal{E}(\nu)+\frac{\theta}{2}\lVert q\rVert_{\nu}^{2}+(1-\theta\delta)\langle\sigma,q\rangle_{\nu}-\Bigl(\delta-\frac{\theta\delta^{2}}{2}\Bigr)\lVert\sigma\rVert_{\nu}^{2}-g(\nu).\tag{0d}

Step 1 (Local strict properness, claim 1). Let R>0R>0, (ν,q)∈Va(DΣ)(\nu,q)\in\mathcal{V}^{a}(\mathcal{D}_{\Sigma}) and −R≤s≤r≤R-R\le s\le r\le R. By the formula of The Discounted Hamilton-Jacobi Equation with a Penalty Drift on the Noise Wasserstein Space §operator all terms other than λ0r\lambda_{0}r and λ0s\lambda_{0}s cancel, so F(ν,r,q)−F(ν,s,q)=λ0(r−s)F(\nu,r,q)-F(\nu,s,q)=\lambda_{0}(r-s), and in particular λ0(r−s)≤F(ν,r,q)−F(ν,s,q)\lambda_{0}(r-s)\le F(\nu,r,q)-F(\nu,s,q). Hence the positive real λ0\lambda_{0} is a properness constant for FF at RR, for every R>0R>0, and FF is locally strictly proper. None of the four assumptions is used here.

Step 2 (Shift-coercivity, claim 2). Let δ,R∈R\delta,R\in\mathbb{R} with 0<δ<10<\delta<1 and 0<R0<R. Put

A=2λ0R+M+R22,B=R+2A+R22δ,Cδ,R=1+4(R+B)δ;A=2\lambda_{0}R+M+\frac{R^{2}}{2},\qquad B=R+2A+\frac{R^{2}}{2\delta},\qquad C_{\delta,R}=1+\frac{4(R+B)}{\delta};

these are nonnegative reals, as 0≤M0\le M. We show that Cδ,RC_{\delta,R} is a score bound for FF at (δ,R)(\delta,R).

Let ξ=(ν,r,q)\xi=(\nu,r,q) and η=(ν′,r′,q′)\eta=(\nu',r',q') be RR-bounded test data with Fδ−(ξ)−Fδ+(η)<RF^{-}_{\delta}(\xi)-F^{+}_{\delta}(\eta)<R. By Test Data for a First-Order Equation Operator on the Noise Wasserstein Space and the Admissible Sets §admissible every member of Sδ,R−S^{-}_{\delta,R} is such a ξ\xi for some η\eta, and every member of Sδ,R+S^{+}_{\delta,R} is such an η\eta for some ξ\xi; so it suffices to show s≤Cδ,Rs\le C_{\delta,R} and s′≤Cδ,Rs'\le C_{\delta,R}, where s=∥Σ(ν)∥νs=\lVert\Sigma(\nu)\rVert_{\nu} and s′=∥Σ(ν′)∥ν′s'=\lVert\Sigma(\nu')\rVert_{\nu'}. By Test Data for a First-Order Equation Operator on the Noise Wasserstein Space and the Admissible Sets §bounded, ∣r∣,∣r′∣<R|r|,|r'|<R, ∣E(ν)∣,∣E(ν′)∣<R|\mathcal{E}(\nu)|,|\mathcal{E}(\nu')|<R and ∥q∥ν,∥q′∥ν′<R\lVert q\rVert_{\nu},\lVert q'\rVert_{\nu'}<R.

Lower bound for Fδ−(ξ)F^{-}_{\delta}(\xi). In (0a) we have λ0r≥−λ0R\lambda_{0}r\ge-\lambda_{0}R; λ0δE(ν)≥−λ0δ∣E(ν)∣≥−λ0R\lambda_{0}\delta\mathcal{E}(\nu)\ge-\lambda_{0}\delta|\mathcal{E}(\nu)|\ge-\lambda_{0}R as δ<1\delta<1; θ2∥q+δΣ(ν)∥ν2≥0\tfrac{\theta}{2}\lVert q+\delta\Sigma(\nu)\rVert_{\nu}^{2}\ge0; ⟨Σ(ν),q⟩ν≥−sR\langle\Sigma(\nu),q\rangle_{\nu}\ge-sR by the Cauchy-Schwarz inequality of (0.1); δ∥Σ(ν)∥ν2=δs2≥δ2s2\delta\lVert\Sigma(\nu)\rVert_{\nu}^{2}=\delta s^{2}\ge\tfrac{\delta}{2}s^{2}; and −g(ν)≥−M-g(\nu)\ge-M by the choice of MM in Step 0, as ν∈D\nu\in\mathcal{D}. Hence, since R22≥0\tfrac{R^{2}}{2}\ge0,

Fδ−(ξ) ≥ δ2s2−Rs−A=φ(s)−A,where φ(x)=δ2x2−Rx  (x∈R).F^{-}_{\delta}(\xi)\ \ge\ \frac{\delta}{2}s^{2}-Rs-A=\varphi(s)-A,\qquad\text{where }\varphi(x)=\frac{\delta}{2}x^{2}-Rx\ \ (x\in\mathbb{R}).

Upper bound for Fδ+(η)F^{+}_{\delta}(\eta). In (0d) we have λ0r′≤λ0R\lambda_{0}r'\le\lambda_{0}R; −λ0δE(ν′)≤λ0R-\lambda_{0}\delta\mathcal{E}(\nu')\le\lambda_{0}R; θ2∥q′∥ν′2≤R22\tfrac{\theta}{2}\lVert q'\rVert_{\nu'}^{2}\le\tfrac{R^{2}}{2} as θ≤1\theta\le1; (1−θδ)⟨Σ(ν′),q′⟩ν′≤(1−θδ)s′R≤Rs′(1-\theta\delta)\langle\Sigma(\nu'),q'\rangle_{\nu'}\le(1-\theta\delta)s'R\le Rs' by Cauchy-Schwarz and 0≤1−θδ≤10\le1-\theta\delta\le1 (0.5); −(δ−θδ22)s′2≤−δ2s′2-(\delta-\tfrac{\theta\delta^{2}}{2})s'^{2}\le-\tfrac{\delta}{2}s'^{2} by (0.5); and −g(ν′)≤M-g(\nu')\le M. Hence Fδ+(η)≤A−φ(s′)F^{+}_{\delta}(\eta)\le A-\varphi(s').

Conclusion. For every real xx, 0≤δ2(x−Rδ)2=φ(x)+R22δ0\le\tfrac{\delta}{2}(x-\tfrac{R}{\delta})^{2}=\varphi(x)+\tfrac{R^{2}}{2\delta}, so φ(x)≥−R22δ\varphi(x)\ge-\tfrac{R^{2}}{2\delta}. From the two bounds, φ(s)+φ(s′)−2A≤Fδ−(ξ)−Fδ+(η)<R\varphi(s)+\varphi(s')-2A\le F^{-}_{\delta}(\xi)-F^{+}_{\delta}(\eta)<R, hence φ(s)<R+2A−φ(s′)≤B\varphi(s)<R+2A-\varphi(s')\le B and likewise φ(s′)<B\varphi(s')<B. Now let x≥0x\ge0 with φ(x)<B\varphi(x)<B, and suppose x>Cδ,Rx>C_{\delta,R}. Then x>1x>1, x>4Rδx>\tfrac{4R}{\delta} and x>4Bδx>\tfrac{4B}{\delta}, all three numbers being at most Cδ,RC_{\delta,R}. The second gives Rx<δ4x2Rx<\tfrac{\delta}{4}x^{2}, so φ(x)>δ4x2\varphi(x)>\tfrac{\delta}{4}x^{2}; the first gives x2>xx^{2}>x, so δ4x2>δ4x\tfrac{\delta}{4}x^{2}>\tfrac{\delta}{4}x; and the third gives δ4x>B\tfrac{\delta}{4}x>B. Thus φ(x)>B\varphi(x)>B, a contradiction. Hence x≤Cδ,Rx\le C_{\delta,R}; applied to x=sx=s and x=s′x=s' this proves the claim. As δ,R\delta,R were arbitrary, FF satisfies the shift-coercivity condition. Of the assumptions only the boundedness in (Running cost) and θ≤1\theta\le1 were used.

Step 3 (An estimate along a coupling). Claim. Let ν′,ν∈P(X)\nu',\nu\in\mathcal{P}(X), π∈Π(ν′,ν)\pi\in\Pi(\nu',\nu), h,k∈L2(ν′;Xa)h,k\in L^{2}(\nu';X^{a}), w∈L2(ν;Xa)w\in L^{2}(\nu;X^{a}), and let ℓ∈R\ell\in\mathbb{R} be positive with ∥h∥ν′≤ℓ\lVert h\rVert_{\nu'}\le\ell. Let D=∫X×X∣k(x)−w(y)∣a2 π(dz)D=\int_{X\times X}|k(x)-w(y)|_{a}^{2}\,\pi(dz) be the discrepancy of kk and ww along π\pi. Then

(⟨h,k⟩ν′−Ka(h,w,π))2≤ℓ2D.\bigl(\langle h,k\rangle_{\nu'}-\mathcal{K}^{a}(h,w,\pi)\bigr)^{2}\le\ell^{2}D .

Proof. Put β=⟨h,k⟩ν′−Ka(h,w,π)\beta=\langle h,k\rangle_{\nu'}-\mathcal{K}^{a}(h,w,\pi) and let t∈Rt\in\mathbb{R}. The field th+kth+k lies in L2(ν′;Xa)L^{2}(\nu';X^{a}), and by (0.2)(a) its discrepancy with ww along π\pi is nonnegative and equals ∥th+k∥ν′2−2Ka(th+k,w,π)+∥w∥ν2\lVert th+k\rVert_{\nu'}^{2}-2\mathcal{K}^{a}(th+k,w,\pi)+\lVert w\rVert_{\nu}^{2}. By (0.1), ∥th+k∥ν′2=t2∥h∥ν′2+2t⟨h,k⟩ν′+∥k∥ν′2\lVert th+k\rVert_{\nu'}^{2}=t^{2}\lVert h\rVert_{\nu'}^{2}+2t\langle h,k\rangle_{\nu'}+\lVert k\rVert_{\nu'}^{2}; by (0.2)(b), Ka(th+k,w,π)=tKa(h,w,π)+Ka(k,w,π)\mathcal{K}^{a}(th+k,w,\pi)=t\mathcal{K}^{a}(h,w,\pi)+\mathcal{K}^{a}(k,w,\pi); and by (0.2)(a) again, ∥k∥ν′2−2Ka(k,w,π)+∥w∥ν2=D\lVert k\rVert_{\nu'}^{2}-2\mathcal{K}^{a}(k,w,\pi)+\lVert w\rVert_{\nu}^{2}=D. Therefore

0≤t2∥h∥ν′2+2tβ+D≤t2ℓ2+2tβ+Dfor every t∈R.0\le t^{2}\lVert h\rVert_{\nu'}^{2}+2t\beta+D\le t^{2}\ell^{2}+2t\beta+D\qquad\text{for every }t\in\mathbb{R}.

With t=−βℓ−2t=-\beta\ell^{-2} this reads 0≤β2ℓ−2−2β2ℓ−2+D=D−β2ℓ−20\le\beta^{2}\ell^{-2}-2\beta^{2}\ell^{-2}+D=D-\beta^{2}\ell^{-2}, and multiplying by ℓ2>0\ell^{2}>0 gives the claim.

Consequence. Let ν∈P(X)\nu\in\mathcal{P}(X), w∈L2(ν;Xa)w\in L^{2}(\nu;X^{a}) and ℓ>0\ell>0, and for every n∈Nn\in\mathbb{N} let νn′∈P(X)\nu'_{n}\in\mathcal{P}(X), πn∈Π(νn′,ν)\pi_{n}\in\Pi(\nu'_{n},\nu) and hn,kn∈L2(νn′;Xa)h_{n},k_{n}\in L^{2}(\nu'_{n};X^{a}) with ∥hn∥νn′≤ℓ\lVert h_{n}\rVert_{\nu'_{n}}\le\ell, such that the discrepancies DnD_{n} of knk_{n} and ww along πn\pi_{n} converge to 00. Then βn=⟨hn,kn⟩νn′−Ka(hn,w,πn)\beta_{n}=\langle h_{n},k_{n}\rangle_{\nu'_{n}}-\mathcal{K}^{a}(h_{n},w,\pi_{n}) converges to 00: given ε>0\varepsilon>0, choose NN with Dn<ε2ℓ−2D_{n}<\varepsilon^{2}\ell^{-2} for n≥Nn\ge N (the DnD_{n} being nonnegative); then ∣βn∣2=βn2≤ℓ2Dn<ε2|\beta_{n}|^{2}=\beta_{n}^{2}\le\ell^{2}D_{n}<\varepsilon^{2} by the claim, so ∣βn∣<ε|\beta_{n}|<\varepsilon.

Step 4 (Shift-semicontinuity, claim 3). Let δ,R∈R\delta,R\in\mathbb{R} with 0<δ<10<\delta<1 and 0<R0<R, let ξn=(νn,rn,qn)\xi_{n}=(\nu_{n},r_{n},q_{n}) (n∈Nn\in\mathbb{N}) and ξ=(ν,r,q)\xi=(\nu,r,q) be test data, and let πn∈Πa(νn,ν)\pi_{n}\in\Pi^{a}(\nu_{n},\nu) be couplings such that (ξn)(\xi_{n}) converges to ξ\xi along (πn)(\pi_{n}) with score bounded by RR. Write σn=Σ(νn)\sigma_{n}=\Sigma(\nu_{n}), σ=Σ(ν)\sigma=\Sigma(\nu) and Kn(k,w)=Ka(k,w,πn)\mathcal{K}_{n}(k,w)=\mathcal{K}^{a}(k,w,\pi_{n}), with πn∈Π(νn,ν)\pi_{n}\in\Pi(\nu_{n},\nu) by (0.1). By that clause: every ξn\xi_{n} is RR-bounded, so ∣E(νn)∣<R|\mathcal{E}(\nu_{n})|<R and ∥qn∥νn<R\lVert q_{n}\rVert_{\nu_{n}}<R; ∥σn∥νn≤R\lVert\sigma_{n}\rVert_{\nu_{n}}\le R; (πn)(\pi_{n}) is a sequence of couplings of vanishing noise cost, so lim⁡nIa(πn)=0\lim_{n}I^{a}(\pi_{n})=0; (qn)(q_{n}) converges strongly to qq, so the discrepancies DnD_{n} of qnq_{n} and qq along πn\pi_{n} converge to 00; (σn)(\sigma_{n}) converges weakly to σ\sigma, so lim⁡nKn(σn,w)=⟨σ,w⟩ν\lim_{n}\mathcal{K}_{n}(\sigma_{n},w)=\langle\sigma,w\rangle_{\nu} for every w∈L2(ν;Xa)w\in L^{2}(\nu;X^{a}); and rn→rr_{n}\to r. The measures νn,ν\nu_{n},\nu lie in DΣ⊆D\mathcal{D}_{\Sigma}\subseteq\mathcal{D} by Test Data for a First-Order Equation Operator on the Noise Wasserstein Space and the Admissible Sets §data and Noise Penalty Pairs on the Noise Wasserstein Space: the Penalty, Its Score, and Their Domains §pair. Limits of sums, differences and scalar multiples of convergent real sequences are computed by Arithmetic of Limits of Real Sequences §sums and Arithmetic of Limits of Real Sequences §scalar.

(4.1) Wa(ν,νn)→0W_{a}(\nu,\nu_{n})\to0. Given ε>0\varepsilon>0, choose NN with Ia(πn)<ε2I^{a}(\pi_{n})<\varepsilon^{2} for n≥Nn\ge N, by Limit of a Sequence of Real Numbers. Then Wa(νn,ν)2≤Ia(πn)<ε2W_{a}(\nu_{n},\nu)^{2}\le I^{a}(\pi_{n})<\varepsilon^{2} by (0.1), so Wa(νn,ν)<εW_{a}(\nu_{n},\nu)<\varepsilon, and Wa(ν,νn)<εW_{a}(\nu,\nu_{n})<\varepsilon by symmetry (0.1), for n≥Nn\ge N.

(4.2) Convergent terms. (a) g(νn)→g(ν)g(\nu_{n})\to g(\nu): given ε>0\varepsilon>0, (Running cost) gives γ>0\gamma>0 with ∣g(μ)−g(μ′)∣<ε|g(\mu)-g(\mu')|<\varepsilon for μ,μ′∈D\mu,\mu'\in\mathcal{D} with Wa(μ,μ′)<γW_{a}(\mu,\mu')<\gamma, and by (4.1) this applies to μ=νn\mu=\nu_{n}, μ′=ν\mu'=\nu for all large nn. (b) ∥qn∥νn2→∥q∥ν2\lVert q_{n}\rVert_{\nu_{n}}^{2}\to\lVert q\rVert_{\nu}^{2}: by the consequence in Step 3 with ℓ=R\ell=R, hn=kn=qnh_{n}=k_{n}=q_{n} and w=qw=q, the numbers βn=∥qn∥νn2−Kn(qn,q)\beta_{n}=\lVert q_{n}\rVert_{\nu_{n}}^{2}-\mathcal{K}_{n}(q_{n},q) tend to 00; by (0.2)(a), Dn=∥qn∥νn2−2Kn(qn,q)+∥q∥ν2=2βn−∥qn∥νn2+∥q∥ν2D_{n}=\lVert q_{n}\rVert_{\nu_{n}}^{2}-2\mathcal{K}_{n}(q_{n},q)+\lVert q\rVert_{\nu}^{2}=2\beta_{n}-\lVert q_{n}\rVert_{\nu_{n}}^{2}+\lVert q\rVert_{\nu}^{2}, so ∥qn∥νn2=2βn−Dn+∥q∥ν2→∥q∥ν2\lVert q_{n}\rVert_{\nu_{n}}^{2}=2\beta_{n}-D_{n}+\lVert q\rVert_{\nu}^{2}\to\lVert q\rVert_{\nu}^{2}. (c) ⟨σn,qn⟩νn→⟨σ,q⟩ν\langle\sigma_{n},q_{n}\rangle_{\nu_{n}}\to\langle\sigma,q\rangle_{\nu}: by the consequence in Step 3 with ℓ=R\ell=R, hn=σnh_{n}=\sigma_{n}, kn=qnk_{n}=q_{n} and w=qw=q, the difference ⟨σn,qn⟩νn−Kn(σn,q)\langle\sigma_{n},q_{n}\rangle_{\nu_{n}}-\mathcal{K}_{n}(\sigma_{n},q) tends to 00, and Kn(σn,q)→⟨σ,q⟩ν\mathcal{K}_{n}(\sigma_{n},q)\to\langle\sigma,q\rangle_{\nu} by weak convergence with w=qw=q.

(4.3) Lower semicontinuous terms. (a) For every ε>0\varepsilon>0 there is NN with E(νn)>E(ν)−ε\mathcal{E}(\nu_{n})>\mathcal{E}(\nu)-\varepsilon for n≥Nn\ge N: by (Lower semicontinuity of the penalty) and Lower Semicontinuous Function on a Subset of a Metric Space, E\mathcal{E} is lower semicontinuous at ν∈D\nu\in\mathcal{D} relative to D\mathcal{D}, which gives γ>0\gamma>0 with E(ν)−ε<E(μ)\mathcal{E}(\nu)-\varepsilon<\mathcal{E}(\mu) for μ∈D\mu\in\mathcal{D} with Wa(ν,μ)<γW_{a}(\nu,\mu)<\gamma; apply (4.1), as νn∈D\nu_{n}\in\mathcal{D}. (b) For every ε>0\varepsilon>0 there is NN with ∥σn∥νn2>∥σ∥ν2−ε\lVert\sigma_{n}\rVert_{\nu_{n}}^{2}>\lVert\sigma\rVert_{\nu}^{2}-\varepsilon for n≥Nn\ge N: by (0.2)(a), 0≤∥σn∥νn2−2Kn(σn,σ)+∥σ∥ν20\le\lVert\sigma_{n}\rVert_{\nu_{n}}^{2}-2\mathcal{K}_{n}(\sigma_{n},\sigma)+\lVert\sigma\rVert_{\nu}^{2}, so ∥σn∥νn2≥2Kn(σn,σ)−∥σ∥ν2\lVert\sigma_{n}\rVert_{\nu_{n}}^{2}\ge2\mathcal{K}_{n}(\sigma_{n},\sigma)-\lVert\sigma\rVert_{\nu}^{2}; the right side converges to 2∥σ∥ν2−∥σ∥ν2=∥σ∥ν22\lVert\sigma\rVert_{\nu}^{2}-\lVert\sigma\rVert_{\nu}^{2}=\lVert\sigma\rVert_{\nu}^{2} by weak convergence with w=σw=\sigma, so it exceeds ∥σ∥ν2−ε\lVert\sigma\rVert_{\nu}^{2}-\varepsilon for large nn.

(4.4) The lower shift. By (0b), Fδ−(ξn)=un+vnF^{-}_{\delta}(\xi_{n})=u_{n}+v_{n} and Fδ−(ξ)=u+vF^{-}_{\delta}(\xi)=u+v, where

un=λ0rn+θ2∥qn∥νn2+(1+θδ)⟨σn,qn⟩νn−g(νn),vn=λ0δ E(νn)+(δ+θδ22)∥σn∥νn2,u_{n}=\lambda_{0}r_{n}+\frac{\theta}{2}\lVert q_{n}\rVert_{\nu_{n}}^{2}+(1+\theta\delta)\langle\sigma_{n},q_{n}\rangle_{\nu_{n}}-g(\nu_{n}),\qquad v_{n}=\lambda_{0}\delta\,\mathcal{E}(\nu_{n})+\Bigl(\delta+\frac{\theta\delta^{2}}{2}\Bigr)\lVert\sigma_{n}\rVert_{\nu_{n}}^{2},

and u,vu,v are the same expressions at ξ\xi. By (4.2) and the limit laws, un→uu_{n}\to u. Let k=λ0δ+δ+θδ22+1>0k=\lambda_{0}\delta+\delta+\tfrac{\theta\delta^{2}}{2}+1>0. Given ε>0\varepsilon>0, (4.3) applied with εk−1\varepsilon k^{-1} gives, for large nn, vn≥v−(λ0δ+δ+θδ22)εk−1≥v−εv_{n}\ge v-(\lambda_{0}\delta+\delta+\tfrac{\theta\delta^{2}}{2})\varepsilon k^{-1}\ge v-\varepsilon, the coefficients being nonnegative by (0.5). Now let c∈Rc\in\mathbb{R} be such that for every ε>0\varepsilon>0 there is NN with Fδ−(ξn)≤c+εF^{-}_{\delta}(\xi_{n})\le c+\varepsilon for n≥Nn\ge N. Fix ε>0\varepsilon>0 and choose nn so large that Fδ−(ξn)≤c+ε3F^{-}_{\delta}(\xi_{n})\le c+\tfrac{\varepsilon}{3}, ∣un−u∣<ε3|u_{n}-u|<\tfrac{\varepsilon}{3} and vn≥v−ε3v_{n}\ge v-\tfrac{\varepsilon}{3} (the largest of three thresholds, each obtained with the positive number ε3\tfrac{\varepsilon}{3} in place of ε\varepsilon). Then

Fδ−(ξ)=u+v<un+ε3+vn+ε3=Fδ−(ξn)+2ε3≤c+ε.F^{-}_{\delta}(\xi)=u+v<u_{n}+\tfrac{\varepsilon}{3}+v_{n}+\tfrac{\varepsilon}{3}=F^{-}_{\delta}(\xi_{n})+\tfrac{2\varepsilon}{3}\le c+\varepsilon .

As ε>0\varepsilon>0 is arbitrary, Comparison of Real Numbers with Arbitrary Positive Slack §slack-above gives Fδ−(ξ)≤cF^{-}_{\delta}(\xi)\le c.

(4.5) The upper shift. By (0d), Fδ+(ξn)=un′−vn′F^{+}_{\delta}(\xi_{n})=u'_{n}-v'_{n} and Fδ+(ξ)=u′−v′F^{+}_{\delta}(\xi)=u'-v', where

un′=λ0rn+θ2∥qn∥νn2+(1−θδ)⟨σn,qn⟩νn−g(νn),vn′=λ0δ E(νn)+(δ−θδ22)∥σn∥νn2,u'_{n}=\lambda_{0}r_{n}+\frac{\theta}{2}\lVert q_{n}\rVert_{\nu_{n}}^{2}+(1-\theta\delta)\langle\sigma_{n},q_{n}\rangle_{\nu_{n}}-g(\nu_{n}),\qquad v'_{n}=\lambda_{0}\delta\,\mathcal{E}(\nu_{n})+\Bigl(\delta-\frac{\theta\delta^{2}}{2}\Bigr)\lVert\sigma_{n}\rVert_{\nu_{n}}^{2},

and u′,v′u',v' are the same expressions at ξ\xi. As in (4.4), un′→u′u'_{n}\to u', and, the coefficients λ0δ\lambda_{0}\delta and δ−θδ22\delta-\tfrac{\theta\delta^{2}}{2} being nonnegative by (0.5), for every ε>0\varepsilon>0 we have vn′≥v′−εv'_{n}\ge v'-\varepsilon for large nn. Let c∈Rc\in\mathbb{R} be such that for every ε>0\varepsilon>0 there is NN with c−ε≤Fδ+(ξn)c-\varepsilon\le F^{+}_{\delta}(\xi_{n}) for n≥Nn\ge N. Fix ε>0\varepsilon>0 and choose nn so large that c−ε3≤Fδ+(ξn)c-\tfrac{\varepsilon}{3}\le F^{+}_{\delta}(\xi_{n}), ∣un′−u′∣<ε3|u'_{n}-u'|<\tfrac{\varepsilon}{3} and vn′≥v′−ε3v'_{n}\ge v'-\tfrac{\varepsilon}{3} (each threshold obtained with ε3\tfrac{\varepsilon}{3} in place of ε\varepsilon). Then

c−ε3≤un′−vn′<u′+ε3−v′+ε3=Fδ+(ξ)+2ε3,c-\tfrac{\varepsilon}{3}\le u'_{n}-v'_{n}<u'+\tfrac{\varepsilon}{3}-v'+\tfrac{\varepsilon}{3}=F^{+}_{\delta}(\xi)+\tfrac{2\varepsilon}{3},

so c−ε≤Fδ+(ξ)c-\varepsilon\le F^{+}_{\delta}(\xi); as ε>0\varepsilon>0 is arbitrary, Comparison of Real Numbers with Arbitrary Positive Slack §slack-below gives c≤Fδ+(ξ)c\le F^{+}_{\delta}(\xi).

By (4.4) and (4.5), FF is shift-semicontinuous at (δ,R)(\delta,R); as δ,R\delta,R were arbitrary, FF satisfies the shift-semicontinuity condition.

Step 5 (First-order structure at uniquely noise-mapped pairs, claim 4). Let T={t∈R:0≤t}T=\{t\in\mathbb{R}:0\le t\}. The pair (ω1,ω2)(\omega_{1},\omega_{2}) below is chosen first; it depends only on g,λ0,θ,Cg,\lambda_{0},\theta,C, and we show it is a structure pair for FF at RR for every positive RR.

(5.1) The modulus ω1\omega_{1}. For s∈Ts\in T let G(s)={∣g(μ′)−g(ν′)∣:μ′,ν′∈D, Wa(μ′,ν′)2≤s}G(s)=\{|g(\mu')-g(\nu')|:\mu',\nu'\in\mathcal{D},\ W_{a}(\mu',\nu')^{2}\le s\}. It contains 0=∣g(μ0)−g(μ0)∣0=|g(\mu_{0})-g(\mu_{0})|, with μ0\mu_{0} from (0.4), since Wa(μ0,μ0)=0W_{a}(\mu_{0},\mu_{0})=0 by (0.1); and it is bounded above by 2M2M. Hence ω1(s)=sup⁡G(s)\omega_{1}(s)=\sup G(s) exists by The Real Numbers: Standing Notation and Background §bounds, and 0≤ω1(s)0\le\omega_{1}(s) because ω1(s)\omega_{1}(s) is an upper bound of G(s)∋0G(s)\ni0 (Upper Bound and Least Upper Bound). Given ε>0\varepsilon>0, (Running cost) gives γ>0\gamma>0 with ∣g(μ′)−g(ν′)∣<ε|g(\mu')-g(\nu')|<\varepsilon whenever μ′,ν′∈D\mu',\nu'\in\mathcal{D} and Wa(μ′,ν′)<γW_{a}(\mu',\nu')<\gamma. Put γ1=γ2/4>0\gamma_{1}=\gamma^{2}/4>0. If t∈Tt\in T and t≤γ1t\le\gamma_{1}, every element of G(t)G(t) comes from μ′,ν′∈D\mu',\nu'\in\mathcal{D} with Wa(μ′,ν′)2≤(γ/2)2W_{a}(\mu',\nu')^{2}\le(\gamma/2)^{2}, so Wa(μ′,ν′)≤γ/2<γW_{a}(\mu',\nu')\le\gamma/2<\gamma and the element is <ε<\varepsilon; thus ε\varepsilon is an upper bound of G(t)G(t) and ω1(t)≤ε\omega_{1}(t)\le\varepsilon, the supremum being the least upper bound (Upper Bound and Least Upper Bound). So ω1\omega_{1} is a modulus of continuity, and by construction ∣g(μ′)−g(ν′)∣≤ω1(s)|g(\mu')-g(\nu')|\le\omega_{1}(s) whenever μ′,ν′∈D\mu',\nu'\in\mathcal{D}, s∈Ts\in T and Wa(μ′,ν′)2≤sW_{a}(\mu',\nu')^{2}\le s.

(5.2) The function ω2\omega_{2}. For t∈Tt\in T and real α>1\alpha>1 put ω2(t,α)=(λ0+4Cθ2α2) t\omega_{2}(t,\alpha)=(\lambda_{0}+4C\theta^{2}\alpha^{2})\,t. For each α>1\alpha>1 the coefficient is nonnegative by (0.4), so t↦ω2(t,α)t\mapsto\omega_{2}(t,\alpha) is a modulus of continuity by Linear Moduli of Continuity §modulus.

(5.3) The inequality. Let R>0R>0, and let α,δ,μ,ν,S,S′,r\alpha,\delta,\mu,\nu,S,S',r be as in The First-Order Structure Condition at Uniquely Noise-Mapped Pairs §pair: 1<α1<\alpha, 0<δ<10<\delta<1, μ,ν∈DΣ\mu,\nu\in\mathcal{D}_{\Sigma} with both ordered pairs (μ,ν)(\mu,\nu) and (ν,μ)(\nu,\mu) uniquely noise-mapped, SS a noise-optimal map from μ\mu to ν\nu and S′S' one from ν\nu to μ\mu, and −R≤r≤R-R\le r\le R. (Neither the unique mapping nor the condition δ(∣E(μ)∣+∣E(ν)∣)≤R\delta(|\mathcal{E}(\mu)|+|\mathcal{E}(\nu)|)\le R will be needed.) Write W=Wa(μ,ν)W=W_{a}(\mu,\nu), σ=Σ(μ)∈L2(μ;Xa)\sigma=\Sigma(\mu)\in L^{2}(\mu;X^{a}), τ=Σ(ν)∈L2(ν;Xa)\tau=\Sigma(\nu)\in L^{2}(\nu;X^{a}), p=α(id−S)=−α(S−id)∈L2(μ;Xa)p=\alpha(\mathrm{id}-S)=-\alpha(S-\mathrm{id})\in L^{2}(\mu;X^{a}), p′=α(S′−id)∈L2(ν;Xa)p'=\alpha(S'-\mathrm{id})\in L^{2}(\nu;X^{a}), e=∣E(μ)∣+∣E(ν)∣e=|\mathcal{E}(\mu)|+|\mathcal{E}(\nu)| and t=δ(e+1)t=\delta(e+1), and let Δ=Fδ−(μ,r,p)−Fδ+(ν,r,p′)\Delta=F^{-}_{\delta}(\mu,r,p)-F^{+}_{\delta}(\nu,r,p') be the difference to be bounded below.

Norms of the displacements. By (0.3), ∥S−id∥μ=W\lVert S-\mathrm{id}\rVert_{\mu}=W and ∥S′−id∥ν=Wa(ν,μ)=W\lVert S'-\mathrm{id}\rVert_{\nu}=W_{a}(\nu,\mu)=W (symmetry, (0.1)); hence, by homogeneity (0.1) with ∣α∣=α|\alpha|=\alpha and ∣−α∣=α|-\alpha|=\alpha, ∥p∥μ=∥p′∥ν=αW\lVert p\rVert_{\mu}=\lVert p'\rVert_{\nu}=\alpha W.

A bound on W2W^{2}. The measures μ,ν,ρ\mu,\nu,\rho lie in Pρa\mathcal{P}^{a}_{\rho}, so by the triangle inequality and symmetry (0.1), 0≤W≤Wa(μ,ρ)+Wa(ν,ρ)0\le W\le W_{a}(\mu,\rho)+W_{a}(\nu,\rho). Hence W2≤(Wa(μ,ρ)+Wa(ν,ρ))2≤2Wa(μ,ρ)2+2Wa(ν,ρ)2W^{2}\le(W_{a}(\mu,\rho)+W_{a}(\nu,\rho))^{2}\le2W_{a}(\mu,\rho)^{2}+2W_{a}(\nu,\rho)^{2}, the last step by (0.5) with m=Wa(μ,ρ)m=W_{a}(\mu,\rho), s=Wa(ν,ρ)s=W_{a}(\nu,\rho). As μ,ν∈D\mu,\nu\in\mathcal{D}, (Growth) yields W2≤2C(1+∣E(μ)∣)+2C(1+∣E(ν)∣)=2C(2+e)W^{2}\le2C(1+|\mathcal{E}(\mu)|)+2C(1+|\mathcal{E}(\nu)|)=2C(2+e). Since 0≤C0\le C (0.4) and 2+e≤2(1+e)2+e\le2(1+e), we get δW2≤2Cδ(2+e)≤4Cδ(1+e)=4Ct\delta W^{2}\le2C\delta(2+e)\le4C\delta(1+e)=4Ct.

Expansion of Δ\Delta. Subtracting (0c) at (ν,r,p′)(\nu,r,p') from (0a) at (μ,r,p)(\mu,r,p), the terms λ0r\lambda_{0}r cancel and

Δ=λ0δ(E(μ)+E(ν))+θ2(∥p+δσ∥μ2−∥p′−δτ∥ν2)+(⟨σ,p⟩μ−⟨τ,p′⟩ν)+δ∥σ∥μ2+δ∥τ∥ν2+g(ν)−g(μ).\Delta=\lambda_{0}\delta\bigl(\mathcal{E}(\mu)+\mathcal{E}(\nu)\bigr)+\frac{\theta}{2}\Bigl(\lVert p+\delta\sigma\rVert_{\mu}^{2}-\lVert p'-\delta\tau\rVert_{\nu}^{2}\Bigr)+\bigl(\langle\sigma,p\rangle_{\mu}-\langle\tau,p'\rangle_{\nu}\bigr)+\delta\lVert\sigma\rVert_{\mu}^{2}+\delta\lVert\tau\rVert_{\nu}^{2}+g(\nu)-g(\mu).

By (0.1), ∥p+δσ∥μ2=∥p∥μ2+2δ⟨p,σ⟩μ+δ2∥σ∥μ2\lVert p+\delta\sigma\rVert_{\mu}^{2}=\lVert p\rVert_{\mu}^{2}+2\delta\langle p,\sigma\rangle_{\mu}+\delta^{2}\lVert\sigma\rVert_{\mu}^{2} and ∥p′−δτ∥ν2=∥p′∥ν2−2δ⟨p′,τ⟩ν+δ2∥τ∥ν2\lVert p'-\delta\tau\rVert_{\nu}^{2}=\lVert p'\rVert_{\nu}^{2}-2\delta\langle p',\tau\rangle_{\nu}+\delta^{2}\lVert\tau\rVert_{\nu}^{2}; as ∥p∥μ2=∥p′∥ν2=α2W2\lVert p\rVert_{\mu}^{2}=\lVert p'\rVert_{\nu}^{2}=\alpha^{2}W^{2},

θ2(∥p+δσ∥μ2−∥p′−δτ∥ν2)=θδ⟨p,σ⟩μ+θδ⟨p′,τ⟩ν+θδ22∥σ∥μ2−θδ22∥τ∥ν2.\frac{\theta}{2}\Bigl(\lVert p+\delta\sigma\rVert_{\mu}^{2}-\lVert p'-\delta\tau\rVert_{\nu}^{2}\Bigr)=\theta\delta\langle p,\sigma\rangle_{\mu}+\theta\delta\langle p',\tau\rangle_{\nu}+\frac{\theta\delta^{2}}{2}\lVert\sigma\rVert_{\mu}^{2}-\frac{\theta\delta^{2}}{2}\lVert\tau\rVert_{\nu}^{2}.

Bounds for the individual terms. (i) Monotonicity of the score. The pair is displacement convex by (Convexity), that is, 00-displacement convex in the sense of Lambda-Displacement Convexity of a Noise Penalty Pair §convex. By (0.3) the couplings πS=(id,S)#μ∈Πa(μ,ν)\pi_{S}=(\mathrm{id},S)_{\#}\mu\in\Pi^{a}(\mu,\nu) and πS′=(id,S′)#ν∈Πa(ν,μ)\pi_{S'}=(\mathrm{id},S')_{\#}\nu\in\Pi^{a}(\nu,\mu) are noise-optimal, with Ja(σ,πS)=⟨σ,S−id⟩μ\mathcal{J}^{a}(\sigma,\pi_{S})=\langle\sigma,S-\mathrm{id}\rangle_{\mu} and Ja(τ,πS′)=⟨τ,S′−id⟩ν\mathcal{J}^{a}(\tau,\pi_{S'})=\langle\tau,S'-\mathrm{id}\rangle_{\nu}. Since μ,ν∈DΣ⊆D\mu,\nu\in\mathcal{D}_{\Sigma}\subseteq\mathcal{D}, the convexity inequality of Lambda-Displacement Convexity of a Noise Penalty Pair §convex with λ=0\lambda=0 applies to μ∈DΣ\mu\in\mathcal{D}_{\Sigma}, ν∈D\nu\in\mathcal{D} and πS\pi_{S}, and to ν∈DΣ\nu\in\mathcal{D}_{\Sigma}, μ∈D\mu\in\mathcal{D} and πS′\pi_{S'}; the cost terms carry the factor 02=0\tfrac{0}{2}=0, so

E(μ)+⟨σ,S−id⟩μ≤E(ν),E(ν)+⟨τ,S′−id⟩ν≤E(μ).\mathcal{E}(\mu)+\langle\sigma,S-\mathrm{id}\rangle_{\mu}\le\mathcal{E}(\nu),\qquad\mathcal{E}(\nu)+\langle\tau,S'-\mathrm{id}\rangle_{\nu}\le\mathcal{E}(\mu).

Adding, the real numbers E(μ)\mathcal{E}(\mu) and E(ν)\mathcal{E}(\nu) cancel and ⟨σ,S−id⟩μ+⟨τ,S′−id⟩ν≤0\langle\sigma,S-\mathrm{id}\rangle_{\mu}+\langle\tau,S'-\mathrm{id}\rangle_{\nu}\le0. By bilinearity (0.1), ⟨σ,p⟩μ−⟨τ,p′⟩ν=−α(⟨σ,S−id⟩μ+⟨τ,S′−id⟩ν)≥0\langle\sigma,p\rangle_{\mu}-\langle\tau,p'\rangle_{\nu}=-\alpha\bigl(\langle\sigma,S-\mathrm{id}\rangle_{\mu}+\langle\tau,S'-\mathrm{id}\rangle_{\nu}\bigr)\ge0. (ii) Cross terms: by Cauchy-Schwarz (0.1) and (0.5) with m=∥σ∥μm=\lVert\sigma\rVert_{\mu}, s=θαWs=\theta\alpha W,

θδ⟨p,σ⟩μ≥−θδ αW∥σ∥μ=−δ2 2ms≥−δ2∥σ∥μ2−δ2θ2α2W2,\theta\delta\langle p,\sigma\rangle_{\mu}\ge-\theta\delta\,\alpha W\lVert\sigma\rVert_{\mu}=-\frac{\delta}{2}\,2ms\ge-\frac{\delta}{2}\lVert\sigma\rVert_{\mu}^{2}-\frac{\delta}{2}\theta^{2}\alpha^{2}W^{2},

and in the same way θδ⟨p′,τ⟩ν≥−δ2∥τ∥ν2−δ2θ2α2W2\theta\delta\langle p',\tau\rangle_{\nu}\ge-\tfrac{\delta}{2}\lVert\tau\rVert_{\nu}^{2}-\tfrac{\delta}{2}\theta^{2}\alpha^{2}W^{2}. (iii) Collecting the score terms: the coefficient of ∥σ∥μ2\lVert\sigma\rVert_{\mu}^{2} becomes δ+θδ22−δ2=δ2+θδ22\delta+\tfrac{\theta\delta^{2}}{2}-\tfrac{\delta}{2}=\tfrac{\delta}{2}+\tfrac{\theta\delta^{2}}{2} and that of ∥τ∥ν2\lVert\tau\rVert_{\nu}^{2} becomes δ−θδ22−δ2=δ2(1−θδ)\delta-\tfrac{\theta\delta^{2}}{2}-\tfrac{\delta}{2}=\tfrac{\delta}{2}(1-\theta\delta), both nonnegative by (0.5) (the second one is where θ≤1\theta\le1 enters), so these terms are ≥0\ge0; the remaining contribution is −θ2α2δW2≥−4Cθ2α2t-\theta^{2}\alpha^{2}\delta W^{2}\ge-4C\theta^{2}\alpha^{2}t. (iv) Penalty terms: λ0δ(E(μ)+E(ν))≥−λ0δe≥−λ0t\lambda_{0}\delta(\mathcal{E}(\mu)+\mathcal{E}(\nu))\ge-\lambda_{0}\delta e\ge-\lambda_{0}t. (v) Running cost: W2≤αW2≤αW2+α−1W^{2}\le\alpha W^{2}\le\alpha W^{2}+\alpha^{-1}, as 1<α1<\alpha, 0≤W20\le W^{2} and 0<α−10<\alpha^{-1}; so (5.1) with s=αW2+α−1s=\alpha W^{2}+\alpha^{-1}, applied to μ,ν∈D\mu,\nu\in\mathcal{D}, gives g(ν)−g(μ)≥−∣g(μ)−g(ν)∣≥−ω1(αW2+α−1)g(\nu)-g(\mu)\ge-|g(\mu)-g(\nu)|\ge-\omega_{1}(\alpha W^{2}+\alpha^{-1}).

Adding (i)-(v) to the expansion of Δ\Delta,

−ω1(αWa(μ,ν)2+α−1)−ω2(δ(∣E(μ)∣+∣E(ν)∣+1),α)≤Fδ−(μ,r,α(id−S))−Fδ+(ν,r,α(S′−id)).-\omega_{1}\bigl(\alpha W_{a}(\mu,\nu)^{2}+\alpha^{-1}\bigr)-\omega_{2}\bigl(\delta(|\mathcal{E}(\mu)|+|\mathcal{E}(\nu)|+1),\alpha\bigr)\le F^{-}_{\delta}\bigl(\mu,r,\alpha(\mathrm{id}-S)\bigr)-F^{+}_{\delta}\bigl(\nu,r,\alpha(S'-\mathrm{id})\bigr).

Hence (ω1,ω2)(\omega_{1},\omega_{2}) is a structure pair for FF at every R>0R>0, and FF satisfies the first-order structure condition at uniquely noise-mapped pairs.

Step 6 (Momentum-continuous shifts, claim 5). A Lipschitz bound. Let δ,R∈R\delta,R\in\mathbb{R} with 0<δ<10<\delta<1 and 0<R0<R, let ν∈DΣ\nu\in\mathcal{D}_{\Sigma} with ∥Σ(ν)∥ν≤R\lVert\Sigma(\nu)\rVert_{\nu}\le R, let r∈Rr\in\mathbb{R}, and let q,q′∈L2(ν;Xa)q,q'\in L^{2}(\nu;X^{a}) with ∥q∥ν≤R\lVert q\rVert_{\nu}\le R and ∥q′∥ν≤R\lVert q'\rVert_{\nu}\le R; write σ=Σ(ν)\sigma=\Sigma(\nu). For w,w′∈L2(ν;Xa)w,w'\in L^{2}(\nu;X^{a}), bilinearity and symmetry (0.1) give ⟨w−w′,w+w′⟩ν=∥w∥ν2−∥w′∥ν2\langle w-w',w+w'\rangle_{\nu}=\lVert w\rVert_{\nu}^{2}-\lVert w'\rVert_{\nu}^{2}. With w=q+δσw=q+\delta\sigma and w′=q′+δσw'=q'+\delta\sigma, so that w−w′=q−q′w-w'=q-q' and w+w′=q+q′+2δσw+w'=q+q'+2\delta\sigma, the formula (0a) shows that the terms λ0r+λ0δE(ν)\lambda_{0}r+\lambda_{0}\delta\mathcal{E}(\nu), δ∥σ∥ν2\delta\lVert\sigma\rVert_{\nu}^{2} and −g(ν)-g(\nu) cancel in the difference, and

Fδ−(ν,r,q)−Fδ−(ν,r,q′)=θ2(∥w∥ν2−∥w′∥ν2)+⟨σ,q−q′⟩ν=θ2 ⟨q−q′, q+q′+2δσ⟩ν+⟨σ,q−q′⟩ν.F^{-}_{\delta}(\nu,r,q)-F^{-}_{\delta}(\nu,r,q')=\frac{\theta}{2}\bigl(\lVert w\rVert_{\nu}^{2}-\lVert w'\rVert_{\nu}^{2}\bigr)+\langle\sigma,q-q'\rangle_{\nu}=\frac{\theta}{2}\,\langle q-q',\,q+q'+2\delta\sigma\rangle_{\nu}+\langle\sigma,q-q'\rangle_{\nu}.

By the triangle inequality and homogeneity (0.1), ∥q+q′+2δσ∥ν≤2R+2δR=2(1+δ)R\lVert q+q'+2\delta\sigma\rVert_{\nu}\le2R+2\delta R=2(1+\delta)R. Taking absolute values, with the triangle inequality for the absolute value and Cauchy-Schwarz (0.1) for each inner product, and using θ>0\theta>0 and ∥σ∥ν≤R\lVert\sigma\rVert_{\nu}\le R,

∣Fδ−(ν,r,q)−Fδ−(ν,r,q′)∣≤(θ(1+δ)+1)R ∥q−q′∥ν≤(2θ+1)R ∥q−q′∥ν,\bigl|F^{-}_{\delta}(\nu,r,q)-F^{-}_{\delta}(\nu,r,q')\bigr|\le\bigl(\theta(1+\delta)+1\bigr)R\,\lVert q-q'\rVert_{\nu}\le(2\theta+1)R\,\lVert q-q'\rVert_{\nu},

since δ<1\delta<1. In the same way, from (0c) with w=q−δσw=q-\delta\sigma and w′=q′−δσw'=q'-\delta\sigma, Fδ+(ν,r,q)−Fδ+(ν,r,q′)=θ2⟨q−q′,q+q′−2δσ⟩ν+⟨σ,q−q′⟩νF^{+}_{\delta}(\nu,r,q)-F^{+}_{\delta}(\nu,r,q')=\tfrac{\theta}{2}\langle q-q',q+q'-2\delta\sigma\rangle_{\nu}+\langle\sigma,q-q'\rangle_{\nu}, ∥q+q′−2δσ∥ν≤2(1+δ)R\lVert q+q'-2\delta\sigma\rVert_{\nu}\le2(1+\delta)R, and ∣Fδ+(ν,r,q)−Fδ+(ν,r,q′)∣≤(2θ+1)R ∥q−q′∥ν|F^{+}_{\delta}(\nu,r,q)-F^{+}_{\delta}(\nu,r,q')|\le(2\theta+1)R\,\lVert q-q'\rVert_{\nu}.

The condition. Let δ,R,η∈R\delta,R,\eta\in\mathbb{R} with 0<δ<10<\delta<1, 0<R0<R and 0<η0<\eta be given; then (2θ+1)R(2\theta+1)R is positive, and we choose κ=η ((2θ+1)R)−1>0\kappa=\eta\,\bigl((2\theta+1)R\bigr)^{-1}>0. For ν∈DΣ\nu\in\mathcal{D}_{\Sigma} with ∥Σ(ν)∥ν≤R\lVert\Sigma(\nu)\rVert_{\nu}\le R, r∈Rr\in\mathbb{R} with ∣r∣≤R|r|\le R, and q,q′∈L2(ν;Xa)q,q'\in L^{2}(\nu;X^{a}) with ∥q∥ν≤R\lVert q\rVert_{\nu}\le R, ∥q′∥ν≤R\lVert q'\rVert_{\nu}\le R and ∥q−q′∥ν<κ\lVert q-q'\rVert_{\nu}<\kappa, the Lipschitz bound gives ∣Fδ∓(ν,r,q)−Fδ∓(ν,r,q′)∣≤(2θ+1)R∥q−q′∥ν<(2θ+1)R κ=η|F^{\mp}_{\delta}(\nu,r,q)-F^{\mp}_{\delta}(\nu,r,q')|\le(2\theta+1)R\lVert q-q'\rVert_{\nu}<(2\theta+1)R\,\kappa=\eta for both shifts. Hence FF has momentum-continuous shifts relative to the pair. None of the four assumptions, nor θ≤1\theta\le1, is used here.

Step 7 (Claim 6). Steps 1, 2, 4, 5 and 6 prove claims 1 to 5, and claim 6 collects them. ■\blacksquare

Citations

Loading…

Dependencies

Uses0

Loading…

Comments

Log in to comment.

Loading…