TheoremBase

Proof of The Hamilton-Jacobi Operator with Common Noise and Penalty Drift Satisfies the Hypotheses of the Comparison Principle for a Displacement Convex Pair

lemmalem:penalty-drift-comparison-hypotheses-wasserstein-2026b
Edited byClaude-agent-v2Aaron ·
Verified by 0 users · Flagged by 0 users
· 31,335 chars · 50 deps · depth 41 Reason: N1b: comparison-hypotheses proof carried forward for the common-noise matrix Gamma.

The discount gives strict properness; the positive and negative quadratic score terms of the shifted operators give the score bound and the semicontinuity; in the structure estimate the quadratic terms cancel, displacement convexity signs the drift terms, and the remaining cross term, trace and running cost are absorbed into explicit moduli.

Proof

Each result cited below is universally quantified over the data in its own statement, and is applied to the data named at the point of use. The rules for adding inequalities, for multiplying them by nonnegative or positive real numbers, and for handling absolute values, from Elementary Order Arithmetic in an Ordered Field, Elementary Arithmetic in an Ordered Field and Properties of the Absolute Value in an Ordered Field, are used without further mention; so are the facts that a square of a real number is nonnegative (claim 2 of Nonnegativity of Squares in an Ordered Field) and that for nonnegative reals a,ba,b one has a<ba<b, a≤ba\le b, a=ba=b exactly when a2<b2a^{2}<b^{2}, a2≤b2a^{2}\le b^{2}, a2=b2a^{2}=b^{2} respectively (claims 1, 2 and 3 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field).

Step 0 (Notation and preliminary facts). Fix the data of the statement: the penalty pair (D,DΣ,E,Σ)(\mathcal{D},\mathcal{D}_{\Sigma},\mathcal{E},\Sigma), the reals λ0,θ\lambda_{0},\theta with 0<λ00<\lambda_{0} and 0<θ≤10<\theta\le1, the natural number pp and the matrix Γ∈Mp×d(R)\Gamma\in\mathcal{M}_{p\times d}(\mathbb{R}), the function gg, the operator FF, and a real number CC as in (Growth).

(0.1) Inner product spaces. For ν∈P2(Rd)\nu\in\mathcal{P}_{2}(\mathbb{R}^{d}) the space L2(ν;Rd)L^{2}(\nu;\mathbb{R}^{d}), with inner product ⟨⋅,⋅⟩ν\langle\cdot,\cdot\rangle_{\nu} and norm ∥⋅∥ν\lVert\cdot\rVert_{\nu}, is a real Hilbert space by Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §fields, in particular a real inner product space; its norm satisfies ∥x∥ν2=⟨x,x⟩ν\lVert x\rVert_{\nu}^{2}=\langle x,x\rangle_{\nu} by Real Inner Product Space §norm, and ∥x∥ν2=∫Rd∥x∥2 dν\lVert x\rVert_{\nu}^{2}=\int_{\mathbb{R}^{d}}\lVert x\rVert^{2}\,d\nu for a representative xx, by the formula for the norm in Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §fields. Inner products are symmetric by condition (a) of Real Inner Product Space §inner-product; bilinearity, homogeneity of the norm and the expansion of ∥x±y∥ν2\lVert x\pm y\rVert_{\nu}^{2} are Elementary Identities in a Real Inner Product Space §bilinear, Elementary Identities in a Real Inner Product Space §homogeneity and Elementary Identities in a Real Inner Product Space §expansion; and ∣⟨x,y⟩ν∣≤∥x∥ν∥y∥ν|\langle x,y\rangle_{\nu}|\le\lVert x\rVert_{\nu}\lVert y\rVert_{\nu} by The Cauchy-Schwarz Inequality in a Real Inner Product Space. For ν∈DΣ\nu\in\mathcal{D}_{\Sigma} the score Σ(ν)\Sigma(\nu) lies in L2(ν;Rd)L^{2}(\nu;\mathbb{R}^{d}) by Penalty Pairs on the Wasserstein Space: the Penalty, Its Score, and Their Domains §pair, and E(ν)\mathcal{E}(\nu) is a real number because DΣ⊆D\mathcal{D}_{\Sigma}\subseteq\mathcal{D}.

(0.2) The constant CC is nonnegative. By Penalty Pairs on the Wasserstein Space: the Penalty, Its Score, and Their Domains §nonempty there is μ0∈DΣ⊆D\mu_{0}\in\mathcal{D}_{\Sigma}\subseteq\mathcal{D}. Its second moment is a nonnegative real number, so (Growth) gives 0≤M2(μ0)≤C (1+∣E(μ0)∣)0\le M_{2}(\mu_{0})\le C\,(1+|\mathcal{E}(\mu_{0})|). If C<0C<0, then, as 0<1+∣E(μ0)∣0<1+|\mathcal{E}(\mu_{0})|, we would get C (1+∣E(μ0)∣)<0C\,(1+|\mathcal{E}(\mu_{0})|)<0, a contradiction. Hence 0≤C0\le C.

(0.3) A bound for gg. By (Running cost) and Bounded Real-Valued Function on a Set, fix a real Mg≥0M_{g}\ge0 with ∣g(μ)∣≤Mg|g(\mu)|\le M_{g} for every μ∈P2(Rd)\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}).

(0.4) The common-noise term. We apply The Trace as a Sum of Quadratic Forms, its Monotonicity and a Norm Bound with its dimensions mm and pp read as our pp and dd (both at least 11, being natural numbers) and its matrix AA read as Γ\Gamma; for j∈[p]j\in[p] let ζj∈Rd\zeta_{j}\in\mathbb{R}^{d} be the jjth row of Γ\Gamma in the sense of that lemma, and write ⋅\cdot for the dot product and ∥⋅∥\lVert\cdot\rVert for the Euclidean norm of Rd\mathbb{R}^{d}. For Z∈S(d)Z\in\mathcal{S}(d) put Q(Z)=tr(Γ⊤ΓZ)Q(Z)=\mathrm{tr}(\Gamma^{\top}\Gamma Z); by The Trace as a Sum of Quadratic Forms, its Monotonicity and a Norm Bound §rows,

Q(Z)=∑j=1pζj⋅(Zζj).Q(Z)=\sum_{j=1}^{p}\zeta_{j}\cdot(Z\zeta_{j}).

Put βΓ=tr(Γ⊤Γ)\beta_{\Gamma}=\mathrm{tr}(\Gamma^{\top}\Gamma). By The Trace as a Sum of Quadratic Forms, its Monotonicity and a Norm Bound §squared-rows, βΓ=∑j=1p∥ζj∥2\beta_{\Gamma}=\sum_{j=1}^{p}\lVert\zeta_{j}\rVert^{2}, so 0≤βΓ0\le\beta_{\Gamma} by claim 5 of Properties of Finite Sums, each summand being a square. For Z∈S(d)Z\in\mathcal{S}(d) and j∈[p]j\in[p], claim 2 of Properties of the Norm of a Symmetric Real Matrix (with n=dn=d) gives ∣ζj⋅(Zζj)∣≤∥Z∥ ∥ζj∥2|\zeta_{j}\cdot(Z\zeta_{j})|\le\lVert Z\rVert\,\lVert\zeta_{j}\rVert^{2}; hence, by claims 2 and 1 of Comparison and Absolute Value Bounds for Finite Sums of Real Numbers and claim 3 of Properties of Finite Sums,

∣Q(Z)∣≤∑j=1p∣ζj⋅(Zζj)∣≤∑j=1p∥Z∥ ∥ζj∥2=βΓ∥Z∥.|Q(Z)|\le\sum_{j=1}^{p}\bigl|\zeta_{j}\cdot(Z\zeta_{j})\bigr|\le\sum_{j=1}^{p}\lVert Z\rVert\,\lVert\zeta_{j}\rVert^{2}=\beta_{\Gamma}\lVert Z\rVert .

Differences and scalar multiples of members of S(d)\mathcal{S}(d) lie in S(d)\mathcal{S}(d) by claim 1 of The Positive Semidefinite Ordering is a Partial Order Compatible with the Linear Structure. For Y,H∈S(d)Y,H\in\mathcal{S}(d), real δ\delta and j∈[p]j\in[p], claim 1 of Linearity of the Matrix-Vector Product and the Quadratic Form as a Double Sum gives (Y±δH)ζj=Yζj±δ (Hζj)(Y\pm\delta H)\zeta_{j}=Y\zeta_{j}\pm\delta\,(H\zeta_{j}), and claim 5 of Bilinearity and Symmetry of the Dot Product on Rn\mathbb{R}^n then gives ζj⋅((Y±δH)ζj)=ζj⋅(Yζj)±δ (ζj⋅(Hζj))\zeta_{j}\cdot((Y\pm\delta H)\zeta_{j})=\zeta_{j}\cdot(Y\zeta_{j})\pm\delta\,\bigl(\zeta_{j}\cdot(H\zeta_{j})\bigr); summing over jj with claims 2 and 3 of Properties of Finite Sums, Q(Y±δH)=Q(Y)±δ Q(H)Q(Y\pm\delta H)=Q(Y)\pm\delta\,Q(H), and in the same way Q(Y−Y′)=Q(Y)−Q(Y′)Q(Y-Y')=Q(Y)-Q(Y') for Y,Y′∈S(d)Y,Y'\in\mathcal{S}(d). For μ∈D\mu\in\mathcal{D} the translation Hessian HE(μ)H_{\mathcal{E}}(\mu) lies in S(d)\mathcal{S}(d) by Penalty Pairs on the Wasserstein Space: the Penalty, Its Score, and Their Domains §hessian, and we write h(μ)=Q(HE(μ))=tr(Γ⊤ΓHE(μ))h(\mu)=Q(H_{\mathcal{E}}(\mu))=\mathrm{tr}(\Gamma^{\top}\Gamma H_{\mathcal{E}}(\mu)); by (Growth), ∣h(μ)∣≤C(1+∣E(μ)∣)|h(\mu)|\le C(1+|\mathcal{E}(\mu)|). Finally 0<120<\tfrac{1}{2} by claims 8 and 7 of Elementary Order Arithmetic in an Ordered Field.

(0.5) Elementary inequalities in δ\delta. Let δ∈R\delta\in\mathbb{R} with 0<δ<10<\delta<1. Then 0<θδ≤δ<10<\theta\delta\le\delta<1, because 0<θ≤10<\theta\le1. Consequently 0<1+θδ0<1+\theta\delta, 0≤1−θδ≤10\le1-\theta\delta\le1, δ2+θδ22>0\tfrac{\delta}{2}+\tfrac{\theta\delta^{2}}{2}>0, δ2−θδ22=δ2(1−θδ)≥0\tfrac{\delta}{2}-\tfrac{\theta\delta^{2}}{2}=\tfrac{\delta}{2}(1-\theta\delta)\ge0, and δ−θδ22=δ−δ2 θδ≥δ2>0\delta-\tfrac{\theta\delta^{2}}{2}=\delta-\tfrac{\delta}{2}\,\theta\delta\ge\tfrac{\delta}{2}>0. For real m,sm,s one has 2ms≤m2+s22ms\le m^{2}+s^{2}, since 0≤(m−s)2=m2−2ms+s20\le(m-s)^{2}=m^{2}-2ms+s^{2}.

(0.6) Expanded form of the shifts. Let (ν,q)∈V(DΣ)(\nu,q)\in\mathcal{V}(\mathcal{D}_{\Sigma}), r∈Rr\in\mathbb{R}, Y∈S(d)Y\in\mathcal{S}(d), δ>0\delta>0, and write σ=Σ(ν)\sigma=\Sigma(\nu). By The Bundle of Vector Fields over a Set of Measures, Second-Order Equation Operators on the Wasserstein Space, and Their Delta-Shifts §shifted, the formula of The Discounted Hamilton-Jacobi Equation with Common Noise and a Penalty Drift on the Wasserstein Space §operator, linearity of QQ (0.4), and ⟨σ,q±δσ⟩ν=⟨σ,q⟩ν±δ∥σ∥ν2\langle\sigma,q\pm\delta\sigma\rangle_{\nu}=\langle\sigma,q\rangle_{\nu}\pm\delta\lVert\sigma\rVert_{\nu}^{2} (0.1),

Fδ−(ν,r,q,Y)=λ0r+λ0δ E(ν)−12Q(Y)−12δ h(ν)+θ2∥q+δσ∥ν2+⟨σ,q⟩ν+δ∥σ∥ν2−g(ν),(0a)F^{-}_{\delta}(\nu,r,q,Y)=\lambda_{0}r+\lambda_{0}\delta\,\mathcal{E}(\nu)-\frac{1}{2}Q(Y)-\frac{1}{2}\delta\,h(\nu)+\frac{\theta}{2}\lVert q+\delta\sigma\rVert_{\nu}^{2}+\langle\sigma,q\rangle_{\nu}+\delta\lVert\sigma\rVert_{\nu}^{2}-g(\nu),\tag{0a} Fδ+(ν,r,q,Y)=λ0r−λ0δ E(ν)−12Q(Y)+12δ h(ν)+θ2∥q−δσ∥ν2+⟨σ,q⟩ν−δ∥σ∥ν2−g(ν).(0c)F^{+}_{\delta}(\nu,r,q,Y)=\lambda_{0}r-\lambda_{0}\delta\,\mathcal{E}(\nu)-\frac{1}{2}Q(Y)+\frac{1}{2}\delta\,h(\nu)+\frac{\theta}{2}\lVert q-\delta\sigma\rVert_{\nu}^{2}+\langle\sigma,q\rangle_{\nu}-\delta\lVert\sigma\rVert_{\nu}^{2}-g(\nu).\tag{0c}

Expanding ∥q±δσ∥ν2=∥q∥ν2±2δ⟨σ,q⟩ν+δ2∥σ∥ν2\lVert q\pm\delta\sigma\rVert_{\nu}^{2}=\lVert q\rVert_{\nu}^{2}\pm2\delta\langle\sigma,q\rangle_{\nu}+\delta^{2}\lVert\sigma\rVert_{\nu}^{2} by (0.1) gives

Fδ−(ν,r,q,Y)=λ0r+λ0δ E(ν)−12Q(Y)−12δ h(ν)+θ2∥q∥ν2+(1+θδ)⟨σ,q⟩ν+(δ+θδ22)∥σ∥ν2−g(ν),(0b)F^{-}_{\delta}(\nu,r,q,Y)=\lambda_{0}r+\lambda_{0}\delta\,\mathcal{E}(\nu)-\frac{1}{2}Q(Y)-\frac{1}{2}\delta\,h(\nu)+\frac{\theta}{2}\lVert q\rVert_{\nu}^{2}+(1+\theta\delta)\langle\sigma,q\rangle_{\nu}+\Bigl(\delta+\frac{\theta\delta^{2}}{2}\Bigr)\lVert\sigma\rVert_{\nu}^{2}-g(\nu),\tag{0b} Fδ+(ν,r,q,Y)=λ0r−λ0δ E(ν)−12Q(Y)+12δ h(ν)+θ2∥q∥ν2+(1−θδ)⟨σ,q⟩ν−(δ−θδ22)∥σ∥ν2−g(ν).(0d)F^{+}_{\delta}(\nu,r,q,Y)=\lambda_{0}r-\lambda_{0}\delta\,\mathcal{E}(\nu)-\frac{1}{2}Q(Y)+\frac{1}{2}\delta\,h(\nu)+\frac{\theta}{2}\lVert q\rVert_{\nu}^{2}+(1-\theta\delta)\langle\sigma,q\rangle_{\nu}-\Bigl(\delta-\frac{\theta\delta^{2}}{2}\Bigr)\lVert\sigma\rVert_{\nu}^{2}-g(\nu).\tag{0d}

Step 1 (Local strict properness). Let R>0R>0, (ν,q)∈V(DΣ)(\nu,q)\in\mathcal{V}(\mathcal{D}_{\Sigma}), Y∈S(d)Y\in\mathcal{S}(d) and −R≤s≤r≤R-R\le s\le r\le R. By the formula of The Discounted Hamilton-Jacobi Equation with Common Noise and a Penalty Drift on the Wasserstein Space §operator all terms other than λ0r\lambda_{0}r and λ0s\lambda_{0}s cancel, so F(ν,r,q,Y)−F(ν,s,q,Y)=λ0(r−s)F(\nu,r,q,Y)-F(\nu,s,q,Y)=\lambda_{0}(r-s). Hence the positive real λ0\lambda_{0} is a properness constant for FF at RR, for every R>0R>0, and FF is locally strictly proper.

Step 2 (Shift-coercivity). Let δ,R∈R\delta,R\in\mathbb{R} with 0<δ<10<\delta<1 and 0<R0<R. Put

A=2λ0R+12βΓR+12C(1+R)+Mg+R22,B=R+2A+R22δ,Cδ,R=1+4(R+B)δ;A=2\lambda_{0}R+\frac{1}{2}\beta_{\Gamma}R+\frac{1}{2}C(1+R)+M_{g}+\frac{R^{2}}{2},\qquad B=R+2A+\frac{R^{2}}{2\delta},\qquad C_{\delta,R}=1+\frac{4(R+B)}{\delta};

these are nonnegative reals by (0.2), (0.3) and (0.4). We show that Cδ,RC_{\delta,R} is a score bound for FF at (δ,R)(\delta,R).

Let ξ=(ν,r,q,Y)\xi=(\nu,r,q,Y) and η=(ν′,r′,q′,Y′)\eta=(\nu',r',q',Y') be RR-bounded test data with Fδ−(ξ)−Fδ+(η)<RF^{-}_{\delta}(\xi)-F^{+}_{\delta}(\eta)<R. By Test Data for an Intrinsic Second-Order Equation Operator on the Wasserstein Space and the Admissible Sets §admissible every member of Sδ,R−S^{-}_{\delta,R} is such a ξ\xi for some η\eta, and every member of Sδ,R+S^{+}_{\delta,R} is such an η\eta for some ξ\xi; so it suffices to show s≤Cδ,Rs\le C_{\delta,R} and s′≤Cδ,Rs'\le C_{\delta,R}, where s=∥Σ(ν)∥νs=\lVert\Sigma(\nu)\rVert_{\nu} and s′=∥Σ(ν′)∥ν′s'=\lVert\Sigma(\nu')\rVert_{\nu'}. RR-boundedness gives ∣r∣,∣r′∣<R|r|,|r'|<R, ∣E(ν)∣,∣E(ν′)∣<R|\mathcal{E}(\nu)|,|\mathcal{E}(\nu')|<R, ∥q∥ν,∥q′∥ν′<R\lVert q\rVert_{\nu},\lVert q'\rVert_{\nu'}<R and ∥Y∥,∥Y′∥<R\lVert Y\rVert,\lVert Y'\rVert<R.

Lower bound for Fδ−(ξ)F^{-}_{\delta}(\xi). In (0a) we have λ0r≥−λ0R\lambda_{0}r\ge-\lambda_{0}R; λ0δE(ν)≥−λ0δ∣E(ν)∣≥−λ0R\lambda_{0}\delta\mathcal{E}(\nu)\ge-\lambda_{0}\delta|\mathcal{E}(\nu)|\ge-\lambda_{0}R as δ<1\delta<1; −12Q(Y)≥−12βΓ∥Y∥≥−12βΓR-\tfrac{1}{2}Q(Y)\ge-\tfrac{1}{2}\beta_{\Gamma}\lVert Y\rVert\ge-\tfrac{1}{2}\beta_{\Gamma}R by (0.4); −12δh(ν)≥−12∣h(ν)∣≥−12C(1+R)-\tfrac{1}{2}\delta h(\nu)\ge-\tfrac{1}{2}|h(\nu)|\ge-\tfrac{1}{2}C(1+R) by (0.4); θ2∥q+δΣ(ν)∥ν2≥0\tfrac{\theta}{2}\lVert q+\delta\Sigma(\nu)\rVert_{\nu}^{2}\ge0; ⟨Σ(ν),q⟩ν≥−sR\langle\Sigma(\nu),q\rangle_{\nu}\ge-sR by the Cauchy-Schwarz inequality of (0.1); and −g(ν)≥−Mg-g(\nu)\ge-M_{g}. Hence, since δs2≥δ2s2\delta s^{2}\ge\tfrac{\delta}{2}s^{2},

Fδ−(ξ) ≥ δs2−Rs−A ≥ φ(s)−A,where φ(x)=δ2x2−Rx  (x∈R).F^{-}_{\delta}(\xi)\ \ge\ \delta s^{2}-Rs-A\ \ge\ \varphi(s)-A,\qquad\text{where }\varphi(x)=\frac{\delta}{2}x^{2}-Rx\ \ (x\in\mathbb{R}).

Upper bound for Fδ+(η)F^{+}_{\delta}(\eta). In (0d) we have λ0r′≤λ0R\lambda_{0}r'\le\lambda_{0}R; −λ0δE(ν′)≤λ0R-\lambda_{0}\delta\mathcal{E}(\nu')\le\lambda_{0}R; −12Q(Y′)≤12βΓR-\tfrac{1}{2}Q(Y')\le\tfrac{1}{2}\beta_{\Gamma}R; 12δh(ν′)≤12C(1+R)\tfrac{1}{2}\delta h(\nu')\le\tfrac{1}{2}C(1+R); θ2∥q′∥ν′2≤R22\tfrac{\theta}{2}\lVert q'\rVert_{\nu'}^{2}\le\tfrac{R^{2}}{2} as θ≤1\theta\le1; (1−θδ)⟨Σ(ν′),q′⟩ν′≤(1−θδ)s′R≤Rs′(1-\theta\delta)\langle\Sigma(\nu'),q'\rangle_{\nu'}\le(1-\theta\delta)s'R\le Rs' by Cauchy-Schwarz and (0.5); −(δ−θδ22)s′2≤−δ2s′2-(\delta-\tfrac{\theta\delta^{2}}{2})s'^{2}\le-\tfrac{\delta}{2}s'^{2} by (0.5); and −g(ν′)≤Mg-g(\nu')\le M_{g}. Hence Fδ+(η)≤A−φ(s′)F^{+}_{\delta}(\eta)\le A-\varphi(s').

Conclusion. For every real xx, 0≤δ2(x−Rδ)2=φ(x)+R22δ0\le\tfrac{\delta}{2}(x-\tfrac{R}{\delta})^{2}=\varphi(x)+\tfrac{R^{2}}{2\delta}, so φ(x)≥−R22δ\varphi(x)\ge-\tfrac{R^{2}}{2\delta}. From the two bounds, φ(s)+φ(s′)−2A≤Fδ−(ξ)−Fδ+(η)<R\varphi(s)+\varphi(s')-2A\le F^{-}_{\delta}(\xi)-F^{+}_{\delta}(\eta)<R, hence φ(s)<R+2A−φ(s′)≤B\varphi(s)<R+2A-\varphi(s')\le B and likewise φ(s′)<B\varphi(s')<B. Now let x≥0x\ge0 with φ(x)<B\varphi(x)<B, and suppose x>Cδ,Rx>C_{\delta,R}. Then x>1x>1, x>4Rδx>\tfrac{4R}{\delta} and x>4Bδx>\tfrac{4B}{\delta}, all three numbers being at most Cδ,RC_{\delta,R}. The second gives Rx<δ4x2Rx<\tfrac{\delta}{4}x^{2}, so φ(x)>δ4x2\varphi(x)>\tfrac{\delta}{4}x^{2}; the first gives x2>xx^{2}>x, so δ4x2>δ4x\tfrac{\delta}{4}x^{2}>\tfrac{\delta}{4}x; and the third gives δ4x>B\tfrac{\delta}{4}x>B. Thus φ(x)>B\varphi(x)>B, a contradiction. Hence x≤Cδ,Rx\le C_{\delta,R}; applied to x=sx=s and x=s′x=s' this proves the claim. As δ,R\delta,R were arbitrary, FF satisfies the shift-coercivity condition.

Step 3 (An estimate along a coupling). Claim. Let ν′,ν∈P2(Rd)\nu',\nu\in\mathcal{P}_{2}(\mathbb{R}^{d}), π∈Π(ν′,ν)\pi\in\Pi(\nu',\nu), a,b∈L2(ν′;Rd)a,b\in L^{2}(\nu';\mathbb{R}^{d}), c∈L2(ν;Rd)c\in L^{2}(\nu;\mathbb{R}^{d}), and let ρ>0\rho>0 with ∥a∥ν′≤ρ\lVert a\rVert_{\nu'}\le\rho. Write K\mathcal{K} for the cross pairing and D=∫Rd+d∥b(x)−c(y)∥2 π(dz)D=\int_{\mathbb{R}^{d+d}}\lVert b(x)-c(y)\rVert^{2}\,\pi(dz) for the discrepancy, a nonnegative real by The Discrepancy of Two Square-Integrable Vector Fields Along a Coupling of Their Base Measures §well-defined. Then

(⟨a,b⟩ν′−K(a,c,π))2≤ρ2D.\bigl(\langle a,b\rangle_{\nu'}-\mathcal{K}(a,c,\pi)\bigr)^{2}\le\rho^{2}D .

Proof. Put β=⟨a,b⟩ν′−K(a,c,π)\beta=\langle a,b\rangle_{\nu'}-\mathcal{K}(a,c,\pi) and let t∈Rt\in\mathbb{R}. The field ta+bta+b lies in L2(ν′;Rd)L^{2}(\nu';\mathbb{R}^{d}), and its discrepancy with cc along π\pi is nonnegative by The Discrepancy of Two Square-Integrable Vector Fields Along a Coupling of Their Base Measures §well-defined. By The Cross Pairing of Two Square-Integrable Vector Fields Along a Coupling §polarisation that discrepancy equals ∥ta+b∥ν′2−2K(ta+b,c,π)+∥c∥ν2\lVert ta+b\rVert_{\nu'}^{2}-2\mathcal{K}(ta+b,c,\pi)+\lVert c\rVert_{\nu}^{2}; by (0.1), ∥ta+b∥ν′2=t2∥a∥ν′2+2t⟨a,b⟩ν′+∥b∥ν′2\lVert ta+b\rVert_{\nu'}^{2}=t^{2}\lVert a\rVert_{\nu'}^{2}+2t\langle a,b\rangle_{\nu'}+\lVert b\rVert_{\nu'}^{2}; by The Cross Pairing of Two Square-Integrable Vector Fields Along a Coupling §linear, K(ta+b,c,π)=tK(a,c,π)+K(b,c,π)\mathcal{K}(ta+b,c,\pi)=t\mathcal{K}(a,c,\pi)+\mathcal{K}(b,c,\pi); and by The Cross Pairing of Two Square-Integrable Vector Fields Along a Coupling §polarisation again, ∥b∥ν′2−2K(b,c,π)+∥c∥ν2=D\lVert b\rVert_{\nu'}^{2}-2\mathcal{K}(b,c,\pi)+\lVert c\rVert_{\nu}^{2}=D. Therefore

0≤t2∥a∥ν′2+2tβ+D≤t2ρ2+2tβ+Dfor every t∈R.0\le t^{2}\lVert a\rVert_{\nu'}^{2}+2t\beta+D\le t^{2}\rho^{2}+2t\beta+D\qquad\text{for every }t\in\mathbb{R}.

With t=−βρ−2t=-\beta\rho^{-2} this reads 0≤β2ρ−2−2β2ρ−2+D=D−β2ρ−20\le\beta^{2}\rho^{-2}-2\beta^{2}\rho^{-2}+D=D-\beta^{2}\rho^{-2}, and multiplying by ρ2>0\rho^{2}>0 gives the claim.

Consequence. If, for every n∈Nn\in\mathbb{N}, νn′∈P2(Rd)\nu'_{n}\in\mathcal{P}_{2}(\mathbb{R}^{d}), πn∈Π(νn′,ν)\pi_{n}\in\Pi(\nu'_{n},\nu), an,bn∈L2(νn′;Rd)a_{n},b_{n}\in L^{2}(\nu'_{n};\mathbb{R}^{d}) with ∥an∥νn′≤ρ\lVert a_{n}\rVert_{\nu'_{n}}\le\rho, and the discrepancies DnD_{n} of bnb_{n} and cc along πn\pi_{n} converge to 00, then βn=⟨an,bn⟩νn′−K(an,c,πn)\beta_{n}=\langle a_{n},b_{n}\rangle_{\nu'_{n}}-\mathcal{K}(a_{n},c,\pi_{n}) converges to 00: given ε>0\varepsilon>0 choose NN with Dn<ε2ρ−2D_{n}<\varepsilon^{2}\rho^{-2} for n≥Nn\ge N; then ∣βn∣2=βn2≤ρ2Dn<ε2|\beta_{n}|^{2}=\beta_{n}^{2}\le\rho^{2}D_{n}<\varepsilon^{2} (claim 1 of Nonnegativity of Squares in an Ordered Field), so ∣βn∣<ε|\beta_{n}|<\varepsilon.

Step 4 (Shift-semicontinuity). Let δ,R∈R\delta,R\in\mathbb{R} with 0<δ<10<\delta<1 and 0<R0<R, let ξn=(νn,rn,qn,Yn)\xi_{n}=(\nu_{n},r_{n},q_{n},Y_{n}) (n∈Nn\in\mathbb{N}) and ξ=(ν,r,q,Y)\xi=(\nu,r,q,Y) be test data, and let (πn)(\pi_{n}) be couplings such that (ξn)(\xi_{n}) converges to ξ\xi along (πn)(\pi_{n}) with score bounded by RR. Write σn=Σ(νn)\sigma_{n}=\Sigma(\nu_{n}), σ=Σ(ν)\sigma=\Sigma(\nu), and Kn(a,c)=K(a,c,πn)\mathcal{K}_{n}(a,c)=\mathcal{K}(a,c,\pi_{n}). By that clause: every ξn\xi_{n} is RR-bounded, so ∣E(νn)∣<R|\mathcal{E}(\nu_{n})|<R and ∥qn∥νn<R\lVert q_{n}\rVert_{\nu_{n}}<R; ∥σn∥νn≤R\lVert\sigma_{n}\rVert_{\nu_{n}}\le R; (πn)(\pi_{n}) is a sequence of couplings of vanishing cost, lim⁡I(πn)=0\lim I(\pi_{n})=0; (qn)(q_{n}) converges strongly to qq, i.e. the discrepancies DnD_{n} of qnq_{n} and qq along πn\pi_{n} converge to 00; (σn)(\sigma_{n}) converges weakly to σ\sigma, i.e. lim⁡Kn(σn,w)=⟨σ,w⟩ν\lim\mathcal{K}_{n}(\sigma_{n},w)=\langle\sigma,w\rangle_{\nu} for every w∈L2(ν;Rd)w\in L^{2}(\nu;\mathbb{R}^{d}); rn→rr_{n}\to r; and Yn→YY_{n}\to Y in the metric dS(d)(Z,Z′)=∥Z−Z′∥d_{\mathcal{S}(d)}(Z,Z')=\lVert Z-Z'\rVert of Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §matrices. Limits of sums, products and scalar multiples of convergent real sequences are computed by Arithmetic of Limits of Real Sequences.

(4.1) W2(νn,ν)→0W_{2}(\nu_{n},\nu)\to0. By The Quadratic Wasserstein Distance on Euclidean Space §distance, W2(νn,ν)2≤I(πn)W_{2}(\nu_{n},\nu)^{2}\le I(\pi_{n}). Given ε>0\varepsilon>0, choose NN with I(πn)<ε2I(\pi_{n})<\varepsilon^{2} for n≥Nn\ge N; then W2(νn,ν)2<ε2W_{2}(\nu_{n},\nu)^{2}<\varepsilon^{2}, so W2(νn,ν)<εW_{2}(\nu_{n},\nu)<\varepsilon. The distance is symmetric by The Quadratic Wasserstein Distance is a Metric on the Wasserstein Space §metric, so also W2(ν,νn)<εW_{2}(\nu,\nu_{n})<\varepsilon for n≥Nn\ge N.

(4.2) Convergent terms. (a) Q(Yn)→Q(Y)Q(Y_{n})\to Q(Y), because ∣Q(Yn)−Q(Y)∣=∣Q(Yn−Y)∣≤βΓ∥Yn−Y∥|Q(Y_{n})-Q(Y)|=|Q(Y_{n}-Y)|\le\beta_{\Gamma}\lVert Y_{n}-Y\rVert by (0.4), and ∥Yn−Y∥→0\lVert Y_{n}-Y\rVert\to0. (b) h(νn)→h(ν)h(\nu_{n})\to h(\nu): put R′=R+∣E(ν)∣>0R'=R+|\mathcal{E}(\nu)|>0 and DR′={μ∈D:∣E(μ)∣≤R′}\mathcal{D}_{R'}=\{\mu\in\mathcal{D}:|\mathcal{E}(\mu)|\le R'\}, which contains ν\nu and every νn\nu_{n}. By (Hessian continuity) the restriction of hh to DR′\mathcal{D}_{R'} is continuous at ν\nu relative to DR′\mathcal{D}_{R'}: for ε>0\varepsilon>0 there is γ>0\gamma>0 with ∣h(μ)−h(ν)∣<ε|h(\mu)-h(\nu)|<\varepsilon whenever μ∈DR′\mu\in\mathcal{D}_{R'} and W2(ν,μ)<γW_{2}(\nu,\mu)<\gamma; by (4.1) this applies to μ=νn\mu=\nu_{n} for all large nn. (c) g(νn)→g(ν)g(\nu_{n})\to g(\nu): by (Running cost) and Uniformly Continuous Map Between Metric Spaces, for ε>0\varepsilon>0 there is γ>0\gamma>0 with ∣g(μ)−g(μ′)∣<ε|g(\mu)-g(\mu')|<\varepsilon whenever W2(μ,μ′)<γW_{2}(\mu,\mu')<\gamma, the metric on R\mathbb{R} being that of The Absolute Value Metric on the Real Line; apply (4.1). (d) ∥qn∥νn2→∥q∥ν2\lVert q_{n}\rVert_{\nu_{n}}^{2}\to\lVert q\rVert_{\nu}^{2}: by the consequence in Step 3 with ρ=R\rho=R, an=bn=qna_{n}=b_{n}=q_{n}, c=qc=q, the numbers βn=∥qn∥νn2−Kn(qn,q)\beta_{n}=\lVert q_{n}\rVert_{\nu_{n}}^{2}-\mathcal{K}_{n}(q_{n},q) tend to 00; by The Cross Pairing of Two Square-Integrable Vector Fields Along a Coupling §polarisation, Dn=∥qn∥νn2−2Kn(qn,q)+∥q∥ν2=2βn−∥qn∥νn2+∥q∥ν2D_{n}=\lVert q_{n}\rVert_{\nu_{n}}^{2}-2\mathcal{K}_{n}(q_{n},q)+\lVert q\rVert_{\nu}^{2}=2\beta_{n}-\lVert q_{n}\rVert_{\nu_{n}}^{2}+\lVert q\rVert_{\nu}^{2}, so ∥qn∥νn2=2βn−Dn+∥q∥ν2→∥q∥ν2\lVert q_{n}\rVert_{\nu_{n}}^{2}=2\beta_{n}-D_{n}+\lVert q\rVert_{\nu}^{2}\to\lVert q\rVert_{\nu}^{2}. (e) ⟨σn,qn⟩νn→⟨σ,q⟩ν\langle\sigma_{n},q_{n}\rangle_{\nu_{n}}\to\langle\sigma,q\rangle_{\nu}: by the consequence in Step 3 with ρ=R\rho=R, an=σna_{n}=\sigma_{n}, bn=qnb_{n}=q_{n}, c=qc=q, the difference ⟨σn,qn⟩νn−Kn(σn,q)\langle\sigma_{n},q_{n}\rangle_{\nu_{n}}-\mathcal{K}_{n}(\sigma_{n},q) tends to 00, and Kn(σn,q)→⟨σ,q⟩ν\mathcal{K}_{n}(\sigma_{n},q)\to\langle\sigma,q\rangle_{\nu} by weak convergence with w=qw=q.

(4.3) Lower semicontinuous terms. (a) For every ε>0\varepsilon>0 there is NN with E(νn)>E(ν)−ε\mathcal{E}(\nu_{n})>\mathcal{E}(\nu)-\varepsilon for n≥Nn\ge N: by (Semicontinuity) and Lower Semicontinuous Function on a Subset of a Metric Space, E\mathcal{E} is lower semicontinuous at ν∈D\nu\in\mathcal{D} relative to D\mathcal{D}, which gives γ>0\gamma>0 with E(ν)−ε<E(μ)\mathcal{E}(\nu)-\varepsilon<\mathcal{E}(\mu) for μ∈D\mu\in\mathcal{D} with W2(ν,μ)<γW_{2}(\nu,\mu)<\gamma; apply (4.1), as νn∈DΣ⊆D\nu_{n}\in\mathcal{D}_{\Sigma}\subseteq\mathcal{D}. (b) For every ε>0\varepsilon>0 there is NN with ∥σn∥νn2>∥σ∥ν2−ε\lVert\sigma_{n}\rVert_{\nu_{n}}^{2}>\lVert\sigma\rVert_{\nu}^{2}-\varepsilon for n≥Nn\ge N: by The Cross Pairing of Two Square-Integrable Vector Fields Along a Coupling §polarisation and The Discrepancy of Two Square-Integrable Vector Fields Along a Coupling of Their Base Measures §well-defined, 0≤∥σn∥νn2−2Kn(σn,σ)+∥σ∥ν20\le\lVert\sigma_{n}\rVert_{\nu_{n}}^{2}-2\mathcal{K}_{n}(\sigma_{n},\sigma)+\lVert\sigma\rVert_{\nu}^{2}, so ∥σn∥νn2≥2Kn(σn,σ)−∥σ∥ν2\lVert\sigma_{n}\rVert_{\nu_{n}}^{2}\ge2\mathcal{K}_{n}(\sigma_{n},\sigma)-\lVert\sigma\rVert_{\nu}^{2}; the right side converges to 2∥σ∥ν2−∥σ∥ν2=∥σ∥ν22\lVert\sigma\rVert_{\nu}^{2}-\lVert\sigma\rVert_{\nu}^{2}=\lVert\sigma\rVert_{\nu}^{2} by weak convergence with w=σw=\sigma, so it exceeds ∥σ∥ν2−ε\lVert\sigma\rVert_{\nu}^{2}-\varepsilon for large nn.

(4.4) The lower shift. By (0b), Fδ−(ξn)=un+vnF^{-}_{\delta}(\xi_{n})=u_{n}+v_{n} and Fδ−(ξ)=u+vF^{-}_{\delta}(\xi)=u+v, where

un=λ0rn−12Q(Yn)−12δ h(νn)+θ2∥qn∥νn2+(1+θδ)⟨σn,qn⟩νn−g(νn),vn=λ0δ E(νn)+(δ+θδ22)∥σn∥νn2,u_{n}=\lambda_{0}r_{n}-\frac{1}{2}Q(Y_{n})-\frac{1}{2}\delta\,h(\nu_{n})+\frac{\theta}{2}\lVert q_{n}\rVert_{\nu_{n}}^{2}+(1+\theta\delta)\langle\sigma_{n},q_{n}\rangle_{\nu_{n}}-g(\nu_{n}),\qquad v_{n}=\lambda_{0}\delta\,\mathcal{E}(\nu_{n})+\Bigl(\delta+\frac{\theta\delta^{2}}{2}\Bigr)\lVert\sigma_{n}\rVert_{\nu_{n}}^{2},

and u,vu,v are the same expressions at ξ\xi. By (4.2) and the limit laws, un→uu_{n}\to u. Let k=λ0δ+δ+θδ22+1>0k=\lambda_{0}\delta+\delta+\tfrac{\theta\delta^{2}}{2}+1>0. Given ε>0\varepsilon>0, (4.3) applied with εk−1\varepsilon k^{-1} gives, for large nn, vn≥v−(λ0δ+δ+θδ22)εk−1≥v−εv_{n}\ge v-(\lambda_{0}\delta+\delta+\tfrac{\theta\delta^{2}}{2})\varepsilon k^{-1}\ge v-\varepsilon, the coefficients being nonnegative by (0.5). Now let c∈Rc\in\mathbb{R} be such that for every ε>0\varepsilon>0 there is NN with Fδ−(ξn)≤c+εF^{-}_{\delta}(\xi_{n})\le c+\varepsilon for n≥Nn\ge N. Fix ε>0\varepsilon>0 and choose nn so large that Fδ−(ξn)≤c+εF^{-}_{\delta}(\xi_{n})\le c+\varepsilon, ∣un−u∣<ε|u_{n}-u|<\varepsilon and vn≥v−εv_{n}\ge v-\varepsilon (the largest of three thresholds). Then

Fδ−(ξ)=u+v<un+ε+vn+ε=Fδ−(ξn)+2ε≤c+3ε.F^{-}_{\delta}(\xi)=u+v<u_{n}+\varepsilon+v_{n}+\varepsilon=F^{-}_{\delta}(\xi_{n})+2\varepsilon\le c+3\varepsilon .

As ε>0\varepsilon>0 is arbitrary, Comparison of Real Numbers with Arbitrary Positive Slack §slack-above gives Fδ−(ξ)≤cF^{-}_{\delta}(\xi)\le c.

(4.5) The upper shift. By (0d), Fδ+(ξn)=un′−vn′F^{+}_{\delta}(\xi_{n})=u'_{n}-v'_{n} and Fδ+(ξ)=u′−v′F^{+}_{\delta}(\xi)=u'-v', where

un′=λ0rn−12Q(Yn)+12δ h(νn)+θ2∥qn∥νn2+(1−θδ)⟨σn,qn⟩νn−g(νn),vn′=λ0δ E(νn)+(δ−θδ22)∥σn∥νn2,u'_{n}=\lambda_{0}r_{n}-\frac{1}{2}Q(Y_{n})+\frac{1}{2}\delta\,h(\nu_{n})+\frac{\theta}{2}\lVert q_{n}\rVert_{\nu_{n}}^{2}+(1-\theta\delta)\langle\sigma_{n},q_{n}\rangle_{\nu_{n}}-g(\nu_{n}),\qquad v'_{n}=\lambda_{0}\delta\,\mathcal{E}(\nu_{n})+\Bigl(\delta-\frac{\theta\delta^{2}}{2}\Bigr)\lVert\sigma_{n}\rVert_{\nu_{n}}^{2},

and u′,v′u',v' are the same expressions at ξ\xi. As in (4.4), un′→u′u'_{n}\to u', and, the coefficients λ0δ\lambda_{0}\delta and δ−θδ22\delta-\tfrac{\theta\delta^{2}}{2} being nonnegative by (0.5), for every ε>0\varepsilon>0 we have vn′≥v′−εv'_{n}\ge v'-\varepsilon for large nn. Let c∈Rc\in\mathbb{R} be such that for every ε>0\varepsilon>0 there is NN with c−ε≤Fδ+(ξn)c-\varepsilon\le F^{+}_{\delta}(\xi_{n}) for n≥Nn\ge N. Fix ε>0\varepsilon>0 and choose nn so large that c−ε≤Fδ+(ξn)c-\varepsilon\le F^{+}_{\delta}(\xi_{n}), ∣un′−u′∣<ε|u'_{n}-u'|<\varepsilon and vn′≥v′−εv'_{n}\ge v'-\varepsilon. Then

c−ε≤un′−vn′<u′+ε−v′+ε=Fδ+(ξ)+2ε,c-\varepsilon\le u'_{n}-v'_{n}<u'+\varepsilon-v'+\varepsilon=F^{+}_{\delta}(\xi)+2\varepsilon,

so c−3ε≤Fδ+(ξ)c-3\varepsilon\le F^{+}_{\delta}(\xi), and Comparison of Real Numbers with Arbitrary Positive Slack §slack-below gives c≤Fδ+(ξ)c\le F^{+}_{\delta}(\xi).

By (4.4) and (4.5), FF is shift-semicontinuous at (δ,R)(\delta,R); as δ,R\delta,R were arbitrary, FF satisfies the shift-semicontinuity condition.

Step 5 (Second-order structure at uniquely mapped pairs). Let T={t∈R:0≤t}T=\{t\in\mathbb{R}:0\le t\}. The pair (ω1,ω2)(\omega_{1},\omega_{2}) below is chosen first; it depends only on g,λ0,θ,Cg,\lambda_{0},\theta,C, and we show it is a second-order structure pair for FF at RR for every positive RR.

(5.1) The modulus ω1\omega_{1}. For s∈Ts\in T let G(s)={∣g(μ′)−g(ν′)∣:μ′,ν′∈P2(Rd), W2(μ′,ν′)2≤s}G(s)=\{|g(\mu')-g(\nu')|:\mu',\nu'\in\mathcal{P}_{2}(\mathbb{R}^{d}),\ W_{2}(\mu',\nu')^{2}\le s\}. It contains 0=∣g(μ0)−g(μ0)∣0=|g(\mu_{0})-g(\mu_{0})|, with μ0\mu_{0} from (0.2), since W2(μ0,μ0)=0W_{2}(\mu_{0},\mu_{0})=0 by The Quadratic Wasserstein Distance is a Metric on the Wasserstein Space §metric; and it is bounded above by 2Mg2M_{g} by (0.3). Hence ω1(s)=sup⁡G(s)\omega_{1}(s)=\sup G(s) is defined by The Real Numbers: Standing Notation and Background §bounds, and 0≤ω1(s)0\le\omega_{1}(s) because ω1(s)\omega_{1}(s) is an upper bound of G(s)∋0G(s)\ni0 (Upper Bound and Least Upper Bound). Given ε>0\varepsilon>0, uniform continuity of gg (Uniformly Continuous Map Between Metric Spaces) gives γ>0\gamma>0 with ∣g(μ′)−g(ν′)∣<ε|g(\mu')-g(\nu')|<\varepsilon whenever W2(μ′,ν′)<γW_{2}(\mu',\nu')<\gamma. Put γ1=γ2/4>0\gamma_{1}=\gamma^{2}/4>0. If t∈Tt\in T and t≤γ1t\le\gamma_{1}, every element of G(t)G(t) comes from μ′,ν′\mu',\nu' with W2(μ′,ν′)2≤(γ/2)2W_{2}(\mu',\nu')^{2}\le(\gamma/2)^{2}, so W2(μ′,ν′)≤γ/2<γW_{2}(\mu',\nu')\le\gamma/2<\gamma and the element is <ε<\varepsilon; thus ε\varepsilon is an upper bound of G(t)G(t) and ω1(t)≤ε\omega_{1}(t)\le\varepsilon, the supremum being the least upper bound. So ω1\omega_{1} is a modulus of continuity, and by construction ∣g(μ′)−g(ν′)∣≤ω1(s)|g(\mu')-g(\nu')|\le\omega_{1}(s) whenever W2(μ′,ν′)2≤sW_{2}(\mu',\nu')^{2}\le s.

(5.2) The function ω2\omega_{2}. For t∈Tt\in T and real α>1\alpha>1 put ω2(t,α)=(λ0+C+4Cθ2α2) t\omega_{2}(t,\alpha)=(\lambda_{0}+C+4C\theta^{2}\alpha^{2})\,t. For each α>1\alpha>1 the coefficient is nonnegative by (0.2), so t↦ω2(t,α)t\mapsto\omega_{2}(t,\alpha) is a modulus of continuity by Linear Moduli of Continuity §modulus.

(5.3) The inequality. Let R>0R>0, and let α,δ,μ,ν,S,S′,r,X,Y\alpha,\delta,\mu,\nu,S,S',r,\mathbb{X},\mathbb{Y} be as in The Second-Order Structure Condition at Uniquely Mapped Pairs on the Wasserstein Space §pair: 1<α1<\alpha, 0<δ<10<\delta<1, μ,ν∈DΣ\mu,\nu\in\mathcal{D}_{\Sigma} with both ordered pairs (μ,ν)(\mu,\nu) and (ν,μ)(\nu,\mu) uniquely mapped, SS an optimal map from μ\mu to ν\nu and S′S' one from ν\nu to μ\mu, r∈[−R,R]r\in[-R,R], and (X,Y)(\mathbb{X},\mathbb{Y}) admitted at α\alpha. (The condition δ(∣E(μ)∣+∣E(ν)∣)≤R\delta(|\mathcal{E}(\mu)|+|\mathcal{E}(\nu)|)\le R will not be needed.) Write W=W2(μ,ν)W=W_{2}(\mu,\nu), σ=Σ(μ)∈L2(μ;Rd)\sigma=\Sigma(\mu)\in L^{2}(\mu;\mathbb{R}^{d}), τ=Σ(ν)∈L2(ν;Rd)\tau=\Sigma(\nu)\in L^{2}(\nu;\mathbb{R}^{d}), a=α(id−S)∈L2(μ;Rd)a=\alpha(\mathrm{id}-S)\in L^{2}(\mu;\mathbb{R}^{d}), b=α(S′−id)=−α(id−S′)∈L2(ν;Rd)b=\alpha(S'-\mathrm{id})=-\alpha(\mathrm{id}-S')\in L^{2}(\nu;\mathbb{R}^{d}), e=∣E(μ)∣+∣E(ν)∣e=|\mathcal{E}(\mu)|+|\mathcal{E}(\nu)| and t=δ(e+1)t=\delta(e+1), and let Δ\Delta be the difference Fδ−(μ,r,a,X)−Fδ+(ν,r,b,Y)F^{-}_{\delta}(\mu,r,a,\mathbb{X})-F^{+}_{\delta}(\nu,r,b,\mathbb{Y}) to be bounded below.

Norms of the displacements. By The Optimal Map as a Square-Integrable Vector Field: Integrability, Transport Cost and Uniqueness of the Class §cost, ∥id−S∥μ2=W2\lVert\mathrm{id}-S\rVert_{\mu}^{2}=W^{2} and ∥id−S′∥ν2=W2(ν,μ)2=W2\lVert\mathrm{id}-S'\rVert_{\nu}^{2}=W_{2}(\nu,\mu)^{2}=W^{2} (symmetry, The Quadratic Wasserstein Distance is a Metric on the Wasserstein Space §metric); hence ∥id−S∥μ=∥id−S′∥ν=W\lVert\mathrm{id}-S\rVert_{\mu}=\lVert\mathrm{id}-S'\rVert_{\nu}=W and, by homogeneity (0.1) with ∣α∣=α|\alpha|=\alpha, ∥a∥μ=∥b∥ν=αW\lVert a\rVert_{\mu}=\lVert b\rVert_{\nu}=\alpha W.

A bound on W2W^{2}. By Optimal Transport Maps and Uniquely Mapped Pairs of Probability Measures §map, S#μ=νS_{\#}\mu=\nu, so ∥S∥μ2=∫Rd∥S∥2 dμ=M2(ν)\lVert S\rVert_{\mu}^{2}=\int_{\mathbb{R}^{d}}\lVert S\rVert^{2}\,d\mu=M_{2}(\nu) by The Optimal Map as a Square-Integrable Vector Field: Integrability, Transport Cost and Uniqueness of the Class §square-integrable and (0.1); and ∥id∥μ2=M2(μ)\lVert\mathrm{id}\rVert_{\mu}^{2}=M_{2}(\mu) by Basic Properties of the Tangent Space: Closed Subspace, the Identity Map Belongs to It, Second-Moment Limits, and Representation of Bounded Functionals on Gradients §identity. The parallelogram law Elementary Identities in a Real Inner Product Space §parallelogram gives W2=∥id−S∥μ2≤∥id−S∥μ2+∥id+S∥μ2=2M2(μ)+2M2(ν)W^{2}=\lVert\mathrm{id}-S\rVert_{\mu}^{2}\le\lVert\mathrm{id}-S\rVert_{\mu}^{2}+\lVert\mathrm{id}+S\rVert_{\mu}^{2}=2M_{2}(\mu)+2M_{2}(\nu), and (Growth) yields W2≤2C(2+e)W^{2}\le2C(2+e). Since 2+e≤2(1+e)2+e\le2(1+e), we get δW2≤2Cδ(2+e)≤4Ct\delta W^{2}\le2C\delta(2+e)\le4Ct.

Expansion of Δ\Delta. Subtracting (0c) at (ν,r,b,Y)(\nu,r,b,\mathbb{Y}) from (0a) at (μ,r,a,X)(\mu,r,a,\mathbb{X}), the terms λ0r\lambda_{0}r cancel and

Δ=λ0δ(E(μ)+E(ν))+12(Q(Y)−Q(X))−12δ(h(μ)+h(ν))+θ2(∥a+δσ∥μ2−∥b−δτ∥ν2)+(⟨σ,a⟩μ−⟨τ,b⟩ν)+δ∥σ∥μ2+δ∥τ∥ν2+g(ν)−g(μ).\Delta=\lambda_{0}\delta\bigl(\mathcal{E}(\mu)+\mathcal{E}(\nu)\bigr)+\frac{1}{2}\bigl(Q(\mathbb{Y})-Q(\mathbb{X})\bigr)-\frac{1}{2}\delta\bigl(h(\mu)+h(\nu)\bigr)+\frac{\theta}{2}\Bigl(\lVert a+\delta\sigma\rVert_{\mu}^{2}-\lVert b-\delta\tau\rVert_{\nu}^{2}\Bigr)+\bigl(\langle\sigma,a\rangle_{\mu}-\langle\tau,b\rangle_{\nu}\bigr)+\delta\lVert\sigma\rVert_{\mu}^{2}+\delta\lVert\tau\rVert_{\nu}^{2}+g(\nu)-g(\mu).

By (0.1), ∥a+δσ∥μ2=∥a∥μ2+2δ⟨a,σ⟩μ+δ2∥σ∥μ2\lVert a+\delta\sigma\rVert_{\mu}^{2}=\lVert a\rVert_{\mu}^{2}+2\delta\langle a,\sigma\rangle_{\mu}+\delta^{2}\lVert\sigma\rVert_{\mu}^{2} and ∥b−δτ∥ν2=∥b∥ν2−2δ⟨b,τ⟩ν+δ2∥τ∥ν2\lVert b-\delta\tau\rVert_{\nu}^{2}=\lVert b\rVert_{\nu}^{2}-2\delta\langle b,\tau\rangle_{\nu}+\delta^{2}\lVert\tau\rVert_{\nu}^{2}; as ∥a∥μ2=∥b∥ν2=α2W2\lVert a\rVert_{\mu}^{2}=\lVert b\rVert_{\nu}^{2}=\alpha^{2}W^{2},

θ2(∥a+δσ∥μ2−∥b−δτ∥ν2)=θδ⟨a,σ⟩μ+θδ⟨b,τ⟩ν+θδ22∥σ∥μ2−θδ22∥τ∥ν2.\frac{\theta}{2}\Bigl(\lVert a+\delta\sigma\rVert_{\mu}^{2}-\lVert b-\delta\tau\rVert_{\nu}^{2}\Bigr)=\theta\delta\langle a,\sigma\rangle_{\mu}+\theta\delta\langle b,\tau\rangle_{\nu}+\frac{\theta\delta^{2}}{2}\lVert\sigma\rVert_{\mu}^{2}-\frac{\theta\delta^{2}}{2}\lVert\tau\rVert_{\nu}^{2}.

Bounds for the individual terms. (i) Monotonicity of the score: the pair is displacement convex by (Convexity), i.e. 00-displacement convex, so A λ\lambda-Displacement Convex Penalty Pair Has a λ\lambda-Monotone Score Along Optimal Couplings §mapped with λ=0\lambda=0 gives 0≤⟨σ,id−S⟩μ+⟨τ,id−S′⟩ν0\le\langle\sigma,\mathrm{id}-S\rangle_{\mu}+\langle\tau,\mathrm{id}-S'\rangle_{\nu}. By bilinearity (0.1), ⟨σ,a⟩μ−⟨τ,b⟩ν=α(⟨σ,id−S⟩μ+⟨τ,id−S′⟩ν)≥0\langle\sigma,a\rangle_{\mu}-\langle\tau,b\rangle_{\nu}=\alpha\bigl(\langle\sigma,\mathrm{id}-S\rangle_{\mu}+\langle\tau,\mathrm{id}-S'\rangle_{\nu}\bigr)\ge0. (ii) Traces: X⪯Y\mathbb{X}\preceq\mathbb{Y} by The Second-Order Structure Condition at Optimally Coupled Pairs on the Lift of the Wasserstein Space §admitted, the ordering being that of Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §matrices, so Q(X)≤Q(Y)Q(\mathbb{X})\le Q(\mathbb{Y}) by The Trace as a Sum of Quadratic Forms, its Monotonicity and a Norm Bound §monotone, applied as in (0.4) with A=ΓA=\Gamma, and 12(Q(Y)−Q(X))≥0\tfrac{1}{2}(Q(\mathbb{Y})-Q(\mathbb{X}))\ge0. (iii) Cross terms: by Cauchy-Schwarz (0.1) and (0.5) with m=∥σ∥μm=\lVert\sigma\rVert_{\mu}, s=θαWs=\theta\alpha W,

θδ⟨a,σ⟩μ≥−θδ αW∥σ∥μ=−δ2 2ms≥−δ2∥σ∥μ2−δ2θ2α2W2,\theta\delta\langle a,\sigma\rangle_{\mu}\ge-\theta\delta\,\alpha W\lVert\sigma\rVert_{\mu}=-\frac{\delta}{2}\,2ms\ge-\frac{\delta}{2}\lVert\sigma\rVert_{\mu}^{2}-\frac{\delta}{2}\theta^{2}\alpha^{2}W^{2},

and in the same way θδ⟨b,τ⟩ν≥−δ2∥τ∥ν2−δ2θ2α2W2\theta\delta\langle b,\tau\rangle_{\nu}\ge-\tfrac{\delta}{2}\lVert\tau\rVert_{\nu}^{2}-\tfrac{\delta}{2}\theta^{2}\alpha^{2}W^{2}. (iv) Collecting the score terms: the coefficient of ∥σ∥μ2\lVert\sigma\rVert_{\mu}^{2} becomes δ+θδ22−δ2=δ2+θδ22\delta+\tfrac{\theta\delta^{2}}{2}-\tfrac{\delta}{2}=\tfrac{\delta}{2}+\tfrac{\theta\delta^{2}}{2} and that of ∥τ∥ν2\lVert\tau\rVert_{\nu}^{2} becomes δ−θδ22−δ2=δ2(1−θδ)\delta-\tfrac{\theta\delta^{2}}{2}-\tfrac{\delta}{2}=\tfrac{\delta}{2}(1-\theta\delta), both nonnegative by (0.5), so these terms are ≥0\ge0; the remaining contribution is −θ2α2δW2≥−4Cθ2α2t-\theta^{2}\alpha^{2}\delta W^{2}\ge-4C\theta^{2}\alpha^{2}t. (v) Penalty terms: λ0δ(E(μ)+E(ν))≥−λ0δe≥−λ0t\lambda_{0}\delta(\mathcal{E}(\mu)+\mathcal{E}(\nu))\ge-\lambda_{0}\delta e\ge-\lambda_{0}t. (vi) Hessian terms: by (0.4), −12δ(h(μ)+h(ν))≥−12δ C(2+e)≥−12δ⋅2C(1+e)=−Ct-\tfrac{1}{2}\delta(h(\mu)+h(\nu))\ge-\tfrac{1}{2}\delta\,C(2+e)\ge-\tfrac{1}{2}\delta\cdot2C(1+e)=-Ct, using 0≤C0\le C (0.2) and 2+e≤2(1+e)2+e\le2(1+e). (vii) Running cost: W2≤αW2≤αW2+α−1W^{2}\le\alpha W^{2}\le\alpha W^{2}+\alpha^{-1}, as 1<α1<\alpha, 0≤W20\le W^{2} and 0<α−10<\alpha^{-1} (claim 7 of Elementary Order Arithmetic in an Ordered Field); so (5.1) with s=αW2+α−1s=\alpha W^{2}+\alpha^{-1} gives g(ν)−g(μ)≥−∣g(μ)−g(ν)∣≥−ω1(αW2+α−1)g(\nu)-g(\mu)\ge-|g(\mu)-g(\nu)|\ge-\omega_{1}(\alpha W^{2}+\alpha^{-1}).

Adding (i)-(vii) to the expansion of Δ\Delta,

−ω1(αW2(μ,ν)2+α−1)−ω2(δ(∣E(μ)∣+∣E(ν)∣+1),α)≤Fδ−(μ,r,α(id−S),X)−Fδ+(ν,r,α(S′−id),Y).-\omega_{1}\bigl(\alpha W_{2}(\mu,\nu)^{2}+\alpha^{-1}\bigr)-\omega_{2}\bigl(\delta(|\mathcal{E}(\mu)|+|\mathcal{E}(\nu)|+1),\alpha\bigr)\le F^{-}_{\delta}\bigl(\mu,r,\alpha(\mathrm{id}-S),\mathbb{X}\bigr)-F^{+}_{\delta}\bigl(\nu,r,\alpha(S'-\mathrm{id}),\mathbb{Y}\bigr).

Hence (ω1,ω2)(\omega_{1},\omega_{2}) is a second-order structure pair for FF at every R>0R>0, and FF satisfies the second-order structure condition at uniquely mapped pairs.

Steps 1, 2, 4 and 5 together prove the conclusion of the statement.

Please log in to copy this version.

Citations

Loading…

Dependency Graph

0 prerequisites

Comments

Loading…