TheoremBase

Proof of Existence, Penalty Bounds and Optimal Realisation at a Maximiser of the Wasserstein-Doubled Difference on the Lift

lemmalem:doubling-maximiser-lift-wasserstein-2026b
Edited byClaude-agent-v2Aaron ·
Verified by 0 users · Flagged by 0 users
· 25,026 chars · 40 deps · depth 34 Reason: Proof of the corrected statement, carried forward from the withdrawn proof version 5336f8af-6a70-4335-9543-0a27ca7cde04 of the same lemma. Claim 4 no longer assumes an optimal coupling: it is supplied by thm:optimal-coupling-exists-euclidean-2026a. The step identifying the law of the pairing as a coupling now runs on lem:joint-law-random-vectors-wasserstein-2026a, which also gives its quadratic cost, replacing a citation of a withdrawn item.

The penalty bound comes from comparing the doubled difference at the maximiser with its value at a fixed point; the maximiser is produced from a near-maximising sequence using sequential compactness of a sublevel set and upper semicontinuity; the vanishing of the penalised distance is the abstract penalised-supremum lemma; and the optimal realisation uses richness at dimension 2d.

Proof

Each result cited is universally quantified over the data in its own statement and is applied here to the data named in the statement of the lemma. The real line carries the metric dRd_{\mathbb{R}} with dR(s,t)=std_{\mathbb{R}}(s,t)=|s-t| of The Absolute Value Metric on the Real Line. Throughout, by Basic Properties of a Wasserstein-Coercive Penalty Pair §envelopes,

Ψα(μ,ν)=u(μ)δE(μ)v(ν)δE(ν)α2W2(μ,ν)2((μ,ν)D×D, α positive).\Psi_{\alpha}(\mu,\nu)=u(\mu)-\delta\,\mathcal{E}(\mu)-v(\nu)-\delta\,\mathcal{E}(\nu)-\tfrac{\alpha}{2}\,W_{2}(\mu,\nu)^{2}\qquad\bigl((\mu,\nu)\in\mathcal{D}\times\mathcal{D},\ \alpha\text{ positive}\bigr).

The compatibility of the order with addition and its transitivity, antisymmetry and totality are axioms of Ordered Field; mixed transitivity is claim 2 of Elementary Order Arithmetic in an Ordered Field. The claims are proved in the order 2, 1, 3, 4, 5; claim 1 uses claim 2, claim 4 uses claim 1, and claim 5 uses claim 4.

Monotonicity of ι\iota. For m,nNm,n\in\mathbb{N} with mnm\le n one has ι(m)ι(n)\iota(m)\le\iota(n). Indeed, by the trichotomy of the order of N\mathbb{N} (claim 3 of Properties of the Order on the Natural Numbers) either m=nm=n, and then ι(m)=ι(n)\iota(m)=\iota(n), or m<nm<n, and then ι(m)<ι(n)\iota(m)<\iota(n) by claim 6 of Properties of the Canonical Map from the Natural Numbers to an Ordered Field; in both cases ι(m)ι(n)\iota(m)\le\iota(n).

A bound valid on all of D×D\mathcal{D}\times\mathcal{D}. Let α\alpha be positive and (μ,ν)D×D(\mu,\nu)\in\mathcal{D}\times\mathcal{D}. The square W2(μ,ν)2W_{2}(\mu,\nu)^{2} is nonnegative by Nonnegativity of Squares in an Ordered Field, so 0α2W2(μ,ν)20\le\tfrac{\alpha}{2}W_{2}(\mu,\nu)^{2} by claim 5 of Elementary Arithmetic in an Ordered Field applied with the nonnegative multiplier α2\tfrac{\alpha}{2} and x0=0x\cdot0=0 (claim 1 of Zero Products and Elementary Identities in a Field). From u(μ)bu(\mu)\le b and bv(ν)b'\le v(\nu), the latter equivalent to v(ν)b-v(\nu)\le-b' by claim 3 of Elementary Arithmetic in an Ordered Field used in both directions, we obtain

()Ψα(μ,ν)u(μ)δE(μ)v(ν)δE(ν)bbδE(μ)δE(ν).(\dagger)\qquad\Psi_{\alpha}(\mu,\nu)\le u(\mu)-\delta\,\mathcal{E}(\mu)-v(\nu)-\delta\,\mathcal{E}(\nu)\le b-b'-\delta\,\mathcal{E}(\mu)-\delta\,\mathcal{E}(\nu).

Also W2(μ0,μ0)=0W_{2}(\mu_{0},\mu_{0})=0 by the metric axioms of Metric Space, so α2W2(μ0,μ0)2=0\tfrac{\alpha}{2}W_{2}(\mu_{0},\mu_{0})^{2}=0 and

()Ψα(μ0,μ0)=u(μ0)v(μ0)2δE(μ0),(\ddagger)\qquad\Psi_{\alpha}(\mu_{0},\mu_{0})=u(\mu_{0})-v(\mu_{0})-2\delta\,\mathcal{E}(\mu_{0}),

the collection of the two equal penalty terms using the distributivity of Field.

Claim 2. Let α\alpha be positive, let η\eta be nonnegative and let (μ^,ν^)D×D(\hat{\mu},\hat{\nu})\in\mathcal{D}\times\mathcal{D} satisfy Ψα(μ0,μ0)ηΨα(μ^,ν^)\Psi_{\alpha}(\mu_{0},\mu_{0})-\eta\le\Psi_{\alpha}(\hat{\mu},\hat{\nu}). Combining this with ()(\dagger) at (μ^,ν^)(\hat{\mu},\hat{\nu}) and with ()(\ddagger) gives

u(μ0)v(μ0)2δE(μ0)ηbbδE(μ^)δE(ν^),u(\mu_{0})-v(\mu_{0})-2\delta\,\mathcal{E}(\mu_{0})-\eta\le b-b'-\delta\,\mathcal{E}(\hat{\mu})-\delta\,\mathcal{E}(\hat{\nu}),

and adding δE(μ^)+δE(ν^)u(μ0)+v(μ0)+2δE(μ0)+η\delta\mathcal{E}(\hat{\mu})+\delta\mathcal{E}(\hat{\nu})-u(\mu_{0})+v(\mu_{0})+2\delta\mathcal{E}(\mu_{0})+\eta to both sides yields the displayed inequality of claim 2,

δE(μ^)+δE(ν^)bbu(μ0)+v(μ0)+2δE(μ0)+η=:Kη.\delta\,\mathcal{E}(\hat{\mu})+\delta\,\mathcal{E}(\hat{\nu})\le b-b'-u(\mu_{0})+v(\mu_{0})+2\delta\,\mathcal{E}(\mu_{0})+\eta=:K_{\eta}.

Since e0E(ν^)e_{0}\le\mathcal{E}(\hat{\nu}), multiplying by the nonnegative δ\delta (claim 5 of Elementary Arithmetic in an Ordered Field) gives δe0δE(ν^)\delta e_{0}\le\delta\mathcal{E}(\hat{\nu}), whence δE(μ^)Kηδe0\delta\mathcal{E}(\hat{\mu})\le K_{\eta}-\delta e_{0}. Multiplying by the positive δ1\delta^{-1} (claim 7 of Elementary Order Arithmetic in an Ordered Field and claim 5 of Elementary Arithmetic in an Ordered Field) and using δ1δ=1\delta^{-1}\delta=1 gives

E(μ^)δ1Kηe0=δ1(bbu(μ0)+v(μ0)+2δE(μ0))e0+δ1η=c0+δ1η,\mathcal{E}(\hat{\mu})\le\delta^{-1}K_{\eta}-e_{0}=\delta^{-1}\bigl(b-b'-u(\mu_{0})+v(\mu_{0})+2\delta\,\mathcal{E}(\mu_{0})\bigr)-e_{0}+\delta^{-1}\eta=c_{0}+\delta^{-1}\eta,

the middle identity by the distributivity and commutativity of Field. Interchanging the roles of μ^\hat{\mu} and ν^\hat{\nu} and using e0E(μ^)e_{0}\le\mathcal{E}(\hat{\mu}) gives E(ν^)c0+δ1η\mathcal{E}(\hat{\nu})\le c_{0}+\delta^{-1}\eta in the same way.

Finally, both μ^\hat{\mu} and ν^\hat{\nu} lie in D\mathcal{D} and have just been shown to satisfy Ec0+δ1η\mathcal{E}\le c_{0}+\delta^{-1}\eta, so Basic Properties of a Wasserstein-Coercive Penalty Pair §moment, read at the level c=c0+δ1ηc=c_{0}+\delta^{-1}\eta, gives M2(μ^)RM_{2}(\hat{\mu})\le R and M2(ν^)RM_{2}(\hat{\nu})\le R for the RR it provides at that level. The level is determined by b,b,u(μ0),v(μ0),E(μ0),δ,e0b,b',u(\mu_{0}),v(\mu_{0}),\mathcal{E}(\mu_{0}),\delta,e_{0} and η\eta alone, so neither it nor RR depends on α\alpha. This completes claim 2.

Claim 1. Let αR\alpha\in\mathbb{R} be positive. By ()(\dagger) and e0Ee_{0}\le\mathcal{E} on D\mathcal{D}, every value of Ψα\Psi_{\alpha} satisfies Ψα(μ,ν)bb2δe0\Psi_{\alpha}(\mu,\nu)\le b-b'-2\delta e_{0}, so the set {Ψα(μ,ν):(μ,ν)D×D}\{\Psi_{\alpha}(\mu,\nu):(\mu,\nu)\in\mathcal{D}\times\mathcal{D}\}, which is nonempty because D\mathcal{D} is, is bounded above and has a supremum MRM\in\mathbb{R} by that clause.

For nNn\in\mathbb{N} the inverse ι(n)1\iota(n)^{-1} exists and is positive by claim 3 of Properties of the Canonical Map from the Natural Numbers to an Ordered Field. Moreover 1n1\le n by claim 4 of Properties of the Order on the Natural Numbers, so ι(1)ι(n)\iota(1)\le\iota(n) by the monotonicity of ι\iota established above, and ι(1)=1\iota(1)=1 by claim 1 of Properties of the Canonical Map from the Natural Numbers to an Ordered Field; multiplying 1ι(n)1\le\iota(n) by the nonnegative ι(n)1\iota(n)^{-1} (claim 5 of Elementary Arithmetic in an Ordered Field) gives

ι(n)11(nN).\iota(n)^{-1}\le1\qquad(n\in\mathbb{N}).

By claim 3 of Approximation Property of the Supremum and the Infimum in R\mathbb{R}, applied with the positive ι(n)1\iota(n)^{-1}, the set of pairs (μ,ν)D×D(\mu,\nu)\in\mathcal{D}\times\mathcal{D} with Mι(n)1<Ψα(μ,ν)M-\iota(n)^{-1}<\Psi_{\alpha}(\mu,\nu) is nonempty for every nNn\in\mathbb{N}, so Axiom of Countable Choice provides a sequence of pairs (μn,νn)D×D(\mu_{n},\nu_{n})\in\mathcal{D}\times\mathcal{D} with

Mι(n)1<Ψα(μn,νn)M,M-\iota(n)^{-1}<\Psi_{\alpha}(\mu_{n},\nu_{n})\le M,

the second inequality because MM is an upper bound of the set of values. As Ψα(μ0,μ0)M\Psi_{\alpha}(\mu_{0},\mu_{0})\le M and ι(n)11\iota(n)^{-1}\le1, we get Ψα(μ0,μ0)1Mι(n)1Ψα(μn,νn)\Psi_{\alpha}(\mu_{0},\mu_{0})-1\le M-\iota(n)^{-1}\le\Psi_{\alpha}(\mu_{n},\nu_{n}), so claim 2 with η=1\eta=1 gives E(μn)c1\mathcal{E}(\mu_{n})\le c_{1} and E(νn)c1\mathcal{E}(\nu_{n})\le c_{1}, where c1=c0+δ1c_{1}=c_{0}+\delta^{-1}. Thus both sequences lie in

K={σD:E(σ)c1},K=\{\sigma\in\mathcal{D}:\mathcal{E}(\sigma)\le c_{1}\},

which is sequentially compact in (P2(Rd),W2)(\mathcal{P}_{2}(\mathbb{R}^{d}),W_{2}) by Wasserstein-Coercive Penalty Pairs §coercive read at the level c1c_{1}.

Hence there are μ^K\hat{\mu}\in K and a strictly increasing sequence (nk)kN(n_{k})_{k\in\mathbb{N}} in N\mathbb{N} such that (μnk)kN(\mu_{n_{k}})_{k\in\mathbb{N}} converges to μ^\hat{\mu} in (P2(Rd),W2)(\mathcal{P}_{2}(\mathbb{R}^{d}),W_{2}); applying sequential compactness once more to the sequence (νnk)kN(\nu_{n_{k}})_{k\in\mathbb{N}} in KK gives ν^K\hat{\nu}\in K and a strictly increasing (kj)jN(k_{j})_{j\in\mathbb{N}} such that (νnkj)jN(\nu_{n_{k_{j}}})_{j\in\mathbb{N}} converges to ν^\hat{\nu}. By A Subsequence of a Convergent Sequence Has the Same Limit, applied to the convergent sequence (μnk)kN(\mu_{n_{k}})_{k\in\mathbb{N}} and the indices (kj)jN(k_{j})_{j\in\mathbb{N}}, the sequence (μnkj)jN(\mu_{n_{k_{j}}})_{j\in\mathbb{N}} converges to μ^\hat{\mu}. Write mj=nkjm_{j}=n_{k_{j}}.

The real sequence (Ψα(μn,νn))nN(\Psi_{\alpha}(\mu_{n},\nu_{n}))_{n\in\mathbb{N}} converges to MM in (R,dR)(\mathbb{R},d_{\mathbb{R}}): from the displayed near-maximising inequality, 0MΨα(μn,νn)ι(n)10\le M-\Psi_{\alpha}(\mu_{n},\nu_{n})\le\iota(n)^{-1}, so dR(Ψα(μn,νn),M)ι(n)1d_{\mathbb{R}}(\Psi_{\alpha}(\mu_{n},\nu_{n}),M)\le\iota(n)^{-1} by claim 6 of Properties of the Absolute Value in an Ordered Field, applied to x=Ψα(μn,νn)Mx=\Psi_{\alpha}(\mu_{n},\nu_{n})-M and c=ι(n)1c=\iota(n)^{-1}, the required bounds ι(n)1x-\iota(n)^{-1}\le x and x0ι(n)1x\le0\le\iota(n)^{-1} being claim 3 of Elementary Arithmetic in an Ordered Field applied to the displayed inequalities. Given a positive ε\varepsilon', claim 2 of The Archimedean Property of the Real Numbers provides NNN\in\mathbb{N} with 1<ι(N)ε1<\iota(N)\varepsilon', whence for NnN\le n one has ι(N)ι(n)\iota(N)\le\iota(n) by the monotonicity of ι\iota established above, then 1<ι(n)ε1<\iota(n)\varepsilon' and, multiplying by the positive ι(n)1\iota(n)^{-1} (claim 10 of Elementary Order Arithmetic in an Ordered Field), ι(n)1<ε\iota(n)^{-1}<\varepsilon'. Applying A Subsequence of a Convergent Sequence Has the Same Limit twice, the sequence (Ψα(μmj,νmj))jN(\Psi_{\alpha}(\mu_{m_{j}},\nu_{m_{j}}))_{j\in\mathbb{N}} converges to MM as well.

The five estimates. Let εR\varepsilon\in\mathbb{R} be positive.

Since uu is upper semicontinuous at μ^\hat{\mu} relative to P2(Rd)\mathcal{P}_{2}(\mathbb{R}^{d}), that definition applied with the positive ε\varepsilon provides a positive r1r_{1} such that u(σ)<u(μ^)+εu(\sigma)<u(\hat{\mu})+\varepsilon for every σP2(Rd)\sigma\in\mathcal{P}_{2}(\mathbb{R}^{d}) with W2(μ^,σ)<r1W_{2}(\hat{\mu},\sigma)<r_{1}. Since vv is lower semicontinuous at ν^\hat{\nu} relative to P2(Rd)\mathcal{P}_{2}(\mathbb{R}^{d}), that definition applied with the positive ε\varepsilon provides a positive r2r_{2} such that v(ν^)ε<v(σ)v(\hat{\nu})-\varepsilon<v(\sigma), equivalently v(σ)<v(ν^)+ε-v(\sigma)<-v(\hat{\nu})+\varepsilon by claim 3 of Elementary Arithmetic in an Ordered Field used in both directions, for every σP2(Rd)\sigma\in\mathcal{P}_{2}(\mathbb{R}^{d}) with W2(ν^,σ)<r2W_{2}(\hat{\nu},\sigma)<r_{2}. No continuity of uu or of vv is used.

By Basic Properties of a Wasserstein-Coercive Penalty Pair §lsc and Lower Semicontinuous Function on a Subset of a Metric Space, applied at μ^\hat{\mu} with the positive δ1ε\delta^{-1}\varepsilon, there is a positive r3r_{3} such that E(μ^)δ1ε<E(σ)\mathcal{E}(\hat{\mu})-\delta^{-1}\varepsilon<\mathcal{E}(\sigma) for every σD\sigma\in\mathcal{D} with W2(μ^,σ)<r3W_{2}(\hat{\mu},\sigma)<r_{3}; multiplying by the positive δ\delta (claim 10 of Elementary Order Arithmetic in an Ordered Field) and using δδ1=1\delta\delta^{-1}=1 gives δE(μ^)ε<δE(σ)\delta\mathcal{E}(\hat{\mu})-\varepsilon<\delta\mathcal{E}(\sigma), that is, δE(σ)<δE(μ^)+ε-\delta\mathcal{E}(\sigma)<-\delta\mathcal{E}(\hat{\mu})+\varepsilon. In the same way there is a positive r4r_{4} with δE(σ)<δE(ν^)+ε-\delta\mathcal{E}(\sigma)<-\delta\mathcal{E}(\hat{\nu})+\varepsilon for every σD\sigma\in\mathcal{D} with W2(ν^,σ)<r4W_{2}(\hat{\nu},\sigma)<r_{4}.

For the last estimate write s=W2(μ^,ν^)s=W_{2}(\hat{\mu},\hat{\nu}), a nonnegative real number, and put

ε0=ε(2α(1+s))1,\varepsilon_{0}=\varepsilon\bigl(2\alpha(1+s)\bigr)^{-1},

which is positive because α\alpha and 1+s1+s are positive (for the latter, 0<11+s0<1\le1+s) and by claims 5 and 7 of Elementary Order Arithmetic in an Ordered Field. Let σ,τP2(Rd)\sigma,\tau\in\mathcal{P}_{2}(\mathbb{R}^{d}) satisfy W2(μ^,σ)<ε0W_{2}(\hat{\mu},\sigma)<\varepsilon_{0} and W2(ν^,τ)<ε0W_{2}(\hat{\nu},\tau)<\varepsilon_{0}. The triangle inequality and the symmetry of W2W_{2} (metric axioms of Metric Space), used twice, give sW2(μ^,σ)+W2(σ,τ)+W2(τ,ν^)s\le W_{2}(\hat{\mu},\sigma)+W_{2}(\sigma,\tau)+W_{2}(\tau,\hat{\nu}) and hence s2ε0<W2(σ,τ)s-2\varepsilon_{0}<W_{2}(\sigma,\tau). We claim that

s24ε0sW2(σ,τ)2.s^{2}-4\varepsilon_{0}s\le W_{2}(\sigma,\tau)^{2}.

If 2ε0s2\varepsilon_{0}\le s, then 0s2ε0<W2(σ,τ)0\le s-2\varepsilon_{0}<W_{2}(\sigma,\tau), so (s2ε0)2<W2(σ,τ)2(s-2\varepsilon_{0})^{2}<W_{2}(\sigma,\tau)^{2} by claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field, and (s2ε0)2=s24ε0s+4ε02(s-2\varepsilon_{0})^{2}=s^{2}-4\varepsilon_{0}s+4\varepsilon_{0}^{2} by claim 5 of Zero Products and Elementary Identities in a Field, the last summand being nonnegative by Nonnegativity of Squares in an Ordered Field, so s24ε0s(s2ε0)2s^{2}-4\varepsilon_{0}s\le(s-2\varepsilon_{0})^{2}. If instead s<2ε0s<2\varepsilon_{0}, then multiplying by the nonnegative ss gives s22ε0ss^{2}\le2\varepsilon_{0}s, so s24ε0s2ε0s0W2(σ,τ)2s^{2}-4\varepsilon_{0}s\le-2\varepsilon_{0}s\le0\le W_{2}(\sigma,\tau)^{2}, the middle step because 2ε0s2\varepsilon_{0}s is nonnegative. Multiplying the claimed inequality by the nonnegative α2\tfrac{\alpha}{2} and using claim 3 of Elementary Arithmetic in an Ordered Field in both directions gives

α2W2(σ,τ)2α2s2+2αε0s.-\tfrac{\alpha}{2}W_{2}(\sigma,\tau)^{2}\le-\tfrac{\alpha}{2}s^{2}+2\alpha\varepsilon_{0}s .

Finally 2αε0s=εs(1+s)1ε2\alpha\varepsilon_{0}s=\varepsilon\,s\,(1+s)^{-1}\le\varepsilon, since s1+ss\le1+s gives s(1+s)11s(1+s)^{-1}\le1 on multiplying by the positive (1+s)1(1+s)^{-1}, and multiplying that by the nonnegative ε\varepsilon preserves the inequality; so

α2W2(σ,τ)2α2s2+ε.-\tfrac{\alpha}{2}W_{2}(\sigma,\tau)^{2}\le-\tfrac{\alpha}{2}s^{2}+\varepsilon .

Conclusion of claim 1. Let rr be the least of r1,r2,r3,r4,ε0r_{1},r_{2},r_{3},r_{4},\varepsilon_{0}, obtained by repeated use of claim 9 of Elementary Order Arithmetic in an Ordered Field; it is positive. Since (μmj)j(\mu_{m_{j}})_{j} converges to μ^\hat{\mu} and (νmj)j(\nu_{m_{j}})_{j} converges to ν^\hat{\nu}, and since (Ψα(μmj,νmj))j(\Psi_{\alpha}(\mu_{m_{j}},\nu_{m_{j}}))_{j} converges to MM, given a positive ε\varepsilon'' there are three thresholds in N\mathbb{N} beyond which the corresponding distances are smaller than rr, rr and ε\varepsilon''; let JJ be the largest of the three, which exists by the trichotomy of the order of N\mathbb{N} (claim 3 of Properties of the Order on the Natural Numbers), and let jNj\in\mathbb{N} satisfy JjJ\le j. Then W2(μ^,μmj)<rW_{2}(\hat{\mu},\mu_{m_{j}})<r and W2(ν^,νmj)<rW_{2}(\hat{\nu},\nu_{m_{j}})<r, so all five estimates apply with σ=μmj\sigma=\mu_{m_{j}} and τ=νmj\tau=\nu_{m_{j}}, and adding them gives

Ψα(μmj,νmj)Ψα(μ^,ν^)+5ε.\Psi_{\alpha}(\mu_{m_{j}},\nu_{m_{j}})\le\Psi_{\alpha}(\hat{\mu},\hat{\nu})+5\varepsilon .

Also dR(Ψα(μmj,νmj),M)<εd_{\mathbb{R}}(\Psi_{\alpha}(\mu_{m_{j}},\nu_{m_{j}}),M)<\varepsilon'', so M<Ψα(μmj,νmj)+εM<\Psi_{\alpha}(\mu_{m_{j}},\nu_{m_{j}})+\varepsilon'' by claim 3 of Properties of the Absolute Value in an Ordered Field, and therefore MΨα(μ^,ν^)+5ε+εM\le\Psi_{\alpha}(\hat{\mu},\hat{\nu})+5\varepsilon+\varepsilon''. As ε\varepsilon'' was an arbitrary positive number, Comparison of Real Numbers with Arbitrary Positive Slack §slack-above gives MΨα(μ^,ν^)+5εM\le\Psi_{\alpha}(\hat{\mu},\hat{\nu})+5\varepsilon. Given now an arbitrary positive ε\varepsilon''', taking ε=ει(5)1\varepsilon=\varepsilon'''\iota(5)^{-1}, positive by claims 3 and 7 quoted above, gives 5ε=ε5\varepsilon=\varepsilon''' and hence MΨα(μ^,ν^)+εM\le\Psi_{\alpha}(\hat{\mu},\hat{\nu})+\varepsilon'''; so MΨα(μ^,ν^)M\le\Psi_{\alpha}(\hat{\mu},\hat{\nu}) by Comparison of Real Numbers with Arbitrary Positive Slack §slack-above once more.

Since KDK\subseteq\mathcal{D}, the pair (μ^,ν^)(\hat{\mu},\hat{\nu}) lies in D×D\mathcal{D}\times\mathcal{D}, so Ψα(μ^,ν^)M\Psi_{\alpha}(\hat{\mu},\hat{\nu})\le M as well. Hence Ψα(μ,ν)M=Ψα(μ^,ν^)\Psi_{\alpha}(\mu,\nu)\le M=\Psi_{\alpha}(\hat{\mu},\hat{\nu}) for every (μ,ν)D×D(\mu,\nu)\in\mathcal{D}\times\mathcal{D}, which is claim 1.

Claim 3. Apply Penalised Suprema: Monotonicity, Near-Maximisers, and Vanishing Penalty along a Doubling Sequence with the nonempty set Z=D×DZ=\mathcal{D}\times\mathcal{D}, with ψ(μ,ν)=uδ(μ)vδ+(ν)\psi(\mu,\nu)=u^{-}_{\delta}(\mu)-v^{+}_{\delta}(\nu) and with D(μ,ν)=12W2(μ,ν)2D(\mu,\nu)=\tfrac{1}{2}W_{2}(\mu,\nu)^{2}. Its hypotheses hold: ψ\psi is bounded above by bb2δe0b-b'-2\delta e_{0}, by ()(\dagger) and e0Ee_{0}\le\mathcal{E} on D\mathcal{D}; DD is nonnegative, being a nonnegative multiple of a square (Nonnegativity of Squares in an Ordered Field and claim 5 of Elementary Arithmetic in an Ordered Field); and D(μ0,μ0)=0D(\mu_{0},\mu_{0})=0 because W2(μ0,μ0)=0W_{2}(\mu_{0},\mu_{0})=0 by the metric axioms of Metric Space. For a positive β\beta the function Ψβ\Psi_{\beta} of that lemma has value ψ(μ,ν)βD(μ,ν)=Ψβ(μ,ν)\psi(\mu,\nu)-\beta D(\mu,\nu)=\Psi_{\beta}(\mu,\nu) here, so the two suprema written M(β)M(\beta) agree and the lemma may be quoted with the present notation.

Let γk=2kα0\gamma_{k}=2^{k}\alpha_{0} for kNk\in\mathbb{N}, so that γn=αn\gamma_{n}=\alpha_{n} for every nNn\in\mathbb{N}. Let εR\varepsilon\in\mathbb{R} be positive. By Penalised Suprema: Monotonicity, Near-Maximisers, and Vanishing Penalty along a Doubling Sequence §vanishing, applied with β0=α0\beta_{0}=\alpha_{0}, the sequence whose kk-th term is M(γk)M(γk+1)M(\gamma_{k})-M(\gamma_{k+1}) converges to 00, so there is NNN\in\mathbb{N} such that

M(γk)M(γk+1)<ε4(kN, Nk),M(\gamma_{k})-M(\gamma_{k+1})<\tfrac{\varepsilon}{4}\qquad(k\in\mathbb{N},\ N\le k),

using claim 3 of Properties of the Absolute Value in an Ordered Field to pass from the distance to the value. Let nNn\in\mathbb{N} satisfy N+1nN+1\le n and write k=n1k=n-1, a natural number with NkN\le k by claim 7 of Properties of the Order on the Natural Numbers; then γk+1=γn=αn\gamma_{k+1}=\gamma_{n}=\alpha_{n} and γk=2n1α0\gamma_{k}=2^{n-1}\alpha_{0}, which is αn2\tfrac{\alpha_{n}}{2} because 2n=22n12^{n}=2\cdot2^{n-1}.

Since (μ^n,ν^n)(\hat{\mu}_{n},\hat{\nu}_{n}) maximises Ψαn\Psi_{\alpha_{n}}, it satisfies M(αn)ηΨαn(μ^n,ν^n)M(\alpha_{n})-\eta\le\Psi_{\alpha_{n}}(\hat{\mu}_{n},\hat{\nu}_{n}) for every positive η\eta, so Penalised Suprema: Monotonicity, Near-Maximisers, and Vanishing Penalty along a Doubling Sequence §near-maximiser, applied at β=αn\beta=\alpha_{n}, gives

αn212W2(μ^n,ν^n)2M(αn2)M(αn)+η\tfrac{\alpha_{n}}{2}\cdot\tfrac{1}{2}W_{2}(\hat{\mu}_{n},\hat{\nu}_{n})^{2}\le M\bigl(\tfrac{\alpha_{n}}{2}\bigr)-M(\alpha_{n})+\eta

for every positive η\eta, whence αn4W2(μ^n,ν^n)2M(γk)M(γk+1)\tfrac{\alpha_{n}}{4}W_{2}(\hat{\mu}_{n},\hat{\nu}_{n})^{2}\le M(\gamma_{k})-M(\gamma_{k+1}) by Comparison of Real Numbers with Arbitrary Positive Slack §slack-above. Combining with the displayed bound and multiplying by the positive 44 gives

αnW2(μ^n,ν^n)2<ε(N+1n).\alpha_{n}\,W_{2}(\hat{\mu}_{n},\hat{\nu}_{n})^{2}<\varepsilon\qquad(N+1\le n).

The left-hand side is nonnegative, so its distance to 00 is itself by Absolute Value in an Ordered Field; hence the sequence converges to 00.

For the last assertion, 2α0αn2\alpha_{0}\le\alpha_{n} for every nNn\in\mathbb{N}, since 22n2\le2^{n} by claim 4 of Properties of the Order on the Natural Numbers and the monotonicity of natural powers of 22; so multiplying W2(μ^n,ν^n)2(2α0)1αnW2(μ^n,ν^n)2W_{2}(\hat{\mu}_{n},\hat{\nu}_{n})^{2}\le(2\alpha_{0})^{-1}\alpha_{n}W_{2}(\hat{\mu}_{n},\hat{\nu}_{n})^{2} shows that the squares converge to 00. Given a positive ε\varepsilon', take nn beyond the threshold for the positive ε2\varepsilon'^{2}; then W2(μ^n,ν^n)2<ε2W_{2}(\hat{\mu}_{n},\hat{\nu}_{n})^{2}<\varepsilon'^{2}, and both W2(μ^n,ν^n)W_{2}(\hat{\mu}_{n},\hat{\nu}_{n}) and ε\varepsilon' are nonnegative, so W2(μ^n,ν^n)<εW_{2}(\hat{\mu}_{n},\hat{\nu}_{n})<\varepsilon' by claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field.

Claim 4. The points μ^\hat{\mu} and ν^\hat{\nu} lie in D\mathcal{D}, hence in P2(Rd)\mathcal{P}_{2}(\mathbb{R}^{d}) by Penalty Pairs on the Wasserstein Space: the Penalty, Its Score, and Their Domains §pair, so Existence of an Optimal Coupling of Two Probability Measures with Finite Second Moment, read with dd in place of the dimension named there, provides an optimal coupling π\pi of μ^\hat{\mu} and ν^\hat{\nu}; in particular π\pi is a probability measure on Rd+d\mathbb{R}^{d+d}. Since (Ω,F,P)(\Omega,\mathcal{F},P) is rich, Rich Probability Space §rich, read with m=d+dm=d+d, provides a random vector WW in Rd+d\mathbb{R}^{d+d} with L(W)=π\mathcal{L}(W)=\pi.

Put X^=pr1W\hat{X}=\mathrm{pr}_{1}\circ W and Y^=pr2W\hat{Y}=\mathrm{pr}_{2}\circ W, random vectors in Rd\mathbb{R}^{d} by Basic Properties of Random Vectors: Coordinates, Borel Images and Arithmetic, Change of Variables, Almost Sure Equality and Pairs §composition, the projections being Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections. By the change-of-variables formula of Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pushforward and claim 1 of The Coordinate Fields of a Coupling, and the Second Moment as a Lipschitz Function of the Wasserstein Distance,

ΩX^2dP=Rd+dpr12dπ=M2(μ^),\int_{\Omega}\lVert\hat{X}\rVert^{2}\,dP=\int_{\mathbb{R}^{d+d}}\lVert\mathrm{pr}_{1}\rVert^{2}\,d\pi=M_{2}(\hat{\mu}),

which is finite, so X^L2(Ω;Rd)\hat{X}\in L^{2}(\Omega;\mathbb{R}^{d}) by Square-Integrable Random Vectors: Coordinates, Operations, Almost Sure Equality and the Mean-Square Form; likewise Y^L2(Ω;Rd)\hat{Y}\in L^{2}(\Omega;\mathbb{R}^{d}). By Basic Properties of Random Vectors: Coordinates, Borel Images and Arithmetic, Change of Variables, Almost Sure Equality and Pairs §composition, L(X^)=L(pr1W)=(pr1)#L(W)=(pr1)#π\mathcal{L}(\hat{X})=\mathcal{L}(\mathrm{pr}_{1}\circ W)=(\mathrm{pr}_{1})_{\#}\mathcal{L}(W)=(\mathrm{pr}_{1})_{\#}\pi, which is μ^\hat{\mu} because π\pi is a coupling of μ^\hat{\mu} and ν^\hat{\nu} by Couplings of Two Probability Measures on Euclidean Space and Their Quadratic Cost §coupling; likewise L(Y^)=(pr2)#π=ν^\mathcal{L}(\hat{Y})=(\mathrm{pr}_{2})_{\#}\pi=\hat{\nu}; in particular X^,Y^DΛ\hat{X},\hat{Y}\in\mathcal{D}^{\Lambda}. The pairing of pr1\mathrm{pr}_{1} with pr2\mathrm{pr}_{2} is the identity of Rd+d\mathbb{R}^{d+d} by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §pairing and Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections, so (X^,Y^)=W(\hat{X},\hat{Y})=W and L((X^,Y^))=π\mathcal{L}((\hat{X},\hat{Y}))=\pi. Finally, by change of variables once more and claim 1 of The Coordinate Fields of a Coupling, and the Second Moment as a Lipschitz Function of the Wasserstein Distance,

X^Y^L22=Rd+dpr1pr22dπ=I(π)=W2(μ^,ν^)2,\lVert\hat{X}-\hat{Y}\rVert_{L^{2}}^{2}=\int_{\mathbb{R}^{d+d}}\lVert\mathrm{pr}_{1}-\mathrm{pr}_{2}\rVert^{2}\,d\pi=I(\pi)=W_{2}(\hat{\mu},\hat{\nu})^{2},

the last step by Optimal Coupling of Two Probability Measures with Finite Second Moment §optimal; both X^Y^L2\lVert\hat{X}-\hat{Y}\rVert_{L^{2}} and W2(μ^,ν^)W_{2}(\hat{\mu},\hat{\nu}) being nonnegative, claim 3 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field gives X^Y^L2=W2(μ^,ν^)\lVert\hat{X}-\hat{Y}\rVert_{L^{2}}=W_{2}(\hat{\mu},\hat{\nu}).

Now let X^,Y^\hat{X},\hat{Y} be any pair with L(X^)=μ^\mathcal{L}(\hat{X})=\hat{\mu}, L(Y^)=ν^\mathcal{L}(\hat{Y})=\hat{\nu} and X^Y^L2=W2(μ^,ν^)\lVert\hat{X}-\hat{Y}\rVert_{L^{2}}=W_{2}(\hat{\mu},\hat{\nu}). The law of the pairing (X^,Y^)(\hat{X},\hat{Y}) is the joint law L(X^,Y^)\mathcal{L}(\hat{X},\hat{Y}) of The Joint Law of Two Classes of Square-Integrable Random Vectors: Marginals, Cost, and Invariance Under Shifts by a Vector Field §joint, and by The Joint Law of Two Classes of Square-Integrable Random Vectors: Marginals, Cost, and Invariance Under Shifts by a Vector Field §marginals it is a coupling of L(X^)=μ^\mathcal{L}(\hat{X})=\hat{\mu} and L(Y^)=ν^\mathcal{L}(\hat{Y})=\hat{\nu} whose quadratic cost is X^Y^L22\lVert\hat{X}-\hat{Y}\rVert_{L^{2}}^{2}; substituting X^Y^L2=W2(μ^,ν^)\lVert\hat{X}-\hat{Y}\rVert_{L^{2}}=W_{2}(\hat{\mu},\hat{\nu}) makes that cost W2(μ^,ν^)2W_{2}(\hat{\mu},\hat{\nu})^{2}, so the coupling is optimal by Optimal Coupling of Two Probability Measures with Finite Second Moment §optimal.

Let X,YDΛX,Y\in\mathcal{D}^{\Lambda}. By The Wasserstein Distance and the Mean-Square Distance of Random Vectors §inequality, W2(L(X),L(Y))XYL2W_{2}(\mathcal{L}(X),\mathcal{L}(Y))\le\lVert X-Y\rVert_{L^{2}}, and both sides are nonnegative, so claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field gives W2(L(X),L(Y))2XYL22W_{2}(\mathcal{L}(X),\mathcal{L}(Y))^{2}\le\lVert X-Y\rVert_{L^{2}}^{2}; multiplying by the nonnegative α2\tfrac{\alpha}{2} (claim 5 of Elementary Arithmetic in an Ordered Field) and using claim 3 of that lemma in both directions gives

α2XYL22α2W2(L(X),L(Y))2.-\tfrac{\alpha}{2}\lVert X-Y\rVert_{L^{2}}^{2}\le-\tfrac{\alpha}{2}W_{2}(\mathcal{L}(X),\mathcal{L}(Y))^{2}.

Adding uδ(L(X))vδ+(L(Y))u^{-}_{\delta}(\mathcal{L}(X))-v^{+}_{\delta}(\mathcal{L}(Y)) to both sides, the left-hand side becomes the expression to be bounded and the right-hand side is Ψα(L(X),L(Y))\Psi_{\alpha}(\mathcal{L}(X),\mathcal{L}(Y)), both L(X)\mathcal{L}(X) and L(Y)\mathcal{L}(Y) lying in D\mathcal{D}. By claim 1 that value is at most Ψα(μ^,ν^)\Psi_{\alpha}(\hat{\mu},\hat{\nu}), which equals uδ(μ^)vδ+(ν^)α2X^Y^L22u^{-}_{\delta}(\hat{\mu})-v^{+}_{\delta}(\hat{\nu})-\tfrac{\alpha}{2}\lVert\hat{X}-\hat{Y}\rVert_{L^{2}}^{2} because X^Y^L2=W2(μ^,ν^)\lVert\hat{X}-\hat{Y}\rVert_{L^{2}}=W_{2}(\hat{\mu},\hat{\nu}). This is the displayed inequality of claim 4.

Claim 5. Let XDΛX\in\mathcal{D}^{\Lambda}. Claim 4 applied to the pair (X,Y^)(X,\hat{Y}), legitimate since Y^DΛ\hat{Y}\in\mathcal{D}^{\Lambda}, gives

uδ(L(X))vδ+(ν^)α2XY^L22uδ(μ^)vδ+(ν^)α2X^Y^L22,u^{-}_{\delta}(\mathcal{L}(X))-v^{+}_{\delta}(\hat{\nu})-\tfrac{\alpha}{2}\lVert X-\hat{Y}\rVert_{L^{2}}^{2}\le u^{-}_{\delta}(\hat{\mu})-v^{+}_{\delta}(\hat{\nu})-\tfrac{\alpha}{2}\lVert\hat{X}-\hat{Y}\rVert_{L^{2}}^{2},

using L(Y^)=ν^\mathcal{L}(\hat{Y})=\hat{\nu}; adding vδ+(ν^)v^{+}_{\delta}(\hat{\nu}) to both sides gives

uδ(L(X))α2XY^L22uδ(μ^)α2X^Y^L22,u^{-}_{\delta}(\mathcal{L}(X))-\tfrac{\alpha}{2}\lVert X-\hat{Y}\rVert_{L^{2}}^{2}\le u^{-}_{\delta}(\hat{\mu})-\tfrac{\alpha}{2}\lVert\hat{X}-\hat{Y}\rVert_{L^{2}}^{2},

and L(X^)=μ^\mathcal{L}(\hat{X})=\hat{\mu}, so the right-hand side is the value of the first function at X^\hat{X}. As XDΛX\in\mathcal{D}^{\Lambda} was arbitrary, that value is at least every value of the function on DΛ\mathcal{D}^{\Lambda}. Symmetrically, claim 4 applied to the pair (X^,Y)(\hat{X},Y) for YDΛY\in\mathcal{D}^{\Lambda} gives

vδ+(L(Y))α2X^YL22vδ+(ν^)α2X^Y^L22-v^{+}_{\delta}(\mathcal{L}(Y))-\tfrac{\alpha}{2}\lVert\hat{X}-Y\rVert_{L^{2}}^{2}\le-v^{+}_{\delta}(\hat{\nu})-\tfrac{\alpha}{2}\lVert\hat{X}-\hat{Y}\rVert_{L^{2}}^{2}

after adding uδ(μ^)-u^{-}_{\delta}(\hat{\mu}) to both sides, and multiplying by 1-1, which reverses the order by claim 4 of Elementary Order Arithmetic in an Ordered Field in its weak form (claim 3 of Elementary Arithmetic in an Ordered Field used in both directions), gives

vδ+(ν^)+α2X^Y^L22vδ+(L(Y))+α2X^YL22,v^{+}_{\delta}(\hat{\nu})+\tfrac{\alpha}{2}\lVert\hat{X}-\hat{Y}\rVert_{L^{2}}^{2}\le v^{+}_{\delta}(\mathcal{L}(Y))+\tfrac{\alpha}{2}\lVert\hat{X}-Y\rVert_{L^{2}}^{2},

so the value of the second function at Y^\hat{Y} is at most every value of it on DΛ\mathcal{D}^{\Lambda}.

For the final assertion, DΛ\mathcal{D}^{\Lambda} is a subset of the metric space (L2(Ω;Rd),dL2)(L^{2}(\Omega;\mathbb{R}^{d}),d_{L^{2}}) and X^,Y^DΛ\hat{X},\hat{Y}\in\mathcal{D}^{\Lambda}; taking the radius 11 in Local Maximum of a Function Relative to a Subset of a Metric Space, the first function has a local maximum at X^\hat{X} relative to DΛ\mathcal{D}^{\Lambda}, since its value there dominates its value at every point of DΛ\mathcal{D}^{\Lambda} and a fortiori at every point of DΛ\mathcal{D}^{\Lambda} within distance 11 of X^\hat{X}; and likewise the second has a local minimum at Y^\hat{Y} by Local Minimum of a Function Relative to a Subset of a Metric Space. This is claim 5.

Please log in to copy this version.

Citations

Loading…

Dependency Graph

0 prerequisites

Comments

Loading…