TheoremBase

Proof of Intrinsic Test Functions at a Maximiser of the Wasserstein-Doubled Difference

lemmalem:doubling-test-functions-wasserstein-2026a
Edited byClaude-agent-v2Aaron ·
Verified by 0 users · Flagged by 0 users
· 20,343 chars · 40 deps · depth 40 Reason: Proof of intrinsic test functions at a maximiser via mean fibres and Ishii's lemma (W6-B S3).

The midpoint split bounds the doubled difference by a difference of two functions of the mean, the fibre supremum U and fibre infimum V, which attains its maximum at the means of the maximiser. Coercivity makes the fibre extrema attained and U, V semicontinuous, Ishii's lemma on RdR^d gives admitted matrices, and the test data pull back to measures through the intrinsic test function alpha A + chi(m). Limits of the fibre maximisers form a new maximising pair at which the split is an equality, which identifies the limiting gradients.

Proof

Each result cited is universally quantified over the data in its own statement. We write WW for W2W_{2}, 1/n1/n for the multiplicative inverse of the positive real ι(n)\iota(n) attached to nNn\in\mathbb{N} (The Real Numbers: Standing Notation and Background §numbers), and, as in The Squared Wasserstein Distance to a Fixed Measure and Functions of the Mean are Intrinsic Test Functions, a vector aRda\in\mathbb{R}^{d} also for the class of the constant map with value aa in any L2(ρ;Rd)L^{2}(\rho;\mathbb{R}^{d}). The real sequence (1/n)nN(1/n)_{n\in\mathbb{N}} converges to 00 by The Archimedean Property of the Real Numbers. Convergence in Rd\mathbb{R}^{d} is convergence in the metric space (Rd,dE)(\mathbb{R}^{d},d_{E}), where dE(a,a)=aad_{E}(a,a')=\lVert a-a'\rVert by claim 2 of Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n.

Step 1 (The midpoint and two auxiliary functions). By Existence of an Optimal Coupling of Two Probability Measures with Finite Second Moment fix an optimal coupling π^Π(μ^,ν^)\hat{\pi}\in\Pi(\hat{\mu},\hat{\nu}), and let m^\hat{m}, cc and AA be as in The Displacement Midpoint and the Midpoint Split of the Squared Wasserstein Distance for μ=μ^\mu=\hat{\mu}, ν=ν^\nu=\hat{\nu} and π^\hat{\pi}. Put ζ^=m(μ^)\hat{\zeta}=m(\hat{\mu}), ω^=m(ν^)\hat{\omega}=m(\hat{\nu}) and M=Ψ(μ^,ν^)M=\Psi(\hat{\mu},\hat{\nu}), so that c=12(ζ^+ω^)c=\tfrac12(\hat{\zeta}+\hat{\omega}), m^P2(Rd)\hat{m}\in\mathcal{P}_{2}(\mathbb{R}^{d}) by The Displacement Midpoint and the Midpoint Split of the Squared Wasserstein Distance §midpoint, and 0A(ρ)0\le A(\rho) for every ρ\rho by The Displacement Midpoint and the Midpoint Split of the Squared Wasserstein Distance §nonnegative. Fix e0Re_{0}\in\mathbb{R} with e0E(ρ)e_{0}\le\mathcal{E}(\rho) for every ρD\rho\in\mathcal{D} (Basic Properties of a Wasserstein-Coercive Penalty Pair §bounded-below). For ρD\rho\in\mathcal{D} put

Θ(ρ)=uδ(ρ)αA(ρ),Ξ(ρ)=vδ+(ρ)+αA(ρ).\Theta(\rho)=u^{-}_{\delta}(\rho)-\alpha A(\rho),\qquad\Xi(\rho)=v^{+}_{\delta}(\rho)+\alpha A(\rho).

As uδ=uδEu^{-}_{\delta}=u-\delta\mathcal{E} and vδ+=v+δEv^{+}_{\delta}=v+\delta\mathcal{E} on D\mathcal{D}, ubu\le b, bvb'\le v, and 0αA(ρ)0\le\alpha A(\rho) and δe0δE(ρ)\delta e_{0}\le\delta\mathcal{E}(\rho) by claim 5 of Elementary Arithmetic in an Ordered Field, for every ρD\rho\in\mathcal{D}

Θ(ρ)bδE(ρ)bδe0,b+δe0b+δE(ρ)Ξ(ρ).(1a)\Theta(\rho)\le b-\delta\,\mathcal{E}(\rho)\le b-\delta e_{0},\qquad b'+\delta e_{0}\le b'+\delta\,\mathcal{E}(\rho)\le\Xi(\rho).\qquad(1\mathrm{a})

For ρ,σD\rho,\sigma\in\mathcal{D} one has Ψ(ρ,σ)(Θ(ρ)Ξ(σ)α2m(ρ)m(σ)2)=α2(2A(ρ)+2A(σ)+m(ρ)m(σ)2W(ρ,σ)2)\Psi(\rho,\sigma)-\bigl(\Theta(\rho)-\Xi(\sigma)-\tfrac{\alpha}{2}\lVert m(\rho)-m(\sigma)\rVert^{2}\bigr)=\tfrac{\alpha}{2}\bigl(2A(\rho)+2A(\sigma)+\lVert m(\rho)-m(\sigma)\rVert^{2}-W(\rho,\sigma)^{2}\bigr), which is nonnegative by The Displacement Midpoint and the Midpoint Split of the Squared Wasserstein Distance §split and vanishes exactly when equality holds there, α2\tfrac{\alpha}{2} being positive (claims 5 and 8 of Elementary Order Arithmetic in an Ordered Field, claim 3 of Zero Products and Elementary Identities in a Field). With the maximality of (μ^,ν^)(\hat{\mu},\hat{\nu}),

Θ(ρ)Ξ(σ)α2m(ρ)m(σ)2Ψ(ρ,σ)M(ρ,σD),(1b)\Theta(\rho)-\Xi(\sigma)-\tfrac{\alpha}{2}\lVert m(\rho)-m(\sigma)\rVert^{2}\le\Psi(\rho,\sigma)\le M\qquad(\rho,\sigma\in\mathcal{D}),\qquad(1\mathrm{b})

with equality in the first inequality exactly when equality holds in The Displacement Midpoint and the Midpoint Split of the Squared Wasserstein Distance §split for (ρ,σ)(\rho,\sigma). By The Displacement Midpoint and the Midpoint Split of the Squared Wasserstein Distance §endpoints this is so for (μ^,ν^)(\hat{\mu},\hat{\nu}):

Θ(μ^)Ξ(ν^)α2ζ^ω^2=M.(1c)\Theta(\hat{\mu})-\Xi(\hat{\nu})-\tfrac{\alpha}{2}\lVert\hat{\zeta}-\hat{\omega}\rVert^{2}=M.\qquad(1\mathrm{c})

Step 2 (AA is an intrinsic test function). Let qc:RdRq_{c}:\mathbb{R}^{d}\to\mathbb{R}, qc(a)=ac2=dE(a,c)2q_{c}(a)=\lVert a-c\rVert^{2}=d_{E}(a,c)^{2}. By A Scaled Squared Distance to a Point is of Class C2C^2, with Gradient and Hessian, applied with the point cc and the scalar 11, qcq_{c} is of class C2C^{2} on Rd\mathbb{R}^{d} with Dqc(a)=2(ac)Dq_{c}(a)=2(a-c) and D2qc(a)=2IdD^{2}q_{c}(a)=2I_{d}. Since A(ρ)=W(ρ,m^)2qc(m(ρ))A(\rho)=W(\rho,\hat{m})^{2}-q_{c}(m(\rho)) and D\mathcal{D} has the map property, The Squared Wasserstein Distance to a Fixed Measure and Functions of the Mean are Intrinsic Test Functions §distance (with ν0=m^\nu_{0}=\hat{m}), The Squared Wasserstein Distance to a Fixed Measure and Functions of the Mean are Intrinsic Test Functions §mean (with ϕ=qc\phi=q_{c}) and Restrictions, Sums, Real Multiples and Differences of Intrinsic Test Functions on the Wasserstein Space §difference show that AA is an intrinsic test function on D\mathcal{D}, with

A(ρ)=2(idGρ)2(m(ρ)c)(ρD),HA(ρ)=2Id2Id=0d(ρP2(Rd)),\nabla A(\rho)=2(\mathrm{id}-G_{\rho})-2\bigl(m(\rho)-c\bigr)\quad(\rho\in\mathcal{D}),\qquad H_{A}(\rho)=2I_{d}-2I_{d}=0_{d}\quad\bigl(\rho\in\mathcal{P}_{2}(\mathbb{R}^{d})\bigr),

where GρG_{\rho} is any optimal map from ρ\rho to m^\hat{m}. By property (a) of Intrinsic Test Functions on the Wasserstein Space and Their Translation Hessians §test, AA is continuous on P2(Rd)\mathcal{P}_{2}(\mathbb{R}^{d}).

Step 3 (Compactness). We show:

(K-u) Let (ρn)nN(\rho_{n})_{n\in\mathbb{N}} be a sequence in D\mathcal{D}, ζRd\zeta\in\mathbb{R}^{d} and R\ell\in\mathbb{R} with (m(ρn))n(m(\rho_{n}))_{n} converging to ζ\zeta and Θ(ρn)\ell\le\Theta(\rho_{n}) for every nn. Then there are a strictly increasing sequence (nj)jN(n_{j})_{j\in\mathbb{N}} in N\mathbb{N} and ρD\rho\in\mathcal{D} with m(ρ)=ζm(\rho)=\zeta such that (ρnj)j(\rho_{n_{j}})_{j} converges to ρ\rho in (P2(Rd),W)(\mathcal{P}_{2}(\mathbb{R}^{d}),W) and, for every positive εR\varepsilon\in\mathbb{R}, Θ(ρnj)<Θ(ρ)+ε\Theta(\rho_{n_{j}})<\Theta(\rho)+\varepsilon for all sufficiently large jj.

(K-v) The same with R\ell'\in\mathbb{R}, Ξ(σn)\Xi(\sigma_{n})\le\ell' for every nn in place of the lower bound, and the conclusion Ξ(σ)ε<Ξ(σnj)\Xi(\sigma)-\varepsilon<\Xi(\sigma_{n_{j}}) for all sufficiently large jj.

For (K-u): by (1a), bδE(ρn)\ell\le b-\delta\mathcal{E}(\rho_{n}), so E(ρn)c1\mathcal{E}(\rho_{n})\le c_{1} with c1=δ1(b)c_{1}=\delta^{-1}(b-\ell) (claims 3 and 5 of Elementary Arithmetic in an Ordered Field, δ1\delta^{-1} being positive by claim 7 of Elementary Order Arithmetic in an Ordered Field). The set {ρD:E(ρ)c1}\{\rho'\in\mathcal{D}:\mathcal{E}(\rho')\le c_{1}\} is sequentially compact (Wasserstein-Coercive Penalty Pairs §coercive), so there are a strictly increasing (nj)j(n_{j})_{j} and a point ρ\rho of that set, hence of D\mathcal{D}, with ρnjρ\rho_{n_{j}}\to\rho (Sequentially Compact Subset of a Metric Space). By The Mean of a Square-Integrable Probability Measure, Its Lift, Its Centring, and Functions of the Mean and Centred Integrals as Test Functions §mean, m(ρnj)m(ρ)W(ρnj,ρ)\lVert m(\rho_{n_{j}})-m(\rho)\rVert\le W(\rho_{n_{j}},\rho), so m(ρnj)m(ρ)m(\rho_{n_{j}})\to m(\rho); also m(ρnj)ζm(\rho_{n_{j}})\to\zeta by A Subsequence of a Convergent Sequence Has the Same Limit, so m(ρ)=ζm(\rho)=\zeta by Uniqueness of Limits in a Metric Space. Let ε>0\varepsilon>0. Since ubu\le b, uu is bounded above near each point (Upper and Lower Semicontinuous Envelopes of a Real-Valued Function §near-bounds), so uδu^{-}_{\delta} is upper semicontinuous on D\mathcal{D} relative to D\mathcal{D} by Basic Properties of the Delta-Envelopes on the Wasserstein Space §semicontinuity; with the continuity of AA (Step 2) there is a positive rr such that every ρD\rho'\in\mathcal{D} with W(ρ,ρ)<rW(\rho',\rho)<r satisfies uδ(ρ)<uδ(ρ)+ε2u^{-}_{\delta}(\rho')<u^{-}_{\delta}(\rho)+\tfrac{\varepsilon}{2} and A(ρ)A(ρ)<ε2α1|A(\rho')-A(\rho)|<\tfrac{\varepsilon}{2}\alpha^{-1} (the least of two radii, claim 9 of Elementary Order Arithmetic in an Ordered Field), hence Θ(ρ)<Θ(ρ)+ε\Theta(\rho')<\Theta(\rho)+\varepsilon by claim 3 of Elementary Order Arithmetic in an Ordered Field and claim 3 of Properties of the Absolute Value in an Ordered Field. As W(ρnj,ρ)<rW(\rho_{n_{j}},\rho)<r for all large jj, (K-u) follows. (K-v) is proved in the same way: (1a) gives E(σn)δ1(b)\mathcal{E}(\sigma_{n})\le\delta^{-1}(\ell'-b'), and vδ+v^{+}_{\delta} is lower semicontinuous on D\mathcal{D} relative to D\mathcal{D} by the second part of Basic Properties of the Delta-Envelopes on the Wasserstein Space §semicontinuity, applied to vv, which is bounded below near each point since bvb'\le v.

Step 4 (The fibre functions). For ζRd\zeta\in\mathbb{R}^{d} let Dζ={ρD:m(ρ)=ζ}\mathcal{D}_{\zeta}=\{\rho\in\mathcal{D}:m(\rho)=\zeta\}. It is nonempty: (τζζ^)#μ^D(\tau_{\zeta-\hat{\zeta}})_{\#}\hat{\mu}\in\mathcal{D} by Penalty Pairs on the Wasserstein Space: the Penalty, Its Score, and Their Domains §translation-invariant, and its mean is ζ^+(ζζ^)=ζ\hat{\zeta}+(\zeta-\hat{\zeta})=\zeta by The Mean of a Square-Integrable Probability Measure, Its Lift, Its Centring, and Functions of the Mean and Centred Integrals as Test Functions §mean. By (1a) the set {Θ(ρ):ρDζ}\{\Theta(\rho):\rho\in\mathcal{D}_{\zeta}\} is bounded above and {Ξ(σ):σDζ}\{\Xi(\sigma):\sigma\in\mathcal{D}_{\zeta}\} bounded below; let U(ζ)RU(\zeta)\in\mathbb{R} be the supremum of the first and V(ζ)RV(\zeta)\in\mathbb{R} the infimum of the second (Approximation Property of the Supremum and the Infimum in R\mathbb{R}).

(4a) The fibre extrema are attained. Let ζRd\zeta\in\mathbb{R}^{d}. By claim 3 of Approximation Property of the Supremum and the Infimum in R\mathbb{R} choose ρnDζ\rho_{n}\in\mathcal{D}_{\zeta} with U(ζ)1/n<Θ(ρn)U(\zeta)-1/n<\Theta(\rho_{n}) for each nn; then U(ζ)1Θ(ρn)U(\zeta)-1\le\Theta(\rho_{n}), as 1/n11/n\le1. (K-u), with the constant sequence of means ζ\zeta and =U(ζ)1\ell=U(\zeta)-1, gives (nj)j(n_{j})_{j} and ρDζ\rho\in\mathcal{D}_{\zeta}. For ε>0\varepsilon>0 take jj so large that Θ(ρnj)<Θ(ρ)+ε\Theta(\rho_{n_{j}})<\Theta(\rho)+\varepsilon and 1/nj<ε1/n_{j}<\varepsilon (the sequence (1/nj)j(1/n_{j})_{j} converges to 00 by A Subsequence of a Convergent Sequence Has the Same Limit); then U(ζ)<Θ(ρ)+2εU(\zeta)<\Theta(\rho)+2\varepsilon. By Comparison of Real Numbers with Arbitrary Positive Slack §slack-above, U(ζ)Θ(ρ)U(ζ)U(\zeta)\le\Theta(\rho)\le U(\zeta), so Θ(ρ)=U(ζ)\Theta(\rho)=U(\zeta). In the same way, with claim 4 of Approximation Property of the Supremum and the Infimum in R\mathbb{R}, (K-v) and Comparison of Real Numbers with Arbitrary Positive Slack §slack-below, there is σDζ\sigma\in\mathcal{D}_{\zeta} with Ξ(σ)=V(ζ)\Xi(\sigma)=V(\zeta).

(4b) UU is upper and VV lower semicontinuous on Rd\mathbb{R}^{d}, in the sense of Upper Semicontinuous Function on a Subset of a Metric Space and Lower Semicontinuous Function on a Subset of a Metric Space in (Rd,dE)(\mathbb{R}^{d},d_{E}), which is the reading of Second-Order Equations on Euclidean Open Sets §extrema. Suppose UU were not upper semicontinuous at ζ\zeta. Then there is ε>0\varepsilon>0 such that for every nn some ζn\zeta_{n} has dE(ζn,ζ)<1/nd_{E}(\zeta_{n},\zeta)<1/n and U(ζ)+εU(ζn)U(\zeta)+\varepsilon\le U(\zeta_{n}); so ζnζ\zeta_{n}\to\zeta. By (4a) choose ρnDζn\rho_{n}\in\mathcal{D}_{\zeta_{n}} with Θ(ρn)=U(ζn)\Theta(\rho_{n})=U(\zeta_{n}). (K-u) with =U(ζ)+ε\ell=U(\zeta)+\varepsilon gives ρDζ\rho\in\mathcal{D}_{\zeta} and, for large jj, U(ζ)+εΘ(ρnj)<Θ(ρ)+ε2U(ζ)+ε2U(\zeta)+\varepsilon\le\Theta(\rho_{n_{j}})<\Theta(\rho)+\tfrac{\varepsilon}{2}\le U(\zeta)+\tfrac{\varepsilon}{2}, which is impossible as ε2<ε\tfrac{\varepsilon}{2}<\varepsilon (claim 8 of Elementary Order Arithmetic in an Ordered Field). The lower semicontinuity of VV follows in the same way from (4a) and (K-v).

Step 5 (Ishii's lemma on the means; claim 2). Let ζ,ωRd\zeta,\omega\in\mathbb{R}^{d} and, by (4a), ρDζ\rho\in\mathcal{D}_{\zeta} and σDω\sigma\in\mathcal{D}_{\omega} with Θ(ρ)=U(ζ)\Theta(\rho)=U(\zeta) and Ξ(σ)=V(ω)\Xi(\sigma)=V(\omega). By (1b),

U(ζ)V(ω)α2ζω2M.(5a)U(\zeta)-V(\omega)-\tfrac{\alpha}{2}\lVert\zeta-\omega\rVert^{2}\le M.\qquad(5\mathrm{a})

Since μ^Dζ^\hat{\mu}\in\mathcal{D}_{\hat{\zeta}} and ν^Dω^\hat{\nu}\in\mathcal{D}_{\hat{\omega}}, Θ(μ^)U(ζ^)\Theta(\hat{\mu})\le U(\hat{\zeta}) and V(ω^)Ξ(ν^)V(\hat{\omega})\le\Xi(\hat{\nu}), so (1c) and (5a) give

U(ζ^)V(ω^)α2ζ^ω^2=M.(5b)U(\hat{\zeta})-V(\hat{\omega})-\tfrac{\alpha}{2}\lVert\hat{\zeta}-\hat{\omega}\rVert^{2}=M.\qquad(5\mathrm{b})

Apply Ishii's Lemma: Test Data and Matrix Bounds at a Maximum of a Quadratically Penalised Difference with n=dn=d, the open set Ω=Rd\Omega=\mathbb{R}^{d} (claim 1 of Euclidean Space is Open in Itself, and CkC^k Maps are Continuous), the functions UU and VV (Step 4), x^=ζ^\hat{x}=\hat{\zeta}, y^=ω^\hat{y}=\hat{\omega} and radius 11: its hypothesis holds by (5a) and (5b) at all points. Let X,YS(d)\mathbb{X},\mathbb{Y}\in\mathcal{S}(d) be the matrices it provides and p=α(ζ^ω^)p=\alpha(\hat{\zeta}-\hat{\omega}). Its claims Ishii's Lemma: Test Data and Matrix Bounds at a Maximum of a Quadratically Penalised Difference §ordering, Ishii's Lemma: Test Data and Matrix Bounds at a Maximum of a Quadratically Penalised Difference §norm-bound and Ishii's Lemma: Test Data and Matrix Bounds at a Maximum of a Quadratically Penalised Difference §quadratic-bound are the four conditions of The Second-Order Structure Condition at Optimally Coupled Pairs on the Lift of the Wasserstein Space §admitted; this is claim 2. By Ishii's Lemma: Test Data and Matrix Bounds at a Maximum of a Quadratically Penalised Difference §test-data, (ζ^,U(ζ^),p,X)(\hat{\zeta},U(\hat{\zeta}),p,\mathbb{X}) is approximable by test data from above for UU and (ω^,V(ω^),p,Y)(\hat{\omega},V(\hat{\omega}),p,\mathbb{Y}) from below for VV, with open set Rd\mathbb{R}^{d}.

Step 6 (Test functions at fibre maximisers and minimisers). Let nNn\in\mathbb{N}. By Quadruple Approximable by Test-Function Data §above with ε=1/n\varepsilon=1/n there are ζnRd\zeta_{n}\in\mathbb{R}^{d}, a function χn\chi_{n} of class C2C^{2} on Rd\mathbb{R}^{d} and a positive rnr_{n} such that U(ζ)χn(ζ)U(ζn)χn(ζn)U(\zeta)-\chi_{n}(\zeta)\le U(\zeta_{n})-\chi_{n}(\zeta_{n}) whenever dE(ζ,ζn)<rnd_{E}(\zeta,\zeta_{n})<r_{n}, and

dE(ζn,ζ^)<1n,U(ζn)U(ζ^)<1n,Dχn(ζn)p<1n,D2χn(ζn)X<1n,d_{E}(\zeta_{n},\hat{\zeta})<\tfrac1n,\qquad|U(\zeta_{n})-U(\hat{\zeta})|<\tfrac1n,\qquad\lVert D\chi_{n}(\zeta_{n})-p\rVert<\tfrac1n,\qquad\lVert D^{2}\chi_{n}(\zeta_{n})-\mathbb{X}\rVert<\tfrac1n,

the last because dS(d)(P,P)=PPd_{\mathcal{S}(d)}(P,P')=\lVert P-P'\rVert (Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §matrices). By (4a) choose ρnDζn\rho_{n}\in\mathcal{D}_{\zeta_{n}} with Θ(ρn)=U(ζn)\Theta(\rho_{n})=U(\zeta_{n}), and let φn=αA+χnm\varphi_{n}=\alpha A+\chi_{n}\circ m. By Step 2, The Squared Wasserstein Distance to a Fixed Measure and Functions of the Mean are Intrinsic Test Functions §mean and Restrictions, Sums, Real Multiples and Differences of Intrinsic Test Functions on the Wasserstein Space §linear, φn\varphi_{n} is an intrinsic test function on D\mathcal{D} with

φn(ρn)=αA(ρn)+Dχn(ζn),Hφn(ρn)=α0d+D2χn(ζn)=D2χn(ζn).\nabla\varphi_{n}(\rho_{n})=\alpha\nabla A(\rho_{n})+D\chi_{n}(\zeta_{n}),\qquad H_{\varphi_{n}}(\rho_{n})=\alpha0_{d}+D^{2}\chi_{n}(\zeta_{n})=D^{2}\chi_{n}(\zeta_{n}).

For ρD\rho'\in\mathcal{D} with W(ρ,ρn)<rnW(\rho',\rho_{n})<r_{n} we have dE(m(ρ),ζn)W(ρ,ρn)<rnd_{E}(m(\rho'),\zeta_{n})\le W(\rho',\rho_{n})<r_{n} by The Mean of a Square-Integrable Probability Measure, Its Lift, Its Centring, and Functions of the Mean and Centred Integrals as Test Functions §mean, so, by the definition of UU and ρDm(ρ)\rho'\in\mathcal{D}_{m(\rho')},

uδ(ρ)φn(ρ)=Θ(ρ)χn(m(ρ))U(m(ρ))χn(m(ρ))U(ζn)χn(ζn)=uδ(ρn)φn(ρn).u^{-}_{\delta}(\rho')-\varphi_{n}(\rho')=\Theta(\rho')-\chi_{n}(m(\rho'))\le U(m(\rho'))-\chi_{n}(m(\rho'))\le U(\zeta_{n})-\chi_{n}(\zeta_{n})=u^{-}_{\delta}(\rho_{n})-\varphi_{n}(\rho_{n}).

So uδφnu^{-}_{\delta}-\varphi_{n} has a local maximum relative to D\mathcal{D} at ρn\rho_{n} (Local Maximum of a Function Relative to a Subset of a Metric Space).

Symmetrically, Quadruple Approximable by Test-Function Data §below with ε=1/n\varepsilon=1/n gives ωn\omega_{n}, χn\chi'_{n} of class C2C^{2} on Rd\mathbb{R}^{d} and rn>0r'_{n}>0 with V(ω)χn(ω)V(ωn)χn(ωn)V(\omega)-\chi'_{n}(\omega)\ge V(\omega_{n})-\chi'_{n}(\omega_{n}) whenever dE(ω,ωn)<rnd_{E}(\omega,\omega_{n})<r'_{n}, and the four bounds with ωn,ω^,V,χn,Y\omega_{n},\hat{\omega},V,\chi'_{n},\mathbb{Y} in place of ζn,ζ^,U,χn,X\zeta_{n},\hat{\zeta},U,\chi_{n},\mathbb{X}. Choose σnDωn\sigma_{n}\in\mathcal{D}_{\omega_{n}} with Ξ(σn)=V(ωn)\Xi(\sigma_{n})=V(\omega_{n}) and let ψn=(α)A+χnm\psi_{n}=(-\alpha)A+\chi'_{n}\circ m, an intrinsic test function on D\mathcal{D} with ψn(σn)=αA(σn)+Dχn(ωn)\nabla\psi_{n}(\sigma_{n})=-\alpha\nabla A(\sigma_{n})+D\chi'_{n}(\omega_{n}) and Hψn(σn)=D2χn(ωn)H_{\psi_{n}}(\sigma_{n})=D^{2}\chi'_{n}(\omega_{n}). For σD\sigma'\in\mathcal{D} with W(σ,σn)<rnW(\sigma',\sigma_{n})<r'_{n}, vδ+(σ)ψn(σ)=Ξ(σ)χn(m(σ))V(m(σ))χn(m(σ))V(ωn)χn(ωn)=vδ+(σn)ψn(σn)v^{+}_{\delta}(\sigma')-\psi_{n}(\sigma')=\Xi(\sigma')-\chi'_{n}(m(\sigma'))\ge V(m(\sigma'))-\chi'_{n}(m(\sigma'))\ge V(\omega_{n})-\chi'_{n}(\omega_{n})=v^{+}_{\delta}(\sigma_{n})-\psi_{n}(\sigma_{n}), so vδ+ψnv^{+}_{\delta}-\psi_{n} has a local minimum relative to D\mathcal{D} at σn\sigma_{n} (Local Minimum of a Function Relative to a Subset of a Metric Space).

Step 7 (The limits ρ\rho^{*}, σ\sigma^{*}; claim 1). We have m(ρn)=ζnζ^m(\rho_{n})=\zeta_{n}\to\hat{\zeta} and Θ(ρn)=U(ζn)>U(ζ^)1/nU(ζ^)1\Theta(\rho_{n})=U(\zeta_{n})>U(\hat{\zeta})-1/n\ge U(\hat{\zeta})-1. (K-u) gives a strictly increasing (nj)j(n_{j})_{j} and ρDζ^\rho^{*}\in\mathcal{D}_{\hat{\zeta}} with ρnjρ\rho_{n_{j}}\to\rho^{*}; exactly as in (4a), Θ(ρ)=U(ζ^)\Theta(\rho^{*})=U(\hat{\zeta}), and then Θ(ρnj)Θ(ρ)=U(ζnj)U(ζ^)<1/nj|\Theta(\rho_{n_{j}})-\Theta(\rho^{*})|=|U(\zeta_{n_{j}})-U(\hat{\zeta})|<1/n_{j}. Symmetrically, (K-v) gives (nj)j(n'_{j})_{j} and σDω^\sigma^{*}\in\mathcal{D}_{\hat{\omega}} with σnjσ\sigma_{n'_{j}}\to\sigma^{*}, Ξ(σ)=V(ω^)\Xi(\sigma^{*})=V(\hat{\omega}) and Ξ(σnj)Ξ(σ)<1/nj|\Xi(\sigma_{n'_{j}})-\Xi(\sigma^{*})|<1/n'_{j}. By (1b) and (5b),

M=Θ(ρ)Ξ(σ)α2ζ^ω^2Ψ(ρ,σ)M,M=\Theta(\rho^{*})-\Xi(\sigma^{*})-\tfrac{\alpha}{2}\lVert\hat{\zeta}-\hat{\omega}\rVert^{2}\le\Psi(\rho^{*},\sigma^{*})\le M,

which is claim 1; and equality in the first inequality means, by Step 1, that equality holds in The Displacement Midpoint and the Midpoint Split of the Squared Wasserstein Distance §split for (ρ,σ)(\rho^{*},\sigma^{*}), hence also for (σ,ρ)(\sigma^{*},\rho^{*}), both sides of that inequality being symmetric in the pair by The Quadratic Wasserstein Distance is a Metric on the Wasserstein Space §symmetry and claim 5 of Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n.

Step 8 (The limiting gradients). As D\mathcal{D} has the map property and ρ,σD\rho^{*},\sigma^{*}\in\mathcal{D}, the pairs (ρ,m^)(\rho^{*},\hat{m}), (ρ,σ)(\rho^{*},\sigma^{*}), (σ,m^)(\sigma^{*},\hat{m}) and (σ,ρ)(\sigma^{*},\rho^{*}) are uniquely mapped (The Map Property of a Set of Probability Measures §map-property). By The Displacement Midpoint and the Midpoint Split of the Squared Wasserstein Distance §equality for (ρ,σ)(\rho^{*},\sigma^{*}) with the optimal maps GρG_{\rho^{*}} and SS, the constant ee there has value 12(ζ^+ω^)c=0Rd\tfrac12(\hat{\zeta}+\hat{\omega})-c=0_{\mathbb{R}^{d}}, so 2(idGρ)=idS2(\mathrm{id}-G_{\rho^{*}})=\mathrm{id}-S in L2(ρ;Rd)L^{2}(\rho^{*};\mathbb{R}^{d}); likewise, for (σ,ρ)(\sigma^{*},\rho^{*}) with GσG_{\sigma^{*}} and SS', 2(idGσ)=idS2(\mathrm{id}-G_{\sigma^{*}})=\mathrm{id}-S' in L2(σ;Rd)L^{2}(\sigma^{*};\mathbb{R}^{d}). As 2(ζ^c)=ζ^ω^2(\hat{\zeta}-c)=\hat{\zeta}-\hat{\omega} and 2(ω^c)=ω^ζ^2(\hat{\omega}-c)=\hat{\omega}-\hat{\zeta}, Step 2 gives A(ρ)=(idS)(ζ^ω^)\nabla A(\rho^{*})=(\mathrm{id}-S)-(\hat{\zeta}-\hat{\omega}) and A(σ)=(idS)(ω^ζ^)\nabla A(\sigma^{*})=(\mathrm{id}-S')-(\hat{\omega}-\hat{\zeta}), hence, with p=α(ζ^ω^)p=\alpha(\hat{\zeta}-\hat{\omega}),

αA(ρ)+p=α(idS) in L2(ρ;Rd),αA(σ)+p=α(Sid) in L2(σ;Rd).(8a)\alpha\nabla A(\rho^{*})+p=\alpha(\mathrm{id}-S)\ \text{in }L^{2}(\rho^{*};\mathbb{R}^{d}),\qquad-\alpha\nabla A(\sigma^{*})+p=\alpha(S'-\mathrm{id})\ \text{in }L^{2}(\sigma^{*};\mathbb{R}^{d}).\qquad(8\mathrm{a})

Step 9 (Claim 3). For each jj let πjΠ(ρnj,ρ)\pi_{j}\in\Pi(\rho_{n_{j}},\rho^{*}) be optimal (Existence of an Optimal Coupling of Two Probability Measures with Finite Second Moment), so I(πj)=W(ρnj,ρ)20I(\pi_{j})=W(\rho_{n_{j}},\rho^{*})^{2}\to0 (Optimal Coupling of Two Probability Measures with Finite Second Moment §optimal). By property (c) of Intrinsic Test Functions on the Wasserstein Space and Their Translation Hessians §test for AA on D\mathcal{D}, the discrepancy DjD_{j} of A(ρnj)\nabla A(\rho_{n_{j}}) and A(ρ)\nabla A(\rho^{*}) along πj\pi_{j} converges to 00. Let EjE_{j} be the discrepancy of φnj(ρnj)\nabla\varphi_{n_{j}}(\rho_{n_{j}}) and α(idS)\alpha(\mathrm{id}-S) along πj\pi_{j}, which is the integral in claim 3. It does not depend on representatives (The Discrepancy of Two Square-Integrable Vector Fields Along a Coupling of Their Base Measures §well-defined), so by (8a) and Step 6 its integrand may be taken to be F1(z)+F2(z)2\lVert F_{1}(z)+F_{2}(z)\rVert^{2} with

F1(z)=α(A(ρnj)(x)A(ρ)(y)),F2(z)=Dχnj(ζnj)p.F_{1}(z)=\alpha\bigl(\nabla A(\rho_{n_{j}})(x)-\nabla A(\rho^{*})(y)\bigr),\qquad F_{2}(z)=D\chi_{n_{j}}(\zeta_{n_{j}})-p .

Both are Borel and square-integrable against πj\pi_{j}, with F1πj=αDj\lVert F_{1}\rVert_{\pi_{j}}=\alpha\sqrt{D_{j}} (claim 5 of Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n) and F2πj=Dχnj(ζnj)p<1/nj\lVert F_{2}\rVert_{\pi_{j}}=\lVert D\chi_{n_{j}}(\zeta_{n_{j}})-p\rVert<1/n_{j}, so the triangle inequality The Norm Metric of a Real Inner Product Space: Triangle Inequalities, Limits and Continuity §triangle in the real Hilbert space L2(πj;Rd)L^{2}(\pi_{j};\mathbb{R}^{d}) (Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §fields) gives Ej<αDj+1/nj\sqrt{E_{j}}<\alpha\sqrt{D_{j}}+1/n_{j}. Moreover uδ(ρnj)uδ(ρ)=(Θ(ρnj)Θ(ρ))+α(A(ρnj)A(ρ))u^{-}_{\delta}(\rho_{n_{j}})-u^{-}_{\delta}(\rho^{*})=\bigl(\Theta(\rho_{n_{j}})-\Theta(\rho^{*})\bigr)+\alpha\bigl(A(\rho_{n_{j}})-A(\rho^{*})\bigr), where the first difference has absolute value below 1/nj1/n_{j} (Step 7) and the second converges to 00 by the continuity of AA.

Let ε>0\varepsilon>0. Each of the real sequences (I(πj))j(I(\pi_{j}))_{j}, (Dj)j(\sqrt{D_{j}})_{j} (claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field), (1/nj)j(1/n_{j})_{j} and (A(ρnj)A(ρ))j(A(\rho_{n_{j}})-A(\rho^{*}))_{j} converges to 00, so we may fix jj with I(πj)<ε2I(\pi_{j})<\varepsilon^{2}, 1/nj<ε21/n_{j}<\tfrac{\varepsilon}{2}, αDj<ε2\alpha\sqrt{D_{j}}<\tfrac{\varepsilon}{2} and αA(ρnj)A(ρ)<ε2\alpha|A(\rho_{n_{j}})-A(\rho^{*})|<\tfrac{\varepsilon}{2}. Put ρ=ρnj\rho=\rho_{n_{j}}, φ=φnj\varphi=\varphi_{n_{j}} and π=πj\pi=\pi_{j}. By Step 6, uδφu^{-}_{\delta}-\varphi has a local maximum relative to D\mathcal{D} at ρ\rho; uδ(ρ)uδ(ρ)<ε|u^{-}_{\delta}(\rho)-u^{-}_{\delta}(\rho^{*})|<\varepsilon by claim 5 of Properties of the Absolute Value in an Ordered Field; Ej<ε\sqrt{E_{j}}<\varepsilon, so Ej<ε2E_{j}<\varepsilon^{2} (claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field); and Hφ(ρ)X=D2χnj(ζnj)X<1/nj<ε\lVert H_{\varphi}(\rho)-\mathbb{X}\rVert=\lVert D^{2}\chi_{n_{j}}(\zeta_{n_{j}})-\mathbb{X}\rVert<1/n_{j}<\varepsilon. This is claim 3.

Step 10 (Claim 4). The same argument, with σnj\sigma_{n'_{j}}, σ\sigma^{*}, optimal couplings γjΠ(σnj,σ)\gamma_{j}\in\Pi(\sigma_{n'_{j}},\sigma^{*}), ψnj\psi_{n'_{j}}, Ξ\Xi, vδ+=ΞαAv^{+}_{\delta}=\Xi-\alpha A, the second identity of (8a), the fields F1(z)=α(A(σnj)(x)A(σ)(y))F_{1}(z)=-\alpha\bigl(\nabla A(\sigma_{n'_{j}})(x)-\nabla A(\sigma^{*})(y)\bigr) and F2(z)=Dχnj(ωnj)pF_{2}(z)=D\chi'_{n'_{j}}(\omega_{n'_{j}})-p, the local minimum of Step 6 and Hψnj(σnj)=D2χnj(ωnj)H_{\psi_{n'_{j}}}(\sigma_{n'_{j}})=D^{2}\chi'_{n'_{j}}(\omega_{n'_{j}}), gives claim 4.

Please log in to copy this version.

Citations

Loading…

Dependency Graph

0 prerequisites

Prerequisites

Loading...

Comments

Loading…