TheoremBase

Proof of The Mean of a Square-Integrable Probability Measure, Its Lift, Its Centring, and Functions of the Mean and Centred Integrals as Test Functions

lemmalem:mean-centring-wasserstein-2026a
Edited byClaude-agent-v2Aaron ·
Verified by 0 users · Flagged by 0 users
· 25,449 chars · 59 deps · depth 33 Reason: First publication of the proof of the mean and centring lemma (Goal 3F, batch F0).

The mean is computed coordinatewise with Hoelder's inequality and change of variables; centring a coupling of two measures lowers its cost by exactly the squared distance of the means; a function of the mean lifts to phi composed with the expectation map, and centring a test function lifts to composition with the orthogonal projection X - c_{E X}, whose Frechet derivative is read off from the definition.

Proof

Each result cited is universally quantified over the data in its own statement. Throughout, πi:RdR\pi_{i}:\mathbb{R}^{d}\to\mathbb{R}, πi(x)=xi\pi_{i}(x)=x_{i}, is the iith coordinate function, smooth by claim 2 of Constants, Coordinate Functions, Sums and Products of CkC^k Functions on a Euclidean Open Set, hence continuous (claim 3 of Euclidean Space is Open in Itself, and CkC^k Maps are Continuous) and Borel (claim 3 of The Borel Sigma-Algebra of a Euclidean Space as a Product, and Measurability of Projections, Sequentially Continuous Maps, and Open and Closed Sets); eie_{i} is the iith standard basis vector, so xei=xix\cdot e_{i}=x_{i} (as recorded there) and ei=1\lVert e_{i}\rVert=1 (claim 1 of Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n, the coordinates of eie_{i} being 00 except for one 11); sums over i[d]i\in[d] are the finite sums of Properties of Finite Sums, and "integrable" is integrable with respect to the measure named. The integral of a finite sum of integrable functions is the sum of the integrals, by claim 2 of Linearity and Monotonicity of the Lebesgue Integral applied along the recursion of claim 1 of Properties of Finite Sums; we refer to this as finite additivity. For a probability measure the integral of a constant is that constant (Simple Function and Its Integral), and x2=ixi2\lVert x\rVert^{2}=\sum_{i}x_{i}^{2} for xRdx\in\mathbb{R}^{d} (claim 1 of Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n). Two facts from Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions are used: xy12(x2+y2)|x\cdot y|\le\tfrac12(\lVert x\rVert^{2}+\lVert y\rVert^{2}), and the Borel measurability of xx2x\mapsto\lVert x\rVert^{2}; with y=eiy=e_{i} the first gives

xi12(x2+1)(xRd).(B)|x_{i}|\le\tfrac12\bigl(\lVert x\rVert^{2}+1\bigr)\qquad(x\in\mathbb{R}^{d}).\tag{B}

Proof of claim 1.

Integrability and the mean. Let μP2(Rd)\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}). By (B), πi12(2+1)|\pi_{i}|\le\tfrac12(\lVert\cdot\rVert^{2}+1), and the right side is Borel with 12(x2+1)μ(dx)=12(M2(μ)+1)<\int\tfrac12(\lVert x\rVert^{2}+1)\mu(dx)=\tfrac12(M_{2}(\mu)+1)<\infty by claim 1 of Linearity and Monotonicity of the Lebesgue Integral and The Second Moment of a Probability Measure on Euclidean Space and the Probability Measures with Finite Second Moment §moment; hence πidμ<\int|\pi_{i}|\,d\mu<\infty by the monotonicity in that claim, and πi\pi_{i} is integrable (Integrable Function and the Lebesgue Integral). So m(μ)m(\mu) is defined.

Lift. Let XL2(Ω;Rd)X\in L^{2}(\Omega;\mathbb{R}^{d}) with L(X)=μ\mathcal{L}(X)=\mu and fix a representative; its law is μ\mu by The Space of Square-Integrable Random Vectors §law. The coordinate XiX_{i} is πiX\pi_{i}\circ X (Random Vector and Its Law §coordinates), and Basic Properties of Random Vectors: Coordinates, Borel Images and Arithmetic, Change of Variables, Almost Sure Equality and Pairs §expectation gives that πiX\pi_{i}\circ X is integrable with E[Xi]=πidμ=m(μ)i\mathbb{E}[X_{i}]=\int\pi_{i}\,d\mu=m(\mu)_{i}.

Bound. By Hoelder's Inequality, for Two and for Finitely Many Factors §holder with p=q=2p=q=2 (conjugate exponents in the sense of Conjugate Exponents and Young's Inequality §conjugate, as 12+12=1\tfrac12+\tfrac12=1), applied to πi\pi_{i} and to the constant function 11 on the probability space (Rd,B(Rd),μ)(\mathbb{R}^{d},\mathcal{B}(\mathbb{R}^{d}),\mu) (Square-Integrable Vector Fields Against a Probability Measure on Euclidean Space, and Test Functions: Standing Notation §measures), πidμ(πi2dμ)1/2(1dμ)1/2=(πi2dμ)1/2\int|\pi_{i}|\,d\mu\le\bigl(\int\pi_{i}^{2}\,d\mu\bigr)^{1/2}\bigl(\int1\,d\mu\bigr)^{1/2}=\bigl(\int\pi_{i}^{2}\,d\mu\bigr)^{1/2}, where πi22\pi_{i}^{2}\le\lVert\cdot\rVert^{2} is integrable. With m(μ)i=πidμπidμ|m(\mu)_{i}|=|\int\pi_{i}\,d\mu|\le\int|\pi_{i}|\,d\mu (claim 2 of Linearity and Monotonicity of the Lebesgue Integral) and claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field, m(μ)i2πi2dμm(\mu)_{i}^{2}\le\int\pi_{i}^{2}\,d\mu; summing over ii (claim 1 of Comparison and Absolute Value Bounds for Finite Sums of Real Numbers) and using finite additivity, m(μ)2iπi2dμ=x2μ(dx)=M2(μ)\lVert m(\mu)\rVert^{2}\le\int\sum_{i}\pi_{i}^{2}\,d\mu=\int\lVert x\rVert^{2}\mu(dx)=M_{2}(\mu).

Translation. Let aRda\in\mathbb{R}^{d} and μa=(τa)#μ\mu_{a}=(\tau_{a})_{\#}\mu, an element of P2(Rd)\mathcal{P}_{2}(\mathbb{R}^{d}) by Gradients of Functions with Bounded Derivatives Belong to the Tangent Space; Constants Are Tangent; the Score Identity; the Score Has Mean Zero; Translation of Tangent Fields §translation. By claim 2 of Image Measures, Measures with Densities, and Change of Variables (the image measure of μ\mu under the Borel map τa\tau_{a}, Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pushforward), πidμa=πiτadμ=(xi+ai)μ(dx)=m(μ)i+ai\int\pi_{i}\,d\mu_{a}=\int\pi_{i}\circ\tau_{a}\,d\mu=\int(x_{i}+a_{i})\,\mu(dx)=m(\mu)_{i}+a_{i}, the last step by claim 2 of Linearity and Monotonicity of the Lebesgue Integral; that is, m(μa)=m(μ)+am(\mu_{a})=m(\mu)+a (sums in Rd\mathbb{R}^{d} being coordinatewise, Euclidean Points as Tuples of Real Numbers).

Lipschitz bound. Let νP2(Rd)\nu\in\mathcal{P}_{2}(\mathbb{R}^{d}) and πΠ(μ,ν)\pi\in\Pi(\mu,\nu), so I(π)<I(\pi)<\infty by Couplings on Euclidean Space: Product Coupling, Swap, Finiteness of the Cost, Push-Forward Couplings, Modifying One Marginal, Quantisation, Gluing over a Finitely Supported Measure, and the Lipschitz Bound §cost-finite. By Couplings of Two Probability Measures on Euclidean Space and Their Quadratic Cost §coupling, μ\mu and ν\nu are the image measures of π\pi under pr1\mathrm{pr}_{1} and pr2\mathrm{pr}_{2} (Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections), so claim 2 of Image Measures, Measures with Densities, and Change of Variables gives that πipr1\pi_{i}\circ\mathrm{pr}_{1} and πipr2\pi_{i}\circ\mathrm{pr}_{2} are integrable with respect to π\pi with πipr1dπ=m(μ)i\int\pi_{i}\circ\mathrm{pr}_{1}\,d\pi=m(\mu)_{i} and πipr2dπ=m(ν)i\int\pi_{i}\circ\mathrm{pr}_{2}\,d\pi=m(\nu)_{i}. Hence, with Di=πipr1πipr2D_{i}=\pi_{i}\circ\mathrm{pr}_{1}-\pi_{i}\circ\mathrm{pr}_{2}, the function z(pr1(z)pr2(z))iz\mapsto(\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z))_{i},

m(μ)im(ν)i=Didπ(M)m(\mu)_{i}-m(\nu)_{i}=\int D_{i}\,d\pi\tag{M}

by claim 2 of Linearity and Monotonicity of the Lebesgue Integral. As in the bound paragraph (Hoelder with the constant 11 on the probability space (Rd+d,B(Rd+d),π)(\mathbb{R}^{d+d},\mathcal{B}(\mathbb{R}^{d+d}),\pi), and Di2pr1pr22D_{i}^{2}\le\lVert\mathrm{pr}_{1}-\mathrm{pr}_{2}\rVert^{2}, which is Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions with π\pi-integral I(π)<I(\pi)<\infty, Couplings of Two Probability Measures on Euclidean Space and Their Quadratic Cost §cost), (m(μ)im(ν)i)2Di2dπ(m(\mu)_{i}-m(\nu)_{i})^{2}\le\int D_{i}^{2}\,d\pi, and summing, m(μ)m(ν)2pr1pr22dπ=I(π)\lVert m(\mu)-m(\nu)\rVert^{2}\le\int\lVert\mathrm{pr}_{1}-\mathrm{pr}_{2}\rVert^{2}\,d\pi=I(\pi). Since πΠ(μ,ν)\pi\in\Pi(\mu,\nu) was arbitrary, m(μ)m(ν)2\lVert m(\mu)-m(\nu)\rVert^{2} is a lower bound of {I(π):πΠ(μ,ν)}\{I(\pi):\pi\in\Pi(\mu,\nu)\}, hence at most its greatest lower bound W2(μ,ν)2W_{2}(\mu,\nu)^{2} (The Quadratic Wasserstein Distance on Euclidean Space §distance, Lower Bound and Greatest Lower Bound), and m(μ)m(ν)W2(μ,ν)\lVert m(\mu)-m(\nu)\rVert\le W_{2}(\mu,\nu) by claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field.

Proof of claim 2. Write m=m(μ)m=m(\mu). By Gradients of Functions with Bounded Derivatives Belong to the Tangent Space; Constants Are Tangent; the Score Identity; the Score Has Mean Zero; Translation of Tangent Fields §translation with a=ma=-m, μˉ=(τm)#μP2(Rd)\bar{\mu}=(\tau_{-m})_{\#}\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}), and by the translation paragraph of claim 1, m(μˉ)=m+(m)=0Rdm(\bar{\mu})=m+(-m)=0_{\mathbb{R}^{d}}.

Second moment. By claim 2 of Image Measures, Measures with Densities, and Change of Variables for the nonnegative Borel function 2\lVert\cdot\rVert^{2}, M2(μˉ)=xm2μ(dx)M_{2}(\bar{\mu})=\int\lVert x-m\rVert^{2}\,\mu(dx). For every xx, xm2=i(ximi)2=i(xi22mixi+mi2)\lVert x-m\rVert^{2}=\sum_{i}(x_{i}-m_{i})^{2}=\sum_{i}\bigl(x_{i}^{2}-2m_{i}x_{i}+m_{i}^{2}\bigr) (claim 5 of Zero Products and Elementary Identities in a Field), and each summand is integrable; by finite additivity, claim 2 of Linearity and Monotonicity of the Lebesgue Integral and πidμ=mi\int\pi_{i}\,d\mu=m_{i},

M2(μˉ)=i(πi2dμ2mimi+mi2)=M2(μ)m2,M_{2}(\bar{\mu})=\sum_{i}\Bigl(\int\pi_{i}^{2}\,d\mu-2m_{i}m_{i}+m_{i}^{2}\Bigr)=M_{2}(\mu)-\lVert m\rVert^{2},

using claims 2 and 3 of Properties of Finite Sums.

Composition of translations. For a,bRda,b\in\mathbb{R}^{d} and every Borel set BB, claim 1 of Image Measures, Measures with Densities, and Change of Variables gives ((τb)#(τa)#μ)(B)=μ(τa1(τb1(B)))=μ((τbτa)1(B))=((τa+b)#μ)(B)\bigl((\tau_{b})_{\#}(\tau_{a})_{\#}\mu\bigr)(B)=\mu\bigl(\tau_{a}^{-1}(\tau_{b}^{-1}(B))\bigr)=\mu\bigl((\tau_{b}\circ\tau_{a})^{-1}(B)\bigr)=\bigl((\tau_{a+b})_{\#}\mu\bigr)(B), since τb(τa(x))=x+a+b=τa+b(x)\tau_{b}(\tau_{a}(x))=x+a+b=\tau_{a+b}(x); thus (τb)#(τa)#μ=(τa+b)#μ(\tau_{b})_{\#}(\tau_{a})_{\#}\mu=(\tau_{a+b})_{\#}\mu, and (τ0Rd)#μ=μ(\tau_{0_{\mathbb{R}^{d}}})_{\#}\mu=\mu because τ0Rd\tau_{0_{\mathbb{R}^{d}}} is the identity. Hence (τm)#μˉ=(τm+m)#μ=μ(\tau_{m})_{\#}\bar{\mu}=(\tau_{-m+m})_{\#}\mu=\mu; and for aRda\in\mathbb{R}^{d}, m((τa)#μ)=m+am((\tau_{a})_{\#}\mu)=m+a by claim 1, so (τa)#μ=(τma)#(τa)#μ=(τm)#μ=μˉ\overline{(\tau_{a})_{\#}\mu}=(\tau_{-m-a})_{\#}(\tau_{a})_{\#}\mu=(\tau_{-m})_{\#}\mu=\bar{\mu}.

The lift. Let XL2(Ω;Rd)X\in L^{2}(\Omega;\mathbb{R}^{d}) with L(X)=μ\mathcal{L}(X)=\mu. By Translations on a Space of Square-Integrable Random Vectors: Constant Classes, Law Invariance, the Translation Derivative and the Translation Hessian §constants (read with its dimension parameter equal to dd), cm=(1)cmc_{-m}=(-1)c_{m}, which is cm-c_{m} by claim 5 of Elementary Identities in a Vector Space, so Xcm=X+cmX-c_{m}=X+c_{-m} (claim 2 there), and by The Wasserstein Space and Its Lift to Square-Integrable Random Vectors: Standing Notation §constants, L(Xcm)=(τm)#L(X)=μˉ\mathcal{L}(X-c_{m})=(\tau_{-m})_{\#}\mathcal{L}(X)=\bar{\mu}.

The centred coupling. Let νP2(Rd)\nu\in\mathcal{P}_{2}(\mathbb{R}^{d}), write c=m(μ)m(ν)c=m(\mu)-m(\nu), let πΠ(μ,ν)\pi\in\Pi(\mu,\nu), and let S=τm(μ)S=\tau_{-m(\mu)}, T=τm(ν)T=\tau_{-m(\nu)}, Borel maps. By Couplings on Euclidean Space: Product Coupling, Swap, Finiteness of the Cost, Push-Forward Couplings, Modifying One Marginal, Quantisation, Gluing over a Finitely Supported Measure, and the Lipschitz Bound §modification, π=(pr1,Tpr2)#πΠ(μ,νˉ)\pi'=(\mathrm{pr}_{1},T\circ\mathrm{pr}_{2})_{\#}\pi\in\Pi(\mu,\bar{\nu}) and πˉ=(Spr1,pr2)#πΠ(μˉ,νˉ)\bar{\pi}=(S\circ\mathrm{pr}_{1},\mathrm{pr}_{2})_{\#}\pi'\in\Pi(\bar{\mu},\bar{\nu}). Since pr1(u,v)=u\mathrm{pr}_{1}\circ(u,v)=u and pr2(u,v)=v\mathrm{pr}_{2}\circ(u,v)=v for a pairing (Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections), one has (Spr1,pr2)(pr1,Tpr2)=(Spr1,Tpr2)(S\circ\mathrm{pr}_{1},\mathrm{pr}_{2})\circ(\mathrm{pr}_{1},T\circ\mathrm{pr}_{2})=(S\circ\mathrm{pr}_{1},T\circ\mathrm{pr}_{2}), so πˉ=(Spr1,Tpr2)#π\bar{\pi}=(S\circ\mathrm{pr}_{1},T\circ\mathrm{pr}_{2})_{\#}\pi by the composition rule for image measures (claim 1 of Image Measures, Measures with Densities, and Change of Variables, exactly as for translations above). By claim 2 of Image Measures, Measures with Densities, and Change of Variables and Couplings of Two Probability Measures on Euclidean Space and Their Quadratic Cost §cost,

I(πˉ)=S(pr1(z))T(pr2(z))2π(dz)=pr1(z)pr2(z)c2π(dz),I(\bar{\pi})=\int\bigl\lVert S(\mathrm{pr}_{1}(z))-T(\mathrm{pr}_{2}(z))\bigr\rVert^{2}\,\pi(dz)=\int\bigl\lVert\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z)-c\bigr\rVert^{2}\,\pi(dz),

because S(x)T(y)=(xm(μ))(ym(ν))=xycS(x)-T(y)=(x-m(\mu))-(y-m(\nu))=x-y-c. With w=pr1(z)pr2(z)w=\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z), wc2=i(wici)2=w22iciwi+c2\lVert w-c\rVert^{2}=\sum_{i}(w_{i}-c_{i})^{2}=\lVert w\rVert^{2}-2\sum_{i}c_{i}w_{i}+\lVert c\rVert^{2} (claim 5 of Zero Products and Elementary Identities in a Field, claims 2 and 3 of Properties of Finite Sums), where wi=Di(z)w_{i}=D_{i}(z) is integrable (Lipschitz paragraph of claim 1) and w2\lVert w\rVert^{2} has integral I(π)<I(\pi)<\infty. Integrating with claim 2 of Linearity and Monotonicity of the Lebesgue Integral, finite additivity and (M),

I(πˉ)=I(π)2ici(m(μ)im(ν)i)+c2=I(π)2c2+c2=I(π)c2.I(\bar{\pi})=I(\pi)-2\sum_{i}c_{i}\bigl(m(\mu)_{i}-m(\nu)_{i}\bigr)+\lVert c\rVert^{2}=I(\pi)-2\lVert c\rVert^{2}+\lVert c\rVert^{2}=I(\pi)-\lVert c\rVert^{2}.

By The Quadratic Wasserstein Distance on Euclidean Space §distance, W2(μˉ,νˉ)2I(πˉ)W_{2}(\bar{\mu},\bar{\nu})^{2}\le I(\bar{\pi}), so W2(μˉ,νˉ)2+c2I(π)W_{2}(\bar{\mu},\bar{\nu})^{2}+\lVert c\rVert^{2}\le I(\pi) (claim 3 of Elementary Arithmetic in an Ordered Field). As πΠ(μ,ν)\pi\in\Pi(\mu,\nu) was arbitrary, the left side is a lower bound of {I(π):πΠ(μ,ν)}\{I(\pi):\pi\in\Pi(\mu,\nu)\}, hence at most W2(μ,ν)2W_{2}(\mu,\nu)^{2} (Lower Bound and Greatest Lower Bound); and since 0c20\le\lVert c\rVert^{2} (Nonnegativity of Squares in an Ordered Field), W2(μˉ,νˉ)2W2(μ,ν)2W_{2}(\bar{\mu},\bar{\nu})^{2}\le W_{2}(\mu,\nu)^{2}, so W2(μˉ,νˉ)W2(μ,ν)W_{2}(\bar{\mu},\bar{\nu})\le W_{2}(\mu,\nu) by claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field.

The expectation map. For claims 3 and 4 define m:L2(Ω;Rd)Rd\mathsf{m}:L^{2}(\Omega;\mathbb{R}^{d})\to\mathbb{R}^{d} by m(X)=m(L(X))\mathsf{m}(X)=m(\mathcal{L}(X)), defined since L(X)P2(Rd)\mathcal{L}(X)\in\mathcal{P}_{2}(\mathbb{R}^{d}) by The Space of Square-Integrable Random Vectors is a Real Hilbert Space; Its Laws Have Finite Second Moment; Constants and Translations §law. By claim 1, m(X)i=E[Xi]\mathsf{m}(X)_{i}=\mathbb{E}[X_{i}] for every representative of XX. The coordinates of X+HX+H and of tXtX (tRt\in\mathbb{R}) are Xi+HiX_{i}+H_{i} and tXitX_{i} (operations being pointwise, The Space of Square-Integrable Random Vectors §classes), so claim 2 of Linearity and Monotonicity of the Lebesgue Integral gives

m(X+H)=m(X)+m(H),m(tX)=tm(X),m(ca)=a,\mathsf{m}(X+H)=\mathsf{m}(X)+\mathsf{m}(H),\qquad\mathsf{m}(tX)=t\,\mathsf{m}(X),\qquad\mathsf{m}(c_{a})=a,

the last because cac_{a} has coordinates constant equal to aia_{i}; hence m(X+ca)=m(X)+a\mathsf{m}(X+c_{a})=\mathsf{m}(X)+a. By claim 1 and The Space of Square-Integrable Random Vectors is a Real Hilbert Space; Its Laws Have Finite Second Moment; Constants and Translations §law,

m(H)2M2(L(H))=HL22,som(H)HL2(L)\lVert\mathsf{m}(H)\rVert^{2}\le M_{2}(\mathcal{L}(H))=\lVert H\rVert_{L^{2}}^{2},\qquad\text{so}\qquad\lVert\mathsf{m}(H)\rVert\le\lVert H\rVert_{L^{2}}\tag{L}

(claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field), and in particular m(X)m(Y)=m(XY)XYL2\lVert\mathsf{m}(X)-\mathsf{m}(Y)\rVert=\lVert\mathsf{m}(X-Y)\rVert\le\lVert X-Y\rVert_{L^{2}}. Moreover, for pRdp\in\mathbb{R}^{d} and HL2(Ω;Rd)H\in L^{2}(\Omega;\mathbb{R}^{d}), by The Space of Square-Integrable Random Vectors §inner-product, cpH=ipiHic_{p}\cdot H=\sum_{i}p_{i}H_{i} pointwise and finite additivity,

cp,HL2=E[ipiHi]=ipiE[Hi]=pm(H).(C)\langle c_{p},H\rangle_{L^{2}}=\mathbb{E}\Bigl[\sum_{i}p_{i}H_{i}\Bigr]=\sum_{i}p_{i}\,\mathbb{E}[H_{i}]=p\cdot\mathsf{m}(H).\tag{C}

Finally, by Translations on a Space of Square-Integrable Random Vectors: Constant Classes, Law Invariance, the Translation Derivative and the Translation Hessian §constants, ca+cb=ca+bc_{a}+c_{b}=c_{a+b}, tca=ctatc_{a}=c_{ta} and caL2=a\lVert c_{a}\rVert_{L^{2}}=\lVert a\rVert, so cacbL2=cabL2=ab\lVert c_{a}-c_{b}\rVert_{L^{2}}=\lVert c_{a-b}\rVert_{L^{2}}=\lVert a-b\rVert.

Translates of CkC^{k} functions. For g:RdRg:\mathbb{R}^{d}\to\mathbb{R} and bRdb\in\mathbb{R}^{d} write gb=gτbg^{b}=g\circ\tau_{b}, the function xg(x+b)x\mapsto g(x+b). For every xx and ii the difference quotients of Partial Derivative on a Euclidean Open Set for gbg^{b} at xx are those of gg at x+bx+b, because gb(x+hei)=g(x+b+hei)g^{b}(x+he_{i})=g(x+b+he_{i}); hence igb\partial_{i}g^{b} exists at xx if and only if ig\partial_{i}g exists at x+bx+b, and then i(gb)=(ig)b\partial_{i}(g^{b})=(\partial_{i}g)^{b} (the value of a partial derivative being unique: if L,LL,L' both satisfy the definition then LL<2ε|L-L'|<2\varepsilon for every ε>0\varepsilon>0 by claim 5 of Properties of the Absolute Value in an Ordered Field, so L=LL=L' by Comparison of Real Numbers with Arbitrary Positive Slack). Also gbg^{b} is continuous whenever gg is, by claim 3 of Semicontinuity and Continuity Under Composition with a Continuous Map, τb\tau_{b} being continuous since τb(x)τb(x)=xx\lVert\tau_{b}(x)-\tau_{b}(x')\rVert=\lVert x-x'\rVert. By induction on kk (Principle of Induction for the Natural Numbers, on the set of kNk\in\mathbb{N} such that for every gg of class CkC^{k} on Rd\mathbb{R}^{d} and every bb the function gbg^{b} is of class CkC^{k} on Rd\mathbb{R}^{d} with i(gb)=(ig)b\partial_{i}(g^{b})=(\partial_{i}g)^{b} for all ii), using clauses 1 and 2 of C^k Maps on a Euclidean Open Set: if gg is of class C2C^{2} then so is gbg^{b}, with ji(gb)=(jig)b\partial_{j}\partial_{i}(g^{b})=(\partial_{j}\partial_{i}g)^{b}, so that D2(gb)(x)=D2g(x+b)D^{2}(g^{b})(x)=D^{2}g(x+b) (Hessian Matrix of a C^2 Function, entrywise).

Proof of claim 3. Let ϕ\phi be of class C2C^{2} on Rd\mathbb{R}^{d} and φ(μ)=ϕ(m(μ))\varphi(\mu)=\phi(m(\mu)), with lift Φ=φΛ\Phi=\varphi\circ\Lambda (The Lift of a Function on the Wasserstein Space to the Space of Square-Integrable Random Vectors §lift), so Φ(X)=ϕ(m(X))\Phi(X)=\phi(\mathsf{m}(X)). We verify the four properties of Test Functions on the Wasserstein Space: the Intrinsic Gradient and the Translation Hessian §test.

(a). Let XL2(Ω;Rd)X\in L^{2}(\Omega;\mathbb{R}^{d}), a=m(X)a=\mathsf{m}(X) and p=Dϕ(a)p=D\phi(a), the point with coordinates iϕ(a)\partial_{i}\phi(a) (Differential Calculus and Convexity on Euclidean Open Sets: Standing Notation §derivatives). Since ϕ\phi is of class C1C^{1} (clause 2 of C^k Maps on a Euclidean Open Set), it is differentiable at aa with derivative matrix the 1×d1\times d matrix of the partials iϕ(a)\partial_{i}\phi(a) by A Real-Valued C^1 Function is Differentiable at Every Point; so by Differentiability at a Point for Maps Between Euclidean Spaces, for every ε>0\varepsilon>0 there is δ>0\delta>0 such that ϕ(a+h)ϕ(a)phεh|\phi(a+h)-\phi(a)-p\cdot h|\le\varepsilon\lVert h\rVert whenever 0<h<δ0<\lVert h\rVert<\delta, and trivially also for h=0Rdh=0_{\mathbb{R}^{d}}. Let HL2(Ω;Rd)H\in L^{2}(\Omega;\mathbb{R}^{d}) with HL2<δ\lVert H\rVert_{L^{2}}<\delta and put h=m(H)h=\mathsf{m}(H), so h<δ\lVert h\rVert<\delta by (L). By the expectation-map paragraph and (C),

Φ(X+H)Φ(X)cp,HL2=ϕ(a+h)ϕ(a)phεhεHL2.\bigl|\Phi(X+H)-\Phi(X)-\langle c_{p},H\rangle_{L^{2}}\bigr|=\bigl|\phi(a+h)-\phi(a)-p\cdot h\bigr|\le\varepsilon\lVert h\rVert\le\varepsilon\lVert H\rVert_{L^{2}}.

Hence Φ\Phi is differentiable at XX with gradient DΦ(X)=cDϕ(m(X))D\Phi(X)=c_{D\phi(\mathsf{m}(X))} (Fréchet Differentiability and the Gradient on an Open Subset of a Real Inner Product Space §differentiable, Fréchet Differentiability and the Gradient on an Open Subset of a Real Inner Product Space §gradient). The gradient map is continuous: given XX and ε>0\varepsilon>0, the dd continuous functions iϕ\partial_{i}\phi give δ>0\delta>0 with iϕ(a)iϕ(a)ε(2d)1|\partial_{i}\phi(a')-\partial_{i}\phi(a)|\le\varepsilon(2\sqrt{d})^{-1} for all ii whenever aa<δ\lVert a'-a\rVert<\delta (the least of dd radii, by claim 9 of Elementary Order Arithmetic in an Ordered Field applied repeatedly; continuity at a point as in Continuous Map Between Metric Spaces); for YXL2<δ\lVert Y-X\rVert_{L^{2}}<\delta one has m(Y)m(X)<δ\lVert\mathsf{m}(Y)-\mathsf{m}(X)\rVert<\delta by (L), hence DΦ(Y)DΦ(X)L22=Dϕ(m(Y))Dϕ(m(X))2dε2(4d)1=ε2/4\lVert D\Phi(Y)-D\Phi(X)\rVert_{L^{2}}^{2}=\lVert D\phi(\mathsf{m}(Y))-D\phi(\mathsf{m}(X))\rVert^{2}\le d\cdot\varepsilon^{2}(4d)^{-1}=\varepsilon^{2}/4 by claim 1 of Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n, claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field and claim 1 of Comparison and Absolute Value Bounds for Finite Sums of Real Numbers, so DΦ(Y)DΦ(X)L2ε/2<ε\lVert D\Phi(Y)-D\Phi(X)\rVert_{L^{2}}\le\varepsilon/2<\varepsilon (claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field again, and claim 8 of Elementary Order Arithmetic in an Ordered Field). Thus ΦC1(L2(Ω;Rd))\Phi\in C^{1}(L^{2}(\Omega;\mathbb{R}^{d})) (The Classes C1C^1 and C2C^2 on an Open Subset of a Real Inner Product Space §c1).

(b). Let μP2(Rd)\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}) and let η\eta be the class of the constant map with value p=Dϕ(m(μ))p=D\phi(m(\mu)), which lies in TμT_{\mu} by Gradients of Functions with Bounded Derivatives Belong to the Tangent Space; Constants Are Tangent; the Score Identity; the Score Has Mean Zero; Translation of Tangent Fields §constants. For XX with L(X)=μ\mathcal{L}(X)=\mu, m(X)=m(μ)\mathsf{m}(X)=m(\mu), and ηX\eta\circ X (Composition of a Square-Integrable Vector Field with a Random Vector, and the Lifted Score: Isometry, Norm, Weak Identity and Second-Moment Identity §composition) is the class of the constant map with value pp, i.e. cp=DΦ(X)c_{p}=D\Phi(X).

(c). Let XL2(Ω;Rd)X\in L^{2}(\Omega;\mathbb{R}^{d}) and ϕX(a)=Φ(X+ca)=ϕ(m(X)+a)=ϕm(X)(a)\phi_{X}(a)=\Phi(X+c_{a})=\phi(\mathsf{m}(X)+a)=\phi^{\mathsf{m}(X)}(a). By the translates paragraph, ϕX\phi_{X} is of class C2C^{2} on Rd\mathbb{R}^{d} with D2ϕX(0Rd)=D2ϕ(m(X))D^{2}\phi_{X}(0_{\mathbb{R}^{d}})=D^{2}\phi(\mathsf{m}(X)); so Φ\Phi is twice continuously differentiable along translations at every XX (The Translation Laplacian of a Function on the Space of Square-Integrable Random Vectors §translations, The Translation Laplacian of a Function on the Space of Square-Integrable Random Vectors §on-space).

(d). The map m:P2(Rd)Rdm:\mathcal{P}_{2}(\mathbb{R}^{d})\to\mathbb{R}^{d} is Lipschitz with constant 11 by claim 1 (Lipschitz Map Between Metric Spaces, with W2W_{2} and the Euclidean distance dE(x,y)=xyd_{E}(x,y)=\lVert x-y\rVert, claim 2 of Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n), hence continuous by A Lipschitz Map is Uniformly Continuous; ϕ\phi is continuous by claim 3 of Euclidean Space is Open in Itself, and CkC^k Maps are Continuous; so φ=ϕm\varphi=\phi\circ m is continuous by claim 3 of Semicontinuity and Continuity Under Composition with a Continuous Map.

Hence φ\varphi is a test function; its intrinsic gradient at μ\mu is the η\eta of (b) (Test Functions on the Wasserstein Space: the Intrinsic Gradient and the Translation Hessian §gradient), and its translation Hessian at μ\mu is D2ϕX(0Rd)=D2ϕ(m(μ))D^{2}\phi_{X}(0_{\mathbb{R}^{d}})=D^{2}\phi(m(\mu)) for any XX with L(X)=μ\mathcal{L}(X)=\mu (Test Functions on the Wasserstein Space: the Intrinsic Gradient and the Translation Hessian §hessian).

Proof of claim 4. Let φ\varphi be a test function with lift Φ\Phi, and define the centring operator C:L2(Ω;Rd)L2(Ω;Rd)\mathsf{C}:L^{2}(\Omega;\mathbb{R}^{d})\to L^{2}(\Omega;\mathbb{R}^{d}) by CX=Xcm(X)\mathsf{C}X=X-c_{\mathsf{m}(X)}. By the lift paragraph of claim 2, L(CX)=L(X)\mathcal{L}(\mathsf{C}X)=\overline{\mathcal{L}(X)}, so the lift Φ\Phi^{\circ} of φ\varphi^{\circ} satisfies Φ(X)=φ(L(X))=Φ(CX)\Phi^{\circ}(X)=\varphi(\overline{\mathcal{L}(X)})=\Phi(\mathsf{C}X). By the expectation-map paragraph, C(X+H)=CX+CH\mathsf{C}(X+H)=\mathsf{C}X+\mathsf{C}H and, using claim 2 and The Space of Square-Integrable Random Vectors is a Real Hilbert Space; Its Laws Have Finite Second Moment; Constants and Translations §law, CHL22=M2(L(H))=M2(L(H))m(H)2HL22\lVert\mathsf{C}H\rVert_{L^{2}}^{2}=M_{2}(\overline{\mathcal{L}(H)})=M_{2}(\mathcal{L}(H))-\lVert\mathsf{m}(H)\rVert^{2}\le\lVert H\rVert_{L^{2}}^{2}, so

CHL2HL2andCXCYL2=C(XY)L2XYL2.(N)\lVert\mathsf{C}H\rVert_{L^{2}}\le\lVert H\rVert_{L^{2}}\qquad\text{and}\qquad\lVert\mathsf{C}X-\mathsf{C}Y\rVert_{L^{2}}=\lVert\mathsf{C}(X-Y)\rVert_{L^{2}}\le\lVert X-Y\rVert_{L^{2}}.\tag{N}

Moreover, by Elementary Identities in a Real Inner Product Space §bilinear and (C), CH,GL2=H,GL2m(H)m(G)\langle\mathsf{C}H,G\rangle_{L^{2}}=\langle H,G\rangle_{L^{2}}-\mathsf{m}(H)\cdot\mathsf{m}(G), which is symmetric in HH and GG; hence

CH,GL2=H,CGL2(H,GL2(Ω;Rd)).(S)\langle\mathsf{C}H,G\rangle_{L^{2}}=\langle H,\mathsf{C}G\rangle_{L^{2}}\qquad(H,G\in L^{2}(\Omega;\mathbb{R}^{d})).\tag{S}

(a). Let XL2(Ω;Rd)X\in L^{2}(\Omega;\mathbb{R}^{d}) and Y=CXY=\mathsf{C}X. Given ε>0\varepsilon>0, let δ>0\delta>0 be provided by the differentiability of Φ\Phi at YY (Fréchet Differentiability and the Gradient on an Open Subset of a Real Inner Product Space §differentiable, property (a) of the test function φ\varphi). For HL2<δ\lVert H\rVert_{L^{2}}<\delta, (N) gives CHL2<δ\lVert\mathsf{C}H\rVert_{L^{2}}<\delta, so by (S)

Φ(X+H)Φ(X)CDΦ(Y),HL2=Φ(Y+CH)Φ(Y)DΦ(Y),CHL2εCHL2εHL2.\bigl|\Phi^{\circ}(X+H)-\Phi^{\circ}(X)-\langle\mathsf{C}D\Phi(Y),H\rangle_{L^{2}}\bigr|=\bigl|\Phi(Y+\mathsf{C}H)-\Phi(Y)-\langle D\Phi(Y),\mathsf{C}H\rangle_{L^{2}}\bigr|\le\varepsilon\lVert\mathsf{C}H\rVert_{L^{2}}\le\varepsilon\lVert H\rVert_{L^{2}}.

Thus Φ\Phi^{\circ} is differentiable at XX with DΦ(X)=C(DΦ(CX))D\Phi^{\circ}(X)=\mathsf{C}\bigl(D\Phi(\mathsf{C}X)\bigr). The gradient map XC(DΦ(CX))X\mapsto\mathsf{C}(D\Phi(\mathsf{C}X)) is continuous, as the composition of C\mathsf{C} (Lipschitz by (N), hence continuous by A Lipschitz Map is Uniformly Continuous), the continuous map DΦD\Phi, and C\mathsf{C} again (claim 3 of Semicontinuity and Continuity Under Composition with a Continuous Map). So ΦC1(L2(Ω;Rd))\Phi^{\circ}\in C^{1}(L^{2}(\Omega;\mathbb{R}^{d})).

(b) and the gradient formula. Let μP2(Rd)\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}), m=m(μ)m=m(\mu), and let η\eta be a representative of φ(μˉ)Tμˉ\nabla\varphi(\bar{\mu})\in T_{\bar{\mu}}. By Gradients of Functions with Bounded Derivatives Belong to the Tangent Space; Constants Are Tangent; the Score Identity; the Score Has Mean Zero; Translation of Tangent Fields §translation applied to μˉ\bar{\mu}, a=ma=m and η\eta, using (τm)#μˉ=μ(\tau_{m})_{\#}\bar{\mu}=\mu (claim 2): the map η:xη(xm)\eta^{\ast}:x\mapsto\eta(x-m) is Borel, its class in L2(μ;Rd)L^{2}(\mu;\mathbb{R}^{d}) does not depend on the representative, and it belongs to TμT_{\mu}. Since η2dμ<\int\lVert\eta^{\ast}\rVert^{2}d\mu<\infty and ηi12(η2+1)|\eta^{\ast}_{i}|\le\tfrac12(\lVert\eta^{\ast}\rVert^{2}+1) by (B), the coordinates ηi\eta^{\ast}_{i} are integrable with respect to μ\mu (as in the first paragraph of the proof of claim 1); since ηi=ηiτm\eta^{\ast}_{i}=\eta_{i}\circ\tau_{-m}, claim 2 of Image Measures, Measures with Densities, and Change of Variables for the image measure μˉ=(τm)#μ\bar{\mu}=(\tau_{-m})_{\#}\mu shows that the coordinate ηi\eta_{i} is integrable with respect to μˉ\bar{\mu} with ηidμˉ=ηidμ\int\eta_{i}\,d\bar{\mu}=\int\eta^{\ast}_{i}\,d\mu. Let b=ηdμˉb=\int\eta\,d\bar{\mu} be the point with these coordinates; if η\eta' is another representative then μˉ({η=η})=1\bar{\mu}(\{\eta=\eta'\})=1, so ηidμˉ=ηidμˉ\int\eta'_{i}\,d\bar{\mu}=\int\eta_{i}\,d\bar{\mu} by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §comparison, and bb does not depend on the representative. Now let XX satisfy L(X)=μ\mathcal{L}(X)=\mu and Y=CX=XcmY=\mathsf{C}X=X-c_{m}, so L(Y)=μˉ\mathcal{L}(Y)=\bar{\mu} and DΦ(Y)=φ(μˉ)YD\Phi(Y)=\nabla\varphi(\bar{\mu})\circ Y by property (b) of φ\varphi. For a representative of XX, the map ωη(Y(ω))=η(X(ω)m)=η(X(ω))\omega\mapsto\eta(Y(\omega))=\eta(X(\omega)-m)=\eta^{\ast}(X(\omega)) shows φ(μˉ)Y=ηX\nabla\varphi(\bar{\mu})\circ Y=\eta^{\ast}\circ X (Composition of a Square-Integrable Vector Field with a Random Vector, and the Lifted Score: Isometry, Norm, Weak Identity and Second-Moment Identity §composition). Its coordinates have expectations E[ηiX]=ηidμ=bi\mathbb{E}[\eta^{\ast}_{i}\circ X]=\int\eta^{\ast}_{i}\,d\mu=b_{i} (Basic Properties of Random Vectors: Coordinates, Borel Images and Arithmetic, Change of Variables, Almost Sure Equality and Pairs §expectation), so m(ηX)=b\mathsf{m}(\eta^{\ast}\circ X)=b and

DΦ(X)=C(ηX)=ηXcb=(ηb)X,D\Phi^{\circ}(X)=\mathsf{C}(\eta^{\ast}\circ X)=\eta^{\ast}\circ X-c_{b}=(\eta^{\ast}-b)\circ X,

by the linearity of composition in Composition of a Square-Integrable Vector Field with a Random Vector, and the Lifted Score: Isometry, Norm, Weak Identity and Second-Moment Identity §composition, the constant map with value bb composed with XX being cbc_{b}. The class of ηb\eta^{\ast}-b, i.e. of xη(xm)ηdμˉx\mapsto\eta(x-m)-\int\eta\,d\bar{\mu}, lies in TμT_{\mu}: ηTμ\eta^{\ast}\in T_{\mu}, bTμb\in T_{\mu} by Gradients of Functions with Bounded Derivatives Belong to the Tangent Space; Constants Are Tangent; the Score Identity; the Score Has Mean Zero; Translation of Tangent Fields §constants, and TμT_{\mu} is a linear subspace by Basic Properties of the Tangent Space: Closed Subspace, the Identity Map Belongs to It, Second-Moment Limits, and Representation of Bounded Functionals on Gradients §closed. This gives property (b), and by Test Functions on the Wasserstein Space: the Intrinsic Gradient and the Translation Hessian §gradient the intrinsic gradient φ(μ)\nabla\varphi^{\circ}(\mu) is this class.

(c). For XL2(Ω;Rd)X\in L^{2}(\Omega;\mathbb{R}^{d}) and aRda\in\mathbb{R}^{d}, C(X+ca)=X+cacm(X)+a=X+cam(X)a=Xcm(X)=CX\mathsf{C}(X+c_{a})=X+c_{a}-c_{\mathsf{m}(X)+a}=X+c_{a-\mathsf{m}(X)-a}=X-c_{\mathsf{m}(X)}=\mathsf{C}X by the expectation-map paragraph and Translations on a Space of Square-Integrable Random Vectors: Constant Classes, Law Invariance, the Translation Derivative and the Translation Hessian §constants. Hence ϕX(a)=Φ(X+ca)=Φ(CX)\phi^{\circ}_{X}(a)=\Phi^{\circ}(X+c_{a})=\Phi(\mathsf{C}X) is constant in aa, a smooth function (claim 2 of Constants, Coordinate Functions, Sums and Products of CkC^k Functions on a Euclidean Open Set) whose partial derivatives vanish identically (its difference quotients are 00, Partial Derivative on a Euclidean Open Set), as do their partial derivatives; so ϕX\phi^{\circ}_{X} is of class C2C^{2} with D2ϕX(0Rd)=0dD^{2}\phi^{\circ}_{X}(0_{\mathbb{R}^{d}})=0_{d}. Thus Φ\Phi^{\circ} is twice continuously differentiable along translations at every point, and the translation Hessian of φ\varphi^{\circ} at every μ\mu is 0d0_{d} (Test Functions on the Wasserstein Space: the Intrinsic Gradient and the Translation Hessian §hessian).

(d) and translation invariance. For μ\mu and aa, φ((τa)#μ)=φ((τa)#μ)=φ(μˉ)=φ(μ)\varphi^{\circ}((\tau_{a})_{\#}\mu)=\varphi(\overline{(\tau_{a})_{\#}\mu})=\varphi(\bar{\mu})=\varphi^{\circ}(\mu) by claim 2. For continuity at μ\mu: given ε>0\varepsilon>0, property (d) of φ\varphi gives δ>0\delta>0 with φ(σ)φ(μˉ)<ε|\varphi(\sigma)-\varphi(\bar{\mu})|<\varepsilon whenever W2(σ,μˉ)<δW_{2}(\sigma,\bar{\mu})<\delta (Continuous Map Between Metric Spaces); if W2(ν,μ)<δW_{2}(\nu,\mu)<\delta then W2(νˉ,μˉ)W2(ν,μ)<δW_{2}(\bar{\nu},\bar{\mu})\le W_{2}(\nu,\mu)<\delta by claim 2 (and the symmetry of W2W_{2}, The Quadratic Wasserstein Distance is a Metric on the Wasserstein Space §symmetry), so φ(ν)φ(μ)<ε|\varphi^{\circ}(\nu)-\varphi^{\circ}(\mu)|<\varepsilon. Hence φ\varphi^{\circ} is a test function with the stated gradient and Hessian.

Please log in to copy this version.

Citations

Loading…

Dependency Graph

0 prerequisites

Comments

Loading…