Each result cited is universally quantified over the data in its own statement. Throughout, π i : R d → R \pi_{i}:\mathbb{R}^{d}\to\mathbb{R} π i : R d → R , π i ( x ) = x i \pi_{i}(x)=x_{i} π i ( x ) = x i , is the i i i th coordinate function, smooth by claim 2 of Constants, Coordinate Functions, Sums and Products of C k C^k C k Functions on a Euclidean Open Set , hence continuous (claim 3 of Euclidean Space is Open in Itself, and C k C^k C k Maps are Continuous ) and Borel (claim 3 of The Borel Sigma-Algebra of a Euclidean Space as a Product, and Measurability of Projections, Sequentially Continuous Maps, and Open and Closed Sets ); e i e_{i} e i is the i i i th standard basis vector , so x ⋅ e i = x i x\cdot e_{i}=x_{i} x ⋅ e i = x i (as recorded there) and ∥ e i ∥ = 1 \lVert e_{i}\rVert=1 ∥ e i ∥ = 1 (claim 1 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n , the coordinates of e i e_{i} e i being 0 0 0 except for one 1 1 1 ); sums over i ∈ [ d ] i\in[d] i ∈ [ d ] are the finite sums of Properties of Finite Sums , and "integrable" is integrable with respect to the measure named. The integral of a finite sum of integrable functions is the sum of the integrals, by claim 2 of Linearity and Monotonicity of the Lebesgue Integral applied along the recursion of claim 1 of Properties of Finite Sums ; we refer to this as finite additivity . For a probability measure the integral of a constant is that constant (Simple Function and Its Integral ), and ∥ x ∥ 2 = ∑ i x i 2 \lVert x\rVert^{2}=\sum_{i}x_{i}^{2} ∥ x ∥ 2 = ∑ i x i 2 for x ∈ R d x\in\mathbb{R}^{d} x ∈ R d (claim 1 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n ). Two facts from Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions are used: ∣ x ⋅ y ∣ ≤ 1 2 ( ∥ x ∥ 2 + ∥ y ∥ 2 ) |x\cdot y|\le\tfrac12(\lVert x\rVert^{2}+\lVert y\rVert^{2}) ∣ x ⋅ y ∣ ≤ 2 1 (∥ x ∥ 2 + ∥ y ∥ 2 ) , and the Borel measurability of x ↦ ∥ x ∥ 2 x\mapsto\lVert x\rVert^{2} x ↦ ∥ x ∥ 2 ; with y = e i y=e_{i} y = e i the first gives
∣ x i ∣ ≤ 1 2 ( ∥ x ∥ 2 + 1 ) ( x ∈ R d ) . (B) |x_{i}|\le\tfrac12\bigl(\lVert x\rVert^{2}+1\bigr)\qquad(x\in\mathbb{R}^{d}).\tag{B} ∣ x i ∣ ≤ 2 1 ( ∥ x ∥ 2 + 1 ) ( x ∈ R d ) . ( B )
Proof of claim 1.
Integrability and the mean. Let μ ∈ P 2 ( R d ) \mu\in\mathcal{P}_{2}(\mathbb{R}^{d}) μ ∈ P 2 ( R d ) . By (B), ∣ π i ∣ ≤ 1 2 ( ∥ ⋅ ∥ 2 + 1 ) |\pi_{i}|\le\tfrac12(\lVert\cdot\rVert^{2}+1) ∣ π i ∣ ≤ 2 1 (∥ ⋅ ∥ 2 + 1 ) , and the right side is Borel with ∫ 1 2 ( ∥ x ∥ 2 + 1 ) μ ( d x ) = 1 2 ( M 2 ( μ ) + 1 ) < ∞ \int\tfrac12(\lVert x\rVert^{2}+1)\mu(dx)=\tfrac12(M_{2}(\mu)+1)<\infty ∫ 2 1 (∥ x ∥ 2 + 1 ) μ ( d x ) = 2 1 ( M 2 ( μ ) + 1 ) < ∞ by claim 1 of Linearity and Monotonicity of the Lebesgue Integral and The Second Moment of a Probability Measure on Euclidean Space and the Probability Measures with Finite Second Moment §moment ; hence ∫ ∣ π i ∣ d μ < ∞ \int|\pi_{i}|\,d\mu<\infty ∫ ∣ π i ∣ d μ < ∞ by the monotonicity in that claim, and π i \pi_{i} π i is integrable (Integrable Function and the Lebesgue Integral ). So m ( μ ) m(\mu) m ( μ ) is defined.
Lift. Let X ∈ L 2 ( Ω ; R d ) X\in L^{2}(\Omega;\mathbb{R}^{d}) X ∈ L 2 ( Ω ; R d ) with L ( X ) = μ \mathcal{L}(X)=\mu L ( X ) = μ and fix a representative; its law is μ \mu μ by The Space of Square-Integrable Random Vectors §law . The coordinate X i X_{i} X i is π i ∘ X \pi_{i}\circ X π i ∘ X (Random Vector and Its Law §coordinates ), and Basic Properties of Random Vectors: Coordinates, Borel Images and Arithmetic, Change of Variables, Almost Sure Equality and Pairs §expectation gives that π i ∘ X \pi_{i}\circ X π i ∘ X is integrable with E [ X i ] = ∫ π i d μ = m ( μ ) i \mathbb{E}[X_{i}]=\int\pi_{i}\,d\mu=m(\mu)_{i} E [ X i ] = ∫ π i d μ = m ( μ ) i .
Bound. By Hoelder's Inequality, for Two and for Finitely Many Factors §holder with p = q = 2 p=q=2 p = q = 2 (conjugate exponents in the sense of Conjugate Exponents and Young's Inequality §conjugate , as 1 2 + 1 2 = 1 \tfrac12+\tfrac12=1 2 1 + 2 1 = 1 ), applied to π i \pi_{i} π i and to the constant function 1 1 1 on the probability space ( R d , B ( R d ) , μ ) (\mathbb{R}^{d},\mathcal{B}(\mathbb{R}^{d}),\mu) ( R d , B ( R d ) , μ ) (Square-Integrable Vector Fields Against a Probability Measure on Euclidean Space, and Test Functions: Standing Notation §measures ), ∫ ∣ π i ∣ d μ ≤ ( ∫ π i 2 d μ ) 1 / 2 ( ∫ 1 d μ ) 1 / 2 = ( ∫ π i 2 d μ ) 1 / 2 \int|\pi_{i}|\,d\mu\le\bigl(\int\pi_{i}^{2}\,d\mu\bigr)^{1/2}\bigl(\int1\,d\mu\bigr)^{1/2}=\bigl(\int\pi_{i}^{2}\,d\mu\bigr)^{1/2} ∫ ∣ π i ∣ d μ ≤ ( ∫ π i 2 d μ ) 1/2 ( ∫ 1 d μ ) 1/2 = ( ∫ π i 2 d μ ) 1/2 , where π i 2 ≤ ∥ ⋅ ∥ 2 \pi_{i}^{2}\le\lVert\cdot\rVert^{2} π i 2 ≤ ∥ ⋅ ∥ 2 is integrable. With ∣ m ( μ ) i ∣ = ∣ ∫ π i d μ ∣ ≤ ∫ ∣ π i ∣ d μ |m(\mu)_{i}|=|\int\pi_{i}\,d\mu|\le\int|\pi_{i}|\,d\mu ∣ m ( μ ) i ∣ = ∣ ∫ π i d μ ∣ ≤ ∫ ∣ π i ∣ d μ (claim 2 of Linearity and Monotonicity of the Lebesgue Integral ) and claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field , m ( μ ) i 2 ≤ ∫ π i 2 d μ m(\mu)_{i}^{2}\le\int\pi_{i}^{2}\,d\mu m ( μ ) i 2 ≤ ∫ π i 2 d μ ; summing over i i i (claim 1 of Comparison and Absolute Value Bounds for Finite Sums of Real Numbers ) and using finite additivity, ∥ m ( μ ) ∥ 2 ≤ ∫ ∑ i π i 2 d μ = ∫ ∥ x ∥ 2 μ ( d x ) = M 2 ( μ ) \lVert m(\mu)\rVert^{2}\le\int\sum_{i}\pi_{i}^{2}\,d\mu=\int\lVert x\rVert^{2}\mu(dx)=M_{2}(\mu) ∥ m ( μ ) ∥ 2 ≤ ∫ ∑ i π i 2 d μ = ∫ ∥ x ∥ 2 μ ( d x ) = M 2 ( μ ) .
Translation. Let a ∈ R d a\in\mathbb{R}^{d} a ∈ R d and μ a = ( τ a ) # μ \mu_{a}=(\tau_{a})_{\#}\mu μ a = ( τ a ) # μ , an element of P 2 ( R d ) \mathcal{P}_{2}(\mathbb{R}^{d}) P 2 ( R d ) by Gradients of Functions with Bounded Derivatives Belong to the Tangent Space; Constants Are Tangent; the Score Identity; the Score Has Mean Zero; Translation of Tangent Fields §translation . By claim 2 of Image Measures, Measures with Densities, and Change of Variables (the image measure of μ \mu μ under the Borel map τ a \tau_{a} τ a , Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pushforward ), ∫ π i d μ a = ∫ π i ∘ τ a d μ = ∫ ( x i + a i ) μ ( d x ) = m ( μ ) i + a i \int\pi_{i}\,d\mu_{a}=\int\pi_{i}\circ\tau_{a}\,d\mu=\int(x_{i}+a_{i})\,\mu(dx)=m(\mu)_{i}+a_{i} ∫ π i d μ a = ∫ π i ∘ τ a d μ = ∫ ( x i + a i ) μ ( d x ) = m ( μ ) i + a i , the last step by claim 2 of Linearity and Monotonicity of the Lebesgue Integral ; that is, m ( μ a ) = m ( μ ) + a m(\mu_{a})=m(\mu)+a m ( μ a ) = m ( μ ) + a (sums in R d \mathbb{R}^{d} R d being coordinatewise, Euclidean Points as Tuples of Real Numbers ).
Lipschitz bound. Let ν ∈ P 2 ( R d ) \nu\in\mathcal{P}_{2}(\mathbb{R}^{d}) ν ∈ P 2 ( R d ) and π ∈ Π ( μ , ν ) \pi\in\Pi(\mu,\nu) π ∈ Π ( μ , ν ) , so I ( π ) < ∞ I(\pi)<\infty I ( π ) < ∞ by Couplings on Euclidean Space: Product Coupling, Swap, Finiteness of the Cost, Push-Forward Couplings, Modifying One Marginal, Quantisation, Gluing over a Finitely Supported Measure, and the Lipschitz Bound §cost-finite . By Couplings of Two Probability Measures on Euclidean Space and Their Quadratic Cost §coupling , μ \mu μ and ν \nu ν are the image measures of π \pi π under p r 1 \mathrm{pr}_{1} pr 1 and p r 2 \mathrm{pr}_{2} pr 2 (Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections ), so claim 2 of Image Measures, Measures with Densities, and Change of Variables gives that π i ∘ p r 1 \pi_{i}\circ\mathrm{pr}_{1} π i ∘ pr 1 and π i ∘ p r 2 \pi_{i}\circ\mathrm{pr}_{2} π i ∘ pr 2 are integrable with respect to π \pi π with ∫ π i ∘ p r 1 d π = m ( μ ) i \int\pi_{i}\circ\mathrm{pr}_{1}\,d\pi=m(\mu)_{i} ∫ π i ∘ pr 1 d π = m ( μ ) i and ∫ π i ∘ p r 2 d π = m ( ν ) i \int\pi_{i}\circ\mathrm{pr}_{2}\,d\pi=m(\nu)_{i} ∫ π i ∘ pr 2 d π = m ( ν ) i . Hence, with D i = π i ∘ p r 1 − π i ∘ p r 2 D_{i}=\pi_{i}\circ\mathrm{pr}_{1}-\pi_{i}\circ\mathrm{pr}_{2} D i = π i ∘ pr 1 − π i ∘ pr 2 , the function z ↦ ( p r 1 ( z ) − p r 2 ( z ) ) i z\mapsto(\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z))_{i} z ↦ ( pr 1 ( z ) − pr 2 ( z ) ) i ,
m ( μ ) i − m ( ν ) i = ∫ D i d π (M) m(\mu)_{i}-m(\nu)_{i}=\int D_{i}\,d\pi\tag{M} m ( μ ) i − m ( ν ) i = ∫ D i d π ( M )
by claim 2 of Linearity and Monotonicity of the Lebesgue Integral . As in the bound paragraph (Hoelder with the constant 1 1 1 on the probability space ( R d + d , B ( R d + d ) , π ) (\mathbb{R}^{d+d},\mathcal{B}(\mathbb{R}^{d+d}),\pi) ( R d + d , B ( R d + d ) , π ) , and D i 2 ≤ ∥ p r 1 − p r 2 ∥ 2 D_{i}^{2}\le\lVert\mathrm{pr}_{1}-\mathrm{pr}_{2}\rVert^{2} D i 2 ≤ ∥ pr 1 − pr 2 ∥ 2 , which is Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions with π \pi π -integral I ( π ) < ∞ I(\pi)<\infty I ( π ) < ∞ , Couplings of Two Probability Measures on Euclidean Space and Their Quadratic Cost §cost ), ( m ( μ ) i − m ( ν ) i ) 2 ≤ ∫ D i 2 d π (m(\mu)_{i}-m(\nu)_{i})^{2}\le\int D_{i}^{2}\,d\pi ( m ( μ ) i − m ( ν ) i ) 2 ≤ ∫ D i 2 d π , and summing, ∥ m ( μ ) − m ( ν ) ∥ 2 ≤ ∫ ∥ p r 1 − p r 2 ∥ 2 d π = I ( π ) \lVert m(\mu)-m(\nu)\rVert^{2}\le\int\lVert\mathrm{pr}_{1}-\mathrm{pr}_{2}\rVert^{2}\,d\pi=I(\pi) ∥ m ( μ ) − m ( ν ) ∥ 2 ≤ ∫ ∥ pr 1 − pr 2 ∥ 2 d π = I ( π ) . Since π ∈ Π ( μ , ν ) \pi\in\Pi(\mu,\nu) π ∈ Π ( μ , ν ) was arbitrary, ∥ m ( μ ) − m ( ν ) ∥ 2 \lVert m(\mu)-m(\nu)\rVert^{2} ∥ m ( μ ) − m ( ν ) ∥ 2 is a lower bound of { I ( π ) : π ∈ Π ( μ , ν ) } \{I(\pi):\pi\in\Pi(\mu,\nu)\} { I ( π ) : π ∈ Π ( μ , ν )} , hence at most its greatest lower bound W 2 ( μ , ν ) 2 W_{2}(\mu,\nu)^{2} W 2 ( μ , ν ) 2 (The Quadratic Wasserstein Distance on Euclidean Space §distance , Lower Bound and Greatest Lower Bound ), and ∥ m ( μ ) − m ( ν ) ∥ ≤ W 2 ( μ , ν ) \lVert m(\mu)-m(\nu)\rVert\le W_{2}(\mu,\nu) ∥ m ( μ ) − m ( ν )∥ ≤ W 2 ( μ , ν ) by claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field .
Proof of claim 2. Write m = m ( μ ) m=m(\mu) m = m ( μ ) . By Gradients of Functions with Bounded Derivatives Belong to the Tangent Space; Constants Are Tangent; the Score Identity; the Score Has Mean Zero; Translation of Tangent Fields §translation with a = − m a=-m a = − m , μ ˉ = ( τ − m ) # μ ∈ P 2 ( R d ) \bar{\mu}=(\tau_{-m})_{\#}\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}) μ ˉ = ( τ − m ) # μ ∈ P 2 ( R d ) , and by the translation paragraph of claim 1, m ( μ ˉ ) = m + ( − m ) = 0 R d m(\bar{\mu})=m+(-m)=0_{\mathbb{R}^{d}} m ( μ ˉ ) = m + ( − m ) = 0 R d .
Second moment. By claim 2 of Image Measures, Measures with Densities, and Change of Variables for the nonnegative Borel function ∥ ⋅ ∥ 2 \lVert\cdot\rVert^{2} ∥ ⋅ ∥ 2 , M 2 ( μ ˉ ) = ∫ ∥ x − m ∥ 2 μ ( d x ) M_{2}(\bar{\mu})=\int\lVert x-m\rVert^{2}\,\mu(dx) M 2 ( μ ˉ ) = ∫ ∥ x − m ∥ 2 μ ( d x ) . For every x x x , ∥ x − m ∥ 2 = ∑ i ( x i − m i ) 2 = ∑ i ( x i 2 − 2 m i x i + m i 2 ) \lVert x-m\rVert^{2}=\sum_{i}(x_{i}-m_{i})^{2}=\sum_{i}\bigl(x_{i}^{2}-2m_{i}x_{i}+m_{i}^{2}\bigr) ∥ x − m ∥ 2 = ∑ i ( x i − m i ) 2 = ∑ i ( x i 2 − 2 m i x i + m i 2 ) (claim 5 of Zero Products and Elementary Identities in a Field ), and each summand is integrable; by finite additivity, claim 2 of Linearity and Monotonicity of the Lebesgue Integral and ∫ π i d μ = m i \int\pi_{i}\,d\mu=m_{i} ∫ π i d μ = m i ,
M 2 ( μ ˉ ) = ∑ i ( ∫ π i 2 d μ − 2 m i m i + m i 2 ) = M 2 ( μ ) − ∥ m ∥ 2 , M_{2}(\bar{\mu})=\sum_{i}\Bigl(\int\pi_{i}^{2}\,d\mu-2m_{i}m_{i}+m_{i}^{2}\Bigr)=M_{2}(\mu)-\lVert m\rVert^{2}, M 2 ( μ ˉ ) = i ∑ ( ∫ π i 2 d μ − 2 m i m i + m i 2 ) = M 2 ( μ ) − ∥ m ∥ 2 ,
using claims 2 and 3 of Properties of Finite Sums .
Composition of translations. For a , b ∈ R d a,b\in\mathbb{R}^{d} a , b ∈ R d and every Borel set B B B , claim 1 of Image Measures, Measures with Densities, and Change of Variables gives ( ( τ b ) # ( τ a ) # μ ) ( B ) = μ ( τ a − 1 ( τ b − 1 ( B ) ) ) = μ ( ( τ b ∘ τ a ) − 1 ( B ) ) = ( ( τ a + b ) # μ ) ( B ) \bigl((\tau_{b})_{\#}(\tau_{a})_{\#}\mu\bigr)(B)=\mu\bigl(\tau_{a}^{-1}(\tau_{b}^{-1}(B))\bigr)=\mu\bigl((\tau_{b}\circ\tau_{a})^{-1}(B)\bigr)=\bigl((\tau_{a+b})_{\#}\mu\bigr)(B) ( ( τ b ) # ( τ a ) # μ ) ( B ) = μ ( τ a − 1 ( τ b − 1 ( B )) ) = μ ( ( τ b ∘ τ a ) − 1 ( B ) ) = ( ( τ a + b ) # μ ) ( B ) , since τ b ( τ a ( x ) ) = x + a + b = τ a + b ( x ) \tau_{b}(\tau_{a}(x))=x+a+b=\tau_{a+b}(x) τ b ( τ a ( x )) = x + a + b = τ a + b ( x ) ; thus ( τ b ) # ( τ a ) # μ = ( τ a + b ) # μ (\tau_{b})_{\#}(\tau_{a})_{\#}\mu=(\tau_{a+b})_{\#}\mu ( τ b ) # ( τ a ) # μ = ( τ a + b ) # μ , and ( τ 0 R d ) # μ = μ (\tau_{0_{\mathbb{R}^{d}}})_{\#}\mu=\mu ( τ 0 R d ) # μ = μ because τ 0 R d \tau_{0_{\mathbb{R}^{d}}} τ 0 R d is the identity. Hence ( τ m ) # μ ˉ = ( τ − m + m ) # μ = μ (\tau_{m})_{\#}\bar{\mu}=(\tau_{-m+m})_{\#}\mu=\mu ( τ m ) # μ ˉ = ( τ − m + m ) # μ = μ ; and for a ∈ R d a\in\mathbb{R}^{d} a ∈ R d , m ( ( τ a ) # μ ) = m + a m((\tau_{a})_{\#}\mu)=m+a m (( τ a ) # μ ) = m + a by claim 1, so ( τ a ) # μ ‾ = ( τ − m − a ) # ( τ a ) # μ = ( τ − m ) # μ = μ ˉ \overline{(\tau_{a})_{\#}\mu}=(\tau_{-m-a})_{\#}(\tau_{a})_{\#}\mu=(\tau_{-m})_{\#}\mu=\bar{\mu} ( τ a ) # μ = ( τ − m − a ) # ( τ a ) # μ = ( τ − m ) # μ = μ ˉ .
The lift. Let X ∈ L 2 ( Ω ; R d ) X\in L^{2}(\Omega;\mathbb{R}^{d}) X ∈ L 2 ( Ω ; R d ) with L ( X ) = μ \mathcal{L}(X)=\mu L ( X ) = μ . By Translations on a Space of Square-Integrable Random Vectors: Constant Classes, Law Invariance, the Translation Derivative and the Translation Hessian §constants (read with its dimension parameter equal to d d d ), c − m = ( − 1 ) c m c_{-m}=(-1)c_{m} c − m = ( − 1 ) c m , which is − c m -c_{m} − c m by claim 5 of Elementary Identities in a Vector Space , so X − c m = X + c − m X-c_{m}=X+c_{-m} X − c m = X + c − m (claim 2 there), and by The Wasserstein Space and Its Lift to Square-Integrable Random Vectors: Standing Notation §constants , L ( X − c m ) = ( τ − m ) # L ( X ) = μ ˉ \mathcal{L}(X-c_{m})=(\tau_{-m})_{\#}\mathcal{L}(X)=\bar{\mu} L ( X − c m ) = ( τ − m ) # L ( X ) = μ ˉ .
The centred coupling. Let ν ∈ P 2 ( R d ) \nu\in\mathcal{P}_{2}(\mathbb{R}^{d}) ν ∈ P 2 ( R d ) , write c = m ( μ ) − m ( ν ) c=m(\mu)-m(\nu) c = m ( μ ) − m ( ν ) , let π ∈ Π ( μ , ν ) \pi\in\Pi(\mu,\nu) π ∈ Π ( μ , ν ) , and let S = τ − m ( μ ) S=\tau_{-m(\mu)} S = τ − m ( μ ) , T = τ − m ( ν ) T=\tau_{-m(\nu)} T = τ − m ( ν ) , Borel maps. By Couplings on Euclidean Space: Product Coupling, Swap, Finiteness of the Cost, Push-Forward Couplings, Modifying One Marginal, Quantisation, Gluing over a Finitely Supported Measure, and the Lipschitz Bound §modification , π ′ = ( p r 1 , T ∘ p r 2 ) # π ∈ Π ( μ , ν ˉ ) \pi'=(\mathrm{pr}_{1},T\circ\mathrm{pr}_{2})_{\#}\pi\in\Pi(\mu,\bar{\nu}) π ′ = ( pr 1 , T ∘ pr 2 ) # π ∈ Π ( μ , ν ˉ ) and π ˉ = ( S ∘ p r 1 , p r 2 ) # π ′ ∈ Π ( μ ˉ , ν ˉ ) \bar{\pi}=(S\circ\mathrm{pr}_{1},\mathrm{pr}_{2})_{\#}\pi'\in\Pi(\bar{\mu},\bar{\nu}) π ˉ = ( S ∘ pr 1 , pr 2 ) # π ′ ∈ Π ( μ ˉ , ν ˉ ) . Since p r 1 ∘ ( u , v ) = u \mathrm{pr}_{1}\circ(u,v)=u pr 1 ∘ ( u , v ) = u and p r 2 ∘ ( u , v ) = v \mathrm{pr}_{2}\circ(u,v)=v pr 2 ∘ ( u , v ) = v for a pairing (Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections ), one has ( S ∘ p r 1 , p r 2 ) ∘ ( p r 1 , T ∘ p r 2 ) = ( S ∘ p r 1 , T ∘ p r 2 ) (S\circ\mathrm{pr}_{1},\mathrm{pr}_{2})\circ(\mathrm{pr}_{1},T\circ\mathrm{pr}_{2})=(S\circ\mathrm{pr}_{1},T\circ\mathrm{pr}_{2}) ( S ∘ pr 1 , pr 2 ) ∘ ( pr 1 , T ∘ pr 2 ) = ( S ∘ pr 1 , T ∘ pr 2 ) , so π ˉ = ( S ∘ p r 1 , T ∘ p r 2 ) # π \bar{\pi}=(S\circ\mathrm{pr}_{1},T\circ\mathrm{pr}_{2})_{\#}\pi π ˉ = ( S ∘ pr 1 , T ∘ pr 2 ) # π by the composition rule for image measures (claim 1 of Image Measures, Measures with Densities, and Change of Variables , exactly as for translations above). By claim 2 of Image Measures, Measures with Densities, and Change of Variables and Couplings of Two Probability Measures on Euclidean Space and Their Quadratic Cost §cost ,
I ( π ˉ ) = ∫ ∥ S ( p r 1 ( z ) ) − T ( p r 2 ( z ) ) ∥ 2 π ( d z ) = ∫ ∥ p r 1 ( z ) − p r 2 ( z ) − c ∥ 2 π ( d z ) , I(\bar{\pi})=\int\bigl\lVert S(\mathrm{pr}_{1}(z))-T(\mathrm{pr}_{2}(z))\bigr\rVert^{2}\,\pi(dz)=\int\bigl\lVert\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z)-c\bigr\rVert^{2}\,\pi(dz), I ( π ˉ ) = ∫ S ( pr 1 ( z )) − T ( pr 2 ( z )) 2 π ( d z ) = ∫ pr 1 ( z ) − pr 2 ( z ) − c 2 π ( d z ) ,
because S ( x ) − T ( y ) = ( x − m ( μ ) ) − ( y − m ( ν ) ) = x − y − c S(x)-T(y)=(x-m(\mu))-(y-m(\nu))=x-y-c S ( x ) − T ( y ) = ( x − m ( μ )) − ( y − m ( ν )) = x − y − c . With w = p r 1 ( z ) − p r 2 ( z ) w=\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z) w = pr 1 ( z ) − pr 2 ( z ) , ∥ w − c ∥ 2 = ∑ i ( w i − c i ) 2 = ∥ w ∥ 2 − 2 ∑ i c i w i + ∥ c ∥ 2 \lVert w-c\rVert^{2}=\sum_{i}(w_{i}-c_{i})^{2}=\lVert w\rVert^{2}-2\sum_{i}c_{i}w_{i}+\lVert c\rVert^{2} ∥ w − c ∥ 2 = ∑ i ( w i − c i ) 2 = ∥ w ∥ 2 − 2 ∑ i c i w i + ∥ c ∥ 2 (claim 5 of Zero Products and Elementary Identities in a Field , claims 2 and 3 of Properties of Finite Sums ), where w i = D i ( z ) w_{i}=D_{i}(z) w i = D i ( z ) is integrable (Lipschitz paragraph of claim 1) and ∥ w ∥ 2 \lVert w\rVert^{2} ∥ w ∥ 2 has integral I ( π ) < ∞ I(\pi)<\infty I ( π ) < ∞ . Integrating with claim 2 of Linearity and Monotonicity of the Lebesgue Integral , finite additivity and (M),
I ( π ˉ ) = I ( π ) − 2 ∑ i c i ( m ( μ ) i − m ( ν ) i ) + ∥ c ∥ 2 = I ( π ) − 2 ∥ c ∥ 2 + ∥ c ∥ 2 = I ( π ) − ∥ c ∥ 2 . I(\bar{\pi})=I(\pi)-2\sum_{i}c_{i}\bigl(m(\mu)_{i}-m(\nu)_{i}\bigr)+\lVert c\rVert^{2}=I(\pi)-2\lVert c\rVert^{2}+\lVert c\rVert^{2}=I(\pi)-\lVert c\rVert^{2}. I ( π ˉ ) = I ( π ) − 2 i ∑ c i ( m ( μ ) i − m ( ν ) i ) + ∥ c ∥ 2 = I ( π ) − 2 ∥ c ∥ 2 + ∥ c ∥ 2 = I ( π ) − ∥ c ∥ 2 .
By The Quadratic Wasserstein Distance on Euclidean Space §distance , W 2 ( μ ˉ , ν ˉ ) 2 ≤ I ( π ˉ ) W_{2}(\bar{\mu},\bar{\nu})^{2}\le I(\bar{\pi}) W 2 ( μ ˉ , ν ˉ ) 2 ≤ I ( π ˉ ) , so W 2 ( μ ˉ , ν ˉ ) 2 + ∥ c ∥ 2 ≤ I ( π ) W_{2}(\bar{\mu},\bar{\nu})^{2}+\lVert c\rVert^{2}\le I(\pi) W 2 ( μ ˉ , ν ˉ ) 2 + ∥ c ∥ 2 ≤ I ( π ) (claim 3 of Elementary Arithmetic in an Ordered Field ). As π ∈ Π ( μ , ν ) \pi\in\Pi(\mu,\nu) π ∈ Π ( μ , ν ) was arbitrary, the left side is a lower bound of { I ( π ) : π ∈ Π ( μ , ν ) } \{I(\pi):\pi\in\Pi(\mu,\nu)\} { I ( π ) : π ∈ Π ( μ , ν )} , hence at most W 2 ( μ , ν ) 2 W_{2}(\mu,\nu)^{2} W 2 ( μ , ν ) 2 (Lower Bound and Greatest Lower Bound ); and since 0 ≤ ∥ c ∥ 2 0\le\lVert c\rVert^{2} 0 ≤ ∥ c ∥ 2 (Nonnegativity of Squares in an Ordered Field ), W 2 ( μ ˉ , ν ˉ ) 2 ≤ W 2 ( μ , ν ) 2 W_{2}(\bar{\mu},\bar{\nu})^{2}\le W_{2}(\mu,\nu)^{2} W 2 ( μ ˉ , ν ˉ ) 2 ≤ W 2 ( μ , ν ) 2 , so W 2 ( μ ˉ , ν ˉ ) ≤ W 2 ( μ , ν ) W_{2}(\bar{\mu},\bar{\nu})\le W_{2}(\mu,\nu) W 2 ( μ ˉ , ν ˉ ) ≤ W 2 ( μ , ν ) by claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field .
The expectation map. For claims 3 and 4 define m : L 2 ( Ω ; R d ) → R d \mathsf{m}:L^{2}(\Omega;\mathbb{R}^{d})\to\mathbb{R}^{d} m : L 2 ( Ω ; R d ) → R d by m ( X ) = m ( L ( X ) ) \mathsf{m}(X)=m(\mathcal{L}(X)) m ( X ) = m ( L ( X )) , defined since L ( X ) ∈ P 2 ( R d ) \mathcal{L}(X)\in\mathcal{P}_{2}(\mathbb{R}^{d}) L ( X ) ∈ P 2 ( R d ) by The Space of Square-Integrable Random Vectors is a Real Hilbert Space; Its Laws Have Finite Second Moment; Constants and Translations §law . By claim 1, m ( X ) i = E [ X i ] \mathsf{m}(X)_{i}=\mathbb{E}[X_{i}] m ( X ) i = E [ X i ] for every representative of X X X . The coordinates of X + H X+H X + H and of t X tX tX (t ∈ R t\in\mathbb{R} t ∈ R ) are X i + H i X_{i}+H_{i} X i + H i and t X i tX_{i} t X i (operations being pointwise, The Space of Square-Integrable Random Vectors §classes ), so claim 2 of Linearity and Monotonicity of the Lebesgue Integral gives
m ( X + H ) = m ( X ) + m ( H ) , m ( t X ) = t m ( X ) , m ( c a ) = a , \mathsf{m}(X+H)=\mathsf{m}(X)+\mathsf{m}(H),\qquad\mathsf{m}(tX)=t\,\mathsf{m}(X),\qquad\mathsf{m}(c_{a})=a, m ( X + H ) = m ( X ) + m ( H ) , m ( tX ) = t m ( X ) , m ( c a ) = a ,
the last because c a c_{a} c a has coordinates constant equal to a i a_{i} a i ; hence m ( X + c a ) = m ( X ) + a \mathsf{m}(X+c_{a})=\mathsf{m}(X)+a m ( X + c a ) = m ( X ) + a . By claim 1 and The Space of Square-Integrable Random Vectors is a Real Hilbert Space; Its Laws Have Finite Second Moment; Constants and Translations §law ,
∥ m ( H ) ∥ 2 ≤ M 2 ( L ( H ) ) = ∥ H ∥ L 2 2 , so ∥ m ( H ) ∥ ≤ ∥ H ∥ L 2 (L) \lVert\mathsf{m}(H)\rVert^{2}\le M_{2}(\mathcal{L}(H))=\lVert H\rVert_{L^{2}}^{2},\qquad\text{so}\qquad\lVert\mathsf{m}(H)\rVert\le\lVert H\rVert_{L^{2}}\tag{L} ∥ m ( H ) ∥ 2 ≤ M 2 ( L ( H )) = ∥ H ∥ L 2 2 , so ∥ m ( H )∥ ≤ ∥ H ∥ L 2 ( L )
(claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field ), and in particular ∥ m ( X ) − m ( Y ) ∥ = ∥ m ( X − Y ) ∥ ≤ ∥ X − Y ∥ L 2 \lVert\mathsf{m}(X)-\mathsf{m}(Y)\rVert=\lVert\mathsf{m}(X-Y)\rVert\le\lVert X-Y\rVert_{L^{2}} ∥ m ( X ) − m ( Y )∥ = ∥ m ( X − Y )∥ ≤ ∥ X − Y ∥ L 2 . Moreover, for p ∈ R d p\in\mathbb{R}^{d} p ∈ R d and H ∈ L 2 ( Ω ; R d ) H\in L^{2}(\Omega;\mathbb{R}^{d}) H ∈ L 2 ( Ω ; R d ) , by The Space of Square-Integrable Random Vectors §inner-product , c p ⋅ H = ∑ i p i H i c_{p}\cdot H=\sum_{i}p_{i}H_{i} c p ⋅ H = ∑ i p i H i pointwise and finite additivity,
⟨ c p , H ⟩ L 2 = E [ ∑ i p i H i ] = ∑ i p i E [ H i ] = p ⋅ m ( H ) . (C) \langle c_{p},H\rangle_{L^{2}}=\mathbb{E}\Bigl[\sum_{i}p_{i}H_{i}\Bigr]=\sum_{i}p_{i}\,\mathbb{E}[H_{i}]=p\cdot\mathsf{m}(H).\tag{C} ⟨ c p , H ⟩ L 2 = E [ i ∑ p i H i ] = i ∑ p i E [ H i ] = p ⋅ m ( H ) . ( C )
Finally, by Translations on a Space of Square-Integrable Random Vectors: Constant Classes, Law Invariance, the Translation Derivative and the Translation Hessian §constants , c a + c b = c a + b c_{a}+c_{b}=c_{a+b} c a + c b = c a + b , t c a = c t a tc_{a}=c_{ta} t c a = c t a and ∥ c a ∥ L 2 = ∥ a ∥ \lVert c_{a}\rVert_{L^{2}}=\lVert a\rVert ∥ c a ∥ L 2 = ∥ a ∥ , so ∥ c a − c b ∥ L 2 = ∥ c a − b ∥ L 2 = ∥ a − b ∥ \lVert c_{a}-c_{b}\rVert_{L^{2}}=\lVert c_{a-b}\rVert_{L^{2}}=\lVert a-b\rVert ∥ c a − c b ∥ L 2 = ∥ c a − b ∥ L 2 = ∥ a − b ∥ .
Translates of C k C^{k} C k functions. For g : R d → R g:\mathbb{R}^{d}\to\mathbb{R} g : R d → R and b ∈ R d b\in\mathbb{R}^{d} b ∈ R d write g b = g ∘ τ b g^{b}=g\circ\tau_{b} g b = g ∘ τ b , the function x ↦ g ( x + b ) x\mapsto g(x+b) x ↦ g ( x + b ) . For every x x x and i i i the difference quotients of Partial Derivative on a Euclidean Open Set for g b g^{b} g b at x x x are those of g g g at x + b x+b x + b , because g b ( x + h e i ) = g ( x + b + h e i ) g^{b}(x+he_{i})=g(x+b+he_{i}) g b ( x + h e i ) = g ( x + b + h e i ) ; hence ∂ i g b \partial_{i}g^{b} ∂ i g b exists at x x x if and only if ∂ i g \partial_{i}g ∂ i g exists at x + b x+b x + b , and then ∂ i ( g b ) = ( ∂ i g ) b \partial_{i}(g^{b})=(\partial_{i}g)^{b} ∂ i ( g b ) = ( ∂ i g ) b (the value of a partial derivative being unique: if L , L ′ L,L' L , L ′ both satisfy the definition then ∣ L − L ′ ∣ < 2 ε |L-L'|<2\varepsilon ∣ L − L ′ ∣ < 2 ε for every ε > 0 \varepsilon>0 ε > 0 by claim 5 of Properties of the Absolute Value in an Ordered Field , so L = L ′ L=L' L = L ′ by Comparison of Real Numbers with Arbitrary Positive Slack ). Also g b g^{b} g b is continuous whenever g g g is, by claim 3 of Semicontinuity and Continuity Under Composition with a Continuous Map , τ b \tau_{b} τ b being continuous since ∥ τ b ( x ) − τ b ( x ′ ) ∥ = ∥ x − x ′ ∥ \lVert\tau_{b}(x)-\tau_{b}(x')\rVert=\lVert x-x'\rVert ∥ τ b ( x ) − τ b ( x ′ )∥ = ∥ x − x ′ ∥ . By induction on k k k (Principle of Induction for the Natural Numbers , on the set of k ∈ N k\in\mathbb{N} k ∈ N such that for every g g g of class C k C^{k} C k on R d \mathbb{R}^{d} R d and every b b b the function g b g^{b} g b is of class C k C^{k} C k on R d \mathbb{R}^{d} R d with ∂ i ( g b ) = ( ∂ i g ) b \partial_{i}(g^{b})=(\partial_{i}g)^{b} ∂ i ( g b ) = ( ∂ i g ) b for all i i i ), using clauses 1 and 2 of C^k Maps on a Euclidean Open Set : if g g g is of class C 2 C^{2} C 2 then so is g b g^{b} g b , with ∂ j ∂ i ( g b ) = ( ∂ j ∂ i g ) b \partial_{j}\partial_{i}(g^{b})=(\partial_{j}\partial_{i}g)^{b} ∂ j ∂ i ( g b ) = ( ∂ j ∂ i g ) b , so that D 2 ( g b ) ( x ) = D 2 g ( x + b ) D^{2}(g^{b})(x)=D^{2}g(x+b) D 2 ( g b ) ( x ) = D 2 g ( x + b ) (Hessian Matrix of a C^2 Function , entrywise).
Proof of claim 3. Let ϕ \phi ϕ be of class C 2 C^{2} C 2 on R d \mathbb{R}^{d} R d and φ ( μ ) = ϕ ( m ( μ ) ) \varphi(\mu)=\phi(m(\mu)) φ ( μ ) = ϕ ( m ( μ )) , with lift Φ = φ ∘ Λ \Phi=\varphi\circ\Lambda Φ = φ ∘ Λ (The Lift of a Function on the Wasserstein Space to the Space of Square-Integrable Random Vectors §lift ), so Φ ( X ) = ϕ ( m ( X ) ) \Phi(X)=\phi(\mathsf{m}(X)) Φ ( X ) = ϕ ( m ( X )) . We verify the four properties of Test Functions on the Wasserstein Space: the Intrinsic Gradient and the Translation Hessian §test .
(a). Let X ∈ L 2 ( Ω ; R d ) X\in L^{2}(\Omega;\mathbb{R}^{d}) X ∈ L 2 ( Ω ; R d ) , a = m ( X ) a=\mathsf{m}(X) a = m ( X ) and p = D ϕ ( a ) p=D\phi(a) p = D ϕ ( a ) , the point with coordinates ∂ i ϕ ( a ) \partial_{i}\phi(a) ∂ i ϕ ( a ) (Differential Calculus and Convexity on Euclidean Open Sets: Standing Notation §derivatives ). Since ϕ \phi ϕ is of class C 1 C^{1} C 1 (clause 2 of C^k Maps on a Euclidean Open Set ), it is differentiable at a a a with derivative matrix the 1 × d 1\times d 1 × d matrix of the partials ∂ i ϕ ( a ) \partial_{i}\phi(a) ∂ i ϕ ( a ) by A Real-Valued C^1 Function is Differentiable at Every Point ; so by Differentiability at a Point for Maps Between Euclidean Spaces , for every ε > 0 \varepsilon>0 ε > 0 there is δ > 0 \delta>0 δ > 0 such that ∣ ϕ ( a + h ) − ϕ ( a ) − p ⋅ h ∣ ≤ ε ∥ h ∥ |\phi(a+h)-\phi(a)-p\cdot h|\le\varepsilon\lVert h\rVert ∣ ϕ ( a + h ) − ϕ ( a ) − p ⋅ h ∣ ≤ ε ∥ h ∥ whenever 0 < ∥ h ∥ < δ 0<\lVert h\rVert<\delta 0 < ∥ h ∥ < δ , and trivially also for h = 0 R d h=0_{\mathbb{R}^{d}} h = 0 R d . Let H ∈ L 2 ( Ω ; R d ) H\in L^{2}(\Omega;\mathbb{R}^{d}) H ∈ L 2 ( Ω ; R d ) with ∥ H ∥ L 2 < δ \lVert H\rVert_{L^{2}}<\delta ∥ H ∥ L 2 < δ and put h = m ( H ) h=\mathsf{m}(H) h = m ( H ) , so ∥ h ∥ < δ \lVert h\rVert<\delta ∥ h ∥ < δ by (L). By the expectation-map paragraph and (C),
∣ Φ ( X + H ) − Φ ( X ) − ⟨ c p , H ⟩ L 2 ∣ = ∣ ϕ ( a + h ) − ϕ ( a ) − p ⋅ h ∣ ≤ ε ∥ h ∥ ≤ ε ∥ H ∥ L 2 . \bigl|\Phi(X+H)-\Phi(X)-\langle c_{p},H\rangle_{L^{2}}\bigr|=\bigl|\phi(a+h)-\phi(a)-p\cdot h\bigr|\le\varepsilon\lVert h\rVert\le\varepsilon\lVert H\rVert_{L^{2}}. Φ ( X + H ) − Φ ( X ) − ⟨ c p , H ⟩ L 2 = ϕ ( a + h ) − ϕ ( a ) − p ⋅ h ≤ ε ∥ h ∥ ≤ ε ∥ H ∥ L 2 .
Hence Φ \Phi Φ is differentiable at X X X with gradient D Φ ( X ) = c D ϕ ( m ( X ) ) D\Phi(X)=c_{D\phi(\mathsf{m}(X))} D Φ ( X ) = c D ϕ ( m ( X )) (Fréchet Differentiability and the Gradient on an Open Subset of a Real Inner Product Space §differentiable , Fréchet Differentiability and the Gradient on an Open Subset of a Real Inner Product Space §gradient ). The gradient map is continuous: given X X X and ε > 0 \varepsilon>0 ε > 0 , the d d d continuous functions ∂ i ϕ \partial_{i}\phi ∂ i ϕ give δ > 0 \delta>0 δ > 0 with ∣ ∂ i ϕ ( a ′ ) − ∂ i ϕ ( a ) ∣ ≤ ε ( 2 d ) − 1 |\partial_{i}\phi(a')-\partial_{i}\phi(a)|\le\varepsilon(2\sqrt{d})^{-1} ∣ ∂ i ϕ ( a ′ ) − ∂ i ϕ ( a ) ∣ ≤ ε ( 2 d ) − 1 for all i i i whenever ∥ a ′ − a ∥ < δ \lVert a'-a\rVert<\delta ∥ a ′ − a ∥ < δ (the least of d d d radii, by claim 9 of Elementary Order Arithmetic in an Ordered Field applied repeatedly; continuity at a point as in Continuous Map Between Metric Spaces ); for ∥ Y − X ∥ L 2 < δ \lVert Y-X\rVert_{L^{2}}<\delta ∥ Y − X ∥ L 2 < δ one has ∥ m ( Y ) − m ( X ) ∥ < δ \lVert\mathsf{m}(Y)-\mathsf{m}(X)\rVert<\delta ∥ m ( Y ) − m ( X )∥ < δ by (L), hence ∥ D Φ ( Y ) − D Φ ( X ) ∥ L 2 2 = ∥ D ϕ ( m ( Y ) ) − D ϕ ( m ( X ) ) ∥ 2 ≤ d ⋅ ε 2 ( 4 d ) − 1 = ε 2 / 4 \lVert D\Phi(Y)-D\Phi(X)\rVert_{L^{2}}^{2}=\lVert D\phi(\mathsf{m}(Y))-D\phi(\mathsf{m}(X))\rVert^{2}\le d\cdot\varepsilon^{2}(4d)^{-1}=\varepsilon^{2}/4 ∥ D Φ ( Y ) − D Φ ( X ) ∥ L 2 2 = ∥ D ϕ ( m ( Y )) − D ϕ ( m ( X )) ∥ 2 ≤ d ⋅ ε 2 ( 4 d ) − 1 = ε 2 /4 by claim 1 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n , claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field and claim 1 of Comparison and Absolute Value Bounds for Finite Sums of Real Numbers , so ∥ D Φ ( Y ) − D Φ ( X ) ∥ L 2 ≤ ε / 2 < ε \lVert D\Phi(Y)-D\Phi(X)\rVert_{L^{2}}\le\varepsilon/2<\varepsilon ∥ D Φ ( Y ) − D Φ ( X ) ∥ L 2 ≤ ε /2 < ε (claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field again, and claim 8 of Elementary Order Arithmetic in an Ordered Field ). Thus Φ ∈ C 1 ( L 2 ( Ω ; R d ) ) \Phi\in C^{1}(L^{2}(\Omega;\mathbb{R}^{d})) Φ ∈ C 1 ( L 2 ( Ω ; R d )) (The Classes C 1 C^1 C 1 and C 2 C^2 C 2 on an Open Subset of a Real Inner Product Space §c1 ).
(b). Let μ ∈ P 2 ( R d ) \mu\in\mathcal{P}_{2}(\mathbb{R}^{d}) μ ∈ P 2 ( R d ) and let η \eta η be the class of the constant map with value p = D ϕ ( m ( μ ) ) p=D\phi(m(\mu)) p = D ϕ ( m ( μ )) , which lies in T μ T_{\mu} T μ by Gradients of Functions with Bounded Derivatives Belong to the Tangent Space; Constants Are Tangent; the Score Identity; the Score Has Mean Zero; Translation of Tangent Fields §constants . For X X X with L ( X ) = μ \mathcal{L}(X)=\mu L ( X ) = μ , m ( X ) = m ( μ ) \mathsf{m}(X)=m(\mu) m ( X ) = m ( μ ) , and η ∘ X \eta\circ X η ∘ X (Composition of a Square-Integrable Vector Field with a Random Vector, and the Lifted Score: Isometry, Norm, Weak Identity and Second-Moment Identity §composition ) is the class of the constant map with value p p p , i.e. c p = D Φ ( X ) c_{p}=D\Phi(X) c p = D Φ ( X ) .
(c). Let X ∈ L 2 ( Ω ; R d ) X\in L^{2}(\Omega;\mathbb{R}^{d}) X ∈ L 2 ( Ω ; R d ) and ϕ X ( a ) = Φ ( X + c a ) = ϕ ( m ( X ) + a ) = ϕ m ( X ) ( a ) \phi_{X}(a)=\Phi(X+c_{a})=\phi(\mathsf{m}(X)+a)=\phi^{\mathsf{m}(X)}(a) ϕ X ( a ) = Φ ( X + c a ) = ϕ ( m ( X ) + a ) = ϕ m ( X ) ( a ) . By the translates paragraph, ϕ X \phi_{X} ϕ X is of class C 2 C^{2} C 2 on R d \mathbb{R}^{d} R d with D 2 ϕ X ( 0 R d ) = D 2 ϕ ( m ( X ) ) D^{2}\phi_{X}(0_{\mathbb{R}^{d}})=D^{2}\phi(\mathsf{m}(X)) D 2 ϕ X ( 0 R d ) = D 2 ϕ ( m ( X )) ; so Φ \Phi Φ is twice continuously differentiable along translations at every X X X (The Translation Laplacian of a Function on the Space of Square-Integrable Random Vectors §translations , The Translation Laplacian of a Function on the Space of Square-Integrable Random Vectors §on-space ).
(d). The map m : P 2 ( R d ) → R d m:\mathcal{P}_{2}(\mathbb{R}^{d})\to\mathbb{R}^{d} m : P 2 ( R d ) → R d is Lipschitz with constant 1 1 1 by claim 1 (Lipschitz Map Between Metric Spaces , with W 2 W_{2} W 2 and the Euclidean distance d E ( x , y ) = ∥ x − y ∥ d_{E}(x,y)=\lVert x-y\rVert d E ( x , y ) = ∥ x − y ∥ , claim 2 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n ), hence continuous by A Lipschitz Map is Uniformly Continuous ; ϕ \phi ϕ is continuous by claim 3 of Euclidean Space is Open in Itself, and C k C^k C k Maps are Continuous ; so φ = ϕ ∘ m \varphi=\phi\circ m φ = ϕ ∘ m is continuous by claim 3 of Semicontinuity and Continuity Under Composition with a Continuous Map .
Hence φ \varphi φ is a test function; its intrinsic gradient at μ \mu μ is the η \eta η of (b) (Test Functions on the Wasserstein Space: the Intrinsic Gradient and the Translation Hessian §gradient ), and its translation Hessian at μ \mu μ is D 2 ϕ X ( 0 R d ) = D 2 ϕ ( m ( μ ) ) D^{2}\phi_{X}(0_{\mathbb{R}^{d}})=D^{2}\phi(m(\mu)) D 2 ϕ X ( 0 R d ) = D 2 ϕ ( m ( μ )) for any X X X with L ( X ) = μ \mathcal{L}(X)=\mu L ( X ) = μ (Test Functions on the Wasserstein Space: the Intrinsic Gradient and the Translation Hessian §hessian ).
Proof of claim 4. Let φ \varphi φ be a test function with lift Φ \Phi Φ , and define the centring operator C : L 2 ( Ω ; R d ) → L 2 ( Ω ; R d ) \mathsf{C}:L^{2}(\Omega;\mathbb{R}^{d})\to L^{2}(\Omega;\mathbb{R}^{d}) C : L 2 ( Ω ; R d ) → L 2 ( Ω ; R d ) by C X = X − c m ( X ) \mathsf{C}X=X-c_{\mathsf{m}(X)} C X = X − c m ( X ) . By the lift paragraph of claim 2, L ( C X ) = L ( X ) ‾ \mathcal{L}(\mathsf{C}X)=\overline{\mathcal{L}(X)} L ( C X ) = L ( X ) , so the lift Φ ∘ \Phi^{\circ} Φ ∘ of φ ∘ \varphi^{\circ} φ ∘ satisfies Φ ∘ ( X ) = φ ( L ( X ) ‾ ) = Φ ( C X ) \Phi^{\circ}(X)=\varphi(\overline{\mathcal{L}(X)})=\Phi(\mathsf{C}X) Φ ∘ ( X ) = φ ( L ( X ) ) = Φ ( C X ) . By the expectation-map paragraph, C ( X + H ) = C X + C H \mathsf{C}(X+H)=\mathsf{C}X+\mathsf{C}H C ( X + H ) = C X + C H and, using claim 2 and The Space of Square-Integrable Random Vectors is a Real Hilbert Space; Its Laws Have Finite Second Moment; Constants and Translations §law , ∥ C H ∥ L 2 2 = M 2 ( L ( H ) ‾ ) = M 2 ( L ( H ) ) − ∥ m ( H ) ∥ 2 ≤ ∥ H ∥ L 2 2 \lVert\mathsf{C}H\rVert_{L^{2}}^{2}=M_{2}(\overline{\mathcal{L}(H)})=M_{2}(\mathcal{L}(H))-\lVert\mathsf{m}(H)\rVert^{2}\le\lVert H\rVert_{L^{2}}^{2} ∥ C H ∥ L 2 2 = M 2 ( L ( H ) ) = M 2 ( L ( H )) − ∥ m ( H ) ∥ 2 ≤ ∥ H ∥ L 2 2 , so
∥ C H ∥ L 2 ≤ ∥ H ∥ L 2 and ∥ C X − C Y ∥ L 2 = ∥ C ( X − Y ) ∥ L 2 ≤ ∥ X − Y ∥ L 2 . (N) \lVert\mathsf{C}H\rVert_{L^{2}}\le\lVert H\rVert_{L^{2}}\qquad\text{and}\qquad\lVert\mathsf{C}X-\mathsf{C}Y\rVert_{L^{2}}=\lVert\mathsf{C}(X-Y)\rVert_{L^{2}}\le\lVert X-Y\rVert_{L^{2}}.\tag{N} ∥ C H ∥ L 2 ≤ ∥ H ∥ L 2 and ∥ C X − C Y ∥ L 2 = ∥ C ( X − Y ) ∥ L 2 ≤ ∥ X − Y ∥ L 2 . ( N )
Moreover, by Elementary Identities in a Real Inner Product Space §bilinear and (C), ⟨ C H , G ⟩ L 2 = ⟨ H , G ⟩ L 2 − m ( H ) ⋅ m ( G ) \langle\mathsf{C}H,G\rangle_{L^{2}}=\langle H,G\rangle_{L^{2}}-\mathsf{m}(H)\cdot\mathsf{m}(G) ⟨ C H , G ⟩ L 2 = ⟨ H , G ⟩ L 2 − m ( H ) ⋅ m ( G ) , which is symmetric in H H H and G G G ; hence
⟨ C H , G ⟩ L 2 = ⟨ H , C G ⟩ L 2 ( H , G ∈ L 2 ( Ω ; R d ) ) . (S) \langle\mathsf{C}H,G\rangle_{L^{2}}=\langle H,\mathsf{C}G\rangle_{L^{2}}\qquad(H,G\in L^{2}(\Omega;\mathbb{R}^{d})).\tag{S} ⟨ C H , G ⟩ L 2 = ⟨ H , C G ⟩ L 2 ( H , G ∈ L 2 ( Ω ; R d )) . ( S )
(a). Let X ∈ L 2 ( Ω ; R d ) X\in L^{2}(\Omega;\mathbb{R}^{d}) X ∈ L 2 ( Ω ; R d ) and Y = C X Y=\mathsf{C}X Y = C X . Given ε > 0 \varepsilon>0 ε > 0 , let δ > 0 \delta>0 δ > 0 be provided by the differentiability of Φ \Phi Φ at Y Y Y (Fréchet Differentiability and the Gradient on an Open Subset of a Real Inner Product Space §differentiable , property (a) of the test function φ \varphi φ ). For ∥ H ∥ L 2 < δ \lVert H\rVert_{L^{2}}<\delta ∥ H ∥ L 2 < δ , (N) gives ∥ C H ∥ L 2 < δ \lVert\mathsf{C}H\rVert_{L^{2}}<\delta ∥ C H ∥ L 2 < δ , so by (S)
∣ Φ ∘ ( X + H ) − Φ ∘ ( X ) − ⟨ C D Φ ( Y ) , H ⟩ L 2 ∣ = ∣ Φ ( Y + C H ) − Φ ( Y ) − ⟨ D Φ ( Y ) , C H ⟩ L 2 ∣ ≤ ε ∥ C H ∥ L 2 ≤ ε ∥ H ∥ L 2 . \bigl|\Phi^{\circ}(X+H)-\Phi^{\circ}(X)-\langle\mathsf{C}D\Phi(Y),H\rangle_{L^{2}}\bigr|=\bigl|\Phi(Y+\mathsf{C}H)-\Phi(Y)-\langle D\Phi(Y),\mathsf{C}H\rangle_{L^{2}}\bigr|\le\varepsilon\lVert\mathsf{C}H\rVert_{L^{2}}\le\varepsilon\lVert H\rVert_{L^{2}}. Φ ∘ ( X + H ) − Φ ∘ ( X ) − ⟨ C D Φ ( Y ) , H ⟩ L 2 = Φ ( Y + C H ) − Φ ( Y ) − ⟨ D Φ ( Y ) , C H ⟩ L 2 ≤ ε ∥ C H ∥ L 2 ≤ ε ∥ H ∥ L 2 .
Thus Φ ∘ \Phi^{\circ} Φ ∘ is differentiable at X X X with D Φ ∘ ( X ) = C ( D Φ ( C X ) ) D\Phi^{\circ}(X)=\mathsf{C}\bigl(D\Phi(\mathsf{C}X)\bigr) D Φ ∘ ( X ) = C ( D Φ ( C X ) ) . The gradient map X ↦ C ( D Φ ( C X ) ) X\mapsto\mathsf{C}(D\Phi(\mathsf{C}X)) X ↦ C ( D Φ ( C X )) is continuous, as the composition of C \mathsf{C} C (Lipschitz by (N), hence continuous by A Lipschitz Map is Uniformly Continuous ), the continuous map D Φ D\Phi D Φ , and C \mathsf{C} C again (claim 3 of Semicontinuity and Continuity Under Composition with a Continuous Map ). So Φ ∘ ∈ C 1 ( L 2 ( Ω ; R d ) ) \Phi^{\circ}\in C^{1}(L^{2}(\Omega;\mathbb{R}^{d})) Φ ∘ ∈ C 1 ( L 2 ( Ω ; R d )) .
(b) and the gradient formula. Let μ ∈ P 2 ( R d ) \mu\in\mathcal{P}_{2}(\mathbb{R}^{d}) μ ∈ P 2 ( R d ) , m = m ( μ ) m=m(\mu) m = m ( μ ) , and let η \eta η be a representative of ∇ φ ( μ ˉ ) ∈ T μ ˉ \nabla\varphi(\bar{\mu})\in T_{\bar{\mu}} ∇ φ ( μ ˉ ) ∈ T μ ˉ . By Gradients of Functions with Bounded Derivatives Belong to the Tangent Space; Constants Are Tangent; the Score Identity; the Score Has Mean Zero; Translation of Tangent Fields §translation applied to μ ˉ \bar{\mu} μ ˉ , a = m a=m a = m and η \eta η , using ( τ m ) # μ ˉ = μ (\tau_{m})_{\#}\bar{\mu}=\mu ( τ m ) # μ ˉ = μ (claim 2): the map η ∗ : x ↦ η ( x − m ) \eta^{\ast}:x\mapsto\eta(x-m) η ∗ : x ↦ η ( x − m ) is Borel, its class in L 2 ( μ ; R d ) L^{2}(\mu;\mathbb{R}^{d}) L 2 ( μ ; R d ) does not depend on the representative, and it belongs to T μ T_{\mu} T μ . Since ∫ ∥ η ∗ ∥ 2 d μ < ∞ \int\lVert\eta^{\ast}\rVert^{2}d\mu<\infty ∫ ∥ η ∗ ∥ 2 d μ < ∞ and ∣ η i ∗ ∣ ≤ 1 2 ( ∥ η ∗ ∥ 2 + 1 ) |\eta^{\ast}_{i}|\le\tfrac12(\lVert\eta^{\ast}\rVert^{2}+1) ∣ η i ∗ ∣ ≤ 2 1 (∥ η ∗ ∥ 2 + 1 ) by (B), the coordinates η i ∗ \eta^{\ast}_{i} η i ∗ are integrable with respect to μ \mu μ (as in the first paragraph of the proof of claim 1); since η i ∗ = η i ∘ τ − m \eta^{\ast}_{i}=\eta_{i}\circ\tau_{-m} η i ∗ = η i ∘ τ − m , claim 2 of Image Measures, Measures with Densities, and Change of Variables for the image measure μ ˉ = ( τ − m ) # μ \bar{\mu}=(\tau_{-m})_{\#}\mu μ ˉ = ( τ − m ) # μ shows that the coordinate η i \eta_{i} η i is integrable with respect to μ ˉ \bar{\mu} μ ˉ with ∫ η i d μ ˉ = ∫ η i ∗ d μ \int\eta_{i}\,d\bar{\mu}=\int\eta^{\ast}_{i}\,d\mu ∫ η i d μ ˉ = ∫ η i ∗ d μ . Let b = ∫ η d μ ˉ b=\int\eta\,d\bar{\mu} b = ∫ η d μ ˉ be the point with these coordinates; if η ′ \eta' η ′ is another representative then μ ˉ ( { η = η ′ } ) = 1 \bar{\mu}(\{\eta=\eta'\})=1 μ ˉ ({ η = η ′ }) = 1 , so ∫ η i ′ d μ ˉ = ∫ η i d μ ˉ \int\eta'_{i}\,d\bar{\mu}=\int\eta_{i}\,d\bar{\mu} ∫ η i ′ d μ ˉ = ∫ η i d μ ˉ by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §comparison , and b b b does not depend on the representative. Now let X X X satisfy L ( X ) = μ \mathcal{L}(X)=\mu L ( X ) = μ and Y = C X = X − c m Y=\mathsf{C}X=X-c_{m} Y = C X = X − c m , so L ( Y ) = μ ˉ \mathcal{L}(Y)=\bar{\mu} L ( Y ) = μ ˉ and D Φ ( Y ) = ∇ φ ( μ ˉ ) ∘ Y D\Phi(Y)=\nabla\varphi(\bar{\mu})\circ Y D Φ ( Y ) = ∇ φ ( μ ˉ ) ∘ Y by property (b) of φ \varphi φ . For a representative of X X X , the map ω ↦ η ( Y ( ω ) ) = η ( X ( ω ) − m ) = η ∗ ( X ( ω ) ) \omega\mapsto\eta(Y(\omega))=\eta(X(\omega)-m)=\eta^{\ast}(X(\omega)) ω ↦ η ( Y ( ω )) = η ( X ( ω ) − m ) = η ∗ ( X ( ω )) shows ∇ φ ( μ ˉ ) ∘ Y = η ∗ ∘ X \nabla\varphi(\bar{\mu})\circ Y=\eta^{\ast}\circ X ∇ φ ( μ ˉ ) ∘ Y = η ∗ ∘ X (Composition of a Square-Integrable Vector Field with a Random Vector, and the Lifted Score: Isometry, Norm, Weak Identity and Second-Moment Identity §composition ). Its coordinates have expectations E [ η i ∗ ∘ X ] = ∫ η i ∗ d μ = b i \mathbb{E}[\eta^{\ast}_{i}\circ X]=\int\eta^{\ast}_{i}\,d\mu=b_{i} E [ η i ∗ ∘ X ] = ∫ η i ∗ d μ = b i (Basic Properties of Random Vectors: Coordinates, Borel Images and Arithmetic, Change of Variables, Almost Sure Equality and Pairs §expectation ), so m ( η ∗ ∘ X ) = b \mathsf{m}(\eta^{\ast}\circ X)=b m ( η ∗ ∘ X ) = b and
D Φ ∘ ( X ) = C ( η ∗ ∘ X ) = η ∗ ∘ X − c b = ( η ∗ − b ) ∘ X , D\Phi^{\circ}(X)=\mathsf{C}(\eta^{\ast}\circ X)=\eta^{\ast}\circ X-c_{b}=(\eta^{\ast}-b)\circ X, D Φ ∘ ( X ) = C ( η ∗ ∘ X ) = η ∗ ∘ X − c b = ( η ∗ − b ) ∘ X ,
by the linearity of composition in Composition of a Square-Integrable Vector Field with a Random Vector, and the Lifted Score: Isometry, Norm, Weak Identity and Second-Moment Identity §composition , the constant map with value b b b composed with X X X being c b c_{b} c b . The class of η ∗ − b \eta^{\ast}-b η ∗ − b , i.e. of x ↦ η ( x − m ) − ∫ η d μ ˉ x\mapsto\eta(x-m)-\int\eta\,d\bar{\mu} x ↦ η ( x − m ) − ∫ η d μ ˉ , lies in T μ T_{\mu} T μ : η ∗ ∈ T μ \eta^{\ast}\in T_{\mu} η ∗ ∈ T μ , b ∈ T μ b\in T_{\mu} b ∈ T μ by Gradients of Functions with Bounded Derivatives Belong to the Tangent Space; Constants Are Tangent; the Score Identity; the Score Has Mean Zero; Translation of Tangent Fields §constants , and T μ T_{\mu} T μ is a linear subspace by Basic Properties of the Tangent Space: Closed Subspace, the Identity Map Belongs to It, Second-Moment Limits, and Representation of Bounded Functionals on Gradients §closed . This gives property (b), and by Test Functions on the Wasserstein Space: the Intrinsic Gradient and the Translation Hessian §gradient the intrinsic gradient ∇ φ ∘ ( μ ) \nabla\varphi^{\circ}(\mu) ∇ φ ∘ ( μ ) is this class.
(c). For X ∈ L 2 ( Ω ; R d ) X\in L^{2}(\Omega;\mathbb{R}^{d}) X ∈ L 2 ( Ω ; R d ) and a ∈ R d a\in\mathbb{R}^{d} a ∈ R d , C ( X + c a ) = X + c a − c m ( X ) + a = X + c a − m ( X ) − a = X − c m ( X ) = C X \mathsf{C}(X+c_{a})=X+c_{a}-c_{\mathsf{m}(X)+a}=X+c_{a-\mathsf{m}(X)-a}=X-c_{\mathsf{m}(X)}=\mathsf{C}X C ( X + c a ) = X + c a − c m ( X ) + a = X + c a − m ( X ) − a = X − c m ( X ) = C X by the expectation-map paragraph and Translations on a Space of Square-Integrable Random Vectors: Constant Classes, Law Invariance, the Translation Derivative and the Translation Hessian §constants . Hence ϕ X ∘ ( a ) = Φ ∘ ( X + c a ) = Φ ( C X ) \phi^{\circ}_{X}(a)=\Phi^{\circ}(X+c_{a})=\Phi(\mathsf{C}X) ϕ X ∘ ( a ) = Φ ∘ ( X + c a ) = Φ ( C X ) is constant in a a a , a smooth function (claim 2 of Constants, Coordinate Functions, Sums and Products of C k C^k C k Functions on a Euclidean Open Set ) whose partial derivatives vanish identically (its difference quotients are 0 0 0 , Partial Derivative on a Euclidean Open Set ), as do their partial derivatives; so ϕ X ∘ \phi^{\circ}_{X} ϕ X ∘ is of class C 2 C^{2} C 2 with D 2 ϕ X ∘ ( 0 R d ) = 0 d D^{2}\phi^{\circ}_{X}(0_{\mathbb{R}^{d}})=0_{d} D 2 ϕ X ∘ ( 0 R d ) = 0 d . Thus Φ ∘ \Phi^{\circ} Φ ∘ is twice continuously differentiable along translations at every point, and the translation Hessian of φ ∘ \varphi^{\circ} φ ∘ at every μ \mu μ is 0 d 0_{d} 0 d (Test Functions on the Wasserstein Space: the Intrinsic Gradient and the Translation Hessian §hessian ).
(d) and translation invariance. For μ \mu μ and a a a , φ ∘ ( ( τ a ) # μ ) = φ ( ( τ a ) # μ ‾ ) = φ ( μ ˉ ) = φ ∘ ( μ ) \varphi^{\circ}((\tau_{a})_{\#}\mu)=\varphi(\overline{(\tau_{a})_{\#}\mu})=\varphi(\bar{\mu})=\varphi^{\circ}(\mu) φ ∘ (( τ a ) # μ ) = φ ( ( τ a ) # μ ) = φ ( μ ˉ ) = φ ∘ ( μ ) by claim 2. For continuity at μ \mu μ : given ε > 0 \varepsilon>0 ε > 0 , property (d) of φ \varphi φ gives δ > 0 \delta>0 δ > 0 with ∣ φ ( σ ) − φ ( μ ˉ ) ∣ < ε |\varphi(\sigma)-\varphi(\bar{\mu})|<\varepsilon ∣ φ ( σ ) − φ ( μ ˉ ) ∣ < ε whenever W 2 ( σ , μ ˉ ) < δ W_{2}(\sigma,\bar{\mu})<\delta W 2 ( σ , μ ˉ ) < δ (Continuous Map Between Metric Spaces ); if W 2 ( ν , μ ) < δ W_{2}(\nu,\mu)<\delta W 2 ( ν , μ ) < δ then W 2 ( ν ˉ , μ ˉ ) ≤ W 2 ( ν , μ ) < δ W_{2}(\bar{\nu},\bar{\mu})\le W_{2}(\nu,\mu)<\delta W 2 ( ν ˉ , μ ˉ ) ≤ W 2 ( ν , μ ) < δ by claim 2 (and the symmetry of W 2 W_{2} W 2 , The Quadratic Wasserstein Distance is a Metric on the Wasserstein Space §symmetry ), so ∣ φ ∘ ( ν ) − φ ∘ ( μ ) ∣ < ε |\varphi^{\circ}(\nu)-\varphi^{\circ}(\mu)|<\varepsilon ∣ φ ∘ ( ν ) − φ ∘ ( μ ) ∣ < ε . Hence φ ∘ \varphi^{\circ} φ ∘ is a test function with the stated gradient and Hessian.