Each result cited below is universally quantified over the data in its own statement, and is applied to the data named at the point of use. The rules for adding, multiplying and comparing inequalities between real numbers in Elementary Order Arithmetic in an Ordered Field and Elementary Arithmetic in an Ordered Field , and Properties of the Absolute Value in an Ordered Field , are used without further mention. Couplings, Π \Pi Π , the quadratic cost I I I , M 2 M_{2} M 2 and W 2 W_{2} W 2 are those of Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §dimensions in dimension m m m ; p r 1 , p r 2 : R m + m → R m \mathrm{pr}_{1},\mathrm{pr}_{2}:\mathbb{R}^{m+m}\to\mathbb{R}^{m} pr 1 , pr 2 : R m + m → R m are the coordinate projections and ι = ι m , m \iota=\iota^{m,m} ι = ι m , m the concatenation map of Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pairs . For real L > 0 L>0 L > 0 put E L = { x ∈ R m : L < ∥ x ∥ } E_{L}=\{x\in\mathbb{R}^{m}:L<\lVert x\rVert\} E L = { x ∈ R m : L < ∥ x ∥} ; it is Borel, being the preimage of the open set { t ∈ R : L < t } \{t\in\mathbb{R}:L<t\} { t ∈ R : L < t } (Borel by claims 4 and 5 of The Borel Sigma-Algebra of a Euclidean Space as a Product, and Measurability of Projections, Sequentially Continuous Maps, and Open and Closed Sets ) under the Borel map x ↦ ∥ x ∥ x\mapsto\lVert x\rVert x ↦ ∥ x ∥ of Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions ; and for ρ ∈ P ( R m ) \rho\in\mathcal{P}(\mathbb{R}^{m}) ρ ∈ P ( R m ) we write ∫ E L ∥ x ∥ 2 ρ ( d x ) \int_{E_{L}}\lVert x\rVert^{2}\,\rho(dx) ∫ E L ∥ x ∥ 2 ρ ( d x ) for ∫ R m ∥ x ∥ 2 1 E L ( x ) ρ ( d x ) ∈ [ 0 , ∞ ] \int_{\mathbb{R}^{m}}\lVert x\rVert^{2}\mathbf{1}_{E_{L}}(x)\,\rho(dx)\in[0,\infty] ∫ R m ∥ x ∥ 2 1 E L ( x ) ρ ( d x ) ∈ [ 0 , ∞ ] . Nonnegative Borel functions are integrated in [ 0 , ∞ ] [0,\infty] [ 0 , ∞ ] as in Probability Measures on Euclidean Space and Random Vectors: Standing Notation §measures , and every bounded Borel function is integrable against every member of P ( R m ) \mathcal{P}(\mathbb{R}^{m}) P ( R m ) by that clause.
Part 1 (clause 1). Let ( μ n ) n ∈ N (\mu_{n})_{n\in\mathbb{N}} ( μ n ) n ∈ N and μ \mu μ be as in clause 1.
Step 1 (A continuous cutoff). For x , x ′ ∈ R m x,x'\in\mathbb{R}^{m} x , x ′ ∈ R m , claims 2, 5 and 6 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n give ∣ ∥ x ∥ − ∥ x ′ ∥ ∣ ≤ ∥ x − x ′ ∥ = d E ( x , x ′ ) \bigl|\lVert x\rVert-\lVert x'\rVert\bigr|\le\lVert x-x'\rVert=d_{E}(x,x') ∥ x ∥ − ∥ x ′ ∥ ≤ ∥ x − x ′ ∥ = d E ( x , x ′ ) , so x ↦ ∥ x ∥ x\mapsto\lVert x\rVert x ↦ ∥ x ∥ is Lipschitz, hence continuous by A Lipschitz Map is Uniformly Continuous . For real K K K define χ K ( x ) = min { 1 , max { 0 , ∥ x ∥ − K } } \chi_{K}(x)=\min\{1,\max\{0,\lVert x\rVert-K\}\} χ K ( x ) = min { 1 , max { 0 , ∥ x ∥ − K }} . By claims 1, 2 and 5 of Continuity of Sums and Products of Real-Valued Functions on a Metric Space and claim 4 of Continuity of the Absolute Value, Maximum and Minimum of Real-Valued Functions on a Metric Space , χ K \chi_{K} χ K is continuous; by Elementary Properties of the Maximum of Two Elements and Elementary Properties of the Minimum of Two Elements , 0 ≤ χ K ≤ 1 0\le\chi_{K}\le1 0 ≤ χ K ≤ 1 , χ K ( x ) = 0 \chi_{K}(x)=0 χ K ( x ) = 0 when ∥ x ∥ ≤ K \lVert x\rVert\le K ∥ x ∥ ≤ K , and χ K ( x ) = 1 \chi_{K}(x)=1 χ K ( x ) = 1 when K + 1 ≤ ∥ x ∥ K+1\le\lVert x\rVert K + 1 ≤ ∥ x ∥ .
Step 2 (Uniform tails, and μ ∈ P 2 ( R m ) \mu\in\mathcal{P}_{2}(\mathbb{R}^{m}) μ ∈ P 2 ( R m ) ). We show:
(⋆ \star ⋆ ) for every real ε > 0 \varepsilon>0 ε > 0 there is a real L > 0 L>0 L > 0 with ∫ E L ∥ x ∥ 2 ρ ( d x ) ≤ ε \int_{E_{L}}\lVert x\rVert^{2}\,\rho(dx)\le\varepsilon ∫ E L ∥ x ∥ 2 ρ ( d x ) ≤ ε for ρ = μ \rho=\mu ρ = μ and for ρ = μ n \rho=\mu_{n} ρ = μ n , every n ∈ N n\in\mathbb{N} n ∈ N ; and M 2 ( μ ) ≤ L 2 + ε M_{2}(\mu)\le L^{2}+\varepsilon M 2 ( μ ) ≤ L 2 + ε .
Let ε > 0 \varepsilon>0 ε > 0 and let K > 0 K>0 K > 0 be given by the hypothesis of clause 1. For j ∈ N j\in\mathbb{N} j ∈ N let f j ( x ) = min { j , ∥ x ∥ 2 χ K ( x ) } f_{j}(x)=\min\{j,\lVert x\rVert^{2}\chi_{K}(x)\} f j ( x ) = min { j , ∥ x ∥ 2 χ K ( x )} ; it is continuous (Step 1 and the continuity results cited there, ∥ x ∥ 2 = ∥ x ∥ ∥ x ∥ \lVert x\rVert^{2}=\lVert x\rVert\,\lVert x\rVert ∥ x ∥ 2 = ∥ x ∥ ∥ x ∥ ) with 0 ≤ f j ≤ j 0\le f_{j}\le j 0 ≤ f j ≤ j , hence bounded . Since χ K \chi_{K} χ K vanishes off E K E_{K} E K and χ K ≤ 1 \chi_{K}\le1 χ K ≤ 1 , f j ≤ ∥ x ∥ 2 1 E K f_{j}\le\lVert x\rVert^{2}\mathbf{1}_{E_{K}} f j ≤ ∥ x ∥ 2 1 E K pointwise, so claim 1 of Linearity and Monotonicity of the Lebesgue Integral gives ∫ f j d μ n ≤ ∫ E K ∥ x ∥ 2 μ n ( d x ) < ε \int f_{j}\,d\mu_{n}\le\int_{E_{K}}\lVert x\rVert^{2}\,\mu_{n}(dx)<\varepsilon ∫ f j d μ n ≤ ∫ E K ∥ x ∥ 2 μ n ( d x ) < ε for every n n n . By the definition of weak convergence , ( ∫ f j d μ n ) n (\int f_{j}\,d\mu_{n})_{n} ( ∫ f j d μ n ) n converges to ∫ f j d μ \int f_{j}\,d\mu ∫ f j d μ , so ∫ f j d μ ≤ ε \int f_{j}\,d\mu\le\varepsilon ∫ f j d μ ≤ ε by claim 1 of Order Properties of Limits of Real Sequences . The sequence ( f j ) j (f_{j})_{j} ( f j ) j is nondecreasing and, for each x x x , claim 1 of The Archimedean Property of the Real Numbers gives j j j with ∥ x ∥ 2 χ K ( x ) < j \lVert x\rVert^{2}\chi_{K}(x)<j ∥ x ∥ 2 χ K ( x ) < j , so sup j f j ( x ) = ∥ x ∥ 2 χ K ( x ) \sup_{j}f_{j}(x)=\lVert x\rVert^{2}\chi_{K}(x) sup j f j ( x ) = ∥ x ∥ 2 χ K ( x ) ; by Monotone Convergence Theorem , ∫ ∥ x ∥ 2 χ K ( x ) μ ( d x ) ≤ ε \int\lVert x\rVert^{2}\chi_{K}(x)\,\mu(dx)\le\varepsilon ∫ ∥ x ∥ 2 χ K ( x ) μ ( d x ) ≤ ε . Put L = K + 1 L=K+1 L = K + 1 . By Step 1, ∥ x ∥ 2 1 E L ( x ) ≤ ∥ x ∥ 2 χ K ( x ) \lVert x\rVert^{2}\mathbf{1}_{E_{L}}(x)\le\lVert x\rVert^{2}\chi_{K}(x) ∥ x ∥ 2 1 E L ( x ) ≤ ∥ x ∥ 2 χ K ( x ) , so ∫ E L ∥ x ∥ 2 μ ( d x ) ≤ ε \int_{E_{L}}\lVert x\rVert^{2}\,\mu(dx)\le\varepsilon ∫ E L ∥ x ∥ 2 μ ( d x ) ≤ ε ; and E L ⊆ E K E_{L}\subseteq E_{K} E L ⊆ E K gives ∫ E L ∥ x ∥ 2 μ n ( d x ) < ε \int_{E_{L}}\lVert x\rVert^{2}\,\mu_{n}(dx)<\varepsilon ∫ E L ∥ x ∥ 2 μ n ( d x ) < ε for every n n n . Finally ∥ x ∥ 2 ≤ L 2 + ∥ x ∥ 2 1 E L ( x ) \lVert x\rVert^{2}\le L^{2}+\lVert x\rVert^{2}\mathbf{1}_{E_{L}}(x) ∥ x ∥ 2 ≤ L 2 + ∥ x ∥ 2 1 E L ( x ) for every x x x (claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field when ∥ x ∥ ≤ L \lVert x\rVert\le L ∥ x ∥ ≤ L ), so M 2 ( μ ) ≤ L 2 + ε < ∞ M_{2}(\mu)\le L^{2}+\varepsilon<\infty M 2 ( μ ) ≤ L 2 + ε < ∞ by claim 1 of Linearity and Monotonicity of the Lebesgue Integral . This proves (⋆ \star ⋆ ); in particular μ ∈ P 2 ( R m ) \mu\in\mathcal{P}_{2}(\mathbb{R}^{m}) μ ∈ P 2 ( R m ) by The Second Moment of a Probability Measure on Euclidean Space and the Probability Measures with Finite Second Moment §space .
Step 3 (A partition of unity; order of choices). Fix a real ε > 0 \varepsilon>0 ε > 0 . Choose L > 0 L>0 L > 0 by (⋆ \star ⋆ ) for this ε \varepsilon ε , and put η = ε > 0 \eta=\sqrt{\varepsilon}>0 η = ε > 0 . Let C = { x : ∥ x ∥ ≤ L } C=\{x:\lVert x\rVert\le L\} C = { x : ∥ x ∥ ≤ L } , the closed ball of ( R m , d E ) (\mathbb{R}^{m},d_{E}) ( R m , d E ) with centre 0 0 0 and radius L L L (claim 2 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n ). It is closed and bounded by claims 3 and 2 of Elementary Properties of the Closed Ball in a Metric Space , hence compact by Heine-Borel Theorem in R n \mathbb{R}^n R n . The open balls B ( z , η / 2 ) B(z,\eta/2) B ( z , η /2 ) , z ∈ C z\in C z ∈ C , are open by Open Ball in a Metric Space is Open and cover C C C , so Compact Subset Criterion via Open Covers in the Ambient Space yields a finite J ⊆ C J\subseteq C J ⊆ C with C ⊆ ⋃ z ∈ J B ( z , η / 2 ) C\subseteq\bigcup_{z\in J}B(z,\eta/2) C ⊆ ⋃ z ∈ J B ( z , η /2 ) ; J ≠ ∅ J\neq\emptyset J = ∅ because 0 ∈ C 0\in C 0 ∈ C , so J J J has k k k elements for some natural number k ≥ 1 k\ge1 k ≥ 1 , which we enumerate as z 1 , … , z k z_{1},\dots,z_{k} z 1 , … , z k . Apply Continuous Partition of Unity Subordinate to a Finite Open Cover of a Compact Set in a Metric Space with X = R m X=\mathbb{R}^{m} X = R m , d = d E d=d_{E} d = d E , this C C C , n = k n=k n = k and U i = B ( z i , η / 2 ) U_{i}=B(z_{i},\eta/2) U i = B ( z i , η /2 ) : it gives continuous φ i : R m → [ 0 , 1 ] \varphi_{i}:\mathbb{R}^{m}\to[0,1] φ i : R m → [ 0 , 1 ] and closed D i ⊆ U i D_{i}\subseteq U_{i} D i ⊆ U i (i ∈ [ k ] i\in[k] i ∈ [ k ] ) with φ i = 0 \varphi_{i}=0 φ i = 0 off D i D_{i} D i (its claims 1 and 2), ∑ i = 1 k φ i ≤ 1 \sum_{i=1}^{k}\varphi_{i}\le1 ∑ i = 1 k φ i ≤ 1 everywhere (claim 3) and ∑ i = 1 k φ i = 1 \sum_{i=1}^{k}\varphi_{i}=1 ∑ i = 1 k φ i = 1 on C C C (claim 4). Put φ 0 = 1 − ∑ i = 1 k φ i \varphi_{0}=1-\sum_{i=1}^{k}\varphi_{i} φ 0 = 1 − ∑ i = 1 k φ i ; it is continuous, 0 ≤ φ 0 ≤ 1 0\le\varphi_{0}\le1 0 ≤ φ 0 ≤ 1 , and φ 0 = 0 \varphi_{0}=0 φ 0 = 0 on C C C , so φ 0 ≤ 1 E L \varphi_{0}\le\mathbf{1}_{E_{L}} φ 0 ≤ 1 E L . Each φ i \varphi_{i} φ i is Borel by claims 3(a) and 5 of The Borel Sigma-Algebra of a Euclidean Space as a Product, and Measurability of Projections, Sequentially Continuous Maps, and Open and Closed Sets .
(P) If i ∈ [ k ] i\in[k] i ∈ [ k ] and φ i ( x ) φ i ( y ) ≠ 0 \varphi_{i}(x)\varphi_{i}(y)\neq0 φ i ( x ) φ i ( y ) = 0 , then x , y ∈ D i ⊆ B ( z i , η / 2 ) x,y\in D_{i}\subseteq B(z_{i},\eta/2) x , y ∈ D i ⊆ B ( z i , η /2 ) , so ∥ x − y ∥ ≤ ∥ x − z i ∥ + ∥ z i − y ∥ < η \lVert x-y\rVert\le\lVert x-z_{i}\rVert+\lVert z_{i}-y\rVert<\eta ∥ x − y ∥ ≤ ∥ x − z i ∥ + ∥ z i − y ∥ < η by claims 2, 5 and 6 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n , and ∥ x − y ∥ 2 ≤ η 2 = ε \lVert x-y\rVert^{2}\le\eta^{2}=\varepsilon ∥ x − y ∥ 2 ≤ η 2 = ε by claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field .
Step 4 (Masses of the cells; choice of N N N ). For n ∈ N n\in\mathbb{N} n ∈ N and i ∈ { 0 , 1 , … , k } i\in\{0,1,\dots,k\} i ∈ { 0 , 1 , … , k } put a i n = ∫ φ i d μ n a_{i}^{n}=\int\varphi_{i}\,d\mu_{n} a i n = ∫ φ i d μ n and a i = ∫ φ i d μ a_{i}=\int\varphi_{i}\,d\mu a i = ∫ φ i d μ , numbers in [ 0 , 1 ] [0,1] [ 0 , 1 ] ; by claim 2 of Linearity and Monotonicity of the Lebesgue Integral , ∑ i = 0 k a i n = 1 = ∑ i = 0 k a i \sum_{i=0}^{k}a_{i}^{n}=1=\sum_{i=0}^{k}a_{i} ∑ i = 0 k a i n = 1 = ∑ i = 0 k a i . For i ∈ [ k ] i\in[k] i ∈ [ k ] let c i n = min { a i n , a i } c_{i}^{n}=\min\{a_{i}^{n},a_{i}\} c i n = min { a i n , a i } and let δ n = 1 − ∑ i = 1 k c i n \delta_{n}=1-\sum_{i=1}^{k}c_{i}^{n} δ n = 1 − ∑ i = 1 k c i n ; since c i n ≤ a i n c_{i}^{n}\le a_{i}^{n} c i n ≤ a i n , δ n ≥ a 0 n ≥ 0 \delta_{n}\ge a_{0}^{n}\ge0 δ n ≥ a 0 n ≥ 0 . By weak convergence ( a i n ) n (a_{i}^{n})_{n} ( a i n ) n converges to a i a_{i} a i for every i i i ; since ∣ c i n − a i ∣ ≤ ∣ a i n − a i ∣ |c_{i}^{n}-a_{i}|\le|a_{i}^{n}-a_{i}| ∣ c i n − a i ∣ ≤ ∣ a i n − a i ∣ (if a i n ≥ a i a_{i}^{n}\ge a_{i} a i n ≥ a i the left side is 0 0 0 , otherwise the two sides agree), claim 3 of Order Properties of Limits of Real Sequences gives c i n → a i c_{i}^{n}\to a_{i} c i n → a i , and claims 1 and 3 of Arithmetic of Limits of Real Sequences give δ n → 1 − ∑ i = 1 k a i = a 0 \delta_{n}\to1-\sum_{i=1}^{k}a_{i}=a_{0} δ n → 1 − ∑ i = 1 k a i = a 0 . As φ 0 ≤ 1 E L ≤ L − 2 ∥ x ∥ 2 1 E L \varphi_{0}\le\mathbf{1}_{E_{L}}\le L^{-2}\lVert x\rVert^{2}\mathbf{1}_{E_{L}} φ 0 ≤ 1 E L ≤ L − 2 ∥ x ∥ 2 1 E L , claim 1 of Linearity and Monotonicity of the Lebesgue Integral and (⋆ \star ⋆ ) give a 0 ≤ ε / L 2 a_{0}\le\varepsilon/L^{2} a 0 ≤ ε / L 2 . Choose N ∈ N N\in\mathbb{N} N ∈ N with ∣ δ n − a 0 ∣ < ε / L 2 |\delta_{n}-a_{0}|<\varepsilon/L^{2} ∣ δ n − a 0 ∣ < ε / L 2 for n ≥ N n\ge N n ≥ N (Limit of a Sequence of Real Numbers ); then L 2 δ n ≤ 2 ε L^{2}\delta_{n}\le2\varepsilon L 2 δ n ≤ 2 ε for n ≥ N n\ge N n ≥ N .
Step 5 (A coupling of μ n \mu_{n} μ n and μ \mu μ ). Fix n ∈ N n\in\mathbb{N} n ∈ N . For i ∈ [ k ] i\in[k] i ∈ [ k ] let w i = c i n / ( a i n a i ) w_{i}=c_{i}^{n}/(a_{i}^{n}a_{i}) w i = c i n / ( a i n a i ) if c i n > 0 c_{i}^{n}>0 c i n > 0 (then a i n , a i ≥ c i n > 0 a_{i}^{n},a_{i}\ge c_{i}^{n}>0 a i n , a i ≥ c i n > 0 ) and w i = 0 w_{i}=0 w i = 0 otherwise; in both cases w i ≥ 0 w_{i}\ge0 w i ≥ 0 , w i a i n a i = c i n w_{i}a_{i}^{n}a_{i}=c_{i}^{n} w i a i n a i = c i n , 0 ≤ w i a i ≤ 1 0\le w_{i}a_{i}\le1 0 ≤ w i a i ≤ 1 and 0 ≤ w i a i n ≤ 1 0\le w_{i}a_{i}^{n}\le1 0 ≤ w i a i n ≤ 1 . Let e = 1 / δ n e=1/\delta_{n} e = 1/ δ n if δ n > 0 \delta_{n}>0 δ n > 0 and e = 0 e=0 e = 0 if δ n = 0 \delta_{n}=0 δ n = 0 , so that e δ n ∈ { 0 , 1 } e\delta_{n}\in\{0,1\} e δ n ∈ { 0 , 1 } and e δ n = 1 e\delta_{n}=1 e δ n = 1 when δ n > 0 \delta_{n}>0 δ n > 0 . Define
g = φ 0 + ∑ i = 1 k ( 1 − w i a i ) φ i , h = φ 0 + ∑ i = 1 k ( 1 − w i a i n ) φ i g=\varphi_{0}+\sum_{i=1}^{k}(1-w_{i}a_{i})\varphi_{i},\qquad h=\varphi_{0}+\sum_{i=1}^{k}(1-w_{i}a_{i}^{n})\varphi_{i} g = φ 0 + i = 1 ∑ k ( 1 − w i a i ) φ i , h = φ 0 + i = 1 ∑ k ( 1 − w i a i n ) φ i
on R m \mathbb{R}^{m} R m . They are Borel with 0 ≤ g ≤ 1 0\le g\le1 0 ≤ g ≤ 1 and 0 ≤ h ≤ 1 0\le h\le1 0 ≤ h ≤ 1 (as φ 0 + ∑ i φ i = 1 \varphi_{0}+\sum_{i}\varphi_{i}=1 φ 0 + ∑ i φ i = 1 ), and
1 − g = ∑ i = 1 k w i a i φ i , 1 − h = ∑ i = 1 k w i a i n φ i , ∫ g d μ n = δ n = ∫ h d μ , (E1) 1-g=\sum_{i=1}^{k}w_{i}a_{i}\varphi_{i},\qquad1-h=\sum_{i=1}^{k}w_{i}a_{i}^{n}\varphi_{i},\qquad\int g\,d\mu_{n}=\delta_{n}=\int h\,d\mu,\tag{E1} 1 − g = i = 1 ∑ k w i a i φ i , 1 − h = i = 1 ∑ k w i a i n φ i , ∫ g d μ n = δ n = ∫ h d μ , ( E1 )
the integrals by claim 2 of Linearity and Monotonicity of the Lebesgue Integral : ∫ g d μ n = a 0 n + ∑ i ( a i n − c i n ) = 1 − ∑ i c i n \int g\,d\mu_{n}=a_{0}^{n}+\sum_{i}(a_{i}^{n}-c_{i}^{n})=1-\sum_{i}c_{i}^{n} ∫ g d μ n = a 0 n + ∑ i ( a i n − c i n ) = 1 − ∑ i c i n , and likewise for h h h . Define G : R m + m → [ 0 , ∞ ) G:\mathbb{R}^{m+m}\to[0,\infty) G : R m + m → [ 0 , ∞ ) by
G ( z ) = ∑ i = 1 k w i φ i ( p r 1 ( z ) ) φ i ( p r 2 ( z ) ) + e g ( p r 1 ( z ) ) h ( p r 2 ( z ) ) , G(z)=\sum_{i=1}^{k}w_{i}\,\varphi_{i}(\mathrm{pr}_{1}(z))\,\varphi_{i}(\mathrm{pr}_{2}(z))+e\,g(\mathrm{pr}_{1}(z))\,h(\mathrm{pr}_{2}(z)), G ( z ) = i = 1 ∑ k w i φ i ( pr 1 ( z )) φ i ( pr 2 ( z )) + e g ( pr 1 ( z )) h ( pr 2 ( z )) ,
Borel because the projections are Borel (Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections ), compositions of Borel maps are Borel (Probability Measures on Euclidean Space and Random Vectors: Standing Notation §borel-maps ), and by claims 2 and 3 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions . Let γ \gamma γ be the measure on B ( R m + m ) \mathcal{B}(\mathbb{R}^{m+m}) B ( R m + m ) with density G G G with respect to μ n ⊠ μ \mu_{n}\boxtimes\mu μ n ⊠ μ , given by claim 3 of Image Measures, Measures with Densities, and Change of Variables . For every Borel F : R m + m → [ 0 , ∞ ] F:\mathbb{R}^{m+m}\to[0,\infty] F : R m + m → [ 0 , ∞ ] , that claim, Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §product and the Tonelli part of Tonelli and Fubini Theorems give
∫ F d γ = ∫ R m ( ∫ R m F ( ι ( x , y ) ) G ( ι ( x , y ) ) μ ( d y ) ) μ n ( d x ) , (E2) \int F\,d\gamma=\int_{\mathbb{R}^{m}}\Bigl(\int_{\mathbb{R}^{m}}F(\iota(x,y))\,G(\iota(x,y))\,\mu(dy)\Bigr)\mu_{n}(dx),\tag{E2} ∫ F d γ = ∫ R m ( ∫ R m F ( ι ( x , y )) G ( ι ( x , y )) μ ( d y ) ) μ n ( d x ) , ( E2 )
and the same with the order of integration reversed; here G ( ι ( x , y ) ) = ∑ i w i φ i ( x ) φ i ( y ) + e g ( x ) h ( y ) G(\iota(x,y))=\sum_{i}w_{i}\varphi_{i}(x)\varphi_{i}(y)+e\,g(x)h(y) G ( ι ( x , y )) = ∑ i w i φ i ( x ) φ i ( y ) + e g ( x ) h ( y ) .
Marginals. Let A ∈ B ( R m ) A\in\mathcal{B}(\mathbb{R}^{m}) A ∈ B ( R m ) and F = 1 A ∘ p r 1 = 1 p r 1 − 1 ( A ) F=\mathbf{1}_{A}\circ\mathrm{pr}_{1}=\mathbf{1}_{\mathrm{pr}_{1}^{-1}(A)} F = 1 A ∘ pr 1 = 1 pr 1 − 1 ( A ) . By claim 1 of Linearity and Monotonicity of the Lebesgue Integral and (E1), the inner integral in (E2) is 1 A ( x ) ( ∑ i w i a i φ i ( x ) + e δ n g ( x ) ) = 1 A ( x ) ( 1 − g ( x ) + e δ n g ( x ) ) \mathbf{1}_{A}(x)\bigl(\sum_{i}w_{i}a_{i}\varphi_{i}(x)+e\delta_{n}g(x)\bigr)=\mathbf{1}_{A}(x)\bigl(1-g(x)+e\delta_{n}g(x)\bigr) 1 A ( x ) ( ∑ i w i a i φ i ( x ) + e δ n g ( x ) ) = 1 A ( x ) ( 1 − g ( x ) + e δ n g ( x ) ) . If δ n > 0 \delta_{n}>0 δ n > 0 this is 1 A ( x ) \mathbf{1}_{A}(x) 1 A ( x ) , and γ ( p r 1 − 1 ( A ) ) = μ n ( A ) \gamma(\mathrm{pr}_{1}^{-1}(A))=\mu_{n}(A) γ ( pr 1 − 1 ( A )) = μ n ( A ) . If δ n = 0 \delta_{n}=0 δ n = 0 it is 1 A ( 1 − g ) \mathbf{1}_{A}(1-g) 1 A ( 1 − g ) , whose integral is μ n ( A ) − ∫ 1 A g d μ n \mu_{n}(A)-\int\mathbf{1}_{A}g\,d\mu_{n} μ n ( A ) − ∫ 1 A g d μ n by claim 2 of Linearity and Monotonicity of the Lebesgue Integral , and 0 ≤ ∫ 1 A g d μ n ≤ ∫ g d μ n = δ n = 0 0\le\int\mathbf{1}_{A}g\,d\mu_{n}\le\int g\,d\mu_{n}=\delta_{n}=0 0 ≤ ∫ 1 A g d μ n ≤ ∫ g d μ n = δ n = 0 ; again γ ( p r 1 − 1 ( A ) ) = μ n ( A ) \gamma(\mathrm{pr}_{1}^{-1}(A))=\mu_{n}(A) γ ( pr 1 − 1 ( A )) = μ n ( A ) . The same computation in the reversed order, with h h h , μ \mu μ and the second identity of (E1), gives γ ( p r 2 − 1 ( A ) ) = μ ( A ) \gamma(\mathrm{pr}_{2}^{-1}(A))=\mu(A) γ ( pr 2 − 1 ( A )) = μ ( A ) . Taking A = R m A=\mathbb{R}^{m} A = R m shows γ ( R m + m ) = 1 \gamma(\mathbb{R}^{m+m})=1 γ ( R m + m ) = 1 , so γ ∈ P ( R m + m ) \gamma\in\mathcal{P}(\mathbb{R}^{m+m}) γ ∈ P ( R m + m ) and γ ∈ Π ( μ n , μ ) \gamma\in\Pi(\mu_{n},\mu) γ ∈ Π ( μ n , μ ) by Couplings of Two Probability Measures on Euclidean Space and Their Quadratic Cost §coupling .
Step 6 (The cost of γ \gamma γ ). By Couplings of Two Probability Measures on Euclidean Space and Their Quadratic Cost §cost and (E2), I ( γ ) = ∫ ∫ ∥ x − y ∥ 2 G ( ι ( x , y ) ) μ ( d y ) μ n ( d x ) I(\gamma)=\int\!\int\lVert x-y\rVert^{2}G(\iota(x,y))\,\mu(dy)\,\mu_{n}(dx) I ( γ ) = ∫ ∫ ∥ x − y ∥ 2 G ( ι ( x , y )) μ ( d y ) μ n ( d x ) , using p r 1 ( ι ( x , y ) ) = x \mathrm{pr}_{1}(\iota(x,y))=x pr 1 ( ι ( x , y )) = x , p r 2 ( ι ( x , y ) ) = y \mathrm{pr}_{2}(\iota(x,y))=y pr 2 ( ι ( x , y )) = y . By (P) and the inequality ∥ x − y ∥ 2 ≤ 2 ∥ x ∥ 2 + 2 ∥ y ∥ 2 \lVert x-y\rVert^{2}\le2\lVert x\rVert^{2}+2\lVert y\rVert^{2} ∥ x − y ∥ 2 ≤ 2 ∥ x ∥ 2 + 2 ∥ y ∥ 2 of Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions ,
∥ x − y ∥ 2 G ( ι ( x , y ) ) ≤ ε ∑ i = 1 k w i φ i ( x ) φ i ( y ) + 2 e ( ∥ x ∥ 2 + ∥ y ∥ 2 ) g ( x ) h ( y ) . \lVert x-y\rVert^{2}G(\iota(x,y))\le\varepsilon\sum_{i=1}^{k}w_{i}\varphi_{i}(x)\varphi_{i}(y)+2e\bigl(\lVert x\rVert^{2}+\lVert y\rVert^{2}\bigr)g(x)h(y). ∥ x − y ∥ 2 G ( ι ( x , y )) ≤ ε i = 1 ∑ k w i φ i ( x ) φ i ( y ) + 2 e ( ∥ x ∥ 2 + ∥ y ∥ 2 ) g ( x ) h ( y ) .
Integrating first in y y y and then in x x x with claim 1 of Linearity and Monotonicity of the Lebesgue Integral and (E1),
I ( γ ) ≤ ε ∑ i = 1 k w i a i n a i + 2 e δ n ( A n + B n ) ≤ ε + 2 ( A n + B n ) , I(\gamma)\le\varepsilon\sum_{i=1}^{k}w_{i}a_{i}^{n}a_{i}+2e\delta_{n}(A_{n}+B_{n})\le\varepsilon+2(A_{n}+B_{n}), I ( γ ) ≤ ε i = 1 ∑ k w i a i n a i + 2 e δ n ( A n + B n ) ≤ ε + 2 ( A n + B n ) ,
where A n = ∫ ∥ x ∥ 2 g ( x ) μ n ( d x ) A_{n}=\int\lVert x\rVert^{2}g(x)\,\mu_{n}(dx) A n = ∫ ∥ x ∥ 2 g ( x ) μ n ( d x ) and B n = ∫ ∥ y ∥ 2 h ( y ) μ ( d y ) B_{n}=\int\lVert y\rVert^{2}h(y)\,\mu(dy) B n = ∫ ∥ y ∥ 2 h ( y ) μ ( d y ) in [ 0 , ∞ ] [0,\infty] [ 0 , ∞ ] , and we used ∑ i w i a i n a i = ∑ i c i n ≤ ∑ i a i n ≤ 1 \sum_{i}w_{i}a_{i}^{n}a_{i}=\sum_{i}c_{i}^{n}\le\sum_{i}a_{i}^{n}\le1 ∑ i w i a i n a i = ∑ i c i n ≤ ∑ i a i n ≤ 1 and e δ n ≤ 1 e\delta_{n}\le1 e δ n ≤ 1 . Since 0 ≤ g ≤ 1 0\le g\le1 0 ≤ g ≤ 1 , ∥ x ∥ 2 g ( x ) ≤ L 2 g ( x ) + ∥ x ∥ 2 1 E L ( x ) \lVert x\rVert^{2}g(x)\le L^{2}g(x)+\lVert x\rVert^{2}\mathbf{1}_{E_{L}}(x) ∥ x ∥ 2 g ( x ) ≤ L 2 g ( x ) + ∥ x ∥ 2 1 E L ( x ) for every x x x (claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field when ∥ x ∥ ≤ L \lVert x\rVert\le L ∥ x ∥ ≤ L ); so A n ≤ L 2 δ n + ε A_{n}\le L^{2}\delta_{n}+\varepsilon A n ≤ L 2 δ n + ε by (E1) and (⋆ \star ⋆ ), and in the same way B n ≤ L 2 δ n + ε B_{n}\le L^{2}\delta_{n}+\varepsilon B n ≤ L 2 δ n + ε . By The Quadratic Wasserstein Distance on Euclidean Space §distance , for n ≥ N n\ge N n ≥ N (Step 4),
W 2 ( μ n , μ ) 2 ≤ I ( γ ) ≤ ε + 4 ( L 2 δ n + ε ) ≤ 13 ε . W_{2}(\mu_{n},\mu)^{2}\le I(\gamma)\le\varepsilon+4(L^{2}\delta_{n}+\varepsilon)\le13\,\varepsilon . W 2 ( μ n , μ ) 2 ≤ I ( γ ) ≤ ε + 4 ( L 2 δ n + ε ) ≤ 13 ε .
Step 7 (Conclusion of Part 1). Let ε ′ > 0 \varepsilon'>0 ε ′ > 0 and run Steps 3 to 6 with ε = ε ′ 2 / 26 \varepsilon=\varepsilon'^{2}/26 ε = ε ′ 2 /26 (the choices being made in the order ε \varepsilon ε , L L L , η \eta η , the cover and φ 0 , … , φ k \varphi_{0},\dots,\varphi_{k} φ 0 , … , φ k , then N N N , and the coupling for each n ≥ N n\ge N n ≥ N ). For n ≥ N n\ge N n ≥ N , W 2 ( μ n , μ ) 2 ≤ ε ′ 2 / 2 < ε ′ 2 W_{2}(\mu_{n},\mu)^{2}\le\varepsilon'^{2}/2<\varepsilon'^{2} W 2 ( μ n , μ ) 2 ≤ ε ′ 2 /2 < ε ′ 2 , so 0 ≤ W 2 ( μ n , μ ) < ε ′ 0\le W_{2}(\mu_{n},\mu)<\varepsilon' 0 ≤ W 2 ( μ n , μ ) < ε ′ by claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field . Hence ( W 2 ( μ n , μ ) ) n (W_{2}(\mu_{n},\mu))_{n} ( W 2 ( μ n , μ ) ) n has limit 0 0 0 ; with Step 2 this proves clause 1.
Part 2 (clause 2). Let h h h , c c c and ( μ n ) n (\mu_{n})_{n} ( μ n ) n be as in clause 2, let b b b be a real number with b ≤ h ( x ) b\le h(x) b ≤ h ( x ) for every x x x , and put b ′ = min { b , 0 } b'=\min\{b,0\} b ′ = min { b , 0 } and C 0 = c − b ′ C_{0}=c-b' C 0 = c − b ′ . For real M > 0 M>0 M > 0 let K M > 0 K_{M}>0 K M > 0 be given by superquadraticity and F M = { x : K M ≤ ∥ x ∥ } F_{M}=\{x:K_{M}\le\lVert x\rVert\} F M = { x : K M ≤ ∥ x ∥} , which is Borel as the complement of the preimage of the open set { t : t < K M } \{t:t<K_{M}\} { t : t < K M } under the Borel norm. Fix M M M and n n n . The functions h 1 F M h\mathbf{1}_{F_{M}} h 1 F M and h 1 R m ∖ F M h\mathbf{1}_{\mathbb{R}^{m}\setminus F_{M}} h 1 R m ∖ F M are Borel (claims 1 and 3 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions ) with absolute values at most ∣ h ∣ |h| ∣ h ∣ , hence μ n \mu_{n} μ n -integrable by Integrable Function and the Lebesgue Integral and claim 1 of Linearity and Monotonicity of the Lebesgue Integral ; their sum is h h h , and h 1 R m ∖ F M ≥ b ′ h\mathbf{1}_{\mathbb{R}^{m}\setminus F_{M}}\ge b' h 1 R m ∖ F M ≥ b ′ pointwise because b ′ ≤ 0 b'\le0 b ′ ≤ 0 and b ′ ≤ b b'\le b b ′ ≤ b . By claim 2 of Linearity and Monotonicity of the Lebesgue Integral ,
∫ h 1 F M d μ n = ∫ h d μ n − ∫ h 1 R m ∖ F M d μ n ≤ c − b ′ = C 0 . \int h\mathbf{1}_{F_{M}}\,d\mu_{n}=\int h\,d\mu_{n}-\int h\mathbf{1}_{\mathbb{R}^{m}\setminus F_{M}}\,d\mu_{n}\le c-b'=C_{0}. ∫ h 1 F M d μ n = ∫ h d μ n − ∫ h 1 R m ∖ F M d μ n ≤ c − b ′ = C 0 .
On F M F_{M} F M , 0 ≤ M ∥ x ∥ 2 ≤ h ( x ) 0\le M\lVert x\rVert^{2}\le h(x) 0 ≤ M ∥ x ∥ 2 ≤ h ( x ) , so 0 ≤ M ∥ x ∥ 2 1 F M ≤ h 1 F M 0\le M\lVert x\rVert^{2}\mathbf{1}_{F_{M}}\le h\mathbf{1}_{F_{M}} 0 ≤ M ∥ x ∥ 2 1 F M ≤ h 1 F M pointwise; the latter is nonnegative, so its integral in [ 0 , ∞ ] [0,\infty] [ 0 , ∞ ] is its Lebesgue integral (Integrable Function and the Lebesgue Integral ), and claim 1 of Linearity and Monotonicity of the Lebesgue Integral gives
M ∫ R m ∥ x ∥ 2 1 F M ( x ) μ n ( d x ) ≤ C 0 ( n ∈ N ) , (E3) M\int_{\mathbb{R}^{m}}\lVert x\rVert^{2}\mathbf{1}_{F_{M}}(x)\,\mu_{n}(dx)\le C_{0}\qquad(n\in\mathbb{N}),\tag{E3} M ∫ R m ∥ x ∥ 2 1 F M ( x ) μ n ( d x ) ≤ C 0 ( n ∈ N ) , ( E3 )
and in particular C 0 ≥ 0 C_{0}\ge0 C 0 ≥ 0 .
(i) With M = 1 M=1 M = 1 : ∥ x ∥ 2 ≤ K 1 2 + ∥ x ∥ 2 1 F 1 ( x ) \lVert x\rVert^{2}\le K_{1}^{2}+\lVert x\rVert^{2}\mathbf{1}_{F_{1}}(x) ∥ x ∥ 2 ≤ K 1 2 + ∥ x ∥ 2 1 F 1 ( x ) for every x x x , so M 2 ( μ n ) ≤ K 1 2 + C 0 M_{2}(\mu_{n})\le K_{1}^{2}+C_{0} M 2 ( μ n ) ≤ K 1 2 + C 0 for every n n n . By Tightness from Bounded Second Moments, and Tightness of the Couplings of Two Measures with Finite Second Moment §moment the set { μ n : n ∈ N } \{\mu_{n}:n\in\mathbb{N}\} { μ n : n ∈ N } is tight in ( R m , d E ) (\mathbb{R}^{m},d_{E}) ( R m , d E ) , that is, the sequence is tight in the sense of Tight Family of Borel Measures on a Metric Space §sequence . By Prokhorov's Theorem on Euclidean Space: a Tight Sequence of Probability Measures Has a Weakly Convergent Subsequence there are a strictly increasing sequence ( n j ) j ∈ N (n_{j})_{j\in\mathbb{N}} ( n j ) j ∈ N in N \mathbb{N} N and μ ∈ P ( R m ) \mu\in\mathcal{P}(\mathbb{R}^{m}) μ ∈ P ( R m ) such that ( μ n j ) j (\mu_{n_{j}})_{j} ( μ n j ) j converges weakly to μ \mu μ .
(ii) Let ε > 0 \varepsilon>0 ε > 0 , put M = ( C 0 + 1 ) / ε > 0 M=(C_{0}+1)/\varepsilon>0 M = ( C 0 + 1 ) / ε > 0 and K = K M K=K_{M} K = K M . Since E K ⊆ F M E_{K}\subseteq F_{M} E K ⊆ F M , (E3) gives ∫ E K ∥ x ∥ 2 μ n j ( d x ) ≤ C 0 / M = C 0 ε / ( C 0 + 1 ) < ε \int_{E_{K}}\lVert x\rVert^{2}\,\mu_{n_{j}}(dx)\le C_{0}/M=C_{0}\varepsilon/(C_{0}+1)<\varepsilon ∫ E K ∥ x ∥ 2 μ n j ( d x ) ≤ C 0 / M = C 0 ε / ( C 0 + 1 ) < ε for every j j j .
Thus the sequence ( μ n j ) j ∈ N (\mu_{n_{j}})_{j\in\mathbb{N}} ( μ n j ) j ∈ N in P 2 ( R m ) \mathcal{P}_{2}(\mathbb{R}^{m}) P 2 ( R m ) and μ \mu μ satisfy the hypotheses of clause 1, and Part 1 gives μ ∈ P 2 ( R m ) \mu\in\mathcal{P}_{2}(\mathbb{R}^{m}) μ ∈ P 2 ( R m ) and that ( W 2 ( μ n j , μ ) ) j ∈ N (W_{2}(\mu_{n_{j}},\mu))_{j\in\mathbb{N}} ( W 2 ( μ n j , μ ) ) j ∈ N has limit 0 0 0 . With ( n j ) j (n_{j})_{j} ( n j ) j as the strictly increasing sequence named ( n k ) k (n_{k})_{k} ( n k ) k in clause 2, this proves clause 2.