Each result cited below is universally quantified over the data in its own statement and is applied with the data named at the point of citation. "The integral theorem" refers to Linearity and Monotonicity of the Lebesgue Integral (claim 1: additivity, homogeneity with a factor in [ 0 , ∞ ) [0,\infty) [ 0 , ∞ ) and monotonicity of the nonnegative integral, additivity extending to finite sums by induction on the number of terms; claim 2: linearity, the bound ∣ ∫ f ∣ ≤ ∫ ∣ f ∣ |\int f|\le\int|f| ∣ ∫ f ∣ ≤ ∫ ∣ f ∣ and monotonicity for integrable functions). "Change of variables" refers to claim 2 of Image Measures, Measures with Densities, and Change of Variables , in force for push-forwards by Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §pushforward , and "the density formula" to claim 3 of the same lemma. Throughout, π 1 , π 2 : X × X → X \pi_{1},\pi_{2}:X\times X\to X π 1 , π 2 : X × X → X are Borel and B ( X × X ) = B ( X ) ⊗ B ( X ) \mathcal{B}(X\times X)=\mathcal{B}(X)\otimes\mathcal{B}(X) B ( X × X ) = B ( X ) ⊗ B ( X ) , by Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §product-sigma ; composites of measurable maps are measurable by claim 4 of Borel Measurability and Bounded Integration on a Metric Space . Four standing facts are used repeatedly.
(a) Norm facts. For x , y ∈ X x,y\in X x , y ∈ X one has ∣ x − y ∣ = d ( x , y ) = d ( y , x ) = ∣ y − x ∣ |x-y|=d(x,y)=d(y,x)=|y-x| ∣ x − y ∣ = d ( x , y ) = d ( y , x ) = ∣ y − x ∣ , by Real Inner Product Space §distance and the symmetry of the metric d d d (The Norm Metric of a Real Inner Product Space: Triangle Inequalities, Limits and Continuity §metric ); and ∣ x + y ∣ ≤ ∣ x ∣ + ∣ y ∣ |x+y|\le|x|+|y| ∣ x + y ∣ ≤ ∣ x ∣ + ∣ y ∣ and ∣ x − y ∣ ≤ ∣ x ∣ + ∣ y ∣ |x-y|\le|x|+|y| ∣ x − y ∣ ≤ ∣ x ∣ + ∣ y ∣ by The Norm Metric of a Real Inner Product Space: Triangle Inequalities, Limits and Continuity §triangle . For real r , s , t ≥ 0 r,s,t\ge0 r , s , t ≥ 0 with r ≤ s + t r\le s+t r ≤ s + t one has r 2 ≤ ( s + t ) 2 r^{2}\le(s+t)^{2} r 2 ≤ ( s + t ) 2 by claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field , and ( s + t ) 2 ≤ 2 s 2 + 2 t 2 (s+t)^{2}\le2s^{2}+2t^{2} ( s + t ) 2 ≤ 2 s 2 + 2 t 2 since 2 s 2 + 2 t 2 − ( s + t ) 2 = ( s − t ) 2 ≥ 0 2s^{2}+2t^{2}-(s+t)^{2}=(s-t)^{2}\ge0 2 s 2 + 2 t 2 − ( s + t ) 2 = ( s − t ) 2 ≥ 0 .
(b) Borel functions. The map x ↦ ∣ x ∣ x\mapsto|x| x ↦ ∣ x ∣ on X X X is Lipschitz with constant 1 1 1 by The Norm Metric of a Real Inner Product Space: Triangle Inequalities, Limits and Continuity §lipschitz , hence continuous by A Lipschitz Map is Uniformly Continuous and Borel by claims 2 and 3 of Borel Measurability and Bounded Integration on a Metric Space ; the map x ↦ ∣ x ∣ 2 x\mapsto|x|^{2} x ↦ ∣ x ∣ 2 is Borel as recorded in The Second Moment of a Borel Probability Measure on a Hilbert Space and the Probability Measures with Finite Second Moment . Let φ 0 : X × X → R \varphi_{0}:X\times X\to\mathbb{R} φ 0 : X × X → R be φ 0 ( z ) = ∣ π 1 ( z ) − π 2 ( z ) ∣ \varphi_{0}(z)=|\pi_{1}(z)-\pi_{2}(z)| φ 0 ( z ) = ∣ π 1 ( z ) − π 2 ( z ) ∣ and φ = φ 0 2 \varphi=\varphi_{0}^{2} φ = φ 0 2 . The map z ↦ π 1 ( z ) − π 2 ( z ) z\mapsto\pi_{1}(z)-\pi_{2}(z) z ↦ π 1 ( z ) − π 2 ( z ) is Lipschitz with constant 2 2 2 , as recorded in Couplings of Two Borel Probability Measures on a Hilbert Space and Their Quadratic Cost §cost , so its composite φ 0 \varphi_{0} φ 0 with the norm is Lipschitz with constant 2 2 2 , hence continuous and Borel by the same three references; φ \varphi φ is Borel by claim 3 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions , and I ( π ) = ∫ φ d π I(\pi)=\int\varphi\,d\pi I ( π ) = ∫ φ d π for every coupling π \pi π by Couplings of Two Borel Probability Measures on a Hilbert Space and Their Quadratic Cost §cost .
(c) Pairings and push-forwards. For maps S , T S,T S , T from a set Ω \Omega Ω into X X X , π 1 ∘ ( S , T ) = S \pi_{1}\circ(S,T)=S π 1 ∘ ( S , T ) = S and π 2 ∘ ( S , T ) = T \pi_{2}\circ(S,T)=T π 2 ∘ ( S , T ) = T by Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §pairs and The Product of Two Real Inner Product Spaces §notation ; hence φ ( ( S , T ) ( ω ) ) = ∣ S ( ω ) − T ( ω ) ∣ 2 \varphi((S,T)(\omega))=|S(\omega)-T(\omega)|^{2} φ (( S , T ) ( ω )) = ∣ S ( ω ) − T ( ω ) ∣ 2 and φ 0 ( ( S , T ) ( ω ) ) = ∣ S ( ω ) − T ( ω ) ∣ \varphi_{0}((S,T)(\omega))=|S(\omega)-T(\omega)| φ 0 (( S , T ) ( ω )) = ∣ S ( ω ) − T ( ω ) ∣ . For measurable maps G G G from a measure space ( Ω 1 , F 1 , τ ) (\Omega_{1},\mathcal{F}_{1},\tau) ( Ω 1 , F 1 , τ ) into a measurable space ( Ω 2 , F 2 ) (\Omega_{2},\mathcal{F}_{2}) ( Ω 2 , F 2 ) and H H H from ( Ω 2 , F 2 ) (\Omega_{2},\mathcal{F}_{2}) ( Ω 2 , F 2 ) into a measurable space ( Ω 3 , F 3 ) (\Omega_{3},\mathcal{F}_{3}) ( Ω 3 , F 3 ) , the image measures of claim 1 of Image Measures, Measures with Densities, and Change of Variables satisfy ( H ∘ G ) # τ = H # ( G # τ ) (H\circ G)_{\#}\tau=H_{\#}(G_{\#}\tau) ( H ∘ G ) # τ = H # ( G # τ ) , both assigning to B ∈ F 3 B\in\mathcal{F}_{3} B ∈ F 3 the value τ ( G − 1 ( H − 1 ( B ) ) ) \tau(G^{-1}(H^{-1}(B))) τ ( G − 1 ( H − 1 ( B ))) ; and G # τ G_{\#}\tau G # τ has total mass τ ( Ω 1 ) \tau(\Omega_{1}) τ ( Ω 1 ) by the same claim, so push-forwards of members of P ( X ) \mathcal{P}(X) P ( X ) or P ( X × X ) \mathcal{P}(X\times X) P ( X × X ) by Borel maps into X X X or X × X X\times X X × X lie in P ( X ) \mathcal{P}(X) P ( X ) , respectively P ( X × X ) \mathcal{P}(X\times X) P ( X × X ) .
(d) Integrals and mean squares. On a measure space, a nonnegative measurable real function with finite integral is integrable with the same integral, its negative part being 0 0 0 and its positive part itself; this identifies the two readings of such integrals below. On a probability space, square-integrable random variables and the norm ∥ ⋅ ∥ 2 \lVert\cdot\rVert_{2} ∥ ⋅ ∥ 2 are those of Square-Integrable Random Variables and the Mean-Square Inner Product ; a measurable real function f f f with ∫ f 2 < ∞ \int f^{2}<\infty ∫ f 2 < ∞ is square-integrable with ∥ f ∥ 2 = ∫ f 2 \lVert f\rVert_{2}=\sqrt{\int f^{2}} ∥ f ∥ 2 = ∫ f 2 , the expectation of the nonnegative random variable f 2 f^{2} f 2 being its integral by Expectation, Variance, and Moments ; and sums of square-integrable random variables are square-integrable by Square-Integrable Random Variables and the Mean-Square Inner Product .
Claim 1. Members of P ( X ) \mathcal{P}(X) P ( X ) have finite total mass. By Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §product-measure , μ ⊗ ν \mu\otimes\nu μ ⊗ ν is a Borel measure on X × X X\times X X × X with ( π 1 ) # ( μ ⊗ ν ) = ν ( X ) μ = μ (\pi_{1})_{\#}(\mu\otimes\nu)=\nu(X)\,\mu=\mu ( π 1 ) # ( μ ⊗ ν ) = ν ( X ) μ = μ and ( π 2 ) # ( μ ⊗ ν ) = μ ( X ) ν = ν (\pi_{2})_{\#}(\mu\otimes\nu)=\mu(X)\,\nu=\nu ( π 2 ) # ( μ ⊗ ν ) = μ ( X ) ν = ν . Its total mass is ( μ ⊗ ν ) ( X × X ) = μ ( X ) ν ( X ) = 1 (\mu\otimes\nu)(X\times X)=\mu(X)\,\nu(X)=1 ( μ ⊗ ν ) ( X × X ) = μ ( X ) ν ( X ) = 1 by Existence and Uniqueness of the Product Measure , X × X X\times X X × X being a measurable rectangle. Hence μ ⊗ ν ∈ P ( X × X ) \mu\otimes\nu\in\mathcal{P}(X\times X) μ ⊗ ν ∈ P ( X × X ) , and μ ⊗ ν ∈ Π ( μ , ν ) \mu\otimes\nu\in\Pi(\mu,\nu) μ ⊗ ν ∈ Π ( μ , ν ) by Couplings of Two Borel Probability Measures on a Hilbert Space and Their Quadratic Cost §coupling .
Claim 2. The swap σ = ( π 2 , π 1 ) \sigma=(\pi_{2},\pi_{1}) σ = ( π 2 , π 1 ) is Borel, as recorded in the statement, so σ # π ∈ P ( X × X ) \sigma_{\#}\pi\in\mathcal{P}(X\times X) σ # π ∈ P ( X × X ) by (c). By (c), π 1 ∘ σ = π 2 \pi_{1}\circ\sigma=\pi_{2} π 1 ∘ σ = π 2 and π 2 ∘ σ = π 1 \pi_{2}\circ\sigma=\pi_{1} π 2 ∘ σ = π 1 , and σ ( σ ( z ) ) = ( π 2 ( σ ( z ) ) , π 1 ( σ ( z ) ) ) = ( π 1 ( z ) , π 2 ( z ) ) = z \sigma(\sigma(z))=(\pi_{2}(\sigma(z)),\pi_{1}(\sigma(z)))=(\pi_{1}(z),\pi_{2}(z))=z σ ( σ ( z )) = ( π 2 ( σ ( z )) , π 1 ( σ ( z ))) = ( π 1 ( z ) , π 2 ( z )) = z , a pair being determined by its coordinates (The Product of Two Real Inner Product Spaces §notation ). Hence by (c), ( π 1 ) # ( σ # π ) = ( π 1 ∘ σ ) # π = ( π 2 ) # π = ν (\pi_{1})_{\#}(\sigma_{\#}\pi)=(\pi_{1}\circ\sigma)_{\#}\pi=(\pi_{2})_{\#}\pi=\nu ( π 1 ) # ( σ # π ) = ( π 1 ∘ σ ) # π = ( π 2 ) # π = ν and ( π 2 ) # ( σ # π ) = ( π 1 ) # π = μ (\pi_{2})_{\#}(\sigma_{\#}\pi)=(\pi_{1})_{\#}\pi=\mu ( π 2 ) # ( σ # π ) = ( π 1 ) # π = μ , so σ # π ∈ Π ( ν , μ ) \sigma_{\#}\pi\in\Pi(\nu,\mu) σ # π ∈ Π ( ν , μ ) by Couplings of Two Borel Probability Measures on a Hilbert Space and Their Quadratic Cost §coupling ; and σ # ( σ # π ) = ( σ ∘ σ ) # π = π \sigma_{\#}(\sigma_{\#}\pi)=(\sigma\circ\sigma)_{\#}\pi=\pi σ # ( σ # π ) = ( σ ∘ σ ) # π = π . By change of variables, I ( σ # π ) = ∫ φ ∘ σ d π I(\sigma_{\#}\pi)=\int\varphi\circ\sigma\,d\pi I ( σ # π ) = ∫ φ ∘ σ d π , and φ ( σ ( z ) ) = ∣ π 2 ( z ) − π 1 ( z ) ∣ 2 = φ ( z ) \varphi(\sigma(z))=|\pi_{2}(z)-\pi_{1}(z)|^{2}=\varphi(z) φ ( σ ( z )) = ∣ π 2 ( z ) − π 1 ( z ) ∣ 2 = φ ( z ) by (c) and (a); so I ( σ # π ) = I ( π ) I(\sigma_{\#}\pi)=I(\pi) I ( σ # π ) = I ( π ) .
Claim 3. Let π ∈ Π ( μ , ν ) \pi\in\Pi(\mu,\nu) π ∈ Π ( μ , ν ) . For every z ∈ X × X z\in X\times X z ∈ X × X , (a) gives φ 0 ( z ) ≤ ∣ π 1 ( z ) ∣ + ∣ π 2 ( z ) ∣ \varphi_{0}(z)\le|\pi_{1}(z)|+|\pi_{2}(z)| φ 0 ( z ) ≤ ∣ π 1 ( z ) ∣ + ∣ π 2 ( z ) ∣ and hence φ ( z ) ≤ 2 ∣ π 1 ( z ) ∣ 2 + 2 ∣ π 2 ( z ) ∣ 2 \varphi(z)\le2|\pi_{1}(z)|^{2}+2|\pi_{2}(z)|^{2} φ ( z ) ≤ 2∣ π 1 ( z ) ∣ 2 + 2∣ π 2 ( z ) ∣ 2 . The functions z ↦ ∣ π i ( z ) ∣ 2 z\mapsto|\pi_{i}(z)|^{2} z ↦ ∣ π i ( z ) ∣ 2 are Borel by (b), and by change of variables ∫ ∣ π 1 ( z ) ∣ 2 π ( d z ) = ∫ ∣ x ∣ 2 ( ( π 1 ) # π ) ( d x ) = M 2 ( μ ) \int|\pi_{1}(z)|^{2}\,\pi(dz)=\int|x|^{2}\,((\pi_{1})_{\#}\pi)(dx)=M_{2}(\mu) ∫ ∣ π 1 ( z ) ∣ 2 π ( d z ) = ∫ ∣ x ∣ 2 (( π 1 ) # π ) ( d x ) = M 2 ( μ ) and likewise ∫ ∣ π 2 ( z ) ∣ 2 π ( d z ) = M 2 ( ν ) \int|\pi_{2}(z)|^{2}\,\pi(dz)=M_{2}(\nu) ∫ ∣ π 2 ( z ) ∣ 2 π ( d z ) = M 2 ( ν ) , by The Second Moment of a Borel Probability Measure on a Hilbert Space and the Probability Measures with Finite Second Moment §moment . Claim 1 of the integral theorem gives I ( π ) ≤ 2 M 2 ( μ ) + 2 M 2 ( ν ) I(\pi)\le2M_{2}(\mu)+2M_{2}(\nu) I ( π ) ≤ 2 M 2 ( μ ) + 2 M 2 ( ν ) in [ 0 , ∞ ] [0,\infty] [ 0 , ∞ ] , which is finite when μ , ν ∈ P 2 ( X ) \mu,\nu\in\mathcal{P}_{2}(X) μ , ν ∈ P 2 ( X ) by The Second Moment of a Borel Probability Measure on a Hilbert Space and the Probability Measures with Finite Second Moment §space . Conversely, π 2 ( z ) = π 1 ( z ) − ( π 1 ( z ) − π 2 ( z ) ) \pi_{2}(z)=\pi_{1}(z)-(\pi_{1}(z)-\pi_{2}(z)) π 2 ( z ) = π 1 ( z ) − ( π 1 ( z ) − π 2 ( z )) by the axioms of the vector space X X X , so (a) gives ∣ π 2 ( z ) ∣ ≤ ∣ π 1 ( z ) ∣ + φ 0 ( z ) |\pi_{2}(z)|\le|\pi_{1}(z)|+\varphi_{0}(z) ∣ π 2 ( z ) ∣ ≤ ∣ π 1 ( z ) ∣ + φ 0 ( z ) and ∣ π 2 ( z ) ∣ 2 ≤ 2 ∣ π 1 ( z ) ∣ 2 + 2 φ ( z ) |\pi_{2}(z)|^{2}\le2|\pi_{1}(z)|^{2}+2\varphi(z) ∣ π 2 ( z ) ∣ 2 ≤ 2∣ π 1 ( z ) ∣ 2 + 2 φ ( z ) ; integrating as before, M 2 ( ν ) ≤ 2 M 2 ( μ ) + 2 I ( π ) M_{2}(\nu)\le2M_{2}(\mu)+2I(\pi) M 2 ( ν ) ≤ 2 M 2 ( μ ) + 2 I ( π ) , which is finite when μ ∈ P 2 ( X ) \mu\in\mathcal{P}_{2}(X) μ ∈ P 2 ( X ) and I ( π ) < ∞ I(\pi)<\infty I ( π ) < ∞ , so then ν ∈ P 2 ( X ) \nu\in\mathcal{P}_{2}(X) ν ∈ P 2 ( X ) .
Claim 4. The pairing ( S , T ) (S,T) ( S , T ) is Borel by Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §pairing , so ( S , T ) # μ ∈ P ( X × X ) (S,T)_{\#}\mu\in\mathcal{P}(X\times X) ( S , T ) # μ ∈ P ( X × X ) by (c), and by (c) its push-forward by π 1 \pi_{1} π 1 is ( π 1 ∘ ( S , T ) ) # μ = S # μ (\pi_{1}\circ(S,T))_{\#}\mu=S_{\#}\mu ( π 1 ∘ ( S , T ) ) # μ = S # μ and by π 2 \pi_{2} π 2 is T # μ T_{\#}\mu T # μ ; so ( S , T ) # μ ∈ Π ( S # μ , T # μ ) (S,T)_{\#}\mu\in\Pi(S_{\#}\mu,T_{\#}\mu) ( S , T ) # μ ∈ Π ( S # μ , T # μ ) by Couplings of Two Borel Probability Measures on a Hilbert Space and Their Quadratic Cost §coupling . By change of variables and (c), I ( ( S , T ) # μ ) = ∫ φ ∘ ( S , T ) d μ = ∫ ∣ S ( x ) − T ( x ) ∣ 2 μ ( d x ) I((S,T)_{\#}\mu)=\int\varphi\circ(S,T)\,d\mu=\int|S(x)-T(x)|^{2}\,\mu(dx) I (( S , T ) # μ ) = ∫ φ ∘ ( S , T ) d μ = ∫ ∣ S ( x ) − T ( x ) ∣ 2 μ ( d x ) . For S = T = i d X S=T=\mathrm{id}_{X} S = T = id X one has ( i d X ) # μ = μ (\mathrm{id}_{X})_{\#}\mu=\mu ( id X ) # μ = μ , and the integrand is ∣ x − x ∣ 2 = d ( x , x ) 2 = 0 |x-x|^{2}=d(x,x)^{2}=0 ∣ x − x ∣ 2 = d ( x , x ) 2 = 0 by The Norm Metric of a Real Inner Product Space: Triangle Inequalities, Limits and Continuity §metric ; the integral of the zero function is 0 0 0 by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §null-integral with the null set ∅ \varnothing ∅ .
Claim 5. The maps T ∘ π 2 T\circ\pi_{2} T ∘ π 2 and S ∘ π 1 S\circ\pi_{1} S ∘ π 1 are Borel, so the pairings ( π 1 , T ∘ π 2 ) (\pi_{1},T\circ\pi_{2}) ( π 1 , T ∘ π 2 ) and ( S ∘ π 1 , π 2 ) (S\circ\pi_{1},\pi_{2}) ( S ∘ π 1 , π 2 ) are Borel by Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §pairing and π ′ , π ′ ′ ∈ P ( X × X ) \pi',\pi''\in\mathcal{P}(X\times X) π ′ , π ′′ ∈ P ( X × X ) by (c). By (c), the push-forward of π ′ \pi' π ′ by π 1 \pi_{1} π 1 is ( π 1 ) # π = μ (\pi_{1})_{\#}\pi=\mu ( π 1 ) # π = μ and by π 2 \pi_{2} π 2 is ( T ∘ π 2 ) # π = T # ( ( π 2 ) # π ) = T # ν (T\circ\pi_{2})_{\#}\pi=T_{\#}((\pi_{2})_{\#}\pi)=T_{\#}\nu ( T ∘ π 2 ) # π = T # (( π 2 ) # π ) = T # ν ; so π ′ ∈ Π ( μ , T # ν ) \pi'\in\Pi(\mu,T_{\#}\nu) π ′ ∈ Π ( μ , T # ν ) , and likewise π ′ ′ ∈ Π ( S # μ , ν ) \pi''\in\Pi(S_{\#}\mu,\nu) π ′′ ∈ Π ( S # μ , ν ) , by Couplings of Two Borel Probability Measures on a Hilbert Space and Their Quadratic Cost §coupling .
Now assume I ( π ) < ∞ I(\pi)<\infty I ( π ) < ∞ and ∫ ∣ T ( y ) − y ∣ 2 ν ( d y ) < ∞ \int|T(y)-y|^{2}\,\nu(dy)<\infty ∫ ∣ T ( y ) − y ∣ 2 ν ( d y ) < ∞ . Regard ( X × X , B ( X × X ) , π ) (X\times X,\mathcal{B}(X\times X),\pi) ( X × X , B ( X × X ) , π ) as a probability space. Let f = φ 0 f=\varphi_{0} f = φ 0 and g = φ 0 ∘ ( π 2 , T ∘ π 2 ) g=\varphi_{0}\circ(\pi_{2},T\circ\pi_{2}) g = φ 0 ∘ ( π 2 , T ∘ π 2 ) , Borel by (b) and Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §pairing , so that g ( z ) = ∣ π 2 ( z ) − T ( π 2 ( z ) ) ∣ g(z)=|\pi_{2}(z)-T(\pi_{2}(z))| g ( z ) = ∣ π 2 ( z ) − T ( π 2 ( z )) ∣ by (c). Then ∫ f 2 d π = I ( π ) < ∞ \int f^{2}\,d\pi=I(\pi)<\infty ∫ f 2 d π = I ( π ) < ∞ , and by change of variables through π 2 \pi_{2} π 2 , applied to the Borel function y ↦ ∣ y − T ( y ) ∣ 2 y\mapsto|y-T(y)|^{2} y ↦ ∣ y − T ( y ) ∣ 2 (Borel as recorded in the statement), and by (a), ∫ g 2 d π = ∫ ∣ y − T ( y ) ∣ 2 ν ( d y ) = ∫ ∣ T ( y ) − y ∣ 2 ν ( d y ) < ∞ \int g^{2}\,d\pi=\int|y-T(y)|^{2}\,\nu(dy)=\int|T(y)-y|^{2}\,\nu(dy)<\infty ∫ g 2 d π = ∫ ∣ y − T ( y ) ∣ 2 ν ( d y ) = ∫ ∣ T ( y ) − y ∣ 2 ν ( d y ) < ∞ . By (d), f f f and g g g are square-integrable with ∥ f ∥ 2 = I ( π ) \lVert f\rVert_{2}=\sqrt{I(\pi)} ∥ f ∥ 2 = I ( π ) and ∥ g ∥ 2 = ∫ ∣ T ( y ) − y ∣ 2 ν ( d y ) \lVert g\rVert_{2}=\sqrt{\int|T(y)-y|^{2}\,\nu(dy)} ∥ g ∥ 2 = ∫ ∣ T ( y ) − y ∣ 2 ν ( d y ) , and f + g f+g f + g is square-integrable. For every z z z , writing x = π 1 ( z ) x=\pi_{1}(z) x = π 1 ( z ) and y = π 2 ( z ) y=\pi_{2}(z) y = π 2 ( z ) , the vector-space axioms (Vector Space over a Field ) give x − T ( y ) = ( x − y ) + ( y − T ( y ) ) x-T(y)=(x-y)+(y-T(y)) x − T ( y ) = ( x − y ) + ( y − T ( y )) , so by (a), ∣ x − T ( y ) ∣ ≤ f ( z ) + g ( z ) |x-T(y)|\le f(z)+g(z) ∣ x − T ( y ) ∣ ≤ f ( z ) + g ( z ) and ∣ x − T ( y ) ∣ 2 ≤ ( f ( z ) + g ( z ) ) 2 |x-T(y)|^{2}\le(f(z)+g(z))^{2} ∣ x − T ( y ) ∣ 2 ≤ ( f ( z ) + g ( z ) ) 2 . By change of variables, (c) and monotonicity in claim 1 of the integral theorem,
I ( π ′ ) = ∫ φ ∘ ( π 1 , T ∘ π 2 ) d π = ∫ ∣ π 1 ( z ) − T ( π 2 ( z ) ) ∣ 2 π ( d z ) ≤ ∫ ( f + g ) 2 d π = ∥ f + g ∥ 2 2 , I(\pi')=\int\varphi\circ(\pi_{1},T\circ\pi_{2})\,d\pi=\int|\pi_{1}(z)-T(\pi_{2}(z))|^{2}\,\pi(dz)\le\int(f+g)^{2}\,d\pi=\lVert f+g\rVert_{2}^{2}, I ( π ′ ) = ∫ φ ∘ ( π 1 , T ∘ π 2 ) d π = ∫ ∣ π 1 ( z ) − T ( π 2 ( z )) ∣ 2 π ( d z ) ≤ ∫ ( f + g ) 2 d π = ∥ f + g ∥ 2 2 ,
and ∥ f + g ∥ 2 ≤ ∥ f ∥ 2 + ∥ g ∥ 2 \lVert f+g\rVert_{2}\le\lVert f\rVert_{2}+\lVert g\rVert_{2} ∥ f + g ∥ 2 ≤ ∥ f ∥ 2 + ∥ g ∥ 2 by claim 2 of Cauchy-Schwarz and Triangle Inequalities for the Mean-Square Norm . Hence I ( π ′ ) < ∞ I(\pi')<\infty I ( π ′ ) < ∞ , and I ( π ′ ) ≤ ∥ f + g ∥ 2 ≤ ∥ f ∥ 2 + ∥ g ∥ 2 \sqrt{I(\pi')}\le\lVert f+g\rVert_{2}\le\lVert f\rVert_{2}+\lVert g\rVert_{2} I ( π ′ ) ≤ ∥ f + g ∥ 2 ≤ ∥ f ∥ 2 + ∥ g ∥ 2 by claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field , which is the asserted bound.
For π ′ ′ \pi'' π ′′ , assume I ( π ) < ∞ I(\pi)<\infty I ( π ) < ∞ and ∫ ∣ S ( x ) − x ∣ 2 μ ( d x ) < ∞ \int|S(x)-x|^{2}\,\mu(dx)<\infty ∫ ∣ S ( x ) − x ∣ 2 μ ( d x ) < ∞ , and run the same argument with f = φ 0 f=\varphi_{0} f = φ 0 and g ′ ′ = φ 0 ∘ ( S ∘ π 1 , π 1 ) g''=\varphi_{0}\circ(S\circ\pi_{1},\pi_{1}) g ′′ = φ 0 ∘ ( S ∘ π 1 , π 1 ) , so that g ′ ′ ( z ) = ∣ S ( π 1 ( z ) ) − π 1 ( z ) ∣ g''(z)=|S(\pi_{1}(z))-\pi_{1}(z)| g ′′ ( z ) = ∣ S ( π 1 ( z )) − π 1 ( z ) ∣ by (c): change of variables through π 1 \pi_{1} π 1 gives ∫ g ′ ′ 2 d π = ∫ ∣ S ( x ) − x ∣ 2 μ ( d x ) < ∞ \int g''^{2}\,d\pi=\int|S(x)-x|^{2}\,\mu(dx)<\infty ∫ g ′′ 2 d π = ∫ ∣ S ( x ) − x ∣ 2 μ ( d x ) < ∞ ; for every z z z , with x = π 1 ( z ) x=\pi_{1}(z) x = π 1 ( z ) and y = π 2 ( z ) y=\pi_{2}(z) y = π 2 ( z ) , one has S ( x ) − y = ( S ( x ) − x ) + ( x − y ) S(x)-y=(S(x)-x)+(x-y) S ( x ) − y = ( S ( x ) − x ) + ( x − y ) , so ∣ S ( x ) − y ∣ 2 ≤ ( g ′ ′ ( z ) + f ( z ) ) 2 |S(x)-y|^{2}\le(g''(z)+f(z))^{2} ∣ S ( x ) − y ∣ 2 ≤ ( g ′′ ( z ) + f ( z ) ) 2 by (a); and I ( π ′ ′ ) = ∫ ∣ S ( π 1 ( z ) ) − π 2 ( z ) ∣ 2 π ( d z ) ≤ ∥ g ′ ′ + f ∥ 2 2 ≤ ( ∥ f ∥ 2 + ∥ g ′ ′ ∥ 2 ) 2 I(\pi'')=\int|S(\pi_{1}(z))-\pi_{2}(z)|^{2}\,\pi(dz)\le\lVert g''+f\rVert_{2}^{2}\le(\lVert f\rVert_{2}+\lVert g''\rVert_{2})^{2} I ( π ′′ ) = ∫ ∣ S ( π 1 ( z )) − π 2 ( z ) ∣ 2 π ( d z ) ≤ ∥ g ′′ + f ∥ 2 2 ≤ (∥ f ∥ 2 + ∥ g ′′ ∥ 2 ) 2 exactly as above, whence I ( π ′ ′ ) < ∞ I(\pi'')<\infty I ( π ′′ ) < ∞ and I ( π ′ ′ ) ≤ I ( π ) + ∫ ∣ S ( x ) − x ∣ 2 μ ( d x ) \sqrt{I(\pi'')}\le\sqrt{I(\pi)}+\sqrt{\int|S(x)-x|^{2}\,\mu(dx)} I ( π ′′ ) ≤ I ( π ) + ∫ ∣ S ( x ) − x ∣ 2 μ ( d x ) .
Claim 6. Let T : X → X T:X\to X T : X → X be Borel with finite image. Then T ( X ) T(X) T ( X ) is Borel, as recorded in the statement for finite subsets of X X X , and T − 1 ( X ∖ T ( X ) ) = ∅ T^{-1}(X\setminus T(X))=\varnothing T − 1 ( X ∖ T ( X )) = ∅ , so T # ν ( X ∖ T ( X ) ) = ν ( ∅ ) = 0 T_{\#}\nu(X\setminus T(X))=\nu(\varnothing)=0 T # ν ( X ∖ T ( X )) = ν ( ∅ ) = 0 by Measure, Measure Space, and Probability Measure .
Now let ν ∈ P 2 ( X ) \nu\in\mathcal{P}_{2}(X) ν ∈ P 2 ( X ) and 0 < ε 0<\varepsilon 0 < ε , and put η = ε 2 / 4 \eta=\varepsilon^{2}/4 η = ε 2 /4 , a positive real number. The choices are made in the order: first n n n (Step 1), then T E T_{E} T E (Step 3), then T T T (Step 4).
Step 1 (choice of n n n ). For m ∈ N m\in\mathbb{N} m ∈ N let t m ( y ) = ∣ Q m y ∣ 2 t_{m}(y)=|Q_{m}y|^{2} t m ( y ) = ∣ Q m y ∣ 2 ; t m t_{m} t m is Borel, as the composite of Q m Q_{m} Q m , Borel by Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §continuity , with the Borel map x ↦ ∣ x ∣ 2 x\mapsto|x|^{2} x ↦ ∣ x ∣ 2 of (b). By the identity ∣ y ∣ 2 = ∣ P m y ∣ 2 + ∣ Q m y ∣ 2 |y|^{2}=|P_{m}y|^{2}+|Q_{m}y|^{2} ∣ y ∣ 2 = ∣ P m y ∣ 2 + ∣ Q m y ∣ 2 of Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §continuity , 0 ≤ t m ( y ) ≤ ∣ y ∣ 2 0\le t_{m}(y)\le|y|^{2} 0 ≤ t m ( y ) ≤ ∣ y ∣ 2 for every y y y . By Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §coordinates , P m P_{m} P m is the orthogonal projection onto X m X_{m} X m , the sequence ( X m ) m ∈ N (X_{m})_{m\in\mathbb{N}} ( X m ) m ∈ N is exhausting and Q m y = y − P m y Q_{m}y=y-P_{m}y Q m y = y − P m y , so Exhausting Sequences of Finite-Dimensional Subspaces in a Separable Real Hilbert Space, and Their Projections §tail shows that ( Q m y ) m (Q_{m}y)_{m} ( Q m y ) m converges to 0 X 0_{X} 0 X in ( X , d ) (X,d) ( X , d ) for every y ∈ X y\in X y ∈ X . By The Norm Metric of a Real Inner Product Space: Triangle Inequalities, Limits and Continuity §continuity , ( ∣ Q m y ∣ ) m (|Q_{m}y|)_{m} ( ∣ Q m y ∣ ) m then converges to ∣ 0 X ∣ = d ( 0 X , 0 X ) = 0 |0_{X}|=d(0_{X},0_{X})=0 ∣ 0 X ∣ = d ( 0 X , 0 X ) = 0 (The Norm Metric of a Real Inner Product Space: Triangle Inequalities, Limits and Continuity §metric ); so for every real δ > 0 \delta>0 δ > 0 there is m 0 m_{0} m 0 with ∣ Q m y ∣ < δ |Q_{m}y|<\sqrt{\delta} ∣ Q m y ∣ < δ , hence t m ( y ) < δ t_{m}(y)<\delta t m ( y ) < δ by claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field , for all m ≥ m 0 m\ge m_{0} m ≥ m 0 ; that is, ( t m ( y ) ) m (t_{m}(y))_{m} ( t m ( y ) ) m converges to 0 0 0 . The function y ↦ ∣ y ∣ 2 y\mapsto|y|^{2} y ↦ ∣ y ∣ 2 is nonnegative with integral M 2 ( ν ) < ∞ M_{2}(\nu)<\infty M 2 ( ν ) < ∞ , hence integrable by (d), and it dominates every ∣ t m ∣ = t m |t_{m}|=t_{m} ∣ t m ∣ = t m . By claim 3 of Dominated Convergence Theorem , applied on the measure space ( X , B ( X ) , ν ) (X,\mathcal{B}(X),\nu) ( X , B ( X ) , ν ) with f m = t m f_{m}=t_{m} f m = t m , limit 0 0 0 and dominating function y ↦ ∣ y ∣ 2 y\mapsto|y|^{2} y ↦ ∣ y ∣ 2 , the integrals ∫ t m d ν \int t_{m}\,d\nu ∫ t m d ν converge to ∫ 0 d ν = ∫ 1 ∅ d ν = ν ( ∅ ) = 0 \int0\,d\nu=\int\mathbf{1}_{\varnothing}\,d\nu=\nu(\varnothing)=0 ∫ 0 d ν = ∫ 1 ∅ d ν = ν ( ∅ ) = 0 (The Integral of an Indicator Function is the Measure of the Set ), these integrals being the nonnegative ones by (d). Hence there is m 1 m_{1} m 1 with ∫ t m d ν < η \int t_{m}\,d\nu<\eta ∫ t m d ν < η for all m ≥ m 1 m\ge m_{1} m ≥ m 1 . Fix n ∈ N n\in\mathbb{N} n ∈ N with n ≥ m 1 n\ge m_{1} n ≥ m 1 and 1 ≤ n 1\le n 1 ≤ n . Then
∫ X ∣ Q n y ∣ 2 ν ( d y ) < η . \int_{X}|Q_{n}y|^{2}\,\nu(dy)<\eta . ∫ X ∣ Q n y ∣ 2 ν ( d y ) < η .
Step 2 (the projected measure). The map p n : X → R n p_{n}:X\to\mathbb{R}^{n} p n : X → R n is Borel by Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §continuity , so ν E = ( p n ) # ν \nu_{E}=(p_{n})_{\#}\nu ν E = ( p n ) # ν is a measure on ( R n , B ( R n ) ) (\mathbb{R}^{n},\mathcal{B}(\mathbb{R}^{n})) ( R n , B ( R n )) of total mass 1 1 1 by Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §pushforward and (c); here B ( R n ) \mathcal{B}(\mathbb{R}^{n}) B ( R n ) is the Borel σ \sigma σ -algebra of ( R n , d E ) (\mathbb{R}^{n},d_{E}) ( R n , d E ) (Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §measures ), which is the σ \sigma σ -algebra of Probability Measures on Euclidean Space and Random Vectors: Standing Notation §spaces , so ν E ∈ P ( R n ) \nu_{E}\in\mathcal{P}(\mathbb{R}^{n}) ν E ∈ P ( R n ) in the sense of Probability Measures on Euclidean Space and Random Vectors: Standing Notation §measures . By Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §continuity , ∥ p n ( y ) ∥ 2 = ∣ P n y ∣ 2 = ∣ y ∣ 2 − ∣ Q n y ∣ 2 ≤ ∣ y ∣ 2 \lVert p_{n}(y)\rVert^{2}=|P_{n}y|^{2}=|y|^{2}-|Q_{n}y|^{2}\le|y|^{2} ∥ p n ( y ) ∥ 2 = ∣ P n y ∣ 2 = ∣ y ∣ 2 − ∣ Q n y ∣ 2 ≤ ∣ y ∣ 2 . The map z ↦ ∥ z ∥ 2 z\mapsto\lVert z\rVert^{2} z ↦ ∥ z ∥ 2 on R n \mathbb{R}^{n} R n is Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions , so by change of variables and monotonicity,
M 2 ( ν E ) = ∫ R n ∥ z ∥ 2 ν E ( d z ) = ∫ X ∥ p n ( y ) ∥ 2 ν ( d y ) ≤ M 2 ( ν ) < ∞ , M_{2}(\nu_{E})=\int_{\mathbb{R}^{n}}\lVert z\rVert^{2}\,\nu_{E}(dz)=\int_{X}\lVert p_{n}(y)\rVert^{2}\,\nu(dy)\le M_{2}(\nu)<\infty, M 2 ( ν E ) = ∫ R n ∥ z ∥ 2 ν E ( d z ) = ∫ X ∥ p n ( y ) ∥ 2 ν ( d y ) ≤ M 2 ( ν ) < ∞ ,
with M 2 ( ν E ) M_{2}(\nu_{E}) M 2 ( ν E ) as in The Second Moment of a Probability Measure on Euclidean Space and the Probability Measures with Finite Second Moment §moment ; thus ν E ∈ P 2 ( R n ) \nu_{E}\in\mathcal{P}_{2}(\mathbb{R}^{n}) ν E ∈ P 2 ( R n ) by The Second Moment of a Probability Measure on Euclidean Space and the Probability Measures with Finite Second Moment §space .
Step 3 (Euclidean quantisation). By Couplings on Euclidean Space: Product Coupling, Swap, Finiteness of the Cost, Push-Forward Couplings, Modifying One Marginal, Quantisation, Gluing over a Finitely Supported Measure, and the Lipschitz Bound §quantisation , applied with d = n d=n d = n , both measures of its statement equal to ν E \nu_{E} ν E , and the positive real ε / 2 \varepsilon/2 ε /2 , there is a Borel map T E : R n → R n T_{E}:\mathbb{R}^{n}\to\mathbb{R}^{n} T E : R n → R n whose image is a finite set, with
∫ R n ∥ T E ( z ) − z ∥ 2 ν E ( d z ) ≤ ( ε / 2 ) 2 = η . \int_{\mathbb{R}^{n}}\lVert T_{E}(z)-z\rVert^{2}\,\nu_{E}(dz)\le(\varepsilon/2)^{2}=\eta . ∫ R n ∥ T E ( z ) − z ∥ 2 ν E ( d z ) ≤ ( ε /2 ) 2 = η .
Step 4 (the map T T T ). Let T = p n ∗ ∘ T E ∘ p n : X → X T=p_{n}^{*}\circ T_{E}\circ p_{n}:X\to X T = p n ∗ ∘ T E ∘ p n : X → X , Borel as a composite of Borel maps, p n ∗ p_{n}^{*} p n ∗ being Borel by Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §continuity . The image T E ( R n ) T_{E}(\mathbb{R}^{n}) T E ( R n ) is finite and nonempty, so by Finite Set it has k k k elements for some k ∈ N k\in\mathbb{N} k ∈ N ; p n ∗ p_{n}^{*} p n ∗ maps it onto p n ∗ ( T E ( R n ) ) p_{n}^{*}(T_{E}(\mathbb{R}^{n})) p n ∗ ( T E ( R n )) , which is therefore finite by claim 4 of Basic Properties of Finite Sets , and T ( X ) T(X) T ( X ) is a subset of it, hence finite by claim 3 there.
Fix y ∈ X y\in X y ∈ X , and let w = T E ( p n ( y ) ) w=T_{E}(p_{n}(y)) w = T E ( p n ( y )) and x = T ( y ) − y = T ( y ) + ( − 1 ) y x=T(y)-y=T(y)+(-1)y x = T ( y ) − y = T ( y ) + ( − 1 ) y (claims 2 and 5 of Elementary Identities in a Vector Space ). For k ∈ [ n ] k\in[n] k ∈ [ n ] , conditions (a), (b) and (c) of Real Inner Product Space §inner-product give ⟨ x , e k ⟩ = ⟨ T ( y ) , e k ⟩ − ⟨ y , e k ⟩ \langle x,e_{k}\rangle=\langle T(y),e_{k}\rangle-\langle y,e_{k}\rangle ⟨ x , e k ⟩ = ⟨ T ( y ) , e k ⟩ − ⟨ y , e k ⟩ , and ⟨ T ( y ) , e k ⟩ = ⟨ p n ∗ ( w ) , e k ⟩ = w k \langle T(y),e_{k}\rangle=\langle p_{n}^{*}(w),e_{k}\rangle=w_{k} ⟨ T ( y ) , e k ⟩ = ⟨ p n ∗ ( w ) , e k ⟩ = w k because p n ( p n ∗ ( w ) ) = w p_{n}(p_{n}^{*}(w))=w p n ( p n ∗ ( w )) = w by Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §continuity . Hence p n ( x ) = w − p n ( y ) p_{n}(x)=w-p_{n}(y) p n ( x ) = w − p n ( y ) , the difference in R n \mathbb{R}^{n} R n being formed componentwise by claims 2 and 3 of Euclidean Space R n \mathbb{R}^n R n is a Real Vector Space and Sum of Points of R n \mathbb{R}^n R n . Moreover, since P n = p n ∗ ∘ p n P_{n}=p_{n}^{*}\circ p_{n} P n = p n ∗ ∘ p n (Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §coordinates ), Q n T ( y ) = p n ∗ ( w ) − p n ∗ ( p n ( p n ∗ ( w ) ) ) = p n ∗ ( w ) − p n ∗ ( w ) = 0 X Q_{n}T(y)=p_{n}^{*}(w)-p_{n}^{*}(p_{n}(p_{n}^{*}(w)))=p_{n}^{*}(w)-p_{n}^{*}(w)=0_{X} Q n T ( y ) = p n ∗ ( w ) − p n ∗ ( p n ( p n ∗ ( w ))) = p n ∗ ( w ) − p n ∗ ( w ) = 0 X , and Q n Q_{n} Q n is linear by Exhausting Sequences of Finite-Dimensional Subspaces in a Separable Real Hilbert Space, and Their Projections §projections (applicable as in Step 1), so Q n x = Q n T ( y ) + ( − 1 ) Q n y = ( − 1 ) Q n y Q_{n}x=Q_{n}T(y)+(-1)Q_{n}y=(-1)Q_{n}y Q n x = Q n T ( y ) + ( − 1 ) Q n y = ( − 1 ) Q n y ; by conditions (a) and (c) of Real Inner Product Space §inner-product and Real Inner Product Space §norm , ∣ Q n x ∣ 2 = ⟨ ( − 1 ) Q n y , ( − 1 ) Q n y ⟩ = ⟨ Q n y , Q n y ⟩ = ∣ Q n y ∣ 2 |Q_{n}x|^{2}=\langle(-1)Q_{n}y,(-1)Q_{n}y\rangle=\langle Q_{n}y,Q_{n}y\rangle=|Q_{n}y|^{2} ∣ Q n x ∣ 2 = ⟨( − 1 ) Q n y , ( − 1 ) Q n y ⟩ = ⟨ Q n y , Q n y ⟩ = ∣ Q n y ∣ 2 . By the identities ∣ x ∣ 2 = ∣ P n x ∣ 2 + ∣ Q n x ∣ 2 |x|^{2}=|P_{n}x|^{2}+|Q_{n}x|^{2} ∣ x ∣ 2 = ∣ P n x ∣ 2 + ∣ Q n x ∣ 2 and ∣ P n x ∣ = ∥ p n ( x ) ∥ |P_{n}x|=\lVert p_{n}(x)\rVert ∣ P n x ∣ = ∥ p n ( x )∥ of Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §continuity ,
∣ T ( y ) − y ∣ 2 = ∥ p n ( x ) ∥ 2 + ∣ Q n x ∣ 2 = ∥ T E ( p n ( y ) ) − p n ( y ) ∥ 2 + ∣ Q n y ∣ 2 . |T(y)-y|^{2}=\lVert p_{n}(x)\rVert^{2}+|Q_{n}x|^{2}=\lVert T_{E}(p_{n}(y))-p_{n}(y)\rVert^{2}+|Q_{n}y|^{2}. ∣ T ( y ) − y ∣ 2 = ∥ p n ( x ) ∥ 2 + ∣ Q n x ∣ 2 = ∥ T E ( p n ( y )) − p n ( y ) ∥ 2 + ∣ Q n y ∣ 2 .
Step 5 (the estimate). The map z ↦ ∥ T E ( z ) − z ∥ 2 z\mapsto\lVert T_{E}(z)-z\rVert^{2} z ↦ ∥ T E ( z ) − z ∥ 2 on R n \mathbb{R}^{n} R n is Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions (with u = T E u=T_{E} u = T E and v v v the identity), so its composite with p n p_{n} p n is Borel. Integrating the identity of Step 4 against ν \nu ν , by claim 1 of the integral theorem, change of variables through p n p_{n} p n , Step 3 and Step 1,
∫ X ∣ T ( y ) − y ∣ 2 ν ( d y ) = ∫ R n ∥ T E ( z ) − z ∥ 2 ν E ( d z ) + ∫ X ∣ Q n y ∣ 2 ν ( d y ) ≤ η + η = ε 2 2 ≤ ε 2 . \int_{X}|T(y)-y|^{2}\,\nu(dy)=\int_{\mathbb{R}^{n}}\lVert T_{E}(z)-z\rVert^{2}\,\nu_{E}(dz)+\int_{X}|Q_{n}y|^{2}\,\nu(dy)\le\eta+\eta=\tfrac{\varepsilon^{2}}{2}\le\varepsilon^{2}. ∫ X ∣ T ( y ) − y ∣ 2 ν ( d y ) = ∫ R n ∥ T E ( z ) − z ∥ 2 ν E ( d z ) + ∫ X ∣ Q n y ∣ 2 ν ( d y ) ≤ η + η = 2 ε 2 ≤ ε 2 .
Claim 7. Step 1 (the atoms). Finite subsets of X X X and their complements are Borel, as recorded in the statement; in particular every singleton { a } \{a\} { a } is Borel, being finite by claim 2 of Basic Properties of Finite Sets and Finite Set . Let F + = { a ∈ F : 0 < ρ ( { a } ) } F_{+}=\{a\in F:0<\rho(\{a\})\} F + = { a ∈ F : 0 < ρ ({ a })} , finite by claim 3 of Basic Properties of Finite Sets , and for a ∈ F + a\in F_{+} a ∈ F + write r a = ρ ( { a } ) r_{a}=\rho(\{a\}) r a = ρ ({ a }) , so that 0 < r a ≤ ρ ( X ) = 1 0<r_{a}\le\rho(X)=1 0 < r a ≤ ρ ( X ) = 1 by claim 2 of Basic Properties of a Measure . Every b ∈ F ∖ F + b\in F\setminus F_{+} b ∈ F ∖ F + has ρ ( { b } ) = 0 \rho(\{b\})=0 ρ ({ b }) = 0 . The set F ∖ F + F\setminus F_{+} F ∖ F + is finite by claim 3 of Basic Properties of Finite Sets , so it is empty or, by Finite Set , the image of a bijection from some [ k ] [k] [ k ] ; in either case ρ ( F ∖ F + ) = 0 \rho(F\setminus F_{+})=0 ρ ( F ∖ F + ) = 0 , as the sum of the values ρ ( { b } ) = 0 \rho(\{b\})=0 ρ ({ b }) = 0 by claim 1 of Basic Properties of a Measure . Since X ∖ F + X\setminus F_{+} X ∖ F + is the disjoint union of X ∖ F X\setminus F X ∖ F and F ∖ F + F\setminus F_{+} F ∖ F + , the same claim gives ρ ( X ∖ F + ) = ρ ( X ∖ F ) + ρ ( F ∖ F + ) = 0 \rho(X\setminus F_{+})=\rho(X\setminus F)+\rho(F\setminus F_{+})=0 ρ ( X ∖ F + ) = ρ ( X ∖ F ) + ρ ( F ∖ F + ) = 0 , and then ρ ( F + ) = 1 \rho(F_{+})=1 ρ ( F + ) = 1 by claim 3 of Basic Properties of a Measure ; so F + F_{+} F + is nonempty, and by Finite Set we fix a bijection from some [ k ] [k] [ k ] onto F + F_{+} F + , through which sums over a ∈ F + a\in F_{+} a ∈ F + are finite sums.
For a ∈ F + a\in F_{+} a ∈ F + let A a = π 2 − 1 ( { a } ) A_{a}=\pi_{2}^{-1}(\{a\}) A a = π 2 − 1 ({ a }) and B a = π 1 − 1 ( { a } ) B_{a}=\pi_{1}^{-1}(\{a\}) B a = π 1 − 1 ({ a }) , Borel subsets of X × X X\times X X × X . As π 12 ∈ Π ( μ , ρ ) \pi_{12}\in\Pi(\mu,\rho) π 12 ∈ Π ( μ , ρ ) and π 23 ∈ Π ( ρ , λ ) \pi_{23}\in\Pi(\rho,\lambda) π 23 ∈ Π ( ρ , λ ) , Couplings of Two Borel Probability Measures on a Hilbert Space and Their Quadratic Cost §coupling gives π 12 ( A a ) = ρ ( { a } ) = r a = π 23 ( B a ) \pi_{12}(A_{a})=\rho(\{a\})=r_{a}=\pi_{23}(B_{a}) π 12 ( A a ) = ρ ({ a }) = r a = π 23 ( B a ) , and likewise π 12 ( π 2 − 1 ( X ∖ F + ) ) = ρ ( X ∖ F + ) = 0 \pi_{12}(\pi_{2}^{-1}(X\setminus F_{+}))=\rho(X\setminus F_{+})=0 π 12 ( π 2 − 1 ( X ∖ F + )) = ρ ( X ∖ F + ) = 0 and π 23 ( π 1 − 1 ( X ∖ F + ) ) = 0 \pi_{23}(\pi_{1}^{-1}(X\setminus F_{+}))=0 π 23 ( π 1 − 1 ( X ∖ F + )) = 0 . The sets A a A_{a} A a , a ∈ F + a\in F_{+} a ∈ F + , are pairwise disjoint with union π 2 − 1 ( F + ) \pi_{2}^{-1}(F_{+}) π 2 − 1 ( F + ) , so ∑ a ∈ F + 1 A a = 1 π 2 − 1 ( F + ) \sum_{a\in F_{+}}\mathbf{1}_{A_{a}}=\mathbf{1}_{\pi_{2}^{-1}(F_{+})} ∑ a ∈ F + 1 A a = 1 π 2 − 1 ( F + ) ; likewise ∑ a ∈ F + 1 B a = 1 π 1 − 1 ( F + ) \sum_{a\in F_{+}}\mathbf{1}_{B_{a}}=\mathbf{1}_{\pi_{1}^{-1}(F_{+})} ∑ a ∈ F + 1 B a = 1 π 1 − 1 ( F + ) .
Step 2 (the density). Let Ω = ( X × X ) × ( X × X ) \Omega=(X\times X)\times(X\times X) Ω = ( X × X ) × ( X × X ) with the product σ \sigma σ -algebra G = B ( X × X ) ⊗ B ( X × X ) \mathcal{G}=\mathcal{B}(X\times X)\otimes\mathcal{B}(X\times X) G = B ( X × X ) ⊗ B ( X × X ) of Product Sigma-Algebra , and let κ 1 , κ 2 : Ω → X × X \kappa_{1},\kappa_{2}:\Omega\to X\times X κ 1 , κ 2 : Ω → X × X be the factor maps κ 1 ( z , z ′ ) = z \kappa_{1}(z,z')=z κ 1 ( z , z ′ ) = z , κ 2 ( z , z ′ ) = z ′ \kappa_{2}(z,z')=z' κ 2 ( z , z ′ ) = z ′ . They are measurable with respect to G \mathcal{G} G and B ( X × X ) \mathcal{B}(X\times X) B ( X × X ) , since κ 1 − 1 ( E ) = E × ( X × X ) \kappa_{1}^{-1}(E)=E\times(X\times X) κ 1 − 1 ( E ) = E × ( X × X ) and κ 2 − 1 ( E ) = ( X × X ) × E \kappa_{2}^{-1}(E)=(X\times X)\times E κ 2 − 1 ( E ) = ( X × X ) × E are measurable rectangles, which belong to G \mathcal{G} G by Product Sigma-Algebra . The probability measures π 12 \pi_{12} π 12 and π 23 \pi_{23} π 23 are σ \sigma σ -finite (Measure, Measure Space, and Probability Measure ), so Existence and Uniqueness of the Product Measure provides the product measure θ 0 = π 12 ⊗ π 23 \theta_{0}=\pi_{12}\otimes\pi_{23} θ 0 = π 12 ⊗ π 23 on ( Ω , G ) (\Omega,\mathcal{G}) ( Ω , G ) . Let
h = ∑ a ∈ F + r a − 1 1 A a × B a : Ω → [ 0 , ∞ ) , h=\sum_{a\in F_{+}}r_{a}^{-1}\,\mathbf{1}_{A_{a}\times B_{a}}:\Omega\to[0,\infty), h = a ∈ F + ∑ r a − 1 1 A a × B a : Ω → [ 0 , ∞ ) ,
measurable by claims 1 and 2 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions , each A a × B a A_{a}\times B_{a} A a × B a being a measurable rectangle; note 1 A a × B a ( ζ ) = 1 A a ( κ 1 ζ ) 1 B a ( κ 2 ζ ) \mathbf{1}_{A_{a}\times B_{a}}(\zeta)=\mathbf{1}_{A_{a}}(\kappa_{1}\zeta)\,\mathbf{1}_{B_{a}}(\kappa_{2}\zeta) 1 A a × B a ( ζ ) = 1 A a ( κ 1 ζ ) 1 B a ( κ 2 ζ ) . Let θ \theta θ be the measure on ( Ω , G ) (\Omega,\mathcal{G}) ( Ω , G ) with density h h h with respect to θ 0 \theta_{0} θ 0 , so that by the density formula ∫ G d θ = ∫ G h d θ 0 \int G\,d\theta=\int Gh\,d\theta_{0} ∫ G d θ = ∫ G h d θ 0 for every measurable G : Ω → [ 0 , ∞ ] G:\Omega\to[0,\infty] G : Ω → [ 0 , ∞ ] .
Step 3 (two disintegration identities). Let g : X × X → [ 0 , ∞ ) g:X\times X\to[0,\infty) g : X × X → [ 0 , ∞ ) be Borel. We claim, and refer to as ( ∗ ) (\ast) ( ∗ ) ,
∫ Ω g ( κ 1 ζ ) h ( ζ ) θ 0 ( d ζ ) = ∫ X × X g d π 12 and ∫ Ω g ( κ 2 ζ ) h ( ζ ) θ 0 ( d ζ ) = ∫ X × X g d π 23 . \int_{\Omega}g(\kappa_{1}\zeta)\,h(\zeta)\,\theta_{0}(d\zeta)=\int_{X\times X}g\,d\pi_{12}\qquad\text{and}\qquad\int_{\Omega}g(\kappa_{2}\zeta)\,h(\zeta)\,\theta_{0}(d\zeta)=\int_{X\times X}g\,d\pi_{23}. ∫ Ω g ( κ 1 ζ ) h ( ζ ) θ 0 ( d ζ ) = ∫ X × X g d π 12 and ∫ Ω g ( κ 2 ζ ) h ( ζ ) θ 0 ( d ζ ) = ∫ X × X g d π 23 .
For the first, let G a ( ζ ) = g ( κ 1 ζ ) 1 A a ( κ 1 ζ ) 1 B a ( κ 2 ζ ) G_{a}(\zeta)=g(\kappa_{1}\zeta)\mathbf{1}_{A_{a}}(\kappa_{1}\zeta)\mathbf{1}_{B_{a}}(\kappa_{2}\zeta) G a ( ζ ) = g ( κ 1 ζ ) 1 A a ( κ 1 ζ ) 1 B a ( κ 2 ζ ) for a ∈ F + a\in F_{+} a ∈ F + , a nonnegative G \mathcal{G} G -measurable function by claims 1 and 3 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions ; then g ( κ 1 ζ ) h ( ζ ) = ∑ a ∈ F + r a − 1 G a ( ζ ) g(\kappa_{1}\zeta)h(\zeta)=\sum_{a\in F_{+}}r_{a}^{-1}G_{a}(\zeta) g ( κ 1 ζ ) h ( ζ ) = ∑ a ∈ F + r a − 1 G a ( ζ ) , and claim 1 of the integral theorem gives ∫ ( g ∘ κ 1 ) h d θ 0 = ∑ a ∈ F + r a − 1 ∫ G a d θ 0 \int(g\circ\kappa_{1})h\,d\theta_{0}=\sum_{a\in F_{+}}r_{a}^{-1}\int G_{a}\,d\theta_{0} ∫ ( g ∘ κ 1 ) h d θ 0 = ∑ a ∈ F + r a − 1 ∫ G a d θ 0 . By the Tonelli clause of Tonelli and Fubini Theorems ,
∫ Ω G a d θ 0 = ∫ X × X ( ∫ X × X g ( z ) 1 A a ( z ) 1 B a ( z ′ ) π 23 ( d z ′ ) ) π 12 ( d z ) . \int_{\Omega}G_{a}\,d\theta_{0}=\int_{X\times X}\Bigl(\int_{X\times X}g(z)\mathbf{1}_{A_{a}}(z)\mathbf{1}_{B_{a}}(z')\,\pi_{23}(dz')\Bigr)\pi_{12}(dz). ∫ Ω G a d θ 0 = ∫ X × X ( ∫ X × X g ( z ) 1 A a ( z ) 1 B a ( z ′ ) π 23 ( d z ′ ) ) π 12 ( d z ) .
For fixed z z z the inner integrand is the constant c = g ( z ) 1 A a ( z ) ∈ [ 0 , ∞ ) c=g(z)\mathbf{1}_{A_{a}}(z)\in[0,\infty) c = g ( z ) 1 A a ( z ) ∈ [ 0 , ∞ ) times 1 B a \mathbf{1}_{B_{a}} 1 B a , so the inner integral is c π 23 ( B a ) = r a g ( z ) 1 A a ( z ) c\,\pi_{23}(B_{a})=r_{a}\,g(z)\mathbf{1}_{A_{a}}(z) c π 23 ( B a ) = r a g ( z ) 1 A a ( z ) by claim 1 of the integral theorem and The Integral of an Indicator Function is the Measure of the Set ; hence ∫ G a d θ 0 = r a ∫ g 1 A a d π 12 \int G_{a}\,d\theta_{0}=r_{a}\int g\mathbf{1}_{A_{a}}\,d\pi_{12} ∫ G a d θ 0 = r a ∫ g 1 A a d π 12 by the same claim. Summing, and using claim 1 of the integral theorem and Step 1,
∫ Ω ( g ∘ κ 1 ) h d θ 0 = ∑ a ∈ F + ∫ g 1 A a d π 12 = ∫ g 1 π 2 − 1 ( F + ) d π 12 . \int_{\Omega}(g\circ\kappa_{1})h\,d\theta_{0}=\sum_{a\in F_{+}}\int g\mathbf{1}_{A_{a}}\,d\pi_{12}=\int g\,\mathbf{1}_{\pi_{2}^{-1}(F_{+})}\,d\pi_{12}. ∫ Ω ( g ∘ κ 1 ) h d θ 0 = a ∈ F + ∑ ∫ g 1 A a d π 12 = ∫ g 1 π 2 − 1 ( F + ) d π 12 .
The integrands g g g and g 1 π 2 − 1 ( F + ) g\mathbf{1}_{\pi_{2}^{-1}(F_{+})} g 1 π 2 − 1 ( F + ) agree off π 2 − 1 ( X ∖ F + ) \pi_{2}^{-1}(X\setminus F_{+}) π 2 − 1 ( X ∖ F + ) , which is π 12 \pi_{12} π 12 -null by Step 1, so their integrals coincide by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §comparison ; this proves the first identity. The second is proved in the same way with G a ′ ( ζ ) = g ( κ 2 ζ ) 1 A a ( κ 1 ζ ) 1 B a ( κ 2 ζ ) G'_{a}(\zeta)=g(\kappa_{2}\zeta)\mathbf{1}_{A_{a}}(\kappa_{1}\zeta)\mathbf{1}_{B_{a}}(\kappa_{2}\zeta) G a ′ ( ζ ) = g ( κ 2 ζ ) 1 A a ( κ 1 ζ ) 1 B a ( κ 2 ζ ) and the other order of integration in Tonelli's theorem: for fixed z ′ z' z ′ the inner π 12 \pi_{12} π 12 -integral is g ( z ′ ) 1 B a ( z ′ ) π 12 ( A a ) = r a g ( z ′ ) 1 B a ( z ′ ) g(z')\mathbf{1}_{B_{a}}(z')\pi_{12}(A_{a})=r_{a}\,g(z')\mathbf{1}_{B_{a}}(z') g ( z ′ ) 1 B a ( z ′ ) π 12 ( A a ) = r a g ( z ′ ) 1 B a ( z ′ ) , so ∫ ( g ∘ κ 2 ) h d θ 0 = ∑ a ∈ F + ∫ g 1 B a d π 23 = ∫ g 1 π 1 − 1 ( F + ) d π 23 = ∫ g d π 23 \int(g\circ\kappa_{2})h\,d\theta_{0}=\sum_{a\in F_{+}}\int g\mathbf{1}_{B_{a}}\,d\pi_{23}=\int g\mathbf{1}_{\pi_{1}^{-1}(F_{+})}\,d\pi_{23}=\int g\,d\pi_{23} ∫ ( g ∘ κ 2 ) h d θ 0 = ∑ a ∈ F + ∫ g 1 B a d π 23 = ∫ g 1 π 1 − 1 ( F + ) d π 23 = ∫ g d π 23 , the last step because π 1 − 1 ( X ∖ F + ) \pi_{1}^{-1}(X\setminus F_{+}) π 1 − 1 ( X ∖ F + ) is π 23 \pi_{23} π 23 -null.
Step 4 (the glued coupling). By the density formula, The Integral of an Indicator Function is the Measure of the Set and ( ∗ ) (\ast) ( ∗ ) with g g g the constant 1 1 1 , θ ( Ω ) = ∫ h d θ 0 = π 12 ( X × X ) = 1 \theta(\Omega)=\int h\,d\theta_{0}=\pi_{12}(X\times X)=1 θ ( Ω ) = ∫ h d θ 0 = π 12 ( X × X ) = 1 , so ( Ω , G , θ ) (\Omega,\mathcal{G},\theta) ( Ω , G , θ ) is a probability space. Let Φ = ( π 1 ∘ κ 1 , π 2 ∘ κ 2 ) : Ω → X × X \Phi=(\pi_{1}\circ\kappa_{1},\pi_{2}\circ\kappa_{2}):\Omega\to X\times X Φ = ( π 1 ∘ κ 1 , π 2 ∘ κ 2 ) : Ω → X × X , measurable with respect to G \mathcal{G} G and B ( X × X ) \mathcal{B}(X\times X) B ( X × X ) by Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §pairing , and let π 13 = Φ # θ \pi_{13}=\Phi_{\#}\theta π 13 = Φ # θ , which lies in P ( X × X ) \mathcal{P}(X\times X) P ( X × X ) by (c). For A ∈ B ( X ) A\in\mathcal{B}(X) A ∈ B ( X ) , by (c), the density formula and The Integral of an Indicator Function is the Measure of the Set ,
( ( π 1 ) # π 13 ) ( A ) = θ ( κ 1 − 1 ( π 1 − 1 ( A ) ) ) = ∫ Ω 1 π 1 − 1 ( A ) ( κ 1 ζ ) h ( ζ ) θ 0 ( d ζ ) = π 12 ( π 1 − 1 ( A ) ) = μ ( A ) , \bigl((\pi_{1})_{\#}\pi_{13}\bigr)(A)=\theta\bigl(\kappa_{1}^{-1}(\pi_{1}^{-1}(A))\bigr)=\int_{\Omega}\mathbf{1}_{\pi_{1}^{-1}(A)}(\kappa_{1}\zeta)\,h(\zeta)\,\theta_{0}(d\zeta)=\pi_{12}(\pi_{1}^{-1}(A))=\mu(A), ( ( π 1 ) # π 13 ) ( A ) = θ ( κ 1 − 1 ( π 1 − 1 ( A )) ) = ∫ Ω 1 π 1 − 1 ( A ) ( κ 1 ζ ) h ( ζ ) θ 0 ( d ζ ) = π 12 ( π 1 − 1 ( A )) = μ ( A ) ,
the third equality by ( ∗ ) (\ast) ( ∗ ) with the Borel function g = 1 π 1 − 1 ( A ) g=\mathbf{1}_{\pi_{1}^{-1}(A)} g = 1 π 1 − 1 ( A ) (claim 1 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions ). Likewise, for C ∈ B ( X ) C\in\mathcal{B}(X) C ∈ B ( X ) , ( ( π 2 ) # π 13 ) ( C ) = ∫ ( 1 π 2 − 1 ( C ) ∘ κ 2 ) h d θ 0 = π 23 ( π 2 − 1 ( C ) ) = λ ( C ) ((\pi_{2})_{\#}\pi_{13})(C)=\int(\mathbf{1}_{\pi_{2}^{-1}(C)}\circ\kappa_{2})h\,d\theta_{0}=\pi_{23}(\pi_{2}^{-1}(C))=\lambda(C) (( π 2 ) # π 13 ) ( C ) = ∫ ( 1 π 2 − 1 ( C ) ∘ κ 2 ) h d θ 0 = π 23 ( π 2 − 1 ( C )) = λ ( C ) . Hence π 13 ∈ Π ( μ , λ ) \pi_{13}\in\Pi(\mu,\lambda) π 13 ∈ Π ( μ , λ ) by Couplings of Two Borel Probability Measures on a Hilbert Space and Their Quadratic Cost §coupling .
Step 5 (the cost bound). On the probability space ( Ω , G , θ ) (\Omega,\mathcal{G},\theta) ( Ω , G , θ ) let f = φ 0 ∘ κ 1 f=\varphi_{0}\circ\kappa_{1} f = φ 0 ∘ κ 1 and f ′ = φ 0 ∘ κ 2 f'=\varphi_{0}\circ\kappa_{2} f ′ = φ 0 ∘ κ 2 , measurable by (b). By the density formula and ( ∗ ) (\ast) ( ∗ ) with g = φ g=\varphi g = φ , ∫ f 2 d θ = ∫ ( φ ∘ κ 1 ) h d θ 0 = ∫ φ d π 12 = I ( π 12 ) < ∞ \int f^{2}\,d\theta=\int(\varphi\circ\kappa_{1})h\,d\theta_{0}=\int\varphi\,d\pi_{12}=I(\pi_{12})<\infty ∫ f 2 d θ = ∫ ( φ ∘ κ 1 ) h d θ 0 = ∫ φ d π 12 = I ( π 12 ) < ∞ , and similarly ∫ f ′ 2 d θ = I ( π 23 ) < ∞ \int f'^{2}\,d\theta=I(\pi_{23})<\infty ∫ f ′ 2 d θ = I ( π 23 ) < ∞ ; so by (d), f f f , f ′ f' f ′ and f + f ′ f+f' f + f ′ are square-integrable, with ∥ f ∥ 2 = I ( π 12 ) \lVert f\rVert_{2}=\sqrt{I(\pi_{12})} ∥ f ∥ 2 = I ( π 12 ) and ∥ f ′ ∥ 2 = I ( π 23 ) \lVert f'\rVert_{2}=\sqrt{I(\pi_{23})} ∥ f ′ ∥ 2 = I ( π 23 ) . Let u = φ ∘ Φ u=\varphi\circ\Phi u = φ ∘ Φ , so that u ( ζ ) = ∣ π 1 ( κ 1 ζ ) − π 2 ( κ 2 ζ ) ∣ 2 u(\zeta)=|\pi_{1}(\kappa_{1}\zeta)-\pi_{2}(\kappa_{2}\zeta)|^{2} u ( ζ ) = ∣ π 1 ( κ 1 ζ ) − π 2 ( κ 2 ζ ) ∣ 2 by (c). Let ζ = ( z , z ′ ) ∈ Ω \zeta=(z,z')\in\Omega ζ = ( z , z ′ ) ∈ Ω . If h ( ζ ) > 0 h(\zeta)>0 h ( ζ ) > 0 , some term of h h h is nonzero at ζ \zeta ζ , so z ∈ A a z\in A_{a} z ∈ A a and z ′ ∈ B a z'\in B_{a} z ′ ∈ B a for some a ∈ F + a\in F_{+} a ∈ F + , that is, π 2 ( z ) = a = π 1 ( z ′ ) \pi_{2}(z)=a=\pi_{1}(z') π 2 ( z ) = a = π 1 ( z ′ ) ; then π 1 ( z ) − π 2 ( z ′ ) = ( π 1 ( z ) − π 2 ( z ) ) + ( π 1 ( z ′ ) − π 2 ( z ′ ) ) \pi_{1}(z)-\pi_{2}(z')=(\pi_{1}(z)-\pi_{2}(z))+(\pi_{1}(z')-\pi_{2}(z')) π 1 ( z ) − π 2 ( z ′ ) = ( π 1 ( z ) − π 2 ( z )) + ( π 1 ( z ′ ) − π 2 ( z ′ )) by the vector-space axioms (Vector Space over a Field ), so ∣ π 1 ( z ) − π 2 ( z ′ ) ∣ ≤ f ( ζ ) + f ′ ( ζ ) |\pi_{1}(z)-\pi_{2}(z')|\le f(\zeta)+f'(\zeta) ∣ π 1 ( z ) − π 2 ( z ′ ) ∣ ≤ f ( ζ ) + f ′ ( ζ ) and u ( ζ ) ≤ ( f ( ζ ) + f ′ ( ζ ) ) 2 u(\zeta)\le(f(\zeta)+f'(\zeta))^{2} u ( ζ ) ≤ ( f ( ζ ) + f ′ ( ζ ) ) 2 by (a). If h ( ζ ) = 0 h(\zeta)=0 h ( ζ ) = 0 , both u ( ζ ) h ( ζ ) u(\zeta)h(\zeta) u ( ζ ) h ( ζ ) and ( f ( ζ ) + f ′ ( ζ ) ) 2 h ( ζ ) (f(\zeta)+f'(\zeta))^{2}h(\zeta) ( f ( ζ ) + f ′ ( ζ ) ) 2 h ( ζ ) vanish. Hence u h ≤ ( f + f ′ ) 2 h uh\le(f+f')^{2}h u h ≤ ( f + f ′ ) 2 h pointwise, and by change of variables, the density formula, monotonicity in claim 1 of the integral theorem and claim 2 of Cauchy-Schwarz and Triangle Inequalities for the Mean-Square Norm ,
I ( π 13 ) = ∫ φ ∘ Φ d θ = ∫ u h d θ 0 ≤ ∫ ( f + f ′ ) 2 h d θ 0 = ∫ ( f + f ′ ) 2 d θ = ∥ f + f ′ ∥ 2 2 ≤ ( ∥ f ∥ 2 + ∥ f ′ ∥ 2 ) 2 . I(\pi_{13})=\int\varphi\circ\Phi\,d\theta=\int uh\,d\theta_{0}\le\int(f+f')^{2}h\,d\theta_{0}=\int(f+f')^{2}\,d\theta=\lVert f+f'\rVert_{2}^{2}\le\bigl(\lVert f\rVert_{2}+\lVert f'\rVert_{2}\bigr)^{2}. I ( π 13 ) = ∫ φ ∘ Φ d θ = ∫ u h d θ 0 ≤ ∫ ( f + f ′ ) 2 h d θ 0 = ∫ ( f + f ′ ) 2 d θ = ∥ f + f ′ ∥ 2 2 ≤ ( ∥ f ∥ 2 + ∥ f ′ ∥ 2 ) 2 .
So I ( π 13 ) I(\pi_{13}) I ( π 13 ) is finite and I ( π 13 ) ≤ ∥ f ∥ 2 + ∥ f ′ ∥ 2 = I ( π 12 ) + I ( π 23 ) \sqrt{I(\pi_{13})}\le\lVert f\rVert_{2}+\lVert f'\rVert_{2}=\sqrt{I(\pi_{12})}+\sqrt{I(\pi_{23})} I ( π 13 ) ≤ ∥ f ∥ 2 + ∥ f ′ ∥ 2 = I ( π 12 ) + I ( π 23 ) by claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field .
Claim 8. A Lipschitz f f f is continuous by A Lipschitz Map is Uniformly Continuous , hence Borel by claims 2 and 3 of Borel Measurability and Bounded Integration on a Metric Space ; being bounded, with a bound M M M as in Bounded Real-Valued Function on a Set , it is integrable with respect to μ \mu μ and to ν \nu ν by claim 6(b) of Borel Measurability and Bounded Integration on a Metric Space . By change of variables in the integrable case, f ∘ π 1 f\circ\pi_{1} f ∘ π 1 and f ∘ π 2 f\circ\pi_{2} f ∘ π 2 are π \pi π -integrable with ∫ f ∘ π 1 d π = ∫ f d μ \int f\circ\pi_{1}\,d\pi=\int f\,d\mu ∫ f ∘ π 1 d π = ∫ f d μ and ∫ f ∘ π 2 d π = ∫ f d ν \int f\circ\pi_{2}\,d\pi=\int f\,d\nu ∫ f ∘ π 2 d π = ∫ f d ν . For every z z z , ∣ f ( π 1 ( z ) ) − f ( π 2 ( z ) ) ∣ ≤ L φ 0 ( z ) |f(\pi_{1}(z))-f(\pi_{2}(z))|\le L\,\varphi_{0}(z) ∣ f ( π 1 ( z )) − f ( π 2 ( z )) ∣ ≤ L φ 0 ( z ) by the Lipschitz property. By claim 2 of the integral theorem, then monotonicity and homogeneity in its claim 1 (the integral of the integrable function ∣ f ∘ π 1 − f ∘ π 2 ∣ |f\circ\pi_{1}-f\circ\pi_{2}| ∣ f ∘ π 1 − f ∘ π 2 ∣ being its nonnegative integral by (d)),
∣ ∫ f d μ − ∫ f d ν ∣ = ∣ ∫ ( f ∘ π 1 − f ∘ π 2 ) d π ∣ ≤ ∫ ∣ f ∘ π 1 − f ∘ π 2 ∣ d π ≤ L ∫ φ 0 d π . \Bigl|\int f\,d\mu-\int f\,d\nu\Bigr|=\Bigl|\int(f\circ\pi_{1}-f\circ\pi_{2})\,d\pi\Bigr|\le\int|f\circ\pi_{1}-f\circ\pi_{2}|\,d\pi\le L\int\varphi_{0}\,d\pi . ∫ f d μ − ∫ f d ν = ∫ ( f ∘ π 1 − f ∘ π 2 ) d π ≤ ∫ ∣ f ∘ π 1 − f ∘ π 2 ∣ d π ≤ L ∫ φ 0 d π .
On the probability space ( X × X , B ( X × X ) , π ) (X\times X,\mathcal{B}(X\times X),\pi) ( X × X , B ( X × X ) , π ) the function φ 0 \varphi_{0} φ 0 is square-integrable with ∥ φ 0 ∥ 2 = I ( π ) \lVert\varphi_{0}\rVert_{2}=\sqrt{I(\pi)} ∥ φ 0 ∥ 2 = I ( π ) by (d), as φ 0 2 = φ \varphi_{0}^{2}=\varphi φ 0 2 = φ and I ( π ) < ∞ I(\pi)<\infty I ( π ) < ∞ , and the constant 1 1 1 is square-integrable with ∥ 1 ∥ 2 = 1 \lVert1\rVert_{2}=1 ∥ 1 ∥ 2 = 1 by The Integral of an Indicator Function is the Measure of the Set with the set X × X X\times X X × X . So ∫ φ 0 d π = E [ φ 0 ⋅ 1 ] ≤ ∥ φ 0 ∥ 2 ∥ 1 ∥ 2 = I ( π ) \int\varphi_{0}\,d\pi=\mathbb{E}[\varphi_{0}\cdot1]\le\lVert\varphi_{0}\rVert_{2}\lVert1\rVert_{2}=\sqrt{I(\pi)} ∫ φ 0 d π = E [ φ 0 ⋅ 1 ] ≤ ∥ φ 0 ∥ 2 ∥ 1 ∥ 2 = I ( π ) by claim 1 of Cauchy-Schwarz and Triangle Inequalities for the Mean-Square Norm , the expectation being the integral by (d) and Expectation, Variance, and Moments . As 0 ≤ L 0\le L 0 ≤ L , combining gives ∣ ∫ f d μ − ∫ f d ν ∣ ≤ L I ( π ) |\int f\,d\mu-\int f\,d\nu|\le L\sqrt{I(\pi)} ∣ ∫ f d μ − ∫ f d ν ∣ ≤ L I ( π ) .
Claim 9. Let μ , ν ∈ P 2 ( X ) \mu,\nu\in\mathcal{P}_{2}(X) μ , ν ∈ P 2 ( X ) and π ∈ Π ( μ , ν ) \pi\in\Pi(\mu,\nu) π ∈ Π ( μ , ν ) . By claim 3, I ( π ) < ∞ I(\pi)<\infty I ( π ) < ∞ . On the probability space ( X × X , B ( X × X ) , π ) (X\times X,\mathcal{B}(X\times X),\pi) ( X × X , B ( X × X ) , π ) let a ( z ) = ∣ π 1 ( z ) ∣ a(z)=|\pi_{1}(z)| a ( z ) = ∣ π 1 ( z ) ∣ and b ( z ) = ∣ π 2 ( z ) ∣ b(z)=|\pi_{2}(z)| b ( z ) = ∣ π 2 ( z ) ∣ , Borel by (b). By change of variables and The Second Moment of a Borel Probability Measure on a Hilbert Space and the Probability Measures with Finite Second Moment §moment , ∫ a 2 d π = M 2 ( μ ) < ∞ \int a^{2}\,d\pi=M_{2}(\mu)<\infty ∫ a 2 d π = M 2 ( μ ) < ∞ and ∫ b 2 d π = M 2 ( ν ) < ∞ \int b^{2}\,d\pi=M_{2}(\nu)<\infty ∫ b 2 d π = M 2 ( ν ) < ∞ , and ∫ φ 0 2 d π = I ( π ) < ∞ \int\varphi_{0}^{2}\,d\pi=I(\pi)<\infty ∫ φ 0 2 d π = I ( π ) < ∞ ; so by (d), a a a , b b b and φ 0 \varphi_{0} φ 0 are square-integrable, with ∥ a ∥ 2 = M 2 ( μ ) \lVert a\rVert_{2}=\sqrt{M_{2}(\mu)} ∥ a ∥ 2 = M 2 ( μ ) , ∥ b ∥ 2 = M 2 ( ν ) \lVert b\rVert_{2}=\sqrt{M_{2}(\nu)} ∥ b ∥ 2 = M 2 ( ν ) and ∥ φ 0 ∥ 2 = I ( π ) \lVert\varphi_{0}\rVert_{2}=\sqrt{I(\pi)} ∥ φ 0 ∥ 2 = I ( π ) , and so are b + φ 0 b+\varphi_{0} b + φ 0 and a + φ 0 a+\varphi_{0} a + φ 0 . For every z z z , π 1 ( z ) = π 2 ( z ) + ( π 1 ( z ) − π 2 ( z ) ) \pi_{1}(z)=\pi_{2}(z)+(\pi_{1}(z)-\pi_{2}(z)) π 1 ( z ) = π 2 ( z ) + ( π 1 ( z ) − π 2 ( z )) by the vector-space axioms (Vector Space over a Field ), so (a) gives a ( z ) ≤ b ( z ) + φ 0 ( z ) a(z)\le b(z)+\varphi_{0}(z) a ( z ) ≤ b ( z ) + φ 0 ( z ) and a ( z ) 2 ≤ ( b ( z ) + φ 0 ( z ) ) 2 a(z)^{2}\le(b(z)+\varphi_{0}(z))^{2} a ( z ) 2 ≤ ( b ( z ) + φ 0 ( z ) ) 2 . By monotonicity in claim 1 of the integral theorem, ∥ a ∥ 2 2 ≤ ∥ b + φ 0 ∥ 2 2 \lVert a\rVert_{2}^{2}\le\lVert b+\varphi_{0}\rVert_{2}^{2} ∥ a ∥ 2 2 ≤ ∥ b + φ 0 ∥ 2 2 , so ∥ a ∥ 2 ≤ ∥ b + φ 0 ∥ 2 ≤ ∥ b ∥ 2 + ∥ φ 0 ∥ 2 \lVert a\rVert_{2}\le\lVert b+\varphi_{0}\rVert_{2}\le\lVert b\rVert_{2}+\lVert\varphi_{0}\rVert_{2} ∥ a ∥ 2 ≤ ∥ b + φ 0 ∥ 2 ≤ ∥ b ∥ 2 + ∥ φ 0 ∥ 2 by claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field and claim 2 of Cauchy-Schwarz and Triangle Inequalities for the Mean-Square Norm ; that is, M 2 ( μ ) − M 2 ( ν ) ≤ I ( π ) \sqrt{M_{2}(\mu)}-\sqrt{M_{2}(\nu)}\le\sqrt{I(\pi)} M 2 ( μ ) − M 2 ( ν ) ≤ I ( π ) . Symmetrically, π 2 ( z ) = π 1 ( z ) + ( π 2 ( z ) − π 1 ( z ) ) \pi_{2}(z)=\pi_{1}(z)+(\pi_{2}(z)-\pi_{1}(z)) π 2 ( z ) = π 1 ( z ) + ( π 2 ( z ) − π 1 ( z )) and ∣ π 2 ( z ) − π 1 ( z ) ∣ = φ 0 ( z ) |\pi_{2}(z)-\pi_{1}(z)|=\varphi_{0}(z) ∣ π 2 ( z ) − π 1 ( z ) ∣ = φ 0 ( z ) by (a), so b ( z ) ≤ a ( z ) + φ 0 ( z ) b(z)\le a(z)+\varphi_{0}(z) b ( z ) ≤ a ( z ) + φ 0 ( z ) , and the same argument gives M 2 ( ν ) − M 2 ( μ ) ≤ I ( π ) \sqrt{M_{2}(\nu)}-\sqrt{M_{2}(\mu)}\le\sqrt{I(\pi)} M 2 ( ν ) − M 2 ( μ ) ≤ I ( π ) . The two inequalities together give ∣ M 2 ( μ ) − M 2 ( ν ) ∣ ≤ I ( π ) \bigl|\sqrt{M_{2}(\mu)}-\sqrt{M_{2}(\nu)}\bigr|\le\sqrt{I(\pi)} M 2 ( μ ) − M 2 ( ν ) ≤ I ( π ) .