Each result cited is universally quantified over the data in its own statement. "The integral theorem" refers to Linearity and Monotonicity of the Lebesgue Integral (claim 1: additivity, homogeneity with a factor in [ 0 , ∞ ) [0,\infty) [ 0 , ∞ ) , and monotonicity of the nonnegative integral; claim 2: linearity, ∣ ∫ f ∣ ≤ ∫ ∣ f ∣ |\int f|\le\int|f| ∣ ∫ f ∣ ≤ ∫ ∣ f ∣ and monotonicity for integrable functions), and "change of variables" to the formula of Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pushforward . Let φ : R d + d → R \varphi:\mathbb{R}^{d+d}\to\mathbb{R} φ : R d + d → R be the Borel map z ↦ ∥ p r 1 ( z ) − p r 2 ( z ) ∥ 2 z\mapsto\lVert\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z)\rVert^{2} z ↦ ∥ pr 1 ( z ) − pr 2 ( z ) ∥ 2 and φ 0 \varphi_{0} φ 0 the Borel map z ↦ ∥ p r 1 ( z ) − p r 2 ( z ) ∥ z\mapsto\lVert\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z)\rVert z ↦ ∥ pr 1 ( z ) − pr 2 ( z )∥ , both of Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions , so that I ( π ) = ∫ φ d π I(\pi)=\int\varphi\,d\pi I ( π ) = ∫ φ d π for every π ∈ P ( R d + d ) \pi\in\mathcal{P}(\mathbb{R}^{d+d}) π ∈ P ( R d + d ) by Couplings of Two Probability Measures on Euclidean Space and Their Quadratic Cost §cost , each such π \pi π being a coupling of its own push-forwards by Couplings of Two Probability Measures on Euclidean Space and Their Quadratic Cost §coupling . Two facts are used repeatedly. (i) For Borel G : R m → R l G:\mathbb{R}^{m}\to\mathbb{R}^{l} G : R m → R l and H : R l → R k H:\mathbb{R}^{l}\to\mathbb{R}^{k} H : R l → R k and λ ∈ P ( R m ) \lambda\in\mathcal{P}(\mathbb{R}^{m}) λ ∈ P ( R m ) , one has ( H ∘ G ) # λ = H # ( G # λ ) (H\circ G)_{\#}\lambda=H_{\#}(G_{\#}\lambda) ( H ∘ G ) # λ = H # ( G # λ ) , since both sides assign to a Borel B B B the value λ ( G − 1 ( H − 1 ( B ) ) ) \lambda(G^{-1}(H^{-1}(B))) λ ( G − 1 ( H − 1 ( B ))) . (ii) For u , v : R m → R d u,v:\mathbb{R}^{m}\to\mathbb{R}^{d} u , v : R m → R d Borel, p r 1 ∘ ( u , v ) = u \mathrm{pr}_{1}\circ(u,v)=u pr 1 ∘ ( u , v ) = u and p r 2 ∘ ( u , v ) = v \mathrm{pr}_{2}\circ(u,v)=v pr 2 ∘ ( u , v ) = v by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections , and φ ( ( u , v ) ( x ) ) = ∥ u ( x ) − v ( x ) ∥ 2 \varphi((u,v)(x))=\lVert u(x)-v(x)\rVert^{2} φ (( u , v ) ( x )) = ∥ u ( x ) − v ( x ) ∥ 2 , φ 0 ( ( u , v ) ( x ) ) = ∥ u ( x ) − v ( x ) ∥ \varphi_{0}((u,v)(x))=\lVert u(x)-v(x)\rVert φ 0 (( u , v ) ( x )) = ∥ u ( x ) − v ( x )∥ . Finally, if λ \lambda λ is a probability measure on a measurable space, a nonnegative measurable real function with finite integral is integrable with the same integral, by Integrable Function and the Lebesgue Integral (its negative part is 0 0 0 ); this identifies the two readings of such integrals below.
Claim 1. By Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pairs and Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §product (with q = p = d q=p=d q = p = d ), μ ⊠ ν ∈ P ( R d + d ) \mu\boxtimes\nu\in\mathcal{P}(\mathbb{R}^{d+d}) μ ⊠ ν ∈ P ( R d + d ) and its push-forwards by p r 1 \mathrm{pr}_{1} pr 1 and p r 2 \mathrm{pr}_{2} pr 2 are μ \mu μ and ν \nu ν ; so μ ⊠ ν ∈ Π ( μ , ν ) \mu\boxtimes\nu\in\Pi(\mu,\nu) μ ⊠ ν ∈ Π ( μ , ν ) by Couplings of Two Probability Measures on Euclidean Space and Their Quadratic Cost §coupling .
Claim 2. The swap σ \sigma σ is Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §pairing , so σ # π ∈ P ( R d + d ) \sigma_{\#}\pi\in\mathcal{P}(\mathbb{R}^{d+d}) σ # π ∈ P ( R d + d ) . By Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections , p r 1 ∘ σ = p r 2 \mathrm{pr}_{1}\circ\sigma=\mathrm{pr}_{2} pr 1 ∘ σ = pr 2 and p r 2 ∘ σ = p r 1 \mathrm{pr}_{2}\circ\sigma=\mathrm{pr}_{1} pr 2 ∘ σ = pr 1 , and σ ( σ ( z ) ) = ι ( p r 2 ( σ ( z ) ) , p r 1 ( σ ( z ) ) ) = ι ( p r 1 ( z ) , p r 2 ( z ) ) = z \sigma(\sigma(z))=\iota(\mathrm{pr}_{2}(\sigma(z)),\mathrm{pr}_{1}(\sigma(z)))=\iota(\mathrm{pr}_{1}(z),\mathrm{pr}_{2}(z))=z σ ( σ ( z )) = ι ( pr 2 ( σ ( z )) , pr 1 ( σ ( z ))) = ι ( pr 1 ( z ) , pr 2 ( z )) = z . Hence by (i), ( p r 1 ) # ( σ # π ) = ( p r 1 ∘ σ ) # π = ( p r 2 ) # π = ν (\mathrm{pr}_{1})_{\#}(\sigma_{\#}\pi)=(\mathrm{pr}_{1}\circ\sigma)_{\#}\pi=(\mathrm{pr}_{2})_{\#}\pi=\nu ( pr 1 ) # ( σ # π ) = ( pr 1 ∘ σ ) # π = ( pr 2 ) # π = ν and ( p r 2 ) # ( σ # π ) = μ (\mathrm{pr}_{2})_{\#}(\sigma_{\#}\pi)=\mu ( pr 2 ) # ( σ # π ) = μ , so σ # π ∈ Π ( ν , μ ) \sigma_{\#}\pi\in\Pi(\nu,\mu) σ # π ∈ Π ( ν , μ ) ; and σ # ( σ # π ) = ( σ ∘ σ ) # π = π \sigma_{\#}(\sigma_{\#}\pi)=(\sigma\circ\sigma)_{\#}\pi=\pi σ # ( σ # π ) = ( σ ∘ σ ) # π = π . By change of variables, I ( σ # π ) = ∫ φ ∘ σ d π I(\sigma_{\#}\pi)=\int\varphi\circ\sigma\,d\pi I ( σ # π ) = ∫ φ ∘ σ d π , and φ ( σ ( z ) ) = ∥ p r 2 ( z ) − p r 1 ( z ) ∥ 2 = φ ( z ) \varphi(\sigma(z))=\lVert\mathrm{pr}_{2}(z)-\mathrm{pr}_{1}(z)\rVert^{2}=\varphi(z) φ ( σ ( z )) = ∥ pr 2 ( z ) − pr 1 ( z ) ∥ 2 = φ ( z ) because p r 2 ( z ) − p r 1 ( z ) = ( − 1 ) ( p r 1 ( z ) − p r 2 ( z ) ) \mathrm{pr}_{2}(z)-\mathrm{pr}_{1}(z)=(-1)(\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z)) pr 2 ( z ) − pr 1 ( z ) = ( − 1 ) ( pr 1 ( z ) − pr 2 ( z )) by claims 2 and 3 of Euclidean Space R n \mathbb{R}^n R n is a Real Vector Space and claims 2 and 5 of Elementary Identities in a Vector Space , and ∥ ( − 1 ) x ∥ = ∥ x ∥ \lVert(-1)x\rVert=\lVert x\rVert ∥( − 1 ) x ∥ = ∥ x ∥ by claim 5 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n ; so I ( σ # π ) = I ( π ) I(\sigma_{\#}\pi)=I(\pi) I ( σ # π ) = I ( π ) . The map π ↦ σ # π \pi\mapsto\sigma_{\#}\pi π ↦ σ # π thus sends Π ( μ , ν ) \Pi(\mu,\nu) Π ( μ , ν ) into Π ( ν , μ ) \Pi(\nu,\mu) Π ( ν , μ ) , and the same map on Π ( ν , μ ) \Pi(\nu,\mu) Π ( ν , μ ) is a two-sided inverse of it, so it is a bijection.
Claim 3. By Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions , φ ( z ) ≤ 2 ∥ p r 1 ( z ) ∥ 2 + 2 ∥ p r 2 ( z ) ∥ 2 \varphi(z)\le2\lVert\mathrm{pr}_{1}(z)\rVert^{2}+2\lVert\mathrm{pr}_{2}(z)\rVert^{2} φ ( z ) ≤ 2 ∥ pr 1 ( z ) ∥ 2 + 2 ∥ pr 2 ( z ) ∥ 2 for every z z z . The maps z ↦ ∥ p r i ( z ) ∥ 2 z\mapsto\lVert\mathrm{pr}_{i}(z)\rVert^{2} z ↦ ∥ pr i ( z ) ∥ 2 are Borel, being compositions of p r i \mathrm{pr}_{i} pr i with x ↦ ∥ x ∥ 2 x\mapsto\lVert x\rVert^{2} x ↦ ∥ x ∥ 2 (Probability Measures on Euclidean Space and Random Vectors: Standing Notation §borel-maps ), and by change of variables ∫ ∥ p r 1 ( z ) ∥ 2 π ( d z ) = ∫ ∥ x ∥ 2 ( ( p r 1 ) # π ) ( d x ) = M 2 ( μ ) \int\lVert\mathrm{pr}_{1}(z)\rVert^{2}\,\pi(dz)=\int\lVert x\rVert^{2}\,((\mathrm{pr}_{1})_{\#}\pi)(dx)=M_{2}(\mu) ∫ ∥ pr 1 ( z ) ∥ 2 π ( d z ) = ∫ ∥ x ∥ 2 (( pr 1 ) # π ) ( d x ) = M 2 ( μ ) , and likewise the second integral is M 2 ( ν ) M_{2}(\nu) M 2 ( ν ) , by The Second Moment of a Probability Measure on Euclidean Space and the Probability Measures with Finite Second Moment §moment . Claim 1 of the integral theorem gives I ( π ) ≤ 2 M 2 ( μ ) + 2 M 2 ( ν ) I(\pi)\le2M_{2}(\mu)+2M_{2}(\nu) I ( π ) ≤ 2 M 2 ( μ ) + 2 M 2 ( ν ) in [ 0 , ∞ ] [0,\infty] [ 0 , ∞ ] , which is finite when μ , ν ∈ P 2 ( R d ) \mu,\nu\in\mathcal{P}_{2}(\mathbb{R}^{d}) μ , ν ∈ P 2 ( R d ) . Conversely, Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions with x = p r 2 ( z ) x=\mathrm{pr}_{2}(z) x = pr 2 ( z ) and y = p r 1 ( z ) y=\mathrm{pr}_{1}(z) y = pr 1 ( z ) gives ∥ p r 2 ( z ) ∥ 2 ≤ 2 ∥ p r 1 ( z ) ∥ 2 + 2 φ ( z ) \lVert\mathrm{pr}_{2}(z)\rVert^{2}\le2\lVert\mathrm{pr}_{1}(z)\rVert^{2}+2\varphi(z) ∥ pr 2 ( z ) ∥ 2 ≤ 2 ∥ pr 1 ( z ) ∥ 2 + 2 φ ( z ) , using ∥ p r 2 ( z ) − p r 1 ( z ) ∥ = ∥ p r 1 ( z ) − p r 2 ( z ) ∥ \lVert\mathrm{pr}_{2}(z)-\mathrm{pr}_{1}(z)\rVert=\lVert\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z)\rVert ∥ pr 2 ( z ) − pr 1 ( z )∥ = ∥ pr 1 ( z ) − pr 2 ( z )∥ as in claim 2; integrating as before, M 2 ( ν ) ≤ 2 M 2 ( μ ) + 2 I ( π ) M_{2}(\nu)\le2M_{2}(\mu)+2I(\pi) M 2 ( ν ) ≤ 2 M 2 ( μ ) + 2 I ( π ) , which is finite under the stated hypotheses, so ν ∈ P 2 ( R d ) \nu\in\mathcal{P}_{2}(\mathbb{R}^{d}) ν ∈ P 2 ( R d ) .
Claim 4. The pairing ( S , T ) (S,T) ( S , T ) is Borel by Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pairs , so ( S , T ) # μ ∈ P ( R d + d ) (S,T)_{\#}\mu\in\mathcal{P}(\mathbb{R}^{d+d}) ( S , T ) # μ ∈ P ( R d + d ) , and by (i) and (ii) its push-forward by p r 1 \mathrm{pr}_{1} pr 1 is ( p r 1 ∘ ( S , T ) ) # μ = S # μ (\mathrm{pr}_{1}\circ(S,T))_{\#}\mu=S_{\#}\mu ( pr 1 ∘ ( S , T ) ) # μ = S # μ , and by p r 2 \mathrm{pr}_{2} pr 2 is T # μ T_{\#}\mu T # μ . By change of variables and (ii), I ( ( S , T ) # μ ) = ∫ φ ∘ ( S , T ) d μ = ∫ ∥ S ( x ) − T ( x ) ∥ 2 μ ( d x ) I((S,T)_{\#}\mu)=\int\varphi\circ(S,T)\,d\mu=\int\lVert S(x)-T(x)\rVert^{2}\,\mu(dx) I (( S , T ) # μ ) = ∫ φ ∘ ( S , T ) d μ = ∫ ∥ S ( x ) − T ( x ) ∥ 2 μ ( d x ) . For S = T = i d S=T=\mathrm{id} S = T = id , i d # μ = μ \mathrm{id}_{\#}\mu=\mu id # μ = μ and the integrand is ∥ x − x ∥ 2 = 0 \lVert x-x\rVert^{2}=0 ∥ x − x ∥ 2 = 0 by claim 2 of Elementary Identities in a Vector Space and claim 3 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n , whose integral is 0 0 0 by claim 1 of the integral theorem with the factor 0 0 0 .
Claim 5. The maps T ∘ p r 2 T\circ\mathrm{pr}_{2} T ∘ pr 2 and S ∘ p r 1 S\circ\mathrm{pr}_{1} S ∘ pr 1 are Borel by Probability Measures on Euclidean Space and Random Vectors: Standing Notation §borel-maps , so the two pairings are Borel and the push-forwards lie in P ( R d + d ) \mathcal{P}(\mathbb{R}^{d+d}) P ( R d + d ) . By (i) and (ii), the push-forward of π ′ = ( p r 1 , T ∘ p r 2 ) # π \pi'=(\mathrm{pr}_{1},T\circ\mathrm{pr}_{2})_{\#}\pi π ′ = ( pr 1 , T ∘ pr 2 ) # π by p r 1 \mathrm{pr}_{1} pr 1 is ( p r 1 ) # π = μ (\mathrm{pr}_{1})_{\#}\pi=\mu ( pr 1 ) # π = μ and by p r 2 \mathrm{pr}_{2} pr 2 is ( T ∘ p r 2 ) # π = T # ( ( p r 2 ) # π ) = T # ν (T\circ\mathrm{pr}_{2})_{\#}\pi=T_{\#}((\mathrm{pr}_{2})_{\#}\pi)=T_{\#}\nu ( T ∘ pr 2 ) # π = T # (( pr 2 ) # π ) = T # ν ; so π ′ ∈ Π ( μ , T # ν ) \pi'\in\Pi(\mu,T_{\#}\nu) π ′ ∈ Π ( μ , T # ν ) , and likewise ( S ∘ p r 1 , p r 2 ) # π ∈ Π ( S # μ , ν ) (S\circ\mathrm{pr}_{1},\mathrm{pr}_{2})_{\#}\pi\in\Pi(S_{\#}\mu,\nu) ( S ∘ pr 1 , pr 2 ) # π ∈ Π ( S # μ , ν ) . Now assume I ( π ) < ∞ I(\pi)<\infty I ( π ) < ∞ and ∫ ∥ T ( y ) − y ∥ 2 ν ( d y ) < ∞ \int\lVert T(y)-y\rVert^{2}\,\nu(dy)<\infty ∫ ∥ T ( y ) − y ∥ 2 ν ( d y ) < ∞ . Regard ( R d + d , B ( R d + d ) , π ) (\mathbb{R}^{d+d},\mathcal{B}(\mathbb{R}^{d+d}),\pi) ( R d + d , B ( R d + d ) , π ) as a probability space, so that Borel real functions on R d + d \mathbb{R}^{d+d} R d + d are random variables on it; let f = φ 0 f=\varphi_{0} f = φ 0 and g = φ 0 ∘ ( p r 2 , T ∘ p r 2 ) g=\varphi_{0}\circ(\mathrm{pr}_{2},T\circ\mathrm{pr}_{2}) g = φ 0 ∘ ( pr 2 , T ∘ pr 2 ) , Borel by Probability Measures on Euclidean Space and Random Vectors: Standing Notation §borel-maps , so that f ( z ) = ∥ p r 1 ( z ) − p r 2 ( z ) ∥ f(z)=\lVert\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z)\rVert f ( z ) = ∥ pr 1 ( z ) − pr 2 ( z )∥ and g ( z ) = ∥ p r 2 ( z ) − T ( p r 2 ( z ) ) ∥ g(z)=\lVert\mathrm{pr}_{2}(z)-T(\mathrm{pr}_{2}(z))\rVert g ( z ) = ∥ pr 2 ( z ) − T ( pr 2 ( z ))∥ by (ii). Then ∫ f 2 d π = I ( π ) < ∞ \int f^{2}\,d\pi=I(\pi)<\infty ∫ f 2 d π = I ( π ) < ∞ , and ∫ g 2 d π = ∫ ∥ p r 2 ( z ) − T ( p r 2 ( z ) ) ∥ 2 π ( d z ) = ∫ ∥ y − T ( y ) ∥ 2 ν ( d y ) < ∞ \int g^{2}\,d\pi=\int\lVert\mathrm{pr}_{2}(z)-T(\mathrm{pr}_{2}(z))\rVert^{2}\,\pi(dz)=\int\lVert y-T(y)\rVert^{2}\,\nu(dy)<\infty ∫ g 2 d π = ∫ ∥ pr 2 ( z ) − T ( pr 2 ( z )) ∥ 2 π ( d z ) = ∫ ∥ y − T ( y ) ∥ 2 ν ( d y ) < ∞ by change of variables through p r 2 \mathrm{pr}_{2} pr 2 , the integrand being ∥ y − T ( y ) ∥ 2 = ∥ T ( y ) − y ∥ 2 \lVert y-T(y)\rVert^{2}=\lVert T(y)-y\rVert^{2} ∥ y − T ( y ) ∥ 2 = ∥ T ( y ) − y ∥ 2 (claim 5 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n as in claim 2); so f f f and g g g are square-integrable random variables with ∥ f ∥ 2 = I ( π ) \lVert f\rVert_{2}=\sqrt{I(\pi)} ∥ f ∥ 2 = I ( π ) and ∥ g ∥ 2 = ∫ ∥ T ( y ) − y ∥ 2 ν ( d y ) \lVert g\rVert_{2}=\sqrt{\int\lVert T(y)-y\rVert^{2}\,\nu(dy)} ∥ g ∥ 2 = ∫ ∥ T ( y ) − y ∥ 2 ν ( d y ) . For every z z z , writing x = p r 1 ( z ) x=\mathrm{pr}_{1}(z) x = pr 1 ( z ) and y = p r 2 ( z ) y=\mathrm{pr}_{2}(z) y = pr 2 ( z ) , one has x − T ( y ) = ( x − y ) + ( y − T ( y ) ) x-T(y)=(x-y)+(y-T(y)) x − T ( y ) = ( x − y ) + ( y − T ( y )) by Euclidean Space R n \mathbb{R}^n R n is a Real Vector Space , so ∥ x − T ( y ) ∥ ≤ f ( z ) + g ( z ) \lVert x-T(y)\rVert\le f(z)+g(z) ∥ x − T ( y )∥ ≤ f ( z ) + g ( z ) by claim 6 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n , and ∥ x − T ( y ) ∥ 2 ≤ ( f ( z ) + g ( z ) ) 2 \lVert x-T(y)\rVert^{2}\le(f(z)+g(z))^{2} ∥ x − T ( y ) ∥ 2 ≤ ( f ( z ) + g ( z ) ) 2 by claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field . By change of variables and (ii), I ( π ′ ) = ∫ φ ∘ ( p r 1 , T ∘ p r 2 ) d π = ∫ ∥ p r 1 ( z ) − T ( p r 2 ( z ) ) ∥ 2 π ( d z ) ≤ ∫ ( f + g ) 2 d π = ∥ f + g ∥ 2 2 I(\pi')=\int\varphi\circ(\mathrm{pr}_{1},T\circ\mathrm{pr}_{2})\,d\pi=\int\lVert\mathrm{pr}_{1}(z)-T(\mathrm{pr}_{2}(z))\rVert^{2}\,\pi(dz)\le\int(f+g)^{2}\,d\pi=\lVert f+g\rVert_{2}^{2} I ( π ′ ) = ∫ φ ∘ ( pr 1 , T ∘ pr 2 ) d π = ∫ ∥ pr 1 ( z ) − T ( pr 2 ( z )) ∥ 2 π ( d z ) ≤ ∫ ( f + g ) 2 d π = ∥ f + g ∥ 2 2 by monotonicity, and ∥ f + g ∥ 2 ≤ ∥ f ∥ 2 + ∥ g ∥ 2 \lVert f+g\rVert_{2}\le\lVert f\rVert_{2}+\lVert g\rVert_{2} ∥ f + g ∥ 2 ≤ ∥ f ∥ 2 + ∥ g ∥ 2 by claim 2 of Cauchy-Schwarz and Triangle Inequalities for the Mean-Square Norm ; hence I ( π ′ ) I(\pi') I ( π ′ ) is finite and I ( π ′ ) ≤ ∥ f + g ∥ 2 ≤ ∥ f ∥ 2 + ∥ g ∥ 2 \sqrt{I(\pi')}\le\lVert f+g\rVert_{2}\le\lVert f\rVert_{2}+\lVert g\rVert_{2} I ( π ′ ) ≤ ∥ f + g ∥ 2 ≤ ∥ f ∥ 2 + ∥ g ∥ 2 by claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field , which is the asserted bound. For the second assertion, one checks from Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections that σ ∘ ( p r 1 , S ∘ p r 2 ) ∘ σ = ( S ∘ p r 1 , p r 2 ) \sigma\circ(\mathrm{pr}_{1},S\circ\mathrm{pr}_{2})\circ\sigma=(S\circ\mathrm{pr}_{1},\mathrm{pr}_{2}) σ ∘ ( pr 1 , S ∘ pr 2 ) ∘ σ = ( S ∘ pr 1 , pr 2 ) pointwise, so by (i) ( S ∘ p r 1 , p r 2 ) # π = σ # ( ( p r 1 , S ∘ p r 2 ) # ( σ # π ) ) (S\circ\mathrm{pr}_{1},\mathrm{pr}_{2})_{\#}\pi=\sigma_{\#}\bigl((\mathrm{pr}_{1},S\circ\mathrm{pr}_{2})_{\#}(\sigma_{\#}\pi)\bigr) ( S ∘ pr 1 , pr 2 ) # π = σ # ( ( pr 1 , S ∘ pr 2 ) # ( σ # π ) ) ; by claim 2, σ # π ∈ Π ( ν , μ ) \sigma_{\#}\pi\in\Pi(\nu,\mu) σ # π ∈ Π ( ν , μ ) with I ( σ # π ) = I ( π ) I(\sigma_{\#}\pi)=I(\pi) I ( σ # π ) = I ( π ) , so the first assertion applied to σ # π \sigma_{\#}\pi σ # π and S S S bounds the cost of ( p r 1 , S ∘ p r 2 ) # ( σ # π ) (\mathrm{pr}_{1},S\circ\mathrm{pr}_{2})_{\#}(\sigma_{\#}\pi) ( pr 1 , S ∘ pr 2 ) # ( σ # π ) by the square of I ( π ) + ∫ ∥ S ( x ) − x ∥ 2 μ ( d x ) \sqrt{I(\pi)}+\sqrt{\int\lVert S(x)-x\rVert^{2}\,\mu(dx)} I ( π ) + ∫ ∥ S ( x ) − x ∥ 2 μ ( d x ) , and claim 2 again shows that the swap preserves this cost.
Claim 6. If T T T has finite image F = T ( R d ) F=T(\mathbb{R}^{d}) F = T ( R d ) , then T − 1 ( R d ∖ F ) = ∅ T^{-1}(\mathbb{R}^{d}\setminus F)=\varnothing T − 1 ( R d ∖ F ) = ∅ , so T # ν ( R d ∖ F ) = ν ( ∅ ) = 0 T_{\#}\nu(\mathbb{R}^{d}\setminus F)=\nu(\varnothing)=0 T # ν ( R d ∖ F ) = ν ( ∅ ) = 0 by Measure, Measure Space, and Probability Measure . Now let ν ∈ P 2 ( R d ) \nu\in\mathcal{P}_{2}(\mathbb{R}^{d}) ν ∈ P 2 ( R d ) and 0 < ε 0<\varepsilon 0 < ε ; put η = ε 2 ⋅ 2 − 1 \eta=\varepsilon^{2}\cdot2^{-1} η = ε 2 ⋅ 2 − 1 and ε ′ = ε ⋅ 2 − 1 \varepsilon'=\varepsilon\cdot2^{-1} ε ′ = ε ⋅ 2 − 1 , so that η + η = ε 2 \eta+\eta=\varepsilon^{2} η + η = ε 2 , 0 < η 0<\eta 0 < η , 0 < ε ′ 0<\varepsilon' 0 < ε ′ by claim 8 of Elementary Order Arithmetic in an Ordered Field , and ε ′ 2 = ε 2 ⋅ 2 − 1 ⋅ 2 − 1 ≤ η \varepsilon'^{2}=\varepsilon^{2}\cdot2^{-1}\cdot2^{-1}\le\eta ε ′ 2 = ε 2 ⋅ 2 − 1 ⋅ 2 − 1 ≤ η because 2 − 1 < 1 2^{-1}<1 2 − 1 < 1 (claim 8 of Elementary Order Arithmetic in an Ordered Field with ε = 1 \varepsilon=1 ε = 1 , using claim 6 there) and claim 5 of Elementary Arithmetic in an Ordered Field . Let κ : N → R \kappa:\mathbb{N}\to\mathbb{R} κ : N → R be the canonical map . For n ∈ N n\in\mathbb{N} n ∈ N let K n = { y : ∥ y ∥ ≤ κ ( n ) } K_{n}=\{y:\lVert y\rVert\le\kappa(n)\} K n = { y : ∥ y ∥ ≤ κ ( n )} , the complement of { y : ∥ y ∥ > κ ( n ) } \{y:\lVert y\rVert>\kappa(n)\} { y : ∥ y ∥ > κ ( n )} , which is Borel by Measure Spaces and the Lebesgue Integral: Standing Notation §measurable since y ↦ ∥ y ∥ y\mapsto\lVert y\rVert y ↦ ∥ y ∥ is Borel; and let h n = ∥ ⋅ ∥ 2 1 K n h_{n}=\lVert\cdot\rVert^{2}\mathbf{1}_{K_{n}} h n = ∥ ⋅ ∥ 2 1 K n , Borel by claims 1 and 3 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions . As κ \kappa κ is increasing (claim 6 of Properties of the Canonical Map from the Natural Numbers to an Ordered Field ), K n ⊆ K n + 1 K_{n}\subseteq K_{n+1} K n ⊆ K n + 1 and ( h n ( y ) ) n (h_{n}(y))_{n} ( h n ( y ) ) n is nondecreasing for each y y y ; and by claim 1 of The Archimedean Property of the Real Numbers there is n n n with ∥ y ∥ < κ ( n ) \lVert y\rVert<\kappa(n) ∥ y ∥ < κ ( n ) , so h m ( y ) = ∥ y ∥ 2 h_{m}(y)=\lVert y\rVert^{2} h m ( y ) = ∥ y ∥ 2 for all m ≥ n m\ge n m ≥ n and sup n h n ( y ) = ∥ y ∥ 2 \sup_{n}h_{n}(y)=\lVert y\rVert^{2} sup n h n ( y ) = ∥ y ∥ 2 . By Monotone Convergence Theorem , M 2 ( ν ) = sup n ∫ h n d ν M_{2}(\nu)=\sup_{n}\int h_{n}\,d\nu M 2 ( ν ) = sup n ∫ h n d ν , each ∫ h n d ν \int h_{n}\,d\nu ∫ h n d ν being a real number at most M 2 ( ν ) < ∞ M_{2}(\nu)<\infty M 2 ( ν ) < ∞ by monotonicity. By claim 3 of Approximation Property of the Supremum and the Infimum in R \mathbb{R} R there is n n n with M 2 ( ν ) − η < ∫ h n d ν M_{2}(\nu)-\eta<\int h_{n}\,d\nu M 2 ( ν ) − η < ∫ h n d ν . Fix this n n n , put R = κ ( n ) R=\kappa(n) R = κ ( n ) and K = K n K=K_{n} K = K n , and let g = ∥ ⋅ ∥ 2 1 R d ∖ K g=\lVert\cdot\rVert^{2}\mathbf{1}_{\mathbb{R}^{d}\setminus K} g = ∥ ⋅ ∥ 2 1 R d ∖ K , Borel likewise. Since ∥ y ∥ 2 = h n ( y ) + g ( y ) \lVert y\rVert^{2}=h_{n}(y)+g(y) ∥ y ∥ 2 = h n ( y ) + g ( y ) for every y y y , claim 1 of the integral theorem gives M 2 ( ν ) = ∫ h n d ν + ∫ g d ν M_{2}(\nu)=\int h_{n}\,d\nu+\int g\,d\nu M 2 ( ν ) = ∫ h n d ν + ∫ g d ν with both terms real, whence ∫ g d ν = M 2 ( ν ) − ∫ h n d ν < η \int g\,d\nu=M_{2}(\nu)-\int h_{n}\,d\nu<\eta ∫ g d ν = M 2 ( ν ) − ∫ h n d ν < η .
K K K is the closed ball of ( R d , d E ) (\mathbb{R}^{d},d_{E}) ( R d , d E ) with centre 0 R d 0_{\mathbb{R}^{d}} 0 R d and radius R R R , since d E ( 0 R d , y ) = ∥ y ∥ d_{E}(0_{\mathbb{R}^{d}},y)=\lVert y\rVert d E ( 0 R d , y ) = ∥ y ∥ by claim 2 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n and symmetry of the metric; it is compact by claim 2 of A Closed Euclidean Ball is Convex and Compact , the radius R = κ ( n ) R=\kappa(n) R = κ ( n ) being nonnegative by claim 3 of Properties of the Canonical Map from the Natural Numbers to an Ordered Field . The family ( B ( c , ε ′ ) ) c ∈ K (B(c,\varepsilon'))_{c\in K} ( B ( c , ε ′ ) ) c ∈ K of open balls , which are open by Open Ball in a Metric Space is Open , covers K K K , as c ∈ B ( c , ε ′ ) c\in B(c,\varepsilon') c ∈ B ( c , ε ′ ) ; by Compact Subset Criterion via Open Covers in the Ambient Space there is a finite J ⊆ K J\subseteq K J ⊆ K with K ⊆ ⋃ c ∈ J B ( c , ε ′ ) K\subseteq\bigcup_{c\in J}B(c,\varepsilon') K ⊆ ⋃ c ∈ J B ( c , ε ′ ) . Since 0 R d ∈ K 0_{\mathbb{R}^{d}}\in K 0 R d ∈ K , J J J is nonempty, so by Finite Set there are n ′ ∈ N n'\in\mathbb{N} n ′ ∈ N and a bijection [ n ′ ] → J [n']\to J [ n ′ ] → J , j ↦ c j j\mapsto c_{j} j ↦ c j . For j ∈ [ n ′ ] j\in[n'] j ∈ [ n ′ ] let B j = B ( c j , ε ′ ) B_{j}=B(c_{j},\varepsilon') B j = B ( c j , ε ′ ) and E j = ( K ∩ B j ) ∖ ⋃ i ∈ [ n ′ ] , i < j B i E_{j}=(K\cap B_{j})\setminus\bigcup_{i\in[n'],\,i<j}B_{i} E j = ( K ∩ B j ) ∖ ⋃ i ∈ [ n ′ ] , i < j B i , a Borel set (open sets being Borel by Probability Measures on Euclidean Space and Random Vectors: Standing Notation §spaces , and finite intersections, unions and differences of Borel sets being Borel by Sigma-Algebra and Measurable Space ). For y ∈ K y\in K y ∈ K the set { j ∈ [ n ′ ] : y ∈ B j } \{j\in[n']:y\in B_{j}\} { j ∈ [ n ′ ] : y ∈ B j } is nonempty, so it has a least element j ( y ) j(y) j ( y ) by The Natural Numbers Are Well Ordered , and y ∈ E j ( y ) y\in E_{j(y)} y ∈ E j ( y ) while y ∉ E j y\notin E_{j} y ∈ / E j for j ≠ j ( y ) j\ne j(y) j = j ( y ) ; thus the sets E j E_{j} E j are pairwise disjoint with union K K K . Define T : R d → R d T:\mathbb{R}^{d}\to\mathbb{R}^{d} T : R d → R d by T ( y ) = c j ( y ) T(y)=c_{j(y)} T ( y ) = c j ( y ) for y ∈ K y\in K y ∈ K and T ( y ) = 0 R d T(y)=0_{\mathbb{R}^{d}} T ( y ) = 0 R d for y ∉ K y\notin K y ∈ / K . For B ∈ B ( R d ) B\in\mathcal{B}(\mathbb{R}^{d}) B ∈ B ( R d ) ,
T − 1 ( B ) = ⋃ j ∈ [ n ′ ] , c j ∈ B E j ∪ N B , N B = R d ∖ K if 0 R d ∈ B , N B = ∅ otherwise , T^{-1}(B)=\bigcup_{j\in[n'],\,c_{j}\in B}E_{j}\ \cup\ N_{B},\qquad N_{B}=\mathbb{R}^{d}\setminus K\ \text{if }0_{\mathbb{R}^{d}}\in B,\quad N_{B}=\varnothing\ \text{otherwise}, T − 1 ( B ) = j ∈ [ n ′ ] , c j ∈ B ⋃ E j ∪ N B , N B = R d ∖ K if 0 R d ∈ B , N B = ∅ otherwise ,
a finite union of Borel sets, hence Borel; so T T T is Borel. Its image is contained in { c 1 , … , c n ′ } ∪ { 0 R d } \{c_{1},\dots,c_{n'}\}\cup\{0_{\mathbb{R}^{d}}\} { c 1 , … , c n ′ } ∪ { 0 R d } , the image of the finite set [ n ′ + 1 ] [n'+1] [ n ′ + 1 ] under the map sending j ≤ n ′ j\le n' j ≤ n ′ to c j c_{j} c j and n ′ + 1 n'+1 n ′ + 1 to 0 R d 0_{\mathbb{R}^{d}} 0 R d , which is finite by claims 1 and 4 of Basic Properties of Finite Sets ; so T ( R d ) T(\mathbb{R}^{d}) T ( R d ) is finite by claim 3 there. For y ∈ K y\in K y ∈ K , ∥ T ( y ) − y ∥ = d E ( c j ( y ) , y ) < ε ′ \lVert T(y)-y\rVert=d_{E}(c_{j(y)},y)<\varepsilon' ∥ T ( y ) − y ∥ = d E ( c j ( y ) , y ) < ε ′ since y ∈ B j ( y ) y\in B_{j(y)} y ∈ B j ( y ) , so ∥ T ( y ) − y ∥ 2 ≤ ε ′ 2 ≤ η \lVert T(y)-y\rVert^{2}\le\varepsilon'^{2}\le\eta ∥ T ( y ) − y ∥ 2 ≤ ε ′ 2 ≤ η by claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field ; for y ∉ K y\notin K y ∈ / K , T ( y ) − y = ( − 1 ) y T(y)-y=(-1)y T ( y ) − y = ( − 1 ) y by claims 2 and 3 of Euclidean Space R n \mathbb{R}^n R n is a Real Vector Space and Vector Space over a Field , so ∥ T ( y ) − y ∥ 2 = ∥ y ∥ 2 = g ( y ) \lVert T(y)-y\rVert^{2}=\lVert y\rVert^{2}=g(y) ∥ T ( y ) − y ∥ 2 = ∥ y ∥ 2 = g ( y ) by claim 5 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n . Hence ∥ T ( y ) − y ∥ 2 ≤ η 1 K ( y ) + g ( y ) \lVert T(y)-y\rVert^{2}\le\eta\mathbf{1}_{K}(y)+g(y) ∥ T ( y ) − y ∥ 2 ≤ η 1 K ( y ) + g ( y ) for every y y y , and by claim 1 of the integral theorem and The Integral of an Indicator Function is the Measure of the Set ,
∫ ∥ T ( y ) − y ∥ 2 ν ( d y ) ≤ η ν ( K ) + ∫ g d ν ≤ η + η = ε 2 , \int\lVert T(y)-y\rVert^{2}\,\nu(dy)\le\eta\,\nu(K)+\int g\,d\nu\le\eta+\eta=\varepsilon^{2}, ∫ ∥ T ( y ) − y ∥ 2 ν ( d y ) ≤ η ν ( K ) + ∫ g d ν ≤ η + η = ε 2 ,
using ν ( K ) ≤ 1 \nu(K)\le1 ν ( K ) ≤ 1 (claim 2 of Basic Properties of a Measure ) and claim 5 of Elementary Arithmetic in an Ordered Field .
Claim 7. Write F + = { y ∈ F : 0 < ρ ( { y } ) } F_{+}=\{y\in F:0<\rho(\{y\})\} F + = { y ∈ F : 0 < ρ ({ y })} , a subset of F F F , finite by claim 3 of Basic Properties of Finite Sets (singletons being Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §finite-sets , so that ρ ( { y } ) \rho(\{y\}) ρ ({ y }) is defined), and let w : R d → R w:\mathbb{R}^{d}\to\mathbb{R} w : R d → R be w ( y ) = ρ ( { y } ) − 1 w(y)=\rho(\{y\})^{-1} w ( y ) = ρ ({ y } ) − 1 for y ∈ F + y\in F_{+} y ∈ F + and w ( y ) = 0 w(y)=0 w ( y ) = 0 otherwise, Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §finite-sets applied to the finite set F + F_{+} F + . Then w ( y ) ρ ( { y } ) = 1 F + ( y ) w(y)\rho(\{y\})=\mathbf{1}_{F_{+}}(y) w ( y ) ρ ({ y }) = 1 F + ( y ) for every y y y . Moreover ρ ( R d ∖ F + ) = 0 \rho(\mathbb{R}^{d}\setminus F_{+})=0 ρ ( R d ∖ F + ) = 0 : the set F ∖ F + F\setminus F_{+} F ∖ F + is finite by claim 3 of Basic Properties of Finite Sets , so it is empty or the image of a bijection [ k ] → F ∖ F + [k]\to F\setminus F_{+} [ k ] → F ∖ F + , and then ρ ( F ∖ F + ) \rho(F\setminus F_{+}) ρ ( F ∖ F + ) is the sum of the k k k values ρ ( { y } ) = 0 \rho(\{y\})=0 ρ ({ y }) = 0 by claim 1 of Basic Properties of a Measure , hence 0 0 0 ; and ρ ( R d ∖ F + ) = ρ ( R d ∖ F ) + ρ ( F ∖ F + ) = 0 \rho(\mathbb{R}^{d}\setminus F_{+})=\rho(\mathbb{R}^{d}\setminus F)+\rho(F\setminus F_{+})=0 ρ ( R d ∖ F + ) = ρ ( R d ∖ F ) + ρ ( F ∖ F + ) = 0 by the same claim.
Let q = d + d q=d+d q = d + d and consider R q + q \mathbb{R}^{q+q} R q + q with the projections P 1 = p r 1 q , q P_{1}=\mathrm{pr}^{q,q}_{1} P 1 = pr 1 q , q , P 2 = p r 2 q , q P_{2}=\mathrm{pr}^{q,q}_{2} P 2 = pr 2 q , q and the concatenation ι q , q \iota^{q,q} ι q , q of Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pairs ; for ζ ∈ R q + q \zeta\in\mathbb{R}^{q+q} ζ ∈ R q + q write x ( ζ ) = p r 1 ( P 1 ζ ) x(\zeta)=\mathrm{pr}_{1}(P_{1}\zeta) x ( ζ ) = pr 1 ( P 1 ζ ) , y ( ζ ) = p r 2 ( P 1 ζ ) y(\zeta)=\mathrm{pr}_{2}(P_{1}\zeta) y ( ζ ) = pr 2 ( P 1 ζ ) , y ′ ( ζ ) = p r 1 ( P 2 ζ ) y'(\zeta)=\mathrm{pr}_{1}(P_{2}\zeta) y ′ ( ζ ) = pr 1 ( P 2 ζ ) and v ( ζ ) = p r 2 ( P 2 ζ ) v(\zeta)=\mathrm{pr}_{2}(P_{2}\zeta) v ( ζ ) = pr 2 ( P 2 ζ ) , four Borel maps R q + q → R d \mathbb{R}^{q+q}\to\mathbb{R}^{d} R q + q → R d (Probability Measures on Euclidean Space and Random Vectors: Standing Notation §borel-maps ). Let θ 0 = π 12 ⊠ π 23 ∈ P ( R q + q ) \theta_{0}=\pi_{12}\boxtimes\pi_{23}\in\mathcal{P}(\mathbb{R}^{q+q}) θ 0 = π 12 ⊠ π 23 ∈ P ( R q + q ) (Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §product ). The set D = { ζ : y ( ζ ) = y ′ ( ζ ) } D=\{\zeta:y(\zeta)=y'(\zeta)\} D = { ζ : y ( ζ ) = y ′ ( ζ )} is Borel: it is the complement of { ζ : ∥ y ( ζ ) − y ′ ( ζ ) ∥ 2 > 0 } \{\zeta:\lVert y(\zeta)-y'(\zeta)\rVert^{2}>0\} { ζ : ∥ y ( ζ ) − y ′ ( ζ ) ∥ 2 > 0 } , the function ζ ↦ ∥ y ( ζ ) − y ′ ( ζ ) ∥ 2 \zeta\mapsto\lVert y(\zeta)-y'(\zeta)\rVert^{2} ζ ↦ ∥ y ( ζ ) − y ′ ( ζ ) ∥ 2 being Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions and vanishing exactly on D D D by claim 3 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n , claim 3 of Zero Products and Elementary Identities in a Field and claim 2 of Elementary Identities in a Vector Space . Let h = ( w ∘ y ) 1 D h=(w\circ y)\,\mathbf{1}_{D} h = ( w ∘ y ) 1 D , a nonnegative Borel function on R q + q \mathbb{R}^{q+q} R q + q (claims 1 and 3 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions ), and let θ \theta θ be the measure with density h h h with respect to θ 0 \theta_{0} θ 0 , so that ∫ G d θ = ∫ G h d θ 0 \int G\,d\theta=\int Gh\,d\theta_{0} ∫ G d θ = ∫ G h d θ 0 for every Borel G : R q + q → [ 0 , ∞ ] G:\mathbb{R}^{q+q}\to[0,\infty] G : R q + q → [ 0 , ∞ ] by claim 3 of that lemma.
Two disintegration identities. Let g : R q → [ 0 , ∞ ) g:\mathbb{R}^{q}\to[0,\infty) g : R q → [ 0 , ∞ ) be Borel. We claim the following two identities, referred to below as ( ∗ ) (\ast) ( ∗ ) :
∫ g ( P 1 ζ ) h ( ζ ) θ 0 ( d ζ ) = ∫ g d π 12 and ∫ g ( P 2 ζ ) h ( ζ ) θ 0 ( d ζ ) = ∫ g d π 23 . \int g(P_{1}\zeta)\,h(\zeta)\,\theta_{0}(d\zeta)=\int g\,d\pi_{12}\qquad\text{and}\qquad\int g(P_{2}\zeta)\,h(\zeta)\,\theta_{0}(d\zeta)=\int g\,d\pi_{23}. ∫ g ( P 1 ζ ) h ( ζ ) θ 0 ( d ζ ) = ∫ g d π 12 and ∫ g ( P 2 ζ ) h ( ζ ) θ 0 ( d ζ ) = ∫ g d π 23 .
For the first, the integrand G = ( g ∘ P 1 ) h G=(g\circ P_{1})h G = ( g ∘ P 1 ) h is Borel, so by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §product the function G ∘ ι q , q G\circ\iota^{q,q} G ∘ ι q , q is measurable for B ( R q ) ⊗ B ( R q ) \mathcal{B}(\mathbb{R}^{q})\otimes\mathcal{B}(\mathbb{R}^{q}) B ( R q ) ⊗ B ( R q ) and ∫ G d θ 0 = ∫ G ∘ ι q , q d ( π 12 ⊗ π 23 ) \int G\,d\theta_{0}=\int G\circ\iota^{q,q}\,d(\pi_{12}\otimes\pi_{23}) ∫ G d θ 0 = ∫ G ∘ ι q , q d ( π 12 ⊗ π 23 ) ; here, for z ∈ R q z\in\mathbb{R}^{q} z ∈ R q and z ′ ∈ R q z'\in\mathbb{R}^{q} z ′ ∈ R q , G ( ι q , q ( z , z ′ ) ) = g ( z ) w ( p r 2 z ) 1 [ p r 2 z = p r 1 z ′ ] G(\iota^{q,q}(z,z'))=g(z)\,w(\mathrm{pr}_{2}z)\,\mathbf{1}[\mathrm{pr}_{2}z=\mathrm{pr}_{1}z'] G ( ι q , q ( z , z ′ )) = g ( z ) w ( pr 2 z ) 1 [ pr 2 z = pr 1 z ′ ] by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections . By the Tonelli clause of Tonelli and Fubini Theorems (both factors being probability measures, hence σ \sigma σ -finite), this equals ∫ R q ( ∫ R q g ( z ) w ( p r 2 z ) 1 [ p r 2 z = p r 1 z ′ ] π 23 ( d z ′ ) ) π 12 ( d z ) \int_{\mathbb{R}^{q}}\bigl(\int_{\mathbb{R}^{q}}g(z)w(\mathrm{pr}_{2}z)\mathbf{1}[\mathrm{pr}_{2}z=\mathrm{pr}_{1}z']\,\pi_{23}(dz')\bigr)\pi_{12}(dz) ∫ R q ( ∫ R q g ( z ) w ( pr 2 z ) 1 [ pr 2 z = pr 1 z ′ ] π 23 ( d z ′ ) ) π 12 ( d z ) . For fixed z z z , with y = p r 2 z y=\mathrm{pr}_{2}z y = pr 2 z , the inner integrand is the real constant g ( z ) w ( y ) g(z)w(y) g ( z ) w ( y ) times the indicator of p r 1 − 1 ( { y } ) \mathrm{pr}_{1}^{-1}(\{y\}) pr 1 − 1 ({ y }) , so the inner integral is g ( z ) w ( y ) π 23 ( p r 1 − 1 ( { y } ) ) = g ( z ) w ( y ) ρ ( { y } ) = g ( z ) 1 F + ( y ) g(z)w(y)\pi_{23}(\mathrm{pr}_{1}^{-1}(\{y\}))=g(z)w(y)\rho(\{y\})=g(z)\mathbf{1}_{F_{+}}(y) g ( z ) w ( y ) π 23 ( pr 1 − 1 ({ y })) = g ( z ) w ( y ) ρ ({ y }) = g ( z ) 1 F + ( y ) by claim 1 of the integral theorem, The Integral of an Indicator Function is the Measure of the Set and π 23 ∈ Π ( ρ , ν ) \pi_{23}\in\Pi(\rho,\nu) π 23 ∈ Π ( ρ , ν ) . Thus the first integral in ( ∗ ) (\ast) ( ∗ ) equals ∫ g ( z ) 1 F + ( p r 2 z ) π 12 ( d z ) \int g(z)\mathbf{1}_{F_{+}}(\mathrm{pr}_{2}z)\,\pi_{12}(dz) ∫ g ( z ) 1 F + ( pr 2 z ) π 12 ( d z ) . The integrands g g g and g ⋅ ( 1 F + ∘ p r 2 ) g\cdot(\mathbf{1}_{F_{+}}\circ\mathrm{pr}_{2}) g ⋅ ( 1 F + ∘ pr 2 ) agree off the set p r 2 − 1 ( R d ∖ F + ) \mathrm{pr}_{2}^{-1}(\mathbb{R}^{d}\setminus F_{+}) pr 2 − 1 ( R d ∖ F + ) , which has π 12 \pi_{12} π 12 -measure ρ ( R d ∖ F + ) = 0 \rho(\mathbb{R}^{d}\setminus F_{+})=0 ρ ( R d ∖ F + ) = 0 as π 12 ∈ Π ( μ , ρ ) \pi_{12}\in\Pi(\mu,\rho) π 12 ∈ Π ( μ , ρ ) ; so they agree π 12 \pi_{12} π 12 -almost everywhere and their integrals coincide by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §comparison , proving the first identity. The second is proved in the same way with the order of integration in Tonelli's theorem reversed: for fixed z ′ z' z ′ , with y ′ = p r 1 z ′ y'=\mathrm{pr}_{1}z' y ′ = pr 1 z ′ , the inner integrand g ( z ′ ) w ( p r 2 z ) 1 [ p r 2 z = y ′ ] g(z')w(\mathrm{pr}_{2}z)\mathbf{1}[\mathrm{pr}_{2}z=y'] g ( z ′ ) w ( pr 2 z ) 1 [ pr 2 z = y ′ ] equals g ( z ′ ) w ( y ′ ) g(z')w(y') g ( z ′ ) w ( y ′ ) times the indicator of p r 2 − 1 ( { y ′ } ) \mathrm{pr}_{2}^{-1}(\{y'\}) pr 2 − 1 ({ y ′ }) , whose π 12 \pi_{12} π 12 -integral is g ( z ′ ) w ( y ′ ) ρ ( { y ′ } ) = g ( z ′ ) 1 F + ( y ′ ) g(z')w(y')\rho(\{y'\})=g(z')\mathbf{1}_{F_{+}}(y') g ( z ′ ) w ( y ′ ) ρ ({ y ′ }) = g ( z ′ ) 1 F + ( y ′ ) , and the outer integral is ∫ g d π 23 \int g\,d\pi_{23} ∫ g d π 23 by the same almost-everywhere argument with π 23 ∈ Π ( ρ , ν ) \pi_{23}\in\Pi(\rho,\nu) π 23 ∈ Π ( ρ , ν ) .
The glued coupling. Taking g g g the constant 1 1 1 in ( ∗ ) (\ast) ( ∗ ) gives θ ( R q + q ) = ∫ h d θ 0 = π 12 ( R q ) = 1 \theta(\mathbb{R}^{q+q})=\int h\,d\theta_{0}=\pi_{12}(\mathbb{R}^{q})=1 θ ( R q + q ) = ∫ h d θ 0 = π 12 ( R q ) = 1 , so θ ∈ P ( R q + q ) \theta\in\mathcal{P}(\mathbb{R}^{q+q}) θ ∈ P ( R q + q ) . Let Φ = ( x , v ) : R q + q → R d + d \Phi=(x,v):\mathbb{R}^{q+q}\to\mathbb{R}^{d+d} Φ = ( x , v ) : R q + q → R d + d , the pairing of the Borel maps x x x and v v v , Borel by Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pairs , and put π 13 = Φ # θ ∈ P ( R d + d ) \pi_{13}=\Phi_{\#}\theta\in\mathcal{P}(\mathbb{R}^{d+d}) π 13 = Φ # θ ∈ P ( R d + d ) . For A ∈ B ( R d ) A\in\mathcal{B}(\mathbb{R}^{d}) A ∈ B ( R d ) , by (i), (ii) and the density formula, ( p r 1 ) # π 13 ( A ) = θ ( x − 1 ( A ) ) = ∫ 1 x − 1 ( A ) h d θ 0 (\mathrm{pr}_{1})_{\#}\pi_{13}(A)=\theta(x^{-1}(A))=\int\mathbf{1}_{x^{-1}(A)}h\,d\theta_{0} ( pr 1 ) # π 13 ( A ) = θ ( x − 1 ( A )) = ∫ 1 x − 1 ( A ) h d θ 0 , and 1 x − 1 ( A ) = g ∘ P 1 \mathbf{1}_{x^{-1}(A)}=g\circ P_{1} 1 x − 1 ( A ) = g ∘ P 1 with g = 1 p r 1 − 1 ( A ) g=\mathbf{1}_{\mathrm{pr}_{1}^{-1}(A)} g = 1 pr 1 − 1 ( A ) , a Borel map R q → [ 0 , ∞ ) \mathbb{R}^{q}\to[0,\infty) R q → [ 0 , ∞ ) ; so by ( ∗ ) (\ast) ( ∗ ) this equals π 12 ( p r 1 − 1 ( A ) ) = μ ( A ) \pi_{12}(\mathrm{pr}_{1}^{-1}(A))=\mu(A) π 12 ( pr 1 − 1 ( A )) = μ ( A ) . Likewise ( p r 2 ) # π 13 ( C ) = ∫ ( 1 p r 2 − 1 ( C ) ∘ P 2 ) h d θ 0 = π 23 ( p r 2 − 1 ( C ) ) = ν ( C ) (\mathrm{pr}_{2})_{\#}\pi_{13}(C)=\int(\mathbf{1}_{\mathrm{pr}_{2}^{-1}(C)}\circ P_{2})h\,d\theta_{0}=\pi_{23}(\mathrm{pr}_{2}^{-1}(C))=\nu(C) ( pr 2 ) # π 13 ( C ) = ∫ ( 1 pr 2 − 1 ( C ) ∘ P 2 ) h d θ 0 = π 23 ( pr 2 − 1 ( C )) = ν ( C ) . Hence π 13 ∈ Π ( μ , ν ) \pi_{13}\in\Pi(\mu,\nu) π 13 ∈ Π ( μ , ν ) .
The cost bound. Regard ( R q + q , B ( R q + q ) , θ ) (\mathbb{R}^{q+q},\mathcal{B}(\mathbb{R}^{q+q}),\theta) ( R q + q , B ( R q + q ) , θ ) as a probability space, and let f = φ 0 ∘ P 1 f=\varphi_{0}\circ P_{1} f = φ 0 ∘ P 1 and f ′ = φ 0 ∘ P 2 f'=\varphi_{0}\circ P_{2} f ′ = φ 0 ∘ P 2 , Borel, so that f ( ζ ) = ∥ x ( ζ ) − y ( ζ ) ∥ f(\zeta)=\lVert x(\zeta)-y(\zeta)\rVert f ( ζ ) = ∥ x ( ζ ) − y ( ζ )∥ and f ′ ( ζ ) = ∥ y ′ ( ζ ) − v ( ζ ) ∥ f'(\zeta)=\lVert y'(\zeta)-v(\zeta)\rVert f ′ ( ζ ) = ∥ y ′ ( ζ ) − v ( ζ )∥ by (ii). By the density formula and ( ∗ ) (\ast) ( ∗ ) with g = φ g=\varphi g = φ , ∫ f 2 d θ = ∫ ( φ ∘ P 1 ) h d θ 0 = ∫ φ d π 12 = I ( π 12 ) < ∞ \int f^{2}\,d\theta=\int(\varphi\circ P_{1})h\,d\theta_{0}=\int\varphi\,d\pi_{12}=I(\pi_{12})<\infty ∫ f 2 d θ = ∫ ( φ ∘ P 1 ) h d θ 0 = ∫ φ d π 12 = I ( π 12 ) < ∞ , and similarly ∫ f ′ 2 d θ = I ( π 23 ) < ∞ \int f'^{2}\,d\theta=I(\pi_{23})<\infty ∫ f ′ 2 d θ = I ( π 23 ) < ∞ ; so f , f ′ f,f' f , f ′ are square-integrable random variables on this probability space with ∥ f ∥ 2 = I ( π 12 ) \lVert f\rVert_{2}=\sqrt{I(\pi_{12})} ∥ f ∥ 2 = I ( π 12 ) and ∥ f ′ ∥ 2 = I ( π 23 ) \lVert f'\rVert_{2}=\sqrt{I(\pi_{23})} ∥ f ′ ∥ 2 = I ( π 23 ) . Let u ( ζ ) = ∥ x ( ζ ) − v ( ζ ) ∥ 2 u(\zeta)=\lVert x(\zeta)-v(\zeta)\rVert^{2} u ( ζ ) = ∥ x ( ζ ) − v ( ζ ) ∥ 2 , Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions . For ζ ∈ D \zeta\in D ζ ∈ D one has y ( ζ ) = y ′ ( ζ ) y(\zeta)=y'(\zeta) y ( ζ ) = y ′ ( ζ ) , so x − v = ( x − y ) + ( y ′ − v ) x-v=(x-y)+(y'-v) x − v = ( x − y ) + ( y ′ − v ) at ζ \zeta ζ by Euclidean Space R n \mathbb{R}^n R n is a Real Vector Space , whence ∥ x ( ζ ) − v ( ζ ) ∥ ≤ f ( ζ ) + f ′ ( ζ ) \lVert x(\zeta)-v(\zeta)\rVert\le f(\zeta)+f'(\zeta) ∥ x ( ζ ) − v ( ζ )∥ ≤ f ( ζ ) + f ′ ( ζ ) by claim 6 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n and u ( ζ ) ≤ ( f ( ζ ) + f ′ ( ζ ) ) 2 u(\zeta)\le(f(\zeta)+f'(\zeta))^{2} u ( ζ ) ≤ ( f ( ζ ) + f ′ ( ζ ) ) 2 by claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field ; for ζ ∉ D \zeta\notin D ζ ∈ / D , h ( ζ ) = 0 h(\zeta)=0 h ( ζ ) = 0 . Hence u h ≤ ( f + f ′ ) 2 h uh\le(f+f')^{2}h u h ≤ ( f + f ′ ) 2 h pointwise, and by change of variables, (ii), the density formula, monotonicity and claim 2 of Cauchy-Schwarz and Triangle Inequalities for the Mean-Square Norm ,
I ( π 13 ) = ∫ φ ∘ Φ d θ = ∫ u d θ = ∫ u h d θ 0 ≤ ∫ ( f + f ′ ) 2 h d θ 0 = ∫ ( f + f ′ ) 2 d θ = ∥ f + f ′ ∥ 2 2 ≤ ( ∥ f ∥ 2 + ∥ f ′ ∥ 2 ) 2 . I(\pi_{13})=\int\varphi\circ\Phi\,d\theta=\int u\,d\theta=\int uh\,d\theta_{0}\le\int(f+f')^{2}h\,d\theta_{0}=\int(f+f')^{2}\,d\theta=\lVert f+f'\rVert_{2}^{2}\le\bigl(\lVert f\rVert_{2}+\lVert f'\rVert_{2}\bigr)^{2}. I ( π 13 ) = ∫ φ ∘ Φ d θ = ∫ u d θ = ∫ u h d θ 0 ≤ ∫ ( f + f ′ ) 2 h d θ 0 = ∫ ( f + f ′ ) 2 d θ = ∥ f + f ′ ∥ 2 2 ≤ ( ∥ f ∥ 2 + ∥ f ′ ∥ 2 ) 2 .
So I ( π 13 ) I(\pi_{13}) I ( π 13 ) is finite and I ( π 13 ) ≤ ∥ f ∥ 2 + ∥ f ′ ∥ 2 = I ( π 12 ) + I ( π 23 ) \sqrt{I(\pi_{13})}\le\lVert f\rVert_{2}+\lVert f'\rVert_{2}=\sqrt{I(\pi_{12})}+\sqrt{I(\pi_{23})} I ( π 13 ) ≤ ∥ f ∥ 2 + ∥ f ′ ∥ 2 = I ( π 12 ) + I ( π 23 ) by claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field .
Claim 8. A Lipschitz f f f is continuous by A Lipschitz Map is Uniformly Continuous , hence Borel by Probability Measures on Euclidean Space and Random Vectors: Standing Notation §borel-maps , and being bounded it is integrable with respect to μ \mu μ and to ν \nu ν by Probability Measures on Euclidean Space and Random Vectors: Standing Notation §measures . By change of variables in the integrable case, f ∘ p r 1 f\circ\mathrm{pr}_{1} f ∘ pr 1 and f ∘ p r 2 f\circ\mathrm{pr}_{2} f ∘ pr 2 are π \pi π -integrable with ∫ f ∘ p r 1 d π = ∫ f d μ \int f\circ\mathrm{pr}_{1}\,d\pi=\int f\,d\mu ∫ f ∘ pr 1 d π = ∫ f d μ and ∫ f ∘ p r 2 d π = ∫ f d ν \int f\circ\mathrm{pr}_{2}\,d\pi=\int f\,d\nu ∫ f ∘ pr 2 d π = ∫ f d ν . For every z z z , ∣ f ( p r 1 ( z ) ) − f ( p r 2 ( z ) ) ∣ ≤ L φ 0 ( z ) |f(\mathrm{pr}_{1}(z))-f(\mathrm{pr}_{2}(z))|\le L\,\varphi_{0}(z) ∣ f ( pr 1 ( z )) − f ( pr 2 ( z )) ∣ ≤ L φ 0 ( z ) by the Lipschitz property. By claim 2 of the integral theorem, then monotonicity and homogeneity of the nonnegative integral,
∣ ∫ f d μ − ∫ f d ν ∣ = ∣ ∫ ( f ∘ p r 1 − f ∘ p r 2 ) d π ∣ ≤ ∫ ∣ f ∘ p r 1 − f ∘ p r 2 ∣ d π ≤ L ∫ φ 0 d π . \Bigl|\int f\,d\mu-\int f\,d\nu\Bigr|=\Bigl|\int(f\circ\mathrm{pr}_{1}-f\circ\mathrm{pr}_{2})\,d\pi\Bigr|\le\int|f\circ\mathrm{pr}_{1}-f\circ\mathrm{pr}_{2}|\,d\pi\le L\int\varphi_{0}\,d\pi . ∫ f d μ − ∫ f d ν = ∫ ( f ∘ pr 1 − f ∘ pr 2 ) d π ≤ ∫ ∣ f ∘ pr 1 − f ∘ pr 2 ∣ d π ≤ L ∫ φ 0 d π .
Finally, on the probability space ( R d + d , B ( R d + d ) , π ) (\mathbb{R}^{d+d},\mathcal{B}(\mathbb{R}^{d+d}),\pi) ( R d + d , B ( R d + d ) , π ) the random variable φ 0 \varphi_{0} φ 0 is square-integrable with ∥ φ 0 ∥ 2 = I ( π ) \lVert\varphi_{0}\rVert_{2}=\sqrt{I(\pi)} ∥ φ 0 ∥ 2 = I ( π ) , as φ 0 2 = φ \varphi_{0}^{2}=\varphi φ 0 2 = φ , and the constant 1 1 1 is square-integrable with ∥ 1 ∥ 2 = 1 \lVert1\rVert_{2}=1 ∥ 1 ∥ 2 = 1 (The Integral of an Indicator Function is the Measure of the Set with A = R d + d A=\mathbb{R}^{d+d} A = R d + d ); so ∫ φ 0 d π = E π [ φ 0 ⋅ 1 ] ≤ ∥ φ 0 ∥ 2 ∥ 1 ∥ 2 = I ( π ) \int\varphi_{0}\,d\pi=\mathbb{E}_{\pi}[\varphi_{0}\cdot1]\le\lVert\varphi_{0}\rVert_{2}\lVert1\rVert_{2}=\sqrt{I(\pi)} ∫ φ 0 d π = E π [ φ 0 ⋅ 1 ] ≤ ∥ φ 0 ∥ 2 ∥ 1 ∥ 2 = I ( π ) by claim 1 of Cauchy-Schwarz and Triangle Inequalities for the Mean-Square Norm , the expectation being the nonnegative integral by the identification recorded at the start. Combining, ∣ ∫ f d μ − ∫ f d ν ∣ ≤ L I ( π ) |\int f\,d\mu-\int f\,d\nu|\le L\sqrt{I(\pi)} ∣ ∫ f d μ − ∫ f d ν ∣ ≤ L I ( π ) by claim 5 of Elementary Arithmetic in an Ordered Field .