Each result cited is universally quantified over the data in its own statement. Throughout, n ∈ N n\in\mathbb{N} n ∈ N is fixed and the notation is that of A Diagonal Gaussian Reference Measure on the Noise Wasserstein Space, Rescaled Heads and Gaussian Tails: Standing Notation ; R n × X \mathbb{R}^{n}\times X R n × X carries B ( R n ) ⊗ B ( X ) \mathcal{B}(\mathbb{R}^{n})\otimes\mathcal{B}(X) B ( R n ) ⊗ B ( X ) , its Borel σ \sigma σ -algebra by A Diagonal Gaussian Reference Measure on the Noise Wasserstein Space, Rescaled Heads and Gaussian Tails: Standing Notation §tails , and L n \mathcal{L}^{n} L n denotes Lebesgue measure on B ( R n ) \mathcal{B}(\mathbb{R}^{n}) B ( R n ) , written λ d \lambda_{d} λ d with d = n d=n d = n in Change of Variables for the Lebesgue Integral under a Continuously Differentiable Bijection of Euclidean Space with Symmetric Positive Definite Jacobian Matrix, and the Density of a Push-Forward . Push-forwards are the image measures of claim 1 of Image Measures, Measures with Densities, and Change of Variables , and integrals against them are computed by claim 2 of that lemma, which we call the transfer formula . Composites of measurable maps are measurable, preimages composing (Measurable Function and Real-Valued Measurable Function ); for the same reason ( T ∘ S ) # ν = T # ( S # ν ) (T\circ S)_{\#}\nu=T_{\#}(S_{\#}\nu) ( T ∘ S ) # ν = T # ( S # ν ) whenever S S S and T T T are measurable and composable. Densities are those of The Radon-Nikodym Theorem for a Finite Measure and a Sigma-Finite Measure, and Uniqueness of Densities , and ϕ ( s ) = s log s \phi(s)=s\log s ϕ ( s ) = s log s is the function used in Relative Entropy of Probability Measures §relative-entropy .
Step 0 (Preliminaries). (P1) For u ∈ R n u\in\mathbb{R}^{n} u ∈ R n put δ n − 1 ( u ) = ( a 1 1 / 2 u 1 , … , a n 1 / 2 u n ) \delta_{n}^{-1}(u)=(a_{1}^{1/2}u_{1},\dots,a_{n}^{1/2}u_{n}) δ n − 1 ( u ) = ( a 1 1/2 u 1 , … , a n 1/2 u n ) ; since a k 1 / 2 a k − 1 / 2 = 1 a_{k}^{1/2}a_{k}^{-1/2}=1 a k 1/2 a k − 1/2 = 1 , this is the inverse map of the bijection δ n \delta_{n} δ n of A Diagonal Gaussian Reference Measure on the Noise Wasserstein Space, Rescaled Heads and Gaussian Tails: Standing Notation §heads . With the synthesis map p n ∗ p_{n}^{*} p n ∗ of Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §coordinates , the map Ψ n \Psi_{n} Ψ n of A Diagonal Gaussian Reference Measure on the Noise Wasserstein Space, Rescaled Heads and Gaussian Tails: Standing Notation §tails reads
Ψ n ( u , w ) = p n ∗ ( δ n − 1 ( u ) ) + w ( u ∈ R n , w ∈ X ) , \Psi_{n}(u,w)=p_{n}^{*}\bigl(\delta_{n}^{-1}(u)\bigr)+w\qquad(u\in\mathbb{R}^{n},\ w\in X), Ψ n ( u , w ) = p n ∗ ( δ n − 1 ( u ) ) + w ( u ∈ R n , w ∈ X ) ,
and, by linearity of the inner product and orthonormality of ( e k ) k ∈ N (e_{k})_{k\in\mathbb{N}} ( e k ) k ∈ N , its coordinates are Ψ n ( u , w ) k = a k 1 / 2 u k + w k \Psi_{n}(u,w)_{k}=a_{k}^{1/2}u_{k}+w_{k} Ψ n ( u , w ) k = a k 1/2 u k + w k for k ∈ [ n ] k\in[n] k ∈ [ n ] and Ψ n ( u , w ) k = w k \Psi_{n}(u,w)_{k}=w_{k} Ψ n ( u , w ) k = w k for k > n k>n k > n . The map Ψ n \Psi_{n} Ψ n is Borel by The Gaussian-Tail Extension of a Probability Measure on the Rescaled Head §borel .
(P2) The maps p n p_{n} p n and p n ∗ p_{n}^{*} p n ∗ are linear, p n ( p n ∗ ( y ) ) = y p_{n}(p_{n}^{*}(y))=y p n ( p n ∗ ( y )) = y for y ∈ R n y\in\mathbb{R}^{n} y ∈ R n by Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §continuity , and Q n x = x − p n ∗ ( p n ( x ) ) Q_{n}x=x-p_{n}^{*}(p_{n}(x)) Q n x = x − p n ∗ ( p n ( x )) by Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §coordinates . Hence Q n Q_{n} Q n is linear, p n ( Q n x ) = p n ( x ) − p n ( x ) = 0 p_{n}(Q_{n}x)=p_{n}(x)-p_{n}(x)=0 p n ( Q n x ) = p n ( x ) − p n ( x ) = 0 for x ∈ X x\in X x ∈ X , and Q n ( p n ∗ ( y ) ) = p n ∗ ( y ) − p n ∗ ( y ) = 0 Q_{n}(p_{n}^{*}(y))=p_{n}^{*}(y)-p_{n}^{*}(y)=0 Q n ( p n ∗ ( y )) = p n ∗ ( y ) − p n ∗ ( y ) = 0 for y ∈ R n y\in\mathbb{R}^{n} y ∈ R n . Let X 0 = { w ∈ X : p n ( w ) = 0 } X_{0}=\{w\in X:p_{n}(w)=0\} X 0 = { w ∈ X : p n ( w ) = 0 } , a Borel set as the preimage under the Borel map p n p_{n} p n (Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §continuity ) of the Borel set { 0 } \{0\} { 0 } (Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §finite-sets ). For w ∈ X 0 w\in X_{0} w ∈ X 0 one has w k = 0 w_{k}=0 w k = 0 for k ∈ [ n ] k\in[n] k ∈ [ n ] and Q n w = w − p n ∗ ( 0 ) = w Q_{n}w=w-p_{n}^{*}(0)=w Q n w = w − p n ∗ ( 0 ) = w . Since Q n x ∈ X 0 Q_{n}x\in X_{0} Q n x ∈ X 0 for every x ∈ X x\in X x ∈ X , we get Q n ( Q n x ) = Q n x Q_{n}(Q_{n}x)=Q_{n}x Q n ( Q n x ) = Q n x and τ n ( X 0 ) = γ c ( Q n − 1 ( X 0 ) ) = γ c ( X ) = 1 \tau_{n}(X_{0})=\gamma_{c}(Q_{n}^{-1}(X_{0}))=\gamma_{c}(X)=1 τ n ( X 0 ) = γ c ( Q n − 1 ( X 0 )) = γ c ( X ) = 1 , so τ n ( X ∖ X 0 ) = 0 \tau_{n}(X\setminus X_{0})=0 τ n ( X ∖ X 0 ) = 0 .
(P3) For u ∈ R n u\in\mathbb{R}^{n} u ∈ R n and w ∈ X w\in X w ∈ X , (P1) and (P2) give Q n ( Ψ n ( u , w ) ) = Q n w Q_{n}(\Psi_{n}(u,w))=Q_{n}w Q n ( Ψ n ( u , w )) = Q n w and p n ( Ψ n ( u , w ) ) = δ n − 1 ( u ) + p n ( w ) p_{n}(\Psi_{n}(u,w))=\delta_{n}^{-1}(u)+p_{n}(w) p n ( Ψ n ( u , w )) = δ n − 1 ( u ) + p n ( w ) . Since r n = δ n ∘ p n r_{n}=\delta_{n}\circ p_{n} r n = δ n ∘ p n by A Diagonal Gaussian Reference Measure on the Noise Wasserstein Space, Rescaled Heads and Gaussian Tails: Standing Notation §heads , for every u ∈ R n u\in\mathbb{R}^{n} u ∈ R n and w ∈ X 0 w\in X_{0} w ∈ X 0
r n ( Ψ n ( u , w ) ) = u , Q n ( Ψ n ( u , w ) ) = w . r_{n}\bigl(\Psi_{n}(u,w)\bigr)=u,\qquad Q_{n}\bigl(\Psi_{n}(u,w)\bigr)=w . r n ( Ψ n ( u , w ) ) = u , Q n ( Ψ n ( u , w ) ) = w .
(P4) Let ( Y , Y ) (Y,\mathcal{Y}) ( Y , Y ) be a measurable space (below it is ( R n , B ( R n ) ) (\mathbb{R}^{n},\mathcal{B}(\mathbb{R}^{n})) ( R n , B ( R n )) , except in Step 12, where it is ( R n + n , B ( R n + n ) ) (\mathbb{R}^{n+n},\mathcal{B}(\mathbb{R}^{n+n})) ( R n + n , B ( R n + n )) ), let α \alpha α be a probability measure on Y \mathcal{Y} Y and β \beta β one on B ( X ) \mathcal{B}(X) B ( X ) ; being finite they are σ \sigma σ -finite, and by Existence and Uniqueness of the Product Measure α ⊗ β \alpha\otimes\beta α ⊗ β is the only measure on Y ⊗ B ( X ) \mathcal{Y}\otimes\mathcal{B}(X) Y ⊗ B ( X ) with ( α ⊗ β ) ( A × B ) = α ( A ) β ( B ) (\alpha\otimes\beta)(A\times B)=\alpha(A)\beta(B) ( α ⊗ β ) ( A × B ) = α ( A ) β ( B ) for all A ∈ Y A\in\mathcal{Y} A ∈ Y , B ∈ B ( X ) B\in\mathcal{B}(X) B ∈ B ( X ) . We call this the rectangle principle : a measure on that σ \sigma σ -algebra taking these values on rectangles is α ⊗ β \alpha\otimes\beta α ⊗ β . The coordinate projections p r Y \mathrm{pr}_{Y} pr Y and p r X \mathrm{pr}_{X} pr X of Y × X Y\times X Y × X are measurable by claim 5 of Sections of Product-Measurable Sets and Maps Are Measurable, and Insertion Maps into Products Are Measurable , and ( p r Y ) # ( α ⊗ β ) = α (\mathrm{pr}_{Y})_{\#}(\alpha\otimes\beta)=\alpha ( pr Y ) # ( α ⊗ β ) = α , ( p r X ) # ( α ⊗ β ) = β (\mathrm{pr}_{X})_{\#}(\alpha\otimes\beta)=\beta ( pr X ) # ( α ⊗ β ) = β , since ( α ⊗ β ) ( A × X ) = α ( A ) β ( X ) = α ( A ) (\alpha\otimes\beta)(A\times X)=\alpha(A)\beta(X)=\alpha(A) ( α ⊗ β ) ( A × X ) = α ( A ) β ( X ) = α ( A ) and ( α ⊗ β ) ( Y × B ) = α ( Y ) β ( B ) = β ( B ) (\alpha\otimes\beta)(Y\times B)=\alpha(Y)\beta(B)=\beta(B) ( α ⊗ β ) ( Y × B ) = α ( Y ) β ( B ) = β ( B ) ; for Y = R n Y=\mathbb{R}^{n} Y = R n we write p r R n \mathrm{pr}_{\mathbb{R}^{n}} pr R n for p r Y \mathrm{pr}_{Y} pr Y . The set Y × ( X ∖ X 0 ) Y\times(X\setminus X_{0}) Y × ( X ∖ X 0 ) has ( α ⊗ τ n ) (\alpha\otimes\tau_{n}) ( α ⊗ τ n ) -measure α ( Y ) τ n ( X ∖ X 0 ) = 0 \alpha(Y)\,\tau_{n}(X\setminus X_{0})=0 α ( Y ) τ n ( X ∖ X 0 ) = 0 by (P2); thus a property holding at every ( y , w ) ∈ Y × X (y,w)\in Y\times X ( y , w ) ∈ Y × X with w ∈ X 0 w\in X_{0} w ∈ X 0 holds ( α ⊗ τ n ) (\alpha\otimes\tau_{n}) ( α ⊗ τ n ) -almost everywhere.
(P5) Let α , α ′ \alpha,\alpha' α , α ′ be probability measures on B ( R n ) \mathcal{B}(\mathbb{R}^{n}) B ( R n ) and β , β ′ \beta,\beta' β , β ′ probability measures on B ( X ) \mathcal{B}(X) B ( X ) , let f f f be a density of α ′ \alpha' α ′ with respect to α \alpha α and g g g a density of β ′ \beta' β ′ with respect to β \beta β . Then F ( u , w ) = f ( u ) g ( w ) F(u,w)=f(u)g(w) F ( u , w ) = f ( u ) g ( w ) is a density of α ′ ⊗ β ′ \alpha'\otimes\beta' α ′ ⊗ β ′ with respect to α ⊗ β \alpha\otimes\beta α ⊗ β . Indeed, F F F is nonnegative and measurable, as the product (claim 3 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions ) of f ∘ p r R n f\circ\mathrm{pr}_{\mathbb{R}^{n}} f ∘ pr R n and g ∘ p r X g\circ\mathrm{pr}_{X} g ∘ pr X ; let ν \nu ν be the measure with density F F F with respect to α ⊗ β \alpha\otimes\beta α ⊗ β (claim 3 of Image Measures, Measures with Densities, and Change of Variables ). For a rectangle A × B A\times B A × B , Tonelli's theorem , applied to the nonnegative measurable function 1 A × B F \mathbf{1}_{A\times B}F 1 A × B F , gives
ν ( A × B ) = ∫ R n 1 A ( u ) f ( u ) ( ∫ X 1 B g d β ) α ( d u ) = ∫ R n 1 A ( u ) f ( u ) β ′ ( B ) α ( d u ) = β ′ ( B ) ∫ R n 1 A f d α = α ′ ( A ) β ′ ( B ) : \nu(A\times B)=\int_{\mathbb{R}^{n}}\mathbf{1}_{A}(u)f(u)\Bigl(\int_{X}\mathbf{1}_{B}\,g\,d\beta\Bigr)\alpha(du)=\int_{\mathbb{R}^{n}}\mathbf{1}_{A}(u)f(u)\,\beta'(B)\,\alpha(du)=\beta'(B)\int_{\mathbb{R}^{n}}\mathbf{1}_{A}\,f\,d\alpha=\alpha'(A)\,\beta'(B): ν ( A × B ) = ∫ R n 1 A ( u ) f ( u ) ( ∫ X 1 B g d β ) α ( d u ) = ∫ R n 1 A ( u ) f ( u ) β ′ ( B ) α ( d u ) = β ′ ( B ) ∫ R n 1 A f d α = α ′ ( A ) β ′ ( B ) :
the inner integral is β ′ ( B ) \beta'(B) β ′ ( B ) because g g g is a density of β ′ \beta' β ′ with respect to β \beta β ; the constant β ′ ( B ) ∈ [ 0 , 1 ] \beta'(B)\in[0,1] β ′ ( B ) ∈ [ 0 , 1 ] is taken out of the outer integral by the homogeneity of the integral of nonnegative functions (Linearity and Monotonicity of the Lebesgue Integral §nonnegative ); and ∫ 1 A f d α = α ′ ( A ) \int\mathbf{1}_{A}f\,d\alpha=\alpha'(A) ∫ 1 A f d α = α ′ ( A ) because f f f is a density of α ′ \alpha' α ′ with respect to α \alpha α . So ν = α ′ ⊗ β ′ \nu=\alpha'\otimes\beta' ν = α ′ ⊗ β ′ by the rectangle principle. The constant function 1 1 1 is a density of every probability measure with respect to itself.
(P6) Let N , n ′ ∈ N N,n'\in\mathbb{N} N , n ′ ∈ N , ψ ∈ C b 1 ( R N ) \psi\in C^{1}_{b}(\mathbb{R}^{N}) ψ ∈ C b 1 ( R N ) (Bounded Continuously Differentiable Functions with Bounded Partial Derivatives on Euclidean Space §bounded ), and let A : R n ′ → R N A:\mathbb{R}^{n'}\to\mathbb{R}^{N} A : R n ′ → R N have components A ( u ) j = s j + β j u j A(u)_{j}=s_{j}+\beta_{j}u_{j} A ( u ) j = s j + β j u j for j ≤ min ( n ′ , N ) j\le\min(n',N) j ≤ min ( n ′ , N ) and A ( u ) j = s j A(u)_{j}=s_{j} A ( u ) j = s j for n ′ < j ≤ N n'<j\le N n ′ < j ≤ N , with real constants s j , β j s_{j},\beta_{j} s j , β j . Then ψ ∘ A ∈ C b 1 ( R n ′ ) \psi\circ A\in C^{1}_{b}(\mathbb{R}^{n'}) ψ ∘ A ∈ C b 1 ( R n ′ ) , with ∂ k ( ψ ∘ A ) ( u ) = β k ∂ k ψ ( A ( u ) ) \partial_{k}(\psi\circ A)(u)=\beta_{k}\,\partial_{k}\psi(A(u)) ∂ k ( ψ ∘ A ) ( u ) = β k ∂ k ψ ( A ( u )) for k ≤ min ( n ′ , N ) k\le\min(n',N) k ≤ min ( n ′ , N ) and ∂ k ( ψ ∘ A ) ( u ) = 0 \partial_{k}(\psi\circ A)(u)=0 ∂ k ( ψ ∘ A ) ( u ) = 0 for N < k ≤ n ′ N<k\le n' N < k ≤ n ′ . Indeed, R n ′ \mathbb{R}^{n'} R n ′ and R N \mathbb{R}^{N} R N are open by claim 1 of Euclidean Space is Open in Itself, and C k C^k C k Maps are Continuous . For j ∈ [ N ] j\in[N] j ∈ [ N ] the component A j A_{j} A j has the form u ↦ q ( j ) ⋅ u + s j u\mapsto q^{(j)}\cdot u+s_{j} u ↦ q ( j ) ⋅ u + s j , where q ( j ) ∈ R n ′ q^{(j)}\in\mathbb{R}^{n'} q ( j ) ∈ R n ′ is β j \beta_{j} β j times the j j j -th standard basis vector of R n ′ \mathbb{R}^{n'} R n ′ if j ≤ min ( n ′ , N ) j\le\min(n',N) j ≤ min ( n ′ , N ) , and q ( j ) = 0 q^{(j)}=0 q ( j ) = 0 if j > n ′ j>n' j > n ′ . By Quadratic and Affine Functions of Class C 2 C^2 C 2 , Translation, and Quadratic Perturbation of Semiconvexity §quadratic , applied in dimension n ′ n' n ′ with M = 0 n ′ M=0_{n'} M = 0 n ′ , q = q ( j ) q=q^{(j)} q = q ( j ) and c = s j c=s_{j} c = s j , the function A j A_{j} A j is of class C 2 C^{2} C 2 on R n ′ \mathbb{R}^{n'} R n ′ with gradient q ( j ) q^{(j)} q ( j ) at every point; hence it is of class C 1 C^{1} C 1 on R n ′ \mathbb{R}^{n'} R n ′ by claim 2 of Euclidean Space is Open in Itself, and C k C^k C k Maps are Continuous , and by Gradient of a Real-Valued Function on a Euclidean Open Set its partial derivatives are the entries of q ( j ) q^{(j)} q ( j ) : ∂ k A j ( u ) = β j \partial_{k}A_{j}(u)=\beta_{j} ∂ k A j ( u ) = β j if k = j ≤ min ( n ′ , N ) k=j\le\min(n',N) k = j ≤ min ( n ′ , N ) , and ∂ k A j ( u ) = 0 \partial_{k}A_{j}(u)=0 ∂ k A j ( u ) = 0 otherwise. Since clause 1 of C^k Maps on a Euclidean Open Set is a condition on each coordinate function separately, A A A is of class C 1 C^{1} C 1 on R n ′ \mathbb{R}^{n'} R n ′ ; and ψ \psi ψ , read as a map into R 1 \mathbb{R}^{1} R 1 by clause 3 there, is of class C 1 C^{1} C 1 on R N \mathbb{R}^{N} R N . Claim 1 of A Composition of C k C^k C k Maps Between Euclidean Open Sets is of Class C k C^k C k , with U = R n ′ U=\mathbb{R}^{n'} U = R n ′ , V = R N V=\mathbb{R}^{N} V = R N , F = A F=A F = A , G = ψ G=\psi G = ψ and p = 1 p=1 p = 1 , gives for k ∈ [ n ′ ] k\in[n'] k ∈ [ n ′ ] and u ∈ R n ′ u\in\mathbb{R}^{n'} u ∈ R n ′
∂ k ( ψ ∘ A ) ( u ) = ∑ l = 1 N ∂ l ψ ( A ( u ) ) ∂ k A l ( u ) , \partial_{k}(\psi\circ A)(u)=\sum_{l=1}^{N}\partial_{l}\psi\bigl(A(u)\bigr)\,\partial_{k}A_{l}(u), ∂ k ( ψ ∘ A ) ( u ) = l = 1 ∑ N ∂ l ψ ( A ( u ) ) ∂ k A l ( u ) ,
in which every term with l ≠ k l\ne k l = k vanishes, and the term with l = k l=k l = k , present only when k ≤ N k\le N k ≤ N , equals β k ∂ k ψ ( A ( u ) ) \beta_{k}\,\partial_{k}\psi(A(u)) β k ∂ k ψ ( A ( u )) ; so the sum is β k ∂ k ψ ( A ( u ) ) \beta_{k}\,\partial_{k}\psi(A(u)) β k ∂ k ψ ( A ( u )) for k ≤ min ( n ′ , N ) k\le\min(n',N) k ≤ min ( n ′ , N ) and 0 0 0 for N < k ≤ n ′ N<k\le n' N < k ≤ n ′ . Claim 2 of A Composition of C k C^k C k Maps Between Euclidean Open Sets is of Class C k C^k C k , with order 1 1 1 , shows that ψ ∘ A \psi\circ A ψ ∘ A is of class C 1 C^{1} C 1 on R n ′ \mathbb{R}^{n'} R n ′ . Finally, if b 0 b_{0} b 0 bounds ∣ ψ ∣ |\psi| ∣ ψ ∣ and b k b_{k} b k bounds ∣ ∂ k ψ ∣ |\partial_{k}\psi| ∣ ∂ k ψ ∣ for k ∈ [ N ] k\in[N] k ∈ [ N ] (Bounded Continuously Differentiable Functions with Bounded Partial Derivatives on Euclidean Space §bounded ), then b 0 b_{0} b 0 bounds ∣ ψ ∘ A ∣ |\psi\circ A| ∣ ψ ∘ A ∣ , and ∣ β k ∣ b k |\beta_{k}|\,b_{k} ∣ β k ∣ b k , respectively 0 0 0 , bounds ∣ ∂ k ( ψ ∘ A ) ∣ |\partial_{k}(\psi\circ A)| ∣ ∂ k ( ψ ∘ A ) ∣ ; so ψ ∘ A ∈ C b 1 ( R n ′ ) \psi\circ A\in C^{1}_{b}(\mathbb{R}^{n'}) ψ ∘ A ∈ C b 1 ( R n ′ ) .
Step 1 (Claim 1). The components u ↦ a k − 1 / 2 u k u\mapsto a_{k}^{-1/2}u_{k} u ↦ a k − 1/2 u k of δ n \delta_{n} δ n and u ↦ a k 1 / 2 u k u\mapsto a_{k}^{1/2}u_{k} u ↦ a k 1/2 u k of δ n − 1 \delta_{n}^{-1} δ n − 1 are of the form u ↦ q ⋅ u u\mapsto q\cdot u u ↦ q ⋅ u , with q ∈ R n q\in\mathbb{R}^{n} q ∈ R n equal to a k − 1 / 2 a_{k}^{-1/2} a k − 1/2 , respectively a k 1 / 2 a_{k}^{1/2} a k 1/2 , times the k k k -th standard basis vector. The set R n \mathbb{R}^{n} R n is open by claim 1 of Euclidean Space is Open in Itself, and C k C^k C k Maps are Continuous . By Quadratic and Affine Functions of Class C 2 C^2 C 2 , Translation, and Quadratic Perturbation of Semiconvexity §quadratic , applied with M = 0 n M=0_{n} M = 0 n , this q q q and c = 0 c=0 c = 0 , these components are of class C 2 C^{2} C 2 on R n \mathbb{R}^{n} R n with constant gradient q q q , hence of class C 1 C^{1} C 1 on R n \mathbb{R}^{n} R n by claim 2 of Euclidean Space is Open in Itself, and C k C^k C k Maps are Continuous ; by Gradient of a Real-Valued Function on a Euclidean Open Set the entries of the gradient are the partial derivatives, so the partial derivative of the k k k -th component of δ n \delta_{n} δ n with respect to the j j j -th variable is a k − 1 / 2 a_{k}^{-1/2} a k − 1/2 if j = k j=k j = k and 0 0 0 otherwise, and likewise for δ n − 1 \delta_{n}^{-1} δ n − 1 with a k 1 / 2 a_{k}^{1/2} a k 1/2 . As clause 1 of C^k Maps on a Euclidean Open Set is a condition on each coordinate function, δ n \delta_{n} δ n and δ n − 1 \delta_{n}^{-1} δ n − 1 are of class C 1 C^{1} C 1 on R n \mathbb{R}^{n} R n . So the Jacobian matrix of δ n \delta_{n} δ n at every point is the diagonal matrix D D D with diagonal entries a 1 − 1 / 2 , … , a n − 1 / 2 a_{1}^{-1/2},\dots,a_{n}^{-1/2} a 1 − 1/2 , … , a n − 1/2 , and that of δ n − 1 \delta_{n}^{-1} δ n − 1 at every point is the diagonal matrix D ′ D' D ′ with diagonal entries a 1 1 / 2 , … , a n 1 / 2 a_{1}^{1/2},\dots,a_{n}^{1/2} a 1 1/2 , … , a n 1/2 . The matrix D D D is symmetric, x ⋅ ( D x ) = ∑ k = 1 n a k − 1 / 2 x k 2 > 0 x\cdot(Dx)=\sum_{k=1}^{n}a_{k}^{-1/2}x_{k}^{2}>0 x ⋅ ( D x ) = ∑ k = 1 n a k − 1/2 x k 2 > 0 for x ≠ 0 x\ne0 x = 0 , so D D D is positive definite (Symmetric, Positive Semidefinite, and Positive Definite Real Matrices ), and D D ′ = D ′ D = I n DD'=D'D=I_{n} D D ′ = D ′ D = I n , so D ′ D' D ′ is the inverse matrix of D D D (Inverse Matrix and Invertible Real Square Matrix ). Thus Change of Variables for the Lebesgue Integral under a Continuously Differentiable Bijection of Euclidean Space with Symmetric Positive Definite Jacobian Matrix, and the Density of a Push-Forward applies to F = δ n F=\delta_{n} F = δ n ; by its claim 1 δ n \delta_{n} δ n and δ n − 1 \delta_{n}^{-1} δ n − 1 are Borel, and det D = ∏ k = 1 n a k − 1 / 2 \det D=\prod_{k=1}^{n}a_{k}^{-1/2} det D = ∏ k = 1 n a k − 1/2 by claim 1 of The Determinant of a Triangular Matrix is the Product of its Diagonal Entries , a diagonal matrix being lower triangular.
Since r n = δ n ∘ p n r_{n}=\delta_{n}\circ p_{n} r n = δ n ∘ p n and ( p n ) # γ c = γ c ( n ) (p_{n})_{\#}\gamma_{c}=\gamma_{c^{(n)}} ( p n ) # γ c = γ c ( n ) by Diagonal Gaussian Measures on a Hilbert Space §measure , we have ( r n ) # γ c = ( δ n ) # γ c ( n ) (r_{n})_{\#}\gamma_{c}=(\delta_{n})_{\#}\gamma_{c^{(n)}} ( r n ) # γ c = ( δ n ) # γ c ( n ) . The measure γ c ( n ) \gamma_{c^{(n)}} γ c ( n ) has the density ρ c ( n ) \rho_{c^{(n)}} ρ c ( n ) with respect to L n \mathcal{L}^{n} L n by Diagonal Gaussian Measures on Euclidean Space §measure (read with n n n for d d d , Variance Sequences and Their Truncations §truncations ). By Change of Variables for the Lebesgue Integral under a Continuously Differentiable Bijection of Euclidean Space with Symmetric Positive Definite Jacobian Matrix, and the Density of a Push-Forward §densities , the function y ↦ ρ c ( n ) ( δ n − 1 ( y ) ) / det D y\mapsto\rho_{c^{(n)}}(\delta_{n}^{-1}(y))/\det D y ↦ ρ c ( n ) ( δ n − 1 ( y )) / det D is a density of ( δ n ) # γ c ( n ) (\delta_{n})_{\#}\gamma_{c^{(n)}} ( δ n ) # γ c ( n ) with respect to L n \mathcal{L}^{n} L n . We compute it with The Diagonal Gaussian Density on Euclidean Space and Its Notation §density and The Diagonal Gaussian Density on Euclidean Space and Its Notation §scaling . For y ∈ R n y\in\mathbb{R}^{n} y ∈ R n and x = δ n − 1 ( y ) x=\delta_{n}^{-1}(y) x = δ n − 1 ( y ) , x k 2 / c k = a k y k 2 / c k = y k 2 / ( c k / a k ) x_{k}^{2}/c_{k}=a_{k}y_{k}^{2}/c_{k}=y_{k}^{2}/(c_{k}/a_{k}) x k 2 / c k = a k y k 2 / c k = y k 2 / ( c k / a k ) , so ∣ x ∣ c ( n ) 2 = ∣ y ∣ c ~ ( n ) 2 |x|^{2}_{c^{(n)}}=|y|^{2}_{\tilde{c}^{(n)}} ∣ x ∣ c ( n ) 2 = ∣ y ∣ c ~ ( n ) 2 . Next, 1 / det D = ∏ k = 1 n a k 1 / 2 = exp ( ∑ k = 1 n log a k 1 / 2 ) 1/\det D=\prod_{k=1}^{n}a_{k}^{1/2}=\exp\bigl(\sum_{k=1}^{n}\log a_{k}^{1/2}\bigr) 1/ det D = ∏ k = 1 n a k 1/2 = exp ( ∑ k = 1 n log a k 1/2 ) and
Z c ( n ) − ∑ k = 1 n log a k 1 / 2 = ∑ k = 1 n log ( κ c k a k 1 / 2 ) = ∑ k = 1 n log ( κ c k / a k ) = Z c ~ ( n ) , Z_{c^{(n)}}-\sum_{k=1}^{n}\log a_{k}^{1/2}=\sum_{k=1}^{n}\log\Bigl(\frac{\kappa\sqrt{c_{k}}}{a_{k}^{1/2}}\Bigr)=\sum_{k=1}^{n}\log\Bigl(\kappa\sqrt{c_{k}/a_{k}}\Bigr)=Z_{\tilde{c}^{(n)}}, Z c ( n ) − k = 1 ∑ n log a k 1/2 = k = 1 ∑ n log ( a k 1/2 κ c k ) = k = 1 ∑ n log ( κ c k / a k ) = Z c ~ ( n ) ,
by the identities exp ( s + t ) = exp ( s ) exp ( t ) \exp(s+t)=\exp(s)\exp(t) exp ( s + t ) = exp ( s ) exp ( t ) (claim 1 of Basic Properties of the Exponential Function ), exp ( log t ) = t \exp(\log t)=t exp ( log t ) = t and log ( s t ) = log s + log t \log(st)=\log s+\log t log ( s t ) = log s + log t (The Natural Logarithm ); from the last, log ( s / t ) = log s − log t \log(s/t)=\log s-\log t log ( s / t ) = log s − log t for s , t > 0 s,t>0 s , t > 0 , since log ( s / t ) + log t = log ( ( s / t ) t ) = log s \log(s/t)+\log t=\log((s/t)t)=\log s log ( s / t ) + log t = log (( s / t ) t ) = log s , and from the first, by induction on the number of summands, exp ( ∑ k = 1 n x k ) = ∏ k = 1 n exp ( x k ) \exp\bigl(\sum_{k=1}^{n}x_{k}\bigr)=\prod_{k=1}^{n}\exp(x_{k}) exp ( ∑ k = 1 n x k ) = ∏ k = 1 n exp ( x k ) for real x 1 , … , x n x_{1},\dots,x_{n} x 1 , … , x n ; and c k / a k 1 / 2 = c k / a k \sqrt{c_{k}}/a_{k}^{1/2}=\sqrt{c_{k}/a_{k}} c k / a k 1/2 = c k / a k , both sides being nonnegative with square c k / a k c_{k}/a_{k} c k / a k (Existence and Uniqueness of the Nonnegative Square Root ). Hence
ρ c ( n ) ( δ n − 1 ( y ) ) det D = exp ( − 1 2 ∣ y ∣ c ~ ( n ) 2 − Z c ( n ) + ∑ k = 1 n log a k 1 / 2 ) = ρ c ~ ( n ) ( y ) . \frac{\rho_{c^{(n)}}(\delta_{n}^{-1}(y))}{\det D}=\exp\Bigl(-\tfrac12|y|^{2}_{\tilde{c}^{(n)}}-Z_{c^{(n)}}+\sum_{k=1}^{n}\log a_{k}^{1/2}\Bigr)=\rho_{\tilde{c}^{(n)}}(y). det D ρ c ( n ) ( δ n − 1 ( y )) = exp ( − 2 1 ∣ y ∣ c ~ ( n ) 2 − Z c ( n ) + k = 1 ∑ n log a k 1/2 ) = ρ c ~ ( n ) ( y ) .
So for every B ∈ B ( R n ) B\in\mathcal{B}(\mathbb{R}^{n}) B ∈ B ( R n ) , ( δ n ) # γ c ( n ) ( B ) = ∫ 1 B ρ c ~ ( n ) d L n = γ ~ n ( B ) (\delta_{n})_{\#}\gamma_{c^{(n)}}(B)=\int\mathbf{1}_{B}\,\rho_{\tilde{c}^{(n)}}\,d\mathcal{L}^{n}=\tilde{\gamma}_{n}(B) ( δ n ) # γ c ( n ) ( B ) = ∫ 1 B ρ c ~ ( n ) d L n = γ ~ n ( B ) by Diagonal Gaussian Measures on Euclidean Space §measure , that is,
( r n ) # γ c = ( δ n ) # γ c ( n ) = γ ~ n . (r_{n})_{\#}\gamma_{c}=(\delta_{n})_{\#}\gamma_{c^{(n)}}=\tilde{\gamma}_{n}. ( r n ) # γ c = ( δ n ) # γ c ( n ) = γ ~ n .
Finally γ c ∈ P 2 ( X ) \gamma_{c}\in\mathcal{P}_{2}(X) γ c ∈ P 2 ( X ) by Moments of a Diagonal Gaussian Measure on a Hilbert Space: Coordinate Covariances, Finite Second Moment, and Exponential Moments of the Squared Norm §moment , so γ ~ n = ( r n ) # γ c ∈ P 2 ( R n ) \tilde{\gamma}_{n}=(r_{n})_{\#}\gamma_{c}\in\mathcal{P}_{2}(\mathbb{R}^{n}) γ ~ n = ( r n ) # γ c ∈ P 2 ( R n ) by Lifting Euclidean Tangent Fields of the Rescaled Head to Noise Tangent Fields on a Hilbert Space §head . Nothing in this step depended on the particular n n n ; so for every m ∈ N m\in\mathbb{N} m ∈ N , δ m \delta_{m} δ m and δ m − 1 \delta_{m}^{-1} δ m − 1 are Borel and ( r m ) # γ c = ( δ m ) # γ c ( m ) = γ ~ m (r_{m})_{\#}\gamma_{c}=(\delta_{m})_{\#}\gamma_{c^{(m)}}=\tilde{\gamma}_{m} ( r m ) # γ c = ( δ m ) # γ c ( m ) = γ ~ m .
Step 2 (Claim 2). Let Δ n : R n × X → R n × X \Delta_{n}:\mathbb{R}^{n}\times X\to\mathbb{R}^{n}\times X Δ n : R n × X → R n × X , Δ n ( u , w ) = ( δ n ( u ) , w ) \Delta_{n}(u,w)=(\delta_{n}(u),w) Δ n ( u , w ) = ( δ n ( u ) , w ) . Its components δ n ∘ p r R n \delta_{n}\circ\mathrm{pr}_{\mathbb{R}^{n}} δ n ∘ pr R n and p r X \mathrm{pr}_{X} pr X are measurable (Step 1 and (P4)), so Δ n \Delta_{n} Δ n is measurable by claim 1 of Pairings into a Product, the Graph of a Measurable Map, and Couplings Concentrated on a Graph , applied with Y = R n Y=\mathbb{R}^{n} Y = R n and Z = X Z=X Z = X . For rectangles, Δ n − 1 ( A × B ) = δ n − 1 ( A ) × B \Delta_{n}^{-1}(A\times B)=\delta_{n}^{-1}(A)\times B Δ n − 1 ( A × B ) = δ n − 1 ( A ) × B , so by Step 1
( Δ n ) # ( γ c ( n ) ⊗ τ n ) ( A × B ) = γ c ( n ) ( δ n − 1 ( A ) ) τ n ( B ) = γ ~ n ( A ) τ n ( B ) , (\Delta_{n})_{\#}(\gamma_{c^{(n)}}\otimes\tau_{n})(A\times B)=\gamma_{c^{(n)}}(\delta_{n}^{-1}(A))\,\tau_{n}(B)=\tilde{\gamma}_{n}(A)\,\tau_{n}(B), ( Δ n ) # ( γ c ( n ) ⊗ τ n ) ( A × B ) = γ c ( n ) ( δ n − 1 ( A )) τ n ( B ) = γ ~ n ( A ) τ n ( B ) ,
and the rectangle principle (P4) gives ( Δ n ) # ( γ c ( n ) ⊗ τ n ) = γ ~ n ⊗ τ n (\Delta_{n})_{\#}(\gamma_{c^{(n)}}\otimes\tau_{n})=\tilde{\gamma}_{n}\otimes\tau_{n} ( Δ n ) # ( γ c ( n ) ⊗ τ n ) = γ ~ n ⊗ τ n . By (P1), Ψ n ( Δ n ( u , w ) ) = p n ∗ ( δ n − 1 ( δ n ( u ) ) ) + w = p n ∗ ( u ) + w = Φ n ( u , w ) \Psi_{n}(\Delta_{n}(u,w))=p_{n}^{*}(\delta_{n}^{-1}(\delta_{n}(u)))+w=p_{n}^{*}(u)+w=\Phi_{n}(u,w) Ψ n ( Δ n ( u , w )) = p n ∗ ( δ n − 1 ( δ n ( u ))) + w = p n ∗ ( u ) + w = Φ n ( u , w ) , with Φ n \Phi_{n} Φ n the map of Head and Tail of a Diagonal Gaussian Measure on a Hilbert Space are Independent . Since τ n = ( Q n ) # γ c \tau_{n}=(Q_{n})_{\#}\gamma_{c} τ n = ( Q n ) # γ c (A Diagonal Gaussian Reference Measure on the Noise Wasserstein Space, Rescaled Heads and Gaussian Tails: Standing Notation §tails ), The Gaussian-Tail Extension of a Probability Measure on the Rescaled Head §extension , the composition rule for push-forwards and Head and Tail of a Diagonal Gaussian Measure on a Hilbert Space are Independent §synthesis give
E n ( γ ~ n ) = ( Ψ n ) # ( γ ~ n ⊗ τ n ) = ( Ψ n ) # ( ( Δ n ) # ( γ c ( n ) ⊗ τ n ) ) = ( Ψ n ∘ Δ n ) # ( γ c ( n ) ⊗ τ n ) = ( Φ n ) # ( γ c ( n ) ⊗ τ n ) = ( Φ n ) # ( γ c ( n ) ⊗ ( Q n ) # γ c ) = γ c . E_{n}(\tilde{\gamma}_{n})=(\Psi_{n})_{\#}(\tilde{\gamma}_{n}\otimes\tau_{n})=(\Psi_{n})_{\#}\bigl((\Delta_{n})_{\#}(\gamma_{c^{(n)}}\otimes\tau_{n})\bigr)=(\Psi_{n}\circ\Delta_{n})_{\#}(\gamma_{c^{(n)}}\otimes\tau_{n})=(\Phi_{n})_{\#}(\gamma_{c^{(n)}}\otimes\tau_{n})=(\Phi_{n})_{\#}\bigl(\gamma_{c^{(n)}}\otimes(Q_{n})_{\#}\gamma_{c}\bigr)=\gamma_{c}. E n ( γ ~ n ) = ( Ψ n ) # ( γ ~ n ⊗ τ n ) = ( Ψ n ) # ( ( Δ n ) # ( γ c ( n ) ⊗ τ n ) ) = ( Ψ n ∘ Δ n ) # ( γ c ( n ) ⊗ τ n ) = ( Φ n ) # ( γ c ( n ) ⊗ τ n ) = ( Φ n ) # ( γ c ( n ) ⊗ ( Q n ) # γ c ) = γ c .
Step 3 (Claim 3). Let λ ∈ P ( R n ) \lambda\in\mathcal{P}(\mathbb{R}^{n}) λ ∈ P ( R n ) and B ∈ B ( R n ) B\in\mathcal{B}(\mathbb{R}^{n}) B ∈ B ( R n ) , and let G = Ψ n − 1 ( r n − 1 ( B ) ) G=\Psi_{n}^{-1}(r_{n}^{-1}(B)) G = Ψ n − 1 ( r n − 1 ( B )) , a measurable set since r n r_{n} r n is Borel (Lifting Euclidean Tangent Fields of the Rescaled Head to Noise Tangent Fields on a Hilbert Space §head ). By (P3), G ∩ ( R n × X 0 ) = B × X 0 G\cap(\mathbb{R}^{n}\times X_{0})=B\times X_{0} G ∩ ( R n × X 0 ) = B × X 0 , and the rest of G G G lies in the ( λ ⊗ τ n ) (\lambda\otimes\tau_{n}) ( λ ⊗ τ n ) -null set R n × ( X ∖ X 0 ) \mathbb{R}^{n}\times(X\setminus X_{0}) R n × ( X ∖ X 0 ) of (P4). Hence
( r n ) # E n ( λ ) ( B ) = ( λ ⊗ τ n ) ( G ) = ( λ ⊗ τ n ) ( B × X 0 ) = λ ( B ) τ n ( X 0 ) = λ ( B ) . (r_{n})_{\#}E_{n}(\lambda)(B)=(\lambda\otimes\tau_{n})(G)=(\lambda\otimes\tau_{n})(B\times X_{0})=\lambda(B)\,\tau_{n}(X_{0})=\lambda(B). ( r n ) # E n ( λ ) ( B ) = ( λ ⊗ τ n ) ( G ) = ( λ ⊗ τ n ) ( B × X 0 ) = λ ( B ) τ n ( X 0 ) = λ ( B ) .
Step 4 (Claim 4). By (P3), Q n ∘ Ψ n = Q n ∘ p r X Q_{n}\circ\Psi_{n}=Q_{n}\circ\mathrm{pr}_{X} Q n ∘ Ψ n = Q n ∘ pr X on R n × X \mathbb{R}^{n}\times X R n × X . Hence, by The Gaussian-Tail Extension of a Probability Measure on the Rescaled Head §extension , the composition rule for push-forwards, (P4) and (P2),
( Q n ) # E n ( λ ) = ( Q n ∘ Ψ n ) # ( λ ⊗ τ n ) = ( Q n ∘ p r X ) # ( λ ⊗ τ n ) = ( Q n ) # ( ( p r X ) # ( λ ⊗ τ n ) ) = ( Q n ) # τ n = ( Q n ∘ Q n ) # γ c = ( Q n ) # γ c = τ n . (Q_{n})_{\#}E_{n}(\lambda)=(Q_{n}\circ\Psi_{n})_{\#}(\lambda\otimes\tau_{n})=(Q_{n}\circ\mathrm{pr}_{X})_{\#}(\lambda\otimes\tau_{n})=(Q_{n})_{\#}\bigl((\mathrm{pr}_{X})_{\#}(\lambda\otimes\tau_{n})\bigr)=(Q_{n})_{\#}\tau_{n}=(Q_{n}\circ Q_{n})_{\#}\gamma_{c}=(Q_{n})_{\#}\gamma_{c}=\tau_{n}. ( Q n ) # E n ( λ ) = ( Q n ∘ Ψ n ) # ( λ ⊗ τ n ) = ( Q n ∘ pr X ) # ( λ ⊗ τ n ) = ( Q n ) # ( ( pr X ) # ( λ ⊗ τ n ) ) = ( Q n ) # τ n = ( Q n ∘ Q n ) # γ c = ( Q n ) # γ c = τ n .
Step 5 (Claim 5). Let λ ∈ P 2 ( R n ) \lambda\in\mathcal{P}_{2}(\mathbb{R}^{n}) λ ∈ P 2 ( R n ) . For ( u , w ) ∈ R n × X (u,w)\in\mathbb{R}^{n}\times X ( u , w ) ∈ R n × X , (P1), the triangle inequality and ∣ p n ∗ ( y ) ∣ = ∥ y ∥ |p_{n}^{*}(y)|=\lVert y\rVert ∣ p n ∗ ( y ) ∣ = ∥ y ∥ (Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §continuity ) give ∣ Ψ n ( u , w ) ∣ ≤ ∥ δ n − 1 ( u ) ∥ + ∣ w ∣ |\Psi_{n}(u,w)|\le\lVert\delta_{n}^{-1}(u)\rVert+|w| ∣ Ψ n ( u , w ) ∣ ≤ ∥ δ n − 1 ( u )∥ + ∣ w ∣ , and ∥ δ n − 1 ( u ) ∥ 2 = ∑ k = 1 n a k u k 2 ≤ a ˉ ∥ u ∥ 2 \lVert\delta_{n}^{-1}(u)\rVert^{2}=\sum_{k=1}^{n}a_{k}u_{k}^{2}\le\bar{a}\lVert u\rVert^{2} ∥ δ n − 1 ( u ) ∥ 2 = ∑ k = 1 n a k u k 2 ≤ a ˉ ∥ u ∥ 2 with a ˉ \bar{a} a ˉ of Probability Measures on a Hilbert Space Transported in the Noise Norm: Standing Notation §weights ; since ( s + t ) 2 ≤ 2 s 2 + 2 t 2 (s+t)^{2}\le2s^{2}+2t^{2} ( s + t ) 2 ≤ 2 s 2 + 2 t 2 ,
∣ Ψ n ( u , w ) ∣ 2 ≤ 2 a ˉ ∥ u ∥ 2 + 2 ∣ w ∣ 2 . |\Psi_{n}(u,w)|^{2}\le2\bar{a}\lVert u\rVert^{2}+2|w|^{2}. ∣ Ψ n ( u , w ) ∣ 2 ≤ 2 a ˉ ∥ u ∥ 2 + 2∣ w ∣ 2 .
By the transfer formula, monotonicity and additivity of the integral of nonnegative functions (Linearity and Monotonicity of the Lebesgue Integral §nonnegative ), and the transfer formula again with the projections of (P4),
M 2 ( E n ( λ ) ) = ∫ ∣ Ψ n ∣ 2 d ( λ ⊗ τ n ) ≤ 2 a ˉ ∫ R n ∥ u ∥ 2 λ ( d u ) + 2 ∫ X ∣ w ∣ 2 τ n ( d w ) , M_{2}(E_{n}(\lambda))=\int|\Psi_{n}|^{2}\,d(\lambda\otimes\tau_{n})\le2\bar{a}\int_{\mathbb{R}^{n}}\lVert u\rVert^{2}\,\lambda(du)+2\int_{X}|w|^{2}\,\tau_{n}(dw), M 2 ( E n ( λ )) = ∫ ∣ Ψ n ∣ 2 d ( λ ⊗ τ n ) ≤ 2 a ˉ ∫ R n ∥ u ∥ 2 λ ( d u ) + 2 ∫ X ∣ w ∣ 2 τ n ( d w ) ,
with M 2 M_{2} M 2 of The Second Moment of a Borel Probability Measure on a Hilbert Space and the Probability Measures with Finite Second Moment §moment ; here x ↦ ∣ x ∣ 2 x\mapsto|x|^{2} x ↦ ∣ x ∣ 2 on X X X is Borel by The Second Moment of a Borel Probability Measure on a Hilbert Space and the Probability Measures with Finite Second Moment , so ∣ Ψ n ∣ 2 |\Psi_{n}|^{2} ∣ Ψ n ∣ 2 and ( u , w ) ↦ ∣ w ∣ 2 (u,w)\mapsto|w|^{2} ( u , w ) ↦ ∣ w ∣ 2 are measurable as composites with the measurable maps Ψ n \Psi_{n} Ψ n and p r X \mathrm{pr}_{X} pr X , and u ↦ ∥ u ∥ 2 u\mapsto\lVert u\rVert^{2} u ↦ ∥ u ∥ 2 on R n \mathbb{R}^{n} R n is Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions , so ( u , w ) ↦ ∥ u ∥ 2 (u,w)\mapsto\lVert u\rVert^{2} ( u , w ) ↦ ∥ u ∥ 2 is measurable as its composite with p r R n \mathrm{pr}_{\mathbb{R}^{n}} pr R n . The first integral is finite as λ ∈ P 2 ( R n ) \lambda\in\mathcal{P}_{2}(\mathbb{R}^{n}) λ ∈ P 2 ( R n ) (The Second Moment of a Probability Measure on Euclidean Space and the Probability Measures with Finite Second Moment §space ); the second equals ∫ ∣ Q n x ∣ 2 γ c ( d x ) ≤ M 2 ( γ c ) < ∞ \int|Q_{n}x|^{2}\,\gamma_{c}(dx)\le M_{2}(\gamma_{c})<\infty ∫ ∣ Q n x ∣ 2 γ c ( d x ) ≤ M 2 ( γ c ) < ∞ , by the transfer formula, ∣ Q n x ∣ ≤ ∣ x ∣ |Q_{n}x|\le|x| ∣ Q n x ∣ ≤ ∣ x ∣ (Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §continuity ) and γ c ∈ P 2 ( X ) \gamma_{c}\in\mathcal{P}_{2}(X) γ c ∈ P 2 ( X ) . So E n ( λ ) ∈ P 2 ( X ) E_{n}(\lambda)\in\mathcal{P}_{2}(X) E n ( λ ) ∈ P 2 ( X ) by The Second Moment of a Borel Probability Measure on a Hilbert Space and the Probability Measures with Finite Second Moment §space .
Step 6 (Claim 6). Let f f f be a density of λ \lambda λ with respect to γ ~ n \tilde{\gamma}_{n} γ ~ n . Then f ∘ r n f\circ r_{n} f ∘ r n is Borel and nonnegative. By (P5), with α = γ ~ n \alpha=\tilde{\gamma}_{n} α = γ ~ n , α ′ = λ \alpha'=\lambda α ′ = λ , β = β ′ = τ n \beta=\beta'=\tau_{n} β = β ′ = τ n and g = 1 g=1 g = 1 , the function ( u , w ) ↦ f ( u ) (u,w)\mapsto f(u) ( u , w ) ↦ f ( u ) is a density of λ ⊗ τ n \lambda\otimes\tau_{n} λ ⊗ τ n with respect to γ ~ n ⊗ τ n \tilde{\gamma}_{n}\otimes\tau_{n} γ ~ n ⊗ τ n . Let B ∈ B ( X ) B\in\mathcal{B}(X) B ∈ B ( X ) . For w ∈ X 0 w\in X_{0} w ∈ X 0 , (P3) gives 1 B ( Ψ n ( u , w ) ) f ( u ) = ( 1 B ⋅ ( f ∘ r n ) ) ( Ψ n ( u , w ) ) \mathbf{1}_{B}(\Psi_{n}(u,w))f(u)=\bigl(\mathbf{1}_{B}\cdot(f\circ r_{n})\bigr)(\Psi_{n}(u,w)) 1 B ( Ψ n ( u , w )) f ( u ) = ( 1 B ⋅ ( f ∘ r n ) ) ( Ψ n ( u , w )) , so the two sides agree ( γ ~ n ⊗ τ n ) (\tilde{\gamma}_{n}\otimes\tau_{n}) ( γ ~ n ⊗ τ n ) -almost everywhere by (P4). Hence, by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §comparison , the transfer formula and Step 2,
E n ( λ ) ( B ) = ∫ 1 B ( Ψ n ( u , w ) ) f ( u ) ( γ ~ n ⊗ τ n ) ( d u d w ) = ∫ ( 1 B ⋅ ( f ∘ r n ) ) ∘ Ψ n d ( γ ~ n ⊗ τ n ) = ∫ X 1 B ( f ∘ r n ) d γ c . E_{n}(\lambda)(B)=\int\mathbf{1}_{B}(\Psi_{n}(u,w))\,f(u)\,(\tilde{\gamma}_{n}\otimes\tau_{n})(du\,dw)=\int\bigl(\mathbf{1}_{B}\cdot(f\circ r_{n})\bigr)\circ\Psi_{n}\,d(\tilde{\gamma}_{n}\otimes\tau_{n})=\int_{X}\mathbf{1}_{B}\,(f\circ r_{n})\,d\gamma_{c}. E n ( λ ) ( B ) = ∫ 1 B ( Ψ n ( u , w )) f ( u ) ( γ ~ n ⊗ τ n ) ( d u d w ) = ∫ ( 1 B ⋅ ( f ∘ r n ) ) ∘ Ψ n d ( γ ~ n ⊗ τ n ) = ∫ X 1 B ( f ∘ r n ) d γ c .
So f ∘ r n f\circ r_{n} f ∘ r n is a density of E n ( λ ) E_{n}(\lambda) E n ( λ ) with respect to γ c \gamma_{c} γ c .
Step 7 (Claim 7). Suppose first that λ \lambda λ has finite relative entropy with respect to γ ~ n \tilde{\gamma}_{n} γ ~ n , with a density f f f such that ϕ ∘ f \phi\circ f ϕ ∘ f is integrable with respect to γ ~ n \tilde{\gamma}_{n} γ ~ n . By Step 6, f ∘ r n f\circ r_{n} f ∘ r n is a density of E n ( λ ) E_{n}(\lambda) E n ( λ ) with respect to γ c \gamma_{c} γ c , and ϕ ∘ ( f ∘ r n ) = ( ϕ ∘ f ) ∘ r n \phi\circ(f\circ r_{n})=(\phi\circ f)\circ r_{n} ϕ ∘ ( f ∘ r n ) = ( ϕ ∘ f ) ∘ r n . Since ( r n ) # γ c = γ ~ n (r_{n})_{\#}\gamma_{c}=\tilde{\gamma}_{n} ( r n ) # γ c = γ ~ n (Step 1), the transfer formula shows that ( ϕ ∘ f ) ∘ r n (\phi\circ f)\circ r_{n} ( ϕ ∘ f ) ∘ r n is integrable with respect to γ c \gamma_{c} γ c with ∫ ( ϕ ∘ f ) ∘ r n d γ c = ∫ ϕ ∘ f d γ ~ n \int(\phi\circ f)\circ r_{n}\,d\gamma_{c}=\int\phi\circ f\,d\tilde{\gamma}_{n} ∫ ( ϕ ∘ f ) ∘ r n d γ c = ∫ ϕ ∘ f d γ ~ n . By Relative Entropy of Probability Measures §relative-entropy , E n ( λ ) E_{n}(\lambda) E n ( λ ) has finite relative entropy with respect to γ c \gamma_{c} γ c and H ( E n ( λ ) ∣ γ c ) = H ( λ ∣ γ ~ n ) H(E_{n}(\lambda)\,|\,\gamma_{c})=H(\lambda\,|\,\tilde{\gamma}_{n}) H ( E n ( λ ) ∣ γ c ) = H ( λ ∣ γ ~ n ) . Conversely, suppose E n ( λ ) E_{n}(\lambda) E n ( λ ) has finite relative entropy with respect to γ c \gamma_{c} γ c . By Relative Entropy on a Measurable Space: the Gibbs Inequality, the Variational Criterion and Formula, the Entropy Inequality, Small Sets and Data Processing §data-processing with the Borel map T = r n T=r_{n} T = r n , the measure ( r n ) # E n ( λ ) (r_{n})_{\#}E_{n}(\lambda) ( r n ) # E n ( λ ) , which is λ \lambda λ by Step 3, has finite relative entropy with respect to ( r n ) # γ c (r_{n})_{\#}\gamma_{c} ( r n ) # γ c , which is γ ~ n \tilde{\gamma}_{n} γ ~ n by Step 1. The equality of the two entropies then follows from the first part.
Step 8 (Claim 8). As recorded in claim 8, γ ~ n ⊗ σ \tilde{\gamma}_{n}\otimes\sigma γ ~ n ⊗ σ is the product measure of Existence and Uniqueness of the Product Measure , a probability measure since ( γ ~ n ⊗ σ ) ( R n × X ) = γ ~ n ( R n ) σ ( X ) = 1 (\tilde{\gamma}_{n}\otimes\sigma)(\mathbb{R}^{n}\times X)=\tilde{\gamma}_{n}(\mathbb{R}^{n})\,\sigma(X)=1 ( γ ~ n ⊗ σ ) ( R n × X ) = γ ~ n ( R n ) σ ( X ) = 1 ; so (P4) and (P5) apply to it. Let f σ f_{\sigma} f σ be a density of σ \sigma σ with respect to τ n \tau_{n} τ n with ϕ ∘ f σ \phi\circ f_{\sigma} ϕ ∘ f σ integrable with respect to τ n \tau_{n} τ n . By (P5), with α = α ′ = γ ~ n \alpha=\alpha'=\tilde{\gamma}_{n} α = α ′ = γ ~ n , f = 1 f=1 f = 1 , β = τ n \beta=\tau_{n} β = τ n , β ′ = σ \beta'=\sigma β ′ = σ and g = f σ g=f_{\sigma} g = f σ , the function F = f σ ∘ p r X F=f_{\sigma}\circ\mathrm{pr}_{X} F = f σ ∘ pr X is a density of γ ~ n ⊗ σ \tilde{\gamma}_{n}\otimes\sigma γ ~ n ⊗ σ with respect to γ ~ n ⊗ τ n \tilde{\gamma}_{n}\otimes\tau_{n} γ ~ n ⊗ τ n . Since ϕ ∘ F = ( ϕ ∘ f σ ) ∘ p r X \phi\circ F=(\phi\circ f_{\sigma})\circ\mathrm{pr}_{X} ϕ ∘ F = ( ϕ ∘ f σ ) ∘ pr X and ( p r X ) # ( γ ~ n ⊗ τ n ) = τ n (\mathrm{pr}_{X})_{\#}(\tilde{\gamma}_{n}\otimes\tau_{n})=\tau_{n} ( pr X ) # ( γ ~ n ⊗ τ n ) = τ n by (P4), the transfer formula shows that ϕ ∘ F \phi\circ F ϕ ∘ F is integrable with respect to γ ~ n ⊗ τ n \tilde{\gamma}_{n}\otimes\tau_{n} γ ~ n ⊗ τ n with integral ∫ ϕ ∘ f σ d τ n \int\phi\circ f_{\sigma}\,d\tau_{n} ∫ ϕ ∘ f σ d τ n . So, on the measurable space ( R n × X , B ( R n ) ⊗ B ( X ) ) (\mathbb{R}^{n}\times X,\mathcal{B}(\mathbb{R}^{n})\otimes\mathcal{B}(X)) ( R n × X , B ( R n ) ⊗ B ( X )) , γ ~ n ⊗ σ \tilde{\gamma}_{n}\otimes\sigma γ ~ n ⊗ σ has finite relative entropy with respect to γ ~ n ⊗ τ n \tilde{\gamma}_{n}\otimes\tau_{n} γ ~ n ⊗ τ n and H ( γ ~ n ⊗ σ ∣ γ ~ n ⊗ τ n ) = H ( σ ∣ τ n ) H(\tilde{\gamma}_{n}\otimes\sigma\,|\,\tilde{\gamma}_{n}\otimes\tau_{n})=H(\sigma\,|\,\tau_{n}) H ( γ ~ n ⊗ σ ∣ γ ~ n ⊗ τ n ) = H ( σ ∣ τ n ) (Relative Entropy of Probability Measures §relative-entropy ). By Relative Entropy on a Measurable Space: the Gibbs Inequality, the Variational Criterion and Formula, the Entropy Inequality, Small Sets and Data Processing §data-processing with the Borel map T = Ψ n T=\Psi_{n} T = Ψ n , and since ( Ψ n ) # ( γ ~ n ⊗ τ n ) = E n ( γ ~ n ) = γ c (\Psi_{n})_{\#}(\tilde{\gamma}_{n}\otimes\tau_{n})=E_{n}(\tilde{\gamma}_{n})=\gamma_{c} ( Ψ n ) # ( γ ~ n ⊗ τ n ) = E n ( γ ~ n ) = γ c by Step 2, the measure ( Ψ n ) # ( γ ~ n ⊗ σ ) (\Psi_{n})_{\#}(\tilde{\gamma}_{n}\otimes\sigma) ( Ψ n ) # ( γ ~ n ⊗ σ ) has finite relative entropy with respect to γ c \gamma_{c} γ c and
H ( ( Ψ n ) # ( γ ~ n ⊗ σ ) ∣ γ c ) ≤ H ( γ ~ n ⊗ σ ∣ γ ~ n ⊗ τ n ) = H ( σ ∣ τ n ) . H\bigl((\Psi_{n})_{\#}(\tilde{\gamma}_{n}\otimes\sigma)\,\big|\,\gamma_{c}\bigr)\le H(\tilde{\gamma}_{n}\otimes\sigma\,|\,\tilde{\gamma}_{n}\otimes\tau_{n})=H(\sigma\,|\,\tau_{n}). H ( ( Ψ n ) # ( γ ~ n ⊗ σ ) γ c ) ≤ H ( γ ~ n ⊗ σ ∣ γ ~ n ⊗ τ n ) = H ( σ ∣ τ n ) .
Step 9 (Claim 9). Let μ ∈ P ( X ) \mu\in\mathcal{P}(X) μ ∈ P ( X ) have finite relative entropy with respect to γ c \gamma_{c} γ c , and let m ∈ N m\in\mathbb{N} m ∈ N . By Relative Entropy of the Finite-Dimensional Projections of Borel Probability Measures on a Hilbert Space: Monotonicity and Approximation §projections , applied with γ = γ c \gamma=\gamma_{c} γ = γ c , for which ( p m ) # γ c = γ c ( m ) (p_{m})_{\#}\gamma_{c}=\gamma_{c^{(m)}} ( p m ) # γ c = γ c ( m ) by Diagonal Gaussian Measures on a Hilbert Space §measure , the measure ( p m ) # μ (p_{m})_{\#}\mu ( p m ) # μ has finite relative entropy with respect to γ c ( m ) \gamma_{c^{(m)}} γ c ( m ) . By Step 1 at level m m m , δ m \delta_{m} δ m and δ m − 1 \delta_{m}^{-1} δ m − 1 are Borel and ( δ m ) # γ c ( m ) = γ ~ m (\delta_{m})_{\#}\gamma_{c^{(m)}}=\tilde{\gamma}_{m} ( δ m ) # γ c ( m ) = γ ~ m , hence also ( δ m − 1 ) # γ ~ m = γ c ( m ) (\delta_{m}^{-1})_{\#}\tilde{\gamma}_{m}=\gamma_{c^{(m)}} ( δ m − 1 ) # γ ~ m = γ c ( m ) ; moreover μ ~ m = ( r m ) # μ = ( δ m ) # ( ( p m ) # μ ) \tilde{\mu}_{m}=(r_{m})_{\#}\mu=(\delta_{m})_{\#}((p_{m})_{\#}\mu) μ ~ m = ( r m ) # μ = ( δ m ) # (( p m ) # μ ) and so ( δ m − 1 ) # μ ~ m = ( p m ) # μ (\delta_{m}^{-1})_{\#}\tilde{\mu}_{m}=(p_{m})_{\#}\mu ( δ m − 1 ) # μ ~ m = ( p m ) # μ . By Relative Entropy on a Measurable Space: the Gibbs Inequality, the Variational Criterion and Formula, the Entropy Inequality, Small Sets and Data Processing §data-processing with T = δ m T=\delta_{m} T = δ m , μ ~ m \tilde{\mu}_{m} μ ~ m has finite relative entropy with respect to γ ~ m \tilde{\gamma}_{m} γ ~ m and H ( μ ~ m ∣ γ ~ m ) ≤ H ( ( p m ) # μ ∣ γ c ( m ) ) H(\tilde{\mu}_{m}\,|\,\tilde{\gamma}_{m})\le H((p_{m})_{\#}\mu\,|\,\gamma_{c^{(m)}}) H ( μ ~ m ∣ γ ~ m ) ≤ H (( p m ) # μ ∣ γ c ( m ) ) ; by the same claim with T = δ m − 1 T=\delta_{m}^{-1} T = δ m − 1 , H ( ( p m ) # μ ∣ γ c ( m ) ) ≤ H ( μ ~ m ∣ γ ~ m ) H((p_{m})_{\#}\mu\,|\,\gamma_{c^{(m)}})\le H(\tilde{\mu}_{m}\,|\,\tilde{\gamma}_{m}) H (( p m ) # μ ∣ γ c ( m ) ) ≤ H ( μ ~ m ∣ γ ~ m ) . Hence H ( μ ~ m ∣ γ ~ m ) = H ( ( p m ) # μ ∣ γ c ( m ) ) H(\tilde{\mu}_{m}\,|\,\tilde{\gamma}_{m})=H((p_{m})_{\#}\mu\,|\,\gamma_{c^{(m)}}) H ( μ ~ m ∣ γ ~ m ) = H (( p m ) # μ ∣ γ c ( m ) ) for every m m m , and by Relative Entropy of the Finite-Dimensional Projections of Borel Probability Measures on a Hilbert Space: Monotonicity and Approximation §limit , again with γ = γ c \gamma=\gamma_{c} γ = γ c , this sequence is nondecreasing and converges to H ( μ ∣ γ c ) H(\mu\,|\,\gamma_{c}) H ( μ ∣ γ c ) .
Step 10 (Claim 10). Let λ \lambda λ and g 1 , … , g n g_{1},\dots,g_{n} g 1 , … , g n be as in claim 10 and put μ = E n ( λ ) \mu=E_{n}(\lambda) μ = E n ( λ ) , which lies in P 2 ( X ) \mathcal{P}_{2}(X) P 2 ( X ) by Step 5. The integrals and inner products of L 2 ( λ ) L^{2}(\lambda) L 2 ( λ ) and L 2 ( μ ) L^{2}(\mu) L 2 ( μ ) are computed on representatives (The Lebesgue Space of Square-Integrable Functions is a Real Hilbert Space §inner-product ).
(a) Square-integrability. For k ∈ [ n ] k\in[n] k ∈ [ n ] , g k ∘ r n g_{k}\circ r_{n} g k ∘ r n is Borel, and by the transfer formula and Step 3, ∫ ( a k − 1 / 2 g k ∘ r n ) 2 d μ = a k − 1 ∫ g k 2 d λ < ∞ \int(a_{k}^{-1/2}g_{k}\circ r_{n})^{2}\,d\mu=a_{k}^{-1}\int g_{k}^{2}\,d\lambda<\infty ∫ ( a k − 1/2 g k ∘ r n ) 2 d μ = a k − 1 ∫ g k 2 d λ < ∞ . If g k ′ g'_{k} g k ′ is another Borel representative, then { g k ∘ r n ≠ g k ′ ∘ r n } = r n − 1 ( { g k ≠ g k ′ } ) \{g_{k}\circ r_{n}\ne g'_{k}\circ r_{n}\}=r_{n}^{-1}(\{g_{k}\ne g'_{k}\}) { g k ∘ r n = g k ′ ∘ r n } = r n − 1 ({ g k = g k ′ }) has μ \mu μ -measure λ ( { g k ≠ g k ′ } ) = 0 \lambda(\{g_{k}\ne g'_{k}\})=0 λ ({ g k = g k ′ }) = 0 by Step 3. So the class ζ k ∈ L 2 ( μ ) \zeta_{k}\in L^{2}(\mu) ζ k ∈ L 2 ( μ ) of a k − 1 / 2 ( g k ∘ r n ) a_{k}^{-1/2}(g_{k}\circ r_{n}) a k − 1/2 ( g k ∘ r n ) is well defined; put ζ k = 0 \zeta_{k}=0 ζ k = 0 for k > n k>n k > n . It remains to show, for every k ∈ N k\in\mathbb{N} k ∈ N and every φ ∈ F C b 1 ( X ) \varphi\in\mathcal{F}C^{1}_{b}(X) φ ∈ F C b 1 ( X ) , that ⟨ ζ k , φ ⟩ L 2 ( μ ) = ∫ h k d μ \langle\zeta_{k},\varphi\rangle_{L^{2}(\mu)}=\int h_{k}\,d\mu ⟨ ζ k , φ ⟩ L 2 ( μ ) = ∫ h k d μ , where h k ( x ) = x k φ ( x ) / c k − ∂ k φ ( x ) h_{k}(x)=x_{k}\varphi(x)/c_{k}-\partial_{k}\varphi(x) h k ( x ) = x k φ ( x ) / c k − ∂ k φ ( x ) ; then μ \mu μ has the relative score ( ζ k ) k ∈ N (\zeta_{k})_{k\in\mathbb{N}} ( ζ k ) k ∈ N by The Relative Score with Respect to a Diagonal Gaussian Measure on a Hilbert Space §score .
(b) Setup. Fix φ ∈ F C b 1 ( X ) \varphi\in\mathcal{F}C^{1}_{b}(X) φ ∈ F C b 1 ( X ) with a representation ( N , ψ ) (N,\psi) ( N , ψ ) , so ψ ∈ C b 1 ( R N ) \psi\in C^{1}_{b}(\mathbb{R}^{N}) ψ ∈ C b 1 ( R N ) and φ = ψ ∘ p N \varphi=\psi\circ p_{N} φ = ψ ∘ p N ; let K ≥ 0 K\ge0 K ≥ 0 bound ∣ φ ∣ |\varphi| ∣ φ ∣ (Bounded C^1 Cylindrical Functions: Linear Structure, the Gradient and the Partial Derivatives, Bounds and Integrability, and Density in the Square-Integrable Functions §bounded-borel ). By Bounded C^1 Cylindrical Functions: Linear Structure, the Gradient and the Partial Derivatives, Bounds and Integrability, and Density in the Square-Integrable Functions §partial , ∂ k φ = ( ∂ k ψ ) ∘ p N \partial_{k}\varphi=(\partial_{k}\psi)\circ p_{N} ∂ k φ = ( ∂ k ψ ) ∘ p N for k ≤ N k\le N k ≤ N and ∂ k φ = 0 \partial_{k}\varphi=0 ∂ k φ = 0 for k > N k>N k > N . The function h k h_{k} h k is Borel (the coordinate x ↦ x k x\mapsto x_{k} x ↦ x k is continuous by Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §continuity , φ \varphi φ and ∂ k φ \partial_{k}\varphi ∂ k φ are Borel by Bounded C^1 Cylindrical Functions: Linear Structure, the Gradient and the Partial Derivatives, Bounds and Integrability, and Density in the Square-Integrable Functions §bounded-borel , and Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions applies) and is integrable with respect to μ \mu μ , as recorded in The Relative Score with Respect to a Diagonal Gaussian Measure on a Hilbert Space ; by the transfer formula h k ∘ Ψ n h_{k}\circ\Psi_{n} h k ∘ Ψ n is integrable with respect to λ ⊗ τ n \lambda\otimes\tau_{n} λ ⊗ τ n and ∫ h k d μ = ∫ h k ∘ Ψ n d ( λ ⊗ τ n ) \int h_{k}\,d\mu=\int h_{k}\circ\Psi_{n}\,d(\lambda\otimes\tau_{n}) ∫ h k d μ = ∫ h k ∘ Ψ n d ( λ ⊗ τ n ) .
(c) The sections φ w \varphi_{w} φ w . For w ∈ X w\in X w ∈ X let A w : R n → R N A_{w}:\mathbb{R}^{n}\to\mathbb{R}^{N} A w : R n → R N have j j j -th component Ψ n ( u , w ) j \Psi_{n}(u,w)_{j} Ψ n ( u , w ) j , that is, u ↦ a j 1 / 2 u j + w j u\mapsto a_{j}^{1/2}u_{j}+w_{j} u ↦ a j 1/2 u j + w j for j ≤ min ( n , N ) j\le\min(n,N) j ≤ min ( n , N ) and the constant w j w_{j} w j for n < j ≤ N n<j\le N n < j ≤ N (P1); thus p N ( Ψ n ( u , w ) ) = A w ( u ) p_{N}(\Psi_{n}(u,w))=A_{w}(u) p N ( Ψ n ( u , w )) = A w ( u ) , and φ w ( u ) : = φ ( Ψ n ( u , w ) ) = ψ ( A w ( u ) ) \varphi_{w}(u):=\varphi(\Psi_{n}(u,w))=\psi(A_{w}(u)) φ w ( u ) := φ ( Ψ n ( u , w )) = ψ ( A w ( u )) . By (P6), with n ′ = n n'=n n ′ = n , s j = w j s_{j}=w_{j} s j = w j and β j = a j 1 / 2 \beta_{j}=a_{j}^{1/2} β j = a j 1/2 , we have φ w = ψ ∘ A w ∈ C b 1 ( R n ) \varphi_{w}=\psi\circ A_{w}\in C^{1}_{b}(\mathbb{R}^{n}) φ w = ψ ∘ A w ∈ C b 1 ( R n ) , with ∂ k φ w ( u ) = a k 1 / 2 ∂ k ψ ( A w ( u ) ) \partial_{k}\varphi_{w}(u)=a_{k}^{1/2}\partial_{k}\psi(A_{w}(u)) ∂ k φ w ( u ) = a k 1/2 ∂ k ψ ( A w ( u )) for k ≤ min ( n , N ) k\le\min(n,N) k ≤ min ( n , N ) and ∂ k φ w ( u ) = 0 \partial_{k}\varphi_{w}(u)=0 ∂ k φ w ( u ) = 0 for N < k ≤ n N<k\le n N < k ≤ n . With the formula for ∂ k φ \partial_{k}\varphi ∂ k φ in (b), in all cases
∂ k φ w ( u ) = a k 1 / 2 ( ∂ k φ ) ( Ψ n ( u , w ) ) ( k ∈ [ n ] , u ∈ R n ) . \partial_{k}\varphi_{w}(u)=a_{k}^{1/2}\,(\partial_{k}\varphi)\bigl(\Psi_{n}(u,w)\bigr)\qquad(k\in[n],\ u\in\mathbb{R}^{n}). ∂ k φ w ( u ) = a k 1/2 ( ∂ k φ ) ( Ψ n ( u , w ) ) ( k ∈ [ n ] , u ∈ R n ) .
(d) The case k ∈ [ n ] k\in[n] k ∈ [ n ] . For w ∈ X 0 w\in X_{0} w ∈ X 0 we have w k = 0 w_{k}=0 w k = 0 (P2), so Ψ n ( u , w ) k = a k 1 / 2 u k \Psi_{n}(u,w)_{k}=a_{k}^{1/2}u_{k} Ψ n ( u , w ) k = a k 1/2 u k , and with (c) and a k 1 / 2 / c k = a k − 1 / 2 / c ~ k a_{k}^{1/2}/c_{k}=a_{k}^{-1/2}/\tilde{c}_{k} a k 1/2 / c k = a k − 1/2 / c ~ k , where c ~ k = c k / a k \tilde{c}_{k}=c_{k}/a_{k} c ~ k = c k / a k is the k k k -th entry of c ~ ( n ) \tilde{c}^{(n)} c ~ ( n ) ,
h k ( Ψ n ( u , w ) ) = a k − 1 / 2 ( u k c ~ k φ w ( u ) − ∂ k φ w ( u ) ) ( u ∈ R n , w ∈ X 0 ) . h_{k}\bigl(\Psi_{n}(u,w)\bigr)=a_{k}^{-1/2}\Bigl(\frac{u_{k}}{\tilde{c}_{k}}\,\varphi_{w}(u)-\partial_{k}\varphi_{w}(u)\Bigr)\qquad(u\in\mathbb{R}^{n},\ w\in X_{0}). h k ( Ψ n ( u , w ) ) = a k − 1/2 ( c ~ k u k φ w ( u ) − ∂ k φ w ( u ) ) ( u ∈ R n , w ∈ X 0 ) .
By Fubini's theorem , integrating first in u u u , there is a τ n \tau_{n} τ n -null set N 1 ∈ B ( X ) N_{1}\in\mathcal{B}(X) N 1 ∈ B ( X ) off which u ↦ h k ( Ψ n ( u , w ) ) u\mapsto h_{k}(\Psi_{n}(u,w)) u ↦ h k ( Ψ n ( u , w )) is λ \lambda λ -integrable, and ∫ h k d μ = ∫ X I d τ n \int h_{k}\,d\mu=\int_{X}I\,d\tau_{n} ∫ h k d μ = ∫ X I d τ n , where I ( w ) = ∫ h k ( Ψ n ( u , w ) ) λ ( d u ) I(w)=\int h_{k}(\Psi_{n}(u,w))\,\lambda(du) I ( w ) = ∫ h k ( Ψ n ( u , w )) λ ( d u ) off N 1 N_{1} N 1 and I = 0 I=0 I = 0 on N 1 N_{1} N 1 . Apply Finite Fisher Information Relative to a Diagonal Gaussian Measure on Euclidean Space as Componentwise Integration by Parts against Bounded C^1 Functions with d = n d=n d = n , the variance vector c ~ ( n ) \tilde{c}^{(n)} c ~ ( n ) and ν = λ \nu=\lambda ν = λ , which lies in P 2 ( R n ) \mathcal{P}_{2}(\mathbb{R}^{n}) P 2 ( R n ) and has finite Fisher information relative to γ ~ n \tilde{\gamma}_{n} γ ~ n : by Finite Fisher Information Relative to a Diagonal Gaussian Measure on Euclidean Space as Componentwise Integration by Parts against Bounded C^1 Functions §agreement the component ( ζ λ c ~ ( n ) ) k (\zeta^{\tilde{c}^{(n)}}_{\lambda})_{k} ( ζ λ c ~ ( n ) ) k , of which g k g_{k} g k is a representative, satisfies the identity of Finite Fisher Information Relative to a Diagonal Gaussian Measure on Euclidean Space as Componentwise Integration by Parts against Bounded C^1 Functions §componentwise for every function in C b 1 ( R n ) C^{1}_{b}(\mathbb{R}^{n}) C b 1 ( R n ) , in particular for φ w \varphi_{w} φ w . Hence for w ∈ X 0 ∖ N 1 w\in X_{0}\setminus N_{1} w ∈ X 0 ∖ N 1
I ( w ) = a k − 1 / 2 ∫ R n ( u k c ~ k φ w ( u ) − ∂ k φ w ( u ) ) λ ( d u ) = a k − 1 / 2 ∫ R n g k ( u ) φ ( Ψ n ( u , w ) ) λ ( d u ) = : J ( w ) . I(w)=a_{k}^{-1/2}\int_{\mathbb{R}^{n}}\Bigl(\frac{u_{k}}{\tilde{c}_{k}}\varphi_{w}(u)-\partial_{k}\varphi_{w}(u)\Bigr)\lambda(du)=a_{k}^{-1/2}\int_{\mathbb{R}^{n}}g_{k}(u)\,\varphi(\Psi_{n}(u,w))\,\lambda(du)=:J(w). I ( w ) = a k − 1/2 ∫ R n ( c ~ k u k φ w ( u ) − ∂ k φ w ( u ) ) λ ( d u ) = a k − 1/2 ∫ R n g k ( u ) φ ( Ψ n ( u , w )) λ ( d u ) =: J ( w ) .
Now let G ( u , w ) = a k − 1 / 2 g k ( u ) φ ( Ψ n ( u , w ) ) G(u,w)=a_{k}^{-1/2}g_{k}(u)\varphi(\Psi_{n}(u,w)) G ( u , w ) = a k − 1/2 g k ( u ) φ ( Ψ n ( u , w )) , measurable on R n × X \mathbb{R}^{n}\times X R n × X . Since ∣ G ( u , w ) ∣ ≤ a k − 1 / 2 K ∣ g k ( u ) ∣ ≤ a k − 1 / 2 K ( 1 + g k ( u ) 2 ) / 2 |G(u,w)|\le a_{k}^{-1/2}K|g_{k}(u)|\le a_{k}^{-1/2}K(1+g_{k}(u)^{2})/2 ∣ G ( u , w ) ∣ ≤ a k − 1/2 K ∣ g k ( u ) ∣ ≤ a k − 1/2 K ( 1 + g k ( u ) 2 ) /2 , monotonicity of the integral and the transfer formula along p r R n \mathrm{pr}_{\mathbb{R}^{n}} pr R n (P4) give ∫ ∣ G ∣ d ( λ ⊗ τ n ) ≤ a k − 1 / 2 K ( 1 + ∫ g k 2 d λ ) / 2 < ∞ \int|G|\,d(\lambda\otimes\tau_{n})\le a_{k}^{-1/2}K(1+\int g_{k}^{2}\,d\lambda)/2<\infty ∫ ∣ G ∣ d ( λ ⊗ τ n ) ≤ a k − 1/2 K ( 1 + ∫ g k 2 d λ ) /2 < ∞ , so G G G is integrable. By Fubini's theorem again, there is a τ n \tau_{n} τ n -null set N 2 N_{2} N 2 off which u ↦ G ( u , w ) u\mapsto G(u,w) u ↦ G ( u , w ) is λ \lambda λ -integrable (so J ( w ) J(w) J ( w ) is defined), and ∫ G d ( λ ⊗ τ n ) = ∫ X J ′ d τ n \int G\,d(\lambda\otimes\tau_{n})=\int_{X}J'\,d\tau_{n} ∫ G d ( λ ⊗ τ n ) = ∫ X J ′ d τ n with J ′ = J J'=J J ′ = J off N 2 N_{2} N 2 and J ′ = 0 J'=0 J ′ = 0 on N 2 N_{2} N 2 . The functions I I I and J ′ J' J ′ agree on X 0 ∖ ( N 1 ∪ N 2 ) X_{0}\setminus(N_{1}\cup N_{2}) X 0 ∖ ( N 1 ∪ N 2 ) , whose complement is τ n \tau_{n} τ n -null by (P2) and The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §null-union ; so ∫ h k d μ = ∫ G d ( λ ⊗ τ n ) \int h_{k}\,d\mu=\int G\,d(\lambda\otimes\tau_{n}) ∫ h k d μ = ∫ G d ( λ ⊗ τ n ) by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §comparison . Finally, by (P3), G = ( a k − 1 / 2 ( g k ∘ r n ) φ ) ∘ Ψ n G=\bigl(a_{k}^{-1/2}(g_{k}\circ r_{n})\,\varphi\bigr)\circ\Psi_{n} G = ( a k − 1/2 ( g k ∘ r n ) φ ) ∘ Ψ n at every ( u , w ) (u,w) ( u , w ) with w ∈ X 0 w\in X_{0} w ∈ X 0 , hence ( λ ⊗ τ n ) (\lambda\otimes\tau_{n}) ( λ ⊗ τ n ) -almost everywhere (P4); by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §comparison and the transfer formula,
∫ X h k d μ = ∫ X a k − 1 / 2 ( g k ∘ r n ) φ d μ = ⟨ ζ k , φ ⟩ L 2 ( μ ) . \int_{X}h_{k}\,d\mu=\int_{X}a_{k}^{-1/2}(g_{k}\circ r_{n})\,\varphi\,d\mu=\langle\zeta_{k},\varphi\rangle_{L^{2}(\mu)}. ∫ X h k d μ = ∫ X a k − 1/2 ( g k ∘ r n ) φ d μ = ⟨ ζ k , φ ⟩ L 2 ( μ ) .
(e) The case k > n k>n k > n . Fix u ∈ R n u\in\mathbb{R}^{n} u ∈ R n and let B u : R N → R N B_{u}:\mathbb{R}^{N}\to\mathbb{R}^{N} B u : R N → R N have j j j -th component the constant a j 1 / 2 u j a_{j}^{1/2}u_{j} a j 1/2 u j for j ≤ min ( n , N ) j\le\min(n,N) j ≤ min ( n , N ) and y ↦ y j y\mapsto y_{j} y ↦ y j for n < j ≤ N n<j\le N n < j ≤ N . For x ∈ X x\in X x ∈ X , Q n x ∈ X 0 Q_{n}x\in X_{0} Q n x ∈ X 0 and ( Q n x ) j = x j − p n ∗ ( p n ( x ) ) j = x j (Q_{n}x)_{j}=x_{j}-p_{n}^{*}(p_{n}(x))_{j}=x_{j} ( Q n x ) j = x j − p n ∗ ( p n ( x ) ) j = x j for j > n j>n j > n (P2), so by (P1) Ψ n ( u , Q n x ) j = a j 1 / 2 u j \Psi_{n}(u,Q_{n}x)_{j}=a_{j}^{1/2}u_{j} Ψ n ( u , Q n x ) j = a j 1/2 u j for j ≤ n j\le n j ≤ n and Ψ n ( u , Q n x ) j = x j \Psi_{n}(u,Q_{n}x)_{j}=x_{j} Ψ n ( u , Q n x ) j = x j for j > n j>n j > n . Hence p N ( Ψ n ( u , Q n x ) ) = B u ( p N ( x ) ) p_{N}(\Psi_{n}(u,Q_{n}x))=B_{u}(p_{N}(x)) p N ( Ψ n ( u , Q n x )) = B u ( p N ( x )) and
φ u ( x ) : = φ ( Ψ n ( u , Q n x ) ) = ( ψ ∘ B u ) ( p N ( x ) ) . \varphi^{u}(x):=\varphi\bigl(\Psi_{n}(u,Q_{n}x)\bigr)=(\psi\circ B_{u})\bigl(p_{N}(x)\bigr). φ u ( x ) := φ ( Ψ n ( u , Q n x ) ) = ( ψ ∘ B u ) ( p N ( x ) ) .
By (P6), with n ′ = N n'=N n ′ = N , s j = a j 1 / 2 u j s_{j}=a_{j}^{1/2}u_{j} s j = a j 1/2 u j and β j = 0 \beta_{j}=0 β j = 0 for j ≤ min ( n , N ) j\le\min(n,N) j ≤ min ( n , N ) , and s j = 0 s_{j}=0 s j = 0 and β j = 1 \beta_{j}=1 β j = 1 for n < j ≤ N n<j\le N n < j ≤ N , we have ψ ∘ B u ∈ C b 1 ( R N ) \psi\circ B_{u}\in C^{1}_{b}(\mathbb{R}^{N}) ψ ∘ B u ∈ C b 1 ( R N ) with ∂ k ( ψ ∘ B u ) = ( ∂ k ψ ) ∘ B u \partial_{k}(\psi\circ B_{u})=(\partial_{k}\psi)\circ B_{u} ∂ k ( ψ ∘ B u ) = ( ∂ k ψ ) ∘ B u for n < k ≤ N n<k\le N n < k ≤ N ; so φ u ∈ F C b 1 ( X ) \varphi^{u}\in\mathcal{F}C^{1}_{b}(X) φ u ∈ F C b 1 ( X ) with representation ( N , ψ ∘ B u ) (N,\psi\circ B_{u}) ( N , ψ ∘ B u ) (Bounded C^1 Cylindrical Functions on a Hilbert Space with an Orthonormal Basis §cylindrical ). For k > n k>n k > n , Bounded C^1 Cylindrical Functions: Linear Structure, the Gradient and the Partial Derivatives, Bounds and Integrability, and Density in the Square-Integrable Functions §partial applied to φ u \varphi^{u} φ u and to φ \varphi φ gives ∂ k φ u ( x ) = ∂ k ψ ( B u ( p N ( x ) ) ) = ( ∂ k φ ) ( Ψ n ( u , Q n x ) ) \partial_{k}\varphi^{u}(x)=\partial_{k}\psi(B_{u}(p_{N}(x)))=(\partial_{k}\varphi)(\Psi_{n}(u,Q_{n}x)) ∂ k φ u ( x ) = ∂ k ψ ( B u ( p N ( x ))) = ( ∂ k φ ) ( Ψ n ( u , Q n x )) if k ≤ N k\le N k ≤ N , and both sides vanish if k > N k>N k > N ; also x k = Ψ n ( u , Q n x ) k x_{k}=\Psi_{n}(u,Q_{n}x)_{k} x k = Ψ n ( u , Q n x ) k . Therefore
h k ( Ψ n ( u , Q n x ) ) = x k c k φ u ( x ) − ∂ k φ u ( x ) ( x ∈ X ) . h_{k}\bigl(\Psi_{n}(u,Q_{n}x)\bigr)=\frac{x_{k}}{c_{k}}\,\varphi^{u}(x)-\partial_{k}\varphi^{u}(x)\qquad(x\in X). h k ( Ψ n ( u , Q n x ) ) = c k x k φ u ( x ) − ∂ k φ u ( x ) ( x ∈ X ) .
By The Relative Score on a Hilbert Space: the Gaussian Measure Has Score Zero, and the Score as a Square-Integrable Field in the Weighted Sequence Space §gaussian , γ c ∈ P 2 ( X ) \gamma_{c}\in\mathcal{P}_{2}(X) γ c ∈ P 2 ( X ) has a relative score with respect to γ c \gamma_{c} γ c whose components are all zero; by The Relative Score with Respect to a Diagonal Gaussian Measure on a Hilbert Space §score applied to γ c \gamma_{c} γ c and φ u \varphi^{u} φ u , the right-hand side is integrable with respect to γ c \gamma_{c} γ c with integral ⟨ 0 , φ u ⟩ L 2 ( γ c ) = 0 \langle0,\varphi^{u}\rangle_{L^{2}(\gamma_{c})}=0 ⟨ 0 , φ u ⟩ L 2 ( γ c ) = 0 . The function w ↦ h k ( Ψ n ( u , w ) ) w\mapsto h_{k}(\Psi_{n}(u,w)) w ↦ h k ( Ψ n ( u , w )) is Borel (claim 3 of Sections of Product-Measurable Sets and Maps Are Measurable, and Insertion Maps into Products Are Measurable ), so by the transfer formula with τ n = ( Q n ) # γ c \tau_{n}=(Q_{n})_{\#}\gamma_{c} τ n = ( Q n ) # γ c it is τ n \tau_{n} τ n -integrable with ∫ X h k ( Ψ n ( u , w ) ) τ n ( d w ) = 0 \int_{X}h_{k}(\Psi_{n}(u,w))\,\tau_{n}(dw)=0 ∫ X h k ( Ψ n ( u , w )) τ n ( d w ) = 0 , for every u ∈ R n u\in\mathbb{R}^{n} u ∈ R n . By Fubini's theorem, integrating first in w w w , ∫ h k d μ = ∫ h k ∘ Ψ n d ( λ ⊗ τ n ) = 0 = ⟨ ζ k , φ ⟩ L 2 ( μ ) \int h_{k}\,d\mu=\int h_{k}\circ\Psi_{n}\,d(\lambda\otimes\tau_{n})=0=\langle\zeta_{k},\varphi\rangle_{L^{2}(\mu)} ∫ h k d μ = ∫ h k ∘ Ψ n d ( λ ⊗ τ n ) = 0 = ⟨ ζ k , φ ⟩ L 2 ( μ ) .
By (d) and (e), μ = E n ( λ ) \mu=E_{n}(\lambda) μ = E n ( λ ) has the relative score ( ζ k ) k ∈ N (\zeta_{k})_{k\in\mathbb{N}} ( ζ k ) k ∈ N with respect to γ c \gamma_{c} γ c described in claim 10.
Step 11 (Claim 11). In the situation of Step 10, for k ∈ [ n ] k\in[n] k ∈ [ n ] the transfer formula and Step 3 give
a k ∥ ζ k ∥ L 2 ( μ ) 2 = a k a k − 1 ∫ X ( g k ∘ r n ) 2 d μ = ∫ R n g k 2 d λ = ∥ ( ζ λ c ~ ( n ) ) k ∥ L 2 ( λ ) 2 , a_{k}\lVert\zeta_{k}\rVert_{L^{2}(\mu)}^{2}=a_{k}\,a_{k}^{-1}\int_{X}(g_{k}\circ r_{n})^{2}\,d\mu=\int_{\mathbb{R}^{n}}g_{k}^{2}\,d\lambda=\bigl\lVert(\zeta^{\tilde{c}^{(n)}}_{\lambda})_{k}\bigr\rVert_{L^{2}(\lambda)}^{2}, a k ∥ ζ k ∥ L 2 ( μ ) 2 = a k a k − 1 ∫ X ( g k ∘ r n ) 2 d μ = ∫ R n g k 2 d λ = ( ζ λ c ~ ( n ) ) k L 2 ( λ ) 2 ,
and a k ∥ ζ k ∥ L 2 ( μ ) 2 = 0 a_{k}\lVert\zeta_{k}\rVert_{L^{2}(\mu)}^{2}=0 a k ∥ ζ k ∥ L 2 ( μ ) 2 = 0 for k > n k>n k > n . So the partial sums of the series ∑ k = 1 ∞ a k ∥ ζ k ∥ L 2 ( μ ) 2 \sum_{k=1}^{\infty}a_{k}\lVert\zeta_{k}\rVert_{L^{2}(\mu)}^{2} ∑ k = 1 ∞ a k ∥ ζ k ∥ L 2 ( μ ) 2 are constant from the index n n n on; the series converges, μ \mu μ has finite Fisher information relative to γ c \gamma_{c} γ c with weights a a a (Weight Sequences and the Weighted Fisher Information Relative to a Diagonal Gaussian Measure on a Hilbert Space §information ), and by Finite Fisher Information Relative to a Diagonal Gaussian Measure on Euclidean Space as Componentwise Integration by Parts against Bounded C^1 Functions §agreement
I a ( E n ( λ ) ∣ γ c ) = ∑ k = 1 n ∥ ( ζ λ c ~ ( n ) ) k ∥ L 2 ( λ ) 2 = I ( λ ∣ γ ~ n ) . \mathcal{I}_{a}(E_{n}(\lambda)\,|\,\gamma_{c})=\sum_{k=1}^{n}\bigl\lVert(\zeta^{\tilde{c}^{(n)}}_{\lambda})_{k}\bigr\rVert_{L^{2}(\lambda)}^{2}=\mathcal{I}(\lambda\,|\,\tilde{\gamma}_{n}). I a ( E n ( λ ) ∣ γ c ) = k = 1 ∑ n ( ζ λ c ~ ( n ) ) k L 2 ( λ ) 2 = I ( λ ∣ γ ~ n ) .
Step 12 (Claim 12). (a) A coupling. Let λ , λ ′ ∈ P 2 ( R n ) \lambda,\lambda'\in\mathcal{P}_{2}(\mathbb{R}^{n}) λ , λ ′ ∈ P 2 ( R n ) . By Existence of an Optimal Coupling of Two Probability Measures with Finite Second Moment (with m = n m=n m = n ) there is a coupling π ∈ P ( R n + n ) \pi\in\mathcal{P}(\mathbb{R}^{n+n}) π ∈ P ( R n + n ) of λ \lambda λ and λ ′ \lambda' λ ′ with quadratic cost I ( π ) = W 2 ( λ , λ ′ ) 2 I(\pi)=W_{2}(\lambda,\lambda')^{2} I ( π ) = W 2 ( λ , λ ′ ) 2 . Let p r 1 , p r 2 : R n + n → R n \mathrm{pr}_{1},\mathrm{pr}_{2}:\mathbb{R}^{n+n}\to\mathbb{R}^{n} pr 1 , pr 2 : R n + n → R n be the Borel projections of Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections , and equip R n + n × X \mathbb{R}^{n+n}\times X R n + n × X with B ( R n + n ) ⊗ B ( X ) \mathcal{B}(\mathbb{R}^{n+n})\otimes\mathcal{B}(X) B ( R n + n ) ⊗ B ( X ) and the product measure π ⊗ τ n \pi\otimes\tau_{n} π ⊗ τ n (Existence and Uniqueness of the Product Measure ). For i = 1 , 2 i=1,2 i = 1 , 2 let L i ( z , w ) = ( p r i ( z ) , w ) L_{i}(z,w)=(\mathrm{pr}_{i}(z),w) L i ( z , w ) = ( pr i ( z ) , w ) , a measurable map into R n × X \mathbb{R}^{n}\times X R n × X by claim 1 of Pairings into a Product, the Graph of a Measurable Map, and Couplings Concentrated on a Graph and claim 5 of Sections of Product-Measurable Sets and Maps Are Measurable, and Insertion Maps into Products Are Measurable . Since L 1 − 1 ( A × B ) = p r 1 − 1 ( A ) × B L_{1}^{-1}(A\times B)=\mathrm{pr}_{1}^{-1}(A)\times B L 1 − 1 ( A × B ) = pr 1 − 1 ( A ) × B has ( π ⊗ τ n ) (\pi\otimes\tau_{n}) ( π ⊗ τ n ) -measure π ( p r 1 − 1 ( A ) ) τ n ( B ) = λ ( A ) τ n ( B ) \pi(\mathrm{pr}_{1}^{-1}(A))\tau_{n}(B)=\lambda(A)\tau_{n}(B) π ( pr 1 − 1 ( A )) τ n ( B ) = λ ( A ) τ n ( B ) , the rectangle principle (P4) gives ( L 1 ) # ( π ⊗ τ n ) = λ ⊗ τ n (L_{1})_{\#}(\pi\otimes\tau_{n})=\lambda\otimes\tau_{n} ( L 1 ) # ( π ⊗ τ n ) = λ ⊗ τ n , and likewise ( L 2 ) # ( π ⊗ τ n ) = λ ′ ⊗ τ n (L_{2})_{\#}(\pi\otimes\tau_{n})=\lambda'\otimes\tau_{n} ( L 2 ) # ( π ⊗ τ n ) = λ ′ ⊗ τ n . Let S = Ψ n ∘ L 1 S=\Psi_{n}\circ L_{1} S = Ψ n ∘ L 1 and T = Ψ n ∘ L 2 T=\Psi_{n}\circ L_{2} T = Ψ n ∘ L 2 , measurable maps into X X X ; the map ( S , T ) (S,T) ( S , T ) into X × X X\times X X × X is measurable by Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §pairing , and Π = ( S , T ) # ( π ⊗ τ n ) ∈ P ( X × X ) \Pi=(S,T)_{\#}(\pi\otimes\tau_{n})\in\mathcal{P}(X\times X) Π = ( S , T ) # ( π ⊗ τ n ) ∈ P ( X × X ) . With the coordinate maps π 1 , π 2 \pi_{1},\pi_{2} π 1 , π 2 of Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §pairs ,
( π 1 ) # Π = S # ( π ⊗ τ n ) = ( Ψ n ) # ( λ ⊗ τ n ) = E n ( λ ) , ( π 2 ) # Π = E n ( λ ′ ) , (\pi_{1})_{\#}\Pi=S_{\#}(\pi\otimes\tau_{n})=(\Psi_{n})_{\#}(\lambda\otimes\tau_{n})=E_{n}(\lambda),\qquad(\pi_{2})_{\#}\Pi=E_{n}(\lambda'), ( π 1 ) # Π = S # ( π ⊗ τ n ) = ( Ψ n ) # ( λ ⊗ τ n ) = E n ( λ ) , ( π 2 ) # Π = E n ( λ ′ ) ,
so Π \Pi Π is a coupling of E n ( λ ) E_{n}(\lambda) E n ( λ ) and E n ( λ ′ ) E_{n}(\lambda') E n ( λ ′ ) .
(b) Its noise cost. For ( z , w ) ∈ R n + n × X (z,w)\in\mathbb{R}^{n+n}\times X ( z , w ) ∈ R n + n × X put t = p r 2 ( z ) − p r 1 ( z ) t=\mathrm{pr}_{2}(z)-\mathrm{pr}_{1}(z) t = pr 2 ( z ) − pr 1 ( z ) . By (P1), T ( z , w ) − S ( z , w ) = h : = ∑ k = 1 n a k 1 / 2 t k e k T(z,w)-S(z,w)=h:=\sum_{k=1}^{n}a_{k}^{1/2}t_{k}e_{k} T ( z , w ) − S ( z , w ) = h := ∑ k = 1 n a k 1/2 t k e k , whose coordinates are h k = a k 1 / 2 t k h_{k}=a_{k}^{1/2}t_{k} h k = a k 1/2 t k for k ∈ [ n ] k\in[n] k ∈ [ n ] and h k = 0 h_{k}=0 h k = 0 for k > n k>n k > n . So the terms a k − 1 h k 2 a_{k}^{-1}h_{k}^{2} a k − 1 h k 2 equal t k 2 t_{k}^{2} t k 2 for k ∈ [ n ] k\in[n] k ∈ [ n ] and 0 0 0 for k > n k>n k > n ; the series ∑ k a k − 1 h k 2 \sum_{k}a_{k}^{-1}h_{k}^{2} ∑ k a k − 1 h k 2 converges with sum ∥ t ∥ 2 \lVert t\rVert^{2} ∥ t ∥ 2 , so h ∈ X a h\in X^{a} h ∈ X a and ∣ h ∣ a 2 = ∥ t ∥ 2 |h|_{a}^{2}=\lVert t\rVert^{2} ∣ h ∣ a 2 = ∥ t ∥ 2 by The Noise Space of a Weight Sequence on a Hilbert Space with an Orthonormal Basis §space and The Noise Space of a Weight Sequence on a Hilbert Space with an Orthonormal Basis §inner-product . Hence ( S , T ) ( z , w ) (S,T)(z,w) ( S , T ) ( z , w ) lies in the set D a D_{a} D a of The Noise Space is a Real Hilbert Space: Orthonormal Basis, Continuous Embedding, Partial Sums, Closed Balls and Borel Measurability §pairs for every ( z , w ) (z,w) ( z , w ) , and c a ( ( S , T ) ( z , w ) ) = ∣ h ∣ a 2 = ∥ p r 1 ( z ) − p r 2 ( z ) ∥ 2 c_{a}((S,T)(z,w))=|h|_{a}^{2}=\lVert\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z)\rVert^{2} c a (( S , T ) ( z , w )) = ∣ h ∣ a 2 = ∥ pr 1 ( z ) − pr 2 ( z ) ∥ 2 by the definition of c a c_{a} c a and n a n_{a} n a there and in The Noise Space is a Real Hilbert Space: Orthonormal Basis, Continuous Embedding, Partial Sums, Closed Balls and Borel Measurability §borel . Therefore Π ( D a ) = ( π ⊗ τ n ) ( R n + n × X ) = 1 \Pi(D_{a})=(\pi\otimes\tau_{n})(\mathbb{R}^{n+n}\times X)=1 Π ( D a ) = ( π ⊗ τ n ) ( R n + n × X ) = 1 , and by the transfer formula (once along ( S , T ) (S,T) ( S , T ) , once along the projection onto R n + n \mathbb{R}^{n+n} R n + n , whose image of π ⊗ τ n \pi\otimes\tau_{n} π ⊗ τ n is π \pi π by (P4) with Y = R n + n Y=\mathbb{R}^{n+n} Y = R n + n ; the function z ↦ ∥ p r 1 ( z ) − p r 2 ( z ) ∥ 2 z\mapsto\lVert\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z)\rVert^{2} z ↦ ∥ pr 1 ( z ) − pr 2 ( z ) ∥ 2 on R n + n \mathbb{R}^{n+n} R n + n is Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions ) and Couplings of Two Probability Measures on Euclidean Space and Their Quadratic Cost §cost ,
∫ X × X c a d Π = ∫ ∥ p r 1 ( z ) − p r 2 ( z ) ∥ 2 ( π ⊗ τ n ) ( d z d w ) = ∫ R n + n ∥ p r 1 ( z ) − p r 2 ( z ) ∥ 2 π ( d z ) = I ( π ) = W 2 ( λ , λ ′ ) 2 < ∞ . \int_{X\times X}c_{a}\,d\Pi=\int\lVert\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z)\rVert^{2}\,(\pi\otimes\tau_{n})(dz\,dw)=\int_{\mathbb{R}^{n+n}}\lVert\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z)\rVert^{2}\,\pi(dz)=I(\pi)=W_{2}(\lambda,\lambda')^{2}<\infty . ∫ X × X c a d Π = ∫ ∥ pr 1 ( z ) − pr 2 ( z ) ∥ 2 ( π ⊗ τ n ) ( d z d w ) = ∫ R n + n ∥ pr 1 ( z ) − pr 2 ( z ) ∥ 2 π ( d z ) = I ( π ) = W 2 ( λ , λ ′ ) 2 < ∞.
So Π ∈ Π a ( E n ( λ ) , E n ( λ ′ ) ) \Pi\in\Pi^{a}(E_{n}(\lambda),E_{n}(\lambda')) Π ∈ Π a ( E n ( λ ) , E n ( λ ′ )) with noise cost I a ( Π ) = W 2 ( λ , λ ′ ) 2 I^{a}(\Pi)=W_{2}(\lambda,\lambda')^{2} I a ( Π ) = W 2 ( λ , λ ′ ) 2 (Couplings of Finite Noise Cost and Their Noise Cost §finite , Couplings of Finite Noise Cost and Their Noise Cost §cost , Couplings of Finite Noise Cost and Their Noise Cost §couplings ), and the ordered pair ( E n ( λ ) , E n ( λ ′ ) ) (E_{n}(\lambda),E_{n}(\lambda')) ( E n ( λ ) , E n ( λ ′ )) is noise-connected (Couplings of Finite Noise Cost and Their Noise Cost §connected ).
(c) Conclusion. Applying (a) and (b) with λ ′ \lambda' λ ′ replaced by γ ~ n \tilde{\gamma}_{n} γ ~ n , which lies in P 2 ( R n ) \mathcal{P}_{2}(\mathbb{R}^{n}) P 2 ( R n ) by Step 1 and satisfies E n ( γ ~ n ) = γ c = ρ E_{n}(\tilde{\gamma}_{n})=\gamma_{c}=\rho E n ( γ ~ n ) = γ c = ρ by Step 2 and A Diagonal Gaussian Reference Measure on the Noise Wasserstein Space, Rescaled Heads and Gaussian Tails: Standing Notation §gaussian , the pair ( E n ( λ ) , ρ ) (E_{n}(\lambda),\rho) ( E n ( λ ) , ρ ) is noise-connected, so E n ( λ ) ∈ P ρ a E_{n}(\lambda)\in\mathcal{P}^{a}_{\rho} E n ( λ ) ∈ P ρ a by The Measures Noise-Connected to the Reference Measure §space ; likewise E n ( λ ′ ) ∈ P ρ a E_{n}(\lambda')\in\mathcal{P}^{a}_{\rho} E n ( λ ′ ) ∈ P ρ a . Since ( E n ( λ ) , E n ( λ ′ ) ) (E_{n}(\lambda),E_{n}(\lambda')) ( E n ( λ ) , E n ( λ ′ )) is noise-connected, The Noise Wasserstein Distance §distance gives W a ( E n ( λ ) , E n ( λ ′ ) ) 2 ≤ I a ( Π ) = W 2 ( λ , λ ′ ) 2 W_{a}(E_{n}(\lambda),E_{n}(\lambda'))^{2}\le I^{a}(\Pi)=W_{2}(\lambda,\lambda')^{2} W a ( E n ( λ ) , E n ( λ ′ ) ) 2 ≤ I a ( Π ) = W 2 ( λ , λ ′ ) 2 , and as both distances are nonnegative,
W a ( E n ( λ ) , E n ( λ ′ ) ) ≤ W 2 ( λ , λ ′ ) . W_{a}\bigl(E_{n}(\lambda),E_{n}(\lambda')\bigr)\le W_{2}(\lambda,\lambda'). W a ( E n ( λ ) , E n ( λ ′ ) ) ≤ W 2 ( λ , λ ′ ) .