TheoremBase

Claim 1: optimality of the head map, and the image of the noise-optimal graph coupling under the rescaled heads is a coupling of the heads of cost at most Wa2W_a^2. Claim 2: isometry of the lift. Claim 3: the graph couplings of the lifts have second marginals whose heads agree with those of nu and whose tails vanish, so they converge to nu in W2W_2; tightness, Prokhorov, lower semicontinuity of the noise cost and uniqueness identify every subsequential weak limit as the graph coupling of T, and the graph-coupling strong-convergence lemma, applied along subsequences, gives strong convergence.

Proof

Each result cited is universally quantified over the data in its own statement.

Throughout, μ\mu, ν\nu, TT and the maps SnS_{n} are those of the statement. By The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §moments, μ,ν∈P2(X)\mu,\nu\in\mathcal{P}_{2}(X), and by The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §connected the ordered pair (μ,ν)(\mu,\nu) is noise-connected, so W:=Wa(μ,ν)W:=W_{a}(\mu,\nu) is a nonnegative real number. Since TT is a noise-optimal map from μ\mu to ν\nu, the transport-cost identity gives ∥T−id∥μ2=W2\lVert T-\mathrm{id}\rVert_{\mu}^{2}=W^{2}, hence ∥T−id∥μ=W\lVert T-\mathrm{id}\rVert_{\mu}=W, both numbers being nonnegative; and since T(x)−x∈XaT(x)-x\in X^{a} and ∣T(x)−x∣a2=na(T(x)−x)|T(x)-x|_{a}^{2}=n_{a}(T(x)-x) for every x∈Xx\in X by Noise-Optimal Maps and Uniquely Noise-Mapped Pairs §map, the norm formula of The Space of Square-Integrable Maps from a Measure Space into a Hilbert Space with an Orthonormal Basis §operations gives

∫Xna(T(x)−x) μ(dx)=W2.(0)\int_{X}n_{a}\bigl(T(x)-x\bigr)\,\mu(dx)=W^{2}. \qquad (0)

For Borel maps gg and hh that can be composed and a Borel measure λ\lambda on the domain of gg one has (h∘g)#λ=h#(g#λ)(h\circ g)_{\#}\lambda=h_{\#}(g_{\#}\lambda), directly from the formula B↦λ(g−1(h−1(B)))B\mapsto\lambda(g^{-1}(h^{-1}(B))) for push-forwards in Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §pushforward and Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pushforward; this is used without further comment. Composites of Borel maps are Borel by claim 4 of Borel Measurability and Bounded Integration on a Metric Space.

Step 1 (the equality in claim 1). Fix n∈Nn\in\mathbb{N} and let ηn:Rn→Rn\eta_{n}:\mathbb{R}^{n}\to\mathbb{R}^{n} be the map ηn(u)=Sn(u)−u\eta_{n}(u)=S_{n}(u)-u. Its kk-th component u↦(Sn(u))k−uku\mapsto(S_{n}(u))_{k}-u_{k} is the difference of two Borel real functions, hence Borel, so ηn\eta_{n} is Borel by the componentwise criterion. The pairing (id,Sn):Rn→Rn+n(\mathrm{id},S_{n}):\mathbb{R}^{n}\to\mathbb{R}^{n+n} is Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §pairing, and pr1∘(id,Sn)=id\mathrm{pr}_{1}\circ(\mathrm{id},S_{n})=\mathrm{id}, pr2∘(id,Sn)=Sn\mathrm{pr}_{2}\circ(\mathrm{id},S_{n})=S_{n}. Put γn=(id,Sn)#μ~n\gamma_{n}=(\mathrm{id},S_{n})_{\#}\tilde{\mu}_{n}, which by hypothesis is an optimal coupling of μ~n\tilde{\mu}_{n} and ν~n\tilde{\nu}_{n}. In particular

(Sn)#μ~n=(pr2)#γn=ν~n(1)(S_{n})_{\#}\tilde{\mu}_{n}=(\mathrm{pr}_{2})_{\#}\gamma_{n}=\tilde{\nu}_{n} \qquad (1)

by Couplings of Two Probability Measures on Euclidean Space and Their Quadratic Cost §coupling. By Couplings of Two Probability Measures on Euclidean Space and Their Quadratic Cost §cost and the change-of-variables formula of Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pushforward,

I(γn)=∫Rn∥u−Sn(u)∥2 μ~n(du)=∫Rn∥Sn(u)−u∥2 μ~n(du),I(\gamma_{n})=\int_{\mathbb{R}^{n}}\lVert u-S_{n}(u)\rVert^{2}\,\tilde{\mu}_{n}(du)=\int_{\mathbb{R}^{n}}\lVert S_{n}(u)-u\rVert^{2}\,\tilde{\mu}_{n}(du),

and I(γn)=W2(μ~n,ν~n)2I(\gamma_{n})=W_{2}(\tilde{\mu}_{n},\tilde{\nu}_{n})^{2} by Optimal Coupling of Two Probability Measures with Finite Second Moment §optimal. This is the equality in claim 1; in particular ∫Rn∥ηn∥2 dμ~n=W2(μ~n,ν~n)2<∞\int_{\mathbb{R}^{n}}\lVert\eta_{n}\rVert^{2}\,d\tilde{\mu}_{n}=W_{2}(\tilde{\mu}_{n},\tilde{\nu}_{n})^{2}<\infty.

Step 2 (the inequality in claim 1). Fix n∈Nn\in\mathbb{N}. The map rnr_{n} is Borel by Lifting Euclidean Tangent Fields of the Rescaled Head to Noise Tangent Fields on a Hilbert Space §head and TT is Borel, so rn∘Tr_{n}\circ T is Borel, and the pairing Gn=(rn,rn∘T):X→Rn+nG_{n}=(r_{n},r_{n}\circ T):X\to\mathbb{R}^{n+n} is measurable with respect to B(X)\mathcal{B}(X) by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §pairing. Let γn′=(Gn)#μ\gamma'_{n}=(G_{n})_{\#}\mu. Then (pr1)#γn′=(rn)#μ=μ~n(\mathrm{pr}_{1})_{\#}\gamma'_{n}=(r_{n})_{\#}\mu=\tilde{\mu}_{n} and (pr2)#γn′=(rn)#(T#μ)=(rn)#ν=ν~n(\mathrm{pr}_{2})_{\#}\gamma'_{n}=(r_{n})_{\#}(T_{\#}\mu)=(r_{n})_{\#}\nu=\tilde{\nu}_{n}, because T#μ=νT_{\#}\mu=\nu by Noise-Optimal Maps and Uniquely Noise-Mapped Pairs §map; so γn′\gamma'_{n} is a coupling of μ~n\tilde{\mu}_{n} and ν~n\tilde{\nu}_{n}. By change of variables,

I(γn′)=∫X∥rn(x)−rn(T(x))∥2 μ(dx).I(\gamma'_{n})=\int_{X}\bigl\lVert r_{n}(x)-r_{n}(T(x))\bigr\rVert^{2}\,\mu(dx).

Fix x∈Xx\in X and put w=T(x)−xw=T(x)-x, which lies in XaX^{a} by Noise-Optimal Maps and Uniquely Noise-Mapped Pairs §map. Coordinates being linear, the kk-th component of rn(T(x))−rn(x)r_{n}(T(x))-r_{n}(x) is ak−1/2wka_{k}^{-1/2}w_{k}, so

∥rn(x)−rn(T(x))∥2=∑k=1nak−1wk2≤∣w∣a2=na(T(x)−x),\bigl\lVert r_{n}(x)-r_{n}(T(x))\bigr\rVert^{2}=\sum_{k=1}^{n}a_{k}^{-1}w_{k}^{2}\le|w|_{a}^{2}=n_{a}\bigl(T(x)-x\bigr),

the inequality because ∣w∣a2|w|_{a}^{2} is the supremum of these partial sums by The Noise Space is a Real Hilbert Space: Orthonormal Basis, Continuous Embedding, Partial Sums, Closed Balls and Borel Measurability §partial-sums, and the last equality by the definition of nan_{a} in The Noise Space is a Real Hilbert Space: Orthonormal Basis, Continuous Embedding, Partial Sums, Closed Balls and Borel Measurability §borel. Integrating this inequality between nonnegative Borel functions and using (0), I(γn′)≤W2I(\gamma'_{n})\le W^{2}. By The Quadratic Wasserstein Distance on Euclidean Space §distance, W2(μ~n,ν~n)2≤I(γn′)≤W2W_{2}(\tilde{\mu}_{n},\tilde{\nu}_{n})^{2}\le I(\gamma'_{n})\le W^{2}. Together with Step 1 this proves claim 1.

Step 3 (claim 2). Fix n∈Nn\in\mathbb{N}. By Step 1, ηn=Sn−id\eta_{n}=S_{n}-\mathrm{id} is Borel with ∫Rn∥ηn∥2 dμ~n<∞\int_{\mathbb{R}^{n}}\lVert\eta_{n}\rVert^{2}\,d\tilde{\mu}_{n}<\infty, and μ∈P2(X)\mu\in\mathcal{P}_{2}(X), so Lifting Euclidean Tangent Fields of the Rescaled Head to Noise Tangent Fields on a Hilbert Space §lift applies: the map Dn:=Λnηn:X→XD_{n}:=\Lambda_{n}\eta_{n}:X\to X is Borel, takes its values in XaX^{a}, is measurable as a map into XaX^{a}, satisfies ∣Dn(x)∣a2=∥ηn(rn(x))∥2|D_{n}(x)|_{a}^{2}=\lVert\eta_{n}(r_{n}(x))\rVert^{2} for every x∈Xx\in X, and is square-integrable with respect to μ\mu; its class is the element Λn(Sn−id)\Lambda_{n}(S_{n}-\mathrm{id}) of L2(μ;Xa)L^{2}(\mu;X^{a}), which is therefore defined. By the norm formula of The Space of Square-Integrable Maps from a Measure Space into a Hilbert Space with an Orthonormal Basis §operations, change of variables along rnr_{n} (Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §pushforward) and claim 1,

∥Λn(Sn−id)∥μ2=∫X∥ηn(rn(x))∥2 μ(dx)=∫Rn∥ηn∥2 dμ~n=W2(μ~n,ν~n)2≤W2.\lVert\Lambda_{n}(S_{n}-\mathrm{id})\rVert_{\mu}^{2}=\int_{X}\lVert\eta_{n}(r_{n}(x))\rVert^{2}\,\mu(dx)=\int_{\mathbb{R}^{n}}\lVert\eta_{n}\rVert^{2}\,d\tilde{\mu}_{n}=W_{2}(\tilde{\mu}_{n},\tilde{\nu}_{n})^{2}\le W^{2}.

Taking nonnegative square roots gives claim 2. Since Dn(x)∈XaD_{n}(x)\in X^{a}, na(Dn(x))=∣Dn(x)∣a2n_{a}(D_{n}(x))=|D_{n}(x)|_{a}^{2} for every xx by The Noise Space is a Real Hilbert Space: Orthonormal Basis, Continuous Embedding, Partial Sums, Closed Balls and Borel Measurability §borel, so

∫Xna(Dn(x)) μ(dx)=∥Dn∥μ2≤W2.(2)\int_{X}n_{a}\bigl(D_{n}(x)\bigr)\,\mu(dx)=\lVert D_{n}\rVert_{\mu}^{2}\le W^{2}. \qquad (2)

Step 4 (graph couplings). By Step 3 and (2), each DnD_{n} is a noise displacement for μ\mu in the sense of the preamble of Strong Convergence of Noise Displacement Fields from Weak Convergence of Their Graph Couplings and an Upper Bound on Their Norms; by that preamble, id+Dn\mathrm{id}+D_{n} and (id,id+Dn)(\mathrm{id},\mathrm{id}+D_{n}) are Borel. Put

λn=(id+Dn)#μ∈P(X),πn=(id,id+Dn)#μ∈P(X×X).\lambda_{n}=(\mathrm{id}+D_{n})_{\#}\mu\in\mathcal{P}(X),\qquad\pi_{n}=(\mathrm{id},\mathrm{id}+D_{n})_{\#}\mu\in\mathcal{P}(X\times X).

The Borel map Σn=id+Dn\Sigma_{n}=\mathrm{id}+D_{n} satisfies Σn(x)−x=Dn(x)∈Xa\Sigma_{n}(x)-x=D_{n}(x)\in X^{a} for every xx and, by (2), ∫Xna(Σn(x)−x) μ(dx)<∞\int_{X}n_{a}(\Sigma_{n}(x)-x)\,\mu(dx)<\infty; so by Couplings of Finite Noise Cost: the Support Bound, Swap, Displacement Couplings and Lower Semicontinuity of the Noise Cost §displacement, πn∈Πa(μ,λn)\pi_{n}\in\Pi^{a}(\mu,\lambda_{n}) and

Ia(πn)=∫Xna(Dn(x)) μ(dx)=∥Dn∥μ2≤W2.(3)I^{a}(\pi_{n})=\int_{X}n_{a}\bigl(D_{n}(x)\bigr)\,\mu(dx)=\lVert D_{n}\rVert_{\mu}^{2}\le W^{2}. \qquad (3)

Since μ∈P2(X)\mu\in\mathcal{P}_{2}(X) and Πa(μ,λn)\Pi^{a}(\mu,\lambda_{n}) is nonempty, λn∈P2(X)\lambda_{n}\in\mathcal{P}_{2}(X) by Couplings of Finite Noise Cost: the Support Bound, Swap, Displacement Couplings and Lower Semicontinuity of the Noise Cost §support-bound. Likewise the map v:X→Xv:X\to X, v(x)=T(x)−xv(x)=T(x)-x, is Borel (preamble of Noise-Optimal Maps and Uniquely Noise-Mapped Pairs), takes values in XaX^{a} and satisfies ∫Xna(v(x)) μ(dx)<∞\int_{X}n_{a}(v(x))\,\mu(dx)<\infty by Noise-Optimal Maps and Uniquely Noise-Mapped Pairs §map, so it is a noise displacement for μ\mu, whose class is T−idT-\mathrm{id}. Since x+v(x)=T(x)x+v(x)=T(x) for every xx, (id,id+v)#μ=(id,T)#μ=:πT(\mathrm{id},\mathrm{id}+v)_{\#}\mu=(\mathrm{id},T)_{\#}\mu=:\pi_{T}, and πT\pi_{T} is a noise-optimal coupling of μ\mu and ν\nu by Noise-Optimal Maps and Uniquely Noise-Mapped Pairs §map.

Step 5 (λn→ν\lambda_{n}\to\nu). Fix n∈Nn\in\mathbb{N}. By Lifting Euclidean Tangent Fields of the Rescaled Head to Noise Tangent Fields on a Hilbert Space §shift, applied to the Borel map ηn\eta_{n}, for every x∈Xx\in X

rn(x+Dn(x))=rn(x)+ηn(rn(x))=Sn(rn(x)),Qn(x+Dn(x))=Qnx.r_{n}\bigl(x+D_{n}(x)\bigr)=r_{n}(x)+\eta_{n}(r_{n}(x))=S_{n}(r_{n}(x)),\qquad Q_{n}\bigl(x+D_{n}(x)\bigr)=Q_{n}x .

The first identity says rn∘(id+Dn)=Sn∘rnr_{n}\circ(\mathrm{id}+D_{n})=S_{n}\circ r_{n}, so (rn)#λn=(Sn)#μ~n=ν~n(r_{n})_{\#}\lambda_{n}=(S_{n})_{\#}\tilde{\mu}_{n}=\tilde{\nu}_{n} by (1). Let An:Rn→RnA_{n}:\mathbb{R}^{n}\to\mathbb{R}^{n}, An(u)=(a11/2u1,…,an1/2un)A_{n}(u)=(a_{1}^{1/2}u_{1},\dots,a_{n}^{1/2}u_{n}); each component is a real multiple of a coordinate, hence continuous and Borel, so AnA_{n} is Borel by the componentwise criterion. Since ak1/2ak−1/2=1a_{k}^{1/2}a_{k}^{-1/2}=1, the coordinate map pnp_{n} of Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §coordinates equals An∘rnA_{n}\circ r_{n}, so

(pn)#λn=(An)#(rn)#λn=(An)#ν~n=(An)#(rn)#ν=(pn)#ν.(p_{n})_{\#}\lambda_{n}=(A_{n})_{\#}(r_{n})_{\#}\lambda_{n}=(A_{n})_{\#}\tilde{\nu}_{n}=(A_{n})_{\#}(r_{n})_{\#}\nu=(p_{n})_{\#}\nu .

The function x↦∣Qnx∣2x\mapsto|Q_{n}x|^{2} is continuous, Borel and nonnegative, with ∣Qnx∣2≤∣x∣2|Q_{n}x|^{2}\le|x|^{2}, by Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §continuity. By change of variables and the second identity,

∫X∣Qnx∣2 λn(dx)=∫X∣Qn(x+Dn(x))∣2 μ(dx)=∫X∣Qnx∣2 μ(dx).\int_{X}|Q_{n}x|^{2}\,\lambda_{n}(dx)=\int_{X}\bigl|Q_{n}\bigl(x+D_{n}(x)\bigr)\bigr|^{2}\,\mu(dx)=\int_{X}|Q_{n}x|^{2}\,\mu(dx).

By Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §coordinates, (Xn)n∈N(X_{n})_{n\in\mathbb{N}} is an exhausting sequence for XX and PnP_{n} is the orthogonal projection onto XnX_{n}, so by Exhausting Sequences of Finite-Dimensional Subspaces in a Separable Real Hilbert Space, and Their Projections §tail, Qnx→0XQ_{n}x\to0_{X} for every x∈Xx\in X, that is ∣Qnx∣→0|Q_{n}x|\to0 and hence ∣Qnx∣2→0|Q_{n}x|^{2}\to0. The function x↦∣x∣2x\mapsto|x|^{2} is integrable with respect to μ\mu because M2(μ)<∞M_{2}(\mu)<\infty (The Second Moment of a Borel Probability Measure on a Hilbert Space and the Probability Measures with Finite Second Moment §space). The dominated convergence theorem, applied on (X,B(X),μ)(X,\mathcal{B}(X),\mu) to fn(x)=∣Qnx∣2f_{n}(x)=|Q_{n}x|^{2} with limit 00 and dominating function x↦∣x∣2x\mapsto|x|^{2}, gives lim⁡n→∞∫X∣Qnx∣2 μ(dx)=0\lim_{n\to\infty}\int_{X}|Q_{n}x|^{2}\,\mu(dx)=0, hence lim⁡n→∞∫X∣Qnx∣2 λn(dx)=0\lim_{n\to\infty}\int_{X}|Q_{n}x|^{2}\,\lambda_{n}(dx)=0. Since ν∈P2(X)\nu\in\mathcal{P}_{2}(X) and every λn∈P2(X)\lambda_{n}\in\mathcal{P}_{2}(X) (Step 4), Measures Whose Heads Agree with a Fixed Measure and Whose Tails Vanish Converge in the Quadratic Wasserstein Distance §convergence gives lim⁡n→∞W2(λn,ν)=0\lim_{n\to\infty}W_{2}(\lambda_{n},\nu)=0; hence λn⇒ν\lambda_{n}\Rightarrow\nu by Wasserstein Convergence on a Hilbert Space: Weak Convergence, Integrals of Continuous Functions of Quadratic Growth, Convergence from Weak Convergence with Uniformly Integrable Second Moments, and Compactness §weak, and the sequence (λn)n∈N(\lambda_{n})_{n\in\mathbb{N}} is tight in (X,d)(X,d) by Cauchy Sequences in the Quadratic Wasserstein Space and Weakly Convergent Sequences on a Hilbert Space are Tight §weakly-convergent.

Step 6 (subsequential limits of the graph couplings). We show: for every strictly increasing sequence (mi)i∈N(m_{i})_{i\in\mathbb{N}} in N\mathbb{N} there is a strictly increasing sequence (ij)j∈N(i_{j})_{j\in\mathbb{N}} in N\mathbb{N} with πmij⇒πT\pi_{m_{i_{j}}}\Rightarrow\pi_{T} as j→∞j\to\infty. By Tight Family of Borel Measures on a Metric Space §sequence, tightness of the sequence (λn)n∈N(\lambda_{n})_{n\in\mathbb{N}} (Step 5) means that the set {λn:n∈N}\{\lambda_{n}:n\in\mathbb{N}\} of its terms is tight in (X,d)(X,d). The set {λmi:i∈N}\{\lambda_{m_{i}}:i\in\mathbb{N}\} is contained in {λn:n∈N}\{\lambda_{n}:n\in\mathbb{N}\}, so it is tight in (X,d)(X,d), every compact set witnessing tightness of the larger set witnessing it for the smaller; and {μ}\{\mu\} is tight by Ulam's Theorem: a Finite Borel Measure on a Complete Separable Metric Space is Tight §tight. By Couplings on a Hilbert Space: Tightness, Closedness under Weak Convergence, and Lower Semicontinuity of the Quadratic Cost §tight, the set of all couplings in Π(μ,λmi)\Pi(\mu,\lambda_{m_{i}}), i∈Ni\in\mathbb{N}, is tight in X×XX\times X; since πmi∈Πa(μ,λmi)⊆Π(μ,λmi)\pi_{m_{i}}\in\Pi^{a}(\mu,\lambda_{m_{i}})\subseteq\Pi(\mu,\lambda_{m_{i}}) by Couplings of Finite Noise Cost and Their Noise Cost §couplings, the set of terms of the sequence (πmi)i∈N(\pi_{m_{i}})_{i\in\mathbb{N}} is tight, that is, this sequence is tight in X×XX\times X (Tight Family of Borel Measures on a Metric Space §sequence), a metric space with the distance of Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §pairs. By Prokhorov's Theorem on a Metric Space: a Tight Sequence of Borel Probability Measures Has a Weakly Convergent Subsequence §subsequence there are a strictly increasing sequence (ij)j∈N(i_{j})_{j\in\mathbb{N}} and π′∈P(X×X)\pi'\in\mathcal{P}(X\times X) with πmij⇒π′\pi_{m_{i_{j}}}\Rightarrow\pi'. Put μj=μ\mu_{j}=\mu and νj=λmij\nu_{j}=\lambda_{m_{i_{j}}}. Then μj⇒μ\mu_{j}\Rightarrow\mu, the integrals in Weak Convergence of Finite Borel Measures on a Metric Space being constant in jj; and νj⇒ν\nu_{j}\Rightarrow\nu, because for every bounded continuous f:X→Rf:X\to\mathbb{R} the sequence (∫Xf dνj)j(\int_{X}f\,d\nu_{j})_{j} is a subsequence, along the strictly increasing indices j↦mijj\mapsto m_{i_{j}}, of the sequence (∫Xf dλn)n(\int_{X}f\,d\lambda_{n})_{n}, which converges to ∫Xf dν\int_{X}f\,d\nu by Step 5, and so converges to the same limit. By (3), 0≤Ia(πmij)≤W20\le I^{a}(\pi_{m_{i_{j}}})\le W^{2} for every jj, so this sequence is bounded and its limit inferior is at most W2W^{2}. By Couplings of Finite Noise Cost: the Support Bound, Swap, Displacement Couplings and Lower Semicontinuity of the Noise Cost §lsc, π′∈Πa(μ,ν)\pi'\in\Pi^{a}(\mu,\nu) and Ia(π′)≤lim inf⁡jIa(πmij)≤W2I^{a}(\pi')\le\liminf_{j}I^{a}(\pi_{m_{i_{j}}})\le W^{2}, while W2≤Ia(π′)W^{2}\le I^{a}(\pi') by The Noise Wasserstein Distance §distance. Hence Ia(π′)=W2I^{a}(\pi')=W^{2}, and π′\pi' is noise-optimal. Since (μ,ν)(\mu,\nu) is uniquely noise-mapped, there is a noise-optimal map T∗T^{*} from μ\mu to ν\nu such that every noise-optimal coupling of μ\mu and ν\nu equals (id,T∗)#μ(\mathrm{id},T^{*})_{\#}\mu. Both π′\pi' and πT\pi_{T} are noise-optimal couplings of μ\mu and ν\nu (Step 4), so π′=(id,T∗)#μ=πT\pi'=(\mathrm{id},T^{*})_{\#}\mu=\pi_{T}.

Step 7 (claim 3). Suppose that the real sequence (∥Λn(Sn−id)−(T−id)∥μ)n∈N(\lVert\Lambda_{n}(S_{n}-\mathrm{id})-(T-\mathrm{id})\rVert_{\mu})_{n\in\mathbb{N}} does not converge to 00. Then there is a real ε>0\varepsilon>0 such that for every N∈NN\in\mathbb{N} some n≥Nn\ge N has ∥Dn−(T−id)∥μ≥ε\lVert D_{n}-(T-\mathrm{id})\rVert_{\mu}\ge\varepsilon; choosing such indices recursively, each larger than the previous one, gives a strictly increasing sequence (mi)i∈N(m_{i})_{i\in\mathbb{N}} with ∥Dmi−(T−id)∥μ≥ε\lVert D_{m_{i}}-(T-\mathrm{id})\rVert_{\mu}\ge\varepsilon for every ii. Let (ij)j∈N(i_{j})_{j\in\mathbb{N}} be given by Step 6, so that πmij⇒πT\pi_{m_{i_{j}}}\Rightarrow\pi_{T}. Apply Strong Convergence of Noise Displacement Fields from Weak Convergence of Their Graph Couplings and an Upper Bound on Their Norms §strong-displacement to the noise displacements vj=Dmijv_{j}=D_{m_{i_{j}}} and v=T−idv=T-\mathrm{id} for μ\mu (Step 4). Its first hypothesis holds because (id,id+vj)#μ=πmij⇒πT=(id,id+v)#μ(\mathrm{id},\mathrm{id}+v_{j})_{\#}\mu=\pi_{m_{i_{j}}}\Rightarrow\pi_{T}=(\mathrm{id},\mathrm{id}+v)_{\#}\mu. Its second hypothesis holds with any JJ: for every jj, ∥vj∥μ≤W=∥v∥μ\lVert v_{j}\rVert_{\mu}\le W=\lVert v\rVert_{\mu} by claim 2 and the transport-cost identity recalled at the start, so ∥vj∥μ≤∥v∥μ+ε′\lVert v_{j}\rVert_{\mu}\le\lVert v\rVert_{\mu}+\varepsilon' for every positive real ε′\varepsilon'. Hence lim⁡j→∞∥Dmij−(T−id)∥μ=0\lim_{j\to\infty}\lVert D_{m_{i_{j}}}-(T-\mathrm{id})\rVert_{\mu}=0, contradicting ∥Dmij−(T−id)∥μ≥ε\lVert D_{m_{i_{j}}}-(T-\mathrm{id})\rVert_{\mu}\ge\varepsilon for every jj. Therefore, the class of DnD_{n} being Λn(Sn−id)\Lambda_{n}(S_{n}-\mathrm{id}) by Step 3, ∥Λn(Sn−id)−(T−id)∥μ→0\lVert\Lambda_{n}(S_{n}-\mathrm{id})-(T-\mathrm{id})\rVert_{\mu}\to0 as n→∞n\to\infty, which is claim 3.

Citations

Loading…

Dependencies

Uses0

Loading…

Comments

Log in to comment.

Loading…