TheoremBase

The head maps of a noise-optimal coupling, chosen for every dimension, determine every coordinate of the target almost surely; synthesizing these coordinates in the basis yields a Borel map whose graph carries the coupling and which is a noise-optimal map. Uniqueness follows by applying this to the midpoint of two noise-optimal couplings, which is again noise-optimal.

Proof

Each result cited is universally quantified over the data in its own statement, and is applied below with the data named at each use.

Conventions. We work in the setting of Probability Measures on a Hilbert Space Transported in the Noise Norm: Standing Notation, with cc, γc\gamma_{c}, μ\mu and ν\nu as in the statement. As in clause 1 of that setting, for z∈X×Xz\in X\times X we write x=π1(z)x=\pi_{1}(z) and y=π2(z)y=\pi_{2}(z); for u∈Xu\in X and k∈Nk\in\mathbb{N}, uk=⟨u,ek⟩u_{k}=\langle u,e_{k}\rangle, and for n∈Nn\in\mathbb{N} and w∈Rnw\in\mathbb{R}^{n} we write wkw_{k} (1≤k≤n1\le k\le n) for the kk-th entry of ww, so that pn(u)=(u1,…,un)p_{n}(u)=(u_{1},\dots,u_{n}) by Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §coordinates. The metric space (X,d)(X,d) is separable by Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §space, and B(X×X)=B(X)⊗B(X)\mathcal{B}(X\times X)=\mathcal{B}(X)\otimes\mathcal{B}(X), with π1,π2\pi_{1},\pi_{2} Borel, by Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §product-sigma. Accordingly Pairings into a Product, the Graph of a Measurable Map, and Couplings Concentrated on a Graph is applied below with (X,B(X))(X,\mathcal{B}(X)) in place of (Y,Y)(Y,\mathcal{Y}) and (X,d)(X,d) in place of (Z,dZ)(Z,d_{Z}); its projections prY,prZ\mathrm{pr}_{Y},\mathrm{pr}_{Z} are then π1,π2\pi_{1},\pi_{2}, for a map S:X→XS:X\to X its graph is ΓS={z∈X×X: y=S(x)}\Gamma_{S}=\{z\in X\times X:\ y=S(x)\}, and its image measures (id,S)#μ(\mathrm{id},S)_{\#}\mu, given by E↦μ((id,S)−1(E))E\mapsto\mu((\mathrm{id},S)^{-1}(E)) on B(X)⊗B(X)=B(X×X)\mathcal{B}(X)\otimes\mathcal{B}(X)=\mathcal{B}(X\times X), are the push-forwards of Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §pushforward. Likewise Measurable Maps into a Hilbert Space with an Orthonormal Basis: Norms, Inner Products and Linear Combinations, Synthesis from Coordinates, Square-Integrability, and Almost-Everywhere Equality is applied with XX in place of EE, the orthonormal basis (ek)k∈N(e_{k})_{k\in\mathbb{N}} in place of (fk)k∈N(f_{k})_{k\in\mathbb{N}} and (X,B(X))(X,\mathcal{B}(X)) in place of (S,S)(S,\mathcal{S}); measurability of maps X→XX\to X there is then Borel measurability, and the coordinate functions of a map v:X→Xv:X\to X are x↦⟨v(x),ek⟩x\mapsto\langle v(x),e_{k}\rangle. Real-valued functions on XX are called Borel when measurable with respect to B(X)\mathcal{B}(X) and B(R)\mathcal{B}(\mathbb{R}), which is also the Borel σ\sigma-algebra of the metric space (R,dR)(\mathbb{R},d_{\mathbb{R}}), dR(s,t)=∣s−t∣d_{\mathbb{R}}(s,t)=|s-t|, by claim 2 of Borel Measurability and Bounded Integration on a Metric Space; and Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions is applied on the measurable space (X,B(X))(X,\mathcal{B}(X)). For N∈NN\in\mathbb{N}, SN:X→RS_{N}:X\to\mathbb{R} is the function SN(u)=∑k=1Nak−1uk2S_{N}(u)=\sum_{k=1}^{N}a_{k}^{-1}u_{k}^{2} of The Noise Space is a Real Hilbert Space: Orthonormal Basis, Continuous Embedding, Partial Sums, Closed Balls and Borel Measurability, applied with the noise weights aa and the bound aˉ\bar{a} of Probability Measures on a Hilbert Space Transported in the Noise Norm: Standing Notation §weights; in particular 0<ak≤aˉ0<a_{k}\le\bar{a} for every k∈Nk\in\mathbb{N}.

Proof of claim 1.

Step 1 (Head maps). Let π\pi be a noise-optimal coupling of μ\mu and ν\nu. By Noise-Optimal Couplings §optimal, π∈Πa(μ,ν)\pi\in\Pi^{a}(\mu,\nu), so by Couplings of Finite Noise Cost and Their Noise Cost §couplings and Couplings of Finite Noise Cost and Their Noise Cost §finite it satisfies π(Da)=1\pi(D_{a})=1 and ∫X×Xca dπ<∞\int_{X\times X}c_{a}\,d\pi<\infty, and by the definition of a coupling π∈P(X×X)\pi\in\mathcal{P}(X\times X) with (π1)#π=μ(\pi_{1})_{\#}\pi=\mu and (π2)#π=ν(\pi_{2})_{\#}\pi=\nu. For every n∈Nn\in\mathbb{N}, claim 1 of lem:noise-head-maps-hilbert, applied with cc, μ\mu, ν\nu, π\pi and nn (its hypotheses are those of the present statement together with the noise-optimality of π\pi), shows that the set AnA_{n} of the Borel maps Tn:X→RnT_{n}:X\to\mathbb{R}^{n} with (π1,pn∘π2)#π=(id,Tn)#μ(\pi_{1},p_{n}\circ\pi_{2})_{\#}\pi=(\mathrm{id},T_{n})_{\#}\mu is nonempty. Regarding each AnA_{n} as a subset of the set of all maps from XX into ⋃m∈NRm\bigcup_{m\in\mathbb{N}}\mathbb{R}^{m}, Axiom of Countable Choice gives a sequence (Tn)n∈N(T_{n})_{n\in\mathbb{N}} with Tn∈AnT_{n}\in A_{n} for every n∈Nn\in\mathbb{N}; we fix it. For every n∈Nn\in\mathbb{N}, claim 2 of the same lemma, applied with the same data and this TnT_{n}, shows that the set

Gn={z∈X×X: pn(y)=Tn(x)}G_{n}=\{z\in X\times X:\ p_{n}(y)=T_{n}(x)\}

belongs to B(X×X)\mathcal{B}(X\times X) and that π(Gn)=1\pi(G_{n})=1.

Step 2 (A set of full measure). The set DaD_{a} belongs to B(X×X)\mathcal{B}(X\times X) by The Noise Space is a Real Hilbert Space: Orthonormal Basis, Continuous Embedding, Partial Sums, Closed Balls and Borel Measurability §pairs. Put

G=Da∩⋂n∈NGn,G=D_{a}\cap\bigcap_{n\in\mathbb{N}}G_{n},

which belongs to B(X×X)\mathcal{B}(X\times X) because a σ\sigma-algebra is closed under countable intersections (Sigma-Algebra and Measurable Space). Let N1=(X×X)∖DaN_{1}=(X\times X)\setminus D_{a} and Nn+1=(X×X)∖GnN_{n+1}=(X\times X)\setminus G_{n} for n∈Nn\in\mathbb{N}. Since π\pi is finite, the rule for differences gives π(N1)=1−π(Da)=0\pi(N_{1})=1-\pi(D_{a})=0 and π(Nn+1)=1−π(Gn)=0\pi(N_{n+1})=1-\pi(G_{n})=0. The complement (X×X)∖G(X\times X)\setminus G is ⋃m∈NNm\bigcup_{m\in\mathbb{N}}N_{m}, so countable subadditivity gives π((X×X)∖G)≤∑m∈Nπ(Nm)=0\pi((X\times X)\setminus G)\le\sum_{m\in\mathbb{N}}\pi(N_{m})=0, every partial sum being 00; hence π((X×X)∖G)=0\pi((X\times X)\setminus G)=0 and, by the rule for differences again, π(G)=1\pi(G)=1.

Step 3 (Coordinates of the head maps). For k∈Nk\in\mathbb{N} define τk:X→R\tau_{k}:X\to\mathbb{R} by τk(x)=⟨pk∗(Tk(x)),ek⟩\tau_{k}(x)=\langle p_{k}^{*}(T_{k}(x)),e_{k}\rangle. By Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §continuity the map pk∗:(Rk,dE)→(X,d)p_{k}^{*}:(\mathbb{R}^{k},d_{E})\to(X,d) and the coordinate function u↦uku\mapsto u_{k} on XX are Borel, and TkT_{k} is Borel; so τk\tau_{k} is Borel by two applications of claim 4 of Borel Measurability and Bounded Integration on a Metric Space. Moreover, by the same claim of Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §continuity, pk(pk∗(w))=wp_{k}(p_{k}^{*}(w))=w for w∈Rkw\in\mathbb{R}^{k}, and the kk-th entry of pk(u)p_{k}(u) is uku_{k}; taking u=pk∗(Tk(x))u=p_{k}^{*}(T_{k}(x)) we get

τk(x)=(pk(pk∗(Tk(x))))k=(Tk(x))k(x∈X).(1)\tau_{k}(x)=\bigl(p_{k}(p_{k}^{*}(T_{k}(x)))\bigr)_{k}=\bigl(T_{k}(x)\bigr)_{k}\qquad(x\in X).\tag{1}

Put gk(x)=τk(x)−xkg_{k}(x)=\tau_{k}(x)-x_{k}; then gkg_{k} is Borel by claim 2 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions. For N∈NN\in\mathbb{N} put

RN(x)=∑k=1Nak−1gk(x)2(x∈X),R_{N}(x)=\sum_{k=1}^{N}a_{k}^{-1}g_{k}(x)^{2}\qquad(x\in X),

a Borel function by claims 2 and 3 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions, with 0≤RN(x)≤RN+1(x)0\le R_{N}(x)\le R_{N+1}(x), the added term aN+1−1gN+1(x)2a_{N+1}^{-1}g_{N+1}(x)^{2} being nonnegative. Let

B=⋃M∈N ⋂N∈N{x∈X: RN(x)≤M},B=\bigcup_{M\in\mathbb{N}}\ \bigcap_{N\in\mathbb{N}}\{x\in X:\ R_{N}(x)\le M\},

the natural numbers MM being read as real numbers. Each set {x∈X:RN(x)≤M}\{x\in X:R_{N}(x)\le M\} is the complement of {x∈X:RN(x)>M}\{x\in X:R_{N}(x)>M\}, which belongs to B(X)\mathcal{B}(X) by the criterion of Measure Spaces and the Lebesgue Integral: Standing Notation §measurable (instantiated at the measurable space (X,B(X))(X,\mathcal{B}(X))); so B∈B(X)B\in\mathcal{B}(X) by Sigma-Algebra and Measurable Space. Thus x∈Bx\in B exactly when there is M∈NM\in\mathbb{N} with RN(x)≤MR_{N}(x)\le M for every N∈NN\in\mathbb{N}.

Step 4 (The map TT). Let x∈Bx\in B and choose M∈NM\in\mathbb{N} as just described. For every N∈NN\in\mathbb{N}, since 0<ak≤aˉ0<a_{k}\le\bar{a} and ak−1gk(x)2≥0a_{k}^{-1}g_{k}(x)^{2}\ge0,

∑k=1Ngk(x)2=∑k=1Nak ak−1gk(x)2≤aˉ∑k=1Nak−1gk(x)2=aˉ RN(x)≤aˉ M,\sum_{k=1}^{N}g_{k}(x)^{2}=\sum_{k=1}^{N}a_{k}\,a_{k}^{-1}g_{k}(x)^{2}\le\bar{a}\sum_{k=1}^{N}a_{k}^{-1}g_{k}(x)^{2}=\bar{a}\,R_{N}(x)\le\bar{a}\,M,

so the series ∑k=1∞gk(x)2\sum_{k=1}^{\infty}g_{k}(x)^{2} of nonnegative terms converges by the criterion for nonnegative terms. Let g^k=1B gk\hat{g}_{k}=\mathbf{1}_{B}\,g_{k}, which is Borel by claims 1 and 3 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions; thus g^k(x)=gk(x)\hat{g}_{k}(x)=g_{k}(x) for x∈Bx\in B and g^k(x)=0\hat{g}_{k}(x)=0 for x∉Bx\notin B. For every x∈Xx\in X the series ∑k=1∞g^k(x)2\sum_{k=1}^{\infty}\hat{g}_{k}(x)^{2} converges: for x∈Bx\in B it is the series just treated, and for x∉Bx\notin B all its terms are 00. Hence the synthesis claim of lem:measurable-maps-hilbert-valued-2026a, applied to the functions g^k\hat{g}_{k}, shows that for every x∈Xx\in X the series

h(x)=∑k=1∞g^k(x) ekh(x)=\sum_{k=1}^{\infty}\hat{g}_{k}(x)\,e_{k}

converges in XX, that h:X→Xh:X\to X is Borel, and that ⟨h(x),ek⟩=g^k(x)\langle h(x),e_{k}\rangle=\hat{g}_{k}(x) for all x∈Xx\in X and k∈Nk\in\mathbb{N}. The identity map of XX is continuous, hence Borel by claim 3 of Borel Measurability and Bounded Integration on a Metric Space; so

T:X→X,T(x)=x+h(x),T:X\to X,\qquad T(x)=x+h(x),

is Borel by the operations claim of the same lemma.

For every x∈Xx\in X one has T(x)−x=h(x)∈XaT(x)-x=h(x)\in X^{a}. Indeed, for N∈NN\in\mathbb{N},

SN(h(x))=∑k=1Nak−1g^k(x)2,S_{N}(h(x))=\sum_{k=1}^{N}a_{k}^{-1}\hat{g}_{k}(x)^{2},

which equals RN(x)≤MR_{N}(x)\le M if x∈Bx\in B (with MM chosen for xx as above) and equals 00 if x∉Bx\notin B; so the sequence (SN(h(x)))N∈N(S_{N}(h(x)))_{N\in\mathbb{N}} is bounded above and h(x)∈Xah(x)\in X^{a} by The Noise Space is a Real Hilbert Space: Orthonormal Basis, Continuous Embedding, Partial Sums, Closed Balls and Borel Measurability §partial-sums. Moreover, for x∈Bx\in B and k∈Nk\in\mathbb{N}, by linearity of the inner product,

⟨T(x),ek⟩=xk+g^k(x)=xk+gk(x)=τk(x).(2)\langle T(x),e_{k}\rangle=x_{k}+\hat{g}_{k}(x)=x_{k}+g_{k}(x)=\tau_{k}(x).\tag{2}

Step 5 (π\pi is concentrated on the graph of TT). Let z∈Gz\in G, with x=π1(z)x=\pi_{1}(z) and y=π2(z)y=\pi_{2}(z). For every k∈Nk\in\mathbb{N} we have z∈Gkz\in G_{k}, that is pk(y)=Tk(x)p_{k}(y)=T_{k}(x), so by (1) τk(x)=(pk(y))k=yk\tau_{k}(x)=(p_{k}(y))_{k}=y_{k}, and therefore

gk(x)=yk−xk=⟨y−x,ek⟩.g_{k}(x)=y_{k}-x_{k}=\langle y-x,e_{k}\rangle .

Consequently RN(x)=∑k=1Nak−1⟨y−x,ek⟩2=SN(y−x)R_{N}(x)=\sum_{k=1}^{N}a_{k}^{-1}\langle y-x,e_{k}\rangle^{2}=S_{N}(y-x) for every N∈NN\in\mathbb{N}. Since z∈Daz\in D_{a}, y−x∈Xay-x\in X^{a}, so by The Noise Space is a Real Hilbert Space: Orthonormal Basis, Continuous Embedding, Partial Sums, Closed Balls and Borel Measurability §partial-sums the sequence (SN(y−x))N∈N(S_{N}(y-x))_{N\in\mathbb{N}} is bounded above with least upper bound ∣y−x∣a2|y-x|_{a}^{2}. By claim 1 of The Archimedean Property of the Real Numbers there is M∈NM\in\mathbb{N} with ∣y−x∣a2<M|y-x|_{a}^{2}<M, and then RN(x)≤MR_{N}(x)\le M for every N∈NN\in\mathbb{N}; so x∈Bx\in B. By (2), ⟨T(x),ek⟩=τk(x)=yk\langle T(x),e_{k}\rangle=\tau_{k}(x)=y_{k}, so ⟨T(x)−y,ek⟩=0\langle T(x)-y,e_{k}\rangle=0 for every k∈Nk\in\mathbb{N}, and T(x)−y=0XT(x)-y=0_{X} because (ek)k∈N(e_{k})_{k\in\mathbb{N}} is an orthonormal basis. Thus y=T(x)y=T(x), that is z∈ΓTz\in\Gamma_{T}; we have shown G⊆ΓTG\subseteq\Gamma_{T}.

Since TT is Borel and (X,d)(X,d) is separable, Pairings into a Product, the Graph of a Measurable Map, and Couplings Concentrated on a Graph §graph-measurable gives ΓT∈B(X)⊗B(X)=B(X×X)\Gamma_{T}\in\mathcal{B}(X)\otimes\mathcal{B}(X)=\mathcal{B}(X\times X), and monotonicity gives 1=π(G)≤π(ΓT)≤π(X×X)=11=\pi(G)\le\pi(\Gamma_{T})\le\pi(X\times X)=1, so π(ΓT)=1\pi(\Gamma_{T})=1.

Step 6 (TT is a noise-optimal map inducing π\pi). The measure π\pi is a probability measure on (X×X,B(X)⊗B(X))(X\times X,\mathcal{B}(X)\otimes\mathcal{B}(X)) with (π1)#π=μ(\pi_{1})_{\#}\pi=\mu and π(ΓT)=1\pi(\Gamma_{T})=1, and TT is Borel; so Pairings into a Product, the Graph of a Measurable Map, and Couplings Concentrated on a Graph §graph, applied with μ\mu, TT in place of SS, and π\pi, gives

π=(id,T)#μ,ν=(π2)#π=T#μ.\pi=(\mathrm{id},T)_{\#}\mu,\qquad \nu=(\pi_{2})_{\#}\pi=T_{\#}\mu .

The map (id,T):X→X×X(\mathrm{id},T):X\to X\times X is Borel by Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §pairing, and the function cac_{a} of The Noise Space is a Real Hilbert Space: Orthonormal Basis, Continuous Embedding, Partial Sums, Closed Balls and Borel Measurability §pairs is Borel and nonnegative with ca((id,T)(x))=na(T(x)−x)c_{a}((\mathrm{id},T)(x))=n_{a}(T(x)-x) for every x∈Xx\in X, by the definition of cac_{a}. Claim 2 of Image Measures, Measures with Densities, and Change of Variables, applied to the measure space (X,B(X),μ)(X,\mathcal{B}(X),\mu), the map (id,T)(\mathrm{id},T) and the function cac_{a}, therefore gives

∫Xna(T(x)−x) μ(dx)=∫Xca∘(id,T) dμ=∫X×Xca d((id,T)#μ)=∫X×Xca dπ<∞,\int_{X}n_{a}\bigl(T(x)-x\bigr)\,\mu(dx)=\int_{X}c_{a}\circ(\mathrm{id},T)\,d\mu=\int_{X\times X}c_{a}\,d\bigl((\mathrm{id},T)_{\#}\mu\bigr)=\int_{X\times X}c_{a}\,d\pi<\infty,

the final inequality by Step 1. Finally (id,T)#μ=π(\mathrm{id},T)_{\#}\mu=\pi is noise-optimal. Together with Step 4 (TT Borel and T(x)−x∈XaT(x)-x\in X^{a} for every x∈Xx\in X) and T#μ=νT_{\#}\mu=\nu, this shows that TT is a noise-optimal map from μ\mu to ν\nu in the sense of Noise-Optimal Maps and Uniquely Noise-Mapped Pairs §map, and π=(id,T)#μ\pi=(\mathrm{id},T)_{\#}\mu. This proves claim 1.

Proof of claim 2.

Step 7 (A reference coupling and its map). The ordered pair (μ,ν)(\mu,\nu) is noise-connected by The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §connected, so The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §optimal gives a noise-optimal coupling π0∈Πa(μ,ν)\pi_{0}\in\Pi^{a}(\mu,\nu); we fix it. By claim 1, applied to π0\pi_{0}, we fix a noise-optimal map T0T_{0} from μ\mu to ν\nu with π0=(id,T0)#μ\pi_{0}=(\mathrm{id},T_{0})_{\#}\mu. Now let π\pi be an arbitrary noise-optimal coupling of μ\mu and ν\nu; we show π=π0\pi=\pi_{0}. Write W=Wa(μ,ν)W=W_{a}(\mu,\nu), the noise Wasserstein distance, so that ∫X×Xca dπ=Ia(π)=W2\int_{X\times X}c_{a}\,d\pi=I^{a}(\pi)=W^{2} and ∫X×Xca dπ0=Ia(π0)=W2\int_{X\times X}c_{a}\,d\pi_{0}=I^{a}(\pi_{0})=W^{2} by Couplings of Finite Noise Cost and Their Noise Cost §cost and Noise-Optimal Couplings §optimal.

Step 8 (An elementary fact about least upper bounds). Let (pm)m∈N(p_{m})_{m\in\mathbb{N}} and (qm)m∈N(q_{m})_{m\in\mathbb{N}} be nondecreasing sequences of real numbers, bounded above, with least upper bounds pp and qq. Then 12(p+q)\tfrac12(p+q) is the least upper bound of {12(pm+qm):m∈N}\{\tfrac12(p_{m}+q_{m}):m\in\mathbb{N}\}. Indeed, it is an upper bound since pm≤pp_{m}\le p and qm≤qq_{m}\le q. Given a real ε>0\varepsilon>0, choose m1,m2∈Nm_{1},m_{2}\in\mathbb{N} with pm1>p−εp_{m_{1}}>p-\varepsilon and qm2>q−εq_{m_{2}}>q-\varepsilon, and let mm be the larger of m1,m2m_{1},m_{2}; by monotonicity pm≥pm1p_{m}\ge p_{m_{1}} and qm≥qm2q_{m}\ge q_{m_{2}}, so 12(pm+qm)>12(p+q)−ε\tfrac12(p_{m}+q_{m})>\tfrac12(p+q)-\varepsilon, and no number below 12(p+q)\tfrac12(p+q) is an upper bound. (†)(\dagger)

Step 9 (The midpoint coupling is noise-optimal). Define πˉ:B(X×X)→[0,∞]\bar{\pi}:\mathcal{B}(X\times X)\to[0,\infty] by

πˉ(E)=12(π(E)+π0(E))(E∈B(X×X)),\bar{\pi}(E)=\tfrac12\bigl(\pi(E)+\pi_{0}(E)\bigr)\qquad(E\in\mathcal{B}(X\times X)),

a real number in [0,1][0,1] since π,π0∈P(X×X)\pi,\pi_{0}\in\mathcal{P}(X\times X). Then πˉ(∅)=0\bar{\pi}(\varnothing)=0. Let (Em)m∈N(E_{m})_{m\in\mathbb{N}} be pairwise disjoint members of B(X×X)\mathcal{B}(X\times X) with union EE. The partial sums Pn=∑m=1nπ(Em)P_{n}=\sum_{m=1}^{n}\pi(E_{m}) are real and nondecreasing, and Pn=π(⋃m=1nEm)≤1P_{n}=\pi(\bigcup_{m=1}^{n}E_{m})\le1 by finite additivity and monotonicity; so by Measure, Measure Space, and Probability Measure the countable additivity of π\pi says that π(E)\pi(E) is the least upper bound of {Pn:n∈N}\{P_{n}:n\in\mathbb{N}\}. Likewise π0(E)\pi_{0}(E) is the least upper bound of the partial sums Qn=∑m=1nπ0(Em)Q_{n}=\sum_{m=1}^{n}\pi_{0}(E_{m}). The partial sums of ∑mπˉ(Em)\sum_{m}\bar{\pi}(E_{m}) are 12(Pn+Qn)\tfrac12(P_{n}+Q_{n}), real and bounded above by 11, and by (†)(\dagger) their least upper bound is 12(π(E)+π0(E))=πˉ(E)\tfrac12(\pi(E)+\pi_{0}(E))=\bar{\pi}(E). So πˉ\bar{\pi} is a measure on B(X×X)\mathcal{B}(X\times X), that is a Borel measure, with πˉ(X×X)=1\bar{\pi}(X\times X)=1; thus πˉ∈P(X×X)\bar{\pi}\in\mathcal{P}(X\times X). For A∈B(X)A\in\mathcal{B}(X), πˉ(π1−1(A))=12(μ(A)+μ(A))=μ(A)\bar{\pi}(\pi_{1}^{-1}(A))=\tfrac12(\mu(A)+\mu(A))=\mu(A) and πˉ(π2−1(A))=12(ν(A)+ν(A))=ν(A)\bar{\pi}(\pi_{2}^{-1}(A))=\tfrac12(\nu(A)+\nu(A))=\nu(A), so πˉ∈Π(μ,ν)\bar{\pi}\in\Pi(\mu,\nu) by Couplings of Two Borel Probability Measures on a Hilbert Space and Their Quadratic Cost §coupling; and πˉ(Da)=12(1+1)=1\bar{\pi}(D_{a})=\tfrac12(1+1)=1.

We compute ∫X×Xca dπˉ\int_{X\times X}c_{a}\,d\bar{\pi}. By the approximation of nonnegative measurable functions by simple functions, applied on (X×X,B(X×X))(X\times X,\mathcal{B}(X\times X)) to the nonnegative Borel function cac_{a}, there is a sequence (sm)m∈N(s_{m})_{m\in\mathbb{N}} of nonnegative simple functions with sm≤sm+1≤cas_{m}\le s_{m+1}\le c_{a} and ca(z)c_{a}(z) the least upper bound of {sm(z):m∈N}\{s_{m}(z):m\in\mathbb{N}\} for every zz; this sequence does not involve the measure. Let λ\lambda be any of π,π0,πˉ\pi,\pi_{0},\bar{\pi}, and let s=∑i=1rti1Ais=\sum_{i=1}^{r}t_{i}\mathbf{1}_{A_{i}} be the standard representation (Simple Function and Its Integral) of a nonnegative simple function ss, so that each ti≥0t_{i}\ge0 and Ai∈B(X×X)A_{i}\in\mathcal{B}(X\times X). By claim 1 of Linearity and Monotonicity of the Lebesgue Integral (applied r−1r-1 times to sums and rr times to nonnegative multiples) and The Integral of an Indicator Function is the Measure of the Set,

∫X×Xs dλ=∑i=1rti λ(Ai),\int_{X\times X}s\,d\lambda=\sum_{i=1}^{r}t_{i}\,\lambda(A_{i}),

a real number. Since πˉ(Ai)=12(π(Ai)+π0(Ai))\bar{\pi}(A_{i})=\tfrac12(\pi(A_{i})+\pi_{0}(A_{i})), this gives

∫X×Xsm dπˉ=12(∫X×Xsm dπ+∫X×Xsm dπ0)(m∈N).\int_{X\times X}s_{m}\,d\bar{\pi}=\tfrac12\Bigl(\int_{X\times X}s_{m}\,d\pi+\int_{X\times X}s_{m}\,d\pi_{0}\Bigr)\qquad(m\in\mathbb{N}).

By the monotonicity part of claim 1 of Linearity and Monotonicity of the Lebesgue Integral, the sequences (∫sm dπ)m(\int s_{m}\,d\pi)_{m} and (∫sm dπ0)m(\int s_{m}\,d\pi_{0})_{m} are nondecreasing and bounded above by ∫ca dπ=W2\int c_{a}\,d\pi=W^{2} and ∫ca dπ0=W2\int c_{a}\,d\pi_{0}=W^{2}, and by Monotone Convergence Theorem their least upper bounds are ∫ca dπ=W2\int c_{a}\,d\pi=W^{2} and ∫ca dπ0=W2\int c_{a}\,d\pi_{0}=W^{2}. By (†)(\dagger) the least upper bound of {∫sm dπˉ:m∈N}\{\int s_{m}\,d\bar{\pi}:m\in\mathbb{N}\} is 12(W2+W2)=W2\tfrac12(W^{2}+W^{2})=W^{2}, and by Monotone Convergence Theorem applied with πˉ\bar{\pi} that least upper bound is ∫ca dπˉ\int c_{a}\,d\bar{\pi}. Hence

∫X×Xca dπˉ=W2<∞.\int_{X\times X}c_{a}\,d\bar{\pi}=W^{2}<\infty .

Together with πˉ(Da)=1\bar{\pi}(D_{a})=1 this shows πˉ∈Πa(μ,ν)\bar{\pi}\in\Pi^{a}(\mu,\nu) with Ia(πˉ)=Wa(μ,ν)2I^{a}(\bar{\pi})=W_{a}(\mu,\nu)^{2} (Couplings of Finite Noise Cost and Their Noise Cost §finite, Couplings of Finite Noise Cost and Their Noise Cost §cost), so πˉ\bar{\pi} is noise-optimal by Noise-Optimal Couplings §optimal.

Step 10 (π=π0\pi=\pi_{0}). By claim 1, applied to πˉ\bar{\pi}, there is a noise-optimal map Tˉ\bar{T} from μ\mu to ν\nu with πˉ=(id,Tˉ)#μ\bar{\pi}=(\mathrm{id},\bar{T})_{\#}\mu; we fix it. Its graph ΓTˉ\Gamma_{\bar{T}} belongs to B(X×X)\mathcal{B}(X\times X) by Pairings into a Product, the Graph of a Measurable Map, and Couplings Concentrated on a Graph §graph-measurable, and (id,Tˉ)(x)=(x,Tˉ(x))∈ΓTˉ(\mathrm{id},\bar{T})(x)=(x,\bar{T}(x))\in\Gamma_{\bar{T}} for every x∈Xx\in X, so πˉ(ΓTˉ)=μ(X)=1\bar{\pi}(\Gamma_{\bar{T}})=\mu(X)=1 and, by the rule for differences, πˉ(F)=0\bar{\pi}(F)=0 for F=(X×X)∖ΓTˉF=(X\times X)\setminus\Gamma_{\bar{T}}. Now π(F)≤π(F)+π0(F)=2πˉ(F)=0\pi(F)\le\pi(F)+\pi_{0}(F)=2\bar{\pi}(F)=0 and likewise π0(F)≤2πˉ(F)=0\pi_{0}(F)\le2\bar{\pi}(F)=0, so π(ΓTˉ)=π0(ΓTˉ)=1\pi(\Gamma_{\bar{T}})=\pi_{0}(\Gamma_{\bar{T}})=1 by the rule for differences. Since (π1)#π=(π1)#π0=μ(\pi_{1})_{\#}\pi=(\pi_{1})_{\#}\pi_{0}=\mu, Pairings into a Product, the Graph of a Measurable Map, and Couplings Concentrated on a Graph §graph, applied with μ\mu, Tˉ\bar{T} in place of SS, and π\pi, respectively π0\pi_{0}, gives

π=(id,Tˉ)#μ=π0.\pi=(\mathrm{id},\bar{T})_{\#}\mu=\pi_{0}.

Step 11 (Conclusion). By Steps 7 and 10, T0T_{0} is a noise-optimal map from μ\mu to ν\nu, fixed before π\pi was chosen, and every noise-optimal coupling π\pi of μ\mu and ν\nu equals π0=(id,T0)#μ\pi_{0}=(\mathrm{id},T_{0})_{\#}\mu. By Noise-Optimal Maps and Uniquely Noise-Mapped Pairs §uniquely-mapped, the ordered pair (μ,ν)(\mu,\nu) is uniquely noise-mapped. This proves claim 2.

Citations

Loading…

Dependencies

Uses0

Loading…

Comments

Log in to comment.

Loading…