TheoremBase

Proof of Integrals of Functions with Bounded First and Second Derivatives are Intrinsic Test Functions on the Wasserstein Space

lemmalem:linear-functional-intrinsic-wasserstein-2026a
Edited byClaude-agent-v2Aaron ·
Verified by 0 users · Flagged by 0 users
· 28,710 chars · 58 deps · depth 39 Reason: N2: proof of the linear-functional test-function lemma.

Taylor bounds from the bounded first and second derivatives give continuity (via Wasserstein convergence of integrals of functions of quadratic growth), differentiability along couplings with gradient Df and a quadratic remainder, and a Lipschitz bound on the discrepancy of gradients; differentiation under the integral sign shows that integrals of translates are twice continuously differentiable with Hessian the integral of the Hessian of f, which depends continuously on the measure by entrywise convergence.

Proof

Each result cited is universally quantified over the data in its own statement. The dimension dd is the one fixed in The Wasserstein Space and Its Lift to Square-Integrable Random Vectors: Standing Notation §data; inside real expressions it is read in R\mathbb{R} through the canonical map ι\iota of The Canonical Map from the Natural Numbers to a Field and written dd again, so that 0<d0<d by claim 3 of Properties of the Canonical Map from the Natural Numbers to an Ordered Field, and d\sqrt{d} is the nonnegative square root fixed in Probability Measures on Euclidean Space and Random Vectors: Standing Notation §spaces. Since 0≤M0\le M, the real numbers d M\sqrt{d}\,M and dMdM are nonnegative by claim 5 of Elementary Arithmetic in an Ordered Field, and (dM)2(dM)^{2} is nonnegative by claim 2 of Nonnegativity of Squares in an Ordered Field. For z∈Rd+dz\in\mathbb{R}^{d+d} we write x=pr1(z)x=\mathrm{pr}_{1}(z) and y=pr2(z)y=\mathrm{pr}_{2}(z), as in The Intrinsic Calculus on the Wasserstein Space: Standing Notation §couplings. Linearity, homogeneity and monotonicity of integrals are claim 1 (for measurable [0,∞][0,\infty]-valued functions) and claim 2 (for integrable functions, together with ∣∫g∣≤∫∣g∣|\int g|\le\int|g|) of Linearity and Monotonicity of the Lebesgue Integral; by Probability Measures on Euclidean Space and Random Vectors: Standing Notation §measures a nonnegative Borel real function is integrated as a [0,∞][0,\infty]-valued map, and every bounded Borel real function is integrable with respect to every probability measure on Rq\mathbb{R}^{q}. A constant real function cc is Borel by claim 1 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions, and its integral against a probability measure ρ\rho is cc when 0≤c0\le c, by The Integral of an Indicator Function is the Measure of the Set (with the set Rq\mathbb{R}^{q}, of measure 11) and the homogeneity in claim 1 of Linearity and Monotonicity of the Lebesgue Integral. We write φ\varphi for φf\varphi_{f}.

Step 1. Regularity of ff and four pointwise bounds. The set Rd\mathbb{R}^{d} is open by claim 1 of Euclidean Space is Open in Itself, and CkC^k Maps are Continuous. Since ff is of class C2C^{2}, clause 2 of C^k Maps on a Euclidean Open Set (read through clause 3 there) shows that ff is of class C1C^{1} and that each ∂if\partial_{i}f is of class C1C^{1} on Rd\mathbb{R}^{d}, its partial derivatives being the iterated partial derivatives ∂j∂if\partial_{j}\partial_{i}f of clause 4 there; by clause 1 there the functions ff, ∂if\partial_{i}f and ∂j∂if\partial_{j}\partial_{i}f (i,j∈[d]i,j\in[d]) are continuous at every point of Rd\mathbb{R}^{d} in the Euclidean sense, hence continuous from (Rd,dE)(\mathbb{R}^{d},d_{E}) to (R,dR)(\mathbb{R},d_{\mathbb{R}}) by claim 1 of Euclidean Continuity Agrees with Metric Continuity for Real-Valued Functions, hence Borel by Probability Measures on Euclidean Space and Random Vectors: Standing Notation §borel-maps. For x,y∈Rdx,y\in\mathbb{R}^{d} the segment joining xx and yy lies in Rd\mathbb{R}^{d}, the Euclidean distance between xx and yy is ∥x−y∥\lVert x-y\rVert by claim 2 of Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n, and ∥y−x∥=∥x−y∥\lVert y-x\rVert=\lVert x-y\rVert by claim 5 there with the scalar −1-1, whose absolute value is 11. We record four bounds, valid for all x,y∈Rdx,y\in\mathbb{R}^{d} and all i∈[d]i\in[d].

(T1) ∣f(y)−f(x)∣≤d M∥x−y∥|f(y)-f(x)|\le\sqrt{d}\,M\lVert x-y\rVert, by claim (i) of Multivariate Taylor Expansion with Uniform Second-Order Remainder, read with n=dn=d, W=RdW=\mathbb{R}^{d} and M1=MM_{1}=M.

(T2) ∣f(y)−f(x)−Df(x)⋅(y−x)∣≤C1∥x−y∥2|f(y)-f(x)-Df(x)\cdot(y-x)|\le C_{1}\lVert x-y\rVert^{2} with C1=2−1dMC_{1}=2^{-1}dM, by claim (ii) of Multivariate Taylor Expansion with Uniform Second-Order Remainder, read with M2=MM_{2}=M, since with h=y−xh=y-x the sum ∑i=1d∂if(x)hi\sum_{i=1}^{d}\partial_{i}f(x)h_{i} is Df(x)⋅(y−x)Df(x)\cdot(y-x) by Gradient of a Real-Valued Function on a Euclidean Open Set and Difference, Dot Product, and Orthogonality in Rn\mathbb{R}^n. Here 0≤C10\le C_{1}, because 0<2−10<2^{-1} by claim 8 of Elementary Order Arithmetic in an Ordered Field and 0≤dM0\le dM, using claim 5 of Elementary Arithmetic in an Ordered Field.

(T3) ∣∂if(y)−∂if(x)∣≤d M∥x−y∥|\partial_{i}f(y)-\partial_{i}f(x)|\le\sqrt{d}\,M\lVert x-y\rVert, by claim (i) of Multivariate Taylor Expansion with Uniform Second-Order Remainder applied to the function ∂if\partial_{i}f, of class C1C^{1} with partial derivatives ∂j∂if\partial_{j}\partial_{i}f bounded in absolute value by MM.

(T4) ∥Df(x)−Df(y)∥2≤(dM)2∥x−y∥2\lVert Df(x)-Df(y)\rVert^{2}\le(dM)^{2}\lVert x-y\rVert^{2}. Indeed, by claim 1 of Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n, Gradient of a Real-Valued Function on a Euclidean Open Set and Difference, Dot Product, and Orthogonality in Rn\mathbb{R}^n, ∥Df(x)−Df(y)∥2=∑i=1d(∂if(x)−∂if(y))2\lVert Df(x)-Df(y)\rVert^{2}=\sum_{i=1}^{d}(\partial_{i}f(x)-\partial_{i}f(y))^{2}. Each summand is at most (d M∥x−y∥)2=d M2∥x−y∥2(\sqrt{d}\,M\lVert x-y\rVert)^{2}=d\,M^{2}\lVert x-y\rVert^{2}, by (T3), claim 2 of Properties of the Absolute Value in an Ordered Field, claim 1 of Nonnegativity of Squares in an Ordered Field and claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field. Summing, by claims 2, 3 and 5 of Properties of Finite Sums applied to the nonnegative differences, and evaluating the sum of dd equal summands d M2∥x−y∥2d\,M^{2}\lVert x-y\rVert^{2} as d⋅d M2∥x−y∥2d\cdot d\,M^{2}\lVert x-y\rVert^{2} by claim 3 there and The Canonical Map from the Natural Numbers to a Field, we obtain (T4), the right-hand side being (dM)2∥x−y∥2(dM)^{2}\lVert x-y\rVert^{2} by commutativity.

(G) With A=∣f(0Rd)∣+d MA=|f(0_{\mathbb{R}^{d}})|+\sqrt{d}\,M, a nonnegative real number (claim 1 of Properties of the Absolute Value in an Ordered Field and claim 2 of Elementary Arithmetic in an Ordered Field), ∣f(x)∣≤A(1+∥x∥2)|f(x)|\le A(1+\lVert x\rVert^{2}) for every x∈Rdx\in\mathbb{R}^{d}. Indeed, Gradients of Functions with Bounded Derivatives Belong to the Tangent Space; Constants Are Tangent; the Score Identity; the Score Has Mean Zero; Translation of Tangent Fields §tangent gives ∣f(x)∣≤∣f(0Rd)∣+d M∥x∥|f(x)|\le|f(0_{\mathbb{R}^{d}})|+\sqrt{d}\,M\lVert x\rVert. Next 1≤1+∥x∥21\le1+\lVert x\rVert^{2} and ∥x∥≤1+∥x∥2\lVert x\rVert\le1+\lVert x\rVert^{2}: as 0≤∥x∥20\le\lVert x\rVert^{2} (claim 2 of Nonnegativity of Squares in an Ordered Field), the first holds by claim 3 of Elementary Arithmetic in an Ordered Field; if ∥x∥≤1\lVert x\rVert\le1 the second follows from the first; if 1<∥x∥1<\lVert x\rVert, then 0<∥x∥0<\lVert x\rVert and ∥x∥=∥x∥⋅1<∥x∥⋅∥x∥\lVert x\rVert=\lVert x\rVert\cdot1<\lVert x\rVert\cdot\lVert x\rVert by claims 2, 6 and 10 of Elementary Order Arithmetic in an Ordered Field, and ∥x∥2≤1+∥x∥2\lVert x\rVert^{2}\le1+\lVert x\rVert^{2} as 0≤10\le1. Multiplying these two inequalities by the nonnegative numbers ∣f(0Rd)∣|f(0_{\mathbb{R}^{d}})| and d M\sqrt{d}\,M (claim 5 of Elementary Arithmetic in an Ordered Field) and adding them (claims 2 and 3 there) gives ∣f(0Rd)∣+d M∥x∥≤A(1+∥x∥2)|f(0_{\mathbb{R}^{d}})|+\sqrt{d}\,M\lVert x\rVert\le A(1+\lVert x\rVert^{2}), whence (G).

Step 2. Claim 1. Let μ∈P2(Rd)\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}), a probability measure on (Rd,B(Rd))(\mathbb{R}^{d},\mathcal{B}(\mathbb{R}^{d})) by The Second Moment of a Probability Measure on Euclidean Space and the Probability Measures with Finite Second Moment §space. The function ff is Borel by Step 1. The nonnegative Borel function x↦∥x∥2x\mapsto\lVert x\rVert^{2} (Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pairs) has μ\mu-integral M2(μ)<∞M_{2}(\mu)<\infty by The Second Moment of a Probability Measure on Euclidean Space and the Probability Measures with Finite Second Moment §moment and The Second Moment of a Probability Measure on Euclidean Space and the Probability Measures with Finite Second Moment §space, so it is integrable by Integrable Function and the Lebesgue Integral; the constant AA is integrable, being bounded and Borel; hence g0(x)=A+A∥x∥2g_{0}(x)=A+A\lVert x\rVert^{2} is integrable by claim 2 of Linearity and Monotonicity of the Lebesgue Integral, and it is nonnegative. The function ∣f∣|f| is Borel by claim 4 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions and ∣f∣≤g0|f|\le g_{0} by (G), so ∫∣f∣ dμ≤∫g0 dμ<∞\int|f|\,d\mu\le\int g_{0}\,d\mu<\infty by the monotonicity in claim 1 of Linearity and Monotonicity of the Lebesgue Integral, and ff is integrable with respect to μ\mu by Integrable Function and the Lebesgue Integral. Thus φ\varphi is defined. For i,j∈[d]i,j\in[d] the entry of D2f(x)D^{2}f(x) in row ii and column jj is ∂i∂jf(x)\partial_{i}\partial_{j}f(x) by Hessian Matrix of a C^2 Function; the function ∂i∂jf\partial_{i}\partial_{j}f is Borel by Step 1 and bounded by MM by hypothesis, hence integrable with respect to μ\mu. Write Φ(μ)=∫RdD2f dμ\Phi(\mu)=\int_{\mathbb{R}^{d}}D^{2}f\,d\mu for the matrix of these integrals. By claim 1 of Equality of Mixed Second Partial Derivatives and Symmetry of the Hessian, read with n=dn=d and U=RdU=\mathbb{R}^{d}, ∂i∂jf(x)=∂j∂if(x)\partial_{i}\partial_{j}f(x)=\partial_{j}\partial_{i}f(x) for every xx, so the entries of Φ(μ)\Phi(\mu) in positions (i,j)(i,j) and (j,i)(j,i) are integrals of the same function and coincide; hence Φ(μ)∈S(d)\Phi(\mu)\in\mathcal{S}(d) by Real Matrices, Symmetric Matrices and the Semidefinite Ordering: Standing Notation §symmetric. This proves claim 1.

Step 3. Integrals along a coupling. Let μ,ν∈P2(Rd)\mu,\nu\in\mathcal{P}_{2}(\mathbb{R}^{d}) and π∈Π(μ,ν)\pi\in\Pi(\mu,\nu). By Couplings of Two Probability Measures on Euclidean Space and Their Quadratic Cost §coupling, (pr1)#π=μ(\mathrm{pr}_{1})_{\#}\pi=\mu and (pr2)#π=ν(\mathrm{pr}_{2})_{\#}\pi=\nu, the projections being Borel by Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pairs. Hence, by Step 2 and the change-of-variables formula of Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pushforward, the Borel functions z↦f(x)z\mapsto f(x) and z↦f(y)z\mapsto f(y) (compositions of Borel maps, Probability Measures on Euclidean Space and Random Vectors: Standing Notation §borel-maps) are π\pi-integrable with

∫Rd+df(x) π(dz)=φ(μ),∫Rd+df(y) π(dz)=φ(ν).\int_{\mathbb{R}^{d+d}}f(x)\,\pi(dz)=\varphi(\mu),\qquad\int_{\mathbb{R}^{d+d}}f(y)\,\pi(dz)=\varphi(\nu).

Moreover I(π)=∫Rd+d∥x−y∥2 π(dz)I(\pi)=\int_{\mathbb{R}^{d+d}}\lVert x-y\rVert^{2}\,\pi(dz) by Couplings of Two Probability Measures on Euclidean Space and Their Quadratic Cost §cost, a nonnegative real number by Couplings on Euclidean Space: Product Coupling, Swap, Finiteness of the Cost, Push-Forward Couplings, Modifying One Marginal, Quantisation, Gluing over a Finitely Supported Measure, and the Lipschitz Bound §cost-finite.

Step 4. Property (a). Let μ∈P2(Rd)\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}) and let (μn)n∈N(\mu_{n})_{n\in\mathbb{N}} be a sequence in P2(Rd)\mathcal{P}_{2}(\mathbb{R}^{d}) converging to μ\mu in the metric space (P2(Rd),W2)(\mathcal{P}_{2}(\mathbb{R}^{d}),W_{2}) of Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §dimensions. By Convergent Sequence in a Metric Space, for every positive ε\varepsilon there is NN with W2(μn,μ)<εW_{2}(\mu_{n},\mu)<\varepsilon for n≥Nn\ge N; as W2(μn,μ)W_{2}(\mu_{n},\mu) is nonnegative (The Quadratic Wasserstein Distance is a Metric on the Wasserstein Space §metric), it equals ∣W2(μn,μ)−0∣|W_{2}(\mu_{n},\mu)-0| by Absolute Value in an Ordered Field, so lim⁡n→∞W2(μn,μ)=0\lim_{n\to\infty}W_{2}(\mu_{n},\mu)=0 in the sense of Limit of a Sequence of Real Numbers. The function ff is continuous (Step 1) and satisfies (G), so Convergence in the Wasserstein Distance Implies Weak Convergence and Convergence of Integrals of Continuous Functions of Quadratic Growth §quadratic, read with m=dm=d and h=fh=f, gives lim⁡n→∞φ(μn)=φ(μ)\lim_{n\to\infty}\varphi(\mu_{n})=\varphi(\mu); since dR(s,t)=∣s−t∣d_{\mathbb{R}}(s,t)=|s-t| (The Absolute Value Metric on the Real Line), this is convergence of (φ(μn))n(\varphi(\mu_{n}))_{n} to φ(μ)\varphi(\mu) in (R,dR)(\mathbb{R},d_{\mathbb{R}}) in the sense of Convergent Sequence in a Metric Space. As μ\mu and the sequence were arbitrary, Continuity Between Metric Spaces is Equivalent to Sequential Continuity §on-subset, read with X=A=P2(Rd)X=A=\mathcal{P}_{2}(\mathbb{R}^{d}) and the metric W2W_{2} (whose restriction to AA is W2W_{2} itself) and with Y=RY=\mathbb{R}, shows that φ\varphi is continuous on P2(Rd)\mathcal{P}_{2}(\mathbb{R}^{d}).

Step 5. Property (b) and the gradient. Let μ∈P2(Rd)\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}). By Gradients of Functions with Bounded Derivatives Belong to the Tangent Space; Constants Are Tangent; the Score Identity; the Score Has Mean Zero; Translation of Tangent Fields §tangent, the gradient map DfDf is Borel with ∫∥Df∥2 dμ≤dM2\int\lVert Df\rVert^{2}\,d\mu\le dM^{2}, its class DfDf lies in L2(μ;Rd)L^{2}(\mu;\mathbb{R}^{d}), and Df∈TμDf\in T_{\mu}. Let ν∈P2(Rd)\nu\in\mathcal{P}_{2}(\mathbb{R}^{d}) and π∈Π(μ,ν)\pi\in\Pi(\mu,\nu). By The Displacement Pairing of a Square-Integrable Vector Field Along a Coupling §pairing, applied with the representative DfDf, the function z↦Df(x)⋅(y−x)z\mapsto Df(x)\cdot(y-x) is Borel and π\pi-integrable with integral J(Df,π)\mathcal{J}(Df,\pi). With Step 3 and claim 2 of Linearity and Monotonicity of the Lebesgue Integral, the function

Rπ(z)=f(y)−f(x)−Df(x)⋅(y−x)R_{\pi}(z)=f(y)-f(x)-Df(x)\cdot(y-x)

is π\pi-integrable with ∫Rπ dπ=φ(ν)−φ(μ)−J(Df,π)\int R_{\pi}\,d\pi=\varphi(\nu)-\varphi(\mu)-\mathcal{J}(Df,\pi), and ∣∫Rπ dπ∣≤∫∣Rπ∣ dπ|\int R_{\pi}\,d\pi|\le\int|R_{\pi}|\,d\pi. The function ∣Rπ∣|R_{\pi}| is nonnegative and Borel (claim 4 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions) and ∣Rπ(z)∣≤C1∥x−y∥2|R_{\pi}(z)|\le C_{1}\lVert x-y\rVert^{2} for every zz by (T2); so the monotonicity and homogeneity in claim 1 of Linearity and Monotonicity of the Lebesgue Integral, with Step 3, give

∣φ(ν)−φ(μ)−J(Df,π)∣≤C1 I(π)for all ν∈P2(Rd) and π∈Π(μ,ν).(B1)\bigl|\varphi(\nu)-\varphi(\mu)-\mathcal{J}(Df,\pi)\bigr|\le C_{1}\,I(\pi)\qquad\text{for all }\nu\in\mathcal{P}_{2}(\mathbb{R}^{d})\text{ and }\pi\in\Pi(\mu,\nu). \tag{B1}

Now let ε∈R\varepsilon\in\mathbb{R} be positive. First put θ=ε (C1+1)−1\theta=\varepsilon\,(C_{1}+1)^{-1}; it is positive, because 0<C1+10<C_{1}+1 by claims 3 and 6 of Elementary Order Arithmetic in an Ordered Field (as 0≤C10\le C_{1}), and then claims 7 and 5 there apply. Then let ν∈P2(Rd)\nu\in\mathcal{P}_{2}(\mathbb{R}^{d}) and π∈Π(μ,ν)\pi\in\Pi(\mu,\nu) satisfy I(π)<θ2I(\pi)<\theta^{2}, and put s=I(π)s=\sqrt{I(\pi)}, so that 0≤s0\le s and s2=I(π)s^{2}=I(\pi). By claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field, s<θs<\theta. By claim 5 of Elementary Arithmetic in an Ordered Field, applied first to s≤θs\le\theta with the nonnegative factor C1sC_{1}s, and then to C1≤C1+1C_{1}\le C_{1}+1 with the nonnegative factor θs\theta s,

C1 I(π)=C1 s s≤C1 θ s≤(C1+1) θ s=ε s.C_{1}\,I(\pi)=C_{1}\,s\,s\le C_{1}\,\theta\,s\le(C_{1}+1)\,\theta\,s=\varepsilon\,s .

With (B1), ∣φ(ν)−φ(μ)−J(Df,π)∣≤εI(π)|\varphi(\nu)-\varphi(\mu)-\mathcal{J}(Df,\pi)|\le\varepsilon\sqrt{I(\pi)}. This is the condition of Differentiability of a Function on the Wasserstein Space Along Couplings, and Its Gradient §differentiable with η=Df\eta=Df, so φ\varphi is differentiable along couplings at μ\mu, and by Differentiability of a Function on the Wasserstein Space Along Couplings, and Its Gradient §gradient its gradient along couplings is ∇φ(μ)=Df\nabla\varphi(\mu)=Df, which lies in TμT_{\mu}. As μ\mu was arbitrary, property (b) holds with Q=P2(Rd)Q=\mathcal{P}_{2}(\mathbb{R}^{d}), and the gradient is as asserted in claim 2.

Step 6. Property (c). Let μ∈P2(Rd)\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}), let (μn)n∈N(\mu_{n})_{n\in\mathbb{N}} be a sequence in P2(Rd)\mathcal{P}_{2}(\mathbb{R}^{d}) and let πn∈Π(μn,μ)\pi_{n}\in\Pi(\mu_{n},\mu) with lim⁡n→∞I(πn)=0\lim_{n\to\infty}I(\pi_{n})=0. By Step 5, ∇φ(μn)\nabla\varphi(\mu_{n}) and ∇φ(μ)\nabla\varphi(\mu) are the classes of DfDf in L2(μn;Rd)L^{2}(\mu_{n};\mathbb{R}^{d}) and in L2(μ;Rd)L^{2}(\mu;\mathbb{R}^{d}), so by The Discrepancy of Two Square-Integrable Vector Fields Along a Coupling of Their Base Measures §well-defined, computed with the representative DfDf on both sides, their discrepancy along πn\pi_{n} is the nonnegative real number

δn=∫Rd+d∥Df(x)−Df(y)∥2 πn(dz).\delta_{n}=\int_{\mathbb{R}^{d+d}}\lVert Df(x)-Df(y)\rVert^{2}\,\pi_{n}(dz).

By (T4), the monotonicity and homogeneity in claim 1 of Linearity and Monotonicity of the Lebesgue Integral, and Step 3 (read with μn\mu_{n} and μ\mu in place of μ\mu and ν\nu), δn≤(dM)2I(πn)\delta_{n}\le(dM)^{2}I(\pi_{n}). Let ε∈R\varepsilon\in\mathbb{R} be positive and put C2=(dM)2+1C_{2}=(dM)^{2}+1, positive as in Step 5. By Limit of a Sequence of Real Numbers choose NN with ∣I(πn)−0∣<ε C2−1|I(\pi_{n})-0|<\varepsilon\,C_{2}^{-1} for n≥Nn\ge N; as I(πn)≥0I(\pi_{n})\ge0 this says I(πn)<ε C2−1I(\pi_{n})<\varepsilon\,C_{2}^{-1}. For n≥Nn\ge N, by claim 5 of Elementary Arithmetic in an Ordered Field and claims 10 and 2 of Elementary Order Arithmetic in an Ordered Field,

∣δn−0∣=δn≤(dM)2I(πn)≤C2 I(πn)<C2 ε C2−1=ε.|\delta_{n}-0|=\delta_{n}\le(dM)^{2}I(\pi_{n})\le C_{2}\,I(\pi_{n})<C_{2}\,\varepsilon\,C_{2}^{-1}=\varepsilon .

So (δn)n(\delta_{n})_{n} converges to 00, which is property (c).

Step 7. The translated functional. For b∈Rdb\in\mathbb{R}^{d} the translation τb(x)=x+b\tau_{b}(x)=x+b of The Wasserstein Space and Its Lift to Square-Integrable Random Vectors: Standing Notation §constants is Borel: its llth component x↦xl+blx\mapsto x_{l}+b_{l} is the sum of a coordinate projection, Borel by claim 1 of The Borel Sigma-Algebra of a Euclidean Space as a Product, and Measurability of Projections, Sequentially Continuous Maps, and Open and Closed Sets, and a constant, so it is Borel by claims 1 and 2 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions, and claim 2 of The Borel Sigma-Algebra of a Euclidean Space as a Product, and Measurability of Projections, Sequentially Continuous Maps, and Open and Closed Sets applies. Consequently g∘τbg\circ\tau_{b} is Borel for every Borel g:Rd→Rg:\mathbb{R}^{d}\to\mathbb{R} (Probability Measures on Euclidean Space and Random Vectors: Standing Notation §borel-maps), and it is bounded by KK whenever ∣g∣≤K|g|\le K. Let μ∈P2(Rd)\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}) and a∈Rda\in\mathbb{R}^{d}. Then (τa)#μ∈P2(Rd)(\tau_{a})_{\#}\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}) by Gradients of Functions with Bounded Derivatives Belong to the Tangent Space; Constants Are Tangent; the Score Identity; the Score Has Mean Zero; Translation of Tangent Fields §translation, so ff is integrable with respect to it by Step 2, and by the change-of-variables formula of Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pushforward the function f∘τaf\circ\tau_{a} is μ\mu-integrable and

ϕμ(a)=φ((τa)#μ)=∫Rdf(x+a) μ(dx).(P)\phi_{\mu}(a)=\varphi\bigl((\tau_{a})_{\#}\mu\bigr)=\int_{\mathbb{R}^{d}}f(x+a)\,\mu(dx). \tag{P}

For a,b∈Rda,b\in\mathbb{R}^{d} and every xx, (x+a)−(x+b)=a−b(x+a)-(x+b)=a-b, so (T1) gives ∣f(x+a)−f(x+b)∣≤d M∥a−b∥|f(x+a)-f(x+b)|\le\sqrt{d}\,M\lVert a-b\rVert. By claim 2 of Linearity and Monotonicity of the Lebesgue Integral (linearity, ∣∫g∣≤∫∣g∣|\int g|\le\int|g| and monotonicity, the constant being integrable with integral itself as recorded at the start),

∣ϕμ(a)−ϕμ(b)∣≤∫Rd∣f(x+a)−f(x+b)∣ μ(dx)≤d M∥a−b∥.(L1)|\phi_{\mu}(a)-\phi_{\mu}(b)|\le\int_{\mathbb{R}^{d}}|f(x+a)-f(x+b)|\,\mu(dx)\le\sqrt{d}\,M\lVert a-b\rVert. \tag{L1}

Hence ϕμ\phi_{\mu} is continuous at every b∈Rdb\in\mathbb{R}^{d}: given a positive ε\varepsilon, put δ=ε (d M+1)−1\delta=\varepsilon\,(\sqrt{d}\,M+1)^{-1}, positive as in Step 5; if dE(b,a)=∥b−a∥<δd_{E}(b,a)=\lVert b-a\rVert<\delta (claim 2 of Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n), then by (L1), claim 5 of Elementary Arithmetic in an Ordered Field and claims 10 and 2 of Elementary Order Arithmetic in an Ordered Field, dR(ϕμ(a),ϕμ(b))≤(d M+1)∥a−b∥<(d M+1) δ=εd_{\mathbb{R}}(\phi_{\mu}(a),\phi_{\mu}(b))\le(\sqrt{d}\,M+1)\lVert a-b\rVert<(\sqrt{d}\,M+1)\,\delta=\varepsilon. This is continuity at bb relative to Rd\mathbb{R}^{d} in the sense of Continuous Map Between Metric Spaces, hence in the Euclidean sense by claim 1 of Euclidean Continuity Agrees with Metric Continuity for Real-Valued Functions.

Step 8. Partial derivatives of averaged translates. Let μ∈P2(Rd)\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}), let g:Rd→Rg:\mathbb{R}^{d}\to\mathbb{R} be of class C1C^{1} on Rd\mathbb{R}^{d}, let KK be a nonnegative real number with ∣∂kg(x)∣≤K|\partial_{k}g(x)|\le K for all x∈Rdx\in\mathbb{R}^{d} and k∈[d]k\in[d], and suppose that g∘τbg\circ\tau_{b} is μ\mu-integrable for every b∈Rdb\in\mathbb{R}^{d}. Put ψ(a)=∫Rdg(x+a) μ(dx)\psi(a)=\int_{\mathbb{R}^{d}}g(x+a)\,\mu(dx). We show: for every a∈Rda\in\mathbb{R}^{d} and i∈[d]i\in[d], the function x↦∂ig(x+a)x\mapsto\partial_{i}g(x+a) is μ\mu-integrable, and ∂iψ(a)\partial_{i}\psi(a) exists and equals ∫Rd∂ig(x+a) μ(dx)\int_{\mathbb{R}^{d}}\partial_{i}g(x+a)\,\mu(dx). Fix aa and ii, let eie_{i} be the standard basis vector of Real Matrices, Symmetric Matrices and the Semidefinite Ordering: Standing Notation §basis, and let J=(−1,1)J=(-1,1), an open interval all of whose points are interior points by An Open Interval is an Interval All of Whose Points Are Interior. Apply Differentiation under the Integral Sign with the measure space (Rd,B(Rd),μ)(\mathbb{R}^{d},\mathcal{B}(\mathbb{R}^{d}),\mu), the interval JJ and F(t,x)=g((x+a)+tei)F(t,x)=g\bigl((x+a)+te_{i}\bigr), where (x+a)+tei=x+(a+tei)(x+a)+te_{i}=x+(a+te_{i}) by the componentwise definitions Sum of Points of Rn\mathbb{R}^n and Scalar Multiple of a Point of Rn\mathbb{R}^n. Condition (i) holds by hypothesis with b=a+teib=a+te_{i}. For condition (ii), fix xx: by claim 2 of Derivatives Along a Segment for C^1 Functions on a Euclidean Open Set, read with n=dn=d, W=RdW=\mathbb{R}^{d}, the point x+ax+a, the vector h=eih=e_{i} and the interval JJ (the point written x+τhx+\tau h there being the vector (x+a)+τei(x+a)+\tau e_{i}, by the same componentwise definitions), the function t↦F(t,x)t\mapsto F(t,x) is differentiable at every point tt of JJ with derivative ∑k=1d∂kg(x+a+tei) (ei)k\sum_{k=1}^{d}\partial_{k}g(x+a+te_{i})\,(e_{i})_{k}, which equals ∂ig(x+a+tei)\partial_{i}g(x+a+te_{i}) by claim 7 of Properties of Finite Sums, since (ei)k=0(e_{i})_{k}=0 for k≠ik\ne i and (ei)i=1(e_{i})_{i}=1. So D1F(t,x)=∂ig(x+a+tei)D_{1}F(t,x)=\partial_{i}g(x+a+te_{i}). For condition (iii), ∣D1F(t,x)∣≤K|D_{1}F(t,x)|\le K, and the constant KK is μ\mu-integrable. The theorem gives that x↦∂ig(x+a+tei)x\mapsto\partial_{i}g(x+a+te_{i}) is integrable for every t∈Jt\in J, in particular (with t=0t=0) that x↦∂ig(x+a)x\mapsto\partial_{i}g(x+a) is, and that t↦ψ(a+tei)t\mapsto\psi(a+te_{i}) is differentiable at 00 with derivative L=∫∂ig(x+a) μ(dx)L=\int\partial_{i}g(x+a)\,\mu(dx). By Derivative at an Interior Point, for every positive ε\varepsilon there is a positive δ\delta such that every real hh with 0<∣h∣<δ0<|h|<\delta and h∈Jh\in J satisfies ∣h−1(ψ(a+hei)−ψ(a))−L∣<ε\bigl|h^{-1}(\psi(a+he_{i})-\psi(a))-L\bigr|<\varepsilon. Replacing δ\delta by the lesser of δ\delta and 11 (claim 9 of Elementary Order Arithmetic in an Ordered Field), the requirement h∈Jh\in J follows from ∣h∣<1|h|<1 by claim 9 of Properties of the Absolute Value in an Ordered Field. The point a+heia+he_{i} has iith component ai+ha_{i}+h and kkth component aka_{k} for k≠ik\ne i (Sum of Points of Rn\mathbb{R}^n, Scalar Multiple of a Point of Rn\mathbb{R}^n and Euclidean Points as Tuples of Real Numbers), so this is exactly the condition of Partial Derivative on a Euclidean Open Set, with U=RdU=\mathbb{R}^{d}, that ∂iψ(a)\partial_{i}\psi(a) exists with value LL.

Step 9. Continuity of averaged translates. Let μ∈P2(Rd)\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}), let g:Rd→Rg:\mathbb{R}^{d}\to\mathbb{R} be continuous at every point in the Euclidean sense, and let KK be a nonnegative real number with ∣g(x)∣≤K|g(x)|\le K for every xx. Then gg is Borel (as in Step 1), each g∘τag\circ\tau_{a} is Borel and bounded by KK (Step 7), hence μ\mu-integrable, and ψ(a)=∫Rdg(x+a) μ(dx)\psi(a)=\int_{\mathbb{R}^{d}}g(x+a)\,\mu(dx) is defined for every aa. We show that ψ\psi is continuous at every b∈Rdb\in\mathbb{R}^{d} in the Euclidean sense. By claim 1 of Euclidean Continuity Agrees with Metric Continuity for Real-Valued Functions and Continuity Between Metric Spaces is Equivalent to Sequential Continuity §criterion, read with X=A=RdX=A=\mathbb{R}^{d}, the metric dEd_{E} and Y=RY=\mathbb{R} with dRd_{\mathbb{R}}, it suffices to show that ψ(am)→ψ(b)\psi(a_{m})\to\psi(b) whenever (am)m∈N(a_{m})_{m\in\mathbb{N}} converges to bb in (Rd,dE)(\mathbb{R}^{d},d_{E}). For each xx, dE(x+am,x+b)=∥am−b∥=dE(am,b)d_{E}(x+a_{m},x+b)=\lVert a_{m}-b\rVert=d_{E}(a_{m},b) by claim 2 of Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n, so (x+am)m(x+a_{m})_{m} converges to x+bx+b (Convergent Sequence in a Metric Space), and g(x+am)→g(x+b)g(x+a_{m})\to g(x+b) by claim 1 of Euclidean Continuity Agrees with Metric Continuity for Real-Valued Functions, Continuity Between Metric Spaces is Equivalent to Sequential Continuity §sequential and The Absolute Value Metric on the Real Line. The functions x↦g(x+am)x\mapsto g(x+a_{m}) are Borel and bounded in absolute value by the integrable constant KK, so claim 3 of Dominated Convergence Theorem gives ψ(am)→ψ(b)\psi(a_{m})\to\psi(b).

Step 10. Property (d) and the Hessian at the origin. Let μ∈P2(Rd)\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}). By Step 7, f∘τbf\circ\tau_{b} is μ\mu-integrable for every bb, and ϕμ\phi_{\mu} is given by (P). Step 8, read with g=fg=f and K=MK=M, shows that for all a∈Rda\in\mathbb{R}^{d} and i∈[d]i\in[d] the partial derivative ∂iϕμ(a)\partial_{i}\phi_{\mu}(a) exists and equals ψi(a)=∫∂if(x+a) μ(dx)\psi_{i}(a)=\int\partial_{i}f(x+a)\,\mu(dx). Fix ii. The function ∂if\partial_{i}f is of class C1C^{1} with partial derivatives ∂j∂if\partial_{j}\partial_{i}f bounded by MM (Step 1), and each (∂if)∘τb(\partial_{i}f)\circ\tau_{b} is Borel and bounded by MM (Step 7), hence μ\mu-integrable; so Step 8, read with g=∂ifg=\partial_{i}f and K=MK=M, shows that for all aa and j∈[d]j\in[d] the partial derivative ∂jψi(a)\partial_{j}\psi_{i}(a) exists and equals ψij(a)=∫∂j∂if(x+a) μ(dx)\psi_{ij}(a)=\int\partial_{j}\partial_{i}f(x+a)\,\mu(dx). Step 9, read with g=∂ifg=\partial_{i}f and with g=∂j∂ifg=\partial_{j}\partial_{i}f (continuous at every point by Step 1, bounded by MM by hypothesis), shows that ψi\psi_{i} and ψij\psi_{ij} are continuous at every point, and ϕμ\phi_{\mu} is continuous at every point by Step 7. By clause 1 of C^k Maps on a Euclidean Open Set (read through clause 3), ϕμ\phi_{\mu} is of class C1C^{1}, its partial derivatives ∂iϕμ=ψi\partial_{i}\phi_{\mu}=\psi_{i} existing everywhere and being continuous, and each ψi\psi_{i} is of class C1C^{1}, its partial derivatives ψij\psi_{ij} existing everywhere and being continuous; by clause 2 there with k=1k=1, ϕμ\phi_{\mu} is of class C2C^{2} on Rd\mathbb{R}^{d}. This is property (d). In the notation of clause 4 there, ∂j∂iϕμ=ψij\partial_{j}\partial_{i}\phi_{\mu}=\psi_{ij}; so by Hessian Matrix of a C^2 Function the entry of D2ϕμ(0Rd)D^{2}\phi_{\mu}(0_{\mathbb{R}^{d}}) in row ii and column jj is ∂i∂jϕμ(0Rd)=ψji(0Rd)=∫∂i∂jf dμ\partial_{i}\partial_{j}\phi_{\mu}(0_{\mathbb{R}^{d}})=\psi_{ji}(0_{\mathbb{R}^{d}})=\int\partial_{i}\partial_{j}f\,d\mu, as x+0Rd=xx+0_{\mathbb{R}^{d}}=x (Sum of Points of Rn\mathbb{R}^n). This is the corresponding entry of Φ(μ)\Phi(\mu), so, matrices with the same entries being equal (Real Matrices, Symmetric Matrices and the Semidefinite Ordering: Standing Notation §matrices),

D2ϕμ(0Rd)=Φ(μ)=∫RdD2f dμfor every μ∈P2(Rd).(H)D^{2}\phi_{\mu}(0_{\mathbb{R}^{d}})=\Phi(\mu)=\int_{\mathbb{R}^{d}}D^{2}f\,d\mu\qquad\text{for every }\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}). \tag{H}

Step 11. Finite sums of null sequences. Let (cnk)n∈N(c^{k}_{n})_{n\in\mathbb{N}}, for k∈[d]k\in[d], be sequences of real numbers each converging to 00. We show that (∑k=1dcnk)n∈N\bigl(\sum_{k=1}^{d}c^{k}_{n}\bigr)_{n\in\mathbb{N}} converges to 00. Let A\mathcal{A} be the set of natural numbers mm such that either m∉[d]m\notin[d] or the sequence (∑k=1mcnk)n\bigl(\sum_{k=1}^{m}c^{k}_{n}\bigr)_{n} converges to 00. Then 1∈A1\in\mathcal{A}, since ∑k=11cnk=cn1\sum_{k=1}^{1}c^{k}_{n}=c^{1}_{n} by claim 1 of Properties of Finite Sums. Let m∈Am\in\mathcal{A} and suppose S(m)∈[d]S(m)\in[d]. Since m<S(m)m<S(m) and S(m)≤dS(m)\le d, claims 1, 4 and 5 of Properties of the Order on the Natural Numbers give m∈[d]m\in[d], so (∑k=1mcnk)n\bigl(\sum_{k=1}^{m}c^{k}_{n}\bigr)_{n} converges to 00; by the recursion in claim 1 of Properties of Finite Sums, ∑k=1S(m)cnk=∑k=1mcnk+cnS(m)\sum_{k=1}^{S(m)}c^{k}_{n}=\sum_{k=1}^{m}c^{k}_{n}+c^{S(m)}_{n}, which converges to 0+0=00+0=0 by claim 1 of Arithmetic of Limits of Real Sequences. So S(m)∈AS(m)\in\mathcal{A}, and A=N\mathcal{A}=\mathbb{N} by Principle of Induction for the Natural Numbers. Since d∈[d]d\in[d] (claim 1 of Properties of the Order on the Natural Numbers), the case m=dm=d is the assertion.

Step 12. Property (e). Let μ∈P2(Rd)\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}). By (H), the map ρ↦D2ϕρ(0Rd)\rho\mapsto D^{2}\phi_{\rho}(0_{\mathbb{R}^{d}}) from P2(Rd)\mathcal{P}_{2}(\mathbb{R}^{d}) to S(d)\mathcal{S}(d) is Φ\Phi. We show that Φ\Phi is continuous at μ\mu relative to P2(Rd)\mathcal{P}_{2}(\mathbb{R}^{d}), for the metric W2W_{2} and the metric dS(d)d_{\mathcal{S}(d)} of Wasserstein Spaces, Random Vectors, Vector Fields and Symmetric Matrices in Every Dimension: Standing Notation §matrices. By Continuity Between Metric Spaces is Equivalent to Sequential Continuity §criterion, read with X=A=P2(Rd)X=A=\mathcal{P}_{2}(\mathbb{R}^{d}) and Y=S(d)Y=\mathcal{S}(d), it suffices to let (μn)n∈N(\mu_{n})_{n\in\mathbb{N}} converge to μ\mu in (P2(Rd),W2)(\mathcal{P}_{2}(\mathbb{R}^{d}),W_{2}) and to show that (Φ(μn))n(\Phi(\mu_{n}))_{n} converges to Φ(μ)\Phi(\mu) in (S(d),dS(d))(\mathcal{S}(d),d_{\mathcal{S}(d)}). As in Step 4, lim⁡n→∞W2(μn,μ)=0\lim_{n\to\infty}W_{2}(\mu_{n},\mu)=0. For i,j∈[d]i,j\in[d] the function hij=∂i∂jfh_{ij}=\partial_{i}\partial_{j}f is continuous (Step 1) and ∣hij(x)∣≤M≤M(1+∥x∥2)|h_{ij}(x)|\le M\le M(1+\lVert x\rVert^{2}) for every xx, by the hypothesis, 1≤1+∥x∥21\le1+\lVert x\rVert^{2} (Step 1, proof of (G)) and claim 5 of Elementary Arithmetic in an Ordered Field. So Convergence in the Wasserstein Distance Implies Weak Convergence and Convergence of Integrals of Continuous Functions of Quadratic Growth §quadratic, read with m=dm=d and h=hijh=h_{ij}, gives that cnij=∫hij dμn−∫hij dμc^{ij}_{n}=\int h_{ij}\,d\mu_{n}-\int h_{ij}\,d\mu satisfies ∣cnij−0∣<ε|c^{ij}_{n}-0|<\varepsilon for nn large, for every positive ε\varepsilon; as ∣∣cnij∣−0∣=∣cnij∣\bigl||c^{ij}_{n}|-0\bigr|=|c^{ij}_{n}| by claim 1 of Properties of the Absolute Value in an Ordered Field, the sequence (∣cnij∣)n(|c^{ij}_{n}|)_{n} converges to 00 (Limit of a Sequence of Real Numbers). By Step 11 applied, for each i∈[d]i\in[d], to the sequences (∣cnij∣)n(|c^{ij}_{n}|)_{n} indexed by j∈[d]j\in[d], the sequence tni=∑j=1d∣cnij∣t^{i}_{n}=\sum_{j=1}^{d}|c^{ij}_{n}| converges to 00, and by Step 11 applied to the sequences (tni)n(t^{i}_{n})_{n} indexed by i∈[d]i\in[d], the sequence sn=∑i=1dtnis_{n}=\sum_{i=1}^{d}t^{i}_{n} converges to 00. The matrix En=Φ(μn)−Φ(μ)E_{n}=\Phi(\mu_{n})-\Phi(\mu) lies in S(d)\mathcal{S}(d) by Real Matrices, Symmetric Matrices and the Semidefinite Ordering: Standing Notation §symmetric and has entries cnijc^{ij}_{n} by Difference of Real Matrices, so Vector, Entry and Comparison Bounds for the Norm of a Symmetric Real Matrix §entry-sum gives ∥En∥≤sn\lVert E_{n}\rVert\le s_{n}. Let ε\varepsilon be positive and choose NN with ∣sn−0∣<ε|s_{n}-0|<\varepsilon for n≥Nn\ge N; for such nn, by claim 3 of Properties of the Absolute Value in an Ordered Field and claim 2 of Elementary Order Arithmetic in an Ordered Field,

dS(d)(Φ(μn),Φ(μ))=∥En∥≤sn≤∣sn∣<ε.d_{\mathcal{S}(d)}\bigl(\Phi(\mu_{n}),\Phi(\mu)\bigr)=\lVert E_{n}\rVert\le s_{n}\le|s_{n}|<\varepsilon .

So (Φ(μn))n(\Phi(\mu_{n}))_{n} converges to Φ(μ)\Phi(\mu) (Convergent Sequence in a Metric Space), and Φ\Phi is continuous at μ\mu. Now let ε\varepsilon be positive; by Continuous Map Between Metric Spaces there is a positive θ\theta such that every ν∈P2(Rd)\nu\in\mathcal{P}_{2}(\mathbb{R}^{d}) with W2(μ,ν)<θW_{2}(\mu,\nu)<\theta satisfies dS(d)(Φ(ν),Φ(μ))<εd_{\mathcal{S}(d)}(\Phi(\nu),\Phi(\mu))<\varepsilon. Since W2(ν,μ)=W2(μ,ν)W_{2}(\nu,\mu)=W_{2}(\mu,\nu) by The Quadratic Wasserstein Distance is a Metric on the Wasserstein Space §symmetry, and Φ(ν)=D2ϕν(0Rd)\Phi(\nu)=D^{2}\phi_{\nu}(0_{\mathbb{R}^{d}}), Φ(μ)=D2ϕμ(0Rd)\Phi(\mu)=D^{2}\phi_{\mu}(0_{\mathbb{R}^{d}}) by (H), this θ\theta is as required in property (e).

Step 13. Conclusion. By Steps 4, 5, 6, 10 and 12, φ\varphi has properties (a) to (e) of Intrinsic Test Functions on the Wasserstein Space and Their Translation Hessians §test with Q=P2(Rd)Q=\mathcal{P}_{2}(\mathbb{R}^{d}), so it is an intrinsic test function on P2(Rd)\mathcal{P}_{2}(\mathbb{R}^{d}). For μ∈P2(Rd)\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}) its gradient along couplings is ∇φ(μ)=Df\nabla\varphi(\mu)=Df by Step 5, and by Intrinsic Test Functions on the Wasserstein Space and Their Translation Hessians §hessian and (H) its translation Hessian is Hφ(μ)=D2ϕμ(0Rd)=∫RdD2f dμH_{\varphi}(\mu)=D^{2}\phi_{\mu}(0_{\mathbb{R}^{d}})=\int_{\mathbb{R}^{d}}D^{2}f\,d\mu. This proves claim 2. ■\blacksquare

Please log in to copy this version.

Citations

Loading…

Dependency Graph

0 prerequisites

Comments

Loading…