TheoremBase

Nonnegativity is the Gibbs inequality. The first variation splits into the Gaussian entropy part, handled by the published noise-gradient perturbation lemma, and the potential part, differentiated through the tangent inequality and dominated convergence. Density uses tail replacement, radial retraction of the rescaled head onto a ball and Gaussian smoothing, which gives finite Fisher information and integrability of the potential and its squared slope.

Proof

Each result cited is universally quantified over the data in its own statement.

Real order and arithmetic are those of The Real Numbers: Standing Notation and Background §background, and absolute values of real numbers obey Properties of the Absolute Value in an Ordered Field. Throughout, ρ=γc\rho=\gamma_{c} as in A Diagonal Gaussian Reference Measure on the Noise Wasserstein Space, Rescaled Heads and Gaussian Tails: Standing Notation §gaussian, and D\mathcal{D}, DΣ\mathcal{D}_{\Sigma}, E\mathcal{E} and Σ\Sigma are as in The Gibbs Entropy Pair on the Noise Wasserstein Space: Relative Entropy to the Gibbs Measure and the Gibbs Score Field §pair: D\mathcal{D} is the set of the μ∈P(X)\mu\in\mathcal{P}(X) of finite relative entropy with respect to γβV\gamma^{V}_{\beta}, DΣ\mathcal{D}_{\Sigma} the set of the μ∈D\mu\in\mathcal{D} that have a relative score with respect to γβV\gamma^{V}_{\beta} and finite Fisher information relative to γβV\gamma^{V}_{\beta} with weights aa (The Gibbs Entropy Pair on the Noise Wasserstein Space: Relative Entropy to the Gibbs Measure and the Gibbs Score Field §penalty-domain, The Gibbs Entropy Pair on the Noise Wasserstein Space: Relative Entropy to the Gibbs Measure and the Gibbs Score Field §score-domain), E(μ)=β H(μ ∣ γβV)\mathcal{E}(\mu)=\beta\,H(\mu\,|\,\gamma^{V}_{\beta}) and Σ(μ)=βZμa+∇aV\Sigma(\mu)=\beta Z^{a}_{\mu}+\nabla_{a}V. Write V=v∘pdV=v\circ p_{d} with head dimension dd, profile vv and semiconvexity constant K≥0K\ge0, as in Admissible Cylindrical Potentials on a Hilbert Space in the Noise Geometry §admissible, with noise gradient ∇aV\nabla_{a}V and functions ∂kV\partial_{k}V of Admissible Cylindrical Potentials on a Hilbert Space in the Noise Geometry §gradient. Fix b∈Rb\in\mathbb{R} as in Admissible Cylindrical Potentials on a Hilbert Space in the Noise Geometry §below and C∈RC\in\mathbb{R} as in Admissible Cylindrical Potentials on a Hilbert Space in the Noise Geometry §slope; by Basic Properties of an Admissible Cylindrical Potential: Continuity, Growth under Noise Translations, Integrability, the Tangent Inequality and Tangency of the Noise Gradient §translation, C≥0C\ge0, and we let C1/2≥0C^{1/2}\ge0 be its square root and L=C1/2(∣b+1∣+2)≥0L=C^{1/2}(|b+1|+2)\ge0 the number of that clause. exp⁡\exp obeys Basic Properties of the Exponential Function; in particular it is increasing (claim 4), positive (claim 2) and exp⁡(u)2=exp⁡(2u)\exp(u)^{2}=\exp(2u) (claim 1). For k∈Nk\in\mathbb{N}, ak1/2>0a_{k}^{1/2}>0 is the square root of aka_{k} and ak−1/2a_{k}^{-1/2} its inverse, as in Lifting Euclidean Tangent Fields of the Rescaled Head to Noise Tangent Fields on a Hilbert Space. The weak form below is claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field: for real α,α′≥0\alpha,\alpha'\ge0, α≤α′\alpha\le\alpha' if and only if α2≤α′2\alpha^{2}\le\alpha'^{2}. Integrals, their linearity and monotonicity are those of Linearity and Monotonicity of the Lebesgue Integral §nonnegative and Linearity and Monotonicity of the Lebesgue Integral §integrable; push-forwards on XX and Rn\mathbb{R}^{n} and the change of variables formula, including the transfer of integrability, are those of Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §pushforward and claim 2 of Image Measures, Measures with Densities, and Change of Variables.

Step 0 (pointwise bounds for the potential). Put U(x)=V(x)+b+1U(x)=V(x)+b+1 for x∈Xx\in X; UU is continuous, hence Borel, since VV is by Basic Properties of an Admissible Cylindrical Potential: Continuity, Growth under Noise Translations, Integrability, the Tangent Inequality and Tangency of the Noise Gradient §continuity.

(0.a) Since −b≤V(x)-b\le V(x) by Basic Properties of an Admissible Cylindrical Potential: Continuity, Growth under Noise Translations, Integrability, the Tangent Inequality and Tangency of the Noise Gradient §continuity, 1≤U(x)1\le U(x) for every x∈Xx\in X.

(0.b) For x∈Xx\in X, ∣V(x)∣=∣U(x)−(b+1)∣≤∣U(x)∣+∣b+1∣=U(x)+∣b+1∣|V(x)|=|U(x)-(b+1)|\le|U(x)|+|b+1|=U(x)+|b+1| by the triangle inequality and (0.a); as 1≤U(x)1\le U(x) and 0≤1+∣b+1∣0\le1+|b+1|, 1+∣b+1∣≤(1+∣b+1∣)U(x)1+|b+1|\le(1+|b+1|)U(x), so

1+∣V(x)∣≤(2+∣b+1∣) U(x)(x∈X).1+|V(x)|\le(2+|b+1|)\,U(x)\qquad(x\in X).

(0.c) For x∈Xx\in X, Basic Properties of an Admissible Cylindrical Potential: Continuity, Growth under Noise Translations, Integrability, the Tangent Inequality and Tangency of the Noise Gradient §continuity, the definition ∂kV(x)=∂kv(pd(x))\partial_{k}V(x)=\partial_{k}v(p_{d}(x)) for k≤dk\le d of Admissible Cylindrical Potentials on a Hilbert Space in the Noise Geometry §gradient, V(x)=v(pd(x))V(x)=v(p_{d}(x)) and Admissible Cylindrical Potentials on a Hilbert Space in the Noise Geometry §slope at u=pd(x)u=p_{d}(x) give

∣∇aV(x)∣a2=∑i=1dai ∂iV(x)2≤C(1+∣V(x)∣)2.|\nabla_{a}V(x)|_{a}^{2}=\sum_{i=1}^{d}a_{i}\,\partial_{i}V(x)^{2}\le C\bigl(1+|V(x)|\bigr)^{2}.

Let k≤dk\le d. Each summand being nonnegative, ak ∂kV(x)2≤C(1+∣V(x)∣)2≤C(2+∣b+1∣)2U(x)2a_{k}\,\partial_{k}V(x)^{2}\le C(1+|V(x)|)^{2}\le C(2+|b+1|)^{2}U(x)^{2}, by (0.b), the weak form and C≥0C\ge0. Multiplying by ak−1=(ak−1/2)2a_{k}^{-1}=(a_{k}^{-1/2})^{2} and putting Γk=ak−1/2C1/2(2+∣b+1∣)≥0\Gamma_{k}=a_{k}^{-1/2}C^{1/2}(2+|b+1|)\ge0, we get ∣∂kV(x)∣2≤(ΓkU(x))2|\partial_{k}V(x)|^{2}\le(\Gamma_{k}U(x))^{2}, and the weak form gives

∣∂kV(x)∣≤Γk U(x)(x∈X, k≤d).|\partial_{k}V(x)|\le\Gamma_{k}\,U(x)\qquad(x\in X,\ k\le d).

(0.d) For x∈Xx\in X and h∈Xah\in X^{a}, Basic Properties of an Admissible Cylindrical Potential: Continuity, Growth under Noise Translations, Integrability, the Tangent Inequality and Tangency of the Noise Gradient §translation reads U(x+h)≤U(x)exp⁡(L∣h∣a)U(x+h)\le U(x)\exp(L|h|_{a}).

Step 1 (clause 1). Let μ∈D\mu\in\mathcal{D}. By The Gibbs Measure of an Admissible Cylindrical Potential Relative to a Diagonal Gaussian Measure §gibbs, γβV\gamma^{V}_{\beta} is a probability measure on (X,B(X))(X,\mathcal{B}(X)), and so is μ\mu; μ\mu has finite relative entropy with respect to γβV\gamma^{V}_{\beta}, so the Gibbs inequality Relative Entropy on a Measurable Space: the Gibbs Inequality, the Variational Criterion and Formula, the Entropy Inequality, Small Sets and Data Processing §gibbs gives 0≤H(μ ∣ γβV)0\le H(\mu\,|\,\gamma^{V}_{\beta}). Since β>0\beta>0, 0≤β H(μ ∣ γβV)=E(μ)0\le\beta\,H(\mu\,|\,\gamma^{V}_{\beta})=\mathcal{E}(\mu).

Step 2 (clause 2). Let μ∈D\mu\in\mathcal{D} and let ε∈R\varepsilon\in\mathbb{R} be positive. The Euclidean items cited for measures on Rn\mathbb{R}^{n} are read with nn in place of the dimension written dd (or mm) there, as fixed in A Diagonal Gaussian Reference Measure on the Noise Wasserstein Space, Rescaled Heads and Gaussian Tails: Standing Notation §euclidean; the letter dd keeps its meaning as the head dimension of VV.

(2.a) Tail replacement. By Entropy and Score Relative to the Gibbs Measure Split into Their Gaussian Parts and the Potential, with a Fisher Information Bound §entropy, μ\mu has finite relative entropy with respect to γc\gamma_{c}, and μ∈P2(X)\mu\in\mathcal{P}_{2}(X) by Entropy and Score Relative to the Gibbs Measure Split into Their Gaussian Parts and the Potential, with a Fisher Information Bound §domain. By the hypothesis ck≤κ akc_{k}\le\kappa\,a_{k} and Tail Replacement in the Noise Wasserstein Distance: a Measure of Finite Relative Entropy is Approximated by the Gaussian-Tail Extensions of Its Rescaled Heads §convergence, Wa(μ,Em(μ~m))→0W_{a}(\mu,E_{m}(\tilde{\mu}_{m}))\to0 as m→∞m\to\infty, so by Limit of a Sequence of Real Numbers there is N0∈NN_{0}\in\mathbb{N} with Wa(μ,Em(μ~m))<ε/3W_{a}(\mu,E_{m}(\tilde{\mu}_{m}))<\varepsilon/3 for every m≥N0m\ge N_{0}, where μ,Em(μ~m)∈Pρa\mu,E_{m}(\tilde{\mu}_{m})\in\mathcal{P}^{a}_{\rho} by Tail Replacement in the Noise Wasserstein Distance: a Measure of Finite Relative Entropy is Approximated by the Gaussian-Tail Extensions of Its Rescaled Heads §membership. Let nn be the larger of N0N_{0} and dd, and put λ=μ~n=(rn)#μ\lambda=\tilde{\mu}_{n}=(r_{n})_{\#}\mu; then Wa(μ,En(λ))<ε/3W_{a}(\mu,E_{n}(\lambda))<\varepsilon/3, and λ∈P2(Rn)\lambda\in\mathcal{P}_{2}(\mathbb{R}^{n}) by Lifting Euclidean Tangent Fields of the Rescaled Head to Noise Tangent Fields on a Hilbert Space §head.

(2.b) Retraction of the head onto a ball. By Radial Retraction onto a Closed Ball: Compactly Supported Approximation in the Wasserstein Distance §approximation with m=n≥1m=n\ge1, applied to λ\lambda and ε/3\varepsilon/3, there is a positive R0∈RR_{0}\in\mathbb{R} with W2((PR)#λ,λ)<ε/3W_{2}((P_{R})_{\#}\lambda,\lambda)<\varepsilon/3 for every real R≥R0R\ge R_{0}; put R=R0R=R_{0} and λ′=(PR)#λ\lambda'=(P_{R})_{\#}\lambda. Then λ′∈P2(Rn)\lambda'\in\mathcal{P}_{2}(\mathbb{R}^{n}) by the same clause, W2(λ′,λ)<ε/3W_{2}(\lambda',\lambda)<\varepsilon/3, and λ′(Bˉ(0Rn,R))=1\lambda'(\bar{B}(0_{\mathbb{R}^{n}},R))=1 by Radial Retraction onto a Closed Ball: Compactly Supported Approximation in the Wasserstein Distance §retraction.

(2.c) Gaussian smoothing. Let ss be the lesser of 12\tfrac12 and ε2/(18n)\varepsilon^{2}/(18n), with nn read in R\mathbb{R}; then 0<s≤120<s\le\tfrac12. Let σ\sigma be the measure written μs\mu_{s} in Gaussian Smoothing of a Measure with Finite Second Moment: Positive Density, Finite Entropy, a Tangent Score, Non-Increasing Fisher Information and Exponential Integrability, built from λ′\lambda' (in the role of μ\mu there) and ss. By Gaussian Smoothing of a Measure with Finite Second Moment: Positive Density, Finite Entropy, a Tangent Score, Non-Increasing Fisher Information and Exponential Integrability §density, σ∈P2(Rn)\sigma\in\mathcal{P}_{2}(\mathbb{R}^{n}) and W2(σ,λ′)2≤n s≤ε2/18<(ε/3)2W_{2}(\sigma,\lambda')^{2}\le n\,s\le\varepsilon^{2}/18<(\varepsilon/3)^{2}, so W2(σ,λ′)<ε/3W_{2}(\sigma,\lambda')<\varepsilon/3 by claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field, both numbers being nonnegative. By Gaussian Smoothing of a Measure with Finite Second Moment: Positive Density, Finite Entropy, a Tangent Score, Non-Increasing Fisher Information and Exponential Integrability §entropy and Gaussian Smoothing of a Measure with Finite Second Moment: Positive Density, Finite Entropy, a Tangent Score, Non-Increasing Fisher Information and Exponential Integrability §score, σ∈P2Ent(Rn)∩P2I(Rn)\sigma\in\mathcal{P}_{2}^{\mathrm{Ent}}(\mathbb{R}^{n})\cap\mathcal{P}_{2}^{\mathcal{I}}(\mathbb{R}^{n}). Since c~(n)\tilde{c}^{(n)} is a variance vector and γ~n\tilde{\gamma}_{n} the diagonal Gaussian measure with variances c~(n)\tilde{c}^{(n)} (A Diagonal Gaussian Reference Measure on the Noise Wasserstein Space, Rescaled Heads and Gaussian Tails: Standing Notation §heads), Relative Entropy and Relative Score with Respect to a Diagonal Gaussian Measure: Comparison with the Entropy, the Score and the Relative Free Energy of a Quadratic Potential §entropy and Relative Entropy and Relative Score with Respect to a Diagonal Gaussian Measure: Comparison with the Entropy, the Score and the Relative Free Energy of a Quadratic Potential §score, applied with the variance vector c~(n)\tilde{c}^{(n)} to σ∈P2(Rn)\sigma\in\mathcal{P}_{2}(\mathbb{R}^{n}), show that σ\sigma has finite relative entropy with respect to γ~n\tilde{\gamma}_{n} and finite Fisher information relative to γ~n\tilde{\gamma}_{n}.

(2.d) The Gaussian-tail extension. Put ν=En(σ)\nu=E_{n}(\sigma), the Gaussian-tail extension of σ\sigma at level nn. By Gaussian-Tail Extensions: the Gaussian Reference Measure, Marginals, Densities, Relative Entropy, Relative Score and the Noise Wasserstein Distance §second-moment, ν∈P2(X)\nu\in\mathcal{P}_{2}(X), in particular ν∈P(X)\nu\in\mathcal{P}(X); by Gaussian-Tail Extensions: the Gaussian Reference Measure, Marginals, Densities, Relative Entropy, Relative Score and the Noise Wasserstein Distance §entropy, ν\nu has finite relative entropy with respect to γc\gamma_{c}; by Gaussian-Tail Extensions: the Gaussian Reference Measure, Marginals, Densities, Relative Entropy, Relative Score and the Noise Wasserstein Distance §score and Gaussian-Tail Extensions: the Gaussian Reference Measure, Marginals, Densities, Relative Entropy, Relative Score and the Noise Wasserstein Distance §fisher, ν\nu has a relative score with respect to γc\gamma_{c} and finite Fisher information relative to γc\gamma_{c} with weights aa; and (rn)#ν=σ(r_{n})_{\#}\nu=\sigma by Gaussian-Tail Extensions: the Gaussian Reference Measure, Marginals, Densities, Relative Entropy, Relative Score and the Noise Wasserstein Distance §head-marginal.

(2.e) The potential under ν\nu. Let A0=v(0Rd)+b+1A_{0}=v(0_{\mathbb{R}^{d}})+b+1, which satisfies 1≤A01\le A_{0} by Admissible Cylindrical Potentials on a Hilbert Space in the Noise Geometry §below at u=0Rdu=0_{\mathbb{R}^{d}}, and let A=(2+∣b+1∣)2A02≥0A=(2+|b+1|)^{2}A_{0}^{2}\ge0. Define h:Rn→Rh:\mathbb{R}^{n}\to\mathbb{R} by h(y)=Aexp⁡(2L∥y∥)h(y)=A\exp(2L\lVert y\rVert); then 0≤h(y)=∣h(y)∣0\le h(y)=|h(y)|. The map y↦2L∥y∥y\mapsto2L\lVert y\rVert is Lipschitz with constant 2L2L from (Rn,dE)(\mathbb{R}^{n},d_{E}) to R\mathbb{R} with the absolute-value metric: for y,y′∈Rny,y'\in\mathbb{R}^{n}, Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n §distance, Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n §homogeneity and Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n §triangle give ∥y∥≤∥y−y′∥+∥y′∥\lVert y\rVert\le\lVert y-y'\rVert+\lVert y'\rVert and ∥y′∥≤∥y−y′∥+∥y∥\lVert y'\rVert\le\lVert y-y'\rVert+\lVert y\rVert, whence ∣2L∥y∥−2L∥y′∥∣≤2L dE(y,y′)|2L\lVert y\rVert-2L\lVert y'\rVert|\le2L\,d_{E}(y,y') by Properties of the Absolute Value in an Ordered Field §two-sided and Properties of the Absolute Value in an Ordered Field §multiplicative; so it is continuous by A Lipschitz Map is Uniformly Continuous. exp⁡\exp is smooth on R=R1\mathbb{R}=\mathbb{R}^{1} by claim 3 of Basic Properties of the Exponential Function, hence continuous on (R1,dE)(\mathbb{R}^{1},d_{E}) by claim 3 of Euclidean Space is Open in Itself, and CkC^k Maps are Continuous; dEd_{E} on R1\mathbb{R}^{1} is the absolute-value metric by The Euclidean Distance on the Real Line is the Absolute Value Metric, so exp⁡\exp is continuous on (R,dR)(\mathbb{R},d_{\mathbb{R}}). By claim 3 of Semicontinuity and Continuity Under Composition with a Continuous Map, y↦exp⁡(2L∥y∥)y\mapsto\exp(2L\lVert y\rVert) is continuous, hence Borel by Probability Measures on Euclidean Space and Random Vectors: Standing Notation §borel-maps, and so is its constant multiple hh. Since λ′(Bˉ(0Rn,R))=1\lambda'(\bar{B}(0_{\mathbb{R}^{n}},R))=1 and ∣h(y)∣≤Aexp⁡(2L∥y∥)|h(y)|\le A\exp(2L\lVert y\rVert) with AA and 2L2L nonnegative, Gaussian Smoothing of a Measure with Finite Second Moment: Positive Density, Finite Entropy, a Tangent Score, Non-Increasing Fisher Information and Exponential Integrability §exponential shows that hh is integrable with respect to σ\sigma, so ∫Rnh dσ<∞\int_{\mathbb{R}^{n}}h\,d\sigma<\infty.

Let x∈Xx\in X. By Lifting Euclidean Tangent Fields of the Rescaled Head to Noise Tangent Fields on a Hilbert Space, rn(x)r_{n}(x) has coordinates ak−1/2xka_{k}^{-1/2}x_{k} (k≤nk\le n), so Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n §square gives ∥rn(x)∥2=∑k=1nxk2/ak\lVert r_{n}(x)\rVert^{2}=\sum_{k=1}^{n}x_{k}^{2}/a_{k}. Since d≤nd\le n and each summand is nonnegative, ∣pd(x)∣a2=∑k=1dxk2/ak≤∥rn(x)∥2|p_{d}(x)|_{a}^{2}=\sum_{k=1}^{d}x_{k}^{2}/a_{k}\le\lVert r_{n}(x)\rVert^{2}, and the weak form gives ∣pd(x)∣a≤∥rn(x)∥|p_{d}(x)|_{a}\le\lVert r_{n}(x)\rVert, both being nonnegative. As L≥0L\ge0 and exp⁡\exp is increasing, Basic Properties of an Admissible Cylindrical Potential: Continuity, Growth under Noise Translations, Integrability, the Tangent Inequality and Tangency of the Noise Gradient §exponential gives U(x)≤A0exp⁡(L∣pd(x)∣a)≤A0exp⁡(L∥rn(x)∥)U(x)\le A_{0}\exp(L|p_{d}(x)|_{a})\le A_{0}\exp(L\lVert r_{n}(x)\rVert). With (0.b), 1+∣V(x)∣≤(2+∣b+1∣)A0exp⁡(L∥rn(x)∥)1+|V(x)|\le(2+|b+1|)A_{0}\exp(L\lVert r_{n}(x)\rVert), both sides being nonnegative, so the weak form and exp⁡(u)2=exp⁡(2u)\exp(u)^{2}=\exp(2u) give

(1+∣V(x)∣)2≤Aexp⁡(2L∥rn(x)∥)=h(rn(x))(x∈X).\bigl(1+|V(x)|\bigr)^{2}\le A\exp\bigl(2L\lVert r_{n}(x)\rVert\bigr)=h\bigl(r_{n}(x)\bigr)\qquad(x\in X).

The function (1+∣V∣)2(1+|V|)^{2} is Borel and nonnegative, VV being Borel, and h∘rnh\circ r_{n} is Borel, rnr_{n} being Borel by Lifting Euclidean Tangent Fields of the Rescaled Head to Noise Tangent Fields on a Hilbert Space §head. By monotonicity and the change of variables formula for the nonnegative Borel function hh and (rn)#ν=σ(r_{n})_{\#}\nu=\sigma,

∫X(1+∣V∣)2 dν≤∫Xh∘rn dν=∫Rnh dσ<∞.\int_{X}\bigl(1+|V|\bigr)^{2}\,d\nu\le\int_{X}h\circ r_{n}\,d\nu=\int_{\mathbb{R}^{n}}h\,d\sigma<\infty .

Since 1≤1+∣V(x)∣1\le1+|V(x)|, ∣V(x)∣≤(1+∣V(x)∣)2|V(x)|\le(1+|V(x)|)^{2}, so the Borel function VV is integrable with respect to ν\nu. By (0.c), ∣∇aV∣a2≤C(1+∣V∣)2|\nabla_{a}V|_{a}^{2}\le C(1+|V|)^{2}, where ∣∇aV∣a2=∑i=1dai(∂iV)2|\nabla_{a}V|_{a}^{2}=\sum_{i=1}^{d}a_{i}(\partial_{i}V)^{2} is Borel and nonnegative, each ∂iV\partial_{i}V being continuous by Basic Properties of an Admissible Cylindrical Potential: Continuity, Growth under Noise Translations, Integrability, the Tangent Inequality and Tangency of the Noise Gradient §continuity; so ∫X∣∇aV∣a2 dν≤C∫X(1+∣V∣)2 dν<∞\int_{X}|\nabla_{a}V|_{a}^{2}\,d\nu\le C\int_{X}(1+|V|)^{2}\,d\nu<\infty.

(2.f) Membership. By (2.d) and (2.e), ν∈P(X)\nu\in\mathcal{P}(X) has finite relative entropy with respect to γc\gamma_{c} and VV is integrable with respect to ν\nu, so ν\nu has finite relative entropy with respect to γβV\gamma^{V}_{\beta} by Entropy and Score Relative to the Gibbs Measure Split into Their Gaussian Parts and the Potential, with a Fisher Information Bound §entropy, that is, ν∈D\nu\in\mathcal{D}. Moreover ν∈P2(X)\nu\in\mathcal{P}_{2}(X), VV is integrable with respect to ν\nu, ν\nu has a relative score with respect to γc\gamma_{c} and finite Fisher information relative to γc\gamma_{c} with weights aa, and ∫X∣∇aV∣a2 dν<∞\int_{X}|\nabla_{a}V|_{a}^{2}\,d\nu<\infty; by the ``if'' direction of Entropy and Score Relative to the Gibbs Measure Split into Their Gaussian Parts and the Potential, with a Fisher Information Bound §splitting, ν\nu has a relative score with respect to γβV\gamma^{V}_{\beta} and finite Fisher information relative to γβV\gamma^{V}_{\beta} with weights aa. Hence ν∈DΣ\nu\in\mathcal{D}_{\Sigma} by The Gibbs Entropy Pair on the Noise Wasserstein Space: Relative Entropy to the Gibbs Measure and the Gibbs Score Field §score-domain.

(2.g) Distance. The measures σ\sigma, λ′\lambda' and λ\lambda lie in P2(Rn)\mathcal{P}_{2}(\mathbb{R}^{n}), so the triangle inequality The Quadratic Wasserstein Distance is a Metric on the Wasserstein Space §triangle gives W2(σ,λ)≤W2(σ,λ′)+W2(λ′,λ)<2ε/3W_{2}(\sigma,\lambda)\le W_{2}(\sigma,\lambda')+W_{2}(\lambda',\lambda)<2\varepsilon/3. By Gaussian-Tail Extensions: the Gaussian Reference Measure, Marginals, Densities, Relative Entropy, Relative Score and the Noise Wasserstein Distance §distance, applied to σ,λ∈P2(Rn)\sigma,\lambda\in\mathcal{P}_{2}(\mathbb{R}^{n}), ν\nu and En(λ)E_{n}(\lambda) belong to Pρa\mathcal{P}^{a}_{\rho} and Wa(ν,En(λ))≤W2(σ,λ)<2ε/3W_{a}(\nu,E_{n}(\lambda))\le W_{2}(\sigma,\lambda)<2\varepsilon/3. By the triangle inequality The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §triangle and the symmetry The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §symmetry for ν\nu, En(λ)E_{n}(\lambda), μ∈Pρa\mu\in\mathcal{P}^{a}_{\rho}, and (2.a),

Wa(ν,μ)≤Wa(ν,En(λ))+Wa(μ,En(λ))<2ε3+ε3=ε.W_{a}(\nu,\mu)\le W_{a}(\nu,E_{n}(\lambda))+W_{a}(\mu,E_{n}(\lambda))<\tfrac{2\varepsilon}{3}+\tfrac{\varepsilon}{3}=\varepsilon .

Since ν∈DΣ\nu\in\mathcal{D}_{\Sigma} by (2.f), this proves clause 2.

Step 3 (the first variation). Let μ∈DΣ\mu\in\mathcal{D}_{\Sigma} and ψ∈FCb2(X)\psi\in\mathcal{F}C^{2}_{b}(X). By Bounded C^2 Cylindrical Functions on a Hilbert Space with an Orthonormal Basis §cylindrical there are n∈Nn\in\mathbb{N} and g∈Cb2(Rn)g\in C^{2}_{b}(\mathbb{R}^{n}) with ψ=g∘pn\psi=g\circ p_{n}; as recorded in Relative Entropy along a Noise-Gradient Perturbation of the Identity: the Formula, the Second-Order Expansion and the First Variation, ψ∈FCb1(X)\psi\in\mathcal{F}C^{1}_{b}(X), with representation (n,g)(n,g). By Entropy and Score Relative to the Gibbs Measure Split into Their Gaussian Parts and the Potential, with a Fisher Information Bound §entropy, μ\mu has finite relative entropy with respect to γc\gamma_{c} and VV is integrable with respect to μ\mu; by The Gibbs Entropy Pair on the Noise Wasserstein Space: Relative Entropy to the Gibbs Measure and the Gibbs Score Field §score-domain, μ∈P2(X)\mu\in\mathcal{P}_{2}(X), μ\mu has a relative score with respect to γc\gamma_{c} and finite Fisher information relative to γc\gamma_{c} with weights aa, and ∫X∣∇aV∣a2 dμ<∞\int_{X}|\nabla_{a}V|_{a}^{2}\,d\mu<\infty.

(3.a) The noise gradient of ψ\psi. By The Noise Gradient of a Cylindrical Function and the Noise Tangent Space at a Probability Measure on a Hilbert Space §gradient, for x∈Xx\in X the kk-th coordinate ηk(x)\eta_{k}(x) of ∇aψ(x)∈Xa\nabla_{a}\psi(x)\in X^{a} is ak∂kψ(x)a_{k}\partial_{k}\psi(x) for k≤nk\le n and 00 for k>nk>n, each ηk\eta_{k} is Borel, and there are real Bk≥0B_{k}\ge0 (k≤nk\le n) with ∣∂kψ∣≤Bk|\partial_{k}\psi|\le B_{k} on XX and ∣∇aψ(x)∣a2≤Bψ|\nabla_{a}\psi(x)|_{a}^{2}\le B_{\psi} for every xx, where Bψ=∑k=1nakBk2≥0B_{\psi}=\sum_{k=1}^{n}a_{k}B_{k}^{2}\ge0. Put Nk=akBkN_{k}=a_{k}B_{k} for k≤nk\le n and Nk=0N_{k}=0 for k>nk>n, so that ∣ηk(x)∣≤Nk|\eta_{k}(x)|\le N_{k} for all x∈Xx\in X and k∈Nk\in\mathbb{N}; let bψ≥0b_{\psi}\ge0 be the square root of BψB_{\psi}, so that ∣∇aψ(x)∣a≤bψ|\nabla_{a}\psi(x)|_{a}\le b_{\psi} by the weak form; and put eψ=exp⁡(L bψ)e_{\psi}=\exp(L\,b_{\psi}). For t∈Rt\in\mathbb{R} let Tt=id+t∇aψT_{t}=\mathrm{id}+t\nabla_{a}\psi, a Borel map X→XX\to X by Noise Penalty Pairs on the Noise Wasserstein Space: the Penalty, Its Score, and Their Domains (preamble), and μt=(Tt)#μ∈P(X)\mu_{t}=(T_{t})_{\#}\mu\in\mathcal{P}(X); T0T_{0} is the identity, so μ0=μ\mu_{0}=\mu. For x∈Xx\in X, Tt(x)−x=t∇aψ(x)T_{t}(x)-x=t\nabla_{a}\psi(x) lies in the linear subspace XaX^{a} (The Noise Space is a Real Hilbert Space: Orthonormal Basis, Continuous Embedding, Partial Sums, Closed Balls and Borel Measurability §hilbert), has coordinates t ηk(x)t\,\eta_{k}(x) and satisfies ∣t∇aψ(x)∣a=∣t∣ ∣∇aψ(x)∣a≤∣t∣ bψ|t\nabla_{a}\psi(x)|_{a}=|t|\,|\nabla_{a}\psi(x)|_{a}\le|t|\,b_{\psi}; and x−Tt(x)=−t∇aψ(x)∈Xax-T_{t}(x)=-t\nabla_{a}\psi(x)\in X^{a}.

(3.b) The potential part is finite. Let t∈Rt\in\mathbb{R} with ∣t∣≤1|t|\le1 and x∈Xx\in X. By (0.d) with h=t∇aψ(x)h=t\nabla_{a}\psi(x), L ∣t∇aψ(x)∣a≤L bψL\,|t\nabla_{a}\psi(x)|_{a}\le L\,b_{\psi} (as L≥0L\ge0) and monotonicity of exp⁡\exp, with U(x)>0U(x)>0,

U(Tt(x))≤eψ U(x).(3.1)U(T_{t}(x))\le e_{\psi}\,U(x).\tag{3.1}

With (0.b), ∣V(Tt(x))∣≤1+∣V(Tt(x))∣≤(2+∣b+1∣) eψ U(x)|V(T_{t}(x))|\le1+|V(T_{t}(x))|\le(2+|b+1|)\,e_{\psi}\,U(x). The function V∘TtV\circ T_{t} is Borel, as a composition of Borel maps, and U=V+b+1U=V+b+1 is integrable with respect to μ\mu by linearity, VV being integrable and constants integrable against the probability measure μ\mu; so V∘TtV\circ T_{t} is integrable with respect to μ\mu by monotonicity, and by the change of variables formula VV is integrable with respect to μt\mu_{t}, with

F(t)=∫XV dμt=∫XV∘Tt dμ.F(t)=\int_{X}V\,d\mu_{t}=\int_{X}V\circ T_{t}\,d\mu .

(3.c) The Gaussian part. By Relative Entropy along a Noise-Gradient Perturbation of the Identity: the Formula, the Second-Order Expansion and the First Variation §variation, applied to μ\mu, nn and gg, there is a positive t0∈Rt_{0}\in\mathbb{R} such that μt\mu_{t} has finite relative entropy with respect to γc\gamma_{c} for every t∈(−t0,t0)t\in(-t_{0},t_{0}) and f(t)=H(μt ∣ γc)f(t)=H(\mu_{t}\,|\,\gamma_{c}) defines a function on (−t0,t0)(-t_{0},t_{0}) differentiable at 00 with derivative Lμa(g)L^{a}_{\mu}(g); there, id+t∇aψ\mathrm{id}+t\nabla_{a}\psi is the same map x↦x+t∇aψ(x)x\mapsto x+t\nabla_{a}\psi(x). By The Noise Score Field of a Measure of Finite Weighted Fisher Information: Existence, Norm, Pairing with Noise Gradients, Head Approximation and Tangency §pairing-functional, Lμa(g)=⟨Zμa,∇aψ⟩μL^{a}_{\mu}(g)=\langle Z^{a}_{\mu},\nabla_{a}\psi\rangle_{\mu}. Let t1t_{1} be the lesser of t0t_{0} and 11; then t1>0t_{1}>0, and 00 is an interior point of the interval (−t1,t1)(-t_{1},t_{1}), as recorded in Noise Penalty Pairs on the Noise Wasserstein Space: the Penalty, Its Score, and Their Domains §variation. The restriction f1f_{1} of ff to (−t1,t1)(-t_{1},t_{1}) is differentiable at 00 with derivative Lμa(g)L^{a}_{\mu}(g): in Derivative at an Interior Point, any δ\delta that serves for ff and a given ε\varepsilon serves for f1f_{1}, since every hh with 0+h∈(−t1,t1)0+h\in(-t_{1},t_{1}) has 0+h∈(−t0,t0)0+h\in(-t_{0},t_{0}) and the difference quotients agree.

(3.d) The penalty along the perturbation. Let t∈(−t1,t1)t\in(-t_{1},t_{1}). Then μt∈P(X)\mu_{t}\in\mathcal{P}(X) has finite relative entropy with respect to γc\gamma_{c} by (3.c) and VV is integrable with respect to μt\mu_{t} by (3.b), so by Entropy and Score Relative to the Gibbs Measure Split into Their Gaussian Parts and the Potential, with a Fisher Information Bound §entropy μt\mu_{t} has finite relative entropy with respect to γβV\gamma^{V}_{\beta}, that is, μt∈D\mu_{t}\in\mathcal{D}, and, multiplying the formula of that clause by β\beta,

E(μt)=β f1(t)+F(t)+βlog⁡ZV,β,(3.2)\mathcal{E}(\mu_{t})=\beta\,f_{1}(t)+F(t)+\beta\log Z_{V,\beta},\tag{3.2}

ZV,βZ_{V,\beta} being a positive real number by The Gibbs Measure of an Admissible Cylindrical Potential Relative to a Diagonal Gaussian Measure §normaliser.

(3.e) Differentiability of the potential part. Put

G=∫X∑k=1d∂kV ηk dμ.G=\int_{X}\sum_{k=1}^{d}\partial_{k}V\,\eta_{k}\,d\mu .

This is well defined: each ∂kV\partial_{k}V (k≤dk\le d) is integrable with respect to μ\mu by Basic Properties of an Admissible Cylindrical Potential: Continuity, Growth under Noise Translations, Integrability, the Tangent Inequality and Tangency of the Noise Gradient §integrable, and ∂kV ηk\partial_{k}V\,\eta_{k} is Borel with ∣∂kV ηk∣≤Nk∣∂kV∣|\partial_{k}V\,\eta_{k}|\le N_{k}|\partial_{k}V|. By Basic Properties of an Admissible Cylindrical Potential: Continuity, Growth under Noise Translations, Integrability, the Tangent Inequality and Tangency of the Noise Gradient §continuity with h=∇aψ(x)h=\nabla_{a}\psi(x), ⟨∇aV(x),∇aψ(x)⟩a=∑k=1d∂kV(x) ηk(x)\langle\nabla_{a}V(x),\nabla_{a}\psi(x)\rangle_{a}=\sum_{k=1}^{d}\partial_{k}V(x)\,\eta_{k}(x). The class of ∇aV\nabla_{a}V lies in L2(μ;Xa)L^{2}(\mu;X^{a}) by Basic Properties of an Admissible Cylindrical Potential: Continuity, Growth under Noise Translations, Integrability, the Tangent Inequality and Tangency of the Noise Gradient §tangent, and ∇aψ\nabla_{a}\psi is square-integrable with respect to μ\mu by The Noise Gradient of a Cylindrical Function and the Noise Tangent Space at a Probability Measure on a Hilbert Space §gradient; so by The Space of Square-Integrable Maps from a Measure Space into a Hilbert Space with an Orthonormal Basis §operations (through Probability Measures on a Hilbert Space Transported in the Noise Norm: Standing Notation §fields),

G=∫X⟨∇aV,∇aψ⟩a dμ=⟨∇aV,∇aψ⟩μ.G=\int_{X}\langle\nabla_{a}V,\nabla_{a}\psi\rangle_{a}\,d\mu=\langle\nabla_{a}V,\nabla_{a}\psi\rangle_{\mu}.

Pointwise estimate. Let x∈Xx\in X and t∈Rt\in\mathbb{R}, and write y=Tt(x)y=T_{t}(x). By (3.a) and Basic Properties of an Admissible Cylindrical Potential: Continuity, Growth under Noise Translations, Integrability, the Tangent Inequality and Tangency of the Noise Gradient §continuity, ⟨∇aV(x),y−x⟩a=t∑k=1d∂kV(x)ηk(x)\langle\nabla_{a}V(x),y-x\rangle_{a}=t\sum_{k=1}^{d}\partial_{k}V(x)\eta_{k}(x), ⟨∇aV(y),x−y⟩a=−t∑k=1d∂kV(y)ηk(x)\langle\nabla_{a}V(y),x-y\rangle_{a}=-t\sum_{k=1}^{d}\partial_{k}V(y)\eta_{k}(x), and ∣y−x∣a2=∣x−y∣a2=t2∣∇aψ(x)∣a2≤t2Bψ|y-x|_{a}^{2}=|x-y|_{a}^{2}=t^{2}|\nabla_{a}\psi(x)|_{a}^{2}\le t^{2}B_{\psi}. The tangent inequality Basic Properties of an Admissible Cylindrical Potential: Continuity, Growth under Noise Translations, Integrability, the Tangent Inequality and Tangency of the Noise Gradient §tangent-inequality, applied to the pair (x,y)(x,y) and to the pair (y,x)(y,x), gives with K≥0K\ge0

Dt(x):=V(y)−V(x)−t∑k=1d∂kV(x)ηk(x) ≥ −K2t2Bψ,Dt(x) ≤ t∑k=1d(∂kV(y)−∂kV(x))ηk(x)+K2t2Bψ.D_{t}(x):=V(y)-V(x)-t\sum_{k=1}^{d}\partial_{k}V(x)\eta_{k}(x)\ \ge\ -\frac{K}{2}t^{2}B_{\psi}, \qquad D_{t}(x)\ \le\ t\sum_{k=1}^{d}\bigl(\partial_{k}V(y)-\partial_{k}V(x)\bigr)\eta_{k}(x)+\frac{K}{2}t^{2}B_{\psi}.

Writing φt,k(x)=∣∂kV(Tt(x))−∂kV(x)∣\varphi_{t,k}(x)=|\partial_{k}V(T_{t}(x))-\partial_{k}V(x)|, the triangle inequality, multiplicativity and ∣ηk(x)∣≤Nk|\eta_{k}(x)|\le N_{k} bound the last sum by ∣t∣∑k=1dNkφt,k(x)|t|\sum_{k=1}^{d}N_{k}\varphi_{t,k}(x), and Properties of the Absolute Value in an Ordered Field §two-sided yields

∣Dt(x)∣≤∣t∣∑k=1dNk φt,k(x)+K2t2Bψ(x∈X, t∈R).(3.3)|D_{t}(x)|\le|t|\sum_{k=1}^{d}N_{k}\,\varphi_{t,k}(x)+\frac{K}{2}t^{2}B_{\psi}\qquad(x\in X,\ t\in\mathbb{R}).\tag{3.3}

Integration. Let t∈(−t1,t1)t\in(-t_{1},t_{1}) with t≠0t\ne0. By (3.b) and linearity, DtD_{t} is integrable with respect to μ\mu with ∫XDt dμ=F(t)−F(0)−tG\int_{X}D_{t}\,d\mu=F(t)-F(0)-tG. For k≤dk\le d, φt,k\varphi_{t,k} is Borel, ∂kV\partial_{k}V being continuous by Basic Properties of an Admissible Cylindrical Potential: Continuity, Growth under Noise Translations, Integrability, the Tangent Inequality and Tangency of the Noise Gradient §continuity and TtT_{t} Borel, and by (0.c) and (3.1) (as ∣t∣≤1|t|\le1), 0≤φt,k≤∣∂kV∘Tt∣+∣∂kV∣≤Γk(eψ+1) U0\le\varphi_{t,k}\le|\partial_{k}V\circ T_{t}|+|\partial_{k}V|\le\Gamma_{k}(e_{\psi}+1)\,U, an integrable function; so φt,k\varphi_{t,k} is integrable. Since −∣Dt∣≤Dt≤∣Dt∣-|D_{t}|\le D_{t}\le|D_{t}|, monotonicity, (3.3) and division by ∣t∣>0|t|>0 give

∣F(t)−F(0)t−G∣≤∑k=1dNk∫Xφt,k dμ+K2∣t∣ Bψ.(3.4)\Bigl|\frac{F(t)-F(0)}{t}-G\Bigr|\le\sum_{k=1}^{d}N_{k}\int_{X}\varphi_{t,k}\,d\mu+\frac{K}{2}|t|\,B_{\psi}.\tag{3.4}

Limit. Let (tj)j∈N(t_{j})_{j\in\mathbb{N}} be a sequence that is admissible for FF on (−t1,t1)(-t_{1},t_{1}) at 00 in the sense of Sequential Criterion for Differentiability at an Interior Point: tj≠0t_{j}\ne0, tj∈(−t1,t1)t_{j}\in(-t_{1},t_{1}), and tj→0t_{j}\to0. Fix k≤dk\le d. For x∈Xx\in X, the distance in XX from Ttj(x)T_{t_{j}}(x) to xx is ∣tj∣ ∣∇aψ(x)∣|t_{j}|\,|\nabla_{a}\psi(x)|, which tends to 00 by claim 3 of Arithmetic of Limits of Real Sequences; so Ttj(x)→xT_{t_{j}}(x)\to x in (X,d)(X,d), and since ∂kV\partial_{k}V is continuous at xx, Continuity Between Metric Spaces is Equivalent to Sequential Continuity §sequential gives ∂kV(Ttj(x))→∂kV(x)\partial_{k}V(T_{t_{j}}(x))\to\partial_{k}V(x), that is, φtj,k(x)→0\varphi_{t_{j},k}(x)\to0. The Borel functions φtj,k\varphi_{t_{j},k} are dominated by the integrable function Γk(eψ+1)U\Gamma_{k}(e_{\psi}+1)U, so the dominated convergence theorem (its claim 3, with limit function 00 and the measure space (X,B(X),μ)(X,\mathcal{B}(X),\mu)) gives ∫Xφtj,k dμ→0\int_{X}\varphi_{t_{j},k}\,d\mu\to0. By claims 1 and 3 of Arithmetic of Limits of Real Sequences, the right side rjr_{j} of (3.4) at t=tjt=t_{j} tends to 00. Given a real ε>0\varepsilon>0, choose JJ with rj<εr_{j}<\varepsilon for j≥Jj\ge J (Limit of a Sequence of Real Numbers); then (3.4) gives ∣qj−G∣<ε|q_{j}-G|<\varepsilon for j≥Jj\ge J, where qj=(F(tj)−F(0))/tjq_{j}=(F(t_{j})-F(0))/t_{j}. So qj→Gq_{j}\to G for every admissible sequence, and claim 2 of Sequential Criterion for Differentiability at an Interior Point shows that FF is differentiable at 00 on (−t1,t1)(-t_{1},t_{1}) with derivative GG.

(3.f) Conclusion. By (3.2), on (−t1,t1)(-t_{1},t_{1}) the function t↦E(μt)t\mapsto\mathcal{E}(\mu_{t}) is the sum of βf1\beta f_{1}, FF and the constant βlog⁡ZV,β\beta\log Z_{V,\beta}. By claims 1 and 2 of Sum, Constant Multiple, and Product Rules for One-Dimensional Derivatives, (3.c) and (3.e), it is differentiable at 00 with derivative

β⟨Zμa,∇aψ⟩μ+⟨∇aV,∇aψ⟩μ=⟨βZμa+∇aV,∇aψ⟩μ=⟨Σ(μ),∇aψ⟩μ,\beta\langle Z^{a}_{\mu},\nabla_{a}\psi\rangle_{\mu}+\langle\nabla_{a}V,\nabla_{a}\psi\rangle_{\mu}=\langle\beta Z^{a}_{\mu}+\nabla_{a}V,\nabla_{a}\psi\rangle_{\mu}=\langle\Sigma(\mu),\nabla_{a}\psi\rangle_{\mu},

by bilinearity of the inner product of the real Hilbert space L2(μ;Xa)L^{2}(\mu;X^{a}) (Probability Measures on a Hilbert Space Transported in the Noise Norm: Standing Notation §fields). Together with μt∈D\mu_{t}\in\mathcal{D} for t∈(−t1,t1)t\in(-t_{1},t_{1}) from (3.d), this is the requirement of Noise Penalty Pairs on the Noise Wasserstein Space: the Penalty, Its Score, and Their Domains §variation, with t1t_{1} in the role of t0t_{0} there.

Step 4 (clause 3). By The Gibbs Entropy Pair on the Noise Wasserstein Space: Relative Entropy to the Gibbs Measure and the Gibbs Score Field §pair, DΣ⊆D⊆Pρa\mathcal{D}_{\Sigma}\subseteq\mathcal{D}\subseteq\mathcal{P}^{a}_{\rho}, E\mathcal{E} is a real-valued function on D\mathcal{D}, and Σ(μ)∈Tμa\Sigma(\mu)\in T^{a}_{\mu} for every μ∈DΣ\mu\in\mathcal{D}_{\Sigma} by The Gibbs Entropy Pair on the Noise Wasserstein Space: Relative Entropy to the Gibbs Measure and the Gibbs Score Field §score-domain; the setting A Diagonal Gaussian Reference Measure on the Noise Wasserstein Space, Rescaled Heads and Gaussian Tails: Standing Notation §background carries that of Probability Measures on a Hilbert Space Transported in the Noise Norm: Standing Notation, with reference measure ρ=γc\rho=\gamma_{c}. We verify the four conditions of Noise Penalty Pairs on the Noise Wasserstein Space: the Penalty, Its Score, and Their Domains §pair.

Nonempty score domain (Noise Penalty Pairs on the Noise Wasserstein Space: the Penalty, Its Score, and Their Domains §nonempty). γc∈P(X)\gamma_{c}\in\mathcal{P}(X) by A Diagonal Gaussian Reference Measure on the Noise Wasserstein Space, Rescaled Heads and Gaussian Tails: Standing Notation §gaussian. The constant function 11 on XX is Borel and nonnegative, and γc(A)=∫X1A⋅1 dγc\gamma_{c}(A)=\int_{X}\mathbf{1}_{A}\cdot1\,d\gamma_{c} for every Borel AA, so it is a density of γc\gamma_{c} with respect to γc\gamma_{c} in the sense of The Radon-Nikodym Theorem for a Finite Measure and a Sigma-Finite Measure, and Uniqueness of Densities; and ϕ(1)=1⋅log⁡1=0\phi(1)=1\cdot\log1=0 for the function ϕ(s)=slog⁡s\phi(s)=s\log s of Relative Entropy of Probability Measures §relative-entropy, since log⁡1=log⁡(exp⁡(0))=0\log1=\log(\exp(0))=0 by claim 1 of Basic Properties of the Exponential Function and The Natural Logarithm. Hence ϕ∘1=0\phi\circ1=0 is integrable with respect to γc\gamma_{c} and γc\gamma_{c} has finite relative entropy with respect to γc\gamma_{c}. By Basic Properties of an Admissible Cylindrical Potential: Continuity, Growth under Noise Translations, Integrability, the Tangent Inequality and Tangency of the Noise Gradient §gaussian, VV is integrable with respect to γc\gamma_{c}, so γc\gamma_{c} has finite relative entropy with respect to γβV\gamma^{V}_{\beta} by Entropy and Score Relative to the Gibbs Measure Split into Their Gaussian Parts and the Potential, with a Fisher Information Bound §entropy, that is, γc∈D\gamma_{c}\in\mathcal{D}. Step 2, applied to γc\gamma_{c} and ε=1\varepsilon=1, gives an element of DΣ\mathcal{D}_{\Sigma}.

Lower bound (Noise Penalty Pairs on the Noise Wasserstein Space: the Penalty, Its Score, and Their Domains §bound). The nonnegative constant of that condition (not the slope constant CC fixed above) may be taken to be 00: for μ∈D\mu\in\mathcal{D}, −0⋅(1+Wa(μ,ρ)2)=0≤E(μ)-0\cdot(1+W_{a}(\mu,\rho)^{2})=0\le\mathcal{E}(\mu) by Step 1.

First variation (Noise Penalty Pairs on the Noise Wasserstein Space: the Penalty, Its Score, and Their Domains §variation). This is Step 3.

Density (Noise Penalty Pairs on the Noise Wasserstein Space: the Penalty, Its Score, and Their Domains §dense). This is Step 2.

Hence (D,DΣ,E,Σ)(\mathcal{D},\mathcal{D}_{\Sigma},\mathcal{E},\Sigma) is a noise penalty pair on Pρa\mathcal{P}^{a}_{\rho}, which is clause 3; clauses 1 and 2 are Steps 1 and 2.

Citations

Loading…

Dependencies

Uses0

Loading…

Comments

Log in to comment.

Loading…