TheoremBase

The Gaussian entropy pair with temperature 1 is a (1/kappa)-displacement convex noise penalty pair for which gammacgamma_c has zero score, so the abstract log-Sobolev and HWI inequalities give claims 1 and 2. For Gross's form, F2+epsF^2+eps is cylindrical and (F2+eps)/ZF^2+eps)/Z is the density of a measure whose entropy is Ent(F2+eps)/ZEnt(F^2+eps)/Z and whose relative score dkd_k g/g (checked by Gaussian integration by parts) has Fisher information at most 4/Z times the Dirichlet integral; then eps -> 0 by dominated convergence. Poincare follows from Gross applied to 1+sF with a second-order Peano expansion of s log s at 1.

Proof

Each result cited is universally quantified over the data in its own statement.

By Gaussian Analysis Relative to a Diagonal Gaussian Reference Measure with Noise Weights: Standing Notation §background the setting A Diagonal Gaussian Reference Measure on the Noise Wasserstein Space, Rescaled Heads and Gaussian Tails: Standing Notation is in force, with reference measure ρ=γc\rho=\gamma_{c}; it is layered on Probability Measures on a Hilbert Space Transported in the Noise Norm: Standing Notation, as is First-Order Equations on the Noise Wasserstein Space Relative to a Noise Penalty Pair: Standing Notation, so the results stated in the latter setting apply with ρ=γc\rho=\gamma_{c}. Throughout, ϕ\phi is the function of The Function slog⁡ss\log s: Continuity, Young's Inequality and Lower Bounds, with the Elementary Bounds for the Exponential and the Logarithm and κ\kappa is as in the statement.

Step 1 (The Gaussian entropy pair). Let (D,DΣ,E,Σ)(\mathcal{D},\mathcal{D}_{\Sigma},\mathcal{E},\Sigma) be the Gaussian entropy pair with temperature β=1\beta=1; its hypothesis holds with the given κ\kappa. By The Gaussian Entropy Pair on the Noise Wasserstein Space: Relative Entropy and the Noise Score Field §penalty-domain, D\mathcal{D} is the set of the μ∈P(X)\mu\in\mathcal{P}(X) of finite relative entropy with respect to γc\gamma_{c}; by The Gaussian Entropy Pair on the Noise Wasserstein Space: Relative Entropy and the Noise Score Field §score-domain, DΣ\mathcal{D}_{\Sigma} is the set of the μ∈D\mu\in\mathcal{D} that have a relative score with respect to γc\gamma_{c} and finite Fisher information relative to γc\gamma_{c} with weights aa; and by The Gaussian Entropy Pair on the Noise Wasserstein Space: Relative Entropy and the Noise Score Field §pair, E(μ)=H(μ ∣ γc)\mathcal{E}(\mu)=H(\mu\,|\,\gamma_{c}) for μ∈D\mu\in\mathcal{D} and Σ(μ)=Zμa\Sigma(\mu)=Z^{a}_{\mu} for μ∈DΣ\mu\in\mathcal{D}_{\Sigma}, the noise score field of μ\mu. By The Noise Score Field of a Measure of Finite Weighted Fisher Information: Existence, Norm, Pairing with Noise Gradients, Head Approximation and Tangency §field, ∥Σ(μ)∥μ2=Ia(μ ∣ γc)\lVert\Sigma(\mu)\rVert_{\mu}^{2}=\mathcal{I}_{a}(\mu\,|\,\gamma_{c}) for μ∈DΣ\mu\in\mathcal{D}_{\Sigma}. The quadruple is a noise penalty pair on Pρa\mathcal{P}^{a}_{\rho} by The Gaussian Entropy Pair is a Noise Penalty Pair, with Nonnegative Penalty and Dense Score Domain §pair, and it is λ\lambda-displacement convex with λ=1/κ\lambda=1/\kappa, a positive real number, by The Gaussian Entropy Pair is Uniformly Displacement Convex, with Modulus the Temperature over the Variance-to-Noise Bound §convex with β=1\beta=1.

Step 2 (The reference measure is a point of zero score). By The Integral of an Indicator Function is the Measure of the Set, ∫X1A⋅1 dγc=γc(A)\int_{X}\mathbf{1}_{A}\cdot1\,d\gamma_{c}=\gamma_{c}(A) for A∈B(X)A\in\mathcal{B}(X), so the constant function 11 is a density of γc\gamma_{c} with respect to γc\gamma_{c} in the sense of The Radon-Nikodym Theorem for a Finite Measure and a Sigma-Finite Measure, and Uniqueness of Densities. Since exp⁡(0)=1\exp(0)=1 by claim 1 of Basic Properties of the Exponential Function, log⁡1=log⁡(exp⁡(0))=0\log1=\log(\exp(0))=0 by The Natural Logarithm, so ϕ∘1\phi\circ1 is the constant ϕ(1)=1⋅log⁡1=0\phi(1)=1\cdot\log1=0, which is integrable with integral 00: by The Integral of an Indicator Function is the Measure of the Set with A=XA=X, the constant 1=1X1=\mathbf{1}_{X} has ∫X1 dγc=γc(X)=1\int_{X}1\,d\gamma_{c}=\gamma_{c}(X)=1, so it is integrable, and its multiple by 00 is integrable with integral 00 by Linearity and Monotonicity of the Lebesgue Integral §integrable. By Relative Entropy of Probability Measures §relative-entropy, γc\gamma_{c} has finite relative entropy with respect to γc\gamma_{c} and H(γc ∣ γc)=0H(\gamma_{c}\,|\,\gamma_{c})=0; thus γc∈D\gamma_{c}\in\mathcal{D} and E(γc)=0\mathcal{E}(\gamma_{c})=0. By The Relative Score on a Hilbert Space: the Gaussian Measure Has Score Zero, and the Score as a Square-Integrable Field in the Weighted Sequence Space §gaussian, γc∈P2(X)\gamma_{c}\in\mathcal{P}_{2}(X) has a relative score with respect to γc\gamma_{c} and finite Fisher information relative to γc\gamma_{c} with weights aa, with Ia(γc ∣ γc)=0\mathcal{I}_{a}(\gamma_{c}\,|\,\gamma_{c})=0. Hence γc∈DΣ\gamma_{c}\in\mathcal{D}_{\Sigma} and, by Step 1, ∥Σ(γc)∥γc2=0\lVert\Sigma(\gamma_{c})\rVert_{\gamma_{c}}^{2}=0; as the norm of the real Hilbert space L2(γc;Xa)L^{2}(\gamma_{c};X^{a}) of Probability Measures on a Hilbert Space Transported in the Noise Norm: Standing Notation §fields vanishes only at the zero vector, Σ(γc)\Sigma(\gamma_{c}) is the zero vector of L2(γc;Xa)L^{2}(\gamma_{c};X^{a}).

Step 3 (Claim 1). Let μ\mu be as in claim 1. Then μ∈D\mu\in\mathcal{D}, μ∈P2(X)\mu\in\mathcal{P}_{2}(X) by Relative Entropy with Respect to a Diagonal Gaussian Measure on a Hilbert Space: the Moment Bound, the Cutoff Projections, and Bounded, Tight, Weakly Closed, Wasserstein-Closed and Weakly Sequentially Compact Sublevel Sets §moment, and μ\mu has a relative score and finite Fisher information, so μ∈DΣ\mu\in\mathcal{D}_{\Sigma}. By Steps 1 and 2, Talagrand, HWI and Log-Sobolev Inequalities for a Uniformly Displacement Convex Noise Penalty Pair with a Point of Zero Score §log-sobolev applies to the pair of Step 1 with λ=1/κ\lambda=1/\kappa and μ∗=γc\mu_{*}=\gamma_{c}, and gives

H(μ ∣ γc)−0≤κ2 ∥Σ(μ)∥μ2=κ2 Ia(μ ∣ γc).H(\mu\,|\,\gamma_{c})-0\le\frac{\kappa}{2}\,\lVert\Sigma(\mu)\rVert_{\mu}^{2}=\frac{\kappa}{2}\,\mathcal{I}_{a}(\mu\,|\,\gamma_{c}).

Step 4 (Claim 2). Let μ\mu be as in claim 1. Then μ∈D⊆Pρa\mu\in\mathcal{D}\subseteq\mathcal{P}^{a}_{\rho} by The Entropy Domain of a Diagonal Gaussian Reference Measure Has the Noise Map Property §inclusion, and γc=ρ∈Pρa\gamma_{c}=\rho\in\mathcal{P}^{a}_{\rho} by The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §reference. As in Step 3, μ∈DΣ\mu\in\mathcal{D}_{\Sigma}, and Talagrand, HWI and Log-Sobolev Inequalities for a Uniformly Displacement Convex Noise Penalty Pair with a Point of Zero Score §hwi with λ=1/κ\lambda=1/\kappa and μ∗=γc\mu_{*}=\gamma_{c} gives H(μ ∣ γc)−0≤∥Σ(μ)∥μWa(μ,γc)−12κWa(μ,γc)2H(\mu\,|\,\gamma_{c})-0\le\lVert\Sigma(\mu)\rVert_{\mu}W_{a}(\mu,\gamma_{c})-\frac{1}{2\kappa}W_{a}(\mu,\gamma_{c})^{2}. Since ∥Σ(μ)∥μ\lVert\Sigma(\mu)\rVert_{\mu} is nonnegative with square Ia(μ ∣ γc)\mathcal{I}_{a}(\mu\,|\,\gamma_{c}) by Step 1, it equals Ia(μ ∣ γc)1/2\mathcal{I}_{a}(\mu\,|\,\gamma_{c})^{1/2} by Existence and Uniqueness of the Nonnegative Square Root, which proves claim 2.

Step 5 (Quadratic expressions in a cylindrical function). Let F∈FCb1(X)F\in\mathcal{F}C^{1}_{b}(X) have a representation (n,ψ)(n,\psi), and let α,s,b0∈R\alpha,s,b_{0}\in\mathbb{R}. We show that α+sF+b0F2∈FCb1(X)\alpha+sF+b_{0} F^{2}\in\mathcal{F}C^{1}_{b}(X) and that for every k∈Nk\in\mathbb{N}

∂k(α+sF+b0F2)=(s+2b0F) ∂kFon X,and both sides vanish for k>n.\partial_{k}\bigl(\alpha+sF+b_{0} F^{2}\bigr)=(s+2b_{0} F)\,\partial_{k}F\quad\text{on }X,\qquad\text{and both sides vanish for }k>n .

By Bounded Continuously Differentiable Functions with Bounded Partial Derivatives on Euclidean Space §bounded, ψ\psi is of class C1C^{1} on Rn\mathbb{R}^{n} and there are real B≥0B\ge0 and Bi≥0B_{i}\ge0 with ∣ψ∣≤B|\psi|\le B and ∣∂iψ∣≤Bi|\partial_{i}\psi|\le B_{i} for i∈[n]i\in[n]; by clause 1 of C^k Maps on a Euclidean Open Set, read through its clause 3 as in Differential Calculus and Convexity on Euclidean Open Sets: Standing Notation §derivatives, ψ\psi and each ∂iψ\partial_{i}\psi are continuous at every point of Rn\mathbb{R}^{n} in the sense of Continuity at a Point for Maps Between Euclidean Spaces, which for real-valued maps reads: for every ε>0\varepsilon>0 there is δ>0\delta>0 with ∣f(y)−f(x)∣<ε|f(y)-f(x)|<\varepsilon whenever ∥y−x∥<δ\lVert y-x\rVert<\delta. Let χ=α+sψ+b0ψ2:Rn→R\chi=\alpha+s\psi+b_{0}\psi^{2}:\mathbb{R}^{n}\to\mathbb{R}. It is bounded by ∣α∣+∣s∣B+∣b0∣B2|\alpha|+|s|B+|b_{0}|B^{2}. It is continuous at every xx: since ψ(y)2−ψ(x)2=(ψ(y)+ψ(x))(ψ(y)−ψ(x))\psi(y)^{2}-\psi(x)^{2}=(\psi(y)+\psi(x))(\psi(y)-\psi(x)), we have ∣χ(y)−χ(x)∣≤(∣s∣+2∣b0∣B) ∣ψ(y)−ψ(x)∣|\chi(y)-\chi(x)|\le(|s|+2|b_{0}|B)\,|\psi(y)-\psi(x)|, and it suffices to take δ\delta for ψ\psi at xx with ε/(∣s∣+2∣b0∣B+1)\varepsilon/(|s|+2|b_{0}|B+1) in place of ε\varepsilon. Fix x∈Rnx\in\mathbb{R}^{n} and i∈[n]i\in[n], let eie_{i} be the ii-th standard basis vector of Rn\mathbb{R}^{n}, and let u:R→Ru:\mathbb{R}\to\mathbb{R}, u(t)=ψ(x+tei)u(t)=\psi(x+te_{i}). The real line is an interval of which 00 is an interior point by claim 1 of One-Dimensional Derivatives, Partial Derivatives, and Smoothness on the Real Line, and the condition of Partial Derivative on a Euclidean Open Set for ψ\psi at xx in the ii-th variable is, word for word, the condition of Derivative at an Interior Point for uu at 00, the extra clauses of the two definitions being vacuous here: the requirement x+hei∈Rnx+he_{i}\in\mathbb{R}^{n} of the former holds for every hh, the open set being all of Rn\mathbb{R}^{n}, and the restriction 0+h∈R0+h\in\mathbb{R} of the latter excludes no hh, the interval being all of R\mathbb{R}; so uu is differentiable at 00 with u′(0)=∂iψ(x)u'(0)=\partial_{i}\psi(x). By claims 1 to 3 of Sum, Constant Multiple, and Product Rules for One-Dimensional Derivatives, v=α+su+b0uuv=\alpha+su+b_{0} uu is differentiable at 00 with v′(0)=su′(0)+2b0u(0)u′(0)v'(0)=su'(0)+2b_{0} u(0)u'(0), and since v(t)=χ(x+tei)v(t)=\chi(x+te_{i}), reading the same identification backwards gives that ∂iχ(x)\partial_{i}\chi(x) exists and equals (s+2b0ψ(x)) ∂iψ(x)(s+2b_{0}\psi(x))\,\partial_{i}\psi(x). The function ∂iχ=(s+2b0ψ)∂iψ\partial_{i}\chi=(s+2b_{0}\psi)\partial_{i}\psi is bounded by (∣s∣+2∣b0∣B)Bi(|s|+2|b_{0}|B)B_{i}, and is continuous at every xx because

∣∂iχ(y)−∂iχ(x)∣≤(∣s∣+2∣b0∣B) ∣∂iψ(y)−∂iψ(x)∣+2∣b0∣Bi ∣ψ(y)−ψ(x)∣,|\partial_{i}\chi(y)-\partial_{i}\chi(x)|\le(|s|+2|b_{0}|B)\,|\partial_{i}\psi(y)-\partial_{i}\psi(x)|+2|b_{0}|B_{i}\,|\psi(y)-\psi(x)|,

so it suffices to take the smaller of the two δ\delta's for ∂iψ\partial_{i}\psi and ψ\psi at xx with ε/(2(∣s∣+2∣b0∣B)+1)\varepsilon/(2(|s|+2|b_{0}|B)+1) and ε/(4∣b0∣Bi+1)\varepsilon/(4|b_{0}|B_{i}+1) in place of ε\varepsilon. Hence χ\chi is of class C1C^{1} on Rn\mathbb{R}^{n} by clause 1 of C^k Maps on a Euclidean Open Set, and χ∈Cb1(Rn)\chi\in C^{1}_{b}(\mathbb{R}^{n}) by Bounded Continuously Differentiable Functions with Bounded Partial Derivatives on Euclidean Space §bounded. Since χ∘pn=α+sF+b0F2\chi\circ p_{n}=\alpha+sF+b_{0} F^{2}, this function lies in FCb1(X)\mathcal{F}C^{1}_{b}(X) by Bounded C^1 Cylindrical Functions on a Hilbert Space with an Orthonormal Basis §cylindrical, with representation (n,χ)(n,\chi), and Bounded C^1 Cylindrical Functions: Linear Structure, the Gradient and the Partial Derivatives, Bounds and Integrability, and Density in the Square-Integrable Functions §partial, applied to both representations (n,χ)(n,\chi) and (n,ψ)(n,\psi), gives for k≤nk\le n that ∂k(χ∘pn)=(∂kχ)∘pn=(s+2b0F) (∂kψ)∘pn=(s+2b0F) ∂kF\partial_{k}(\chi\circ p_{n})=(\partial_{k}\chi)\circ p_{n}=(s+2b_{0} F)\,(\partial_{k}\psi)\circ p_{n}=(s+2b_{0} F)\,\partial_{k}F, and for k>nk>n that both ∂k(χ∘pn)\partial_{k}(\chi\circ p_{n}) and ∂kF\partial_{k}F vanish identically. In particular, by the display of The Noise Gradient of a Cylindrical Function and the Noise Tangent Space at a Probability Measure on a Hilbert Space §gradient computed with these representations, ∣∇a(α+sF+b0F2)∣a2=∑k=1nak(s+2b0F)2(∂kF)2|\nabla_{a}(\alpha+sF+b_{0} F^{2})|_{a}^{2}=\sum_{k=1}^{n}a_{k}(s+2b_{0} F)^{2}(\partial_{k}F)^{2}.

Step 6 (Entropies of bounded functions). Let G:X→RG:X\to\mathbb{R} be Borel with values in an interval [0,b][0,b]. The function ϕ\phi is continuous on [0,∞)[0,\infty) by The Function slog⁡ss\log s: Continuity, Young's Inequality and Lower Bounds, with the Elementary Bounds for the Exponential and the Logarithm §continuous, hence on [0,b][0,b], so Extreme Value Theorem on a Closed Real Interval, applied to the restriction of ϕ\phi to [0,b][0,b], gives rmin⁡,rmax⁡∈[0,b]r_{\min},r_{\max}\in[0,b] with ϕ(rmin⁡)≤ϕ(r)≤ϕ(rmax⁡)\phi(r_{\min})\le\phi(r)\le\phi(r_{\max}) for r∈[0,b]r\in[0,b]; with Mb=max⁡(∣ϕ(rmin⁡)∣,∣ϕ(rmax⁡)∣)M_{b}=\max(|\phi(r_{\min})|,|\phi(r_{\max})|) we get ∣ϕ(r)∣≤Mb|\phi(r)|\le M_{b} for r∈[0,b]r\in[0,b]; ϕ∘G\phi\circ G is Borel by The Function slog⁡ss\log s: Continuity, Young's Inequality and Lower Bounds, with the Elementary Bounds for the Exponential and the Logarithm §continuous. Thus GG and ϕ∘G\phi\circ G are bounded and Borel, hence integrable with respect to every Borel probability measure on XX by Relative Entropy on a Measurable Space: the Gibbs Inequality, the Variational Criterion and Formula, the Entropy Inequality, Small Sets and Data Processing §functional, and Ent⁡γc(G)\operatorname{Ent}_{\gamma_{c}}(G) is defined by The Entropy of a Nonnegative Function with Respect to a Probability Measure §entropy. For a real constant rr, ∫Xr dγc=r\int_{X}r\,d\gamma_{c}=r by The Integral of an Indicator Function is the Measure of the Set and Linearity and Monotonicity of the Lebesgue Integral §integrable, γc\gamma_{c} being a probability measure. Now let F∈FCb1(X)F\in\mathcal{F}C^{1}_{b}(X) with representation (n,ψ)(n,\psi), and fix a real B≥1B\ge1 with ∣ψ∣≤B|\psi|\le B, so that ∣F∣≤B|F|\le B on XX. Since FF is Borel by Bounded C^1 Cylindrical Functions: Linear Structure, the Gradient and the Partial Derivatives, Bounds and Integrability, and Density in the Square-Integrable Functions §bounded-borel, F2F^{2} is Borel by claim 3 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions, nonnegative and bounded by B2B^{2}; so F2F^{2} and ϕ∘F2\phi\circ F^{2} are integrable with respect to γc\gamma_{c}, which is the first part of claim 3. We write D=∫X∣∇aF∣a2 dγcD=\int_{X}|\nabla_{a}F|_{a}^{2}\,d\gamma_{c}, a nonnegative real number by The Noise Gradient of a Cylindrical Function and the Noise Tangent Space at a Probability Measure on a Hilbert Space §gradient.

Step 7 (The inequality for F2+εF^{2}+\varepsilon). Let FF, (n,ψ)(n,\psi), BB and DD be as in Step 6, and let ε\varepsilon be real with 0<ε≤10<\varepsilon\le1. Let g=ε+F2g=\varepsilon+F^{2}. By Step 5 (with α=ε\alpha=\varepsilon, s=0s=0, b0=1b_{0}=1), g∈FCb1(X)g\in\mathcal{F}C^{1}_{b}(X) and ∂kg=2F ∂kF\partial_{k}g=2F\,\partial_{k}F for every kk, with ∂kg=0\partial_{k}g=0 for k>nk>n; and ε≤g≤B2+1\varepsilon\le g\le B^{2}+1. Let Nε=∫Xg dγc=ε+∫XF2 dγcN_{\varepsilon}=\int_{X}g\,d\gamma_{c}=\varepsilon+\int_{X}F^{2}\,d\gamma_{c}, so ε≤Nε≤B2+1\varepsilon\le N_{\varepsilon}\le B^{2}+1. The function g/Nε:X→[0,∞)g/N_{\varepsilon}:X\to[0,\infty) is Borel, so by claim 3 of Image Measures, Measures with Densities, and Change of Variables the measure μ\mu with density g/Nεg/N_{\varepsilon} with respect to γc\gamma_{c}, μ(A)=∫X1A (g/Nε) dγc\mu(A)=\int_{X}\mathbf{1}_{A}\,(g/N_{\varepsilon})\,d\gamma_{c}, is a measure on B(X)\mathcal{B}(X), with μ(X)=Nε−1∫Xg dγc=1\mu(X)=N_{\varepsilon}^{-1}\int_{X}g\,d\gamma_{c}=1; so μ∈P(X)\mu\in\mathcal{P}(X) and g/Nεg/N_{\varepsilon} is a density of μ\mu with respect to γc\gamma_{c} in the sense of The Radon-Nikodym Theorem for a Finite Measure and a Sigma-Finite Measure, and Uniqueness of Densities. We use repeatedly that, by claim 3 of Image Measures, Measures with Densities, and Change of Variables, a Borel f:X→Rf:X\to\mathbb{R} with f g/Nεf\,g/N_{\varepsilon} integrable with respect to γc\gamma_{c} is integrable with respect to μ\mu, with ∫Xf dμ=Nε−1∫Xf g dγc\int_{X}f\,d\mu=N_{\varepsilon}^{-1}\int_{X}f\,g\,d\gamma_{c}.

Entropy. The function g/Nεg/N_{\varepsilon} is Borel with values in [0,(B2+1)/ε][0,(B^{2}+1)/\varepsilon], so ϕ∘(g/Nε)\phi\circ(g/N_{\varepsilon}) is integrable with respect to γc\gamma_{c} by Step 6, and μ\mu has finite relative entropy with respect to γc\gamma_{c}, with H(μ ∣ γc)=∫Xϕ∘(g/Nε) dγcH(\mu\,|\,\gamma_{c})=\int_{X}\phi\circ(g/N_{\varepsilon})\,d\gamma_{c}, by Relative Entropy of Probability Measures §relative-entropy. For r>0r>0, log⁡r=log⁡(r/Nε)+log⁡Nε\log r=\log(r/N_{\varepsilon})+\log N_{\varepsilon} by The Natural Logarithm, so ϕ(r/Nε)=(r/Nε)log⁡(r/Nε)=Nε−1(ϕ(r)−rlog⁡Nε)\phi(r/N_{\varepsilon})=(r/N_{\varepsilon})\log(r/N_{\varepsilon})=N_{\varepsilon}^{-1}\bigl(\phi(r)-r\log N_{\varepsilon}\bigr). Applying this with r=g(x)>0r=g(x)>0 and integrating, using Linearity and Monotonicity of the Lebesgue Integral §integrable, Step 6 for gg, and The Entropy of a Nonnegative Function with Respect to a Probability Measure §entropy,

H(μ ∣ γc)=Nε−1(∫Xϕ∘g dγc−Nεlog⁡Nε)=Nε−1(∫Xϕ∘g dγc−ϕ(Nε))=Nε−1Ent⁡γc(g).H(\mu\,|\,\gamma_{c})=N_{\varepsilon}^{-1}\Bigl(\int_{X}\phi\circ g\,d\gamma_{c}-N_{\varepsilon}\log N_{\varepsilon}\Bigr)=N_{\varepsilon}^{-1}\Bigl(\int_{X}\phi\circ g\,d\gamma_{c}-\phi(N_{\varepsilon})\Bigr)=N_{\varepsilon}^{-1}\operatorname{Ent}_{\gamma_{c}}(g).

In particular μ∈P2(X)\mu\in\mathcal{P}_{2}(X) by Relative Entropy with Respect to a Diagonal Gaussian Measure on a Hilbert Space: the Moment Bound, the Cutoff Projections, and Bounded, Tight, Weakly Closed, Wasserstein-Closed and Weakly Sequentially Compact Sublevel Sets §moment.

Relative score. For k∈Nk\in\mathbb{N} let ζk=∂kg⋅(1/g)\zeta_{k}=\partial_{k}g\cdot(1/g). The function 1/g1/g is Borel by the criterion of Measure Spaces and the Lebesgue Integral: Standing Notation §measurable: {1/g>r}\{1/g>r\} is XX for r≤0r\le0 and {−g>−1/r}\{-g>-1/r\} for r>0r>0, Borel by claim 2 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions. As ∂kg\partial_{k}g is bounded and Borel by Bounded C^1 Cylindrical Functions: Linear Structure, the Gradient and the Partial Derivatives, Bounds and Integrability, and Density in the Square-Integrable Functions §bounded-borel and 0<1/g≤1/ε0<1/g\le1/\varepsilon, ζk\zeta_{k} is bounded and Borel (claim 3 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions); so ζk2\zeta_{k}^{2} is integrable with respect to μ\mu by Relative Entropy on a Measurable Space: the Gibbs Inequality, the Variational Criterion and Formula, the Entropy Inequality, Small Sets and Data Processing §functional, and ζk\zeta_{k} defines an element of L2(μ)L^{2}(\mu), again written ζk\zeta_{k}; for k>nk>n, ζk=0\zeta_{k}=0. Let φ∈FCb1(X)\varphi\in\mathcal{F}C^{1}_{b}(X) and put q(x)=xkφ(x)/ck−∂kφ(x)q(x)=x_{k}\varphi(x)/c_{k}-\partial_{k}\varphi(x). By The Lebesgue Space of Square-Integrable Functions is a Real Hilbert Space §inner-product, then the density formula (the function ζkφ g/Nε=Nε−1∂kg φ\zeta_{k}\varphi\,g/N_{\varepsilon}=N_{\varepsilon}^{-1}\partial_{k}g\,\varphi being bounded and Borel), then Gaussian Integration by Parts for Products of Cylindrical Functions, and Closability of the Noise Gradient in the Gaussian Lebesgue Space §ibp applied with gg in place of FF and with φ\varphi, then Linearity and Monotonicity of the Lebesgue Integral §integrable and the density formula once more (the function q g/Nεq\,g/N_{\varepsilon} being integrable with respect to γc\gamma_{c} as Nε−1N_{\varepsilon}^{-1} times the integrable function g qg\,q of Gaussian Integration by Parts for Products of Cylindrical Functions, and Closability of the Noise Gradient in the Gaussian Lebesgue Space §ibp),

⟨ζk,φ⟩L2(μ)=∫Xζkφ dμ=1Nε∫X∂kg φ dγc=1Nε∫Xg q dγc=∫X(xkck φ(x)−∂kφ(x))μ(dx).\langle\zeta_{k},\varphi\rangle_{L^{2}(\mu)}=\int_{X}\zeta_{k}\varphi\,d\mu=\frac{1}{N_{\varepsilon}}\int_{X}\partial_{k}g\,\varphi\,d\gamma_{c}=\frac{1}{N_{\varepsilon}}\int_{X}g\,q\,d\gamma_{c}=\int_{X}\Bigl(\frac{x_{k}}{c_{k}}\,\varphi(x)-\partial_{k}\varphi(x)\Bigr)\mu(dx).

By The Relative Score with Respect to a Diagonal Gaussian Measure on a Hilbert Space §score, μ\mu has a relative score with respect to γc\gamma_{c}, with components ζk\zeta_{k}.

Fisher information. For k>nk>n the term ak∥ζk∥L2(μ)2a_{k}\lVert\zeta_{k}\rVert_{L^{2}(\mu)}^{2} is 00, so the partial sums of ∑k=1∞ak∥ζk∥L2(μ)2\sum_{k=1}^{\infty}a_{k}\lVert\zeta_{k}\rVert_{L^{2}(\mu)}^{2} are constant from the nn-th on; the series converges with sum ∑k=1nak∥ζk∥L2(μ)2\sum_{k=1}^{n}a_{k}\lVert\zeta_{k}\rVert_{L^{2}(\mu)}^{2}, and by Weight Sequences and the Weighted Fisher Information Relative to a Diagonal Gaussian Measure on a Hilbert Space §information μ\mu has finite Fisher information relative to γc\gamma_{c} with weights aa and Ia(μ ∣ γc)\mathcal{I}_{a}(\mu\,|\,\gamma_{c}) is that finite sum. For k≤nk\le n, by The Lebesgue Space of Square-Integrable Functions is a Real Hilbert Space §inner-product and the density formula, ∥ζk∥L2(μ)2=∫Xζk2 dμ=Nε−1∫X(∂kg)2/g dγc\lVert\zeta_{k}\rVert_{L^{2}(\mu)}^{2}=\int_{X}\zeta_{k}^{2}\,d\mu=N_{\varepsilon}^{-1}\int_{X}(\partial_{k}g)^{2}/g\,d\gamma_{c}, and pointwise

(∂kg)2g=4F2(∂kF)2ε+F2≤4(∂kF)2,\frac{(\partial_{k}g)^{2}}{g}=\frac{4F^{2}(\partial_{k}F)^{2}}{\varepsilon+F^{2}}\le4(\partial_{k}F)^{2},

since 0≤F2≤ε+F20\le F^{2}\le\varepsilon+F^{2}. By the monotonicity and linearity of Linearity and Monotonicity of the Lebesgue Integral §integrable and the display of The Noise Gradient of a Cylindrical Function and the Noise Tangent Space at a Probability Measure on a Hilbert Space §gradient for the representation (n,ψ)(n,\psi),

Ia(μ ∣ γc)≤4Nε∑k=1nak∫X(∂kF)2 dγc=4Nε∫X∣∇aF∣a2 dγc=4DNε.\mathcal{I}_{a}(\mu\,|\,\gamma_{c})\le\frac{4}{N_{\varepsilon}}\sum_{k=1}^{n}a_{k}\int_{X}(\partial_{k}F)^{2}\,d\gamma_{c}=\frac{4}{N_{\varepsilon}}\int_{X}|\nabla_{a}F|_{a}^{2}\,d\gamma_{c}=\frac{4D}{N_{\varepsilon}}.

Conclusion. By claim 1 (Step 3) applied to μ\mu, Nε−1Ent⁡γc(g)=H(μ ∣ γc)≤κ2⋅4DNεN_{\varepsilon}^{-1}\operatorname{Ent}_{\gamma_{c}}(g)=H(\mu\,|\,\gamma_{c})\le\frac{\kappa}{2}\cdot\frac{4D}{N_{\varepsilon}}, and multiplying by Nε>0N_{\varepsilon}>0,

Ent⁡γc(F2+ε)≤2κD(0<ε≤1).\operatorname{Ent}_{\gamma_{c}}(F^{2}+\varepsilon)\le2\kappa D\qquad(0<\varepsilon\le1).

Step 8 (Claim 3: letting ε→0\varepsilon\to0). Let FF, BB, DD be as in Step 6, and for j∈Nj\in\mathbb{N} let gj=F2+1/jg_{j}=F^{2}+1/j, with values in [0,B2+1][0,B^{2}+1]. The sequence (1/j)j(1/j)_{j} converges to 00 in the sense of Limit of a Sequence of Real Numbers: given a real θ>0\theta>0, claim 3 of The Archimedean Property of the Real Numbers gives j0∈Nj_{0}\in\mathbb{N} with 0<1/j0<θ0<1/j_{0}<\theta, and then 0<1/j≤1/j0<θ0<1/j\le1/j_{0}<\theta for j≥j0j\ge j_{0}. As a constant sequence converges to its value, Arithmetic of Limits of Real Sequences §sums shows that for each x∈Xx\in X the sequence (F(x)2+1/j)j(F(x)^{2}+1/j)_{j} in [0,∞)[0,\infty) converges to F(x)2F(x)^{2}, so (ϕ(gj(x)))j(\phi(g_{j}(x)))_{j} converges to ϕ(F(x)2)\phi(F(x)^{2}) by Continuity Between Metric Spaces is Equivalent to Sequential Continuity §sequential and the continuity of ϕ\phi (The Function slog⁡ss\log s: Continuity, Young's Inequality and Lower Bounds, with the Elementary Bounds for the Exponential and the Logarithm §continuous); moreover ∣ϕ∘gj∣≤MB2+1|\phi\circ g_{j}|\le M_{B^{2}+1} (Step 6), a constant, which is integrable. By Dominated Convergence Theorem, ∫Xϕ∘gj dγc→∫Xϕ∘F2 dγc\int_{X}\phi\circ g_{j}\,d\gamma_{c}\to\int_{X}\phi\circ F^{2}\,d\gamma_{c}. Also ∫Xgj dγc=∫XF2 dγc+1/j→∫XF2 dγc\int_{X}g_{j}\,d\gamma_{c}=\int_{X}F^{2}\,d\gamma_{c}+1/j\to\int_{X}F^{2}\,d\gamma_{c} by the same sum clause, so ϕ(∫Xgj dγc)→ϕ(∫XF2 dγc)\phi\bigl(\int_{X}g_{j}\,d\gamma_{c}\bigr)\to\phi\bigl(\int_{X}F^{2}\,d\gamma_{c}\bigr) by the same continuity. By Arithmetic of Limits of Real Sequences §scalar, Ent⁡γc(gj)→Ent⁡γc(F2)\operatorname{Ent}_{\gamma_{c}}(g_{j})\to\operatorname{Ent}_{\gamma_{c}}(F^{2}); since Ent⁡γc(gj)≤2κD\operatorname{Ent}_{\gamma_{c}}(g_{j})\le2\kappa D for every jj by Step 7, claim 1 of Order Properties of Limits of Real Sequences gives Ent⁡γc(F2)≤2κD\operatorname{Ent}_{\gamma_{c}}(F^{2})\le2\kappa D, which completes claim 3.

Step 9 (A second-order expansion of ϕ\phi at 11). Let U=(0,∞)U=(0,\infty), open in R1\mathbb{R}^{1} by Open Subset of Euclidean Space (for t∈Ut\in U, every yy with (y−t)2<t2(y-t)^{2}<t^{2} is positive), and let f:U→Rf:U\to\mathbb{R}, f(t)=tlog⁡t=ϕ(t)f(t)=t\log t=\phi(t). Every t∈Ut\in U is an interior point of the interval UU, as t/2<t<2tt/2<t<2t. We first show that log⁡\log, defined in The Natural Logarithm, is differentiable at every t∈Ut\in U in the sense of Derivative at an Interior Point, with derivative 1/t1/t. Fix t∈Ut\in U and a real hh with 0<∣h∣<t/20<|h|<t/2, and put τ=1+h/t\tau=1+h/t, so that τ>1/2\tau>1/2 and t+h=tτt+h=t\tau. By the identity log⁡(tτ)=log⁡t+log⁡τ\log(t\tau)=\log t+\log\tau recorded in The Natural Logarithm, log⁡(t+h)−log⁡t=log⁡τ\log(t+h)-\log t=\log\tau, and by The Function slog⁡ss\log s: Continuity, Young's Inequality and Lower Bounds, with the Elementary Bounds for the Exponential and the Logarithm §log, 1−τ−1≤log⁡τ≤τ−11-\tau^{-1}\le\log\tau\le\tau-1, that is, h/(t+h)≤log⁡(t+h)−log⁡t≤h/th/(t+h)\le\log(t+h)-\log t\le h/t. Dividing by hh, the difference quotient Qh=(log⁡(t+h)−log⁡t)/hQ_{h}=(\log(t+h)-\log t)/h lies between 1/(t+h)1/(t+h) and 1/t1/t (in this order if h>0h>0, in the reverse order if h<0h<0), so, using t+h>t/2t+h>t/2,

∣Qh−1t∣≤∣1t+h−1t∣=∣h∣t (t+h)≤2∣h∣t2.\Bigl|Q_{h}-\frac{1}{t}\Bigr|\le\Bigl|\frac{1}{t+h}-\frac{1}{t}\Bigr|=\frac{|h|}{t\,(t+h)}\le\frac{2|h|}{t^{2}}.

Given a real ε>0\varepsilon>0, put δt=min⁡(t/2,εt2/2)\delta_{t}=\min(t/2,\varepsilon t^{2}/2); every hh with 0<∣h∣<δt0<|h|<\delta_{t} then satisfies t+h∈Ut+h\in U and ∣Qh−1/t∣<2δt/t2≤ε|Q_{h}-1/t|<2\delta_{t}/t^{2}\le\varepsilon, which is the condition of Derivative at an Interior Point with L=1/tL=1/t. The identity map of UU has derivative 11 (its difference quotients equal 11); so by claims 1 to 3 of Sum, Constant Multiple, and Product Rules for One-Dimensional Derivatives, ff is differentiable at every t∈Ut\in U with f′(t)=log⁡t+1f'(t)=\log t+1, and f′f' is differentiable at every tt with derivative ω(t)=1/t\omega(t)=1/t, which is differentiable at every t∈Ut\in U by claim 2 of Reciprocal Rule for One-Dimensional Derivatives, 00 not lying in UU. Two elementary observations transfer this to the notions used by C^k Maps on a Euclidean Open Set. First, if κ^:U→R\hat\kappa:U\to\mathbb{R} is differentiable at tt with derivative LL in the sense of Derivative at an Interior Point, then the partial derivative of κ^\hat\kappa with respect to the first variable exists at tt with value LL in the sense of Partial Derivative on a Euclidean Open Set: the two conditions concern the same difference quotient, and replacing δ\delta by min⁡(δ,t)\min(\delta,t) ensures t+h∈Ut+h\in U whenever 0<∣h∣<δ0<|h|<\delta. Second, such a κ^\hat\kappa is continuous at tt in the sense of Continuity at a Point for Maps Between Euclidean Spaces: taking δ1≤t\delta_{1}\le t for the value 11 in the derivative condition, ∣κ^(t+h)−κ^(t)∣≤∣h∣(∣L∣+1)|\hat\kappa(t+h)-\hat\kappa(t)|\le|h|(|L|+1) for 0<∣h∣<δ10<|h|<\delta_{1}, which is less than a given ε>0\varepsilon>0 when also ∣h∣<ε/(∣L∣+1)|h|<\varepsilon/(|L|+1). Hence ff, ∂1f=f′\partial_{1}f=f' and ∂1∂1f=ω\partial_{1}\partial_{1}f=\omega exist and are continuous at every point of UU; by clause 1 of C^k Maps on a Euclidean Open Set, ff and ∂1f\partial_{1}f are of class C1C^{1} on UU, and by its clause 2 (with k=1k=1), ff is of class C2C^{2} on UU. Moreover f(1)=0f(1)=0, ∂1f(1)=log⁡1+1=1\partial_{1}f(1)=\log1+1=1 and ∂1∂1f(1)=1\partial_{1}\partial_{1}f(1)=1, using log⁡1=0\log1=0 from Step 2. By Second-Order Taylor Expansion with Peano Remainder with n=1n=1 and x=1x=1, where the Euclidean norm of h∈R1h\in\mathbb{R}^{1} is ∣h∣|h|: for every real η>0\eta>0 there is a real δ>0\delta>0 such that every real hh with ∣h∣<δ|h|<\delta satisfies 1+h>01+h>0 and

∣ϕ(1+h)−h−12h2∣≤η h2.\Bigl|\phi(1+h)-h-\tfrac{1}{2}h^{2}\Bigr|\le\eta\,h^{2}.

For h>−1h>-1 write R(h)=ϕ(1+h)−h−12h2R(h)=\phi(1+h)-h-\frac{1}{2}h^{2}.

Step 10 (Claim 4). Let FF, (n,ψ)(n,\psi), B≥1B\ge1 and DD be as in Step 6, and put m=∫XF dγcm=\int_{X}F\,d\gamma_{c}, v=∫XF2 dγcv=\int_{X}F^{2}\,d\gamma_{c}, m3=∫XF3 dγcm_{3}=\int_{X}F^{3}\,d\gamma_{c} and m4=∫XF4 dγcm_{4}=\int_{X}F^{4}\,d\gamma_{c}, integrals of bounded Borel functions (claim 3 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions); by Linearity and Monotonicity of the Lebesgue Integral §integrable, ∣m∣≤B|m|\le B, 0≤v≤B20\le v\le B^{2}, ∣m3∣≤B3|m_{3}|\le B^{3} and m4≥0m_{4}\ge0. Let η>0\eta>0, let δ\delta be as in Step 9, and let s=min⁡(η,1/B,δ/(6B))s=\min(\eta,1/B,\delta/(6B)), so that 0<s≤η0<s\le\eta, sB≤1sB\le1 and s≤1s\le1. Let Gs=1+sFG_{s}=1+sF. By Step 5 (with α=1\alpha=1, ss, b0=0b_{0}=0), Gs∈FCb1(X)G_{s}\in\mathcal{F}C^{1}_{b}(X) with ∣∇aGs∣a2=s2∑k=1nak(∂kF)2=s2∣∇aF∣a2|\nabla_{a}G_{s}|_{a}^{2}=s^{2}\sum_{k=1}^{n}a_{k}(\partial_{k}F)^{2}=s^{2}|\nabla_{a}F|_{a}^{2}, so claim 3 applied to GsG_{s} gives, with Linearity and Monotonicity of the Lebesgue Integral §integrable,

Ent⁡γc(Gs2)≤2κs2D.\operatorname{Ent}_{\gamma_{c}}(G_{s}^{2})\le2\kappa s^{2}D .

Now Gs2=1+uG_{s}^{2}=1+u with u=2sF+s2F2u=2sF+s^{2}F^{2}, and ∣u∣≤2sB+s2B2≤3sB<δ|u|\le2sB+s^{2}B^{2}\le3sB<\delta on XX. Let w=∫Xu dγc=2sm+s2vw=\int_{X}u\,d\gamma_{c}=2sm+s^{2}v; then ∣w∣≤∫X∣u∣ dγc≤3sB<δ|w|\le\int_{X}|u|\,d\gamma_{c}\le3sB<\delta and ∫XGs2 dγc=1+w\int_{X}G_{s}^{2}\,d\gamma_{c}=1+w. By Step 9, ϕ∘Gs2=u+12u2+R∘u\phi\circ G_{s}^{2}=u+\frac{1}{2}u^{2}+R\circ u on XX, where R∘uR\circ u is integrable as a difference of integrable functions (Step 6) and satisfies R∘u≥−ηu2R\circ u\ge-\eta u^{2}; and ϕ(1+w)=w+12w2+R(w)\phi(1+w)=w+\frac{1}{2}w^{2}+R(w) with R(w)≤ηw2R(w)\le\eta w^{2}. By The Entropy of a Nonnegative Function with Respect to a Probability Measure §entropy and Linearity and Monotonicity of the Lebesgue Integral §integrable,

Ent⁡γc(Gs2)=12(∫Xu2 dγc−w2)+∫XR∘u dγc−R(w)≥12(∫Xu2 dγc−w2)−η(∫Xu2 dγc+w2).\operatorname{Ent}_{\gamma_{c}}(G_{s}^{2})=\frac{1}{2}\Bigl(\int_{X}u^{2}\,d\gamma_{c}-w^{2}\Bigr)+\int_{X}R\circ u\,d\gamma_{c}-R(w)\ge\frac{1}{2}\Bigl(\int_{X}u^{2}\,d\gamma_{c}-w^{2}\Bigr)-\eta\Bigl(\int_{X}u^{2}\,d\gamma_{c}+w^{2}\Bigr).

Here ∫Xu2 dγc≤9s2B2\int_{X}u^{2}\,d\gamma_{c}\le9s^{2}B^{2} and w2≤9s2B2w^{2}\le9s^{2}B^{2}, while expanding the squares,

12(∫Xu2 dγc−w2)=2s2(v−m2)+2s3(m3−mv)+s42(m4−v2)≥2s2(v−m2)−4s3B3−s42B4.\frac{1}{2}\Bigl(\int_{X}u^{2}\,d\gamma_{c}-w^{2}\Bigr)=2s^{2}(v-m^{2})+2s^{3}(m_{3}-mv)+\frac{s^{4}}{2}(m_{4}-v^{2})\ge2s^{2}(v-m^{2})-4s^{3}B^{3}-\frac{s^{4}}{2}B^{4}.

Combining the last three displays and dividing by 2s2>02s^{2}>0,

v−m2≤κD+2sB3+s24B4+9ηB2≤κD+η(2B3+B44+9B2),v-m^{2}\le\kappa D+2sB^{3}+\frac{s^{2}}{4}B^{4}+9\eta B^{2}\le\kappa D+\eta\Bigl(2B^{3}+\frac{B^{4}}{4}+9B^{2}\Bigr),

using s2≤s≤ηs^{2}\le s\le\eta. Let C0=2B3+B44+9B2C_{0}=2B^{3}+\frac{B^{4}}{4}+9B^{2} be the bracket, which does not depend on η\eta and is positive as B≥1B\ge1. Given a positive real ε\varepsilon, the argument above with η=ε/C0\eta=\varepsilon/C_{0} gives v−m2≤κD+εv-m^{2}\le\kappa D+\varepsilon; as ε\varepsilon was arbitrary, Comparison of Real Numbers with Arbitrary Positive Slack §slack-above gives v−m2≤κDv-m^{2}\le\kappa D, which is claim 4.

Citations

Loading…

Dependencies

Uses0

Loading…

Comments

Log in to comment.

Loading…