Each result cited is universally quantified over the data in its own statement.
By Gaussian Analysis Relative to a Diagonal Gaussian Reference Measure with Noise Weights: Standing Notation §background the setting A Diagonal Gaussian Reference Measure on the Noise Wasserstein Space, Rescaled Heads and Gaussian Tails: Standing Notation is in force, with reference measure ρ = γ c \rho=\gamma_{c} ρ = γ c ; it is layered on Probability Measures on a Hilbert Space Transported in the Noise Norm: Standing Notation , as is First-Order Equations on the Noise Wasserstein Space Relative to a Noise Penalty Pair: Standing Notation , so the results stated in the latter setting apply with ρ = γ c \rho=\gamma_{c} ρ = γ c . Throughout, ϕ \phi ϕ is the function of The Function s log s s\log s s log s : Continuity, Young's Inequality and Lower Bounds, with the Elementary Bounds for the Exponential and the Logarithm and κ \kappa κ is as in the statement.
Step 1 (The Gaussian entropy pair). Let ( D , D Σ , E , Σ ) (\mathcal{D},\mathcal{D}_{\Sigma},\mathcal{E},\Sigma) ( D , D Σ , E , Σ ) be the Gaussian entropy pair with temperature β = 1 \beta=1 β = 1 ; its hypothesis holds with the given κ \kappa κ . By The Gaussian Entropy Pair on the Noise Wasserstein Space: Relative Entropy and the Noise Score Field §penalty-domain , D \mathcal{D} D is the set of the μ ∈ P ( X ) \mu\in\mathcal{P}(X) μ ∈ P ( X ) of finite relative entropy with respect to γ c \gamma_{c} γ c ; by The Gaussian Entropy Pair on the Noise Wasserstein Space: Relative Entropy and the Noise Score Field §score-domain , D Σ \mathcal{D}_{\Sigma} D Σ is the set of the μ ∈ D \mu\in\mathcal{D} μ ∈ D that have a relative score with respect to γ c \gamma_{c} γ c and finite Fisher information relative to γ c \gamma_{c} γ c with weights a a a ; and by The Gaussian Entropy Pair on the Noise Wasserstein Space: Relative Entropy and the Noise Score Field §pair , E ( μ ) = H ( μ ∣ γ c ) \mathcal{E}(\mu)=H(\mu\,|\,\gamma_{c}) E ( μ ) = H ( μ ∣ γ c ) for μ ∈ D \mu\in\mathcal{D} μ ∈ D and Σ ( μ ) = Z μ a \Sigma(\mu)=Z^{a}_{\mu} Σ ( μ ) = Z μ a for μ ∈ D Σ \mu\in\mathcal{D}_{\Sigma} μ ∈ D Σ , the noise score field of μ \mu μ . By The Noise Score Field of a Measure of Finite Weighted Fisher Information: Existence, Norm, Pairing with Noise Gradients, Head Approximation and Tangency §field , ∥ Σ ( μ ) ∥ μ 2 = I a ( μ ∣ γ c ) \lVert\Sigma(\mu)\rVert_{\mu}^{2}=\mathcal{I}_{a}(\mu\,|\,\gamma_{c}) ∥ Σ ( μ ) ∥ μ 2 = I a ( μ ∣ γ c ) for μ ∈ D Σ \mu\in\mathcal{D}_{\Sigma} μ ∈ D Σ . The quadruple is a noise penalty pair on P ρ a \mathcal{P}^{a}_{\rho} P ρ a by The Gaussian Entropy Pair is a Noise Penalty Pair, with Nonnegative Penalty and Dense Score Domain §pair , and it is λ \lambda λ -displacement convex with λ = 1 / κ \lambda=1/\kappa λ = 1/ κ , a positive real number, by The Gaussian Entropy Pair is Uniformly Displacement Convex, with Modulus the Temperature over the Variance-to-Noise Bound §convex with β = 1 \beta=1 β = 1 .
Step 2 (The reference measure is a point of zero score). By The Integral of an Indicator Function is the Measure of the Set , ∫ X 1 A ⋅ 1 d γ c = γ c ( A ) \int_{X}\mathbf{1}_{A}\cdot1\,d\gamma_{c}=\gamma_{c}(A) ∫ X 1 A ⋅ 1 d γ c = γ c ( A ) for A ∈ B ( X ) A\in\mathcal{B}(X) A ∈ B ( X ) , so the constant function 1 1 1 is a density of γ c \gamma_{c} γ c with respect to γ c \gamma_{c} γ c in the sense of The Radon-Nikodym Theorem for a Finite Measure and a Sigma-Finite Measure, and Uniqueness of Densities . Since exp ( 0 ) = 1 \exp(0)=1 exp ( 0 ) = 1 by claim 1 of Basic Properties of the Exponential Function , log 1 = log ( exp ( 0 ) ) = 0 \log1=\log(\exp(0))=0 log 1 = log ( exp ( 0 )) = 0 by The Natural Logarithm , so ϕ ∘ 1 \phi\circ1 ϕ ∘ 1 is the constant ϕ ( 1 ) = 1 ⋅ log 1 = 0 \phi(1)=1\cdot\log1=0 ϕ ( 1 ) = 1 ⋅ log 1 = 0 , which is integrable with integral 0 0 0 : by The Integral of an Indicator Function is the Measure of the Set with A = X A=X A = X , the constant 1 = 1 X 1=\mathbf{1}_{X} 1 = 1 X has ∫ X 1 d γ c = γ c ( X ) = 1 \int_{X}1\,d\gamma_{c}=\gamma_{c}(X)=1 ∫ X 1 d γ c = γ c ( X ) = 1 , so it is integrable, and its multiple by 0 0 0 is integrable with integral 0 0 0 by Linearity and Monotonicity of the Lebesgue Integral §integrable . By Relative Entropy of Probability Measures §relative-entropy , γ c \gamma_{c} γ c has finite relative entropy with respect to γ c \gamma_{c} γ c and H ( γ c ∣ γ c ) = 0 H(\gamma_{c}\,|\,\gamma_{c})=0 H ( γ c ∣ γ c ) = 0 ; thus γ c ∈ D \gamma_{c}\in\mathcal{D} γ c ∈ D and E ( γ c ) = 0 \mathcal{E}(\gamma_{c})=0 E ( γ c ) = 0 . By The Relative Score on a Hilbert Space: the Gaussian Measure Has Score Zero, and the Score as a Square-Integrable Field in the Weighted Sequence Space §gaussian , γ c ∈ P 2 ( X ) \gamma_{c}\in\mathcal{P}_{2}(X) γ c ∈ P 2 ( X ) has a relative score with respect to γ c \gamma_{c} γ c and finite Fisher information relative to γ c \gamma_{c} γ c with weights a a a , with I a ( γ c ∣ γ c ) = 0 \mathcal{I}_{a}(\gamma_{c}\,|\,\gamma_{c})=0 I a ( γ c ∣ γ c ) = 0 . Hence γ c ∈ D Σ \gamma_{c}\in\mathcal{D}_{\Sigma} γ c ∈ D Σ and, by Step 1, ∥ Σ ( γ c ) ∥ γ c 2 = 0 \lVert\Sigma(\gamma_{c})\rVert_{\gamma_{c}}^{2}=0 ∥ Σ ( γ c ) ∥ γ c 2 = 0 ; as the norm of the real Hilbert space L 2 ( γ c ; X a ) L^{2}(\gamma_{c};X^{a}) L 2 ( γ c ; X a ) of Probability Measures on a Hilbert Space Transported in the Noise Norm: Standing Notation §fields vanishes only at the zero vector, Σ ( γ c ) \Sigma(\gamma_{c}) Σ ( γ c ) is the zero vector of L 2 ( γ c ; X a ) L^{2}(\gamma_{c};X^{a}) L 2 ( γ c ; X a ) .
Step 3 (Claim 1). Let μ \mu μ be as in claim 1. Then μ ∈ D \mu\in\mathcal{D} μ ∈ D , μ ∈ P 2 ( X ) \mu\in\mathcal{P}_{2}(X) μ ∈ P 2 ( X ) by Relative Entropy with Respect to a Diagonal Gaussian Measure on a Hilbert Space: the Moment Bound, the Cutoff Projections, and Bounded, Tight, Weakly Closed, Wasserstein-Closed and Weakly Sequentially Compact Sublevel Sets §moment , and μ \mu μ has a relative score and finite Fisher information, so μ ∈ D Σ \mu\in\mathcal{D}_{\Sigma} μ ∈ D Σ . By Steps 1 and 2, Talagrand, HWI and Log-Sobolev Inequalities for a Uniformly Displacement Convex Noise Penalty Pair with a Point of Zero Score §log-sobolev applies to the pair of Step 1 with λ = 1 / κ \lambda=1/\kappa λ = 1/ κ and μ ∗ = γ c \mu_{*}=\gamma_{c} μ ∗ = γ c , and gives
H ( μ ∣ γ c ) − 0 ≤ κ 2 ∥ Σ ( μ ) ∥ μ 2 = κ 2 I a ( μ ∣ γ c ) . H(\mu\,|\,\gamma_{c})-0\le\frac{\kappa}{2}\,\lVert\Sigma(\mu)\rVert_{\mu}^{2}=\frac{\kappa}{2}\,\mathcal{I}_{a}(\mu\,|\,\gamma_{c}). H ( μ ∣ γ c ) − 0 ≤ 2 κ ∥ Σ ( μ ) ∥ μ 2 = 2 κ I a ( μ ∣ γ c ) .
Step 4 (Claim 2). Let μ \mu μ be as in claim 1. Then μ ∈ D ⊆ P ρ a \mu\in\mathcal{D}\subseteq\mathcal{P}^{a}_{\rho} μ ∈ D ⊆ P ρ a by The Entropy Domain of a Diagonal Gaussian Reference Measure Has the Noise Map Property §inclusion , and γ c = ρ ∈ P ρ a \gamma_{c}=\rho\in\mathcal{P}^{a}_{\rho} γ c = ρ ∈ P ρ a by The Noise Wasserstein Distance is a Metric on the Measures Noise-Connected to the Reference Measure: Existence of Noise-Optimal Couplings, Comparison with the Quadratic Wasserstein Distance and Lower Semicontinuity §reference . As in Step 3, μ ∈ D Σ \mu\in\mathcal{D}_{\Sigma} μ ∈ D Σ , and Talagrand, HWI and Log-Sobolev Inequalities for a Uniformly Displacement Convex Noise Penalty Pair with a Point of Zero Score §hwi with λ = 1 / κ \lambda=1/\kappa λ = 1/ κ and μ ∗ = γ c \mu_{*}=\gamma_{c} μ ∗ = γ c gives H ( μ ∣ γ c ) − 0 ≤ ∥ Σ ( μ ) ∥ μ W a ( μ , γ c ) − 1 2 κ W a ( μ , γ c ) 2 H(\mu\,|\,\gamma_{c})-0\le\lVert\Sigma(\mu)\rVert_{\mu}W_{a}(\mu,\gamma_{c})-\frac{1}{2\kappa}W_{a}(\mu,\gamma_{c})^{2} H ( μ ∣ γ c ) − 0 ≤ ∥ Σ ( μ ) ∥ μ W a ( μ , γ c ) − 2 κ 1 W a ( μ , γ c ) 2 . Since ∥ Σ ( μ ) ∥ μ \lVert\Sigma(\mu)\rVert_{\mu} ∥ Σ ( μ ) ∥ μ is nonnegative with square I a ( μ ∣ γ c ) \mathcal{I}_{a}(\mu\,|\,\gamma_{c}) I a ( μ ∣ γ c ) by Step 1, it equals I a ( μ ∣ γ c ) 1 / 2 \mathcal{I}_{a}(\mu\,|\,\gamma_{c})^{1/2} I a ( μ ∣ γ c ) 1/2 by Existence and Uniqueness of the Nonnegative Square Root , which proves claim 2.
Step 5 (Quadratic expressions in a cylindrical function). Let F ∈ F C b 1 ( X ) F\in\mathcal{F}C^{1}_{b}(X) F ∈ F C b 1 ( X ) have a representation ( n , ψ ) (n,\psi) ( n , ψ ) , and let α , s , b 0 ∈ R \alpha,s,b_{0}\in\mathbb{R} α , s , b 0 ∈ R . We show that α + s F + b 0 F 2 ∈ F C b 1 ( X ) \alpha+sF+b_{0} F^{2}\in\mathcal{F}C^{1}_{b}(X) α + s F + b 0 F 2 ∈ F C b 1 ( X ) and that for every k ∈ N k\in\mathbb{N} k ∈ N
∂ k ( α + s F + b 0 F 2 ) = ( s + 2 b 0 F ) ∂ k F on X , and both sides vanish for k > n . \partial_{k}\bigl(\alpha+sF+b_{0} F^{2}\bigr)=(s+2b_{0} F)\,\partial_{k}F\quad\text{on }X,\qquad\text{and both sides vanish for }k>n . ∂ k ( α + s F + b 0 F 2 ) = ( s + 2 b 0 F ) ∂ k F on X , and both sides vanish for k > n .
By Bounded Continuously Differentiable Functions with Bounded Partial Derivatives on Euclidean Space §bounded , ψ \psi ψ is of class C 1 C^{1} C 1 on R n \mathbb{R}^{n} R n and there are real B ≥ 0 B\ge0 B ≥ 0 and B i ≥ 0 B_{i}\ge0 B i ≥ 0 with ∣ ψ ∣ ≤ B |\psi|\le B ∣ ψ ∣ ≤ B and ∣ ∂ i ψ ∣ ≤ B i |\partial_{i}\psi|\le B_{i} ∣ ∂ i ψ ∣ ≤ B i for i ∈ [ n ] i\in[n] i ∈ [ n ] ; by clause 1 of C^k Maps on a Euclidean Open Set , read through its clause 3 as in Differential Calculus and Convexity on Euclidean Open Sets: Standing Notation §derivatives , ψ \psi ψ and each ∂ i ψ \partial_{i}\psi ∂ i ψ are continuous at every point of R n \mathbb{R}^{n} R n in the sense of Continuity at a Point for Maps Between Euclidean Spaces , which for real-valued maps reads: for every ε > 0 \varepsilon>0 ε > 0 there is δ > 0 \delta>0 δ > 0 with ∣ f ( y ) − f ( x ) ∣ < ε |f(y)-f(x)|<\varepsilon ∣ f ( y ) − f ( x ) ∣ < ε whenever ∥ y − x ∥ < δ \lVert y-x\rVert<\delta ∥ y − x ∥ < δ . Let χ = α + s ψ + b 0 ψ 2 : R n → R \chi=\alpha+s\psi+b_{0}\psi^{2}:\mathbb{R}^{n}\to\mathbb{R} χ = α + s ψ + b 0 ψ 2 : R n → R . It is bounded by ∣ α ∣ + ∣ s ∣ B + ∣ b 0 ∣ B 2 |\alpha|+|s|B+|b_{0}|B^{2} ∣ α ∣ + ∣ s ∣ B + ∣ b 0 ∣ B 2 . It is continuous at every x x x : since ψ ( y ) 2 − ψ ( x ) 2 = ( ψ ( y ) + ψ ( x ) ) ( ψ ( y ) − ψ ( x ) ) \psi(y)^{2}-\psi(x)^{2}=(\psi(y)+\psi(x))(\psi(y)-\psi(x)) ψ ( y ) 2 − ψ ( x ) 2 = ( ψ ( y ) + ψ ( x )) ( ψ ( y ) − ψ ( x )) , we have ∣ χ ( y ) − χ ( x ) ∣ ≤ ( ∣ s ∣ + 2 ∣ b 0 ∣ B ) ∣ ψ ( y ) − ψ ( x ) ∣ |\chi(y)-\chi(x)|\le(|s|+2|b_{0}|B)\,|\psi(y)-\psi(x)| ∣ χ ( y ) − χ ( x ) ∣ ≤ ( ∣ s ∣ + 2∣ b 0 ∣ B ) ∣ ψ ( y ) − ψ ( x ) ∣ , and it suffices to take δ \delta δ for ψ \psi ψ at x x x with ε / ( ∣ s ∣ + 2 ∣ b 0 ∣ B + 1 ) \varepsilon/(|s|+2|b_{0}|B+1) ε / ( ∣ s ∣ + 2∣ b 0 ∣ B + 1 ) in place of ε \varepsilon ε . Fix x ∈ R n x\in\mathbb{R}^{n} x ∈ R n and i ∈ [ n ] i\in[n] i ∈ [ n ] , let e i e_{i} e i be the i i i -th standard basis vector of R n \mathbb{R}^{n} R n , and let u : R → R u:\mathbb{R}\to\mathbb{R} u : R → R , u ( t ) = ψ ( x + t e i ) u(t)=\psi(x+te_{i}) u ( t ) = ψ ( x + t e i ) . The real line is an interval of which 0 0 0 is an interior point by claim 1 of One-Dimensional Derivatives, Partial Derivatives, and Smoothness on the Real Line , and the condition of Partial Derivative on a Euclidean Open Set for ψ \psi ψ at x x x in the i i i -th variable is, word for word, the condition of Derivative at an Interior Point for u u u at 0 0 0 , the extra clauses of the two definitions being vacuous here: the requirement x + h e i ∈ R n x+he_{i}\in\mathbb{R}^{n} x + h e i ∈ R n of the former holds for every h h h , the open set being all of R n \mathbb{R}^{n} R n , and the restriction 0 + h ∈ R 0+h\in\mathbb{R} 0 + h ∈ R of the latter excludes no h h h , the interval being all of R \mathbb{R} R ; so u u u is differentiable at 0 0 0 with u ′ ( 0 ) = ∂ i ψ ( x ) u'(0)=\partial_{i}\psi(x) u ′ ( 0 ) = ∂ i ψ ( x ) . By claims 1 to 3 of Sum, Constant Multiple, and Product Rules for One-Dimensional Derivatives , v = α + s u + b 0 u u v=\alpha+su+b_{0} uu v = α + s u + b 0 uu is differentiable at 0 0 0 with v ′ ( 0 ) = s u ′ ( 0 ) + 2 b 0 u ( 0 ) u ′ ( 0 ) v'(0)=su'(0)+2b_{0} u(0)u'(0) v ′ ( 0 ) = s u ′ ( 0 ) + 2 b 0 u ( 0 ) u ′ ( 0 ) , and since v ( t ) = χ ( x + t e i ) v(t)=\chi(x+te_{i}) v ( t ) = χ ( x + t e i ) , reading the same identification backwards gives that ∂ i χ ( x ) \partial_{i}\chi(x) ∂ i χ ( x ) exists and equals ( s + 2 b 0 ψ ( x ) ) ∂ i ψ ( x ) (s+2b_{0}\psi(x))\,\partial_{i}\psi(x) ( s + 2 b 0 ψ ( x )) ∂ i ψ ( x ) . The function ∂ i χ = ( s + 2 b 0 ψ ) ∂ i ψ \partial_{i}\chi=(s+2b_{0}\psi)\partial_{i}\psi ∂ i χ = ( s + 2 b 0 ψ ) ∂ i ψ is bounded by ( ∣ s ∣ + 2 ∣ b 0 ∣ B ) B i (|s|+2|b_{0}|B)B_{i} ( ∣ s ∣ + 2∣ b 0 ∣ B ) B i , and is continuous at every x x x because
∣ ∂ i χ ( y ) − ∂ i χ ( x ) ∣ ≤ ( ∣ s ∣ + 2 ∣ b 0 ∣ B ) ∣ ∂ i ψ ( y ) − ∂ i ψ ( x ) ∣ + 2 ∣ b 0 ∣ B i ∣ ψ ( y ) − ψ ( x ) ∣ , |\partial_{i}\chi(y)-\partial_{i}\chi(x)|\le(|s|+2|b_{0}|B)\,|\partial_{i}\psi(y)-\partial_{i}\psi(x)|+2|b_{0}|B_{i}\,|\psi(y)-\psi(x)|, ∣ ∂ i χ ( y ) − ∂ i χ ( x ) ∣ ≤ ( ∣ s ∣ + 2∣ b 0 ∣ B ) ∣ ∂ i ψ ( y ) − ∂ i ψ ( x ) ∣ + 2∣ b 0 ∣ B i ∣ ψ ( y ) − ψ ( x ) ∣ ,
so it suffices to take the smaller of the two δ \delta δ 's for ∂ i ψ \partial_{i}\psi ∂ i ψ and ψ \psi ψ at x x x with ε / ( 2 ( ∣ s ∣ + 2 ∣ b 0 ∣ B ) + 1 ) \varepsilon/(2(|s|+2|b_{0}|B)+1) ε / ( 2 ( ∣ s ∣ + 2∣ b 0 ∣ B ) + 1 ) and ε / ( 4 ∣ b 0 ∣ B i + 1 ) \varepsilon/(4|b_{0}|B_{i}+1) ε / ( 4∣ b 0 ∣ B i + 1 ) in place of ε \varepsilon ε . Hence χ \chi χ is of class C 1 C^{1} C 1 on R n \mathbb{R}^{n} R n by clause 1 of C^k Maps on a Euclidean Open Set , and χ ∈ C b 1 ( R n ) \chi\in C^{1}_{b}(\mathbb{R}^{n}) χ ∈ C b 1 ( R n ) by Bounded Continuously Differentiable Functions with Bounded Partial Derivatives on Euclidean Space §bounded . Since χ ∘ p n = α + s F + b 0 F 2 \chi\circ p_{n}=\alpha+sF+b_{0} F^{2} χ ∘ p n = α + s F + b 0 F 2 , this function lies in F C b 1 ( X ) \mathcal{F}C^{1}_{b}(X) F C b 1 ( X ) by Bounded C^1 Cylindrical Functions on a Hilbert Space with an Orthonormal Basis §cylindrical , with representation ( n , χ ) (n,\chi) ( n , χ ) , and Bounded C^1 Cylindrical Functions: Linear Structure, the Gradient and the Partial Derivatives, Bounds and Integrability, and Density in the Square-Integrable Functions §partial , applied to both representations ( n , χ ) (n,\chi) ( n , χ ) and ( n , ψ ) (n,\psi) ( n , ψ ) , gives for k ≤ n k\le n k ≤ n that ∂ k ( χ ∘ p n ) = ( ∂ k χ ) ∘ p n = ( s + 2 b 0 F ) ( ∂ k ψ ) ∘ p n = ( s + 2 b 0 F ) ∂ k F \partial_{k}(\chi\circ p_{n})=(\partial_{k}\chi)\circ p_{n}=(s+2b_{0} F)\,(\partial_{k}\psi)\circ p_{n}=(s+2b_{0} F)\,\partial_{k}F ∂ k ( χ ∘ p n ) = ( ∂ k χ ) ∘ p n = ( s + 2 b 0 F ) ( ∂ k ψ ) ∘ p n = ( s + 2 b 0 F ) ∂ k F , and for k > n k>n k > n that both ∂ k ( χ ∘ p n ) \partial_{k}(\chi\circ p_{n}) ∂ k ( χ ∘ p n ) and ∂ k F \partial_{k}F ∂ k F vanish identically. In particular, by the display of The Noise Gradient of a Cylindrical Function and the Noise Tangent Space at a Probability Measure on a Hilbert Space §gradient computed with these representations, ∣ ∇ a ( α + s F + b 0 F 2 ) ∣ a 2 = ∑ k = 1 n a k ( s + 2 b 0 F ) 2 ( ∂ k F ) 2 |\nabla_{a}(\alpha+sF+b_{0} F^{2})|_{a}^{2}=\sum_{k=1}^{n}a_{k}(s+2b_{0} F)^{2}(\partial_{k}F)^{2} ∣ ∇ a ( α + s F + b 0 F 2 ) ∣ a 2 = ∑ k = 1 n a k ( s + 2 b 0 F ) 2 ( ∂ k F ) 2 .
Step 6 (Entropies of bounded functions). Let G : X → R G:X\to\mathbb{R} G : X → R be Borel with values in an interval [ 0 , b ] [0,b] [ 0 , b ] . The function ϕ \phi ϕ is continuous on [ 0 , ∞ ) [0,\infty) [ 0 , ∞ ) by The Function s log s s\log s s log s : Continuity, Young's Inequality and Lower Bounds, with the Elementary Bounds for the Exponential and the Logarithm §continuous , hence on [ 0 , b ] [0,b] [ 0 , b ] , so Extreme Value Theorem on a Closed Real Interval , applied to the restriction of ϕ \phi ϕ to [ 0 , b ] [0,b] [ 0 , b ] , gives r min , r max ∈ [ 0 , b ] r_{\min},r_{\max}\in[0,b] r m i n , r m a x ∈ [ 0 , b ] with ϕ ( r min ) ≤ ϕ ( r ) ≤ ϕ ( r max ) \phi(r_{\min})\le\phi(r)\le\phi(r_{\max}) ϕ ( r m i n ) ≤ ϕ ( r ) ≤ ϕ ( r m a x ) for r ∈ [ 0 , b ] r\in[0,b] r ∈ [ 0 , b ] ; with M b = max ( ∣ ϕ ( r min ) ∣ , ∣ ϕ ( r max ) ∣ ) M_{b}=\max(|\phi(r_{\min})|,|\phi(r_{\max})|) M b = max ( ∣ ϕ ( r m i n ) ∣ , ∣ ϕ ( r m a x ) ∣ ) we get ∣ ϕ ( r ) ∣ ≤ M b |\phi(r)|\le M_{b} ∣ ϕ ( r ) ∣ ≤ M b for r ∈ [ 0 , b ] r\in[0,b] r ∈ [ 0 , b ] ; ϕ ∘ G \phi\circ G ϕ ∘ G is Borel by The Function s log s s\log s s log s : Continuity, Young's Inequality and Lower Bounds, with the Elementary Bounds for the Exponential and the Logarithm §continuous . Thus G G G and ϕ ∘ G \phi\circ G ϕ ∘ G are bounded and Borel, hence integrable with respect to every Borel probability measure on X X X by Relative Entropy on a Measurable Space: the Gibbs Inequality, the Variational Criterion and Formula, the Entropy Inequality, Small Sets and Data Processing §functional , and Ent γ c ( G ) \operatorname{Ent}_{\gamma_{c}}(G) Ent γ c ( G ) is defined by The Entropy of a Nonnegative Function with Respect to a Probability Measure §entropy . For a real constant r r r , ∫ X r d γ c = r \int_{X}r\,d\gamma_{c}=r ∫ X r d γ c = r by The Integral of an Indicator Function is the Measure of the Set and Linearity and Monotonicity of the Lebesgue Integral §integrable , γ c \gamma_{c} γ c being a probability measure. Now let F ∈ F C b 1 ( X ) F\in\mathcal{F}C^{1}_{b}(X) F ∈ F C b 1 ( X ) with representation ( n , ψ ) (n,\psi) ( n , ψ ) , and fix a real B ≥ 1 B\ge1 B ≥ 1 with ∣ ψ ∣ ≤ B |\psi|\le B ∣ ψ ∣ ≤ B , so that ∣ F ∣ ≤ B |F|\le B ∣ F ∣ ≤ B on X X X . Since F F F is Borel by Bounded C^1 Cylindrical Functions: Linear Structure, the Gradient and the Partial Derivatives, Bounds and Integrability, and Density in the Square-Integrable Functions §bounded-borel , F 2 F^{2} F 2 is Borel by claim 3 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions , nonnegative and bounded by B 2 B^{2} B 2 ; so F 2 F^{2} F 2 and ϕ ∘ F 2 \phi\circ F^{2} ϕ ∘ F 2 are integrable with respect to γ c \gamma_{c} γ c , which is the first part of claim 3. We write D = ∫ X ∣ ∇ a F ∣ a 2 d γ c D=\int_{X}|\nabla_{a}F|_{a}^{2}\,d\gamma_{c} D = ∫ X ∣ ∇ a F ∣ a 2 d γ c , a nonnegative real number by The Noise Gradient of a Cylindrical Function and the Noise Tangent Space at a Probability Measure on a Hilbert Space §gradient .
Step 7 (The inequality for F 2 + ε F^{2}+\varepsilon F 2 + ε ). Let F F F , ( n , ψ ) (n,\psi) ( n , ψ ) , B B B and D D D be as in Step 6, and let ε \varepsilon ε be real with 0 < ε ≤ 1 0<\varepsilon\le1 0 < ε ≤ 1 . Let g = ε + F 2 g=\varepsilon+F^{2} g = ε + F 2 . By Step 5 (with α = ε \alpha=\varepsilon α = ε , s = 0 s=0 s = 0 , b 0 = 1 b_{0}=1 b 0 = 1 ), g ∈ F C b 1 ( X ) g\in\mathcal{F}C^{1}_{b}(X) g ∈ F C b 1 ( X ) and ∂ k g = 2 F ∂ k F \partial_{k}g=2F\,\partial_{k}F ∂ k g = 2 F ∂ k F for every k k k , with ∂ k g = 0 \partial_{k}g=0 ∂ k g = 0 for k > n k>n k > n ; and ε ≤ g ≤ B 2 + 1 \varepsilon\le g\le B^{2}+1 ε ≤ g ≤ B 2 + 1 . Let N ε = ∫ X g d γ c = ε + ∫ X F 2 d γ c N_{\varepsilon}=\int_{X}g\,d\gamma_{c}=\varepsilon+\int_{X}F^{2}\,d\gamma_{c} N ε = ∫ X g d γ c = ε + ∫ X F 2 d γ c , so ε ≤ N ε ≤ B 2 + 1 \varepsilon\le N_{\varepsilon}\le B^{2}+1 ε ≤ N ε ≤ B 2 + 1 . The function g / N ε : X → [ 0 , ∞ ) g/N_{\varepsilon}:X\to[0,\infty) g / N ε : X → [ 0 , ∞ ) is Borel, so by claim 3 of Image Measures, Measures with Densities, and Change of Variables the measure μ \mu μ with density g / N ε g/N_{\varepsilon} g / N ε with respect to γ c \gamma_{c} γ c , μ ( A ) = ∫ X 1 A ( g / N ε ) d γ c \mu(A)=\int_{X}\mathbf{1}_{A}\,(g/N_{\varepsilon})\,d\gamma_{c} μ ( A ) = ∫ X 1 A ( g / N ε ) d γ c , is a measure on B ( X ) \mathcal{B}(X) B ( X ) , with μ ( X ) = N ε − 1 ∫ X g d γ c = 1 \mu(X)=N_{\varepsilon}^{-1}\int_{X}g\,d\gamma_{c}=1 μ ( X ) = N ε − 1 ∫ X g d γ c = 1 ; so μ ∈ P ( X ) \mu\in\mathcal{P}(X) μ ∈ P ( X ) and g / N ε g/N_{\varepsilon} g / N ε is a density of μ \mu μ with respect to γ c \gamma_{c} γ c in the sense of The Radon-Nikodym Theorem for a Finite Measure and a Sigma-Finite Measure, and Uniqueness of Densities . We use repeatedly that, by claim 3 of Image Measures, Measures with Densities, and Change of Variables , a Borel f : X → R f:X\to\mathbb{R} f : X → R with f g / N ε f\,g/N_{\varepsilon} f g / N ε integrable with respect to γ c \gamma_{c} γ c is integrable with respect to μ \mu μ , with ∫ X f d μ = N ε − 1 ∫ X f g d γ c \int_{X}f\,d\mu=N_{\varepsilon}^{-1}\int_{X}f\,g\,d\gamma_{c} ∫ X f d μ = N ε − 1 ∫ X f g d γ c .
Entropy. The function g / N ε g/N_{\varepsilon} g / N ε is Borel with values in [ 0 , ( B 2 + 1 ) / ε ] [0,(B^{2}+1)/\varepsilon] [ 0 , ( B 2 + 1 ) / ε ] , so ϕ ∘ ( g / N ε ) \phi\circ(g/N_{\varepsilon}) ϕ ∘ ( g / N ε ) is integrable with respect to γ c \gamma_{c} γ c by Step 6, and μ \mu μ has finite relative entropy with respect to γ c \gamma_{c} γ c , with H ( μ ∣ γ c ) = ∫ X ϕ ∘ ( g / N ε ) d γ c H(\mu\,|\,\gamma_{c})=\int_{X}\phi\circ(g/N_{\varepsilon})\,d\gamma_{c} H ( μ ∣ γ c ) = ∫ X ϕ ∘ ( g / N ε ) d γ c , by Relative Entropy of Probability Measures §relative-entropy . For r > 0 r>0 r > 0 , log r = log ( r / N ε ) + log N ε \log r=\log(r/N_{\varepsilon})+\log N_{\varepsilon} log r = log ( r / N ε ) + log N ε by The Natural Logarithm , so ϕ ( r / N ε ) = ( r / N ε ) log ( r / N ε ) = N ε − 1 ( ϕ ( r ) − r log N ε ) \phi(r/N_{\varepsilon})=(r/N_{\varepsilon})\log(r/N_{\varepsilon})=N_{\varepsilon}^{-1}\bigl(\phi(r)-r\log N_{\varepsilon}\bigr) ϕ ( r / N ε ) = ( r / N ε ) log ( r / N ε ) = N ε − 1 ( ϕ ( r ) − r log N ε ) . Applying this with r = g ( x ) > 0 r=g(x)>0 r = g ( x ) > 0 and integrating, using Linearity and Monotonicity of the Lebesgue Integral §integrable , Step 6 for g g g , and The Entropy of a Nonnegative Function with Respect to a Probability Measure §entropy ,
H ( μ ∣ γ c ) = N ε − 1 ( ∫ X ϕ ∘ g d γ c − N ε log N ε ) = N ε − 1 ( ∫ X ϕ ∘ g d γ c − ϕ ( N ε ) ) = N ε − 1 Ent γ c ( g ) . H(\mu\,|\,\gamma_{c})=N_{\varepsilon}^{-1}\Bigl(\int_{X}\phi\circ g\,d\gamma_{c}-N_{\varepsilon}\log N_{\varepsilon}\Bigr)=N_{\varepsilon}^{-1}\Bigl(\int_{X}\phi\circ g\,d\gamma_{c}-\phi(N_{\varepsilon})\Bigr)=N_{\varepsilon}^{-1}\operatorname{Ent}_{\gamma_{c}}(g). H ( μ ∣ γ c ) = N ε − 1 ( ∫ X ϕ ∘ g d γ c − N ε log N ε ) = N ε − 1 ( ∫ X ϕ ∘ g d γ c − ϕ ( N ε ) ) = N ε − 1 Ent γ c ( g ) .
In particular μ ∈ P 2 ( X ) \mu\in\mathcal{P}_{2}(X) μ ∈ P 2 ( X ) by Relative Entropy with Respect to a Diagonal Gaussian Measure on a Hilbert Space: the Moment Bound, the Cutoff Projections, and Bounded, Tight, Weakly Closed, Wasserstein-Closed and Weakly Sequentially Compact Sublevel Sets §moment .
Relative score. For k ∈ N k\in\mathbb{N} k ∈ N let ζ k = ∂ k g ⋅ ( 1 / g ) \zeta_{k}=\partial_{k}g\cdot(1/g) ζ k = ∂ k g ⋅ ( 1/ g ) . The function 1 / g 1/g 1/ g is Borel by the criterion of Measure Spaces and the Lebesgue Integral: Standing Notation §measurable : { 1 / g > r } \{1/g>r\} { 1/ g > r } is X X X for r ≤ 0 r\le0 r ≤ 0 and { − g > − 1 / r } \{-g>-1/r\} { − g > − 1/ r } for r > 0 r>0 r > 0 , Borel by claim 2 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions . As ∂ k g \partial_{k}g ∂ k g is bounded and Borel by Bounded C^1 Cylindrical Functions: Linear Structure, the Gradient and the Partial Derivatives, Bounds and Integrability, and Density in the Square-Integrable Functions §bounded-borel and 0 < 1 / g ≤ 1 / ε 0<1/g\le1/\varepsilon 0 < 1/ g ≤ 1/ ε , ζ k \zeta_{k} ζ k is bounded and Borel (claim 3 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions ); so ζ k 2 \zeta_{k}^{2} ζ k 2 is integrable with respect to μ \mu μ by Relative Entropy on a Measurable Space: the Gibbs Inequality, the Variational Criterion and Formula, the Entropy Inequality, Small Sets and Data Processing §functional , and ζ k \zeta_{k} ζ k defines an element of L 2 ( μ ) L^{2}(\mu) L 2 ( μ ) , again written ζ k \zeta_{k} ζ k ; for k > n k>n k > n , ζ k = 0 \zeta_{k}=0 ζ k = 0 . Let φ ∈ F C b 1 ( X ) \varphi\in\mathcal{F}C^{1}_{b}(X) φ ∈ F C b 1 ( X ) and put q ( x ) = x k φ ( x ) / c k − ∂ k φ ( x ) q(x)=x_{k}\varphi(x)/c_{k}-\partial_{k}\varphi(x) q ( x ) = x k φ ( x ) / c k − ∂ k φ ( x ) . By The Lebesgue Space of Square-Integrable Functions is a Real Hilbert Space §inner-product , then the density formula (the function ζ k φ g / N ε = N ε − 1 ∂ k g φ \zeta_{k}\varphi\,g/N_{\varepsilon}=N_{\varepsilon}^{-1}\partial_{k}g\,\varphi ζ k φ g / N ε = N ε − 1 ∂ k g φ being bounded and Borel), then Gaussian Integration by Parts for Products of Cylindrical Functions, and Closability of the Noise Gradient in the Gaussian Lebesgue Space §ibp applied with g g g in place of F F F and with φ \varphi φ , then Linearity and Monotonicity of the Lebesgue Integral §integrable and the density formula once more (the function q g / N ε q\,g/N_{\varepsilon} q g / N ε being integrable with respect to γ c \gamma_{c} γ c as N ε − 1 N_{\varepsilon}^{-1} N ε − 1 times the integrable function g q g\,q g q of Gaussian Integration by Parts for Products of Cylindrical Functions, and Closability of the Noise Gradient in the Gaussian Lebesgue Space §ibp ),
⟨ ζ k , φ ⟩ L 2 ( μ ) = ∫ X ζ k φ d μ = 1 N ε ∫ X ∂ k g φ d γ c = 1 N ε ∫ X g q d γ c = ∫ X ( x k c k φ ( x ) − ∂ k φ ( x ) ) μ ( d x ) . \langle\zeta_{k},\varphi\rangle_{L^{2}(\mu)}=\int_{X}\zeta_{k}\varphi\,d\mu=\frac{1}{N_{\varepsilon}}\int_{X}\partial_{k}g\,\varphi\,d\gamma_{c}=\frac{1}{N_{\varepsilon}}\int_{X}g\,q\,d\gamma_{c}=\int_{X}\Bigl(\frac{x_{k}}{c_{k}}\,\varphi(x)-\partial_{k}\varphi(x)\Bigr)\mu(dx). ⟨ ζ k , φ ⟩ L 2 ( μ ) = ∫ X ζ k φ d μ = N ε 1 ∫ X ∂ k g φ d γ c = N ε 1 ∫ X g q d γ c = ∫ X ( c k x k φ ( x ) − ∂ k φ ( x ) ) μ ( d x ) .
By The Relative Score with Respect to a Diagonal Gaussian Measure on a Hilbert Space §score , μ \mu μ has a relative score with respect to γ c \gamma_{c} γ c , with components ζ k \zeta_{k} ζ k .
Fisher information. For k > n k>n k > n the term a k ∥ ζ k ∥ L 2 ( μ ) 2 a_{k}\lVert\zeta_{k}\rVert_{L^{2}(\mu)}^{2} a k ∥ ζ k ∥ L 2 ( μ ) 2 is 0 0 0 , so the partial sums of ∑ k = 1 ∞ a k ∥ ζ k ∥ L 2 ( μ ) 2 \sum_{k=1}^{\infty}a_{k}\lVert\zeta_{k}\rVert_{L^{2}(\mu)}^{2} ∑ k = 1 ∞ a k ∥ ζ k ∥ L 2 ( μ ) 2 are constant from the n n n -th on; the series converges with sum ∑ k = 1 n a k ∥ ζ k ∥ L 2 ( μ ) 2 \sum_{k=1}^{n}a_{k}\lVert\zeta_{k}\rVert_{L^{2}(\mu)}^{2} ∑ k = 1 n a k ∥ ζ k ∥ L 2 ( μ ) 2 , and by Weight Sequences and the Weighted Fisher Information Relative to a Diagonal Gaussian Measure on a Hilbert Space §information μ \mu μ has finite Fisher information relative to γ c \gamma_{c} γ c with weights a a a and I a ( μ ∣ γ c ) \mathcal{I}_{a}(\mu\,|\,\gamma_{c}) I a ( μ ∣ γ c ) is that finite sum. For k ≤ n k\le n k ≤ n , by The Lebesgue Space of Square-Integrable Functions is a Real Hilbert Space §inner-product and the density formula, ∥ ζ k ∥ L 2 ( μ ) 2 = ∫ X ζ k 2 d μ = N ε − 1 ∫ X ( ∂ k g ) 2 / g d γ c \lVert\zeta_{k}\rVert_{L^{2}(\mu)}^{2}=\int_{X}\zeta_{k}^{2}\,d\mu=N_{\varepsilon}^{-1}\int_{X}(\partial_{k}g)^{2}/g\,d\gamma_{c} ∥ ζ k ∥ L 2 ( μ ) 2 = ∫ X ζ k 2 d μ = N ε − 1 ∫ X ( ∂ k g ) 2 / g d γ c , and pointwise
( ∂ k g ) 2 g = 4 F 2 ( ∂ k F ) 2 ε + F 2 ≤ 4 ( ∂ k F ) 2 , \frac{(\partial_{k}g)^{2}}{g}=\frac{4F^{2}(\partial_{k}F)^{2}}{\varepsilon+F^{2}}\le4(\partial_{k}F)^{2}, g ( ∂ k g ) 2 = ε + F 2 4 F 2 ( ∂ k F ) 2 ≤ 4 ( ∂ k F ) 2 ,
since 0 ≤ F 2 ≤ ε + F 2 0\le F^{2}\le\varepsilon+F^{2} 0 ≤ F 2 ≤ ε + F 2 . By the monotonicity and linearity of Linearity and Monotonicity of the Lebesgue Integral §integrable and the display of The Noise Gradient of a Cylindrical Function and the Noise Tangent Space at a Probability Measure on a Hilbert Space §gradient for the representation ( n , ψ ) (n,\psi) ( n , ψ ) ,
I a ( μ ∣ γ c ) ≤ 4 N ε ∑ k = 1 n a k ∫ X ( ∂ k F ) 2 d γ c = 4 N ε ∫ X ∣ ∇ a F ∣ a 2 d γ c = 4 D N ε . \mathcal{I}_{a}(\mu\,|\,\gamma_{c})\le\frac{4}{N_{\varepsilon}}\sum_{k=1}^{n}a_{k}\int_{X}(\partial_{k}F)^{2}\,d\gamma_{c}=\frac{4}{N_{\varepsilon}}\int_{X}|\nabla_{a}F|_{a}^{2}\,d\gamma_{c}=\frac{4D}{N_{\varepsilon}}. I a ( μ ∣ γ c ) ≤ N ε 4 k = 1 ∑ n a k ∫ X ( ∂ k F ) 2 d γ c = N ε 4 ∫ X ∣ ∇ a F ∣ a 2 d γ c = N ε 4 D .
Conclusion. By claim 1 (Step 3) applied to μ \mu μ , N ε − 1 Ent γ c ( g ) = H ( μ ∣ γ c ) ≤ κ 2 ⋅ 4 D N ε N_{\varepsilon}^{-1}\operatorname{Ent}_{\gamma_{c}}(g)=H(\mu\,|\,\gamma_{c})\le\frac{\kappa}{2}\cdot\frac{4D}{N_{\varepsilon}} N ε − 1 Ent γ c ( g ) = H ( μ ∣ γ c ) ≤ 2 κ ⋅ N ε 4 D , and multiplying by N ε > 0 N_{\varepsilon}>0 N ε > 0 ,
Ent γ c ( F 2 + ε ) ≤ 2 κ D ( 0 < ε ≤ 1 ) . \operatorname{Ent}_{\gamma_{c}}(F^{2}+\varepsilon)\le2\kappa D\qquad(0<\varepsilon\le1). Ent γ c ( F 2 + ε ) ≤ 2 κ D ( 0 < ε ≤ 1 ) .
Step 8 (Claim 3: letting ε → 0 \varepsilon\to0 ε → 0 ). Let F F F , B B B , D D D be as in Step 6, and for j ∈ N j\in\mathbb{N} j ∈ N let g j = F 2 + 1 / j g_{j}=F^{2}+1/j g j = F 2 + 1/ j , with values in [ 0 , B 2 + 1 ] [0,B^{2}+1] [ 0 , B 2 + 1 ] . The sequence ( 1 / j ) j (1/j)_{j} ( 1/ j ) j converges to 0 0 0 in the sense of Limit of a Sequence of Real Numbers : given a real θ > 0 \theta>0 θ > 0 , claim 3 of The Archimedean Property of the Real Numbers gives j 0 ∈ N j_{0}\in\mathbb{N} j 0 ∈ N with 0 < 1 / j 0 < θ 0<1/j_{0}<\theta 0 < 1/ j 0 < θ , and then 0 < 1 / j ≤ 1 / j 0 < θ 0<1/j\le1/j_{0}<\theta 0 < 1/ j ≤ 1/ j 0 < θ for j ≥ j 0 j\ge j_{0} j ≥ j 0 . As a constant sequence converges to its value, Arithmetic of Limits of Real Sequences §sums shows that for each x ∈ X x\in X x ∈ X the sequence ( F ( x ) 2 + 1 / j ) j (F(x)^{2}+1/j)_{j} ( F ( x ) 2 + 1/ j ) j in [ 0 , ∞ ) [0,\infty) [ 0 , ∞ ) converges to F ( x ) 2 F(x)^{2} F ( x ) 2 , so ( ϕ ( g j ( x ) ) ) j (\phi(g_{j}(x)))_{j} ( ϕ ( g j ( x )) ) j converges to ϕ ( F ( x ) 2 ) \phi(F(x)^{2}) ϕ ( F ( x ) 2 ) by Continuity Between Metric Spaces is Equivalent to Sequential Continuity §sequential and the continuity of ϕ \phi ϕ (The Function s log s s\log s s log s : Continuity, Young's Inequality and Lower Bounds, with the Elementary Bounds for the Exponential and the Logarithm §continuous ); moreover ∣ ϕ ∘ g j ∣ ≤ M B 2 + 1 |\phi\circ g_{j}|\le M_{B^{2}+1} ∣ ϕ ∘ g j ∣ ≤ M B 2 + 1 (Step 6), a constant, which is integrable. By Dominated Convergence Theorem , ∫ X ϕ ∘ g j d γ c → ∫ X ϕ ∘ F 2 d γ c \int_{X}\phi\circ g_{j}\,d\gamma_{c}\to\int_{X}\phi\circ F^{2}\,d\gamma_{c} ∫ X ϕ ∘ g j d γ c → ∫ X ϕ ∘ F 2 d γ c . Also ∫ X g j d γ c = ∫ X F 2 d γ c + 1 / j → ∫ X F 2 d γ c \int_{X}g_{j}\,d\gamma_{c}=\int_{X}F^{2}\,d\gamma_{c}+1/j\to\int_{X}F^{2}\,d\gamma_{c} ∫ X g j d γ c = ∫ X F 2 d γ c + 1/ j → ∫ X F 2 d γ c by the same sum clause, so ϕ ( ∫ X g j d γ c ) → ϕ ( ∫ X F 2 d γ c ) \phi\bigl(\int_{X}g_{j}\,d\gamma_{c}\bigr)\to\phi\bigl(\int_{X}F^{2}\,d\gamma_{c}\bigr) ϕ ( ∫ X g j d γ c ) → ϕ ( ∫ X F 2 d γ c ) by the same continuity. By Arithmetic of Limits of Real Sequences §scalar , Ent γ c ( g j ) → Ent γ c ( F 2 ) \operatorname{Ent}_{\gamma_{c}}(g_{j})\to\operatorname{Ent}_{\gamma_{c}}(F^{2}) Ent γ c ( g j ) → Ent γ c ( F 2 ) ; since Ent γ c ( g j ) ≤ 2 κ D \operatorname{Ent}_{\gamma_{c}}(g_{j})\le2\kappa D Ent γ c ( g j ) ≤ 2 κ D for every j j j by Step 7, claim 1 of Order Properties of Limits of Real Sequences gives Ent γ c ( F 2 ) ≤ 2 κ D \operatorname{Ent}_{\gamma_{c}}(F^{2})\le2\kappa D Ent γ c ( F 2 ) ≤ 2 κ D , which completes claim 3.
Step 9 (A second-order expansion of ϕ \phi ϕ at 1 1 1 ). Let U = ( 0 , ∞ ) U=(0,\infty) U = ( 0 , ∞ ) , open in R 1 \mathbb{R}^{1} R 1 by Open Subset of Euclidean Space (for t ∈ U t\in U t ∈ U , every y y y with ( y − t ) 2 < t 2 (y-t)^{2}<t^{2} ( y − t ) 2 < t 2 is positive), and let f : U → R f:U\to\mathbb{R} f : U → R , f ( t ) = t log t = ϕ ( t ) f(t)=t\log t=\phi(t) f ( t ) = t log t = ϕ ( t ) . Every t ∈ U t\in U t ∈ U is an interior point of the interval U U U , as t / 2 < t < 2 t t/2<t<2t t /2 < t < 2 t . We first show that log \log log , defined in The Natural Logarithm , is differentiable at every t ∈ U t\in U t ∈ U in the sense of Derivative at an Interior Point , with derivative 1 / t 1/t 1/ t . Fix t ∈ U t\in U t ∈ U and a real h h h with 0 < ∣ h ∣ < t / 2 0<|h|<t/2 0 < ∣ h ∣ < t /2 , and put τ = 1 + h / t \tau=1+h/t τ = 1 + h / t , so that τ > 1 / 2 \tau>1/2 τ > 1/2 and t + h = t τ t+h=t\tau t + h = t τ . By the identity log ( t τ ) = log t + log τ \log(t\tau)=\log t+\log\tau log ( t τ ) = log t + log τ recorded in The Natural Logarithm , log ( t + h ) − log t = log τ \log(t+h)-\log t=\log\tau log ( t + h ) − log t = log τ , and by The Function s log s s\log s s log s : Continuity, Young's Inequality and Lower Bounds, with the Elementary Bounds for the Exponential and the Logarithm §log , 1 − τ − 1 ≤ log τ ≤ τ − 1 1-\tau^{-1}\le\log\tau\le\tau-1 1 − τ − 1 ≤ log τ ≤ τ − 1 , that is, h / ( t + h ) ≤ log ( t + h ) − log t ≤ h / t h/(t+h)\le\log(t+h)-\log t\le h/t h / ( t + h ) ≤ log ( t + h ) − log t ≤ h / t . Dividing by h h h , the difference quotient Q h = ( log ( t + h ) − log t ) / h Q_{h}=(\log(t+h)-\log t)/h Q h = ( log ( t + h ) − log t ) / h lies between 1 / ( t + h ) 1/(t+h) 1/ ( t + h ) and 1 / t 1/t 1/ t (in this order if h > 0 h>0 h > 0 , in the reverse order if h < 0 h<0 h < 0 ), so, using t + h > t / 2 t+h>t/2 t + h > t /2 ,
∣ Q h − 1 t ∣ ≤ ∣ 1 t + h − 1 t ∣ = ∣ h ∣ t ( t + h ) ≤ 2 ∣ h ∣ t 2 . \Bigl|Q_{h}-\frac{1}{t}\Bigr|\le\Bigl|\frac{1}{t+h}-\frac{1}{t}\Bigr|=\frac{|h|}{t\,(t+h)}\le\frac{2|h|}{t^{2}}. Q h − t 1 ≤ t + h 1 − t 1 = t ( t + h ) ∣ h ∣ ≤ t 2 2∣ h ∣ .
Given a real ε > 0 \varepsilon>0 ε > 0 , put δ t = min ( t / 2 , ε t 2 / 2 ) \delta_{t}=\min(t/2,\varepsilon t^{2}/2) δ t = min ( t /2 , ε t 2 /2 ) ; every h h h with 0 < ∣ h ∣ < δ t 0<|h|<\delta_{t} 0 < ∣ h ∣ < δ t then satisfies t + h ∈ U t+h\in U t + h ∈ U and ∣ Q h − 1 / t ∣ < 2 δ t / t 2 ≤ ε |Q_{h}-1/t|<2\delta_{t}/t^{2}\le\varepsilon ∣ Q h − 1/ t ∣ < 2 δ t / t 2 ≤ ε , which is the condition of Derivative at an Interior Point with L = 1 / t L=1/t L = 1/ t . The identity map of U U U has derivative 1 1 1 (its difference quotients equal 1 1 1 ); so by claims 1 to 3 of Sum, Constant Multiple, and Product Rules for One-Dimensional Derivatives , f f f is differentiable at every t ∈ U t\in U t ∈ U with f ′ ( t ) = log t + 1 f'(t)=\log t+1 f ′ ( t ) = log t + 1 , and f ′ f' f ′ is differentiable at every t t t with derivative ω ( t ) = 1 / t \omega(t)=1/t ω ( t ) = 1/ t , which is differentiable at every t ∈ U t\in U t ∈ U by claim 2 of Reciprocal Rule for One-Dimensional Derivatives , 0 0 0 not lying in U U U . Two elementary observations transfer this to the notions used by C^k Maps on a Euclidean Open Set . First, if κ ^ : U → R \hat\kappa:U\to\mathbb{R} κ ^ : U → R is differentiable at t t t with derivative L L L in the sense of Derivative at an Interior Point , then the partial derivative of κ ^ \hat\kappa κ ^ with respect to the first variable exists at t t t with value L L L in the sense of Partial Derivative on a Euclidean Open Set : the two conditions concern the same difference quotient, and replacing δ \delta δ by min ( δ , t ) \min(\delta,t) min ( δ , t ) ensures t + h ∈ U t+h\in U t + h ∈ U whenever 0 < ∣ h ∣ < δ 0<|h|<\delta 0 < ∣ h ∣ < δ . Second, such a κ ^ \hat\kappa κ ^ is continuous at t t t in the sense of Continuity at a Point for Maps Between Euclidean Spaces : taking δ 1 ≤ t \delta_{1}\le t δ 1 ≤ t for the value 1 1 1 in the derivative condition, ∣ κ ^ ( t + h ) − κ ^ ( t ) ∣ ≤ ∣ h ∣ ( ∣ L ∣ + 1 ) |\hat\kappa(t+h)-\hat\kappa(t)|\le|h|(|L|+1) ∣ κ ^ ( t + h ) − κ ^ ( t ) ∣ ≤ ∣ h ∣ ( ∣ L ∣ + 1 ) for 0 < ∣ h ∣ < δ 1 0<|h|<\delta_{1} 0 < ∣ h ∣ < δ 1 , which is less than a given ε > 0 \varepsilon>0 ε > 0 when also ∣ h ∣ < ε / ( ∣ L ∣ + 1 ) |h|<\varepsilon/(|L|+1) ∣ h ∣ < ε / ( ∣ L ∣ + 1 ) . Hence f f f , ∂ 1 f = f ′ \partial_{1}f=f' ∂ 1 f = f ′ and ∂ 1 ∂ 1 f = ω \partial_{1}\partial_{1}f=\omega ∂ 1 ∂ 1 f = ω exist and are continuous at every point of U U U ; by clause 1 of C^k Maps on a Euclidean Open Set , f f f and ∂ 1 f \partial_{1}f ∂ 1 f are of class C 1 C^{1} C 1 on U U U , and by its clause 2 (with k = 1 k=1 k = 1 ), f f f is of class C 2 C^{2} C 2 on U U U . Moreover f ( 1 ) = 0 f(1)=0 f ( 1 ) = 0 , ∂ 1 f ( 1 ) = log 1 + 1 = 1 \partial_{1}f(1)=\log1+1=1 ∂ 1 f ( 1 ) = log 1 + 1 = 1 and ∂ 1 ∂ 1 f ( 1 ) = 1 \partial_{1}\partial_{1}f(1)=1 ∂ 1 ∂ 1 f ( 1 ) = 1 , using log 1 = 0 \log1=0 log 1 = 0 from Step 2. By Second-Order Taylor Expansion with Peano Remainder with n = 1 n=1 n = 1 and x = 1 x=1 x = 1 , where the Euclidean norm of h ∈ R 1 h\in\mathbb{R}^{1} h ∈ R 1 is ∣ h ∣ |h| ∣ h ∣ : for every real η > 0 \eta>0 η > 0 there is a real δ > 0 \delta>0 δ > 0 such that every real h h h with ∣ h ∣ < δ |h|<\delta ∣ h ∣ < δ satisfies 1 + h > 0 1+h>0 1 + h > 0 and
∣ ϕ ( 1 + h ) − h − 1 2 h 2 ∣ ≤ η h 2 . \Bigl|\phi(1+h)-h-\tfrac{1}{2}h^{2}\Bigr|\le\eta\,h^{2}. ϕ ( 1 + h ) − h − 2 1 h 2 ≤ η h 2 .
For h > − 1 h>-1 h > − 1 write R ( h ) = ϕ ( 1 + h ) − h − 1 2 h 2 R(h)=\phi(1+h)-h-\frac{1}{2}h^{2} R ( h ) = ϕ ( 1 + h ) − h − 2 1 h 2 .
Step 10 (Claim 4). Let F F F , ( n , ψ ) (n,\psi) ( n , ψ ) , B ≥ 1 B\ge1 B ≥ 1 and D D D be as in Step 6, and put m = ∫ X F d γ c m=\int_{X}F\,d\gamma_{c} m = ∫ X F d γ c , v = ∫ X F 2 d γ c v=\int_{X}F^{2}\,d\gamma_{c} v = ∫ X F 2 d γ c , m 3 = ∫ X F 3 d γ c m_{3}=\int_{X}F^{3}\,d\gamma_{c} m 3 = ∫ X F 3 d γ c and m 4 = ∫ X F 4 d γ c m_{4}=\int_{X}F^{4}\,d\gamma_{c} m 4 = ∫ X F 4 d γ c , integrals of bounded Borel functions (claim 3 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions ); by Linearity and Monotonicity of the Lebesgue Integral §integrable , ∣ m ∣ ≤ B |m|\le B ∣ m ∣ ≤ B , 0 ≤ v ≤ B 2 0\le v\le B^{2} 0 ≤ v ≤ B 2 , ∣ m 3 ∣ ≤ B 3 |m_{3}|\le B^{3} ∣ m 3 ∣ ≤ B 3 and m 4 ≥ 0 m_{4}\ge0 m 4 ≥ 0 . Let η > 0 \eta>0 η > 0 , let δ \delta δ be as in Step 9, and let s = min ( η , 1 / B , δ / ( 6 B ) ) s=\min(\eta,1/B,\delta/(6B)) s = min ( η , 1/ B , δ / ( 6 B )) , so that 0 < s ≤ η 0<s\le\eta 0 < s ≤ η , s B ≤ 1 sB\le1 s B ≤ 1 and s ≤ 1 s\le1 s ≤ 1 . Let G s = 1 + s F G_{s}=1+sF G s = 1 + s F . By Step 5 (with α = 1 \alpha=1 α = 1 , s s s , b 0 = 0 b_{0}=0 b 0 = 0 ), G s ∈ F C b 1 ( X ) G_{s}\in\mathcal{F}C^{1}_{b}(X) G s ∈ F C b 1 ( X ) with ∣ ∇ a G s ∣ a 2 = s 2 ∑ k = 1 n a k ( ∂ k F ) 2 = s 2 ∣ ∇ a F ∣ a 2 |\nabla_{a}G_{s}|_{a}^{2}=s^{2}\sum_{k=1}^{n}a_{k}(\partial_{k}F)^{2}=s^{2}|\nabla_{a}F|_{a}^{2} ∣ ∇ a G s ∣ a 2 = s 2 ∑ k = 1 n a k ( ∂ k F ) 2 = s 2 ∣ ∇ a F ∣ a 2 , so claim 3 applied to G s G_{s} G s gives, with Linearity and Monotonicity of the Lebesgue Integral §integrable ,
Ent γ c ( G s 2 ) ≤ 2 κ s 2 D . \operatorname{Ent}_{\gamma_{c}}(G_{s}^{2})\le2\kappa s^{2}D . Ent γ c ( G s 2 ) ≤ 2 κ s 2 D .
Now G s 2 = 1 + u G_{s}^{2}=1+u G s 2 = 1 + u with u = 2 s F + s 2 F 2 u=2sF+s^{2}F^{2} u = 2 s F + s 2 F 2 , and ∣ u ∣ ≤ 2 s B + s 2 B 2 ≤ 3 s B < δ |u|\le2sB+s^{2}B^{2}\le3sB<\delta ∣ u ∣ ≤ 2 s B + s 2 B 2 ≤ 3 s B < δ on X X X . Let w = ∫ X u d γ c = 2 s m + s 2 v w=\int_{X}u\,d\gamma_{c}=2sm+s^{2}v w = ∫ X u d γ c = 2 s m + s 2 v ; then ∣ w ∣ ≤ ∫ X ∣ u ∣ d γ c ≤ 3 s B < δ |w|\le\int_{X}|u|\,d\gamma_{c}\le3sB<\delta ∣ w ∣ ≤ ∫ X ∣ u ∣ d γ c ≤ 3 s B < δ and ∫ X G s 2 d γ c = 1 + w \int_{X}G_{s}^{2}\,d\gamma_{c}=1+w ∫ X G s 2 d γ c = 1 + w . By Step 9, ϕ ∘ G s 2 = u + 1 2 u 2 + R ∘ u \phi\circ G_{s}^{2}=u+\frac{1}{2}u^{2}+R\circ u ϕ ∘ G s 2 = u + 2 1 u 2 + R ∘ u on X X X , where R ∘ u R\circ u R ∘ u is integrable as a difference of integrable functions (Step 6) and satisfies R ∘ u ≥ − η u 2 R\circ u\ge-\eta u^{2} R ∘ u ≥ − η u 2 ; and ϕ ( 1 + w ) = w + 1 2 w 2 + R ( w ) \phi(1+w)=w+\frac{1}{2}w^{2}+R(w) ϕ ( 1 + w ) = w + 2 1 w 2 + R ( w ) with R ( w ) ≤ η w 2 R(w)\le\eta w^{2} R ( w ) ≤ η w 2 . By The Entropy of a Nonnegative Function with Respect to a Probability Measure §entropy and Linearity and Monotonicity of the Lebesgue Integral §integrable ,
Ent γ c ( G s 2 ) = 1 2 ( ∫ X u 2 d γ c − w 2 ) + ∫ X R ∘ u d γ c − R ( w ) ≥ 1 2 ( ∫ X u 2 d γ c − w 2 ) − η ( ∫ X u 2 d γ c + w 2 ) . \operatorname{Ent}_{\gamma_{c}}(G_{s}^{2})=\frac{1}{2}\Bigl(\int_{X}u^{2}\,d\gamma_{c}-w^{2}\Bigr)+\int_{X}R\circ u\,d\gamma_{c}-R(w)\ge\frac{1}{2}\Bigl(\int_{X}u^{2}\,d\gamma_{c}-w^{2}\Bigr)-\eta\Bigl(\int_{X}u^{2}\,d\gamma_{c}+w^{2}\Bigr). Ent γ c ( G s 2 ) = 2 1 ( ∫ X u 2 d γ c − w 2 ) + ∫ X R ∘ u d γ c − R ( w ) ≥ 2 1 ( ∫ X u 2 d γ c − w 2 ) − η ( ∫ X u 2 d γ c + w 2 ) .
Here ∫ X u 2 d γ c ≤ 9 s 2 B 2 \int_{X}u^{2}\,d\gamma_{c}\le9s^{2}B^{2} ∫ X u 2 d γ c ≤ 9 s 2 B 2 and w 2 ≤ 9 s 2 B 2 w^{2}\le9s^{2}B^{2} w 2 ≤ 9 s 2 B 2 , while expanding the squares,
1 2 ( ∫ X u 2 d γ c − w 2 ) = 2 s 2 ( v − m 2 ) + 2 s 3 ( m 3 − m v ) + s 4 2 ( m 4 − v 2 ) ≥ 2 s 2 ( v − m 2 ) − 4 s 3 B 3 − s 4 2 B 4 . \frac{1}{2}\Bigl(\int_{X}u^{2}\,d\gamma_{c}-w^{2}\Bigr)=2s^{2}(v-m^{2})+2s^{3}(m_{3}-mv)+\frac{s^{4}}{2}(m_{4}-v^{2})\ge2s^{2}(v-m^{2})-4s^{3}B^{3}-\frac{s^{4}}{2}B^{4}. 2 1 ( ∫ X u 2 d γ c − w 2 ) = 2 s 2 ( v − m 2 ) + 2 s 3 ( m 3 − m v ) + 2 s 4 ( m 4 − v 2 ) ≥ 2 s 2 ( v − m 2 ) − 4 s 3 B 3 − 2 s 4 B 4 .
Combining the last three displays and dividing by 2 s 2 > 0 2s^{2}>0 2 s 2 > 0 ,
v − m 2 ≤ κ D + 2 s B 3 + s 2 4 B 4 + 9 η B 2 ≤ κ D + η ( 2 B 3 + B 4 4 + 9 B 2 ) , v-m^{2}\le\kappa D+2sB^{3}+\frac{s^{2}}{4}B^{4}+9\eta B^{2}\le\kappa D+\eta\Bigl(2B^{3}+\frac{B^{4}}{4}+9B^{2}\Bigr), v − m 2 ≤ κ D + 2 s B 3 + 4 s 2 B 4 + 9 η B 2 ≤ κ D + η ( 2 B 3 + 4 B 4 + 9 B 2 ) ,
using s 2 ≤ s ≤ η s^{2}\le s\le\eta s 2 ≤ s ≤ η . Let C 0 = 2 B 3 + B 4 4 + 9 B 2 C_{0}=2B^{3}+\frac{B^{4}}{4}+9B^{2} C 0 = 2 B 3 + 4 B 4 + 9 B 2 be the bracket, which does not depend on η \eta η and is positive as B ≥ 1 B\ge1 B ≥ 1 . Given a positive real ε \varepsilon ε , the argument above with η = ε / C 0 \eta=\varepsilon/C_{0} η = ε / C 0 gives v − m 2 ≤ κ D + ε v-m^{2}\le\kappa D+\varepsilon v − m 2 ≤ κ D + ε ; as ε \varepsilon ε was arbitrary, Comparison of Real Numbers with Arbitrary Positive Slack §slack-above gives v − m 2 ≤ κ D v-m^{2}\le\kappa D v − m 2 ≤ κ D , which is claim 4.