Each result cited is universally quantified over the data in its own statement. Throughout, R n \mathbb{R}^{n} R n is open by claim 1 of Euclidean Space is Open in Itself, and C k C^k C k Maps are Continuous , and partial derivatives, the class C k C^{k} C k , gradients and Hessian matrices are those of Differential Calculus and Convexity on Euclidean Open Sets: Standing Notation §derivatives . Let δ n − 1 : R n → R n \delta_{n}^{-1}:\mathbb{R}^{n}\to\mathbb{R}^{n} δ n − 1 : R n → R n be the map u ↦ ( a 1 1 / 2 u 1 , … , a n 1 / 2 u n ) u\mapsto(a_{1}^{1/2}u_{1},\dots,a_{n}^{1/2}u_{n}) u ↦ ( a 1 1/2 u 1 , … , a n 1/2 u n ) ; it is the inverse of the map δ n \delta_{n} δ n of A Diagonal Gaussian Reference Measure on the Noise Wasserstein Space, Rescaled Heads and Gaussian Tails: Standing Notation §heads , since a k 1 / 2 a k − 1 / 2 = 1 a_{k}^{1/2}a_{k}^{-1/2}=1 a k 1/2 a k − 1/2 = 1 , and as r n = δ n ∘ p n r_{n}=\delta_{n}\circ p_{n} r n = δ n ∘ p n there, p n = δ n − 1 ∘ r n p_{n}=\delta_{n}^{-1}\circ r_{n} p n = δ n − 1 ∘ r n . Let h = g ∘ δ n − 1 h=g\circ\delta_{n}^{-1} h = g ∘ δ n − 1 , and for x ∈ X x\in X x ∈ X and k ∈ [ n ] k\in[n] k ∈ [ n ] let β k ( x ) = ∂ k g ( p n ( x ) ) \beta_{k}(x)=\partial_{k}g(p_{n}(x)) β k ( x ) = ∂ k g ( p n ( x )) .
Step 1 (Partial derivatives along δ n − 1 \delta_{n}^{-1} δ n − 1 ). Let f : R n → R f:\mathbb{R}^{n}\to\mathbb{R} f : R n → R , y ∈ R n y\in\mathbb{R}^{n} y ∈ R n and i ∈ [ n ] i\in[n] i ∈ [ n ] , and suppose that the partial derivative of f f f with respect to the i i i th variable exists at δ n − 1 ( y ) \delta_{n}^{-1}(y) δ n − 1 ( y ) with value L L L . For s ∈ R s\in\mathbb{R} s ∈ R , the point obtained from y y y by adding s s s to its i i i th coordinate is mapped by δ n − 1 \delta_{n}^{-1} δ n − 1 to the point obtained from δ n − 1 ( y ) \delta_{n}^{-1}(y) δ n − 1 ( y ) by adding a i 1 / 2 s a_{i}^{1/2}s a i 1/2 s to its i i i th coordinate. Given ε > 0 \varepsilon>0 ε > 0 , Partial Derivative on a Euclidean Open Set gives η > 0 \eta>0 η > 0 such that the difference quotient of f f f at δ n − 1 ( y ) \delta_{n}^{-1}(y) δ n − 1 ( y ) in the i i i th variable with increment σ \sigma σ differs from L L L by less than ε / a i 1 / 2 \varepsilon/a_{i}^{1/2} ε / a i 1/2 whenever 0 < ∣ σ ∣ < η 0<|\sigma|<\eta 0 < ∣ σ ∣ < η . If 0 < ∣ s ∣ < η / a i 1 / 2 0<|s|<\eta/a_{i}^{1/2} 0 < ∣ s ∣ < η / a i 1/2 , then σ = a i 1 / 2 s \sigma=a_{i}^{1/2}s σ = a i 1/2 s satisfies 0 < ∣ σ ∣ < η 0<|\sigma|<\eta 0 < ∣ σ ∣ < η , and the difference quotient of f ∘ δ n − 1 f\circ\delta_{n}^{-1} f ∘ δ n − 1 at y y y in the i i i th variable with increment s s s is a i 1 / 2 a_{i}^{1/2} a i 1/2 times the former one with increment σ \sigma σ , so it differs from a i 1 / 2 L a_{i}^{1/2}L a i 1/2 L by less than ε \varepsilon ε . Hence ∂ i ( f ∘ δ n − 1 ) ( y ) \partial_{i}(f\circ\delta_{n}^{-1})(y) ∂ i ( f ∘ δ n − 1 ) ( y ) exists and equals a i 1 / 2 ∂ i f ( δ n − 1 ( y ) ) a_{i}^{1/2}\,\partial_{i}f(\delta_{n}^{-1}(y)) a i 1/2 ∂ i f ( δ n − 1 ( y )) . Moreover δ n − 1 \delta_{n}^{-1} δ n − 1 is continuous at every point: with A = ∑ k = 1 n a k A=\sum_{k=1}^{n}a_{k} A = ∑ k = 1 n a k , each ( y k − y k ′ ) 2 (y_{k}-y'_{k})^{2} ( y k − y k ′ ) 2 is at most ∑ j = 1 n ( y j − y j ′ ) 2 \sum_{j=1}^{n}(y_{j}-y'_{j})^{2} ∑ j = 1 n ( y j − y j ′ ) 2 , so ∑ k = 1 n a k ( y k − y k ′ ) 2 ≤ A ∑ j = 1 n ( y j − y j ′ ) 2 \sum_{k=1}^{n}a_{k}(y_{k}-y'_{k})^{2}\le A\sum_{j=1}^{n}(y_{j}-y'_{j})^{2} ∑ k = 1 n a k ( y k − y k ′ ) 2 ≤ A ∑ j = 1 n ( y j − y j ′ ) 2 , which is less than ε 2 \varepsilon^{2} ε 2 when ∑ j ( y j − y j ′ ) 2 < ( ε / ( 1 + A ) ) 2 \sum_{j}(y_{j}-y'_{j})^{2}<(\varepsilon/(1+A))^{2} ∑ j ( y j − y j ′ ) 2 < ( ε / ( 1 + A ) ) 2 . So if f f f is continuous at every point, so is f ∘ δ n − 1 f\circ\delta_{n}^{-1} f ∘ δ n − 1 , by Composition of Continuous Euclidean Maps . Consequently, if f f f is of class C 1 C^{1} C 1 on R n \mathbb{R}^{n} R n (C^k Maps on a Euclidean Open Set , clauses 1 and 3), then so is f ∘ δ n − 1 f\circ\delta_{n}^{-1} f ∘ δ n − 1 , with ∂ i ( f ∘ δ n − 1 ) = a i 1 / 2 ( ∂ i f ) ∘ δ n − 1 \partial_{i}(f\circ\delta_{n}^{-1})=a_{i}^{1/2}\,(\partial_{i}f)\circ\delta_{n}^{-1} ∂ i ( f ∘ δ n − 1 ) = a i 1/2 ( ∂ i f ) ∘ δ n − 1 , a constant multiple of a function continuous at every point. Such a multiple is again continuous at every point: if φ : R n → R \varphi:\mathbb{R}^{n}\to\mathbb{R} φ : R n → R is continuous at y y y and ε > 0 \varepsilon>0 ε > 0 , choose δ > 0 \delta>0 δ > 0 for φ \varphi φ at y y y and ε / ( 1 + a i 1 / 2 ) \varepsilon/(1+a_{i}^{1/2}) ε / ( 1 + a i 1/2 ) in place of ε \varepsilon ε ; then ∑ j ( y j ′ − y j ) 2 < δ 2 \sum_{j}(y'_{j}-y_{j})^{2}<\delta^{2} ∑ j ( y j ′ − y j ) 2 < δ 2 gives ( a i 1 / 2 φ ( y ′ ) − a i 1 / 2 φ ( y ) ) 2 = a i ( φ ( y ′ ) − φ ( y ) ) 2 < a i ε 2 / ( 1 + a i 1 / 2 ) 2 < ε 2 \bigl(a_{i}^{1/2}\varphi(y')-a_{i}^{1/2}\varphi(y)\bigr)^{2}=a_{i}\bigl(\varphi(y')-\varphi(y)\bigr)^{2}<a_{i}\varepsilon^{2}/(1+a_{i}^{1/2})^{2}<\varepsilon^{2} ( a i 1/2 φ ( y ′ ) − a i 1/2 φ ( y ) ) 2 = a i ( φ ( y ′ ) − φ ( y ) ) 2 < a i ε 2 / ( 1 + a i 1/2 ) 2 < ε 2 .
Step 2 (h h h is of class C 2 C^{2} C 2 with Hessian M M M ). The function g g g is of class C 2 C^{2} C 2 , hence of class C 1 C^{1} C 1 by claim 2 of Euclidean Space is Open in Itself, and C k C^k C k Maps are Continuous , and each ∂ i g \partial_{i}g ∂ i g is of class C 1 C^{1} C 1 by clause 2 of C^k Maps on a Euclidean Open Set . By Step 1 with f = g f=g f = g , h h h is of class C 1 C^{1} C 1 with ∂ i h = a i 1 / 2 ( ∂ i g ) ∘ δ n − 1 \partial_{i}h=a_{i}^{1/2}(\partial_{i}g)\circ\delta_{n}^{-1} ∂ i h = a i 1/2 ( ∂ i g ) ∘ δ n − 1 . By Step 1 with f = ∂ i g f=\partial_{i}g f = ∂ i g , the function ( ∂ i g ) ∘ δ n − 1 (\partial_{i}g)\circ\delta_{n}^{-1} ( ∂ i g ) ∘ δ n − 1 is of class C 1 C^{1} C 1 with j j j th partial derivative a j 1 / 2 ( ∂ j ∂ i g ) ∘ δ n − 1 a_{j}^{1/2}(\partial_{j}\partial_{i}g)\circ\delta_{n}^{-1} a j 1/2 ( ∂ j ∂ i g ) ∘ δ n − 1 ; so by claims 3 and 1 of Constants, Coordinate Functions, Sums and Products of C k C^k C k Functions on a Euclidean Open Set , ∂ i h \partial_{i}h ∂ i h is of class C 1 C^{1} C 1 and
∂ j ∂ i h ( u ) = a i 1 / 2 a j 1 / 2 ∂ j ∂ i g ( δ n − 1 ( u ) ) ( u ∈ R n , i , j ∈ [ n ] ) . \partial_{j}\partial_{i}h(u)=a_{i}^{1/2}a_{j}^{1/2}\,\partial_{j}\partial_{i}g(\delta_{n}^{-1}(u))\qquad(u\in\mathbb{R}^{n},\ i,j\in[n]). ∂ j ∂ i h ( u ) = a i 1/2 a j 1/2 ∂ j ∂ i g ( δ n − 1 ( u )) ( u ∈ R n , i , j ∈ [ n ]) .
Hence h h h is of class C 2 C^{2} C 2 by clause 2 of C^k Maps on a Euclidean Open Set . By Hessian Matrix of a C^2 Function , the entry of D 2 h ( u ) D^{2}h(u) D 2 h ( u ) in row i i i and column j j j is ∂ i ∂ j h ( u ) = a i 1 / 2 a j 1 / 2 ∂ i ∂ j g ( δ n − 1 ( u ) ) \partial_{i}\partial_{j}h(u)=a_{i}^{1/2}a_{j}^{1/2}\partial_{i}\partial_{j}g(\delta_{n}^{-1}(u)) ∂ i ∂ j h ( u ) = a i 1/2 a j 1/2 ∂ i ∂ j g ( δ n − 1 ( u )) , which is at most b b b in absolute value by the hypothesis on b b b (with the indices interchanged). For x ∈ X x\in X x ∈ X and u = r n ( x ) u=r_{n}(x) u = r n ( x ) we have δ n − 1 ( u ) = p n ( x ) \delta_{n}^{-1}(u)=p_{n}(x) δ n − 1 ( u ) = p n ( x ) , so the entry of D 2 h ( r n ( x ) ) D^{2}h(r_{n}(x)) D 2 h ( r n ( x )) in row i i i and column j j j is the entry of M ( x ) M(x) M ( x ) in row j j j and column i i i ; since D 2 h ( r n ( x ) ) D^{2}h(r_{n}(x)) D 2 h ( r n ( x )) is symmetric by claim 2 of Equality of Mixed Second Partial Derivatives and Symmetry of the Hessian , it follows that D 2 h ( r n ( x ) ) = M ( x ) D^{2}h(r_{n}(x))=M(x) D 2 h ( r n ( x )) = M ( x ) . Also ∂ i h ( r n ( x ) ) = a i 1 / 2 β i ( x ) \partial_{i}h(r_{n}(x))=a_{i}^{1/2}\beta_{i}(x) ∂ i h ( r n ( x )) = a i 1/2 β i ( x ) .
Step 3 (The potential). Let q ( u ) = 1 2 ∥ u ∥ 2 = 1 2 ∑ l = 1 n π l ( u ) 2 q(u)=\tfrac12\lVert u\rVert^{2}=\tfrac12\sum_{l=1}^{n}\pi_{l}(u)^{2} q ( u ) = 2 1 ∥ u ∥ 2 = 2 1 ∑ l = 1 n π l ( u ) 2 (Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n §square ), with the coordinate functions π l ( u ) = u l \pi_{l}(u)=u_{l} π l ( u ) = u l . By claims 2 and 3 of Constants, Coordinate Functions, Sums and Products of C k C^k C k Functions on a Euclidean Open Set , q q q is smooth, hence of class C 2 C^{2} C 2 (Smooth Map on a Euclidean Open Set ). Directly from Partial Derivative on a Euclidean Open Set , ∂ i π l \partial_{i}\pi_{l} ∂ i π l is the constant 1 1 1 if l = i l=i l = i and 0 0 0 otherwise, the difference quotients being constant; so claim 1 of Constants, Coordinate Functions, Sums and Products of C k C^k C k Functions on a Euclidean Open Set gives ∂ i q = π i \partial_{i}q=\pi_{i} ∂ i q = π i and ∂ j ∂ i q \partial_{j}\partial_{i}q ∂ j ∂ i q equal to 1 1 1 if j = i j=i j = i and 0 0 0 otherwise, that is, D 2 q ( u ) = I n D^{2}q(u)=I_{n} D 2 q ( u ) = I n . For t ∈ R t\in\mathbb{R} t ∈ R let Φ t = q + t h \Phi_{t}=q+t\,h Φ t = q + t h . By claims 3 and 1 of Constants, Coordinate Functions, Sums and Products of C k C^k C k Functions on a Euclidean Open Set (applied to q , h q,h q , h and to their first partial derivatives), Φ t \Phi_{t} Φ t is of class C 2 C^{2} C 2 , ∂ i Φ t ( u ) = u i + t ∂ i h ( u ) \partial_{i}\Phi_{t}(u)=u_{i}+t\,\partial_{i}h(u) ∂ i Φ t ( u ) = u i + t ∂ i h ( u ) , and D 2 Φ t ( u ) = I n + t D 2 h ( u ) D^{2}\Phi_{t}(u)=I_{n}+t\,D^{2}h(u) D 2 Φ t ( u ) = I n + t D 2 h ( u ) for every u u u .
Step 4 (Pinching). Let t ∈ R t\in\mathbb{R} t ∈ R satisfy 2 n ∣ t ∣ b ≤ 1 2n|t|b\le1 2 n ∣ t ∣ b ≤ 1 , and let u , z ∈ R n u,z\in\mathbb{R}^{n} u , z ∈ R n and A = D 2 h ( u ) A=D^{2}h(u) A = D 2 h ( u ) . By Step 2 and 2 ∣ z i ∣ ∣ z j ∣ ≤ z i 2 + z j 2 2|z_{i}||z_{j}|\le z_{i}^{2}+z_{j}^{2} 2∣ z i ∣∣ z j ∣ ≤ z i 2 + z j 2 ,
∣ z ⋅ ( A z ) ∣ = ∣ ∑ i , j = 1 n z i A i j z j ∣ ≤ b ∑ i , j = 1 n ∣ z i ∣ ∣ z j ∣ ≤ b 2 ∑ i , j = 1 n ( z i 2 + z j 2 ) = n b ∥ z ∥ 2 , |z\cdot(Az)|=\Bigl|\sum_{i,j=1}^{n}z_{i}A_{ij}z_{j}\Bigr|\le b\sum_{i,j=1}^{n}|z_{i}||z_{j}|\le\frac{b}{2}\sum_{i,j=1}^{n}(z_{i}^{2}+z_{j}^{2})=n\,b\,\lVert z\rVert^{2}, ∣ z ⋅ ( A z ) ∣ = i , j = 1 ∑ n z i A ij z j ≤ b i , j = 1 ∑ n ∣ z i ∣∣ z j ∣ ≤ 2 b i , j = 1 ∑ n ( z i 2 + z j 2 ) = n b ∥ z ∥ 2 ,
using Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n §square . Since z ⋅ ( D 2 Φ t ( u ) z ) = z ⋅ z + t z ⋅ ( A z ) = ∥ z ∥ 2 + t z ⋅ ( A z ) z\cdot(D^{2}\Phi_{t}(u)z)=z\cdot z+t\,z\cdot(Az)=\lVert z\rVert^{2}+t\,z\cdot(Az) z ⋅ ( D 2 Φ t ( u ) z ) = z ⋅ z + t z ⋅ ( A z ) = ∥ z ∥ 2 + t z ⋅ ( A z ) and n ∣ t ∣ b ≤ 1 2 n|t|b\le\tfrac12 n ∣ t ∣ b ≤ 2 1 , we get 1 2 ∥ z ∥ 2 ≤ z ⋅ ( D 2 Φ t ( u ) z ) ≤ 3 2 ∥ z ∥ 2 \tfrac12\lVert z\rVert^{2}\le z\cdot(D^{2}\Phi_{t}(u)z)\le\tfrac32\lVert z\rVert^{2} 2 1 ∥ z ∥ 2 ≤ z ⋅ ( D 2 Φ t ( u ) z ) ≤ 2 3 ∥ z ∥ 2 , that is, z ⋅ ( ( 1 2 I n ) z ) ≤ z ⋅ ( D 2 Φ t ( u ) z ) ≤ z ⋅ ( ( 3 2 I n ) z ) z\cdot((\tfrac12I_{n})z)\le z\cdot(D^{2}\Phi_{t}(u)z)\le z\cdot((\tfrac32I_{n})z) z ⋅ (( 2 1 I n ) z ) ≤ z ⋅ ( D 2 Φ t ( u ) z ) ≤ z ⋅ (( 2 3 I n ) z ) . The matrix D 2 Φ t ( u ) D^{2}\Phi_{t}(u) D 2 Φ t ( u ) lies in S ( n ) \mathcal{S}(n) S ( n ) by Differential Calculus and Convexity on Euclidean Open Sets: Standing Notation §derivatives , so by The Positive Semidefinite Ordering on Symmetric Matrices
1 2 I n ⪯ D 2 Φ t ( u ) ⪯ 3 2 I n ( u ∈ R n ) . \tfrac12I_{n}\preceq D^{2}\Phi_{t}(u)\preceq\tfrac32I_{n}\qquad(u\in\mathbb{R}^{n}). 2 1 I n ⪯ D 2 Φ t ( u ) ⪯ 2 3 I n ( u ∈ R n ) .
Step 5 (The head map). Keep t t t as in Step 4 and let F t = ∇ Φ t F_{t}=\nabla\Phi_{t} F t = ∇ Φ t , the gradient map u ↦ D Φ t ( u ) u\mapsto D\Phi_{t}(u) u ↦ D Φ t ( u ) , so that by Steps 2 and 3 its k k k th coordinate is ( F t ) k ( u ) = u k + t a k 1 / 2 ∂ k g ( δ n − 1 ( u ) ) (F_{t})_{k}(u)=u_{k}+t\,a_{k}^{1/2}\partial_{k}g(\delta_{n}^{-1}(u)) ( F t ) k ( u ) = u k + t a k 1/2 ∂ k g ( δ n − 1 ( u )) . By The Gradient of a Twice Continuously Differentiable Function with Hessian Pinched between Two Positive Multiples of the Identity is a Bi-Lipschitz Bijection of Euclidean Space with Continuously Differentiable Inverse with d = n d=n d = n , Φ = Φ t \Phi=\Phi_{t} Φ = Φ t , ε = 1 2 \varepsilon=\tfrac12 ε = 2 1 and L = 3 2 L=\tfrac32 L = 2 3 , which applies by Step 4: F t F_{t} F t has components of class C 1 C^{1} C 1 (as recorded in its preamble) and is a bijection of R n \mathbb{R}^{n} R n (The Gradient of a Twice Continuously Differentiable Function with Hessian Pinched between Two Positive Multiples of the Identity is a Bi-Lipschitz Bijection of Euclidean Space with Continuously Differentiable Inverse §bijection ); and by The Gradient of a Twice Continuously Differentiable Function with Hessian Pinched between Two Positive Multiples of the Identity is a Bi-Lipschitz Bijection of Euclidean Space with Continuously Differentiable Inverse §inverse , for every u u u the matrix D F t ( u ) = D 2 Φ t ( u ) DF_{t}(u)=D^{2}\Phi_{t}(u) D F t ( u ) = D 2 Φ t ( u ) is symmetric and positive definite, the inverse map F t − 1 F_{t}^{-1} F t − 1 has components of class C 1 C^{1} C 1 , and D ( F t − 1 ) ( F t ( u ) ) D(F_{t}^{-1})(F_{t}(u)) D ( F t − 1 ) ( F t ( u )) is the inverse matrix of D F t ( u ) DF_{t}(u) D F t ( u ) . So F t F_{t} F t satisfies the hypotheses imposed on F F F in Moving the Rescaled Head of the Diagonal Gaussian Measure by a Diffeomorphism: the Image Has an Explicit Positive Density Depending on the Head Only , and 0 < det D F t ( u ) 0<\det DF_{t}(u) 0 < det D F t ( u ) for every u u u , as recorded there.
Step 6 (The perturbation is T F t T_{F_{t}} T F t ). For x ∈ X x\in X x ∈ X and u = r n ( x ) u=r_{n}(x) u = r n ( x ) , Step 5 and δ n − 1 ( u ) = p n ( x ) \delta_{n}^{-1}(u)=p_{n}(x) δ n − 1 ( u ) = p n ( x ) give ( F t ( u ) − u ) k = t a k 1 / 2 β k ( x ) (F_{t}(u)-u)_{k}=t\,a_{k}^{1/2}\beta_{k}(x) ( F t ( u ) − u ) k = t a k 1/2 β k ( x ) , so by the definition of the lift in Lifting Euclidean Tangent Fields of the Rescaled Head to Noise Tangent Fields on a Hilbert Space and a k 1 / 2 a k 1 / 2 = a k a_{k}^{1/2}a_{k}^{1/2}=a_{k} a k 1/2 a k 1/2 = a k ,
Λ n ( F t − i d ) ( x ) = ∑ k = 1 n a k 1 / 2 t a k 1 / 2 β k ( x ) e k = t ∑ k = 1 n a k ∂ k g ( p n ( x ) ) e k = t ∇ a ψ ( x ) , \Lambda_{n}(F_{t}-\mathrm{id})(x)=\sum_{k=1}^{n}a_{k}^{1/2}\,t\,a_{k}^{1/2}\beta_{k}(x)\,e_{k}=t\sum_{k=1}^{n}a_{k}\,\partial_{k}g(p_{n}(x))\,e_{k}=t\,\nabla_{a}\psi(x), Λ n ( F t − id ) ( x ) = k = 1 ∑ n a k 1/2 t a k 1/2 β k ( x ) e k = t k = 1 ∑ n a k ∂ k g ( p n ( x )) e k = t ∇ a ψ ( x ) ,
the last equality being the formula for ∇ a ψ \nabla_{a}\psi ∇ a ψ recorded in the statement. Hence T F t = i d + t ∇ a ψ T_{F_{t}}=\mathrm{id}+t\nabla_{a}\psi T F t = id + t ∇ a ψ . By Moving the Rescaled Head of the Diagonal Gaussian Measure by a Diffeomorphism: the Image Has an Explicit Positive Density Depending on the Head Only §bijection , this map is a bijection of X X X , it and its inverse T F t − 1 T_{F_{t}^{-1}} T F t − 1 are Borel, and r n ( T F t ( x ) ) = F t ( r n ( x ) ) r_{n}(T_{F_{t}}(x))=F_{t}(r_{n}(x)) r n ( T F t ( x )) = F t ( r n ( x )) for every x ∈ X x\in X x ∈ X .
Step 7 (The log-determinant). For x ∈ X x\in X x ∈ X , Steps 3, 2 and 5 give D F t ( r n ( x ) ) = D 2 Φ t ( r n ( x ) ) = I n + t M ( x ) DF_{t}(r_{n}(x))=D^{2}\Phi_{t}(r_{n}(x))=I_{n}+t\,M(x) D F t ( r n ( x )) = D 2 Φ t ( r n ( x )) = I n + t M ( x ) , so 0 < det ( I n + t M ( x ) ) 0<\det(I_{n}+tM(x)) 0 < det ( I n + tM ( x )) by Step 5. By Determinants of Positive Definite Matrices: Positivity, the Bound log det A ≤ t r A − d \log\det A\le\mathrm{tr}\,A-d log det A ≤ tr A − d , Bounds under Pinching, and the Expansion of det ( I + t B ) \det(I+tB) det ( I + tB ) §pinching with d = n d=n d = n , ε = 1 2 \varepsilon=\tfrac12 ε = 2 1 and L = 3 2 L=\tfrac32 L = 2 3 , which applies by Step 4, n − 2 n ≤ log det ( I n + t M ( x ) ) ≤ 3 2 n − n n-2n\le\log\det(I_{n}+tM(x))\le\tfrac32n-n n − 2 n ≤ log det ( I n + tM ( x )) ≤ 2 3 n − n ; so ∣ log det ( I n + t M ( x ) ) ∣ ≤ n |\log\det(I_{n}+tM(x))|\le n ∣ log det ( I n + tM ( x )) ∣ ≤ n . The function u ↦ det D F t ( u ) u\mapsto\det DF_{t}(u) u ↦ det D F t ( u ) is continuous on R n \mathbb{R}^{n} R n by Change of Variables for the Lebesgue Integral under a Continuously Differentiable Bijection of Euclidean Space with Symmetric Positive Definite Jacobian Matrix, and the Density of a Push-Forward §regularity , and it takes values in ( 0 , ∞ ) (0,\infty) ( 0 , ∞ ) by Step 5; log \log log is smooth on the open set ( 0 , ∞ ) (0,\infty) ( 0 , ∞ ) by The Natural Logarithm , hence continuous relative to ( 0 , ∞ ) (0,\infty) ( 0 , ∞ ) at every point of ( 0 , ∞ ) (0,\infty) ( 0 , ∞ ) by claim 3 of Euclidean Space is Open in Itself, and C k C^k C k Maps are Continuous , the Euclidean distance on the real line being the absolute-value metric by The Euclidean Distance on the Real Line is the Absolute Value Metric ; and r n r_{n} r n is continuous by Lifting Euclidean Tangent Fields of the Rescaled Head to Noise Tangent Fields on a Hilbert Space §head . So x ↦ log det ( I n + t M ( x ) ) = log det D F t ( r n ( x ) ) x\mapsto\log\det(I_{n}+tM(x))=\log\det DF_{t}(r_{n}(x)) x ↦ log det ( I n + tM ( x )) = log det D F t ( r n ( x )) is continuous, as a composite of continuous maps between metric spaces, the inner ones taking values in the sets on which the outer ones are continuous (two applications of Continuous Map Between Metric Spaces ), hence Borel by claim 3 of Borel Measurability and Bounded Integration on a Metric Space , and bounded.
Step 8 (Integrability). Let k ∈ [ n ] k\in[n] k ∈ [ n ] . The function ∂ k g \partial_{k}g ∂ k g is of class C 1 C^{1} C 1 by clause 2 of C^k Maps on a Euclidean Open Set and is bounded with bounded partial derivatives by Bounded Twice Continuously Differentiable Functions with Bounded First and Second Partial Derivatives on Euclidean Space §bounded , so it lies in C b 1 ( R n ) C^{1}_{b}(\mathbb{R}^{n}) C b 1 ( R n ) (Bounded Continuously Differentiable Functions with Bounded Partial Derivatives on Euclidean Space §bounded ) and β k = ( ∂ k g ) ∘ p n ∈ F C b 1 ( X ) \beta_{k}=(\partial_{k}g)\circ p_{n}\in\mathcal{F}C^{1}_{b}(X) β k = ( ∂ k g ) ∘ p n ∈ F C b 1 ( X ) by Bounded C^1 Cylindrical Functions on a Hilbert Space with an Orthonormal Basis §cylindrical . Since μ ∈ P 2 ( X ) \mu\in\mathcal{P}_{2}(X) μ ∈ P 2 ( X ) , as recorded in the statement, the function x ↦ x k β k ( x ) x\mapsto x_{k}\beta_{k}(x) x ↦ x k β k ( x ) is integrable with respect to μ \mu μ by Bounded C^1 Cylindrical Functions: Linear Structure, the Gradient and the Partial Derivatives, Bounds and Integrability, and Density in the Square-Integrable Functions §coordinate-integrable . The function ∂ k g \partial_{k}g ∂ k g , being of class C 1 C^{1} C 1 , is continuous on R n \mathbb{R}^{n} R n as a map from ( R n , d E ) (\mathbb{R}^{n},d_{E}) ( R n , d E ) into ( R , d R ) (\mathbb{R},d_{\mathbb{R}}) ( R , d R ) , where d R ( s , s ′ ) = ∣ s − s ′ ∣ d_{\mathbb{R}}(s,s')=|s-s'| d R ( s , s ′ ) = ∣ s − s ′ ∣ , by claim 3 of Euclidean Space is Open in Itself, and C k C^k C k Maps are Continuous . The function ∂ k ∂ k g \partial_{k}\partial_{k}g ∂ k ∂ k g is continuous at every point in the Euclidean sense by clause 1 of C^k Maps on a Euclidean Open Set , applied to the function ∂ k g \partial_{k}g ∂ k g of class C 1 C^{1} C 1 (g g g itself is only of class C 2 C^{2} C 2 , so claim 3 of Euclidean Space is Open in Itself, and C k C^k C k Maps are Continuous is not applied to ∂ k ∂ k g \partial_{k}\partial_{k}g ∂ k ∂ k g ). Since d E ( y , y ′ ) = ∥ y − y ′ ∥ d_{E}(y,y')=\lVert y-y'\rVert d E ( y , y ′ ) = ∥ y − y ′ ∥ is nonnegative with d E ( y , y ′ ) 2 = ∑ j = 1 n ( y j − y j ′ ) 2 d_{E}(y,y')^{2}=\sum_{j=1}^{n}(y_{j}-y'_{j})^{2} d E ( y , y ′ ) 2 = ∑ j = 1 n ( y j − y j ′ ) 2 by Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n §distance and Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n §square , the condition ∑ j ( y j − y j ′ ) 2 < δ 2 \sum_{j}(y_{j}-y'_{j})^{2}<\delta^{2} ∑ j ( y j − y j ′ ) 2 < δ 2 is equivalent to d E ( y , y ′ ) < δ d_{E}(y,y')<\delta d E ( y , y ′ ) < δ , and for real s , s ′ s,s' s , s ′ the condition ( s − s ′ ) 2 < ε 2 (s-s')^{2}<\varepsilon^{2} ( s − s ′ ) 2 < ε 2 is equivalent to ∣ s − s ′ ∣ < ε |s-s'|<\varepsilon ∣ s − s ′ ∣ < ε ; so ∂ k ∂ k g \partial_{k}\partial_{k}g ∂ k ∂ k g is continuous on R n \mathbb{R}^{n} R n as a map from ( R n , d E ) (\mathbb{R}^{n},d_{E}) ( R n , d E ) into ( R , d R ) (\mathbb{R},d_{\mathbb{R}}) ( R , d R ) . The map p n : ( X , d ) → ( R n , d E ) p_{n}:(X,d)\to(\mathbb{R}^{n},d_{E}) p n : ( X , d ) → ( R n , d E ) is continuous by Borel Sets of a Hilbert Space with an Orthonormal Basis: Coordinates, Determination by Finite-Dimensional Projections, and Pairs §continuity , being Lipschitz there (A Lipschitz Map is Uniformly Continuous ); so β k \beta_{k} β k , β k 2 \beta_{k}^{2} β k 2 (by Continuity of Sums and Products of Real-Valued Functions on a Metric Space §on-set ) and x ↦ ∂ k ∂ k g ( p n ( x ) ) x\mapsto\partial_{k}\partial_{k}g(p_{n}(x)) x ↦ ∂ k ∂ k g ( p n ( x )) are continuous, as composites of continuous maps between metric spaces, hence Borel by claim 3 of Borel Measurability and Bounded Integration on a Metric Space , and they are bounded, by B B B , B 2 B^{2} B 2 and a bound for ∂ k ∂ k g \partial_{k}\partial_{k}g ∂ k ∂ k g respectively. By claim 6(b) of Borel Measurability and Bounded Integration on a Metric Space they, and the bounded Borel function of Step 7, are integrable with respect to the probability measure μ \mu μ . Hence, by Linearity and Monotonicity of the Lebesgue Integral §integrable , the function
Θ t ( x ) = ∑ k = 1 n ( t a k x k c k β k ( x ) + t 2 a k 2 2 c k β k ( x ) 2 ) − log det ( I n + t M ( x ) ) \Theta_{t}(x)=\sum_{k=1}^{n}\Bigl(\frac{t\,a_{k}\,x_{k}}{c_{k}}\,\beta_{k}(x)+\frac{t^{2}a_{k}^{2}}{2c_{k}}\,\beta_{k}(x)^{2}\Bigr)-\log\det\bigl(I_{n}+tM(x)\bigr) Θ t ( x ) = k = 1 ∑ n ( c k t a k x k β k ( x ) + 2 c k t 2 a k 2 β k ( x ) 2 ) − log det ( I n + tM ( x ) )
is integrable with respect to μ \mu μ , and so is each summand.
Step 9 (The logarithm of the density along the map). Let G = G F t G=G_{F_{t}} G = G F t be the function of Moving the Rescaled Head of the Diagonal Gaussian Measure by a Diffeomorphism: the Image Has an Explicit Positive Density Depending on the Head Only for F = F t F=F_{t} F = F t . By Moving the Rescaled Head of the Diagonal Gaussian Measure by a Diffeomorphism: the Image Has an Explicit Positive Density Depending on the Head Only §density , G G G is Borel and positive, and G ∘ r n G\circ r_{n} G ∘ r n , which is Borel by claim 4 of Borel Measurability and Bounded Integration on a Metric Space and positive, is a density of ( T F t ) # γ c (T_{F_{t}})_{\#}\gamma_{c} ( T F t ) # γ c with respect to γ c \gamma_{c} γ c . Let x ∈ X x\in X x ∈ X and u = r n ( x ) u=r_{n}(x) u = r n ( x ) , so that a k 1 / 2 u k = x k a_{k}^{1/2}u_{k}=x_{k} a k 1/2 u k = x k and F t ( u ) k = u k + t a k 1 / 2 β k ( x ) F_{t}(u)_{k}=u_{k}+t\,a_{k}^{1/2}\beta_{k}(x) F t ( u ) k = u k + t a k 1/2 β k ( x ) . By Step 6 and Moving the Rescaled Head of the Diagonal Gaussian Measure by a Diffeomorphism: the Image Has an Explicit Positive Density Depending on the Head Only §log-density ,
log G ( r n ( T F t ( x ) ) ) = log G ( F t ( u ) ) = 1 2 ∣ F t ( u ) ∣ c ~ ( n ) 2 − 1 2 ∣ u ∣ c ~ ( n ) 2 − log det D F t ( u ) . \log G\bigl(r_{n}(T_{F_{t}}(x))\bigr)=\log G(F_{t}(u))=\tfrac12\,|F_{t}(u)|^{2}_{\tilde{c}^{(n)}}-\tfrac12\,|u|^{2}_{\tilde{c}^{(n)}}-\log\det DF_{t}(u). log G ( r n ( T F t ( x )) ) = log G ( F t ( u )) = 2 1 ∣ F t ( u ) ∣ c ~ ( n ) 2 − 2 1 ∣ u ∣ c ~ ( n ) 2 − log det D F t ( u ) .
By The Diagonal Gaussian Density on Euclidean Space and Its Notation §scaling and c ~ k ( n ) = c k / a k \tilde{c}^{(n)}_{k}=c_{k}/a_{k} c ~ k ( n ) = c k / a k (A Diagonal Gaussian Reference Measure on the Noise Wasserstein Space, Rescaled Heads and Gaussian Tails: Standing Notation §heads ), ∣ v ∣ c ~ ( n ) 2 = ∑ k = 1 n a k v k 2 / c k |v|^{2}_{\tilde{c}^{(n)}}=\sum_{k=1}^{n}a_{k}v_{k}^{2}/c_{k} ∣ v ∣ c ~ ( n ) 2 = ∑ k = 1 n a k v k 2 / c k for v ∈ R n v\in\mathbb{R}^{n} v ∈ R n , and for each k k k
a k 2 c k ( ( u k + t a k 1 / 2 β k ( x ) ) 2 − u k 2 ) = a k c k t a k 1 / 2 u k β k ( x ) + t 2 a k 2 2 c k β k ( x ) 2 = t a k x k c k β k ( x ) + t 2 a k 2 2 c k β k ( x ) 2 . \frac{a_{k}}{2c_{k}}\Bigl(\bigl(u_{k}+t\,a_{k}^{1/2}\beta_{k}(x)\bigr)^{2}-u_{k}^{2}\Bigr)=\frac{a_{k}}{c_{k}}\,t\,a_{k}^{1/2}u_{k}\,\beta_{k}(x)+\frac{t^{2}a_{k}^{2}}{2c_{k}}\,\beta_{k}(x)^{2}=\frac{t\,a_{k}\,x_{k}}{c_{k}}\,\beta_{k}(x)+\frac{t^{2}a_{k}^{2}}{2c_{k}}\,\beta_{k}(x)^{2}. 2 c k a k ( ( u k + t a k 1/2 β k ( x ) ) 2 − u k 2 ) = c k a k t a k 1/2 u k β k ( x ) + 2 c k t 2 a k 2 β k ( x ) 2 = c k t a k x k β k ( x ) + 2 c k t 2 a k 2 β k ( x ) 2 .
Together with det D F t ( u ) = det ( I n + t M ( x ) ) \det DF_{t}(u)=\det(I_{n}+tM(x)) det D F t ( u ) = det ( I n + tM ( x )) (Step 7), this gives log G ( r n ( T F t ( x ) ) ) = Θ t ( x ) \log G\bigl(r_{n}(T_{F_{t}}(x))\bigr)=\Theta_{t}(x) log G ( r n ( T F t ( x )) ) = Θ t ( x ) for every x ∈ X x\in X x ∈ X .
Step 10 (Claim 1). Apply Relative Entropy of the Image under a Bimeasurable Bijection that Moves the Reference Measure by a Positive Density §entropy on ( S , S ) = ( X , B ( X ) ) (S,\mathcal{S})=(X,\mathcal{B}(X)) ( S , S ) = ( X , B ( X )) with γ = γ c \gamma=\gamma_{c} γ = γ c , ν = μ \nu=\mu ν = μ , T = T F t T=T_{F_{t}} T = T F t and the density G ∘ r n G\circ r_{n} G ∘ r n : T T T is a bijection with T T T and T − 1 T^{-1} T − 1 Borel (Step 6); G ∘ r n G\circ r_{n} G ∘ r n is measurable, positive and a density of T # γ c T_{\#}\gamma_{c} T # γ c with respect to γ c \gamma_{c} γ c (Step 9); μ \mu μ has finite relative entropy with respect to γ c \gamma_{c} γ c ; and s ↦ log ( ( G ∘ r n ) ( T ( s ) ) ) = Θ t ( s ) s\mapsto\log\bigl((G\circ r_{n})(T(s))\bigr)=\Theta_{t}(s) s ↦ log ( ( G ∘ r n ) ( T ( s )) ) = Θ t ( s ) is integrable with respect to μ \mu μ (Steps 8 and 9). The image measure T # μ T_{\#}\mu T # μ there is the push-forward ( i d + t ∇ a ψ ) # μ (\mathrm{id}+t\nabla_{a}\psi)_{\#}\mu ( id + t ∇ a ψ ) # μ of Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §pushforward , by Step 6. Hence ( i d + t ∇ a ψ ) # μ (\mathrm{id}+t\nabla_{a}\psi)_{\#}\mu ( id + t ∇ a ψ ) # μ has finite relative entropy with respect to γ c \gamma_{c} γ c and
H ( ( i d + t ∇ a ψ ) # μ ∣ γ c ) = H ( μ ∣ γ c ) + ∫ X Θ t d μ , H\bigl((\mathrm{id}+t\nabla_{a}\psi)_{\#}\mu\,\big|\,\gamma_{c}\bigr)=H(\mu\,|\,\gamma_{c})+\int_{X}\Theta_{t}\,d\mu , H ( ( id + t ∇ a ψ ) # μ γ c ) = H ( μ ∣ γ c ) + ∫ X Θ t d μ ,
and splitting ∫ X Θ t d μ \int_{X}\Theta_{t}\,d\mu ∫ X Θ t d μ into the integrals of its summands by Linearity and Monotonicity of the Lebesgue Integral §integrable (Step 8) gives the displayed formula of claim 1. The remaining assertions of claim 1 are Step 7 (positivity of the determinant; the log-determinant is Borel and bounded) and Step 6 (the map i d + t ∇ a ψ \mathrm{id}+t\nabla_{a}\psi id + t ∇ a ψ is Borel).
Step 11 (Claim 2). Let t t t satisfy 2 n ∣ t ∣ b ≤ 1 2n|t|b\le1 2 n ∣ t ∣ b ≤ 1 and ∣ t ∣ b ≤ θ n |t|b\le\theta_{n} ∣ t ∣ b ≤ θ n . By The Noise Ornstein-Uhlenbeck Functional of a Probability Measure at a Bounded C^2 Function of Finitely Many Coordinates §functional and Linearity and Monotonicity of the Lebesgue Integral §integrable (with the integrable functions of Step 8),
t L μ a ( g ) = ∑ k = 1 n ∫ X t a k x k c k β k ( x ) μ ( d x ) − ∫ X t t r M ( x ) μ ( d x ) , t\,L^{a}_{\mu}(g)=\sum_{k=1}^{n}\int_{X}\frac{t\,a_{k}\,x_{k}}{c_{k}}\,\beta_{k}(x)\,\mu(dx)-\int_{X}t\,\mathrm{tr}\,M(x)\,\mu(dx), t L μ a ( g ) = k = 1 ∑ n ∫ X c k t a k x k β k ( x ) μ ( d x ) − ∫ X t tr M ( x ) μ ( d x ) ,
since ∑ k = 1 n a k ∂ k ∂ k g ( p n ( x ) ) = ∑ k = 1 n M ( x ) k k = t r M ( x ) \sum_{k=1}^{n}a_{k}\,\partial_{k}\partial_{k}g(p_{n}(x))=\sum_{k=1}^{n}M(x)_{kk}=\mathrm{tr}\,M(x) ∑ k = 1 n a k ∂ k ∂ k g ( p n ( x )) = ∑ k = 1 n M ( x ) kk = tr M ( x ) by Trace of a Real Square Matrix and a k 1 / 2 a k 1 / 2 = a k a_{k}^{1/2}a_{k}^{1/2}=a_{k} a k 1/2 a k 1/2 = a k . Subtracting this from the formula of claim 1, again by linearity of the integral,
H ( ( i d + t ∇ a ψ ) # μ ∣ γ c ) − H ( μ ∣ γ c ) − t L μ a ( g ) = ∑ k = 1 n ∫ X t 2 a k 2 2 c k β k 2 d μ − ∫ X ( log det ( I n + t M ( x ) ) − t t r M ( x ) ) μ ( d x ) . H\bigl((\mathrm{id}+t\nabla_{a}\psi)_{\#}\mu\,\big|\,\gamma_{c}\bigr)-H(\mu\,|\,\gamma_{c})-t\,L^{a}_{\mu}(g)=\sum_{k=1}^{n}\int_{X}\frac{t^{2}a_{k}^{2}}{2c_{k}}\,\beta_{k}^{2}\,d\mu-\int_{X}\Bigl(\log\det\bigl(I_{n}+tM(x)\bigr)-t\,\mathrm{tr}\,M(x)\Bigr)\mu(dx). H ( ( id + t ∇ a ψ ) # μ γ c ) − H ( μ ∣ γ c ) − t L μ a ( g ) = k = 1 ∑ n ∫ X 2 c k t 2 a k 2 β k 2 d μ − ∫ X ( log det ( I n + tM ( x ) ) − t tr M ( x ) ) μ ( d x ) .
Since 0 ≤ β k 2 ≤ B 2 0\le\beta_{k}^{2}\le B^{2} 0 ≤ β k 2 ≤ B 2 and μ ( X ) = 1 \mu(X)=1 μ ( X ) = 1 , monotonicity of the integral (Linearity and Monotonicity of the Lebesgue Integral §integrable ) and claim 6(a) of Borel Measurability and Bounded Integration on a Metric Space give 0 ≤ ∫ X t 2 a k 2 2 c k β k 2 d μ ≤ t 2 B 2 a k 2 2 c k 0\le\int_{X}\frac{t^{2}a_{k}^{2}}{2c_{k}}\beta_{k}^{2}\,d\mu\le t^{2}B^{2}\frac{a_{k}^{2}}{2c_{k}} 0 ≤ ∫ X 2 c k t 2 a k 2 β k 2 d μ ≤ t 2 B 2 2 c k a k 2 . For each x ∈ X x\in X x ∈ X , the matrix M ( x ) M(x) M ( x ) has entries of absolute value at most b b b by the hypothesis on b b b , and ∣ t ∣ b ≤ θ n ≤ 1 |t|\,b\le\theta_{n}\le1 ∣ t ∣ b ≤ θ n ≤ 1 , θ n \theta_{n} θ n being the constant c n ≤ 1 c_{n}\le1 c n ≤ 1 of Determinants of Positive Definite Matrices: Positivity, the Bound log det A ≤ t r A − d \log\det A\le\mathrm{tr}\,A-d log det A ≤ tr A − d , Bounds under Pinching, and the Expansion of det ( I + t B ) \det(I+tB) det ( I + tB ) §expansion with d = n d=n d = n , as in the statement; so Determinants of Positive Definite Matrices: Positivity, the Bound log det A ≤ t r A − d \log\det A\le\mathrm{tr}\,A-d log det A ≤ tr A − d , Bounds under Pinching, and the Expansion of det ( I + t B ) \det(I+tB) det ( I + tB ) §expansion with d = n d=n d = n , the matrix M ( x ) M(x) M ( x ) in place of B B B and m = b m=b m = b gives ∣ log det ( I n + t M ( x ) ) − t t r M ( x ) ∣ ≤ K n t 2 b 2 \bigl|\log\det(I_{n}+tM(x))-t\,\mathrm{tr}\,M(x)\bigr|\le K_{n}t^{2}b^{2} log det ( I n + tM ( x )) − t tr M ( x ) ≤ K n t 2 b 2 . The function x ↦ t t r M ( x ) = ∑ k = 1 n t a k ∂ k ∂ k g ( p n ( x ) ) x\mapsto t\,\mathrm{tr}\,M(x)=\sum_{k=1}^{n}t\,a_{k}\,\partial_{k}\partial_{k}g(p_{n}(x)) x ↦ t tr M ( x ) = ∑ k = 1 n t a k ∂ k ∂ k g ( p n ( x )) is continuous by Continuity of Sums and Products of Real-Valued Functions on a Metric Space §on-set , applied to the continuous functions x ↦ ∂ k ∂ k g ( p n ( x ) ) x\mapsto\partial_{k}\partial_{k}g(p_{n}(x)) x ↦ ∂ k ∂ k g ( p n ( x )) of Step 8; so the integrand of the last integral, its difference with the continuous function of Step 7, is continuous by the same theorem, hence Borel by claim 3 of Borel Measurability and Bounded Integration on a Metric Space , and bounded by K n t 2 b 2 K_{n}t^{2}b^{2} K n t 2 b 2 ; and so claim 6(b) of Borel Measurability and Bounded Integration on a Metric Space bounds the absolute value of that integral by K n t 2 b 2 K_{n}t^{2}b^{2} K n t 2 b 2 . The triangle inequality now gives claim 2.
Step 12 (Claim 3). If b > 0 b>0 b > 0 , let t 0 t_{0} t 0 be the lesser of 1 / ( 2 n b ) 1/(2nb) 1/ ( 2 nb ) and θ n / b \theta_{n}/b θ n / b ; if b = 0 b=0 b = 0 , let t 0 = 1 t_{0}=1 t 0 = 1 . Then t 0 > 0 t_{0}>0 t 0 > 0 , θ n \theta_{n} θ n being positive, and every t t t in the open interval I = ( − t 0 , t 0 ) I=(-t_{0},t_{0}) I = ( − t 0 , t 0 ) satisfies 2 n ∣ t ∣ b ≤ 1 2n|t|b\le1 2 n ∣ t ∣ b ≤ 1 and ∣ t ∣ b ≤ θ n |t|b\le\theta_{n} ∣ t ∣ b ≤ θ n (if b > 0 b>0 b > 0 because ∣ t ∣ < t 0 |t|<t_{0} ∣ t ∣ < t 0 , and if b = 0 b=0 b = 0 trivially). The set I I I is an interval : if s 1 , s 3 ∈ I s_{1},s_{3}\in I s 1 , s 3 ∈ I and s 1 ≤ s 2 ≤ s 3 s_{1}\le s_{2}\le s_{3} s 1 ≤ s 2 ≤ s 3 , then − t 0 < s 1 ≤ s 2 ≤ s 3 < t 0 -t_{0}<s_{1}\le s_{2}\le s_{3}<t_{0} − t 0 < s 1 ≤ s 2 ≤ s 3 < t 0 , so − t 0 < s 2 < t 0 -t_{0}<s_{2}<t_{0} − t 0 < s 2 < t 0 and s 2 ∈ I s_{2}\in I s 2 ∈ I . By claim 1, ( i d + t ∇ a ψ ) # μ (\mathrm{id}+t\nabla_{a}\psi)_{\#}\mu ( id + t ∇ a ψ ) # μ has finite relative entropy with respect to γ c \gamma_{c} γ c for every t ∈ I t\in I t ∈ I ; let H ( t ) \mathcal{H}(t) H ( t ) be its relative entropy. Since i d + 0 ∇ a ψ \mathrm{id}+0\nabla_{a}\psi id + 0 ∇ a ψ is the identity map of X X X , whose push-forward of μ \mu μ is μ \mu μ , H ( 0 ) = H ( μ ∣ γ c ) \mathcal{H}(0)=H(\mu\,|\,\gamma_{c}) H ( 0 ) = H ( μ ∣ γ c ) . The point 0 0 0 is an interior point of I I I , as − t 0 / 2 < 0 < t 0 / 2 -t_{0}/2<0<t_{0}/2 − t 0 /2 < 0 < t 0 /2 with both bounds in I I I . Let C = K n b 2 + B 2 ∑ k = 1 n a k 2 / ( 2 c k ) C=K_{n}b^{2}+B^{2}\sum_{k=1}^{n}a_{k}^{2}/(2c_{k}) C = K n b 2 + B 2 ∑ k = 1 n a k 2 / ( 2 c k ) , a nonnegative number, and L = L μ a ( g ) L=L^{a}_{\mu}(g) L = L μ a ( g ) . For s ∈ I s\in I s ∈ I with s ≠ 0 s\ne0 s = 0 , claim 2 divided by ∣ s ∣ |s| ∣ s ∣ gives
∣ H ( s ) − H ( 0 ) s − L ∣ ≤ C ∣ s ∣ . \Bigl|\frac{\mathcal{H}(s)-\mathcal{H}(0)}{s}-L\Bigr|\le C\,|s| . s H ( s ) − H ( 0 ) − L ≤ C ∣ s ∣.
Given ε > 0 \varepsilon>0 ε > 0 , let δ \delta δ be the lesser of t 0 t_{0} t 0 and ε / ( C + 1 ) \varepsilon/(C+1) ε / ( C + 1 ) . If 0 < ∣ s ∣ < δ 0<|s|<\delta 0 < ∣ s ∣ < δ , then s ∈ I s\in I s ∈ I and C ∣ s ∣ ≤ C ε / ( C + 1 ) < ε C|s|\le C\varepsilon/(C+1)<\varepsilon C ∣ s ∣ ≤ Cε / ( C + 1 ) < ε . Hence, by Derivative at an Interior Point , the function H \mathcal{H} H on I I I is differentiable at 0 0 0 with derivative L μ a ( g ) L^{a}_{\mu}(g) L μ a ( g ) , which is claim 3.