Each result cited below is universally quantified over the data in its own statement, and is applied to the data named at the point of use. The rules for adding and scaling inequalities of Elementary Order Arithmetic in an Ordered Field and Elementary Arithmetic in an Ordered Field , and the elementary properties of the absolute value in Properties of the Absolute Value in an Ordered Field (claim 4, multiplicativity; claim 5, the triangle inequality; claim 9, the strict two-sided bound), are used without further mention.
Let μ \mu μ , ρ \rho ρ and L L L be as in the statement. Here λ = λ 1 \lambda=\lambda_{1} λ = λ 1 is the Lebesgue measure of Euclidean Space and Lebesgue Measure: Standing Notation §measure , which by Lebesgue Measure on R n \mathbb{R}^n R n (case n = 1 n=1 n = 1 ) is the Lebesgue measure of Existence of Lebesgue Measure on the Real Line ; by claim 4 of the latter, λ ( ( a , b ) ) = b − a \lambda((a,b))=b-a λ (( a , b )) = b − a for real a ≤ b a\le b a ≤ b . For c ∈ R c\in\mathbb{R} c ∈ R and positive r ∈ R r\in\mathbb{R} r ∈ R write I ( c , r ) = ( c − r , c + r ) = { y ∈ R : ∣ y − c ∣ < r } I(c,r)=(c-r,c+r)=\{y\in\mathbb{R}:|y-c|<r\} I ( c , r ) = ( c − r , c + r ) = { y ∈ R : ∣ y − c ∣ < r } , an open, hence Borel, subset of R \mathbb{R} R (Euclidean Space and Lebesgue Measure: Standing Notation §borel ). Every z ∈ R 2 z\in\mathbb{R}^{2} z ∈ R 2 is ι ( x , y ) \iota(x,y) ι ( x , y ) with x = p r 1 ( z ) x=\mathrm{pr}_{1}(z) x = pr 1 ( z ) , y = p r 2 ( z ) y=\mathrm{pr}_{2}(z) y = pr 2 ( z ) by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections ; ℓ \ell ℓ is the kernel of The Logarithmic Kernel on the Real Line: Borel Measurability, a Linear Lower Bound, and the Null Diagonal of an Atomless Measure , and Δ ⊆ R 2 \Delta\subseteq\mathbb{R}^{2} Δ ⊆ R 2 its diagonal. For a set A A A , 1 A \mathbf{1}_{A} 1 A is its indicator function, and ∫ 1 A d ν = ν ( A ) \int\mathbf{1}_{A}\,d\nu=\nu(A) ∫ 1 A d ν = ν ( A ) for a measure ν \nu ν and a measurable A A A by Simple Function and Its Integral .
Step 0 (three standing facts). (D) The hypothesis says μ ( B ) = ∫ R 1 B ρ d λ \mu(B)=\int_{\mathbb{R}}\mathbf{1}_{B}\,\rho\,d\lambda μ ( B ) = ∫ R 1 B ρ d λ for every Borel B B B , so μ \mu μ is the measure with density ρ \rho ρ with respect to λ \lambda λ of claim 3 of Image Measures, Measures with Densities, and Change of Variables (ρ \rho ρ being Borel with values in [ 0 , ∞ ) [0,\infty) [ 0 , ∞ ) ). By that claim, ∫ R g d μ = ∫ R g ρ d λ \int_{\mathbb{R}}g\,d\mu=\int_{\mathbb{R}}g\rho\,d\lambda ∫ R g d μ = ∫ R g ρ d λ for every Borel g : R → [ 0 , ∞ ] g:\mathbb{R}\to[0,\infty] g : R → [ 0 , ∞ ] , and a Borel g : R → R g:\mathbb{R}\to\mathbb{R} g : R → R is μ \mu μ -integrable if and only if g ρ g\rho g ρ is λ \lambda λ -integrable, in which case the same identity holds. (M) By Square-Integrable Vector Fields Against a Probability Measure on Euclidean Space, and Test Functions: Standing Notation §measures , ( R , B ( R ) , μ ) (\mathbb{R},\mathcal{B}(\mathbb{R}),\mu) ( R , B ( R ) , μ ) and ( R 2 , B ( R 2 ) , μ ⊠ μ ) (\mathbb{R}^{2},\mathcal{B}(\mathbb{R}^{2}),\mu\boxtimes\mu) ( R 2 , B ( R 2 ) , μ ⊠ μ ) are probability spaces, and Borel real functions on R \mathbb{R} R , resp. R 2 \mathbb{R}^{2} R 2 , are exactly the random variables on them (Probability Space, Event, and Random Variable ); so by the preliminaries of Square-Integrable Random Variables and the Mean-Square Inner Product , sums, scalar multiples and products of Borel real functions on R \mathbb{R} R or on R 2 \mathbb{R}^{2} R 2 are Borel. Continuous real maps are Borel and compositions of Borel maps are Borel, by Probability Measures on Euclidean Space and Random Vectors: Standing Notation §borel-maps . (L) By The Difference Quotient of a Function with Bounded Continuous Derivative is Bounded, Symmetric and Continuous on the Plane §bound , applied to ϕ = ρ \phi=\rho ϕ = ρ (differentiable at every point, with continuous derivative bounded by L L L ), the difference quotient of ρ \rho ρ is bounded by L L L in absolute value; multiplying by ∣ x − y ∣ |x-y| ∣ x − y ∣ gives
∣ ρ ( x ) − ρ ( y ) ∣ ≤ L ∣ x − y ∣ ( x , y ∈ R ) , ( 1 ) |\rho(x)-\rho(y)|\le L|x-y|\qquad(x,y\in\mathbb{R}),\qquad(1) ∣ ρ ( x ) − ρ ( y ) ∣ ≤ L ∣ x − y ∣ ( x , y ∈ R ) , ( 1 )
trivially also for x = y x=y x = y . In particular 0 ≤ ∣ ρ ′ ( 0 ) ∣ ≤ L 0\le|\rho'(0)|\le L 0 ≤ ∣ ρ ′ ( 0 ) ∣ ≤ L .
Step 1 (ρ \rho ρ is bounded, and small intervals have small mass). Put L 1 = L + 1 ≥ 1 L_{1}=L+1\ge1 L 1 = L + 1 ≥ 1 and K = 2 L 1 K=2L_{1} K = 2 L 1 . We show ρ ( x 0 ) ≤ K \rho(x_{0})\le K ρ ( x 0 ) ≤ K for every x 0 ∈ R x_{0}\in\mathbb{R} x 0 ∈ R . Let M = ρ ( x 0 ) M=\rho(x_{0}) M = ρ ( x 0 ) ; if M = 0 M=0 M = 0 there is nothing to prove, so let M > 0 M>0 M > 0 , and put r = M / ( 2 L 1 ) > 0 r=M/(2L_{1})>0 r = M / ( 2 L 1 ) > 0 . For x ∈ I ( x 0 , r ) x\in I(x_{0},r) x ∈ I ( x 0 , r ) , (1) gives ∣ ρ ( x ) − ρ ( x 0 ) ∣ ≤ L 1 ∣ x − x 0 ∣ < L 1 r = M / 2 |\rho(x)-\rho(x_{0})|\le L_{1}|x-x_{0}|<L_{1}r=M/2 ∣ ρ ( x ) − ρ ( x 0 ) ∣ ≤ L 1 ∣ x − x 0 ∣ < L 1 r = M /2 when x ≠ x 0 x\ne x_{0} x = x 0 (and 0 < M / 2 0<M/2 0 < M /2 when x = x 0 x=x_{0} x = x 0 ), so ρ ( x ) > M / 2 \rho(x)>M/2 ρ ( x ) > M /2 . Hence 1 I ( x 0 , r ) ρ ≥ M 2 1 I ( x 0 , r ) \mathbf{1}_{I(x_{0},r)}\rho\ge\tfrac{M}{2}\mathbf{1}_{I(x_{0},r)} 1 I ( x 0 , r ) ρ ≥ 2 M 1 I ( x 0 , r ) pointwise, and by (D) and claim 1 of Linearity and Monotonicity of the Lebesgue Integral
1 = μ ( R ) = ∫ R ρ d λ ≥ ∫ R 1 I ( x 0 , r ) ρ d λ ≥ M 2 λ ( I ( x 0 , r ) ) = M 2 ⋅ 2 r = M 2 2 L 1 . 1=\mu(\mathbb{R})=\int_{\mathbb{R}}\rho\,d\lambda\ge\int_{\mathbb{R}}\mathbf{1}_{I(x_{0},r)}\rho\,d\lambda\ge\frac{M}{2}\,\lambda\bigl(I(x_{0},r)\bigr)=\frac{M}{2}\cdot2r=\frac{M^{2}}{2L_{1}} . 1 = μ ( R ) = ∫ R ρ d λ ≥ ∫ R 1 I ( x 0 , r ) ρ d λ ≥ 2 M λ ( I ( x 0 , r ) ) = 2 M ⋅ 2 r = 2 L 1 M 2 .
So M 2 ≤ 2 L 1 = K M^{2}\le2L_{1}=K M 2 ≤ 2 L 1 = K . If M ≥ 1 M\ge1 M ≥ 1 then M ≤ M ⋅ M ≤ K M\le M\cdot M\le K M ≤ M ⋅ M ≤ K ; if M < 1 M<1 M < 1 then M < 1 ≤ K M<1\le K M < 1 ≤ K . Thus 0 ≤ ρ ≤ K 0\le\rho\le K 0 ≤ ρ ≤ K on R \mathbb{R} R . Consequently, for c ∈ R c\in\mathbb{R} c ∈ R and r > 0 r>0 r > 0 , by (D) and claim 1 of Linearity and Monotonicity of the Lebesgue Integral ,
μ ( I ( c , r ) ) = ∫ R 1 I ( c , r ) ρ d λ ≤ K λ ( I ( c , r ) ) = 2 K r , ( 2 ) \mu\bigl(I(c,r)\bigr)=\int_{\mathbb{R}}\mathbf{1}_{I(c,r)}\rho\,d\lambda\le K\,\lambda\bigl(I(c,r)\bigr)=2Kr,\qquad(2) μ ( I ( c , r ) ) = ∫ R 1 I ( c , r ) ρ d λ ≤ K λ ( I ( c , r ) ) = 2 Kr , ( 2 )
and likewise μ ( { c } ) = ∫ 1 { c } ρ d λ ≤ K λ ( { c } ) = 0 \mu(\{c\})=\int\mathbf{1}_{\{c\}}\rho\,d\lambda\le K\lambda(\{c\})=0 μ ({ c }) = ∫ 1 { c } ρ d λ ≤ K λ ({ c }) = 0 , by Countable Sets are Null for an Atomless Measure, One-Point Sets are Lebesgue Null, and an Absolutely Continuous Measure is Atomless §singleton (with q = 1 q=1 q = 1 ). So μ ( { c } ) = 0 \mu(\{c\})=0 μ ({ c }) = 0 for every c c c , and therefore, by The Logarithmic Kernel on the Real Line: Borel Measurability, a Linear Lower Bound, and the Null Diagonal of an Atomless Measure §diagonal ,
( μ ⊠ μ ) ( Δ ) = 0. ( 3 ) (\mu\boxtimes\mu)(\Delta)=0.\qquad(3) ( μ ⊠ μ ) ( Δ ) = 0. ( 3 )
Step 2 (strips around the diagonal). For r > 0 r>0 r > 0 let S r = { z ∈ R 2 : ∣ p r 1 ( z ) − p r 2 ( z ) ∣ < r } S_{r}=\{z\in\mathbb{R}^{2}:|\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z)|<r\} S r = { z ∈ R 2 : ∣ pr 1 ( z ) − pr 2 ( z ) ∣ < r } . The map z ↦ ∥ p r 1 ( z ) − p r 2 ( z ) ∥ z\mapsto\lVert\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z)\rVert z ↦ ∥ pr 1 ( z ) − pr 2 ( z )∥ is Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions and equals z ↦ ∣ p r 1 ( z ) − p r 2 ( z ) ∣ z\mapsto|\mathrm{pr}_{1}(z)-\mathrm{pr}_{2}(z)| z ↦ ∣ pr 1 ( z ) − pr 2 ( z ) ∣ by One-Dimensional Test Functions: Scalars, Derivatives, and the Difference Quotient of the Derivative §scalars ; so S r S_{r} S r , its preimage of the open set ( − ∞ , r ) (-\infty,r) ( − ∞ , r ) , is Borel. By Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §product and the Tonelli part of Tonelli and Fubini Theorems ,
( μ ⊠ μ ) ( S r ) = ∫ R × R 1 S r ∘ ι d ( μ ⊗ μ ) = ∫ R ( ∫ R 1 S r ( ι ( x , y ) ) μ ( d y ) ) μ ( d x ) = ∫ R μ ( I ( x , r ) ) μ ( d x ) ≤ 2 K r , ( 4 ) (\mu\boxtimes\mu)(S_{r})=\int_{\mathbb{R}\times\mathbb{R}}\mathbf{1}_{S_{r}}\circ\iota\,d(\mu\otimes\mu)=\int_{\mathbb{R}}\Bigl(\int_{\mathbb{R}}\mathbf{1}_{S_{r}}(\iota(x,y))\,\mu(dy)\Bigr)\mu(dx)=\int_{\mathbb{R}}\mu\bigl(I(x,r)\bigr)\,\mu(dx)\le2Kr,\qquad(4) ( μ ⊠ μ ) ( S r ) = ∫ R × R 1 S r ∘ ι d ( μ ⊗ μ ) = ∫ R ( ∫ R 1 S r ( ι ( x , y )) μ ( d y ) ) μ ( d x ) = ∫ R μ ( I ( x , r ) ) μ ( d x ) ≤ 2 Kr , ( 4 )
since ι ( x , y ) ∈ S r \iota(x,y)\in S_{r} ι ( x , y ) ∈ S r exactly when y ∈ I ( x , r ) y\in I(x,r) y ∈ I ( x , r ) , and by (2) and claim 1 of Linearity and Monotonicity of the Lebesgue Integral with μ ( R ) = 1 \mu(\mathbb{R})=1 μ ( R ) = 1 .
Step 3 (μ ∈ D log \mu\in\mathcal{D}_{\log} μ ∈ D l o g ). By hypothesis μ ∈ P 2 ( R ) \mu\in\mathcal{P}_{2}(\mathbb{R}) μ ∈ P 2 ( R ) , and μ ( { c } ) = 0 \mu(\{c\})=0 μ ({ c }) = 0 for every c c c by Step 1; by The Logarithmic Energy of a Probability Measure on the Real Line §energy it remains to show that ℓ \ell ℓ , which is Borel by The Logarithmic Kernel on the Real Line: Borel Measurability, a Linear Lower Bound, and the Null Diagonal of an Atomless Measure §borel , is μ ⊠ μ \mu\boxtimes\mu μ ⊠ μ -integrable, that is (Integrable Function and the Lebesgue Integral ), ∫ ∣ ℓ ∣ d ( μ ⊠ μ ) < ∞ \int|\ell|\,d(\mu\boxtimes\mu)<\infty ∫ ∣ ℓ ∣ d ( μ ⊠ μ ) < ∞ .
(a) Let N ( z ) = ∣ p r 1 ( z ) ∣ + ∣ p r 2 ( z ) ∣ N(z)=|\mathrm{pr}_{1}(z)|+|\mathrm{pr}_{2}(z)| N ( z ) = ∣ pr 1 ( z ) ∣ + ∣ pr 2 ( z ) ∣ , a nonnegative Borel function by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions and One-Dimensional Test Functions: Scalars, Derivatives, and the Difference Quotient of the Derivative §scalars . By Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §product the image measures of μ ⊠ μ \mu\boxtimes\mu μ ⊠ μ under p r 1 \mathrm{pr}_{1} pr 1 and p r 2 \mathrm{pr}_{2} pr 2 are μ \mu μ , so by the change-of-variables formula of Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pushforward and claim 1 of Linearity and Monotonicity of the Lebesgue Integral , ∫ N d ( μ ⊠ μ ) = 2 ∫ R ∣ x ∣ μ ( d x ) \int N\,d(\mu\boxtimes\mu)=2\int_{\mathbb{R}}|x|\,\mu(dx) ∫ N d ( μ ⊠ μ ) = 2 ∫ R ∣ x ∣ μ ( d x ) . The last inequality of Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions with q = 1 q=1 q = 1 and the second point equal to 1 1 1 gives ∣ x ∣ ≤ 1 2 ( x 2 + 1 ) |x|\le\tfrac12(x^{2}+1) ∣ x ∣ ≤ 2 1 ( x 2 + 1 ) , so, with The Second Moment of a Probability Measure on Euclidean Space and the Probability Measures with Finite Second Moment §moment and One-Dimensional Test Functions: Scalars, Derivatives, and the Difference Quotient of the Derivative §scalars ,
∫ R 2 N d ( μ ⊠ μ ) ≤ ∫ R ( x 2 + 1 ) μ ( d x ) = M 2 ( μ ) + 1 < ∞ . \int_{\mathbb{R}^{2}}N\,d(\mu\boxtimes\mu)\le\int_{\mathbb{R}}(x^{2}+1)\,\mu(dx)=M_{2}(\mu)+1<\infty . ∫ R 2 N d ( μ ⊠ μ ) ≤ ∫ R ( x 2 + 1 ) μ ( d x ) = M 2 ( μ ) + 1 < ∞.
(b) For m ∈ N m\in\mathbb{N} m ∈ N let U m = ∑ k = 0 m − 1 1 S exp ( − k ) U_{m}=\sum_{k=0}^{m-1}\mathbf{1}_{S_{\exp(-k)}} U m = ∑ k = 0 m − 1 1 S e x p ( − k ) (so U 0 = 0 U_{0}=0 U 0 = 0 ), a nonnegative Borel function with U m ≤ U m + 1 U_{m}\le U_{m+1} U m ≤ U m + 1 , and let U = sup m U m : R 2 → [ 0 , ∞ ] U=\sup_{m}U_{m}:\mathbb{R}^{2}\to[0,\infty] U = sup m U m : R 2 → [ 0 , ∞ ] . By Monotone Convergence Theorem , U U U is measurable and ∫ U d ( μ ⊠ μ ) = sup m ∫ U m d ( μ ⊠ μ ) \int U\,d(\mu\boxtimes\mu)=\sup_{m}\int U_{m}\,d(\mu\boxtimes\mu) ∫ U d ( μ ⊠ μ ) = sup m ∫ U m d ( μ ⊠ μ ) . By claim 1 of Linearity and Monotonicity of the Lebesgue Integral and (4), ∫ U m d ( μ ⊠ μ ) ≤ 2 K ∑ k = 0 m − 1 exp ( − k ) \int U_{m}\,d(\mu\boxtimes\mu)\le2K\sum_{k=0}^{m-1}\exp(-k) ∫ U m d ( μ ⊠ μ ) ≤ 2 K ∑ k = 0 m − 1 exp ( − k ) . By claim 4 of Basic Properties of the Exponential Function , exp ( 1 ) ≥ 1 + 1 = 2 \exp(1)\ge1+1=2 exp ( 1 ) ≥ 1 + 1 = 2 , so by claim 2 there 0 < exp ( − 1 ) = 1 / exp ( 1 ) ≤ 1 2 0<\exp(-1)=1/\exp(1)\le\tfrac12 0 < exp ( − 1 ) = 1/ exp ( 1 ) ≤ 2 1 ; by claim 1 there and induction on k k k , exp ( − k ) = exp ( − 1 ) k ≤ ( 1 2 ) k \exp(-k)=\exp(-1)^{k}\le(\tfrac12)^{k} exp ( − k ) = exp ( − 1 ) k ≤ ( 2 1 ) k . By Series of Nonnegative Real Numbers, Comparison, and the Geometric Series §geometric with r = 1 2 r=\tfrac12 r = 2 1 , ∑ k = 1 n ( 1 2 ) k = 1 − ( 1 2 ) n ≤ 1 \sum_{k=1}^{n}(\tfrac12)^{k}=1-(\tfrac12)^{n}\le1 ∑ k = 1 n ( 2 1 ) k = 1 − ( 2 1 ) n ≤ 1 for n ≥ 1 n\ge1 n ≥ 1 ; hence ∑ k = 0 m − 1 exp ( − k ) ≤ 2 \sum_{k=0}^{m-1}\exp(-k)\le2 ∑ k = 0 m − 1 exp ( − k ) ≤ 2 for every m m m , and
∫ R 2 U d ( μ ⊠ μ ) ≤ 4 K . \int_{\mathbb{R}^{2}}U\,d(\mu\boxtimes\mu)\le4K . ∫ R 2 U d ( μ ⊠ μ ) ≤ 4 K .
(c) We show ∣ ℓ ( z ) ∣ ≤ U ( z ) + N ( z ) |\ell(z)|\le U(z)+N(z) ∣ ℓ ( z ) ∣ ≤ U ( z ) + N ( z ) for every z = ι ( x , y ) z=\iota(x,y) z = ι ( x , y ) . By The Logarithmic Kernel on the Real Line: Borel Measurability, a Linear Lower Bound, and the Null Diagonal of an Atomless Measure §lower-bound , − ℓ ( z ) ≤ N ( z ) -\ell(z)\le N(z) − ℓ ( z ) ≤ N ( z ) ; so if ℓ ( z ) ≤ 0 \ell(z)\le0 ℓ ( z ) ≤ 0 then ∣ ℓ ( z ) ∣ = − ℓ ( z ) ≤ N ( z ) |\ell(z)|=-\ell(z)\le N(z) ∣ ℓ ( z ) ∣ = − ℓ ( z ) ≤ N ( z ) . Let ℓ ( z ) > 0 \ell(z)>0 ℓ ( z ) > 0 ; then x ≠ y x\ne y x = y (as ℓ \ell ℓ vanishes on Δ \Delta Δ ), and w = ℓ ( z ) = − log ∣ x − y ∣ w=\ell(z)=-\log|x-y| w = ℓ ( z ) = − log ∣ x − y ∣ is positive with exp ( − w ) = ∣ x − y ∣ \exp(-w)=|x-y| exp ( − w ) = ∣ x − y ∣ by The Natural Logarithm . As exp \exp exp is strictly increasing (claim 4 of Basic Properties of the Exponential Function ), for k ∈ N k\in\mathbb{N} k ∈ N we have z ∈ S exp ( − k ) z\in S_{\exp(-k)} z ∈ S e x p ( − k ) , i.e. exp ( − w ) < exp ( − k ) \exp(-w)<\exp(-k) exp ( − w ) < exp ( − k ) , exactly when k < w k<w k < w . Let c m c_{m} c m be the number of k ∈ { 0 , … , m − 1 } k\in\{0,\dots,m-1\} k ∈ { 0 , … , m − 1 } with k < w k<w k < w , so U m ( z ) = c m U_{m}(z)=c_{m} U m ( z ) = c m . By induction on m m m , c m ≥ min { m , w } c_{m}\ge\min\{m,w\} c m ≥ min { m , w } : for m = 0 m=0 m = 0 both sides are 0 0 0 since w > 0 w>0 w > 0 ; if w ≤ m w\le m w ≤ m then min { m + 1 , w } = w = min { m , w } ≤ c m ≤ c m + 1 \min\{m+1,w\}=w=\min\{m,w\}\le c_{m}\le c_{m+1} min { m + 1 , w } = w = min { m , w } ≤ c m ≤ c m + 1 ; if w > m w>m w > m then every k ≤ m k\le m k ≤ m satisfies k < w k<w k < w , so c m + 1 = m + 1 ≥ min { m + 1 , w } c_{m+1}=m+1\ge\min\{m+1,w\} c m + 1 = m + 1 ≥ min { m + 1 , w } . By claim 1 of The Archimedean Property of the Real Numbers choose m ∈ N m\in\mathbb{N} m ∈ N with w < m w<m w < m ; then U ( z ) ≥ U m ( z ) = c m ≥ w = ∣ ℓ ( z ) ∣ U(z)\ge U_{m}(z)=c_{m}\ge w=|\ell(z)| U ( z ) ≥ U m ( z ) = c m ≥ w = ∣ ℓ ( z ) ∣ .
(d) By (c), claim 1 of Linearity and Monotonicity of the Lebesgue Integral , (a) and (b), ∫ ∣ ℓ ∣ d ( μ ⊠ μ ) ≤ 4 K + M 2 ( μ ) + 1 < ∞ \int|\ell|\,d(\mu\boxtimes\mu)\le4K+M_{2}(\mu)+1<\infty ∫ ∣ ℓ ∣ d ( μ ⊠ μ ) ≤ 4 K + M 2 ( μ ) + 1 < ∞ . Hence μ ∈ D log \mu\in\mathcal{D}_{\log} μ ∈ D l o g .
Step 4 (a regularised kernel). For real ε \varepsilon ε with 0 < ε ≤ 1 0<\varepsilon\le1 0 < ε ≤ 1 define w ε : R → R w_{\varepsilon}:\mathbb{R}\to\mathbb{R} w ε : R → R by w ε ( r ) = r / ( r 2 + ε 2 ) w_{\varepsilon}(r)=r/(r^{2}+\varepsilon^{2}) w ε ( r ) = r / ( r 2 + ε 2 ) , the denominator being at least ε 2 > 0 \varepsilon^{2}>0 ε 2 > 0 . Then: (W1) w ε ( − r ) = − w ε ( r ) w_{\varepsilon}(-r)=-w_{\varepsilon}(r) w ε ( − r ) = − w ε ( r ) and w ε ( 0 ) = 0 w_{\varepsilon}(0)=0 w ε ( 0 ) = 0 . (W2) ∣ w ε ( r ) ∣ ≤ 1 / ( 2 ε ) |w_{\varepsilon}(r)|\le1/(2\varepsilon) ∣ w ε ( r ) ∣ ≤ 1/ ( 2 ε ) for every r r r , because r 2 + ε 2 − 2 ∣ r ∣ ε = ( ∣ r ∣ − ε ) 2 ≥ 0 r^{2}+\varepsilon^{2}-2|r|\varepsilon=(|r|-\varepsilon)^{2}\ge0 r 2 + ε 2 − 2∣ r ∣ ε = ( ∣ r ∣ − ε ) 2 ≥ 0 ; and ∣ w ε ( r ) ∣ ≤ ∣ r ∣ / r 2 = 1 / ∣ r ∣ |w_{\varepsilon}(r)|\le|r|/r^{2}=1/|r| ∣ w ε ( r ) ∣ ≤ ∣ r ∣/ r 2 = 1/∣ r ∣ for r ≠ 0 r\ne0 r = 0 . (W3) For all r , s r,s r , s , w ε ( r ) − w ε ( s ) = ( r − s ) ( ε 2 − r s ) / D w_{\varepsilon}(r)-w_{\varepsilon}(s)=(r-s)(\varepsilon^{2}-rs)/D w ε ( r ) − w ε ( s ) = ( r − s ) ( ε 2 − rs ) / D with D = ( r 2 + ε 2 ) ( s 2 + ε 2 ) ≥ ε 2 ( r 2 + s 2 + ε 2 ) ≥ ε 2 ( ∣ r s ∣ + ε 2 ) D=(r^{2}+\varepsilon^{2})(s^{2}+\varepsilon^{2})\ge\varepsilon^{2}(r^{2}+s^{2}+\varepsilon^{2})\ge\varepsilon^{2}(|rs|+\varepsilon^{2}) D = ( r 2 + ε 2 ) ( s 2 + ε 2 ) ≥ ε 2 ( r 2 + s 2 + ε 2 ) ≥ ε 2 ( ∣ rs ∣ + ε 2 ) (using r 2 + s 2 ≥ 2 ∣ r s ∣ r^{2}+s^{2}\ge2|rs| r 2 + s 2 ≥ 2∣ rs ∣ ), so ∣ w ε ( r ) − w ε ( s ) ∣ ≤ ∣ r − s ∣ / ε 2 |w_{\varepsilon}(r)-w_{\varepsilon}(s)|\le|r-s|/\varepsilon^{2} ∣ w ε ( r ) − w ε ( s ) ∣ ≤ ∣ r − s ∣/ ε 2 ; hence w ε w_{\varepsilon} w ε is continuous for d R d_{\mathbb{R}} d R (given η > 0 \eta>0 η > 0 take δ = ε 2 η \delta=\varepsilon^{2}\eta δ = ε 2 η in Continuous Map Between Metric Spaces ), and Borel by (M). (W4) For r ≠ 0 r\ne0 r = 0 , r w ε ( r ) = r 2 / ( r 2 + ε 2 ) ∈ [ 0 , 1 ] r\,w_{\varepsilon}(r)=r^{2}/(r^{2}+\varepsilon^{2})\in[0,1] r w ε ( r ) = r 2 / ( r 2 + ε 2 ) ∈ [ 0 , 1 ] and 1 − r w ε ( r ) = ε 2 / ( r 2 + ε 2 ) ≤ ε 2 / r 2 1-r\,w_{\varepsilon}(r)=\varepsilon^{2}/(r^{2}+\varepsilon^{2})\le\varepsilon^{2}/r^{2} 1 − r w ε ( r ) = ε 2 / ( r 2 + ε 2 ) ≤ ε 2 / r 2 .
For x ∈ R x\in\mathbb{R} x ∈ R the map y ↦ w ε ( x − y ) y\mapsto w_{\varepsilon}(x-y) y ↦ w ε ( x − y ) satisfies ∣ w ε ( x − y ) − w ε ( x − y ′ ) ∣ ≤ ∣ y − y ′ ∣ / ε 2 |w_{\varepsilon}(x-y)-w_{\varepsilon}(x-y')|\le|y-y'|/\varepsilon^{2} ∣ w ε ( x − y ) − w ε ( x − y ′ ) ∣ ≤ ∣ y − y ′ ∣/ ε 2 by (W3), so it is continuous, hence Borel, and bounded by (W2), hence μ \mu μ -integrable (Probability Measures on Euclidean Space and Random Vectors: Standing Notation §measures ). Put
G ε ( x ) = ∫ R w ε ( x − y ) μ ( d y ) ( x ∈ R ) . G_{\varepsilon}(x)=\int_{\mathbb{R}}w_{\varepsilon}(x-y)\,\mu(dy)\qquad(x\in\mathbb{R}). G ε ( x ) = ∫ R w ε ( x − y ) μ ( d y ) ( x ∈ R ) .
By claim 2 of Linearity and Monotonicity of the Lebesgue Integral and (W3), ∣ G ε ( x ) − G ε ( x ′ ) ∣ ≤ ∫ ∣ w ε ( x − y ) − w ε ( x ′ − y ) ∣ μ ( d y ) ≤ ∣ x − x ′ ∣ / ε 2 |G_{\varepsilon}(x)-G_{\varepsilon}(x')|\le\int|w_{\varepsilon}(x-y)-w_{\varepsilon}(x'-y)|\,\mu(dy)\le|x-x'|/\varepsilon^{2} ∣ G ε ( x ) − G ε ( x ′ ) ∣ ≤ ∫ ∣ w ε ( x − y ) − w ε ( x ′ − y ) ∣ μ ( d y ) ≤ ∣ x − x ′ ∣/ ε 2 , so G ε G_{\varepsilon} G ε is continuous, hence Borel.
Step 5 (∣ G ε ∣ ≤ 2 L + 1 |G_{\varepsilon}|\le2L+1 ∣ G ε ∣ ≤ 2 L + 1 ). Fix x ∈ R x\in\mathbb{R} x ∈ R and ε ∈ ( 0 , 1 ] \varepsilon\in(0,1] ε ∈ ( 0 , 1 ] , and let J = I ( x , 1 ) J=I(x,1) J = I ( x , 1 ) . Define Borel functions (by (M)) on R \mathbb{R} R :
C ( y ) = 1 J ( y ) w ε ( x − y ) , A ( y ) = C ( y ) ( ρ ( y ) − ρ ( x ) ) , B ( y ) = 1 R ∖ J ( y ) w ε ( x − y ) ρ ( y ) , C(y)=\mathbf{1}_{J}(y)\,w_{\varepsilon}(x-y),\qquad A(y)=C(y)\bigl(\rho(y)-\rho(x)\bigr),\qquad B(y)=\mathbf{1}_{\mathbb{R}\setminus J}(y)\,w_{\varepsilon}(x-y)\,\rho(y), C ( y ) = 1 J ( y ) w ε ( x − y ) , A ( y ) = C ( y ) ( ρ ( y ) − ρ ( x ) ) , B ( y ) = 1 R ∖ J ( y ) w ε ( x − y ) ρ ( y ) ,
so that w ε ( x − y ) ρ ( y ) = A ( y ) + ρ ( x ) C ( y ) + B ( y ) w_{\varepsilon}(x-y)\rho(y)=A(y)+\rho(x)C(y)+B(y) w ε ( x − y ) ρ ( y ) = A ( y ) + ρ ( x ) C ( y ) + B ( y ) for every y y y . By (W2), ∣ C ∣ ≤ 1 2 ε 1 J |C|\le\frac{1}{2\varepsilon}\mathbf{1}_{J} ∣ C ∣ ≤ 2 ε 1 1 J , so ∫ ∣ C ∣ d λ ≤ 1 2 ε λ ( J ) < ∞ \int|C|\,d\lambda\le\frac{1}{2\varepsilon}\lambda(J)<\infty ∫ ∣ C ∣ d λ ≤ 2 ε 1 λ ( J ) < ∞ and C C C is λ \lambda λ -integrable. For y ∈ J y\in J y ∈ J with y ≠ x y\ne x y = x , (W2) and (1) give ∣ A ( y ) ∣ ≤ 1 ∣ x − y ∣ L ∣ x − y ∣ = L |A(y)|\le\frac{1}{|x-y|}\,L|x-y|=L ∣ A ( y ) ∣ ≤ ∣ x − y ∣ 1 L ∣ x − y ∣ = L , while A ( x ) = 0 A(x)=0 A ( x ) = 0 and A = 0 A=0 A = 0 off J J J ; so ∣ A ∣ ≤ L 1 J |A|\le L\,\mathbf{1}_{J} ∣ A ∣ ≤ L 1 J , A A A is λ \lambda λ -integrable, and ∣ ∫ A d λ ∣ ≤ ∫ ∣ A ∣ d λ ≤ L λ ( J ) = 2 L |\int A\,d\lambda|\le\int|A|\,d\lambda\le L\lambda(J)=2L ∣ ∫ A d λ ∣ ≤ ∫ ∣ A ∣ d λ ≤ L λ ( J ) = 2 L (claims 1 and 2 of Linearity and Monotonicity of the Lebesgue Integral ). By (D), y ↦ w ε ( x − y ) ρ ( y ) y\mapsto w_{\varepsilon}(x-y)\rho(y) y ↦ w ε ( x − y ) ρ ( y ) is λ \lambda λ -integrable with integral G ε ( x ) G_{\varepsilon}(x) G ε ( x ) ; hence B B B is λ \lambda λ -integrable by claim 2 of Linearity and Monotonicity of the Lebesgue Integral , and
G ε ( x ) = ∫ R A d λ + ρ ( x ) ∫ R C d λ + ∫ R B d λ . G_{\varepsilon}(x)=\int_{\mathbb{R}}A\,d\lambda+\rho(x)\int_{\mathbb{R}}C\,d\lambda+\int_{\mathbb{R}}B\,d\lambda . G ε ( x ) = ∫ R A d λ + ρ ( x ) ∫ R C d λ + ∫ R B d λ .
The function g = 1 R ∖ J w ε ( x − ⋅ ) g=\mathbf{1}_{\mathbb{R}\setminus J}\,w_{\varepsilon}(x-\cdot) g = 1 R ∖ J w ε ( x − ⋅ ) is Borel with g ρ = B g\rho=B g ρ = B , and ∣ g ∣ ≤ 1 |g|\le1 ∣ g ∣ ≤ 1 , since y ∉ J y\notin J y ∈ / J means ∣ x − y ∣ ≥ 1 |x-y|\ge1 ∣ x − y ∣ ≥ 1 and then ∣ w ε ( x − y ) ∣ ≤ 1 / ∣ x − y ∣ ≤ 1 |w_{\varepsilon}(x-y)|\le1/|x-y|\le1 ∣ w ε ( x − y ) ∣ ≤ 1/∣ x − y ∣ ≤ 1 by (W2); by (D) and claim 6(b) of Borel Measurability and Bounded Integration on a Metric Space , ∣ ∫ B d λ ∣ = ∣ ∫ g d μ ∣ ≤ 1 |\int B\,d\lambda|=|\int g\,d\mu|\le1 ∣ ∫ B d λ ∣ = ∣ ∫ g d μ ∣ ≤ 1 . Finally, for every y y y , the point 2 x − y 2x-y 2 x − y lies in J J J exactly when y y y does (as ∣ ( 2 x − y ) − x ∣ = ∣ x − y ∣ |(2x-y)-x|=|x-y| ∣ ( 2 x − y ) − x ∣ = ∣ x − y ∣ ), and w ε ( x − ( 2 x − y ) ) = w ε ( y − x ) = − w ε ( x − y ) w_{\varepsilon}(x-(2x-y))=w_{\varepsilon}(y-x)=-w_{\varepsilon}(x-y) w ε ( x − ( 2 x − y )) = w ε ( y − x ) = − w ε ( x − y ) by (W1); so C ( 2 x − y ) = − C ( y ) C(2x-y)=-C(y) C ( 2 x − y ) = − C ( y ) . Claim 3 of Translation and Reflection Invariance of Lebesgue Measure on R n \mathbb{R}^n R n (with n = 1 n=1 n = 1 and a = 2 x a=2x a = 2 x ) gives ∫ C ( 2 x − y ) λ ( d y ) = ∫ C d λ \int C(2x-y)\,\lambda(dy)=\int C\,d\lambda ∫ C ( 2 x − y ) λ ( d y ) = ∫ C d λ , i.e. − ∫ C d λ = ∫ C d λ -\int C\,d\lambda=\int C\,d\lambda − ∫ C d λ = ∫ C d λ , so ∫ C d λ = 0 \int C\,d\lambda=0 ∫ C d λ = 0 . Therefore
∣ G ε ( x ) ∣ ≤ 2 L + 1 ( x ∈ R , 0 < ε ≤ 1 ) . ( 5 ) |G_{\varepsilon}(x)|\le2L+1\qquad(x\in\mathbb{R},\ 0<\varepsilon\le1).\qquad(5) ∣ G ε ( x ) ∣ ≤ 2 L + 1 ( x ∈ R , 0 < ε ≤ 1 ) . ( 5 )
Step 6 (the regularised double integral). Fix ψ ∈ C c ∞ ( R ) \psi\in C_{c}^{\infty}(\mathbb{R}) ψ ∈ C c ∞ ( R ) . By One-Dimensional Test Functions: Scalars, Derivatives, and the Difference Quotient of the Derivative §derivatives , ψ ′ \psi' ψ ′ is continuous and bounded, say ∣ ψ ′ ∣ ≤ B ψ |\psi'|\le B_{\psi} ∣ ψ ′ ∣ ≤ B ψ , and by One-Dimensional Test Functions: Scalars, Derivatives, and the Difference Quotient of the Derivative §quotient fix L ψ ≥ 0 L_{\psi}\ge0 L ψ ≥ 0 with ∣ F ψ ∣ ≤ L ψ |F_{\psi}|\le L_{\psi} ∣ F ψ ∣ ≤ L ψ on R 2 \mathbb{R}^{2} R 2 ; F ψ F_{\psi} F ψ is Borel. Put C ∗ = 2 ( 2 L + 1 ) C_{*}=2(2L+1) C ∗ = 2 ( 2 L + 1 ) , a nonnegative number not depending on ψ \psi ψ . For ε ∈ ( 0 , 1 ] \varepsilon\in(0,1] ε ∈ ( 0 , 1 ] define Φ 1 , Φ 2 , Φ ε : R 2 → R \Phi^{1},\Phi^{2},\Phi_{\varepsilon}:\mathbb{R}^{2}\to\mathbb{R} Φ 1 , Φ 2 , Φ ε : R 2 → R by
Φ 1 ( ι ( x , y ) ) = w ε ( x − y ) ψ ′ ( x ) , Φ 2 ( ι ( x , y ) ) = w ε ( x − y ) ψ ′ ( y ) , Φ ε = Φ 1 − Φ 2 . \Phi^{1}(\iota(x,y))=w_{\varepsilon}(x-y)\psi'(x),\qquad\Phi^{2}(\iota(x,y))=w_{\varepsilon}(x-y)\psi'(y),\qquad\Phi_{\varepsilon}=\Phi^{1}-\Phi^{2}. Φ 1 ( ι ( x , y )) = w ε ( x − y ) ψ ′ ( x ) , Φ 2 ( ι ( x , y )) = w ε ( x − y ) ψ ′ ( y ) , Φ ε = Φ 1 − Φ 2 .
They are Borel by (M) (p r 1 \mathrm{pr}_{1} pr 1 , p r 2 \mathrm{pr}_{2} pr 2 being Borel by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §projections , ψ ′ \psi' ψ ′ and w ε w_{\varepsilon} w ε being continuous), and bounded by B ψ / ε B_{\psi}/\varepsilon B ψ / ε in absolute value by (W2). Since μ ⊠ μ \mu\boxtimes\mu μ ⊠ μ is the image measure of μ ⊗ μ \mu\otimes\mu μ ⊗ μ under ι \iota ι , which is measurable (Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §product ), claim 2 of Image Measures, Measures with Densities, and Change of Variables shows that Φ 1 ∘ ι \Phi^{1}\circ\iota Φ 1 ∘ ι is μ ⊗ μ \mu\otimes\mu μ ⊗ μ -integrable with ∫ Φ 1 d ( μ ⊠ μ ) = ∫ Φ 1 ∘ ι d ( μ ⊗ μ ) \int\Phi^{1}\,d(\mu\boxtimes\mu)=\int\Phi^{1}\circ\iota\,d(\mu\otimes\mu) ∫ Φ 1 d ( μ ⊠ μ ) = ∫ Φ 1 ∘ ι d ( μ ⊗ μ ) (Φ 1 \Phi^{1} Φ 1 being μ ⊠ μ \mu\boxtimes\mu μ ⊠ μ -integrable as a bounded Borel function, Probability Measures on Euclidean Space and Random Vectors: Standing Notation §measures ). By the Fubini part of Tonelli and Fubini Theorems there is N 1 ∈ B ( R ) N_{1}\in\mathcal{B}(\mathbb{R}) N 1 ∈ B ( R ) with μ ( N 1 ) = 0 \mu(N_{1})=0 μ ( N 1 ) = 0 such that the function H H H equal to ∫ Φ 1 ( ι ( x , y ) ) μ ( d y ) \int\Phi^{1}(\iota(x,y))\,\mu(dy) ∫ Φ 1 ( ι ( x , y )) μ ( d y ) off N 1 N_{1} N 1 and to 0 0 0 on N 1 N_{1} N 1 is μ \mu μ -integrable with ∫ Φ 1 ∘ ι d ( μ ⊗ μ ) = ∫ H d μ \int\Phi^{1}\circ\iota\,d(\mu\otimes\mu)=\int H\,d\mu ∫ Φ 1 ∘ ι d ( μ ⊗ μ ) = ∫ H d μ . For every x x x , claim 2 of Linearity and Monotonicity of the Lebesgue Integral gives ∫ Φ 1 ( ι ( x , y ) ) μ ( d y ) = ψ ′ ( x ) G ε ( x ) \int\Phi^{1}(\iota(x,y))\,\mu(dy)=\psi'(x)G_{\varepsilon}(x) ∫ Φ 1 ( ι ( x , y )) μ ( d y ) = ψ ′ ( x ) G ε ( x ) ; thus H H H agrees with the Borel function ψ ′ G ε \psi'G_{\varepsilon} ψ ′ G ε (by (M)) off the null set N 1 N_{1} N 1 , and The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §comparison yields ∫ Φ 1 d ( μ ⊠ μ ) = ∫ ψ ′ G ε d μ \int\Phi^{1}\,d(\mu\boxtimes\mu)=\int\psi'G_{\varepsilon}\,d\mu ∫ Φ 1 d ( μ ⊠ μ ) = ∫ ψ ′ G ε d μ . In the same way, using the Fubini identity in the other order and, for every y y y , ∫ w ε ( x − y ) ψ ′ ( y ) μ ( d x ) = − ψ ′ ( y ) ∫ w ε ( y − x ) μ ( d x ) = − ψ ′ ( y ) G ε ( y ) \int w_{\varepsilon}(x-y)\psi'(y)\,\mu(dx)=-\psi'(y)\int w_{\varepsilon}(y-x)\,\mu(dx)=-\psi'(y)G_{\varepsilon}(y) ∫ w ε ( x − y ) ψ ′ ( y ) μ ( d x ) = − ψ ′ ( y ) ∫ w ε ( y − x ) μ ( d x ) = − ψ ′ ( y ) G ε ( y ) by (W1), we get ∫ Φ 2 d ( μ ⊠ μ ) = − ∫ ψ ′ G ε d μ \int\Phi^{2}\,d(\mu\boxtimes\mu)=-\int\psi'G_{\varepsilon}\,d\mu ∫ Φ 2 d ( μ ⊠ μ ) = − ∫ ψ ′ G ε d μ . Hence
∫ R 2 Φ ε d ( μ ⊠ μ ) = 2 ∫ R ψ ′ G ε d μ . \int_{\mathbb{R}^{2}}\Phi_{\varepsilon}\,d(\mu\boxtimes\mu)=2\int_{\mathbb{R}}\psi'\,G_{\varepsilon}\,d\mu . ∫ R 2 Φ ε d ( μ ⊠ μ ) = 2 ∫ R ψ ′ G ε d μ .
By claims 1 and 2 of Linearity and Monotonicity of the Lebesgue Integral and (5), ∣ ∫ ψ ′ G ε d μ ∣ ≤ ( 2 L + 1 ) ∫ ∣ ψ ′ ∣ d μ |\int\psi'G_{\varepsilon}\,d\mu|\le(2L+1)\int|\psi'|\,d\mu ∣ ∫ ψ ′ G ε d μ ∣ ≤ ( 2 L + 1 ) ∫ ∣ ψ ′ ∣ d μ . The functions ∣ ψ ′ ∣ |\psi'| ∣ ψ ′ ∣ and 1 1 1 are bounded, hence square-integrable random variables on ( R , B ( R ) , μ ) (\mathbb{R},\mathcal{B}(\mathbb{R}),\mu) ( R , B ( R ) , μ ) , so claim 1 of Cauchy-Schwarz and Triangle Inequalities for the Mean-Square Norm gives ∫ ∣ ψ ′ ∣ d μ = E [ ∣ ψ ′ ∣ ⋅ 1 ] ≤ ∥ ∣ ψ ′ ∣ ∥ 2 ∥ 1 ∥ 2 \int|\psi'|\,d\mu=\mathbb{E}[|\psi'|\cdot1]\le\lVert\,|\psi'|\,\rVert_{2}\,\lVert1\rVert_{2} ∫ ∣ ψ ′ ∣ d μ = E [ ∣ ψ ′ ∣ ⋅ 1 ] ≤ ∥ ∣ ψ ′ ∣ ∥ 2 ∥ 1 ∥ 2 , where by Square-Integrable Random Variables and the Mean-Square Inner Product ∥ 1 ∥ 2 = 1 = 1 \lVert1\rVert_{2}=\sqrt{1}=1 ∥ 1 ∥ 2 = 1 = 1 (Existence and Uniqueness of the Nonnegative Square Root ) and ∥ ∣ ψ ′ ∣ ∥ 2 = ∫ ( ψ ′ ) 2 d μ = ∥ ∇ ψ ∥ μ \lVert\,|\psi'|\,\rVert_{2}=\sqrt{\int(\psi')^{2}\,d\mu}=\lVert\nabla\psi\rVert_{\mu} ∥ ∣ ψ ′ ∣ ∥ 2 = ∫ ( ψ ′ ) 2 d μ = ∥ ∇ ψ ∥ μ by One-Dimensional Test Functions: Scalars, Derivatives, and the Difference Quotient of the Derivative §derivatives . Therefore
∣ ∫ R 2 Φ ε d ( μ ⊠ μ ) ∣ ≤ C ∗ ∥ ∇ ψ ∥ μ ( 0 < ε ≤ 1 ) . ( 6 ) \Bigl|\int_{\mathbb{R}^{2}}\Phi_{\varepsilon}\,d(\mu\boxtimes\mu)\Bigr|\le C_{*}\lVert\nabla\psi\rVert_{\mu}\qquad(0<\varepsilon\le1).\qquad(6) ∫ R 2 Φ ε d ( μ ⊠ μ ) ≤ C ∗ ∥ ∇ ψ ∥ μ ( 0 < ε ≤ 1 ) . ( 6 )
Step 7 (removing the regularisation; μ ∈ P 2 Φ ∗ ( R ) \mu\in\mathcal{P}_{2}^{\Phi^{*}}(\mathbb{R}) μ ∈ P 2 Φ ∗ ( R ) ). For m ∈ N m\in\mathbb{N} m ∈ N put ε m = 1 / ( m + 1 ) ∈ ( 0 , 1 ] \varepsilon_{m}=1/(m+1)\in(0,1] ε m = 1/ ( m + 1 ) ∈ ( 0 , 1 ] . Let z = ι ( x , y ) z=\iota(x,y) z = ι ( x , y ) . If x = y x=y x = y then Φ ε ( z ) = 0 \Phi_{\varepsilon}(z)=0 Φ ε ( z ) = 0 by (W1). If x ≠ y x\ne y x = y , then ψ ′ ( x ) − ψ ′ ( y ) = ( x − y ) F ψ ( z ) \psi'(x)-\psi'(y)=(x-y)F_{\psi}(z) ψ ′ ( x ) − ψ ′ ( y ) = ( x − y ) F ψ ( z ) by the definition of F ψ F_{\psi} F ψ , so Φ ε ( z ) = ( x − y ) w ε ( x − y ) F ψ ( z ) \Phi_{\varepsilon}(z)=(x-y)w_{\varepsilon}(x-y)\,F_{\psi}(z) Φ ε ( z ) = ( x − y ) w ε ( x − y ) F ψ ( z ) and by (W4)
∣ Φ ε ( z ) ∣ ≤ ∣ F ψ ( z ) ∣ ≤ L ψ , ∣ F ψ ( z ) − Φ ε ( z ) ∣ ≤ L ψ ε 2 ( x − y ) 2 ≤ L ψ ( m + 1 ) ( x − y ) 2 ( ε = ε m ) , |\Phi_{\varepsilon}(z)|\le|F_{\psi}(z)|\le L_{\psi},\qquad|F_{\psi}(z)-\Phi_{\varepsilon}(z)|\le L_{\psi}\,\frac{\varepsilon^{2}}{(x-y)^{2}}\le\frac{L_{\psi}}{(m+1)(x-y)^{2}}\quad(\varepsilon=\varepsilon_{m}), ∣ Φ ε ( z ) ∣ ≤ ∣ F ψ ( z ) ∣ ≤ L ψ , ∣ F ψ ( z ) − Φ ε ( z ) ∣ ≤ L ψ ( x − y ) 2 ε 2 ≤ ( m + 1 ) ( x − y ) 2 L ψ ( ε = ε m ) ,
using ε m 2 ≤ ε m \varepsilon_{m}^{2}\le\varepsilon_{m} ε m 2 ≤ ε m . Given η > 0 \eta>0 η > 0 , claim 2 of The Archimedean Property of the Real Numbers yields m 0 m_{0} m 0 with L ψ < m 0 ( x − y ) 2 η L_{\psi}<m_{0}(x-y)^{2}\eta L ψ < m 0 ( x − y ) 2 η , and then ∣ F ψ ( z ) − Φ ε m ( z ) ∣ < η |F_{\psi}(z)-\Phi_{\varepsilon_{m}}(z)|<\eta ∣ F ψ ( z ) − Φ ε m ( z ) ∣ < η for all m ≥ m 0 m\ge m_{0} m ≥ m 0 ; so Φ ε m ( z ) → F ψ ( z ) \Phi_{\varepsilon_{m}}(z)\to F_{\psi}(z) Φ ε m ( z ) → F ψ ( z ) for every z ∉ Δ z\notin\Delta z ∈ / Δ , that is, almost everywhere by (3). Moreover ∣ Φ ε m ∣ ≤ L ψ |\Phi_{\varepsilon_{m}}|\le L_{\psi} ∣ Φ ε m ∣ ≤ L ψ everywhere, and the constant L ψ L_{\psi} L ψ is μ ⊠ μ \mu\boxtimes\mu μ ⊠ μ -integrable (claim 6(a) of Borel Measurability and Bounded Integration on a Metric Space ). By The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §dominated ,
lim m → ∞ ∫ R 2 Φ ε m d ( μ ⊠ μ ) = ∫ R 2 F ψ d ( μ ⊠ μ ) . \lim_{m\to\infty}\int_{\mathbb{R}^{2}}\Phi_{\varepsilon_{m}}\,d(\mu\boxtimes\mu)=\int_{\mathbb{R}^{2}}F_{\psi}\,d(\mu\boxtimes\mu). m → ∞ lim ∫ R 2 Φ ε m d ( μ ⊠ μ ) = ∫ R 2 F ψ d ( μ ⊠ μ ) .
If we had ∣ ∫ F ψ d ( μ ⊠ μ ) ∣ > C ∗ ∥ ∇ ψ ∥ μ \bigl|\int F_{\psi}\,d(\mu\boxtimes\mu)\bigr|>C_{*}\lVert\nabla\psi\rVert_{\mu} ∫ F ψ d ( μ ⊠ μ ) > C ∗ ∥ ∇ ψ ∥ μ , then with η = ∣ ∫ F ψ d ( μ ⊠ μ ) ∣ − C ∗ ∥ ∇ ψ ∥ μ > 0 \eta=\bigl|\int F_{\psi}\,d(\mu\boxtimes\mu)\bigr|-C_{*}\lVert\nabla\psi\rVert_{\mu}>0 η = ∫ F ψ d ( μ ⊠ μ ) − C ∗ ∥ ∇ ψ ∥ μ > 0 some m m m would satisfy ∣ ∫ Φ ε m d ( μ ⊠ μ ) − ∫ F ψ d ( μ ⊠ μ ) ∣ < η \bigl|\int\Phi_{\varepsilon_{m}}\,d(\mu\boxtimes\mu)-\int F_{\psi}\,d(\mu\boxtimes\mu)\bigr|<\eta ∫ Φ ε m d ( μ ⊠ μ ) − ∫ F ψ d ( μ ⊠ μ ) < η , whence ∣ ∫ Φ ε m d ( μ ⊠ μ ) ∣ > C ∗ ∥ ∇ ψ ∥ μ \bigl|\int\Phi_{\varepsilon_{m}}\,d(\mu\boxtimes\mu)\bigr|>C_{*}\lVert\nabla\psi\rVert_{\mu} ∫ Φ ε m d ( μ ⊠ μ ) > C ∗ ∥ ∇ ψ ∥ μ , contradicting (6). Hence
∣ ∫ R 2 F ψ d ( μ ⊠ μ ) ∣ ≤ C ∗ ∥ ∇ ψ ∥ μ for every ψ ∈ C c ∞ ( R ) , \Bigl|\int_{\mathbb{R}^{2}}F_{\psi}\,d(\mu\boxtimes\mu)\Bigr|\le C_{*}\,\lVert\nabla\psi\rVert_{\mu}\qquad\text{for every }\psi\in C_{c}^{\infty}(\mathbb{R}), ∫ R 2 F ψ d ( μ ⊠ μ ) ≤ C ∗ ∥ ∇ ψ ∥ μ for every ψ ∈ C c ∞ ( R ) ,
with C ∗ = 2 ( 2 L + 1 ) ≥ 0 C_{*}=2(2L+1)\ge0 C ∗ = 2 ( 2 L + 1 ) ≥ 0 independent of ψ \psi ψ . As μ ∈ P 2 ( R ) \mu\in\mathcal{P}_{2}(\mathbb{R}) μ ∈ P 2 ( R ) , Finite Free Fisher Information, the Free Score and the Free Fisher Information of a Probability Measure on the Real Line §finite gives μ ∈ P 2 Φ ∗ ( R ) \mu\in\mathcal{P}_{2}^{\Phi^{*}}(\mathbb{R}) μ ∈ P 2 Φ ∗ ( R ) . Together with Step 3 this proves the lemma.