Throughout, λ, B(R), Bl, λl and the insertion maps Ψi are those of Finite Products of Lebesgue Measure and Coordinate Integration on Rl; write κ=λl⊗μ. Expectations are those of Expectation, Variance, and Moments, and integrals those of Lebesgue Integral of a Nonnegative Measurable Function and Integrable Function and the Lebesgue Integral. For l=1 we use throughout the conventions of Finite Products of Lebesgue Measure and Coordinate Integration on Rl: a pair (θ′,y) reads as y, λl−1⊗μ as μ, and Ψ1(t,y)=(t,y); every step below then applies verbatim, the case i=j in Step 4(b) being vacuous.
Step 1 (change of variables). (Θ,D) is measurable and Q is its image measure, and by assumption (i), Q is the measure with density p with respect to κ. Hence claims 2 and 3 of Image Measures, Measures with Densities, and Change of Variables give: for every Bl⊗G-measurable g:Rl×Y→[0,∞],
∫Ωg∘(Θ,D)dP=∫gdQ=∫gpdκin [0,∞];
and for measurable g:Rl×Y→R, the random variable g∘(Θ,D) is integrable with respect to P if and only if gp is integrable with respect to κ, in which case E[g(Θ,D)]=∫gpdκ. Taking g≡1: ∫pdκ=Q(Rl×Y)=P(Ω)=1.
Step 2 (a one-dimensional vanishing lemma). Let g:R→R be differentiable at every point with g and its derivative g′ continuous, and let g and g′ be integrable with respect to λ. (They are measurable: for continuous u:R→R and real a, each point of {u>a} has an open interval around it inside the set, by continuity, so {u>a} is open and hence lies in the Borel σ-algebra.) Then:
(i) ∫Rg′dλ=0.
(ii) If moreover t↦tg(t) and t↦tg′(t) are integrable with respect to λ, then ∫Rtg′(t)dλ(t)=−∫Rgdλ.
Proof of (i). For real a<b, the function g~(s)=g(a+s) on [0,b−a] has difference quotients at s coinciding with those of g at a+s, so g~ is continuous and differentiable at every point of [0,b−a] with continuous derivative g~′(s)=g′(a+s), i.e. g~∈C1([0,b−a]), and the Fundamental Theorem of Calculus Fundamental Theorem of Calculus in One Dimension gives g(b)−g(a)=∫0b−ag′(a+s)ds, read as the Riemann integral of the continuous integrand g~′, which exists by Continuous Functions on a Closed Interval are Riemann Integrable. By Agreement of the Riemann and Lebesgue Integrals for Continuous Functions on a Closed Interval this equals the Lebesgue integral ∫Rg′(a+s)1[0,b−a](s)dλ(s), which equals ∫Rg′1[a,b]dλ by claim 2 of Translation Invariance of Lebesgue Measure and the Lebesgue Integral applied to the positive and negative parts of g′1[a,b] (whose translates by a are the positive and negative parts of g′(a+⋅)1[0,b−a]; each part is bounded — continuous functions on a compact interval are bounded, Continuous Real-Valued Functions on a Compact Interval are Bounded — and supported in an interval of finite measure, hence of finite integral, so the two nonnegative identities may be subtracted). Thus
g(b)−g(a)=∫Rg′1[a,b]dλ(a<b).(∗)
As n→∞ through the natural numbers, g′1[0,n]→g′1[0,∞) pointwise, dominated by the integrable ∣g′∣, so Dominated Convergence Theorem and (∗) give g(n)→c+:=g(0)+∫g′1[0,∞)dλ. Moreover, by (∗) and monotonicity, supt∈[n,n+1]∣g(t)−g(n)∣≤∫∣g′∣1[n,∞)dλ→0 (dominated convergence again), so g(t)→c+ as t→∞. If c+=0, there is R0>0 with ∣g(t)∣≥∣c+∣/2 for all t≥R0; then for every natural n>R0, monotonicity, the integral of simple functions, and Existence of Lebesgue Measure on the Real Line give ∫∣g∣dλ≥(∣c+∣/2)λ([R0,n])=(∣c+∣/2)(n−R0)→∞, contradicting integrability of g; so c+=0. Symmetrically, g(−n)=g(0)−∫g′1[−n,0]dλ converges to a limit c−, g(t)→c− as t→−∞, and c−=0. Finally, by dominated convergence and (∗),
∫Rg′dλ=nlim∫g′1[−n,n]dλ=nlim(g(n)−g(−n))=c+−c−=0.
Proof of (ii). h(t)=tg(t) is differentiable at every point with h′(t)=g(t)+tg′(t) by the product rule of Sum and Product Rules for One-Dimensional Derivatives and Continuity (the map t↦t has derivative 1 from the definition of the derivative); h and h′ are continuous, h is integrable by assumption, and h′ is integrable as a sum of integrable functions (Linearity and Monotonicity of the Lebesgue Integral). Part (i) applied to h gives ∫(g(t)+tg′(t))dλ(t)=0, and linearity gives (ii).
Step 3 (slices of the density). Fix i∈{1,…,l}. For (θ′,y)∈Rl−1×Y and t∈R put gθ′,y(t)=p(Ψi(t,(θ′,y))). The map t↦(θ1′,…,θi−1′,t,θi′,…,θl−1′) from R to Rl is continuous, its coordinate functions being constant or the identity (Coordinatewise Characterization of Continuity for Euclidean Maps). By assumption (ii) and Slice Function and the Partial Derivative, applied to the C1 function p(⋅,y) on Rl at each inserted point in the ith variable, gθ′,y is strictly positive and differentiable at every t∈R with
gθ′,y′(t)=∂ip(Ψi(t,(θ′,y))).
(The cited lemma yields, at each fixed t0, an interval I around t0 on which its slice function coincides with gθ′,y, and equates the slice derivative at the interior point t0 with the partial derivative; differentiability at t0 and the value of the derivative depend only on the restriction to I, so the identity holds at every t0∈R.) Both gθ′,y and gθ′,y′ are continuous, being compositions of the continuous insertion with the continuous functions p(⋅,y) and ∂ip(⋅,y) (Composition of Continuous Euclidean Maps; continuity of these is part of the C1 property).
Step 4 (the score identities). Fix i,j∈{1,…,l} and let δij=1 if i=j and δij=0 otherwise. We show
E[mj(D)Si]=0andE[ΘjSi]=−δij.
We use three elementary facts on a measure space, from Lebesgue Integral of a Nonnegative Measurable Function, Simple Function and Its Integral, and Linearity and Monotonicity of the Lebesgue Integral: (F1) a [0,∞]-valued measurable u with finite integral is finite off a set of measure zero (on N={u=∞} one has u≥L1N for every L, so Lm(N)≤∫udm for every L); (F2) a [0,∞]-valued measurable u vanishing off a set of measure zero has ∫udm=0 (every simple s with 0≤s≤u is bounded by a multiple of the indicator of that set, so ∫sdm=0; take the supremum); (F3) consequently, integrable real functions agreeing off a set of measure zero have equal integrals, and likewise [0,∞]-valued measurable functions (split by the exceptional set and use additivity with (F2)).
(a) E[mj(D)Si]=0. The function G(θ,y)=mj(y)∂ip(θ,y)/p(θ,y) is Bl⊗G-measurable: (θ,y)↦mj(y) is measurable (the preimage of a Borel set A is the measurable rectangle Rl×mj−1(A)), ∂ip/p is measurable as in assumption (iv), and products of real-valued measurable functions are measurable by Sequentially Continuous Functions of Measurable Euclidean Maps are Measurable (the map (u,v)↦uv is continuous on R2). The random variable G(Θ,D)=mj(D)Si is integrable, being a product of the square-integrable random variables mj(D) and Si (Square-Integrable Random Variables and the Mean-Square Inner Product). By Step 1, Gp is κ-integrable and
E[mj(D)Si]=∫Gpdκ=∫mj(y)∂ip(θ,y)dκ(θ,y),
the second equality holding pointwise because p>0 everywhere (assumption (ii)). Apply claim 4 of Finite Products of Lebesgue Measure and Coordinate Integration on Rl (coordinate Fubini at coordinate i) to the κ-integrable (θ,y)↦mj(y)∂ip(θ,y): there is a set N1 of measure zero off which t↦mj(y)gθ′,y′(t) is λ-integrable (Step 3 identifies the integrand), and ∫mj∂ipdκ equals the integral of the function F given off N1 by F(θ′,y)=∫Rmj(y)gθ′,y′(t)dλ(t) and by 0 on N1. Now apply claim 3 of Finite Products of Lebesgue Measure and Coordinate Integration on Rl (coordinate Tonelli) to the nonnegative measurable functions p and ∣∂ip∣: since ∫pdκ=1<∞ (Step 1) and ∫∣∂ip∣dκ<∞ (assumption (iii)), the measurable [0,∞]-valued functions (θ′,y)↦∫Rgθ′,ydλ and (θ′,y)↦∫R∣gθ′,y′∣dλ have finite integrals, hence by (F1) are finite off sets N2, N3 of measure zero. For (θ′,y)∈/N1∪N2∪N3, the function gθ′,y satisfies all hypotheses of Step 2(i) (continuity and differentiability from Step 3, integrability off N2∪N3), so ∫Rgθ′,y′dλ=0 and, pulling out the constant mj(y) by linearity (licensed off N3, where gθ′,y′ is λ-integrable), F(θ′,y)=0. Thus F vanishes off a set of measure zero, and E[mj(D)Si]=∫Fd(λl−1⊗μ)=0 by (F3).
(b) E[ΘjSi]=−δij. Take G(θ,y)=θj∂ip(θ,y)/p(θ,y); the coordinate map (θ,y)↦θj is measurable (preimages are measurable rectangles), G(Θ,D)=ΘjSi is integrable as a product of square-integrable random variables, and Step 1 gives
E[ΘjSi]=∫θj∂ip(θ,y)dκ(θ,y),
the integrand being κ-integrable also directly from assumption (iii), since ∣θj∂ip∣≤(1+∑j′∣θj′∣)∣∂ip∣ pointwise.
Case i=j. The jth coordinate of Ψi(t,(θ′,y)) does not depend on t: it equals θj′ if j<i and θj−1′ if j>i; call it θ(j)′. Coordinate Fubini applied to θj∂ip represents the integral through inner integrals ∫Rθ(j)′gθ′,y′(t)dλ(t)=θ(j)′∫Rgθ′,y′dλ=0 off a set of measure zero, exactly as in (a) with the constant θ(j)′ in place of mj(y). Hence E[ΘjSi]=0.
Case i=j. Note that the ith coordinate of Ψi(t,(θ′,y)) is exactly t. Coordinate Fubini applied to the κ-integrable θi∂ip represents E[ΘiSi] as the integral of the function F equal, off a set N1′ of measure zero, to F(θ′,y)=∫Rtgθ′,y′(t)dλ(t) and to 0 on N1′. Two further applications of coordinate Tonelli give sets of measure zero off which ∫R∣t∣gθ′,y(t)dλ(t)<∞ and ∫R∣t∣∣gθ′,y′(t)∣dλ(t)<∞: the first because ∫∣θi∣pdκ=E[∣Θi∣]<∞ by Step 1 (square-integrable random variables are integrable, Square-Integrable Random Variables and the Mean-Square Inner Product), the second from assumption (iii) with the weight ∣θi∣, both followed by (F1). Off the union of all these sets and N2, N3 of part (a), Step 2(ii) applies to gθ′,y and gives
F(θ′,y)=∫Rtgθ′,y′(t)dλ(t)=−∫Rgθ′,ydλ.
Let G0(θ′,y)=∫Rgθ′,ydλ, the [0,∞]-valued measurable function of coordinate Tonelli applied to p, with ∫G0d(λl−1⊗μ)=∫pdκ=1. Then F is integrable (claim 4), F+ vanishes off a set of measure zero, and F−=G0 off a set of measure zero; so by (F2) and (F3), ∫Fd(λl−1⊗μ)=−∫G0d(λl−1⊗μ)=−1. Hence E[ΘiSi]=−1.
Combining (a) and (b) with linearity of the expectation (Linearity and Monotonicity of the Lebesgue Integral),
E[(mj(D)−Θj)Si]=δij(1≤i,j≤l).(∗∗)
Step 5 (Cauchy-Schwarz step). Each mj(D)−Θj is square-integrable (Square-Integrable Random Variables and the Mean-Square Inner Product), so every product (mi(D)−Θi)(mj(D)−Θj) is integrable, R is well defined, and Rij=Rji by commutativity of pointwise multiplication. J is symmetric positive definite by assumption (iv), so J−1 exists and is symmetric positive definite by Invertibility of Symmetric Positive Definite Matrices. We use the dot product and the matrix-vector product on Rl; componentwise, a⋅(Ma)=∑i,jaiMijaj for any l×l matrix M; A(Bx)=(AB)x for l×l matrices, since (A(Bx))i=∑jAij∑mBjmxm=∑m(AB)imxm with the matrix product; and Ix=x for the identity matrix I of Inverse Matrix and Invertible Real Square Matrix, since (Ix)i=∑jIijxj=xi.
Fix a∈Rl and set b=J−1a, and
X=j=1∑laj(mj(D)−Θj),W=i=1∑lbiSi,
both square-integrable as linear combinations of square-integrable random variables. Expanding the products and using linearity of the expectation:
E[X2]=i,j∑aiajRij=a⋅(Ra);E[XW]=i,j∑ajbiE[(mj(D)−Θj)Si]=j∑ajbj=a⋅(J−1a)
by (∗∗); and
E[W2]=i,i′∑bibi′Jii′=b⋅(Jb)=b⋅a=a⋅(J−1a),
since Jb=J(J−1a)=(JJ−1)a=Ia=a by the identities above and the definition of the inverse, and the dot product is symmetric. Since (X−W)2≥0 pointwise, monotonicity and linearity of the expectation give
0≤E[(X−W)2]=E[X2]−2E[XW]+E[W2]=a⋅(Ra)−a⋅(J−1a)=a⋅((R−J−1)a),
the last step by componentwise bilinearity. As a was arbitrary and R−J−1 is symmetric, R−J−1 is positive semidefinite; that is, R⪰J−1 in the semidefinite order. ■