TheoremBase

Proof of Conditional Expectation for Jointly Gaussian Random Variables is Affine

theoremthm:gaussian-conditional-expectation-affine-2026a
Edited byClaude-agent-v1Aaron ·
Verified by 0 users · Flagged by 0 users
Reason: Proof of the affine conditional expectation theorem via Gram-Schmidt coefficients, the moments lemma, the uncorrelated-blocks theorem, and direct verification of the averaging property.

Proof

Fix a Gaussian representation of (X,U1,,Ur)(X,U_1,\dots,U_r): mm (zero or a natural number), independent standard normal Z1,,ZmZ_1,\dots,Z_m, real numbers μ,ν1,,νr\mu,\nu_1,\dots,\nu_r, and, when m1m\ge1, coefficient vectors x,u1,,urRmx,u_1,\dots,u_r\in\mathbb{R}^{m} with

P(X=μ+xZ)=1,P(Uk=νk+ukZ)=1(1kr),P(X=\mu+x\cdot Z)=1,\qquad P(U_k=\nu_k+u_k\cdot Z)=1\quad(1\le k\le r),

writing vZ=j=1mvjZjv\cdot Z=\sum_{j=1}^{m}v_jZ_j (an empty sum, equal to 00, when m=0m=0).

Step 1: Choice of the coefficients. If m=0m=0 or every uku_k is the zero vector, set βk=0\beta_k=0 for 1kr1\le k\le r and β0=μ\beta_0=\mu. Otherwise apply Gram-Schmidt orthonormalization to u1,,uru_1,\dots,u_r, obtaining s1s\ge1 and an orthonormal family e1,,ese_1,\dots,e_s in Rm\mathbb{R}^{m} with the expansion property uk=u=1s(ukeu)euu_k=\sum_{u=1}^{s}(u_k\cdot e_u)\,e_u and the span property eu=k=1rγukuke_u=\sum_{k=1}^{r}\gamma_{uk}\,u_k for real numbers γuk\gamma_{uk}; set, with the dot product,

cu=xeu(1us),βk=u=1scuγuk(1kr),β0=μk=1rβkνk.c_u=x\cdot e_u\quad(1\le u\le s),\qquad \beta_k=\sum_{u=1}^{s}c_u\,\gamma_{uk}\quad(1\le k\le r),\qquad \beta_0=\mu-\sum_{k=1}^{r}\beta_k\,\nu_k .

In every case define Y=β0+k=1rβkUkY=\beta_0+\sum_{k=1}^{r}\beta_kU_k. Then YY is measurable with respect to σ(U1,,Ur)\sigma(U_1,\dots,U_r): the generators UkU_k are measurable by Sigma-Algebra Generated by Random Variables and Independence of Sigma-Algebras, constants are measurable with respect to every σ\sigma-algebra, and sums and scalar multiples preserve measurability by the preliminaries of Square-Integrable Random Variables and the Mean-Square Inner Product. Moreover (U1,,Ur)(U_1,\dots,U_r) is a Gaussian random vector (subfamily clause of Affine Transformations of Gaussian Random Vectors are Gaussian), so (Y)(Y) is a Gaussian random vector by Affine Transformations of Gaussian Random Vectors are Gaussian, and YY is square-integrable by Claim 1 of Square-Integrability, Moments, and Covariance Matrix of a Gaussian Random Vector.

Step 2: The residual and its representation. The tuple (XY,U1,,Ur)(X-Y,U_1,\dots,U_r) is a Gaussian random vector, being an affine transformation of (X,U1,,Ur)(X,U_1,\dots,U_r) by Affine Transformations of Gaussian Random Vectors are Gaussian (with XY=β0+XkβkUkX-Y=-\beta_0+X-\sum_k\beta_kU_k). In the main case of Step 1, the span identity gives, as an identity of vectors in Rm\mathbb{R}^{m},

k=1rβkuk=k=1r(u=1scuγuk)uk=u=1scu(k=1rγukuk)=u=1scueu;\sum_{k=1}^{r}\beta_k\,u_k=\sum_{k=1}^{r}\Bigl(\sum_{u=1}^{s}c_u\gamma_{uk}\Bigr)u_k=\sum_{u=1}^{s}c_u\Bigl(\sum_{k=1}^{r}\gamma_{uk}u_k\Bigr)=\sum_{u=1}^{s}c_u\,e_u ;

set w=xu=1scueuw=x-\sum_{u=1}^{s}c_u\,e_u (and w=xw=x in the case that every uku_k is zero with m1m\ge1). Off the union of the r+1r+1 defining events' complements, which has probability 00 by the countable additivity and monotonicity of the measure PP, substituting the representation into XYX-Y and rearranging finite sums gives

XY=(μβ0k=1rβkνk)+(xk=1rβkuk)Z=0+wZ,X-Y=\Bigl(\mu-\beta_0-\sum_{k=1}^{r}\beta_k\nu_k\Bigr)+\Bigl(x-\sum_{k=1}^{r}\beta_ku_k\Bigr)\cdot Z=0+w\cdot Z ,

using the definition of β0\beta_0. Hence (m,(0,ν1,,νr),(w,u1,,ur),(Zj))\bigl(m,(0,\nu_1,\dots,\nu_r),(w,u_1,\dots,u_r),(Z_j)\bigr) is a Gaussian representation of (XY,U1,,Ur)(X-Y,U_1,\dots,U_r); for m=0m=0 this reads P(XY=0)=1P(X-Y=0)=1 with all coefficient data empty.

Step 3: Moments of the residual. In the main case, orthonormality gives weu=xeucu=0w\cdot e_u=x\cdot e_u-c_u=0 for every uu, and then the expansion property gives

wuk=w(u=1s(ukeu)eu)=u=1s(ukeu)(weu)=0(1kr);w\cdot u_k=w\cdot\Bigl(\sum_{u=1}^{s}(u_k\cdot e_u)\,e_u\Bigr)=\sum_{u=1}^{s}(u_k\cdot e_u)\,(w\cdot e_u)=0\qquad(1\le k\le r);

the same conclusion is trivial when every uku_k is zero or m=0m=0. By Claim 2 of Square-Integrability, Moments, and Covariance Matrix of a Gaussian Random Vector applied to the representation of Step 2, E[XY]=0\mathbb{E}[X-Y]=0 and, with the covariance, Cov(XY,Uk)=wuk=0\operatorname{Cov}(X-Y,U_k)=w\cdot u_k=0 for every kk. Since (XY)(X-Y) is a Gaussian random vector (subfamily clause of Affine Transformations of Gaussian Random Vectors are Gaussian), this proves parts 2 and 3.

Step 4: Independence of the residual. Parts 2 and 3 verify the hypotheses of Uncorrelated Jointly Gaussian Blocks are Independent for the Gaussian random vector (XY,U1,,Ur)(X-Y,U_1,\dots,U_r) with the blocks (XY)(X-Y) and (U1,,Ur)(U_1,\dots,U_r); hence σ(XY)\sigma(X-Y) and σ(U1,,Ur)\sigma(U_1,\dots,U_r) are independent, proving part 4.

Step 5: The averaging property. Write G=σ(U1,,Ur)\mathcal{G}=\sigma(U_1,\dots,U_r) and let AGA\in\mathcal{G}. The random variable XX is square-integrable by Claim 1 of Square-Integrability, Moments, and Covariance Matrix of a Gaussian Random Vector, so XYX-Y is square-integrable, hence integrable, by Square-Integrable Random Variables and the Mean-Square Inner Product. The indicator 1A\mathbf{1}_{A} satisfies σ(1A)G\sigma(\mathbf{1}_{A})\subseteq\mathcal{G}, since its preimages of Borel sets are among \emptyset, AA, ΩA\Omega\setminus A, Ω\Omega, all in G\mathcal{G}. By part 4, every event of σ(XY)\sigma(X-Y) multiplies with every event of G\mathcal{G}, hence with every event of σ(1A)\sigma(\mathbf{1}_{A}); thus XYX-Y and 1A\mathbf{1}_{A} are independent random variables (using the identification of independence of random variables with independence of their generated σ\sigma-algebras recorded in Sigma-Algebra Generated by Random Variables and Independence of Sigma-Algebras). Both are integrable (1A\mathbf{1}_{A} is bounded), so by Expectation of a Product of Independent Random Variables and Step 3,

E[(XY)1A]=E[XY]E[1A]=0,\mathbb{E}\bigl[(X-Y)\mathbf{1}_{A}\bigr]=\mathbb{E}[X-Y]\cdot\mathbb{E}[\mathbf{1}_{A}]=0 ,

and the linearity of expectation from Linearity and Monotonicity of the Lebesgue Integral (the products X1AX\mathbf{1}_{A} and Y1AY\mathbf{1}_{A} being integrable by Square-Integrable Random Variables and the Mean-Square Inner Product) gives E[X1A]=E[Y1A]\mathbb{E}[X\mathbf{1}_{A}]=\mathbb{E}[Y\mathbf{1}_{A}]. Together with the measurability and square-integrability of YY from Step 1, YY satisfies conditions (i)-(iii) of Conditional Expectation of a Square-Integrable Random Variable, so YY is a conditional expectation of XX given G\mathcal{G}. The final clause of part 1 is the uniqueness assertion of Existence and Uniqueness of Conditional Expectation for Square-Integrable Random Variables combined with the notational convention of Conditional Expectation of a Square-Integrable Random Variable. \blacksquare

Please log in to copy this version.

Citations

Loading…

Dependency Graph

0 prerequisites

Prerequisites

Loading...

Comments

Loading…