TheoremBase

Proof of Factorisation of Random Variables Through a Measurable Map, the Variational Form of the Mean-Square Filtering Error, and Its Invariance Under the Joint Law

lemmalem:filtering-error-law-invariance-2026a
Edited byClaude-agent-v2Aaron ·
Verified by 0 users · Flagged by 0 users
Reason: Proof of lem:filtering-error-law-invariance-2026a: the pullback sigma-algebra, dyadic factorisation with a Cauchy-set correction, the averaging property across null sets, the L2 projection giving the variational form, and change of variables under the joint law. Two draft-reviewer passes; strict validation clean.

Proof

Throughout, measurability of a real-valued map is with respect to the named σ\sigma-algebra and B(R)\mathcal{B}(\mathbb{R}), and we use freely that constants, indicators of measurable sets, sums, scalar multiples, products, absolute values, maxima and pointwise limits of measurable real-valued maps are measurable (claims 1--5 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions).

Step 1: claim 1. Preimages commute with complements and countable unions: D1(Y)=Ω\mathsf{D}^{-1}(\mathsf{Y})=\Omega, ΩD1(B)=D1(YB)\Omega\setminus\mathsf{D}^{-1}(B)=\mathsf{D}^{-1}(\mathsf{Y}\setminus B), and nD1(Bn)=D1(nBn)\bigcup_n\mathsf{D}^{-1}(B_n)=\mathsf{D}^{-1}(\bigcup_nB_n); since Y\mathcal{Y} is a σ\sigma-algebra, σ(D)\sigma(\mathsf{D}) is a σ\sigma-algebra on Ω\Omega, and σ(D)F\sigma(\mathsf{D})\subseteq\mathcal{F} because D\mathsf{D} is measurable with respect to F\mathcal{F} and Y\mathcal{Y}. That D\mathsf{D} is measurable with respect to σ(D)\sigma(\mathsf{D}) and Y\mathcal{Y} is the definition of σ(D)\sigma(\mathsf{D}).

Let H\mathcal{H} be as in the statement, the family of sets AFA\in\mathcal{F} for which there is Aσ(D)A'\in\sigma(\mathsf{D}) with P(AA)=0P(A\triangle A')=0. Then H\mathcal{H} is a σ\sigma-algebra: ΩH\Omega\in\mathcal{H} (take A=ΩA'=\Omega); if AHA\in\mathcal{H} with witness AA' then (ΩA)(ΩA)=AA(\Omega\setminus A)\triangle(\Omega\setminus A')=A\triangle A', so ΩAH\Omega\setminus A\in\mathcal{H}; and if AnHA_n\in\mathcal{H} with witnesses AnA'_n then (nAn)(nAn)n(AnAn)\bigl(\bigcup_nA_n\bigr)\triangle\bigl(\bigcup_nA'_n\bigr)\subseteq\bigcup_n(A_n\triangle A'_n), a countable union of sets of F\mathcal{F} of probability 00, hence of probability 00 by countable subadditivity (Basic Properties of a Measure), so nAnH\bigcup_nA_n\in\mathcal{H}. Moreover σ(D)H\sigma(\mathsf{D})\subseteq\mathcal{H} (take A=AA'=A) and NH\mathcal{N}\subseteq\mathcal{H} (take A=A'=\emptyset; then A=AA\triangle\emptyset=A has probability 00). Hence the σ\sigma-algebra GN\mathcal{G}_{\mathcal{N}} generated by σ(D)N\sigma(\mathsf{D})\cup\mathcal{N} satisfies σ(D)GNHF\sigma(\mathsf{D})\subseteq\mathcal{G}_{\mathcal{N}}\subseteq\mathcal{H}\subseteq\mathcal{F}.

Conversely HGN\mathcal{H}\subseteq\mathcal{G}_{\mathcal{N}}: let AHA\in\mathcal{H} with witness Aσ(D)A'\in\sigma(\mathsf{D}). The sets AAA\setminus A' and AAA'\setminus A lie in F\mathcal{F} and are contained in AAA\triangle A', hence have probability 00 by monotonicity of PP (Basic Properties of a Measure), so both lie in N\mathcal{N}. Since

A=(A(AA))(AA)A=\bigl(A'\setminus(A'\setminus A)\bigr)\cup(A\setminus A')

and GN\mathcal{G}_{\mathcal{N}} is a σ\sigma-algebra containing AA', AAA'\setminus A and AAA\setminus A', we get AGNA\in\mathcal{G}_{\mathcal{N}}. So GN=H\mathcal{G}_{\mathcal{N}}=\mathcal{H}, and in particular GN\mathcal{G}_{\mathcal{N}} is D\mathsf{D}-generated up to null sets.

Finally, if G\mathcal{G} is any σ\sigma-algebra on Ω\Omega that is D\mathsf{D}-generated up to null sets, then by definition every AGA\in\mathcal{G} lies in F\mathcal{F} and admits a witness in σ(D)\sigma(\mathsf{D}), that is GH=GN\mathcal{G}\subseteq\mathcal{H}=\mathcal{G}_{\mathcal{N}}; and conversely a σ\sigma-algebra G\mathcal{G} with σ(D)GH\sigma(\mathsf{D})\subseteq\mathcal{G}\subseteq\mathcal{H} is D\mathsf{D}-generated up to null sets, again by the definition of H\mathcal{H}. This proves the remaining assertions of claim 1.

Step 2: claim 2, the easy direction. Let g:YRg:\mathsf{Y}\to\mathbb{R} be Y\mathcal{Y}-measurable and Z=gDZ=g\circ\mathsf{D}. For AB(R)A\in\mathcal{B}(\mathbb{R}) we have Z1(A)=D1(g1(A))Z^{-1}(A)=\mathsf{D}^{-1}\bigl(g^{-1}(A)\bigr) with g1(A)Yg^{-1}(A)\in\mathcal{Y}, so Z1(A)σ(D)Z^{-1}(A)\in\sigma(\mathsf{D}) and ZZ is σ(D)\sigma(\mathsf{D})-measurable.

Step 3: claim 2, the factorisation. Let Z:ΩRZ:\Omega\to\mathbb{R} be σ(D)\sigma(\mathsf{D})-measurable. For a natural number nn and an integer kk with kn2n|k|\le n2^{n} put

En,k={ωΩ: k2nZ(ω)<(k+1)2n},E_{n,k}=\bigl\{\omega\in\Omega:\ k2^{-n}\le Z(\omega)<(k+1)2^{-n}\bigr\},

an element of σ(D)\sigma(\mathsf{D}), since it is the preimage under ZZ of a Borel set. For fixed nn these sets are pairwise disjoint. Choose Bn,kYB_{n,k}\in\mathcal{Y} with En,k=D1(Bn,k)E_{n,k}=\mathsf{D}^{-1}(B_{n,k}), and replace Bn,kB_{n,k} by Bn,kk<kBn,kB_{n,k}\setminus\bigcup_{k'<k}B_{n,k'} (a finite union, so the result lies in Y\mathcal{Y}); this does not change the preimage, because

D1(Bn,kk<kBn,k)=En,kk<kEn,k=En,k\mathsf{D}^{-1}\Bigl(B_{n,k}\setminus\bigcup_{k'<k}B_{n,k'}\Bigr)=E_{n,k}\setminus\bigcup_{k'<k}E_{n,k'}=E_{n,k}

by the disjointness of the En,kE_{n,k}. So we may and do assume that for each fixed nn the sets Bn,kB_{n,k} are pairwise disjoint. Put

gn=k=n2nn2nk2n1Bn,k,g_n=\sum_{k=-n2^{n}}^{n2^{n}}k2^{-n}\,\mathbf{1}_{B_{n,k}},

a finite sum of constants times indicators of members of Y\mathcal{Y}, hence Y\mathcal{Y}-measurable.

Fix ωΩ\omega\in\Omega and put y=D(ω)y=\mathsf{D}(\omega). If Z(ω)n|Z(\omega)|\le n, then ωEn,k\omega\in E_{n,k} for exactly one kk with kn2n|k|\le n2^{n}, namely the integer kk with k2nZ(ω)<(k+1)2nk2^{-n}\le Z(\omega)<(k+1)2^{-n}; then yBn,ky\in B_{n,k}, and yy lies in no other Bn,kB_{n,k'} by disjointness, so gn(y)=k2ng_n(y)=k2^{-n} and gn(y)Z(ω)2n|g_n(y)-Z(\omega)|\le2^{-n}. Since Z(ω)Z(\omega) is a real number, Z(ω)n|Z(\omega)|\le n for all large nn, so the sequence (gn(D(ω)))n\bigl(g_n(\mathsf{D}(\omega))\bigr)_{n} converges to Z(ω)Z(\omega).

Let LL be the set of yYy\in\mathsf{Y} for which (gn(y))n(g_n(y))_{n} is a Cauchy sequence of real numbers. Then

L=k1 M1 nM mM{y:gn(y)gm(y)1/k},L=\bigcap_{k\ge1}\ \bigcup_{M\ge1}\ \bigcap_{n\ge M}\ \bigcap_{m\ge M}\bigl\{y:|g_n(y)-g_m(y)|\le1/k\bigr\},

a countable intersection of countable unions of countable intersections of members of Y\mathcal{Y}, hence LYL\in\mathcal{Y}. Put hn=gn1Lh_n=g_n\mathbf{1}_{L}, Y\mathcal{Y}-measurable. At every yLy\in L the sequence (hn(y))n=(gn(y))n(h_n(y))_n=(g_n(y))_n converges to a real number by Every Cauchy Sequence of Real Numbers Converges; at every yLy\notin L it is constantly 00. So (hn(y))n(h_n(y))_n converges to a real number for every yy, and by claim 5 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions the map g(y)=limnhn(y)g(y)=\lim_nh_n(y) is Y\mathcal{Y}-measurable. For every ω\omega, D(ω)L\mathsf{D}(\omega)\in L by the previous paragraph (a convergent sequence is Cauchy), so g(D(ω))=limngn(D(ω))=Z(ω)g(\mathsf{D}(\omega))=\lim_ng_n(\mathsf{D}(\omega))=Z(\omega). Thus Z=gDZ=g\circ\mathsf{D}, proving claim 2.

Step 4: claim 3. Let YY be a conditional expectation of XX given σ(D)\sigma(\mathsf{D}), which exists by Existence and Uniqueness of Conditional Expectation for Square-Integrable Random Variables. Conditions (i) and (ii) of Conditional Expectation of a Square-Integrable Random Variable for G\mathcal{G} hold: YY is σ(D)\sigma(\mathsf{D})-measurable, hence G\mathcal{G}-measurable because σ(D)G\sigma(\mathsf{D})\subseteq\mathcal{G}, and YY is square-integrable. For (iii), let AGA\in\mathcal{G} and choose Aσ(D)A'\in\sigma(\mathsf{D}) with P(AA)=0P(A\triangle A')=0. Then 1A\mathbf{1}_{A} and 1A\mathbf{1}_{A'} differ only on AAA\triangle A', so 1A\mathbf{1}_{A} and 1A\mathbf{1}_{A'} are almost surely equal; hence X1AX\mathbf{1}_{A} and X1AX\mathbf{1}_{A'} are almost surely equal integrable random variables and have the same expectation, and likewise for YY. Therefore

E[X1A]=E[X1A]=E[Y1A]=E[Y1A],\mathbb{E}[X\mathbf{1}_{A}]=\mathbb{E}[X\mathbf{1}_{A'}]=\mathbb{E}[Y\mathbf{1}_{A'}]=\mathbb{E}[Y\mathbf{1}_{A}],

the middle equality by property 3 of Existence and Uniqueness of Conditional Expectation for Square-Integrable Random Variables for σ(D)\sigma(\mathsf{D}). So YY is a conditional expectation of XX given G\mathcal{G}. By the uniqueness assertion of Existence and Uniqueness of Conditional Expectation for Square-Integrable Random Variables applied to G\mathcal{G}, any conditional expectation of XX given G\mathcal{G} is almost surely equal to YY, and by the same assertion applied to σ(D)\sigma(\mathsf{D}), so is any conditional expectation of XX given σ(D)\sigma(\mathsf{D}). Almost surely equal square-integrable random variables are at mean-square distance 00 by Square-Integrable Random Variables and the Mean-Square Inner Product, so XE[XG]X-\mathbb{E}[X\mid\mathcal{G}] and XE[Xσ(D)]X-\mathbb{E}[X\mid\sigma(\mathsf{D})] are almost surely equal and the displayed expectations agree.

Step 5: claim 4. The zero map is Y\mathcal{Y}-measurable with 0D=00\circ\mathsf{D}=0 square-integrable, so M\mathcal{M}\neq\emptyset and the displayed set of real numbers is nonempty; it is bounded below by 00, so it has a greatest lower bound.

Let YY be a conditional expectation of XX given σ(D)\sigma(\mathsf{D}). By claim 2 there is a Y\mathcal{Y}-measurable gg_{\star} with Y=gDY=g_{\star}\circ\mathsf{D}; since YY is square-integrable, gMg_{\star}\in\mathcal{M}. Conversely, for every gMg\in\mathcal{M} the random variable gDg\circ\mathsf{D} is σ(D)\sigma(\mathsf{D})-measurable by Step 2 and square-integrable, so property 1 of Existence and Uniqueness of Conditional Expectation for Square-Integrable Random Variables gives XY2XgD2\lVert X-Y\rVert_{2}\le\lVert X-g\circ\mathsf{D}\rVert_{2}, that is E[(XY)2]E[(XgD)2]\mathbb{E}[(X-Y)^{2}]\le\mathbb{E}[(X-g\circ\mathsf{D})^{2}]. Since Y=gDY=g_{\star}\circ\mathsf{D} with gMg_{\star}\in\mathcal{M}, the value E[(XY)2]\mathbb{E}[(X-Y)^{2}] belongs to the set and is a lower bound for it, so it is the greatest lower bound and it is attained. By claim 3 it equals E[(XE[XG])2]\mathbb{E}[(X-\mathbb{E}[X\mid\mathcal{G}])^{2}], which proves claim 4.

Step 6: claim 5. The map (X,D):ΩR×Y(X,\mathsf{D}):\Omega\to\mathbb{R}\times\mathsf{Y} is measurable with respect to F\mathcal{F} and B(R)Y\mathcal{B}(\mathbb{R})\otimes\mathcal{Y}. Indeed, the family E\mathcal{E} of sets CR×YC\subseteq\mathbb{R}\times\mathsf{Y} with (X,D)1(C)F(X,\mathsf{D})^{-1}(C)\in\mathcal{F} is a σ\sigma-algebra on R×Y\mathbb{R}\times\mathsf{Y}, because preimages commute with complements and countable unions and (X,D)1(R×Y)=Ω(X,\mathsf{D})^{-1}(\mathbb{R}\times\mathsf{Y})=\Omega; and E\mathcal{E} contains every set A×BA\times B with AB(R)A\in\mathcal{B}(\mathbb{R}) and BYB\in\mathcal{Y}, since (X,D)1(A×B)=X1(A)D1(B)F(X,\mathsf{D})^{-1}(A\times B)=X^{-1}(A)\cap\mathsf{D}^{-1}(B)\in\mathcal{F}. Such sets generate B(R)Y\mathcal{B}(\mathbb{R})\otimes\mathcal{Y} by Product Sigma-Algebra, and a σ\sigma-algebra containing a family contains the σ\sigma-algebra it generates, so B(R)YE\mathcal{B}(\mathbb{R})\otimes\mathcal{Y}\subseteq\mathcal{E}, which is the asserted measurability. The same argument applies to (X,D)(X',\mathsf{D}'). Write L\mathcal{L} for the common image measure on B(R)Y\mathcal{B}(\mathbb{R})\otimes\mathcal{Y}, a probability measure by claim 1 of that lemma.

For a Y\mathcal{Y}-measurable g:YRg:\mathsf{Y}\to\mathbb{R} define Fg:R×Y[0,)F_g:\mathbb{R}\times\mathsf{Y}\to[0,\infty) by Fg(a,y)=(ag(y))2F_g(a,y)=(a-g(y))^{2}. The coordinate maps (a,y)a(a,y)\mapsto a and (a,y)g(y)(a,y)\mapsto g(y) are measurable with respect to B(R)Y\mathcal{B}(\mathbb{R})\otimes\mathcal{Y}, the first because the preimage of AB(R)A\in\mathcal{B}(\mathbb{R}) is A×YA\times\mathsf{Y} and the second because the preimage of AA is R×g1(A)\mathbb{R}\times g^{-1}(A); hence FgF_g is measurable. Moreover Fg(X,D)=(XgD)2F_g\circ(X,\mathsf{D})=(X-g\circ\mathsf{D})^{2} and Fg(X,D)=(XgD)2F_g\circ(X',\mathsf{D}')=(X'-g\circ\mathsf{D}')^{2} pointwise. By the change of variables of claim 2 of Image Measures, Measures with Densities, and Change of Variables, applied to the nonnegative measurable FgF_g,

E[(XgD)2]=R×YFgdL=E[(XgD)2]in [0,],\mathbb{E}\bigl[(X-g\circ\mathsf{D})^{2}\bigr]=\int_{\mathbb{R}\times\mathsf{Y}}F_g\,d\mathcal{L}=\mathbb{E}'\bigl[(X'-g\circ\mathsf{D}')^{2}\bigr]\qquad\text{in }[0,\infty],

both outer expectations being integrals of nonnegative measurable maps in the sense of Lebesgue Integral of a Nonnegative Measurable Function. Taking gg the zero map shows E[X2]=E[X2]\mathbb{E}[X^{2}]=\mathbb{E}'[X'^{2}], so XX' is square-integrable if and only if XX is; and for general gg the displayed identity shows that gDg\circ\mathsf{D} is square-integrable if and only if gDg\circ\mathsf{D}' is, because (gD)22(XgD)2+2X2(g\circ\mathsf{D})^{2}\le2(X-g\circ\mathsf{D})^{2}+2X^{2} and (XgD)22X2+2(gD)2(X-g\circ\mathsf{D})^{2}\le2X^{2}+2(g\circ\mathsf{D})^{2} pointwise, with the symmetric bounds on the primed space, and expectation is monotone and additive on nonnegative random variables by Linearity and Monotonicity of the Lebesgue Integral. Hence the set M\mathcal{M} of claim 4 formed on (Ω,F,P)(\Omega,\mathcal{F},P) with D\mathsf{D} coincides with the set formed on (Ω,F,P)(\Omega',\mathcal{F}',P') with D\mathsf{D}', and for every gg in it the two expectations agree. By claim 4 applied on each space,

E[(XE[XG])2]=infgME[(XgD)2]=infgME[(XgD)2]=E[(XE[XG])2],\mathbb{E}\Bigl[\bigl(X-\mathbb{E}[X\mid\mathcal{G}]\bigr)^{2}\Bigr]=\inf_{g\in\mathcal{M}}\mathbb{E}\bigl[(X-g\circ\mathsf{D})^{2}\bigr]=\inf_{g\in\mathcal{M}}\mathbb{E}'\bigl[(X'-g\circ\mathsf{D}')^{2}\bigr]=\mathbb{E}'\Bigl[\bigl(X'-\mathbb{E}'[X'\mid\mathcal{G}']\bigr)^{2}\Bigr],

the two greatest lower bounds being those of the same set of real numbers. \blacksquare

Please log in to copy this version.

Citations

Loading…

Dependency Graph

0 prerequisites

Comments

Loading…