Step 0 (An almost surely convergent subsequence). Choose a strictly increasing sequence ( k j ) j β N (k_j)_{j\in\mathbb{N}} ( k j β ) j β N β such that for every j j j and every i β { 1 , β¦ , p } i\in\{1,\dots,p\} i β { 1 , β¦ , p } ,
P ( β£ X i k j β X i β£ > 2 β j ) β€ 2 β j ; P\bigl(|X^{k_j}_i-X_i|>2^{-j}\bigr)\le 2^{-j}; P ( β£ X i k j β β β X i β β£ > 2 β j ) β€ 2 β j ;
this is possible because, for each of the finitely many i i i , convergence in probability provides a threshold beyond which P ( β£ X i k β X i β£ β₯ 2 β j ) β€ 2 β j P(|X^{k}_i-X_i|\ge2^{-j})\le2^{-j} P ( β£ X i k β β X i β β£ β₯ 2 β j ) β€ 2 β j , and { β£ X i k β X i β£ > 2 β j } β { β£ X i k β X i β£ β₯ 2 β j } \{|X^k_i-X_i|>2^{-j}\}\subseteq\{|X^k_i-X_i|\ge2^{-j}\} { β£ X i k β β X i β β£ > 2 β j } β { β£ X i k β β X i β β£ β₯ 2 β j } . For fixed i i i , the events A j i = { β£ X i k j β X i β£ > 2 β j } A^i_j=\{|X^{k_j}_i-X_i|>2^{-j}\} A j i β = { β£ X i k j β β β X i β β£ > 2 β j } satisfy β j P ( A j i ) < β \sum_j P(A^i_j)<\infty β j β P ( A j i β ) < β , so by the first Borel-Cantelli lemma P ( limβsup β‘ j A j i ) = 0 P(\limsup_j A^i_j)=0 P ( lim sup j β A j i β ) = 0 . Off limβsup β‘ j A j i \limsup_j A^i_j lim sup j β A j i β , one has β£ X i k j β X i β£ β€ 2 β j |X^{k_j}_i-X_i|\le2^{-j} β£ X i k j β β β X i β β£ β€ 2 β j for all large j j j , hence X i k j β X i X^{k_j}_i\to X_i X i k j β β β X i β in the sense of Limit of a Sequence of Real Numbers . Let Ξ© 0 = β i = 1 p ( Ξ© β limβsup β‘ j A j i ) \Omega_0=\bigcap_{i=1}^{p}(\Omega\setminus\limsup_j A^i_j) Ξ© 0 β = β i = 1 p β ( Ξ© β lim sup j β A j i β ) ; then P ( Ξ© 0 ) = 1 P(\Omega_0)=1 P ( Ξ© 0 β ) = 1 (finite intersection of events of probability one) and on Ξ© 0 \Omega_0 Ξ© 0 β all p p p convergences hold.
Step 1 (Factorization over bounded continuous functions). Let f 1 , β¦ , f p : R β R f_1,\dots,f_p:\mathbb{R}\to\mathbb{R} f 1 β , β¦ , f p β : R β R be bounded continuous functions, say β£ f i β£ β€ C i |f_i|\le C_i β£ f i β β£ β€ C i β . Continuous functions are measurable for the Borel Ο \sigma Ο -algebra (preimages of open sets are open, and open sets generate the Borel Ο \sigma Ο -algebra), so each f i ( X i k ) f_i(X^k_i) f i β ( X i k β ) is a random variable with Ο ( f i ( X i k ) ) β Ο ( X i k ) \sigma(f_i(X^k_i))\subseteq\sigma(X^k_i) Ο ( f i β ( X i k β )) β Ο ( X i k β ) , since ( f i β X i k ) β 1 ( B ) = ( X i k ) β 1 ( f i β 1 ( B ) ) (f_i\circ X^k_i)^{-1}(B)=(X^k_i)^{-1}(f_i^{-1}(B)) ( f i β β X i k β ) β 1 ( B ) = ( X i k β ) β 1 ( f i β 1 β ( B )) (elementary set algebra) with f i β 1 ( B ) f_i^{-1}(B) f i β 1 β ( B ) Borel, and Ο ( X i k ) = { ( X i k ) β 1 ( B β² ) : B β² Β Borel } \sigma(X^k_i)=\{(X^k_i)^{-1}(B'):B'\ \text{Borel}\} Ο ( X i k β ) = {( X i k β ) β 1 ( B β² ) : B β² Β Borel } by Sigma-Algebra Generated by Random Variables and Independence of Sigma-Algebras .
Fix k k k and abbreviate Y i = f i ( X i k ) Y_i=f_i(X^k_i) Y i β = f i β ( X i k β ) . We claim E [ Y 1 β― Y p ] = E [ Y 1 ] β― E [ Y p ] \mathbb{E}[Y_1\cdots Y_p]=\mathbb{E}[Y_1]\cdots\mathbb{E}[Y_p] E [ Y 1 β β― Y p β ] = E [ Y 1 β ] β― E [ Y p β ] , by induction on p p p . For p = 1 p=1 p = 1 the claim is trivial. Bounded random variables are integrable on a probability space (dominated by an integrable constant, using monotonicity from Linearity and Monotonicity of the Lebesgue Integral ). For the induction step with p β₯ 2 p\ge2 p β₯ 2 , the grouping lemma applied to the independent family ( X 1 k , β¦ , X p k ) (X^k_1,\dots,X^k_p) ( X 1 k β , β¦ , X p k β ) with the nonempty disjoint blocks { 1 } \{1\} { 1 } and { 2 , β¦ , p } \{2,\dots,p\} { 2 , β¦ , p } shows that Y 1 Y_1 Y 1 β (which is Ο ( X 1 k ) \sigma(X^k_1) Ο ( X 1 k β ) -measurable) and Y 2 β― Y p Y_2\cdots Y_p Y 2 β β― Y p β (which is Ο ( X 2 k , β¦ , X p k ) \sigma(X^k_2,\dots,X^k_p) Ο ( X 2 k β , β¦ , X p k β ) -measurable, products of measurable functions being measurable by the argument recorded in the preliminaries of Square-Integrable Random Variables and the Mean-Square Inner Product ) are independent . Both are bounded, hence integrable, so Expectation of a Product of Independent Random Variables gives E [ Y 1 ( Y 2 β― Y p ) ] = E [ Y 1 ] β E [ Y 2 β― Y p ] \mathbb{E}[Y_1(Y_2\cdots Y_p)]=\mathbb{E}[Y_1]\,\mathbb{E}[Y_2\cdots Y_p] E [ Y 1 β ( Y 2 β β― Y p β )] = E [ Y 1 β ] E [ Y 2 β β― Y p β ] , and the induction hypothesis finishes the claim.
Now pass to the limit along ( k j ) (k_j) ( k j β ) . Define
g ~ j = 1 Ξ© 0 β i = 1 p f i ( X i k j ) + 1 Ξ© β Ξ© 0 β i = 1 p f i ( X i ) , \tilde g_j=\mathbf{1}_{\Omega_0}\prod_{i=1}^{p}f_i(X^{k_j}_i)+\mathbf{1}_{\Omega\setminus\Omega_0}\prod_{i=1}^{p}f_i(X_i), g ~ β j β = 1 Ξ© 0 β β i = 1 β p β f i β ( X i k j β β ) + 1 Ξ© β Ξ© 0 β β i = 1 β p β f i β ( X i β ) ,
a random variable with β£ g ~ j β£ β€ C 1 β― C p |\tilde g_j|\le C_1\cdots C_p β£ g ~ β j β β£ β€ C 1 β β― C p β everywhere. On Ξ© 0 \Omega_0 Ξ© 0 β , continuity of each f i f_i f i β gives f i ( X i k j ) β f i ( X i ) f_i(X^{k_j}_i)\to f_i(X_i) f i β ( X i k j β β ) β f i β ( X i β ) , and products of convergent real sequences converge to the product of the limits; off Ξ© 0 \Omega_0 Ξ© 0 β the sequence is constant. Hence g ~ j β β i f i ( X i ) \tilde g_j\to\prod_i f_i(X_i) g ~ β j β β β i β f i β ( X i β ) at every point of Ξ© \Omega Ξ© , and the dominated convergence theorem (with the constant dominating function) yields E [ g ~ j ] β E [ β i f i ( X i ) ] \mathbb{E}[\tilde g_j]\to\mathbb{E}[\prod_i f_i(X_i)] E [ g ~ β j β ] β E [ β i β f i β ( X i β )] . Moreover β£ E [ g ~ j ] β E [ β i f i ( X i k j ) ] β£ β€ 2 C 1 β― C p β P ( Ξ© β Ξ© 0 ) = 0 \bigl|\mathbb{E}[\tilde g_j]-\mathbb{E}[\prod_i f_i(X^{k_j}_i)]\bigr|\le 2C_1\cdots C_p\,P(\Omega\setminus\Omega_0)=0 β E [ g ~ β j β ] β E [ β i β f i β ( X i k j β β )] β β€ 2 C 1 β β― C p β P ( Ξ© β Ξ© 0 β ) = 0 . The same argument applied to a single factor gives E [ f i ( X i k j ) ] β E [ f i ( X i ) ] \mathbb{E}[f_i(X^{k_j}_i)]\to\mathbb{E}[f_i(X_i)] E [ f i β ( X i k j β β )] β E [ f i β ( X i β )] . Combining with the stage-k j k_j k j β factorization:
E [ β i = 1 p f i ( X i ) ] = lim β‘ j β i = 1 p E [ f i ( X i k j ) ] = β i = 1 p E [ f i ( X i ) ] . \mathbb{E}\Bigl[\prod_{i=1}^{p}f_i(X_i)\Bigr]=\lim_j\prod_{i=1}^{p}\mathbb{E}[f_i(X^{k_j}_i)]=\prod_{i=1}^{p}\mathbb{E}[f_i(X_i)]. E [ i = 1 β p β f i β ( X i β ) ] = j lim β i = 1 β p β E [ f i β ( X i k j β β )] = i = 1 β p β E [ f i β ( X i β )] .
Step 2 (Half-line factorization). Fix a 1 , β¦ , a p β R a_1,\dots,a_p\in\mathbb{R} a 1 β , β¦ , a p β β R . For m β₯ 1 m\ge1 m β₯ 1 let f a m ( x ) = 1 f^m_{a}(x)=1 f a m β ( x ) = 1 for x β€ a x\le a x β€ a , f a m ( x ) = 1 β m ( x β a ) f^m_a(x)=1-m(x-a) f a m β ( x ) = 1 β m ( x β a ) for a < x < a + 1 / m a<x<a+1/m a < x < a + 1/ m , and f a m ( x ) = 0 f^m_a(x)=0 f a m β ( x ) = 0 for x β₯ a + 1 / m x\ge a+1/m x β₯ a + 1/ m ; each f a m f^m_a f a m β is continuous and bounded by 1 1 1 . As m β β m\to\infty m β β , f a m ( x ) β 1 ( β β , a ] ( x ) f^m_a(x)\to\mathbf{1}_{(-\infty,a]}(x) f a m β ( x ) β 1 ( β β , a ] β ( x ) at every x x x . Applying Step 1 with f i = f a i m f_i=f^m_{a_i} f i β = f a i β m β and letting m β β m\to\infty m β β with the dominated convergence theorem (dominating constant 1 1 1 ; the product of the indicators is the indicator of the intersection) gives
P ( X 1 β€ a 1 , β¦ , X p β€ a p ) = β i = 1 p P ( X i β€ a i ) forΒ allΒ a 1 , β¦ , a p β R . (*) P\bigl(X_1\le a_1,\dots,X_p\le a_p\bigr)=\prod_{i=1}^{p}P(X_i\le a_i)\qquad\text{for all }a_1,\dots,a_p\in\mathbb{R}. \tag{*} P ( X 1 β β€ a 1 β , β¦ , X p β β€ a p β ) = i = 1 β p β P ( X i β β€ a i β ) forΒ allΒ a 1 β , β¦ , a p β β R . ( * )
Step 3 (From half-lines to Borel sets). We show by induction on r β { 0 , 1 , β¦ , p } r\in\{0,1,\dots,p\} r β { 0 , 1 , β¦ , p } the claim C r C_r C r β : for all Borel sets B 1 , β¦ , B r B_1,\dots,B_r B 1 β , β¦ , B r β and all reals a r + 1 , β¦ , a p a_{r+1},\dots,a_p a r + 1 β , β¦ , a p β ,
P ( β i β€ r { X i β B i } β© β i > r { X i β€ a i } ) = β i β€ r P ( X i β B i ) β i > r P ( X i β€ a i ) . P\Bigl(\bigcap_{i\le r}\{X_i\in B_i\}\cap\bigcap_{i>r}\{X_i\le a_i\}\Bigr)=\prod_{i\le r}P(X_i\in B_i)\prod_{i>r}P(X_i\le a_i). P ( i β€ r β β { X i β β B i β } β© i > r β β { X i β β€ a i β } ) = i β€ r β β P ( X i β β B i β ) i > r β β P ( X i β β€ a i β ) .
C 0 C_0 C 0 β is (*). Assume C r β 1 C_{r-1} C r β 1 β and fix Borel B 1 , β¦ , B r β 1 B_1,\dots,B_{r-1} B 1 β , β¦ , B r β 1 β and reals a r + 1 , β¦ , a p a_{r+1},\dots,a_p a r + 1 β , β¦ , a p β . Let
Ξ = { B β B ( R ) : P ( { X r β B } β© D ) = P ( X r β B ) β
c } , \Lambda=\Bigl\{B\in\mathcal{B}(\mathbb{R}):P\Bigl(\{X_r\in B\}\cap D\Bigr)=P(X_r\in B)\cdot c\Bigr\}, Ξ = { B β B ( R ) : P ( { X r β β B } β© D ) = P ( X r β β B ) β
c } ,
where D = β i < r { X i β B i } β© β i > r { X i β€ a i } D=\bigcap_{i<r}\{X_i\in B_i\}\cap\bigcap_{i>r}\{X_i\le a_i\} D = β i < r β { X i β β B i β } β© β i > r β { X i β β€ a i β } and c = β i < r P ( X i β B i ) β i > r P ( X i β€ a i ) c=\prod_{i<r}P(X_i\in B_i)\prod_{i>r}P(X_i\le a_i) c = β i < r β P ( X i β β B i β ) β i > r β P ( X i β β€ a i β ) . By C r β 1 C_{r-1} C r β 1 β , Ξ \Lambda Ξ contains every half-line ( β β , a ] (-\infty,a] ( β β , a ] ; it contains R \mathbb{R} R by letting a β β a\to\infty a β β along ( β β , n ] β R (-\infty,n]\uparrow\mathbb{R} ( β β , n ] β R and using continuity of P P P from below (countable additivity) on both sides, which yields P ( D ) = c P(D)=c P ( D ) = c . Both B β¦ P ( { X r β B } β© D ) B\mapsto P(\{X_r\in B\}\cap D) B β¦ P ({ X r β β B } β© D ) and B β¦ P ( X r β B ) β c B\mapsto P(X_r\in B)\,c B β¦ P ( X r β β B ) c are finitely additive and continuous from below in B B B , so Ξ \Lambda Ξ is closed under proper differences and increasing countable unions; together with R β Ξ \mathbb{R}\in\Lambda R β Ξ this makes Ξ \Lambda Ξ a Ξ» \lambda Ξ» -system in the sense of Dynkin's Pi-Lambda Theorem . The half-lines form a Ο \pi Ο -system generating B ( R ) \mathcal{B}(\mathbb{R}) B ( R ) (every open interval is a countable combination of half-lines, ( a , b ) = β n β₯ 1 ( ( β β , b β ( b β a ) / ( 2 n ) ] β ( β β , a ] ) (a,b)=\bigcup_{n\ge1}\bigl((-\infty,b-(b-a)/(2n)]\setminus(-\infty,a]\bigr) ( a , b ) = β n β₯ 1 β ( ( β β , b β ( b β a ) / ( 2 n )] β ( β β , a ] ) ; every open set is a countable union of open intervals with rational endpoints by the density of the rationals ; and open sets generate the Borel Ο \sigma Ο -algebra by Borel Sigma-Algebra on the Real Line ). By Dynkin's theorem , Ξ = B ( R ) \Lambda=\mathcal{B}(\mathbb{R}) Ξ = B ( R ) , which is C r C_r C r β .
Step 4 (Conclusion). C p C_p C p β states that P ( β i = 1 p { X i β B i } ) = β i = 1 p P ( X i β B i ) P\bigl(\bigcap_{i=1}^{p}\{X_i\in B_i\}\bigr)=\prod_{i=1}^{p}P(X_i\in B_i) P ( β i = 1 p β { X i β β B i β } ) = β i = 1 p β P ( X i β β B i β ) for all Borel B 1 , β¦ , B p B_1,\dots,B_p B 1 β , β¦ , B p β . For any nonempty subfamily of indices, choose B i = R B_i=\mathbb{R} B i β = R for the omitted indices; the corresponding factors equal 1 1 1 , so the product identity holds for every subfamily. By Independence of Events and of Random Variables (see also the discussion in Sigma-Algebra Generated by Random Variables and Independence of Sigma-Algebras of inserting Ξ© \Omega Ξ© for omitted factors), the random variables X 1 , β¦ , X p X_1,\dots,X_p X 1 β , β¦ , X p β are independent. β‘ \square β‘