TheoremBase

Proof

Fix tt as in the statement and abbreviate R=RtR=R_t, W=WtW=W_t, G=R−1W⊤\mathsf{G}=R^{-1}W^{\top} (a real m×lm\times l matrix), s=st\mathfrak{s}=\mathfrak{s}_t, a=at\mathfrak{a}=\mathfrak{a}_t, s^=s^t\hat{\mathfrak{s}}=\hat{\mathfrak{s}}_t, ε=εt\varepsilon=\varepsilon_t, and u=ut=a+Gsu=u_t=\mathfrak{a}+\mathsf{G}\mathfrak{s}. The sans-serif G\mathsf{G} is distinct from the terminal cost function GG of the adopted setting, and the plain RR from the cost remainder RNR_N there; the letter MM is kept for the generic matrices of the entry notation.

Step 1 (Square-integrability). Each sγ\mathfrak{s}^\gamma is a random variable: Σtγ\Sigma^\gamma_t is one by the solution definition, and sγ=N(Σtγ−Stγ)\mathfrak{s}^\gamma=\sqrt{N}(\Sigma^\gamma_t-S^\gamma_t) is a sequentially continuous (affine) function of it, measurable by measurability of sequentially continuous functions of measurable maps. Moreover (sγ)2≤2N(\mathfrak{s}^\gamma)^2\le2N at every point of Ω\Omega: first, (sγ)2≤∑γ′(sγ′)2=∣s∣2(\mathfrak{s}^\gamma)^2\le\sum_{\gamma'}(\mathfrak{s}^{\gamma'})^2=|\mathfrak{s}|^2 by claim 1 of the norm properties lemma, and ∣s∣2=N d(Σt,St)2|\mathfrak{s}|^2=N\,d(\Sigma_t,S_t)^2 by claims 5 and 2 of the same lemma. Second, all coordinates of a point xx of the probability simplex lie in [0,1][0,1]: they are nonnegative, and each is at most the total sum 11, the remaining summands being nonnegative. For s,s′∈[0,1]s,s'\in[0,1] we have (s−s′)2=s s−2 s s′+s′ s′≤s s+s′ s′≤s+s′(s-s')^2=s\,s-2\,s\,s'+s'\,s'\le s\,s+s'\,s'\le s+s', since 0≤s s′0\le s\,s' (either a factor is 00, or both are positive and claim 5 of the order arithmetic lemma applies) and since r r≤rr\,r\le r for r∈[0,1]r\in[0,1] (for r=0r=0 or r=1r=1 this holds with equality; for 0<r<10<r<1, claim 10 of the same lemma gives r r<r⋅1=rr\,r<r\cdot1=r). Summing over coordinates, any x,y∈Δlx,y\in\Delta^l satisfy d(x,y)2=∑γ(xγ−yγ)2≤∑γ(xγ+yγ)=2d(x,y)^2=\sum_\gamma(x^\gamma-y^\gamma)^2\le\sum_\gamma(x^\gamma+y^\gamma)=2. Hence E[(sγ)2]≤2N<∞\mathbb{E}[(\mathfrak{s}^\gamma)^2]\le2N<\infty by monotonicity of the integral, and sγ\mathfrak{s}^\gamma is square-integrable. Each s^γ\hat{\mathfrak{s}}^\gamma is Gt\mathcal{G}_t-measurable and square-integrable by conditions (i)--(ii) of the conditional-expectation definition, and each εγ=sγ−s^γ\varepsilon^\gamma=\mathfrak{s}^\gamma-\hat{\mathfrak{s}}^\gamma is square-integrable by the closure properties of the square-integrability definition. Define tuples X=(X1,…,Xm)X=(X^1,\dots,X^m) and X^=(X^1,…,X^m)\hat{X}=(\hat{X}^1,\dots,\hat{X}^m) by

Xj=−∑γ=1lGjγ sγ,X^j=−∑γ=1lGjγ s^γ(j∈{1,…,m}),X^j=-\sum_{\gamma=1}^{l}\mathsf{G}^{j\gamma}\,\mathfrak{s}^\gamma,\qquad \hat{X}^j=-\sum_{\gamma=1}^{l}\mathsf{G}^{j\gamma}\,\hat{\mathfrak{s}}^\gamma\qquad(j\in\{1,\dots,m\}),

so that X=−GsX=-\mathsf{G}\mathfrak{s} and X^=−Gs^\hat{X}=-\mathsf{G}\hat{\mathfrak{s}} componentwise in the entry notation. Each XjX^j and each X^j\hat{X}^j is a finite linear combination of square-integrable random variables, hence square-integrable by the closure properties.

Step 2 (X^\hat{X} is a conditional expectation of XX). By claim 1 (linearity) of the basic properties of conditional expectation, applied repeatedly --- induction on the number of summands, the claim being stated for two --- each X^j\hat{X}^j, a finite linear combination of the conditional expectations s^γ\hat{\mathfrak{s}}^\gamma with coefficients −Gjγ-\mathsf{G}^{j\gamma}, is a conditional expectation of the corresponding combination XjX^j of the sγ\mathfrak{s}^\gamma given Gt\mathcal{G}_t.

Step 3 (Admissibility of Y=aY=\mathfrak{a}). Each aj\mathfrak{a}^j is square-integrable by hypothesis, and by conclusion 3 of the adaptedness lemma it is almost surely equal to a Gt\mathcal{G}_t-measurable random variable, which is square-integrable with the same mean-square norm by conclusions 2--3 of that lemma. Thus the tuple Y=(a1,…,am)Y=(\mathfrak{a}^1,\dots,\mathfrak{a}^m) is admissible in the sense of the conditional mean-square optimality lemma: each component is square-integrable and almost surely equal to a Gt\mathcal{G}_t-measurable square-integrable random variable.

Step 4 (Optimality). The matrix RR is symmetric positive definite by conclusion (a) of the completion-of-squares theorem under (H1); in particular it is positive semidefinite, the quadratic form being 00 at the origin and positive elsewhere. By claim 2 of the optimality lemma, applied with k=mk=m, the tuple XX, the conditional expectations X^\hat{X} of Step 2, and the admissible tuple YY of Step 3 (the tuple written ε\varepsilon in that lemma being our X−X^X-\hat{X}, not our ε\varepsilon),

E[(Y−X)⋅(R (Y−X))] ≥ E[(X−X^)⋅(R (X−X^))],\mathbb{E}\big[(Y-X)\cdot\big(R\,(Y-X)\big)\big]\ \ge\ \mathbb{E}\big[(X-\hat{X})\cdot\big(R\,(X-\hat{X})\big)\big],

both sides being defined and finite by the integrability of such forms established in that lemma's preamble (the components of Y−XY-X and X−X^X-\hat{X} being square-integrable by the closure properties). Componentwise, Y−X=a+Gs=uY-X=\mathfrak{a}+\mathsf{G}\mathfrak{s}=u and X−X^=−G(s−s^)=−GεX-\hat{X}=-\mathsf{G}(\mathfrak{s}-\hat{\mathfrak{s}})=-\mathsf{G}\varepsilon.

Step 5 (Matrix algebra). At every point of Ω\Omega, expanding (Gε)j=∑γGjγεγ(\mathsf{G}\varepsilon)^j=\sum_\gamma \mathsf{G}^{j\gamma}\varepsilon^\gamma and rearranging the finite sums (the two sign factors cancelling),

(−Gε)⋅(R (−Gε))=∑j,kRjk(Gε)j(Gε)k=∑γ,δ(∑j,kGjγRjkGkδ)εγεδ=ε⋅((G⊤RG) ε),(-\mathsf{G}\varepsilon)\cdot\big(R\,(-\mathsf{G}\varepsilon)\big)=\sum_{j,k}R^{jk}(\mathsf{G}\varepsilon)^j(\mathsf{G}\varepsilon)^k=\sum_{\gamma,\delta}\Big(\sum_{j,k}\mathsf{G}^{j\gamma}R^{jk}\mathsf{G}^{k\delta}\Big)\varepsilon^\gamma\varepsilon^\delta=\varepsilon\cdot\big((\mathsf{G}^{\top}R\mathsf{G})\,\varepsilon\big),

since (G⊤RG)γδ=∑j,k(G⊤)γjRjkGkδ=∑j,kGjγRjkGkδ(\mathsf{G}^{\top}R\mathsf{G})^{\gamma\delta}=\sum_{j,k}(\mathsf{G}^{\top})^{\gamma j}R^{jk}\mathsf{G}^{k\delta}=\sum_{j,k}\mathsf{G}^{j\gamma}R^{jk}\mathsf{G}^{k\delta} in the entry notation. Moreover G⊤=WR−1\mathsf{G}^{\top}=WR^{-1}: entrywise, (G⊤)γj=Gjγ=∑k(R−1)jk(W⊤)kγ=∑k(R−1)kjWγk=(WR−1)γj(\mathsf{G}^{\top})^{\gamma j}=\mathsf{G}^{j\gamma}=\sum_{k}(R^{-1})^{jk}(W^{\top})^{k\gamma}=\sum_{k}(R^{-1})^{kj}W^{\gamma k}=(WR^{-1})^{\gamma j}, using the symmetry of R−1R^{-1} (inverse of a positive definite matrix). Hence, by associativity of the matrix product and the defining property R−1R R−1=R−1R^{-1}R\,R^{-1}=R^{-1} of the inverse,

G⊤RG=WR−1 R R−1W⊤=W R−1 W⊤.\mathsf{G}^{\top}R\mathsf{G}=WR^{-1}\,R\,R^{-1}W^{\top}=W\,R^{-1}\,W^{\top} .

Combining with Step 4 yields conclusion 1:

E[u⋅Ru] ≥ E[ε⋅(WR−1W⊤) ε].\mathbb{E}\big[u\cdot Ru\big]\ \ge\ \mathbb{E}\big[\varepsilon\cdot(WR^{-1}W^{\top})\,\varepsilon\big].

Step 6 (Trace form). Claim 1 of the expected bilinear forms lemma, applied to the tuple ε\varepsilon (components square-integrable by Step 1) and the matrix WR−1W⊤WR^{-1}W^{\top}, gives

E[ε⋅(WR−1W⊤ε)]=∑γ=1l∑δ=1l(WR−1W⊤)γδ E[εγεδ].\mathbb{E}\big[\varepsilon\cdot(WR^{-1}W^{\top}\varepsilon)\big]=\sum_{\gamma=1}^{l}\sum_{\delta=1}^{l}\big(WR^{-1}W^{\top}\big)^{\gamma\delta}\,\mathbb{E}\big[\varepsilon^\gamma\varepsilon^\delta\big].

Step 7 (Independence of the choice). Let s^′γ\hat{\mathfrak{s}}'^{\gamma} be any other conditional expectations of the sγ\mathfrak{s}^\gamma given Gt\mathcal{G}_t, and set ε′γ=sγ−s^′γ\varepsilon'^{\gamma}=\mathfrak{s}^\gamma-\hat{\mathfrak{s}}'^{\gamma}, square-integrable as in Step 1. By the uniqueness part of the existence and uniqueness theorem, s^′γ\hat{\mathfrak{s}}'^{\gamma} is almost surely equal to s^γ\hat{\mathfrak{s}}^{\gamma} for each γ\gamma. Hence ε′γ−εγ=s^γ−s^′γ\varepsilon'^{\gamma}-\varepsilon^{\gamma}=\hat{\mathfrak{s}}^{\gamma}-\hat{\mathfrak{s}}'^{\gamma} is a square-integrable random variable almost surely equal to 00, so its mean-square norm vanishes, ∥ε′γ−εγ∥2=0\lVert\varepsilon'^{\gamma}-\varepsilon^{\gamma}\rVert_2=0, by the null-equivalence property recorded in the square-integrability definition. Fix γ,δ\gamma,\delta. All products below are integrable, as products of square-integrable random variables (closure properties of the same definition), and by linearity of the integral,

E[ε′γε′δ]−E[εγεδ]=E[ε′γ (ε′δ−εδ)]+E[(ε′γ−εγ) εδ].\mathbb{E}\big[\varepsilon'^{\gamma}\varepsilon'^{\delta}\big]-\mathbb{E}\big[\varepsilon^{\gamma}\varepsilon^{\delta}\big]=\mathbb{E}\big[\varepsilon'^{\gamma}\,(\varepsilon'^{\delta}-\varepsilon^{\delta})\big]+\mathbb{E}\big[(\varepsilon'^{\gamma}-\varepsilon^{\gamma})\,\varepsilon^{\delta}\big].

By the mean-square Cauchy-Schwarz inequality, the first term satisfies ∣E[ε′γ(ε′δ−εδ)]∣≤∥ε′γ∥2 ∥ε′δ−εδ∥2=0|\mathbb{E}[\varepsilon'^{\gamma}(\varepsilon'^{\delta}-\varepsilon^{\delta})]|\le\lVert\varepsilon'^{\gamma}\rVert_2\,\lVert\varepsilon'^{\delta}-\varepsilon^{\delta}\rVert_2=0, and likewise ∣E[(ε′γ−εγ)εδ]∣≤∥ε′γ−εγ∥2 ∥εδ∥2=0|\mathbb{E}[(\varepsilon'^{\gamma}-\varepsilon^{\gamma})\varepsilon^{\delta}]|\le\lVert\varepsilon'^{\gamma}-\varepsilon^{\gamma}\rVert_2\,\lVert\varepsilon^{\delta}\rVert_2=0. Hence E[ε′γε′δ]=E[εγεδ]\mathbb{E}[\varepsilon'^{\gamma}\varepsilon'^{\delta}]=\mathbb{E}[\varepsilon^{\gamma}\varepsilon^{\delta}], which proves conclusion 2. ■\blacksquare

Citations

Loading…

Dependencies

Uses0

Loading…

Comments

Log in to comment.

Loading…