Reason: Proof of the conditional-variance lower bound: apply weighted mean-square optimality of conditional expectation with the observation-adapted control as competitor, then identify the resulting weight as W R^{-1} W^T.
Proof
Fix t as in the statement and abbreviate R=Rt, W=Wt, G=R−1W⊤ (a real m×l matrix), s=st, a=at, s^=s^t, ε=εt, and u=ut=a+Gs. The sans-serif G is distinct from the terminal cost function G of the adopted setting, and the plain R from the cost remainder RN there; the letter M is kept for the generic matrices of the entry notation.
Step 1 (Square-integrability). Each sγ is a random variable: Σtγ is one by the solution definition, and sγ=N(Σtγ−Stγ) is a sequentially continuous (affine) function of it, measurable by measurability of sequentially continuous functions of measurable maps. Moreover (sγ)2≤2N at every point of Ω: first, (sγ)2≤∑γ′(sγ′)2=∣s∣2 by claim 1 of the norm properties lemma, and ∣s∣2=Nd(Σt,St)2 by claims 5 and 2 of the same lemma. Second, all coordinates of a point x of the probability simplex lie in [0,1]: they are nonnegative, and each is at most the total sum 1, the remaining summands being nonnegative. For s,s′∈[0,1] we have (s−s′)2=ss−2ss′+s′s′≤ss+s′s′≤s+s′, since 0≤ss′ (either a factor is 0, or both are positive and claim 5 of the order arithmetic lemma applies) and since rr≤r for r∈[0,1] (for r=0 or r=1 this holds with equality; for 0<r<1, claim 10 of the same lemma gives rr<r⋅1=r). Summing over coordinates, any x,y∈Δl satisfy d(x,y)2=∑γ(xγ−yγ)2≤∑γ(xγ+yγ)=2. Hence E[(sγ)2]≤2N<∞ by monotonicity of the integral, and sγ is square-integrable. Each s^γ is Gt-measurable and square-integrable by conditions (i)--(ii) of the conditional-expectation definition, and each εγ=sγ−s^γ is square-integrable by the closure properties of the square-integrability definition. Define tuples X=(X1,…,Xm) and X^=(X^1,…,X^m) by
Xj=−γ=1∑lGjγsγ,X^j=−γ=1∑lGjγs^γ(j∈{1,…,m}),
so that X=−Gs and X^=−Gs^ componentwise in the entry notation. Each Xj and each X^j is a finite linear combination of square-integrable random variables, hence square-integrable by the closure properties.
Step 2 (X^ is a conditional expectation of X). By claim 1 (linearity) of the basic properties of conditional expectation, applied repeatedly --- induction on the number of summands, the claim being stated for two --- each X^j, a finite linear combination of the conditional expectations s^γ with coefficients −Gjγ, is a conditional expectation of the corresponding combination Xj of the sγ given Gt.
Step 3 (Admissibility of Y=a). Each aj is square-integrable by hypothesis, and by conclusion 3 of the adaptedness lemma it is almost surely equal to a Gt-measurable random variable, which is square-integrable with the same mean-square norm by conclusions 2--3 of that lemma. Thus the tuple Y=(a1,…,am) is admissible in the sense of the conditional mean-square optimality lemma: each component is square-integrable and almost surely equal to a Gt-measurable square-integrable random variable.
Step 4 (Optimality). The matrix R is symmetric positive definite by conclusion (a) of the completion-of-squares theorem under (H1); in particular it is positive semidefinite, the quadratic form being 0 at the origin and positive elsewhere. By claim 2 of the optimality lemma, applied with k=m, the tuple X, the conditional expectations X^ of Step 2, and the admissible tuple Y of Step 3 (the tuple written ε in that lemma being our X−X^, not our ε),
E[(Y−X)⋅(R(Y−X))]≥E[(X−X^)⋅(R(X−X^))],
both sides being defined and finite by the integrability of such forms established in that lemma's preamble (the components of Y−X and X−X^ being square-integrable by the closure properties). Componentwise, Y−X=a+Gs=u and X−X^=−G(s−s^)=−Gε.
Step 5 (Matrix algebra). At every point of Ω, expanding (Gε)j=∑γGjγεγ and rearranging the finite sums (the two sign factors cancelling),
since (G⊤RG)γδ=∑j,k(G⊤)γjRjkGkδ=∑j,kGjγRjkGkδ in the entry notation. Moreover G⊤=WR−1: entrywise, (G⊤)γj=Gjγ=∑k(R−1)jk(W⊤)kγ=∑k(R−1)kjWγk=(WR−1)γj, using the symmetry of R−1 (inverse of a positive definite matrix). Hence, by associativity of the matrix product and the defining property R−1RR−1=R−1 of the inverse,
G⊤RG=WR−1RR−1W⊤=WR−1W⊤.
Combining with Step 4 yields conclusion 1:
E[u⋅Ru]≥E[ε⋅(WR−1W⊤)ε].
Step 6 (Trace form). Claim 1 of the expected bilinear forms lemma, applied to the tuple ε (components square-integrable by Step 1) and the matrix WR−1W⊤, gives
E[ε⋅(WR−1W⊤ε)]=γ=1∑lδ=1∑l(WR−1W⊤)γδE[εγεδ].
Step 7 (Independence of the choice). Let s^′γ be any other conditional expectations of the sγ given Gt, and set ε′γ=sγ−s^′γ, square-integrable as in Step 1. By the uniqueness part of the existence and uniqueness theorem, s^′γ is almost surely equal to s^γ for each γ. Hence ε′γ−εγ=s^γ−s^′γ is a square-integrable random variable almost surely equal to 0, so its mean-square norm vanishes, ∥ε′γ−εγ∥2=0, by the null-equivalence property recorded in the square-integrability definition. Fix γ,δ. All products below are integrable, as products of square-integrable random variables (closure properties of the same definition), and by linearity of the integral,
E[ε′γε′δ]−E[εγεδ]=E[ε′γ(ε′δ−εδ)]+E[(ε′γ−εγ)εδ].
By the mean-square Cauchy-Schwarz inequality, the first term satisfies ∣E[ε′γ(ε′δ−εδ)]∣≤∥ε′γ∥2∥ε′δ−εδ∥2=0, and likewise ∣E[(ε′γ−εγ)εδ]∣≤∥ε′γ−εγ∥2∥εδ∥2=0. Hence E[ε′γε′δ]=E[εγεδ], which proves conclusion 2. ■