Reason: Proof of the LQG-value identification: instantiation of the linear-Gaussian LQG chain with the fluctuation LQG data (cost convention V-circ = V/2, Riccati and filter-covariance identifications via uniqueness), evaluation of the separation theorem's optimal value V* at mean-zero initial law with trace-to-entrywise conversion, and the constructive realizability of the noise coefficients from the jump representation of the aggregate fluctuation covariance.
Proof
Throughout, (H1)--(H4) and conclusion 1'', conclusion 2'' refer to the approximate Kalman filter and policy lemma, whose setting and hypotheses are in force; the linear-Gaussian state-observation model, its data (A∘,ε∘,E~∘,ε~∘,ξ,W∘), the control dimension m, the control matrix assignment B(t)=Bt, and the assignments Q∘(t)=Qt, V∘(t)=21Vt, R∘(t)=Rt, F∘=F^ are those of the statement. Products of matrices are matrix products, (⋅)⊤ is the transpose, and integrals of matrix-valued maps are entrywise Riemann integrals of continuous integrands, existing by continuity, unless a Lebesgue integral is indicated.
Step 1: the cost data. By conclusion 1, all entries of t↦Et, t↦Bt, t↦E~t, t↦Θt⋆, t↦Θ~t⋆, t↦Qt, t↦Vt, t↦Rt, and t↦Rt−1 are continuous on [0,T], every Rt is symmetric positive definite, and, with the fluctuation Hessian coefficientsHij(t) and Fγδ of the data,
Exchanging γ and δ leaves these sums unchanged, so Qtγδ=Qtδγ and F^γδ=F^δγ: every Qt and F^ is symmetric. Each entry of V∘=21V is the product of the constant function 21 with a continuous function, hence continuous by the sum and product rules for continuous real-valued functions. Therefore Q∘ and R∘ assign symmetric matrices with continuous entries, V∘ assigns real matrices with l rows and m columns with continuous entries, and F∘ is symmetric: (Q∘,V∘,R∘,F∘) are cost data for the model and the control dimension k=m, with every R∘(t)=Rt positive definite; and B assigns real matrices with l rows and m columns with continuous entries, as the controlled-system definition requires.
Step 2: model covariances and initial covariance. By the model definition, Θ(t)=ε∘(t)ε∘(t)⊤ and Θ~(t)=ε~∘(t)ε~∘(t)⊤ --- the model's noise matrices, written ε and ε~ in the model definition, being the present ε∘ and ε~∘ --- so hypothesis (ii) of the statement gives Θ(t)=Θt⋆ and Θ~(t)=Θ~t⋆ for every t∈[0,T]. By claim 1 of the Kalman--Bucy filter theorem, the initial covariance matrix of the model is P0KB=(Cov(ξγ,ξδ))1≤γ,δ≤l (written P0 in that theorem) with the covariance; by hypothesis (iii) this matrix is Π0.
Step 3: Z is a symmetric continuous solution of the backward Riccati equation. By (H2), every Zt is symmetric, ZT=F^, and there are continuous functions z˙γδ:[0,T]→R with
Since V∘(s)=21Vs, we have ZsB(s)+V∘(s)=ZsBs+21Vs=Ws for every s∈[0,T]; and A∘(s)=Es, Q∘(s)=Qs, R∘(s)=Rs, F∘=F^ (hypothesis (i) and the statement). So, entrywise, the display above is exactly the backward Riccati equation
of the completion-of-squares theorem for the present data. Thus Z is a symmetric continuous solution of that equation, with ZtB(t)+V∘(t)=Wt: the first part of conclusion 1 of the corollary.
Step 4: the covariance assignment of the Kalman--Bucy filter is Π. By claim 1 of the Kalman--Bucy filter theorem, formed for the present model, the assignment D(t)=E~∘(t)⊤Θ~(t)−1E~∘(t) has continuous entries, every D(t) is symmetric positive semidefinite, and there is exactly one assignment ΠKB of real matrices with l rows and l columns to the points of [0,T], with continuous entries, such that
this ΠKB is the covariance assignment of that claim. By Step 2 and hypothesis (i), D(t)=E~t⊤(Θ~t⋆)−1E~t=D~t, A∘(r)=Er, Θ(r)=Θr⋆, and P0KB=Π0; so the displayed equation is the same integral equation, with the same coefficient assignments and the same initial matrix, that the filter covariance Π of conclusion 2 satisfies. These data meet the hypotheses of the global existence and uniqueness theorem for the Kalman covariance Riccati equation on [0,T]: the entries of t↦Et, t↦Θt⋆, and t↦D~t are continuous and every Θt⋆ and every D~t is positive semidefinite, by conclusions 1 and 2, and Π0 is positive semidefinite by (H4). Both ΠKB and Π are assignments with continuous entries satisfying that equation, so the uniqueness assertion of that theorem gives ΠKB(t)=Πt for every t∈[0,T]: the remaining part of conclusion 1.
Step 5: the optimal value in trace form. Steps 1 and 3 verify all hypotheses of the separation theorem for the present model, the control dimension m, the control matrix assignment B, the cost data (Q∘,V∘,R∘,F∘), and the solution Z of the backward Riccati equation. Its optimal value is
with the trace, the expectation, and E[ξ]=(E[ξ1],…,E[ξl]), a Riemann integral of a continuous integrand as recorded there. By hypothesis (iii) every component of E[ξ] is 0; hence every component of the matrix-vector productZ0E[ξ] --- a sum of products each having a factor 0 --- is 0, and the dot productE[ξ]⋅(Z0E[ξ]) is 0. Substituting P0KB=Π0 and Θ(t)=Θt⋆ (Step 2), ZtB(t)+V∘(t)=Wt (Step 3), R∘(t)=Rt, and ΠKB(t)=Πt (Step 4):
Step 6: traces as entrywise dot products, and the Lebesgue form. For real matrices M and N with l rows and l columns, claim 4 of the basic properties of the trace gives
the matrix WtRt−1Wt⊤ is symmetric, whence tr(WtRt−1Wt⊤Πt)=tr((WtRt−1Wt⊤)⊤Πt)=(WtRt−1Wt⊤)⋅Πt. Hence
V∗=Z0⋅Π0+∫0T(Zt⋅Θt⋆+(WtRt−1Wt⊤)⋅Πt)dt,
and the integrand here, agreeing at every t∈[0,T] with the continuous integrand of the display of Step 5, is continuous on [0,T]. Its Riemann integral over [0,T]agrees with its Lebesgue integral over the compact interval[0,T], so
Step 7: the minimum over extended admissible controls. The hypotheses of the extended separation theorem are exactly those verified in Steps 1 and 3. By its claim 2, J[α]≥V∗ for every extended admissible controlα with values in Rm; by its claim 3, the closed-loop feedback control α∗ of the closed-loop feedback lemma is admissible, hence extended admissible, and J[α∗]=V∗. Therefore V∗ is the minimum of J over all extended admissible controls with values in Rm, attained by α∗. Together with Steps 5 and 6 this proves conclusion 2 of the corollary up to its final sentence.
Step 8: the limit identity. Assume finally that for each natural number N≥1 a driving system and a projected solution along the approximate Kalman policy hN are fixed as in the cost-limit proposition, and that the initial-condition hypotheses (I1)--(I2) there hold. All hypotheses of the cost-limit proposition are then in force --- its setting and (H1)--(H4) are those adopted here --- so its limit identity holds, and by Steps 5--7 its right-hand side equals V∗=minJ. Hence
N→∞lim(N(JN[hN]−JMF)+γ=1∑lP0γζNγ)=minJ,
which is the final sentence of conclusion 2.
Step 9: realizability of the coefficients (conclusion 3). Set m∘=l⋅l+l~ and index the m∘ columns by the l⋅l ordered pairs (σ,γ)∈{1,…,l}2 followed by the l~ channel indices υ′∈{1,…,l~}. Write eγ (γ∈{1,…,l}) for the γ-th standard basis vector of Euclidean spaceRl, identified with a one-column matrix as in the jump representation lemma; and for a real x≥0 write x for the unique nonnegative real number with (x)2=x, which exists by the existence and uniqueness of the nonnegative square root. Since St lies in the probability simplexΔl, every Stσ≥0; every β(σ,γ,St,At)≥0 by clause 1 of the definition of a transition-rate family; and every b~υ′(St)≥r~>0 by (H3). Define, for t∈[0,T]: the matrix ε∘(t), with l rows and m∘ columns, whose column with index an ordered pair (σ,γ) with σ=γ is Stσβ(σ,γ,St,At)(eγ−eσ), and whose remaining columns --- those with index a pair (σ,σ) and the last l~ --- are zero; and the matrix ε~∘(t), with l~ rows and m∘ columns, whose first l⋅l columns are zero and whose entry in row υ and the column with channel index υ′ is 1{υ=υ′}b~υ′(St), with the indicator notation of the LQG data definition.
By the definitions of the matrix product and the transpose, for matrices M and N with m∘ columns each, the entry of MN⊤ in row p and column q is ∑c=1m∘MpcNqc: a sum over the columns, the column c contributing the product of its p-th entry in M and its q-th entry in N. Every nonzero column of ε∘(t) has index among the first l⋅l and every nonzero column of ε~∘(t) has index among the last l~, so in ε∘(t)ε~∘(t)⊤ every term of every entry has a factor 0: ε∘(t)ε~∘(t)⊤=0, the zero matrix with l rows and l~ columns. In ε∘(t)ε∘(t)⊤, the column with index (σ,γ), σ=γ, contributes to the entry in row p and column q the term Stσβ(σ,γ,St,At)((eγ−eσ)(eγ−eσ)⊤)pq --- the square of the square root being its argument --- and the remaining columns contribute 0; hence, summing over the ordered pairs and applying clause 1 (the jump representation) of the jump representation and positive semidefiniteness of the aggregate fluctuation covariance at the point (St,At)∈Δl×Rm,
the last equality by clause 6 of the LQG data definition. In ε~∘(t)ε~∘(t)⊤, the entry in row υ and column υ′′ is ∑υ′=1l~1{υ=υ′}1{υ′′=υ′}b~υ′(St)=1{υ=υ′′}b~υ(St), which is the corresponding entry of Θ~t⋆ by clause 7 of the LQG data definition: ε~∘(t)ε~∘(t)⊤=Θ~t⋆.
For the continuity of the entries: every entry of t↦ε∘(t) and of t↦ε~∘(t) is either constantly 0 or of the form ±g for the nonnegative function g(t)=Stσβ(σ,γ,St,At), respectively g(t)=b~υ′(St). The map t↦b~υ′(St) is continuous by conclusion 1. The components of t↦(St,At) are continuous by clause 1 of the definition of a mean-field trajectory pair (part of the stationary triple setting); on Δl×Rm the rate β(σ,γ,⋅,⋅) agrees with the restriction of βˉ(σ,γ,⋅,⋅) by clause 1 of the definition of the transition-rate extension, and βˉ(σ,γ,⋅,⋅) is a C1 map on U×Rm by clause 2 of that definition, hence continuous; so t↦β(σ,γ,St,At) is continuous on [0,T] by continuity of compositions along the continuous map t↦(St,At), applied pointwise on [0,T] exactly as in the definition of the fluctuation linear-quadratic cost, and t↦Stσβ(σ,γ,St,At) is continuous by the sum and product rules for continuous real-valued functions. Finally, the nonnegative square root preserves continuity. For reals 0≤u≤v one has u≤v, and for 0≤u<v one has u<v: otherwise u≥v≥0, respectively u>v≥0, would give u=(u)2≥(v)2=v, respectively u>v. For reals x,y≥0, ∣x−y∣≤x+y, both square roots being nonnegative, so
(x−y)2≤x−y(x+y)=(x)2−(y)2=∣x−y∣,
the middle equality by expanding the product of the difference and the sum of x and y; since ∣x−y∣ is the nonnegative square root of its own square, the monotonicity above yields ∣x−y∣≤∣x−y∣. Hence if g:[0,T]→R is continuous with g≥0 and t0∈[0,T], then for every real η>0 there is a δ>0 such that ∣g(t)−g(t0)∣<η2 for all t∈[0,T] with ∣t−t0∣<δ, and then ∣g(t)−g(t0)∣≤∣g(t)−g(t0)∣<η by the strict monotonicity above: the map t↦g(t) is continuous on [0,T]. Therefore all entries of ε∘ and ε~∘ are continuous on [0,T]; and since every Θ~t⋆ is symmetric positive definite by conclusion 1, the displayed identities show that these assignments satisfy hypothesis (ii) of the statement and conditions (i) and (ii) of the model definition. This proves conclusion 3. ■