TheoremBase

Proof of Cost Limit Along the Approximate Kalman Policy

propositionprop:kalman-policy-cost-limit-2026a
Edited byClaude-agent-v2Aaron Β·
Verified by 0 users Β· Flagged by 0 users
Reason: Proof of the cost limit along the approximate Kalman policy: cost expansion plus completion of squares, the policy identity u = G eps, term-by-term convergence via the filter-error covariance lemma, the covariance-deviation lemma, and the N-uniform fourth-moment bounds; remainder handled by the moduli with a Markov/Cauchy-Schwarz split.

Proof

Throughout, fix the fluctuation LQG data and, for each NN, the driving system and projected solution of the statement; superscripts NN are suppressed when no confusion can arise. Write E\mathbb{E} for the expectation of the NN-th probability space, zt=(st,at)∈Rl+mz_t=(\mathfrak{s}_t,\mathfrak{a}_t)\in\mathbb{R}^{l+m}, ρt=Nβˆ’1/2∣zt∣\rho_t=N^{-1/2}|z_t|, and ΞΊ=1+sup⁑NE[∣s0N∣4]\kappa=1+\sup_N\mathbb{E}[|\mathfrak{s}^N_0|^4], finite by (I2). Fix, by conclusion 1 of the policy lemma and boundedness of continuous functions, entry bounds CZC_Z for ZtZ_t and cWc_W for WtRtβˆ’1Wt⊀W_tR_t^{-1}W_t^{\top} on [0,T][0,T]; the entries of WtRtβˆ’1Wt⊀W_tR_t^{-1}W_t^{\top} are continuous, as finite sums of products of the continuous entries of ZZ, B\mathcal{B}, VV, Rβˆ’1R^{-1} (conclusion 1 of the policy lemma and (H2)). By conclusion 1 of the policy lemma, its matrices QtQ_t, VtV_t, RtR_t, WtW_t, and F^\hat{F} are given by the same defining formulas as those of the completion-of-squares theorem, and the family ZZ of (H2) satisfies the Riccati equation of hypothesis (H2) of that theorem; so both that theorem and the cost expansion theorem are available for each fixed NN, their common hypothesis A=∫[0,T]E[∣at∣2]dt<∞\mathcal{A}=\int_{[0,T]}\mathbb{E}[|\mathfrak{a}_t|^2]dt<\infty holding by part (a) of the filter error covariance lemma. By part (d) of that lemma there is a real C1C_1, independent of NN, with

sup⁑t∈[0,T](E[∣st∣4]+E[∣s^tN∣4]+E[∣ΡtN∣4]+E[∣at∣4])≀C1ΞΊandsup⁑t(E[∣st∣2]+E[∣at∣2])≀2(4+C1)ΞΊ(0.1)\sup_{t\in[0,T]}\Big(\mathbb{E}[|\mathfrak{s}_t|^4]+\mathbb{E}[|\hat{\mathfrak{s}}^N_t|^4]+\mathbb{E}[|\varepsilon^N_t|^4]+\mathbb{E}[|\mathfrak{a}_t|^4]\Big)\le C_1\kappa\quad\text{and}\quad\sup_{t}\Big(\mathbb{E}[|\mathfrak{s}_t|^2]+\mathbb{E}[|\mathfrak{a}_t|^2]\Big)\le2(4+C_1)\kappa\tag{0.1}

for every NN (the second bound obtained from the first, together with E[βˆ£β‹…βˆ£2]≀1+E[βˆ£β‹…βˆ£4]\mathbb{E}[|\cdot|^2]\le1+\mathbb{E}[|\cdot|^4] applied to the four summands of the first display and ΞΊβ‰₯1\kappa\ge1, as in part (d) of that lemma); in particular E[∣zt∣2]≀2(4+C1)ΞΊ\mathbb{E}[|z_t|^2]\le2(4+C_1)\kappa and E[∣zt∣4]≀8(E[∣st∣4]+E[∣at∣4])≀16C1ΞΊ\mathbb{E}[|z_t|^4]\le8(\mathbb{E}[|\mathfrak{s}_t|^4]+\mathbb{E}[|\mathfrak{a}_t|^4])\le16C_1\kappa for all tt and NN. Finally, the limit integrand t↦Ztβ‹…Ξ˜t⋆+(WtRtβˆ’1Wt⊀)β‹…Ξ tt\mapsto Z_t\cdot\Theta^\star_t+(W_tR_t^{-1}W_t^{\top})\cdot\Pi_t is continuous, being a finite sum of products of continuous functions (conclusions 1 and 2 of the policy lemma and (H2)), which proves the first assertion.

Step 1: expansion and completion of squares. Fix NN. By parts (b) and (c) of the cost expansion theorem, JN[hN]J^N[h^N] is finite, LQG[(s),(a)]LQG[(\mathfrak{s}),(\mathfrak{a})] is a well-defined real number, and

N(JN[hN]βˆ’JMF)+βˆ‘Ξ³P0Ξ³ΞΆNΞ³=LQG[(s),(a)]+RN,N\big(J^N[h^N]-J^{MF}\big)+\sum_{\gamma}P^\gamma_0\zeta^\gamma_N=LQG\big[(\mathfrak{s}),(\mathfrak{a})\big]+R_N ,

with RNR_N obeying the modulus bound of its part (c). By part (c) of the completion-of-squares theorem, with ut=at+Rtβˆ’1Wt⊀stu_t=\mathfrak{a}_t+R_t^{-1}W_t^{\top}\mathfrak{s}_t and ese_s its linearization residual,

LQG[(s),(a)]=E[s0β‹…Z0s0]+∫[0,T]E[usβ‹…Rsus]ds+∫[0,T](2 E[ssβ‹…Zses]+βˆ‘Ξ³,Ξ΄Zsγδ E[Θγδ(Ξ£s,Ξ±s)])ds,LQG\big[(\mathfrak{s}),(\mathfrak{a})\big]=\mathbb{E}\big[\mathfrak{s}_0\cdot Z_0\mathfrak{s}_0\big]+\int_{[0,T]}\mathbb{E}\big[u_s\cdot R_su_s\big]ds+\int_{[0,T]}\Big(2\,\mathbb{E}\big[\mathfrak{s}_s\cdot Z_se_s\big]+\sum_{\gamma,\delta}Z^{\gamma\delta}_s\,\mathbb{E}\big[\Theta^{\gamma\delta}(\Sigma_s,\alpha_s)\big]\Big)ds,

all terms finite. In particular each N(JN[hN]βˆ’JMF)+βˆ‘Ξ³P0Ξ³ΞΆNΞ³N(J^N[h^N]-J^{MF})+\sum_\gamma P^\gamma_0\zeta^\gamma_N is a well-defined real number. We treat the four pieces in Steps 2--5.

Step 2: initial term. E[s0β‹…Z0s0]=βˆ‘Ξ³,Ξ΄Z0γδ E[s0N,Ξ³s0N,Ξ΄]β†’βˆ‘Ξ³,Ξ΄Z0γδΠ0Ξ³Ξ΄=Z0β‹…Ξ 0\mathbb{E}[\mathfrak{s}_0\cdot Z_0\mathfrak{s}_0]=\sum_{\gamma,\delta}Z^{\gamma\delta}_0\,\mathbb{E}[\mathfrak{s}^{N,\gamma}_0\mathfrak{s}^{N,\delta}_0]\to\sum_{\gamma,\delta}Z^{\gamma\delta}_0\Pi^{\gamma\delta}_0=Z_0\cdot\Pi_0 as Nβ†’βˆžN\to\infty, by (I1) and linearity (a finite sum of convergent sequences).

Step 3: control term. By conclusion 4(a) of the policy lemma, at=N1/2(Ξ±tβˆ’At)=βˆ’Rtβˆ’1Wt⊀s^tN\mathfrak{a}_t=N^{1/2}(\alpha_t-A_t)=-R_t^{-1}W_t^{\top}\hat{\mathfrak{s}}^N_t on the regular event, hence almost surely ut=Rtβˆ’1Wt⊀(stβˆ’s^tN)=Gt ΡtNu_t=R_t^{-1}W_t^{\top}(\mathfrak{s}_t-\hat{\mathfrak{s}}^N_t)=\mathcal{G}_t\,\varepsilon^N_t. Therefore, for every ss,

E[usβ‹…Rsus]=βˆ‘i,jRsijβˆ‘Ξ³,Ξ΄GsiΞ³Gsjδ E[Ξ΅sN,Ξ³Ξ΅sN,Ξ΄]=(Gs⊀RsGs)β‹…Ξ sN=(WsRsβˆ’1Ws⊀)β‹…Ξ sN,\mathbb{E}\big[u_s\cdot R_su_s\big]=\sum_{i,j}R^{ij}_s\sum_{\gamma,\delta}\mathcal{G}^{i\gamma}_s\mathcal{G}^{j\delta}_s\,\mathbb{E}\big[\varepsilon^{N,\gamma}_s\varepsilon^{N,\delta}_s\big]=\big(\mathcal{G}_s^{\top}R_s\mathcal{G}_s\big)\cdot\Pi^N_s=\big(W_sR_s^{-1}W_s^{\top}\big)\cdot\Pi^N_s ,

where the last equality uses Gs=Rsβˆ’1Ws⊀\mathcal{G}_s=R_s^{-1}W_s^{\top}, the symmetry of RsR_s and of Rsβˆ’1R_s^{-1} (invertibility of symmetric positive definite matrices), and RsRsβˆ’1=IR_sR_s^{-1}=I: G⊀RG=WRβˆ’1RRβˆ’1W⊀=WRβˆ’1W⊀\mathcal{G}^{\top}R\mathcal{G}=WR^{-1}RR^{-1}W^{\top}=WR^{-1}W^{\top}, with the reversal rule for the transpose. Consequently, by linearity and monotonicity of the integral (both integrands bounded and measurable in ss: part (a) of the filter error covariance lemma for Ξ N\Pi^N, continuity for Ξ \Pi),

∣∫[0,T]E[usβ‹…Rsus] dsβˆ’βˆ«[0,T](WsRsβˆ’1Ws⊀)β‹…Ξ s dsβˆ£Β β‰€Β T l2 cW sup⁑s∈[0,T]max⁑γ,δ∣ΠsN,Ξ³Ξ΄βˆ’Ξ sγδ∣.\Big|\int_{[0,T]}\mathbb{E}[u_s\cdot R_su_s]\,ds-\int_{[0,T]}\big(W_sR_s^{-1}W_s^{\top}\big)\cdot\Pi_s\,ds\Big|\ \le\ T\,l^2\,c_W\,\sup_{s\in[0,T]}\max_{\gamma,\delta}\big|\Pi^{N,\gamma\delta}_s-\Pi^{\gamma\delta}_s\big| .

By part (e) of the filter error covariance lemma, the supremum is at most C(max⁑γ,δ∣E[s0N,Ξ³s0N,Ξ΄]βˆ’Ξ 0γδ∣+Nβˆ’1/2(1+E[∣s0N∣4]))≀C(max⁑γ,δ∣E[s0N,Ξ³s0N,Ξ΄]βˆ’Ξ 0γδ∣+Nβˆ’1/2ΞΊ)C(\max_{\gamma,\delta}|\mathbb{E}[\mathfrak{s}^{N,\gamma}_0\mathfrak{s}^{N,\delta}_0]-\Pi^{\gamma\delta}_0|+N^{-1/2}(1+\mathbb{E}[|\mathfrak{s}^N_0|^4]))\le C(\max_{\gamma,\delta}|\mathbb{E}[\mathfrak{s}^{N,\gamma}_0\mathfrak{s}^{N,\delta}_0]-\Pi^{\gamma\delta}_0|+N^{-1/2}\kappa), which tends to 00 by (I1) and (I2). Hence ∫E[usβ‹…Rsus]dsβ†’βˆ«[0,T](WsRsβˆ’1Ws⊀)β‹…Ξ s ds\int\mathbb{E}[u_s\cdot R_su_s]ds\to\int_{[0,T]}(W_sR_s^{-1}W_s^{\top})\cdot\Pi_s\,ds.

Step 4: covariance term. By clause 6 of the fluctuation LQG data, (Θs⋆)Ξ³Ξ΄=Θγδ(Ss,As)(\Theta^\star_s)^{\gamma\delta}=\Theta^{\gamma\delta}(S_s,A_s). By part (c) of the covariance deviation lemma together with (0.1) and monotonicity of the integral,

∣∫[0,T]βˆ‘Ξ³,Ξ΄ZsΞ³Ξ΄(E[Θγδ(Ξ£s,Ξ±s)]βˆ’Ξ˜Ξ³Ξ΄(Ss,As))dsβˆ£Β β‰€Β l2CZ cΘN T (S+A)1/2 ≀ l2CZ cΞ˜β€‰T (2T(4+C1)ΞΊ)1/2 Nβˆ’1/2,\Big|\int_{[0,T]}\sum_{\gamma,\delta}Z^{\gamma\delta}_s\Big(\mathbb{E}[\Theta^{\gamma\delta}(\Sigma_s,\alpha_s)]-\Theta^{\gamma\delta}(S_s,A_s)\Big)ds\Big|\ \le\ l^2C_Z\,\frac{c_\Theta}{\sqrt{N}}\,\sqrt{T}\,\big(\mathcal{S}+\mathcal{A}\big)^{1/2}\ \le\ l^2C_Z\,c_\Theta\,\sqrt{T}\,\big(2T(4+C_1)\kappa\big)^{1/2}\,N^{-1/2},

with the constant cΘc_\Theta of that lemma and S+A≀2T(4+C1)ΞΊ\mathcal{S}+\mathcal{A}\le2T(4+C_1)\kappa by (0.1). This tends to 00, and ∫[0,T]βˆ‘Ξ³,Ξ΄ZsγδΘγδ(Ss,As)ds=∫[0,T]Zsβ‹…Ξ˜s⋆ ds\int_{[0,T]}\sum_{\gamma,\delta}Z^{\gamma\delta}_s\Theta^{\gamma\delta}(S_s,A_s)ds=\int_{[0,T]}Z_s\cdot\Theta^\star_s\,ds by the definition of the pairing.

Step 5: residual term. By part (b) of the completion-of-squares theorem, ∣esβˆ£β‰€ceNβˆ’1/2(∣ss∣2+∣as∣2)=ceNβˆ’1/2∣zs∣2|e_s|\le c_eN^{-1/2}(|\mathfrak{s}_s|^2+|\mathfrak{a}_s|^2)=c_eN^{-1/2}|z_s|^2 with ce=32l3/2(l+m)Kc_e=\tfrac32l^{3/2}(l+m)K. Since ∣xβ‹…Zsyβˆ£β‰€CZ(βˆ‘Ξ³βˆ£xγ∣)(βˆ‘Ξ΄βˆ£yδ∣)≀lCZ∣x∣∣y∣|x\cdot Z_sy|\le C_Z(\sum_\gamma|x^\gamma|)(\sum_\delta|y^\delta|)\le lC_Z|x||y|, and pointwise ∣ssβˆ£β€‰βˆ£zs∣2β‰€βˆ£zs∣3≀1+∣zs∣4|\mathfrak{s}_s|\,|z_s|^2\le|z_s|^3\le1+|z_s|^4,

∣∫[0,T]2 E[ssβ‹…Zses] dsβˆ£Β β‰€Β 2lCZ ce Nβˆ’1/2∫[0,T]E[1+∣zs∣4]ds ≀ 2lCZ ce T (1+16C1ΞΊ) Nβˆ’1/2 ⟢ 0.\Big|\int_{[0,T]}2\,\mathbb{E}[\mathfrak{s}_s\cdot Z_se_s]\,ds\Big|\ \le\ 2lC_Z\,c_e\,N^{-1/2}\int_{[0,T]}\mathbb{E}\big[1+|z_s|^4\big]ds\ \le\ 2lC_Z\,c_e\,T\,(1+16C_1\kappa)\,N^{-1/2}\ \longrightarrow\ 0 .

Step 6: the remainder vanishes. Let ΞΈ>0\theta>0. By part (a) of the cost expansion theorem there is δ∘>0\delta^\circ>0 with Ο‰L(u)≀θ\omega_L(u)\le\theta, Ο‰b(u)≀θ\omega_b(u)\le\theta, and Ο‰G(u)≀θ\omega_G(u)\le\theta for all u∈[0,δ∘]u\in[0,\delta^\circ], and globally Ο‰L≀2Kc\omega_L\le2K_c, Ο‰b≀6lK\omega_b\le6lK, Ο‰G≀2Kc\omega_G\le2K_c. Let CPC_P be the co-state bound of that theorem and write Ο‰Λ‰=2Kc+CP 6lK\bar{\omega}=2K_c+C_P\,6lK. Splitting the time integral of its part (c) on the events {ρtβ‰€Ξ΄βˆ˜}\{\rho_t\le\delta^\circ\} and {ρt>δ∘}\{\rho_t>\delta^\circ\} (the moduli being nondecreasing) and using monotonicity,

∣RNβˆ£Β β‰€Β l+m2((1+CP)θ∫[0,T]E[∣zt∣2]dt+Ο‰Λ‰βˆ«[0,T]E[1{ρt>δ∘}∣zt∣2]dt)+l2(θ E[∣sT∣2]+2Kc E[1{∣sT∣>δ∘N}∣sT∣2]),|R_N|\ \le\ \frac{l+m}{2}\Big((1+C_P)\theta\int_{[0,T]}\mathbb{E}[|z_t|^2]dt+\bar{\omega}\int_{[0,T]}\mathbb{E}\big[\mathbf{1}_{\{\rho_t>\delta^\circ\}}|z_t|^2\big]dt\Big)+\frac{l}{2}\Big(\theta\,\mathbb{E}[|\mathfrak{s}_T|^2]+2K_c\,\mathbb{E}\big[\mathbf{1}_{\{|\mathfrak{s}_T|>\delta^\circ\sqrt{N}\}}|\mathfrak{s}_T|^2\big]\Big),

where 1E\mathbf{1}_E denotes the indicator of the event EE and we used d(Ξ£T,ST)=Nβˆ’1/2∣sT∣d(\Sigma_T,S_T)=N^{-1/2}|\mathfrak{s}_T|. By the Cauchy-Schwarz inequality and Markov's inequality applied to the nonnegative variable ∣zt∣2|z_t|^2 at level N(δ∘)2N(\delta^\circ)^2,

E[1{ρt>δ∘}∣zt∣2] ≀ (E[∣zt∣4])1/2 P(∣zt∣2β‰₯N(δ∘)2)1/2 ≀ (16C1ΞΊ)1/2(2(4+C1)ΞΊN(δ∘)2)1/2Β =Β 42 κ C1(4+C1)δ∘N ≀ 6(4+C1)κδ∘N,\mathbb{E}\big[\mathbf{1}_{\{\rho_t>\delta^\circ\}}|z_t|^2\big]\ \le\ \big(\mathbb{E}[|z_t|^4]\big)^{1/2}\,\mathbb{P}\big(|z_t|^2\ge N(\delta^\circ)^2\big)^{1/2}\ \le\ \big(16C_1\kappa\big)^{1/2}\Big(\frac{2(4+C_1)\kappa}{N(\delta^\circ)^2}\Big)^{1/2}\ =\ \frac{4\sqrt{2}\,\kappa\,\sqrt{C_1(4+C_1)}}{\delta^\circ\sqrt{N}}\ \le\ \frac{6(4+C_1)\kappa}{\delta^\circ\sqrt{N}} ,

using C1(4+C1)≀4+C1\sqrt{C_1(4+C_1)}\le4+C_1 (since C1≀4+C1C_1\le4+C_1) and 42≀64\sqrt{2}\le6; using (0.1), and the same bound holds for the terminal term with ∣sT∣|\mathfrak{s}_T| in place of ∣zt∣|z_t|. Hence, with (0.1) again,

lim sup⁑Nβ†’βˆžβˆ£RNβˆ£Β β‰€Β (l+m2(1+CP) 2T(4+C1)ΞΊ+l2 2(4+C1)ΞΊ) θ,\limsup_{N\to\infty}|R_N|\ \le\ \Big(\frac{l+m}{2}(1+C_P)\,2T(4+C_1)\kappa+\frac{l}{2}\,2(4+C_1)\kappa\Big)\,\theta ,

and since θ>0\theta>0 was arbitrary, RN→0R_N\to0.

Step 7: conclusion. Combining Steps 1--6, the sequence N(JN[hN]βˆ’JMF)+βˆ‘Ξ³P0Ξ³ΞΆNΞ³N(J^N[h^N]-J^{MF})+\sum_\gamma P^\gamma_0\zeta^\gamma_N is, for each NN, the sum of four terms converging respectively to Z0β‹…Ξ 0Z_0\cdot\Pi_0, ∫[0,T](WsRsβˆ’1Ws⊀)β‹…Ξ s ds\int_{[0,T]}(W_sR_s^{-1}W_s^{\top})\cdot\Pi_s\,ds, ∫[0,T]Zsβ‹…Ξ˜s⋆ ds\int_{[0,T]}Z_s\cdot\Theta^\star_s\,ds, and 00, plus the remainder RNβ†’0R_N\to0. The limit identity follows, the two integrals combining by linearity of the integral into the integral of the continuous integrand recorded in the statement. β– \blacksquare

Please log in to copy this version.

Citations

Loading…

Dependency Graph

0 prerequisites

Prerequisites

Loading...

Comments

Loading…