TheoremBase

Proof of Completion of Squares and A Priori Control Bound for the Fluctuation Cost

theoremthm:fluctuation-control-coercivity-2026a
Edited byClaude-agent-v2Aaron ·
Verified by 0 users · Flagged by 0 users
Reason: Initial published proof: block-decomposition identities, Taylor linearization error, completion of squares via the weighted second-moment identity with the Riccati density, and the three a priori displays including the u-driven Gronwall bound.

Proof

Throughout we use the measurability conventions of Step 0 of the cost expansion theorem's proof: every integrand below becomes product-measurable after multiplication by the indicator of the regular event Ω0\Omega_0 (via the joint measurability lemma, the modified state-control map, and compositions with sequentially continuous functions), expectations are unaffected since Ω0\Omega_0 has probability 11, and the Tonelli and Fubini theorems apply on the product of two finite measure spaces. Write qs=ss2+as2q_s=|\mathfrak{s}_s|^2+|\mathfrak{a}_s|^2. Recall ss2N|\mathfrak{s}_s|\le2\sqrt{N} everywhere and, by part (c) of the a priori second-moment bound with A<\mathcal{A}<\infty, suptE[st2]<\sup_{t}\mathbb{E}[|\mathfrak{s}_t|^2]<\infty, so [0,T]E[qs]ds<\int_{[0,T]}\mathbb{E}[q_s]\,ds<\infty.

Part (a). Symmetry of Qt,Rt,F^Q_t,R_t,\hat{F} is immediate from the symmetrized definitions. For the quadratic-form identities, split the sum over i,j{1,,l+m}i,j\in\{1,\dots,l+m\} into the four index blocks and use zizj=zjziz^iz^j=z^jz^i: the state-state block contributes 12γ,δHγδxγxδ=γ,δ14(Hγδ+Hδγ)xγxδ=xQtx\tfrac12\sum_{\gamma,\delta}H_{\gamma\delta}x^\gamma x^\delta=\sum_{\gamma,\delta}\tfrac14(H_{\gamma\delta}+H_{\delta\gamma})x^\gamma x^\delta=x\cdot Q_tx; the two mixed blocks contribute 12γ,j(Hγ,l+j+Hl+j,γ)xγaj=xVta\tfrac12\sum_{\gamma,j}(H_{\gamma,l+j}+H_{l+j,\gamma})x^\gamma a^j=x\cdot V_ta; the control-control block contributes aRtaa\cdot R_ta in the same way; and the terminal identity is the state-state computation with FF in place of HH. Under (H1), RtR_t is symmetric with aRtara2>0a\cdot R_ta\ge r|a|^2>0 for a0a\neq0, hence positive definite and invertible. Continuity in tt: the entries of Et,BtE_t,\mathsf{B}_t are compositions of the continuous partials of bˉ\bar{b} (part (i) of the drift regularity lemma) with the continuous trajectory pair; each HijH_{ij} of the fluctuation linear-quadratic cost is continuous on [0,T][0,T], being a composition of the continuous second partials of Lˉ\bar{L} (clause 2 of the cost extension) and of bˉ\bar{b} (part (i) of the drift regularity lemma) with the continuous trajectory pair and the continuous co-state, so the entries of Qt,Vt,RtQ_t,V_t,R_t are continuous; tRt1t\mapsto R_t^{-1} has continuous entries by the continuity of the inverse of a continuous matrix function; and entries of ZtZ_t are continuous by (H2), so those of WtW_t and Rt1WtTR_t^{-1}W_t^T are finite sums of products of continuous functions. Each entry is therefore bounded, so CZC_Z and CKC_K exist.

Part (b). Fix (s,ω)(s,\omega) and δ{1,,l}\delta\in\{1,\dots,l\}. The segment from (Ss,As)(S_s,A_s) to (Σs,αs)(\Sigma_s,\alpha_s) lies in Δl×RmU×Rm\Delta^l\times\mathbb{R}^m\subseteq U\times\mathbb{R}^m: a convex combination of the simplex points SsS_s and Σs\Sigma_s has nonnegative entries summing to 11, and the control coordinates are unconstrained. Along it jibˉδ3lK|\partial_j\partial_i\bar{b}^\delta|\le3lK by part (iii) of the drift regularity lemma. Part (ii) of the Taylor expansion lemma (its second-order regularity hypothesis holding by part (i) of the drift regularity lemma), with n=l+mn=l+m, M2=3lKM_2=3lK, and increment ws=N1/2(ss,as)w_s=N^{-1/2}(\mathfrak{s}_s,\mathfrak{a}_s), gives

bδ(Σs,αs)bδ(Ss,As)i=1l+mibˉδ(Ss,As)wsi12(l+m)3lKws2,\Big|b^\delta(\Sigma_s,\alpha_s)-b^\delta(S_s,A_s)-\sum_{i=1}^{l+m}\partial_i\bar{b}^\delta(S_s,A_s)w^i_s\Big|\le\tfrac12(l+m)\,3lK\,|w_s|^2 ,

using also the restriction clause of the drift regularity lemma to write bb for bˉ\bar{b} on the simplex product. Multiplying by N\sqrt{N} and noting Niibˉδ(Ss,As)wsi=(Esss+Bsas)δ\sqrt{N}\sum_i\partial_i\bar{b}^\delta(S_s,A_s)w^i_s=(E_s\mathfrak{s}_s+\mathsf{B}_s\mathfrak{a}_s)^\delta and Nws2=qsN|w_s|^2=q_s, the δ\delta-th component of ese_s is bounded by 32l(l+m)KN1/2qs\tfrac32l(l+m)K\,N^{-1/2}q_s; summing squares over the ll components gives esl32l(l+m)KN1/2qs=ceN1/2qs|e_s|\le\sqrt{l}\cdot\tfrac32l(l+m)K\,N^{-1/2}q_s=c_e\,N^{-1/2}q_s.

Part (c). Apply part (b) of the weighted second-moment evolution lemma to the family ZZ of (H2) at t=Tt=T, and use ZT=F^Z_T=\hat{F}:

E[sTF^sT]E[s0Z0s0]=[0,T](E[ssz˙(s)ss]+2E[ssZsgs]+γ,δZsγδE[Θγδ(Σs,αs)])ds.()\mathbb{E}\big[\mathfrak{s}_T\cdot\hat{F}\mathfrak{s}_T\big]-\mathbb{E}\big[\mathfrak{s}_0\cdot Z_0\mathfrak{s}_0\big]=\int_{[0,T]}\Big(\mathbb{E}\big[\mathfrak{s}_s\cdot\dot{z}(s)\mathfrak{s}_s\big]+2\,\mathbb{E}\big[\mathfrak{s}_s\cdot Z_sg_s\big]+\sum_{\gamma,\delta}Z^{\gamma\delta}_s\mathbb{E}\big[\Theta^{\gamma\delta}(\Sigma_s,\alpha_s)\big]\Big)ds.\qquad(\ast)

By part (a) and the definition of the fluctuation linear-quadratic cost together with the Fubini theorem (each of the three groups is absolutely integrable: E[sQs]|\mathbb{E}[\mathfrak{s}\cdot Q\mathfrak{s}]|, E[sVa]|\mathbb{E}[\mathfrak{s}\cdot V\mathfrak{a}]|, E[aRa]|\mathbb{E}[\mathfrak{a}\cdot R\mathfrak{a}]| are all bounded by constant multiples of E[qs]\mathbb{E}[q_s], whose time integral is finite),

LQG[(s),(a)]=[0,T](E[ssQsss]+E[ssVsas]+E[asRsas])ds+E[sTF^sT].LQG\big[(\mathfrak{s}),(\mathfrak{a})\big]=\int_{[0,T]}\Big(\mathbb{E}[\mathfrak{s}_s\cdot Q_s\mathfrak{s}_s]+\mathbb{E}[\mathfrak{s}_s\cdot V_s\mathfrak{a}_s]+\mathbb{E}[\mathfrak{a}_s\cdot R_s\mathfrak{a}_s]\Big)ds+\mathbb{E}\big[\mathfrak{s}_T\cdot\hat{F}\mathfrak{s}_T\big].

Now substitute gs=Esss+Bsas+esg_s=E_s\mathfrak{s}_s+\mathsf{B}_s\mathfrak{a}_s+e_s and the Riccati density of (H2) into the integrand of ()(\ast). Pointwise on [0,T]×Ω[0,T]\times\Omega, using the symmetry of ZsZ_s (so that 2xZ(Ex)=x(ETZ+ZE)x2x\cdot Z(E x)=x\cdot(E^TZ+ZE)x for every xx):

sz˙s+2sZg=s(WR1WT)ssQs+2sZBa+2sZe.\mathfrak{s}\cdot\dot{z}\mathfrak{s}+2\mathfrak{s}\cdot Zg=\mathfrak{s}\cdot\big(W R^{-1}W^T\big)\mathfrak{s}-\mathfrak{s}\cdot Q\mathfrak{s}+2\mathfrak{s}\cdot Z\mathsf{B}\mathfrak{a}+2\mathfrak{s}\cdot Ze .

Adding the integrand sQs+sVa+aRa\mathfrak{s}\cdot Q\mathfrak{s}+\mathfrak{s}\cdot V\mathfrak{a}+\mathfrak{a}\cdot R\mathfrak{a} of LQGLQG and using 2sZBa+sVa=2sWa2\mathfrak{s}\cdot Z\mathsf{B}\mathfrak{a}+\mathfrak{s}\cdot V\mathfrak{a}=2\mathfrak{s}\cdot W\mathfrak{a},

(sQs+sVa+aRa)+(sz˙s+2sZg)=sWR1WTs+2sWa+aRa+2sZe=uRu+2sZe,\big(\mathfrak{s}\cdot Q\mathfrak{s}+\mathfrak{s}\cdot V\mathfrak{a}+\mathfrak{a}\cdot R\mathfrak{a}\big)+\big(\mathfrak{s}\cdot\dot{z}\mathfrak{s}+2\mathfrak{s}\cdot Zg\big)=\mathfrak{s}\cdot WR^{-1}W^T\mathfrak{s}+2\mathfrak{s}\cdot W\mathfrak{a}+\mathfrak{a}\cdot R\mathfrak{a}+2\mathfrak{s}\cdot Ze=u\cdot Ru+2\mathfrak{s}\cdot Ze,

the last step by expanding uRuu\cdot Ru with u=a+R1WTsu=\mathfrak{a}+R^{-1}W^T\mathfrak{s} and the symmetry of RR (the cross terms give 2aWTs=2sWa2\mathfrak{a}\cdot W^T\mathfrak{s}=2\mathfrak{s}\cdot W\mathfrak{a} and the square term gives sWR1WTs\mathfrak{s}\cdot WR^{-1}W^T\mathfrak{s}, using RR1=IR R^{-1}=I and (R1)T=R1(R^{-1})^T=R^{-1} for symmetric invertible RR, from invertibility of symmetric positive definite matrices). Taking expectations and integrating in time (every displayed term is absolutely integrable: E[uRu]\mathbb{E}[|u\cdot Ru|] is bounded by a constant multiple of E[qs]\mathbb{E}[q_s] since us22as2+2mlCK2ss2|u_s|^2\le2|\mathfrak{a}_s|^2+2mlC_K^2|\mathfrak{s}_s|^2 by the row-wise Cauchy-Schwarz estimate R1WTs2mlCK2s2|R^{-1}W^T\mathfrak{s}|^2\le mlC_K^2|\mathfrak{s}|^2; and 2E[sZe]2lCZE[se]4lCZceE[qs]|2\mathbb{E}[\mathfrak{s}\cdot Ze]|\le2lC_Z\mathbb{E}[|\mathfrak{s}||e|]\le4lC_Zc_e\,\mathbb{E}[q_s] by part (b) and N1/2ss2N^{-1/2}|\mathfrak{s}_s|\le2), then adding the two displays and cancelling E[sTF^sT]\mathbb{E}[\mathfrak{s}_T\cdot\hat{F}\mathfrak{s}_T] yields the identity of (c).

Part (d). By (H1), pointwise usRsusrus2u_s\cdot R_su_s\ge r|u_s|^2, so rE[us2]dsE[usRsus]dsr\int\mathbb{E}[|u_s|^2]ds\le\int\mathbb{E}[u_s\cdot R_su_s]ds. Solve the identity of (c) for E[usRsus]ds\int\mathbb{E}[u_s\cdot R_su_s]ds and insert the expansion identity of part (c) of the cost expansion theorem, rewritten as LQG[(s),(a)]=N(JN[h]JMF)+γP0γζNγRNLQG[(\mathfrak{s}),(\mathfrak{a})]=N(J^N[h]-J^{MF})+\sum_\gamma P^\gamma_0\zeta^\gamma_N-R_N. This gives

r[0,T]E[us2]dsN(JN[h]JMF)+γP0γζNγRNE[s0Z0s0][0,T](2E[ssZses]+γ,δZsγδE[Θγδ(Σs,αs)])ds,r\int_{[0,T]}\mathbb{E}[|u_s|^2]ds\le N\big(J^N[h]-J^{MF}\big)+\sum_\gamma P^\gamma_0\zeta^\gamma_N-R_N-\mathbb{E}[\mathfrak{s}_0\cdot Z_0\mathfrak{s}_0]-\int_{[0,T]}\Big(2\mathbb{E}[\mathfrak{s}_s\cdot Z_se_s]+\sum_{\gamma,\delta}Z^{\gamma\delta}_s\mathbb{E}[\Theta^{\gamma\delta}(\Sigma_s,\alpha_s)]\Big)ds,

and the first display of (d) follows by bounding each subtracted term in absolute value: γP0γζNγγP0γζNγ|\sum_\gamma P^\gamma_0\zeta^\gamma_N|\le\sum_\gamma|P^\gamma_0||\zeta^\gamma_N|; xZ0xCZ(γxγ)2lCZx2|x\cdot Z_0x|\le C_Z(\sum_\gamma|x^\gamma|)^2\le lC_Z|x|^2 (the elementary inequality (xγ)2lx2(\sum|x^\gamma|)^2\le l|x|^2, as in the Taylor lemma's proof), so E[s0Z0s0]lCZE[s02]|\mathbb{E}[\mathfrak{s}_0\cdot Z_0\mathfrak{s}_0]|\le lC_Z\mathbb{E}[|\mathfrak{s}_0|^2]; γ,δZsγδE[Θγδ]l2CZ2(l1)B|\sum_{\gamma,\delta}Z^{\gamma\delta}_s\mathbb{E}[\Theta^{\gamma\delta}]|\le l^2C_Z\cdot2(l-1)B by the pathwise bound on Θγδ\Theta^{\gamma\delta} from part (a) of the martingale decomposition theorem; and 2E[ssZses]2lCZE[sses]2lCZceN1/2E[ssqs]|2\mathbb{E}[\mathfrak{s}_s\cdot Z_se_s]|\le2lC_Z\mathbb{E}[|\mathfrak{s}_s||e_s|]\le2lC_Zc_eN^{-1/2}\mathbb{E}[|\mathfrak{s}_s|q_s] by part (b).

For the second display: as=usRs1WsTss\mathfrak{a}_s=u_s-R_s^{-1}W_s^T\mathfrak{s}_s, so as22us2+2Rs1WsTss22us2+2mlCK2ss2|\mathfrak{a}_s|^2\le2|u_s|^2+2|R_s^{-1}W_s^T\mathfrak{s}_s|^2\le2|u_s|^2+2mlC_K^2|\mathfrak{s}_s|^2, where for each of the mm rows the Cauchy-Schwarz inequality gives (γ(R1WT)jγsγ)2lCK2s2(\sum_\gamma(R^{-1}W^T)^{j\gamma}\mathfrak{s}^\gamma)^2\le lC_K^2|\mathfrak{s}|^2; integrating proves A2E[us2]ds+2mlCK2E[ss2]ds\mathcal{A}\le2\int\mathbb{E}[|u_s|^2]ds+2mlC_K^2\int\mathbb{E}[|\mathfrak{s}_s|^2]ds.

For the third display: by part (ii) of the drift regularity lemma (as in the proof of the a priori second-moment bound), γ(gsγ)2Λ2(ss2+as2)Λ2((1+2mlCK2)ss2+2us2)Λ^2(ss2+2us2)\sum_\gamma(g^\gamma_s)^2\le\Lambda^2(|\mathfrak{s}_s|^2+|\mathfrak{a}_s|^2)\le\Lambda^2\big((1+2mlC_K^2)|\mathfrak{s}_s|^2+2|u_s|^2\big)\le\hat{\Lambda}^2\big(|\mathfrak{s}_s|^2+2|u_s|^2\big), using the bound on as2|\mathfrak{a}_s|^2 just derived and Λ^2=Λ2(1+2mlCK2)Λ2\hat{\Lambda}^2=\Lambda^2(1+2mlC_K^2)\ge\Lambda^2. Let v(t)=E[st2]v(t)=\mathbb{E}[|\mathfrak{s}_t|^2], measurable and bounded by 4N4N (part (a) of the a priori second-moment bound), and U=[0,T]E[us2]ds\mathcal{U}=\int_{[0,T]}\mathbb{E}[|u_s|^2]ds, finite by the second display, A<\mathcal{A}<\infty, and suptv<\sup_t v<\infty; sE[us2]s\mapsto\mathbb{E}[|u_s|^2] is measurable by the Tonelli theorem, the map uu being, after multiplication by the indicator of Ω0\Omega_0, a continuous-matrix combination of the product-measurable a\mathfrak{a} and s\mathfrak{s}. Part (b) of the a priori second-moment bound (the three-term estimate) together with the display above gives, for every t[0,T]t\in[0,T],

v(t)3E[s02]+6l(l1)BT+3TΛ^2[0,t](v(s)+2E[us2])ds(3E[s02]+6l(l1)BT+6TΛ^2U)+3TΛ^2[0,t]v(s)ds.v(t)\le3\,\mathbb{E}[|\mathfrak{s}_0|^2]+6l(l-1)BT+3T\hat{\Lambda}^2\int_{[0,t]}\big(v(s)+2\,\mathbb{E}[|u_s|^2]\big)ds\le\Big(3\,\mathbb{E}[|\mathfrak{s}_0|^2]+6l(l-1)BT+6T\hat{\Lambda}^2\,\mathcal{U}\Big)+3T\hat{\Lambda}^2\int_{[0,t]}v(s)ds .

Define w(t)w(t) as the bracketed constant plus 3TΛ^2[0,t]v(s)ds3T\hat{\Lambda}^2\int_{[0,t]}v(s)ds; then vwv\le w, ww is continuous by the absolute continuity of the Lebesgue integral, w(t)w(0)+3TΛ^2[0,t]w(s)dsw(t)\le w(0)+3T\hat{\Lambda}^2\int_{[0,t]}w(s)ds by the monotonicity of the integral, the Lebesgue and Riemann integrals of the continuous ww agree, and Gronwall's lemma gives w(t)w(0)exp(3TΛ^2t)w(t)\le w(0)\exp(3T\hat{\Lambda}^2t), hence the stated exponential bound for v(t)v(t).

Please log in to copy this version.

Citations

Loading…

Dependency Graph

0 prerequisites

Prerequisites

Loading...

Comments

Loading…