Adopt the full setting of the second-order expansion of the N N N -agent cost : the fluctuation processes s t \mathfrak{s}_t s t , a t \mathfrak{a}_t a t of a solution about a mean-field trajectory pair ( S , A ) (S,A) ( S , A ) , the extension ( U , V , β ˉ ) (U,V,\bar{\beta}) ( U , V , β ˉ ) of β \beta β with derivative bound K K K and its extended aggregate state drift b ˉ \bar{b} b ˉ , the extension of ( L , G ) (L,G) ( L , G ) - whose open set, written W W W in the cost expansion theorem, we here write U c U_c U c , so the extension is ( U c , L ˉ , G ˉ ) (U_c,\bar{L},\bar{G}) ( U c , L ˉ , G ˉ ) , freeing the letter W W W ; the time-indexed matrices V t V_t V t and W t W_t W t below are unrelated to the control-side open set V V V of ( U , V , β ˉ ) (U,V,\bar{\beta}) ( U , V , β ˉ ) - a stationary co-state P P P , the fluctuation linear-quadratic cost with Hessian coefficients H i j ( t ) H_{ij}(t) H ij ( t ) and F γ δ F_{\gamma\delta} F γ δ , the N N N -agent cost J N [ h ] J^N[h] J N [ h ] , the mean-field cost J M F = J M F [ ( S ) , ( A ) ] J^{MF}=J^{MF}[(S),(A)] J MF = J MF [( S ) , ( A )] , the quantity ζ N \zeta_N ζ N , and the remainder R N R_N R N of that theorem, under its hypotheses - the control set A \mathcal{A} A being convex and A 2 = ∫ [ 0 , T ] E [ ∣ a t ∣ 2 ] d t < ∞ \mathcal{A}_2=\int_{[0,T]}\mathbb{E}[|\mathfrak{a}_t|^2]\,dt<\infty A 2 = ∫ [ 0 , T ] E [ ∣ a t ∣ 2 ] d t < ∞ - with E \mathbb{E} E the expectation and ∣ ⋅ ∣ |\cdot| ∣ ⋅ ∣ the Euclidean norm (Euclidean distance to the origin). Throughout, a real-valued function on a subinterval I I I of the real numbers R \mathbb{R} R is called continuous on I I I when it is continuous relative to I I I , both I I I and the codomain R \mathbb{R} R carrying the metric of the real line . Let Θ \Theta Θ be the aggregate fluctuation covariance of β \beta β , and let Λ = l ( B + K ) l ( l + m ) \Lambda=l(B+K)\sqrt{l(l+m)} Λ = l ( B + K ) l ( l + m ) and g s = N ( b ( Σ s , α s ) − b ( S s , A s ) ) g_s=\sqrt{N}(b(\Sigma_s,\alpha_s)-b(S_s,A_s)) g s = N ( b ( Σ s , α s ) − b ( S s , A s )) be as in the a priori second-moment bound , b b b being the aggregate state drift of β \beta β . For real matrices with any index ranges we use entry notation: x ⋅ M y = ∑ p , q M p q x p y q x\cdot My=\sum_{p,q}M^{pq}x^py^q x ⋅ M y = ∑ p , q M pq x p y q over the matching index ranges (extending the square-matrix convention of the weighted second-moment evolution lemma ), ( M M ′ ) p q = ∑ r M p r M ′ r q (MM')^{pq}=\sum_{r}M^{pr}M'^{rq} ( M M ′ ) pq = ∑ r M p r M ′ r q , and ( M T ) q p = M p q (M^T)^{qp}=M^{pq} ( M T ) qp = M pq . Define, for t ∈ [ 0 , T ] t\in[0,T] t ∈ [ 0 , T ] :
E t δ γ = ∂ γ b ˉ δ ( S t , A t ) , B t δ j = ∂ l + j b ˉ δ ( S t , A t ) ( γ , δ ∈ { 1 , … , l } , j ∈ { 1 , … , m } ) , E^{\delta\gamma}_t=\partial_\gamma\bar{b}^\delta(S_t,A_t),\quad \mathsf{B}^{\delta j}_t=\partial_{l+j}\bar{b}^\delta(S_t,A_t)\quad(\gamma,\delta\in\{1,\dots,l\},\ j\in\{1,\dots,m\}), E t δ γ = ∂ γ b ˉ δ ( S t , A t ) , B t δ j = ∂ l + j b ˉ δ ( S t , A t ) ( γ , δ ∈ { 1 , … , l } , j ∈ { 1 , … , m }) ,
where the sans-serif B t \mathsf{B}_t B t is distinct from the rate bound B B B , and
Q t γ δ = 1 4 ( H γ δ ( t ) + H δ γ ( t ) ) , V t γ j = 1 2 ( H γ , l + j ( t ) + H l + j , γ ( t ) ) , R t i j = 1 4 ( H l + i , l + j ( t ) + H l + j , l + i ( t ) ) , F ^ γ δ = 1 4 ( F γ δ + F δ γ ) . Q^{\gamma\delta}_t=\tfrac{1}{4}\big(H_{\gamma\delta}(t)+H_{\delta\gamma}(t)\big),\quad V^{\gamma j}_t=\tfrac{1}{2}\big(H_{\gamma,l+j}(t)+H_{l+j,\gamma}(t)\big),\quad R^{ij}_t=\tfrac{1}{4}\big(H_{l+i,l+j}(t)+H_{l+j,l+i}(t)\big),\quad \hat{F}^{\gamma\delta}=\tfrac{1}{4}\big(F_{\gamma\delta}+F_{\delta\gamma}\big). Q t γ δ = 4 1 ( H γ δ ( t ) + H δ γ ( t ) ) , V t γj = 2 1 ( H γ , l + j ( t ) + H l + j , γ ( t ) ) , R t ij = 4 1 ( H l + i , l + j ( t ) + H l + j , l + i ( t ) ) , F ^ γ δ = 4 1 ( F γ δ + F δ γ ) .
Hypotheses. (H1) There is a real r > 0 r>0 r > 0 with a ⋅ R t a ≥ r ∣ a ∣ 2 a\cdot R_ta\ge r|a|^2 a ⋅ R t a ≥ r ∣ a ∣ 2 for every t ∈ [ 0 , T ] t\in[0,T] t ∈ [ 0 , T ] and a ∈ R m a\in\mathbb{R}^m a ∈ R m ; by (H1) and conclusion (a) below, whose proof uses only (H1), each R t R_t R t is invertible. (H2) There is a family Z = ( Z t ) t ∈ [ 0 , T ] Z=(Z_t)_{t\in[0,T]} Z = ( Z t ) t ∈ [ 0 , T ] of symmetric real l × l l\times l l × l matrices, continuously differentiable in integral form as in the weighted second-moment evolution lemma , whose densities are
z ˙ γ δ ( t ) = − ( E t T Z t + Z t E t − W t R t − 1 W t T + Q t ) γ δ with W t = Z t B t + 1 2 V t , \dot{z}^{\gamma\delta}(t)=-\Big(E_t^TZ_t+Z_tE_t-W_tR_t^{-1}W_t^T+Q_t\Big)^{\gamma\delta}\qquad\text{with}\qquad W_t=Z_t\mathsf{B}_t+\tfrac{1}{2}V_t, z ˙ γ δ ( t ) = − ( E t T Z t + Z t E t − W t R t − 1 W t T + Q t ) γ δ with W t = Z t B t + 2 1 V t ,
and which satisfies the terminal condition Z T = F ^ Z_T=\hat{F} Z T = F ^ (a backward Riccati equation).
Define u t = a t + R t − 1 W t T s t u_t=\mathfrak{a}_t+R_t^{-1}W_t^T\mathfrak{s}_t u t = a t + R t − 1 W t T s t and e s = g s − E s s s − B s a s e_s=g_s-E_s\mathfrak{s}_s-\mathsf{B}_s\mathfrak{a}_s e s = g s − E s s s − B s a s . Then:
(a) (Coefficients.) Q t Q_t Q t , R t R_t R t , and F ^ \hat{F} F ^ are symmetric; for every t t t and every ( x , a ) ∈ R l × R m (x,a)\in\mathbb{R}^l\times\mathbb{R}^m ( x , a ) ∈ R l × R m , writing z = ( x , a ) ∈ R l + m z=(x,a)\in\mathbb{R}^{l+m} z = ( x , a ) ∈ R l + m ,
1 2 ∑ i , j = 1 l + m H i j ( t ) z i z j = x ⋅ Q t x + x ⋅ V t a + a ⋅ R t a and 1 2 ∑ γ , δ = 1 l F γ δ x γ x δ = x ⋅ F ^ x ; \tfrac{1}{2}\sum_{i,j=1}^{l+m}H_{ij}(t)\,z^iz^j=x\cdot Q_tx+x\cdot V_ta+a\cdot R_ta\qquad\text{and}\qquad \tfrac{1}{2}\sum_{\gamma,\delta=1}^{l}F_{\gamma\delta}\,x^\gamma x^\delta=x\cdot\hat{F}x ; 2 1 i , j = 1 ∑ l + m H ij ( t ) z i z j = x ⋅ Q t x + x ⋅ V t a + a ⋅ R t a and 2 1 γ , δ = 1 ∑ l F γ δ x γ x δ = x ⋅ F ^ x ;
under (H1) each R t R_t R t is symmetric positive definite , hence invertible ; and all entries of E t , B t , Q t , V t , R t , R t − 1 , W t E_t,\mathsf{B}_t,Q_t,V_t,R_t,R_t^{-1},W_t E t , B t , Q t , V t , R t , R t − 1 , W t , and R t − 1 W t T R_t^{-1}W_t^T R t − 1 W t T are continuous in t t t (using the continuity of the matrix inverse ), as are the entries of Z t Z_t Z t by (H2); all are hence bounded on [ 0 , T ] [0,T] [ 0 , T ] by the extreme value theorem : fix reals C Z C_Z C Z and C K C_K C K with ∣ Z t γ δ ∣ ≤ C Z |Z^{\gamma\delta}_t|\le C_Z ∣ Z t γ δ ∣ ≤ C Z and ∣ ( R t − 1 W t T ) j γ ∣ ≤ C K |(R_t^{-1}W_t^T)^{j\gamma}|\le C_K ∣ ( R t − 1 W t T ) jγ ∣ ≤ C K for all indices and t t t .
(b) (Linearization error.) With c e = 3 2 l 3 / 2 ( l + m ) K c_e=\tfrac{3}{2}\,l^{3/2}(l+m)\,K c e = 2 3 l 3/2 ( l + m ) K , at every point of [ 0 , T ] × Ω [0,T]\times\Omega [ 0 , T ] × Ω ,
∣ e s ∣ ≤ c e N − 1 / 2 ( ∣ s s ∣ 2 + ∣ a s ∣ 2 ) . |e_s|\le c_e\,N^{-1/2}\big(|\mathfrak{s}_s|^2+|\mathfrak{a}_s|^2\big). ∣ e s ∣ ≤ c e N − 1/2 ( ∣ s s ∣ 2 + ∣ a s ∣ 2 ) .
(c) (Completion of squares.) All integrals and expectations below are finite, and
L Q G [ ( s ) , ( a ) ] = E [ s 0 ⋅ Z 0 s 0 ] + ∫ [ 0 , T ] E [ u s ⋅ R s u s ] d s + ∫ [ 0 , T ] ( 2 E [ s s ⋅ Z s e s ] + ∑ γ , δ = 1 l Z s γ δ E [ Θ γ δ ( Σ s , α s ) ] ) d s . LQG\big[(\mathfrak{s}),(\mathfrak{a})\big]=\mathbb{E}\big[\mathfrak{s}_0\cdot Z_0\mathfrak{s}_0\big]+\int_{[0,T]}\mathbb{E}\big[u_s\cdot R_su_s\big]\,ds+\int_{[0,T]}\Big(2\,\mathbb{E}\big[\mathfrak{s}_s\cdot Z_se_s\big]+\sum_{\gamma,\delta=1}^{l}Z^{\gamma\delta}_s\,\mathbb{E}\big[\Theta^{\gamma\delta}(\Sigma_s,\alpha_s)\big]\Big)ds . L QG [ ( s ) , ( a ) ] = E [ s 0 ⋅ Z 0 s 0 ] + ∫ [ 0 , T ] E [ u s ⋅ R s u s ] d s + ∫ [ 0 , T ] ( 2 E [ s s ⋅ Z s e s ] + γ , δ = 1 ∑ l Z s γ δ E [ Θ γ δ ( Σ s , α s ) ] ) d s .
(d) (A priori control bound.)
r ∫ [ 0 , T ] E [ ∣ u s ∣ 2 ] d s ≤ N ( J N [ h ] − J M F ) + ∑ γ = 1 l ∣ P 0 γ ∣ ∣ ζ N γ ∣ + ∣ R N ∣ + l C Z E [ ∣ s 0 ∣ 2 ] + 2 ( l − 1 ) B l 2 C Z T + 2 l C Z c e N − 1 / 2 ∫ [ 0 , T ] E [ ∣ s s ∣ ( ∣ s s ∣ 2 + ∣ a s ∣ 2 ) ] d s ; r\int_{[0,T]}\mathbb{E}\big[|u_s|^2\big]ds\ \le\ N\big(J^N[h]-J^{MF}\big)+\sum_{\gamma=1}^{l}|P^\gamma_0||\zeta^\gamma_N|+|R_N|+l\,C_Z\,\mathbb{E}\big[|\mathfrak{s}_0|^2\big]+2(l-1)B\,l^2C_Z\,T+2\,l\,C_Z\,c_e\,N^{-1/2}\int_{[0,T]}\mathbb{E}\Big[|\mathfrak{s}_s|\big(|\mathfrak{s}_s|^2+|\mathfrak{a}_s|^2\big)\Big]ds ; r ∫ [ 0 , T ] E [ ∣ u s ∣ 2 ] d s ≤ N ( J N [ h ] − J MF ) + γ = 1 ∑ l ∣ P 0 γ ∣∣ ζ N γ ∣ + ∣ R N ∣ + l C Z E [ ∣ s 0 ∣ 2 ] + 2 ( l − 1 ) B l 2 C Z T + 2 l C Z c e N − 1/2 ∫ [ 0 , T ] E [ ∣ s s ∣ ( ∣ s s ∣ 2 + ∣ a s ∣ 2 ) ] d s ;
moreover A 2 ≤ 2 ∫ [ 0 , T ] E [ ∣ u s ∣ 2 ] d s + 2 m l C K 2 ∫ [ 0 , T ] E [ ∣ s s ∣ 2 ] d s \mathcal{A}_2\le2\int_{[0,T]}\mathbb{E}[|u_s|^2]ds+2\,m\,l\,C_K^2\int_{[0,T]}\mathbb{E}[|\mathfrak{s}_s|^2]ds A 2 ≤ 2 ∫ [ 0 , T ] E [ ∣ u s ∣ 2 ] d s + 2 m l C K 2 ∫ [ 0 , T ] E [ ∣ s s ∣ 2 ] d s , and, with Λ ^ = Λ 1 + 2 m l C K 2 \hat{\Lambda}=\Lambda\sqrt{1+2mlC_K^2} Λ ^ = Λ 1 + 2 m l C K 2 , for every t ∈ [ 0 , T ] t\in[0,T] t ∈ [ 0 , T ] ,
E [ ∣ s t ∣ 2 ] ≤ ( 3 E [ ∣ s 0 ∣ 2 ] + 6 l ( l − 1 ) B T + 6 T Λ ^ 2 ∫ [ 0 , T ] E [ ∣ u s ∣ 2 ] d s ) exp ( 3 T Λ ^ 2 t ) . \mathbb{E}\big[|\mathfrak{s}_t|^2\big]\ \le\ \Big(3\,\mathbb{E}\big[|\mathfrak{s}_0|^2\big]+6\,l(l-1)BT+6\,T\hat{\Lambda}^2\int_{[0,T]}\mathbb{E}\big[|u_s|^2\big]ds\Big)\exp\big(3\,T\hat{\Lambda}^2t\big). E [ ∣ s t ∣ 2 ] ≤ ( 3 E [ ∣ s 0 ∣ 2 ] + 6 l ( l − 1 ) BT + 6 T Λ ^ 2 ∫ [ 0 , T ] E [ ∣ u s ∣ 2 ] d s ) exp ( 3 T Λ ^ 2 t ) .