Adopt the setting of the \reftext{def:n-agent-fluctuation-processes-2026a}{fluctuation processes of the controlled N N N -agent dynamics}: a \reftext{def:transition-rate-family-2026a}{transition-rate family} Ξ² \beta Ξ² with rate bound B B B on l l l states with control dimension m m m , an \reftext{def:observation-rate-family-2026a}{observation-rate family} Ξ² ~ \tilde{\beta} Ξ² ~ β , a horizon T > 0 T>0 T > 0 , an \reftext{def:n-agent-driving-system-2026a}{N N N -agent driving system}, an \reftext{def:observation-driven-control-policy-2026a}{observation-driven control policy} h h h , a \reftext{def:n-agent-controlled-dynamics-2026a}{solution} on [ 0 , T ] [0,T] [ 0 , T ] with regular event Ξ© 0 \Omega_0 Ξ© 0 β , empirical state measure Ξ£ t \Sigma_t Ξ£ t β , and control Ξ± t \alpha_t Ξ± t β , a \reftext{def:mean-field-trajectory-pair-2026a}{mean-field trajectory pair} ( S , A ) (S,A) ( S , A ) for Ξ² \beta Ξ² with horizon T T T , and the fluctuation processes s t = N ( Ξ£ t β S t ) \mathfrak{s}_t=\sqrt{N}(\Sigma_t-S_t) s t β = N β ( Ξ£ t β β S t β ) and a t = N ( Ξ± t β A t ) \mathfrak{a}_t=\sqrt{N}(\alpha_t-A_t) a t β = N β ( Ξ± t β β A t β ) . Let ( L , G ) (L,G) ( L , G ) be \reftext{def:population-cost-data-2026a}{population cost data} on l l l states with control dimension m m m , with \reftext{def:n-agent-cost-2026a}{N N N -agent cost} J N [ h ] J^N[h] J N [ h ] and \reftext{def:mean-field-cost-2026a}{mean-field cost} J M F [ ( S ) , ( A ) ] J^{MF}[(S),(A)] J MF [( S ) , ( A )] . Let ( U , Ξ² Λ ) (U,\bar{\beta}) ( U , Ξ² Λ β ) be a \reftext{def:c2-transition-rate-extension-2026a}{twice continuously differentiable extension} of Ξ² \beta Ξ² with derivative bound K K K and extended aggregate state drift b Λ \bar{b} b Λ , let ( V , L Λ , G Λ ) (V,\bar{L},\bar{G}) ( V , L Λ , G Λ ) be a \reftext{def:c2-population-cost-extension-2026a}{twice continuously differentiable extension} of ( L , G ) (L,G) ( L , G ) with second-derivative bound K c K_c K c β , and let P P P be a \reftext{def:stationary-mean-field-triple-2026a}{stationary co-state} for these data, so that ( S , A , P ) (S,A,P) ( S , A , P ) is a stationary mean-field triple. Adopt the partial-derivative notation β j β i \partial_j\partial_i β j β β i β of the extension definitions, write d d d for the \reftext{def:euclidean-distance-rn-2026a}{Euclidean distance}, β£ β
β£ |\cdot| β£ β
β£ for the Euclidean norm, E \mathbb{E} E for the \reftext{def:expectation-variance-2026a}{expectation}, and Ξ l \Delta^l Ξ l for the \reftext{def:probability-simplex-2026a}{probability simplex}, and let C P C_P C P β be a real number with β Ξ΄ = 1 l β£ P t Ξ΄ β£ β€ C P \sum_{\delta=1}^{l}|P^\delta_t|\le C_P β Ξ΄ = 1 l β β£ P t Ξ΄ β β£ β€ C P β for all t β [ 0 , T ] t\in[0,T] t β [ 0 , T ] , which exists because each component of P P P is continuous and hence \reftext{lem:continuous-compact-interval-bounded-2026a}{bounded}.
\textbf{Hypothesis.} Assume A = β« [ 0 , T ] E [ β£ a t β£ 2 ] β d t < β \mathcal{A}=\int_{[0,T]}\mathbb{E}[|\mathfrak{a}_t|^2]\,dt<\infty A = β« [ 0 , T ] β E [ β£ a t β β£ 2 ] d t < β , this integral being well defined by part (a) of the \reftext{lem:fluctuation-state-moment-bound-2026a}{a priori second-moment bound}.
For u β₯ 0 u\ge0 u β₯ 0 define
Ο L ( u ) = sup β‘ { β£ β j β i L Λ ( x ) β β j β i L Λ ( y ) β£ : Β i , j β { 1 , β¦ , l + m } , Β x , y β Ξ l Γ R m , Β d ( x , y ) β€ u } , \omega_L(u)=\sup\big\{|\partial_j\partial_i\bar{L}(x)-\partial_j\partial_i\bar{L}(y)|:\ i,j\in\{1,\dots,l+m\},\ x,y\in\Delta^l\times\mathbb{R}^m,\ d(x,y)\le u\big\}, Ο L β ( u ) = sup { β£ β j β β i β L Λ ( x ) β β j β β i β L Λ ( y ) β£ : Β i , j β { 1 , β¦ , l + m } , Β x , y β Ξ l Γ R m , Β d ( x , y ) β€ u } ,
Ο b ( u ) = sup β‘ { β£ β j β i b Λ Ξ³ ( x ) β β j β i b Λ Ξ³ ( y ) β£ : Β Ξ³ β { 1 , β¦ , l } , Β i , j β { 1 , β¦ , l + m } , Β x , y β Ξ l Γ R m , Β d ( x , y ) β€ u } , \omega_b(u)=\sup\big\{|\partial_j\partial_i\bar{b}^\gamma(x)-\partial_j\partial_i\bar{b}^\gamma(y)|:\ \gamma\in\{1,\dots,l\},\ i,j\in\{1,\dots,l+m\},\ x,y\in\Delta^l\times\mathbb{R}^m,\ d(x,y)\le u\big\}, Ο b β ( u ) = sup { β£ β j β β i β b Λ Ξ³ ( x ) β β j β β i β b Λ Ξ³ ( y ) β£ : Β Ξ³ β { 1 , β¦ , l } , Β i , j β { 1 , β¦ , l + m } , Β x , y β Ξ l Γ R m , Β d ( x , y ) β€ u } ,
Ο G ( u ) = sup β‘ { β£ β Ξ΄ β Ξ³ G Λ ( Ξ£ ) β β Ξ΄ β Ξ³ G Λ ( Ξ£ β² ) β£ : Β Ξ³ , Ξ΄ β { 1 , β¦ , l } , Β Ξ£ , Ξ£ β² β Ξ l , Β d ( Ξ£ , Ξ£ β² ) β€ u } . \omega_G(u)=\sup\big\{|\partial_\delta\partial_\gamma\bar{G}(\Sigma)-\partial_\delta\partial_\gamma\bar{G}(\Sigma')|:\ \gamma,\delta\in\{1,\dots,l\},\ \Sigma,\Sigma'\in\Delta^l,\ d(\Sigma,\Sigma')\le u\big\}. Ο G β ( u ) = sup { β£ β Ξ΄ β β Ξ³ β G Λ ( Ξ£ ) β β Ξ΄ β β Ξ³ β G Λ ( Ξ£ β² ) β£ : Β Ξ³ , Ξ΄ β { 1 , β¦ , l } , Β Ξ£ , Ξ£ β² β Ξ l , Β d ( Ξ£ , Ξ£ β² ) β€ u } .
\textbf{(a) (Moduli.)} Ο L \omega_L Ο L β , Ο b \omega_b Ο b β , and Ο G \omega_G Ο G β are nondecreasing functions from [ 0 , β ) [0,\infty) [ 0 , β ) to [ 0 , β ) [0,\infty) [ 0 , β ) with Ο L β€ 2 K c \omega_L\le2K_c Ο L β β€ 2 K c β , Ο b β€ 6 β l β K \omega_b\le6\,l\,K Ο b β β€ 6 l K , and Ο G β€ 2 K c \omega_G\le2K_c Ο G β β€ 2 K c β everywhere, and for every Ξ΅ > 0 \varepsilon>0 Ξ΅ > 0 there is Ξ΄ > 0 \delta>0 Ξ΄ > 0 such that Ο L ( u ) β€ Ξ΅ \omega_L(u)\le\varepsilon Ο L β ( u ) β€ Ξ΅ , Ο b ( u ) β€ Ξ΅ \omega_b(u)\le\varepsilon Ο b β ( u ) β€ Ξ΅ , and Ο G ( u ) β€ Ξ΅ \omega_G(u)\le\varepsilon Ο G β ( u ) β€ Ξ΅ for all u β [ 0 , Ξ΄ ] u\in[0,\delta] u β [ 0 , Ξ΄ ] .
\textbf{(b) (Finiteness and eligibility.)} J N [ h ] J^N[h] J N [ h ] is finite; the pair ( ( s ) , ( a ) ) ((\mathfrak{s}),(\mathfrak{a})) (( s ) , ( a )) satisfies requirements (i)-(iii) of the \reftext{def:fluctuation-lqg-cost-2026a}{fluctuation linear-quadratic cost} (requirement (i) holding with Ξ© 1 = Ξ© 0 \Omega_1=\Omega_0 Ξ© 1 β = Ξ© 0 β by the \reftext{lem:n-agent-joint-measurability-2026a}{joint measurability of the state and control}), so that L Q G [ ( s ) , ( a ) ] LQG[(\mathfrak{s}),(\mathfrak{a})] L QG [( s ) , ( a )] is a well-defined real number; and all expectations and integrals appearing in (c) are well defined and finite.
\textbf{(c) (Expansion with quantitative remainder.)} Set ΞΆ N = N β ( E [ Ξ£ 0 ] β S 0 ) β R l \zeta_N=N\,(\mathbb{E}[\Sigma_0]-S_0)\in\mathbb{R}^l ΞΆ N β = N ( E [ Ξ£ 0 β ] β S 0 β ) β R l , with the componentwise expectation, and set Ο t = d ( ( Ξ£ t , Ξ± t ) , ( S t , A t ) ) \rho_t=d\big((\Sigma_t,\alpha_t),(S_t,A_t)\big) Ο t β = d ( ( Ξ£ t β , Ξ± t β ) , ( S t β , A t β ) ) for t β [ 0 , T ] t\in[0,T] t β [ 0 , T ] , so that Ο t = N β 1 / 2 ( β£ s t β£ 2 + β£ a t β£ 2 ) 1 / 2 \rho_t=N^{-1/2}\big(|\mathfrak{s}_t|^2+|\mathfrak{a}_t|^2\big)^{1/2} Ο t β = N β 1/2 ( β£ s t β β£ 2 + β£ a t β β£ 2 ) 1/2 and d ( Ξ£ T , S T ) = N β 1 / 2 β£ s T β£ d(\Sigma_T,S_T)=N^{-1/2}|\mathfrak{s}_T| d ( Ξ£ T β , S T β ) = N β 1/2 β£ s T β β£ . Then the real number R N R_N R N β defined by the identity
N ( J N [ h ] β J M F [ ( S ) , ( A ) ] ) = L Q G [ ( s ) , ( a ) ] β β Ξ³ = 1 l P 0 Ξ³ β ΞΆ N Ξ³ + R N N\Big(J^N[h]-J^{MF}[(S),(A)]\Big)=LQG\big[(\mathfrak{s}),(\mathfrak{a})\big]-\sum_{\gamma=1}^{l}P^\gamma_0\,\zeta^\gamma_N+R_N N ( J N [ h ] β J MF [( S ) , ( A )] ) = L QG [ ( s ) , ( a ) ] β Ξ³ = 1 β l β P 0 Ξ³ β ΞΆ N Ξ³ β + R N β
satisfies
β£ R N β£ Β β€ Β l + m 2 β« [ 0 , T ] E [ ( Ο L ( Ο t ) + C P β Ο b ( Ο t ) ) ( β£ s t β£ 2 + β£ a t β£ 2 ) ] β d t Β + Β l 2 β E [ Ο G ( d ( Ξ£ T , S T ) ) β β£ s T β£ 2 ] . |R_N|\ \le\ \frac{l+m}{2}\int_{[0,T]}\mathbb{E}\Big[\big(\omega_L(\rho_t)+C_P\,\omega_b(\rho_t)\big)\big(|\mathfrak{s}_t|^2+|\mathfrak{a}_t|^2\big)\Big]\,dt\ +\ \frac{l}{2}\,\mathbb{E}\Big[\omega_G\big(d(\Sigma_T,S_T)\big)\,|\mathfrak{s}_T|^2\Big]. β£ R N β β£ Β β€ Β 2 l + m β β« [ 0 , T ] β E [ ( Ο L β ( Ο t β ) + C P β Ο b β ( Ο t β ) ) ( β£ s t β β£ 2 + β£ a t β β£ 2 ) ] d t Β + Β 2 l β E [ Ο G β ( d ( Ξ£ T β , S T β ) ) β£ s T β β£ 2 ] .