Let l l l and m m m be natural numbers with l ≥ 2 l\ge2 l ≥ 2 and m ≥ 1 m\ge1 m ≥ 1 , let ( β 0 , β 1 ) (\beta_{0},\beta_{1}) ( β 0 , β 1 ) be an affine-controlled transition-rate family on l l l states with control set A ⊆ R m \mathcal{A}\subseteq\mathbb{R}^{m} A ⊆ R m and let β \beta β be its transition-rate family , with rate bound B B B , aggregate state drift b b b and state-Lipschitz constant Λ b \Lambda_{b} Λ b , so that ∣ b ( Σ , a ) − b ( Σ ′ , a ) ∣ ≤ Λ b ∣ Σ − Σ ′ ∣ |b(\Sigma,a)-b(\Sigma',a)|\le\Lambda_{b}|\Sigma-\Sigma'| ∣ b ( Σ , a ) − b ( Σ ′ , a ) ∣ ≤ Λ b ∣Σ − Σ ′ ∣ for all Σ , Σ ′ \Sigma,\Sigma' Σ , Σ ′ in the probability simplex Δ l \Delta^{l} Δ l and every a ∈ A a\in\mathcal{A} a ∈ A by claim 4 there, where ∣ ⋅ ∣ |\cdot| ∣ ⋅ ∣ is the Euclidean norm . As part of the data of the rate family, A \mathcal{A} A is a nonempty compact convex subset of Euclidean space R m \mathbb{R}^{m} R m .
Let β ~ \tilde{\beta} β ~ be an observation-rate family on l l l states with l ~ \tilde{l} l ~ channels, where l ~ \tilde{l} l ~ is a natural number with l ~ ≥ 1 \tilde{l}\ge1 l ~ ≥ 1 , let T > 0 T>0 T > 0 be a real number , and let ( L , G ) (L,G) ( L , G ) be population cost data on l l l states with control dimension m m m that is convex in the control on A \mathcal{A} A , meaning that for every Σ ∈ Δ l \Sigma\in\Delta^{l} Σ ∈ Δ l the function sending a a a to L ( Σ , a ) L(\Sigma,a) L ( Σ , a ) is convex on A \mathcal{A} A ; these are the standing hypotheses of the boundedness, lower-semicontinuity and attainment theorem for the mean-field cost . Let C C C be the bound of claim 1 of the lemma on cost data over a compact control set , so that ∣ L ( Σ , a ) ∣ ≤ C |L(\Sigma,a)|\le C ∣ L ( Σ , a ) ∣ ≤ C and ∣ G ( Σ ) ∣ ≤ C |G(\Sigma)|\le C ∣ G ( Σ ) ∣ ≤ C for every Σ ∈ Δ l \Sigma\in\Delta^{l} Σ ∈ Δ l and every a ∈ A a\in\mathcal{A} a ∈ A , and let C F C_{F} C F be the bound of claim 1 of the boundedness, lower-semicontinuity and attainment theorem . Let S ( x 0 , ξ ) S(x_{0},\xi) S ( x 0 , ξ ) be the mean-field flow of claim 2 of the flow stability lemma and let F F F be the mean-field cost of a control from an initial state under ( L , G ) (L,G) ( L , G ) . Let U A \mathcal{U}_{\mathcal{A}} U A be the set of A \mathcal{A} A -valued controls and let ρ \rho ρ , d Δ d_{\Delta} d Δ , X X X and d X d_{X} d X be as in the compactness lemma for the simplex, the control set and their product . Write λ [ 0 , T ] \lambda_{[0,T]} λ [ 0 , T ] for the restricted Lebesgue measure on [ 0 , T ] [0,T] [ 0 , T ] .
Let N N N be a natural number with N ≥ 1 N\ge1 N ≥ 1 , let ( Ω , F , P ) (\Omega,\mathcal{F},P) ( Ω , F , P ) be an N N N -agent driving system , let h h h be an observation-driven control policy with horizon T T T , control dimension m m m and l ~ \tilde{l} l ~ channels that is A \mathcal{A} A -valued , and let ( σ i , Υ υ , α ) (\sigma^{i},\Upsilon^{\upsilon},\alpha) ( σ i , Υ υ , α ) be a solution of the controlled N N N -agent dynamics on [ 0 , T ] [0,T] [ 0 , T ] for β \beta β , β ~ \tilde{\beta} β ~ , this driving system and this policy, with regular event Ω 0 \Omega_{0} Ω 0 and empirical state measure Σ \Sigma Σ . Let α ^ \hat{\alpha} α ^ be the realized control of the realized-control lemma , so that α ^ ( ω ) ∈ U A \hat{\alpha}(\omega)\in\mathcal{U}_{\mathcal{A}} α ^ ( ω ) ∈ U A for every ω ∈ Ω \omega\in\Omega ω ∈ Ω , the pair ( Σ 0 , α ^ ) (\Sigma_{0},\hat{\alpha}) ( Σ 0 , α ^ ) is a random element of ( X , d X ) (X,d_{X}) ( X , d X ) , and α ^ ( t , ω ) = α t ( ω ) \hat{\alpha}(t,\omega)=\alpha_{t}(\omega) α ^ ( t , ω ) = α t ( ω ) for every t ∈ [ 0 , T ] t\in[0,T] t ∈ [ 0 , T ] and every ω ∈ Ω 0 \omega\in\Omega_{0} ω ∈ Ω 0 .
Let M M M be the martingale part of the martingale decomposition of the empirical state measure , and let Ω ∗ ⊆ Ω 0 \Omega_{*}\subseteq\Omega_{0} Ω ∗ ⊆ Ω 0 and M ‾ \overline{M} M be the event of probability 1 1 1 and the random variable furnished by the martingale bound for the empirical state measure , so that ∣ M t ( ω ) ∣ ≤ M ‾ ( ω ) |M_{t}(\omega)|\le\overline{M}(\omega) ∣ M t ( ω ) ∣ ≤ M ( ω ) for every t ∈ [ 0 , T ] t\in[0,T] t ∈ [ 0 , T ] and every ω ∈ Ω ∗ \omega\in\Omega_{*} ω ∈ Ω ∗ , the path t ↦ Σ t ( ω ) t\mapsto\Sigma_{t}(\omega) t ↦ Σ t ( ω ) is right-continuous on [ 0 , T ) [0,T) [ 0 , T ) for every ω ∈ Ω ∗ \omega\in\Omega_{*} ω ∈ Ω ∗ , and E [ M ‾ 2 ] ≤ 8 l ( l − 1 ) B T / N \mathbb{E}\bigl[\overline{M}^{\,2}\bigr]\le 8l(l-1)BT/N E [ M 2 ] ≤ 8 l ( l − 1 ) BT / N . Let J N [ h ] J^{N}[h] J N [ h ] be the N N N -agent cost of the policy h h h .
Then the following hold.
1. (Pathwise comparison of the flows.) Let ω ∈ Ω ∗ \omega\in\Omega_{*} ω ∈ Ω ∗ and write S ω = S ( Σ 0 ( ω ) , α ^ ( ω ) ) S^{\omega}=S\bigl(\Sigma_{0}(\omega),\hat{\alpha}(\omega)\bigr) S ω = S ( Σ 0 ( ω ) , α ^ ( ω ) ) for the mean-field flow started at the realized initial state and driven by the realized control. Then
∣ Σ t ( ω ) − S t ω ∣ ≤ e Λ b T M ‾ ( ω ) for every t ∈ [ 0 , T ] . \bigl|\Sigma_{t}(\omega)-S^{\omega}_{t}\bigr|\le e^{\Lambda_{b}T}\,\overline{M}(\omega)\qquad\text{for every }t\in[0,T]. Σ t ( ω ) − S t ω ≤ e Λ b T M ( ω ) for every t ∈ [ 0 , T ] .
2. (The realized mean-field cost, and finiteness of the N N N -agent cost.) The map ω ↦ F ( Σ 0 ( ω ) , α ^ ( ω ) ) \omega\mapsto F\bigl(\Sigma_{0}(\omega),\hat{\alpha}(\omega)\bigr) ω ↦ F ( Σ 0 ( ω ) , α ^ ( ω ) ) is a random variable , and ∣ F ( Σ 0 ( ω ) , α ^ ( ω ) ) ∣ ≤ C F \bigl|F\bigl(\Sigma_{0}(\omega),\hat{\alpha}(\omega)\bigr)\bigr|\le C_{F} F ( Σ 0 ( ω ) , α ^ ( ω ) ) ≤ C F for every ω ∈ Ω \omega\in\Omega ω ∈ Ω . Moreover, fix a point e ∈ Δ l e\in\Delta^{l} e ∈ Δ l and let Σ ∗ : [ 0 , T ] × Ω → R l \Sigma^{*}:[0,T]\times\Omega\to\mathbb{R}^{l} Σ ∗ : [ 0 , T ] × Ω → R l be the map with Σ t ∗ ( ω ) = Σ t ( ω ) \Sigma^{*}_{t}(\omega)=\Sigma_{t}(\omega) Σ t ∗ ( ω ) = Σ t ( ω ) for ω ∈ Ω ∗ \omega\in\Omega_{*} ω ∈ Ω ∗ and Σ t ∗ ( ω ) = e \Sigma^{*}_{t}(\omega)=e Σ t ∗ ( ω ) = e for ω ∉ Ω ∗ \omega\notin\Omega_{*} ω ∈ / Ω ∗ , so that Σ t ∗ ( ω ) ∈ Δ l \Sigma^{*}_{t}(\omega)\in\Delta^{l} Σ t ∗ ( ω ) ∈ Δ l for every t t t and every ω \omega ω . Then
W ( ω ) = ∫ [ 0 , T ] L ( Σ t ∗ ( ω ) , α ^ ( t , ω ) ) d λ [ 0 , T ] ( t ) + G ( Σ T ∗ ( ω ) ) W(\omega)=\int_{[0,T]}L\bigl(\Sigma^{*}_{t}(\omega),\hat{\alpha}(t,\omega)\bigr)\,d\lambda_{[0,T]}(t)+G\bigl(\Sigma^{*}_{T}(\omega)\bigr) W ( ω ) = ∫ [ 0 , T ] L ( Σ t ∗ ( ω ) , α ^ ( t , ω ) ) d λ [ 0 , T ] ( t ) + G ( Σ T ∗ ( ω ) )
defines a random variable with ∣ W ( ω ) ∣ ≤ C ( T + 1 ) |W(\omega)|\le C(T+1) ∣ W ( ω ) ∣ ≤ C ( T + 1 ) for every ω ∈ Ω \omega\in\Omega ω ∈ Ω , the N N N -agent cost J N [ h ] J^{N}[h] J N [ h ] is a real number, and J N [ h ] = E [ W ] J^{N}[h]=\mathbb{E}[W] J N [ h ] = E [ W ] .
3. (Uniform comparison of the costs.) For every real ε > 0 \varepsilon>0 ε > 0 there is a natural number N 0 N_{0} N 0 , depending only on ε \varepsilon ε , l l l , B B B , T T T , Λ b \Lambda_{b} Λ b , A \mathcal{A} A , L L L and G G G , and not on N N N , on the driving system, on the policy or on the solution, such that
∣ J N [ h ] − E [ F ( Σ 0 , α ^ ) ] ∣ ≤ ε \Bigl|J^{N}[h]-\mathbb{E}\bigl[F\bigl(\Sigma_{0},\hat{\alpha}\bigr)\bigr]\Bigr|\le\varepsilon J N [ h ] − E [ F ( Σ 0 , α ^ ) ] ≤ ε
holds for every N ≥ N 0 N\ge N_{0} N ≥ N 0 , every N N N -agent driving system, every A \mathcal{A} A -valued observation-driven control policy h h h with horizon T T T , control dimension m m m and l ~ \tilde{l} l ~ channels, and every solution of the controlled N N N -agent dynamics for those data, with Σ 0 \Sigma_{0} Σ 0 and α ^ \hat{\alpha} α ^ the initial empirical state measure and the realized control of that solution.