TheoremBase

Uniform Convergence in Probability of the N-Agent Empirical State Measure to the Unique Optimal Mean-Field Trajectory

Statement

Adopt the setting, hypotheses and notation of the convergence theorem for the optimal NN-agent value. In particular: (β0,β1)(\beta_{0},\beta_{1}) is an affine-controlled transition-rate family on ll states with control set A⊆Rm\mathcal{A}\subseteq\mathbb{R}^{m}, with transition-rate family β\beta, rate bound BB and state-Lipschitz constant Λb\Lambda_{b}; β~\tilde{\beta}, T>0T>0 and the population cost data (L,G)(L,G) are as there; Δl\Delta^{l} is the probability simplex, ∣⋅∣|\cdot| the Euclidean norm, UA\mathcal{U}_{\mathcal{A}} the set of A\mathcal{A}-valued controls and S(z0,ξ)S(z_{0},\xi) the mean-field flow of claim 2 of the flow stability lemma; x0∈Δlx_{0}\in\Delta^{l} is the initial state, and Mx0∗\mathcal{M}^{*}_{x_{0}} and Sx0∗\mathcal{S}^{*}_{x_{0}} are the set of optimal mean-field controls and the set of optimal mean-field trajectories from x0x_{0} in the sense of the definition of the optimal value, the optimal controls and the optimal trajectories of the mean-field problem; and DD is the deviation from the optimal set of claim 2 of the optimal-set structure lemma. For every natural number NN there are given the NN-agent driving system (ΩN,FN,PN)(\Omega^{N},\mathcal{F}^{N},P^{N}) with expectation EN\mathbb{E}^{N}, the A\mathcal{A}-valued policy hNh^{N}, the solution of the controlled NN-agent dynamics with empirical state measure ΣN\Sigma^{N}, and the realized control α^N\hat{\alpha}^{N} of the realized-control lemma; the sequence (ηN)N∈N(\eta_{N})_{N\in\mathbb{N}} and the hypothesis that (Σ0N)N∈N(\Sigma^{N}_{0})_{N\in\mathbb{N}} converges in distribution to Yx0Y_{x_{0}} are as there, and DND_{N} is the random variable of claim 4 of that theorem, so that DN(ω)=D(Σ0N(ω),α^N(ω))D_{N}(\omega)=D(\Sigma^{N}_{0}(\omega),\hat{\alpha}^{N}(\omega)).

Assume in addition that the set Sx0∗\mathcal{S}^{*}_{x_{0}} has exactly one element, and write S∗S^{*} for that element, so that Sx0∗={S∗}\mathcal{S}^{*}_{x_{0}}=\{S^{*}\}; it is a map on [0,T][0,T] with St∗∈ΔlS^{*}_{t}\in\Delta^{l} for every t∈[0,T]t\in[0,T], by claim 2 of the flow stability lemma.

For every natural number NN let Ω∗N∈FN\Omega^{N}_{*}\in\mathcal{F}^{N} and M‾N\overline{M}^{N} be the event and the random variable furnished by the martingale bound for the empirical state measure, applied to the NN-th solution: thus PN(Ω∗N)=1P^{N}(\Omega^{N}_{*})=1, the path t↦ΣtN(ω)t\mapsto\Sigma^{N}_{t}(\omega) is right-continuous at every t∈[0,T)t\in[0,T) for every ω∈Ω∗N\omega\in\Omega^{N}_{*}, and

EN[(M‾N)2]≤8 l (l−1) B TN.\mathbb{E}^{N}\bigl[(\overline{M}^{N})^{2}\bigr]\le\frac{8\,l\,(l-1)\,B\,T}{N}.

Write 1Ω∗N\mathbf{1}_{\Omega^{N}_{*}} for the function on ΩN\Omega^{N} equal to 11 on Ω∗N\Omega^{N}_{*} and to 00 elsewhere.

For a natural number nn put

Qn={0}∪{j2n T  :  j∈{1,…,2n}}⊆[0,T],Q_{n}=\{0\}\cup\Bigl\{\tfrac{j}{2^{n}}\,T\;:\;j\in\{1,\dots,2^{n}\}\Bigr\}\subseteq[0,T],

a finite nonempty set, where a natural number is read in R\mathbb{R} through its image in the real field. For natural numbers NN and nn let gnNg^{N}_{n} be the map on ΩN\Omega^{N} given by

gnN(ω)=(max⁡{ ∣ΣtN(ω)−St∗∣  :  t∈Qn }) 1Ω∗N(ω),g^{N}_{n}(\omega)=\Bigl(\max\bigl\{\,\bigl|\Sigma^{N}_{t}(\omega)-S^{*}_{t}\bigr|\;:\;t\in Q_{n}\,\bigr\}\Bigr)\,\mathbf{1}_{\Omega^{N}_{*}}(\omega),

the maximum of a finite nonempty set of real numbers. For every ω∈ΩN\omega\in\Omega^{N} the set {gnN(ω):n∈N}\bigl\{g^{N}_{n}(\omega):n\in\mathbb{N}\bigr\} is a nonempty set of real numbers bounded above by 22: the vectors ΣtN(ω)\Sigma^{N}_{t}(\omega) and St∗S^{*}_{t} lie in Δl\Delta^{l} and hence have norm at most 11 by claim 1 of the simplex compactness lemma, so ∣ΣtN(ω)−St∗∣≤2|\Sigma^{N}_{t}(\omega)-S^{*}_{t}|\le2 and therefore 0≤gnN(ω)≤20\le g^{N}_{n}(\omega)\le2. Hence that set has a least upper bound; write

WN(ω)=sup⁡{ gnN(ω)  :  n∈N }W_{N}(\omega)=\sup\bigl\{\,g^{N}_{n}(\omega)\;:\;n\in\mathbb{N}\,\bigr\}

for it.

Then the following hold.

1. (The deviation measures the distance to the optimal trajectory.) S(x0,ζ)=S∗S(x_{0},\zeta)=S^{*} for every ζ∈Mx0∗\zeta\in\mathcal{M}^{*}_{x_{0}}. Moreover, for every z0∈Δlz_{0}\in\Delta^{l} and every ξ∈UA\xi\in\mathcal{U}_{\mathcal{A}} the supremum

sup⁡{ ∣St(z0,ξ)−St∗∣  :  t∈[0,T] }\sup\bigl\{\,\bigl|S_{t}(z_{0},\xi)-S^{*}_{t}\bigr|\;:\;t\in[0,T]\,\bigr\}

exists and is equal to D(z0,ξ)D(z_{0},\xi).

2. (The uniform deviation is a random variable.) For all natural numbers NN and nn the map gnNg^{N}_{n} is a random variable on (ΩN,FN,PN)(\Omega^{N},\mathcal{F}^{N},P^{N}) with 0≤gnN(ω)≤20\le g^{N}_{n}(\omega)\le2 for every ω∈ΩN\omega\in\Omega^{N}. For every natural number NN the supremum defining WN(ω)W_{N}(\omega) exists for every ω∈ΩN\omega\in\Omega^{N}, and WNW_{N} is a random variable with 0≤WN(ω)≤20\le W_{N}(\omega)\le2 for every ω∈ΩN\omega\in\Omega^{N}. Moreover, for every ω∈Ω∗N\omega\in\Omega^{N}_{*} the supremum

sup⁡{ ∣ΣtN(ω)−St∗∣  :  t∈[0,T] }\sup\bigl\{\,\bigl|\Sigma^{N}_{t}(\omega)-S^{*}_{t}\bigr|\;:\;t\in[0,T]\,\bigr\}

exists and is equal to WN(ω)W_{N}(\omega).

3. (Uniform convergence to the optimal trajectory in probability.) For every real ε>0\varepsilon>0 and every natural number NN,

PN({ω∈ΩN  :  ε≤WN(ω)})  ≤  32 l (l−1) B T e2ΛbTN ε2  +  PN({ω∈ΩN  :  ε/2≤DN(ω)}),P^{N}\bigl(\{\omega\in\Omega^{N}\;:\;\varepsilon\le W_{N}(\omega)\}\bigr)\;\le\;\frac{32\,l\,(l-1)\,B\,T\,e^{2\Lambda_{b}T}}{N\,\varepsilon^{2}}\;+\;P^{N}\bigl(\{\omega\in\Omega^{N}\;:\;\varepsilon/2\le D_{N}(\omega)\}\bigr),

and the sequence

(PN({ω∈ΩN  :  ε≤WN(ω)}))N∈N\Bigl(P^{N}\bigl(\{\omega\in\Omega^{N}\;:\;\varepsilon\le W_{N}(\omega)\}\bigr)\Bigr)_{N\in\mathbb{N}}

converges to 00.

Proofs

Log in to submit a proof.

Loading...

Citations

Loading…

Dependencies

Loading…

Related

0 relations

Curated associations between results. These are editable and subjective — they do not replace the dependency graph, which is derived from the references in the text.

No relations recorded yet.

Comments

Log in to comment.

Loading…