TheoremBase

Uniform Mean-Square Bound for the Martingale Part of the Empirical State Measure

Statement

Adopt the setting of the controlled NN-agent dynamics with NN agents, ll states and control dimension mm, and let A\mathcal{A} be a nonempty subset of Euclidean space Rm\mathbb{R}^m: a transition-rate family β\beta on ll states with control set A\mathcal{A} and rate bound BB, an observation-rate family β~\tilde{\beta} with l~\tilde{l} channels, a horizon T>0T>0, an NN-agent driving system (Ω,F,P)(\Omega,\mathcal{F},P), an observation-driven control policy hh which is A\mathcal{A}-valued, and a solution on [0,T][0,T] with empirical state measure Σ\Sigma, control α\alpha and regular event Ω0\Omega_0, which exists by the existence and uniqueness theorem. Let bb be the aggregate state drift of β\beta and let M=(M1,…,Ml)M=(M^1,\dots,M^l) be the martingale part of the martingale decomposition of the empirical state measure, and write 1Ω0\mathbf{1}_{\Omega_0} for the function equal to 11 on Ω0\Omega_0 and 00 off it, so that

Mtγ=Σtγ−Σ0γ−∫[0,t]1Ω0 bγ(Σs,αs) ds(t∈[0,T], γ∈{1,…,l}).M^\gamma_t=\Sigma^\gamma_t-\Sigma^\gamma_0-\int_{[0,t]}\mathbf{1}_{\Omega_0}\,b^\gamma(\Sigma_s,\alpha_s)\,ds\qquad(t\in[0,T],\ \gamma\in\{1,\dots,l\}) .

Let DD be the set of dyadic partition points of [0,T][0,T] as in the supremum lemma for bounded right-continuous processes, and set KM=2+2(l−1)BTK_M=2+2(l-1)BT.

Then there is an event Ω∗∈F\Omega_*\in\mathcal{F} with Ω∗⊆Ω0\Omega_*\subseteq\Omega_0 and P(Ω∗)=1P(\Omega_*)=1 such that, for every ω∈Ω∗\omega\in\Omega_* and every γ\gamma, the path t↦Σtγ(ω)t\mapsto\Sigma^\gamma_t(\omega) is right-continuous at every t∈[0,T)t\in[0,T), and the path t↦Mtγ(ω)t\mapsto M^\gamma_t(\omega) satisfies ∣Mtγ(ω)∣≤KM|M^\gamma_t(\omega)|\le K_M for all t∈[0,T]t\in[0,T] and is right-continuous at every t∈[0,T)t\in[0,T). Moreover, for any such event Ω∗\Omega_*, writing 1Ω∗\mathbf{1}_{\Omega_*} for the function equal to 11 on Ω∗\Omega_* and 00 elsewhere, and

Mγ‾=sup⁡t∈D(∣Mtγ∣ 1Ω∗),M‾=(∑γ=1l(Mγ‾)2)1/2,\overline{M^\gamma}=\sup_{t\in D}\big(|M^\gamma_t|\,\mathbf{1}_{\Omega_*}\big),\qquad \overline{M}=\Big(\sum_{\gamma=1}^l\big(\overline{M^\gamma}\big)^2\Big)^{1/2},

each Mγ‾\overline{M^\gamma} and M‾\overline{M} is a random variable, one has ∣Mt(ω)∣≤M‾(ω)|M_t(\omega)|\le\overline{M}(\omega) for every t∈[0,T]t\in[0,T] and every ω∈Ω∗\omega\in\Omega_*, where ∣Mt∣|M_t| denotes the Euclidean norm of the vector Mt=(Mt1,…,Mtl)M_t=(M^1_t,\dots,M^l_t) whereas ∣Mtγ∣|M^\gamma_t| above denotes the absolute value of a real number, and

E[M‾ 2]≤8 l (l−1) B TN.\mathbb{E}\big[\overline{M}^{\,2}\big]\le\frac{8\,l\,(l-1)\,B\,T}{N} .

Proofs

Log in to submit a proof.

Loading...

Citations

Loading…

Dependencies

Loading…

Related

0 relations

Curated associations between results. These are editable and subjective — they do not replace the dependency graph, which is derived from the references in the text.

No relations recorded yet.

Comments

Log in to comment.

Loading…