TheoremBase

Proof

Let Mγ‾\overline{M^\gamma} and M‾\overline{M} be as in the martingale bound for the empirical state measure for the given event Ω∗\Omega_*, so that P(Ω∗)=1P(\Omega_*)=1, Ω∗⊆Ω0\Omega_*\subseteq\Omega_0 with Ω0\Omega_0 the regular event, and ∣Mt(ω)∣≤M‾(ω)|M_t(\omega)|\le\overline{M}(\omega) for every t∈[0,T]t\in[0,T] and ω∈Ω∗\omega\in\Omega_*.

Step 1: Ψ\Psi is a random variable. Both Σt\Sigma_t and StS_t lie in the probability simplex, so ∣Σt∣≤1|\Sigma_t|\le1 and ∣St∣≤1|S_t|\le1, whence ∣Σt−St∣≤2|\Sigma_t-S_t|\le2 for every tt and every ω\omega, by the triangle inequality. For ω∈Ω∗\omega\in\Omega_* each path t↦Σtγ(ω)t\mapsto\Sigma^\gamma_t(\omega) is right-continuous at every t∈[0,T)t\in[0,T) by the martingale bound for the empirical state measure, hence so is t↦Σt(ω)t\mapsto\Sigma_t(\omega), and t↦Stt\mapsto S_t is continuous; hence t↦Σt(ω)−Stt\mapsto\Sigma_t(\omega)-S_t is right-continuous there, and so is t↦∣Σt(ω)−St∣t\mapsto|\Sigma_t(\omega)-S_t|. Applying the supremum lemma for bounded right-continuous processes to the family of random variables Zt=∣Σt−St∣Z_t=|\Sigma_t-S_t| with K=2K=2 and the event Ω∗\Omega_*, the map Ψ\Psi is a random variable with 0≤Ψ≤20\le\Psi\le2, and Ψ(ω)=sup⁡t∈[0,T]∣Σt(ω)−St∣\Psi(\omega)=\sup_{t\in[0,T]}|\Sigma_t(\omega)-S_t| for every ω∈Ω∗\omega\in\Omega_*.

Step 2: a pathwise integral inequality. Fix ω∈Ω∗\omega\in\Omega_* and write Ψt=sup⁡s∈[0,t]∣Σs(ω)−Ss∣\Psi_t=\sup_{s\in[0,t]}|\Sigma_s(\omega)-S_s| for t∈[0,T]t\in[0,T], so that ΨT=Ψ(ω)\Psi_T=\Psi(\omega) by Step 1. The function t↦Ψtt\mapsto\Psi_t is nondecreasing with values in [0,2][0,2], hence bounded and measurable by measurability of monotone functions.

By the open-loop policy lemma, αs(ω)=As\alpha_s(\omega)=A_s for every s∈[0,T]s\in[0,T], since ω∈Ω0\omega\in\Omega_0. Subtracting the dynamics of the generalized mean-field trajectory pair from the martingale decomposition of Σ\Sigma componentwise,

Σtγ−Stγ=(Σ0γ−S0γ)+∫[0,t](bγ(Σs,As)−bγ(Ss,As))ds+Mtγ.\Sigma^\gamma_t-S^\gamma_t=\big(\Sigma^\gamma_0-S^\gamma_0\big)+\int_{[0,t]}\Big(b^\gamma(\Sigma_s,A_s)-b^\gamma(S_s,A_s)\Big)ds+M^\gamma_t .

The integrand, as a map into Rl\mathbb{R}^l, has measurable components and is bounded: the first term is measurable and bounded by 2(l−1)B2(l-1)B on Ω∗\Omega_* by clause (a) of the martingale decomposition, and the second by the definition of a generalized mean-field trajectory pair, whose drift is that of the projected extension, which coincides with bb on Δl\Delta^l by claim 6 of the lemma on affine-controlled data. Hence the norm bound for vector-valued integrals and the triangle inequality give

∣Σt−St∣≤∣Σ0−S0∣+∫[0,t]∣b(Σs,As)−b(Ss,As)∣ ds+∣Mt∣.|\Sigma_t-S_t|\le|\Sigma_0-S_0|+\int_{[0,t]}\big|b(\Sigma_s,A_s)-b(S_s,A_s)\big|\,ds+|M_t| .

By the state-Lipschitz bound of the lemma on affine-controlled data, valid since Σs\Sigma_s and SsS_s lie in Δl\Delta^l,

∣b(Σs,As)−b(Ss,As)∣≤Λb ∣Σs−Ss∣≤Λb Ψs,\big|b(\Sigma_s,A_s)-b(S_s,A_s)\big|\le\Lambda_b\,|\Sigma_s-S_s|\le\Lambda_b\,\Psi_s ,

so by monotonicity of the integral and ∣Mt∣≤M‾(ω)|M_t|\le\overline{M}(\omega),

∣Σt−St∣≤X+Λb∫[0,t]Ψs ds,X=∣Σ0(ω)−S0∣+M‾(ω).|\Sigma_t-S_t|\le X+\Lambda_b\int_{[0,t]}\Psi_s\,ds,\qquad X=|\Sigma_0(\omega)-S_0|+\overline{M}(\omega) .

The right-hand side is nondecreasing in tt, so taking the supremum over the times in [0,t][0,t] on the left gives

Ψt≤X+Λb∫[0,t]Ψs ds(t∈[0,T]).\Psi_t\le X+\Lambda_b\int_{[0,t]}\Psi_s\,ds\qquad(t\in[0,T]).

Step 3: Gronwall and conclusion. The function t↦Ψtt\mapsto\Psi_t is bounded and measurable, so Gronwall's lemma for bounded measurable functions, applied with a=Xa=X and c=Λbc=\Lambda_b, gives Ψt≤X eΛbt\Psi_t\le X\,e^{\Lambda_bt} for every t∈[0,T]t\in[0,T]. At t=Tt=T,

Ψ(ω)=ΨT≤eΛbT(∣Σ0(ω)−S0∣+M‾(ω)).\Psi(\omega)=\Psi_T\le e^{\Lambda_bT}\Big(|\Sigma_0(\omega)-S_0|+\overline{M}(\omega)\Big).

Squaring and using (p+q)2≤2p2+2q2(p+q)^2\le2p^2+2q^2, valid because (p−q)2≥0(p-q)^2\ge0,

Ψ(ω)2≤2e2ΛbT(∣Σ0(ω)−S0∣2+M‾(ω)2)for every ω∈Ω∗.\Psi(\omega)^2\le2e^{2\Lambda_bT}\Big(|\Sigma_0(\omega)-S_0|^2+\overline{M}(\omega)^2\Big)\qquad\text{for every }\omega\in\Omega_* .

Both sides of this inequality are bounded random variables: Ψ≤2\Psi\le2 and ∣Σ0−S0∣≤2|\Sigma_0-S_0|\le2 as in Step 1, while, with KM=2+2(l−1)BTK_M=2+2(l-1)BT as in the martingale bound lemma, M‾≤l KM\overline{M}\le\sqrt{l}\,K_M because each Mγ‾≤KM\overline{M^\gamma}\le K_M by the supremum lemma. Since P(Ω∗)=1P(\Omega_*)=1, taking expectations by the passage of almost sure inequalities between bounded random variables to expectations, using linearity of the integral and then the bound E[M‾ 2]≤8l(l−1)BT/N\mathbb{E}[\overline{M}^{\,2}]\le8l(l-1)BT/N of the martingale bound lemma gives

E[Ψ2]≤2e2ΛbT(E[∣Σ0−S0∣2]+8l(l−1)BTN).■\mathbb{E}\big[\Psi^2\big]\le2e^{2\Lambda_bT}\Big(\mathbb{E}\big[|\Sigma_0-S_0|^2\big]+\frac{8l(l-1)BT}{N}\Big). \qquad\blacksquare

Citations

Loading…

Dependencies

Uses0

Loading…

Comments

Log in to comment.

Loading…