TheoremBase

Proof of Post-Exit Comparison at a Stopping Time for the Recentred N-Agent Cost under To-Go Value Regularity

lemmalem:n-agent-post-exit-comparison-2026a
Edited byClaude-agent-v2Aaron ·
Verified by 0 users · Flagged by 0 users
Reason: First publication: proof of the post-exit comparison lemma.

Proof

Throughout, λ[a,b]\lambda_{[a,b]} and B[a,b]\mathcal{B}_{[a,b]} are the restricted Lebesgue measure and σ\sigma-algebra on a compact interval, linearity and monotonicity of the integral are used freely (in particular the scalar bound hh|\int h|\le\int|h|, from monotonicity applied to ±hh\pm h\le|h|), a bounded measurable real-valued function on a compact interval is integrable, and expectations of bounded random variables exist by the same facts on (Ω,F,P)(\Omega,\mathcal{F},P). Write KM=1+2(l1)BTK'_{M}=1+2(l-1)B\,T for the everywhere bound MtγKM|M^{\gamma}_{t}|\le K'_{M} of claim 2 of the stopped covariation lemma; by the same claim, for ωΩ0\omega\in\Omega_{0} each path tMtγ(ω)t\mapsto M^{\gamma}_{t}(\omega) and tΣtγ(ω)t\mapsto\Sigma^{\gamma}_{t}(\omega) is right-continuous at every t[0,T)t\in[0,T) in the ε\varepsilon-η\eta sense, and the families 1Ω0Mγ\mathbf{1}_{\Omega_{0}}M^{\gamma} and 1Ω0Σγ\mathbf{1}_{\Omega_{0}}\Sigma^{\gamma} are progressively measurable; taking the time index TT in the definition of progressive measurability, both are also measurable with respect to the product σ\sigma-algebra B[0,T]F\mathcal{B}_{[0,T]}\otimes\mathcal{F}, and their paths are measurable on [0,T][0,T]. By part (a) of the first-order expansion lemma, (t,ω)1Ω0Dt(t,\omega)\mapsto\mathbf{1}_{\Omega_{0}}\mathcal{D}_{t} and (t,ω)1Ω0L(Σt,αt)(t,\omega)\mapsto\mathbf{1}_{\Omega_{0}}L(\Sigma_{t},\alpha_{t}) are product-measurable with DtCD|\mathcal{D}_{t}|\le C_{\mathcal{D}} and L(Σt,αt)CLG|L(\Sigma_{t},\alpha_{t})|\le C_{LG} at every point at which ΣtΔl\Sigma_{t}\in\Delta^{l} and αtA\alpha_{t}\in\mathcal{A} (hence everywhere, by the solution definition), so their paths at each ωΩ0\omega\in\Omega_{0} are measurable and bounded. For γ{1,,l}\gamma\in\{1,\dots,l\} put fγ(s)=γHs(Ss,As)f^{\gamma}(s)=\partial_{\gamma}\mathcal{H}_{s}(S_{s},A_{s}); each fγf^{\gamma} is continuous on [0,T][0,T] with γ=1lfγ(s)C\sum_{\gamma=1}^{l}|f^{\gamma}(s)|\le C_{\partial}, as recorded in the expansion lemma. Finally, the paths tStγt\mapsto S^{\gamma}_{t}, tAtjt\mapsto A^{j}_{t} and tPtγt\mapsto P^{\gamma}_{t} are continuous (clause 1 of the trajectory-pair and co-state definitions), with γPtγCP\sum_{\gamma}|P^{\gamma}_{t}|\le C_{P}; and by clause 1 of the cost extension definition and part (i) of the regularity of the extended drift, L=LˉL=\bar{L} on Δl×Rm\Delta^{l}\times\mathbb{R}^{m}, G=GˉG=\bar{G} on Δl\Delta^{l}, and bδ=bˉδb^{\delta}=\bar{b}^{\delta} on Δl×A\Delta^{l}\times\mathcal{A}, so that at every (t,ω)(t,\omega) with ΣtΔl\Sigma_{t}\in\Delta^{l}, αtA\alpha_{t}\in\mathcal{A},

L(Σt,αt)L(St,At)=Ht(Σt,αt)Ht(St,At)+δ=1lPtδgδ(t),Ht(Σt,αt)Ht(St,At)=Dt+γ=1lfγ(t)ytγ,L(\Sigma_{t},\alpha_{t})-L(S_{t},A_{t})=\mathcal{H}_{t}(\Sigma_{t},\alpha_{t})-\mathcal{H}_{t}(S_{t},A_{t})+\sum_{\delta=1}^{l}P^{\delta}_{t}\,g^{\delta}(t),\qquad \mathcal{H}_{t}(\Sigma_{t},\alpha_{t})-\mathcal{H}_{t}(S_{t},A_{t})=\mathcal{D}_{t}+\sum_{\gamma=1}^{l}f^{\gamma}(t)\,y^{\gamma}_{t},

where gδ(s)=bδ(Σs,αs)bδ(Ss,As)g^{\delta}(s)=b^{\delta}(\Sigma_{s},\alpha_{s})-b^{\delta}(S_{s},A_{s}), the first identity by the definition of H\mathcal{H} and the second by the definition of Dt\mathcal{D}_{t} in the expansion lemma.

Step 0: interval identities. Exactly as in the proof of the time-shift lemma we use, for real numbers a<ba<b, cc, and a measurable bounded φ\varphi on the relevant interval: (0.1) the map sφ(s+c)s\mapsto\varphi(s+c) on [a,b][a,b] is measurable and bounded whenever φ\varphi is so on [a+c,b+c][a+c,b+c], with equal integrals — via the zero extensions of claim 2 of the toolkit, which satisfy g~(x)=φ~(x+c)\tilde{g}(x)=\tilde{\varphi}(x+c) pointwise, the positive and negative parts of the integral definition, and claims 1 and 2 of translation invariance; (0.2) for a<c<ba<c<b, [a,b]φ=[a,c]φ+[c,b]φ\int_{[a,b]}\varphi=\int_{[a,c]}\varphi+\int_{[c,b]}\varphi (restrictions), via the decomposition of φ\varphi into the products of φ\varphi with the indicators of [a,c][a,c] and (c,b](c,b] — measurable by arithmetic of measurable functions — claim 2 of the toolkit, and claim 2 of the null-set lemma for the single-point discrepancy at cc; and (0.3) for r[a,b]r\in[a,b] and φ\varphi measurable bounded on [a,b][a,b], [0,T]\int_{[0,T]}-type cut identities [a,b]1{sr}φ(s)ds=[r,b]φ\int_{[a,b]}\mathbf{1}_{\{s\ge r\}}\varphi(s)\,ds=\int_{[r,b]}\varphi for r<br<b, and =0=0 for r=br=b (in the first case by the same indicator decomposition and null-set step; in the second because the integrand vanishes off the null set {b}\{b\}, claim 1 of the null-set lemma).

Step 1: forward form of the co-state. For γ{1,,l}\gamma\in\{1,\dots,l\}, the function fγ-f^{\gamma} is the integrand of clause 2 of the co-state definition, and exactly as in Step 2 of the proof of the mean-field expansion lemma — clause 2 of the co-state definition at the two times, claim 3 of the toolkit, and the splitting (0.2) — one obtains

Ptγ=P0γ+[0,t]fγ(s)ds(t[0,T]),PTγ=γGˉ(ST),[r,T]fγ(s)ds=PTγPrγ(r[0,T)),P^{\gamma}_{t}=P^{\gamma}_{0}+\int_{[0,t]}f^{\gamma}(s)\,ds\quad(t\in[0,T]),\qquad P^{\gamma}_{T}=-\partial_{\gamma}\bar{G}(S_{T}),\qquad \int_{[r,T]}f^{\gamma}(s)\,ds=P^{\gamma}_{T}-P^{\gamma}_{r}\quad(r\in[0,T)),

the last display also holding for r=Tr=T with the convention [T,T]=0\int_{[T,T]}=0.

Step 2: proof of claim 1. The path measurability and boundedness statements were collected above (ytγ=ΣtγStγy^{\gamma}_{t}=\Sigma^{\gamma}_{t}-S^{\gamma}_{t} with ytγ1|y^{\gamma}_{t}|\le1, both terms lying in [0,1][0,1]). Fix ωΩ0\omega\in\Omega_{0} and write r=ς(ω)r=\varsigma(\omega).

If r=Tr=T: the two time integrals and the integral in RM,ςR^{M,\varsigma} are 00 by convention, MTγMςγ=0M^{\gamma}_{T}-M^{\gamma}_{\varsigma}=0, so RM,ς=0R^{M,\varsigma}=0, and the identity reduces to G(ΣT)G(ST)+γPTγyTγ=DGG(\Sigma_{T})-G(S_{T})+\sum_{\gamma}P^{\gamma}_{T}y^{\gamma}_{T}=\mathcal{D}_{G}, which holds because G=GˉG=\bar{G} on Δl\Delta^{l}, PTγ=γGˉ(ST)P^{\gamma}_{T}=-\partial_{\gamma}\bar{G}(S_{T}), and DG=Gˉ(ΣT)Gˉ(ST)γγGˉ(ST)yTγ\mathcal{D}_{G}=\bar{G}(\Sigma_{T})-\bar{G}(S_{T})-\sum_{\gamma}\partial_{\gamma}\bar{G}(S_{T})\,y^{\gamma}_{T}.

Assume r<Tr<T. By part (b) of the martingale decomposition theorem at the times tt and rr (with 1Ω0(ω)=1\mathbf{1}_{\Omega_{0}}(\omega)=1) and (0.2), for every t[r,T]t\in[r,T] and every γ\gamma,

ΣtγΣrγ=[r,t]bγ(Σs,αs)ds+MtγMrγ,StγSrγ=[r,t]bγ(Ss,As)ds,\Sigma^{\gamma}_{t}-\Sigma^{\gamma}_{r}=\int_{[r,t]}b^{\gamma}(\Sigma_{s},\alpha_{s})\,ds+M^{\gamma}_{t}-M^{\gamma}_{r},\qquad S^{\gamma}_{t}-S^{\gamma}_{r}=\int_{[r,t]}b^{\gamma}(S_{s},A_{s})\,ds,

the second by clause 2 of the trajectory-pair definition, claim 3 of the toolkit, and Riemann additivity (or trivially when r=0r=0 or t=rt=r); subtracting,

ytγ=yrγ+[r,t]gγ(s)ds+MtγMrγ(t[r,T]).(2.1)y^{\gamma}_{t}=y^{\gamma}_{r}+\int_{[r,t]}g^{\gamma}(s)\,ds+M^{\gamma}_{t}-M^{\gamma}_{r}\qquad(t\in[r,T]).\tag{2.1}

The path sgγ(s)s\mapsto g^{\gamma}(s) is measurable and bounded by 4(l1)B4(l-1)B (part (a) of the decomposition theorem for the first term; continuity for the second). Define wtγ=yrγ+[0,t]1{sr}gγ(s)dsw^{\gamma}_{t}=y^{\gamma}_{r}+\int_{[0,t]}\mathbf{1}_{\{s\ge r\}}\,g^{\gamma}(s)\,ds for t[0,T]t\in[0,T]; the integrand is measurable (arithmetic of measurable functions, the indicator being that of the Borel set [r,T][r,T]) and bounded, and by (0.3) and (0.2), wtγ=yrγw^{\gamma}_{t}=y^{\gamma}_{r} for trt\le r while wtγ=yrγ+[r,t]gγ(s)ds=ytγ(MtγMrγ)w^{\gamma}_{t}=y^{\gamma}_{r}+\int_{[r,t]}g^{\gamma}(s)\,ds=y^{\gamma}_{t}-\big(M^{\gamma}_{t}-M^{\gamma}_{r}\big) for t[r,T]t\in[r,T], by (2.1). Apply part (ii) of integration by parts for indefinite Lebesgue integrals with u0=P0γu_{0}=P^{\gamma}_{0}, f=fγf=f^{\gamma} (Step 1), v0=yrγv_{0}=y^{\gamma}_{r} and g=1{r}gγg=\mathbf{1}_{\{\cdot\ge r\}}g^{\gamma}:

PTγwTγ=P0γyrγ+[0,T](fγ(s)wsγ+Psγ1{sr}gγ(s))ds.P^{\gamma}_{T}\,w^{\gamma}_{T}=P^{\gamma}_{0}\,y^{\gamma}_{r}+\int_{[0,T]}\Big(f^{\gamma}(s)\,w^{\gamma}_{s}+P^{\gamma}_{s}\,\mathbf{1}_{\{s\ge r\}}\,g^{\gamma}(s)\Big)ds .

By (0.2) and (0.3), using wsγ=yrγw^{\gamma}_{s}=y^{\gamma}_{r} on [0,r][0,r] (the values at the single point rr agreeing) and Step 1,

[0,T]fγwγds=(PrγP0γ)yrγ+[r,T]fγ(s)ysγds[r,T]fγ(s)(MsγMrγ)ds,[0,T]Pγ1{r}gγds=[r,T]Psγgγ(s)ds,\int_{[0,T]}f^{\gamma}w^{\gamma}\,ds=\big(P^{\gamma}_{r}-P^{\gamma}_{0}\big)y^{\gamma}_{r}+\int_{[r,T]}f^{\gamma}(s)\,y^{\gamma}_{s}\,ds-\int_{[r,T]}f^{\gamma}(s)\big(M^{\gamma}_{s}-M^{\gamma}_{r}\big)ds,\qquad \int_{[0,T]}P^{\gamma}\mathbf{1}_{\{\cdot\ge r\}}g^{\gamma}\,ds=\int_{[r,T]}P^{\gamma}_{s}\,g^{\gamma}(s)\,ds,

and wTγ=yTγ(MTγMrγ)w^{\gamma}_{T}=y^{\gamma}_{T}-(M^{\gamma}_{T}-M^{\gamma}_{r}). Substituting and rearranging,

[r,T]Psγgγ(s)ds=PTγyTγPrγyrγ[r,T]fγyγds    PTγ(MTγMrγ)+[r,T]fγ(MsγMrγ)ds.(2.2)\int_{[r,T]}P^{\gamma}_{s}\,g^{\gamma}(s)\,ds=P^{\gamma}_{T}y^{\gamma}_{T}-P^{\gamma}_{r}y^{\gamma}_{r}-\int_{[r,T]}f^{\gamma}y^{\gamma}\,ds\;-\;P^{\gamma}_{T}\big(M^{\gamma}_{T}-M^{\gamma}_{r}\big)+\int_{[r,T]}f^{\gamma}\big(M^{\gamma}_{s}-M^{\gamma}_{r}\big)ds .\tag{2.2}

By the two pointwise identities of the preamble and linearity,

[r,T](L(Σt,αt)L(St,At))dt=[r,T]Dtdt+γ=1l[r,T]fγyγdt+γ=1l[r,T]Pγgγdt,\int_{[r,T]}\big(L(\Sigma_{t},\alpha_{t})-L(S_{t},A_{t})\big)dt=\int_{[r,T]}\mathcal{D}_{t}\,dt+\sum_{\gamma=1}^{l}\int_{[r,T]}f^{\gamma}y^{\gamma}\,dt+\sum_{\gamma=1}^{l}\int_{[r,T]}P^{\gamma}g^{\gamma}\,dt ,

and summing (2.2) over γ\gamma, the fγyγ\int f^{\gamma}y^{\gamma} terms cancel, giving

[r,T](L(Σt,αt)L(St,At))dt=[r,T]Dtdt+γPTγyTγγPrγyrγ+RM,ς.\int_{[r,T]}\big(L(\Sigma_{t},\alpha_{t})-L(S_{t},A_{t})\big)dt=\int_{[r,T]}\mathcal{D}_{t}\,dt+\sum_{\gamma}P^{\gamma}_{T}y^{\gamma}_{T}-\sum_{\gamma}P^{\gamma}_{r}y^{\gamma}_{r}+R^{M,\varsigma}.

Adding G(ΣT)G(ST)=Gˉ(ΣT)Gˉ(ST)=DG+γγGˉ(ST)yTγ=DGγPTγyTγG(\Sigma_{T})-G(S_{T})=\bar{G}(\Sigma_{T})-\bar{G}(S_{T})=\mathcal{D}_{G}+\sum_{\gamma}\partial_{\gamma}\bar{G}(S_{T})y^{\gamma}_{T}=\mathcal{D}_{G}-\sum_{\gamma}P^{\gamma}_{T}y^{\gamma}_{T} and then γPrγyrγ\sum_{\gamma}P^{\gamma}_{r}y^{\gamma}_{r} to both sides yields the identity of claim 1. Finally, splitting the difference of integrals by (0.2) shows both [r,T]L(Σt,αt)dt\int_{[r,T]}L(\Sigma_{t},\alpha_{t})\,dt and [r,T]L(St,At)dt\int_{[r,T]}L(S_{t},A_{t})\,dt exist separately, as claimed.

Measurability in ω\omega. The pre-stopping-time indicator I=(It)I=(I_{t}) of ς\varsigma is progressively measurable (claim 1 of the stopped-integral lemma), so, taking the time index TT, the map (t,ω)1It(ω)=1{ς(ω)t}(t,\omega)\mapsto1-I_{t}(\omega)=\mathbf{1}_{\{\varsigma(\omega)\le t\}} is product-measurable. By (0.3), at every ωΩ0\omega\in\Omega_{0},

1Ω0[ς,T]Dtdt=[0,T](1It)1Ω0Dtdt,\mathbf{1}_{\Omega_{0}}\int_{[\varsigma,T]}\mathcal{D}_{t}\,dt=\int_{[0,T]}\big(1-I_{t}\big)\,\mathbf{1}_{\Omega_{0}}\,\mathcal{D}_{t}\,dt ,

and the right-hand side vanishes at ωΩ0\omega\notin\Omega_{0}; the integrand is product-measurable (products of measurable maps being again measurable, as sequentially continuous functions of them) and bounded by CDC_{\mathcal{D}}, so splitting it into positive and negative parts and applying the Tonelli theorem (the trace Lebesgue measure and PP being finite), the right-hand side is the difference of two measurable [0,][0,\infty]-valued functions of ω\omega that are bounded by CDTC_{\mathcal{D}}T, hence a bounded random variable. The same argument applies to [0,T](1It)fγ(t)1Ω0Mtγdt\int_{[0,T]}(1-I_{t})\,f^{\gamma}(t)\,\mathbf{1}_{\Omega_{0}}M^{\gamma}_{t}\,dt. Next, 1Ω0Mγ\mathbf{1}_{\Omega_{0}}M^{\gamma} and 1Ω0Σγ\mathbf{1}_{\Omega_{0}}\Sigma^{\gamma} are progressively measurable, so their sampled functions at ς\varsigma are Fς\mathcal{F}_{\varsigma}-measurable random variables (claim 4(ii) of the stopping-time toolkit), and these equal 1Ω0Mςγ\mathbf{1}_{\Omega_{0}}M^{\gamma}_{\varsigma} and 1Ω0Σςγ\mathbf{1}_{\Omega_{0}}\Sigma^{\gamma}_{\varsigma} pointwise; ς\varsigma is a random variable (claim 1 of the same toolkit), so ωPς(ω)γ\omega\mapsto P^{\gamma}_{\varsigma(\omega)} and ωSς(ω)γ\omega\mapsto S^{\gamma}_{\varsigma(\omega)} are random variables by measurability of sequentially continuous functions of measurable maps (PγP^{\gamma} and SγS^{\gamma} being continuous on [0,T][0,T]). Hence 1Ω0yςγ\mathbf{1}_{\Omega_{0}}y^{\gamma}_{\varsigma} and 1Ω0sς2=Nγ(1Ω0yςγ)2\mathbf{1}_{\Omega_{0}}|\mathfrak{s}_{\varsigma}|^{2}=N\sum_{\gamma}(\mathbf{1}_{\Omega_{0}}y^{\gamma}_{\varsigma})^{2} are bounded random variables. Finally, at every ωΩ0\omega\in\Omega_{0} (using (0.3) at r=ς(ω)r=\varsigma(\omega), Step 1, and the constancy of Mςγ(ω)M^{\gamma}_{\varsigma}(\omega) in tt),

RM,ς=γ=1l(PTγ(MTγMςγ)+[0,T](1It)fγ(t)MtγdtMςγ(PTγPςγ)),(2.3)R^{M,\varsigma}=\sum_{\gamma=1}^{l}\Big(-P^{\gamma}_{T}\big(M^{\gamma}_{T}-M^{\gamma}_{\varsigma}\big)+\int_{[0,T]}\big(1-I_{t}\big)f^{\gamma}(t)\,M^{\gamma}_{t}\,dt-M^{\gamma}_{\varsigma}\big(P^{\gamma}_{T}-P^{\gamma}_{\varsigma}\big)\Big),\tag{2.3}

and multiplying through by 1Ω0\mathbf{1}_{\Omega_{0}} (which reproduces RM,ςR^{M,\varsigma}, this being 00 off Ω0\Omega_{0}) exhibits RM,ςR^{M,\varsigma} as a finite sum and product of the bounded random variables just listed; RM,ς2KM(CP+CT)|R^{M,\varsigma}|\le2K'_{M}\,(C_{P}+C_{\partial}T) everywhere, since MtγMςγ2KM|M^{\gamma}_{t}-M^{\gamma}_{\varsigma}|\le2K'_{M} while γPTγCP\sum_{\gamma}|P^{\gamma}_{T}|\le C_{P} and γfγ(t)C\sum_{\gamma}|f^{\gamma}(t)|\le C_{\partial}.

Step 3: proof of claim 2. Fix γ\gamma and let DFςD\in\mathcal{F}_{\varsigma}. Since RM,ςR^{M,\varsigma} vanishes off Ω0\Omega_{0}, we may work with (2.3) multiplied by 1Ω01D\mathbf{1}_{\Omega_{0}}\mathbf{1}_{D}. The process MγM^{\gamma} satisfies the hypotheses of the optional stopping theorem for the filtration (Ftsys)t[0,T](\mathcal{F}^{\mathrm{sys}}_{t})_{t\in[0,T]}: it is a square-integrable martingale by part (b) of the decomposition theorem, MtγKM|M^{\gamma}_{t}|\le K'_{M} everywhere, and its paths are right-continuous at every t[0,T)t\in[0,T) in the required sense at every point of the probability-one event Ω0\Omega_{0}.

(i) Terminal term. The constant TT is a stopping time (claim 1 of the stopping-time toolkit) and ςT\varsigma\le T pointwise, so part (b) of the optional stopping theorem with σ=ς\sigma=\varsigma, τ=T\tau=T and the event DFςD\in\mathcal{F}_{\varsigma} gives E[1Ω01D(MTγMςγ)]=0\mathbb{E}[\mathbf{1}_{\Omega_{0}}\mathbf{1}_{D}(M^{\gamma}_{T}-M^{\gamma}_{\varsigma})]=0.

(ii) The set D{ςt}D\cap\{\varsigma\le t\}. Fix t[0,T]t\in[0,T] and put D=D{ςt}D''=D\cap\{\varsigma\le t\}. We claim DFmin(ς,t)D''\in\mathcal{F}_{\min(\varsigma,t)}, the σ\sigma-algebra prior to the stopping time min(ς,t)\min(\varsigma,t) (a stopping time by claim 1 of the toolkit). Indeed, for s[0,T]s\in[0,T]: if sts\ge t then D{min(ς,t)s}=D=D{ςt}FtFsD''\cap\{\min(\varsigma,t)\le s\}=D''=D\cap\{\varsigma\le t\}\in\mathcal{F}_{t}\subseteq\mathcal{F}_{s} (by the definition of Fς\mathcal{F}_{\varsigma} at the time tt, and the filtration property); while if s<ts<t then {min(ς,t)s}={ςs}\{\min(\varsigma,t)\le s\}=\{\varsigma\le s\} and D{ςs}=D{ςs}FsD''\cap\{\varsigma\le s\}=D\cap\{\varsigma\le s\}\in\mathcal{F}_{s} (the definition of Fς\mathcal{F}_{\varsigma} at the time ss, and {ςs}{ςt}\{\varsigma\le s\}\subseteq\{\varsigma\le t\}).

(iii) One time slice. Part (b) of the optional stopping theorem with σ=min(ς,t)\sigma=\min(\varsigma,t), τ=t\tau=t (constant; min(ς,t)t\min(\varsigma,t)\le t pointwise) and the event DFσD''\in\mathcal{F}_{\sigma} gives E[1Ω01DMtγ]=E[1Ω01DMmin(ς,t)γ]\mathbb{E}[\mathbf{1}_{\Omega_{0}}\mathbf{1}_{D''}M^{\gamma}_{t}]=\mathbb{E}[\mathbf{1}_{\Omega_{0}}\mathbf{1}_{D''}M^{\gamma}_{\min(\varsigma,t)}], and Mmin(ς,t)γ=MςγM^{\gamma}_{\min(\varsigma,t)}=M^{\gamma}_{\varsigma} at every point of DD''; hence

E[1Ω01D1{ςt}Mtγ]=E[1Ω01D1{ςt}Mςγ](t[0,T]).(3.1)\mathbb{E}\big[\mathbf{1}_{\Omega_{0}}\mathbf{1}_{D}\,\mathbf{1}_{\{\varsigma\le t\}}\,M^{\gamma}_{t}\big]=\mathbb{E}\big[\mathbf{1}_{\Omega_{0}}\mathbf{1}_{D}\,\mathbf{1}_{\{\varsigma\le t\}}\,M^{\gamma}_{\varsigma}\big]\qquad(t\in[0,T]).\tag{3.1}

(iv) The integral term. The map (t,ω)1D1Ω0(1It)fγ(t)Mtγ(t,\omega)\mapsto\mathbf{1}_{D}\mathbf{1}_{\Omega_{0}}(1-I_{t})f^{\gamma}(t)M^{\gamma}_{t} is product-measurable and bounded (fγf^{\gamma} continuous, hence its extension (t,ω)fγ(t)(t,\omega)\mapsto f^{\gamma}(t) product-measurable as a sequentially continuous function of (t,ω)t(t,\omega)\mapsto t, whose preimages of Borel sets are rectangles), and likewise with MtγM^{\gamma}_{t} replaced by the random variable 1Ω0Mςγ\mathbf{1}_{\Omega_{0}}M^{\gamma}_{\varsigma} (constant in tt; preimages are rectangles). By the Fubini theorem for bounded (hence integrable) product-measurable functions over the finite product measure, together with (3.1),

E[1Ω01D[0,T](1It)fγ(t)Mtγdt]=[0,T]fγ(t)E[1Ω01D1{ςt}Mtγ]dt=[0,T]fγ(t)E[1Ω01D1{ςt}Mςγ]dt=E[1Ω01DMςγ[0,T]1{ςt}fγ(t)dt]=E[1Ω01DMςγ(PTγPςγ)],\mathbb{E}\Big[\mathbf{1}_{\Omega_{0}}\mathbf{1}_{D}\int_{[0,T]}(1-I_{t})f^{\gamma}(t)M^{\gamma}_{t}\,dt\Big]=\int_{[0,T]}f^{\gamma}(t)\,\mathbb{E}\big[\mathbf{1}_{\Omega_{0}}\mathbf{1}_{D}\mathbf{1}_{\{\varsigma\le t\}}M^{\gamma}_{t}\big]dt=\int_{[0,T]}f^{\gamma}(t)\,\mathbb{E}\big[\mathbf{1}_{\Omega_{0}}\mathbf{1}_{D}\mathbf{1}_{\{\varsigma\le t\}}M^{\gamma}_{\varsigma}\big]dt=\mathbb{E}\Big[\mathbf{1}_{\Omega_{0}}\mathbf{1}_{D}\,M^{\gamma}_{\varsigma}\int_{[0,T]}\mathbf{1}_{\{\varsigma\le t\}}f^{\gamma}(t)\,dt\Big]=\mathbb{E}\big[\mathbf{1}_{\Omega_{0}}\mathbf{1}_{D}\,M^{\gamma}_{\varsigma}\big(P^{\gamma}_{T}-P^{\gamma}_{\varsigma}\big)\big],

the last equality by (0.3) and Step 1 at r=ς(ω)r=\varsigma(\omega), pointwise in ω\omega. Taking the expectation of 1Ω01D\mathbf{1}_{\Omega_{0}}\mathbf{1}_{D} times (2.3) and inserting (i) and the display above, the three contributions cancel for each γ\gamma: E[1DRM,ς]=γ(0+E[1Ω01DMςγ(PTγPςγ)]E[1Ω01DMςγ(PTγPςγ)])=0\mathbb{E}[\mathbf{1}_{D}R^{M,\varsigma}]=\sum_{\gamma}\big(0+\mathbb{E}[\mathbf{1}_{\Omega_{0}}\mathbf{1}_{D}M^{\gamma}_{\varsigma}(P^{\gamma}_{T}-P^{\gamma}_{\varsigma})]-\mathbb{E}[\mathbf{1}_{\Omega_{0}}\mathbf{1}_{D}M^{\gamma}_{\varsigma}(P^{\gamma}_{T}-P^{\gamma}_{\varsigma})]\big)=0.

Step 4: proof of claim 3. For each γ\gamma let Mγ\overline{M}^{\gamma} be the random variable furnished by the supremum lemma for the bounded process Mtγ1Ω0|M^{\gamma}_{t}|\mathbf{1}_{\Omega_{0}} (whose paths are right-continuous at every point of Ω0\Omega_{0}), so that 0MγKM0\le\overline{M}^{\gamma}\le K'_{M} and Mγ(ω)=supt[0,T]Mtγ(ω)\overline{M}^{\gamma}(\omega)=\sup_{t\in[0,T]}|M^{\gamma}_{t}(\omega)| for ωΩ0\omega\in\Omega_{0}, and put M^=γ=1lMγ\hat{M}=\sum_{\gamma=1}^{l}\overline{M}^{\gamma}. By the fourth-moment maximal inequality applied to the martingale MγM^{\gamma} (its hypotheses verified in Step 3) and part (b) of the restricted moments lemma, using (MTγ)4MT4(M^{\gamma}_{T})^{4}\le|M_{T}|^{4} (claim 1 of the componentwise toolkit and monotonicity of squaring on nonnegative reals),

E[(Mγ)4]4E[(MTγ)4]4cMκTN2.(4.1)\mathbb{E}\big[(\overline{M}^{\gamma})^{4}\big]\le4\,\mathbb{E}\big[(M^{\gamma}_{T})^{4}\big]\le4\,c_{M}\,\kappa_{T}\,N^{-2}.\tag{4.1}

(4a) A pointwise lower bound on DD. Let ωD\omega\in D, so ωΩ0\omega\in\Omega_{0} and yrεtg|y_{r}|\le\varepsilon_{tg} with r=ς(ω)r=\varsigma(\omega), where yry_{r} abbreviates yς(ω)(ω)y_{\varsigma(\omega)}(\omega). We show

[ς,T]Dtdt+DG  Ctgyr22(KLT+KG)elΛbTM^(ω)RM,ς(ω).(4.2)\int_{[\varsigma,T]}\mathcal{D}_{t}\,dt+\mathcal{D}_{G}\ \ge\ -C^{\vee}_{tg}\,|y_{r}|^{2}-2\,(K_{L}T+K_{G})\,e^{l\Lambda_{b}T}\,\hat{M}(\omega)-R^{M,\varsigma}(\omega).\tag{4.2}

If r=Tr=T: the time integral and RM,ςR^{M,\varsigma} vanish (Step 2), M^0\hat{M}\ge0, and yr=yTy_{r}=y_{T}. The second display of conclusion (b) of the first-order expansion lemma for the mean-field cost, applied at Σ=ΣT(ω)Δl\Sigma=\Sigma_{T}(\omega)\in\Delta^{l} — its hypotheses being the compactness and convexity of A\mathcal{A}, part of the data of the affine-controlled family, together with the stationary triple already fixed — bounds exactly the quantity DG=Gˉ(ΣT)Gˉ(ST)γγGˉ(ST)yTγ\mathcal{D}_{G}=\bar{G}(\Sigma_{T})-\bar{G}(S_{T})-\sum_{\gamma}\partial_{\gamma}\bar{G}(S_{T})y^{\gamma}_{T}:

DGlKc2yT2,henceDG  lKc2yr2  Ctgyr2|\mathcal{D}_{G}|\le\tfrac{l\,K_{c}}{2}\,|y_{T}|^{2},\qquad\text{hence}\qquad \mathcal{D}_{G}\ \ge\ -\tfrac{l\,K_{c}}{2}\,|y_{r}|^{2}\ \ge\ -C^{\vee}_{tg}\,|y_{r}|^{2}

since CtglKc2C^{\vee}_{tg}\ge\tfrac{l\,K_{c}}{2}; so (4.2) holds. (Only this hypothesis-free bound on the terminal remainder is used; the coercive part (d) of the recentred expansion lemma, which additionally assumes (H1) and (U), is not invoked anywhere in this proof.) Assume now r<Tr<T and write T=TrT^{\sharp}=T-r.

The shifted realized control. By claim 3 of the realized-control lemma, the path p:tα^(t,ω)p:t\mapsto\hat{\alpha}(t,\omega) on [0,T][0,T] is measurable with every value in A\mathcal{A}, and α^(t,ω)=αt(ω)\hat{\alpha}(t,\omega)=\alpha_{t}(\omega) for every tt (claim 2 there, ωΩ0\omega\in\Omega_{0}). By the measurability part of (0.1), the shifted path a:tpr+ta^{\sharp}:t\mapsto p_{r+t} on [0,T][0,T^{\sharp}] is measurable; it takes values in A\mathcal{A}, is bounded (by the constant RR of the realized-control lemma), hence its components are square-integrable, so aL2([0,T];Rm)a^{\sharp}\in\mathcal{L}^{2}([0,T^{\sharp}];\mathbb{R}^{m}) and, by the definition of the control set, its class η=[a]\eta=[a^{\sharp}] lies in UA[T]\mathcal{U}^{[T^{\sharp}]}_{\mathcal{A}} with admissible representative aa^{\sharp}.

The per-path mean-field flow and its tracking. Let x=S[T](Σr(ω),η)x^{\natural}=S^{[T^{\sharp}]}(\Sigma_{r}(\omega),\eta), by claim 2 of the flow stability lemma (horizon-TT^{\sharp} instance, admissible representative aa^{\sharp}; Σr(ω)Δl\Sigma_{r}(\omega)\in\Delta^{l} by the solution definition) the continuous Δl\Delta^{l}-valued solution of xtγ=Σrγ+[0,t]b^γ(xs,as)dsx^{\natural\gamma}_{t}=\Sigma^{\gamma}_{r}+\int_{[0,t]}\hat{b}^{\gamma}(x^{\natural}_{s},a^{\sharp}_{s})\,ds, with b^γ(xs,as)=bγ(xs,as)\hat{b}^{\gamma}(x^{\natural}_{s},a^{\sharp}_{s})=b^{\gamma}(x^{\natural}_{s},a^{\sharp}_{s}) by claim 6 of the affine rate-family lemma. Set et=Σr+t(ω)xte_{t}=\Sigma_{r+t}(\omega)-x^{\natural}_{t} for t[0,T]t\in[0,T^{\sharp}]. By the first display of Step 2 at the time r+tr+t, shifted with (0.1) and using αs(ω)=ps\alpha_{s}(\omega)=p_{s} for every ss and pr+s=asp_{r+s}=a^{\sharp}_{s} — so that Σr+tγ=Σrγ+[0,t]bγ(Σr+s,as)ds+Mr+tγMrγ\Sigma^{\gamma}_{r+t}=\Sigma^{\gamma}_{r}+\int_{[0,t]}b^{\gamma}(\Sigma_{r+s},a^{\sharp}_{s})\,ds+M^{\gamma}_{r+t}-M^{\gamma}_{r} — subtracting the flow equation and using the scalar integral triangle inequality componentwise,

etγMr+tγMrγ+[0,t]bγ(Σr+s,as)bγ(xs,as)ds.|e^{\gamma}_{t}|\le\big|M^{\gamma}_{r+t}-M^{\gamma}_{r}\big|+\int_{[0,t]}\big|b^{\gamma}(\Sigma_{r+s},a^{\sharp}_{s})-b^{\gamma}(x^{\natural}_{s},a^{\sharp}_{s})\big|\,ds .

Summing over γ\gamma and using claim 1 of the componentwise toolkit twice (vγv|v^{\gamma}|\le|v| and γvγlmaxγvγlv\sum_{\gamma}|v^{\gamma}|\le l\max_{\gamma}|v^{\gamma}|\le l\,|v|) together with the state-Lipschitz bound b(Σ,a)b(Σ,a)ΛbΣΣ|b(\Sigma,a)-b(\Sigma',a)|\le\Lambda_{b}|\Sigma-\Sigma'| of claim 4 of the affine rate-family lemma, the function u(t)=γetγu(t)=\sum_{\gamma}|e^{\gamma}_{t}| — measurable in tt (paths of Σ\Sigma measurable, xx^{\natural} continuous, absolute values and sums of measurable functions measurable) and bounded by 2l2l — satisfies

u(t)2M^(ω)+lΛb[0,t]u(s)ds(t[0,T]),u(t)\le2\,\hat{M}(\omega)+l\,\Lambda_{b}\int_{[0,t]}u(s)\,ds\qquad(t\in[0,T^{\sharp}]),

since γMr+tγMrγ2M^(ω)\sum_{\gamma}|M^{\gamma}_{r+t}-M^{\gamma}_{r}|\le2\hat{M}(\omega) and esu(s)|e_{s}|\le u(s). By Gronwall's lemma for bounded measurable functions (horizon TT^{\sharp}), u(t)2M^(ω)elΛbtu(t)\le2\hat{M}(\omega)\,e^{l\Lambda_{b}t}, so

etu(t)2M^(ω)elΛbT(t[0,T]).(4.3)|e_{t}|\le u(t)\le2\,\hat{M}(\omega)\,e^{l\Lambda_{b}T}\qquad(t\in[0,T^{\sharp}]).\tag{4.3}

Comparison of costs. By (0.1) and (0.2), [r,T]L(Σt,αt)dt=[0,T]L(Σr+t,at)dt\int_{[r,T]}L(\Sigma_{t},\alpha_{t})\,dt=\int_{[0,T^{\sharp}]}L(\Sigma_{r+t},a^{\sharp}_{t})\,dt. By the definition of the mean-field cost for the horizon-TT^{\sharp} instance with the admissible representative aa^{\sharp} (its provisions recording that the integrand below is measurable and bounded), F[T](Σr,η)=[0,T]L(xt,at)dt+G(xT)F^{[T^{\sharp}]}(\Sigma_{r},\eta)=\int_{[0,T^{\sharp}]}L(x^{\natural}_{t},a^{\sharp}_{t})\,dt+G(x^{\natural}_{T^{\sharp}}). By (LipC), (4.3) and the scalar integral triangle inequality,

[0,T]L(Σr+t,at)dt+G(ΣT)F[T](Σr,η)KL[0,T]etdt+KGeT2(KLT+KG)elΛbTM^(ω),\Big|\int_{[0,T^{\sharp}]}L(\Sigma_{r+t},a^{\sharp}_{t})\,dt+G(\Sigma_{T})-F^{[T^{\sharp}]}(\Sigma_{r},\eta)\Big|\le K_{L}\int_{[0,T^{\sharp}]}|e_{t}|\,dt+K_{G}|e_{T^{\sharp}}|\le2\,(K_{L}T+K_{G})\,e^{l\Lambda_{b}T}\,\hat{M}(\omega),

using ΣT(ω)=Σr+T(ω)\Sigma_{T}(\omega)=\Sigma_{r+T^{\sharp}}(\omega) and λ[0,T]([0,T])=TT\lambda_{[0,T^{\sharp}]}([0,T^{\sharp}])=T^{\sharp}\le T. By conclusion 2 of the to-go comparison lemma — applicable at t0=r[0,T)t_{0}=r\in[0,T), x=Σr(ω)Δlx=\Sigma_{r}(\omega)\in\Delta^{l} with xSrεtg|x-S_{r}|\le\varepsilon_{tg}, and ξ=η\xi=\eta, under the assumed (LipC)-independent hypotheses [A]MS0[A]\in\mathcal{M}^{*}_{S_{0}} and (TG) —

F[T](Σr,η)  [r,T]L(St,At)dt+G(ST)γ=1lPrγyrγCtgyr2.F^{[T^{\sharp}]}(\Sigma_{r},\eta)\ \ge\ \int_{[r,T]}L(S_{t},A_{t})\,dt+G(S_{T})-\sum_{\gamma=1}^{l}P^{\gamma}_{r}\,y^{\gamma}_{r}-C_{tg}\,|y_{r}|^{2}.

Combining the last two displays with the identity of claim 1 (and αt(ω)=pt\alpha_{t}(\omega)=p_{t} for every tt),

[ς,T]Dtdt+DG  Ctgyr22(KLT+KG)elΛbTM^(ω)RM,ς(ω),\int_{[\varsigma,T]}\mathcal{D}_{t}\,dt+\mathcal{D}_{G}\ \ge\ -C_{tg}|y_{r}|^{2}-2(K_{L}T+K_{G})e^{l\Lambda_{b}T}\hat{M}(\omega)-R^{M,\varsigma}(\omega),

which implies (4.2) because CtgCtgC_{tg}\le C^{\vee}_{tg}.

(4b) Expectations. By (4.2), the function

Z=1D([ς,T]Dtdt+DG)+Ctg1Dyς2+2(KLT+KG)elΛbT1DM^+1DRM,ςZ=\mathbf{1}_{D}\Big(\int_{[\varsigma,T]}\mathcal{D}_{t}\,dt+\mathcal{D}_{G}\Big)+C^{\vee}_{tg}\,\mathbf{1}_{D}\,|y_{\varsigma}|^{2}+2(K_{L}T+K_{G})e^{l\Lambda_{b}T}\,\mathbf{1}_{D}\,\hat{M}+\mathbf{1}_{D}\,R^{M,\varsigma}

is nonnegative at every point of Ω\Omega (every term vanishes off DD, and on DD the nonnegativity is exactly (4.2)), and it is a bounded random variable: the first and last terms by claim 1 (DG\mathcal{D}_{G} is a random variable bounded by CDGC_{\mathcal{D}G} at every point, by part (a) of the expansion lemma and the solution definition, and DΩ0D\subseteq\Omega_{0} allows the insertion of 1Ω0\mathbf{1}_{\Omega_{0}} throughout), the middle terms by claim 1 and Step 4's construction of M^\hat{M}. Hence E[NZ]0\mathbb{E}[N\,Z]\ge0, and by linearity, claim 2 (with the event DFςD\in\mathcal{F}_{\varsigma}), and Nyς2=sς2N|y_{\varsigma}|^{2}=|\mathfrak{s}_{\varsigma}|^{2} on Ω0\Omega_{0},

E[1D(N[ς,T]Dtdt+NDG)]  CtgE[1Dsς2]2(KLT+KG)elΛbTNE[1DM^].\mathbb{E}\Big[\mathbf{1}_{D}\Big(N\int_{[\varsigma,T]}\mathcal{D}_{t}\,dt+N\mathcal{D}_{G}\Big)\Big]\ \ge\ -C^{\vee}_{tg}\,\mathbb{E}\big[\mathbf{1}_{D}|\mathfrak{s}_{\varsigma}|^{2}\big]-2(K_{L}T+K_{G})e^{l\Lambda_{b}T}\,N\,\mathbb{E}\big[\mathbf{1}_{D}\hat{M}\big].

By part (a) of the restricted moments lemma applied to each Mγ\overline{M}^{\gamma}, (4.1), and the multiplicativity and monotonicity of xx1/4x\mapsto x^{1/4} recorded there,

E[1DM^]=γ=1lE[1DMγ]γ=1lP(D)3/4E[(Mγ)4]1/4l(4cMκT)1/4N1/2P(D)3/4,\mathbb{E}\big[\mathbf{1}_{D}\hat{M}\big]=\sum_{\gamma=1}^{l}\mathbb{E}\big[\mathbf{1}_{D}\overline{M}^{\gamma}\big]\le\sum_{\gamma=1}^{l}P(D)^{3/4}\,\mathbb{E}\big[(\overline{M}^{\gamma})^{4}\big]^{1/4}\le l\,\big(4c_{M}\kappa_{T}\big)^{1/4}\,N^{-1/2}\,P(D)^{3/4},

and NN1/2=NN\cdot N^{-1/2}=\sqrt{N}, which gives the bound of claim 3 with Cns=2l(KLT+KG)elΛbT(4cMκT)1/4C_{ns}=2\,l\,(K_{L}T+K_{G})\,e^{l\Lambda_{b}T}(4c_{M}\kappa_{T})^{1/4}. \blacksquare

Please log in to copy this version.

Citations

Loading…

Dependency Graph

0 prerequisites

Comments

Loading…