Adopt the setting, hypotheses and notation of the pathwise tracking lemma : the affine-controlled transition-rate family ( β 0 , β 1 ) (\beta_{0},\beta_{1}) ( β 0 , β 1 ) on l l l states with control set A ⊆ R m \mathcal{A}\subseteq\mathbb{R}^{m} A ⊆ R m , its transition-rate family β \beta β with rate bound B B B , aggregate state drift b b b , projected drift b ^ \hat{b} b ^ and state-Lipschitz constant Λ b \Lambda_{b} Λ b ; the observation-rate family β ~ \tilde{\beta} β ~ ; the horizon T > 0 T>0 T > 0 ; the population cost data ( L , G ) (L,G) ( L , G ) , convex in the control on A \mathcal{A} A ; the probability simplex Δ l \Delta^{l} Δ l ; the set U A \mathcal{U}_{\mathcal{A}} U A of A \mathcal{A} A -valued controls ; the mean-field flow S ( z 0 , ξ ) S(z_{0},\xi) S ( z 0 , ξ ) of claim 2 of the flow stability lemma (the two-argument flow map, distinguished throughout from the trajectory S S S of the mean-field trajectory pair below by the presence of its arguments); the mean-field cost F F F ; the N N N -agent driving system ( Ω , F , P ) (\Omega,\mathcal{F},P) ( Ω , F , P ) ; the A \mathcal{A} A -valued policy h h h ; the solution with regular event Ω 0 \Omega_{0} Ω 0 , empirical state measure Σ t \Sigma_{t} Σ t , control α t \alpha_{t} α t and system filtration ( F t s y s ) t ∈ [ 0 , T ] (\mathcal{F}^{\mathrm{sys}}_{t})_{t\in[0,T]} ( F t sys ) t ∈ [ 0 , T ] ; the realized control α ^ \hat{\alpha} α ^ of the realized-control lemma ; and the martingale part M = ( M 1 , … , M l ) M=(M^{1},\dots,M^{l}) M = ( M 1 , … , M l ) of the martingale decomposition theorem . The hypothesis (LipC) of the tracking lemma, with constants K L K_{L} K L and K G K_{G} K G , is assumed only in claim 3.
Let further ( S , A ) (S,A) ( S , A ) be a mean-field trajectory pair for β \beta β with horizon T T T , let ( U , V , β ˉ ) (U,V,\bar{\beta}) ( U , V , β ˉ ) be a twice continuously differentiable extension of β \beta β with derivative bound K K K and extended aggregate state drift b ˉ \bar{b} b ˉ , let ( U c , L ˉ , G ˉ ) (U_{c},\bar{L},\bar{G}) ( U c , L ˉ , G ˉ ) be a twice continuously differentiable extension of ( L , G ) (L,G) ( L , G ) with second-derivative bound K c K_{c} K c , and let P P P be a stationary co-state for these data, so that the setting of the first-order expansion lemma for the recentred N N N -agent cost is in force; adopt from it the fixed constant C P C_{P} C P , the mean-field Hamiltonian H t ( Σ , α ) \mathcal{H}_{t}(\Sigma,\alpha) H t ( Σ , α ) and its state derivative coefficients ∂ γ H t ( Σ , α ) \partial_{\gamma}\mathcal{H}_{t}(\Sigma,\alpha) ∂ γ H t ( Σ , α ) , the quantities D t \mathcal{D}_{t} D t and D G \mathcal{D}_{G} D G with their bounds C D C_{\mathcal{D}} C D and C D G C_{\mathcal{D}G} C D G , the fluctuation process s t = N ( Σ t − S t ) \mathfrak{s}_{t}=\sqrt{N}\,(\Sigma_{t}-S_{t}) s t = N ( Σ t − S t ) , and the abbreviation y t = Σ t − S t y_{t}=\Sigma_{t}-S_{t} y t = Σ t − S t . Adopt also the setting and notation of the time-shift lemma and of the to-go comparison lemma for these data (the affine family, cost data and extensions being the same, and the triple ( S , A , P ) (S,A,P) ( S , A , P ) that of this statement), including the horizon-θ \theta θ instance notation superscripted by [ θ ] [\theta] [ θ ] and, in claim 3, the hypothesis (TG) with constants ε t g \varepsilon_{tg} ε t g and C t g C_{tg} C t g .
Stopping times are those of ( F t s y s ) t ∈ [ 0 , T ] (\mathcal{F}^{\mathrm{sys}}_{t})_{t\in[0,T]} ( F t sys ) t ∈ [ 0 , T ] ; for a stopping time ς \varsigma ς , F ς \mathcal{F}_{\varsigma} F ς denotes the σ \sigma σ -algebra of events prior to ς \varsigma ς , sampled functions such as M ς γ M^{\gamma}_{\varsigma} M ς γ and Σ ς \Sigma_{\varsigma} Σ ς are the sampled functions , and 1 { t < ς } \mathbf{1}_{\{t<\varsigma\}} 1 { t < ς } is the pre-stopping-time indicator . For a family X = ( X t ) t ∈ [ 0 , T ] X=(X_{t})_{t\in[0,T]} X = ( X t ) t ∈ [ 0 , T ] , a point ω ∈ Ω \omega\in\Omega ω ∈ Ω at which the path t ↦ X t ( ω ) t\mapsto X_{t}(\omega) t ↦ X t ( ω ) is measurable and bounded, and a stopping time ς \varsigma ς , write ∫ [ ς , T ] X t d t \int_{[\varsigma,T]}X_{t}\,dt ∫ [ ς , T ] X t d t for the Lebesgue integral of the path over [ ς ( ω ) , T ] [\varsigma(\omega),T] [ ς ( ω ) , T ] when ς ( ω ) < T \varsigma(\omega)<T ς ( ω ) < T , and 0 0 0 when ς ( ω ) = T \varsigma(\omega)=T ς ( ω ) = T . Write 1 D \mathbf{1}_{D} 1 D for the indicator of a set D D D , E \mathbb{E} E for the expectation , ∣ ⋅ ∣ |\cdot| ∣ ⋅ ∣ for the Euclidean norm , and e x e^{x} e x for the real exponential function . Adopt the constants c M = 6 l 2 ( 2 ( l − 1 ) ) 4 c_{M}=6\,l^{2}(2(l-1))^{4} c M = 6 l 2 ( 2 ( l − 1 ) ) 4 and κ T = B T + ( B T ) 2 \kappa_{T}=BT+(BT)^{2} κ T = BT + ( BT ) 2 of the restricted moments lemma , and for a nonnegative real x x x let x 1 / 4 x^{1/4} x 1/4 and x 3 / 4 x^{3/4} x 3/4 be as defined there. Define
C t g ∨ = max ( C t g , l K c 2 ) , C n s = 2 l ( K L T + K G ) e l Λ b T ( 4 c M κ T ) 1 / 4 . C^{\vee}_{tg}=\max\Big(C_{tg},\ \tfrac{l\,K_{c}}{2}\Big),\qquad C_{ns}=2\,l\,(K_{L}T+K_{G})\,e^{l\,\Lambda_{b}T}\,\big(4\,c_{M}\,\kappa_{T}\big)^{1/4}. C t g ∨ = max ( C t g , 2 l K c ) , C n s = 2 l ( K L T + K G ) e l Λ b T ( 4 c M κ T ) 1/4 .
Let ς \varsigma ς be a stopping time. Define, at every ω ∈ Ω 0 \omega\in\Omega_{0} ω ∈ Ω 0 ,
R M , ς = − ∑ γ = 1 l P T γ ( M T γ − M ς γ ) + ∫ [ ς , T ] ∑ γ = 1 l ∂ γ H t ( S t , A t ) ( M t γ − M ς γ ) d t , R^{M,\varsigma}=-\sum_{\gamma=1}^{l}P^{\gamma}_{T}\,\big(M^{\gamma}_{T}-M^{\gamma}_{\varsigma}\big)+\int_{[\varsigma,T]}\sum_{\gamma=1}^{l}\partial_{\gamma}\mathcal{H}_{t}(S_{t},A_{t})\,\big(M^{\gamma}_{t}-M^{\gamma}_{\varsigma}\big)\,dt , R M , ς = − γ = 1 ∑ l P T γ ( M T γ − M ς γ ) + ∫ [ ς , T ] γ = 1 ∑ l ∂ γ H t ( S t , A t ) ( M t γ − M ς γ ) d t ,
and R M , ς = 0 R^{M,\varsigma}=0 R M , ς = 0 at every ω ∉ Ω 0 \omega\notin\Omega_{0} ω ∈ / Ω 0 .
Then the following hold.
1. (Anchored pathwise identity.) For every ω ∈ Ω 0 \omega\in\Omega_{0} ω ∈ Ω 0 , the paths on [ 0 , T ] [0,T] [ 0 , T ] of t ↦ L ( Σ t , α t ) t\mapsto L(\Sigma_{t},\alpha_{t}) t ↦ L ( Σ t , α t ) , t ↦ D t t\mapsto\mathcal{D}_{t} t ↦ D t , t ↦ y t γ t\mapsto y^{\gamma}_{t} t ↦ y t γ and t ↦ M t γ t\mapsto M^{\gamma}_{t} t ↦ M t γ are measurable and bounded, all integrals below exist, and, writing r = ς ( ω ) r=\varsigma(\omega) r = ς ( ω ) ,
∫ [ ς , T ] L ( Σ t , α t ) d t + G ( Σ T ) − ∫ [ ς , T ] L ( S t , A t ) d t − G ( S T ) + ∑ γ = 1 l P r γ y r γ = ∫ [ ς , T ] D t d t + D G + R M , ς . \int_{[\varsigma,T]}L(\Sigma_{t},\alpha_{t})\,dt+G(\Sigma_{T})-\int_{[\varsigma,T]}L(S_{t},A_{t})\,dt-G(S_{T})+\sum_{\gamma=1}^{l}P^{\gamma}_{r}\,y^{\gamma}_{r}=\int_{[\varsigma,T]}\mathcal{D}_{t}\,dt+\mathcal{D}_{G}+R^{M,\varsigma}. ∫ [ ς , T ] L ( Σ t , α t ) d t + G ( Σ T ) − ∫ [ ς , T ] L ( S t , A t ) d t − G ( S T ) + γ = 1 ∑ l P r γ y r γ = ∫ [ ς , T ] D t d t + D G + R M , ς .
Moreover the functions 1 Ω 0 ∫ [ ς , T ] D t d t \mathbf{1}_{\Omega_{0}}\int_{[\varsigma,T]}\mathcal{D}_{t}\,dt 1 Ω 0 ∫ [ ς , T ] D t d t (set to 0 0 0 off Ω 0 \Omega_{0} Ω 0 ), R M , ς R^{M,\varsigma} R M , ς , 1 Ω 0 M ς γ \mathbf{1}_{\Omega_{0}}M^{\gamma}_{\varsigma} 1 Ω 0 M ς γ , 1 Ω 0 y ς γ \mathbf{1}_{\Omega_{0}}y^{\gamma}_{\varsigma} 1 Ω 0 y ς γ , 1 Ω 0 ∣ s ς ∣ 2 \mathbf{1}_{\Omega_{0}}|\mathfrak{s}_{\varsigma}|^{2} 1 Ω 0 ∣ s ς ∣ 2 and ω ↦ P ς ( ω ) γ \omega\mapsto P^{\gamma}_{\varsigma(\omega)} ω ↦ P ς ( ω ) γ are bounded random variables.
2. (Vanishing of the anchored martingale remainder.) For every event D ∈ F ς D\in\mathcal{F}_{\varsigma} D ∈ F ς ,
E [ 1 D R M , ς ] = 0. \mathbb{E}\big[\mathbf{1}_{D}\,R^{M,\varsigma}\big]=0 . E [ 1 D R M , ς ] = 0.
3. (Post-exit comparison.) Assume (LipC), assume [ A ] ∈ M S 0 ∗ [A]\in\mathcal{M}^{*}_{S_{0}} [ A ] ∈ M S 0 ∗ (the class of A A A is an optimal mean-field control from S 0 S_{0} S 0 , as in claim 5 of the time-shift lemma ), and assume hypothesis (TG) of the to-go comparison lemma with constants ε t g > 0 \varepsilon_{tg}>0 ε t g > 0 and C t g ≥ 0 C_{tg}\ge0 C t g ≥ 0 . Let D ∈ F ς D\in\mathcal{F}_{\varsigma} D ∈ F ς be an event with D ⊆ Ω 0 D\subseteq\Omega_{0} D ⊆ Ω 0 such that ∣ Σ ς ( ω ) ( ω ) − S ς ( ω ) ∣ ≤ ε t g |\Sigma_{\varsigma(\omega)}(\omega)-S_{\varsigma(\omega)}|\le\varepsilon_{tg} ∣ Σ ς ( ω ) ( ω ) − S ς ( ω ) ∣ ≤ ε t g for every ω ∈ D \omega\in D ω ∈ D . Then
E [ 1 D ( N ∫ [ ς , T ] D t d t + N D G ) ] ≥ − C t g ∨ E [ 1 D ∣ s ς ∣ 2 ] − C n s N P ( D ) 3 / 4 , \mathbb{E}\Big[\mathbf{1}_{D}\Big(N\int_{[\varsigma,T]}\mathcal{D}_{t}\,dt+N\,\mathcal{D}_{G}\Big)\Big]\ \ge\ -\,C^{\vee}_{tg}\,\mathbb{E}\big[\mathbf{1}_{D}\,|\mathfrak{s}_{\varsigma}|^{2}\big]-C_{ns}\,\sqrt{N}\,P(D)^{3/4}, E [ 1 D ( N ∫ [ ς , T ] D t d t + N D G ) ] ≥ − C t g ∨ E [ 1 D ∣ s ς ∣ 2 ] − C n s N P ( D ) 3/4 ,
all expectations being defined, the integrands being bounded random variables by claim 1 and part (a) of the first-order expansion lemma .