Reason: First published version. Proof of the pathwise Gronwall comparison with the mean-field flow driven by the realized control, of measurability and boundedness of the mean-field cost along the realized data, and of the uniform comparison between the N-agent cost and the expected mean-field cost.
Proof
Write G=B[0,T]⊗F for the product σ-algebra on [0,T]×Ω, and λ=λ[0,T]. Real-valued maps are called measurable when they are measurable for the relevant σ-algebra and the Borel σ-algebra of the real line. We use repeatedly that sums, differences, products, absolute values and maxima of finitely many measurable real-valued maps are measurable, and that a product of a measurable map with the indicator of a measurable set is measurable: in each case the map in question is a sequentially continuous function of finitely many measurable real-valued maps, so Sequentially Continuous Functions of Measurable Euclidean Maps are Measurable applies, the indicator of a measurable set being measurable because its preimages are ∅, that set, its complement, or the whole space. For ω∈Ω write Sω=S(Σ0(ω),α^(ω)), which is defined since Σ0(ω)∈Δl and α^(ω)∈UA.
Step 0: joint measurability of the empirical state measure on Ω∗. We show that every component of the map Σ∗ of claim 2 is G-measurable. For a natural number n let Dn={kT2−n:k∈{0,1,…,2n}} and, for t∈[0,T], let tn be the least element of Dn with tn≥t. Define Σt∗,n(ω)=Σtn(ω) for ω∈Ω∗ and Σt∗,n(ω)=e for ω∈/Ω∗. For each γ the map (t,ω)↦(Σt∗,n)γ(ω) is a finite sum, over the finitely many elements s∈Dn, of the product of the indicator of {t∈[0,T]:tn=s}×Ω, a set in G because {t:tn=s} is an interval, with the random variable Σsγ1Ω∗, plus eγ1[0,T]×(Ω∖Ω∗); hence it is G-measurable.
Claim 1. Fix ω∈Ω∗ and abbreviate Σt=Σt(ω), S=Sω, and let u be the path t↦α^(t,ω), an admissible representative of α^(ω) by claim 3 of the realized-control lemma. By claim 2 of the flow stability lemma, S is the map furnished by claim 1 of the existence and uniqueness theorem for the initial value Σ0(ω) and the control u; in particular St∈Δl for every t, the path S satisfies ∣St−Sr∣≤Kb∣t−r∣, and
St=Σ0(ω)+∫[0,t]b^(Ss,u(s))ds(t∈[0,T]),
where b^ is the projected drift of the projected extension lemma, since that existence theorem is stated with the projected drift. Because Ss∈Δl for every s, claim 6 of that lemma gives b^(Ss,u(s))=b(Ss,u(s)), so the integrand may equally be written with the aggregate state drift:
Next, put g(t,ω)=L(Σt∗(ω),α^(t,ω)). The components of (t,ω)↦(Σt∗(ω),α^(t,ω)) are G-measurable by Step 0 and claim 2 of the realized-control lemma, this map takes values in the nonempty subset Δl×Rm of Rl+m, and L is sequentially continuous there by condition 1 of Population Cost Data; so g is G-measurable by Sequentially Continuous Functions of Measurable Euclidean Maps are Measurable. Since Σt∗(ω)∈Δl and α^(t,ω)∈A, claim 1 of the lemma on cost data over a compact control set gives ∣g∣≤C. Writing g+=max(g,0) and g−=max(−g,0), the Tonelli statement of Tonelli and Fubini Theorems makes the map ω↦∫[0,T]g±(t,ω)dλ(t) measurable with respect to F, with values in [0,∞], and both are at most CT, hence real; their difference is ∫[0,T]g(t,ω)dλ(t), which is therefore a random variable bounded in absolute value by CT. Also ω↦G(ΣT∗(ω)) is a random variable by Sequentially Continuous Functions of Measurable Euclidean Maps are Measurable, bounded by C. Hence W is a random variable with ∣W∣≤C(T+1).
Finally let Ξ(ω)=∫[0,T]L(Σt,αt)dt+G(ΣT) be the extended-real-valued map whose expectation is JN[h] by The N-Agent Cost Functional. By claim (v) of the existence and uniqueness theorem, Ξ is measurable as an extended-real-valued map, is bounded below by −CLT−CG, and E[Ξ] is the limit of the expectations of the truncations Ξn=min(Ξ,n). For ω∈Ω∗ we have Σt∗(ω)=Σt(ω) and αt(ω)=α^(t,ω) for every t, so Ξ(ω)=W(ω) and hence ∣Ξ(ω)∣≤C(T+1). Therefore, for every natural number n≥C(T+1), the random variables Ξn and W agree at every point of Ω∗, and both are bounded on all of Ω, the first between −CLT−CG and n. Claim 2 of Almost Sure Inequalities Between Bounded Random Variables Pass to Expectations, applied with the event Ω∗ of probability 1, gives E[Ξn]=E[W] for every such n. Consequently JN[h]=E[Ξ]=E[W], a real number.
whenever Σ,Σ′∈Δl satisfy ∣Σ−Σ′∣≤δ and a∈A. Put η=δe−ΛbT>0 and let BN={M>η}, an event since M is a random variable.
For every ω∈Ω, claim 2 of the attainment theorem shows that Sω is an admissible state path and that F(Σ0(ω),α^(ω))=ΦSω(α^(ω))+G(STω); since the path t↦α^(t,ω) is everywhere A-valued, claim 1 of the running-cost lemma evaluates the first term as an integral, so
Let ω∈Ω∗ with ω∈/BN. Then M(ω)≤η, so claim 1 gives ∣Σt(ω)−Stω∣≤eΛbTη=δ for every t∈[0,T]; as Σt∗(ω)=Σt(ω) and both points lie in Δl, the choice of δ and the monotonicity of the integral, claim 1 of Linearity and Monotonicity of the Lebesgue Integral, give
Then C2P(BN)≤ε/2 for every N≥N0, and therefore, using JN[h]=E[W] from claim 2, the linearity of the expectation and the bound ∣E[V]∣≤E[∣V∣] for a bounded random variable V, both from Linearity and Monotonicity of the Lebesgue Integral,
It remains to note the dependence of N0. The constants C and CF, hence C2, depend only on L, G, Δl, A and T; the number δ depends only on L, G, Δl, A and ε, and η only on those together with Λb and T; and the displayed lower bound for N0 involves besides these only l, B, T and ε. None of them refers to N, to the driving system, to the policy or to the solution, so a single N0 serves for all of them simultaneously. ■