Reason: Proof of thm:n-agent-cost-expansion-2026c, carried forward from the verified proof of the 2026b version: Taylor segments re-argued inside Delta^l x V, off-Omega_0 placeholder moved into Delta^l x A, martingale decomposition used in its 2026c indicator form, Riemann/Lebesgue conversions rerouted to lem:interval-lebesgue-toolkit-2026b claim 3, boundedness rerouted to the extreme value theorem; no change to the mathematical content otherwise.
Proof
Throughout, write 1=1Ω0, wt=(Σt,αt)−(St,At)∈Rl+m with components wti, so that wt=N−1/2(st,at) and ρt=∣wt∣, and abbreviate mt=∣st∣2+∣at∣2=Nρt2. Let b denote the aggregate state drift of β, which agrees with the extended aggregate state driftbˉ on Δl×A by part (i) of the drift regularity lemma; every state-control point at which this agreement is invoked below lies in Δl×A, both αt and At taking values in A.
Step 0 (measurability conventions). Every Σt lies in Δl (each agent occupies exactly one state, by the derived notation of the solution definition) and every αt lies in A (the policy being A-valued), at every point of Ω. Define the modified state-control map Z on [0,T]×Ω by Z=(Σt,αt) for ω∈Ω0 and Z=(e1,a0) off Ω0, where e1=(1,0,…,0)∈Δl and a0 is a fixed point of the nonempty control set A, so that Z takes values in Δl×A; by the joint measurability lemma each component of Z is product-measurable with respect to the product σ-algebra of the trace Borel σ-algebra and F. The map (t,ω)↦(St,At) is product-measurable, being the composition of (t,ω)↦t with the continuous trajectory components, by measurability of sequentially continuous functions of measurable maps. Every integrand appearing below is built from these maps, the co-state, and restrictions of Lˉ to Δl×Rm, of bˉδ to Δl×V (which contains the range Δl×A of Z, A being a subset of V by the extension definition), and of Gˉ to Δl, together with their first and second partial derivatives - all sequentially continuous on their domains, continuity of the C1 maps and of their partials being part of the extensiondefinitions and of the drift regularity lemma - together with sums, products, and the moduli. Compositions with sequentially continuous functions preserve product-measurability by the composition lemma; sums and products are compositions with the continuous arithmetic operations; and each modulus is nondecreasing on [0,∞), so its sublevel sets are intervals, hence Borel, and ωL(ρt), ωb(ρt), ωG-compositions are product-measurable. Consequently every integrand below is product-measurable after multiplication by 1 (equivalently, after substituting Z); since Ω0 has probability 1, no expectation is affected, and we use this silently. Sections and partial integrals of nonnegative product-measurable maps are measurable by the Tonelli theorem, whose applications here are on the product of two finite, hence σ-finite, measure spaces (the restricted Lebesgue measure of [0,T] has total mass T by the toolkit, and (Ω,F,P) is a probability space).
Step 1 (part (a)). Each modulus is nondecreasing because the supremum is over a set that grows with u, and nonnegative. By clause 3 of the cost extension, any two values of a second partial of Lˉ (or of Gˉ) differ by at most 2Kc, so ωL≤2Kc and ωG≤2Kc; by part (iii) of the drift regularity lemma, ∣∂j∂ibˉγ∣≤3lK on Δl×V, so ωb≤6lK. Given ε>0: clause 4 of the cost extension yields δ>0 making all oscillations of the second partials of Lˉ and Gˉ at distance at most δ no larger than ε, whence ωL(u)≤ε and ωG(u)≤ε for u∈[0,δ]; part (iii) of the drift regularity lemma yields the same for ωb.
Step 2 (pointwise Taylor expansions). Fix (t,ω). The segment from (St,At) to (Σt,αt) lies in Δl×V: a convex combination of two simplex points has nonnegative entries summing to 1, and both At and αt lie in A⊆V, the set V being convex by the extension definition. Hence the segment lies in Δl×Rm - the domain of the ωL supremum - and in the open sets W×Rm and U×V; likewise the segment from ST to ΣT lies in Δl⊂W. By the definition of the moduli, along these segments the oscillation of each second partial of Lˉ, bˉδ, Gˉ against its value at the base point is at most ωL(ρt), ωb(ρt), ωG(d(ΣT,ST)) respectively. Part (iii) of the Taylor expansion lemma (with n=l+m for Lˉ and bˉδ, whose iterated-partial regularity is supplied by clause 2 of the extensions and by part (i) of the drift regularity lemma, and n=l for Gˉ), together with clause 1 of both extensions and the restriction clause of the drift regularity lemma to rewrite values on the simplex product in terms of L, b, G, gives
where all sums over i,j run over {1,…,l+m} and those over γ,δ over {1,…,l}.
Step 3 (integrability and part (b)). Each map t↦∂iLˉ(St,At) is continuous relative to [0,T] with the metric of the real line (a composition of continuous maps, the Euclidean and metric notions of continuity agreeing for real-valued maps by claim 1 of the continuity agreement lemma) and therefore attains a maximum and a minimum on [0,T] by the extreme value theorem; let C1 bound them all in absolute value. By the a priori second-moment bound and the hypothesis A2<∞, there is C2<∞ with E[∣st∣2]≤C2 for all t, so ∫[0,T]E[mt]dt≤TC2+A2<∞; also E[∣sT∣2]≤4N by part (a) of that lemma. Since ∣wti∣≤ρt≤21(1+ρt2) and Nρt2=mt, we get ∫[0,T]E[ρt]dt<∞. Each of the three pieces on the right of (1) then has finite ∫[0,T]E∣⋅∣dt: the linear piece is at most C1∑i∣wti∣≤C1(l+m)ρt; the quadratic piece at most 21(l+m)2Kcρt2 (clause 3); and ∣rtL∣≤(l+m)Kcρt2 since ωL≤2Kc. As t↦L(St,At) is continuous and bounded, ∫[0,T]E∣L(Σt,αt)∣dt<∞. Similarly, using (3) with the pointwise bounds ∣∂γGˉ(ST)∣ finite, ωG≤2Kc, and E∣sT∣2≤4N, the terminal difference is integrable. Hence the random variable ∫[0,T]L(Σt,αt)dt+G(ΣT), whose expectation defines JN[h] in (−∞,+∞] by the N-agent cost definition, is integrable, JN[h] is finite, and by linearity of the expectation and the Fubini theorem,
where the deterministic parts of JMF (mean-field cost, a Riemann integral) were converted to Lebesgue integrals by claim 3 of the integral toolkit on a compact interval, the integrand t↦L(St,At) being continuous as recorded in the mean-field cost definition. Splitting (4) by (1) and (3) into linear, quadratic, and remainder groups is legitimate since each group is absolutely integrable. Eligibility of ((s),(a)) for the fluctuation linear-quadratic cost: requirement (i) holds with Ω1=Ω0 (Step 0), (ii) is ∫[0,T]E[mt]dt<∞, and (iii) is E[∣sT∣2]≤4N. This proves (b), the finiteness of the expressions in (c) being contained in the estimates above and below.
Step 4 (linear terms). Set ψγ(t)=E[Σtγ]−Stγ=E[wtγ]. By part (b) of the martingale decomposition theorem, Mtγ=Σtγ−Σ0γ−∫[0,t]1bγ(Σs,αs)ds is a square-integrable martingale with M0γ=0, so its defining property with r=0 and D=Ω gives E[Mtγ]=E[M0γ]=0. Taking expectations and applying the Fubini theorem to the bounded integrand 1bγ(Σs,αs) (bounded by 2(l−1)B by part (a) of the decomposition theorem; dropping the indicator changes no expectation, Ω0 having probability 1), and using condition 2 of the mean-field trajectory pair (its Riemann integral converted to a Lebesgue integral over [0,t] by claim 3 of the integral toolkit, applied for t>0 to the restriction of the continuous integrand, continuous by claim 1 of restriction stability, the case t=0 being trivial),
with g^γ measurable (Tonelli) and bounded. By clause 2 of the stationary co-state definition and additivity of the Riemann integral, with the continuous function pγ(s)=∂γLˉ(Ss,As)−∑δ∂γbˉδ(Ss,As)Psδ we have Ptγ=P0γ+∫[0,t]pγ(s)ds for all t (additivity applied for 0<t<T, the cases t=0 and t=T holding by the zero-integral conventions of clause 2 of the co-state definition; the Riemann integrals over [0,t] are converted to Lebesgue integrals as in the preceding conversion). The integration by parts lemma applied to u=Pγ, v=ψγ gives
where the state part was written with the expectation inside the time integral by the Fubini theorem. By clause 2 of the co-state definition at t=T, ∂γGˉ(ST)=−PTγ, so the terminal term equals −∑γPTγψγ(T). By the stationarity clause 3, pointwise in (t,ω), ∑j∂l+jLˉ(St,At)wtl+j=∑j∑δ∂l+jbˉδ(St,At)Ptδwtl+j. Substituting these and (5), and observing that ∂γLˉ(St,At)=pγ(t)+∑δ∂γbˉδ(St,At)Ptδ makes the ∫∑γpγψγ terms cancel,
Moving the expectation back inside the first integral (Fubini), combining the three integrals under one E∫ (each absolutely integrable by Step 3 and boundedness of the coefficients), and using ∑γ∂γbˉδwγ+∑j∂l+jbˉδwl+j=∑i=1l+m∂ibˉδwi together with the definition of g^δ,
Multiply by N and use wt=N−1/2(st,at): the first term becomes −∑γP0γζNγ since ζNγ=Nψγ(0); the two quadratic terms become exactly LQG[(s),(a)] with the fluctuation Hessian coefficients Hij(t) and Fγδ of the fluctuation linear-quadratic cost evaluated at zt=(st,at); and thus the identity of part (c) holds with
RN=NE[∫[0,T](rtL−δ∑Ptδrtb,δ)dt]+NE[rG].
By the bounds in (1)-(3) and ∑δ∣Ptδ∣≤CP, pointwise
and N∣rG∣≤21lωG(d(ΣT,ST))∣sT∣2 since N∣ΣT−ST∣2=∣sT∣2. Taking absolute values inside the expectations (triangle inequality) and applying the Tonelli theorem to the nonnegative product-measurable dominating integrand gives