Reason: Republication of the cost-expansion proof against thm:n-agent-cost-expansion-2026b with inline references repointed to def:fluctuation-lqg-cost-2026b and def:c2-population-cost-extension-2026b; no mathematical change from the proof of the 2026a version.
Proof
Throughout, write 1=1Ω0, wt=(Σt,αt)−(St,At)∈Rl+m with components wti, so that wt=N−1/2(st,at) and ρt=∣wt∣, and abbreviate mt=∣st∣2+∣at∣2=Nρt2. Let b denote the aggregate state drift of β, which agrees with the extended aggregate state driftbˉ on Δl×Rm by part (i) of the drift regularity lemma.
Step 0 (measurability conventions). Every Σt lies in Δl and every αt in Rm, at every point of Ω (each agent occupies exactly one state, by the derived notation of the solution definition). Define the modified state-control map Z on [0,T]×Ω by Z=(Σt,αt) for ω∈Ω0 and Z=(e1,0) off Ω0, where e1=(1,0,…,0)∈Δl; by the joint measurability lemma each component of Z is product-measurable with respect to the product σ-algebra of the trace Borel σ-algebra and F. The map (t,ω)↦(St,At) is product-measurable, being the composition of (t,ω)↦t with the continuous trajectory components, by measurability of sequentially continuous functions of measurable maps. Every integrand appearing below is built from these maps, the co-state, and restrictions to Δl×Rm (or Δl) of Lˉ, bˉδ, Gˉ and their first and second partial derivatives - all sequentially continuous on their domains, continuity of the C1 maps and of their partials being part of the extensiondefinitions and of the drift regularity lemma - together with sums, products, and the moduli. Compositions with sequentially continuous functions preserve product-measurability by the composition lemma; sums and products are compositions with the continuous arithmetic operations; and each modulus is nondecreasing on [0,∞), so its sublevel sets are intervals, hence Borel, and ωL(ρt), ωb(ρt), ωG-compositions are product-measurable. Consequently every integrand below is product-measurable after multiplication by 1 (equivalently, after substituting Z); since Ω0 has probability 1, no expectation is affected, and we use this silently. Sections and partial integrals of nonnegative product-measurable maps are measurable by the Tonelli theorem, whose applications here are on the product of two finite, hence σ-finite, measure spaces (the restricted Lebesgue measure of [0,T] has total mass T by the toolkit, and (Ω,F,P) is a probability space).
Step 1 (part (a)). Each modulus is nondecreasing because the supremum is over a set that grows with u, and nonnegative. By clause 3 of the cost extension, any two values of a second partial of Lˉ (or of Gˉ) differ by at most 2Kc, so ωL≤2Kc and ωG≤2Kc; by part (iii) of the drift regularity lemma, ∣∂j∂ibˉγ∣≤3lK on Δl×Rm, so ωb≤6lK. Given ε>0: clause 4 of the cost extension yields δ>0 making all oscillations of the second partials of Lˉ and Gˉ at distance at most δ no larger than ε, whence ωL(u)≤ε and ωG(u)≤ε for u∈[0,δ]; part (iii) of the drift regularity lemma yields the same for ωb.
Step 2 (pointwise Taylor expansions). Fix (t,ω). The segment from (St,At) to (Σt,αt) lies in Δl×Rm (a convex combination of two simplex points has nonnegative entries summing to 1), hence in the open sets V×Rm and U×Rm; likewise the segment from ST to ΣT lies in Δl⊂V. By the definition of the moduli, along these segments the oscillation of each second partial of Lˉ, bˉδ, Gˉ against its value at the base point is at most ωL(ρt), ωb(ρt), ωG(d(ΣT,ST)) respectively. Part (iii) of the Taylor expansion lemma (with n=l+m for Lˉ and bˉδ, whose iterated-partial regularity is supplied by clause 2 of the extensions and by part (i) of the drift regularity lemma, and n=l for Gˉ), together with clause 1 of both extensions and the restriction clause of the drift regularity lemma to rewrite values on the simplex product in terms of L, b, G, gives
where all sums over i,j run over {1,…,l+m} and those over γ,δ over {1,…,l}.
Step 3 (integrability and part (b)). Each map t↦∂iLˉ(St,At) is continuous (a composition of continuous maps) and hence bounded; let C1 bound them all. By the a priori second-moment bound and the hypothesis A<∞, there is C2<∞ with E[∣st∣2]≤C2 for all t, so ∫[0,T]E[mt]dt≤TC2+A<∞; also E[∣sT∣2]≤4N by part (a) of that lemma. Since ∣wti∣≤ρt≤21(1+ρt2) and Nρt2=mt, we get ∫[0,T]E[ρt]dt<∞. Each of the three pieces on the right of (1) then has finite ∫[0,T]E∣⋅∣dt: the linear piece is at most C1∑i∣wti∣≤C1(l+m)ρt; the quadratic piece at most 21(l+m)2Kcρt2 (clause 3); and ∣rtL∣≤(l+m)Kcρt2 since ωL≤2Kc. As t↦L(St,At) is continuous and bounded, ∫[0,T]E∣L(Σt,αt)∣dt<∞. Similarly, using (3) with the pointwise bounds ∣∂γGˉ(ST)∣ finite, ωG≤2Kc, and E∣sT∣2≤4N, the terminal difference is integrable. Hence the random variable ∫[0,T]L(Σt,αt)dt+G(ΣT), whose expectation defines JN[h] in (−∞,+∞] by the N-agent cost definition, is integrable, JN[h] is finite, and by linearity of the expectation and the Fubini theorem,
where the deterministic parts of JMF (mean-field cost, a Riemann integral) were converted by the agreement of the Riemann and Lebesgue integrals. Splitting (4) by (1) and (3) into linear, quadratic, and remainder groups is legitimate since each group is absolutely integrable. Eligibility of ((s),(a)) for the fluctuation linear-quadratic cost: requirement (i) holds with Ω1=Ω0 (Step 0), (ii) is ∫[0,T]E[mt]dt<∞, and (iii) is E[∣sT∣2]≤4N. This proves (b), the finiteness of the expressions in (c) being contained in the estimates above and below.
Step 4 (linear terms). Set ψγ(t)=E[Σtγ]−Stγ=E[wtγ]. By part (b) of the martingale decomposition theorem, Mtγ=Σtγ−Σ0γ−∫[0,t]bγ(Σs,αs)ds is a square-integrable martingale with M0γ=0, so its defining property with r=0 and D=Ω gives E[Mtγ]=E[M0γ]=0. Taking expectations and applying the Fubini theorem to the bounded integrand bγ(Σs,αs) (bounded by 2(l−1)B by part (a) of the decomposition theorem), and using condition 2 of the mean-field trajectory pair,
with g^γ measurable (Tonelli) and bounded. By clause 2 of the stationary co-state definition and additivity of the Riemann integral, with the continuous function pγ(s)=∂γLˉ(Ss,As)−∑δ∂γbˉδ(Ss,As)Psδ we have Ptγ=P0γ+∫[0,t]pγ(s)ds for all t (Riemann and Lebesgue integrals agreeing for continuous integrands). The integration by parts lemma applied to u=Pγ, v=ψγ gives
where the state part was written with the expectation inside the time integral by the Fubini theorem. By clause 2 of the co-state definition at t=T, ∂γGˉ(ST)=−PTγ, so the terminal term equals −∑γPTγψγ(T). By the stationarity clause 3, pointwise in (t,ω), ∑j∂l+jLˉ(St,At)wtl+j=∑j∑δ∂l+jbˉδ(St,At)Ptδwtl+j. Substituting these and (5), and observing that ∂γLˉ(St,At)=pγ(t)+∑δ∂γbˉδ(St,At)Ptδ makes the ∫∑γpγψγ terms cancel,
Moving the expectation back inside the first integral (Fubini), combining the three integrals under one E∫ (each absolutely integrable by Step 3 and boundedness of the coefficients), and using ∑γ∂γbˉδwγ+∑j∂l+jbˉδwl+j=∑i=1l+m∂ibˉδwi together with the definition of g^δ,
Multiply by N and use wt=N−1/2(st,at): the first term becomes −∑γP0γζNγ since ζNγ=Nψγ(0); the two quadratic terms become exactly LQG[(s),(a)] with the fluctuation Hessian coefficients Hij(t) and Fγδ of the fluctuation linear-quadratic cost evaluated at zt=(st,at); and thus the identity of part (c) holds with
RN=NE[∫[0,T](rtL−δ∑Ptδrtb,δ)dt]+NE[rG].
By the bounds in (1)-(3) and ∑δ∣Ptδ∣≤CP, pointwise
and N∣rG∣≤21lωG(d(ΣT,ST))∣sT∣2 since N∣ΣT−ST∣2=∣sT∣2. Taking absolute values inside the expectations (triangle inequality) and applying the Tonelli theorem to the nonnegative product-measurable dominating integrand gives