Reason: Initial published proof: slice-interval product-rule derivation of the derivative formulas, explicit simplex bounds, restriction identity, and the quantitative uniform-continuity estimate.
Proof
Throughout fix γ∈{1,…,l}, write x=(Σ,α) for points of U×Rm, and recall from the extension definition its clauses 1-4. On Δl×Rm we use, from the probability simplex: Σσ≥0 for all σ and ∑σ=1lΣσ=1; and from clause 1 together with the transition-rate family bounds: 0≤βˉ(σ,γ′,x)≤B for x∈Δl×Rm and all admissible index pairs.
(i). For σ∈{1,…,l} let πσ:U×Rm→R be the coordinate function πσ(x)=Σσ. Directly from the definition of the partial derivative (the difference quotients being constant), ∂iπσ exists and equals the constant δiσ, so πσ is a C1 map, and each constant function is again C1 with vanishing partial derivatives. By the definition of the extended aggregate state drift,
bˉγ=σ:σ=γ∑(πσ⋅βˉ(σ,γ,⋅,⋅)−πγ⋅βˉ(γ,σ,⋅,⋅)).
Fix i and a point of U×Rm. The set of slice parameters s for which the point shifted by s along the i-th coordinate remains in the open set U×Rm is an open subset of R containing the given parameter value; choose an open interval I0 around that value contained in it. On I0, every factor above restricts to a one-variable function whose derivative exists and is given by the corresponding partial derivative (this is the definition of the partial derivative), so the one-dimensional sum and product rules, stated for functions on an interval, give that ∂ibˉγ exists at the given point and equals the first displayed formula of the statement. Each term of that formula is continuous: βˉ(σ,γ,⋅,⋅) is continuous because it is differentiable at every point (clause 2 with C1 implies differentiable, and the inequality in the definition of differentiability at a point forces f(a+h)→f(a) as h→0); its partial derivatives are continuous by clause 2 and the definition of a C1 map; the coordinate functions and constants are continuous; and finite sums and products of continuous real functions are continuous (immediate from the sequential formulation of continuity at a point). Hence ∂ibˉγ is continuous for every i, so bˉγ is a C1 map. Applying the same slice-interval argument to the first displayed formula - a finite sum of products of coordinate functions, constants, the functions βˉ(⋅), and their first partials, all of which are C1 by clause 2 - gives that ∂j∂ibˉγ exists and equals the second displayed formula (using ∂jπσ=δjσ and ∂jδiσ=0), and this expression is continuous by the same reasoning; hence each ∂ibˉγ is a C1 map. Finally, for x∈Δl×Rm, clause 1 allows replacing every βˉ by β in the defining formula of bˉγ(x), which then coincides with the defining formula of the aggregate state drift of β at x; hence bˉ agrees with it on Δl×Rm.
(ii). Let x∈Δl×Rm and i∈{1,…,l+m}. Estimating the four groups of terms of the first displayed formula separately: ∑σ=γδiσ∣βˉ(σ,γ,x)∣≤B (at most one σ equals i); ∑σ=γΣσ∣∂iβˉ(σ,γ,x)∣≤K∑σ=γΣσ≤K by clause 3; δiγ∑σ=γ∣βˉ(γ,σ,x)∣≤(l−1)B; and Σγ∑σ=γ∣∂iβˉ(γ,σ,x)∣≤(l−1)K. Altogether ∣∂ibˉγ(x)∣≤lB+lK=l(B+K).
For the Lipschitz estimate let x,y∈Δl×Rm. The segment from x to y stays in Δl×Rm: a convex combination of two points of the simplex has nonnegative entries summing to 1, hence lies in the simplex. The segment therefore lies in the open set U×Rm, and the Taylor expansion lemma, part (i), applied to the C1 map bˉγ with n=l+m and M1=l(B+K) gives ∣bˉγ(x)−bˉγ(y)∣≤l+ml(B+K)d(x,y).
(iii). For x∈Δl×Rm, estimating the six groups of the second displayed formula with clause 3 and ∑σ=γΣσ≤1, Σγ≤1: the terms with δiσ and δjσ contribute at most K each; the terms Σσ∂j∂iβˉ(σ,γ,x) contribute at most K in total; the terms with δiγ and δjγ contribute at most (l−1)K each; and the terms Σγ∂j∂iβˉ(γ,σ,x) contribute at most (l−1)K in total. Hence ∣∂j∂ibˉγ(x)∣≤3K+3(l−1)K=3lK.
For the uniform continuity claim, let ε>0 and set C∗=2ll+mK+2lK (a convenient over-bound: the exact coefficient collected below is 2ll+mK+2(l−1)K, and we over-estimate l−1 by l). By clause 4 there is δ1>0 such that ∣∂j∂iβˉ(σ,γ′,x)−∂j∂iβˉ(σ,γ′,y)∣≤ε/(2l) for all admissible indices whenever d(x,y)≤δ1. Set δ=δ1 if C∗=0 and δ=min(δ1,ε/(2C∗)) otherwise. Let x,y∈Δl×Rm with d(x,y)≤δ, and take the difference of the second displayed formula at x and at y term by term. For the first-derivative factors: each ∂jβˉ(σ,γ′,⋅,⋅) is a C1 map (clause 2) whose partial derivatives are bounded by K on U×Rm (clause 3), so part (i) of the Taylor expansion lemma along the segment from x to y (which lies in the simplex product, as above) gives ∣∂jβˉ(σ,γ′,x)−∂jβˉ(σ,γ′,y)∣≤l+mKd(x,y). For the product terms, writing y=(Σ′,α′):
and similarly with γ in place of σ. Summing all contributions: the δiσ and δjσ groups give at most 2l+mKd(x,y); the δiγ and δjγ groups give at most 2(l−1)l+mKd(x,y); the two product groups give at most 2(l−1)Kd(x,y) from the first summands (using ∣Σσ−Σ′σ∣≤d(x,y) for every σ) plus 2lε(∑σ=γΣ′σ+(l−1)Σ′γ)≤2lε⋅l=2ε from the second summands. Altogether the difference is at most C∗d(x,y)+ε/2≤ε/2+ε/2=ε, uniformly over i,j,γ, as claimed.