Proof of Convergence of the Optimal N-Agent Value to the Optimal Mean-Field Value and Concentration of Asymptotically Optimal Controls
theoremthm:n-agent-optimal-value-mean-field-limit-2026aThroughout, is the set of natural numbers, and we use freely the elementary order arithmetic of Elementary Order Arithmetic in an Ordered Field and the properties of the absolute value of Properties of the Absolute Value in an Ordered Field.
Step 0. Four elementary devices.
(a) Eventual lower bounds and the limit inferior. Let be a bounded sequence of real numbers, let be a real number, and let be such that for every . Then .
In the notation of the definition of the limit inferior, write . The number is a lower bound of , so part (ii) of the definition of the greatest lower bound gives . The number belongs to the set , whose least upper bound is by definition ; a least upper bound is in particular an upper bound, so . Transitivity of the order gives the assertion.
(b) Removing a doubled epsilon. Let and be real numbers such that for every real . Then .
Suppose not, so that , and put , a positive real number. By claim 8 of Elementary Order Arithmetic in an Ordered Field the number is positive and ; applying that claim to shows that is positive and satisfies . The hypothesis applied to this gives , which is impossible.
(c) Changing the policy does not change the law of the initial empirical state. Let and let two solutions of the controlled -agent dynamics on for , and the -th driving system be given, for policies in , with regular events and and empirical state measures and . Then and agree at every point of an event of probability , and they have the same law as random elements of .
Write for the initial states of the driving system. Condition 1 of the definition of a solution holds at every point of the regular event, so for every and every , and the same identity holds for the second solution at every . The empirical state measure at time is determined by the states at time through the formula recalled in that definition, so for every in , and because a -algebra is closed under finite intersections. Since , claim 3 of the basic properties of a measure gives . The complement of is the union of those two sets, so claim 4 there (countable subadditivity), applied to the sequence whose first two terms are those sets and all of whose further terms are empty, gives , and claim 3 again gives .
Let be a Borel set of . Then . The sets and are disjoint with union , so claim 1 of the basic properties (finite additivity) gives , and the last term vanishes by claim 2 (monotonicity) applied to the inclusion . The same computation applies to , so the two probabilities agree; by the definition of the law, the laws coincide.
Consequently, since converges in distribution to , and convergence in distribution is by definition weak convergence of the laws, the same convergence holds for the empirical state measures at time of any family consisting, for each , of a solution for the -th driving system and some policy in .
(d) Composites of strictly increasing sequences. If and are strictly increasing sequences of natural numbers, then is strictly increasing: an induction shows that implies , and for every .
Step 1. Proof of claim 1. By the preamble, is a nonempty set of real numbers of which is a lower bound, and . Part (ii) of the definition of the greatest lower bound gives . By part (i) the number is a lower bound of ; since we have , so . Moreover , so by claim 6 of the absolute-value lemma, and by transitivity. From and claim 6 again, .
Step 2. The open-loop policy is admissible. By claim 1 of the optimal-set structure lemma the set is nonempty. Let and let be an admissible representative of , which exists by claim 2 of the flow stability lemma; thus for every , and, being a representative of an element of the Lebesgue space of square-integrable vector-valued functions on , the map has components measurable with respect to the trace Borel -algebra on and the Borel -algebra of the real line. Hence the open-loop policy lemma applies, and its claim 1 shows that is an observation-driven control policy with horizon , control dimension and channels. By the formulas defining it in that lemma, every value of is of the form with , hence lies in ; this is exactly the requirement of the definition of an -valued policy. Therefore , and in particular is a real number lying in for every .
Step 3. The initial states converge in mean square. For put .
First, is a random variable. Indeed is a random element of by claim 5 of the realized-control lemma; the coordinate projections are Borel measurable by claim 1 of the lemma on Borel measurability in Euclidean space, so each component of is measurable by claim 4 (composition) of the Borel toolkit on a metric space; and the map sending to is sequentially continuous, so the lemma on sequentially continuous functions of measurable Euclidean maps applies.
Second, everywhere: for every both and lie in , so they have Euclidean norm at most by claim 1 of the compactness lemma, and claims 5 and 6 of the lemma on the Euclidean norm give ; if then by two applications of claim 10 of the order-arithmetic lemma in the nonstrict form obtained by adjoining the case of equality, the multipliers and being nonnegative.
Now let be real and put . By claims 2 and 3 of the lemma on convergence in distribution to a constant, each is an event and the sequence converges to ; and , because is the restriction to of the Euclidean distance on , as fixed in the compactness lemma, and that distance is expressed through the norm by claim 2 of the norm lemma. Let be the function equal to on and to elsewhere. For we get ; for we get and hence, by the same two applications of claim 10, . Thus at every point of .
Both sides are random variables bounded everywhere, so claim 1 of the lemma on almost sure inequalities between bounded random variables gives , and claim 2 of the linearity and monotonicity theorem for the integral together with the lemma on the integral of an indicator function gives
Let be real. Choose to be the smaller of and , a positive real number; then , the first inequality by claim 10 in the nonstrict form applied to with the nonnegative multiplier . Choose with for every . For such we get , the lower bound because everywhere. As was arbitrary, converges to .
Step 4. Proof of claim 2. Retain , and from Step 2. By part (ii) of the existence, uniqueness and regularity theorem, for every the set of solutions of the controlled -agent dynamics on for , , the -th driving system and the policy is nonempty; these sets all lie in the set of all such solutions for all , so Axiom of Countable Choice provides one for every . Write for the empirical state measure of the chosen solution. By device (c) the maps and agree at every point of an event of probability , so the random variables and do as well; both are bounded everywhere, so claim 2 of the almost sure expectation lemma gives , and by Step 3 this converges to .
By claim 2 of the flow stability lemma the mean-field flow is the map furnished by claim 1 of the existence and uniqueness theorem for the generalized mean-field trajectory applied to the initial value and the control ; hence is the generalized mean-field trajectory pair with value at that appears in the setting of the mean-square tracking proposition. All the data required there are now in place, with and with the map and the open-loop policy of Step 2. Therefore claim 2 of the corollary on convergence of the -agent cost under an open-loop control applies and shows that converges to the generalized mean-field cost of that pair. (The cost has the same value for every solution, as recorded in the definition of the -agent cost, so no ambiguity arises from the choice made above.)
Finally, by the definition of the mean-field cost of a control from an initial state, evaluated with the admissible representative of , that generalized mean-field cost is ; and because . This proves claim 2.
Step 5. Claim 3: the upper estimate. The sequences and are bounded, both being bounded in absolute value by , so that is a strictly positive bound as the definition of a bounded sequence requires. For every the number lies in , of which is a lower bound, so . Claim 2 of the basic properties of the limit inferior and the limit superior therefore gives , while claim 5 there, applied to the sequence , which converges to by claim 2, gives . Hence .
Step 6. Claim 3: the lower estimate. Let be real. Fix . By claim 1 of the order-arithmetic lemma , so is not a lower bound of : otherwise part (ii) of the definition of the greatest lower bound would give . Hence there is a policy with , and, by part (ii) of the existence theorem, a solution for the -th driving system and that policy. The set of all pairs consisting of such a policy and such a solution is therefore nonempty, and these sets lie in a common set; so Axiom of Countable Choice provides a pair for every . Write for the chosen policy and for the empirical state measure of the chosen solution, so that for every .
By device (c) the sequence converges in distribution to . Hence every hypothesis of the asymptotic lower bound theorem is met by the data consisting of the given driving systems, the policies and the chosen solutions, and claim 3 there gives that is bounded with
By claim 3 of the basic properties of the limit inferior and the limit superior there is with for every . For such ,
so and hence, adding to both sides by claim 1 of the order-arithmetic lemma, ; in particular for every . Device (a), applied to the bounded sequence , gives , that is, . Since was arbitrary, device (b) gives .
Step 7. Proof of claim 3. By claim 1 of the basic properties of the limit inferior and the limit superior, . With Steps 5 and 6 this yields
so both the limit inferior and the limit superior equal , and claim 5 there shows that converges to .
For the second sequence, claim 1 and the hypothesis on give , hence by claim 6 of the absolute-value lemma. Claim 5 of that lemma (the triangle inequality) then gives
The sequence converges to , directly by the definition of the limit applied to , and converges to by hypothesis, so their sum converges to by claim 1 of the arithmetic of limits of real sequences. Claim 3 of the order properties of limits of real sequences now gives that converges to .
Step 8. Claim 4: measurability. By claim 4 of the optimal-set structure lemma the function on is sequentially continuous: if is a sequence in converging to in , then converges to ; by claim 3 of the arithmetic of limits, converges to . Consequently both and are lower semicontinuous on : given such a sequence and a real , the definition of the limit provides with for , whence by claim 6 of the absolute-value lemma, and similarly for ; claims 1 and 3 of the sequential characterization of lower semicontinuity, applied with the subset of , give the assertion.
By claim 5 of the Borel toolkit the function is measurable with respect to and the Borel -algebra of the real line. By claim 5 of the realized-control lemma the pair is a random element of , that is, measurable with respect to and . Claim 4 (composition) of the Borel toolkit therefore shows that is measurable, that is, a random variable.
Fix a real and put . Since and is lower semicontinuous on , claim 3 of the lemma on sublevel and superlevel sets, applied with the subset of , shows that is closed in the topology of open subsets of ; hence by claim 1 of the Borel toolkit. Moreover is the preimage of under , so by the definition of the law
Step 9. Claim 4: the limit. Suppose, for a contradiction, that does not converge to . Since for every , negating the definition of the limit provides a real such that for every there is with .
Let , which is nonempty by the case , and let be the set of pairs with . For the case provides with , so ; thus is a binary relation on such that every element of is related to some element of . Let be an element of , which exists because is nonempty. Then Axiom of Dependent Choice, applied to , and , provides a sequence in with and , that is , for every . It is strictly increasing and satisfies for every .
By claim 1 of the asymptotic lower bound theorem each is a Borel measure on with , and by claim 4 of the compactness lemma is compact. Hence the weak sequential compactness theorem applies to and provides a strictly increasing sequence of natural numbers and a Borel measure on with such that converges weakly to . By device (d) the sequence is strictly increasing.
By claim 3 the sequence converges to , so claim 4 of the asymptotic lower bound theorem applies to the strictly increasing sequence and to : the set belongs to and .
The sets and are disjoint: for claim 2 of the optimal-set structure lemma gives , and is false because . Hence . Since is finite, claim 3 of the basic properties of a measure gives , and claim 2 (monotonicity) gives ; as a measure takes values in , this forces .
The measures and are probability measures on and is nonempty, so claim 4 of the portmanteau theorem, applied to the closed set , gives
On the other hand for every , and the sequence is bounded, with strictly positive bound , because by monotonicity. Hence claim 2 of the basic properties of the limit inferior and the limit superior gives , and claim 1 there gives . Combining, , contradicting .
Therefore converges to , which completes the proof of claim 4 and of the theorem.
Loading…
Prerequisites
0fa69dd2-ae7b-417a-8b3a-1ee6bd0f9cc6