Convergence of the Optimal N-Agent Value to the Optimal Mean-Field Value and Concentration of Asymptotically Optimal Controls
theoremAnalysisProbabilitythm:n-agent-optimal-value-mean-field-limit-2026aAdopt the setting, hypotheses and notation of the asymptotic lower bound theorem for the -agent cost. In particular: is an affine-controlled transition-rate family on states with control set and projected extension , with rate bound , aggregate state drift and state-Lipschitz constant ; is an observation-rate family on states with channels; is a real number; is population cost data on states with control dimension that is convex in the control on ; is the bound of claim 1 of the lemma on cost data over a compact control set; is the mean-field flow of claim 2 of the flow stability lemma; is the mean-field cost of a control from an initial state under ; is the set of -valued controls; and , , and are the metrics and the product space of the compactness lemma for the simplex, the control set and their product, where is the probability simplex. Write for the Borel -algebra of and for the Euclidean norm.
The initial state is , and and are the optimal mean-field value and the set of optimal mean-field controls from in the sense of the definition of the optimal value, the optimal controls and the optimal trajectories of the mean-field problem. Let be the deviation from the optimal set furnished by claim 2 of the optimal-set structure lemma, a function on .
For every natural number there are given an -agent driving system , an observation-driven control policy with horizon , control dimension and channels that is -valued, and a solution of the controlled -agent dynamics on for , , that driving system and that policy, with empirical state measure ; and is the realized control attached to those data by the realized-control lemma. Write for the expectation on , write for the -agent cost under of an observation-driven control policy with horizon , control dimension and channels for the -th driving system, and let be the law of the random element of furnished by claim 5 of the realized-control lemma. Finally, is the map with constant value on a probability space, a random element of by claim 1 of the lemma on convergence in distribution to a constant, and it is assumed that converges in distribution to .
Write for the set of all -valued observation-driven control policies with horizon , control dimension and channels, and for a natural number put
The set is nonempty, because . For every , part (ii) of the existence, uniqueness and regularity theorem for the controlled -agent dynamics furnishes a solution on for , , the -th driving system and the policy , and claim 2 of the comparison lemma for the -agent system and the mean-field flow, applied to that solution, shows that is a real number with ; so by claim 6 of the absolute-value lemma the number is a lower bound of . Hence the infimum of exists by Existence of the Infimum of a Nonempty Subset of Bounded Below and is unique by Uniqueness of the Supremum and of the Infimum. The optimal -agent value is the real number
Let be a sequence of real numbers with for every , converging to , and such that
Then the following hold.
1. (The optimal values are bounded reals.) For every natural number one has and .
2. (Open-loop policies built from optimal mean-field controls.) The set is nonempty. Let , let be an admissible representative of in the sense of claim 2 of the flow stability lemma, and let be the open-loop policy determined by . Then , and the sequence converges to .
3. (Convergence of the values.) The sequences and both converge to .
4. (Concentration on the optimal controls.) For every natural number the map on defined by is a random variable, and for every real the sequence
converges to .
Loading…
Prerequisites
No prerequisites tracked.
Dependents
No dependents yet.
Dependent proofs
No dependent proofs yet.
No relations recorded yet.