Existence of Asymptotically Optimal-Value Observation-Driven Policies
theoremProbabilitythm:asymptotically-optimal-value-policies-2026bAdopt the setting, hypotheses (H1)--(H4) and (C), and notation of the approximate Kalman filter and policy lemma, for the fluctuation LQG data of the stationary mean-field triple , whose stationary co-state has value at time : in particular the natural numbers , , , the control set (a nonempty subset of ), the transition-rate family , with control set , and the observation-rate family , together with their extensions, the cost extension of the population cost data , the horizon , the matrices , , , , of the fluctuation LQG data, the symmetrized coefficient matrices , , , --- each symmetric positive definite with continuous inverse by conclusion 1 of the policy lemma --- the Riccati family of (H2), fixed throughout, with , the matrix of (H4), and the filter covariance of conclusion 2 of the policy lemma. Throughout, a real-valued function on a subinterval of the real numbers is called continuous on when it is continuous relative to , both and the codomain carrying the metric of the real line. For conclusion 2 below, which invokes the cost-limit proposition, and for the final sentence of conclusion 3, which combines conclusion 2 with the minimum identity, assume in addition the two hypotheses that the proposition carries beyond (H1)--(H4) and (C): that the control set is convex, and hypothesis (H5) there, namely that there is a real such that every with lies in , for every (as that proposition records, these two together with (C) and the boundedness of the control-side open set of the transition-rate extension --- the letter here being unrelated to the matrices and to the number defined below --- make a closed, bounded, convex subset of ). The definition of , conclusion 1, and all of conclusion 3 except its final sentence do not use these two hypotheses. All entries of the maps (part of (H2)), , , , and (conclusion 1 of the policy lemma), and (conclusion 2 there) are continuous on ; hence, with the matrix pairing of the cost-limit proposition, the map is continuous on by sums and products of continuous functions. Define the real number
the integral being the Lebesgue integral over the compact interval of this continuous integrand, which agrees with its Riemann integral by claim 3 of that interval toolkit. Then:
1. (Existence of the policies.) For each natural number the approximate Kalman policy of conclusion 3 of the policy lemma is an observation-driven control policy with horizon , control dimension , and channels, and it is -valued; and for every -agent driving system with states and observation channels there is a projected solution of the controlled -agent dynamics on for , , that driving system, and , as in conclusion 4(a) of the policy lemma --- the state processes, observation processes, and regular event of a solution for the modified family and policy of conclusion 3 there, with the control process replaced by the vector of its first components --- and every solution for , , that driving system, and is indistinguishable from it in the sense of the uniqueness theorem.
2. (Asymptotic value.) Suppose that for each such a driving system and projected solution are fixed, with empirical state measure as in the solution definition and state fluctuation process , and that the initial-condition hypotheses (I1)--(I2) of the cost-limit proposition hold for these solutions: as for all , and , with the expectation. Let be the -agent cost of the solution at level , let be the mean-field cost of , and set with the componentwise expectation, as in the cost-limit proposition. Then each is a well-defined real number and
3. (Optimality of the value.) For every linear-Gaussian state-observation model on with state dimension , observation dimension , and any Brownian dimension , its data written as in the identification corollary and matched to the fluctuation LQG data as in hypotheses (i)--(iii) there --- a class of coefficient data that is nonempty, by conclusion 3 of that corollary, already for Brownian dimension --- with the control dimension of the controlled system taken to be , the control matrix assignment sending each to , and the cost data , (entrywise), , , all as in the corollary, the linear-quadratic-Gaussian cost attains a minimum over the extended admissible controls with values in , and
In particular, under the hypotheses of conclusion 2 --- the driving systems and projected solutions fixed there, (I1)--(I2), and the two additional hypotheses --- the recentred limiting cost of conclusion 2 along the approximate Kalman policies equals this minimal LQG cost.
Loading…
Prerequisites
No prerequisites tracked.
Dependents
No dependents yet.
Dependent proofs
No dependent proofs yet.
No relations recorded yet.