Proof of Existence of Asymptotically Optimal-Value Observation-Driven Policies
theoremthm:asymptotically-optimal-value-policies-2026aWell-definedness of . All entries of are continuous on as part of hypothesis (H2); all entries of , , , and are continuous by conclusion 1 of the approximate Kalman filter and policy lemma, and all entries of are continuous by conclusion 2 there. The entries of and the map are built from these by finitely many sums and products of continuous functions, hence are continuous on ; the Lebesgue integral over the compact interval of the latter map therefore exists and agrees with its Riemann integral, so is a well-defined real number.
Conclusion 1. By conclusion 3 of the approximate Kalman filter and policy lemma, is an observation-driven control policy with horizon , control dimension , and channels, the family defined there is a transition-rate family on states with control dimension and rate bound , and the family defined there is an observation-driven control policy with horizon , control dimension , and channels. Given an -agent driving system with states and observation channels, solutions of the controlled -agent dynamics on for , , that driving system, and exist by the existence and uniqueness theorem for the controlled -agent dynamics, as recorded in conclusion 4 of the policy lemma. By conclusion 4(a) of the policy lemma, the projection of such a solution --- the same state and observation processes and regular event, with the control process replaced by the vector of its first components --- is a solution of the controlled -agent dynamics on for , , the same driving system, and , and every solution for these data is indistinguishable from it in the sense of the uniqueness theorem. This proves conclusion 1.
Conclusion 2. Fix, for each , a driving system and a projected solution as in conclusion 1, with empirical state measure and state fluctuation process , and assume (I1)--(I2). The hypotheses of the cost-limit proposition then hold: its setting, hypotheses (H1)--(H4), and notation are those adopted here, its per-level data --- the driving system, the projected solution as in conclusion 4 of the policy lemma, the -agent cost , the mean-field cost , and the vector --- are those fixed in the statement, and its initial-condition hypotheses (I1)--(I2), spelled out in the statement of conclusion 2, are assumed. The proposition asserts that each is a well-defined real number and that
and the right-hand side is by definition. This proves conclusion 2.
Conclusion 3. Let a linear-Gaussian state-observation model on with state dimension , observation dimension , and Brownian dimension , matched to the fluctuation LQG data as in hypotheses (i)--(iii) of the identification corollary, be given; take the control dimension of the controlled system to be , the control matrix assignment sending each to , and the cost data , (entrywise), , --- exactly the choices made in the corollary. By conclusion 2 of the corollary, the minimum of the linear-quadratic-Gaussian cost over the extended admissible controls with values in exists (it is attained by the closed-loop feedback control of the closed-loop feedback lemma) and
By conclusion 3 of the corollary, the class of matched coefficient data is nonempty. The final sentence of conclusion 3 follows by combining the displayed identity with the limit of conclusion 2. This proves conclusion 3.
Loading…
Prerequisites
3c48547b-1d4c-4534-a8b4-fcc21b423676