Proof of Existence of Asymptotically Optimal-Value Observation-Driven Policies
theoremthm:asymptotically-optimal-value-policies-2026bWell-definedness of . All entries of are continuous on as part of hypothesis (H2), continuity throughout this proof being understood in the sense fixed in the statement; all entries of , , , and are continuous by conclusion 1 of the approximate Kalman filter and policy lemma, and all entries of are continuous by conclusion 2 there. The entries of and the map are built from these by finitely many sums and products of continuous functions, hence are continuous on ; the Lebesgue integral over the compact interval of the latter map therefore exists and agrees with its Riemann integral by claim 3 of that interval toolkit, so is a well-defined real number.
Conclusion 1. By conclusion 3 of the approximate Kalman filter and policy lemma, is an observation-driven control policy with horizon , control dimension , and channels, and is -valued; the family defined there is a transition-rate family on states with control set and rate bound ; and the family defined there is an observation-driven control policy with horizon , control dimension , and channels, and is -valued. Given an -agent driving system with states and observation channels, solutions of the controlled -agent dynamics on for , , that driving system, and exist by the existence and uniqueness theorem for the controlled -agent dynamics, as recorded in conclusion 4 of the policy lemma. By conclusion 4(a) of the policy lemma, the projection of such a solution --- the same state and observation processes and regular event, with the control process replaced by the vector of its first components --- is a solution of the controlled -agent dynamics on for , , the same driving system, and , and every solution for these data is indistinguishable from it in the sense of the uniqueness theorem. This proves conclusion 1.
Conclusion 2. Fix, for each , a driving system and a projected solution as in conclusion 1, with empirical state measure and state fluctuation process , and assume (I1)--(I2). The hypotheses of the cost-limit proposition then hold: its setting, hypotheses (H1)--(H4) and (C), and notation are those adopted here, its two further hypotheses --- convexity of the control set and (H5) --- are assumed in the statement for this conclusion, its per-level data --- the driving system, the projected solution as in conclusion 4 of the policy lemma, the -agent cost , the mean-field cost , and the vector --- are those fixed in the statement, and its initial-condition hypotheses (I1)--(I2), spelled out in the statement of conclusion 2, are assumed. The proposition asserts that each is a well-defined real number and that
and the right-hand side is by definition. This proves conclusion 2.
Conclusion 3. Let a linear-Gaussian state-observation model on with state dimension , observation dimension , and Brownian dimension , matched to the fluctuation LQG data as in hypotheses (i)--(iii) of the identification corollary, be given; take the control dimension of the controlled system to be , the control matrix assignment sending each to , and the cost data , (entrywise), , --- exactly the choices made in the corollary. By conclusion 2 of the corollary --- the minimum identity being the part of conclusion 2 preceding its final assertion, the corollary scoping convexity of and (H5) to that final assertion alone --- the minimum of the linear-quadratic-Gaussian cost over the extended admissible controls with values in exists (it is attained by the closed-loop feedback control of the closed-loop feedback lemma) and
By conclusion 3 of the corollary, the class of matched coefficient data is nonempty --- already for Brownian dimension . The final sentence of conclusion 3 follows by combining the displayed identity with the limit of conclusion 2. This proves conclusion 3.
Loadingβ¦
Prerequisites
b83e3b76-99b3-4c6e-8b21-cc6b52acd68f