The Fluctuation LQG Value Is the Asymptotically Optimal Value of the Recentred N-Agent Cost over Admissible Families of Observation-Driven Policies, and Is Attained by the Approximate Kalman Policies, under Deterministic Initial States
theoremAnalysisProbabilitythm:lqg-value-is-asymptotically-optimal-2026aCommon data. Adopt the common data and the standing hypotheses of Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses, and the notation of Injection Certificates on the Trimmed Synthetic Copy for the Family of N-Agent Solutions: the Law-Transported Van Trees Certificate Hypothesis Holds under Deterministic Initial States, Uniform Observation Positivity and a C^2 Observation Extension for them: the natural numbers , , ; the affine-controlled transition-rate family on states with nonempty convex control set , compact for the topology of the Euclidean distance (the collection of subsets of that are open in , a topology by Metric Open Sets Form a Topology; this is the topology with respect to which compactness of the control set is understood in Affine-Controlled Transition-Rate Family and throughout the adopted setting), with control bound , and its transition-rate family with rate bound ; the twice continuously differentiable extension of with derivative bound and its extended aggregate state drift ; the observation-rate family with channels and rate bound and its aggregate observation drift ; the horizon ; the population cost data , convex in the control on , with their twice continuously differentiable extension, written here (its open set is written in Fluctuation LQG Data of a Stationary Mean-Field Triple and in The Approximate Kalman Filter and Policy for the Controlled N-Agent Dynamics; the letter denotes the observation record below); the stationary mean-field triple for these data with initial point in the probability simplex , whose co-state has the components at time ; the aggregate fluctuation covariance of with ; the matrices , , , , and of the completion-of-squares theorem ( and formed from the partial derivatives of along , and , , , from the fluctuation Hessian coefficients of , and ), its entry pairing , and its hypothesis (H2) with the Riccati family , part of the common data; and the standing hypotheses of Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses — hypotheses (A) and (U) of Quadratic Growth of the Mean-Field Hamiltonian in the Control along a Stationary Mean-Field Triple, (JC) of Localized Joint Coercivity of the Recentred N-Agent Cost Integrand under a Positive-Definite Fluctuation Hessian with constant , (H2), (LipC) of Pointwise-in-Time Tracking of the Mean-Field Flow and Cost along the Realized Control of the Controlled N-Agent Dynamics, the optimality and (TG) of Quadratic Expansion Bounds for the Mean-Field To-Go Value along a Stationary Mean-Field Triple, and the standing hypothesis on of The Block Cascade of Anchored Good-Set Clocks: Adapted Good Sets, Matched Escape Bounds, and the Energy Ledger — each of which, on inspection of the cited items, is a condition on the common data and the triple only, so that it is satisfied by every family of solutions for these data (the proof of claim 2 uses only this). Assume in addition:
(OC) is a real number with for all and all (hypothesis (OC) of Injection Certificates on the Trimmed Synthetic Copy for the Family of N-Agent Solutions: the Law-Transported Van Trees Certificate Hypothesis Holds under Deterministic Initial States, Uniform Observation Positivity and a C^2 Observation Extension);
(X) is a twice continuously differentiable extension of with derivative bound ;
(H5) there is a real such that every with lies in , for every (hypothesis (H5) of Mean-Square Error Covariance of the Approximate Kalman Filter, assumed in Cost Limit Along the Approximate Kalman Policy; the subscript distinguishes from the near-field radius and the certificate measure of the cited settings).
Hypothesis (H1) of the completion-of-squares theorem — there is a real with for every and — holds with under (JC), as Ledger Decomposition of the Recentred N-Agent Cost over the Block Cascade and Its Near-Field Filtering Lower Bound records.
The fluctuation LQG data and the value. Let , , , and () be the state matrix, control matrix, observation matrix, state noise covariance and observation noise covariance of the fluctuation LQG data of relative to the extensions , and (the state noise covariance is the matrix already named, the state matrix coincides with and the control matrix with , by the identical defining formulas), and put, for ,
with the matrix product, the transpose and the matrix inverse ( is invertible under (H1), and under (OC); claim 1 below records both). Let be the solution of the Kalman covariance Riccati equation on with the data , , and the zero matrix as initial value (the datum written in that theorem; claim 1 below records that the theorem applies), and define the real number
the Lebesgue integral over the compact interval of an integrand that claim 1 records to be continuous. Write for the Euclidean norm and for the limit of a real sequence.
Families of solutions; admissibility. A family of solutions for the common data consists of, for every natural number : an -agent driving system with states and observation channels, with expectation ; an -valued observation-driven control policy with horizon , control dimension and channels; and a solution of the controlled -agent dynamics on for , , that driving system and , with regular event , empirical state measure , observation record , observation filtration , state fluctuation , , and recentred cost
of First-Order Expansion of the Recentred N-Agent Cost about a Stationary Mean-Field Triple and Its Coercive Lower Bound, formed with the -agent cost of the -th solution (finite by conclusion (b) of that lemma under (A)) and the mean-field cost of ; the dependence of these objects on is suppressed. Such a family is called admissible if it satisfies the hypotheses that Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses and Injection Certificates on the Trimmed Synthetic Copy for the Family of N-Agent Solutions: the Law-Transported Van Trees Certificate Hypothesis Holds under Deterministic Initial States, Uniform Observation Positivity and a C^2 Observation Extension impose on a family beyond the common data and beyond hypothesis (I) of the former, which claim 2 below derives from (D0):
(I) there is a real with for every ;
(CB) there is a real with for every ;
(D0) there are points in the aggregate lattice with for every , such that the real sequence has limit .
(Claim 2 below records that under (D0) hypothesis (I) of Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses holds with the zero matrix as , so that an admissible family satisfies every hypothesis that theorem and the certificate lemma impose on a family.) Write for the set of all sequences of recentred costs of admissible families, a set of real sequences (claim 3 below records that is nonempty).
The approximate Kalman family. For every natural number let (written in the cited items) be the approximate Kalman policy at level of the data , , , , , , , , , , , with the Riccati family and with the zero matrix in the role of the matrix of its hypothesis (H4) (claim 1 below records that the lemma applies and that its filter covariance is ); let an -agent driving system with states and observation channels, with expectation , be given, and let a projected solution of the controlled -agent dynamics on for , , this driving system and , as in conclusion 1 of Existence of Asymptotically Optimal-Value Observation-Driven Policies, be fixed, with empirical state measure , state fluctuation and recentred cost (the dependence on again suppressed; the objects named for a family of solutions above are written with the superscript for this one); claim 3 below records that these objects form a family of solutions in the sense above, the Kalman family. Assume:
(DK) there are points with for every such that the real sequence has limit .
Notational cautions: the matrix carries a time subscript, the observation record does not; , are probability measures, with a time subscript is the co-state, and in the instantiation of the Riccati theorem is the name of its initial-value datum; in that instantiation , , are the names of the data of the Riccati theorem, whereas with a time subscript is the mean-field control of the triple; is the control set, whereas with subscript and argument is an information functional of the cited items; is the derivative bound of , whereas the block count of the ledger lemma, written there, is not named here; with a time subscript is a coefficient matrix, whereas the noise majorant written in the cited settings is not named here; is the rate bound and , are control matrices; is the control bound and a coefficient matrix; and are the state matrix, whereas the path set written and the energy written in the cited settings are not named here; the roman superscript marks the objects of the Kalman family and is unrelated to ; is the observation filtration, whereas the feedback gain of The Approximate Kalman Filter and Policy for the Controlled N-Agent Dynamics is not named here; the letter in is an open set, unrelated to the matrices and to ; is a set of real sequences, unrelated to the Riccati datum ; denotes the realized mean-field flow of the cited settings, the fundamental solution written (with inverse ) in The Approximate Kalman Filter and Policy for the Controlled N-Agent Dynamics not being named here; and (DK) is a hypothesis on the Kalman family, the existence of driving systems with prescribed deterministic initial aggregate states not being asserted here.
Then the following hold.
1. (The Kalman-policy lemma and the attainment theorem apply; the value is well defined.) Each and each is invertible. The hypotheses (H1)--(H4) and (C) of The Approximate Kalman Filter and Policy for the Controlled N-Agent Dynamics hold for the data named above, with the zero matrix as and with in (H3); its coefficient matrices , , , are the matrices so named above, its state and control matrices are and , its Riccati family can be taken to be , whose matrices are those defined above, and its filter covariance is : the Riccati theorem applies, exists and is unique with continuous entries, and every is symmetric positive semidefinite. The control set is convex and hypothesis (H5) of Cost Limit Along the Approximate Kalman Policy holds with , so Existence of Asymptotically Optimal-Value Observation-Driven Policies applies to these data; the integrand defining is continuous on , and the real number defined in that theorem equals the defined above.
2. (Every admissible family satisfies the LQG lower bound.) Let a family of solutions be admissible. Then hypothesis (I) of Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses holds for it with the zero matrix as ; the setting and hypotheses of Injection Certificates on the Trimmed Synthetic Copy for the Family of N-Agent Solutions: the Law-Transported Van Trees Certificate Hypothesis Holds under Deterministic Initial States, Uniform Observation Positivity and a C^2 Observation Extension are instantiated by the common data, (OC), (X) and this family, once the auxiliary objects fixed in that setting (among them the natural number , the dense sequences and the reconstruction data) are chosen, which is possible; hypothesis (MS) of Asymptotic Lower Bound for the Recentred N-Agent Cost by the Fluctuation LQG Value, under Law-Transported Injection Certificates holds for this family; every hypothesis of Asymptotic Lower Bound for the Recentred N-Agent Cost by the Fluctuation LQG Value, under Law-Transported Injection Certificates holds for the common data and this family, its being the above and its real number being ; and consequently, for every real there is a natural number with
3. (The Kalman family is admissible and attains the value.) For every , for all and , so that, with the limit assumed in (DK), the initial-condition hypotheses (I1) (with the zero matrix as ) and (I2) of Cost Limit Along the Approximate Kalman Policy hold for the Kalman family; each is a well-defined real number and
Moreover the Kalman family satisfies (I), (CB) and (D0) — the last with — and is therefore admissible; in particular is nonempty and contains .
4. (The fluctuation LQG value is the asymptotically optimal value.) is an asymptotically optimal value of , and it is the only real number with this property: if a real number is an asymptotically optimal value of then . Thus is the asymptotically optimal value of the recentred -agent cost over admissible families. Moreover is the minimal linear-quadratic-Gaussian cost of the fluctuation problem: for every linear-Gaussian state-observation model on with state dimension , observation dimension and any Brownian dimension , matched to the fluctuation LQG data as in conclusion 3 of Existence of Asymptotically Optimal-Value Observation-Driven Policies (through hypotheses (i)--(iii) of The Limiting Cost Along the Approximate Kalman Policy Is the Optimal Value of the Fluctuation LQG Problem, with control dimension , control matrix assignment and cost data , , , ), the linear-quadratic-Gaussian cost formed there attains a minimum over the extended admissible controls with values in , and .
Loading…
Prerequisites
No prerequisites tracked.
Dependents
No dependents yet.
Dependent proofs
No dependent proofs yet.
No relations recorded yet.