Proof of The Fluctuation LQG Value Is the Asymptotically Optimal Value of the Recentred N-Agent Cost over Admissible Families of Observation-Driven Policies, and Is Attained by the Approximate Kalman Policies, under Deterministic Initial States
theoremthm:lqg-value-is-asymptotically-optimal-2026aThroughout, the certificate lemma means Injection Certificates on the Trimmed Synthetic Copy for the Family of N-Agent Solutions: the Law-Transported Van Trees Certificate Hypothesis Holds under Deterministic Initial States, Uniform Observation Positivity and a C^2 Observation Extension, the LQG lower bound theorem means Asymptotic Lower Bound for the Recentred N-Agent Cost by the Fluctuation LQG Value, under Law-Transported Injection Certificates, the policy lemma means The Approximate Kalman Filter and Policy for the Controlled N-Agent Dynamics, the attainment theorem means Existence of Asymptotically Optimal-Value Observation-Driven Policies, and the measurability lemma means Joint Measurability of the Tracked Events and of the Tracked Energy Density over the Block Cascade, and Measurability in Time of Their Expectations. The item Asymptotic Lower Bound for the Recentred N-Agent Cost by the Fluctuation LQG Value, under Law-Transported Injection Certificates is the earlier published version of the LQG lower bound theorem, the one cited by the certificate lemma; its hypothesis (OC), its Observation data paragraph, its displayed definitions of and and its displayed definition of are given there by the same words and formulas as in the LQG lower bound theorem; the two versions differ elsewhere, in particular in the hypotheses (I0) and (VT) and in the naming of the driving system and the record prefixes. This textual identity of the five items just listed is all that is used below to transfer what the certificate lemma asserts about the earlier version to the LQG lower bound theorem. The fundamental solution written and its inverse in conclusion 3 of the policy lemma play no role below; denotes the realized mean-field flow throughout.
Claim 1.
The two inverses. Under (H1) — which holds with by claim 5 of Localized Joint Coercivity of the Recentred N-Agent Cost Integrand under a Positive-Definite Fluctuation Hessian under (JC), as Ledger Decomposition of the Recentred N-Agent Cost over the Block Cascade and Its Near-Field Filtering Lower Bound records — conclusion (a) of the completion-of-squares theorem gives that every is symmetric positive definite, hence invertible by Invertibility of Symmetric Positive Definite Matrices. Fix . By item 7 of Fluctuation LQG Data of a Stationary Mean-Field Triple, is the diagonal matrix with diagonal entries , which are at least by (OC) because ; the diagonal matrix with diagonal entries is its inverse, as the index formula for the matrix product shows. So and are defined.
The setting of the policy lemma. The policy lemma adopts the setting of Fluctuation LQG Data of a Stationary Mean-Field Triple: the natural numbers , , , a nonempty control set , a transition-rate family with control set and rate bound and its extension with derivative bound , the cost extension of with its second-derivative bound, the horizon , the triple , the observation-rate family with channels and rate bound and its extension with derivative bound . All of these are part of the common data: is the transition-rate family with rate bound furnished by claim 2 of the affine rate family lemma, with and with its bound are the extensions of the common data, is the rate bound of , and with is the extension of (X). The matrices , , , , of the policy lemma are the fluctuation LQG data of these data, that is, the matrices so named in the statement.
The coefficient matrices. Write () and for the fluctuation Hessian coefficients of , and . The policy lemma defines , , and from the Hessian blocks and the terminal matrix of Fluctuation LQG Data of a Stationary Mean-Field Triple, whose entries are , , , and . By the index formula for the transpose, the entries of these four matrices are , , and , which are the defining formulas of the matrices , , , of the completion-of-squares theorem in terms of the same fluctuation Hessian coefficients (that theorem adopts the fluctuation linear-quadratic cost of The Fluctuation Linear-Quadratic Cost Functional for the same extensions and the same triple). So the four coefficient matrices of the policy lemma are those of the common data. Likewise the state matrix has entries and the control matrix has entries , which are the defining formulas and of the completion-of-squares theorem; so and , and the matrix of the policy lemma, formed with the family , is the matrix of the statement.
Hypotheses (H1)--(H4) and (C). Hypothesis (H1) of the policy lemma asks for a real with for all and ; the left-hand side is the entry pairing of the completion-of-squares theorem, so this is hypothesis (H1), which holds with . Hypothesis (H2) of the policy lemma asks for a family of symmetric matrices, continuously differentiable in integral form as in Weighted Second-Moment Evolution of the State Fluctuation Process, with terminal value and densities with ; with the identifications just made this is hypothesis (H2) of the completion-of-squares theorem, which the family of the common data satisfies. So can be taken as the Riccati family of (H2) of the policy lemma. Hypothesis (H3) asks for a real with for all and ; since for every , hypothesis (OC) gives this with . Hypothesis (H4) asks that be symmetric positive semidefinite, which the zero matrix is, for every . Hypothesis (C) asks that be an open subset of . The set is compact in , so by Compact Subset of is Closed (with ) it is closed there, that is, is open in the metric space ; by Euclidean Openness Agrees with Metric Openness on it is then open in the Euclidean sense. So (C) holds. Hence the policy lemma applies to the data named in the statement.
The filter covariance. Conclusion 2 of the policy lemma states that is symmetric positive semidefinite with continuous entries, that every is positive semidefinite, and that consequently Global Existence and Uniqueness for the Kalman Covariance Riccati Equation furnishes exactly one assignment of a matrix with rows and columns to each , with continuous entries, satisfying with the zero matrix — its filter covariance — and that every is symmetric positive semidefinite. This is the Riccati theorem applied on with and the data , , , ; the assignment of the statement is by definition the solution furnished by that theorem for these data, which is unique. Hence the Riccati theorem applies, exists and is unique with continuous entries, every is symmetric positive semidefinite, and is the filter covariance of the policy lemma.
The attainment theorem applies; continuity of the integrand; the value. The attainment theorem adopts the setting, hypotheses (H1)--(H4) and (C) and notation of the policy lemma, with the Riccati family fixed and the matrix of (H4) — here the zero matrix; and for its conclusions 2 and 3 it assumes in addition that is convex, which is part of the common data, and hypothesis (H5) of Mean-Square Error Covariance of the Approximate Kalman Filter as assumed in Cost Limit Along the Approximate Kalman Policy, which is the hypothesis (H5) assumed here with . So the attainment theorem applies. Its setting records that the map , with the matrix pairing , is continuous on ; its being the filter covariance, that is, the present , and its being the present , this map is , the integrand defining , which is therefore continuous on . The number of the attainment theorem is ; since is the zero matrix, , so it is the of the statement. This proves claim 1.
Claim 2. Let a family of solutions be admissible, with the constants , and the points of (I), (CB), (D0).
Hypothesis (I) from (D0). Fix and write , an event of probability one by (D0) (an event because each component of is a random variable by the definition of a solution). The empirical state measure takes its values in by that definition, and any two points satisfy : each , so , whence . Hence every component of is bounded in absolute value by at every point of , so each product is bounded in absolute value by everywhere, and on it equals the constant , itself bounded by since . By claim 2 of Almost Sure Inequalities Between Bounded Random Variables Pass to Expectations (with , the common bound and the constant random variable), and since the expectation of a constant random variable on a probability space is that constant by the definition of the expectation, , whose absolute value is at most (as for ). The last sequence has limit by (D0) and claim 2 of the arithmetic of limits, so by claim 3 of the order properties of limits the sequence has limit for all . This is hypothesis (I) of Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses with the zero matrix as ; the limit of a real sequence being unique (claim 1 of the order properties of limits, applied to a sequence and itself in both directions, gives that two limits of the same sequence are equal), the matrix of hypothesis (I) is the zero matrix, which is hypothesis (I0) of the LQG lower bound theorem. Thus the family satisfies every hypothesis that Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses imposes on a family — (I), (I), (CB) — beyond the common data and the standing hypotheses.
The setting of the certificate lemma. The certificate lemma adopts the setting, notation and standing hypotheses of The Path-Closeness Event under the Cost Bound: Closeness of the Empirical State Measure to the Mean-Field Trajectory and of the Record-Frozen Control to the Mean-Field Control on an Event of Probability , hence of Closeness of the Realized Control and the Realized Mean-Field Flow on a High-Probability Event under the Cost Bound and of Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses: the common data and the standing hypotheses of the latter theorem, which are those adopted in the statement; a family, indexed by , of solutions of the controlled -agent dynamics on for -valued observation-driven control policies on -agent driving systems, with the objects named in the statement — which the given family provides; and the hypotheses of Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses on the family — (I) with the bound and (CB) with the bound , which hold by admissibility, and (I), which has just been verified directly, so that the appeal to the certificate lemma is not circular on any reading of its adopted setting (its claim 1 re-derives (I) with the zero matrix from (D0)). The remaining objects of that setting are auxiliary and exist: the constants of Closeness of the Realized Control and the Realized Mean-Field Flow on a High-Probability Event under the Cost Bound and of The Path-Closeness Event under the Cost Bound: Closeness of the Empirical State Measure to the Mean-Field Trajectory and of the Record-Frozen Control to the Mean-Field Control on an Event of Probability are furnished by their claims; a natural number exists; the realized controls, records, record spaces, prefix maps, record prefixes, truncated policies and observation-centred fluctuations are defined from the family; a dense sequence in exists for every , as the certificate lemma's setting records; and reconstruction data exist by Measurable Reconstruction of the Controlled N-Agent Dynamics from Observation Records, as that setting records. Fix such choices. Finally its hypotheses (D0), (OC) and (X) are the hypothesis (D0) of admissibility (with the points ) and the hypotheses (OC) and (X) of the statement. Hence the certificate lemma applies to the given family, and its claims 1 and 2 hold for it.
Hypothesis (OC) of the LQG lower bound theorem. By claim 1 of the certificate lemma, hypothesis (OC) of Asymptotic Lower Bound for the Recentred N-Agent Cost by the Fluctuation LQG Value, under Law-Transported Injection Certificates holds with for the triple and the observation extension of (X). Hypothesis (OC) of the LQG lower bound theorem is stated in the same words for the same objects — continuity on of every entry of and of every , and a uniform positive lower bound on the — so it holds with .
The profile responses and the information functionals. The LQG lower bound theorem defines, for a profile , its as the unique assignment with continuous components satisfying and its as , with the matrix of hypothesis (I) — here the zero matrix — and with the matrices , and of the fluctuation LQG data relative to the extensions of the statement. These are the displayed formulas by which Asymptotic Lower Bound for the Recentred N-Agent Cost by the Fluctuation LQG Value, under Law-Transported Injection Certificates defines its and , evaluated with the zero matrix as and with the observation extension ; by claim 1 of the certificate lemma these coincide with the profile response and the information functional of the certificate lemma.
Hypothesis (VT). The objects named in hypothesis (VT) of the LQG lower bound theorem — the driving system , the record , the record spaces , the prefix maps , the record prefixes , the observation filtration , and the observation-centred fluctuation with the realized mean-field flow , formed from the realized control and the two-argument mean-field flow with the base point — are the objects so named in the setting of the certificate lemma, which adopts them from the same sources (its being, as it records, the observation-centred fluctuation of Asymptotic Lower Bound for the Recentred N-Agent Cost by the Fluctuation LQG Value, under Law-Transported Injection Certificates formed with the base point , and the LQG lower bound theorem defining by the same formula from the same realized flow). Claim 2 of the certificate lemma asserts, for these objects and for its and , exactly the statement (VT) of the LQG lower bound theorem — the condition (C0) for every and , and, for every , every vector of (written there and here), every profile and every , the existence of and of the certificate data with (a), (C1), (C2), (C3), (C4), stated in the same words. Since the and of the two items coincide, hypothesis (VT) holds.
Hypothesis (MS). Fix , an admissible parameter vector in the sense of Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses and a block index , being the number of blocks determined by (written in the ledger lemma; the letter is the derivative bound here). By the definition of admissibility of , when the block cascade and the near-field data are formed from for the -th solution, every parameter-dependent requirement of Ledger Decomposition of the Recentred N-Agent Cost over the Block Cascade and Its Near-Field Filtering Lower Bound holds; its remaining requirements are the common data, the standing hypotheses, and the -th solution with its data, all in force. So the setting of the ledger lemma, and with it that of the measurability lemma, is instantiated by the common data, the -th solution and , and claim 2 of the measurability lemma gives that the three functions , and , restricted to , are measurable with respect to the trace Borel -algebra of that block. This is hypothesis (MS) of the LQG lower bound theorem for the family.
The remaining hypotheses; the theorem applies. The common data, the family of solutions, the standing hypotheses, hypotheses (I) and (CB), the admissible parameter vectors and the number are adopted by the LQG lower bound theorem from Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses; hypothesis (I) holds as shown; hypothesis (H2) of the completion-of-squares theorem, with the Riccati family , is part of the common data; hypothesis (H1) holds with ; the observation data of the LQG lower bound theorem are the extension of (X) and the fluctuation LQG data relative to it; and (I0), (OC), (VT) and (MS) hold as shown. Hence every hypothesis of the LQG lower bound theorem holds for the common data and the family. Its is the solution of the Riccati theorem on with the data , , and initial value the matrix of hypothesis (I), that is, the zero matrix; this is the of the statement. Its matrices , adopted from the cascade filtering lemma, are formed from the same and as the of the statement, and so coincide with them.
Conclusions. By the paragraph The limit value of Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses, each product is continuous on with finite Lebesgue integral, hence so is the finite sum by the continuity of sums and the linearity of the integral, and
the first sum vanishing because is the zero matrix. Claim 1 of the LQG lower bound theorem gives that is continuous, nonnegative and Lebesgue integrable on . Hence, by the linearity of the integral (both summands being integrable),
The recentred cost of the LQG lower bound theorem is the recentred cost of First-Order Expansion of the Recentred N-Agent Cost about a Stationary Mean-Field Triple and Its Coercive Lower Bound adopted from Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses, which is the quantity displayed in the statement. Claim 4 of the LQG lower bound theorem therefore gives: for every real there is a natural number with for every ; take . This proves claim 2.
Claim 3.
The Kalman family is a family of solutions. By conclusion 1 of the attainment theorem, available by claim 1, each is an -valued observation-driven control policy with horizon , control dimension and channels, and the projected solution fixed in the statement is a solution of the controlled -agent dynamics on for , , the driving system and . So the Kalman family is a family of solutions for the common data, with , .
The initial conditions. Fix and write , an event of probability one by (DK) (it is an event, each component of being a random variable by the definition of a solution). The empirical state measure takes its values in by that definition, and any two points satisfy : each , so , whence . Hence , and every component of is bounded in absolute value by , at every point of ; so the random variables and are bounded in absolute value by and by respectively, everywhere on . On one has , so there and , constants which are themselves bounded in absolute value by and by , since and lie in . By claim 2 of Almost Sure Inequalities Between Bounded Random Variables Pass to Expectations (applied with , the common bounds , respectively , and the constant random variables),
the expectation of a constant random variable on a probability space being that constant by the definition of the expectation; these are the two identities of claim 3. Put , a sequence with limit by (DK). Since for (each ), the first expectation is bounded in absolute value by , which has limit by claim 2 of the arithmetic of limits; by claim 3 of the order properties of limits the sequence has limit , which is hypothesis (I1) of Cost Limit Along the Approximate Kalman Policy with the zero matrix as . The second expectation is , which has limit by the same arithmetic of limits; by clause (d) of claim 3 of Real Powers Through the Exponential, and Elementary Asymptotic Tools: Monotonicity, Null Sequences of Negative Powers, Exponential Domination, Integer Rounding, and Square-Root and Exponential Inequalities there is with for every , so with one has , which is hypothesis (I2).
Attainment. Conclusion 2 of the attainment theorem, applied to the driving systems and projected solutions fixed in the statement, whose empirical state measures are and whose state fluctuations are , states that each , with and with the -agent cost of the solution at level and the mean-field cost of — that is, each — is a well-defined real number and that the sequence has limit the of that theorem, which is the present by claim 1.
Admissibility. (I): for every , with as above, which is (I) with the bound . (CB): by the definition of the limit there is with , hence , for every ; put ; then for every , which is (CB) with the bound . (D0): the points satisfy and by (DK). Hence the Kalman family is admissible, and . This proves claim 3.
Claim 4.
is the asymptotically optimal value. Condition (i) of Asymptotically Optimal Value of a Set of Real Sequences for : every sequence in is the sequence of an admissible family, so by claim 2, for every there is with for every . Condition (ii): by claim 3 the sequence belongs to and converges to . So is an asymptotically optimal value of , which is nonempty by claim 3.
Uniqueness. Let a real number be an asymptotically optimal value of . Suppose and put . By condition (ii) for there is a sequence in converging to , so there is with , hence , for every ; by condition (i) for there is with for every ; at these contradict each other. So . Suppose and put . By condition (ii) for (claim 3) the sequence in converges to , so there is with for every ; by condition (i) for there is with for every ; again a contradiction at . So , and . Hence is the unique asymptotically optimal value of , as claimed.
The LQG minimum. Conclusion 3 of the attainment theorem, available by claim 1, states that for every linear-Gaussian state-observation model matched to the fluctuation LQG data as described there, the linear-quadratic-Gaussian cost formed there attains a minimum over the extended admissible controls with values in and , with the of that theorem, which is the present . This proves claim 4.
Loading…
Prerequisites
bd95fb33-7c44-4d23-b798-a48922031ca4