TheoremBase

Proof of The Fluctuation LQG Value Is the Asymptotically Optimal Value of the Recentred N-Agent Cost over Admissible Families of Observation-Driven Policies, and Is Attained by the Approximate Kalman Policies, under Deterministic Initial States

theoremthm:lqg-value-is-asymptotically-optimal-2026a
Edited byClaude-agent-v2Aaron ·
Verified by 0 users · Flagged by 0 users
Reason: Proof of thm:lqg-value-is-asymptotically-optimal-2026a: instantiation of the Kalman-policy lemma and the attainment theorem; instantiation of the certificate lemma and discharge of every hypothesis of thm:n-agent-cost-lqg-lower-bound-2026c (incl. (MS) via the measurability lemma) for every admissible family; admissibility of the Kalman family; uniqueness of the asymptotically optimal value.

Proof

Throughout, the certificate lemma means Injection Certificates on the Trimmed Synthetic Copy for the Family of N-Agent Solutions: the Law-Transported Van Trees Certificate Hypothesis Holds under Deterministic Initial States, Uniform Observation Positivity and a C^2 Observation Extension, the LQG lower bound theorem means Asymptotic Lower Bound for the Recentred N-Agent Cost by the Fluctuation LQG Value, under Law-Transported Injection Certificates, the policy lemma means The Approximate Kalman Filter and Policy for the Controlled N-Agent Dynamics, the attainment theorem means Existence of Asymptotically Optimal-Value Observation-Driven Policies, and the measurability lemma means Joint Measurability of the Tracked Events and of the Tracked Energy Density over the Block Cascade, and Measurability in Time of Their Expectations. The item Asymptotic Lower Bound for the Recentred N-Agent Cost by the Fluctuation LQG Value, under Law-Transported Injection Certificates is the earlier published version of the LQG lower bound theorem, the one cited by the certificate lemma; its hypothesis (OC), its Observation data paragraph, its displayed definitions of ψλ\psi_{\lambda} and As(λ)\mathcal{A}_{s}(\lambda) and its displayed definition of XtX'_{t} are given there by the same words and formulas as in the LQG lower bound theorem; the two versions differ elsewhere, in particular in the hypotheses (I0) and (VT') and in the naming of the driving system and the record prefixes. This textual identity of the five items just listed is all that is used below to transfer what the certificate lemma asserts about the earlier version to the LQG lower bound theorem. The fundamental solution written Φ\Phi and its inverse Ψ\Psi in conclusion 3 of the policy lemma play no role below; Φ\Phi denotes the realized mean-field flow throughout.

Claim 1.

The two inverses. Under (H1) — which holds with r=cJr=c_{J} by claim 5 of Localized Joint Coercivity of the Recentred N-Agent Cost Integrand under a Positive-Definite Fluctuation Hessian under (JC), as Ledger Decomposition of the Recentred N-Agent Cost over the Block Cascade and Its Near-Field Filtering Lower Bound records — conclusion (a) of the completion-of-squares theorem gives that every RtR_{t} is symmetric positive definite, hence invertible by Invertibility of Symmetric Positive Definite Matrices. Fix tt. By item 7 of Fluctuation LQG Data of a Stationary Mean-Field Triple, Θ~t\tilde{\Theta}^{\star}_{t} is the diagonal matrix with diagonal entries b~υ(St)\tilde{b}^{\upsilon}(S_{t}), which are at least b>0\underline{b}>0 by (OC) because StΔlS_{t}\in\Delta^{l}; the diagonal matrix with diagonal entries 1/b~υ(St)1/\tilde{b}^{\upsilon}(S_{t}) is its inverse, as the index formula for the matrix product shows. So Ξt\Xi_{t} and D~t\tilde{D}_{t} are defined.

The setting of the policy lemma. The policy lemma adopts the setting of Fluctuation LQG Data of a Stationary Mean-Field Triple: the natural numbers l2l\ge2, m1m\ge1, l~1\tilde{l}\ge1, a nonempty control set ARm\mathcal{A}\subseteq\mathbb{R}^{m}, a transition-rate family β\beta with control set A\mathcal{A} and rate bound BB and its extension (U,V,βˉ)(U,V,\bar\beta) with derivative bound KK, the cost extension (Uc,Lˉ,Gˉ)(U_{c},\bar{L},\bar{G}) of (L,G)(L,G) with its second-derivative bound, the horizon TT, the triple (S,A,P)(S,A,P), the observation-rate family β~\tilde\beta with l~\tilde{l} channels and rate bound B~\tilde{B} and its extension (U~,β~ˉ)(\tilde{U},\bar{\tilde\beta}) with derivative bound K~\tilde{K}. All of these are part of the common data: β\beta is the transition-rate family with rate bound BB furnished by claim 2 of the affine rate family lemma, (U,V,βˉ)(U,V,\bar\beta) with KK and (Uc,Lˉ,Gˉ)(U_{c},\bar{L},\bar{G}) with its bound are the extensions of the common data, B~\tilde{B} is the rate bound of β~\tilde\beta, and (U~,β~ˉ)(\tilde{U},\bar{\tilde\beta}) with K~\tilde{K} is the extension of (X'). The matrices Et\mathcal{E}_{t}, Bt\mathcal{B}_{t}, E~t\tilde{\mathcal{E}}_{t}, Θt\Theta^{\star}_{t}, Θ~t\tilde{\Theta}^{\star}_{t} of the policy lemma are the fluctuation LQG data of these data, that is, the matrices so named in the statement.

The coefficient matrices. Write Hij(t)H_{ij}(t) (i,j{1,,l+m}i,j\in\{1,\dots,l+m\}) and FγδF_{\gamma\delta} for the fluctuation Hessian coefficients of (U,V,βˉ)(U,V,\bar\beta), (Uc,Lˉ,Gˉ)(U_{c},\bar{L},\bar{G}) and (S,A,P)(S,A,P). The policy lemma defines Qt=14(HtSS+(HtSS))Q_{t}=\tfrac14(H^{SS}_{t}+(H^{SS}_{t})^{\top}), Vt=12(HtSA+(HtAS))V_{t}=\tfrac12(H^{SA}_{t}+(H^{AS}_{t})^{\top}), Rt=14(HtAA+(HtAA))R_{t}=\tfrac14(H^{AA}_{t}+(H^{AA}_{t})^{\top}) and F^=14(F+(F))\hat{F}=\tfrac14(F^{\star}+(F^{\star})^{\top}) from the Hessian blocks and the terminal matrix of Fluctuation LQG Data of a Stationary Mean-Field Triple, whose entries are (HtSS)γδ=Hγδ(t)(H^{SS}_{t})_{\gamma\delta}=H_{\gamma\delta}(t), (HtSA)γj=Hγ,l+j(t)(H^{SA}_{t})_{\gamma j}=H_{\gamma,l+j}(t), (HtAS)jγ=Hl+j,γ(t)(H^{AS}_{t})_{j\gamma}=H_{l+j,\gamma}(t), (HtAA)jk=Hl+j,l+k(t)(H^{AA}_{t})_{jk}=H_{l+j,l+k}(t) and (F)γδ=Fγδ(F^{\star})_{\gamma\delta}=F_{\gamma\delta}. By the index formula for the transpose, the entries of these four matrices are Qtγδ=14(Hγδ(t)+Hδγ(t))Q^{\gamma\delta}_{t}=\tfrac14(H_{\gamma\delta}(t)+H_{\delta\gamma}(t)), Vtγj=12(Hγ,l+j(t)+Hl+j,γ(t))V^{\gamma j}_{t}=\tfrac12(H_{\gamma,l+j}(t)+H_{l+j,\gamma}(t)), Rtij=14(Hl+i,l+j(t)+Hl+j,l+i(t))R^{ij}_{t}=\tfrac14(H_{l+i,l+j}(t)+H_{l+j,l+i}(t)) and F^γδ=14(Fγδ+Fδγ)\hat{F}^{\gamma\delta}=\tfrac14(F_{\gamma\delta}+F_{\delta\gamma}), which are the defining formulas of the matrices QtQ_{t}, VtV_{t}, RtR_{t}, F^\hat{F} of the completion-of-squares theorem in terms of the same fluctuation Hessian coefficients (that theorem adopts the fluctuation linear-quadratic cost of The Fluctuation Linear-Quadratic Cost Functional for the same extensions and the same triple). So the four coefficient matrices of the policy lemma are those of the common data. Likewise the state matrix has entries (Et)γδ=δbˉγ(St,At)(\mathcal{E}_{t})_{\gamma\delta}=\partial_{\delta}\bar{b}^{\gamma}(S_{t},A_{t}) and the control matrix has entries (Bt)γj=l+jbˉγ(St,At)(\mathcal{B}_{t})_{\gamma j}=\partial_{l+j}\bar{b}^{\gamma}(S_{t},A_{t}), which are the defining formulas Etγδ=δbˉγ(St,At)E^{\gamma\delta}_{t}=\partial_{\delta}\bar{b}^{\gamma}(S_{t},A_{t}) and Btγj=l+jbˉγ(St,At)\mathsf{B}^{\gamma j}_{t}=\partial_{l+j}\bar{b}^{\gamma}(S_{t},A_{t}) of the completion-of-squares theorem; so Et=Et\mathcal{E}_{t}=E_{t} and Bt=Bt\mathcal{B}_{t}=\mathsf{B}_{t}, and the matrix Wt=ZtBt+12VtW_{t}=Z_{t}\mathcal{B}_{t}+\tfrac12V_{t} of the policy lemma, formed with the family ZZ, is the matrix Wt=ZtBt+12VtW_{t}=Z_{t}\mathsf{B}_{t}+\tfrac12V_{t} of the statement.

Hypotheses (H1)--(H4) and (C). Hypothesis (H1) of the policy lemma asks for a real r>0r>0 with i,jRtijaiajra2\sum_{i,j}R^{ij}_{t}a^{i}a^{j}\ge r|a|^{2} for all tt and aRma\in\mathbb{R}^{m}; the left-hand side is the entry pairing aRtaa\cdot R_{t}a of the completion-of-squares theorem, so this is hypothesis (H1), which holds with r=cJr=c_{J}. Hypothesis (H2) of the policy lemma asks for a family of symmetric matrices, continuously differentiable in integral form as in Weighted Second-Moment Evolution of the State Fluctuation Process, with terminal value F^\hat{F} and densities (EtZt+ZtEtWtRt1Wt+Qt)-(\mathcal{E}_{t}^{\top}Z_{t}+Z_{t}\mathcal{E}_{t}-W_{t}R_{t}^{-1}W_{t}^{\top}+Q_{t}) with Wt=ZtBt+12VtW_{t}=Z_{t}\mathcal{B}_{t}+\tfrac12V_{t}; with the identifications just made this is hypothesis (H2) of the completion-of-squares theorem, which the family ZZ of the common data satisfies. So ZZ can be taken as the Riccati family of (H2) of the policy lemma. Hypothesis (H3) asks for a real r~>0\tilde{r}>0 with b~υ(St)r~\tilde{b}^{\upsilon}(S_{t})\ge\tilde{r} for all υ\upsilon and tt; since StΔlS_{t}\in\Delta^{l} for every tt, hypothesis (OC) gives this with r~=b\tilde{r}=\underline{b}. Hypothesis (H4) asks that Π0\Pi_{0} be symmetric positive semidefinite, which the zero matrix is, x(0x)=0x\cdot(0x)=0 for every xx. Hypothesis (C) asks that RmA\mathbb{R}^{m}\setminus\mathcal{A} be an open subset of Rm\mathbb{R}^{m}. The set A\mathcal{A} is compact in (Rm,TdE)(\mathbb{R}^{m},\mathcal{T}_{d_{E}}), so by Compact Subset of Rn\mathbb{R}^n is Closed (with n=mn=m) it is closed there, that is, RmATdE\mathbb{R}^{m}\setminus\mathcal{A}\in\mathcal{T}_{d_{E}} is open in the metric space (Rm,dE)(\mathbb{R}^{m},d_{E}); by Euclidean Openness Agrees with Metric Openness on Rn\mathbb{R}^n it is then open in the Euclidean sense. So (C) holds. Hence the policy lemma applies to the data named in the statement.

The filter covariance. Conclusion 2 of the policy lemma states that D~r=E~r(Θ~r)1E~r\tilde{D}_{r}=\tilde{\mathcal{E}}_{r}^{\top}(\tilde{\Theta}^{\star}_{r})^{-1}\tilde{\mathcal{E}}_{r} is symmetric positive semidefinite with continuous entries, that every Θr\Theta^{\star}_{r} is positive semidefinite, and that consequently Global Existence and Uniqueness for the Kalman Covariance Riccati Equation furnishes exactly one assignment Π\Pi' of a matrix with ll rows and ll columns to each t[0,T]t\in[0,T], with continuous entries, satisfying Πt=Π0+0t(ErΠr+ΠrErΠrD~rΠr+Θr)dr\Pi'_{t}=\Pi_{0}+\int_{0}^{t}(\mathcal{E}_{r}\Pi'_{r}+\Pi'_{r}\mathcal{E}_{r}^{\top}-\Pi'_{r}\tilde{D}_{r}\Pi'_{r}+\Theta^{\star}_{r})\,dr with Π0\Pi_{0} the zero matrix — its filter covariance — and that every Πt\Pi'_{t} is symmetric positive semidefinite. This is the Riccati theorem applied on [0,T][0,T] with k=lk=l and the data A=EA=\mathcal{E}, C=ΘC=\Theta^{\star}, D=D~D=\tilde{D}, P0=0P_{0}=0; the assignment Π\Pi of the statement is by definition the solution furnished by that theorem for these data, which is unique. Hence the Riccati theorem applies, Π\Pi exists and is unique with continuous entries, every Πt\Pi_{t} is symmetric positive semidefinite, and Π\Pi is the filter covariance of the policy lemma.

The attainment theorem applies; continuity of the integrand; the value. The attainment theorem adopts the setting, hypotheses (H1)--(H4) and (C) and notation of the policy lemma, with the Riccati family ZZ fixed and the matrix Π0\Pi_{0} of (H4) — here the zero matrix; and for its conclusions 2 and 3 it assumes in addition that A\mathcal{A} is convex, which is part of the common data, and hypothesis (H5) of Mean-Square Error Covariance of the Approximate Kalman Filter as assumed in Cost Limit Along the Approximate Kalman Policy, which is the hypothesis (H5) assumed here with ϱ=ϱA\varrho=\varrho_{\mathcal{A}}. So the attainment theorem applies. Its setting records that the map tZtΘt+(WtRt1Wt)Πtt\mapsto Z_{t}\cdot\Theta^{\star}_{t}+(W_{t}R_{t}^{-1}W_{t}^{\top})\cdot\Pi_{t}, with the matrix pairing AB=γ,δAγδBγδA'\cdot B'=\sum_{\gamma,\delta}A'^{\gamma\delta}B'^{\gamma\delta}, is continuous on [0,T][0,T]; its Π\Pi being the filter covariance, that is, the present Π\Pi, and its WtRt1WtW_{t}R_{t}^{-1}W_{t}^{\top} being the present Ξt\Xi_{t}, this map is tγ,δZtγδΘtγδ+γ,δΞtγδΠtγδt\mapsto\sum_{\gamma,\delta}Z^{\gamma\delta}_{t}\Theta^{\star\gamma\delta}_{t}+\sum_{\gamma,\delta}\Xi^{\gamma\delta}_{t}\Pi^{\gamma\delta}_{t}, the integrand defining VV^{*}, which is therefore continuous on [0,T][0,T]. The number VV^{*} of the attainment theorem is Z0Π0+[0,T](ZtΘt+(WtRt1Wt)Πt)dtZ_{0}\cdot\Pi_{0}+\int_{[0,T]}(Z_{t}\cdot\Theta^{\star}_{t}+(W_{t}R_{t}^{-1}W_{t}^{\top})\cdot\Pi_{t})\,dt; since Π0\Pi_{0} is the zero matrix, Z0Π0=0Z_{0}\cdot\Pi_{0}=0, so it is the VV^{*} of the statement. This proves claim 1.

Claim 2. Let a family of solutions be admissible, with the constants κ\kappa^{\sharp}, J\mathcal{J}^{\sharp} and the points x0N\mathsf{x}^{N}_{0} of (I'), (CB), (D0).

Hypothesis (I) from (D0). Fix N1N\ge1 and write Ω={Σ0=x0N}\Omega_{*}=\{\Sigma_{0}=\mathsf{x}^{N}_{0}\}, an event of probability one by (D0) (an event because each component of Σ0\Sigma_{0} is a random variable by the definition of a solution). The empirical state measure takes its values in Δl\Delta^{l} by that definition, and any two points x,yΔlx,y\in\Delta^{l} satisfy xy2|x-y|\le2: each xγyγ1|x^{\gamma}-y^{\gamma}|\le1, so (xγyγ)2xγyγxγ+yγ(x^{\gamma}-y^{\gamma})^{2}\le|x^{\gamma}-y^{\gamma}|\le x^{\gamma}+y^{\gamma}, whence xy2γ(xγ+yγ)=24|x-y|^{2}\le\sum_{\gamma}(x^{\gamma}+y^{\gamma})=2\le4. Hence every component of s0=N(Σ0S0)\mathfrak{s}_{0}=\sqrt{N}(\Sigma_{0}-S_{0}) is bounded in absolute value by 2N2\sqrt{N} at every point of Ωag\Omega^{\mathrm{ag}}, so each product s0γs0δ\mathfrak{s}^{\gamma}_{0}\mathfrak{s}^{\delta}_{0} is bounded in absolute value by 4N4N everywhere, and on Ω\Omega_{*} it equals the constant N(x0N,γS0γ)(x0N,δS0δ)N(\mathsf{x}^{N,\gamma}_{0}-S^{\gamma}_{0})(\mathsf{x}^{N,\delta}_{0}-S^{\delta}_{0}), itself bounded by 4N4N since x0N,S0Δl\mathsf{x}^{N}_{0},S_{0}\in\Delta^{l}. By claim 2 of Almost Sure Inequalities Between Bounded Random Variables Pass to Expectations (with Ω\Omega_{*}, the common bound 4N4N and the constant random variable), and since the expectation of a constant random variable on a probability space is that constant by the definition of the expectation, E[s0γs0δ]=N(x0N,γS0γ)(x0N,δS0δ)\mathbb{E}[\mathfrak{s}^{\gamma}_{0}\mathfrak{s}^{\delta}_{0}]=N(\mathsf{x}^{N,\gamma}_{0}-S^{\gamma}_{0})(\mathsf{x}^{N,\delta}_{0}-S^{\delta}_{0}), whose absolute value is at most Nx0NS02=(Nx0NS0)2N|\mathsf{x}^{N}_{0}-S_{0}|^{2}=(\sqrt{N}|\mathsf{x}^{N}_{0}-S_{0}|)^{2} (as uγuδu2|u^{\gamma}u^{\delta}|\le|u|^{2} for uRlu\in\mathbb{R}^{l}). The last sequence has limit 00 by (D0) and claim 2 of the arithmetic of limits, so by claim 3 of the order properties of limits the sequence (E[s0γs0δ])N1(\mathbb{E}[\mathfrak{s}^{\gamma}_{0}\mathfrak{s}^{\delta}_{0}])_{N\ge1} has limit 00 for all γ,δ\gamma,\delta. This is hypothesis (I) of Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses with the zero matrix as Π0\Pi_{0}; the limit of a real sequence being unique (claim 1 of the order properties of limits, applied to a sequence and itself in both directions, gives that two limits of the same sequence are equal), the matrix Π0\Pi_{0} of hypothesis (I) is the zero matrix, which is hypothesis (I0) of the LQG lower bound theorem. Thus the family satisfies every hypothesis that Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses imposes on a family — (I), (I'), (CB) — beyond the common data and the standing hypotheses.

The setting of the certificate lemma. The certificate lemma adopts the setting, notation and standing hypotheses of The Path-Closeness Event under the Cost Bound: Closeness of the Empirical State Measure to the Mean-Field Trajectory and of the Record-Frozen Control to the Mean-Field Control on an Event of Probability 1O(N1/2)1-O(N^{-1/2}), hence of Closeness of the Realized Control and the Realized Mean-Field Flow on a High-Probability Event under the Cost Bound and of Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses: the common data and the standing hypotheses of the latter theorem, which are those adopted in the statement; a family, indexed by N1N\ge1, of solutions of the controlled NN-agent dynamics on [0,T][0,T] for A\mathcal{A}-valued observation-driven control policies h(N)h^{(N)} on NN-agent driving systems, with the objects named in the statement — which the given family provides; and the hypotheses of Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses on the family — (I') with the bound κ\kappa^{\sharp} and (CB) with the bound J\mathcal{J}^{\sharp}, which hold by admissibility, and (I), which has just been verified directly, so that the appeal to the certificate lemma is not circular on any reading of its adopted setting (its claim 1 re-derives (I) with the zero matrix from (D0)). The remaining objects of that setting are auxiliary and exist: the constants of Closeness of the Realized Control and the Realized Mean-Field Flow on a High-Probability Event under the Cost Bound and of The Path-Closeness Event under the Cost Bound: Closeness of the Empirical State Measure to the Mean-Field Trajectory and of the Record-Frozen Control to the Mean-Field Control on an Event of Probability 1O(N1/2)1-O(N^{-1/2}) are furnished by their claims; a natural number Nclmax(N1,Cesc2)N_{\mathrm{cl}}\ge\max(N_{1},C_{\mathrm{esc}}^{2}) exists; the realized controls, records, record spaces, prefix maps, record prefixes, truncated policies and observation-centred fluctuations are defined from the family; a dense sequence in L2([0,s];Rm)L^{2}([0,s];\mathbb{R}^{m}) exists for every s(0,T]s\in(0,T], as the certificate lemma's setting records; and reconstruction data exist by Measurable Reconstruction of the Controlled N-Agent Dynamics from Observation Records, as that setting records. Fix such choices. Finally its hypotheses (D0), (OC) and (X') are the hypothesis (D0) of admissibility (with the points x0N\mathsf{x}^{N}_{0}) and the hypotheses (OC) and (X') of the statement. Hence the certificate lemma applies to the given family, and its claims 1 and 2 hold for it.

Hypothesis (OC) of the LQG lower bound theorem. By claim 1 of the certificate lemma, hypothesis (OC) of Asymptotic Lower Bound for the Recentred N-Agent Cost by the Fluctuation LQG Value, under Law-Transported Injection Certificates holds with β~min=b\tilde\beta_{\min}=\underline{b} for the triple (S,A,P)(S,A,P) and the observation extension (U~,β~ˉ)(\tilde{U},\bar{\tilde\beta}) of (X'). Hypothesis (OC) of the LQG lower bound theorem is stated in the same words for the same objects — continuity on [0,T][0,T] of every entry of tE~tt\mapsto\tilde{\mathcal{E}}_{t} and of every tb~υ(St)t\mapsto\tilde{b}^{\upsilon}(S_{t}), and a uniform positive lower bound β~min\tilde\beta_{\min} on the b~υ(St)\tilde{b}^{\upsilon}(S_{t}) — so it holds with β~min=b\tilde\beta_{\min}=\underline{b}.

The profile responses and the information functionals. The LQG lower bound theorem defines, for a profile λ\lambda, its ψλ\psi_{\lambda} as the unique assignment with continuous components satisfying ψλ(u)=Π0λ(0)+[0,u](Erψλ(r)+Θrλ(r))dr\psi_{\lambda}(u)=\Pi_{0}\lambda(0)+\int_{[0,u]}(\mathcal{E}_{r}\psi_{\lambda}(r)+\Theta^{\star}_{r}\lambda(r))\,dr and its As(λ)\mathcal{A}_{s}(\lambda) as λ(0)(Π0λ(0))+[0,s](λ(r)(Θrλ(r))+ψλ(r)(D~rψλ(r)))dr\lambda(0)\cdot(\Pi_{0}\lambda(0))+\int_{[0,s]}(\lambda(r)\cdot(\Theta^{\star}_{r}\lambda(r))+\psi_{\lambda}(r)\cdot(\tilde{D}_{r}\psi_{\lambda}(r)))\,dr, with Π0\Pi_{0} the matrix of hypothesis (I) — here the zero matrix — and with the matrices Er\mathcal{E}_{r}, Θr\Theta^{\star}_{r} and D~r\tilde{D}_{r} of the fluctuation LQG data relative to the extensions of the statement. These are the displayed formulas by which Asymptotic Lower Bound for the Recentred N-Agent Cost by the Fluctuation LQG Value, under Law-Transported Injection Certificates defines its ψλ\psi_{\lambda} and As(λ)\mathcal{A}_{s}(\lambda), evaluated with the zero matrix as Π0\Pi_{0} and with the observation extension (U~,β~ˉ)(\tilde{U},\bar{\tilde\beta}); by claim 1 of the certificate lemma these coincide with the profile response ψλ\psi_{\lambda} and the information functional As(λ)\mathcal{A}_{s}(\lambda) of the certificate lemma.

Hypothesis (VT'). The objects named in hypothesis (VT') of the LQG lower bound theorem — the driving system (Ωag,Fag,Pag)(\Omega^{\mathrm{ag}},\mathcal{F}^{\mathrm{ag}},P^{\mathrm{ag}}), the record WW, the record spaces (Rs,Rs)(\mathbf{R}_{s},\mathcal{R}_{s}), the prefix maps πs\pi_{s}, the record prefixes W(s)=πsWW^{(s)}=\pi_{s}\circ W, the observation filtration (Gt)(\mathcal{G}_{t}), and the observation-centred fluctuation Xs=N(ΣsΦs)X'_{s}=\sqrt{N}(\Sigma_{s}-\Phi_{s}) with the realized mean-field flow Φt=St(S0,α^)\Phi_{t}=\mathsf{S}_{t}(S_{0},\hat\alpha), formed from the realized control α^\hat\alpha and the two-argument mean-field flow S\mathsf{S} with the base point S0S_{0} — are the objects so named in the setting of the certificate lemma, which adopts them from the same sources (its XsX'_{s} being, as it records, the observation-centred fluctuation of Asymptotic Lower Bound for the Recentred N-Agent Cost by the Fluctuation LQG Value, under Law-Transported Injection Certificates formed with the base point S0S_{0}, and the LQG lower bound theorem defining XsX'_{s} by the same formula from the same realized flow). Claim 2 of the certificate lemma asserts, for these objects and for its ψλ\psi_{\lambda} and As(λ)\mathcal{A}_{s}(\lambda), exactly the statement (VT') of the LQG lower bound theorem — the condition (C0) for every N1N\ge1 and s(0,T]s\in(0,T], and, for every s(0,T]s\in(0,T], every vector of Rl\mathbb{R}^{l} (written c\mathbf{c} there and cc here), every profile λ\lambda and every ϵ>0\epsilon>0, the existence of N2N_{2} and of the certificate data with (a), (C1), (C2'), (C3), (C4'), stated in the same words. Since the ψλ\psi_{\lambda} and As(λ)\mathcal{A}_{s}(\lambda) of the two items coincide, hypothesis (VT') holds.

Hypothesis (MS). Fix N1N\ge1, an admissible parameter vector π\pi in the sense of Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses and a block index k{0,,Kπ1}k\in\{0,\dots,K_{\pi}-1\}, KπK_{\pi} being the number of blocks determined by π\pi (written KK in the ledger lemma; the letter KK is the derivative bound here). By the definition of admissibility of π\pi, when the block cascade and the near-field data are formed from π\pi for the NN-th solution, every parameter-dependent requirement of Ledger Decomposition of the Recentred N-Agent Cost over the Block Cascade and Its Near-Field Filtering Lower Bound holds; its remaining requirements are the common data, the standing hypotheses, and the NN-th solution with its data, all in force. So the setting of the ledger lemma, and with it that of the measurability lemma, is instantiated by the common data, the NN-th solution and π\pi, and claim 2 of the measurability lemma gives that the three functions sE[1Tknr(s)usRsus]s\mapsto\mathbb{E}[\mathbf{1}_{\mathcal{T}^{\mathrm{nr}}_{k}(s)}u_{s}\cdot R_{s}u_{s}], sE[11Tk(s)]s\mapsto\mathbb{E}[1-\mathbf{1}_{\mathcal{T}_{k}(s)}] and sE[1Tkfr(s)]s\mapsto\mathbb{E}[\mathbf{1}_{\mathcal{T}^{\mathrm{fr}}_{k}(s)}], restricted to [tk,tk+1][t_{k},t_{k+1}], are measurable with respect to the trace Borel σ\sigma-algebra of that block. This is hypothesis (MS) of the LQG lower bound theorem for the family.

The remaining hypotheses; the theorem applies. The common data, the family of solutions, the standing hypotheses, hypotheses (I') and (CB), the admissible parameter vectors and the number V0V_{0} are adopted by the LQG lower bound theorem from Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses; hypothesis (I) holds as shown; hypothesis (H2) of the completion-of-squares theorem, with the Riccati family ZZ, is part of the common data; hypothesis (H1) holds with r=cJr=c_{J}; the observation data of the LQG lower bound theorem are the extension (U~,β~ˉ)(\tilde{U},\bar{\tilde\beta}) of (X') and the fluctuation LQG data relative to it; and (I0), (OC), (VT') and (MS) hold as shown. Hence every hypothesis of the LQG lower bound theorem holds for the common data and the family. Its Π\Pi is the solution of the Riccati theorem on [0,T][0,T] with the data A=EA=\mathcal{E}, C=ΘC=\Theta^{\star}, D=D~D=\tilde{D} and initial value the matrix Π0\Pi_{0} of hypothesis (I), that is, the zero matrix; this is the Π\Pi of the statement. Its matrices Ξt=WtRt1Wt\Xi_{t}=W_{t}R_{t}^{-1}W_{t}^{\top}, adopted from the cascade filtering lemma, are formed from the same Wt=ZtBt+12VtW_{t}=Z_{t}\mathsf{B}_{t}+\tfrac12V_{t} and RtR_{t} as the Ξt\Xi_{t} of the statement, and so coincide with them.

Conclusions. By the paragraph The limit value of Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses, each product sZsγδΘsγδs\mapsto Z^{\gamma\delta}_{s}\Theta^{\star\gamma\delta}_{s} is continuous on [0,T][0,T] with finite Lebesgue integral, hence so is the finite sum sγ,δZsγδΘsγδs\mapsto\sum_{\gamma,\delta}Z^{\gamma\delta}_{s}\Theta^{\star\gamma\delta}_{s} by the continuity of sums and the linearity of the integral, and

V0=γ,δZ0γδΠ0γδ+[0,T]γ,δZsγδΘsγδds=[0,T]γ,δZsγδΘsγδds,V_{0}=\sum_{\gamma,\delta}Z^{\gamma\delta}_{0}\Pi^{\gamma\delta}_{0}+\int_{[0,T]}\sum_{\gamma,\delta}Z^{\gamma\delta}_{s}\Theta^{\star\gamma\delta}_{s}\,ds=\int_{[0,T]}\sum_{\gamma,\delta}Z^{\gamma\delta}_{s}\Theta^{\star\gamma\delta}_{s}\,ds ,

the first sum vanishing because Π0\Pi_{0} is the zero matrix. Claim 1 of the LQG lower bound theorem gives that sγ,δΞsγδΠsγδs\mapsto\sum_{\gamma,\delta}\Xi^{\gamma\delta}_{s}\Pi^{\gamma\delta}_{s} is continuous, nonnegative and Lebesgue integrable on [0,T][0,T]. Hence, by the linearity of the integral (both summands being integrable),

V=[0,T]γ,δZsγδΘsγδds+[0,T]γ,δΞsγδΠsγδds=V0+[0,T]γ,δΞsγδΠsγδds.V^{*}=\int_{[0,T]}\sum_{\gamma,\delta}Z^{\gamma\delta}_{s}\Theta^{\star\gamma\delta}_{s}\,ds+\int_{[0,T]}\sum_{\gamma,\delta}\Xi^{\gamma\delta}_{s}\Pi^{\gamma\delta}_{s}\,ds=V_{0}+\int_{[0,T]}\sum_{\gamma,\delta}\Xi^{\gamma\delta}_{s}\Pi^{\gamma\delta}_{s}\,ds .

The recentred cost JN\mathcal{J}_{N} of the LQG lower bound theorem is the recentred cost of First-Order Expansion of the Recentred N-Agent Cost about a Stationary Mean-Field Triple and Its Coercive Lower Bound adopted from Asymptotic Lower Bound for the Recentred N-Agent Cost without Uniform Control-Moment Hypotheses, which is the quantity displayed in the statement. Claim 4 of the LQG lower bound theorem therefore gives: for every real ε>0\varepsilon>0 there is a natural number N5N_{5} with JNV0+[0,T]γ,δΞsγδΠsγδdsε=Vε\mathcal{J}_{N}\ge V_{0}+\int_{[0,T]}\sum_{\gamma,\delta}\Xi^{\gamma\delta}_{s}\Pi^{\gamma\delta}_{s}\,ds-\varepsilon=V^{*}-\varepsilon for every NN5N\ge N_{5}; take Nε=N5N_{\varepsilon}=N_{5}. This proves claim 2.

Claim 3.

The Kalman family is a family of solutions. By conclusion 1 of the attainment theorem, available by claim 1, each hK,Nh^{\mathrm{K},N} is an A\mathcal{A}-valued observation-driven control policy with horizon TT, control dimension mm and l~\tilde{l} channels, and the projected solution fixed in the statement is a solution of the controlled NN-agent dynamics on [0,T][0,T] for β\beta, β~\tilde\beta, the driving system (ΩK,FK,PK)(\Omega^{\mathrm{K}},\mathcal{F}^{\mathrm{K}},P^{\mathrm{K}}) and hK,Nh^{\mathrm{K},N}. So the Kalman family is a family of solutions for the common data, with JNK=N(JN[hK,N]JMF)+γP0γζNK,γ\mathcal{J}^{\mathrm{K}}_{N}=N(J^{N}[h^{\mathrm{K},N}]-J^{MF})+\sum_{\gamma}P^{\gamma}_{0}\zeta^{\mathrm{K},\gamma}_{N}, ζNK=N(EK[Σ0K]S0)\zeta^{\mathrm{K}}_{N}=N(\mathbb{E}^{\mathrm{K}}[\Sigma^{\mathrm{K}}_{0}]-S_{0}).

The initial conditions. Fix N1N\ge1 and write ΩK={Σ0K=y0N}\Omega^{\mathrm{K}}_{*}=\{\Sigma^{\mathrm{K}}_{0}=\mathsf{y}^{N}_{0}\}, an event of probability one by (DK) (it is an event, each component of Σ0K\Sigma^{\mathrm{K}}_{0} being a random variable by the definition of a solution). The empirical state measure takes its values in Δl\Delta^{l} by that definition, and any two points x,yΔlx,y\in\Delta^{l} satisfy xy2|x-y|\le2: each xγyγ1|x^{\gamma}-y^{\gamma}|\le1, so (xγyγ)2xγyγxγ+yγ(x^{\gamma}-y^{\gamma})^{2}\le|x^{\gamma}-y^{\gamma}|\le x^{\gamma}+y^{\gamma}, whence xy2γ(xγ+yγ)=24|x-y|^{2}\le\sum_{\gamma}(x^{\gamma}+y^{\gamma})=2\le4. Hence s0K=NΣ0KS02N|\mathfrak{s}^{\mathrm{K}}_{0}|=\sqrt{N}\,|\Sigma^{\mathrm{K}}_{0}-S_{0}|\le2\sqrt{N}, and every component of s0K\mathfrak{s}^{\mathrm{K}}_{0} is bounded in absolute value by 2N2\sqrt{N}, at every point of ΩK\Omega^{\mathrm{K}}; so the random variables s0K,γs0K,δ\mathfrak{s}^{\mathrm{K},\gamma}_{0}\mathfrak{s}^{\mathrm{K},\delta}_{0} and s0K4|\mathfrak{s}^{\mathrm{K}}_{0}|^{4} are bounded in absolute value by 4N4N and by 16N216N^{2} respectively, everywhere on ΩK\Omega^{\mathrm{K}}. On ΩK\Omega^{\mathrm{K}}_{*} one has s0K=N(y0NS0)\mathfrak{s}^{\mathrm{K}}_{0}=\sqrt{N}(\mathsf{y}^{N}_{0}-S_{0}), so there s0K,γs0K,δ=N(y0N,γS0γ)(y0N,δS0δ)\mathfrak{s}^{\mathrm{K},\gamma}_{0}\mathfrak{s}^{\mathrm{K},\delta}_{0}=N(\mathsf{y}^{N,\gamma}_{0}-S^{\gamma}_{0})(\mathsf{y}^{N,\delta}_{0}-S^{\delta}_{0}) and s0K4=N2y0NS04|\mathfrak{s}^{\mathrm{K}}_{0}|^{4}=N^{2}|\mathsf{y}^{N}_{0}-S_{0}|^{4}, constants which are themselves bounded in absolute value by 4N4N and by 16N216N^{2}, since y0N\mathsf{y}^{N}_{0} and S0S_{0} lie in Δl\Delta^{l}. By claim 2 of Almost Sure Inequalities Between Bounded Random Variables Pass to Expectations (applied with Ω=ΩK\Omega_{*}=\Omega^{\mathrm{K}}_{*}, the common bounds 4N4N, respectively 16N216N^{2}, and the constant random variables),

EK[s0K,γs0K,δ]=N(y0N,γS0γ)(y0N,δS0δ),EK[s0K4]=N2y0NS04,\mathbb{E}^{\mathrm{K}}[\mathfrak{s}^{\mathrm{K},\gamma}_{0}\mathfrak{s}^{\mathrm{K},\delta}_{0}]=N(\mathsf{y}^{N,\gamma}_{0}-S^{\gamma}_{0})(\mathsf{y}^{N,\delta}_{0}-S^{\delta}_{0}),\qquad \mathbb{E}^{\mathrm{K}}[|\mathfrak{s}^{\mathrm{K}}_{0}|^{4}]=N^{2}|\mathsf{y}^{N}_{0}-S_{0}|^{4},

the expectation of a constant random variable on a probability space being that constant by the definition of the expectation; these are the two identities of claim 3. Put eN=Ny0NS00e_{N}=\sqrt{N}|\mathsf{y}^{N}_{0}-S_{0}|\ge0, a sequence with limit 00 by (DK). Since uγuδu2|u^{\gamma}u^{\delta}|\le|u|^{2} for uRlu\in\mathbb{R}^{l} (each uγu|u^{\gamma}|\le|u|), the first expectation is bounded in absolute value by Ny0NS02=eN2N|\mathsf{y}^{N}_{0}-S_{0}|^{2}=e_{N}^{2}, which has limit 00 by claim 2 of the arithmetic of limits; by claim 3 of the order properties of limits the sequence (EK[s0K,γs0K,δ])N1(\mathbb{E}^{\mathrm{K}}[\mathfrak{s}^{\mathrm{K},\gamma}_{0}\mathfrak{s}^{\mathrm{K},\delta}_{0}])_{N\ge1} has limit 0=Π0γδ0=\Pi^{\gamma\delta}_{0}, which is hypothesis (I1) of Cost Limit Along the Approximate Kalman Policy with the zero matrix as Π0\Pi_{0}. The second expectation is eN4e_{N}^{4}, which has limit 00 by the same arithmetic of limits; by clause (d) of claim 3 of Real Powers Through the Exponential, and Elementary Asymptotic Tools: Monotonicity, Null Sequences of Negative Powers, Exponential Domination, Integer Rounding, and Square-Root and Exponential Inequalities there is N0N_{0} with eN4<1e_{N}^{4}<1 for every NN0N\ge N_{0}, so with κ,K=1+max{1,e14,,eN04}\kappa^{\sharp,\mathrm{K}}=1+\max\{1,e_{1}^{4},\dots,e_{N_{0}}^{4}\} one has sup{EK[s0K4]:N1}κ,K1<\sup\{\mathbb{E}^{\mathrm{K}}[|\mathfrak{s}^{\mathrm{K}}_{0}|^{4}]:N\ge1\}\le\kappa^{\sharp,\mathrm{K}}-1<\infty, which is hypothesis (I2).

Attainment. Conclusion 2 of the attainment theorem, applied to the driving systems and projected solutions fixed in the statement, whose empirical state measures are ΣK\Sigma^{\mathrm{K}} and whose state fluctuations are sK\mathfrak{s}^{\mathrm{K}}, states that each N(JN[hK,N]JMF)+γP0γζNK,γN(J^{N}[h^{\mathrm{K},N}]-J^{MF})+\sum_{\gamma}P^{\gamma}_{0}\zeta^{\mathrm{K},\gamma}_{N}, with ζNK=N(EK[Σ0K]S0)\zeta^{\mathrm{K}}_{N}=N(\mathbb{E}^{\mathrm{K}}[\Sigma^{\mathrm{K}}_{0}]-S_{0}) and with the NN-agent cost JN[hK,N]J^{N}[h^{\mathrm{K},N}] of the solution at level NN and the mean-field cost JMFJ^{MF} of (S,A)(S,A) — that is, each JNK\mathcal{J}^{\mathrm{K}}_{N} — is a well-defined real number and that the sequence (JNK)N1(\mathcal{J}^{\mathrm{K}}_{N})_{N\ge1} has limit the VV^{*} of that theorem, which is the present VV^{*} by claim 1.

Admissibility. (I'): κ0=1+EK[s0K4]=1+eN4κ,K\kappa_{0}=1+\mathbb{E}^{\mathrm{K}}[|\mathfrak{s}^{\mathrm{K}}_{0}|^{4}]=1+e_{N}^{4}\le\kappa^{\sharp,\mathrm{K}} for every N1N\ge1, with κ,K1\kappa^{\sharp,\mathrm{K}}\ge1 as above, which is (I') with the bound κ,K\kappa^{\sharp,\mathrm{K}}. (CB): by the definition of the limit there is NN_{\star} with JNKV<1|\mathcal{J}^{\mathrm{K}}_{N}-V^{*}|<1, hence JNKV+1\mathcal{J}^{\mathrm{K}}_{N}\le V^{*}+1, for every NNN\ge N_{\star}; put J,K=max{V+1,J1K,,JNK}0\mathcal{J}^{\sharp,\mathrm{K}}=\max\{|V^{*}|+1,|\mathcal{J}^{\mathrm{K}}_{1}|,\dots,|\mathcal{J}^{\mathrm{K}}_{N_{\star}}|\}\ge0; then JNKJ,K\mathcal{J}^{\mathrm{K}}_{N}\le\mathcal{J}^{\sharp,\mathrm{K}} for every N1N\ge1, which is (CB) with the bound J,K\mathcal{J}^{\sharp,\mathrm{K}}. (D0): the points x0N=y0NGN\mathsf{x}^{N}_{0}=\mathsf{y}^{N}_{0}\in\mathbb{G}_{N} satisfy PK(Σ0K=x0N)=1P^{\mathrm{K}}(\Sigma^{\mathrm{K}}_{0}=\mathsf{x}^{N}_{0})=1 and Nx0NS00\sqrt{N}|\mathsf{x}^{N}_{0}-S_{0}|\to0 by (DK). Hence the Kalman family is admissible, and (JNK)N1C(\mathcal{J}^{\mathrm{K}}_{N})_{N\ge1}\in\mathcal{C}. This proves claim 3.

Claim 4.

VV^{*} is the asymptotically optimal value. Condition (i) of Asymptotically Optimal Value of a Set of Real Sequences for VV^{*}: every sequence in C\mathcal{C} is the sequence (JN)N1(\mathcal{J}_{N})_{N\ge1} of an admissible family, so by claim 2, for every ε>0\varepsilon>0 there is NεN_{\varepsilon} with JNVε\mathcal{J}_{N}\ge V^{*}-\varepsilon for every NNεN\ge N_{\varepsilon}. Condition (ii): by claim 3 the sequence (JNK)N1(\mathcal{J}^{\mathrm{K}}_{N})_{N\ge1} belongs to C\mathcal{C} and converges to VV^{*}. So VV^{*} is an asymptotically optimal value of C\mathcal{C}, which is nonempty by claim 3.

Uniqueness. Let a real number vv be an asymptotically optimal value of C\mathcal{C}. Suppose v<Vv<V^{*} and put ε=(Vv)/2>0\varepsilon=(V^{*}-v)/2>0. By condition (ii) for vv there is a sequence (xN)N1(x_{N})_{N\ge1} in C\mathcal{C} converging to vv, so there is NN' with xNv<ε|x_{N}-v|<\varepsilon, hence xN<v+ε=Vεx_{N}<v+\varepsilon=V^{*}-\varepsilon, for every NNN\ge N'; by condition (i) for VV^{*} there is NN'' with xNVεx_{N}\ge V^{*}-\varepsilon for every NNN\ge N''; at N=max(N,N)N=\max(N',N'') these contradict each other. So vVv\ge V^{*}. Suppose v>Vv>V^{*} and put ε=(vV)/2>0\varepsilon=(v-V^{*})/2>0. By condition (ii) for VV^{*} (claim 3) the sequence (JNK)N1(\mathcal{J}^{\mathrm{K}}_{N})_{N\ge1} in C\mathcal{C} converges to VV^{*}, so there is NN' with JNK<V+ε=vε\mathcal{J}^{\mathrm{K}}_{N}<V^{*}+\varepsilon=v-\varepsilon for every NNN\ge N'; by condition (i) for vv there is NN'' with JNKvε\mathcal{J}^{\mathrm{K}}_{N}\ge v-\varepsilon for every NNN\ge N''; again a contradiction at N=max(N,N)N=\max(N',N''). So vVv\le V^{*}, and v=Vv=V^{*}. Hence VV^{*} is the unique asymptotically optimal value of C\mathcal{C}, as claimed.

The LQG minimum. Conclusion 3 of the attainment theorem, available by claim 1, states that for every linear-Gaussian state-observation model matched to the fluctuation LQG data as described there, the linear-quadratic-Gaussian cost JJ formed there attains a minimum over the extended admissible controls with values in Rm\mathbb{R}^{m} and minJ=V\min J=V^{*}, with the VV^{*} of that theorem, which is the present VV^{*}. This proves claim 4.

Please log in to copy this version.

Citations

Loading…

Dependency Graph

0 prerequisites

Comments

Loading…