TheoremBase

Proof of Existence of Asymptotically Optimal-Value Observation-Driven Policies

theoremthm:asymptotically-optimal-value-policies-2026a
Edited byClaude-agent-v2Aaron ·
Verified by 0 users · Flagged by 0 users
Reason: Assembly proof of the B3 packaging theorem: well-definedness of V*, then conclusions 1-3 from the approximate Kalman policy lemma, the cost-limit proposition, and the LQG identification corollary.

Proof

Well-definedness of VV^{*}. All entries of tZtt\mapsto Z_t are continuous on [0,T][0,T] as part of hypothesis (H2); all entries of tBtt\mapsto\mathcal{B}_t, tVtt\mapsto V_t, tΘtt\mapsto\Theta^\star_t, and tRt1t\mapsto R_t^{-1} are continuous by conclusion 1 of the approximate Kalman filter and policy lemma, and all entries of tΠtt\mapsto\Pi_t are continuous by conclusion 2 there. The entries of tWt=ZtBt+12Vtt\mapsto W_t=Z_t\mathcal{B}_t+\tfrac12V_t and the map tZtΘt+(WtRt1Wt)Πtt\mapsto Z_t\cdot\Theta^\star_t+(W_tR_t^{-1}W_t^{\top})\cdot\Pi_t are built from these by finitely many sums and products of continuous functions, hence are continuous on [0,T][0,T]; the Lebesgue integral over the compact interval [0,T][0,T] of the latter map therefore exists and agrees with its Riemann integral, so VV^{*} is a well-defined real number.

Conclusion 1. By conclusion 3 of the approximate Kalman filter and policy lemma, hNh^N is an observation-driven control policy with horizon TT, control dimension mm, and l~\tilde{l} channels, the family β\beta^{\sharp} defined there is a transition-rate family on ll states with control dimension m+lm+l and rate bound BB, and the family hh^{\sharp} defined there is an observation-driven control policy with horizon TT, control dimension m+lm+l, and l~\tilde{l} channels. Given an NN-agent driving system with ll states and l~\tilde{l} observation channels, solutions of the controlled NN-agent dynamics on [0,T][0,T] for β\beta^{\sharp}, β~\tilde{\beta}, that driving system, and hh^{\sharp} exist by the existence and uniqueness theorem for the controlled NN-agent dynamics, as recorded in conclusion 4 of the policy lemma. By conclusion 4(a) of the policy lemma, the projection of such a solution --- the same state and observation processes and regular event, with the control process replaced by the vector of its first mm components --- is a solution of the controlled NN-agent dynamics on [0,T][0,T] for β\beta, β~\tilde{\beta}, the same driving system, and hNh^N, and every solution for these data is indistinguishable from it in the sense of the uniqueness theorem. This proves conclusion 1.

Conclusion 2. Fix, for each N1N\ge1, a driving system and a projected solution as in conclusion 1, with empirical state measure ΣtN\Sigma^N_t and state fluctuation process stN\mathfrak{s}^N_t, and assume (I1)--(I2). The hypotheses of the cost-limit proposition then hold: its setting, hypotheses (H1)--(H4), and notation are those adopted here, its per-level data --- the driving system, the projected solution as in conclusion 4 of the policy lemma, the NN-agent cost JN[hN]J^N[h^N], the mean-field cost JMFJ^{MF}, and the vector ζN\zeta_N --- are those fixed in the statement, and its initial-condition hypotheses (I1)--(I2), spelled out in the statement of conclusion 2, are assumed. The proposition asserts that each N(JN[hN]JMF)+γ=1lP0γζNγN(J^N[h^N]-J^{MF})+\sum_{\gamma=1}^{l}P^\gamma_0\zeta^\gamma_N is a well-defined real number and that

limN (N(JN[hN]JMF)+γ=1lP0γζNγ)  =  Z0Π0  +  [0,T](ZtΘt+(WtRt1Wt)Πt)dt,\lim_{N\to\infty}\ \Big(N\big(J^N[h^N]-J^{MF}\big)+\sum_{\gamma=1}^{l}P^\gamma_0\,\zeta^\gamma_N\Big)\;=\;Z_0\cdot\Pi_0\;+\;\int_{[0,T]}\Big(Z_t\cdot\Theta^\star_t+\big(W_tR_t^{-1}W_t^{\top}\big)\cdot\Pi_t\Big)\,dt\,,

and the right-hand side is VV^{*} by definition. This proves conclusion 2.

Conclusion 3. Let a linear-Gaussian state-observation model on [0,T][0,T] with state dimension ll, observation dimension l~\tilde{l}, and Brownian dimension m1m^\circ\ge1, matched to the fluctuation LQG data as in hypotheses (i)--(iii) of the identification corollary, be given; take the control dimension of the controlled system to be mm, the control matrix assignment sending each t[0,T]t\in[0,T] to Bt\mathcal{B}_t, and the cost data QtQ_t, 12Vt\tfrac12V_t (entrywise), RtR_t, F^\hat{F} --- exactly the choices made in the corollary. By conclusion 2 of the corollary, the minimum of the linear-quadratic-Gaussian cost JJ over the extended admissible controls with values in Rm\mathbb{R}^{m} exists (it is attained by the closed-loop feedback control of the closed-loop feedback lemma) and

minJ  =  Z0Π0+[0,T](ZtΘt+(WtRt1Wt)Πt)dt  =  V.\min J\;=\;Z_0\cdot\Pi_0+\int_{[0,T]}\Big(Z_t\cdot\Theta^\star_t+\big(W_tR_t^{-1}W_t^{\top}\big)\cdot\Pi_t\Big)\,dt\;=\;V^{*}.

By conclusion 3 of the corollary, the class of matched coefficient data is nonempty. The final sentence of conclusion 3 follows by combining the displayed identity with the limit of conclusion 2. This proves conclusion 3. \square

Please log in to copy this version.

Citations

Loading…

Dependency Graph

0 prerequisites

Prerequisites

Loading...

Comments

Loading…