(a) Eventual lower bounds and the limit inferior.Let (an)n∈N be a bounded sequence of real numbers, let M be a real number, and let N1∈N be such that M≤am for every m≥N1. Then M≤liminfnan.
In the notation of the definition of the limit inferior, write Ak={am:m∈N,m≥k}. The number M is a lower bound of AN1, so part (ii) of the definition of the greatest lower bound gives M≤infAN1. The number infAN1 belongs to the set {infAk:k∈N}, whose least upper bound is by definition liminfnan; a least upper bound is in particular an upper bound, so infAN1≤liminfnan. Transitivity of the order gives the assertion.
(b) Removing a doubled epsilon.Let a and b be real numbers such that a≤b+ε+ε for every real ε>0. Then a≤b.
Suppose not, so that b<a, and put δ=a−b, a positive real number. By claim 8 of Elementary Order Arithmetic in an Ordered Field the number δ/2 is positive and δ/2<δ; applying that claim to δ/2 shows that ε=δ/4 is positive and satisfies ε+ε=δ/2. The hypothesis applied to this ε gives a≤b+δ/2<b+δ=a, which is impossible.
(c) Changing the policy does not change the law of the initial empirical state.Let N∈N and let two solutions of the controlled N-agent dynamics on [0,T] for β, β~ and the N-th driving system be given, for policies in H, with regular events Ω0 and Ω0′ and empirical state measures Σ and Σ′. Then Σ0 and Σ0′ agree at every point of an event of probability 1, and they have the same law as random elements of (Δl,dΔ).
Write ς01,…,ς0N for the initial states of the driving system. Condition 1 of the definition of a solution holds at every point of the regular event, so σ0i(ω)=ς0i(ω) for every i and every ω∈Ω0, and the same identity holds for the second solution at every ω∈Ω0′. The empirical state measure at time 0 is determined by the states at time 0 through the formula recalled in that definition, so Σ0(ω)=Σ0′(ω) for every ω in E=Ω0∩Ω0′, and E∈FN because a σ-algebra is closed under finite intersections. Since PN(Ω0)=PN(Ω0′)=1, claim 3 of the basic properties of a measure gives PN(ΩN∖Ω0)=PN(ΩN∖Ω0′)=0. The complement of E is the union of those two sets, so claim 4 there (countable subadditivity), applied to the sequence whose first two terms are those sets and all of whose further terms are empty, gives PN(ΩN∖E)=0, and claim 3 again gives PN(E)=1.
Let Γ be a Borel set of (Δl,dΔ). Then Σ0−1(Γ)∩E=(Σ0′)−1(Γ)∩E. The sets Σ0−1(Γ)∩E and Σ0−1(Γ)∖E are disjoint with union Σ0−1(Γ), so claim 1 of the basic properties (finite additivity) gives PN(Σ0−1(Γ))=PN(Σ0−1(Γ)∩E)+PN(Σ0−1(Γ)∖E), and the last term vanishes by claim 2 (monotonicity) applied to the inclusion Σ0−1(Γ)∖E⊆ΩN∖E. The same computation applies to Σ0′, so the two probabilities agree; by the definition of the law, the laws coincide.
Consequently, since (Σ0N)N∈N converges in distribution to Yx0, and convergence in distribution is by definition weak convergence of the laws, the same convergence holds for the empirical state measures at time 0 of any family consisting, for each N, of a solution for the N-th driving system and some policy in H.
(d) Composites of strictly increasing sequences. If (Nj)j∈N and (jk)k∈N are strictly increasing sequences of natural numbers, then k↦Njk is strictly increasing: an induction shows that p<q implies Np<Nq, and jk<jk+1 for every k.
Step 1. Proof of claim 1. By the preamble, JN is a nonempty set of real numbers of which −C(T+1) is a lower bound, and VN=infJN. Part (ii) of the definition of the greatest lower bound gives −C(T+1)≤VN. By part (i) the number VN is a lower bound of JN; since hN∈H we have JN[hN]∈JN, so VN≤JN[hN]. Moreover ∣JN[hN]∣≤C(T+1), so JN[hN]≤C(T+1) by claim 6 of the absolute-value lemma, and VN≤C(T+1) by transitivity. From −C(T+1)≤VN≤C(T+1) and claim 6 again, ∣VN∣≤C(T+1).
Step 2. The open-loop policy is admissible. By claim 1 of the optimal-set structure lemma the set Mx0∗ is nonempty. Let ξ∗∈Mx0∗ and let A∗ be an admissible representative of ξ∗, which exists by claim 2 of the flow stability lemma; thus At∗∈A for every t∈[0,T], and, being a representative of an element of the Lebesgue space of square-integrable vector-valued functions on [0,T], the map A∗ has components measurable with respect to the trace Borel σ-algebra on [0,T] and the Borel σ-algebra of the real line. Hence the open-loop policy lemma applies, and its claim 1 shows that hA∗ is an observation-driven control policy with horizon T, control dimension m and l~ channels. By the formulas defining it in that lemma, every value of hA∗ is of the form At∗ with t∈[0,T], hence lies in A; this is exactly the requirement of the definition of an A-valued policy. Therefore hA∗∈H, and in particular JN[hA∗] is a real number lying in JN for every N.
Step 3. The initial states converge in mean square. For ω∈ΩN put ZN(ω)=∣Σ0N(ω)−x0∣2.
Second, 0≤ZN≤4 everywhere: for every ω both Σ0N(ω) and x0 lie in Δl, so they have Euclidean norm at most 1 by claim 1 of the compactness lemma, and claims 5 and 6 of the lemma on the Euclidean norm give ∣Σ0N(ω)−x0∣≤∣Σ0N(ω)∣+∣x0∣≤2; if 0≤u≤2 then u⋅u≤u⋅2≤2⋅2 by two applications of claim 10 of the order-arithmetic lemma in the nonstrict form obtained by adjoining the case of equality, the multipliers u and 2 being nonnegative.
Now let ε>0 be real and put EN,ε={ω∈ΩN:ε≤dΔ(Σ0N(ω),x0)}. By claims 2 and 3 of the lemma on convergence in distribution to a constant, each EN,ε is an event and the sequence (PN(EN,ε))N∈N converges to 0; and dΔ(Σ0N(ω),x0)=∣Σ0N(ω)−x0∣, because dΔ is the restriction to Δl of the Euclidean distance on Rl, as fixed in the compactness lemma, and that distance is expressed through the norm by claim 2 of the norm lemma. Let 1EN,ε be the function equal to 1 on EN,ε and to 0 elsewhere. For ω∈EN,ε we get ZN(ω)≤4=41EN,ε(ω)≤ε2+41EN,ε(ω); for ω∈/EN,ε we get ∣Σ0N(ω)−x0∣<ε and hence, by the same two applications of claim 10, ZN(ω)≤ε2=ε2+41EN,ε(ω). Thus ZN≤ε2+41EN,ε at every point of ΩN.
Let δ>0 be real. Choose ε to be the smaller of 1 and δ/4, a positive real number; then ε2≤ε≤δ/4, the first inequality by claim 10 in the nonstrict form applied to ε≤1 with the nonnegative multiplier ε. Choose N1 with PN(EN,ε)<δ/16 for every N≥N1. For such N we get 0≤EN[ZN]≤δ/4+δ/4=δ/2<δ, the lower bound because 0≤ZN everywhere. As δ>0 was arbitrary, (EN[ZN])N∈N converges to 0.
Step 4. Proof of claim 2. Retain ξ∗, A∗ and hA∗ from Step 2. By part (ii) of the existence, uniqueness and regularity theorem, for every N the set of solutions of the controlled N-agent dynamics on [0,T] for β, β~, the N-th driving system and the policy hA∗ is nonempty; these sets all lie in the set of all such solutions for all N, so Axiom of Countable Choice provides one for every N. Write ΣA,N for the empirical state measure of the chosen solution. By device (c) the maps Σ0A,N and Σ0N agree at every point of an event of probability 1; the map Σ0A,N is a random element of (Δl,dΔ) by claim 5 of the realized-control lemma, so ∣Σ0A,N−x0∣2 is a random variable by the argument of Step 3, and ∣Σ0A,N−x0∣2 and ZN agree at every point of that event; both are bounded everywhere, so claim 2 of the almost sure expectation lemma gives EN[∣Σ0A,N−x0∣2]=EN[ZN], and by Step 3 this converges to 0.
Step 5. Claim 3: the upper estimate. The sequences (VN)N∈N and (JN[hA∗])N∈N are bounded, both being bounded in absolute value by C(T+1), so that C(T+1)+1 is a strictly positive bound as the definition of a bounded sequence requires. For every N the number JN[hA∗] lies in JN, of which VN is a lower bound, so VN≤JN[hA∗]. Claim 2 of the basic properties of the limit inferior and the limit superior therefore gives limsupNVN≤limsupNJN[hA∗], while claim 5 there, applied to the sequence (JN[hA∗])N∈N, which converges to Jx0∗ by claim 2, gives limsupNJN[hA∗]=Jx0∗. Hence limsupNVN≤Jx0∗.
Step 6. Claim 3: the lower estimate. Let ε>0 be real. Fix N. By claim 1 of the order-arithmetic lemma VN<VN+ε, so VN+ε is not a lower bound of JN: otherwise part (ii) of the definition of the greatest lower bound would give VN+ε≤VN. Hence there is a policy h∈H with JN[h]<VN+ε, and, by part (ii) of the existence theorem, a solution for the N-th driving system and that policy. The set PN of all pairs consisting of such a policy and such a solution is therefore nonempty, and these sets lie in a common set; so Axiom of Countable Choice provides a pair for every N. Write h~N for the chosen policy and Σ~N for the empirical state measure of the chosen solution, so that JN[h~N]<VN+ε for every N.
By device (c) the sequence (Σ~0N)N∈N converges in distribution to Yx0. Hence every hypothesis of the asymptotic lower bound theorem is met by the data consisting of the given driving systems, the policies h~N and the chosen solutions, and claim 3 there gives that (JN[h~N])N∈N is bounded with
Jx0∗≤NliminfJN[h~N].
By claim 3 of the basic properties of the limit inferior and the limit superior there is N1∈N with liminfnJn[h~n]−ε<Jm[h~m] for every m≥N1. For such m,
Jx0∗−ε≤nliminfJn[h~n]−ε<Jm[h~m]<Vm+ε,
so Jx0∗−ε<Vm+ε and hence, adding ε to both sides by claim 1 of the order-arithmetic lemma, Jx0∗<Vm+ε+ε; in particular Jx0∗−ε−ε≤Vm for every m≥N1. Device (a), applied to the bounded sequence (VN)N∈N, gives Jx0∗−ε−ε≤liminfNVN, that is, Jx0∗≤liminfNVN+ε+ε. Since ε>0 was arbitrary, device (b) gives Jx0∗≤liminfNVN.
Step 7. Proof of claim 3. By claim 1 of the basic properties of the limit inferior and the limit superior, liminfNVN≤limsupNVN. With Steps 5 and 6 this yields
Jx0∗≤NliminfVN≤NlimsupVN≤Jx0∗,
so both the limit inferior and the limit superior equal Jx0∗, and claim 5 there shows that (VN)N∈N converges to Jx0∗.
For the second sequence, claim 1 and the hypothesis on (ηN)N∈N give 0≤JN[hN]−VN≤ηN, hence ∣JN[hN]−VN∣≤ηN by claim 6 of the absolute-value lemma. Claim 5 of that lemma (the triangle inequality) then gives
Step 8. Claim 4: measurability. By claim 4 of the optimal-set structure lemma the function D on X is sequentially continuous: if (yj)j∈N is a sequence in Xconverging to x in (X,dX), then (D(yj))j∈N converges to D(x); by claim 3 of the arithmetic of limits, (−D(yj))j∈N converges to −D(x). Consequently both D and −D are lower semicontinuous on X: given such a sequence and a real ε>0, the definition of the limit provides N with ∣D(yj)−D(x)∣<ε for j≥N, whence D(x)−ε<D(yj) by claim 6 of the absolute-value lemma, and similarly for −D; claims 1 and 3 of the sequential characterization of lower semicontinuity, applied with the subset X of X, give the assertion.
By claim 5 of the Borel toolkit the function D is measurable with respect to B(X) and the Borel σ-algebra of the real line. By claim 5 of the realized-control lemma the pair (Σ0N,α^N) is a random element of (X,dX), that is, measurable with respect to FN and B(X). Claim 4 (composition) of the Borel toolkit therefore shows that DN is measurable, that is, a random variable.
Fix a real ε>0 and put Dε↑={x∈X:ε≤D(x)}. Since Dε↑={x∈X:−D(x)≤−ε} and −D is lower semicontinuous on X, claim 3 of the lemma on sublevel and superlevel sets, applied with the subset X of X, shows that Dε↑ is closed in the topology of open subsets of (X,dX); hence Dε↑∈B(X) by claim 1 of the Borel toolkit. Moreover {ω∈ΩN:ε≤DN(ω)} is the preimage of Dε↑ under (Σ0N,α^N), so by the definition of the law
aN:=PN({ω∈ΩN:ε≤DN(ω)})=μN(Dε↑).
Step 9. Claim 4: the limit. Suppose, for a contradiction, that (aN)N∈N does not converge to 0. Since 0≤aN for every N, negating the definition of the limit provides a real ε0>0 such that for every K∈N there is m≥K with ε0≤am.
Let S={n∈N:ε0≤an}, which is nonempty by the case K=1, and let R be the set of pairs (n,n′)∈S×S with n<n′. For n∈S the case K=n+1 provides n′∈S with n<n′, so (n,n′)∈R; thus R is a binary relation on S such that every element of S is related to some element of S. Let s be an element of S, which exists because S is nonempty. Then Axiom of Dependent Choice, applied to S, R and s, provides a sequence (Nj)j∈N in S with N1=s and (Nj,Nj+1)∈R, that is Nj<Nj+1, for every j. It is strictly increasing and satisfies ε0≤aNj for every j.
By claim 1 of the asymptotic lower bound theorem each μN is a Borel measure on (X,dX) with μN(X)=1, and by claim 4 of the compactness lemma (X,dX) is compact. Hence the weak sequential compactness theorem applies to (μNj)j∈N and provides a strictly increasing sequence (jk)k∈N of natural numbers and a Borel measure μ on (X,dX) with μ(X)=1 such that (μNjk)k∈N converges weakly to μ. By device (d) the sequence (Njk)k∈N is strictly increasing.
By claim 3 the sequence (JN[hN])N∈N converges to Jx0∗, so claim 4 of the asymptotic lower bound theorem applies to the strictly increasing sequence (Njk)k∈N and to μ: the set Π={x0}×Mx0∗ belongs to B(X) and μ(Π)=1.
The sets Dε↑ and Π are disjoint: for ξ∗∈Mx0∗ claim 2 of the optimal-set structure lemma gives D(x0,ξ∗)=0, and ε≤0 is false because ε>0. Hence Dε↑⊆X∖Π. Since μ(X)=1 is finite, claim 3 of the basic properties of a measure gives μ(X∖Π)=μ(X)−μ(Π)=0, and claim 2 (monotonicity) gives μ(Dε↑)≤0; as a measure takes values in [0,∞], this forces μ(Dε↑)=0.
The measures μNjk and μ are probability measures on (X,B(X)) and X is nonempty, so claim 4 of the portmanteau theorem, applied to the closed set Dε↑, gives
klimsupμNjk(Dε↑)≤μ(Dε↑)=0.
On the other hand ε0≤aNjk=μNjk(Dε↑) for every k, and the sequence (μNjk(Dε↑))k∈N is bounded, with strictly positive bound 2, because 0≤μN(Dε↑)≤μN(X)=1 by monotonicity. Hence claim 2 of the basic properties of the limit inferior and the limit superior gives ε0≤liminfkμNjk(Dε↑), and claim 1 there gives liminfkμNjk(Dε↑)≤limsupkμNjk(Dε↑). Combining, ε0≤0, contradicting ε0>0.
Therefore (aN)N∈N converges to 0, which completes the proof of claim 4 and of the theorem.