Each result cited is universally quantified over the data in its own statement.
Throughout, the notation is that of First-Order Equations on the Noise Wasserstein Space Relative to a Noise Penalty Pair: Standing Notation ; elementary order and arithmetic of real numbers, including absolute values and the least of two positive reals, is carried by The Real Numbers: Standing Notation and Background §background , in force through Probability Measures on a Hilbert Space Transported in the Noise Norm: Standing Notation §background and Borel Probability Measures on a Real Hilbert Space with an Orthonormal Basis: Standing Notation §background . For μ ∈ P ( X ) \mu\in\mathcal{P}(X) μ ∈ P ( X ) the space L 2 ( μ ; X a ) L^{2}(\mu;X^{a}) L 2 ( μ ; X a ) is a real Hilbert space, in particular a real inner product space , whose norm ∥ ⋅ ∥ μ \lVert\cdot\rVert_{\mu} ∥ ⋅ ∥ μ is the norm of its inner product, by The Space of Square-Integrable Hilbert-Valued Maps is a Real Hilbert Space: Coordinates and Synthesis §hilbert . The two claims are proved together: fix s ∈ { − 1 , 1 } s\in\{-1,1\} s ∈ { − 1 , 1 } and a point μ ^ ∈ Q ∩ D Σ \hat{\mu}\in Q\cap\mathcal{D}_{\Sigma} μ ^ ∈ Q ∩ D Σ at which the function χ + s δ E \chi+s\delta\mathcal{E} χ + sδ E on D \mathcal{D} D , with value χ ( μ ) + s δ E ( μ ) \chi(\mu)+s\delta\,\mathcal{E}(\mu) χ ( μ ) + sδ E ( μ ) at μ \mu μ , has a local maximum relative to D \mathcal{D} D if s = − 1 s=-1 s = − 1 , and a local minimum relative to D \mathcal{D} D if s = 1 s=1 s = 1 . For s = − 1 s=-1 s = − 1 this is the hypothesis of claim 1 and for s = 1 s=1 s = 1 that of claim 2. Then μ ^ ∈ D Σ ⊆ D \hat{\mu}\in\mathcal{D}_{\Sigma}\subseteq\mathcal{D} μ ^ ∈ D Σ ⊆ D by Noise Penalty Pairs on the Noise Wasserstein Space: the Penalty, Its Score, and Their Domains §pair . By Local Maximum of a Function Relative to a Subset of a Metric Space , respectively Local Minimum of a Function Relative to a Subset of a Metric Space , fix a positive R ∈ R R\in\mathbb{R} R ∈ R such that for every μ ∈ D \mu\in\mathcal{D} μ ∈ D with W a ( μ ^ , μ ) < R W_{a}(\hat{\mu},\mu)<R W a ( μ ^ , μ ) < R
χ ( μ ) + s δ E ( μ ) ≤ χ ( μ ^ ) + s δ E ( μ ^ ) ( s = − 1 ) , χ ( μ ) + s δ E ( μ ) ≥ χ ( μ ^ ) + s δ E ( μ ^ ) ( s = 1 ) . ( ∗ ) \chi(\mu)+s\delta\,\mathcal{E}(\mu)\le\chi(\hat{\mu})+s\delta\,\mathcal{E}(\hat{\mu})\quad(s=-1),\qquad\chi(\mu)+s\delta\,\mathcal{E}(\mu)\ge\chi(\hat{\mu})+s\delta\,\mathcal{E}(\hat{\mu})\quad(s=1). \tag{$\ast$} χ ( μ ) + sδ E ( μ ) ≤ χ ( μ ^ ) + sδ E ( μ ^ ) ( s = − 1 ) , χ ( μ ) + sδ E ( μ ) ≥ χ ( μ ^ ) + sδ E ( μ ^ ) ( s = 1 ) . ( ∗ )
We show that ∇ χ ( μ ^ ) = − s δ Σ ( μ ^ ) \nabla\chi(\hat{\mu})=-s\delta\,\Sigma(\hat{\mu}) ∇ χ ( μ ^ ) = − sδ Σ ( μ ^ ) , which is ∇ χ ( μ ^ ) = δ Σ ( μ ^ ) \nabla\chi(\hat{\mu})=\delta\,\Sigma(\hat{\mu}) ∇ χ ( μ ^ ) = δ Σ ( μ ^ ) for s = − 1 s=-1 s = − 1 and ∇ χ ( μ ^ ) = − δ Σ ( μ ^ ) \nabla\chi(\hat{\mu})=-\delta\,\Sigma(\hat{\mu}) ∇ χ ( μ ^ ) = − δ Σ ( μ ^ ) for s = 1 s=1 s = 1 .
Step 1 (the perturbed measures). Let ψ ∈ F C b 1 ( X ) \psi\in\mathcal{F}C^{1}_{b}(X) ψ ∈ F C b 1 ( X ) . Its noise gradient ∇ a ψ : X → X a \nabla_{a}\psi:X\to X^{a} ∇ a ψ : X → X a is measurable and square-integrable with respect to μ ^ \hat{\mu} μ ^ by The Noise Gradient of a Cylindrical Function and the Noise Tangent Space at a Probability Measure on a Hilbert Space §gradient ; its class, again written ∇ a ψ \nabla_{a}\psi ∇ a ψ , belongs to the set G μ ^ a G^{a}_{\hat{\mu}} G μ ^ a of The Noise Gradient of a Cylindrical Function and the Noise Tangent Space at a Probability Measure on a Hilbert Space §gradients . Write n ψ = ∥ ∇ a ψ ∥ μ ^ ≥ 0 n_{\psi}=\lVert\nabla_{a}\psi\rVert_{\hat{\mu}}\ge0 n ψ = ∥ ∇ a ψ ∥ μ ^ ≥ 0 . For t ∈ R t\in\mathbb{R} t ∈ R let S t = i d + t ∇ a ψ S_{t}=\mathrm{id}+t\,\nabla_{a}\psi S t = id + t ∇ a ψ and ν t = ( S t ) # μ ^ \nu_{t}=(S_{t})_{\#}\hat{\mu} ν t = ( S t ) # μ ^ . By Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §displacement , applied with μ ^ \hat{\mu} μ ^ in place of its μ \mu μ , with h h h the class of ∇ a ψ \nabla_{a}\psi ∇ a ψ and its representative ∇ a ψ \nabla_{a}\psi ∇ a ψ , the map S t S_{t} S t is Borel, π t = ( i d , S t ) # μ ^ ∈ Π a ( μ ^ , ν t ) \pi_{t}=(\mathrm{id},S_{t})_{\#}\hat{\mu}\in\Pi^{a}(\hat{\mu},\nu_{t}) π t = ( id , S t ) # μ ^ ∈ Π a ( μ ^ , ν t ) , and
I a ( π t ) = t 2 n ψ 2 , J a ( η , π t ) = t ⟨ η , ∇ a ψ ⟩ μ ^ for every η ∈ L 2 ( μ ^ ; X a ) ; I^{a}(\pi_{t})=t^{2}n_{\psi}^{2},\qquad\mathcal{J}^{a}(\eta,\pi_{t})=t\,\langle\eta,\nabla_{a}\psi\rangle_{\hat{\mu}}\quad\text{for every }\eta\in L^{2}(\hat{\mu};X^{a}); I a ( π t ) = t 2 n ψ 2 , J a ( η , π t ) = t ⟨ η , ∇ a ψ ⟩ μ ^ for every η ∈ L 2 ( μ ^ ; X a ) ;
and since μ ^ ∈ D Σ ⊆ P ρ a \hat{\mu}\in\mathcal{D}_{\Sigma}\subseteq\mathcal{P}^{a}_{\rho} μ ^ ∈ D Σ ⊆ P ρ a (Noise Penalty Pairs on the Noise Wasserstein Space: the Penalty, Its Score, and Their Domains §pair ), ν t ∈ P ρ a \nu_{t}\in\mathcal{P}^{a}_{\rho} ν t ∈ P ρ a by Noise Displacement and Cross Pairings Along Couplings: Bounds, Linearity, Displacement Couplings, Polarisation and a Vanishing Criterion §displacement-connected . As ∣ t ∣ n ψ ≥ 0 |t|\,n_{\psi}\ge0 ∣ t ∣ n ψ ≥ 0 and ( ∣ t ∣ n ψ ) 2 = t 2 n ψ 2 (|t|\,n_{\psi})^{2}=t^{2}n_{\psi}^{2} ( ∣ t ∣ n ψ ) 2 = t 2 n ψ 2 , the uniqueness in Existence and Uniqueness of the Nonnegative Square Root gives I a ( π t ) = ∣ t ∣ n ψ \sqrt{I^{a}(\pi_{t})}=|t|\,n_{\psi} I a ( π t ) = ∣ t ∣ n ψ . By The Noise Wasserstein Distance §distance , W a ( μ ^ , ν t ) 2 ≤ I a ( π t ) = ( ∣ t ∣ n ψ ) 2 W_{a}(\hat{\mu},\nu_{t})^{2}\le I^{a}(\pi_{t})=(|t|\,n_{\psi})^{2} W a ( μ ^ , ν t ) 2 ≤ I a ( π t ) = ( ∣ t ∣ n ψ ) 2 , so claim 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field , both numbers being nonnegative, gives
W a ( μ ^ , ν t ) ≤ ∣ t ∣ n ψ . (1) W_{a}(\hat{\mu},\nu_{t})\le|t|\,n_{\psi}. \tag{1} W a ( μ ^ , ν t ) ≤ ∣ t ∣ n ψ . ( 1 )
For t = 0 t=0 t = 0 the map S 0 S_{0} S 0 is i d \mathrm{id} id , so ν 0 = μ ^ \nu_{0}=\hat{\mu} ν 0 = μ ^ , because i d − 1 ( B ) = B \mathrm{id}^{-1}(B)=B id − 1 ( B ) = B for every Borel set B B B .
Step 2 (the two one-variable functions). By Noise Penalty Pairs on the Noise Wasserstein Space: the Penalty, Its Score, and Their Domains §variation , applied at μ ^ ∈ D Σ \hat{\mu}\in\mathcal{D}_{\Sigma} μ ^ ∈ D Σ with this ψ \psi ψ , there is a positive t 0 ∈ R t_{0}\in\mathbb{R} t 0 ∈ R such that ν t ∈ D \nu_{t}\in\mathcal{D} ν t ∈ D for every t ∈ ( − t 0 , t 0 ) t\in(-t_{0},t_{0}) t ∈ ( − t 0 , t 0 ) and the function e : ( − t 0 , t 0 ) → R e:(-t_{0},t_{0})\to\mathbb{R} e : ( − t 0 , t 0 ) → R , e ( t ) = E ( ν t ) e(t)=\mathcal{E}(\nu_{t}) e ( t ) = E ( ν t ) , is differentiable at 0 0 0 with e ′ ( 0 ) = ⟨ Σ ( μ ^ ) , ∇ a ψ ⟩ μ ^ e'(0)=\langle\Sigma(\hat{\mu}),\nabla_{a}\psi\rangle_{\hat{\mu}} e ′ ( 0 ) = ⟨ Σ ( μ ^ ) , ∇ a ψ ⟩ μ ^ . The point 0 0 0 is an interior point of ( − t 0 , t 0 ) (-t_{0},t_{0}) ( − t 0 , t 0 ) by Basic Facts about Intervals of the Real Line and Their Interior Points §open-interval , since − t 0 < 0 < t 0 -t_{0}<0<t_{0} − t 0 < 0 < t 0 .
Let k : ( − t 0 , t 0 ) → R k:(-t_{0},t_{0})\to\mathbb{R} k : ( − t 0 , t 0 ) → R be k ( t ) = χ ( ν t ) k(t)=\chi(\nu_{t}) k ( t ) = χ ( ν t ) , so that k ( 0 ) = χ ( μ ^ ) k(0)=\chi(\hat{\mu}) k ( 0 ) = χ ( μ ^ ) . We show that k k k is differentiable at 0 0 0 with k ′ ( 0 ) = ⟨ ∇ χ ( μ ^ ) , ∇ a ψ ⟩ μ ^ k'(0)=\langle\nabla\chi(\hat{\mu}),\nabla_{a}\psi\rangle_{\hat{\mu}} k ′ ( 0 ) = ⟨ ∇ χ ( μ ^ ) , ∇ a ψ ⟩ μ ^ . By property (b) of Noise Intrinsic Test Functions on the Noise Wasserstein Space §test , χ \chi χ is differentiable along noise couplings at μ ^ ∈ Q \hat{\mu}\in Q μ ^ ∈ Q with gradient ∇ χ ( μ ^ ) \nabla\chi(\hat{\mu}) ∇ χ ( μ ^ ) . Put c = 1 + n ψ c=1+n_{\psi} c = 1 + n ψ , a positive real number. Let ε ∈ R \varepsilon\in\mathbb{R} ε ∈ R be positive, and let θ \theta θ be as in Differentiability of a Function on the Noise-Connected Measures Along Noise Couplings, and Its Gradient §differentiable for the positive number ε ( 2 c ) − 1 \varepsilon\,(2c)^{-1} ε ( 2 c ) − 1 in place of its ε \varepsilon ε . Let t ∈ ( − t 0 , t 0 ) t\in(-t_{0},t_{0}) t ∈ ( − t 0 , t 0 ) satisfy 0 < ∣ t ∣ < θ c − 1 0<|t|<\theta\,c^{-1} 0 < ∣ t ∣ < θ c − 1 . Then I a ( π t ) = ∣ t ∣ n ψ ≤ ∣ t ∣ c < θ \sqrt{I^{a}(\pi_{t})}=|t|\,n_{\psi}\le|t|\,c<\theta I a ( π t ) = ∣ t ∣ n ψ ≤ ∣ t ∣ c < θ , so I a ( π t ) < θ 2 I^{a}(\pi_{t})<\theta^{2} I a ( π t ) < θ 2 by claim 1 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field . As ν t ∈ P ρ a \nu_{t}\in\mathcal{P}^{a}_{\rho} ν t ∈ P ρ a and π t ∈ Π a ( μ ^ , ν t ) \pi_{t}\in\Pi^{a}(\hat{\mu},\nu_{t}) π t ∈ Π a ( μ ^ , ν t ) , the estimate of that clause and Step 1, with η = ∇ χ ( μ ^ ) \eta=\nabla\chi(\hat{\mu}) η = ∇ χ ( μ ^ ) , give
∣ k ( t ) − k ( 0 ) − t ⟨ ∇ χ ( μ ^ ) , ∇ a ψ ⟩ μ ^ ∣ ≤ ε ( 2 c ) − 1 ∣ t ∣ n ψ ≤ ε 2 ∣ t ∣ , \Bigl|k(t)-k(0)-t\,\bigl\langle\nabla\chi(\hat{\mu}),\nabla_{a}\psi\bigr\rangle_{\hat{\mu}}\Bigr|\le\varepsilon\,(2c)^{-1}\,|t|\,n_{\psi}\le\tfrac{\varepsilon}{2}\,|t| , k ( t ) − k ( 0 ) − t ⟨ ∇ χ ( μ ^ ) , ∇ a ψ ⟩ μ ^ ≤ ε ( 2 c ) − 1 ∣ t ∣ n ψ ≤ 2 ε ∣ t ∣ ,
the last step because n ψ ≤ c n_{\psi}\le c n ψ ≤ c . Multiplying by the positive ∣ t ∣ − 1 |t|^{-1} ∣ t ∣ − 1 gives ∣ k ( t ) − k ( 0 ) t − ⟨ ∇ χ ( μ ^ ) , ∇ a ψ ⟩ μ ^ ∣ ≤ ε 2 < ε \bigl|\tfrac{k(t)-k(0)}{t}-\langle\nabla\chi(\hat{\mu}),\nabla_{a}\psi\rangle_{\hat{\mu}}\bigr|\le\tfrac{\varepsilon}{2}<\varepsilon t k ( t ) − k ( 0 ) − ⟨ ∇ χ ( μ ^ ) , ∇ a ψ ⟩ μ ^ ≤ 2 ε < ε . As ε \varepsilon ε was arbitrary, with the positive number θ c − 1 \theta\,c^{-1} θ c − 1 in the role of the radius of Derivative at an Interior Point and 0 + t = t ∈ ( − t 0 , t 0 ) 0+t=t\in(-t_{0},t_{0}) 0 + t = t ∈ ( − t 0 , t 0 ) , this is the differentiability of k k k at 0 0 0 with the stated derivative.
Step 3 (the penalised function along the curve). Let h : ( − t 0 , t 0 ) → R h:(-t_{0},t_{0})\to\mathbb{R} h : ( − t 0 , t 0 ) → R be h = k + ( s δ ) e h=k+(s\delta)e h = k + ( sδ ) e , with value h ( t ) = χ ( ν t ) + s δ E ( ν t ) h(t)=\chi(\nu_{t})+s\delta\,\mathcal{E}(\nu_{t}) h ( t ) = χ ( ν t ) + sδ E ( ν t ) ; in particular h ( 0 ) = χ ( μ ^ ) + s δ E ( μ ^ ) h(0)=\chi(\hat{\mu})+s\delta\,\mathcal{E}(\hat{\mu}) h ( 0 ) = χ ( μ ^ ) + sδ E ( μ ^ ) . By claim 2 of Sum, Constant Multiple, and Product Rules for One-Dimensional Derivatives , h h h is differentiable at 0 0 0 with
h ′ ( 0 ) = ⟨ ∇ χ ( μ ^ ) , ∇ a ψ ⟩ μ ^ + s δ ⟨ Σ ( μ ^ ) , ∇ a ψ ⟩ μ ^ = ⟨ ∇ χ ( μ ^ ) + s δ Σ ( μ ^ ) , ∇ a ψ ⟩ μ ^ , h'(0)=\langle\nabla\chi(\hat{\mu}),\nabla_{a}\psi\rangle_{\hat{\mu}}+s\delta\,\langle\Sigma(\hat{\mu}),\nabla_{a}\psi\rangle_{\hat{\mu}}=\bigl\langle\nabla\chi(\hat{\mu})+s\delta\,\Sigma(\hat{\mu}),\nabla_{a}\psi\bigr\rangle_{\hat{\mu}}, h ′ ( 0 ) = ⟨ ∇ χ ( μ ^ ) , ∇ a ψ ⟩ μ ^ + sδ ⟨ Σ ( μ ^ ) , ∇ a ψ ⟩ μ ^ = ⟨ ∇ χ ( μ ^ ) + sδ Σ ( μ ^ ) , ∇ a ψ ⟩ μ ^ ,
the second equality by the additivity and homogeneity of the inner product in its first argument (Real Inner Product Space §inner-product ).
Let r r r be the lesser of t 0 t_{0} t 0 and R c − 1 R\,c^{-1} R c − 1 , a positive real number. Let t ∈ ( − t 0 , t 0 ) t\in(-t_{0},t_{0}) t ∈ ( − t 0 , t 0 ) satisfy ∣ 0 − t ∣ < r |0-t|<r ∣0 − t ∣ < r . Then ν t ∈ D \nu_{t}\in\mathcal{D} ν t ∈ D by Step 2, and by (1), W a ( μ ^ , ν t ) ≤ ∣ t ∣ n ψ ≤ ∣ t ∣ c < r c ≤ R W_{a}(\hat{\mu},\nu_{t})\le|t|\,n_{\psi}\le|t|\,c<r\,c\le R W a ( μ ^ , ν t ) ≤ ∣ t ∣ n ψ ≤ ∣ t ∣ c < r c ≤ R . So ( ∗ ) (\ast) ( ∗ ) applies to μ = ν t \mu=\nu_{t} μ = ν t and gives h ( t ) ≤ h ( 0 ) h(t)\le h(0) h ( t ) ≤ h ( 0 ) if s = − 1 s=-1 s = − 1 and h ( t ) ≥ h ( 0 ) h(t)\ge h(0) h ( t ) ≥ h ( 0 ) if s = 1 s=1 s = 1 . Since the metric of The Absolute Value Metric on the Real Line gives d R ( 0 , t ) = ∣ 0 − t ∣ d_{\mathbb{R}}(0,t)=|0-t| d R ( 0 , t ) = ∣0 − t ∣ , the function h h h has a local maximum at 0 0 0 relative to ( − t 0 , t 0 ) (-t_{0},t_{0}) ( − t 0 , t 0 ) if s = − 1 s=-1 s = − 1 , and a local minimum there if s = 1 s=1 s = 1 . By Vanishing of the Derivative at an Interior Local Extremum , applied with p = − t 0 p=-t_{0} p = − t 0 , q = t 0 q=t_{0} q = t 0 , g = h g=h g = h and the point 0 ∈ ( − t 0 , t 0 ) 0\in(-t_{0},t_{0}) 0 ∈ ( − t 0 , t 0 ) , h ′ ( 0 ) = 0 h'(0)=0 h ′ ( 0 ) = 0 . As ψ \psi ψ was arbitrary,
⟨ v , ∇ a ψ ⟩ μ ^ = 0 for every ψ ∈ F C b 1 ( X ) , where v = ∇ χ ( μ ^ ) + s δ Σ ( μ ^ ) . (2) \bigl\langle v,\nabla_{a}\psi\bigr\rangle_{\hat{\mu}}=0\qquad\text{for every }\psi\in\mathcal{F}C^{1}_{b}(X),\qquad\text{where }v=\nabla\chi(\hat{\mu})+s\delta\,\Sigma(\hat{\mu}). \tag{2} ⟨ v , ∇ a ψ ⟩ μ ^ = 0 for every ψ ∈ F C b 1 ( X ) , where v = ∇ χ ( μ ^ ) + sδ Σ ( μ ^ ) . ( 2 )
Step 4 (vanishing in the tangent space). The fields ∇ χ ( μ ^ ) \nabla\chi(\hat{\mu}) ∇ χ ( μ ^ ) and Σ ( μ ^ ) \Sigma(\hat{\mu}) Σ ( μ ^ ) lie in T μ ^ a T^{a}_{\hat{\mu}} T μ ^ a , by property (b) of Noise Intrinsic Test Functions on the Noise Wasserstein Space §test and by Noise Penalty Pairs on the Noise Wasserstein Space: the Penalty, Its Score, and Their Domains §pair , and T μ ^ a T^{a}_{\hat{\mu}} T μ ^ a is a linear subspace of L 2 ( μ ^ ; X a ) L^{2}(\hat{\mu};X^{a}) L 2 ( μ ^ ; X a ) by Linearity of the Noise Gradient, and the Noise Tangent Space is a Closed Linear Subspace §subspace ; hence v ∈ T μ ^ a v\in T^{a}_{\hat{\mu}} v ∈ T μ ^ a by Linear Subspace . Every element of G μ ^ a G^{a}_{\hat{\mu}} G μ ^ a is the class of ∇ a ψ \nabla_{a}\psi ∇ a ψ for some ψ ∈ F C b 1 ( X ) \psi\in\mathcal{F}C^{1}_{b}(X) ψ ∈ F C b 1 ( X ) (The Noise Gradient of a Cylindrical Function and the Noise Tangent Space at a Probability Measure on a Hilbert Space §gradients ), so by (2) and the symmetry of the inner product (Real Inner Product Space §inner-product ), ⟨ g , v ⟩ μ ^ = 0 \langle g,v\rangle_{\hat{\mu}}=0 ⟨ g , v ⟩ μ ^ = 0 for every g ∈ G μ ^ a g\in G^{a}_{\hat{\mu}} g ∈ G μ ^ a .
Suppose, for a contradiction, that ∥ v ∥ μ ^ ≠ 0 \lVert v\rVert_{\hat{\mu}}\ne0 ∥ v ∥ μ ^ = 0 , so ∥ v ∥ μ ^ > 0 \lVert v\rVert_{\hat{\mu}}>0 ∥ v ∥ μ ^ > 0 . By The Noise Gradient of a Cylindrical Function and the Noise Tangent Space at a Probability Measure on a Hilbert Space §tangent , T μ ^ a T^{a}_{\hat{\mu}} T μ ^ a is the closure of G μ ^ a G^{a}_{\hat{\mu}} G μ ^ a in the metric space ( L 2 ( μ ^ ; X a ) , d ) (L^{2}(\hat{\mu};X^{a}),d) ( L 2 ( μ ^ ; X a ) , d ) of the distance d ( ξ , ζ ) = ∥ ξ − ζ ∥ μ ^ d(\xi,\zeta)=\lVert\xi-\zeta\rVert_{\hat{\mu}} d ( ξ , ζ ) = ∥ ξ − ζ ∥ μ ^ (Real Hilbert Space §topology , Real Inner Product Space §distance ). By Sequential Characterization of the Closure in a Metric Space there is a sequence in G μ ^ a G^{a}_{\hat{\mu}} G μ ^ a converging to v v v , so by Convergent Sequence in a Metric Space there is g ∈ G μ ^ a g\in G^{a}_{\hat{\mu}} g ∈ G μ ^ a with d ( v , g ) < 1 2 ∥ v ∥ μ ^ d(v,g)<\tfrac12\lVert v\rVert_{\hat{\mu}} d ( v , g ) < 2 1 ∥ v ∥ μ ^ . By The Norm Metric of a Real Inner Product Space: Triangle Inequalities, Limits and Continuity §lipschitz , the map ξ ↦ ⟨ ξ , v ⟩ μ ^ \xi\mapsto\langle\xi,v\rangle_{\hat{\mu}} ξ ↦ ⟨ ξ , v ⟩ μ ^ is Lipschitz with constant ∥ v ∥ μ ^ \lVert v\rVert_{\hat{\mu}} ∥ v ∥ μ ^ , so
∥ v ∥ μ ^ 2 = ∣ ⟨ v , v ⟩ μ ^ − ⟨ g , v ⟩ μ ^ ∣ ≤ ∥ v ∥ μ ^ d ( v , g ) < 1 2 ∥ v ∥ μ ^ 2 , \lVert v\rVert_{\hat{\mu}}^{2}=\bigl|\langle v,v\rangle_{\hat{\mu}}-\langle g,v\rangle_{\hat{\mu}}\bigr|\le\lVert v\rVert_{\hat{\mu}}\,d(v,g)<\tfrac12\lVert v\rVert_{\hat{\mu}}^{2}, ∥ v ∥ μ ^ 2 = ⟨ v , v ⟩ μ ^ − ⟨ g , v ⟩ μ ^ ≤ ∥ v ∥ μ ^ d ( v , g ) < 2 1 ∥ v ∥ μ ^ 2 ,
using ⟨ v , v ⟩ μ ^ = ∥ v ∥ μ ^ 2 \langle v,v\rangle_{\hat{\mu}}=\lVert v\rVert_{\hat{\mu}}^{2} ⟨ v , v ⟩ μ ^ = ∥ v ∥ μ ^ 2 (Real Inner Product Space §norm ) and ⟨ g , v ⟩ μ ^ = 0 \langle g,v\rangle_{\hat{\mu}}=0 ⟨ g , v ⟩ μ ^ = 0 . This contradicts ∥ v ∥ μ ^ > 0 \lVert v\rVert_{\hat{\mu}}>0 ∥ v ∥ μ ^ > 0 . Hence ∥ v ∥ μ ^ = 0 \lVert v\rVert_{\hat{\mu}}=0 ∥ v ∥ μ ^ = 0 , and v v v is the zero vector of L 2 ( μ ^ ; X a ) L^{2}(\hat{\mu};X^{a}) L 2 ( μ ^ ; X a ) by Elementary Identities in a Real Inner Product Space §vanishing .
Step 5 (conclusion). Adding − s δ Σ ( μ ^ ) -s\delta\,\Sigma(\hat{\mu}) − sδ Σ ( μ ^ ) to both sides of ∇ χ ( μ ^ ) + s δ Σ ( μ ^ ) = 0 \nabla\chi(\hat{\mu})+s\delta\,\Sigma(\hat{\mu})=0 ∇ χ ( μ ^ ) + sδ Σ ( μ ^ ) = 0 in the vector space L 2 ( μ ^ ; X a ) L^{2}(\hat{\mu};X^{a}) L 2 ( μ ^ ; X a ) gives ∇ χ ( μ ^ ) = ( − s δ ) Σ ( μ ^ ) \nabla\chi(\hat{\mu})=(-s\delta)\,\Sigma(\hat{\mu}) ∇ χ ( μ ^ ) = ( − sδ ) Σ ( μ ^ ) . For s = − 1 s=-1 s = − 1 this is ∇ χ ( μ ^ ) = δ Σ ( μ ^ ) \nabla\chi(\hat{\mu})=\delta\,\Sigma(\hat{\mu}) ∇ χ ( μ ^ ) = δ Σ ( μ ^ ) , which is claim 1. For s = 1 s=1 s = 1 it is ∇ χ ( μ ^ ) = ( − δ ) Σ ( μ ^ ) \nabla\chi(\hat{\mu})=(-\delta)\,\Sigma(\hat{\mu}) ∇ χ ( μ ^ ) = ( − δ ) Σ ( μ ^ ) , and ( − δ ) Σ ( μ ^ ) = ( − 1 ) ( δ Σ ( μ ^ ) ) = − ( δ Σ ( μ ^ ) ) (-\delta)\,\Sigma(\hat{\mu})=(-1)\bigl(\delta\,\Sigma(\hat{\mu})\bigr)=-\bigl(\delta\,\Sigma(\hat{\mu})\bigr) ( − δ ) Σ ( μ ^ ) = ( − 1 ) ( δ Σ ( μ ^ ) ) = − ( δ Σ ( μ ^ ) ) by the scalar-multiplication axioms and claim 5 of Elementary Identities in a Vector Space ; this is claim 2. ■ \blacksquare ■