Reason: First published version. Derives a single basic inequality from the second-order Taylor expansion with Peano remainder along the rays h = tz, reads off the vanishing of the gradient at scale t and the semidefiniteness of the Hessian at scale t^2 by explicit choices of the parameters rather than limits, and reduces the local minimum case to the local maximum case through the difference of the zero function and w.
(b) For all y,h∈Rn, dE(y,y+h)=∥h∥. Indeed (yi−(yi+hi))2=(−hi)⋅(−hi)=hi2 for every i, so both sides are the nonnegative square root of ∑i=1nhi2.
(d) If 0<t and z∈Rn then ∥tz∥=t∥z∥. Indeed by (a) and field arithmetic ∥tz∥2=∑i=1n(tzi)2=t2∑i=1nzi2=t2∥z∥2=(t∥z∥)2. Both ∥tz∥ and t∥z∥ are nonnegative: the first by (a), and for the second, either ∥z∥=0 and then t∥z∥=0, or 0<∥z∥ and then 0<t∥z∥ by claim 5. Claim 3 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field now gives the asserted equality.
(e) If p≤q and 0<c then cp≤cq. Indeed, if p=q this is an equality, and if p<q then cp<cq by claim 10.
(f) If p≤q and r≤s then p+r≤q+s. This follows from the compatibility of ≤ with addition in an ordered field, which gives p+r≤q+r and q+r≤q+s, together with transitivity of ≤.
Step 1 (the basic inequality).Assume that w has a local maximum at x relative to U. Then for every ε∈R with 0<ε and every z∈Rn with z=0Rn there is τ∈R with 0<τ such that
L(z)+tQ(z)≤εt∥z∥2for every t∈R with 0<t<τ.
By the definition of a local maximum relative to U there is δ1∈R with 0<δ1 such that every y∈U with dE(x,y)<δ1 satisfies w(y)≤w(x). Let ε be given with 0<ε. By Second-Order Taylor Expansion with Peano Remainder, applied to w at x, there is δ2∈R with 0<δ2 such that every h∈Rn with ∥h∥<δ2 satisfies x+h∈U and
w(x+h)−w(x)−L(h)−Q(h)≤ε∥h∥2,
where ∣⋅∣ is the absolute value on R; the two sums appearing in that theorem are exactly L(h) and Q(h). By claim 9 let δ3 be the smaller of δ1 and δ2, so 0<δ3.
Let z=0Rn. By (c), 0<∥z∥, so by claim 7 the inverse ∥z∥−1 exists and is positive, and by claim 5 the element τ=δ3∥z∥−1 satisfies 0<τ. Let t∈R with 0<t<τ and put h=tz. By claim 10, ∥z∥t<∥z∥δ3∥z∥−1=δ3, so by (d) we get ∥h∥=t∥z∥<δ3, whence ∥h∥<δ1 and ∥h∥<δ2 by claim 2.
Since ∥h∥<δ2, we have x+h∈U and the displayed Taylor estimate holds for this h. Since dE(x,x+h)=∥h∥<δ1 by (b), the choice of δ1 gives w(x+h)≤w(x), hence w(x+h)−w(x)≤0 by claim 1.
By (d) and (a), ∥h∥2=(t∥z∥)2=t2∥z∥2, while L(h)=tL(z) and Q(h)=t2Q(z). Thus
tL(z)+t2Q(z)≤εt2∥z∥2.
Since 0<t, claim 7 gives 0<t−1, so multiplying by t−1 and simplifying by field arithmetic, using (e), yields L(z)+tQ(z)≤εt∥z∥2. This proves Step 1.
Step 2 (the gradient vanishes). Keep the hypothesis of claim 1 and let z=0Rn. Apply Step 1 with ε=1, which is admissible since 0<1 by claim 6, and let τ be as there. Adding −tQ(z) to both sides of the inequality of Step 1 (compatibility of ≤ with addition in an ordered field) gives
L(z)≤t(∥z∥2−Q(z))for every t with 0<t<τ.
Write M=∥z∥2−Q(z) and suppose, for contradiction, that 0<L(z).
If M≤0, take t=τ⋅2−1, which satisfies 0<t<τ by claim 8. Then tM≤t⋅0=0 by (e), so L(z)≤0 by transitivity of ≤, contradicting 0<L(z).
If 0<M, then M−1 exists and 0<M−1 by claim 7, and 0<L(z)M−1⋅2−1 by claims 5 and 8. By claim 9 let t be the smaller of τ⋅2−1 and L(z)M−1⋅2−1; then 0<t and t≤τ⋅2−1<τ, so t is admissible by claim 2. By (e), tM≤L(z)M−1⋅2−1⋅M=L(z)⋅2−1, and L(z)⋅2−1<L(z) by claim 8. Hence L(z)≤tM≤L(z)⋅2−1<L(z), so L(z)<L(z) by claim 2, which is impossible.
Therefore L(z)≤0 for every z=0Rn; and L(0Rn)=0≤0. So L(z)≤0 for every z∈Rn. Applying this to −z and using L(−z)=−L(z) gives −L(z)≤0, hence 0≤L(z) by claim 4. Since ≤ is a total order and therefore antisymmetric, L(z)=0 for every z∈Rn.
For i∈{1,…,n} let ei∈Rn be the point whose ith coordinate is 1 and whose other coordinates are 0. Then L(ei)=αi, so αi=0. By the definition of the gradient, Dw(x)=(α1,…,αn)=0Rn.
Step 3 (the Hessian is negative semidefinite). Keep the hypothesis of claim 1, let z=0Rn and let ε∈R with 0<ε. Let τ be as in Step 1 for this ε and z, and take t=τ⋅2−1, so 0<t<τ by claim 8. By Step 2 we have L(z)=0, so Step 1 gives tQ(z)≤εt∥z∥2. Multiplying by t−1, which is positive by claim 7, and using (e) and field arithmetic, we get
Q(z)≤ε∥z∥2.
Suppose, for contradiction, that 0<Q(z). By (c), 0<∥z∥2, so (∥z∥2)−1 exists and is positive by claim 7, and ε0=Q(z)(∥z∥2)−1⋅2−1 satisfies 0<ε0 by claims 5 and 8. Applying the previous inequality with ε0 in place of ε gives Q(z)≤Q(z)⋅2−1, while Q(z)⋅2−1<Q(z) by claim 8; by claim 2 this yields Q(z)<Q(z), which is impossible. Hence Q(z)≤0 for every z=0Rn, and Q(0Rn)=0≤0, so Q(z)≤0 for every z∈Rn.
Since Q(z)=2−1z⋅(D2w(x)z) and 0<2 by claim 8, multiplying by 2 and using (e) gives z⋅(D2w(x)z)≤0 for every z∈Rn. By the definition of the matrix-vector product, (0nz)i=∑j=1n0⋅zj=0 for every i, so 0nz=0Rn and, by the definition of the dot product, z⋅(0nz)=0. Hence z⋅(D2w(x)z)≤z⋅(0nz) for every z∈Rn, which by the definition of the positive semidefinite ordering says D2w(x)⪯0n. Together with Step 2 this proves claim 1.
Step 4 (the local minimum case). Assume now that w has a local minimum at x relative to U. Let k0:U→R be the function with constant value 0. By claim 2 of Differences and Constants for Functions of Class C2 on a Euclidean Open Set, k0 is of class C2 on U with Dk0(y)=0Rn and D2k0(y)=0n for every y∈U. By claim 1 of that lemma the function v=k0−w, whose value at y∈U is 0−w(y)=−w(y), is of class C2 on U, and for every y∈U
Dv(y)=0Rn−Dw(y),D2v(y)=0n−D2w(y).
By the definition of a local minimum relative to U there is δ∈R with 0<δ such that every y∈U with dE(x,y)<δ satisfies w(x)≤w(y); by claim 4 this gives −w(y)≤−w(x), that is, v(y)≤v(x). Hence v has a local maximum at x relative to U.
Applying claim 1, already proved, to v in place of w gives Dv(x)=0Rn and D2v(x)⪯0n. By the definition of the difference of points of Rn, the ith coordinate of Dv(x) is 0−∂w/∂xi(x); since it is 0, we get ∂w/∂xi(x)=0 for every i, and hence Dw(x)=0Rn by the definition of the gradient.
Since D2v(x)⪯0n and z⋅(0nz)=0, this gives −(z⋅(D2w(x)z))≤0, hence 0≤z⋅(D2w(x)z) by claim 4, that is, z⋅(0nz)≤z⋅(D2w(x)z) for every z∈Rn. By the definition of the positive semidefinite ordering, 0n⪯D2w(x). This proves claim 2.