Throughout, ∥ ⋅ ∥ \lVert\,\cdot\,\rVert ∥ ⋅ ∥ denotes the Euclidean norm , used on R n \mathbb{R}^n R n , R m \mathbb{R}^m R m and R p \mathbb{R}^p R p alike; sums, differences and scalar multiples of points are those of Sum of Points of R n \mathbb{R}^n R n , Difference, Dot Product, and Orthogonality in R n \mathbb{R}^n R n and Scalar Multiple of a Point of R n \mathbb{R}^n R n , and each of R n \mathbb{R}^n R n , R m \mathbb{R}^m R m , R p \mathbb{R}^p R p is regarded as a real vector space by Euclidean Space R n \mathbb{R}^n R n is a Real Vector Space , with origin 0 R n 0_{\mathbb{R}^n} 0 R n and so on as in The Origin of R n \mathbb{R}^n R n . The order arithmetic of the ordered field R \mathbb{R} R of real numbers is that of Elementary Order Arithmetic in an Ordered Field and Elementary Arithmetic in an Ordered Field ; in particular we use repeatedly that non-strict inequalities may be added, that is, that x ≤ y x\le y x ≤ y and x ′ ≤ y ′ x'\le y' x ′ ≤ y ′ imply x + x ′ ≤ y + y ′ x+x'\le y+y' x + x ′ ≤ y + y ′ , which follows from claims 2 and 3 of Elementary Arithmetic in an Ordered Field , and that a non-strict inequality may be multiplied by a nonnegative factor, which is claim 5 of that lemma. Write 2 = 1 + 1 2=1+1 2 = 1 + 1 , so that 0 < 2 0<2 0 < 2 and ε ⋅ 2 − 1 + ε ⋅ 2 − 1 = ε \varepsilon\cdot 2^{-1}+\varepsilon\cdot 2^{-1}=\varepsilon ε ⋅ 2 − 1 + ε ⋅ 2 − 1 = ε for every real ε \varepsilon ε with 0 < ε 0<\varepsilon 0 < ε , by claim 8 of Elementary Order Arithmetic in an Ordered Field .
Setup. Put b = f ( a ) b=f(a) b = f ( a ) ; then b ∈ V b\in V b ∈ V . By claim 3 of Linearity, Compatibility with the Matrix Product, and a Norm Bound for the Matrix-Vector Product there are real numbers C A C_A C A and C B C_B C B with 0 ≤ C A 0\le C_A 0 ≤ C A and 0 ≤ C B 0\le C_B 0 ≤ C B such that
∥ A z ∥ ≤ C A ∥ z ∥ for every z ∈ R n , ∥ B u ∥ ≤ C B ∥ u ∥ for every u ∈ R m . \lVert A\,z\rVert\le C_A\lVert z\rVert\ \ \text{for every }z\in\mathbb{R}^n,\qquad \lVert B\,u\rVert\le C_B\lVert u\rVert\ \ \text{for every }u\in\mathbb{R}^m . ∥ A z ∥ ≤ C A ∥ z ∥ for every z ∈ R n , ∥ B u ∥ ≤ C B ∥ u ∥ for every u ∈ R m .
Put K = C A + 1 K=C_A+1 K = C A + 1 and L = C B + 1 L=C_B+1 L = C B + 1 . Since 0 < 1 0<1 0 < 1 by claim 6 of Elementary Order Arithmetic in an Ordered Field , claim 3 of that lemma gives 0 < K 0<K 0 < K and 0 < L 0<L 0 < L ; and since 0 ≤ 1 0\le 1 0 ≤ 1 by claim 1 of Elementary Arithmetic in an Ordered Field , claim 3 of that lemma gives C A ≤ K C_A\le K C A ≤ K and C B ≤ L C_B\le L C B ≤ L . By claim 7 of Elementary Order Arithmetic in an Ordered Field the inverses K − 1 K^{-1} K − 1 and L − 1 L^{-1} L − 1 exist and are positive.
Let ε \varepsilon ε be a real number with 0 < ε 0<\varepsilon 0 < ε . By claim 8 of Elementary Order Arithmetic in an Ordered Field we have 0 < ε ⋅ 2 − 1 0<\varepsilon\cdot 2^{-1} 0 < ε ⋅ 2 − 1 , and then by claim 5 of that lemma
η = ε ⋅ 2 − 1 ⋅ L − 1 and ε 2 = ε ⋅ 2 − 1 ⋅ K − 1 \eta=\varepsilon\cdot 2^{-1}\cdot L^{-1}\quad\text{and}\quad \varepsilon_2=\varepsilon\cdot 2^{-1}\cdot K^{-1} η = ε ⋅ 2 − 1 ⋅ L − 1 and ε 2 = ε ⋅ 2 − 1 ⋅ K − 1
are positive. By claim 9 of Elementary Order Arithmetic in an Ordered Field there is a real number ε 1 \varepsilon_1 ε 1 with ε 1 ≤ 1 \varepsilon_1\le 1 ε 1 ≤ 1 , ε 1 ≤ η \varepsilon_1\le\eta ε 1 ≤ η , and ε 1 \varepsilon_1 ε 1 equal to 1 1 1 or to η \eta η ; since both 1 1 1 and η \eta η are positive, 0 < ε 1 0<\varepsilon_1 0 < ε 1 .
Applying Differentiability at a Point for Maps Between Euclidean Spaces to g g g at b b b with the tolerance ε 2 \varepsilon_2 ε 2 , there is a real δ 2 \delta_2 δ 2 with 0 < δ 2 0<\delta_2 0 < δ 2 such that every u ∈ R m u\in\mathbb{R}^m u ∈ R m with 0 < ∥ u ∥ < δ 2 0<\lVert u\rVert<\delta_2 0 < ∥ u ∥ < δ 2 satisfies b + u ∈ V b+u\in V b + u ∈ V and
∥ g ( b + u ) − g ( b ) − B u ∥ ≤ ε 2 ∥ u ∥ . ( ∗ ) \lVert g(b+u)-g(b)-B\,u\rVert\le\varepsilon_2\lVert u\rVert . \qquad (\ast) ∥ g ( b + u ) − g ( b ) − B u ∥ ≤ ε 2 ∥ u ∥ . ( ∗ )
Applying the same definition to f f f at a a a with the tolerance ε 1 \varepsilon_1 ε 1 , there is a real δ 1 \delta_1 δ 1 with 0 < δ 1 0<\delta_1 0 < δ 1 such that every h ∈ R n h\in\mathbb{R}^n h ∈ R n with 0 < ∥ h ∥ < δ 1 0<\lVert h\rVert<\delta_1 0 < ∥ h ∥ < δ 1 satisfies a + h ∈ U a+h\in U a + h ∈ U and
∥ f ( a + h ) − f ( a ) − A h ∥ ≤ ε 1 ∥ h ∥ . ( ∗ ∗ ) \lVert f(a+h)-f(a)-A\,h\rVert\le\varepsilon_1\lVert h\rVert . \qquad (\ast\ast) ∥ f ( a + h ) − f ( a ) − A h ∥ ≤ ε 1 ∥ h ∥ . ( ∗ ∗ )
Since 0 < K − 1 0<K^{-1} 0 < K − 1 , claim 5 of Elementary Order Arithmetic in an Ordered Field gives 0 < δ 2 K − 1 0<\delta_2K^{-1} 0 < δ 2 K − 1 , so by claim 9 of that lemma there is a real δ \delta δ with δ ≤ δ 1 \delta\le\delta_1 δ ≤ δ 1 , δ ≤ δ 2 K − 1 \delta\le\delta_2K^{-1} δ ≤ δ 2 K − 1 , and δ \delta δ equal to δ 1 \delta_1 δ 1 or to δ 2 K − 1 \delta_2K^{-1} δ 2 K − 1 ; in either case 0 < δ 0<\delta 0 < δ .
The estimate. Fix h ∈ R n h\in\mathbb{R}^n h ∈ R n with 0 < ∥ h ∥ < δ 0<\lVert h\rVert<\delta 0 < ∥ h ∥ < δ , and put
u = f ( a + h ) − f ( a ) ∈ R m . u=f(a+h)-f(a)\in\mathbb{R}^m . u = f ( a + h ) − f ( a ) ∈ R m .
By claim 2 of Elementary Order Arithmetic in an Ordered Field and δ ≤ δ 1 \delta\le\delta_1 δ ≤ δ 1 we have ∥ h ∥ < δ 1 \lVert h\rVert<\delta_1 ∥ h ∥ < δ 1 , so a + h ∈ U a+h\in U a + h ∈ U and ( ∗ ∗ ) (\ast\ast) ( ∗ ∗ ) holds; note also that 0 ≤ ∥ h ∥ 0\le\lVert h\rVert 0 ≤ ∥ h ∥ by claim 1 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n . In the notation just introduced, ( ∗ ∗ ) (\ast\ast) ( ∗ ∗ ) reads ∥ u − A h ∥ ≤ ε 1 ∥ h ∥ \lVert u-A\,h\rVert\le\varepsilon_1\lVert h\rVert ∥ u − A h ∥ ≤ ε 1 ∥ h ∥ .
Step 1: ∥ u ∥ ≤ K ∥ h ∥ \lVert u\rVert\le K\lVert h\rVert ∥ u ∥ ≤ K ∥ h ∥ and ∥ u ∥ < δ 2 \lVert u\rVert<\delta_2 ∥ u ∥ < δ 2 . By claim 3 of Euclidean Space R n \mathbb{R}^n R n is a Real Vector Space we have u = ( u − A h ) + A h u=(u-A\,h)+A\,h u = ( u − A h ) + A h , so the triangle inequality, claim 6 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n , gives
∥ u ∥ ≤ ∥ u − A h ∥ + ∥ A h ∥ ≤ ε 1 ∥ h ∥ + C A ∥ h ∥ = ( ε 1 + C A ) ∥ h ∥ , \lVert u\rVert\le\lVert u-A\,h\rVert+\lVert A\,h\rVert\le\varepsilon_1\lVert h\rVert+C_A\lVert h\rVert=(\varepsilon_1+C_A)\lVert h\rVert , ∥ u ∥ ≤ ∥ u − A h ∥ + ∥ A h ∥ ≤ ε 1 ∥ h ∥ + C A ∥ h ∥ = ( ε 1 + C A ) ∥ h ∥ ,
adding the two non-strict inequalities and using distributivity. From ε 1 ≤ 1 \varepsilon_1\le 1 ε 1 ≤ 1 and claim 3 of Elementary Arithmetic in an Ordered Field we get ε 1 + C A ≤ 1 + C A = K \varepsilon_1+C_A\le 1+C_A=K ε 1 + C A ≤ 1 + C A = K , and multiplying by the nonnegative factor ∥ h ∥ \lVert h\rVert ∥ h ∥ gives ( ε 1 + C A ) ∥ h ∥ ≤ K ∥ h ∥ (\varepsilon_1+C_A)\lVert h\rVert\le K\lVert h\rVert ( ε 1 + C A ) ∥ h ∥ ≤ K ∥ h ∥ . Hence ∥ u ∥ ≤ K ∥ h ∥ \lVert u\rVert\le K\lVert h\rVert ∥ u ∥ ≤ K ∥ h ∥ .
Furthermore ∥ h ∥ < δ 2 K − 1 \lVert h\rVert<\delta_2K^{-1} ∥ h ∥ < δ 2 K − 1 by claim 2 of Elementary Order Arithmetic in an Ordered Field , so multiplying by the positive factor K K K (claim 10 of that lemma) gives K ∥ h ∥ < K δ 2 K − 1 = δ 2 K\lVert h\rVert<K\delta_2K^{-1}=\delta_2 K ∥ h ∥ < K δ 2 K − 1 = δ 2 . Combining with the previous inequality by claim 2 of that lemma yields ∥ u ∥ < δ 2 \lVert u\rVert<\delta_2 ∥ u ∥ < δ 2 .
Step 2: the decomposition. Since a + h ∈ U a+h\in U a + h ∈ U we have f ( a + h ) ∈ V f(a+h)\in V f ( a + h ) ∈ V , and by claim 3 of Euclidean Space R n \mathbb{R}^n R n is a Real Vector Space
b + u = f ( a ) + ( f ( a + h ) − f ( a ) ) = f ( a + h ) , b+u=f(a)+\bigl(f(a+h)-f(a)\bigr)=f(a+h) , b + u = f ( a ) + ( f ( a + h ) − f ( a ) ) = f ( a + h ) ,
so g ( b + u ) = g ( f ( a + h ) ) = ( g ∘ f ) ( a + h ) g(b+u)=g\bigl(f(a+h)\bigr)=(g\circ f)(a+h) g ( b + u ) = g ( f ( a + h ) ) = ( g ∘ f ) ( a + h ) and g ( b ) = ( g ∘ f ) ( a ) g(b)=(g\circ f)(a) g ( b ) = ( g ∘ f ) ( a ) . By claim 2 of Linearity, Compatibility with the Matrix Product, and a Norm Bound for the Matrix-Vector Product we have ( B A ) h = B ( A h ) (BA)h=B(A\,h) ( B A ) h = B ( A h ) , and by claim 1 of that lemma B u − B ( A h ) = B ( u − A h ) B\,u-B(A\,h)=B(u-A\,h) B u − B ( A h ) = B ( u − A h ) . Working in the vector space R p \mathbb{R}^p R p of Euclidean Space R n \mathbb{R}^n R n is a Real Vector Space ,
( g ∘ f ) ( a + h ) − ( g ∘ f ) ( a ) − ( B A ) h = ( g ( b + u ) − g ( b ) − B u ) + B ( u − A h ) , (g\circ f)(a+h)-(g\circ f)(a)-(BA)h=\bigl(g(b+u)-g(b)-B\,u\bigr)+B(u-A\,h) , ( g ∘ f ) ( a + h ) − ( g ∘ f ) ( a ) − ( B A ) h = ( g ( b + u ) − g ( b ) − B u ) + B ( u − A h ) ,
since the two sides differ only by the regrouping of the term B u B\,u B u with its additive inverse. The triangle inequality therefore gives
∥ ( g ∘ f ) ( a + h ) − ( g ∘ f ) ( a ) − ( B A ) h ∥ ≤ ∥ g ( b + u ) − g ( b ) − B u ∥ + ∥ B ( u − A h ) ∥ . ( † ) \bigl\lVert (g\circ f)(a+h)-(g\circ f)(a)-(BA)h\bigr\rVert\le\lVert g(b+u)-g(b)-B\,u\rVert+\lVert B(u-A\,h)\rVert . \qquad (\dagger) ( g ∘ f ) ( a + h ) − ( g ∘ f ) ( a ) − ( B A ) h ≤ ∥ g ( b + u ) − g ( b ) − B u ∥ + ∥ B ( u − A h )∥ . ( † )
Step 3: the second term of ( † ) (\dagger) ( † ) . By the choice of C B C_B C B and then ( ∗ ∗ ) (\ast\ast) ( ∗ ∗ ) multiplied by the nonnegative factor C B C_B C B ,
∥ B ( u − A h ) ∥ ≤ C B ∥ u − A h ∥ ≤ C B ε 1 ∥ h ∥ . \lVert B(u-A\,h)\rVert\le C_B\lVert u-A\,h\rVert\le C_B\varepsilon_1\lVert h\rVert . ∥ B ( u − A h )∥ ≤ C B ∥ u − A h ∥ ≤ C B ε 1 ∥ h ∥ .
Since 0 ≤ C B 0\le C_B 0 ≤ C B and ε 1 ≤ η \varepsilon_1\le\eta ε 1 ≤ η , multiplying by C B C_B C B gives C B ε 1 ≤ C B η C_B\varepsilon_1\le C_B\eta C B ε 1 ≤ C B η ; since 0 ≤ η 0\le\eta 0 ≤ η and C B ≤ L C_B\le L C B ≤ L , multiplying by η \eta η gives C B η ≤ L η = L ⋅ ε ⋅ 2 − 1 ⋅ L − 1 = ε ⋅ 2 − 1 C_B\eta\le L\eta=L\cdot\varepsilon\cdot2^{-1}\cdot L^{-1}=\varepsilon\cdot 2^{-1} C B η ≤ L η = L ⋅ ε ⋅ 2 − 1 ⋅ L − 1 = ε ⋅ 2 − 1 . Hence C B ε 1 ≤ ε ⋅ 2 − 1 C_B\varepsilon_1\le\varepsilon\cdot2^{-1} C B ε 1 ≤ ε ⋅ 2 − 1 , and multiplying by the nonnegative factor ∥ h ∥ \lVert h\rVert ∥ h ∥ gives
∥ B ( u − A h ) ∥ ≤ ε ⋅ 2 − 1 ∥ h ∥ . \lVert B(u-A\,h)\rVert\le\varepsilon\cdot2^{-1}\,\lVert h\rVert . ∥ B ( u − A h )∥ ≤ ε ⋅ 2 − 1 ∥ h ∥ .
Step 4: the first term of ( † ) (\dagger) ( † ) . Suppose first that u ≠ 0 R m u\ne 0_{\mathbb{R}^m} u = 0 R m . Then ∥ u ∥ ≠ 0 \lVert u\rVert\ne 0 ∥ u ∥ = 0 by claim 3 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n , and 0 ≤ ∥ u ∥ 0\le\lVert u\rVert 0 ≤ ∥ u ∥ by claim 1 of that lemma, so 0 < ∥ u ∥ 0<\lVert u\rVert 0 < ∥ u ∥ . Together with ∥ u ∥ < δ 2 \lVert u\rVert<\delta_2 ∥ u ∥ < δ 2 from step 1, the estimate ( ∗ ) (\ast) ( ∗ ) applies and gives ∥ g ( b + u ) − g ( b ) − B u ∥ ≤ ε 2 ∥ u ∥ \lVert g(b+u)-g(b)-B\,u\rVert\le\varepsilon_2\lVert u\rVert ∥ g ( b + u ) − g ( b ) − B u ∥ ≤ ε 2 ∥ u ∥ . Multiplying ∥ u ∥ ≤ K ∥ h ∥ \lVert u\rVert\le K\lVert h\rVert ∥ u ∥ ≤ K ∥ h ∥ by the nonnegative factor ε 2 \varepsilon_2 ε 2 gives ε 2 ∥ u ∥ ≤ ε 2 K ∥ h ∥ = ε ⋅ 2 − 1 ⋅ K − 1 ⋅ K ∥ h ∥ = ε ⋅ 2 − 1 ∥ h ∥ \varepsilon_2\lVert u\rVert\le\varepsilon_2K\lVert h\rVert=\varepsilon\cdot2^{-1}\cdot K^{-1}\cdot K\,\lVert h\rVert=\varepsilon\cdot2^{-1}\lVert h\rVert ε 2 ∥ u ∥ ≤ ε 2 K ∥ h ∥ = ε ⋅ 2 − 1 ⋅ K − 1 ⋅ K ∥ h ∥ = ε ⋅ 2 − 1 ∥ h ∥ , whence
∥ g ( b + u ) − g ( b ) − B u ∥ ≤ ε ⋅ 2 − 1 ∥ h ∥ . \lVert g(b+u)-g(b)-B\,u\rVert\le\varepsilon\cdot2^{-1}\,\lVert h\rVert . ∥ g ( b + u ) − g ( b ) − B u ∥ ≤ ε ⋅ 2 − 1 ∥ h ∥ .
Suppose instead that u = 0 R m u=0_{\mathbb{R}^m} u = 0 R m . Then b + u = b b+u=b b + u = b , so g ( b + u ) − g ( b ) = 0 R p g(b+u)-g(b)=0_{\mathbb{R}^p} g ( b + u ) − g ( b ) = 0 R p by claim 3 of Additive Cancellation and Elementary Additive Identities in a Field applied coordinatewise, and B u = B 0 R m = 0 R p B\,u=B\,0_{\mathbb{R}^m}=0_{\mathbb{R}^p} B u = B 0 R m = 0 R p by claim 1 of Linearity, Compatibility with the Matrix Product, and a Norm Bound for the Matrix-Vector Product ; hence g ( b + u ) − g ( b ) − B u = 0 R p g(b+u)-g(b)-B\,u=0_{\mathbb{R}^p} g ( b + u ) − g ( b ) − B u = 0 R p and its norm is 0 0 0 by claim 3 of Elementary Properties of the Euclidean Norm on R n \mathbb{R}^n R n . Since 0 ≤ ε ⋅ 2 − 1 0\le\varepsilon\cdot2^{-1} 0 ≤ ε ⋅ 2 − 1 and 0 ≤ ∥ h ∥ 0\le\lVert h\rVert 0 ≤ ∥ h ∥ , claim 5 of Elementary Arithmetic in an Ordered Field gives 0 ≤ ε ⋅ 2 − 1 ∥ h ∥ 0\le\varepsilon\cdot2^{-1}\lVert h\rVert 0 ≤ ε ⋅ 2 − 1 ∥ h ∥ , so the displayed bound holds in this case too.
Step 5: conclusion. Adding the bounds of steps 3 and 4 in ( † ) (\dagger) ( † ) and using distributivity together with ε ⋅ 2 − 1 + ε ⋅ 2 − 1 = ε \varepsilon\cdot2^{-1}+\varepsilon\cdot2^{-1}=\varepsilon ε ⋅ 2 − 1 + ε ⋅ 2 − 1 = ε ,
∥ ( g ∘ f ) ( a + h ) − ( g ∘ f ) ( a ) − ( B A ) h ∥ ≤ ε ⋅ 2 − 1 ∥ h ∥ + ε ⋅ 2 − 1 ∥ h ∥ = ε ∥ h ∥ . \bigl\lVert (g\circ f)(a+h)-(g\circ f)(a)-(BA)h\bigr\rVert\le\varepsilon\cdot2^{-1}\lVert h\rVert+\varepsilon\cdot2^{-1}\lVert h\rVert=\varepsilon\,\lVert h\rVert . ( g ∘ f ) ( a + h ) − ( g ∘ f ) ( a ) − ( B A ) h ≤ ε ⋅ 2 − 1 ∥ h ∥ + ε ⋅ 2 − 1 ∥ h ∥ = ε ∥ h ∥ .
Thus for every real ε \varepsilon ε with 0 < ε 0<\varepsilon 0 < ε there is a real δ \delta δ with 0 < δ 0<\delta 0 < δ such that every h ∈ R n h\in\mathbb{R}^n h ∈ R n with 0 < ∥ h ∥ < δ 0<\lVert h\rVert<\delta 0 < ∥ h ∥ < δ satisfies a + h ∈ U a+h\in U a + h ∈ U and the displayed inequality. Since B A BA B A is a real matrix with p p p rows and n n n columns, Differentiability at a Point for Maps Between Euclidean Spaces says exactly that g ∘ f g\circ f g ∘ f is differentiable at a a a with derivative matrix B A BA B A .
The Jacobian form. Claim 2 of A Derivative Matrix is the Jacobian Matrix, and is Unique , applied to f f f at a a a , to g g g at b = f ( a ) b=f(a) b = f ( a ) , and to g ∘ f g\circ f g ∘ f at a a a , gives that D f ( a ) Df(a) D f ( a ) , D g ( f ( a ) ) Dg(f(a)) D g ( f ( a )) and D ( g ∘ f ) ( a ) D(g\circ f)(a) D ( g ∘ f ) ( a ) are defined and equal A A A , B B B and B A BA B A respectively. Hence D ( g ∘ f ) ( a ) = D g ( f ( a ) ) D f ( a ) D(g\circ f)(a)=Dg(f(a))\,Df(a) D ( g ∘ f ) ( a ) = D g ( f ( a )) D f ( a ) .