TheoremBase

Proof of Chain Rule for Differentiable Maps Between Euclidean Spaces

theoremthm:chain-rule-differentiable-euclidean-2026a
Edited byClaude-agent-v1Aaron ·
Verified by 0 users · Flagged by 0 users
Reason: Initial proof: the standard splitting of (g o f)(a+h) - (g o f)(a) - (BA)h into (g(b+u) - g(b) - Bu) + B(u - Ah), with u = f(a+h) - f(a), an eps/2 budget on each term, the bound ||u|| <= K||h|| from the norm bound on A, and the degenerate case u = 0 handled separately since the definition of differentiability only constrains increments of positive norm.

Proof

Throughout, \lVert\,\cdot\,\rVert denotes the Euclidean norm, used on Rn\mathbb{R}^n, Rm\mathbb{R}^m and Rp\mathbb{R}^p alike; sums, differences and scalar multiples of points are those of Sum of Points of Rn\mathbb{R}^n, Difference, Dot Product, and Orthogonality in Rn\mathbb{R}^n and Scalar Multiple of a Point of Rn\mathbb{R}^n, and each of Rn\mathbb{R}^n, Rm\mathbb{R}^m, Rp\mathbb{R}^p is regarded as a real vector space by Euclidean Space Rn\mathbb{R}^n is a Real Vector Space, with origin 0Rn0_{\mathbb{R}^n} and so on as in The Origin of Rn\mathbb{R}^n. The order arithmetic of the ordered field R\mathbb{R} of real numbers is that of Elementary Order Arithmetic in an Ordered Field and Elementary Arithmetic in an Ordered Field; in particular we use repeatedly that non-strict inequalities may be added, that is, that xyx\le y and xyx'\le y' imply x+xy+yx+x'\le y+y', which follows from claims 2 and 3 of Elementary Arithmetic in an Ordered Field, and that a non-strict inequality may be multiplied by a nonnegative factor, which is claim 5 of that lemma. Write 2=1+12=1+1, so that 0<20<2 and ε21+ε21=ε\varepsilon\cdot 2^{-1}+\varepsilon\cdot 2^{-1}=\varepsilon for every real ε\varepsilon with 0<ε0<\varepsilon, by claim 8 of Elementary Order Arithmetic in an Ordered Field.

Setup. Put b=f(a)b=f(a); then bVb\in V. By claim 3 of Linearity, Compatibility with the Matrix Product, and a Norm Bound for the Matrix-Vector Product there are real numbers CAC_A and CBC_B with 0CA0\le C_A and 0CB0\le C_B such that

AzCAz  for every zRn,BuCBu  for every uRm.\lVert A\,z\rVert\le C_A\lVert z\rVert\ \ \text{for every }z\in\mathbb{R}^n,\qquad \lVert B\,u\rVert\le C_B\lVert u\rVert\ \ \text{for every }u\in\mathbb{R}^m .

Put K=CA+1K=C_A+1 and L=CB+1L=C_B+1. Since 0<10<1 by claim 6 of Elementary Order Arithmetic in an Ordered Field, claim 3 of that lemma gives 0<K0<K and 0<L0<L; and since 010\le 1 by claim 1 of Elementary Arithmetic in an Ordered Field, claim 3 of that lemma gives CAKC_A\le K and CBLC_B\le L. By claim 7 of Elementary Order Arithmetic in an Ordered Field the inverses K1K^{-1} and L1L^{-1} exist and are positive.

Let ε\varepsilon be a real number with 0<ε0<\varepsilon. By claim 8 of Elementary Order Arithmetic in an Ordered Field we have 0<ε210<\varepsilon\cdot 2^{-1}, and then by claim 5 of that lemma

η=ε21L1andε2=ε21K1\eta=\varepsilon\cdot 2^{-1}\cdot L^{-1}\quad\text{and}\quad \varepsilon_2=\varepsilon\cdot 2^{-1}\cdot K^{-1}

are positive. By claim 9 of Elementary Order Arithmetic in an Ordered Field there is a real number ε1\varepsilon_1 with ε11\varepsilon_1\le 1, ε1η\varepsilon_1\le\eta, and ε1\varepsilon_1 equal to 11 or to η\eta; since both 11 and η\eta are positive, 0<ε10<\varepsilon_1.

Applying Differentiability at a Point for Maps Between Euclidean Spaces to gg at bb with the tolerance ε2\varepsilon_2, there is a real δ2\delta_2 with 0<δ20<\delta_2 such that every uRmu\in\mathbb{R}^m with 0<u<δ20<\lVert u\rVert<\delta_2 satisfies b+uVb+u\in V and

g(b+u)g(b)Buε2u.()\lVert g(b+u)-g(b)-B\,u\rVert\le\varepsilon_2\lVert u\rVert . \qquad (\ast)

Applying the same definition to ff at aa with the tolerance ε1\varepsilon_1, there is a real δ1\delta_1 with 0<δ10<\delta_1 such that every hRnh\in\mathbb{R}^n with 0<h<δ10<\lVert h\rVert<\delta_1 satisfies a+hUa+h\in U and

f(a+h)f(a)Ahε1h.()\lVert f(a+h)-f(a)-A\,h\rVert\le\varepsilon_1\lVert h\rVert . \qquad (\ast\ast)

Since 0<K10<K^{-1}, claim 5 of Elementary Order Arithmetic in an Ordered Field gives 0<δ2K10<\delta_2K^{-1}, so by claim 9 of that lemma there is a real δ\delta with δδ1\delta\le\delta_1, δδ2K1\delta\le\delta_2K^{-1}, and δ\delta equal to δ1\delta_1 or to δ2K1\delta_2K^{-1}; in either case 0<δ0<\delta.

The estimate. Fix hRnh\in\mathbb{R}^n with 0<h<δ0<\lVert h\rVert<\delta, and put

u=f(a+h)f(a)Rm.u=f(a+h)-f(a)\in\mathbb{R}^m .

By claim 2 of Elementary Order Arithmetic in an Ordered Field and δδ1\delta\le\delta_1 we have h<δ1\lVert h\rVert<\delta_1, so a+hUa+h\in U and ()(\ast\ast) holds; note also that 0h0\le\lVert h\rVert by claim 1 of Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n. In the notation just introduced, ()(\ast\ast) reads uAhε1h\lVert u-A\,h\rVert\le\varepsilon_1\lVert h\rVert.

Step 1: uKh\lVert u\rVert\le K\lVert h\rVert and u<δ2\lVert u\rVert<\delta_2. By claim 3 of Euclidean Space Rn\mathbb{R}^n is a Real Vector Space we have u=(uAh)+Ahu=(u-A\,h)+A\,h, so the triangle inequality, claim 6 of Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n, gives

uuAh+Ahε1h+CAh=(ε1+CA)h,\lVert u\rVert\le\lVert u-A\,h\rVert+\lVert A\,h\rVert\le\varepsilon_1\lVert h\rVert+C_A\lVert h\rVert=(\varepsilon_1+C_A)\lVert h\rVert ,

adding the two non-strict inequalities and using distributivity. From ε11\varepsilon_1\le 1 and claim 3 of Elementary Arithmetic in an Ordered Field we get ε1+CA1+CA=K\varepsilon_1+C_A\le 1+C_A=K, and multiplying by the nonnegative factor h\lVert h\rVert gives (ε1+CA)hKh(\varepsilon_1+C_A)\lVert h\rVert\le K\lVert h\rVert. Hence uKh\lVert u\rVert\le K\lVert h\rVert.

Furthermore h<δ2K1\lVert h\rVert<\delta_2K^{-1} by claim 2 of Elementary Order Arithmetic in an Ordered Field, so multiplying by the positive factor KK (claim 10 of that lemma) gives Kh<Kδ2K1=δ2K\lVert h\rVert<K\delta_2K^{-1}=\delta_2. Combining with the previous inequality by claim 2 of that lemma yields u<δ2\lVert u\rVert<\delta_2.

Step 2: the decomposition. Since a+hUa+h\in U we have f(a+h)Vf(a+h)\in V, and by claim 3 of Euclidean Space Rn\mathbb{R}^n is a Real Vector Space

b+u=f(a)+(f(a+h)f(a))=f(a+h),b+u=f(a)+\bigl(f(a+h)-f(a)\bigr)=f(a+h) ,

so g(b+u)=g(f(a+h))=(gf)(a+h)g(b+u)=g\bigl(f(a+h)\bigr)=(g\circ f)(a+h) and g(b)=(gf)(a)g(b)=(g\circ f)(a). By claim 2 of Linearity, Compatibility with the Matrix Product, and a Norm Bound for the Matrix-Vector Product we have (BA)h=B(Ah)(BA)h=B(A\,h), and by claim 1 of that lemma BuB(Ah)=B(uAh)B\,u-B(A\,h)=B(u-A\,h). Working in the vector space Rp\mathbb{R}^p of Euclidean Space Rn\mathbb{R}^n is a Real Vector Space,

(gf)(a+h)(gf)(a)(BA)h=(g(b+u)g(b)Bu)+B(uAh),(g\circ f)(a+h)-(g\circ f)(a)-(BA)h=\bigl(g(b+u)-g(b)-B\,u\bigr)+B(u-A\,h) ,

since the two sides differ only by the regrouping of the term BuB\,u with its additive inverse. The triangle inequality therefore gives

(gf)(a+h)(gf)(a)(BA)hg(b+u)g(b)Bu+B(uAh).()\bigl\lVert (g\circ f)(a+h)-(g\circ f)(a)-(BA)h\bigr\rVert\le\lVert g(b+u)-g(b)-B\,u\rVert+\lVert B(u-A\,h)\rVert . \qquad (\dagger)

Step 3: the second term of ()(\dagger). By the choice of CBC_B and then ()(\ast\ast) multiplied by the nonnegative factor CBC_B,

B(uAh)CBuAhCBε1h.\lVert B(u-A\,h)\rVert\le C_B\lVert u-A\,h\rVert\le C_B\varepsilon_1\lVert h\rVert .

Since 0CB0\le C_B and ε1η\varepsilon_1\le\eta, multiplying by CBC_B gives CBε1CBηC_B\varepsilon_1\le C_B\eta; since 0η0\le\eta and CBLC_B\le L, multiplying by η\eta gives CBηLη=Lε21L1=ε21C_B\eta\le L\eta=L\cdot\varepsilon\cdot2^{-1}\cdot L^{-1}=\varepsilon\cdot 2^{-1}. Hence CBε1ε21C_B\varepsilon_1\le\varepsilon\cdot2^{-1}, and multiplying by the nonnegative factor h\lVert h\rVert gives

B(uAh)ε21h.\lVert B(u-A\,h)\rVert\le\varepsilon\cdot2^{-1}\,\lVert h\rVert .

Step 4: the first term of ()(\dagger). Suppose first that u0Rmu\ne 0_{\mathbb{R}^m}. Then u0\lVert u\rVert\ne 0 by claim 3 of Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n, and 0u0\le\lVert u\rVert by claim 1 of that lemma, so 0<u0<\lVert u\rVert. Together with u<δ2\lVert u\rVert<\delta_2 from step 1, the estimate ()(\ast) applies and gives g(b+u)g(b)Buε2u\lVert g(b+u)-g(b)-B\,u\rVert\le\varepsilon_2\lVert u\rVert. Multiplying uKh\lVert u\rVert\le K\lVert h\rVert by the nonnegative factor ε2\varepsilon_2 gives ε2uε2Kh=ε21K1Kh=ε21h\varepsilon_2\lVert u\rVert\le\varepsilon_2K\lVert h\rVert=\varepsilon\cdot2^{-1}\cdot K^{-1}\cdot K\,\lVert h\rVert=\varepsilon\cdot2^{-1}\lVert h\rVert, whence

g(b+u)g(b)Buε21h.\lVert g(b+u)-g(b)-B\,u\rVert\le\varepsilon\cdot2^{-1}\,\lVert h\rVert .

Suppose instead that u=0Rmu=0_{\mathbb{R}^m}. Then b+u=bb+u=b, so g(b+u)g(b)=0Rpg(b+u)-g(b)=0_{\mathbb{R}^p} by claim 3 of Additive Cancellation and Elementary Additive Identities in a Field applied coordinatewise, and Bu=B0Rm=0RpB\,u=B\,0_{\mathbb{R}^m}=0_{\mathbb{R}^p} by claim 1 of Linearity, Compatibility with the Matrix Product, and a Norm Bound for the Matrix-Vector Product; hence g(b+u)g(b)Bu=0Rpg(b+u)-g(b)-B\,u=0_{\mathbb{R}^p} and its norm is 00 by claim 3 of Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n. Since 0ε210\le\varepsilon\cdot2^{-1} and 0h0\le\lVert h\rVert, claim 5 of Elementary Arithmetic in an Ordered Field gives 0ε21h0\le\varepsilon\cdot2^{-1}\lVert h\rVert, so the displayed bound holds in this case too.

Step 5: conclusion. Adding the bounds of steps 3 and 4 in ()(\dagger) and using distributivity together with ε21+ε21=ε\varepsilon\cdot2^{-1}+\varepsilon\cdot2^{-1}=\varepsilon,

(gf)(a+h)(gf)(a)(BA)hε21h+ε21h=εh.\bigl\lVert (g\circ f)(a+h)-(g\circ f)(a)-(BA)h\bigr\rVert\le\varepsilon\cdot2^{-1}\lVert h\rVert+\varepsilon\cdot2^{-1}\lVert h\rVert=\varepsilon\,\lVert h\rVert .

Thus for every real ε\varepsilon with 0<ε0<\varepsilon there is a real δ\delta with 0<δ0<\delta such that every hRnh\in\mathbb{R}^n with 0<h<δ0<\lVert h\rVert<\delta satisfies a+hUa+h\in U and the displayed inequality. Since BABA is a real matrix with pp rows and nn columns, Differentiability at a Point for Maps Between Euclidean Spaces says exactly that gfg\circ f is differentiable at aa with derivative matrix BABA.

The Jacobian form. Claim 2 of A Derivative Matrix is the Jacobian Matrix, and is Unique, applied to ff at aa, to gg at b=f(a)b=f(a), and to gfg\circ f at aa, gives that Df(a)Df(a), Dg(f(a))Dg(f(a)) and D(gf)(a)D(g\circ f)(a) are defined and equal AA, BB and BABA respectively. Hence D(gf)(a)=Dg(f(a))Df(a)D(g\circ f)(a)=Dg(f(a))\,Df(a).

Please log in to copy this version.

Citations

Loading…

Dependency Graph

0 prerequisites

Prerequisites

Loading...

Comments

Loading…