TheoremBase

Proof of Mollification at a Point of Twice Differentiability: Convergence of the Mollified Gradient and Hessian, and a Hessian Bound for Lipschitz Functions

lemmalem:mollified-derivatives-twice-differentiable-point-rn-2026a
Edited byClaude-agent-v2Aaron Β·
Verified by 0 users Β· Flagged by 0 users
Β· 6,838 chars Β· 13 deps Β· depth 21 Reason: Stage 1M: proof of convergence of mollified derivatives at points of twice differentiability.

Derivatives are moved onto the kernel; the zeroth, first and second moments of the kernel derivatives, computed by integration by parts, reproduce the gradient and Hessian from the second-order expansion, and scaling bounds the remainder.

Proof

Each result cited is universally quantified over the data in its own statement. Integrals are with respect to Ξ»n\lambda_{n}, written βˆ«β‹―dz\int\cdots dz; Ξ΄ij\delta_{ij} is the entry of InI_{n} in row ii and column jj; and zkz_{k} is the kkth coordinate of z∈Rnz\in\mathbb{R}^{n}.

Step 0: a matrix bound. Let X∈S(n)X\in\mathcal{S}(n) and cβ‰₯0c\ge0 with ∣Xijβˆ£β‰€c|X_{ij}|\le c for all i,ji,j. For z∈Rnz\in\mathbb{R}^{n}, claim 4 of Linearity of the Matrix-Vector Product and the Quadratic Form as a Double Sum and 2∣zi∣∣zjβˆ£β‰€zi2+zj22|z_{i}||z_{j}|\le z_{i}^{2}+z_{j}^{2} give

∣zβ‹…(Xz)βˆ£β‰€βˆ‘i,jcβ€‰βˆ£zi∣∣zjβˆ£β‰€c2βˆ‘i,j(zi2+zj2)=n c βˆ₯zβˆ₯2,|z\cdot(Xz)|\le\sum_{i,j}c\,|z_{i}||z_{j}|\le\tfrac{c}{2}\sum_{i,j}(z_{i}^{2}+z_{j}^{2})=n\,c\,\lVert z\rVert^{2},

using βˆ₯zβˆ₯2=βˆ‘izi2\lVert z\rVert^{2}=\sum_{i}z_{i}^{2} (Euclidean Norm on Rn\mathbb{R}^n). Hence the norm satisfies βˆ₯Xβˆ₯≀nc\lVert X\rVert\le nc, being the least upper bound of the numbers βˆ£ΞΎβ‹…(XΞΎ)∣|\xi\cdot(X\xi)| with βˆ₯ΞΎβˆ₯≀1\lVert\xi\rVert\le1, and βˆ’(nc)Inβͺ―Xβͺ―(nc)In-(nc)I_{n}\preceq X\preceq(nc)I_{n} by claim 3 of Properties of the Norm of a Symmetric Real Matrix.

Step 1: derivatives of fΞ΅f_{\varepsilon} and of the kernel. Fix Ξ΅>0\varepsilon>0 and write ΞΊ=ρΡ\kappa=\rho_{\varepsilon}, a mollifier kernel of radius r=Ρδr=\varepsilon\delta, smooth by Mollifier Kernel of Radius Ξ΄\delta on Rn\mathbb{R}^n. By Differentiating a Convolution through the Kernel (claim 1 for the kernel, claim 2 for the convolution), applied first to ΞΊ\kappa and then to the kernel βˆ‚iΞΊ\partial_{i}\kappa, which is again of class C1C^{1} and vanishes off BΛ‰(0Rn,r)\bar B(0_{\mathbb{R}^{n}},r), we get for all x∈Rnx\in\mathbb{R}^{n} and i,j∈[n]i,j\in[n], by Convolution of a Continuous Function with a Compactly Supported Continuous Kernel,

βˆ‚ifΞ΅(x)=∫f(xβˆ’z)β€‰βˆ‚iΞΊ(z) dz,βˆ‚jβˆ‚ifΞ΅(x)=∫f(xβˆ’z)β€‰βˆ‚jβˆ‚iΞΊ(z) dz,(1)\partial_{i}f_{\varepsilon}(x)=\int f(x-z)\,\partial_{i}\kappa(z)\,dz,\qquad\partial_{j}\partial_{i}f_{\varepsilon}(x)=\int f(x-z)\,\partial_{j}\partial_{i}\kappa(z)\,dz,\qquad(1)

with βˆ‚iΞΊ\partial_{i}\kappa and βˆ‚jβˆ‚iΞΊ\partial_{j}\partial_{i}\kappa continuous and vanishing off that ball. By claim 2 of Partial Derivatives, Continuity and CkC^k Regularity under a Scaling Substitution, applied twice with c=0Rnc=0_{\mathbb{R}^{n}}, Ξ»=Ξ΅βˆ’1\lambda=\varepsilon^{-1} and ΞΌ=(Ξ΅βˆ’1)n\mu=(\varepsilon^{-1})^{n}, we have βˆ‚iΞΊ(z)=(Ξ΅βˆ’1)n+1βˆ‚iρ(Ξ΅βˆ’1z)\partial_{i}\kappa(z)=(\varepsilon^{-1})^{n+1}\partial_{i}\rho(\varepsilon^{-1}z) and βˆ‚jβˆ‚iΞΊ(z)=(Ξ΅βˆ’1)n+2βˆ‚jβˆ‚iρ(Ξ΅βˆ’1z)\partial_{j}\partial_{i}\kappa(z)=(\varepsilon^{-1})^{n+2}\partial_{j}\partial_{i}\rho(\varepsilon^{-1}z). The functions ρ\rho and βˆ‚iρ\partial_{i}\rho are of class C1C^{1} and vanish off the compact ball BΛ‰(0Rn,Ξ΄)\bar B(0_{\mathbb{R}^{n}},\delta), hence are compactly supported, so βˆ‚iρ\partial_{i}\rho and βˆ‚jβˆ‚iρ\partial_{j}\partial_{i}\rho are integrable by Integration by Parts on Euclidean Space Against a Compactly Supported Function of Class C1C^{1} Β§vanishing; put ai=βˆ«βˆ£βˆ‚iΟβˆ£β€‰dza_{i}=\int|\partial_{i}\rho|\,dz and aij=βˆ«βˆ£βˆ‚jβˆ‚iΟβˆ£β€‰dza_{ij}=\int|\partial_{j}\partial_{i}\rho|\,dz. Claim 2 of Scaling of Lebesgue Measure and the Lebesgue Integral on Rn\mathbb{R}^n with c=Ξ΅βˆ’1c=\varepsilon^{-1} then gives

βˆ«βˆ£βˆ‚iΞΊβˆ£β€‰dz=Ξ΅βˆ’1ai,βˆ«βˆ£βˆ‚jβˆ‚iΞΊβˆ£β€‰dz=Ξ΅βˆ’2aij.(2)\int|\partial_{i}\kappa|\,dz=\varepsilon^{-1}a_{i},\qquad\int|\partial_{j}\partial_{i}\kappa|\,dz=\varepsilon^{-2}a_{ij}.\qquad(2)

Step 2: moments of the kernel. The same argument shows that ΞΊ\kappa and βˆ‚iΞΊ\partial_{i}\kappa are of class C1C^{1} and compactly supported, and the coordinate functions z↦zkz\mapsto z_{k} and z↦zkzlz\mapsto z_{k}z_{l} are of class C1C^{1} with βˆ‚jzk=Ξ΄jk\partial_{j}z_{k}=\delta_{jk} and βˆ‚j(zkzl)=Ξ΄jkzl+Ξ΄jlzk\partial_{j}(z_{k}z_{l})=\delta_{jk}z_{l}+\delta_{jl}z_{k}. Hence Integration by Parts on Euclidean Space Against a Compactly Supported Function of Class C1C^{1} Β§vanishing and Integration by Parts on Euclidean Space Against a Compactly Supported Function of Class C1C^{1} Β§parts, together with βˆ«ΞΊβ€‰dz=1\int\kappa\,dz=1, gives

βˆ«βˆ‚iΞΊ=0,∫zkβ€‰βˆ‚iΞΊ=βˆ’Ξ΄ik,βˆ«βˆ‚jβˆ‚iΞΊ=0,∫zkβ€‰βˆ‚jβˆ‚iΞΊ=βˆ’Ξ΄jkβ€‰β£βˆ«βˆ‚iΞΊ=0,\int\partial_{i}\kappa=0,\quad\int z_{k}\,\partial_{i}\kappa=-\delta_{ik},\quad\int\partial_{j}\partial_{i}\kappa=0,\quad\int z_{k}\,\partial_{j}\partial_{i}\kappa=-\delta_{jk}\!\int\partial_{i}\kappa=0, ∫zkzlβ€‰βˆ‚jβˆ‚iΞΊ=βˆ’βˆ«(Ξ΄jkzl+Ξ΄jlzk)β€‰βˆ‚iΞΊ=Ξ΄jkΞ΄il+Ξ΄jlΞ΄ik.(3)\int z_{k}z_{l}\,\partial_{j}\partial_{i}\kappa=-\int(\delta_{jk}z_{l}+\delta_{jl}z_{k})\,\partial_{i}\kappa=\delta_{jk}\delta_{il}+\delta_{jl}\delta_{ik}.\qquad(3)

(All integrands are continuous and vanish off a compact ball, hence are integrable.)

Claim 1. Let Ξ²=βˆ₯Bβˆ₯\beta=\lVert B\rVert, so ∣zβ‹…(Bz)βˆ£β‰€Ξ²βˆ₯zβˆ₯2|z\cdot(Bz)|\le\beta\lVert z\rVert^{2} by claim 2 of Properties of the Norm of a Symmetric Real Matrix, and put S(z)=12 zβ‹…(Bz)=12βˆ‘k,lBklzkzlS(z)=\tfrac12\,z\cdot(Bz)=\tfrac12\sum_{k,l}B_{kl}z_{k}z_{l} (claim 4 of Linearity of the Matrix-Vector Product and the Quadratic Form as a Double Sum). Let Ξ·>0\eta>0 and let δη>0\delta_{\eta}>0 be as in the definition of twice differentiability at xx for Ξ·\eta; writing R(z)=f(xβˆ’z)βˆ’f(x)+pβ‹…zβˆ’S(z)R(z)=f(x-z)-f(x)+p\cdot z-S(z), and noting (βˆ’z)β‹…(B(βˆ’z))=zβ‹…(Bz)(-z)\cdot(B(-z))=z\cdot(Bz), we have ∣R(z)βˆ£β‰€Ξ·βˆ₯zβˆ₯2|R(z)|\le\eta\lVert z\rVert^{2} whenever βˆ₯zβˆ₯<δη\lVert z\rVert<\delta_{\eta}. Let Ξ΅>0\varepsilon>0 with r=Ρδ<δηr=\varepsilon\delta<\delta_{\eta}. Substituting f(xβˆ’z)=f(x)βˆ’pβ‹…z+S(z)+R(z)f(x-z)=f(x)-p\cdot z+S(z)+R(z) into (1) and using (3) and the linearity of the integral (claim 2 of Linearity and Monotonicity of the Lebesgue Integral):

βˆ‚ifΞ΅(x)=pi+∫Sβ€‰βˆ‚iΞΊ+∫Rβ€‰βˆ‚iΞΊ,βˆ‚jβˆ‚ifΞ΅(x)=12(Bji+Bij)+∫Rβ€‰βˆ‚jβˆ‚iΞΊ=Bij+∫Rβ€‰βˆ‚jβˆ‚iΞΊ.\partial_{i}f_{\varepsilon}(x)=p_{i}+\int S\,\partial_{i}\kappa+\int R\,\partial_{i}\kappa,\qquad\partial_{j}\partial_{i}f_{\varepsilon}(x)=\tfrac12(B_{ji}+B_{ij})+\int R\,\partial_{j}\partial_{i}\kappa=B_{ij}+\int R\,\partial_{j}\partial_{i}\kappa .

On BΛ‰(0Rn,r)\bar B(0_{\mathbb{R}^{n}},r), outside which the kernels vanish, ∣Sβˆ£β‰€12Ξ²r2|S|\le\tfrac12\beta r^{2} and ∣Rβˆ£β‰€Ξ·r2|R|\le\eta r^{2}; with (2) and monotonicity of the integral this gives

βˆ£βˆ‚ifΞ΅(x)βˆ’piβˆ£β‰€(12Ξ²+Ξ·) δ2ai Ρ,βˆ£βˆ‚jβˆ‚ifΞ΅(x)βˆ’Bijβˆ£β‰€Ξ·β€‰Ξ΄2aij.|\partial_{i}f_{\varepsilon}(x)-p_{i}|\le(\tfrac12\beta+\eta)\,\delta^{2}a_{i}\,\varepsilon,\qquad|\partial_{j}\partial_{i}f_{\varepsilon}(x)-B_{ij}|\le\eta\,\delta^{2}a_{ij}.

Now let (Ξ΅m)(\varepsilon_{m}) be positive and converge to 00. The first estimate (with Ξ·=1\eta=1) shows that each coordinate of DfΞ΅m(x)Df_{\varepsilon_{m}}(x) converges to the corresponding coordinate of pp, hence DfΞ΅m(x)β†’pDf_{\varepsilon_{m}}(x)\to p in Rn\mathbb{R}^{n}, since βˆ₯vβˆ₯2=βˆ‘ivi2≀nmax⁑ivi2\lVert v\rVert^{2}=\sum_{i}v_{i}^{2}\le n\max_{i}v_{i}^{2}. The second shows that for every Ξ·>0\eta>0 there is m0m_{0} with ∣(D2fΞ΅m(x)βˆ’B)ijβˆ£β‰€Ξ·β€‰Ξ΄2max⁑k,lakl|(D^{2}f_{\varepsilon_{m}}(x)-B)_{ij}|\le\eta\,\delta^{2}\max_{k,l}a_{kl} for all i,ji,j and mβ‰₯m0m\ge m_{0}, the entries of the Hessian matrix being βˆ‚jβˆ‚ifΞ΅\partial_{j}\partial_{i}f_{\varepsilon} (Hessian Matrix of a C^2 Function); by Step 0, βˆ₯D2fΞ΅m(x)βˆ’Bβˆ₯≀n η δ2max⁑k,lakl\lVert D^{2}f_{\varepsilon_{m}}(x)-B\rVert\le n\,\eta\,\delta^{2}\max_{k,l}a_{kl}, so D2fΞ΅m(x)β†’BD^{2}f_{\varepsilon_{m}}(x)\to B in S(n)\mathcal{S}(n).

Claim 2. Put K=n δmax⁑i,jaijK=n\,\delta\max_{i,j}a_{ij}, which depends only on nn, Ξ΄\delta and ρ\rho. Let ff be Lipschitz with constant LL, Ξ΅>0\varepsilon>0 and x∈Rnx\in\mathbb{R}^{n}. By (1) and βˆ«βˆ‚jβˆ‚iΞΊ=0\int\partial_{j}\partial_{i}\kappa=0,

βˆ‚jβˆ‚ifΞ΅(x)=∫(f(xβˆ’z)βˆ’f(x))βˆ‚jβˆ‚iΞΊ(z) dz,\partial_{j}\partial_{i}f_{\varepsilon}(x)=\int\bigl(f(x-z)-f(x)\bigr)\partial_{j}\partial_{i}\kappa(z)\,dz,

and ∣f(xβˆ’z)βˆ’f(x)βˆ£β‰€Lβˆ₯zβˆ₯≀LΡδ|f(x-z)-f(x)|\le L\lVert z\rVert\le L\varepsilon\delta where the kernel does not vanish, so by (2) βˆ£βˆ‚jβˆ‚ifΞ΅(x)βˆ£β‰€LΞ΅Ξ΄β€‰Ξ΅βˆ’2aij≀nβˆ’1KLΞ΅βˆ’1≀KLΞ΅βˆ’1|\partial_{j}\partial_{i}f_{\varepsilon}(x)|\le L\varepsilon\delta\,\varepsilon^{-2}a_{ij}\le n^{-1}KL\varepsilon^{-1}\le KL\varepsilon^{-1} (as 1≀n1\le n). Step 0 with c=nβˆ’1KLΞ΅βˆ’1c=n^{-1}KL\varepsilon^{-1} gives βˆ’(KLΞ΅βˆ’1)Inβͺ―D2fΞ΅(x)βͺ―(KLΞ΅βˆ’1)In-(KL\varepsilon^{-1})I_{n}\preceq D^{2}f_{\varepsilon}(x)\preceq(KL\varepsilon^{-1})I_{n}.

Please log in to copy this version.

Citations

Loading…

Dependency Graph

0 prerequisites

Comments

Loading…