TheoremBase

The Lipschitz bound gives |Dw| <= L at every differentiability point via a difference quotient along the gradient, and the Borel clause of the weak-form lemma (with zero potential and running cost) gives measurability; Rademacher's theorem and absolute continuity make the non-differentiability set mu-null. For semiconvex w, w + K|x|^2/2 is convex with subdifferential {Dw + Kx} on EwE_w, so grad w + K id is tangent and subtracting K id stays in the linear subspace TmuT_mu; the semiconcave case applies this to -w.

Proof

Each result cited below is universally quantified over the data in its own statement. Elementary order and arithmetic in R\mathbb{R}, including rules for finite sums and for squares of nonnegative numbers, are provided by The Real Numbers: Standing Notation and Background §background and are not cited individually. Steps 1 to 4 use only that ww is bounded and Lipschitz with constant LL.

Step 0 (Conventions). By Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n §distance the Euclidean distance is dE(x,y)=∥x−y∥d_{E}(x,y)=\lVert x-y\rVert, and the metric of The Absolute Value Metric on the Real Line is ∣s−t∣|s-t|; so by Lipschitz Map Between Metric Spaces

∣w(x)−w(y)∣≤L ∥x−y∥for all x,y∈Rd.(0.1)|w(x)-w(y)|\le L\,\lVert x-y\rVert\qquad\text{for all }x,y\in\mathbb{R}^{d}.\tag{0.1}

The set Rd\mathbb{R}^{d} is open (every open ball about one of its points lies in it) and convex (it contains every point tx+(1−t)ytx+(1-t)y). For x∈Ewx\in E_{w} let JJ be a derivative matrix of ww at xx; by claim 1 of A Derivative Matrix is the Jacobian Matrix, and is Unique its entries are J1i=∂iw(x)J_{1i}=\partial_{i}w(x), so by the definitions of the matrix-vector product and of the dot product JhJh has the single coordinate Dw(x)⋅hDw(x)\cdot h, and the Euclidean norm of a point of R1\mathbb{R}^{1} is the absolute value of its coordinate by Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n §square. Thus Differentiability at a Point for Maps Between Euclidean Spaces (with U=RdU=\mathbb{R}^{d}) gives: for x∈Ewx\in E_{w} and every positive ε\varepsilon there is a positive δ\delta with

∣w(x+h)−w(x)−Dw(x)⋅h∣≤ε∥h∥whenever 0<∥h∥<δ.(0.2)\bigl|w(x+h)-w(x)-Dw(x)\cdot h\bigr|\le\varepsilon\lVert h\rVert\qquad\text{whenever }0<\lVert h\rVert<\delta .\tag{0.2}

Step 1 (Continuity). Let x∈Rdx\in\mathbb{R}^{d}, let ε\varepsilon be positive and put δ=εL+1\delta=\tfrac{\varepsilon}{L+1}. By (0.1), every y∈Rdy\in\mathbb{R}^{d} with ∥x−y∥<δ\lVert x-y\rVert<\delta satisfies ∣w(y)−w(x)∣≤Lδ=LεL+1<ε|w(y)-w(x)|\le L\delta=\tfrac{L\varepsilon}{L+1}<\varepsilon. So ww is continuous on Rd\mathbb{R}^{d}, in the sense fixed in Differential Calculus and Convexity on Euclidean Open Sets: Standing Notation §extrema.

Step 2 (The gradient bound). Let x∈Ewx\in E_{w} and v=Dw(x)v=Dw(x). If v=0Rdv=0_{\mathbb{R}^{d}}, then ∥v∥=0≤L\lVert v\rVert=0\le L by Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n §vanishing. Otherwise 0<∥v∥0<\lVert v\rVert. Let ε\varepsilon be positive, let δ\delta be as in (0.2), and put t=δ2∥v∥t=\tfrac{\delta}{2\lVert v\rVert} and h=tvh=tv. Then ∥h∥=t∥v∥=δ2\lVert h\rVert=t\lVert v\rVert=\tfrac{\delta}{2} by Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n §homogeneity, and v⋅h=t∥v∥2v\cdot h=t\lVert v\rVert^{2} by Bilinearity and Symmetry of the Dot Product on Rn\mathbb{R}^n and Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n §square. By (0.2) and (0.1),

t∥v∥2=v⋅h≤w(x+h)−w(x)+ε∥h∥≤(L+ε) t∥v∥,t\lVert v\rVert^{2}=v\cdot h\le w(x+h)-w(x)+\varepsilon\lVert h\rVert\le(L+\varepsilon)\,t\lVert v\rVert ,

and dividing by the positive number t∥v∥t\lVert v\rVert gives ∥v∥≤L+ε\lVert v\rVert\le L+\varepsilon. As ε\varepsilon was arbitrary, ∥Dw(x)∥≤L\lVert Dw(x)\rVert\le L. Since ∇w\nabla w equals DwDw on EwE_{w} and the origin, of norm 00, off EwE_{w},

∥∇w(x)∥≤Lfor every x∈Rd.(2.1)\lVert\nabla w(x)\rVert\le L\qquad\text{for every }x\in\mathbb{R}^{d}.\tag{2.1}

Step 3 (Claim 1). We apply Integrating a Semiconvex Viscosity Subsolution of the Penalty-Drift Equation against a Measure of Finite Relative Fisher Information §borel with the following data from its Data paragraph: the open convex set Rd\mathbb{R}^{d} (Step 0) in place of its DD; UU the zero function on Rd\mathbb{R}^{d}, which is smooth by claim 2 of Constants, Coordinate Functions, Sums and Products of CkC^k Functions on a Euclidean Open Set and hence of class C2C^{2} by Smooth Map on a Euclidean Open Set; λ=κ=1\lambda=\kappa=1 and θ′=0\theta'=0; g~\tilde{g} the zero function on Rd\mathbb{R}^{d}, which is bounded, its extension gˉ\bar{g} being the zero function on Rd\mathbb{R}^{d}, which is continuous (being constant) and hence Borel by Probability Measures on Euclidean Space and Random Vectors: Standing Notation §borel-maps; the bounded function ww; and A=L2A=L^{2}, B=0B=0, p0=0p_{0}=0. With these data the set EwE_{w} and the map ∇w\nabla w of that lemma are those of the present statement. Its gradient hypothesis holds: for x∈Ewx\in E_{w}, 0≤∥Dw(x)∥≤L0\le\lVert Dw(x)\rVert\le L by (2.1), so ∥Dw(x)∥2≤L2=A+B(U(x)−p0)\lVert Dw(x)\rVert^{2}\le L^{2}=A+B\bigl(U(x)-p_{0}\bigr). As ww is continuous on Rd\mathbb{R}^{d} (Step 1), the clause gives Ew∈B(Rd)E_{w}\in\mathcal{B}(\mathbb{R}^{d}) and that ∇w\nabla w is Borel. Together with (2.1) this gives the measurability statements and the bound of claim 1. The function x↦∥∇w(x)∥2x\mapsto\lVert\nabla w(x)\rVert^{2} is Borel, as the composition of ∇w\nabla w with the Borel map z↦∥z∥2z\mapsto\lVert z\rVert^{2} (Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pairs and Probability Measures on Euclidean Space and Random Vectors: Standing Notation §borel-maps), and it is bounded by L2L^{2} by (2.1). Let μ∈P2(Rd)\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}). The function x↦∥∇w(x)∥2x\mapsto\lVert\nabla w(x)\rVert^{2} is integrable with respect to μ\mu by Probability Measures on Euclidean Space and Random Vectors: Standing Notation §measures, that is ∫Rd∥∇w∥2 dμ<∞\int_{\mathbb{R}^{d}}\lVert\nabla w\rVert^{2}\,d\mu<\infty, and the class of ∇w\nabla w belongs to L2(μ;Rd)L^{2}(\mu;\mathbb{R}^{d}) by Square-Integrable Vector Fields Against a Probability Measure on Euclidean Space, and Test Functions: Standing Notation §l2mu, for every μ∈P2(Rd)\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}).

Step 4 (Claim 1: the non-differentiability set is null for absolutely continuous measures). Let μ∈P2(Rd)\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}) be absolutely continuous. By Rademacher's Theorem in Rn\mathbb{R}^n §ae (with n=dn=d and f=wf=w, Lipschitz with constant LL by (0.1)) there is N∈B(Rd)N\in\mathcal{B}(\mathbb{R}^{d}) with λd(N)=0\lambda_{d}(N)=0 such that every x∈Rd∖Nx\in\mathbb{R}^{d}\setminus N is a point as in that clause, at which ww is differentiable by Rademacher's Theorem in Rn\mathbb{R}^n §derivative. Hence Rd∖Ew⊆N\mathbb{R}^{d}\setminus E_{w}\subseteq N. The set Rd∖Ew\mathbb{R}^{d}\setminus E_{w} belongs to B(Rd)\mathcal{B}(\mathbb{R}^{d}), a σ\sigma-algebra, by Step 3. Since μ\mu is absolutely continuous and NN is a Borel set with λd(N)=0\lambda_{d}(N)=0, μ(N)=0\mu(N)=0; by claim 2 of Basic Properties of a Measure, μ(Rd∖Ew)≤μ(N)=0\mu(\mathbb{R}^{d}\setminus E_{w})\le\mu(N)=0, and by claim 3 there, as μ(Rd)=1\mu(\mathbb{R}^{d})=1,

μ(Rd∖Ew)=0,μ(Ew)=1.(4.1)\mu(\mathbb{R}^{d}\setminus E_{w})=0,\qquad\mu(E_{w})=1 .\tag{4.1}

As μ\mu was an arbitrary absolutely continuous member of P2(Rd)\mathcal{P}_{2}(\mathbb{R}^{d}), this completes the proof of claim 1.

Step 5 (Claim 2). Suppose that ww is semiconvex with constant KK, fix an absolutely continuous μ∈P2(Rd)\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}), so that (4.1) holds for μ\mu by claim 1 (Step 4), and let ϕ:Rd→R\phi:\mathbb{R}^{d}\to\mathbb{R}, ϕ(x)=w(x)+K2∥x∥2\phi(x)=w(x)+\tfrac{K}{2}\lVert x\rVert^{2}, which is convex on Rd\mathbb{R}^{d} by Semiconvex Function on a Convex Subset of Rn\mathbb{R}^n. Let x∈Ewx\in E_{w}. For h∈Rdh\in\mathbb{R}^{d}, Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n §square and Bilinearity and Symmetry of the Dot Product on Rn\mathbb{R}^n give ∥x+h∥2=∥x∥2+2 x⋅h+∥h∥2\lVert x+h\rVert^{2}=\lVert x\rVert^{2}+2\,x\cdot h+\lVert h\rVert^{2}, hence

ϕ(x+h)−ϕ(x)−(Dw(x)+Kx)⋅h=(w(x+h)−w(x)−Dw(x)⋅h)+K2∥h∥2.\phi(x+h)-\phi(x)-\bigl(Dw(x)+Kx\bigr)\cdot h=\bigl(w(x+h)-w(x)-Dw(x)\cdot h\bigr)+\tfrac{K}{2}\lVert h\rVert^{2}.

Given a positive ε\varepsilon, let δ0\delta_{0} be given by (0.2) for ε2\tfrac{\varepsilon}{2} and let δ\delta be the lesser of δ0\delta_{0} and εK+1\tfrac{\varepsilon}{K+1}. For 0<∥h∥<δ0<\lVert h\rVert<\delta the absolute value of the left-hand side is at most ε2∥h∥+K2⋅εK+1∥h∥≤ε∥h∥\tfrac{\varepsilon}{2}\lVert h\rVert+\tfrac{K}{2}\cdot\tfrac{\varepsilon}{K+1}\lVert h\rVert\le\varepsilon\lVert h\rVert. So ϕ\phi is differentiable at xx with derivative matrix the row whose iith entry is the iith coordinate of Dw(x)+KxDw(x)+Kx, and the subdifferential at a point of differentiability (with U=RdU=\mathbb{R}^{d} and f=ϕf=\phi) gives

∂Rdϕ(x)={Dw(x)+Kx}(x∈Ew),(5.1)\partial_{\mathbb{R}^{d}}\phi(x)=\{Dw(x)+Kx\}\qquad(x\in E_{w}),\tag{5.1}

∂Rd\partial_{\mathbb{R}^{d}} being the subdifferential relative to Rd\mathbb{R}^{d}.

Let id\mathrm{id} be the identity map of Rd\mathbb{R}^{d}; by Basic Properties of the Tangent Space: Closed Subspace, the Identity Map Belongs to It, Second-Moment Limits, and Representation of Bounded Functionals on Gradients §identity it is Borel, its class belongs to L2(μ;Rd)L^{2}(\mu;\mathbb{R}^{d}), and id∈Tμ\mathrm{id}\in T_{\mu}. Let T=∇w+K idT=\nabla w+K\,\mathrm{id}, the pointwise sum. As ∇w\nabla w and id\mathrm{id} are Borel and square-integrable against μ\mu (Step 3), The Space of Square-Integrable Random Vectors §space, applied on the probability space (Rd,B(Rd),μ)(\mathbb{R}^{d},\mathcal{B}(\mathbb{R}^{d}),\mu) as in Square-Integrable Vector Fields Against a Probability Measure on Euclidean Space, and Test Functions: Standing Notation §l2mu, shows that TT is Borel with ∫Rd∥T∥2 dμ<∞\int_{\mathbb{R}^{d}}\lVert T\rVert^{2}\,d\mu<\infty, and by The Space of Square-Integrable Random Vectors §classes its class is ∇w+K id\nabla w+K\,\mathrm{id}. By (5.1), ∂Rdϕ(x)={T(x)}\partial_{\mathbb{R}^{d}}\phi(x)=\{T(x)\} for every x∈Ewx\in E_{w}. We apply A Square-Integrable Selection of the Subdifferential of a Convex Potential Belongs to the Tangent Space §tangent with G=RdG=\mathbb{R}^{d} (open and convex by Step 0), the convex function ϕ\phi, its Borel set DD taken to be EwE_{w} (Step 3), which satisfies Ew⊆RdE_{w}\subseteq\mathbb{R}^{d} and μ(Ew)=1\mu(E_{w})=1 by (4.1), and this TT: it gives T∈TμT\in T_{\mu}. By Basic Properties of the Tangent Space: Closed Subspace, the Identity Map Belongs to It, Second-Moment Limits, and Representation of Bounded Functionals on Gradients §closed, TμT_{\mu} is a linear subspace of L2(μ;Rd)L^{2}(\mu;\mathbb{R}^{d}), so it contains T+(−K) idT+(-K)\,\mathrm{id}, which is the class of ∇w\nabla w by The Space of Square-Integrable Random Vectors §classes, since T(x)−Kx=∇w(x)T(x)-Kx=\nabla w(x) for every xx. Hence ∇w∈Tμ\nabla w\in T_{\mu}; as μ\mu was an arbitrary absolutely continuous member of P2(Rd)\mathcal{P}_{2}(\mathbb{R}^{d}), claim 2 is proved.

Step 6 (Claim 3). Suppose that −w-w is semiconvex with constant KK. The function −w-w is bounded, since ∣−w(x)∣=∣w(x)∣|-w(x)|=|w(x)|, and Lipschitz with constant LL, since ∣(−w)(x)−(−w)(y)∣=∣w(x)−w(y)∣|(-w)(x)-(-w)(y)|=|w(x)-w(y)|; so Steps 1 to 5 apply to −w-w in place of ww. Let x∈Rdx\in\mathbb{R}^{d}. If ww is differentiable at xx with derivative matrix JJ, let −J-J be the row with entries −J1i-J_{1i}; then (−J)h=−(Jh)(-J)h=-(Jh) for h∈Rdh\in\mathbb{R}^{d} by Matrix-Vector Product and Bilinearity and Symmetry of the Dot Product on Rn\mathbb{R}^n (the single coordinate of JhJh being the dot product of the vector of entries of JJ with hh, as in Step 0), so (−w)(x+h)−(−w)(x)−(−J)h=−(w(x+h)−w(x)−Jh)(-w)(x+h)-(-w)(x)-(-J)h=-\bigl(w(x+h)-w(x)-Jh\bigr) has the same absolute value, and −w-w is differentiable at xx with derivative matrix −J-J. Applying this to −w-w, whose negative is ww, gives the converse; so E−w=EwE_{-w}=E_{w}. By claim 1 of A Derivative Matrix is the Jacobian Matrix, and is Unique, ∂i(−w)(x)=−J1i=−∂iw(x)\partial_{i}(-w)(x)=-J_{1i}=-\partial_{i}w(x) for x∈Ewx\in E_{w}, so D(−w)(x)=−Dw(x)D(-w)(x)=-Dw(x) there, and hence ∇(−w)(x)=−∇w(x)\nabla(-w)(x)=-\nabla w(x) for every x∈Rdx\in\mathbb{R}^{d} (both sides being the origin off EwE_{w}). Let μ∈P2(Rd)\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}) be absolutely continuous. Claim 2 for −w-w gives ∇(−w)∈Tμ\nabla(-w)\in T_{\mu}. By The Space of Square-Integrable Random Vectors §classes the class of ∇w\nabla w is (−1)(-1) times that of ∇(−w)\nabla(-w), and TμT_{\mu} is a linear subspace by Basic Properties of the Tangent Space: Closed Subspace, the Identity Map Belongs to It, Second-Moment Limits, and Representation of Bounded Functionals on Gradients §closed; hence ∇w∈Tμ\nabla w\in T_{\mu}, which proves claim 3.

Citations

Loading…

Dependencies

Uses0

Loading…

Comments

Log in to comment.

Loading…