TheoremBase

Perturb the extremum along mutmu_t = (id + t grad psi)_# mu-hat: semiconvexity gives a one-sided quadratic expansion of W at the a.e. differentiability points, phi and E are differentiable along the curve, and the vanishing first variation bounds the Ornstein-Uhlenbeck functional (finite relative Fisher information) and shows that grad w - grad phi - (+/-) delta Sigma is a tangent field orthogonal to all test gradients, hence zero.

Proof

Each result cited below is universally quantified over the data in its own statement. Elementary arithmetic and order facts for real numbers (The Real Numbers: Standing Notation and Background §background) are used without citation, and linearity of integrals of integrable functions is Linearity and Monotonicity of the Lebesgue Integral §integrable. By The Gaussian Free-Energy Pair: Relative Entropy and Relative Score with Respect to a Diagonal Gaussian Measure §pair, D\mathcal{D} is the set of μ∈P2(Rd)\mu\in\mathcal{P}_{2}(\mathbb{R}^{d}) of finite relative entropy with respect to γc\gamma_{c}, DΣ\mathcal{D}_{\Sigma} is the set of μ∈D\mu\in\mathcal{D} of finite Fisher information relative to γc\gamma_{c}, and Σ(μ)=a ζμc∈Tμ\Sigma(\mu)=a\,\zeta^{c}_{\mu}\in T_{\mu} for μ∈DΣ\mu\in\mathcal{D}_{\Sigma}. We write ℓ=ℓμ^c\ell=\ell^{c}_{\hat{\mu}} for the Ornstein-Uhlenbeck functional of μ^\hat{\mu}.

Reduction. We treat both clauses at once, with σ=1\sigma=1 for clause 1 and σ=−1\sigma=-1 for clause 2. In either case the function σw\sigma w (value σw(x)\sigma w(x) at xx) is semiconvex on Rd\mathbb{R}^{d} with constant KK. Let Φ:D→R\Phi:\mathcal{D}\to\mathbb{R}, Φ(μ)=σ(W(μ)−φ(μ))−δ E(μ)\Phi(\mu)=\sigma\bigl(W(\mu)-\varphi(\mu)\bigr)-\delta\,\mathcal{E}(\mu). For σ=1\sigma=1 this is the function of clause 1; for σ=−1\sigma=-1 it is the negative of the function of clause 2, so by Local Minimum of a Function Relative to a Subset of a Metric Space and Local Maximum of a Function Relative to a Subset of a Metric Space it has a local maximum at μ^\hat{\mu} relative to D\mathcal{D}. Thus in both cases there is a positive r0∈Rr_{0}\in\mathbb{R} with Φ(μ)≤Φ(μ^)\Phi(\mu)\le\Phi(\hat{\mu}) for every μ∈D\mu\in\mathcal{D} with W2(μ^,μ)<r0W_{2}(\hat{\mu},\mu)<r_{0}. We prove that μ^∈DΣ\hat{\mu}\in\mathcal{D}_{\Sigma} and ∇w=∇φ(μ^)+σδ Σ(μ^)\nabla w=\nabla\varphi(\hat{\mu})+\sigma\delta\,\Sigma(\hat{\mu}) in L2(μ^;Rd)L^{2}(\hat{\mu};\mathbb{R}^{d}), which is clause 1 for σ=1\sigma=1 and clause 2 for σ=−1\sigma=-1.

Step 1 (the gradient of ww at μ^\hat{\mu}). As μ^\hat{\mu} has finite relative entropy with respect to γc\gamma_{c}, it has finite entropy by Relative Entropy and Relative Score with Respect to a Diagonal Gaussian Measure: Comparison with the Entropy, the Score and the Relative Free Energy of a Quadratic Potential §entropy, hence is absolutely continuous by Basic Properties of the Entropy on the Wasserstein Space: Comparison with the Gaussian Relative Entropy, Lower Bound, Translation Invariance, Absolute Continuity, Closed Sublevel Sets and Lower Semicontinuity §absolutely-continuous. Let EwE_{w} and Dw(x)Dw(x) (x∈Ewx\in E_{w}) be as in The Almost-Everywhere Gradient of a Lipschitz Semiconvex or Semiconcave Function Is Tangent at Every Absolutely Continuous Measure. By The Almost-Everywhere Gradient of a Lipschitz Semiconvex or Semiconcave Function Is Tangent at Every Absolutely Continuous Measure §borel, EwE_{w} is Borel, ∇w\nabla w is Borel and ∥∇w(x)∥≤L\lVert\nabla w(x)\rVert\le L for every xx, and, as μ^∈D⊆P2(Rd)\hat{\mu}\in\mathcal{D}\subseteq\mathcal{P}_{2}(\mathbb{R}^{d}) is absolutely continuous, μ^(Rd∖Ew)=0\hat{\mu}(\mathbb{R}^{d}\setminus E_{w})=0; by The Almost-Everywhere Gradient of a Lipschitz Semiconvex or Semiconcave Function Is Tangent at Every Absolutely Continuous Measure §semiconvex (for σ=1\sigma=1) or The Almost-Everywhere Gradient of a Lipschitz Semiconvex or Semiconcave Function Is Tangent at Every Absolutely Continuous Measure §semiconcave (for σ=−1\sigma=-1), ∇w∈Tμ^\nabla w\in T_{\hat{\mu}}.

Step 2 (a pointwise inequality). Let x∈Ewx\in E_{w} and h∈Rdh\in\mathbb{R}^{d}. We show

σ(w(x+h)−w(x)−Dw(x)⋅h)≥−K2∥h∥2.\sigma\bigl(w(x+h)-w(x)-Dw(x)\cdot h\bigr)\ge-\tfrac{K}{2}\lVert h\rVert^{2}.

This is clear if h=0h=0; let h≠0h\ne0. By claims 1 and 2 of A Derivative Matrix is the Jacobian Matrix, and is Unique and Gradient of a Real-Valued Function on a Euclidean Open Set, ww is differentiable at xx with a derivative matrix JJ satisfying Jk=Dw(x)⋅kJk=Dw(x)\cdot k for k∈Rdk\in\mathbb{R}^{d}. Let ε>0\varepsilon>0; choose r>0r>0 for ε\varepsilon as in Differentiability at a Point for Maps Between Euclidean Spaces, and then s∈(0,1]s\in(0,1] with s∥h∥<rs\lVert h\rVert<r. By Quadratic Increment Characterisation of Semiconvexity, applied to σw\sigma w on Rd\mathbb{R}^{d} with constant KK, the points x+hx+h and xx and the parameter ss (note s(x+h)+(1−s)x=x+shs(x+h)+(1-s)x=x+sh),

σ(w(x+sh)−w(x))≤s σ(w(x+h)−w(x))+K2s(1−s)∥h∥2.\sigma\bigl(w(x+sh)-w(x)\bigr)\le s\,\sigma\bigl(w(x+h)-w(x)\bigr)+\tfrac{K}{2}s(1-s)\lVert h\rVert^{2}.

The point k=shk=sh satisfies 0<∥k∥=s∥h∥<r0<\lVert k\rVert=s\lVert h\rVert<r by Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n §homogeneity (with ∥h∥>0\lVert h\rVert>0 by Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n §square and Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n §vanishing, as h≠0h\ne0), so the choice of rr gives ∥w(x+k)−w(x)−Jk∥≤ε∥k∥\lVert w(x+k)-w(x)-Jk\rVert\le\varepsilon\lVert k\rVert, the norm on the left being that of R1\mathbb{R}^{1}. Here Jk=Dw(x)⋅(sh)=s Dw(x)⋅hJk=Dw(x)\cdot(sh)=s\,Dw(x)\cdot h by Bilinearity and Symmetry of the Dot Product on Rn\mathbb{R}^n, and the Euclidean norm of a point of R1\mathbb{R}^{1} is the absolute value of its coordinate by Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n §square; hence ∣w(x+sh)−w(x)−s Dw(x)⋅h∣≤εs∥h∥|w(x+sh)-w(x)-s\,Dw(x)\cdot h|\le\varepsilon s\lVert h\rVert, so, as ∣σ∣=1|\sigma|=1, the left side is at least s σDw(x)⋅h−εs∥h∥s\,\sigma Dw(x)\cdot h-\varepsilon s\lVert h\rVert. Dividing by s>0s>0 and using 1−s≤11-s\le1,

σDw(x)⋅h−ε∥h∥≤σ(w(x+h)−w(x))+K2∥h∥2.\sigma Dw(x)\cdot h-\varepsilon\lVert h\rVert\le\sigma\bigl(w(x+h)-w(x)\bigr)+\tfrac{K}{2}\lVert h\rVert^{2}.

As ε>0\varepsilon>0 was arbitrary, the claim follows.

Step 3 (the curve). Fix ψ∈Cc∞(Rd)\psi\in C_{c}^{\infty}(\mathbb{R}^{d}), let Gt=id+t∇ψG_{t}=\mathrm{id}+t\nabla\psi and μt=(Gt)#μ^\mu_{t}=(G_{t})_{\#}\hat{\mu}, and put P=∥∇ψ∥μ^P=\lVert\nabla\psi\rVert_{\hat{\mu}}, A=⟨∇w,∇ψ⟩μ^A=\langle\nabla w,\nabla\psi\rangle_{\hat{\mu}} and B=⟨∇φ(μ^),∇ψ⟩μ^B=\langle\nabla\varphi(\hat{\mu}),\nabla\psi\rangle_{\hat{\mu}}; here ∇φ(μ^)∈Tμ^\nabla\varphi(\hat{\mu})\in T_{\hat{\mu}} is the gradient along couplings of φ\varphi at μ^\hat{\mu}, which exists by Intrinsic Test Functions on the Wasserstein Space and Their Translation Hessians §differentiability and is unique by Differentiability of a Function on the Wasserstein Space Along Couplings, and Its Gradient §gradient. Let t0t_{0} be as in The Gaussian Free-Energy Pair is a Penalty Pair: Translations, the First Variation, Lower Semicontinuity, the Moment Bound and the Map Property §variation for μ^\hat{\mu} and ψ\psi, and t2t_{2} the lesser of t0t_{0} and r0/(P+1)r_{0}/(P+1). For t∈(−t2,t2)t\in(-t_{2},t_{2}): μt∈D\mu_{t}\in\mathcal{D} by that clause, W2(μ^,μt)≤∣t∣P<r0W_{2}(\hat{\mu},\mu_{t})\le|t|P<r_{0} by The Push-Forward of a Probability Measure with Finite Second Moment by the Identity Perturbed along the Gradient of a Test Function: Borel, Finite Second Moment, the Diagonal Coupling and the Wasserstein Bound §distance, and therefore Φ(μt)≤Φ(μ^)\Phi(\mu_{t})\le\Phi(\hat{\mu}); moreover μ0=μ^\mu_{0}=\hat{\mu} by The Push-Forward of a Probability Measure with Finite Second Moment by the Identity Perturbed along the Gradient of a Test Function: Borel, Finite Second Moment, the Diagonal Coupling and the Wasserstein Bound §borel.

Step 4 (WW along the curve). We show that for t∈(−t2,t2)t\in(-t_{2},t_{2}),

σ(W(μt)−W(μ^))≥tσA−K2t2P2.\sigma\bigl(W(\mu_{t})-W(\hat{\mu})\bigr)\ge t\sigma A-\tfrac{K}{2}t^{2}P^{2}.

The map GtG_{t} is Borel (The Push-Forward of a Probability Measure with Finite Second Moment by the Identity Perturbed along the Gradient of a Test Function: Borel, Finite Second Moment, the Diagonal Coupling and the Wasserstein Bound §borel) and ww is Borel and bounded, so w∘Gtw\circ G_{t} is Borel (Probability Measures on Euclidean Space and Random Vectors: Standing Notation §borel-maps) and bounded, and W(μt)=∫w∘Gt dμ^W(\mu_{t})=\int w\circ G_{t}\,d\hat{\mu} by the change-of-variables formula of Probability Measures on Euclidean Space and Random Vectors: Standing Notation §pushforward. By The Gradient of a Test Function is Bounded and Square-Integrable, and Its Laplacian Bounded and Integrable, Against Every Probability Measure §gradient there is Kψ≥0K_{\psi}\ge0 with ∥∇ψ(x)∥≤Kψ\lVert\nabla\psi(x)\rVert\le K_{\psi} for all xx; by Pairs of Euclidean Points: Coordinate Projections, Pairings, the Product Measure on a Euclidean Space, Borel Norm Functions and Finite Sets §functions the function ∇w⋅∇ψ\nabla w\cdot\nabla\psi is Borel with ∣∇w⋅∇ψ∣≤LKψ|\nabla w\cdot\nabla\psi|\le LK_{\psi}, and ∥∇ψ∥2\lVert\nabla\psi\rVert^{2} is Borel and bounded; their integrals against μ^\hat{\mu} are AA and P2P^{2} (Square-Integrable Vector Fields Against a Probability Measure on Euclidean Space, and Test Functions: Standing Notation §l2mu). Hence

ft=σ(w∘Gt−w−t ∇w⋅∇ψ)+K2t2∥∇ψ∥2f_{t}=\sigma\bigl(w\circ G_{t}-w-t\,\nabla w\cdot\nabla\psi\bigr)+\tfrac{K}{2}t^{2}\lVert\nabla\psi\rVert^{2}

is bounded and Borel (claim 2 of Arithmetic, Absolute Values, and Pointwise Limits of Measurable Real-Valued Functions), so integrable with respect to μ^\hat{\mu} (Probability Measures on Euclidean Space and Random Vectors: Standing Notation §measures), with ∫ft dμ^=σ(W(μt)−W(μ^)−tA)+K2t2P2\int f_{t}\,d\hat{\mu}=\sigma\bigl(W(\mu_{t})-W(\hat{\mu})-tA\bigr)+\tfrac{K}{2}t^{2}P^{2}. For x∈Ewx\in E_{w} one has ∇w(x)=Dw(x)\nabla w(x)=Dw(x), and Step 2 with h=t∇ψ(x)h=t\nabla\psi(x) gives ft(x)≥0f_{t}(x)\ge0 (using Bilinearity and Symmetry of the Dot Product on Rn\mathbb{R}^n and Elementary Properties of the Euclidean Norm on Rn\mathbb{R}^n §homogeneity). So the negative part ft−f_{t}^{-} vanishes off the null set Rd∖Ew\mathbb{R}^{d}\setminus E_{w} (Step 1), whence ∫ft− dμ^=0\int f_{t}^{-}\,d\hat{\mu}=0 by The Lebesgue Integral and Null Sets: Almost-Everywhere Comparison, Markov's Inequality, and Dominated Convergence Almost Everywhere §null-integral, and ∫ft dμ^=∫ft+ dμ^−∫ft− dμ^≥0\int f_{t}\,d\hat{\mu}=\int f_{t}^{+}\,d\hat{\mu}-\int f_{t}^{-}\,d\hat{\mu}\ge0 (Integrable Function and the Lebesgue Integral). This is the claim.

Step 5 (φ\varphi and E\mathcal{E} along the curve). Let u(t)=φ(μt)u(t)=\varphi(\mu_{t}) and e(t)=E(μt)e(t)=\mathcal{E}(\mu_{t}) for t∈(−t2,t2)t\in(-t_{2},t_{2}). Let ε>0\varepsilon>0 and let θ>0\theta>0 be given by Differentiability of a Function on the Wasserstein Space Along Couplings, and Its Gradient §differentiable at μ^\hat{\mu}, for the gradient ∇φ(μ^)\nabla\varphi(\hat{\mu}) (Differentiability of a Function on the Wasserstein Space Along Couplings, and Its Gradient §gradient), with ε/(P+1)\varepsilon/(P+1) in place of ε\varepsilon. For t∈(−t2,t2)t\in(-t_{2},t_{2}) with 0<∣t∣<θ/(P+1)0<|t|<\theta/(P+1), the coupling πt=(id,Gt)#μ^∈Π(μ^,μt)\pi_{t}=(\mathrm{id},G_{t})_{\#}\hat{\mu}\in\Pi(\hat{\mu},\mu_{t}) has I(πt)=t2P2<θ2I(\pi_{t})=t^{2}P^{2}<\theta^{2} (The Push-Forward of a Probability Measure with Finite Second Moment by the Identity Perturbed along the Gradient of a Test Function: Borel, Finite Second Moment, the Diagonal Coupling and the Wasserstein Bound §coupling), since ∣t∣P<θ|t|P<\theta, and by The Displacement Pairing of a Square-Integrable Vector Field Along a Coupling §displacement with S=GtS=G_{t}, whose displacement t∇ψt\nabla\psi is square-integrable against μ^\hat{\mu} (The Gradient of a Test Function is Bounded and Square-Integrable, and Its Laplacian Bounded and Integrable, Against Every Probability Measure §gradient), J(∇φ(μ^),πt)=⟨∇φ(μ^),t∇ψ⟩μ^=tB\mathcal{J}(\nabla\varphi(\hat{\mu}),\pi_{t})=\langle\nabla\varphi(\hat{\mu}),t\nabla\psi\rangle_{\hat{\mu}}=tB. Here, and in Steps 6 to 8, we use that the inner product of the real Hilbert space L2(μ^;Rd)L^{2}(\hat{\mu};\mathbb{R}^{d}) (Square-Integrable Vector Fields Against a Probability Measure on Euclidean Space, and Test Functions: Standing Notation §l2mu) is symmetric, additive and homogeneous in each argument, by Real Inner Product Space §inner-product. Hence ∣u(t)−u(0)−tB∣≤εP+1∣t∣P<ε∣t∣|u(t)-u(0)-tB|\le\tfrac{\varepsilon}{P+1}|t|P<\varepsilon|t|, so uu is differentiable at 00 with u′(0)=Bu'(0)=B (Derivative at an Interior Point). By The Gaussian Free-Energy Pair is a Penalty Pair: Translations, the First Variation, Lower Semicontinuity, the Moment Bound and the Map Property §variation and claim 2 of Restriction Stability of Continuity and of the Derivative, ee is differentiable at 00 with e′(0)=−a ℓ(ψ)e'(0)=-a\,\ell(\psi). Finally q(t)=−tσA+K2t2P2q(t)=-t\sigma A+\tfrac{K}{2}t^{2}P^{2} satisfies q(0)=0q(0)=0 and ∣q(t)/t−(−σA)∣=K2∣t∣P2<ε|q(t)/t-(-\sigma A)|=\tfrac{K}{2}|t|P^{2}<\varepsilon whenever 0<∣t∣<2ε/(KP2+1)0<|t|<2\varepsilon/(KP^{2}+1), so q′(0)=−σAq'(0)=-\sigma A (Derivative at an Interior Point).

Step 6 (the first variation). Let g=σu+δe+qg=\sigma u+\delta e+q on (−t2,t2)(-t_{2},t_{2}). By claim 2 of Sum, Constant Multiple, and Product Rules for One-Dimensional Derivatives, applied repeatedly, gg is differentiable at 00 with g′(0)=σB−δa ℓ(ψ)−σAg'(0)=\sigma B-\delta a\,\ell(\psi)-\sigma A. For t∈(−t2,t2)t\in(-t_{2},t_{2}), Step 3 gives σ(W(μt)−u(t))−δe(t)≤σ(W(μ^)−u(0))−δe(0)\sigma\bigl(W(\mu_{t})-u(t)\bigr)-\delta e(t)\le\sigma\bigl(W(\hat{\mu})-u(0)\bigr)-\delta e(0), that is σ(W(μt)−W(μ^))≤σ(u(t)−u(0))+δ(e(t)−e(0))\sigma\bigl(W(\mu_{t})-W(\hat{\mu})\bigr)\le\sigma\bigl(u(t)-u(0)\bigr)+\delta\bigl(e(t)-e(0)\bigr); with Step 4 this yields g(t)≥g(0)g(t)\ge g(0). So gg has a local minimum at 00 relative to (−t2,t2)(-t_{2},t_{2}), and g′(0)=0g'(0)=0 by Vanishing of the Derivative at an Interior Local Extremum. Multiplying by σ\sigma (σ2=1\sigma^{2}=1),

⟨∇w−∇φ(μ^),∇ψ⟩μ^=A−B=−σδa ℓ(ψ)for every ψ∈Cc∞(Rd).(6.1)\langle\nabla w-\nabla\varphi(\hat{\mu}),\nabla\psi\rangle_{\hat{\mu}}=A-B=-\sigma\delta a\,\ell(\psi)\qquad\text{for every }\psi\in C_{c}^{\infty}(\mathbb{R}^{d}).\tag{6.1}

Step 7 (finite Fisher information). By (6.1) and The Cauchy-Schwarz Inequality in a Real Inner Product Space in L2(μ^;Rd)L^{2}(\hat{\mu};\mathbb{R}^{d}), ∣ℓ(ψ)∣≤(δa)−1∥∇w−∇φ(μ^)∥μ^ ∥∇ψ∥μ^|\ell(\psi)|\le(\delta a)^{-1}\lVert\nabla w-\nabla\varphi(\hat{\mu})\rVert_{\hat{\mu}}\,\lVert\nabla\psi\rVert_{\hat{\mu}} for every ψ\psi, the constant not depending on ψ\psi. So μ^\hat{\mu} has finite Fisher information relative to γc\gamma_{c} (Finite Fisher Information Relative to a Diagonal Gaussian Measure: the Ornstein-Uhlenbeck Functional, the Relative Score and the Relative Fisher Information §finite), and as μ^∈D\hat{\mu}\in\mathcal{D}, μ^∈DΣ\hat{\mu}\in\mathcal{D}_{\Sigma}.

Step 8 (the first-order condition). By Finite Fisher Information Relative to a Diagonal Gaussian Measure: the Ornstein-Uhlenbeck Functional, the Relative Score and the Relative Fisher Information §score, ⟨Σ(μ^),∇ψ⟩μ^=a⟨ζμ^c,∇ψ⟩μ^=−a ℓ(ψ)\langle\Sigma(\hat{\mu}),\nabla\psi\rangle_{\hat{\mu}}=a\langle\zeta^{c}_{\hat{\mu}},\nabla\psi\rangle_{\hat{\mu}}=-a\,\ell(\psi), so by (6.1) the field η=∇w−∇φ(μ^)−σδ Σ(μ^)\eta=\nabla w-\nabla\varphi(\hat{\mu})-\sigma\delta\,\Sigma(\hat{\mu}) satisfies ⟨η,∇ψ⟩μ^=0\langle\eta,\nabla\psi\rangle_{\hat{\mu}}=0 for every ψ∈Cc∞(Rd)\psi\in C_{c}^{\infty}(\mathbb{R}^{d}). Since Tμ^T_{\hat{\mu}} is a linear subspace of L2(μ^;Rd)L^{2}(\hat{\mu};\mathbb{R}^{d}) (Basic Properties of the Tangent Space: Closed Subspace, the Identity Map Belongs to It, Second-Moment Limits, and Representation of Bounded Functionals on Gradients §closed) containing ∇w\nabla w (Step 1), ∇φ(μ^)\nabla\varphi(\hat{\mu}) (Step 3) and Σ(μ^)\Sigma(\hat{\mu}) (The Gaussian Free-Energy Pair: Relative Entropy and Relative Score with Respect to a Diagonal Gaussian Measure §pair), η∈Tμ^\eta\in T_{\hat{\mu}}. By Basic Properties of the Tangent Space: Closed Subspace, the Identity Map Belongs to It, Second-Moment Limits, and Representation of Bounded Functionals on Gradients §representation, applied to the zero functional with C=0C=0, exactly one ξ∈Tμ^\xi\in T_{\hat{\mu}} satisfies ⟨ξ,∇ψ⟩μ^=0\langle\xi,\nabla\psi\rangle_{\hat{\mu}}=0 for every ψ\psi; both η\eta and 00 do, so η=0\eta=0, that is ∇w=∇φ(μ^)+σδ Σ(μ^)\nabla w=\nabla\varphi(\hat{\mu})+\sigma\delta\,\Sigma(\hat{\mu}). ■\blacksquare

Citations

Loading…

Dependencies

Uses0

Loading…

Comments

Log in to comment.

Loading…