of lem:noise-penalty-drift-hypotheses-hilbert-2026b
Verified by 0 users · Flagged by 0 users
Expand the shifts of the operator; properness is immediate, coercivity follows from completing a square, semicontinuity from strong and weak convergence along couplings and lower semicontinuity of the penalty, the structure condition from score monotonicity derived from displacement convexity along optimal maps, and momentum continuity from a Lipschitz bound.
Proof
Each result cited is universally quantified over the data in its own statement.
Elementary order and arithmetic of real numbers (adding inequalities, multiplying them by nonnegative or positive reals, handling absolute values, the nonnegativity of squares) is carried by The Real Numbers: Standing Notation and Background §background and is used without further mention; so is the fact that for nonnegative reals x,y one has x<y, respectively x≤y, exactly when x2<y2, respectively x2≤y2 (claims 1 and 2 of Monotonicity of Squaring on the Nonnegative Elements of an Ordered Field). The four assumptions of the statement are referred to as (Convexity), (Lower semicontinuity of the penalty), (Growth) and (Running cost).
Step 0 (Notation and preliminary facts). Fix the data of the statement: the noise penalty pair (D,DΣ,E,Σ), the reals λ0,θ with 0<λ0 and 0<θ≤1, the function g, the operator F, a real number C as in (Growth), and a real number M≥0 as in (Running cost), so that ∣g(μ)∣≤M for every μ∈D.
(0.5) Elementary inequalities in δ. Let δ∈R with 0<δ<1. Then 0<θδ≤δ<1, because 0<θ≤1. Consequently 0<1+θδ, 0≤1−θδ≤1, 2δ+2θδ2>0, 2δ−2θδ2=2δ(1−θδ)≥0, and δ−2θδ2=δ−2δθδ≥2δ>0. For real m,s one has 2ms≤m2+s2, since 0≤(m−s)2=m2−2ms+s2.
Lower bound for Fδ−(ξ). In (0a) we have λ0r≥−λ0R; λ0δE(ν)≥−λ0δ∣E(ν)∣≥−λ0R as δ<1; 2θ∥q+δΣ(ν)∥ν2≥0; ⟨Σ(ν),q⟩ν≥−sR by the Cauchy-Schwarz inequality of (0.1); δ∥Σ(ν)∥ν2=δs2≥2δs2; and −g(ν)≥−M by the choice of M in Step 0, as ν∈D. Hence, since 2R2≥0,
Upper bound for Fδ+(η). In (0d) we have λ0r′≤λ0R; −λ0δE(ν′)≤λ0R; 2θ∥q′∥ν′2≤2R2 as θ≤1; (1−θδ)⟨Σ(ν′),q′⟩ν′≤(1−θδ)s′R≤Rs′ by Cauchy-Schwarz and 0≤1−θδ≤1 (0.5); −(δ−2θδ2)s′2≤−2δs′2 by (0.5); and −g(ν′)≤M. Hence Fδ+(η)≤A−φ(s′).
Conclusion. For every real x, 0≤2δ(x−δR)2=φ(x)+2δR2, so φ(x)≥−2δR2. From the two bounds, φ(s)+φ(s′)−2A≤Fδ−(ξ)−Fδ+(η)<R, hence φ(s)<R+2A−φ(s′)≤B and likewise φ(s′)<B. Now let x≥0 with φ(x)<B, and suppose x>Cδ,R. Then x>1, x>δ4R and x>δ4B, all three numbers being at most Cδ,R. The second gives Rx<4δx2, so φ(x)>4δx2; the first gives x2>x, so 4δx2>4δx; and the third gives 4δx>B. Thus φ(x)>B, a contradiction. Hence x≤Cδ,R; applied to x=s and x=s′ this proves the claim. As δ,R were arbitrary, F satisfies the shift-coercivity condition. Of the assumptions only the boundedness in (Running cost) and θ≤1 were used.
Step 3 (An estimate along a coupling).Claim. Let ν′,ν∈P(X), π∈Π(ν′,ν), h,k∈L2(ν′;Xa), w∈L2(ν;Xa), and let ℓ∈R be positive with ∥h∥ν′≤ℓ. Let D=∫X×X∣k(x)−w(y)∣a2π(dz) be the discrepancy of k and w along π. Then
(⟨h,k⟩ν′−Ka(h,w,π))2≤ℓ2D.
Proof. Put β=⟨h,k⟩ν′−Ka(h,w,π) and let t∈R. The field th+k lies in L2(ν′;Xa), and by (0.2)(a) its discrepancy with w along π is nonnegative and equals ∥th+k∥ν′2−2Ka(th+k,w,π)+∥w∥ν2. By (0.1), ∥th+k∥ν′2=t2∥h∥ν′2+2t⟨h,k⟩ν′+∥k∥ν′2; by (0.2)(b), Ka(th+k,w,π)=tKa(h,w,π)+Ka(k,w,π); and by (0.2)(a) again, ∥k∥ν′2−2Ka(k,w,π)+∥w∥ν2=D. Therefore
0≤t2∥h∥ν′2+2tβ+D≤t2ℓ2+2tβ+Dfor every t∈R.
With t=−βℓ−2 this reads 0≤β2ℓ−2−2β2ℓ−2+D=D−β2ℓ−2, and multiplying by ℓ2>0 gives the claim.
Consequence. Let ν∈P(X), w∈L2(ν;Xa) and ℓ>0, and for every n∈N let νn′∈P(X), πn∈Π(νn′,ν) and hn,kn∈L2(νn′;Xa) with ∥hn∥νn′≤ℓ, such that the discrepancies Dn of kn and w along πnconverge to 0. Then βn=⟨hn,kn⟩νn′−Ka(hn,w,πn) converges to 0: given ε>0, choose N with Dn<ε2ℓ−2 for n≥N (the Dn being nonnegative); then ∣βn∣2=βn2≤ℓ2Dn<ε2 by the claim, so ∣βn∣<ε.
(4.1) Wa(ν,νn)→0. Given ε>0, choose N with Ia(πn)<ε2 for n≥N, by Limit of a Sequence of Real Numbers. Then Wa(νn,ν)2≤Ia(πn)<ε2 by (0.1), so Wa(νn,ν)<ε, and Wa(ν,νn)<ε by symmetry (0.1), for n≥N.
(4.2) Convergent terms. (a) g(νn)→g(ν): given ε>0, (Running cost) gives γ>0 with ∣g(μ)−g(μ′)∣<ε for μ,μ′∈D with Wa(μ,μ′)<γ, and by (4.1) this applies to μ=νn, μ′=ν for all large n. (b) ∥qn∥νn2→∥q∥ν2: by the consequence in Step 3 with ℓ=R, hn=kn=qn and w=q, the numbers βn=∥qn∥νn2−Kn(qn,q) tend to 0; by (0.2)(a), Dn=∥qn∥νn2−2Kn(qn,q)+∥q∥ν2=2βn−∥qn∥νn2+∥q∥ν2, so ∥qn∥νn2=2βn−Dn+∥q∥ν2→∥q∥ν2. (c) ⟨σn,qn⟩νn→⟨σ,q⟩ν: by the consequence in Step 3 with ℓ=R, hn=σn, kn=qn and w=q, the difference ⟨σn,qn⟩νn−Kn(σn,q) tends to 0, and Kn(σn,q)→⟨σ,q⟩ν by weak convergence with w=q.
(4.3) Lower semicontinuous terms. (a) For every ε>0 there is N with E(νn)>E(ν)−ε for n≥N: by (Lower semicontinuity of the penalty) and Lower Semicontinuous Function on a Subset of a Metric Space, E is lower semicontinuous at ν∈D relative to D, which gives γ>0 with E(ν)−ε<E(μ) for μ∈D with Wa(ν,μ)<γ; apply (4.1), as νn∈D. (b) For every ε>0 there is N with ∥σn∥νn2>∥σ∥ν2−ε for n≥N: by (0.2)(a), 0≤∥σn∥νn2−2Kn(σn,σ)+∥σ∥ν2, so ∥σn∥νn2≥2Kn(σn,σ)−∥σ∥ν2; the right side converges to 2∥σ∥ν2−∥σ∥ν2=∥σ∥ν2 by weak convergence with w=σ, so it exceeds ∥σ∥ν2−ε for large n.
(4.4) The lower shift. By (0b), Fδ−(ξn)=un+vn and Fδ−(ξ)=u+v, where
and u,v are the same expressions at ξ. By (4.2) and the limit laws, un→u. Let k=λ0δ+δ+2θδ2+1>0. Given ε>0, (4.3) applied with εk−1 gives, for large n, vn≥v−(λ0δ+δ+2θδ2)εk−1≥v−ε, the coefficients being nonnegative by (0.5). Now let c∈R be such that for every ε>0 there is N with Fδ−(ξn)≤c+ε for n≥N. Fix ε>0 and choose n so large that Fδ−(ξn)≤c+3ε, ∣un−u∣<3ε and vn≥v−3ε (the largest of three thresholds, each obtained with the positive number 3ε in place of ε). Then
and u′,v′ are the same expressions at ξ. As in (4.4), un′→u′, and, the coefficients λ0δ and δ−2θδ2 being nonnegative by (0.5), for every ε>0 we have vn′≥v′−ε for large n. Let c∈R be such that for every ε>0 there is N with c−ε≤Fδ+(ξn) for n≥N. Fix ε>0 and choose n so large that c−3ε≤Fδ+(ξn), ∣un′−u′∣<3ε and vn′≥v′−3ε (each threshold obtained with 3ε in place of ε). Then
Step 5 (First-order structure at uniquely noise-mapped pairs, claim 4). Let T={t∈R:0≤t}. The pair (ω1,ω2) below is chosen first; it depends only on g,λ0,θ,C, and we show it is a structure pair for F at R for every positive R.
(5.1) The modulus ω1. For s∈T let G(s)={∣g(μ′)−g(ν′)∣:μ′,ν′∈D,Wa(μ′,ν′)2≤s}. It contains 0=∣g(μ0)−g(μ0)∣, with μ0 from (0.4), since Wa(μ0,μ0)=0 by (0.1); and it is bounded above by 2M. Hence ω1(s)=supG(s) exists by The Real Numbers: Standing Notation and Background §bounds, and 0≤ω1(s) because ω1(s) is an upper bound of G(s)∋0 (Upper Bound and Least Upper Bound). Given ε>0, (Running cost) gives γ>0 with ∣g(μ′)−g(ν′)∣<ε whenever μ′,ν′∈D and Wa(μ′,ν′)<γ. Put γ1=γ2/4>0. If t∈T and t≤γ1, every element of G(t) comes from μ′,ν′∈D with Wa(μ′,ν′)2≤(γ/2)2, so Wa(μ′,ν′)≤γ/2<γ and the element is <ε; thus ε is an upper bound of G(t) and ω1(t)≤ε, the supremum being the least upper bound (Upper Bound and Least Upper Bound). So ω1 is a modulus of continuity, and by construction ∣g(μ′)−g(ν′)∣≤ω1(s) whenever μ′,ν′∈D, s∈T and Wa(μ′,ν′)2≤s.
(5.2) The function ω2. For t∈T and real α>1 put ω2(t,α)=(λ0+4Cθ2α2)t. For each α>1 the coefficient is nonnegative by (0.4), so t↦ω2(t,α) is a modulus of continuity by Linear Moduli of Continuity §modulus.
(5.3) The inequality. Let R>0, and let α,δ,μ,ν,S,S′,r be as in The First-Order Structure Condition at Uniquely Noise-Mapped Pairs §pair: 1<α, 0<δ<1, μ,ν∈DΣ with both ordered pairs (μ,ν) and (ν,μ) uniquely noise-mapped, S a noise-optimal map from μ to ν and S′ one from ν to μ, and −R≤r≤R. (Neither the unique mapping nor the condition δ(∣E(μ)∣+∣E(ν)∣)≤R will be needed.) Write W=Wa(μ,ν), σ=Σ(μ)∈L2(μ;Xa), τ=Σ(ν)∈L2(ν;Xa), p=α(id−S)=−α(S−id)∈L2(μ;Xa), p′=α(S′−id)∈L2(ν;Xa), e=∣E(μ)∣+∣E(ν)∣ and t=δ(e+1), and let Δ=Fδ−(μ,r,p)−Fδ+(ν,r,p′) be the difference to be bounded below.
Norms of the displacements. By (0.3), ∥S−id∥μ=W and ∥S′−id∥ν=Wa(ν,μ)=W (symmetry, (0.1)); hence, by homogeneity (0.1) with ∣α∣=α and ∣−α∣=α, ∥p∥μ=∥p′∥ν=αW.
A bound on W2. The measures μ,ν,ρ lie in Pρa, so by the triangle inequality and symmetry (0.1), 0≤W≤Wa(μ,ρ)+Wa(ν,ρ). Hence W2≤(Wa(μ,ρ)+Wa(ν,ρ))2≤2Wa(μ,ρ)2+2Wa(ν,ρ)2, the last step by (0.5) with m=Wa(μ,ρ), s=Wa(ν,ρ). As μ,ν∈D, (Growth) yields W2≤2C(1+∣E(μ)∣)+2C(1+∣E(ν)∣)=2C(2+e). Since 0≤C (0.4) and 2+e≤2(1+e), we get δW2≤2Cδ(2+e)≤4Cδ(1+e)=4Ct.
Expansion of Δ. Subtracting (0c) at (ν,r,p′) from (0a) at (μ,r,p), the terms λ0r cancel and
Bounds for the individual terms. (i) Monotonicity of the score. The pair is displacement convex by (Convexity), that is, 0-displacement convex in the sense of Lambda-Displacement Convexity of a Noise Penalty Pair §convex. By (0.3) the couplings πS=(id,S)#μ∈Πa(μ,ν) and πS′=(id,S′)#ν∈Πa(ν,μ) are noise-optimal, with Ja(σ,πS)=⟨σ,S−id⟩μ and Ja(τ,πS′)=⟨τ,S′−id⟩ν. Since μ,ν∈DΣ⊆D, the convexity inequality of Lambda-Displacement Convexity of a Noise Penalty Pair §convex with λ=0 applies to μ∈DΣ, ν∈D and πS, and to ν∈DΣ, μ∈D and πS′; the cost terms carry the factor 20=0, so
E(μ)+⟨σ,S−id⟩μ≤E(ν),E(ν)+⟨τ,S′−id⟩ν≤E(μ).
Adding, the real numbers E(μ) and E(ν) cancel and ⟨σ,S−id⟩μ+⟨τ,S′−id⟩ν≤0. By bilinearity (0.1), ⟨σ,p⟩μ−⟨τ,p′⟩ν=−α(⟨σ,S−id⟩μ+⟨τ,S′−id⟩ν)≥0. (ii) Cross terms: by Cauchy-Schwarz (0.1) and (0.5) with m=∥σ∥μ, s=θαW,
and in the same way θδ⟨p′,τ⟩ν≥−2δ∥τ∥ν2−2δθ2α2W2. (iii) Collecting the score terms: the coefficient of ∥σ∥μ2 becomes δ+2θδ2−2δ=2δ+2θδ2 and that of ∥τ∥ν2 becomes δ−2θδ2−2δ=2δ(1−θδ), both nonnegative by (0.5) (the second one is where θ≤1 enters), so these terms are ≥0; the remaining contribution is −θ2α2δW2≥−4Cθ2α2t. (iv) Penalty terms: λ0δ(E(μ)+E(ν))≥−λ0δe≥−λ0t. (v) Running cost: W2≤αW2≤αW2+α−1, as 1<α, 0≤W2 and 0<α−1; so (5.1) with s=αW2+α−1, applied to μ,ν∈D, gives g(ν)−g(μ)≥−∣g(μ)−g(ν)∣≥−ω1(αW2+α−1).
Step 6 (Momentum-continuous shifts, claim 5).A Lipschitz bound. Let δ,R∈R with 0<δ<1 and 0<R, let ν∈DΣ with ∥Σ(ν)∥ν≤R, let r∈R, and let q,q′∈L2(ν;Xa) with ∥q∥ν≤R and ∥q′∥ν≤R; write σ=Σ(ν). For w,w′∈L2(ν;Xa), bilinearity and symmetry (0.1) give ⟨w−w′,w+w′⟩ν=∥w∥ν2−∥w′∥ν2. With w=q+δσ and w′=q′+δσ, so that w−w′=q−q′ and w+w′=q+q′+2δσ, the formula (0a) shows that the terms λ0r+λ0δE(ν), δ∥σ∥ν2 and −g(ν) cancel in the difference, and
By the triangle inequality and homogeneity (0.1), ∥q+q′+2δσ∥ν≤2R+2δR=2(1+δ)R. Taking absolute values, with the triangle inequality for the absolute value and Cauchy-Schwarz (0.1) for each inner product, and using θ>0 and ∥σ∥ν≤R,
since δ<1. In the same way, from (0c) with w=q−δσ and w′=q′−δσ, Fδ+(ν,r,q)−Fδ+(ν,r,q′)=2θ⟨q−q′,q+q′−2δσ⟩ν+⟨σ,q−q′⟩ν, ∥q+q′−2δσ∥ν≤2(1+δ)R, and ∣Fδ+(ν,r,q)−Fδ+(ν,r,q′)∣≤(2θ+1)R∥q−q′∥ν.
The condition. Let δ,R,η∈R with 0<δ<1, 0<R and 0<η be given; then (2θ+1)R is positive, and we choose κ=η((2θ+1)R)−1>0. For ν∈DΣ with ∥Σ(ν)∥ν≤R, r∈R with ∣r∣≤R, and q,q′∈L2(ν;Xa) with ∥q∥ν≤R, ∥q′∥ν≤R and ∥q−q′∥ν<κ, the Lipschitz bound gives ∣Fδ∓(ν,r,q)−Fδ∓(ν,r,q′)∣≤(2θ+1)R∥q−q′∥ν<(2θ+1)Rκ=η for both shifts. Hence F has momentum-continuous shifts relative to the pair. None of the four assumptions, nor θ≤1, is used here.
Step 7 (Claim 6). Steps 1, 2, 4, 5 and 6 prove claims 1 to 5, and claim 6 collects them. ■