Research
Robust Risk Under Evolving Uncertainty: A Wasserstein Counterpart of the Entropic Value-at-Risk
Overview Research area: Risk-averse sequential decision-making, combining coherent risk measures (entropic value-at-risk), distributionally robust optimization, optimal transport / Wasserstein ambigui
- arXiv
- 2608.19073
- Published
- 2026-08-19
- Authors
- Deep Kumar Ganguly, Jan Křetínský
AI summary
Overview
- Research area: Risk-averse sequential decision-making, combining coherent risk measures (entropic value-at-risk), distributionally robust optimization, optimal transport / Wasserstein ambiguity sets, and safe reinforcement learning.
- Technical level: Advanced. The paper is primarily a theory paper (definitions, dualities, coherence axioms, contraction proofs) with a numerical verification section solved using Gurobi 12.
- Scope in one sentence: The paper replaces the relative-entropy ball underlying the entropic value-at-risk with a 1-Wasserstein ball to define a new coherent risk measure, the Wasserstein entropic value-at-risk (WEVaR), and then drives that transport radius with the agent's belief entropy to obtain a closed-form robust dynamic-programming operator with a certified safety sandwich.
What This Paper Is About
An agent learning its environment should be maximally cautious while ignorant and bold once confident, so its risk attitude must be a function of its evolving epistemic state. The entropic value-at-risk already formalizes this intuition—a confidence level α sets the radius −ln α of a relative-entropy ball of alternative models—but that ball contains only laws absolutely continuous with respect to the nominal, so it cannot represent a catastrophe the nominal deems impossible, exactly the case a safe agent must hedge. The authors keep the robust-optimization template and swap the relative-entropy ball for a 1-Wasserstein ball, defining the Wasserstein entropic value-at-risk, proving its variational dual, coherence, and place in the risk hierarchy, and then setting the transport radius to β times the belief entropy to obtain a closed-form robust dynamic-programming operator with a safety switch.
Key Contributions
- A new risk measure with a mirroring dual. The authors define the Wasserstein entropic value-at-risk (WEVaR) as the worst-case expectation over a 1-Wasserstein ball of radius ε, and prove a Kantorovich–Rubinstein variational dual (Theorem 1) that mirrors the entropic formula term for term, is convex and well-posed (the transport analogue of the guarantee of Ganguly et al., 2025), and admits a closed "mean-plus-Lipschitz" form.
- Coherence and a definite place in the risk hierarchy. They show WEVaR is coherent (Proposition 1) and place it relative to the entropic measure (Theorem 2): it is sandwiched against the entropic measure by a transport–entropy inequality, sweeps the full hierarchy from the mean to the essential supremum, and strictly accounts for zero-nominal-probability catastrophes the entropic measure ignores.
- Numerical verification of both robust programs and their dualities. Using Gurobi 12 on a five-state metric space, they verify the entropic primal/dual pair, the transport linear program against its Kantorovich–Rubinstein dual, and the comparison between the two measures.
- An evolving-uncertainty operator. Driving the transport radius by belief entropy yields a closed-form robust dynamic-programming operator (Theorem 3), a γ-contraction with a unique fixed point, a certified safety sandwich, and a computable "safety switch" under evolving belief.
Main Findings
- Variational dual and closed form (Theorem 1): WEVaR_ε(X) = inf_{λ ≥ 0} {λ ε + E_P[X^c_λ]}, where X^c_λ(s) = max_{s'}(X(s') − λ d(s', s)) is the c-transform of X. The objective is convex in λ, so (4) is a well-posed one-dimensional convex program; the optimal transport price satisfies λ* ≤ Lip_d(X); ε ↦ WEVaR_ε(X) is concave and nondecreasing; and WEVaR_ε(X) = E_P[X] + ε Lip_d(X) for ε below a saturation threshold, in particular exactly for two-point supports.
- Coherence (Proposition 1): WEVaR_ε is a coherent risk measure for every ε ≥ 0—it is the support function of the convex, compact set {Q : W_1(Q, P) ≤ ε}, giving positive homogeneity, subadditivity, monotonicity, and translation invariance.
- Risk hierarchy (Theorem 2(i)): WEVaR_0(X) = E_P[X] and WEVaR_ε(X) increases to the essential supremum of X as ε approaches the diameter D, so WEVaR sweeps the full hierarchy.
- Sandwich against the entropic measure (Theorem 2(ii)): EVaR_α(X) ≤ WEVaR_{D√(−½ ln α)}(X) for all α in (0, 1], derived via Pinsker's inequality, the bound W_1 ≤ D · TV, and the containment of the relative-entropy ball in a Wasserstein ball.
- Catastrophe blind spot versus transport reach (Theorem 2(iii)): If P(s*) = 0 and X(s*) > E_P[X], then EVaR_α(X) is independent of X(s*) for every α, whereas WEVaR_ε(X) strictly increases in X(s*) once ε > dist_d(s*, supp P).
- Numerical verification (Section 4): On a five-state metric space with X = [0, 2, 4, 8, 100] and d(s, s') = |s − s'|, the entropic primal (1) and the relative-entropy program (2) agree to 1.2 × 10^−4, and the transport linear program agrees with its Kantorovich–Rubinstein dual to 9 × 10^−7. The closed form (5) is exact in the unsaturated regime (here ε ≤ 0.1) and otherwise upper-bounded by the exact one-dimensional dual.
- Catastrophe quantified: With a zero-nominal-probability disaster whose loss is grown from 50 to 1000, the entropic value-at-risk changes by exactly 0 at fixed confidence, while WEVaR at a fixed radius changes by 807.5.
- Belief-entropy radius (Section 5): The radius is set to ε(b) = β H(b), where H(b) = −Σ_z b(z) ln b(z) is the belief entropy, giving the ambiguity set U(b) = {Q : W_1(Q, P̄_b) ≤ β H(b)}. As observations sharpen b, the entropy falls, the ball contracts and drifts toward the nominal, and the risk attitude slides from worst-case to risk-neutral.
- Closed-form robust update (Theorem 3): The operator (𝔗V)(s, b) = max_a [r(s, a) + γ inf_{Q ∈ U(b)} E_Q V(s', b')] has a closed-form inner problem with no coupling linear program, reducing the per-(s, a, b) cost to O(|S|²); 𝔗 is a γ-contraction in the sup norm when γ < 1 and a contraction on proper MDPs when γ = 1, with a unique fixed point V*.
- Safety sandwich (Theorem 3(iii)): V*wc(s, b) ≤ V*(s, b) ≤ E{z~b}[V*_opt(s, z)] between the always-maximally-cautious value and the type-aware oracle, and as H(b) → 0 the radius vanishes and V*(s, b) → V*_opt(s, z*).
- Calibration of the sensitivity: Since H(b) ≤ ln|Z|, the worst-case radius is β ln|Z|; requiring one action to remain viable under maximal ignorance gives β < D / ln|Z|, described as the transport analogue of choosing a confidence level.
- Evolving safety switch: On the Ambiguous Bridge with γ = 1, Sprint reaches the goal (+100) under a benign type but falls to catastrophe (−1000) under an adversarial type, while Crawl is always safe but costly. With belief b = [α, 1 − α], d(S_G, S_F) = 1, and Lip_d(V) = 1100, the closed forms give Q(Sprint) = −1 + 1100(α − β H(α))^+ − 1000 and Q(Crawl) = 80 − 1100 min(β H(α), 1), reproduced by a Gurobi transport solve to 10^−13. Equating them yields a safety switch independent of β: the agent Crawls until α ≥ α* = 1081/1100 ≈ 0.983, then Sprints.
- Canyon illustration: The same creep-then-cruise switch governs the drone example of Section 1—the drone hovers until its confidence that the air is calm crosses ≈ 0.63, then cruises.
- Illustrated guarantees (Figure 3): On a five-state corridor with a hidden "storm"/slip type (γ = 0.9, Gurobi-checked), the safety sandwich holds at every belief; the price of robustness is temporary, with the ceiling gap falling from 8.4 at the uniform belief to 4.4 once the belief reaches b(calm) = 0.95, and to zero at identification. Value iteration with the closed-form operator converges from arbitrary initializations at empirical rate 0.90 = γ.
- Geometric contrast (Figure 1): On the next-state simplex at a confident belief where the nominal P̄_b places zero mass on "Fail," the relative-entropy ball is trapped on the support edge and blind to the catastrophe, while the 1-Wasserstein ball reaches the Fail vertex, with the worst-case adversary on its boundary.
- Not reported: The paper does not report evaluation on any standard reinforcement-learning benchmark suite, dataset sizes, or wall-clock timings for the experiments; the numerical work is a five-state metric-space verification and a five-state corridor computed with Gurobi 12.
Methodology in Plain English
The authors take an existing robust-optimization recipe and change one ingredient. The entropic value-at-risk is defined as the worst-case expected loss over all probability models within a certain "relative-entropy" distance from the nominal model. That distance measures informational surprise, and it has a structural flaw: if the nominal model assigns zero probability to a disaster, every model within any finite relative-entropy distance also assigns it zero probability, so the risk measure never sees that disaster no matter how large the loss.
The authors instead measure model distance by how far probability mass has to physically move across a ground metric—the 1-Wasserstein distance. A disaster can then be reached at a cost proportional to how far it is from the nominal's support: reachable, but expensive. They define WEVaR as the worst-case expected loss over this ball. Using linear-programming duality for Wasserstein distributionally robust optimization, they turn the inner worst-case problem into an infimum over a single scalar "transport price" λ of an expected smoothed loss plus λ times the radius—the exact structural analogue of the entropic formula, where the inverse temperature t becomes the transport price λ and the exponential tilt becomes a c-transform.
They then prove the resulting risk measure satisfies the standard coherence axioms, and compare it to the entropic measure using Pinsker's inequality plus the bound that Wasserstein distance is at most the diameter times total variation, which shows the relative-entropy ball sits inside a suitably sized Wasserstein ball.
For the learning setting, they consider a hidden environment type drawn once, about which the agent maintains a Bayesian belief. The agent forms a belief-weighted nominal kernel and sets the radius to β times the entropy of its belief: high entropy (ignorance) means a large ambiguity set, low entropy (confidence) means a small one. Because the inner worst-case problem has a closed form, the robust Bellman update requires no nested optimization—just a Lipschitz constant of the value function and a scalar computation. They prove this operator contracts, has a unique fixed point, and is sandwiched between an always-maximally-cautious value and a type-aware oracle, then check the whole construction numerically with Gurobi 12 on small finite metric spaces.
Why This Matters
- Impact on research: The paper argues that risk attitude should be read off the agent's evolving epistemic state rather than fixed in advance, and it supplies the optimal-transport geometry that safety arguments appear to require. It links two normally separate lines of work—entropic risk measures and Wasserstein distributionally robust optimization—by showing the two formulas are the same template under different ambiguity geometries, connected by entropic optimal transport.
- Real-world applications:
- Autonomous flight and navigation under uncertain conditions, such as the paper's drone crossing a canyon under uncertain wind, where a high-cruise policy is fast in calm air but fatal in a gust, and a hover-and-creep policy is safe but slow.
- Safe reinforcement learning and control for systems with rare but catastrophic failure modes, where the blind spot of relative-entropy ambiguity sets makes the safety check vacuous precisely when overconfidence is most dangerous.
- Quantitative risk management in finance, where the worst-case loss over a Wasserstein ball can be computed in closed form as mean plus radius times the Lipschitz constant of the loss.
- Verification of learning agents during deployment, since the safety sandwich gives a valid bound at every iterate of value iteration, not only at convergence.
- Industry relevance: The closed-form inner update reduces a per-state-action-belief computation to order |S|² instead of solving a transportation linear program with order |S|² variables, which matters for embedding robust risk into planning loops. The calibration rule β < D / ln|Z| gives practitioners a concrete way to set the caution slider—described as the transport analogue of choosing a confidence level—and the safety switch on the Ambiguous Bridge shows the resulting policy thresholds can be set purely by reward asymmetry.
Future Directions
- Scaling to continuous state spaces. The discussion names scaling the closed-form operator to continuous spaces via Lipschitz critics as a natural next step.
- Active information gathering. The authors list active information gathering as a second natural next step, connecting the belief-entropy-driven radius to decisions about which observations to seek.
- Exploiting the entropic-optimal-transport bridge. The discussion frames the entropic and Wasserstein robust risks as two faces of one idea bridged by entropic optimal transport (Cuturi, 2013; Peyré and Cuturi, 2019), suggesting interpolation between the two geometries as an open direction.
- Open questions the paper leaves implicit: how the sensitivity parameter β should be chosen in practice beyond the stated bound β < D / ln|Z|; whether the safety sandwich can be tightened beyond the maximal-ignorance floor V*_wc; and whether the approach can be validated at scale, since no evaluation on standard reinforcement-learning benchmarks or datasets is reported here.
Target Audience
This paper is aimed at researchers and graduate students working on distributionally robust optimization, coherent risk measures, optimal transport, and safe or risk-sensitive reinforcement learning. It will also interest theoretically inclined practitioners in autonomous systems, robotics, and quantitative finance who need worst-case guarantees that remain meaningful for events the nominal model considers impossible. The density of proofs and duality arguments makes it most accessible to readers comfortable with convex analysis, linear-programming duality, and Markov decision processes; beginners will find the motivating drone and bridge examples useful but the theorems demanding.
Authors’ abstract
An agent still learning its environment should be cautious while ignorant and bold once confident. The entropic value-at-risk captures this through a robust-optimization identity---a confidence level fixes the radius of a relative-entropy ball of alternative models---but that ball cannot reach catastrophes the nominal deems impossible, precisely what a safe agent must hedge. We instead use an optimal-transport ball and study the coherent risk measure it induces, the Wasserstein entropic value-at-risk. It has a variational dual mirroring the entropic formula (an inverse temperature becomes a transport price), occupies a definite place in the risk hierarchy, and provably accounts for the reachable catastrophes the entropic measure ignores; we verify both dualities numerically. Driving the transport radius by belief entropy then yields a closed-form robust dynamic-programming operator whose caution contracts as the belief sharpens, with a certified safety sandwich and a sharp safety switch.