Research
Soft-Radial Projection for Constrained End-to-End Learning
Soft-Radial Projection for Constrained End-to-End Learning Overview Research area: Constrained machine learning / differentiable optimization layers; safe and feasibility-guaranteeing neural network a

- arXiv
- 2602.03461
- Published
- 2026-02-03
- Authors
- Philipp J. Schneider, Daniel Kuhn
AI summary
Soft-Radial Projection for Constrained End-to-End LearningOverview
- Research area: Constrained machine learning / differentiable optimization layers; safe and feasibility-guaranteeing neural network architectures for decision-making.
- Technical level: Advanced. The paper relies on convex analysis, homeomorphism and Lipschitz arguments, Jacobian invertibility, Clarke subgradients, and Polyak–Łojasiewicz (PL) inequalities, although the geometric intuition behind the method is accessible.
- Authors and venue: Philipp J. Schneider and Daniel Kuhn (Risk Analytics, Optimization Chair, EPFL); arXiv:2602.03461v1 [cs.LG], 03 Feb 2026, licensed CC BY 4.0.
- Scope in one sentence: The paper proposes a differentiable projection layer that maps any unconstrained neural network output into the strict interior of a convex feasible set while keeping the Jacobian full rank, avoiding the gradient saturation that afflicts orthogonal projection layers.
What This Paper Is About
Neural networks are not inherently constraint-aware, so in safety-critical or operational settings they can output decisions that violate hard constraints. A common fix is to append a layer that projects raw outputs onto the feasible set, but the standard Euclidean (orthogonal) projection collapses all infeasible points onto the lower-dimensional boundary. This makes the projection's Jacobian rank-deficient, zeros out gradients in directions orthogonal to active constraints, and causes optimization to stall. The paper's goal is a drop-in, closed-form, differentiable layer that enforces feasibility by construction without sacrificing gradient signal or model expressivity.
Key Contributions
- Soft-Radial Projection layer. A closed-form layer for convex sets with nonempty interior that maps the entire ambient space into the strict interior of the feasible set via a radial transformation along rays from an anchor point, differentiable almost everywhere.
- Geometric and gradient-dynamics analysis. Proofs that the layer is a homeomorphism from Euclidean space onto the interior, that its Jacobian is invertible wherever it exists (full rank almost everywhere), and hence that the rank deficiency of orthogonal projection is avoided.
- Optimization guarantees. Results on equivalence of optimal values, correspondence of stationary points, global optimality of interior stationary points for convex losses, a negative result showing a global PL inequality cannot hold in general, and standard SGD/subgradient convergence rates under bounded iterates.
- Preserved expressivity and empirical evaluation. A universal approximation theorem for constrained predictors into the feasible set, plus experiments on portfolio optimization and ride-sharing dispatch against softmax, orthogonal projection, DC3, and HardNet baselines.
Main Findings
- Gradient saturation is the bottleneck of projection layers. Orthogonal projection maps the exterior onto the boundary surface, so infinitesimal variations orthogonal to the boundary produce zero change in the output, nullifying those gradient components and causing optimization to stall or crawl along the boundary.
- Soft-radial projection is a homeomorphism. Under radial monotonicity, the map p: R^n → Int(C) is a homeomorphism (Theorem 2.8), giving a one-to-one parametrization of the feasible interior instead of the many-to-one collapse of orthogonal projection.
- Full-rank Jacobian almost everywhere. The soft-radial projection is differentiable almost everywhere and its Jacobian is invertible wherever it exists (Theorem 2.9), so gradients remain usable even when the raw output is far outside the feasible set.
- Optimal values are preserved. The infimum of the composite objective equals the infimum of the loss over the interior, which equals the infimum over the feasible set; minimizers correspond exactly, provided the constrained minimizer set intersects the interior (Theorem 3.1).
- No spurious interior stationary points. By the chain rule, the gradient of the composite objective is the transposed Jacobian times the loss gradient, so stationarity of the composite objective is equivalent to stationarity of the loss at the corresponding interior point (Proposition 3.2), implying global optimality of interior local minimizers for convex, C^1 losses (Corollary 3.3).
- Boundary optima push iterates outward. If a constrained minimizer lies on the boundary with nonzero loss gradient, no corresponding stationary point exists, and minimizing sequences must diverge in norm, pushing the projected point toward the boundary.
- A global PL inequality is impossible in general. Even for C = B_1(0), ℓ(x) = ||x||^2 and anchor u_0 = 0, no radial contraction r satisfying the stated assumptions makes the composite objective satisfy a global PL inequality (Lemma 3.4). This rules out direct global linear convergence arguments.
- Convergence still holds under bounded iterates. With iterates confined to a compact set, SGD with step size proportional to T^{-1/2} achieves O(T^{-1/2}) squared gradient norm in the smooth regime; stochastic subgradient descent with diminishing steps converges asymptotically to Clarke stationary points in the nonsmooth regime (Proposition 3.5).
- Expressivity is preserved. The constrained class {p ∘ g} is a universal approximator for continuous targets into the feasible set (Theorem 4.1), and uniform approximation error of the base class transfers to the constrained class scaled by the local Lipschitz constant of the projection.
- Closed-form computation for common sets. For polyhedral sets the ray-intersection scalar is computed in O(m) per sample and is fully parallelizable, in contrast to the constrained quadratic programs (typically solved iteratively, e.g. Dykstra's algorithm or ADMM) required by orthogonal projection onto intersections such as the capped simplex.
- Empirical numbers are not reported in the available content. The abstract states improved convergence behavior and solution quality over optimization- and projection-based baselines, but the truncated text ends mid-way through the portfolio optimization section, so no benchmark values, dataset sizes, or quantitative comparisons could be extracted. Extended experimental details are referenced to Appendix C.
Methodology in Plain English
The authors keep the model architecture unchanged and insert a special layer after the network's raw output. Instead of pushing a point to the nearest feasible point, they pick an anchor point strictly inside the feasible region and shoot a ray from that anchor through the network's raw output. The ray hits the boundary at one unique point; the layer then returns a point on the segment between the anchor and that boundary hit, positioned by a smooth scaling function of the distance from the anchor. Because the scaling factor never reaches 1, the output is always strictly inside the feasible region, and because the scaling is strictly increasing, distinct inputs map to distinct interior points — so no information is lost, unlike when an entire exterior region is squashed onto the boundary. They then analyze this map mathematically: showing it is a continuous bijection with a continuous inverse, that its Jacobian is invertible almost everywhere, and that solving the unconstrained problem through this layer recovers the constrained problem's optimum. They also test three concrete choices of the scaling function — rational, exponential, and hyperbolic forms parameterized by ε in (0,1) and λ > 0 — and evaluate with MLP and LSTM base networks on portfolio optimization and ride-sharing dispatch, backpropagating gradients straight through the projection.
Why This Matters
- Impact on research: The paper shifts attention from merely guaranteeing constraint satisfaction to the geometry of the resulting optimization landscape, identifying rank-deficient Jacobians as the mechanism behind stalled training in projection layers. It supplies both positive results (homeomorphism, invertibility, universal approximation) and an honest negative result (no global PL inequality), which sets realistic expectations for convergence theory of constructive constraint layers.
- Real-world applications named in the paper:
- Safety envelopes in autonomous driving.
- Actuator limits in robotics.
- Budget and capacity constraints in operations.
- Portfolio optimization under portfolio-weight constraints, and resource allocation for demand sharing / ride-sharing dispatch.
- Industry relevance: Because the layer is closed-form and vectorized for polyhedra, balls, and general convex level sets, it is cheaper than iterative QP-based projection layers and can be bolted onto existing architectures without penalty terms, post-hoc repair, or solver-in-the-loop training. That makes feasibility-by-construction more practical for deployment in regulated or safety-critical pipelines where infeasible outputs are unacceptable.
Future Directions
- Extending beyond convex feasible sets. The construction assumes the feasible set is closed, convex, and has nonempty interior; nonconvex constraint sets, common in robotics and control, would require a genuinely different radial or local argument.
- Characterizing when useful convergence conditions hold. Since a global PL inequality provably fails, an open question is whether PL-type or sharpness conditions hold on bounded regions, and how the choice among the rational, exponential, and hyperbolic contractions affects conditioning and convergence in practice.
- Handling boundary optima. The analysis shows that constrained optima on the boundary correspond to divergent minimizing sequences rather than stationary points, so practical schemes for finite-step boundary attainment and stopping criteria remain to be worked out.
- Anchor and hyperparameter selection. The anchor point and the contraction parameters ε and λ are design choices with no reported selection strategy, and their interaction with network width, depth, and learning-rate schedules is not addressed in the available content.
Target Audience
Researchers and practitioners working on constrained deep learning, differentiable optimization layers, decision-focused learning, and safe machine learning, as well as graduate students with background in convex analysis and optimization. Readers mainly interested in applied results should note that the quantitative experimental section is not included in the available text, while readers interested in the theoretical treatment of projection geometry, Jacobian conditioning, and universal approximation will find the core of the contribution here.
Authors’ abstract
Integrating hard constraints into deep learning is essential for safety-critical systems. Yet existing constructive layers that project predictions onto constraint boundaries face a fundamental bottleneck: gradient saturation. By collapsing exterior points onto lower-dimensional surfaces, standard orthogonal projections induce rank-deficient Jacobians, which nullify gradients orthogonal to active constraints and hinder optimization. We introduce Soft-Radial Projection, a differentiable reparameterization layer that circumvents this issue through a radial mapping from Euclidean space into the interior of the feasible set. This construction guarantees strict feasibility while preserving a full-rank Jacobian almost everywhere, thereby preventing the optimization stalls typical of boundary-based methods. We theoretically prove that the architecture retains the universal approximation property and empirically show improved convergence behavior and solution quality over state-of-the-art optimization- and projection-based baselines.