Skip to content
AI.info

Research

V-OCBF: Learning Safety Filters from Offline Data via Value-Guided Offline Control Barrier Functions

V-OCBF: Learning Safety Filters from Offline Data via Value-Guided Offline Control Barrier Functions Overview Research area: Safe reinforcement learning and control theory — specifically offline safe

V-OCBF: Learning Safety Filters from Offline Data via Value-Guided Offline Control Barrier Functions
arXiv
2512.10822
Published
2025-12-11
Authors
Mumuksh Tayal, Manan Tayal, Aditya Singh, Shishir Kolathaya, Ravi Prakash

AI summary

V-OCBF: Learning Safety Filters from Offline Data via Value-Guided Offline Control Barrier Functions

Overview

  • Research area: Safe reinforcement learning and control theory — specifically offline safe RL, neural Control Barrier Functions (CBFs), and Hamilton–Jacobi reachability.
  • Technical level: Advanced. The paper assumes familiarity with control-affine dynamics, Lie derivatives, Control Barrier Functions, Quadratic Programs, and offline RL machinery such as expectile regression (Implicit Q-Learning style objectives).
  • Scope: The paper proposes and empirically evaluates V-OCBF, a method that learns a neural Control Barrier Value Function purely from a static offline dataset and uses it inside a CBF-QP safety filter, without online interaction and without hand-designed barriers.

What This Paper Is About

Autonomous systems need controllers that never violate state-wise safety constraints, but safe offline reinforcement learning methods usually only enforce soft expected-cost constraints, which permit occasional violations. Control Barrier Functions can enforce strict forward invariance, but traditionally require either hand-crafted barrier functions or knowledge of the system dynamics. V-OCBF bridges these two lines of work: it learns a neural barrier function directly from offline demonstrations using a value-style recursion, then uses that learned barrier inside a Quadratic Program to filter a reference controller's actions at deployment time.

Key Contributions

  1. V-OCBF framework: A method for learning safe controllers entirely from offline demonstrations, with no online rollouts.
  2. Model-free finite-difference barrier recursion: The authors derive a recursion over consecutive dataset transitions and prove that, in the idealized setting without learning or approximation errors, adhering to this update provides a one-step forward-invariance safety guarantee for any control-affine system.
  3. Expectile-based learning objective: An objective that lets the learned barrier improve over the behavior policy in the dataset while never querying the barrier on out-of-distribution actions, restricting updates to the dataset-supported action set.
  4. Empirical gains across diverse systems: Across a low-dimensional Dubins' car task and high-dimensional Safety Gymnasium tasks, V-OCBF is reported to outperform constrained offline RL and neural CBF baselines in both safety and reward.

Main Findings

  • AGV collision avoidance results: On the 3-dimensional Dubins' car collision-avoidance task, BC+V-OCBF achieves 98.28 ± 0.54 percent safe episodes, an episode reward of 54.93 ± 0.46, and a safe set volume of 92.57 percent. This is the strongest result reported in Table 1 across all three metrics.
  • Comparison to offline RL baselines: BC reaches 48.92 ± 1.69 percent safe episodes with reward 20.45 ± 1.84; BEAR-Lag reaches 65.12 ± 0.24 with reward 13.85 ± 0.81; COptiDICE reaches 68.91 ± 0.32 with reward 15.33 ± 0.67; FISOR reaches 95.78 ± 0.2 with reward 52.33 ± 0.93.
  • Comparison to neural CBF baselines: BC+NCBF reaches 92.48 ± 0.60 percent safe episodes with reward 44.61 ± 2.58; BC+iDBF reaches 92.87 ± 0.73 with reward 48.23 ± 2.01; BC+CCBF reaches 93.56 ± 0.56 with reward 49.66 ± 2.34.
  • Safe set volume: Reported safe set volume is 42.51 percent for BC, 58.21 for BEAR-Lag, 62.32 for COptiDICE, 81.92 for BC+NCBF, 83.32 for BC+iDBF, 90.94 for BC+CCBF, 90.14 for FISOR, and 92.57 for BC+V-OCBF.
  • The CBF-QP layer matters: The authors attribute the higher safety of CBF-based methods to the QP filtering out unsafe actions even when the nominal controller is imperfect, and note that FISOR's lack of explicit safety filtering leads to lower safety rates than V-OCBF.
  • MuJoCo Safety Gymnasium scaling: Evaluated on Hopper, Swimmer, Half-Cheetah, Walker2D, and Ant with unknown dynamics, V-OCBF achieves the lowest safety violation rates across all tasks and near-zero violations on some tasks, while preserving reward levels compared to BC and outperforming iDBF, NCBF, and CCBF. The paper reports that NCBF suffers from optimization difficulties and iDBF enforces overly restrictive boundaries that suppress task performance in higher dimensions.
  • Qualitative Hopper behavior: In Hopper rollouts, the V-OCBF-controlled agent adapts its gait to keep forward velocity within the allowable safety threshold, indicating the learned barrier actively regulates unsafe accelerations.
  • Evaluation protocol: AGV results were evaluated over 500 episodes and 5 seed values; the MuJoCo Safety Gymnasium results used 500 randomly sampled initial states per environment and 5 seed values.
  • Guarantee caveat: The one-step forward-invariance guarantee holds only under idealized assumptions (exact barrier and exact Lie derivatives). The authors state that function approximation, finite-data estimation, and model error can break the closed-loop guarantee in practice.
  • Dynamics model placement: Using learned dynamics inside the CBVF learning loop would require evaluating terms under actions outside the dataset-supported action set, so the learned dynamics are used only at inference for Lie derivative computation. Section 5.2.2 is said to provide analysis demonstrating the limitations of using learned dynamics inside the CBVF learning loop, but the specific numbers from that analysis are not included in the available content.

Methodology in Plain English

The starting point is a standard control-affine system, ẋ = f(x) + g(x)u, with a safe set defined as the region where a barrier function B is non-negative and an unsafe region defined by a Lipschitz specification function ℓ.

Instead of solving the Hamilton–Jacobi–Bellman Variational Inequality directly, the authors write a finite-difference version of it: a state is safe if it is safe according to ℓ, or if its successor state is safe. This recursion, B(x_t) = min{ ℓ(x_t), max_{u_t} B(x_{t+1}) }, can be evaluated using only consecutive transitions from an offline dataset, so it needs no knowledge of f and g.

A naive fit of this recursion can collapse to a trivially small constant, so the authors use a discounted target, B_Target = (1 − γ)ℓ(x) + γ min{ ℓ(x), max_u B(x') } with γ → 1, which promotes contraction and stable learning.

The maximization over actions is the problem: in offline data you never see how alternative actions would have turned out, and querying the barrier on unobserved actions risks distortion. The authors therefore drop the online maximization and instead use expectile regression, with loss L^τ(y) = |τ − 1(y < 0)| y², over the actual dataset transitions. A higher expectile level τ puts more weight on underestimation errors than overestimation errors, pushing the learned barrier toward the upper envelope of safety values supported by the dataset. This borrows the intuition of Implicit Q-Learning (IQL).

The resulting network B_θ is the Value-guided Offline Control Barrier Function. Separately, a neural surrogate of the dynamics, x_{t+1} = x_t + (f_φ(x_t) + g_φ(x_t)u_t)Δt, is trained with a mean-squared-error one-step prediction loss on the same offline data — but only for use at inference.

At deployment, the learned dynamics supply the Lie derivatives L_f B(x) = ∇_x B_θ(x)ᵀ f_φ(x) and L_g B(x) = ∇_x B_θ(x)ᵀ g_φ(x), which feed a Quadratic Program that minimally modifies a reference action: minimize ||u − π_ref||² subject to L_f B(x) + L_g B(x)u + κ(B(x)) ≥ 0. In all experiments, Behavior Cloning serves as the nominal reference controller.

Why This Matters

  • Research impact: The paper links two previously separate research threads — distributionally-aware offline RL and control-theoretic safety certificates. It shows that a value-style backup (expectile regression) can substitute for online maximization when learning a barrier, and it argues for a clean separation between training-time (model-free) and inference-time (model-based) use of dynamics.
  • Practical safety guarantee framing: The work targets strict state-wise safety rather than soft expected-cost constraints, which matters for settings where a single violation is unacceptable.
  • Real-world applications:
    • Autonomous ground vehicles performing collision avoidance around static obstacles in bounded environments, as demonstrated with Dubins' car dynamics.
    • Legged and articulated robots (Hopper, Walker2D, Ant, Half-Cheetah, Swimmer) where velocity limits must be respected while maximizing task reward.
    • Aerial robotics and adaptive cruise control, which the paper cites as classic CBF application domains.
    • Household service robots, autonomous vehicles, and drones operating in complex, unstructured environments, cited by the authors as motivation.
  • Industry relevance: Because the method needs no online interaction and no hand-engineered barrier functions, it suits domains where high-fidelity simulators are unavailable or where real-system exploration is too costly or dangerous, and it plugs into a standard QP safety-filter architecture that already has real-time solver support.

Future Directions

  • Closing the guarantee gap: The one-step forward-invariance proof holds only under idealized, error-free assumptions. Addressing the combined effects of function approximation error, finite-data estimation error, and learned-dynamics model error is an open problem that the authors explicitly flag.
  • Better dynamics surrogates for inference: Since the QP depends on the learned f_φ and g_φ only at inference, improving the fidelity or robustness of that surrogate — and quantifying how its error propagates into safety — is a natural next step.
  • Understanding and tuning the expectile level: The paper defers a sensitivity analysis of τ to Appendix C.3 and a full derivation of the recursion to Appendix A.1. How τ should be selected across environments, and how it trades off conservatism against safe-set coverage, remains an open practical question.
  • Beyond behavior-induced barriers: The authors note that dropping the action maximization yields a behavior-induced barrier that reflects the safety profile of the data-generating policy and is typically sub-optimal, since other admissible actions could yield a larger safe set. Recovering that lost coverage without querying out-of-distribution actions is an unresolved direction.
  • Broader empirical scope: Evaluation covers the AGV task and five Safety Gymnasium tasks with the DSRL dataset. Whether the approach transfers to real hardware, to systems with different control-affine structure, or to datasets with sparse safety-critical transitions is not established in the available content.

Target Audience

  • Robotics and control researchers working on CBF-based safety filters, neural barrier functions, and Hamilton–Jacobi reachability, who will care about the model-free finite-difference recursion and the expectile-style barrier backup.
  • Safe reinforcement learning researchers interested in moving from soft expected-cost constraints to strict state-wise constraints in offline settings.
  • Practitioners deploying safety filters on real platforms, who need controllers synthesized without online interaction and who already use CBF-QP safety-filter architectures.
  • Graduate students and advanced readers with a background in nonlinear control, Lie derivatives, and offline RL, for whom the paper provides a concrete bridge between value-based offline learning and control-theoretic invariance guarantees. Beginners will likely need to consult the cited background on CBFs (Ames et al.) and CBVFs (Choi et al.) first.

Authors’ abstract

Ensuring safety in autonomous systems requires controllers that aim to satisfy state-wise constraints without relying on online interaction.While existing Safe Offline RL methods typically enforce soft expected-cost constraints, they struggle to ensure strict state-wise safety. Conversely, Control Barrier Functions (CBFs) offer a principled mechanism to enforce forward invariance, but often rely on expert-designed barrier functions or knowledge of the system dynamics. We introduce Value-Guided Offline Control Barrier Functions (V-OCBF), a framework that learns a neural CBF entirely from offline demonstrations. Unlike prior approaches, V-OCBF does not assume access to the dynamics model; instead, it derives a recursive finite-difference barrier update, enabling model-free learning of a barrier that propagates safety information over time. Moreover, V-OCBF incorporates an expectile-based objective that avoids querying the barrier on out-of-distribution actions and restricts updates to the dataset-supported action set. The learned barrier is then used with a Quadratic Program (QP) formulation to synthesize real-time safe control. Across multiple case studies, V-OCBF yields substantially fewer safety violations than baseline methods while maintaining strong task performance, highlighting its scalability for offline synthesis of safety-critical controllers without online interaction or hand-engineered barriers.

Read the original paper