Publication Type

Conference Proceeding Article

Version

publishedVersion

Publication Date

7-2026

Abstract

Sequential decision-making using Markov Decision Process underpins many real-world applications. Both model-based and model-free methods have achieved strong results in these settings. However, real-world tasks must balance reward maximization with safety constraints, often conflicting objectives, that can lead to unstable min–max, adversarial optimization. A promising alternative is safety reachability analysis, which precomputes a forward-invariant safe state–action set, ensuring that an agent starting inside this set remains safe indefinitely. Yet, most reachability-based methods address only hard safety constraints, and little work extends reachability to cumulative cost constraints. To address this, first, we define a safety-conditioned reachability set that decouples reward maximization from cumulative safety cost constraints. Second, we show how this set enforces safety constraints without unstable min–max or Lagrangian optimization, yielding a novel offline safe RL algorithm that learns a safe policy from a fixed dataset without environment interaction. Finally, experiments on standard offline safe-RL benchmarks, and a real-world maritime navigation task demonstrate that our method matches or outperforms state-of-the-art baselines while maintaining safety.

Discipline

Artificial Intelligence and Robotics

Research Areas

Intelligent Systems and Optimization

Areas of Excellence

Sustainability

Publication

Proceedings of the Thirty-Sixth International Conference on Automated Planning and Scheduling, Dublin, Ireland, 2026 June 27 - July 2

Volume

36

First Page

581

Last Page

590

ISBN

9781577359104

Identifier

10.1609/icaps.v36i1.42876

Publisher

AAAI Press

City or Country

Dublin

Additional URL

https://doi.org/10.1609/icaps.v36i1.42876

Share

COinS