← Back to overview

Paper Atlas

The verified bibliography behind the Safe Learning notes — 249 papers. Every entry was checked against a primary source (arXiv, DBLP, OpenReview or the publisher) on 2026-09-29; stars are an editorial importance rating (5 = landmark).

Start with the notes for your topic. The study guide gives a reading route and prerequisites. Use this atlas to find sources after studying a module; each module’s Key Papers table explains which references support its results.

Scroll the table horizontally to read all columns on a narrow screen.

PaperYearVenueTopicsImp.Contribution
LipKernel: Lipschitz-Bounded Convolutional Neural Networks via Dissipative Layers
Patricia Pauli, Ruigang Wang et al.
2026 Automatica 188 (2026) 112959, DOI 10.1016/j.automatica.20… PauliCertified NNs ★★★★★ Layer-wise direct parameterization of 1-D and 2-D CNNs (pooling, stride, dilation, zero padding) in which every layer satisfies an LMI implying incremental dissipativity, so the network is (Q,R)-Lipschitz by construction. Kernels come from a Roesser realization, a Cayley transform and closed-form 2-D Gramian-type solutions, and are evaluated in standard (kernel) form at inference.
Novel Quadratic Constraints for Extending LipSDP beyond Slope-Restricted Activations
Patricia Pauli, Aaron Havens et al.
2024 ICLR 2024 (conference paper; arXiv 2401.14033) PauliCertified NNs ★★★★★ Derives quadratic constraints for gradient-norm-preserving activations that are not slope-restricted (GroupSort, MaxMin, Householder), exploiting 1-Lipschitzness and sum preservation within each group. Yields SDPs (LipSDP-NSR) for l2 and l_inf-to-l1 Lipschitz bounds of feedforward, residual and implicit networks. Motivating example: LipSDP applied to the residual-ReLU rewrite of MaxMin only certifies rho=2, i.e. sqrt(2).
On Safety in Safe Bayesian Optimization
Christian Fiedler, Johanna Menn et al.
2024 Transactions on Machine Learning Research (TMLR), 2024 (a… Trimpe groupSafe exploration & BO ★★★★★ Shows that SafeOpt-type algorithms lose their guarantees in practice for two reasons: implementations swap the frequentist bound for a heuristic β (β≡2 in Berkenkamp 2016/Turchetta 2016, β≡3 in Helwa 2019/GoSafe, β≡4 in GoSafeOpt), and an RKHS-norm bound cannot be obtained from engineering knowledge. It proposes three fixes: Real-β-SafeOpt (the Abbasi-Yadkori-type bound is actually evaluated), LoSBO (safe sets built only from a Lipschitz constant and a hard noise bound, so safety is deterministic) and LoS-GP-UCB (a gridless variant over unions of balls for moderately high dimensions).
Recurrent Equilibrium Networks: Flexible Dynamic Models With Guaranteed Stability and Robustness
Max Revay, Ruigang Wang, Ian R. Manchester
2024 IEEE Transactions on Automatic Control 69(5):2855-2870 (2… PauliCertified NNs ★★★★★ Recurrent equilibrium networks: LTI dynamics in feedback with an implicit equilibrium layer, directly parameterized so every parameter value gives a contracting model satisfying prescribed incremental IQCs.
Bayesian Optimization with Safety Constraints: Safe and Automatic Parameter Tuning in Robotics
Felix Berkenkamp, Andreas Krause, Angela P. Schoellig
2023 Machine Learning 112(10):3713-3747 (print issue October 2… Safe exploration & BO ★★★★★ SafeOpt-MC: generalizes SafeOpt to several unknown safety constraints that differ from the objective by modeling objective and constraints as one GP over an index-augmented input, and adds contextual safe BO for transfer across tasks. Demonstrated on quadrotor controller tuning.
Data-Driven Safety Filters: Hamilton-Jacobi Reachability, Control Barrier Functions, and Predictive Methods for Uncertain Systems
Kim P. Wabersich, Andrew J. Taylor et al.
2023 IEEE Control Systems Magazine 43(5):137–177 Control-theoretic safety ★★★★★ Tutorial and survey treating safety filters as approximations of an ideal minimally invasive filter; develops HJ-reachability, CBF and predictive safety filters side by side with data-driven and uncertain variants and hardware case studies, and shows how the three relate and combine.
Direct Parameterization of Lipschitz-Bounded Deep Networks
Ruigang Wang, Ian R. Manchester
2023 ICML 2023 (arXiv 2301.11526, 'accepted to ICML 2023') PauliCertified NNs ★★★★★ Smooth and complete ('direct') parameterization of all feedforward networks satisfying the LipSDP certificate, extended to circular convolutions via the FFT; the 'sandwich layer' makes Lipschitz-bounded training unconstrained.
GoSafeOpt: Scalable safe exploration for global optimization of dynamical systems
Bhavya Sukhija, Matteo Turchetta et al.
2023 Artificial Intelligence, vol. 320, art. 103922 (2023); ar… Trimpe groupSafe exploration & BO ★★★★★ Alternates local safe exploration (SafeOpt in parameter space, LSE) with global exploration of possibly unsafe parameters (GE). Global steps stay safe through backup policies that are learned passively from every state visited in safe rollouts (Markov property, Prop. A.3) and triggered by a cheap online boundary condition. It is the first model-free method with safety and global-optimality guarantees that scales to high-dimensional state spaces (tested on a 7-DoF Franka Emika Panda arm).
Safe Planning in Dynamic Environments Using Conformal Prediction
Lars Lindemann, Matthew Cleaveland et al.
2023 IEEE Robotics and Automation Letters, vol. 8, no. 8, pp. … Certified NNs ★★★★★ Wraps any learned trajectory predictor (e.g. LSTM) of dynamic agents with split-conformal prediction regions calibrated on offline trajectories and inserts them into MPC constraints, giving distribution-free probabilistic safety guarantees for open- and closed-loop planning (CARLA pedestrian case study).
Safe Value Functions
Pierre-François Massiani, Steve Heim et al.
2023 IEEE Transactions on Automatic Control 68(5):2743-2757 (s… Trimpe groupConstrained RLControl-theoretic safety ★★★★★ Formalizes when penalizing failures yields value functions that are both optimal and safe (safe value functions). It proves a finite threshold penalty always exists (given a uniform bound on the time-to-failure outside the viability kernel), that larger penalties do not harm optimality, and shows how the threshold depends on rewards, discounting and time-to-failure.
Safe Learning in Robotics: From Learning-Based Control to Safe Reinforcement Learning
Lukas Brunke, Melissa Greeff et al.
2022 Annual Review of Control, Robotics, and Autonomous System… Constrained RLControl-theoretic safety ★★★★★ Unifying review of safe learning in robotics across learning-based control, safe RL and certification methods, organized by three safety levels (hard, probabilistic, encouraged constraints).
Safety Verification and Robustness Analysis of Neural Networks via Quadratic Constraints and Semidefinite Programming
Mahyar Fazlyab, Manfred Morari, George J. Pappas
2022 IEEE Transactions on Automatic Control 67(1):1-15 (2022),… PauliCertified NNs ★★★★★ Introduces the quadratic-constraint abstraction of activation functions and input sets, and verifies output-set containment and robustness of feedforward NNs by SDP via the S-procedure (DeepSDP).
Stability Analysis Using Quadratic Constraints for Systems With Neural Network Controllers
He Yin, Peter Seiler, Murat Arcak
2022 IEEE Transactions on Automatic Control 67(4):1980-1987 (2… PauliCertified NNs ★★★★★ Two stability theorems for LTI plants with NN controllers: (1) Lyapunov function plus local sector quadratic constraints from interval bounds on pre-activations, giving an SDP for an ellipsoidal inner approximation of the ROA; (2) extension to perturbations (unmodeled dynamics, slope-restricted nonlinearities, time delays) via IQCs.
Training Robust Neural Networks Using Lipschitz Bounds
Patricia Pauli, Anne Koch et al.
2022 IEEE Control Systems Letters 6:121-126 (2022), DOI 10.110… PauliCertified NNs ★★★★★ Puts the LipSDP certificate into NN training with ADMM: a backprop loss step alternates with an SDP 'Lipschitz step' that either penalizes the certified bound or enforces a prescribed bound L_des. The paper also gives counterexamples showing that the coupled-multiplier variant of LipSDP (Fazlyab et al. 2019, 'LipSDP-Network') is invalid, so only diagonal multipliers are admissible.
A predictive safety filter for learning-based control of constrained nonlinear dynamical systems
Kim P. Wabersich, Melanie N. Zeilinger
2021 Automatica 129:109597 (DOI 10.1016/j.automatica.2021.109597) Control-theoretic safety ★★★★★ Introduces the predictive safety filter: an MPC-like problem searches for a backup input sequence reaching a terminal safe set with first input as close as possible to the RL input, turning a constrained system into an unconstrained 'safe system' that any RL algorithm can run on. Extended to data-driven models with state- and input-dependent uncertainty, with constraint satisfaction at a prescribed probability.
Beta-CROWN: Efficient Bound Propagation with Per-neuron Split Constraints for Neural Network Robustness Verification
Shiqi Wang, Huan Zhang et al.
2021 NeurIPS 2021 (Advances in Neural Information Processing S… Certified NNs ★★★★★ Encodes branch-and-bound ReLU split constraints directly inside bound propagation through optimizable multipliers beta (jointly with the alpha relaxation slopes), matching the LP-with-splits bound at GPU speed; sound and complete with BaB and up to three orders of magnitude faster than LP-based BaB; engine of alpha,beta-CROWN (VNN-COMP 2021 winner).
Linear systems with neural network nonlinearities: Improved stability analysis via acausal Zames-Falb multipliers
Patricia Pauli, Dennis Gramlich et al.
2021 60th IEEE Conference on Decision and Control (CDC 2021), … Pauli ★★★★★ Abstracts an NN in discrete-time feedback with an LTI plant using dynamic IQCs: full-block circle/Yakubovich multipliers combined with acausal FIR Zames-Falb multipliers, for which a hard-IQC factorization is proved. Yields LMI certificates for local stability and ellipsoidal inner approximations of the region of attraction; repeated-nonlinearity variants (ZF-R, ZF-RL) for bias-free networks.
Practical and Rigorous Uncertainty Bounds for Gaussian Process Regression
Christian Fiedler, Carsten W. Scherer, Sebastian Trimpe
2021 Proceedings of the AAAI Conference on Artificial Intellig… Trimpe groupSafe exploration & BO ★★★★★ Replaces the information-gain scaling of Chowdhury & Gopalan with an a-posteriori, data-dependent β_N that can be evaluated numerically and comes close to common heuristics. It also gives a log-det-free bound for independent noise (Prop. 2) and robustness results for misspecified kernels (Props. 3-4, Thm 5) that degrade gracefully.
Learning-Based Model Predictive Control: Toward Safe Learning in Control
Lukas Hewing, Kim P. Wabersich et al.
2020 Annual Review of Control, Robotics, and Autonomous System… Control-theoretic safety ★★★★★ Review of learning-based MPC in three categories: (i) learning the prediction model from data, (ii) inferring MPC parameters (cost, constraints) for closed-loop performance, (iii) using MPC to give learning-based controllers constraint-satisfaction guarantees. Emphasizes robust and stochastic formulations.
Responsive Safety in Reinforcement Learning by PID Lagrangian Methods
Adam Stooke, Joshua Achiam, Pieter Abbeel
2020 ICML 2020 Constrained RL ★★★★★ Recasts the Lagrange-multiplier update as integral control of the cost and adds proportional and derivative terms to damp oscillation and overshoot. Adds reward-scale invariance through a 1/(1+lambda) rescaled objective; implemented as constraint-controlled PPO (CPPO-PID).
A General Safety Framework for Learning-Based Control in Uncertain Robotic Systems
Jaime F. Fisac, Anayo K. Akametalu et al.
2019 IEEE Transactions on Automatic Control 64(7):2737–2752 Control-theoretic safety ★★★★★ Wraps an arbitrary learning controller in an HJ-reachability safety supervisor. A least-restrictive law intervenes only at the boundary of the computed safe set; a GP-based Bayesian mechanism updates disturbance bounds and triggers intervention when confidence in the model-based guarantee drops. Policy-gradient RL on a quadrotor never crashes and retracts from a strong unmodeled fan disturbance, although the safety analysis uses a simplified point-mass model.
Benchmarking Safe Exploration in Deep Reinforcement Learning
Alex Ray, Joshua Achiam, Dario Amodei
2019 OpenAI technical report / preprint (Safety Gym); not on a… Constrained RL ★★★★★ Introduced the Safety Gym MuJoCo suite (Point/Car/Doggo robots; Goal/Button/Push tasks with hazards) and standardized safe-exploration metrics, notably the training cost rate as a safety-regret measure. Benchmarked PPO/TRPO, their Lagrangian variants and CPO.
Bridging Hamilton-Jacobi Safety Analysis and Reinforcement Learning
Jaime F. Fisac, Neil F. Lugovoy et al.
2019 2019 International Conference on Robotics and Automation … Control-theoretic safety ★★★★★ Shows that time-discounting the min-over-time safety objective yields a Bellman backup that is a contraction, so tabular and deep RL can approximate HJ safe sets and safety policies without a model or a grid; the undiscounted safety value is recovered as gamma->1.
Certified Adversarial Robustness via Randomized Smoothing
Jeremy M. Cohen, Elan Rosenfeld, J. Zico Kolter
2019 ICML 2019 (PMLR 97, pp. 1310-1320) Certified NNs ★★★★★ Proves a tight l2 robustness radius for the Gaussian-smoothed version of any base classifier and gives Monte-Carlo PREDICT/CERTIFY procedures with statistical guarantees; first certified defense at ImageNet scale (49% certified top-1 at l2 radius 0.5).
Constrained Reinforcement Learning Has Zero Duality Gap
Santiago Paternain, Luiz F. O. Chamon et al.
2019 NeurIPS 2019 Constrained RL ★★★★★ Proves that despite non-convexity in policy space, CMDPs have zero duality gap under Slater's condition, so they can be solved exactly in the convex dual. Bounds the dual suboptimality of epsilon-universal policy parametrizations linearly in epsilon and analyzes a primal-dual algorithm.
Control Barrier Functions: Theory and Applications
Aaron D. Ames, Samuel Coogan et al.
2019 2019 18th European Control Conference (ECC), pp. 3420–343… Control-theoretic safety ★★★★★ Tutorial overview of CBF theory: barrier functions and Nagumo-type invariance, zeroing CBFs with extended class-K functions, CLF-CBF QPs, exponential CBFs for high relative degree, and multi-robot applications. The standard entry point for the CBF-QP 'safety filter' view u*=argmin||u-k(x)||.
Efficient and Accurate Estimation of Lipschitz Constants for Deep Neural Networks
Mahyar Fazlyab, Alexander Robey et al.
2019 NeurIPS 2019 (arXiv 1906.04893; v2 dated 14 Jan 2023) PauliCertified NNs ★★★★★ Treats slope-restricted activations (gradients of convex potentials) with incremental quadratic constraints, turning l2-Lipschitz estimation into an SDP (LipSDP) with Neuron and Layer variants and network splitting for scalability.
End-to-End Safe Reinforcement Learning through Barrier Functions for Safety-Critical Continuous Control Tasks
Richard Cheng, Gábor Orosz et al.
2019 Proceedings of the AAAI Conference on Artificial Intellig… Control-theoretic safety ★★★★★ RL-CBF adds a CBF-QP compensator, computed from a GP model of the unknown dynamics, to a model-free RL policy (TRPO or DDPG), giving high-probability safety during learning; past CBF corrections are distilled into the policy to guide exploration. Validated on an inverted pendulum and car-following.
On the Effectiveness of Interval Bound Propagation for Training Verifiably Robust Models
Sven Gowal, Krishnamurthy Dvijotham et al.
2019 ICCV 2019, pp. 4841-4850, DOI 10.1109/ICCV.2019.00494 (pu… Certified NNs ★★★★★ Shows that cheap interval bound propagation (IBP), with a mixed nominal/worst-case loss and epsilon/kappa curricula, trains networks whose IBP bounds become tight, beating more sophisticated relaxations in verified accuracy and scaling to downscaled ImageNet.
Reward Constrained Policy Optimization
Chen Tessler, Daniel J. Mankowitz, Shie Mannor
2019 ICLR 2019 (poster) Constrained RL ★★★★★ Multi-timescale actor-critic trained on a lambda-penalized reward (so a TD critic works even for general constraints) while lambda is updated on the slowest timescale from Monte-Carlo constraint estimates. Proves a.s. convergence to a feasible fixed point, assuming local minima of the discounted penalty are feasible.
Efficient Neural Network Robustness Certification with General Activation Functions
Huan Zhang, Tsui-Wei Weng et al.
2018 NeurIPS 2018 (Advances in Neural Information Processing S… Certified NNs ★★★★★ CROWN: sandwiches each activation between two linear functions on its pre-activation interval (adaptive, possibly different slopes for upper and lower bound) and back-substitutes layer by layer to obtain closed-form linear lower/upper bounds on every output for general activations (ReLU, tanh, sigmoid, arctan); generalizes Fast-Lin.
Risk-Constrained Reinforcement Learning with Percentile Risk Criteria
Yinlam Chow, Mohammad Ghavamzadeh et al.
2018 Journal of Machine Learning Research 18(167):1-51 (2018) Constrained RL ★★★★★ Formulates percentile-risk-constrained MDPs with chance (VaR) or CVaR constraints on the cumulative cost, derives gradients of the Lagrangian (with the VaR parameter nu optimized jointly) and gives policy-gradient and actor-critic algorithms with multi-timescale convergence to locally optimal (saddle) points.
Safe Reinforcement Learning via Shielding
Mohammed Alshiekh, Roderick Bloem et al.
2018 Proceedings of the AAAI Conference on Artificial Intellig… Control-theoretic safety ★★★★★ Introduces shields: reactive systems synthesized from a temporal-logic safety specification and a finite abstraction of the environment. A shield sits before the agent (restricting the action set) or after it (overriding unsafe actions); shields are correct and minimally interfering, and conditions are given under which the learner keeps its convergence guarantees.
Constrained Policy Optimization
Joshua Achiam, David Held et al.
2017 ICML 2017 (PMLR 70) Constrained RL ★★★★★ First general-purpose trust-region policy search for CMDPs with near-constraint-satisfaction guarantees at every iteration. Proves a policy-performance bound in terms of expected TV divergence under the old state distribution and solves the linearized QCQP through its low-dimensional dual (with a recovery step when infeasible).
Control Barrier Function Based Quadratic Programs for Safety Critical Systems
Aaron D. Ames, Xiangru Xu et al.
2017 IEEE Transactions on Automatic Control 62(8):3861–3876 (D… Control-theoretic safety ★★★★★ Defines reciprocal and zeroing control barrier functions ('two novel generalizations of barrier functions'), whose Lyapunov-like conditions are affine in u and imply forward invariance of C={h>=0}. Unifies them with exponentially stabilizing CLFs in one QP (CLF softened by a slack, CBF hard), proves the QP controller is locally Lipschitz under a relative-degree-one condition, and demonstrates adaptive cruise control and lane keeping with actuator bounds.
Hamilton-Jacobi Reachability: A Brief Overview and Recent Advances
Somil Bansal, Mo Chen et al.
2017 2017 IEEE 56th Conference on Decision and Control (CDC), … Control-theoretic safety ★★★★★ Tutorial on HJ reachability: differential-game formulation with non-anticipative strategies, reachable sets (BRS) versus tubes (BRT), forward versus backward sets, control and disturbance roles, numerical tools (level-set toolbox, helperOC, GPU implementation), and decomposition / high-dimensional techniques.
On Kernelized Multi-armed Bandits
Sayak Ray Chowdhury, Aditya Gopalan
2017 ICML 2017 (PMLR 70:844-853) Safe exploration & BO ★★★★★ Derives a self-normalized concentration inequality for martingales in (possibly infinite-dimensional) RKHSs via a 'double mixture' argument and turns it into a confidence bound that holds simultaneously for all x and t. Proposes IGP-UCB and GP-TS with frequentist regret bounds.
Reluplex: An Efficient SMT Solver for Verifying Deep Neural Networks
Guy Katz, Clark Barrett et al.
2017 CAV 2017 (LNCS 10426, pp. 97-117, DOI 10.1007/978-3-319-6… Certified NNs ★★★★★ Extends the simplex method with lazy case-splitting on ReLU constraints, giving a sound and complete SMT procedure for linear input-output properties of ReLU networks; verified ACAS Xu collision-avoidance networks an order of magnitude larger than before and notes the verification problem is NP-complete.
Safe Model-based Reinforcement Learning with Stability Guarantees
Felix Berkenkamp, Matteo Turchetta et al.
2017 NeurIPS (NIPS) 2017 Safe exploration & BOControl-theoretic safety ★★★★★ Certifies an inner approximation of the region of attraction (a Lyapunov level set) of a policy using GP confidence intervals on the dynamics plus a discretization/Lipschitz argument, optimizes the policy to enlarge it, and proves that data can be collected safely inside it to expand the certified region.
Automatic LQR Tuning Based on Gaussian Process Global Optimization
Alonso Marco, Philipp Hennig et al.
2016 IEEE International Conference on Robotics and Automation … Trimpe group ★★★★★ Parametrizes a state-feedback controller through LQR weights on a nominal linear model and tunes those weights directly from experimental cost using Entropy Search. Demonstrated with 2D and 4D tuning on a 7-DoF arm balancing an inverted pole.
Safe Exploration in Finite Markov Decision Processes with Gaussian Processes
Matteo Turchetta, Felix Berkenkamp, Andreas Krause
2016 NeurIPS (NIPS) 2016, pp. 4305-4313 Constrained RLSafe exploration & BO ★★★★★ SafeMDP: safe exploration of deterministic finite MDPs with an unknown safety function modeled by a GP; expands a Lipschitz-certified safe set from lower confidence bounds, restricted to states that are reachable and from which the agent can return.
A Comprehensive Survey on Safe Reinforcement Learning
Javier García, Fernando Fernández
2015 Journal of Machine Learning Research 16:1437-1480 (2015) Constrained RL ★★★★★ Classic taxonomy of safe RL: (i) modifying the optimality criterion (worst-case, risk-sensitive, constrained) and (ii) modifying the exploration process (external knowledge, risk-directed exploration).
Safe Exploration for Optimization with Gaussian Processes
Yanan Sui, Alkis Gotovos et al.
2015 ICML 2015 (PMLR 37:997-1005) Safe exploration & BO ★★★★★ Formulates safe BO (every query must satisfy f(x_t) >= h for an unknown f) and proposes SafeOpt, which grows a Lipschitz-certified safe set from a safe seed while localizing the optimum by uncertainty sampling over expanders and potential maximizers. Proves safety w.h.p. and convergence to the epsilon-reachable optimum with an explicit sample complexity.
Provably safe and robust learning-based model predictive control
Anil Aswani, Humberto Gonzalez et al.
2013 Automatica 49(5):1216–1226 (DOI 10.1016/j.automatica.2013… Control-theoretic safety ★★★★★ LBMPC keeps two models: a nominal linear model with bounded uncertainty enters tube-MPC constraints (deterministic robust feasibility and constraint satisfaction), and a learned 'oracle' model enters only the cost (performance). With sufficient excitation the LBMPC input converges to that of MPC with the true model.
Improved Algorithms for Linear Stochastic Bandits
Yasin Abbasi-Yadkori, David Pal, Csaba Szepesvari
2011 NeurIPS (NIPS) 2011 (Advances in Neural Information Proce… Safe exploration & BO ★★★★★ Proves a self-normalized tail inequality for vector-valued martingales (method of mixtures) and derives anytime confidence ellipsoids for regularized least squares. Uses them to improve OFUL-type linear bandit regret by a log factor and to obtain constant regret for a modified UCB.
Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design
Niranjan Srinivas, Andreas Krause et al.
2010 ICML 2010 (extended version: IEEE Transactions on Informa… Safe exploration & BO ★★★★★ Gives the first regret bounds for GP-UCB when f is a GP sample or has bounded RKHS norm, expressing cumulative regret through the maximum information gain gamma_T. Bounds gamma_T for linear, squared-exponential and Matern kernels via operator spectra, connecting BO to experimental design.
The Exact Feasibility of Randomized Solutions of Uncertain Convex Programs
Marco C. Campi, Simone Garatti
2008 SIAM Journal on Optimization, vol. 19, no. 3, pp. 1211-12… Certified NNs ★★★★★ For convex scenario programs with d decision variables and N i.i.d. sampled constraints, gives the exact, distribution-free bound on the probability that the scenario solution violates more than an epsilon-mass of constraints (tight for fully-supported problems), sharpening the Calafiore-Campi bound.
A time-dependent Hamilton-Jacobi formulation of reachable sets for continuous dynamic games
Ian M. Mitchell, Alexandre M. Bayen, Claire J. Tomlin
2005 IEEE Transactions on Automatic Control 50(7):947–957 Control-theoretic safety ★★★★★ Proves that the backward reachable set of a nonlinear two-player differential game is the zero sublevel set of the viscosity solution of a time-dependent HJI PDE; a min[0,H] term 'freezes' trajectories once they enter the target. Makes level-set methods directly usable for safety verification, with aircraft conflict detection as application.
Approximately Optimal Approximate Reinforcement Learning
Sham M. Kakade, John Langford
2002 Proceedings of the 19th International Conference on Machi… Constrained RL ★★★★★ Introduces the performance-difference lemma and conservative policy iteration (mixture policy updates with a guaranteed improvement bound). This lemma is the starting point of every trust-region surrogate bound used in TRPO, CPO, PCPO, FOCOPS and CUP.
Constrained Markov Decision Processes
Eitan Altman
1999 Book, Chapman & Hall/CRC (Stochastic Modeling series), 19… Constrained RL ★★★★★ Unified theory of CMDPs: occupation-measure linear programs, Lagrangian/saddle-point duality, and existence of optimal stationary randomized policies. Shows some optimal stationary policy needs at most K randomizations (K = number of constraints).
Linear Matrix Inequalities in System and Control Theory
Stephen Boyd, Laurent El Ghaoui et al.
1994 SIAM Studies in Applied Mathematics vol. 15, SIAM, Philad… Pauli ★★★★★ The canonical reference for LMIs in control: Schur complements, the S-procedure, Lyapunov and Riccati inequalities as LMIs, invariant ellipsoids, positive-real/bounded-real lemmas, and the reduction of robust-control tests to convex feasibility. LipSDP, Pauli's training-under-SDP-constraints, Yin's ROA analysis and Fiedler's learning-enhanced synthesis all use exactly these manipulations.
Dissipative dynamical systems part I: General theory
Jan C. Willems
1972 Archive for Rational Mechanics and Analysis 45(5):321-351… Pauli ★★★★★ Founding paper of dissipativity theory: supply rates, storage functions, available storage and required supply, existence and convexity of storage functions, quadratic storage for linear systems. Every dissipativity-based Lipschitz/stability certificate in the Pauli/Manchester/Scherer line (LipKernel, 1-D CNN layers, RENs, closed-loop dissipativity synthesis) instantiates this framework.
A Decade of Bayesian Optimization for Controller Tuning and Robot Learning: Tutorial, Review, and Future Prospects
David Stenger, Paul Brunzema et al.
2026 arXiv preprint 2609.09403 (8 Sep 2026; under review) Trimpe group ★★★★ A practitioner-oriented tutorial and review of BO for controller tuning and robot learning, positioning BO against deep RL and data-driven control. It surveys the diverse BO variants (constrained/safe, crash-constraint, multi-objective, preferential, contextual, time-varying, local, multi-fidelity), and starts a lightweight benchmark suite with metrics and best practices for control.
Breaking Safety Paradox with Feasible Dual Policy Iteration
Yujie Yang, Jinglin Teh et al.
2026 ICLR 2026 (Tsinghua University / Westlake University / Su… Constrained RL ★★★★ Identifies the 'safety paradox': as a policy gets safer, constraint-violating samples become rare, so the estimated feasibility function (constraint decay function, CDF) degrades, which in turn harms the safety of the next policy update. Proves (Thm. 1) that the relative Monte-Carlo estimation error of the CDF at infeasible states grows with the mean and variance of the time-to-first-violation N. Proposes feasible dual policy iteration (FDPI): a second 'dual' policy pi_d is trained to maximize constraint violation (maximize its own action-feasibility value G_d) while staying KL-close to the primal policy pi_p; data from both policies is pooled and re-weighted by truncated per-trajectory importance-sampling ratios. Instantiated on top of feasible policy iteration + SAC as SAC-FDPI.
Convolutional neural networks as 2-D systems
Dennis Gramlich, Patricia Pauli et al.
2026 Automatica 187 (2026) 112876, DOI 10.1016/j.automatica.20… PauliCertified NNs ★★★★ 2-D convolutional layers have Roesser realizations, so a fully convolutional CNN is a 2-D Lur'e system; develops 2-D dissipativity/IQC analysis and an LMI Lipschitz bound valid for images of arbitrary size, extended to CNNs with fully connected heads.
Dyna-Style Safety Augmented Reinforcement Learning: Staying Safe in the Face of Uncertainty
Artur Eisele, Bernd Frauenknecht et al.
2026 arXiv preprint (28 Apr 2026) Trimpe group ★★★★ Dyna-SAuR learns a control policy and a scalable safety filter together in a Dyna loop with an uncertainty-aware model; the filter is a learned state-dependent hyperplane constraint on actions (exact for control-affine dynamics). It avoids both failures and model-uncertain regions, and a better model enlarges the safe-and-certain set, making the filter less conservative.
LipNeXt: Scaling up Lipschitz-based Certified Robustness to Billion-parameter Models
Kai Hu, Haoqi Hu, Matt Fredrikson
2026 ICLR 2026 (arXiv 2601.18513) PauliCertified NNs ★★★★ First constraint-free and convolution-free 1-Lipschitz architecture: orthogonal weights optimized directly on the orthogonal manifold, a Spatial Shift Module instead of convolutions, beta-Abs activations and L2 spatial pooling; scales to 1-2B parameters.
Neural Networks in the Loop: Learning with Stability and Robustness Guarantees
Ian R. Manchester, Ruigang Wang, Nicholas H. Barbara
2026 Annual Review of Control, Robotics, and Autonomous System… Certified NNs ★★★★ Review of static and dynamic NN model structures with built-in certificates (Lipschitz bounds, contraction/stability, robust invertibility) obtained from IQC-based certification plus direct parameterizations for unconstrained training, and their use as observers, stabilizing policy parameterizations and Lyapunov/storage/value functions.
Safe Exploration via Policy Priors
Manuel Wendl, Yarden As et al.
2026 ICLR 2026 Constrained RLSafe exploration & BO ★★★★ SOOPER uses a conservative prior policy (from offline data or simulation) as a pessimistic fallback: the agent explores optimistically with a probabilistic dynamics model and hands control to the prior whenever realized cost plus the prior's pessimistic cost-to-go would exceed the budget.
Synthesizing Neural Network Controllers with Closed-Loop Dissipativity Guarantees
Neelay Junnarkar, Murat Arcak, Peter Seiler
2026 Automatica 193 (2026) 113185, DOI 10.1016/j.automatica.20… Pauli ★★★★ Synthesizes recurrent implicit NN controllers that maximize RL reward subject to closed-loop dissipativity for uncertain plants (LTI plus IQC-described uncertainty), using a convexified LMI inside projection-based training.
Uncertainty-Aware Predictive Safety Filters for Probabilistic Neural Network Dynamics
Bernd Frauenknecht, Lukas Kesper et al.
2026 Reinforcement Learning Journal 2026 / Reinforcement Learn… Trimpe groupControl-theoretic safety ★★★★ UPSi is a predictive safety filter built on the same probabilistic-ensemble (PE) neural dynamics used in Dyna-style MBRL. It over-approximates robust reachable sets with ellipsoidal tubes and adds a certainty constraint (average Kalman gain of aleatoric vs. total variance) that keeps plans where the model is accurate. Inside MBPO it gives much safer exploration than prior NN-based PSFs at comparable return.
ActSafe: Active Exploration with Safety Constraints for Reinforcement Learning
Yarden As, Bhavya Sukhija et al.
2025 ICLR 2025 Constrained RLSafe exploration & BO ★★★★ Model-based safe exploration that keeps a pessimistic (with respect to a calibrated model set) safe set of policies, expands it by optimistic intrinsic exploration of epistemic uncertainty, then exploits; a practical Dreamer-style version handles visual control.
Approximate non-linear model predictive control with safety-augmented neural networks
Henrik Hose, Johannes Köhler et al.
2025 IEEE Transactions on Control Systems Technology 33(6):249… Trimpe group ★★★★ The NN predicts the whole MPC input sequence. Online, a single NN evaluation plus forward simulation checks whether it is feasible and no worse than a shifted safe candidate with terminal controller; otherwise the candidate is applied. This gives deterministic constraint satisfaction and convergence despite approximation errors.
SDP-CROWN: Efficient Bound Propagation for Neural Network Verification with Tightness of Semidefinite Programming
Hong-Ming Chiu, Hao Chen et al.
2025 ICML 2025; arXiv 2506.06665 Certified NNs ★★★★ Derives from SDP principles a new linear relaxation for l2-ball input sets that captures inter-neuron coupling with a single extra parameter per layer (provably up to a sqrt(n) factor tighter than per-neuron box-based bounds); inside alpha-CROWN it approaches SDP tightness on models with up to 65k neurons / 2.47M parameters.
Viability of Future Actions: Robust Safety in Reinforcement Learning via Entropy Regularization
Pierre-François Massiani, Alexander von Rohr et al.
2025 ECML-PKDD 2025, LNCS 'Machine Learning and Knowledge Disc… Trimpe group ★★★★ Shows empirically that entropy regularization in constrained RL biases optimal policies toward states with many future viable actions, which gives robustness to action noise. It proves that failure penalties approximate the entropy-regularized constrained problem arbitrarily closely (at the price of δ-safety instead of exact viability), so standard model-free algorithms such as SAC can be used.
A Review of Safe Reinforcement Learning: Methods, Theories, and Applications
Shangding Gu, Long Yang et al.
2024 IEEE Transactions on Pattern Analysis and Machine Intelli… Constrained RL ★★★★ Comprehensive review of safe RL methods (model-free, model-based, multi-agent), theory (including a comparison of sample complexities of CMDP algorithms), applications and benchmarks, organized around five questions called '2H3W'.
Datasets and Benchmarks for Offline Safe Reinforcement Learning
Zuxin Liu, Zijian Guo et al.
2024 Journal of Data-centric Machine Learning Research (DMLR),… Constrained RL ★★★★ Offline safe RL benchmark: expert safe policies, D4RL-style datasets with post-processing filters for 38 tasks (Safety-Gymnasium, Bullet-Safety-Gym, MetaDrive), and reference implementations (OSRL) of BC, CDT, BCQ-Lag, BEAR-Lag, CPQ and COptiDICE.
Event-Triggered Safe Bayesian Optimization on Quadcopters
Antonia Holzapfel, Paul Brunzema, Sebastian Trimpe
2024 6th Annual Learning for Dynamics & Control Conference (L4… Trimpe group ★★★★ Event-Triggered SafeOpt (ETSO) monitors whether new cost observations are consistent with the GP surrogate. When a time-variation such as wear causes a significant deviation, it reverts to a safe backup controller (Assumption 1: safe for all system modes), resets data and safe set, and restarts safe exploration; tested on quadcopter tuning in simulation and hardware.
Lipschitz constant estimation for general neural network architectures using control tools
Patricia Pauli, Dennis Gramlich, Frank Allgöwer
2024 arXiv preprint 2405.01125 (v2, 25 Nov 2024); no journal r… PauliCertified NNs ★★★★ Treats a feedforward NN as a time-varying system (layer index as time) and uses a dynamic-programming chain of quadratic value functions to split LipSDP into one LMI per layer (GLipSDP); covers 1-D/2-D/N-D convolutions (Roesser), state-space-model layers, residual, pooling, GroupSort and fully connected layers.
Lipschitz Safe Bayesian Optimization for Automotive Control
Johanna Menn, Pietro Pelizzari et al.
2024 63rd IEEE Conference on Decision and Control (CDC 2024), … Trimpe group ★★★★ MCLoSBO extends LoSBO to multiple safety constraints (following SafeOpt-MC) with a Lipschitz constant and noise bound per constraint. It tuned a self-driving car's trajectory-tracking controller in simulation and on a real test vehicle without leaving the track or violating other constraints.
Log Barriers for Safe Black-box Optimization with Application to Safe Reinforcement Learning
Ilnura Usmanova, Yarden As et al.
2024 Journal of Machine Learning Research 25(171):1-54 (2024) Safe exploration & BO ★★★★ LB-SGD runs SGD with a carefully chosen adaptive step size on a log-barrier surrogate so that every iterate stays feasible with high probability; complete convergence analysis for non-convex, convex and strongly convex problems with first- and zeroth-order (one-point) feedback, plus safe policy search.
Lyapunov-stable Neural Control for State and Output Feedback: A Novel Formulation
Lujie Yang, Hongkai Dai et al.
2024 Proceedings of the 41st International Conference on Machi… Control-theoretic safety ★★★★ Trains NN controllers (and NN observers for output feedback) with Lyapunov certificates using fast empirical falsification; a new region-of-attraction formulation requires Lyapunov decrease only inside an invariant sublevel set, verified afterwards with alpha,beta-CROWN branch-and-bound. Larger verified ROAs and, per the authors, the first formally verified Lyapunov-stable neural output-feedback control.
Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model
Yinan Zheng, Jianxiong Li et al.
2024 ICLR 2024 Constrained RL ★★★★ FISOR: turns hard (zero-violation) offline safe RL into identifying the largest feasible region by Hamilton-Jacobi reachability (reversed expectile regression), maximizing reward inside it and minimizing violation outside; the decoupled optimum is weighted behavior cloning, extracted with an energy-guided diffusion policy.
Safe RLHF: Safe Reinforcement Learning from Human Feedback
Josef Dai, Xuehai Pan et al.
2024 ICLR 2024 (spotlight) Constrained RL ★★★★ Decouples human preferences into helpfulness (reward model) and harmlessness (cost model with an extra sign-classification term) and fine-tunes an LLM (Alpaca-7B) with a Lagrangian-relaxed cost constraint.
SafeDreamer: Safe Reinforcement Learning with World Models
Weidong Huang, Jiaming Ji et al.
2024 ICLR 2024 Constrained RL ★★★★ Integrates Lagrangian methods into DreamerV3 world-model planning: online safety-reward planning with a constrained CEM (OSRP), its PID-Lagrangian version (OSRP-Lag), and background planning with an augmented-Lagrangian actor loss (BSRP-Lag).
The Safety Filter: A Unified View of Safety-Critical Control in Autonomous Systems
Kai-Chieh Hsu, Haimin Hu, Jaime F. Fisac
2024 Annual Review of Control, Robotics, and Autonomous System… Control-theoretic safety ★★★★ Unifies HJ least-restrictive filters, CBF filters, model-predictive shielding, forward-reachable-set filters and learned filters. Each filter is split into a fallback policy, a safety monitor and an intervention scheme; guarantees follow as corollaries of a 'Universal Safety Filter Theorem'.
Transductive Active Learning: Theory and Applications
Jonas Huebotter, Bhavya Sukhija et al.
2024 NeurIPS 2024 Safe exploration & BO ★★★★ Introduces transductive active learning (ITL/VTL: sample in an accessible region to reduce uncertainty about prediction targets possibly outside it) with variance-convergence and RKHS approximation guarantees. Instantiated for safe BO with provable safety and convergence to the optimum of the largest reachable safe set, without Lipschitz constants.
(Certified!!) Adversarial Robustness for Free!
Nicholas Carlini, Florian Tramèr et al.
2023 ICLR 2023; arXiv 2206.10550 Certified NNs ★★★★ Diffusion denoised smoothing: an off-the-shelf diffusion model used as a one-shot denoiser in front of an off-the-shelf classifier inside randomized smoothing, with no training or fine-tuning; SOTA certified l2 robustness on ImageNet (71% at radius 0.5, +14 pp over the prior SOTA).
A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification
Anastasios N. Angelopoulos, Stephen Bates
2023 Foundations and Trends in Machine Learning, vol. 16, no. … Certified NNs ★★★★ Tutorial on split conformal prediction and its extensions (conformal risk control, covariate-shift weighting, distribution drift, full CP) with code; the standard reference for distribution-free prediction sets used in learning-based control.
A Unified Algebraic Perspective on Lipschitz Neural Networks
Alexandre Araujo, Aaron Havens et al.
2023 ICLR 2023 (spotlight; arXiv 2303.03169) PauliCertified NNs ★★★★ Shows spectral normalization, orthogonal layers, AOL and CPL are analytic solutions of one SDP condition W^T W <= T (T diagonal) and derives SDP-based Lipschitz Layers (SLL).
Constrained Decision Transformer for Offline Safe Reinforcement Learning
Zuxin Liu, Zijian Guo et al.
2023 ICML 2023 Constrained RL ★★★★ Treats offline safe RL as multi-objective optimization: a decision transformer conditioned on reward-to-go and cost-to-go, with a stochastic entropy-regularized policy and Pareto-frontier-based return relabeling, enabling zero-shot adaptation to different cost thresholds.
Constrained Efficient Global Optimization of Expensive Black-box Functions
Wenjie Xu, Yuning Jiang et al.
2023 ICML 2023 (PMLR 202:38485-38498) Safe exploration & BO ★★★★ CONFIG minimizes the lower confidence bound of the objective subject to lower confidence bounds of the constraints; proves cumulative regret and cumulative constraint-violation bounds, convergence rates, and an infeasibility-declaration scheme.
Lipschitz constant estimation for 1D convolutional neural networks
Patricia Pauli, Dennis Gramlich, Frank Allgöwer
2023 Proceedings of the 5th Annual Learning for Dynamics and C… Pauli ★★★★ Realizes 1-D convolutions as causal FIR state-space systems and certifies incremental (Q,S,R)-dissipativity layer by layer (convolutional, fully connected, average/max pooling via new incremental QCs); layers are coupled through a block-tridiagonal LMI whose size depends on channels and kernel size, not on signal length.
Lipschitz-Bounded 1D Convolutional Neural Networks using the Cayley Transform and the Controllability Gramian
Patricia Pauli, Ruigang Wang et al.
2023 62nd IEEE Conference on Decision and Control (CDC 2023), … Pauli ★★★★ First direct parameterization of end-to-end Lipschitz-bounded 1-D CNNs. The storage matrix of each convolutional layer is the inverse of a controllability Gramian (closed form because the FIR state matrix is nilpotent), and kernels come from a Cayley transform, so the layer LMIs hold for all free parameters. Demonstrated on MIT-BIH arrhythmia data.
On the Sublinear Regret of GP-UCB
Justin Whitehouse, Aaditya Ramdas, Zhiwei Steven Wu
2023 NeurIPS 2023 (main conference track) Safe exploration & BO ★★★★ Shows that vanilla GP-UCB attains nearly optimal, sublinear regret for kernels with polynomial eigendecay (including all Matern kernels) by regularizing the kernel ridge estimator in proportion to kernel smoothness. Relies on a simplified derivation of Abbasi-Yadkori's (2013) Hilbert-space self-normalized bound.
Safe Control With Learned Certificates: A Survey of Neural Lyapunov, Barrier, and Contraction Methods for Robotics and Control
Charles Dawson, Sicun Gao, Chuchu Fan
2023 IEEE Transactions on Robotics 39(3):1749–1767 (DOI 10.110… Control-theoretic safety ★★★★ Survey and tutorial on learning neural Lyapunov, barrier and contraction-metric certificates together with controllers: losses, sampling, verification, robustness to model and state-estimation error, hardware case studies, open-source code (neural_clbf).
Safety-Gymnasium: A Unified Safe Reinforcement Learning Benchmark
Jiaming Ji, Borong Zhang et al.
2023 NeurIPS 2023 (Datasets and Benchmarks Track) Constrained RL ★★★★ Unified environment suite of safety-critical single- and multi-agent tasks with vector and vision-only inputs, plus the SafePO library of 16 SafeRL algorithms for standardized evaluation.
Almost-Orthogonal Layers for Efficient General-Purpose Lipschitz Networks
Bernd Prach, Christoph H. Lampert
2022 ECCV 2022 (LNCS, pp. 350-365, DOI 10.1007/978-3-031-19803… Certified NNs ★★★★ AOL: a rescaling-based parameterization that makes any fully-connected or convolutional layer provably 1-Lipschitz without orthogonalization or matrix inversion; the learned weights end up almost orthogonal.
Constrained Policy Optimization via Bayesian World Models
Yarden As, Ilnura Usmanova et al.
2022 ICLR 2022 Constrained RL ★★★★ LAMBDA: a model-based CMDP agent with a Bayesian world model; policy gradients are back-propagated through imagined rollouts, maximizing optimistic bounds on the task objective and pessimistic bounds on the constraints, solved with an augmented Lagrangian.
Constrained Variational Policy Optimization for Safe Reinforcement Learning
Zuxin Liu, Zhepeng Cen et al.
2022 ICML 2022 Constrained RL ★★★★ Casts safe RL as probabilistic inference solved by EM: a convex E-step gives the optimal feasible nonparametric variational policy in closed form (strong duality under Slater), and a supervised trust-region M-step fits the parametric policy; off-policy and more stable than primal-dual updates.
Dissipativity and Integral Quadratic Constraints: Tailored Computational Robustness Tests for Complex Interconnections
Carsten W. Scherer
2022 IEEE Control Systems (Magazine) 42(3):115-139 (DOI 10.110… Pauli ★★★★ Tutorial that re-derives the IQC stability theorem through dissipativity (storage functions and dynamic multipliers) instead of the original homotopy argument, and shows how to build tailored LMI robustness tests for interconnections. Scherer is a co-author of the 2-D-systems CNN paper; this is the most accessible rigorous route to the IQC/LMI machinery used throughout Pauli's work.
High-Order Control Barrier Functions
Wei Xiao, Calin Belta
2022 IEEE Transactions on Automatic Control 67(7):3655–3662 (c… Control-theoretic safety ★★★★ High-order CBFs handle constraints of any relative degree via a chain of derivatives each augmented by a class-K term; the intersection of the nested sets is made forward invariant by a constraint affine in u, usable in a CLF-HOCBF QP. Also proposes penalizing the class-K functions to recover feasibility under input bounds.
Information-Theoretic Safe Exploration with Gaussian Processes
Alessandro G. Bottero, Carlos E. Luis et al.
2022 NeurIPS 2022 Safe exploration & BO ★★★★ ISE selects safe parameters that maximize a closed-form approximation of the mutual information between a new observation and the safety indicator of any other parameter. Works on continuous domains, needs no Lipschitz constant, and comes with safety and exploration guarantees.
Neural network training under semidefinite constraints
Patricia Pauli, Niklas Funcke et al.
2022 61st IEEE Conference on Decision and Control (CDC 2022), … Pauli ★★★★ Replaces ADMM with a log-det barrier formulation trained by backprop, exploiting the block-tridiagonal structure of the Lipschitz LMI for fast Cholesky factorizations and inverses. Multipliers can be decision variables (bilinear variant), and the method scales to Lipschitz-constrained WGAN critics.
Reachability Constrained Reinforcement Learning
Dongjie Yu, Haitong Ma et al.
2022 ICML 2022, PMLR 162:25636-25655 Constrained RL ★★★★ Brings Hamilton-Jacobi reachability into constrained policy optimization: a learned safety value function (worst-case constraint along the trajectory) characterizes the largest feasible set, and a state-dependent Lagrange multiplier enforces the reachability constraint; proves convergence to a local optimum that preserves the largest feasible set. Direct ancestor of FISOR and of Feasible Dual Policy Iteration (ICLR 2026), both already in the index.
Saute RL: Almost Surely Safe Reinforcement Learning Using State Augmentation
Aivar Sootla, Alexander I. Cowen-Rivers et al.
2022 ICML 2022 Constrained RL ★★★★ Removes the constraint by augmenting the state with the (rescaled) remaining safety budget and reshaping the cost, giving a Sauté MDP that satisfies a Bellman equation and approximates almost-surely constrained safe RL. Any RL algorithm can be 'Sautéed', and policies generalize across budgets.
DeepReach: A Deep Learning Approach to High-Dimensional Reachability
Somil Bansal, Claire J. Tomlin
2021 2021 IEEE International Conference on Robotics and Automa… Control-theoretic safety ★★★★ Neural PDE solver for HJ reachability: a sine-activation network V_theta(x,t) is trained self-supervised on the HJI variational inequality with no ground-truth values, so cost scales with value-function complexity rather than grid size. Shown on 9-D multi-vehicle collision avoidance and a 10-D narrow-passage problem; also yields a safety controller.
Fast and Complete: Enabling Complete Neural Network Verification with Rapid and Massively Parallel Incomplete Verifiers
Kaidi Xu, Huan Zhang et al.
2021 ICLR 2021; arXiv 2011.13824 Certified NNs ★★★★ alpha-CROWN: replaces the LP in branch-and-bound by backward linear relaxation (LiRPA/CROWN) whose relaxation slopes alpha are optimized by gradient ascent, with batched splits on GPU/TPU; an order of magnitude faster than LP-based complete verifiers. The 'alpha' half of alpha,beta-CROWN.
GoSafe: Globally Optimal Safe Robot Learning
Dominik Baumann, Alonso Marco et al.
2021 IEEE International Conference on Robotics and Automation … Trimpe group ★★★★ Extends SafeOpt's safe set to the joint space of policy parameters and (discretized) initial conditions, learning from where backup policies can recover the system. This lets it safely evaluate parameters outside the safe region reachable from the initial seed (stages S1-S3), with conditions for convergence to the global optimum; validated on a Furuta pendulum.
Learning-enhanced robust controller synthesis with rigorous statistical and control-theoretic guarantees
Christian Fiedler, Carsten W. Scherer, Sebastian Trimpe
2021 60th IEEE Conference on Decision and Control (CDC 2021), … Trimpe group ★★★★ Pulls unknown static nonlinearities into the Δ-block of a linear fractional representation and learns them with GP regression plus rigorous RKHS error bounds. The uncertainty tube is converted into sector bounds and IQC multipliers for modern robust synthesis, so guarantees hold at every stage while performance improves with data.
Offset-free setpoint tracking using neural network controllers
Patricia Pauli, Johannes Köhler et al.
2021 Proceedings of the 3rd Conference on Learning for Dynamic… PauliControl-theoretic safety ★★★★ Adds integral action to an NN controller for an LTI plant and gives LMI conditions for global and local exponential stability of the tracking error, for one reference or for all references in a set, with ellipsoidal ROA estimates; a reference governor enlarges the certified region.
On Information Gain and Regret Bounds in Gaussian Process Bandits
Sattar Vakili, Kia Khezeli, Victor Picheny
2021 AISTATS 2021 (PMLR 130:82-90) Safe exploration & BO ★★★★ Bounds the maximum information gain via a finite-dimensional RKHS projection plus tail control, for general polynomial or exponential eigendecay; the Matern bound is tight up to log factors relative to known lower bounds.
Orthogonalizing Convolutional Layers with the Cayley Transform
Asher Trockman, J. Zico Kolter
2021 ICLR 2021 (arXiv 2104.07167) PauliCertified NNs ★★★★ Parameterizes orthogonal (1-Lipschitz) convolutions by applying the Cayley transform to skew-symmetric convolutions in the Fourier domain.
Robot Learning with Crash Constraints
Alonso Marco, Dominik Baumann et al.
2021 IEEE Robotics and Automation Letters 6(2):1439-1446 Trimpe group ★★★★ Targets settings where failures are tolerable but return no data. It proposes GPCR, a constraint model that mixes binary crash labels with continuous observations and learns the unknown constraint threshold (by MAP with the hyperparameters), used inside expected improvement with crash constraints (EIC²); demonstrated on a jumping quadruped.
Safety and Liveness Guarantees through Reach-Avoid Reinforcement Learning
Kai-Chieh Hsu, Vicenç Rubies-Royo et al.
2021 Robotics: Science and Systems (RSS) XVII, 2021 Control-theoretic safety ★★★★ Extends discounted safety to reach-avoid problems: the time-discounted reach-avoid Bellman operator is a contraction, reach-avoid Q-learning converges, and discounted reach-avoid sets under-approximate the true set and converge to it as gamma->1. Deep RL solutions serve as untrusted oracles inside a supervisory (MPC-checked) filter that keeps zero-violation guarantees.
Data-driven Economic NMPC using Reinforcement Learning
Sébastien Gros, Mario Zanon
2020 IEEE Transactions on Automatic Control 65(2):636–648 Control-theoretic safety ★★★★ Shows that an (economic) NMPC scheme built on a wrong model can still deliver the true optimal policy and value functions if stage and terminal costs are modified suitably, so a parametrized MPC can serve as the function approximator in RL; tuned with Q-learning and deterministic policy gradients and connected to dissipativity.
Event-triggered Learning
Friedrich Solowjow, Sebastian Trimpe
2020 Automatica 117:109009 Trimpe group ★★★★ Introduces event-triggered learning. In event-triggered state estimation, empirical inter-communication times are compared with those predicted by the model, and re-learning is triggered only when the mismatch is statistically significant. Triggers come from Hoeffding (expectation-based) and DKW (CDF-based) bounds for linear Gaussian dynamics.
First Order Constrained Optimization in Policy Space
Yiming Zhang, Quan Vuong, Keith W. Ross
2020 NeurIPS 2020 Constrained RL ★★★★ Solves the CPO-style trust-region problem exactly in nonparametric policy space (closed-form Gibbs policy), then projects back to the parametric class by minimizing KL with first-order methods; inherits a CPO-type worst-case violation bound.
Learning Control Barrier Functions from Expert Demonstrations
Alexander Robey, Haimin Hu et al.
2020 2020 59th IEEE Conference on Decision and Control (CDC), … Control-theoretic safety ★★★★ Learns a CBF from safe expert trajectories via margin constraints on sampled safe and unsafe points and on the CBF derivative; Lipschitz arguments on epsilon-nets prove validity of the learned local CBF, and the problem is convex for models linear in parameters. Claimed as the first provably safe CBFs learned from data.
Learning for Safety-Critical Control with Control Barrier Functions
Andrew Taylor, Andrew Singletary et al.
2020 Proceedings of the 2nd Conference on Learning for Dynamic… Control-theoretic safety ★★★★ Learns the model-uncertainty part of the CBF time derivative from data in an episodic data-aggregation loop and uses the learned derivative inside the CBF-QP, keeping the filter safe despite model error. Validated in simulation and on a Segway.
Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision Processes
Dongsheng Ding, Kaiqing Zhang et al.
2020 NeurIPS 2020 (journal version with Jiali Duan: 'Convergen… Constrained RL ★★★★ Primal natural-policy-gradient ascent (a multiplicative-weights update under softmax) plus projected dual subgradient descent for CMDPs with a utility constraint. First non-asymptotic, dimension-free global convergence for a policy-based primal-dual method, with extensions to general parametrizations and sample-based variants.
Projection-Based Constrained Policy Optimization
Tsung-Yen Yang, Justinian Rosca et al.
2020 ICLR 2020 Constrained RL ★★★★ Two-step update: an unconstrained TRPO reward-improvement step followed by projection onto the linearized constraint set (L2 or KL). Gives worst-case reward-degradation and constraint-violation bounds for feasible and infeasible iterates plus convergence analysis for both projections.
Reach-SDP: Reachability Analysis of Closed-Loop Systems with Neural Network Controllers via Semidefinite Programming
Haimin Hu, Mahyar Fazlyab et al.
2020 2020 59th IEEE Conference on Decision and Control (CDC), … Certified NNs ★★★★ Over-approximates finite-horizon forward reachable sets of linear time-varying plants with (projected) ReLU NN controllers by propagating polytopes or ellipsoids with one SDP per step built from QCs of the input set, the NN and the output set; certifies constraint satisfaction of an NN-approximated quadrotor MPC.
A Learnable Safety Measure
Steve Heim, Alexander von Rohr et al.
2019 Conference on Robot Learning (CoRL 2019), PMLR 100:627-63… Trimpe group ★★★★ Defines a safety measure as the volume of viable actions at each state, which implicitly encodes both the dynamics and the failure set. The viable set is recovered as its strictly positive level set (Λ_Q > 0), and the measure is learned model-free by active sampling with a GP; using the estimated measure already reduces failures during learning.
Adaptive and Safe Bayesian Optimization in High Dimensions via One-Dimensional Subspaces
Johannes Kirschner, Mojmir Mutny et al.
2019 ICML 2019 (PMLR 97:3429-3438) Safe exploration & BO ★★★★ LineBO solves a sequence of one-dimensional BO sub-problems along directions from an oracle (random, coordinate or descent), with global convergence, fast local rates for strongly convex objectives and adaptation to the effective dimension. Using SafeOpt on each line yields the first safe BO with guarantees applicable in high dimensions; deployed on the SwissFEL with up to 40 parameters.
Input-to-State Safety With Control Barrier Functions
Shishir Kolathaya, Aaron D. Ames
2019 IEEE Control Systems Letters 3(1):108–113 Control-theoretic safety ★★★★ Input-to-state safety: under bounded input disturbances an ISSf-CBF guarantees forward invariance of a slightly larger set whose inflation is a class-K function of the disturbance bound. Gives a way to build ISSf-CBFs from CBFs, a universal (Sontag-like) control law, and CLF-ISSf-CBF QPs.
Linear Stochastic Bandits Under Safety Constraints
Sanae Amani, Mahnoosh Alizadeh, Christos Thrampoulidis
2019 NeurIPS 2019 Safe exploration & BO ★★★★ Safe-LUCB: linear bandits where every action must satisfy a linear constraint that depends on the unknown parameter; a random safe pure-exploration phase is followed by LUCB that is optimistic in cost and pessimistic in the constraint. Gives a general regret bound and a problem-dependent bound governed by the safety gap at the optimum.
Neural Lyapunov Control
Ya-Chien Chang, Nima Roohi, Sicun Gao
2019 Advances in Neural Information Processing Systems 32 (Neu… Control-theoretic safety ★★★★ Learns a neural Lyapunov function jointly with a controller by minimizing an empirical 'Lyapunov risk'; an SMT falsifier (dReal) returns counterexamples until none remain, proving stability on the verified domain. Certified regions of attraction are larger than LQR and SOS/SDP on benchmarks.
No-Regret Bayesian Optimization with Unknown Hyperparameters
Felix Berkenkamp, Angela P. Schoellig, Andreas Krause
2019 Journal of Machine Learning Research 20(50):1-24 Safe exploration & BO ★★★★ First GP-UCB-type algorithm (A-GP-UCB) with sublinear regret when kernel hyperparameters and the RKHS norm bound are unknown, by slowly expanding the function class (shrinking lengthscales, growing norm bound). It is the constructive counterpart to Fiedler et al.'s critique of hyperparameter-dependent safety in safe BO and the basis for PACSBO-type adaptivity.
Safe Exploration for Interactive Machine Learning
Matteo Turchetta, Felix Berkenkamp, Andreas Krause
2019 NeurIPS 2019 Safe exploration & BO ★★★★ GoOSE (Goal-Oriented Safe Exploration) wraps any unsafe interactive-ML oracle: the oracle proposes decisions within an optimistic safe set and GoOSE learns the safety of only those decisions via pessimistic/optimistic expansion operators with ergodicity. Keeps SafeMDP-style safety guarantees while avoiding exhaustive safe-set exploration.
Uniform Error Bounds for Gaussian Process Regression with Application to Safe Control
Armin Lederer, Jonas Umlauft, Sandra Hirche
2019 NeurIPS 2019 Safe exploration & BO ★★★★ Derives a uniform (over a compact continuous set) GP error bound in the Bayesian setting from a covering/grid argument plus Lipschitz continuity of kernel, posterior mean and posterior std, and derives probabilistic Lipschitz constants for GP sample paths. Uses it for safe feedback-linearizing control of a manipulator.
A Lyapunov-based Approach to Safe Reinforcement Learning
Yinlam Chow, Ofir Nachum et al.
2018 Advances in Neural Information Processing Systems 31 (Neu… Control-theoretic safety ★★★★ Constructs Lyapunov functions for CMDPs (via an LP over an auxiliary constraint cost) whose local linear constraints keep every policy update feasible; derives safe policy iteration and safe value iteration plus scalable safe DQN and safe DPI.
Certified Defenses against Adversarial Examples
Aditi Raghunathan, Jacob Steinhardt, Percy Liang
2018 ICLR 2018; arXiv 1801.09344 Certified NNs ★★★★ For one-hidden-layer networks, upper-bounds the worst-case l_inf adversarial margin through a gradient-norm bound relaxed to an SDP (MAXCUT-style) and trains jointly against the differentiable dual certificate (SDP-NN): no attack at eps=0.1 on MNIST can exceed 35% error.
Gaussian Processes and Kernel Methods: A Review on Connections and Equivalences
Motonobu Kanagawa, Philipp Hennig et al.
2018 arXiv preprint 1807.02582 (July 2018, 64 pp.) Safe exploration & BO ★★★★ The standard rigorous account of the GP-regression / kernel-ridge-regression equivalence, of the posterior variance as a worst-case RKHS error, and of the fact that GP sample paths almost surely lie outside the RKHS of the kernel. Exactly the material needed to explain why frequentist (RKHS) bounds (Fiedler, Chowdhury-Gopalan) and Bayesian sample-path bounds (Lederer) rest on different assumptions.
Learning an Approximate Model Predictive Controller with Guarantees
Michael Hertneck, Johannes Köhler et al.
2018 IEEE Control Systems Letters 2(3):543-548 Trimpe groupPauli ★★★★ Designs a robust MPC that tolerates bounded input errors and approximates it by supervised learning (for example an NN). The learned controller is validated along closed-loop trajectories with Hoeffding's inequality, giving a statistical guarantee of stability and constraint satisfaction.
Learning-Based Model Predictive Control for Safe Exploration
Torsten Koller, Felix Berkenkamp et al.
2018 2018 IEEE Conference on Decision and Control (CDC), pp. 6… Control-theoretic safety ★★★★ Safe exploration with learning-based MPC: GP confidence intervals are propagated as ellipsoids through multi-step predictions under affine feedback, all constraints are checked on these ellipsoids, and a terminal constraint into a safe set with a backup controller keeps the scheme recursively safe with high probability.
Lipschitz regularity of deep neural networks: analysis and efficient estimation
Aladin Virmaux, Kevin Scaman
2018 NeurIPS 2018 (Advances in Neural Information Processing S… Certified NNs ★★★★ Proves exact Lipschitz computation is NP-hard even for 2-layer ReLU MLPs (Theorem 2), introduces AutoLip (autodiff + power-method upper bound for any differentiable program) and SeqLip, which exploits misalignment of consecutive singular vectors to tighten the product-of-norms bound.
Lipschitz-Margin Training: Scalable Certification of Perturbation Invariance for Deep Neural Networks
Yusuke Tsuzuku, Issei Sato, Masashi Sugiyama
2018 NeurIPS 2018 (Advances in Neural Information Processing S… Certified NNs ★★★★ Links Lipschitz constants and prediction margins into a cheap certificate for large networks and proposes LMT, which inflates non-target logits by the Lipschitz-scaled margin during training to enlarge certified l2 regions.
Provable Defenses against Adversarial Examples via the Convex Outer Adversarial Polytope
Eric Wong, J. Zico Kolter
2018 ICML 2018 (PMLR 80, pp. 5286-5295); arXiv 1711.00851 Certified NNs ★★★★ Trains provably robust ReLU networks by bounding the worst-case loss over a convex outer approximation (LP relaxation, triangle relaxation of ReLU) of the set of reachable activations under l_inf perturbations; the LP dual is itself a backward-pass network, so the bound is differentiable and cheap. MNIST: provable test error < 5.8% at eps = 0.1.
Safe Exploration in Continuous Action Spaces
Gal Dalal, Krishnamurthy Dvijotham et al.
2018 arXiv preprint 1801.08757 (26 Jan 2018), DeepMind Constrained RL ★★★★ The 'safety layer': a learned linearized constraint model per cost signal and a closed-form projection of the policy action onto the constraint set, giving zero-violation (state-wise) safe exploration in continuous control. It is the most cited action-projection baseline and the template for later projection-based safety filters in deep RL.
Stagewise Safe Bayesian Optimization with Gaussian Processes
Yanan Sui, Vincent Zhuang et al.
2018 ICML 2018 (PMLR 80:4781-4789) Safe exploration & BO ★★★★ StageOpt separates safe-region expansion (driven only by the safety GPs) from utility maximization (GP-UCB, or dueling-bandit preference feedback) inside the certified region, with finite-time guarantees for both stages. Applied clinically to spinal cord stimulation.
Virtual vs. Real: Trading Off Simulations and Physical Experiments in Reinforcement Learning with Bayesian Optimization
Alonso Marco, Felix Berkenkamp et al.
2017 IEEE International Conference on Robotics and Automation … Trimpe group ★★★★ Multi-fidelity Entropy Search (MF-ES) decides at each iteration whether to evaluate a policy in simulation or on the robot, maximizing information about the real-system optimum per unit of evaluation effort. The simulator error is modeled explicitly in the kernel, and cart-pole experiments needed fewer physical trials than BO on the physical system alone.
Exponential Control Barrier Functions for enforcing high relative-degree safety-critical constraints
Quan Nguyen, Koushil Sreenath
2016 2016 American Control Conference (ACC), pp. 322–328 Control-theoretic safety ★★★★ Exponential CBFs enforce state constraints of relative degree r through a condition on the r-th derivative that is linear in u, with gains chosen by pole placement. Demonstrated on dynamic walking over stepping stones and other robots.
Safe Controller Optimization for Quadrotors with Gaussian Processes
Felix Berkenkamp, Angela P. Schoellig, Andreas Krause
2016 ICRA 2016, pp. 491-496 Safe exploration & BO ★★★★ First application of SafeOpt to automatic controller tuning on hardware (quadrotor), guaranteeing that only parameters above a safe performance threshold are evaluated. Introduces a Lipschitz-free SafeOpt variant in which safety and expanders are defined directly from GP confidence bounds.
Robustness of Control Barrier Functions for Safety Critical Control
Xiangru Xu, Paulo Tabuada et al.
2015 IFAC-PapersOnLine 48(27):54-61 (ADHS 2015), DOI 10.1016/j… Control-theoretic safety ★★★★ Introduces the now-standard CBF definition with an extended class-K function and proves that CBF-based controllers make the safe set both forward invariant and asymptotically stable (hence robust to perturbations and model error). It is the mathematical core behind the 2017 TAC paper and behind ISSf.
Safe Exploration for Optimization with Gaussian Processes (SafeOpt)
Yanan Sui, Alkis Gotovos et al.
2015 ICML 2015, PMLR 37:997-1005 Trimpe group ★★★★ External foundation, not a Trimpe-group paper. GP confidence bounds plus a Lipschitz constant define a growing safe set, and queries are restricted to potential maximizers and expanders. This guarantees safety with high probability and ε-optimality within the safely reachable region.
Trust Region Policy Optimization
John Schulman, Sergey Levine et al.
2015 ICML 2015, PMLR 37:1889-1897 Constrained RL ★★★★ Establishes the surrogate-objective/KL-penalty bound and the trust-region (natural-gradient) update that CPO, PCPO, FOCOPS, CUP and SB-TRPO extend with constraints. CPO's proof is literally the TRPO bound applied to both reward and cost.
Control barrier function based quadratic programs with application to adaptive cruise control
Aaron D. Ames, Jessy W. Grizzle, Paulo Tabuada
2014 53rd IEEE Conference on Decision and Control (CDC 2014), … Control-theoretic safety ★★★★ The first CBF-QP paper: introduces control barrier functions (reciprocal form, B->infinity at the boundary) and unifies them with control Lyapunov functions in a quadratic program, with the CBF as hard constraint and the CLF relaxed, applied to adaptive cruise control with force bounds.
Safe Exploration in Markov Decision Processes
Teodor Mihai Moldovan, Pieter Abbeel
2012 ICML 2012 (Proceedings of the 29th International Conferen… Safe exploration & BO ★★★★ First formalization of safe exploration as preserving ergodicity (the ability to return to the start state) with high probability; shows the exact problem is hard and gives a constrained-optimization algorithm with guaranteed-safe but possibly suboptimal exploration. Turchetta et al. (2016) and the 'reachability and returnability operators' concept descend directly from it.
Viability Theory
Jean-Pierre Aubin
2009 Birkhauser Boston, Modern Birkhauser Classics reprint (DO… Trimpe group ★★★★ Defines the viability kernel of a constraint set under a differential inclusion and proves the viability theorem (tangential condition) and kernel-computation results. Heim's safety measure, Massiani's viable set / critical set and safe value functions, and the 'viability of future actions' paper all use these definitions.
The Scenario Approach to Robust Control Design
Giuseppe C. Calafiore, Marco C. Campi
2006 IEEE Transactions on Automatic Control, vol. 51, no. 5, p… Certified NNs ★★★★ Introduces the scenario approach to robust control: replace a semi-infinite robust constraint by N randomly sampled scenarios, solve the resulting convex program, and certify, independently of the uncertainty distribution, that the design violates at most an epsilon-fraction of the uncertainty with confidence 1-beta.
Robust model predictive control of constrained linear systems with bounded disturbances
David Q. Mayne, Maria M. Seron, Sasa V. Rakovic
2005 Automatica 41(2):219-224 (DOI 10.1016/j.automatica.2004.0… Trimpe group ★★★★ The tube-MPC paper: nominal MPC plus an ancillary feedback keeps the true state in a robust positively invariant tube around the nominal trajectory, with constraints tightened by the tube cross-section. Hertneck's certified approximate MPC, Nubert's robust MPC + NN, Hose's safety augmentation, Wabersich's predictive safety filter and Zanon's robust-MPC safe RL all rely on this construction.
Set invariance in control
Franco Blanchini
1999 Automatica 35(11):1747-1767 (DOI 10.1016/S0005-1098(99)00… Control-theoretic safety ★★★★ The survey that fixes the definitions and main theorems of (robust) positive invariance, Nagumo's theorem, ellipsoidal and polyhedral invariant sets, and their LMI/LP computation. Provides the invariance vocabulary that CBFs, tube MPC terminal sets and predictive safety filters assume.
System analysis via integral quadratic constraints
Alexandre Megretski, Anders Rantzer
1997 IEEE Transactions on Automatic Control 42(6):819-830 (Jun… Pauli ★★★★ Unifies small-gain, passivity, circle/Popov and multiplier criteria as IQCs and proves the IQC stability theorem by a homotopy argument.
On the Kalman-Yakubovich-Popov lemma
Anders Rantzer
1996 Systems & Control Letters 28(1):7-10 (DOI 10.1016/0167-69… Pauli ★★★★ Gives the modern, short proof of the KYP lemma (frequency-domain inequality equals an LMI) via convexity/separation, in the general form used by IQC theory. It is the bridge between Megretski-Rantzer frequency-domain IQCs and the finite-dimensional LMIs solved in Pauli's Zames-Falb analysis, Yin's ROA analysis and Scherer's dissipativity tests.
Demystifying Lipschitz verification: positive matrices, negative results
Simon Kuang, Yuezhu Xu et al.
2026 arXiv preprint 2603.28113 (Mar 2026; v2 May 2026, 'reduce… Pauli ★★★ Argues Lipschitz verification is structurally hard because it needs hidden-state reachability (NP-hard); constructs instances where SDP/QC bounds inherit the qualitative failures of the product bound; proposes regularizing the trivial bound with bias-free trigonometric layers.
From Demonstrations to Safe Deployment: Path-Consistent Safety Filtering for Diffusion Policies
Ralf Römer, Julian Balletshofer et al.
2026 IEEE International Conference on Robotics and Automation … Control-theoretic safety ★★★ PACS filters diffusion-policy action chunks: it builds a path from the generated action sequence and applies path-consistent braking verified with set-based reachability analysis, so the executed motion stays close to the training distribution while giving formal safety guarantees in dynamic environments. Reports up to 68% higher task success than reactive CBF-based filtering.
Preferential Bayesian Optimization with Crash Feedback
Johanna Menn, David Stenger, Sebastian Trimpe
2026 IEEE Robotics and Automation Letters 11(4):4299-4306 (DOI… Trimpe group ★★★ CrashPBO lets users give pairwise preferences and also report crashes in preferential BO. Crashes become virtual comparisons that rank crashed parameters below all non-crashed ones, a hyperparameter-free mechanism evaluated on three robot platforms.
Provably Safe Generative Sampling with Constricting Barrier Functions
Darshan Gadginmath, Ahmed Allibhoy, Fabio Pasqualetti
2026 arXiv:2602.21429 (v1 24 Feb 2026, v3 31 Jul 2026, 26 page… Control-theoretic safety ★★★ Applies control-barrier-function reasoning to the reverse-time sampling dynamics of pretrained flow/diffusion models without retraining. A 'constricting' time-varying barrier defines a safety tube that is loose at the initial noise level and shrinks to the target safe set C at the final sampling step; a minimum-norm guidance control computed by a (linearized) QP enforces a discrete-time barrier inequality. Proves discrete-time reverse invariance (final sample in C) under the exact inequality, bounds the KL divergence between the guided and original sample distributions by the accumulated control energy weighted by 1/g(t_k)^2, and reports 100% constraint satisfaction experimentally.
R2DN: Scalable Parameterization of Contracting and Lipschitz Recurrent Deep Networks
Nicholas H. Barbara, Ruigang Wang, Ian R. Manchester
2026 IEEE CDC 2026 (accepted per arXiv v3, Sept 2026; arXiv 25… Pauli ★★★ Recurrent models built as an LTI system in feedback with a 1-Lipschitz deep feedforward network, directly parameterized to be contracting and Lipschitz without REN's equilibrium solve; up to an order of magnitude faster training/inference.
Rethinking Evaluation Paradigms in IBP-based Certified Training
Konstantin Kaulen, Hadar Shavit, Holger H. Hoos
2026 ICML 2026 (per the arXiv comments field 'ICML 2026'; the … Certified NNs ★★★ Argues that certified-training methods should be compared on the whole natural-vs-certified accuracy Pareto front rather than a single tuned configuration. Uses multi-objective Bayesian optimisation (Gaussian-process surrogate, EHVI acquisition, BoTorch via Optuna, 3 seeds) on the CNN7 architecture of Shi et al. 2021 for IBP, CROWN-IBP, SABR and MTL-IBP (CC-IBP and Exp-IBP in appendices) on CIFAR-10 (eps = 2/255, 8/255), Tiny-ImageNet (1/255) and MNIST. Certified accuracy on the final fronts is measured by complete verification with alpha-beta-CROWN (1000 s cutoff); during tuning an under-approximation via IBP/CROWN-IBP/CROWN is optimised. Findings: IBP and CROWN-IBP perform substantially better than previously reported when properly tuned; SABR gives the highest clean accuracies with strong certified robustness; MTL-IBP is best where high certified accuracy is desired at small radii; no method dominates, so combined fronts mix all four.
Scalable Gaussian Process Regression via Deterministic Trigonometric Features: Uniform Bounds for Safe Model Predictive Control
Julius Jagdt, Johanna Menn et al.
2026 arXiv preprint 2608.16415 (17 Aug 2026) Trimpe group ★★★ DTF-GP uses deterministic trigonometric features on a fixed frequency grid (from Bochner's theorem), reducing GP regression to Bayesian linear regression. It admits a high-probability uniform error bound (closed form for SE kernels; projection error decays exponentially for RBF when the period exceeds the domain), enabling scalable learning-based MPC with safety guarantees.
Scalable Incremental Robustness Analysis of Neural Network Feedback Systems
Zichen Wang, Peter Seiler et al.
2026 arXiv preprint 2609.22576 (18 Sept 2026) Pauli ★★★ Decomposes the full-order SDP for feedback loops with deep NNs and unmodeled dynamics and combines it with scalable Lipschitz estimation; reduced LMIs certify incremental convergence and incremental l2-gain with size depending only on the last two NN layers.
Adaptable Safe Policy Learning from Multi-task Data with Constraint Prioritized Decision Transformer
Ruiqi Xue, Ziqian Zhang et al.
2025 NeurIPS 2025 (Advances in Neural Information Processing S… Constrained RL ★★★ Extends the Constrained Decision Transformer (CDT) line to multi-task, multi-constraint, multi-budget offline safe RL. Two components: (1) a constraint-prioritized prompt encoder that uses the sparse binary cost signal of a reference trajectory to route segments to safe/unsafe sub-encoders and produce a constraint embedding z; (2) a constraint-prioritized return-to-go (CPRTG) generator q_phi(R|C_hat, s) that replaces the hand-specified target return with a return sampled conditional on the current cost-to-go and state, selected at a CTG-dependent quantile beta_t, so the RTG never conflicts with the safety budget. One unified DT policy is then trained on all tasks.
Automatic nonlinear MPC approximation with closed-loop guarantees
Abdullah Tokmak, Christian Fiedler et al.
2025 IEEE Transactions on Automatic Control 70(10):6388-6403 (… Trimpe group ★★★ ALKIA-X (Adaptive and Localized Kernel Interpolation Algorithm with eXtrapolated RKHS norm) non-iteratively computes an explicit kernel-interpolation approximation of a nonlinear MPC law on adaptively partitioned sub-domains (cubes with samples at the vertices), guaranteed to meet a user-chosen uniform error bound. A robust MPC design then yields closed-loop guarantees.
CTBench: A Library and Benchmark for Certified Training
Yuhao Mao, Stefan Balauca, Martin Vechev
2025 arXiv 2406.04848 (v4 May 2025); reported as ICML 2025 in … Certified NNs ★★★ Unified library and fair re-evaluation of l_inf certified-training methods (IBP, CROWN-IBP, SABR, TAPS, STAPS, MTL-IBP): tuned baselines beat their published numbers and most recent claimed gains shrink; certified models have less fragmented loss surfaces and better OOD generalization potential.
Discrete-Time Stability Analysis of ReLU Feedback Systems via Integral Quadratic Constraints
Sahel Vahedi Noori, Bin Hu et al.
2025 arXiv preprint 2511.12826 (16 Nov 2025) Pauli ★★★ Hard IQCs specific to scalar ReLU, built from FIR filters and structured matrices, for discrete-time ReLU feedback loops (e.g., RNNs), combined with a dissipation inequality into an LMI for internal stability.
Embedding Safety into RL: A New Take on Trust Region Methods
Nikola Milosevic, Johannes Müller, Nico Scherf
2025 ICML 2025 (PMLR 267) Constrained RL ★★★ Reshapes policy geometry with a Bregman divergence built from a barrier on the constraint margin so that trust regions contain only safe policies, giving C-TRPO (surrogate divergence plus hysteresis-based recovery) and C-NPG (safe set invariant under the flow).
Enhancing Certified Robustness via Block Reflector Orthogonal Layers and Logit Annealing Loss (BRONet)
Bo-Han Lai, Pin-Han Huang et al.
2025 ICML 2025 (spotlight; arXiv 2505.15174) PauliCertified NNs ★★★ Block Reflector Orthogonal (BRO) layers, a low-rank Householder-type orthogonal parameterization for 1-Lipschitz networks, plus a logit annealing loss; the previous state of the art in deterministic l2 certified robustness before LipNeXt.
Event-Triggered Time-Varying Bayesian Optimization
Paul Brunzema, Alexander von Rohr et al.
2025 Transactions on Machine Learning Research (TMLR), publish… Trimpe group ★★★ ET-GP-UCB treats the objective as static until an event trigger based on a GP uniform error bound detects a change, then resets the dataset, so the rate of change ε need not be known. It derives regret bounds for any reset strategy that resets between a lower and an upper reset time.
Formal Verification and Control with Conformal Prediction
Lars Lindemann, Yiqi Zhao et al.
2025 arXiv preprint 2409.00536 (v1 Aug 2024, v3 Aug 2025) Certified NNs ★★★ Survey of conformal prediction for verification and control of learning-enabled autonomous systems (navigation to temporal-logic specifications), contrasting CP with scenario optimization and PAC-Bayes and listing open directions (feedback, distribution shift).
Gaussian Process Upper Confidence Bound Achieves Nearly-Optimal Regret in Noise-Free Gaussian Process Bandits
Shogo Iwazaki
2025 NeurIPS 2025 (39th Conference on Neural Information Proce… Safe exploration & BO ★★★ Refined regret analysis of (noise-free) GP-UCB for a fixed function in a bounded-norm RKHS with exact observations. The key technical result (Lemma 3) bounds the minimum and the cumulative posterior standard deviation along any input sequence for SE and Matérn kernels, via an elliptical-potential count lemma (from Flynn & Reeb) and the Vakili et al. MIG bounds; combined with the deterministic noise-free confidence bound |f(x)-mu(x)| <= B sigma(x), this yields the first constant cumulative regret for SE kernels and for Matérn kernels with nu > d, matching Li-Scarlett lower bounds up to log factors, and resolves an open problem on adaptive noise-free algorithms.
Generalizing Safety Beyond Collision-Avoidance via Latent-Space Reachability Analysis
Kensuke Nakamura, Lasse Peters, Andrea Bajcsy
2025 Proceedings of Robotics: Science and Systems (RSS) XXI, 2025 Control-theoretic safety ★★★ Latent Safety Filters run HJ reachability in the latent space of a generative world model trained on RGB observations; failure sets come from a classifier over latent states, so constraints like 'do not spill the bag' need no hand-written signed distance. Safeguards teleoperation and learned policies on a Franka manipulator.
Kernel conditional tests from learning-theoretic bounds
Pierre-François Massiani, Christian Fiedler et al.
2025 NeurIPS 2025 (Advances in Neural Information Processing S… Trimpe group ★★★ Turns confidence bounds of learning methods into conditional hypothesis tests that locate the inputs where conditional expectations (or, via kernel mean embeddings, whole conditional distributions) differ. It proves time-uniform KRR confidence bounds for infinite-dimensional outputs with non-trace-class kernels and non-i.i.d. data, plus bootstrapped thresholds.
Online Optimization for Offline Safe Reinforcement Learning
Yassine Chemingui, Aryan Deshwal et al.
2025 NeurIPS 2025 (Advances in Neural Information Processing S… Constrained RL ★★★ Introduces O3SRL: offline safe RL is posed as a minimax problem over distributions D over policies and a Lagrange multiplier lambda, and solved by alternating (i) an offline RL oracle run on a lambda-shaped reward and (ii) a no-regret online update of lambda. Theorem 1: the averaged iterates are an eps-approximate minimax equilibrium with eps = eps_offlineRL(n) + R_T(Lambda)/T. The practical variant discretizes lambda into K arms and runs EXP3 (a multi-armed bandit), which removes the need for off-policy evaluation at every round; Theorem 2 gives eps = O(eps_offlineRL(n) + sqrt(K/T) + 1/K). The final practical algorithm warm-starts the offline RL learner across rounds and returns a single policy instead of a mixture.
Robust Neural Networks: Analysis, Synthesis, and Control
Patricia Pauli
2025 PhD dissertation, University of Stuttgart; published by L… Pauli ★★★ Monograph unifying SDP-based Lipschitz estimation (incl. structure-exploiting CNN analysis), training with prescribed Lipschitz bounds via semidefinite constraints and direct parameterizations, and closed-loop stability certification with dynamic IQCs.
Robust Representation Consistency Model via Contrastive Denoising
Jiachen Lei, Julius Berner et al.
2025 ICLR 2025 (proceedings PDF header: 'Published as a confer… Certified NNs ★★★ Replaces explicit diffusion-denoising inside randomized smoothing by a single classifier (rRCM) pre-trained so that representations are consistent along diffusion (PF-ODE) trajectories: temporally adjacent noisy points x_{t_n}, x_{t_{n-1}} sharing the same Gaussian noise form positive pairs in an InfoNCE-style objective, then the encoder plus a linear head is fine-tuned per noise level sigma with a consistency-regularized cross-entropy. The certificate itself is unchanged Cohen-style Gaussian smoothing; the gain is accuracy/cost. On ImageNet (500 test images, 100,000 noise samples, 99.9% confidence) rRCM-B-Deep reports certified accuracy 77.4/64.0/51.2/40.0/32.6/25.0 % at l2 radii 0/0.5/1.0/1.5/2.0/2.5, vs. DDS (Carlini et al. 2022) 76.2/61.0/41.4/28.0/21.2/17.2, with ~85x lower inference cost than diffusion-based smoothing on average.
Safe exploration in reproducing kernel Hilbert spaces
Abdullah Tokmak, Kiran G. Krishnan et al.
2025 AISTATS 2025, PMLR 258:784-792 Trimpe groupSafe exploration & BO ★★★ A follow-up from alumnus Baumann's Aalto group, not a Trimpe paper. It over-estimates the unknown RKHS norm from data by comparing with random RKHS functions through a sampling-and-discarding scenario approach (Campi & Garatti), plugs the estimate into Abbasi-Yadkori-type confidence intervals, and proves safety of the resulting SafeOpt variant.
Safe Guaranteed Exploration for Non-linear Systems
Manish Prajapat, Johannes Koehler et al.
2025 IEEE Transactions on Automatic Control (2025), DOI 10.110… Safe exploration & BO ★★★ A safe guaranteed-exploration framework based on optimal control for nonlinear systems with an a priori unknown constraint set, with finite-time sample complexity and safety w.h.p.; SageMPC adds Lipschitz bounds, goal-directed exploration and receding-horizon replanning while keeping the guarantees.
SafeDiffuser: Safe Planning with Diffusion Probabilistic Models
Wei Xiao, Tsun-Hsuan Wang et al.
2025 ICLR 2025 (arXiv v1 2023 lists four authors: Xiao, Wang, … Control-theoretic safety ★★★ Adds formal safety to diffusion planners: the denoising process is treated as a controlled dynamical system and CBF-type constraints are enforced at every denoising step through QPs, in robust-safe, relaxed-safe and time-varying-safe variants that avoid local traps, guaranteeing finite-time diffusion invariance. Evaluated on maze planning, legged locomotion and manipulation.
Scalable Synthesis of Formally Verified Neural Value Function for Hamilton-Jacobi Reachability Analysis
Yujie Yang, Hanjiang Hu et al.
2025 Journal of Artificial Intelligence Research (JAIR) 83, Ar… Control-theoretic safety ★★★ Synthesizes neural HJ-reachability value functions for a FIXED closed-loop policy in discrete time and formally verifies (via MILP over ReLU networks) that the zero-sublevel set is a constraint-satisfying forward-invariant region. Training combines pre-training, adversarial training and verification-guided (counterexample) training, plus three scalability techniques: boundary-guided backtracking (BGB), entering-state regularization (ESR) and activation-pattern alignment (APA). Introduces the Cersyve-9 benchmark (nine tasks, state dimension 2-6) and reports verified certificates on all nine.
Simulation-Aided Policy Tuning for Black-Box Robot Learning
Shiming He, Alexander von Rohr et al.
2025 IEEE Transactions on Robotics 41:2533-2548 (DOI 10.1109/T… Trimpe group ★★★ Proposes HCI-GIBO (High-Confidence Improvement Gradient-Information BO) and its simulator-aided variant S-HCI-GIBO: a local, zero-order policy-search method that keeps querying a derivative-GP model until it can certify, with probability at least alpha, that a gradient step improves the policy, and that first exhausts gradient information from an imperfect simulator (modeled as robot objective plus a GP reality gap) before querying the robot. Guarantees are on performance improvement per update, not on trajectory/constraint safety; validated on synthetic benchmarks and a manipulator balancing a pendulum with an imperfect simulator.
The 6th International Verification of Neural Networks Competition (VNN-COMP 2025): Summary and Results
Konstantin Kaulen, Tobias Ladner et al.
2025 arXiv 2512.19007 (Dec 2025); report of the competition he… Certified NNs ★★★ Competition report: 8 teams, 16 regular and 9 extended benchmarks, standardized ONNX/VNN-LIB formats and equal-cost AWS hardware. Overall score (Table 6): 1. alpha,beta-CROWN 1566.9, 2. NeuralSAT 1430.2, 3. PyRAT 1228.4, 4. CORA 987.2, 5. NNV 796.4, 6. nnenum 740.3.
Tighter Confidence Bounds for Sequential Kernel Regression
Hamish Flynn, David Reeb
2025 AISTATS 2025 (PMLR 258:3844-3852) Safe exploration & BO ★★★ Builds confidence sequences for sequential kernel regression from martingale tail bounds; the exact UCB is a finite-dimensional second-order cone program and its dual reduces to a 1-D minimization. Proven always tighter than Abbasi-Yadkori and Chowdhury-Gopalan bounds, and improves KernelUCB/GP-UCB empirically with matching worst-case guarantees.
1-Lipschitz Layers Compared: Memory, Speed, and Certifiable Robustness
Bernd Prach, Fabio Brau et al.
2024 CVPR 2024, pp. 24574-24583, DOI 10.1109/CVPR52733.2024.02… Certified NNs ★★★ Systematic theoretical and empirical comparison of 1-Lipschitz layer constructions (AOL, SOC, BCOP, Cayley, SLL, CPL, LOT, ...) in memory, speed and certified robust accuracy, with practical selection guidance and code.
A Survey of Constraint Formulations in Safe Reinforcement Learning
Akifumi Wachi, Xun Shen, Yanan Sui
2024 IJCAI 2024 (Survey Track) Constrained RL ★★★ Surveys constraint formulations (expected cumulative, state constraints, joint chance, expected instantaneous, almost-sure cumulative/instantaneous) with representative algorithms and proves transformability and conservative-approximation relations among them.
ECLipsE: Efficient Compositional Lipschitz Constant Estimation for Deep Neural Networks
Yuezhu Xu, S. Sivaranjani
2024 NeurIPS 2024 (Advances in Neural Information Processing S… PauliCertified NNs ★★★ Exact decomposition of the LipSDP certificate into a sequence of layer-size matrix inequalities, giving ECLipsE (small SDPs) and ECLipsE-Fast (closed form); up to thousands of times faster with near-LipSDP tightness.
Efficiently Computable Safety Bounds for Gaussian Processes in Active Learning
Jörn Tebbe, Christoph Zimmer et al.
2024 AISTATS 2024 (PMLR 238:1333-1341) Safe exploration & BO ★★★ For safe active learning along input trajectories (as in Zimmer et al. 2018), the probability that a GP posterior trajectory violates a safety threshold must be estimated for every candidate; naive Monte-Carlo needs many samples for small violation probabilities. The paper gives (i) an adaptive sequential MC test with controlled misclassification probability (Theorem 1, 'AMC'), (ii) an analytical upper bound on the trajectory violation probability via the Borell-TIS inequality applied to a centered/normalized GP, in terms of the median of the supremum and the maximal pointwise variance (Theorem 4), and (iii) adaptive estimation of that median with finite-sample order-statistic confidence intervals ('AB'/'ABM', Theorem 5, Corollary 6). Validated in simulation and on a high-pressure fluid engine system.
Gameplay Filters: Robust Zero-Shot Safety through Adversarial Imagination
Duy P. Nguyen, Kai-Chieh Hsu et al.
2024 Conference on Robot Learning (CoRL 2024), Proceedings of … Control-theoretic safety ★★★ A predictive safety filter that plays out hypothetical 'gameplay' between a simulation-trained safety policy and a co-trained virtual adversary invoking worst-case events; intervention is triggered when the imagined game is lost. First full-order safety filter for 36-D quadrupedal dynamics, demonstrated zero-shot on hardware under tugging and unmodeled terrain.
How to Train Your Neural Control Barrier Function: Learning Safety Filters for Complex Input-Constrained Systems
Oswin So, Zachary Serlin et al.
2024 2024 IEEE International Conference on Robotics and Automa… Control-theoretic safety ★★★ Diagnoses why neural CBF training struggles with input constraints and high dimension and proposes policy neural CBFs (PNCBF): learn the maximum-over-time constraint value of a nominal policy, which is itself a valid CBF, optionally iterated like policy iteration. Demonstrated up to a 16-D F-16 model and two-quadcopter hardware.
Information-Theoretic Safe Bayesian Optimization
Alessandro G. Bottero, Carlos E. Luis et al.
2024 arXiv preprint (arXiv:2402.15347, v2 May 2024) Safe exploration & BO ★★★ ISE-BO extends ISE from pure safe exploration to safe optimization by taking, over the certified safe set, the maximum of the ISE safe-exploration criterion and a max-value entropy search (MES) criterion for the safe optimum; proves safety and, on finite domains, that the largest safely reachable region is classified safe and the uncertainty at the safe optimum vanishes.
Learning Soft Constrained MPC Value Functions: Efficient MPC Design and Implementation providing Stability and Safety Guarantees
Nicolas Chatzikiriakos, Kim P. Wabersich et al.
2024 Proceedings of the 6th Annual Learning for Dynamics & Con… Pauli ★★★ Learns the value function of a tightened soft-constrained MPC rather than its possibly discontinuous policy; proves local Lipschitz continuity of that value function and ISS w.r.t. the approximation error, with constraint satisfaction under sufficient tightening.
Local Bayesian optimization for controller tuning with crash constraints
Alexander von Rohr, David Stenger et al.
2024 at - Automatisierungstechnik 72(4):281-292 (DOI 10.1515/a… Trimpe group ★★★ VDP-GIBO extends gradient-information BO (GIBO) to crash constraints using adaptive virtual observations at crashed parameters, batch design-of-experiments that minimizes gradient uncertainty, and resets to feasible points. Simulation and hardware tests show efficient tuning in larger parameter spaces.
Marabou 2.0: A Versatile Formal Analyzer of Neural Networks
Haoze Wu, Omri Isac et al.
2024 CAV 2024 (LNCS, pp. 249-264, DOI 10.1007/978-3-031-65630-… Certified NNs ★★★ Second-generation successor of Reluplex/Marabou (CAV 2019, LNCS pp. 443-452): an SMT-style complete verifier combining simplex-based LP reasoning, abstract interpretation / bound tightening (DeepPoly-type), case splitting with conflict learning, proof production and support for many piecewise-linear constraints and Python/ONNX front-ends.
Off-Policy Primal-Dual Safe Reinforcement Learning
Zifan Wu, Bo Tang et al.
2024 ICLR 2024 (poster) Constrained RL ★★★ CAL: shows off-policy primal-dual methods underestimate cumulative cost and therefore violate constraints; fixes it with conservative policy optimization (an upper-confidence cost estimate from an ensemble of cost critics) and local policy convexification (an augmented-Lagrangian-style quadratic term) that shrinks the estimation uncertainty over training.
OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning Research
Jiaming Ji, Jiayi Zhou et al.
2024 Journal of Machine Learning Research 25(285):1-6 (2024) Constrained RL ★★★ Modular PyTorch framework implementing dozens of SafeRL algorithms across on-policy, off-policy, model-based and offline settings, with adapters (e.g., Saute MDP) and torch.distributed parallel acceleration.
On the Scalability and Memory Efficiency of Semidefinite Programs for Lipschitz Constant Estimation of Neural Networks: Scaling the Computation for ImageNet
Zi Wang, Bin Hu et al.
2024 ICLR 2024 (conference paper; OpenReview id dwzLn78jq7; no… Pauli ★★★ Rewrites LipSDP as an exact-penalty eigenvalue optimization (EP-LipSDP) by adding a redundant trace constraint to the dual, so first-order autodiff solvers (LipDiff) can be used; scales SDP-based Lipschitz estimation to CIFAR-10 and ImageNet networks.
PACSBO: Probably approximately correct safe Bayesian optimization
Abdullah Tokmak, Thomas B. Schön, Dominik Baumann
2024 Symposium on Systems Theory in Data and Optimization (Sys… Safe exploration & BO ★★★ Addresses the practical weak point that Fiedler et al. (2024) identified in safe BO (the need for a known RKHS-norm bound): estimates a local RKHS-norm bound from data in a PAC sense and plugs it into the confidence bounds, with hardware experiments. Baumann (GoSafe, first author) and Tokmak (safe exploration in RKHS, in the index) are Trimpe-group alumni/collaborators.
Safe Bayesian Optimization for the Control of High-Dimensional Embodied Systems
Yunyue Wei, Zeji Yi et al.
2024 CoRL 2024 Safe exploration & BO ★★★ HdSafeBO combines isometric embeddings, trust-region local search and optimistic (UCB-based) safety identification with a per-step probabilistic safety guarantee and a cumulative-violation bound, scaling safe BO to hundreds or thousands of dimensions.
Safe reinforcement learning in uncertain contexts
Dominik Baumann, Thomas B. Schön
2024 IEEE Transactions on Robotics 40:1828-1841 (DOI 10.1109/T… Trimpe group ★★★ By a Trimpe alumnus; Trimpe is not a co-author. Performs safe learning when discrete contexts (payloads, surfaces) cannot be measured: contexts are identified with MMD tests on excitation trajectories and estimated online with frequentist multi-class classification bounds built on conditional mean embeddings; validated on a Furuta pendulum with camera images.
State Augmented Constrained Reinforcement Learning: Overcoming the Limitations of Learning With Rewards
Miguel Calvo-Fullana, Santiago Paternain et al.
2024 IEEE Transactions on Automatic Control 69(7):4275-4290 (J… Constrained RL ★★★ Exhibits CMDPs whose optimal policy cannot be induced by any fixed weighted combination of rewards (so regularized/classical primal-dual can fail at primal recovery), and proposes augmenting the state with the Lagrange multipliers, so that the dual dynamics become part of the state and the augmented policy provably attains the constrained optimum.
State space representations of the Roesser type for convolutional layers
Patricia Pauli, Dennis Gramlich, Frank Allgöwer
2024 IFAC-PapersOnLine (2024), DOI 10.1016/j.ifacol.2024.10.19… Pauli ★★★ Explicit Roesser realization of 2-D convolutional layers, shown to be minimal when c_in = c_out and K[r1,r2] has full rank; extended to dilated, strided and N-D convolutions.
Trust the Model Where It Trusts Itself - Model-Based Actor-Critic with Uncertainty-Aware Rollout Adaption
Bernd Frauenknecht, Artur Eisele et al.
2024 ICML 2024, PMLR 235:13973-14005 (arXiv 2405.19014) Trimpe group ★★★ Trimpe-group MBRL paper (MACURA) that adapts model-rollout length per step from the learned model's own uncertainty estimate, with a data-driven threshold; state-of-the-art data efficiency on MuJoCo. It is the methodological base for the group's 2026 uncertainty-aware predictive safety filters and Dyna-style safety-augmented RL already in the index.
Efficient Bound of Lipschitz Constant for Convolutional Layers by Gram Iteration
Blaise Delattre, Quentin Barthélemy et al.
2023 ICML 2023, PMLR 202:7513-7532 (arXiv 2305.16173) Pauli ★★★ A precise, fast and differentiable upper bound on the spectral norm of convolutional layers (Gram iteration on the circulant/FFT structure) that converges super-linearly and is usable as a training regularizer; it is the main non-SDP competitor that LipKernel, the 2-D-systems paper and the 1-D CNN Gramian approach are compared with.
Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs
Dongsheng Ding, Chen-Yu Wei et al.
2023 NeurIPS 2023 Constrained RL ★★★ Formulates the discounted CMDP as a constrained saddle-point problem and gives two single-time-scale algorithms with last-iterate guarantees: RPG-PD (entropy-regularized policy gradient + quadratically regularized dual ascent; sublinear last-iterate convergence) and OPG-PD (optimistic primal and dual updates; linear rate to a saddle point containing an optimal constrained policy).
Provably Safe Reinforcement Learning: Conceptual Analysis, Survey, and Benchmarking
Hanna Krasowski, Jakob Thumm et al.
2023 Transactions on Machine Learning Research (2023); arXiv 2… Constrained RL ★★★ Defines the taxonomy of provably safe RL (action replacement, action projection, action masking), analyses when each gives hard guarantees, and benchmarks them on control tasks (action replacement performs best). It is the standard organizing reference for shielding/safety-filter-style RL, complementing the CMDP-centric surveys already in the index.
Unlocking Deterministic Robustness Certification on ImageNet
Kai Hu, Andy Zou et al.
2023 NeurIPS 2023 (Advances in Neural Information Processing S… Certified NNs ★★★ LiResNet: a residual architecture (linear residual blocks) whose global Lipschitz bound stays tight and cheap, plus the EMMA loss (Efficient Margin MAximization, penalizing worst-case margins against all classes at once); the first Lipschitz-based deterministic l2 certification at ImageNet scale.
Bounding the difference between model predictive control and neural networks
Ross Drummond, Stephen R. Duncan et al.
2022 Proceedings of the 4th Annual Learning for Dynamics and C… Pauli ★★★ SDP bound on the worst-case gap between an MPC law and an NN policy over a box of states, combining activation QCs, QP/KKT constraints of MPC and box-set constraints; the same machinery synthesizes MPCs robustly approximating a given NN.
Chordal Sparsity for Lipschitz Constant Estimation of Deep Neural Networks
Anton Xue, Lars Lindemann et al.
2022 61st IEEE Conference on Decision and Control (CDC 2022), … PauliCertified NNs ★★★ Shows the LipSDP constraint has chordal sparsity and decomposes it into an equivalent set of smaller LMIs (Chordal-LipSDP) plus a tunable sparsity parameter.
Constrained Update Projection Approach to Safe Policy Optimization
Long Yang, Jiaming Ji et al.
2022 NeurIPS 2022 Constrained RL ★★★ Derives GAE-based surrogate performance bounds that unify earlier bounds, and a two-step (improve, then project) update solved by first-order primal-dual optimization without convex approximations; with worst-case improvement and violation guarantees.
Constraints Penalized Q-learning for Safe Offline Reinforcement Learning
Haoran Xu, Xianyuan Zhan, Xiangyu Zhu
2022 AAAI 2022 Constrained RL ★★★ CPQ makes out-of-distribution actions look unsafe by inflating their cost values (a CQL-like penalty on CVAE-detected OOD actions) and backs up and optimizes reward only through state-actions deemed safe.
COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation
Jongmin Lee, Cosmin Paduraru et al.
2022 ICLR 2022 (spotlight) Constrained RL ★★★ Optimizes directly over stationary distribution corrections w = d^pi/d^D with f-divergence regularization, reducing offline constrained RL to one convex minimization in (lambda, nu); adds cost upper-bound estimation for conservative constraint satisfaction.
Gaussian Process Uniform Error Bounds with Unknown Hyperparameters for Safety-Critical Applications
Alexandre Capone, Armin Lederer, Sandra Hirche
2022 ICML 2022 (PMLR 162:2609-2624) Safe exploration & BO ★★★ Robust GP uniform error bounds when kernel lengthscales are unknown: build a posterior confidence region over lengthscales and bound the error for every lengthscale in it using the posterior std of the smallest lengthscale, inflated by an explicit factor.
General Cutting Planes for Bound-Propagation-Based Neural Network Verification
Huan Zhang, Shiqi Wang et al.
2022 NeurIPS 2022 (Advances in Neural Information Processing S… Certified NNs ★★★ GCP-CROWN: generalizes beta-CROWN so that arbitrary linear cutting-plane constraints (e.g. from a MIP solver run in parallel on CPU) can be added to GPU bound propagation, completely solving the oval20 benchmark and verifying twice as many oval21 instances as prior tools; part of alpha,beta-CROWN (VNN-COMP 2022 winner).
Meta-Learning Priors for Safe Bayesian Optimization
Jonas Rothfuss, Christopher Koenig et al.
2022 CoRL 2022 (PMLR 205:237-265, published 2023) Safe exploration & BO ★★★ Meta-learns GP priors for safe BO from offline related tasks (F-PACOH) and selects safety-compliant hyperparameters by maximizing sharpness subject to empirical calibration, solved with an efficient frontier search (SaMBO).
On Controller Tuning with Time-Varying Bayesian Optimization
Paul Brunzema, Alexander von Rohr, Sebastian Trimpe
2022 61st IEEE Conference on Decision and Control (CDC 2022), … Trimpe group ★★★ UI-TVBO models incremental, lasting changes such as wear by injecting uncertainty through a Wiener-process temporal kernel instead of back-to-prior forgetting. It encodes convexity of the tuning objective via GPs with linear inequality constraints, which reduces regret and the number of unstable controllers tried.
Penalized Proximal Policy Optimization for Safe Reinforcement Learning
Linrui Zhang, Li Shen et al.
2022 IJCAI 2022 Constrained RL ★★★ P3O replaces constrained policy iteration by one unconstrained minimization with a ReLU exact penalty and PPO clipping, proves exactness for a finite penalty factor, and extends to multi-constraint and multi-agent settings.
Robustness analysis and training of recurrent neural networks using dissipativity theory
Patricia Pauli, Julian Berberich, Frank Allgöwer
2022 at - Automatisierungstechnik 70(8):730-739 (2022), DOI 10… Pauli ★★★ Comprehensive introduction to robustness analysis of RNNs with robust control and dissipativity theory, using H2 performance and l2-gain as robustness measures w.r.t. input perturbations, and LMI constraints to enforce them in training (abstract verified via Crossref).
safe-control-gym: a Unified Benchmark Suite for Safe Learning-based Control and Reinforcement Learning in Robotics
Zhaocong Yuan, Adam W. Hall et al.
2022 IEEE Robotics and Automation Letters 7(4):11142-11149 (Oc… Constrained RL ★★★ Open-source PyBullet benchmark (cart-pole, 1D and 2D quadrotor; stabilization and trajectory tracking) that extends the Gym API with symbolic a-priori dynamics, queryable constraints and injectable disturbances, to compare control, learning-based control and RL on performance, data efficiency and safety.
Controller Design via Experimental Exploration with Robustness Guarantees
Tobias Holicki, Carsten W. Scherer, Sebastian Trimpe
2021 IEEE Control Systems Letters 5(2):641-646 (2021; DOI 10.1… Trimpe group ★★★ For a partially unknown linear plant described by an LFR P_0 = Δ_0 ⋆ P with Δ_0 in a known set Δ, the set Δ is partitioned into Δ_1,...,Δ_N and, for each member, a robustly stabilizing controller is synthesized by IQC/LMI relaxation. The resulting low-dimensional family of test controllers F(θ) is guaranteed to stabilize the unknown plant, so data-driven (e.g. BO-type) exploration of the closed-loop cost over θ is safe by construction, and the coarseness of the partition trades optimality for exploration cost.
Event-triggered Learning for Linear Quadratic Control
Henning Schlüter, Friedrich Solowjow, Sebastian Trimpe
2021 IEEE Transactions on Automatic Control 66(10):4485-4498 Trimpe group ★★★ Derives the exact moment-generating function of the finite-horizon LQ cost under the model. Chernoff bounds on it give a learning trigger that detects model mismatch from observed costs with a prescribed false-alarm probability and triggers re-identification.
Globally-Robust Neural Networks
Klas Leino, Zifan Wang, Matt Fredrikson
2021 ICML 2021 (PMLR 139, pp. 6212-6222); arXiv 2102.08452 Certified NNs ★★★ GloRo nets: add a 'bottom' logit equal to the best competing logit plus the Lipschitz-scaled radius, so a global Lipschitz bound yields certification at inference cost of one forward pass and training directly maximizes verifiable accuracy; shows tight global bounds are attainable and that the maximum verifiable accuracy is the same for local and global bounds.
Misspecified Gaussian Process Bandit Optimization
Ilija Bogunovic, Andreas Krause
2021 NeurIPS 2021 Safe exploration & BO ★★★ Formalizes epsilon-misspecified kernelized bandits (f* is only epsilon-close in sup norm to a bounded-norm RKHS function) and gives the enlarged-confidence EC-GP-UCB, a phased-elimination variant (Phased GP Uncertainty Sampling) that adapts to unknown epsilon, and an epsilon-agnostic regret-balancing scheme for the contextual case.
On exploration requirements for learning safety constraints
Pierre-François Massiani, Steve Heim, Sebastian Trimpe
2021 3rd Conference on Learning for Dynamics and Control (L4DC… Trimpe group ★★★ Asks which state-action constraints must be learned to safely constrain a given nominal policy. They need not be accurate everywhere: they only have to contain the optimal viable policy OPT(Q_V) and exclude the 'critical set' of unviable pairs that are closer to the nominal policy than the optimal viable action.
Probabilistic Robust Linear Quadratic Regulators with Gaussian Processes
Alexander von Rohr, Matthias Neumann-Brosig, Sebastian Trimpe
2021 3rd Conference on Learning for Dynamics and Control (L4DC… Trimpe group ★★★ Synthesizes LQR controllers for linearized GP dynamics that are robust with respect to a probabilistic stability margin. It combines a scenario-based common-Lyapunov LMI initialization with a-priori guarantees, iterative improvement of the expected cost, and a-posteriori sample-based (Hoeffding) validation.
Robust Control Barrier-Value Functions for Safety-Critical Control
Jason J. Choi, Donggun Lee et al.
2021 2021 60th IEEE Conference on Decision and Control (CDC), … Control-theoretic safety ★★★ Unifies HJ reachability and CBFs through the control barrier-value function (CBVF), the viscosity solution of a discounted HJI variational inequality; its zero-superlevel set recovers the viability kernel, and a QP yields a disturbance-robust safety controller smoother than the least-restrictive switch.
Safe Reinforcement Learning Using Robust MPC
Mario Zanon, Sébastien Gros
2021 IEEE Transactions on Automatic Control 66(8):3638–3652 (D… Control-theoretic safety ★★★ Uses a robust (tube) MPC scheme as the RL policy/value approximator so that policy updates by Q-learning or policy gradient stay within the class of controllers that are safe and stable by construction, even with an imperfect model; a projection keeps the learned parameters inside the set where the robust MPC guarantees hold.
Skew Orthogonal Convolutions
Sahil Singla, Soheil Feizi
2021 ICML 2021 (PMLR 139, pp. 9756-9766); arXiv 2105.11417 Certified NNs ★★★ Constructs orthogonal convolutions as the exponential of a skew-symmetric convolution, approximated by a truncated Taylor series with a provable error bound; faster and more accurate than earlier gradient-norm-preserving convolutions.
WCSAC: Worst-Case Soft Actor Critic for Safety-Constrained Reinforcement Learning
Qisong Yang, Thiago D. Simão et al.
2021 AAAI 2021 (Proc. AAAI 35(12)) Constrained RL ★★★ Extends SAC-Lagrangian with a distributional (Gaussian) safety critic that learns the mean and variance of the cost return and constrains its closed-form CVaR at risk level alpha instead of the expected cost.
Cautious Model Predictive Control Using Gaussian Process Regression
Lukas Hewing, Juraj Kabzan, Melanie N. Zeilinger
2020 IEEE Transactions on Control Systems Technology 28(6):273… Control-theoretic safety ★★★ The reference GP-based learning MPC: a nominal model plus a GP residual, approximate propagation of the state distribution, chance constraints tightened by the GP variance, and sparse GP approximations for real-time use (autonomous racing). It is the model most learning-MPC and predictive-safety-filter papers compare against.
IPO: Interior-point Policy Optimization under Constraints
Yongshuai Liu, Jiaxin Ding, Xin Liu
2020 AAAI 2020 (Proc. AAAI 34(04):4940-4947) Constrained RL ★★★ Interior-point-inspired first-order method that augments PPO's clipped objective with logarithmic barriers of the cumulative constraints; handles multiple and general cumulative constraints.
Lipschitz Bounded Equilibrium Networks
Max Revay, Ruigang Wang, Ian R. Manchester
2020 arXiv preprint 2010.01732 (Oct 2020) Certified NNs ★★★ First direct (unconstrained) parameterization of equilibrium networks with prescribed Lipschitz bounds and guaranteed well-posedness under weaker conditions than prior work; connects LipSDP-type certificates to convex optimization and neural ODEs. Cited by Wang & Manchester (ICML 2023) as the source of the direct-parameterization idea and of the explanation of the LipSDP-Network proof error.
Safe and Fast Tracking on a Robot Manipulator: Robust MPC and Neural Network Control
Julian Nubert, Johannes Köhler et al.
2020 IEEE Robotics and Automation Letters 5(2):3050-3057 Trimpe group ★★★ A robust setpoint-tracking MPC (nonlinear constraint tightening via incremental stability, artificial steady states) that unifies planning and control, plus an NN approximation with statistical validation. It is the first demonstration of both on a real KUKA LBR4+ manipulator.
Safe Linear Stochastic Bandits
Kia Khezeli, Eilyan Bitar
2020 AAAI 2020 (Proceedings of the AAAI Conference on Artifici… Safe exploration & BO ★★★ Safe linear bandits in which the expected reward itself must exceed a threshold at every stage, given a known safe baseline arm; SEGE explores with convex combinations of the baseline and random exploratory arms within the exploration budget b0 - b and exploits the certainty-equivalent (greedy) arm only when its lower confidence bound is above the threshold.
Safe Reinforcement Learning in Constrained Markov Decision Processes
Akifumi Wachi, Yanan Sui
2020 ICML 2020 (PMLR 119:9797-9806) Safe exploration & BO ★★★ SNO-MDP first expands a pessimistic (SafeMDP-style) safe region and then optimizes the GP-modeled reward inside it, with an early-stopping rule (ES2) that halts safety exploration once further expansion cannot improve the return. Proves safety and near-optimality (PAC-MDP style); introduces GP-Safety-Gym.
Lyapunov-based Safe Policy Optimization for Continuous Control
Yinlam Chow, Ofir Nachum et al.
2019 arXiv preprint arXiv:1901.10031 (v1 28 Jan 2019, v2 11 Fe… Control-theoretic safety ★★★ Extends the Lyapunov CMDP approach of Chow et al. (NeurIPS 2018) to continuous actions and policy-gradient methods (DDPG, PPO). Two mechanisms: θ-projection (policy-parameter projection onto the feasible set induced by linearized Lyapunov constraints, shown to coincide with the CPO constraint minus the line search) and a-projection (a differentiable 'Lyapunov safety layer', in the spirit of Dalal et al. 2018, that projects the unconstrained action onto a linearized Lyapunov half-space with a closed-form solution). Evaluated on MuJoCo safety tasks and real-robot navigation.
Preventing Gradient Attenuation in Lipschitz Constrained Convolutional Networks
Qiyang Li, Saminul Haque et al.
2019 NeurIPS 2019 (Advances in Neural Information Processing S… Certified NNs ★★★ BCOP: an expressive parameterization of orthogonal (gradient-norm-preserving) convolutions built from block convolutions of symmetric projectors, enabling large provably 1-Lipschitz CNNs (with GroupSort) for certified robustness and Wasserstein estimation; analyzes the connected components of orthogonal convolutions.
Sorting out Lipschitz function approximation
Cem Anil, James Lucas, Roger Grosse
2019 ICML 2019 (arXiv 1811.05381) Pauli ★★★ Identifies gradient-norm preservation as necessary for expressive 1-Lipschitz networks and proposes GroupSort (MaxMin) with norm-constrained weights; such networks are universal approximators of Lipschitz functions.
Learning-Based Robust Model Predictive Control with State-Dependent Uncertainty
Raffaele Soloperto, Matthias A. Müller et al.
2018 6th IFAC Conference on Nonlinear Model Predictive Control… Trimpe group ★★★ Robust MPC for linear systems subject to bounded, state-dependent uncertainty: instead of a uniform worst-case disturbance set, the scheme explicitly exploits that the uncertainty set W(x) is smaller where a learned model is more confident, which is the natural setting for learning-based MPC where the model comes from data. Illustrated with a Gaussian-process regression example. (Full text could not be retrieved during verification; bibliographic data confirmed via Crossref/OpenAlex/Semantic Scholar, affiliations Stuttgart + MPI-IS.)
Safe Active Learning for Time-Series Modeling with Gaussian Processes
Christoph Zimmer, Mona Meister, Duy Nguyen-Tuong
2018 NeurIPS 2018 Safe exploration & BO ★★★ Safe active learning of GP NX (nonlinear exogenous) dynamics models: piecewise input trajectories are chosen to maximize an optimal-design criterion of the predictive covariance subject to a GP-based probabilistic safety constraint on the whole trajectory; the determinant criterion is linked to information gain and gamma_n.
Semidefinite relaxations for certifying robustness to adversarial examples
Aditi Raghunathan, Jacob Steinhardt, Percy Liang
2018 NeurIPS 2018 (Advances in Neural Information Processing S… Certified NNs ★★★ General SDP relaxation for certifying arbitrary ReLU networks (not just one hidden layer), tighter than LP relaxations and giving meaningful certificates on 'foreign' networks trained without the relaxation in the loop.
Spectral Normalization for Generative Adversarial Networks
Takeru Miyato, Toshiki Kataoka et al.
2018 ICLR 2018; arXiv 1802.05957 Certified NNs ★★★ Normalizes every weight matrix by its spectral norm, estimated with one power-iteration step per update, to bound the Lipschitz constant of GAN discriminators; became the default cheap layer-wise Lipschitz control.
Parseval Networks: Improving Robustness to Adversarial Examples
Moustapha Cisse, Piotr Bojanowski et al.
2017 ICML 2017 (PMLR 70, pp. 854-863); arXiv 1704.08847 Certified NNs ★★★ Keeps linear/convolutional weight matrices approximately Parseval tight frames (Lipschitz <= 1) and aggregation layers as convex combinations, maintained by a cheap retraction after every SGD step; an early Lipschitz-control defense.
Zames-Falb multipliers for absolute stability: From O'Shea's contribution to convex searches
Joaquin Carrasco, Matthew C. Turner, William P. Heath
2016 European Journal of Control 28:1-19 (March 2016), DOI 10.… Pauli ★★★ Tutorial/survey of Zames-Falb multipliers for slope-restricted nonlinearities, their history and convex search methods (incl. FIR and discrete-time versions).
Bayesian Optimization with Inequality Constraints
Jacob R. Gardner, Matt J. Kusner et al.
2014 ICML 2014 (PMLR 32(2):937-945) Safe exploration & BO ★★★ Extends BO to expensive unknown inequality constraints by weighting expected improvement over the best feasible point with the GP probability of feasibility (constrained EI).
Bayesian Optimization with Unknown Constraints
Michael A. Gelbart, Jasper Snoek, Ryan P. Adams
2014 UAI 2014 (30th Conference on Uncertainty in Artificial In… Safe exploration & BO ★★★ Formulates constrained BO with noisy, a-priori unknown constraints as a stochastic program with probabilistic (chance) constraints, models objective and each constraint with independent GPs via latent constraint functions g_k with C_k(x) <=> g_k(x) >= 0, and proposes constraint-weighted expected improvement, including a feasibility-search fallback when no point satisfies the probabilistic constraint and a decoupled variant where objective and constraints can be evaluated separately. Provides no guarantee that individual evaluations are safe; it is the standard 'constrained BO' baseline to contrast with SafeOpt-style safe exploration.
ECLipsE-Gen-Local: Efficient Compositional Local Lipschitz Estimates for Deep Neural Networks
Yuezhu Xu, S. Sivaranjani
2026 Transactions on Machine Learning Research, 2026 (arXiv 25… Pauli ★★ Extends the compositional ECLipsE recursion to local Lipschitz estimates over input regions, keeping the layer-size matrix inequalities and closed-form variants.
LipSSM: Structurally Lipschitz-Bounded Cascaded State-Space Model via Metric Transfer between Consecutive SSM Layers
Natsuki Yoshino, Ren Uchida et al.
2026 arXiv preprint 2609.30973 (25 Sept 2026; submitted to IEE… Pauli ★★ Extends the dissipative-layer / metric-transfer idea of LipKernel to cascaded state-space-model layers, giving structurally Lipschitz-bounded SSMs.
Approximate Nonlinear Model Predictive Control With Safety-Augmented Neural Networks
Henrik Hose, Johannes Köhler et al.
2025 IEEE Transactions on Control Systems Technology, vol. 33,… Certified NNs ★★ Approximates a nonlinear MPC's whole input sequence by a neural network and augments it with an online feasibility/safety check (constraint satisfaction and a decrease condition) with fallback to a safe candidate sequence, giving deterministic closed-loop safety and convergence guarantees despite approximation error.
Improved Scalable Lipschitz Bounds for Deep Neural Networks
Usman Syed, Bin Hu
2025 arXiv preprint 2503.14297 (Mar 2025) Pauli ★★ Improves the scalable compositional (ECLipsE-type) Lipschitz bounds for deep networks.
Robust direct data-driven control for probabilistic systems
Alexander von Rohr, Dmitrii Likhachev, Sebastian Trimpe
2025 Systems & Control Letters 196:106011 (2025; DOI 10.1016/j… Trimpe group ★★ Combines the scenario approach with direct data-driven (Willems'-lemma / informativity) LMI synthesis. Trajectories from several realizations of a system with aleatoric variation, such as a robot fleet, yield a controller that quadratically stabilizes unseen realizations with high probability, with a lower bound on the number of trajectories needed.
Safety in safe Bayesian optimization and its ramifications for control
Christian Fiedler, Johanna Menn, Sebastian Trimpe
2025 arXiv preprint 2501.13697 (23 Jan 2025); extended abstrac… Trimpe group ★★ Control-audience extended abstract of the TMLR 2024 paper 'On Safety in Safe Bayesian Optimization'. It states explicitly: 'This extended abstract, which has been presented as a poster at ... SysDO 2024, disseminates results from the journal paper [7]. The content of all sections is adapted and all plots and results are taken verbatim from [7].' Two practical obstacles are highlighted: (i) SafeOpt-type implementations replace theoretically valid beta_t by heuristics (usually beta_t = 2), losing all guarantees; (ii) valid bounds need an RKHS-norm upper bound that engineering prior knowledge cannot supply. LoSBO (safety from a known Lipschitz bound and bounded noise only) and LoS-GP-UCB (gridding-free variant) are the proposed fixes. Not an independent algorithmic advance; useful as a citation for the control framing.
SB-TRPO: Towards Safe Reinforcement Learning with Hard Constraints
Dominik Wagner, Ankit Kanwar, Luke Ong
2025 arXiv preprint arXiv:2512.23770 (Dec 2025, revised Sep 2026) Constrained RL ★★ Safety-Biased TRPO for zero-cost (hard) constraints: each trust-region step must achieve a fixed fraction beta of the maximal cost reduction attainable inside the trust region, leaving the rest for reward; converges to zero-cost solutions in finite MDPs and a gradient-based practical variant improves safety and reward on Safety-Gymnasium level-2 tasks.
Towards a Practical Understanding of Lagrangian Methods in Safe Reinforcement Learning
Lindsay Spoor, Álvaro Serra-Gómez et al.
2025 arXiv preprint arXiv:2510.17564 (Oct 2025, v2 Mar 2026) Constrained RL ★★ Empirical study of the return-cost trade-off as a function of the Lagrange multiplier on eight Safety-Gymnasium tasks: builds empirical Pareto frontiers over fixed lambda, compares multiplier update rules (fixed, gradient ascent, PI control), shows strong lambda sensitivity and task-dependent constraint restrictiveness, and recommends per-task cost limits.
Last-Iterate Global Convergence of Policy Gradients for Constrained Reinforcement Learning
Alessandro Montenegro, Marco Mussi et al.
2024 NeurIPS 2024 Constrained RL ★★ C-PG: alternating primal ascent / dual descent on a dual-regularized Lagrangian with dimension-free last-iterate global convergence under (weak) gradient-domination assumptions; action-based (C-PGAE) and parameter-based (C-PGPE) variants and an extension to risk-measure constraints.
OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning
Yihang Yao, Zhepeng Cen et al.
2024 NeurIPS 2024 Constrained RL ★★ Uses a conditional diffusion model to synthesize an offline dataset whose distribution is shaped toward high-reward, constraint-satisfying trajectories, then trains an offline safe RL policy on the shaped data; improves data efficiency and robustness to imperfect datasets.
Neural Network Verification in Control
Michael Everett
2021 2021 60th IEEE Conference on Decision and Control (CDC), … Certified NNs ★★ Tutorial connecting open-loop NN robustness verification (bound propagation, relaxations) to closed-loop reachability and robust deep RL for neural feedback loops.
A Self-Tuning LQR Approach Demonstrated on an Inverted Pendulum
Sebastian Trimpe, Alexander Millane et al.
2014 19th IFAC World Congress (Cape Town, 2014); IFAC Proceedi… Trimpe group ★★ Pre-Bayesian-optimization origin of the group's controller-tuning line: the LQR design weights are iteratively adapted from experimentally measured closed-loop cost on an inverted pendulum, using Simultaneous Perturbation Stochastic Approximation (SPSA) as a gradient-estimate optimizer, while the nominal linear model is kept fixed and only used to synthesize the gain. As stated by Marco et al. (ICRA 2016), SPSA finds local minima and does not exploit past data, which motivated replacing it by GP-based Entropy Search. No safe-exploration guarantee.
Robust and optimal predictive control of the COVID-19 outbreak
Johannes Köhler, Lukas Schwenkel et al.
2020 Annual Reviews in Control (2020) (arXiv 2005.03580) Pauli ★ Robust and optimal MPC for epidemic (COVID-19) mitigation; Pauli's only 2020-2026 arXiv paper outside neural-network robustness.