Controlling a hybrid system to a target set under parameter uncertainty can require informative actions that temporarily drive the system state away from the specified set. We propose a belief-informed dual-control algorithm that combines belief-space receding-horizon selection with an expected-decrease constraint on a nonnegative target-set progress function. To accommodate exploratory deviations, a scalar controller state bounds the constraint's cumulative slack. Our algorithm's two-step lookahead selection uses predicted observations to update the parameter belief before evaluating the subsequent action's admissibility. Under distance-comparison bounds, correct conditional prediction, and recursive feasibility, we prove almost-sure target-set convergence at decision times and bound both the sum of expected progress-function values and the expected neighborhood-entry time. Upper confidence bounds extend these convergence guarantees to bounded model samples with summable error probabilities. In planar regulation with an unknown control direction, we verify recursive feasibility: the proposed two-step selection reduces the Euclidean state norm below 0.01 within 11 decisions for either sign, whereas a myopic one-step selection loses admissibility immediately after zero input. In a simulated bimanual assembly task, our algorithm's one-step implementation uses clearance feedback to complete the assembly under nominal friction, yielding a 96.4% lower infinity-norm relative-position error than the reference-tracking baseline at the method's completion time. At lower friction, its posterior conditioning successfully detects an empty admissible set, whereas a fixed-prior alternative admits an action that violates the conditional decrease constraint. Project page: https://clintonenwerem.com/belief-hybrid-control/.
Figures & tables
Fig. 1: Part Assembly Outcomes. We compare controllers on a contact-rich assembly task in which a bimanual robot with anthropomorphic end-effectors must align and insert a key-like part into a close-fitting housing (or socket) while accommodating uncertainty in the part–housing interaction. From the same initial state at θ=0.7 , our belief-informed control algorithm completes the assembly with 96.4% lower infinity-norm position error than a reference-tracking baseline. Upper panels show task endpoints with magnified insets. Lower panels show the key/socket sidewalls at a common scale together with dashed lines marking the desired part pose. Here pHP and p∗ denote the current and desired part-origin positions in the housing frame, dhand is the minimum hand–assembly clearance, and LH is the summed magnitude of hand–housing contact forces.
Fig. 2: Planar Regulation with Unknown Control Direction. In state coordinates (x1,x2) , ★ marks the target at the origin and the black dot marks x0 . The dotted circle shows rotation under the drift ωJx with zero input. Two-step selection first applies the test action; the change in radius identifies the control-direction sign θ . We then select the contracting radial feedback gain. Blue solid and green dashed curves show the trajectories for θ=+1 and θ=−1 , respectively; colored dots mark the test-interval endpoints.
Fig. 3: Assembly Task and Frames. Panel (a) shows the bimanual RealHand A7 robot with L6 hands relative to the world frame W . Panel (b) identifies the part frame P , housing frame H , and insertion axis aW=RH(1,0,0)T . At finger–part contact i , the world-expressed contact position and force are ciW and fiW . The part and housing poses are XP=(RP,pP) and XH=(RH,pH) , respectively.
Fig. 4: Action Selection Across Initial Beliefs and Slack Bounds. The panels show the first actions selected over a 19-by-21 grid of prior probabilities p=b0(−1) and normalized cumulative slack bounds β=B0/∥x0∥2 . Colors identify the selected action. Diagonal hatching denotes an empty admissible set, while crosshatching denotes admissible first actions with no feasible second decision. Across panels, we hold the plant, gains, initial state, and expected-decrease constraint fixed. Stars mark the nominal case (p,β)=(0.5,0.12) .
j
Residual cj(s)≤0
Scale λj
1–3
∣pHP,j−p∗,j∣−ϵp,j
(0.1,0.05,0.05)j
4
eR−0.06
0.2
5, 6
∥vi∥−0.005
0.1
7, 8
∥ωi∥−0.05
1
9
0.005−dhand
0.005
10
F∗−Ftable
F∗
TABLE I: Completion Residuals and Normalization Scales. Values use SI units. For paired object rows, i=P,H in that order.
qc
Mode
Forward Predicate
0
Housing acquisition
HH ; housing speeds ≤(0.05m/s,0.5rad/s)
1
Part acquisition
HP∧HH ; part speeds ≤(0.05m/s,0.5rad/s)
2
Lift
pP,z≥−0.4441m , FP≤0.12N
3
Transfer
∣pHP−pT∣≤ϵT , eR≤0.12rad
4
Alignment
∣pHP−pA∣≤ϵA , eR≤0.10rad
5
Insertion
gins(s) , Ftable≥F∗
TABLE II: Forward Guards for Ordered Assembly Modes
Parameter
Value
Friction hypotheses; initial weights
{0.35,0.7} ; (1/2,1/2)
c,B0,Hp
10−3,0.1,1
Plant step, hsim
1ms
Arm/hand stiffness (N m/rad)
180 / 40
Arm/hand damping (N m s/rad)
28 / 1.4
Action limits
0.1s , 8s
TABLE III: Controller and Simulation Parameters
Action
vσ
δf (rad)
Available qc
Hold
0
0
0–9
Advance
1
0
0–7
Tighten
0
0.03
1–5
Relax
0
−0.03
1–5
Clearance
1
0
6, 7
Clearance
0
0
8, 9
TABLE IV: Feedback Action Catalog
Fig. 5: Assembly Sequence. Both one-step selectors produce these states at friction θ=0.7 . Physical finger contacts support acquisition and transport. Completion follows table support, hand withdrawal, and two seconds of stationary dwell.
Quantity
Selector
Baseline
Final progress, V
0
3.94
Position error, ∥pHP−p∗∥∞ (mm)
0.188
5.20
Rotation error (rad)
0.00194
0.0453
Hand clearance (mm)
7.14
0
Hand–housing load , LH (N)
0
0.704
Total table support (N)
4.61
4.30
TABLE V: Assembly Outcomes at θ=0.7 . We evaluate all methods at 30.96 s, when the robot reaches A under either selector with identical endpoints. Mint denotes selection; salmon denotes the reference-tracking baseline.
Endpoint mean, μ
Required slack, e
Action
Posterior
Fixed prior
Posterior
Fixed prior
Hold
5.7805
5.7221
0.1220
0.0636
Advance
8.9906
6.8842
3.3321
1.2258
Tighten
5.7728
5.7184
0.1143
0.0600
Relax
5.7900
5.7267
0.1315
0.0682
TABLE VI: Action Evaluation at θ=0.35 , V≈5.6641 , and c=10−3 . Remaining slack is 0.0936 (posterior) and 0.09318 (fixed prior). Mint and salmon distinguish their tightening evaluations.
Hybrid dynamical systems provide a powerful modeling framework for robotic systems, particularly in contact-rich environments. However, ensuring safety and performance in such systems remains challenging due to the intricate coupling between continuous dynamics and discrete mode transitions. In this work, we extend classical Hamilton-Jacobi (HJ) reachability analysis, a formal verification method for continuous-time nonlinear systems, to hybrid dynamical systems. Our framework characterizes safe sets for hybrid systems through a generalized value function defined over both discrete and continuous states while accounting for control constraints and model uncertainty. We additionally provide a numerical algorithm to compute this value function. Building on these safe sets, we propose two different mechanisms to integrate performance objectives. First, we introduce a hybrid least-restrictive safety filter that intervenes on both the discrete and continuous components of a nominal controller only when necessary to avoid unsafe states, thereby preserving nominal behavior whenever possible. Second, we formulate and compute hybrid backward reach-avoid tubes, enabling the simultaneous enforcement of safety and goal-reaching behavior, an extension not previously addressed within hybrid HJ reachability. This enables the synthesis of continuous and discrete control policies that guarantee both safety and task completion. We validate our framework through simulation studies and real-world experiments on a quadrupedal robot, demonstrating its effectiveness in hybrid mode planning and safety-critical applications.
Javier Borquez, Shuang Peng, Somil Bansal
Universidad de Santiago de Chile · University of Southern California · Stanford University
Ensuring safety for black-box hybrid dynamical systems presents significant challenges due to their instantaneous state jumps and unknown explicit nonlinear dynamics. Existing solutions for strict safety constraint satisfaction, like control barrier functions (CBFs) and reachability analysis, rely on direct knowledge of the dynamics. Similarly, safe reinforcement learning (RL) approaches often rely on known system dynamics or merely discourage safety violations through reward shaping. In this work, we want to learn RL policies which provably satisfy affine state constraints in closed loop for black-box hybrid dynamical systems with affine reset maps. Our key insight is forcing the RL policy to be affine and repulsive near the constraint boundaries for the unknown nonlinear dynamics of the system, providing guarantees that the trajectories will not violate the constraint. We further account for constraint violation due to instantaneous state jumps that occur due to impacts or reset maps in the hybrid system by introducing a second repulsive affine region before the reset that prevents post-reset states from violating the constraint. We derive sufficient conditions under which these policies satisfy safety constraints in closed loop. We also compare our approach with state-of-the-art reward shaping and learned-CBF methods on hybrid dynamical systems like the constrained pendulum and paddle juggler environments. In both scenarios, we show that our methodology learns higher quality policies while always satisfying the safety constraints.
Model-based control can achieve reliable task performance, but its effectiveness depends on the accuracy of the underlying model. Robots operating under model uncertainty must often adapt to previously unseen payloads, objects, and interaction dynamics to complete a task successfully. Conventional approaches typically rely on a dedicated task- and control-agnostic excitation phase for estimating physics parameters, delaying task execution and collecting potentially irrelevant data. In this paper, we instead consider zero-shot task execution under parametric uncertainty, where online model learning and control proceed concurrently during task execution. Our approach formulates reference generation within a dual control framework to produce task-relevant, informative trajectories that reduce parameter uncertainty in directions critical to task success. We predefine a feedback policy with an explicit parameter adaptation law and optimize the reference through the resulting adaptive closed-loop dynamics. We propose two formulations: 1) minimizing expected task cost under parameter uncertainty, and 2) minimizing optimality loss, which quantifies the degradation in task performance caused by planning with an incorrect parameter estimate. By actively optimizing for informative trajectories via a natural Fisher information measure, we tightly approach a lower bound on task-relevant parameter uncertainty while simultaneously achieving reliable task execution. Across diverse tasks, controllers, and adaptation laws, targeted exploration enables the identification required for successful task execution, allowing robots to learn relevant physical parameters while completing the task in a single attempt.
Victor Vantilborgh, Hrishikesh Sathyanarayan, Guillaume Crevecoeur +2
Department of Electromechanical, Systems and Metal Engineering, Ghent University, Belgium · Department of Mechanical Engineering, Yale University, USA