Controlling a hybrid system to a target set under parameter uncertainty can require informative actions that temporarily drive the system state away from the specified set. We propose a belief-informed dual-control algorithm that combines belief-space receding-horizon selection with an expected-decrease constraint on a nonnegative target-set progress function. To accommodate exploratory deviations, a scalar controller state bounds the constraint's cumulative slack. Our algorithm's two-step lookahead selection uses predicted observations to update the parameter belief before evaluating the subsequent action's admissibility. Under distance-comparison bounds, correct conditional prediction, and recursive feasibility, we prove almost-sure target-set convergence at decision times and bound both the sum of expected progress-function values and the expected neighborhood-entry time. Upper confidence bounds extend these convergence guarantees to bounded model samples with summable error probabilities. In planar regulation with an unknown control direction, we verify recursive feasibility: the proposed two-step selection reduces the Euclidean state norm below 0.01 within 11 decisions for either sign, whereas a myopic one-step selection loses admissibility immediately after zero input. In a simulated bimanual assembly task, our algorithm's one-step implementation uses clearance feedback to complete the assembly under nominal friction, yielding a 96.4% lower infinity-norm relative-position error than the reference-tracking baseline at the method's completion time. At lower friction, its posterior conditioning successfully detects an empty admissible set, whereas a fixed-prior alternative admits an action that violates the conditional decrease constraint. Project page: https://clintonenwerem.com/belief-hybrid-control/.
Figures & tables
Fig. 1: Part Assembly Outcomes. We compare controllers on a contact-rich assembly task in which a bimanual robot with anthropomorphic end-effectors must align and insert a key-like part into a close-fitting housing (or socket) while accommodating uncertainty in the part–housing interaction. From the same initial state at θ=0.7 , our belief-informed control algorithm completes the assembly with 96.4% lower infinity-norm position error than a reference-tracking baseline. Upper panels show task endpoints with magnified insets. Lower panels show the key/socket sidewalls at a common scale together with dashed lines marking the desired part pose. Here pHP and p∗ denote the current and desired part-origin positions in the housing frame, dhand is the minimum hand–assembly clearance, and LH is the summed magnitude of hand–housing contact forces.
Fig. 2: Planar Regulation with Unknown Control Direction. In state coordinates (x1,x2) , ★ marks the target at the origin and the black dot marks x0 . The dotted circle shows rotation under the drift ωJx with zero input. Two-step selection first applies the test action; the change in radius identifies the control-direction sign θ . We then select the contracting radial feedback gain. Blue solid and green dashed curves show the trajectories for θ=+1 and θ=−1 , respectively; colored dots mark the test-interval endpoints.
Fig. 3: Assembly Task and Frames. Panel (a) shows the bimanual RealHand A7 robot with L6 hands relative to the world frame W . Panel (b) identifies the part frame P , housing frame H , and insertion axis aW=RH(1,0,0)T . At finger–part contact i , the world-expressed contact position and force are ciW and fiW . The part and housing poses are XP=(RP,pP) and XH=(RH,pH) , respectively.
Fig. 4: Action Selection Across Initial Beliefs and Slack Bounds. The panels show the first actions selected over a 19-by-21 grid of prior probabilities p=b0(−1) and normalized cumulative slack bounds β=B0/∥x0∥2 . Colors identify the selected action. Diagonal hatching denotes an empty admissible set, while crosshatching denotes admissible first actions with no feasible second decision. Across panels, we hold the plant, gains, initial state, and expected-decrease constraint fixed. Stars mark the nominal case (p,β)=(0.5,0.12) .
j
Residual cj(s)≤0
Scale λj
1–3
∣pHP,j−p∗,j∣−ϵp,j
(0.1,0.05,0.05)j
4
eR−0.06
0.2
5, 6
∥vi∥−0.005
0.1
7, 8
∥ωi∥−0.05
1
9
0.005−dhand
0.005
10
F∗−Ftable
F∗
TABLE I: Completion Residuals and Normalization Scales. Values use SI units. For paired object rows, i=P,H in that order.
qc
Mode
Forward Predicate
0
Housing acquisition
HH ; housing speeds ≤(0.05m/s,0.5rad/s)
1
Part acquisition
HP∧HH ; part speeds ≤(0.05m/s,0.5rad/s)
2
Lift
pP,z≥−0.4441m , FP≤0.12N
3
Transfer
∣pHP−pT∣≤ϵT , eR≤0.12rad
4
Alignment
∣pHP−pA∣≤ϵA , eR≤0.10rad
5
Insertion
gins(s) , Ftable≥F∗
TABLE II: Forward Guards for Ordered Assembly Modes
Parameter
Value
Friction hypotheses; initial weights
{0.35,0.7} ; (1/2,1/2)
c,B0,Hp
10−3,0.1,1
Plant step, hsim
1ms
Arm/hand stiffness (N m/rad)
180 / 40
Arm/hand damping (N m s/rad)
28 / 1.4
Action limits
0.1s , 8s
TABLE III: Controller and Simulation Parameters
Action
vσ
δf (rad)
Available qc
Hold
0
0
0–9
Advance
1
0
0–7
Tighten
0
0.03
1–5
Relax
0
−0.03
1–5
Clearance
1
0
6, 7
Clearance
0
0
8, 9
TABLE IV: Feedback Action Catalog
Fig. 5: Assembly Sequence. Both one-step selectors produce these states at friction θ=0.7 . Physical finger contacts support acquisition and transport. Completion follows table support, hand withdrawal, and two seconds of stationary dwell.
Quantity
Selector
Baseline
Final progress, V
0
3.94
Position error, ∥pHP−p∗∥∞ (mm)
0.188
5.20
Rotation error (rad)
0.00194
0.0453
Hand clearance (mm)
7.14
0
Hand–housing load , LH (N)
0
0.704
Total table support (N)
4.61
4.30
TABLE V: Assembly Outcomes at θ=0.7 . We evaluate all methods at 30.96 s, when the robot reaches A under either selector with identical endpoints. Mint denotes selection; salmon denotes the reference-tracking baseline.
Endpoint mean, μ
Required slack, e
Action
Posterior
Fixed prior
Posterior
Fixed prior
Hold
5.7805
5.7221
0.1220
0.0636
Advance
8.9906
6.8842
3.3321
1.2258
Tighten
5.7728
5.7184
0.1143
0.0600
Relax
5.7900
5.7267
0.1315
0.0682
TABLE VI: Action Evaluation at θ=0.35 , V≈5.6641 , and c=10−3 . Remaining slack is 0.0936 (posterior) and 0.09318 (fixed prior). Mint and salmon distinguish their tightening evaluations.