Learning to Leverage Compliance:
A Policy–Admittance Learning Framework for Robotic Insertion
School of Astronautics
Northwestern Polytechnical University
Abstract
Policy learning and compliant control offer a promising route to reliable autonomous assembly under pose errors and contact uncertainty. However, combining them does not ensure coordination: the policy may continue pushing against contact while the controller yields, producing sustained loading with limited progress. To address this problem, we propose LeCo (Leverage Compliance), a policy–admittance learning framework that trains a visual policy to leverage fixed admittance for effective insertion with reduced contact loads. A multirate feedback mechanism aggregates high-rate contact-interaction records into policy-transition rewards. An integrated conflict cost then characterizes sustained policy-loading/controller-unloading opposition, while a directional high-force tail cost captures continued-loading events within a transition. Combined with a task-completion reward, these costs encourage the policy to leverage compliance with less unproductive loading. We evaluate LeCo on four real connector-assembly tasks, obtaining an aggregate success rate of 94%. Across tasks, mean successful-trial resultant-force and torque peaks decrease by approximately 30% and 64% relative to the comparison baseline. Reward ablation further shows that adding conflict shaping reduces median successful-trial contact-conditioned conflict density by approximately 53%. These results support integrating multirate policy–admittance execution with interaction-based reward shaping to achieve effective, lower-load insertion.
Method
A 10-Hz visual policy operates with fixed 100-Hz admittance. High-rate contact records are aggregated into transition-level rewards to guide policy–admittance cooperation.
Experiments
We evaluate LeCo on four real connector-assembly tasks.
Aviation-connector insertion
Early pose correction avoids jamming; axial motion completes final seating.
Single-trial force and torque
Selected successful trials with recorded resultant force and torque.
Aviation connector
Results
LeCo achieves 94% aggregate success across four tasks, with lower mean successful-trial lateral-force, resultant-force, and torque peaks than the evaluated SOTA baselines.
| Assembly task | Method | Success | Lateral force Fxy peak (N) | Resultant force peak (N) | Resultant torque peak (N·m) |
|---|---|---|---|---|---|
| Aviation connector | HIL-SERL | 14 / 20 | 9.809 ± 1.239 | 11.950 ± 0.965 | 2.304 ± 0.252 |
| ConRFT | 17 / 20 | 6.658 ± 1.833 | 10.701 ± 1.693 | 1.569 ± 0.433 | |
| LeCo | 19 / 20 | 3.121 ± 0.669 | 7.760 ± 0.807 | 0.722 ± 0.147 | |
| Anderson head | HIL-SERL | 26 / 30 | 6.147 ± 1.056 | 7.856 ± 1.119 | 1.701 ± 0.242 |
| ConRFT | 30 / 30 | 5.702 ± 0.862 | 7.293 ± 1.409 | 1.563 ± 0.249 | |
| LeCo | 30 / 30 | 2.236 ± 0.654 | 6.188 ± 1.309 | 0.596 ± 0.172 | |
| Anderson side | HIL-SERL | 4 / 20 | 10.822 ± 3.539 | 12.883 ± 1.969 | 2.165 ± 0.667 |
| ConRFT | 13 / 20 | 9.531 ± 0.887 | 12.437 ± 1.049 | 2.089 ± 0.137 | |
| LeCo | 17 / 20 | 2.849 ± 0.651 | 7.379 ± 1.691 | 0.640 ± 0.151 | |
| ATX 20+4-pin | HIL-SERL | 22 / 30 | 7.876 ± 1.015 | 9.009 ± 1.041 | 1.601 ± 0.178 |
| ConRFT | 18 / 30 | 7.617 ± 1.289 | 9.096 ± 1.555 | 1.537 ± 0.357 | |
| LeCo | 28 / 30 | 2.931 ± 0.742 | 6.454 ± 0.929 | 0.483 ± 0.147 |