Learning to Leverage Compliance:
A Policy–Admittance Learning Framework for Robotic Insertion

Chongren Wang*Minghe Li*Honghua Dai†Zhicheng LinShiyang WeiXiaokui Yue

School of Astronautics
Northwestern Polytechnical University

* Equal contribution   † Corresponding author

LeCo learns to leverage fixed admittance for robotic insertion across four connector-assembly tasks.

Abstract

Policy learning and compliant control offer a promising route to reliable autonomous assembly under pose errors and contact uncertainty. However, combining them does not ensure coordination: the policy may continue pushing against contact while the controller yields, producing sustained loading with limited progress. To address this problem, we propose LeCo (Leverage Compliance), a policy–admittance learning framework that trains a visual policy to leverage fixed admittance for effective insertion with reduced contact loads. A multirate feedback mechanism aggregates high-rate contact-interaction records into policy-transition rewards. An integrated conflict cost then characterizes sustained policy-loading/controller-unloading opposition, while a directional high-force tail cost captures continued-loading events within a transition. Combined with a task-completion reward, these costs encourage the policy to leverage compliance with less unproductive loading. We evaluate LeCo on four real connector-assembly tasks, obtaining an aggregate success rate of 94%. Across tasks, mean successful-trial resultant-force and torque peaks decrease by approximately 30% and 64% relative to the comparison baseline. Reward ablation further shows that adding conflict shaping reduces median successful-trial contact-conditioned conflict density by approximately 53%. These results support integrating multirate policy–admittance execution with interaction-based reward shaping to achieve effective, lower-load insertion.

Method

A 10-Hz visual policy operates with fixed 100-Hz admittance. High-rate contact records are aggregated into transition-level rewards to guide policy–admittance cooperation.

LeCo architecture: robot observations feed a 10-Hz visual policy, fixed 100-Hz admittance adjusts execution, and contact interaction provides transition-level rewards.
An integrated conflict cost captures sustained policy-loading/controller-unloading opposition, while a directional high-force tail cost captures continued-loading events within a transition.

Experiments

We evaluate LeCo on four real connector-assembly tasks.

Aviation-connector insertion

Early pose correction avoids jamming; axial motion completes final seating.

Single-trial force and torque

Selected successful trials with recorded resultant force and torque.

Aviation connector

Successful episode 03 · 8.66 s

Results

LeCo achieves 94% aggregate success across four tasks, with lower mean successful-trial lateral-force, resultant-force, and torque peaks than the evaluated SOTA baselines.

Autonomous evaluation against SOTA baselines
Assembly taskMethodSuccessLateral force
Fxy peak (N)
Resultant force
peak (N)
Resultant torque
peak (N·m)
Aviation connectorHIL-SERL14 / 209.809 ± 1.23911.950 ± 0.9652.304 ± 0.252
ConRFT17 / 206.658 ± 1.83310.701 ± 1.6931.569 ± 0.433
LeCo19 / 203.121 ± 0.6697.760 ± 0.8070.722 ± 0.147
Anderson headHIL-SERL26 / 306.147 ± 1.0567.856 ± 1.1191.701 ± 0.242
ConRFT30 / 305.702 ± 0.8627.293 ± 1.4091.563 ± 0.249
LeCo30 / 302.236 ± 0.6546.188 ± 1.3090.596 ± 0.172
Anderson sideHIL-SERL4 / 2010.822 ± 3.53912.883 ± 1.9692.165 ± 0.667
ConRFT13 / 209.531 ± 0.88712.437 ± 1.0492.089 ± 0.137
LeCo17 / 202.849 ± 0.6517.379 ± 1.6910.640 ± 0.151
ATX 20+4-pinHIL-SERL22 / 307.876 ± 1.0159.009 ± 1.0411.601 ± 0.178
ConRFT18 / 307.617 ± 1.2899.096 ± 1.5551.537 ± 0.357
LeCo28 / 302.931 ± 0.7426.454 ± 0.9290.483 ± 0.147
Online learning curves for LeCo, ConRFT and HIL-SERL on the four tasks, showing lateral-force peaks, intervention ratios and autonomous success.
Online learning curves. Rows show recorded XY-force peaks, intervention ratios, and autonomous success. Online time excludes offline training.