Steer2Grasp

Inference-Time Embodiment-Aware Steering for Diverse Physically Feasible Grasp Diffusion

Anonymous submission

Steer2Grasp overview: the robot embodiment and the environment signed distance field enter a frozen grasp diffusion model only at inference time, where a reward is calculated, weights are assigned and the particle population is resampled, yielding embodiment-aware grasps.

A grasp diffusion model is trained with no knowledge of the robot or the scene. Steer2Grasp leaves it frozen and introduces the embodiment and the environment only at sampling time: a deployment reward is evaluated on each particle, the population is reweighted by that reward and resampled, and probability mass moves onto grasp modes the robot can actually reach and execute.

Abstract

Current grasp diffusion models provide rich priors for generation, yet their object-centric approach can violate the kinematic and collision constraints imposed by the embodiment and the environment. Existing embodiment-aware methods primarily perform local corrections around generated grasps through gradient guidance or optimization, making it difficult to recover from fundamentally infeasible modes. We present Steer2Grasp, a training-free, embodiment-agnostic framework for inference-time grasp steering that adapts a frozen Cartesian grasp diffusion model using deployment-specific rewards. Through Feynman-Kac (FK) inspired particle reweighting and resampling, the method reallocates population mass from infeasible to high-reward grasp modes, enabling population-level mode transitions without modifying the pretrained diffusion model or requiring differentiable constraints. The framework enables a unified treatment for single and dual arm grasping through reachability and collision aware rewards, followed by gradient free gripper level local refinement. Across diverse objects, robot embodiments, and constrained environments, our method substantially improves feasible grasp generation while maintaining proximity to the underlying grasp prior.

Video

Method

Overview of the proposed method: reverse diffusion produces grasp hypotheses which are mapped to clean grasp estimates, evaluated against deployment-specific feasibility rewards, and used to reweight and resample the particle population.

Given an full object point cloud $P$ and a pretrained Cartesian space grasp diffusion model, the reverse diffusion produces grasp hypotheses $H_{t-1}$ which are mapped to their clean grasp estimates $\hat{H}_0(H_t)$ using the Tweedie estimate, which is evaluated against deployment specific feasibility rewards. The resulting reward $\mathbf{r}(\hat{H}_0;P,\mathcal{R},\mathcal{E})$ combine embodiment and environment information through multi-start IK, reachability and signed geometric clearances. These rewards instantiate the Feynman--Kac twisting potential used to re-weight and resample ($\mathbf{H}_{b}$) the particle population in turn removing particles ($\mathbf{H}_{a}$). Thus embodiment and environment constraints act through population level reweighting and resampling rather than modifying individual denoising transitions, enabling mode level steering toward feasible regions while retaining the multi modal structure of the pretrained grasp prior.

Qualitative Comparisons

Grasps generated by gradient guidance and by our method in the same scene, executed in MuJoCo.

Real World Experiments

Cylinder

2×

Easy — 3/3 successful trials

2×

Medium — 3/3 successful trials

2×

Hard — 2/3 successful trials

Bowl

2×

Easy — 3/3 successful trials

2×

Medium — 3/3 successful trials

2×

Hard — 3/3 successful trials

30 of 36 trials succeed overall. Failures are concentrated on the bucket in the hard setting, due to asymmetric loading and the lack of compliant torque control.

Results

We spawn up to four walls around an object and randomise the base position of each Franka Emika Panda, retaining 50 scenes per object that admit at least one executable grasp. Using 18 single-arm and 12 dual-arm objects, this gives 45,000 single-arm grasps and 60,000 dual-arm grasp pairs, all evaluated in MuJoCo. O denotes steering alone and O+LR adds local refinement.

Single-arm GraspGen prior
Dual-arm DAGDiff prior

Reachability vs collision rate better ↖

Both must hold: a valid joint configuration and an arm that clears the scene.

Single-arm
Dual-arm

Feasible vs stability better ↗

Clearing both constraints makes a grasp feasible; it still has to survive the pull test.

Single-arm
Dual-arm
Ours Baselines

Large-scale scene evaluation. Top row: A grasp needs a valid joint configuration and a collision-free arm. Satisfying both makes it feasible. Bottom row: feasible against stability, whether the grasp still holds the object under a 3 N pull along every object-frame axis.