IMIGRASP: Zero-shot 6-DOF pose estimation via generative AI imagination and cross-image geometric transfer
DOI:
https://doi.org/10.31224/8042Keywords:
Generative AI for pose estimation, zero-shot robotic grasping, 6-DOF grasp generation, foundation models for manipulation, robotic affordance learning.Abstract
Robotic grasping in unstructured environments remains a fundamental challenge, particularly for novel objects not encountered during training. Traditional approaches rely on large scale annotated datasets, extensive training, and hand crafted features. This work introduces IMIGRASP (Imagine to Grasp), a zero-shot framework for 6-DOF pose estimation that leverages generative artificial intelligence for imagination guided pose estimation. The key insight is that modern generative models, trained on vast internet scale visual data, implicitly encode rich knowledge about human grasping behavior and object affordances. Given a single RGB image of a target object, the proposed pipeline employs GPT-5.6 Luna to generate a corresponding image with a realistic hand grasping the object, complete with colored landmarks at the thumb tip, index fingertip, and wrist. A state-of-the-art segmentation model, SAM3, is utilized for precise segmentation of objects and hands, while a bounding box transfer method is introduced for cross-image geometric transfer between generated and real images. The complete 3D grasp frame is reconstructed using RGB-D data and robot kinematics, and executed on a Franka Panda robot. The approach is validated on 130 diverse household objects, achieving a success rate of 86.9% without any prior training or object specific tuning. The results demonstrate that generative AI models can serve as powerful zero-shot priors for robotic manipulation, opening new directions for imagination guided robotics.
Downloads
Downloads
Posted
License
Copyright (c) 2026 Muhammad Aun Raza

This work is licensed under a Creative Commons Attribution 4.0 International License.