Amazing stuff! Less clutter is better for robots! 😊
This is a new research paper by Alan Yuille and his team published by CVPR 2026!
From the abstract:
"Occlusion remains a core challenge in vision, as projecting a 3D world into 2D inevitably hides much of the scene geometry.
We present FoundationDeOcclusion, a fast generative framework for fast occlusion recovery that reconstructs hidden geometry and appearance from partial visual observations. FoundationDeOcclusion first identifies occluded objects from monocular image sequences using Grounded-SAM, then matches them across views using depth cues estimated by 3D reconstruction models.
We introduce a novel geometry-aware linear de-occlusion Transformer (GL-DoT) that synthesizes the missing regions, which are then integrated into the 3D scene through depth-aware fusion.
As a result, FoundationDeOcclusion improves 3D scene reconstruction under occlusion. Notably, GL-DoT attains strong performance with only four denoising steps and runs at near real-time speed (6 FPS),
challenging the prevailing belief that generative models are impractical for time-critical robotic tasks.
Finally, we demonstrate the effectiveness of FoundationDeOcclusion in real-robot navigation and manipulation. Without bells and whistles, it boosts the prior-art AnyGrasp by 43.8% in success rate without task-specific tuning, setting a new state of the art for mobile manipulation under occlusion."
No comments:
Post a Comment