This could be an interesting new paper by Trevor Darrell and his team. However, this paper is only three pages long, possibly a preliminary paper.
From the abstract:
"Two recent developments in computer vision have brought us closer to human-like perception.
First, multiview models trained on visual-spatial data have closed a longstanding gap between human and machine 3D perception; remarkably, these models demonstrate an emergent alignment to human error patterns and reaction times. Second, autoregressive gazing models trained for general-purpose reconstruction objectives exhibit an emergent alignment with human fixation patterns - despite no exposure to eye-tracking data.
Here we ask whether these developments can be integrated into a foveated multiview transformer.
We construct a model without any re-training, determine its zero-shot performance on 3D vision benchmarks, and evaluate its alignment to human behavior.
The model retains meaningful task performance across all benchmarks even with sparse, low-resolution inputs, and its gaze patterns correlate with human gaze despite no training on eye-tracking data.
These zero-shot results establish foveated multiview models as a promising direction for vision systems that are performant on 3D tasks and grounded in the mechanisms of human visual perception."
No comments:
Post a Comment