AIBullisharXiv – CS AI · Jun 47/10
🧠
VISTA: Vision-Grounded and Physics-Validated Adaptation of UMI data for VLA Training
VISTA is a new framework that improves robot learning by adapting real-world manipulation data collected via Universal Manipulation Interface (UMI) for training Vision-Language-Action (VLA) models. The framework addresses two key challenges: making distorted wrist-mounted camera views compatible with pre-trained vision models and filtering out physically infeasible trajectories before training, resulting in significantly better policy performance.