SAVLA: Symmetry-Aware Vision-Language-Action Models for Robotic Manipulation

9月 16, 2026·
Junle Li
Weixian Waylon Li
Weixian Waylon Li
,
Fuxiang Wu
,
Fusheng Hao
,
Fengxiang He
· 1 分钟阅读时长
摘要
Vision-language-action (VLA) models have become the dominant paradigm for language-conditioned robot manipulation. However, although images and language instructions inherently encode geometric information, VLAs acquire their spatial competence purely from demonstrations. As a result, they are reliable only within the range of scene poses that the demonstrations cover. We propose SAVLA, an end-to-end symmetry-aware VLA model for robust and data-efficient policy learning. Our approach keeps the pretrained vision-language backbone entirely frozen while combining it with an equivariant flow-matching action head and a learned canonicalizer. The head decomposes its state, action, and conditioning inputs into invariant and equivariant channels, and preserves this typing throughout all of its layers. The canonicalizer transforms oblique-view images into a canonical frame and rotates the geometric conditions consistently. We evaluate our model on LIBERO. Compared with the GR00T N1.5 baseline, SAVLA improves the success rate averaged over all four LIBERO suites by 5.1 points and increases the mean success rate under rotation on LIBERO-Goal from 41.5% to 90.4%.
类型
出版物
Preprint 2026

Citation

@misc{li2026savlasymmetryawarevisionlanguageactionmodels,
      title={SAVLA: Symmetry-Aware Vision-Language-Action Models for Robotic Manipulation}, 
      author={Junle Li and Weixian Waylon Li and Fuxiang Wu and Fusheng Hao and Fengxiang He},
      year={2026},
      eprint={2609.16641},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2609.16641}, 
}