Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision–Language–Action Models
2026
Open publication workspace · Sign in to read the full PDF.
AI-generated summary
This paper introduces ACT2ANSWER, a novel embodied evaluation protocol to measure commonsense and world knowledge retention in Vision-Language-Action (VLA) models, addressing a critical gap in current VLA research.
* **ACT2ANSWER Protocol:** An embodied evaluation framework that adapts existing VLM knowledge benchmarks into action-based simulated episodes, enabling knowledge probing through simple selection actions.
* **Diverse Benchmark Suite:** A curated collection of 1,720 binary questions across 12 categories, designed to systematically evaluate commonsense and world knowledge in VLA models.
* **Empirical Study and Analysis:** A large-scale study of modern VLA systems and VLM baselines, revealing performance disparities across knowledge categories and introducing layerwise intent probing to analyze knowledge distribution within models.
Tags: VLA, Commonsense, World Knowledge, Embodied AI, Evaluation
Check the original publication for accuracy and context.