Skip to content
dotdock

Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision–Language–Action Models

2026

Publication cover

Open publication workspace · Sign in to read the full PDF.

AI-generated summary

This paper introduces ACT2ANSWER, a novel embodied evaluation protocol to measure commonsense and world knowledge retention in Vision-Language-Action (VLA) models, addressing a critical gap in current VLA research.

* **ACT2ANSWER Protocol:** An embodied evaluation framework that adapts existing VLM knowledge benchmarks into action-based simulated episodes, enabling knowledge probing through simple selection actions.
* **Diverse Benchmark Suite:** A curated collection of 1,720 binary questions across 12 categories, designed to systematically evaluate commonsense and world knowledge in VLA models.
* **Empirical Study and Analysis:** A large-scale study of modern VLA systems and VLM baselines, revealing performance disparities across knowledge categories and introducing layerwise intent probing to analyze knowledge distribution within models.

Tags: VLA, Commonsense, World Knowledge, Embodied AI, Evaluation

Check the original publication for accuracy and context.