Meta releases OpenEQA, a benchmark to measure an AI agent's understanding of physical spaces by probing it with questions about the environment
that on difficult tasks, VLMs regress to being nearly blind! Visual content provides minor improvement to a VLM over an LLM, even when these are questions about visual content. Language does not contain the answer, only vision does. Why? Because language provides easy priors about the world and “seeing” is hard... Yoav Artzi / @yoavartzi : We notice similar trends, @anne_youw put up a short report saying the same on NLVR ¯\_(ツ)_/¯ https://arxiv.org/... Nice thing about NLVR, it's tiny and super accessible to work with (even if you don't have Meta's resources 😉) Jesse Thomason / @_jessethomason_ : It's true, and the need for single-modality ablations during model building and dataset curation extends beyond classification tasks into actions too. A few years ago we found that many “embodied” agents end up either ignoring language OR ignoring vision. https://arxiv.org/... @jalexinee : it's happening! this is how, us women, we're gonna be replace by AI gf Hawkins Entrekin / @hawkinsentrekin : @DhruvBatraDB Decently quick rate of improvement though for the models. Wonder if with enough data / next gen models this gap can be closed anytime soon or if this will be more analogous to driverless cars where improvement rate slows @teortaxestex : This is also what @ylecun means when saying that GPT-4 or “GPT-5000” doesn't have a cat's understanding of the world and can't anticipate simple physical causal chains not present in texts. It's easy to mock his dismissive takes - but, incredibly, there is evidence behind them. [image]
Context & Ripple Effects
OpenEQA extends Meta's sequence of work on machine perception: I-JEPA’s world-knowledge approach to vision was followed by V-JEPA’s video-based prediction of missing visual information. The new evaluation focus matters because it tests whether those representations translate into answers about real environments rather than just stronger visual generation or prediction.
The reported difficulty of these tasks also sharpens a distinction between language-derived priors and visual grounding. A benchmark targeted at that gap gives researchers a shared way to identify where vision-language models are not reliably using what they see.
First-order effects
- Meta and other model developers gain a targeted test for measuring an agent’s ability to answer questions whose evidence is in a physical scene.
- Vision-language models that perform well on language-heavy evaluations may show weaker results when prompts require spatial or environmental understanding.
Second-order effects
- Model teams will face pressure to report performance on grounded visual tasks alongside broader language benchmarks, making failures in perception easier to compare.
- Developers of agent and embodied-AI systems can use a common evaluation target to distinguish improvements in visual grounding from gains driven mainly by linguistic priors.
Third-order effects
- If such evaluations become widely used, AI progress claims will increasingly be separated by whether models can ground answers in sensory evidence, not merely produce plausible text.
- The field may move toward more task-specific benchmark suites for agents, with physical-world understanding treated as a distinct capability rather than an extension of general language reasoning.
The trend: AI evaluation is shifting from broad model scores toward benchmarks that isolate the perceptual and grounding capabilities agents need to act reliably beyond text.