/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

Meta releases OpenEQA, a benchmark to measure an AI agent's understanding of physical spaces by probing it with questions about the environment

that on difficult tasks, VLMs regress to being nearly blind!  Visual content provides minor improvement to a VLM over an LLM, even when these are questions about visual content.  Language does not contain the answer, only vision does.  Why?  Because language provides easy priors about the world and “seeing” is hard... Yoav Artzi / @yoavartzi : We notice similar trends, @anne_youw put up a short report saying the same on NLVR ¯\_(ツ)_/¯ https://arxiv.org/... Nice thing about NLVR, it's tiny and super accessible to work with (even if you don't have Meta's resources 😉) Jesse Thomason / @_jessethomason_ : It's true, and the need for single-modality ablations during model building and dataset curation extends beyond classification tasks into actions too. A few years ago we found that many “embodied” agents end up either ignoring language OR ignoring vision. https://arxiv.org/... @jalexinee : it's happening! this is how, us women, we're gonna be replace by AI gf Hawkins Entrekin / @hawkinsentrekin : @DhruvBatraDB Decently quick rate of improvement though for the models. Wonder if with enough data / next gen models this gap can be closed anytime soon or if this will be more analogous to driverless cars where improvement rate slows @teortaxestex : This is also what @ylecun means when saying that GPT-4 or “GPT-5000” doesn't have a cat's understanding of the world and can't anticipate simple physical causal chains not present in texts. It's easy to mock his dismissive takes - but, incredibly, there is evidence behind them. [image]

VentureBeat Michael Nuñez

Context & Ripple Effects

OpenEQA extends Meta's sequence of work on machine perception: I-JEPA’s world-knowledge approach to vision was followed by V-JEPA’s video-based prediction of missing visual information. The new evaluation focus matters because it tests whether those representations translate into answers about real environments rather than just stronger visual generation or prediction.

The reported difficulty of these tasks also sharpens a distinction between language-derived priors and visual grounding. A benchmark targeted at that gap gives researchers a shared way to identify where vision-language models are not reliably using what they see.

First-order effects

  • Meta and other model developers gain a targeted test for measuring an agent’s ability to answer questions whose evidence is in a physical scene.
  • Vision-language models that perform well on language-heavy evaluations may show weaker results when prompts require spatial or environmental understanding.

Second-order effects

  • Model teams will face pressure to report performance on grounded visual tasks alongside broader language benchmarks, making failures in perception easier to compare.
  • Developers of agent and embodied-AI systems can use a common evaluation target to distinguish improvements in visual grounding from gains driven mainly by linguistic priors.

Third-order effects

  • If such evaluations become widely used, AI progress claims will increasingly be separated by whether models can ground answers in sensory evidence, not merely produce plausible text.
  • The field may move toward more task-specific benchmark suites for agents, with physical-world understanding treated as a distinct capability rather than an extension of general language reasoning.

The trend: AI evaluation is shifting from broad model scores toward benchmarks that isolate the perceptual and grounding capabilities agents need to act reliably beyond text.

Discussion

  • @yoavartzi Yoav Artzi on x
    We notice similar trends, @anne_youw put up a short report saying the same on NLVR ¯\_(ツ)_/¯ https://arxiv.org/... Nice thing about NLVR, it's tiny and super accessible to work with (even if you don't have Meta's resources 😉)
  • @_jessethomason_ Jesse Thomason on x
    It's true, and the need for single-modality ablations during model building and dataset curation extends beyond classification tasks into actions too. A few years ago we found that many “embodied” agents end up either ignoring language OR ignoring vision. https://arxiv.org/...
  • @jalexinee @jalexinee on x
    it's happening! this is how, us women, we're gonna be replace by AI gf
  • @hawkinsentrekin Hawkins Entrekin on x
    @DhruvBatraDB Decently quick rate of improvement though for the models. Wonder if with enough data / next gen models this gap can be closed anytime soon or if this will be more analogous to driverless cars where improvement rate slows
  • @teortaxestex @teortaxestex on x
    This is also what @ylecun means when saying that GPT-4 or “GPT-5000” doesn't have a cat's understanding of the world and can't anticipate simple physical causal chains not present in texts. It's easy to mock his dismissive takes - but, incredibly, there is evidence behind them. […
  • @dhruvbatradb Dhruv Batra on x
    I have been working on vision+language models (VLMs) for a decade.  And every few years, this community re-discovers the same lesson — that on difficult tasks, VLMs regress to being nearly blind!  Visual content provides minor improvement to a VLM over an LLM, even when these are…