Apple researchers detail an AI system that can resolve references to elements displayed on a screen, in some cases better than GPT-4 can when given screenshots
new AI model could make Siri way faster and smarter Nickie Louise / Tech Startups : Apple's new AI model ReALM outperforms GPT-4 Tyler Lee / Phandroid : Apple might have just given us a sneak peek at its AI Justinas Vainilavičius / Cybernews.com : Apple experimenting with AI models that can “see” Alex Blake / TechRadar : Apple researchers reveal AI breakthrough that could make Siri much smarter James Kinoti / Cryptopolitan : Apple Introduces Cutting-Edge AI Model Surpassing GPT-4 Abhinav Anand / The Mac Observer : Apple Researchers Unveil ReALM To Upgrade Siri and Take on GPT-4 James Lee Taylor / Sammy Fans : ReALM: Apple races to close AI gap as Samsung, Google soar DigiTimes : Apple unveils new AI model boosting Siri, rivals GPT-4 in performance Rounak Jain / Benzinga : Apple Quietly Unveils New AI Model ReALM That Outperforms OpenAI's GPT-4 Despite Being ‘Much Lighter And Faster’ Boone Ashworth / Wired : How an iPhone Powered by Google's Gemini AI Might Work Zac Hall / 9to5Mac : Apple AI researchers boast useful on-device model that ‘substantially outperforms’ GPT-4 Wesley Hilliard / AppleInsider : Apple AI research: ReALM is smaller, faster than GPT-4 when parsing contextual data José Adorno / BGR : Apple says its latest AI model ReALM is even better than OpenAI's GPT4 X: Brian Roemmele / @brianroemmele : Apple says its latest AI model ReALM is even “better than OpenAI's GPT4”. It likely is as GPT4 has regressed because of “alignment”. The ReALM war begins at WWDC 2024. Paper: https://arxiv.org/... [image] Parker Ortolani / @parkerortolani : this could be big, it definitely seems to align with recent reporting on what types of AI tools are coming in iOS 18 not focused on chat, but far more contextually aware and able to better understand complex requests whether they're direct or ambient @_akhaliq : Apple announces ReALM Reference Resolution As Language Modeling Reference resolution is an important problem, one that is essential to understand and successfully handle context of different kinds. This context includes both previous turns and context that pertains to [image] Forums: Hacker News : Apple AI researcher boast on-device model that ‘substantially outperforms’ GPT-4 r/artificial : Apple researchers develop AI that can ‘see’ and understand screen context MacRumors Forums : Apple Researchers Reveal New AI System That Can Beat GPT-4
Context & Ripple Effects
Apple’s reported testing of Ajax models for Siri and Messages had already indicated that it was evaluating in-house language models against ChatGPT. ReALM adds a narrower but consequential capability: interpreting references to what is currently visible on a device screen.
The work also sits alongside Apple’s subsequent Ferret-UI research on mobile-screen understanding, suggesting a research track focused on assistants that can act on interface context rather than only respond to text prompts. Apple had not announced a ReALM product integration.
First-order effects
- ReALM gives Apple a research-backed approach to resolving ambiguous references to on-screen items, a key prerequisite for more context-aware assistant interactions.
- Its reported performance against GPT-4 in some screenshot-based reference-resolution tests strengthens the case for smaller, lighter models where screen context and response speed matter.
Second-order effects
- A Siri implementation would make UI-aware inference a more important competitive benchmark for assistant makers, rather than treating general conversational quality as the sole measure.
- The result supports Apple’s apparent split approach: develop specialized internal capabilities while separately testing external models for Siri and Messages, as indicated by the Ajax-versus-ChatGPT testing.
Third-order effects
- If specialized, on-device-capable models continue to handle device context well, assistant competition may shift toward tightly integrated system features rather than a single general-purpose model.
- This points toward more ambient assistants whose usefulness depends on access to interface state; the practical boundary will be determined by product integration and user controls, neither of which ReALM’s paper establishes.
The trend: ReALM is one data point in the move from standalone chatbots toward context-aware assistants optimized for the devices and interfaces they inhabit.