Google demos new Gemini features, such as video search and Screenshare, which lets users ask questions based on their phone's screen, coming later in March 2025
Google is adding new features to its AI assistant, Gemini, that let users ask it questions using video and content on the screen in real time.
Context & Ripple Effects
Google had already positioned Gemini 2.0 for use in Search and AI Overviews, with an emphasis on systems that can plan and act; this demo extends that direction from text prompts toward live visual context. It also follows an earlier Gemini Live experience that was judged stronger conversationally than functionally because it depended on the cloud.
The announcement became a concrete device feature when Google rolled out Gemini Live video capabilities later that month, followed by camera and screenshare availability on Pixel and Galaxy flagships. That sequence matters because it turns a demo into a distribution question across Android hardware and services.
First-order effects
- Gemini users gain a new way to ask questions about what a phone camera captures or what is currently displayed, reducing the need to describe visual context in a text query.
- Google makes Gemini Live a more capable real-time assistant surface, while initially tying the experience to its own rollout and supported-device path.
Second-order effects
- Android device partners have a stronger reason to feature Gemini-enabled hardware and software experiences, as illustrated by the subsequent Pixel and Galaxy rollout to those flagship lines.
- Search, app, and support tasks that begin with a screen or camera view can shift toward conversational assistance, raising the importance of response quality and reliability in live, multimodal interactions.
Third-order effects
- If these capabilities spread beyond flagship launches, the phone assistant can become an always-available interpretation layer over apps and the physical environment rather than a destination users open only to type queries.
- That shift would intensify competition over who controls the assistant layer on devices, while making cloud dependence and access to device context enduring product constraints.
The trend: This is part of the move from chat-based AI toward multimodal assistants that operate across a user’s live device and visual context.