Google announces “multisearch” for Lens, letting users ask questions about photos and screenshots, available in beta in the US on iOS and Android
One of the most interesting announcements at Google's Search-focused keynote in September was a big upgrade to Lens that lets you take a photo and ask questions about it.
Context & Ripple Effects
Lens began as an in-photo computer-vision feature for Google Photos and Assistant, then reached iOS with a live viewfinder for object and text analysis. Multisearch adds text questions to that visual input, turning Lens from object recognition into a combined image-and-query interface.
The beta is an early step in a rollout path that later includes expansion to more than 70 languages, indicating that Google treated image-plus-text search as a broadly deployable Search capability rather than a one-platform experiment.
First-order effects
- US English users on Android and iOS can submit a photo or screenshot alongside a question in Lens, giving Google a new mobile-search entry point beyond typed queries.
- Google extends Lens across both major mobile platforms while the feature remains in beta, concentrating visual-query feedback and usage within its Search product.
Second-order effects
- Google’s planned language expansion makes localized visual queries the next operational requirement for Lens, broadening the set of users and screenshots the product must interpret.
- Traditional text-query discovery faces a more image-led starting point as users can frame searches around photographed objects, screens, and accompanying questions.
Third-order effects
- If Lens continues adding modalities—as in its later video-and-voice search update—mobile search is likely to organize around live camera context and multimodal prompts rather than a single search box.
- The pattern points toward sensor-level intelligence becoming a core search interface, with the quality of answers depending increasingly on Google’s ability to connect visual context to search results.
The trend: Google is moving Search toward multimodal, camera-centered queries that combine what a user sees with what they ask.