Google says Gemini 3 Pro sets new vision AI benchmark records, including in complex visual reasoning, beating Claude Opus 4.5 and GPT-5.1 in some categories
Raising Concerns for Real-World Use Will McCurdy / PCMag : ChatGPT Overtakes Amazon, X, Reddit, WhatsApp, and Wikipedia in Visitors X: Demis Hassabis / @demishassabis : Gemini has always had exceptionally strong multimodal capabilities. Gemini 3 Pro is an incredible vision AI model and is SOTA across all main vision & multimodal benchmarks. It's great for document, screen, image, video & spatial understanding tasks - try now in the @GeminiApp! [image] Rohan Paul / @rohanpaul_ai : Google just published a deep dive on how they pushed Gemini 3 Pro's vision capabilities across document, spatial, screen and video understanding. They upgraded the whole vision pipeline, from perception to reasoning. The model now “derenders” messy scans into structured code [image] Jeff Dean / @jeffdean : One aspect of our Gemini 3 Pro model to look at is how it performs in multimodal capabilities. We've worked on making it perform really well across a variety of multimodal use cases, like understanding of documents, videos, spatial characteristics, biomedical data, and computer [image] LinkedIn: Tuan Nguyen, Ph.D : Our latest work has significantly elevated Gemini 3.0 Pro's spatial understanding of digital interfaces! … Apostol : Video is the richest, most complex data format we interact with, but for AI, the challenge has always been moving beyond simple recognition to true reasoning. …