Google adds image-to-video generation to Veo 3 in the Gemini app for Pro and Ultra subscribers, and says users have created 40M+ videos since Veo 3's May launch
Google said on Thursday it's adding an image-to-video generation feature to its Veo 3 AI video generator through its Gemini app.
Context & Ripple Effects
Veo 3 was introduced alongside Imagen 4 and Flow, positioning Google’s video model within a broader creative-tool stack. A subsequent rollout to AI Pro subscribers across 159+ countries widened access shortly before this update.
Adding image-to-video to Gemini shifts Veo 3 from prompt-only clip creation toward a workflow in which users can animate an existing visual asset. Google’s reported 40M-plus videos provides an early usage signal, though it does not by itself establish sustained creator or business adoption.
First-order effects
- Pro and Ultra subscribers can use a source image as the starting point for Veo 3 output in Gemini, giving them more control over a clip’s visual starting point.
- The feature makes Gemini a more central access point for Google’s generative-video offering, while keeping the capability tied to paid subscription tiers.
Second-order effects
- Creators and teams can test image-led video creation without moving a visual asset into a separate video-generation product, increasing the value of an integrated Veo, Imagen and Flow creative stack.
- Competing AI video services face added pressure to pair text-to-video quality with asset conditioning and workflow convenience, rather than compete on generation alone.
Third-order effects
- If image-to-video becomes a standard subscriber feature, generative video products may differentiate less through a single model prompt and more through control over source assets, editing, and distribution workflows.
- The pattern supports a broader shift toward synthetic-media tools being sold as recurring-access product suites; usage figures will matter most when paired with evidence of repeat, paid workflow adoption.
The trend: Generative video is moving from stand-alone prompt experiments toward subscription-based creative workflows that turn existing images into controllable video inputs.