Google unveils two text-to-video AI generators: Imagen Video, a higher image quality system, and Phenaki, which prioritizes coherency and length over quality
Not to be outdone by Meta's Make-A-Video, Google today detailed its work on Imagen Video, an AI system that can generate video clips given …
The announcement also begins an Imagen product arc that later included Imagen 2’s image-quality upgrade and ImageFX’s rollout on top of Imagen 2. That makes the early video work relevant as part of Google’s broader generative-media model development, even though the related coverage does not establish a product launch for either video system.
First-order effects
Google gains two distinct text-to-video research positions: Imagen Video for higher visual quality and Phenaki for longer, more coherent output.
Meta’s Make-A-Video is immediately measured against Google on the limits Meta had already exposed—brief clips and no audio—while Google frames a different set of output trade-offs.
Second-order effects
Developers and prospective creative-tool partners must evaluate text-to-video systems by the intended workflow—visual fidelity versus sustained sequence coherence—instead of treating generation as a single capability.
Google’s decision to develop video alongside the Imagen family gives it a path to connect image-generation advances with later creative interfaces such as ImageFX, increasing pressure on Meta to pair model announcements with usable editing and generation tools.
Third-order effects
If leading labs continue to optimize different generation objectives, synthetic-media competition will organize around specialized models and product workflows rather than one universally best text-to-video engine.
The pattern points toward generative-media platforms differentiating through control and integration layers as much as raw model output, a foundation for [[a:concept|synthetic media control plane]]—but the supplied coverage does not show how either 2022 video model will be commercialized.
The trend: Text-to-video is evolving from isolated model demonstrations into a broader generative-media stack in which quality, temporal coherence and creative-tool integration are separate competitive levers.
last week, meta unveiled its project to generate an entire video from a short text prompt. this week, google is doing the same thing. h/t @_akhaliq https://imagen.research.google/ video/ https://imagen.research.google/ ... https://twitter.com/...
Cannot full wrap my mind around the idea that we're on the verge of a world where you can write a script, press a button and have a movie https://twitter.com/...
if someone can explain to me how this differs from meta's method, i'd love to hear (i've looked through both research papers and it seems like the methods differ, but i am just a simple AI reporter...) https://twitter.com/...
If you're keeping score at home, two text-to-video AI generators — a category that did not exist a month ago — have come out of Google *today*. https://twitter.com/...
All media, all creative industries, will be a first line casualty of the AI dawn. Millions of jobs, are completely hosed within what, 5 years? Universal Basic Income needs to be a massive priority moving forward. We can't adapt this quickly. https://twitter.com/...
Absolutely worth digging into the Imagen Video paper for more examples, some of these are really extraordinary https://imagen.research.google/ ... https://twitter.com/...
Cannot really emphasize enough how fast AI is moving. There have been several major releases *in the month since I wrote a column about how fast AI is moving*, including OpenAI's Whisper (speech-to-text transcription) and now text-to-video. https://www.nytimes.com/... https://twi…
(2) Imagen Video (https://imagen.research.google/ video/) It's so elegant how time is just another dimension in the diffusion process, and by “inpainting” in time we can do text-to-video prediction. Amazing how the we can now use this technology to generate ~720p videos, in just …
meta ai: we have a text-to-video system! google ai: hold my beer in a glistening pint glass, the glass has precipitation dripping down the sides, on a wooden table, golden hour, bokeh. https://twitter.com/...
Amazing that Google Brain drops not just one, but *two* text-to-video papers, at the same time! (1) Phenaki: Variable Length Video Generation from Open Domain Textual Descriptions https://phenaki.github.io/ https://twitter.com/...