Stability AI releases Stable Video Diffusion in research preview, its first foundation model for generative video based on the Stable Diffusion image model
a latent video diffusion model for high-resolution, state-of-the-art text-to-video and image-to-video generation.... [video] Tanishq Mathew Abraham, PhD / @iscienceluvr : Stability AI is releasing Stable Video Diffusion! 🔥 This is a new image-to-video model that can produce 14-25 frames at resolution 576x1024 given a context frame of the same size. Text2video model coming soon. Announcement: https://stability.ai/... Paper:... [video] @stabilityai : Today, we are releasing Stable Video Diffusion, our first foundation model for generative AI video based on the image model, @StableDiffusion. As part of this research preview, the code, weights, and research paper are now available. Additionally, today you can sign up for our... [video]
Context & Ripple Effects
Stable Video Diffusion extends Stability AI's image-model line into video through a research release with code, weights, and a paper. The launch establishes the base model that later supported Stable Video 3D and Stable Video 4D work.
The arc matters because it turns image-generation research into a reusable video foundation rather than a one-off demonstration. Later API availability for Stable Diffusion 3 shows the company also pursued a route from model previews to developer-facing access.
First-order effects
- Researchers and developers can inspect and experiment with Stability AI's video model using the released materials, centered on image-to-video generation at the stated output format.
- Stability AI gains a video counterpart to Stable Diffusion, giving its model ecosystem a foundation for subsequent video-specific tools.
Second-order effects
- Open research materials lower the barrier for third parties to test video-generation workflows and build specialized layers on top of the model, while competing model providers face pressure to show comparable developer access.
- The shared foundation creates a technical path for adjacent formats: later coverage ties it directly to multi-perspective video generation and 3D-video tooling.
Third-order effects
- If foundation models continue to be released with reusable artifacts, differentiation in generative media is likely to shift toward product integration, workflow controls, and distribution rather than model access alone.
- The progression from a video foundation model to derivative 3D and multi-view tools suggests generative-video competition may fragment into specialized production capabilities rather than remain a single text-to-video category.
The trend: This is an early instance of generative-media vendors turning image-model ecosystems into reusable video foundations that can support increasingly specialized creation tools.