The large models remaking software.
Generative AI is organized around foundation models: large systems that can generate and interpret language, code, images, audio and video, then serve as a base for products and workflows. Competition now extends beyond headline model releases to reasoning, multimodality, context length, deployment cost, open-weight access, infrastructure capacity and the governance of increasingly consequential systems.
Foundation models are broad AI systems designed to support many downstream tasks rather than one narrowly defined application. Large language models are the most visible category, generating and analyzing text and code, while multimodal models extend that scope to combinations of text, images, audio and video. Transformer architectures directly enabled the recent generation of powerful generative models.
The public expansion of the field was propelled by products including OpenAI's ChatGPT and DALL·E, alongside image-generation systems such as Midjourney and Stable Diffusion. OpenAI's GPT-4 was introduced as a large multimodal model, illustrating the shift from text-only interfaces toward systems that can work across media. These models have affected technology platforms including Apple, Amazon, Meta, Google and Microsoft.
Model development has moved toward systems with wider input and output capabilities. Google introduced Gemini 1.5 with context windows reaching up to 1 million tokens, and later offered a version of Gemini 1.5 Pro with up to 2 million tokens in private preview. Long context is one route to making models useful with large bodies of text, code and other material.
Video, audio and interactive-world generation broaden the definition of generative AI beyond chat and static images. Meta announced Movie Gen models for realistic video and audio clips, while Google DeepMind released Genie 3, which can generate 3D worlds from a prompt and maintain visual memory through a few minutes of continuous interaction. Google's Gemini Omni was presented as a multimodal model beginning with video generation, reflecting an effort to consolidate multiple media capabilities within a common model layer.
Frontier-model competition is expressed through model releases, benchmark results, context capacity and claims of improved reasoning and coding. Google said Gemini 3 Pro set records on vision benchmarks and outperformed Claude Opus 4.5 and GPT-5.1 in some categories, while OpenAI described GPT-4.5 as its most knowledgeable model but cautioned that it was not a frontier model and could perform below o1 or o3-mini on some measures. Such comparisons show that leadership is increasingly task-specific rather than reducible to one universal ranking.
A parallel emphasis on reasoning has emerged as developers seek gains beyond conventional prediction-oriented language models. OpenAI's o1 was characterized as a notable shift toward reasoning models, and reports on Gemini 3 attributed its gains to better pre-training and post-training. Model claims remain closely tied to the evaluation design, the task category and whether results come from internal or external testing.
Scale remains a visible dimension of the field: Google's PaLM was described as a 540 billion-parameter dense decoder-only Transformer, Meta's LLaMA was released in sizes from 7 billion to 65 billion parameters, and Mistral Large 2 was announced with 123 billion parameters. Yet model size is not the sole route to usefulness. Researchers are also developing smaller open-source generative models that can outperform much larger models on particular tasks.
Open and open-weight releases have created a second competitive track alongside proprietary frontier systems. Meta positioned LLaMA as a foundational model for researchers, Google released Gemma-family models including the multimodal Gemma 3n designed to run with as little as 2GB of memory, and Z.ai released GLM models with open-source or open-weight positioning. Licensing, hardware requirements, inference cost and local deployment therefore shape adoption alongside raw capability.
As foundation models become operating inputs for enterprises, buyers can treat providers as increasingly substitutable while comparing performance, integration, reliability and inference cost. Workflow-native AI raises the bar further: value depends on fitting tools, data, hardware and established production habits, not simply on a general-purpose conversational interface. Uses in advertising and in Broadway's large-scale narrative video production show how generative capability can be incorporated into existing creative systems.
The field also faces constraints that are technical, legal and institutional. Researchers have found that several models can reproduce long excerpts from training data when prompted strategically, while companies have lagged worker adoption in setting employee-use guidelines for generative AI. At the frontier, access and development are shaped not only by research and product execution, but by compute and power capacity, safety assurance, rights and consent controls for synthetic media, regulatory legitimacy, and relationships with governments.
Grounded in the archive and knowledge graph. Browse all topic guides, the concept reference, or the posts.