OpenAI makes its Realtime API generally available with features like MCP support and debuts gpt-realtime, its most advanced speech-to-speech model, in the API
[video] @liodakis : Congrats to @pbbakkum on shipping gpt-realtime! It's been awesome watching him and the multimodal team sweat the details and get to a GA quality multimodal model. It's crazy to think we shipped a preview of realtime API less than a year ago and today we are sharing integrations Justin Uberti / @juberti : Huge Realtime API release today! Details below, but TLDR: - GA (out of beta) - better instruction following, naturalness, audio - MCP support - new voices - SIP (telephony) support - new WebRTC APIs and video support Demos: https://hello-realtime.val.run/, or call 425-800-0042 @vapi_ai : OpenAI's new GPT Realtime model is now live on Vapi! We got early access from OpenAI and were impressed: latency improvements are significant, voices are more natural, multilingual switching is great, accent imitation is impressive, and instruction following is more reliable. [image] Justin Uberti / @juberti : Yes, the Realtime API WebRTC interface supports video input, just add “video: true” to your getUserMedia call. Will make a demo of this soon! Sierra Catalina / @sierracatalina1 : openAI just shipped some major upgrades to the realtimeAPI [for those building with realtime voice], here's what matters: • voice: model speaks with human-like quality, shifts emotion on command [sad → excited] & switches languages mid-sentence [english, spanish, japanese] • @kwindla : Voice-only programming with the new OpenAI Realtime API ... I spend a lot of time these days pair programming with LLMs. Often I'm talking rather than typing. This “voice dictation” use case has become an important vibe benchmark for me. Being able to create text input just by [video] @openaidevs : We've added two API voices, Cedar and Marin, and improved voice quality overall for more humanlike speech that can adapt to tone. And we've cut pricing by 20%. Prompting guide on speed, tone, hand-offs, and more: https://cookbook.openai.com/ ... Blog post: https://openai.com/... @openaidevs : The Realtime API is officially out of beta and ready for your production voice agents! We're also introducing gpt-realtime—our most advanced speech-to-speech model yet—plus new voices and API capabilities: 🔌 Remote MCPs 🖼️ Image input 📞 SIP phone calling ♻️ Reusable prompts [video] Greg Brockman / @gdb : new speech-to-speech model and other platform improvements: @openaidevs : gpt-realtime was trained with customers to excel at real-world tasks like support, personal assistance, and education. It's better at: • Following instructions • Calling tools • Natural, expressive speech • Understanding cues (like laughs) • Switching languages [video] Rita Kozlov / @ritakozlov_ : we (@cloudflare) created this internal channel late last year, assuming realtime would play a major part in making agents im mind. cool to see this play out ps: might have an announcement here tomorrow — stay tuned! [image] @swyx : Voice is the OG modality. So excited for image inputs, function calling & MCP support in the Realtime API GA! ‘gpt-realtime’ is a lot more natural and expressive, and every time a SOTA voice model is released, you know what I gotta do... Here is the new voice Marin, on KaraokeBench! scores 3 out of 7 again from me Olivia Moore / @omooretweets : I've spent the last two years studying consumer AI trends. Yesterday, our team @a16z published our latest report on the top 100 AI products (by usage). My biggest surprises - and what to learn from them ⬇️ Olivia Moore / @omooretweets : OpenAI just dropped major upgrades to the realtime API - and a new speech model (GPT-realtime) Here's what they announced + what it means for startups 👇 [image] Ryan Carson / @ryancarson : Pricing details for the upgraded gpt-realtime model: Reduced prices for gpt-realtime by 20% compared to gpt-4o-realtime-preview ... $32 / 1M audio input tokens ($0.40 for cached input tokens) $64 / 1M audio output tokens @openai : Introducing gpt-realtime — our best speech-to-speech model for developers, and updates to the Realtime API https://openai.com/live/ LinkedIn: Andriana Lazana : Imagine building voice agents that sound truly human, handling complex queries, switching languages seamlessly, and even integrating images for richer interactions! … Charles Wang : The Realtime API is now GA! This feels like a monumental moment for us at OpenAI and a full-circle moment for me, as it builds on a lot … Kirby Thornton : T-Mobile's AI voice assistant is here. I am so proud to finally be able to share this one. Our multimodal assistant has now begun its roll out to customers to support upgrades in the T-Life app. … Forums: Hacker News : GPT-realtime and Realtime API updates
Context & Ripple Effects
OpenAI’s voice push moved from voice prompting in ChatGPT to GPT-4o’s native multimodality and then dedicated API speech tools, including more expressive text-to-speech and speech-to-text models. This release turns that progression into a generally available, developer-facing real-time stack.
The significance is not simply a stronger voice model: MCP, WebRTC video, and SIP put conversational AI closer to the systems where support, assistance, and education interactions already occur.
First-order effects
- Developers can deploy gpt-realtime on a generally available Realtime API, with improved speech naturalness, instruction following, multilingual switching, and interpretation of conversational cues.
- Teams building voice experiences can connect real-time agents to remote MCP integrations, video input, and telephone infrastructure through SIP, reducing the need to assemble those interfaces separately.
Second-order effects
- Customer-service and assistant vendors can test more capable phone and web-based agents against existing scripted IVR and chat workflows; differentiation shifts toward integration quality, task design, and escalation handling.
- MCP support makes the API more useful in tool-connected applications, increasing pressure on adjacent voice platforms to offer comparable interoperability rather than compete solely on voice quality.
Third-order effects
- If adoption follows, real-time AI will increasingly be bought as a workflow layer spanning voice, video, tools, and telephony—not as a standalone transcription or text-to-speech component.
- General availability creates a more stable base for production deployments, but the lasting market shift depends on whether these agents reliably complete real-world tasks in support, personal-assistance, and education settings.
The trend: This is part of the shift from isolated speech features toward full-duplex, tool-connected AI agents embedded in customer and operational workflows.