A look at startups like WellSaid Labs and VocaliD, which are building custom AI voice actors for digital assistants, video game characters, and corporate videos
A new wave of startups are using deep learning to build synthetic voice actors for digital assistants, video-game characters, and corporate videos. Tweets: @kathyreid , @eagerbeavertech , @martinwaxman , @ttallon , @eagerbeavertech , @erikphoel , @eagerbeavertech , @resembleai , @_karenhao , @paulnemitz , @_karenhao , @vocalidinc , and @mattnavarra Tweets: @kathyreid : Another excellent piece by @_KarenHao for @techreview about #TTS - or voice synthesis, and its repercussions - both positive and negative. Interesting to see the new business models. #AI #voice actors sound more human than ever—and they're ready to hire https://www.technologyreview.com/ ... Eager Beaver / @eagerbeavertech : https://www.technologyreview.com/ ... Each one is based on a real voice actor, whose likeness has been preserved using AI. Martin Waxman / @martinwaxman : #AI voices are sounding a lot more like people. Companies can now use a synthetic brand voice for customer service, videos and training. But the AI still needs human actors for training data. And compensation is an issue. cc @alexsevigny #AIinPR https://www.technologyreview.com/ ... Tina Tallon, Ph.D. / @ttallon : “If they're not afraid of being automated away by AI, [voice actors are] worried about being compensated unfairly or losing control over their voices...” A great overview of current concerns in AI speech synthesis from @techreview: https://www.technologyreview.com/ ... Eager Beaver / @eagerbeavertech : https://www.technologyreview.com/ ... Whereas companies used to have to hire different voice actors for different markets-the Northeast versus Southern US, or France versus Mexico-some voice AI firms can manipulate the accent or switch the language of a single voice in different ways. Erik Hoel / @erikphoel : Goodbye voice actors https://www.technologyreview.com/ ... Eager Beaver / @eagerbeavertech : https://www.technologyreview.com/ ... Unlike a recording of a human voice actor, synthetic voices can also update their script in real time, opening up new opportunities to personalize advertising. @resembleai : At @resembleai, we are using deep learning to build synthetic voices for various use cases, such as: digital assistants, video-game characters, and virtual call center agents etc. Great overview of the landscape done by MIT Technology Review @techreview -https://www.technologyreview.com/ ... Karen Hao / @_karenhao : Each of the voices I've embedded in my piece represent impressive advancements in deep learning. But the rise of hyperrealistic fake voices has left human voice actors to wonder what this means for their livelihoods. https://www.technologyreview.com/ ... Paul Nemitz / @paulnemitz : That's why in the #AIA #ArtificailIntelligenceAct we need a comprehensive duty to signal that a machine is speaking: #AI #voice actors sound more human than ever—and are ready to hire | MIT Technology Review #PrinzipMensch https://www.technologyreview.com/ ... Karen Hao / @_karenhao : Can you tell that this voice isn't human? It's actually an AI voice actor made by @wellsaidlabs. The startup is one of many that now offer such voices to hire for digital assistants, corporate videos, & even video-game characters. https://www.technologyreview.com/ ... https://soundcloud.com/... VocaliD / @vocalidinc : “... for VocaliD's @TweetRupal the point of AI voices is ultimately not to replicate human performance or to automate away existing voice-over work. Instead, the promise is that they could open up entirely new possibilities.”-by @_KarenHao @techreview https://www.technologyreview.com/ ... Matt Navarra / @mattnavarra : AI voice actors sound more human than ever—and they're ready to hire Bad news for voice over actors https://www.technologyreview.com/ ...
Context & Ripple Effects
This 2021 profile captures the first commercial phase of synthetic speech: WellSaid Labs, VocaliD, and Resemble AI building bespoke AI voice actors sold into digital assistants, games, and corporate video — an extension of the corporate-friendly deepfake market Wired had already spotted at Synthesia for multilingual training content. The business question then was licensing and new revenue models, as the Twitter reaction to Karen Hao's piece notes.
The arc since has moved fast in both directions: toward consent-based likeness markets — Tel Aviv's Hour One paying people to lend their faces and voices to AI characters weeks later — and toward trivially cheap cloning, culminating in OpenAI's Voice Engine generating a voice copy from a 15-second sample. The fraud side surfaced in between, when a WSJ columnist's ElevenLabs-and-Synthesia-built clone defeated her own bank's voice biometrics.
First-order effects
- Studios, assistant makers, and corporate video teams gain a new supply channel for voice work, while human voice actors — per the reporting — push back on unfair compensation and losing control over their own voices.
- WellSaid Labs, VocaliD, and Resemble AI turn TTS from a utility feature into a product category: named, licensable synthetic performers rather than generic robotic output.
Second-order effects
- Hour One's paid-likeness model shows the knock-on labor structure this creates — people monetizing their voices and faces as recurring digital assets instead of one-off session fees.
- As cloning gets cheap enough that OpenAI's Voice Engine needs only a 15-second sample, banks and other voice-biometric users are forced to treat synthetic speech as an authentication threat, not just a content tool — the exact failure mode demonstrated by the WSJ columnist's bank-defeating clone.
Third-order effects
- If hyperrealistic synthesis keeps spreading, disclosure becomes the regulatory battleground: Paul Nemitz is already urging that the proposed EU AI Act include a duty to signal when a machine is speaking, which would make provenance labeling a compliance requirement across every voice-first product.
- The industry splits structurally between consented, licensed voice libraries (VocaliD-style marketplaces) and uncontrolled scraping — with payment norms for voice talent set by whichever side regulators and platforms favor.
The trend: Synthetic voice is moving from bespoke licensed studio work to instant low-sample cloning, with consent markets, voice-biometric defenses, and machine-disclosure regulation racing to catch up.