How Apple added support for automatic transcript generation in the Podcast app, a feature highly requested by both disabled users and podcast creators
As seen in MacRumors, iOS 18 beta testers can set a new Siri wake … X: Dave Winer / @davewiner : @johnspurlock John, I just heard about transcriptions a couple of days ago. Do I have to publish my podcast through Apple in some way in order for my feed to have these transcriptions? Ari Saperstein / @ari_saperstein : I reported an in-depth look at accessibility & podcasts for @GuardianUS. Through talking with experts, users with disabilities and audio platforms themselves, I looked at the history of transcription in audio (or, rather, lack thereof) & Apple's new game-changing initiative: @jamescridland : @davewiner @johnspurlock Important to note that you can submit your own transcriptions via the podcast namespace in RSS, which Apple then ingests. That is the only way to do multi-speaker transcripts (since Apple's systems don't know who the voices are). Those RSS transcripts also used by other apps. @jamescridland : @davewiner @johnspurlock Both Morning Coffee Notes and Trade Secrets have auto transcripts in the RSS feed; but both fall back to Apple's transcripts since they're better than the ones I roughly produced. But we have the power to overwrite them! @zeldman : ‘"I was knocked out on how accurate it was," says Larry Goldberg, a media and technology accessibility pioneer who created the first closed captioning system for movie theaters.’ https://www.theguardian.com/ ... #a11y LinkedIn: Christopher Patnoe : This is important and useful for so I want people to know about it. — AND I also want to take this opportunity to remind that for years Android … See also Mediagazer
Context & Ripple Effects
Apple’s podcast transcript work moved from an iOS 17.4 beta feature that generated text after episodes were published to a broader accessibility-focused implementation. That progression sits alongside Apple’s earlier accessibility releases, including new cognitive, vision, and speech accessibility tools.
The key implementation detail is that Apple can ingest publisher transcripts through the podcast namespace in RSS, while retaining the ability to use its own version when it judges that transcript better. This makes the feature both an accessibility layer and a contested piece of podcast metadata.
First-order effects
- Listeners gain text access to podcast episodes through Apple’s app, while creators can submit RSS transcripts rather than rely solely on Apple’s automated output.
- Publishers of multi-speaker shows face a practical constraint: Apple’s automatic system does not identify speakers, so speaker-attributed transcripts must be supplied by the publisher.
Second-order effects
- Transcript production becomes a more important publishing workflow for creators seeking control over accuracy, speaker labels, and the version shown to listeners; Apple’s earlier automatic transcript rollout established the default option they now can override.
- Podcast platforms and hosting tools have an incentive to support the podcast-namespace transcript format, since RSS submission is the route through which publishers can provide a preferred transcript to Apple.
Third-order effects
- If major listening apps treat transcripts as first-class feed data, podcast distribution may evolve from audio-only delivery toward structured, machine-readable episode publishing.
- The split between platform-generated and publisher-supplied text may make transcript provenance and editorial control durable issues, particularly where automated text is less useful for complex conversations.
The trend: Podcast platforms are turning speech-to-text from an optional creator service into core accessibility and metadata infrastructure.