Twitter says it is expanding voice tweets to more iOS users and plans to add transcriptions to voice tweets to improve accessibility
Voice tweets will come to Android and the web in 2021 — Twitter has just expanded its voice tweets feature, which lets you record a snippet of audio to include with a tweet, to more users on iOS.
Context & Ripple Effects
Twitter introduced voice tweets in June 2020 as a 140-second audio clip attached to a tweet, rolling out to iOS users first (the original launch). This update widens that iOS rollout and adds a commitment the launch lacked: transcriptions to make voice tweets accessible, with Android and web versions promised for 2021.
The move sits alongside a second audio track Twitter opened days earlier — audio DMs, beginning tests in Brazil — showing the company pushing recorded audio into both its public timeline and private messaging at once.
First-order effects
- More iOS users gain the ability to record 140-second audio tweets immediately, while deaf and hard-of-hearing users get a stated path to usable voice tweets once transcriptions ship.
- Android and web users remain locked out until 2021, leaving the feature's audience skewed toward one platform.
Second-order effects
- Twitter's audio DM experiments in Brazil — and later India and Japan — give it two parallel audio surfaces to iterate on, so lessons from private voice messages can feed back into public voice tweets before the Android/web launch.
- Committing publicly to transcriptions raises the bar for how quickly the feature must become accessible; the eventual payoff came when Twitter shipped auto-captions for voice tweets in 11 languages more than a year after the iOS debut.
Third-order effects
- If the pattern holds, speech-to-text stops being an optional add-on and becomes baseline infrastructure for any audio feature a consumer platform ships — the year-long gap between promise and delivered captions shows how heavy that lift is.
- Audio expanding across both tweets and DMs points toward Twitter treating voice as a first-class content format rather than a novelty, which reshapes moderation, discovery, and accessibility work across the product.
The trend: Consumer social platforms are building recorded audio into their core text products, with automatic transcription emerging as the accessibility layer those features are expected to carry.