Amazon's “Speech Synthesis Markup Language” now lets Alexa developers add personality to Skills' responses by controlling pronunciation, pitch, timing, and more
Amazon's Alexa is going to sound more human. The company announced this week the addition of a new set …
Context & Ripple Effects
This 2017 SSML update is the earliest move in a nine-year arc of Amazon teaching Alexa to sound less like a machine. At the time, the Alexa platform was still young enough that adding Skills by voice counted as a headline feature, and developer control over pronunciation, pitch, and timing gave Skill builders their first real lever over how the assistant felt rather than just what it said.
First-order effects
- Skill developers gain programmatic control of pronunciation, pitch, and timing, so third-party responses can carry distinct vocal character instead of Alexa's single default delivery.
Second-order effects
- Amazon keeps compounding this surface for developers: eight free Polly voices arrive in 2018, happy/excited and disappointed/empathetic tones in 2019, and a long-form news and music speaking style with conversational pauses in 2020.
Third-order effects
- By the time personality styles including an adult-only "Sassy" option reach Alexa+, expressive voice control has become a product-line strategy rather than a developer nicety — raising governance questions about how far synthetic personalities should go.
The trend: Voice assistants are evolving from fixed synthetic voices toward developer- and user-selectable personalities, with each expressive capability layer adding both engagement upside and content-boundary risk.