Researchers say they can embed audio commands to Siri, Alexa, and Google's Assistant in music and spoken text in such a way that humans can't detect it
Researchers can now send secret audio instructions undetectable to the human ear to Apple's Siri, Amazon's Alexa and Google's Assistant.
Context & Ripple Effects
This finding extends a line of research that began with cheaper hardware: researchers had already shown in 2017 that Siri and Alexa respond to ultrasonic commands generated by about $3 of equipment, and would later show that lasers can trigger the same assistants from a distance. The new work closes the gap between those demonstrations and everyday media — the attack vector is now ordinary music and spoken text, not specialized transmitters.
The disclosure lands while all three companies are under separate scrutiny over what their assistants record and who hears it, including the 1,000+ Google Assistant clips a Belgian broadcaster obtained from a contractor, some apparently captured inadvertently. Together they frame voice assistants as surfaces that both listen too much and obey too readily.
First-order effects
- Apple, Amazon, and Google face immediate pressure to distinguish human-audible speech from machine-decodable commands in their wake-word and speech-recognition pipelines, since their devices currently act on audio their owners cannot hear.
- Owners of always-listening speakers, phones, and tablets are exposed to commands embedded in content they choose to play — a video, song, or podcast becomes a potential remote control.
Second-order effects
- The three vendors will likely be pushed toward countermeasures that trade convenience for safety, such as requiring on-device confirmation or speaker authentication before sensitive actions, which changes the product experience every assistant user gets.
- Content platforms and advertisers gain a new liability question: audio distributed through their channels could carry instructions to listeners' devices, making media distribution a security surface as well as a content business.
Third-order effects
- If covert-command research keeps escalating from ultrasonics to lasers to embedded media, voice interfaces may converge on cryptographic or biometric verification of who — or what — is speaking, restructuring how assistants decide to trust an instruction.
- Regulators already examining assistant recordings for privacy violations have a second hook: always-on microphones that execute inaudible commands give data-protection authorities a security rationale alongside the privacy one.
The trend: Voice assistants are becoming attack surfaces that researchers can reach through ever-cheaper physical channels — ultrasound, light, now ordinary audio — forcing trust and verification into the core of ambient computing.