In a significant upgrade to Amazon’s voice assistant, Alexa is poised to become more naturally expressive and context-aware, thanks to a generative AI-powered enhancement. This transformative update includes the ability for Alexa to continue conversations seamlessly without the need for the wakeword “Alexa.” Additionally, Amazon has introduced an improved “speech-to-speech” engine that pays close attention to the user’s emotions and the tone of their voice, enabling Alexa to respond with matching emotional variations in its output.
During a demonstration, Amazon showcased the new voice, which presents a less robotic-sounding Alexa and exhibits greater expressiveness. This enhancement is made possible by leveraging large transformers trained on various languages and accents. For instance, if a customer inquired about their favorite sports team and they had just won a game, Alexa would respond with a joyful tone. Conversely, if the team had lost, Alexa would express empathy.
Rohit Prasad, Senior Vice President of Alexa, explained, “And we’re working on a new model — which we refer to as speech-to-speech — again powered by massive transformers. Instead of first converting a customer’s audio request into text using speech recognition, and then using an LLM to generate a text response or an action, and then text-to-speech to produce audio back — this new model will unify these tasks, creating a much richer conversational experience.”
Amazon has harnessed its Large Text-to-Speech (LTTS) and Speech-to-Speech (S2S) technologies to make this possible. LTTS allows Alexa to adapt its responses based on textual input, such as user requests or discussion topics, while S2S integrates audio input with text, enabling Alexa to offer more conversational richness, including attributes like laughter, surprise, and encouraging cues like “uh-huh” to keep the conversation flowing smoothly.


