Bloomberg/Getty Images
- Amazon's Alexa+ looks for ways to reduce its use of Anthropic's Claude models to lower AI costs.
- Amazon is also trying to use GPUs more efficiently to make Alexa+ cheaper to run.
- Internal forecasts projected Alexa+'s AWS cloud costs would nearly triple this year.
Amazon has redesigned Alexa to rely less on Anthropic models, part of a sweeping effort to lower the cost of running its AI-powered voice assistant, according to internal documents reviewed by Business Insider.
The documents, which span late last year through early this year, show Amazon pursuing a series of changes in how Alexa generates answers by routing more requests to its in-house AI models, avoiding unnecessary calls to Anthropic's Claude models, and squeezing more work from each GPU.
Together, the initiatives were expected to more than quadruple the number of customer transactions each unit of computing capacity could support.
The effort offers a glimpse into AI's next battleground.
As frontier models become more capable, competition is shifting from building smarter AI to making them cheaper to run. Google has promoted lower-cost AI through Gemini Flash, while companies including OpenAI and Cursor have introduced techniques that automatically send simpler requests to lower-cost models.
Amazon's financial projections underscore why the company has devoted so much effort to this challenge.
Internal forecasts from early this year showed AWS cloud costs for the upgraded, AI-powered Alexa+ were on pace to reach roughly $1.7 billion in 2026, nearly triple the previous year.
Alexa+ was also projected to run about 60% above Amazon's target for AWS cloud cost per monthly active user. Even after identifying roughly $450 million in potential savings, internal reviews concluded the business would not hit its financial targets. Amazon declined to comment.
A costly new Alexa
Unlike earlier versions of Alexa, Alexa+ generates many responses with large language models running on GPU-intensive cloud services. That turned relatively inexpensive voice requests into AI workloads that cost far more to serve.
Those costs became more important as Amazon worked through a difficult launch. Business Insider previously reported that the company delayed Alexa+ multiple times as engineers grappled with AI hallucinations and questions about whether the service was ready for customers. Alexa+ expanded its availability in the US earlier this year.
Scaling the service only increased the financial pressure, a sign of how different generative AI is from more traditional software services.
As Alexa+ rolled out to more users, Amazon projected sharply higher AWS cloud spending as demand for AI computing capacity grew.
The company even weighed delaying some of its most expensive AI initiatives. Business Insider previously reported that Project Moonraker, Amazon's effort to give Alexa more advanced AI agent capabilities, was expected to become the service's largest AI expense this year, and the company considered delaying parts of the project as it searched for savings.
Reducing unnecessary calls to Claude
Andrej Sokolow/picture alliance via Getty Images
One of Amazon's priorities was narrowing where Anthropic's Claude models would be used inside Alexa+.
Internal roadmaps called for moving specialized Alexa "Experts" from Claude Sonnet to Amazon's own AI models while reducing other use of Claude across the digital-assistant service.
Amazon also sought to avoid inference whenever possible. Inference is how AI models are run, and one way to limit the cost of this is to use caching, which stores answers to common requests so the AI doesn't have to do the same work again.
One Amazon roadmap called for Alexa+ to stop calling Claude models when suitable answers were already available in cache, and expand "deterministic" handling, which enables Alexa to answer more predictable requests without tapping a large language model.
The strategy is notable given Amazon's deep ties to Anthropic. Amazon has invested billions in the AI startup, partners closely with it, and stands to reap a significant windfall from Anthropic's IPO, if that goes ahead.
Yet the official internal documents reviewed by Business Insider show Amazon has been looking for ways to reduce how often Alexa relies on Anthropic's models.
Amazon's approach mirrors a growing trend across the AI industry. Investment firm William Blair wrote in a recent report that software companies are starting to reserve frontier models for difficult, high-stakes reasoning while routing less complex requests to cheaper models. That lowers inference costs without changing the customer experience.
"Multi-model routing is becoming standard architecture in software," analysts at William Blair wrote in the report.
Delivering more with fewer GPUs
Reducing model costs was only one part of the strategy. Amazon also focused on increasing how much work each GPU could perform.
Rather than simply adding more Nvidia GPUs, Amazon wanted to process more customer requests from the same computing gear. One roadmap projected software upgrades would increase available computing capacity by roughly 50% while cutting response times by about 40%. Internal planning dashboards tracked projected customer growth, GPU utilization, available capacity and inference efficiency as Amazon prepared to scale Alexa+.
Amazon's cost-saving efforts extended beyond software. Planning documents show the company evaluating both Nvidia GPUs and its own Trainium chips to further lower the cost of running Alexa+.
More broadly, the documents show Amazon treating frontier AI models and GPU capacity as expensive resources to be deployed selectively rather than by default.
That philosophy echoes a point CEO Andy Jassy has made publicly. In his shareholder letter last year, Jassy argued there's an "urgency" to make AI inference dramatically less expensive.
"Reducing the cost per unit in AI will unleash AI being used as expansively as customers desire, and also lead to more overall AI spending," Jassy wrote.
Have a tip? Contact this reporter via email at ekim@businessinsider.com or Signal, Telegram, or WhatsApp at 650-942-3061. Use a personal email address, a nonwork WiFi network, and a nonwork device; here's our guide to sharing information securely.
Read the original article on Business Insider