31 de julio de 2026

Leaked Amazon Documents Detail $1.8 Million Overrun on a Single Claude AI Task Missed for Five Months

Internal Amazon documents released this week reveal several AI cost overruns. The largest involved a $1.8 million project using Anthropic's Claude Sonnet model, which exceeded its budget by 860% and was never launched.

Amazon did not detect the overspending for five months. Senior engineers reported the overruns at a staff meeting on Tuesday, July 28. The tool was designed to automatically match author details to product listings on Amazon's site.

Two additional projects also exceeded their budgets during this period, bringing total unplanned AI spending to approximately $2.5 million.

Largest Amazon AI Overrun Exposed in Internal Documents

The Claude Sonnet author-matching tool accounted for the largest overrun at $1.8 million, exceeding its budget by 860% and never launching.

A financial auditing tool exceeded its budget by approximately $541,000, while a logistics system designed to accelerate deliveries ran $134,000 over budget.

Together, the three projects resulted in approximately $2.5 million in unplanned AI spending. This figure is small relative to Amazon's scale, citing $181.5 billion in revenue and $23.9 billion in operating income for the first quarter of 2026.

Engineers attributed the overruns to a shift in AI provider billing models. Many providers now bill based on tokens, or units of processed text, rather than flat monthly subscriptions.

This can cause costs to rise quickly if tasks generate more computing activity than expected. One presentation described coding mistakes that were once inexpensive to fix as "catastrophically expensive" once AI agents took over the work.

Amazon responded in a statement. "As with any new technology, we're experimenting, learning and improving how we use it, including how we drive cost efficiencies," the company said.

"Cherry-picking small, isolated examples where teams are learning from one another and portraying them as business as usual doesn't reflect how teams across Amazon are using AI."Previous AI Tool Outages, KiroRank Leaderboard, and "Tokenmaxxing""

The overruns followed two earlier outages involving Amazon's AI coding tools. In mid-December 2025, AWS engineers authorized the Kiro coding assistant to make changes to Cost Explorer, an internal system for reviewing AWS billing.

The agent deleted and recreated its environment, resulting in a 13-hour outage that affected AWS customers in parts of mainland China. This was the second outage involving one of the company's AI tools within months, following an earlier disruption connected to Amazon's Q Developer chatbot.

In both cases, the AI tools carried the same permissions as the human engineers operating them, and the changes went through without the second-person approval normally required.

Amazon said the December incident was a user access control issue rather than an AI autonomy issue, and that the same disruption could have occurred with any developer tool or through manual action.

Amazon ran an internal dashboard called KiroRank that tracked how many tokens individual employees used on the company's Kiro and MeshClaw platforms.

The company had set a goal for more than 80 percent of its developers to use AI tools every week. According to the documents, several employees began assigning AI agents to unnecessary tasks specifically to raise their standing on the leaderboard, a practice workers called tokenmaxxing.

Amazon senior vice president Dave Treadwell addressed the leaderboard in a message to staff. "Please don't use AI just for the sake of using AI," he wrote.

"Use AI to help you solve customer problems, to help you solve business problems, to innovate." Amazon has since discontinued KiroRank, stating it was built informally by employees and was never intended to reward AI use for its own sake.

What Amazon’s AI Cost Overruns Mean for Other Companies

BetaNews reports that similar patterns have surfaced at other companies using the same AI models. Uber's chief technology officer, Praveen Neppalli Naga, said in April that Uber had already used up its entire 2026 budget for Claude Code.

Uber's president and chief operating officer, Andrew Macdonald, said afterward that the company had not found a clear connection between rising AI token usage and the number of useful features it produced for customers, telling colleagues that relationship had not yet been established.

For organizations deploying token-billed AI agents, these disclosures highlight several controls to consider:

  1. Set strict budget caps and automated spend alerts for token-billed AI tasks, as Amazon did not detect the $1.8 million overrun for five months.
  2. Require secondary approval for AI agents with the same permissions as human engineers, especially for actions that could affect production systems.
  3. Avoid internal leaderboards that reward token consumption, as KiroRank reportedly encouraged tokenmaxxing behavior.
  4. Monitor whether increased token usage leads to useful output, a relationship Uber's COO said his company had not established.
  5. Limit agent permissions rather than granting environment-level delete and recreate access, which led to the 13-hour Cost Explorer outage.

Amazon characterized the overruns as isolated examples of teams learning from one another rather than standard practice, and did not confirm whether additional projects exceeded budgets during the same period.

The documents were presented internally on July 28 and made public this week, but Amazon has not released the full figures or said what changes it is making to budget monitoring for token-billed AI projects beyond discontinuing KiroRank.

Thank you for being a Ghacks reader. The post Leaked Amazon Documents Detail $1.8 Million Overrun on a Single Claude AI Task Missed for Five Months appeared first on gHacks.



☞ El artículo completo original de Arthur Kay lo puedes ver aquí

No hay comentarios.:

Publicar un comentario