Microsoft Finds Cheaper AI Model Rates Can Still Raise Agent Costs
Microsoft said lower token prices for Claude Sonnet 5 did not remove AI agent cost spikes when it compared Claude models inside GitHub Copilot.

Microsoft’s evaluation of Claude Sonnet 4.6 and Claude Sonnet 5 found that lower token prices did not guarantee lower costs for AI coding agents.
Sonnet 5 was cheaper per token, but it sometimes consumed dramatically more tokens while carrying out the same engineering instructions inside GitHub Copilot Chat in Visual Studio Code on Windows.
The assessment covered 150 agent tasks across 15 technical scenarios.
Engineers ran each scenario five times with each model, then evaluated every run against a binary Select gate and separate quality dimensions.
An LLM judge, calibrated for consistency, scored the results.
Costs came from actual per-turn token usage priced at GitHub Copilot rates.
The tests covered two different developer workloads: Azure architecture design grounded in Microsoft Learn documentation and complex SharePoint Framework project upgrades.
That split produced sharply different results.
Architecture work exposed volatile spending
Sonnet 5’s rate card showed a 33 percent reduction across every token category.
Input tokens cost $2 per million, compared with $3 for Sonnet 4.6.
Cached input fell from $0.30 to $0.20 per million, while output tokens dropped from $15 to $10.
Those listed prices did not predict the amount of work each agent performed.
Across 12 architecture scenarios and 60 runs per model, Sonnet 5 used 12 times more tokens at the median.
One run consumed 47 times the typical baseline token volume.
The higher usage did not always produce a higher bill.
Sonnet 5 averaged $0.47 per architecture run, compared with $0.54 for Sonnet 4.6, because the price reduction outweighed the additional consumption in that workload.
Its spending was less consistent, however.
Most Sonnet 4.6 architecture runs clustered between 14,000 and 45,000 tokens, while Sonnet 5 showed much wider variance.
The difference was accompanied by a decline on one quality measure.
Both models achieved a 75 percent success rate on the Select gate, but Sonnet 4.6 scored 90 percent on the Idiomatic dimension across the nine scenarios where both models produced usable output.
Sonnet 5 scored 78 percent, and the older model matched or exceeded it in eight of those nine comparisons.
An IoT analytics architecture scenario illustrated both gaps.
Sonnet 4.6 passed the idiomatic check in four of five runs; Sonnet 5 passed it once.
Sonnet 5 also used 16,000 tokens in one run and 6.6 million in another on the same baseline prompt.
Code upgrades improved execution but raised the bill
The SharePoint Framework tests reversed the quality and reliability pattern.
Three upgrade scenarios included moving a build system from gulp to Heft and converting a legacy ESLint configuration to flat config.
Sonnet 5 passed the Select gate in 100 percent of the runs, compared with 60 percent for Sonnet 4.6.
The clearest instruction-following difference appeared in an upgrade from SPFx version 1.21.1 to 1.22.0.
Sonnet 4.6 failed all five attempts by overriding the requested version and choosing 1.22.1 based on its Microsoft Learn grounding context.
Sonnet 5 followed the version instruction in every attempt.
That improvement came with substantially higher execution costs.
Token consumption differed by a factor of 10 across the 15 code-upgrade runs per model.
Sonnet 5 cost $2.01 per run, making it 3.7 times more expensive than Sonnet 4.6’s $0.55 median.
One Sonnet 5 run consumed 69 million tokens while conducting extensive web fetching to find undocumented migration steps.
It met 21 of 30 strict evaluation criteria, but four of five runs in each scenario failed to reach that level of analytical depth.
Neither model solved the underlying configuration problem.
Configuration correctness remained at zero percent across all SharePoint Framework scenarios.
The engineering team identified seven missing configuration changes involving build-tool flags, package manifests, deprecated files and configuration formats.
Because those steps were not centralized in the available documentation, changing the model did not resolve the migration gap.




















