SendTech Times
News
SYSTEMS SHIFT:

Microsoft Finds Cheaper AI Model Rates Can Still Raise Agent Costs

Newsroom brief

Microsoft said lower token prices for Claude Sonnet 5 did not remove AI agent cost spikes when it compared Claude models inside GitHub Copilot.

Verified against source materialEdited by SendTech Times AI & Enterprise DeskSource: Developer Tech
Microsoft Finds Cheaper AI Model Rates Can Still Raise Agent Costs

Microsoft’s evaluation of Claude Sonnet 4.6 and Claude Sonnet 5 found that lower token prices did not guarantee lower costs for AI coding agents.

Sonnet 5 was cheaper per token, but it sometimes consumed dramatically more tokens while carrying out the same engineering instructions inside GitHub Copilot Chat in Visual Studio Code on Windows.

The assessment covered 150 agent tasks across 15 technical scenarios.

Engineers ran each scenario five times with each model, then evaluated every run against a binary Select gate and separate quality dimensions.

An LLM judge, calibrated for consistency, scored the results.

Costs came from actual per-turn token usage priced at GitHub Copilot rates.

The tests covered two different developer workloads: Azure architecture design grounded in Microsoft Learn documentation and complex SharePoint Framework project upgrades.

That split produced sharply different results.

Architecture work exposed volatile spending

Sonnet 5’s rate card showed a 33 percent reduction across every token category.

Input tokens cost $2 per million, compared with $3 for Sonnet 4.6.

Cached input fell from $0.30 to $0.20 per million, while output tokens dropped from $15 to $10.

Those listed prices did not predict the amount of work each agent performed.

Across 12 architecture scenarios and 60 runs per model, Sonnet 5 used 12 times more tokens at the median.

One run consumed 47 times the typical baseline token volume.

The higher usage did not always produce a higher bill.

Sonnet 5 averaged $0.47 per architecture run, compared with $0.54 for Sonnet 4.6, because the price reduction outweighed the additional consumption in that workload.

Its spending was less consistent, however.

Most Sonnet 4.6 architecture runs clustered between 14,000 and 45,000 tokens, while Sonnet 5 showed much wider variance.

The difference was accompanied by a decline on one quality measure.

Both models achieved a 75 percent success rate on the Select gate, but Sonnet 4.6 scored 90 percent on the Idiomatic dimension across the nine scenarios where both models produced usable output.

Sonnet 5 scored 78 percent, and the older model matched or exceeded it in eight of those nine comparisons.

An IoT analytics architecture scenario illustrated both gaps.

Sonnet 4.6 passed the idiomatic check in four of five runs; Sonnet 5 passed it once.

Sonnet 5 also used 16,000 tokens in one run and 6.6 million in another on the same baseline prompt.

Code upgrades improved execution but raised the bill

The SharePoint Framework tests reversed the quality and reliability pattern.

Three upgrade scenarios included moving a build system from gulp to Heft and converting a legacy ESLint configuration to flat config.

Sonnet 5 passed the Select gate in 100 percent of the runs, compared with 60 percent for Sonnet 4.6.

The clearest instruction-following difference appeared in an upgrade from SPFx version 1.21.1 to 1.22.0.

Sonnet 4.6 failed all five attempts by overriding the requested version and choosing 1.22.1 based on its Microsoft Learn grounding context.

Sonnet 5 followed the version instruction in every attempt.

That improvement came with substantially higher execution costs.

Token consumption differed by a factor of 10 across the 15 code-upgrade runs per model.

Sonnet 5 cost $2.01 per run, making it 3.7 times more expensive than Sonnet 4.6’s $0.55 median.

One Sonnet 5 run consumed 69 million tokens while conducting extensive web fetching to find undocumented migration steps.

It met 21 of 30 strict evaluation criteria, but four of five runs in each scenario failed to reach that level of analytical depth.

Neither model solved the underlying configuration problem.

Configuration correctness remained at zero percent across all SharePoint Framework scenarios.

The engineering team identified seven missing configuration changes involving build-tool flags, package manifests, deprecated files and configuration formats.

Because those steps were not centralized in the available documentation, changing the model did not resolve the migration gap.

Share this article
inXf

Related articles

More
Microsoft Puts Agentic Cloud Ops Behind Azure Copilot And FinOps Tools
AI

Microsoft Puts Agentic Cloud Ops Behind Azure Copilot And FinOps Tools

Microsoft said Azure Copilot observability agent is generally available and Azure Resource Manager MCP Server is in public preview, tying agentic cloud operations to governance, cost visibility and human approval.

Sapiom Raises $35M As AI Agent Costs Face First Hard Audit
AI

Sapiom Raises $35M As AI Agent Costs Face First Hard Audit

TNW reported that Sapiom raised a $35 million Series A for software that routes AI-agent calls to lower-cost models and tools, turning agent deployment from a capability race into a budget-control problem.

Microsoft Uses Build 2026 to Push Agents Beyond Copilot
AI

Microsoft Uses Build 2026 to Push Agents Beyond Copilot

Microsoft used its Build 2026 keynote to introduce MAI models, Project Soltera and Microsoft Scout as part of a broader agent strategy. MAI-Thinking-1 is described as a 35-billion-parameter reasoning model with a 128,000-context window for multi-step instructions, long-context reasoning and code generation. The announcement gives Microsoft a clearer agent roadmap, but the source does not provide customer rollout data, pricing or enterprise adoption evidence.

Microsoft Positions Its AI Stack Against OpenAI And Anthropic
AI

Microsoft Positions Its AI Stack Against OpenAI And Anthropic

TechCrunch reported that Satya Nadella used Microsoft's latest analyst call to argue enterprises should keep AI harnesses separate from any single model provider, as the company sells Copilot agents, MAI models and Maya chips alongside its OpenAI and Anthropic relationships.

1Password Gives Claude Session-Scoped Logins Without Model Access
AI

1Password Gives Claude Session-Scoped Logins Without Model Access

SiliconANGLE wrote that 1Password launched a Claude browser integration that lets Anthropic’s agent use approved logins without exposing passwords or one-time codes to the model. The release starts on Mac and leaves payment-card support, personal-detail autofill and independent security validation outside the launch record.

OpenAI Keeps GPT-Red Attack Model Private After Prompt-Injection Tests
AI

OpenAI Keeps GPT-Red Attack Model Private After Prompt-Injection Tests

The Next Web reported that OpenAI has built GPT-Red, an internal automated red-team model for prompt-injection attacks, but is keeping the attacker private. The report cited attack success rates above 90% against an older GPT-5 and below 23% against GPT-5.6, while noting that human testers still catch cases GPT-Red misses.

Sakana Fugu Pitches Model Orchestration As Export-Control Hedge
AI

Sakana Fugu Pitches Model Orchestration As Export-Control Hedge

Sakana AI launched Fugu and Fugu Ultra as generally available orchestration models, but its benchmark and beta claims still depend on enterprise users accepting a routed multi-model system.

Niteshift Targets Enterprise AI Coding With A Model-Neutral Infrastructure Layer
AI

Niteshift Targets Enterprise AI Coding With A Model-Neutral Infrastructure Layer

Niteshift has raised $7 million to build an AI coding cloud that routes across models, pitching enterprise buyers on control, verification and lower dependence on frontier AI labs.

Keep Reading

More Stories

Latest
India and Saudi Arabia Put Port Investment Talks on Strategic Maritime RouteEconomyOct 9, 2026India and Saudi Arabia Put Port Investment Talks on Strategic Maritime RouteIndia and Saudi Arabia discussed possible Saudi participation in Vadhavan and Galathea Bay port projects, linking Gulf trade, container transshipment and maritime capacity building.MIT Expands STEM Workforce Training Through AI And Design ProgramsAIOct 9, 2026MIT Expands STEM Workforce Training Through AI And Design ProgramsMIT for America will scale existing university education programs nationwide, combining calculus support, responsible AI resources and hands-on fabrication to address STEM workforce access gaps.Google Maps Expands Restaurant Search Into Food OrderingCapital & PolicyOct 9, 2026Google Maps Expands Restaurant Search Into Food OrderingGoogle Maps is adding more dining workflow features, with Gemini-powered Ask Maps ordering links through Toast, Square and Uber Eats plus city trend lists and practical restaurant review signals.Lumio Raises $12 Million To Expand Its Consumer-Electronics PortfolioChips & SemiconductorsOct 9, 2026Lumio Raises $12 Million To Expand Its Consumer-Electronics PortfolioLumio raised a $12 million Series A led by Blume Ventures to add products, software, offline retail and support capacity for its smart TV, projector and speaker business.Apple Touchscreen MacBook Report Points To Late-October Hardware SlateChips & SemiconductorsOct 9, 2026Apple Touchscreen MacBook Report Points To Late-October Hardware SlateApple is reportedly preparing a late-October touchscreen MacBook debut alongside an OLED iPad Mini and M6 updates for the entry-level MacBook Pro and iMac.Australia Health Data Modernisation Reaches $32.8 Million on Google CloudCloud & Data CentersOct 9, 2026Australia Health Data Modernisation Reaches $32.8 Million on Google CloudAustralia’s Health department has lifted disclosed spending on a Google Cloud data and analytics modernisation program to $32.8 million across Accenture, NTT, Deloitte and Google contracts.Kioxia LD4 E1.L SSD Targets Dense 1U Hyperscale StorageCloud & Data CentersOct 9, 2026Kioxia LD4 E1.L SSD Targets Dense 1U Hyperscale StorageKioxia’s LD4 Series brings BiCS FLASH generation 8 QLC NAND into the E1.L form factor, starting with 15.36TB and 30.72TB models for read-intensive hyperscale servers while the architecture is validated up to 122.88TB.NeoFleet Raises $4 Million to Expand Taxi-Fleet FinancingReal EstateOct 9, 2026NeoFleet Raises $4 Million to Expand Taxi-Fleet FinancingNeoFleet Capital secured a $4 million pre-seed round to scale a taxi-fleet financing model that combines vehicle credit, telematics, insurance and operating controls across Africa and other emerging markets.Brahma AI Raises US$150 Million Toward Digital-Human LaunchAIOct 8, 2026Brahma AI Raises US$150 Million Toward Digital-Human LaunchBrahma AI raised US$150 million through preferred shares, including US$100 million from Multiples Alternate Asset Management, to expand an enterprise AI content platform and prepare interactive digital humans.OKX Draws Circle and Ripple Backing at $25 Billion ValuationFintech & Digital PaymentsOct 8, 2026OKX Draws Circle and Ripple Backing at $25 Billion ValuationOKX took undisclosed new investment from Circle, Ripple, Standard Chartered’s venture arm and Qube Research as it pushes beyond exchange trading into tokenized shares, stablecoins and consumer finance.Microsoft Shows Windows AI Agents With Nvidia-Powered Surface Laptop UltraAIOct 8, 2026Microsoft Shows Windows AI Agents With Nvidia-Powered Surface Laptop UltraMicrosoft used its Windows and Surface event to show how Copilot agents could work with local files while introducing a premium Surface Laptop Ultra built on Nvidia’s RTX Spark chip.Cybersecurity buyers use 39 September deals to fill AI and OT gapsCybersecurityOct 8, 2026Cybersecurity buyers use 39 September deals to fill AI and OT gapsSecurityWeek counted 39 cybersecurity M&A deals in September, with buyers using acquisitions to add OT visibility, AI-security controls, offensive-testing scale, sovereign-technology work and compliance reach.