Writer Says Harness Changes Cut AI Costs 40% Across Multi-Model Tests
Writer says small changes to its agentic harness cut costs by an average of 40% across multi-model tests, while Palmyra X6 and the broader upgrade could reduce basic-task costs by as much as 50%.

Writer is betting that enterprise AI costs can fall in the orchestration layer before customers replace a model.
TechCrunch reported that the company launched Palmyra X6 and an upgraded agentic harness that Writer says can reduce basic-task costs by as much as 50%.
Palmyra X6 is a post-training variation on Z.ai's open-source GLM-5.2.
It joins Writer's existing models and models imported through Azure or Amazon Bedrock, keeping customers inside the same operating environment while they route different tasks across providers.
The harness is the common layer between those models and the work they perform.
It selects a model for each task, coordinates multi-step jobs and controls token use, giving customers a cost lever that does not depend on migrating every workload to one new model.
Both parts of the release became available on Thursday.
Across tests involving multiple models, Writer researchers found that small harness-efficiency changes reduced costs by an average of 40% and, in many cases, produced more consistent savings than model choice alone.
The architecture also separates model selection from task orchestration.
A workload can move between an in-house Writer model and an outside model without changing the layer that manages steps, context and token consumption.
That makes the harness upgrade relevant to customers already running mixed model portfolios, not only buyers willing to standardize on Palmyra X6.
In comments to TechCrunch, chief executive May Habib described customers' demand for flatter token costs and rising frustration with major AI labs whose business models benefit when token use grows.
Enterprise buyers, by contrast, are trying to make complex, multi-step work faster and less expensive.
Writer's launch therefore gives a CIO two separate decisions: whether Palmyra X6 fits a workload and whether the harness can lower bills across models already in production.
The release will be measured by whether its routing and token controls reproduce the reported savings on real enterprise tasks without forcing another model migration.




















