Across the AI industry, users are becoming more conscious of just how expensive their deployments can be, and feeling new urgency to cut costs. Open source models offer significantly lower per-token costs — but finding the right model for a given job is difficult.
Writer, which offers AI tools and agents for marketers, has launched a new flagship model, Palmyra X6, aimed at that problem.
Built on an open model
Palmyra X6 is a post-training variation on Z.ai's open source GLM-5.2. Writer says the system delivers deployment-ready capabilities at a much lower price.
The company estimates that the new model, combined with changes to its harness infrastructure, will cut customer costs by as much as 50 percent for basic tasks. The model shipped alongside significant upgrades to Writer's standard agentic harness.
The real claim: the harness
The interesting part of the announcement is not the model but the software around it. A paper from Writer researchers tested small changes in harness efficiency across multiple different models.
The finding: in many cases, changes in the harness were a more reliable way to reduce costs than model choice, with costs falling an average of 40 percent across their testing.
The researchers' formulation captures the point: "The harness is the one component whose efficiency multiplies across every model an organization runs — present and future."
The new approach puts particular emphasis on complex, multi-step tasks executed faster and with fewer tokens.
Why enterprises are tired
CEO May Habib's assessment is blunt: "I think the enterprise is absolutely sick of chasing the next benchmark. They want flattening cost, and it seems like nobody can deliver that."
Habib also sees the push to cut costs feeding a broader distrust toward the major AI labs, which have a financial incentive to drive up token use.
"The cost explosion here is just unprecedented for customers, and so is the degree to which CIOs are giving up on the labs," she said, adding that the labs "don't deeply understand right how to help an enterprise get benefit from AI."
Staying model-agnostic
For Writer's clients the experience remains model-agnostic. Palmyra X6 sits alongside the company's other models, or alongside outside models imported through Azure or Amazon Bedrock.
That choice follows naturally from the harness argument: if the efficiency comes from the layer around the model rather than the model itself, which model runs underneath becomes a secondary question.
Why this argument surfaces now
The cost of an AI deployment is invisible during setup. The bill grows once the system is genuinely used and multi-step agent tasks come into play: every step is another model call, and every call is more tokens.
Habib's "benchmark fatigue" observation grows from the same soil. A model rising two points in a ranking changes nothing on the invoice of the organisation running it. What the organisation measures is not the score but what the same job costs.
The harness argument has a testable side, too: the claim can be checked independently, because running the same model under different harnesses and measuring the cost is entirely possible.
By the numbers
- Base model: Z.ai's open source GLM-5.2, with post-training on top
- Promised saving: up to 50 percent on basic tasks
- Research finding: an average of 40 percent from harness changes alone
- Access: available to Writer clients; runs alongside outside models imported via Azure or Amazon Bedrock