AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A Wagtail team’s month-long effort to use GLM 5.3 Flash for coding reached about half its planned token use: 1 billion of 2 billion tokens went to the target model. The team reported good results from the model, but prototype costs, infrastructure limits and research needs led it to use alternatives; it plans tighter tracking and budgeting for a renewed effort in October.

Wagtail’s month-long coding experiment with GLM 5.3 Flash fell short of its target: the team used the model for about 1 billion of 2 billion tokens, or half the total, before shifting substantial work to other models. In its report, the team said the model performed well, but prototype spending, limited inference-provider capacity and ongoing research needs complicated the plan.

The experiment aimed to use GLM 5.3 Flash for a full month of coding. Wagtail said it managed to use the model exclusively for the first half of the month, keeping that use within its budget. It reported spending $68 on that model during the period, with about 4 kilowatt-hours of energy use and 365 grams of carbon emissions. The second half brought a change: about 1 billion tokens went to other models.

A separate cost came from building an experimental Wagtail MCP server, described by the team as a “vibe-coded prototype.” Wagtail said it chose a model that was not well suited to the task, leading to 450 million tokens being used almost overnight. The team put that episode at $150 and 5 kWh. It said the prototype works and provides a useful demonstration, but estimated similar results could probably have been achieved at about one-fifth the cost with somewhat more effort.

Wagtail also reported performance degradation with GLM 5.3 Flash from its selected inference providers, which it attributed most likely to capacity constraints. The team said it switched to similar alternatives, including DeepSeek V4.1 Flash and Qwen 3.8 Flash. Its comparison focused on open models it could confirm were available in a European data center; the report cautions that model choices and availability can differ for users willing to run inference elsewhere.

At a glance
reportWhen: Report published ahead of a planned Oct…
The developmentWagtail published a report on a month-long coding experiment with GLM 5.3 Flash, detailing why the team used the model for only half of its planned AI token use.

What the Trial Revealed About AI Costs

The report illustrates that choosing a lower-cost model for routine coding does not, by itself, control the total cost of AI-assisted development. Prototyping and experimentation can use large amounts of inference, and Wagtail’s server example shows how model selection and agentic workflows can quickly affect a project’s budget and energy use. The team said it needs to budget for research work separately from day-to-day engineering.

Its experience also points to a practical constraint for developers: provider capacity affects access and performance, even when a model is a good fit technically. Wagtail said switching models was straightforward, but the availability issue disrupted its plan. For teams setting usage or sustainability goals, the report argues for tracking cost and energy alongside tokens, rather than treating token counts alone as a measure of efficiency.

The findings are one organization’s experience, not a controlled comparison establishing that GLM 5.3 Flash will be cheaper or more reliable for every team. Still, Wagtail’s account gives developers concrete figures and operational issues to factor into their own trials.

Amazon

AI coding assistant software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Wagtail Set Its Model Goal

Wagtail framed the exercise as an effort to use one model for ordinary engineering work, while continuing to test other models for research and development. The report says GLM 5.3 Flash was used on Wagtail itself, websites built with the CMS, user-interface tasks, AI research, documentation and evaluations. The team praised its 1-million-token context window, vision support and availability across multiple providers as useful features for varied coding tasks.

The experiment was not designed to eliminate all other model use. Wagtail said it considered broader testing necessary to compare performance on its own tasks and keep pace with new provider releases. It is developing benchmarks and tools, including a CLI prototype, intended to help make leaner model choices more viable for agent-based work. The source presents these as ongoing efforts rather than completed results.

Wagtail characterized the month as a technical failure against its original target, while saying it had learned from the attempt. It reported roughly 35 kWh of total energy use, compared with a planned 10 kWh, and said it wants a future goal focused on having efficient “flash-tier” models handle most day-to-day inference, measured by cost or energy rather than raw token totals.

““The MCP server itself works well and we now have a great demo of the capabilities, so it’s not for nothing.””

— Wagtail team

Amazon

AI model inference provider

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Usage Figures Do Not Show

The report does not provide a complete breakdown of tokens, cost and energy by task, model or provider. It also does not specify the exact dates covered, the accounting method behind its energy and emissions estimates, or the circumstances of the reported provider degradation. Wagtail said limited capacity was the most likely explanation, rather than identifying a confirmed cause.

The cost comparison for the MCP server is an estimate by the team: it said similar results could probably have been reached for about five times less cost. The report does not describe a controlled rerun validating that estimate. Nor does it establish that the same model mix, provider availability or results would apply to other organizations, workloads or locations.

Amazon

energy-efficient AI server hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Wagtail Plans an October Retry

Wagtail said it wants to repeat the effort in October, this time targeting efficient models for more than half of normal day-to-day production work rather than limiting all research and development to one model. Its planned changes include continuous local tracking of usage, spend and energy, and a separate budget for experimentation and prototype work.

The team also plans to make more deliberate decisions about which prototypes to build and to improve its use of agent roles, such as orchestrators, scouts, implementers and reviewers. It said it will keep testing efficient models and techniques while developing benchmarks on Wagtail-specific tasks. The report points to a Wagtail Space 2026 event in November as a place to discuss how the October effort went; it does not provide an event date or publish benchmark results yet.

Amazon

AI development budget tracking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Did Wagtail use GLM 5.3 Flash for the entire month?

No. Wagtail said it used the model exclusively during the first half of the month. Overall, it used GLM 5.3 Flash for about 1 billion of 2 billion tokens, while the second half included substantial use of other models.

How much did the target model use cost?

Wagtail reported that use of GLM 5.3 Flash stayed within its budget of $68, with about 4 kWh of energy use and 365 grams of carbon emissions. These are figures from the team’s report; it does not detail the calculation method.

Why did the team switch to other models?

Wagtail cited performance degradation with its inference providers and the need to keep testing models for research and benchmarks. It said it switched to DeepSeek V4.1 Flash and Qwen 3.8 Flash, while identifying limited provider capacity as the likely explanation for the degradation.

What does Wagtail plan to change for its next experiment?

The team plans to track cost and energy more consistently, budget separately for experimentation, choose prototypes more carefully and refine its use of agent roles. It said its next effort is planned for October, with a focus on efficient models for most routine production work.

Source: hn

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Pro Se Litigants: How Grammarly Makes Legal Writing Easier

A new AI-powered tool aims to help self-represented litigants draft court documents more accurately, verifying citations and formalities to reduce errors and rejections.

xAI’s Grok 4.6 Is Now Available In Amazon Bedrock | Artificial Intelligence – Amazon Web Services (AWS)

AWS customers can access xAI’s Grok 4.6 through Bedrock. Regional availability, platform pricing and independent performance comparisons remain unclear.

So You Want To Use OpenRouter?

A comprehensive guide on using OpenRouter, including current developments, what is confirmed, and what remains uncertain about this emerging platform.