Wednesday, September 9, 2026
venfeedSubscribe

Long agentic tasks may carry 10,000 times the energy footprint of a simple query

A Vals AI study found some models used the equivalent of 2.5 hours of home electricity to build a single web app. The comparison everyone quotes is the wrong one.

Venfeed Editor2 min read
ShareXBlueskyLinkedInHNRedditEmail

A study by Vals AI estimates that long agentic tasks carry roughly 10,000 times the energy footprint of a simple query, with web app building on some models matching about 2.5 hours of household electricity use, according to Bloomberg.

The figure reframes a debate that has been conducted almost entirely on the wrong terms. Public argument about AI energy use has centred on the cost of a single chatbot query — a number that is small, frequently cited, and increasingly irrelevant to how the technology is actually used.

Why the multiplier is so large

A chat response is one forward pass over a few thousand tokens. An agentic task is thousands of model calls, executed serially, each carrying the accumulated context of everything before it.

Three factors compound. Volume: an agent working for an hour makes far more calls than a person having a conversation. Context: each call carries a growing history, and attention cost rises faster than linearly with context length. Exploration: agents try approaches, fail, and retry, and the discarded work consumes the same energy as the work that succeeds.

A 10,000-fold difference between a query and a long agentic run is therefore not surprising. It is what happens when the unit of work changes from a sentence to a project.

The industry is moving decisively towards the expensive case

Every significant product decision this fortnight increases the share of inference that looks like the expensive case rather than the cheap one.

OpenAI's Astra coordinates 16 agents. Cognition raised $2 billion at a $48 billion valuation selling an autonomous coding agent, with cash burn that could reach $800 million this year — a figure that is mostly leased Nvidia capacity. xAI expanded Grok Bot into a persistent agent that keeps working while the user is offline. Meta launched Muse, a consumer agent connected to email, calendars and payments. OpenAI reports researchers getting 3.1 agent-workdays per human workday.

Each of those converts occasional queries into continuous background computation.

The efficiency argument does not survive it

The standard response to AI energy concern is that cost per token is falling fast, and it is. Anthropic cut cached-context costs by 75 percent with Fable 5.1. Google shipped its third Flash model in six weeks. Microsoft's HydraFusion in GitHub Copilot claims comparable outcomes at up to 67 percent lower cost.

But efficiency gains are being consumed by capability expansion, not banked. Cheaper tokens make longer agentic runs economically viable, which increases total consumption. This is Jevons' paradox operating on a two-year cycle, and the aggregate infrastructure numbers show which force is winning: five gigawatt-scale AI data centres expected online this year, South Korea committing $919 billion towards 8.4GW by 2029, Australia softening renewables rules against a forecast sevenfold rise in demand.

What the study does not do

These are estimates, not measurements. Vals AI does not have access to providers' actual power draw, hardware utilisation or data centre efficiency, and the methodology behind a 10,000-fold multiplier will vary enormously with assumptions about model size, context handling and hardware.

The household electricity comparison is also rhetorical rather than analytical — it makes the number legible without making it precise.

What the study establishes is directional and useful: the per-query framing understates consumption by orders of magnitude for the workloads the industry is now selling, and no provider publishes per-task energy figures that would allow a better estimate.

Venfeed Editor
Editor in chief

Runs the newsroom. Rename this profile in the studio to your own byline.

The Feed · weekdays, 6:30am ET

Every weekday, the AI stories that moved money or shipped code.

No cross-posting, unsubscribe anytime. See all newsletters