AI Infrastructure

Tokens Per Watt: Why AI Infrastructure Is Moving Beyond GPU-Hour Pricing

Tokens Per Watt: Why AI Infrastructure Is Moving Beyond GPU-Hour Pricing

For quite some time now, GPU-hour pricing has been one of the most straightforward approaches for comparing AI infrastructure prices.

One company lists its price for H100 per hour. The other company offers its price for H200 per hour. The customer can then make a comparison and calculate the approximate cost of their workload.

This approach works rather well for determining the price of hardware rental.

However, it may not be sufficient for calculating the efficiency of using such infrastructure.

New Research

The AI Compute Threshold Report

We analyzed pricing from 150+ GPU cloud providers to find the exact threshold where an AI startup's OpenAI API bill eclipses the cost of a dedicated H100 cluster.

Read the Full Report

This gets increasingly pertinent because of the larger inference workloads. The vendor Qualcomm views itself as the key player for the metric of tokens-per-watt in the agentic AI era, while NVIDIA prioritizes the metrics of cost per token and tokens per watt.

This does not mean that GPU-hour pricing is obsolete. Rather, this means that clients need to have one more tier of pricing options.

GPU-hour allows measuring the price of rented hardware, whereas tokens per watt allows calculating its efficiency.

Key takeaways

  • GPU-hour pricing gauges the hardware rental fee without regard to the actual inference output generated by such hardware.
  • Tokens-per-watt represent efficiency; however, because it depends on the computation performed, engineers should not use it as a benchmark for comparing hardware.
  • Throughput, latency, context length, batching, memory, software utilization, and power measurement will be significant factors in determining the result.
  • Cost per token is the economic parameter that connects infrastructure costs to results. Tokens per watt explains one of these components.
  • In GPU cloud purchasing, benchmark vendors based on the same model, workloads, latency, power limit, and power measurement technique.

What is tokens per watt?

Tokens per watt measures inference per watt of energy, indicating how many tokens AI uses.

At a basic level:

Tokens per watt = number of tokens produced Γ· power used

Suppose you have two inference systems. System A generates 10,000 tokens per second using 5,000 watts, or 2 tokens per watt. System B generates 12,000 tokens per second using 4,000 watts, or 3 tokens per watt.

System B produces greater outputs with reduced energy consumption.

This is not to say that it is cheaper or superior in every way. However, it provides data that GPU-hour prices will not provide: how much inference output your computing system produces relative to its energy consumption.

We ought to treat tokens per watt as a metric of efficiency for AI computations rather than facility metrics like PUE.

Why GPU-hour pricing is no longer enough

GPU-hour pricing provides an obvious answer to the simple question: What is the cost of renting this GPU for one hour?

The inference buyer needs to answer a more troublesome question: What is the amount of useful AI processing that I can do with that infrastructure based on the cost and energy spent?

Here is an example comparing two providers. Provider A charges $2.00 per GPU-hour. Provider B charges $2.50 per GPU-hour. On the surface, Provider A appears to be cheaper.

However, what if Provider B provides significantly higher throughput for the same task because of memory optimisations, batching, software optimisation, networking, and other factors? This means that the cheaper GPU-hour will not be cheaper.

Utilisation is another key aspect, since we must distinguish a GPU that sits idle waiting for memory and other parts of the serving infrastructure from one that produces useful results.

The economics of inference involve many additional factors like latency, model type and size, context length, precision, memory size and bandwidth, networking, energy consumption, power management, cooling, and more.

Inference is increasingly becoming a problem of system-level optimisation.

From FLOPS to tokens per watt

FLOPS still have value. They quantify theoretical computing capabilities and enable consumers to know what a processor is capable of.

However, inference consumers rarely buy FLOPS because they are useful in themselves. They buy infrastructure that will provide services to users.

For an LLM use case, this is about producing tokens within some latency target.

The following path makes sense:

FLOPS β†’ tokens/second β†’ tokens/watt β†’ cost/token

FLOPS tells you how much theoretical computing capability the hardware can provide. Tokens/second tells you how many results the system can produce. Tokens/watt tells you how efficient this production is in terms of power. Cost/token tells you what this result costs.

The point here is not that FLOPS is no longer relevant. The point is that FLOPS doesn’t tell us anything about the economics of production inference.

The most recent NVIDIA material on token economics uses the same trick by combining throughput, efficiency, and the cost of tokens in terms of AI infrastructure.

What actually affects tokens per watt?

TPW is helpful information, but it does not have an absolute value for any piece of hardware since it can show completely different performance when used differently.

Model architecture

Models require various amounts of computation power per token generated.

Different dense, mixture-of-experts, and multimodal models can have significantly varying compute, memory, and network demands.

Thus, TPW comparison between different models is not quite right without workload control.

Context length

TPW heavily depends on the context for LLMs.

According to “The 1/W Law” study from 2026, identical hardware could show a nearly 40x variation in tokens per watt performance between specific 4K and 64K context cases. Authors of the paper note that these are analytical findings, not new hardware experiments.

More context leads to more KV-cache use, which lowers the number of concurrent sequences supported by a GPU while keeping similar power usage.

This way, it generates fewer tokens while using the same amount of power.

Batch size

However, the best batch size will be determined by both the workload and the latency objective.

An optimal batch configuration based on achieving maximum throughput would not work for an interactive application that requires quick response time.

Here, the service provider claiming high tokens per watt must specify the batching and latency assumptions used in the measurement.

Prefill and decode

LLM inference is not a single process.

The LLM first takes in the input context, referred to as the prefill, and outputs tokens in the decode phase.

These phases can put different pressure on the infrastructure.

Reporting a single efficiency metric in the benchmark is insufficient to capture some aspects of the workload.

Memory

Inference in the modern world is becoming more reliant on efficient movement and access to data.

The size of the memory, the memory bandwidth, and the KV-cache performance may affect the performance achieved.

That is one reason system architecture is also important along with the accelerator.

Software optimization

Software also influences hardware efficiency.

Kernel improvements, batching, quantisation, speculative decoding, scheduling, and frameworks for serving can all affect the quantity of meaningful work that the hardware produces.

For instance, MLPerf Inference v6.1 includes the use of modern serving and agentic workloads instead of treating inference as a one-off benchmark for an accelerator.

Engineers must see tokens per watt as a measure of workloads and stacks, and not a forever specification etched into a GPU.

Tokens per watt vs. cost per token

Token per watt is a measure of energy efficiency, while cost per token is an indicator of economic efficiency. While these two indicators go together, they are not interchangeable.

The simplified formula for cost-per-token calculation is as follows:

Cost per token = infrastructure cost Γ· useful tokens

The infrastructure cost might consist of GPU cost, CPU and RAM, storage space, network, electricity supply, software, and all other costs associated with inefficient use of resources.

In case the first supplier shows higher token per watt, while the cost of the infrastructure he uses is rather high, then his cost per token would be higher than that of a supplier who demonstrates lower token per watt but whose GPU price is lower.

That is why efficiency calculations should not play the only role in the procurement process.

A good hierarchy would be the following:

GPU hour β†’ throughput β†’ power efficiency β†’ cost per token

Can you compare GPU providers using tokens per watt?

Yes, but only when providers use the same test conditions.

Providers that report 5 tokens per watt do not necessarily perform better than those that report 3 tokens per watt.

Before making any comparisons, normalise the benchmark.

Benchmark conditions to record

Model: Different models use different amounts of computation.

Context length affects memory and KV-cache needs.

Input-to-output ratio: prefill and decoding have unique characteristics.

Batch size affects utilisation and throughput.

Precision: affects compute and memory needs.

Framework: Software can significantly impact performance.

Tokens per second: provides information about output throughput.

Latency: prevents high throughput from obscuring low responsiveness.

Power boundary: defines the power being tested.

Utilisation: determines the efficiency of capacity usage.

What buyers should ask GPU cloud providers

It has become important to disclose more information about GPUs besides their type and hourly price. Buyers should also understand what a GPU provider’s pricing page does and does not reveal.

Find out whether a particular workload can support the throughput on this system.

Identify the target latency objective. A system optimised for maximum throughput may differ significantly from a system optimised for low latency.

Check the specific model and context length used during testing. Tokens per Watt measurement without reference to the workload is difficult to understand.

Find out what precision and the specific software framework powering the system. There could be significant differences in performance between FP16, BF16, FP8, and quantised.

Determine how the provider tracks power consumption. Power consumption at the level of the accelerator, server, rack, and entire facility differs.

Verify the baseline utilisation levels. The saturated lab-based benchmark is different from the production-like workload.

And finally, find out what the cost per million tokens is. It will link efficiency measurement to the economic considerations that matter to the procurement teams.

Therefore, a useful comparison of providers could include the following: GPU hour + tokens per second + latency + tokens per watt + cost per million tokens.

What tokens per watt means for GPU cloud pricing

Tokens-per-watt measurement could redefine how buyers evaluate GPU infrastructure.

Currently, cloud marketplaces mostly market resources based on their GPU type together with the hourly cost.

For heavy inference workloads, buyers require a more detailed marketing breakdown: GPU type, throughput, latency, efficiency, and the cost per token.

This does not imply that GPU clouds will not measure their costs per hour. Measuring costs per hour is still a valuable tool for flexible use cases, development, training, and reserved infrastructure.

However, those who need inference services will prioritise attention to the output economy of the infrastructure, rather than just its cost per hour.

Marketplaces will eventually begin measuring tokens per second, tokens per watt, cost per million tokens, latency, GPU usage, benchmark workload, and availability.

The limitations of tokens per watt

Tokens per watt is an important measure, but it must not become another shortcut measure.

Firstly, not all AI use cases have token counts as their natural output metrics. Vision models, recommendation systems, speech workloads, and retrieval pipelines may have other types of output metrics.

Secondly, the agentic workload makes the definition of output usefulness complicated because the agent is creating tons of tokens while performing tool calls and actions. This does not provide a quantified task performance value.

Thirdly, the output efficiency depends heavily on the workload configuration. Context length, batching, architecture, and routing may play a role in the output metrics.

MLCommons is moving towards agentic and end-to-end RAG benchmarking for this very reason. In the MLCommons’ RAG benchmarks, we see that the workload, which includes retrieval, re-ranking, and multiple models, cannot adequately capture workloads in terms of the tokens per second metric.

This is why tokens per watt should remain one component of the infrastructure benchmark.

FAQs

Does tokens per watt include cooling and facility power?

No, not necessarily. Providers can measure tokens per watt at various levels of power consumption, such as the accelerator level, server level, rack, or facility level. Before making any comparisons between two numbers, the buyer needs to know about the power level being used.

Why can short AI responses distort energy-per-token measurements?

The energy per request for a brief response can contain a relatively higher proportion of the fixed request energy. Research on inference energy has shown that there can be a difference in direction between request energy and token energy. Therefore, a tokens-per-watt calculation for very brief responses may not reflect a production scenario.

Can more output tokens make an inference run look more efficient?

Sometimes. Studies of inference energy have shown that the energy per output token could decrease as output length increases, although the total energy required for the query increases. The measure will then be better on a per-token basis without the task requiring less total energy.

How often should a GPU provider refresh its tokens-per-watt benchmark?

There isn’t a universal refresh interval in the industry. A benchmark needs to be redone in case the provider makes any updates to their model, serving infrastructure, precision, batching parameters, accelerator generation, etc. Otherwise, an outdated result would no longer represent what the service customers are paying for.

Should buyers compare tokens per watt at steady state or under real traffic?

The two measurements answer different questions. The steady-state measurement helps isolate infrastructure efficiency, whereas production-like traffic will help in understanding the impact of bursts, queueing, mix of requests, and idle time. The most important metric for procurement is the metric that is closest to the planned workload.

Can tokens per watt be used to estimate data-centre power requirements?

Only as an input to another model. A token-efficiency measurement requires the token volume, concurrency, utilisation, and power limit to make sense. Engineers should never use it to estimate facility demand without these assumptions about the workload.

Conclusion

However, the GPU-hour is not going anywhere, except that it is simply becoming less and less relevant as a single point of measurement by which buyers judge the value of inference infrastructure.

As AI use cases evolve toward real-time inference, long-term context, reasoning, and agency, buyers will want to understand how effectively their infrastructure can turn the precious resources they have available into valuable output. That also makes it important to choose the right GPU for the workload rather than judging hardware in isolation.

This involves taking things further than simply asking, “How expensive is it per hour?” to instead considering, “How many valuable AI computations is this infrastructure performing per unit of energy per dollar?”

Tokens per watt represent efficiency. Cost per token represents the economics.

Buyers must evaluate both measurements alongside issues of throughput, latency, utilisation, memory, networking, and workloads.

What buyers should do is actually straightforward: consider what infrastructure is delivering, rather than how expensive the underlying GPU is per hour.

Share this article
Find the best GPU cloud for your workload

Get personalised, no-commitment quotes from top AI infrastructure providers in under 2 minutes.

Get Free Quotes β†’