
AI Infrastructure Cost by GPU Hour, Token, Project, Tenant, Model
Companies can measure AI infrastructure cost by combining resource metering and service metering under the same ownership model. Accelerator hours measure how long expensive compute is allocated, Token metering measures model-service consumption, and project, tenant, and model identifiers explain who used the resources and what produced the cost.
The arithmetic is easy; allocation is where it gets hard. If the scheduler, API gateway, energy system, and billing system use different project names or time windows, the same workload can produce several conflicting cost numbers. Cost accounting needs one shared definition of resource, owner, usage window, and unit.
What is a GPU hour?
A GPU hour is one accelerator card assigned or consumed for one hour according to the metering policy. A simple calculation is:
GPU hours = Number of GPUs × Allocation time in hours
If a training job uses 8 GPUs for 3 hours:
8 × 3 = 24 GPU hours
This metric measures resource time, and it does not automatically measure useful work. A job can consume 24 GPU hours while average utilization sits at only 40 percent. The distinction matters because allocated capacity still has economic value when the workload is inefficient, so the cost system should keep GPU hours and utilization as separate metrics.
GPU hours answer "How much accelerator capacity did this workload occupy?" Utilization answers "How busy was the accelerator while it was occupied?" Put the two together and idle cost becomes visible.
Should GPU hours use allocated time or active time?
Use the definition that matches your business question, and label it clearly. Allocated GPU hours count the full stretch in which a card is reserved for the workload. Active GPU hours count only the time that meets a defined activity rule.
For internal capacity planning, allocated hours are often more useful, because the scheduler could not give the same card to another tenant during that time. For performance analysis, active time can show how much of the reservation actually did useful work. For billing, the commercial model decides. If a tenant reserves a dedicated GPU for ten hours, charging only for moments above a utilization threshold may not reflect the capacity held for that tenant.
A useful reporting model can show allocated GPU hours, estimated active GPU hours, idle rate, and average utilization side by side, which separates capacity consumption from efficiency.
How do you calculate cost per GPU hour?
Cost per GPU hour is the cost allocated to accelerator capacity divided by the GPU hours over the same span of time. A simple formula is:
Cost per GPU hour = Allocated accelerator-related cost / GPU hours
The numerator can be narrow or broad. A narrow model may include only hardware depreciation and electricity. A broader one may include:
- Accelerator depreciation
- Server chassis cost
- Power
- Cooling overhead
- Rack or colocation cost
- Network
- Storage
- Software
- Operations labor
- Maintenance
No numerator is universally correct, but it has to be consistent. If you call the result "fully loaded cost per GPU hour," define every cost category it includes. If you call it "energy cost per GPU hour," include energy cost only. And do not compare two unit costs that use different boundaries.
How does Token metering fit into infrastructure cost?
Token metering connects model-service output to the infrastructure that produced it. A model-serving platform can record input and output Tokens by project, tenant, API key, model, and time window, while the infrastructure layer records accelerator hours, energy, and other resource cost over that same window.
With both in place, the system can calculate cost per 1,000 Tokens, cost per 1 million Tokens, energy cost per Token, infrastructure cost per Token, and cost broken down by model, project, and tenant. AI cost stops being an equipment number and becomes a service-unit number.
For the service architecture behind those records, what MaaS is and how model repositories, inference instances, API gateways, and Token metering work together explains the full chain.
How do you calculate cost per Token?
Cost per Token is the cost attributed to a model service divided by the Token volume attributed to the same service and time window.
Say a model service is allocated $600 of infrastructure cost during one day and processes 300 million billable Tokens:
$600 / 300,000,000 = $0.000002 per Token
That can also be reported as:
$2 per 1 million Tokens
The example is illustrative, and your result depends entirely on the cost boundary and the Token definition. A production system should define how it treats input versus output Tokens, failed requests, retries, cached results, shared-instance allocation, time zone, billing cycle, model version, and project and tenant ownership. The number is only useful when that definition is stable.
Should input and output Tokens be costed separately?
They can be, especially when output generation behaves quite differently from input processing in compute terms. Separate pricing is a business decision, though, and infrastructure operations do not require it.
At minimum, record input and output Tokens separately wherever the serving system exposes them. You can always combine them later, but you cannot split them later if you never recorded the distinction.
Keeping both also helps explain cost changes. Suppose request count stays flat but average output length doubles. Total Token volume rises even though the number of calls has not changed, and that can push up inference compute and cost. A dashboard that only counted requests would hide the cause.
How do you allocate cost by project?
Link every resource reservation and service call to a stable project identifier. For training, the scheduler or job submission process should require the project identity. For inference, the API key, service route, tenant, or application identity should map to a project.
Dedicated resources are easy to allocate. Shared resources need a documented allocation rule, based on something like GPU reservation time, actual execution time, Token volume, request volume, reserved capacity share, or measured utilization share. Pick the rule that best reflects how the shared resource is consumed and keep it the same from one reporting period to the next.
The platform should also keep the raw usage records so finance or operations can explain the result.
How do you allocate cost by tenant?
Tenant allocation works the same way, with the tenant as the primary ownership dimension. A tenant may contain several projects, so the data model has to support a hierarchy. For example:
- Tenant A
- Project A1
- Project A2
- Model Service A
- Training Job A3
Costs can then be reported at each level without counting usage twice. If Project A1 consumes 100 GPU hours and Project A2 consumes 50, the tenant total is 150, and Token usage rolls up the same way.
That is why one shared ownership model is essential. Do not let the scheduler use department names while the gateway uses account IDs and finance uses yet another project code, with no mapping between them.
How do you allocate cost by model?
Model-level cost needs a relationship between the model version, its inference instances, and the compute they consume. For dedicated model instances this is simple: if Model X occupies four GPUs for ten hours, those 40 GPU hours can be attributed directly.
Shared serving is harder. If several models share one accelerator, the platform needs an allocation rule based on execution time, Token volume, reserved share, or another measurable driver.
Model-level cost should also carry service quality context. A cheaper model is not automatically better if it misses the required SLO or produces output of unacceptable quality. Cost and service quality are separate operating dimensions, and the comparison worth making is cost at the required level of quality.
How should energy be included?
Use measured accelerator or server power where it exists, with a clear rule for facility overhead. At accelerator level, integrate power over time to get kWh. At server level, use server metering where available. To estimate facility-attributed energy, you can apply a documented facility overhead factor such as a relevant PUE measurement, as long as you label it as an allocation method. Then apply the electricity tariff for the matching time window.
This matters when electricity prices change by time of day. A flexible training job that runs off peak may cost less in energy even though it uses the same number of GPU hours.
For the formulas, how a data center can calculate PUE, WUE, GPU energy consumption, and energy cost per Token gives the detailed method.
How do you measure idle cost?
Idle cost is the cost of allocated or owned capacity that is not producing the intended amount of useful work, under the organization's own definition. So start by defining idle. A completely unallocated card is one type, a card allocated to a job at 5 percent utilization is another, a card waiting on storage is a third, and a fragmented pool with free cards that cannot satisfy any queued job is yet another.
These causes should not be lumped into one number and blamed on users. Classify the idle reason where you can, for example:
- No demand
- Quota restriction
- Fragmentation
- Network bottleneck
- Storage bottleneck
- Data loading
- Degraded hardware
- Oversized resource request
Then calculate the accelerator hours and energy tied to each cause. An idle-rate KPI becomes an optimization backlog.
How should shared infrastructure cost be allocated?
Shared cost, with network and storage as the usual examples, should use a stable allocation driver that people understand and that is hard to game. Possible drivers include bandwidth consumed, storage capacity reserved, I/O volume, GPU hours, project share, or a fixed base charge plus variable usage.
No single method is best. What counts is whether the allocation helps someone make a decision. If storage cost is tiny next to GPU cost, a complicated IOPS-based allocation may create more administrative work than value. If storage is a major cost and certain projects consume a disproportionate share of bandwidth, a more detailed rule may be justified. Start simple, and add precision only when the result would change a real decision.
How should cost data connect to quotas?
Cost and quota definitions should share the same resource model. If a tenant has a quota for eight high-memory GPUs, the metering system should identify those same resources consistently. If the quota system treats a partitioned GPU as one resource while billing treats it as a fraction, document the conversion, because misaligned definitions create disputes.
The source operating model calls this out directly: accelerator hours, energy consumption, and Token metering need aligned allocation definitions so different modules do not produce inconsistent figures. That makes it a data-governance issue as much as a billing one.
What should an AI cost dashboard show?
A useful cost dashboard shows cost, usage, efficiency, and ownership together. At minimum:
- Total accelerator hours
- Available versus allocated capacity
- Idle rate
- Cost per GPU hour
- Token volume
- Cost per Token
- Energy consumption
- Project ranking
- Tenant ranking
- Model ranking
- Resource type or accelerator model
- Trend over time
Every summary number should let you drill down to the usage records that produced it.
A platform example using one shared metering model across projects, tenants, models, and accelerator resources is Sensaka.
If I were building AI cost accounting, I would start with one rule: every GPU hour and every Token must carry an owner and a time interval. Once that holds, project, tenant, and model reporting is a grouping problem. Without it, even a beautiful cost dashboard is mostly an estimate.
Frequently Asked Questions
What is a GPU hour?
A GPU hour is one GPU allocated or used for one hour under the organization's metering definition. Eight GPUs allocated for three hours equal 24 GPU hours, even if utilization during those hours is lower than 100 percent.
How do you allocate AI infrastructure cost to a project or tenant?
Link workload and API usage records to project and tenant identities, then allocate direct and shared costs using documented rules. Keep the same identifiers across scheduling, metering, service gateways, and billing.
What is the difference between cost per GPU hour and cost per Token?
Cost per GPU hour measures the cost of infrastructure capacity over time. Cost per Token measures the cost associated with a unit of model-service output, so it connects infrastructure cost with service consumption.