
Cut Infrastructure Costs With Peak and Off-Peak Electricity Pricing
IT teams can reduce infrastructure operating costs by matching flexible workload schedules to time-of-use electricity prices. The source v3.2 energy-management design explicitly tracks peak and off-peak tariffs, identifies interruptible tasks that can move to lower-price periods, and calculates the saving for each shift. Its example contains four price periods with a fourfold difference between the lowest and highest tariff.
Workload flexibility sets the limit. Electricity price should influence when suitable work runs, but it should not override service SLOs, deadlines, capacity limits, or business priority.
What is peak and off-peak electricity pricing?
Peak and off-peak pricing means the electricity price changes by time period. The exact tariff can include peak, shoulder, off-peak, and other utility-defined windows.
The source v3.2 example uses four pricing periods. That example comes from the platform design and is not a universal market structure, so the enterprise should load the actual electricity contract or tariff applicable to each site. The platform can then calculate energy cost using the correct price for the time the energy was consumed.
Why does time-of-use pricing matter for AI infrastructure?
AI workloads can be energy intensive, and some of them are flexible in time. A large training task may run for hours, a model evaluation may not need to start immediately, and a batch inference workload may have a completion deadline instead of a real-time latency requirement.
If the electricity price varies significantly during the day, moving these workloads can reduce cost without changing the hardware. The source energy model connects electricity price with scheduling for exactly this reason.
Which workloads are good candidates for shifting?
Good candidates are workloads whose start time can move without violating service requirements, such as batch training, model evaluation, nonurgent fine tuning, backfill tasks, data processing, and some maintenance jobs.
The source specifically refers to interruptible tasks, and the word choice matters: a task that can be delayed or interrupted is easier to shift. The organization should define that flexibility when the workload is submitted, not after the resource is already running.
Which workloads should usually stay where they are?
Online or time-sensitive services should generally be governed by their service requirement first. That covers production inference, customer-facing APIs, urgent incident recovery, deadline-critical training, and critical batch jobs with a fixed completion window.
A high electricity price does not justify missing an SLO. The wider source operations model includes service quality, priority, quotas, and SRE, and energy optimization must stay inside those rules.
How should the tariff be represented?
Store the actual price by site and time window. Useful fields include the data center, effective date, time window, electricity price per kWh, weekend or holiday rules where applicable, and contract version.
The source does not define a tariff data schema. In practice, the cost calculation has to know which rate applies to each energy interval. If the electricity contract changes, version the new tariff instead of rewriting historical cost.
How should workload energy be measured?
Use measured device or allocated workload energy where available. The source energy model includes GPU power, device-level energy, project attribution, and unit Token energy.
For a workload, integrate the relevant power over its execution window and map that energy to the tariff periods it crossed. If a job spans peak and off-peak time, divide the energy across those windows instead of applying one price to the whole job. The exact granularity depends on the telemetry.
How should the potential saving be calculated?
Compare the estimated cost at the current schedule with the estimated cost at the proposed schedule. Conceptually, you take the current schedule's energy by tariff window and the proposed schedule's energy by tariff window, and the difference is the projected saving.
The source v3.2 design calculates each shifted workload's saving individually. That helps because not every move saves the same amount. A short job with low power may not justify the operational complexity, while a long high-power job can produce material savings.
What assumptions should be shown?
Show the assumptions behind the savings estimate: that workload energy and job duration remain similar, that the same accelerator type is available, that no higher-priority task displaces the job, that the target off-peak window has enough capacity, and that the job still finishes before its deadline.
The source does not provide one universal forecasting method, so transparent assumptions are what make the recommendation reviewable.
How should queue and capacity be considered?
Moving work to off-peak windows can create a new capacity peak. Suppose many teams shift training jobs to midnight. Electricity gets cheaper, but the GPU pool becomes oversubscribed, and jobs queue and miss deadlines.
The source scheduler already tracks quota, priority, queue, resource availability, and fragmentation. Energy-aware scheduling should therefore check whether the lower-price window has deployable capacity, since cheap electricity does nothing for a workload that cannot start.
How should deadlines be handled?
A flexible workload still needs a completion requirement. A scheduling rule can consider the earliest start, latest finish, estimated duration, resource requirement, priority, and electricity price.
The source does not define a deadline-aware scheduler algorithm. The principle it does support is shifting interruptible work without violating the service or project requirement, and the platform should make the reason for the selected window visible.
For the scheduling controls around quotas, priorities, reservations, and preemption, what GPU quotas, reservations, priorities, and preemption mean and when each should be used explains how workload urgency and eligibility should remain separate from energy price.
How should preemption affect energy-aware scheduling?
Preemptible work is easier to move, but interruption can waste compute. The source scheduler supports preemption and checkpoint protection.
If a training job has a recent checkpoint, the platform may be able to pause or reschedule it with limited lost work. If the job cannot recover safely, interrupting it for a lower electricity rate can cost more than it saves, so energy optimization should account for checkpoint and restart cost.
How should idle power affect the decision?
A workload moved to off-peak time can still leave hardware powered and idle during peak hours. The source v3.2 energy model also tracks GPU idle power, so a scheduling decision should weigh both the workload energy shifted and the idle energy that remains.
If the infrastructure can safely enter a lower-power state during the unused window, the saving may be larger. If the hardware stays at high idle power, the benefit may be smaller. The exact power-management action depends on the infrastructure and service requirements.
How should cooling cost be considered?
Moving a large workload changes heat load as well as IT power. The source energy model connects power and cooling at the facility level through PUE and zone energy, so a shift to off-peak electricity can also shift cooling demand and affect total facility energy.
The source does not define a predictive cooling-cost model for each workload. A practical calculation can begin with IT energy and then use the approved facility allocation method if the organization includes cooling overhead in the cost model.
For the underlying measurement hierarchy, how organizations can monitor and manage server power consumption at device, rack, and data center level explains how device energy, rack load, and facility energy stay connected.
How should different data centers be compared?
If the enterprise has multiple sites with different tariffs, energy-aware scheduling can become a question of where to run as well as when. The source supports multi-data-center resource views and energy data, but it does not define an automatic cross-site energy scheduler as a completed capability in every deployment.
The source-grounded approach is to compare resource availability, electricity price, network requirements, data location, service latency, capacity, and cooling efficiency before moving work between sites. Do not shift a workload to a cheaper site if data transfer or service constraints make the move impractical.
How should projects see electricity savings?
Attribute the shifted workload and saving to the project. The source metering model already assigns compute and Token consumption by project and tenant, and the v3.2 energy view ranks idle power by project and calculates off-peak savings.
A project report can therefore show the original schedule cost, the shifted schedule cost, the estimated or realized saving, the energy consumed, and the resources used. That gives teams evidence that schedule flexibility has financial value.
How should realized savings be verified?
Compare the forecast with the actual run. After the task completes, measure the actual start and finish, energy, tariff windows, cost, and service outcome.
The source operations model emphasizes closing the loop by feeding execution results back into monitoring and analysis, and that applies here: a planned saving should not be reported as realized until actual usage confirms it.
How should energy-aware recommendations be governed?
Recommendations can be automated, while execution should follow scheduling and authorization rules. The source platform repeatedly separates analysis from production action. A recommendation can say:
"Move Training Job A from the 18:00 peak window to the 23:00 off-peak window. Expected saving: X."
The project owner or policy can approve the change, and the scheduler then executes the approved plan, which preserves accountability.
What should an energy-aware scheduling dashboard show?
A practical view can show the current tariff window and the next off-peak window, the electricity price, flexible queued workloads with their estimated duration and required resources, current and proposed cost, expected saving, deadline, capacity in the target window, approval state, and realized saving.
A platform example that connects tariff periods, GPU power, project attribution, and interruptible scheduling is Sensaka.
If I were implementing time-of-use optimization, I would start with jobs that are already flexible and easy to move. Measure their real energy first, load the actual tariff, calculate the saving transparently, and verify the result after execution. The goal is to move only the work whose timing is flexible enough that the lower electricity price produces a real saving without creating a new capacity or reliability problem.
Frequently Asked Questions
What does the source support for peak and off-peak energy optimization?
The v3.2 energy design includes four electricity-price periods with a fourfold price difference in the example, moves interruptible tasks toward low-price periods, and calculates the saving for each scheduling decision.
Which workloads are suitable for time shifting?
Flexible, interruptible, or deadline-based batch work is the best candidate. Online services, urgent workloads, and tasks with strict completion or latency requirements should not be shifted solely for electricity price.
Should electricity price override reliability?
No. The source treats energy optimization as part of a wider operations model with resource capacity, service quality, scheduling policy, and authorization. Cost optimization remains inside those constraints.