From PUE to TPW: SenseTime's AI Infrastructure Unit Proposes a New Efficiency Standard
From PUE to TPW: SenseTime's AI Infrastructure Unit Proposes a New Efficiency Standard
Producing more tokens from limited energy and cost is the defining challenge of AI infrastructure in the token economy.
Token consumption is rising exponentially, with compute demand following suit and AIDC electricity usage claiming an ever-larger share of national power consumption. Industry attention has shifted from who owns the most compute to who can produce the most high-quality tokens at the lowest cost and highest efficiency.
Lin Hai, general manager of SenseTime's AIDC, argues that in the token economy, AIDCs should be reframed as value platforms linking electricity to tokens. The depth of joint compute-energy optimization, he contends, is now a decisive variable in commercial competitiveness. [IMAGE:0]
[Lin Hai, General Manager, SenseTime AIDC]
01 Why PUE Falls Short in the Token Economy
How should AI infrastructure efficiency be measured, and how can it be improved? For years, the answer was PUE. The ratio of total to IT energy consumption in a data center, PUE has served as the industry's gold standard for energy efficiency. For more than a decade, virtually every energy-saving initiative—airflow optimization, cooling architecture redesign, free-cooling adoption, scaled liquid-cooling deployment—has targeted a lower PUE.
AIDC business models, however, are moving from selling GPU-hours to selling tokens. As the commercial unit changes, so must the efficiency metric. PUE measures electricity consumed; it cannot answer what that electricity produced in effective value—a blind spot the industry now confronts.
At the root lies a fractured data link between power input and token output, a legacy of the industry's existing division of labor.
• The grid feeds IDCs yet remains blind to the workloads inside the racks;
• IDCs bill by rack, capacity, or kilowatt-hour, yet often cannot confirm effective GPU utilization;
• Compute clouds sell card-hours but seldom track the token output those resources generate;
• Model services bill by token but rarely consider the electricity price underlying each workload.
Each layer optimizes its own metrics in isolation; global optimality is therefore unattainable.
The deeper contradiction stems from the old metric's limitations: falling PUE does not guarantee falling total energy use or token cost. Raising supply-air temperature, Lin Hai notes, may trim cooling load and flatter the PUE ledger, yet lift GPU and server-fan draw—leaving total consumption flat or token costs higher.
Because PUE cannot price the value of the electricity consumed, a new measure is required.
02 From PUE to TPW: Reframing AIDC Efficiency
SenseTime's AI infrastructure unit has therefore introduced TPW (Tokens Per Watt), taking effective token output per unit of electricity cost as the core measure, unifying compute efficiency, energy efficiency, and business value.
PUE measures electricity utilization; TPW asks how many effective tokens each kilowatt-hour yields, incorporating time-of-use tariffs into the economic calculus. Attention shifts accordingly from electricity consumed to effective output produced.
TPW neither rejects nor merely replaces PUE; it extends PUE into a full-chain assessment from power input to token output. Disparate levers—tariffs, storage, PUE, GPU utilization, model efficiency—are thereby aligned on a single value metric: token output. [IMAGE:1]
TPW's value lies less in the metric itself than in the shared value language it creates. Utilities, IDCs, compute platforms, and model providers can now optimize against a common standard, rendering compute-energy coordination measurable, replicable, and commercializable.
03 System-Level Optimization Through Compute-Energy Coordination
Given a new yardstick, the question becomes how an AIDC can produce more tokens within a fixed power envelope.
SenseTime's answer is neither more compute nor more capacity. A compute-energy coordination Agent jointly orchestrates workloads, energy supply, and infrastructure, extracting the time value of time-of-use tariffs, demand control, and demand response—and wringing more from each kilowatt-hour.
The framework proceeds along three layers: data integration, workload scheduling, and compute-energy coordination.
Step one: data integration and observability
Optimization begins with observability.
In a conventional IDC, the power system monitors physical assets—buildings, racks, servers—while the compute platform tracks cloud-native entities—Jobs, Pods, Nodes. Long independent, the two domains cannot answer two critical questions: who consumes the power, and are the loads schedulable?
SenseTime has accordingly built an eight-tier data-penetration architecture, mapping workloads, compute resources, and physical consumption onto a single timeline. GPU utilization, task-level power, rack load, and building draw are reconciled so that every kilowatt-hour maps to a specific task—the data foundation for forecasting, scheduling, and optimization. [IMAGE:2]
Zhang Xu, operations director of SenseTime's AIDC, notes: "At the Lingang AIDC control wall, this mapping now operates in production: real-time task power, node GPU utilization, rack and floor load, building IT draw, and feeder-side consumption share a single observation chain—rendering every kilowatt-hour traceable in both destination and output."
Step two: workload tiering for resource elasticity
With data unified, optimization proper begins.
Rather than relocating workloads or adding hardware, SenseTime schedules compute by task-level timing tolerance.
Real-time workloads—online inference, for example—demand instant response; training, evaluation, and batch jobs tolerate minute-, hour-, or day-scale deferral. Using business priority and latency tolerance, the system shifts deferrable workloads to windows of lower tariffs, lighter load, or greater resource slack, lifting utilization without breaching service quality.
Total consumption is unchanged; the gain comes from timing. Idle capacity and power are better exploited so that identical capacity and energy yield more effective tokens—and a higher TPW.
Step three: multi-axis coordination, scaling load-flexibility value
SenseTime further pairs capacity mining, storage dispatch, and grid demand response in dynamic compute-energy coordination. Load forecasting, capacity optimization, and storage shaving unlock existing infrastructure's latent potential, allowing fixed power to carry more effective compute without new supply capacity.
Zhang Xu: "Measured results show effective IT carrying capacity up roughly 60% under existing supply. The gain derives not from new capacity but from system-level coordination, enabling the same megawatts to carry more effective compute—sustained efficiency gains across AI infrastructure." [IMAGE:3]
[Zhang Xu, Operations Director, SenseTime AIDC]
04 Lingang AIDC: Scaling Compute-Energy Coordination
The methodology has now been validated at scale at SenseTime's Lingang AIDC.
In Shanghai's July 10 summer peak-shaving demand-response event, the Lingang AIDC was first to deploy compute-energy coordination: midday baseline load cut 75%, dispatchable load of 23MW, and 46,000 kWh released to the grid over two hours via suspended non-critical tasks, infrastructure drawdown, and storage discharge.
The AIDC-grid relationship is thus changing. Once rigid loads deserving priority protection, AIDCs now treat load flexibility as an operable, dispatchable asset as tasks become identifiable, tierable, and schedulable. Workloads are timed by business priority; storage, compute, and grid demand move in concert in demand-response participation.
The shift from PUE to TPW moves optimization from equipment efficiency to whole-system operating efficiency. SenseTime's gains derive not from added compute or power, but from system efficiency latent in task elasticity, resource scheduling, and energy coordination.
05 The Compute-Energy Coordination Agent: Putting TPW Into Operation
A new metric alone is insufficient; TPW demands a system engineered to improve it continuously.
Built on years of AIDC operations, SenseTime's compute-energy coordination Agent is an intelligent scheduling system that perceives, reasons, decides, executes, and evolves. Where traditional energy platforms see only the power side, the Agent links compute, energy, and business in end-to-end optimization, enabling each kilowatt-hour to yield more tokens. [IMAGE:4]
Architecturally, the Agent relies on an eight-tier data-penetration system, five intelligent decision chains, and six operational capabilities, dismantling silos among workloads, compute, energy, and infrastructure in a closed loop spanning sensing, forecasting, decision, execution, and measurement. Its decisions cover load forecasting, tariff analysis, storage optimization, capacity control, HVAC coordination, and compute scheduling, with TPW as the single objective that unifies previously fragmented optimization actions.
The Agent is now in full production at the Lingang AIDC: token output per unit of electricity cost up 80%, average tariffs roughly 10% below regional peers, and compute-load forecast accuracy of 96%.
To enable replication, Zhang Xu notes, the Agent adopts a platform, modular design, converting years of operating experience into reusable capability, with an FDE (Field Domain Expert) system standardizing delivery. New scenarios deploy in as little as one week, fast-scaling the capability.
For national computing hubs and ten-thousand-card AIDCs, the Agent unifies scheduling of compute, energy, and storage, lifting cluster-wide compute efficiency and improving renewable off-take and total energy cost.
For compute operators, AI cloud providers, and token factories, the system pairs tariff forecasting with load optimization and compute scheduling to drive per-token electricity costs lower, raising utilization and margins without breaching SLA commitments.
For enterprise AIDCs, research institutions, and vertical AI platforms, modular architecture and standardized delivery enable rapid deployment without infrastructure rework—yielding smarter, more efficient, and greener operations.
Lin Hai, general manager of SenseTime's AIDC, argues that the move from PUE to TPW alters not only the metric but the operating model of AI infrastructure itself. Compute-energy coordination, he predicts, will evolve from an AIDC energy-saving tactic into a core capability linking compute, energy, and business. Through its Agent, SenseTime intends to standardize, productize, and scale optimization capability, bringing AI infrastructure into the age of system-level compute efficiency.