Beyond GPUs: How token economics are reshaping the modern data center

Zziwa Raymond
Full-stack engineer and member of the Andela Talent Network
Aug 7, 2026
9 min
Beyond GPUs: How token economics are reshaping the modern data center

A few years ago, few engineers knew the term tokens-per-watt. Today, it is one of the defining metrics in AI infrastructure.

Data centers historically relied on metrics like server utilization, floating-point operations per second (FLOPS), and power usage effectiveness (PUE). Those measurements fit an era when facilities functioned as business support centers housing applications, storing data, and processing predictable traffic. That era has ended.

The scale of this transition is unprecedented. Omdia’s Global AI Factory Market Landscape report projects global tech companies will spend more than $600 billion on AI infrastructure alone, with cumulative data center investment reaching $1.6 trillion by 2030. Goldman Sachs projects global token usage will multiply 24-fold between 2026 and 2030, reaching 120 quadrillion tokens per month.

This buildout represents the largest concentrated capital investment in computing history. Gartner projects AI infrastructure will become a $3.3 trillion market opportunity by 2029. (Gartner, Forecast Alert: AI in IT Spending, August 19, 2025) Meanwhile, the Linux Foundation launched the Tokenomics Foundation alongside the FinOps Foundation to establish open standards for tracking AI token economics.

What is an "AI factory"?

The term "AI factory" signals a fundamental architectural and operational shift rather than a simple data center rebrand. Omdia defines an AI factory as a new class of industrial infrastructure where intelligence itself is the primary product, with digital tokens and model performance defining output.

Regardless of facility scale, AI factories rely on a four-layer architecture:

  1. Energy and physical infrastructure: Power, cooling, and facilities.
  2. Hardware and network interconnect: GPUs, custom silicon, and networking fabrics.
  3. Scheduling and virtualization orchestration: Resource management and workload distribution.
  4. Model-as-a-Service and AI application ecosystem: The models and applications delivering business value.

AI factories operate as manufacturing plants for intelligence. Their efficiency depends on how effectively they convert power, cooling, and silicon into useful tokens.

Tokenomics: a new measurement framework

For decades, FLOPS measured raw compute throughput. However, FLOPS fails to account for memory bandwidth, system utilization, or energy efficiency under real-world AI workloads.

The industry is adopting a new evaluation framework: tokenomics (token economics). In AI factories, the primary operational metrics have shifted from server utilization to three core yield figures:

  • Tokens-per-dollar
  • Tokens-per-watt
  • Tokens-per-second

What is a token?

A token is the fundamental unit of information that an AI model processes. In text models, tokens represent word fragments, punctuation, or short character sequences. In vision and audio systems, tokens represent image patches, sound segments, or compressed video frames.

AI providers bill usage per million input and output tokens. Because tokens serve as the foundational unit of exchange in AI infrastructure, consumption requires strict governance.

The enterprise reality check

Financial realities are forcing enterprises to adopt tokenomics. As AI budgets surge, executive oversight has intensified.

In early 2026, enterprise adopters like Meta, Amazon, Uber, and Salesforce pushed for rapid AI adoption, setting up internal leaderboards to track token consumption. Uncapped usage led to budget exhaustion: Uber burned through its entire 2026 AI budget in four months, while Salesforce ran up an estimated $300 million annual bill for Anthropic services.

A swift industry correction followed. Companies abandoned gamified consumption metrics in favor of governance frameworks that treat tokens as finite operational resources—akin to engineering hours or cloud compute budgets. When enterprise CFOs discuss "tokenomics," they mean AI cost containment, not cryptocurrency.

The stack tax: why integration matters

The "stack tax" describes the compounding inefficiency that occurs when engineers overbuild individual infrastructure layers independently.

According to data presented at Data Center World 2026, uncoordinated layer design can add roughly 35% overhead at each tier. An optimized 1 GW AI factory running 300,000 GPUs achieves maximum token throughput. That same facility, if poorly integrated across its stack, may realize only 65% of its throughput capacity—wasting massive capital and power resources.

Full-stack optimization

Nvidia frames AI infrastructure as a five-layer stack spanning energy, physical facility, silicon, models, and applications. All five layers must scale in tandem.

Cooling is no longer merely a facility management task; it directly dictates chip performance and thermal throttling. Similarly, advanced packaging decisions determine rack density, while local power limits constrain facility placement. Modern optimization focuses on two core ratios:

  • Watts-per-token: Energy efficiency per unit of intelligence generated.
  • Performance-per-token: Throughput maximization under operational latency and memory constraints.

The four technology levers

A McKinsey analysis of 13 inference cost-reduction technologies highlighted four high-impact levers that, when combined, can reduce inference costs by up to two orders of magnitude:

  1. Model optimization: Techniques like low-precision quantization, pruning, and model distillation offer an 85% to 95% cost reduction without degrading response quality.
  2. Advanced packaging: Placing high-bandwidth memory closer to compute logic widens interconnect buses and reduces data movement latency.
  3. Custom silicon (inference ASICs): Application-specific chips optimized purely for inference workloads reduce power consumption and per-unit costs compared to general-purpose GPUs.
  4. Co-packaged optics (CPO): Integrating optical interfaces directly with compute silicon lowers the power and latency of high-speed interconnects over longer physical distances.

Model optimization remains the most immediate path to efficiency. DeepSeek's R1 model demonstrated this principle by delivering competitive reasoning capabilities at a fraction of the compute cost of contemporary models, proving that tight architectural constraints often force superior engineering.

Custom silicon: inference-optimized chips

While GPUs dominate model training, inference workloads demand specialized silicon. Application-specific integrated circuits (ASICs) tailored for inference lower power consumption and unit costs.

Deloitte estimates revenues for inference-optimized chips will surpass $50 billion in 2026, up from $20 billion in 2025. Major cloud providers and chipmakers—including Meta, Google, Amazon, Intel, AMD, Qualcomm, Groq, SambaNova, and Cerebras—are expanding custom silicon lines to capture this demand.

The inference boom

The AI market has entered its second phase: inference at scale.

Deloitte's TMT forecast projects global AI data center capex will reach $400 billion to $450 billion in 2026, with chips accounting for more than half of that expenditure. Total capex could hit $1 trillion by 2028. Although training growth is stabilizing, operational compute demands from post-training scaling, test-time reasoning, and live production queries continue to multiply.

Inference everywhere

AI workloads are moving beyond hyperscale data centers into edge environments, local hardware, and physical robotics systems. Three primary drivers accelerate this distribution:

  • Edge AI: Extended reality, robotics, and local sensors require low-latency processing at the edge.
  • Physical AI: Autonomous vehicles and industrial automation demand real-time, on-device inference.
  • Sovereign AI: National governments are deploying domestic infrastructure to guarantee data sovereignty and reduce vendor dependence.

Omdia predicts regional and industry-specific infrastructure providers will represent the fastest-growing market segment by 2031.

The "burst-first" strategy

To balance compute efficiency with cost control, Gartner recommends a "burst-first" infrastructure approach: (Gartner, Burst Your AI Bubble: Prove Outcomes Before Scaling AI Infrastructure, November 11, 2024.)

  1. Leverage public cloud capacity for temporary traffic spikes.
  2. Deploy dedicated on-premises infrastructure for predictable baselines and data sovereignty.
  3. Transition physical AI models systematically from simulation environments to production floors.

This hybrid approach allows engineering teams to move quickly from pilot to production without incurring unmanaged infrastructure debt.

Gartner's top three infrastructure trends

Gartner highlights three primary trends redefining infrastructure strategy: (Gartner, Top Strategic Technology Trends for 2025, October 21, 2024])

1. Building AI supercomputing platforms

Modern supercomputing platforms require deliberate alignment with token economics. Engineering leaders should:

  • Avoid speculative capacity expansion; align hardware directly with model parameter size and query volume.
  • Design collaboratively across physical, network, and application layers.
  • Audit token yields continuously to verify high-cost compute instances deliver proportional business value.

2. Deploying AI everywhere

Inference is scaling across cloud, regional data centers, and edge devices. Infrastructure teams must optimize compute scheduling across heterogeneous hardware (GPUs, NPUs, and TPUs) while embedding FinOps discipline into workload orchestration.

3. Automated operations and AI security

As AI transitions from experimental software to core infrastructure, systems demand enterprise-grade fault tolerance, automated failovers, and robust security controls.

Conclusion

AI infrastructure is undergoing its most significant transformation since the dawn of cloud computing. Data centers are converting into AI factories—industrial-scale production systems evaluated on tokens-per-dollar, tokens-per-watt, and tokens-per-second rather than peak FLOPS.

The operational pressure is already present. Fragmented infrastructure designs waste up to 35% of capacity, while unmanaged query volumes exhaust enterprise budgets. Achieving efficiency requires holistic system design: teams must optimize power, cooling, silicon, and software as a unified stack.

The next generation of engineering leaders will not be defined by the size of their models or total GPU counts, but by how efficiently their architecture converts energy and silicon into useful intelligence.

Zziwa Raymond
Full-stack engineer and member of the Andela Talent Network
No items found.
No items found.
No items found.

Recent articles