Close Menu
Daily Guardian
  • Home
  • News
  • Politics
  • Business
  • Entertainment
  • Lifestyle
  • Health
  • Sports
  • Technology
  • Climate
  • Auto
  • Travel
  • Web Stories
What's On

Fox ESS launches AirGate in Australia, simplifying whole-home backup

September 10, 2026

Charges not laid against RCMP officer who broke woman’s arm in Saskatchewan

September 10, 2026

Slack can now vibe-code interactive charts and reports inside chats

September 10, 2026

“A Decade in the Field”|GUILD Gem Laboratories Reflects on 10 Years of Field Gemology Research in Bangkok

September 10, 2026

Solving the “Board Doesn’t Fit the Enclosure” Problem: RapidDirect Integrates PCB and Mechanical Manufacturing

September 10, 2026
Facebook X (Twitter) Instagram
Finance Pro
Facebook X (Twitter) Instagram
Daily Guardian
Subscribe
  • Home
  • News
  • Politics
  • Business
  • Entertainment
  • Lifestyle
  • Health
  • Sports
  • Technology
  • Climate
  • Auto
  • Travel
  • Web Stories
Daily Guardian
Home » The Next AI Infrastructure Challenge Is Before the First Token
Press Release

The Next AI Infrastructure Challenge Is Before the First Token

By News RoomSeptember 10, 20266 Mins Read
The Next AI Infrastructure Challenge Is Before the First Token
Share
Facebook Twitter LinkedIn Pinterest Email

As the Industry Separates Prefill from Decode, Lumai Says the Next Step Is to Rethink the Compute Architecture Powering Each Stage

Lumai Iris Nova

Lumai’s Iris Nova is purpose-built to accelerate AI prefill workloads, freeing GPU capacity for token generation and making existing infrastructure more productive without requiring wholesale replacement. In Lumai testing, Iris Nova ran billion-parameter LLMs in real time and demonstrated approximately 10× more compute per watt than GPU-based equivalents on prefill workloads.

OXFORD, United Kingdom, Sept. 10, 2026 (GLOBE NEWSWIRE) — AI infrastructure has spent the past several years optimizing for the moment a model generates an answer. But as AI applications evolve, the harder problem is increasingly likely to occur before even the first token is generated.

Frontier AI companies project a roughly 1,000x increase in effective compute demand over the next five years. Delivering this using conventional digital accelerators would require an estimated $100 trillion in infrastructure investment and around 1,000 GW of additional electrical capacity. Data centers have limited power budgets, so improving how AI infrastructure addresses incoming tasks matters when every watt counts. Understanding that AI inference is not a single workload helps identify opportunities for optimization.

Prefill, which processes the input context before generation begins, and decode, which generates output tokens, have fundamentally different computational characteristics. Prefill is dominated by highly parallel matrix multiplication and is compute-bound, while decode is driven more heavily by memory bandwidth and data movement. As leading AI infrastructure providers move toward disaggregating prefill and decode into separate infrastructure pools, Lumai sees the next logical step as specializing the hardware powering each stage.

“If the workloads are fundamentally different, it makes sense to stop asking the same hardware to do both jobs,” said Phil Burr, Head of Product at Lumai. “Disaggregating prefill and decode is an important step. But the bigger opportunity is to match the compute architecture to the workload.”

The Next Bottleneck May Be Processing Context, Not Generating Tokens

Much of the AI infrastructure conversation has focused on generation: how quickly can a model produce the next token?

But AI applications are changing the economics of inference. Longer context windows, retrieval-augmented generation, multimodal inputs and agentic workflows are increasing the information that must be processed before token generation begins. A single user request can now trigger multiple model calls, each carrying accumulated context from everything that came before it. That makes prefill as important as decode. Each long prompt, retrieved knowledge set, codebase, document collection, or multimodal context must first be processed before the model can generate its response.

Today, a growing share of inference compute runs on hardware built to do two jobs adequately, rather than one well. As context length grows or an agentic workflow adds another interaction or call, more of the available power is consumed by prefill, widening the gap between the inference capacity a watt could deliver and what it actually delivers. The problem can grow unnoticed until it becomes large enough to constrain both infrastructure capacity and economics.

Prefill, therefore, sits directly on the critical path for time-to-first-token and makes its efficiency increasingly important to the overall economics of inference.

The Cost of Processing Context

More of the work is happening before the first token is generated. At scale, this creates three infrastructure challenges.

1. Power becomes the binding constraint. Data centers effectively convert a fixed megawatt budget into inference capacity. The efficiency of the prefill tier determines the quantity of useful compute that power can deliver before a model generates its first token.

2. GPU capacity ends up stranded. GPUs are highly capable inference processors, but as the same fleet absorbs more prefill computation, that capacity increasingly processes context instead of generating tokens – for which GPUs are better suited.

3. Inference economics becomes a constraint. As context grows and agentic applications trigger more model calls, the cost of processing that context can make long-context and agentic applications increasingly difficult to price competitively. That can ultimately limit not only the economics of existing applications, but what will get built in the first place.

“The real issue is what happens at scale,” said Burr. “Power is fixed, GPU capacity is finite, and every interaction adds more context to process. We have gotten to the point where the economics of inference comes down to the efficiency of the hardware running it.”

Energy Efficiency Becomes Increasingly Important

As AI demand grows, power is increasingly constraining how much inference infrastructure data centers can deploy.

Prefill requires large amounts of highly parallel matrix computation, making its energy efficiency increasingly important. Running prefill on general-purpose GPUs uses capacity that could otherwise support token generation in the decode stage, where the GPU’s memory bandwidth is better utilized. Separating the workloads creates an opportunity to improve both GPU utilization and energy efficiency – and to consider compute architectures designed specifically for prefill.

Lumai Uses Light to Address the Prefill Problem

Lumai Iris Nova uses light rather than electricity to perform the matrix multiplications that define the prefill workload, completing each vector-matrix multiplication in a single optical cycle.

The significance is not merely that optical computing can perform matrix multiplication. It is that optical computing creates the possibility of treating prefill as a purpose-built infrastructure workload rather than running it on general-purpose silicon.

With prefill more efficient, longer context windows and increasingly complex agentic workflows can become more practical and financially viable without proportionally increasing compute and power requirements.

Moving prefill onto purpose-built hardware frees GPU capacity for token generation, enabling existing infrastructure to be more productive, without wholesale replacement. In Lumai testing, Iris Nova ran billion-parameter LLMs in real time and demonstrated approximately 10× more compute per watt than GPU-based equivalents on prefill workloads.

“Our view is not that every accelerator needs to be replaced – it is that we should stop asking every accelerator to solve every problem,” said Burr. “The future AI data center will be increasingly heterogeneous, with different technologies optimized for different stages of inference.”

For a more detailed view on prefill, download the Lumai white paper “The Half of AI Inference Nobody Had Optimized”

To learn more about Lumai’s groundbreaking optical AI technology, visit lumai.ai. To request an evaluation of the Lumai Iris Nova Inference Server, visit lumai.ai/eval.

About Lumai

Lumai, the optical compute company, is building the next-generation AI infrastructure for the Inference Era. Spun out of world-leading optics research at the University of Oxford in 2021, Lumai’s mission is to unlock sustainable intelligence at global scale – delivering materially faster inference, significantly higher execution efficiency, and up to 90% lower energy consumption than conventional GPU architecture.

Lumai is the recipient of the Falling Walls Award for Science Breakthrough of the Year 2025. The company was part of Intel Ignite’s first London cohort and won ‘Best Overall Technology’ at the OCP Future Technologies Symposium. Lumai’s CEO, Dr. Xianxin Guo, is an alumnus of the Royal Academy of Engineering’s prestigious Shott Accelerator program – and CTO Dr. James Spall was recognized in the 2025 Photonics 100.

For more information, visit lumai.ai or follow the company on LinkedIn.

Media Contact:
Stephanie Olsen
Lages & Associates
(949) 453-8080
[email protected]

A photo accompanying this announcement is available at https://www.globenewswire.com/NewsRoom/AttachmentNg/396a17e8-a275-4d2f-bef0-2b218aaa3955

Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Keep Reading

Fox ESS launches AirGate in Australia, simplifying whole-home backup

“A Decade in the Field”|GUILD Gem Laboratories Reflects on 10 Years of Field Gemology Research in Bangkok

Solving the “Board Doesn’t Fit the Enclosure” Problem: RapidDirect Integrates PCB and Mechanical Manufacturing

Find Your Further // iPhone 18 and Urban Armor Gear

Industry leaders gather in Stuttgart to explore how AI, Modular Storage and DC Coupling Are Redefining Renewable Energy Economics

BookBreak Announces First Annual Culture of Reading Virtual Conference

2XO Introduces “Rapper’s Delight” – The Hip-Hop Blend, a Collaboration Between Renowned Blender Dixon Dedman and The Sugarhill Gang

ROHM’s New High Anti-Surge Chip Resistors SDR01 Series Achieve 0.33W Rated Power in the 0402 Size

Two Best-Selling Favorites Take Center Stage: NEXA Wraps Up a Successful CHAMPS Trade Show 2026

Editors Picks

Charges not laid against RCMP officer who broke woman’s arm in Saskatchewan

September 10, 2026

Slack can now vibe-code interactive charts and reports inside chats

September 10, 2026

“A Decade in the Field”|GUILD Gem Laboratories Reflects on 10 Years of Field Gemology Research in Bangkok

September 10, 2026

Solving the “Board Doesn’t Fit the Enclosure” Problem: RapidDirect Integrates PCB and Mechanical Manufacturing

September 10, 2026

Latest News

Find Your Further // iPhone 18 and Urban Armor Gear

September 10, 2026

Industry leaders gather in Stuttgart to explore how AI, Modular Storage and DC Coupling Are Redefining Renewable Energy Economics

September 10, 2026

BookBreak Announces First Annual Culture of Reading Virtual Conference

September 10, 2026
Facebook X (Twitter) Pinterest TikTok Instagram
© 2026 Daily Guardian Canada. All Rights Reserved.
  • Privacy Policy
  • Terms
  • Advertise
  • Contact

Type above and press Enter to search. Press Esc to cancel.

Go to mobile version