ComputerWeekly

Why the next AI race will be won at the inference layer


As enterprises transition generative artificial intelligence (GenAI) from pilot projects to production systems, attention is shifting from training large language models (LLMs) to managing the growing cost and complexity of inference.

According to a report from Core42, the next competitive advantage in enterprise AI will come not from deploying the biggest models on the most powerful hardware, but from intelligently matching every workload to the most appropriate combination of model, accelerator and deployment environment.

The report, Efficiency wins the inference era, argues that relying on a single hardware ecosystem or defaulting to the largest available model can significantly increase infrastructure costs while reducing overall performance. Instead, enterprises should adopt workload-aware AI architectures capable of dynamically routing inference requests according to latency, throughput, utilisation, governance and cost requirements.

The findings reflect a broader shift in enterprise AI priorities as organisations move beyond experimentation towards large-scale deployments that must balance performance with operational efficiency.

Core42 believes the challenge will become even more pronounced as agentic AI gains traction. Unlike conventional AI applications, where one user request typically results in a single inference event, autonomous AI agents can generate multiple model calls while reasoning through tasks, accessing external tools and executing multi-step workflows.

“Consumption stops scaling with headcount and starts scaling with autonomy,” said Raghu Chakravarthi, chief product and technology officer at Core42. “That is exactly the dynamic that produces bills far larger than anyone modelled at the pilot stage.”

He added that enterprises must begin planning infrastructure around complete AI workflows rather than simply estimating the number of users or prompts: “The most common planning mistake is capacity math based on user counts rather than inference events per task. A proof of concept that looks trivial with a handful of users can become expensive very quickly once agents are running multi-step workflows at scale.”

Chakravarthi also warns against postponing governance and cost controls until after deployment. “Spend caps, budgets, alerts and routing logic are often treated as things to add later, by then the consumption has already compounded.”

Rather than treating infrastructure as a fixed environment, Core42 advocates making infrastructure selection an active part of AI orchestration. AI workloads vary considerably. Interactive copilots require extremely low latency, while document summarisation, indexing and classification workloads often prioritise throughput and efficient resource utilisation. Likewise, many routine enterprise tasks can be handled effectively by smaller models, reserving frontier models for applications where higher reasoning capabilities justify their additional cost.

“Multi-silicon routing allows infrastructure to become a workload-placement decision,” said Chakravarthi. “Efficiency in the inference era comes from treating hardware diversity as an economic asset rather than a procurement inconvenience.”

Core42’s Compass platform implements this approach by routing workloads across more than 60 open and proprietary AI models while supporting multiple accelerator architectures, including Nvidia, AMD, Qualcomm and Cerebras. Routing decisions consider latency sensitivity, throughput requirements, utilisation, governance policies and overall cost before determining the optimal execution path.

The platform also balances deployment across cloud, on-premise, shared and sovereign environments according to operational requirements.

Production-scale optimisation

Core42 said Compass is already operating at production scale, processing more than seven million API requests and over 100 billion tokens each week while offering a 99.5% availability commitment.

The company reports throughput improvements of up to 20 times on its fastest Cerebras inference path, although Chakravarthi stressed that this should not be interpreted as a universal cost reduction.

Instead, the primary benefit comes from reducing the amount of premium compute consumed for each completed business task. “The goal is to deliver the required business outcome at the right speed and cost,” he said.

For organisations deploying AI across sectors such as government, financial services and healthcare, intelligent workload routing can improve responsiveness for latency-sensitive applications while increasing infrastructure utilisation for batch processing workloads.

The company evaluates infrastructure efficiency using a metric it describes as “tokens per second per dollar”, reflecting the balance between performance and operational cost rather than focusing solely on raw model capability.

Sovereignty becomes an operational advantage

AI sovereignty is primarily a regulatory requirement. “When sovereignty is added after an AI system has already been designed, organisations often need to rebuild data flows, introduce separate monitoring systems or create additional approval processes,” said Chakravarthi.

By integrating governance into the AI platform itself, organisations can automatically determine where workloads execute, which models are permitted, who can access data and how usage is monitored.

For organisations operating in highly regulated sectors across the Gulf, Chakravarthi argued that integrating sovereignty with workload optimisation avoids treating governance and efficiency as competing priorities.

“The same observability that shows finance teams which department generated a cost can show governance teams which model was used, where the request was processed and under which policy,” he said.

Chakravarthi believes the Middle East is following a different AI infrastructure trajectory from many global markets. “In many regions, organisations adopted AI first and deployed data residency and governance later,” he said. “In the Gulf, sovereign control has been a starting requirement.”

He stated that building sovereign AI infrastructure from the outset allows organisations to avoid costly redesigns while supporting AI deployment at scale. However, he cautions that sovereign infrastructure must also preserve technological choice.

“A sovereign platform still has to offer diverse silicon and a broad model library,” he said. “Otherwise it simply trades one form of lock-in for another.”

As enterprise AI enters what Core42 describes as the inference era, the company believes competitive advantage will increasingly depend on how efficiently organisations execute AI workloads rather than simply which models they deploy.



Source link