Realise the full potential of AI with inference economics and Penguin Solutions

Organisations transitioning from AI training to inference can control costs more effectively by moving from cloud-only to hybrid environments and locating workloads based on economic, performance and governance criteria, according to Penguin Solution

Realise the full potential of AI with inference economics and Penguin Solutions

Realise the full potential of AI with inference economics and Penguin Solutions

Organisations transitioning from AI training to inference can control costs more effectively by moving from cloud-only to hybrid environments and locating workloads based on economic, performance and governance criteria, according to Penguin Solutions.




Realise the full potential of AI with inference economics and Penguin Solutions










 

Exploding, unpredictable AI costs

Inference costs compound with every user, prompt and query, meaning pay per token costs can climb quickly when sustained inference workloads run under variable, consumption-based cloud pricing models. “This may lead to a runaway bill when inference goes live across different teams,’ says Lin Hoe Foong, Vice President and Managing Director, APeJ, Penguin Solutions. “As AI initiatives mature, multi-step agentic AI workflows and massive context windows can create severe pressure on memory resources and cause non-linear cost spikes, making cost prediction impossible and compromising AI operationalisation programs.

“It becomes hard for organisations to defend sustained investment in AI when they can’t reliably forecast the quarterly inference infrastructure bill.”

Compounding these problems at organisations in Australia is indecision and inexperience in selecting the right environment–on-premises, hybrid or public cloud–to scale key AI use cases, says Foong.

“Scaling from successful pilots to production means consolidating and expanding AI workloads into centralised, shared infrastructure and, without the right platform and operational expertise, that transition can stall.”

In practice, he adds, this centralised infrastructure brings together GPU compute clusters, high-speed interconnects, shared high-performance storage, and cluster management software into a single, production-ready environment. Memory-optimised systems can also be added to ease inference bottlenecks.

“The best-practice question isn’t just what the infrastructure comprises, but who runs it,” says Foong. “Typically, the enterprise IT or AI platform team owns and governs the environment, while a specialist partner handles the heavy lifting of design, build, deployment, and ongoing management.

“At Penguin Solutions, we take on the complexity of designing, building, deploying, and managing AI factories, so customer teams retain control while focusing on outcomes.”

Leveraging inference economics to overcome challenges

According to Penguin Solutions, leveraging inference economics is key to overcoming cost and infrastructure impediments to AI take-up. “This means understanding that the true cost per token is not set by a cloud provider–it is determined by how organisations engineer compute, memory and networking to combine for their specific workloads,” says Foong. “Well designed infrastructure enables organisations to optimise token throughput and, when done at scale, reduce costs dramatically.” 

“For sustained, high-utilisation workloads, the strategic move may be to complement public cloud with dedicated capacity, swapping volatile operating expenses for fixed, amortised infrastructure and securing a predictable total cost of ownership.

“This is where hybrid by design matters. Public cloud remains valuable for elastic, experimental, or lower-volume workloads, while sustained production inference, particularly at scale, often justifies dedicated infrastructure.

The point is to match each workload to the environment that best balances cost, performance, and control.”

A recent TCO report prepared for Penguin Solutions pointed out that in modelled, high-utilisation scenarios, moving sustained inference workloads to dedicated hybrid or on-premises clusters could deliver four to six times lower costs over five years than running those workloads on variable public clouds. 

Tracking the three levers of AI success

Leveraging inference economics to power AI efficiently and cost-effectively means tracking what Foong describes as the three key metrics that act as direct economic levers to AI success: Time to First Token (TTFT), Time Per Output Token (TPOT), and Token Throughput (TPS).

“TTFT is a measure of the time between the first prompt to the first response, while TPOT is the average time between the generation of subsequent tokens,” says Foong. “TTFT is a responsiveness metric that ties directly to user experience, while TPOT is a measure of the sharing speed and fluidity of the end-user experience in different AI use cases.

“TPS measures the volume of tokens that an organisation’s infrastructure processes under a concurrent load and works as a scalability and cost efficiency metric.

Together with automated payload routing, these metrics connect technical design choices directly to user experience, concurrency capacity, and cost per inference. “Tracking all these metrics provides control by enabling an organisation to translate technical performance into financial outcomes and manage AI spend with the same rigour applied to any other operational costs,” says Foong.

Australian awareness patchy but positive

Foong describes awareness of these issues in the Australian marketplace as rising but uneven. “Australian leaders are seeing their cloud bills rising but not a lot of them are connecting these to inference economics,” he says. “Many organisations track traditional cloud spend well, but TTFT, TPOT and even TPS rarely appear on executive dashboards.

“So while there is a lingering assumption that cloud is always cheaper and more flexible, our data shows that while this equation holds for variable low-volume workloads, it flips when it comes to sustained production, particularly at scale.

“The encouraging thing is that Australian executives are very pragmatic and data driven, so once they see the number, the conversation shifts very quickly.”

Building advanced AI infrastructures for organisations  

Penguin Solutions’ AI infrastructure capabilities are exemplified by its work with organisations across a range of industries.

For example, the vendor has collaborated with Korea’s Ministry of Science and Technology and SK Telecom to build one of the world’s most advanced sovereign AI infrastructures.

The Haein platform integrates more than 1,000 NVIDIA Blackwell GPUs into a single cluster to enable SK Telecom to provide GPU as a Service to Korean developers. The project delivers a stable and scalable computing environment optimised for large-scale AI model training and inference workloads. Penguin Solutions maintains and evolves the platform, providing enterprise-grade reliability, predictable performance, clear upgrade paths and help desk support services.

“At Penguin Solutions, we believe we can help Australian enterprises turn AI ambition into operational reality,” concludes Foong. “We take on the complexity of designing, building, deploying and managing AI factories so that organisation teams can focus on outcomes.”



About Author

What do you feel about this?

Subscribe To InfoSec Today News

You have successfully subscribed to the newsletter

There was an error while trying to send your request. Please try again.

World Wide Crypto will use the information you provide on this form to be in touch with you and to provide updates and marketing.