Broadcom launches VMware Private AI Cloud for on-premises inference
Broadcom introduced VMware Private AI Cloud at VMware Explore to run inference, agent applications and traditional workloads on one private cloud. Its foundation is VMware AI Factory. The company says customers can run more than 150 models on premises. Microsoft's Foundry model router spreads requests by cost, quality or latency and has expanded from 2 regions to 28 for global standard deployments and to 21 for data zone deployments, adding Claude Opus 4.8 and the GPT-5.6 family.
Artificial Intelligence··Night
More than 150 models on premises
Broadcom introduced VMware Private AI Cloud at VMware Explore, a stack that runs inference workloads, agent applications and traditional enterprise workloads together on one private cloud platform. Its software-defined foundation is VMware AI Factory, announced at the same conference. The company says VMware Cloud Foundation 9 lowers hardware cost through NVMe memory tiering and cluster-wide deduplication, and adds token monitoring, multi-tenant model sharing and a metrics dashboard. Customers can run more than 150 open source and commercial models on their own premises, among them Nemotron 3, Gemma 4, cotomi, Qwen 3.7-Max and GLM 5.2. On the security side there is vDefend microsegmentation, the Avi Load Balancer web application firewall, and TrueSource, a line of verifiably built open source packages. The cost and performance descriptions are the company's own and were released with the Las Vegas conference.[1]
Foundry router from 2 regions to 28
The Foundry model router, which spreads requests across different models by cost, quality or latency, has expanded from 2 regions to 28 for global standard deployments and to 21 for data zone deployments. Its model pool was refreshed at the same time. Anthropic's Claude Opus 4.8 and the GPT-5.6 family were added, while gpt-5-chat, gpt-5.2-chat, gpt-5.3-chat and DeepSeek-V3.1 were removed after reaching end of life. Microsoft says the endpoint stays stable, so teams do not need to redeploy the router to receive the update. An Azure MVP says API stability and behavioural stability are separate things, and that new models can change response patterns. Claude models must still be deployed separately to the same Foundry account before the router can reach them.[2]
A private-cloud stack and an Azure router
VMware Private AI Cloud keeps more than 150 models, inference workloads and agent applications on the customer's premises in a private cloud stack. The Foundry model router spreads requests by cost, quality or latency across 28 global standard regions and has added Claude Opus 4.8 and the GPT-5.6 family to its pool. The two may both sit as enterprise control planes in front of many models; the VMware stack may instead remain an on-premises private cloud while the Foundry router stays an Azure endpoint. Claude models still have to be deployed separately to the same Foundry account before that router can reach them.[1], [2]
Related columns
For more information on this topic, you can read the related columns.