The 15-Cent AI Query: Broadcom's Paul Turner on Rebuilding Enterprise AI Economics
A Six Five study found an eight-to-one cost gap between running an AI agent through a leading frontier model and running it on modest on-prem hardware. Paul Turner, Chief Product Officer of the VMware Cloud Foundation Division at Broadcom, joins Daniel Newman and Patrick Moorhead at VMware Explore 2026 to explain how VMware's AI Factory turns GPU provisioning, model management, and security into a repeatable path from infrastructure to production AI.
More than half of enterprise IT leaders say they need to run AI workloads on a private cloud, and setup complexity has been the main obstacle.
At VMware Explore 2026 in Las Vegas, Patrick Moorhead and Daniel Newman welcomed Paul Turner, Chief Product Officer of the VMware Cloud Foundation Division at Broadcom, to discuss what it takes to move enterprise AI from pilot projects into production.
In a recent Broadcom survey of 1,800 IT professionals, 56% said they need to run AI workloads on a private cloud, citing security and cost as the reasons why. Turner ties that cost math directly to what customers are telling Broadcom, pointing to setup complexity as a barrier that prevented many of those customers from acting on that need.
VMware's AI Factory approach automates GPU provisioning, networking, storage, and compute configuration into what Turner calls a repeatable path from metal to model, using a Kubernetes substrate and an open model registry to let customers swap or share models as workloads change.
Key Takeaways:
🔹 VCF 9.1 uses open-source LLM-D and vLLM components to let customers swap models without rebuilding infrastructure. A built-in registry based on Harbor gives IT teams a governed catalog of models to draw from.
🔹 Customers are actively moving from vSphere-based deployments to full VCF deployments. Turner says that shift is what unlocks private AI at scale, bringing virtualized storage, networking, compute, and Kubernetes together as a baseline.
🔹 Security is shifting from perimeter-based to intrinsic and agent-level. Turner points to an agentic attack on Hugging Face that ran more than 100,000 autonomous operations over several days as the kind of threat that requires wrapping every agent and application in its own controls.
🔹 Model sizing is becoming a bigger lever than model scale. Turner expects more verticalized, right-sized models this year as enterprises match model capability to the actual task.
Turner frames that shift as evidence that AI infrastructure decisions are now landing on the CFO's desk alongside the technical team's.
For more Six Five coverage from VMware Explore 2026, visit sixfivemedia.com.
Watch the full video at sixfivemedia.com, and subscribe to our YouTube channel so you never miss an episode.
Disclaimer: Six Five Media is for information and entertainment purposes only. Over the course of this video, we may discuss companies that are publicly traded, and we may reference their equity share prices. Nothing discussed during this webcast should be considered investment advice or a recommendation to buy or sell any security. We are not investment advisors, and you should not rely on this content as financial advice. Six Five Media collaborates with technology companies and industry leaders to produce research-driven interviews and multimedia programming for enterprise technology audiences.
Paul Turner:
What we've done with the VMware AI factory, we've made it very easy for them to go from effectively metal to model. We basically look at how do you provision, how do you provide the GPUs, how do you pull them, how do you provide the configuration of the networking, the storage, the compute that you need associated with that.
Patrick Moorhead:
The Six Five is on the road here in Las Vegas for VMware Explorer 2026. It's been a great event so far, really hitting some hot buttons for our community, Daniel. And that's how to move, how if you're an enterprise, do you move into scale and production with a lot of these challenges around cost? Like how do you protect your IP? How do you secure your data? How do I flex between on-prem and the cloud?
Daniel Newman:
Yeah there's a lot going on here. And of course we've watched the world very very rapidly go from trying to use a I at any cost to trying to use a I at a cost that they can afford. And even more importantly a cost they can afford while getting that productivity and efficiency in those gains that have been promised. The split between those that love a I and those that doubt a I are the ones that are driving ROI today.
Patrick Moorhead:
Yeah, and I can't imagine a better person to have this conversation up here on the stage than Paul Turner, Chief Product Officer of VCF at VMware. Welcome back to the show.
Paul Turner:
Thank you. Delighted to be back here. Yeah, we had a great conversation.
Daniel Newman:
All right, so you heard the setup, right? I mean, you're working, you build the product, but you're also talking to the customers. And you know, we use this, it's kind of becoming maybe a little bit of cliche to say like, you know, experimentation production. It feels like we've been saying this for like three years, right? I know. But I would still say we're pretty early with AI. And I think most enterprises are still trying to bring those workloads to production at scale. And then you've got public and private, right? You're very focused on those that are building in private environments. Talk a little bit about what customers are doing to succeed in driving those workloads to production in the private cloud. And what are they doing that maybe in the last year under your leadership or under the product development, what can they do now that's different?
Paul Turner:
Well, I think you're right. I think some customers are moving much quicker than other customers. And we learn a lot from those kind of lighthouse customers. One of the things we shared at this event is that we're actually delivering our private AI factory, and the VMware private AI factory. The reason for that, that AI factory, is because our customers see the need for private cloud hosted models. We did a survey out on 1,800 IT professionals, and of those IT professionals, 56% of them have actually said they need to run their AI workloads on a private cloud. And the reason for them running it on the private cloud is, surprise, surprise, security and cost. No rocket science there. But the challenge was just too difficult for them to set up. So what we've done with the VMware AI factory, we've made it very easy for them to go from effectively metal to model. We basically look at how do you provision, how do you provide the GPUs, how do you pull them, how do you provide the configuration of the networking, the storage, the compute that you need associated with that. How do you, of course, Kubernetes is required because that's how you're going to run all of your AI models and applications. So you deliver that, you deliver model, then as a service, you're able to allocate the models out in a dynamic way, out onto the GPU pools, you're able to share the models. And then the customer can use their you know, NVIDIA NEMO framework that they want, you know, or the NBIE, the NVIDIA AI enterprise kind of framework, or they can use their open source frameworks, or they can use AMD's governed set of AI interfaces all on top of the shared models. So I think customers are getting real. We've learned from our customer and install base. We learn that it's too difficult, too hard, make it easier. That's what we're doing with the AI factory. And then it's about enablement of the model vendors, which we're seeing customers choose models. And they're choosing not just the common models that are, you know, by the biggest vendors that you know in the AI space, but they are choosing sometimes the Chinese models. They are looking at Z.I.A.I. They are looking at other options that are in the market. And that's pretty interesting.
Patrick Moorhead:
Yeah, it's interesting. You ask enterprises about, hey, is this all about cost per token? They're like, no, it's about getting the right outcome at the right price and having it secured. And it sounds like that's exactly what VCF 9.1 does and whether they want to ship between infrastructures, whether it's AMD and NVIDIA. different CPU choices, but different models, like you said. In fact, a study that we just announced today clearly showed the difference between the intelligence and the cost of the different models, including the Chinese models, but also Frontier through CLI. And it was amazing. I think on a certain on a certain agent it was a dollar twenty six to get the right answer through Frontier CLI. It was 15 cents on prem with the least least least piece of hardware. And that definitely clicks the cost button for sure. But talk a little bit. How does VCF nine point one do this. Is it. you know, the magic portal, is it, I like to call auto magic with AI, where it just, you tell it what to do and it automatically configures?
Paul Turner:
Yeah, well, I think, first off, I mean, just it's interesting, the data point, I was trying to do my math in my head, that's an eight to one ratio, right? That's one eighth the cost, 15% of the cost of a normal cloud infrastructure. So people are going to move when it's that dramatic a change. So what do we do? I'll tell you how the AI factory works.
Patrick Moorhead:
It's almost the easy button you talked about in a way.
Paul Turner:
So a lot of it is day zero. How do we actually go and configure the server, wrap it in, we've rapid enablement, we take those GPUs, we're able to pull those GPUs, so we basically effectively virtualize those GPUs, bring them in, they're understood then by our DRS and vMotion capability inside in the platform. We then, importantly, we fully configure a Kubernetes substrate on top of it, because that's what you're actually using to allocate out resources and GPU resources out to tenants. On top of it, we have LLMD and VLLM, so those are two open source components that many people use. We use them so that we can actually go and switch the model, so you've got choice of models and you're able to swap the model out. So you can write size to the model, just like you said, or you also can allocate, you can sub-slice a model or share a model. So those are the two uses for that. And then we have the built-in registry for models, which is based on Harbor. Again, open source that we build, but it's fully governed and supported by us. Now you have a model registry that's available. So all of that difficulty is all done for you, and that's what we do with VMware AI Factory. That's great.
Daniel Newman:
Have you assessed much ease versus the other alternatives? I think I saw some data points during the keynote, but how much faster getting this stuff spun up? Because I think part of it was the data point he gave. Part of it is getting to the point where you can even deploy something, right?
Paul Turner:
Yeah, well, we haven't actually benchmarked it. What we do know is that if we don't make it easy, because we've always done this as VMware, if we don't make it easy, people don't do it. I will say that through experience, we have had private AI for a while. We had made it actually, honestly, a little too difficult for people to use and to consume and to set up and to configure. So we kind of re-pivoted. We learned from our customers and we went and built, I think what's right now, to build much simpler enablement, much more… You made it easier than the cloud.
Patrick Moorhead:
I think your initial benchmark was, hey, let's make it as easy as the cloud. But the cloud is not always easy.
Paul Turner:
The infrastructure in the cloud is just there. So what you need to do is think about, well, what do I need to do? So long as it's a few days to set up. That's good. I now have an infrastructure as a service ready to go. I can run my model of choice against it. And you mentioned about right-sizing models. I think this will be the year of right-sizing models. I don't know if we need to go from 400 billion parameter models to a trillion parameter models to two trillion parameter models. Actually, if you're doing a chatbot, the 10, 15, 17, 20 billion parameter model is going to work just fine. I mean, those are pretty intelligent. So I think we've had this, this industry has kind of gone leaping into the future of, hey, it's got to be run on the cloud and I'm going to throw more and more GPU capacity to build the most trained model. We should get more targeted on use cases. I think you'll see more, this year, more verticalization, more specialist type of models. I think it'll be much more interesting. And then, of course, you see what's NVIDIA and their actions in the market. I think they see it happening too. I think they expect to see more deployments into real enterprises. And that's why they're making some of the moves that they've just announced.
Daniel Newman:
Certainly value in the Frontier and with the leading edge models, but I think using them for so many different use cases. The analogy someone gave me is like, you'd be like taking an F1 car and driving it to get coffee at Starbucks. That's a nice way of putting it. You know, just something, you know, it's like you wouldn't need to, right? You're perfectly fine taking your everyday car.
Paul Turner:
Yeah.
Daniel Newman:
And so I think… You don't have to be in a Trabant either.
Paul Turner:
You can be in a reasonable car. These are pretty good, powerful models. I'm telling you, they're not bad.
Daniel Newman:
Absolutely, and I think that's going to be all about matching the task to the model, to the infrastructure, and getting to that point where, you know, again, you're going to get the rubber sealed. It's not going to be just a technology decision. It's going to be the CFO's desk. It's going to hit the CEO's desk. The intelligence becomes part of the workflow. Yes. And so just, you know, you started to touch on this, and I started to ask a question, and I just kept talking.
Patrick Moorhead:
You're such an analyst.
Daniel Newman:
That's why I'm here. So all these capabilities that you're creating this cloud-like experience, right? The cloud on-prem. You've made more things available. How is that changing the way the teams that you're seeing are deploying and managing and consuming AI? What are you seeing there?
Paul Turner:
So, I mean, I think the biggest thing, right, we deploy this with VCF9. And what we're seeing is customers are actively moving forward from vSphere kind of base deployment actually to VCF deployments. Now, when they've done that, they've suddenly got all the virtualized storage, networking, compute, the Kubernetes service that they need, everything that's enablement pieces. So that's what we're seeing happen is this massive change is actually enabling now basically private AI at scale because they've got the baseline. Our customers have been re-baselining to BCF over the last three years. So it's quite powerful. I think that's what we're going to see. I think as we move forward, if you're going to run this on the private cloud and you're going to have it as a shared infrastructure, you've got to be able to do still the monitoring, the service monitoring, the performance monitoring. You've got to have the dashboards that actually show how your models are doing. And that's why we bring it in as like a primary citizen into the way that we show our operations and reporting and monitoring. So not just can you deploy the model, can you share it out to tenants, they can use it just like they could any other cloud service. They also get the ability to monitor, to report, to performance manage, which is what you also need.
Patrick Moorhead:
So I think it's important to talk in this conversation with security between you and Ram and we talked about cost, we talked about control. I think Ram said control was primary cost, IP protection, but also securing it. And AI has really compressed the entire chain of open source security. in particular, and I'm curious, how should customers rethink their security strategies in a world where you're mixing and matching infrastructures, applications, models, and agents very, very frequently because you're going to have a ton of workloads that you're going to want different capabilities across the sphere?
Paul Turner:
Well, I think, I mean, it's dramatically changed. If you look just in the last six months, what have we seen happen in security? We've seen attacks happen. Right. So we we look at the open AI model was actually used to attack a hugging face. Right. Hugging face the whole show of other models. What did they get access. They got access to customer tokens basically privileges and rights that those customers had. And they got access to the internal network of hugging face. That was all an agentic attack onto the infrastructure. That was more than 100,000 operations that that agent did over a few days, over three, four days. And at the end, it was able to get access to tokens. Now, you cannot protect and run an environment if there's 160,000 autonomous operations and you're trying to actually stop them and break them. It just doesn't work. So you've got to change security to be intrinsic security. So one reason is the threat of A.I. And I hate talking about how good A.I. is and the potential of A.I. When we also talk in that other one it's like black and white it's like left brain right brain. You've got this. Oh well it's also a threat. It is a threat, but it's a threat we can mitigate. But if you're going to deliver agents as well, think about an agent. An agent is an autonomous service. An agent has to have guardrails and controls. An agent has to be kept in effectively a sandbox. They don't have civilizations? They don't. Actually, be careful of that, because the agent might, right? Because as we build out agents, agents are going to have little civilizations that will talk to each other. I know. But you've got to control it.
Daniel Newman:
A lot of hyperbole.
Paul Turner:
Exactly, it's hyperbole, but you've got to control how they're going to talk, who they're going to talk to, what they're going to talk to. And then also, you do that at two layers. You do that within the network layer, which means intrinsic security means you've got to change it from perimeter security to application-based security. You wrap every agent, you wrap every application. You have to do it because of AI threat. You have to do it because you're trying to control agents. So that will happen in the network. And then secondly, you look at how you're developing agents. And you develop agents then in a much more secure, trusted way. In fact, you don't use just IaaS, infrastructure as a service, and let developers build what they want. You've got to use a much more prescriptive way to build agents. That's how we think it will happen.
Daniel Newman:
Well, Paul, I want to thank you so much for joining us here. Of course. It's been a lot of fun. We've got to do this at least once a year. Yeah, we do have to do it. We should do it more often, but great to catch up. Congrats on all the progress. Thank you very much. And thank you, everybody, for being part of this 6.5. We're on the road here at VMware Explorer 2026 in Las Vegas. Check out all of our coverage here at the event and, of course, be part of the 6.5 community. Thanks for tuning in. See you all later.
MORE VIDEOS

Broadcom's Ram Velaga on the AI Economics Pulling Enterprise Workloads Back On-Prem
Enterprise AI economics are pulling workloads back toward on-premises infrastructure, with control over data as the primary driver. Ram Velaga, President of Broadcom's Infrastructure Software Group, joins Patrick Moorhead and Daniel Newman at VMware Explore 2026 to detail how Broadcom is helping customers route AI workloads, govern agents, and rebuild VMware's role underneath the AI buildout.
Claudeforce, AWS's 2 Million GPU Bet, and the Earnings Week That Buried the SaaSpocalypse
Salesforce and Anthropic launch Claudeforce as Marc Benioff and Dario Amodei explain the collaboration together on CNBC, AWS commits to 2 million more NVIDIA GPUs on top of its GTC pledge, plus six new earnings this week from NVIDIA, Salesforce, Synopsys, HP, Everpure, and Marvell test every bear thesis on AI infrastructure and software at once. Patrick Moorhead and Daniel Newman also cover Hot Chips 2026, the NVIDIA-Hugging Face acquisition rumor, and debate whether Anthropic's SaaS reassurances hold up on Ep. 317 of The Six Five Pod.
Marvell's $120B Google Win, Modular's Open Compiler Play, and the Fight Over NVIDIA's Ohio Scale-Back
Patrick Moorhead and Daniel Newman break down Marvell's $120 billion custom-silicon deal with Google, Qualcomm-owned Modular's open-source compiler push at ModCon 2026, and state-level pushback against AI data center construction in Texas and Pennsylvania. The hosts also flip roles to debate whether NVIDIA's reduced Ohio financing guarantee signals cracks in the AI infrastructure buildout.
Other Categories
CYBERSECURITY

Threat Intelligence: Insights on Cybersecurity from Secureworks
Alex Rose from Secureworks joins Shira Rubinoff on the Cybersphere to share his insights on the critical role of threat intelligence in modern cybersecurity efforts, underscoring the importance of proactive, intelligence-driven defense mechanisms.
QUANTUM

Quantum in Action: Insights and Applications with Matt Kinsella
Quantum is no longer a technology of the future; the quantum opportunity is here now. During this keynote conversation, Infleqtion CEO, Matt Kinsella will explore the latest quantum developments and how organizations can best leverage quantum to their advantage.

Accelerating Breakthrough Quantum Applications with Neutral Atoms
Our planet needs major breakthroughs for a more sustainable future and quantum computing promises to provide a path to new solutions in a variety of industry segments. This talk will explore what it takes for quantum computers to be able to solve these significant computational challenges, and will show that the timeline to addressing valuable applications may be sooner than previously thought.

