Home

Broadcom's Ram Velaga on the AI Economics Pulling Enterprise Workloads Back On-Prem

Broadcom's Ram Velaga on the AI Economics Pulling Enterprise Workloads Back On-Prem

Enterprise AI economics are pulling workloads back toward on-premises infrastructure, with control over data as the primary driver. Ram Velaga, President of Broadcom's Infrastructure Software Group, joins Patrick Moorhead and Daniel Newman at VMware Explore 2026 to detail how Broadcom is helping customers route AI workloads, govern agents, and rebuild VMware's role underneath the AI buildout.

Enterprise AI budgets are colliding with a cost problem few companies planned for. As token consumption scales past pilot programs, rising hardware costs are pushing enterprises to reconsider where AI workloads actually belong.

At VMware Explore 2026 in Las Vegas, Patrick Moorhead and Daniel Newman welcomed Ram Velaga, President of Broadcom's Infrastructure Software Group, for a look at what Broadcom is hearing from customers navigating the next phase of enterprise AI.

Velaga unpacks the current shift that began in early 2026, when large language models became reliable enough to accelerate production-grade coding at scale. That change arrived alongside a reversal in cloud planning. Enterprises that once expected 70-80% of workloads to move to the cloud are now keeping the majority on-prem, citing data ownership and token cost as the deciding factors.

Velaga points to one enterprise customer whose two-year infrastructure review concluded with a commitment to build tens of megawatts of new data center capacity, a decision he frames as proof that control over infrastructure is now driving enterprise investment. He also repositions VMware's role in that buildout. The company name signals virtual machines. Velaga notes that 90% of today's containers already run on VMware VMs, placing the platform underneath an AI infrastructure shift many assumed would bypass it.

Key Takeaways:
🔹 An internal AI routing layer sends simpler tasks to in-house, open-weight models. Broadcom reserves frontier models for the work that actually requires them, treating token cost as a design constraint from the start.
🔹 AgentMinder assigns every AI agent an identity and an access policy. The tool is part of Broadcom's broader push to govern what agents can see and do inside enterprise environments.
🔹 Micro-segmentation and continuous patching are now baseline infrastructure requirements. Customers are asking for security built directly into the platform, along with the ability to patch systems without moving workloads.
🔹 VMware Kubernetes Service lets containers run on a virtualized layer instead of bare metal. Velaga points to rising memory and CPU costs as the reason virtualization remains essential even as AI workloads shift toward containers.
🔹 Tanzu now functions as Broadcom's platform-as-a-service layer. Velaga compares the addition to the way public cloud providers pair infrastructure services with a platform layer on top.

Velaga frames that shift as a multi-year infrastructure commitment now central to enterprise AI strategy.

For more Six Five coverage from VMware Explore 2026, visit sixfivemedia.com.

Watch the full video at sixfivemedia.com, and subscribe to our YouTube channel so you never miss an episode.

Disclaimer: Six Five Media is for information and entertainment purposes only. Over the course of this video, we may discuss companies that are publicly traded, and we may reference their equity share prices. Nothing discussed during this webcast should be considered investment advice or a recommendation to buy or sell any security. We are not investment advisors, and you should not rely on this content as financial advice. Six Five Media collaborates with technology companies and industry leaders to produce research-driven interviews and multimedia programming for enterprise technology audiences.

Transcript

Ram Velaga:

People generally would associate, just even by virtue of a name, VM is equal to virtual machines, right? But when you look at AI and all of these new kinds of applications, they're all in the context of how things are getting deployed in containers. So people then think, oh, well, if it's containers, is it on bare metal? The reality is 90% of the containers today, they're actually on our VMs.

Daniel Newman: 

Welcome back to The Six Five. We are on the road in Las Vegas at VMware Explorer 2026. It is awesome. I am here with my bestie, Daniel Newman, and we're here going to talk about a topic that we explore a lot on the show, and that is the challenges that enterprises are having. But we're going to hone in on Frontier AI, the challenges they're having with like token costs. like keeping their alpha or their IP. That's been a huge topic out there today. And the investments required to make that work. The good news is there's a lot of on-prem solutions to help companies do that. And I'm here with Ram Vallega, President of Broadcom's Infrastructure Software Group. Ram, welcome back to The Six Five.

Ram Velaga: 

Hey guys.

Patrick Moorhead: 

Good to see you. Good to see you too. New role, congratulations. I mean, not that new, but. Running those six months is kind of new. Yep. I don't know. You want us. Yes, exactly Well, we've had a number of sit-downs with rom over the years and you know, he's done a lot of different things Yeah, he's you know doing great though. And of course, this was a big challenge real, uh big priority a year ago I I was here sitting, uh with hawk and we had the conversation and you know This has been a big focus of the company put you in a in a really important role to drive one of the biggest priorities that Broadcom has. Everything lately has been all about chip, chip, chip, chip. But running applications, running your business, and the actual agentic tools, a lot of things sit in VMware here across the different business. So let's just start there and talk about you're here, you're spending time with your customers, frontier AI, you can say near frontier AI, they're all trying to get outcomes. They're trying to get away from just token maxing and they're trying to actually invest. What are you hearing from your customers about what their focus is and how they're investing and how are they actually doing in getting that AI into production?

Ram Velaga: 

Yeah, actually, I think there was a turning point sometime towards late last year, early this year, with the whole idea of being able to use these large language models to significantly improve the pace and quality of coding. At about the same time, we also started to see something else happening, which is companies that previously were kind of looking at it and saying we're going to be all in on cloud, they were going to be in a 5 or 10 year journey, but everything is going to be deployed in cloud, 70 to 80% of the workloads. We slowly started to hear, hey, that's not necessarily the case. Cloud has a place that it would play, but for a majority of the more stable workloads, maybe 60%, 70%, 80% of it actually stays on-prem. And that happened along with this whole idea of you're going to start doing more LLM-based coding. And if you also think about it, sometime in March this year, there was a whole, as you call it, a SaaS apocalypse. Yes, thank you very much. And it's happening, right? Because people are starting to say, okay, these LLM companies are also vertically integrating up potentially and disrupting more than just coding, that they were potentially going after vertical businesses. So when you look at all of that, I think it kind of brought to bear a few things. One is, There is a balance between what's in cloud, what's on-prem. You have to be very careful with your data because data may be the key IP and a differentiation for you as an enterprise. You can't just put everything in the cloud. And three, if I'm going to use a lot of these large language model, what's the cost of using this model? So a combination of those three is starting to put a lot of focus on How do they think about infrastructure going forward for the next three to five years where, if I may, not long ago some of this compute and IT was treated as a context to companies. Now they're kind of thinking about it as, hey, if the compute is going to replace humans, it's no longer a context, it's core to my business. And so the idea of keeping things on-prem and then amongst other things, keeping data with you. And then lastly, but definitely not the least, how do I cost optimize this is all coming to bear.

Daniel Newman: 

Yeah. So many things coming together at the same time. And we had four point six that changed the game on development. And even recently it's funny it seemed like we were all token maxing for about 18 months. And then enterprises realized hey you know I've used you know my entire years worth of tokens in one or months. And we're going to have to dial this back. And that combined with the IP protection really seemed to make a shift. So we've outlined the challenge. What are the solutions on a cost side that people are rethinking how to scale their infrastructure.

Ram Velaga: 

Yeah, I think, look, to your point, even Broadcom is a heavy user of tokens. As you can imagine, last five, six months, a lot of us have been using these tokens to actually find vulnerabilities in our software, fix our software, make sure that it is robust enough for any attacks, AI-accelerated attacks, and so on and so forth. When you look at it, it is a significant cost, especially if you're not careful about which model do you use for what task, right? So then when you think about in the context of what we would do as an enterprise, we would say, okay, let's have some kind of a gateway so that we can look at what models we can use for majority of the work, and then what models we would actually initiate for the work that cannot be done with simpler models which are on-prem, potentially on-prem. So that's how we've been thinking about it. When we built our own framework inside Broadcom, if you go talk to our Broadcom CIO, we've developed a framework called Maveneer. And what it does, it's built on a foundation of VCF, which is first you virtualize the compute, the networking, and then the storage. And then on top of that, you have automation and you have cloud. And then sitting on top of that is agent sandbox. And of course, these agents are also then authenticated using something called an agent minder. So an agent actually is associated with policies and access to what it can do, what it cannot do. Then we have this multi-context protocol, you know, gateways, MCP gateways. And, you know, we're kind of using all of that to say what and where can we use, you know, open weight, you know, LLMs in-house and where do we punch out externally to be able to, you know, leverage the more capable frontier models.

Patrick Moorhead: 

It's really interesting. It feels like what his customers are going through is exactly why we stood up Pinnacle, right? We stood that up because everybody wanted to figure out which model for what specific persona or use case on which infrastructure could that token be delivered most efficiently. And so we're seeing the same thing. And almost everything out there is actually just measuring like how many tokens. It's like, does it matter? I only care about the token that actually delivered me value. So you're very focused on working and helping customers figure out how to make VMware part of their AI infrastructure strategy. And as they do that, what are they asking you to deliver them? And when they ask you, how does that sort of impact the way, now that you're leading the charge here, you're thinking about R&D, because the development cycles are so fast now.

Ram Velaga: 

If you look at our journey, people generally would associate, just even by virtue of our name, VM is equal to virtual machines, right? But when you look at AI and all of these new kinds of applications, they're all in the context of how things are getting deployed on containers. So people then think, oh, well, if it's containers, is it on bare metal? The reality is 90% of the containers today, they're actually on our VMs. Now, we look at it and say, OK, so there are containers. Containers, well, they come in and they go, how do you actually manage these containers? How do you scale up and scale down? Which is where we've actually released, for the last couple of years ago, something called VKS, right? You still need a virtualized platform because you cannot run containers economically on bare metal. And you see that with the cost of the hardware going up right now, the cost of memory is going up, the CPUs and everything else. So that's the baseline capabilities that we're doing. Now, above and beyond that, being able to support containers and then we look at the underlying silicon, being able to support any CPU and then GPUs so that you potentially provide that hardware abstraction between the workloads and the CPUs and the GPUs. And then we actually are also looking at it beyond infrastructure, right? Because if you think about a cloud, if you look at somebody, one of these cloud providers, they provide infrastructure as a service, but they also provide platform as a service. That is where we have our Tanzu team now focused on providing platform as a service. The other thing that's happening is security is no longer a separate pillar. Almost any conversations we're having with our customers is security is now embedded deep into the platforms because these LLMs are telling you, you have to do micro-segmentation everywhere. Another one worth actually calling out is customers can no longer take downtime to move workloads in and out of a physical host to deal with applying patches and keeping their code current. They're basically saying the rate at which these patches are coming in, we have to be able to patch while not moving workloads around. So a lot of our focus on R&D is providing that seamless upgrade paths and integrated security so that you can actually have this infrastructure running just like they would expect external clouds to be running is number one priority. And number two, obviously, is can we accelerate the rate of products coming out? And we're kind of seeing our engineering teams using more and more tokens and compressing timelines for roadmaps. So there's some goodness coming out of it too.

Daniel Newman: 

Sourav, great conversation. A lot of moving parts and it seems like the conversation shifts every year to something new. But when you look back at what we were talking about with enterprises, some of that is actually pretty stable in what they're looking for, right? They're looking for Governance, they're looking for security. They want to get results of their investment. And quite frankly, if you go back in the last 30 years of IT, that's been the mantra. But with all these things moving so quickly, how do customers turn AI into something that supports a scalable foundation, that supports frontier AI and agentic AI?

Ram Velaga: 

Look, I think coming from a world in my previous role where we served the hyperscalers versus coming into the enterprise, as much as we would like to see the enterprise move much faster, there is a timeline and a journey that you have to kind of follow. where a lot of these enterprises and customers are in their journey today is they're kind of coming to the realization that they have to get an equivalent of a public cloud experience on-prem. They're working towards making sure they're able to get that, which is largely first automation operations on an infrastructure that is virtualized, not just compute, but also storage and networking with integrated security. Once they have that as a baseline platform, then they're looking at it and saying, OK, how do I have my agents create sandboxes for them and kind of secure them? I think this is a journey that's going to happen for the next year or so. And while they are doing that, then they're kind of starting to look at and saying, OK, what are my economics with regards to where do I punch out, where do I run it in-house? But I will tell you, generally, most of them that we speak with, there is a conviction there that more of it is going to be on-prem to deploy this infrastructure, both for just regular cloud as well as AI-enabled on-prem cloud.

Patrick Moorhead: 

Yeah, we're hearing the same thing for sure. So just quickly, what do you think is driving that conviction, right? Of course, you know, we've had this debate through the pre AI era about cloud and what would move on-prem and clearly you and Hawk and the leadership team here believe that this on-prem thing was substantial heading into this AI transformation, but is the conviction dollar driven? Is it data driven? Is it, you know, security? Like, do you think in the end, or I mean, it's easy to say all of these things, but like, what's going to really drive that? Because I think right now there's still a lot of opportunity but also a lot of speculation and cloud of course like you said they innovate very fast.

Ram Velaga: 

So actually, I was surprised. One large institution that I was talking to, they were actually planning this for almost two plus years ago. And they were looking at it and saying, hey, look, AI is coming, and we are in a regulated industry. Our journey to cloud is not what we expected it to be, which we thought was going to be five years ago. we have to own our infrastructure and we have to start thinking about it as how do you place applications together, which is which applications have to be co-located for latency reasons, which applications have to be further apart for disaster recovery reasons, how do I secure the data that I need, and they've actually done a lot of work actually ground up from the first principles. Based on that, they've concluded in the next few years, they're going to be building many tens of megawatts of data centers. This is an enterprise. And they have a sense for how these workloads are going to be placed, what machines that they're going to be using in there, what fiber is going to be connecting all of this stuff together. But all of this just goes back to the belief that In the future, there's going to be far more compute being used by the enterprises than they have envisioned they were going to be using a few years ago. With that in mind, they no longer can think of this as, oh, I can just let somebody else do this for me. Now it's more about, I've got to do this myself. So cost almost is a secondary observation to, I've got to own this. And then they're also confident that they will actually have a better cost structure than somebody else doing it with their own CAPEX requirements and margin profile.

Patrick Moorhead: 

So it's control. Control and cost. Control first. Yeah, cost second. Ram, thank you so much for spending some time with us here at VMware Explorer. I know you probably have quite a few meetings, a lot going on here, so have a great rest of your event.

Ram Velaga: 

Yeah, thank you.

Patrick Moorhead: 

Thank you for hosting us. We appreciate it. Thank you. Thank you, everybody, for checking out our coverage here at VMware Explorer 2026 here in Las Vegas. Check out all of our coverage here on The Six Five. Appreciate you. Talk to you all soon.

MORE VIDEOS

The 15-Cent AI Query: Broadcom's Paul Turner on Rebuilding Enterprise AI Economics

A Six Five study found an eight-to-one cost gap between running an AI agent through a leading frontier model and running it on modest on-prem hardware. Paul Turner, Chief Product Officer of the VMware Cloud Foundation Division at Broadcom, joins Daniel Newman and Patrick Moorhead at VMware Explore 2026 to explain how VMware's AI Factory turns GPU provisioning, model management, and security into a repeatable path from infrastructure to production AI.

Claudeforce, AWS's 2 Million GPU Bet, and the Earnings Week That Buried the SaaSpocalypse

Salesforce and Anthropic launch Claudeforce as Marc Benioff and Dario Amodei explain the collaboration together on CNBC, AWS commits to 2 million more NVIDIA GPUs on top of its GTC pledge, plus six new earnings this week from NVIDIA, Salesforce, Synopsys, HP, Everpure, and Marvell test every bear thesis on AI infrastructure and software at once. Patrick Moorhead and Daniel Newman also cover Hot Chips 2026, the NVIDIA-Hugging Face acquisition rumor, and debate whether Anthropic's SaaS reassurances hold up on Ep. 317 of The Six Five Pod.

Marvell's $120B Google Win, Modular's Open Compiler Play, and the Fight Over NVIDIA's Ohio Scale-Back

Patrick Moorhead and Daniel Newman break down Marvell's $120 billion custom-silicon deal with Google, Qualcomm-owned Modular's open-source compiler push at ModCon 2026, and state-level pushback against AI data center construction in Texas and Pennsylvania. The hosts also flip roles to debate whether NVIDIA's reduced Ohio financing guarantee signals cracks in the AI infrastructure buildout.

See more

Other Categories

CYBERSECURITY

QUANTUM