Google Cloud on Why AI Demands a New Infrastructure Operating Model

Google Cloud is now processing more than three quadrillion tokens a month. 7x more than a year ago.

Daniel Newman is joined by Mark Lohmeyer, VP and GM of AI and Computing Infrastructure at Google Cloud, for an AI Infrastructure Spotlight at the Six Five Summit 2026, exploring how rapidly accelerating AI demand is reshaping the infrastructure stack.

Lohmeyer traces the current infrastructure strain to the shift from chat-based AI to agentic AI, where a single user intent can spin off hundreds or thousands of parallel machine-to-machine interactions, increasing inference transaction volume by 50-100x compared to just a few years ago. He outlines three areas Google Cloud is investing in to meet that demand: hardware choice across NVIDIA GPUs, Google's own TPUs, and CPU platforms including its Axion ARM processors.

They’re prioritizing cross-platform compatibility, such as native PyTorch and vLLM support on TPUs, so workloads can move between hardware without being rebuilt, and dynamic reallocation through Google Kubernetes Engine's custom compute classes, which allows workloads to spill over from GPUs to TPUs automatically as demand scales.

Lohmeyer cites internal research showing 90% of enterprises want to deploy agentic AI, but only 17% believe their current infrastructure can handle the load they expect, a gap he calls a defining test of architectural flexibility.

Key Insights:

🔹 Agentic AI is driving a 50–100× increase in inference transactions as a single user request fans out into hundreds—or even thousands—of parallel machine-to-machine interactions.

🔹 Google Cloud is betting on hardware choice, offering NVIDIA GPUs, Google TPUs, Intel and AMD CPUs, and its custom Axion Arm processors—which Lohmeyer says deliver up to 30% better price-performance than comparable hyperscaler options.

🔹 Native PyTorch and vLLM support gives customers greater workload portability between GPUs and TPUs, while Google Kubernetes Engine can automatically shift demand across platforms through custom compute classes.

🔹 The infrastructure readiness gap is stark: 90% of enterprises want to deploy agentic AI, but only 17% believe their current architecture can support the expected workload—making flexibility a defining enterprise test.

🔹 Google Cloud now processes more than three quadrillion tokens per month, seven times last year’s volume, while Google has accelerated its TPU release cycle from once every two years to twice annually with TPU 8t and TPU 8i.

Lohmeyer points to Google's co-designed stack, built with partners including Anthropic and used by customers like Citadel Securities and Mercedes-Benz, as the model he expects enterprises to increasingly demand, with models, software, and hardware built together from the start.

Watch the full video at sixfivemedia.com, and subscribe to our YouTube channel so you never miss an episode.

Explore more sessions from Six Five Summit: AI Unleashed 2026 at sixfivemedia.com/summit.

Disclaimer: Six Five Media is for information and entertainment purposes only. Over the course of this video, we may discuss companies that are publicly traded, and we may reference their equity share prices. Nothing discussed during this webcast should be considered investment advice or a recommendation to buy or sell any security. We are not investment advisors, and you should not rely on this content as financial advice. Six Five Media collaborates with technology companies and industry leaders to produce research-driven interviews and multimedia programming for enterprise technology audiences.

Daniel Newman:
Hey everyone, welcome back. We're at the Six Five Summit 2026, AI Unleashed. We've got another great session here in the AI infrastructure track. A spotlight keynote with Mark Lohmeyer. Mark, not the first time on the show, but we're gonna be talking about what is going on and how AI is demanding a new infrastructure operating model. Mark, welcome back. It's good to see you again.

Mark Lohmeyer: 

Hey, Daniel. Yeah, great to be here. Thank you for having me.

Daniel Newman: 

Yeah, I think last time we had you, we were in person though. It's always a little bit more fun, but this summit has been amazing. And as much as I'd love to have everyone in person getting this group of people together in one place at one time on one day, that'd be a monumental lift. It's like Davos, it's like getting everyone to the middle of the mountains. We'll figure that out someday, but I'm so glad you're back. I mean, what a great topic right now, so much going on. feels like every day there's a new model, every week there's a new breakthrough in infrastructure and in terms of memory walls, networking connectivity, compute capabilities, heterogeneous architecture. So we got a ton to talk about. So let's just start out talking about what's going on with demand. Like AI demand continues to be I call it insatiable, but let's just agree that it's accelerating, right? But anyone that sort of came out of cloud era one and is into the AI era has realized that infrastructure was not, the last era was not designed for what we're dealing with now. Like just for everyone out there, kind of talk a little bit about the fundamental changes that this AI era has created and what it demands from enterprises that are trying to build infrastructure for what's happening.

Mark Lohmeyer: 

Yeah, you know, I think it's certainly an exciting time in infrastructure, probably one of the most exciting times in my career. And I think within that, you know, it's just even over the last few years, I think changing so rapidly. And I'd see the big change we're seeing right now is sort of the shift from, let's call it the chat phase for maybe a couple years ago to the agentic era. And, you know, if you think back a few years ago, it was, hey, chat, you ask a question, you get an answer. Now in this agentic era, it's getting much more complex. You express your intent, something you want to accomplish, that spins off potentially multiple agents, sub-agents, each of those are working in parallel, preserving state as they go. And so now that single intent can trigger hundreds or maybe thousands of machine-to-machine interactions that need to happen in parallel. And as a result, this is increasing the amount of inference requests and inference transactions by 50x, 100x, maybe more, compared to even just just a few years ago. So you have these agents kind of is one big piece. And then I think the other kind of part here, I think is quite interesting is, you know, this is not just about the accelerators, because those agents actually now need to interact with existing applications and existing data. They don't sit on an island just by themselves doing it. For instance, they're leveraging tools, they're calling existing applications, they're accessing existing data resources. And so at the same time that they're putting pressure on the GPUs and TPUs, let's call it, they're also placing increased pressure on traditional infrastructure, traditional compute storage and networking. that is the foundation of all those existing tools and applications. So you put these two things together and it's really driving the need for a much more dynamic and much more, let's call it workload optimized infrastructure. It's not one infrastructure that meets all those needs. You need to have just the right type of hardware, software, full stack solutions that's mapped back to the unique needs of each of those workload types, both for the agents, as well as those existing enterprise workloads. So, having that level of elasticity, that level of dynamic capability within the infrastructure is more important than ever.

Daniel Newman:

Yeah, absolutely. You hit on a couple of really interesting things, especially how traditional infrastructure still has a place, but it's at such a. exponential scale to what it was. Sort of found that too, by the way, as we started the whole sort of AI boom, really focused on compute architecture. And it was all about, you know, basically GPU and then XPU, right? Custom, because obviously Google has built its own, one of the first and done it for the longest time. But, and then we realized, like, I don't know, somewhere Feels like maybe late last year, earlier this year, we're like, oh, my gosh, we need CPUs to like, you know, because you kind of went from like. Single turn, you know. queries on a chatbot to like multi-turn threads. And then you went to inference driven applications. Then you went to agentic. And then of course, we've been iterating on agentic to your point with all the tools called harnesses. And then of course, loops. So that you got all these agents that are running, but they're running perpetually to kind of continuously work on. And all of this created, you know, and we've seen innovations like KB cache, where we, you know, we keep memory accessible for longer periods of time. because obviously each query would require so much compute. If it didn't remember anything, it couldn't. But I mean, it's wild. But to your point, like, this has also put constraint on memory, it's put constraint on storage, it's put constraint on networking. And now we're trying to figure out how to kind of have these two things work together for workloads that are parallel versus workloads that are, you know, more serial, that are done more serially. just so much constraint. So, you talk about kind of this idea of infrastructure becoming dynamic, right? More of an intelligent resource. And I'd like to get your take on this, because if everything I just said and you just said, you kind of put it together, right? Organizations really can't predict the way they maybe once did. Like, you used to buy a certain amount of storage, or you'd get a certain amount of compute. You could sort of Think about how to predict that into the future. Like nowadays things are so exponential, they change so quickly. How are you sort of thinking about dynamic infrastructure to help so that enterprises aren't going to have to constantly rebuild?

Mark Lohmeyer:

Yeah, yeah. It's a really, really important question. And what you described is the heart of it, right? You have these highly variable demands that are being placed on these systems. So how can you sort of provide the responsiveness those applications, those users need, but do so at the lowest possible cost and also for an enterprise in a way that's, you know, doesn't add too much operational complexity. So, you know, this is something we're really focused on, I would say, within Google and Google Cloud. I would highlight kind of three key areas that we think are really important here and that we're investing in. The first is, you know, to give customers choice at the hardware layer, right? Because not every workload is the same. You need different types of hardware for different classes of workloads. And so giving customers the choice across both best-in-class NVIDIA GPUs, as well as our own in-house TPUs, so they can mix and match and choose the right platform based on the specific needs of their workloads. Um, and then you mentioned CPU, right? It's kind of like the return of the, uh, the CPUs, right. And so for all the reasons you described, and so, but I think we think there are two it's important, right? It's not going to be one size fits all. It's not like there's going to be a single, uh, just like in the past, you need a different. VM families, different Google Compute Engine instance types, spanning Intel CPUs, AMD CPUs, our own in-house Axion, ARM CPUs, just like you needed those in the past for traditional workloads. As agents, basically, you need to be orchestrated, call tools, call applications. You need choice of CPUs across those different workload types as well. That's one big thing we're investing in. The second is, now that you have choice of different types of underlying hardware platforms, you don't necessarily want the workload to be locked into any one type of that underlying hardware. So if I take GPUs and TPUs, we're investing a lot to enable compatibility, basically, across GPUs and TPUs. So make it easy for a customer that has maybe optimized a model or a workload for a GPU in the past to be able to now also just as easily run that workload on top of TPUs. So as a specific example of this, we're investing in native PyTorch support on TPUs. Huge ecosystem around PyTorch on GPUs. It's fantastic, making it easy to bring those same workloads and all the things they're used to with the native PyTorch environment, eager mode, for example, to TPUs as well. Another example of that, if you think about the inference side of things, think about agents, you know, the inference engine software is a really critical piece there. Something called VLLM is an open source inference engine, very popular amongst many customers on GPUs. So we enabled that to run, you know, equally well on top of TPUs. So this gives you a lot of flexibility, right? Now I can mix or match just the right GPU or the TPU for the workload. I can move between them easily. And so, you know, customers really, really like that level of flexibility. And then maybe the third thing I would say is, you know, once you've got kind of this breadth of hardware options, you make it compatible across those hardware options through software, you actually need to be able to do the reallocation of those resources to the workloads based on what the workload needs. And so one example of something we've done here is in Google Kubernetes Engine on Google Cloud, something called custom compute classes, where you can basically say, let's go back to that VLM example running on top of GPUs or TPUs. You could say, hey, for this workload, I want to start serving it on GPUs, but if I you know, if the demand scales such that the GPUs that have available can no longer meet that total demand, let's spill over and run that and scale that on TPUs next or vice versa.

Daniel Newman: 

As long as there's something to allocate, right? Because the demand is so high right now that… I've watched those spot prices, but, you know, on a nuanced basis, Mark, I want to really appreciate what you said kind of when you started that whole response, which was, you know, you talked about Nvidia, you know, world class compute. And, of course, you talked about yours. You know, I don't know why sometimes. I listened to outside opinions, investors, media, and they really want it to be zero-sum. They always talk about how like anytime Google comes out with a new TPU, it's their attempt to eradicate, you know, or anytime a new architecture or any event. And I like that you called out because what we always talk about is we're in the abundance era. We're in the era right now where You know there are no idle sitting GPUs or TPUs or CPUs really anywhere in Google's cloud. And so I think a year ago a lot of us were looking at models thinking the models are the most and I think increasingly we're seeing that the compute. And not just the computer, the compute infrastructure stack is the most, of course, to your point, having as much compute is required. And then, of course, having the software that enables you to thread that all together as needed. And I think that's increasingly become the most because once you have that. compute at your disposal, you can train the models, you can build the applications, you can leverage all the data. But obviously we've seen really any time where we've stalled out over the last year, AI overall, it's kind of stalled out when something that people want to use badly didn't have enough compute allocated to it. And we could point to a few examples, but I think everyone's been pained by this at some point, so no need to call out any names. But that brings me to kind of another thing, Mark, is really about you know, cost. You know, we, I have this quote I've used and I say, I tend to say the same things a lot. It's like my TV talk track, but I'm always like, went from token maxing to token optimization inside of 90 days. You know, right now there, you know, you do not need to have you know, the formula one race car to drive to Starbucks to get coffee. Um, sometimes you want that very, very leading edge frontier model run on the most advanced infrastructure. And sometimes you have a workload that can very much run on a TPU V five with a flash or open source flash Gemini or an open source, right? It's just, it's not, it's like, how are you seeing enterprise leaders sort of balance that right? Performance utilization costs, simplicity, you know, how are they balancing it? What are you seeing them in terms of, you know, making these architectural decisions and then kind of what's separating those that seem to be getting it more right than not?

Mark Lohmeyer: 

Yeah, it's a great question. I think, you know, some of the leading customers we work with, one thing they're often talking to us about is the importance of sort of architectural flexibility, let's call it. You think about how dynamic the environment is and the examples you used, right? you know, that flexibility of the architecture that can adapt based on the needs changing rapidly is really important. You know, we did a little bit of internal research and, you know, we found that 90% of enterprises, you know, want to deploy agentic AI, maybe not surprised over the next few years, but only 17% of them feel that their current setup can actually handle the load that they foresee coming. And they're probably actually even underestimating the load, even at that. And so this architectural flexibility is, we think, going to be a key factor that determines their success. As we talked about before, that flexibility starts with the hardware and the hardware choices that you give underneath that you can map back to the workloads. Different horses for different courses, I guess, as they say sometimes. You gave the car analogy, many out there. Same applies to infrastructure, right? And maybe the big surprise of this year to some was the return of the CPU. That's what we're talking about.

Daniel Newman: 

That and DRAM.

Mark Lohmeyer

Yeah, and memory. Who would have guessed? But this is something, again, that we're investing in Google Cloud, as you can imagine, right? having a broad range of Google compute instances, enabling Google Kubernetes on top of that, great partnerships with Intel, with AMD around their platforms, and we're scaling those massively in Google Cloud. But also we have our own Axion ARM processors, which deliver up to 30% better price performance than other hyperscalers. So having that hardware flexibility is an important part of the architecture. I'd say the second piece is open frameworks are really important here. You know, Google's been a leader in open source, open software for many, many years. Think back to the early days of Kubernetes. And we continue that investment going forward. So if you take JAX, for example, developed within Google, used by Gemini, but also made available externally to customers, also supporting things like PyTorch and Keras and VLM and LLMD. you know, all open software capabilities that can enable sort of that portability and flexibility across different hardware types. And then if you sort of work your way up the stack in terms of flexibility, you know, the next layer obviously would be the orchestration itself, right? Like, how do you map the workloads to the infrastructure? How do you scale that infrastructure dynamically as the needs of those workloads, business needs change? And so, Kubernetes here has been sort of the orchestration platform of choice for cloud-native workloads. That will continue to be true, of course. But we're also investing to make Kubernetes the orchestration platform of choice for agent-native workloads. So a lot of really important work happening here. because agents argue require more dynamic, dynamicism in their use of the infrastructure, maybe even than cloud native did in the past. So we think all those pieces are important together at a platform level to enable that level of flexibility to meet the known demands of the future and maybe the unknown demands as well.

Daniel Newman:  

Yeah, no question. And every time I think Every time I hear people kind of talk about having some understanding of what demand really looks like, I just am reminded that people don't understand exponential very well at all. You know, I talk about, you know, people are very good at linear, you know, saying, I think it was one of your peers I was talking to and I basically said, like, it's interesting how a model that was state-of-the-art a year ago within like by this point we would look at that and be like that's not even that's not it's a toy it's not usable it's not something you know and how quickly and I think to some extent even infrastructure you know you used to have a three or four year cycle and now it's you could argue it's it's annual you could argue it's twice annually right now probably at the very bleeding edge of it and I think the beginning of this year, we probably had the most notable inflection of the architectures when we saw what the first kind of wave of agent and inference models were doing, where I think we all kind of went, holy crap, this stuff, like we can really do a lot of stuff with these models. And by the way, those models between then and now in August, and we're doing this conversation on the Six Five Summit, you would argue the ones that we were excited about are kind of what we considered You know, aging at this point, we wouldn't use them for things. But anyways, it's fascinating. So, so let's, let's kind of wrap this up just going to the future a little bit. You know, we've. built a lot about trying to be more dynamic, building frameworks, giving capacity that can scale across networks, and of course, all these things. But what are the infrastructure capabilities that you think are going to become even more essential as this AI adoption continues to mature? And what do you recommend to these enterprise leaders to prepare for?

Mark Lohmeyer: 

Yeah it's a great question. I think the pace of innovation and the pace of value that these applications and services can deliver is only going to keep growing, keep growing exponentially. You know, you talked about the pace of innovation and infrastructure. Just to put that in a specific context, you think about TPUs within Google, right? Previously, we were at a pace maybe of a new TPU generation, one new TPU generation every two years. Then we accelerated it to one new TPU generation every year. And then in this last round, we actually are delivering two new platforms every year with TPU-8i and TPU-8t. So we think that pace will continue. But to bring it back to the customers and what we're hearing from them, I think you know, navigating the growth in demand, we think is going to continue to be really important. We see an industry level demand continuing to outpace supply for the foreseeable future. Even within, you know, within Google Cloud, the number of tokens that we're processing per month is now over three quadrillion. That's a one with 15 zeros behind it, if you're wondering, and that's seven times more than just a year ago, right? So, So that's within Google. Um, but if you're an enterprise leader, you know, you need to, is that what you said?

Daniel Newman: 

Three quadrillion a month. Is that right?

Mark Lohmeyer: 

Yeah. Three quadrillion tokens a month.

Daniel Newman: 

If you wouldn't have this conversation in a year, Mark, we might be saying three quadrillion a week.

Mark Lohmeyer: 

Yeah, we probably will be actually, if you just project the curve. Right. So yeah, but, but you know, that's in Google, right? But now you think you're an enterprise customer. You also need to, you know, scale, uh, your application, scale your infrastructure. Um, maybe not at that, that, 7x rate, but certainly at a very high rate. So this requires thinking deeply about the partners that you want to work with, the platforms you want to build in, build on, and make sure those platforms are ones that can scale with you as your business needs grow and change. So that's something we're investing a lot in Google Cloud, of course. Within that, we think one really critical thing is this idea of a fully co-designed stack, especially for AI energetic workloads, from the models to the software to the hardware. We do that, of course, with our own Gemini models. make those available to the world. But also, we co-design with many of our large customers, like Anthropic. And ultimately, this enables Google Cloud as a great platform that you can come to as an enterprise customer. You can get access to the leading models. You can get access to leading infrastructure in a nicely integrated way, which makes it easier for enterprise customers to adopt and really get the benefits. And as a result, we're just seeing really, really great traction, of course, Virtually all, if not all, of the leading AI frontier labs are running on Google Cloud, but not just those. Financial services, capital markets firms like Citadel Security, enterprise customers like Mercedes-Benz, and many others. You know, it's sort of at this point, I think every enterprise is going looking to drive this type of innovation, innovation and going through a transformation. You know, really, we want to be a strong partner to, you know, to everyone and really help support their growth, support their growth and demand. But maybe even more importantly, help accelerate their own innovation, because that's how they're going to, you know, ultimately all win in the markets that they're that they're playing in.

Daniel Newman:

Yeah, and I think you really bring it together nicely there, that it really is all the above. I know I made the comment that compute is the moat, and I'll stand by that, that you can't do the other things, it's kind of at the foundation of it all, but you do make great points. Because when you want the absolute agent performance, and we know this from doing our own benchmarking, that it is the whole stack. It is software. It is what model you're running on which infrastructure with which software for what use case. And that's also where we bring in the governed compliant. you know, the business rules and all those things, because in the end, like, and then, of course, trying to create an experience for the user, which is sometimes not scored at all in our world of benchmarking. But again, it has to also work really, really well. So, look, very exciting times, Mark. Appreciate you coming on, sharing a little bit of your perspective, giving us a little bit of a view into the future of what you're doing, what you're building. Congratulations on all the success so far, and thanks for joining Six Five Summit.

Mark Lohmeyer: 

Yeah, thanks, Daniel. Great to spend some time with you here. And as always, really enjoy the deep and insightful discussions.

Daniel Newman: 

Yeah, let's do it again. And everyone out there, stick with us. Don't forget to subscribe. We're here at the Six Five Summit. It's an AI infrastructure track. Follow us on all the socials. Check out all our summit coverage. More coming soon.

Speaker

Mark Lohmeyer
Vice President and General Manager, AI
Google Cloud

Mark Lohmeyer leads the Compute and AI Infrastructure business for Google Cloud. In this role, he is responsible for the Google Cloud Compute Engine, AI/ML infrastructure (Cloud TPU and GPU), Core ML services, block storage (Persistent Disk and Hyperdisk), and enterprise solutions (SAP on GCP, Google Cloud VMware Engine, etc.)
Mark’s background includes leadership roles in general management, product management, marketing, business development, and engineering management, across a wide range of core infrastructure technologies, including compute, storage, and networking.


Prior to joining Google, Mark was the SVP/GM of VMware’s Cloud Infrastructure Business Group. In this role, he led a large-scale, global organization, spanning engineering, operations, product management, and product marketing for the VMware infrastructure portfolio across Private Clouds, Public Clouds, and Cloud Provider Partners / Sovereign Clouds.


Prior to VMware, Mark led the product team for Enterprise WAN and Routing at Cisco and was the GM for HA/DR and Storage solutions at Veritas Software. Earlier in his career, he worked on storage I/O hardware at Adaptec, and digital imaging research and hardware design at the Sarnoff Research Labs, and holds a patent based on this work.


Mark holds a Bachelor's and Master's degree in Electrical Engineering and Computer Science from the Massachusetts Institute of Technology, where he also served as the head teaching assistant for Computational Structures.

Mark Lohmeyer
Vice President and General Manager, AI