Home

Fleet-Scale Deployment Is the Physical AI Benchmark

Fleet-Scale Deployment Is the Physical AI Benchmark

Physical AI is shifting from lab prototypes to fleets of machines operating reliably in the real world, and Ambarella's Muneyb Minhazuddin joins Six Five On The Road at AI Infra 2026 to explain what that shift demands of silicon, memory and developer workflows. Minhazuddin traces the case back to a 22-year-old chip architecture and forward to a developer cloud built to compress model porting time from months to days, and he identifies enterprise ROI proof points as the biggest hurdle left before physical AI scales into millions of deployed units.

Physical AI prototypes are multiplying across trade show floors. Industrial buyers need a different bar: years of uptime, continuous upgrades, and predictable performance across thousands of deployed units.

Muneyb Minhazudn, Customer Growth Officer at Ambarella, joined Brendan Burke and Matt Kimball at AI Infra 2026 in Santa Clara, CA. Together, they unpack what separates a physical AI prototype from a deployable product, and what that separation demands from silicon, memory architecture, and developer tooling.

Minhazuddin argues physical AI's decision loop: perceive, reason, act, and learn, now closes entirely in real-time on the device itself, folding functions that used to route between edge and cloud into a single local cycle. He pointed to Ambarella's SRAM-based chip architecture, a design choice made 22 years ago for GoPro-era power constraints, as the reason its chips use up to 10 times less memory and 13 times less power than competing designs running the same transformer and vision-language models. He also detailed how Ambarella's developer cloud lets independent software vendors test workloads on live silicon before committing to hardware, cutting porting timelines from as long as a year to a few days, and pointed to a hospital patient-monitoring deployment where moving inference into the camera itself eliminated the need for a dedicated server for every 50 camera streams.

Key Takeaways:

🔹 Ambarella has deployed more than 50 million chips across cameras, vehicles, drones, and robots. That scale positions its physical AI strategy as an extension of two decades of edge silicon experience—and gives the company a real-world reliability track record that competitors without deployed fleets cannot easily replicate.

🔹 Ambarella’s system-level SRAM architecture, designed 22 years ago, uses up to 10 times less memory and 13 times less power than competing chips running the same models. That advantage becomes even more valuable as memory supply tightens and data centers and edge devices compete for capacity.

🔹 Ambarella’s developer cloud allows independent software vendors to test and port models onto live silicon within days—a process that previously took some partners eight to 12 months.

🔹 Integrations with Google Cloud, Gemini, Capgemini, and Ultralytics expand Ambarella’s ecosystem beyond chip sales into applications, systems integration, and model distribution. Its new X7 accelerator also allows customers to add AI capabilities to existing deployments without replacing their installed hardware.

🔹 In one hospital deployment, Ambarella’s technology eliminated the need for a dedicated server for every 50 monitored rooms—demonstrating the potential TCO benefits of physical AI at scale. Minhazuddin identifies consistent proof of ROI and TCO across industries such as retail, manufacturing, and healthcare as the biggest remaining barrier to fleet-wide adoption.

The chip decisions that survived GoPro's early power constraints are becoming Ambarella's differentiator in a market where every AI accelerator company is discovering memory and power constraints at once.

Watch the full video at sixfivemedia.com, and subscribe to our YouTube channel so you never miss an episode.

Disclaimer: Six Five Media is for information and entertainment purposes only. Over the course of this video, we may discuss companies that are publicly traded, and we may reference their equity share prices. Nothing discussed during this webcast should be considered investment advice or a recommendation to buy or sell any security. We are not investment advisors, and you should not rely on this content as financial advice. Six Five Media collaborates with technology companies and industry leaders to produce research-driven interviews and multimedia programming for enterprise technology audiences.

Transcript

Muneyb Minhazuddin:

That closed loop used to be a TikTok model. Perceive in the edge, think in the cloud, act in the edge, train in the cloud. Physical AI is closing the loop all in real time.

Matt Kimball:
Hello and welcome to 6.5 On The Road from the AI Infra Summit 2026 in Santa Clara. I'm Matt Kimball with More Insights & Strategy and I'm joined by Brendan Burke, the Futurum. And today we're exploring physical AI beyond the lab and at scale. What does that mean? So we're joined by Muneeb Hazidin from Ambarella, Chief Customer Growth Officer, to unpack what this really means and what that means in terms of the shift in demands on silicon. So thank you for joining us, and let's have a fun conversation. Thanks for having me.

Brendan Burke:
Many physical A.I. has become one of the most used buzzwords in the A.I. industry over the past year. But a lot of what gets shown ends up being a prototype instead of a production scale deployment. And so industrial buyers are evaluating this technology. They can't necessarily risk investing in a prototype. So what can a developer in an industrial setting look for from physical AI? What can they measure to really prove that it's ready and that they can use physical AI today?

Muneyb Minhazuddin:
Absolutely. I think it's been a big buzzword. I think the key thing they should look for is deployments, right? So it's easy to do a prototype, do a one-off. But in reality, you want to be rolling tens of thousands of these things and do the boring stuff of make them reliable, run for years, upgrade them, make them continuously intelligent. How do you do that? And I think people should go beyond experimentation and to real full-scale deployment. So people assessing this should be looking at, what does it take to deploy at scale? And I think that comes with its own challenges, right? Because people have done a great job of doing AI in the cloud. We all know this. And that development is useful because a lot of big brain function is being developed over there. So massive power, massive data center footprint and capacity. However, scaling to a few hundred data center locations, few cloud locations is very different, which are highly dense, no matter, to millions of devices across the whole world or a geography. And managing them is super hard. So trying to squeeze a large model to do something that is Very nimble at the edge is very different. So you got me. It's not like squeezing big models and infrastructure for physically. I actually almost have to reengineer everything to meet the physical requirements and I think that's where the big challenges. And I would say from an umbrella perspective great opportunity that we've been doing physically I before physically I was calling. We've deployed 50 million plus chips across. What is the definition of physical eye in in variables cameras cars drones robots. for a long time. And so we've kind of done that at scale. So when buyers are looking at it, they have to consider what does it all take not to do just one demo, but deploy millions of these units and operationalize them. And then how do you drive the efficiency from it? Because that's where the rubber meets the road.

Brendan Burke:
Figuring out what the gigawatt unit is for physical AI, because I think we're learning how to scale those data centers. But what is that right unit to extend it into the real world?

Muneyb Minhazuddin:
Oh, geez. The performance per watt, per dollar, per, you know, and the deterministic nature of it. Because, you know, this is the difference between always being between the cloud and the edge is when you're going on an e-commerce link and it takes 30, 40 milliseconds for you to do a web click to do a transaction, it's a small nuance you don't notice. You apply that same delay and latency in a real-time system, it could be huge implications. Like, you know, it could be a fast train applying a brake, seeing something in real time. And those milliseconds need to be microseconds in order to avoid an accident. You cannot afford for that. So this is where I say entire re-engineering is required. for not just squeezing models and infrastructure, but you gotta do it in real time, in a deterministic way, deterministic as it has to do it a million times again and again, and at scale, right?

Matt Kimball:
So when you think about that, right, because you're hitting on a lot of kind of what we want to dig into here, you know, as you move from, as we see physical AI move from, that's really cool and it's almost science fiction, into this is real world application, real world reasoning, Those teams that are deploying all of that physical AI, how do they need to look at this differently? Because they're not necessarily used to this. We understand some of the compute requirements, how they change, but what else changes in the real world for those teams?

Muneyb Minhazuddin:
I would say, you know, if you take the basic principle of compute, compute, networking, storage, data, right? Fundamentals, right? Go back to your first year of engineering and computer science. They teach them that. So compute looks very different. It's not massive compute. It's lightweight compute. Storage, I don't have a lot of storage. So where does data go? I have to do streaming data and I have to inference the data in real time. And networking, I have to not rely on networking for everything. So cloud's not gonna come to my rescue. So I have to be really autonomous. Most likely I can talk to my peering device in my neighborhood, but not gonna go to the cloud. So when I talk about the re-engineering aspect of it, you have to start with, Where's the data because a lot of the times the data is the gold mine for a lot of the business is trying to apply physically. I they're trying to deal with some amount of data take that data and then process it inference or didn't the data so I don't have to do a lot of storage and then I have to do compute on that in order to receive the data. people have been using sensors. So vision is a sensor, you know, LIDAR, radar, temperature, audio, like, you know, all the sensor types. So sensors are kind of the analog to the digital converters. And then you aggregate that and then you process it at the edge. So you kind of like, you almost have to start not from cloud and models data. You have to start from the physical aspects of what type of things am I collecting? How am I redoing it? Then you've got to estimate my power, my throughput, my efficiency. So what's evolved in the physical AI space is we've gone from sensors giving a level of perception. From perception, which is all the sensors, and call it application modules in the chips that process that. Could be digital processing, signal processing, image processing, video processing. From there, you moved on to now to say, I have to do some level of cognition. That runs on a, where does it run? Does it run on an AI accelerator or a CPU? Maybe it's an AI accelerator that's doing the brain thinking of, I've got all these sensors, I'm fusing them, I'm putting a cortex around my scene. Could be vision, temperature, whatever it is. Once I've done the cognition part, I have to take action. That decision-making may not be in the AI pod, could be on the CPU, could be agentic. The agents are now making that closed-loop decision-making, and then the actuation happens. Once the actuation happens, there's learning. The learning is like, oh, I did this right, I did this wrong. That learning is fed back into that autonomous device. So that closed-loop,

Brendan Burke:
Used to be a tick tock model perceive in the edge think in the cloud act in the edge train in the cloud yet physically I is closing the loop all in real time so it's impressive to see the way you're engaging with the AI LM community and being able to deploy models closer to where these decisions are being made and. You know that on an edge device. You might only use a say 1.5 billion parameter model. I mean it's a fairly small compared to the frontier models we have, but when you look at edge devices and the yellow detectors that edge developers are used to deploying on a device, there's still a order of magnitude difference in the memory you need. And we're here in the Ambarella booth. I can see a robot over there that's having its DRAM bandwidth measured in real time in terms of the memory requirements it uses. But at a time when we don't have as much memory available and data centers are taking it up, how does Ambarella decide how to leverage the memory that is on a device to actually make use of the models that are coming out?

Muneyb Minhazuddin:
Yeah, this is probably a trillion dollar question on memory availability for the market. Everybody's after memory. If I was giving away memory in the booth, instead of the yo-yos, everybody will be queuing up. But it's a matter of fact that the memory requirements have gone through the roof, of course, through data center build-outs and cloud build-outs for a lot of folks. Again, the edge in the problem is different. We've optimized, and again, this is a little technical to get into the umbrella architecture across our 40 chips, across our 400 million deployed chips, it's the same. So it's like, you know, we decided on this architecture 22 years ago when we found it. It wasn't for here and now that we're going to have a memory shortage or power shortage. So I always ask the founders, you knew this day was coming, right? So our architecture has a CPU block. Has a you know application blocks for ISP for image signal processing, you know video processing. Then we started adding our NPU block for the acceleration part, but we always you know typically when you think about all these blocks on a chip. They're traversing IO and memory they hit HBM and you know when they're going and going and when you're building this data pipelines they go in and out of these blocks to common memory channels and hitting IO all the time and receiving it. Now the genius of the founders 22 years ago was to build an SRM on a system level and connect these data pipelines directly. So we kind of, yes we use SRAM and memory, but efficiently always in our 40 different tape outs where our memory utilization is super low, our power utilization is super low. Because we were first designs building for the GoPros with the world, already consumer battery operated. So that design was always there. Now fast forward, today we're able to do CNN and transformer networks, LLMs, VLMs, VLAs. in that same architecture, adds significantly lower memory. Sometimes 10 times less of memory than other chips. Sometimes power is 13 times lower than other major competitors. But the design was done 22 years ago. And I think we started adding AI in 2015 after AlexNet. Because a lot of AI chip startup companies are into AI because now, we're like, well, I think we should do this about 12 years ago.

Brendan Burke:
SRM accelerator bets are starting to pay off in the data center space and some of those only got around to it five ten years ago. But you've had a longer term bet in the space and it seems that the architecture is now catching up.

Muneyb Minhazuddin: 

I think it's always the timing you know there's all big companies big GPU companies starting with gaming pivoting to crypto that going after cloud like and it took 20 plus years for those. Moment to come I think this is the moment for umbrella where I think we've been designing for physical edge AI systems all along. And the timing is right, because I think all the elements of compute, network storage, AI, and agentics also, because agentics is not that old. And people are, I am also very, agentics is happening, it's too early. But one realization was, that is a true statement for agentics in the cloud. Because when you take large ERP, CRM systems in the cloud, agentics has to go through millions or billions of permutations and combinations of different logic and systems. Coming down to an edge to a physical system, it's a fixed function. There's only a few things you have to do. Therefore, the models are much smaller. The logic is very fixed. Doing agentics and decision makings in those small constraints is way better and controllable than unleashing it in a cloud with billions of permutations and combinations. So agentics in a physical AI is a lot more practical and controlled and doable rather than doing it in the world wide web.

Matt Kimball:
Hey so speaking of the cloud and kind of getting into you know amber Ella supporting the ecosystem right so you've got a cloud that allows developers to test their software on your silicon live right fantastic right. How does that impact or maybe it compresses like that decision making process for this is what I choose this what I'm selecting for my silicon my hardware to to implement designs. What's the impact of that in the real world.

Muneyb Minhazuddin: 

It's huge by the way. Right. It's time to market. Everybody wants to get to market fast. Yeah. Traditional way. You know it's like hey you know card. I want to look at a car. I research my car online. I go go to a dealership look at the look I do a test drive and I you know it's a you know silicon buying this kind of like that you read up on it then you buy a kid and then you evaluate the kid you put your things and do this now the time to market with certain certain things there's a there's a time lag sure people want to move faster. But they also want to do these evaluations somewhere where it doesn't cost them anything that's right. So what we did beginning this year we started saying hey we have a consistent software SDK API across all our 40 chips, but it was always closed to our big customers. So we decided to open that out into the cloud so we launched our developer cloud. But when you usually open up your APIs and SDKs, there's a huge effort you have to do to build communities. You have to do meetups, shirt coffees, hackathons, build a community, train them, enable them on APIs to write good code. Now, here comes the genetics and what coding. So, when we actually brought it to market, we said one, Well, this ramp is going to take a long time, so we're going to put an agentic wrapper on our SDK, and then we're going to make this available to independent software developers, ISVs, application builders. And we went to a few application developers in the evaluation, so we opened it as early access. We said, we want some feedback. And in my past, I've worked with some of these ISVs, porting their models and applications from one platform to the other takes me eight months, nine months, 12 months. Now, when we brought them to our agentic platform, and this is where I went from, call it, not sure to a believer, right? Where these guys were able to, I was a convert because in three days they came back and said, Muneeb, I took everything, I put it over to yourself using your agentic layer. I don't know anything about our SDK, I don't know anything about our APIs, but my app, my model now runs efficiently on a chip. I was blown away. I'm like, oh my God, this thing really works. Um, so agentics is kind of happening at the real. I said, Why do you think it works again? What? I reiterate it. It's very small parameters and very well controlled. His fist functions. It's not going crazy. It knows exactly what it needs to do. That was a lot more, you know, tight that they could pull it off. And I think what we decided subsequently with press release this morning, we were doing it on our own. Now we took that developer cloud, and we're hosting our developer kits in Google Cloud, where we're integrating with Gemini services. And enterprises that are coming through a Google interface can start leveraging. They can have our kits as a target to build their applications in physical AI and test it. In a couple of days. Yeah, like maybe faster, but we're also along the way building an ecosystem. So we announced some relationships with, you know, umbrella has always been a direct business resellers with Magneta system integrators like Capgemini. Because those are the ones will come and actually do the deployment for large scale enterprises. And then also our tech stack. You know, we did this morning announcements with ultralytics for Yolo models. You know we move you put it over liquid models to our model garden so we publish a whole bunch of models you know we have the data partnership announced this morning where we you know we also you know are able to bring them on there able to run evil as their hypervisor on our chip and manage a fleet of hundreds of thousands or millions of edge stuff which is what enterprises will need for that's for large scale deployment. And we launched our new chip, it's a X7 accelerator chip. And that acts as a, we have a 400 million install base, but 50 million AI, if the others want to add AI, they don't have to go through a refresher of a chip, they can just add an extra accelerator and make it more AI enabled.

Brendan Burke:
It sounds like we're ready for scale in physical AI. Do you think that now developers can deploy an AI model across tens of thousands, or like you said, a million devices, and have a base model run everywhere? And if that is the goal, what's the biggest hurdle to achieving it? Amber was ready for this.

Muneyb Minhazuddin:
The ecosystem's getting there. I think a lot of the enterprise use cases need to come to life. People have again done the perception solving the problem in the cloud they need to kind of create. ROI TCO based outcomes for enterprises to drive that option because everybody's focus on data center in cloud. The reality is line of business will start touching these elements and they wanted in a coffee shop they wanted in a manufacturing floor they want a hospital room those use cases are still stringent where people are unsure because they still are getting the cloud and data center high. They've been honed in on, I have a use case. I have a scenario. Here's a ROI TCO. So what Ambrel has been doing by enrolling a lot of ISVs in the retail, in manufacturing, in health care, in transportation, is coming up with end-to-end solutions that enterprises can realize the value for it. I can talk for ages on RYTC of any one of these use cases. Simple one, hospitals, hospital rooms, they're putting cameras, and they were looking at monitoring patients who fall over. And when they fall over, they inform the nurse. It's a simple use case. But they were taking the video streams. They had a huge server and a big GPU monitoring for every 50 streams if this is happening. And the server will decode it and then inform the nurse station. Umbrella eliminated that entire box. So there's no box. The AI runs in the camera. You watch somebody fall down. The LLM and the agentic stuff notifies the nurse directly. So now you're getting two things. One, the TCO is that 10, 12,000 server per 50 cameras. If you are 200 room, four servers, that cost of ownership is gone. Second, it can scale because for every 50 I need a server. Now I can have a thousand rooms. I only need those intelligent cameras. That's incredible.

Brendan Burke:
That's a sign to put up in your booth here that you're saving people on memory. If you can reduce the memory they need.

Muneyb Minhazuddin:
I think you don't need any other servers and that all that memory is freed up now, right?

Brendan Burke:
It's otherwise the language of the enterprise and it's crazy here that you're speaking it from the supplier perspective. Because that suggests that there's a clear end to end stack and a clear value proposition downstream to get more of those solution bleep runs out there.

Muneyb Minhazuddin:
So our mission to the developer zone is to create a chess board. I call it. You know verticals and horizontals is like hey anomaly detection. applied in a retail store is, you know, defect, like, you know, it's like, hey, shoplifting or fraud detection, applied in a manufacturing is defect detection. So you want to take this horizontal use cases, apply them verticals, have ISVs, who have opinionated models, small models, working on the edge, solving for it. So we're building out this catalog of applications to populate that entire thing and give outcomes to enterprises. And that's where system integrators, distributors, resellers are going to be able to take those blueprints and catalogs and go to play.

Brendan Burke:
Really exciting to see. Really appreciate you sharing how much progress you've made with a lot of great news that the developer community should check out. Absolutely. Thank you. Thanks for tuning in to 6.5 On The Road from AI InfraSummit 2026. Be sure to follow us on socials and check out our coverage on 6.5media.com. We'll see you next time.

MORE VIDEOS

AI Safety Alarms, China's Distillation Reckoning, and Qualcomm's AWS Breakthrough: A Pivotal Week for Enterprise AI

A viral AI extinction claim, a federal advisory naming six Chinese AI labs for industrial-scale distillation, and a $60 billion Qualcomm-AWS silicon pact define a week when enterprise AI risk and opportunity collided head-on. Patrick Moorhead and Daniel Newman weigh in on what's signal and what's noise across AI safety, chip supply deals, and a market rally still betting big on compute.

The 15-Cent AI Query: Broadcom's Paul Turner on Rebuilding Enterprise AI Economics

A Six Five study found an eight-to-one cost gap between running an AI agent through a leading frontier model and running it on modest on-prem hardware. Paul Turner, Chief Product Officer of the VMware Cloud Foundation Division at Broadcom, joins Daniel Newman and Patrick Moorhead at VMware Explore 2026 to explain how VMware's AI Factory turns GPU provisioning, model management, and security into a repeatable path from infrastructure to production AI.

VMware Explore 2026 Wrap-Up: What’s Next for Private Cloud and Enterprise AI

What infrastructure model makes the most sense when AI workloads move from experimentation into production?

In this analyst recap of VMware Explore 2026, Patrick Moorhead and Daniel Newman examine this central question facing enterprise technology leaders and how Broadcom and VMware are positioning private cloud as a critical foundation for bringing enterprise AI into production.

See more

Other Categories

CYBERSECURITY

QUANTUM