Home

AMD Instinct MI355X vs NVIDIA HGX B200 GPU platforms AMD vs NVIDIA: A Three-Level Inference Infrastructure Evaluation

AMD Instinct MI355X vs NVIDIA HGX B200 GPU platforms AMD vs NVIDIA: A Three-Level Inference Infrastructure Evaluation

Signal65 President Ryan Shrout and Senior Performance Analyst Mitch Lewis review a new three-level inference infrastructure evaluation of AMD's Instinct MI355X against NVIDIA's HGX B200 with AMD's Andrew Dieckmann and Anush Elangovan. The MI355X's memory bandwidth advantage compounds as context and concurrency scale, delivering up to 2.25x throughput gains and roughly 40 percent lower cost per GPU hour.

Enterprise AI infrastructure buyers have spent the past year evaluating GPU platforms through raw benchmarks and FLOPs-to-token ratios. But these metrics often fail to capture how systems perform at production scale. Signal65’s latest research moves beyond headline numbers, comparing AMD Instinct™ MI355X and NVIDIA HGX B200 across three dimensions: performance, infrastructure cost, and modeled business outcomes.

Andrew Dieckmann, Corporate Vice President and General Manager of Data Center GPU at AMD, and Anush Elangovan, Corporate Vice President of Software Development at AMD, join Signal65 President Ryan Shrout and Senior Performance Analyst Mitch Lewis to unpack the findings.

The research shows that the MI355X’s memory bandwidth and HBM capacity advantages become more pronounced as context length and concurrency increase. In one GPT-OSS-120B test, the platform delivered 2.25x the throughput of the B200. Signal65 also measured about 40% lower cost per GPU hour across cloud providers, including TensorWave, Vultr, and Oracle.

In interactive-serving tests, the MI355X supported nearly twice as many concurrent users within a defined 15-second P99 service-level agreement while delivering 42% lower tail latency. Elangovan explains why tail latency is critical to real-world user experience: strong average throughput can still mask unacceptable worst-case response times.

The conversation also explores AMD's rapidly evolving software ecosystem. ROCm now operates on a six-week release cadence, while AMD's ATOM inference engine led most tests against vLLM and SGLang deployments. Dieckmann connects the platform roadmap to AMD's co-design work with OpenAI, Anthropic, and Meta, and previews how the company's next-generation Helios rack-scale platform will extend its memory advantage.

Key Takeaways Include:

🔹 The MI355X’s performance advantage increases with context length and concurrency, reaching 2.25x the throughput of the B200 on GPT-OSS-120B as HBM capacity and bandwidth become critical constraints.

🔹 Signal65 measured approximately 40% lower GPU-hour costs for the MI355X across TensorWave, Vultr, and Oracle, with the cost advantage holding across all three models tested.

🔹 In interactive-serving tests, the MI355X supported nearly twice as many concurrent users within a 15-second P99 SLA while delivering 42% lower tail latency than the B200.

🔹 AMD’s ATOM inference engine led most production-model tests against vLLM and SGLang, challenging the perception that software remains a weakness for the platform.

🔹 AMD now releases ROCm updates every six weeks and is co-designing its next-generation roadmap with OpenAI, Anthropic, and Meta.

Read the full research, executive summary, and infographic at https://signal65.com/research/ai/amd-instinct-mi355x-vs-nvidia-b200-inference-evaluation/.

IMPORTANT: Note to Uploader: VIDEO EMBED – added by uploader.

Disclaimer: Six Five Media is for information and entertainment purposes only. Over the course of this video, we may discuss companies that are publicly traded, and we may reference their equity share prices. Nothing discussed during this webcast should be considered investment advice or a recommendation to buy or sell any security. We are not investment advisors, and you should not rely on this content as financial advice. Six Five Media collaborates with technology companies and industry leaders to produce research-driven interviews and multimedia programming for enterprise technology audiences.

Transcript

Andrew Dieckmann:

We want to offer our customers more in terms of performance, performance per watt, performance per dollar, and continue to drive efficiencies in their business because the tokenomics of these clusters really matters.

Ryan Shrout: 

Hey everybody, welcome to a Signal 65 Video Insights. I'm your host, Ryan Shrout, President at Signal 65, joined by my partner in crime, Mitch Lewis. Mitch, how are you today? 

Mitch Lewis:

Doing well, happy to be here. 

Ryan Shrout:

We are here today to talk about a report that we published just a little while ago called the AMD Instinct MI355X GPU Platform, comparing it to the NVIDIA B200 GPU. And we looked at it as a three level inference infrastructure evaluation. It's a lot of words. We're going to dive into it. We really have two of the best possible people that you could have with us. We're very fortunate to have them to talk all about this report and kind of the ecosystem and market in general. Andrew and Anoush from AMD, welcome to the show and thanks for joining us. Really appreciate it.

Andrew Dieckmann: 

Thanks very much. Happy to be here.

Ryan Shrout: 

Thank you for having us. Look, we talk a lot, all of us, about AI inference and the performance capabilities and kind of the changing landscape of how all of this works. We've seen, I think, a very important move over the last, say, year or so, where AI infrastructure purchasing decisions were often made on kind of like raw benchmark, I'll call them raw benchmark numbers alone, kind of your flops to tokens per second ratios. I'm curious, maybe, you know, either one of you, I'll start with you, Andrew, like, why do you think that might be an incomplete picture? What should enterprises actually be talking about or evaluating when considering kind of different hardware choices?

Andrew Dieckmann: 

Yeah, I mean obviously performance is a key buying criteria for when people make infrastructure decisions. Raw performance, performance per watt, but also performance per dollar all come into the picture. And then I would say another equally important buying criteria is the ability of the platform to continue to evolve with the workloads as things change, because we live in a rapidly changing environment in terms of various models or various inferencing techniques that people want to employ. And having a platform that continues to be scalable and provide good utility for many years is an important part of the equation as well.

Anush Elangovan: 

It definitely raw performance is one indicator, but then you want to actually see what, what, or how does the platform, you know, perform for like the outcomes that customers need right so if you're actually looking for an agent take workload and how does that actually execute because it's a. It's a synchrony of networking, CPU, GPU, and you put all of those together and a raw performance benchmark would just be one aspect of it. But you want to look at it in a system and platform level. And so we spend a lot of time trying to optimize the entire platform to be able to demonstrate measurable outcomes for customers.

Ryan Shrout: 

Yeah, and I think it's to your point, Andrew, about how quickly all of this changes as the infrastructure changes, as the software layer changes, you know, the hardware changes that we generally think of as very fast, that one, this kind of one year cadence are almost like the slowest piece of that chain. And so I think it's important for the way that we talk about performance, the way we measure and kind of look at it as well, kind of adapts in the same way. But maybe I want to start a little bit higher level than even that. Andrew, could you, you know, for the audience, it's maybe, you know, a little bit more naive on this type of stuff, or they're just figuring it out. Like, explain where the Instinct MI355X sits today in the AMD kind of GPU strategy.

Andrew Dieckmann: 

We've been working this market for several generations. We just recently launched our first rack scale infrastructure with Helios. The report really focused on 355 which is our eight way offering, which is available in both direct liquid cooled as well as air cooled offerings. That is the platform that we have been predominantly shipping for the last year or so. You know, we found a lot of great utility with that platform, particularly with a lot of pickup in the inferencing workloads with a variety of customers. And we really show into some of our earlier discussion and ability to continue to scale the performance of that platform as new models and techniques evolve, be able to get multiples of the initial performance out of that platform as we work collaboratively with customers. So we remain committed to enabling at-scale inferencing, that's the predominant workload even today in the market, as well as enabling our customers with large-scale training clusters based upon our platform. So we're really operating in both of those categories and continuing to expand our platform as we move forward into Helios and rack-scale infrastructure moving forward.

Ryan Shrout: 

Some of the testing that Mitch found was that the MI-355X advantage grew quite a bit as the context and concurrency increased on some of the models we looked at, up to 2.25x on GPT-OSS, for example. How much of that is the memory architecture design and why does that HBM design and capacity matter so much for modern inference today?

Anush Elangovan:

 Yeah, I can. So one of the key design points that AMD has like doubled down on is memory bandwidth and memory capacity. So if you look at every generation of product, we've had a tremendous lead over competition in both fronts, right? And leading up all the way to 55 years. And, you know, the industry is definitely now keyed in on like decode being bandwidth heavy and memory bandwidth heavy and and then obviously the model sizes are getting, you know, to the trillions of parameters. So both capacity and bandwidth are just exactly the sweet spot that, you know, AMD has been designing for multiple generations and now. Pin that with competitive flops, you get a system that, like I mentioned, is holistically set up for large language models and inferencing at scale. And so it isn't just that it isn't just one aspect of it, but it is designed at a whole system level. And memory bandwidth is key to decode.

Mitch Lewis: 

Maybe just to jump in there real quick and put some quick numbers as context. So, we we can get more into this, but we ran a couple models and you know there was only one model where we saw. a little bit of a B200 advantage, around 17%. But then as the context grew and the concurrencies grew, the MI-355X actually went up to around a 1.96x advantage. So you can kind of see that memory architecture on the larger HBM come into play really clearly in that kind of situation.

Anush Elangovan: 

Yeah, and I'll just add that, you know, those are impressive numbers in terms of performance. But then once you add like, you know, the PCO benefits, etc, right, you have a very, very compelling platform for large scale inference.

Ryan Shrout: 

Yeah, I mean, that was one of the key data points that we're able to pull out of it, right? As Mitch, we kind of looked across like blended rates of different areas where you could, you know, rent out an MI355X node versus a B200 node. And I think the numbers we came up to was roughly 40% lower per GPU hour for the MI355X platform, right? And, you know, we were seeing more vendors kind of coming on board with 355X support. You know, you've got TensorWave Vulture and Oracle as an example. And I'm curious, like, to your point, any performance advantages that you have get accentuated as you look at it from a performance per dollar perspective. But Andrew, I don't know if like, is there any kind of underlying commentary about how AMD thinks about that combination of offering tremendous performance, but still offering it at a great value to kind of like bring that TCO kind of down? Like what's kind of the business strategy behind it?

Andrew Dieckmann: 

Yeah, how we think about that, we start off by absolutely leaning into performance of the architecture. And our goal when we conceive of a product is to deliver performance leadership. We certainly think that our competitiveness has been steadily improving over time when you benchmark us against leading offerings in the market. And we think particularly with Helios, we are offering the highest performing GPU and the highest performing rack in the market, given our memory. We're extending our memory advantage again for another generation, offering a lot more capacity, a lot more throughput versus competing platforms. And we know that as a challenger within this market and the growing footprint, we are competing aggressively to gain share. So the first step is offering a great product. We're not offering a lesser product. We're offering, actually, we think, a leadership product. In the case of 355, that product, from a specification and workload perspective, is very competitive with the Blackwell offerings. We're extending that forward and then offering it at excellent value to our customers. So, you know, that's our goal. That's our recipe. We want to offer our customers more. in terms of performance, performance per watt, performance per dollar, and continue to drive efficiencies in their business. Because particularly as the market moves into more and more inferencing and the tokenomics of these clusters really matters in terms of the underlying businesses being profitable and maximizing their profitability. So that's our recipe that we're bringing to the market.

Ryan Shrout: 

I'm going to dive into a little bit of a software piece here and let Anoush gloat for a minute, but I'm curious, right? So I would say if you go back a couple of years, several years now, software has historically kind of been the narrative used against AMD and the AI infrastructure space. But in this report, we actually ran production models on RockM across both VLLM and Atom, and the MI-355X led on the majority of our testing. What would you say has been the biggest change from an AMD software approach that's built this up?

Anush Elangovan: 

Yeah, I think what we've done is over the years, we have now become software first. We ship software like a software company. I think in the past, AMD had grown organically and had hardware that required BSPs and you shipped for a local maxima of like, okay, you can unlock the CPU, you can unlock the GPU, you can unlock the FPGA. I can unlock parts of it. But what we're doing with Rokum is that we're trying to bring it as a unified platform that allows for customers to invest, not just for the current generation of hardware. And it provides a surface for what is available today, but it also extends out into the future so that you build on Rokum and then every generation you can upgrade your GPUs, your CPUs, whatever it is. And that starts to live on as a platform that customers can rely on. And as part of that, we've, you know, we've modernized how fast we ship RockM too, right? Like, and these are very core, like deep infrastructural changes that we've done that went from like four month release cycles that could be gated on like some feature coming in here or there, versus now it's like clockwork every six weeks, right? The train's going to leave. If your feature lands, great. If it doesn't, catch the next train, right? So we're modeling it on the Chrome OS release cycles. And that's actionable now, right? Like ROKM 10 was released, I think it was like August 26. Then ROKM 10.1 is already in RC. And six weeks from then, ROKM 10.2, 10.3, it's just going to go every six weeks. And so that's really a refreshing change that the entire AMD team has been able to like rally behind. And so now software is, you know, we'd like to think of it as transparent and the capabilities of the platform are just there for you to tap in and software becomes a mechanism for you to like tap into those capabilities of the platform.

Ryan Shrout: 

Some of the testing we did looked at Atom versus TensorRT as an example. I'm curious, from your point of view, how do you view these native stacks and their importance and value versus broadening or deepening the support for frameworks like VLLM?

Anush Elangovan: 

Very, very good question. So number one, our commitment to open source is unwavering and 100%. are like fully committed to all of the inference stacks, training stacks, so it's like VLLM, SGLang, we are just, you know, we will do what it takes and we do full upstream development of all of that. Separately, we also have our, like, you know, our hardware platforms that are evolving, right? So, for example, we may be adding new capabilities in the upcoming 450 generation or something like that, right? And those, even though we have full, you know, execution of VLLM and SGLANG, and we're trying to make sure that that's catching up, There are cases where speed of light of what we're doing in like co-design phases, et cetera, are tied to just having ownership of like, okay, I need to make this change plus this change. And you may have to take some shortcuts here or there. So the strategy you could think of is like you have a tentpole kind of like speed of light solution, but then you want to lift that up and then you get the breadth of all of the open source systems. The idea is that this just paves the way for us to get all of the VLLM, SGLANG, all of them up to what we think the platform can achieve. But the initial iterative loop can be unlocked with having something like Atom that we just go in and have to do some changes here or there. And it educates how we integrate into VLLM and SGLANG.

Ryan Shrout:
Now, Mitch, I want to come to you because I actually want to have you dive into a little bit of just a quick background on what the testing and the economics and the business outcomes of this report actually look like. I think the idea of this three-level approach of looking at things was kind of interesting. If you could just frame it up for us.

Mitch Lewis: 

Yeah, so we looked at three levels, the first being raw performance, and then infrastructure costs, and business outcomes, three different business outcomes. So they all kind of build on each other, right? And so just real quickly, on the raw performance, we ran three different models, GPT, OSS, 120B, Quan3Next, ADB, and KimiK2.6. just a raw performance throughput advantage on two to three models. And then on the infrastructure costs, that includes not just the 40% number that you're talking about before, dollar per tokens, or tokens per dollar, I should say. And that advantage extends to all three models, even the one that was a bit closer on the on the raw throughput. And then on top of that, we modeled three different business outcomes and found some pretty interesting and consistent results across all three of those as well.

Ryan Shrout: 

Yeah, it was interesting. I think we saw, we looked at like financial, was it financial document analysis, report generation, and kind of an interactive rag chat as the three different kind of use cases. And so the idea is to build out what are the scenarios for that, you know, kind of generalizing input output shapes and sizes and concurrency levels in order to simulate what would be deployed outcomes, right, for business use cases, right, which I think are super, it's a super interesting way to look at performance and maybe, you know, it's one step higher level, but I think more applicable to a broader range and audience of kind of, you know, procurement and IT buyers, which I think is pretty interesting. There was one data point, right, in the interactive serving side where the MI355X had, you know, roughly 2X the concurrent users within the SLA that we had defined, I think 15 seconds for a P99, right? And it had a 42% lower tail latency. I'm curious, Anush or Andrew, like why is tail latency such a interesting or critical metric for these types of user-facing AI applications in your mind?

Anush Elangovan: 

Yeah, so I think the important thing about tail latency is that it comes back to what is the worst case SLA that you could get in terms of an experience, right? So you want to define your bounding box of what would your experience be, right? And if your P99 or P95 gets to a point where it's unacceptable, then you could have the the, you know, banner numbers be good. But then your pay latency kind of like, you know, falls off and then you have issues that, like, just just experience issues right and so it gets you a good metric on how you want to measure the overall system experience. It's not just about one peak that you measured and then everything else is just jaggedy or not a good experience because then the end customer is going to try to use the system and they're like, oh, it's slow or it's bad. But then you just look at that one slice of it and then you're like, oh, this is Benchmark wise is good, but experience wise is bad, right? And that's why you really want to push on the P 90 plus on it being making nothing itself is representative of the experience that people would have.

Andrew Dieckmann: 

And I would add that ultimately partitioning to the SLA that you require, I mean, that dictates the economics of the solution ultimately, right? And so people are going to provision a solution for a given SLA and our ability to improve the economics of that is part of the value proposition of the platform. That is just part of us delivering the economic value that we intend to, to our user base.

Ryan Shrout: 

Yeah, excellent. I would say, you know, coming away from this, from this testing and this report, that it's, you know, it looks, the MI355X is an incredibly strong platform and continues to be, even though we've, we're talking about next generation platforms and future solutions, Andrew. Right. And so I would encourage everybody to go to Signal65.com and check out this report. But I always want to end and ask about some kind of forward looking type of question. Right. And as we talk about as these models, you know, they get larger, the context windows are getting longer. Memory capacity looks like an increasingly decisive trait. How do you view, you know, this shaping AMD's hardware and software roadmap? We've talked a little bit about it with with Helios, but I'm curious if there's anything else you can you can add about kind of how you how you view this future as as all these things shift.

Andrew Dieckmann: 

Yeah, I'll kick it off. I mean, it is a rapidly evolving landscape, right? And part of our strategy is we partner really closely with a number of leaders within the market. So, I mean, we have public partnerships with OpenAI, with Anthropic, with Meta. Just to talk about those three customers, I mean, the amount of co-design that we do across the board with those key partners understanding where their workloads are going, understanding how we can map our key advantages of the platform to derive the most benefit for them, for their workloads as they evolve. Then make sure that our platform can continue to evolve with their needs because the flexibility of the platform, the programmability of the platform, and ultimately being able to partition the workloads in different ways on the infrastructure, is an incredibly important part of our strategy moving forward. And so we've revealed Helios to the market. We have a really strong roadmap. We're working on actively the next two or three things following that for the coming two or three years. And we're excited to be able to share more about those as we get closer to the deployment of those systems as well.

Ryan Shrout:

 In Anoush, maybe I tilt the question a little bit for you of, you're often talking about the moat and how the moat has shifted. Do you see that shifting again and changing as we see these kind of shifts in AI and models and hardware capacity, or are you sticking with speed as the most important aspect?

Anush Elangovan: 

First, before I get to the speed, I would say the co-design part is super important, right, at the hardware layer. So at the hardware layer, co-design is a closed loop right now. It's a very aggressive closed loop. So it's like partners obviously have an input, but between software and hardware, it's a flywheel that, you know, it's on an execution cadence that is very impressive for what we're pulling off. And popping up to the software level, obviously, you know, I think this agentic AI is just, it's not that RockM is powering agentic AI, agentic AI is like rewriting RockM itself. And so RockM.ai, what we launched is like, I was setting up RockM on one of my systems and I just said, hey, install Rokum and optimize the Quen model. It was all natural language, right? And it knew how to pull in Rokum.ai skills. It knew how to install it, how to test it, how to fully optimize it. And it said, OK, what budget do you want to give? I said, OK, optimize for an hour. And it was running on the same Quen Next model that we were talking about. But for my client platform, right, and running Rockham, and it was able to like, get improvements that was meaningful. And I didn't open anything, you know, other than just my cloud code, text user interface. So increasingly, you know, the platform and what we're shipping is not just to humans, it's shipping to agents and then agents are able to like consume that in a way that that we haven't been able to do it. So the loop is closed even before it's a human. And you can start targeting that outcome, you know, those agents towards your outcomes, which is great, because now I have a very performant quantity running locally to, you know, consume some tokens, right? But popping that up a little bit, you know, coming back to, I do strongly believe speed is the mode that is, you know, it's like change is constant kind of thing. Uh, you just have to be able to be on your toes for what is coming next. Uh, because a year ago we weren't talking about agents in this context, right? Like it was, it wasn't like until December that it actually clicked. And a year ago, my, my perception was agents are like, you know prompts in a cron job right which is like okay yeah um but now it's like oh i mean everything is agentic there are like 10-12 agents that are running doing like very complex tasks that i just check in once every few hours or like once a day uh and it's continuing to do its thing so so all i'd say is like speed is the moat in terms of like you know a mindset and and then everything else is like ephemeral and and and you know um You just got to go with the flow to execute.

Ryan Shrout: 

It's been really interesting, even as we've pushed out our new Pinnacle benchmark and working with your team on those platforms, using agents for the optimization steps along the way for some of that has been very interesting to see. Again, to your point, just what's changed in the last six months. It's been pretty awesome. Well, Anoush, Andrew, I really appreciate you joining us to talk about this. It's been great. I think we could go on and on about this, not just this report, but a whole bunch of other things for a long time. But I know time is of the essence. You guys are super busy, so we appreciate it. Thanks for joining us.

Andrew Dieckmann: 

Ryan and Mitch, thank you.

Ryan Shrout: 

Thank you for having us. And for myself and Mitch, we'll close out this Signal 65 video insights. Make sure you go to Signal65.com to find this report and all of our other content. Thanks so much.

MORE VIDEOS

Anthropic’s AI Slowdown, OpenAI’s $1.2T Valuation & Salesforce’s AI Bet

Anthropic's "Pace the Frontier" essay set off a week in which nearly every major AI leader, and two heads of state, weighed in on whether the frontier should slow down. Patrick Moorhead and Daniel Newman read the incident as something larger than a safety debate. Whoever sets the pace of the frontier also sets the pace for open-weight models, enterprise custom models, and the trillion-dollar valuations riding on top of them. That contest, not the essay itself, is the through-line of Episode 320 of The Six Five Pod.

WD's Tim Rausch: AI Storage Demand Compounds With Every GPU Cycle

Western Digital's Tim Rausch argues that AI's compounding data growth, more than GPU capacity, sets the real ceiling on infrastructure scaling. He joins Matt Kimball at AI InfraSummit 2026 to walk through how total cost of ownership, archived data reuse, and WD's capacity roadmap are reshaping enterprise plans to scale AI over the next several years.

The Main Scoop Ep. 45: Attackers Only Need to Be Right Once. AI Just Made That Easier.

Chad Rikansrud, R&D Software Security Engineer at Broadcom, explains how frontier AI models are finding decades-old mainframe vulnerabilities in seconds, collapsing the complexity advantage that once protected legacy platforms. He outlines why enterprises need faster Security Intelligence (SECINT) ingestion and disciplined human oversight to keep pace with adversaries wielding the same tools.

See more

Other Categories

CYBERSECURITY

QUANTUM