Redefining AI's Limits: Samsung on Memory, Power, and System Design

AI has moved into production, and the constraint everyone assumed was compute is turning out to be memory bandwidth. Enterprise leaders are already measuring token throughput per employee, but few understand that throughput traces directly back to how fast data moves between memory and compute.

Opening the Semiconductor Track at Six Five Summit: AI Unleashed 2026, Patrick Moorhead is joined by Paul Cho, President and Head of Samsung Semiconductor U.S. and Corporate EVP of Samsung Electronics, to unpack how memory, packaging, and power have stopped being separate engineering problems and started functioning as one coupled system.

Cho walks through why Samsung builds its HBM4 base die in-house rather than outsourcing it, arguing that owning the full integration chain delivers power and signal-integrity advantages that a handoff model can't match. He details Samsung's progress on HCB technology, which has produced a 20% improvement in thermal resistance, and previews zHBM, a 3D z-axis stacking architecture that couples memory directly to logic and targets products for 2029 and beyond. Cho also explains why production AI workloads compound inefficiency daily in a way training never did, making memory requirements the starting constraint that determines which compute and packaging choices are even viable. The conversation closes on Samsung's newly announced second fab in Taylor, Texas, built to expand U.S. capacity for sub-2 nanometer nodes.

Key Insights:

🔹 Token throughput is a memory bandwidth problem, not a compute problem. The number of tokens an AI system can serve per second, and how many concurrent users it can support, is determined by how fast data moves between memory and compute, a relationship Cho says remains poorly understood outside the semiconductor industry.

🔹 In-house integration is Samsung's stated differentiator in HBM4. Building the base die on Samsung's own 4-nanometer process and managing the full integration internally avoids the handoff problem competitors face when outsourcing the base die.

🔹 Usable bandwidth per watt has replaced raw bandwidth as the metric that matters. HCB technology has already delivered a 20% thermal resistance improvement, and Samsung's forthcoming zHBM architecture aims to push sustained bandwidth further within the same power envelope.

🔹 Production AI has inverted the system design sequence. Where compute used to get sized first with memory and packaging worked out afterward, memory requirements now determine which compute and packaging choices are viable at all, because inefficiencies compound continuously rather than once per training run.

🔹 Samsung's second Taylor, Texas fab expands U.S. advanced-node capacity. The facility will produce nodes at 2 nanometers and below to support American AI infrastructure customers and is targeted to come online within three years.

Cho's closing argument reframes the industry's blind spot: tooling and benchmarks are still optimized for a compute-bound world, while the systems that will define the next decade of AI infrastructure are the ones built around memory-centric compute from the start.

Watch the full video at sixfivemedia.com, and subscribe to our YouTube channel so you never miss an episode.

Check out all of the Six Five Summit: AI Unleashed 2026 content at sixfivemedia.com/summit.

Disclaimer: Six Five Media is for information and entertainment purposes only. Over the course of this video, we may discuss companies that are publicly traded, and we may reference their equity share prices. Nothing discussed during this webcast should be considered investment advice or a recommendation to buy or sell any security. We are not investment advisors, and you should not rely on this content as financial advice. Six Five Media collaborates with technology companies and industry leaders to produce research-driven interviews and multimedia programming for enterprise technology audiences.

Paul Cho:
The thing is, memory bandwidth is what determines what an AI system can actually deliver and what it costs to deliver. In 40 years, the applications that win won't be the ones with the most memory. They'll be the ones that figured out how to use the bandwidth they have.

Patrick Moorhead: 

Hey, everybody. Welcome to the Six Five Summit 2026. The theme of this year is AI Unleashed. And if you've been paying attention at the beginning part of the year, you know exactly what I'm talking about. We're going to kick off our semiconductor track with a conversation about infrastructure and how it is changing as AI moves deeper and deeper into production, pretty much spreading everywhere on every continent in every form you can even imagine. It's no longer just about adding more compute. It's really becoming a game of memory bandwidth, power, packaging, and how the entire system comes together is increasingly defining what AI can deliver. And we are joined by Paul Cho. Welcome back, president and head of Samsung Semiconductor US. and Corporate EVP of Samsung Electronics. Welcome back to the show. Thank you for having me. Yeah, it's interesting. Since the last time we talked, I mean, memory was important, but it's amazing how much more memory has come into the foresight here. And you have a huge job at Samsung. You lead memory, foundry, and systems LSI. So it's great to get your perspective on here. So you've talked about every CEO will be thinking about the number of tokens available to their people. And it's interesting, you have people that are token maxing and some who are focused on token efficiency. That's been a conversation that has recently come up. And what's interesting is that's directly correlated And I'm not sure people who are buying AI fully understand how correlated this is to memory. And I'm curious, why does this disconnect still exist?

Paul Cho: Well, you know, first of all, Pat, you are very right when you say that CEOs are measuring the use of tokens. They are already doing that. I'm doing that. You know, they look at individuals, individual teams, you know, in terms of number of tokens or the budget allocated to them. This is happening already. So we are in the full AI era. You know, what's not fully understood in the industry is that, you know, what really matters in AI inference service, the number of tokens produced per second as needed by the organizations is determined by a memory bandwidth. The bandwidth here is what determines the token throughput, how many concurrent users or requests a system can serve. How fast it responds, how much context it can hold. All of it traces back to how quickly data moves between memory and compute or memory and network. Yeah, this relationship is not well understood outside of semiconductor industry right now.

Patrick Moorhead: 

Do you think that disconnect exists? Is it more of a marketing and education thing or is it something different or it's just so new that it's just going to take time for people to understand that?

Paul Cho: 

This is a really under the hood thing. There are people out there building systems. They understand the bandwidth at the system-level concept, but it all boils down when it goes into under the hood, especially inside the package of a system packaging and inside the chips. And what really matters is the data bandwidth between different components. And here, because data are all stored in memory, the memory bandwidth is what really matters from our viewpoint. Yeah, I see the same thing.

Patrick Moorhead: 

So, hey, I want to talk about packaging. So when I was in the semiconductor industry, maybe about 15 years ago, packaging was a little bit of an afterthought. And the thought of putting different devices in one package, it would be like, okay, it's going to be too high power or too low performance. And boy, have things changed, right? And as memory and logic are getting more tightly integrated, we are seeing the performance gains and the efficiency gains at the packaging level. What's driving advanced packaging to become a first-party citizen, I will say, of AI system performance?

Paul Cho: 

So process node scaling to increase the amount of compute or increase the amount of bandwidth, this alone can no longer deliver the power and performance gains AI demands. So this packaging is where the system gets connected and composed, mixing dyes, nodes, and memory into one coherent system. So the gains are increasingly at the integration layer, not individual transistor layer. So this is important, and we understand that at the concept level, but when it comes to memory, we at Samsung have this position that We want to provide the whole components of an advanced high-bandwidth memory in-house, and the way we integrate them together also in-house. So Samsung's HPM4, I'll take it as an example, the base die there is built on Samsung's advanced 4nano technology in-house. So that way, we put together everything in-house. This gives us edge in power, signal integrity, performance, bandwidth, and everything. You know, there are companies out there that have to outsource the base die, and they manage a handoff, whereas Samsung, we manage an integration to realize the design objectives at the system level, chip level, memory level, and everything.

Patrick Moorhead: 

Yeah, I think it does give you an advantage, and I brought this up in our video that we did before. It is a differentiator. Now I want to talk about bandwidth here. So bandwidth has roughly doubled each generation of HBM, and power demands have gone along right with it. You can build more capacity, build more chips, but you can't just instantaneously get more power. So how is the industry keep scaling the bandwidth when the power keeps increasingly becoming that ceiling?

Paul Cho: 

So here, the number that matters is no longer raw bandwidth, and it is not watts per rack. It is the usable bandwidth per watt. How much useful work the system delivers inside a fixed power budget or envelope. That connects directly back to our previous throughput question. And we are making progress in different fronts. Let's take HCB as an example. Last time, in our last interview, we talked about HCB a little bit there. We are working very hard to realize HCB in our HBM product in the following generations. And it gives us a near-term answer here, and it's confirmed. We see 20% thermal resistance improvement there. A better thermal headroom translates into higher usable and sustained bandwidth within the same power envelope, which is exactly the level it matters. And as a next step, we are looking very seriously into the new memory architecture called zHBM. It's a directional architecture. It's a 3D z-axis stacking. And we have a lot of traction right now because all the customers come and tell us that this is the way to go. We want to put high-bandwidth, low-power memory on top the logic they are designing. So by tightly coupling the memory bits and logic transistors in the shortest amount of distance between them, that highlights the importance of the CHL-HBM concept. And we are working very hard to realize this as a real product right now.

Patrick Moorhead: 

By the way, the industry thanks you for new innovations like this. And I don't think people will be surprised that they're is a different modified architecture for HBM. Just for our listeners, Paul, are you going to be talking about when a device with zHBM might be in the market?

Paul Cho: 

You know, this is a work in progress. We are actively defining the specification for the first generation of CHBM. And I do expect the future products based on the concept of CHBM will be, you know, for 2029 and beyond.

Patrick Moorhead:

For everybody out there, that is not a long time in semiconductor time, given, first of all, defining the spec, making it manufacturable, making sure it's performant, and then doing it at these incredible high volumes that Samsung does. Yeah, I was thinking 2030. You talked about 2029, so you did a pull-in on camera, so that's great.

Paul Cho: You know, everybody wants likes pull-in. 2030 to 2029, maybe 2028, you know, but everybody has to work together really hard to realize the benefits promised by the HBA.

Patrick Moorhead: 

Yeah, I always like to say you can't vibe code a foundry, and you can't vibe code a semiconductor architecture. So I think sometimes it just takes a while. Once people step back and look at the value chain and the fact that it takes three to four years to even build a foundry, will look at different architectures probably 10 years out. You pick some that look like the leading new technologies and then you run with them. So, it surprises a lot of people out there. Let's pull this all together, right? Memory bandwidth we've talked about, we've talked about packaging and we've talked about power, right? How… Is AI moving into production changing the way that we need to think about the total system design? It used to be, okay, I'm going to build this compute and there's a JEDEC memory standard that we're all going to agree on. And it was not a serial process necessarily, but it wasn't done as a complete system, bringing in memory bandwidth, power, and even packaging.

Paul Cho: 

So here, the real shift isn't the three constraints individually. It's that they are now very well connected and coupled. So training tolerated inefficiency to a certain degree. finite runs, batchable, overprovision, and pay the cost once. But production runs continuously at unpredictable demand. So every inefficiency in bandwidth, packaging, or power compounds daily instead of once. And as you properly said, they can't be solved in sequence anymore. More bandwidth needs denser stacking. Denser stacking generates more heat. Heat caps how much bandwidth you can actually sustain. So solve one in isolation, and you've just moved the constraint to the next one. HCB, as we discussed, is the proof point already in the deck. It wasn't a packaging fix or a power fix. It was rather both problems forcing a single answer. The design sequence has inverted. It used to be pick compute, size memory to fit, figure out packaging and power afterwards. Now, memory requirements are often the starting constraint that determines which compute and packaging choices are even viable.

Patrick Moorhead: 

Yeah, I mean, that forces this in parallel system design for sure. And as I look back in history of a lot of the different leaps we've taken, and whether it's mainframes to minis, minis to client server, mobile, local, social, the internet, at some points the industry comes together and has to work in parallel to get increasingly complex things done quicker. And as opposed to this serialized event and then the industry, we disaggregate, okay? And then we go back to piece parts for economy. But there is no doubt in my mind, in addition to memory and packaging and power correlated with the compute, but also the networking. you know that entire system. I'm seeing modern designs and co-innovation even even with customers.

Paul Cho: 

You are absolutely right. Our customers are right now asking Samsung in earlier into the system design itself. So that's the market catching up to what I have been talking about. So we already talk about the memory architecture and system architecture for 2029, 2030 and beyond. So that's how it's being realized right now.

Patrick Moorhead: 

Yeah, it makes sense. And yes, this co-design is somewhat of a new concept here and everybody's getting used to it. I remember, Paul, remember when we used to have a new GPU compute every three years? And then it was two years, and then it's every year. Every year, multiple trips, yes. Yeah, and we bifurcated too into, it's not just a GPU world as well, it's an XPU world. And the complexity that goes into this, I'm amazed that as an industry, we can keep up, but the industry is doing a pretty good job and getting quicker as we get in there. So I want to talk about, I always love these type of questions. I always like the, What's the industry getting wrong? If we look forward three to five years, what is the industry confidently getting wrong right now that will look obvious in hindsight related to how people are building AI infrastructure today?

Paul Cho: 

We are still building software for a world where compute is the bottleneck. The frameworks optimize for flops or tops, if you will, for modern AI. The benchmarks measure tops. Engineers are trained to squeeze to extract more tops. And the hardware has already moved. The thing is, memory bandwidth is what determines what an AI system can actually deliver and what it costs to deliver it. In 40 years, the applications that win won't be the ones with the most memory. They will be the ones that figured out how to use the bandwidth they have. And the companies building the tools for that will define the next decade of AI infrastructure. So the term that I love to reiterate is the memory-centric compute, where you start from the memory, define the bandwidth, and how to use that bandwidth in the most effective way. That's one way to look into it.

Patrick Moorhead: 

already talking with companies who are looking into this. I think they recognize this, particularly in the age of agents. It was one thing to do in LLM, but agents changed the game, I believe, related to memory bandwidth. And your point that, hey, more memory isn't always the best, that is true. There are things you can do inside when you've got better memory bandwidth to give you better answers, right? And that's sometimes, like if you look downstream, Paul, getting the best answers, whether you're an enterprise or you're a consumer services company, having better memory bandwidth gives you more agentic, better agentic outcomes. I can get a line of sight even to what you're doing in memory to the reason that we're spending all this money on AI which is the outcomes.

Paul Cho: 

So already we at Samsung are selling not only selling bits but we are selling bandwidth today and that will enable the next generation of AI infrastructure.

Patrick Moorhead: So Paul, I'd love to end on a manufacturing high point here. I am here in my studio in Austin, Texas. And while I know Taylor isn't Austin, Texas, even though I like to pretend it is, because I'm so proud of it, you have a huge facility here, but you also made a huge announcement about a second facility. Can you tell our viewers a little bit about it?

Paul Cho: 

Right. So we announced last month during our Q2 earnings call that we have a plan to break ground of our second fab in Taylor, Texas, very close to Austin. So this fab is to expand our capacity to produce advanced technology node, including two nanometer and beyond to support a lot of US based customers. So we are thrilled to be at this stage to announce and start working on this new fab. And we are so excited, and we'll be working very hard to make this facility online in the next three years.

Patrick Moorhead: Yeah, kudos to Samsung and not just because it's near Austin and makes me proud, but also it's a lot more difficult and more expensive to build a foundry here in the United States. So that's a huge commitment. And based on a lot of your customers' desires to balance out their supply chain, you're delivering the goods. Thank you. On behalf of your customers, I work with all of your major customers and they thank you here. But very exciting.

Paul Cho: 

I'm excited, too.

Patrick Moorhead: 

Thank you for kicking off the semiconductor track of the Six Five Summit. Really appreciate your time. And it's always great to have you on the show. Thank you, Pat. To our viewers, don't forget to hit subscribe, follow us on social media, check out all of our Six Five Summit content, all of the memory content, also all the Samsung content, and stick around for the next semiconductor track session. See you next time.

Speaker

Paul Cho
President and Corporate EVP, Samsung Electronics
Samsung Semiconductor

Sangyeun “Paul” Cho serves as President of Samsung Semiconductor and Corporate EVP of Samsung Electronics, responsible for Samsung's U.S. semiconductor business, which includes Memory, Foundry, and System LSI. 

With over 18 years at Samsung, Cho has held leadership roles in Korea and the US. He founded the Memory Solutions Lab, pioneering innovations like in-storage computing and multi-streamed solid-state drives, which are now industry standards. Cho also led server-class SSD software engineering and security initiatives. Previously, he spent nearly a decade in academia as a professor at the University of Pittsburgh, focusing on computer architecture. Cho holds a Ph.D. from the University of Minnesota and a B.S. from Seoul National University.

Paul Cho
President and Corporate EVP, Samsung Electronics