NVIDIA on How Open Models Are Defining the Next Frontier

Nader Khalil runs an open model on his desk that delivers two to three times the speed of Claude or ChatGPT for the workloads he gives it.

Patrick Moorhead visited NVIDIA headquarters to speak with Khalil, Director of Developer Technology, for an AI Platforms, Ops, and Models Spotlight at the Six Five Summit: AI Unleashed 2026.

Khalil traces the open-model era to a lesson the market learned quickly from Llama: weights are ephemeral, but data is durable. That principle shapes NVIDIA’s Nemotron strategy, which provides open weights, architecture, and training data so enterprises can post-train models on their own proprietary information—where the lasting value lies.

On an NVIDIA DGX Station at his desk, Khalil runs the open-weight GLM 5.2 model at roughly 100 tokens per second, reserving frontier closed models for work outside clearly defined domains.

But the model itself may not be the biggest security concern. Khalil identifies the agent harness as the new battleground because its reasoning traces can expose an organization’s intellectual property. NVIDIA is addressing that risk through the proposed Secure AI Findings Exchange, or SAFE, developed with the Linux Foundation and the Open Source Security Foundation around a core principle: agent traces belong to the organization that generated them.

Key Insights:
🔹 Llama 1 and Llama 2 exposed a defining truth of the open-model era: model weights are reproducible, but the data behind them holds the durable value.
🔹 NVIDIA released Nemotron with an open-weights architecture and training data because Khalil calls weights ephemeral and data durable, giving enterprises a foundation to post-train against their own proprietary data.
🔹 Khalil runs the open-weight model GLM 5.2 on an NVIDIA DGX Station at his desk, generating around 100 tokens per second, faster than he gets using cloud-hosted frontier models for the same tasks.
🔹The agent harness, not the underlying model, may be the real security battleground, since reasoning traces flowing through the harness often contain an enterprise's own intellectual property.
🔹NVIDIA’s proposed Secure AI Findings Exchange (SAFE), developed with the Linux Foundation and the Open Secure AI Alliance, starts from a clear principle: agent traces belong to the organization that generated them.

Khalil expects closed-model token consumption and open-model adoption to keep rising together as expanding AI usage creates room for both to grow.

Watch the full video at sixfivemedia.com, and be sure to subscribe to our YouTube channel so you never miss an episode.

Explore more sessions from Six Five Summit: AI Unleashed 2026 at sixfivemedia.com/summit.

Disclaimer: Six Five Media is for information and entertainment purposes only. Over the course of this video, we may discuss companies that are publicly traded, and we may reference their equity share prices. Nothing discussed during this webcast should be considered investment advice or a recommendation to buy or sell any security. We are not investment advisors, and you should not rely on this content as financial advice. Six Five Media collaborates with technology companies and industry leaders to produce research-driven interviews and multimedia programming for enterprise technology audiences.

Nader Khalil:
We know that this is going to be a multi-model world, and we see where there's the bottleneck. The bottleneck is having great models that are fully open. We open everything, not just the weights and the model, but the data set as well.

Patrick Moorhead: 

Welcome back to the Six Five Summit 2026. The theme of this year is AI Unleashed, and it is amazing at all of the innovation that's being unleashed, and also the requirement to secure your agents. We are here talking AI platforms, open models, and pretty much anywhere the conversation goes here. It's really key, a lot of conversation about open models. It's funny, it seemed like one camp there was some consternation, but a whole lot more were supportive of it, given certain guardrails here, and I can't imagine a better Guy to have this conversation with. Nader, great to see you. Good to see you too. Thanks for having me. Yes, you run, you're director of developer tech, but I also knew you as CEO and co-founder of Brev that was acquired, that NVIDIA acquired about two years ago. So post congratulations on that. It's great to see you at NVIDIA.

Nader Khalil: 

Good to see you too. And yeah, it's an amazing place to be. We're really happy. Brev's been doing really well. We're getting to focus a lot on open source, which has always been the mission. So this has been awesome.

Patrick Moorhead: Yeah, totally. It's been great to get to know you outside of the video and break some bread with you as well, get to know you. Absolutely. I appreciate that. So, hey, I want to start. From Red Pajama to Llama 2, it's like every release drew the line between open and closed, right? I remember when Lama was going to solve all the world's problems. There were no very limited, even Chinese open source models. What did open models change about the industry?

Nader Khalil: 

Yeah. You know, open source broadly is the reason the industry innovates, right? It's kind of like a catalyst for both open and closed innovation. We've seen this through the internet. We've seen this through OS's. We've seen this through cloud. So it's no different here, of course. And I think sometimes it's worth reminding folks, especially in the broader conversation, that open source is actually just a way that software itself is shared. It's not an AI specific thing, but of course it applies to AI software as well. When you look at the early LLMs, the early large language models, it's interesting to see how the market was trying to figure out what to do. When Red Pajama came out, we had the data, the weights, the architecture, all open sourced. And then StabilityLM came out, and StabilityLM notably did not release the architecture. And I think stability AI was maybe scarred from not being able to capture so much value from stable diffusion, which was an incredibly valuable contribution and really helped the entire industry. And even if you just look at stable diffusion, that model, not an LLM, but how many startups were created because of that? I know a lot of the early Brev users were looking to train, fine tune custom stable diffusion variants for whatever their use case were. At the time, even, DALL-E really surprised us. DALL-E was the first time people were able to generate an image. But it was very limited. I got access because I was a YC founder previously. It was really hard to get off the wait list. And even then, you couldn't render faces. You could only use it through their playground. If you had an NVIDIA GPU, you could simply download the weights. You could make a custom fine tune. You could do something really cool and unique for your use case. And that's when we really saw a proliferation of these use cases. Open source really helped deliver that to a broad audience. And so going back to the language models, StabilityLM, I think, was trying to figure out how to capture more of the value. So they didn't release the architecture. Then LLAMA1 came out. And when LLAMA1 came out, they did not release the data or the weights. They just released the architecture, right? The weights were really hard to get. And then with Lama 2, they released the weights, and they released the architecture, but they didn't release the data. And I think the market very rapidly discovered where the value was accruing, right? Value was accruing at the data.

Patrick Moorhead: 

Yeah, it is amazing, even the definition. It's not just simple, is it open?

Nader Khalil: 

Yeah, totally.

Patrick Moorhead: 

There might be four or five different things that you can get access to. And it is, I find it humorous sometimes too, just humorous, but I guess an admiration of how open software is replaying itself with open models. So one thing, and maybe I incorrectly framed this, I mean, I just went with the meme, which was there's open models and closed models. So first of all, you gave a great explanation of what differentiates different types of openness. And I talk to CIOs who, they would like an open model, possibly to run on-prem, or where their intellectual property is sitting, and then using a frontier model for a lot of other things. Or maybe they've got data sharing agreements with the frontier models. Does that split still hold? Is the honest answer that it's opened and closed?

Nader Khalil: 

Yeah, so, you know, it's just to give a very tangible example. I have GLM 5.2 now running on a DGX station. Lucky you. I know. It is amazing, right? I have it running at 100 tokens per second. It's like two to three times faster than using Clawed or ChatGPT. But it's not, so while I'm using it a lot more, I am consuming a lot more tokens there. My consumption on Clawed and ChatGPT is still high. One example, something I really like doing, I've been using chat GPT to generate images of different brand assets, of different themes. Since it can generate images, I can interact with it and iterate very quickly on very visual footprint things, generate those as skills and go leverage those for these closed source models. In the case of enterprises, Enterprises have a lot of IP. You can use that and train a, you can fine tune, you can post train a model. And now you have a smaller model that can do tasks for cheaper. It can also do it way faster because it's a smaller model, because it's a smaller footprint. That does not mean that you don't use a closed source model. I think it ultimately comes down to is the task that you're trying to accomplish within the domain, right? If it's an in-domain task, it'll benefit a lot to have a smaller footprint model that's optimized, it's faster, that takes less resources to get that same task accomplished. But when you have an unknown task, if it's not clear if that's in the domain, that's a perfect opportunity to go use a closed source or honestly frontier intelligence.

Patrick Moorhead: 

It's funny, when you're an analyst, we make calls early on, and then we do victory laps on all the ones we got right, and then we ignore the ones we got wrong. And three years ago, it was pretty clear to my firm that there was going to be… heterogeneous models, like small models, open models, frontier models, and that's the way it was going to end up. Even understanding how enterprise and CIOs and even boards of directors, I mean, it's amazing. Some of the stuff that we think are new conversations around governance, we were having three years ago. around the data, so it's great to see.

Nader Khalil: 

Totally, and it's interesting, there are also use cases that simply don't work in the cloud, right? If you think about anything where latency matters, right? It's not just intellectual property or unique data that you have, but an ESPN replay. at, say, Wimbledon or for the NBA. Latency would matter so much there that you can't use frontier intelligence. You have to do something that is on-prem. So that's where edge hardware, edge AI, is really important. And if you're going to run something on the edge, that means you have to have access to the weights. You have to be able to download it and run it there. And so the only option is to take an open-source model and train it. It would be prohibitive for some of these companies to have to train a foundational model

Patrick Moorhead: 

I've watched NVIDIA operate since 1997. That's where I met Jensen when I was a hardware OEM. And watching the different businesses that you got in is really fun to watch. from GPUs to full stack infrastructure solutions to essentially architecting data centers. But you also are a heavy investor in models and open models and open weights, open data, open architecture, everything. And sometimes the, by the way, congrats on Nemotron 3.4 Lightning and Nemo Switchyard.

Nader Khalil: 

Yeah, it's super exciting stuff.

Patrick Moorhead: 

Love Switchyard. But sometimes I get the question, like, why does NVIDIA give this stuff away? Right? And I think I know the question, but I want to hear the answer. I want to hear in your words.

Nader Khalil

Yeah, absolutely. You know, earlier when you were talking about the different models and how the market quickly found value at the data, I think when Llama 2 came out, those Llama 1 weights that weren't released, people weren't really looking for them anymore. In a sense, the weights were ephemeral, but data was this truly durable asset. So you ask yourself, where does data accrue? And it accrues at enterprises, and it accrues for you as a person, right? When I open up Instagram on a fresh account, or if I open up TikTok on a fresh account, it really quickly finds out what I'm into, right? Do you remember old social media, like 10 years ago? It would be like, hey, do you like basketball? Or you like politics? You'd have to pick what your interests were. We don't do that anymore because it's very easy for them through the aggregate data that they have to figure out what my interests are based off of some usage patterns really quickly. So if we know that there is this interesting data, it is not enough data to train a foundational model. So the only option is to post-train something. So we build Nemotron because we know that this is going to be a multi-model world. And we see where there's the bottleneck. The bottleneck is having great models that are fully open. We open everything, not just the weights and the model, but the data set as well. And if you think about that, that means we're spending money to go and accumulate this data set that we also release. Lots and lots of money. Yeah. Yeah, absolutely. But it's the big bottleneck, right, is people have data. What can they do with it? And if we can make sure that there's a model that's fully open so that people feel comfortable customizing or taking any part of it. By the way, you can just take the data and train your own model. I believe it was ServiceNow who did this. We celebrated that. That was awesome. We know that there's so much value to be unlocked when people are able to train models off of the data that they have, when enterprises and even individuals. We released this model and tooling around it, because post-training is also not the most straightforward thing. It's a little complicated. knowing whether a model performs better is also a little complicated. Every time there's a new model out, the first thing you see on X is the big vibe test of people trying it and being like, I don't have a very strong eval, but I think it's better. Exactly.

Patrick Moorhead: 

You're doing it, though, to accelerate innovation, maybe with end customers or end developers who can't do it all themselves.

Nader Khalil: 

Is that fair? Absolutely. Yeah. And there's a lot of valuable data that You know, I think that it's not sufficient to train a foundational model. Right. And maybe there are not enough, like it doesn't justify the spend of training a foundation model. As you mentioned, we're spending a lot of money doing this. Right. Yes. And so you get a little discount on the hardware, but it's still a tremendous investment. Absolutely, yeah. And so, you know, if we're able to make sure that there's a really good model, you know, lightning is really fast. I think what's really interesting, there's this kind of Pareto frontier of on one axis you have intelligence and on the other you have, you know, speed. You want the, sometimes you might need the fastest, like latency matters so much that it's okay if it's not as smart at everything. And sometimes you just need the most intelligence and you're okay with a latency hit. everything comes down to the use case specific, like it needs to be use case specific. And so, you know, where you are on that Pareto frontier is ultimately where you would pick the right model for the right task. And from that lens, it becomes really clear that we need different models. It might not even just be one post-trained model. You might be using a frontier model, but then also an array of models. So that way, if you know that latency is really what matters here, you could start there, right?

Patrick Moorhead: 

Yeah. What's exciting, I think, is you've got models for the data center, the industrial edge, even automotive. And that's pretty exciting stuff. And definitely seeing acceleration of innovation with your models, which is quite frankly, some people were saying there's no way you can do this. But given some of the numbers you're putting up there, by the way, I've also talked to some of your customers that said, listen, NVIDIA's models actually outperformed what any benchmark could ever tell me, right? And I thought that was really interesting, that real world usage was better than even the numbers.

Nader Khalil: 

Absolutely. I think this is something that is missed a bit with the open source models is it's not, you know, just use this out of the box. It's, you know, this is a tool for you so that you can customize it. If you go and take Nemotron, one, it is amazing out of the box. It is really fast. It is really smart. But also when you take your IP and apply it, this model is going to be more performant than anything else there. And that's a unique advantage that you have as an enterprise.

Patrick Moorhead: So hey, let's dive into agents. And you know, it's interesting, agents is actually a system, a collection of different things from the model, the harness skills, runtime, and you know, you're on record calling harness the underappreciated piece. And quite frankly, we've also watched agents, not yours, of course, leak out into the environment. There's been a lot of news about that. We saw it at OpenAI and even AISI. Why is the harness the ultimate security front or battleground?

Nader Khalil: 

The harness is where the model gets used. To your point, the agent is comprised of a few things. It's the harness. It's the model. It's the tools and the skills it has access to. It's the runtime. You need to look at this thing as a system. It might be helpful to even think about how we got to harnesses. ChatGPT innovated a lot outside of the model. It was not just an LLM. they made it very easy to prompt the LLM, right? That chat interface. They made it really easy to use multimodal prompts. You could send photos. The first time I sent a photo to ChatGPT, that felt crazy. I remember I was at, I think it was the Denny's or something where they had like, guess how many marbles are in the jar? And I took a photo with ChatGPT. I was like, oh my God, this thing, this game is dead now. But slowly they added more. The second they added memory, that felt amazing. That means- That game changer.

Patrick Moorhead: 

Memory game changer.

Nader Khalil: 

Game changer. Yes. Now I don't have to remind it of my previous chats. It just knew about what we talked about, right? You could actually ask ChatGPT, like, tell me, find my 10 strengths and weaknesses. And it does an unfairly good job because you've talked to it a bit. Then they added more. They added web search. That was huge. Before it had access. And that was the first time it had this amazing tool call. Before it could use the internet, there was this perspective that maybe we would never have another programming language. If you introduced a new programming language, then developers wouldn't be able to use AI to help them code it unless it was trained on that. But the internet web tool changed everything. Now, ChatGPT could go and search the web, and in fact, it would prefer to because it wants to get you the right answer. It's not just going to depend on what it had already been trained on. Then my experience at the time when I'm coding, I would go to my code base and I would copy code files. Then I would go to chat.gpt and I'd paste them. If it needed more context, I'd have to go manually back and forth. Cursor was an amazing innovation where they brought the file system directly. They put chat GPT into my file system. Now, when it needed new files, I didn't have to go and put them. It could just go and find them. So if you think about all the innovations that are happening here, that's all the harness. And that's kind of why this is the new frontier. The harness is what made the model so useful. It wasn't just the model. Of course, we need amazing new models. But part of what happened this year is we reached an inflection point. It felt like it was around December to February where we got really good models. Of course, we had Opus and 5.5, GPT 5.5. But then we also got really good harnesses. We got Cloud Code, and we got Codex, and we got OpenCode, and we got Pi, and OpenClaw. And it's them working together that really unlocked all this. And so to kind of answer your question about security, something goes wrong. If there's a bug on our server, the first thing I do is check the logs. The harness, the traces, the reasoning traces, if something happens, if an agent escapes the sandbox, we need to collect that information and share it quickly. That way we can all go and learn from it.

Patrick Moorhead: 

It's funny, Jensen on social media, okay? The first time he weighed in was on supporting open models. And kudos to Nvidia for leading that charge. I do a lot of stuff in Washington, D.C. I know you had made a couple visits there as well. A lot of conversations going on. And then my biggest question coming out of that, well, what about security? How do you secure something that you can't actually see the source code like you can in open software? And then boom, NVIDIA comes out with a coalition of companies on security. And you're proposing secure AI finding exchange with Linux Foundation, with the Open Source Alliance. And I think this hit at the right time related to IP when you had a lot of senior executives coming out saying, hey, I don't know if I'm going to use this model or this service because through agent traces, right, they can essentially see my IP. And we've seen some frontier model companies stand up completely new businesses. And yeah, that got the board of directors and CEO conversations hitting again. But let's get practical. What should organizations be doing today? How do they get the incredible benefits of all this amazing AI technology and innovation, but also protect their IP? Where do you start?

Nader Khalil: 

Yeah. I mean, I think the multi-model world is key for this, right? A lot of times you might need to go externally to go tap into an expert. And that's what you can do with these frontier intelligence models. But I think in a lot of ways, the industry is kind of learning some best practices that we already knew. And one of them is making sure that you can own and run your stack. It matters where my code runs. It matters where the services that I use run. And if your harness is going to have access, that's where you're going to put a lot of your IP in, so that you can use AI for the best use cases. Then it does matter that you own the harness traces, or those agent traces. Making sure that if you're doing the more sensitive things, you can do it on open models that you have customized or not. Like the GLM 5.2 example I gave that's running on my DGX Station. I didn't customize it. It's great. But I still have my high level planning with Claude or chat GPT before I then go to the GLM 5.2.

Patrick Moorhead: 

Yeah, I'm glad you brought a DGX Station. By the way, for about a year I had a DGX Spark sitting in my closet. And then when my son moved out, he took it with him. Not good. But he does give me special SSH abilities on that thing. But it is amazing how this whole notion, I call it distributed AI, really whatever modality, whether it's on a computer, on a workstation, on-prem data centers, Colo, NeoCloud, hyperscaler data center, and pretty much everything in between. But what does frontier intelligence running on a black well in your office actually look like? You talked about it a little bit, but how do you partition what you do here versus maybe what you do in a big NVIDIA data center versus Frontier?

Nader Khalil: 

It's a great question. I think the first, again, it all comes down to the use case. And when you talked about Switchyard, which is a model router.

Patrick Moorhead: Glad you're going this direction.

Nader Khalil: 

This is good. When there are multiple models, the next frontier becomes how do you know which one to use and when? And you're seeing a lot of companies now create model routers. I think that's clearly the answer to how to leverage the best of the multiple models. And so when you have something that's sitting under your desk, I think one, it's nice. It's there. You know what IP is. is on the box, you know where it's going, it's just under your desk. The latency also is great. So for any use case where latency matters, having the edge hardware is going to matter a lot. The other thing is it's cheaper than getting a data center, but it is a data center GPU that sits under your desk. And so I think, again, it goes back to whether it's picking the model or picking the hardware target, all of that comes down to the use case. And so you can batch multiple users on the station. So what we've done is we've created API keys for the GLM 5.2 model that's running on it, and different team members are all able to use the same DGX station. So one of my goals is that the core Brev team, which is about eight-ish people, are all going to be able to use one DGX station. So that's going to be that's kind of the experiment that we're running right now. But we, of course, also leverage our internal compute cluster. We also leverage the cloud. We also leverage these frontier intelligence models. And so the the harness and the router are definitely what's taking the most or what I'm thinking about the most now to see how can we more seamlessly leverage all of this.

Patrick Moorhead: 

CIOs are super excited about this. And where this is different from, you know, there was this thought in the early days of the cloud, hey, I'm going to burst to the cloud. We never actually bursted to the cloud ever. What we did is we ran the same application with the same data on-prem, let's say if you were a retailer, and in the public cloud. But the cool part about, you know, you add Switchyard, you add genetic orchestrators, and it is possible, right? Totally. You have a security policy, a data policy to determine any reasoning, and this is also in Switchyard, kind of a logic of hey, where does this need to go to run to get the best response? I think you had talked a little bit about workflows that you know well work really well on local. And it literally is bursting to a frontier model or to a gigantic on-prem or Neo cloud.

Nader Khalil: 

Totally. And there's actually a lot here of, you know, when you burst, you can do PII redaction, right? does. Your model router can identify what is the IP so that when you burst, you make sure you only burst with what's safe to go. And so there's a lot of that's a great point. I haven't thought about how much we've talked about burst to cloud in the past, but how much I haven't actually gotten to experience it until maybe now.

Patrick Moorhead: 

But it's like, no, it's real now. Yeah, I've never had the tools. Totally breaking up an application didn't make sense. But when the application starts in a certain place, and it's more about the logic flow, totally of it, you can totally, totally do that. I get these questions a lot. Pat, okay, we heard the same thing 10 years ago, go about about the cloud. And you know, bursting never, never actually Yeah. Hey, I want to dial. It's been a great conversation so far, but I've got one more for you. I want you to put your futures cap on. Let's say we're sitting here in Endeavor in 2027 for the summit. What's the one thing about open models that you think will have surprised everybody?

Nader Khalil: 

One, I think it's going to be really clear in a year from now. However many tokens you've consumed from the closed source, from the foundational frontier intelligence models, you're going to consume more than those from open source. I think we're all going to have felt that within a year. But that doesn't mean that you're not going to be using these closed source frontier models. You're going to be using more of both. I think something that's really hard for us as just generally as humans is when things are nonlinear. We're in a huge scale up right now. And so the pie is exploding. And so any sort of use that moves to one doesn't mean it had to take from another. The pie itself is expanding. And so I think we're going to see that. I think we're also going to credit open source for making sure that we're able to leverage agents safely and securely. Open source, honestly, is just the concept that it's a way of sharing ideas. If you look at the computing industry, compute was invented in Cambridge on the East Coast, but hippies on the West Coast started in Silicon Valley. And a lot of that was just hippies trying to share ideas and thinking that the ideas themselves should be free through software. And so the learnings that we're going to have as an industry, making sure that we can get those collectively together really quickly. We're going to look at open source models of being able to have delivered that.

Patrick Moorhead: 

Yeah, I buy in 100%. And whether it's sovereignty, whether it's data protection, whether it's cost, or even innovation, right? I mean, the capabilities of these open models, it's not like, oh, you know, let's wait nine months. Totally. Right, afterwards. So it is keeping up. And what that does is give people confidence that they can rely on open model innovation, which then will diffuse out to the robotics and industrial edge. Absolutely. As well. And you know, it's funny, I have these debates a lot. I mean, it's classic Jevons paradox, right? You increase the capability, you lower the cost, more people are going to use it. 2% of consumers pay for AI today. I saw that. That's crazy. And a typical, let's say, bank, they have 50,000, some have 50, but I'm just going to say 10. How many workflows are in there? Probably hundreds of thousands of different workflows. And they're barely scratching the surface. They might have 200 useful workflows. in there right now. So they have a long way to go as well. And India, I agree with, is frontier models will still have a huge place. I mean, I remember when virtualization was invented, that was going to collapse the number of servers. No, it went up by 10X, right? Smartphones were going to take out PCs. Tablets were going to take out PCs. No, they just created new markets and new things that added value. So open models being the dominant thing that people use. By the way, it's an easy one when you look at the client, okay? Like a DGX capability. Because once people start, you know, loading that on, you know, NVIDIA Notebook with that same capability in a few years, that's going to blow people away. So, man, great conversation, Nader. This went great. I really appreciate you being part of this year's Summit.

Nader Khalil: 

Thank you so much for having me. We're a huge fan of you. The industry is lucky to have you. Thank you so much. I appreciate it. That's kind.

Patrick Moorhead: 

So this is Patrick Moorhead here at NVIDIA headquarters, Endeavor building. Hit all of our Six Five Summit 2026 content and also look at all the NVIDIA content on the More Insights and Strategy website and all the Six Five podcasts that we cover. Take in, hit that subscribe button.

Speaker

Nader Khalil
Director of Developer Tech
NVIDIA

Nader Khalil is Director of Developer Tech at NVIDIA where he leads Open Source & Agent Marketing. He was the CEO & co-founder of Brev.dev, which was acquired by NVIDIA in July 2024. 

Nader Khalil
Director of Developer Tech