Home

QumulusAI CEO Michael Maniscalco on Nasdaq Listing, Hyperspeed Infrastructure, and the Limits of Hyperscale

QumulusAI CEO Michael Maniscalco on Nasdaq Listing, Hyperspeed Infrastructure, and the Limits of Hyperscale

QumulusAI CEO Michael Maniscalco joins Six Five following the company’s NASDAQ debut to explain why traditional gigawatt-scale infrastructure cannot keep pace with AI demand and how a portfolio of smaller, faster-to-deploy compute pockets could reshape enterprise infrastructure strategy.

AI demand is growing faster than infrastructure can keep up. While hyperscalers invest billions in massive data center campuses that may take years to build, a new generation of infrastructure providers is focused on speed rather than scale. QumulusAI calls the approach "hyperspeed," deploying distributed AI compute in months instead of years to close the widening gap between enterprise demand and available capacity.

Daniel Newman, CEO and Chief Analyst at Futurum, spoke with Michael Maniscalco, CEO of QumulusAI, following the company’s NASDAQ debut under the stock ticker $QMLS about why traditional gigawatt-scale infrastructure development cannot move at the speed enterprise AI now requires.

Maniscalco explained why QumulusAI is building a portfolio of smaller, distributed compute pockets rather than relying exclusively on massive campus deployments. By targeting sub-10-megawatt capacity, the company can pursue power availability, data center space, and site approvals that are often easier to secure than the resources required for gigawatt-scale projects.

The conversation also examined the economics of inference, the growing role of hybrid and distributed infrastructure, and why organizations should match each workload to the right model and token economics. Maniscalco argued that most enterprise queries do not require the most expensive frontier model available, making flexibility, cost, latency, security, compliance, and data locality increasingly important infrastructure considerations.

Key Takeaways:

🔹 AI infrastructure development is not keeping pace with demand. QumulusAI built its operating model around “hyperspeed” to narrow that gap and deliver capacity closer to the timelines customers now expect.

🔹 Smaller compute pockets can move faster than gigawatt campuses. QumulusAI has been scaling primarily in the sub-10-megawatt range, where power availability and data center capacity can be easier to identify and secure. 

🔹 Not every workload needs the most capable or expensive model. Maniscalco compared frontier models to an F1 car: powerful, specialized, and unnecessary for many everyday jobs. Enterprise infrastructure strategy should account for the value of each token based on workload requirements such as model quality, latency, geography, security, compliance, and timeliness.

🔹 Flexibility, cost, trust, and speed are shaping compute placement. Enterprises are evaluating where to run workloads based on architecture, duration, capital requirements, access timelines, token economics, data privacy, and locality.

🔹 Distributed capacity can support future sovereignty and compliance needs. A footprint assembled through smaller compute pockets can give customers more options for placing inference workloads near required data, users, or regulatory jurisdictions as privacy and localization requirements become more demanding.

🔹 The AI infrastructure market remains in its early innings. Maniscalco expects demand for capacity to continue rising as models improve and AI expands into more production workloads. Newman pointed to more than $200 million in contracts signed by QumulusAI since June as a signal that its differentiated supply strategy is gaining commercial traction.

Watch the full video at sixfivemedia.com, and be sure to subscribe to our YouTube channel so you never miss an episode.

Disclaimer: Six Five Media is for information and entertainment purposes only. Over the course of this video, we may discuss companies that are publicly traded, and we may reference their equity share prices. Nothing discussed during this webcast should be considered investment advice or a recommendation to buy or sell any security. We are not investment advisors, and you should not rely on this content as financial advice. Six Five Media collaborates with technology companies and industry leaders to produce research-driven interviews and multimedia programming for enterprise technology audiences.

Transcript

Michael Maniscalco:
I look at the pace that infrastructure has been developed. I look at the pace that compute and data center capacity has been brought online. That doesn't oftentimes line up with the speed that the market, the AI market specifically, is moving at. To us, this is all about hyperspeed.

Daniel Newman: 

Hey everyone, welcome to The Six Five. I'm Daniel Newman. Today, I am joined here by Michael Maniscalco, CEO, Cumulus AI, just went public on the NASDAQ. Mike, welcome to The Six Five. Excited to be here with you today. I mean, what an exciting time.

Michael Maniscalco: 

Yeah, absolutely. Thanks for having us today and a really exciting time for us, for the company, for all the work we put in, but also an exciting milestone for us as we move forward. I think this space, this AI space is moving at an incredible velocity and it feels good to get to this milestone to help increase our velocity and our pace because AI is moving just so fast and there's a lot more to it.

Daniel Newman: 

It's like you're searching for the word. It's exponential times. feels like infinite demand. I call it insatiable right now. And it's interesting because you sort of have a tale of two markets, right? You have the people that are implementing, deploying, and using AI, and all the words I just used apply. And then you sort of have the investor public market side, which you've now gotten yourself neck deep into, that seems to be vacillating between all the words we just said and all the risks that they want to constantly, you know, China risks or financing risks. So it's been really interesting to kind of watch. But what I can tell you, you know, being on the industry side, Mike, and working so closely with the hyperscalers, the infrastructure companies, you know, the capacity engines like Cumulus, is that people that are using and deploying AI And you can see this in spot pricing demand. You can see this in the energy build out. You can see this in the demand and revenues of companies like NVIDIA. There is no slowing down there. So despite the fact that, like I said, I think it makes for great TV to pull the panic switch at times, like realistically, I think it's still really early.

Michael Maniscalco: 

I completely agree. And the one thing I will say is going back to that speed and velocity comment earlier, our model has always been built around hyperspeed because the If I look at the pace that infrastructure has been developed historically, I look at the pace that compute and data center capacity has been brought online, that doesn't oftentimes line up with the speed that the market, the AI market specifically, is moving at. To us, this is all about hyperspeed. How do we bring compute to the demands of our customers who want They're literally, if you ask them when they want their compute, they would say yesterday. So how do we help line up reality with the fastest possible delivery timelines? And that's what it's all about to me. I see, as you said, a massive amount of demand and supply that's struggling to keep pace with that demand.

Daniel Newman: 

Yeah, totally. Turn up the volume to 11 is the old adage, right? How fast can we make this happen? But, you know, I think you would say that for like the past couple of years, right? So kind of 2023 was the chat GPT inflection. Now we're in 2026. Enterprises have been sort of in this continuum and now we're in this full head down sprint race to adopt AI. you guys are trying to go hyperspeed because you're really trying to keep up, in my opinion, with the customers and they're trying to go hyperspeed. You're out there, you're talking to customers, you're signing new deals, you're expanding your capacity. How, in your conversations, are you seeing the sort of, we're trying to do AI to we are doing AI at scale, enterprise-wide, at production level? Is that happening? And what is that creating in terms of a change for what organizations need from their infrastructure?

Michael Maniscalco: 

Yeah, I think I like to separate it out, because a lot of the customers we're providing infrastructure to a lot of the customers are out looking for for infrastructure AI infrastructure at scale are usually the ones providing the a combination of things. They're providing the models, so they're the model developers. They're providing the application layers or software interface layers to interface with those models and get value out of those models. That's a lot of the customers we're servicing today, or they have what I would define as some killer app use case, and the easiest one to point to that everybody, I think, can just associate with most simply is the code development applications. So those are the types of customers we're servicing today. That's where we're seeing people really pick up and scale. But I also think you're right. And a lot of the enterprises and even in the long tail of the SMB market are still trying to figure out what their killer app use cases are. And what I like about that is And the way I position this with our team, the way I position this when I talk to others who are trying to understand the space and the opportunity here is we're still in the early innings in this game. And we've got a long way to go. And all of that requires more and more infrastructure. And that's the service that we're providing at Cumulus AI.

Daniel Newman: 

You'll have a laugh. I was on CNBC closing bell and they asked me about that because I think someone had said like we're in the third inning or the fourth inning. And I said, I think we're tailgating in the parking lot. We're still, you know, we're cracking a beer and cooking a hot dog right now. I said, people feel like we're really far along. I think it's still super early. You know, we've only really had what I would call usable models that were capable of delivering enterprise outcomes for less than a year. It's probably been six months. Realistically, like maybe end of last year, Opus 4.6 kind of was like when people started to feel like, oh my gosh, I can do things like autonomously, like agentically in my enterprise.

Michael Maniscalco: 

So to think that we're that far along just feels like a stretch right yeah I my my line was we're in the parking lot waiting for the game to start that was probably a year ago and and now the question is, are you are we in the parking lot are we in the early innings I think understanding where you're on the curve is hard, but I think it's hard to deny that we have a long way to go and you're absolutely right I think. The early indications to me that got me reinvigorated about the infrastructure side of things before I came back into cumulus AI were just looking at some of those use cases and looking at those models, looking at the innovation that was happening around getting the most out of those models, reasoning mixed with experts, and then seeing the output, and seeing the output get really, really good really quickly. I've got a software development background, computer science background, so watching AI get better at writing code was pretty eye opening to me. And my rationale is, well, if that's happening, and it continues to happen, all these results are being driven by better models, more reasoning, which is more compute, which is more GPU infrastructure, more power shell, more power to deliver that intelligence. And that got me really excited about infrastructure.

Daniel Newman: 

Yeah, you sort of look at the curve, right? And it started out with like chatbots. And then, you know, you kind of went from chatbots to, you know, multi-turn, multi-modal, then you went to reasoning, and then you went to a gentic. By the way, just wait till we go to physical and robotic humanoid. And that actually creates another many orders of magnitude of compute demand. I don't think people have put that into their model yet, how big that is going to be. So I think the opportunities are huge. You sort of said this, maybe just have a chance to reiterate this a little bit, but you talk about the model game. Models are getting bigger, but there's also a lot of talk about sort of model economics. There was a period of time where it was like, I think the world thought one model was going to take over every industry. I feel like we've almost 180 now. And now we are talking about sort of open sourcing, open weighting everything. And then people panic when they say that. They're like, oh my gosh, what does that mean for compute and infrastructure? It seems like every time that story comes out, everyone panics and says we're overbuilding. But what I'm finding is that when we actually hear these things, it's the Jevons paradox thing. Models become more efficient, become cheaper. We need more infrastructure. Is that what you're seeing?

Michael Maniscalco: 

Yeah, I think so. And I As an infrastructure provider I think there's a lot of room for these other models, not every application and use case needs the most expensive most capable model with the deepest amount of reasoning. Not every use case requires the latest and greatest model and I think. Jensen puts this really well, is you have to think about the value of the token and the different values of the token depending on your goal. And you may need something that needs a very, very low value token. And that token value could be based on a variety of factors, quality of the models, latency, geographical location, security, compliance, timeliness, all these things matter. And then there on the other side of things, you may have a requirement for the most expensive token in the world. And you need that answer as fast as possible, no matter what. And we can talk about the real world use cases, but there are plenty. So that's how I view things. And I think the open source models today give you a a value-based token. And I think that's great, especially for the model providers, especially for those hosting the open source models. That just means more compute.

Daniel Newman: 

Yeah. I mean, there's a lot of debate to be had about where the open source models come from, securing those models. Are they truly open source versus open weight? You know, we're not truly showing the code, but this is a whole nother topic than infrastructure for another day. But it's really interesting that and by the way, I think I heard a great analogy. Maybe you'll appreciate this. I like F1. I heard like basically using a frontier for every workload is like having a Formula One race car used for delivering Amazon packages. It's like It's not, you don't need that for every case, right? And I think we're starting to see that shine through.

Michael Maniscalco: 

I use that with our own team as well, the same one. I mean, Lewis Hamilton can drive a Ferrari, right? An F1 Ferrari. You put the average Joe off the street in that car and it's not going to go well. And I think you can think of the state-of-the-art largest models a lot in the same way. Not every query needs that much horsepower behind it.

Daniel Newman: 

So, you know, we've kind of talked a lot about the model itself. We talked about the infrastructure deployment, you know, There's not a single deployment model for AI and the enterprise. I think we kind of have come full circle there was this period where it felt like everything would just be kind of open, or, you know, cloud, I guess you would say like you just go to the cloud and use a model. And now we're sort of seeing. on-prem deployments. We've got a lot of governance, compliance, sovereignty, all these different concerns. What do you see? How does hybrid infrastructure evolve? What role do reserve, compute, and distribute infrastructure play alongside traditional hyperscale cloud environments?

Michael Maniscalco: 

So I think about our differentiations and that's the view I take on it. And when I look at our customers and what they're asking for and what we're offering and why we're valuable to them, it's flexibility. And I think flexibility is really, really important for how quickly do you need that compute? What architecture do you need? How long do you need it? Do you need to do a heavy CapEx investment in that or do you want to operationalize it? you have that flexibility with the cloud providers, especially with the new cloud providers. So that's a big piece. How quickly can I get access to that compute is another one. Because going to buy small quantities of compute can be challenging as far as timelines go, even large quantities, just buying compute in general, depending on the supply chain can be challenging. So we're providing that access without a lot of that complexity. The next one is cost. because in certain environments, you're paying more for those tokens because of the supply demand situation right now. So we want to deliver that compute at our customers at the right cost for them to do the analysis of does this make sense to do on-prem and a hyperscaler environment or other. And then there's the trust factor, which becomes bigger and bigger, I think, part of the equation as enterprise thinker where they want to keep their data, how they want to keep their data and information safe because data is becoming the moat. And the last one, as I mentioned before, is speed. And just what can we do to bring that speed, that compute online for our customers as quickly as possible? And based on those factors, our customers will make decisions on where, when and how they want to place their compute infrastructure in the world.

Daniel Newman: 

By the way, very interesting that you point that out. You know, all I can think to myself is we're in a period, going back to the beginning of our conversation, where if you can build demand or deliver a chip or build memory or create a network switch or topology that works, you can almost sell it right now, because there's just so much demand, right? That probably won't be forever. At some point, capacity will catch up. And despite the fact that I do think there's an exponential curve for compute and for demand, we will build. That's what we do. That's what the world does. That's what economies do. You seem to be leaning in on hyperspeed, like hyperspeed versus hyperscale. Is that going to be sort of your mode? Is that sort of what you see as the cumulus moat long term is instead of maybe trying to bring gigawatts online, you're going to bring more you know, achievable, rapid compute online that can be accessed by enterprises and put into use in, I think, months rather than years. Is that kind of what you see as a long-term differentiator with Cumulus to some of the other players in this space?

Michael Maniscalco: 

I do. I believe so. In my career, in my history, I've seen what it takes to build large-scale infrastructure, how long that takes just from design to financing to construction to delivery. And then there's the risk of operations that you have to consider. And I think that that game is very important. That compute, that infrastructure is very important in the market. There are certain players in the market who have the expertise, have built the teams, have built the partnerships around the financing side of things and development side of things, that they can do that really well. But that's reserved for a few companies in the world at this point in time. And I think moving forward as training requirements get into the gigawatt scale, that remains true. So I saw a big gap, a lot of white space in the market. for bringing compute online in a different way. And that was, how do we think in smaller chunks of compute to match up with the pace that AI is moving, match up with the timelines that our customers want that compute online? And to me, that speaks to smaller pockets of compute, which are easier to come by. So smaller pockets of Power cell data center capacity easier to find in the world, smaller pockets of power land are easier to find in the US and abroad. And we believe we can run those projects in parallel to bring compute online faster for our customers. And that's the differentiated model that cumulus is pursuing. And then we also believe that there's some value in the distributed nature of that compute because as the world continues to change, somebody may have stricter requirements around federation and where they keep their data. And if we have a distributed footprint by matching up the demand to supply as quick as possible through these smaller pockets, fast forward, and now we have a distributed footprint, which has a lot of flexibility on how discerning that inference workload may become in the future.

Daniel Newman: 

Let's jump into the future. Now that you have had your big moment, Kumos is a public company. I think I read you signed another deal today, a pretty sizable deal. Congratulations on that. For everyone out there, I don't know what day you're watching this, beauty of on demand, but definitely check it out. Cumulus is signing deals consistently, and that's been one of the exciting things to have you in the market. But what do you see this looking like in, I don't know, three, five years? How does this industry… evolve? What should the CIOs, CEOs, board members, leaders, investors, what should we be watching as organizations are building for what looks like an AI-powered future?

Michael Maniscalco: 

I go back to that early innings comment. If we're in the early innings, I think that the demand for this capacity is still going to be high and we have to think about new ways to bring supply online and think about innovative ways to do that that aren't necessarily the big shiny object of another gigawatt campus. I think that can't reinforce the gigawatt campuses are important, but we think a portfolio of smaller pockets will go a long way. So I would be looking at partners that can keep up with that pace, if I'm honest, and keep up with that pace for the next few years, but also partners who are working with their customers to understand what they need in their future inference portfolio, if security and compliance are top concerns, or if data privacy and the locality of that data are big concerns, or if it's really heavily around the latest and greatest models or budget. I think all those factors are important. And I think certain partners will have a little more flexibility on where you can run those workloads to meet all those factors most effectively. And if you are looking at ways to maximize your spend on AI. I think having all those levers to pull is going to be important to the future.

Daniel Newman: 

Yeah, I think you're right. Now, just curious before I send everyone home, what do you consider that smaller chunk? Do you guys have sort of a guide? Is it 20, 50, 100? You know, what's the What are those smaller chunks that can be cut down to two months versus, you know, say years. And obviously the one red tape is all raw. I say politics, policy, regulation and red tape are the one thing we probably didn't talk much about that can. But I think part of doing smaller deals is also to get it right. It's a lot easier to get approval to do 20 or 50. than it is to get a gigawatt data center approved, right?

Michael Maniscalco: 

Absolutely. It's a lot less disruptive. It requires a lot less work on the power side to understand power availability. So I'll say that there are some selling points to the smaller chunks and that finding those pockets of power are a little easier navigating that. data center capacity is a little easier. So there are lots of selling points to the smaller capacity. We've been growing really nicely in the sub 10 megawatt range today. The challenge becomes as you scale, you want to get into a larger pocket. So if you look at most of our contracts today, there are a few megawatts and measurable data center capacity. And then we're looking at our roadmap into the future to determine what the right size is moving forward.

Daniel Newman: 

Mike, it's been a lot of fun chatting with you. Congratulations on the news this week. I know you rang the bell. That's always such a big moment. I'm really excited to continue to follow the journey. I've been tracking you guys for a couple of years. I've been tracking this space. It's been wild to watch all this demand. And I love seeing, you know, you guys taking a slightly different approach, having a, an idea of how you can grow this business and stand out from, you know, the, the, the large growth of sort of capacity shell power, uh, Neo cloud. There's a, there's a whole continuum of what these companies are. but you've got a story, you've got a differentiation, you're leaning into something. And I think it's looking like based on the 200 million plus of contracts you've recently signed since June, like it's working out. So congratulations. Love having you here on The Six Five, Mike, and let's do it again sometime soon. Thanks. Always enjoy talking to you. All right. And thank you everybody for tuning in to this episode of The Six Five virtual webcast. Check out everything Kumulos AI is doing and check out everything we have here on The Six Five. Appreciate you being part of our community. See you all later.

MORE VIDEOS

OpenAI's IPO Filing Meets a Hugging Face Breach as AMD Claims 75% Hyperscaler CPU Share

OpenAI files for a trillion-dollar IPO the same week a pre-release model breaches Hugging Face and 42 state attorneys general open a coordinated investigation. Patrick Moorhead and Daniel Newman also break down AMD's hyperscaler CPU numbers from Advancing AI, the Moonshot distillation accusations, a wave of coordinated AI governance moves in Washington, and a stacked earnings slate spanning TSMC, Alphabet, IBM, ServiceNow, and Intel.

Adobe's Vision for Redefining Creative Workflows in the AI Era

Enterprises are rethinking creative workflows as AI reshapes how content gets produced, governed, and measured. Elliot Sedegah, Director of B2B Product Marketing and Strategy at Adobe, joins Keith Kirkpatrick of Futurum to unpack what separates organizations capturing measurable ROI from generative AI from those still stuck in fragmented experimentation.

From Pilot to Production: How Lenovo XIQ Is Bringing Agentic AI to Retail at Scale

Ryan Shrout and Mitch Lewis of Signal65 walk through their independent validation of the Lenovo xIQ platform and Lenovo Super Agent for Retail with Paige Grady, Director of AI Solutions at Lenovo SSG. The conversation covers what stalls retail AI at the pilot stage, how full-stack integration addresses the operational challenges that custom builds cannot, and what Signal65 found across deployment time, agent accuracy improvement, and lifecycle management in testing.

See more

Other Categories

CYBERSECURITY

QUANTUM