Broadcom Transcript: Custom XPU Silicon Revenue to Exceed $100 Billion as Tech Giants Shift from Nvidia to Tailored Modular Architectures
August 19, 2026 - RAISE Summit 2026
Introduction and the High Cost of General AI Compute
Host: All right, great. I think this is, is this the last one, or probably? Oh, everyone is tired. Okay, we will try to make it more entertaining. Well, great to see you here. You know, we are in France, in Paris, the land of haute couture, and I think that Broadcom in a way is the haute couture of chips.
Charlie Kawwas: Thank you. I think so.
Host: I think so. Obviously, with AI and the rise of the XPU, everyone wants to build XPUs, and they all come to the house of haute couture at Broadcom. So, why don't we discuss why the XPU has become such a pivotal thing for basically every hyperscaler and every foundation lab in the world?
Charlie Kawwas: First of all, thank you for having me. We are going to have some show and tell today. Yes, thank you for having me here. This is actually a great introduction about the haute couture. I never thought of it, but in a way, it is a phenomenal idea. Look, at the end of the day, this AI wave was really started and driven by what we call general AI compute, driven by Nvidia. Nvidia has done a great job to build these big racks of custom and proprietary platforms, which is a great platform. But the people who are spending more than 80% of the money are the frontier labs. They are the ones who are consuming all of this hardware, and then they build the software that sits on top. What they have realized is that this general compute comes with a tax.
Charlie Kawwas: The tax is first the efficiency of that compute platform. Meaning, it is built for everybody. If you have a specific workload in your AI frontier model, guess what? You are paying that extra tax. The other tax is the high margins that come with it. Given the big spend that they are doing, when it was hundreds of millions of dollars to a few billion dollars, it made a lot of sense to use a general-purpose platform that is pre-built for you in a rack. But when you start spending $50 billion, $100 billion, or maybe $150 billion—as you know, the top four in the US are each spending $150 billion to $200 billion—that tax for each player is about $100 billion to $150 billion that they could have gotten in tokens. As a result, they realized that if there is a way to build their own platform fairly fast, at the same rapid pace that Nvidia does it, which is once a year—which is quite difficult—they would love to do that.
Charlie Kawwas: Our journey started 13 years ago, starting with the TPUs from Google. Their initial idea was that they did search using text, and they wondered if they could actually do search using voice. To do that required the creation of the TPU. But as AI developed, we realized that these chips are getting bigger and bigger.
Show and Tell: The Evolution of Custom Silicon from 2025 to 2028
Charlie Kawwas: Charlie reaches down and picks up several physical semiconductor packages, placing them on the table to demonstrate the progression of Broadcom's custom silicon design. I am actually going to show you these for a reason, and I am going to pass them to you. But you have to promise to give them back to me because they are very expensive. This is, for example, an XPU that was built and shipped last year in 2025. As you can tell, it is hard to see, but it is made of different chiplets. Some are memory-based, some are networking-based, and some are compute.
Charlie Kawwas: Charlie picks up a second, visibly larger chip and hands both to the host. What we are building and shipping this year in 2026 is this chip. As you can see, it got bigger. So, I am going to ask you to hold one. This is last year's, and this is this year's. You can see from last year to this year, you are now getting three times more memory. In the 2025 version, you have two cubes of memory. In the 2026 version, you have six cubes of memory, and the compute that is in the middle is almost three times stronger.
Charlie Kawwas: Charlie picks up a third, significantly larger modular board representing the 2027 design. What ended up happening is we said for next year, we actually have to build this. So now, you have even more memory.
Host: How many more times memory is this?
Charlie Kawwas: This is one, three, nine. In two years, it is almost an order of magnitude. When you look at the compute, in the earlier version you had one large compute block, and here you have four of them. Then you go to the new technology in semiconductors that allows you to shrink the transistors, and then you add a lot more. I am going to give you this one as well.
Charlie Kawwas: Charlie picks up a massive, highly complex multi-layered system board representing the 2028 design. And then, this is what we are working on for 2028. This is 2027, and this 2028 version is literally going to be the size of the platform. If you put them next to each other, you can see the scaling from 2025, 2026, 2027, to 2028.
Charlie Kawwas: The reason I want you to see that is I want you to understand two things. One is the speed at which this technology is moving, and hence the innovation that it needs to go through has to be extremely rapid. Two, let me open this up for you. Literally, that is how we build these chips. These are actually modular. This is why people come to the haute couture, as you said, because we actually take these platforms and pre-build them. These little chiplets that go in the chip, we pre-build them, make sure they are production-ready, work with partners like Samsung and others, and create these memory blocks. This memory integration is one of the reasons why the prices are going higher, because look, we are putting so much more in them. We create these input-output Ethernet-based networking capabilities, and then we work with our partners. Remember, we used to have half of one of these. Now we have one, two, three, four, five, six, seven, eight, times two. This is a double-decker. We are actually stacking these chips back-to-back, or actually the correct term is face-to-face. So you will have 16 times what we had two years ago, 16 times in a single chip. And of the first chip, we deployed a million units.
Host: This does not even look like a chip, Charlie. This looks like a monstrosity of a system. It is a beast.
Charlie Kawwas: Yes, it is. Actually, that is what I call it at work. This is the beast. We are building four of these for the largest four Large Language Model players. It is modular. The beauty of this, to your point, is that it is haute couture. You come to us, you tell us what kind of tuxedo or dress you want, and we will customize it and build it for you, removing that general compute tax. This will be purpose-built for, let us say, your inference workload or your decode workload. This is what makes this fast and great. As long as Nvidia continues to develop great technology every year to match that cycle, there is only one other place to go to.
The Rapid Cadence of Nvidia and the Scale of the Big Five
Host: You mentioned the one-year cadence that Nvidia is on. It used to be two or three years, right? And then they accelerated it to one year, every year. So the pressure is on. For most companies, they cannot keep up with that cadence. The only one that can is maybe you, with this modular Lego architecture.
Charlie Kawwas: You are exactly right. If you are one of the big five in the US, which today are spending maybe 90% of the capital expenditure in the world—which are OpenAI, Anthropic, Google with Gemini, Meta, and xAI—that is it. Amongst these five is literally 80% to 90% plus of the market. Four out of the five have realized that while the Nvidia technology is so good, they cannot keep buying this general compute platform that is custom only at the rack level. So they come to us, each one of them, and they say, "What can we do together to get a purpose-built platform for me that is tailored to my workload, and that is not at an 80% margin?" That is what we do for a living, because we focus just on building this. We make sure each of these Lego blocks is production-ready before you come and talk to me. When you come and talk to me, it becomes an integration exercise. What we announced with Jalapeno, for example, is the fastest chip that was ever built, completed in nine months. It is actually something similar to this. They did that in less than a year.
Host: So you have basically weaponized or enabled these foundation labs to build their own silicon in a year. Five years ago, your AI chip business was zero, and now it is going to be over $100 billion next year. Zero to over $100 billion in five years.
Charlie Kawwas: More than $100 billion. Remember, $100 billion is not good enough. It is more than $100 billion.
Host: More than $100 billion in five years. Yet, how many of these are really scaling? Of those four, in the 2027 period, it is still in early formation, right?
Charlie Kawwas: It is, but just to give you an order of magnitude view, think of the 2025 chip as being about 1 million units a year. Think of the 2026 chip as being about 2 million units a year. But remember, the 2026 chip is three times the capacity of the 2025 chip, so that is a six-fold increase in compute capacity in one year. Then you go to the 2027 chip, which is probably 4 million to 5 million units, and this is nine times the capacity of the 2025 chip. Now you are at 40 to 50 times the scale of the first one. And then the beast in 2028 is three times that again. When you think of the scale at which people are building these things, it is a race toward superintelligence and Artificial General Intelligence. These four to five frontier labs believe that the first company in the world that would enable this platform is what I call the agentic AI platform. The first company that would enable a fully integrated GPU, XPU, and CPU in a single platform potentially can be the first to reach AGI. With this 2028 technology, I think we are going to start seeing a lot of capable agents, experts in certain domains, that would finally have the technology to enable this. This is the exciting piece about what we are doing, and why people keep coming back and saying, "I would love to use the Nvidia chip, but I have a general compute tax and a margin tax." If they want to drop their token cost per dollar, they can get two to three times the compute for the same dollar, which means the token price comes down massively. That is ultimately what they are driving for.
Host: So you are effectively a haute couture design house with low taxes.
Charlie Kawwas: Correct.
Co-Design at the Rack Level and Open Standards
Host: The other topic du jour in the AI world is co-design. You brought up Jalapeno and OpenAI, and how the tight integration of model development helps define the chip architecture, and then the chip performance is optimized to the model architecture. This tight integration of co-development seems like it is only possible if you have your own XPU.
Charlie Kawwas: That is absolutely right. There are two portions to that haute couture approach. One is exactly the deep co-design for the XPU that you described. The second is at the rack level. Not only do we have to do it at the XPU level for their Large Language Models and their specific use cases, but we also engage heavily in making sure that the entire rack is open, and co-designed for them. Part of that is where we bring our networking and Ethernet capabilities, making sure it is open. We started off by making the front-end networks Ethernet-based and open. Then we went to scale-out, and now everybody is open, including Nvidia. The next battle is scale-up, where a lot of these guys use these blue chips from us, these little chiplets we call them, which are all Ethernet-based. When we co-design all of this, it is co-design at the rack level, starting from the XPU with each of these blocks, then the entire rack, which gives them a choice. Honestly, if Broadcom does not give you the best technology, they have a choice to go with somebody else. But this has been multigenerational and quite successful for us because, as you said, we started from less than $1 billion two years ago in this space, and next year, in three years, we are going to be over $100 billion just in this space.
Host: You mentioned the openness of the interconnect standards, and you guys are full-stack. You do networking, packaging, custom XPU design, and IP blocks. They want that open standard. What about open-source? Is there an idea of an open-source foundation lab? Do you see a market for XPU development for open-source foundation labs or open-source models, rather than just the closed foundation labs?
Charlie Kawwas: Yes, actually, you are absolutely right. One good example that we discussed earlier today is SambaNova. How many of you have heard of a company called SambaNova? Just raise your hand. That is good, a bunch of people here, maybe 15% of you. SambaNova Systems has actually been co-developed through deep engineering with Rodrigo Liang and his team. We have done three generations of this. What Rodrigo does is he uses his XPU to run a closed frontier model, which is state-of-the-art, but he also brings an open-source model like Mixtral or DeepSeek and runs it on the same chip. Hence, he enables a hybrid solution. If you need to use Claude or ChatGPT for specific use cases, you can run that on that XPU, but if you suddenly need to use something that does not require state-of-the-art closed models and you want to use an open-source stack, you use it and you pay zero for those tokens. That model already exists today, and we have proven it with a few startups. I think we will start seeing that more on XPUs as we move forward.
Future Outlook: Power Demands, Edge AI, and Broadcom's 16 Franchises
Host: In the last couple of minutes here, how do you see the next couple of years and the path forward? You are showing the 2028 chip, and showing something that looks massive. What do you anticipate are going to be the main topics for the next couple of years and beyond?
Charlie Kawwas: That is the $2 trillion question.
Host: There is more than just Jalapeno, right?
Charlie Kawwas: Yes, there is more than Jalapeno. There are definitely more spicy chips and a lot more XPUs that will be coming out, both from the large labs and a lot of startups. But what I see is an insatiable demand in gigawatts. For example, we have announced for one of the top five customers that this year we are putting in about 1.5 gigawatts with one of the chips I showed you. Next year, we are going to put in over 5 gigawatts, which we have already contracted with that same customer. So we are going to go from 1.5 gigawatts to 6.5 gigawatts total. The year after that, I am pretty sure it will be over 10 gigawatts incremental. If you look at the next three years for each of these lab players, they are going to go from 1 to 2 gigawatts this year, to 5 gigawatts, to over 10 gigawatts incremental. Over that three-year journey, they will go from 2 to 5 to 10 gigawatts, which is 17 gigawatts total, starting from just 1 gigawatt today.
Charlie Kawwas: This means there is a massive challenge in terms of the amount of power that is needed, the amount of silicon wafers that are needed, and the amount of memory that is needed. It is actually going to pose a challenge for the rest of the non-AI businesses, which we also do at Broadcom. Remember, I have 16 franchises. Five of them are in AI, and 11 are in non-AI. Part of the challenge that I think we all have to live through as we experience the exciting life of AI scaling up into AGI is making sure that the rest of the businesses survive. AI is not just going to live in data centers; AI has to get to the edge, and these other businesses have to actually enable this edge AI. We have a great, exciting journey on the AI track with power and all of these things. We are even building our own factories in Singapore to build these big beasts and big chips because existing technologies cannot enable them. But as we do this, the challenge for all of us is that AI has to migrate, just like the internet started from data centers and moved all the way to smart devices today. The same thing has to happen with AI. I think for us at Broadcom, it will be exciting because we will play in both pieces as the haute couture house of silicon.
Host: Thanks, Charlie.
Charlie Kawwas: Thank you. Thank you.