Cerebras Unveils CS4 System With 30x GPU Speed Advantage as OpenAI Confirms Compute Demand "Vastly Outpaces" Capacity
Post-IPO product event details next-generation architecture, 2027 roadmap, and expanding customer roster including OpenAI, AMD, Arista and CrowdStrike
Cerebras Systems held its first major public event since going public earlier this year, a product launch dubbed "Supernova" in San Francisco on August 18, 2026. CEO Andrew Feldman used the venue to unveil the company's next-generation CS4 system and rack-scale architecture, while parading a lineup of partners and customers, including OpenAI, AMD, Arista Networks and CrowdStrike, that together sketch out a broader disaggregated-inference ecosystem forming around the company's wafer-scale chips.
CS4 architecture: the headline technical disclosure
The most consequential new information for investors is the specification of the CS4 system, which Cerebras says delivers 6x higher system-level performance than its current CS3 generation, 2x faster token generation, and up to 10x more tokens per watt. Chief Technology Officer Sean Lie said the system is built around a new chip called the WSE-3 Turbo and is the first Cerebras system to integrate three wafers into a single unit, yielding what the company describes as 43 petabytes per second of memory bandwidth per wafer. Lie claimed this is "2,000x more memory bandwidth than Rubin, Nvidia's next-generation GPU," a figure that, if accurate, underscores the structural memory-bandwidth advantage Cerebras believes it holds in the decode phase of inference.
The company is also introducing a new rack-scale architecture called Nexus, built around modular power supplies and "pluggable backpacks" that each house one wafer-scale engine. Cerebras says the modular design cuts component count by 50%, increases manufacturing automation by 60%, and reduces data center deployment times from days to hours, a claim aimed squarely at addressing the capital intensity and deployment friction that have dogged large-scale AI infrastructure buildouts. The CS4 is in early access now and is expected to reach general availability later this quarter.
Perhaps more important for long-term modeling purposes is the roadmap commitment: Cerebras is guiding to a doubling of raw performance every year, with a specific target of 20x higher throughput by 2027, alongside a next-generation wafer-scale engine slated for the CS5 system that year. This is a multi-year architecture commitment, not a one-off product refresh, and it gives investors a concrete cadence against which to measure execution.
OpenAI partnership scales up, and demand commentary is the real tell
Cerebras disclosed that OpenAI last week announced "GPT-5.6 Sol Ultrafast mode," a new service tier running OpenAI's most intelligent model at up to 14x standard speed, available exclusively on Cerebras hardware. Thibault, OpenAI's Head of Core Products and Platforms, appeared on stage and stated plainly that "we would love for Ultrafast to just be the default. I think it is a glimpse of what's to come," a comment Feldman flagged directly to the audience as noteworthy for the investment community.
More significant than the product tier itself was Thibault's commentary on capacity constraints across OpenAI's fleet. Asked about how the company thinks about scale, he said that "at some point, you reach a point where you run the model on the GPU and it's providing more value than the cost of running it on the GPU. And at that point, you want to have all the capacity in the world to just run it... this is a thing that we talk about all the time... it's like the demand fast outpaces the capacity that we have across our entire fleet. I don't think there is a reason to believe that will slow down or change." That is a direct, on-record data point on the demand side of the AI infrastructure buildout, coming from one of the largest buyers of compute in the industry, and it reinforces the bull case for hardware providers broadly, not just Cerebras.
Thibault also noted OpenAI now has 15 million weekly users running Codex and ChatGPT agents, with an updated, higher figure expected within the week, while cautioning that agentic adoption remains "a very small fraction" of ChatGPT's roughly 1 billion weekly users, suggesting significant headroom for inference volume growth industry-wide.
Disaggregated inference: the AMD partnership formalizes a new architecture pattern
Cerebras used the event to detail a partnership with AMD built around what it calls disaggregated inference, splitting the prefill stage of inference (processing user input, which is parallelizable and memory-bandwidth-light) onto AMD's GPUs, specifically the upcoming Helios rack, while routing the decode stage (sequential token generation, which is memory-bandwidth-intensive) to Cerebras wafers. Feldman said this combination can deliver 10x faster performance than GPUs alone and 5x more throughput than Cerebras alone. AMD CTO Mark Papermaster called it "an innovation driven out of necessity," reflecting an industry-wide shift in infrastructure spend from training-dominated buildouts toward inference-optimized architecture. The CS4's I/O module was explicitly designed as a "programmable universal disaggregation interface" using RoCE-based networking, positioning Cerebras as hardware-agnostic rather than trying to displace GPUs outright, a notable strategic signal given the size of Nvidia's installed base.
Data center buildout and the networking dependency
Feldman disclosed that Cerebras has brought online or under contract 600 megawatts of data center capacity for delivery by the end of next year, spanning Santa Clara, Toronto, Dallas, Minneapolis, Montreal, Oklahoma City, Alabama, Lyon, and sites in Norway and Finland. He was candid that this "is not nearly enough," acknowledging the capacity race remains a persistent constraint rather than a solved problem.
Arista Networks CEO Jayshree Ullal appeared to detail the networking layer underpinning these clusters, noting that Arista has built differentiated traffic classes for AI workloads distinct from traditional cloud networking, and flagging that enterprise adoption of agentic AI, as distinct from hyperscaler and neo-cloud deployment, "has hardly come in yet," implying a large uncaptured demand pool ahead. She also emphasized network resiliency as the next major bottleneck, stating plainly that "the biggest challenge going forward will be how you retain the availability of your compute" when hardware fails or cables are pulled, a reminder that scaling reliability, not just raw speed, will be a differentiator.
Customer validation: Figma, Cognition, Armis and CrowdStrike quantify the speed thesis
Beyond the infrastructure partnerships, Cerebras brought forward real deployment data. Cognition, maker of the Devin coding agent, disclosed that its in-house SWE-grep model runs on Cerebras at up to 2,800 tokens per second, more than 10x faster than comparable large language models, and that its newer SWE-1.7 model matches GPT-5.5 and Opus 4.8 in quality while running at close to 1,000 tokens per second at meaningfully lower cost than comparable models like Kimi K2.7 Code or GLM 5.2. Figma detailed a newly shipped, Cerebras-powered design agent embedded directly in its multiplayer canvas, arguing that fast inference is what makes iterative, non-verifiable creative work like design tractable for agents in a way slower inference cannot support.
Cerebras also cited Armis, a code-security scanning customer, completing scans in roughly one-third the time of benchmark frontier models while finding more vulnerabilities at lower cost, and CrowdStrike, whose VP of Engineering Keith Culley framed the security case for speed starkly: with faster inference, security tools can perform "5x to 10x more inspection inside the same acceptable window of time," which lowers friction and increases adoption of security controls broadly. He pointed to a industry benchmark of 27 seconds as the fastest recorded network breakout time by an adversary, arguing that decision windows in cybersecurity are now measured in single-digit seconds, a use case where inference latency has a direct and quantifiable dollar impact rather than just a UX benefit.
Taken together, the event functions less as a single product reveal and more as a systems-level thesis defense: Cerebras is betting that memory-bandwidth-bound decode workloads remain structurally advantaged on wafer-scale silicon, that disaggregated architectures with GPU partners (not GPU replacement) is the commercially viable path, and that a widening set of paying customers across coding, design, and security verticals validates the willingness to pay for speed as a distinct product attribute rather than a benchmark curiosity.