Cerebras Systems Stock price
📊 Peer Group
📈 What is it?
The peer group consists of the companies with the most similar business model. They serve as a benchmark for putting a stock into context.
🧮 How is it selected?
Based on similarity of business model, meaning companies from the same industry with comparable products and a similar customer base. That's the only way to compare apples to apples.
🏛️ Why does it matter?
Whether a stock is cheap or expensive is best judged by comparison. A P/E of 18 or an EV/FCF of 20 can look cheap or expensive depending on the yardstick. The peer group gives you the most accurate one: companies with a similar business model that operate under the same conditions.
🎯 What does it mean for investors?
When a metric sits below the peer average, the stock is valued more cheaply relative to its competitors, and above the average more expensively. A discount to the peer group can be an opportunity, but it can also have a reason (for example lower growth). The comparison is a starting point, not a verdict.
Is Cerebras Systems a Top Scorer Stock based on the Dividend, High-Growth-Investing or Leverman Strategy?
As a Free StocksGuide user, you can view scores for all 9,127 stocks worldwide.
StocksGuide Premium
StocksGuide Unlimited
Key metrics
📘 Market Capitalization
📈 What is it?
Market capitalization shows how much a company is currently worth on the stock market.
🧮 How is it calculated?
🏛️ Why is it important?
It helps classify companies by size (Large, Mid, Small Cap) and indicates their market presence and relative stability.
🧮 Calculation
🎯 What does this mean for investors?
- Large-cap companies tend to be more stable, often pay dividends, but may grow more slowly.
- Smaller firms may offer higher growth potential but come with more volatility.
- Market capitalization is a useful indicator of company size — but not a measure of whether a stock is undervalued or overvalued.
📘 Enterprise Value (EV)
📈 What is it?
Enterprise Value represents the total cost to acquire a company — including its debt and excluding its cash reserves.
🧮 How is it calculated?
(= Market Cap + Net Debt)
🏛️ Why is it important?
EV gives a more complete picture of a company's value than market cap alone and is used in key valuation ratios like EV/FCF or EV/Sales.
🧮 Calculation
🎯 What does this mean for investors?
- Enterprise Value shows the true cost of buying a company, including all financial obligations.
- It is more accurate than just looking at market cap, especially when comparing companies with different levels of debt or cash.
- Professional investors prefer EV-based multiples because they better reflect the company’s full financial footprint.
📘 Net Debt
📈 What is it?
Net Debt shows how much debt remains after subtracting a company’s available cash reserves.
🧮 How is it calculated?
🏛️ Why is it important?
It indicates how dependent a company is on borrowed money and how easily it can service its debt in the short term.
🧮 Calculation
🎯 What does this mean for investors?
- Low or negative net debt signals financial strength and flexibility.
- Companies with strong cash positions are better positioned in crises.
- High net debt increases financial risk — especially in environments with rising interest rates or economic downturns.
📘 Cash
📈 What is it?
Cash represents all liquid assets a company can access immediately — including cash, bank deposits, and short-term investments.
🧮 How is it calculated?
🏛️ Why is it important?
It reflects a company’s financial flexibility and resilience — enabling investments, buybacks, or buffer in downturns.
🧮 Calculation
🎯 What does this mean for investors?
- A strong cash position means greater room for maneuver and crisis resistance.
- Cash-rich companies can invest, pay down debt, or repurchase shares.
- But excess idle cash might indicate a lack of growth opportunities.
📘 Shares Outstanding
📈 What is it?
Shares outstanding represent the total number of a company’s shares currently held by investors — excluding treasury stock.
🧮 How is it calculated?
🏛️ Why is it important?
It’s the basis for key metrics like Earnings Per Share (EPS), Market Capitalization, or the Price/Earnings ratio (P/E).
🧮 Calculation
🎯 What does this mean for investors?
- Fewer shares in circulation typically increase earnings per share — making each share more valuable.
- Share buybacks reduce the number of shares and boost per-share metrics.
- Issuing new shares does the opposite — diluting shareholder value and lowering per-share figures.
📘 Price-to-Earnings Ratio (P/E)
📈 What is it?
The P/E ratio shows how many times a company's earnings per share are reflected in its current share price — in other words, how "expensive" the stock appears relative to its profits.
🧮 How is it calculated?
🏛️ Why is it important?
The P/E ratio is one of the most widely used valuation metrics. It helps investors assess whether a stock appears cheap or expensive compared to its earnings power.
🎯 What does this mean for investors?
- A low P/E may indicate undervaluation — or signal underlying issues.
- A high P/E may reflect strong growth expectations — or an overvalued stock.
📘 Price-to-Sales Ratio (P/S)
📈 What is it?
The P/S ratio shows how much investors are paying for $1 of the company’s revenue – regardless of profitability.
🧮 How is it calculated?
🏛️ Why is it important?
P/S is especially useful for evaluating growth companies or businesses not yet profitable. It reflects how the market values the company’s sales.
🧮 Calculation
Market Cap = $46.28b | Estimated Revenue = $904.26m
🎯 What does this mean for investors?
- A low P/S may indicate undervaluation — or low profitability.
- A high P/S can reflect strong growth expectations — or excessive optimism.
- Especially helpful when evaluating companies where profits are low, volatile, or negative.
📘 Enterprise Value to Sales (EV/Sales)
📈 What is it?
EV/Sales shows how much investors are paying for $1 of revenue — considering not just equity, but also debt and cash. It’s the capital structure–adjusted version of the P/S ratio.
🧮 How is it calculated?
🏛️ Why is it important?
It’s ideal for comparing companies with different levels of debt. It reflects a company's true cost relative to its revenue.
🧮 Calculation
Enterprise Value = $39.28b | Forward Revenue = $904.26m
🎯 What does this mean for investors?
- EV/Sales allows for capital structure–neutral company comparisons.
- A lower ratio may indicate undervaluation; a higher one may signal strong growth expectations or overvaluation.
- Especially helpful when evaluating high-growth companies with low or negative earnings.
📘 Enterprise Value to Free Cash Flow (EV/FCF)
📈 What is it?
EV/FCF shows how many years it would take for a company to "pay back" its enterprise value using its free cash flow.
🧮 How is it calculated?
🏛️ Why is it important?
It focuses on real cash generation, ignoring accounting noise — ideal for assessing profitability and value based on liquidity, not earnings.
🧮 Calculation
🎯 What does this mean for investors?
- A low EV/FCF may signal undervaluation and strong cash generation.
- A high EV/FCF might reflect weak recent cash flow or aggressive growth expectations.
- Best suited for stable, mature businesses with predictable free cash flows.
📘 Price-to-Book Ratio (P/B)
📈 What is it?
The P/B ratio compares a company’s market value to its book value — showing how much investors are paying for each dollar of net assets.
🧮 How is it calculated?
🏛️ Why is it important?
P/B is commonly used for asset-heavy industries like banks or industrials. It helps assess whether a stock is trading above or below its net asset value.
🧮 Calculation
🎯 What does this mean for investors?
- A P/B below 1 may signal undervaluation — or weak profitability.
- A P/B above 1 implies the market expects future value creation (e.g., brand, IP, growth).
- Best used for companies with tangible assets and strong balance sheets.
📘 Equity Ratio
📈 What is it?
The equity ratio indicates what portion of a company’s total assets is financed by shareholders’ equity – in other words, how much it relies on its own capital.
🧮 How is it calculated?
🏛️ Why is it important?
A high equity ratio reflects financial strength and stability, especially during downturns. It’s a key indicator of a company’s solvency and long-term risk profile.
🧮 Calculation
🎯 What does this mean for investors?
- Companies with high equity ratios are generally more resilient and less dependent on external debt.
- Low equity ratios can signal higher risk or aggressive financial strategies.
- Important: Always assess the equity ratio in combination with the return on equity (ROE). This shows not just how stable the company is – but also how efficiently it uses shareholder capital.
📘 Return on Equity (ROE)
📈 What is it?
Return on equity (ROE) shows how efficiently a company uses its shareholders’ equity to generate profit. In other words: how much net income is earned per dollar of equity.
🧮 How is it calculated?
🏛️ Why is it important?
ROE is a core profitability metric. It helps investors understand whether a company delivers attractive returns on the capital provided by its shareholders.
🧮 Calculation
🎯 What does this mean for investors?
- A high ROE indicates that the company is using its capital efficiently and profitably.
- It’s especially meaningful for capital-intensive businesses or firms with high equity bases.
- Important: A very high ROE can also result from high debt levels – always interpret it alongside the equity ratio to assess financial health.
📘 Return on Capital Employed (ROCE)
📈 What is it?
ROCE measures how efficiently a company generates profits from its total capital – including both equity and interest-bearing debt.
🧮 How is it calculated?
It evaluates the return on all capital employed, regardless of how it’s financed.
🏛️ Why is it important?
ROCE is ideal for comparing companies with different financing structures. It shows how well management uses capital to create value for both shareholders and creditors.
🎯 What does this mean for investors?
- A high ROCE means the company uses its capital efficiently – regardless of whether it's funded by debt or equity.
- The higher the ROCE compared to peers, the more value the company creates with its invested capital.
- Especially relevant for capital-intensive sectors like industrials, energy, or infrastructure.
📘 Return on Invested Capital (ROIC)
📈 What is it?
ROIC measures how efficiently a company generates returns from the capital invested in its core operations – regardless of whether the capital comes from equity or debt.
🧮 How is it calculated?
- NOPAT = Net Operating Profit After Taxes
- Invested Capital = Operating assets minus non-interest-bearing liabilities
🏛️ Why is it important?
ROIC is one of the most accurate indicators of capital efficiency. Unlike return on equity, it is not distorted by leverage and shows how much value is created for all capital providers.
🎯 What does this mean for investors?
- A high ROIC shows how effectively a company uses the capital that is truly invested in its core operations.
- Unlike ROCE, ROIC focuses only on the capital that is actively used to run the business – and that requires a return (i.e. interest-bearing).
- Especially useful when comparing companies with large amounts of excess cash or non-interest-bearing liabilities – giving a more realistic picture of capital efficiency.
📘 Leverage Ratio (Debt-to-Equity)
📈 What is it?
The leverage ratio indicates how much a company relies on interest-bearing debt (such as loans and bonds) relative to its shareholders’ equity.
🧮 How is it calculated?
🏛️ Why is it important?
This ratio helps assess a company’s financial structure and risk profile. High leverage can enhance returns – but also increases exposure to interest rate changes and financial stress.
🧮 Calculation
🎯 What does this mean for investors?
- A low leverage ratio signals financial strength and independence.
- A higher ratio can improve returns in good times but increases risk during downturns or rising interest rate periods.
- 👉 Always interpret in the context of industry, capital intensity, and interest rate environment.
📘 Revenue
📈 What is it?
Revenue shows how much a company earns in total from selling its products and services – the gross income before any costs are deducted.
🧮 How is it calculated?
🏛️ Why is it important?
Revenue is one of the key figures to assess a company’s size, market position, and growth potential.
🧮 Calculation
🎯 What does this mean for investors?
- Growing revenue indicates rising demand and can be an early signal of future earnings growth.
- Comparing actual and expected revenue reveals trends in the market environment and analyst sentiment.
- Note: Strong revenue alone isn’t enough – margins and profitability matter just as much.
📘 EBITDA
📈 What is it?
EBITDA stands for “Earnings Before Interest, Taxes, Depreciation, and Amortization.” It reflects a company’s operating profit before the effects of financing, taxes, and accounting depreciation.
🧮 How is it calculated?
🏛️ Why is it important?
EBITDA is widely used to evaluate a company’s operating performance – especially across capital-intensive sectors or international comparisons.
🎯 What does this mean for investors?
- A high or growing EBITDA indicates strong operational profitability – independent of taxes, interest, or accounting methods.
- It’s especially useful for comparing companies across sectors or geographies.
- Important: EBITDA is not a net income figure – it excludes key costs like depreciation and interest.
📘 EBIT
📈 What is it?
EBIT stands for “Earnings Before Interest and Taxes.” It reflects a company’s operating profit after depreciation, but before interest and tax expenses.
🧮 How is it calculated?
🏛️ Why is it important?
EBIT is a core profitability metric that shows how well the company performs in its main business operations – independent of capital structure and tax environment.
🎯 What does this mean for investors?
- A high EBIT indicates strong profitability from the company’s core business – before financial and tax effects.
- It allows better comparison between companies with different debt levels or tax structures.
- Compared to EBITDA, EBIT already accounts for depreciation and reflects capital intensity more clearly.
📘 Net Income
📈 What is it?
Net income is the company’s total profit – the amount left after all expenses, taxes, interest, and depreciation have been deducted.
🧮 How is it calculated?
🏛️ Why is it important?
Net income is the most comprehensive measure of a company’s profitability – showing how much actual profit remains after all business and financing costs.
🎯 What does this mean for investors?
- Growing net income indicates that the company is managing all of its costs efficiently.
- It directly influences valuation metrics like P/E ratio and the company’s dividend capacity.
- Over time, net income trends reveal how resilient and profitable the business model really is.
📘 Free Cash Flow (FCF)
📈 What is it?
Free Cash Flow shows how much actual cash remains after a company covers its operating expenses and capital expenditures.
🧮 How is it calculated?
🏛️ Why is it important?
FCF reflects a company’s real financial strength – regardless of accounting profits. It shows how much flexibility a company has for dividends, share buybacks, or debt reduction.
🧮 Calculation
🎯 What does this mean for investors?
- High free cash flow means the company generates real, usable cash – independent of reported net income.
- It’s often the most reliable base for sustainable dividends and buybacks.
- Declining FCF can be an early warning sign – even when profits appear stable.
📘 Revenue Growth
📈 What is it?
Revenue growth shows how much a company’s sales have changed compared to the previous year – both on a trailing basis (TTM) and based on forward projections.
🧮 How is it calculated?
Forward = (Expected revenue ÷ Revenue in prior year − 1) × 100
Forward growth is based on analyst estimates for the current fiscal year.
🏛️ Why is it important?
Rising revenue signals growing demand, business expansion, and market share gains – especially important for growth-oriented companies.
🧮 Calculation
🎯 What does this mean for investors?
- Growth is the engine of long-term value creation – especially in tech and growth sectors.
- What matters is not just current growth, but its sustainability.
- Forward projections reflect whether analysts expect continued momentum – or a slowdown.
📘 EBITDA Growth
📈 What is it?
EBITDA growth shows how much a company’s operating profit (before interest, taxes, depreciation, and amortization) has increased or decreased compared to the previous year.
🧮 How is it calculated?
Forward = (Expected EBITDA ÷ EBITDA from prior year − 1) × 100
The forward estimate is based on analyst projections for the current fiscal year.
🏛️ Why is it important?
Growing EBITDA indicates improving operational profitability – regardless of financing or accounting effects.
🎯 What does this mean for investors?
- Strong EBITDA growth signals operational efficiency and scalability – especially during growth phases.
- EBITDA growth can be an early indicator of margin and earnings expansion – but should be assessed alongside revenue and EBIT.
📘 EBIT Growth
📈 What is it?
EBIT growth shows how much a company’s operating profit (after depreciation, but before interest and taxes) has increased compared to the previous year.
🧮 How is it calculated?
Forward = (Expected EBIT ÷ EBIT from prior year − 1) × 100
The forward estimate is based on analyst projections for the current fiscal year.
🏛️ Why is it important?
EBIT growth is a direct indicator of a company’s business performance – taking into account capital intensity through depreciation.
🎯 What does this mean for investors?
- Rising EBIT signals improving operating profitability – even after accounting for depreciation.
- It’s especially important for evaluating companies with significant capital expenditures.
- Combined with revenue and EBITDA growth, EBIT growth provides a well-rounded view of operational progress.
📘 Net Income Growth
📈 What is it?
Net income growth shows how much a company’s bottom-line profit has increased or decreased compared to the previous year – both on a trailing basis (TTM) and based on analyst projections.
🧮 How is it calculated?
Forward = (Expected net income ÷ Net income from prior year − 1) × 100
The forward estimate reflects analysts’ expectations for the current fiscal year.
🏛️ Why is it important?
Net income is the ultimate measure of profitability. Growing net income signals stronger efficiency, cost control, and sustainable earnings power.
🎯 What does this mean for investors?
- Stronger net income boosts valuation, dividend potential, and investor confidence.
- If profits stall while revenue grows, it may signal margin pressure.
📘 Free Cash Flow Growth
📈 What is it?
Free cash flow (FCF) growth shows how a company’s available cash – after covering operating expenses and capital expenditures – has changed compared to the previous year.
🧮 How is it calculated?
🏛️ Why is it important?
Free cash flow reflects real financial strength. Growing FCF indicates more flexibility for dividends, share buybacks, and reinvestment.
🎯 What does this mean for investors?
- Declining FCF may point to rising investments, increasing costs, or weaker operating performance.
- Especially for dividend investors, FCF growth is critical – since dividends are paid from actual available cash.
- A negative trend isn't always bad, but it deserves closer attention.
📘 Gross Margin
📈 What is it?
Gross margin shows how much of a company’s revenue remains after deducting the direct costs of goods sold (like materials and production). It represents the company’s “raw profit” before fixed costs, taxes, and interest.
🧮 How is it calculated?
Or simply: Gross Margin = Gross Profit ÷ Revenue × 100
🏛️ Why is it important?
Gross margin indicates how efficiently a company can produce or procure what it sells. It is a key measure of product-level profitability and pricing power.
🎯 What does this mean for investors?
- A high gross margin suggests strong pricing power and efficient production.
- Falling margins may signal rising input costs or competitive pressure.
- Compared to peers, gross margin offers insights into the quality of a business model.
📘 EBITDA Margin
📈 What is it?
The EBITDA margin shows how much of a company’s revenue remains as operating profit before interest, taxes, depreciation, and amortization.It reflects operating efficiency without being distorted by financing or accounting factors.
🧮 How is it calculated?
🏛️ Why is it important?
The EBITDA margin reveals how much operating income a company generates per dollar of revenue – independent of capital structure and tax effects.
🎯 What does this mean for investors?
- A high EBITDA margin reflects strong core profitability – before accounting distortions.
- It allows for effective comparisons across companies and sectors.
- A stable or growing margin signals efficient cost control and business scalability.
📘 EBIT Margin
📈 What is it?
The EBIT margin shows what percentage of revenue remains as operating profit after depreciation but before interest and taxes.
🧮 How is it calculated?
🏛️ Why is it important?
The EBIT margin reflects a company’s core profitability while accounting for capital intensity (e.g. machinery, infrastructure). It’s especially useful for comparing businesses with different levels of depreciation.
🎯 What does this mean for investors?
- A high EBIT margin shows that the company remains efficient even after factoring in depreciation.
- It’s especially relevant for capital-intensive industries.
- Stable or rising EBIT margins over time are a strong indicator of pricing power and business quality.
📘 Net margin
📈 What is it?
Net margin shows how much of a company’s revenue remains as bottom-line profit after deducting all costs, interest, taxes, and depreciation.
🧮 How is it calculated?
🏛️ Why is it important?
Net margin reflects a company’s overall efficiency – across operations, financing, and taxation. It shows how much actual profit is generated from each dollar of revenue.
🎯 What does this mean for investors?
- A high net margin means the company is not only strong operationally but also manages financing and taxes efficiently.
- Peer comparisons reveal business quality and competitiveness.
- Declining margins despite revenue growth can be a red flag for rising costs or inefficiencies.
📘 Free cash flow margin
📈 What is it?
The free cash flow (FCF) margin shows how much of a company’s revenue remains as actual free cash after covering all operating expenses and capital expenditures.
🧮 How is it calculated?
🏛️ Why is it important?
This margin reflects the true liquidity generated by the business – independent of accounting rules or depreciation. It’s especially relevant for dividends, buybacks, and reinvestment decisions.
🎯 What does this mean for investors?
- A high FCF margin means a company consistently generates strong cash flow.
- It’s a positive signal for financial stability and shareholder returns.
- The long-term trend is key – a declining margin may indicate rising investments or weakening operating efficiency.
📘 Earnings per share (EPS)
📈 What is it?
Earnings per Share (EPS) shows how much profit is attributable to a single share – and is one of the most important metrics for evaluating a company's performance.
🧮 How is it calculated?
The diluted share count reflects potential new shares that could be issued through options, convertible bonds, or other rights.
🏛️ Why is it important?
EPS is the basis for many key valuation metrics like P/E ratio, PEG ratio, or payout ratio. It enables comparisons of profitability across companies, regardless of their size.
🎯 What does this mean for investors?
- EPS captures per-share profitability and is especially useful for comparisons over time or with analyst estimates.
- Rising EPS may signal consistent growth or share buybacks.
- Important: Always use diluted EPS for more realistic valuations – especially in companies with stock-based compensation.
📘 Free cash flow per share (FCF per share)
📈 What is it?
Free Cash Flow per Share shows how much free cash flow a company generates per outstanding share – after investments, but before dividends or debt repayments.
🧮 How is it calculated?
Free cash flow is calculated as operating cash flow minus capital expenditures (CapEx).
🏛️ Why is it important?
FCF per Share reveals how much real cash is available per share – useful for dividends, buybacks, or reducing debt. Unlike net income, free cash flow is harder to manipulate and often seen as a more reliable metric.
🧮 Calculation
🎯 What does this mean for investors?
- High FCF per share signals strong financial flexibility.
- It shows how much capital the company can effectively reinvest or return to shareholders.
- Particularly relevant for dividend payers and capital-efficient businesses.
📘 Short interest
📈 What is it?
Short interest indicates how many shares of a company are currently sold short – that is, borrowed and sold by investors who expect the price to decline.
🧮 How is it calculated?
It reflects the percentage of a company’s shares that are being shorted relative to the total shares available.
🏛️ Why is it important?
Short interest serves as a sentiment indicator: A high value may signal skepticism or bearish expectations – but also increases the potential for a short squeeze if prices rise unexpectedly.
🧮 Calculation
🎯 What does this mean for investors?
- Low short interest usually indicates market confidence in the company.
- High short interest can be a warning sign – or an opportunity if sentiment shifts.
- Especially relevant in volatile markets or ahead of key earnings releases.
📘 Employees
📈 What is it?
The employee count shows how many people a company employs worldwide – offering insights into its size, structure, and business model.
🧮 How is it calculated?
🏛️ Why is it important?
It helps assess operational scale, labor intensity, and cost structure. Combined with revenue and profit, it enables key metrics like revenue per employee or productivity.
🎯 What does this mean for investors?
- A high headcount can signal operational complexity – but also significant growth capacity.
- Revenue per employee is a key indicator of efficiency.
- Especially useful for comparing tech, industrial, or service-heavy companies.
📘 Turnover per employee
📈 What is it?
Revenue per employee indicates how much revenue a company generates on average per employee – a key measure of efficiency and productivity.
🧮 How is it calculated?
The employee count is typically taken from the most recent annual report.
🏛️ Why is it important?
This metric helps compare business models – especially between labor-intensive and technology-driven companies. A high value suggests automation, operational efficiency, or strong value creation per head.
🎯 What does this mean for investors?
- A high revenue per employee indicates a scalable and margin-strong business model.
- A low figure may reflect labor-intensive operations or lower value-add.
- Especially helpful when comparing tech companies to industrial or service sectors.
Cerebras Systems Stock Analysis
Analyst Opinions
18 Analysts have issued a Cerebras Systems forecast:
Analyst Opinions
18 Analysts have issued a Cerebras Systems forecast:
Cerebras Systems Events
Past Events
|
AUG
18
Special Call - Cerebras Systems Inc.
about one month ago
|
|
AUG
12
Q2 2026 Earnings Call
about 2 months ago
|
|
JUN
23
Q1 2026 Earnings Call
3 months ago
|
StocksGuide Free
Cerebras Systems — Special Call - Cerebras Systems Inc.
1. Management Discussion
Please welcome Cerebras CEO and Co-Founder, Andrew Feldman.
Welcome, everyone. Thank you so much for coming. It is with such pride and joy we see all of you here today. And to point out that while we have years of experience in building the fastest AI, managing queues out in front of our building is new to us. So forgive us for the delays there. This has been a truly extraordinary last 3 or 4 months for us. The penalty of going public. Let's all stare at that for a little while, and we can understand the cost of going public, right?
This is it. Yes. Right. Right. Going public is a truly extraordinary event. It allowed me to see engineers I've been working with for 20 years. It allowed me to know they actually own a coat and tie, we didn't know this. It was a time to share with our families after a decade of work. And it was sort of humbling in the sense that we're now able to step forward in a different way and participate in the biggest of the big leagues and be part of the show.
Now we were able to do this in part because we rode this enormous wave. All right. If you close your eyes and think back 3 years ago, right, AI was a novelty. All right? It was sort of a parlor trick. It was something you showed your friends and didn't use at work. And then it became useful. All right. And now it's a necessity. And in that necessity, speed, latency, shapes what we can build. And real-time AI feels like an active collaborator when you're engaged with it.
In AI, speed is productivity. And as you know, what we do at Cerebras is deliver the fastest inference speed in the industry. Now when you have that speed, AI responds in real time, users do more with AI, they stay longer, they run more interesting workloads. They solve more interesting problems. Speed makes new markets and allows new ideas to flourish. But until recently, there was a trade-off. There was a trade-off between speed and throughput. Everybody wants speed, right? Nobody says give me something slower.
Or said differently, if you really want to punish a naughty 12-year-old, don't take away their phone, set it to dial-up speed, right, for a week, right? This drives good behavior. All right? You had a trade-off. You had a trade-off between smart and fast, all right? If you went fast, you'd go to smaller models, more specialized models. All right. But when speed is no longer a challenge, you can have both speed and intelligence. And then you can move AI to time-sensitive work, then you can move it to new parts of the business.
You can all sorts of new things are possible. The underpinning of our performance is our wafer-scale engine. And I can't share with you how many people along this journey told us it would never work. And when it worked, they said you could never make it in volume. And when we made it at volume, they said, you'd never package it, and then JP figured out how to package it. And then they said you only had 1 customer in the government. And then they said after we had a sovereign cloud, they said you don't have a hyperscaler. Now we have hyperscaler. They said, oh, you don't have a frontier lab, and now we have frontier lab. All right.
That's the journey, not just of Cerebras, but for everybody who does work, that's not obvious. It is hard. It is challenging. This is at the foundation of our work. But the foundation is insufficient, right? To deliver fast AI right now, you need a chip, you need system and a rack, you need to build massive clusters and you need software to tie it together seamlessly. And a lot of what we're going to be talking about in the next hour and a bit, all right, is how we build systems, how those systems tied together into clusters to deliver industry-leading performance and set the new bar.
Now speed is no longer sort of an infrastructure metric. I want you guys to really think about that. All right. For a long time, it was a benchmark. Look how fast we are. But right now, it shapes how products are built and it changes fundamentally the user experience. AI as it becomes mainstream, users expect it to feel like the best software they've ever worked with. They expect AI not to be different. It's going to be in the same category as their favorite software tools.
They want to be immediate, fluid and responsive. And when the response arrives before your attention can move, you see in the problem. All right, and you keep building. And as an infrastructure builder when you allow someone to run a software that allows them to build interesting things, you're proud. That's why you build infrastructure. All right. And that's why we do this Okay.
Last week, OpenAI announced a first look at GPT-5.6 Sol ultrafast mode. Ultrafast is a new service tier that runs OpenAI's most intelligent models at up to 14x faster than standard. We're testing this with select customers to start what this brings the industry for the first time. It's frontier intelligence instantly, right? 14x speed removes the speed, intelligence trade-off. You can now have both. There is only one place you can get this, frontier speed frontier Intelligence and instantaneous speed.
Customers now complete more useful work per second. Now I'm going to show you a little demo because that's what I want to do, Oh, new slide. Keep the CEO on his toes. Reorder the slide deck. This is a cool slide. On the x-axis, you have speed. And on the Y axis, you have intelligence. All right. You want to be up and to the right on both. You want to be smarter and faster, and that puts you up here. And the only thing up here is GPT-5.6 Sol running on Cerebras. In the entire industry, the only thing here.
Now I'm going to show you a number. What you're about to see is a speed run of a benchmark called humanity last exam. And for those of you who worked up PhD, this will be depressant because it will be fast. This is a leading industry benchmark, okay? And it's designed to show graduate-level reasoning. And we will be on the left. There are 2,500 questions. We're going to show one to start. And then we're going to go to the entire set. Let's go.
Okay. We're done. First question. All right. Now they're done. Now what we're going to do is take the entire exam. We're done. It took 11 hours, 11 minutes and 26 seconds, less than half a day. The other guys took more than 3 days. As my old soccer coach could tell you, this is the only thing I've ever been fast at. So I really appreciate the applause.
Okay. What I'd like to do right now is invite on stage one of the leaders in the leading frontier lab, I'd like to invite Thibault on stage. Thibault is Head of core products and platforms at OpenAI. And we'll go through a few questions and hear a little bit about what OpenAI is thinking around AI and ultrafast AI.
Thibault?
It's good to see.
So if you ever think you're having a good year or a good several years, OpenAI success can disabuse you of that fact. When you're growing as fast as any company in the history of capitalism has grown, right? You have a pretty good run. And what I understand right now is that you have 15 million people a week running Codex and ChatGPT agents, awesome numbers. And these products have sort of made it an enormous impact on us.
I mean if you walk around our lab, and I think this is true in most companies now, everybody's got a GPT screen up. All right. Talk a little bit about your strategy. Tell us a little bit about how things have been evolving and how you think of the landscape.
Sure. I mean first of all, great to be here. It seems like we coordinated on the naming of our models in Supernova, Sol Supernova.
We are. We're trying to follow with the celestial theme.
it sees more coordinated than it is.
Not that coordinated. I have no fear.
For me, the strategy has been like very simple, like build the best models and then start them at scale. And then recently, we also realized, well, it's all going to be about agents. So you've got to built like the infrastructure you have agents, run at scale, run safely, have like very aligned models and then build delightful interfaces to those agents so that they can do actual work for you.
And up until recently, we kind of have these trade-offs where if you wanted to get something done fast, you have to serve like a smaller model and you have to serve it at lower latency. And then you always have like this decision to make up like, hey, maybe I want something done now, so I'm going to use like a Luna-size model or like, okay, we can think a little bit more, so we're going to use Sol and Ultrafast is kind of don't have to make that decision anymore. It kind of gives a glimpse of what's to come, I think reducing the cognitive load and just building something that is delightful, just works as fast when it needs to and then can also run slower in the background, if you don't need at larger scale.
Yes. I think that's super exciting, and we're really proud of our partnership. Tell us a little bit about how Ultrafast sort of fits in with your product vision and strategy.
Yes. We would love to -- we're not there yet, but we would love for Ultrafast to just be the default. I think it is a glimpse of what's to come.
Write that down investment community.
As you saw in the videos, it's quite spectacular when you're in front of it. The first time I demoed to someone, they were like this must be fake. You just planted some fake data and like showed me the back end. I was like, no, it's just running on this hardware over there. This is just a little awesome company, Cerebras, they build great chips. So we're pushing very hard on reliability. Obviously, this infrastructure that powers a very large chunk of -- like in the future, like a very large chunk of the GDP must be reliable.
But then also we're pushing on the frontier of like how fast we can make it. And this requires rethinking the entire approach. It's not just from the inference. I think the inference matters, but also like the model is like how do we train them? How do we think about tool calls, where do we see that new bottlenecks sort of like emerging. So we rewrote a large part of a stack for Ultrafast and it was a delightful partnership.
That's really cool when sort of your compute infrastructure so fast, it pushes the software guys up the stack and then that now there's headroom and now the hardware can get faster, right, that interplay is really a pleasure.
When it started to work, the first thing that we did is like we gave ultrafast to the engineers working on making Ultrafast work and then just cranking that loop.
And you told me something back there that when your engineers have a really important problem, they'll ask for Ultrafast.
Yes. Yes. Everyone wants to access the Ultrafast. We reserve it for incidents, be it like auditors or security incidents. We have those just as any other company and like Ultrafast comes in handy. Top priority projects, we always have Ultrafast like provision for those engineers and then like important research efforts as well get Ultrafast. We don't yet we reserve a lot of our capacity as well like for serving it to our customers over the API. So we don't just gobble up all of the Ultrafast capacity for ourselves.
What do you think is going to change in consumer behavior when they get AI at sort of this speed and this intelligence?
Yes, it makes you sort of realize that before you were compromising and you had to switch tasks or you have to sort of like break the illusion that it wasn't really real-time and it wasn't operating sort of like at the speed of thought. And with these kinds of speeds and future speeds, this changes, right? You can stay in the loop, stay in the flow. It becomes like this real-time interaction. You can get a lot more done much more quickly.
And then you just sort of like every time you switch back to like normal speeds, they're like, "Oh, wow, I forgot how slow this was"
Yes. For those of you who remember, right, I mean speed transformed the entertainment industry, right? We used to go to Blockbuster and then Netflix delivered in envelopes, DVDs, right? But when the Internet got fast, right, Netflix didn't become better at delivering DVDs. They become a movie studio, right? That speed enabled an entirely new creation of new markets. They ended up buying the studio, right? And I think speed when put in the hands of entrepreneurs and developers, it allows them to create and make new things, not just be faster the things they've been building for a while.
Yes, we're excited that it's on the API because it will enable different kinds of approaches, different types of products that we haven't yet come up with. But definitely, when you're sort of like in the creation, in the flow, it feels like, okay, like now you can try something maybe draw something and then get like a few different proposals like in real time, you can sort of explore the space in different designs, so you can implement entire variations of back-end like all in real time. And it just feels like a new world is sort of like opening up.
One of the things I think OpenAI has been one of the many things that you guys have been absolutely best in the world that is sort of imagining the scale that AI could be right? You guys were out building Stargate when everybody thought it was roughly big, and now everybody said, "We need a 10x bigger, it's too small." Right? Tell us a little bit about how -- I mean, educate us is how an organization thinks about scale in the way you guys do.
I think we just have deep conviction that, and we had deep conviction that the models would get better. And at some point, you reach a point where you run the model on the GPU and it's providing more value than the cost of running it on the GPU. And at that point, you want to have all the capacity in the world to just like run it. And this will continue to be the case. Like we don't see a slowdown in the level of capabilities that we're able to develop with our models.
And so every time we have a new generation of models, we're like, well, like the utility that we produce on the current compute that we have increases and already vastly outpaces like the costs. And then this is a thing that we talk about all the time, right, which is like, can we get more capacity? Can We get more capacity? It's like the demand...
They talk about it all the time. believe him right?
[indiscernible] fast outpaces the capacity that we have like across our entire fleet. And I don't think there is a reason to believe that, that will slow down or change.
All right. And last question. We're just sort of in the first few months, for 6 or 8 months of a multiyear, your, partnership, and we announced sort of the frontier model at speed as a first look. As you look forward, sort of what are you most excited about? I mean you've got maybe the best view of our industry of anybody.
What I'm excited about is like it still feels extremely early and maybe going back to like why we were building all of this capacity and why we continue to invest and we were early. Like if you look at the adoption of ChatGPT is like above like 1 billion users right now. But really the fraction of those that use those more sophisticated agentic workflows to help them day-to-day, now those are the numbers that we have published like 15 million. That's like the number we published last week, like we'll probably publish another number end of the week, even more impressive than that.
But it's growing fast, but it's still like a very small fraction of that 1 billion. And so it feels extremely early. And there's sort of like this thing right now is like you use it and if you're technical, you're kind of like used to it being clunky. But the illusion is not quite perfect yet. You don't yet have this personal AGI in your pocket that knows everything about your goals, your schedule for the week, like anything that [indiscernible].
How you preferred I mean ...
Every time I try to write something, I was just like, oh, remember my tone of voice is like I have this somewhere in the file, in a skill, it's like -- that is clunky. So all of that is going to sort of reduce in something much simpler, safer as well and like readily understandable by everyone in the world, and then that's going to be something quite different to what we have today.
What an exciting vision. Thibault, thank you so much for the partnership and coming on stage and sharing. Ladies and gentlemen, Thibault. Okay. We have a little video of some what customers think, both inside of OpenAI and elsewhere. Let's go.
[Presentation]
Okay. Let's race forward. While we love our partners in the closed source community, we also are fastest on the full range of open source models. All right, from large to small, U.S., non-U.S. on these, like just about everything we do, we're up to 15x faster than the competition. Now one of the things Thibault talked about -- one of the things that's extremely sort of top of mind for us is the acquisition of data centers. And we have data centers rolling in. Here's our data center in Santa Clara. Here's our data center in Toronto; Dallas, Texas; Minneapolis, Minnesota; Montreal, for those of you who are Canadian, I got the little thing over [indiscernible] it was pointed out to me twice in presentations previously.
Oklahoma City. And for those of you who don't know what they look like when you're building them, here's our data center in Alabama that's going up. Here's our data center in Lyon, France going up. Here's our data center in a part of Norway that I can't pronounce. And here's our data center in Michelle Finland, one of several that are going up, right? We now have data centers across North America and Europe. And we've brought on 600 megawatts, right, that's online or under contract for delivery by the end of next year.
It's a huge effort and it's just the start. It's not nearly enough, but we have a whole team chasing this every single day. Now what you put in data centers, all right, is much more all right, than a fast AI accelerator, right? AI is generated by a cluster, all right? It's generated by a combination of equipment, some of ours. And some of others. And what I'd like to do right now is invite on to stage someone who's sort of had an extraordinary career. At Cisco, she was, in my view, singularly responsible for the rise of Cisco in the late '90s into one of the great companies of that era. She ran the Catalysts division, which was Cisco's monster. All right. She's been CEO of Arista now for, I think, 14 or 15 years and has led that company to extraordinary success. Please join me in welcoming Jayshree Ullal to the stage.
Good to see you.
Congratulations, Andrew. What a great company. What do you guys think? Cerebras.
Okay. As we -- Well, first, I began my career in networking, competing against Jayshree, and I would advise against that. For those of you thinking about building switches, do not compete with Jayshree, that is some sort of unhappy stuff.
it's a lot more fun partnering...
It is fun partnering. So look, one thing that's become clear as we sort of think about the AI landscape is the amount of equipment that goes into these clusters. And the sort of the need for it to be coordinated and work together that if your network or your fabric can't deliver the goods, it doesn't matter how fast your accelerator is you're going to get bitten. And so as we think about that, sort of how do you see Arista sort of working with Cerebras to deliver these sort of extraordinary solutions.
Absolutely. First of all, I think you're only as good on the compute side as the network. Imagine if all your Cerebras stuff were idling and waiting for the network, right? So I'm hoping my network can pay for itself by making your compute much faster. We live in a world of you talked about 600 megawatts, you're going to go gigawatt, terawatts. We live in a world of more and more thousands of tokens. We just have to deal with trillions of parameters, but you just can't do that single-handedly.
So it's really about building the best-of-breed stack together. And other companies would have you believe that everything is built by one. I think the power of all of us together is far greater than each one alone. So it starts at layer one. I'm creating my own OSI model here...
No, Let carry on.
For those of us who learned the textbook that had 7 layers, the physical power and cooling and infrastructure is powerful. I heard you talk about how you're running around trying to find this because compute power cooling is the scarcest commodity, right? And if you can get that, you grab that in any form and shape. And then, of course, is the compute accelerators.
And nobody does this better than Cerebras from an inference perspective, of course, there's a few others we all know about that does some training. But here, I'm going to focus on the inference. But to make this all hum, we've had to really think about what we do differently with AI than we did in the past with cloud. In cloud, it was easy, to just say, okay, we'll just throw a bunch of capacity, do low latency, have lots of bandwidth and deal with one set of traffic, which was cloud computing.
In AI, you really have different forms of traffic. The fidelity is different, the type of traffic is different, the any-to-any is different. So we had to build different types of traffic to deal with networks to deal with different classes of frontier models and application agents. And I think this whole area, especially of frontier models and agents is evolving too, because the enterprise has hardly come in yet.
I mean, that's a really important point that the enterprise has been on the sideline to date.
Yes. So we mostly talk about the neo clouds and the hyperscalers, et cetera. But I believe when you're here next time, you're going to be talking about a lot more agents, agentic AI even going into our phones, and this is going to be powerful. And behind all of that is, of course, the importance of a network.
I have no questions. Keep going.
Oh, you don't. Okay. So given our engineering background, I like to think in terms of X and Y and Z axes, right? And frankly, as a student I was terrible at the third dimension, but it's important. So when we think of this and how we work with Cerebras accelerators, we think scale up, how do we connect to as many of your, I know you don't build stands you only build dinner plate, right?
That's right.
So as much of that to get the greatest [ Radix ], whether it's 64 or 128, how much of that can we connect until we run out of compute capacity and network capacity. And then we go from scale up to scale out, how do we connect these racks together. And this is where we come in, and I think we worked very closely together. But there's another emerging market, which is how do you distribute the compute. You're never going to get enough just to be in within 1 surface area.
And this is where you can have a multi-tenant capacity, multi-tenant engineering, where you can secure different connections of compute in a distributed fashion across distance. And this is what we call the scale across that goes across locations. So that's a little tutorial and networking for me.
Look, I think as we sort of grown our partnership. Jayshree supported us when we weren't buying very much at all. And now we're buying a fair bit. We appreciate it. Thank you. And maybe by way of final question, what do you see in the next 2 or 3 years for the networking industry and for the sort of demands placed on it by this new type of compute?
Well, I think at one level, we're all going to be pushing the envelope on bandwidth and latency. Just to give you guys a perspective, we were all on 10 gigabit back when you were doing networking for 10 years, right? But now we've gone from 100 gigabit to 400 gigabit to 800 gigabit...
We won't tell them that you and I begin building fast Ethernet switches.
The rate and pace of throughput and capacity is now every 18 months. We've got to keep up with your computing processes, right? So in the next 3 years, I fully see this going in completely wild directions of 3.2 terabits or 6.4 terabits it's not stopping. Now when it doesn't stop like that, you also need the physical connectivity. So the rate at which optics or cable or any kind of connections happen at good distances has to also keep up. And that's a nontrivial challenge when you go at high speeds, whether it's coherent optics or co-package or even co-package copper when you stay within Iraq. So I think the future of that is very significant.
But there is one other thing I want to bring up, which comes back to, I don't want any of your processes idling. The biggest challenge going forward will be how do you keep -- retain the availability of your compute by making -- building a good network. In other words, suppose these compute cycles go away or they break down or somebody pulls out a cable, how do I restore and recover. And this is where while hardware is very important, software has to really help the recovery of this. So smart system upgrade, high availability, level of automation, analytics will be super important in the future as well.
Ladies and gentlemen, Jayshree.
Thank you. Thank you very much Andrew.
Good to see you. Thank you so much. I'll say it again, do not compete against her Okay. I'm going to take 5 minutes right now and show you NVIDIA's road map for the next 5 years and ours.
Okay. Let me explain. The X-axis is tokens per second per user, okay? This is how fast you experience AI. All right. The y-axis is how many tokens total is delivered through the solution, right? What that means is, it's the number of users that can simultaneously get the speed that's on the x-axis. We call that throughput. Everybody understands it's really important. Speed per user, number of users times speed per user, right. Okay. They are extraordinary in this domain. This is the graph for us.
We are extraordinary in this domain. Okay. Now this should tell you exactly what the future is. This is what the road map for GPUs will be. This is what our road map will be. They will try and get faster without giving up throughput, and we will try and get more throughput without giving up speed. That is the competitive landscape. Now to this end, there is a solution that can be achieved through partnership.
And that solution is called disaggregation. And by using the GPU to do part of inference called prefill, and using Cerebras to do part of inference called decode, right? You can get 10x faster than the GPU and 5x more throughput than Cerebras. And this is an extraordinarily compelling solution. Now I'd like to invite to the stage someone who I've admired a great deal in my career. He began his career at IBM, where he worked on the PowerPC and on their first blade servers, it this next part. He then went to work for Steve Jobs. He report to Steve Jobs and ran the iPhone business.
He then went to AMD, where he was the technical sort of visionary behind the transformation that has produced the company that they are today. I'd like to welcome Mark Papermaster to the stage. Andrew?
Great to see you. Congratulations on this incredible event.
One of the things that I forgot to say about Mark is he also is 1 of the true gentleman in the industry. Really, you should applaud that because there are not many of them.
Thank you. And to you as well.
Okay. Let's talk a little bit about disaggregated solutions. Why now? Why was now a good time to build disaggregated solutions?
Well, Andrew, I think it's really a statement of this inflection point that we are at right now because you think about the massive infrastructure, and you were talking a moment ago about the whole GPU build-out, it's been so dominated by training and it had to, to get the kind of foundational model capabilities we have. And that's what's fueling this transition right now because with that kind of capability, people are finding the boundless applications that we can inference on and get real work done. And so it's a huge demand for all of us in the industry to figure out how then to accelerate inferencing tokonimics that have to be optimized, helping people get their job done. And so I think disaggregation is an innovation driven out of necessity, right?
I think that's right. I think 1 of the things we forget was so you make AI smart. You make it with the training, right? But once it's made, once it's smart, we're going to use it. And we use it with the inference. And as that sort of gets more mature, we're able to attack it with specialized solutions with solutions and partnership. Tell me a little bit about how you think about the sort of disaggregated solution bringing together the best of both worlds. I've got a slide here. You have a slide here. Let's see here. How we bring together sort of the best of both worlds?
Well, that's what we love about this partnership. We know each other well. And when this problem really need to be solved it was a partnership of 2 solutions that are better together. So what have we been focused on at AMD. You see it with the Helios rack coming out at the end of this year. It's a throughput monster.
Monster. Here we could show you, it's a monster.
Right. So it's 72 GPU rack, super high bandwidth, and it's everything about it in terms of how the CPU, GPU, the networking is all about just incredibly efficient throughput. And so it is a perfect solution for a broad range of computing. But when you want to also have a low latency response, and we know you guys are really serving the need of these many applications, growing applications that need latency than what better solution that disaggregated where the teams have really worked together to separate out hidden from the end users, right?
It's really hide under the covers, [indiscernible] prefill where you've got to take a very broad context, and it really demands a parallel computation that the GPUs are perfect at, right? And so you can just -- you can get that economic throughput. And then when you think about decode, when you're really processing that token throughput and optimizing for ultrafast response and low latency, right, send that decode to the wafer scale engine. And this combination is a win for everybody. The users get better economics and low latency and that quick response.
So it's -- again, I think this kind of disaggregated solutions, really, meeting a market need.
I think that's right. I think one sort of reasonable way to think about the chart I showed you before, is that throughput drives economics and speed drives user experience. And when we can bring them together, we're in this position of sort of you can have both.
Exactly.
And that's enormously powerful. Now first, Sean and I have had a lot of fun working with your team and working with Mark's teams is always a joy. So we appreciate that. But you're bringing AI across AMD, and you've got lots of parts. Tell us a little bit about how you're doing that.
Well, it's an extension of what we just said. I mean AI is being used for so many diverse needs, and that's what we focus on our portfolio. So it is a Helios rack at the top of our offering for those toughest both training and inference and this huge inference throughput that we just talked about. So that's our Instinct product line. We have our epic product line with high-performance CPU servers.
We use those in our cluster.
And that's where we've been -- developed this tight partnership for years between AMD and Cerebras. And then, of course, local AI. So when you need to run, we're seeing more and more where we're getting [indiscernible] models that can run efficiently locally, often on open weight models. And then, of course, the adaptive compute and embedded. So the way we think about it at AMD is a diverse set of workloads and an open ecosystem.
That's sort of essential to us. It's a rack and AI stack that's open and the ability to bring a diverse set of solutions. More and more we need innovations exactly like we're, I think, paving the way together with the disaggregated solution of AMD and Cerebras. This is a really fun partnership.
Just to wrap up, and Mark came, they had a Board meeting this morning. He's racing back to dinner. So we thank him for making time. When you look out into the future as sort of the CTO of AMD, and you think about the compute needs in the future around AI and other. What do you see?
Well, I used the word this insatiable demand for more computing. And so what we see now is the fact that we're still at the early days, of AI applications, inferencing application, it's actually scary what that insatiable demand is going to be. And so what I see going forward is we need more and more innovations like we're doing together. And I think the algorithms are going to evolve. And what we're going to see is a more and more integration of diverse technologies to optimize on today's algorithms and the algorithms of tomorrow.
I think that's exactly right. What I tried to think about sharing with you guys today was that we use AMD CPUs, right, to manage the software that runs on the cluster. We've partnered with AMD to build a disaggregated solutions. We use NVIDIA, excuse me, we use NVIDIA. We test on NVIDIA sometimes to see what they're doing. Now we use Arista, we use Arista to tie everything together, all right? That's the start of a solution, right? And then we're partnering with different software vendors, and we use both closed source and open source models at the top that it's taking a village. Mark, I want to thank you for coming and thank you so much.
Andrew, thank you very much.
Thank you. Appreciate it. Ladies and gentlemen, Mark Papermaster. Okay. Here's our road map for the next several years. That's it. Any questions right? We are going to get 4x faster. We're going to get 20x more throughput, right? This is what we're going to be doing. Our speed will double every year. And in the second half of next year, we'll be at 20x the throughput we are today. Now to tell you a little bit about how this is going to work. You're going to hear from a collection of different people over the next little while.
Sean is going to talk a little bit, Jessica is going to talk about. We've got some very interesting things for you. But I did want to sort of leave you with this, that this is where we're going as a company. All right, better user experience, better economics, all right? That's what we're focused on. Okay. And with that, I'm going to hand things over to Jessica, and we'll go to the next stage. Thank you so much, everybody.
Please welcome Cerebras SVP of Product, Jessica Liu.
Good afternoon, everyone. Now Andrew has just shared our trajectory to double raw performance every year for the next several years. So now what we're going to do is take all of that speed and accelerate the hell out of AI agents. So in last year, we've already seen Cerebras speed transform all kinds of interactive AI applications. It made AI search instant, it kept developers in the flow while coding, and it made AI voice conversations feel natural.
And now with AI agents, we are capable of even more complex tasks. Because instead of just giving an agent a question, you can actually give it a goal and trust that it's going to figure out its way to succeeding in the goal you've given it. Agent can autonomously make a plan, call models, call tools, and it can revise this plan in a loop over and over again until it achieves the outcome that you want. So it is actually this thinking and this iteration that makes agents so capable.
But now underneath the hood, each one of those dozens of calls incurs a latency cost. So then the smarter your agent the better your outcome. But the smarter your agent, the longer it also takes for you to get to a result. And now what if you want something faster? Well, historically, when you want something faster, you would just use a smaller model. And then you could get to your real-time interactive voice agent, and it just means that sometimes, Siri, is not going to do what you want.
So on a given time budget on GPUs, you always have to choose. You can either use a faster dumber model or you can use a smarter, slower one. It's not a great trade-off to make. But importantly, it's also a false trade-off. Because at Cerebras, you can do both. You can have a smart and fast agent. And this is why inference speed is so important to agents. Because when you can run the whole system 15x faster, you don't end up with a latency debt. You actually get latency credit. You have extra time to spend on using a smarter model on more loops, even while the overall system ends up still running faster, you can do both.
So now let's look at a couple of examples. Agents today can already show great outcomes on GPUs. Harvey's legal benchmark completed a whole set of common client tasks in just 22 minutes. OpenAI use Codex to completely build a design tool from scratch in 25 hours. And Cursor ran hundreds of parallel agents and was able to build an entire web browser in under a week. So those are some really impressive results. But it's still kind of a long time.
Any time you're spending tens of minutes, tens of hours or multiple days for something, it is a long time before you get to your response. And so think about how much more you could do. If you could achieve the same outcomes in 2 minutes or 2 hours or a single work day. If your agents could do all these tasks 15x more quickly, you could do 15x as many tasks in the same amount of time. You can serve 15x as many clients, you could complete 15x as many projects, or you can take the speed and actually just simply use it to make working with the agent 15x were magical of an experience.
You can also just use speed and speed. You can have that buttery iteration loop with the agent even if you are using a large smart model instead of waiting 7 minutes to see if the agent even did what you wanted. We can use speed to make a agentic work delightful. But the best way to see this impact is from real builders who are deploying fast agents to work in the real world today. And so we have leaders with us from Figma and cognition who are going to show us what's possible when agents can move at the speed of the people working with them. So please welcome to the stage me, [indiscernible], AI research product lead from Figma.
Cool. Thanks, Jessica. You're going to hear fast inference a lot today. I want to spend the next 8 minutes talking about why it matters for a problem space AI hasn't quite mastered yet, one that pushes on agents in ways, code and reasoning don't, design. Hi, everyone. My name is Pavy, and I lead product for Figma's AI research team. In May of this year, we shipped Figma's design agent, an agent that knows your canvas, knows Figma and works right alongside of you. Let's meet the agent.
There was so much that we thought about in exploring this idea. But next, I want to touch on a particular philosophy that went into building this agent. AI should adapt to your workloads. Some of our favorite tools today are where the interface melts away, and you can focus on doing actual work. They are embedded into your workflow in seamless ways would appear when you need it the most. When we set out to build this agent, we knew designers needed purpose-built tools that serve the essentials. As teams adopted Agentic tools to build products more quickly, false choices were emerging, speed or precision. AI generation or direct manipulation. You shouldn't have to choose.
We needed to create an agent fluent in Figma and native to the way teams work. We also knew that design is not a linear process. You start with an idea, build some prototype, [ hate it ], throw it out, try it again. The design process is rooted in exploration, feedback and refinement. And we have built this brilliant multiplayer canvas that supports the messy middle of a very messy process. That's our product. But many of our aI tools today have started shifting that experience. Designers are used to getting an open canvas with lots of exploration getting divergent ideas very quickly.
But AI tools today focus on helping you get the highest fidelity functionality and really zero in on a singular idea. Designers are used to collaborating and [indiscernible] out in the open together. But the AI tools have created silos where the work is stuck on one person's computer and it's hard to share ideas really quickly, and be inspired by one another. And so we at Figma don't think it should be one way or another. There is room for all of it at different times, and that is exactly what we wanted to build. An AI collaborator that can help you when you need it and is embedded into your workflows.
Unlike the MCP server, the agent lives directly on the multiplayer campus, no separate setup or context switching required. All right. Let's dive into the technical details. The bar for design agent is fundamentally different than most coding agents. If I ask a coding agent to write a function, there's the right answer. It compiles or it doesn't. It passes unit tests or it doesn't.
If I ask a design agent to make this feel more premium, there is no compile step. There is no unit test. The answer is right when a designer looks and says, "Yes, that's what I meant. When evaluating design, it's also subjective. What looks good varies by viewer, brand and audience. Second, it's multidimensional. A design can have the perfect alignment, but terrible color contrast or brilliant typography in the wrong hierarchy. Third, it's task dependent.
A social media post and a pitch deck have different quality bars. And lastly, it's contextual. The same design can be great for one audience and wrong for another Design, quality, resist reduction to a single metric. It's a bundle of competing signals and the challenge is turning that bundle of soft judgments into something we can measure in our research team. Also, as I mentioned, design not being a non-verifiable domain, it also has no ground tooth and the ground keeps changing.
What's great in design today might not be great tomorrow. And that's exactly what makes this such a difficult area for our research team to work on. So we decided to help solve this problem by building our own model. Specifically, a model fine-tuned for editing Figma files, that making Figma itself legible to model in ways that aren't possible with third-party tools. With deep context of your designs, your team standards and your best practices. This model helps power our agent, which brings me to why Figma is uniquely positioned to build our own in-house custom model.
While foundational models are getting better every day, they often still lack the ability to judge design quality in a reliable way. They might still have biases towards certain stylistic patterns, and we can build a specialized model just for design. Second, we have the corpus. Off-the-shelf models tend to also want to generate the average of their weight. So need to be steered. We have the data and context that allows for this to happen.
For the past decade, Figma has watched billions of designs get built layer by layer as the world's designers work at their craft and our products. And lastly, we have the Canvas. One of our beta users put it perfectly. And I quote "I'm just happy that the agent is an environment I know and love versus having to go somewhere else with MCPs and figure out how to link it together."
And last but not least, with inference hardware from Cerebras, speed can become our strategy. We can build an agent in your canvas that is cheaper, faster, specialized for the way designers actually work. So why exactly is speed so important for agent design? Firstly, design agents are call heavy. Earlier, I talked about how design is a messy, complicated process. The messy squittle wasn't just a metaphor for human design. It's also the architecture of the agent. Slow inference forces the agent to flatten that process into generation instead of designing fast inference lets the agent preserve the loop, and those loops are where design quality comes from.
Second, you can do more in less time with fast inference. Design is full of tedious work, none really hard on your own, but together, they can [indiscernible]. When the agent makes those instant, we're not saving seconds. We are returning the day so designers can get back to the fun part of design. And last but not least, designers work in loops of seconds. Every extra second between try and see a big sum from that loop, fast inference keeps you in the flow, no more contact switching. And that's the difference between AI as a tool and AI as the way you work.
Working with Cerebras hasn't just been an optimization for us. It is what helps make this product possible. Figma agent is available in beta today and will [ GA ] soon. This is a really fun way, design and collaborate and it's cheaper, faster and easier than anything ever seen before. Thank you all.
Please welcome Cognition SVP of Research and Founding Engineer, Silas Alberti.
I'm Silas, and I'm super excited to be here, and I'm going to talk to you a little bit about how Cognition and Cerebras have collaborated on building superfast, coding agents. And also, I'm going to highlight some of our research that powers. First of all, Devon, you remember Devon so you have the billboards in the city. This is Devon, our product suite, and we are mainly known for Devin Cloud, which in 2024 was the first -- the world's first software engineering agent. And especially in the last 6 months, people have gotten really on the cloud agent train and like running hundreds of agents in parallel, agents for hours at a time.
But we have way more than that, our product suite covers the entire software engineering life cycle. We have DeepWiki for understanding and planning code basis. We have Devin Review for reviewing code and Devin Automation for maintaining code. And all this is possible due to the incredible advances in AI models. And at Cognition, use a mix of models. So we use frontier models like OpenAI and Anthropic, but we've also increasingly invested in training our own models, and I'm going to talk to you a little bit about that today, 2 of our model training projects and how we're able to run them at lighting speeds using Cerebras.
The first project I want to talk about has a special place in my heart. It's [indiscernible]. And it's actually almost a year old, which is an eternity in AI era, but the cool thing about it is it was actually the first model that Cognition and Cerebras launched together and also one of the first models that our model training team builds. At the time, we had much less compute than we have today. So we have to really pick our problem as well. And we noticed that coding agents spend a long, long time, even just like finding the right files on the code base to add it again and again. And so we thought, okay, how can we make that super fast.
And so we trained a small model precisely on this code-based search task. So imagine, you give the model a question, for example, here, how does VsCode efficiently implement, file watching. And then the model's task is to explore the code base and return just the list of files that are relevant to this question. And for an researcher, this is like incredible task because it's verifiable, right? Like it's like there's an objective, a list of files that are relevant, you can just grade, did it find the correct files? So we went and trained the model. And so here on this code search eval that we built for like code-based search task, our model SWE-grep and SWE-grep-mini as you see performed on par or better than the frontier models at the time. It's kind of crazy how much happened in a year, but at the time, Sonnet 4.5 was the best model in the market.
And then we put this model on Cerebras and it ran at incredible speed. So here you see a SWE-grep at the time ran at 680 tokens per second. And so we've got many even at 2,800 tokens per second. So more than 10x faster than any LLM model. We optimize not just the tokents per second, but also really the end-to-end time of the agent took to complete the task, which also meant doing more things in parallel. So also at the time, a Sonnet 4.5 was the first model to do parallel tool calls. And those models would still do like one tool call at a time.
And we thought SWE-grep, okay, how can we push this further and push the model to do 5, 6, 7 or even 8 tool calls and searches in parallel. And so here, you can see some of the graphs from our training run. So the -- on the right side, you see the training rewards went nicely, beautifully up over the course of the run. But the cool thing in terms of the parallel tool calling is that at the beginning of the training run, the model would maybe do like 3 tool calls in parallel max and then through our training, it goes up. And at the end, it does many, many tool calls in parallel.
And this really pays off. So what we basically did is we deployed SWE-grep into our product together with Claude as the main agent. So you have to imagine like the user asked the question to Claude. And then Claude can call SWE-grep as the tool to search the code base, but then based on the results, Claude will give the final answer. And what this means is there's no compromise in terms of the quality of the answer, but you save a lot in end-to-end latency.
So interestingly, as you see here, in some cases, maybe for this question, we would get like a 2x, 3x, even 4x speed up in terms of end-to-end latency for the same quality of answer. After SWE-grep, we became more ambitious and said, okay, now let's start training real frontier coding models. So here, you see one of our charts that we released a while ago. So this is like SWE 1.5 a couple of months after SWE-grep. And it was the first frontier model that we deployed on Cerebras. I think you've seen the graphs like this already earlier today, but it's just incredible the speed, the quality trade-off. You basically get quality and products equivalent to the contemporaneous frontier models, but speeds that are like more than 5x faster than anything else.
And we've continued investing into this. So this is a more recent results on our in-house coding eval frontier code 1.1, and our latest model SWE 1.7 performs in terms of quality on par with GPT 5.5 and Opus 4.8. But only can run it at close to 1,000 per second, but it's also a lot more cost efficient. So what you see here is actually a cost versus quality trade-off chart. And SWE 1.7 is at the same cost and scales like a Kimi K2.7 Code, the Composer 2.5 or GLM 5.2, but achieve close to frontier in quality on the SWE.
So to recap our phenomenal Collaboration with Cerebras. We're able to serve frontier models at close to 1,000 tokens per second. We're serving small models like SWE-grep at multiple thousands of tokens per second and are able to achieve a 4x speed up an end-to-end task completion. And this is not -- this is just the beginning, much more to come. And thank you so much for having me.
Please welcome Cerebras SVP of Product, Angela Yeung.
Hello, San Francisco, I'm Angela Yeung, SVP of Product here at Cerebras. The next competitive advantage in AI isn't bigger models, it's time. For many years, we've asked the question, how smart can these models get? These models are already smart enough today to impact our daily lives and to make decisions in our businesses.
But no matter how smart a model is, it doesn't matter unless the answer is delivered in enough time to change an outcome. Every application, every mission-critical system has a time budget. That budget could be a few hours. It could be a few minutes or it could be milliseconds. You can think of it as a window. Within the window, an answer is useful. The answer can still influence what happens. It can shape a decision, and it can change an outcome. But if an answer arrives too late, it doesn't matter how correct that answer is. It's no longer useful and the value of that answer drops to 0.
Let's take some examples from our everyday lives, Priority Uber, same-day shipping from Amazon and Disney Fast Pass. If my Uber arrives too late and I miss my flight, it's game over, even if it got me to the right destination. So we buy time back. AI is no different. Every AI application also has a time budget. And once the useful window for that application is understood, the business value of fast inference becomes very clear. Take Armis. Armis is a company that runs a code scanning service. It looks for security vulnerabilities inside code. Powered by Cerebras, Armis completed a code security scan in roughly 1/3 of time, beating benchmark frontier models and finding more vulnerabilities and at a fraction of the cost.
That's not just a faster code scanner, that's a better product. And it's a product that customers are willing to pay for. It's a perfect example of when speed, quality and cost come together, you get an unbeatable product. In many cases, the clock is not negotiable. In payments, we have 50 milliseconds to accept or decline a transaction. In voice operations, 200 milliseconds before the human hangs up on an AI engine.
And in cybersecurity, 27 seconds is the fastest breakout adversary attack recorded on record, and it's getting faster every year. Let's think about that for a second. 7 seconds to prevent a company-wide data breach. The smartest answer that arrives after those 27 seconds has no value because the decision window has closed. That's why our partnership with CrowdStrike is so important. Cybersecurity is a perfect example of an industry where speed is mission-critical. And frontier models only matter if the decisions show up in time.
So please welcome on stage, Keith Culley, VP of Engineering at CrowdStrike.
Thank you, Angela. Great to be here. Excited to talk about the partnership. So unifying security and AI. To tell you about the partnership, I really have to go back about 17 years. So bear with me, grab a drink, it really starts with the founding of CrowdStrike, and that was actually on a plane. So 17 years ago, our CEO and Founder, George Kurtz, is on an airplane. He just took over the role of CTO of a very large computer security company. And he sees another passenger turn their laptop. And they go into something known as a mandatory boot scan. And if you're too young to remember that, it was awful. It was a 15-minute process that you had to wait as your computers scanned every single file before you were allowed to do anything.
So he notices that and says, "Well, this isn't going to work for very long. People are going to come -- this is prime for disruption." And that led to the birth of CrowdStrike. So a couple of years later, George creates CrowdStrike. In our industry, time is measured in milliseconds. Your success and failure can be between milliseconds. Thinking about real-world examples.
So if you have a lock on your door, you can open in a couple of seconds. That's an acceptable trade-off, right? It keeps your house secure. Little bit of friction, not a big deal. If that lock took 5 minutes, you probably aren't going to use the lock. You're going to find a reason to say, it's fine. I don't really need this lock. It's too much of a pain.
So it leads to poor security hygiene. And that's the same thing when you think about security and especially AI security. With computers frustrating experience leaves the bad outcomes. With that in mind, if we start from the baseline of an acceptable window to perform inspection, what can you do when you get -- where you need more time? Well, the answer is you need faster inference. And that's where Cerebras comes. In. With fast inference, you get 5x to 10x more inspection inside the same acceptable window of time.
What is that unlock? It lowers friction. It increases adoption of security. It increases the indicators, the telemetry coming in so we can correlate indicators of attack, indicators of compromise. Better security, leads to more AI adoption. And when AI is adopted properly with security, now you've realized the true value of AI at scale.
So with that in mind, the metric that matters here for us is actually time to decision. It's how long it takes to decide, should I allow this action or not? Should I warn about this auction? Should I prevent this option? Less friction, more security enabled at all the layers. That's why I'm so excited about this partnership. In just a few years, we've seen AI create a sea change across the behaviors in the entire industry. But it's still just the tip of the iceberg for what has real potential.
We think one of the biggest factors slowing AI adoption is the general anxiety around the security and safety of AI and agents. The key to getting past these concerns is faster inference being able to do more detection, response and decision quickly. That enables you to have low friction, adoption across the entire stack and security enabled across all your assets. And that's why we're so excited to partner with Cerebras.
Thank you.
Please welcome Cerebras CTO and Co-Founder, Sean Lie.
Hi, everyone. Thank you so much for being here today. When we started Cerebras, we had a vision to drastically change the landscape of compute for AI. And we did that by building the world's first and only wafer scale chip. Now this is a really big deal because we solve a fundamental problem that was limiting the entire semiconductor industry for decades. But this was only possible because we codesigned a system architecture for wafer scale, a system that could power and that could cool a chip the size of a wafer. This is our first generation wafer scale system architecture. This is the system architecture on which our current product is based.
And this is the architecture that brought Ultrafast inference to the world. It was codesigned for wafer scale from day 1, and it has served our first 3 generations of the products. the CS1, the CS2 and our current generation CSI. It's a 16 RU server that sits in a standard data center rack. And today, it's deployed at scale at our customers worldwide running production workloads every single day. Now over the last few years, we, as an industry, we have learned a lot. We've all learned that scaling AI inference is no longer just a server-level problem.
In fact, it's a rack scale problem. It's a cluster scale problem. It's a data center level problem. And we've all learned that to really take AI inference to the next level. We need more performance, we need faster interconnects and we need greater scale. And this is exactly what we designed our next-generation system to solve.
[Presentation]
The CS4 is our next-generation system that will push the frontier of ultrafast inference to the next level. Because with -- the fastest gets even faster by providing up to 2x faster tokens. And the CS4 was built for hyperscale, providing solutions I can provide up to 10x more tokens per watt. Now remember, today, our current generation is already running up to 15x faster than GPUs. The CS4 will be 2x even faster than that. Let me show you what that looks like.
I'm going to ask a model to perform a task. I'm going to ask it to create an HTML file for the periodic table of all the elements. This might be something that you or your agent might ask a model to do when you're creating a web page. We have our next-generation CS4 on the left. We have our current generation in the middle, and we have GPUs on the right. Now before I press enter watch really carefully because you might miss it because CS4 is done, and now CS3 is done and we are waiting on the GPU. Still waiting. Now it's still going in the background.
I won't make you guys wait through all of this. It is painfully slow. Now what you noticed right off the bat is that our current generation CS3 is blazingly fast at over 2,300 tokens per second. But what's crazy is that next to the next-generation CS4, it actually felt slow. Our current generation has enabled Cerebras to already be the undisputed leader in ultrafast inference.
And with CS4, we will widen that gap and we will push the frontier even further. By the way, this is still going. This is possible because the CS4 was designed from ground up to run faster and larger models with up to 2x more performance per wafer and 3x more density per rack. The CS4 was designed from ground up for heterogeneous disaggregation by adding 2x higher I/O bandwidth and 2x faster latency. And lastly, the CSI was designed from ground up for hyperscale, with 50% fewer components and up to 3x faster data center deployments. Now the way we did this is with a brand-new rack scale platform architecture that we call Nexus.
With the Nexus rack-scale platform, we designed it for modularity, so that it could be simpler to build, faster to deploy. And it is modular so we can innovate across power, compute and I/O all independently. In the Nexus rack scale platform, in the front of the rack is the power. And this is done with modular power supplies. And in the back of the rack, what you'll see is what we call pluggable backpacks. And there are 3 of them in the back of the Nexus platform.
Each one contains one wafer scale engine. Now if we look at the backpack, what you'll see is something that looks a little bit unique. The CS4 backpack is the completely reimagined server. This is now the fastest AI server in the world. The backpack is a vertical modular enclosure that provides all of the power, the cooling and the I/O to the wafer. And we've designed it specifically to be simpler than our current generation so that it can be manufactured more efficiently, with 50% fewer components and 60% more manufacturing automation.
Now what's more is that this entire backpack architecture was conceived for more rapid data center deployment because we can deploy the front of the rack, all of the power supplies upfront in the data center and then we can just drop in the pluggable backpacks on site. This allows up to 3x faster data center deployments bringing deployment times down from days to hours. Now let's zoom into the back. I love this picture. This is where the wafer scale magic happens.
Because inside the CS4 backpack, we have a brand new wafer package that has a brand-new direct vertical power delivery system that provides 2x more power and 2x more cooling to the wafer. Additionally, we also have a brand-new wafer I/O module that has 2x more I/O bandwidth and 2x faster latency, and it was designed to be modular to support future upgrades as networking standards evolve. Now if we pull this all together, what you see is that the CS4 system delivers 6x higher system-level performance than our current generation. higher performance.
Now this is possible because the CS4 is the first system to use our faster wafer scale engine called the WSC 3 Turbo. And the CS4 is the first system to integrate 3 wafers into a single system. With these wafers, the CS4 has 6x more memory bandwidth, 6x more compute, 6x more fabric bandwidth, 6x more I/O bandwidth at half the latency. All of it enabled with a brand-new system architecture.
Now these are some truly mind boggling performance numbers. But in inference, the number that matters the most is memory bandwidth. And the WSE 3 turbo chip in the CS4 system, each wafer, each chip has 43 petabytes per second of memory bandwidth. That's 2,000x more memory bandwidth than Ruben, 2,000x more memory bandwidth in NVIDIA's next-generation GPU. Now the reason why this matters is because in inference, all of the model weights need to be read from memory over and over for every single output.
And on GPUs, the weights are stored off chip. They're stored in a separate memory device called HBM. And they need to traverse this very thin and narrow memory bus to reach the compute. Now Cerebras on the other hand, because our chip is so massive, we can fit all the model weights in the on-chip memory. And we pait it with a ton of compute. And by doing so, we completely remove the memory bandwidth bottleneck. Now in the GPU world, they try really hard to work around this memory bandwidth limitation.
And the way they do that is they take the model and they try to distribute it across multiple chips. They take the model experts and they distribute it across multiple GPUs in many forms of parallism, tensor parallelism, expert parallelism, all in an attempt to aggregate the memory bandwidth from multiple chips by accessing them in parallel. But that is super complicated. It has a tremendous amount of performance overhead because there's so much complicated communication between all of these chips because in the end, it's just one problem that needs to be brought back together.
So the result is actually slower performance, but not just that, it's higher power, it's higher cost. On Cerebras other hand, because we have so much memory bandwidth, we can run all of the more experts on a single chip. All the experts are interleaved on the wafer memory. There's no cross-chip communication. There's no complex routing. All you get is ultrafast performance because it's simple and efficient. Now all of that complexity on the GPUs, this is what all of that complexity looks like physically.
This is a Rubin MVL72 rack. And if you look under the covers, what you'll see is that you'll see thousands and thousands of cables. What a mess. Now what's really funny is that NVIDIA would have you believe that this is a really good thing, right? They talk about how they have 5,000 cables in every single rack connecting together all their GPUs, and they can provide more bandwidth in the entire Internet running through these cables. That's a good thing, really. I mean, how much does these cables cost in terms of performance overhead, in terms of power, in terms of actual dollar cost, in terms of reliability of communication.
On Cerebras, all of the communication in that cable set is done on the wafer. All of the communication is done with no cables because it's all on chip. And what's more is that on the wafer -- we have 200x more communication bandwidth than all of those cables combined because it's all on chip.
Now what happens when you have to off chip. To go off chip, in the CS4, we designed a brand-new next-generation wafer I/O interface. This has a new wafer I/O module that extends the fabric from the edges of the wafer. And it's designed to be modular and programmable so that we can extend it in the future as networking standards evolve. Here, we've actually taken a page out of the chiplet playbook by separating the I/O from the compute silicon, we can innovate on each independently. We have higher bandwidth, and we have lower latency because we have a brand-new direct wafer link interface.
And this new wafer I/O module continues our commitment to standards based networking with RoCE RDMA over Ethernet. With this new wafer I/O module, we can run the wafer links faster which gives us 2x more bandwidth per wafer, 2.4 terabits per second compared to 1.2 terabits in our current generation. This new wafer module, wafer I/O module, also has a brand-new low-latency packet processing pipeline that gives us 1.7x faster latency through the network down to 3 microseconds from 5 microseconds in our current generation.
And lastly, the new wafer I/O module has brand-new direct wafer links, which lets us connect wafers directly to one and another, bypassing the traditional network. This improves wafer-to-wafer latency by 2.5x and brings the latency down to mere 2 microseconds. Now why is all this important? To understand why the wafer performance -- the wafer I/O performance is important. We need to look at how it's used to run large models across multiple wafers.
So to do that, let me first start with our cogeneration CS3. Today, our current generation, in production CS3, already runs the largest frontier models, and it runs it really fast. GPT-5.6 Sol, as an example, today already runs on our current generation CS3 at ultrafast speeds. GPT-5.6 Sol is OpenAI's leading largest, most intelligent model already running on our current generation. Now how do we do that? It's actually pretty simple.
We do this by mapping the model as a pipeline on to the wafers. And this mapping is very natural because it maps to the model architecture directly, making it seamless and fast because we can keep all of the high communication bandwidth on the wafer, where we have all of that memory bandwidth where we have all that fabric bandwidth. And we're only transmitting activations between wafers. So now you're starting to see why the CS4's I/O is important because we can already run the largest frontier models on our current generation. But with CS4's higher performance and lower latency, we will be able to run even larger models of the future.
Here's a graph that shows the wafer-to-wafer I/O latency versus the model size. On the x-axis, is the model size in trillions of parameters. And on the y-axis is the total I/O latency across that entire wafer plan, all added up, all combined. And this line is the CS3. And this is CS4, 2.5x faster with 2.5x lower latency because of the new I/O module. And what you can see is that even 10 trillion parameter models have a mere 0.2 milliseconds of I/O latency aggregate across the entire pipeline of wafers, 0.2 milliseconds. That's just a fraction of a millisecond. And recall that if the entire round trip latency is one millisecond, that equals 1,000 tokens per second of generation performance.
So what this means is that even 10 trillion parameter models can run at 1,000 tokens per second, and the I/O is not the bottleneck. Now what's more is that all of these numbers don't even include speculative decode. So this will push even higher performance numbers. This means that with CS4's advanced I/O, we will enable sub-millisecond latency or more than 1,000 tokens per second on frontier models with 10 trillion parameters or even more in the future, all because of our optimized wafer-to-wafer I/O latency.
Now I/O is important beyond just wafer-to-wafer I/O communication. In fact, we designed CS4 so that it can connect to other hardware infrastructures. We designed CS4 for disaggregated inference. At Cerebras, we believe very strongly in a disaggregated heterogeneous ecosystem where the user can choose the best hardware for the job. And this is the reason that we have hardware partnerships with AMD and with AWS to bring this aggregated solutions to the market.
But to understand why disaggregation is valuable, let me show you how it works. Inference has 2 parts. The first is called prefill. This is when the model is processing the user's inputs. And the second part is called decode. This is when the model is generating the output. Now what's actually happening under the covers is that during prefill, the model is processing all of those inputs and it's trying to make sense of it by creating an internal representation of what it all means.
We call this the context. Now this context is really, really important because the context is what is used during decode to generate output because during decode, the model takes that context and generate output, one token at a time. And while it's generating output one token a time, it's extending that context until the end of the output. So now if we step back a little bit and we look at what's actually happening in each of these phases, you can see that their properties are very different.
First of all, during prefill, since the model knows all of the input upfront, it can process all of those tokens in parallel. This means that it can reuse the model weights over and over and over. And as a consequence, it has very low memory bandwidth requirements. Now decode on the other hand, is completely different. It's completely opposite. Because during decode, every single output requires reading all of those model weights from memory over and over and over. And so decode, because of its serial nature, requires significantly higher memory bandwidth.
Now Cerebras can run both prefilled and decode. But as we all saw, Cerebras runs decode ultrafast. And similarly, the GPU can run both prefill and decode, but GPUs run decode slowly. So since prefilled doesn't need high memory bandwidth, it can be offloaded to the GPUs. And since decode needs high memory bandwidth, it can be offloaded to Cerebras and we get the best of both worlds. This is the value and the power of disaggregated inference, the best hardware for the job, efficient prefill, come on. Efficient prefill on GPUs and the fastest decode on Cerebras.
But remember that context, that context that was generated by the prefill but is used by decode. Well, when everything is running on the same hardware, that context can be generated locally and used locally. But in disaggregated inference, it needs to transfer from the GPU to Cerebras. And the time to transfer that context directly impacts your TTFT because the decode can start until it has all of the context. And for very large models today, that context could be tens of gigabytes in size. That's like transferring multiple HD movies worth of data on every single request.
And so now you see why the CS4 is important because with 2x higher wafer bandwidth with 2x faster latency, we directly reduce the transfer time, which directly reduces the TTFT and improve throughput. Now what I just explained is a very common form of disaggregation called prefill decode disaggregation. But it turns out there's many other forms of disaggregation. For example, attention-FFN Disaggregation. And there's many of forms that are being invented every single day.
When we designed the we anticipated this. So we designed the CS4 I/O module to be a programmable universal disaggregation interface. So that is designed to support all forms of disaggregation and provide a flexible integration point with all other hardware. And we do this in 2 ways. The first is our commitment to [indiscernible] based networking for universal compatibility. And the second is by making the module programmable. We support future protocol extensions as networking standards evolve.
So now let's pull together everything that we just talked about today. And let's look at the throughput interactivity landscape today. We're very familiar with this graph now, right? GPUs, can run at high throughput, but they're slow. This is our current generation CS3, up to 15x faster than GPUs. This is what invented the ultrafast segment. But with CS4, we are pushing this frontier even further by providing up to 2x faster tokens and providing solutions that can provide up to 10x more token capacity.
With the CS4, we are forging a brand new frontier for ultrafast inference. And the CS4 is where the fastest gets even faster, and it's built for hyperscale. Now let me show you what this looks like in terms of -- so a little bit of IT issues up here. Let me show you what's going on here across multiple different models. This is our production performance today across many different models. Small models like Gemma, to medium-sized models like GLM and Kimi, all the way to the largest, most intelligent frontier models like GPT-5.6. Now this is the CS3 performance today with our current generation product where we are already up to 15x faster than GPUs.
This is CS4, up to 30x faster than GPU solutions. This level of performance is transformative because it will completely transform user experience, up to 30x faster tokens means significantly more interactive, engaging applications. It means offline applications now can become interactive. It will completely transform our agents because of the 30x faster tokens means 30x more reasoning means 30x more agentic calls, the CS4 enables a new era of more capable and more intelligent agents. And it's built for hyperscale with higher performance and higher density, faster manufacturing, faster to deploy, so that we can bring more ultrafast tokens to the world because with less power per token, it means you can get more tokens per data center with less cost per token, you can get more profitable data centers. This is transformative and it's available now. The Cerebras next-generation CS4 is in early access right now and will be generally available later this quarter.
But we're not done yet because we designed the Nexus Rack scale platform architecture from day 1 for multiple generations of products from CS4 to CS5 and CS6. This is enabled by the modular design that allows us to independently innovate on power, compute and I/O rapidly. This is what enabled us to codesign the system architecture with our next wafer scale engine, which will be in the CS5 in 2027. Our brand new Nexus rack scale platform architecture is the foundation of our road map commitment for 2x more speed every single year.
And this rack scale platform architecture is the foundation for our road map commitment to provide solutions with up to 20x higher throughput by 2027. With this level of performance and this scale, we can provide ultra-fast inference to everyone to more customers, to more developers to more users to all of you. And I personally believe that we are just scratching the surface. And so I am so excited for the future of ultrafast inference. Thank you very much.
Please welcome Cerebras Chief Marketing Officer, Julie Shin Choi.
How's everyone doing? All right. I'm Julie. I'm the Chief marketer here at Cerebras. And on behalf of our team, thank you so much for being the best crowd ever. And wafer and I appreciate you so much, okay? So today, we had a lot of announcements, but the main takeaway here is that Cerebras, we live to serve the bestest AI for each of you, right? We're partnering with the absolute best customers, companies, developers, partners in the world, and we want to bring the fastest AI tokens to each of you ASAP. Who here wants to build with CS4. Can I see a raise of hand. Makes them noise guys, make some noise. We want speed. We want speed Okay. So we're going to work on that. So after Thibault left, I stopped him. He's in a rush because he has to go back to work at OpenAI down the street.
And Thibault and I had a conversation and Thibault like Julie, I really love the crowd there, the vibes were just insane. And we need to do more. And I said, "Thibault, I think we need to -- we need to just give these people, we need to just help them fly. So we're going to work on ways to open up more and more of this amazing frontier level speed for all of you. And folks that came to Supernova will be sending you special ways to get on early access lists and just keep us honest, all right?
Okay. A little bit of logistics. After this, no more talking. This room is going to be turned into a party space. So the Midway is very famous. This is known for good sound system and the floor here is known for that thing. So we have some amazing musical talent coming this evening. So we'll come back here at 7:30 to listen to DJs, including Lucy Guo, an amazing technical founder and very talented musician. And then -- so everyone exit here after I leave this stage, go through those doors. And then we have plenty of demos from Deep Mind, cognition CrowdStrike, AMD, like OpenAI, us, all, everyone. So we have demos, go check those out, get some food and then go to CAFE Compute and make a new friend.
Okay. Cerebras is here to answer all your questions about fast inference. So don't be a stranger. All right. Let's have some fun. I'll see you back in here at 7:30.
[Break]
Cerebras Systems — Q2 2026 Earnings Call
1. Management Discussion
Good afternoon, and welcome to the Cerebras Systems Second Quarter Fiscal Year 2026 Earnings Conference Call. [Operator Instructions] Please note that today's call is being recorded. I will now turn the call over to Sean Dorsey, Head of Investor Relations. Please go ahead.
Thank you, operator. Good afternoon, everyone, and welcome to Cerebras Systems' Q2 2026 Earnings Call. Earlier today, we issued our press release and posted our supplemental earnings presentation to the Investor Relations section of our website. A replay of this webcast will also be available on our Investor Relations website following the call.
Joining me today are Andrew Feldman, our Co-Founder, Chief Executive Officer and President; and Bob Komin, our Chief Financial Officer. Before we begin, I would like to remind everyone that today's discussion will include forward-looking statements under the safe harbor of the Private Securities Litigation Reform Act of 1995.
These statements include, but are not limited to, statements regarding our future financial performance, business strategy, market opportunity, customer demand, product road map, technology leadership, supply chain, operating model and outlook for Q3 and full year 2026.
Forward-looking statements are based on our current expectations and assumptions and are subject to risks and uncertainties that could cause actual results to differ materially from those expressed or implied. These risks are described in our SEC filings, including our final prospectus related to our IPO and our future periodic filings with the SEC.
We undertake no obligation to update these forward-looking statements, except as required by law. During today's call, we will also discuss certain non-GAAP financial measures. Reconciliations between GAAP and non-GAAP results are included in today's press release and supplemental materials, which are available on the Investor Relations page of our website. With that, I'll turn the call over to Andrew.
Thank you, Sean. Thank you all for joining us today. Q2 was a strong quarter. We completed our public offering, but we did not let that distract us from execution. We delivered record core revenue and beat guidance on all metrics, core revenue, core gross margins and core operating margin.
Looking forward, we see unbound demand for fast inference. The market is realizing that speed is not a benchmark item. Speed changes user engagement, it changes Agentic genic performance, and it changes AI productivity. Fast inference unlocks new applications and new markets. As we've shared with you previously, 2026 is a foundation-building year for Cerebras.
We've made excellent progress on multiple fronts in the past 7 weeks since our last earnings call, preparing us for a massive 2027, 2028 and 2029 as we deliver on the $25 billion of RPO we currently have on our books. With the benefit of that progress, we expect to more than triple our core revenues in '27 and continue to grow at multiples in the years following.
We think of progress in terms of capacity capabilities and customers. We're expanding capacity by adding new contracts for data centers around the world, expanding manufacturing capabilities and collaborating with our vendors to ensure supply and to support our extraordinary growth. We're advancing our capabilities by inventing new technology that extends our performance and throughput and our power efficiency.
And we're expanding our customer base by accelerating AI productivity in existing markets like coating and Agentic flows and pioneering new areas like security, where speed opens up entirely new opportunities.
On the capacity front, data center space continues to be the bottleneck for the entire industry, and we are no exception. The faster we and our customers bring on new data centers, the faster we grow. So over the last 7 months, we've had an all-out push to secure and build out data centers.
We have 2 advantages. First, because we are serving inference, we do not need gigawatt footprint locations like those needed for training clusters. This gives us much more flexibility to scale up capacity across a multitude of locations around the world.
Second, we built a repeatable process for site selection, cluster deployment and customer activation which is an operational muscle required to turn gigawatts into production tokens at a global scale. I'm pleased to report that our push has been very successful. We now have data centers either up or under contract in Alabama, Dallas, Denver, Minneapolis, Santa Clara, Stockton and outside the U.S. in France, Finland, Manitoba, Montreal, Norway, Saskatchewan and Toronto.
In total, over the last 7 months, we have secured more than 600 megawatts of data center capacity that is either live now or will be delivered by the end of 2027. And while this isn't nearly enough to meet our demand, our data center pipeline of new opportunities for expansion continues to grow and is now measured in gigawatts -- to put it in perspective, as we continue to build out our first-party cloud, it will be among the largest non-hyperscale AI cloud.
And whereas at the end of 2025, we're on a steep learning curve. Today, I'm happy to report that we're pretty good at data center buildout with a clear path to becoming excellent. Other key dimensions of capacity include manufacturing and supply chain. Here, we've successfully increased our manufacturing capacity and are building up new factories with Flex and [indiscernible] and expect to increase our manufacturing capacity by more than 10x in 2026 and continue that expansion in 2027.
Again, preparing us for the exceptional growth expected in the years ahead. Our partnership with our supply chain vendors has also turned into a significant advantage. TSMC has once again come through, and we have the wafers needed to fuel our growth. Our ability to get wafer supply also benefits from the fact that we were able to deliver industry-leading performance while running on TSMC's 5-nanometer node, where wafers are less expensive and supply is less constrained.
Our decades-long relationships with our supply partners reinforces our confidence that we can deliver on our growth plans going forward. These relationships are rare and valuable, particularly in times of short supply. Finally, recall that most of the critical supply chain constraints currently faced by the industry don't apply to us.
For example, we don't use HBM memory co-op packaging or require 3-nanometer fab capacity. On the capabilities front, in the second quarter, we delivered support for OpenAI's GPT 56 Soul the largest and most capable of the Frontier models. In fact, Cerebras serves 56 soul at a speed that is 10x faster with GPT 56 souls this lays to rest any of the remaining concerns regarding our ability to support large frontier model.
Being a partner for the delivery of GPT 56 soul and serving it to our cloud speaks to the maturity of our software stack. It takes millions of system hours of production hardening to get to the point where one can deliver hyperscale quality and reliability. We're proud that our inference cloud can meet the requirements of the most demanding customers.
Our collaboration on serving models at the Frontier has opened up new and significant strategic advantage previously only available to NVIDIA. Closed source frontier models include a continual stream of new insights and new AI techniques, serving these models allows us to see into the future and to prepare for it.
Our road map from the hardware through the software stack now reflects what we're seeing and will give us a compounding advantage in the years to come. Continuing on the theme of capabilities, let's turn to disaggregation. We now have disaggregated inference solutions with 2 of the leading chip companies, AMD with their Helios and AWS with Trainium.
This aggregation expands the market for both the GPU provider and for Cerebras, this aggregation enables GPUs to participate in a market currently foreclose to them, namely Fast inference. This aggregation enables Cerebras to expand our opportunity to those customers who are more price sensitive and expands the profitability of our data centers.
Let's see how this works. As with any compute market, as inference grows and matures, opportunities for specialization emerge. This aggregation is a form of specialization that is particularly well suited for workloads with well-known traffic patterns.
In these cases, disaggregation delivers advantage by separating inference into 2 stages. Prefill and decode and using different processors for each stage. Prefill processes the input from the user or agent. It is a paralyzable workload. As a result, prefill is well suited for GPUs and their HBM-based memory architectures.
Decode generates the output tokens. It's the harder technical problem and is the bulk of the computational work in a disaggregated solution. It is sequential and memory bandwidth intensive and is particularly well suited for our wafer scale engine.
The prefill and decode processors need to be linked to create the end-to-end solution. And this is where standards-based IO and open engagement strategy has made integration easy and straightforward for Cerebras. A few weeks ago, we announced a partnership with AMD to build disaggregated inference solutions.
The solutions combine their Helios racks with RCS systems. The combined solution maintains Cerebras speed while increasing throughput by 5x. To understand how powerful this is, it's important to understand the difference between speed and throughput. Speed is a measure per user. It's measured in tokens per second per user.
It is how fast your query is answered or how long it takes an agent to finish a task. Here it is on the x-axis. Throughput on the other hand, is the total number of tokens the solution can produce per second. It is measured by adding up all the tokens across all the simultaneous users. Here is shown as it's generally done on the Y axis.
Speed is critical for user experience. Throughput is critical for inference economics. GPU solutions can support high throughput, but only at low speeds. When configured to support even moderate speeds, GPU throughput drops precipitously.
This is true not just for GPUs, but also for ASICs and all solutions that use HBM. The HBM memory architecture forces a trade-off between throughput and speed SRAM-based architectures like Cerebras' are the exact opposite. We support blisteringly fast tokens, but at moderate throughput. So GPU's to get faster without giving up throughput.
Cerebras wants more throughput without giving up speed. Herein is the strength of our disaggregated solution. It delivers 3 bit speed with 5x higher throughput, increasing throughput by 5x while keeping our industry-leading speed has a profound impact on the economics of token generation.
It means up to 5x as many high-speed, high-value tokens are made by each Cerebras system. More tokens per system at lower cost means more revenue and more gross margin. More tokens generated per CS system also means more tokens per watt, making each data center more profitable. Perhaps most important in a data center constrained environment, the disaggregated solution allows us to serve more of the demand that we have in RPO.
Finally, we believe this disaggregation approach makes performance and economic sense with any GP. For operators who have already deployed large footprints of GPUs, disaggregation with Cerebras offers them an opportunity to create meaningful leverage built on their existing investments by pairing some portion of those GPUs with Cerebras solutions, dramatically improving the value and usefulness of their data center footprint.
Continuing on the capabilities theme, let's turn to our road map. Our engineering execution is continuing at pace. We expect to deliver new systems that double our speed each year for the next several years. Remember, we're doubling our performance, starting with a 15x performance advantage over everyone else in the industry.
In addition, while keeping the performance Crown, over the next 18 months, we plan to deliver solutions that increase throughput by more than 20x. Next week at our annual Super Nova conference, we will be unveiling the CSI, our fourth generation system. It will be a great event with lots of product announcements. So I recommend you attend.
Finally, we are currently on track to launch our CSV in the second half of 2027. Looking even further out, our invention engine is humming. We have significant partnerships with the U.S. government for delivery of stacked memory solutions as well as integrated wafer scale optical solutions. In the years ahead, you can expect to see inventions from us in chip and chip architecture as well as all elements of system design, including packaging, I/O and power delivery.
To summarize the capability section, we expect to continue to deliver pioneering advances in product and technology. to drive up speed and throughput, reduce the power use per token and flesh the cost per token of our solution. Now let's turn to the customer front.
Fast tokens are in demand and command a premium at market and fast tokens with Frontier intelligence are only available through open AI thru Cerebras partnership. Our work with AWS continues, and we expect to have solutions generally available in Q1 2027 through AWS' Bedrock platform. This AWS partnership expands our market opportunity and provides us with global reach through an industry leader who is trusted by nearly every enterprise in the world.
Our discussions with other hyperscalers are also going well. We expect to produce first revenue starting in mid-2027 and ramp through 2028 and beyond. And with all of this progress, I think it's important to keep in mind that our $25 billion in RPO does not reflect any backlog of business from AWS or any other hyperscaler at this time.
Our business outside of OpenAI and the hyperscalers continues to grow nicely. For example, in Q2, we signed 6 deals north of $30 million. AI coding continues its rapid rate of growth. In our experience, no one says I'm happy with flow tokens on coating. So not surprisingly, in the coating category, our footprint continues to grow.
We signed new agreements with public companies such as Sigma and start-up leaders such as Cognition, and we extended our presence in Europe, the major win at loveable. Agentic flows are growing quickly in the value of speed compounds as Agentic operations rapidly evolve toward multistep multi-agent solutions.
Companies as diverse as block, Alphasense and GSK signed new agreements during the second quarter with Cerebras to leverage fast inference to provide their customer-made agents. Fast AI also opens up new markets, extending the TAM for Cerebras. Security is one such example. Our recent win with CrowdStrike is an application that only exists if AI is fast.
Fast AI enables AI and security devices to sit in line with enterprise traffic and use LLM to secure traffic so quickly that nobody notices. Fast AI enables an LLM to provide security that is invisible to users. The AI provides its security.
The speed creates the invisibility that enables the security to avoid delay and disruption. We expect this type of security to become the norm given the rapidly evolving threat landscape. Enterprises will soon expect vast swaths of their traffic to be inspected in this way, creating massive new opportunities made possible exclusively through Fast AI.
Frontier labs, hyperscalers, leading chip makers, the fastest-growing startups and massive enterprises are all now customers and partners of Cerebras and benefit from our blazing fast inference. To summarize, overall, a strong quarter. We went public in a successful IPO. We beat on all metrics, core revenue, core margins and core operating margins, we made progress in each of our key domains, capacity, capability and customers.
These are the foundations on which we'll achieve our goals with massive growth in 2027 and 2028 and continue this exceptional rate of growth in 2019 and beyond. And with that, I'll turn things over to Bob. Bob?
Thank you, Andrew, and good afternoon, everyone. We made tremendous progress in the first half of 2026. As we described, 2026 is the foundation for multiples of growth over the years ahead. We entered the year having won one of the largest technology deals ever, creating RPO of more than $25 billion.
This required us to immediately work on major increases in 3 critical components of capacity. First, we need to increase our wafer supply. As Andrew described, due to our strong relationship and support from TSMC, we did that and are now well positioned not just for the remainder of this year but for the next year as well.
Second, we needed to scale our manufacturing capacity. We are already 4x above where we were in the first half of 2025. And we will have increased the manufacturing capacity more than 10x in 2026. So we're making great progress here.
And third, we need to substantially increase our data center capacity. We've made significant progress with over 600 megawatts of now up or under contract expected for delivery by the end of 2027, plus a pipeline in gigawatts. So the foundation is in place to support a tripling or better in our core revenue in 2027 and additional multiples in future years.
This growth is also setting us up for significant margin expansion in 2027 and beyond. Turning to Q2 financial results. We had another quarter of strong results, beating expectations across each element of our guidance. We delivered record core revenue, we beat on core gross margin and on core operating margin.
I will be using the same core business framework introduced last quarter to describe our progress. The definition of our core business metrics and reconciliations of all of them to GAAP are included in today's earnings release and on our website. Core revenue was $209.9 million up 103% year-over-year. Our private cloud business is growing at an extraordinary pace.
Core cloud and other services revenue was $127.7 million, up 287% year-over-year. This nearly fourfold increase reflects the tremendous demand we have for Cerebras' fast inference service. Core hardware revenue was $82.1 million in the quarter, up 17% compared to last year.
We focus on total core revenue not the mix between the 2, which can vary significantly quarter-to-quarter due to the timing of large new cloud capacity additions and hardware shipments. In Q2, most of the total core revenue was attributable to increases in our core cloud offering, reflecting the ramp in our open AI deployment increases in our other cloud customers usage; and finally, hardware customers who are also wrestling with the timing of new data center capacity.
The demand for fast inference continues to be strong with several late-stage hardware deals, representing hundreds of millions of dollars in the pipeline from new customers as well as significant new cloud deals for 2027. Existing fast inference markets are growing and new ones are getting started. We see disaggregation as an important unlock to drive new use cases since it dramatically improves the economics of inference and up data center ownership.
Today, this means that up to 5x more tokens are produced per CS system, so power and costs are significantly reduced for token. And by continuing to invest heavily in R&D and our product road map, Cerebras will quadruple our current industry-leading speed and increased throughput by more than 20x through the end of 2027, drastically improving our performance and the economics of inference.
Turning now to gross margin. Year-over-year, core gross margins improved substantially. Q2 core gross margin was 40.6%, approximately 940 basis points higher than Q2 '25. This reflects the increase in value of fast inference by the market, our continuous stream of product improvements and additional economies of scale.
Breaking the total core gross margin into its components. Core cloud and other services gross margin was 41.8%, 1,600 basis points better than Q2 '25. Core hardware gross margin was 38.8%, 510 basis points higher than a year ago. As we described last quarter, we are meeting some of the overwhelming demand for our fast inference service by temporarily renting some of our own systems back from our cloud customers and making it available through the Cerebras cloud.
Serving this inference demand sooner strengthens our ability to meet the needs of our cloud customers and to grow with them over time. We believe this will create additional long-term value for Cerebras and its shareholders. In the short term, it reduces gross margin as we have a higher cost for this rented capacity.
As a result, sequentially, core gross margin was 40.6% and versus 46.5% in Q1 '26. Had we not had higher costs due to increasing our private cloud capacity by renting back more of our systems. Core gross margins would have been approximately 500 basis points higher and more similar to last quarter.
Looking forward, we expect Q3 to be the low point for core gross margin before improving significantly in Q4 '26 as we bring on more data centers built with lower-cost Cerebras own system. This will cause core cloud gross margin to step back up. Core gross margin will also continue to improve in 2027 and trend towards our target of 60% plus for several reasons.
The market has recognized that fast tokens are more valuable tokens. This supports higher pricing, which is reflected in hardware and cloud deals that will be recognized over the next several quarters. Over the next few quarters, we will roll off higher-cost rented systems and replace them with lower cost owned systems in our private cloud.
Our product road map has us increasing throughput by 20x over the next 18 months. This reduces the cost to produce tokens per system and per unit of power. As our scale grows, our bill of material costs in our supply chain will improve more. By being on the 5-nanometer node, our wafer costs are lower than others, who need to be on the 3- or 2-nanometer node.
We're also purchasing wafers in much higher volumes. Finally, we do not rely on HBM, which pressures those who use it to either raise prices or lose margin points. We are not exposed to that risk, which we believe will improve our value proposition and pricing flexibility.
Turning to operating margin. Core operating loss was $33.6 million. Core operating margin was negative 16% compared to negative 42% a year ago, an improvement of approximately 2,600 basis points year-over-year. Our ability to deliver this significant improvement in core operating margin, while more than doubling revenues and stepping up our investments in all areas demonstrates the strong operating leverage inherent in our business model.
Today, we are investing in world-class people, manufacturing and data center capacity and company infrastructure to support the significant increase in scale we expect to deliver over the next several years. Remaining performance obligations at June 30, 2026, are $25.4 billion. This backlog provides us visibility to have high confidence in future revenue growth and to invest as needed ahead of it.
Our existing large strategic customers provide validation, contractual visibility, and the economic support required to build new capacity at scale. At the same time, we're having success expanding our addressable market and customer base. OpenAI provides, among other things, scale and Frontier Insight.
AWS provides global enterprise reach Recent collaboration with AMD expands the market opportunity to include disaggregated inference and fast inferences cracking open more new markets like security. We ended Q2 with more than $8.6 billion in cash, cash equivalents, restricted cash and marketable securities. We also have a revolving credit facility of up to $850 million that has been unused to date.
Our liquidity and balance sheet position is strong and was enhanced by our IPO in Q2. It is a significant advantage that provides us with flexibility to invest and adjust opportunities in these very dynamic and high-growth market conditions. In addition, we have the advantage of much lower net capital expenditures per megawatt in the vast majority of AI cloud providers for 2 key reasons.
First, we primarily incur CapEx for the deployment of our own hardware and our data centers at much lower BOM cost. It does not include the high profit margins many others must pay. Second, we are reimbursed for a meaningful portion of the remaining CapEx for data center fit out as data center pass-through cost reimbursement from our largest customer.
Now turning to our outlook. For Q3 2026, we expect core revenue to be in the range of $214 million to $216 million. Core gross margin in the range of 38% to 40% and core operating margin in the range of minus 25% to minus 23%. For the full year 2026, we're raising core revenue to the range of $880 million to $890 million.
We're raising core gross margin to the range of 41% to 43%, and we're raising core operating margin to the range of negative 19% to negative 17%. In closing, Q2 was a very strong quarter of continued execution and growth for Cerebras. We delivered record core revenue cloud and services revenue nearly quadrupled and gross margin and operating margin were also significantly better than our guidance.
We improved our guidance for each of these items for the full year. We've made great progress building our capabilities, capacity and customers and ended the quarter with more than $8.6 billion in cash and cash equivalents and investments to continue to execute our growth plans.
We're well positioned to grow revenue by more than 3x in 2027 and for tremendous additional growth in the following years while also significantly expanding gross and operating margins towards our targets.
I'll now turn this over to Andrew for final thoughts.
Thank you, Bob. More than 10 years ago, we started Cerebras with the belief that we could build a better processor for AI and the belief that to deliver the processor, we would need to build a full accelerator system and racks. Today, Cerebras is 1 of only 4 companies, Google, Amazon, NVIDIA and Cerebras to build processors, systems, data centers and deliver AI-based cloud services to customers.
Thank you for listening to our prepared remarks. And with that, I ask the operator to please open the line for questions.
[Operator Instructions] Our first question comes from Timothy Arcuri with UBS.
2. Question Answer
Andrew, I wanted to ask about customer concentration. So you did say that revenue would be up more than 3x next year. And obviously, we know that Open AI is ramping right now. So that's a big piece of your incremental revenue today. I would think that AWS could be $1 billion next year, something like that, maybe more. So how do you think about customer concentration when you look at next year? Like is it going to be 2/3 of your revenue is like those 2 customers? And then can you also speak to your talks with some of the other folks, Google and Microsoft and folks like that.
Sure. I think it's a good question, and I think some historical perspective might be worthwhile, right? In 2021, people complained that we only had government customers. And then when we won a sovereign cloud at $1 billion, there were concerns we only had a sovereign cloud then we won the largest lab, frontier lab, then there were concerns that we didn't have a hyperscaler. And then we want AWS.
And so I think in each of those cases, we were able to use the momentum that the previous step gave us to expand our business. I think OpenAI is an enormous customer and they're enormous part of not just our business but on everybody's business in the sector. And I think they'll say a big part next year.
But you're absolutely right that AWS and others, whether they're rapidly growing coating companies or some of the use cases around security, they will be a larger portion and open AI will shrink as a percentage of our revenue over time. But I think you're going to expect for next year them to still be a meaningful portion of our revenue.
Great. And then just as a quick follow-up. So I know, Bob, you said that capacity, I think you said it's going up 10x this year, year-over-year. Is there any sense of how much it's going to grow next year? I know that Andrew said that revenue is going to grow 3x, but is there any sense in terms of how much your manufacturing capacity when should grow next year year-over-year?
Yes. We're going to end the year well over 10x our manufacturing capacity. And we already have contracted facilities 3 or 4x more for growth in 2027, and we still have some time to contract for more. So we're looking at enormous growth over a several year period.
Our next question comes from Joshua Buchalter with TD Cowen.
Maybe following up on Tim's previous one. Can you maybe just walk us through how the economics of the Amazon dealer are going to work? Is the plan that it will be available next year and offered in AWS cloud and then we'll basically see how much demand is and so it's difficult to forecast right now.
Sure. It will be available. It is deployed in Amazon data centers. It will be delivered through Amazon's API service bedrock. And we are in the process right now of organizing deployments. So I think that's sort of the way to think about it. We expect the service to be live in Q1.
Got it. Andrew. And then maybe with the AMD engagement, any more color you can give on the go-to-market as you connect with the Helios rack and time line you would expect to revenue. And then regarding the EMD engagement, they made an acquisition of an instrumenting hardware company recently.
Could you maybe speak to how that fits in with what you guys are offering as we think about the broader suite?
Sure. I think a couple of things. I think that we will be announcing additional parts of our arrangement with open -- with AMD over time. But I think the joint solution of Helios racks in front of Cerebras systems, the Helios Rack's doing prefilled cerebral student decode is an extremely strong offering. Right, an offering in which we deliver vastly faster speed than Helios can deliver and vastly more throughput than rivers can deliver a lot.
And that solution is enormously compelling and we have buyers for it already. The second question is of recent acquisition by AMD. Look, I think the company they acquired was interesting and innovative, and no one is more excited about innovative hardware than we are. I think buying hardware start-ups.
There's a lot of time between when you buy them when they deliver. I think we were impressed by what those guys were working on. And think there are many applications in AMD's portfolio for them. I don't see the first application there being data center inference.
Our next question comes from Tom O'Malley with Barclays.
This is Kyle Blostein on for Tom O'Malley. I wanted to go back to Josh's question on the economics with the AMD deal. Is the way this kind of works you buy an AMD Helios Rack install it in your cloud and then all the revenue that comes from customers renting out with disaggregated infra solution goes to you? Or is there some sort of revenue sharing agreement that happens here?
Yes, to the first part of the question.
Okay. And then for my follow-up, the AWS deal is getting installed in their clouds first. Do you see an eventual path to you hosting Trainium and the CSG together in your cloud or another hyperscale cloud. Just trying to think about how segregated inference can evolve in terms of future deployments?
I think we are very interested in that approach. I think that as you've seen with Google, there is a an opportunity for hyperscalers with their own parts to seek to deploy those parts outside of the boundaries of their own data centers. that's something we'd be interested in, not just with AWS but with others.
And so I think that's very much on the table for the future with AWS.
Our next question comes from Quinn Bolton with Needham & Co.
I just wanted to -- just a quick clarification on the AMD deal. Andrew, if you purchase and stand at the Helios racks in your Cerebras Cloud, but that service is delivered to open AI under your contract. Does that represent an additional revenue opportunity? Or how should we think about the potential for revenue in that instance?
And then I've got a follow-up. I think that whenever you increase throughput while keeping your performance the same, you increase your opportunity for revenue, right? Throughput is the number of customers that you can simultaneously support. And if you can do that without giving up speed, you've got more revenue per system.
I don't want to go into the specifics of our relationship with Open AI. But one of the things that makes us so excited about this partnership is that you keep our speed and you increase throughput, which makes each system more profitable. Each system is generating more tokens. That means tokens cost less tokens use less power.
So not only do we make more on top line, but our margins improve. Not only do our margins improve and our top line improvement, but it makes each data center investment more valuable because data center is a power envelope. And if you can get more tokens out of that power envelope, that converts to more dollars.
So it's an enormously powerful thing and with AWS doing this aggregation with us and with AMD doing this aggregation with us. We've tied up about half the leading chip makers. And so it's a very, very powerful story.
And then a follow-on question. You mentioned in the script a couple of times that you'll increase your throughput of the wafer scale engine by a factor of 20 by the end of 2027. Is that remove the need for some of this disaggregated compute or heterogeneous [indiscernible] that you're talking about? Or does that just make the entire throughput of the heterogeneous solutions just that much faster? Or higher throughput.
I think we're exploring all sorts of ways to drive throughput up, right? If your throughput increases 20x and your costs stay the same, you're in pretty darn good shape, right? So we -- our systems are improving throughput. We're looking for ways to improve the throughput of disaggregated solutions.
We're looking at all sorts of different inventions, technologies, partnerships that continue our sort of pattern of industry-leading performance and vastly increasing throughput. That's a really good question. I mean that is what we're thinking about in our road map, that exact point.
Our next question comes from Joe Moore with Morgan Stanley.
On the lines of what you were just talking about, when you talk about disaggregated decode, where are you in terms of commercialization of that? Like do you -- is there -- we know you can do fast inference at scale, you've done it when it comes to disaggregation, is that ready to deploy now?
And what work needs to be done over the next kind of year to get to the types of improvements that you guys are talking about?
Sure. I think -- we have decimated inference with GPUs running in our labs right now. I think it will be deployed and available in Q4.
Okay. And then you talked in your script about the ability to work with the installed base of GPUs. Is there -- when you work closely with Amazon, more closely with AMD. Is that stuff going to work better than kind of what you would be able to do with like NVIDIA installed base GPUs that are out there?
I think it's fair to say, though, we haven't done it yet. I think it's certainly fair to say that Helios racks will give us a bigger and better solution than if we were to use the 355, right? And Trainium 3s will give us better solutions than if we were to use Trainium 2s is -- and if we were to use other GPUs that the current generation, the top of tree generation will give us better performance than the minus the top of 3 minus 1 generation.
But I think it's also fair to say that the minus one generation will be vastly better in a disaggregated solution and not in a disaggregated solution. All right? And in an environment where everyone is trying to extend the life of their hardware and continue to keep it delivering valuable tokens. This is an important option. Does that make sense?
Our next question comes from Vijay Rakesh with Mizuho.
Andrew. Just a couple of quick questions. On the -- you mentioned the 600 megawatts signed capacity in the Power and then 10x increase in capacity by end of the year. Do you think that should help you accelerate some of the ramps in 2027?
Yes. Yes, of course. I think that we are pursuing data center capacity around the world every day. And we're doing it because we have tremendous demand for fast inference and the faster we can deploy the faster revenue grows.
And that's true not just for our cloud business, but it turns out to be true for our customers on-prem business. right, that the faster they can get data centers, the faster we can ship the hardware. And so it is top of mind. It is something I spend an enormous amount of time on.
We have a whole team now -- we're pretty darn good at chasing down data centers around the world. And once you've signed them, your job isn't done as you well know. We have people on site every day. We are engaged with the developer, the construction firms at every stage to do our best to keep them on track.
And so the faster we can do that, I think the faster we can ramp our revenue.
Got it. And then as you look at partnering, I know you mentioned hyperscalers but there's a whole emerging new cloud group that's coming up, getting financing. There's a lot of financing structures being developed across Wall Street, I guess. How is that pipeline developing for you?
Sure. I think early on, the Neo clouds were very focused on NVIDIA. I think as the business has become clearer to investors, there are no clouds that are diversifying and find themselves less dependent on one hardware vendor. And so the opportunities for us in that category are large.
They're neo clouds that are multi-vendor. They're neo clouds that are AMD-only their neo clouds that are coming up out of people who have power assets. And I think in 2027, that will be an important part of our business.
Thank you. I'm showing no further questions at this time. This concludes today's conference call. Thank you for participating. You may now disconnect.
Cerebras Systems — Q1 2026 Earnings Call
1. Management Discussion
Good afternoon, and welcome to Cerebras Systems' First Quarter Fiscal Year 2026 Earnings Conference Call. [Operator Instructions] Please note that today's call is being recorded.
I will now turn the call over to Sean Dorsey, Head of Investor Relations. Please go ahead.
Thank you, operator. Good afternoon, everyone, and welcome to Cerebras Systems' first earnings call as a public company. Earlier today, we issued our press release and posted our supplemental earnings presentation to the Investor Relations section of our website. A replay of this webcast will also be available on our Investor Relations website following the call.
Joining me today are Andrew Feldman, our Co-Founder, Chief Executive and President; and Bob Komin, our Chief Financial Officer.
Before we begin, I would like to remind everyone that today's discussion will include forward-looking statements under the safe harbor of the Private Securities Litigation Reform Act of 1995. These statements include, but are not limited to, statements regarding our future financial performance, business strategy, market opportunity, customer demand, product road map, technology leadership, supply chain, operating model and outlook for Q2 and full year 2026.
Forward-looking statements are based on current expectations and assumptions and are subject to risks and uncertainties that could cause actual results to differ materially from those expressed or implied. These risks are described in our SEC filings, including our final prospectus related to our IPO and our future periodic filings with the SEC. We undertake no obligation to update these forward-looking statements, except as required by law.
During today's call, we will also discuss certain non-GAAP financial measures. Reconciliations between GAAP and non-GAAP results are included in today's press release and supplemental materials, which are available on the Investor Relations page of our website.
With that, I'll turn the call over to Andrew.
Thank you, Sean, and thank you, everyone, for joining us today. This has been an extraordinary several months, and I want to begin by thanking our customers, our partners, suppliers, employees and shareholders. We would not be here without your trust and your support.
Earlier today, we posted our Q1 2026 results, and we delivered a strong quarter. We delivered core revenue of $191.3 million, up 92% year-over-year. Core hardware revenue contributed $111.6 million, up 60% year-over-year, while core cloud and services revenue contributed $79.8 million, up 167% year-over-year. Bob will share more color on our financial results shortly.
Before Bob digs into that, I'd like to say a few things about the market. I'll divide my comments into several sections. I'll begin by spending a few minutes sharing my views on the larger drivers underpinning the AI revolution, their impact on the compute market and why speed wins. I'll then turn to our successes in Q1 with special attention to our progress with OpenAI and AWS. And finally, I will talk about how we expect to avoid many of the supply chain challenges that bedevil others in our space.
To understand the dynamics in the compute market, it's important to realize that AI provides new capabilities to computers. AI gives computers purchase on whole swaths of the world that had previously been foreclosed. This is why AI is so transformative and why its impact is so profound and why we believe it increases the size of the market addressable to compute by many thousands of times.
Computers have historically been good at math, very good, but they were relatively poor at everything else. They did not provide much insight into text or images. For these modalities, all they could do is store and retrieve. Computers were at their best in a 2D world of numbers. In a real world of 3 dimensions, they were challenged. AI opens up the world of human experience to computers. As a result, the size of the market increases exponentially.
It is as if prior to AI, computers worked in black and white and in 2 dimensions and after AI, they address a world of color in many dimensions. This is why AI has spurred an explosion in the demand for compute. Computers can now do things they have never done before and why, in our opinion, demand will continue to accelerate for many years to come. Text, images, video, agents, robotics, these are all part of how AI expands the computer's ability to understand, participate and take actions in the world. These all represent opportunities for Cerebras.
Let's look at the specifics of how this is unfolding. Prior to 2025, AI was a parlor trick, a novelty, interesting, but not useful, cool, but not valuable. AI is now valuable because it has become profoundly useful. Led by OpenAI, the foundation model providers pioneered the way, the foundation model makers and shortly thereafter, the open source models made models smart enough to be useful across many domains. And once something is useful, people use it. And once people start using the technology, speed determines its productivity.
Fast is productive and slow is unproductive. Speed provides answers in less time, providing competitive advantage. Speed makes the largest and smartest frontier models interactive. Speed enables agents to complete tasks faster. Fast tokens are the most valuable tokens because they get more work done in less time. And today, Cerebras delivers the fastest AI in the world, bar none, not by a little bit, but by an order of magnitude. And we do this for small models, for medium models and for the largest models in the industry. We do this for models with small KV cache, with medium KV cache and with giant KV caches.
We generate tokens faster than anyone else. What I'd like to show you right now is a quick demo. Just how much faster we are than GPUs on [ Kimi K2 ], a trillion parameter open source model. We're going to run the exact same prompts. On the left, it's Cerebras. On the right is a leading GPU. The only difference, same model, we're finished already. Same model, right, same prompt, the difference is hardware, and we're finished. It took us 21 seconds. We're now waiting on the GPU. Still waiting.
Now we've increased the speed 5x in the video to not make you wait as long as you otherwise would. Still waiting. Okay. Cerebras did in 21 seconds. It took 4 minutes and 37 seconds for the GPU to do. The same model, the same prompt. That's what it means to be 13x faster. In AI, inference speed is productivity. Slow isn't productive. But this should not come as a surprise. It is in line with each of our everyday experience.
How big is the market for slow search? How big is the market for slow Internet access? Any of you still use dial-up? How long will you wait for a website to resolve? Why would it be different for AI? In fact, not only does speed increase the value of tokens, but speed accelerates the adoption of AI. When AI is fast, it's more fun to use. People use it. They use it more often for more things, and they use it to solve more important problems. With fast AI, users invent things that never existed before. They solve problems in new ways. They develop new offerings, new business models. This is what speed does. And this is what Cerebras' speed enables.
A final point on speed. There recently has been a great deal of focus, especially at the frontier model level on safety and the importance of guardrails. How do guardrails work? Guardrails add a layer of compute on top of the AI to create a safer experience. This compute takes time, and it takes more time on slow infrastructure. Traditionally, guardrails force the trade-off between safety and user experience, between safe and fast. Cerebras eliminates this trade-off. Fast AI inference allows guardrails to work without inserting crippling delays. AI is safer with these guardrails and AI is safer and more productive when it's [ extremely ] fast.
Our performance advantage is borne of our wafer-scale architecture. We're more than an order of magnitude faster than GPUs because we solve problems that haven't been solved or couldn't be solved by others. The problems of yield, cross-reticle connectivity, mismatches in thermal expansion, power delivery and cooling are all problems that the industry struggles with, but the Cerebras solved years ago. Moreover, the advantages of wafer-scale are durable. By building chips that are 58x larger than the largest competitor, we're able to use SRAM and benefits from its blistering speed, while competitive offerings use HBM, which is slow, expensive and in short supply.
We see the advantage of wafer-scale technology expanding our performance lead as we bring next-generation solutions to market. In fact, the technology underpinning of wafer-scale fundamentally advantages additional technologies in the future. For example, wafer-scale technology brings profound advantage to memory stacking and optical integration. And as we look further into the future, data centers in space are also advantaged by wafer-scale integration. Not only does wafer-scale compute deliver faster speeds and for latency-sensitive workloads, less power per unit compute than do GPUs. But most importantly, it requires less chip-to-chip communication. And chip-to-chip communication is one of the fundamental limitations of terrestrial data centers and a yet to be solved problem for data centers in space.
So with this as a backdrop, in the first quarter of 2026, how did we meet this extraordinary market? And how do we leave Q1 even better positioned? In this section, I'll focus on our partnership with OpenAI and AWS as they took shape in this quarter. We signed a definitive agreement with OpenAI on December 24, 2025, for the purchase of more than $20 billion of Cerebras compute over the next several years. By February 1, we were in production, running a model we've never before seen, 35 days from signature to production deployment.
Beyond the transformative revenue ramifications, our collaboration with OpenAI gives us a direct view into frontier model development and the direction it is moving. By pairing frontier model intelligence with the world's fastest inference, we build products and technologies that others simply can't. In fact, the boundaries of these capabilities have yet to be fully explored. OpenAI and Cerebras are excited that GPT 5.4 is now running on Cerebras. This collaboration brings together OpenAI's frontier models with Cerebras' wafer-scale inference infrastructure to enable highly responsive model interactions. GPT 5.4 on Cerebras is currently available to OpenAI engineers and to select OpenAI customers as part of OpenAI's strategic rollout. OpenAI and Cerebras are also actively working to bring GPT 5.5 onto Cerebras as part of the next phase of this rollout and expect to share more shortly.
In March, continuing this trend, we signed a binding term sheet with AWS to deploy Cerebras and AWS data centers. Our solutions will combine AWS' leading Trainium 3 chips with Cerebras' CS-3 in a disaggregated solution that is expected to be an order of magnitude faster. Trainium will do prefill and Cerebras will be decode. And together, the solution is expected to deliver the fastest tokens at massive throughput.
Remember, disaggregated solutions are a significant opportunity for Cerebras. The technical strategy is one of divide and conquer. It is based on the recognition that inference has 2 computational components. The first is where we process the prompt. This is called prefill and it's highly parallelizable. The second is where we generate the response. This is called decode and it's strictly sequential. By using different processors for the prefill and for the decode, we can deliver truly exceptional results.
We are also proud to announce that we have, as of this week, completed a definitive agreement with AWS and we will begin our technical collaboration as well as prepare for deployments in their data centers. As you all know, AWS is a leading cloud compute company and one of the most important providers in the world for developers and enterprises. And many enterprises want to run AI where they store their data and where they have existing agreements and where the environment is familiar and is secure. As a result, AWS provides an easy way for Cerebras solutions to meet the world's enterprises where they already are.
Let's for a minute now turn to supply chain. Keeping up with this extraordinary market growth has brought supply chain challenges to many in our industry. At Cerebras, we have several fundamental advantages. First, the binding constraint in the market right now is HBM memory. It's in short supply, it's expensive, and we don't use it. So we avoid this constraint entirely. We use SRAM. And SRAM is printed on our logic wafer. It's not a separate chip. As long as you can make the chip, you can make SRAM. Its supply is approximately infinite.
The second binding constraint is the CoWoS process at TSMC. We don't use it. So again, we sidestep this constraint. Third, 3-nanometer capacity at TSMC is a constraint. And again, we don't use it. We're the fastest in the world and happily at the 5-nanometer node, where there is less contention for fab resources and where manufacturing is less expensive.
Our partnership with TSMC deserves special mention as they know more about chip making than just about anyone else on earth. They believed in the wafer-scale approach from the time we were a tiny team with nothing but a PowerPoint slide, and they've been with us along the way. They have proven themselves to be an extraordinary partner. Just as a reminder, our salable unit is not our wafer, but our CS-3 system. We sell the CS-3 for on-premise deployments or time on the CS-3 through our Cerebras cloud or through our partners' cloud. We manufacture our CS-3s in the U.S. And in fact, to the best of my knowledge, we are the only accelerator maker to manufacture exclusively in the U.S.
We have added hundreds of thousands of square feet of manufacturing and clean room space to support our growth. We've expanded our partnership with Flextronics and are proud to have added Sanmina as our second major contract manufacturer to assist us in managing our expansion.
Finally, it's no secret that data center capacity is at a premium. It's a dog fight out there. Despite this, we've added data centers around the world. We've added data centers across the U.S. and Canada, Europe, including France and the Nordics, and we're in early discussions for data centers in Israel, the UAE, Australia, Singapore, India and Indonesia. We're expanding the capacity. We need to serve customers, and we're doing it with urgency. The demand environment is strong, but this is just not -- this is not just about demand. It's about building the infrastructure required for the next phase of AI.
So to wrap up, there is a tectonic shift in compute demand brought about by AI's ability to make the world around us tractable for computers. As a result, the market will need vastly more compute, in my view, for decades. AI power users represent today a tiny fraction of the world's population, by some estimates, less than 1% and compute and memory is already in tight supply. Just imagine. To this AI revolution, we bring leadership technology, which in turn enables us to deliver the fastest AI inference in the world by more than an order of magnitude. Fast tokens are more valuable tokens and Cerebras tokens are the fastest. The result was a record quarter.
With that, I'll turn things over to Bob, and he can provide more color on the financial results. Bob?
Thank you, Andrew, and good afternoon, everyone. I want to also add my thanks to our customers, partners, team Cerebras and the investment community, both new and who have gotten to know us over the last several years. Cerebras is more than 10 years into the journey, and we're still just at the very beginning. I want to thank everyone for joining us today on our first earnings call operating as a public company. Opportunities we see ahead for us with fast AI are massive, and we appreciate everyone who has chosen to join us for the road ahead.
Today, I want to describe the financial framework we will use to discuss our results. It's the same way that we evaluate our financial performance and make resource allocation decisions internally, provides additional visibility to amounts that are embedded in our reported GAAP revenue and cost of revenue that we believe provide more transparency as well as direct comparability to our prior historical results to better analyze our trends.
Beginning in Q1 '26, we have data center costs, which are contract with OpenAI has us pass through to them with a 3% markup. These data center pass-through items are reported gross, so they increase both our cloud and other services revenue and cost of services, but are at a significantly lower margin than the rest of our business. These amounts start out small in Q1, but they'll become more significant over time. Also, OpenAI has the option to choose whether to receive its future committed amounts in our cloud or in its own data centers, which would mean there would be no future corresponding pass-through amounts for that capacity. Because these amounts can be highly variable and are outside of our control, we're excluding them from our core business metrics.
We also now have noncash amortization of customer warrants that is recorded as a reduction in revenue for both our hardware and cloud and other services GAAP revenue line items, depending on the related services the customer is purchasing. So we're adjusting our GAAP numbers to exclude the impact of these items and a few other common ones like stock-based compensation and onetime items, and we define the resulting non-GAAP amounts as our core business metrics. I will only be discussing these core metrics today. Reconciliations to GAAP for all of our non-GAAP items are available in today's earnings material and on our website.
Let's start with revenues. Q1 was another record quarter for Cerebras. Our core total revenue was $191.3 million, representing 92% year-over-year growth. Now looking at revenue by type. Core cloud and other services revenue reached $79.8 million and grew 167% year-over-year. Market demand for Cerebras Inference Cloud remains incredibly strong. We are ramping our capacity rapidly, and we saw a meaningful pickup in revenue across Q1 as we began our ramp with OpenAI in February as well as from other customers using the Cerebras Cloud.
We expect increasing year-over-year growth rates for each quarter in 2026 with more of this revenue coming later in the year as the ramp in our cloud capacity deployments accelerates. Core hardware revenue was $111.6 million, up 60% year-over-year. We plan to see decreasing hardware revenue for the next few quarters as our existing POs are delivered and our mix shifts towards the majority of our hardware production being deployed in Cerebras Cloud to fulfill our significant contracts. This trend could change relatively quickly, however, as OpenAI and AWS as well as other customers make decisions about when and how they prefer to deploy our hardware solutions in our data centers or theirs.
Now moving on to gross margin. Core gross margin was 46.5% in the quarter compared to 42.1% in the prior year period and 41% last quarter. Core cloud and services margin improved significantly to 52.9% in the quarter from lower levels we saw last year as we launched the Cerebras Cloud service. The primary reasons for the increase were higher pricing as the market is now valuing higher speed inference at a premium and market demand exceeds supply. The utilization of our systems that we began to deploy in late 2025 improved quickly. And there was a small amount of [ ramp back ], relatively speaking, to increase capacity from a customer.
For the rest of 2026, in order to accelerate our ability to service the significant near-term demand in our contracted backlog, we've chosen to make more capacity available sooner by temporarily renting our own systems back from an existing customer while we aggressively build out and deploy our own data center capacity. The additional cost of renting third-party capacity will depress core cloud and other services margin temporarily from current levels. We expect the impact to be a decrease of 10 to 15 margin points based on the volumes we are now anticipating before beginning to [ ramp back ] towards our target margin of 60% plus as we transition away from our rented systems.
Core hardware margin was 42% compared to 30.6% in Q1 '25. Over the last few quarters, we benefited from the timing of incremental performance-based incentive pricing after the target was achieved, but was recognized prospectively for the remaining systems that have not yet been shipped. We expect core hardware margin to be more similar to the first half of 2025 and return to the low 30s as this contract pricing normalizes.
As a reminder, when we sell hardware systems and recognize that revenue upfront, we also include support and other services, which have significantly higher margins. As a result, total profitability over the life of the individual contracts is much closer to our target overall gross margin. These additional elements of revenue are required to be recognized over the contracted life of the services and are recorded as core cloud and other services, so are not included in our core hardware revenue and gross margin.
We are focused on improving gross margin over time through scale economies, improved product throughput and performance, manufacturing efficiency, utilization of cloud capacity and performance-driven pricing improvements to achieve our long-term overall gross margin target of 60%. At the same time, we will continue to be aggressive and creative, including potentially investing ahead of demand when we see attractive long-term opportunities to gain key customers, accelerate revenues and drive gains in market share.
Now I'm going to talk about operating expenses. Our non-GAAP operating expenses were $92.6 million, up 51% from a year ago at just more than half the rate of core revenue growth of 92%, demonstrating the strong operating leverage available as we grow our business. R&D was our largest area of investment at $69.8 million. We believe sustained R&D investment is essential to maintaining our technology leadership and requires being at the frontier of AI across silicon, systems, software, models and cloud infrastructure to deliver the fastest performance.
We have an exciting product road map to bring to market over the next several years, including near-term innovations such as the implementation of disaggregated inference solutions, with multiple hardware partners, which we expect to begin to deliver in the second half of this year. Sales and marketing expense was $12.9 million, reflecting continued investment in customer engagement, field capacity, developer adoption and go-to-market infrastructure to support increasing market demand. G&A expense was $9.9 million and will continue to step up significantly next quarter due to incremental costs associated with operating as a public company and rapid growth in the size of the business.
Moving on to profitability. Core non-GAAP operating loss improved to near breakeven at minus $3.5 million with operating margin of negative 2%, a significant improvement from a year ago when core operating loss was minus $19.3 million and operating margin was negative 19%. There was also a nice improvement sequentially from Q4 '25 when operating margin was minus 10%.
Core non-GAAP net loss was $2.5 million, while the temporary reduction in gross margin I described earlier that will result from renting back our systems until we deploy significant capacity in our own data centers will cause these metrics to regress somewhat for the next few quarters. We believe the steady improvement that we delivered over the past several quarters highlights our ability to achieve our target profitability profile of approximately 60% gross margin and 40% operating margin in the medium to long term.
Moving on to our current cash position. We ended the quarter with $3.3 billion in cash, cash equivalents, restricted cash and marketable securities. We've accelerated the pace of our fundraising over the last several quarters to support our increasing growth rate and provide us with the liquidity we need to scale. As a reminder, we raised $1 billion in Series G equity in September 2025, another $1 billion in Series H equity in February 2026, added a revolving credit facility for up to $850 million in April '26. And then just a few weeks ago, completed the largest semiconductor IPO in history, raising another $6.4 billion. We are well positioned with the financial flexibility to accelerate the sourcing and deployment of data centers and our supply chain to support significant near-term growth of our cloud business.
Now turning to our outlook. We'll typically provide quarterly guidance, but since this is our first earnings call, we'll also provide some color on the year. In our core business in Q2, we expect core revenue of approximately $194 million, representing year-over-year growth of 88%. Core gross margin in the range of 36% to 38%; core operating margin in the range of minus 30% to minus 32%. And for the full year 2026, we currently project core revenue in the range of $855 million to $865 million, representing year-over-year growth of 69% at the midpoint. Core gross margin in the range of 38% to 41% and core operating margin in the range of minus 28% to minus 32%.
In summary, we made significant progress in our business during the first quarter. We delivered strong revenue growth, gross margin improvement and meaningful customer momentum. We significantly strengthened our balance sheet through our IPO and our fundraising activities, and we're poised to continue executing on the enormous amount of opportunity we see. We're working hard to bring more data center capacity online as soon as possible to meet robust demand.
With that, I'll turn the call back to Andrew for closing remarks. Andrew?
Thank you, Bob. Cerebras was founded on the belief that AI infrastructure needed a new approach, one that was built from a clean sheet. The progress we report today reinforce this belief, the world needs faster AI. Faster AI like faster versions of all technologies before it drive adoption, usage and customer experience. When given the choice, who wants slow. And we're built to deliver fast AI. That's what we do.
As AI continues to expand its footprint, so will we. We're proud to be a public company, and we're redoubling our effort on the work ahead. We continue to fuel our culture with fearless engineering and with the ability to delight our customers with experiences that are unavailable elsewhere. We also will work diligently to communicate with our stakeholders and our investors and to do so with transparency and with discipline.
We thank you for joining us today. Operator, please open the line for questions.
[Operator Instructions] Our first question comes from the line of Timothy Arcuri of UBS.
2. Question Answer
Andrew, now that you have the definitive agreement with AWS, can you just sort of help us to think about the timing on this and your ability to supply that customer? I know you had to put in your wafer orders back in February. So can you just give us a little bit of help in terms of when you could start to ship to them?
Sure. I think TSMC has been extremely good to us. We are in the happy position of having supply for our plan and beyond in 2026. I think you should expect to see AWS' impact in 2027.
Got it. And then if I can ask a quick follow-up. I also heard, Andrew, you talked about multiple partners for disaggregated solutions. Does this imply that there's another customer beyond AWS? And I guess I asked because I did see that Cerebras had a presence at Microsoft Build. So I'm just wondering what you mean by the multiple partners.
I think the opportunity to provide decode for people who have GPUs is real and in front of us. I think that's exciting. I think that the GPU as an architecture struggles with the sequential nature of decode, and we are extraordinary at it. So it makes sense to explore partnerships on that vector.
Our next question comes from the line of Tom O'Malley of Barclays.
Congrats on the nice results. Andrew, I wanted to ask you a question on your TAM. I think that during the process, there was a lot of conversation about your ability to handle larger models. When you look at Kimi, that's one example of a large model. You're again showing a demonstration today about attacking larger models as well. Jensen spent time talking about 25% of the inferencing market is fast inferencing and maybe even took a step back on that on the last call. But what do you think your TAM is when you look at the broader AI market? Would love to get your opinion there.
Thanks for the question. We look out into technologies and can't find examples of where slow has owned meaningful portions of the market over medium periods of time. And I think you should think very carefully about the example of search, right? There is no slow search because nobody wants it, right? There's no more dial-up because nobody wants it. And I think when given the choice on the same model between fast and slow, I don't think it's a very hard decision. And so when we look out at the space, we see the entire inference market as available to us for fast inference.
I mean who doesn't want answers in less time? And who doesn't want more productive agents? So that's what we see. I know that's at odds with GPU makers. And both of our arguments are, I think, in some way self-interested. We build fast and I think the market is big for fast. So I'm not surprised at that.
Super helpful. And then we might find this out in the filings, but just wanted to give it a crack on the call. Did you have any top 10% customers? And are you willing to share on the call how large they were?
I don't think we should share on the call. I think you'll see in the filings.
Our next question comes from the line of Quinn Bolton of Needham & Company.
Andrew, Bob, congratulations on your first call as a public company.
Andrew, I wanted to follow up on the inference TAM question. Just obviously you guys are addressing the fast inference portion of the market, which you think allows you to address the entire market. But your tokens may be more expensive. And so I was just wondering if you could address the higher token cost for fast inference. How much of the market do you think is willing to pay a premium for fast inference? And then I've got a follow-up on the road map.
I think there -- today, in many instances, fast is priced at a premium. I think you saw Anthropic offer a service. In fact, most now offer services in which fast tokens are sold at a premium. I think they're sold at a premium because they're more valuable, right? And I think you can look to your own experience with your Internet provider. If dial-up were free, do you want it? I think the answer there is quite the contrary. You have to pay quite a bit of money to get someone to take dial-up. And so I think that the reason right now that there's a premium is because people prefer fast. It's more valuable. I think we'll see over time how that shapes out.
Got it. And then the question just with the AWS definitive agreement now signed, if you look across the compute spectrum, oftentimes, these AI compute deals can extend into the gigawatt range. Just wondering, can you give us any sense of the scale? Is this tens of megawatts, hundreds of megawatts? Could it reach a gigawatt? Just any sense on the size of the AWS partnership and definitive agreement?
I don't think we're going to -- we're sharing that at this time.
Our next question comes from the line of Atif Malik of Citi.
Congratulations on the debut. Andrew, on the OpenAI and AWS partnerships, what is the decision tree for them to take the future commitments in cloud or as hardware and data centers?
So first, greeting Atif, good to hear from you, I guess. Second, with AWS, they are deployed in AWS data centers. That's the deal. I think OpenAI has a choice. They can deploy it in their data centers in a model where they buy the hardware or they can receive the compute via cloud service. I think it will depend on OpenAI sort of a portfolio decision of their data centers and their various capacity versus what we can bring in data centers. I think that's likely to be the determining factor, but I think that's really an important question for them.
Got it. And Bob, as a follow-up, I mean Andrew talked about the dog fight in terms of data centers and power availability and whatnot. When you look at your full year outlook, and thank you for providing that on this call, how much of that are new data centers or new power shells versus renting back from your existing G42 customer or your Cerebras Cloud?
This is Andrew, Atif. We're trying to add data center space as fast as we can. I mean we are engaged with builders throughout North America, data center operators in Europe, in the Middle East. We have new data centers coming on board in Q3, Q4, Q1, Q2, Q3, Q4 of next year and are adding more. We're in discussions with literally dozens of different data center owner operators. And so I think the answer is all of the above. We are going to -- the demand for our product right now is so significant. We are seeking data center capacity around the world as quickly as we can.
Our next question comes from the line of Joe Moore of Morgan Stanley.
On the same lines as the last question, is the constraint on your growth 5-nanometer wafer capacity? Is it space and power and the kind of build-out of your cloud? Or are there some other constraints that we should be thinking of? It feels like demand is not the constraint here. It's how quickly you can ramp.
Demand is not the constraint. Supply is not the constraint. The constraint is data centers.
Okay. That's helpful. And to the extent that your gross margins are better than we had modeled, is that a function of sort of a quicker ramp of that internal capacity versus the G42 rental? Or just what are the dynamics of gross margin through the rest of this year?
Thanks, Joe. There's a few things going on. One is actually higher pricing. So we -- because there's tremendous demand, we've been able to see higher pricing from existing customers. So even as OpenAI is starting to ramp, that's been an upside to our gross margin profile and something that we're reflecting now in the outlook for the rest of the year.
Another way to think about it is the competition has also increased in price. They have higher cost for HBM and other things. So I think the floor in the marketplace has come up a bit. And then we've been able to look at the timing of the amount of capacity that we need to bring on and the economics around it, which we were estimating a couple of quarters ago. And that's also turned out to be a bit more favorable, both in terms of how much is coming on when and also the amount that we're paying. So I think all of those factors as they play out for the rest of the year will allow us to be at higher gross margins than what we had predicted at the beginning.
Our next question comes from the line of Joshua Buchalter of TD Cowen.
Welcome to the fun world of earnings calls. Maybe -- sorry to keep pulling at this thread, I wanted to follow up on sort of Tom and Quinn's earlier questions about the ability to service some of these -- the larger models. Maybe using the demo that you guys provided of the supporting the trillion parameter Kimi model, like any details you can give on the specs that were in that benchmark you showed, like how many CS-3s were used to support Kimi and maybe what the competing GPU-based rack architecture was?
We used the leading -- by way of comparison, we used a leading inference cloud. So we try to do our best to compare top of tree to top of tree. My understanding is that they're using B300s to serve as an endpoint for this model, but I can double check that for you. I think there is a fundamental misunderstanding propagated by some analysts who just didn't understand that our architecture was perfectly suited for these models of large size, small size, medium with big caches, small caches and that we can do them and are doing them not just in this demo, but for OpenAI at frontier models. right?
There are only 2 hardware vendors that currently serve OpenAI models, and we're one of them. And so it is sort of a proof point, right, an empirical validation that big models work just fine on us, and we have the same advantage as small models.
Okay. Understood. And then maybe for Bob, as we think about the annual guide you gave, I think it implies sort of 20% plus half-over-half growth. Any help you can give us on how much of the second half growth is from pricing or maybe OpenAI contribution that we should expect for that first -- as you build up to the first 250-megawatt build?
Yes. Look, I think this initial guide coming out in the -- which is really focused on the first quarter and looking forward for the rest of the year, where we have data centers coming on largely in the back end of the year. A lot of the improvement is going to come from OpenAI being deployed in our cloud, and it's back-end loaded.
As I mentioned in my remarks at the beginning, we actually have in the forecast that hardware will come down a little bit sequentially for the rest of the year. So I'm being conservative for the second half as we're still pretty early in the year, data center capacity is coming on. And as we move throughout the year, we'll update you as we have more information about the progress and timing.
Our next question comes from the line of Matt Bryson of Wedbush Securities.
Just going back to trying to figure out the market, it sounds like there's some more opportunity for what we're seeing with Amazon, where they're using Cerebras solutions to decode. We're thinking about the amount of value that you're capturing in that type of architecture versus [ prefetch ]. Is there any chance you could take a swag at kind of what portion of the value is in the Cerebras system?
Not exactly. Let me share maybe a different crack at the problem. A decode prefill, a disaggregated solution, is really good in some instances. And in particular, if you know the shape of the work, it's intended to support. When you specialize, right, when you buy some hardware for prefill and some for decode, you embed in your hardware deployment an assumption about the shape of the traffic. And if the traffic looks different, then you have stranded compute and low utilization and higher cost.
This is obviously a huge opportunity for a hyperscaler like AWS because they have technology that can drive traffic, right, of the shape they expected to their disaggregated solution and route it to other solutions if it's different from that assumption. Right? So the value of the solution is highest to a hyperscaler. The exact split of value between us and Trainium is very difficult to say. And as nobody has yet has deployed a true disaggregated solution, we have a lot to learn in the market still.
Understood. That's helpful. And then just one for you, Bob. We're thinking about you renting out capacity from a customer to fill that OpenAI demand. Is the full rental requirement baked into your quarterly guide? And -- or is there any chance that there's a further impact on gross margins in Q3? Basically, I'm trying to figure out if gross margins in Q2 are trough.
So the rental costs that we're assuming for the rest of the year are baked into Q2 and the annual guide.
Our next question comes from the line of Vijay Rakesh of Mizuho.
Congratulations on a good quarter and guide. Just wondering, you mentioned 50 megawatts per month ramp into 4Q '26. I'm just wondering how that is going? And how do you see that scaling into 2027? And I have a quick follow-up.
I don't think I mentioned that. I'm -- maybe I didn't hear the question right. Could you repeat the question?
I think you had talked about 50 -- I believe you had talked about a 50 megawatts per month ramp into 4Q '26. And then just wondering how that is going and how you see that beyond -- how that capacity ramping into '27?
Yes. Okay. I don't remember giving specifics on the monthly ramp. We are seeking, on average, a huge amount of capacity in through the end of '26 and into '27. As you know, we signed our agreement with OpenAI at the end of '25, which means you probably need 6 or 8 or 10 months at a minimum to bring on vastly more capacity. And as our business ramps, we are signing large deals as well, many of which will come on in the first part of 2027.
I think we announced a 120-megawatt deal with Bell Canada, for example, in a facility there that does have room to expand. So I think the -- while we haven't given specifics, we are working our hardest to add as much capacity as we can between now and the end of '27.
Got it. And then obviously, you mentioned fast inference is very disruptive. You see a lot of LLM frontier model guys try to move to fast inferencing. Just wondering on how you see your customer pipeline broadening out into '27 if you were to look out beyond OpenAI and AWS?
Sure. Look, we're pleased with the way the customer pipeline is going. I think, obviously, deals of the size of OpenAI or the size that AWS could do are few and far between. But the business is robust, and we're happy at the rate at which we're signing new customers. We're also happy at the rate at which existing customers are doubling down, growing their footprint and the rate at which sort of their token consumption is up and to the right. And so on all fronts, we're pretty pleased.
Our next question comes from the line of Richard Shannon of Craig-Hallum Capital Group.
Congrats on the first quarter call here. Andrew, my first question is following up on one of your -- a couple of your prepared remarks regarding OpenAI. You talked about stepping up a new model under 35 days here. Then you also mentioned about doing some work with GPT 5.4. I'd love to hear about your experience in bringing up the [ serving ] model, the Codex-Spark, and what you've learned from that and how you apply that to working with the [ GPT 5.4 ] that you might see going forward with OpenAI and/or other customers?
I think foundation model providers are fundamentally different. They are at the absolute cutting edge. What you see when you engage with them is really quite extraordinary. And the amount of work that goes into a foundation model and the visibility that we have is really one of the exceptional advantages that we get from this partnership.
So I think beginning with Spark, we got better. I think it improved us. It challenged us. We were up to the task. We very much enjoy working with their engineering team. And I think from the feedback we've gotten, they found a kindred spirit and enjoy working with our team as well. And so I think the way to temper metal is with fire. And I think we're proud of our work with them and our continued work. And so I think it's a really thoughtful question. I think having access to extraordinary customers and partners is a fundamental and long-term differentiator.
Andrew, my follow-on question is regarding AWS. There are media reports out there that Amazon may be trying to sell the Trainium-based hardware externally, not just in their own data centers. Do you view this as an opportunity for Cerebras?
I do.
Thank you. And with that, I think we'll wrap up.
Yes, sir. We have reached the end of the Q&A session, and that does conclude today's conference call. Thank you for participating. You may now disconnect.
Financial data from Cerebras Systems
Revenue
Revenue is the sum of all sales generated by a company, e.g. for its products or services.
Revenue (TTM) metric explainedDirect Costs
Direct costs are the costs incurred directly in connection with the manufacture of the product or service.
Gross Profit
Gross Profit indicates how much of the revenue remains in the company after deducting direct production costs. If the percentage share of sales is calculated, this is referred to as the gross margin.
Gross Profit metric explainedSelling and Administrative Expenses
Selling, general and administrative expenses (SG&A) include all expenses for marketing and sales as well as the general administration of the company.
Research and Development Expense
Research and development costs (R&D) provide information on how much the company invests in the research and development of its products. The costs are particularly interesting as a percentage of revenue and in comparison to direct competitors.
EBITDA
EBITDA (Earnings Before Interest, Taxes, Depreciation and Amortization) is the company's earnings before interest, taxes, depreciation and amortization. The EBITDA margin is calculated as a percentage of sales.
Depreciation and Amortization
Depreciation represents reductions in the value of the company's assets (e.g. due to wear and tear on machinery).
EBIT (Operating Income)
EBIT (Earnings Before Interest and Taxes) is the company's profit before interest and taxes, also known as the operating income. The EBIT Margin is calculated as a percentage of sales at
.
Net Profit
Net Profit represents the profit or loss after deduction of all costs.
Net Profit metric explainedStocksGuide Premium
| Jun '26 |
+/-
%
|
||
| Revenue | 180 180 |
-
100%
|
|
| - Direct Costs | 155 155 |
-
86%
|
|
| Gross Profit | 26 26 |
-
14%
|
|
| - Selling and Administrative Expenses | 183 183 |
-
101%
|
|
| - Research and Development Expense | 320 320 |
-
178%
|
|
| EBITDA | -453 -453 |
-
-251%
|
|
| - Depreciation and Amortization | 24 24 |
-
14%
|
|
| EBIT (Operating Income) EBIT | -477 -477 |
-
-265%
|
|
| Net Profit | -451 -451 |
-
-250%
|
|
In millions USD.
Don't miss a Thing! We will send you all news about Cerebras Systems directly to your mailbox free of charge.
If you wish, we will send you an e-mail every morning with news on stocks of your portfolios.
Cerebras Systems Stock News
Company Profile
Cerebras Systems, Inc. engages in the designing and provision of processors for artificial intelligence (AI) training and inference. The company is headquartered in Sunnyvale, California and currently employs 708 full-time employees. The company went IPO on 2026-05-14. The firm's products include inference Cloud, Training Cloud, CS-3 system, AI supercomputer, Wafer Scale Engine and model development. The firm's pioneering Wafer-Scale Engine (WSE), a chip encompassing an entire silicon wafer, was specifically designed to enable higher performance and speeds than GPUs for the computational demands of inference, Generative AI (GenAI), and other AI applications. The company offers deployment services to assist customers with data preparation, model architecture design, training management, inference optimization, and, in select cases, ongoing system operations and management. The company also offers a subscription service providing access to an ongoing stream of software updates and upgrades for purchasers of its hardware.
StocksGuide Premium
| Head office | United States |
| Website | www.cerebras.ai |


