Introduction
“The war games say we’re gonna run out of munitions in eight days in a fight with China.”
Palmer Luckey, Co-founder of Anduril, May 2025 (60 Minutes)
For decades, Western defence procurement optimised for technological superiority: exquisite aircraft, ships and missiles, built in small numbers over long development cycles by an ever-smaller group of giant contractors. The war in Ukraine has exposed the limits of that approach. Drones costing a few hundred dollars destroy multimillion-dollar vehicles, designs are revised within weeks as countermeasures evolve, and sustained combat power depends on how fast weapons can be produced, updated and replaced. This new dimension of warfare is not what the old model was designed for.
This mismatch has created an opening for a new guard of defence-technology companies that operate more like commercial technology businesses, designing their products, factories and supply chains for speed, affordability and scale. The new guard is also scaling into powerful tailwinds. Global military spending is rising, Europe is rearming, and Washington has requested a US$1.5 trillion defence budget for fiscal 2027, up from the roughly US$1 trillion enacted for 2026. This follows the 2025 White House executive order, which called for rapid reform of America’s “antiquated” defence acquisition processes with an emphasis on speed and flexibility. Most of the new guard remains private today, but as these tailwinds help the sector mature, we expect a wave of IPOs in the coming years that will attract significant investor interest.
In this note, we first examine how warfare is changing and what that means for technology, affordability and procurement. We then take a closer look at Anduril, a leader in this space, as well as Saronic, to illustrate how the new guard is built from the ground up for this new era. Finally, we assess the size of the opportunity.
The Demands of Modern Warfare
Western defence procurement was shaped by the Cold War strategy of using technological superiority to offset the Soviet Union’s numerical advantage. Over time, this contributed to a focus on “exquisite” systems. These highly sophisticated platforms were designed to maximise performance, but typically came with high costs, development cycles that could stretch over a decade and small production runs.
By the middle of the 2010s, however, two challenges facing the defence sector were becoming clearer. One concerned access to commercial technology. Much of the expertise in software, artificial intelligence and autonomy was concentrated in technology companies that were unwilling to work with the military, making it harder for the traditional defence sector to keep pace with commercial innovation. Palmer Luckey, co-founder of Anduril, described this divide as the “national divorce” between the technology industry and the national-security apparatus. The second challenge concerned building systems for scale. Militaries increasingly needed larger numbers of lower-cost, modular and interoperable systems that could be readily upgraded.
The war in Ukraine brought these challenges into much sharper focus. Defending against relatively inexpensive drones could require interceptors costing around US$1 million. Meanwhile, Ukraine’s drone production rose from several thousand units in 2022 to four million in 2025, as high rates of use and attrition made production capacity and stockpile depth part of combat power. In July 2025, the Pentagon acknowledged that US forces lacked the quantities of low-cost drones required for modern warfare and directed the military to treat small drones as cheap, replaceable and consumable systems, closer to munitions than high-end aircraft.
Figure 1: Iranian-made drones being used by Russia (2023)

Source: ABC News
Producing affordable mass requires redesigning supply chains, not simply building larger factories. Weapons designed for production at scale need common components, commercial manufacturing methods and broad supplier bases; one bespoke part with a long lead time can otherwise constrain an entire production line. Weapons must also become smarter. Their effectiveness depends not only on physical performance but on how well they interpret sensor data, communicate with other systems and respond to changing conditions. Artificial intelligence can combine sensor data, help operators identify potential threats and coordinate uncrewed systems, while software updates can improve capabilities after the hardware has been deployed.
“Never have battlefield outcomes been more susceptible to rapid technological innovation. Our entire acquisition model must adapt to product life cycles measured in weeks and days.”
Maj. Gen. Curtis Taylor, U.S. Army, December 2025 (Army University Press)
The systems must also be able to adapt. In Ukraine, drone designs can be revised within weeks as each side develops new tactics, jamming techniques and other countermeasures. Modular, open architectures allow software, sensors and payloads to be changed without redesigning the whole platform. Modularity therefore reconciles the standardisation required for mass production with the need for continued adaptation, allowing systems to be manufactured at scale without permanently freezing their capabilities.
“Why would a company propose a $1 million solution when they could propose a $100 million solution?”
Palmer Luckey, Co-founder of Anduril, March 2026 (60 Minutes Australia)
“If it’s more expensive, you get more profit. If it is less reliable, you get more profit.”
Brian Schimpf, Co-founder and CEO of Anduril, January 2025 (a16z Podcast)
A further obstacle to producing affordable mass has been the way defence programmes are procured. Some contracts are fixed-price, but complex development programmes have often used cost-plus contracts, which reimburse contractors’ costs and add a fee based on the programme’s estimated cost. Higher estimated costs can therefore support larger fees, leaving the industry with limited incentive to invest in reducing costs. Consolidation has compounded the problem by weakening competitive pressure, with the number of major defence primes falling from 51 after the Cold War to just five today.
These structural weaknesses have, however, created an opening for a new guard of defence-technology companies, sometimes called neo-primes. They are developing systems suited to this kind of warfare: strike and reconnaissance drones, counter-drone systems, crewless surface and underwater vessels, autonomous aircraft, low-cost missiles and the AI software that connects them. As important as what they build is how they build it: designing for affordability from the outset, setting up factories and supply chains to produce at scale. Anduril has been at the forefront of this shift, alongside a widening field that includes Saronic, Castelion, Shield AI, Neros, Epirus and Skydio (Figure 2). Incumbent primes are responding through partnerships, selective acquisitions and their own efforts to deliver affordable mass, leaving the competitive landscape fluid.
Figure 2: Overview of selected new-guard companies

Source: Practical Venture Capital
Anduril – Leading the Charge
“Remember that when we started Anduril, there hadn’t been a significant new defence company of any size for 30 years.”
Palmer Luckey, Co-founder of Anduril, March 2026 (60 Minutes Australia)
At 19, Palmer Luckey founded the virtual-reality company Oculus, which Facebook later acquired in 2014 for approximately US$2 billion. During his time in Silicon Valley, he became concerned about what he saw as a divide between the technology industry and national security. Traditional defence primes, he believed, were good at iterating on existing systems but ill-equipped to build weapon systems around autonomy and artificial intelligence. He responded in 2017 by co-founding Anduril, a defence company built around precisely those technologies, with a mission to transform US and allied military capabilities.
Beyond technology, Anduril’s business model was another important early differentiator. It positioned itself as a defence product company rather than a traditional contractor. Instead of waiting for the government to define and fund a bespoke development programme, Anduril anticipates military needs and uses its own capital to develop products and build production capacity before securing orders, then competes to sell the finished systems. Compared with conventional cost-plus programmes, this creates stronger incentives to reduce costs and development times, while leaving Anduril to absorb the losses when it backs the wrong product.
Anduril’s first product was an autonomous surveillance tower that could spot threats at a distance without requiring a person to watch a screen. It has since expanded across air, land and sea, with products spanning counter-drone defence, autonomous aircraft, missiles and undersea systems. It has also developed working prototypes of subterranean systems. At the centre of the portfolio is Anduril’s core product, Lattice, its AI-powered software platform. It fuses data from sensors into a single real-time model and serves as both the connective layer and command-and-control backbone across the portfolio, linking Anduril’s products with government and third-party systems.
Within its portfolio, Fury, Anduril’s semi-autonomous fighter aircraft, provides a concrete example of how the company can now compete directly with traditional defence primes for major programmes (Figure 3). In June 2026, the US Air Force selected Fury as one of two aircraft for full-scale manufacturing, with Anduril reportedly beating proposals from Boeing, Lockheed Martin and Northrop Grumman. Anduril says the award made it the first new company to win a US fighter-aircraft programme since the 1970s, following a 26-month progression from prototype award to production contract that it describes as the fastest for a fighter aircraft in more than 50 years.
Figure 3: Anduril’s Autonomous Air Vehicle – Fury

Source: Anduril
In 2025, Anduril generated US$2.2 billion of revenue, more than double the previous year. In March 2026, the US Army awarded it a 10-year enterprise agreement with a ceiling of US$20 billion, streamlining procurement by consolidating more than 120 separate procurement actions into a single framework. Anduril projects to reach US$4.3 billion for 2026, though it is expected to generate a loss of more than US$1 billion, according to The Information. Earlier in 2026, Luckey said that Anduril’s mature products generate margins of approximately 40% despite selling for roughly one-tenth of competitors’ prices. However, they remain unprofitable because they reinvest all the cash generated into R&D. In May 2026, the company raised US$5 billion at a US$61 billion valuation. It was later reported in July that the company was seeking a new financing round at a ~US$100 billion valuation.
Ultimately, Anduril plans to go public. Luckey has argued that the company cannot win a programme on the scale of the F-35 while remaining private. In March 2026, Anduril President Matthew Steckman also said that becoming public was important because public companies command greater trust within the US national-security apparatus. However, he estimated that an IPO was still a couple of years away because only around a quarter of Anduril’s 20 products were in rate production and generating cash for the business.
Saronic – Revitalising US Shipbuilding
“The United States can build 100,000 gross tons of ships every year… the Chinese… can build 23 million gross tons. So they can outbuild the US 230 to 1.”
Dino Mavrookas, Co-founder and CEO of Saronic, August 2026 (All-In Podcast)
The United States currently builds fewer than five large commercial ships a year, down from over 70 in 1975. A single Chinese state shipbuilder, CSSC, delivered more than 250 in 2024. The gap is not just a commercial issue, but also a defence issue, because shipbuilding capacity is redirected towards naval production in wartime. The US Navy’s own fleet, at around 290 ships, also sits below the 355-ship target Congress set in the 2018 defence bill.
Saronic was founded in 2022 to address this gap by building unmanned surface vehicles for defence and commercial maritime applications. Among its co-founders are CEO Dino Mavrookas, a former Navy SEAL who later worked in private equity, and CTO Vibhav Altekar, an early Anduril engineer. Saronic has raised about US$2.5 billion in total, including a US$1.75 billion round led by Kleiner Perkins in March 2026 that valued the company at US$9.25 billion.
Figure 4: Saronic’s Autonomous Vessels

Source: Saronic
Saronic’s vessels can operate autonomously or under remote human supervision and can be configured for surveillance, logistics or strike missions (FIgure 4). Its flagship product is Corsair, a 24-foot autonomous speedboat that the company says its Austin factory can produce at a rate of 2,000 per year. In 2025, the Navy signed a US$392 million production agreement with Saronic, reportedly for Corsair vessels. In June 2026, a Navy-operated Corsair made headlines when it located and picked up the two crew members of a US Army Apache helicopter downed near the Strait of Hormuz. Saronic described the operation as the first known at-sea rescue by an autonomous boat.
“Because we’re building for software and autonomy rather than putting people on the ship, we can strip 90% of the complexity out of the ship itself.”
Dino Mavrookas, Co-founder and CEO of Saronic, March 2026 (McKinsey)
By building boats that do not need to cater for a crew, there are a range of benefits. The boat becomes simpler to build, as dedicated features designed for humans such as doors, stairs, plumbing and displays are no longer needed. Additionally, the boat does not need to be designed around human safety. This simplicity can significantly reduce costs while allowing designers to prioritise speed, range and firepower over crew accommodation. Mavrookas argues that, relative to cost, his Marauders can put significantly more firepower to sea than a $3 billion Destroyer. Simplicity also accelerates product timelines, with Saronic going from prototype to serial production of its Corsairs in under a year.
“When you look at the overall Department of War’s budget, 1% goes to autonomous systems. That’s not nearly enough. Make that 5%.”
Dino Mavrookas, Co-founder and CEO of Saronic, August 2026 (All-In Podcast)
To expand production beyond its existing sites in Austin and Louisiana, Saronic announced in July 2026 that it would build Port Alpha, a next-generation shipyard in Brownsville, Texas. The company plans to invest more than US$3 billion in the site, which it says will use software-defined production methods to build vessels at scale and strengthen US shipbuilding capacity. Saronic expects construction to begin in 2026 and operations to start in 2028, with the project expected to create up to 10,000 direct jobs. The initial Port Alpha facility is designed to build vessels up to 850 feet long, while a future expansion could support ships longer than 1,200 feet. By comparison, the largest vessel in Saronic’s current product range is its 180-foot Marauder. Saronic says Port Alpha is intended to build both commercial and military vessels.
Market sizing
Although some new guard companies also pursue commercial opportunities, defence budgets principally define the market. According to SIPRI, global military spending reached US$2,887 billion in 2025, up 2.9% in real terms year-over-year. This represented 2.5% of global GDP, the highest share since 2009. The US remained the largest spender at US$954 billion, representing about a third of the global total, followed by China at US$336 billion (Figure 5). Europe was the main contributor to the increase in global spending, with expenditure across the region rising 14% in 2025 to US$864 billion, including a 24% increase in Germany. This reflected both the war in Ukraine and broader rearmament among European NATO members amid pressure from Washington to shoulder more of the defence burden. NATO members have pledged to invest 5% of GDP annually in defence by 2035. That rearmament is already visible in order books. Raytheon booked more than US$10 billion of international awards in the first half of 2026, more than double the figure in the same period a year earlier, including more than US$7 billion from European customers. International programmes now account for 48% of its backlog.
Figure 5: World military spending, 2025

Source: SIPRI (April 2026)
Note: Percentages are real-terms changes versus 2024.
The new guard is also gaining traction in international markets. Anduril has disclosed a US$1.1 billion Australian award for the Ghost Shark submarine programme and has agreed with Poland’s PGZ to build its Barracuda missiles locally. Even so, the US remains the new guard’s largest market, and Washington sets the pace. For fiscal 2027, the administration has requested congressional approval for a US$1.5 trillion defence budget, up from roughly US$1 trillion enacted for 2026. About half of this, or US$756.8 billion, is investment funding for buying and developing weapons, comprising US$413.1 billion for procurement and US$343.7 billion for research and development. The department describes this nearly 75% year-over-year increase as historic. The stated priorities behind this funding align closely with the new guard’s focus: advancing high-tech capabilities and autonomous systems, rebuilding weapon stockpiles, revitalising US shipbuilding and expanding the defence industrial base.
A large share of the funds has typically flowed to the incumbents. For context, Lockheed Martin booked US$75 billion of sales in 2025, compared to Anduril’s US$2.2 billion (Figure 6). A September 2025 Bain & Company report estimated, based on valuations at the time, that the new guard’s market share would rise from about 1% to 5-7% by 2030, equivalent to roughly US$15-20 billion of annual revenue. Although any such projection is highly speculative, and traditional primes are likely to retain a substantial share of spending, at current growth rates and with the tailwinds we have described, we think this target should not be difficult to achieve.
Figure 6: The new guard’s projected 2030 revenue versus the 2025 market

Source: SIPRI (April 2026), Lockheed Martin, Anduril, Bain & Company (September 2025).
Conclusion
Modern warfare is creating demand for systems that are smarter, cheaper and available in far greater numbers. The new guard has seized this opportunity and is now benefiting from rising budgets and procurement reforms that prioritise speed and flexibility. The new guard has already made the incumbent primes dance, a sign of its growing influence, but also a warning that the incumbents will not surrender market share without a fight. Defence procurement also remains difficult for newcomers to navigate, particularly in programmes that may support only one winner. Execution risk around scale-up, software performance in contested environments and capital intensity remains material. Nevertheless, Anduril and Saronic have shown that new entrants can win major contracts and reach meaningful scale, offering an early glimpse of how some of today’s challengers could become the primes of tomorrow.
At AlphaTarget, we invest our capital in some of the most promising disruptive businesses at the forefront of secular trends; and utilise stage analysis and other technical tools to continuously monitor our holdings and manage our investment portfolio. AlphaTarget produces cutting edge research and our subscribers gain exclusive access to information such as the holdings in our investment portfolio, our in-depth fundamental and technical analysis of each company, our portfolio management moves and details of our proprietary systematic trend following hedging strategy to reduce portfolio drawdowns. To learn more about our research service, please visit https://alphatarget.com/subscriptions/.
Introduction
“Whoever has presence on the edge is going to win. The edge is where the humans are.”
Cristiano Amon, CEO of Qualcomm, January 2026 (Time)
The AI boom has largely centred on the data centre, where the largest and most powerful models are trained and run using vast amounts of computing power. Yet a growing share of AI is now executing directly on the devices that people and machines use every day such as phones, laptops, cars, cameras, industrial sensors and robots.This is edge AI: running inference (and increasingly lighter forms of reasoning) locally rather than sending every request to the cloud.
Edge devices cannot match the raw scale of a data centre. However, many useful tasks require only modest compute, and these tasks can benefit from lower latency, stronger privacy and lower costs provided by edge AI. Two technical advances have made this shift practical: smaller, more efficient models that fit within the memory and power budgets of real devices, and steadily more capable on-device hardware (especially neural processing units and unified-memory architectures).
The result is an expanding set of workloads that can live at the edge, creating opportunities across semiconductors, software and device makers. Against this backdrop, the global edge AI market is projected to grow from roughly US$30 billion in 2026 to US$119 billion by 2033.
In this note, we first outline the benefits and use cases of edge AI before examining developments in both hardware and models. We then take a closer look at Qualcomm, one of the leading players in the market, and consider its growth ambitions across handsets, automotive and IoT. We also highlight other listed companies with exposure to the theme before finally discussing the size of the opportunity.
The case for edge AI
Running models locally brings several benefits. The first is speed, as data does not need to make a round trip to a data centre. This is critical in applications such as self-driving cars, which must react within milliseconds to changing road conditions, and humanoid robots, which must adjust their grip on malleable objects in real time. Reducing dependence on a data centre can also improve safety and reliability, as devices can continue operating if their internet connection is lost during an outage or when a machine enters a dead zone. Edge AI can also reduce costs by limiting networking and cloud-processing requirements, while keeping data on the device can improve privacy.
Long before the current wave of interest, AI workloads were already running locally on everyday devices. In 2017, smartphone makers began adding dedicated AI hardware, with Huawei and Apple each introducing their first neural processing units (NPUs). NPUs are small, highly power-efficient accelerators designed for the repetitive calculations used by AI (in contrast to the larger, more powerful and more general-purpose GPUs). NPUs have since worked quietly in the background on tasks such as facial recognition, speech transcription and organising photographs based on their content. Edge AI has also extended far beyond the smartphone. Security cameras classify what they see locally, factory sensors monitor motor vibrations to identify potential failures and inspection drones use onboard vision AI to avoid obstacles their human pilots miss while examining power lines. Over time, the range and complexity of these workloads have grown, allowing devices to perform increasingly sophisticated tasks locally.
Edge AI proliferation
Two of the main drivers of edge AI progress have been improvements in hardware and models. When Apple introduced its first Neural Engine (NPU) in the A11 Bionic chip in 2017, it could perform up to 600 billion operations per second. By 2024, the Neural Engine in the M4 chip was capable of 38 trillion operations per second, over 60 times the A11 Bionic’s figure. As far as we know, Apple has not disclosed comparable Neural Engine figures for the more recent A19 Pro or M5, though both added a Neural Accelerator to each GPU core, with the M5 delivering over 4x the peak GPU compute of the M4.
“Our models provide the quality of frontier LLMs on specialized applications but with LFMs, which are up to 1,000 times smaller.”
Ramin Hasani, Liquid AI CEO and Co-founder, January 2026 (McKinsey)
Models are also advancing alongside the hardware, with smaller models becoming increasingly capable. Liquid AI, an MIT spin-out that raised US$250 million in December 2024 with AMD among its backers, is one of the companies pioneering device-native foundation models. It calls its models Liquid Foundation Models (LFMs), which are essentially a form of small language models (SLMs). Liquid says its differentiator is that efficiency is built into its models from the ground up, rather than achieved solely by compressing massive LLMs into a smaller footprint. Its models are designed around the memory, processing and power limits of real devices, bringing reasoning, vision and speech to phones, vehicles and embedded systems. Liquid says its 1.2-billion-parameter reasoning model delivers the fastest inference and best quality in its size class while fitting within 900 MB of memory on a phone. Its smallest model can run on a Raspberry Pi 5. Liquid AI and Mercedes-Benz announced a multi-year partnership in April 2026 to bring speech, language understanding and reasoning onto the vehicle’s own hardware in North America, with a first production deployment for advanced speech technology targeted as early as the second half of 2026.
Edge AI has also become far more prevalent on personal computers, with broadly capable models, including large language models (LLMs), now running locally. An unlikely hero of this shift has been Apple’s Mac mini, given its low cost and Apple’s unified memory design. Running a language model is above all a memory problem, as the entire model must sit in memory the chip can access. On a conventional PC, the CPU uses system RAM, while the discrete GPU has its own onboard memory, which is faster but usually much smaller. The GPU works most efficiently on data held in its own memory, so anything sitting in system RAM generally must first be copied across a narrower connection. Apple’s unified memory removes this split entirely: the CPU and GPU draw on one shared pool, allowing the GPU to access far more memory than most consumer graphics cards contain. Large unified-memory configurations therefore enable computers to run larger and more capable models locally. The industry is now building for this shift deliberately. NVIDIA’s DGX Spark, launched in October 2025, is a palm-sized machine built on the same architectural idea. It has 128 GB of unified memory and can run models of up to 200 billion parameters locally (Figure 1).
Figure1: Evolution of AI computing: Nvidia DGX-1 to DGX Spark

Source: Nvidia
We have also seen OpenAI lay the groundwork for a move towards the edge. In 2025, OpenAI acquired io, a device start-up co-founded by former Apple design chief Jony Ive, for ~US$6.5 billion. Bloomberg reported in July 2026 that its first product will be a portable, screen-free smart speaker with a camera and sensors, designed to act as a ChatGPT companion in the home. OpenAI’s job listings indicate an edge strategy, with roles for an inference technical lead focused on on-device transformers and an SoC architect developing custom AI silicon for edge deployments. Reuters reported in December 2025 that its earliest devices will be cloud-based, with more capable local inference likely to follow in later products. According to a court filing, OpenAI does not expect its first hardware device to ship before the end of February 2027.
Qualcomm
“We find ourselves at this inflection point, and as this matures, we see a massive edge content upgrade cycle hit us.”
Nakul Duggal, Qualcomm Executive Vice President, June 2026 (Investor Day)
One of the largest semiconductor players in the edge AI space is Qualcomm. Its handset segment generated US$27.8 billion in FY25, accounting for 63% of company revenue. This segment is underpinned by its Snapdragon platform used in Android devices, which combine a CPU, GPU and NPU. However, its growth expectations for Android handsets are modest. At its June 2026 Investor Day, Qualcomm forecast a 5% CAGR through FY29, though this assumed no uplift from AI and no improvement in the memory market, two areas management identified as potential sources of upside.
By contrast, it expects much stronger growth from its other segments. In automotive, Qualcomm’s compute content per vehicle grew 8x between FY22 and FY26 and management expects demand to continue to remain robust (Figure 2). This will be driven by richer digital cockpit capabilities, more sensors, higher levels of driver assistance and autonomy (robotaxis), and generative AI. Automotive revenue reached US$4.0 billion in FY25 and is forecast to reach US$10 billion by FY29, representing a 26% CAGR, supported by a US$65 billion design-win pipeline.
Figure 2: Qualcomm’s Automotive Platform

Source: Qualcomm Investor Day 2026
“When I have glasses that I’m wearing all the time, the amount of information [gathered] is going to be so much bigger, that whoever is present at the edge is actually going to have a better model over time.”
Cristiano Amon, CEO of Qualcomm, January 2026 (Time)
Qualcomm’s IoT revenue is also expected to grow strongly, from US$6.6 billion in FY25 to more than US$14 billion by FY29, representing a 20% CAGR (Figure 3). Qualcomm splits IoT into two buckets. Personal AI and Compute, which includes devices such as smart glasses and PCs, is forecast to reach US$6 billion by FY29. Industrial, Networking and Robotics, which serves sectors such as oil and gas and utilities, is forecast to reach US$8 billion.
Figure 3: Qualcomm IoT Revenue Growth Forecast

Source: Qualcomm Investor Day 2026
Finally, it is worth noting that Qualcomm has announced its entry into the AI data centre market, where it is betting that performance per watt, a key advantage in its devices business, will carry over. This is not its first attempt: its Centriq server CPUs were discontinued within roughly a year of launch, and its earlier Cloud AI 100 inference accelerators found limited commercial traction. Nevertheless, Qualcomm forecasts revenue from data centres will exceed US$15 billion by FY29 (Figure 4), broadly comparable to its forecast for its IoT segment. The move underlines that the future is not a choice between the edge and the data centre, as Qualcomm expects both to grow meaningfully, with some workloads best run locally and others in the cloud.
Figure 4: Qualcomm revenue targets for FY29

Source: Qualcomm Investor day 2026
Other listed companies with edge AI exposure
There are many other listed companies that offer direct or indirect exposure to edge AI. The examples below are not exhaustive but illustrate the range of companies participating in the theme.
- Semtech. Provides high-performance semiconductors for data centre networking, IoT connectivity and cellular infrastructure. Its long-range, low-power wireless platform, LoRa, provides connectivity for edge AI in IoT applications.
- Synaptics. Provides embedded compute, wireless connectivity and multimodal sensing solutions. In June 2026, it was announced that Onsemi would acquire the company at an enterprise valuation of ~US$7 billion.
- Ambiq Micro. Provides ultra-low-power chips that run AI on battery-powered devices such as smartwatches, healthcare monitors and sensors.
- Mobileye. Provides the chips and software for driver assistance and autonomous driving.
Market sizing
According to Grand View Research, the global edge AI market was valued at US$24.9 billion in 2025. Hardware accounted for the largest share at 51.8%, while North America was the largest regional market with a 36% share. By end use, consumer electronics held the largest share, driven by AI chips embedded in smartphones, wearables and smart home devices. The market is forecast to grow from US$29.9 billion in 2026 to US$118.7 billion by 2033, a CAGR of 21.7% (Figure 5), driven by the expansion of connected devices, demand for real-time processing, AI-enabled automation and an increasing focus on data privacy.
Figure 5: Edge AI market forecast

Source: Grand View Research
Conclusion
We expect the future of AI to be hybrid. Data centres will still be required for the most demanding workloads, such as frontier-model training and scientific discovery, and benefit from the higher utilisation that comes with scale. Edge AI, on the other hand, brings intelligence closer to where data is generated and decisions are made. The balance between the two will depend on the economics and technical requirements of each workload and will keep shifting as models and hardware improve. Edge AI will not replace the data centre, but it will continue to expand what can be done without one.
At AlphaTarget, we invest our capital in some of the most promising disruptive businesses at the forefront of secular trends; and utilise stage analysis and other technical tools to continuously monitor our holdings and manage our investment portfolio. AlphaTarget produces cutting edge research and our subscribers gain exclusive access to information such as the holdings in our investment portfolio, our in-depth fundamental and technical analysis of each company, our portfolio management moves and details of our proprietary systematic trend following hedging strategy to reduce portfolio drawdowns. To learn more about our research service, please visit https://alphatarget.com/subscriptions/.
Introduction
“How big is the market for slow search?”
Andrew Feldman, Cerebras Founder and CEO, Q1 2026
Over the past decade, a new category of specialised AI hardware has been quietly developing that is now coming to the forefront: fast inference accelerators. The appeal of these systems is fairly intuitive (Figure 1). Higher inference speeds improve the productivity of both humans and agents by allowing more work to be done in less time; they also deliver the instant, low-friction responses that users expect and prefer. Two of the most well-known companies in this space are Cerebras and Groq, with Cerebras claiming its system can deliver inference speeds up to 15 times faster than leading GPU-based alternatives. The company recently went public, raising over US$6 billion in the largest semiconductor IPO of all time and has a multi-year deal with OpenAI worth more than US$20 billion. Groq has also received major validation through a licensing agreement with Nvidia, whose CEO Jensen Huang estimated that the new Groq product would add an incremental 25% to Nvidia’s revenue.
Figure 1: Exchange between Paul Graham and Sam Altman on the need for fast inference

Source: X
In this note, we first discuss the inference bottlenecks that fast inference architectures are designed to address and how this is leading to an unbundling of AI hardware accelerators. We then profile Groq and Cerebras, outlining the foundational technologies behind their approaches, as well as their commercial partnerships and prospects. Finally, we discuss the size of the opportunity.
Unbundling AI Hardware
“There’s only two ways I know of to make money – bundling, and unbundling.”
Jim Barksdale, former Netscape CEO, 1995
AI inference has two separate stages: prefill and decode (Figure 2). Prefill is the input stage, where the model processes the user’s prompt. This stage is well suited to standard AI accelerators such as Nvidia GPUs because much of this work can be done in parallel, which is where they thrive. Decode is the output stage, where the model generates the answer one token at a time. Because each new token depends on the one before it, the process is inherently sequential and cannot be parallelised. At this stage, the bottleneck shifts from compute to memory bandwidth and latency, often referred to as the memory wall: the system is limited by how quickly it can move data between memory and compute.
Figure 2: Prefill and decode

Source: Cerebras
Several companies, most of which remain private, are currently working to address this bottleneck. Groq and Cerebras are two of the more well-known companies in this space, which are building their systems around a particular type of memory called SRAM. SRAM has significantly more bandwidth than HBM, the primary memory used in GPUs. SRAM also sits on the compute chip, whereas HBM sits off the chip, meaning data has to travel much shorter distances. This reduces both latency and power use.
However, SRAM has historically been difficult to use as primary memory because it takes up far more silicon area than HBM, which has meant SRAM capacity has been too limited. As a result, SRAM has traditionally been used only as a small, fast cache sitting above much larger DRAM/HBM main memory. This new generation of AI accelerators has been able to alleviate this capacity constraint, unlocking SRAM’s bandwidth advantage and resulting in their systems achieving far higher token output speeds than GPUs.
This, in turn, has also allowed inference to be disaggregated across hardware types, where GPUs can handle the prefill stage, while SRAM-based accelerators target the decode stage. Each architecture can therefore focus on the part of the workload it is best suited to. As a result, inference hardware is beginning to unbundle.
Groq – The Software-First Approach
Groq was founded in 2016 by Jonathan Ross, a former Google engineer who helped start what became Google’s Tensor Processing Unit (TPU). Its core technology is the Language Processing Unit (LPU). Unlike the versatile GPU, which can handle many different compute tasks, the LPU was designed specifically for linear algebra calculations, which is the primary requirement for AI inference. The Groq team also took a software-first approach, with execution decisions made in software rather than hardware. The team even developed the software (compiler) before they began chip design. Because all execution planning happens in software rather than hardware, this allows for a simpler and more efficient hardware architecture that executes predetermined scripts. This determinism helps reduce delays and improve latency stability, while the simpler hardware architecture frees up space for additional memory bandwidth and transistors for performance.
Figure 3: Rubin GPU vs Groq 3 LPU comparison – Capacity vs Bandwidth trade-off

Source: Nvidia
However, even with the gains noted above, the amount of SRAM that fits on a single LPU remains limited. The Groq 3 LPU still only has a 500 MB SRAM, a tiny amount relative to 288 GB of HBM, an Nvidia Rubin GPU (Figure 3). This is far less than large models require, which means many LPUs must be connected together to create a larger effective pool of memory. This introduces synchronisation challenges. Groq solves this in two ways: its software maps out in advance exactly when each chip sends data to the others, while its chip-to-chip protocol cancels clock drift and keeps hundreds of LPUs aligned. Together, this allows the system to behave like one coordinated chip.
In December 2025, Groq and Nvidia entered into a non-exclusive licensing agreement. As part of the agreement, members of Groq’s team, including its CEO Jonathan Ross, joined Nvidia. News outlets reported that the deal was worth US$17–20 billion. (The structure has since drawn scrutiny from Senators Elizabeth Warren and Richard Blumenthal, who questioned whether it was designed to avoid antitrust laws.)
At Nvidia’s GTC event in March 2026, Jensen Huang unveiled the Nvidia Groq 3 LPX. The Groq 3 LPX is a rack of 256 interconnected LPUs, offering 128 GB of SRAM capacity and 40 PB/s of bandwidth (640 TB/s of scale-up bandwidth across the rack), positioned as a low-latency inference layer for Nvidia’s Vera Rubin platform. By combining Rubin GPUs with Groq’s LPUs, inference is disaggregated, with Groq handling the bandwidth-heavy decode, enabling the platform to reach much higher tokens per second (TPS) (Figure 4). However, it was explained that the price for this performance boost would be set “quite high” because it has a lower TPS per MW, and was intended for high value workloads and users like software engineers who could afford it. Nvidia expects to start shipping the Groq 3 LPX in the second half of 2026.
“If you extended this chart way out here and you said you wanted to have services that delivers not 400 tokens per second, but a 1000 tokens per second, all of a sudden, NVLink 72 runs out of steam and simply can’t get there. We just don’t have enough bandwidth. And this is where Groq comes in.”
Jensen Huang, CEO of Nvidia, March 2026 (GTC 2026)
Figure 4: Vera Rubin NV72 + Groq 3 LPX enabling High-Throughput and Low-Latency

Source: Nvidia
Cerebras – The Massive Chip Approach
“Fast tokens are more valuable tokens and Cerebras tokens are the fastest.”
Andrew Feldman, Cerebras Founder and CEO, Q1 2026
Cerebras was co-founded in 2016 by its CEO Andrew Feldman on the founding bet that the age of AI would demand a new kind of compute, just as PCs needed x86, graphics needed GPUs and mobile needed ARM. Like Groq, Cerebras uses SRAM as its primary memory, but it takes a very different approach. Rather than connecting many smaller chips together, Cerebras turned an entire silicon wafer into one massive processor, called the Wafer-Scale Engine (WSE). The latest WSE-3 is 58 times larger than Nvidia’s B200, which gives Cerebras far more space to place SRAM directly alongside compute (Figure 5). In total, the chip carries 900,000 AI-optimised cores and 44GB of on-chip SRAM, resulting in 21PB/s of memory bandwidth, 2,625 times more than an Nvidia B200 package. Overall, Cerebras says this architecture enables inference speeds up to 15 times faster than leading GPU-based systems.
Figure 5: Cerebras Wafer Scale Engine 3 vs Nvidia GPU B200 – 58x larger

Source: Cerebras
The key benefit of wafer-scale is that far more communication can stay on-chip. Even when a model is too large to fit entirely on a single wafer, the large chip reduces how often data needs to travel over slower interconnects between separate chips or systems. Achieving this chip scale required Cerebras to develop two foundational semiconductor technologies:
- Multi-die interconnect: A die is a region of silicon containing an integrated circuit, individually stamped onto a silicon wafer and then normally diced into small, separate chips. Cerebras invented a technology to interconnect these otherwise independent die at the wafer level. This lets adjacent die communicate at the same bandwidth as within a single die, so the whole wafer works as one chip and avoids the slowdown of going off-chip.
- Fault-tolerant architecture: Wafers typically have defects, but normally when a wafer is diced into smaller chips, the defective ones are discarded. With a WSE this cannot be done, because the whole wafer is effectively a single chip. Cerebras’ answer was to build redundancy into the design, so that flaws are recognised, shut down and routed around.
(For a deeper dive on Cerebras’ technology, see: https://www.youtube.com/watch?v=7GV_OdqzmIU.)
Cerebras initially spent much of its existence as a relatively obscure hardware company, building a radical chip the market was not ready for. That began to change in 2023, when the UAE’s G42 commissioned Condor Galaxy, a series of AI supercomputers built on Cerebras hardware. Momentum has built since then, with revenue reaching US$510 million in 2025, up 76% year-on-year. In January 2026, OpenAI signed a multi-year agreement worth more than US$20 billion, under which it agreed to deploy 750MW of Cerebras compute and to co-design future models for Cerebras hardware. This partnership is already beginning to translate into customer access, with OpenAI saying its upcoming GPT-5.6 Sol model will be available on Cerebras in July 2026 at up to 750 tokens per second for select customers as capacity expands. In March 2026, Cerebras also launched a multi-year partnership with AWS to bring its fast inference to a broader enterprise and developer market.
Market Sizing
“When we look out at the space, we see the entire inference market as available to us for fast inference. I mean, who doesn’t want answers in less time?”
Andrew Feldman, Cerebras Founder and CEO, Q1 2026
In terms of market demand for fast inference, we expect this to be strong given the wide range of use cases in which it can boost productivity, from coding agents and document analysis to drug discovery. Additionally, fast inference can also improve user experiences, such as cutting dead time while users wait for responses, helping them stay in flow rather than switching context. Voice is another area where fast inference is very beneficial, as it can help address the current trade-off between response quality and speed, as even small delays can make the interaction feel unnatural.
At the March 2026 GTC, Jensen Huang estimated that Nvidia’s Groq product would add approximately 25% in incremental total revenue. For context, Huang said Nvidia had US$1 trillion in high-confidence demand for Blackwell and Rubin through the end of 2027, which means its Groq product would add an additional US$250 billion for the period. Cerebras reported US$193 million in revenue for FY26 Q1, up 92% year-on-year. Growth is expected to remain strong as the company scales production under its OpenAI and AWS partnerships, which provide significant visibility into future demand.
Figure 6: AI Cloud Semiconductor Forecast

Source: Morgan Stanley Research
Looking further out, Morgan Stanley estimates that the fast inference market could capture 20% of the inference semiconductor market by 2030, representing an US$80 billion opportunity (Figure 6). This is significantly lower than a simple extrapolation of Huang’s 25% figure across the broader AI semiconductor market. However, given that the commercialisation of fast inference is still in its infancy, forecasts should be treated with wide error bars. There are still major unknowns, including how quickly supply can scale, what capacity can be secured, how price competitive fast inference becomes and how much customers are willing to pay for speed. There is also uncertainty around the pace of improvement in SRAM-based architectures relative to GPUs and other inference solutions. Ultimately, the question is not whether customers will want faster inference, but whether supply and demand converge at a point where fast inference becomes ubiquitous, similar to fast internet, or instead occupies a smaller segment that is only affordable for high-value workloads.
Limitations and Risks
Large on-chip SRAM architectures such as Groq and Cerebras excel in low-batch, latency-sensitive decode scenarios where memory bandwidth is the primary constraint. While demand for faster inference is robust in interactive and agentic workloads, there are scenarios where their advantages are eroded. Model efficiency improvements represent the most significant long-term risk. Techniques such as distillation, aggressive quantisation, Mixture-of-Experts routing, and speculative decoding can materially reduce both model size and the memory bandwidth required during decode. As these methods mature, a growing share of inference workloads may achieve acceptable latency on conventional GPU infrastructure at lower cost and higher utilisation, narrowing the addressable market for specialized SRAM-based systems.
Furthermore, large-scale batch inference, offline processing, and throughput-oriented jobs continue to favor GPUs, which offer superior utilisation and more mature software ecosystems. Extremely long context lengths can also erode the advantage, as KV cache requirements eventually exceed even the substantial on-chip SRAM capacity, forcing data movement over slower interconnects.
Finally, power efficiency and total cost of ownership remain important considerations. Both Groq and Cerebras architectures carry higher power draw relative to their performance in certain configurations, and the premium pricing required to justify this may limit adoption outside high-value use cases.
Conclusion
“We don’t know all the things that can be done with fast AI because we haven’t had it yet.”
Jonathan Ross, Groq founder, March 2026 (Nvidia GTC 2026)
In large, growing markets, unbundling is a natural part of the cycle as specialists emerge to solve persistent pain points. Although GPUs remain foundational to AI infrastructure, the memory-bandwidth constraint has created an opening for specialised SRAM-based architectures. Groq and Cerebras are both pursuing this approach but through meaningfully different designs: Groq relies on a software-first, deterministic model across many smaller chips, while Cerebras uses a monolithic wafer-scale architecture to maximise on-chip bandwidth. Both appear well positioned, supported by strong commercial agreements, though competition is intensifying as more players emerge. Nvidia, the incumbent, has also moved quickly to bundle these capabilities back into the broader AI infrastructure stack in order to defend its dominant market position. Given the nascent stage of the market, there is still a limited line of sight on how large it could ultimately become. However, we expect demand to be robust, given the productivity gains and improved user experiences that fast inference can deliver, as well as its potential to unlock new AI use cases that have yet to be imagined.
At AlphaTarget, we invest our capital in some of the most promising disruptive businesses at the forefront of secular trends; and utilise stage analysis and other technical tools to continuously monitor our holdings and manage our investment portfolio. AlphaTarget produces cutting edge research and our subscribers gain exclusive access to information such as the holdings in our investment portfolio, our in-depth fundamental and technical analysis of each company, our portfolio management moves and details of our proprietary systematic trend following hedging strategy to reduce portfolio drawdowns. To learn more about our research service, please visit https://alphatarget.com/subscriptions/.
Introduction
“There’s this bizarre 10x per year growth in revenue that we’ve seen… you would think it would slow down, but we added another few billion to revenue in January. ”
Dario Amodei, CEO of Anthropic, February 2026 (Dwarkesh Podcast)
Anthropic, the creator of the Claude AI models, has seen unprecedented growth in recent years. In early April 2026, its annualised revenue run-rate surpassed US$30 billion, up around 30x since January 2025. This puts it roughly in line with its major rival OpenAI, though Anthropic continues to grow at a materially faster pace.
Much of Anthropic’s success has been driven by Claude’s strength in coding and enterprise workflows. However, the company now faces several hurdles to overcome which will shape its next phase, including challenging unit economics and the path to self-funding profitability. Furthermore, Anthropic appears to be increasingly compute constrained, with some cracks already beginning to emerge. Last but not least, competition is intensifying, with OpenAI appearing better positioned on the compute front, while Chinese open-weight models continue to rapidly advance.
In this note, we first provide an overview of Anthropic’s products and models and discuss the drivers of its success. We then examine future growth hurdles, assess Anthropic’s competitive position against OpenAI and explore the broader competitive landscape. Finally, we discuss how investors can gain exposure to the company today through public-market proxies.
Claude
Claude is Anthropic’s core AI platform, available through a chatbot and a growing range of dedicated products. It is also offered via an API, allowing its models to be integrated into other applications and workflows. Users can choose between a free tier, a Pro subscription at US$20 a month or Max plans starting at US$100 a month, while API usage is priced separately on a per-token basis.
Much of Anthropic’s product expansion has centred on agentic offerings. Claude Code, which became generally available in May 2025, is an agentic coding tool designed to help developers complete complex tasks more autonomously. Claude Cowork, which became generally available in April 2026, is aimed at non-technical users and is designed to handle multi-step tasks across a user’s computer, files and applications. Anthropic has continued to roll out new products and features at a rapid pace, spanning everything from design and productivity tools to developer infrastructure.
The company also offers a tiered model lineup. Opus 4.7 is the most capable broadly available model and is positioned for the most complex coding and reasoning tasks (Figure 1). Sonnet 4.6 offers the best balance of speed and intelligence, while Haiku 4.5 is the fastest and most cost-efficient option for lighter or high-volume workloads.
Figure 1: Claude Models via API – Intelligence vs Speed vs Cost

Source: Anthropic
Anthropic’s Rise
Anthropic was founded in early 2021 by a group of former OpenAI employees led by siblings Dario and Daniela Amodei. It initially operated largely out of the spotlight, but that changed as Claude gained significant traction. This was reflected in its rapid revenue growth, with its annualised run-rate increasing by roughly 10x in both 2024 and 2025 before surpassing US$30 billion by early April 2026 (Figure 2).
Figure 2: Anthropic’s exponential run-rate revenue growth

Source: Anthropic
Anthropic’s success has been driven by a confluence of factors. A primary driver was its early focus on coding, where it regularly topped benchmarks. Coding turned out to be a task where AI models particularly excelled, given the huge amount of training data available and the clear feedback loops. This enabled large productivity gains, which in turn drove significant demand from developers within enterprises.
User experience has been another key advantage. Anthropic lead engineer, Felix Rieseberg, recently argued on The MAD Podcast that raw model performance alone is often not enough to win a market. He pointed to Claude Code as an example: by embedding Claude directly into the terminal, Anthropic made the model more accessible and more deeply integrated into developers’ existing workflows, which helped boost adoption.
Broader availability has also helped. Anthropic has been the only frontier-model company to make its models available across all three major cloud platforms: AWS, Google Cloud and Azure. Anthropic also benefited from the viral rise of OpenClaw, an agent platform first released in November 2025 (under the name Clawd) and powered by Claude, which drove further demand.
Anthropic’s brand has also been a key asset. We have come across many anecdotes of enterprises choosing Anthropic because it is seen as more reliable and trustworthy than its competitors. Additionally, Business Insider reported in early March 2026 that Claude saw an influx of users following Anthropic’s stance against the Pentagon’s proposed use of its models. Many of those new users were defectors from ChatGPT, as OpenAI had subsequently agreed to deploy its models for the Department of War.
Figure 3: Model Adoption – Share of U.S. businesses with paid subscriptions to AI models, platforms and tools

Source: Ramp
The commercial base remains heavily enterprise-oriented. In October 2025, Reuters reported that, according to an Anthropic spokesperson, roughly 80% of the company’s revenue came from businesses. On 12 February 2026, Anthropic said the number of customers spending more than US$100,000 annually on Claude had grown 7x in the past year, while eight of the Fortune 10 were already customers. It also said that more than 500 customers were spending over US$1 million annually. By 6 April 2026, that figure had exceeded 1000, doubling in less than two months.
Looking Ahead
Anthropic is now trying to build on its enterprise foothold by broadening adoption with Claude Cowork. Early signs appear encouraging, with Bloomberg reporting in April 2026 that, according to Anthropic’s Chief Commercial Officer, Cowork had seen stronger adoption in its first few weeks than Claude Code did over a comparable period a year earlier. The executive also suggested Cowork could ultimately reach a broader market than Claude Code, noting that engineers typically account for only 2% to 5% of staff at a large company, while Cowork is designed to appeal to a much wider group of non-technical users.
“There is something both impressive but also slightly terrifying about seeing a model that is so much smarter than the last model.”
Felix Rieseberg, engineering lead for Claude Cowork & Claude Code Desktop, April 2026 (The MAD Podcast)
Anthropic is also continuing to make strong progress at the model level. Opus 4.7, which was released in mid-April, was a solid improvement on Opus 4.6. Earlier that month, Anthropic claimed to achieve an even bigger leap with Mythos, a general-purpose frontier model that people within the company described as a “GPT-3 moment.” However, they discovered that it had outsized capabilities in cybersecurity with far-reaching implications for the safety of software and infrastructure. As such, Anthropic does not plan to make Mythos generally available, reflecting its view that releasing a model with these capabilities into the public domain today would be too dangerous.
Instead, Anthropic is providing limited access to a select group of organisations through Project Glasswing, an invitation-only programme for defensive cybersecurity work. The programme is intended to help organisations responsible for critical software and infrastructure strengthen their defences and prepare for a future in which more powerful models with Mythos-like capabilities become more widely available. Many have applauded Anthropic for acting responsibly by choosing not to commercialise the model more broadly. Others have argued the risks were overstated and that the move was more of a marketing exercise. Some have also suggested that, given Anthropic’s current compute constraints, it may not have had the capacity to support broader commercial deployment, with Mythos in its current form costing 5x as much per token compared to Opus 4.7.
Nevertheless, these developments show that the core opportunity is expanding on two fronts: models continue to improve rapidly in both capability and real-world relevance, while products like Claude Cowork make those capabilities accessible to far more users. This combination should sustain demand growth well beyond current levels.
The Growth Tightrope
“The amount of compute the industry is building this year is probably, call it, 10-15 gigawatts. It goes up by roughly 3x a year. So next year’s 30-40 gigawatts. 2028 might be 100 gigawatts. 2029 might be like 300 gigawatts.”
Dario Amodei, CEO of Anthropic, February 2026 (Dwarkesh Podcast)
Although Anthropic’s exceptional growth rate is expected to continue, the rate of growth remains highly unpredictable. This is because growth is not just a question of demand, but it also depends on how much compute the company commits to in advance. As outlined in a January 2026 OpenAI blog, compute capacity (GW) and run-rate revenue are highly correlated, with both tripling in 2024 and 2025 and each GW of compute translating into roughly US$10 billion of run-rate revenue (Figure 4). While OpenAI is not a perfect proxy for Anthropic, the chart illustrates the extent to which AI revenue growth is dependent on compute availability.
Figure 4: OpenAI revenue and GW correlation

Source: OpenAI
Compute availability is only one side of the equation; the margin on each dollar of revenue is the other, and here Anthropic’s position is more nuanced. In January 2026 it was reported that Anthropic had revised its 2025 gross margin projection down from 50% to roughly 40%, after inference costs on third-party clouds came in higher than internally forecast. This is no doubt a sharp improvement from a gross margin of negative 94% in 2024, but leaves the company well short of the 77% gross margin that management is reportedly targeting for 2028. Obviously, a gigawatt of compute generating revenue at 40% margins is a very different business to the same gigawatt at 77% margins, and much of Anthropic’s current valuation rests on how convincingly that margin ramp can be delivered.
Earlier in the year, Amodei noted that Anthropic was aiming to maintain or even exceed its 10x annual growth rate, though he thought realistically that this trajectory could begin to bend somewhat in 2026. So far, that bending has not materialised. By early April 2026, the run rate had more than tripled in just over three months, implying annualised growth closer to 100x. However, most expect a run-rate of US$80-100 billion by year end.
Even so, Amodei pointed out that a 10x growth rate is not sustainable over longer time periods. This is because each GW of compute is estimated to cost roughly US$50 billion in data centre CAPEX, and therefore the level of investment required to sustain 10x revenue growth is too risky. The danger is that even a slight shortfall in demand and revenue could lead to bankruptcy. As he put it: “If my revenue is not US$1 trillion, if it’s even US$800 billion, there’s no force on Earth, there’s no hedge on Earth that could stop me from going bankrupt if I buy that much compute.” This risk is compounded by the fact that Anthropic is not yet self-funding at the operating level; the margin ramp has to materialise alongside the compute commitment for the overall financial structure to hold together.
Amodei assumes that industry AI compute capacity will continue to roughly triple annually, reaching ~100 GW by 2028. This is broadly in line with forecasts from analysts such as Morgan Stanley, who project 126 GW by then. If that 3x pace continued into 2029, it would imply roughly 300 GW and a staggering US$3 trillion in industry revenue. Anthropic’s own share of that market would depend to a large extent on its ability to secure enough capacity to keep its models at the frontier while simultaneously serving customer demand.
Anthropic vs OpenAI
On 31 March 2026, OpenAI said in a press release that it was generating US$2 billion of revenue per month, which annualises to ~US$24 billion. That is around US$6 billion below Anthropic’s reported run-rate on a headline basis. However, the gap may be narrower than it appears. Multiple reports have said that Anthropic presents its revenue on a gross basis, meaning its headline figure does not deduct the revenue share paid to cloud partners, whereas OpenAI is said to report revenue on a net basis. Therefore, with Anthropic’s revenue share estimated at between 5% and 27%, the two companies appear to be operating at broadly similar run-rate levels.
However, Anthropic’s growth trajectory is clearly ahead, as OpenAI “only” grew ~20% over the past 3 months, versus Anthropic’s growth of ~233%. Furthermore, Anthropic has higher revenue quality because the majority of its customer base is enterprise. On this basis, it would appear that Anthropic is now in the dominant position. Nevertheless, OpenAI could still mount a meaningful comeback.
OpenAI’s Codex, a rival to Claude Code, is now broadly competitive, with more than 2 million weekly users and growth of more than 70% month-on-month as of 31 March 2026. Additionally, enterprise now contributes more than 40% of revenue and is on track to reach parity with consumer revenue by the end of 2026. OpenAI’s management also argues that its consumer reach of 900 million weekly users gives it a distribution edge, with familiarity in daily life helping drive adoption at work. Hedge fund manager Brad Gerstner also recently predicted that OpenAI’s imminent new model, Spud, would drive an inflection in revenue.
“I think it is true we’re spending somewhat less than some of the other players… I get the impression that some of the other companies have not written down the spreadsheet, that they don’t really understand the risks they’re taking.”
Dario Amodei, CEO of Anthropic, February 2026 (Dwarkesh Podcast)
Compute, however, is the area where the competitive dynamic is most contested. Dylan Patel of SemiAnalysis has argued that Anthropic’s cautious approach has left it at a relative disadvantage, while OpenAI was willing to “just sign these crazy deals”. He added that Anthropic has now had to turn to “lower-quality providers that they would not have gone to before”, noting that Anthropic historically enjoyed access to the best quality providers like Google and Amazon. The framing is provocative but warrants more nuance. Dylan’s observation centres on Anthropic needing to diversify beyond its traditional premium partners into newer or smaller suppliers to address its capacity shortfall. However, this does not diminish the strategic value of Anthropic’s deep, co-designed partnership with Amazon on custom Trainium silicon, which stands as a high-quality, deliberate choice rather than a concession to inferior options.
Trainium does not match Nvidia’s latest GPUs across every dimension. The software ecosystem around CUDA is more mature, and Nvidia retains the edge on flexibility for cutting-edge research workloads. However, Trainium has evolved into a genuine co-design programme between Anthropic and AWS: Anthropic is now running Claude on more than one million Trainium2 chips via Project Rainier, and the two companies’ engineering teams are shaping how Anthropic will use Trainium4 and beyond. Furthermore, Trainium’s memory-bandwidth-per-dollar advantage is particularly well-suited to reinforcement learning workloads. Together with Google DeepMind, Anthropic is now one of only two frontier labs with meaningful hardware-software co-design integration.
We can hypothesise that Anthropic has accepted a different set of tradeoffs: less software flexibility and a thinner research-workload ecosystem in exchange for a better cost-per-token profile on inference, which is where most of the volume and likely margin dollars sit. For a company whose path to a 77% gross margin depends on inference economics rather than training flexibility, that is arguably an advantage rather than a compromise.
That said, capacity constraints are a separate and more immediate issue. The pressure has been acknowledged by Anthropic lead engineer, Felix Rieseberg, who recently said that the overwhelming demand for the company’s products remained a challenge. In early April, Anthropic cancelled subscription-based support for third-party tools such as OpenClaw, with a spokesperson saying they placed an “outsized strain” on the company’s systems. Such tools are now supported only via the API. There has also been mounting frustration among Claude users, with many claiming that output quality has recently deteriorated significantly. One of the most high-profile complaints came from AMD’s AI Director, who reportedly said that “Claude cannot be trusted to perform complex engineering tasks.”
Although the cause of the deterioration remains disputed, with some pointing to a product update, we think it is highly plausible that this reflects capacity constraints following the recent surge in demand. Many high-profile online personalities, as well as power users we spoke with, are now moving to competitors, at least temporarily, despite the costs and frictions involved in switching providers (Figure 5). However, it remains unclear whether these examples point to a broader migration or a more limited reaction among certain users.
Figure 5: High-profile Claude user cancelling his subscription due to degraded output quality

Source: X
Although OpenAI is currently well placed to capture some of Anthropic’s momentum, it would be premature to forecast a winner. Anthropic retains a strong brand, deep enterprise traction and highly regarded models. It is also taking steps to address its weakness in compute, notably through a deal announced with Amazon on 20 April 2026 that secures up to 5 GW of capacity to train and deploy Claude. Anthropic also said that significant new Trainium2 capacity would begin coming online in Q2 2026, with nearly 1 GW of combined Trainium2 and Trainium3 capacity expected by the end of 2026. Under the agreement, Anthropic also committed more than US$100 billion over the next decade to AWS technologies, while Amazon invested US$5 billion in Anthropic, with up to a further US$20 billion available over time. Earlier in the month, Anthropic also secured approximately 3.5 GW of additional capacity through Broadcom and Google, with that capacity expected to begin coming online in 2027. Taken together, these agreements should help Anthropic ease some of its near-term capacity pressure while also expanding its longer-term capacity base.
Beyond compute, both OpenAI and Anthropic are expected to soon release their next-generation models, which could swing momentum in either direction. Additionally, both companies face legal overhangs whose ultimate impact remains uncertain, with Anthropic still contesting a Pentagon supply-chain-risk designation and OpenAI heading into trial in Elon Musk’s case over its restructuring. In sum, with compute, product releases and legal outcomes all still in flux, the race remains far from over.
The Broader Competitive Landscape
The Anthropic-OpenAI competition captures the frontier-lab dynamic but understates the pressure coming from two other directions: rapidly advancing Chinese open-weight models, and vertically integrated competitors with advantages across the stack.
The open-weight gap is closing faster than what most commentators acknowledge. Take Moonshot AI’s Kimi K2.6, for example, which was released in April 2026. On the industry’s most-watched coding benchmark, SWE-Bench Verified, Kimi is at parity versus Opus, while being priced at a fraction of the cost. It also supports native orchestration of up to 300 parallel sub-agents. Kimi is one of several Chinese open-weight models nipping at the heels of frontier labs, alongside Z.ai’s GLM-5.1 and DeepSeek V4, and the cadence of releases has been relentless.
Benchmark parity is one, albeit important, dimension although factors such as long-horizon agentic reliability, polish of tooling, depth of first-party tooling (e.g., Claude Code & MCP), and safety alignment play an equally important role.
Additionally, there is the geopolitical dimension. Chinese open-weight models face real headwinds in US regulated enterprise, where procurement rules, data-sovereignty concerns and export considerations favour domestic providers. However, the rest of the world is largely open territory, and several buyers (e.g., European finance, healthcare, non-US government) increasingly have reasons to prefer self-hostable open weights on their own infrastructure over any frontier-lab API, including Anthropic’s.
The pressure from vertically integrated players with huge distribution advantages is also real. Google runs frontier models on its own TPUs, distributes through Workspace and Android, and owns the underlying cloud, a stack that is difficult to match on unit economics alone. xAI has moved aggressively on compute, with the Colossus cluster giving it scale that few pure-play labs can match. Meta continues to anchor the open-weight ecosystem through Llama, which, even if it does not lead on benchmarks, exerts downward pressure on pricing across the industry.
Ultimately, Anthropic’s moat rests on its software layer, namely Claude Code, Cowork, MCP, and the embedded workflows these products create. The durability of those embedded workflows, rather than any specific benchmark lead, is arguably the more important long-term moat to watch.
Investing in Anthropic
While Anthropic remains private, a range of listed companies hold stakes in it, giving public-market investors indirect exposure. Among the most closely watched are Zoom and SK Telecom, whose holdings in Anthropic are meaningful relative to their market capitalisations.
Zoom invested in Anthropic’s Series C in May 2023. The total raise was US$450 million and Reuters reported that the valuation was nearly US$5 billion. Although we do not know the exact amount Zoom invested, the company disclosed that during that period it had made US$51 million of “strategic investments in equity securities of private companies.” Some analysts believe all or the vast majority of this investment was in Anthropic. If true, that would imply a roughly 1% ownership stake before dilution from future rounds. More recently, Zoom’s CFO also said in an investor call that the company’s balance sheet included a US$1.6 billion line item “of which the most significant portion is Anthropic” for the quarter ending January 2026. The company also recorded a pre-tax gain of US$532 million, which predominantly reflected its Anthropic stake. At that point in time, Anthropic’s most recent funding round was its Series F, which valued the company at US$183 billion post-money.
The other popular market proxy, SK Telecom, announced a US$100 million investment in Anthropic in August 2023. This was in addition to a previous amount it had invested, though we are unaware of the size of that earlier investment. Anthropic’s valuation was also not disclosed at the time of the US$100 million investment, though some speculate that it was at a similar valuation to the Series C of nearly US$5 billion.
Both holdings have become increasingly important to the equity story for Zoom and SK Telecom as Anthropic’s valuation has continued to rise. In its 12 February 2026 Series G funding round, Anthropic was valued at US$380 billion post-money, having raised US$30 billion. Bloomberg later reported in mid-April that the company had attracted investor interest at a valuation of US$800 billion. OpenAI’s 31 March 2026 US$122 billion funding round, which valued it at US$852 billion post-money, also provides a useful reference point given that the two companies are now operating at broadly similar annualised revenue levels. However, Anthropic’s valuation could ultimately surpass OpenAI’s if it maintains its faster growth rate, although that should not be taken for granted given its compute constraints.
Conclusion
Anthropic has built one of the strongest positions in frontier AI through its coding excellence, agentic products, and deep enterprise relationships, driving explosive revenue growth to a ~US$30 billion run-rate. However, the company now faces a more complex test: it must simultaneously scale compute at unprecedented speed, improve gross margins and defend against intensifying competition on multiple fronts from OpenAI, rapidly advancing Chinese open-weight models and vertically integrated players. OpenAI remains formidable, while models such as Kimi K2.6 are achieving benchmark parity on coding tasks while offering lower cost and self-hosting flexibility, while Google, xAI, and Meta apply pressure through distribution and scale. In this environment, Anthropic’s most durable advantage will likely come from the embedded workflows in Claude Code, Cowork, and MCP rather than benchmark leadership alone. The battle for AI dominance is far from over, but it has entered a more demanding phase where execution on compute, unit economics, and product moats will determine the winners.
At AlphaTarget, we invest our capital in some of the most promising disruptive businesses at the forefront of secular trends; and utilise stage analysis and other technical tools to continuously monitor our holdings and manage our investment portfolio. AlphaTarget produces cutting edge research and our subscribers gain exclusive access to information such as the holdings in our investment portfolio, our in-depth fundamental and technical analysis of each company, our portfolio management moves and details of our proprietary systematic trend following hedging strategy to reduce portfolio drawdowns. To learn more about our research service, please visit https://alphatarget.com/subscriptions/.
Introduction
“We see an acceleration because of AI workloads. More data is being used for training of models, more data is being used for inference and that by itself creates more data that needs to be stored.”
Kris Sennesael, WD CFO (February 2026)
As AI continues its rapid advancement, new bottlenecks are emerging across the ecosystem. Much of the investor focus has centred on compute, energy and memory, but storage has also emerged as a critical constraint. AI workloads both consume vast amounts of data and generate enormous new volumes that must be stored, moved and accessed efficiently. This demand is primarily met by two technologies: hard disk drives (HDDs), which offer the lowest cost per unit of storage, and NAND flash-based solid-state drives (SSDs), which offer higher performance. Together, they form complementary layers of the storage stack, with HDDs handling bulk capacity and SSDs serving higher-speed workloads.
“In the last 150 years … 15 billion images were created. With AI, that same number of images was created only in the last one and a half years.”
B.S. Teh., Seagate COO (May 2025)
The sharp rise in demand has supported stronger revenues and profitability for storage manufacturers and driven significant share price appreciation. SanDisk, a NAND/SSD pure-play, has seen its shares rise roughly 20x over the past year, while HDD leaders Western Digital and Seagate are up approximately 9x and 7x, respectively. At the same time, the sector’s historic cyclicality remains an obvious concern, though management teams argue that AI is creating a more secular long-term demand backdrop rather than a traditional boom-and-bust cycle. This is reflected in customers placing greater emphasis on security of supply and increasingly signing multi-year agreements, which should provide greater price and demand stability going forwards. However, whether this marks a lasting shift in industry structure over the longer-term remains an open question.
In this note, we trace how AI-driven demand is reshaping the storage value chain, from the component manufacturers producing HDDs and NAND flash through to the enterprise platforms that package storage into solutions for end customers. We use SanDisk and Everpure (formerly Pure Storage) as case studies to illustrate how this demand is flowing through at each level.
Hard Disk Drives
HDDs are relatively slow in terms of bandwidth, but they remain the most economical option for storing data at scale. Currently, they account for approximately 80% of all data storage and are predominantly sold into data centers. As AI workloads consume and generate vast amounts of data, these drives are essential for keeping that volume affordable and persistent. This includes checkpoint datasets used to train models and maintain model integrity. It also includes growing volumes of inference-related data, particularly from emerging agentic AI systems, which rely on persistent access to large volumes of historical data to support planning, reasoning and autonomous decision-making.
Figure 1: 2025 HDD Market Share

Source: Forbes, Coughlin Associates
This surge in AI-related data volumes is already translating into strong demand for HDDs. The two largest manufacturers (Figure 1), Western Digital (WD) and Seagate, have seen revenues and profits accelerate sharply. For CY26, both companies have already largely sold out their planned capacity, and WD said that three of its top five customers had entered into long-term agreements, with two running through to the end of CY27 and one through to the end of CY28. They are also in conversations regarding CY29 and CY30 capacity. These expanding multi-year customer commitments highlight the growing strategic importance of storage in the AI buildout.
Figure 2: WD Long-term Financial Model (3-5 years)

Source: WD Innovation Date 2026
Against this backdrop, HDD manufacturers are guiding for solid growth going forward, with AI-related storage demand catalysing innovation and reshaping their technology roadmaps. During WD’s 2026 Innovation Day, the company guided to an annual revenue CAGR of more than 20% over the next 3-5 years, with data centre HDD capacity shipments (nearline exabytes) guided at a CAGR in the mid-20%s (Figure 2). WD also said pricing is expected to remain stable. IDC has also forecasted a similar exabyte CAGR in the mid-20%s through 2028 for the overall enterprise HDD market (Figure 3).
Figure 3: Worldwide Enterprise HDD Exabyte (EB) Forecast

Source: IDC, Seagate Analyst Day 2025
The growth in capacity is expected to come primarily from increases in capacity per drive rather than from greater unit volumes. For instance, WD recently announced a new 40TB HDD, up from 32TB, as well as a longer-term roadmap to 100TB by 2029. The company is also innovating in areas such as power optimisation and performance, and recently announced new drive technologies that offer 2x bandwidth, with a roadmap to 8x by 2030. These innovations should meaningfully improve HDD performance for AI workloads.
Taken together, HDDs will remain a critical component of AI infrastructure for the foreseeable future, with cost-effective scaling and ongoing innovation continuing to support demand.
It is worth noting that the boundary between HDDs and SSDs is not static. As AI workloads increasingly require higher-speed access, some tiers of storage that were historically served by HDDs may migrate to SSDs over time. However, the sheer volume of data being generated means that HDDs are likely to retain their role as the backbone of bulk storage for the foreseeable future, even as SSDs capture a growing share of higher-performance tiers.
NAND flash and Solid-State Drives
“We are now seeing NAND demand significantly in excess of our available supply for the foreseeable future.”
Sanjay Mehrotra, Micron Chairman, President and CEO (March 2026)
While NAND flash has many use cases, it is currently seeing especially strong demand due to its use in SSDs for AI data centres. SSDs offer several advantages over HDDs, including higher bandwidth, greater density and lower latency, making them well-suited to high-performance workloads. Demand drivers include vector databases and KV cache offloading, with the latter involving moving the KV cache (the working memory used during inference) from scarce and expensive GPU and system RAM onto SSDs. Additionally, shortages of HDD capacity for bulk storage is driving further SSD demand in some parts of the storage stack.
Figure 4: Q4 CY25 Revenue Ranking for Top Five Branded NAND Flash Suppliers

Source: TrendForce
The NAND market is also fairly concentrated, with the top five players accounting for roughly 90% of the market (Figure 4), the balance largely held by China’s domestically focussed YMTC. Overall revenue for these top five suppliers grew 23.8% in Q4 CY25 quarter-over-quarter to cUS$21 billion, driven largely by price increases. This price trend is forecast to accelerate in Q1 CY26, with Counterpoint expecting a 90% surge quarter-over-quarter (Figure 5). This was reflected in Micron’s recent Q2 FY26 results (ended 26 February 2026), where its NAND revenue grew 169% year-over-year and 82% quarter-over-quarter. For CY26, TrendForce forecasts that NAND revenue will grow 112% year-over-year to US$147 billion.
Figure 5: NAND and DRAM Price Trends Q2 2025 – Q2 2026E

Source: Counterpoint
SanDisk Case Study
SanDisk is a NAND/SSD pure-play that was spun out of WD in February 2025. As of Q4 CY25, it held a 12.8% market share. The company has benefited significantly from the stronger market backdrop: Q2 FY26 revenue (ended 2 January) rose to US$3 billion, up 61% year-over-year, while GAAP net income margins reached 26.5%. Growth was driven primarily by higher average selling prices rather than volume, with the data centre segment posting the strongest sequential growth at 64% (Figure 6).
Figure 6: SanDisk Revenue Trends by Segment

Source: SanDisk
“There’s a whole new category of storage systems and the industry is so excited because this is a pain point for just about everybody who does a lot of token generation today.”
Jensen Huang, Nvidia CEO, CES 2026 (January 2026)
The company has noted that it is in discussions with Nvidia regarding a new KV cache opportunity. At CES 2026, Nvidia unveiled its Inference Context Memory Storage Platform, which adds a dedicated SSD memory layer for KV cache offloading and is set to be available in H2 2026.
Beyond the current demand environment, SanDisk’s technology roadmap positions it for the next phase of AI storage. Its new UltraQLC enterprise SSD platform stores four bits per memory cell compared with three in the current standard, delivering higher capacity, lower latency and improved power efficiency. Management noted in January that UltraQLC is now in customer qualification with two hyperscalers, with revenue shipments expected within the next couple of quarters. The company is also developing High Bandwidth Flash (HBF), designed to deliver bandwidth comparable to GPU High Bandwidth Memory but with 8-16x greater capacity at a similar cost, with the first AI inference devices using HBF targeted for early 2027.
“It is our view that this structural evolution is sustainable and should reduce the cyclicality of our NAND business, creating higher average long-term margins and returns.”
Luis Visoso, SanDisk CFO (January 2026)
The more significant question for investors, however, is whether the current demand environment marks a structural break from NAND’s past cyclical pattern. Historically, NAND has traded as a commodity through quarterly auctions, leaving suppliers exposed to sharp price swings. But management is now seeing a behavioural shift: similar to the HDD market, SSD buyers are increasingly focused on securing supply rather than negotiating spot prices. SanDisk has signed one long-term agreement, with several more in the queue, and some customers are sharing capacity plans through CY29 and CY30. If the industry can transition towards longer multi-year agreements, it would allow suppliers to plan capacity more effectively while making demand and pricing more predictable – potentially reducing the cyclicality that still weighs on valuations despite the sector’s sharp re-rating.
A recent development that is worrying investors is Google Research’s recent TurboQuant breakthrough, announced on 24 March 2026. This is a compression algorithm that the company says can reduce KV-cache memory size by at least 6x and deliver up to 8x performance. On the face of it, this could dampen some of the urgency around NAND demand tied to KV-cache offloading.
However, such efficiency gains cut both ways. If TurboQuant materially improves the return on investment of AI data centres, it could accelerate broader AI deployment and ultimately drive additional infrastructure spending, including on storage. This dynamic (known as Jevons paradox, where efficiency improvements lead to greater overall consumption rather than less) has played out repeatedly across technology cycles. Whether it applies here will depend on the pace and breadth of AI adoption relative to the efficiency gains themselves. This tension between efficiency and demand expansion is arguably the most important variable for the storage thesis more broadly, extending well beyond any single algorithm, and is worth monitoring closely.
Everpure (Pure Storage) Case Study
Everpure, formerly Pure Storage, is an enterprise storage platform built on NAND flash. It buys raw NAND and packages it into its own custom DirectFlash Modules rather than relying on off-the-shelf SSDs like many of its competitors (Figure 7). Its software, Purity, is the operating system that communicates directly with the DirectFlash Modules. The benefit of this architecture is that it allows flash management functions such as wear levelling, garbage collection and overprovisioning to be handled at the array level rather than within individual SSDs. Everpure argues that this architecture supports lower power consumption, higher density and greater capacity utilisation of the underlying NAND.
Figure 7: Everpure’s Technology Stack

Source: Everpure
The company is well-regarded in its space, ranking highest on both Ability to Execute and Completeness of Vision in Gartner’s Magic Quadrant for Enterprise Storage Platforms (Figure 8), and reporting a Net Promoter Score of 84.
Figure 8: Magic Quadrant for Enterprise Storage Platforms

Source: Gartner
In terms of market share, it ranks 4th in the enterprise storage platform market at 6.8% of the market in Q3 CY25 (Figure 9), though it grew faster than its peers at 15.5% versus 2.1% for the rest of the market.
Figure 9: External Enterprise Storage Systems Market, Q3 CY25

Source: IDC
In February 2026, the company rebranded to Everpure to reflect its evolution towards a broader data management platform, and recently agreed to acquire 1touch, a data intelligence and orchestration platform, to strengthen its data discovery and AI-readiness capabilities. In terms of its latest financials, Everpure reported US$3.66 billion in FY26 revenue (ended 1 February 2026), representing 16% growth (Figure 10). It generates revenue from product sales (integrated storage hardware) and subscriptions (Storage-as-a-Service). For FY27, management is guiding for 17%-20% revenue growth, with operating income growing faster at 23%-29%.
Figure 10: Everpure Revenue by Segment

Source: Everpure
The current environment of extreme NAND demand creates both risks and opportunities. For instance, management has warned of supply constraints, though it noted that it has a highly distributed and resilient supply chain to help mitigate these risks. On the positive side, hyperscalers are very eager for capacity. Everpure has seen a footprint expansion with its current hyperscaler customer, Meta, beyond its expectations, and is in engineering test environments with multiple other hyperscalers. Everpure has also recently standardised its hyperscaler business model, under which it will procure some of the components needed by hyperscalers to build the solution in their own environment, while the hyperscalers themselves will procure the NAND through their own supply chains. We expect this will reduce some of the pressure on Everpure’s own supply chain, as hyperscalers are likely better positioned to procure NAND given their greater purchasing scale. Management mentioned that this new business model would be accretive to gross margins.
Overall, amid NAND supply constraints and elevated prices, the ability to maximise usable capacity from each unit of NAND becomes increasingly valuable, potentially putting Everpure in a favourable position in the current environment.
Conclusion
“With some new applications that are coming, we believe the data will become more valuable over time, and that data is going to grow like crazy.”
Dr. Dave Mosley, Seagate CEO (May 2025)
Storage manufacturers have benefited significantly from the ongoing AI data centre buildout, as AI workloads both consume and generate vast volumes of data. We expect this tailwind to continue, particularly as emerging use cases such as AI agents, AI video generation, autonomous vehicles and robotics drive additional demand. This view is reinforced by sustained growth in hyperscaler CapEx budgets, signalling that AI infrastructure investment remains robust.
At the same time, supply growth appears relatively measured; HDD capacity is forecast to grow at around the mid-20%s, while NAND supply is constrained by the multi-year lead times required to bring new fab capacity online. The supply picture is further complicated by YMTC’s trajectory and the potential for shifts in trade policy. On the demand side, efficiency breakthroughs such as TurboQuant could either dampen or amplify storage demand depending on whether Jevon’s paradox applies. These supply and demand crosscurrents make the outlook more nuanced than a simple extrapolation of current trends
For the long-term structural thesis to hold, several conditions would need to be met, and not all are within the industry’s control. AI would need to prove a durable source of demand that delivers clear ROI, sustaining infrastructure investment beyond the current cycle. The shift towards multi-year supply agreements would need to deepen beyond early signings and letters of intent into firm, binding commitments that genuinely lock in volumes and pricing. Capital discipline across the industry, from both established players and emerging Chinese entrants, would need to hold, without a wave of new capacity overwhelming demand. And the current assumption that AI inference workloads will remain heavily concentrated in data centres would need to broadly hold; a faster-than-expected shift towards on-device inference at the edge and endpoint level could meaningfully reduce the volume of data flowing through centralised storage infrastructure. Some of these conditions are observable today, while others such as the trajectory of AI ROI and the long-term architecture of inference will only become clear over a period of years, which is why the market continues to price meaningful cyclical risk into the sector despite the strong near-term fundamentals.
The signposts to watch are: the pace and breadth of long-term agreement signings, the trajectory of hyperscaler CapEx, the commercial impact of efficiency breakthroughs on overall AI deployment, and any shifts in the competitive supply landscape. If these conditions hold, valuations that still embed significant cyclical risk could have further room to re-rate. If they do not, the sector’s history suggests that the correction can be sharp.
At AlphaTarget, we invest our capital in some of the most promising disruptive businesses at the forefront of secular trends; and utilise stage analysis and other technical tools to continuously monitor our holdings and manage our investment portfolio. AlphaTarget produces cutting edge research and our subscribers gain exclusive access to information such as the holdings in our investment portfolio, our in-depth fundamental and technical analysis of each company, our portfolio management moves and details of our proprietary systematic trend following hedging strategy to reduce portfolio drawdowns. To learn more about our research service, please visit https://alphatarget.com/subscriptions/.
“I think in terms of the AI, the biggest challenge I think for a lot of my customers is memory. Memory actually there’s no relief as far as I know when I talked to the you know, only three key players, two of them I talked to very frequently, and then they told me, ‘Lip-Bu, there’s no relief until 2028.’”
Lip-Bu Tan, Intel CEO, February 2026 (Cisco AI Summit)
Introduction
Memory has emerged as one of the hottest sectors in technology, driven by the explosive growth in AI infrastructure build-out. Memory has become a critical bottleneck in AI systems across both training and inference. With demand rising far faster than supply can respond, memory prices have risen sharply, driving strong growth in revenues and profits across the memory industry. The three main players, SK Hynix, Samsung and Micron, have all benefited tremendously, with their share prices rising roughly 3-4x over the past year. All three have also sold out their entire AI GPU memory production for 2026, and the overall memory market is now forecast to grow by 134%, to US$552 billion (TrendForce forecast), by the end of the year.
“They’ve seen boom and bust 10 times. That’s a lot of layers of scar tissue. During the boom times, it looks like everything is going to be great forever. Then the crash happens and they’re desperately trying to avoid bankruptcy.”
Elon Musk, February 2026 (Dwarkesh Patel Podcast)
Despite the strong industry tailwinds, however, both memory producers and investors are still haunted by the boom-and-bust cycles of the past. The key question now is whether history will repeat once again, or whether this is a new normal, with the AI secular trend sustaining elevated economics and ushering in a memory golden age.
In this note, we first provide a market overview of the memory industry, highlighting different memory types and company market shares. Next, we examine why memory is essential to AI training and inference and how it has emerged as a major bottleneck. We then outline some steps taken to mitigate these constraints. We next present a case study of SK Hynix, focusing on its recent performance and future growth. We also discuss the spillover effects from memory shortages on other sectors, including smartphones and PCs. Finally, we provide a market outlook and discuss the sector risks.
The memory landscape
Different types of memory can be viewed as a hierarchy, often illustrated as a pyramid (Figure 1). At the top sits the fastest memory, with the highest bandwidth, where bandwidth refers to how much data can be moved per second and is typically measured in GB/s or TB/s. However, this top-tier memory also has the highest cost and the lowest capacity, where capacity refers to how much data it can hold and is typically measured in GB or TB. As you move down the pyramid, bandwidth falls while capacity rises and cost per bit declines.
Figure 1: The Memory Pyramid Hierarchy (Simplified)

Source: Sam Mokhtari
The memory market is dominated by two categories: DRAM and NAND flash. DRAM is the principal memory technology used for active computing workloads and is much faster than NAND, but it is also more expensive per bit and typically offers less capacity. NAND, by contrast, is used primarily for storage applications such as SSDs and offers much higher capacity at lower cost, but with much lower speed.
Figure 2: Global DRAM Market Share by Revenue

Source: Counterpoint Research
Within DRAM, there are several sub-types. Conventional DRAM typically serves as CPU-attached system memory, while High-Bandwidth Memory (HBM) is a specialised variant of DRAM used in AI GPUs. HBM provides much higher bandwidth than conventional DRAM and is essential for keeping GPUs fed with data during active computation. SK Hynix is currently the leader in HBM and a major supplier to Nvidia (Figure 3).
Figure 3: Global HBM Market Share by Revenue

Source: Counterpoint Research
The memory bottleneck
The reason memory has become such a critical focus in AI is that it is now a key bottleneck to further progress, due to both bandwidth and capacity constraints.
The bandwidth bottleneck stems from GPU compute power having improved at a faster rate than memory bandwidth over the past decades. This divergence has compounded considerably, leading to a significant gap. This is illustrated in the chart below, where compute speed (hardware FLOPS) outgrew memory speed (DRAM bandwidth) by a factor of 600x over a 20-year period (Figure 4).
In practice, this means GPUs frequently finish their calculations and then sit idle waiting for the next batch of data to arrive from memory. The result is expensive GPUs left underutilised and slower model training. The bandwidth bottleneck is also known as the “memory wall,” a term originally coined in 1995 by William Wulf and Sally McKee to describe the growing gap between processor speeds and memory performance. While it originally described the growing gap between processor speeds and memory performance, the term is now used more broadly to encompass capacity constraints as well.
Figure 4: The memory wall

Source: Gholami et al., 2024, “AI and Memory Wall”
The capacity bottleneck has also become a major limiting factor. Training large models requires storing not just the parameters (the core learned weights), but also gradients (signals for updating weights), optimiser states (extra data used to enhance and stabilise updates) and activations (temporary intermediate results used to compute gradients) (Figure 5). Therefore, limited memory capacity acts as a bottleneck for building more powerful models. Similarly, it can also limit inference performance, as longer context windows exhaust available GPU memory.
Figure 5: Memory breakdown example of a 7 billion parameter model

Source: Sam Mokhtari
Solving the bottlenecks
A range of optimisations have been adopted to address these bottlenecks. This includes “Precision Reduction,” which entails storing data in GPU memory in lower precision formats (e.g. FP32 vs FP16), reducing both memory capacity and bandwidth required. A further optimisation is the use of “Parallelism”, which allows the memory that would normally sit on a single GPU to be split across multiple GPUs, enabling larger models.
A third and particularly consequential development is “Key-Value (KV) cache offloading.” KV cache is a data structure created during inference and grows linearly with prompt length. For use cases such as multi-turn conversations, deep research and code generation, limited and costly GPU memory becomes a significant constraint, especially when the KV cache must be retained in memory for extended periods. KV cache offloading solves this by progressively moving less-active portions of the cache from the GPU’s limited HBM first to CPU DRAM (system memory) and then to SSD storage as the cache grows. This hierarchical approach eases the capacity bottleneck for longer contexts without requiring additional GPUs, while keeping the most frequently used (“hot”) data in the fastest memory tier. The surge in inference workloads and wider adoption of KV cache offloading have, in turn, driven a sharp rise in demand for conventional DRAM and NAND.
Beyond these system-level techniques, architectural innovations are also reducing memory intensity per unit of AI capability. Mixture of Experts (MoE) models, such as those used by DeepSeek and Mistral, activate only a fraction of the model’s total parameters on any given forward pass. This means a model with hundreds of billions of parameters may only require memory bandwidth for a small subset during each computation, significantly easing both bandwidth and capacity demands relative to a dense model of equivalent capability. Similarly, distillation, namely the process of training smaller, more efficient models to replicate the outputs of larger ones, is producing compact models that deliver strong performance with a fraction of the memory footprint. Together, these developments mean that useful AI capability is growing faster than raw memory consumption, which has important implications for the demand outlook.
On the hardware side, Processing-in-Memory (PIM) represents an emerging approach that differs fundamentally from the software-level optimisations above. Rather than accepting the separation between compute and memory and working around it, PIM integrates computational capabilities directly into the memory itself, reducing the need to move data back and forth. SK Hynix (see next section) showcased several PIM-related technologies at CES 2026, including an accelerator card prototype specialised for large language models and a Compute-using-DRAM product. The HBM4 standard itself is a stepping stone in this direction, introducing a logic base die manufactured using a logic process rather than a traditional DRAM process, enabling basic computational tasks to be performed on-die. Industry experts view this as a pivotal early step toward fuller PIM integration, with specialised AI processing units expected to be embedded directly into HBM logic dies by 2027.
Although these optimisations help reduce bottlenecks, they still come with trade-offs, such as potential accuracy losses or added complexity. And still, the memory bottleneck persists, as the appetite for higher bandwidth speeds and more capacity remains insatiable. While the memory suppliers are investing heavily to meet these demands, it takes time to respond, as bringing new state of the art fabrication plants (fabs) online is typically a five-year period when factoring in the time for construction, tool installation and yield ramp. As such, there does not appear to be any near-term solutions to fully solving the current bottlenecks.
SK Hynix case study
SK Hynix saw strong growth in FY2025, with overall revenue up 47% year-over-year (Figure 6) and HBM revenue more than doubling and making a significant contribution. There was also strong growth in conventional DRAM and NAND, driven to a large extent by growth in inference and the use of KV cache offloading. Operating profit for the year grew 101%, driven largely by price increases.
Figure 6: SK Hynix revenue and profit growth

Source: SK Hynix
Zooming in on the Q4 2025 results (Figure 7), there was particularly strong acceleration here, with revenue up 66% year-over-year and up 34% sequentially quarter-over-quarter. The Q4 sequential growth was driven largely by DRAM average selling price (ASP) increasing in the mid-20s%, and to a lesser extent by DRAM shipment growth (bit growth), which only grew in the low single digits. Shipment growth was driven by both HBM3E products and DDR5 for servers, where shipment of high-density DDR5 modules were up by roughly 50% quarter-over-quarter. NAND also saw strong revenue growth, with sequential ASP growth in the low 30s% and shipments up roughly 10%.
Figure 7: SK Hynix revenue by product and application

Source: SK Hynix
For FY26, SK Hynix has already secured full customer demand for its entire DRAM and NAND production and remains capacity constrained. DRAM bit shipments are guided to grow over 20% and NAND is guided to grow in the high teens%. Its new Cheongju M15X fab is currently on track to begin mass HBM production in H1 2026 and FY26 CAPEX is also set to increase considerably as it continues to invest in new fabs to expand production capacity. SK Hynix is also expected to maintain its leadership in HBM3E whilst simultaneously ramping up its next-generation GPU memory, HBM4, which started mass production in September 2025. HBM4 can process over 2 TB/s versus 1.2 TB/s for HBM3E and has a power efficiency improvement of more than 40%. Overall, analysts still expect HBM3E to make up two-thirds of total HBM shipments in 2026, with SK Hynix maintaining its market leading position. Analysts also expect SK Hynix to achieve roughly 70% market share for HBM4 in Nvidia’s next-generation Rubin platform.
Spillover effects in other sectors
“The AI-driven growth in the data centre has led to a surge in demand for memory and storage. Micron has made the difficult decision to exit the Crucial consumer business in order to improve supply and support for our larger, strategic customers in faster-growing segments.”
Sumit Sadana, Micron EVP and Chief Business Officer (December 2025)
“PCs and mobile devices are expected to see short-term shipment adjustments due to rising component costs and weakened consumer sentiment. Memory content per device is expected to grow at a slower pace due to price increases and supply constraints.”
Song Hyun Jong, SK Hynix President (Q4 2025)
The memory supply-demand imbalance is also impacting other sectors, including smartphones and PCs. Memory manufacturers are prioritising lucrative AI-grade DRAM over traditional DRAM, creating pricing pressure across consumer electronics. As a result, IDC estimates that the smartphone market will shrink 2.9% in terms of shipments year-over-year in 2026. It also expects that prices will have to rise significantly or specifications will have to be cut, or both. Furthermore, it expects the lower-end smartphones with thin margins to be impacted most severely. For high-end smartphones like Apple and Samsung there is also expected to be pressure, though they will likely be more insulated due to long-term supply contracts and market power. For the PC market, IDC expects even deeper disruption, with shipments forecasted to fall 4.9% (Figure 8). IDC also expects PC vendors with larger volumes to be better positioned and to take share away from the smaller, more vulnerable brands.
Figure 8: PC Market Forecast Scenarios

Source: IDC
Short to medium term outlook
In the short- to medium-term, market forecasts for the memory sector differ materially, largely reflecting the high uncertainty around pricing levels. Additionally, given how rapidly the space is evolving, forecasts are frequently being revised. In its most recent forecast, TrendForce revised its numbers upward significantly and now expects Q1 2026 conventional DRAM contract prices to rise 90-95% and blended HBM price to be up 80-85%, quarter-over-quarter (Figure 9).
Figure 9: Memory Price Growth Forecasts 4Q25-1Q26

Source: TrendForce
For the overall memory market (DRAM + NAND) in 2026, TrendForce forecasts 134% revenue growth, reaching a market size of US$551.6 billion (Figure 10). For 2027, it forecasts 53% growth, reaching a market size of US$842.7 billion.
Figure 10: DRAM and NAND Flash Revenue Projections

Source: TrendForce
On HBM bit shipments, SemiAnalysis forecasts that Nvidia will still dominate demand, while Broadcom, Amazon and AMD are also expected to grow their shares significantly (Figure 11).
Figure 11: HBM Bit (shipments) Demand Forecast

Source: SemiAnalysis
Longer term supply outlook
“I’d say my biggest concern actually is memory. The path to creating logic chips is more obvious than the path to having sufficient memory to support logic chips. That’s why you see DDR prices going ballistic.”
“They’re building fabs as fast as they can. So is Samsung. They’re pedal to the metal. They’re going balls to the wall, as fast as they can. It’s still not fast enough.”
Elon Musk, February 2026 (Dwarkesh Patel Podcast)
While memory market forecasts for 2028 and beyond are even more uncertain, if AI proves to be a secular trend that becomes integrated across all facets of society, this should sustain strong memory demand over the longer term. Demand would not only stem from data centres on Earth, but also potentially data centres in space, as well as self-driving cars and humanoids.
The key question then lies on the supply side, where figures like Elon Musk have voiced concerns that production capacity is likely to fail to keep pace with surging demand. Musk has urged both Samsung and Micron to build fabs faster and has said he would guarantee to purchase the output of those fabs. However, he does not think production will match his needs, which is why his proposed TeraFab ambitions extend beyond logic and packaging to include memory.
Nevertheless, TeraFab should not be viewed as a guaranteed solution to ease supply constraints any time soon. Memory manufacturing is highly specialised and requires deep domain expertise, process know-how and execution capability, making the industry extremely difficult to enter. While Musk has a strong track record of entering complex industries which makes success more likely than not, it is not assured. Eventual success would also take considerable time, not only because building, equipping and ramping a fab is a multi-year process, but also because a new entrant must assemble a world-class team, develop process expertise and work through a steep learning curve. Moreover, output may be used primarily to support Musk’s own companies rather than being supplied broadly to third parties.
New supply could also come from existing Chinese suppliers such as CXMT. However, they are generally considered to be several years behind the frontier and geopolitical considerations may make it difficult to participate in existing supply chains. Therefore, it currently looks unlikely that there will be a surprise supply shock in AI-grade memory. As such, if AI demand remains robust and the current build-out continues, supply constraints are likely to persist for many years to come.
Risks to memory demand
One of the larger risks to both the memory sector and the broader AI build-out is energy availability. AI data centres are extremely power-intensive and, given limited growth in power generation and grid capacity across the West, this could constrain deployment. Musk believes that already by the end of 2026 there will be an excess of chips, as there will not be enough available power to turn them all on. It is, however, unclear how power constraints might impact demand and pricing over time, though the effect may be more moderating than destructive, as newer chips are significantly more power efficient, resulting in continued demand as hyperscalers replace older inefficient chips.
In the long run, if orbital data centres become viable, which Musk now believes is three years away, the power constraint would be largely removed. That said, Musk’s timelines are often optimistic, and he himself has said that his timelines are typically set with only a 50% probability of being achieved. Other leaders in the AI space tend to view orbital data centres as something more like a decade away. There may also be other pathways to easing the energy constraints over time, including a faster build-out of nuclear capacity or meaningful breakthroughs in fusion, but these too are likely to take time to scale. As a result, power-related constraints could remain a long-term headwind to AI infrastructure growth and, by extension, the memory market.
Aside from energy, efficiency improvements on both the software and hardware side could moderate memory demand growth. As discussed earlier, architectural innovations such as MoE and distillation are reducing the memory required per unit of useful AI output. If these trends accelerate, for instance, if future models achieve frontier-level performance at a fraction of current parameter counts, the rate of growth in memory demand could slow materially, even if total demand continues to rise. Aggressive quantisation techniques, which have already moved well beyond FP16 to INT8 and INT4 in production inference, further compress memory requirements. On the hardware front, PIM could reduce the need for extreme memory bandwidth by performing certain computations within the memory itself. However, since PIM is being developed by the incumbent memory producers, most notably SK Hynix, it is more likely to enhance their competitive positioning than to disrupt their economics, at least in the medium term.
Finally, there could also be diminishing returns to scaling models, as well as challenges with further AI adoption and monetisation, which could subsequently reduce CAPEX. A scenario in which AI monetisation disappoints and hyperscalers pull back spending, coinciding with the arrival of new fab capacity currently under construction, could produce the kind of oversupply bust that has historically plagued the memory industry.
Conclusion
If the AI secular trend persists, driven by widespread adoption and positive ROI, the memory market could enter a prolonged era of structural tightness rather than repeating historical boom-and-bust cycles. In this scenario, demand would continue to be led by hyperscale data centres, both on Earth and potentially in space, as well as adjacent compute-intensive categories such as full self-driving vehicles and humanoid robotics. This continued demand expansion would keep supply struggling to catch up, compounded by the lengthy timelines for new fabs to become fully operational, thereby maintaining elevated prices and high margins for many years to come. That said, meaningful risks remain, such as energy constraints and potential AI monetisation challenges, which could reduce aggregate memory demand, leading to an oversupply bust. Barring such disruptions, however, the memory market may well be entering a golden age.
At AlphaTarget, we invest our capital in some of the most promising disruptive businesses at the forefront of secular trends; and utilise stage analysis and other technical tools to continuously monitor our holdings and manage our investment portfolio. AlphaTarget produces cutting edge research and our subscribers gain exclusive access to information such as the holdings in our investment portfolio, our in-depth fundamental and technical analysis of each company, our portfolio management moves and details of our proprietary systematic trend following hedging strategy to reduce portfolio drawdowns. To learn more about our research service, please visit https://alphatarget.com/subscriptions/.