Inference has become the fastest-growing part of the artificial intelligence (AI) infrastructure market, and Bloomberg Intelligence projects it will double the size of the AI training market by 2032, reaching $1.3 trillion. With so much at stake, both leading chipmakers and upstarts are jockeying to grab a slice of this huge, fast-growing market.
Nvidia (NVDA -0.02%), Cerebras (CBRS +2.06%), and Advanced Micro Devices (AMD +1.01%) are all tackling this market in different ways. Let's see how these AI stocks stack up and why they could all be winners, given the size and growth of the inference market.
Image source: Getty Images.
1. Nvidia
Already the winner in AI model training, Nvidia now has its sights on the inference market. The company's big move to capture share was its "acquisition" of Groq and its language processing units (LPUs). Inference is more about fast memory access and low latency than raw compute power, and LPUs help address this by having SRAM (static random-access memory) embedded directly on their chips.

NASDAQ: NVDA
Key Data Points
LPUs are particularly useful during the decode phase of inference, which is when large language models (LLMs) answer queries. As such, Nvidia now offers complete systems designed specifically for inference, where its graphics processing units (GPUs) handle the pre-fill phase (reading the prompt) while its LPUs handle the decode phase, thereby speeding up response times.
This is a nice solution and positions Nvidia to remain an AI infrastructure leader, even if it doesn't capture the same market share it does in training.
2. Cerebras
Like Nvidia, Cerebras is tackling inference with SRAM-based chips. However, because SRAM is so bulky, instead of just embedding a small amount onto its chips and stringing them together, Cerebras has created huge wafer-sized chips that are five to six times faster than LPUs.
The physical size of Cerebras' chips comes with some trade-offs. They require specialized cooling and energy management solutions and, as such, are only sold or rented as part of the Cerebras CS-3 systems. They also come at a very premium price tag.

NASDAQ: CBRS
Key Data Points
However, the company has inked major deals with OpenAI and Amazon's AWS, and it recently announced a partnership with AMD that should help reduce the cost of ownership. The two companies will offer an inference solution in which AMD's Helios rack-scale solution will handle the pre-fill phase of inference, which it can do more cheaply, while Cerebras' Wafer-Scale Engine will perform the decode phase, which it can do more quickly. It's a nice way for companies to better compete with Nvidia's offerings.
Given the high cost of its systems, Cerebras has been more of a premium, niche solution, but it now looks set to become a major player in the humongous inference market.
3. AMD
After losing out on the LLM training market to Nvidia, AMD has been aggressively pursuing the inference market to make sure it doesn't get left behind again. Its chiplet design is better suited for inference, as it allows its GPUs to be packaged with more high-bandwidth memory (HBM) and to act as part of an entire unit to reduce latency. Meanwhile, its partnership with Cerebras looks like a smart move to help it better compete with Nvidia's complete inference system.
However, the company has not stopped there. It recently acquired memory optimization company MEXT and chip start-up Taalas to boost its inference offering. Memory is one of the biggest AI bottlenecks right now, and MEXT's solution can offload seldom-accessed data from DRAM to unused flash and then, using predictive AI, can transfer it back into DRAM before an application even requests it. This can reduce the need for more expensive DRAM and help save costs.

NASDAQ: AMD
Key Data Points
Meanwhile, Taalas has developed chips in which AI models are hardwired directly to bolster inference performance. Since the chips are model-specific, they aren't as flexible, but they are cheaper and much faster. The company plans to use them as part of a complete system where its GPUs would handle the pre-fill phase and Taalas' chips would handle the decode phase.
AMD is tackling inference from a couple of different angles, which should position it to capture a nice share of this huge market.





