Nvidia's Groq 3 LPU Benchmarks Show 4x Inference Speedup, but Questions Remain on Scalability

First independent tests of the $20B acquisition reveal impressive throughput on small models, but the architecture's reliance on SRAM limits large-model performance

edit
By LineZotpaper
Published
Read Time3 min
Nvidia has released the first independent benchmarks of its Groq 3 LPU-based LPX rack systems, showing a throughput of 3,400 tokens per second on Google's Gemma 4 31B model — four times faster than the nearest competitor, Cerebras. However, analysts caution that the results are a best-case scenario for a small model and that the architecture's limited SRAM capacity (500 MB per LPU) may hamper performance on larger, more complex models crucial for enterprise AI workloads.

Nvidia's $20 billion bet on Groq's LPU technology appears to be paying off, at least on small models. On Monday, the GPU giant released benchmarks from independent tester Artificial Analysis showing that its LPX rack systems can churn out 3,400 tokens per second (tok/s) on Google's Gemma 4 31B model with a 100,000-token input sequence. Nvidia claims this makes it 4x faster than the nearest alternative platform, which according to Artificial Analysis' leaderboard would be Cerebras, which achieved 882 tok/s under the same conditions.

Groq's LPUs, acquired by Nvidia in late December, use an SRAM-heavy dataflow architecture designed specifically for high-performance inference. Unlike traditional datacenter GPUs that rely on fast DRAM such as GDDR7 or HBM4, Groq's chips depend entirely on on-die SRAM, which offers bandwidth of around 2.75 TB/s per chip. The third-generation Groq 3 LPU, launched as part of Nvidia's Vera Rubin platform earlier this year, boasts 150 TB/s of memory bandwidth.

However, the trade-off is stark: SRAM consumes significant die area, limiting each Groq 3 LPU to just 500 MB of memory — 576 times less than Nvidia's top-spec Rubin GPU (288 GB). To run a 31B-parameter model, Nvidia's architecture distributes it across multiple LPUs using Ethernet, with each LPX rack housing up to 256 LPUs for 128 GB of high-bandwidth SRAM. For larger models, multiple racks can be ganged together.

Nvidia argues that the extreme speed is essential for AI agents and code assistants, where faster token generation enables longer reasoning, more turns, and greater information processing in the same time window. The company has already announced that Netherlands-based neocloud Nebius will be among the first to deploy the combined GPU and LPU systems.

Critics, however, point out that the Gemma 4 31B model is a best-case scenario: it fits neatly into a single LPX rack at FP8 precision (requiring about 64 LPUs). The model is small enough to run on a high-end consumer GPU like an RTX 3090/4090 at 4-bit precision. The architecture's performance on larger, more complex MoE models remains untested, and Nvidia has not disclosed how it distributes the model across chips — likely pipeline parallelism with possible data parallelism for higher concurrency.

§

Analysis

Why This Matters

  • Nvidia's $20B acquisition of Groq represents a bet that ultra-fast inference will be a premium feature in the age of AI agents — if the benchmarks hold up, it could reshape the inference hardware market.
  • The 4x speedup over Cerebras demonstrates that SRAM-based architectures can deliver dramatic latency improvements, but the small model size tested raises questions about real-world applicability for large-scale enterprise deployments.
  • Nvidia's integration of LPUs into its datacenter ecosystem (alongside GPUs) could create a hybrid offering that combines high throughput for small models with high memory for large ones, potentially pressuring competitors like AMD, Intel, and Cerebras.

Background

Nvidia's acquisition of Groq in late December 2024 stunned the industry: the GPU giant paid $20 billion for a startup that had never shipped a product at scale. Groq's LPU (Language Processing Unit) uses a dataflow architecture that eliminates the need for external DRAM by relying entirely on on-die SRAM, which is orders of magnitude faster but far more expensive and limited in capacity. The third-generation Groq 3 LPU was launched earlier this year as part of Nvidia's Vera Rubin platform. Independent benchmarks have been scarce, making Monday's Artificial Analysis results the first credible public measure of the technology's performance. The test used Google's Gemma 4 31B, a relatively small open-weight model, running at FP8 precision.

Key Perspectives

Nvidia: The company touts the 3,400 tok/s result as validation of its $20B investment. An Nvidia spokesperson stated that the LPX system is "4x faster than the nearest alternative platform," positioning it as essential for agentic AI workloads where latency is critical. Nvidia points to customer interest from Nebius as evidence of market demand.

Cerebras (competitor): The results place Cerebras in second place at 882 tok/s, but the company may argue that its architecture can handle larger models more efficiently. Cerebras' wafer-scale engine also uses SRAM, but with a different design philosophy focused on scaling to massive models.

Critics/Skeptics: Analysts note that the Gemma 4 31B benchmark is a best-case scenario. The model is small enough to fit in a single LPX rack, and real-world deployments often involve much larger models (e.g., 70B-400B parameters) that would require multiple racks and complex parallelism. The architecture's reliance on Ethernet for inter-chip communication could introduce latency penalties at scale. Additionally, the cost per chip is high due to SRAM's die area consumption, and power efficiency comparisons have not been disclosed.

What to Watch

  • Independent benchmarks on larger models, particularly 70B+ MoE models like Llama 3.1 405B or Mixtral 8x22B, to assess scalability.
  • Nvidia's disclosure of pricing and power consumption for the LPX racks — these will determine whether the speed premium is economically viable.
  • Adoption by major cloud providers: Nebius is a niche neocloud; watch for AWS, Azure, or Google Cloud to deploy LPU-based instances.
  • The impact on Cerebras' IPO plans: the company filed confidentially earlier this year, and Nvidia's results could either validate or challenge its market position.

Sources

newspaper

Zotpaper

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.