Qualcomm’s AI Initiative: A Risky Gamble or a Strategic Leap?
Qualcomm is making a serious commitment to AI infrastructure, with its venture into the datacenter dependent on a bold near-memory compute architecture aimed at providing superior inference economics compared to current GPUs. Revealed during its 2026 investor day last week, this technology will involve Qualcomm layering multiple DRAM units atop its XPUs to create a cohesive compute and memory module dubbed high-bandwidth compute (HBC). “We deliver all the performance benefits of SRAM while achieving the density and memory capacity that HBM (high-bandwidth memory) stacks provide,” stated Tony Pialis, Qualcomm’s EVP of datacenter, in last week’s investor briefing.
High-Bandwidth Compute: Qualcomm’s Disruptor?
This innovation is expected to debut next year as part of Qualcomm’s AI250-series Dragonfly rack systems and signifies a notable alteration in Qualcomm’s AI infrastructure approach. The smartphone giant is well-acquainted with AI accelerators, with virtually every Snapdragon processor currently available integrating an NPU. However, in the datacenter realm, the firm has faced challenges in generating the same enthusiasm as Nvidia, AMD, and emerging players like Cerebras. Relative to the GPUs of the top two competitors, Qualcomm’s AI-series accelerators have not fared as well, but this may soon change as the company strives to establish itself in the datacenter domain.
Qualcomm’s Assertions: Unbelievably Impressive?
With the AI250, the system-on-chip (SoC) manufacturer is asserting 768 GB of memory capacity and up to 133 TB/s of effective memory bandwidth per card. For comparison, Nvidia’s Groq 3 LPUs provide merely 500 MB of SRAM and 150 TB/s of bandwidth. If that seems overly ambitious, it’s because it likely is. Qualcomm is putting heavy emphasis on the term “effective.” This is evident as for the AI200-based Dragonfly systems launching this year, they claimed 414 TB/s of “effective” memory bandwidth across 56 chips. At first glance, this seems plausible, but achieving this with just 8800 MT/s LPDDR5x would necessitate a 6,720-bit-wide bus, which is almost certain not to be present.
Exploring High-Bandwidth Compute
Boosting “effective” bandwidth is not the only feature of these HBC-based accelerators. Qualcomm contends that by situating some of the XPU’s compute beneath the DRAM, they can greatly minimize the power consumption of their chips. In traditional datacenter GPUs, data is swiftly exchanged between HBM and the compute dies. Even with advanced packaging techniques like TSMC’s CoWoS, the energy expended to facilitate this data movement is considerable.
The Advantages of Stacked DRAM
By stacking the DRAM directly atop certain logic components and linking them via through-silicon vias (TSVs), the distance between compute and memory is significantly reduced. “Think of it as living in the same building where you work, so you only need to travel a few floors,” Pialis remarked. “What does this imply for the roads linking suburban areas to the city? The roads are clear. The benefits to the industry include reduced power usage, lower heat, and the expensive silicon interposer road typically needed by HBM solutions is no longer necessary.”
Qualcomm’s Rivals: Who Will Emerge Victorious?
Does HBC truly offer a competitive edge? While Qualcomm is an early voice among chip manufacturers to emphasize near-memory or HBC, it is neither the first nor is the technology out of reach for Nvidia or AMD. In fact, rumors suggest that both Nvidia and AMD are collaborating with HBM suppliers and TSMC to develop customized base dies to enhance the performance of their upcoming chips, although it’s unclear how much compute has been integrated into these designs. Qualcomm informs us that its HBC “incorporates LPDDR memory into a tailor-made near-memory computing architecture that amalgamates compute with highly-accelerated memory bandwidth in a 3D-stacked silicon design.”
Qualcomm Recaptures its Essence
Alongside teasing its forthcoming AI250 and AI300 accelerators, Qualcomm’s investor day was also marked by the acquisition of AI software startup Modular. Modular was co-founded by Tim Davis and Chris Lattner, the latter being recognized as the creator of LLVM, Clang, the Swift programming language, and the multi-level intermediate representation (MLIR) compiler framework. At Modular, Lattner and his team developed Mojo, a low-level programming interface for GPUs, providing a high-performance alternative to Nvidia’s CUDA and AMD’s HIP and ROCm frameworks. The core idea is that users should be able to create highly efficient AI applications that will operate irrespective of the hardware in use. For Qualcomm, Mojo represents an opportunity to bypass the CUDA barrier that has constrained AMD for an extended period.
A Future Beyond Hardware
With Mojo, Qualcomm’s clients will not have to commit to a single platform; they can design their applications and execute them on whatever computing resources are available at the time. It is not an all-or-nothing situation either. Modular is expected to enable support for heterogeneous deployments akin to Nvidia’s Groq LPU technology, where GPUs could be employed for prefill tasks, while AI250s handle decoding in any ratio that best suits the specific application.
Conclusion: Qualcomm’s Strategy – Deciphering the Hype
In any case, we won’t have to wait much longer to witness the HBC in operation. Following the rollout of its AI200-series racks later this year, Qualcomm plans to launch its inaugural HBC-based AI250 starting in 2027, with its second-generation HBC platform scheduled for 2028. While you await this, consider reading more about Qualcomm’s new datacenter CPU, which we examined in greater detail last week.