Qualcomm is preparing significant neural processing unit (NPU) upgrades for its upcoming Snapdragon Elite mobile processors. On September 10, 2026, the company detailed the next generation of its Hexagon NPU, which is engineered specifically to run multimodal AI applications and agentic workflows directly on smartphones.
These NPU improvements are slated to follow the recently launched Snapdragon 8 Elite Gen 5 and the regular Snapdragon 8 Gen 5 chips. Qualcomm intends for these Elite-branded processors to define the next generation of Android flagships.
Last year, the company's Snapdragon 8 Gen 3 was heavily promoted to compete for AI dominance, boasting a 3.5 times bump in AI performance over the preceding Snapdragon 8 Gen 2. That older chip was pitched directly against Google's Tensor G3 as part of an ongoing battle for native generative AI capabilities in the Android ecosystem.
The current push for higher on-device AI performance also comes as Qualcomm faces direct competition from hardware rivals like MediaTek. MediaTek recently unveiled its Dimensity 9500 processor featuring an "All Big Core" architecture that notably breaks the 4GHz mark.
The new Element Accelerator
The upgraded Hexagon NPU introduces a specialized hardware component called the Element Accelerator. Sitting alongside the existing scalar, vector, and tensor units, this new accelerator is built to speed up transformer inference.
Qualcomm claims this design helps on-device AI agents respond and reason much faster without adding extra power costs. This builds on Qualcomm's massive existing footprint in AI data processing; the company claims it has already shipped more than 150 billion chips that handle AI on devices.
To accommodate these complex tasks, Qualcomm has implemented significantly larger shared memory across all four NPU accelerator units. This expanded memory reduces the processor's dependence on the phone's main RAM when exchanging data from the KV-cache. As a result, the NPU can hold more context internally, making room for heavier agentic workloads.
Local execution for 30B parameter MoE models
The revised Hexagon architecture is designed to support Mixture-of-Experts (MoE) models, allowing AI systems with up to 30 billion parameters to operate locally. This represents a significant leap from earlier architectures like the Snapdragon 8 Gen 3, which was promoted to run generative AI models up to 13 billion parameters on-device.
MoE models achieve this local processing by only activating a fraction of their total parameters. The architecture routes specific queries to specialized "expert" models based on the input, effectively reducing both compute demands and memory bandwidth needs.
For AI models utilizing INT4 precision, the chip maker states the new NPU provides 50% higher pre-fill performance, faster decoding throughput, enhanced speculative decoding, and a higher overall token-per-second rate. However, Qualcomm clarifies that these claims apply specifically to INT4 models and should not be taken as a direct 50% increase in overall AI performance.
Qualcomm notes that these specific enhancements will make multi-step agentic workflows far more responsive on upcoming Snapdragon Elite chips. Beyond mobile devices, Qualcomm is also pushing its Hexagon NPUs into personal computers. The next-generation Hexagon NPU in Snapdragon X2 Series PC processors is capable of up to 80 TOPS, delivering a 78% peak performance uplift over the first generation's 45 TOPS baseline.