Skip to content
AI.info

The Pulse

Qualcomm Puts 30B-Parameter AI Models on Its New Phone Chips

Qualcomm has introduced the Snapdragon 8 Elite Extreme Gen 6 and Snapdragon 8 Elite Gen 6 smartphone platforms. The Extreme model is designed to run 30-billion-parameter mixture-of-experts models locally, while both chips add new hardware f

Qualcomm Puts 30B-Parameter AI Models on Its New Phone Chips

AI.info Team ·

Qualcomm’s 30B model target moves into the phone

A smartphone processor designed to run a 30-billion-parameter AI model locally is now part of Qualcomm’s flagship mobile lineup. The company announced the Snapdragon 8 Elite Extreme Gen 6 and Snapdragon 8 Elite Gen 6 on September 22 at Snapdragon Summit in Maui, positioning both chips around on-device AI agents rather than conventional voice assistants.

Qualcomm says the Extreme version can run a 30-billion-parameter mixture-of-experts model on the phone. Such models contain many available parameters but activate only a subset for each token, reducing the compute and memory demands of each response. Qualcomm’s technical explanation says a 30B model can activate about 3B routed parameters at a time through the new Hexagon neural processing unit.

The announcement does not mean every 30B model will run on every phone using the chip. The capability depends on model architecture, quantization, memory capacity and software support. Qualcomm’s claim applies to the Snapdragon 8 Elite Extreme Gen 6 and its support for mixture-of-experts workloads.

Two chips, one push toward local agents

The new processors form a two-tier flagship strategy. Qualcomm describes the Snapdragon 8 Elite Extreme Gen 6 as its most powerful mobile platform, while the standard Snapdragon 8 Elite Gen 6 brings many of the same AI, camera, gaming and connectivity features to a broader set of premium devices.

“Today marks a milestone as we introduce not one, but two new flagship mobile platforms, delivering agentic AI experiences that are immediate, personal, and adaptive,” Chris Patrick, senior vice president and general manager of mobile handsets at Qualcomm Technologies, said in the company’s launch announcement.

Both platforms use Qualcomm’s next-generation Hexagon NPU and a new Element Accelerator aimed at transformer workloads. Qualcomm says the NPU has a 50% larger memory subsystem, allowing more model data, activations and intermediate results to stay close to the processor instead of moving repeatedly through the phone’s main memory.

The Hexagon NPU is built around selective computation

Qualcomm outlined the architecture in a September 10 post by Vinesh Sukumar, the company’s vice president of product management for AI and generative AI. The post says the NPU supports INT2, INT4, INT8, FP8 and FP16 precision, giving developers different ways to trade model quality, speed, memory use and power consumption.

Qualcomm also claims up to 50% faster prefill performance for INT4 models. Prefill is the stage in which a model processes the user’s existing prompt or conversation before generating new tokens. Faster prefill can reduce the delay before an agent begins responding, although the effect on a finished product will depend on the model and software stack used by each phone maker.

The company’s Hexagon NPU explanation frames mixture-of-experts models as a way to make larger AI systems fit within mobile power and memory limits. Instead of running every expert for every request, the system routes each token to the specialists it needs. That makes the 30B headline number different from the amount of computation performed during any one step.

Small models handle context close to the sensors

The chips also add a sensing hub designed for lower-power, always-available tasks. Qualcomm says it can run models of up to 200 million parameters locally, support a personal scribe and distinguish between speakers. The company says the hub can also build a usage-based memory to improve suggestions for automating tasks.

Qualcomm says the processor can run a complete voice-in, voice-out agent on the device. Keeping parts of that interaction local could reduce dependence on a network connection and limit the amount of personal audio or context sent to a cloud service, although Qualcomm has not described a single standard software experience that will appear across all phones.

The company is presenting those capabilities as personal agents that can understand context, maintain information about a user and act with permission. The practical test will be whether manufacturers expose the functions in everyday software, rather than limiting them to demonstrations or proprietary applications.

Camera and gaming hardware share the billing

AI is not the only focus of the new platforms. Qualcomm says the Extreme model supports 8K video at 60 frames per second and 4K video at 240 frames per second, along with its Advanced Professional Video codec. Intelligent Pixel Control is intended to give the camera pipeline more detailed control over individual parts of a scene.

The company is also adding Adreno Neural Fusion, which uses dedicated AI processing in the GPU for functions such as super resolution, frame generation and image reconstruction. Qualcomm says the Extreme model’s Oryon CPU reaches 5 GHz, while the redesigned GPU delivers a 44% performance increase and a 40% improvement in power efficiency compared with the previous generation.

Connectivity receives a major upgrade as well. Qualcomm lists peak download speeds of 14.8 gigabits per second through the X105 5G modem-RF system and up to 11.6 gigabits per second through the FastConnect 8800 wireless system on the Extreme platform.

Motorola names the first announced phone

Qualcomm says devices using the two platforms will come from HONOR, iQOO, Motorola, OnePlus, OPPO, REDMI, RedMagic, vivo and Xiaomi. Motorola separately announced the Motorola Signature 27, powered by the Snapdragon 8 Elite Extreme Gen 6, with general availability planned for later in 2026.

The launch gives Qualcomm a clear technical argument for local AI: a phone can host a model with tens of billions of available parameters while activating only a smaller expert set for each request. Whether that produces useful, battery-friendly agents will depend on model availability, phone memory, thermal design and the software choices made by each manufacturer. For now, the concrete promise is defined by the chip: a 30B mixture-of-experts model is intended to run on the handset, not in a remote data center.

Source

Qualcomm

Explore

More articles