ailia LLM NPU (QNN) Guide

How to run fast LLM / VLM inference on the NPU of Qualcomm SoCs using the QNN backend of ailia LLM.

Overview

ailia LLM 1.5.0 and later support fast inference on the NPU (Hexagon) of Qualcomm SoCs. For the NPU, ailia LLM uses a proprietary implementation developed by ailia Inc. that is separate from llama.cpp.

Using the NPU runs an LLM quickly and with low power consumption. In addition to text LLMs and VLMs (Vision Language Models) that take images as input, ailia LLM also supports audio input, with processing on the NPU. Gemma 4 E2B and E4B also run on mid-range devices such as the Snapdragon 7s.

In our measurements, prefill is up to 4.2x faster on Android devices, with CPU use during prefill kept below one thirtieth of the CPU path (Snapdragon 8+ Gen 1, 2,048-token input, compared with the CPU on the same device).

QNN (Qualcomm AI Engine Direct) is a backend API provided by Qualcomm. ailia LLM achieves NPU inference by converting GGUF, the model format of llama.cpp, into a format that can be executed by QNN.

Note: The QNN backend is available in ailia LLM 1.5.0 and later.

Comparison with Other Approaches

The table below compares ailia LLM with the main approaches for running LLMs on Qualcomm SoCs on Android. ailia LLM can use the NPU, runs Gemma 4 on the NPU, and supports relatively old SoCs with Hexagon v69 or later.

ApproachNPUGemma 4Supported SoCs (Android)Notes
llama.cppNoYesNo restriction (CPU execution)Runs on the CPU. An experimental Hexagon backend exists, but its NPU-side libraries (Skel) are unsigned, so it cannot run on regular devices. The image encoder runs on the CPU.
QNN (Qualcomm AI Engine Direct)YesNoHexagon v65 or later (Snapdragon 845 or later)Low-level inference API; no LLM runtime is provided.
QNN GenAI (Qualcomm Genie / GenieX)YesNoHexagon v79 or later (Snapdragon 8 Elite or later)The officially supported Android SoCs are limited to the Snapdragon 8 Elite family, and Hexagon v79 or later is required. No NPU (QAIRT) model is provided for Gemma 4.
LiteRT-LM NPUYesNoSnapdragon 8 Gen 2 / 8 Gen 3 / 8 Elite (SM8550 / SM8650 / SM8750)The only NPU model published for Qualcomm is Gemma3-1B (4-bit, 1,280 context). NPU models for Gemma 4 are provided only for Google Tensor and Intel. It is text only and does not support image or audio input.
ailia LLMYesYesHexagon v69 or later (Snapdragon 8 Gen 1 or later, 7s Gen 3, etc.)Uses the NPU even on older SoCs such as Snapdragon 8+ Gen 1, and runs Gemma 4 E2B / E4B on the NPU for text, VLM and ALM alike. Int16 quantization is supported, so it also runs on the Snapdragon 7s, which has no FP16 support.

Supported Platforms

Note: The supported QNN (QAIRT) version is 2.47.0.260601, the same as ailia SDK. QNN Models also depend on the QNN SDK version, so do not combine them with a different version of the QNN runtime.

Workflow

A GGUF file is converted into a QNN Model on a PC, and the converted QNN Model is executed by ailia LLM on the device. The GGUF file is not needed on the device; only the QNN Model is deployed.

GGUF
llama.cpp compatible model file (Q4 quantization)
↓
ailia LLM Compiler (PC)
Builds the graph with the QNN API and generates a per-SoC context binary
↓
QNN Model (.qnn)
SoC-specific, self-contained model file
↓
ailia LLM (Device)
Restores the context binary and runs prefill / decode on the Hexagon NPU

A QNN Model stores a fixed-length prefill graph (AR-N) and a decode graph (AR-1) in the same QNN context, sharing the weights and the KV cache between them. Weights are quantized to 4 bits (W4) at conversion time. In FP16 models, activations and the KV cache are processed in FP16.

ailia LLM Compiler supports both FP16 and Int16, which can be selected at conversion time according to the target SoC. The Snapdragon 7s Gen 3 NPU uses Hexagon v73 but does not support FP16, so ailia LLM uses an Int16-quantized model on this SoC. The Qualcomm Snapdragon 7s Gen 3 product brief (PDF, Artificial Intelligence section on page 2) lists INT4, INT8, and INT16 as supported NPU precisions; FP16 is not included. Select a model based on the precisions supported by the specific SoC, as well as its Hexagon version.

Supported Models

Gemma 4 E2B and Gemma 4 E4B (both supporting text, image, and audio input) are currently supported. For image and audio input, a QNN Model for the image / audio encoder (mmproj) is used in addition to the text QNN Model. The supported models depend on the Hexagon version of the SoC.

Hexagon versionExample SoCsSupported modelsContext length
v68 (Experimental)QCS6490Gemma 4 E2B8K
v69Snapdragon 8 Gen 1 / 8+ Gen 1Gemma 4 E2B8K
v73 or laterSnapdragon 7s Gen 3, 8 Gen 2 or laterGemma 4 E2B / E4B8K
Note: Hexagon v68 / v69 have a 2 GB size limit, so only E2B is supported there. Hexagon v73 and later have no such limit, so E4B is also supported. The pre-converted models all use a context length of 8K, which can be changed through the conversion settings of ailia LLM Compiler.

QNN Models are separate files for each model and each Snapdragon model (SoC). Pre-converted model files are provided, so please use the file that matches the SoC of your target device.

FileContents
gemma4-<model>-<soc>.qnn (e.g. gemma4-e2b-sm8475.qnn)Text model (prefill + decode)
gemma4-<model>-<soc>-mmproj.qnn (e.g. gemma4-e2b-sm8475-mmproj.qnn)Image / audio encoder (mmproj)

Model downloads

Pre-converted QNN Models for Gemma 4 can be downloaded below. To use VLM or audio input, download both the text model and the mmproj file. E4B requires Hexagon v73 or later.

SoCExample deviceModelFileContentsSize
sm8475
(Hexagon v69 / FP16)
Snapdragon 8+ Gen 1E2Bgemma4-e2b-sm8475.qnnText model (prefill + decode)3.60 GB
gemma4-e2b-sm8475-mmproj.qnnImage / audio encoder (mmproj)1.10 GB
sm7635
(Hexagon v73 / Int16)
Snapdragon 7s Gen 3E2Bgemma4-e2b-sm7635.qnnText model (prefill + decode)3.56 GB
gemma4-e2b-sm7635-mmproj.qnnImage / audio encoder (mmproj)0.63 GB
E4Bgemma4-e4b-sm7635.qnnText model (prefill + decode)5.54 GB
gemma4-e4b-sm7635-mmproj.qnnImage / audio encoder (mmproj)0.64 GB
qcs6490
(Hexagon v68 / Int16)
Experimental
QCS6490E2Bgemma4-e2b-qcs6490.qnnText model (prefill + decode)3.60 GB
gemma4-e2b-qcs6490-mmproj.qnnImage / audio encoder (mmproj)0.63 GB

All of these are for ailia LLM 1.5.0 (QAIRT 2.47.0.260601), with a context length of 8192 and the AR-256 / AR-1 configuration.

Experimental: The qcs6490 models have not been validated on a device for operation, accuracy or speed. Only structural checks and numerical comparison on a PC have been performed.
Other SoCs: Model files are currently published only for the two SoCs above (sm8475 and sm7635, plus the experimental qcs6490), but we can also convert models for other SoCs. Let us know the target SoC (the name returned by the ailiaLLMGetQNNModelName API) and we will compile a QNN Model for it with ailia LLM Compiler. Please contact ailia for details.
Note: A QNN Model depends not only on the Hexagon version but also on the QNN SDK version and the SoC model ID. Do not reuse a file built for a different SoC. The SoC name of the running device can be obtained with the ailiaLLMGetQNNModelName API (e.g. "sm8475").

Usage

Integrating into an Android app

In addition to libailia_llm.so, bundle libailia_llm_qnn.so in jniLibs/arm64-v8a of your application. The app loads only libailia_llm.so; libailia_llm_qnn.so is lazily loaded as a plug-in only when a QNN Model is opened.

You need to declare the use of libcdsprpc.so in AndroidManifest.xml. Also enable extractNativeLibs, because the NPU-side libraries (libQnnHtpV*Skel.so) must be reachable through a regular filesystem path.

<application android:extractNativeLibs="true" ...>
    <uses-native-library android:name="libcdsprpc.so" android:required="false"/>
</application>

Add the following to build.gradle so that the QNN runtime is automatically bundled into the apk.

implementation 'com.qualcomm.qti:qnn-runtime:2.47.0'
Note: QNN Models are specified by filesystem path, so they cannot be opened directly from apk assets. Copy them to the app's private storage or download them on first launch before opening.

Loading a QNN Model

A QNN Model (.qnn) is passed to the same model loading API as a GGUF file. When the context size is set to 0, the fixed context length embedded in the QNN Model is used. For image or audio input, open the mmproj QNN Model as the projector after opening the text model.

#include "ailia_llm.h"
struct AILIALLM *llm = nullptr;
ailiaLLMCreate(&llm);
ailiaLLMOpenModelFileA(llm, "gemma4-e2b-sm8475.qnn", /*ctx_size=*/0);
ailiaLLMOpenMultimodalProjectorFileA(llm, "gemma4-e2b-sm8475-mmproj.qnn"); // image / audio input only
// The prompt / generate APIs are the same as for GGUF
ailiaLLMDestroy(llm);

In Kotlin (JNI), pass the path of the QNN Model instead of the GGUF path in the same way.

val llm = AiliaLLM()
llm.openModelFile(qnnModelPath, 0)            // n_ctx=0 uses the fixed length in the QNN Model
llm.openMultimodalProjectorFile(mmprojQnnPath) // image / audio input only

Evaluation app

An evaluation apk for Android can be downloaded from the link below. The QNN-enabled build is ailia-models-kotlin-qnn.apk.

https://github.com/ailia-ai/ailia-models-kotlin/releases/tag/v260927

Architecture

ailia LLM maps the QNN Model with mmap, restores the QNN context, and executes it on the Hexagon NPU (HTP). Prefill is processed with the fixed-length AR-N graph; when the prompt length is not a multiple of AR, the last slice is padded. Decode uses the AR-1 graph in the same context, so no KV cache copy occurs after prefill. For VLM, image patching and position embedding are performed on the CPU, and the 16-layer Vision Transformer, pooling and projection are executed in a single QNN graph. Audio input uses an Audio graph to run the audio encoder on the NPU. The per-layer embedding (PLE) frontend also runs inside each transformer graph, leaving only token-row lookup and media embedding replacement on the CPU. This improves prefill for the Int16 model compared with running that input processing on the CPU.

QNN Model (.qnn)SoC-specific, self-contained model file
AR-N graph
Prefill (transformer + LM head)
+
AR-1 graph
Decode (transformer + LM head)
+
Shared static weights (W4)
Referenced by both AR-N and AR-1
+
Vision graph
Image encoder (mmproj.qnn)
+
Audio graph
Audio encoder (mmproj.qnn)
→
mmap / context restore
ailia LLM (Device)Runtime software stack
ailia LLM
Tokenization / prompt slicing / sampling
↓
ailia_llm_qnn
Context restore, KV cache management (plug-in)
↓
QNN (HTP backend)
Qualcomm AI Engine Direct
↓
libcdsprpc.so
FastRPC (RPC interface between CPU and NPU)
↓
Hexagon NPU
HVX (vector unit) / HMX (matrix unit)

Benchmark

The following results were measured with Gemma 4 on two devices. Snapdragon 8+ Gen 1 (SM8475 / Hexagon v69) uses the E2B FP16 model, and Snapdragon 7s Gen 3 (SM7635 / Hexagon v73) uses the E2B and E4B Int16 models. All were measured with QAIRT 2.47, a context length of 8192 and the AR-256 / AR-1 configuration.

The CPU numbers run the same Q4_0 GGUF through the standard llama.cpp path on the same device (8 decoder workers, context 8192). The shipping configuration is the build with OpenMP disabled, which is the "CPU" column. "CPU (OpenMP)" is a comparison build made with GGML_OPENMP=ON and OMP_NUM_THREADS=2.

Decode throughput

Throughput with 81 input tokens and 1,000 generated tokens.

DeviceModelCPUCPU (OpenMP)NPU (QNN)Effect of the NPU
SM8475 (Hexagon v69)E2B / FP169.971 tokens/s8.482 tokens/s8.062 tokens/s0.81x
SM7635 (Hexagon v73)E2B / Int165.805 tokens/s4.195 tokens/s6.410 tokens/s1.10x faster
SM7635 (Hexagon v73)E4B / Int163.351 tokens/s3.203 tokens/s3.128 tokens/s0.93x

With E2B on SM7635, NPU decode is 1.10x faster than the CPU (1.53x faster than the OpenMP build). On the other hand, E2B on SM8475 (0.81x) and E4B on SM7635 (0.93x) are slower than the CPU. Decode processes one token at a time, so it is dominated by weight loading and does not benefit much from the parallel compute of the NPU.

SM8475 E2B per-token decode throughput over 1000 generated tokens (CPU / CPU OpenMP / NPU)
SM8475 (E2B / FP16) per-token decode throughput over 1000 generated tokens
SM7635 E2B per-token decode throughput over 1000 generated tokens (CPU / CPU OpenMP / NPU)
SM7635 (E2B / Int16) per-token decode throughput over 1000 generated tokens
SM7635 E4B per-token decode throughput over 1000 generated tokens (CPU / CPU OpenMP / NPU)
SM7635 (E4B / Int16) per-token decode throughput over 1000 generated tokens

Prefill throughput

Throughput (tokens/s) for input lengths from 64 to 2,048 tokens. After a two-token warmup, SetPrompt is called twice without evaluation to clear the KV cache, and exactly one token is generated at each input length.

Device / model / path6412825651210242048
SM8475 E2B CPU109.35995.72288.24385.66379.94071.244
SM8475 E2B CPU (OpenMP)54.18455.33361.47660.31955.50051.064
SM8475 E2B NPU (FP16)76.244144.406277.775302.667301.791299.827
SM7635 E2B CPU91.55077.06480.49676.18769.75759.192
SM7635 E2B CPU (OpenMP)27.44149.99962.79370.96569.22856.059
SM7635 E2B NPU (Int16)33.60166.523130.329133.404131.367130.252
SM7635 E4B CPU1.69729.10528.41762.66364.38360.897
SM7635 E4B CPU (OpenMP)6.74627.82440.09761.87766.98054.372
SM7635 E4B NPU (Int16)20.30440.89080.70083.75883.31478.258
SM8475 E2B prefill throughput and host CPU use during prefill (CPU / CPU OpenMP / NPU)
SM8475 (E2B / FP16) prefill throughput (left) and host process CPU use during a 2,048-token prefill (right)
SM7635 E2B prefill throughput and host CPU use during prefill (CPU / CPU OpenMP / NPU)
SM7635 (E2B / Int16) prefill throughput (left) and host process CPU use during a 2,048-token prefill (right)
SM7635 E4B prefill throughput and host CPU use during prefill (CPU / CPU OpenMP / NPU)
SM7635 (E4B / Int16) prefill throughput (left) and host process CPU use during a 2,048-token prefill (right)

CPU use during prefill

Process CPU time measured around the first generate call of a 2,048-token prefill, expressed as a share of all 8 cores fully occupied (8 cores = 100%). This is the utilization of the application process, not of the whole device.

Device / modelCPUCPU (OpenMP)NPU (QNN)
SM8475 E2B98.7%86.3%3.0%
SM7635 E2B96.4%97.3%2.4%
SM7635 E4B97.5%96.7%2.6%

On the NPU, CPU use during prefill stays around 3%, leaving the CPU available for the rest of the application. Running on the CPU occupies nearly all 8 cores, which raises heat and interferes with other work.

VLM (image input) latency

Medians over the same 10 images at 768x768 with 100 output tokens each. TTFT is the prompt setup including the image encoder plus the first generate call, excluding model opening.

Device / model / pathTTFT (median)Decode (median)
SM8475 E2B CPU83.63 s12.081 tokens/s
SM8475 E2B CPU (OpenMP)128.22 s7.491 tokens/s
SM8475 E2B NPU (FP16)6.60 s7.033 tokens/s
SM7635 E2B CPU95.01 s5.891 tokens/s
SM7635 E2B CPU (OpenMP)97.45 s3.569 tokens/s
SM7635 E2B NPU (Int16)8.35 s6.444 tokens/s
SM7635 E4B CPU115.03 s3.332 tokens/s
SM7635 E4B CPU (OpenMP)96.90 s2.642 tokens/s
SM7635 E4B NPU (Int16)15.01 s3.098 tokens/s

Image TTFT is about 12.7x faster for E2B on SM8475, about 11.4x faster for E2B on SM7635 and about 7.7x faster for E4B. Both the CPU and the NPU consume the same 768x768 pixels and the same 256 image embeddings.

Audio input latency

Medians over 10 FLEURS recordings (2 each in 5 languages, generated to natural EOS with a 256-token cap). TTFT is the prompt setup including the audio encoder plus the first generate call, excluding model opening.

Device / model / pathTTFT (median)Decode (median)Output tokens
SM8475 E2B CPU37.29 s10.416 tokens/s21-89
SM8475 E2B CPU (OpenMP)36.49 s10.539 tokens/s21-89
SM8475 E2B NPU (FP16)4.06 s6.495 tokens/s21-90
SM7635 E2B CPU47.22 s5.166 tokens/s21-89
SM7635 E2B CPU (OpenMP)35.37 s3.846 tokens/s21-89
SM7635 E2B NPU (Int16)6.98 s6.356 tokens/s22-60
SM7635 E4B CPU89.91 s3.053 tokens/s21-91
SM7635 E4B CPU (OpenMP)45.99 s2.947 tokens/s21-91
SM7635 E4B NPU (Int16)12.37 s3.018 tokens/s21-91

Audio TTFT is about 9.2x faster for E2B on SM8475, about 6.8x faster for E2B on SM7635 and about 7.3x faster for E4B.

SM8475 E2B VLM and audio TTFT breakdown (CPU / CPU OpenMP / NPU)
SM8475 (E2B / FP16) VLM and audio TTFT breakdown (encoder, and decoder prefill plus the first token)
SM7635 E2B VLM and audio TTFT breakdown (CPU / CPU OpenMP / NPU)
SM7635 (E2B / Int16) VLM and audio TTFT breakdown (encoder, and decoder prefill plus the first token)
SM7635 E4B VLM and audio TTFT breakdown (CPU / CPU OpenMP / NPU)
SM7635 (E4B / Int16) VLM and audio TTFT breakdown (encoder, and decoder prefill plus the first token)

Memory usage

E2B: memory usage on SM7635 (Hexagon v73) while running the LLM, VLM and ALM configurations with 8 generated tokens (in MiB, context length 8192). The CPU path uses the Q4_0 GGUF with the BF16 mmproj, and the NPU path uses the Int16 QNN Model. Models are loaded with mmap, so anonymous memory (RssAnon, the working memory actually allocated) and file-backed mappings (RssFile, reusable page cache) are listed separately. Peak is the peak resident size of the process (VmHWM).

ModelWorkloadPathVmHWMRssAnonRssFile
E2BLLMCPU4574.21524.12896.5
E2BLLMNPU (Int16)1666.8386.41276.2
E2BVLMCPU4881.22512.440.3
E2BVLMNPU (Int16)2161.2673.31483.7
E2BALMCPU4873.72623.01509.1
E2BALMNPU (Int16)2202.5591.61606.7

E4B was measured on the same device as a single LLM run per path with 81 input tokens and 100 output tokens (in MiB). The RssAnon / RssFile split was not captured, so live PSS is listed instead. These conditions differ from the E2B table.

ModelWorkloadPathVmHWMLive PSS
E4BLLMCPU5569.84432.7
E4BLLMCPU (OpenMP)5518.23922.2
E4BLLMNPU (Int16)1976.01801.8

With E4B as well, the NPU peak is about a third of the CPU path. This measurement was taken with the device at Thermal Status 1 and is affected by Android page reclamation.

Limitations