Laya encoder for the Qualcomm QCS6490 (Hexagon V68) NPU โ w8a16
Quantized QNN/QAIRT DLC builds of the encoder from
convaiinnovations/laya (421M ModernBERT
typed-decision model), ready to run on a Qualcomm QCS6490 / Hexagon V68 HTP at w8a16
(int8 weights, int16 activations โ V68 has no FP16 path).
Recipe, benchmarks, and host-side code: https://github.com/nameissakthi25/laya-qcs6490
Files
| file | context | calibration | use |
|---|---|---|---|
laya_encoder_mm.dlc |
128 tokens | min-max | near-lossless at short context (NPUโfloat r 0.87, 100% decision agreement) |
laya_encoder_256_sqnr.dlc |
256 tokens | sqnr | long-context typed decisions; required โ min-max collapses at 256 (see below) |
mask_only_128.onnx / mask_only_256.onnx |
โ | โ | host helper: builds the gmask/smask attention-bias tensors from an attention_mask |
ctx/laya_encoder_mm_ctx.bin / ctx/laya_encoder_256_sqnr_ctx.bin |
128 / 256 | โ | pre-built HTP context binaries for QCS6490 + QAIRT 2.37.1 (skip the ~33 s finalize) |
These are the encoder only. The design is a host/NPU split: tokenize + embedding lookup + the
marker/scorer head + temperature calibration run in float on the host; the encoder runs on the
NPU. (V68 ignores runtime integer Gather indices on-chip, so the lookups must stay on the host.)
The DLC takes inputs_embeds (B,L,1024) + gmask + smask, and returns the final-LayerNorm hidden
states.
The key finding: calibration method at long context
Typed decisions read marker tokens whose meaning is a small RoPE positional signal. At 256
tokens, min-max calibration sets activation ranges from outliers, making the 16-bit step too coarse
to preserve that signal โ two adjacent markers quantize to the same value and every decision
collapses to 0.5 (AUROC 0.00). --act_quantizer_calibration sqnr keeps a fine enough step: the
markers separate (0.00 โ 23.4 on-board) and the encoder stays faithful to float (hidden-state Pearson
0.88). So: min-max โค128, sqnr 256+.
Run it
# On a QCS6490 + QAIRT 2.37.1 you can use the bundled context binary directly:
export ADSP_LIBRARY_PATH=/usr/lib/rfsa/adsp
qnn-net-run --backend /usr/lib/libQnnHtp.so \
--retrieve_context ctx/laya_encoder_256_sqnr_ctx.bin \
--input_list il.txt --output_dir out
# On any other V68 / QAIRT version, regenerate it from the DLC first (~33 s, then ~1 s/load):
qnn-context-binary-generator --backend /usr/lib/libQnnHtp.so --model /usr/lib/libQnnModelDlc.so \
--dlc_path laya_encoder_256_sqnr.dlc --binary_file laya_encoder_256_sqnr_ctx --output_dir ctx
Each il.txt line is one inference: inputs_embeds:=ie_0.raw gmask:=gm_0.raw smask:=sm_0.raw.
Batch many inputs per qnn-net-run (each --retrieve_context creates/destroys an HTP context; doing
it per-inference can exhaust the DSP).
Built with QAIRT / QNN 2.37.1.250807. The HTP context binaries are board- and QAIRT-version-specific; regenerate them from these DLCs with the command above.
License
Apache-2.0, inherited from the base model
convaiinnovations/laya.
Model tree for nameissakthi/laya-qcs6490-encoder
Base model
convaiinnovations/laya