Laya encoder for the Qualcomm QCS6490 (Hexagon V68) NPU โ€” w8a16

Quantized QNN/QAIRT DLC builds of the encoder from convaiinnovations/laya (421M ModernBERT typed-decision model), ready to run on a Qualcomm QCS6490 / Hexagon V68 HTP at w8a16 (int8 weights, int16 activations โ€” V68 has no FP16 path).

Recipe, benchmarks, and host-side code: https://github.com/nameissakthi25/laya-qcs6490

Files

file context calibration use
laya_encoder_mm.dlc 128 tokens min-max near-lossless at short context (NPUโ†”float r 0.87, 100% decision agreement)
laya_encoder_256_sqnr.dlc 256 tokens sqnr long-context typed decisions; required โ€” min-max collapses at 256 (see below)
mask_only_128.onnx / mask_only_256.onnx โ€” โ€” host helper: builds the gmask/smask attention-bias tensors from an attention_mask
ctx/laya_encoder_mm_ctx.bin / ctx/laya_encoder_256_sqnr_ctx.bin 128 / 256 โ€” pre-built HTP context binaries for QCS6490 + QAIRT 2.37.1 (skip the ~33 s finalize)

These are the encoder only. The design is a host/NPU split: tokenize + embedding lookup + the marker/scorer head + temperature calibration run in float on the host; the encoder runs on the NPU. (V68 ignores runtime integer Gather indices on-chip, so the lookups must stay on the host.) The DLC takes inputs_embeds (B,L,1024) + gmask + smask, and returns the final-LayerNorm hidden states.

The key finding: calibration method at long context

Typed decisions read marker tokens whose meaning is a small RoPE positional signal. At 256 tokens, min-max calibration sets activation ranges from outliers, making the 16-bit step too coarse to preserve that signal โ€” two adjacent markers quantize to the same value and every decision collapses to 0.5 (AUROC 0.00). --act_quantizer_calibration sqnr keeps a fine enough step: the markers separate (0.00 โ†’ 23.4 on-board) and the encoder stays faithful to float (hidden-state Pearson 0.88). So: min-max โ‰ค128, sqnr 256+.

Run it

# On a QCS6490 + QAIRT 2.37.1 you can use the bundled context binary directly:
export ADSP_LIBRARY_PATH=/usr/lib/rfsa/adsp
qnn-net-run --backend /usr/lib/libQnnHtp.so \
  --retrieve_context ctx/laya_encoder_256_sqnr_ctx.bin \
  --input_list il.txt --output_dir out

# On any other V68 / QAIRT version, regenerate it from the DLC first (~33 s, then ~1 s/load):
qnn-context-binary-generator --backend /usr/lib/libQnnHtp.so --model /usr/lib/libQnnModelDlc.so \
  --dlc_path laya_encoder_256_sqnr.dlc --binary_file laya_encoder_256_sqnr_ctx --output_dir ctx

Each il.txt line is one inference: inputs_embeds:=ie_0.raw gmask:=gm_0.raw smask:=sm_0.raw. Batch many inputs per qnn-net-run (each --retrieve_context creates/destroys an HTP context; doing it per-inference can exhaust the DSP).

Built with QAIRT / QNN 2.37.1.250807. The HTP context binaries are board- and QAIRT-version-specific; regenerate them from these DLCs with the command above.

License

Apache-2.0, inherited from the base model convaiinnovations/laya.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for nameissakthi/laya-qcs6490-encoder

Quantized
(59)
this model